Результаты исследований: Публикации в книгах, отчётах, сборниках, трудах конференций › статья в сборнике › Рецензирование
Semantic-Episodic Robustness of Large Language Models: Formalization and Evaluation Dataset Generation. / Грудинин, Михаил Артемович; Вейбер, Евгения Николаевна; Абрамов, Максим Викторович.
Proceedings of 2026 29th International Conference on Soft Computing and Measurements, SCM 2026. Saint Petersburg : Institute of Electrical and Electronics Engineers Inc., 2026. стр. 166-169.Результаты исследований: Публикации в книгах, отчётах, сборниках, трудах конференций › статья в сборнике › Рецензирование
}
TY - CHAP
T1 - Semantic-Episodic Robustness of Large Language Models: Formalization and Evaluation Dataset Generation
AU - Грудинин, Михаил Артемович
AU - Вейбер, Евгения Николаевна
AU - Абрамов, Максим Викторович
N1 - Conference code: 29
PY - 2026/5/21
Y1 - 2026/5/21
N2 - Although large language models (LLMs) are capable of solving tasks that involve general reasoning patterns and factual knowledge, the reliability of their outputs remains highly dependent on the input text they receive. Increasing the length of the input introduces additional difficulty, as relevant facts may be distributed across different turns of a dialogue, making their integration more challenging and potentially reducing overall performance. This work introduces a method for constructing a dataset intended to assess semantic–episodic robustness in conversational settings. Central to the approach is the notion of a semantic–episodic task, defined as an evaluation unit in which a correct answer can only be produced by jointly leveraging episodic information provided exclusively within the dialogue and semantic knowledge that is not explicitly stated in the text but instead resides within the model itself. To support this, we develop a multi-stage pipeline for synthetic data generation that incorporates automated validation, enforces strict typing of outputs, and applies a graded difficulty scheme reflecting both the growth of input sequences and the increasing structural complexity of dialogues. The resulting benchmark is scalable, programmatically verifiable, and suitable for systematic comparison of models, prompting approaches, and memory mechanisms.
AB - Although large language models (LLMs) are capable of solving tasks that involve general reasoning patterns and factual knowledge, the reliability of their outputs remains highly dependent on the input text they receive. Increasing the length of the input introduces additional difficulty, as relevant facts may be distributed across different turns of a dialogue, making their integration more challenging and potentially reducing overall performance. This work introduces a method for constructing a dataset intended to assess semantic–episodic robustness in conversational settings. Central to the approach is the notion of a semantic–episodic task, defined as an evaluation unit in which a correct answer can only be produced by jointly leveraging episodic information provided exclusively within the dialogue and semantic knowledge that is not explicitly stated in the text but instead resides within the model itself. To support this, we develop a multi-stage pipeline for synthetic data generation that incorporates automated validation, enforces strict typing of outputs, and applies a graded difficulty scheme reflecting both the growth of input sequences and the increasing structural complexity of dialogues. The resulting benchmark is scalable, programmatically verifiable, and suitable for systematic comparison of models, prompting approaches, and memory mechanisms.
KW - large language models
KW - input sequence
KW - semantic knowledge
KW - episodic memory
KW - dataset generation
KW - benchmark
KW - long context
KW - strict answer formats
KW - automatic validation
KW - large language models
KW - input sequence
KW - semantic knowledge
KW - episodic memory
KW - dataset generation
KW - benchmark
KW - long context
KW - strict answer formats
KW - automatic validation
UR - https://ieeexplore.ieee.org/document/11625881
U2 - 10.1109/SCM71573.2026.11625881
DO - 10.1109/SCM71573.2026.11625881
M3 - Article in an anthology
SP - 166
EP - 169
BT - Proceedings of 2026 29th International Conference on Soft Computing and Measurements, SCM 2026
PB - Institute of Electrical and Electronics Engineers Inc.
CY - Saint Petersburg
Y2 - 20 May 2026 through 22 May 2026
ER -
ID: 160336398