Standard

Semantic-Episodic Robustness of Large Language Models: Formalization and Evaluation Dataset Generation. / Грудинин, Михаил Артемович; Вейбер, Евгения Николаевна; Абрамов, Максим Викторович.

Proceedings of 2026 29th International Conference on Soft Computing and Measurements, SCM 2026. Saint Petersburg : Institute of Electrical and Electronics Engineers Inc., 2026. стр. 166-169.

Результаты исследований: Публикации в книгах, отчётах, сборниках, трудах конференцийстатья в сборникеРецензирование

Harvard

Грудинин, МА, Вейбер, ЕН & Абрамов, МВ 2026, Semantic-Episodic Robustness of Large Language Models: Formalization and Evaluation Dataset Generation. в Proceedings of 2026 29th International Conference on Soft Computing and Measurements, SCM 2026. Institute of Electrical and Electronics Engineers Inc., Saint Petersburg, стр. 166-169, XXIX Международная конференция по мягким вычислениям и измерениям (SCM-2026), Санкт-Петербург, Российская Федерация, 20/05/26. https://doi.org/10.1109/SCM71573.2026.11625881

APA

Грудинин, М. А., Вейбер, Е. Н., & Абрамов, М. В. (2026). Semantic-Episodic Robustness of Large Language Models: Formalization and Evaluation Dataset Generation. в Proceedings of 2026 29th International Conference on Soft Computing and Measurements, SCM 2026 (стр. 166-169). Institute of Electrical and Electronics Engineers Inc.. https://doi.org/10.1109/SCM71573.2026.11625881

Vancouver

Грудинин МА, Вейбер ЕН, Абрамов МВ. Semantic-Episodic Robustness of Large Language Models: Formalization and Evaluation Dataset Generation. в Proceedings of 2026 29th International Conference on Soft Computing and Measurements, SCM 2026. Saint Petersburg: Institute of Electrical and Electronics Engineers Inc. 2026. стр. 166-169 https://doi.org/10.1109/SCM71573.2026.11625881

Author

Грудинин, Михаил Артемович ; Вейбер, Евгения Николаевна ; Абрамов, Максим Викторович. / Semantic-Episodic Robustness of Large Language Models: Formalization and Evaluation Dataset Generation. Proceedings of 2026 29th International Conference on Soft Computing and Measurements, SCM 2026. Saint Petersburg : Institute of Electrical and Electronics Engineers Inc., 2026. стр. 166-169

BibTeX

@inbook{7079bca64e764db59f2d722b1a4ad4a1,
title = "Semantic-Episodic Robustness of Large Language Models: Formalization and Evaluation Dataset Generation",
abstract = "Although large language models (LLMs) are capable of solving tasks that involve general reasoning patterns and factual knowledge, the reliability of their outputs remains highly dependent on the input text they receive. Increasing the length of the input introduces additional difficulty, as relevant facts may be distributed across different turns of a dialogue, making their integration more challenging and potentially reducing overall performance. This work introduces a method for constructing a dataset intended to assess semantic–episodic robustness in conversational settings. Central to the approach is the notion of a semantic–episodic task, defined as an evaluation unit in which a correct answer can only be produced by jointly leveraging episodic information provided exclusively within the dialogue and semantic knowledge that is not explicitly stated in the text but instead resides within the model itself. To support this, we develop a multi-stage pipeline for synthetic data generation that incorporates automated validation, enforces strict typing of outputs, and applies a graded difficulty scheme reflecting both the growth of input sequences and the increasing structural complexity of dialogues. The resulting benchmark is scalable, programmatically verifiable, and suitable for systematic comparison of models, prompting approaches, and memory mechanisms.",
keywords = "large language models, input sequence, semantic knowledge, episodic memory, dataset generation, benchmark, long context, strict answer formats, automatic validation, large language models, input sequence, semantic knowledge, episodic memory, dataset generation, benchmark, long context, strict answer formats, automatic validation",
author = "Грудинин, {Михаил Артемович} and Вейбер, {Евгения Николаевна} and Абрамов, {Максим Викторович}",
note = "Grudinin M. A., Vejber Y. N., Abramov M. V. Semantic-Episodic Robustness of Large Language Models: Formalization and Evaluation Dataset Generation // Proceedings of 2026 29th International Conference on Soft Computing and Measurements, SCM 2026. IEEE, 2026. P. 166–169. DOI: 10.1109/SCM71573.2026.11625881.; null ; Conference date: 20-05-2026 Through 22-05-2026",
year = "2026",
month = may,
day = "21",
doi = "10.1109/SCM71573.2026.11625881",
language = "English",
pages = "166--169",
booktitle = "Proceedings of 2026 29th International Conference on Soft Computing and Measurements, SCM 2026",
publisher = "Institute of Electrical and Electronics Engineers Inc.",
address = "United States",
url = "https://scm.etu.ru/2026/ru/",

}

RIS

TY - CHAP

T1 - Semantic-Episodic Robustness of Large Language Models: Formalization and Evaluation Dataset Generation

AU - Грудинин, Михаил Артемович

AU - Вейбер, Евгения Николаевна

AU - Абрамов, Максим Викторович

N1 - Conference code: 29

PY - 2026/5/21

Y1 - 2026/5/21

N2 - Although large language models (LLMs) are capable of solving tasks that involve general reasoning patterns and factual knowledge, the reliability of their outputs remains highly dependent on the input text they receive. Increasing the length of the input introduces additional difficulty, as relevant facts may be distributed across different turns of a dialogue, making their integration more challenging and potentially reducing overall performance. This work introduces a method for constructing a dataset intended to assess semantic–episodic robustness in conversational settings. Central to the approach is the notion of a semantic–episodic task, defined as an evaluation unit in which a correct answer can only be produced by jointly leveraging episodic information provided exclusively within the dialogue and semantic knowledge that is not explicitly stated in the text but instead resides within the model itself. To support this, we develop a multi-stage pipeline for synthetic data generation that incorporates automated validation, enforces strict typing of outputs, and applies a graded difficulty scheme reflecting both the growth of input sequences and the increasing structural complexity of dialogues. The resulting benchmark is scalable, programmatically verifiable, and suitable for systematic comparison of models, prompting approaches, and memory mechanisms.

AB - Although large language models (LLMs) are capable of solving tasks that involve general reasoning patterns and factual knowledge, the reliability of their outputs remains highly dependent on the input text they receive. Increasing the length of the input introduces additional difficulty, as relevant facts may be distributed across different turns of a dialogue, making their integration more challenging and potentially reducing overall performance. This work introduces a method for constructing a dataset intended to assess semantic–episodic robustness in conversational settings. Central to the approach is the notion of a semantic–episodic task, defined as an evaluation unit in which a correct answer can only be produced by jointly leveraging episodic information provided exclusively within the dialogue and semantic knowledge that is not explicitly stated in the text but instead resides within the model itself. To support this, we develop a multi-stage pipeline for synthetic data generation that incorporates automated validation, enforces strict typing of outputs, and applies a graded difficulty scheme reflecting both the growth of input sequences and the increasing structural complexity of dialogues. The resulting benchmark is scalable, programmatically verifiable, and suitable for systematic comparison of models, prompting approaches, and memory mechanisms.

KW - large language models

KW - input sequence

KW - semantic knowledge

KW - episodic memory

KW - dataset generation

KW - benchmark

KW - long context

KW - strict answer formats

KW - automatic validation

KW - large language models

KW - input sequence

KW - semantic knowledge

KW - episodic memory

KW - dataset generation

KW - benchmark

KW - long context

KW - strict answer formats

KW - automatic validation

UR - https://ieeexplore.ieee.org/document/11625881

U2 - 10.1109/SCM71573.2026.11625881

DO - 10.1109/SCM71573.2026.11625881

M3 - Article in an anthology

SP - 166

EP - 169

BT - Proceedings of 2026 29th International Conference on Soft Computing and Measurements, SCM 2026

PB - Institute of Electrical and Electronics Engineers Inc.

CY - Saint Petersburg

Y2 - 20 May 2026 through 22 May 2026

ER -

ID: 160336398