Although large language models (LLMs) are capable of solving tasks that involve general reasoning patterns and factual knowledge, the reliability of their outputs remains highly dependent on the input text they receive. Increasing the length of the input introduces additional difficulty, as relevant facts may be distributed across different turns of a dialogue, making their integration more challenging and potentially reducing overall performance. This work introduces a method for constructing a dataset intended to assess semantic–episodic robustness in conversational settings. Central to the approach is the notion of a semantic–episodic task, defined as an evaluation unit in which a correct answer can only be produced by jointly leveraging episodic information provided exclusively within the dialogue and semantic knowledge that is not explicitly stated in the text but instead resides within the model itself. To support this, we develop a multi-stage pipeline for synthetic data generation that incorporates automated validation, enforces strict typing of outputs, and applies a graded difficulty scheme reflecting both the growth of input sequences and the increasing structural complexity of dialogues. The resulting benchmark is scalable, programmatically verifiable, and suitable for systematic comparison of models, prompting approaches, and memory mechanisms.
Translated title of the contributionСемантико-эпизодическая устойчивость больших языковых моделей: формализация и генерация оценочного набора данных
Original languageEnglish
Title of host publicationProceedings of 2026 29th International Conference on Soft Computing and Measurements, SCM 2026
Place of PublicationSaint Petersburg
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages166-169
Number of pages4
ISBN (Electronic)979-8-3195-4898-6
DOIs
StatePublished - 21 May 2026
EventXXIX Международная конференция по мягким вычислениям и измерениям (SCM-2026) - Санкт-Петербургский государственный электротехнический университет «ЛЭТИ» им. В.И. Ульянова (Ленина), Санкт-Петербург, Russian Federation
Duration: 20 May 202622 May 2026
Conference number: 29
https://scm.etu.ru/2026/ru/

Conference

ConferenceXXIX Международная конференция по мягким вычислениям и измерениям (SCM-2026)
Abbreviated titleSCM'2026
Country/TerritoryRussian Federation
CityСанкт-Петербург
Period20/05/2622/05/26
Internet address

    Research areas

  • large language models, input sequence, semantic knowledge, episodic memory, dataset generation, benchmark, long context, strict answer formats, automatic validation

ID: 160336398