DOI

Although large language models (LLMs) are capable of solving tasks that involve general reasoning patterns and factual knowledge, the reliability of their outputs remains highly dependent on the input text they receive. Increasing the length of the input introduces additional difficulty, as relevant facts may be distributed across different turns of a dialogue, making their integration more challenging and potentially reducing overall performance. This work introduces a method for constructing a dataset intended to assess semantic–episodic robustness in conversational settings. Central to the approach is the notion of a semantic–episodic task, defined as an evaluation unit in which a correct answer can only be produced by jointly leveraging episodic information provided exclusively within the dialogue and semantic knowledge that is not explicitly stated in the text but instead resides within the model itself. To support this, we develop a multi-stage pipeline for synthetic data generation that incorporates automated validation, enforces strict typing of outputs, and applies a graded difficulty scheme reflecting both the growth of input sequences and the increasing structural complexity of dialogues. The resulting benchmark is scalable, programmatically verifiable, and suitable for systematic comparison of models, prompting approaches, and memory mechanisms.
Переведенное названиеСемантико-эпизодическая устойчивость больших языковых моделей: формализация и генерация оценочного набора данных
Язык оригиналаанглийский
Название основной публикацииProceedings of 2026 29th International Conference on Soft Computing and Measurements, SCM 2026
Место публикацииSaint Petersburg
ИздательInstitute of Electrical and Electronics Engineers Inc.
Страницы166-169
Число страниц4
ISBN (электронное издание)979-8-3195-4898-6
DOI
СостояниеОпубликовано - 21 мая 2026
СобытиеXXIX Международная конференция по мягким вычислениям и измерениям (SCM-2026) - Санкт-Петербургский государственный электротехнический университет «ЛЭТИ» им. В.И. Ульянова (Ленина), Санкт-Петербург, Российская Федерация
Продолжительность: 20 мая 202622 мая 2026
Номер конференции: 29
https://scm.etu.ru/2026/ru/

конференция

конференцияXXIX Международная конференция по мягким вычислениям и измерениям (SCM-2026)
Сокращенное названиеSCM'2026
Страна/TерриторияРоссийская Федерация
ГородСанкт-Петербург
Период20/05/2622/05/26
Сайт в сети Internet

    Области исследований

  • large language models, input sequence, semantic knowledge, episodic memory, dataset generation, benchmark, long context, strict answer formats, automatic validation

ID: 160336398