Ссылки

DOI

Reference-free evaluation of LLM-generated code is essential when execution-based testing is unavailable or costly. We compare two paradigms: $explicit LLM-as-a-Judge$ scoring, which assigns a quality score to a solution, and $log-probability scoring$, which uses $code mid task)$ as an instruction-free signal.Across HumanEval-X, we find that the two approaches capture $qualitatively different aspects$ of code correctness. Explicit judges --- particularly larger models --- perform strongly on generated code, reflecting their ability to reason about task-solution alignment, but fail to distinguish correct solutions from minimally mutated ones. Log-probability exhibits the opposite pattern: weaker performance on generated code, but consistent pairwise separation of canonical from mutated solutions.These results reveal a $discrimination-ranking dissociation$ and show that the two paradigms provide complementary, non-interchangeable signals: explicit judges capture semantic correctness, while log-probability captures local structural consistency.
Язык оригиналаанглийский
Название основной публикацииProceedings of the Fifth Workshop on Generation, Evaluation and Metrics (GEM)
РедакторыSimon Mille, Sebastian Gehrmann, Patrícia Schmidtová, Ondřej Dušek, Marzieh Fadaee, Kyle Lo, Enrico Santus, Gabriel Stanovsky
Место публикацииSan Diego, California, USA
ИздательAssociation for Computational Linguistics
Страницы574-581
Число страниц8
ISBN (печатное издание)979-8-89176-423-1
DOI
СостояниеОпубликовано - 1 июл 2026
СобытиеACL The Fifth Generation, Evaluation & Metrics Workshop (GEM) - San Diego, Соединенные Штаты Америки
Продолжительность: 1 июл 2026 → …
https://gem-workshop.com/

конференция

конференцияACL The Fifth Generation, Evaluation & Metrics Workshop (GEM)
Сокращенное названиеGEM at ACL 2026
Страна/TерриторияСоединенные Штаты Америки
ГородSan Diego
Период1/07/26 → …
Сайт в сети Internet

ID: 156757179