Research output: Chapter in Book/Report/Conference proceeding › Conference contribution › Research › peer-review
Semantic vs. Structural Signals: Log-Probability and LLM-as-a-Judge for Reference-Free Code Evaluation. / Fedrushkov, Dmitriy; He, Yulong; Smirnov, Ivan; Aliev, Artem; Kovalchuk, Sergey.
Proceedings of the Fifth Workshop on Generation, Evaluation and Metrics (GEM). ed. / Simon Mille; Sebastian Gehrmann; Patrícia Schmidtová; Ondřej Dušek; Marzieh Fadaee; Kyle Lo; Enrico Santus; Gabriel Stanovsky. San Diego, California, USA : Association for Computational Linguistics, 2026. p. 574-581.Research output: Chapter in Book/Report/Conference proceeding › Conference contribution › Research › peer-review
}
TY - GEN
T1 - Semantic vs. Structural Signals: Log-Probability and LLM-as-a-Judge for Reference-Free Code Evaluation
AU - Fedrushkov, Dmitriy
AU - He, Yulong
AU - Smirnov, Ivan
AU - Aliev, Artem
AU - Kovalchuk, Sergey
PY - 2026/7/1
Y1 - 2026/7/1
N2 - Reference-free evaluation of LLM-generated code is essential when execution-based testing is unavailable or costly. We compare two paradigms: $explicit LLM-as-a-Judge$ scoring, which assigns a quality score to a solution, and $log-probability scoring$, which uses $code mid task)$ as an instruction-free signal.Across HumanEval-X, we find that the two approaches capture $qualitatively different aspects$ of code correctness. Explicit judges --- particularly larger models --- perform strongly on generated code, reflecting their ability to reason about task-solution alignment, but fail to distinguish correct solutions from minimally mutated ones. Log-probability exhibits the opposite pattern: weaker performance on generated code, but consistent pairwise separation of canonical from mutated solutions.These results reveal a $discrimination-ranking dissociation$ and show that the two paradigms provide complementary, non-interchangeable signals: explicit judges capture semantic correctness, while log-probability captures local structural consistency.
AB - Reference-free evaluation of LLM-generated code is essential when execution-based testing is unavailable or costly. We compare two paradigms: $explicit LLM-as-a-Judge$ scoring, which assigns a quality score to a solution, and $log-probability scoring$, which uses $code mid task)$ as an instruction-free signal.Across HumanEval-X, we find that the two approaches capture $qualitatively different aspects$ of code correctness. Explicit judges --- particularly larger models --- perform strongly on generated code, reflecting their ability to reason about task-solution alignment, but fail to distinguish correct solutions from minimally mutated ones. Log-probability exhibits the opposite pattern: weaker performance on generated code, but consistent pairwise separation of canonical from mutated solutions.These results reveal a $discrimination-ranking dissociation$ and show that the two paradigms provide complementary, non-interchangeable signals: explicit judges capture semantic correctness, while log-probability captures local structural consistency.
U2 - 10.18653/v1/2026.gem-main.55
DO - 10.18653/v1/2026.gem-main.55
M3 - Conference contribution
SN - 979-8-89176-423-1
SP - 574
EP - 581
BT - Proceedings of the Fifth Workshop on Generation, Evaluation and Metrics (GEM)
A2 - Mille, Simon
A2 - Gehrmann, Sebastian
A2 - Schmidtová, Patrícia
A2 - Dušek, Ondřej
A2 - Fadaee, Marzieh
A2 - Lo, Kyle
A2 - Santus, Enrico
A2 - Stanovsky, Gabriel
PB - Association for Computational Linguistics
CY - San Diego, California, USA
T2 - ACL The Fifth Generation, Evaluation & Metrics Workshop (GEM)
Y2 - 1 July 2026
ER -
ID: 156757179