Standard

Semantic vs. Structural Signals: Log-Probability and LLM-as-a-Judge for Reference-Free Code Evaluation. / Fedrushkov, Dmitriy; He, Yulong; Smirnov, Ivan; Aliev, Artem; Kovalchuk, Sergey.

Proceedings of the Fifth Workshop on Generation, Evaluation and Metrics (GEM). ed. / Simon Mille; Sebastian Gehrmann; Patrícia Schmidtová; Ondřej Dušek; Marzieh Fadaee; Kyle Lo; Enrico Santus; Gabriel Stanovsky. San Diego, California, USA : Association for Computational Linguistics, 2026. p. 574-581.

Research output: Chapter in Book/Report/Conference proceedingConference contributionResearchpeer-review

Harvard

Fedrushkov, D, He, Y, Smirnov, I, Aliev, A & Kovalchuk, S 2026, Semantic vs. Structural Signals: Log-Probability and LLM-as-a-Judge for Reference-Free Code Evaluation. in S Mille, S Gehrmann, P Schmidtová, O Dušek, M Fadaee, K Lo, E Santus & G Stanovsky (eds), Proceedings of the Fifth Workshop on Generation, Evaluation and Metrics (GEM). Association for Computational Linguistics, San Diego, California, USA, pp. 574-581, ACL The Fifth Generation, Evaluation & Metrics Workshop (GEM), San Diego, California, United States, 1/07/26. https://doi.org/10.18653/v1/2026.gem-main.55

APA

Fedrushkov, D., He, Y., Smirnov, I., Aliev, A., & Kovalchuk, S. (2026). Semantic vs. Structural Signals: Log-Probability and LLM-as-a-Judge for Reference-Free Code Evaluation. In S. Mille, S. Gehrmann, P. Schmidtová, O. Dušek, M. Fadaee, K. Lo, E. Santus, & G. Stanovsky (Eds.), Proceedings of the Fifth Workshop on Generation, Evaluation and Metrics (GEM) (pp. 574-581). Association for Computational Linguistics. https://doi.org/10.18653/v1/2026.gem-main.55

Vancouver

Fedrushkov D, He Y, Smirnov I, Aliev A, Kovalchuk S. Semantic vs. Structural Signals: Log-Probability and LLM-as-a-Judge for Reference-Free Code Evaluation. In Mille S, Gehrmann S, Schmidtová P, Dušek O, Fadaee M, Lo K, Santus E, Stanovsky G, editors, Proceedings of the Fifth Workshop on Generation, Evaluation and Metrics (GEM). San Diego, California, USA: Association for Computational Linguistics. 2026. p. 574-581 https://doi.org/10.18653/v1/2026.gem-main.55

Author

Fedrushkov, Dmitriy ; He, Yulong ; Smirnov, Ivan ; Aliev, Artem ; Kovalchuk, Sergey. / Semantic vs. Structural Signals: Log-Probability and LLM-as-a-Judge for Reference-Free Code Evaluation. Proceedings of the Fifth Workshop on Generation, Evaluation and Metrics (GEM). editor / Simon Mille ; Sebastian Gehrmann ; Patrícia Schmidtová ; Ondřej Dušek ; Marzieh Fadaee ; Kyle Lo ; Enrico Santus ; Gabriel Stanovsky. San Diego, California, USA : Association for Computational Linguistics, 2026. pp. 574-581

BibTeX

@inproceedings{7c626d6179c74f80a4a46156c53d956d,
title = "Semantic vs. Structural Signals: Log-Probability and LLM-as-a-Judge for Reference-Free Code Evaluation",
abstract = "Reference-free evaluation of LLM-generated code is essential when execution-based testing is unavailable or costly. We compare two paradigms: $explicit LLM-as-a-Judge$ scoring, which assigns a quality score to a solution, and $log-probability scoring$, which uses $code mid task)$ as an instruction-free signal.Across HumanEval-X, we find that the two approaches capture $qualitatively different aspects$ of code correctness. Explicit judges --- particularly larger models --- perform strongly on generated code, reflecting their ability to reason about task-solution alignment, but fail to distinguish correct solutions from minimally mutated ones. Log-probability exhibits the opposite pattern: weaker performance on generated code, but consistent pairwise separation of canonical from mutated solutions.These results reveal a $discrimination-ranking dissociation$ and show that the two paradigms provide complementary, non-interchangeable signals: explicit judges capture semantic correctness, while log-probability captures local structural consistency.",
author = "Dmitriy Fedrushkov and Yulong He and Ivan Smirnov and Artem Aliev and Sergey Kovalchuk",
year = "2026",
month = jul,
day = "1",
doi = "10.18653/v1/2026.gem-main.55",
language = "English",
isbn = "979-8-89176-423-1",
pages = "574--581",
editor = "Simon Mille and Sebastian Gehrmann and Patr{\'i}cia Schmidtov{\'a} and Ond{\v r}ej Du{\v s}ek and Marzieh Fadaee and Kyle Lo and Enrico Santus and Gabriel Stanovsky",
booktitle = "Proceedings of the Fifth Workshop on Generation, Evaluation and Metrics (GEM)",
publisher = "Association for Computational Linguistics",
address = "United States",
note = "ACL The Fifth Generation, Evaluation & Metrics Workshop (GEM), GEM at ACL 2026 ; Conference date: 01-07-2026",
url = "https://gem-workshop.com/",

}

RIS

TY - GEN

T1 - Semantic vs. Structural Signals: Log-Probability and LLM-as-a-Judge for Reference-Free Code Evaluation

AU - Fedrushkov, Dmitriy

AU - He, Yulong

AU - Smirnov, Ivan

AU - Aliev, Artem

AU - Kovalchuk, Sergey

PY - 2026/7/1

Y1 - 2026/7/1

N2 - Reference-free evaluation of LLM-generated code is essential when execution-based testing is unavailable or costly. We compare two paradigms: $explicit LLM-as-a-Judge$ scoring, which assigns a quality score to a solution, and $log-probability scoring$, which uses $code mid task)$ as an instruction-free signal.Across HumanEval-X, we find that the two approaches capture $qualitatively different aspects$ of code correctness. Explicit judges --- particularly larger models --- perform strongly on generated code, reflecting their ability to reason about task-solution alignment, but fail to distinguish correct solutions from minimally mutated ones. Log-probability exhibits the opposite pattern: weaker performance on generated code, but consistent pairwise separation of canonical from mutated solutions.These results reveal a $discrimination-ranking dissociation$ and show that the two paradigms provide complementary, non-interchangeable signals: explicit judges capture semantic correctness, while log-probability captures local structural consistency.

AB - Reference-free evaluation of LLM-generated code is essential when execution-based testing is unavailable or costly. We compare two paradigms: $explicit LLM-as-a-Judge$ scoring, which assigns a quality score to a solution, and $log-probability scoring$, which uses $code mid task)$ as an instruction-free signal.Across HumanEval-X, we find that the two approaches capture $qualitatively different aspects$ of code correctness. Explicit judges --- particularly larger models --- perform strongly on generated code, reflecting their ability to reason about task-solution alignment, but fail to distinguish correct solutions from minimally mutated ones. Log-probability exhibits the opposite pattern: weaker performance on generated code, but consistent pairwise separation of canonical from mutated solutions.These results reveal a $discrimination-ranking dissociation$ and show that the two paradigms provide complementary, non-interchangeable signals: explicit judges capture semantic correctness, while log-probability captures local structural consistency.

U2 - 10.18653/v1/2026.gem-main.55

DO - 10.18653/v1/2026.gem-main.55

M3 - Conference contribution

SN - 979-8-89176-423-1

SP - 574

EP - 581

BT - Proceedings of the Fifth Workshop on Generation, Evaluation and Metrics (GEM)

A2 - Mille, Simon

A2 - Gehrmann, Sebastian

A2 - Schmidtová, Patrícia

A2 - Dušek, Ondřej

A2 - Fadaee, Marzieh

A2 - Lo, Kyle

A2 - Santus, Enrico

A2 - Stanovsky, Gabriel

PB - Association for Computational Linguistics

CY - San Diego, California, USA

T2 - ACL The Fifth Generation, Evaluation & Metrics Workshop (GEM)

Y2 - 1 July 2026

ER -

ID: 156757179