Reference-free evaluation of LLM-generated code is essential when execution-based testing is unavailable or costly. We compare two paradigms: $explicit LLM-as-a-Judge$ scoring, which assigns a quality score to a solution, and $log-probability scoring$, which uses $code mid task)$ as an instruction-free signal.Across HumanEval-X, we find that the two approaches capture $qualitatively different aspects$ of code correctness. Explicit judges --- particularly larger models --- perform strongly on generated code, reflecting their ability to reason about task-solution alignment, but fail to distinguish correct solutions from minimally mutated ones. Log-probability exhibits the opposite pattern: weaker performance on generated code, but consistent pairwise separation of canonical from mutated solutions.These results reveal a $discrimination-ranking dissociation$ and show that the two paradigms provide complementary, non-interchangeable signals: explicit judges capture semantic correctness, while log-probability captures local structural consistency.
Original languageEnglish
Title of host publicationProceedings of the Fifth Workshop on Generation, Evaluation and Metrics (GEM)
EditorsSimon Mille, Sebastian Gehrmann, Patrícia Schmidtová, Ondřej Dušek, Marzieh Fadaee, Kyle Lo, Enrico Santus, Gabriel Stanovsky
Place of PublicationSan Diego, California, USA
PublisherAssociation for Computational Linguistics
Pages574-581
Number of pages8
ISBN (Print)979-8-89176-423-1
DOIs
StatePublished - 1 Jul 2026
EventACL The Fifth Generation, Evaluation & Metrics Workshop (GEM) - San Diego, United States
Duration: 1 Jul 2026 → …
https://gem-workshop.com/

Conference

ConferenceACL The Fifth Generation, Evaluation & Metrics Workshop (GEM)
Abbreviated titleGEM at ACL 2026
Country/TerritoryUnited States
CitySan Diego
Period1/07/26 → …
Internet address

ID: 156757179