Publication Type

PhD Dissertation

Version

publishedVersion

Publication Date

6-2026

Abstract

The rapid advancement of large language models (LLMs) has created an urgent need for evaluation methodologies that are timely, scalable, reliable, and informative. Conventional evaluation benchmarks, although essential for measuring model capabilities and guiding model development, are often constructed and maintained through labor-intensive human annotation. As LLMs continue to improve through increases in model scale, training data, and computational resources, static benchmarks may quickly lose discriminative power. Moreover, the growing use of large and diverse training corpora increases the risk of benchmark leakage, which can inflate evaluation results and obscure the true capabilities of models. These challenges call for a transition from traditional manualevaluation towardautomaticevaluationframeworks thatcandynamically construct evaluation data, assess model outputs efficiently, and provide deeper diagnostic feedback.

This thesis investigates automatic evaluation methodologies for large language models under a general question-based evaluation paradigm. It focuses on three closely related challenges: automatic evaluation dataset construction, automatic evaluation pipelines, and evaluation beyond output-level performance. First, this thesis studies how LLMs can be used to automatically construct, modify, and update evaluation data. By leveraging existing knowledge-intensive benchmarks and LLM-based data transformation, it proposes automatic robustness evaluation and dataset updating methods that support scalable benchmark construction, controllable difficulty adjustment, and mitigation of benchmark leakage. Second, this thesis develops an end-to-end automatic evaluation pipeline, termed Language-Model-as-an-Examiner, in which LLMs serve as reference-free evaluators for judging model responses, conducting comparative assessment, and generating model rankings. Through rubric-based evaluation and multi-model peer assessment, the proposed framework improves the scalability, robustness, and alignment of automatic judgment with human preferences. Third, this thesis extends evaluation beyond external performance scores by incorporating internal model signals. It introduces mechanism-aware indicators and the Model Utilization Index (MUI) to quantify the internal effort expended by models during inference, providing a complementary perspective for model comparison, training diagnostics, and dataset characterization.

Overall, this thesis contributes to the development of reliable auto-evaluation systems for large language models. The proposed methods demonstrate that automatic evaluation can reduce reliance on manual annotation, sustain benchmark effectiveness under rapid model advancement, and provide more informative feedback for understanding and improving LLMs. By integrating automatic data construction, automatic answer judgment, and mechanism-aware analysis, this work offers both practical tools and theoretical insights for advancing evaluation practices in the era of increasingly capable language models.

Degree Awarded

PhD in Computer Science

Discipline

Artificial Intelligence and Robotics | Software Engineering

Supervisor(s)

CAO, Yixin; SUN, Qianru

First Page

1

Last Page

188

Publisher

Singapore Management University

City or Country

Singapore

Copyright Owner and License

Author

Share

COinS