0

DeepSurvey-Bench: Evaluating Academic Value of Automatically Generated Scientific Surveys

The rapid development of automated survey generation technology has made it increasingly important to establish a comprehensive benchmark to evaluate the quality of generated surveys.

Preview
Year
2026
Hosting
Abstract onlyARXIV-DEFAULT

Cite

Notes

Only stored in your browser.

Attribution

Abstract & full text
arxiv.org/abs/2601.15307ARXIV-DEFAULT
TL;DR
Semantic Scholar
Attribution policy →

Abstract

The rapid development of automated survey generation technology has made it increasingly important to establish a comprehensive benchmark to evaluate the quality of generated surveys. Most existing benchmarks first construct ground-truth datasets by selecting human-written surveys based on limited selection criteria, such as citation counts and structural coherence, and evaluate generated surveys primarily based on conventional quality dimensions, including structural quality and reference relevance. However, these benchmarks have two key issues: (1) the datasets are insufficiently reliable because the selection criteria only identify highly cited or structurally coherent surveys without verifying their academic value; (2) the evaluation metrics mainly reflect the surface-level quality of generated surveys and are insufficient to assess their academic value. Together, these issues prevent existing benchmarks from effectively assessing the academic value of generated surveys. To address the above problems, we propose DeepSurvey-Bench, a comprehensive benchmark for evaluating the academic value of automatically generated surveys. Specifically, our proposed benchmark introduces a set of academic value evaluation criteria covering three dimensions: informational value, scholarly communication value, and research guidance value. We first construct a reliable dataset with academic value annotations based on these criteria, and then evaluate the academic value of generated surveys according to these criteria through a multi-LLM-as-a-judge approach. Extensive experiments demonstrate that DeepSurvey-Bench not only aligns closely with human assessments in evaluating the academic value of surveys, but also reveals underlying academic value beyond the reach of surface-level quality metrics, providing a foundation for fine-grained diagnosis and iterative improvement of generated surveys.