We introduce SciTrek, a synthetic question-answering dataset for assessing and improving long-context numerical reasoning in large language models (LLMs). Existing long-context datasets with inputs beyond 64K tokens either target simple information retrieval or, when they do involve reasoning, rely on artificial contexts, while numerical reasoning remains largely overlooked in both cases. SciTrek addresses these limitations with questions that require numerical operations (e.g., counting, sorting, aggregation, and comparison) over collections of full-text scientific articles. Questions are generated automatically by formulating them as SQL queries over a database of article metadata (titles, authors, and references), and ground-truth answers are obtained by executing these queries. The underlying SQL provides a transparent, verifiable specification of the reasoning each question requires, enabling fine-grained error analysis, while the fully automated pipeline scales to arbitrary context lengths and dataset sizes with minimal human supervision, mitigating data contamination and supplying abundant data for post-training. Extensive experiments show that frontier open-weight and proprietary LLMs struggle even with ostensibly simple questions: the best-performing model achieves only 46.5% exact match at 128K tokens, and performance degrades steadily as contexts grow. Fine-grained analysis reveals systematic weaknesses on citation-related questions and on compound logical conditions, particularly those involving negation. Finally, post-training open-weight models on SciTrek improves their numerical reasoning in ways that generalise to out-of-domain long-context tasks.
SciTrek: Evaluating and Improving Long-Context Numerical Reasoning over Scientific Articles
We introduce SciTrek, a synthetic question-answering dataset for assessing and improving long-context numerical reasoning in large language models (LLMs). Existing long-context datasets with inputs beyond 64K tokens either target simple information retrieval or, when they do…
- Preview

- Year
- 2025
- Hosting
- Full text hostedCC-BY-4.0
Cite
Notes
Only stored in your browser.
Attribution
- Abstract & full text
- arxiv.org/abs/2509.21028CC-BY-4.0
- TL;DR
- Semantic Scholar