5 Oct 2026
A recurring design principle in modern optimizers is to decouple update magnitude from the raw gradient norm, yet its consequences for learning-curve and resource scaling remain unclear.
Trending research and the full catalog - each paper linked to the benchmarks, methods, and models it introduces.
Filtering here covers the 2,000 most recent papers, as much as one page can hold in memory. See the full index of 22,422 papers.
5 Oct 2026
A recurring design principle in modern optimizers is to decouple update magnitude from the raw gradient norm, yet its consequences for learning-curve and resource scaling remain unclear.
5 Oct 2026
We develop and test a theory of language model representations in which there exist atomic features. Our main theoretical insight is that in such a model, sparse dictionaries (e.g., SAEs) of increasing size recover an increasing prefix of the most prevalent atoms in the training…
5 Oct 2026
Representations of translated sentences are similar in the inner layers of multilingual language models -- an observation connected to the platonic representation hypothesis, yet unexplained theoretically.
5 Oct 2026
Deterministic benchmark scores show that an agent received credit, but not whether that credit was earned, reported honestly, or would hold on a second run. We introduce a Trust Layer for Agent Evaluation, an additive post-hoc framework that reports, beside each recorded score,…
5 Oct 2026
Systems engineers use several diagrams to describe the structure and behavior of systems. Engineers create these diagrams together to make sure that they use the same elements and remain consistent with one another.
5 Oct 2026
As AI agents increasingly tackle complex repository-level coding tasks, distributing work across multiple agents is a natural way to scale beyond the capabilities of a single agent.
4 Oct 2026
An electromagnetic inversion scheme that integrates stochastic optimization (STO) and plug-and-play (PNP) regularization into contrast-source inversion (CSI), termed STO-PNP-CSI, is developed.
4 Oct 2026
This paper presents the workflow for building an automated ArcGIS Pro tool using ArcPy to extract the Land Surface Temperature (LST) from Landsat 7, 8 and 9 datasets. The tool eliminates the need for manual band selection and repetitive raster computations by automating the…
4 Oct 2026
We ask whether physically meaningful correspondences discovered inside a trained scientific model remain meaningful outside the conditions under which they were discovered.
4 Oct 2026
Pointwise Mutual Information (PMI) is a core measure of testing word association, and Skip-gram with Negative Sampling (SGNS) is essentially a method that implicitly factorizes a shifted PMI matrix.
4 Oct 2026
Evaluating vision encoders requires metrics that reliably predict their downstream performance in multimodal large language models (MLLMs). Although recent studies have shown that cross-modal metrics can better capture such performance, unimodal metrics remain the dominant…
4 Oct 2026
This study analyzes the predictive power of LM-derived, token-level surprisal on Mandarin Chinese reading times. We first propose the Shortest Matching Sequence (SMS), an alignment scheme that maps between the word segmentation assumed by eye-tracking corpora and the LMs'…
4 Oct 2026
Reinforcement learning controllers for underwater vehicles are usually trained against one physics representation and deployed against another, so reported performance does not always describe behavior outside training.
4 Oct 2026
We develop a Unified Scaling Law and a Unified Theory of Time Series Learning to understand how model capacity and historical information support forecasting. Across different lookback lengths and forecast horizons, we analyze 18,768 experimental cells from 21 checkpoints on 23…
4 Oct 2026
Large language model (LLM)-based multi-agent systems increasingly rely on memory to transform execution trajectories into reusable procedural knowledge. Yet repeated retrieval also makes memory errors persistent: memory pollution arises when outdated, weakly supported, or…
3 Oct 2026
We propose $\mathrm{TRIZ}^{a}$ (TRIZ exponentiated by an agent), a general R\&D automation paradigm that combines TRIZ inventive theory with LLM-driven agent evolutionary search.
3 Oct 2026
Designing effective reward functions in RL typically requires substantial expertise and trial and error. Recent work automates this process with LLM-based systems that generate and iteratively improve reward code using policy feedback.
3 Oct 2026
Exactly computing the maximum function is a standard test case for studying depth in ReLU networks. Two hidden layers are known to suffice for up to twelve inputs through computer-assisted constructions.
3 Oct 2026
Reservoir computing (RC) is a machine learning technique for data-driven modeling of dynamics for forecasting and control, primarily studied in computer science and physics literature with promising applications in neural network learning, physical computing, and neuroscience.
3 Oct 2026
Extreme winter weather has repeatedly disrupted gas-fired power generation in the United States, yet the plant-level data needed to systematically quantify outage risk remain proprietary.
3 Oct 2026
Fine-tuning a language model on new text degrades what it already does. Replay-free projectors such as Adam-NSCL and GPM forbid one shared subspace of a layer's inputs in every row of the update. The tropical geometry of a ReLU layer shows why this is too coarse.
2 Oct 2026
We study constrained online convex optimization with adversarial convex losses and constraints ($\mathsf{COCO}$). At each round \(t\in[T]\), a learner selects \(x_t\) from a \(d\)-dimensional convex decision set \(\mathcal X\), after which an adaptive adversary reveals a convex…
2 Oct 2026
To deploy deep neural networks on edge hardware, highly efficient inference schemes are necessary that retain high accuracy. This work presents W16A16, a high precision (16-bit), fast speed, low energy quantization method.
2 Oct 2026
We introduce 4DCodeBench, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs.