All papers
Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.
The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development?2026SparDA: Sparse Decoupled Attention for Efficient Long-Context LLM Inference2026GENEB: Why Genomic Models Are Hard to Compare2026Rethinking Continual Experience Internalization for Self-Evolving LLM Agents2026GRAIL: Gradient-Reweighted Advantages for Reinforcement Learning with Verifiable Rewards2026Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning2026M$^3$Eval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks2026DAR: Deontic Reasoning with Agentic Harnesses2026Evaluating Large Language Models in Dynamic Clinical Decision-Making with Standardized Patient Cases2026Audio Interaction Model2026Reinforcement Learning from Rich Feedback with Distributional DAgger2026Streaming Communication in Multi-Agent Reasoning2026Agents' Last Exam2026Almieyar-Oryx-BloomBench: A Bilingual Multimodal Benchmark for Cognitively Informed Evaluation of Vision-Language Models2026AURA: Intent-Directed Probing for Implicit-Need Surfacing in Situated LLM Agents2026AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraints2026SubtleMemory: A Benchmark for Fine-Grained Relational Memory Discrimination in Long-Horizon AI Agents2026Evolving Agents in the Dark: Retrospective Harness Optimization via Self-Preference2026LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents2026LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs2026Reinforcement Learning Elicits Contextual Learning of Unseen Language Translation2026MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery2026Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectories2026Cosmos 3: Omnimodal World Models for Physical AI2026The Shadow Price of Reasoning: Economic Perspective on Optimal Budget Allocation for LLMs2026ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning2026Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models2026EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management2026Agent libOS: A Runtime Substrate for Capability-Controlled Self-Evolving LLM Agents2026Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking2026The Road Ahead in Autonomous Driving: The KITScenes Multimodal Dataset2026KVarN: Variance-Normalized KV-Cache Quantization Mitigates Error Accumulation in Reasoning Tasks2026SEAOTTER: Sensor Embedded Autoencoding with One-Time Transcode for Efficient Reconstruction2026MedForge: Interpretable Medical Deepfake Detection via Forgery-aware Reasoning2026AutoMedBench: Towards Medical AutoResearch with Agentic AI Models2026Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models2026StandardE2E: A Unified Framework for End-to-End Autonomous Driving Datasets2026Online Skill Learning for Web Agents via State-Grounded Dynamic Retrieval2026MapAgent: An Industrial-Grade Agentic Framework for City-scale Lane-level Map Generation2026AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks?2026Personal AI Agent for Camera Roll VQA2026DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models2026Toto 2.0: Time Series Forecasting Enters the Scaling Era2026When Tools Fail: Benchmarking Dynamic Replanning and Anomaly Recovery in LLM Agents2026Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models2026World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis2026OPRD: On-Policy Representation Distillation2026Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents2026Self-Distilled Policy Gradient2026Measuring Model Robustness via Fisher Information: Spectral Bounds, Theoretical Guarantees, and Practical Algorithms2026UnpredictaBench: A Benchmark for Evaluating Distributional Randomness in LLMs2026Data-Efficient Autoregressive-to-Diffusion Language Models via On-Policy Distillation2026SigmaScale: LLM Compression with SVD-based Low-Rank Decomposition and Learned Scaling Matrices2026MMAE: A Massive Multitask Audio Editing Benchmark2026SWE-Explore: Benchmarking How Coding Agents Explore Repositories2026Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings2026In-Context Multiple Instance Learning2026Benchmark Everything Everywhere All at Once2026Lean4Agent: Formal Modeling and Verification for Agent Workflow and Trajectory2026Direct 3D-Aware Object Insertion via Decomposed Visual Proxies2026dots.tts Technical Report2026Where Flow Matching Leaks: Characterising Membership Signals Along the Interpolation Path2026DuMate-DeepResearch: An Auditable Multi-Agent System with Recursive Search and Rubric-Grounded Reasoning2026Watch, Remember, Reason: Human-View Video Understanding with MLLMs2026PaperFlow: Profiling, Recommending, and Adapting Across Daily Paper Streams2026ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research2026Poise: Position-Aware One-Instruction Skill Injection for Silent Execution on LLM Agents2026Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding?2026Set-Based Transformer for Atmospheric Compensation in Standoff LWIR Hyperspectral Imaging2026CoVEBench: Can Video Editing Models Handle Complex Instructions?2026Trajectory-Refined Distillation2026PIPE-Cypher: Automatic Enterprise Benchmark Generation for Text-to-Cypher Systems2026AlloSpatial: Agentic Harness Framework for Spatial Reasoning in Foundation Models2026Hardening Agent Benchmarks with Adversarial Hacker-Fixer Loops2026FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention2026TRL-Bench: Standardizing Cross-Paradigm Representation-Level Evaluation of Tabular Encoders2026WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces2026Optical Reasoning: Rethinking Images as an Expressive Reasoning Medium Beyond Text2026End-to-End Context Compression at Scale2026SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks2026Correlation Is Not Enough: Embedding Human Metadata for Individual Causal Discovery2026SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Research2026Collaborative Human-Agent Protocol (CHAP)2026Tangram: Unlocking Non-Uniform KV Cache Compression for Efficient Multi-turn LLM Serving2026Skip a Layer or Loop It? Learning Program-of-Layers in LLMs2026Robotic Policy Adaptation via Weight-Space Meta-Learning2026How Far Can Chord-Symbol Time-Series Adaptation Carry Genre Identity? Capabilities and Boundaries in Multi-Genre Chord-Symbol Modeling2026SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision History2026BrainSurgery: Reproducible and Reliable Declarative Weight Manipulations for Model Editing and Upcycling2026Echo-Memory: A Controlled Study of Memory in Action World Models2026Rethinking the Divergence Regularization in LLM RL2026The Cold-Start Safety Gap in LLM Agents2026Bayesian-Agent: Posterior-Guided Skill Evolution Across LLM Agent Harnesses2026PaperMentor: A Human-Centered Multi-Agent Writing Tutor for AI Research Papers on Overleaf2026Bridging the Agent-World Gap: Text World Models for LLM-based Agents2026Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating2026Precision Is Not Faithfulness: Coverage-Aware Evaluation of Grounded Generation with a Complete Oracle2026PsychoSafe: Eliciting Psychologically-Informed Refusals in Large Language Models2026$τ$-Rec: A Verifiable Benchmark for Agentic Recommender Systems2026WebChallenger: A Reliable and Efficient Generalist Web Agent2026Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories2026RedAct: Redacting Agent Capability Traces for Procedural Skill Protection2026Building Social World Models with Large Language Models2026When is Your LLM Steerable?2026ICA Lens: Interpreting Language Models Without Training Another Dictionary2026Grammar-Constrained Decoding Can Jailbreak LLMs into Generating Malicious Code2026Toward Generalist Autonomous Research via Hypothesis-Tree Refinement2026FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents2026Adaptive Multi-Resolution Procedural Knowledge Compression for Large Language Models2026Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks2026Which Models Are Our Models Built On? Auditing Invisible Dependencies in Modern LLMs2026Redesign Mixture-of-Experts Routers with Manifold Power Iteration2026Quickest Detection of Hallucination Onset: Delay Bounds and Learned CUSUM Statistics2026Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior2026Zero-source LLM Hallucination Detection with Human-like Criteria Probing2026No Hidden Prompts Needed! You Can Game AI Peer Review with Presentation-Only Revisions2026Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents2026LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories2026One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders2026EurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific Discovery2026Orchestra-o1: Omnimodal Agent Orchestration2026AdaSR: Adaptive Streaming Reasoning with Hierarchical Relative Policy Optimization2026ClinHallu: A Benchmark for Diagnosing Stage-Wise Hallucinations in Medical MLLM Reasoning2026Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models2026Risk Under Pressure: Compute-Aware Evaluation of Adversarial Robustness in Language Models2026APPO: Agentic Procedural Policy Optimization2026Pythagoras-Prover: Advancing Efficient Formal Proving via Augmented Lean Formalisation2026Benchmarking AI Agents for Addressing Scientific Challenges Across Scales2026HarnessBridge: Learnable Bidirectional Controller for LLM Agent Harness2026OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data2026SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning2026MiniMax Sparse Attention2026A Benchmark and Framework for Evaluating Next Action Predictions in Spreadsheets2026AFFORDANCE20Q: Evaluating Affordance Reasoning from Physical Properties2026Squeeze-Release: Iterative Pruning with Exact Structural Minimization2026Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack2026VISTA: View-Consistent Self-Verified Training for GUI Grounding2026RSRCC: A Remote Sensing Regional Change Comprehension Benchmark Constructed via Retrieval-Augmented Best-of-N Ranking2026JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence2026Dr-DCI: Scaling Direct Corpus Interaction via Dynamic Workspace Expansion2026FastMix: Fast Data Mixture Optimization via Gradient Descent2026Harnessing cortical geometry, wiring, and function as inductive biases for recurrent neural networks2026Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning2026Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale2026CODA-BENCH: Can Code Agents Handle Data-Intensive Tasks?2026Selective Synergistic Learning for Video Object-Centric Learning2026LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies2026You Don't Need Strong Assumptions: Visual Representation Learning via Temporal Differences2026Thinking with Visual Grounding2026VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models2026PianoKontext: Expressive Performance Rendering from Deadpan Context2026VideoMDM: Towards 3D Human Motion Generation From 2D Supervision2026Dense Supervision, Sparse Updates: On the Sparsity and Geometry of On-Policy Distillation2026SpikF-GO: Spiking Fourier Graph Operators for Multivariate Time Series Forecasting2026Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments2026Rethinking the Role of Efficient Attention in Hybrid Architectures2026Reinforcement Learning-Guided Retrieval with Soft Fusion for Robust Multimodal Imitation Learning under Missing Modalities2026The Data Manifold under the Microscope2026A Gradient Perspective on RLVR Stability and Winner Advantage Policy Optimization2026PhoneHarness: Harnessing Phone-Use Agents through Mixed GUI, CLI, and Tool Actions2026Distilling Examples into Task Instructions: Enhanced In-Context Learning for Real-World B2B Conversations2026SciOrch: Learning to Orchestrate Expert LLMs for Solving Frontier Multimodal Scientific Reasoning Tasks2026Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence2026Who Flips? Self- and Cross-Model Counterarguments Reveal Answer Instability in LLMs2026VisualClaw: A Real-Time, Personalized Agent for the Physical World2026Taylor-Calibrate: Principled Initialization for Hybrid Linear Attention Distillation2026MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents2026Understanding the Behaviors of Environment-aware Information Retrieval2026Speaking the Language of Science: Toward a General-Purpose Generative Foundation Model for the Natural Sciences2026Context-Aware RL for Agentic and Multimodal LLMs2026CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent Economies2026TuneJury: An Open Metric for Improving Music Generation Preference Alignment2026Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI2026Reinforcing Dual-Path Reasoning in Spatial Vision Language Models2026GD$^2$PO: Mitigating Multi-Reward Conflicts via Group-Dynamic reward-Decoupled Policy Optimization2026ExpRL: Exploratory RL for LLM Mid-Training2026Geometric Action Model for Robot Policy Learning2026Verified Detection and Prevention of Concurrency Anomalies in Multi-Agent Large Language Model Systems2026MemSlides: A Hierarchical Memory Driven Agent Framework for Personalized Slide Generation with Multi-turn Local Revision2026RepSelect: Robust LLM Unlearning via Representation Selectivity2026OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation2026From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning2026EComAgentBench: Benchmarking Shopping Agents on Long-Horizon Tasks with Distributed Hidden Intent2026GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?2026ChLogic: Evaluating Robustness of Logical Reasoning in Chinese Expressions2026LegalHalluLens: Typed Hallucination Auditing and Calibrated Multi-Agent Debate for Trustworthy Legal AI2026Learning from the Self-future: On-policy Self-distillation for dLLMs2026Darshana Graph: A Parallel Commentary Corpus for Comparative Indian Philosophy, with Stylometric and Exploratory Graph Analyses2026Variable-Width Transformers2026SAE Interventions are Unreliable: Post-Intervention Recovery of Suppressed Behavior2026CEO-Bench: Can Agents Play the Long Game?2026Morpheus: A Morphology-Aware Neural Tokenizer and Word Embedder for Turkish2026Skill-MAS: Evolving Meta-Skill for Automatic Multi-Agent Systems2026Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models2026GateMem: Benchmarking Memory Governance in Multi-Principal Shared-Memory Agents2026SproutRAG: Attention-Guided Tree Search with Progressive Embeddings for Long-Document RAG2026JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Drafting2026REVES: REvision and VErification--Augmented Training for Test-Time Scaling2026Sumi: Open Uniform Diffusion Language Model from Scratch2026STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability2026