All papers
Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.
Learning User Simulators with Turing Rewards2026Native Active Perception as Reasoning for Omni-Modal Understanding2026Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness2026EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts2026PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models2026Uncertainty Decomposition for Clarification Seeking in LLM Agents2026Think Again or Think Longer? Selective Verification for Budget-Aware Reasoning2026DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis2026Human Universal Grasping2026Playful Agentic Robot Learning2026FAPO: Fully Automated Prompt Optimization of Multi-Step LLM Pipelines2026Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents2026Seeing Before Reasoning: Decoupling Perception and Reasoning for Shortcut-Resilient Multimodal On-Policy Self-Distillation2026Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models2026SkillHarness: Harnessing Safe Skills for Computer-Use Agents2026Fara-1.5: Scalable Learning Environments for Computer Use Agents2026PoLAR: Factorizing Extent and Mode in Latent Actions for Robot Policy Learning2026DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams2026A Verifiable Search Is Not a Learnable Chain-of-Thought2026BioMatrix: Towards a Comprehensive Biological Foundation Model Spanning the Modality Matrix of Sequences, Structures, and Language2026PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems2026PhoneBuddy: Training Open Models for Agentic Phone Use2026ReNIO: Reweighting Negative Trajectory Importance for LLM On-Policy Distillation2026Capable but Careless: Do Computer-Use Agents Follow Contextual Integrity?2026AOHP: An Open-Source OS-Level Agent Harness for Personalized, Efficient and Secure Interaction2026VeriEvol: Scaling Multimodal Mathematical Reasoning via Verifiable Evol-Instruct2026Causal Discovery in the Era of Agents2026Go-with-the-Track: Video Compositing and Motion Control with Point Tracking2026TROPT: An Open Framework for Unifying and Advancing Discrete Text Optimization2026Toward Open Weight Models Without Risks: Separating Public and Private Capabilities in LLMs2026EvoEmbedding: Evolvable Representations for Long-Context Retrieval and Agentic Memory2026Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding2026OpenBioRQ: Unsolved Biomedical Research Questions for Agents2026HAKARI-Bench: A Lightweight Benchmark for Comparing Retrieval Architectures and Efficiency Settings under Unified Conditions2026SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning2026Unlimited OCR Works2026Tmax: A simple recipe for terminal agents2026EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions2026Learning to Trigger: Reinforcement Learning at the Large Hadron Collider2026CAVEWOMAN: How Large Language Models Behave Under Linguistic Input and Output Compression2026ReMMD: Realistic Multilingual Multi-Image Agentic Verification for Multimodal Misinformation Detection2026FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning2026GUI vs. CLI: Execution Bottlenecks in Screen-Only and Skill-Mediated Computer-Use Agents2026OpenThoughts-Agent: Data Recipes for Agentic Models2026InSight: Self-Guided Skill Acquisition via Steerable VLAs2026RoPE-Aware Bit Allocation for KV-Cache Quantization2026Holistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement Learning2026AsyncOPD: How Stale Can On-Policy Distillation Be?2026Real vs. Complex Spectral Bases for Neural Operators: The Role of Green's Function Alignment2026NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?2026Qwen-AgentWorld: Language World Models for General Agents2026DREAM: Dense Retrieval Embeddings via Autoregressive Modeling2026Are We Ready For An Agent-Native Memory System?2026Do Thinking Tokens Help with Safety?2026Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR2026What Intermediate Layers Know: Detecting Jailbreaks from Entropy Dynamics2026Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation2026Improved Large Language Diffusion Models2026Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents2026Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It2026Constraint Tax in Open-Weight LLMs: An Empirical Study of Tool Calling Suppression Under Structured Output Constraints2026OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning2026Play2Perfect: What Matters in Dexterous Play Pretraining for Precise Assembly?2026Confidence-Aware Tool Orchestration for Robust Video Understanding2026Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)2026EO-WM: A Physically Informed World Model for Probabilistic Earth Observation Forecasting2026ReFreeKV: Towards Threshold-Free KV Cache Compression2025ELF: Embedded Language Flows2026What the LLM Should Not Say: Boundary-Aware Context Grounding for A Seven-Channel EEG Agent2026Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models?2026ProMSA:Progressive Multimodal Search Agents for Knowledge-Based Visual Question Answering2026Parallel Rollout Approximation for Pixel-Space Autoregressive Image Generation2026MultiHashFormer: Hash-based Generative Language Models2026Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction2026Fast LeWorldModel2026Scalable Message-Passing Quantum Graph Neural Networks in the Weisfeiler-Leman Hierarchy2026How Good Can Linear Models Be for Time-Series Forecasting?2026Hallucination in World Models is Predictable and Preventable2026Delayed Verification Destabilizes Multi-Agent LLM Belief: Instability Thresholds and Optimal Corrector Placement2026TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents2026When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling2026Agentic Abstention: Do Agents Know When to Stop Instead of Act?2026Flow Matching in Feature Space for Stochastic World Modeling2026PolicyGuard: A Dialogue-Grounded Sub-Agent Verifier for Policy Adherence in LLM Agents2026Hierarchical Experimentalist Agents2026Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction2026Rank-Aware Hyperbolic Alignment for Vision-Language Dataset Distillation2026OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks2026RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources2026One Scene, Two Depths: Probing Geometric Ambiguity in Monocular Foundation Models2026MatMMExtract: An Open-Source Pipeline for Panel-Level Extraction of Grounded Image-Text Pairs from Materials Science Literature2026GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots2026SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailing2026SWE-Together: Evaluating Coding Agents in Interactive User Sessions2026MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs2026Little Brains, Big Feats: Exploring Compact Language Models2026Automating the Design of Embodied Agent Architectures2026Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) Benchmark2026Beyond IID: How General Are Tabular Foundation Models, Really?2026DOPD: Dual On-policy Distillation2026DataComp-VLM: Improved Open Datasets for Vision-Language Models2026Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks2026Multi-Block Diffusion Language Models2026SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions2026Morphing into Hybrid Attention Models2026Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent2026LUMOS: A Semantic Operating-System Layer for Accessibility-Grounded AI Agents2026LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent2026HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents20263D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance2026QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents2026SeKV: Resolution-Adaptive KV Cache with Hierarchical Semantic Memory for Long-Context LLM Inference2026BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding2026AutoTrainess: Teaching Language Models to Improve Language Models Autonomously2026When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing Errors2026Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs2026Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning2026GRPO, Dr. GRPO, and DAPO Are Three Operations on One Number: The Group-Standard-Deviation Identity2026VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement2026Domain Arithmetic: One-Shot VLA Adaptation under Environmental Shifts2026Valdi: Value Diffusion World Models2026Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual Recombination2026TiRex-2: Generalizing TiRex to Multivariate Data and Streaming2026The State-Prediction Separation Hypothesis2026Multi-Turn Agentic Scientific Literature Search via Workflow Induction2026AutoMem: Automated Learning of Memory as a Cognitive Skill2026Measuring the Gap Between Human and LLM Research Ideas2026LeVLJEPA: End-to-End Vision-Language Pretraining Without Negatives2026Scaling Laws for Grid-Based Approximate Nearest Neighbor Search in High Dimensions2026RuleChef: Grounding LLM Task Knowledge in Human-Editable Rules2026AgenticDataBench: A Comprehensive Benchmark for Data Agents2026Denser $\neq$ Better: Limits of On-Policy Self-Distillation for Continual Post-Training2026PACE: A Proxy for Agentic Capability Evaluation2026AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents2026EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments2026Program-as-Weights: A Programming Paradigm for Fuzzy Functions2026MemSyco-Bench: Benchmarking Sycophancy in Agent Memory2026WARP: Weight-Space Analysis for Recovering Training Data Portfolios2026Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification2026AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models2026Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs2026Gemma 4 Technical Report2026Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning2026Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling2026SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe2026Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process2026CGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-Centric 3D Scene Generation2026OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers2026Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization2026UI-MOPD: Multi-Platform On-Policy Distillation for Unified GUI Agents2026dOPSD: On-Policy Self-Distillation for Diffusion Language Models2026RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies2026ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes2026Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval2026Multi-Turn On-Policy Distillation with Prefix Replay2026DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation2026Multiplayer Interactive World Models with Representation Autoencoders2026GaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation Tasks2026Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation2026LLM-as-a-Verifier: A General-Purpose Verification Framework2026Weak-to-Strong Generalization via Direct On-Policy Distillation2026Vidu S1: A Real-Time Interactive Video Generation Model2026Taste-aware music retrieval from audio embeddings2026TESSERA v2: Scaling Pixel-wise Earth Foundation Models2026MANCE: Manifold Aware Concept Erasure2026PAST-TIDE: Prototype-Anchored Statement Tuning with Topic-Invariant Normalization for Stance Detection2026EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments2026PraMem: Practice-derived Experiential Memory for Long-horizon Behavior Prediction2026Can Dialects Be Steered Like Languages? Sparse Neurons and Distributed Directions in Arabic LLMs2026AI Wizards at EXIST 2026: Hierarchical Soft-Label Learning for Multimodal Sexism Identification in Memes2026PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails2026PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages2026UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation2026Token-Based Dual-view Fusion and Adaptation of Large Vision Models for Breast Cancer Classification2026Where to cut, how deep: BPE and Unigram-LM on chemistry SMILES2026AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation2026SPEAR: A Simulator for Photorealistic Embodied AI Research2026Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning2026MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models2026Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE2026Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing2026CausalDS: Benchmarking Causal Reasoning in Data-Science Agents2026KronQ: LLM Quantization via Kronecker-Factored Hessian2026MuScriptor: An Open Model for Multi-Instrument Music Transcription2026DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment2026Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents2026UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks2026Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models2026DrugGen 2: A disease-aware language model for enhancing drug discovery2026Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation2026Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading2026CLAP: Direct VLM-to-VLA Adaptation via Language-Action Grounding2026Phone Segmentation and Recognition through Phonological Activation Mapping2026A Sovereign, Open-Source Foundation Model for German and English2026ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Streams2026SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding2026Towards Autonomous and Auditable Medical Imaging Model Development2026SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL2026Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos2026MAGIC: Transition-Aware Generation of Navigable Multi-Scene Game Worlds with Large Language Models2026