All papers
Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.
Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model2026RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM2026Edge Cluster Expansion with Radial Rotary Attention for Interatomic Potentials2026Modernizing HEBO: a robust Bayesian optimization baseline for practical heteroskedastic and non-stationary problems2026GigaAM Multilingual: Foundation Model for Underrepresented Languages2026From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World2026Rethinking the Evaluation of Harness Evolution for Agents2026Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models2026Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation2026PalmClaw: A Native On-Device Agent Framework for Mobile Phones2026KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill2026Self-Improvements in Modern Agentic Systems: A Survey2026Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget2026Discrete Diffusion Models: A Unified Framework from Tokenization to Generation2026AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities2026Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Code2026RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination2026Chat2Scenic: An Iterative RAG-Based Framework for Scenario Generation in Autonomous Driving2026SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration2026xHC: Expanded Hyper-Connections2026LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget2026On-Policy Delta Distillation2026BadWAM: When World-Action Models Dream Right but Act Wrong2026MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators2026SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning2026Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models2026Cura 1T: Specialized Model for Agentic Healthcare2026SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction2026S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation2026Understanding Reasoning from Pretraining to Post-Training2026An Exam for Active Observers2026Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification2026Agent psychometrics: Task-level performance prediction in agentic coding benchmarks2026DepthART: Scaling Foundation Monocular Depth to Tiny Models2026Distilled Reinforcement Learning for LLM Post-training2026Uncovering Latent Reasoning Strategies in Language Models2026ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video2026WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting2026Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints2026Differentiable Logic Gate Networks for Low-Latency EEG Classification on Edge Devices2026SkillRouter: Skill Routing for LLM Agents at Scale2026Three-Body Scattering for Generative Modeling2026Patch Policy: Efficient Embodied Control via Dense Visual Representations2026Can Multimodal Large Language Models Understand OCT?2026EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World2026AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report2026EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration2026AutoIndex: Learning Representation Programs for Retrieval2026AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents2026Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training2026Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing2026ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU2026ISO: An RLVR-Native Optimization Stack2026Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning2026HPD-Parsing: Hierarchical Parallel Document Parsing2026Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing2026SLPO: Scaling Latent Reasoning via a Surrogate Policy2026DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations2026G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object Detection2026Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model2026SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD2026K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs2026Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking2026DataPrep-Bench: Benchmarking LLMs as Training Data Preparators2026LLMs Get Lost in Evolving User Intent2026ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders2026Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems2026Visual Contrastive Self-Distillation2026AREX: Towards a Recursively Self-Improving Agent for Deep Research2026FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills2026TriGlue: a Biology-Inspired Generative Model for Generating Molecular Glue-Induced Ternary Complex2026dRAE: Representation Autoencoder with Hyper-Spherical Codes2026IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation2026SceneActBench: Can Agents Act on the 3D Scenes They See?2026Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning2026Progress Reward Modeling for Robotic Learning: A Comprehensive Survey2026Leveraging External Knowledge for Historical Document Restoration via Retrieval-Augmented Large Language Models2026Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills2026Context-Aware Concept Distillation for Trustworthy Flood Prediction2026RareLens: Towards End-to-End Rare Disease Care via Aligning Divergent Large Language Model Reasoning2026Novel Claim or Déjà Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking2026From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement2026A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever2026Grading the Narrators: An Isnad-Rijal Framework for Claim-Level Provenance in Multi-Agent Knowledge Systems2026FilmBench: A Film-Grade Benchmark for Cinematic Video Generation2026The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation2026Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation2026ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding2026Bitcoin Price Direction Prediction via Regime-Aware Multi-Modal Fusion of Social Sentiment and Technical Features2026Chamaileon: Cross-Context Binder Design with Contextualized Modeling and Mixed Sampling2026DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification2026Kimi K3: Open Frontier Intelligence2026A New Role for Relevance: Guiding Corpus Interaction in Agentic Search2026Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory2026Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels2026Agent Retrieval Bench: Evaluating Repository Context Retrieval for Coding Agents2026Towards Robust Reinforcement Learning for Small-Scale Language Model Agents2026OPERA: Offline Policy-guided Expert Routing and Adaptation for Universal Biomedical Image Analysis2026CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition2026CAST: Game Solvers as Turn-Level Teachers for LLM Agents2026MemSFT: Mitigating Alignment Tax with an External Parametric Memory2026MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities2026Parallel Decoding Distillation for Fast Image and Video Generation2026$π\mathbf{R}^2$: Reactive Real-time Flow Policies2026Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model2026VisualPatchWorld: Code World Models as Latent Structured Representations for Planning2026Temporal-Distance JEPA: Plan-Aware Representation Learning for Latent World Model Predictive Control2026Pass the Baton: Trajectory-Relayed On-Policy Distillation2026StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents2026Voice Memory for Agentic Speech Recognition2026Constitutional Midtraining: Content Presence Drives Alignment Gains2026See2Think: Do Multimodal Models Really Use Intermediate Visual States?2026SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution2026SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response2026OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding2026Weak-to-Strong On-Policy Distillation2026Metis: Memory Foundation Model2026Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes2026DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search2026Mental World Modeling2026Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation2026ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow2026Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale2026EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents2026Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations2026Harness-G: A Graph-Structured Harness for Search Agents2026Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation2026Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory2026Can Large Language Models Execute Parent Orders?2026Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning2026Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering2026VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation2026AISPA: User-Centric System Prompt Auditing for Large Language Model Applications2026AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis2026Multi-Head Attention Residuals2026VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System2026SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them2026ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow2026ReToken: One Token to Improve Vision-Language Models for Visual Retrieval2026MemHarness: Memory Is Reconstructed, Not Replayed2026MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations2026RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems2026WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning2026AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?2026Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts2026Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI2026StreamTalk: Streaming Co-Speech Gesture Generation with Key-Pose Anchoring2026DAPD: Dual-Anchored Policy Distillation2026Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs2026Learning Compositional Meta-Routing for Agentic Workflows: An Executable Benchmark2026Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors20263DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering2026Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval2026Progressive Agent Skill Generation via Reinforcement Learning2026GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning2026StoryScope: Investigating idiosyncrasies in AI fiction2026OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution2026ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step2026SWE-Touch: Benchmarking Coding Agents When Users Touch the Code2026UEmbed: Unified Sparse and Dense Multimodal Embeddings2026AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling2026DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents2026PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs2026GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience2026ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads2026SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs2026ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?2026Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling2026Search, Inspect, Fetch: Exploiting Structure-Aware Boolean Retrieval for Deep-Search Agents2026Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning2026CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning2026PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents2026TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning2026Quo Vadis, World Modeling?2026Multi-Task Multi-Frame Visual Piano Transcription2026When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation2026GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks2026ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning2026OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents2026NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap2026MameLoshnLM: Yiddish Language Model and Evaluation Benchmark2026MatrAIx: Simulating the World with 8.3 Billion Persona Agents2026Small Foundation Models of Human Cognition and Behaviour2026Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay2026AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning2026From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models2026EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning2026The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images2026Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection2026Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory2026$A^2E$ : An End-to-End Agent Auditing Engine2026KVAE: Family of Tokenizers for Multimodal Generative Models2026On-Policy Self-Distillation without Any Supervision2026MirrorWorld: Taming Video Diffusion Models for Mirror Reflection Generation2026LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers2026Scaling Inherently Interpretable Language Models2026Mendel Gödel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution2026Vision-Language Grounding as Bidirectional Concept Correspondence2026Thought-Level Beam Search for Reasoning2026Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives2026