All papers
Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.
SuperLocalMemory 4.0: The Governed Memory Operating System for AI Agents2026Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution2026Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses2026Business Arena: Benchmarking LLM Agents in a Realistic Marketplace2026360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents2026How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review2026A Hybrid Nested Harness for Decoupling Structure and Parameters in LLM-Driven Optimization2026Parameter Exploration for RLVR via Variational Learning2026RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance2026BDH-CQ: In-Context Learning with Recurrent Latent Reasoning2026Evo-Bench: Can Language Models Improve Agent Harness?2026Intent Speaks Louder: Controllable User Simulation Beyond Response Imitation2026Motif 3: Technical Report2026Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains2026Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers2026Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference2026InSight-doc: Agentic Visual Perception for Long-Document Understanding2026ReRound: Reconstructive Rounding to Resolve Midpoint Ambiguity in Calibration-Free LLM Quantization2026Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design2026DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?2026DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation2026Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence2026Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation2026VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?2026SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models2026Persistent Recursive Worlds Enable Autonomous Software Evolution2026Dion3: Full-Stack Orthogonal Updates2026MBA: Multimodal Benchmark and Agents for Real-World Business Ideation2026Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence2026ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents2026Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill2026Claim-Level Reliability Assessment for Efficient Test-Time Reasoning2026Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus2026AVA-Encoder: Towards Agent-Native Video Representation Learning2026From Atomic Evidence to Logical Composition: Structured Compositional Reasoning over Compound Answer Options2026The Embedder's Dilemma: LLMs Are Better, but at What Cost?2026LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation2026Latent On-Policy Self-Distillation2026Intern-S2-Preview: Scientific Agentic Foundation Model2026DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data2026OmniScientist: An Omni-Modal Omni-Discipline AI Scientist2026AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design2026CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation2026Gaze Target Estimation Anywhere with Concepts2026Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control2026Scaling Automatic Research Agents via World Models2026CW-BASS v2: Saturation-Aware Pseudo-Label Selection for Semi-Supervised Segmentation under Foundation-Model Teachers2026Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence2026TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents2026Who Speaks Matters: Authority-Aware Multi-View RAG over Italian Parliamentary Proceedings2026HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark2026Agentic Transaction: Towards ACID-Compliant Agent Systems2026Nanbeige4.2-3B on Apple Silicon: Fixing Deployment Bugs and Decreasing Looped Transformer Memory Overhead2026ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models2026Demystifying Agent Skills: Why They Work-Until They Don't2026A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images2026MathForm: Scaling Mathematical Autoformalization with Knowledge Retrieval and Verification-Guided Refinement2026SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning2026Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning2026Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination2026Knowing When to Stop: Bayesian Optimal Stopping for LLM Evaluations2026PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments2026Marionette: Predicting World States, Rendering Geometry, Painting Appearance2026HarmProfile: Characterizing Harmful Distributions in Frontier LLMs2026Gathered, Not Admitted: How Attention Brings a Latent Variable into Verbalizable Form2026StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling2026VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?2026Dynamic Multi-Byte Prediction With Hierarchical Language Models2026TinyCast: Probabilistic Zero-Shot Forecasting with Computed Periodicity2026Bounded Agents: Delegation Security for Multi-Agent AI Systems2026UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations2026From Sequence to Structure: Relational Uncertainty Propagation for LLM Agents2026Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency2026Drive, Pack, Fly: The Travelling Thief Problem with Drone2026Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift2026Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation2026Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search2026Advancing Open and Reproducible Relational Learning: RelArena-$α$, TabPFN-Rel and RPI2026ClawGym II: Exploring Black-Box RL on Agent Harness2026VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding2026How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks2026$R^3$-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets2026TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation2026The Problem Is the Problem: Towards Scalable Mathematical Discovery2026Cross-Model Memory Transfer via Target-Side Reader Adaptation2026Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models2026Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL2026ASI-Bench: At the Dawn of Artificial Superintelligence2026PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX2026LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents2026SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation2026Agent Lightning v1.0: Towards Harnessed Agentic RL2026HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety2026AutoResearch: Insight In, Hallucination Out2026Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements2026Evaluating Music Context Preservation: A Multi-facet Framework for Music Editing Systems2025FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents2026Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models2026FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis2026SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents2026SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution2026Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation2026SPADE: Self-Play in Adaptive Synthetic Executable Environments2026Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability2026PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents2026EnvHarness: Awakening Static Worlds for Agent Learning2026Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection2026One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows2026FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving2026SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?2026MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use2026Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference2026VGI-Bench: Probing Visual Intelligence in Video Generation Models2026GOAG: Generative and Object-Agnostic Grasp Planner for Dexterous Robotic Manipulation2026Repo0: Design-Driven Zero-to-All Code Generation2026Listening Forward: Next Patch Embedding Prediction Enables Scalable Audio Learners2026Towards Quantifying Benchmark Optimization in ASR Models2026CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning2026Peer-Voted LLM-Agent Stress Tests Find Feed-Induced Lexical Convergence but No Reliable Matched-Exposure Advantage for Distributed Sources2026FlavourBench: Executable Culinary Reward Maps for Language Model Evaluation and Post-Training2026TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming2026A Modular Agent for Reliable and Auditable Spatial Relation Verification in CT Scans2026Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence2026SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation2026Length-Adaptive Decoding for Masked Diffusion Machine Translation2026ClawProBench: Trace-Aware Evaluation of AI Agents with Runtime Coverage and Frozen Workplace-Style Holdouts2026MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks2026AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces2026From Generation to Simulation: How Far Are World Models from Being True Simulators?2026Apodex 1.1: Scaling Agentic Intelligence for Complex Work2026Agent-G$^2$: Gaussian Guidance for Agentic Reinforcement Learning2026Towards a Densing Law for User Representation Learning at Billion-Scale Capacity2026Prime Agent: A Self-Improving RLM Harness2026SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?2026ReWorld: An Interactive World Model with Long-Horizon Memory2026Best Practice Critic Optimization2026RIBOSPAN: A Long-Context RNA Foundation Model for Versatile RNA Modeling2026The Laws of Context Allocation: Causal Measurement and Closed-Loop Orchestration in Generative Search2026GUI-Primitives: Diagnosing Spatial Reasoning Failures in Vision-Language GUI Grounding2026Industrial-Instruction: An End-to-End Framework for Building Instruction-Tuning and Benchmark Datasets from Industrial Technical Reports2026Better Retrieval, Worse Robustness: How Multi-hop RAG Amplifies Upstream ASR Errors2026CaRGo-T: Causal Reasoning Graph-of-Thought improves Multimodal Humor Comprehension2026CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild2026Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment2026MARS: Multi-Specialist LLM Relay System for Competitive Programming2026Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial Environments2026PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control2026On-policy Distillation with Verifiable Reward2026StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing2026Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses2026WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation2026WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report2026MemUse: Moving Memory Evaluation from Direct QA to Natural Integration in Long-Term Human-AI Conversation2026FrontierChallenge: Evaluating Scientific Workflow Completion2026DataKernelBench: Can LLMs Optimize Database Queries on GPUs?2026Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents2026CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval2026Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models2026LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale2026JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution2026Skill Issue: Are Skills Language-Invariant in LLMs?2026Prefix Sliding for efficient test-time scaling2026VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning2026GGSS: Geodesic-Gated Spherical Steering for Inference-Time Debiasing of Generative Vision-Language Models2026Code World Model: Coding Agent as World Brain2026V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning2026GameWAM: A World Action Model for Video Games2026J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data2026Generative Semantic Scene Completion2026Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $50902026TTPO: Test-Time Policy Optimization2026VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction2026LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics2026PAWBench: How Far Are We from Probabilistically Aligned World Modeling?2026RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests2026OpenStamp: A Watermark for Open-Source Language Models2026LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering2026Fast Weight Attention for Continual Learning2026ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL2026Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge2026Evaluating the Hidden Costs of Personalization in Large Language Models2026SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models2026Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space2026Dynamic Important Example Mining for Reinforcement Finetuning2026QCell: Recombining and Aligning Cell Queries for Overlapping Instance Segmentation2026Super Library Agent: Joint Generation and Maintenance of Multiple Applications Beyond the Single Codebase2026EvoGenUI-Bench: Evaluating LLMs as Multi-Turn Generative UI Assistants2026Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered2026Cross-lingual Functional Vectors for Emotion Detection in Large Language Models2026E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation2026Normalized Low-Rank Adaptation2026Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement2026Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents2026SpanCalib-VLM: Calibrated Hallucination Span Detection in Vision-Language Models2026On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability2026CogEvol: Towards Efficient and Reliable Learning Environment Generation2026PaperGym: Rubric-Centered Evolution for Research-Plan Generation2026LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation2026UI-Venus-2 Technical Report2026Safin-1: Safety from Within through Memory-Native State Evolution2026