0

All papers

Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.

WHALE: A Simple Recipe for Joint Harness-Weight Optimization2026Dr. Claw: An AI Scientist Workspace for Vibe Research2026EM^2Mem: Event-Centric Multimodal Memory for Large Language Models2026Let Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc Calibration2026GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent Environments2026Enoki: Efficient Multi-Level Hallucination Detection2026Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs2026Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall2026StudentSim: Training LLM-based Student Simulators2026Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence2026Recursive Criticality of AI Self-Improvement2026EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents2026Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching2026Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens2026A Common Measure of Communication for Speech Brain-Computer Interfaces2026Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills2026ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding2026EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction2026Unifying Conformal Language Tasks with In-Context Ensembles2026What Else Needs Fixing? Exploring Cost-Effective Test-Time Compute for Revision Propagation in Artifacts Generated Through Conversation2026Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs2026Editable Visual Design2026When Models Edit Too Much: On the Fidelity of Minimal Code Edits2026Last Translation Benchmark2026Causal Foundation Models2026H3-World: Turning Language Understanding into World Control2026FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience2026Unlocking Lossless Speedups in LLMs via Discrete Diffusion2026DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training2026Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems2026The Attention Triangle in Audio-Video Models2026ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation2026LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes2026One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing2026Iris: Climbing to the Search Frontier2026MaxKernel: Agentic Kernel Generation for TPUs2026$τ^τ$-Bench: An Environment for End-To-End, Realistic Agent Construction2026Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models2026Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models2026Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization2026RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?2026Privacy Failure in Split-LLM Training, The Returned Gradient Nullifies the Decoys2026BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference2026UniMate: One Unified Model to Animate Diverse Skeletons2026Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs2026WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data2026SQS: Bayesian DNN Compression through Sparse Quantized Sub-distributions2025What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets2026Diffs vs. Whole Files: An Empirical Comparison of Iterative Edit-Based and Direct Generation for Flutter/Dart Code Models2026Steering Geometry: Validating Human Value Geometry in LLM Steering Space2026OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution2026Reason Through the Latent! Making Latent Visual Reasoning Necessary2026SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem2026RelightFormer: Feed-forward Generative Transformer for Multiview Object Relighting2026Harnessing CLIP and DINO: An Uncertainty-Aware Cascaded Fusion Network for Generalizable Deepfake Image Detection2026Kalman Delta Networks: Uncertainty-aware Associative Memory2026A*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of LLM2026SchemeArena: Factorized Stress Testing of Scheming in LLM Agents2026SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents2026Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks2026Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR2026PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving2026Omni Interaction Agent Technical Report2026SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?2026SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models2026RenderFormer-V2: Neural Rendering with Heterogeneous Scene Primitives2026Cadence: Error-Bounded Lossy Compression of Demand Time Series with a Time-Series Foundation Model2026Train Smarter, Not Harder: Switching Signal-Guided Training in Active Learning2026MOLE: Detecting Insider Threats in AI Agents2026Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training2026Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation2026CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs2026Miles v0.1: Production-Level Post-Training2026Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation2026PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents2026Revisiting Complete Reasoning Traces for Post-Training2026NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness2026AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing2026ActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation2026Studying Image Tokenizers as Visual Languages in Unified Multimodal Models2026Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents2026RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems2026MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes2026Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States2026Show-Harness: Just a VLM Agent Can Play Robots2026$Φ$-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?2026Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs2026IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications2026NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction2026Think Before You Link: Rarity, Reasoning, and Retrieval in Multilingual Entity Linking2026Negative Self-Distillation: Learning to Reason by Avoiding Flaws2026

Back to Papers