All papers
Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.
AA-Omniscience: A Benchmark for Long-Tail Factual KnowledgepreprintAider's Polyglot Coding BenchmarkblogAIME as an LLM Evaluation BenchmarkblogThe Arcade Learning Environment: An Evaluation Platform for General AgentsJAIRALFWorld: Aligning Text and Embodied Environments for Interactive LearningICLRLength-Controlled AlpacaEval: A Simple Way to Debias Automatic EvaluatorsCOLMAPEX: An Expert-Authored Benchmark for Real-World Expert WorkflowspreprintFrom Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder PipelinepreprintAutoEnv: Automated Environments for Measuring Cross-Environment Agent LearningpreprintAya Model: An Instruction Finetuned Open-Access Multilingual Language ModelACLChallenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve ThemACLBeyond the Imitation Game: Quantifying and extrapolating the capabilities of language modelsTMLRBrowseComp: A Simple Yet Challenging Benchmark for Browsing AgentspreprintChatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceICMLEvaluating Large Language Models Trained on CodepreprintDeepDive: Reinforcement Learning Environments for Deep Research AgentsblogDeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learningpreprintdistilabel: AI Feedback for Building High-Quality DatasetsblogDeepMind Control SuitepreprintFinRL: A Deep Reinforcement Learning Library for Automated Stock Trading in Quantitative FinancepreprintFrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AIpreprintGAIA: A Benchmark for General AI AssistantsICLRGDPval: Evaluating AI Model Performance on Real-World Economically Valuable TaskspreprintGPQA: A Graduate-Level Google-Proof Q&A BenchmarkCOLMTraining Verifiers to Solve Math Word ProblemspreprintHolistic Evaluation of Language ModelsTMLRHelpSteer2: Open-source Dataset for Training Top-Performing Reward ModelsNeurIPSHermes 4 Technical ReportblogTraining a Helpful and Harmless Assistant with Reinforcement Learning from Human FeedbackpreprintHumanity's Last ExampreprintEvaluating Large Language Models Trained on CodepreprintTraining language models to follow instructions with human feedbackNeurIPSJupyter Agents: Building and Evaluating Agents That Use NotebooksblogLeRobot EnvHub: Standardized Environments for Robot LearningblogLiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for CodeNeurIPSMagpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with NothingpreprintMeasuring Mathematical Problem Solving With the MATH DatasetNeurIPSProgram Synthesis with Large Language Modelspreprintmini-SWE-agent: A Minimal Reference Agent for SWE-benchblogMeasuring Massive Multitask Language UnderstandingICLRMMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding BenchmarkNeurIPSJudging LLM-as-a-Judge with MT-Bench and Chatbot ArenaNeurIPSNeMo-Gym: NVIDIA's Framework for LLM Reinforcement Learning EnvironmentsblogNeedle In A Haystack - Pressure Testing LLMsblogNuminaMath: The Largest Public Dataset in AI4Maths with 860k Pairs of Competition Math Problems and SolutionsblogOpenEnv: An Open Standard for Agent EnvironmentsrfcOpenSpiel: A Framework for Reinforcement Learning in GamespreprintOpen Thoughts: Curating Reasoning Datasets for Open-Source R1 ReplicationsblogOSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer EnvironmentsNeurIPSOSWorld-Verified: A Cleaner, More Reliable Computer-Use BenchmarkblogEvaluating Large Language Models Trained on Code (pass@k formulation)preprintLet's Verify Step by StepICLRLet's Verify Step by StepICLRQwen3 Technical ReportpreprintRewardBench 2: Advancing Reward Model Evaluationpreprints1: Simple Test-Time ScalingpreprintBeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference DatasetNeurIPSStarling-7B: Improving Helpfulness and Harmlessness with RLAIFICMLSUMO-RL: Traffic Signal Control via Reinforcement LearningblogSWE-bench: Can Language Models Resolve Real-World GitHub Issues?ICLRSWE-Gym: An Open Environment for Training Software Engineering Agents and VerifierspreprintSWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?preprintτ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World DomainspreprintIntroducing Terminal-BenchblogTerminal-Bench: A Benchmark for Real-World Terminal-Based AgentsblogTextArena: Multi-Agent Text-Based Games for LLM EvaluationpreprintMeasuring AI Ability to Complete Long TaskspreprintTruthfulQA: Measuring How Models Mimic Human FalsehoodsACLTulu 3: Pushing Frontiers in Open Language Model Post-TrainingpreprintEnhancing Chat Language Models by Scaling High-quality Instructional ConversationsEMNLPUltraFeedback: Boosting Language Models with High-quality FeedbackICMLIntroducing the Environments Hubblogverifiers: Reinforcement Learning with LLMs in Verifiable EnvironmentsblogVicuna: An Open-Source Chatbot Impressing GPT-4 with 90% ChatGPT QualityblogVisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web TasksACLWebArena: A Realistic Web Environment for Building Autonomous AgentsICLRWildChat: 1M ChatGPT Interaction Logs in the WildICLRWizardLM: Empowering Large Language Models to Follow Complex InstructionsICLRWorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?ICMLRCT: Random Consistency Training for Semi-supervised Sound Event DetectionarXiv 2021Distributional Offline Policy Evaluation with Predictive Error GuaranteesarXiv 2023Graph Neural Prompting with Large Language ModelsarXiv 2023SentinelLMs: Encrypted Input Adaptation and Fine-tuning of Language Models for Private and Secure InferencearXiv 2023Addressing Negative Transfer in Diffusion Modelsaddressing-negative-transfer-in-diffusionDistilled Feature Fields Enable Few-Shot Language-Guided ManipulationarXiv 2023Hybrid Internal Model: Learning Agile Legged Locomotion with Simulated Robot ResponsearXiv 2023Large Language Models Are Not Strong Abstract ReasonersarXiv 2023TIAM -- A Metric for Evaluating Alignment in Text-to-Image GenerationarXiv 2023PLLaMa: An Open-source Large Language Model for Plant SciencearXiv 2024Gaussian processes at the Helm(holtz): A more fluid model for ocean currentsarXiv 2023Spikformer V2: Join the High Accuracy Club on ImageNet with an SNN TicketarXiv 2024Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body TeleoperationarXiv 2024Training Diffusion Models with Reinforcement LearningarXiv 2023PoET: A generative model of protein families as sequences-of-sequencesNeurIPS 2023 11FITS: Modeling Time Series with $10k$ ParametersarXiv 2023Geo2SigMap: High-Fidelity RF Signal Mapping Using Geographic DatabasesarXiv 2023DeepSeek LLM: Scaling Open-Source Language Models with LongtermismarXiv 2024Improving Visual Grounding by Encouraging Consistent Gradient-based ExplanationsCVPR 2023 1Is Complexity Required for Neural Network Pruning? A Case Study on Global Magnitude PruningarXiv 2022Improved Representation of Asymmetrical Distances with Interval Quasimetric EmbeddingsarXiv 2022Impossibility Theorems for Feature AttributionarXiv 2022LEXTREME: A Multi-Lingual and Multi-Task Benchmark for the Legal DomainarXiv 2023Evaluating Self-Supervised Learning via Risk DecompositionarXiv 2023SkillOpt: Executive Strategy for Self-Evolving Agent SkillsarXiv 2026stable-worldmodel-v1: Reproducible World Modeling Research and EvaluationarXiv 2026MOSS-TTS Technical ReportarXiv 2026MMSkills: Towards Multimodal Skills for General Visual AgentsarXiv 2026Docling: An Efficient Open-Source Toolkit for AI-driven Document ConversionarXiv 2025Docling Technical ReportarXiv 2024Agent Lightning: Train ANY AI Agents with Reinforcement LearningarXiv 2025MinerU2.5: A Decoupled Vision-Language Model for Efficient
High-Resolution Document ParsingarXiv 2025GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation ModelsarXiv 2025PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B
Ultra-Compact Vision-Language ModelarXiv 2025Self-Supervised Prompt OptimizationarXiv 2025FlashVSR: Towards Real-Time Diffusion-Based Streaming Video
Super-ResolutionarXiv 2025Single-stream Policy OptimizationarXiv 2025LightRAG: Simple and Fast Retrieval-Augmented GenerationarXiv 2024SimWorld: An Open-ended Realistic Simulator for Autonomous Agents in Physical and Social WorldsarXiv 2025Large Language Models for Cyber Security: A Systematic Literature ReviewarXiv 2024Correcting Negative Bias in Large Language Models through Negative Attention Score AlignmentarXiv 2024ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM AgentsarXiv 2026VibeVoice Technical ReportarXiv 2025MemOS: A Memory OS for AI SystemarXiv 2025The Unreasonable Effectiveness of Scaling Agents for Computer UsearXiv 2025Easy Dataset: A Unified and Extensible Framework for Synthesizing LLM Fine-Tuning Data from Unstructured DocumentsarXiv 2025IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech SystemarXiv 2025I-BERT: Integer-only BERT QuantizationarXiv 2021Improving the Domain Adaptation of Retrieval Augmented Generation (RAG) Models for Open Domain Question AnsweringarXiv 2022Fine-tune the Entire RAG Architecture (including DPR retriever) for Question-Answeringfine-tune-the-entire-rag-architecture-1Pre-trained Summarization DistillationarXiv 2020Improve Transformer Models with Better Relative Position EmbeddingsFindings of the Association for Computational Linguistics 2020SqueezeBERT: What can computer vision teach NLP about efficient neural networks?EMNLP (sustainlp) 2020 11Movement Pruning: Adaptive Sparsity by Fine-TuningNeurIPS 2020 12MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent ResearcharXiv 2026NanoResearch: Co-Evolving Skills, Memory, and Policy for Personalized Research AutomationarXiv 2026Robust Speech Recognition via Large-Scale Weak SupervisionPreprint 2022 9Qwen3-TTS Technical ReportarXiv 2026OmniFlatten: An End-to-end GPT Model for Seamless Voice ConversationarXiv 2024Fara-7B: An Efficient Agentic Model for Computer UsearXiv 2025Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line InterfacesarXiv 2026PyTorch FSDP: Experiences on Scaling Fully Sharded Data ParallelarXiv 2023Magentic-One: A Generalist Multi-Agent System for Solving Complex TasksarXiv 2024QuantaAlpha: An Evolutionary Framework for LLM-Driven Alpha MiningarXiv 2026HRM-Text: Efficient Pretraining Beyond ScalingarXiv 2026Lance: Unified Multimodal Modeling by Multi-Task SynergyarXiv 2026CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-trainingarXiv 2025EverMemOS: A Self-Organizing Memory Operating System for Structured Long-Horizon ReasoningarXiv 2026Zep: A Temporal Knowledge Graph Architecture for Agent MemoryarXiv 2025TriSplat: Simulation-Ready Feed-Forward 3D Scene ReconstructionarXiv 2026SeACo-Paraformer: A Non-Autoregressive ASR System with Flexible and Effective Hotword Customization AbilityarXiv 2023Fish Audio S2 Technical ReportarXiv 2026SmolVLA: A Vision-Language-Action Model for Affordable and Efficient RoboticsarXiv 2025OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache QuantizationarXiv 2026StableAvatar: Infinite-Length Audio-Driven Avatar Video GenerationarXiv 2025Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal
Generation and UnderstandingarXiv 2025Agentic Reinforced Policy OptimizationarXiv 2025DINOv3arXiv 2025SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated ReasoningarXiv 2025DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research SynthesisarXiv 2025Reinforced Visual Perception with ToolsarXiv 2025DeepAgent: A General Reasoning Agent with Scalable ToolsetsarXiv 2025AutoTriton: Automatic Triton Programming with Reinforcement Learning in LLMsarXiv 2025Real-Time Object Detection Meets DINOv3arXiv 2025A Survey of Scientific Large Language Models: From Data Foundations to Agent FrontiersarXiv 2025HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or PixelsarXiv 2025SWE-SQL: Illuminating LLM Pathways to Solve User SQL Issues in Real-World ApplicationsarXiv 2025SimpleVLA-RL: Scaling VLA Training via Reinforcement LearningarXiv 2025Thought Anchors: Which LLM Reasoning Steps Matter?arXiv 2025Qwen-Image Technical ReportarXiv 2025VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool UsearXiv 2025VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe ZooarXiv 2025Code2Video: A Code-centric Paradigm for Educational Video GenerationarXiv 2025Self-Rewarding Vision-Language Model via Reasoning DecompositionarXiv 2025MiDashengLM: Efficient Audio Understanding with General Audio CaptionsarXiv 2025Paper2Video: Automatic Video Generation from Scientific PapersarXiv 2025Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial
RepresentationsarXiv 2025Qwen3-Omni Technical ReportarXiv 2025Detect Anything via Next Point PredictionarXiv 2025VLA-0: Building State-of-the-Art VLAs with Zero ModificationarXiv 2025Is Diversity All You Need for Scalable Robotic Manipulation?arXiv 2025ToolUniverse: An open platform for democratizing AI scientistsarXiv 2025Stable Video Infinity: Infinite-Length Video Generation with Error
RecyclingarXiv 2025Matrix-3D: Omnidirectional Explorable 3D World GenerationarXiv 2025Scaling RL to Long VideosarXiv 2025Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term MemoryarXiv 2025PersonaLive! Expressive Portrait Image Animation for Live StreamingarXiv 2025OpenHands: An Open Platform for AI Software Developers as Generalist AgentsarXiv 2024Structured 3D Latents for Scalable and Versatile 3D GenerationCVPR 2025 1DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language ModelsarXiv 2026MolmoAct2: Action Reasoning Models for Real-world DeploymentarXiv 2026Very Large-Scale Multi-Agent Simulation in AgentScopearXiv 2024Attention Mesh: High-fidelity Face Mesh Prediction in Real-timearXiv 2020FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precisionarXiv 2024Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any ResolutionarXiv 2024Qwen2.5-VL Technical ReportarXiv 2025SPEED-Bench: A Unified and Diverse Benchmark for Speculative DecodingarXiv 2026EvoScientist: Towards Multi-Agent Evolving AI Scientists for End-to-End Scientific DiscoveryarXiv 2026Adam's Law: Textual Frequency Law on Large Language ModelsarXiv 2026SAGE: Scalable Agentic 3D Scene Generation for Embodied AIarXiv 20263D Gaussian Splatting for Real-Time Radiance Field RenderingarXiv 2023