All papers
Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL OptimizationarXiv 2026TideGS: Scalable Training of Over One Billion 3D Gaussian Splatting Primitives via Out-of-Core OptimizationarXiv 2026Pseudo-Simulation for Autonomous DrivingarXiv 2025FullStack Bench: Evaluating LLMs as Full Stack CodersarXiv 2024SpatialLM: Training Large Language Models for Structured Indoor ModelingarXiv 2025A Multiscale Visualization of Attention in the Transformer Modela-multiscale-visualization-of-attention-in-1I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion ModelsarXiv 2023DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable ConstraintsarXiv 2026Cradle: Empowering Foundation Agents Towards General Computer ControlarXiv 2024LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt CompressionarXiv 2024Foundations of Large Language ModelsarXiv 2025LLM Post-Training: A Deep Dive into Reasoning Large Language ModelsarXiv 2025OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and CleaningarXiv 2025Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language ModelsCVPR 2025 1Emerging Properties in Self-Supervised Vision TransformersICCV 2021 10ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language ModelsarXiv 2025Efficient Track AnythingICCV 2025Llama 2: Open Foundation and Fine-Tuned Chat ModelsarXiv 2023Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining DatasetarXiv 2024AndroidWorld: A Dynamic Benchmarking Environment for Autonomous AgentsarXiv 2024Moonshot: Towards Controllable Video Generation and Editing with Multimodal ConditionsarXiv 2024The Well: a Large-Scale Collection of Diverse Physics Simulations for Machine LearningarXiv 2024Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech ModelarXiv 2025VLM-R1: A Stable and Generalizable R1-style Large Vision-Language ModelarXiv 2025GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera ControlCVPR 2025 1Geometry-Aware Representation Denoising for Robust Multi-view 3D ReconstructionarXiv 2026U$^2$-Net: Going Deeper with Nested U-Structure for Salient Object DetectionarXiv 2020NNsight and NDIF: Democratizing Access to Open-Weight Foundation Model InternalsarXiv 2024StreamDiffusion: A Pipeline-level Solution for Real-time Interactive GenerationICCV 2025AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning datasetarXiv 2025A Survey of Text-to-SQL in the Era of LLMs: Where are we, and where are we going?arXiv 2024UniVLA: Learning to Act Anywhere with Task-centric Latent ActionsarXiv 2025Mask2Former for Video Instance SegmentationarXiv 2021Decompile-Bench: Million-Scale Binary-Source Function Pairs for
Real-World Binary DecompilationarXiv 2025CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language ModelsarXiv 2024AutoFigure-Edit: Generating Editable Scientific IllustrationarXiv 2026DeepEyes: Incentivizing "Thinking with Images" via Reinforcement LearningarXiv 2025ClawBench: Can AI Agents Complete Everyday Online Tasks?arXiv 2026Signals: Trajectory Sampling and Triage for Agentic InteractionsarXiv 2026SpatialBench: Is Your Spatial Foundation Model an All-Round Player?arXiv 2026Restructuring Vector Quantization with the Rotation TrickarXiv 2024How to Unleash the Power of Large Language Models for Few-shot Relation Extraction?arXiv 2023LightNER: A Lightweight Tuning Paradigm for Low-resource NER via Pluggable PromptingCOLING 2022 10VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic PlanningarXiv 2024ASAP: Aligning Simulation and Real-World Physics for Learning Agile Humanoid Whole-Body SkillsarXiv 2025InspireMusic: Integrating Super Resolution and Large Language Model for High-Fidelity Long-Form Music GenerationarXiv 2025MASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction PriorsCVPR 2025 1CNN Explainer: Learning Convolutional Neural Networks with Interactive VisualizationarXiv 2020UniTok: A Unified Tokenizer for Visual Generation and UnderstandingarXiv 2025Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and RankingarXiv 2026DreamLite: A Lightweight On-Device Unified Model for Image Generation and EditingarXiv 2026Learning Dynamics of LLM FinetuningarXiv 2024SkVM: Compiling Skills for Efficient Execution EverywherearXiv 2026MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M TokensarXiv 2026The Last Human-Written Paper: Agent-Native Research ArtifactsarXiv 2026AgenticPay: A Multi-Agent LLM Negotiation System for Buyer-Seller TransactionsarXiv 2026Generative Refinement Networks for Visual SynthesisarXiv 2026ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and WatchersarXiv 2026TimeGPT-1arXiv 2023Kimi-Audio Technical ReportarXiv 2025MiniRAG: Towards Extremely Simple Retrieval-Augmented GenerationarXiv 2025Token Reduction Should Go Beyond Efficiency in Generative Models -- From Vision, Language to MultimodalityarXiv 2025LLMs as Hackers: Autonomous Linux Privilege Escalation AttacksarXiv 20232D Gaussian Splatting for Geometrically Accurate Radiance FieldsarXiv 2024Harnessing the Power of LLMs in Practice: A Survey on ChatGPT and BeyondarXiv 2023Reinforcement Learning from Human FeedbackarXiv 2025RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement LearningarXiv 2025OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action ModelarXiv 2025Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation ControlarXiv 2025MedRAX: Medical Reasoning Agent for Chest X-rayarXiv 2025Generative agent-based modeling with actions grounded in physical, social, or digital space using ConcordiaarXiv 2023Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?arXiv 2025VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool InvocationarXiv 2026Reasoning Beyond Language: A Comprehensive Survey on Latent Chain-of-Thought ReasoningarXiv 2025AgentCPM-GUI: Building Mobile-Use Agents with Reinforcement Fine-TuningarXiv 2025RIFE: Real-Time Intermediate Flow Estimation for Video Frame InterpolationarXiv 2020Show-o2: Improved Native Unified Multimodal ModelsarXiv 2025A foundation model for atomistic materials chemistryarXiv 2023Post-Trained MoE Can Skip Half Experts via Self-DistillationarXiv 2026BlenderLLM: Training Large Language Models for Computer-Aided Design with Self-improvementarXiv 2024SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to EvolutionarXiv 2026Gaussian Grouping: Segment and Edit Anything in 3D ScenesarXiv 2023OCR-free Document Understanding TransformerarXiv 2021Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech GenerationarXiv 2024Metis: A Foundation Speech Generation Model with Masked Generative Pre-trainingarXiv 2025HUGSIM: A Real-Time, Photo-Realistic and Closed-Loop Simulator for Autonomous DrivingarXiv 2024FlashRAG: A Modular Toolkit for Efficient Retrieval-Augmented Generation ResearcharXiv 2024Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical PerspectivesICCV 2023 1Token Merging: Your ViT But FasterarXiv 2022Captum: A unified and generic model interpretability library for PyTorcharXiv 2020End-to-end Autonomous Driving: Challenges and FrontiersarXiv 2023Learning Bipedal Walking On Planned Footsteps For Humanoid RobotsarXiv 2022D-FINE: Redefine Regression Task in DETRs as Fine-grained Distribution RefinementarXiv 2024HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesishifi-gan-generative-adversarial-networks-for-1Nougat: Neural Optical Understanding for Academic DocumentsarXiv 2023PP-DocLayout: A Unified Document Layout Detection Model to Accelerate Large-Scale Data ConstructionarXiv 2025DSVT: Dynamic Sparse Voxel Transformer with Rotated SetsCVPR 2023 1MambaVision: A Hybrid Mamba-Transformer Vision BackboneCVPR 2025 1OASIS: Open Agent Social Interaction Simulations with One Million AgentsarXiv 2024High-Resolution Image Synthesis and Semantic Manipulation with Conditional GANshigh-resolution-image-synthesis-and-semantic-1EgoLife: Towards Egocentric Life AssistantCVPR 2025 1Comet: Fine-grained Computation-communication Overlapping for Mixture-of-ExpertsarXiv 2025FrameSkip: Learning from Fewer but More Informative Frames in VLA TrainingarXiv 2026EVA-Bench: A New End-to-end Framework for Evaluating Voice AgentsarXiv 2026Qwen-Image-VAE-2.0 Technical ReportarXiv 2026AI CFD Scientist: Toward Open-Ended Computational Fluid Dynamics Discovery with Physics-Aware AI AgentsarXiv 2026KernelBench-X: A Comprehensive Benchmark for Evaluating LLM-Generated GPU KernelsarXiv 2026Grounding Image Matching in 3D with MASt3RarXiv 2024WebThinker: Empowering Large Reasoning Models with Deep Research CapabilityarXiv 2025VILA: On Pre-training for Visual Language ModelsCVPR 2024 1Seed-Coder: Let the Code Model Curate Data for ItselfarXiv 2025SuperAnimal pretrained pose estimation models for behavioral analysisarXiv 2022MineDojo: Building Open-Ended Embodied Agents with Internet-Scale KnowledgearXiv 2022EASYTOOL: Enhancing LLM-based Agents with Concise Tool InstructionarXiv 2024MiniCPM4: Ultra-Efficient LLMs on End DevicesarXiv 2025Trinity-RFT: A General-Purpose and Unified Framework for Reinforcement Fine-Tuning of Large Language ModelsarXiv 2025Old Photo Restoration via Deep Latent Space TranslationarXiv 2020Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation ModelsarXiv 2024ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level CodearXiv 2023DreamGen: Unlocking Generalization in Robot Learning through Video World ModelsarXiv 2025Linformer: Self-Attention with Linear ComplexityarXiv 2020Scaling A Simple Approach to Zero-Shot Speech RecognitionarXiv 2024Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and LanguagearXiv 2022DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detectiondino-detr-with-improved-denoising-anchorFew-shot Learning with Multilingual Language ModelsarXiv 2021Cross-lingual Retrieval for Iterative Self-Supervised TrainingNeurIPS 2020 12XLS-R: Self-supervised Cross-lingual Speech Representation Learning at ScalearXiv 2021Generative Spoken Language Modeling from Raw AudioarXiv 2021Multilingual Denoising Pre-training for Neural Machine TranslationarXiv 2020data2vec: A General Framework for Self-supervised Learning in Speech, Vision and LanguagePreprint 2022 1Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-TuningarXiv 2025Deformable DETR: Deformable Transformers for End-to-End Object Detectiondeformable-detr-deformable-transformers-forRULER: What's the Real Context Size of Your Long-Context Language Models?arXiv 2024HART: Efficient Visual Generation with Hybrid Autoregressive TransformerarXiv 2024RouteLLM: Learning to Route LLMs with Preference DataarXiv 2024On the Measure of IntelligencearXiv 2019MMaDA: Multimodal Large Diffusion Language ModelsarXiv 2025Habitat 2.0: Training Home Assistants to Rearrange their HabitatNeurIPS 2021 12Efficient Few-Shot Learning Without PromptsarXiv 2022High Fidelity Neural Audio CompressionarXiv 2022TimeMixer: Decomposable Multiscale Mixing for Time Series Forecastingtimemixer-decomposable-multiscale-mixing-forA ConvNet for the 2020sCVPR 2022 1Torchreid: A Library for Deep Learning Person Re-Identification in PytorcharXiv 2019PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion TransformersarXiv 2025MedSAM2: Segment Anything in 3D Medical Images and VideosarXiv 2025How NeRFs and 3D Gaussian Splatting are Reshaping SLAM: a SurveyarXiv 2024A Survey of Large Language ModelsarXiv 2023InternVideo2.5: Empowering Video MLLMs with Long and Rich Context ModelingarXiv 2025Robust High-Resolution Video Matting with Temporal GuidancearXiv 2021Generating Synergistic Formulaic Alpha Collections via Reinforcement LearningarXiv 2023Simple Online and Realtime Tracking with a Deep Association MetricarXiv 2017InstantID: Zero-shot Identity-Preserving Generation in SecondsarXiv 2024TAPNext: Tracking Any Point (TAP) as Next Token PredictionICCV 2025Zero-1-to-3: Zero-shot One Image to 3D ObjectICCV 2023 1Is Sora a World Simulator? A Comprehensive Survey on General World Models and BeyondarXiv 2024BEVFusion: Multi-Task Multi-Sensor Fusion with Unified Bird's-Eye View RepresentationarXiv 2022MiMo-V2-Flash Technical ReportarXiv 2026Aligning benchmark datasets for table structure recognitionarXiv 2023Live Music Diffusion Models: Efficient Fine-Tuning and Post-Training of Interactive Diffusion Music GeneratorsarXiv 2026MemSkill: Learning and Evolving Memory Skills for Self-Evolving AgentsarXiv 2026D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion ModelsarXiv 2026PP-MobileSeg: Explore the Fast and Accurate Semantic Segmentation Model on Mobile DevicesarXiv 2023Morphological Prototyping for Unsupervised Slide Representation Learning in Computational PathologyCVPR 2024 1SuperGlue: Learning Feature Matching with Graph Neural Networkssuperglue-learning-feature-matching-with-1speechocean762: An Open-Source Non-native English Speech Corpus For Pronunciation AssessmentarXiv 2021TED-LIUM 3: twice as much data and corpus repartition for experiments on speaker adaptationarXiv 2018Caffe: Convolutional Architecture for Fast Feature EmbeddingarXiv 2014Masked Feature Prediction for Self-Supervised Visual Pre-TrainingCVPR 2022 1Diffusion Forcing: Next-token Prediction Meets Full-Sequence DiffusionarXiv 2024FastText.zip: Compressing text classification modelsarXiv 2016Datasets for Large Language Models: A Comprehensive SurveyarXiv 2024Segment Anything in Medical Images and Videos: Benchmark and DeploymentarXiv 2024OmniSVG: A Unified Scalable Vector Graphics Generation ModelarXiv 2025Diffusion for World Modeling: Visual Details Matter in AtariarXiv 2024GenericAgent: A Token-Efficient Self-Evolving LLM Agent via Contextual Information Density Maximization (V1.0)arXiv 2026KAN: Kolmogorov-Arnold NetworksarXiv 2024DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsarXiv 2024Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language ModelsarXiv 2025Dolphin: A Large-Scale Automatic Speech Recognition Model for Eastern LanguagesarXiv 2025VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the WildarXiv 2024Sharp Monocular View Synthesis in Less Than a SecondarXiv 2025Fast Text-to-Audio Generation with Adversarial Post-TrainingarXiv 2025LightGlue: Local Feature Matching at Light SpeedICCV 2023 1Flux Already Knows -- Activating Subject-Driven Image Generation without TrainingarXiv 2025EmbodiedGen: Towards a Generative 3D World Engine for Embodied IntelligencearXiv 2025Revisiting Image Pyramid Structure for High Resolution Salient Object DetectionarXiv 2022MAISI: Medical AI for Synthetic ImagingarXiv 2024FILM: Frame Interpolation for Large MotionarXiv 2022MambaIRv2: Attentive State Space RestorationCVPR 2025 1RAFT: Recurrent All-Pairs Field Transforms for Optical FlowECCV 2020 8Infinite Photorealistic Worlds using Procedural Generationinfinite-photorealistic-worlds-usingDeep Patch Visual SLAMarXiv 2024DeepResearch Bench: A Comprehensive Benchmark for Deep Research AgentsarXiv 2025Real-Time High-Resolution Background MattingCVPR 2021 1LocAgent: Graph-Guided LLM Agents for Code LocalizationarXiv 2025Towards Total Recall in Industrial Anomaly DetectionCVPR 2022 1Cosmos-Reason1: From Physical Common Sense To Embodied ReasoningarXiv 2025Scaffold-GS: Structured 3D Gaussians for View-Adaptive RenderingCVPR 2024 1Sekai: A Video Dataset towards World ExplorationarXiv 2025Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion TransformerCVPR 2025 1