All papers
Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.
SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion RefinementarXiv 2025Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation GenerationarXiv 2026DINO-X: A Unified Vision Model for Open-World Object Detection and UnderstandingarXiv 2024RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPOarXiv 20264D Gaussian Splatting for Real-Time Dynamic Scene RenderingCVPR 2024 1Describe Anything: Detailed Localized Image and Video CaptioningICCV 2025WebVoyager: Building an End-to-End Web Agent with Large Multimodal ModelsarXiv 2024SSL4EO-L: Datasets and Foundation Models for Landsat Imageryssl4eo-l-datasets-and-foundation-models-forLangSplat: 3D Language Gaussian SplattingCVPR 2024 1JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching AgentarXiv 2025Transformer in TransformerNeurIPS 2021 12Self-Supervised Learning from Images with a Joint-Embedding Predictive ArchitectureCVPR 2023 1ImageBind: One Embedding Space To Bind Them AllCVPR 2023 1DDT: Decoupled Diffusion Transformerddt-decoupled-diffusion-transformerYOLOX: Exceeding YOLO Series in 2021arXiv 2021ConvNeXt V2: Co-designing and Scaling ConvNets with Masked AutoencodersCVPR 2023 1Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining ResearcharXiv 2024MoMask: Generative Masked Modeling of 3D Human MotionsCVPR 2024 1Normalizing Flows are Capable Generative ModelsarXiv 2024Intent-based Prompt Calibration: Enhancing prompt optimization with synthetic boundary casesarXiv 2024RenderFormer: Transformer-based Neural Rendering of Triangle Meshes with Global IlluminationarXiv 2025ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM InferencearXiv 2025lmgame-Bench: How Good are LLMs at Playing Games?arXiv 2025TAPIR: Tracking Any Point with per-frame Initialization and temporal RefinementICCV 2023 1emotion2vec: Self-Supervised Pre-Training for Speech Emotion RepresentationarXiv 2023EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark ConditionsarXiv 2024Marigold: Affordable Adaptation of Diffusion-Based Image Generators for Image AnalysisarXiv 2025Zero-Reference Deep Curve Estimation for Low-Light Image Enhancementzero-reference-deep-curve-estimation-for-low-1ConceptAttention: Diffusion Transformers Learn Highly Interpretable FeaturesarXiv 2025When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrieval-Augmented GenerationarXiv 2025RLCard: A Toolkit for Reinforcement Learning in Card GamesarXiv 2019Resources for Brewing BEIR: Reproducible Reference Models and an Official LeaderboardarXiv 2023TransUNet: Transformers Make Strong Encoders for Medical Image SegmentationarXiv 2021Decision Transformer: Reinforcement Learning via Sequence ModelingNeurIPS 2021 12OpenThoughts: Data Recipes for Reasoning ModelsarXiv 2025Trajectory Prediction Meets Large Language Models: A SurveyarXiv 2025Block Diffusion: Interpolating Between Autoregressive and Diffusion Language ModelsarXiv 2025ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement LearningarXiv 2025Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual TokensarXiv 2025Human Motion Diffusion ModelarXiv 2022Learning Humanoid Standing-up Control across Diverse PosturesarXiv 2025OpenOOD v1.5: Enhanced Benchmark for Out-of-Distribution DetectionarXiv 2023RT-1: Robotics Transformer for Real-World Control at ScalearXiv 2022Taming Transformers for High-Resolution Image SynthesisCVPR 2021 1Generating and Imputing Tabular Data via Diffusion and Flow-based Gradient-Boosted TreesarXiv 2023JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language ModelsarXiv 2024Inpaint Anything: Segment Anything Meets Image InpaintingarXiv 2023PINNsFormer: A Transformer-Based Framework For Physics-Informed Neural NetworksarXiv 2023Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-shot Cross-dataset TransferarXiv 2019OminiControl2: Efficient Conditioning for Diffusion TransformersarXiv 2025Darwin Godel Machine: Open-Ended Evolution of Self-Improving AgentsarXiv 2025HVI: A New color space for Low-light Image EnhancementCVPR 2025 1Chronos: Learning the Language of Time SeriesarXiv 2024Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localizationgrad-cam-visual-explanations-from-deep-1Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined LevelsarXiv 2023Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric PerspectivesICCV 2025RepoAgent: An LLM-Powered Open-Source Framework for Repository-level Code Documentation GenerationarXiv 2024AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent BehaviorsarXiv 2023River: machine learning for streaming data in PythonarXiv 2020GRUtopia: Dream General Robots in a City at ScalearXiv 2024Gated Delta Networks: Improving Mamba2 with Delta RulearXiv 2024SWE-smith: Scaling Data for Software Engineering AgentsarXiv 2025Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentationmask-dino-towards-a-unified-transformer-basedWatermark Anything with Localized MessagesarXiv 2024BoT-SORT: Robust Associations Multi-Pedestrian TrackingarXiv 2022MM-Agent: LLM as Agents for Real-world Mathematical Modeling ProblemarXiv 2025Real Time Speech Enhancement in the Waveform DomainarXiv 2020NovelSeek: When Agent Becomes the Scientist -- Building Closed-Loop System from Hypothesis to VerificationarXiv 2025VMamba: Visual State Space ModelarXiv 2024RL4CO: an Extensive Reinforcement Learning for Combinatorial Optimization BenchmarkarXiv 2023A Survey on Inference Optimization Techniques for Mixture of Experts ModelsarXiv 2024MoBA: Mixture of Block Attention for Long-Context LLMsarXiv 2025Deep Compression Autoencoder for Efficient High-Resolution Diffusion ModelsarXiv 2024All-In-One Metrical And Functional Structure Analysis With Neighborhood Attentions on Demixed AudioarXiv 2023SpatialVLA: Exploring Spatial Representations for Visual-Language-Action ModelarXiv 2025TERA: Self-Supervised Learning of Transformer Encoder Representation for SpeecharXiv 2020When AI Meets Finance (StockAgent): Large Language Model-based Stock Trading in Simulated Real-world EnvironmentsarXiv 2024CityGaussianV2: Efficient and Geometrically Accurate Reconstruction for Large-Scale ScenesarXiv 2024CityGaussian: Real-time High-quality Large-Scale Scene Rendering with GaussiansarXiv 2024DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal UnderstandingarXiv 2024Medical Graph RAG: Towards Safe Medical Large Language Model via Graph Retrieval-Augmented GenerationarXiv 2024REPA-E: Unlocking VAE for End-to-End Tuning of Latent Diffusion TransformersarXiv 2025RLVR-World: Training World Models with Reinforcement LearningarXiv 2025MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio SynthesisCVPR 2025 1MT3: Multi-Task Multitrack Music Transcriptionmt3-multi-task-multitrack-music-transcriptionVisuoThink: Empowering LVLM Reasoning with Multimodal Tree SearcharXiv 2025ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous DrivingarXiv 2025InstructSAM: Segment Any Instance with Any InstructionsarXiv 2026Autoregressive Image Generation without Vector QuantizationarXiv 2024Avalanche: an End-to-End Library for Continual LearningarXiv 2021LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy ModelarXiv 2025Formalizing and Benchmarking Prompt Injection Attacks and DefensesarXiv 2023AdaWorld: Learning Adaptable World Models with Latent ActionsarXiv 2025LiveBench: A Challenging, Contamination-Limited LLM BenchmarkarXiv 2024TabICL: A Tabular Foundation Model for In-Context Learning on Large DataarXiv 2025LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationarXiv 2020PyCIL: A Python Toolbox for Class-Incremental LearningarXiv 2021Activating More Pixels in Image Super-Resolution TransformerCVPR 2023 1MatAnyone: Stable Video Matting with Consistent Memory PropagationCVPR 2025 1OOTDiffusion: Outfitting Fusion based Latent Diffusion for Controllable Virtual Try-onarXiv 2024STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-ResolutionICCV 2025Kubric: A scalable dataset generatorCVPR 2022 1DistilHuBERT: Speech Representation Learning by Layer-wise Distillation of Hidden-unit BERTarXiv 2021S3PRL-VC: Open-source Voice Conversion Framework with Self-supervised Speech RepresentationsarXiv 2021Defeating Prompt Injections by DesignarXiv 2025RemoteCLIP: A Vision Language Foundation Model for Remote SensingarXiv 2023LBM: Latent Bridge Matching for Fast Image-to-Image TranslationICCV 2025Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified FlowarXiv 2022DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot PlanningarXiv 2024SkillClaw: Let Skills Evolve Collectively with Agentic EvolverarXiv 2026History-Guided Video DiffusionarXiv 2025HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple CharactersarXiv 2025Vision Language Models in Autonomous Driving: A Survey and OutlookarXiv 2023DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world EnvironmentsarXiv 2025Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionarXiv 2023LSNet: See Large, Focus SmallCVPR 2025 1MiDaS v3.1 -- A Model Zoo for Robust Monocular Relative Depth EstimationarXiv 2023Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding HeadsarXiv 2024VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement LearningarXiv 2025Navigation World ModelsCVPR 2025 1OpenUnlearning: Accelerating LLM Unlearning via Unified Benchmarking of Methods and MetricsarXiv 2025Visual Agentic Reinforcement Fine-TuningarXiv 2025Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generationis-your-code-generated-by-chatgpt-reallyFrom Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language ModelsarXiv 2026rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified DatasetarXiv 2025Video-LLaVA: Learning United Visual Representation by Alignment Before Projectionvideo-llava-learning-united-visualChatlaw: A Multi-Agent Legal Assistant based on a Role-Aligned Mixture-of-Experts ArchitecturearXiv 2023Beyond Outlining: Heterogeneous Recursive Planning for Adaptive Long-form Writing with Language ModelsarXiv 2025An Improved RaftStereo Trained with A Mixed Dataset for the Robust Vision Challenge 2022arXiv 2022Spherical Channels for Modeling Atomic InteractionsarXiv 2022GS-LIVO: Real-Time LiDAR, Inertial, and Visual Multi-sensor Fused Odometry with Gaussian MappingarXiv 2025ERNIE-Doc: A Retrospective Long-Document Modeling TransformerACL 2021 5ERNIE 3.0 Titan: Exploring Larger-scale Knowledge Enhanced Pre-training for Language Understanding and GenerationarXiv 2021ERNIE-Code: Beyond English-Centric Cross-lingual Pretraining for Programming LanguagesarXiv 2022ERNIE-Layout: Layout Knowledge Enhanced Pre-training for Visually-rich Document UnderstandingarXiv 2022UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive LearningACL 2021 5Building Chinese Biomedical Language Models via Multi-Level Text DiscriminationarXiv 2021ERNIE: Enhanced Representation through Knowledge IntegrationarXiv 2019ERNIE-Gram: Pre-Training with Explicitly N-Gram Masked Language Modeling for Natural Language UnderstandingNAACL 2021 4Agentless: Demystifying LLM-based Software Engineering AgentsarXiv 2024Meta Learning Text-to-Speech Synthesis in over 7000 LanguagesarXiv 2024Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial IntelligencearXiv 2025OneFlow: Redesign the Distributed Deep Learning Framework from ScratcharXiv 2021ERNIE 2.0: A Continual Pre-training Framework for Language UnderstandingarXiv 2019DeepSeek-VL: Towards Real-World Vision-Language UnderstandingarXiv 2024MotionGPT: Human Motion as a Foreign LanguageNeurIPS 2023 11xLSTM 7B: A Recurrent LLM for Fast and Efficient InferencearXiv 2025XCube: Large-Scale 3D Generative Modeling using Sparse Voxel Hierarchiesxcube-large-scale-3d-generative-modelingOmniRe: Omni Urban Scene ReconstructionarXiv 2024HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge AdaptationarXiv 2025Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning AbilitiesarXiv 2025Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue AbilitiesarXiv 2024Reflexion: Language Agents with Verbal Reinforcement LearningNeurIPS 2023 11Observation-Centric SORT: Rethinking SORT for Robust Multi-Object TrackingCVPR 2023 1Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphsarXiv 2016MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic VideosarXiv 2024ScreenAgent: A Vision Language Model-driven Computer Control AgentarXiv 2024BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and DatasetarXiv 2025DepthSplat: Connecting Gaussian Splatting and DepthCVPR 2025 1AI2-THOR: An Interactive 3D Environment for Visual AIarXiv 2017Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language ModelsarXiv 2025DiffSinger: Singing Voice Synthesis via Shallow Diffusion MechanismarXiv 2021DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous GraspingarXiv 20252D3D-MATR: 2D-3D Matching Transformer for Detection-free Registration between Images and Point CloudsICCV 2023 1Bidirectional Copy-Paste for Semi-Supervised Medical Image SegmentationCVPR 2023 1Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative DecodingarXiv 2024Multiple Object Tracking as ID PredictionCVPR 2025 1MING-MOE: Enhancing Medical Multi-Task Learning in Large Language Models with Sparse Mixture of Low-Rank Adapter ExpertsarXiv 2024Generalized Neighborhood Attention: Multi-dimensional Sparse Attention at the Speed of LightarXiv 2025AerialVLN: Vision-and-Language Navigation for UAVsICCV 2023 1TiRex: Zero-Shot Forecasting Across Long and Short Horizons with Enhanced In-Context LearningarXiv 2025LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?arXiv 2025SCOPE: Simulating Cross-game Operations in Playable Environments for FPS World ModelsarXiv 2026OpenClaw-RL: Train Any Agent Simply by TalkingarXiv 2026World Action Models: The Next Frontier in Embodied AIarXiv 2026Exploring Intrinsic Normal Prototypes within a Single Image for Universal Anomaly DetectionCVPR 2025 1Scaling Up Models and Data with $\texttt{t5x}$ and $\texttt{seqio}$arXiv 2022SkillMimic: Learning Basketball Interaction Skills from DemonstrationsCVPR 2025 1PathRAG: Pruning Graph-based Retrieval Augmented Generation with Relational PathsarXiv 2025DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic ModelsarXiv 2022From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgearXiv 2024RepoMaster: Autonomous Exploration and Understanding of GitHub Repositories for Complex Task SolvingarXiv 2025Retrieval-Augmented Generation with Graphs (GraphRAG)arXiv 2024A Self-Improving Coding AgentarXiv 2025Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataCVPR 2024 1Open-CD: A Comprehensive Toolbox for Change DetectionarXiv 2024HyperGraphRAG: Retrieval-Augmented Generation with Hypergraph-Structured Knowledge RepresentationarXiv 2025MobileMamba: Lightweight Multi-Receptive Visual Mamba NetworkCVPR 2025 1LeanDojo: Theorem Proving with Retrieval-Augmented Language Modelsleandojo-theorem-proving-with-retrievalLayoutParser: A Unified Toolkit for Deep Learning Based Document Image AnalysisarXiv 2021On the limits of agency in agent-based modelsarXiv 2024RSPrompter: Learning to Prompt for Remote Sensing Instance Segmentation based on Visual Foundation ModelarXiv 2023Conditional Prompt Learning for Vision-Language ModelsCVPR 2022 1Ming-Omni: A Unified Multimodal Model for Perception and GenerationarXiv 2025Deep Learning for Camera Calibration and Beyond: A SurveyarXiv 2023Harnessing Multiple Large Language Models: A Survey on LLM EnsemblearXiv 2025Uni3C: Unifying Precisely 3D-Enhanced Camera and Human Motion Controls for Video GenerationarXiv 2025Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry PriorsarXiv 2025Seamless: Multilingual Expressive and Streaming Speech TranslationarXiv 2023Visual Large Language Models for Generalized and Specialized ApplicationsarXiv 2025