All papers
Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.
Prompting Depth Anything for 4K Resolution Accurate Metric Depth EstimationCVPR 2025 1A Survey of Graph Retrieval-Augmented Generation for Customized Large Language ModelsarXiv 2025WebLLM: A High-Performance In-Browser LLM Inference EnginearXiv 2024DanceGRPO: Unleashing GRPO on Visual GenerationarXiv 2025V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and PlanningarXiv 2025FoundationStereo: Zero-Shot Stereo MatchingCVPR 2025 1Lens: Rethinking Training Efficiency for Foundational Text-to-Image ModelsarXiv 2026Paper2Poster: Towards Multimodal Poster Automation from Scientific PapersarXiv 2025SANA-Sprint: One-Step Diffusion with Continuous-Time Consistency DistillationarXiv 2025SantaCoder: don't reach for the stars!arXiv 2023Automated head and neck tumor segmentation from 3D PET/CTarXiv 2022LoRA: Low-Rank Adaptation of Large Language Modelslora-low-rank-adaptation-of-large-language-1Grounded SAM: Assembling Open-World Models for Diverse Visual TasksarXiv 2024Hybrid Transformers for Music Source SeparationarXiv 2022Towards Automated Circuit Discovery for Mechanistic Interpretabilitytowards-automated-circuit-discovery-forYuE: Scaling Open Foundation Models for Long-Form Music GenerationarXiv 2025MMA-Diffusion: MultiModal Attack on Diffusion ModelsCVPR 2024 1OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity FieldsarXiv 2018MLVU: Benchmarking Multi-task Long Video UnderstandingCVPR 2025 1Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200karXiv 2025End-to-End Object Detection with TransformersECCV 2020 8Tutel: Adaptive Mixture-of-Experts at ScalearXiv 2022RTMDet: An Empirical Study of Designing Real-Time Object DetectorsarXiv 2022An Open and Comprehensive Pipeline for Unified Object Grounding and DetectionarXiv 2024Resolution-robust Large Mask Inpainting with Fourier ConvolutionsarXiv 2021DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous DrivingCVPR 2025 1SeedVR2: One-Step Video Restoration via Diffusion Adversarial Post-TrainingarXiv 2025FastPitch: Parallel Text-to-speech with Pitch PredictionarXiv 2020Collaborating Action by Action: A Multi-agent LLM Framework for Embodied ReasoningarXiv 2025Datasets: A Community Library for Natural Language ProcessingEMNLP (ACL) 2021 11gRNAde: Geometric Deep Learning for 3D RNA inverse designarXiv 2023DeepFaceLab: Integrated, flexible and extensible face-swapping frameworkarXiv 2020FG-CLIP: Fine-Grained Visual and Textual AlignmentarXiv 2025Group-in-Group Policy Optimization for LLM Agent TrainingarXiv 2025RTMW: Real-Time Multi-Person 2D and 3D Whole-body Pose EstimationarXiv 2024Revisiting Feature Prediction for Learning Visual Representations from VideoarXiv preprint 2024 2SAM3-LiteText: An Anatomical Study of the SAM3 Text Encoder for Efficient Vision-Language SegmentationarXiv 2026VACE: All-in-One Video Creation and EditingICCV 2025Interpretable Machine Learning for Science with PySR and SymbolicRegression.jlarXiv 2023nnU-Net Revisited: A Call for Rigorous Validation in 3D Medical Image SegmentationarXiv 2024IMAGGarment-1: Fine-Grained Garment Generation for Controllable Fashion DesignarXiv 2025Eliza: A Web3 friendly AI Agent Operating SystemarXiv 2025Denoising Diffusion Probabilistic ModelsNeurIPS 2020 12Benchmarking Multimodal AutoML for Tabular Data with Text FieldsarXiv 2021OLMo: Accelerating the Science of Language ModelsarXiv 2024Convolution Meets LoRA: Parameter Efficient Finetuning for Segment Anything ModelarXiv 2024Fast Graph Representation Learning with PyTorch GeometricarXiv 2019CodexGraph: Bridging Large Language Models and Code Repositories via Code Graph DatabasesarXiv 2024One-Shot Diffusion Mimicker for Handwritten Text GenerationarXiv 2024Voyager: An Open-Ended Embodied Agent with Large Language ModelsarXiv 2023Skywork R1V2: Multimodal Hybrid Reinforcement Learning for ReasoningarXiv 2025From RAG to Memory: Non-Parametric Continual Learning for Large Language ModelsarXiv 2025Dinomaly: The Less Is More Philosophy in Multi-Class Unsupervised Anomaly DetectionCVPR 2025 1MonkeyOCR: Document Parsing with a Structure-Recognition-Relation Triplet ParadigmarXiv 20253DGUT: Enabling Distorted Cameras and Secondary Rays in Gaussian SplattingCVPR 2025 1Flow-OPD: On-Policy Distillation for Flow Matching ModelsarXiv 2026Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelarXiv 2024AppAgent: Multimodal Agents as Smartphone UsersarXiv 2023LTX-Video: Realtime Video Latent DiffusionarXiv 2024TinyLlama: An Open-Source Small Language ModelarXiv 2024LiT: Zero-Shot Transfer with Locked-image text TuningCVPR 2022 1Perceiver IO: A General Architecture for Structured Inputs & Outputsperceiver-io-a-general-architecture-for-1Towards mental time travel: a hierarchical memory for reinforcement learning agentsNeurIPS 2021 12Bootstrap your own latent: A new approach to self-supervised LearningarXiv 2020High-Performance Large-Scale Image Recognition Without NormalizationarXiv 2021A Lip Sync Expert Is All You Need for Speech to Lip Generation In The WildarXiv 2020SwinIR: Image Restoration Using Swin TransformerarXiv 2021QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement LearningarXiv 2025Data-Juicer: A One-Stop Data Processing System for Large Language ModelsarXiv 2023Universal and Transferable Adversarial Attacks on Aligned Language ModelsarXiv 2023WeNet: Production oriented Streaming and Non-streaming End-to-End Speech Recognition ToolkitarXiv 2021Liger Kernel: Efficient Triton Kernels for LLM TrainingarXiv 2024Unified Multimodal Understanding and Generation Models: Advances, Challenges, and OpportunitiesarXiv 20253D-Speaker: A Large-Scale Multi-Device, Multi-Distance, and Multi-Dialect Corpus for Speech Representation DisentanglementarXiv 2023Proactive Detection of Voice Cloning with Localized WatermarkingarXiv 2024EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time TestarXiv 2025TorchTitan: One-stop PyTorch native solution for production ready LLM pre-trainingarXiv 2024RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationarXiv 2025Slicing Aided Hyper Inference and Fine-tuning for Small Object DetectionarXiv 2022NeRF-Det: Learning Geometry-Aware Volumetric Representation for Multi-View 3D Object DetectionICCV 2023 1FastVLM: Efficient Vision Encoding for Vision Language ModelsCVPR 2025 1A Comprehensive Survey of Mixture-of-Experts: Algorithms, Theory, and ApplicationsarXiv 2025ClearerVoice-Studio: Bridging Advanced Speech Processing Research and
Practical DeploymentarXiv 2025ByteTrack: Multi-Object Tracking by Associating Every Detection Boxbytetrack-multi-object-tracking-byVideoRAG: Retrieval-Augmented Generation with Extreme Long-Context VideosarXiv 2025Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity AsymmetryarXiv 2026MLE-bench: Evaluating Machine Learning Agents on Machine Learning EngineeringarXiv 2024DiffRhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent DiffusionarXiv 2025WaDi: Weight Direction-aware Distillation for One-step Image SynthesisarXiv 2026UFO2: The Desktop AgentOSarXiv 2025UFO: A UI-Focused Agent for Windows OS InteractionarXiv 2024Moshi: a speech-text foundation model for real-time dialoguearXiv 2024Open-Vocabulary Camouflaged Object SegmentationarXiv 2023QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction AgentsarXiv 2026What Drives Success in Physical Planning with Joint-Embedding Predictive World Models?arXiv 2025Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation ModelsarXiv 2025ReCamMaster: Camera-Controlled Generative Rendering from A Single VideoICCV 2025Agentic Retrieval-Augmented Generation: A Survey on Agentic RAGarXiv 2025The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree SearcharXiv 2025AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningarXiv 2023Large Language Model Agent: A Survey on Methodology, Applications and ChallengesarXiv 2025Pixel-SAIL: Single Transformer For Pixel-Grounded UnderstandingarXiv 2025Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and VideosarXiv 2025Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidatessubword-regularization-improving-neural-1Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative RecommendationsarXiv 2024Training Large Language Models to Reason in a Continuous Latent SpacearXiv 2024TADA: A Generative Framework for Speech Modeling via Text-Acoustic Dual AlignmentarXiv 2026Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation ModelsarXiv 2025VOID: Video Object and Interaction DeletionarXiv 2026EasyJailbreak: A Unified Framework for Jailbreaking Large Language ModelsarXiv 2024A Survey of Large Language Models in Medicine: Progress, Application, and ChallengearXiv 2023A-MEM: Agentic Memory for LLM AgentsarXiv 2025TOPIQ: A Top-down Approach from Semantics to Distortions for Image Quality AssessmentarXiv 2023Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Modelsreconstruction-vs-generation-tamingNAFSSR: Stereo Image Super-Resolution Using NAFNetarXiv 2022Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head SynthesisarXiv 2024Toward Native Multimodal Modeling: A RoadmaparXiv 2026Zero-shot Voice Conversion with Diffusion TransformersarXiv 2024BLAST: Balanced Sampling Time Series Corpus for Universal Forecasting ModelsarXiv 2025XGrammar: Flexible and Efficient Structured Generation Engine for Large Language ModelsarXiv 2024OS-Copilot: Towards Generalist Computer Agents with Self-ImprovementarXiv 2024SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersNeurIPS 2021 12GLM-5: from Vibe Coding to Agentic EngineeringarXiv 2026UniK3D: Universal Camera Monocular 3D EstimationCVPR 2025 1dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language ModelarXiv 2025Aviary: training language agents on challenging scientific tasksarXiv 2024RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use AgentsarXiv 2025Qwen3-ASR Technical ReportarXiv 2026Llama Guard 3 Vision: Safeguarding Human-AI Image Understanding ConversationsarXiv 2024YOLOv13: Real-Time Object Detection with Hypergraph-Enhanced Adaptive Visual PerceptionarXiv 2025LeVo: High-Quality Song Generation with Multi-Preference AlignmentarXiv 2025World Guidance: World Modeling in Condition Space for Action GenerationarXiv 2026CoTracker: It is Better to Track TogetherarXiv 2023Deep Industrial Image Anomaly Detection: A SurveyarXiv 2023Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and BeyondarXiv 2023MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon AgentsarXiv 2025DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive PerceptionarXiv 2024Tree of Thoughts: Deliberate Problem Solving with Large Language Modelstree-of-thoughts-deliberate-problem-solvingMobileSAMv2: Faster Segment Anything to EverythingarXiv 2023Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe SystemsarXiv 2025NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable RailsarXiv 2023Learning to Discover at Test TimearXiv 2026DUSt3R: Geometric 3D Vision Made EasyCVPR 2024 1Muon is Scalable for LLM TrainingarXiv 2025The Python Simulations of Chemistry Framework: 10 years of an open-source quantum chemistry projectarXiv 2026TaskWeaver: A Code-First Agent FrameworkarXiv 2023SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense FeaturesarXiv 2025SuGaR: Surface-Aligned Gaussian Splatting for Efficient 3D Mesh Reconstruction and High-Quality Mesh RenderingCVPR 2024 1EchoMimicV2: Towards Striking, Simplified, and Semi-Body Human AnimationCVPR 2025 1DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code IntelligencearXiv 2024DFlash: Block Diffusion for Flash Speculative DecodingarXiv 2026BERTopic: Neural topic modeling with a class-based TF-IDF procedurearXiv 2022FireRedASR: Open-Source Industrial-Grade Mandarin Speech Recognition Models from Encoder-Decoder to LLM IntegrationarXiv 2025Let Them Talk: Audio-Driven Multi-Person Conversational Video GenerationarXiv 2025Challenges in Trustworthy Human Evaluation of ChatbotsarXiv 2024Chinese CLIP: Contrastive Vision-Language Pretraining in ChinesearXiv 2022Soap2Soap: Long Cinematic Video Remaking via Multi-Agent CollaborationarXiv 2026GPT4All: An Ecosystem of Open Source Compressed Language ModelsarXiv 2023Masked Autoencoders Are Scalable Vision LearnersCVPR 2022 1Unifying the Perspectives of NLP and Software Engineering: A Survey on Language Models for CodearXiv 2023YOLO-World: Real-Time Open-Vocabulary Object DetectionCVPR 2024 1RWKV: Reinventing RNNs for the Transformer EraarXiv 2023Dive into Claude Code: The Design Space of Today's and Future AI Agent SystemsarXiv 2026YOLOE: Real-Time Seeing AnythingICCV 2025XDoc: Unified Pre-training for Cross-Format Document UnderstandingarXiv 2022TrOCR: Transformer-based Optical Character Recognition with Pre-trained ModelsarXiv 2021BEATs: Audio Pre-Training with Acoustic TokenizersarXiv 2022Kosmos-G: Generating Images in Context with Multimodal Large Language ModelsarXiv 2023DiT: Self-supervised Pre-training for Document Image TransformerarXiv 2022Preference Optimization for Reasoning with Pseudo FeedbackarXiv 2024You Only Cache Once: Decoder-Decoder Architectures for Language ModelsarXiv 2024MathScale: Scaling Instruction Tuning for Mathematical ReasoningarXiv 2024Multimodal Latent Language Modeling with Next-Token DiffusionarXiv 2024LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document UnderstandingarXiv 2021BEiT v2: Masked Image Modeling with Vector-Quantized Visual TokenizersarXiv 2022DeltaLM: Encoder-Decoder Pre-training for Language Generation and Translation by Augmenting Pretrained Multilingual EncodersarXiv 2021VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-ExpertsarXiv 2021LayoutLMv3: Pre-training for Document AI with Unified Text and Image MaskingarXiv 2022Unlocking Dense Metric Depth Estimation in VLMsarXiv 2026Fast Segment AnythingarXiv 2023High-Fidelity Audio Compression with Improved RVQGANNeurIPS 2023 11MetaClaw: Just Talk -- An Agent That Meta-Learns and Evolves in the WildarXiv 2026Continuous 3D Perception Model with Persistent StateCVPR 2025 1Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy DistillationarXiv 2026ThunderKittens: Simple, Fast, and Adorable AI KernelsarXiv 2024Causal-JEPA: Learning World Models through Object-Level Latent MaskingarXiv 2026ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action LooparXiv 2026OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive AnnotationsCVPR 2025 1Planning-oriented Autonomous DrivingCVPR 2023 1On Neural Differential EquationsarXiv 2022MAGI-1: Autoregressive Video Generation at ScalearXiv 2025Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosarXiv 20223D Scene Generation: A SurveyarXiv 2025Deep Speech 2: End-to-End Speech Recognition in English and MandarinarXiv 2015World Action Models are Zero-shot PoliciesarXiv 2026Colossal-AI: A Unified Deep Learning System For Large-Scale Parallel TrainingarXiv 2021LISA: Reasoning Segmentation via Large Language ModelCVPR 2024 14D-RGPT: Toward Region-level 4D Understanding via Perceptual DistillationarXiv 2025Scaling Multiagent Systems with Process RewardsarXiv 2026PixARMesh: Autoregressive Mesh-Native Single-View Scene ReconstructionarXiv 2026