0

All papers

Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.

MatchAnything: Universal Cross-Modality Image Matching with Large-Scale Pre-TrainingarXiv 2025AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM AgentsarXiv 2024Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent EvolutionarXiv 2025TerraMind: Large-Scale Generative Multimodality for Earth ObservationICCV 2025Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive ReinforcementarXiv 2025StrongSORT: Make DeepSORT Great AgainarXiv 2022MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View ImagesarXiv 2024DreamGaussian: Generative Gaussian Splatting for Efficient 3D Content CreationarXiv 2023ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View SynthesisarXiv 2024Instruction Tuning with Human CurriculumarXiv 2023Separate Anything You DescribearXiv 2023From Bytes to Ideas: Language Modeling with Autoregressive U-NetsarXiv 2025DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language ModelsarXiv 2024StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task LearningarXiv 2024Exploiting Diffusion Prior for Real-World Image Super-ResolutionarXiv 2023Depth Anything with Any PriorarXiv 2025AdaptCLIP: Adapting CLIP for Universal Visual Anomaly DetectionarXiv 2025WORLDMEM: Long-term Consistent World Simulation with MemoryarXiv 2025DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking ProcessarXiv 2025Mamba YOLO: A Simple Baseline for Object Detection with State Space ModelarXiv 2024AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and ReasoningarXiv 2025Corrective Retrieval Augmented GenerationarXiv 2024LHM: Large Animatable Human Reconstruction Model from a Single Image in SecondsarXiv 2025Marco-o1: Towards Open Reasoning Models for Open-Ended SolutionsarXiv 2024Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse AttentionarXiv 2025MagCache: Fast Video Generation with Magnitude-Aware CachearXiv 2025CubiCasa5K: A Dataset and an Improved Multi-Task Model for Floorplan Image AnalysisarXiv 2019LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation ModelsICCV 2025General OCR Theory: Towards OCR-2.0 via a Unified End-to-end ModelarXiv 2024DeepSeek-OCR 2: Visual Causal FlowarXiv 2026Colorful Diffuse Intrinsic Image Decomposition in the Wildcolorful-diffuse-intrinsic-imageMimic Intent, Not Just TrajectoriesarXiv 2026Octree-GS: Towards Consistent Real-time Rendering with LOD-Structured 3D GaussiansarXiv 2024DarkIR: Robust Low-Light Image RestorationCVPR 2025 1Evaluate & Evaluation on the Hub: Better Best Practices for Data and Model MeasurementsarXiv 2022Diffusion Policy Policy OptimizationarXiv 2024RefAlign: Representation Alignment for Reference-to-Video GenerationarXiv 2026Tails Tell Tales: Chapter-Wide Manga Transcriptions with Character NamesarXiv 2024D4RL: Datasets for Deep Data-Driven Reinforcement LearningarXiv 2020EfficientNetV2: Smaller Models and Faster TrainingarXiv 2021EfficientDet: Scalable and Efficient Object Detectionefficientdet-scalable-and-efficient-object-1Learning to Reason under Off-Policy GuidancearXiv 2025Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language ModelsarXiv 2023Qwen2-Audio Technical ReportarXiv 2024DDColor: Towards Photo-Realistic Image Colorization via Dual DecodersICCV 2023 1OmniSQL: Synthesizing High-quality Text-to-SQL Data at ScalearXiv 2025SLAM3R: Real-Time Dense Scene Reconstruction from Monocular RGB VideosCVPR 2025 1Evaluating Object Hallucination in Large Vision-Language ModelsarXiv 2023DoWhy-GCM: An extension of DoWhy for causal inference in graphical causal modelsarXiv 2022Verbs in Action: Improving verb understanding in video-language modelsICCV 2023 1Streaming Dense Video CaptioningCVPR 2024 1Attention Bottlenecks for Multimodal FusionNeurIPS 2021 12VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality DocumentsarXiv 2024GPT4Scene: Understand 3D Scenes from Videos with Vision-Language ModelsarXiv 2025ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsarXiv 2023Scene as OccupancyICCV 2023 1Align Anything: Training All-Modality Models to Follow Instructions with Language FeedbackarXiv 2024LiteWebAgent: The Open-Source Suite for VLM-Based Web-Agent ApplicationsarXiv 2025LLMDFA: Analyzing Dataflow in Code with Large Language ModelsarXiv 2024Diffusion Models Beat GANs on Image SynthesisNeurIPS 2021 12Foundation Models in Robotics: Applications, Challenges, and the FuturearXiv 2023RoMa: Robust Dense Feature MatchingCVPR 2024 1Step1X-3D: Towards High-Fidelity and Controllable Generation of Textured 3D AssetsarXiv 2025Autoregressive Models in Vision: A SurveyarXiv 2024OmAgent: A Multi-modal Agent Framework for Complex Video Understanding with Task Divide-and-ConquerarXiv 2024Improving Pixel-based MIM by Reducing Wasted Modeling CapabilityICCV 2023 1Don't Do RAG: When Cache-Augmented Generation is All You Need for Knowledge TasksarXiv 2024Elucidating the Design Space of Diffusion-Based Generative ModelsarXiv 2022nuScenes: A multimodal dataset for autonomous drivingnuscenes-a-multimodal-dataset-for-autonomous-1Retinexformer: One-stage Retinex-based Transformer for Low-light Image EnhancementICCV 2023 1AutoCodeRover: Autonomous Program ImprovementarXiv 2024RAVE: A variational autoencoder for fast and high-quality neural audio synthesisrave-a-variational-autoencoder-for-fast-andHLLM: Enhancing Sequential Recommendations via Hierarchical Large Language Models for Item and User ModelingarXiv 2024SkillX: Automatically Constructing Skill Knowledge Bases for AgentsarXiv 2026ZeroSearch: Incentivize the Search Capability of LLMs without SearchingarXiv 2025ImLoc: Revisiting Visual Localization with Image-based RepresentationarXiv 2026AudioLDM 2: Learning Holistic Audio Generation with Self-supervised PretrainingarXiv 2023Swin2SR: SwinV2 Transformer for Compressed Image Super-Resolution and RestorationarXiv 2022DrivAerNet++: A Large-Scale Multimodal Car Dataset with Computational Fluid Dynamics Simulations and Deep Learning BenchmarksarXiv 2024Kimi-VL Technical ReportarXiv 2025TinyTL: Reduce Activations, Not Trainable Parameters for Efficient On-Device Learningtinytl-reduce-memory-not-parameters-forEfficient Streaming Language Models with Attention SinksarXiv 2023XAttention: Block Sparse Attention with Antidiagonal ScoringarXiv 2025QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM ServingarXiv 20243D Dynamic Scene Graphs: Actionable Spatial Perception with Places, Objects, and HumansarXiv 2020MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-MakingarXiv 2024LServe: Efficient Long-sequence LLM Serving with Unified Sparse AttentionarXiv 2025TorchXRayVision: A library of chest X-ray datasets and modelsarXiv 2021Mistral 7BarXiv 2023Pixtral 12BarXiv 2024Exploring the Evolution of Physics Cognition in Video Generation: A SurveyarXiv 2025ExT5: Towards Extreme Multi-Task Scaling for Transfer Learningext5-towards-extreme-multi-task-scaling-forvAttention: Dynamic Memory Management for Serving LLMs without PagedAttentionarXiv 2024SPANN: Highly-efficient Billion-scale Approximate Nearest Neighbor Searchspann-highly-efficient-billion-scale-1From Reusing to Forecasting: Accelerating Diffusion Models with TaylorSeersICCV 2025EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied AgentsarXiv 2025TinyCLIP: CLIP Distillation via Affinity Mimicking and Weight InheritanceICCV 2023 1EfficientViT: Memory Efficient Vision Transformer with Cascaded Group AttentionCVPR 2023 1DeepCFD: Efficient Steady-State Laminar Flow Approximation with Deep Convolutional Neural NetworksarXiv 2020SegMamba: Long-range Sequential Modeling Mamba For 3D Medical Image SegmentationarXiv 2024GEO: Generative Engine OptimizationarXiv 2023PETRv2: A Unified Framework for 3D Perception from Multi-Camera ImagesICCV 2023 1LLM2Vec: Large Language Models Are Secretly Powerful Text EncodersarXiv 2024Exploring the structure of a real-time, arbitrary neural artistic stylization networkarXiv 2017HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMsarXiv 2024MIDI: Multi-Instance Diffusion for Single Image to 3D Scene GenerationCVPR 2025 1GS-IR: 3D Gaussian Splatting for Inverse RenderingCVPR 2024 1OLMoE: Open Mixture-of-Experts Language ModelsarXiv 2024KVzip: Query-Agnostic KV Cache Compression with Context ReconstructionarXiv 2025USP: A Unified Sequence Parallelism Approach for Long Context Generative AIarXiv 2024MiMo-VL Technical ReportarXiv 2025UMAP: Uniform Manifold Approximation and Projection for Dimension ReductionarXiv 2018Router-R1: Teaching LLMs Multi-Round Routing and Aggregation via Reinforcement LearningarXiv 2025DiffuEraser: A Diffusion Model for Video InpaintingarXiv 2025FlowEdit: Inversion-Free Text-Based Editing Using Pre-Trained Flow ModelsICCV 2025A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and TrustworthinessarXiv 2024Agile But Safe: Learning Collision-Free High-Speed Legged LocomotionarXiv 2024NodeRAG: Structuring Graph-based RAG with Heterogeneous NodesarXiv 2025Improving Autoformalization using Type CheckingarXiv 2024Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward PassCVPR 2025 1DAIR-V2X: A Large-Scale Dataset for Vehicle-Infrastructure Cooperative 3D Object DetectionCVPR 2022 1Navigating Text-To-Image Customization: From LyCORIS Fine-Tuning to Model EvaluationarXiv 2023The Surprising Effectiveness of Negative Reinforcement in LLM ReasoningarXiv 2025Exploring Object-Centric Temporal Modeling for Efficient Multi-View 3D Object DetectionICCV 2023 1Simple and Effective Masked Diffusion Language ModelsarXiv 2024VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion ModelsCVPR 2024 1Locating and Editing Factual Associations in GPTarXiv 2022Deep Residual Learning for Image Recognitiondeep-residual-learning-for-image-recognition-1ToolRL: Reward is All Tool Learning NeedsarXiv 2025CoMotion: Concurrent Multi-person 3D MotionarXiv 2025Seed1.5-VL Technical ReportarXiv 2025OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?CVPR 2025 1Learning to Reason and Memorize with Self-Noteslearning-to-reason-and-memorize-with-selfPutting People in their Place: Monocular Regression of 3D People in DepthCVPR 2022 1Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and OpportunitiesarXiv 2024OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and BeyondarXiv 2026WildClawBench: A Benchmark for Real-World, Long-Horizon Agent EvaluationarXiv 2026DexJoCo: A Benchmark and Toolkit for Task-Oriented Dexterous Manipulation on MuJoCoarXiv 2026Gated DeltaNet-2: Decoupling Erase and Write in Linear AttentionarXiv 2026DeepCode: Open Agentic CodingarXiv 2025Audio-Visual Intelligence in Large Foundation ModelsarXiv 2026SAM Audio: Segment Anything in AudioarXiv 2025Object-Centric Learning with Slot AttentionNeurIPS 2020 12Region-centric Image-Language Pretraining for Open-Vocabulary DetectionarXiv 2023Large Action Models: From Inception to ImplementationarXiv 2024Kosmos-2: Grounding Multimodal Large Language Models to the WorldarXiv 2023Self-training and Pre-training are Complementary for Speech RecognitionarXiv 2020Intrinsic Image Decomposition via Ordinal Shadingintrinsic-image-decomposition-via-ordinalSimple Open-Vocabulary Object Detection with Vision TransformersarXiv 2022Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video CaptioningCVPR 2023 1WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language ModelingarXiv 2024VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement LearningarXiv 2025Vision-Language Models for Vision Tasks: A SurveyarXiv 2023Pyramid Stereo Matching Networkpyramid-stereo-matching-network-1Free-Form Image Inpainting with Gated Convolutionfree-form-image-inpainting-with-gated-1OpenOccupancy: A Large Scale Benchmark for Surrounding Semantic Occupancy PerceptionICCV 2023 1GMT: General Motion Tracking for Humanoid Whole-Body ControlarXiv 2025Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain PerspectivearXiv 2025OmniPaint: Mastering Object-Oriented Editing via Disentangled Insertion-Removal InpaintingICCV 2025MV-Adapter: Multi-view Consistent Image Generation Made EasyarXiv 2024GPTQ: Accurate Post-Training Quantization for Generative Pre-trained TransformersarXiv 2022pixelSplat: 3D Gaussian Splats from Image Pairs for Scalable Generalizable 3D ReconstructionCVPR 2024 1ZoeDepth: Zero-shot Transfer by Combining Relative and Metric DeptharXiv 2023LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language ModelsCVPR 2025 1LVSM: A Large View Synthesis Model with Minimal 3D Inductive BiasarXiv 2024VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation ModelsarXiv 2025BioReason: Incentivizing Multimodal Biological Reasoning within a DNA-LLM ModelarXiv 2025Unpaired Image-to-Image Translation via Neural Schrödinger BridgearXiv 2023HUGS: Human Gaussian SplatsCVPR 2024 1LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation MethodsarXiv 2024FilmAgent: A Multi-Agent Framework for End-to-End Film Automation in Virtual 3D SpacesarXiv 2025RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion TransformersarXiv 2025Code Generation with AlphaCodium: From Prompt Engineering to Flow EngineeringarXiv 2024Effective Whole-body Pose Estimation with Two-stages DistillationarXiv 2023This Time is Different: An Observability Perspective on Time Series Foundation ModelsarXiv 2025SimVPv2: Towards Simple yet Powerful Spatiotemporal Predictive LearningarXiv 2022OpenSTL: A Comprehensive Benchmark of Spatio-Temporal Predictive Learningopenstl-a-comprehensive-benchmark-of-spatioSoK: Evaluating Jailbreak Guardrails for Large Language ModelsarXiv 2025CMX: Cross-Modal Fusion for RGB-X Semantic Segmentation with TransformersarXiv 2022HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalarXiv 2024CPsyCoun: A Report-based Multi-turn Dialogue Reconstruction and Evaluation Framework for Chinese Psychological CounselingarXiv 2024GhostNet: More Features from Cheap Operationsghostnet-more-features-from-cheap-operations-1Augmented Shortcuts for Vision TransformersNeurIPS 2021 12PaSa: An LLM Agent for Comprehensive Academic Paper SearcharXiv 2025In-depth Analysis of Graph-based RAG in a Unified FrameworkarXiv 2025MeshAnything V2: Artist-Created Mesh Generation With Adjacent Mesh TokenizationICCV 2025TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion ModelsICCV 2025From Automation to Autonomy: A Survey on Large Language Models in Scientific DiscoveryarXiv 2025Distil-Whisper: Robust Knowledge Distillation via Large-Scale Pseudo LabellingarXiv 2023ECG-FM: An Open Electrocardiogram Foundation ModelarXiv 2024SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the WildarXiv 2025Joint Feature Learning and Relation Modeling for Tracking: A One-Stream FrameworkarXiv 2022LLaDA-V: Large Language Diffusion Models with Visual Instruction TuningarXiv 2025SIMPL: A Simple and Efficient Multi-agent Motion Prediction Baseline for Autonomous DrivingarXiv 2024Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of ExpertsarXiv 2024StereoCrafter: Diffusion-based Generation of Long and High-fidelity Stereoscopic 3D from Monocular VideosarXiv 2024Retrieval-Augmented Generation with Hierarchical KnowledgearXiv 2025MedCLIP-SAMv2: Towards Universal Text-Driven Medical Image SegmentationarXiv 2024GenMol: A Drug Discovery Generalist with Discrete DiffusionarXiv 2025TheoremExplainAgent: Towards Video-based Multimodal Explanations for LLM Theorem UnderstandingarXiv 2025

Back to Papers