All papers
Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.
Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement LearningarXiv 2025Cross-modal Learning for Image-Guided Point Cloud Shape CompletionarXiv 2022Argmax Flows and Multinomial Diffusion: Learning Categorical DistributionsNeurIPS 2021 12Transformer-based Planning for Symbolic Regressiontransformer-based-planning-for-symbolicOS-Harm: A Benchmark for Measuring Safety of Computer Use AgentsarXiv 2025Mamba as a Bridge: Where Vision Foundation Models Meet Vision Language Models for Domain-Generalized Semantic SegmentationCVPR 2025 1GRAM: A Generative Foundation Reward Model for Reward GeneralizationarXiv 2025On the Cross-lingual Transferability of Monolingual Representationson-the-cross-lingual-transferability-of-1Adaptively Sparse Transformersadaptively-sparse-transformers-1Teaching Machines to Read and Comprehendteaching-machines-to-read-and-comprehend-1MOS: Towards Scaling Out-of-distribution Detection for Large Semantic SpaceCVPR 2021 1ResUNet++: An Advanced Architecture for Medical Image SegmentationarXiv 2019Model Quantization and Hardware Acceleration for Vision Transformers: A Comprehensive SurveyarXiv 2024Top2Vec: Distributed Representations of TopicsarXiv 2020BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-DistillationarXiv 2024Pathformer: Multi-scale Transformers with Adaptive Pathways for Time Series ForecastingarXiv 2024Scalable Training of Artificial Neural Networks with Adaptive Sparse Connectivity inspired by Network SciencearXiv 2017Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial ReasoningarXiv 2024MM-BigBench: Evaluating Multimodal Models on Multimodal Content Comprehension TasksarXiv 2023REGNav: Room Expert Guided Image-Goal NavigationarXiv 2025DisCO: Reinforcing Large Reasoning Models with Discriminative Constrained OptimizationarXiv 2025RoboFAC: A Comprehensive Framework for Robotic Failure Analysis and CorrectionarXiv 2025Unlocking Feature Visualization for Deeper Networks with MAgnitude Constrained OptimizationarXiv 2023DoNet: Deep De-overlapping Network for Cytology Instance SegmentationCVPR 2023 1YOLACT++: Better Real-time Instance SegmentationarXiv 2019Bamboo: Building Mega-Scale Vision Dataset Continually with Human-Machine SynergyarXiv 2022Understanding the Role of Individual Units in a Deep Neural NetworkarXiv 2020PlanGenLLMs: A Modern Survey of LLM Planning CapabilitiesarXiv 2025Understanding the Repeat Curse in Large Language Models from a Feature PerspectivearXiv 20253DFlowAction: Learning Cross-Embodiment Manipulation from 3D Flow World ModelarXiv 2025TransFace: Calibrating Transformer Training for Face Recognition from a Data-Centric PerspectiveICCV 2023 1An LLM-Assisted Easy-to-Trigger Backdoor Attack on Code Completion Models: Injecting Disguised Vulnerabilities against Strong DetectionarXiv 2024CaRtGS: Computational Alignment for Real-Time Gaussian Splatting SLAMarXiv 2024INTERS: Unlocking the Power of Large Language Models in Search with Instruction TuningarXiv 2024IRCAN: Mitigating Knowledge Conflicts in LLM Generation via Identifying and Reweighting Context-Aware NeuronsarXiv 2024Gotta be SAFE: A New Framework for Molecular DesignarXiv 2023LLM Maybe LongLM: Self-Extend LLM Context Window Without TuningarXiv 2024DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM WorkflowsarXiv 2024Benchmarking Language Models for Code Syntax UnderstandingarXiv 2022A Simulation Benchmark for Autonomous Racing with Large-Scale Human DataarXiv 2024Composer: Creative and Controllable Image Synthesis with Composable ConditionsarXiv 2023Stabilize the Latent Space for Image Autoregressive Modeling: A Unified PerspectivearXiv 2024VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLMCVPR 2025 1DanceTrack: Multi-Object Tracking in Uniform Appearance and Diverse MotionCVPR 2022 1Harnessing Vision Models for Time Series Analysis: A SurveyarXiv 2025How do Large Language Models Handle Multilingualism?arXiv 2024A Closer Look into Automatic Evaluation Using Large Language ModelsarXiv 2023SVDFormer: Complementing Point Cloud via Self-view Augmentation and Self-structure Dual-generatorICCV 2023 1MonoScene: Monocular 3D Semantic Scene CompletionCVPR 2022 1TEXGen: a Generative Diffusion Model for Mesh TexturesarXiv 2024Prompt-Singer: Controllable Singing-Voice-Synthesis with Natural Language PromptarXiv 2024Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow MatchingarXiv 2024SlimSAM: 0.1% Data Makes Segment Anything SlimarXiv 2023CySecBench: Generative AI-based CyberSecurity-focused Prompt Dataset for Benchmarking Large Language ModelsarXiv 2025Collaborative Decoding Makes Visual Auto-Regressive Modeling EfficientCVPR 2025 1SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse AutoencodersarXiv 2025Query-Key Normalization for TransformersFindings of the Association for Computational Linguistics 2020USP: Unified Self-Supervised Pretraining for Image Generation and UnderstandingICCV 2025Parametric Classification for Generalized Category Discovery: A Baseline StudyICCV 2023 1PLA: Language-Driven Open-Vocabulary 3D Scene UnderstandingCVPR 2023 1Local All-Pair Correspondence for Point TrackingarXiv 2024Learning To Count EverythingCVPR 2021 1ZoomLDM: Latent Diffusion Model for multi-scale image generationCVPR 2025 1Raidar: geneRative AI Detection viA RewritingarXiv 2024pix2gestalt: Amodal Segmentation by Synthesizing WholesCVPR 2024 1AutoP2C: An LLM-Based Agent Framework for Code Repository Generation from Multimodal Content in Academic PapersarXiv 2025Visual Programming for Zero-shot Open-Vocabulary 3D Visual GroundingCVPR 2024 1GLACE: Global Local Accelerated Coordinate EncodingCVPR 2024 1LeanAgent: Lifelong Learning for Formal Theorem ProvingarXiv 2024Label-free Node Classification on Graphs with Large Language Models (LLMS)arXiv 2023Who is ChatGPT? Benchmarking LLMs' Psychological Portrayal Using PsychoBencharXiv 2023Large Language Models as Tool MakersarXiv 2023RISurConv: Rotation Invariant Surface Attention-Augmented Convolutions for 3D Point Cloud Classification and SegmentationarXiv 2024Towards Robust and Parameter-Efficient Knowledge Unlearning for LLMsarXiv 2024Efficient neural networks for real-time modeling of analog dynamic range compressionarXiv 2021LLM+P: Empowering Large Language Models with Optimal Planning ProficiencyarXiv 2023CLIP2Video: Mastering Video-Text Retrieval via Image CLIParXiv 2021SceneSmith: Agentic Generation of Simulation-Ready Indoor ScenesarXiv 2026Unified Perceptual Parsing for Scene Understandingunified-perceptual-parsing-for-scene-1Deep Graph Contrastive Representation LearningarXiv 2020Ouroboros3D: Image-to-3D Generation via 3D-aware Recursive DiffusionCVPR 2025 1FFHQ-UV: Normalized Facial UV-Texture Dataset for 3D Face ReconstructionCVPR 2023 1QuIP: 2-Bit Quantization of Large Language Models With GuaranteesNeurIPS 2023 11QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice CodebooksarXiv 2024PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy ReductionarXiv 2024Self-Detoxifying Language Models via Toxification ReversalarXiv 2023MARLIN: Masked Autoencoder for facial video Representation LearnINgCVPR 2023 1NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic ReasoningICCV 2025Generative Representational Instruction TuningarXiv 2024Anchored Preference Optimization and Contrastive Revisions: Addressing Underspecification in AlignmentarXiv 2024RLDX-1 Technical ReportarXiv 2026From Big to Small: Multi-Scale Local Planar Guidance for Monocular Depth EstimationarXiv 2019MAPF-GPT: Imitation Learning for Multi-Agent Pathfinding at Scalemapf-gpt-imitation-learning-for-multi-agentZigMa: A DiT-style Zigzag Mamba Diffusion ModelarXiv 2024CleanDIFT: Diffusion Features without NoiseCVPR 2025 1Cactus: Towards Psychological Counseling Conversations using Cognitive Behavioral TheoryarXiv 2024CodeRAG-Bench: Can Retrieval Augment Code Generation?arXiv 2024Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language ModelsarXiv 2024Reinforcing Multi-Turn Reasoning in LLM Agents via Fine-Grained Reward Structure and Credit AssignmentarXiv 2025Panoptic Segmentationpanoptic-segmentation-1SocialCircle: Learning the Angle-based Social Interaction Representation for Pedestrian Trajectory PredictionCVPR 2024 1Foresight -- Generative Pretrained Transformer (GPT) for Modelling of Patient Timelines using EHRsarXiv 2022Efficient Attention: Attention with Linear ComplexitiesarXiv 2018GaussianProperty: Integrating Physical Properties to 3D Gaussians with LMMsICCV 2025Hierarchical Prior Mining for Non-local Multi-View StereoICCV 2023 1TokenPacker: Efficient Visual Projector for Multimodal LLMarXiv 2024SupplyGraph: A Benchmark Dataset for Supply Chain Planning using Graph Neural NetworksarXiv 2024Enhancing Perceptual Quality in Video Super-Resolution through Temporally-Consistent Detail Synthesis using Diffusion ModelsarXiv 2023Integrating Earth Observation Data into Causal Inference: Challenges and OpportunitiesarXiv 2023MM-Soc: Benchmarking Multimodal Large Language Models in Social Media PlatformsarXiv 2024Interpreting Attention Layer Outputs with Sparse AutoencodersarXiv 2024Vector Quantized Diffusion Model for Text-to-Image SynthesisCVPR 2022 1UDC: A Unified Neural Divide-and-Conquer Framework for Large-Scale Combinatorial Optimization ProblemsarXiv 2024Sentence-level Prompts Benefit Composed Image RetrievalarXiv 2023Click: Controllable Text Generation with Sequence Likelihood Contrastive LearningarXiv 2023ATM: Adversarial Tuning Multi-agent System Makes a Robust Retrieval-Augmented GeneratorarXiv 2024Scalable Reinforcement Learning Policies for Multi-Agent ControlarXiv 2020Generalized and Efficient 2D Gaussian Splatting for Arbitrary-scale Super-ResolutionICCV 2025Synergy between 3DMM and 3D Landmarks for Accurate 3D Facial GeometryarXiv 2021Contrastive Difference Predictive CodingarXiv 2023Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning CapabilityarXiv 2024RuleRAG: Rule-guided retrieval-augmented generation with language models for question answeringarXiv 2024LLaGA: Large Language and Graph AssistantarXiv 2024LasUIE: Unifying Information Extraction with Latent Adaptive Structure-aware Generative Language ModelarXiv 2023Unifying Diffusion Models' Latent Space, with Applications to CycleDiffusion and GuidancearXiv 2022Res-VMamba: Fine-Grained Food Category Visual Classification Using Selective State Space Models with Deep Residual LearningarXiv 2024Marching-Primitives: Shape Abstraction from Signed Distance FunctionCVPR 2023 1SaMam: Style-aware State Space Model for Arbitrary Image Style TransferCVPR 2025 1FreeU: Free Lunch in Diffusion U-NetCVPR 2024 1FateZero: Fusing Attentions for Zero-shot Text-based Video EditingICCV 2023 1Holistically-Attracted Wireframe Parsing: From Supervised to Self-Supervised LearningarXiv 2022Spoken SQuAD: A Study of Mitigating the Impact of Speech Recognition Errors on Listening ComprehensionarXiv 2018Blind Video Deflickering by Neural Filtering with a Flawed AtlasCVPR 2023 1Learning Span-Level Interactions for Aspect Sentiment Triplet ExtractionACL 2021 5Debiasing Vision-Language Models via Biased PromptsarXiv 2023Debiased Contrastive LearningNeurIPS 2020 12ChangeMamba: Remote Sensing Change Detection With Spatiotemporal State Space ModelarXiv 2024Denoising Likelihood Score Matching for Conditional Score-based Data Generationdenoising-likelihood-score-matching-forBRIGHT: A globally distributed multimodal building damage assessment dataset with very-high-resolution for all-weather disaster responsearXiv 2025DeMamba: AI-Generated Video Detection on Million-Scale GenVideo BenchmarkarXiv 2024WeatherQA: Can Multimodal Language Models Reason about Severe Weather?arXiv 2024Motion Anything: Any to Motion GenerationarXiv 2025PAL: Persona-Augmented Emotional Support Conversation GenerationarXiv 2022Benchmarking Multi-modal Semantic Segmentation under Sensor Failures: Missing and Noisy Modality RobustnessarXiv 2025What can Large Language Models do in chemistry? A comprehensive benchmark on eight tasksNeurIPS 2023 11Zero-Shot ECG Classification with Multimodal Learning and Test-time Clinical Knowledge EnhancementarXiv 2024Robust Scene Change Detection Using Visual Foundation Models and Cross-Attention MechanismsarXiv 2024Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache ConsumptionarXiv 2024Neural-PIL: Neural Pre-Integrated Lighting for Reflectance DecompositionNeurIPS 2021 12NeRD: Neural Reflectance Decomposition from Image CollectionsICCV 2021 10Parallelized Autoregressive Visual GenerationCVPR 2025 1BossNAS: Exploring Hybrid CNN-transformers with Block-wisely Self-supervised Neural Architecture SearchICCV 2021 10Frustum PointNets for 3D Object Detection from RGB-D Datafrustum-pointnets-for-3d-object-detection-1Reliable and Efficient Concept Erasure of Text-to-Image Diffusion ModelsarXiv 2024Semi-Supervised Semantic Segmentation with Cross Pseudo SupervisionCVPR 2021 1Parameter-Efficient Fine-Tuning with Discrete Fourier TransformarXiv 2024Q-VLM: Post-training Quantization for Large Vision-Language ModelsarXiv 2024Identifying Linear Relational Concepts in Large Language ModelsarXiv 2023Self-Supervised Video Forensics by Audio-Visual Anomaly DetectionCVPR 2023 1Iterative Refinement of Project-Level Code Context for Precise Code Generation with Compiler FeedbackarXiv 2024This Looks Like That: Deep Learning for Interpretable Image Recognitionthis-looks-like-that-deep-learning-for-1Fed-SB: A Silver Bullet for Extreme Communication Efficiency and Performance in (Private) Federated LoRA Fine-TuningarXiv 2025AdaMoE: Token-Adaptive Routing with Null Experts for Mixture-of-Experts Language ModelsarXiv 2024SQLFixAgent: Towards Semantic-Accurate Text-to-SQL Parsing via Consistency-Enhanced Multi-Agent CollaborationarXiv 2024Mastering Memory Tasks with World ModelsarXiv 2024A Simple LLM Framework for Long-Range Video Question-AnsweringarXiv 2023Question Decomposition Tree for Answering Complex Questions over Knowledge BasesarXiv 2023MarkQA: A large scale KBQA dataset with numerical reasoningarXiv 2023ALOcc: Adaptive Lifting-based 3D Semantic Occupancy and Cost Volume-based Flow PredictionICCV 2025CLadder: Assessing Causal Reasoning in Language Modelscladder-a-benchmark-to-assess-causalWhat the DAAM: Interpreting Stable Diffusion Using Cross AttentionarXiv 2022Causal Diffusion Transformers for Generative ModelingarXiv 2024Learning Multi-modal Representations by Watching Hundreds of Surgical Video LecturesarXiv 2023Fairer Preferences Elicit Improved Human-Aligned Large Language Model JudgmentsarXiv 2024FlexiAct: Towards Flexible Action Control in Heterogeneous ScenariosarXiv 2025Deep Equilibrium Diffusion Restoration with Parallel SamplingCVPR 2024 1Virtual Personas for Language Models via an Anthology of BackstoriesarXiv 2024What Matters in Transformers? Not All Attention is NeededarXiv 2024TweetEval: Unified Benchmark and Comparative Evaluation for Tweet ClassificationFindings of the Association for Computational Linguistics 2020Learning to generate line drawings that convey geometry and semanticsCVPR 2022 1Everybody Dance Noweverybody-dance-now-1Regional Attention for Shadow RemovalarXiv 2024HybridDepth: Robust Metric Depth Fusion by Leveraging Depth from Focus and Single-Image PriorsarXiv 2024MedIAnomaly: A comparative study of anomaly detection in medical imagesarXiv 2024Adaptation of Whisper models to child speech recognitionarXiv 2023vesselFM: A Foundation Model for Universal 3D Blood Vessel SegmentationCVPR 2025 1Towards Real-World Prohibited Item Detection: A Large-Scale X-ray BenchmarkICCV 2021 10CoLLaVO: Crayon Large Language and Vision mOdelarXiv 2024Phantom of Latent for Large Language and Vision ModelsarXiv 2024LM4LV: A Frozen Large Language Model for Low-level Vision TasksarXiv 2024Valley2: Exploring Multimodal Models with Scalable Vision-Language DesignarXiv 2025Shot2Story20K: A New Benchmark for Comprehensive Understanding of Multi-shot VideosarXiv 2023RealHiTBench: A Comprehensive Realistic Hierarchical Table Benchmark for Evaluating LLM-Based Table AnalysisarXiv 2025MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning AttentionarXiv 2025DynaGuide: Steering Diffusion Polices with Active Dynamic GuidancearXiv 2025xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World EvaluationsarXiv 2025MotionPro: A Precise Motion Controller for Image-to-Video GenerationCVPR 2025 1ReasonPlan: Unified Scene Prediction and Decision Reasoning for Closed-loop Autonomous DrivingarXiv 2025The Avengers: A Simple Recipe for Uniting Smaller Language Models to Challenge Proprietary GiantsarXiv 2025LLVIP: A Visible-infrared Paired Dataset for Low-light VisionarXiv 2021