All papers
Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.
Real-Time Scene Text Detection with Differentiable Binarization and Adaptive Scale FusionarXiv 2022DialoGPT: Large-Scale Generative Pre-training for Conversational Response GenerationarXiv 2019Audio Mamba: Bidirectional State Space Model for Audio Representation LearningarXiv 2024Neural Video Compression with Feature ModulationCVPR 2024 1CvT: Introducing Convolutions to Vision TransformersICCV 2021 10The GigaMIDI Dataset with Features for Expressive Music Performance DetectionarXiv 2025DexTrack: Towards Generalizable Neural Tracking Control for Dexterous Manipulation from Human ReferencesarXiv 2025Edit Temporal-Consistent Videos with Image Diffusion ModelarXiv 2023Subgraph-Aware Training of Language Models for Knowledge Graph Completion Using Structure-Aware Contrastive LearningarXiv 2024MapCoder: Multi-Agent Code Generation for Competitive Problem SolvingarXiv 2024HiDe-LLaVA: Hierarchical Decoupling for Continual Instruction Tuning of Multimodal Large Language ModelarXiv 2025Dual-Expert Consistency Model for Efficient and High-Quality Video GenerationICCV 20253DSAM-adapter: Holistic adaptation of SAM from 2D to 3D for promptable tumor segmentationarXiv 2023Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM ReasoningarXiv 2025FlexWorld: Progressively Expanding 3D Scenes for Flexiable-View SynthesisarXiv 2025Native-Resolution Image SynthesisarXiv 2025MindEye2: Shared-Subject Models Enable fMRI-To-Image With 1 Hour of DataarXiv 2024Twins: Revisiting the Design of Spatial Attention in Vision TransformersNeurIPS 2021 12SparseBEV: High-Performance Sparse 3D Object Detection from Multi-Camera VideosICCV 2023 14D Panoptic LiDAR SegmentationCVPR 2021 1MeMOTR: Long-Term Memory-Augmented Transformer for Multi-Object TrackingICCV 2023 1MedMNIST v2 -- A large-scale lightweight benchmark for 2D and 3D biomedical image classificationarXiv 2021Extracting Motion and Appearance via Inter-Frame Attention for Efficient Video Frame InterpolationCVPR 2023 1Deep Equilibrium Object DetectionICCV 2023 1WebLINX: Real-World Website Navigation with Multi-Turn DialoguearXiv 2024DOT: A Distillation-Oriented TrainerICCV 2023 1Practical Stereo Matching via Cascaded Recurrent Network with Adaptive CorrelationCVPR 2022 1Evaluating Correctness and Faithfulness of Instruction-Following Models for Question AnsweringarXiv 2023VideoGPT+: Integrating Image and Video Encoders for Enhanced Video UnderstandingarXiv 2024UniMed-CLIP: Towards a Unified Image-Text Pretraining Paradigm for Diverse Medical Imaging ModalitiesarXiv 2024G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement LearningarXiv 2025GeoPixel: Pixel Grounding Large Multimodal Model in Remote SensingarXiv 2025All Languages Matter: Evaluating LMMs on Culturally Diverse 100 LanguagesCVPR 2025 1OpenFactCheck: A Unified Framework for Factuality Evaluation of LLMsarXiv 2024Learning to See by Looking at NoiseNeurIPS 2021 12GenRL: Multimodal-foundation world models for generalization in embodied agentsarXiv 2024Generative Action Description Prompts for Skeleton-based Action RecognitionICCV 2023 1Do Vision and Language Encoders Represent the World Similarly?CVPR 2024 1MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical CodearXiv 2024TEMOS: Generating diverse human motions from textual descriptionsarXiv 2022Follow-Your-Canvas: Higher-Resolution Video Outpainting with Extensive Content GenerationarXiv 2024Feynman-Kac Correctors in Diffusion: Annealing, Guidance, and Product of ExpertsarXiv 2025From Zero to Turbulence: Generative Modeling for 3D Flow SimulationarXiv 2023Understanding and Improving Knowledge Distillation for Quantization-Aware Training of Large Transformer EncodersarXiv 2022From an Image to a Scene: Learning to Imagine the World from a Million 360 VideosarXiv 2024Efficient Modulation for Vision NetworksarXiv 2024Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and ReactionCVPR 2025 1XMem++: Production-level Video Segmentation From Few Annotated FramesICCV 2023 1Brouhaha: multi-task training for voice activity detection, speech-to-noise ratio, and C50 room acoustics estimationarXiv 2022HEST-1k: A Dataset for Spatial Transcriptomics and Histology Image AnalysisarXiv 2024Neural Face Identification in a 2D Wireframe Projection of a Manifold Objectneural-face-identification-in-a-2d-wireframePretraining task diversity and the emergence of non-Bayesian in-context learning for regressionpretraining-task-diversity-and-the-emergencePLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioningpllava-parameter-free-llava-extension-fromLightningDrag: Lightning Fast and Accurate Drag-based Image Editing Emerging from VideosarXiv 2024RoCo: Dialectic Multi-Robot Collaboration with Large Language ModelsarXiv 2023Compact 3D Gaussian Representation for Radiance FieldCVPR 2024 1RaTEScore: A Metric for Radiology Report GenerationarXiv 2024DDSP: Differentiable Digital Signal ProcessingICLR 2020 1Do Large Language Model Benchmarks Test Reliability?arXiv 2025TRAK: Attributing Model Behavior at ScalearXiv 2023QuoTA: Query-oriented Token Assignment via CoT Query Decouple for Long Video ComprehensionarXiv 2025PsycoLLM: Enhancing LLM for Psychological Understanding and EvaluationarXiv 2024PoseBusters: AI-based docking methods fail to generate physically valid poses or generalise to novel sequencesarXiv 2023Towards RAW Object Detection in Diverse ConditionsCVPR 2025 1CPPO: Accelerating the Training of Group Relative Policy Optimization-Based Reasoning ModelsarXiv 2025Implicit In-context LearningarXiv 2024EmoLLMs: A Series of Emotional Large Language Models and Annotation Tools for Comprehensive Affective AnalysisarXiv 2024Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text IntegrationarXiv 2023LCM-LoRA: A Universal Stable-Diffusion Acceleration ModulearXiv 2023CoachLM: Automatic Instruction Revisions Improve the Data Quality in LLM Instruction TuningarXiv 2023FoodieQA: A Multimodal Dataset for Fine-Grained Understanding of Chinese Food CulturearXiv 2024OMG-Seg: Is One Model Good Enough For All Segmentation?CVPR 2024 1SwinFace: A Multi-task Transformer for Face Recognition, Expression Recognition, Age Estimation and Attribute EstimationarXiv 2023Pseudo Numerical Methods for Diffusion Models on Manifoldspseudo-numerical-methods-for-diffusion-modelsRobust Latent Matters: Boosting Image Generation with Sampling ErrorarXiv 2025BizFinBench: A Business-Driven Real-World Financial Benchmark for Evaluating LLMsarXiv 2025ControlVAR: Exploring Controllable Visual Autoregressive ModelingarXiv 2024Improving Multi-turn Emotional Support Dialogue Generation with Lookahead Strategy PlanningarXiv 2022One Map to Find Them All: Real-time Open-Vocabulary Mapping for Zero-shot Multi-Object NavigationarXiv 2024The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"arXiv 2023Guess What I Think: Streamlined EEG-to-Image Generation with Latent Diffusion ModelsarXiv 2024Recent Advances of Multimodal Continual Learning: A Comprehensive SurveyarXiv 2024Learning Fine-Grained Grounded Citations for Attributed Large Language ModelsarXiv 2024A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open QuestionsarXiv 2023Memorizing Transformersmemorizing-transformersStructured State Space Models for In-Context Reinforcement Learningstructured-state-space-models-for-in-contextReturn of Unconditional Generation: A Self-supervised Representation Generation MethodarXiv 2023SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code AgentsarXiv 2024Idiosyncrasies in Large Language ModelsarXiv 2025Prompt-In-Prompt Learning for Universal Image RestorationarXiv 2023OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task SynthesisarXiv 2024Empowering Cross-lingual Abilities of Instruction-tuned Large Language Models by Translation-following demonstrationsarXiv 2023Learning to Estimate 3D Hand Pose from Single RGB Imageslearning-to-estimate-3d-hand-pose-from-single-1TempCompass: Do Video LLMs Really Understand Videos?arXiv 2024SAM-DiffSR: Structure-Modulated Diffusion Model for Image Super-ResolutionarXiv 2024VIKI-R: Coordinating Embodied Multi-Agent Cooperation via Reinforcement LearningarXiv 2025MIRAGE: Multimodal foundation model and benchmark for comprehensive retinal OCT image analysisarXiv 2025FaithfulRAG: Fact-Level Conflict Modeling for Context-Faithful Retrieval-Augmented GenerationarXiv 2025MegaMath: Pushing the Limits of Open Math CorporaarXiv 2025LLaVA-Plus: Learning to Use Tools for Creating Multimodal AgentsarXiv 2023Asymmetric Mask Scheme for Self-Supervised Real Image DenoisingarXiv 2024COSMOS: A Hybrid Adaptive Optimizer for Memory-Efficient Training of LLMsarXiv 2025CLIP-Driven Universal Model for Organ Segmentation and Tumor DetectionICCV 2023 1Fréchet Video Motion Distance: A Metric for Evaluating Motion Consistency in VideosarXiv 2024FaceVerse: a Fine-grained and Detail-controllable 3D Face Morphable Model from a Hybrid DatasetCVPR 2022 1StyleAvatar: Real-time Photo-realistic Portrait Avatar from a Single VideoarXiv 2023LatestEval: Addressing Data Contamination in Language Model Evaluation through Dynamic and Time-Sensitive Test ConstructionarXiv 2023Graph Diffusion Transformers for Multi-Conditional Molecular GenerationarXiv 2024Federated Optimization in Heterogeneous NetworksarXiv 2018Infini-gram: Scaling Unbounded n-gram Language Models to a Trillion TokensarXiv 2024Du-IN: Discrete units-guided mask modeling for decoding speech from Intracranial Neural signalsarXiv 2024Reduce Information Loss in Transformers for Pluralistic Image InpaintingCVPR 2022 1MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMsarXiv 2024Text-Guided Texturing by Synchronized Multi-View DiffusionarXiv 2023Instance Segmentation in the DarkarXiv 2023Dialogue Action Tokens: Steering Language Models in Goal-Directed Dialogue with a Multi-Turn PlannerarXiv 2024Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelNeurIPS 2023 11Dynamic Mixture of Experts: An Auto-Tuning Approach for Efficient Transformer ModelsarXiv 2024Ray Denoising: Depth-aware Hard Negative Sampling for Multi-view 3D Object DetectionarXiv 2024Rec-R1: Bridging Generative Large Language Models and User-Centric Recommendation Systems via Reinforcement LearningarXiv 2025AutoAgents: A Framework for Automatic Agent GenerationarXiv 2023NoisyRollout: Reinforcing Visual Reasoning with Data AugmentationarXiv 2025Towards Visual Grounding: A SurveyarXiv 2024Against The Achilles' Heel: A Survey on Red Teaming for Generative ModelsarXiv 2024Kanana: Compute-efficient Bilingual Language ModelsarXiv 2025Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMsarXiv 2023OneChart: Purify the Chart Structural Extraction via One Auxiliary TokenarXiv 2024Evaluation of Segment Anything Model 2: The Role of SAM2 in the Underwater EnvironmentarXiv 2024Universal Score-based Speech Enhancement with High Content PreservationarXiv 2024LogEval: A Comprehensive Benchmark Suite for Large Language Models In Log AnalysisarXiv 2024RORem: Training a Robust Object Remover with Human-in-the-LoopCVPR 2025 1DAFormer: Improving Network Architectures and Training Strategies for Domain-Adaptive Semantic SegmentationCVPR 2022 1SegEarth-OV: Towards Training-Free Open-Vocabulary Segmentation for Remote Sensing ImagesCVPR 2025 1Hier-SLAM: Scaling-up Semantics in SLAM with a Hierarchically Categorical Gaussian SplattingarXiv 2024Panoptic Video Scene Graph Generationpanoptic-video-scene-graph-generationMultiMed: Multilingual Medical Speech Recognition via Attention Encoder DecoderarXiv 2024Blockwise Parallel Transformer for Large Context ModelsarXiv 2023A Survey on Hallucination in Large Vision-Language ModelsarXiv 2024Detecting Line Segments in Motion-blurred Images with EventsarXiv 2022Petals: Collaborative Inference and Fine-tuning of Large ModelsarXiv 2022learn2learn: A Library for Meta-Learning ResearcharXiv 2020ChaosBench: A Multi-Channel, Physics-Based Benchmark for Subseasonal-to-Seasonal Climate PredictionarXiv 2024SafeGen: Mitigating Sexually Explicit Content Generation in Text-to-Image ModelsarXiv 2024Agent Attention: On the Integration of Softmax and Linear AttentionarXiv 2023DAT++: Spatially Dynamic Vision Transformer with Deformable AttentionarXiv 2023Accelerating Scientific Discovery with Generative Knowledge Extraction, Graph-Based Representation, and Multimodal Intelligent Graph ReasoningarXiv 2024Comparing Retrieval-Augmentation and Parameter-Efficient Fine-Tuning for Privacy-Preserving Personalization of Large Language ModelsarXiv 2024SciAgents: Automating scientific discovery through multi-agent intelligent graph reasoningarXiv 2024Brain Latent Progression: Individual-based Spatiotemporal Disease Progression on 3D Brain MRIs via Latent DiffusionarXiv 2025Ranger21: a synergistic deep learning optimizerarXiv 2021Delving into Masked Autoencoders for Multi-Label Thorax Disease ClassificationarXiv 2022AGIQA-3K: An Open Database for AI-Generated Image Quality AssessmentarXiv 2023A Comprehensive Survey on Long Context Language ModelingarXiv 2025Modeling Long- and Short-Term Temporal Patterns with Deep Neural NetworksarXiv 2017ReMamba: Equip Mamba with Effective Long-Sequence ModelingarXiv 2024FastMoE: A Fast Mixture-of-Expert Training SystemarXiv 2021AGLA: Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local AttentionarXiv 2024Transparent Image Layer Diffusion using Latent TransparencyarXiv 2024No More Strided Convolutions or Pooling: A New CNN Building Block for Low-Resolution Images and Small ObjectsarXiv 2022AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and PruningICCV 2025DurLAR: A High-fidelity 128-channel LiDAR Dataset with Panoramic Ambient and Reflectivity Imagery for Multi-modal Autonomous Driving Applicationsdurlar-a-high-fidelity-128-channel-lidarHigh-Fidelity Simultaneous Speech-To-Speech TranslationarXiv 2025CSS10: A Collection of Single Speaker Speech Datasets for 10 LanguagesarXiv 2019Large Language Models Know Your Contextual Search Intent: A Prompting Framework for Conversational SearcharXiv 2023Language Model Analysis for Ontology Subsumption InferencearXiv 2023AvatarMe++: Facial Shape and BRDF Inference with Photorealistic Rendering-Aware GANsarXiv 2021Endo-4DGS: Endoscopic Monocular Scene Reconstruction with 4D Gaussian SplattingarXiv 2024A is for Absorption: Studying Feature Splitting and Absorption in Sparse AutoencodersarXiv 2024Understanding Expressivity of GNN in Rule LearningarXiv 2023Cautious Optimizers: Improving Training with One Line of CodearXiv 2024A Large-Scale Benchmark for Food Image SegmentationarXiv 2021Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning TasksarXiv 2024Robust Prompt Optimization for Defending Language Models Against Jailbreaking AttacksarXiv 2024Time Travelling Pixels: Bitemporal Features Integration with Foundation Model for Remote Sensing Image Change DetectionarXiv 2023Evading Forensic Classifiers with Attribute-Conditioned Adversarial Facesevading-forensic-classifiers-with-attributeVoice Disorder Analysis: a Transformer-based ApproacharXiv 2024DRT-o1: Optimized Deep Reasoning Translation via Long Chain-of-ThoughtarXiv 2024Skip Tuning: Pre-trained Vision-Language Models are Effective and Efficient Adapters ThemselvesCVPR 2025 1Medical World Model: Generative Simulation of Tumor Evolution for Treatment PlanningarXiv 2025SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse ViewpointsarXiv 2024One Step Diffusion via Shortcut ModelsarXiv 2024Exploiting Foundation Models and Speech Enhancement for Parkinson's Disease Detection from Speech in Real-World Operative ConditionsarXiv 2024Large Language Models are Zero-Shot ReasonersarXiv 2022From Pixels to Insights: A Survey on Automatic Chart Understanding in the Era of Large Foundation ModelsarXiv 2024Effective control of two-dimensional Rayleigh--Bénard convection: invariant multi-agent reinforcement learning is all you needarXiv 2023SeFlow: A Self-Supervised Scene Flow Method in Autonomous DrivingarXiv 2024Designing BERT for Convolutional Networks: Sparse and Hierarchical Masked ModelingarXiv 2023EdgeGaussians -- 3D Edge Mapping via Gaussian SplattingarXiv 2024Ref-Diff: Zero-shot Referring Image Segmentation with Generative ModelsarXiv 2023KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for CodingarXiv 2025Mimic In-Context Learning for Multimodal TasksCVPR 2025 1Caduceus: Bi-Directional Equivariant Long-Range DNA Sequence ModelingarXiv 2024Visual Prompt TuningarXiv 2022KLUE: Korean Language Understanding EvaluationarXiv 2021T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video GenerationCVPR 2025 1DWIE: an entity-centric dataset for multi-task document-level information extractionarXiv 2020DyCoke: Dynamic Compression of Tokens for Fast Video Large Language ModelsCVPR 2025 1The CoT Collection: Improving Zero-shot and Few-shot Learning of Language Models via Chain-of-Thought Fine-TuningarXiv 2023Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language ModelsarXiv 2024VCISR: Blind Single Image Super-Resolution with Video Compression Synthetic DataarXiv 2023