0

All papers

Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.

MetaFormer: A Unified Meta Framework for Fine-Grained RecognitionarXiv 2022BEVBert: Multimodal Map Pre-training for Language-guided NavigationICCV 2023 1Rectified Diffusion: Straightness Is Not Your Need in Rectified FlowarXiv 2024Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step QuestionsarXiv 2022REaLTabFormer: Generating Realistic Relational and Tabular Data using TransformersarXiv 2023Using Human Feedback to Fine-tune Diffusion Models without Any Reward ModelCVPR 2024 1PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal RetrieversarXiv 2024LST: Ladder Side-Tuning for Parameter and Memory Efficient Transfer LearningarXiv 2022FlashSplat: 2D to 3D Gaussian Splatting Segmentation Solved OptimallyarXiv 2024MeshSDF: Differentiable Iso-Surface ExtractionNeurIPS 2020 12Music2Latent: Consistency Autoencoders for Latent Audio CompressionarXiv 2024Temporal Graph Benchmark for Machine Learning on Temporal Graphstemporal-graph-benchmark-for-machine-learningDDDM-VC: Decoupled Denoising Diffusion Models with Disentangled Representation and Prior Mixup for Verified Robust Voice ConversionarXiv 2023Global Features are All You Need for Image Retrieval and RerankingICCV 2023 1Self-Distilled StyleGAN: Towards Generation from Internet PhotosarXiv 2022Unsupervised CNN for Single View Depth Estimation: Geometry to the RescuearXiv 2016DriveAdapter: Breaking the Coupling Barrier of Perception and Planning in End-to-End Autonomous DrivingICCV 2023 13DHumanGAN: 3D-Aware Human Image Generation with 3D Pose MappingICCV 2023 1MotionStreamer: Streaming Motion Generation via Diffusion-based Autoregressive Model in Causal Latent SpaceICCV 2025PsyQA: A Chinese Dataset for Generating Long Counseling Text for Mental Health SupportFindings (ACL) 2021 8Panoptic SegFormer: Delving Deeper into Panoptic Segmentation with TransformersCVPR 2022 1GreaseLM: Graph REASoning Enhanced Language Models for Question AnsweringarXiv 2022From Knowledge Distillation to Self-Knowledge Distillation: A Unified Approach with Normalized Loss and Customized Soft LabelsICCV 2023 1SweetDreamer: Aligning Geometric Priors in 2D Diffusion for Consistent Text-to-3DarXiv 2023One for All: Towards Training One Graph Model for All Classification TasksarXiv 2023A Simple Language Model for Task-Oriented DialogueNeurIPS 2020 12Reformulating Unsupervised Style Transfer as Paraphrase GenerationEMNLP 2020 11Towards Foundational Models for Molecular Learning on Large-Scale Multi-Task DatasetsarXiv 2023Neural Prompt SearcharXiv 2022Diffusion Models are Evolutionary AlgorithmsarXiv 2024Keypoint Promptable Re-IdentificationarXiv 2024Language Models can Solve Computer TasksNeurIPS 2023 11A Text-guided Protein Design FrameworkarXiv 2023TableLLM: Enabling Tabular Data Manipulation by LLMs in Real Office Usage ScenariosarXiv 2024ABQ-LLM: Arbitrary-Bit Quantized Inference Acceleration for Large Language ModelsarXiv 2024EfficientVMamba: Atrous Selective Scan for Light Weight Visual MambaarXiv 2024Context-aware Feature Generation for Zero-shot Semantic SegmentationarXiv 2020PUG: Photorealistic and Semantically Controllable Synthetic Data for Representation Learningpug-photorealistic-and-semanticallyProactive Interaction Framework for Intelligent Social Receptionist RobotsarXiv 2020Learning 3D Human Shape and Pose from Dense Body PartsarXiv 2019MANTIS: Interleaved Multi-Image Instruction TuningarXiv 2024UniRef++: Segment Every Reference Object in Spatial and Temporal SpacesarXiv 2023Tiny and Efficient Model for the Edge Detection GeneralizationarXiv 2023Monash Time Series Forecasting ArchivearXiv 2021DartControl: A Diffusion-Based Autoregressive Motion Model for Real-Time Text-Driven Motion ControlarXiv 2024Neural Message Passing for Quantum Chemistryneural-message-passing-for-quantum-chemistry-1Long-LRM: Long-sequence Large Reconstruction Model for Wide-coverage Gaussian SplatsarXiv 2024Multi-Head RAG: Solving Multi-Aspect Problems with LLMsarXiv 2024EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt OptimizersarXiv 2023Agentic Deep Graph Reasoning Yields Self-Organizing Knowledge NetworksarXiv 2025FaceTalk: Audio-Driven Motion Diffusion for Neural Parametric Head ModelsCVPR 2024 1SODA: Million-scale Dialogue Distillation with Social Commonsense ContextualizationarXiv 2022BioCLIP: A Vision Foundation Model for the Tree of LifeCVPR 2024 1HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-trainingEMNLP 2020 11The ArtBench Dataset: Benchmarking Generative Models with ArtworksarXiv 2022LoftQ: LoRA-Fine-Tuning-Aware Quantization for Large Language ModelsarXiv 2023Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image SynthesisCVPR 2024 1Small LLMs Are Weak Tool Learners: A Multi-LLM AgentarXiv 2024Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of GeneralizationarXiv 2024Q-Instruct: Improving Low-level Visual Abilities for Multi-modality Foundation ModelsCVPR 2024 1Bridging Textual and Tabular Data for Cross-Domain Text-to-SQL Semantic ParsingFindings of the Association for Computational Linguistics 2020Co-Scale Conv-Attentional Image TransformersICCV 2021 10AERO: Audio Super Resolution in the Spectral DomainarXiv 2022Interpreting CLIP's Image Representation via Text-Based DecompositionarXiv 2023DiC: Rethinking Conv3x3 Designs in Diffusion ModelsCVPR 2025 1Language-Driven Representation Learning for RoboticsarXiv 2023xBD: A Dataset for Assessing Building Damage from Satellite ImageryarXiv 2019MoMA: Multimodal LLM Adapter for Fast Personalized Image GenerationarXiv 2024UDiffText: A Unified Framework for High-quality Text Synthesis in Arbitrary Images via Character-aware Diffusion ModelsarXiv 2023HPNet: Dynamic Trajectory Forecasting with Historical Prediction AttentionCVPR 2024 1Efficient Inference for Large Reasoning Models: A SurveyarXiv 2025A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural NetworksarXiv 2016UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight DetectionCVPR 2022 1FluidLab: A Differentiable Environment for Benchmarking Complex Fluid ManipulationarXiv 2023AvaTaR: Optimizing LLM Agents for Tool Usage via Contrastive ReasoningarXiv 2024Valley: Video Assistant with Large Language model Enhanced abilitYarXiv 2023AutoAct: Automatic Agent Learning from Scratch for QA via Self-PlanningarXiv 2024Distribution Matching for Crowd CountingNeurIPS 2020 12Segment and Caption AnythingCVPR 2024 1UnIVAL: Unified Model for Image, Video, Audio and Language TasksarXiv 2023Nerfbusters: Removing Ghostly Artifacts from Casually Captured NeRFsICCV 2023 1ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical ReasoningFindings (ACL) 2022 5ControlVideo: Conditional Control for One-shot Text-driven Video Editing and BeyondarXiv 2023Gather-Excite: Exploiting Feature Context in Convolutional Neural Networksgather-excite-exploiting-feature-context-in-1Lighting Every Darkness with 3DGS: Fast Training and Real-Time Rendering for HDR View SynthesisarXiv 2024Free4D: Tuning-free 4D Scene Generation with Spatial-Temporal ConsistencyICCV 2025OpenKiwi: An Open Source Framework for Quality Estimationopenkiwi-an-open-source-framework-for-quality-1StructureFlow: Image Inpainting via Structure-aware Appearance Flowstructureflow-image-inpainting-via-structure-1GraspXL: Generating Grasping Motions for Diverse Objects at ScalearXiv 2024VideoTetris: Towards Compositional Text-to-Video GenerationarXiv 2024Cleaner Pretraining Corpus Curation with Neural Web ScrapingarXiv 2024LexGLUE: A Benchmark Dataset for Legal Language Understanding in EnglishACL 2022 5GES: Generalized Exponential Splatting for Efficient Radiance Field RenderingarXiv 2024Dynamic Cheatsheet: Test-Time Learning with Adaptive MemoryarXiv 2025PMC-CLIP: Contrastive Language-Image Pre-training using Biomedical DocumentsarXiv 2023OptiMUS: Scalable Optimization Modeling with (MI)LP Solvers and Large Language ModelsarXiv 2024HiFTNet: A Fast High-Quality Neural Vocoder with Harmonic-plus-Noise Filter and Inverse Short Time Fourier TransformarXiv 2023BigTranslate: Augmenting Large Language Models with Multilingual Translation Capability over 100 LanguagesarXiv 2023LoRA+: Efficient Low Rank Adaptation of Large ModelsarXiv 2024Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation ModelsarXiv 2024I2V-Adapter: A General Image-to-Video Adapter for Diffusion ModelsarXiv 2023DiM: Diffusion Mamba for Efficient High-Resolution Image SynthesisarXiv 2024Unsupervised Universal Image SegmentationCVPR 2024 1U-DiTs: Downsample Tokens in U-Shaped Diffusion TransformersarXiv 2024HateXplain: A Benchmark Dataset for Explainable Hate Speech DetectionarXiv 2020KUIELab-MDX-Net: A Two-Stream Neural Network for Music DemixingarXiv 2021OneRestore: A Universal Restoration Framework for Composite DegradationarXiv 2024ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use CapabilitiesarXiv 2024Chain of Hindsight Aligns Language Models with FeedbackarXiv 2023WildDeepfake: A Challenging Real-World Dataset for Deepfake DetectionarXiv 2021BiLLM: Pushing the Limit of Post-Training Quantization for LLMsarXiv 2024GNeRF: GAN-based Neural Radiance Field without Posed CameraICCV 2021 10GameFormer: Game-theoretic Modeling and Learning of Transformer-based Interactive Prediction and Planning for Autonomous DrivingarXiv 2023Automatic Data Curation for Self-Supervised Learning: A Clustering-Based ApproacharXiv 2024Vision Transformer with Quadrangle AttentionarXiv 2023Learning to Use Tools via Cooperative and Interactive AgentsarXiv 2024DrivingWorld: Constructing World Model for Autonomous Driving via Video GPTarXiv 2024GFTE: Graph-based Financial Table ExtractionarXiv 2020DynamicStereo: Consistent Dynamic Depth from Stereo VideosCVPR 2023 1Neural Architecture Design for GPU-Efficient NetworksarXiv 2020Multimodal Prompting with Missing Modalities for Visual RecognitionCVPR 2023 1Model and Data Transfer for Cross-Lingual Sequence Labelling in Zero-Resource SettingsarXiv 2022HIT-UAV: A high-altitude infrared thermal dataset for Unmanned Aerial Vehicle-based object detectionarXiv 2022WRENCH: A Comprehensive Benchmark for Weak SupervisionarXiv 2021ModuleFormer: Modularity Emerges from Mixture-of-ExpertsarXiv 2023TACO: Topics in Algorithmic COde generation datasetarXiv 2023LLaRA: Supercharging Robot Learning Data for Vision-Language PolicyarXiv 2024Text-Driven Image Editing via Learnable RegionsCVPR 2024 1SeqGPT: An Out-of-the-box Large Language Model for Open Domain Sequence UnderstandingarXiv 2023Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Trainingdeep-gradient-compression-reducing-the-1Constrained Decision Transformer for Offline Safe Reinforcement LearningarXiv 2023WorkArena++: Towards Compositional Planning and Reasoning-based Common Knowledge Work TasksarXiv 2024BYOL for Audio: Self-Supervised Learning for General-Purpose Audio RepresentationarXiv 2021Soundwave: Less is More for Speech-Text Alignment in LLMsarXiv 2025Learning 3D Representations from 2D Pre-trained Models via Image-to-Point Masked AutoencodersCVPR 2023 1Deep Fashion3D: A Dataset and Benchmark for 3D Garment Reconstruction from Single ImagesECCV 2020 8Retrieval Head Mechanistically Explains Long-Context FactualityarXiv 2024DYffusion: A Dynamics-informed Diffusion Model for Spatiotemporal Forecastingdyffusion-a-dynamics-informed-diffusion-modelHawk: Learning to Understand Open-World Video AnomaliesarXiv 2024Ultra-High-Definition Low-Light Image Enhancement: A Benchmark and Transformer-Based MethodarXiv 2022Robust Multiview Point Cloud Registration with Reliable Pose Graph Initialization and History ReweightingCVPR 2023 1Controlled Text Generation via Language Model ArithmeticarXiv 2023SMPLest-X: Ultimate Scaling for Expressive Human Pose and Shape EstimationarXiv 2025DreamPhysics: Learning Physics-Based 3D Dynamics with Video Diffusion PriorsarXiv 2024Reconstructing Personalized Semantic Facial NeRF Models From Monocular VideoarXiv 2022PVO: Panoptic Visual OdometryCVPR 2023 1ImageInWords: Unlocking Hyper-Detailed Image DescriptionsarXiv 2024PMC-VQA: Visual Instruction Tuning for Medical Visual Question AnsweringarXiv 2023CLIP Itself is a Strong Fine-tuner: Achieving 85.7% and 88.0% Top-1 Accuracy with ViT-B and ViT-L on ImageNetarXiv 202275 Languages, 1 Model: Parsing Universal Dependencies Universally75-languages-1-model-parsing-universal-1On Evaluating Adversarial Robustness of Large Vision-Language ModelsNeurIPS 2023 11RETA-LLM: A Retrieval-Augmented Large Language Model ToolkitarXiv 2023SceneCraft: Layout-Guided 3D Scene GenerationarXiv 2024All You Need is DAGarXiv 2021Gaussian Splatting on the Move: Blur and Rolling Shutter Compensation for Natural Camera MotionarXiv 2024A Provable Defense for Deep Residual NetworksarXiv 2019Scatterbrain: Unifying Sparse and Low-rank Attention Approximationscatterbrain-unifying-sparse-and-low-rank-1The boundary of neural network trainability is fractalarXiv 2024PanopticNeRF-360: Panoramic 3D-to-2D Label Transfer in Urban ScenesarXiv 2023Locate Anything on Earth: Advancing Open-Vocabulary Object Detection for Remote Sensing CommunityarXiv 2024A Survey of Large Language Models AttributionarXiv 2023TKAN: Temporal Kolmogorov-Arnold NetworksarXiv 2024CharacterFactory: Sampling Consistent Characters with GANs for Diffusion ModelsarXiv 2024Pixelated Butterfly: Simple and Efficient Sparse training for Neural Network Modelspixelated-butterfly-simple-and-efficientLERT: A Linguistically-motivated Pre-trained Language ModelarXiv 2022IT3D: Improved Text-to-3D Generation with Explicit View SynthesisarXiv 2023Foundation Models for Music: A SurveyarXiv 2024Monarch: Expressive Structured Matrices for Efficient and Accurate TrainingarXiv 2022ACR: Attention Collaboration-based Regressor for Arbitrary Two-Hand ReconstructionCVPR 2023 1Network Dissection: Quantifying Interpretability of Deep Visual Representationsnetwork-dissection-quantifying-1Generative Diffusion Models on Graphs: Methods and ApplicationsarXiv 2023multiGradICON: A Foundation Model for Multimodal Medical Image RegistrationarXiv 2024Efficient conformer: Progressive downsampling and grouped attention for automatic speech recognitionarXiv 2021Shepherd: A Critic for Language Model GenerationarXiv 2023CVSS Corpus and Massively Multilingual Speech-to-Speech TranslationLREC 2022 6UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond ScalingarXiv 2024Reward Design with Language ModelsarXiv 2023Multimodal Table UnderstandingarXiv 2024Classification Done Right for Vision-Language Pre-TrainingarXiv 2024Can Long-Context Language Models Subsume Retrieval, RAG, SQL, and More?arXiv 2024MAGE: Machine-generated Text Detection in the WildarXiv 2023Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language ProcessingNeurIPS 2020 12ThemeStation: Generating Theme-Aware 3D Assets from Few ExemplarsarXiv 2024Measuring the Intrinsic Dimension of Objective Landscapesmeasuring-the-intrinsic-dimension-of-1Recurrent Drafter for Fast Speculative Decoding in Large Language ModelsarXiv 2024LLM-FP4: 4-Bit Floating-Point Quantized TransformersarXiv 2023Similarity Reasoning and Filtration for Image-Text MatchingarXiv 2021DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs TrainingarXiv 2023Towards Learning a Generalist Model for Embodied NavigationCVPR 2024 1ShieldLM: Empowering LLMs as Aligned, Customizable and Explainable Safety DetectorsarXiv 2024SimplyRetrieve: A Private and Lightweight Retrieval-Centric Generative AI ToolarXiv 2023FLASK: Fine-grained Language Model Evaluation based on Alignment Skill SetsarXiv 2023Scaling Vision Pre-Training to 4K ResolutionCVPR 2025 1One-shot Implicit Animatable Avatars with Model-based PriorsICCV 2023 1Collaborative Novel Object Discovery and Box-Guided Cross-Modal Alignment for Open-Vocabulary 3D Object DetectionarXiv 2024ETH-XGaze: A Large Scale Dataset for Gaze Estimation under Extreme Head Pose and Gaze VariationECCV 2020 8MMedAgent: Learning to Use Medical Tools with Multi-modal AgentarXiv 2024NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving ScenarioarXiv 2023Git-Theta: A Git Extension for Collaborative Development of Machine Learning ModelsarXiv 2023An Extensible Framework for Open Heterogeneous Collaborative PerceptionarXiv 2024

Back to Papers