0

All papers

Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.

OccFormer: Dual-path Transformer for Vision-based 3D Semantic Occupancy PredictionICCV 2023 1VisionLLaMA: A Unified LLaMA Backbone for Vision TasksarXiv 2024SampleNet: Differentiable Point Cloud Samplingsamplenet-differentiable-point-cloud-sampling-1Zero-Shot Surgical Tool Segmentation in Monocular Video Using Segment Anything Model 2arXiv 2024The Truth is in There: Improving Reasoning in Language Models with Layer-Selective Rank ReductionarXiv 2023ChatGPT Empowered Long-Step Robot Control in Various Environments: A Case ApplicationarXiv 2023OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-AllocationCVPR 2024 1A Discourse-Aware Attention Model for Abstractive Summarization of Long Documentsa-discourse-aware-attention-model-for-1Google Crowdsourced Speech Corpora and Related Open-Source Resources for Low-Resource Languages and Dialects: An OverviewarXiv 2020BARS: Towards Open Benchmarking for Recommender SystemsarXiv 2022Zhongjing: Enhancing the Chinese Medical Capabilities of Large Language Model through Expert Feedback and Real-world Multi-turn DialoguearXiv 2023GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented UnderstandingarXiv 2024SimAlign: High Quality Word Alignments without Parallel Training Data using Static and Contextualized EmbeddingsFindings of the Association for Computational Linguistics 2020ID-Animator: Zero-Shot Identity-Preserving Human Video GenerationarXiv 2024Dojo: A Differentiable Physics Engine for RoboticsarXiv 2022Sparse Networks from Scratch: Faster Training without Losing Performancesparse-networks-from-scratch-faster-training-1Animate-X: Universal Character Image Animation with Enhanced Motion RepresentationarXiv 2024Offsite-Tuning: Transfer Learning without Full ModelarXiv 2023When Counting Meets HMER: Counting-Aware Network for Handwritten Mathematical Expression RecognitionarXiv 2022Inst-Inpaint: Instructing to Remove Objects with Diffusion ModelsarXiv 2023Learning Structured Sparsity in Deep Neural Networkslearning-structured-sparsity-in-deep-neural-1OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective FusionarXiv 20243D-SIS: 3D Semantic Instance Segmentation of RGB-D Scans3d-sis-3d-semantic-instance-segmentation-of-1LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and EditingarXiv 2023Deep Floor Plan Recognition Using a Multi-Task Network with Room-Boundary-Guided Attentiondeep-floor-plan-recognition-using-a-multi-1Point-SAM: Promptable 3D Segmentation Model for Point CloudsarXiv 2024PFGM++: Unlocking the Potential of Physics-Inspired Generative ModelsarXiv 2023BinaryConnect: Training Deep Neural Networks with binary weights during propagationsbinaryconnect-training-deep-neural-networks-1MemoryBank: Enhancing Large Language Models with Long-Term MemoryarXiv 2023StyleGAN of All Trades: Image Manipulation with Only Pretrained StyleGANarXiv 2021Fast Matrix Multiplications for Lookup Table-Quantized LLMsarXiv 2024Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalICLR 2021 1DeCLUTR: Deep Contrastive Learning for Unsupervised Textual RepresentationsACL 2021 5DiffuseVAE: Efficient, Controllable and High-Fidelity Generation from Low-Dimensional LatentsarXiv 2022SceneRF: Self-Supervised Monocular 3D Scene Reconstruction with Radiance FieldsICCV 2023 1Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalICCV 2021 10AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionarXiv 2022DIRE for Diffusion-Generated Image DetectionICCV 2023 1ViewDiff: 3D-Consistent Image Generation with Text-to-Image ModelsCVPR 2024 1iSeeBetter: Spatio-temporal video super-resolution using recurrent generative back-projection networksarXiv 2020Localization, Detection and Tracking of Multiple Moving Sound Sources with a Convolutional Recurrent Neural NetworkarXiv 2019Knowledge Enhanced Contextual Word Representationsknowledge-enhanced-contextual-word-1ALIKE: Accurate and Lightweight Keypoint Detection and Descriptor ExtractionarXiv 2021Diff2Lip: Audio Conditioned Diffusion Models for Lip-SynchronizationarXiv 2023Arm-Constrained Curriculum Learning for Loco-Manipulation of the Wheel-Legged RobotarXiv 2024Robust Watermarking Using Generative Priors Against Image Editing: From Benchmarking to AdvancesarXiv 2024DisPose: Disentangling Pose Guidance for Controllable Human Image AnimationarXiv 2024Rethinking Few-Shot Image Classification: a Good Embedding Is All You Need?ECCV 2020 8GaussianImage: 1000 FPS Image Representation and Compression by 2D Gaussian SplattingarXiv 2024Unifying Vision-and-Language Tasks via Text GenerationarXiv 2021Rethinking Image Inpainting via a Mutual Encoder-Decoder with Feature EqualizationsECCV 2020 8UniVTG: Towards Unified Video-Language Temporal GroundingICCV 2023 1Direct Inversion: Boosting Diffusion-based Editing with 3 Lines of CodearXiv 2023Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarizationdont-give-me-the-details-just-the-summary-1Towards a Reinforcement Learning Environment Toolbox for Intelligent Electric Motor ControlarXiv 2019Single-Image Piece-wise Planar 3D Reconstruction via Associative Embeddingsingle-image-piece-wise-planar-3d-1DaViT: Dual Attention Vision TransformersarXiv 2022RETURNN as a Generic Flexible Neural Toolkit with Application to Translation and Speech Recognitionreturnn-as-a-generic-flexible-neural-toolkit-1A Survey of Deep Learning for Mathematical ReasoningarXiv 2022RETURNN as a Generic Flexible Neural Toolkit with Application to Translation and Speech Recognitionreturnn-as-a-generic-flexible-neural-toolkit-1SMIRK: 3D Facial Expressions through Analysis-by-Neural-SynthesisarXiv 2024Lookahead: An Inference Acceleration Framework for Large Language Model with Lossless Generation AccuracyarXiv 2023HarDNet: A Low Memory Traffic Networkhardnet-a-low-memory-traffic-network-1Taming Visually Guided Sound GenerationarXiv 2021A Self-Supervised Descriptor for Image Copy DetectionCVPR 2022 1Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language ModelarXiv 2023Lost in the Middle: How Language Models Use Long ContextsarXiv 2023ReMoDiffuse: Retrieval-Augmented Motion Diffusion ModelICCV 2023 1Dodrio: Exploring Transformer Models with Interactive VisualizationACL 2021 5LLark: A Multimodal Instruction-Following Language Model for MusicarXiv 2023FRNet: Frustum-Range Networks for Scalable LiDAR SegmentationarXiv 2023Whistle: Data-Efficient Multilingual and Crosslingual Speech Recognition via Weakly Phonetic SupervisionarXiv 2024AutoRecon: Automated 3D Object Discovery and ReconstructionCVPR 2023 1Deep Hough Transform for Semantic Line DetectionECCV 2020 8Talk-to-Edit: Fine-Grained Facial Editing via DialogICCV 2021 10ByT5 model for massively multilingual grapheme-to-phoneme conversionarXiv 2022Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive AttacksarXiv 2024Mel-Band RoFormer for Music Source SeparationarXiv 2023Mask-Free Video Instance SegmentationCVPR 2023 1Accelerating Material Design with the Generative Toolkit for Scientific DiscoveryarXiv 2022VideoPrism: A Foundational Visual Encoder for Video UnderstandingarXiv 2024Universal Source Separation with Weakly Labelled DataarXiv 2023Measuring Information Propagation in Literary Social NetworksEMNLP 2020 11TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style ControlarXiv 2024Generative Time Series Forecasting with Diffusion, Denoise, and DisentanglementarXiv 2023Side Adapter Network for Open-Vocabulary Semantic SegmentationCVPR 2023 1AudioMNIST: Exploring Explainable Artificial Intelligence for Audio Analysis on a Simple BenchmarkarXiv 2018Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question ComplexityarXiv 2024Q-Diffusion: Quantizing Diffusion ModelsICCV 2023 1Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical ReasoningarXiv 2023Selective Reflection-Tuning: Student-Selected Data Recycling for LLM Instruction-TuningarXiv 2024MindMap: Knowledge Graph Prompting Sparks Graph of Thoughts in Large Language ModelsarXiv 2023Learning Deep Time-index Models for Time Series ForecastingarXiv 2022Igniting Language Intelligence: The Hitchhiker's Guide From Chain-of-Thought Reasoning to Language AgentsarXiv 2023Reflection-Tuning: Data Recycling Improves LLM Instruction-TuningarXiv 2023V-IRL: Grounding Virtual Intelligence in Real LifearXiv 2024GaussianOcc: Fully Self-supervised and Efficient 3D Occupancy Estimation with Gaussian SplattingarXiv 2024vMAP: Vectorised Object Mapping for Neural Field SLAMCVPR 2023 1Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation ModelarXiv 2025Machine Reading Comprehension: The Role of Contextualized Language Models and BeyondarXiv 2020Towards Multimodal Sarcasm Detection (An _Obviously_ Perfect Paper)arXiv 2019AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-TuningarXiv 2023IRIS: LLM-Assisted Static Analysis for Detecting Security VulnerabilitiesarXiv 2024Generating Holistic 3D Human Motion from SpeechCVPR 2023 1DISK: Learning local features with policy gradientNeurIPS 2020 12Unsupervised Learning of Video Representations using LSTMsarXiv 2015Deep Fusion Network for Image CompletionarXiv 2019Efficient Training of Audio Transformers with PatchoutarXiv 2021DesignEdit: Multi-Layered Latent Decomposition and Fusion for Unified & Accurate Image EditingarXiv 2024Word Alignment by Fine-tuning Embeddings on Parallel CorporaEACL 2021 2EVEv2: Improved Baselines for Encoder-Free Vision-Language ModelsICCV 2025DeepFool: a simple and accurate method to fool deep neural networksdeepfool-a-simple-and-accurate-method-to-fool-1KV-Edit: Training-Free Image Editing for Precise Background PreservationarXiv 2025EscherNet: A Generative Model for Scalable View SynthesisCVPR 2024 1Geometry-Aware Learning of Maps for Camera Localizationgeometry-aware-learning-of-maps-for-camera-1QuadTree Attention for Vision Transformersquadtree-attention-for-vision-transformersHow Far Are We From AGI: Are LLMs All We Need?arXiv 2024AnimeSR: Learning Real-World Super-Resolution Models for Animation VideosarXiv 2022Reconstructing the Mind's Eye: fMRI-to-Image with Contrastive Learning and Diffusion PriorsNeurIPS 2023 11RLPrompt: Optimizing Discrete Text Prompts with Reinforcement LearningarXiv 2022Recovering 3D Human Mesh from Monocular Images: A SurveyarXiv 2022LiLT: A Simple yet Effective Language-Independent Layout Transformer for Structured Document UnderstandingACL 2022 5HumanTOMATO: Text-aligned Whole-body Motion GenerationarXiv 2023Deep Geometrized Cartoon Line Inbetweeningdeep-geometrized-cartoon-line-inbetweeningSelf-Supervised Learning for Time Series Analysis: Taxonomy, Progress, and ProspectsarXiv 2023MMICL: Empowering Vision-language Model with Multi-Modal In-Context LearningarXiv 2023TextDescriptives: A Python package for calculating a large variety of metrics from textarXiv 2023BEVPlace++: Fast, Robust, and Lightweight LiDAR Global Localization for Unmanned Ground VehiclesarXiv 2024D3Net: Densely connected multidilated DenseNet for music source separationarXiv 2020GiT: Towards Generalist Vision Transformer through Universal Language InterfacearXiv 2024Is Mamba Effective for Time Series Forecasting?arXiv 2024Spatially-Adaptive Feature Modulation for Efficient Image Super-ResolutionICCV 2023 1MCVD: Masked Conditional Video Diffusion for Prediction, Generation, and InterpolationarXiv 2022Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D ModelsarXiv 2024VQMIVC: Vector Quantization and Mutual Information-Based Unsupervised Speech Representation Disentanglement for One-shot Voice ConversionarXiv 2021HyperSIGMA: Hyperspectral Intelligence Comprehension Foundation ModelarXiv 20243D Face Reconstruction with the Geometric Guidance of Facial Part SegmentationCVPR 2024 1Propagate Yourself: Exploring Pixel-Level Consistency for Unsupervised Visual Representation LearningCVPR 2021 1Towards Language Models That Can See: Computer Vision Through the LENS of Natural LanguagearXiv 2023Zero-Shot Video Editing Using Off-The-Shelf Image Diffusion ModelsarXiv 2023Pandora3D: A Comprehensive Framework for High-Quality 3D Shape and Texture GenerationarXiv 2025SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual ComprehensionarXiv 2024Maintaining Plasticity in Deep Continual LearningarXiv 2023Code to Think, Think to Code: A Survey on Code-Enhanced Reasoning and Reasoning-Driven Code Intelligence in LLMsarXiv 2025Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMsCVPR 2024 1Tabular Transformers for Modeling Multivariate Time SeriesarXiv 2020Wav2CLIP: Learning Robust Audio Representations From CLIParXiv 2021VQGAN-CLIP: Open Domain Image Generation and Editing with Natural Language GuidancearXiv 2022GLOBEM Dataset: Multi-Year Datasets for Longitudinal Human Behavior Modeling GeneralizationarXiv 2022How Attentive are Graph Attention Networks?how-attentive-are-graph-attention-networks-1Training Socially Aligned Language Models on Simulated Social InteractionsarXiv 2023VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video GenerationarXiv 2024Revisiting Self-Supervised Visual Representation Learningrevisiting-self-supervised-visual-1Inversion-Free Image Editing with Natural LanguagearXiv 2023SoccerNet Game State Reconstruction: End-to-End Athlete Tracking and Identification on a MinimaparXiv 2024Odyssey: Empowering Minecraft Agents with Open-World SkillsarXiv 2024Smooth Diffusion: Crafting Smooth Latent Spaces in Diffusion ModelsCVPR 2024 1CoSeR: Bridging Image and Language for Cognitive Super-ResolutionCVPR 2024 1SAM2Point: Segment Any 3D as Videos in Zero-shot and Promptable MannersarXiv 2024Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative InstructionsarXiv 2023Simple and Fast Distillation of Diffusion ModelsarXiv 2024AnimateZero: Video Diffusion Models are Zero-Shot Image AnimatorsarXiv 2023Language as Queries for Referring Video Object SegmentationCVPR 2022 1Learning Spatio-Temporal Representation with Pseudo-3D Residual Networkslearning-spatio-temporal-representation-with-1Any6D: Model-free 6D Pose Estimation of Novel ObjectsCVPR 2025 1Crystal Diffusion Variational Autoencoder for Periodic Material Generationcrystal-diffusion-variational-autoencoder-forDynamic Typography: Bringing Text to Life via Video Diffusion PriorICCV 2025DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretrainingdoremi-optimizing-data-mixtures-speeds-upCellViT: Vision Transformers for Precise Cell Segmentation and ClassificationarXiv 2023HD-Painter: High-Resolution and Prompt-Faithful Text-Guided Image Inpainting with Diffusion ModelsarXiv 2023Faster Diffusion: Rethinking the Role of the Encoder for Diffusion Model InferencearXiv 2023Database Reasoning Over TextACL 2021 5Multimodality Helps Unimodality: Cross-Modal Few-Shot Learning with Multimodal ModelsCVPR 2023 1A Simple Unified Framework for Detecting Out-of-Distribution Samples and Adversarial Attacksa-simple-unified-framework-for-detecting-out-1UniPC: A Unified Predictor-Corrector Framework for Fast Sampling of Diffusion Modelsunipc-a-unified-predictor-corrector-frameworkSMILES Transformer: Pre-trained Molecular Fingerprint for Low Data Drug DiscoveryarXiv 2019Understanding Deep Networks via Extremal Perturbations and Smooth Masksunderstanding-deep-networks-via-extremal-1Unleashing the Emergent Cognitive Synergy in Large Language Models: A Task-Solving Agent through Multi-Persona Self-CollaborationarXiv 2023Time-Travel RephotographyarXiv 2020DRCT: Saving Image Super-resolution away from Information BottleneckarXiv 2024GPT-Fathom: Benchmarking Large Language Models to Decipher the Evolutionary Path towards GPT-4 and BeyondarXiv 2023Arabic Text Diacritization Using Deep Neural NetworksarXiv 2019A Tale of Two Features: Stable Diffusion Complements DINO for Zero-Shot Semantic CorrespondenceNeurIPS 2023 11Learning to Learn with Generative Models of Neural Network CheckpointsarXiv 2022ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech DetectionACL 2022 5MS MARCO Web Search: a Large-scale Information-rich Web Dataset with Millions of Real Click LabelsarXiv 2024pfl-research: simulation framework for accelerating research in Private Federated LearningarXiv 2024PromptKD: Unsupervised Prompt Distillation for Vision-Language ModelsCVPR 2024 1Evaluating Pixel Language Models on Non-Standardized LanguagesarXiv 2024Learning to Ask: Neural Question Generation for Reading Comprehensionlearning-to-ask-neural-question-generation-1MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video UnderstandingCVPR 2024 1HumanVid: Demystifying Training Data for Camera-controllable Human Image AnimationarXiv 2024Efficient Certification of Spatial RobustnessarXiv 2020Unleashing Vanilla Vision Transformer with Masked Image Modeling for Object DetectionICCV 2023 1Unmasked Teacher: Towards Training-Efficient Video Foundation ModelsICCV 2023 1Enhancing End-to-End Autonomous Driving with Latent World ModelarXiv 2024TEACHTEXT: CrossModal Generalized Distillation for Text-Video RetrievalICCV 2021 10PromptBERT: Improving BERT Sentence Embeddings with PromptsarXiv 2022PhysTwin: Physics-Informed Reconstruction and Simulation of Deformable Objects from VideosICCV 2025XPhoneBERT: A Pre-trained Multilingual Model for Phoneme Representations for Text-to-SpeecharXiv 2023

Back to Papers