0

All papers

Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.

X-MOBILITY: End-To-End Generalizable Navigation via World ModelingarXiv 2024SceneFormer: Indoor Scene Generation with TransformersarXiv 2020Rethinking Self-supervised Correspondence Learning: A Video Frame-level Similarity PerspectiveICCV 2021 10G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language ModelarXiv 2023CTRLsum: Towards Generic Controllable Text Summarizationctrlsum-towards-generic-controllable-textInstructScene: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph PriorarXiv 2024Policy-Guided DiffusionarXiv 2024The Archives Unleashed Project: Technology, Process, and Community to Improve Scholarly Access to Web ArchivesarXiv 2020Superposition Prompting: Improving and Accelerating Retrieval-Augmented GenerationarXiv 2024TriAAN-VC: Triple Adaptive Attention Normalization for Any-to-Any Voice ConversionarXiv 2023DART: Articulated Hand Model with Diverse Accessories and Rich TexturesarXiv 2022See More Details: Efficient Image Super-Resolution by Experts MiningarXiv 2024Tuning-Free Image Customization with Image and Text GuidancearXiv 2024SWARM Parallelism: Training Large Models Can Be Surprisingly Communication-Efficientswarm-parallelism-training-large-models-can3D-aware Blending with Generative NeRFsICCV 2023 1A Systematic Review on the Evaluation of Large Language Models in Theory of Mind TasksarXiv 2025Partial Order Pruning: for Best Speed/Accuracy Trade-off in Neural Architecture Searchpartial-order-pruning-for-best-speedaccuracy-1Strivec: Sparse Tri-Vector Radiance FieldsICCV 2023 1Shortformer: Better Language Modeling using Shorter InputsACL 2021 5Can Graph Learning Improve Planning in LLM-based Agents?arXiv 2024Parameter-Efficient Sparsity Crafting from Dense to Mixture-of-Experts for Instruction Tuning on General TasksarXiv 2024Latent Diffusion for Language Generationlatent-diffusion-for-language-generationSOLO: A Single Transformer for Scalable Vision-Language ModelingarXiv 2024Reduce, Reuse, Recycle: Compositional Generation with Energy-Based Diffusion Models and MCMCarXiv 2023EgoMimic: Scaling Imitation Learning via Egocentric VideoarXiv 2024Cross Aggregation Transformer for Image RestorationarXiv 2022Autonomous Evaluation and Refinement of Digital AgentsarXiv 2024AlignScore: Evaluating Factual Consistency with a Unified Alignment FunctionarXiv 2023CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model GenerationarXiv 2023Adversarial Attacks and Defences CompetitionarXiv 2018OneGen: Efficient One-Pass Unified Generation and Retrieval for LLMsarXiv 2024Understanding and Improving Transformer From a Multi-Particle Dynamic System Point of Viewunderstanding-and-improving-transformer-from-1Simple Entity-Centric Questions Challenge Dense RetrieversEMNLP 2021 11Photo-Realistic Monocular Gaze Redirection Using Generative Adversarial Networksphoto-realistic-monocular-gaze-redirection-1Rethinking RL Scaling for Vision Language Models: A Transparent, From-Scratch Framework and Comprehensive Evaluation SchemearXiv 2025Zero-Shot Image Harmonization with Generative Model PriorarXiv 2023QA-LoRA: Quantization-Aware Low-Rank Adaptation of Large Language ModelsarXiv 2023TimeDART: A Diffusion Autoregressive Transformer for Self-Supervised Time Series RepresentationarXiv 2024Sparse Sampling Transformer with Uncertainty-Driven Ranking for Unified Removal of Raindrops and Rain StreaksICCV 2023 1VAR-CLIP: Text-to-Image Generator with Visual Auto-Regressive ModelingarXiv 2024OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel ViewsarXiv 2024Many-Shot In-Context Learning in Multimodal Foundation ModelsarXiv 2024Better Diffusion Models Further Improve Adversarial TrainingarXiv 2023Cross-modal Orthogonal High-rank Augmentation for RGB-Event Transformer-trackersICCV 2023 1OmnimatteRF: Robust Omnimatte with 3D Background ModelingICCV 2023 1Mamba3D: Enhancing Local Features for 3D Point Cloud Analysis via State Space ModelarXiv 2024SlimmeRF: Slimmable Radiance FieldsarXiv 2023A Survey on Benchmarks of Multimodal Large Language ModelsarXiv 2024mMARCO: A Multilingual Version of the MS MARCO Passage Ranking DatasetarXiv 2021Point Cloud Mamba: Point Cloud Learning via State Space ModelarXiv 2024CoIR: A Comprehensive Benchmark for Code Information Retrieval ModelsarXiv 2024Exploring Human-Like Translation Strategy with Large Language ModelsarXiv 2023List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMsarXiv 2024AlignDet: Aligning Pre-training and Fine-tuning in Object DetectionICCV 2023 1AutoPresent: Designing Structured Visuals from ScratchCVPR 2025 1SmartPlay: A Benchmark for LLMs as Intelligent AgentsarXiv 2023L2MAC: Large Language Model Automatic Computer for Extensive Code GenerationarXiv 2023GPT4Motion: Scripting Physical Motions in Text-to-Video Generation via Blender-Oriented GPT PlanningarXiv 2023UV Volumes for Real-time Rendering of Editable Free-view Human PerformanceCVPR 2023 1RWKV-CLIP: A Robust Vision-Language Representation LearnerarXiv 2024Thermostat: A Large Collection of NLP Model Explanations and Analysis ToolsEMNLP (ACL) 2021 11Neural Question Generation from Text: A Preliminary StudyarXiv 2017Towards Large-Scale Training of Pathology Foundation ModelsarXiv 2024Pose Recognition with Cascade TransformersCVPR 2021 1TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful SpacearXiv 2024Synthetic continued pretrainingarXiv 2024InstructERC: Reforming Emotion Recognition in Conversation with Multi-task Retrieval-Augmented Large Language ModelsarXiv 2023Reason for Future, Act for Now: A Principled Framework for Autonomous LLM Agents with Provable Sample EfficiencyarXiv 2023CausalTime: Realistically Generated Time-series for Benchmarking of Causal DiscoveryarXiv 2023Hierarchical Neural Coding for Controllable CAD Model GenerationarXiv 2023A Closer Look at Self-Supervised Lightweight Vision TransformersarXiv 2022BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-modelsACL 2022 5ShineOn: Illuminating Design Choices for Practical Video-based Virtual Clothing Try-onarXiv 2020Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video GenerationarXiv 2024Self-Supervised Relational Reasoning for Representation LearningNeurIPS 2020 12RECOMP: Improving Retrieval-Augmented LMs with Compression and Selective AugmentationarXiv 2023EditEval: An Instruction-Based Benchmark for Text ImprovementsarXiv 2022KazakhTTS: An Open-Source Kazakh Text-to-Speech Synthesis DatasetarXiv 2021DALL-Eval: Probing the Reasoning Skills and Social Biases of Text-to-Image Generation ModelsICCV 2023 1One Token to Seg Them All: Language Instructed Reasoning Segmentation in VideosarXiv 2024DialogLM: Pre-trained Model for Long Dialogue Understanding and SummarizationarXiv 2021DEYOLO: Dual-Feature-Enhancement YOLO for Cross-Modality Object DetectionarXiv 2024OmniSeg3D: Omniversal 3D Segmentation via Hierarchical Contrastive LearningCVPR 2024 1V2X-R: Cooperative LiDAR-4D Radar Fusion for 3D Object Detection with Denoising DiffusionarXiv 2024Representing Schema Structure with Graph Neural Networks for Text-to-SQL ParsingarXiv 2019LLM-CXR: Instruction-Finetuned LLM for CXR Image Understanding and GenerationarXiv 2023Self-Supervised Vision Transformers Learn Visual Concepts in HistopathologyarXiv 2022UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language ModelarXiv 2023Q-Insight: Understanding Image Quality via Visual Reinforcement LearningarXiv 2025UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement LearningarXiv 2025MoEfication: Transformer Feed-forward Layers are Mixtures of ExpertsFindings (ACL) 2022 5HiFormer: Hierarchical Multi-scale Representations Using Transformers for Medical Image SegmentationarXiv 2022GPT Takes the Bar ExamarXiv 2022Self-playing Adversarial Language Game Enhances LLM ReasoningarXiv 2024SGLC: Semantic Graph-Guided Coarse-Fine-Refine Full Loop Closing for LiDAR SLAMarXiv 2024ProSpect: Prompt Spectrum for Attribute-Aware Personalization of Diffusion ModelsarXiv 2023SLiMe: Segment Like MearXiv 2023Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention MapsarXiv 2024PROB: Probabilistic Objectness for Open World Object DetectionCVPR 2023 1Zero-Shot Tokenizer TransferarXiv 2024Cousins Of The Vendi Score: A Family Of Similarity-Based Diversity Metrics For Science And Machine LearningarXiv 2023Open-RAG: Enhanced Retrieval-Augmented Reasoning with Open-Source Large Language ModelsarXiv 2024Playing with Words at the National Library of Sweden -- Making a Swedish BERTarXiv 2020ConvAI3: Generating Clarifying Questions for Open-Domain Dialogue Systems (ClariQ)arXiv 2020Zero-Shot Entity Linking by Reading Entity Descriptionszero-shot-entity-linking-by-reading-entity-1Symbolic Knowledge Distillation: from General Language Models to Commonsense ModelsNAACL 2022 7PID: Physics-Informed Diffusion Model for Infrared Image GenerationarXiv 2024Learning Optimized Risk ScoresarXiv 2016Visual Dexterity: In-Hand Reorientation of Novel and Complex Object ShapesarXiv 2022Representational dissimilarity metric spaces for stochastic neural networksarXiv 2022BERTje: A Dutch BERT ModelarXiv 2019Split-Brain Autoencoders: Unsupervised Learning by Cross-Channel Predictionsplit-brain-autoencoders-unsupervised-1DexArt: Benchmarking Generalizable Dexterous Manipulation with Articulated ObjectsCVPR 2023 1Generating Synthetic Documents for Cross-Encoder Re-Rankers: A Comparative Study of ChatGPT and Human ExpertsarXiv 2023NodePiece: Compositional and Parameter-Efficient Representations of Large Knowledge Graphsnodepiece-compositional-and-parameter-1Benchmarking Agentic Workflow GenerationarXiv 2024Adaptive Rotated Convolution for Rotated Object DetectionICCV 2023 1Scalable Autoregressive Image Generation with MambaarXiv 2024SalUn: Empowering Machine Unlearning via Gradient-based Weight Saliency in Both Image Classification and GenerationarXiv 2023MultiQA: An Empirical Investigation of Generalization and Transfer in Reading Comprehensionmultiqa-an-empirical-investigation-of-1On Feature Normalization and Data AugmentationCVPR 2021 1QMSum: A New Benchmark for Query-based Multi-domain Meeting SummarizationNAACL 2021 4Normalized Loss Functions for Deep Learning with Noisy LabelsICML 2020 1Post-training Quantization on Diffusion ModelsCVPR 2023 1Ground-A-Video: Zero-shot Grounded Video Editing using Text-to-image Diffusion ModelsarXiv 2023Constrained Graphic Layout Generation via Latent OptimizationarXiv 2021Self-supervised Learning of Adversarial Example: Towards Good Generalizations for Deepfake DetectionCVPR 2022 1An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLMarXiv 2024FastPathology: An open-source platform for deep learning-based research and decision support in digital pathologyarXiv 2020Sample4Geo: Hard Negative Sampling For Cross-View Geo-LocalisationICCV 2023 1Message Passing Neural PDE Solversmessage-passing-neural-pde-solversVIOLET : End-to-End Video-Language Transformers with Masked Visual-token ModelingarXiv 2021Natural scene reconstruction from fMRI signals using generative latent diffusionarXiv 2023StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video UnderstandingarXiv 2024Equivariant Multi-Modality Image FusionCVPR 2024 1SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and SynthesisarXiv 2024Pile of Law: Learning Responsible Data Filtering from the Law and a 256GB Open-Source Legal DatasetarXiv 2022Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%arXiv 2024Image-to-Image Translation via Group-wise Deep Whitening-and-Coloring Transformationimage-to-image-translation-via-group-wise-1Generative Compositional Augmentations for Scene Graph PredictionICCV 2021 10A-Bench: Are LMMs Masters at Evaluating AI-generated Images?arXiv 2024Implicit Chain of Thought Reasoning via Knowledge DistillationarXiv 2023Telling Left from Right: Identifying Geometry-Aware Semantic CorrespondenceCVPR 2024 1MC-LLaVA: Multi-Concept Personalized Vision-Language ModelarXiv 2024A Comparative Study of Image Restoration Networks for General Backbone Network DesignarXiv 2023DM-VTON: Distilled Mobile Real-time Virtual Try-OnarXiv 2023Filtering, Distillation, and Hard Negatives for Vision-Language Pre-TrainingCVPR 2023 1Deep Reinforcement Learning: An OverviewarXiv 2017Retrieval-Augmented Layout Transformer for Content-Aware Layout GenerationCVPR 2024 1Sin3DM: Learning a Diffusion Model from a Single 3D Textured ShapearXiv 2023PoRF: Pose Residual Field for Accurate Neural Surface ReconstructionarXiv 2023Decomposing and Editing Predictions by Modeling Model ComputationarXiv 2024Symmetry and Uncertainty-Aware Object SLAM for 6DoF Object Pose EstimationarXiv 2022RepQ-ViT: Scale Reparameterization for Post-Training Quantization of Vision TransformersICCV 2023 1Hidden Gems: 4D Radar Scene Flow Learning Using Cross-Modal SupervisionCVPR 2023 1Twin-Merging: Dynamic Integration of Modular Expertise in Model MergingarXiv 2024Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language ModelsarXiv 2024Large Language Models are Efficient Learners of Noise-Robust Speech RecognitionarXiv 2024Language modeling via stochastic processeslanguage-modeling-via-stochastic-processesRankGen: Improving Text Generation with Large Ranking ModelsarXiv 2022KQA Pro: A Dataset with Explicit Compositional Programs for Complex Question Answering over Knowledge BaseACL 2022 5Described Object Detection: Liberating Object Detection with Flexible Expressionsdescribed-object-detection-liberating-objectLogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical ReasoningarXiv 2020CAT-SAM: Conditional Tuning for Few-Shot Adaptation of Segment Anything ModelarXiv 2024RegFormer: An Efficient Projection-Aware Transformer Network for Large-Scale Point Cloud RegistrationICCV 2023 1Self-Distillation Bridges Distribution Gap in Language Model Fine-TuningarXiv 2024The Impact of Positional Encoding on Length Generalization in Transformersthe-impact-of-positional-encoding-on-lengthBeyond Text: Frozen Large Language Models in Visual Signal ComprehensionCVPR 2024 1Copiloting the Copilots: Fusing Large Language Models with Completion Engines for Automated Program RepairarXiv 2023GNFactor: Multi-Task Real Robot Learning with Generalizable Neural Feature FieldsarXiv 2023EditWorld: Simulating World Dynamics for Instruction-Following Image EditingarXiv 2024ZeroShape: Regression-based Zero-shot Shape ReconstructionCVPR 2024 1UnitSpeech: Speaker-adaptive Speech Synthesis with Untranscribed DataarXiv 2023Evaluating Quantized Large Language ModelsarXiv 2024MiniGPT-Med: Large Language Model as a General Interface for Radiology DiagnosisarXiv 2024Parallel Speculative Decoding with Adaptive Draft LengtharXiv 2024BitMoD: Bit-serial Mixture-of-Datatype LLM AccelerationarXiv 2024If your data distribution shifts, use self-learningif-your-data-distribution-shifts-use-selfTowards Dense and Accurate Radar Perception Via Efficient Cross-Modal Diffusion ModelarXiv 2024Decoupling the Depth and Scope of Graph Neural Networksdecoupling-the-depth-and-scope-of-graphLarge Raw Emotional Dataset with Aggregation MechanismarXiv 2022EvoAgent: Towards Automatic Multi-Agent Generation via Evolutionary AlgorithmsarXiv 2024Deductive Verification of Chain-of-Thought Reasoningdeductive-verification-of-chain-of-thoughtGlancing Transformer for Non-Autoregressive Neural Machine TranslationACL 2021 5CPF: Learning a Contact Potential Field to Model the Hand-Object InteractionICCV 2021 10Latent diffusion models for generative precipitation nowcasting with accurate uncertainty quantificationarXiv 2023DrugAssist: A Large Language Model for Molecule OptimizationarXiv 2023Towards Squeezing-Averse Virtual Try-On via Sequential DeformationarXiv 2023Putting NeRF on a Diet: Semantically Consistent Few-Shot View SynthesisICCV 2021 10Estimating the Contamination Factor's Distribution in Unsupervised Anomaly DetectionarXiv 2022Building Efficient Universal Classifiers with Natural Language InferencearXiv 2023GAN-Control: Explicitly Controllable GANsICCV 2021 10FOLIO: Natural Language Reasoning with First-Order LogicarXiv 2022How Do Large Language Models Capture the Ever-changing World Knowledge? A Review of Recent AdvancesarXiv 2023RLTF: Reinforcement Learning from Unit Test FeedbackarXiv 2023Instant Multi-View Head Capture through Learnable Registrationinstant-multi-view-head-capture-throughCAT-DM: Controllable Accelerated Virtual Try-on with Diffusion ModelCVPR 2024 1$\texttt{DINO-Foresight}$: Looking into the Future with DINOarXiv 2024The Ubuntu Dialogue Corpus: A Large Dataset for Research in Unstructured Multi-Turn Dialogue Systemsthe-ubuntu-dialogue-corpus-a-large-datasetImproved Universal Sentence Embeddings with Prompt-based Contrastive Learning and Energy-based LearningarXiv 2022

Back to Papers