All papers
Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.
Locate and Verify: A Two-Stream Network for Improved Deepfake DetectionarXiv 2023Looking Through the Glass: Neural Surface Reconstruction Against High Specular ReflectionsCVPR 2023 1TiZero: Mastering Multi-Agent Football with Curriculum Learning and Self-PlayarXiv 2023ExpMRC: Explainability Evaluation for Machine Reading ComprehensionarXiv 2021Kangaroo: Lossless Self-Speculative Decoding via Double Early ExitingarXiv 2024Online Speculative DecodingarXiv 2023MPI-Flow: Learning Realistic Optical Flow with Multiplane ImagesICCV 2023 1GPT or BERT: why not both?arXiv 2024Self-Exploring Language Models: Active Preference Elicitation for Online AlignmentarXiv 2024DASpeech: Directed Acyclic Transformer for Fast and High-quality Speech-to-Speech Translationdaspeech-directed-acyclic-transformer-for3D-SPS: Single-Stage 3D Visual Grounding via Referred Point Progressive SelectionCVPR 2022 1Object-Aware Distillation Pyramid for Open-Vocabulary Object DetectionCVPR 2023 1A Survey on the Honesty of Large Language ModelsarXiv 2024Bridging the Gap Between Vision Transformers and Convolutional Neural Networks on Small DatasetsarXiv 2022Agent-based Learning of Materials Datasets from Scientific LiteraturearXiv 2023Efficient Visual Pretraining with Contrastive DetectionICCV 2021 10Expressing Visual Relationships via Languageexpressing-visual-relationships-via-language-1HEAR: Holistic Evaluation of Audio RepresentationsarXiv 2022Liger: Linearizing Large Language Models to Gated Recurrent StructuresarXiv 2025Visual Semantic Role Labeling for Video UnderstandingCVPR 2021 1An Internal Learning Approach to Video Inpaintingan-internal-learning-approach-to-video-1Compressed Context Memory For Online Language Model InteractionarXiv 2023Improving 3D Imaging with Pre-Trained Perpendicular 2D Diffusion ModelsICCV 2023 1ReCo: Retrieve and Co-segment for Zero-shot Transferreco-retrieve-and-co-segment-for-zero-shotCall for Customized Conversation: Customized Conversation Grounding Persona and KnowledgearXiv 2021Transformer Embeddings of Irregularly Spaced Events and Their Participantstransformer-embeddings-of-irregularly-spacedSelf-supervised Spatio-temporal Representation Learning for Videos by Predicting Motion and Appearance Statisticsself-supervised-spatio-temporal-1Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewardsrewarded-soups-towards-pareto-optimalCheetah: Bridging the Gap Between Machine Learning and Particle Accelerator Physics with High-Speed, Differentiable SimulationsarXiv 2024NoiseCollage: A Layout-Aware Text-to-Image Diffusion Model Based on Noise Cropping and MergingCVPR 2024 1Compute Better Spent: Replacing Dense Layers with Structured MatricesarXiv 2024Plug-and-Play Knowledge Injection for Pre-trained Language ModelsarXiv 2023AutoDefense: Multi-Agent LLM Defense against Jailbreak AttacksarXiv 2024Is ChatGPT a Good Recommender? A Preliminary StudyarXiv 2023AttentionHTR: Handwritten Text Recognition Based on Attention Encoder-Decoder NetworksarXiv 2022DropPos: Pre-Training Vision Transformers by Reconstructing Dropped Positionsdroppos-pre-training-vision-transformers-byEfficient Large Multi-modal Models via Visual Context CompressionarXiv 2024BARThez: a Skilled Pretrained French Sequence-to-Sequence ModelEMNLP 2021 11GPT3Mix: Leveraging Large-scale Language Models for Text AugmentationFindings (EMNLP) 2021 11Revisiting Multimodal Representation in Contrastive Learning: From Patch and Token Embeddings to Finite Discrete TokensCVPR 2023 1Flames: Benchmarking Value Alignment of LLMs in ChinesearXiv 2023FP-Age: Leveraging Face Parsing Attention for Facial Age Estimation in the WildarXiv 2021Questions Are All You Need to Train a Dense Passage RetrieverarXiv 2022Joint Learning of Deep Retrieval Model and Product Quantization based Embedding IndexarXiv 2021Skill Expansion and Composition in Parameter SpacearXiv 2025RegionDrag: Fast Region-Based Image Editing with Diffusion ModelsarXiv 2024How Does Information Bottleneck Help Deep Learning?arXiv 2023HybridGS: Decoupling Transients and Statics with 2D and 3D Gaussian SplattingCVPR 2025 1INSTRUCTSCORE: Explainable Text Generation Evaluation with Finegrained FeedbackarXiv 2023EfficientRAG: Efficient Retriever for Multi-Hop Question AnsweringarXiv 2024Ag2Manip: Learning Novel Manipulation Skills with Agent-Agnostic Visual and Action RepresentationsarXiv 2024Applications of Spiking Neural Networks in Visual Place RecognitionarXiv 2023LongReward: Improving Long-context Large Language Models with AI FeedbackarXiv 2024LoRAPrune: Structured Pruning Meets Low-Rank Parameter-Efficient Fine-TuningarXiv 2023PAPILLON: Privacy Preservation from Internet-based and Local Language Model EnsemblesarXiv 2024Mask-Align: Self-Supervised Neural Word AlignmentACL 2021 5From Molecules to Materials: Pre-training Large Generalizable Models for Atomic Property PredictionarXiv 2023Pretrained Language Models as Visual Planners for Human AssistanceICCV 2023 1Cross-Ray Neural Radiance Fields for Novel-view Synthesis from Unconstrained Image CollectionsICCV 2023 1Large Language Models for Next Point-of-Interest RecommendationarXiv 2024Embodied Multi-Modal Agent trained by an LLM from a Parallel TextWorldCVPR 2024 1Mixture of LoRA ExpertsarXiv 2024Ragnarök: A Reusable RAG Framework and Baselines for TREC 2024 Retrieval-Augmented Generation TrackarXiv 2024UniSumm and SummZoo: Unified Model and Diverse Benchmark for Few-Shot SummarizationarXiv 2022Efficient Image Pre-Training with Siamese Cropped Masked AutoencodersarXiv 2024MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language UnderstandingarXiv 2024Long-Term Rhythmic Video SoundtrackerarXiv 2023Stable Neural Stochastic Differential Equations in Analyzing Irregular Time Series DataarXiv 2024Unnoticeable Backdoor Attacks on Graph Neural NetworksarXiv 2023Improved baselines for vision-language pre-trainingarXiv 2023Instagram Fake and Automated Account DetectionarXiv 2019Generative Modeling with Optimal Transport Mapsgenerative-modeling-with-optimal-transport-1Otter-Knowledge: benchmarks of multimodal knowledge graph representation learning from different sources for drug discoveryarXiv 2023Mining Discourse Markers for Unsupervised Sentence Representation Learningmining-discourse-markers-for-unsupervised-1ARoFace: Alignment Robustness to Improve Low-Quality Face RecognitionarXiv 2024A Large-Scale Multi-Document Summarization Dataset from the Wikipedia Current Events Portala-large-scale-multi-document-summarization-1Using a Waffle Iron for Automotive Point Cloud Semantic SegmentationICCV 2023 1RadioTalk: a large-scale corpus of talk radio transcriptsarXiv 2019Building a Role Specified Open-Domain Dialogue System Leveraging Large-Scale Language ModelsNAACL 2022 7Airavata: Introducing Hindi Instruction-tuned LLMarXiv 2024Deep Reinforcement Learning Guided Improvement Heuristic for Job Shop SchedulingarXiv 2022GraCo: Granularity-Controllable Interactive SegmentationCVPR 2024 1Waffling around for Performance: Visual Classification with Random Words and Broad ConceptsICCV 2023 1Gradient Starvation: A Learning Proclivity in Neural NetworksNeurIPS 2021 12M3DBench: Let's Instruct Large Models with Multi-modal 3D PromptsarXiv 2023A Practitioner's Guide to Continual Multimodal PretrainingarXiv 2024TorchLean: Formalizing Neural Networks in LeanarXiv 2026Multiple View Geometry Transformers for 3D Human Pose EstimationCVPR 2024 1ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement LearningarXiv 2025Semantic Role Labeling as Dependency Parsing: Exploring Latent Tree Structures Inside ArgumentsCOLING 2022 10Composing Parameter-Efficient Modules with Arithmetic OperationsarXiv 2023Towards Galaxy Foundation Models with Hybrid Contrastive LearningarXiv 2022Cross-Modal Translation and Alignment for Survival AnalysisICCV 2023 1From Commit Message Generation to History-Aware Commit Message CompletionarXiv 2023SUTD-TrafficQA: A Question Answering Benchmark and an Efficient Network for Video Reasoning over Traffic Eventssutd-trafficqa-a-question-answering-benchmarkDynamic Inertial Poser (DynaIP): Part-Based Motion Dynamics Learning for Enhanced Human Pose Estimation with Sparse Inertial SensorsCVPR 2024 1Trustworthy Long-Tailed ClassificationCVPR 2022 1Inducing Positive Perspectives with Text ReframingACL 2022 5Light Schrödinger BridgearXiv 2023Bootstrap Latent Representations for Multi-modal RecommendationarXiv 2022INQUIRE: A Natural World Text-to-Image Retrieval BenchmarkarXiv 2024SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference AccelerationarXiv 2024Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning ModelsarXiv 2025LAVENDER: Unifying Video-Language Understanding as Masked Language ModelingCVPR 2023 1PA-LLaVA: A Large Language-Vision Assistant for Human Pathology Image UnderstandingarXiv 2024Emotion Recognition From Speech With Recurrent Neural NetworksarXiv 2017Recasting Self-Attention with Holographic Reduced RepresentationsarXiv 2023Improving Knowledge Graph Embedding Using Simple Constraintsimproving-knowledge-graph-embedding-using-1GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient AnalysisarXiv 2024MagicClay: Sculpting Meshes With Generative Neural FieldsarXiv 2024Unified Vision and Language Prompt LearningarXiv 2022Unmasking Anomalies in Road-Scene SegmentationICCV 2023 1iFormer: Integrating ConvNet and Transformer for Mobile ApplicationarXiv 2025Deep Clustering via Joint Convolutional Autoencoder Embedding and Relative Entropy Minimizationdeep-clustering-via-joint-convolutional-1DOCBENCH: A Benchmark for Evaluating LLM-based Document Reading SystemsarXiv 2024MoCha: Towards Movie-Grade Talking Character SynthesisarXiv 2025Jetfire: Efficient and Accurate Transformer Pretraining with INT8 Data Flow and Per-Block QuantizationarXiv 2024EP2P-Loc: End-to-End 3D Point to 2D Pixel Localization for Large-Scale Visual LocalizationICCV 2023 1Learning to Route in Similarity GraphsarXiv 2019Sense Vocabulary Compression through the Semantic Knowledge of WordNet for Neural Word Sense DisambiguationGWC 2019 7LiveScene: Language Embedding Interactive Radiance Fields for Physical Scene Rendering and ControlarXiv 2024CRAFT: Customizing LLMs by Creating and Retrieving from Specialized ToolsetsarXiv 2023On the Role of Attention Heads in Large Language Model SafetyarXiv 2024CTR-Driven Advertising Image Generation with Multimodal Large Language ModelsarXiv 2025FLIP: A Provable Defense Framework for Backdoor Mitigation in Federated LearningarXiv 2022UMBRAE: Unified Multimodal Brain DecodingarXiv 2024Are Local Features All You Need for Cross-Domain Visual Place Recognition?arXiv 2023PD-Quant: Post-Training Quantization based on Prediction Difference MetricCVPR 2023 1CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent BenchmarkarXiv 2024Revisiting DocRED -- Addressing the False Negative Problem in Relation ExtractionarXiv 2022Augmentation-Adapted Retriever Improves Generalization of Language Models as Generic Plug-InarXiv 2023Revisiting Weighted Aggregation in Federated Learning with Neural NetworksarXiv 2023N-ImageNet: Towards Robust, Fine-Grained Object Recognition with Event Camerasn-imagenet-towards-robust-fine-grained-objectPSIMiner: A Tool for Mining Rich Abstract Syntax Trees from CodearXiv 2021A Few Thousand Translations Go a Long Way! Leveraging Pre-trained Models for African News TranslationNAACL 2022 7A Closer Look at Few-shot Classification AgainarXiv 2023DiffGraph: Heterogeneous Graph Diffusion ModelarXiv 2025DataMUX: Data Multiplexing for Neural NetworksarXiv 2022Defending LLMs against Jailbreaking Attacks via BacktranslationarXiv 2024NeuZip: Memory-Efficient Training and Inference with Dynamic Compression of Neural NetworksarXiv 2024Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language Modelsdoes-localization-inform-editing-surprisingMulti-View Masked World Models for Visual Robotic ManipulationarXiv 2023Semantic Image Manipulation Using Scene Graphssemantic-image-manipulation-using-scene-1VLKEB: A Large Vision-Language Model Knowledge Editing BenchmarkarXiv 2024SpecReason: Fast and Accurate Inference-Time Compute via Speculative ReasoningarXiv 2025On Learning Multi-Modal Forgery Representation for Diffusion Generated Video DetectionarXiv 2024BirdSet: A Large-Scale Dataset for Audio Classification in Avian BioacousticsarXiv 2024SirLLM: Streaming Infinite Retentive LLMarXiv 2024How to Evaluate Reward Models for RLHFarXiv 2024A Dataset for Interactive Vision-Language Navigation with Unknown Command FeasibilityarXiv 2022Rethinking the Role of Token Retrieval in Multi-Vector Retrievalrethinking-the-role-of-token-retrieval-inHow Useful is Self-Supervised Pretraining for Visual Tasks?how-useful-is-self-supervised-pretraining-for-1SYENet: A Simple Yet Effective Network for Multiple Low-Level Vision Tasks with Real-time Performance on Mobile DeviceICCV 2023 1STARSS23: An Audio-Visual Dataset of Spatial Recordings of Real Scenes with Spatiotemporal Annotations of Sound EventsNeurIPS 2023 11Bot or Human? Detecting ChatGPT Imposters with A Single QuestionarXiv 2023STUNT: Few-shot Tabular Learning with Self-generated Tasks from Unlabeled TablesarXiv 2023DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image GenerationarXiv 2023Weakly-supervised 3D Pose Transfer with KeypointsICCV 2023 1Tell Me More! Towards Implicit User Intention Understanding of Language Model Driven AgentsarXiv 2024SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM AgentsarXiv 2024Unified Generative Modeling of 3D Molecules via Bayesian Flow NetworksarXiv 2024OphNet: A Large-Scale Video Benchmark for Ophthalmic Surgical Workflow UnderstandingarXiv 2024Efficiently Teaching an Effective Dense Retriever with Balanced Topic Aware SamplingarXiv 2021Learning From Mistakes Makes LLM Better ReasonerarXiv 2023FABLES: Evaluating faithfulness and content selection in book-length summarizationarXiv 2024A Cheaper and Better Diffusion Language Model with Soft-Masked NoisearXiv 2023From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfacesfrom-pixels-to-ui-actions-learning-to-followTowards Quantifiable Dialogue Coherence EvaluationACL 2021 5Observational Scaling Laws and the Predictability of Language Model PerformancearXiv 2024Evaluating and Inducing Personality in Pre-trained Language ModelsarXiv 2022Delta-CoMe: Training-Free Delta-Compression with Mixed-Precision for Large Language ModelsarXiv 2024Why do Nearest Neighbor Language Models Work?arXiv 2023STICKERCONV: Generating Multimodal Empathetic Responses from ScratcharXiv 2024Perceiving and Modeling Density is All You Need for Image DehazingarXiv 2021The Life Cycle of Knowledge in Big Language Models: A SurveyarXiv 2023Unsupervised Semantic Correspondence Using Stable Diffusionunsupervised-semantic-correspondence-usingScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality DebiasingarXiv 2025Fast and Accurate Network Embeddings via Very Sparse Random ProjectionarXiv 2019PKU-DyMVHumans: A Multi-View Video Benchmark for High-Fidelity Dynamic Human ModelingCVPR 2024 1GSGAN: Adversarial Learning for Hierarchical Generation of 3D Gaussian SplatsarXiv 2024GLM-Dialog: Noise-tolerant Pre-training for Knowledge-grounded Dialogue GenerationarXiv 2023A Dataset of Information-Seeking Questions and Answers Anchored in Research PapersNAACL 2021 4HateCheck: Functional Tests for Hate Speech Detection ModelsACL 2021 5Community Detection in Bipartite Networks with Stochastic BlockmodelsarXiv 2020Spatio-Temporal Few-Shot Learning via Diffusive Neural Network GenerationarXiv 2024Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions FollowingarXiv 2024LangProp: A code optimization framework using Large Language Models applied to drivingarXiv 2024Distilling Large Vision-Language Model with Out-of-Distribution GeneralizabilityICCV 2023 1Reversible Decoupling Network for Single Image Reflection RemovalCVPR 2025 1Embodied BERT: A Transformer Model for Embodied, Language-guided Visual Task CompletionarXiv 2021ViTGaze: Gaze Following with Interaction Features in Vision TransformersarXiv 2024Towards Flexible Multi-modal Document ModelsCVPR 2023 1TrojDiff: Trojan Attacks on Diffusion Models with Diverse TargetsCVPR 2023 1Revisiting Link Prediction: A Data PerspectivearXiv 2023IvyGPT: InteractiVe Chinese pathwaY language model in medical domainarXiv 2023Continuous Deep Equilibrium Models: Training Neural ODEs faster by integrating them to InfinityarXiv 2022Rethinking Negative Instances for Generative Named Entity RecognitionarXiv 2024PPTC Benchmark: Evaluating Large Language Models for PowerPoint Task CompletionarXiv 2023Channel Importance Matters in Few-Shot Image ClassificationarXiv 2022Toward effective protection against diffusion based mimicry through score distillationarXiv 2023