All papers
Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.
FreGrad: Lightweight and Fast Frequency-aware Diffusion VocoderarXiv 2024LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and LocatingarXiv 2024The devil is in the object boundary: towards annotation-free instance segmentation using Foundation ModelsarXiv 2024Controllable Dialogue Simulation with In-Context LearningarXiv 2022Human-Robot Gym: Benchmarking Reinforcement Learning in Human-Robot
CollaborationarXiv 2023Bootstrapping Autonomous Driving Radars with Self-Supervised LearningCVPR 2024 1ACE : Off-Policy Actor-Critic with Causality-Aware Entropy RegularizationarXiv 2024Sharpness-Aware Training for FreearXiv 2022Benchmarking Mobile Device Control Agents across Diverse ConfigurationsarXiv 2024CIPHER: Cybersecurity Intelligent Penetration-testing Helper for Ethical ResearcherarXiv 2024MSAGPT: Neural Prompting Protein Structure Prediction via MSA Generative Pre-TrainingarXiv 2024Overcoming Common Flaws in the Evaluation of Selective Classification SystemsarXiv 2024High Performance Unstructured SpMM Computation Using Tensor CoresarXiv 2024D2LLM: Decomposed and Distilled Large Language Models for Semantic SearcharXiv 2024HardCoRe-NAS: Hard Constrained diffeRentiable Neural Architecture SearcharXiv 2021Flooding Spread of Manipulated Knowledge in LLM-Based Multi-Agent CommunitiesarXiv 2024Just Rank: Rethinking Evaluation with Word and Sentence SimilaritiesACL 2022 5BAM-DETR: Boundary-Aligned Moment Detection Transformer for Temporal Sentence Grounding in VideosarXiv 2023Open-Vocabulary Audio-Visual Semantic SegmentationarXiv 2024NuScenes-MQA: Integrated Evaluation of Captions and QA for Autonomous Driving Datasets using Markup AnnotationsarXiv 2023Classifier-guided Gradient Modulation for Enhanced Multimodal LearningarXiv 2024Unified Embedding Alignment for Open-Vocabulary Video Instance SegmentationarXiv 2024AdapMoE: Adaptive Sensitivity-based Expert Gating and Management for Efficient MoE InferencearXiv 2024TimberVision: A Multi-Task Dataset and Framework for Log-Component Segmentation and Tracking in Autonomous Forestry OperationsarXiv 2025Improving anatomical plausibility in medical image segmentation via hybrid graph neural networks: applications to chest x-ray analysisarXiv 2022REALTALK: A 21-Day Real-World Dataset for Long-Term ConversationarXiv 2025Multilingual Machine Translation with Large Language Models: Empirical Results and AnalysisarXiv 2023MathFusion: Enhancing Mathematic Problem-solving of LLM through Instruction FusionarXiv 2025Brazilian Portuguese Speech Recognition Using Wav2vec 2.0arXiv 2021KPTimes: A Large-Scale Dataset for Keyphrase Generation on News Documentskptimes-a-large-scale-dataset-for-keyphraseCan Machines Help Us Answering Question 16 in Datasheets, and In Turn Reflecting on Inappropriate Content?arXiv 2022Process-Driven Autoformalization in Lean 4arXiv 2024Semi-Supervised Neural System for Tagging, Parsing and Lematizationsemi-supervised-neural-system-for-taggingIdealGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language ModelsarXiv 2023Teaching Matters: Investigating the Role of Supervision in Vision TransformersCVPR 2023 1Fuse It More Deeply! A Variational Transformer with Layer-Wise Latent Variable Inference for Text Generationfuse-it-more-deeply-a-variational-transformerIntent Contrastive Learning with Cross Subsequences for Sequential RecommendationarXiv 2023The LLM SurgeonarXiv 2023Intriguing Properties of Data Attribution on Diffusion ModelsarXiv 2023CoCo-Agent: A Comprehensive Cognitive MLLM Agent for Smartphone GUI AutomationarXiv 2024HumanEdit: A High-Quality Human-Rewarded Dataset for Instruction-based Image EditingarXiv 2024Removing Bias in Multi-modal Classifiers: Regularization by Maximizing Functional EntropiesNeurIPS 2020 12Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?arXiv 2025MultiModN- Multimodal, Multi-Task, Interpretable Modular NetworksarXiv 2023ConsNet: Learning Consistency Graph for Zero-Shot Human-Object Interaction DetectionarXiv 2020Bayesian Estimation of Differential PrivacyarXiv 2022Contrastive Energy Prediction for Exact Energy-Guided Diffusion Sampling in Offline Reinforcement LearningarXiv 2023MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language ModelsarXiv 2024Interpretable Image Classification with Adaptive Prototype-based Vision TransformersarXiv 2024A Hybrid ANN-SNN Architecture for Low-Power and Low-Latency Visual PerceptionarXiv 2023Repeat After Me: Transformers are Better than State Space Models at CopyingarXiv 2024Rapid Response: Mitigating LLM Jailbreaks with a Few ExamplesarXiv 2024Point-GCC: Universal Self-supervised 3D Scene Pre-training via Geometry-Color ContrastarXiv 2023Computationally-Efficient Neural Image Compression with Shallow Decoderscomputationally-efficient-neural-image-1Oscillation-free Quantization for Low-bit Vision TransformersarXiv 2023An Efficient Memory-Augmented Transformer for Knowledge-Intensive NLP TasksarXiv 2022MMGDreamer: Mixed-Modality Graph for Geometry-Controllable 3D Indoor Scene Generationmmgdreamer-mixed-modality-graph-for-geometryText is no more Enough! A Benchmark for Profile-based Spoken Language UnderstandingarXiv 2021SADGA: Structure-Aware Dual Graph Aggregation Network for Text-to-SQLNeurIPS 2021 12DUAL: Discrete Spoken Unit Adaptive Learning for Textless Spoken Question AnsweringarXiv 2022Text Detoxification using Large Pre-trained Neural ModelsEMNLP 2021 11XFT: Unlocking the Power of Code Instruction Tuning by Simply Merging Upcycled Mixture-of-ExpertsarXiv 2024VisualWebInstruct: Scaling up Multimodal Instruction Data through Web SearcharXiv 2025QAmeleon: Multilingual QA with Only 5 ExamplesarXiv 2022Convex Aggregation for Opinion SummarizationFindings (EMNLP) 2021 11ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language ModelsarXiv 2024ESCoT: Towards Interpretable Emotional Support Dialogue SystemsarXiv 2024MOSO: Decomposing MOtion, Scene and Object for Video PredictionCVPR 2023 1Distillation Contrastive Decoding: Improving LLMs Reasoning with Contrastive Decoding and DistillationarXiv 2024RoBERTuito: a pre-trained language model for social media text in SpanishLREC 2022 6Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-Based DecodingarXiv 2024Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?arXiv 2025Learning Cooperative Trajectory Representations for Motion ForecastingarXiv 2023Constant Acceleration Flowconstant-acceleration-flowTowards Rationality in Language and Multimodal Agents: A SurveyarXiv 2024A Self-Training Method for Machine Reading Comprehension with Soft Evidence Extractiona-self-training-method-for-machine-reading-1Multi-step Jailbreaking Privacy Attacks on ChatGPTarXiv 2023SparQ Attention: Bandwidth-Efficient LLM InferencearXiv 2023BioTrove: A Large Curated Image Dataset Enabling AI for BiodiversityarXiv 2024Type-Aware Decomposed Framework for Few-Shot Named Entity RecognitionarXiv 2023Blackout Diffusion: Generative Diffusion Models in Discrete-State SpacesarXiv 2023In-Context Alignment: Chat with Vanilla Language Models Before Fine-TuningarXiv 2023DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code RepositoriesarXiv 2024Frame Interpolation with Consecutive Brownian Bridge DiffusionarXiv 2024Prompted LLMs as Chatbot Modules for Long Open-domain ConversationarXiv 2023Supervised Fine-tuning in turn Improves Visual Foundation ModelsarXiv 2024End-to-End Semi-Supervised Learning for Video Action DetectionCVPR 2022 1THQA: A Perceptual Quality Assessment Database for Talking HeadsarXiv 2024MedCoT: Medical Chain of Thought via Hierarchical ExpertarXiv 2024Injecting Domain Knowledge in Language Models for Task-Oriented Dialogue SystemsarXiv 2022Bias Loss for Mobile Neural NetworksICCV 2021 10LaserHuman: Language-guided Scene-aware Human Motion Generation in Free EnvironmentarXiv 2024LoLI-Street: Benchmarking Low-Light Image Enhancement and BeyondarXiv 2024HAGRID: A Human-LLM Collaborative Dataset for Generative Information-Seeking with AttributionarXiv 2023ChatEDA: A Large Language Model Powered Autonomous Agent for EDAarXiv 2023Generative causal explanations of black-box classifiersNeurIPS 2020 12TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation ModelsarXiv 2024Improved off-policy training of diffusion samplersarXiv 2024Large Language Model Unlearning via Embedding-Corrupted PromptsarXiv 2024Enhancing Conversational Search: Large Language Model-Aided Informative Query RewritingarXiv 2023Self-Attention for Audio Super-ResolutionarXiv 2021GAIA: Rethinking Action Quality Assessment for AI-Generated VideosarXiv 2024Bi-directional Distribution Alignment for Transductive Zero-Shot LearningCVPR 2023 1Automatic High Resolution Wire Segmentation and RemovalCVPR 2023 1PartGlot: Learning Shape Part Segmentation from Language Reference GamesCVPR 2022 1Extract-and-Adaptation Network for 3D Interacting Hand Mesh RecoveryarXiv 2023Individualizing Glioma Radiotherapy Planning by Optimization of Data and
Physics-Informed Discrete LossarXiv 2023Semi-Supervised Raw-to-Raw MappingarXiv 2021Adaptive Advantage-Guided Policy Regularization for Offline Reinforcement LearningarXiv 2024ADNet: Lane Shape Prediction via Anchor DecompositionICCV 2023 1ILIAS: Instance-Level Image retrieval At ScaleCVPR 2025 1CorpusBrain: Pre-train a Generative Retrieval Model for Knowledge-Intensive Language TasksarXiv 2022OTSeg: Multi-prompt Sinkhorn Attention for Zero-Shot Semantic SegmentationarXiv 2024Sarcasm Detection using Hybrid Neural NetworkarXiv 2019How to Continually Adapt Text-to-Image Diffusion Models for Flexible Customization?arXiv 2024Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis RefinementarXiv 2023Self-Correcting Self-Consuming Loops for Generative Model TrainingarXiv 2024Multimodal Situational SafetyarXiv 2024Towards Understanding Mixture of Experts in Deep LearningarXiv 2022Transformer Meets Boundary Value Inverse ProblemsarXiv 2022CAMEL-Bench: A Comprehensive Arabic LMM BenchmarkarXiv 2024Advancing Molecular Machine Learning Representations with Stereoelectronics-Infused Molecular GraphsarXiv 2024FusionDTI: Fine-grained Binding Discovery with Token-level Fusion for Drug-Target InteractionarXiv 2024Trained on 100 million words and still in shape: BERT meets British National CorpusarXiv 2023Exploring Architectural Ingredients of Adversarially Robust Deep Neural NetworksNeurIPS 2021 12UniFashion: A Unified Vision-Language Model for Multimodal Fashion Retrieval and GenerationarXiv 2024Interactive Text-to-Image Retrieval with Large Language Models: A Plug-and-Play ApproacharXiv 2024ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language ModelsarXiv 2024Emergence of In-Context Reinforcement Learning from Noise DistillationarXiv 2023In-Context Explainers: Harnessing LLMs for Explaining Black Box ModelsarXiv 2023MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer DecodingarXiv 2024BeHonest: Benchmarking Honesty in Large Language ModelsarXiv 2024A Framework for Fast and Stable Representations of Multiparameter Persistent Homology Decompositionsa-framework-for-fast-and-stableReconstructed Convolution Module Based Look-Up Tables for Efficient Image Super-ResolutionICCV 2023 1$\textit{latent}$-GLAT: Glancing at Latent Variables for Parallel Text GenerationarXiv 2022EG4D: Explicit Generation of 4D Object without Score DistillationarXiv 2024OPD: Single-view 3D Openable Part DetectionarXiv 2022Interactive Task Planning with Language ModelsarXiv 2023Multi-Label Knowledge DistillationICCV 2023 1Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal DataarXiv 2024Review of deep learning models for crypto price prediction: implementation and evaluationarXiv 2024NegBERT: A Transfer Learning Approach for Negation Detection and Scope Resolutionnegbert-a-transfer-learning-approach-for-1Larimar: Large Language Models with Episodic Memory ControlarXiv 2024Plug-and-Play Regulators for Image-Text MatchingarXiv 2023ProtoQA: A Question Answering Dataset for Prototypical Common-Sense ReasoningEMNLP 2020 11IncreLoRA: Incremental Parameter Allocation Method for Parameter-Efficient Fine-tuningarXiv 2023Generalization in NLI: Ways (Not) To Go Beyond Simple HeuristicsEMNLP (insights) 2021 11Low-rank passthrough neural networkslow-rank-passthrough-neural-networks-1AD-YOLO: You Look Only Once in Training Multiple Sound Event Localization and DetectionarXiv 2023Addressing "Documentation Debt" in Machine Learning Research: A Retrospective Datasheet for BookCorpusarXiv 2021Uncertainty-aware Unsupervised Multi-Object TrackingICCV 2023 1Lottery Jackpots Exist in Pre-trained ModelsarXiv 2021A Step Towards Worldwide Biodiversity Assessment: The BIOSCAN-1M Insect Dataseta-step-towards-worldwide-biodiversity-1APE: Faster and Longer Context-Augmented Generation via Adaptive Parallel EncodingarXiv 2025Practical No-box Adversarial Attacks against DNNspractical-no-box-adversarial-attacks-againstCommunication-Efficient Collaborative Perception via Information Filling with CodebookCVPR 2024 1Bootstrapping Complete The Look at PinterestarXiv 2020A Two-Stage Adaptation of Large Language Models for Text RankingarXiv 2023RP-DNN: A Tweet level propagation context based deep neural networks for early rumor detection in Social Mediarp-dnn-a-tweet-level-propagation-context-1Aligning Large Language Models with Representation Editing: A Control PerspectivearXiv 2024Can Large Vision Language Models Read Maps Like a Human?arXiv 2025Unified Normalization for Accelerating and Stabilizing TransformersarXiv 2022BARTSmiles: Generative Masked Language Models for Molecular RepresentationsarXiv 2022Making Vision Transformers Efficient from A Token Sparsification ViewCVPR 2023 1PeaCoK: Persona Commonsense Knowledge for Consistent and Engaging NarrativesarXiv 2023Jatmo: Prompt Injection Defense by Task-Specific FinetuningarXiv 2023AfroLID: A Neural Language Identification Tool for African LanguagesarXiv 2022Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller, Faster, and BetterarXiv 2021Longitudinal Segmentation of MS Lesions via Temporal Difference WeightingarXiv 2024PhysVLM: Enabling Visual Language Models to Understand Robotic Physical ReachabilityCVPR 2025 1Domain Adaptive Video Segmentation via Temporal Pseudo SupervisionarXiv 2022Extend Model Merging from Fine-Tuned to Pre-Trained Large Language Models via Weight DisentanglementarXiv 2024QuestBench: Can LLMs ask the right question to acquire information in reasoning tasks?arXiv 2025TaxaBind: A Unified Embedding Space for Ecological ApplicationsarXiv 2024DailyDVS-200: A Comprehensive Benchmark Dataset for Event-Based Action RecognitionarXiv 2024Domain-General Crowd Counting in Unseen ScenariosarXiv 2022Text2FX: Harnessing CLAP Embeddings for Text-Guided Audio EffectsarXiv 2024Learning with Mixture of Prototypes for Out-of-Distribution DetectionarXiv 2024AlignGPT: Multi-modal Large Language Models with Adaptive Alignment CapabilityarXiv 2024SVFT: Parameter-Efficient Fine-Tuning with Singular VectorsarXiv 2024FinEAS: Financial Embedding Analysis of SentimentarXiv 2021Equiangular Basis VectorsCVPR 2023 1Tell What You Hear From What You See -- Video to Audio Generation Through TextarXiv 2024MotionMix: Weakly-Supervised Diffusion for Controllable Motion GenerationarXiv 2024Your Transformer May Not be as Powerful as You ExpectarXiv 2022Abstractive Meeting Summarization: A SurveyarXiv 2022DeltaZip: Efficient Serving of Multiple Full-Model-Tuned LLMsarXiv 2023Sample-Efficient Automated Deep Reinforcement LearningICLR 2021 1ProtoGCD: Unified and Unbiased Prototype Learning for Generalized Category DiscoveryarXiv 2025ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question AnsweringarXiv 2025Scalable Attentive Sentence-Pair Modeling via Distilled Sentence EmbeddingarXiv 2019Towards Generic Image Manipulation Detection with Weakly-Supervised Self-Consistency Learningtowards-generic-image-manipulation-detectionPromptCARE: Prompt Copyright Protection by Watermark Injection and
VerificationarXiv 2023REAR: A Relevance-Aware Retrieval-Augmented Framework for Open-Domain Question AnsweringarXiv 2024Precipitation nowcasting with generative diffusion modelsarXiv 2023Redefining Temporal Modeling in Video Diffusion: The Vectorized Timestep ApproacharXiv 2024ContractNLI: A Dataset for Document-level Natural Language Inference for ContractsFindings (EMNLP) 2021 11A Tale of Two Graphs: Freezing and Denoising Graph Structures for Multimodal RecommendationarXiv 2022Multimodal Clustering Networks for Self-supervised Learning from Unlabeled VideosICCV 2021 10Deep Active Learning in Remote Sensing for data efficient Change DetectionarXiv 2020