All papers
Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.
Pitfalls in Language Models for Code Intelligence: A Taxonomy and SurveyarXiv 2023Unsupervised Deep Learning-based Pansharpening with Jointly-Enhanced Spectral and Spatial FidelityarXiv 2023Retro-FPN: Retrospective Feature Pyramid Network for Point Cloud Semantic SegmentationICCV 2023 1DotaMath: Decomposition of Thought with Code Assistance and Self-correction for Mathematical ReasoningarXiv 2024The Magic of IF: Investigating Causal Reasoning Abilities in Large Language Models of CodearXiv 2023Visually Descriptive Language Model for Vector Graphics ReasoningarXiv 2024Dancing with Still Images: Video Distillation via Static-Dynamic DisentanglementCVPR 2024 1Attentive WaveBlock: Complementarity-enhanced Mutual Networks for Unsupervised Domain Adaptation in Person Re-identification and BeyondarXiv 2020Deep Learning for Functional Data Analysis with Adaptive Basis LayersarXiv 2021MAPL: Parameter-Efficient Adaptation of Unimodal Pre-Trained Models for Vision-Language Few-Shot PromptingarXiv 2022Improving Retrieval-Augmented Large Language Models via Data Importance LearningarXiv 2023Improving Hateful Meme Detection through Retrieval-Guided Contrastive LearningarXiv 2023MagDiff: Multi-Alignment Diffusion for High-Fidelity Video Generation and EditingarXiv 2023Merging Multi-Task Models via Weight-Ensembling Mixture of ExpertsarXiv 2024Self-supervised learning of Split Invariant Equivariant representationsarXiv 2023On the Evaluation of Commit Message Generation Models: An Experimental StudyarXiv 2021IT5: Text-to-text Pretraining for Italian Language Understanding and GenerationarXiv 2022AMO Sampler: Enhancing Text Rendering with OvershootingCVPR 2025 1The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-DeterminismarXiv 2024FrustumFormer: Adaptive Instance-aware Resampling for Multi-view 3D DetectionCVPR 2023 1Astroformer: More Data Might not be all you need for ClassificationarXiv 2023EvalTree: Profiling Language Model Weaknesses via Hierarchical Capability TreesarXiv 2025InstructAny2Pix: Flexible Visual Editing via Multimodal Instruction FollowingarXiv 2023MAtch, eXpand and Improve: Unsupervised Finetuning for Zero-Shot Action Recognition with Language KnowledgeICCV 2023 1Isomer: Isomerous Transformer for Zero-shot Video Object SegmentationICCV 2023 1Comparing Dataset Characteristics that Favor the Apriori, Eclat or FP-Growth Frequent Itemset Mining AlgorithmsarXiv 2017LiST: Lite Prompted Self-training Makes Parameter-Efficient Few-shot Learnerslist-lite-self-training-makes-efficient-fewOmniTab: Pretraining with Natural and Synthetic Data for Few-shot Table-based Question Answeringomnitab-pretraining-with-natural-andNet2Vec: Quantifying and Explaining how Concepts are Encoded by Filters in Deep Neural Networksnet2vec-quantifying-and-explaining-how-1Learning Neural Templates for Recommender Dialogue SystemEMNLP 2021 11Fault-Aware Neural Code RankersarXiv 2022Vlogger: Make Your Dream A VlogCVPR 2024 1A Survey of Medical Vision-and-Language Applications and Their TechniquesarXiv 2024Robust Adaptation of Large Multimodal Models for Retrieval Augmented Hateful Meme Detectionimproved-fine-tuning-of-large-multimodalFast and Effective Weight Update for Pruned Large Language ModelsarXiv 2024Pseudo-label Alignment for Semi-supervised Instance SegmentationICCV 2023 1ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and RobustnessarXiv 2025HoneyBee: Progressive Instruction Finetuning of Large Language Models for Materials SciencearXiv 2023Optimistic Curiosity Exploration and Conservative Exploitation with Linear Reward ShapingarXiv 2022Disentangled Representation Learning for RF Fingerprint Extraction under Unknown Channel StatisticsarXiv 2022Do Pedestrians Pay Attention? Eye Contact Detection in the WildarXiv 2021Multimodal-Conditioned Latent Diffusion Models for Fashion Image EditingarXiv 2024Measuring Data Diversity for Instruction Tuning: A Systematic Analysis and A Reliable MetricarXiv 2025Prompt Switch: Efficient CLIP Adaptation for Text-Video RetrievalICCV 2023 1Automatic Model Selection with Large Language Models for ReasoningarXiv 2023Meta-Learning with Fewer Tasks through Task Interpolationmeta-learning-with-fewer-tasks-through-task-1Multimodal Detection of Unknown Objects on Roads for Autonomous DrivingarXiv 2022MuMiN: A Large-Scale Multilingual Multimodal Fact-Checked Misinformation Social Network DatasetarXiv 2022The Role of Entropy and Reconstruction in Multi-View Self-Supervised LearningarXiv 2023Optimizing DDPM Sampling with Shortcut Fine-TuningarXiv 2023EQ-Net: Elastic Quantization Neural NetworksICCV 2023 1Can Deep Learning be Applied to Model-Based Multi-Object Tracking?arXiv 20223D VR Sketch Guided 3D Shape Prototyping and ExplorationICCV 2023 1SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of ExpertsICCV 2025RF-ULM: Ultrasound Localization Microscopy Learned from Radio-Frequency WavefrontsarXiv 2023MQAG: Multiple-choice Question Answering and Generation for Assessing Information Consistency in SummarizationarXiv 2023Chatting Makes Perfect: Chat-based Image Retrievalchatting-makes-perfect-chat-based-imageASCEND: A Spontaneous Chinese-English Dataset for Code-switching in Multi-turn ConversationLREC 2022 6Region-Adaptive Transform with Segmentation Prior for Image CompressionarXiv 2024Decoupled Textual Embeddings for Customized Image GenerationarXiv 2023Refusal in LLMs is an Affine FunctionarXiv 2024Conditional GANs with Auxiliary Discriminative Classifierconditional-gans-with-auxiliaryElastic Feature Consolidation for Cold Start Exemplar-Free Incremental LearningarXiv 2024Cross-Domain Few-Shot Segmentation via Iterative Support-Query Correspondence MiningCVPR 2024 1Model Surgery: Modulating LLM's Behavior Via Simple Parameter EditingarXiv 2024How Good is Google Bard's Visual Understanding? An Empirical Study on Open ChallengesarXiv 2023Generative Modeling of Regular and Irregular Time Series Data via Koopman VAEsarXiv 2023Deblurring Masked Autoencoder is Better Recipe for Ultrasound Image RecognitionarXiv 2023Intuitive Fine-Tuning: Towards Simplifying Alignment into a Single ProcessarXiv 2024Content-Based Collaborative Generation for Recommender SystemsarXiv 2024MixPath: A Unified Approach for One-shot Neural Architecture SearchICCV 2023 1Hydra: Multi-head Low-rank Adaptation for Parameter Efficient Fine-tuningarXiv 2023Hierarchical Verbalizer for Few-Shot Hierarchical Text ClassificationarXiv 2023Lifting the Curse of Capacity Gap in Distilling Language ModelsarXiv 2023SciEval: A Multi-Level Large Language Model Evaluation Benchmark for Scientific ResearcharXiv 2023X-Mesh: Towards Fast and Accurate Text-driven 3D Stylization via Dynamic Textual GuidanceICCV 2023 1An Empirical Study of GPT-4o Image Generation CapabilitiesarXiv 2025A Closer Look at Fourier Spectrum Discrepancies for CNN-generated Images DetectionCVPR 2021 1CWEval: Outcome-driven Evaluation on Functionality and Security of LLM Code GenerationarXiv 2025Plain-Det: A Plain Multi-Dataset Object DetectorarXiv 2024Robust Self-Augmentation for Named Entity Recognition with Meta ReweightingNAACL 2022 7Rethinking Model Selection and Decoding for Keyphrase Generation with Pre-trained Sequence-to-Sequence ModelsarXiv 2023Strategic Preys Make Acute Predators: Enhancing Camouflaged Object Detectors by Generating Camouflaged ObjectsarXiv 2023Diffusion Guided Language ModelingarXiv 2024Torch.manual_seed(3407) is all you need: On the influence of random seeds in deep learning architectures for computer visionarXiv 2021Multimodal Pretraining for Dense Video CaptioningAsian Chapter of the Association for Computational Linguistics 2020InsViE-1M: Effective Instruction-based Video Editing with Elaborate Dataset ConstructionarXiv 2025WaterBench: Towards Holistic Evaluation of Watermarks for Large Language ModelsarXiv 2023Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop QueriesarXiv 2024Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-AlignmentarXiv 2023Watset: Local-Global Graph Clustering with Applications in Sense and Frame Inductionwatset-local-global-graph-clustering-with-1Physics-Informed Deep Neural Network Method for Limited Observability
State EstimationarXiv 2019VNHSGE: VietNamese High School Graduation Examination Dataset for Large Language ModelsarXiv 2023CLIM: Contrastive Language-Image Mosaic for Region RepresentationarXiv 2023VersiCode: Towards Version-controllable Code GenerationarXiv 2024Learning to Maximize Mutual Information for Dynamic Feature SelectionarXiv 2023Enhancing CLIP with GPT-4: Harnessing Visual Descriptions as PromptsarXiv 2023LaLaLoc: Latent Layout Localisation in Dynamic, Unvisited EnvironmentsICCV 2021 10ZS4IE: A toolkit for Zero-Shot Information Extraction with simple VerbalizationsNAACL (ACL) 2022 7FedEx-LoRA: Exact Aggregation for Federated and Efficient Fine-Tuning of Foundation ModelsarXiv 2024Online Domain Adaptation for Semantic Segmentation in Ever-Changing ConditionsarXiv 2022D3: A Massive Dataset of Scholarly Metadata for Analyzing the State of Computer Science ResearchLREC 2022 6Reward Reports for Reinforcement LearningarXiv 2022Improving performance of real-time full-band blind packet-loss
concealment with predictive networkarXiv 2022Developing Retrieval Augmented Generation (RAG) based LLM Systems from PDFs: An Experience ReportarXiv 2024ProSA: Assessing and Understanding the Prompt Sensitivity of LLMsarXiv 2024RoScenes: A Large-scale Multi-view 3D Dataset for Roadside PerceptionarXiv 2024Reward-Free Curricula for Training Robust World ModelsarXiv 2023Probabilistic Emulation of a Global Climate Model with Spherical DYffusionarXiv 2024ERASE: Benchmarking Feature Selection Methods for Deep Recommender SystemsarXiv 2024Hierarchical Point-based Active Learning for Semi-supervised Point Cloud Semantic SegmentationICCV 2023 1Datasets for Studying Generalization from Easy to Hard ExamplesarXiv 2021Influence Selection for Active LearningICCV 2021 10ALOHA: Artificial Learning of Human Attributes for Dialogue AgentsarXiv 2019GIST: Generating Image-Specific Text for Fine-grained Object ClassificationarXiv 2023GenView: Enhancing View Quality with Pretrained Generative Model for Self-Supervised LearningarXiv 2024VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?CVPR 2025 1LLaVA-Chef: A Multi-modal Generative Model for Food RecipesarXiv 2024CREATOR: Tool Creation for Disentangling Abstract and Concrete Reasoning of Large Language ModelsarXiv 2023Reasoning with Reinforced Functional Token TuningarXiv 2025Online Continual Learning For Interactive Instruction Following AgentsarXiv 2024RuMedBench: A Russian Medical Language Understanding BenchmarkarXiv 2022TDAG: A Multi-Agent Framework based on Dynamic Task Decomposition and Agent GenerationarXiv 2024MagMax: Leveraging Model Merging for Seamless Continual LearningarXiv 2024CR3DT: Camera-RADAR Fusion for 3D Detection and TrackingarXiv 2024SCP-116K: A High-Quality Problem-Solution Dataset and a Generalized Pipeline for Automated Extraction in the Higher Education Science DomainarXiv 2025Quality Not Quantity: On the Interaction between Dataset Design and Robustness of CLIParXiv 2022SeFAR: Semi-supervised Fine-grained Action Recognition with Temporal Perturbation and Learning StabilizationarXiv 2025Learning From Failure: Integrating Negative Examples when Fine-tuning Large Language Models as AgentsarXiv 2024GLEN: Generative Retrieval via Lexical Index LearningarXiv 2023RLVF: Learning from Verbal Feedback without OvergeneralizationarXiv 2024CAMIL: Context-Aware Multiple Instance Learning for Cancer Detection and Subtyping in Whole Slide ImagesarXiv 2023A Whisper transformer for audio captioning trained with synthetic captions and transfer learningarXiv 2023POUF: Prompt-oriented unsupervised fine-tuning for large pre-trained modelsarXiv 2023Ambient Diffusion Omni: Training Good Models with Bad DataarXiv 2025Comprehensive Layer-wise Analysis of SSL Models for Audio Deepfake DetectionarXiv 2025STaR-GATE: Teaching Language Models to Ask Clarifying QuestionsarXiv 2024DyFADet: Dynamic Feature Aggregation for Temporal Action DetectionarXiv 2024A multi-centre polyp detection and segmentation dataset for generalisability assessmentarXiv 2021Everybody Prune Now: Structured Pruning of LLMs with only Forward PassesarXiv 2024Efficient and Scalable Estimation of Tool Representations in Vector SpacearXiv 2024Forecasting Bitcoin volatility spikes from whale transactions and CryptoQuant data using Synthesizer Transformer modelsarXiv 2022ThinK: Thinner Key Cache by Query-Driven PruningarXiv 2024Learning to Identify Critical States for Reinforcement Learning from VideosICCV 2023 1MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in VideosarXiv 2024Learning Support and Trivial Prototypes for Interpretable Image ClassificationICCV 2023 1PromptIQA: Boosting the Performance and Generalization for No-Reference Image Quality Assessment via PromptsarXiv 2024View-Invariant Policy Learning via Zero-Shot Novel View SynthesisarXiv 2024Improving Pretraining Data Using Perplexity CorrelationsarXiv 2024OptiBench Meets ReSocratic: Measure and Improve LLMs for Optimization ModelingarXiv 2024Adapting and Evaluating Influence-Estimation Methods for Gradient-Boosted Decision TreesarXiv 2022Every child should have parents: a taxonomy refinement algorithm based on hyperbolic term embeddingsevery-child-should-have-parents-a-taxonomy-1GenRC: Generative 3D Room Completion from Sparse Image CollectionsarXiv 2024Can Brain Signals Reveal Inner Alignment with Human Languages?arXiv 2022GateLoop: Fully Data-Controlled Linear Recurrence for Sequence ModelingarXiv 2023CLIP: A Dataset for Extracting Action Items for Physicians from Hospital Discharge NotesACL 2021 5Fine-tuning large language models for domain adaptation: Exploration of training strategies, scaling, model merging and synergistic capabilitiesarXiv 2024HonestLLM: Toward an Honest and Helpful Large Language ModelarXiv 2024Quantized Spike-driven TransformerarXiv 2025Scaling Up Dataset Distillation to ImageNet-1K with Constant MemoryarXiv 2022Spanish Legalese Language Model and CorporaarXiv 2021Pix2Shape: Towards Unsupervised Learning of 3D Scenes from Images using a View-based RepresentationarXiv 2020No Task Left Behind: Isotropic Model Merging with Common and Task-Specific SubspacesarXiv 2025Exposing Numeracy Gaps: A Benchmark to Evaluate Fundamental Numerical Abilities in Large Language ModelsarXiv 2025Overthinking the Truth: Understanding how Language Models Process False DemonstrationsarXiv 2023FlipNeRF: Flipped Reflection Rays for Few-shot Novel View SynthesisICCV 2023 1Time Does Tell: Self-Supervised Time-Tuning of Dense Image RepresentationsICCV 2023 1LongGenBench: Benchmarking Long-Form Generation in Long Context LLMsarXiv 2024Peering Through Preferences: Unraveling Feedback Acquisition for Aligning Large Language ModelsarXiv 2023LaProp: Separating Momentum and Adaptivity in AdamarXiv 2020TALC: Time-Aligned Captions for Multi-Scene Text-to-Video GenerationarXiv 2024Question Decomposition Improves the Faithfulness of Model-Generated ReasoningarXiv 2023ChangeChip: A Reference-Based Unsupervised Change Detection for PCB Defect DetectionarXiv 2021BACKTIME: Backdoor Attacks on Multivariate Time Series ForecastingarXiv 2024Meta-Learning Dynamics Forecasting Using Task Inferencemeta-learning-dynamics-forecasting-using-task-1OPT-Tree: Speculative Decoding with Adaptive Draft Tree StructurearXiv 2024MARIO: MAth Reasoning with code Interpreter Output -- A Reproducible PipelinearXiv 2024Entropy is not Enough for Test-Time Adaptation: From the Perspective of Disentangled FactorsarXiv 2024TimeX++: Learning Time-Series Explanations with Information BottleneckarXiv 2024Entity-Centric Reinforcement Learning for Object Manipulation from PixelsarXiv 2024Defending Large Language Models Against Jailbreaking Attacks Through Goal PrioritizationarXiv 2023BENO: Boundary-embedded Neural Operators for Elliptic PDEsarXiv 2024No Parameter Left Behind: How Distillation and Model Size Affect Zero-Shot RetrievalarXiv 2022PanNuke Dataset Extension, Insights and BaselinesarXiv 2020Language-guided Human Motion Synthesis with Atomic ActionsarXiv 2023AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task GenerationarXiv 2024Q-Refine: A Perceptual Quality Refiner for AI-Generated ImagearXiv 2024Mitigating Data Sparsity for Short Text Topic Modeling by Topic-Semantic Contrastive LearningarXiv 2022OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring ModelingarXiv 2024Towards Latent Masked Image Modeling for Self-Supervised Visual Representation LearningarXiv 2024Enhancing LLM Safety Through a Theoretical Minimax Game LensarXiv 2025Troika: Multi-Path Cross-Modal Traction for Compositional Zero-Shot LearningCVPR 2024 1SLAck: Semantic, Location, and Appearance Aware Open-Vocabulary TrackingarXiv 2024Model Merging by Uncertainty-Based Gradient MatchingarXiv 2023The DEVIL is in the Details: A Diagnostic Evaluation Benchmark for Video InpaintingCVPR 2022 1ForgerySleuth: Empowering Multimodal Large Language Models for Image Manipulation DetectionarXiv 2024TwHIN-BERT: A Socially-Enriched Pre-trained Language Model for Multilingual Tweet Representations at TwitterarXiv 2022Multi-Task Recommendations with Reinforcement LearningarXiv 2023Marked Personas: Using Natural Language Prompts to Measure Stereotypes in Language ModelsarXiv 2023Posterior Sampling for Deep Reinforcement LearningarXiv 2023