0

All papers

Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.

On the Potential of Lexico-logical Alignments for Semantic Parsing to SQL QueriesarXiv 2020Learning Procedure-aware Video Representation from Instructional Videos and Their NarrationsCVPR 2023 1Uncertainty-guided Perturbation for Image Super-Resolution Diffusion ModelCVPR 2025 1PlantSeg: A Large-Scale In-the-wild Dataset for Plant Disease SegmentationarXiv 2024AMU-Tuning: Effective Logit Bias for CLIP-based Few-shot LearningCVPR 2024 1RNb-NeuS: Reflectance and Normal-based Multi-View 3D ReconstructionCVPR 2024 1Towards Self-Assembling Artificial Neural Networks through Neural Developmental ProgramsarXiv 2023Bioformer: an efficient transformer language model for biomedical text miningarXiv 2023LinkAlign: Scalable Schema Linking for Real-World Large-Scale Multi-Database Text-to-SQLarXiv 2025RaLLe: A Framework for Developing and Evaluating Retrieval-Augmented Large Language ModelsarXiv 2023QuEST: Low-bit Diffusion Model Quantization via Efficient Selective FinetuningICCV 2025AttT2M: Text-Driven Human Motion Generation with Multi-Perspective Attention Mechanismattt2m-text-driven-human-motion-generationRI3D: Few-Shot Gaussian Splatting With Repair and Inpainting Diffusion PriorsICCV 2025Video Diffusion Models: A SurveyarXiv 2024BAMM: Bidirectional Autoregressive Motion ModelarXiv 2024SSM-DTA: Breaking the Barriers of Data Scarcity in Drug-Target Affinity PredictionarXiv 2022PMET: Precise Model Editing in a TransformerarXiv 2023MRN: Multiplexed Routing Network for Incremental Multilingual Text RecognitionICCV 2023 1Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language ModelsarXiv 2025DragVideo: Interactive Drag-style Video EditingarXiv 2023Generative Disco: Text-to-Video Generation for Music VisualizationarXiv 2023How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden StatesarXiv 2024Fast Machine Unlearning Without Retraining Through Selective Synaptic DampeningarXiv 2023Retrieval-Augmented Dynamic Prompt Tuning for Incomplete Multimodal LearningarXiv 2025Dreamer XL: Towards High-Resolution Text-to-3D Generation via Trajectory Score MatchingarXiv 2024QAFactEval: Improved QA-Based Factual Consistency Evaluation for SummarizationNAACL 2022 7ExPLoRA: Parameter-Efficient Extended Pre-Training to Adapt Vision Transformers under Domain ShiftsarXiv 2024Visual Perception by Large Language Model's WeightsarXiv 2024FIRST: Faster Improved Listwise Reranking with Single Token DecodingarXiv 2024Weak-to-Strong Diffusion with ReflectionarXiv 2025Struc-Bench: Are Large Language Models Really Good at Generating Complex Structured Data?arXiv 2023UniMD: Towards Unifying Moment Retrieval and Temporal Action DetectionarXiv 2024Large Scale Crowdsourcing and Characterization of Twitter Abusive BehaviorarXiv 2018OnlineRefer: A Simple Online Baseline for Referring Video Object SegmentationICCV 2023 1DeepFace-EMD: Re-ranking Using Patch-wise Earth Mover's Distance Improves Out-Of-Distribution Face IdentificationCVPR 2022 1Unified Discrete Diffusion for Simultaneous Vision-Language GenerationarXiv 2022diffGrad: An Optimization Method for Convolutional Neural NetworksarXiv 2019DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated TextarXiv 2023DVERGE: Diversifying Vulnerabilities for Enhanced Robust Generation of EnsemblesNeurIPS 2020 12HyDe: The First Open-Source, Python-Based, GPU-Accelerated Hyperspectral Denoising PackagearXiv 2022Evaluating the Ripple Effects of Knowledge Editing in Language ModelsarXiv 2023Empowering Diffusion Models on the Embedding Space for Text GenerationarXiv 2022Style-A-Video: Agile Diffusion for Arbitrary Text-based Video Style TransferarXiv 2023Is Disentanglement all you need? Comparing Concept-based & Disentanglement ApproachesarXiv 2021SEAL: A Framework for Systematic Evaluation of Real-World Super-ResolutionarXiv 2023Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language ModelsarXiv 2024BLSP-Emo: Towards Empathetic Large Speech-Language ModelsarXiv 2024Re-TACRED: Addressing Shortcomings of the TACRED DatasetarXiv 2021Spatial Implicit Neural Representations for Global-Scale Species MappingarXiv 2023A Low-Shot Object Counting Network With Iterative Prototype AdaptationICCV 2023 1Large-Scale Spatio-Temporal Person Re-identification: Algorithms and BenchmarkarXiv 2021ZBS: Zero-shot Background Subtraction via Instance-level Background Modeling and Foreground SelectionCVPR 2023 1FitMe: Deep Photorealistic 3D Morphable Model AvatarsCVPR 2023 1LMPT: Prompt Tuning with Class-Specific Embedding Loss for Long-tailed Multi-Label Visual RecognitionarXiv 2023Knowing Where to Focus: Event-aware Transformer for Video GroundingICCV 2023 1DEMix Layers: Disentangling Domains for Modular Language ModelingNAACL 2022 7EmbodiedEval: Evaluate Multimodal LLMs as Embodied AgentsarXiv 2025Is ChatGPT a Good Sentiment Analyzer? A Preliminary StudyarXiv 2023GLEU Without TuningarXiv 2016Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic SegmentationarXiv 2024Zero-shot Composed Text-Image RetrievalarXiv 2023Combining Modular Skills in Multitask LearningarXiv 2022High-fidelity Person-centric Subject-to-Image SynthesisCVPR 2024 1Robust Attentional Aggregation of Deep Feature Sets for Multi-view 3D ReconstructionarXiv 2018ICL-D3IE: In-Context Learning with Diverse Demonstrations Updating for Document Information ExtractionICCV 2023 1Dynamic Slate Recommendation with Gated Recurrent Units and Thompson SamplingarXiv 2021Stochastic Multi-Person 3D Motion ForecastingarXiv 2023Retrieval-Augmented Score Distillation for Text-to-3D GenerationarXiv 2024Towards High-Quality Specular Highlight Removal by Leveraging Large-Scale Synthetic DataICCV 2023 1Read-only Prompt Optimization for Vision-Language Few-shot LearningICCV 2023 1WavReward: Spoken Dialogue Models With Generalist Reward EvaluatorsarXiv 2025Activation Functions in Deep Learning: A Comprehensive Survey and BenchmarkarXiv 2021YAGO 4.5: A Large and Clean Knowledge Base with a Rich TaxonomyarXiv 2023GeCoNeRF: Few-shot Neural Radiance Fields via Geometric ConsistencyarXiv 2023IKEA Manuals at Work: 4D Grounding of Assembly Instructions on Internet VideosarXiv 2024OpenFMNav: Towards Open-Set Zero-Shot Object Navigation via Vision-Language Foundation ModelsarXiv 2024Attributed Question Answering: Evaluation and Modeling for Attributed Large Language ModelsarXiv 2022Multi-Vector Models with Textual Guidance for Fine-Grained Scientific Document SimilarityNAACL 2022 7ARIES: A Corpus of Scientific Paper Edits Made in Response to Peer ReviewsarXiv 2023Self-Supervised Any-Point Tracking by Contrastive Random WalksarXiv 2024EVREAL: Towards a Comprehensive Benchmark and Analysis Suite for Event-based Video ReconstructionarXiv 2023All Languages Matter: On the Multilingual Safety of Large Language ModelsarXiv 2023OmniVid: A Generative Framework for Universal Video UnderstandingCVPR 2024 1LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long VideosCVPR 2025 1InterFusion: Text-Driven Generation of 3D Human-Object InteractionarXiv 2024Multimodal Needle in a Haystack: Benchmarking Long-Context Capability of Multimodal Large Language ModelsarXiv 2024MatFuse: Controllable Material Generation with Diffusion ModelsCVPR 2024 1M2D2: A Massively Multi-domain Language Modeling DatasetarXiv 2022Grounding-IQA: Multimodal Language Grounding Model for Image Quality AssessmentarXiv 2024Fast Full-frame Video Stabilization with Iterative OptimizationICCV 2023 1Well-classified Examples are Underestimated in Classification with Deep Neural NetworksarXiv 2021Adaptive Parametric ActivationarXiv 2024A Synthetic Dataset for Personal Attribute InferencearXiv 2024SemiCD-VL: Visual-Language Model Guidance Makes Better Semi-supervised Change DetectorarXiv 2024Generative Dual Adversarial Network for Generalized Zero-shot Learninggenerative-dual-adversarial-network-for-1GuardT2I: Defending Text-to-Image Models from Adversarial PromptsarXiv 2024Diff9D: Diffusion-Based Domain-Generalized Category-Level 9-DoF Object Pose EstimationarXiv 2025Generative Models from the perspective of Continual Learninggenerative-models-from-the-perspective-of-1UnifiedGesture: A Unified Gesture Synthesis Model for Multiple SkeletonsarXiv 2023Learning to Retrieve Passages without SupervisionNAACL 2022 7When is Tree Search Useful for LLM Planning? It Depends on the DiscriminatorarXiv 2024Audio Retrieval with Natural Language Queries: A Benchmark StudyarXiv 2021Instruct and Extract: Instruction Tuning for On-Demand Information ExtractionarXiv 2023Effective Test Generation Using Pre-trained Large Language Models and Mutation TestingarXiv 2023Learning a Room with the Occ-SDF Hybrid: Signed Distance Function Mingled with Occupancy Aids Scene RepresentationICCV 2023 1Squeezed Attention: Accelerating Long Context Length LLM InferencearXiv 2024The Interplay of Variant, Size, and Task Type in Arabic Pre-trained Language ModelsEACL (WANLP) 2021 4Fine-tuning Global Model via Data-Free Knowledge Distillation for Non-IID Federated LearningCVPR 2022 1SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond WordsarXiv 2024VideoXum: Cross-modal Visual and Textural Summarization of VideosarXiv 2023MedSyn: Text-guided Anatomy-aware Synthesis of High-Fidelity 3D CT ImagesarXiv 2023Symbol as Points: Panoptic Symbol Spotting via Point-based RepresentationarXiv 2024Interweaved Graph and Attention Network for 3D Human Pose EstimationarXiv 2023Fast Encoder-Based 3D from Casual Videos via Point Track ProcessingarXiv 2024Distilled Decoding 1: One-step Sampling of Image Auto-regressive Models with Flow MatchingarXiv 2024MMInA: Benchmarking Multihop Multimodal Internet AgentsarXiv 2024DUET: Cross-modal Semantic Grounding for Contrastive Zero-shot LearningarXiv 2022ViewFusion: Towards Multi-View Consistency via Interpolated DenoisingCVPR 2024 1Black Box Adversarial Prompting for Foundation ModelsarXiv 2023Facial-Sketch Synthesis: A New ChallengearXiv 2021Mind with Eyes: from Language Reasoning to Multimodal ReasoningarXiv 2025Sat2Density: Faithful Density Learning from Satellite-Ground Image PairsICCV 2023 1PointNorm: Dual Normalization is All You Need for Point Cloud AnalysisarXiv 2022Learning without Forgetting for Vision-Language ModelsarXiv 2023ProtoCLIP: Prototypical Contrastive Language Image PretrainingarXiv 2022Benchmarking Representations for Speech, Music, and Acoustic EventsarXiv 2024OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMsarXiv 2024Seal-Tools: Self-Instruct Tool Learning Dataset for Agent Tuning and Detailed BenchmarkarXiv 2024Interactive Agents: Simulating Counselor-Client Psychological Counseling via Role-Playing LLM-to-LLM InteractionsarXiv 2024Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional UnderstandingCVPR 2024 1Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview imagesarXiv 2025CodeJudge: Evaluating Code Generation with Large Language ModelsarXiv 2024BPE-Dropout: Simple and Effective Subword Regularizationbpe-dropout-simple-and-effective-subwordActive Vision Might Be All You Need: Exploring Active Vision in Bimanual Robotic ManipulationarXiv 2024Generalized Few-shot 3D Point Cloud Segmentation with Vision-Language ModelCVPR 2025 1Beyond MOT: Semantic Multi-Object TrackingarXiv 2024Pedagogical Alignment of Large Language ModelsarXiv 2024TESS 2: A Large-Scale Generalist Diffusion Language ModelarXiv 2025Edge Representation Learning with HypergraphsNeurIPS 2021 12A Survey of Graph Neural Networks for Social Recommender SystemsarXiv 2022COMEDIAN: Self-Supervised Learning and Knowledge Distillation for Action Spotting using TransformersarXiv 2023NVFi: Neural Velocity Fields for 3D Physics Learning from Dynamic Videosnvfi-neural-velocity-fields-for-3d-physicsThink Before Recommend: Unleashing the Latent Reasoning Power for Sequential RecommendationarXiv 2025Disentangling Orthogonal Planes for Indoor Panoramic Room Layout Estimation with Cross-Scale Distortion AwarenessCVPR 2023 1StegoGAN: Leveraging Steganography for Non-Bijective Image-to-Image TranslationCVPR 2024 1Linearly Mapping from Image to Text SpacearXiv 2022Towards Exploiting Background Knowledge for Building Conversation Systemstowards-exploiting-background-knowledge-for-1Cross-view Masked Diffusion Transformers for Person Image SynthesisarXiv 2024ScaleDreamer: Scalable Text-to-3D Synthesis with Asynchronous Score DistillationarXiv 2024Measuring The Impact Of Programming Language DistributionarXiv 2023Unleashing the Power of Pre-trained Language Models for Offline Reinforcement LearningarXiv 2023Emotion Recognition from SpeecharXiv 2019Unveiling the Ignorance of MLLMs: Seeing Clearly, Answering IncorrectlyCVPR 2025 1Efficient Semantic Segmentation by Altering Resolutions for Compressed VideosCVPR 2023 1CrossSum: Beyond English-Centric Cross-Lingual Summarization for 1,500+ Language PairsarXiv 2021Extracting Prompts by Inverting LLM OutputsarXiv 2024In-Context Language Learning: Architectures and AlgorithmsarXiv 2024Mapping Global Floods with 10 Years of Satellite Radar DataarXiv 2024Are Neural Topic Models Broken?arXiv 2022On the Emergence of Thinking in LLMs I: Searching for the Right IntuitionarXiv 2025Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM InferencearXiv 2024PadChest: A large chest x-ray image dataset with multi-label annotated reportsarXiv 2019Flow: Modularized Agentic Workflow AutomationarXiv 2025Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail ModerationarXiv 2025ADELIE: Aligning Large Language Models on Information ExtractionarXiv 2024Robustness and Generalizability of Deepfake Detection: A Study with Diffusion ModelsarXiv 2023ResumeAtlas: Revisiting Resume Classification with Large-Scale Datasets and Large Language ModelsarXiv 2024Text-Video Retrieval with Disentangled Conceptualization and Set-to-Set AlignmentarXiv 2023Hybrid Spectral Denoising Transformer with Guided AttentionICCV 2023 1Improving Transformer-based Image Matching by Cascaded Capturing Spatially Informative KeypointsICCV 2023 1Scalable Transformer for PDE Surrogate Modelingscalable-transformer-for-pde-surrogateJointly Optimizing Query Encoder and Product Quantization to Improve Retrieval PerformancearXiv 2021SINet: Extreme Lightweight Portrait Segmentation Networks with Spatial Squeeze Modules and Information Blocking DecoderarXiv 2019Use Your INSTINCT: INSTruction optimization for LLMs usIng Neural bandits Coupled with TransformersarXiv 2023HelloBench: Evaluating Long Text Generation Capabilities of Large Language ModelsarXiv 2024From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language ModelsarXiv 2024CURLoRA: Stable LLM Continual Fine-Tuning and Catastrophic Forgetting MitigationarXiv 2024There and Back Again: Revisiting Backpropagation Saliency Methodsthere-and-back-again-revisiting-1Inherently Interpretable Time Series Classification via Multiple Instance LearningarXiv 2023One Thousand and One Pairs: A "novel" challenge for long-context language modelsarXiv 2024QLASS: Boosting Language Agent Inference via Q-Guided Stepwise SearcharXiv 2025eCeLLM: Generalizing Large Language Models for E-commerce from Large-scale, High-quality Instruction DataarXiv 2024vid-TLDR: Training Free Token merging for Light-weight Video TransformerCVPR 2024 1Learning Sequential Descriptors for Sequence-based Visual Place RecognitionarXiv 2022Generating Benchmarks for Factuality Evaluation of Language ModelsarXiv 2023ChildMandarin: A Comprehensive Mandarin Speech Dataset for Young Children Aged 3-5arXiv 2024Mono-ViFI: A Unified Learning Framework for Self-supervised Single- and Multi-frame Monocular Depth EstimationarXiv 2024Prompting4Debugging: Red-Teaming Text-to-Image Diffusion Models by Finding Problematic PromptsarXiv 2023TEG-DB: A Comprehensive Dataset and Benchmark of Textual-Edge GraphsarXiv 2024Cross-Modal and Uni-Modal Soft-Label Alignment for Image-Text RetrievalarXiv 2024Towards Real-world Event-guided Low-light Video Enhancement and DeblurringarXiv 2024Mitigate the Gap: Investigating Approaches for Improving Cross-Modal Alignment in CLIParXiv 2024Breaking the Ceiling of the LLM Community by Treating Token Generation as a Classification for EnsemblingarXiv 2024Large Language Models are Strong Audio-Visual Speech Recognition LearnersarXiv 2024CHAMPAGNE: Learning Real-world Conversation from Large-Scale Web VideosICCV 2023 1TextWorldExpress: Simulating Text Games at One Million Steps Per SecondarXiv 2022SQuADDS: A validated design database and simulation workflow for superconducting qubit designarXiv 2023Online-LoRA: Task-free Online Continual Learning via Low Rank AdaptationarXiv 2024PEARL: Prompting Large Language Models to Plan and Execute Actions Over Long DocumentsarXiv 2023Derm1M: A Million-scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for DermatologyICCV 2025

Back to Papers