All papers
Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.
Image Captioning: Transforming Objects into Wordsimage-captioning-transforming-objects-into-1Adaptive Super Resolution For One-Shot Talking-Head GenerationarXiv 2024Neural Graph Reasoning: Complex Logical Query Answering Meets Graph DatabasesarXiv 2023Progressive Knowledge Distillation Of Stable Diffusion XL Using Layer Level LossarXiv 2024LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal ModelsarXiv 2024Explicit Correspondence Matching for Generalizable Neural Radiance FieldsarXiv 2023V2X-Seq: A Large-Scale Sequential Dataset for Vehicle-Infrastructure Cooperative Perception and ForecastingCVPR 2023 1Sherpa3D: Boosting High-Fidelity Text-to-3D Generation via Coarse 3D PriorCVPR 2024 1VSSD: Vision Mamba with Non-Causal State Space DualityICCV 2025The KiTS21 Challenge: Automatic segmentation of kidneys, renal tumors, and renal cysts in corticomedullary-phase CTarXiv 2023Anchor3DLane: Learning to Regress 3D Anchors for Monocular 3D Lane DetectionCVPR 2023 1CORD-19: The COVID-19 Open Research DatasetACL 2020 7Music Mixing Style Transfer: A Contrastive Learning Approach to Disentangle Audio EffectsarXiv 2022TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question AnsweringICCV 2023 1EleGANt: Exquisite and Locally Editable GAN for Makeup TransferarXiv 2022Heron-Bench: A Benchmark for Evaluating Vision Language Models in JapanesearXiv 2024An Empirical Study of Example Forgetting during Deep Neural Network Learningan-empirical-study-of-example-forgetting-1Infinite Mobility: Scalable High-Fidelity Synthesis of Articulated Objects via Procedural GenerationarXiv 2025Relighting Neural Radiance Fields with Shadow and Highlight HintsarXiv 2023ARNOLD: A Benchmark for Language-Grounded Task Learning With Continuous States in Realistic 3D ScenesICCV 2023 1GrammarGPT: Exploring Open-Source LLMs for Native Chinese Grammatical Error Correction with Supervised Fine-TuningarXiv 2023Goat: Fine-tuned LLaMA Outperforms GPT-4 on Arithmetic TasksarXiv 2023Bottom-Up Abstractive Summarizationbottom-up-abstractive-summarization-1Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?arXiv 2022DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMsarXiv 2024Hash3D: Training-free Acceleration for 3D GenerationCVPR 2025 1Large Language Models for Automated Data Science: Introducing CAAFE for Context-Aware Automated Feature Engineeringlarge-language-models-for-automated-dataExact Feature Distribution Matching for Arbitrary Style Transfer and Domain GeneralizationCVPR 2022 1TTS-Portuguese Corpus: a corpus for speech synthesis in Brazilian PortuguesearXiv 2020Contrastive Preference Learning: Learning from Human Feedback without RLarXiv 2023HyperSeg: Towards Universal Visual Segmentation with Large Language ModelarXiv 2024Joint-task Self-supervised Learning for Temporal Correspondencejoint-task-self-supervised-learning-for-1KAN or MLP: A Fairer ComparisonarXiv 2024SpeechCraft: A Fine-grained Expressive Speech Dataset with Natural Language DescriptionarXiv 2024F-coref: Fast, Accurate and Easy to Use Coreference ResolutionarXiv 2022General Image-to-Image Translation with One-Shot Image GuidanceICCV 2023 1Multi-Granularity Cross-modal Alignment for Generalized Medical Visual Representation LearningarXiv 2022When Do Neural Nets Outperform Boosted Trees on Tabular Data?when-do-neural-nets-outperform-boosted-treesAVID: Any-Length Video Inpainting with Diffusion ModelCVPR 2024 1PointOdyssey: A Large-Scale Synthetic Dataset for Long-Term Point TrackingICCV 2023 1Aligning Pretraining for Detection via Object-Level Contrastive LearningNeurIPS 2021 12Recursive Generalization Transformer for Image Super-ResolutionarXiv 2023Learned Initializations for Optimizing Coordinate-Based Neural RepresentationsCVPR 2021 1ODIN: A Single Model for 2D and 3D SegmentationCVPR 2024 1Quilt-1M: One Million Image-Text Pairs for Histopathologyquilt-1m-one-million-image-text-pairs-forLearning De-biased Representations with Biased RepresentationsICML 2020 1Smooth ECE: Principled Reliability Diagrams via Kernel SmoothingarXiv 2023ImagenHub: Standardizing the evaluation of conditional image generation modelsarXiv 2023Self-Supervised Generalisation with Meta Auxiliary Learningself-supervised-generalisation-with-meta-1OpenStreetView-5M: The Many Roads to Global Visual GeolocationCVPR 2024 1FlexSP: Accelerating Large Language Model Training via Flexible Sequence ParallelismarXiv 20244D-Rotor Gaussian Splatting: Towards Efficient Novel View Synthesis for Dynamic ScenesarXiv 2024Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual FeedbackarXiv 2025Dataset for Automatic Summarization of Russian NewsarXiv 2020Learning to Upsample by Learning to SampleICCV 2023 1Visual Classification via Description from Large Language ModelsarXiv 2022Uncovering ChatGPT's Capabilities in Recommender SystemsarXiv 2023Efficient Passage Retrieval with Hashing for Open-domain Question AnsweringACL 2021 5Inference-Time Scaling for Complex Tasks: Where We Stand and What Lies AheadarXiv 2025Multiple Heads are Better than One: Few-shot Font Generation with Multiple Localized ExpertsICCV 2021 10Deep Neural Networks are Easily Fooled: High Confidence Predictions for Unrecognizable Imagesdeep-neural-networks-are-easily-fooled-high-1A Scalable Communication Protocol for Networks of Large Language ModelsarXiv 2024DreamMatcher: Appearance Matching Self-Attention for Semantically-Consistent Text-to-Image PersonalizationCVPR 2024 1Unified Concept Editing in Diffusion ModelsarXiv 2023Towards Building an Intelligent Anti-Malware System: A Deep Learning Approach using Support Vector Machine (SVM) for Malware ClassificationarXiv 2017A Modern Self-Referential Weight Matrix That Learns to Modify ItselfarXiv 2022COLD-Attack: Jailbreaking LLMs with Stealthiness and ControllabilityarXiv 2024Ghostbuster: Detecting Text Ghostwritten by Large Language ModelsarXiv 2023SeaLLMs -- Large Language Models for Southeast AsiaarXiv 2023A Neural Network Architecture Combining Gated Recurrent Unit (GRU) and Support Vector Machine (SVM) for Intrusion Detection in Network Traffic DataarXiv 2017Few-Shot Segmentation Without Meta-Learning: A Good Transductive Inference Is All You Need?CVPR 2021 1A Dataset of German Legal Documents for Named Entity Recognitiona-dataset-of-german-legal-documents-for-named-1ParaFold: Paralleling AlphaFold for Large-Scale PredictionsarXiv 2021Mixed Neural Voxels for Fast Multi-view Video SynthesisICCV 2023 1Cross-Attention of Disentangled Modalities for 3D Human Mesh Recovery with TransformersarXiv 2022Composite Motion Learning with Task ControlarXiv 2023Rethinking Interpretability in the Era of Large Language ModelsarXiv 2024IterMVS: Iterative Probability Estimation for Efficient Multi-View StereoCVPR 2022 1Learning multiple visual domains with residual adapterslearning-multiple-visual-domains-with-1Molecule3D: A Benchmark for Predicting 3D Geometries from Molecular GraphsarXiv 2021Arbitrary-Scale Video Super-Resolution with Structural and Textural PriorsarXiv 2024VisDA: The Visual Domain Adaptation ChallengearXiv 2017MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory GuidanceICCV 2025Continual learning with hypernetworksICLR 2020 1Law of Vision Representation in MLLMsarXiv 2024Internal Consistency and Self-Feedback in Large Language Models: A SurveyarXiv 2024StaQC: A Systematically Mined Question-Code Dataset from Stack OverflowarXiv 2018EAGLES: Efficient Accelerated 3D Gaussians with Lightweight EncodingSarXiv 2023Widely Applicable Strong Baseline for Sports Ball Detection and TrackingarXiv 2023BEHAVE: Dataset and Method for Tracking Human Object InteractionsCVPR 2022 1Multimodal Deep LearningarXiv 2023Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic ReasoningACL 2021 5GPG: A Simple and Strong Reinforcement Learning Baseline for Model ReasoningarXiv 2025DiverseVul: A New Vulnerable Source Code Dataset for Deep Learning Based Vulnerability DetectionarXiv 2023Image Textualization: An Automatic Framework for Creating Accurate and Detailed Image DescriptionsarXiv 2024InsetGAN for Full-Body Image GenerationCVPR 2022 1SwinLSTM:Improving Spatiotemporal Prediction Accuracy using Swin Transformer and LSTMarXiv 2023AutoAD: Movie Description in ContextCVPR 2023 1Unlearnable Examples: Making Personal Data Unexploitableunlearnable-examples-making-personal-dataTimberTrek: Exploring and Curating Sparse Decision Trees with Interactive VisualizationarXiv 2022M2-omni: Advancing Omni-MLLM for Comprehensive Modality Support with Competitive PerformancearXiv 2025SALMON: Self-Alignment with Instructable Reward ModelsarXiv 2023Attention is Not All You Need: Pure Attention Loses Rank Doubly Exponentially with DeptharXiv 2021The Tracking Machine Learning challenge : Throughput phasearXiv 2021VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement LearningarXiv 2025Kolmogorov Arnold Informed neural network: A physics-informed deep learning framework for solving forward and inverse problems based on Kolmogorov Arnold NetworksarXiv 2024FreeCustom: Tuning-Free Customized Image Generation for Multi-Concept CompositionCVPR 2024 1RISE: Randomized Input Sampling for Explanation of Black-box ModelsarXiv 2018Parsing is All You Need for Accurate Gait Recognition in the WildarXiv 2023VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion ModelsarXiv 2024DER: Dynamically Expandable Representation for Class Incremental LearningCVPR 2021 1CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledgecommonsenseqa-a-question-answering-challenge-1PrefixQuant: Eliminating Outliers by Prefixed Tokens for Large Language Models QuantizationarXiv 2024Label Words are Anchors: An Information Flow Perspective for Understanding In-Context LearningarXiv 2023Language Modeling Is CompressionarXiv 2023Understanding Deep Image Representations by Inverting Themunderstanding-deep-image-representations-by-1FP8 Quantization: The Power of the ExponentarXiv 2022MoVA: Adapting Mixture of Vision Experts to Multimodal ContextarXiv 2024CCMB: A Large-scale Chinese Cross-modal BenchmarkarXiv 2022Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse AutoencodersarXiv 2024FinDKG: Dynamic Knowledge Graphs with Large Language Models for Detecting Global Trends in Financial MarketsarXiv 2024Learning Longer Memory in Recurrent Neural NetworksarXiv 2014Free Process Rewards without Process LabelsarXiv 2024A Generalization of ViT/MLP-Mixer to GraphsarXiv 2022Rethinking Performance Gains in Image Dehazing NetworksarXiv 2022C3: Zero-shot Text-to-SQL with ChatGPTarXiv 2023A Closer Look at Invalid Action Masking in Policy Gradient AlgorithmsarXiv 2020CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept MatchingarXiv 2024Resolving Multi-Condition Confusion for Finetuning-Free Personalized Image GenerationarXiv 2024Improving Diffusion Inverse Problem Solving with Decoupled Noise AnnealingCVPR 2025 1Dense X Retrieval: What Retrieval Granularity Should We Use?arXiv 2023DeepMapping2: Self-Supervised Large-Scale LiDAR Map OptimizationCVPR 2023 1InstructABSA: Instruction Learning for Aspect Based Sentiment AnalysisarXiv 2023KG-BART: Knowledge Graph-Augmented BART for Generative Commonsense ReasoningarXiv 2020BoQ: A Place is Worth a Bag of Learnable QueriesCVPR 2024 1Dynamic Parametric Retrieval Augmented Generation for Test-time Knowledge EnhancementarXiv 2025OpenWebMath: An Open Dataset of High-Quality Mathematical Web TextarXiv 2023MedVAE: Efficient Automated Interpretation of Medical Images with Large-Scale Generalizable AutoencodersarXiv 2025PromptDet: Towards Open-vocabulary Detection using Uncurated ImagesarXiv 2022SignAvatars: A Large-scale 3D Sign Language Holistic Motion Dataset and BenchmarkarXiv 2023Generation and Comprehension of Unambiguous Object Descriptionsgeneration-and-comprehension-of-unambiguous-1POCOVID-Net: Automatic Detection of COVID-19 From a New Lung Ultrasound Imaging Dataset (POCUS)arXiv 2020Multimodal Diffusion Transformer: Learning Versatile Behavior from
Multimodal GoalsarXiv 2024Towards Collaborative Autonomous Driving: Simulation Platform and End-to-End SystemarXiv 2024ParsiNLU: A Suite of Language Understanding Challenges for PersianarXiv 2020CricaVPR: Cross-image Correlation-aware Representation Learning for Visual Place RecognitionarXiv 2024Safety Alignment Should Be Made More Than Just a Few Tokens DeeparXiv 2024Prompt-aligned Gradient for Prompt TuningICCV 2023 1Boosting Object Detection with Zero-Shot Day-Night Domain AdaptationCVPR 2024 1Free3D: Consistent Novel View Synthesis without 3D RepresentationCVPR 2024 1Integrating NVIDIA Deep Learning Accelerator (NVDLA) with RISC-V SoC on FireSimarXiv 2019Seal-3D: Interactive Pixel-Level Editing for Neural Radiance FieldsICCV 2023 1Moisesdb: A dataset for source separation beyond 4-stemsarXiv 2023Social NCE: Contrastive Learning of Socially-aware Motion RepresentationsICCV 2021 10SVIT: Scaling up Visual Instruction TuningarXiv 2023Task-aware Retrieval with InstructionsarXiv 2022Self-Supervised Aggregation of Diverse Experts for Test-Agnostic Long-Tailed RecognitionarXiv 2021Faithful Chain-of-Thought ReasoningarXiv 2023Learning Type-Aware Embeddings for Fashion CompatibilityarXiv 2018Position-Aware Tagging for Aspect Sentiment Triplet ExtractionEMNLP 2020 11Logical Natural Language Generation from Open-Domain Tableslogical-natural-language-generation-from-open-1DeepInception: Hypnotize Large Language Model to Be JailbreakerarXiv 2023StableV2V: Stablizing Shape Consistency in Video-to-Video EditingarXiv 2024Taiyi: A Bilingual Fine-Tuned Large Language Model for Diverse Biomedical TasksarXiv 2023What Makes Convolutional Models Great on Long Sequence Modeling?arXiv 2022ResFields: Residual Neural Fields for Spatiotemporal Signalsresfields-residual-neural-fields-forLearning to Generate Explainable Stock Predictions using Self-Reflective Large Language ModelsarXiv 2024Efficient Two-Stage Detection of Human-Object Interactions with a Novel Unary-Pairwise TransformerCVPR 2022 1Diffusion Models without Classifier-free Guidancediffusion-models-without-classifier-freeThe Song Describer Dataset: a Corpus of Audio Captions for Music-and-Language EvaluationarXiv 2023Open-Domain Question Answering Goes Conversational via Question RewritingNAACL 2021 4Generative Photography: Scene-Consistent Camera Control for Realistic Text-to-Image SynthesisCVPR 2025 1Spectrally Pruned Gaussian Fields with Neural CompensationarXiv 2024Evaluating Graph Vulnerability and Robustness using TIGERarXiv 2020Libra: Building Decoupled Vision System on Large Language ModelsarXiv 2024TULIP: Towards Unified Language-Image PretrainingarXiv 2025Compressing Neural Networks: Towards Determining the Optimal Layer-wise DecompositionNeurIPS 2021 12DynaSent: A Dynamic Benchmark for Sentiment AnalysisACL 2021 5Foreground-Aware Relation Network for Geospatial Object Segmentation in High Spatial Resolution Remote Sensing Imageryforeground-aware-relation-network-forBeyond LLaVA-HD: Diving into High-Resolution Large Multimodal ModelsarXiv 2024Internet Explorer: Targeted Representation Learning on the Open WebarXiv 2023ReST: A Reconfigurable Spatial-Temporal Graph Model for Multi-Camera Multi-Object TrackingICCV 2023 1Neural HMMs are all you need (for high-quality attention-free TTS)arXiv 2021GuardReasoner: Towards Reasoning-based LLM SafeguardsarXiv 2025Deeper Text Understanding for IR with Contextual Neural Language ModelingarXiv 2019TREAD: Token Routing for Efficient Architecture-agnostic Diffusion TrainingICCV 2025DS-Fusion: Artistic Typography via Discriminated and Stylized DiffusionICCV 2023 1NeAT: Learning Neural Implicit Surfaces with Arbitrary Topologies from Multi-view ImagesCVPR 2023 1Lodge: A Coarse to Fine Diffusion Network for Long Dance Generation Guided by the Characteristic Dance PrimitivesCVPR 2024 1CorrMatch: Label Propagation via Correlation Matching for Semi-Supervised Semantic SegmentationCVPR 2024 1VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion ModelsarXiv 2024Learning Transformer Programslearning-transformer-programsDiscrete Contrastive Diffusion for Cross-Modal Music and Image GenerationarXiv 2022RealMAN: A Real-Recorded and Annotated Microphone Array Dataset for
Dynamic Speech Enhancement and LocalizationarXiv 2024Shifted Diffusion for Text-to-image GenerationCVPR 2023 1VIOLIN: A Large-Scale Dataset for Video-and-Language Inferenceviolin-a-large-scale-dataset-for-video-and-1Frequency-Adaptive Dilated Convolution for Semantic SegmentationCVPR 2024 1Data-Efficient Reinforcement Learning with Self-Predictive Representationsdata-efficient-reinforcement-learning-with-2Image Super-resolution Via Latent Diffusion: A Sampling-space Mixture Of Experts And Frequency-augmented Decoder ApproacharXiv 2023Block Transformer: Global-to-Local Language Modeling for Fast InferencearXiv 2024