0

All papers

Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.

Cascade-DETR: Delving into High-Quality Universal Object DetectionICCV 2023 1Can ChatGPT Assess Human Personalities? A General Evaluation FrameworkarXiv 2023InstructDial: Improving Zero and Few-shot Generalization in Dialogue through Instruction TuningarXiv 2022ERASER: A Benchmark to Evaluate Rationalized NLP Modelseraser-a-benchmark-to-evaluate-rationalized-1Natural Language Reinforcement LearningarXiv 2024GraphWiz: An Instruction-Following Language Model for Graph ProblemsarXiv 2024DDXPlus: A New Dataset For Automatic Medical DiagnosisarXiv 2022Link-Context Learning for Multimodal LLMsCVPR 2024 1SoundCTM: Unifying Score-based and Consistency Models for Full-band Text-to-Sound GenerationarXiv 2024Breathing New Life into 3D Assets with Generative RepaintingarXiv 2023EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the BackboneICCV 2023 1Spot the Fake: Large Multimodal Model-Based Synthetic Image Detection with Artifact ExplanationarXiv 2025Gated Multimodal Units for Information FusionarXiv 2017RoSteALS: Robust Steganography using Autoencoder Latent SpacearXiv 2023Filter-enhanced MLP is All You Need for Sequential RecommendationarXiv 2022Large Language Models for Software Engineering: A Systematic Literature ReviewarXiv 2023Improving Passage Retrieval with Zero-Shot Question GenerationarXiv 2022Rethinking pose estimation in crowds: overcoming the detection information-bottleneck and ambiguityarXiv 2023Diffusion Bridge Implicit ModelsarXiv 2024MatPlotAgent: Method and Evaluation for LLM-Based Agentic Scientific Data VisualizationarXiv 2024RealCustom++: Representing Images as Real-Word for Real-Time CustomizationarXiv 2024CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and GenerationarXiv 2023PatentSBERTa: A Deep NLP based Hybrid Model for Patent Distance and Classification using Augmented SBERTarXiv 2021AlphaNet: Improved Training of Supernets with Alpha-DivergencearXiv 2021A Critical Evaluation of AI Feedback for Aligning Large Language ModelsarXiv 2024AniGAN: Style-Guided Generative Adversarial Networks for Unsupervised Anime Face GenerationarXiv 2021MV-Map: Offboard HD-Map Generation with Multi-view ConsistencyICCV 2023 1Language models scale reliably with over-training and on downstream tasksarXiv 2024Morphable Diffusion: 3D-Consistent Diffusion for Single-image Avatar CreationCVPR 2024 1GridMM: Grid Memory Map for Vision-and-Language NavigationICCV 2023 1TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton OperatorsarXiv 2025Multiresolution Textual InversionarXiv 2022Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMsICCV 2025CoDi: Conditional Diffusion Distillation for Higher-Fidelity and Faster Image GenerationCVPR 2024 1SeqXGPT: Sentence-Level AI-Generated Text DetectionarXiv 2023A Comprehensive Survey of Continual Learning: Theory, Method and ApplicationarXiv 2023MLCopilot: Unleashing the Power of Large Language Models in Solving Machine Learning TasksarXiv 2023Generated Knowledge Prompting for Commonsense ReasoningACL 2022 5EfficientZero V2: Mastering Discrete and Continuous Control with Limited DataarXiv 2024Syntax-Aware Network for Handwritten Mathematical Expression RecognitionCVPR 2022 1ProtST: Multi-Modality Learning of Protein Sequences and Biomedical TextsarXiv 2023SF-V: Single Forward Video Generation ModelarXiv 2024Brush Your Text: Synthesize Any Scene Text on Images via Diffusion ModelarXiv 2023NeuFlow: Real-time, High-accuracy Optical Flow Estimation on Robots Using Edge DevicesarXiv 2024DePT: Decomposed Prompt Tuning for Parameter-Efficient Fine-tuningarXiv 2023Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference OptimizationarXiv 2023Transformer Feed-Forward Layers Are Key-Value MemoriesEMNLP 2021 11TroL: Traversal of Layers for Large Language and Vision ModelsarXiv 2024VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language ModellingarXiv 2024Keep It SimPool: Who Said Supervised Transformers Suffer from Attention Deficit?ICCV 2023 1Point2Vec for Self-Supervised Representation Learning on Point CloudsarXiv 2023One Model is All You Need: Multi-Task Learning Enables Simultaneous Histology Image Segmentation and ClassificationarXiv 2022GST: Precise 3D Human Body from a Single Image with Gaussian Splatting TransformersarXiv 2024Mask2Map: Vectorized HD Map Construction Using Bird's Eye View Segmentation MasksarXiv 2024PlankAssembly: Robust 3D Reconstruction from Three Orthographic Views with Learnt Shape ProgramsICCV 2023 1Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language ModelsarXiv 2024Harnessing Vision Foundation Models for High-Performance, Training-Free Open Vocabulary SegmentationICCV 2025Large-Scale Multi-Label Text Classification on EU Legislationlarge-scale-multi-label-text-classification-2Lighting up NeRF via Unsupervised Decomposition and EnhancementICCV 2023 1What Language Model Architecture and Pretraining Objective Work Best for Zero-Shot Generalization?arXiv 2022RoboSense: Large-scale Dataset and Benchmark for Egocentric Robot Perception and Navigation in Crowded and Unstructured EnvironmentsCVPR 2025 1SeqDiffuSeq: Text Diffusion with Encoder-Decoder TransformersarXiv 2022LSG Attention: Extrapolation of pretrained Transformers to long sequencesarXiv 2022MasRouter: Learning to Route LLMs for Multi-Agent SystemsarXiv 2025GRANDE: Gradient-Based Decision Tree Ensembles for Tabular DataarXiv 2023OMNI-DC: Highly Robust Depth Completion with Multiresolution Depth IntegrationarXiv 2024RefGPT: Dialogue Generation of GPT, by GPT, and for GPTarXiv 2023Structural Text Segmentation of Legal DocumentsarXiv 2020Collaborative Neural Rendering using Anime Character SheetsarXiv 2022Reducing Hallucinations in Vision-Language Models via Latent Space SteeringarXiv 2024Tracking Meets LoRA: Faster Training, Larger Model, Stronger PerformancearXiv 2024Deep Visual Odometry with Events and FramesarXiv 2023CAGroup3D: Class-Aware Grouping for 3D Object Detection on Point CloudsarXiv 2022RITA: a Study on Scaling Up Generative Protein Sequence ModelsarXiv 2022Rethinking Domain Generalization for Face Anti-spoofing: Separability and AlignmentCVPR 2023 1RAM: Retrieval-Based Affordance Transfer for Generalizable Zero-Shot Robotic ManipulationarXiv 2024Large Language Models to Enhance Bayesian OptimizationarXiv 2024SAN: Inducing Metrizability of GAN with Discriminative Normalized Linear LayerarXiv 2023Segment3D: Learning Fine-Grained Class-Agnostic 3D Segmentation without Manual LabelsarXiv 2023Xiezhi: An Ever-Updating Benchmark for Holistic Domain Knowledge EvaluationarXiv 2023CheXpert Plus: Augmenting a Large Chest X-ray Dataset with Text Radiology Reports, Patient Demographics and Additional Image FormatsarXiv 2024Rejuvenating image-GPT as Strong Visual Representation LearnersarXiv 2023Exploring the Benefits of Training Expert Language Models over Instruction TuningarXiv 2023Compression, Transduction, and Creation: A Unified Framework for Evaluating Natural Language GenerationEMNLP 2021 11Improving Post Training Neural Quantization: Layer-wise Calibration and Integer ProgrammingarXiv 2020Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene UnderstandingarXiv 2024Tabular Benchmarks for Joint Architecture and Hyperparameter OptimizationarXiv 2019SciFive: a text-to-text transformer model for biomedical literaturearXiv 2021Benchmarking Complex Instruction-Following with Multiple Constraints CompositionarXiv 2024ILDC for CJPE: Indian Legal Documents Corpus for Court Judgment Prediction and ExplanationACL 2021 5GlossBERT: BERT for Word Sense Disambiguation with Gloss Knowledgeglossbert-bert-for-word-sense-disambiguation-1DFR: Deep Feature Reconstruction for Unsupervised Anomaly SegmentationarXiv 2020Realistic Synthetic Financial Transactions for Anti-Money Laundering Modelsrealistic-synthetic-financial-transactionsAdaMerging: Adaptive Model Merging for Multi-Task LearningarXiv 2023Multimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear DistillationarXiv 2025Domain Adaptation via Prompt LearningarXiv 2022Dialog Inpainting: Turning Documents into DialogsarXiv 2022FreeV: Free Lunch For Vocoders Through Pseudo Inversed Mel FilterarXiv 2024Cross-video Identity Correlating for Person Re-identification Pre-trainingarXiv 2024Secure Transformer Inference ProtocolarXiv 2023DreamCatalyst: Fast and High-Quality 3D Editing via Controlling Editability and Identity PreservationarXiv 2024Efficient Heterogeneous Graph Learning via Random ProjectionarXiv 2023Datamodels: Predicting Predictions from Training DataarXiv 2022Web2Code: A Large-scale Webpage-to-Code Dataset and Evaluation Framework for Multimodal LLMsarXiv 2024TUDataset: A collection of benchmark datasets for learning with graphsarXiv 2020Face2Diffusion for Fast and Editable Face PersonalizationCVPR 2024 1Learning Performance-Improving Code EditsarXiv 2023Incorporating External Knowledge through Pre-training for Natural Language to Code Generationincorporating-external-knowledge-through-pre-1Full-Atom Peptide Design based on Multi-modal Flow MatchingarXiv 2024UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual EncodingarXiv 2025Free-Bloom: Zero-Shot Text-to-Video Generator with LLM Director and LDM AnimatorNeurIPS 2023 11A Survey on LLM Inference-Time Self-ImprovementarXiv 2024SUM: Saliency Unification through Mamba for Visual Attention ModelingarXiv 2024Less is More: Reducing Task and Model Complexity for 3D Point Cloud Semantic SegmentationCVPR 2023 1BabelCalib: A Universal Approach to Calibrating Central CamerasICCV 2021 10MineDreamer: Learning to Follow Instructions via Chain-of-Imagination for Simulated-World ControlarXiv 2024Self-Supervised Visual Representation Learning with Semantic GroupingarXiv 2022Alibaba LingmaAgent: Improving Automated Issue Resolution via Comprehensive Repository ExplorationarXiv 2024Self-supervised Learning on Graphs: Deep Insights and New DirectionarXiv 2020Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning from Human FeedbackarXiv 2023EHRSQL: A Practical Text-to-SQL Benchmark for Electronic Health Recordsehrsql-a-practical-text-to-sql-benchmark-forDraft, Sketch, and Prove: Guiding Formal Theorem Provers with Informal ProofsarXiv 2022Objaverse++: Curated 3D Object Dataset with Quality AnnotationsarXiv 2025Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM modelsarXiv 2024Follow the Rules: Reasoning for Video Anomaly Detection with Large Language ModelsarXiv 2024Polyjuice: Generating Counterfactuals for Explaining, Evaluating, and Improving ModelsACL 2021 5Multilingual Jailbreak Challenges in Large Language ModelsarXiv 2023CycleNet: Rethinking Cycle Consistency in Text-Guided Diffusion for Image Manipulationcyclenet-rethinking-cycle-consistency-in-textMusic-Driven Group ChoreographyCVPR 2023 1TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance ControlarXiv 2024Automatic Speech Recognition Datasets in Cantonese: A Survey and New DatasetLREC 2022 6CLUTRR: A Diagnostic Benchmark for Inductive Reasoning from Textclutrr-a-diagnostic-benchmark-for-inductive-1Interpreting and Editing Vision-Language Representations to Mitigate HallucinationsarXiv 2024MagicStick: Controllable Video Editing via Control Handle TransformationsarXiv 2023PA-SAM: Prompt Adapter SAM for High-Quality Image SegmentationarXiv 2024An Embarrassingly Simple Approach for Transfer Learning from Pretrained Language Modelsan-embarrassingly-simple-approach-for-1Literature Meets Data: A Synergistic Approach to Hypothesis GenerationarXiv 2024Chupa: Carving 3D Clothed Humans from Skinned Shape Priors using 2D Diffusion Probabilistic ModelsICCV 2023 1SAILER: Structure-aware Pre-trained Language Model for Legal Case RetrievalarXiv 2023GenDoP: Auto-regressive Camera Trajectory Generation as a Director of PhotographyICCV 2025PATIENT-Ψ: Using Large Language Models to Simulate Patients for Training Mental Health ProfessionalsarXiv 2024Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters MorearXiv 2025NaSGEC: a Multi-Domain Chinese Grammatical Error Correction Dataset from Native Speaker TextsarXiv 2023DA-Code: Agent Data Science Code Generation Benchmark for Large Language ModelsarXiv 2024MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction TuningarXiv 2023OBoW: Online Bag-of-Visual-Words Generation for Self-Supervised Learningobow-online-bag-of-visual-words-generationMegaFusion: Extend Diffusion Models towards Higher-resolution Image Generation without Further TuningarXiv 2024Answering Questions by Meta-Reasoning over Multiple Chains of ThoughtarXiv 2023A Systematic Review of Deep Learning-based Research on Radiology Report GenerationarXiv 2023Label, Verify, Correct: A Simple Few Shot Object Detection MethodCVPR 2022 1Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMsCVPR 2024 1Normalizing Flows for Human Pose Anomaly DetectionICCV 2023 1Visual Instruction Inversion: Image Editing via Visual PromptingarXiv 2023Language Representations Can be What Recommenders Need: Findings and PotentialsarXiv 2024VIGC: Visual Instruction Generation and CorrectionarXiv 2023COGMEN: COntextualized GNN based Multimodal Emotion recognitioNNAACL 2022 7VisText: A Benchmark for Semantically Rich Chart CaptioningarXiv 2023One Loss for All: Deep Hashing with a Single Cosine Similarity based Learning ObjectiveNeurIPS 2021 12Fast Diffusion ModelarXiv 2023Towards Open-Vocabulary Video Instance SegmentationICCV 2023 1Decomposed Prompting: A Modular Approach for Solving Complex TasksarXiv 2022Efficiently Learning at Test-Time: Active Fine-Tuning of LLMsarXiv 2024IPRE: a Dataset for Inter-Personal Relationship ExtractionarXiv 2019DyLoRA: Parameter Efficient Tuning of Pre-trained Models using Dynamic Search-Free Low-Rank AdaptationarXiv 2022DIVOTrack: A Novel Dataset and Baseline Method for Cross-View Multi-Object Tracking in DIVerse Open ScenesarXiv 2023Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the KeyCVPR 2025 1GLUCOSE: GeneraLized and COntextualized Story ExplanationsEMNLP 2020 11FALCON: Honest-Majority Maliciously Secure Framework for Private Deep LearningarXiv 2020Recitation-Augmented Language ModelsarXiv 2022StableLLaVA: Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue DataarXiv 2023WAVES: Benchmarking the Robustness of Image WatermarksarXiv 2024SegFace: Face Segmentation of Long-Tail ClassesarXiv 2024Making Sense of Vision and Touch: Self-Supervised Learning of Multimodal Representations for Contact-Rich TasksarXiv 2018Speculative Decoding with Big Little Decoderspeculative-decoding-with-big-little-decoderMemLong: Memory-Augmented Retrieval for Long Text ModelingarXiv 2024Joint Prompt Optimization of Stacked LLMs using Variational Inferencejoint-prompt-optimization-of-stacked-llmsSelective Prompt Anchoring for Code GenerationarXiv 2024LangBridge: Multilingual Reasoning Without Multilingual SupervisionarXiv 2024Fast-HuBERT: An Efficient Training Framework for Self-Supervised Speech Representation LearningarXiv 2023Getting the Ball Rolling: Learning a Dexterous Policy for a Biomimetic Tendon-Driven Hand with Rolling Contact JointsarXiv 2023KARMA: Leveraging Multi-Agent LLMs for Automated Knowledge Graph EnrichmentarXiv 2025StyleDubber: Towards Multi-Scale Style Learning for Movie DubbingarXiv 2024Conversations Are Not Flat: Modeling the Dynamic Information Flow across Dialogue UtterancesACL 2021 5AlpaCare:Instruction-tuned Large Language Models for Medical ApplicationarXiv 2023NILUT: Conditional Neural Implicit 3D Lookup Tables for Image EnhancementarXiv 2023Modeling the Label Distributions for Weakly-Supervised Semantic SegmentationarXiv 2024Lexically Constrained Decoding for Sequence Generation Using Grid Beam Searchlexically-constrained-decoding-for-sequence-1CEM: Commonsense-aware Empathetic Response GenerationarXiv 2021OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented GenerationICCV 2025From Code to Correctness: Closing the Last Mile of Code Generation with Hierarchical DebuggingarXiv 2024Why are Visually-Grounded Language Models Bad at Image Classification?arXiv 2024K-Core based Temporal Graph Convolutional Network for Dynamic GraphsarXiv 2020Language Models can Self-Lengthen to Generate Long TextsarXiv 2024ASAM: Boosting Segment Anything Model with Adversarial TuningCVPR 2024 1TUNA: Taming Unified Visual Representations for Native Unified Multimodal ModelsarXiv 2025Enriched CNN-Transformer Feature Aggregation Networks for Super-ResolutionarXiv 2022Large-scale Bilingual Language-Image Contrastive LearningarXiv 2022ConvLoRA and AdaBN based Domain Adaptation via Self-TrainingarXiv 2024MetaBEV: Solving Sensor Failures for BEV Detection and Map SegmentationarXiv 2023UltraMedical: Building Specialized Generalists in BiomedicinearXiv 2024

Back to Papers