All papers
Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.
EdNet: A Large-Scale Hierarchical Dataset in EducationarXiv 2019LP-MusicCaps: LLM-Based Pseudo Music CaptioningarXiv 2023Intent Prediction-Driven Model Predictive Control for UAV Planning and Navigation in Dynamic EnvironmentsarXiv 2024Flow-Guided Transformer for Video InpaintingarXiv 2022VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language ModelarXiv 2025VulDeePecker: A Deep Learning-Based System for Vulnerability DetectionarXiv 2018How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMsarXiv 2024GPAvatar: Generalizable and Precise Head Avatar from Image(s)arXiv 20243D-Adapter: Geometry-Consistent Multi-View Diffusion for High-Quality 3D GenerationarXiv 2024RetinaMask: Learning to predict masks improves state-of-the-art single-shot detection for freearXiv 2019MedMentions: A Large Biomedical Corpus Annotated with UMLS ConceptsAKBC 2019GroundingGPT:Language Enhanced Multi-modal Grounding ModelarXiv 2024A Toolkit for Generating Code Knowledge GraphsarXiv 2020GenSim: Generating Robotic Simulation Tasks via Large Language ModelsarXiv 2023Gradient Surgery for Multi-Task LearningNeurIPS 2020 12InfoBatch: Lossless Training Speed Up by Unbiased Dynamic Data PruningarXiv 2023A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmarka-large-scale-study-of-representationMaskGWM: A Generalizable Driving World Model with Video Mask ReconstructionCVPR 2025 1Survey on Factuality in Large Language Models: Knowledge, Retrieval and Domain-SpecificityarXiv 2023PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent TasksarXiv 2024The Surprising Effectiveness of Test-Time Training for Few-Shot LearningarXiv 2024DiffusionDepth: Diffusion Denoising Approach for Monocular Depth EstimationarXiv 2023SE(3)-DiffusionFields: Learning smooth cost functions for joint grasp and motion optimization through diffusionarXiv 2022Speech Denoising in the Waveform Domain with Self-AttentionarXiv 2022How is ChatGPT's behavior changing over time?arXiv 2023Probing the 3D Awareness of Visual Foundation ModelsCVPR 2024 1DiffusionBERT: Improving Generative Masked Language Models with Diffusion ModelsarXiv 2022Knowing When to Look: Adaptive Attention via A Visual Sentinel for Image Captioningknowing-when-to-look-adaptive-attention-via-a-1The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language VariantsarXiv 2023GRiT: A Generative Region-to-text Transformer for Object UnderstandingarXiv 2022GPL: Generative Pseudo Labeling for Unsupervised Domain Adaptation of Dense RetrievalNAACL 2022 7An Empirical Study on Cross-X Transfer for Legal Judgment PredictionarXiv 2022AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding AgentsarXiv 2024Where Visual Speech Meets Language: VSP-LLM Framework for Efficient and Context-Aware Visual Speech ProcessingarXiv 2024Tree-Ring Watermarks: Fingerprints for Diffusion Images that are Invisible and RobustarXiv 2023BRIO: Bringing Order to Abstractive SummarizationACL 2022 5HAT: Hardware-Aware Transformers for Efficient Natural Language Processinghat-hardware-aware-transformers-for-efficient-1Model compression via distillation and quantizationmodel-compression-via-distillation-and-1StoryDALL-E: Adapting Pretrained Text-to-Image Transformers for Story ContinuationarXiv 2022GRAND: Graph Neural DiffusionNeurIPS Workshop DLDE 2021 12Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image SynthesisarXiv 2024Surface Representation for Point CloudsarXiv 2022COCO-O: A Benchmark for Object Detectors under Natural Distribution Shiftscoco-o-a-benchmark-for-object-detectors-underData Management For Training Large Language Models: A SurveyarXiv 2023Chain of Draft: Thinking Faster by Writing LessarXiv 2025LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive MemoryarXiv 2024Contrastive Learning of Musical RepresentationsarXiv 2021Evaluation for Weakly Supervised Object Localization: Protocol, Metrics, and DatasetsarXiv 2020Bridging Evolutionary Algorithms and Reinforcement Learning: A Comprehensive Survey on Hybrid AlgorithmsarXiv 2024ViT-CoMer: Vision Transformer with Convolutional Multi-scale Feature Interaction for Dense PredictionsCVPR 2024 1HoloGAN: Unsupervised learning of 3D representations from natural imageshologan-unsupervised-learning-of-3d-14D-fy: Text-to-4D Generation Using Hybrid Score Distillation SamplingCVPR 2024 1Semantic Models for the First-stage Retrieval: A Comprehensive ReviewarXiv 2021Neural Network Verification with Branch-and-Bound for General NonlinearitiesarXiv 2024Scale-Equalizing Pyramid Convolution for Object Detectionscale-equalizing-pyramid-convolution-for-1GCC: Graph Contrastive Coding for Graph Neural Network Pre-TrainingarXiv 2020Fast Certified Robust Training with Short WarmupNeurIPS 2021 12GoalFlow: Goal-Driven Flow Matching for Multimodal Trajectories Generation in End-to-End Autonomous DrivingCVPR 2025 1Efficiently Computing Local Lipschitz Constants of Neural Networks via Bound PropagationarXiv 2022FNSPID: A Comprehensive Financial News Dataset in Time SeriesarXiv 2024Apollo: Band-sequence Modeling for High-Quality Audio RestorationarXiv 2024EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion TransformerarXiv 2024Black-Box Prompt Optimization: Aligning Large Language Models without Model TrainingarXiv 2023AiOS: All-in-One-Stage Expressive Human Pose and Shape EstimationCVPR 2024 1Atom: Low-bit Quantization for Efficient and Accurate LLM ServingarXiv 2023TokenHMR: Advancing Human Mesh Recovery with a Tokenized Pose RepresentationCVPR 2024 1Style Your Hair: Latent Optimization for Pose-Invariant Hairstyle Transfer via Local-Style-Aware Hair AlignmentarXiv 2022FAST-VQA: Efficient End-to-end Video Quality Assessment with Fragment SamplingarXiv 2022MoAI: Mixture of All Intelligence for Large Language and Vision ModelsarXiv 2024Learning to Estimate Hidden Motions with Global Motion AggregationICCV 2021 10DDPM-CD: Denoising Diffusion Probabilistic Models as Feature Extractors for Change DetectionarXiv 2022The Decades Progress on Code-Switching Research in NLP: A Systematic Survey on Trends and ChallengesarXiv 2022MIDI-DDSP: Detailed Control of Musical Performance via Hierarchical Modelingmidi-ddsp-detailed-control-of-musicalVideoScene: Distilling Video Diffusion Model to Generate 3D Scenes in One StepCVPR 2025 1ReaL: Efficient RLHF Training of Large Language Models with Parameter ReallocationarXiv 2024Text2Performer: Text-Driven Human Video GenerationICCV 2023 1Flowformer: Linearizing Transformers with Conservation FlowsarXiv 2022RelBench: A Benchmark for Deep Learning on Relational DatabasesarXiv 2024AutoVFX: Physically Realistic Video Editing from Natural Language InstructionsarXiv 2024NEAT: Neural Attention Fields for End-to-End Autonomous DrivingICCV 2021 10BEVPlace: Learning LiDAR-based Place Recognition using Bird's Eye View ImagesICCV 2023 1CLIP-Art: Contrastive Pre-training for Fine-Grained Art Classificationclip-art-contrastive-pre-training-for-fineGPT Can Solve Mathematical Problems Without a CalculatorarXiv 2023Polygonal Building Segmentation by Frame Field LearningarXiv 2020Gen-LaneNet: A Generalized and Scalable Approach for 3D Lane DetectionECCV 2020 8Gentopia: A Collaborative Platform for Tool-Augmented LLMsarXiv 2023FakeShield: Explainable Image Forgery Detection and Localization via Multi-modal Large Language ModelsarXiv 2024MeshXL: Neural Coordinate Field for Generative 3D Foundation ModelsarXiv 2024Monte Carlo Tree Search Boosts Reasoning via Iterative Preference LearningarXiv 2024Paint Transformer: Feed Forward Neural Painting with Stroke PredictionICCV 2021 10Context-Aware Sentence/Passage Term Importance Estimation For First Stage RetrievalarXiv 2019Flow Matching in Latent SpacearXiv 2023BERT has a Mouth, and It Must Speak: BERT as a Markov Random Field Language Modelbert-has-a-mouth-and-it-must-speak-bert-as-a-1StyleSpace Analysis: Disentangled Controls for StyleGAN Image GenerationCVPR 2021 1Efficient Image Super-Resolution Using Pixel AttentionarXiv 2020Measuring and Narrowing the Compositionality Gap in Language ModelsarXiv 2022STaRK: Benchmarking LLM Retrieval on Textual and Relational Knowledge BasesarXiv 2024PUMA: Secure Inference of LLaMA-7B in Five MinutesarXiv 2023EDICT: Exact Diffusion Inversion via Coupled TransformationsCVPR 2023 1PhysGen: Rigid-Body Physics-Grounded Image-to-Video GenerationarXiv 2024VIMA: General Robot Manipulation with Multimodal PromptsarXiv 2022Let's Think Outside the Box: Exploring Leap-of-Thought in Large Language Models with Creative Humor GenerationCVPR 2024 1Make-A-Protagonist: Generic Video Editing with An Ensemble of ExpertsarXiv 2023Thinking Like Transformersthinking-like-transformersSelf-Rectifying Diffusion Sampling with Perturbed-Attention GuidancearXiv 2024Distance-IoU Loss: Faster and Better Learning for Bounding Box RegressionarXiv 2019FABRIC: Personalizing Diffusion Models with Iterative FeedbackarXiv 2023GDRNPP: A Geometry-guided and Fully Learning-based Object Pose EstimatorCVPR 2021 1MLAgentBench: Evaluating Language Agents on Machine Learning ExperimentationarXiv 2023FruitNeRF: A Unified Neural Radiance Field based Fruit Counting FrameworkarXiv 2024DF-GAN: A Simple and Effective Baseline for Text-to-Image SynthesisCVPR 2022 1Training-Free Structured Diffusion Guidance for Compositional Text-to-Image SynthesisarXiv 2022Advancing LLM Reasoning Generalists with Preference TreesarXiv 2024Rethinking Knowledge Graph Propagation for Zero-Shot Learningrethinking-knowledge-graph-propagation-for-1CLIPA-v2: Scaling CLIP Training with 81.1% Zero-shot ImageNet Accuracy within a \$10,000 Budget; An Extra \$4,000 Unlocks 81.8% AccuracyarXiv 2023A Survey of Efficient Reasoning for Large Reasoning Models: Language, Multimodality, and BeyondarXiv 2025OpenGraph: Towards Open Graph Foundation ModelsarXiv 2024HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language ModelsCVPR 2024 1DeepA2: A Modular Framework for Deep Argument Analysis with Pretrained Neural Text2Text Language Models*SEM (NAACL) 2022 7"Why Should I Trust You?": Explaining the Predictions of Any ClassifierarXiv 2016NeuralLift-360: Lifting An In-the-wild 2D Photo to A 3D Object with 360° ViewsarXiv 2022SHERF: Generalizable Human NeRF from a Single ImageICCV 2023 1RMVPE: A Robust Model for Vocal Pitch Estimation in Polyphonic MusicarXiv 2023PeopleSansPeople: A Synthetic Data Generator for Human-Centric Computer VisionarXiv 2021TyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Diverse Languagestydi-qa-a-benchmark-for-information-seeking-1MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated CapabilitiesarXiv 2024DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video GenerationCVPR 2025 1ChEF: A Comprehensive Evaluation Framework for Standardized Assessment of Multimodal Large Language ModelsarXiv 2023Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language ModelsarXiv 2024LLM-Grounder: Open-Vocabulary 3D Visual Grounding with Large Language Model as an AgentarXiv 2023Radiative Gaussian Splatting for Efficient X-ray Novel View SynthesisarXiv 2024Pengi: An Audio Language Model for Audio Taskspengi-an-audio-language-model-for-audio-tasksDropout Reduces UnderfittingarXiv 2023CoCoCo: Improving Text-Guided Video Inpainting for Better Consistency, Controllability and CompatibilityarXiv 2024BERT-of-Theseus: Compressing BERT by Progressive Module ReplacingEMNLP 2020 11AnyTool: Self-Reflective, Hierarchical Agents for Large-Scale API CallsarXiv 2024Cross-Domain Image Captioning with Discriminative FinetuningCVPR 2023 1HumanMAC: Masked Motion Completion for Human Motion PredictionICCV 2023 1Emotional Speech-Driven Animation with Content-Emotion DisentanglementarXiv 2023Rethinking Benchmark and Contamination for Language Models with Rephrased SamplesarXiv 2023LlamaDuo: LLMOps Pipeline for Seamless Migration from Service LLMs to Small-Scale Local LLMsarXiv 2024Subject-Diffusion:Open Domain Personalized Text-to-Image Generation without Test-time Fine-tuningarXiv 2023YAYI-UIE: A Chat-Enhanced Instruction Tuning Framework for Universal Information ExtractionarXiv 2023Stabilizing Transformer Training by Preventing Attention Entropy CollapsearXiv 2023RAVE: Randomized Noise Shuffling for Fast and Consistent Video Editing with Diffusion ModelsCVPR 2024 1An Extendable, Efficient and Effective Transformer-based Object DetectorarXiv 2022Softmax-free Linear TransformersarXiv 2022Mini-DALLE3: Interactive Text to Image by Prompting Large Language ModelsarXiv 2023PuzzleAvatar: Assembling 3D Avatars from Personal AlbumsarXiv 2024ALIKED: A Lighter Keypoint and Descriptor Extraction Network via Deformable TransformationarXiv 2023Consistency Trajectory Models: Learning Probability Flow ODE Trajectory of DiffusionarXiv 2023Fuzz4All: Universal Fuzzing with Large Language ModelsarXiv 2023Hopular: Modern Hopfield Networks for Tabular Datahopular-modern-hopfield-networks-for-tabularVideoFlow: Exploiting Temporal Cues for Multi-frame Optical Flow EstimationICCV 2023 1DF40: Toward Next-Generation Deepfake DetectionarXiv 2024Simplified State Space Layers for Sequence ModelingarXiv 2022GNM: A General Navigation Model to Drive Any RobotarXiv 2022FaceStudio: Put Your Face Everywhere in SecondsarXiv 2023Automated Concatenation of Embeddings for Structured Predictionautomated-concatenation-of-embeddings-forRelay Diffusion: Unifying diffusion process across resolutions for image synthesisrelay-diffusion-unifying-diffusion-processAddressing Representation Collapse in Vector Quantized Models with One Linear LayerICCV 2025Beyond Static Features for Temporally Consistent 3D Human Pose and Shape from a VideoCVPR 2021 1FaceXFormer: A Unified Transformer for Facial AnalysisICCV 2025KorNLI and KorSTS: New Benchmark Datasets for Korean Natural Language UnderstandingFindings of the Association for Computational Linguistics 2020LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMsarXiv 2025CVEfixes: Automated Collection of Vulnerabilities and Their Fixes from Open-Source SoftwarearXiv 2021Making the Most of Text Semantics to Improve Biomedical Vision--Language ProcessingarXiv 2022Prometheus: Inducing Fine-grained Evaluation Capability in Language ModelsarXiv 2023Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question AnsweringarXiv 2025Robust Invisible Video Watermarking with AttentionarXiv 2019Learning JPEG Compression Artifacts for Image Manipulation Detection and
LocalizationarXiv 2021ZipIt! Merging Models from Different Tasks without TrainingarXiv 2023RepMLP: Re-parameterizing Convolutions into Fully-connected Layers for Image RecognitionarXiv 2021Evolving from Single-modal to Multi-modal Facial Deepfake Detection: Progress and ChallengesarXiv 2024InCoder: A Generative Model for Code Infilling and SynthesisarXiv 2022Multi-view Self-supervised Deep Learning for 6D Pose Estimation in the Amazon Picking ChallengearXiv 2016EVA2.0: Investigating Open-Domain Chinese Dialogue Systems with Large-Scale Pre-TrainingarXiv 2022Bayesian Flow NetworksarXiv 2023Long-VITA: Scaling Large Multi-modal Models to 1 Million Tokens with Leading Short-Context AccurayarXiv 2025Controllable Multi-Interest Framework for RecommendationarXiv 2020SynCode: LLM Generation with Grammar AugmentationarXiv 2024GeoCLIP: Clip-Inspired Alignment between Locations and Images for Effective Worldwide Geo-localizationgeoclip-clip-inspired-alignment-betweenCumulative Reasoning with Large Language ModelsarXiv 2023Personalized Federated Learning with Moreau EnvelopesNeurIPS 2020 12Diffusion360: Seamless 360 Degree Panoramic Image Generation based on Diffusion ModelsarXiv 2023FreeDoM: Training-Free Energy-Guided Conditional Diffusion ModelICCV 2023 1GaussianCity: Generative Gaussian Splatting for Unbounded 3D City GenerationarXiv 2024VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and DatasetarXiv 2023Understanding Self-supervised Learning with Dual Deep NetworksarXiv 2020GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image GenerationarXiv 2025Ford Multi-AV Seasonal DatasetarXiv 2020IML-ViT: Benchmarking Image Manipulation Localization by Vision TransformerarXiv 2023Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-DenoisingarXiv 2023MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence Frontiersmauve-measuring-the-gap-between-neural-textBERTology Meets Biology: Interpreting Attention in Protein Language ModelsICLR 2021 1Parsing R-CNN for Instance-Level Human Analysisparsing-r-cnn-for-instance-level-human-1Automatic Liver and Tumor Segmentation of CT and MRI Volumes using Cascaded Fully Convolutional Neural NetworksarXiv 2017Theory, Analysis, and Best Practices for Sigmoid Self-AttentionarXiv 2024GenWarp: Single Image to Novel Views with Semantic-Preserving Generative WarpingarXiv 2024CAMixerSR: Only Details Need More "Attention"CVPR 2024 1