All papers
Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.
AnimateAnything: Fine-Grained Open Domain Image Animation with Motion GuidancearXiv 2023StyleNeRF: A Style-based 3D-Aware Generator for High-resolution Image Synthesisstylenerf-a-style-based-3d-aware-generatorNeural Body: Implicit Neural Representations with Structured Latent Codes for Novel View Synthesis of Dynamic HumansCVPR 2021 1DeepCache: Accelerating Diffusion Models for FreeCVPR 2024 1Do Adversarially Robust ImageNet Models Transfer Better?NeurIPS 2020 12LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion ModelsarXiv 2023Blind Face Restoration via Deep Multi-scale Component DictionariesECCV 2020 8Performance evaluation of SLAM-ASR: The Good, the Bad, the Ugly, and the Way ForwardarXiv 2024GPT Understands, TooarXiv 2021TinyLLaVA Factory: A Modularized Codebase for Small-scale Large Multimodal ModelsarXiv 2024SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage TrainingarXiv 2024DeepfakeBench: A Comprehensive Benchmark of Deepfake Detectiondeepfakebench-a-comprehensive-benchmark-ofSimPO: Simple Preference Optimization with a Reference-Free RewardarXiv 2024Scaling Up Your Kernels to 31x31: Revisiting Large Kernel Design in CNNsCVPR 2022 1ReWOO: Decoupling Reasoning from Observations for Efficient Augmented Language ModelsarXiv 2023Pivotal Tuning for Latent-based Editing of Real ImagesarXiv 2021Hallucination of Multimodal Large Language Models: A SurveyarXiv 2024ChatGPT for Zero-shot Dialogue State Tracking: A Solution or an Opportunity?arXiv 2023PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning OptimizationarXiv 2023Multimodal Chain-of-Thought Reasoning: A Comprehensive SurveyarXiv 2025FancyVideo: Towards Dynamic and Consistent Video Generation via Cross-frame Textual GuidancearXiv 2024End-to-End Semi-Supervised Object Detection with Soft TeacherICCV 2021 10Defending Against Neural Fake Newsdefending-against-neural-fake-news-1Uformer: A General U-Shaped Transformer for Image RestorationCVPR 2022 1FBNet: Hardware-Aware Efficient ConvNet Design via Differentiable Neural Architecture Searchfbnet-hardware-aware-efficient-convnet-design-1AutoWebGLM: A Large Language Model-based Web Navigating AgentarXiv 2024Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion ModelsarXiv 2025GaussianAvatars: Photorealistic Head Avatars with Rigged 3D GaussiansCVPR 2024 1MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUsarXiv 2024TabM: Advancing Tabular Deep Learning with Parameter-Efficient EnsemblingarXiv 2024Follow-Your-Click: Open-domain Regional Image Animation via Short PromptsarXiv 2024Towards AI-Complete Question Answering: A Set of Prerequisite Toy TasksarXiv 2015TransPixeler: Advancing Text-to-Video Generation with TransparencyCVPR 2025 1A flexible framework for accurate LiDAR odometry, map manipulation, and
localizationarXiv 2024Context Encoders: Feature Learning by Inpaintingcontext-encoders-feature-learning-by-1SpanBERT: Improving Pre-training by Representing and Predicting Spansspanbert-improving-pre-training-by-1T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on EdgearXiv 2024Projected GANs Converge FasterNeurIPS 2021 12MegActor: Harness the Power of Raw Video for Vivid Portrait AnimationarXiv 2024A Survey on Spoken Language Understanding: Recent Advances and New FrontiersarXiv 2021FasterViT: Fast Vision Transformers with Hierarchical AttentionarXiv 2023LEAF: A Benchmark for Federated SettingsarXiv 2018GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibriumgans-trained-by-a-two-time-scale-update-rule-1RDD2022: A multi-national image dataset for automatic Road Damage DetectionarXiv 2022Lite-HRNet: A Lightweight High-Resolution NetworkCVPR 2021 1COCO-Stuff: Thing and Stuff Classes in Contextcoco-stuff-thing-and-stuff-classes-in-context-1UniFormer: Unified Transformer for Efficient Spatiotemporal Representation LearningarXiv 2022Evaluating Real-World Robot Manipulation Policies in SimulationarXiv 2024EduChat: A Large-Scale Language Model-based Chatbot System for Intelligent EducationarXiv 2023Neural Network DiffusionarXiv 2024ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated CasesarXiv 2023PromptFix: You Prompt and We Fix the PhotoarXiv 2024OmniQuant: Omnidirectionally Calibrated Quantization for Large Language ModelsarXiv 2023Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language ModelsarXiv 2024DETR3D: 3D Object Detection from Multi-view Images via 3D-to-2D QueriesarXiv 2021Timer: Generative Pre-trained Transformers Are Large Time Series ModelsarXiv 2024CrisperWhisper: Accurate Timestamps on Verbatim Speech TranscriptionsarXiv 2024A Survey on Large Language Model-Based Game AgentsarXiv 2024SEED-Story: Multimodal Long Story Generation with Large Language ModelarXiv 2024OpenEMMA: Open-Source Multimodal Model for End-to-End Autonomous DrivingarXiv 2024SGPT: GPT Sentence Embeddings for Semantic SearcharXiv 2022A Survey on In-context LearningarXiv 2022Graph Matching with Bi-level Noisy CorrespondenceICCV 2023 1Once Detected, Never Lost: Surpassing Human Performance in Offline LiDAR based 3D Object DetectionICCV 2023 1DiffusionCLIP: Text-Guided Diffusion Models for Robust Image Manipulationdiffusionclip-text-guided-image-manipulation-1MiniGPT-5: Interleaved Vision-and-Language Generation via Generative VokensarXiv 2023Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free LuncharXiv 2023BianQue: Balancing the Questioning and Suggestion Ability of Health LLMs with Multi-turn Health Conversations Polished by ChatGPTarXiv 2023Alpha-CLIP: A CLIP Model Focusing on Wherever You WantCVPR 2024 1A Transformer-based Framework for Multivariate Time Series Representation Learninga-transformer-based-framework-forProAgent: From Robotic Process Automation to Agentic Process AutomationarXiv 2023ChemCrow: Augmenting large-language models with chemistry toolsarXiv 2023ControlVideo: Training-free Controllable Text-to-Video GenerationarXiv 2023HHAvatar: Gaussian Head Avatar with Dynamic HairsCVPR 2024 1VoxelNeXt: Fully Sparse VoxelNet for 3D Object Detection and Trackingvoxelnext-fully-sparse-voxelnet-for-3d-objectSpatial As Deep: Spatial CNN for Traffic Scene UnderstandingarXiv 2017StackGAN++: Realistic Image Synthesis with Stacked Generative Adversarial NetworksarXiv 2017Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of ExpertsarXiv 2024Simplifying Graph Convolutional NetworksarXiv 20193D-GPT: Procedural 3D Modeling with Large Language ModelsarXiv 2023FastFCN: Rethinking Dilated Convolution in the Backbone for Semantic SegmentationarXiv 2019Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by SteparXiv 2025Text-to-3D using Gaussian SplattingCVPR 2024 1One Model, Many Languages: Meta-learning for Multilingual Text-to-SpeecharXiv 2020GIM: Learning Generalizable Image Matcher From Internet VideosarXiv 2024pyvene: A Library for Understanding and Improving PyTorch Models via InterventionsarXiv 2024D2-Net: A Trainable CNN for Joint Detection and Description of Local FeaturesarXiv 2019We don't need no bounding-boxes: Training object class detectors using
only human verificationarXiv 2016Beyond Reward Hacking: Causal Rewards for Large Language Model AlignmentarXiv 2025SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model InferencearXiv 2025Compositional Semantic Parsing on Semi-Structured Tablescompositional-semantic-parsing-on-semi-1Osprey: Pixel Understanding with Visual Instruction TuningCVPR 2024 1SVTRv2: CTC Beats Encoder-Decoder Models in Scene Text RecognitionICCV 2025LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RLarXiv 2025Automated Hate Speech Detection and the Problem of Offensive LanguagearXiv 2017Vista: A Generalizable Driving World Model with High Fidelity and Versatile ControllabilityarXiv 2024Dense Extreme Inception Network for Edge DetectionarXiv 2021DoctorGLM: Fine-tuning your Chinese Doctor is not a Herculean TaskarXiv 2023MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and EditingICCV 2023 1Fast and Eager k-Medoids Clustering: O(k) Runtime Improvement of the PAM, CLARA, and CLARANS AlgorithmsarXiv 2020MINT-1T: Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion TokensarXiv 2024SAM-Med3D: Towards General-purpose Segmentation Models for Volumetric Medical ImagesarXiv 2023LaMini-LM: A Diverse Herd of Distilled Models from Large-Scale InstructionsarXiv 2023Augmenting Language Models with Long-Term Memoryaugmenting-language-models-with-long-termTF-ICON: Diffusion-Based Training-Free Cross-Domain Image CompositionICCV 2023 1VanillaNet: the Power of Minimalism in Deep Learningvanillanet-the-power-of-minimalism-in-deep12-in-1: Multi-Task Vision and Language Representation Learning12-in-1-multi-task-vision-and-language-1How Do Vision Transformers Work?how-do-vision-transformers-workDiffuSeq: Sequence to Sequence Text Generation with Diffusion ModelsarXiv 2022HAWQ-V2: Hessian Aware trace-Weighted Quantization of Neural NetworksNeurIPS 2020 12Modeling Relational Data with Graph Convolutional NetworksarXiv 2017Dataset DistillationarXiv 2018Google Landmarks Dataset v2 -- A Large-Scale Benchmark for Instance-Level Recognition and RetrievalarXiv 2020ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text GenerationarXiv 2024No Language is an Island: Unifying Chinese and English in Financial Large Language Models, Instruction Data, and BenchmarksarXiv 2024SERL: A Software Suite for Sample-Efficient Robotic Reinforcement LearningarXiv 2024Visual Attention NetworkarXiv 2022Benchmarks and leaderboards for sound demixing tasksarXiv 2023TransNet V2: An effective deep network architecture for fast shot transition detectionarXiv 2020DistServe: Disaggregating Prefill and Decoding for Goodput-optimized
Large Language Model ServingarXiv 2024Two at Once: Enhancing Learning and Generalization Capacities via IBN-Nettwo-at-once-enhancing-learning-and-1Adan: Adaptive Nesterov Momentum Algorithm for Faster Optimizing Deep ModelsarXiv 2022Adaptivity without Compromise: A Momentumized, Adaptive, Dual Averaged Gradient Method for Stochastic OptimizationarXiv 2021HandRefiner: Refining Malformed Hands in Generated Images by Diffusion-based Conditional InpaintingarXiv 2023Shikra: Unleashing Multimodal LLM's Referential Dialogue MagicarXiv 2023Expressive Text-to-Image Generation with Rich TextICCV 2023 1Multilingual Autoregressive Entity LinkingarXiv 2021CMMLU: Measuring massive multitask language understanding in ChinesearXiv 2023PhoGPT: Generative Pre-training for VietnamesearXiv 2023TS2Vec: Towards Universal Representation of Time SeriesarXiv 2021CraftsMan3D: High-fidelity Mesh Generation with 3D Native Generation and Interactive Geometry RefinerarXiv 2024TEXTure: Text-Guided Texturing of 3D ShapesarXiv 2023Evaluating Large Language Models: A Comprehensive SurveyarXiv 2023Internet of Agents: Weaving a Web of Heterogeneous Agents for Collaborative IntelligencearXiv 2024NL-Augmenter: A Framework for Task-Sensitive Natural Language AugmentationarXiv 2021Multi-Scale Context Aggregation by Dilated ConvolutionsarXiv 2015Dense Text Retrieval based on Pretrained Language Models: A SurveyarXiv 2022Asymmetric Loss For Multi-Label ClassificationICCV 2021 10FRESCO: Spatial-Temporal Correspondence for Zero-Shot Video TranslationCVPR 2024 1Depth-supervised NeRF: Fewer Views and Faster Training for FreeCVPR 2022 1Video-R1: Reinforcing Video Reasoning in MLLMsarXiv 2025GroupViT: Semantic Segmentation Emerges from Text SupervisionCVPR 2022 1PSGAN: Pose and Expression Robust Spatial-Aware GAN for Customizable Makeup Transferpsgan-pose-and-expression-robust-spatialCenterMask : Real-Time Anchor-Free Instance Segmentationcentermask-real-time-anchor-free-instanceScaling Transformer to 1M tokens and beyond with RMTarXiv 2023Rethinking Keypoint Representations: Modeling Keypoints and Poses as Objects for Multi-Person Human Pose EstimationarXiv 2021PatrickStar: Parallel Training of Pre-trained Models via Chunk-based Memory ManagementarXiv 2021GPT4Tools: Teaching Large Language Model to Use Tools via Self-instructionNeurIPS 2023 11OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and ReasoningarXiv 2024AGIEval: A Human-Centric Benchmark for Evaluating Foundation ModelsarXiv 2023Unsupervised Translation of Programming LanguagesNeurIPS 2020 12Text2SQL is Not Enough: Unifying AI and Databases with TAGarXiv 2024ResAdapter: Domain Consistent Resolution Adapter for Diffusion ModelsarXiv 2024One-Stage 3D Whole-Body Mesh Recovery with Component Aware TransformerCVPR 2023 1Total-Text: A Comprehensive Dataset for Scene Text Detection and RecognitionarXiv 2017Rethinking the Value of Labels for Improving Class-Imbalanced LearningNeurIPS 2020 12MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion ModelarXiv 2024PhoBERT: Pre-trained language models for VietnameseFindings of the Association for Computational Linguistics 2020Attend-and-Excite: Attention-Based Semantic Guidance for Text-to-Image Diffusion ModelsarXiv 2023Towards Deep Learning Models Resistant to Adversarial Attackstowards-deep-learning-models-resistant-to-1LightGaussian: Unbounded 3D Gaussian Compression with 15x Reduction and 200+ FPSarXiv 2023Unsupervised Feature Learning via Non-Parametric Instance-level DiscriminationarXiv 2018Prompt-Free Diffusion: Taking "Text" out of Text-to-Image Diffusion ModelsCVPR 2024 1End-to-End Video Instance Segmentation with TransformersCVPR 2021 1Robust fine-tuning of zero-shot modelsrobust-fine-tuning-of-zero-shot-models-1Multimodal Image Synthesis and Editing: The Generative AI EraarXiv 2021VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsICLR 2020 1FewRel 2.0: Towards More Challenging Few-Shot Relation Classificationfewrel-20-towards-more-challenging-few-shot-1Tracking Anything in High QualityarXiv 2023Open-Vocabulary Semantic Segmentation with Mask-adapted CLIPCVPR 2023 1UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of RerankersarXiv 2023CRITIC: Large Language Models Can Self-Correct with Tool-Interactive CritiquingarXiv 2023Identity-Preserving Talking Face Generation with Landmark and Appearance PriorsCVPR 2023 1Personalizing Text-to-Image Generation via Aesthetic GradientsarXiv 2022Deep Learning for Classical Japanese LiteraturearXiv 2018Random Erasing Data AugmentationarXiv 2017PIDNet: A Real-time Semantic Segmentation Network Inspired by PID ControllersCVPR 2023 1Making Pre-trained Language Models Better Few-shot LearnersACL 2021 5Let 2D Diffusion Model Know 3D-Consistency for Robust Text-to-3D GenerationarXiv 2023mLUKE: The Power of Entity Representations in Multilingual Pretrained Language ModelsACL 2022 5Sparks of Large Audio Models: A Survey and OutlookarXiv 2023WonderWorld: Interactive 3D Scene Generation from a Single ImageCVPR 2025 1Evaluating Protein Transfer Learning with TAPEevaluating-protein-transfer-learning-with-1Joint Learning of Sentence Embeddings for Relevance and Entailmentjoint-learning-of-sentence-embeddings-for-1XGen-7B Technical ReportarXiv 2023SARDet-100K: Towards Open-Source Benchmark and ToolKit for Large-Scale SAR Object DetectionarXiv 2024SqueezeLLM: Dense-and-Sparse QuantizationarXiv 2023Fin-R1: A Large Language Model for Financial Reasoning through Reinforcement LearningarXiv 2025DistriFusion: Distributed Parallel Inference for High-Resolution Diffusion ModelsCVPR 2024 1Label-Efficient Semantic Segmentation with Diffusion Modelslabel-efficient-semantic-segmentation-withTime Series Classification from Scratch with Deep Neural Networks: A Strong BaselinearXiv 2016Latent-NeRF for Shape-Guided Generation of 3D Shapes and TexturesCVPR 2023 1Span-based Joint Entity and Relation Extraction with Transformer Pre-trainingarXiv 2019GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed AudioarXiv 2021Learning Enriched Features for Real Image Restoration and EnhancementECCV 2020 8CrossWOZ: A Large-Scale Chinese Cross-Domain Task-Oriented Dialogue Datasetcrosswoz-a-large-scale-chinese-cross-domain-1Only a Matter of Style: Age Transformation Using a Style-Based Regression ModelarXiv 2021DifFace: Blind Face Restoration with Diffused Error ContractionarXiv 2022SparseFlex: High-Resolution and Arbitrary-Topology 3D Shape ModelingICCV 2025CoBa: Convergence Balancer for Multitask Finetuning of Large Language ModelsarXiv 2024