0

All papers

Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.

DepthCrafter: Generating Consistent Long Depth Sequences for Open-world VideosCVPR 2025 1Privacy-Preserving Face Recognition Using Random Frequency ComponentsICCV 2023 1CCNet: Criss-Cross Attention for Semantic Segmentationccnet-criss-cross-attention-for-semantic-1TripoSG: High-Fidelity 3D Shape Synthesis using Large-Scale Rectified Flow ModelsarXiv 2025Automated Design of Agentic SystemsarXiv 2024Image Super-Resolution Using Very Deep Residual Channel Attention Networksimage-super-resolution-using-very-deep-1MotionCtrl: A Unified and Flexible Motion Controller for Video GenerationarXiv 2023AgentTuning: Enabling Generalized Agent Abilities for LLMsarXiv 2023Beyond Appearance: a Semantic Controllable Self-Supervised Learning Framework for Human-Centric Visual TasksCVPR 2023 1Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You ThinkarXiv 2024Focused Transformer: Contrastive Training for Context ScalingNeurIPS 2023 11pixelNeRF: Neural Radiance Fields from One or Few ImagesCVPR 2021 1Skywork: A More Open Bilingual Foundation ModelarXiv 2023NerfAcc: Efficient Sampling Accelerates NeRFsICCV 2023 1One Transformer Fits All Distributions in Multi-Modal Diffusion at ScalearXiv 2023NaturalSpeech: End-to-End Text to Speech Synthesis with Human-Level QualityarXiv 2022SCAN: Learning to Classify Images without LabelsECCV 2020 8StableVideo: Text-driven Consistency-aware Diffusion Video EditingICCV 2023 1JoJoGAN: One Shot Face Stylizationjojogan-one-shot-face-stylizationHAT: Hybrid Attention Transformer for Image RestorationarXiv 2023BM25S: Orders of magnitude faster lexical search via eager sparse scoringarXiv 2024Boundary-Aware Segmentation Network for Mobile and Web ApplicationsarXiv 2021metric-learn: Metric Learning Algorithms in PythonarXiv 2019TorchGAN: A Flexible Framework for GAN Training and EvaluationarXiv 2019Scaling Synthetic Data Creation with 1,000,000,000 PersonasarXiv 2024WavLLM: Towards Robust and Adaptive Speech Large Language ModelarXiv 2024Secrets of RLHF in Large Language Models Part II: Reward ModelingarXiv 2024SpeechAlign: Aligning Speech Generation to Human PreferencesarXiv 2024AST: Audio Spectrogram TransformerarXiv 2021Deal or No Deal? End-to-End Learning for Negotiation DialoguesarXiv 2017Evolutionary Optimization of Model Merging RecipesarXiv 2024Taiwan LLM: Bridging the Linguistic Divide with a Culturally Aligned Language ModelarXiv 2023Decoupling Magnitude and Phase Estimation with Deep ResUNet for Music Source SeparationarXiv 2021VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic FaithfulnessarXiv 2025RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist RobotsarXiv 2024Visual Attribute Transfer through Deep Image AnalogyarXiv 2017Learning Continuous Image Representation with Local Implicit Image FunctionCVPR 2021 1Matching Anything by Segmenting AnythingCVPR 2024 1DEIM: DETR with Improved Matching for Fast ConvergenceCVPR 2025 1Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free VideosarXiv 2023Multimodal Foundation Models: From Specialists to General-Purpose AssistantsarXiv 2023Path Aggregation Network for Instance Segmentationpath-aggregation-network-for-instance-1Arbitrary-steps Image Super-resolution via Diffusion InversionCVPR 2025 1ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video GenerationarXiv 2024Geometric-aware Pretraining for Vision-centric 3D Object DetectionarXiv 2023Leveraging Vision-Centric Multi-Modal Expertise for 3D Object Detectionleveraging-vision-centric-multi-modalGeneralized Decoding for Pixel, Image, and LanguageCVPR 2023 1MedSegDiff: Medical Image Segmentation with Diffusion Probabilistic ModelarXiv 2022Contrastive Multiview CodingECCV 2020 8Versatile Diffusion: Text, Images and Variations All in One Diffusion ModelICCV 2023 1Designing a Practical Degradation Model for Deep Blind Image Super-ResolutionICCV 2021 10Large Language Models Are Human-Level Prompt EngineersarXiv 2022Enhancing Photorealism EnhancementarXiv 2021Unlimited-Size Diffusion RestorationarXiv 2023Efficient Diffusion Model for Image Restoration by Residual ShiftingarXiv 2024DETRs with Collaborative Hybrid Assignments TrainingICCV 2023 1Voice Separation with an Unknown Number of Multiple SpeakersICML 2020 1ConceptNet 5.5: An Open Multilingual Graph of General KnowledgearXiv 2016BundleSDF: Neural 6-DoF Tracking and 3D Reconstruction of Unknown ObjectsCVPR 2023 1Wide Residual NetworksarXiv 2016PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM CompressionarXiv 2024Tri-Perspective View for Vision-Based 3D Semantic Occupancy PredictionCVPR 2023 1Attention, Learn to Solve Routing Problems!attention-learn-to-solve-routing-problems-1MotionBERT: A Unified Perspective on Learning Human Motion RepresentationsICCV 2023 1TinyGPT-V: Efficient Multimodal Large Language Model via Small BackbonesarXiv 2023Text Summarization with Pretrained Encoderstext-summarization-with-pretrained-encoders-1Image Segmentation Using Text and Image PromptsCVPR 2022 1From System 1 to System 2: A Survey of Reasoning Large Language ModelsarXiv 2025Learning to (Learn at Test Time): RNNs with Expressive Hidden StatesarXiv 2024PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information FunnelingarXiv 2024EMMA: Your Text-to-Image Diffusion Model Can Secretly Accept Multi-Modal PromptsarXiv 2024ELLA: Equip Diffusion Models with LLM for Enhanced Semantic AlignmentarXiv 2024Fast On-device LLM Inference with NPUsarXiv 2024Sequence-Level Knowledge Distillationsequence-level-knowledge-distillation-1Shape of Motion: 4D Reconstruction from a Single VideoICCV 2025Motion Representations for Articulated Animationmotion-representations-for-articulatedMedical SAM Adapter: Adapting Segment Anything Model for Medical Image SegmentationarXiv 2023CutMix: Regularization Strategy to Train Strong Classifiers with Localizable Featurescutmix-regularization-strategy-to-train-1DragDiffusion: Harnessing Diffusion Models for Interactive Point-based Image EditingCVPR 2024 1GIRAFFE: Representing Scenes as Compositional Generative Neural Feature FieldsCVPR 2021 1HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech SynthesisarXiv 2023Self-Play Fine-Tuning Converts Weak Language Models to Strong Language ModelsarXiv 2024A Survey on Knowledge Distillation of Large Language ModelsarXiv 2024Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference OptimizationarXiv 2024Global Context NetworksarXiv 2020Microsoft COCO Captions: Data Collection and Evaluation ServerarXiv 2015Cyclical Learning Rates for Training Neural NetworksarXiv 2015Segment Any Text: A Universal Approach for Robust, Efficient and Adaptable Sentence SegmentationarXiv 2024Fractal Generative ModelsarXiv 2025Diffusion-LM Improves Controllable Text GenerationarXiv 2022Matterport3D: Learning from RGB-D Data in Indoor EnvironmentsarXiv 2017Swin UNETR: Swin Transformers for Semantic Segmentation of Brain Tumors in MRI ImagesarXiv 2022Scalable Zero-shot Entity Linking with Dense Entity RetrievalEMNLP 2020 11mixup: Beyond Empirical Risk Minimizationmixup-beyond-empirical-risk-minimization-1StyleGAN-T: Unlocking the Power of GANs for Fast Large-Scale Text-to-Image SynthesisarXiv 2023Matcha-TTS: A fast TTS architecture with conditional flow matchingarXiv 2023ECON: Explicit Clothed humans Optimized via Normal integrationCVPR 2023 1Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetICCV 2021 10StyleGAN-Human: A Data-Centric Odyssey of Human GenerationarXiv 2022DreamClear: High-Capacity Real-World Image Restoration with Privacy-Safe Dataset CurationarXiv 2024Single Headed Attention RNN: Stop Thinking With Your HeadarXiv 2019Self-Supervised Pre-Training of Swin Transformers for 3D Medical Image AnalysisCVPR 2022 1Neural Architectures for Named Entity Recognitionneural-architectures-for-named-entity-1General Object Foundation Model for Images and Videos at ScaleCVPR 2024 1Large Language Model based Multi-Agents: A Survey of Progress and ChallengesarXiv 2024AirSLAM: An Efficient and Illumination-Robust Point-Line Visual SLAM SystemarXiv 2024Plug and Play Language Models: A Simple Approach to Controlled Text GenerationICLR 2020 1KGAT: Knowledge Graph Attention Network for RecommendationarXiv 2019Nuclei instance segmentation and classification in histopathology images with StarDistarXiv 2022DetectoRS: Detecting Objects with Recursive Feature Pyramid and Switchable Atrous Convolutiondetectors-detecting-objects-with-recursivePrinciple-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervisionprinciple-driven-self-alignment-of-languageAutoInt: Automatic Feature Interaction Learning via Self-Attentive Neural NetworksarXiv 2018Common Objects in 3D: Large-Scale Learning and Evaluation of Real-life 3D Category ReconstructionarXiv 2021MarioGPT: Open-Ended Text2Level Generation through Large Language Modelsmariogpt-open-ended-text2level-generationGenerative Data Augmentation using LLMs improves Distributional Robustness in Question AnsweringarXiv 2023Automated Machine Learning on Graphs: A SurveyarXiv 2021Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video GenerationarXiv 2023GAN Inversion: A SurveyarXiv 2021VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language TasksarXiv 2024Improved Distribution Matching Distillation for Fast Image SynthesisarXiv 2024Improving visual image reconstruction from human brain activity using latent diffusion models via multiple decoded inputsarXiv 2023Search-o1: Agentic Search-Enhanced Large Reasoning ModelsarXiv 2025MASS: Masked Sequence to Sequence Pre-training for Language GenerationarXiv 2019Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model ParametersarXiv 2024Safety Assessment of Chinese Large Language ModelsarXiv 2023Concept Sliders: LoRA Adaptors for Precise Control in Diffusion ModelsarXiv 2023Spectral Normalization for Generative Adversarial Networksspectral-normalization-for-generative-1Allegro: Open the Black Box of Commercial-Level Video Generation ModelarXiv 2024Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with TransformersCVPR 2021 1VITON-HD: High-Resolution Virtual Try-On via Misalignment-Aware NormalizationCVPR 2021 1Cognitive Architectures for Language AgentsarXiv 2023AIDE: AI-Driven Exploration in the Space of CodearXiv 2025LSA: Modeling Aspect Sentiment Coherency via Local Sentiment AggregationarXiv 2021Biomedical Named Entity Recognition at ScalearXiv 2020Channel Pruning for Accelerating Very Deep Neural Networkschannel-pruning-for-accelerating-very-deep-1Bridging the Gap Between Anchor-based and Anchor-free Detection via Adaptive Training Sample Selectionbridging-the-gap-between-anchor-based-and-1Aria: An Open Multimodal Native Mixture-of-Experts ModelarXiv 2024DisCo: Disentangled Control for Realistic Human Dance GenerationCVPR 2024 1TencentPretrain: A Scalable and Flexible Toolkit for Pre-training Models of Different ModalitiesarXiv 2022Real3D-Portrait: One-shot Realistic 3D Talking Portrait SynthesisarXiv 2024Patches Are All You Need?patches-are-all-you-needPyABSA: A Modularized Framework for Reproducible Aspect-based Sentiment AnalysisarXiv 2022Direct Voxel Grid Optimization: Super-fast Convergence for Radiance Fields ReconstructionCVPR 2022 1M2DGR: A Multi-sensor and Multi-scenario SLAM Dataset for Ground RobotsarXiv 2021DeepRobust: A PyTorch Library for Adversarial Attacks and DefensesarXiv 2020Large Multilingual Models Pivot Zero-Shot Multimodal Learning across LanguagesarXiv 2023Splatter Image: Ultra-Fast Single-View 3D ReconstructionCVPR 2024 1AdaBelief Optimizer: Adapting Stepsizes by the Belief in Observed GradientsNeurIPS 2020 12RoFormer: Enhanced Transformer with Rotary Position EmbeddingarXiv 2021VideoGPT: Video Generation using VQ-VAE and TransformersarXiv 2021Unlimiformer: Long-Range Transformers with Unlimited Length InputNeurIPS 2023 11Cascade R-CNN: Delving into High Quality Object Detectioncascade-r-cnn-delving-into-high-quality-1EnlightenGAN: Deep Light Enhancement without Paired SupervisionarXiv 2019LINE: Large-scale Information Network EmbeddingarXiv 2015Scaling Up Your Kernels: Large Kernel Design in ConvNets towards Universal RepresentationsarXiv 2024A Closer Look at Spatiotemporal Convolutions for Action Recognitiona-closer-look-at-spatiotemporal-convolutions-1SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant TransformersarXiv 2024LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context MultitasksarXiv 2024The BrowserGym Ecosystem for Web Agent ResearcharXiv 2024MotionDirector: Motion Customization of Text-to-Video Diffusion ModelsarXiv 2023OpenDelta: A Plug-and-play Library for Parameter-efficient Adaptation of Pre-trained ModelsarXiv 2023RepViT-SAM: Towards Real-Time Segmenting AnythingarXiv 2023Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8BarXiv 2024Off-Policy Primal-Dual Safe Reinforcement LearningarXiv 2024PubLayNet: largest dataset ever for document layout analysisarXiv 2019TinyMPC: Model-Predictive Control on Resource-Constrained MicrocontrollersarXiv 2023Segment Anything Meets Point TrackingarXiv 2023MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesismelgan-generative-adversarial-networks-for-1Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesisarXiv 2023HyperStyle: StyleGAN Inversion with HyperNetworks for Real Image EditingCVPR 2022 1StyleGAN2 Distillation for Feed-forward Image Manipulationstylegan2-distillation-for-feed-forward-image-1PointPillars: Fast Encoders for Object Detection from Point Cloudspointpillars-fast-encoders-for-object-1PatchFusion: An End-to-End Tile-Based Framework for High-Resolution Monocular Metric Depth EstimationCVPR 2024 13DTopia-XL: Scaling High-quality 3D Asset Generation via Primitive DiffusionCVPR 2025 1Self-Attention Generative Adversarial Networksself-attention-generative-adversarial-1IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architecturesimpala-scalable-distributed-deep-rl-with-1Pixel-Aware Stable Diffusion for Realistic Image Super-resolution and Personalized StylizationarXiv 2023A Novel Transformer Based Semantic Segmentation Scheme for Fine-Resolution Remote Sensing ImagesarXiv 2021LidarGait: Benchmarking 3D Gait Recognition with Point CloudsCVPR 2023 1Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Taskspider-a-large-scale-human-labeled-dataset-1Deep Long-Tailed Learning: A SurveyarXiv 2021CLUECorpus2020: A Large-scale Chinese Corpus for Pre-training Language ModelarXiv 2020Plug-and-Play Diffusion Features for Text-Driven Image-to-Image TranslationCVPR 2023 1AdaLomo: Low-memory Optimization with Adaptive Learning RatearXiv 2023Reasoning with Language Model Prompting: A SurveyarXiv 2022StyleGAN-XL: Scaling StyleGAN to Large Diverse DatasetsarXiv 2022MVDream: Multi-view Diffusion for 3D GenerationarXiv 2023Designing an Encoder for StyleGAN Image ManipulationarXiv 2021SurroundOcc: Multi-Camera 3D Occupancy Prediction for Autonomous DrivingICCV 2023 1Principled Instructions Are All You Need for Questioning LLaMA-1/2, GPT-3.5/4arXiv 2023Averaging Weights Leads to Wider Optima and Better GeneralizationarXiv 2018Decoupling Representation and Classifier for Long-Tailed RecognitionICLR 2020 1CNN-generated images are surprisingly easy to spot... for nowcnn-generated-images-are-surprisingly-easy-to-1Large Language Models on Graphs: A Comprehensive SurveyarXiv 2023Text2Mesh: Text-Driven Neural Stylization for MeshesCVPR 2022 1LXMERT: Learning Cross-Modality Encoder Representations from Transformerslxmert-learning-cross-modality-encoder-1LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative FusionarXiv 2023Improved Vector Quantized Diffusion ModelsarXiv 2022Prefix-Tuning: Optimizing Continuous Prompts for GenerationACL 2021 5CogView2: Faster and Better Text-to-Image Generation via Hierarchical TransformersarXiv 2022

Back to Papers