0

All papers

Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.

SinGAN: Learning a Generative Model from a Single Natural Imagesingan-learning-a-generative-model-from-a-1ReAct: Synergizing Reasoning and Acting in Language ModelsarXiv 2022Segmentation Transformer: Object-Contextual Representations for Semantic SegmentationECCV 2020 8Neural Operator: Learning Maps Between Function SpacesarXiv 2021Vision Models Are More Robust And Fair When Pretrained On Uncurated Images Without Supervisionvision-models-are-more-robust-and-fair-when-1DepGraph: Towards Any Structural PruningCVPR 2023 1Glow: Generative Flow with Invertible 1x1 Convolutionsglow-generative-flow-with-invertible-1x1-1OtterHD: A High-Resolution Multi-modality ModelarXiv 2023DocETL: Agentic Query Rewriting and Evaluation for Complex Document ProcessingarXiv 2024ResNeSt: Split-Attention NetworksarXiv 2020Diffusion Models: A Comprehensive Survey of Methods and ApplicationsarXiv 2022TextGrad: Automatic "Differentiation" via TextarXiv 2024Hierarchical Spatio-temporal Decoupling for Text-to-Video GenerationCVPR 2024 1Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation ModelarXiv 2025Visualizing the Loss Landscape of Neural Netsvisualizing-the-loss-landscape-of-neural-nets-1GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken ChatbotarXiv 2024Segment and Track AnythingarXiv 2023Accelerated Hierarchical Density ClusteringarXiv 2017Moonshine: Speech Recognition for Live Transcription and Voice CommandsarXiv 2024Bilateral Reference for High-Resolution Dichotomous Image SegmentationarXiv 2024AgentBench: Evaluating LLMs as AgentsarXiv 2023Video Understanding with Large Language Models: A SurveyarXiv 2023Onesweep: A Faster Least Significant Digit Radix Sort for GPUsarXiv 2022Phoenix: Democratizing ChatGPT across LanguagesarXiv 2023ABSApp: A Portable Weakly-Supervised Aspect-Based Sentiment Extraction Systemabsapp-a-portable-weakly-supervised-aspect-1TacSL: A Library for Visuotactile Sensor Simulation and LearningarXiv 2024Dynamic Routing Between Capsulesdynamic-routing-between-capsules-1Visual Reinforcement Learning with Imagined Goalsvisual-reinforcement-learning-with-imagined-1NUWA-Infinity: Autoregressive over Autoregressive Generation for Infinite Visual SynthesisarXiv 2022Tool Learning with Foundation ModelsarXiv 2023Efficiently Modeling Long Sequences with Structured State Spacesefficiently-modeling-long-sequences-withSemantic-SAM: Segment and Recognize Anything at Any GranularityarXiv 2023Chain-of-Thought Hub: A Continuous Effort to Measure Large Language Models' Reasoning PerformancearXiv 2023LoFTR: Detector-Free Local Feature Matching with TransformersCVPR 2021 1Mastering Atari Games with Limited DataNeurIPS 2021 12LongCoder: A Long-Range Pre-trained Language Model for Code CompletionarXiv 2023Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networksmodel-agnostic-meta-learning-for-fast-1LongLoRA: Efficient Fine-tuning of Long-Context Large Language ModelsarXiv 2023YOLOv12: Attention-Centric Real-Time Object DetectorsarXiv 2025INT2.1: Towards Fine-Tunable Quantized Large Language Models with Error Correction through Low-Rank AdaptationarXiv 2023InfiniteYou: Flexible Photo Recrafting While Preserving Your IdentityICCV 2025Benchmarking Graph Neural NetworksarXiv 2020PipeFusion: Patch-level Pipeline Parallelism for Diffusion Transformers InferencearXiv 2024EdgeConnect: Generative Image Inpainting with Adversarial Edge LearningarXiv 2019GAN Prior Embedded Network for Blind Face Restoration in the WildCVPR 2021 1MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose GuidancearXiv 2024Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsarXiv 2023On the Variance of the Adaptive Learning Rate and BeyondICLR 2020 1mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language ModelsarXiv 2024mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality CollaborationCVPR 2024 1Highly Accurate Dichotomous Image SegmentationarXiv 2022VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech InteractionarXiv 2025EfficientSAM: Leveraged Masked Image Pretraining for Efficient Segment AnythingCVPR 2024 1Contrastive Learning for Unpaired Image-to-Image TranslationarXiv 2020Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement LearningarXiv 2025Learning an Animatable Detailed 3D Face Model from In-The-Wild ImagesarXiv 2020Improved Training of Wasserstein GANsimproved-training-of-wasserstein-gans-1Deep Flow-Guided Video Inpaintingdeep-flow-guided-video-inpainting-1Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingNeurIPS 2021 12V-Express: Conditional Dropout for Progressive Training of Portrait Video GenerationarXiv 2024TotalSegmentator MRI: Robust Sequence-independent Segmentation of Multiple Anatomic Structures in MRIarXiv 2024A Time Series is Worth 64 Words: Long-term Forecasting with TransformersarXiv 2022Restormer: Efficient Transformer for High-Resolution Image RestorationCVPR 2022 1The Natural Language Decathlon: Multitask Learning as Question Answeringthe-natural-language-decathlon-multitask-1The Future of AI: Exploring the Potential of Large Concept ModelsarXiv 2025Supervised Learning of Universal Sentence Representations from Natural Language Inference Datasupervised-learning-of-universal-sentence-1MaskGAN: Towards Diverse and Interactive Facial Image Manipulationmaskgan-towards-diverse-and-interactive-1mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document UnderstandingarXiv 2024Neural Machine Translation of Rare Words with Subword Unitsneural-machine-translation-of-rare-words-with-1SMART: Robust and Efficient Fine-Tuning for Pre-trained Natural Language Models through Principled Regularized Optimizationsmart-robust-and-efficient-fine-tuning-for-1Tracking Everything Everywhere All at OnceICCV 2023 1EasyAnimate: A High-Performance Long Video Generation Method based on Transformer ArchitecturearXiv 2024DiffusionDet: Diffusion Model for Object DetectionICCV 2023 1Analyzing Learned Molecular Representations for Property PredictionarXiv 2019Unsupervised Data Augmentation for Consistency Trainingunsupervised-data-augmentationWhen LLMs step into the 3D World: A Survey and Meta-Analysis of 3D Tasks via Multi-modal Large Language ModelsarXiv 2024BioBERT: a pre-trained biomedical language representation model for biomedical text miningarXiv 2019Revisiting and Advancing Chinese Natural Language Understanding with Accelerated Heterogeneous Knowledge Pre-trainingarXiv 2022Optimal Linear Subspace Search: Learning to Construct Fast and High-Quality Schedulers for Diffusion ModelsarXiv 2023A Survey on Large Language Models for RecommendationarXiv 2023The Optimal BERT Surgeon: Scalable and Accurate Second-Order Pruning for Large Language ModelsarXiv 2022Med3D: Transfer Learning for 3D Medical Image AnalysisarXiv 2019Retrieval-Augmented Generation for Large Language Models: A SurveyarXiv 2023Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption AugmentationarXiv 2022SentEval: An Evaluation Toolkit for Universal Sentence Representationssenteval-an-evaluation-toolkit-for-universal-1Methods for Detoxification of Texts for the Russian LanguagearXiv 2021Center-based 3D Object Detection and TrackingCVPR 2021 1P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and TasksarXiv 2021Pre-trained Models for Natural Language Processing: A SurveyarXiv 2020Metric3Dv2: A Versatile Monocular Geometric Foundation Model for Zero-shot Metric Depth and Surface Normal Estimationmetric3d-v2-a-versatile-monocular-geometricScaled-YOLOv4: Scaling Cross Stage Partial NetworkCVPR 2021 1BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval ModelsarXiv 2021Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt InjectionarXiv 2023One Embedder, Any Task: Instruction-Finetuned Text EmbeddingsarXiv 2022Unsupervised Learning of Depth and Ego-Motion from Videounsupervised-learning-of-depth-and-ego-motion-4YOLOR-Based Multi-Task LearningarXiv 2023Addressing Function Approximation Error in Actor-Critic Methodsaddressing-function-approximation-error-in-1Graph Mixup with Soft AlignmentsarXiv 2023Detecting Twenty-thousand Classes using Image-level SupervisionarXiv 2022Zero123++: a Single Image to Consistent Multi-view Diffusion Base ModelarXiv 2023Regularizing and Optimizing LSTM Language Modelsregularizing-and-optimizing-lstm-language-1DB-GPT-Hub: Towards Open Benchmarking Text-to-SQL Empowered by Large Language ModelsarXiv 2024Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsarXiv 2024Detection and Tracking Meet Drones ChallengearXiv 2020Multi-Concept Customization of Text-to-Image DiffusionCVPR 2023 1SplaTAM: Splat, Track & Map 3D Gaussians for Dense RGB-D SLAMarXiv 2023Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image PyramidarXiv 2024PixelCNN++: Improving the PixelCNN with Discretized Logistic Mixture Likelihood and Other ModificationsarXiv 2017iTransformer: Inverted Transformers Are Effective for Time Series ForecastingarXiv 2023Implicit Neural Representations with Periodic Activation Functionsimplicit-neural-representations-with-periodic-1Aggregated Residual Transformations for Deep Neural Networksaggregated-residual-transformations-for-deep-1A Large-Scale Chinese Short-Text Conversation DatasetarXiv 2020Learning to Embed Time Series Patches IndependentlyarXiv 2023Holistically-Nested Edge Detectionholistically-nested-edge-detection-1Vary: Scaling up the Vision Vocabulary for Large Vision-Language ModelsarXiv 2023S-LoRA: Serving Thousands of Concurrent LoRA AdaptersarXiv 2023Make-It-3D: High-Fidelity 3D Creation from A Single Image with Diffusion PriorICCV 2023 1ViTPose++: Vision Transformer for Generic Body Pose EstimationarXiv 2022PVT v2: Improved Baselines with Pyramid Vision TransformerarXiv 2021OminiControl: Minimal and Universal Control for Diffusion TransformerICCV 2025BERTScore: Evaluating Text Generation with BERTICLR 2020 1The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer UsearXiv 2024Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMsarXiv 2024Sequential Modeling Enables Scalable Learning for Large Vision ModelsCVPR 2024 1A Deep Reinforcement Learning Framework for the Financial Portfolio Management ProblemarXiv 2017Text Classification Algorithms: A SurveyarXiv 2019Training Generative Adversarial Networks with Limited DataNeurIPS 2020 12Show-o: One Single Transformer to Unify Multimodal Understanding and GenerationarXiv 2024Open X-Embodiment: Robotic Learning Datasets and RT-X ModelsarXiv 2023NGBoost: Natural Gradient Boosting for Probabilistic PredictionICML 2020 1An LLM Compiler for Parallel Function CallingarXiv 2023Dynamic Graph CNN for Learning on Point CloudsarXiv 2018CogView: Mastering Text-to-Image Generation via TransformersNeurIPS 2021 12Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learningseq2sql-generating-structured-queries-from-1LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMsarXiv 2024Score-Based Generative Modeling through Stochastic Differential Equationsscore-based-generative-modeling-throughCaption Anything: Interactive Image Description with Diverse Multimodal ControlsarXiv 2023A Survey on Knowledge Graphs: Representation, Acquisition and ApplicationsarXiv 2020Emergent Tool Use From Multi-Agent AutocurriculaICLR 2020 1Learning to Fly -- a Gym Environment with PyBullet Physics for Reinforcement Learning of Multi-agent Quadcopter ControlarXiv 2021Generative Multimodal Models are In-Context LearnersCVPR 2024 1Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionICML 2020 1NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view ReconstructionNeurIPS 2021 12Variational Graph Auto-EncodersarXiv 2016TokenFlow: Consistent Diffusion Features for Consistent Video EditingarXiv 2023Any-to-Any Generation via Composable DiffusionNeurIPS 2023 11Unifying Vision, Text, and Layout for Universal Document ProcessingCVPR 2023 1MeshCNN: A Network with an EdgearXiv 2018Deep Interest Network for Click-Through Rate PredictionarXiv 2017TFB: Towards Comprehensive and Fair Benchmarking of Time Series Forecasting MethodsarXiv 2024BrushNet: A Plug-and-Play Image Inpainting Model with Decomposed Dual-Branch DiffusionarXiv 2024Pastiche Master: Exemplar-Based High-Resolution Portrait Style TransferCVPR 2022 1OneFormer: One Transformer to Rule Universal Image SegmentationCVPR 2023 1ICON: Implicit Clothed humans Obtained from NormalsCVPR 2022 1CodeNet: A Large-Scale AI for Code Dataset for Learning a Diversity of Coding TasksarXiv 2021ESC-Eval: Evaluating Emotion Support Conversations in Large Language ModelsarXiv 2024SymbolicAI: A framework for logic-based approaches combining generative models and solversarXiv 2024Meta-Transformer: A Unified Framework for Multimodal LearningarXiv 2023OpenMoE: An Early Effort on Open Mixture-of-Experts Language ModelsarXiv 2024Octo: An Open-Source Generalist Robot PolicyarXiv 2024Personalize Segment Anything Model with One ShotarXiv 2023ST-MoE: Designing Stable and Transferable Sparse Expert ModelsarXiv 2022It's Not Just Size That Matters: Small Language Models Are Also Few-Shot LearnersNAACL 2021 4StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from TextCVPR 2025 1ControlNeXt: Powerful and Efficient Control for Image and Video GenerationarXiv 2024Video Swin TransformerCVPR 2022 1d3rlpy: An Offline Deep Reinforcement Learning LibraryarXiv 2021ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generationimagereward-learning-and-evaluating-humanWebGLM: Towards An Efficient Web-Enhanced Question Answering System with Human PreferencesarXiv 2023AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion EncodingarXiv 2024SyncTalk: The Devil is in the Synchronization for Talking Head SynthesisCVPR 2024 1LLM4Drive: A Survey of Large Language Models for Autonomous DrivingarXiv 2023ShowUI: One Vision-Language-Action Model for GUI Visual AgentCVPR 2025 1Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by TencentarXiv 2024Few-Shot Unsupervised Image-to-Image Translationfew-shot-unsupervised-image-to-image-1Semantic Operators: A Declarative Model for Rich, AI-based Data ProcessingarXiv 2024CPM: A Large-scale Generative Chinese Pre-trained Language ModelarXiv 2020Network Pruning via Transformable Architecture Searchnetwork-pruning-via-transformable-1Safe RLHF: Safe Reinforcement Learning from Human FeedbackarXiv 2023RDT-1B: a Diffusion Foundation Model for Bimanual ManipulationarXiv 2024CatVTON: Concatenation Is All You Need for Virtual Try-On with Diffusion ModelsarXiv 2024CascadeTabNet: An approach for end to end table detection and structure recognition from image-based documentsarXiv 2020Arbitrary Style Transfer in Real-time with Adaptive Instance Normalizationarbitrary-style-transfer-in-real-time-with-1ReFT: Representation Finetuning for Language ModelsarXiv 2024Stretching Each Dollar: Diffusion Training from Scratch on a Micro-BudgetCVPR 2025 1Neural Fields in Robotics: A SurveyarXiv 2024WikiChat: Stopping the Hallucination of Large Language Model Chatbots by Few-Shot Grounding on WikipediaarXiv 2023All4One: Symbiotic Neighbour Contrastive Learning via Self-Attention and Redundancy ReductionICCV 2023 1YOLOv9: Learning What You Want to Learn Using Programmable Gradient InformationarXiv 2024Magic Clothing: Controllable Garment-Driven Image SynthesisarXiv 2024Stable Virtual Camera: Generative View Synthesis with Diffusion ModelsICCV 2025CLUENER2020: Fine-grained Named Entity Recognition Dataset and Benchmark for ChinesearXiv 2020LucidDreamer: Domain-free Generation of 3D Gaussian Splatting ScenesarXiv 2023cuRobo: Parallelized Collision-Free Minimum-Jerk Robot Motion GenerationarXiv 2023MultiModal-GPT: A Vision and Language Model for Dialogue with HumansarXiv 2023Lag-Llama: Towards Foundation Models for Probabilistic Time Series ForecastingarXiv 2023Data-Copilot: Bridging Billions of Data and Humans with Autonomous WorkflowarXiv 2023MegaBlocks: Efficient Sparse Training with Mixture-of-ExpertsarXiv 2022Fine-tune BERT for Extractive Summarizationfine-tune-bert-for-extractive-summarization-1Humans in 4D: Reconstructing and Tracking Humans with TransformersICCV 2023 1

Back to Papers