All papers
Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.
Do LLMs Possess a Personality? Making the MBTI Test an Amazing Evaluation for Large Language ModelsarXiv 2023PromptSource: An Integrated Development Environment and Repository for Natural Language PromptsACL 2022 5CLoSD: Closing the Loop between Simulation and Diffusion for multi-task character controlarXiv 2024Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific HeadsarXiv 2024Mini-Omni: Language Models Can Hear, Talk While Thinking in StreamingarXiv 2024Big Self-Supervised Models are Strong Semi-Supervised LearnersNeurIPS 2020 12BIRB: A Generalization Benchmark for Information Retrieval in BioacousticsarXiv 2023CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use AgentsarXiv 2026The Entropy Mechanism of Reinforcement Learning for Reasoning Language ModelsarXiv 2025NeuralProphet: Explainable Forecasting at Scaleneuralprophet-explainable-forecasting-at-1DetZero: Rethinking Offboard 3D Object Detection with Long-term Sequential Point CloudsICCV 2023 1PIXART-δ: Fast and Controllable Image Generation with Latent Consistency ModelsarXiv 2024When Attention Sink Emerges in Language Models: An Empirical ViewarXiv 2024Generative AI for Medical Imaging: extending the MONAI FrameworkarXiv 2023OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality InteractionarXiv 2025SpokenWOZ: A Large-Scale Speech-Text Benchmark for Spoken Task-Oriented Dialogue Agentsspokenwoz-a-large-scale-speech-text-benchmarkPreference Ranking Optimization for Human AlignmentarXiv 2023PaCE: Unified Multi-modal Dialogue Pre-training with Progressive and Compositional ExpertsarXiv 2023MoE-LLaVA: Mixture of Experts for Large Vision-Language ModelsarXiv 2024DreamForge: Motion-Aware Autoregressive Video Generation for Multi-View Driving ScenesarXiv 2024StudioGAN: A Taxonomy and Benchmark of GANs for Image SynthesisarXiv 2022OS-ATLAS: A Foundation Action Model for Generalist GUI AgentsarXiv 2024VideoMAE V2: Scaling Video Masked Autoencoders with Dual MaskingCVPR 2023 1TabDDPM: Modelling Tabular Data with Diffusion ModelsarXiv 2022Phantom: Subject-consistent video generation via cross-modal alignmentICCV 2025DiffPose: Toward More Reliable 3D Pose EstimationCVPR 2023 1DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative ModelsarXiv 2022Jailbreaking Black Box Large Language Models in Twenty QueriesarXiv 2023CaRL: Learning Scalable Planning Policies with Simple RewardsarXiv 2025Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environmentsmulti-agent-actor-critic-for-mixed-1LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-trainingarXiv 2024HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration TestingarXiv 2024Improved Denoising Diffusion Probabilistic Modelsimproved-denoising-diffusion-probabilisticLarge Language Models Meet Knowledge Graphs for Question Answering: Synthesis and OpportunitiesarXiv 2025Alita: Generalist Agent Enabling Scalable Agentic Reasoning with Minimal Predefinition and Maximal Self-EvolutionarXiv 2025ImgEdit: A Unified Image Editing Dataset and BenchmarkarXiv 2025VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D ReconstructionarXiv 2025DiffSplat: Repurposing Image Diffusion Models for Scalable Gaussian Splat GenerationarXiv 2025Tensor Product Attention Is All You NeedarXiv 2025EMAGE: Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture ModelingCVPR 2024 1MedViT: A Robust Vision Transformer for Generalized Medical Image ClassificationarXiv 2023Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base ModelarXiv 2025NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive SecurityarXiv 2024One-Step Image Translation with Text-to-Image ModelsarXiv 2024Alias-Free Generative Adversarial NetworksNeurIPS 2021 12DoRA: Weight-Decomposed Low-Rank AdaptationarXiv 2024NVILA: Efficient Frontier Visual Language ModelsCVPR 2025 1Revisiting VerilogEval: A Year of Improvements in Large-Language Models for Hardware Code GenerationarXiv 2024TAPIP3D: Tracking Any Point in Persistent 3D GeometryarXiv 2025Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware PermutationarXiv 2025ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and UnderstandingarXiv 2025M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language ModelsarXiv 2024COAT: Compressing Optimizer states and Activation for Memory-Efficient FP8 TrainingarXiv 2024BigVGAN: A Universal Neural Vocoder with Large-Scale TrainingarXiv 2022video-SALMONN 2: Captioning-Enhanced Audio-Visual Large Language ModelsarXiv 2025DISC-FinLLM: A Chinese Financial Large Language Model based on Multiple Experts Fine-tuningarXiv 2023AutoSurvey: Large Language Models Can Automatically Write SurveysarXiv 2024Olmo 3arXiv 2025Learning Flow Fields in Attention for Controllable Person Image GenerationCVPR 2025 1Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image SynthesisCVPR 2025 1CAD-Coder: An Open-Source Vision-Language Model for Computer-Aided Design Code GenerationarXiv 2025ChatTS: Aligning Time Series with LLMs via Synthetic Data for Enhanced Understanding and ReasoningarXiv 2024YOLO-MS: Rethinking Multi-Scale Representation Learning for Real-time Object DetectionarXiv 2023TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language ModelingarXiv 2025AlphaPose: Whole-Body Regional Multi-Person Pose Estimation and Tracking in Real-TimearXiv 2022UNet++: A Nested U-Net Architecture for Medical Image SegmentationarXiv 2018RadGPT: Constructing 3D Image-Text Tumor DatasetsICCV 2025PyGAD: An Intuitive Genetic Algorithm Python LibraryarXiv 2021Structured prompt interrogation and recursive extraction of semantics (SPIRES): A method for populating knowledge bases using zero-shot learningarXiv 2023DataComp: In search of the next generation of multimodal datasetsNeurIPS 2023 11AdaFace: Quality Adaptive Margin for Face RecognitionCVPR 2022 1TSD-SR: One-Step Diffusion with Target Score Distillation for Real-World Image Super-ResolutionCVPR 2025 1ChatGPT for Robotics: Design Principles and Model AbilitiesarXiv 2023PIKE-RAG: sPecIalized KnowledgE and Rationale Augmented GenerationarXiv 2025Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter TransferarXiv 2022Magma: A Foundation Model for Multimodal AI AgentsCVPR 2025 1Accelerating Goal-Conditioned RL Algorithms and ResearcharXiv 2024Z-Code++: A Pre-trained Language Model Optimized for Abstractive SummarizationarXiv 2022CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and GenerationarXiv 2021Language Agents as Optimizable GraphsarXiv 2024Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning ModelsarXiv 2025Mixed Dimension Embeddings with Application to Memory-Efficient Recommendation SystemsarXiv 2019Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and SoundarXiv 2025Eureka: Human-Level Reward Design via Coding Large Language ModelsarXiv 2023PyTorch Tabular: A Framework for Deep Learning with Tabular DataarXiv 2021Accelerating Data Processing and Benchmarking of AI Models for PathologyarXiv 2025T2M-GPT: Generating Human Motion from Textual Descriptions with Discrete RepresentationsarXiv 2023Building reliable sim driving agents by scaling self-playarXiv 2025VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context ControlarXiv 2025Simple Online and Realtime TrackingarXiv 2016Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts ModelsarXiv 2024Stop Overthinking: A Survey on Efficient Reasoning for Large Language ModelsarXiv 2025MVSplat360: Feed-Forward 360 Scene Synthesis from Sparse ViewsarXiv 2024Pangu-Weather: A 3D High-Resolution Model for Fast and Accurate Global Weather ForecastarXiv 2022MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMsarXiv 2025Large-Scale 3D Medical Image Pre-training with Geometric Context PriorsarXiv 2024TinyLLaVA: A Framework of Small-scale Large Multimodal ModelsarXiv 2024AgentOhana: Design Unified Data and Training Pipeline for Effective Agent LearningarXiv 2024RepVGG: Making VGG-style ConvNets Great AgainCVPR 2021 1ConTextTab: A Semantics-Aware Tabular In-Context LearnerarXiv 2025ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D DataarXiv 2021GameFactory: Creating New Games with Generative Interactive VideosICCV 2025Representing Long Volumetric Video with Temporal Gaussian HierarchyarXiv 2024A Simple and Effective Pruning Approach for Large Language ModelsarXiv 2023MINIMA: Modality Invariant Image MatchingCVPR 2025 1On-Policy RL with Optimal Reward BaselinearXiv 2025Large Language Model for Science: A Study on P vs. NParXiv 2023Inference with Reference: Lossless Acceleration of Large Language ModelsarXiv 2023UniDepthV2: Universal Monocular Metric Depth Estimation Made SimplerarXiv 2025AndroidEnv: A Reinforcement Learning Platform for AndroidarXiv 2021StarCraft II: A New Challenge for Reinforcement LearningarXiv 2017Grounded Language Learning Fast and SlowICLR 2021 1GraphCast: Learning skillful medium-range global weather forecastingarXiv 2022Discrete Diffusion Modeling by Estimating the Ratios of the Data DistributionarXiv 2023MagicQuill: An Intelligent Interactive Image Editing SystemCVPR 2025 1Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented GenerationarXiv 2025Arctic-Text2SQL-R1: Simple Rewards, Strong Reasoning in Text-to-SQLarXiv 2025Frequency-aware Feature Fusion for Dense Image PredictionarXiv 2024Efficient Multimodal Large Language Models: A SurveyarXiv 2024Refusal in Language Models Is Mediated by a Single DirectionarXiv 2024LibCity: A Unified Library Towards Efficient and Comprehensive Urban Spatial-Temporal Predictiontowards-efficient-and-comprehensive-urbanExpeL: LLM Agents Are Experiential LearnersarXiv 2023Pantograph: A Machine-to-Machine Interaction Interface for Advanced Theorem Proving, High Level Reasoning, and Data Extraction in Lean 4arXiv 2024VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video UnderstandingarXiv 2025VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMsarXiv 2024Disentangling Writer and Character Styles for Handwriting GenerationCVPR 2023 1Automatic Chain of Thought Prompting in Large Language ModelsarXiv 2022DeblurGAN-v2: Deblurring (Orders-of-Magnitude) Faster and Betterdeblurgan-v2-deblurring-orders-of-magnitude-1Autoregressive Image Generation using Residual QuantizationCVPR 2022 1A Survey on Graph Neural Networks for Time Series: Forecasting, Classification, Imputation, and Anomaly DetectionarXiv 2023Beyond a Gaussian Denoiser: Residual Learning of Deep CNN for Image DenoisingarXiv 2016SCSegamba: Lightweight Structure-Aware Vision Mamba for Crack Segmentation in StructuresCVPR 2025 1Language Model InversionarXiv 2023Dynamic 3D Gaussians: Tracking by Persistent Dynamic View SynthesisarXiv 2023Lumina-Image 2.0: A Unified and Efficient Image Generative FrameworkICCV 2025Unsupervised Night Image Enhancement: When Layer Decomposition Meets Light-Effects SuppressionarXiv 2022Perspective Fields for Single Image Camera CalibrationCVPR 2023 1FisherRF: Active View Selection and Uncertainty Quantification for Radiance Fields using Fisher InformationarXiv 2023AlphaEdit: Null-Space Constrained Knowledge Editing for Language ModelsarXiv 2024A Novel Unified Architecture for Low-Shot Counting by Detection and SegmentationarXiv 2024Investigating Tradeoffs in Real-World Video Super-ResolutionCVPR 2022 1EasyControl: Adding Efficient and Flexible Control for Diffusion TransformerICCV 2025DeepCAD: A Deep Generative Network for Computer-Aided Design ModelsICCV 2021 10A Review of Safe Reinforcement Learning: Methods, Theory and ApplicationsarXiv 2022DEA-Net: Single image dehazing based on detail-enhanced convolution and content-guided attentionarXiv 2023Grokfast: Accelerated Grokking by Amplifying Slow GradientsarXiv 2024DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal DecompositionarXiv 2025Seed-TTS: A Family of High-Quality Versatile Speech Generation ModelsarXiv 2024SALMONN: Towards Generic Hearing Abilities for Large Language ModelsarXiv 2023LayoutLLM: Layout Instruction Tuning with Large Language Models for Document UnderstandingCVPR 2024 1Multi-Granularity Prediction for Scene Text RecognitionarXiv 2022OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language ModelsarXiv 2025DepthMaster: Taming Diffusion Models for Monocular Depth EstimationarXiv 2025UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and GenerationarXiv 2025CausalPFN: Amortized Causal Effect Estimation via In-Context LearningarXiv 2025ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning AgentsarXiv 2025SDXS: Real-Time One-Step Latent Diffusion Models with Image ConditionsarXiv 2024A Survey on LLM-as-a-JudgearXiv 2024TerraTorch: The Geospatial Foundation Models ToolkitarXiv 2025Torchhd: An Open Source Python Library to Support Research on Hyperdimensional Computing and Vector Symbolic ArchitecturesarXiv 2022Bridging Language and Items for Retrieval and RecommendationarXiv 2024Senna: Bridging Large Vision-Language Models and End-to-End Autonomous DrivingarXiv 2024MDocAgent: A Multi-Modal Multi-Agent Framework for Document UnderstandingarXiv 2025Foundation Models in Autonomous Driving: A Survey on Scenario Generation and Scenario AnalysisarXiv 2025SNAC: Multi-Scale Neural Audio CodecarXiv 2024LLM-Pruner: On the Structural Pruning of Large Language Modelsllm-pruner-on-the-structural-pruning-of-largeSegment Anything Model for Road Network Graph ExtractionarXiv 2024ReEvo: Large Language Models as Hyper-Heuristics with Reflective EvolutionarXiv 2024Audio-Reasoner: Improving Reasoning Capability in Large Audio Language ModelsarXiv 2025Multi-agent Architecture Search via Agentic SupernetarXiv 2025IndicLLMSuite: A Blueprint for Creating Pre-training and Fine-Tuning Datasets for Indian LanguagesarXiv 2024ACE2: Accurately learning subseasonal to decadal atmospheric variability and forced responsesarXiv 2024Time-VLM: Exploring Multimodal Vision-Language Models for Augmented Time Series ForecastingarXiv 2025Structured3D: A Large Photo-realistic Dataset for Structured 3D ModelingECCV 2020 8Unified Video Action ModelarXiv 2025DQ-DETR: DETR with Dynamic Query for Tiny Object DetectionarXiv 2024VLFM: Vision-Language Frontier Maps for Zero-Shot Semantic NavigationarXiv 2023TransMLA: Multi-Head Latent Attention Is All You NeedarXiv 2025Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability CurvaturearXiv 2023Diffusion Models and Semi-Supervised Learners Benefit Mutually with Few Labelsdiffusion-models-and-semi-supervised-learnersSSLRec: A Self-Supervised Learning Framework for RecommendationarXiv 2023Tracking Anything with Decoupled Video SegmentationICCV 2023 1Coloring the Noise: Adversarial Sobolev Alignment for Faithful Image Super ResolutionarXiv 2026ReSearch: Learning to Reason with Search for LLMs via Reinforcement LearningarXiv 2025Emu3: Next-Token Prediction is All You NeedarXiv 2024Masked Autoencoders Are Effective Tokenizers for Diffusion ModelsarXiv 2025Hidden Biases of End-to-End Driving DatasetsarXiv 20243D Reconstruction with Spatial MemoryarXiv 2024Turning Trash into Treasure: Accelerating Inference of Large Language Models with Token RecyclingarXiv 2024DynaMem: Online Dynamic Spatio-Semantic Memory for Open World Mobile ManipulationarXiv 2024BiomedCoOp: Learning to Prompt for Biomedical Vision-Language ModelsCVPR 2025 1MEFLUT: Unsupervised 1D Lookup Tables for Multi-exposure Image FusionICCV 2023 1CameraCtrl: Enabling Camera Control for Text-to-Video GenerationarXiv 2024LION: Linear Group RNN for 3D Object Detection in Point CloudsarXiv 2024Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language ModelsarXiv 2025MonST3R: A Simple Approach for Estimating Geometry in the Presence of MotionarXiv 2024GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLMarXiv 2024A Comprehensive Survey on Composed Image RetrievalarXiv 2025VoiceFixer: A Unified Framework for High-Fidelity Speech RestorationarXiv 2022From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent SkillsarXiv 2026