0

All papers

Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.

QLoRA: Efficient Finetuning of Quantized LLMsNeurIPS 2023 11Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actorsoft-actor-critic-off-policy-maximum-entropy-1MotionCLIP: Exposing Human Motion Generation to CLIP SpacearXiv 2022SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model InferencearXiv 2024CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip RetrievalarXiv 2021Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex CapabilitiesarXiv 2024Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model SizesarXiv 2023Learning Compact Metrics for MTEMNLP 2021 11Mixtures of Experts Unlock Parameter Scaling for Deep RLarXiv 2024The Diffusion DualityarXiv 2025Sundial: A Family of Highly Capable Time Series Foundation ModelsarXiv 2025NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language ModelsarXiv 2023Deep Confident Steps to New Pockets: Strategies for Docking GeneralizationarXiv 2024Halu-J: Critique-Based Hallucination JudgearXiv 2024OpenProteinSet: Training data for structural biology at scaleNeurIPS 2023 11DiffPortrait3D: Controllable Diffusion for Zero-Shot Portrait View SynthesisCVPR 2024 1CAD-Recode: Reverse Engineering CAD Code from Point CloudsarXiv 2024LaViDa: A Large Diffusion Language Model for Multimodal UnderstandingarXiv 2025Paint by Example: Exemplar-based Image Editing with Diffusion ModelsCVPR 2023 1DOVE: Efficient One-Step Diffusion Model for Real-World Video Super-ResolutionarXiv 2025A Pragmatic VLA Foundation ModelarXiv 2026Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons LearnedarXiv 2022Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and InferencearXiv 2024Representation Engineering: A Top-Down Approach to AI TransparencyarXiv 2023Learning Synergies between Pushing and Grasping with Self-supervised Deep Reinforcement LearningarXiv 2018RePaint: Inpainting using Denoising Diffusion Probabilistic ModelsCVPR 2022 1OctoTools: An Agentic Framework with Extensible Tools for Complex ReasoningarXiv 2025Aria Everyday Activities DatasetarXiv 2024Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at ScalearXiv 2025DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga GenerationCVPR 2025 1Focal Loss for Dense Object Detectionfocal-loss-for-dense-object-detection-1DeiT III: Revenge of the ViTarXiv 2022All-atom Diffusion Transformers: Unified generative modelling of molecules and materialsarXiv 2025Large Multi-modal Models Can Interpret Features in Large Multi-modal ModelsICCV 2025From Imitation to Refinement -- Residual RL for Precise AssemblyarXiv 2024FiLM: Visual Reasoning with a General Conditioning LayerarXiv 2017CATANet: Efficient Content-Aware Token Aggregation for Lightweight Image Super-ResolutionCVPR 2025 1r/Fakeddit: A New Multimodal Benchmark Dataset for Fine-grained Fake News DetectionarXiv 2019Datasheet for the PilearXiv 2022Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiTarXiv 2024Human-like Episodic Memory for Infinite Context LLMsarXiv 2024Word-level Deep Sign Language Recognition from Video: A New Large-scale Dataset and Methods ComparisonarXiv 2019Targeted Attack Improves Protection against Unauthorized Diffusion CustomizationarXiv 2023FLOAT: Generative Motion Latent Flow Matching for Audio-driven Talking PortraitICCV 2025PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image GenerationarXiv 2024Diffusion Posterior Sampling for General Noisy Inverse ProblemsarXiv 2022Respecting causality is all you need for training physics-informed neural networksarXiv 2022EdgeFace: Efficient Face Recognition Model for Edge DevicesarXiv 2023Transformers in Time Series: A SurveyarXiv 2022Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI AgentsarXiv 2024FinMem: A Performance-Enhanced LLM Trading Agent with Layered Memory and Character DesignarXiv 2023An Illusion of Progress? Assessing the Current State of Web AgentsarXiv 2025TALENT: A Tabular Analytics and Learning ToolboxarXiv 2024AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMsarXiv 2024ULIP-2: Towards Scalable Multimodal Pre-training for 3D UnderstandingCVPR 2024 1ProGen2: Exploring the Boundaries of Protein Language ModelsarXiv 2022Text2CAD: Generating Sequential CAD Models from Beginner-to-Expert Level Text PromptsarXiv 2024Lifelong Learning of Large Language Model based Agents: A RoadmaparXiv 2025SimpleNet: A Simple Network for Image Anomaly Detection and LocalizationCVPR 2023 1Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsarXiv 2025CodeT5+: Open Code Large Language Models for Code Understanding and GenerationarXiv 2023Recipe for a General, Powerful, Scalable Graph TransformerarXiv 2022Marsellus: A Heterogeneous RISC-V AI-IoT End-Node SoC with 2-to-8b DNN Acceleration and 30%-Boost Adaptive Body BiasingarXiv 2023TikZero: Zero-Shot Text-Guided Graphics Program SynthesisICCV 2025Recent Advances in Speech Language Models: A SurveyarXiv 2024CodeGen2: Lessons for Training LLMs on Programming and Natural LanguagesarXiv 2023Soft Actor-Critic for Discrete Action SettingsarXiv 2019Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual EditingarXiv 2025MMAU: A Massive Multi-Task Audio Understanding and Reasoning BenchmarkarXiv 2024Transformer-Squared: Self-adaptive LLMsarXiv 2025An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language ModelsarXiv 2024PyTorch Frame: A Modular Framework for Multi-Modal Tabular LearningarXiv 2024An Electrocardiogram Foundation Model Built on over 10 Million Recordings with External Evaluation across Multiple DomainsarXiv 2024ImageBind-LLM: Multi-modality Instruction TuningarXiv 2023GAN Lab: Understanding Complex Deep Generative Models using Interactive Visual ExperimentationarXiv 2018HELMET: How to Evaluate Long-Context Language Models Effectively and ThoroughlyarXiv 2024Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic DatasetsarXiv 2025InternImage: Exploring Large-Scale Vision Foundation Models with Deformable ConvolutionsCVPR 2023 1MARS: An Instance-aware, Modular and Realistic Simulator for Autonomous DrivingarXiv 2023OpenLane-V2: A Topology Reasoning Benchmark for Unified 3D HD Mappingopenlane-v2-a-topology-reasoning-benchmarkSAM3D: Segment Anything in 3D ScenesarXiv 2023Shap-E: Generating Conditional 3D Implicit FunctionsarXiv 2023Sample-Efficient Alignment for LLMsarXiv 202470% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length FloatarXiv 2025MDTv2: Masked Diffusion Transformer is a Strong Image SynthesizerICCV 2023 1Graph Attention Networksgraph-attention-networks-1Grokking: Generalization Beyond Overfitting on Small Algorithmic DatasetsarXiv 2022GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsarXiv 2021RAPTOR: Recursive Abstractive Processing for Tree-Organized RetrievalarXiv 2024Detecting and Grounding Multi-Modal Media ManipulationCVPR 2023 1Cube: A Roblox View of 3D IntelligencearXiv 2025Graph-constrained Reasoning: Faithful Reasoning on Knowledge Graphs with Large Language ModelsarXiv 2024BitTensor: A Peer-to-Peer Intelligence MarketarXiv 2020Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement LearningarXiv 2019Boosting Generative Image Modeling via Joint Image-Feature SynthesisarXiv 2025DreamCraft3D: Hierarchical 3D Generation with Bootstrapped Diffusion PriorarXiv 2023DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language ModelarXiv 2024RepoGraph: Enhancing AI Software Engineering with Repository-level Code GrapharXiv 2024Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyarXiv 2023Fine-Tuning Discrete Diffusion Models with Policy Gradient MethodsarXiv 2025The Unreasonable Effectiveness of Deep Features as a Perceptual Metricthe-unreasonable-effectiveness-of-deep-1Point2RBox: Combine Knowledge from Synthetic Visual Patterns for End-to-end Oriented Object Detection with Single Point SupervisionCVPR 2024 1COSMIC: COmmonSense knowledge for eMotion Identification in ConversationsFindings of the Association for Computational Linguistics 2020Posterior-Mean Rectified Flow: Towards Minimum MSE Photo-Realistic Image RestorationarXiv 2024A Style-Based Generator Architecture for Generative Adversarial Networksa-style-based-generator-architecture-for-1FB-BEV: BEV Representation from Forward-Backward View TransformationsICCV 2023 1Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal ControlarXiv 2025Cosmos World Foundation Model Platform for Physical AIarXiv 2025VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language ModelarXiv 2025Training Sparse Mixture Of Experts Text Embedding ModelsarXiv 2025MultiPL-E: A Scalable and Extensible Approach to Benchmarking Neural Code GenerationarXiv 2022Lite-Mono: A Lightweight CNN and Transformer Architecture for Self-Supervised Monocular Depth EstimationCVPR 2023 1Scalable Multi-Agent Reinforcement Learning through Intelligent Information AggregationarXiv 2022DFormerv2: Geometry Self-Attention for RGBD Semantic SegmentationCVPR 2025 1ScienceWorld: Is your Agent Smarter than a 5th Grader?arXiv 2022SeeClick: Harnessing GUI Grounding for Advanced Visual GUI AgentsarXiv 2024Dream to Control: Learning Behaviors by Latent Imaginationdream-to-control-learning-behaviors-by-latent-1RewardBench: Evaluating Reward Models for Language ModelingarXiv 2024SimSwap: An Efficient Framework For High Fidelity Face Swappingsimswap-an-efficient-framework-for-highViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionarXiv 2021Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video UnderstandingarXiv 2023Parallax-Tolerant Unsupervised Deep Image StitchingICCV 2023 1Rotary Position Embedding for Vision TransformerarXiv 2024An Efficiency Study for SPLADE ModelsarXiv 2022CroCo v2: Improved Cross-view Completion Pre-training for Stereo Matching and Optical FlowICCV 2023 1miditok: A Python package for MIDI file tokenizationarXiv 2023FinQA: A Dataset of Numerical Reasoning over Financial DataEMNLP 2021 11EAT: Self-Supervised Pre-Training with Efficient Audio TransformerarXiv 2024Dataset Cartography: Mapping and Diagnosing Datasets with Training DynamicsEMNLP 2020 11Why Do Multi-Agent LLM Systems Fail?arXiv 2025Steering Your Generalists: Improving Robotic Foundation Models via Value GuidancearXiv 2024Residual Denoising Diffusion ModelsCVPR 2024 1GlueStick: Robust Image Matching by Sticking Points and Lines TogetherICCV 2023 1In-Context LoRA for Diffusion TransformersarXiv 2024Speaker Recognition from Raw Waveform with SincNetarXiv 2018SeeSR: Towards Semantics-Aware Real-World Image Super-ResolutionCVPR 2024 1Auto-AVSR: Audio-Visual Speech Recognition with Automatic LabelsarXiv 2023Scalable High-Resolution Pixel-Space Image Synthesis with Hourglass Diffusion TransformersarXiv 2024JLT: Clean-Latent Prediction in Latent Diffusion TransformersarXiv 2026Learning A Sparse Transformer Network for Effective Image DerainingCVPR 2023 1Q-Ensemble for Offline RL: Don't Scale the Ensemble, Scale the Batch SizearXiv 2022TorchSparse: Efficient Point Cloud Inference EnginearXiv 2022DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming HeadsarXiv 2024Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timearXiv 2022BeatNet: CRNN and Particle Filtering for Online Joint Beat Downbeat and Meter TrackingarXiv 2021Reducing Energy Bloat in Large Model TrainingarXiv 2023DepthFM: Fast Monocular Depth Estimation with Flow MatchingarXiv 2024Windows Agent Arena: Evaluating Multi-Modal OS Agents at ScalearXiv 2024Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-ServearXiv 2024MAIRA-2: Grounded Radiology Report GenerationarXiv 2024Web-Bench: A LLM Code Benchmark Based on Web Standards and FrameworksarXiv 2025MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse AttentionarXiv 2025MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse AttentionarXiv 2024Microscaling Data Formats for Deep LearningarXiv 2023MarS: a Financial Market Simulation Engine Powered by Generative Foundation ModelarXiv 2024YOLOv4: Optimal Speed and Accuracy of Object DetectionarXiv 2020KBLaM: Knowledge Base augmented Language ModelarXiv 2024InterpretML: A Unified Framework for Machine Learning InterpretabilityarXiv 2019Pushing the limits of raw waveform speaker recognitionarXiv 2022What Is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model Analysiswhat-is-wrong-with-scene-text-recognition-1The INTERSPEECH 2020 Deep Noise Suppression Challenge: Datasets, Subjective Testing Framework, and Challenge ResultsarXiv 2020DeepSpeed-VisualChat: Multi-Round Multi-Image Interleave Chat via Multi-Modal Causal AttentionarXiv 2023SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video SegmentationCVPR 2025 1Agent-as-a-Judge: Evaluate Agents with AgentsarXiv 2024FMA: A Dataset For Music Analysisfma-a-dataset-for-music-analysis-1HyperSteer: Activation Steering at Scale with HypernetworksarXiv 2025Co-Evolving LLM Coder and Unit Tester via Reinforcement LearningarXiv 2025Towards An End-to-End Framework for Flow-Guided Video InpaintingCVPR 2022 1VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Trainingvideomae-masked-autoencoders-are-dataRSBuilding: Towards General Remote Sensing Image Building Extraction and Change Detection with Foundation ModelarXiv 2024Medical SAM 2: Segment medical images as video via Segment Anything Model 2arXiv 2024VFIMamba: Video Frame Interpolation with State Space ModelsarXiv 2024CO-SPY: Combining Semantic and Pixel Features to Detect Synthetic Images by AICVPR 2025 1SWE-bench Goes Live!arXiv 2025GeoChat: Grounded Large Vision-Language Model for Remote SensingCVPR 2024 1Participatory Research for Low-resourced Machine Translation: A Case Study in African LanguagesFindings of the Association for Computational Linguistics 2020EdgeCrafter: Compact ViTs for Edge Dense Prediction via Task-Specialized DistillationarXiv 2026TimeCMA: Towards LLM-Empowered Multivariate Time Series Forecasting via Cross-Modality AlignmentarXiv 2024Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseparaphrasing-evades-detectors-of-ai-generatedGarmentCodeData: A Dataset of 3D Made-to-Measure Garments With Sewing PatternsarXiv 2024Benchmarking Large Language Models in Retrieval-Augmented GenerationarXiv 2023IDOL: Instant Photorealistic 3D Human Creation from a Single ImageCVPR 2025 1FcaNet: Frequency Channel Attention NetworksICCV 2021 10PIXIU: A Large Language Model, Instruction Data and Evaluation Benchmark for FinancearXiv 2023SuperPoint: Self-Supervised Interest Point Detection and DescriptionarXiv 2017Raising the Cost of Malicious AI-Powered Image EditingarXiv 2023ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency PolicyarXiv 2025CityWalker: Learning Embodied Urban Navigation from Web-Scale VideosCVPR 2025 1ContextCite: Attributing Model Generation to ContextarXiv 2024SLIM: Sparsified Late Interaction for Multi-Vector Retrieval with Inverted IndexesarXiv 2023Poseidon: Efficient Foundation Models for PDEsarXiv 2024OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic ManipulationarXiv 2025AnomalyGPT: Detecting Industrial Anomalies Using Large Vision-Language ModelsarXiv 2023SongEval: A Benchmark Dataset for Song Aesthetics EvaluationarXiv 2025Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point ModelingCVPR 2022 1You Only Sample Once: Taming One-Step Text-to-Image Synthesis by Self-Cooperative Diffusion GANsarXiv 2024Less-to-More Generalization: Unlocking More Controllability by In-Context GenerationICCV 2025TokenFlow: Unified Image Tokenizer for Multimodal Understanding and GenerationCVPR 2025 1Discrete Diffusion in Large Language and Multimodal Models: A SurveyarXiv 2025SCITUNE: Aligning Large Language Models with Scientific Multimodal InstructionsarXiv 2023

Back to Papers