0

All papers

Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.

Debug like a Human: A Large Language Model Debugger via Verifying Runtime Execution Step-by-steparXiv 2024Solver-Informed RL: Grounding Large Language Models for Authentic Optimization ModelingarXiv 2025Autoregressive Model Beats Diffusion: Llama for Scalable Image GenerationarXiv 2024Gaze-LLE: Gaze Target Estimation via Large-Scale Learned EncodersCVPR 2025 1Efficient Region-Aware Neural Radiance Fields for High-Fidelity Talking Portrait SynthesisICCV 2023 1Holodeck: Language Guided Generation of 3D Embodied AI EnvironmentsCVPR 2024 1VERINA: Benchmarking Verifiable Code GenerationarXiv 2025DocScanner: Robust Document Image Rectification with Progressive LearningarXiv 2021MAT: Mask-Aware Transformer for Large Hole Image InpaintingCVPR 2022 1JAFAR: Jack up Any Feature at Any Resolutionjafar-jack-up-any-feature-at-any-resolutionSnapKV: LLM Knows What You are Looking for Before GenerationarXiv 2024StyleMaster: Stylize Your Video with Artistic Generation and TranslationCVPR 2025 1PokéChamp: an Expert-level Minimax Language AgentarXiv 2025Is Space-Time Attention All You Need for Video Understanding?arXiv 2021Intuitive physics understanding emerges from self-supervised pretraining on natural videosarXiv 2025Per-Pixel Classification is Not All You Need for Semantic SegmentationNeurIPS 2021 12Hyperbolic Image-Text RepresentationsarXiv 2023MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 SecondsCVPR 2025 1Improved Baselines with Momentum Contrastive LearningarXiv 2020Questioning Representational Optimism in Deep Learning: The Fractured Entangled Representation HypothesisarXiv 2025MLGym: A New Framework and Benchmark for Advancing AI Research AgentsarXiv 2025AnchorCrafter: Animate CyberAnchors Saling Your Products via Human-Object Interacting Video GenerationarXiv 2024End-to-End Autonomous Driving through V2X CooperationarXiv 2024EdgeTAM: On-Device Track Anything ModelCVPR 2025 1Code Llama: Open Foundation Models for CodearXiv 2023Dress Code: High-Resolution Multi-Category Virtual Try-OnarXiv 2022Sparse Autoencoders Learn Monosemantic Features in Vision-Language ModelsarXiv 2025Interactive Post-Training for Vision-Language-Action ModelsarXiv 2025EQ-Bench: An Emotional Intelligence Benchmark for Large Language ModelsarXiv 2023Sailor: Open Language Models for South-East AsiaarXiv 2024VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool UsearXiv 2025LucidDreamer: Towards High-Fidelity Text-to-3D Generation via Interval Score MatchingCVPR 2024 1Brain-JEPA: Brain Dynamics Foundation Model with Gradient Positioning and Spatiotemporal MaskingarXiv 2024Denoising Diffusion Implicit Modelsdenoising-diffusion-implicit-models3CAD: A Large-Scale Real-World 3C Product Dataset for Unsupervised AnomalyarXiv 2025Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarXiv 2025PSLM: Parallel Generation of Text and Speech with LLMs for Low-Latency Spoken Dialogue SystemsarXiv 2024PowerBEV: A Powerful Yet Lightweight Framework for Instance Prediction in Bird's-Eye ViewarXiv 2023PolyPythias: Stability and Outliers across Fifty Language Model Pre-Training RunsarXiv 2025Simple Cues Lead to a Strong Multi-Object TrackerCVPR 2023 1VisionZip: Longer is Better but Not Necessary in Vision Language ModelsCVPR 2025 1Mini-Gemini: Mining the Potential of Multi-modality Vision Language ModelsarXiv 2024Towards Comprehensive Detection of Chinese Harmful MemesarXiv 2024DynamiCrafter: Animating Open-domain Images with Video Diffusion PriorsarXiv 2023Retrieval-Augmented Perception: High-Resolution Image Perception Meets Visual RAGarXiv 2025SuperSimpleNet: Unifying Unsupervised and Supervised Learning for Fast and Reliable Surface Defect DetectionarXiv 2024Vulnerability Detection with Code Language Models: How Far Are We?arXiv 2024DreamO: A Unified Framework for Image CustomizationarXiv 2025Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and VerificationarXiv 2025Parallel Scaling Law for Language ModelsarXiv 2025Agents in Software Engineering: Survey, Landscape, and VisionarXiv 2024Let the Expert Stick to His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language ModelsarXiv 2024MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversationsmeld-a-multimodal-multi-party-dataset-for-1AutoML-Agent: A Multi-Agent LLM Framework for Full-Pipeline AutoMLarXiv 2024Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought ReasoningarXiv 2024A Systematic Study of Joint Representation Learning on Protein Sequences and StructuresarXiv 2023OpenAGI: When LLM Meets Domain ExpertsNeurIPS 2023 11AdaIR: Adaptive All-in-One Image Restoration via Frequency Mining and ModulationarXiv 2024Personalized Image Generation with Deep Generative Models: A Decade SurveyarXiv 2025Aligning Superhuman AI with Human Behavior: Chess as a Model SystemarXiv 2020Rethinking Memory in AI: Taxonomy, Operations, Topics, and Future DirectionsarXiv 2025Vision Transformers Don't Need Trained RegistersarXiv 2025OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image GenerationarXiv 2025GaussianFusion: Gaussian-Based Multi-Sensor Fusion for End-to-End Autonomous DrivingarXiv 2025Time-MMD: Multi-Domain Multimodal Dataset for Time Series AnalysisarXiv 2024KTO: Model Alignment as Prospect Theoretic OptimizationarXiv 2024Learnable latent embeddings for joint behavioral and neural analysisarXiv 2022HorizonStream: Long-Horizon Attention for Streaming 3D ReconstructionarXiv 2026Improving Factuality and Reasoning in Language Models through Multiagent DebatearXiv 2023Memory, Benchmark & Robots: A Benchmark for Solving Complex Tasks with Reinforcement LearningarXiv 2025Adaptation of Agentic AIarXiv 2025COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence ActarXiv 2024GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game AgentsarXiv 2026Deep Learning-Based Object Pose Estimation: A Comprehensive SurveyarXiv 2024ReasonFlux-PRM: Trajectory-Aware PRMs for Long Chain-of-Thought Reasoning in LLMsarXiv 2025GNN-RAG: Graph Neural Retrieval for Large Language Model ReasoningarXiv 2024LoRAMoE: Alleviate World Knowledge Forgetting in Large Language Models via MoE-Style PluginarXiv 2023Proactive Agent Research Environment: Simulating Active Users to Evaluate Proactive AssistantsarXiv 2026Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration TestingarXiv 2026Identity-Preserving Text-to-Video Generation by Frequency DecompositionCVPR 2025 1A Frame is Worth One Token: Efficient Generative World Modeling with Delta TokensarXiv 2026R-KV: Redundancy-aware KV Cache Compression for Training-Free Reasoning Models AccelerationarXiv 2025Towards Fast, Accurate and Stable 3D Dense Face Alignmenttowards-fast-accurate-and-stable-3d-denseGeneral-Reasoner: Advancing LLM Reasoning Across All DomainsarXiv 2025VisualQuality-R1: Reasoning-Induced Image Quality Assessment via Reinforcement Learning to RankarXiv 2025EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text PromptingarXiv 2025Frequency-Aware Deepfake Detection: Improving Generalizability through Frequency Space LearningarXiv 2024Character-LLM: A Trainable Agent for Role-PlayingarXiv 2023Universal Actions for Enhanced Embodied Foundation ModelsCVPR 2025 1How far are we from solving the 2D & 3D Face Alignment problem? (and a dataset of 230,000 3D facial landmarks)how-far-are-we-from-solving-the-2d-3d-face-1SplatFormer: Point Transformer for Robust 3D Gaussian SplattingarXiv 2024EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level PlanningarXiv 2023PanSplat: 4K Panorama Synthesis with Feed-Forward Gaussian SplattingCVPR 2025 1TaxoAdapt: Aligning LLM-Based Multidimensional Taxonomy Construction to Evolving Research CorporaarXiv 2025BTLM-3B-8K: 7B Parameter Performance in a 3B Parameter ModelarXiv 2023NeoBERT: A Next-Generation BERTarXiv 2025ChatEval: Towards Better LLM-based Evaluators through Multi-Agent DebatearXiv 2023CHGNet: Pretrained universal neural network potential for charge-informed atomistic modelingarXiv 2023detrex: Benchmarking Detection TransformersarXiv 2023SelfPose3d: Self-Supervised Multi-Person Multi-View 3d Pose EstimationCVPR 2024 1AdaCLIP: Adapting CLIP with Hybrid Learnable Prompts for Zero-Shot Anomaly DetectionarXiv 2024SoftCoT++: Test-Time Scaling with Soft Chain-of-Thought ReasoningarXiv 2025ORLM: A Customizable Framework in Training Large Models for Automated Optimization ModelingarXiv 2024HDR-GS: Efficient High Dynamic Range Novel View Synthesis at 1000x Speed via Gaussian SplattingarXiv 2024Structure-informed Language Models Are Protein DesignersarXiv 2023iBOT: Image BERT Pre-Training with Online TokenizerarXiv 2021Reinforcement Learning for Reasoning in Large Language Models with One Training ExamplearXiv 2025FlowTok: Flowing Seamlessly Across Text and Image TokensICCV 2025MeshAnything: Artist-Created Mesh Generation with Autoregressive TransformersarXiv 2024AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-TuningarXiv 2025Diffusion Guidance Is a Controllable Policy Improvement OperatorarXiv 2025Generative Modeling of Molecular Dynamics TrajectoriesarXiv 2024TimeSuite: Improving MLLMs for Long Video Understanding via Grounded TuningarXiv 2024RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for RoboticsCVPR 2025 1CODI: Compressing Chain-of-Thought into Continuous Space via Self-DistillationarXiv 2025BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex InstructionsarXiv 2024EnerVerse-AC: Envisioning Embodied Environments with Action ConditionarXiv 2025DexUMI: Using Human Hand as the Universal Manipulation Interface for Dexterous ManipulationarXiv 2025FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized SoundsarXiv 2024OpenDiLoCo: An Open-Source Framework for Globally Distributed Low-Communication TrainingarXiv 2024ReSimAD: Zero-Shot 3D Domain Transfer for Autonomous Driving with Source Reconstruction and Target SimulationarXiv 2023Noisy-Correspondence Learning for Text-to-Image Person Re-identificationCVPR 2024 1SMILE: Single-turn to Multi-turn Inclusive Language Expansion via ChatGPT for Mental Health SupportarXiv 2023Process Reinforcement through Implicit RewardsarXiv 2025Learning Getting-Up Policies for Real-World Humanoid RobotsarXiv 2025Towards Semantic Segmentation of Urban-Scale 3D Point Clouds: A Dataset, Benchmarks and ChallengesCVPR 2021 1RandLA-Net: Efficient Semantic Segmentation of Large-Scale Point Cloudsrandla-net-efficient-semantic-segmentation-ofHMT: Hierarchical Memory Transformer for Long Context Language ProcessingarXiv 2024Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning IncentivizationarXiv 2025Sim-to-Real Transfer for Mobile Robots with Reinforcement Learning: from NVIDIA Isaac Sim to Gazebo and Real ROS 2 RobotsarXiv 2025BiFormer: Vision Transformer with Bi-Level Routing AttentionCVPR 2023 1IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI SystemsarXiv 2025Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video GamesarXiv 2025TravelPlanner: A Benchmark for Real-World Planning with Language AgentsarXiv 2024M-Prometheus: A Suite of Open Multilingual LLM JudgesarXiv 2025EmoBench: Evaluating the Emotional Intelligence of Large Language ModelsarXiv 2024DenseShift: Towards Accurate and Efficient Low-Bit Power-of-Two QuantizationICCV 2023 1Conversation Graph: Data Augmentation, Training and Evaluation for Non-Deterministic Dialogue ManagementarXiv 2020HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and DetectionarXiv 2022Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and BeyondarXiv 2025Mind2Web: Towards a Generalist Agent for the Webmind2web-towards-a-generalist-agent-for-theMovieChat: From Dense Token to Sparse Memory for Long Video UnderstandingCVPR 2024 1Process Reward Models for LLM Agents: Practical Framework and DirectionsarXiv 2025Interpretable Neural-Symbolic Concept ReasoningarXiv 2023Soft Actor-Critic Algorithms and ApplicationsarXiv 2018RoFL: Robustness of Secure Federated LearningarXiv 2021LogAI: A Library for Log Analytics and IntelligencearXiv 2023Masked Autoencoders for Microscopy are Scalable Learners of Cellular BiologyCVPR 2024 1LLaVA-CoT: Let Vision Language Models Reason Step-by-SteparXiv 2024MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMsarXiv 2025WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image GenerationarXiv 2025Open-Sora Plan: Open-Source Large Video Generation ModelarXiv 2024TOTEM: TOkenized Time Series EMbeddings for General Time Series AnalysisarXiv 2024Graph Retrieval-Augmented Generation: A SurveyarXiv 2024LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA CompositionarXiv 2023Machine Mindset: An MBTI Exploration of Large Language ModelsarXiv 2023SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsarXiv 2023PAL: Pluralistic Alignment Framework for Learning from Heterogeneous PreferencesarXiv 2024CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMsarXiv 2024Before Generation, Align it! A Novel and Effective Strategy for Mitigating Hallucinations in Text-to-SQL GenerationarXiv 2024Liquid Time-constant NetworksarXiv 2020DARTS: Differentiable Architecture Searchdarts-differentiable-architecture-search-1Initializing Models with Larger OnesarXiv 2023QuanTA: Efficient High-Rank Fine-Tuning of LLMs with Quantum-Informed Tensor AdaptationarXiv 2024PonderV2: Pave the Way for 3D Foundation Model with A Universal Pre-training ParadigmarXiv 2023HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language ModelsarXiv 2023StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryICCV 2021 10Video Mamba Suite: State Space Model as a Versatile Alternative for Video UnderstandingarXiv 2024VideoMamba: State Space Model for Efficient Video UnderstandingarXiv 2024Causal Reasoning and Large Language Models: Opening a New Frontier for CausalityarXiv 2023MedCLIP: Contrastive Learning from Unpaired Medical Images and TextarXiv 2022Punica: Multi-Tenant LoRA ServingarXiv 2023VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-TuningarXiv 2025Text-to-SQL Empowered by Large Language Models: A Benchmark EvaluationarXiv 2023WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsarXiv 2022Large Language Models Meet Symbolic Provers for Logical Reasoning EvaluationarXiv 2025EuroSAT: A Novel Dataset and Deep Learning Benchmark for Land Use and Land Cover ClassificationarXiv 2017LMDrive: Closed-Loop End-to-End Driving with Large Language ModelsCVPR 2024 1SimCSE: Simple Contrastive Learning of Sentence EmbeddingsEMNLP 2021 11Understanding R1-Zero-Like Training: A Critical PerspectivearXiv 2025GenAD: Generalized Predictive Model for Autonomous DrivingCVPR 2024 1A Frustratingly Easy Approach for Entity and Relation ExtractionNAACL 2021 4How to Train Long-Context Language Models (Effectively)arXiv 2024High-Quality Entity SegmentationarXiv 2022WizMap: Scalable Interactive Visualization for Exploring Large Machine Learning EmbeddingsarXiv 2023Graph-based Topology Reasoning for Driving ScenesarXiv 2023A Graph-Based Approach for Category-Agnostic Pose EstimationarXiv 2023SafeDreamer: Safe Reinforcement Learning with World ModelsarXiv 2023DLF: Disentangled-Language-Focused Multimodal Sentiment AnalysisarXiv 2024Sheared LLaMA: Accelerating Language Model Pre-training via Structured PruningarXiv 2023OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific ProblemsarXiv 2024LESS: Selecting Influential Data for Targeted Instruction TuningarXiv 2024GeoLLM: Extracting Geospatial Knowledge from Large Language ModelsarXiv 2023$\infty$Bench: Extending Long Context Evaluation Beyond 100K TokensarXiv 2024VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric TasksNeurIPS 2023 11MTGS: Multi-Traversal Gaussian SplattingarXiv 2025Scaling and evaluating sparse autoencodersarXiv 2024Text-to-CadQuery: A New Paradigm for CAD Generation with Scalable Large Model CapabilitiesarXiv 2025Adversarial Latent Autoencodersadversarial-latent-autoencoders-1Exploration by Random Network DistillationICLR 2019

Back to Papers