All papers
Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.
UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement LearningarXiv 2025OctoThinker: Mid-training Incentivizes Reinforcement Learning ScalingarXiv 2025Evolving Language Models without Labels: Majority Drives Selection,
Novelty Promotes VariationarXiv 2025Leaky Thoughts: Large Reasoning Models Are Not Private ThinkersarXiv 2025Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmasarXiv 2025VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in
Real-world ApplicationsarXiv 2025HeuriGym: An Agentic Benchmark for LLM-Crafted Heuristics in Combinatorial OptimizationarXiv 2025Klear-Reasoner: Advancing Reasoning Capability via Gradient-Preserving Clipping Policy OptimizationarXiv 2025PrefPalette: Personalized Preference Modeling with Latent AttributesarXiv 2025Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMsarXiv 2025DSI-Bench: A Benchmark for Dynamic Spatial IntelligencearXiv 2025Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive SurveyarXiv 2025Follow-Your-Shape: Shape-Aware Image Editing via Trajectory-Guided Region ControlarXiv 2025AdaCtrl: Towards Adaptive and Controllable Reasoning via Difficulty-Aware BudgetingarXiv 2025DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion TransformersarXiv 2025DRAGged into Conflicts: Detecting and Addressing Conflicting Sources in Search-Augmented LLMsarXiv 2025Data Whisperer: Efficient Data Selection for Task-Specific LLM Fine-Tuning via Few-Shot In-Context LearningarXiv 2025Token Perturbation Guidance for Diffusion ModelsarXiv 2025Agents of Change: Self-Evolving LLM Agents for Strategic PlanningarXiv 2025REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion TrainingarXiv 2025O$^2$-Searcher: A Searching-based Agent Model for Open-Domain Open-Ended Question AnsweringarXiv 2025Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache CompressionarXiv 2025Enhancing Efficiency and Exploration in Reinforcement Learning for LLMsarXiv 2025ARGUS: Hallucination and Omission Evaluation in Video-LLMsICCV 2025Taming Masked Diffusion Language Models via Consistency Trajectory
Reinforcement Learning with Fewer Decoding SteparXiv 2025Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language NavigationarXiv 2025HAPO: Training Language Models to Reason Concisely via History-Aware Policy OptimizationarXiv 2025Don't Overthink It: A Survey of Efficient R1-style Large Reasoning ModelsarXiv 2025LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in
Mechanism via Multi-Step ReasoningarXiv 2025Do Large Language Models Excel in Complex Logical Reasoning with Formal Language?arXiv 2025Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM ReasoningarXiv 2025M2-Reasoning: Empowering MLLMs with Unified General and Spatial ReasoningarXiv 2025AgentsNet: Coordination and Collaborative Reasoning in Multi-Agent LLMsarXiv 2025UniSkill: Imitating Human Videos via Cross-Embodiment Skill RepresentationsarXiv 2025EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?arXiv 2025Parallel Continuous Chain-of-Thought with Jacobi IterationarXiv 2025Benchmarking Optimizers for Large Language Model PretrainingarXiv 2025The Unreasonable Effectiveness of Entropy Minimization in LLM ReasoningarXiv 2025Visionary-R1: Mitigating Shortcuts in Visual Reasoning with Reinforcement LearningarXiv 2025Reasoning Models Better Express Their ConfidencearXiv 2025EasyText: Controllable Diffusion Transformer for Multilingual Text RenderingarXiv 2025CrossLMM: Decoupling Long Video Sequences from LMMs via Dual Cross-Attention MechanismsarXiv 2025Rank-K: Test-Time Reasoning for Listwise RerankingarXiv 2025MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning BenchmarkarXiv 2025GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI AgentsarXiv 2025Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion ModelarXiv 2025Language Models Can Learn from Verbal Feedback Without Scalar RewardsarXiv 2025OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and EditingarXiv 2025PersonaLens: A Benchmark for Personalization Evaluation in Conversational AI AssistantsarXiv 2025StepFun-Formalizer: Unlocking the Autoformalization Potential of LLMs through Knowledge-Reasoning FusionarXiv 2025Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement LearningarXiv 2025Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context InjectionarXiv 2025LIMOPro: Reasoning Refinement for Efficient and Effective Test-time ScalingarXiv 2025ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMsarXiv 2025ReDit: Reward Dithering for Improved LLM Policy OptimizationarXiv 2025Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVRarXiv 2025More Thought, Less Accuracy? On the Dual Nature of Reasoning in
Vision-Language ModelsarXiv 2025Test-Time Reinforcement Learning for GUI Grounding via Region ConsistencyarXiv 2025DeepTheorem: Advancing LLM Reasoning for Theorem Proving Through Natural Language and Reinforcement LearningarXiv 2025Overclocking LLM Reasoning: Monitoring and Controlling Thinking Path Lengths in LLMsarXiv 2025RealUnify: Do Unified Models Truly Benefit from Unification? A
Comprehensive BenchmarkarXiv 2025Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement LearningarXiv 2025SWE-Flow: Synthesizing Software Engineering Data in a Test-Driven MannerarXiv 2025SiLVR: A Simple Language-based Video Reasoning FrameworkarXiv 2025DeepCritic: Deliberate Critique with Large Language ModelsarXiv 2025FullFront: Benchmarking MLLMs Across the Full Front-End Engineering WorkflowarXiv 2025AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMsarXiv 2025Seq vs Seq: An Open Suite of Paired Encoders and DecodersarXiv 2025UItron: Foundational GUI Agent with Advanced Perception and PlanningarXiv 2025Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive RewardsarXiv 2025Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RLarXiv 2025Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning ModelsarXiv 2025Reinforcing General Reasoning without VerifiersarXiv 2025KaLM-Embedding-V2: Superior Training Techniques and Data Inspire A Versatile Embedding ModelarXiv 2025UI-AGILE: Advancing GUI Agents with Effective Reinforcement Learning and Precise Inference-Time GroundingarXiv 2025Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human StatesarXiv 2025Long-Short Chain-of-Thought Mixture Supervised Fine-Tuning Eliciting Efficient Reasoning in Large Language ModelsarXiv 2025EgoPrivacy: What Your First-Person Camera Says About You?arXiv 2025SPC: Evolving Self-Play Critic via Adversarial Games for LLM ReasoningarXiv 2025ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All ToolsarXiv 2024Efficient Memory Management for Large Language Model Serving with PagedAttentionarXiv 2023Towards Real-World Blind Face Restoration with Generative Facial PriorCVPR 2021 1Drag Your GAN: Interactive Point-based Manipulation on the Generative Image ManifoldarXiv 2023MediaPipe: A Framework for Building Perception PipelinesarXiv 2019Real-ESRGAN: Training Real-World Blind Super-Resolution with Pure Synthetic DataarXiv 2021Demonstrate-Search-Predict: Composing retrieval and language models for knowledge-intensive NLParXiv 2022Self-Instruct: Aligning Language Models with Self-Generated InstructionsarXiv 2022Assisting in Writing Wikipedia-like Articles From Scratch with Large Language ModelsarXiv 2024TradingAgents: Multi-Agents LLM Financial Trading FrameworkarXiv 2024Adversarial Diffusion DistillationarXiv 2023MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsarXiv 2023MiniGPT-v2: large language model as a unified interface for vision-language multi-task learningarXiv 2023Towards Agentic Recommender Systems in the Era of Multimodal Large Language ModelsarXiv 2025SGLang: Efficient Execution of Structured Language Model ProgramsarXiv 2023RET-LLM: Towards a General Read-Write Memory for Large Language ModelsarXiv 2023Efficient and Effective Text Encoding for Chinese LLaMA and AlpacaarXiv 2023Unity: A General Platform for Intelligent AgentsarXiv 2018Augmented SBERT: Data Augmentation Method for Improving Bi-Encoders for Pairwise Sentence Scoring TasksNAACL 2021 4Towards Robust Blind Face Restoration with Codebook Lookup TransformerarXiv 2022Self-Attention with Relative Position Representationsself-attention-with-relative-position-1Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityarXiv 2024First Order Motion Model for Image Animationfirst-order-motion-model-for-image-animationWan: Open and Advanced Large-Scale Video Generative ModelsarXiv 2025A guide to convolution arithmetic for deep learningarXiv 2016OGB-LSC: A Large-Scale Challenge for Machine Learning on GraphsarXiv 2021YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectorsCVPR 2023 1F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow MatchingarXiv 2024SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face AnimationCVPR 2023 1Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning AlgorithmsarXiv 2017GoEX: Perspectives and Designs Towards a Runtime for Autonomous LLM ApplicationsarXiv 2024CogVideoX: Text-to-Video Diffusion Models with An Expert TransformerarXiv 2024pix2code: Generating Code from a Graphical User Interface ScreenshotarXiv 2017Ludwig: a type-based declarative deep learning toolboxarXiv 2019HunyuanVideo: A Systematic Framework For Large Video Generative ModelsarXiv 2024A Closed-form Solution to Photorealistic Image Stylizationa-closed-form-solution-to-photorealistic-1YOLOv10: Real-Time End-to-End Object DetectionarXiv 2024Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech TokensarXiv 2025Pre-Training with Whole Word Masking for Chinese BERTarXiv 2019PhotoMaker: Customizing Realistic Human Photos via Stacked ID EmbeddingCVPR 2024 1Neural Discrete Representation Learningneural-discrete-representation-learning-1WizardCoder: Empowering Code Large Language Models with Evol-InstructarXiv 2023Agent S2: A Compositional Generalist-Specialist Framework for Computer Use AgentsarXiv 2025CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-XarXiv 2023PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPUarXiv 2023The Rise and Potential of Large Language Model Based Agents: A SurveyarXiv 2023Yi: Open Foundation Models by 01.AIarXiv 2024LLM.int8(): 8-bit Matrix Multiplication for Transformers at ScalearXiv 2022GLM-130B: An Open Bilingual Pre-trained ModelarXiv 2022Objects as PointsarXiv 2019A Task is Worth One Word: Learning with Task Prompts for High-Quality Versatile Image InpaintingarXiv 20233D Photography using Context-aware Layered Depth Inpainting3d-photography-using-context-aware-layered-1SAMURAI: Adapting Segment Anything Model for Zero-Shot Visual Tracking with Motion-Aware Memorysamurai-adapting-segment-anything-model-forCogVLM2: Visual Language Models for Image and Video UnderstandingarXiv 2024Know More about Each Other: Evolving Dialogue Strategy via Compound Assessmentknow-more-about-each-other-evolving-dialogueScore Distillation via Reparametrized DDIMarXiv 2024InstructPix2Pix: Learning to Follow Image Editing InstructionsCVPR 2023 1Neural Volume Rendering: NeRF And BeyondarXiv 2020CogAgent: A Visual Language Model for GUI AgentsCVPR 2024 1ESRGAN: Enhanced Super-Resolution Generative Adversarial NetworksarXiv 2018OPT: Open Pre-trained Transformer Language ModelsarXiv 2022MLS: A Large-Scale Multilingual Dataset for Speech ResearcharXiv 2020ProPainter: Improving Propagation and Transformer for Video InpaintingICCV 2023 1scikit-image: Image processing in PythonarXiv 2014Adversarial PatcharXiv 2017Know Your Self-supervised Learning: A Survey on Image-based Generative
and Discriminative TrainingarXiv 2023DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth GradientsarXiv 2016XLNet: Generalized Autoregressive Pretraining for Language Understandingxlnet-generalized-autoregressive-pretraining-1Progressive Growing of GANs for Improved Quality, Stability, and Variationprogressive-growing-of-gans-for-improved-1U-GAT-IT: Unsupervised Generative Attentional Networks with Adaptive Layer-Instance Normalization for Image-to-Image TranslationICLR 2020 1TripoSR: Fast 3D Object Reconstruction from a Single ImagearXiv 2024Bag of Freebies for Training Object Detection Neural NetworksarXiv 2019YOLOv6 v3.0: A Full-Scale ReloadingarXiv 2023Automated Unit Test Improvement using Large Language Models at MetaarXiv 2024GR00T N1: An Open Foundation Model for Generalist Humanoid RobotsarXiv 2025StarVector: Generating Scalable Vector Graphics Code from Images and TextCVPR 2025 1MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal SamplingarXiv 2024Realtime Multi-Person 2D Pose Estimation using Part Affinity Fieldsrealtime-multi-person-2d-pose-estimation-1AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait AnimationarXiv 2024RedPajama: an Open Dataset for Training Large Language ModelsarXiv 2024Densely Connected Convolutional Networksdensely-connected-convolutional-networks-1AnyText: Multilingual Visual Text Generation And EditingarXiv 2023SSD: Single Shot MultiBox DetectorarXiv 2015Improving Diffusion Models for Authentic Virtual Try-on in the WildarXiv 2024OpenPrompt: An Open-source Framework for Prompt-learningACL 2022 5OpenAgents: An Open Platform for Language Agents in the WildarXiv 2023The Prompt Report: A Systematic Survey of Prompting TechniquesarXiv 2024Step-Audio: Unified Understanding and Generation in Intelligent Speech InteractionarXiv 2025A Lightweight Instrument-Agnostic Model for Polyphonic Note Transcription and Multipitch EstimationarXiv 2022OpenNRE: An Open and Extensible Toolkit for Neural Relation Extractionopennre-an-open-and-extensible-toolkit-for-1Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video GenerationICCV 2023 1Instruction Tuning with GPT-4arXiv 2023The Ideal Continual Learner: An Agent That Never ForgetsarXiv 2023Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese UnderstandingarXiv 2024OmniGen: Unified Image GenerationCVPR 2025 1Champ: Controllable and Consistent Human Image Animation with 3D Parametric GuidancearXiv 2024InstructIE: A Bilingual Instruction-based Information Extraction DatasetarXiv 2023MODNet: Real-Time Trimap-Free Portrait Matting via Objective DecompositionarXiv 2020InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction ModelsarXiv 2024Segment Anything in High QualityNeurIPS 2023 11Diffusion Policy: Visuomotor Policy Learning via Action DiffusionarXiv 2023Feature Generation by Convolutional Neural Network for Click-Through Rate PredictionarXiv 2019Generative Adversarial Networksgenerative-adversarial-networks-1DiffBIR: Towards Blind Image Restoration with Generative Diffusion PriorarXiv 2023Snips Voice Platform: an embedded Spoken Language Understanding system for private-by-design voice interfacesarXiv 2018Guiding Instruction-based Image Editing via Multimodal Large Language ModelsarXiv 2023Qwen2.5-Omni Technical ReportarXiv 2025T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion ModelsarXiv 2023Open Deep Search: Democratizing Search with Open-source Reasoning AgentsarXiv 2025PLAID: An Efficient Engine for Late Interaction RetrievalarXiv 2022Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image AnimationarXiv 2024Transformer-XL: Attentive Language Models Beyond a Fixed-Length Contexttransformer-xl-attentive-language-models-1Inductive Representation Learning on Large Graphsinductive-representation-learning-on-large-1Thin-Plate Spline Motion Model for Image AnimationCVPR 2022 1Recognize Anything: A Strong Image Tagging ModelarXiv 2023Squeeze-and-Excitation Networkssqueeze-and-excitation-networks-1Open-Set Image Tagging with Multi-Grained Text SupervisionarXiv 2023PuLID: Pure and Lightning ID Customization via Contrastive AlignmentarXiv 2024"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language ModelsarXiv 2023GLM: General Language Model Pretraining with Autoregressive Blank InfillingACL 2022 5FCOS: Fully Convolutional One-Stage Object Detectionfcos-fully-convolutional-one-stage-object-1