All papers
Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.
InSTA: Towards Internet-Scale Training For AgentsarXiv 2025Multilingual Machine Translation with Open Large Language Models at Practical Scale: An Empirical StudyarXiv 2025The Superposition of Diffusion Models Using the Itô Density EstimatorarXiv 2024InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured CaptionCVPR 2025 1Geometry-Aware Generative Autoencoders for Warped Riemannian Metric Learning and Generative Modeling on Data ManifoldsarXiv 2024Scalable Chain of Thoughts via Elastic ReasoningarXiv 2025Training Language Models to Self-Correct via Reinforcement LearningarXiv 2024Measuring General Intelligence with Generated GamesarXiv 2025SoFlow: Solution Flow Models for One-Step Generative ModelingarXiv 2025ChartEdit: How Far Are MLLMs From Automating Chart Analysis? Evaluating MLLMs' Capability via Chart EditingarXiv 2025The Traitors: Deception and Trust in Multi-Agent Language Model SimulationsarXiv 2025AudioTrust: Benchmarking the Multifaceted Trustworthiness of Audio Large Language ModelsarXiv 2025One RL to See Them All: Visual Triple Unified Reinforcement LearningarXiv 2025One-shot Entropy MinimizationarXiv 2025ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific WorkflowsarXiv 2025DetailFlow: 1D Coarse-to-Fine Autoregressive Image Generation via Next-Detail PredictionarXiv 2025The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to ReasonarXiv 2025Advancing Multimodal Reasoning: From Optimized Cold Start to Staged Reinforcement LearningarXiv 2025PLATO: Pre-trained Dialogue Generation Model with Discrete Latent Variableplato-pre-trained-dialogue-generation-model-1Long Time No See! Open-Domain Conversation with Long-Term Persona MemoryFindings (ACL) 2022 5Evolving Reinforcement Learning Algorithmsevolving-reinforcement-learning-algorithmsPiFold: Toward effective and efficient protein inverse foldingarXiv 2022A recipe for scalable attention-based MLIPs: unlocking long-range accuracy with all-to-all node attentionarXiv 2026o1-Coder: an o1 Replication for CodingarXiv 2024LifelongMemory: Leveraging LLMs for Answering Queries in Long-form Egocentric VideosarXiv 2023A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and DeploymentarXiv 2025Embedding-based classifiers can detect prompt injection attacksarXiv 2024Multiview Scene GrapharXiv 2024IndicTrans2: Towards High-Quality and Accessible Machine Translation Models for all 22 Scheduled Indian LanguagesarXiv 2023SSCBench: A Large-Scale 3D Semantic Scene Completion Benchmark for Autonomous DrivingarXiv 2023RouteFinder: Towards Foundation Models for Vehicle Routing ProblemsarXiv 2024PVBM: A Python Vasculature Biomarker Toolbox Based On Retinal Blood Vessel SegmentationarXiv 2022OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory SynthesisarXiv 2026RedCode: Risky Code Execution and Generation Benchmark for Code AgentsarXiv 2024Pretrained Language Models for Sequential Sentence Classificationpretrained-language-models-for-sequential-1The Wisdom of Crowds: Temporal Progressive Attention for Early Action PredictionCVPR 2023 1Towards Unified Music Emotion Recognition across Dimensional and Categorical ModelsarXiv 2025Reasoning-Table: Exploring Reinforcement Learning for Table ReasoningarXiv 2025PolyFormer: Referring Image Segmentation as Sequential Polygon GenerationCVPR 2023 1Not All Steps are Created Equal: Selective Diffusion Distillation for Image ManipulationICCV 2023 1LeviTor: 3D Trajectory Oriented Image-to-Video SynthesisCVPR 2025 1Universal Guidance for Diffusion ModelsarXiv 2023A Bayesian Flow Network Framework for Chemistry TasksarXiv 2024DuReader_retrieval: A Large-scale Chinese Benchmark for Passage Retrieval from Web Search EnginearXiv 2022SparseLLM: Towards Global Pruning for Pre-trained Language ModelsarXiv 2024Progressive Transformers for End-to-End Sign Language ProductionECCV 2020 8Using millions of emoji occurrences to learn any-domain representations for detecting sentiment, emotion and sarcasmusing-millions-of-emoji-occurrences-to-learn-1Composed Image Retrieval for Remote SensingarXiv 2024Towards Fewer Annotations: Active Learning via Region Impurity and Prediction Uncertainty for Domain Adaptive Semantic SegmentationCVPR 2022 1SePiCo: Semantic-Guided Pixel Contrast for Domain Adaptive Semantic SegmentationarXiv 2022MagicPose: Realistic Human Poses and Facial Expressions Retargeting with Identity-aware DiffusionarXiv 2023Align and Attend: Multimodal Summarization with Dual Contrastive LossesCVPR 2023 1Woodpecker: Hallucination Correction for Multimodal Large Language ModelsarXiv 2023FitDiT: Advancing the Authentic Garment Details for High-fidelity Virtual Try-onarXiv 2024A Comparison of Discrete and Soft Speech Units for Improved Voice ConversionarXiv 2021AGILE: A Novel Reinforcement Learning Framework of LLM AgentsarXiv 2024Convolutions Die Hard: Open-Vocabulary Segmentation with Single Frozen Convolutional CLIPconvolutions-die-hard-open-vocabularyWhere do Large Vision-Language Models Look at when Answering Questions?arXiv 2025Advancing Surgical VQA with Scene Graph KnowledgearXiv 2023ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM InferencearXiv 2024CLIP-NeRF: Text-and-Image Driven Manipulation of Neural Radiance FieldsCVPR 2022 1Mixing Dirichlet Topic Models and Word Embeddings to Make lda2vecarXiv 2016AdvCLIP: Downstream-agnostic Adversarial Examples in Multimodal Contrastive LearningarXiv 2023DarkSAM: Fooling Segment Anything Model to Segment NothingarXiv 2024DiffusionInst: Diffusion Model for Instance SegmentationarXiv 2022Benchmarking Knowledge-driven Zero-shot LearningarXiv 2021Orthogonal Subspace Learning for Language Model Continual LearningarXiv 2023Vision Search Assistant: Empower Vision-Language Models as Multimodal Search EnginesarXiv 2024You Only Look at Screens: Multimodal Chain-of-Action AgentsarXiv 2023Semantics-aware BERT for Language UnderstandingarXiv 20191M-Deepfakes Detection ChallengearXiv 2024On the use of Vision-Language models for Visual Sentiment Analysis: a study on CLIParXiv 2023OneLLM: One Framework to Align All Modalities with LanguageCVPR 2024 1GeoCalib: Learning Single-image Calibration with Geometric OptimizationarXiv 2024DATED: Guidelines for Creating Synthetic Datasets for Engineering Design ApplicationsarXiv 2023OpenCOLE: Towards Reproducible Automatic Graphic Design GenerationarXiv 2024TransNeXt: Robust Foveal Visual Perception for Vision TransformersCVPR 2024 1Data-centric Artificial Intelligence: A SurveyarXiv 2023PEER: A Comprehensive and Multi-Task Benchmark for Protein Sequence UnderstandingarXiv 2022RotatE: Knowledge Graph Embedding by Relational Rotation in Complex Spacerotate-knowledge-graph-embedding-by-1FlowMap: High-Quality Camera Poses, Intrinsics, and Depth via Gradient DescentarXiv 2024Feature Re-Embedding: Towards Foundation Model-Level Performance in Computational PathologyCVPR 2024 1Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word ProblemsarXiv 2017Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMsarXiv 2025TriDet: Temporal Action Detection with Relative Boundary ModelingCVPR 2023 1Group equivariant neural posterior estimationgroup-equivariant-neural-posterior-estimationUnmasking and Improving Data Credibility: A Study with Datasets for Training Harmless Language ModelsarXiv 2023Medical MLLM is Vulnerable: Cross-Modality Jailbreak and Mismatched Attacks on Medical Multimodal Large Language ModelsarXiv 2024SRA-MCTS: Self-driven Reasoning Augmentation with Monte Carlo Tree Search for Code GenerationarXiv 20243D-aware Conditional Image SynthesisCVPR 2023 1Memory-and-Anticipation Transformer for Online Action UnderstandingICCV 2023 1Open-domain Visual Entity Recognition: Towards Recognizing Millions of Wikipedia EntitiesICCV 2023 1Demystifying Long Chain-of-Thought Reasoning in LLMsarXiv 2025PaperRobot: Incremental Draft Generation of Scientific Ideaspaperrobot-incremental-draft-generation-of-1LongEmbed: Extending Embedding Models for Long Context RetrievalarXiv 2024Denoising Diffusion Step-aware ModelsarXiv 2023Defect Spectrum: A Granular Look of Large-Scale Defect Datasets with Rich SemanticsarXiv 2023Ref-NeuS: Ambiguity-Reduced Neural Implicit Surface Learning for Multi-View Reconstruction with ReflectionICCV 2023 1STaR: Bootstrapping Reasoning With ReasoningarXiv 2022Language Control Diffusion: Efficiently Scaling through Space, Time, and TasksarXiv 2022Quiet-STaR: Language Models Can Teach Themselves to Think Before SpeakingarXiv 2024Barlow Twins: Self-Supervised Learning via Redundancy ReductionarXiv 2021Consistent Video Depth EstimationarXiv 2020The Hateful Memes Challenge: Detecting Hate Speech in Multimodal MemesNeurIPS 2020 12Deep Clustering for Unsupervised Learning of Visual Featuresdeep-clustering-for-unsupervised-learning-of-1Movie Gen: A Cast of Media Foundation ModelsarXiv 2024Fast and Accurate Model ScalingCVPR 2021 1LLM-QAT: Data-Free Quantization Aware Training for Large Language ModelsarXiv 2023Non-local Neural Networksnon-local-neural-networks-1Semi-Supervised Offline Reinforcement Learning with Action-Free TrajectoriesarXiv 2022Towards VQA Models That Can Readtowards-vqa-models-that-can-read-1Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsNeurIPS 2020 12LoFiT: Localized Fine-tuning on LLM RepresentationsarXiv 2024Chat-Edit-3D: Interactive 3D Scene Editing via Text PromptsarXiv 2024VFLAIR: A Research Library and Benchmark for Vertical Federated LearningarXiv 2023Pre-training for Speech Translation: CTC Meets Optimal TransportarXiv 2023Real-time Photorealistic Dynamic Scene Representation and Rendering with 4D Gaussian SplattingarXiv 2023Reformatted AlignmentarXiv 2024Multi-event Video-Text RetrievalICCV 2023 1GameGen-X: Interactive Open-world Game Video GenerationarXiv 2024RRHF: Rank Responses to Align Language Models with Human Feedback without tearsarXiv 2023How much is a noisy image worth? Data Scaling Laws for Ambient DiffusionarXiv 2024PokeLLMon: A Human-Parity Agent for Pokemon Battles with Large Language ModelsarXiv 2024Multi-Source Diffusion Models for Simultaneous Music Generation and SeparationarXiv 2023ByT5: Towards a token-free future with pre-trained byte-to-byte modelsarXiv 2021Texts as Images in Prompt Tuning for Multi-Label Image RecognitionCVPR 2023 1Segment Any Mesh: Zero-shot Mesh Part Segmentation via Lifting Segment Anything 2 to 3DarXiv 2024End-to-end Learning of Driving Models from Large-scale Video Datasetsend-to-end-learning-of-driving-models-from-1TokenFormer: Rethinking Transformer Scaling with Tokenized Model ParametersarXiv 2024MedCLIP-SAM: Bridging Text and Image Towards Universal Medical Image SegmentationarXiv 2024The Tatoeba Translation Challenge -- Realistic Data Sets for Low Resource and Multilingual MTarXiv 2020MeViS: A Large-scale Benchmark for Video Segmentation with Motion ExpressionsICCV 2023 1GRES: Generalized Referring Expression Segmentationgres-generalized-referring-expressionConvolutional Networks on Graphs for Learning Molecular Fingerprintsconvolutional-networks-on-graphs-for-learning-1ThinkGrasp: A Vision-Language System for Strategic Part Grasping in ClutterCoRL2024Binding Language Models in Symbolic LanguagesarXiv 2022Aguvis: Unified Pure Vision Agents for Autonomous GUI InteractionarXiv 2024CLEAR: Conv-Like Linearization Revs Pre-Trained Diffusion Transformers UparXiv 2024DriveGEN: Generalized and Robust 3D Detection in Driving via Controllable Text-to-Image Diffusion GenerationCVPR 2025 1EpiDiff: Enhancing Multi-View Synthesis via Localized Epipolar-Constrained DiffusionCVPR 2024 1From Parts to Whole: A Unified Reference Framework for Controllable Human Image GenerationarXiv 2024Rethinking Model Ensemble in Transfer-based Adversarial AttacksarXiv 2023ViG: Linear-complexity Visual Sequence Learning with Gated Linear AttentionarXiv 2024LLMs Learn Task Heuristics from Demonstrations: A Heuristic-Driven Prompting Strategy for Document-Level Event Argument ExtractionarXiv 2023Unitxt: Flexible, Shareable and Reusable Data Preparation and Evaluation for Generative AIarXiv 2024Granite Code Models: A Family of Open Foundation Models for Code IntelligencearXiv 2024ChatRex: Taming Multimodal LLM for Joint Perception and UnderstandingarXiv 2024StyleAvatar3D: Leveraging Image-Text Diffusion Models for High-Fidelity 3D Avatar GenerationarXiv 2023HumanSD: A Native Skeleton-Guided Diffusion Model for Human Image GenerationICCV 2023 1Non-deep Networksnon-deep-networksWebCanvas: Benchmarking Web Agents in Online EnvironmentsarXiv 2024Cross-Domain Aspect Extraction using Transformers Augmented with Knowledge GraphsarXiv 2022From Temporal to Contemporaneous Iterative Causal Discovery in the Presence of Latent ConfoundersarXiv 2023Exploring the Limit of Outcome Reward for Learning Mathematical ReasoningarXiv 2025Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language ModelsarXiv 2024Language-driven Semantic Segmentationlanguage-driven-semantic-segmentationT2V-Turbo-v2: Enhancing Video Generation Model Post-Training through Data, Reward, and Conditional Guidance DesignarXiv 2024Training Turn-by-Turn Verifiers for Dialogue Tutoring Agents: The Curious Case of LLMs as Your Coding TutorsarXiv 2025T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward FeedbackarXiv 2024One Diffusion Step to Real-World Super-Resolution via Flow Trajectory DistillationarXiv 2025Multi-interactive Feature Learning and a Full-time Multi-modality Benchmark for Image Fusion and SegmentationICCV 2023 1Is ChatGPT Fair for Recommendation? Evaluating Fairness in Large Language Model RecommendationarXiv 2023FIFO-Diffusion: Generating Infinite Videos from Text without TrainingarXiv 2024Offline Reinforcement Learning for LLM Multi-Step ReasoningarXiv 2024A Survey on Hardware Accelerators for Large Language ModelsarXiv 2024Bridging Vision and Language Encoders: Parameter-Efficient Tuning for Referring Image SegmentationICCV 2023 1VLind-Bench: Measuring Language Priors in Large Vision-Language ModelsarXiv 2024A multi-path 2.5 dimensional convolutional neural network system for segmenting stroke lesions in brain MRI imagesarXiv 2019ChatDoctor: A Medical Chat Model Fine-Tuned on a Large Language Model Meta-AI (LLaMA) Using Medical Domain KnowledgearXiv 2023Improved Precision and Recall Metric for Assessing Generative Modelsimproved-precision-and-recall-metric-for-1LaserMix for Semi-Supervised LiDAR Semantic SegmentationCVPR 2023 1SPIdepth: Strengthened Pose Information for Self-supervised Monocular Depth EstimationarXiv 2024Step Back to Leap Forward: Self-Backtracking for Boosting Reasoning of Language ModelsarXiv 2025Dynamic Perceiver for Efficient Visual RecognitionICCV 2023 1EfficientTrain: Exploring Generalized Curriculum Learning for Training Visual BackbonesICCV 2023 13DIS-FLUX: simple and efficient multi-instance generation with DiT renderingarXiv 20253DIS: Depth-Driven Decoupled Instance Synthesis for Text-to-Image GenerationarXiv 2024MIC: Masked Image Consistency for Context-Enhanced Domain AdaptationCVPR 2023 1FlowTransformer: A Transformer Framework for Flow-based Network Intrusion Detection SystemsarXiv 2023Pard: Permutation-Invariant Autoregressive Diffusion for Graph GenerationarXiv 2024AI Hospital: Benchmarking Large Language Models in a Multi-agent Medical Interaction SimulatorarXiv 2024Improving CLIP Training with Language Rewritesimproving-clip-training-with-languagePerception-R1: Pioneering Perception Policy with Reinforcement LearningarXiv 2025LiDAR-CS Dataset: LiDAR Point Cloud Dataset with Cross-Sensors for 3D Object DetectionarXiv 2023FBI-LLM: Scaling Up Fully Binarized LLMs from Scratch via Autoregressive DistillationarXiv 2024Solving Linear Inverse Problems Provably via Posterior Sampling with Latent Diffusion Modelssolving-linear-inverse-problems-provably-viaCan Knowledge Editing Really Correct Hallucinations?arXiv 2024Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!arXiv 2023Enhancing Document-level Event Argument Extraction with Contextual Clues and Role Relevanceenhancing-document-level-event-argumentScaling Laws for Data Filtering -- Data Curation cannot be Compute AgnosticarXiv 2024VoCo: A Simple-yet-Effective Volume Contrastive Learning Framework for 3D Medical Image AnalysisCVPR 2024 1Text2Room: Extracting Textured 3D Meshes from 2D Text-to-Image ModelsICCV 2023 1Step-level Value Preference Optimization for Mathematical ReasoningarXiv 2024Lightweight and High-Fidelity End-to-End Text-to-Speech with Multi-Band Generation and Inverse Short-Time Fourier TransformarXiv 2022GANs N' Roses: Stable, Controllable, Diverse Image to Image Translation (works for videos too!)arXiv 2021CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-SpeecharXiv 2025Few-Shot Physically-Aware Articulated Mesh Generation via Hierarchical DeformationICCV 2023 1MetaMath: Bootstrap Your Own Mathematical Questions for Large Language ModelsarXiv 2023Efficient Pipeline for Camera Trap Image ReviewarXiv 2019GIT: A Generative Image-to-text Transformer for Vision and LanguagearXiv 2022