All papers
Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.
Hard-aware Instance Adaptive Self-training for Unsupervised Cross-domain Semantic SegmentationarXiv 2023Diffusion Language Models Are Versatile Protein LearnersarXiv 2024Parameters vs. Context: Fine-Grained Control of Knowledge Reliance in Language ModelsarXiv 2025FluidNexus: 3D Fluid Reconstruction and Prediction from a Single VideoCVPR 2025 1FedSkel: Efficient Federated Learning on Heterogeneous Systems with Skeleton Gradients UpdatearXiv 2021Voice Conversion With Just Nearest NeighborsarXiv 2023V1T: large-scale mouse V1 response prediction using a Vision TransformerarXiv 2023LLMVoX: Autoregressive Streaming Text-to-Speech Model for Any LLMarXiv 2025GLA-GCN: Global-local Adaptive Graph Convolutional Network for 3D Human Pose Estimation from Monocular VideoICCV 2023 1Unleashing Hour-Scale Video Training for Long Video-Language UnderstandingarXiv 2025Kinetics: Rethinking Test-Time Scaling LawsarXiv 2025EMBER2024 -- A Benchmark Dataset for Holistic Evaluation of Malware ClassifiersarXiv 2025A Smooth Sea Never Made a Skilled $\texttt{SAILOR}$: Robust Imitation via Learning to SearcharXiv 2025CAD-GPT: Synthesising CAD Construction Sequence with Spatial Reasoning-Enhanced Multimodal LLMsarXiv 2024Effective Data Augmentation With Diffusion ModelsarXiv 2023Domain Adaptation Through Task Distillationdomain-adaptation-through-task-distillation-1AR-Diffusion: Asynchronous Video Generation with Auto-Regressive DiffusionCVPR 2025 1SemanticFormer: Holistic and Semantic Traffic Scene Representation for Trajectory Prediction using Knowledge GraphsarXiv 2024Divide & Bind Your Attention for Improved Generative Semantic NursingarXiv 2023Rewriting Pre-Training Data Boosts LLM Performance in Math and CodearXiv 2025Topological AutoencodersICML 2020 1Topological Graph Neural Networkstopological-graph-neural-networks-1AutoCast++: Enhancing World Event Prediction with Zero-shot Ranking-based Context RetrievalarXiv 2023GarmentDreamer: 3DGS Guided Garment Synthesis with Diverse Geometry and Texture DetailsarXiv 2024BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-HaystackarXiv 2024MathCoder-VL: Bridging Vision and Code for Enhanced Multimodal Mathematical ReasoningarXiv 2025Hierarchical Document Refinement for Long-context Retrieval-augmented GenerationarXiv 2025Context-Alignment: Activating and Enhancing LLM Capabilities in Time SeriesarXiv 2025Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image SynthesisCVPR 2025 1Lightweight Image Inpainting by Stripe Window Transformer with Joint Attention to CNNarXiv 2023Fusion is Not Enough: Single Modal Attacks on Fusion Models for 3D Object DetectionarXiv 2023Cross-Tokenizer Distillation via Approximate Likelihood MatchingarXiv 2025Vista4D: Video Reshooting with 4D Point CloudsarXiv 2026Learning Latent Dynamic Robust Representations for World ModelsarXiv 2024Towards A Generalizable Pathology Foundation Model via Unified Knowledge DistillationarXiv 2024SeRL: Self-Play Reinforcement Learning for Large Language Models with Limited DataarXiv 2025MolFM: A Multimodal Molecular Foundation ModelarXiv 2023AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge BasesarXiv 2024What Language Model to Train if You Have One Million GPU Hours?arXiv 2022RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time VideoarXiv 2025SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agentsarXiv 2025OctoPack: Instruction Tuning Code Large Language ModelsarXiv 2023An Efficient Recipe for Long Context Extension via Middle-Focused Positional EncodingarXiv 2024Masked Spiking TransformerICCV 2023 1PMIndia -- A Collection of Parallel Corpora of Languages of IndiaarXiv 2020When Do We Not Need Larger Vision Models?arXiv 2024GENIUS: Sketch-based Language Model Pre-training via Extreme and Selective Masking for Text Generation and AugmentationarXiv 2022EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World ModelsarXiv 2025Vamba: Understanding Hour-Long Videos with Hybrid Mamba-TransformersICCV 2025Model Internals-based Answer Attribution for Trustworthy Retrieval-Augmented GenerationarXiv 2024Why think step by step? Reasoning emerges from the locality of experiencewhy-think-step-by-step-reasoning-emerges-fromGuacaMol: Benchmarking Models for De Novo Molecular DesignarXiv 2018SPIN-Bench: How Well Do LLMs Plan Strategically and Reason Socially?arXiv 2025Sample Efficient Preference Alignment in LLMs via Active ExplorationarXiv 2023BLADE: Benchmarking Language Model Agents for Data-Driven SciencearXiv 2024macOSWorld: A Multilingual Interactive Benchmark for GUI AgentsarXiv 2025Theia: Distilling Diverse Vision Foundation Models for Robot LearningarXiv 2024Taming the Power of Diffusion Models for High-Quality Virtual Try-On with Appearance FlowarXiv 2023Making Images Real Again: A Comprehensive Survey on Deep Image CompositionarXiv 2021FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language ModelsarXiv 2023ProtoECGNet: Case-Based Interpretable Deep Learning for Multi-Label ECG Classification with Contrastive LearningarXiv 2025Neural Spline Flowsneural-spline-flows-1BeLFusion: Latent Diffusion for Behavior-Driven Human Motion PredictionICCV 2023 1Planetarium: A Rigorous Benchmark for Translating Text to Structured Planning LanguagesarXiv 2024Recoverable Compression: A Multimodal Vision Token Recovery Mechanism Guided by Text InformationarXiv 2024Dynamic Early Exit in Reasoning ModelsarXiv 2025Guiding Masked Representation Learning to Capture Spatio-Temporal Relationship of ElectrocardiogramarXiv 2024Emotional RAG: Enhancing Role-Playing Agents through Emotional RetrievalarXiv 2024Baichuan 2: Open Large-scale Language ModelsarXiv 2023EHRXQA: A Multi-Modal Question Answering Dataset for Electronic Health Records with Chest X-ray Imagesehrxqa-a-multi-modal-question-answeringRethinking Inductive Biases for Surface Normal EstimationCVPR 2024 1Simul-Whisper: Attention-Guided Streaming Whisper with Truncation DetectionarXiv 2024Capacity, Bandwidth, and Compositionality in Emergent Language LearningarXiv 2019Scalable Best-of-N Selection for Large Language Models via Self-CertaintyarXiv 2025FENICE: Factuality Evaluation of summarization based on Natural language Inference and Claim ExtractionarXiv 2024You See it, You Got it: Learning 3D Creation on Pose-Free Videos at ScaleCVPR 2025 1ALERT: A Comprehensive Benchmark for Assessing Large Language Models' Safety through Red TeamingarXiv 2024SegGPT: Segmenting Everything In ContextarXiv 2023Efficient Multimodal Learning from Data-centric PerspectivearXiv 2024HYTREL: Hypergraph-enhanced Tabular Data Representation Learninghytrel-hypergraph-enhanced-tabular-data-1Diffusion Models for Molecules: A Survey of Methods and TasksarXiv 2025Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI BenchmarkarXiv 2023Active Learning Through a Covering LensarXiv 2022Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary SpacearXiv 2022Parting with Misconceptions about Learning-based Vehicle Motion PlanningarXiv 2023EquiformerV2: Improved Equivariant Transformer for Scaling to Higher-Degree RepresentationsarXiv 2023SpeechCLIP: Integrating Speech with Pre-Trained Vision and Language ModelarXiv 2022FLoRA: Federated Fine-Tuning Large Language Models with Heterogeneous Low-Rank AdaptationsarXiv 2024Unlocking State-Tracking in Linear RNNs Through Negative EigenvaluesarXiv 2024SLEDGE: Synthesizing Driving Environments with Generative Models and Rule-Based TrafficarXiv 2024SoundMind: RL-Incentivized Logic Reasoning for Audio-Language ModelsarXiv 2025ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional DependenciesarXiv 2025Make-A-Shape: a Ten-Million-scale 3D Shape ModelarXiv 2024MM-Verify: Enhancing Multimodal Reasoning with Chain-of-Thought VerificationarXiv 2025DeciMamba: Exploring the Length Extrapolation Potential of MambaarXiv 2024Towards Robust Fidelity for Evaluating Explainability of Graph Neural NetworksarXiv 2023Fine-tune the pretrained ATST model for sound event detectionarXiv 2023Self-supervised Audio Teacher-Student Transformer for Both Clip-level and Frame-level TasksarXiv 2023Scaling physics-informed hard constraints with mixture-of-expertsarXiv 2024UniGaze: Towards Universal Gaze Estimation via Large-scale Pre-TrainingarXiv 2025WavJourney: Compositional Audio Creation with Large Language ModelsarXiv 2023P2P-Bridge: Diffusion Bridges for 3D Point Cloud DenoisingarXiv 2024MDETR -- Modulated Detection for End-to-End Multi-Modal UnderstandingarXiv 2021Delicate Textured Mesh Recovery from NeRF via Adaptive Surface RefinementICCV 2023 1Pretraining Codomain Attention Neural Operators for Solving Multiphysics PDEsarXiv 2024Context Autoencoder for Self-Supervised Representation LearningarXiv 2022MARFT: Multi-Agent Reinforcement Fine-TuningarXiv 2025Named Entity Recognition in Twitter: A Dataset and Analysis on Short-Term Temporal ShiftsarXiv 2022CausalGym: Benchmarking causal interpretability methods on linguistic tasksarXiv 2024Can Language Models Solve Graph Problems in Natural Language?NeurIPS 2023 11OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video GenerationarXiv 2026FreeTraj: Tuning-Free Trajectory Control in Video Diffusion ModelsarXiv 2024FonTS: Text Rendering with Typography and Style ControlsICCV 2025CODESIM: Multi-Agent Code Generation and Problem Solving through Simulation-Driven Planning and Debuggingcodesim-multi-agent-code-generation-andLong-range UAV Thermal Geo-localization with Satellite ImageryarXiv 2023Can LLMs Clean Up Your Mess? A Survey of Application-Ready Data Preparation with LLMsarXiv 2026LLMs Improving LLMs: Agentic Discovery for Test-Time ScalingarXiv 2026From Storage to Experience: A Survey on the Evolution of LLM Agent Memory MechanismsarXiv 2026SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic SegmentationarXiv 2022Warp-as-History: Generalizable Camera-Controlled Video Generation from One Training VideoarXiv 2026Collapsible Linear Blocks for Super-Efficient Super ResolutionarXiv 2021ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive FeedbackarXiv 2025PatientSim: A Persona-Driven Simulator for Realistic Doctor-Patient InteractionsarXiv 2025The Curious Case of Neural Text DegenerationICLR 2020 1Aria-UI: Visual Grounding for GUI InstructionsarXiv 2024Do Large Language Models Perform Latent Multi-Hop Reasoning without Exploiting Shortcuts?arXiv 2024TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion ModelsarXiv 2025ORV: 4D Occupancy-centric Robot Video GenerationarXiv 2025Region-Adaptive Sampling for Diffusion TransformersarXiv 2025SVGenius: Benchmarking LLMs in SVG Understanding, Editing and GenerationarXiv 2025Injecting Domain-Specific Knowledge into Large Language Models: A Comprehensive SurveyarXiv 2025TimeCAP: Learning to Contextualize, Augment, and Predict Time Series Events with Large Language Model AgentsarXiv 2025ECoDepth: Effective Conditioning of Diffusion Models for Monocular Depth EstimationCVPR 2024 1Probabilistic Attention for Interactive SegmentationNeurIPS 2021 12AniMer: Animal Pose and Shape Estimation Using Family Aware TransformerCVPR 2025 1Probing the Multi-turn Planning Capabilities of LLMs via 20 Question GamesarXiv 2023Switti: Designing Scale-Wise Transformers for Text-to-Image SynthesisarXiv 2024CheXWorld: Exploring Image World Modeling for Radiograph Representation LearningCVPR 2025 1FedCompass: Efficient Cross-Silo Federated Learning on Heterogeneous Client Devices using a Computing Power Aware SchedulerarXiv 2023KV Prediction for Improved Time to First TokenarXiv 2024dMel: Speech Tokenization made SimplearXiv 2024The AdEMAMix Optimizer: Better, Faster, OlderarXiv 2024Identifying Functionally Important Features with End-to-End Sparse Dictionary LearningarXiv 2024LLM as Effective Streaming Processor: Bridging Streaming-Batch Mismatches with Group Position EncodingarXiv 2025WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement LearningarXiv 2025SWE-Search: Enhancing Software Agents with Monte Carlo Tree Search and Iterative RefinementarXiv 2024TubeDETR: Spatio-Temporal Video Grounding with TransformersCVPR 2022 1Zero-Shot Video Question Answering via Frozen Bidirectional Language ModelsarXiv 2022CROMA: Remote Sensing Representations with Contrastive Radar-Optical Masked Autoencoderscroma-remote-sensing-representations-withMindGames: Targeting Theory of Mind in Large Language Models with Dynamic Epistemic Modal LogicarXiv 2023Length Generalization of Causal Transformers without Position EncodingarXiv 2024FAST-RIR: Fast neural diffuse room impulse response generatorarXiv 2021Toy Models of SuperpositionarXiv 2022DesignQA: A Multimodal Benchmark for Evaluating Large Language Models' Understanding of Engineering DocumentationarXiv 2024Axiomatic Attribution for Deep Networksaxiomatic-attribution-for-deep-networks-1Layer-wise Analysis of a Self-supervised Speech Representation ModelarXiv 2021Handwriting TransformersICCV 2021 10Forecasting Future World Events with Neural NetworksarXiv 2022Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language ModelsarXiv 2024Post-Training Sparse Attention with Double SparsityarXiv 2024Disentangled Diffusion-Based 3D Human Pose Estimation with Hierarchical Spatial and Temporal DenoiserarXiv 2024Post-pre-training for Modality Alignment in Vision-Language Foundation ModelsCVPR 2025 1Mask Image WatermarkingarXiv 2025MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM TutorsarXiv 2025Lawyer LLaMA Technical ReportarXiv 2023Shape Preserving Facial Landmarks with Graph Attention NetworksarXiv 2022Neural Arithmetic UnitsICLR 2020 1Doc2Graph: a Task Agnostic Document Understanding Framework based on Graph Neural NetworksarXiv 2022Few-Shot Bot: Prompt-Based Learning for Dialogue SystemsarXiv 2021Multi-Behavior Generative RecommendationarXiv 2024GroupMamba: Efficient Group-Based Visual State Space ModelCVPR 2025 1What's "up" with vision-language models? Investigating their struggle with spatial reasoningarXiv 2023ADAHESSIAN: An Adaptive Second Order Optimizer for Machine LearningarXiv 2020Visual Prompting via Image InpaintingarXiv 2022Data Shapley: Equitable Valuation of Data for Machine LearningarXiv 2019Open-YOLO 3D: Towards Fast and Accurate Open-Vocabulary 3D Instance SegmentationarXiv 2024The Hidden Attention of Mamba ModelsarXiv 2024Predict, Refine, Synthesize: Self-Guiding Diffusion Models for Probabilistic Time Series ForecastingNeurIPS 2023 11ReCode: Robustness Evaluation of Code Generation ModelsarXiv 2022GLEAN: Generalized Category Discovery with Diverse and Quality-Enhanced LLM FeedbackarXiv 2025Coarse-to-Fine Amodal Segmentation with Shape PriorICCV 2023 1PEAVS: Perceptual Evaluation of Audio-Visual Synchrony Grounded in Viewers' Opinion ScoresarXiv 2024CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction FollowingarXiv 2025TIIF-Bench: How Does Your T2I Model Follow Your Instructions?arXiv 2025MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document RetrievalarXiv 2025Automatic Evaluation and Analysis of Idioms in Neural Machine TranslationarXiv 2022GSV-Cities: Toward Appropriate Supervised Visual Place RecognitionarXiv 2022Every SAM Drop Counts: Embracing Semantic Priors for Multi-Modality Image Fusion and BeyondCVPR 2025 1What's the Magic Word? A Control Theory of LLM PromptingarXiv 2023MidiCaps: A large-scale MIDI dataset with text captionsarXiv 2024Mustango: Toward Controllable Text-to-Music GenerationarXiv 2023JamendoMaxCaps: A Large Scale Music-caption Dataset with Imputed MetadataarXiv 2025DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning DataarXiv 2024FreeSplatter: Pose-free Gaussian Splatting for Sparse-view 3D ReconstructionICCV 2025P2P: Automated Paper-to-Poster Generation and Fine-Grained BenchmarkarXiv 2025GRIT: Teaching MLLMs to Think with ImagesarXiv 2025Web-Shepherd: Advancing PRMs for Reinforcing Web AgentsarXiv 2025Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation DatasetsarXiv 2025Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative PretrainingarXiv 2024OmniCaptioner: One Captioner to Rule Them AllarXiv 2025