All papers
Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.
JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior SynchronizationarXiv 2025H3WB: Human3.6M 3D WholeBody Dataset and BenchmarkICCV 2023 1Video-T1: Test-Time Scaling for Video GenerationICCV 2025Mega: Moving Average Equipped Gated AttentionarXiv 2022Deep Entity Matching with Pre-Trained Language ModelsarXiv 2020COLD: A Benchmark for Chinese Offensive Language DetectionarXiv 20223D Scene Graph: A Structure for Unified Semantics, 3D Space, and Camera3d-scene-graph-a-structure-for-unified-1DiffSketcher: Text Guided Vector Sketch Synthesis through Latent Diffusion Modelsdiffsketcher-text-guided-vector-sketchRealmDreamer: Text-Driven 3D Scene Generation with Inpainting and Depth DiffusionarXiv 2024SpinNet: Learning a General Surface Descriptor for 3D Point Cloud RegistrationCVPR 2021 1On the Road with GPT-4V(ision): Early Explorations of Visual-Language Model on Autonomous DrivingarXiv 2023HexPlane: A Fast Representation for Dynamic ScenesCVPR 2023 1T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by SteparXiv 2023Map-free Visual Relocalization: Metric Pose Relative to a Single ImagearXiv 2022GS2Mesh: Surface Reconstruction from Gaussian Splatting via Novel Stereo ViewsarXiv 2024Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and BenchmarksarXiv 2023Early-Learning Regularization Prevents Memorization of Noisy LabelsNeurIPS 2020 12Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning TasksarXiv 2022RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human FeedbackCVPR 2024 1TAPEX: Table Pre-training via Learning a Neural SQL Executortapex-table-pre-training-via-learning-a-1ExpertPrompting: Instructing Large Language Models to be Distinguished ExpertsarXiv 2023Towards Interpretable Mental Health Analysis with Large Language ModelsarXiv 2023Refining activation downsampling with SoftPoolICCV 2021 10CNOS: A Strong Baseline for CAD-based Novel Object SegmentationarXiv 2023Music Understanding LLaMA: Advancing Text-to-Music Generation with Question Answering and CaptioningarXiv 2023Open-Set Recognition: a Good Closed-Set Classifier is All You Need?open-set-recognition-a-good-closed-setBridging Different Language Models and Generative Vision Models for Text-to-Image GenerationarXiv 2024FreeReg: Image-to-Point Cloud Registration Leveraging Pretrained Diffusion Models and Monocular Depth EstimatorsarXiv 2023DRAMA: Diverse Augmentation from Large Language Models to Smaller Dense RetrieversarXiv 2025DiffCSE: Difference-based Contrastive Learning for Sentence EmbeddingsNAACL 2022 7Quantum advantage in learning from experimentsarXiv 2021MoH: Multi-Head Attention as Mixture-of-Head AttentionarXiv 2024Towards Emotional Support Dialog SystemsACL 2021 5Joint Unsupervised Learning of Deep Representations and Image Clustersjoint-unsupervised-learning-of-deep-1MPNet: Masked and Permuted Pre-training for Language UnderstandingNeurIPS 2020 12Unleashing Vecset Diffusion Model for Fast Shape GenerationICCV 2025Tetra-NeRF: Representing Neural Radiance Fields Using TetrahedraICCV 2023 1Generalized Zero- and Few-Shot Learning via Aligned Variational AutoencodersarXiv 2018unarXive 2022: All arXiv Publications Pre-Processed for NLP, Including Structured Full-Text and Citation NetworkarXiv 2023DiLu: A Knowledge-Driven Approach to Autonomous Driving with Large Language ModelsarXiv 2023IGEV++: Iterative Multi-range Geometry Encoding Volumes for Stereo MatchingarXiv 2024SRFormerV2: Taking a Closer Look at Permuted Self-Attention for Image Super-ResolutionICCV 2023 1BizGen: Advancing Article-level Visual Text Rendering for Infographics GenerationCVPR 2025 1Cinemo: Consistent and Controllable Image Animation with Motion Diffusion ModelsarXiv 2024Efficient Emotional Adaptation for Audio-Driven Talking-Head GenerationICCV 2023 1MRQA 2019 Shared Task: Evaluating Generalization in Reading Comprehensionmrqa-2019-shared-task-evaluating-1Large Language Models Play StarCraft II: Benchmarks and A Chain of Summarization ApproacharXiv 2023Human Preference Score: Better Aligning Text-to-Image Models with Human PreferenceICCV 2023 1SUQL: Conversational Search over Structured and Unstructured Data with Large Language ModelsarXiv 2023iDisc: Internal Discretization for Monocular Depth EstimationCVPR 2023 1Simplifying Transformer BlocksarXiv 2023MesoNet: a Compact Facial Video Forgery Detection NetworkarXiv 2018The RoboDepth Challenge: Methods and Advancements Towards Robust Depth EstimationarXiv 2023Octopus: Embodied Vision-Language Programmer from Environmental FeedbackarXiv 2023Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction TuningarXiv 2023Continual Pre-training of Language ModelsarXiv 2023Graph Chain-of-Thought: Augmenting Large Language Models by Reasoning on GraphsarXiv 2024Hierarchical Open-vocabulary Universal Image Segmentationhierarchical-open-vocabulary-universal-imageOptimal Transport Aggregation for Visual Place RecognitionCVPR 2024 1SparseNeRF: Distilling Depth Ranking for Few-shot Novel View SynthesisICCV 2023 1A Thorough Examination of the CNN/Daily Mail Reading Comprehension Taska-thorough-examination-of-the-cnndaily-mail-1Fast Vision Transformers with HiLo AttentionarXiv 2022LayoutDM: Discrete Diffusion Model for Controllable Layout GenerationCVPR 2023 1RS5M and GeoRSCLIP: A Large Scale Vision-Language Dataset and A Large Vision-Language Model for Remote SensingarXiv 2023Unfolding the Headline: Iterative Self-Questioning for News Retrieval and Timeline SummarizationarXiv 2025Image Processing Using Multi-Code GAN Priorimage-processing-using-multi-code-gan-prior-1CarDreamer: Open-Source Learning Platform for World Model based Autonomous DrivingarXiv 2024Retrieval Augmented Generation and Understanding in Vision: A Survey and New OutlookarXiv 2025Sparse Autoencoders Find Highly Interpretable Features in Language ModelsarXiv 2023RenderIH: A Large-scale Synthetic Dataset for 3D Interacting Hand Pose EstimationICCV 2023 1Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task AgentsarXiv 2023Generate rather than Retrieve: Large Language Models are Strong Context GeneratorsarXiv 2022PlanT: Explainable Planning Transformers via Object-Level RepresentationsarXiv 2022Mol-Instructions: A Large-Scale Biomolecular Instruction Dataset for Large Language ModelsarXiv 2023ZipNN: Lossless Compression for AI ModelsarXiv 2024DRL-VO: Learning to Navigate Through Crowded Dynamic Scenes Using Velocity ObstaclesarXiv 2023Self-supervised Co-training for Video Representation LearningNeurIPS 2020 12Visual Speech-Aware Perceptual 3D Facial Expression Reconstruction from VideosarXiv 2022Modeling Multi-turn Conversation with Deep Utterance Aggregationmodeling-multi-turn-conversation-with-deep-2When and why vision-language models behave like bags-of-words, and what to do about it?arXiv 2022VideoMind: A Chain-of-LoRA Agent for Long Video ReasoningarXiv 2025Jacobian Descent for Multi-Objective OptimizationarXiv 2024Scenimefy: Learning to Craft Anime Scene via Semi-Supervised Image-to-Image TranslationICCV 2023 1AnyCam: Learning to Recover Camera Poses and Intrinsics from Casual VideosCVPR 2025 1Exploring Visual Prompts for Adapting Large-Scale ModelsarXiv 2022Kiss3DGen: Repurposing Image Diffusion Models for 3D Asset GenerationCVPR 2025 1Towards Realistic Scene Generation with LiDAR Diffusion ModelsarXiv 2024PreRoutGNN for Timing Prediction with Order Preserving Partition: Global Circuit Pre-training, Local Delay Learning and Attentional Cell ModelingarXiv 2024SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language ModelsarXiv 2024FALCON: Fast Autonomous Aerial Exploration using Coverage Path GuidancearXiv 2024Long Video Generation with Time-Agnostic VQGAN and Time-Sensitive TransformerarXiv 2022GridMask Data AugmentationarXiv 2020VoiceFixer: Toward General Speech Restoration with Neural VocoderarXiv 2021You Only Segment Once: Towards Real-Time Panoptic SegmentationCVPR 2023 1Lightplane: Highly-Scalable Components for Neural 3D FieldsarXiv 2024MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector QuantizationarXiv 2025Long-Context Autoregressive Video Modeling with Next-Frame Predictionlong-context-autoregressive-video-modelingUni-Perceiver v2: A Generalist Model for Large-Scale Vision and Vision-Language TasksCVPR 2023 1TSIT: A Simple and Versatile Framework for Image-to-Image TranslationECCV 2020 8StyleCLIPDraw: Coupling Content and Style in Text-to-Drawing TranslationarXiv 2022ToolQA: A Dataset for LLM Question Answering with External Toolstoolqa-a-dataset-for-llm-question-answeringGRF: Learning a General Radiance Field for 3D Representation and Renderinggrf-learning-a-general-radiance-field-for-3dPointCLIP V2: Prompting CLIP and GPT for Powerful 3D Open-world LearningICCV 2023 1Masked Image Training for Generalizable Deep Image DenoisingCVPR 2023 1PixelFlow: Pixel-Space Generative Models with FlowarXiv 2025Contrastive Audio-Visual Masked AutoencoderarXiv 2022NU-Wave: A Diffusion Probabilistic Model for Neural Audio UpsamplingarXiv 2021PySAD: A Streaming Anomaly Detection Framework in PythonarXiv 2020Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion and AviationarXiv 2024Alphazero-like Tree-Search can Guide Large Language Model Decoding and TrainingarXiv 2023DressCode: Autoregressively Sewing and Generating Garments from Text GuidancearXiv 2024Codec-SUPERB: An In-Depth Analysis of Sound Codec ModelsarXiv 2024Self-regulating Prompts: Foundational Model Adaptation without ForgettingICCV 2023 1All in One: Exploring Unified Video-Language Pre-trainingCVPR 2023 1Hi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion ModelsarXiv 2024DEADiff: An Efficient Stylization Diffusion Model with Disentangled RepresentationsCVPR 2024 1Lossless data compression by large modelsarXiv 2024LAMP: Learn A Motion Pattern for Few-Shot-Based Video GenerationarXiv 2023User-Controllable Latent Transformer for StyleGAN Image Layout EditingarXiv 2022VCoder: Versatile Vision Encoders for Multimodal Large Language ModelsCVPR 2024 1SummerTime: Text Summarization Toolkit for Non-expertsEMNLP (ACL) 2021 113DGS-LM: Faster Gaussian-Splatting Optimization with Levenberg-MarquardtarXiv 2024Towards Better Dynamic Graph Learning: New Architecture and Unified Librarytowards-better-dynamic-graph-learning-newA Survey on Detection of LLMs-Generated ContentarXiv 2023GaussianHead: High-fidelity Head Avatars with Learnable Gaussian DerivationarXiv 2023InterDiff: Generating 3D Human-Object Interactions with Physics-Informed DiffusionICCV 2023 1Loc-NeRF: Monte Carlo Localization using Neural Radiance FieldsarXiv 2022Real-Time Drone Detection and Tracking With Visible, Thermal and
Acoustic SensorsarXiv 2020REC-MV: REconstructing 3D Dynamic Cloth from Monocular Videosrec-mv-reconstructing-3d-dynamic-cloth-fromExample-based Motion Synthesis via Generative Motion MatchingarXiv 2023QMoE: Practical Sub-1-Bit Compression of Trillion-Parameter ModelsarXiv 2023ZeroCap: Zero-Shot Image-to-Text Generation for Visual-Semantic ArithmeticCVPR 2022 1ZeroQ: A Novel Zero Shot Quantization Frameworkzeroq-a-novel-zero-shot-quantization-1SceneSplat: Gaussian Splatting-based Scene Understanding with Vision-Language PretrainingICCV 2025CustomNet: Zero-shot Object Customization with Variable-Viewpoints in Text-to-Image Diffusion ModelsarXiv 2023BAdam: A Memory Efficient Full Parameter Optimization Method for Large Language ModelsarXiv 2024D(R,O) Grasp: A Unified Representation of Robot and Object
Interaction for Cross-Embodiment Dexterous GraspingarXiv 2024HDLTex: Hierarchical Deep Learning for Text ClassificationarXiv 2017How to build the best medical image segmentation algorithm using foundation models: a comprehensive empirical study with Segment Anything ModelarXiv 2024Online Analytic Exemplar-Free Continual Learning with Large Models for Imbalanced Autonomous Driving TaskarXiv 2024Scalable 3D Captioning with Pretrained ModelsNeurIPS 2023 11Prophet: Prompting Large Language Models with Complementary Answer Heuristics for Knowledge-based Visual Question AnsweringCVPR 2023 1Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agentslanguage-models-as-zero-shot-plannersCan 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time ScalingarXiv 2025A Review of Large Language Models and Autonomous Agents in ChemistryarXiv 2024M+: Extending MemoryLLM with Scalable Long-Term MemoryarXiv 2025Guiding Language Models of Code with Global Context using MonitorsarXiv 2023BBT-Fin: Comprehensive Construction of Chinese Financial Domain Pre-trained Language Model, Corpus and BenchmarkarXiv 2023Repaint123: Fast and High-quality One Image to 3D Generation with Progressive Controllable 2D RepaintingarXiv 2023DeepSolo: Let Transformer Decoder with Explicit Points Solo for Text SpottingCVPR 2023 1Grounding Large Language Models in Interactive Environments with Online Reinforcement LearningarXiv 2023Learning Temporally Consistent Video Depth from Video Diffusion PriorsCVPR 2025 1AvatarArtist: Open-Domain 4D AvatarizationCVPR 2025 1NusaCrowd: Open Source Initiative for Indonesian NLP ResourcesarXiv 2022Contrastive Multi-View Representation Learning on GraphsICML 2020 1LongRoPE2: Near-Lossless LLM Context Window ScalingarXiv 2025Halton Scheduler For Masked Generative Image TransformerarXiv 20253D Human Mesh Estimation from Virtual Markers3d-human-mesh-estimation-from-virtual-markersDETRs with Hybrid MatchingCVPR 2023 1Simple and Accurate Dependency Parsing Using Bidirectional LSTM Feature Representationssimple-and-accurate-dependency-parsing-using-1Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language ModelarXiv 2024AutoKaggle: A Multi-Agent Framework for Autonomous Data Science CompetitionsarXiv 2024Dice Loss for Data-imbalanced NLP Tasksdice-loss-for-data-imbalanced-nlp-tasks-1Evaluation of Text-to-Video Generation Models: A Dynamics PerspectivearXiv 2024UniPortrait: A Unified Framework for Identity-Preserving Single- and Multi-Human Image PersonalizationarXiv 2024STPLS3D: A Large-Scale Synthetic and Real Aerial Photogrammetry 3D Point Cloud DatasetarXiv 2022SSL4EO-S12: A Large-Scale Multi-Modal, Multi-Temporal Dataset for Self-Supervised Learning in Earth ObservationarXiv 2022AUITestAgent: Automatic Requirements Oriented GUI Function TestingarXiv 2024Towards Building Multilingual Language Model for MedicinearXiv 2024Sparse Mixture-of-Experts are Domain Generalizable LearnersarXiv 2022Embodied Agent Interface: Benchmarking LLMs for Embodied Decision MakingarXiv 2024Residual Flows for Invertible Generative Modelingresidual-flows-for-invertible-generative-2Class-Incremental Learning: A SurveyarXiv 2023BBTv2: Towards a Gradient-Free Future with Large Language ModelsarXiv 2022Long-tailed Recognition by Routing Diverse Distribution-Aware Expertslong-tailed-recognition-by-routing-diverseAutomated Movie Generation via Multi-Agent CoT PlanningarXiv 2025E5-V: Universal Embeddings with Multimodal Large Language ModelsarXiv 2024Space-Time Correspondence as a Contrastive Random WalkNeurIPS 2020 12AndroidLab: Training and Systematic Benchmarking of Android Autonomous AgentsarXiv 2024RESDSQL: Decoupling Schema Linking and Skeleton Parsing for Text-to-SQLarXiv 2023Residual Kolmogorov-Arnold Network for Enhanced Deep LearningarXiv 2024Retiring Adult: New Datasets for Fair Machine LearningNeurIPS 2021 12Pix2NeRF: Unsupervised Conditional $π$-GAN for Single Image to Neural Radiance Fields TranslationarXiv 2022A Tutorial on Bayesian OptimizationarXiv 2018HOPE: A Reinforcement Learning-based Hybrid Policy Path Planner for Diverse Parking ScenariosarXiv 2024Neural Body Fitting: Unifying Deep Learning and Model-Based Human Pose and Shape EstimationarXiv 2018Cam4DOcc: Benchmark for Camera-Only 4D Occupancy Forecasting in Autonomous Driving ApplicationsarXiv 2023Human-Art: A Versatile Human-Centric Dataset Bridging Natural and Artificial ScenesCVPR 2023 1BoxDiff: Text-to-Image Synthesis with Training-Free Box-Constrained DiffusionICCV 2023 1Baichuan-Omni Technical ReportarXiv 2024Non-Autoregressive Neural Machine Translationnon-autoregressive-neural-machine-translation-2VPGTrans: Transfer Visual Prompt Generator across LLMsNeurIPS 2023 11Reactive Diffusion Policy: Slow-Fast Visual-Tactile Policy Learning for Contact-Rich ManipulationarXiv 2025Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn'tarXiv 2025Using Self-Supervised Learning Can Improve Model Robustness and Uncertaintyusing-self-supervised-learning-can-improve-1Sigma: Siamese Mamba Network for Multi-Modal Semantic SegmentationarXiv 2024Packed Levitated Marker for Entity and Relation ExtractionACL 2022 5Every Activation Boosted: Scaling General Reasoner to 1 Trillion Open
Language FoundationarXiv 2025VARGPT-v1.1: Improve Visual Autoregressive Large Unified Model via Iterative Instruction Tuning and Reinforcement LearningarXiv 2025Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive SurveyarXiv 2023