All papers
Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.
Data Augmentation using Pre-trained Transformer ModelsAACL (lifelongnlp) 2020 12VisorGPT: Learning Visual Prior via Generative Pre-TrainingarXiv 2023Back to the Feature: Classical 3D Features are (Almost) All You Need for 3D Anomaly DetectionarXiv 2022Denoised MDPs: Learning World Models Better Than the World ItselfarXiv 2022AdaMix: Mixture-of-Adaptations for Parameter-efficient Model TuningarXiv 2022Hogwild! Inference: Parallel LLM Generation via Concurrent AttentionarXiv 2025Neeko: Leveraging Dynamic LoRA for Efficient Multi-Character Role-Playing AgentarXiv 2024MLLM-Tool: A Multimodal Large Language Model For Tool Agent LearningarXiv 2024RAG-Instruct: Boosting LLMs with Diverse Retrieval-Augmented InstructionsarXiv 2024ExpertQA: Expert-Curated Questions and Attributed AnswersarXiv 2023Scaleformer: Iterative Multi-scale Refining Transformers for Time Series
ForecastingarXiv 2022Behavior Transformers: Cloning $k$ modes with one stonearXiv 2022Are NLP Models really able to Solve Simple Math Word Problems?NAACL 2021 4HAE-RAE Bench: Evaluation of Korean Knowledge in Language ModelsarXiv 2023Evaluating Hallucinations in Chinese Large Language ModelsarXiv 2023VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and RelationsarXiv 2022Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language ModelsarXiv 2023InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated
Large Language Model AgentsarXiv 2024AutoDecoding Latent 3D Diffusion Modelsautodecoding-latent-3d-diffusion-modelsDifferentially Private Optimization on Large Model at Small CostarXiv 2022DreamLIP: Language-Image Pre-training with Long CaptionsarXiv 2024Neural Scene Flow PriorNeurIPS 2021 12Dense 3D Object Reconstruction from a Single Depth ViewarXiv 2018Efficient Test-Time Model Adaptation without ForgettingarXiv 2022GFPose: Learning 3D Human Pose Prior with Gradient FieldsCVPR 2023 1RLIPv2: Fast Scaling of Relational Language-Image Pre-trainingICCV 2023 1Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?arXiv 2024GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI NavigationarXiv 2023Probabilistic Embeddings for Cross-Modal RetrievalCVPR 2021 1Number it: Temporal Grounding Videos like Flipping MangaCVPR 2025 1CLIP-KD: An Empirical Study of CLIP Model DistillationCVPR 2024 1MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative DecodingarXiv 2024SwinJSCC: Taming Swin Transformer for Deep Joint Source-Channel CodingarXiv 2023More Agents Is All You NeedarXiv 2024Visformer: The Vision-friendly TransformerICCV 2021 10SEE-2-SOUND: Zero-Shot Spatial Environment-to-Spatial SoundarXiv 2024DreamComposer: Controllable 3D Object Generation via Multi-View ConditionsCVPR 2024 1Transformers without Tears: Improving the Normalization of Self-AttentionEMNLP (IWSLT) 2019 11Emergent Road Rules In Multi-Agent Driving Environmentsemergent-road-rules-in-multi-agent-drivingu-LLaVA: Unifying Multi-Modal Tasks via Large Language ModelarXiv 2023BlockFusion: Expandable 3D Scene Generation using Latent Tri-plane ExtrapolationarXiv 2024Continuous, Subject-Specific Attribute Control in T2I Models by Identifying Semantic DirectionsCVPR 2025 1MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction TuningarXiv 2022ScreenAI: A Vision-Language Model for UI and Infographics UnderstandingarXiv 2024PartImageNet: A Large, High-Quality Dataset of PartsarXiv 2021Grape detection, segmentation and tracking using deep neural networks and three-dimensional associationarXiv 2019Dataset Distillation via Curriculum Data Synthesis in Large Data EraarXiv 2023AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation,
Recognition and Speaker Diarization in Conference ScenarioarXiv 2021SDF-StyleGAN: Implicit SDF-Based StyleGAN for 3D Shape GenerationarXiv 2022Self-Directed Online Machine Learning for Topology OptimizationarXiv 2020CoEdIT: Text Editing by Task-Specific Instruction TuningarXiv 2023We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?arXiv 2024NaturalProofs: Mathematical Theorem Proving in Natural LanguagearXiv 2021GUICourse: From General Vision Language Models to Versatile GUI AgentsarXiv 2024Predicting Cellular Responses to Novel Drug Perturbations at a Single-Cell ResolutionarXiv 2022EgoLifter: Open-world 3D Segmentation for Egocentric PerceptionarXiv 2024Evaluating Large Language Models at Evaluating Instruction FollowingarXiv 2023LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis Evaluationllmscore-unveiling-the-power-of-largeDiff-Font: Diffusion Model for Robust One-Shot Font GenerationarXiv 2022High-Speed Motion Planning for Aerial Swarms in Unknown and Cluttered
EnvironmentsarXiv 2024Attention-Driven Dynamic Graph Convolutional Network for Multi-Label Image Recognitionattention-driven-dynamic-graph-convolutionalChain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMsarXiv 2024Trans-Encoder: Unsupervised sentence-pair modelling through self- and mutual-distillationstrans-encoder-unsupervised-sentence-pair-1Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 smallarXiv 2022Zenseact Open Dataset: A large-scale and diverse multimodal dataset for autonomous drivingICCV 2023 1Grounding Everything: Emerging Localization Properties in Vision-Language TransformersCVPR 2024 1Grounding 3D Object Affordance from 2D Interactions in ImagesICCV 2023 1Word Embeddings Are Steers for Language ModelsarXiv 2023Eliciting Human Preferences with Language ModelsarXiv 2023BBQ: A Hand-Built Bias Benchmark for Question AnsweringFindings (ACL) 2022 5Language Models Meet World Models: Embodied Experiences Enhance Language Modelslanguage-models-meet-world-models-embodiedClosed-Loop Visuomotor Control with Generative Expectation for Robotic
ManipulationarXiv 2024EcoAssistant: Using LLM Assistant More Affordably and AccuratelyarXiv 2023SSAMBA: Self-Supervised Audio Representation Learning with Mamba State Space ModelarXiv 2024LinkTransformer: A Unified Package for Record Linkage with Transformer Language ModelsarXiv 2023CDeC-Net: Composite Deformable Cascade Network for Table Detection in Document ImagesarXiv 2020Introducing Visual Perception Token into Multimodal Large Language ModelarXiv 2025Whitening for Self-Supervised Representation LearningarXiv 2020Edge-guided Multi-domain RGB-to-TIR image Translation for Training Vision Tasks with Challenging LabelsarXiv 2023CrossNER: Evaluating Cross-Domain Named Entity RecognitionarXiv 2020Smoothed Energy Guidance: Guiding Diffusion Models with Reduced Energy Curvature of AttentionarXiv 2024To See is to Believe: Prompting GPT-4V for Better Visual Instruction TuningarXiv 2023Evidential Deep Learning for Open Set Action RecognitionICCV 2021 10NNV: The Neural Network Verification Tool for Deep Neural Networks and Learning-Enabled Cyber-Physical SystemsarXiv 2020ASH: Animatable Gaussian Splats for Efficient and Photoreal Human RenderingCVPR 2024 1HarmonyView: Harmonizing Consistency and Diversity in One-Image-to-3DCVPR 2024 1Pyramid Diffusion for Fine 3D Large Scene GenerationarXiv 2023Tell Your Model Where to Attend: Post-hoc Attention Steering for LLMsarXiv 2023OpenViDial 2.0: A Larger-Scale, Open-Domain Dialogue Generation Dataset with Visual ContextsarXiv 2021Residual Pattern Learning for Pixel-wise Out-of-Distribution Detection in Semantic SegmentationICCV 2023 1Next Patch Prediction for Autoregressive Visual GenerationarXiv 2024X-Pool: Cross-Modal Language-Video Attention for Text-Video RetrievalCVPR 2022 1On Realization of Intelligent Decision-Making in the Real World: A Foundation Decision Model PerspectivearXiv 2022House price estimation from visual and textual featuresarXiv 2016WebUI: A Dataset for Enhancing Visual UI Understanding with Web
SemanticsarXiv 2023The Jazz Transformer on the Front Line: Exploring the Shortcomings of AI-composed Music through Quantitative MeasuresarXiv 2020Directed Acyclic Transformer Pre-training for High-quality Non-autoregressive Text GenerationarXiv 2023Small Models are Valuable Plug-ins for Large Language ModelsarXiv 2023MultiSports: A Multi-Person Video Dataset of Spatio-Temporally Localized Sports ActionsICCV 2021 10CLIP-VG: Self-paced Curriculum Adapting of CLIP for Visual GroundingarXiv 2023Specializing Smaller Language Models towards Multi-Step ReasoningarXiv 2023Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language ModelsarXiv 2024PLIP: Language-Image Pre-training for Person Representation LearningarXiv 2023Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning StepsCOLING 2020 8Multi-Scale Representations by Varying Window Attention for Semantic SegmentationarXiv 2024SimGRAG: Leveraging Similar Subgraphs for Knowledge Graphs Driven Retrieval-Augmented GenerationarXiv 2024CR-FIQA: Face Image Quality Assessment by Learning Sample Relative ClassifiabilityCVPR 2023 1RectifID: Personalizing Rectified Flow with Anchored Classifier GuidancearXiv 2024AeroGen: Enhancing Remote Sensing Object Detection with Diffusion-Driven Data GenerationCVPR 2025 1Multimodal Analogical Reasoning over Knowledge GraphsarXiv 2022Data Movement Is All You Need: A Case Study on Optimizing TransformersarXiv 2020Differentiable Point-Based Radiance Fields for Efficient View SynthesisarXiv 2022EDA: Explicit Text-Decoupling and Dense Alignment for 3D Visual GroundingCVPR 2023 1Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clustersarXiv 2024Eliciting In-Context Learning in Vision-Language Models for Videos Through Curated Data Distributional PropertiesarXiv 2023LVBench: An Extreme Long Video Understanding BenchmarkICCV 2025DVIS++: Improved Decoupled Framework for Universal Video SegmentationarXiv 2023DACS: Domain Adaptation via Cross-domain Mixed SamplingarXiv 2020DACS: Domain Adaptation via Cross-domain Mixed SamplingarXiv 2020GTA: A Benchmark for General Tool AgentsarXiv 2024Grounded 3D-LLM with Referent TokensarXiv 2024GlyphDraw: Seamlessly Rendering Text with Intricate Spatial Structures in Text-to-Image GenerationarXiv 2023P2P: Tuning Pre-trained Image Models for Point Cloud Analysis with Point-to-Pixel PromptingarXiv 2022LIV: Language-Image Representations and Rewards for Robotic ControlarXiv 2023DPM-Solver-v3: Improved Diffusion ODE Solver with Empirical Model Statisticsdpm-solver-v3-improved-diffusion-ode-solverLearning to Act without ActionsarXiv 2023GaussianWorld: Gaussian World Model for Streaming 3D Occupancy PredictionCVPR 2025 1SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language UnderstandingarXiv 2024NeoRL: A Near Real-World Benchmark for Offline Reinforcement LearningarXiv 2021PUMA: Empowering Unified MLLM with Multi-granular Visual GenerationICCV 2025SnapMix: Semantically Proportional Mixing for Augmenting Fine-grained DataarXiv 2020Compacter: Efficient Low-Rank Hypercomplex Adapter LayersNeurIPS 2021 12Tutela: An Open-Source Tool for Assessing User-Privacy on Ethereum and Tornado CasharXiv 2022PPLLaVA: Varied Video Sequence Understanding With Prompt GuidancearXiv 2024SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language ModelsarXiv 2023Attention-based Point Cloud Edge SamplingCVPR 2023 1Deep Pyramidal Residual Networksdeep-pyramidal-residual-networks-1Sequential Latent Knowledge Selection for Knowledge-Grounded DialogueICLR 2020 1High Fidelity Speech Synthesis with Adversarial NetworksICLR 2020 1GooAQ: Open Question Answering with Diverse Answer TypesFindings (EMNLP) 2021 11OakInk: A Large-scale Knowledge Repository for Understanding Hand-Object InteractionCVPR 2022 1Coarse-to-Fine Vision-Language Pre-training with Fusion in the Backbonecoarse-to-fine-vision-language-pre-training-13D Object Reconstruction from a Single Depth View with Adversarial LearningarXiv 2017V-DETR: DETR with Vertex Relative Position Encoding for 3D Object DetectionarXiv 2023Artistic Glyph Image Synthesis via One-Stage Few-Shot LearningarXiv 2019A Survey of Large Language Models for Healthcare: from Data, Technology, and Applications to Accountability and EthicsarXiv 2023Harmonizing Visual Text Comprehension and GenerationarXiv 2024Graph Inductive Biases in Transformers without Message PassingarXiv 2023Describing Differences in Image Sets with Natural LanguageCVPR 2024 1TOD3Cap: Towards 3D Dense Captioning in Outdoor ScenesarXiv 2024Diverse and Aligned Audio-to-Video Generation via Text-to-Video Model AdaptationarXiv 2023Agent-Pro: Learning to Evolve via Policy-Level Reflection and OptimizationarXiv 2024Graphic Design with Large Multimodal ModelarXiv 2024SHViT: Single-Head Vision Transformer with Memory Efficient Macro DesignCVPR 2024 1Scaling Spherical CNNsarXiv 2023SoTaNa: The Open-Source Software Development AssistantarXiv 2023Active Neural MappingICCV 2023 1ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code GenerationarXiv 2024Diving into Underwater: Segment Anything Model Guided Underwater Salient Instance Segmentation and A Large-scale DatasetarXiv 2024Cascaded Text Generation with Markov TransformersNeurIPS 2020 12The Devil is in Temporal Token: High Quality Video Reasoning SegmentationCVPR 2025 1FEVER: a large-scale dataset for Fact Extraction and VERificationfever-a-large-scale-dataset-for-fact-1ChessGPT: Bridging Policy Learning and Language Modelingchessgpt-bridging-policy-learning-andBTGenBot: Behavior Tree Generation for Robotic Tasks with Lightweight
LLMsarXiv 2024Edge-MoE: Memory-Efficient Multi-Task Vision Transformer Architecture with Task-level Sparsity via Mixture-of-ExpertsarXiv 2023Domain Adaptation for Time Series Under Feature and Label ShiftsarXiv 2023Input Perturbation Reduces Exposure Bias in Diffusion ModelsarXiv 2023Video-Based Human Pose Regression via Decoupled Space-Time AggregationCVPR 2024 1CVTHead: One-shot Controllable Head Avatar with Vertex-feature TransformerarXiv 2023Aligning Language Models with Demonstrated FeedbackarXiv 2024SceneTracker: Long-term Scene Flow Estimation NetworkarXiv 2024Optimal Brain Compression: A Framework for Accurate Post-Training Quantization and PruningarXiv 2022Temporally Consistent Transformers for Video GenerationarXiv 2022EMDB: The Electromagnetic Database of Global 3D Human Pose and Shape in the WildICCV 2023 1Optimizing Dense Retrieval Model Training with Hard NegativesarXiv 2021DBCopilot: Natural Language Querying over Massive Databases via Schema RoutingarXiv 2023Latent Embedding Feedback and Discriminative Features for Zero-Shot ClassificationECCV 2020 8EmoSpeech: Guiding FastSpeech2 Towards Emotional Text to SpeecharXiv 2023Learning to Generate Better Than Your LLMarXiv 2023On-device Sora: Enabling Training-Free Diffusion-based Text-to-Video Generation for Mobile DevicesarXiv 2025MosaicFusion: Diffusion Models as Data Augmenters for Large Vocabulary Instance SegmentationarXiv 2023LQ-LoRA: Low-rank Plus Quantized Matrix Decomposition for Efficient Language Model FinetuningarXiv 2023Dynamic Entity Representations in Neural Language Modelsdynamic-entity-representations-in-neural-1Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?arXiv 2024CityDreamer4D: Compositional Generative Model of Unbounded 4D CitiesarXiv 2025Interpretations are useful: penalizing explanations to align neural networks with prior knowledgeinterpretations-are-useful-penalizing-1AnyControl: Create Your Artwork with Versatile Control on Text-to-Image GenerationarXiv 2024GAM Changer: Editing Generalized Additive Models with Interactive VisualizationarXiv 2021SCNet: Sparse Compression Network for Music Source SeparationarXiv 2024FILM: Following Instructions in Language with Modular Methodsfilm-following-instructions-in-language-withDecouple and Track: Benchmarking and Improving Video Diffusion Transformers for Motion TransferICCV 2025AceGPT, Localizing Large Language Models in ArabicarXiv 2023BooookScore: A systematic exploration of book-length summarization in the era of LLMsarXiv 2023VGCN-BERT: Augmenting BERT with Graph Embedding for Text ClassificationarXiv 2020Transformers are Multi-State RNNsarXiv 2024PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse SamplingarXiv 2024Model Stock: All we need is just a few fine-tuned modelsarXiv 2024Rethinking Patch Dependence for Masked AutoencodersarXiv 2024An Adaptive and Momental Bound Method for Stochastic LearningarXiv 2019Repository-Level Prompt Generation for Large Language Models of CodearXiv 2022