All papers
Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.
Attention Calibration for Disentangled Text-to-Image PersonalizationCVPR 2024 1Masked Autoencoders Enable Efficient Knowledge DistillersCVPR 2023 1TANDA: Transfer and Adapt Pre-Trained Transformer Models for Answer Sentence Selectiontanda-transfer-and-adapt-pre-trained-1ArtiGrasp: Physically Plausible Synthesis of Bi-Manual Dexterous Grasping and ArticulationarXiv 2023SCOREQ: Speech Quality Assessment with Contrastive RegressionarXiv 2024Reasoning Implicit Sentiment with Chain-of-Thought PromptingarXiv 2023SGAligner : 3D Scene Alignment with Scene GraphsarXiv 2023OlympicArena Medal Ranks: Who Is the Most Intelligent AI So Far?arXiv 2024LLM360: Towards Fully Transparent Open-Source LLMsarXiv 2023MMChat: Multi-Modal Chat Dataset on Social MediaLREC 2022 6Segment Anything with Multiple ModalitiesarXiv 2024Preventing Zero-Shot Transfer Degradation in Continual Learning of Vision-Language ModelsICCV 2023 1Dear Sir or Madam, May I introduce the GYAFC Dataset: Corpus, Benchmarks and Metrics for Formality Style Transferdear-sir-or-madam-may-i-introduce-the-gyafc-1Asynchronous Large Language Model Enhanced Planner for Autonomous DrivingarXiv 2024Multi-modal Understanding and Generation for Medical Images and Text via Vision-Language Pre-TrainingarXiv 2021JudgeBench: A Benchmark for Evaluating LLM-based JudgesarXiv 2024Task Arithmetic in the Tangent Space: Improved Editing of Pre-Trained ModelsNeurIPS 2023 11Dense Learning based Semi-Supervised Object DetectionCVPR 2022 1Vocabulary-free Image Classificationvocabulary-free-image-classificationAttention Prompting on Image for Large Vision-Language ModelsarXiv 2024Hierarchical Video-Moment Retrieval and Step-CaptioningCVPR 2023 1Attention-Based Transformers for Instance Segmentation of Cells in MicrostructuresarXiv 2020Autoregressive Structured Prediction with Language ModelsarXiv 2022AID: Attention Interpolation of Text-to-Image DiffusionarXiv 2024BertNet: Harvesting Knowledge Graphs with Arbitrary Relations from Pretrained Language ModelsarXiv 2022ShortcutsBench: A Large-Scale Real-world Benchmark for API-based AgentsarXiv 2024Synthetic Experience Replaysynthetic-experience-replayLongVLM: Efficient Long Video Understanding via Large Language ModelsarXiv 2024MISF: Multi-level Interactive Siamese Filtering for High-Fidelity Image InpaintingCVPR 2022 1AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZarXiv 2023Vanishing Point Estimation in Uncalibrated Images with Prior Gravity DirectionICCV 2023 1The pitfalls of next-token predictionarXiv 2024Monash University, UEA, UCR Time Series Extrinsic Regression ArchivearXiv 2020Towards a Multimodal Large Language Model with Pixel-Level Insight for BiomedicinearXiv 2024Bring Your Own Data! Self-Supervised Evaluation for Large Language ModelsarXiv 2023Word-Level Coreference ResolutionEMNLP 2021 11GeoSynth: Contextually-Aware High-Resolution Satellite Image SynthesisarXiv 2024FairGBM: Gradient Boosting with Fairness ConstraintsarXiv 2022An Explanation of In-context Learning as Implicit Bayesian Inferencean-explanation-of-in-context-learning-asSearching Latent Program SpacesarXiv 2024Hammer: Robust Function-Calling for On-Device Language Models via Function MaskingarXiv 2024Self6D: Self-Supervised Monocular 6D Object Pose EstimationECCV 2020 8Glot500: Scaling Multilingual Corpora and Language Models to 500 LanguagesarXiv 2023Pre-Training to Learn in ContextarXiv 2023Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to SeearXiv 2024Self-Evolving Multi-Agent Simulations for Realistic Clinical InteractionsarXiv 2025Diffusion Models and Representation Learning: A SurveyarXiv 2024A comprehensive evaluation of ChatGPT's zero-shot Text-to-SQL capabilityarXiv 2023Artificial Kuramoto Oscillatory NeuronsarXiv 2024VITON-GAN: Virtual Try-on Image Generator Trained with Adversarial LossarXiv 2019Towards accurate instance segmentation in large-scale LiDAR point cloudsarXiv 2023CLIP model is an Efficient Continual LearnerarXiv 2022Sequence-to-Sequence Knowledge Graph Completion and Question AnsweringACL 2022 5OGC: Unsupervised 3D Object Segmentation from Rigid Dynamics of Point CloudsarXiv 2022SummaC: Re-Visiting NLI-based Models for Inconsistency Detection in SummarizationarXiv 2021$\infty$-Diff: Infinite Resolution Diffusion with Subsampled Mollified StatesarXiv 2023Gated Linear Attention Transformers with Hardware-Efficient TrainingarXiv 2023Enhancing Activity Prediction Models in Drug Discovery with the Ability to Understand Human LanguagearXiv 2023H2RBox: Horizontal Box Annotation is All You Need for Oriented Object DetectionarXiv 2022OmniFusion: 360 Monocular Depth Estimation via Geometry-Aware FusionCVPR 2022 1ViTime: A Visual Intelligence-Based Foundation Model for Time Series ForecastingarXiv 2024Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic LanguagesarXiv 2022SQUID: Deep Feature In-Painting for Unsupervised Anomaly DetectionCVPR 2023 1Half Wavelet Attention on M-Net+ for Low-Light Image EnhancementarXiv 2022NeRFLiX: High-Quality Neural View Synthesis by Learning a Degradation-Driven Inter-viewpoint MiXerCVPR 2023 1Load What You Need: Smaller Versions of Multilingual BERTarXiv 2020Meta Optimal TransportarXiv 2022SELFormer: Molecular Representation Learning via SELFIES Language Modelsselformer-molecular-representation-learning-1OvarNet: Towards Open-vocabulary Object Attribute RecognitionCVPR 2023 1Co-Mixup: Saliency Guided Joint Mixup with Supermodular Diversityco-mixup-saliency-guided-joint-mixup-withSpanish Pre-trained BERT Model and Evaluation DataarXiv 2023Adaptable Logical Control for Large Language ModelsarXiv 2024UPop: Unified and Progressive Pruning for Compressing Vision-Language TransformersarXiv 2023A Comparative Study on Reasoning Patterns of OpenAI's o1 ModelarXiv 2024Memory Gym: Towards Endless Tasks to Benchmark Memory Capabilities of AgentsarXiv 2023Towards Lossless Dataset Distillation via Difficulty-Aligned Trajectory MatchingarXiv 2023SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-ImprovementarXiv 2025Annotating the Tweebank Corpus on Named Entity Recognition and Building NLP Models for Social Media AnalysisLREC 2022 6SuS-X: Training-Free Name-Only Transfer of Vision-Language ModelsICCV 2023 1Aksharantar: Open Indic-language Transliteration datasets and models for the Next Billion UsersarXiv 2022Diffusion Models for Video Prediction and InfillingarXiv 2022Prompt Cache: Modular Attention Reuse for Low-Latency InferencearXiv 2023PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action ChainarXiv 2024Contrastive Test-Time AdaptationCVPR 2022 1Flora: Low-Rank Adapters Are Secretly Gradient CompressorsarXiv 2024Language in a Bottle: Language Model Guided Concept Bottlenecks for Interpretable Image ClassificationCVPR 2023 1Bag of Tricks and A Strong baseline for Image Copy DetectionarXiv 2021SeqNet: Learning Descriptors for Sequence-based Hierarchical Place RecognitionarXiv 2021Self-supervised Deep Reinforcement Learning with Generalized Computation Graphs for Robot NavigationarXiv 2017ESimCSE: Enhanced Sample Building Method for Contrastive Learning of Unsupervised Sentence EmbeddingCOLING 2022 10SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video RestorationCVPR 2025 13DILG: Irregular Latent Grids for 3D Generative ModelingarXiv 2022AttentiveNAS: Improving Neural Architecture Search via Attentive SamplingCVPR 2021 1Faithful Persona-based Conversational Dataset Generation with Large Language ModelsarXiv 2023Gotta Go Fast When Generating Data with Score-Based Modelsgotta-go-fast-when-generating-data-with-score-1SFPNet: Sparse Focal Point Network for Semantic Segmentation on General LiDAR Point CloudsarXiv 2024LLMeBench: A Flexible Framework for Accelerating LLMs BenchmarkingarXiv 2023FunQA: Towards Surprising Video ComprehensionarXiv 2023SleepFM: Multi-modal Representation Learning for Sleep Across Brain Activity, ECG and Respiratory SignalsarXiv 2024NeRF-MAE: Masked AutoEncoders for Self-Supervised 3D Representation Learning for Neural Radiance FieldsarXiv 2024Rethinking Architecture Selection in Differentiable NASrethinking-architecture-selection-inFastSHAP: Real-Time Shapley Value Estimationfastshap-real-time-shapley-value-estimation-1Instruct-MusicGen: Unlocking Text-to-Music Editing for Music Language Models via Instruction TuningarXiv 2024LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context InferencearXiv 2024XCOPA: A Multilingual Dataset for Causal Commonsense ReasoningEMNLP 2020 11VQGraph: Rethinking Graph Representation Space for Bridging GNNs and MLPsarXiv 2023Parameter-Efficient Fine-Tuning for Foundation ModelsarXiv 2025Language-Grounded Dynamic Scene Graphs for Interactive Object Search
with Mobile ManipulationarXiv 2024Neural Symbolic Regression that ScalesarXiv 2021Fourier Position Embedding: Enhancing Attention's Periodic Extension for Length GeneralizationarXiv 2024BARTpho: Pre-trained Sequence-to-Sequence Models for VietnamesearXiv 2021SmartControl: Enhancing ControlNet for Handling Rough Visual ConditionsarXiv 2024DiJiang: Efficient Large Language Models through Compact KernelizationarXiv 2024Rephrase and Respond: Let Large Language Models Ask Better Questions for ThemselvesarXiv 2023FSFM: A Generalizable Face Security Foundation Model via Self-Supervised Facial Representation LearningCVPR 2025 1Adaptive Token Sampling For Efficient Vision TransformersarXiv 2021Foundation Policies with Hilbert RepresentationsarXiv 2024DiffRate : Differentiable Compression Rate for Efficient Vision TransformersICCV 2023 1Rank-DETR for High Quality Object Detectionrank-detr-for-high-quality-object-detectionRoboMatrix: A Skill-centric Hierarchical Framework for Scalable Robot Task Planning and Execution in Open-WorldarXiv 2024Dynamic Evaluation of Neural Sequence Modelsdynamic-evaluation-of-neural-sequence-models-1Automated forest inventory: analysis of high-density airborne LiDAR point clouds with 3D deep learningarXiv 2023Explicit Shape Encoding for Real-Time Instance Segmentationexplicit-shape-encoding-for-real-time-1Learning Neural Causal Models with Active Interventionslearning-neural-causal-models-with-active-1EigenTrajectory: Low-Rank Descriptors for Multi-Modal Trajectory ForecastingICCV 2023 1Deep Reinforcement Learning Based Joint Downlink Beamforming and RIS Configuration in RIS-aided MU-MISO Systems Under Hardware Impairments and Imperfect CSIarXiv 2022MegaScenes: Scene-Level View Synthesis at ScalearXiv 2024FPGA: Fast Patch-Free Global Learning Framework for Fully End-to-End Hyperspectral Image Classificationfpga-fast-patch-free-global-learningUltraPose: Synthesizing Dense Pose with 1 Billion Points by Human-body Decoupling 3D Modelultrapose-synthesizing-dense-pose-with-1Can Large Language Model Agents Simulate Human Trust Behavior?arXiv 2024SALSA: Spatial Cue-Augmented Log-Spectrogram Features for Polyphonic Sound Event Localization and DetectionarXiv 2021M3DeTR: Multi-representation, Multi-scale, Mutual-relation 3D Object Detection with TransformersarXiv 2021Graph Pre-training for AMR Parsing and GenerationACL 2022 5Online normalizer calculation for softmaxarXiv 2018Deepfake Video Detection Using Convolutional Vision TransformerarXiv 2021Autoencoders as Cross-Modal Teachers: Can Pretrained 2D Image Transformers Help 3D Representation Learning?arXiv 2022ChatLog: Carefully Evaluating the Evolution of ChatGPT Across TimearXiv 2023T-Stitch: Accelerating Sampling in Pre-Trained Diffusion Models with Trajectory StitchingarXiv 2024Compositional Exemplars for In-context LearningarXiv 2023HiTab: A Hierarchical Table Dataset for Question Answering and Natural Language GenerationACL 2022 5VCP-CLIP: A visual context prompting model for zero-shot anomaly segmentationarXiv 2024M3KE: A Massive Multi-Level Multi-Subject Knowledge Evaluation Benchmark for Chinese Large Language ModelsarXiv 2023Fast and Robust Dynamic Hand Gesture Recognition via Key Frames Extraction and Feature FusionarXiv 2019Zolly: Zoom Focal Length Correctly for Perspective-Distorted Human Mesh ReconstructionICCV 2023 1Learning from Committee: Reasoning Distillation from a Mixture of Teachers with Peer-ReviewarXiv 2024ROSGPT_Vision: Commanding Robots Using Only Language Models' PromptsarXiv 2023Getting it Right: Improving Spatial Consistency in Text-to-Image ModelsarXiv 2024DR2: Diffusion-based Robust Degradation Remover for Blind Face RestorationCVPR 2023 1Controllable Sentence Simplificationcontrollable-sentence-simplification-2ChiMed-GPT: A Chinese Medical Large Language Model with Full Training Regime and Better Alignment to Human PreferencesarXiv 2023Speech Slytherin: Examining the Performance and Efficiency of Mamba for Speech Separation, Recognition, and SynthesisarXiv 2024Compose and Conquer: Diffusion-Based 3D Depth Aware Composable Image SynthesisarXiv 2024M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language ModelsNeurIPS 2023 11On the De-duplication of LAION-2BarXiv 2023In-Context Learning Creates Task VectorsarXiv 2023UPGPT: Universal Diffusion Model for Person Image Generation, Editing and Pose TransferarXiv 2023Improving Zero-Shot Generalization for CLIP with Synthesized PromptsICCV 2023 1Implicit Autoencoder for Point-Cloud Self-Supervised Representation LearningICCV 2023 1SFace: Sigmoid-Constrained Hypersphere Loss for Robust Face Recognitionsface-sigmoid-constrained-hypersphere-lossRethinking Automatic Evaluation in Sentence SimplificationarXiv 2021FLAIR: Federated Learning Annotated Image RepositoryarXiv 2022Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing PolicyarXiv 2023NeuroBench: A Framework for Benchmarking Neuromorphic Computing Algorithms and SystemsarXiv 2023ReMoE: Fully Differentiable Mixture-of-Experts with ReLU RoutingarXiv 2024TRANSIC: Sim-to-Real Policy Transfer by Learning from Online CorrectionarXiv 2024CLIPAway: Harmonizing Focused Embeddings for Removing Objects via Diffusion ModelsarXiv 2024LitSearch: A Retrieval Benchmark for Scientific Literature SearcharXiv 2024Towards the Law of Capacity Gap in Distilling Language ModelsarXiv 2023Do Large Language Models Know What They Don't Know?arXiv 2023CLIP2Protect: Protecting Facial Privacy using Text-Guided Makeup via Adversarial Latent Searchclip2protect-protecting-facial-privacy-usingSwitchHead: Accelerating Transformers with Mixture-of-Experts AttentionarXiv 2023ProphetFuzz: Fully Automated Prediction and Fuzzing of High-Risk Option
Combinations with Only Documentation via Large Language ModelarXiv 2024ALIP: Adaptive Language-Image Pre-training with Synthetic CaptionICCV 2023 1ManipVQA: Injecting Robotic Affordance and Physically Grounded
Information into Multi-Modal Large Language ModelsarXiv 2024NEVIS'22: A Stream of 100 Tasks Sampled from 30 Years of Computer Vision ResearcharXiv 2022LoCoOp: Few-Shot Out-of-Distribution Detection via Prompt Learninglocoop-few-shot-out-of-distribution-detectionAlias-Free Latent Diffusion Models:Improving Fractional Shift Equivariance of Diffusion Latent SpacearXiv 2025MPCFormer: fast, performant and private Transformer inference with MPCarXiv 2022Learned Token Pruning for TransformersarXiv 2021WorldGUI: An Interactive Benchmark for Desktop GUI Automation from Any Starting PointarXiv 2025EmoBERTa: Speaker-Aware Emotion Recognition in Conversation with RoBERTaarXiv 2021MotionMaster: Training-free Camera Motion Transfer For Video GenerationarXiv 2024Multimodal Pathway: Improve Transformers with Irrelevant Data from Other ModalitiesCVPR 2024 1Analyzing Leakage of Personally Identifiable Information in Language ModelsarXiv 2023SHiNe: Semantic Hierarchy Nexus for Open-vocabulary Object DetectionarXiv 2024Generalized Out-of-Distribution Detection and Beyond in Vision Language Model Era: A SurveyarXiv 2024UmlsBERT: Clinical Domain Knowledge Augmentation of Contextual Embeddings Using the Unified Medical Language System MetathesaurusNAACL 2021 4Contrastive Embedding for Generalized Zero-Shot LearningCVPR 2021 1A Variational Perspective on Solving Inverse Problems with Diffusion ModelsarXiv 2023Capturing and Inferring Dense Full-Body Human-Scene Contactcapturing-and-inferring-dense-full-body-humanTime Evidence Fusion Network: Multi-source View in Long-Term Time Series ForecastingarXiv 2024Lets keep it simple, Using simple architectures to outperform deeper and more complex architecturesarXiv 2016MedChatZH: a Better Medical Adviser Learns from Better InstructionsarXiv 2023SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?arXiv 2024MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to UsearXiv 2023ConvNet vs Transformer, Supervised vs CLIP: Beyond ImageNet AccuracyarXiv 2023REFLECT: Summarizing Robot Experiences for Failure Explanation and CorrectionarXiv 2023Driv3R: Learning Dense 4D Reconstruction for Autonomous DrivingarXiv 2024PartCraft: Crafting Creative Objects by PartsarXiv 2024LidarCLIP or: How I Learned to Talk to Point CloudsarXiv 2022