All papers
Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.
Intent Detection and Slot Filling for VietnamesearXiv 2021Named Entity Recognition in Indian court judgmentsarXiv 2022Attention in Attention Network for Image Super-Resolutionattention-in-attention-network-for-image-1PlainMamba: Improving Non-Hierarchical Mamba in Visual RecognitionarXiv 2024Pseudo-Relevance Feedback for Multiple Representation Dense RetrievalarXiv 2021A disciplined approach to neural network hyper-parameters: Part 1 -- learning rate, batch size, momentum, and weight decayarXiv 20183D Common Corruptions and Data AugmentationCVPR 2022 1DAF:re: A Challenging, Crowd-Sourced, Large-Scale, Long-Tailed Dataset For Anime Character RecognitionarXiv 2021Multilingual Byte2Speech Models for Scalable Low-resource Speech SynthesisarXiv 2021EVP: Enhanced Visual Perception using Inverse Multi-Attentive Feature Refinement and Regularized Image-Text AlignmentarXiv 2023TURNA: A Turkish Encoder-Decoder Language Model for Enhanced Understanding and GenerationarXiv 2024Fast, Expressive SE$(n)$ Equivariant Networks through Weight-Sharing in Position-Orientation SpacearXiv 2023Language Models of Code are Few-Shot Commonsense LearnersarXiv 2022VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and GenerationarXiv 2023ABINet++: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text SpottingarXiv 2022Massive Values in Self-Attention Modules are the Key to Contextual Knowledge UnderstandingarXiv 2025Characteristic Guidance: Non-linear Correction for Diffusion Model at Large Guidance ScalearXiv 2023GNOT: A General Neural Operator Transformer for Operator LearningarXiv 2023SiriuS: Self-improving Multi-agent Systems via Bootstrapped ReasoningarXiv 2025Geometric Representation Learning for Document Image RectificationarXiv 2022B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught ReasonersarXiv 2024BearLLM: A Prior Knowledge-Enhanced Bearing Health Management Framework with Unified Vibration Signal RepresentationarXiv 20243D Segmentation of Humans in Point Clouds with Synthetic DataICCV 2023 1Demystifying MMD GANsdemystifying-mmd-gans-1Elysium: Exploring Object-level Perception in Videos via MLLMarXiv 2024Progressive Autoregressive Video Diffusion ModelsarXiv 2024Exploring Predicate Visual Context in Detecting Human-Object InteractionsICCV 2023 1WalkTheDog: Cross-Morphology Motion Alignment via Phase Manifoldswalkthedog-cross-morphology-motion-alignmentLenna: Language Enhanced Reasoning Detection AssistantarXiv 2023STAIR: Improving Safety Alignment with Introspective ReasoningarXiv 2025BenchLMM: Benchmarking Cross-style Visual Capability of Large Multimodal ModelsarXiv 2023DGInStyle: Domain-Generalizable Semantic Segmentation with Image Diffusion Models and Stylized Semantic ControlarXiv 2023Back to the Source: Diffusion-Driven Test-Time AdaptationarXiv 2022OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution FittingarXiv 2025A Comprehensive Overview of Large Language ModelsarXiv 2023A Review of Modern Recommender Systems Using Generative Models (Gen-RecSys)arXiv 2024Online Normalization for Training Neural Networksonline-normalization-for-training-neural-1NEF: Neural Edge Fields for 3D Parametric Curve Reconstruction from Multi-view ImagesCVPR 2023 1An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQAarXiv 2021Contrastive Learning with Adversarial Perturbations for Conditional Text Generationcontrastive-learning-with-adversarialEnhancing Chain-of-Thoughts Prompting with Iterative Bootstrapping in Large Language ModelsarXiv 2023Fully Hyperbolic Neural NetworksACL 2022 5LivelySpeaker: Towards Semantic-Aware Co-Speech Gesture GenerationICCV 2023 1MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language BenchmarkarXiv 2024Towards Long-Form Video Understandingtowards-long-form-video-understandingConnecting Vision and Language with Localized NarrativesECCV 2020 8Convergent Learning: Do different neural networks learn the same representations?arXiv 2015Generalized Lightness Adaptation with Channel Selective NormalizationICCV 2023 1SwinGNN: Rethinking Permutation Invariance in Diffusion Models for Graph GenerationarXiv 2023Knowledge Unlearning for Mitigating Privacy Risks in Language ModelsarXiv 2022Game-theoretic LLM: Agent Workflow for Negotiation GamesarXiv 2024UnitedHuman: Harnessing Multi-Source Data for High-Resolution Human GenerationICCV 2023 1Self-Evolved Diverse Data Sampling for Efficient Instruction TuningarXiv 2023MoDem: Accelerating Visual Model-Based Reinforcement Learning with DemonstrationsarXiv 2022Scaling Laws and Compute-Optimal Training Beyond Fixed Training DurationsarXiv 2024AdaNPC: Exploring Non-Parametric Classifier for Test-Time AdaptationarXiv 2023RWKU: Benchmarking Real-World Knowledge Unlearning for Large Language ModelsarXiv 2024Wired Perspectives: Multi-View Wire Art Embraces Generative AICVPR 2024 1Exploring Discrete Diffusion Models for Image CaptioningarXiv 2022Separating the "Chirp" from the "Chat": Self-supervised Visual Grounding of Sound and LanguageCVPR 2024 1Rethinking Vision-Language Model in Face Forensics: Multi-Modal Interpretable Forged Face DetectorCVPR 2025 1Data Distributional Properties Drive Emergent In-Context Learning in TransformersarXiv 2022Recovering the Pre-Fine-Tuning Weights of Generative ModelsarXiv 2024Sub-Sentence Encoder: Contrastive Learning of Propositional Semantic RepresentationsarXiv 2023BanglaNLG and BanglaT5: Benchmarks and Resources for Evaluating Low-Resource Natural Language Generation in BanglaarXiv 2022An Empirical Study of End-to-End Temporal Action DetectionCVPR 2022 1PTQ4SAM: Post-Training Quantization for Segment AnythingarXiv 2024SG-Former: Self-guided Transformer with Evolving Token ReallocationICCV 2023 1MonoNeRD: NeRF-like Representations for Monocular 3D Object DetectionICCV 2023 1SciInstruct: a Self-Reflective Instruction Annotated Dataset for Training Scientific Language ModelsarXiv 2024Can AI Assistants Know What They Don't Know?arXiv 2024GlyphDraw2: Automatic Generation of Complex Glyph Posters with Diffusion Models and Large Language ModelsarXiv 2024Quantized GAN for Complex Music Generation from Dance VideosarXiv 2022Cheating Automatic LLM Benchmarks: Null Models Achieve High Win RatesarXiv 2024Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token ReductionarXiv 2024DebugBench: Evaluating Debugging Capability of Large Language ModelsarXiv 2024MoTCoder: Elevating Large Language Models with Modular of Thought for Challenging Programming TasksarXiv 2023Baby Llama: knowledge distillation from an ensemble of teachers trained on a small dataset with no performance penaltyarXiv 2023PSAvatar: A Point-based Shape Model for Real-Time Head Avatar Animation with 3D Gaussian SplattingarXiv 2024WECHSEL: Effective initialization of subword embeddings for cross-lingual transfer of monolingual language modelsNAACL 2022 7Mercury: A Code Efficiency Benchmark for Code Large Language ModelsarXiv 2024Hierarchical VAEs Know What They Don't KnowarXiv 2021CLIP-EBC: CLIP Can Count Accurately through Enhanced Blockwise ClassificationarXiv 2024Hyperparameters in Reinforcement Learning and How To Tune ThemarXiv 2023Sparse Low-rank Adaptation of Pre-trained Language ModelsarXiv 2023MCTS-RAG: Enhancing Retrieval-Augmented Generation with Monte Carlo Tree SearcharXiv 2025PALO: A Polyglot Large Multimodal Model for 5B PeoplearXiv 2024Multimodal Self-Instruct: Synthetic Abstract Image and Visual Reasoning Instruction Using Language ModelarXiv 2024Towards Explainable In-the-Wild Video Quality Assessment: A Database and a Language-Prompted ApproacharXiv 2023Few-Shot Font Generation by Learning Fine-Grained Local StylesCVPR 2022 1GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and A Comprehensive Multimodal Dataset Towards General Medical AIarXiv 2024Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge GraphsarXiv 2024One-Step Diffusion Distillation through Score Implicit MatchingarXiv 2024Symbolic Music Generation with Non-Differentiable Rule Guided DiffusionarXiv 2024AdaMAE: Adaptive Masking for Efficient Spatiotemporal Learning with Masked AutoencodersCVPR 2023 1XLand-100B: A Large-Scale Multi-Task Dataset for In-Context Reinforcement LearningarXiv 2024iHuman: Instant Animatable Digital Humans From Monocular VideosarXiv 2024A Comprehensive Study of Jailbreak Attack versus Defense for Large Language ModelsarXiv 2024MotionAug: Augmentation with Physical Correction for Human Motion PredictionCVPR 2022 1Logical Fallacy DetectionarXiv 2022Kick Back & Relax: Learning to Reconstruct the World by Watching SlowTVICCV 2023 1LexEval: A Comprehensive Chinese Legal Benchmark for Evaluating Large Language ModelsarXiv 20243D Paintbrush: Local Stylization of 3D Shapes with Cascaded Score DistillationCVPR 2024 1LangSuitE: Planning, Controlling and Interacting with Large Language Models in Embodied Text EnvironmentsarXiv 2024Context is Key: A Benchmark for Forecasting with Essential Textual InformationarXiv 2024Predicting Gradient is Better: Exploring Self-Supervised Learning for SAR ATR with a Joint-Embedding Predictive ArchitecturearXiv 2023Curiosity-driven Red-teaming for Large Language ModelsarXiv 2024LiDAR-Camera Panoptic Segmentation via Geometry-Consistent and Semantic-Aware AlignmentICCV 2023 1iFusion: Inverting Diffusion for Pose-Free Reconstruction from Sparse ViewsarXiv 2023AdaPool: Exponential Adaptive Pooling for Information-Retaining DownsamplingarXiv 2021Frequency-Aware Transformer for Learned Image CompressionarXiv 2023Mugs: A Multi-Granular Self-Supervised Learning FrameworkarXiv 2022MAUD: An Expert-Annotated Legal NLP Dataset for Merger Agreement UnderstandingarXiv 2023LLMs Know More Than They Show: On the Intrinsic Representation of LLM HallucinationsarXiv 2024HIR-Diff: Unsupervised Hyperspectral Image Restoration Via Improved Diffusion ModelsCVPR 2024 1A Survey on Federated Fine-tuning of Large Language ModelsarXiv 2025The Foundation Model Transparency IndexarXiv 2023Teaching Arithmetic to Small TransformersarXiv 2023Think While You Generate: Discrete Diffusion with Planned DenoisingarXiv 2024Graph Neural Networks for Learning Equivariant Representations of Neural NetworksarXiv 2024Representative Forgery Mining for Fake Face DetectionCVPR 2021 1Image Captioning with Deep Bidirectional LSTMsarXiv 2016SkeletonMAE: Graph-based Masked Autoencoder for Skeleton Sequence Pre-trainingICCV 2023 1OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and OptimizationarXiv 2024InterControl: Zero-shot Human Interaction Generation by Controlling Every JointarXiv 2023How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMsarXiv 2023Multimodal Learning Without Labeled Multimodal Data: Guarantees and ApplicationsarXiv 2023Robust Distortion-free Watermarks for Language ModelsarXiv 2023TabReD: Analyzing Pitfalls and Filling the Gaps in Tabular Deep Learning BenchmarksarXiv 2024Low-light Image Enhancement via CLIP-Fourier Guided Wavelet DiffusionarXiv 2024LayerTracer: Cognitive-Aligned Layered SVG Synthesis via Diffusion TransformerICCV 2025Can I Trust Your Answer? Visually Grounded Video Question AnsweringCVPR 2024 1Diffusion Language Models Can Perform Many Tasks with Scaling and Instruction-FinetuningarXiv 2023MedleyVox: An Evaluation Dataset for Multiple Singing Voices SeparationarXiv 2022ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique PipelinearXiv 2024Tartarus: A Benchmarking Platform for Realistic And Practical Inverse
Molecular DesignarXiv 2022DQS3D: Densely-matched Quantization-aware Semi-supervised 3D DetectionICCV 2023 1Neural Sheaf Diffusion: A Topological Perspective on Heterophily and Oversmoothing in GNNsarXiv 2022A Comprehensive Survey of Regression Based Loss Functions for Time Series ForecastingarXiv 2022Enhancing High-Resolution 3D Generation through Pixel-wise Gradient ClippingarXiv 2023Towards Localized Fine-Grained Control for Facial Expression GenerationarXiv 2024Zero-Shot Dialogue State Tracking via Cross-Task TransferEMNLP 2021 11Modular Visual Question Answering via Code GenerationarXiv 2023TRACE: A Comprehensive Benchmark for Continual Learning in Large Language ModelsarXiv 2023SILO Language Models: Isolating Legal Risk In a Nonparametric DatastorearXiv 2023Consistency-guided Prompt Learning for Vision-Language ModelsarXiv 2023Multi-modal Vision Pre-training for Medical Image AnalysisarXiv 2024Your Mixture-of-Experts LLM Is Secretly an Embedding Model For FreearXiv 2024LinVT: Empower Your Image-level Large Language Model to Understand VideosarXiv 2024Ctrl-U: Robust Conditional Image Generation via Uncertainty-aware Reward ModelingarXiv 2024REMIND Your Neural Network to Prevent Catastrophic ForgettingECCV 2020 8xCos: An Explainable Cosine Metric for Face Verification TaskarXiv 2020PosFormer: Recognizing Complex Handwritten Mathematical Expression with Position Forest TransformerarXiv 2024RAW-Adapter: Adapting Pre-trained Visual Model to Camera RAW ImagesarXiv 2024MusiConGen: Rhythm and Chord Control for Transformer-Based Text-to-Music GenerationarXiv 2024Randomized Positional Encodings Boost Length Generalization of TransformersarXiv 2023Diversity-Aware Meta Visual PromptingCVPR 2023 1One-Trimap Video MattingarXiv 2022ULSAM: Ultra-Lightweight Subspace Attention Module for Compact Convolutional Neural NetworksarXiv 2020HyperAttention: Long-context Attention in Near-Linear TimearXiv 2023Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMsarXiv 2024State-Free Inference of State-Space Models: The Transfer Function ApproacharXiv 2024mDPO: Conditional Preference Optimization for Multimodal Large Language ModelsarXiv 2024TouchStone: Evaluating Vision-Language Models by Language ModelsarXiv 2023Reinforce Data, Multiply Impact: Improved Model Accuracy and Robustness with Dataset ReinforcementICCV 2023 1Soft Contrastive Learning for Time SeriesarXiv 2023UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and ReasoningarXiv 2023CIE XYZ Net: Unprocessing Images for Low-Level Computer Vision TasksarXiv 2020ReClor: A Reading Comprehension Dataset Requiring Logical ReasoningICLR 2020 1A Frustratingly Simple Yet Highly Effective Attack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1arXiv 2025RealTime QA: What's the Answer Right Now?realtime-qa-what-s-the-answer-right-nowAdvancing Radiograph Representation Learning with Masked Record ModelingarXiv 2023RetrievalAttention: Accelerating Long-Context LLM Inference via Vector RetrievalarXiv 2024Visual Reinforcement Learning with Self-Supervised 3D RepresentationsarXiv 2022End-to-End Goal-Driven Web Navigationend-to-end-goal-driven-web-navigation-1Mellow: a small audio language model for reasoningarXiv 2025ClimateGAN: Raising Climate Change Awareness by Generating Images of Floodsclimategan-raising-climate-change-awareness-1Semantic MapNet: Building Allocentric Semantic Maps and Representations from Egocentric ViewsarXiv 2020Aging with GRACE: Lifelong Model Editing with Discrete Key-Value AdaptorsNeurIPS 2023 11Investigating the Factual Knowledge Boundary of Large Language Models with Retrieval AugmentationarXiv 2023ZONE: Zero-Shot Instruction-Guided Local EditingCVPR 2024 1Large language models surpass human experts in predicting neuroscience resultsarXiv 2024VIMI: Vehicle-Infrastructure Multi-view Intermediate Fusion for Camera-based 3D Object DetectionarXiv 2023Large Language Models as Zero-Shot Conversational RecommendersarXiv 2023AutoFlow: Automated Workflow Generation for Large Language Model AgentsarXiv 2024Adversarial Open Domain Adaptation for Sketch-to-Photo SynthesisarXiv 2021SA-BEV: Generating Semantic-Aware Bird's-Eye-View Feature for Multi-view 3D Object DetectionICCV 2023 1OCR-VQGAN: Taming Text-within-Image GenerationarXiv 2022Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answeringcan-a-suit-of-armor-conduct-electricity-a-new-1Vision-by-Language for Training-Free Compositional Image RetrievalarXiv 2023MedShapeNet -- A Large-Scale Dataset of 3D Medical Shapes for Computer VisionarXiv 2023Margin-aware Preference Optimization for Aligning Diffusion Models without ReferencearXiv 2024Etalon: Holistic Performance Evaluation Framework for LLM Inference SystemsarXiv 2024GenSim2: Scaling Robot Data Generation with Multi-modal and Reasoning LLMsarXiv 2024PARIS: Part-level Reconstruction and Motion Analysis for Articulated ObjectsICCV 2023 1Measuring a hate speech spectrum with faceted Rasch item response theory and perspective-aware, explainable-by-design deep learningarXiv 2020LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language ModelsarXiv 2022WARP: Word-level Adversarial ReProgrammingACL 2021 5LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed PromptsarXiv 2023Alignment faking in large language modelsarXiv 2024