All papers
Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.
On the Copying Behaviors of Pre-Training for Neural Machine TranslationFindings (ACL) 2021 8Harnessing GANs for Zero-shot Learning of New Classes in Visual Speech RecognitionarXiv 2019Supervised Knowledge Makes Large Language Models Better In-context LearnersarXiv 2023ETC-NLG: End-to-end Topic-Conditioned Natural Language GenerationarXiv 2020KOROL: Learning Visualizable Object Feature with Koopman Operator
Rollout for ManipulationarXiv 2024PEFT-U: Parameter-Efficient Fine-Tuning for User PersonalizationarXiv 2024On the Over-Memorization During Natural, Robust and Catastrophic OverfittingarXiv 2023Image Labels Are All You Need for Coarse Seagrass SegmentationarXiv 2023Human Guided Exploitation of Interpretable Attention Patterns in Summarization and Topic SegmentationarXiv 2021GM-DF: Generalized Multi-Scenario Deepfake DetectionarXiv 2024On Learning the Transformer Kernelon-learning-the-transformer-kernelHow many words does ChatGPT know? The answer is ChatWordsarXiv 2023Igeood: An Information Geometry Approach to Out-of-Distribution Detectionigeood-an-information-geometry-approach-toTowards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and FindingsarXiv 2025Unobserved Local Structures Make Compositional Generalization HardarXiv 2022Indoor Scene Generation from a Collection of Semantic-Segmented Depth ImagesICCV 2021 10Boundary-Denoising for Video Activity LocalizationarXiv 2023MURI: High-Quality Instruction Tuning Datasets for Low-Resource Languages via Reverse InstructionsarXiv 2024Focus on what matters: Applying Discourse Coherence Theory to Cross Document CoreferenceEMNLP 2021 11SortedAP: Rethinking evaluation metrics for instance segmentationarXiv 2023Environment-Invariant Curriculum Relation Learning for Fine-Grained Scene Graph GenerationICCV 2023 1GradSign: Model Performance Inference with Theoretical Insightsgradsign-model-performance-inference-withPaRot: Patch-Wise Rotation-Invariant Network via Feature Disentanglement and Pose RestorationarXiv 2023Just ASR + LLM? A Study on Speech Large Language Models' Ability to Identify and Understand Speaker in Spoken DialoguearXiv 2024From Insights to Actions: The Impact of Interpretability and Analysis Research on NLParXiv 2024ViC-MAE: Self-Supervised Representation Learning from Images and Video with Contrastive Masked AutoencodersarXiv 2023Post-hoc Probabilistic Vision-Language ModelsarXiv 2024The Benefits of Model-Based Generalization in Reinforcement LearningarXiv 2022HC4: A New Suite of Test Collections for Ad Hoc CLIRarXiv 2022TCOVIS: Temporally Consistent Online Video Instance SegmentationICCV 2023 1Cross-lingual Prompting: Improving Zero-shot Chain-of-Thought Reasoning across LanguagesarXiv 2023All Keypoints You Need: Detecting Arbitrary Keypoints on the Body of Triple, High, and Long Jump AthletesarXiv 2023On the Efficacy of Differentially Private Few-shot Image ClassificationarXiv 2023Divide, Reweight, and Conquer: A Logit Arithmetic Approach for In-Context LearningarXiv 2024SENetV2: Aggregated dense layer for channelwise and global representationsarXiv 2023Annealing Self-Distillation Rectification Improves Adversarial TrainingarXiv 2023Diversity-Driven Synthesis: Enhancing Dataset Distillation through Directed Weight AdjustmentarXiv 2024Free-Form Variational Inference for Gaussian Process State-Space ModelsarXiv 2023Opportunities and Challenges in Neural Dialog TutoringarXiv 2023QADiscourse -- Discourse Relations as QA Pairs: Representation, Crowdsourcing and BaselinesarXiv 2020Mind Your Format: Towards Consistent Evaluation of In-Context Learning ImprovementsarXiv 2024VALUE: Understanding Dialect Disparity in NLUACL 2022 5Efficient Generation of Structured Objects with Constrained Adversarial NetworksNeurIPS 2020 12Distilling Reasoning Capabilities into Smaller Language ModelsarXiv 2022A Reply to Makelov et al. (2023)'s "Interpretability Illusion" ArgumentsarXiv 2024LUNet: Deep Learning for the Segmentation of Arterioles and Venules in High Resolution Fundus ImagesarXiv 2023Deep-Q Learning with Hybrid Quantum Neural Network on Solving Maze ProblemsarXiv 2023Learn Your Tokens: Word-Pooled Tokenization for Language ModelingarXiv 2023Positive Label Is All You Need for Multi-Label ClassificationarXiv 2023iSEA: An Interactive Pipeline for Semantic Error Analysis of NLP ModelsarXiv 2022A Mechanistic Analysis of a Transformer Trained on a Symbolic Multi-Step Reasoning TaskarXiv 2024GraphCleaner: Detecting Mislabelled Samples in Popular Graph Learning BenchmarksarXiv 2023Better Training of GFlowNets with Local Credit and Incomplete TrajectoriesarXiv 2023LexGPT 0.1: pre-trained GPT-J models with Pile of LawarXiv 2023MDD-Eval: Self-Training on Augmented Data for Multi-Domain Dialogue EvaluationarXiv 2021CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production EditingarXiv 2026hist2RNA: An efficient deep learning architecture to predict gene expression from breast cancer histopathology imagesarXiv 2023On the logistical difficulties and findings of Jopara Sentiment AnalysisNAACL (CALCS) 2021 6Posterior Uncertainty Quantification in Neural Networks using Data AugmentationarXiv 2024Towards Trustworthy GUI Agents: A SurveyarXiv 2025On Speeding Up Language Model EvaluationarXiv 2024Automatic Readability Assessment of German Sentences with Transformer Ensemblesautomatic-readability-assessment-of-germanLinear Log-Normal Attention with Unbiased ConcentrationarXiv 2023Modeling Event Plausibility with Consistent Conceptual AbstractionNAACL 2021 4Fine-tuning deep learning model parameters for improved super-resolution of dynamic MRI with prior-knowledgearXiv 2021Deeper Insights into Weight Sharing in Neural Architecture SearcharXiv 2020In Search of Insights, Not Magic Bullets: Towards Demystification of the Model Selection Dilemma in Heterogeneous Treatment Effect EstimationarXiv 2023Pretraining is All You Need: A Multi-Atlas Enhanced Transformer Framework for Autism Spectrum Disorder ClassificationarXiv 2023From Single to Multi: How LLMs Hallucinate in Multi-Document SummarizationarXiv 2024Fighting Fire with Fire: Can ChatGPT Detect AI-generated Text?arXiv 2023Unintended Impacts of LLM Alignment on Global RepresentationarXiv 2024Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?arXiv 2024What's in a Prior? Learned Proximal Networks for Inverse ProblemsarXiv 2023Enhanced Meta Label Correction for Coping with Label CorruptionICCV 2023 1Controlled Text ReductionarXiv 2022NUBES: A Corpus of Negation and Uncertainty in Spanish Clinical Textsnubes-a-corpus-of-negation-and-uncertainty-in-1ImageNet-OOD: Deciphering Modern Out-of-Distribution Detection AlgorithmsarXiv 2023Learning to Collocate Visual-Linguistic Neural Modules for Image CaptioningarXiv 2022LEXI: Large Language Models Experimentation InterfacearXiv 2024MMDisCo: Multi-Modal Discriminator-Guided Cooperative Diffusion for Joint Audio and Video GenerationarXiv 2024Why Are My Prompts Leaked? Unraveling Prompt Extraction Threats in Customized Large Language ModelsarXiv 2024Men Also Do Laundry: Multi-Attribute Bias AmplificationarXiv 2022DoRA: Enhancing Parameter-Efficient Fine-Tuning with Dynamic Rank DistributionarXiv 2024CDEval: A Benchmark for Measuring the Cultural Dimensions of Large Language ModelsarXiv 2023Making deep neural networks right for the right scientific reasons by interacting with their explanationsarXiv 2020CLEAR: Can Language Models Really Understand Causal Graphs?arXiv 20242018 Robotic Scene Segmentation ChallengearXiv 2020Mitigating Entity-Level Hallucination in Large Language ModelsarXiv 2024Deciphering Hate: Identifying Hateful Memes and Their TargetsarXiv 2024From Trojan Horses to Castle Walls: Unveiling Bilateral Data Poisoning Effects in Diffusion ModelsarXiv 2023Learning to Rank Context for Named Entity Recognition Using a Synthetic DatasetarXiv 2023Understanding In-Context Learning from RepetitionsarXiv 2023Novelty Search makes Evolvability InevitablearXiv 2020Label-Agnostic Forgetting: A Supervision-Free Unlearning in Deep ModelsarXiv 2024Swiss Parliaments Corpus, an Automatically Aligned Swiss German Speech to Standard German Text CorpusarXiv 2020PROST: Physical Reasoning of Objects through Space and TimearXiv 2021How Susceptible are Large Language Models to Ideological Manipulation?arXiv 2024MedSumm: A Multimodal Approach to Summarizing Code-Mixed Hindi-English Clinical QueriesarXiv 2024Neural Weight Search for Scalable Task Incremental LearningarXiv 2022Augmenting Passage Representations with Query Generation for Enhanced Cross-Lingual Dense RetrievalarXiv 2023Code-Switching Red-Teaming: LLM Evaluation for Safety and Multilingual UnderstandingarXiv 2024Online Cascade Learning for Efficient Inference over StreamsarXiv 2024Can LMs Generalize to Future Data? An Empirical Analysis on Text SummarizationarXiv 2023FAMMA: A Benchmark for Financial Domain Multilingual Multimodal Question AnsweringarXiv 2024DP-SGD Without Clipping: The Lipschitz Neural Network WayarXiv 2023MuseChat: A Conversational Music Recommendation System for VideosCVPR 2024 1AmbieGen: A Search-based Framework for Autonomous Systems TestingarXiv 2023Privacy-Aware Energy Consumption Modeling of Connected Battery Electric Vehicles using Federated LearningarXiv 2023Histopathological Image Classification based on Self-Supervised Vision Transformer and Weak LabelsarXiv 2022DAMO-StreamNet: Optimizing Streaming Perception in Autonomous DrivingarXiv 2023Peek Across: Improving Multi-Document Modeling via Cross-Document Question-AnsweringarXiv 2023Adapting to game trees in zero-sum imperfect information gamesarXiv 2022The Curious Decline of Linguistic Diversity: Training Language Models on Synthetic TextarXiv 2023Unsupervised Domain Adaptive Detection with Network Stability AnalysisICCV 2023 1Noise transfer for unsupervised domain adaptation of retinal OCT imagesarXiv 2022An Empirical Comparison of Vocabulary Expansion and Initialization Approaches for Language ModelsarXiv 2024Graph Representation Learning for Road Type ClassificationarXiv 2021Weak Proxies are Sufficient and Preferable for Fairness with Missing Sensitive AttributesarXiv 2022Probabilistic Circuits That Know What They Don't KnowarXiv 2023Privacy- and Utility-Preserving NLP with Anonymized Data: A case study of PseudonymizationarXiv 2023LLM4VV: Developing LLM-Driven Testsuite for Compiler ValidationarXiv 2023Acceptable Use Policies for Foundation ModelsarXiv 2024VLN-PETL: Parameter-Efficient Transfer Learning for Vision-and-Language NavigationICCV 2023 1Leveraging Optimization for Adaptive Attacks on Image WatermarksarXiv 2023Skim-Attention: Learning to Focus via Document LayoutFindings (EMNLP) 2021 11The Unmet Promise of Synthetic Training Images: Using Retrieved Real Images Performs BetterarXiv 2024Fortify the Shortest Stave in Attention: Enhancing Context Awareness of Large Language Models for Effective Tool UsearXiv 2023Scalable Real-Time Recurrent Learning Using Columnar-Constructive NetworksarXiv 2023Causal Direction of Data Collection Matters: Implications of Causal and Anticausal Learning for NLPEMNLP 2021 11APOLLO: An Optimized Training Approach for Long-form Numerical ReasoningarXiv 2022WildRefer: 3D Object Localization in Large-scale Dynamic Scenes with Multi-modal Visual Data and Natural LanguagearXiv 2023MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement LearningarXiv 2025What Can We Learn From Almost a Decade of Food TweetsarXiv 2020Bounds on Representation-Induced Confounding Bias for Treatment Effect EstimationarXiv 2023MELA: Multilingual Evaluation of Linguistic AcceptabilityarXiv 2023DiffuMask: Synthesizing Images with Pixel-level Annotations for Semantic Segmentation Using Diffusion ModelsICCV 2023 1Machine-generated text detection prevents language model collapsearXiv 2025MultiLegalSBD: A Multilingual Legal Sentence Boundary Detection DatasetarXiv 2023Enhancing Small Medical Learners with Privacy-preserving Contextual PromptingarXiv 2023Evaluating the Zero-shot Robustness of Instruction-tuned Language ModelsarXiv 2023Unsupervised Matching of Data and TextarXiv 2021Toward Advancing License Plate Super-Resolution in Real-World Scenarios: A Dataset and BenchmarkarXiv 2025Image Embedding for Denoising Generative ModelsarXiv 2022SAMPart3D: Segment Any Part in 3D ObjectsarXiv 2024Exact Gradients for Stochastic Spiking Neural Networks Driven by Rough SignalsarXiv 2024MiDe22: An Annotated Multi-Event Tweet Dataset for Misinformation DetectionarXiv 2022VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward ModelsICCV 2025Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEsarXiv 2025Beyond Worst-case Attacks: Robust RL with Adaptive Defense via Non-dominated PoliciesarXiv 2024Identifying and Manipulating Personality Traits in LLMs Through Activation EngineeringarXiv 2024Could Thinking Multilingually Empower LLM Reasoning?arXiv 2025Relaxing the Additivity Constraints in Decentralized No-Regret High-Dimensional Bayesian OptimizationarXiv 2023Process-based Self-Rewarding Language ModelsarXiv 2025SAEs Can Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMsarXiv 2025SMART: Submodular Data Mixture Strategy for Instruction TuningarXiv 2024SemEval-2025 Task 5: LLMs4Subjects -- LLM-based Automated Subject Tagging for a National Technical Library's Open-Access CatalogarXiv 2025I&S-ViT: An Inclusive & Stable Method for Pushing the Limit of Post-Training ViTs QuantizationarXiv 2023Stochastic Marginal Likelihood Gradients using Neural Tangent KernelsarXiv 2023Building Flexible, Scalable, and Machine Learning-ready Multimodal Oncology DatasetsarXiv 2023RadAdapt: Radiology Report Summarization via Lightweight Domain Adaptation of Large Language ModelsarXiv 2023Cost-of-Pass: An Economic Framework for Evaluating Language ModelsarXiv 2025Beyond Reward: Offline Preference-guided Policy OptimizationarXiv 2023Causal isotonic calibration for heterogeneous treatment effectsarXiv 2023AdaptiVocab: Enhancing LLM Efficiency in Focused Domains through Lightweight Vocabulary AdaptationarXiv 2025Knowledge of Knowledge: Exploring Known-Unknowns Uncertainty with Large Language ModelsarXiv 2023User Story Tutor (UST) to Support Agile Software DevelopersarXiv 2024NLU on Data Diets: Dynamic Data Subset Selection for NLP Classification TasksarXiv 2023Extending Activation Steering to Broad Skills and Multiple BehavioursarXiv 2024Composition-contrastive Learning for Sentence EmbeddingsarXiv 2023Logical Reasoning over Natural Language as Knowledge Representation: A SurveyarXiv 2023ORAN-Bench-13K: An Open Source Benchmark for Assessing LLMs in Open Radio Access NetworksarXiv 2024MPCODER: Multi-user Personalized Code Generator with Explicit and Implicit Style Representation LearningarXiv 2024Joint Metrics Matter: A Better Standard for Trajectory ForecastingICCV 2023 1The Emergence of Essential Sparsity in Large Pre-trained Models: The Weights that Matterthe-emergence-of-essential-sparsity-in-largeDisparate Vulnerability to Membership Inference AttacksarXiv 2019I-AI: A Controllable & Interpretable AI System for Decoding Radiologists' Intense Focus for Accurate CXR DiagnosesarXiv 2023On Cross-Layer Alignment for Model Fusion of Heterogeneous Neural Networksmodel-fusion-of-heterogeneous-neural-networksAdaptiveLog: An Adaptive Log Analysis Framework with the Collaboration of Large and Small Language ModelarXiv 2025Can Vision-Language Models be a Good Guesser? Exploring VLMs for Times and Location ReasoningarXiv 2023Unleashing Mask: Explore the Intrinsic Out-of-Distribution Detection CapabilityarXiv 2023Good at captioning, bad at counting: Benchmarking GPT-4V on Earth observation dataarXiv 2024Unbiased Learning to Rank Meets Reality: Lessons from Baidu's Large-Scale Search DatasetarXiv 2024Language Models for German Text Simplification: Overcoming Parallel Data Scarcity through Style-specific Pre-trainingarXiv 2023Verifiable Goal Recognition for Autonomous Driving with OcclusionsarXiv 2022Witness Generation for JSON SchemaarXiv 2022Neural Parameter Allocation Searchneural-parameter-allocation-searchSystematic Rectification of Language Models via Dead-end AnalysisarXiv 2023Do Large Language Models Truly Understand Geometric Structures?arXiv 2025Copyright Violations and Large Language ModelsarXiv 2023Zero-shot causal learningzero-shot-causal-learningFAENet: Frame Averaging Equivariant GNN for Materials ModelingarXiv 2023Efficient Document Re-Ranking for Transformers by Precomputing Term RepresentationsarXiv 2020ODE Discovery for Longitudinal Heterogeneous Treatment Effects InferencearXiv 2024"Why did the Model Fail?": Attributing Model Performance Changes to Distribution ShiftsarXiv 2022SynTQA: Synergistic Table-based Question Answering via Mixture of Text-to-SQL and E2E TQAarXiv 2024CLIN-X: pre-trained language models and a study on cross-task transfer for concept extraction in the clinical domainarXiv 2021Skin Deep Unlearning: Artefact and Instrument Debiasing in the Context of Melanoma ClassificationarXiv 2021The Many Dimensions of Truthfulness: Crowdsourcing Misinformation Assessments on a Multidimensional ScalearXiv 2021A Theoretical Explanation for Perplexing Behaviors of Backpropagation-based Visualizationsa-theoretical-explanation-for-perplexing-1LPViT: Low-Power Semi-structured Pruning for Vision TransformersarXiv 2024