All papers
Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.
An Empirical Analysis of Uncertainty in Large Language Model EvaluationsarXiv 2025See What LLMs Cannot Answer: A Self-Challenge Framework for Uncovering LLM WeaknessesarXiv 2024SyntaxShap: Syntax-aware Explainability Method for Text GenerationarXiv 2024ReSpAct: Harmonizing Reasoning, Speaking, and Acting Towards Building Large Language Model-Based Conversational AI AgentsarXiv 2024Sifting through the Chaff: On Utilizing Execution Feedback for Ranking
the Generated Code CandidatesarXiv 2024Signing the Supermask: Keep, Hide, Invertsigning-the-supermask-keep-hide-invertMulti-Agent MDP Homomorphic Networksmulti-agent-mdp-homomorphic-networksAntiLeak-Bench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World KnowledgearXiv 2024Slot Filling for Biomedical Information ExtractionBioNLP (ACL) 2022 5Hope Speech detection in under-resourced Kannada languagearXiv 2021Evaluating Cultural and Social Awareness of LLM Web AgentsarXiv 2024GeLLMO: Generalizing Large Language Models for Multi-property Molecule OptimizationarXiv 2025Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the EdgearXiv 2023Narrative Question Answering with Cutting-Edge Open-Domain QA Techniques: A Comprehensive StudyarXiv 2021On the Generalization of Training-based ChatGPT Detection MethodsarXiv 2023ToMChallenges: A Principle-Guided Dataset and Diverse Evaluation Tasks for Exploring Theory of MindarXiv 2023NevIR: Negation in Neural Information RetrievalarXiv 2023Discovering Novel Biological Traits From Images Using Phylogeny-Guided Neural NetworksarXiv 2023Can LLMs Master Math? Investigating Large Language Models on Math Stack ExchangearXiv 2024μ-Bench: A Vision-Language Benchmark for Microscopy UnderstandingarXiv 2024FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio GenerationarXiv 2024The Open DAC 2023 Dataset and Challenges for Sorbent Discovery in Direct Air CapturearXiv 2023Towards Explaining Distribution ShiftsarXiv 2022Compositional preference models for aligning LMsarXiv 2023Sparse Pairwise Re-ranking with Pre-trained TransformersarXiv 2022Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language ModelsarXiv 2025DiSCo: LLM Knowledge Distillation for Efficient Sparse Retrieval in Conversational SearcharXiv 2024Establishing Strong Baselines for TripClick Health RetrievalarXiv 2022Towards Improved Input Masking for Convolutional Neural NetworksICCV 2023 1Improving (Dis)agreement Detection with Inductive Social Relation Information From Comment-Reply InteractionsarXiv 2023Interpreting Pretrained Language Models via Concept BottlenecksarXiv 2023Negation detection in Dutch clinical texts: an evaluation of rule-based and machine learning methodsarXiv 2022mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text RetrievalarXiv 2024Identity-Consistent Aggregation for Video Object DetectionICCV 2023 1FARE: Provably Fair Representation Learning with Practical CertificatesarXiv 2022Interpretable Clustering: A SurveyarXiv 2024Quality and Quantity of Machine Translation References for Automatic MetricsarXiv 2024Self-supervised Meta-Prompt Learning with Meta-Gradient Regularization for Few-shot GeneralizationarXiv 2023Leveraging Label Non-Uniformity for Node Classification in Graph Neural NetworksarXiv 2023RoleMRC: A Fine-Grained Composite Benchmark for Role-Playing and Instruction-FollowingarXiv 2025Should We Really Edit Language Models? On the Evaluation of Edited Language ModelsarXiv 2024CasiMedicos-Arg: A Medical Question Answering Dataset Annotated with Explanatory Argumentative StructuresarXiv 2024Quantifying Uncertainty in Motion Prediction with Variational Bayesian MixtureCVPR 2024 1Adaptation Strategies for Automated Machine Learning on Evolving DataarXiv 2020Iterative Self-Tuning LLMs for Enhanced Jailbreaking CapabilitiesarXiv 2024Human Learning by Model Feedback: The Dynamics of Iterative Prompting with MidjourneyarXiv 2023DropNAS: Grouped Operation Dropout for Differentiable Architecture SearcharXiv 2022ReasoningRec: Bridging Personalized Recommendations and Human-Interpretable Explanations through LLM ReasoningarXiv 2024Retrieval-Pretrained Transformer: Long-range Language Modeling with Self-retrievalarXiv 2023Double-Weighting for Covariate Shift AdaptationarXiv 2023Detecting Arbitrary Keypoints on Limbs and Skis with Sparse Partly Correct Segmentation MasksarXiv 2022MICDIR: Multi-scale Inverse-consistent Deformable Image Registration using UNetMSS with Self-Constructing Graph LatentarXiv 2022DAFA: Distance-Aware Fair Adversarial TrainingarXiv 2024Understanding the differences in Foundation Models: Attention, State Space Models, and Recurrent Neural NetworksarXiv 2024UniAff: A Unified Representation of Affordances for Tool Usage and Articulation with Vision-Language ModelsarXiv 2024USimAgent: Large Language Models for Simulating Search UsersarXiv 2024Bridging Context Gaps: Leveraging Coreference Resolution for Long Contextual UnderstandingarXiv 2024Repeated Random Sampling for Minimizing the Time-to-Accuracy of LearningarXiv 2023Eliciting and Understanding Cross-Task Skills with Task-Level Mixture-of-ExpertsarXiv 2022Fantastic Bugs and Where to Find Them in AI BenchmarksarXiv 2025How Do Transformers Learn Topic Structure: Towards a Mechanistic UnderstandingarXiv 2023SafeAligner: Safety Alignment against Jailbreak Attacks via Response Disparity GuidancearXiv 2024Generative Principal Component Analysisgenerative-principal-component-analysisSpeechformer: Reducing Information Loss in Direct Speech TranslationEMNLP 2021 11Finding Biological Plausibility for Adversarially Robust Features via Metameric Tasksfinding-biological-plausibility-for-1Sparse Reward Exploration via Novelty Search and EmittersarXiv 2021SEGA: Structural Entropy Guided Anchor View for Graph Contrastive LearningarXiv 2023GreekBART: The First Pretrained Greek Sequence-to-Sequence ModelarXiv 2023Improving Weak-to-Strong Generalization with Reliability-Aware AlignmentarXiv 2024PeFLL: Personalized Federated Learning by Learning to LearnarXiv 2023On the Stability of Iterative Retraining of Generative Models on their own DataarXiv 2023Regions of Reliability in the Evaluation of Multivariate Probabilistic ForecastsarXiv 2023mCoT: Multilingual Instruction Tuning for Reasoning Consistency in Language ModelsarXiv 2024BertaQA: How Much Do Language Models Know About Local Culture?arXiv 2024PPM: Automated Generation of Diverse Programming Problems for Benchmarking Code Generation ModelsarXiv 2024A User-Friendly Framework for Generating Model-Preferred Prompts in Text-to-Image SynthesisarXiv 2024The Data Addition DilemmaarXiv 2024Optimally-Weighted Estimators of the Maximum Mean Discrepancy for Likelihood-Free InferencearXiv 2023Inducing Neural Collapse in Deep Long-tailed LearningarXiv 2023A benchmark for toxic comment classification on Civil Comments datasetarXiv 2023AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific ResearcharXiv 2025Enhancing LLM Robustness to Perturbed Instructions: An Empirical StudyarXiv 2025Exploring Design Choices for Building Language-Specific LLMsarXiv 2024Hyperpolyglot LLMs: Cross-Lingual Interpretability in Token EmbeddingsarXiv 2023Blinded by Generated Contexts: How Language Models Merge Generated and Retrieved Contexts When Knowledge Conflicts?arXiv 2024Insect Identification in the Wild: The AMI DatasetarXiv 2024Social perception of faces in a vision-language modelarXiv 2024Revisiting In-context Learning Inference Circuit in Large Language ModelsarXiv 2024Assessing Social and Intersectional Biases in Contextualized Word Representationsassessing-social-and-intersectional-biases-in-1Blockwise Self-Attention for Long Document UnderstandingFindings of the Association for Computational Linguistics 2020A Comprehensive Evaluation of Quantization Strategies for Large Language ModelsarXiv 2024A cost-benefit analysis of cross-lingual transfer methodsarXiv 2021GanLM: Encoder-Decoder Pre-training with an Auxiliary DiscriminatorarXiv 2022Towards the generation of synchronized and believable non-verbal facial behaviors of a talking virtual agentarXiv 2023ScoNe: Benchmarking Negation Reasoning in Language Models With Fine-Tuning and In-Context LearningarXiv 2023Towards Enhancing Coherence in Extractive Summarization: Dataset and Experiments with LLMsarXiv 2024Anchored Answers: Unravelling Positional Bias in GPT-2's Multiple-Choice QuestionsarXiv 2024Retrieval Augmented Generation using Engineering Design KnowledgearXiv 2023Automatic Evaluation and Moderation of Open-domain Dialogue SystemsarXiv 2021Bridging the Sim-to-Real Gap from the Information Bottleneck PerspectivearXiv 2023Rejuvenating Low-Frequency Words: Making the Most of Parallel Data in Non-Autoregressive TranslationACL 2021 5Chordal Averaging on Flag Manifolds and Its ApplicationsICCV 2023 1FOCUS: Familiar Objects in Common and Uncommon Settingsfocus-familiar-objects-in-common-and-uncommon-1Bilevel Optimization under Unbounded Smoothness: A New Algorithm and Convergence AnalysisarXiv 2024Revisiting Uncertainty-based Query Strategies for Active Learning with TransformersFindings (ACL) 2022 5Conformal inference is (almost) free for neural networks trained with early stoppingarXiv 2023A Scalable AutoML Approach Based on Graph Neural Networksa-scalable-automl-approach-based-on-graph-1ClassActionPrediction: A Challenging Benchmark for Legal Judgment Prediction of Class Action Cases in the USarXiv 2022FSUIE: A Novel Fuzzy Span Mechanism for Universal Information ExtractionarXiv 2023ProcBench: Benchmark for Multi-Step Reasoning and Following ProcedurearXiv 2024An Empirical Analysis of Forgetting in Pre-trained Models with Incremental Low-Rank UpdatesarXiv 2024PhayaThaiBERT: Enhancing a Pretrained Thai Language Model with Unassimilated LoanwordsarXiv 2023LLM Agents in Interaction: Measuring Personality Consistency and Linguistic Alignment in Interacting Populations of Large Language ModelsarXiv 2024SentMix-3L: A Bangla-English-Hindi Code-Mixed Dataset for Sentiment AnalysisarXiv 2023MedExQA: Medical Question Answering Benchmark with Multiple ExplanationsarXiv 2024Deconfounding Legal Judgment Prediction for European Court of Human Rights Cases Towards Better Alignment with ExpertsarXiv 2022Graph-Based Multilingual Label Propagation for Low-Resource Part-of-Speech TaggingarXiv 2022Multilingual Detection of Personal Employment Status on TwitterACL 2022 5Policy Improvement using Language Feedback ModelsarXiv 2024Visual explanation of black-box model: Similarity Difference and Uniqueness (SIDU) methodarXiv 2021In-context Learning and Gradient Descent RevisitedarXiv 2023Deep Anomaly Detection under Labeling Budget ConstraintsarXiv 2023From Principle to Practice: Vertical Data Minimization for Machine LearningarXiv 2023Don't Just Say "I don't know"! Self-aligning Large Language Models for Responding to Unknown Questions with ExplanationsarXiv 2024A Surprisingly Simple yet Effective Multi-Query Rewriting Method for Conversational Passage RetrievalarXiv 2024A Search Engine for Discovery of Scientific Challenges and DirectionsNeurIPS Workshop AI4Scien 2021 12Dual Propagation: Accelerating Contrastive Hebbian Learning with Dyadic NeuronsarXiv 2023Rethinking Channel Dimensions to Isolate Outliers for Low-bit Weight Quantization of Large Language ModelsarXiv 2023Large Language Models of Code Fail at Completing Code with Potential Bugslarge-language-models-of-code-fail-atUnable to Forget: Proactive lnterference Reveals Working Memory Limits in LLMs Beyond Context LengtharXiv 2025Radio Galaxy Zoo: Using semi-supervised learning to leverage large unlabelled data-sets for radio galaxy classification under data-set shiftarXiv 2022PFB-Diff: Progressive Feature Blending Diffusion for Text-driven Image EditingarXiv 2023Evaluating the Instruction-Following Robustness of Large Language Models to Prompt InjectionarXiv 2023Memotion 3: Dataset on Sentiment and Emotion Analysis of Codemixed Hindi-English MemesarXiv 2023The Possible, the Plausible, and the Desirable: Event-Based Modality Detection for Language ProcessingACL 2021 5On the Importance of Backbone to the Adversarial Robustness of Object DetectorsarXiv 2023Does Table Source Matter? Benchmarking and Improving Multimodal Scientific Table Understanding and ReasoningarXiv 2025CoReS: Compatible Representations via StationarityarXiv 2021AdvisorQA: Towards Helpful and Harmless Advice-seeking Question Answering with Collective IntelligencearXiv 2024RNNs of RNNs: Recursive Construction of Stable Assemblies of Recurrent Neural Networksrecursive-construction-of-stable-assemblies-1A Conditional Normalizing Flow for Accelerated Multi-Coil MR ImagingarXiv 2023Emergent Asymmetry of Precision and Recall for Measuring Fidelity and Diversity of Generative Models in High DimensionsarXiv 2023Scaling Down to Scale Up: A Guide to Parameter-Efficient Fine-TuningarXiv 2023Node Proximity Is All You Need: Unified Structural and Positional Node and Graph EmbeddingarXiv 2021Experimental Analysis of Large-scale Learnable Vector Storage CompressionarXiv 2023MATHSENSEI: A Tool-Augmented Large Language Model for Mathematical ReasoningarXiv 2024BLAB: Brutally Long Audio BencharXiv 2025Pandora's Box or Aladdin's Lamp: A Comprehensive Analysis Revealing the Role of RAG Noise in Large Language ModelsarXiv 2024HLT-MT: High-resource Language-specific Training for Multilingual Neural Machine TranslationarXiv 2022CTRL-ALT-LED: Leaking Data from Air-Gapped Computers via Keyboard LEDsarXiv 2019Gene Regulatory Network Inference in the Presence of Dropouts: a Causal ViewarXiv 2024GLU Variants Improve TransformerarXiv 2020IndicIRSuite: Multilingual Dataset and Neural Information Models for Indian LanguagesarXiv 2023Inferring Implicit Relations in Complex Questions with Language ModelsarXiv 2022Topic-oriented Adversarial Attacks against Black-box Neural Ranking ModelsarXiv 2023TableQA: a Large-Scale Chinese Text-to-SQL Dataset for Table-Aware SQL
GenerationarXiv 2020BrackishMOT: The Brackish Multi-Object Tracking DatasetarXiv 2023SciDQA: A Deep Reading Comprehension Dataset over Scientific PapersarXiv 2024Robust model benchmarking and bias-imbalance in data-driven materials
science: a case study on MODNetarXiv 2021Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense EvaluationarXiv 2025CaPo: Cooperative Plan Optimization for Efficient Embodied Multi-Agent CooperationarXiv 2024Reliable and Efficient Amortized Model-based EvaluationarXiv 2025M3GIA: A Cognition Inspired Multilingual and Multimodal General Intelligence Ability BenchmarkarXiv 2024Enhancing Tool Retrieval with Iterative Feedback from Large Language ModelsarXiv 2024Large Language Models Assume People are More Rational than We Really arearXiv 2024ACAT: Adversarial Counterfactual Attention for Classification and Detection in Medical ImagingarXiv 2023Does Chain-of-Thought Reasoning Help Mobile GUI Agent? An Empirical StudyarXiv 2025Task-Agnostic Low-Rank Adapters for Unseen English DialectsarXiv 2023HyperSparse Neural Networks: Shifting Exploration to Exploitation through Adaptive RegularizationarXiv 2023Automated Code-centric Software Vulnerability Assessment: How Far Are We? An Empirical Study in C/C++arXiv 2024A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPTarXiv 2023Machine Learning and Deep Learning -- A review for EcologistsarXiv 2022TIBET: Identifying and Evaluating Biases in Text-to-Image Generative ModelsarXiv 2023Prediction of the motion of chest internal points using a recurrent neural network trained with real-time recurrent learning for latency compensation in lung cancer radiotherapyarXiv 2022MABe22: A Multi-Species Multi-Task Benchmark for Learned Representations
of BehaviorarXiv 2022Distributional Reinforcement Learning for Multi-Dimensional Reward FunctionsNeurIPS 2021 12DPTDR: Deep Prompt Tuning for Dense Passage RetrievalCOLING 2022 10Challenge on Sound Scene Synthesis: Evaluating Text-to-Audio GenerationarXiv 2024Weakly Supervised Semantic Segmentation via Progressive Patch LearningarXiv 2022Data Similarity is Not Enough to Explain Language Model PerformancearXiv 2023Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous WordsarXiv 2025Exploring the Inquiry-Diagnosis Relationship with Advanced Patient SimulatorsarXiv 2025EssayJudge: A Multi-Granular Benchmark for Assessing Automated Essay Scoring Capabilities of Multimodal Large Language ModelsarXiv 2025Improved Visual Fine-tuning with Natural Language SupervisionICCV 2023 1An Open-World, Diverse, Cross-Spatial-Temporal Benchmark for Dynamic Wild Person Re-IdentificationarXiv 2024Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-TuningarXiv 2025Looped Transformers as Programmable ComputersarXiv 2023MoleculeQA: A Dataset to Evaluate Factual Accuracy in Molecular ComprehensionarXiv 2024Bilingual Dual-Head Deep Model for Parkinson's Disease Detection from SpeecharXiv 2025SilverSpeak: Evading AI-Generated Text Detectors using HomoglyphsarXiv 2024A kernel Stein test of goodness of fit for sequential modelsarXiv 2022Reinforcement learning with learned gadgets to tackle hard quantum problems on real hardwarearXiv 2024RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language ModelsarXiv 2024How Should We Enhance the Safety of Large Reasoning Models: An Empirical StudyarXiv 2025A Simple Unified Uncertainty-Guided Framework for Offline-to-Online Reinforcement LearningarXiv 2023Dissecting Sample Hardness: A Fine-Grained Analysis of Hardness Characterization Methods for Data-Centric AIarXiv 2024A Tale of Pronouns: Interpretability Informs Gender Bias Mitigation for Fairer Instruction-Tuned Machine TranslationarXiv 2023ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended CapabilitiesarXiv 2024Curiosity-Driven Exploration via Latent Bayesian SurpriseICLR Workshop SSL-RL 2021 5Detecting Unassimilated Borrowings in Spanish: An Annotated Corpus and Approaches to ModelingACL 2022 5