0

All papers

Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.

Multi-News+: Cost-efficient Dataset Cleansing via LLM-based Data AnnotationarXiv 2024DP-SGD vs PATE: Which Has Less Disparate Impact on Model Accuracy?arXiv 2021SummIt: Iterative Text Summarization via ChatGPTarXiv 2023A Theoretical Analysis of the Learning Dynamics under Class ImbalancearXiv 2022Greedy-DiM: Greedy Algorithms for Unreasonably Effective Face MorphsarXiv 2024Anchor Sampling for Federated Learning with Partial Client ParticipationarXiv 2022Fine-tuning Large Language Models for Improving Factuality in Legal Question AnsweringarXiv 2025Kaleidoscope: In-language Exams for Massively Multilingual Vision EvaluationarXiv 2025All you need is spin: SU(2) equivariant variational quantum circuits based on spin networksarXiv 2023Enhancing textual textbook question answering with large language models and retrieval augmented generationarXiv 2024HaluEval-Wild: Evaluating Hallucinations of Language Models in the WildarXiv 2024Sinogram upsampling using Primal-Dual UNet for undersampled CT and radial MRI reconstructionarXiv 2021Do LLMs Really Think Step-by-step In Implicit Reasoning?arXiv 2024Condensed Gradient BoostingarXiv 2022Inducing Neural Collapse to a Fixed Hierarchy-Aware Frame for Reducing Mistake SeverityICCV 2023 1Are Personalized Stochastic Parrots More Dangerous? Evaluating Persona Biases in Dialogue SystemsarXiv 2023Aspects of human memory and Large Language ModelsarXiv 2023Edge-based sequential graph generation with recurrent neural networksarXiv 2020Attention Is Indeed All You Need: Semantically Attention-Guided Decoding for Data-to-Text NLGINLG (ACL) 2021 8VacancySBERT: the approach for representation of titles and skills for semantic similarity search in the recruitment domainarXiv 2023Protocols for creating and distilling multipartite GHZ states with Bell pairsarXiv 2020A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and RecommendationsarXiv 2024Mixture of Nested Experts: Adaptive Processing of Visual TokensarXiv 2024Free Lunch: Robust Cross-Lingual Transfer via Model Checkpoint AveragingarXiv 2023Fast model inference and training on-board of SatellitesarXiv 2023GROOViST: A Metric for Grounding Objects in Visual StorytellingarXiv 2023AraStance: A Multi-Country and Multi-Domain Dataset of Arabic Stance Detection for Fact CheckingNAACL (NLP4IF) 2021 6Unsupervised Learning under Latent Label ShiftarXiv 2022Sexism Prediction in Spanish and English Tweets Using Monolingual and Multilingual BERT and Ensemble ModelsarXiv 2021Wave to Syntax: Probing spoken language models for syntaxarXiv 2023Structural Similarities Between Language Models and Neural Response MeasurementsarXiv 2023How Well Can LLMs Echo Us? Evaluating AI Chatbots' Role-Play Ability with ECHOarXiv 2024Examining Forgetting in Continual Pre-training of Aligned Large Language ModelsarXiv 2024The Impact of Prompt Programming on Function-Level Code GenerationarXiv 2024Continual Learning for Monolingual End-to-End Automatic Speech RecognitionarXiv 2021Robust Pronoun Fidelity with English LLMs: Are they Reasoning, Repeating, or Just Biased?arXiv 2024Image Inpainting via Tractable Steering of Diffusion ModelsarXiv 2023ImplicitMemBench: Measuring Unconscious Behavioral Adaptation in Large Language ModelsarXiv 2026Bridging Relevance and Reasoning: Rationale Distillation in Retrieval-Augmented GenerationarXiv 2024From Artificial Needles to Real Haystacks: Improving Retrieval Capabilities in LLMs by Finetuning on Synthetic DataarXiv 2024Post-hoc Bias Scoring Is Optimal For Fair ClassificationarXiv 2023Scalable Data Ablation Approximations for Language Models through Modular Training and MergingarXiv 2024Expectation-Complete Graph Representations with HomomorphismsarXiv 2023Towards Efficiently Diversifying Dialogue Generation via Embedding AugmentationarXiv 2021ASSERT: Automated Safety Scenario Red Teaming for Evaluating the Robustness of Large Language ModelsarXiv 2023MalAlgoQA: Pedagogical Evaluation of Counterfactual Reasoning in Large Language Models and Implications for AI in EducationarXiv 2024A Linear Reconstruction Approach for Attribute Inference Attacks against Synthetic DataarXiv 2023It's Not Easy Being Wrong: Large Language Models Struggle with Process of Elimination ReasoningarXiv 2023A Coupled Flow Approach to Imitation LearningarXiv 2023Gibbsian polar slice samplingarXiv 2023LLM Can be a Dangerous Persuader: Empirical Study of Persuasion Safety in Large Language ModelsarXiv 2025Clustering Cluster Algebras with ClustersarXiv 2022Behind the Mask: Demographic bias in name detection for PII maskingLTEDI (ACL) 2022 5A Novel Contrastive Learning Method for Clickbait Detection on RoCliCo: A Romanian Clickbait Corpus of News ArticlesarXiv 2023Towards Collaborative Plan Acquisition through Theory of Mind Modeling in Situated DialoguearXiv 2023The Impact of Cross-Lingual Adjustment of Contextual Word Representations on Zero-Shot TransferarXiv 2022A Contrastive Learning Approach to Mitigate Bias in Speech ModelsarXiv 2024GAVEL: Generating Games Via Evolution and Language ModelsarXiv 2024Transformers Can Represent $n$-gram Language ModelsarXiv 2024GTrans: Grouping and Fusing Transformer Layers for Neural Machine TranslationarXiv 2022Exploring Weight Balancing on Long-Tailed Recognition ProblemarXiv 2023Lipschitzness Is All You Need To Tame Off-policy Generative Adversarial Imitation LearningarXiv 2020Robust AI-Generated Text Detection by Restricted EmbeddingsarXiv 2024Data Representations' Study of Latent Image ManifoldsarXiv 2023Towards Efficient and Explainable Hate Speech Detection via Model DistillationarXiv 2024Mathematical Reasoning in Large Language Models: Assessing Logical and Arithmetic Errors across Wide Numerical RangesarXiv 2025When Neural Code Completion Models Size up the Situation: Attaining Cheaper and Faster Completion through Dynamic Model InferencearXiv 2024Clembench: Using Game Play to Evaluate Chat-Optimized Language Models as Conversational AgentsarXiv 2023ALDi: Quantifying the Arabic Level of Dialectness of TextarXiv 2023SEMICON: A Learning-to-hash Solution for Large-scale Fine-grained Image RetrievalarXiv 2022Academically intelligent LLMs are not necessarily socially intelligentarXiv 2024Learning Conformal Abstention Policies for Adaptive Risk Management in Large Language and Vision-Language ModelsarXiv 2025Learning to Reason via Program Generation, Emulation, and SearcharXiv 2024LAHAJA: A Robust Multi-accent Benchmark for Evaluating Hindi ASR SystemsarXiv 2024Emergent Properties of Foveated Perceptual SystemsarXiv 2020Understanding writing style in social media with a supervised contrastively pre-trained transformerarXiv 2023DISGAN: Wavelet-informed Discriminator Guides GAN to MRI Super-resolution with Noise CleaningarXiv 2023Vi(E)va LLM! A Conceptual Stack for Evaluating and Interpreting Generative AI-based VisualizationsarXiv 2024Some Fundamental Aspects about Lipschitz Continuity of Neural NetworksarXiv 2023Sim2Rec: A Simulator-based Decision-making Approach to Optimize Real-World Long-term User Engagement in Sequential Recommender SystemsarXiv 2023Linear Cross-document Event Coreference Resolution with X-AMRarXiv 2024Shedding More Light on Robust Classifiers under the lens of Energy-based ModelsarXiv 2024Distilling ChatGPT for Explainable Automated Student Answer AssessmentarXiv 2023How to Synthesize Text Data without Model Collapse?arXiv 2024Improving Factuality in Large Language Models via Decoding-Time Hallucinatory and Truthful ComparatorsarXiv 2024JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesisarXiv 2017Whitened CLIP as a Likelihood Surrogate of Images and CaptionsarXiv 2025Rethinking Chunk Size For Long-Document Retrieval: A Multi-Dataset AnalysisarXiv 2025Mapping and Influencing the Political Ideology of Large Language Models using Synthetic PersonasarXiv 2024AROID: Improving Adversarial Robustness Through Online Instance-Wise Data AugmentationarXiv 2023Federated Learning via Plurality Votefederated-learning-via-plurality-vote-1Constrained Decoding for Cross-lingual Label ProjectionarXiv 2024Trans4D: Realistic Geometry-Aware Transition for Compositional Text-to-4D SynthesisarXiv 2024ModSCAN: Measuring Stereotypical Bias in Large Vision-Language Models from Vision and Language ModalitiesarXiv 2024Investigating the Impact of Quantization Methods on the Safety and Reliability of Large Language ModelsarXiv 2025Can Large Language Models Capture Dissenting Human Voices?arXiv 2023BeMap: Balanced Message Passing for Fair Graph Neural NetworkarXiv 2023Visual Search Asymmetry: Deep Nets and Humans Share Similar Inherent BiasesNeurIPS 2021 12CochlScene: Acquisition of acoustic scene data using crowdsourcingarXiv 2022Discrete Prompt Optimization via Constrained Generation for Zero-shot Re-rankerarXiv 2023FACESEC: A Fine-grained Robustness Evaluation Framework for Face Recognition SystemsCVPR 2021 1Counterfactual Plans under Distributional Ambiguitycounterfactual-plans-under-distributionalAQE: Argument Quadruplet Extraction via a Quad-Tagging Augmented Generative ApproacharXiv 2023TMGBench: A Systematic Game Benchmark for Evaluating Strategic Reasoning Abilities of LLMsarXiv 2024Think Before You Speak: Cultivating Communication Skills of Large Language Models via Inner MonologuearXiv 2023Why do These Match? Explaining the Behavior of Image Similarity ModelsECCV 2020 8Dr-LLaVA: Visual Instruction Tuning with Symbolic Clinical GroundingarXiv 2024X-Pruner: eXplainable Pruning for Vision TransformersCVPR 2023 1MEND: Meta dEmonstratioN Distillation for Efficient and Effective In-Context LearningarXiv 2024Lost in Space: Probing Fine-grained Spatial Understanding in Vision and Language ResamplersarXiv 2024Self-Judge: Selective Instruction Following with Alignment Self-EvaluationarXiv 2024Use HiResCAM instead of Grad-CAM for faithful explanations of convolutional neural networksarXiv 2020Are Character-level Translations Worth the Wait? Comparing ByT5 and mT5 for Machine TranslationarXiv 2023PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-RailsarXiv 2024Are LLM-Judges Robust to Expressions of Uncertainty? Investigating the effect of Epistemic Markers on LLM-based EvaluationarXiv 2024Contrastive Attraction and Contrastive Repulsion for Representation Learningcontrastive-attraction-and-contrastiveEfficient Data Selection at Scale via Influence DistillationarXiv 2025CX-ToM: Counterfactual Explanations with Theory-of-Mind for Enhancing Human Trust in Image Recognition ModelsarXiv 2021Do language models plan ahead for future tokens?arXiv 2024RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference DataarXiv 2024JavaBERT: Training a transformer-based model for the Java programming languagearXiv 2021Linguistic Dependencies and Statistical DependenceEMNLP 2021 11Evaluating the Ability of LLMs to Solve Semantics-Aware Process Mining TasksarXiv 2024Bugs in the Data: How ImageNet Misrepresents BiodiversityarXiv 2022Automatic Generation of Contrast Sets from Scene Graphs: Probing the Compositional Consistency of GQANAACL 2021 4Using Natural Language Explanations to Rescale Human JudgmentsarXiv 2023IRCoCo: Immediate Rewards-Guided Deep Reinforcement Learning for Code CompletionarXiv 2024Can Vision-Language Models Evaluate Handwritten Math?arXiv 2025Bayesian Calibration of Win Rate Estimation with LLM EvaluatorsarXiv 2024Can Graph Neural Networks Learn Language with Extremely Weak Text Supervision?arXiv 2024SLIDE: A Framework Integrating Small and Large Language Models for Open-Domain Dialogues EvaluationarXiv 2024Identifying and Adapting Transformer-Components Responsible for Gender Bias in an English Language ModelarXiv 2023Distributional MIPLIB: a Multi-Domain Library for Advancing ML-Guided MILP MethodsarXiv 2024DAMRO: Dive into the Attention Mechanism of LVLM to Reduce Object HallucinationarXiv 2024Learning Optimal Advantage from Preferences and Mistaking it for RewardarXiv 2023Does RAG Introduce Unfairness in LLMs? Evaluating Fairness in Retrieval-Augmented Generation SystemsarXiv 2024Does Reasoning Introduce Bias? A Study of Social Bias Evaluation and Mitigation in LLM ReasoningarXiv 2025Fully Bayesian VIB-DeepSSMarXiv 2023SurfGen: Adversarial 3D Shape Synthesis with Explicit Surface Discriminatorssurfgen-adversarial-3d-shape-synthesis-withYour Attack Is Too DUMB: Formalizing Attacker Scenarios for Adversarial TransferabilityarXiv 2023Demystifying Disagreement-on-the-Line in High DimensionsarXiv 2023Learning Stance Embeddings from Signed Social GraphsarXiv 2022Detecting AI-Generated Sentences in Human-AI Collaborative Hybrid Texts: Challenges, Strategies, and InsightsarXiv 2024J-Guard: Journalism Guided Adversarially Robust Detection of AI-generated NewsarXiv 2023TTIDA: Controllable Generative Data Augmentation via Text-to-Text and Text-to-Image ModelsarXiv 2023CatGCN: Graph Convolutional Networks with Categorical Node FeaturesarXiv 2020NOVUM: Neural Object Volumes for Robust Object ClassificationarXiv 2023IDEA: Enhancing the Rule Learning Ability of Large Language Model Agent through Induction, Deduction, and AbductionarXiv 2024Distillation-based fabric anomaly detectionarXiv 2024Vulnerabilities in AI-generated Image Detection: The Challenge of Adversarial AttacksarXiv 2024Quantile Regression for Distributional Reward Models in RLHFarXiv 2024Large-scale, Language-agnostic Discourse Classification of Tweets During COVID-19arXiv 2020Improving Adversarial Robustness by Putting More Regularizations on Less Robust SamplesarXiv 2022Benchmarking Adversarial Robustness to Bias Elicitation in Large Language Models: Scalable Automated Assessment with LLM-as-a-JudgearXiv 2025Uncertainty-Aware Natural Language Inference with Stochastic Weight AveragingarXiv 2023Mitigating the Backdoor Effect for Multi-Task Model Merging via Safety-Aware SubspacearXiv 2024Learning from Sparse Offline Datasets via Conservative Density EstimationarXiv 2024Accurately and Efficiently Interpreting Human-Robot Instructions of Varying GranularitiesarXiv 2017Living Machines: A study of atypical animacyCOLING 2020 8LLM2: Let Large Language Models Harness System 2 ReasoningarXiv 2024Improving Medical Dialogue Generation with Abstract Meaning RepresentationsarXiv 2023A Unified Model for Reverse Dictionary and Definition ModellingarXiv 2022On Balancing Bias and Variance in Unsupervised Multi-Source-Free Domain AdaptationarXiv 2022Memorization Capacity of Multi-Head Attention in TransformersarXiv 2023MADation: Face Morphing Attack Detection with Foundation ModelsarXiv 2025Efficient and Transferable Adversarial Examples from Bayesian Neural NetworksarXiv 2020Question-Answering Dense Video EventsarXiv 2024Learn from the Learnt: Source-Free Active Domain Adaptation via Contrastive Sampling and Visual PersistencearXiv 2024LLM Chain Ensembles for Scalable and Accurate Data AnnotationarXiv 2024Can Open-Source LLMs Compete with Commercial Models? Exploring the Few-Shot Performance of Current GPT Models in Biomedical TasksarXiv 2024A comparison of Human, GPT-3.5, and GPT-4 Performance in a University-Level Coding CoursearXiv 2024Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks?arXiv 2024VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation ModelarXiv 2024Not (yet) the whole story: Evaluating Visual Storytelling Requires More than Measuring Coherence, Grounding, and RepetitionarXiv 2024Annotation Sensitivity: Training Data Collection Methods Affect Model PerformancearXiv 2023Jamba-1.5: Hybrid Transformer-Mamba Models at ScalearXiv 2024Can Transformers Do Enumerative Geometry?arXiv 2024SimPLe: Similarity-Aware Propagation Learning for Weakly-Supervised Breast Cancer Segmentation in DCE-MRIarXiv 2023AttrSeg: Open-Vocabulary Semantic Segmentation via Attribute Decomposition-Aggregationopen-vocabulary-semantic-segmentation-viaLanguage Models can Exploit Cross-Task In-context Learning for Data-Scarce Novel TasksarXiv 2024Temperature-scaling surprisal estimates improve fit to human reading times -- but does it do so for the "right reasons"?arXiv 2023Why only Micro-F1? Class Weighting of Measures for Relation Classificationwhy-only-micro-f1-class-weighting-of-measuresLMD: Faster Image Reconstruction with Latent Masking DiffusionarXiv 2023TADA: Task-Agnostic Dialect Adapters for EnglisharXiv 2023Neural Collapse in Deep Linear Networks: From Balanced to Imbalanced DataarXiv 2023Large Language Models are In-Context Molecule LearnersarXiv 2024Bangla Handwritten Digit Recognition and GenerationarXiv 2021Dwell in the Beginning: How Language Models Embed Long Documents for Dense RetrievalarXiv 2024Efficient Subgraph GNNs by Learning Effective Selection PoliciesarXiv 2023I Need Help! Evaluating LLM's Ability to Ask for Users' Support: A Case Study on Text-to-SQL GenerationarXiv 2024SimOAP: Improve Coherence and Consistency in Persona-based Dialogue Generation via Over-sampling and Post-evaluationarXiv 2023Interpersonal Memory Matters: A New Task for Proactive Dialogue Utilizing Conversational HistoryarXiv 2025Analysis of learning a flow-based generative model from limited sample complexityarXiv 2023Robust Multi-Objective Controlled Decoding of Large Language ModelsarXiv 2025An Empirical Study of In-context Learning in LLMs for Machine TranslationarXiv 2024StRE: Self Attentive Edit Quality Prediction in Wikipediastre-self-attentive-edit-quality-prediction-1Object-Centric Learning with Slot Mixture ModulearXiv 2023Unveiling the Truth: Exploring Human Gaze Patterns in Fake ImagesarXiv 2024Sem-CS: Semantic CLIPStyler for Text-Based Image Style TransferarXiv 2023Mixtral of ExpertsarXiv 2024

Back to Papers