All papers
Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.
ChartThinker: A Contextual Chain-of-Thought Approach to Optimized Chart SummarizationarXiv 2024Semi-automatic tuning of coupled climate models with multiple intrinsic timescales: lessons learned from the Lorenz96 modelarXiv 2022Multi-sense embeddings through a word sense disambiguation processarXiv 2021Linking In-context Learning in Transformers to Human Episodic MemoryarXiv 2024Shedding Light on Software Engineering-specific Metaphors and IdiomsarXiv 2023Sequence to sequence pretraining for a less-resourced Slovenian languagearXiv 2022Uncovering the Causes of Emotions in Software Developer Communication
Using Zero-shot LLMsarXiv 2023Generating High-Precision Feedback for Programming Syntax Errors using Large Language ModelsarXiv 2023NoiseDiffusion: Correcting Noise for Image Interpolation with Diffusion Models beyond Spherical Linear InterpolationarXiv 2024How Do Humans Write Code? Large Models Do It the Same Way TooarXiv 2024AFRIDOC-MT: Document-level MT Corpus for African LanguagesarXiv 2025LLaSA: Large Language and E-Commerce Shopping AssistantarXiv 2024AutoTransfer: AutoML with Knowledge Transfer -- An Application to Graph Neural NetworksarXiv 2023On the Vulnerability of Skip Connections to Model Inversion AttacksarXiv 2024FIRST: Teach A Reliable Large Language Model Through Efficient Trustworthy DistillationarXiv 2024A Gromov--Wasserstein Geometric View of Spectrum-Preserving Graph CoarseningarXiv 2023MedDet: Generative Adversarial Distillation for Efficient Cervical Disc Herniation DetectionarXiv 2024Kvasir-VQA: A Text-Image Pair GI Tract DatasetarXiv 2024Single-shot Quantum Signal Processing InterferometryarXiv 2023Training-free Zero-shot Composed Image Retrieval via Weighted Modality Fusion and SimilarityarXiv 2024Variance-Aware Regret Bounds for Stochastic Contextual Dueling BanditsarXiv 2023PAC Prediction Sets for Large Language Models of CodearXiv 2023An Extended Study of Human-like Behavior under Adversarial TrainingarXiv 2023A comparative analysis between Conformer-Transducer, Whisper, and wav2vec2 for improving the child speech recognitionarXiv 2023Self-Correction is More than Refinement: A Learning Framework for Visual and Language Reasoning TasksarXiv 2024ReIFE: Re-evaluating Instruction-Following EvaluationarXiv 2024Automatic Generation of Model and Data Cards: A Step Towards Responsible AIarXiv 2024To Revise or Not to Revise: Learning to Detect Improvable Claims for Argumentative Writing SupportarXiv 2023GistScore: Learning Better Representations for In-Context Example Selection with Gist BottlenecksarXiv 2023A Survey on the Role of Crowds in Combating Online Misinformation:
Annotators, Evaluators, and CreatorsarXiv 2023Language Models' Factuality Depends on the Language of InquiryarXiv 2025Beyond Benchmarks: Evaluating Embedding Model Similarity for Retrieval Augmented Generation SystemsarXiv 2024Parameter-Efficient Fine-Tuning of Large Language Models for Unit Test Generation: An Empirical StudyarXiv 2024LGMCTS: Language-Guided Monte-Carlo Tree Search for Executable Semantic
Object RearrangementarXiv 2023Maximum Optimality Margin: A Unified Approach for Contextual Linear Programming and Inverse Linear ProgrammingarXiv 2023Exploring Predictive Uncertainty and Calibration in NLP: A Study on the Impact of Method & Data ScarcityarXiv 2022Differential Privacy has Bounded Impact on Fairness in ClassificationarXiv 2022The Promises and Pitfalls of LLM Annotations in Dataset Labeling: a Case Study on Media Bias DetectionarXiv 2024Discrete Key-Value BottleneckarXiv 2022VaxxHesitancy: A Dataset for Studying Hesitancy towards COVID-19 Vaccination on TwitterarXiv 2023Cross-lingual Transfer of Reward Models in Multilingual AlignmentarXiv 2024Deriving Language Models from Masked Language ModelsarXiv 2023Reshaping Free-Text Radiology Notes Into Structured Reports With Generative TransformersarXiv 2024ANTN: Bridging Autoregressive Neural Networks and Tensor Networks for Quantum Many-Body Simulationautoregressive-neural-tensornet-bridgingAn adaptively inexact first-order method for bilevel optimization with
application to hyperparameter learningarXiv 2023Do GPTs Produce Less Literal Translations?arXiv 2023Wait, that's not an option: LLMs Robustness with Incorrect Multiple-Choice OptionsarXiv 2024Using Motion Forecasting for Behavior-Based Virtual Reality (VR) AuthenticationarXiv 2024IDGI: A Framework to Eliminate Explanation Noise from Integrated GradientsCVPR 2023 1Designing Network Algorithms via Large Language ModelsarXiv 2024PAXQA: Generating Cross-lingual Question Answering Examples at Training ScalearXiv 2023Expedited Training of Visual Conditioned Language Generation via Redundancy ReductionarXiv 2023Novel Policy Seeking with Constrained Optimizationnovel-policy-seeking-with-constrained-1APT-Pipe: A Prompt-Tuning Tool for Social Data Annotation using ChatGPTarXiv 2024Generalist embedding models are better at short-context clinical semantic search than specialized embedding modelsarXiv 2024"My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language ModelsarXiv 2024Helpful assistant or fruitful facilitator? Investigating how personas affect language model behaviorarXiv 2024Understanding Certified Training with Interval Bound PropagationarXiv 2023AMELI: Enhancing Multimodal Entity Linking with Fine-Grained AttributesarXiv 2023Stochastic Parrots Looking for Stochastic Parrots: LLMs are Easy to Fine-Tune and Hard to Detect with other LLMsarXiv 2023The World of an Octopus: How Reporting Bias Influences a Language Model's Perception of ColorarXiv 2021Multi-Task Text Classification using Graph Convolutional Networks for Large-Scale Low Resource LanguagearXiv 2022Automatic Metrics in Natural Language Generation: A Survey of Current Evaluation PracticesarXiv 2024AirBirds: A Large-scale Challenging Dataset for Bird Strike Prevention in Real-world AirportsarXiv 2023Regret-Minimizing Double Oracle for Extensive-Form GamesarXiv 2023Evaluating Step-by-Step Reasoning through Symbolic VerificationarXiv 2022Two Complementary Perspectives to Continual Learning: Ask Not Only What to Optimize, But Also HowarXiv 2023DRiVE: Diffusion-based Rigging Empowers Generation of Versatile and Expressive CharactersCVPR 2025 1CoRe-Sleep: A Multimodal Fusion Framework for Time Series Robust to
Imperfect ModalitiesarXiv 2023When to Retrieve: Teaching LLMs to Utilize Information Retrieval EffectivelyarXiv 2024Rethinking Table Instruction TuningarXiv 2025Llama meets EU: Investigating the European Political Spectrum through the Lens of LLMsarXiv 2024Safety Alignment in NLP Tasks: Weakly Aligned Summarization as an In-Context AttackarXiv 2023MetaCoCo: A New Few-Shot Classification Benchmark with Spurious CorrelationarXiv 2024HERDPhobia: A Dataset for Hate Speech against Fulani in NigeriaarXiv 2022Is Fine-tuning Needed? Pre-trained Language Models Are Near Perfect for Out-of-Domain DetectionarXiv 2023What's in a Name? Auditing Large Language Models for Race and Gender BiasarXiv 2024Weakly-supervised Automated Audio Captioning via text only trainingarXiv 2023Understanding Data Temporality Impact on Large Language Models Pre-trainingarXiv 2026Personalised Distillation: Empowering Open-Sourced LLMs with Adaptive Learning for Code GenerationarXiv 2023Attacking Open-domain Question Answering by Injecting Misinformationcontraqa-question-answering-underMultilingual and Explainable Text Detoxification with Parallel CorporaarXiv 2024MAGID: An Automated Pipeline for Generating Synthetic Multi-modal DatasetsarXiv 2024NusaBERT: Teaching IndoBERT to be Multilingual and MulticulturalarXiv 2024Pre-training with Synthetic Data Helps Offline Reinforcement LearningarXiv 2023PyRadar: Towards Automatically Retrieving and Validating Source Code
Repository Information for PyPI PackagesarXiv 2024Embedding structure matters: Comparing methods to adapt multilingual vocabularies to new languagesarXiv 2023Check_square at CheckThat! 2020: Claim Detection in Social Media via Fusion of Transformer and Syntactic FeaturesarXiv 2020ZipLM: Inference-Aware Structured Pruning of Language Modelsziplm-inference-aware-structured-pruning-ofAnalysing the Noise Model Error for Realistic Noisy Label DataarXiv 2021Know thy corpus! Robust methods for digital curation of Web corporaknow-thy-corpus-robust-methods-for-digital-1Enhancing Language Multi-Agent Learning with Multi-Agent Credit Re-Assignment for Interactive Environment GeneralizationarXiv 2025PanGu-Coder: Program Synthesis with Function-Level Language ModelingarXiv 2022Do Differences in Values Influence Disagreements in Online Discussions?arXiv 2023From Relational Pooling to Subgraph GNNs: A Universal Framework for More Expressive Graph Neural NetworksarXiv 2023MathChat: Converse to Tackle Challenging Math Problems with LLM AgentsarXiv 2023CRAFT: Extracting and Tuning Cultural Instructions from the WildarXiv 2024Crossing the Linguistic Causeway: A Binational Approach for Translating
Soundscape Attributes to Bahasa MelayuarXiv 2022Unveiling Factual Recall Behaviors of Large Language Models through Knowledge NeuronsarXiv 2024How Well Do LLMs Represent Values Across Cultures? Empirical Analysis of LLM Responses Based on Hofstede Cultural DimensionsarXiv 2024Deep Integrated ExplanationsarXiv 2023Generating Data to Mitigate Spurious Correlations in Natural Language
Inference DatasetsarXiv 2022On The Differences Between Song and Speech Emotion Recognition: Effect of Feature Sets, Feature Types, and ClassifiersarXiv 2020How many perturbations break this model? Evaluating robustness beyond adversarial accuracyarXiv 2022SoFA: Shielded On-the-fly Alignment via Priority Rule FollowingarXiv 2024Towards Training Without Depth Limits: Batch Normalization Without Gradient ExplosionarXiv 2023An Evaluation Framework for Legal Document SummarizationLREC 2022 6Counting Carbon: A Survey of Factors Influencing the Emissions of Machine LearningarXiv 2023Exploring the Impact of Model Scaling on Parameter-Efficient TuningarXiv 2023NaviNeRF: NeRF-based 3D Representation Disentanglement by Latent Semantic NavigationICCV 2023 1Bounding the Expected Robustness of Graph Neural Networks Subject to Node Feature AttacksarXiv 2024ZeroUnlearn: Few-Shot Knowledge Unlearning in Large Language ModelsarXiv 2026Integrating Prior Knowledge in Contrastive Learning with KernelarXiv 2022Investigating the Efficacy of Large Language Models for Code Clone DetectionarXiv 2024In-Context Learning Learns Label Relationships but Is Not Conventional LearningarXiv 2023Spectral Co-Distillation for Personalized Federated Learningspectral-co-distillation-for-personalizedGraph Neural Networks for Knowledge Enhanced Visual Representation of PaintingsarXiv 2021CoDeF: Content Deformation Fields for Temporally Consistent Video ProcessingCVPR 2024 1Recurrent Action Transformer with MemoryarXiv 2023Large Language Models as Zero-Shot Human Models for Human-Robot InteractionarXiv 2023Phase-aware Single-stage Speech Denoising and Dereverberation with U-Netphase-aware-single-stage-speech-denoising-andCsFEVER and CTKFacts: Acquiring Czech data for fact verificationarXiv 2022What Makes Instruction Learning Hard? An Investigation and a New Challenge in a Synthetic EnvironmentarXiv 2022Linear Optimal Partial Transport EmbeddingarXiv 2023Can LLMs Solve longer Math Word Problems Better?arXiv 2024Dimensionless Anomaly Detection on Multivariate Streams with Variance Norm and Path SignaturearXiv 2020Semantic Contextualization of Face Forgery: A New Definition, Dataset, and Detection MethodarXiv 2024Persona-judge: Personalized Alignment of Large Language Models via Token-level Self-judgmentarXiv 2025Prompting-based Synthetic Data Generation for Few-Shot Question AnsweringarXiv 2024The Validity of Evaluation Results: Assessing Concurrence Across Compositionality BenchmarksarXiv 2023Towards Consistent Natural-Language Explanations via Explanation-Consistency FinetuningarXiv 2024HoloDetect: Few-Shot Learning for Error DetectionarXiv 2019NLEBench+NorGLM: A Comprehensive Empirical Analysis and Benchmark Dataset for Generative Language Models in NorwegianarXiv 2023A Quantum Algorithm for Solving Linear Differential Equations: Theory
and ExperimentarXiv 2018CSGNet: Neural Shape Parser for Constructive Solid GeometryarXiv 2017Online Recognition of Incomplete Gesture Data to Interface Collaborative RobotsarXiv 2023What Do Llamas Really Think? Revealing Preference Biases in Language Model RepresentationsarXiv 2023A Bayes Factor for Replications of ANOVA ResultsarXiv 2016A large annotated corpus for learning natural language inferencea-large-annotated-corpus-for-learning-natural-1Quantum algorithm for solving linear systems of equationsarXiv 2008Why Random Pruning Is All We Need to Start SparsearXiv 2022Is That Your Final Answer? Test-Time Scaling Improves Selective Question AnsweringarXiv 2025Emergent representations in networks trained with the Forward-Forward algorithmarXiv 2023Socially Aware Bias Measurements for Hindi Language RepresentationsNAACL 2022 7DEsignBench: Exploring and Benchmarking DALL-E 3 for Imagining Visual DesignarXiv 2023SaudiBERT: A Large Language Model Pretrained on Saudi Dialect CorporaarXiv 2024POV: Prompt-Oriented View-Agnostic Learning for Egocentric Hand-Object Interaction in the Multi-View WorldarXiv 2024Large Language Models Meet Open-World Intent Discovery and Recognition: An Evaluation of ChatGPTarXiv 2023CoLiDE: Concomitant Linear DAG EstimationarXiv 2023Are Large Language Models Actually Good at Text Style Transfer?arXiv 2024Implicit regularization of deep residual networks towards neural ODEsarXiv 2023Benchmark Data and Evaluation Framework for Intent Discovery Around COVID-19 Vaccine HesitancyarXiv 2022Learning invariant representations of time-homogeneous stochastic dynamical systemsarXiv 2023Beyond the Surface: Measuring Self-Preference in LLM JudgmentsarXiv 2025AbsInstruct: Eliciting Abstraction Ability from LLMs through Explanation Tuning with Plausibility EstimationarXiv 2024You Are What You Annotate: Towards Better Models through Annotator RepresentationsarXiv 2023Is Temperature the Creativity Parameter of Large Language Models?arXiv 2024Late Stopping: Avoiding Confidently Learning from Mislabeled ExamplesICCV 2023 1TOPFORMER: Topology-Aware Authorship Attribution of Deepfake Texts with Diverse Writing StylesarXiv 2023Beyond Factual Accuracy: Evaluating Coverage of Diverse Factual Information in Long-form Text GenerationarXiv 2025Understanding the Impact of Adversarial Robustness on Accuracy DisparityarXiv 2022Principled Federated Domain Adaptation: Gradient Projection and Auto-WeightingarXiv 2023Humans or LLMs as the Judge? A Study on Judgement BiasesarXiv 2024Combating Adversarial Attacks with Multi-Agent DebatearXiv 2024Language Models are Crossword SolversarXiv 2024yosm: A new yoruba sentiment corpus for movie reviewsarXiv 2022A Comparative Analysis of Conversational Large Language Models in Knowledge-Based Text GenerationarXiv 2024Can Muon Fine-tune Adam-Pretrained Models?arXiv 2026A Formal Perspective on Byte-Pair EncodingarXiv 2023Soft Prompt Tuning for Cross-Lingual Transfer: When Less is MorearXiv 2024A Methodology for Evaluating RAG Systems: A Case Study On Configuration Dependency ValidationarXiv 2024Navigating Cultural Chasms: Exploring and Unlocking the Cultural POV of Text-To-Image ModelsarXiv 2023What Do Compressed Multilingual Machine Translation Models Forget?arXiv 2022When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model LeaderboardsarXiv 2024Lie Group Decompositions for Equivariant Neural NetworksarXiv 2023Federated Learning with Partial Model Personalizationfederated-learning-with-partial-modelAdaptive Multi-head Contrastive LearningarXiv 2023Topology-Preserving Neural Operator Learning via Hodge DecompositionarXiv 2026Robust Counterfactual Explanations for Neural Networks With Probabilistic GuaranteesarXiv 2023Order-Preserving GFlowNetsarXiv 2023Interpretable Proof Generation via Iterative Backward ReasoningNAACL 2022 7Invisible Relevance Bias: Text-Image Retrieval Models Prefer AI-Generated ImagesarXiv 2023When does Privileged Information Explain Away Label Noise?arXiv 2023When Background Matters: Breaking Medical Vision Language Models by Transferable AttackarXiv 2026Indian-BhED: A Dataset for Measuring India-Centric Biases in Large Language ModelsarXiv 2023Deep Learning for Identifying Iran's Cultural Heritage Buildings in Need of Conservation Using Image Classification and Grad-CAMarXiv 2023Does Unlearning Truly Unlearn? A Black Box Evaluation of LLM Unlearning MethodsarXiv 2024Towards Unifying Evaluation of Counterfactual Explanations: Leveraging Large Language Models for Human-Centric AssessmentsarXiv 2024An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4arXiv 2024Benchmarking Linguistic Diversity of Large Language ModelsarXiv 2024Knowing Before Saying: LLM Representations Encode Information About Chain-of-Thought Success Before CompletionarXiv 2025The Elusive Pursuit of Reproducing PATE-GAN: Benchmarking, Auditing, DebuggingarXiv 2024Causal Strategic Classification: A Tale of Two ShiftsarXiv 2023MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation ModelsarXiv 2024Revision Transformers: Instructing Language Models to Change their ValuesarXiv 2022Improving Adversarial Data Collection by Supporting Annotators: Lessons from GAHD, a German Hate Speech DatasetarXiv 2024HaRiM$^+$: Evaluating Summary Quality with Hallucination RiskarXiv 2022Ultra-compact Binary Neural Networks for Human Activity Recognition on RISC-V ProcessorsarXiv 2022Self-supervised visual learning from interactions with objectsarXiv 2024Identifying the Correlation Between Language Distance and Cross-Lingual Transfer in a Multilingual Representation SpacearXiv 2023