0

Understanding or Memorizing? A Case Study of German Definite Articles in Language Models

Gradient-based interpretability reveals that language models rely on memorized associations rather than abstract grammatical rules for German definite articles.

Year
2026
Venue
arXiv 2026
Authors
3
Hosting
Abstract onlyARXIV-DEFAULT

Cite

Notes

Only stored in your browser.

Attribution

Abstract & full text
arxiv.org/abs/2601.09313ARXIV-DEFAULT
TL;DR
Semantic Scholar
Attribution policy →

Abstract

Language models perform well on grammatical agreement, but it is unclear whether this reflects rule-based generalization or memorization. We study this question for German definite singular articles, whose forms depend on gender and case. Using GRADIEND, a gradient-based interpretability method, we learn parameter update directions for gender-case specific article transitions. We find that updates learned for a specific gender-case article transition frequently affect unrelated gender-case settings, with substantial overlap among the most affected neurons across settings. These results argue against a strictly rule-based encoding of German definite articles, indicating that models at least partly rely on memorized associations rather than abstract grammatical rules.

Authors

3