0

Writing Style Similarity Reflects Academic Genealogy

As authorship attribution systems are increasingly deployed to detect ghostwritten and AI-generated papers, their errors can support accusations against legitimate authors. These systems assume each author's style is their own.

Preview
Year
2026
Hosting
Full text hostedCC-BY-4.0

Cite

Notes

Only stored in your browser.

Attribution

Abstract & full text
arxiv.org/abs/2608.14843CC-BY-4.0
TL;DR
Semantic Scholar
Attribution policy →

Abstract

As authorship attribution systems are increasingly deployed to detect ghostwritten and AI-generated papers, their errors can support accusations against legitimate authors. These systems assume each author's style is their own. Researchers, however, study under advisors, and inherit their stylistic quirks. We build a corpus of arXiv authors with \geq 2 solo papers from the Mathematics Genealogy Project graph, giving 5{,}803 total authors and 2{,}501 ground-truth advisor-student pairings. Using embeddings from a fine-tuned model, advisors sit 39.9% closer in cosine distance to their students than a random same-field author does. Two open encoders reproduce the effect at 12.6% and 14.5%. Academic siblings, two students of one advisor who may never have met, sit 30.4% closer across 8{,}360 pairs, even when they studied at different institutions. Pairs who share only an institution and a field show negligible similarity. Given a closed-set attribution task over the same corpus, the system's errors occur on the true author's advisors and academic siblings 11 times more often than chance.