Stylometric similarity in literary corpora: Non-authorship clustering andDeutscher Novellenschatz

Päpcke S, Weitin T, Herget K, Glawion A, Brandes U (2022)


Publication Type: Journal article

Publication year: 2022

Journal

Publisher: Oxford University Press

City/Town: Oxford

Book Volume: 38

Pages Range: 277-295

Issue: 1

DOI: 10.1093/llc/fqac039

Abstract

A distant-reading task in literary corpus analysis is to group stylometrically similar texts. Since there are many ways to define writing style, the result not only depends on the clustering method but even more so on the measure of similarity. With authorship attribution, the predominant application of stylometry, as its benchmark much research has addressed the utility of methods for measuring similarity. We use a corpus of German-language novellas to demonstrate that one may be interested in very different meaningful groups of texts simultaneously, and that these can be recovered from stylometric clustering if the measure is chosen accordingly. As can be expected, different measures do better at recovering groups associated with, for instance, subgenre, author gender, or narrative perspective. As a consequence, it is suggested that corpus analyses should not be based on what is currently considered the most refined measure of stylometric similarity, but rather break down the decisions that yield a specific measure and provide substantively justified arguments for them.

Authors with CRIS profile

Involved external institutions

How to cite

APA:

Päpcke, S., Weitin, T., Herget, K., Glawion, A., & Brandes, U. (2022). Stylometric similarity in literary corpora: Non-authorship clustering and<i>Deutscher Novellenschatz</i>. Digital Scholarship in the Humanities, 38, 277-295. https://doi.org/10.1093/llc/fqac039

MLA:

Päpcke, Simon, et al. "Stylometric similarity in literary corpora: Non-authorship clustering and<i>Deutscher Novellenschatz</i>." Digital Scholarship in the Humanities 38 (2022): 277-295.

BibTeX: Download