← Back to all projects
Conference paper

Gutenberg+: A More Temporally Faithful Corpus for Diachronic NLP

Leon Hammerla, Alexander Mehler

Proceedings Workshop on Structured Linguistic Data and Evaluation (SLiDE 2026), co-located with the Language Resources and Evaluation Conference (LREC 2026) · 2026

View publication Code & archive

Abstract

We introduce Gutenberg+, a temporally more faithful version of the Project Gutenberg (PG) corpus, one of the most widely used resources for diachronic text analysis. Despite its popularity, the PG corpus contains a major yet overlooked flaw: around 15% of its entries are collections (e.g., anthologies of books, letters, or poems) rather than atomic works, which distorts temporal analyses since such collections may span multiple decades. We present an automatic method to detect and split these collections into their constituent works, producing a finer-grained and temporally consistent corpus. We further re-annotate publication years using LLM-based retrieval-augmented generative methods, demonstrating the potential of LLMs to enhance structured linguistic resources. To illustrate the utility of Gutenberg+, we conduct a small-scale diachronic case study on negation, showing that our refined corpus captures more nuanced cross-linguistic variation than the original PG data. Finally, we release the corpus in UIMA format with full metadata and linguistic annotations, providing a standardized resource for future research on diachronic language change.

Keywords

neglab

BibTeX

@inproceedings{Hammerla:Mehler:2026:a,
  title     = {{Gutenberg+}: A More Temporally Faithful Corpus for Diachronic {NLP}},
  author    = {Leon Hammerla and Alexander Mehler},
  booktitle = {Proceedings Workshop on Structured Linguistic Data and Evaluation
               (SLiDE 2026), co-located with the Language Resources and Evaluation
               Conference (LREC 2026)},
  year      = {2026},
  keywords  = {neglab},
  pages     = {86--92},
  address   = {Palma, Mallorca, Spain},
  publisher = {European Language Resources Association (ELRA)},
  editor    = {Erhard Hinrichs (Tübingen University, Germany) and Joakim Nivre (Uppsala University, Sweden)
               and Petya Osenova (Sofia University, Bulgaria) and James Pustejovsky (Brandeis University, USA)
               and Claus Zinn (Tübingen University, Germany)},
  doi       = {10.63317/2kjofgrkkbt9},
  abstract  = {We introduce Gutenberg+, a temporally more faithful version of
               the Project Gutenberg (PG) corpus, one of the most widely used
               resources for diachronic text analysis. Despite its popularity,
               the PG corpus contains a major yet overlooked flaw: around 15%
               of its entries are collections (e.g., anthologies of books, letters,
               or poems) rather than atomic works, which distorts temporal analyses
               since such collections may span multiple decades. We present an
               automatic method to detect and split these collections into their
               constituent works, producing a finer-grained and temporally consistent
               corpus. We further re-annotate publication years using LLM-based
               retrieval-augmented generative methods, demonstrating the potential
               of LLMs to enhance structured linguistic resources. To illustrate
               the utility of Gutenberg+, we conduct a small-scale diachronic
               case study on negation, showing that our refined corpus captures
               more nuanced cross-linguistic variation than the original PG data.
               Finally, we release the corpus in UIMA format with full metadata
               and linguistic annotations, providing a standardized resource
               for future research on diachronic language change.}
}