Using scholarly digital editions as research sources

Using scholarly digital editions as research sources#

Scholarly digital editions are another valuable — and often underused — source of full-text material for media studies research. These are carefully curated, peer-reviewed collections of historical or political texts, usually maintained by universities, libraries, or archives. Scholarly editions often come with rich metadata (dates, authorship, subject tags) and annotations (contextualising comments) that makes them well suited to systematic corpus-building.

What makes digital editions useful — and their limits#

Unlike scraped social media data, digital editions are clean and ready-to-use. They often come in different file formats for different research purposes, including plain text and XML/TEI. Scholarly editions are also typically free of copyright restrictions as many focus on historical documents now in the public domain. However, most digital editions offer only search and filter functions, not bulk downloads, so you cannot usually export the whole collection or selected items as a single dataset. In some cases, the full collection may be available via GitHub, but most collections will require you to download records individually. Of course, this process can typically be sped up with a web-automation script once you know what the collection contains and how it is structured. But you should only do this if you are sure that the researchers you created the collection have shared it through a creative commons license that allows such data reuse.

A note on access#

As with subscription databases, some digital editions overlap with resources UM already licenses. Check the University Library’s database overview if you’re looking for a specific edition, since it may already be accessible through your UM account even if it isn’t fully open-access.