Research strategies for manual data collection#

Finding sufficient data for a topic you are interested in can be challenging, also because relevant data may not include the most obvious keywords. Here is some advice on identifying useful data sets when you are aiming for a document-by-document manual data collection.

Finding relevant search terms#

Regardless of the search engine, platform, or database you use to find content related to a topic, it is important to carefully develop your search terms. A mistake that students often make is that their search terms are either too broad or much too narrow. For example, searching for “history” will give you more results than you can process, and it will not help you to create or answer a clear research question. Instead of searching for “history,” you may want to search for “medieval history in France” to get a more manageable list of results. One example of the opposite challenge is that a student of mine in a past course tried to find podcasts on “genital mutilation,” but the few results that came up had absolutely no reviews they could analyse. In such cases, it is recommended to branch out and consider related aspects such as female bodies, female sexuality, or violence against women.

Use platform-integrated search features#

When searching for content to analyse — whether videos, podcasts, articles, or posts — you can, of course, begin your search with the search features built into that platform. Most people find this straightforward on platforms like YouTube, but not every search interface works the same way, and it’s worth understanding the quirks of whichever one you use. For example, if you’re working with the Apple Podcasts app, note that reviews may only be visible when you open the show or episode website in a browser rather than the app itself — a particular challenge for Mac users, whose machines tend to automatically redirect to the app. Also consider that many platform search engines only index titles and short descriptions, not full content or transcripts: if your keyword isn’t mentioned there, a relevant result may simply never surface. This makes it worth combining platform search with external research tools, aggregators, and databases (see below) rather than relying on a single search box.

Aggregators, directories and recommendation sites#

Depending on your data source, dedicated aggregator and directory sites can help you find popular or niche content more efficiently than a single platform’s search bar. For podcasts, websites like Podchaser, FeedSpot, or Listen Notes list top-rated and specialised shows, often filterable by subject, reviews, and popularity. Blogs and tech news sites also regularly feature curated recommendations, such as the 60 Best Podcasts (2024) list from Wired. Similar approaches apply to other media: for YouTube, tech sites offer guidance such as “9 Ways To Find Trending Topics On YouTube In 2024”. Whatever your topic and medium, it’s worth searching for “best/most popular [medium] about [topic]” as its own research step, and then checking whether the results are available on the platform you intend to analyse.

Explore social media and discussion forums#

Platforms like X, Mastodon, Threads, BlueSky, Reddit, and Facebook groups often have communities that discuss and recommend content relevant to your topic — whether podcasts, videos, books, or news outlets. Becoming part of a social media community engaging with your research topic can, therefore, be very helpful, even when you do not intend to analyse social media discussions. You can also consider asking other social media users directly for recommendations on a specific topic, especially if they are experts.

Academic and subscription databases#

Many data sets suited to social sciences and humanities research are available via academic or commercial databases. Academic databases are usually created by researchers and financed through public funding, so they can be accessed free of charge. Commercial databases, like many newspaper archives, require are subscription. Before you consider paying for paywalled data yourself, always make sure to check if the university offers a collective subscription for staff and students!

Here are some databases that may be of special interest to MA Digital Cultures students:

  • HathiTrust: a large digital library of (older) digitised books, journals, and government documents, useful for historical text analysis, corpus linguistics, and long-run studies of published discourse.

  • Nexis Uni (Lexis Nexis): a major database of news articles, legal documents, and company records, widely used for media analysis and tracking how topics have been covered over time.

To check for which databases Maastricht University holds subscriptions, consult the following overview page:

Maastricht University Databases collections

You can simply browse the entire list or use the search bar on the left to identify databases for particular subject areas like medicine and law.