Too Many; Didn't Read? Classifying Large Datasets with LLMs
About this Event
Convenor
Raphael is a PhD candidate at Cambridge Digital Humanities, researching how AI impacts journalism, information environments, and mediation. His academic work comes after more than a decade as a journalist at places like Folha de S.Paulo and the Guardian.
Description
This workshop covers how to use large language models (LLMs) to classify datasets and extract information at scale, and on knowing when the results can be trusted. Researchers often face corpora too large to read and classify manually. Data from social media and other online platforms pose a greater challenge due to chaotic communication, slang, and shifting meanings across communities. Fixed dictionaries, keyword counts, and older machine learning models can be tricky to implement and often fail at complex labelling in this kind of data. LLMs and multimodal models can code (in the social sciences sense) and extract information at scale. Outputs, however, are probabilistic by design; without testing, there is no way to know whether a classification is reliable. Thus, validation is an essential step. The workshop moves in four parts. First, context: where LLMs are appropriate against manual coding or supervised machine learning. Second, the hands-on core: designing a codebook, writing structured prompts, classifying a real dataset. Third, validation: building gold-standard subsets, measuring agreement, and analysing errors. Fourth, extensions: cross-checking across models and extracting information from images. No programming experience is necessary for this workshop, though Python notebooks will be presented as an option for those who prefer them.
Is any equipment or Software required?
Participants need a laptop, a web browser, and a Google Account. The workshop uses Google's Gemini API, which has a free tier that requires no credit card or payment details. Participants create their own API key through Google AI Studio using a personal Google account. Setup instructions will be circulated in advance. Both tracks run in the browser, with nothing to install: - No-code: a Google Sheets template I provide, which calls the model from a spreadsheet formula. - Code: a Google Colab notebook I provide. Exercises are sized to stay within free-tier rate limits. I will also supply pre-computed model outputs so that the validation exercises — the core of the session — can be completed with no API access at all, should keys or connectivity fail. The methods are not tied to any one provider. The same workflow applies to other models, and the materials note alternatives, including locally-run ones.
Target Audience
Our CDH Methods workshops have limited places and are prioritised for students and staff at the University of Cambridge. However, if space is available, we welcome all participants who want to learn and apply digital methods and use digital tools in their research.
This session may be of particular interest to:
- PhD students in the Arts, Humanities and Social Sciences
- Early Career Researchers in the Arts, Humanities and Social Sciences
Contact CDH
If you have specific accessibility needs for this event, please get in touch. We will do our best to accommodate any requests, however please note that the building is grade 2 listed and has 4 steps into the building so unfortunately wheelchair access is not available.
This workshop is part of our Methods Fellowship programme, which develops and delivers innovative teaching in digital methods. You can read more about the and view the complete series of .
Where is it happening?
Event Location & Nearby Stays:
GBP 0.00



















