NLP Datasets
NLP Datasets is a focused GitHub list that catalogues datasets made of text, the raw material for most retrieval and language work. Entries span general corpora, conversational and dialogue data, sentiment-labelled sets, and summarisation collections, each with a short note and a link to the source.
The alphabetical layout makes it easy to skim, though it means you browse by name rather than by task. As with any community list, check each dataset's licence and last update at its home before you rely on it. It pairs well with broader directories when you specifically need language data rather than tabular or geospatial sets.
Related sources
Awesome Legal Data
A community-maintained list of legal datasets, tools, and resources for legal text processing across jurisdictions, including court records, statutes, contracts, and legal NLP benchmarks. A useful map for anyone building a legal RAG system.
Awesome Public Datasets
A community-curated list of high-quality open datasets on GitHub, organised by topic: agriculture, biology, climate, economics, education, finance, government, healthcare, and more. A good starting point when you need RAG-ready data for a specific domain and do not yet know where to look.
DataKind UK Open Data Sets
A curated list of UK-focused open datasets from DataKind UK, covering government, health, crime, housing, and social data. A quick way into British public data when you are building a RAG system with a UK focus.
UK Data Service
The UK's largest collection of economic, population, and social research data for teaching, learning, and public benefit. Many datasets need registration or an institutional login, so plan for an access step before you build with them.