Awesome Public Datasets
Awesome Public Datasets is one of the longest-running "awesome" lists on GitHub, gathering links to open datasets under clear topic headings. Rather than hosting data itself, it points you to the primary source for each dataset, which fits the RAG Repo principle of always linking to the official home.
Because it is maintained by the community through pull requests, coverage is broad but uneven, and some links age faster than others. Treat it as a map for discovery: scan the topic that matches your use case, follow the link, then check the licence and freshness at the source before you build on it.
Related sources
Awesome Legal Data
A community-maintained list of legal datasets, tools, and resources for legal text processing across jurisdictions, including court records, statutes, contracts, and legal NLP benchmarks. A useful map for anyone building a legal RAG system.
DataKind UK Open Data Sets
A curated list of UK-focused open datasets from DataKind UK, covering government, health, crime, housing, and social data. A quick way into British public data when you are building a RAG system with a UK focus.
NLP Datasets
An alphabetical list of free and public domain text datasets for natural language processing, covering corpora, dialogue, sentiment, and summarisation. Handy when you want text-heavy data to build or evaluate a RAG system.
UK Data Service
The UK's largest collection of economic, population, and social research data for teaching, learning, and public benefit. Many datasets need registration or an institutional login, so plan for an access step before you build with them.