Browse all sources
Every source in the directory, in one place. Search by name, description, or tag, and filter by category, access type, and how ready the data is for retrieval.
More filters: RAG readiness
78 sources
- Academic & scientific literatureOpen
arXiv
An open-access preprint server for physics, mathematics, computer science, quantitative biology, statistics, and more, holding over 2.5 million papers. You can pull the full archive in bulk from Amazon S3, or harvest the metadata through OAI-PMH, a standard protocol for sharing records between repositories.
preprintsphysicsmathematics - Curated lists & meta-resourcesOpen
Awesome Legal Data
A community-maintained list of legal datasets, tools, and resources for legal text processing across jurisdictions, including court records, statutes, contracts, and legal NLP benchmarks. A useful map for anyone building a legal RAG system.
awesome-listlegaldirectory - Curated lists & meta-resourcesOpen
Awesome Public Datasets
A community-curated list of high-quality open datasets on GitHub, organised by topic: agriculture, biology, climate, economics, education, finance, government, healthcare, and more. A good starting point when you need RAG-ready data for a specific domain and do not yet know where to look.
awesome-listdirectoryopen-data - Data platforms & marketplacesCommercial
AWS Data Exchange
A marketplace for finding, subscribing to, and using third-party data inside the AWS cloud. It carries both free open datasets and paid commercial data products, so you can pull licensed data straight into your AWS workflows without setting up separate transfers.
data-platformmarketplacecommercial - Data platforms & marketplacesOpen
AWS Open Data Registry
A registry of high-value datasets hosted on AWS and made publicly available, covering genomics, geospatial data, climate, satellite imagery, and more. Over 300 PB of data in total, free to access: you pay only for the compute you use to process it.
data-platformcloudgeospatial - Web corporaOpen
C4 (Colossal Clean Crawled Corpus)
A 750 GB English corpus derived from Common Crawl using heuristic cleaning to keep natural language and drop gibberish, boilerplate, and placeholder text. Built to train Google's T5 and later used for MPT-7B and others.
web-crawlenglishpretraining - Legal & regulatoryLimited
Cambridge Law Corpus
A research dataset of more than 250,000 UK court cases, mostly from the 21st century but with some reaching back to the 16th century. Built by the University of Cambridge for legal natural language processing work, it is available under restricted access for research use.
ukcase-lawlegal - News, events & mediaOpen
CC-News
A subset of Common Crawl focused on news, containing millions of articles pulled from news websites around the world. It has fed several large language model training pipelines and gives you a ready news corpus without crawling sites yourself.
newsweb-crawlpretraining - Biomedical & healthOpen
ClinicalTrials.gov
A US National Institutes of Health database of privately and publicly funded clinical studies. It holds structured records on more than 450,000 studies, including protocols, conditions, interventions, outcomes, and results, all in the public domain.
clinical-trialsuspublic-domain - Web corporaOpen
Common Crawl
A nonprofit that crawls the web and freely provides its archives and datasets. Petabytes of raw web data from billions of pages, with a new snapshot each month. Used in the training of GPT-3, LLaMA, T5, and many other large language models.
web-crawlenglishmultilingual - Government & public sectorOpen
Companies House
The UK's official register of companies. Search every registered company, its directors, filing history, and accounts, or pull the data in bulk. A reliable source for UK corporate facts in a RAG system.
ukgovernmentcompanies - Knowledge graphs & structured dataOpen
ConceptNet
A multilingual common-sense knowledge graph that links words and phrases with labelled connections, such as "a cat is a pet" or "rain causes wet". It captures the everyday relationships between ideas that plain text rarely spells out.
knowledge-graphcommon-sensemultilingual - Academic & scientific literatureOpen
CORE
An aggregator of open-access research papers that harvests from thousands of repositories and journals worldwide. It holds over 300 million metadata records and more than 40 million full-text articles, all reachable through one search API.
open-accessaggregatorfull-text - Academic & scientific literatureOpen
Crossref
A nonprofit DOI registration agency that publishes metadata for over 150 million scholarly works, including journal articles, books, and conference proceedings. A DOI is the permanent identifier assigned to each work, and Crossref shares its metadata through a free API.
metadatadoicitations - Government & public sectorOpen
data.europa.eu
The European Data Portal, a single point of access to open data from 36 European countries and the EU institutions. Over 1.6 million datasets span every policy area, giving RAG systems broad, multilingual coverage of official European data.
eueuropegovernment - Government & public sectorOpen
data.gov
The US federal government's open data portal, gathering more than 300,000 datasets from federal agencies. Coverage spans agriculture, climate, education, energy, finance, health, and public safety, making it a broad source of official US data for RAG systems.
usgovernmentopen-data - Government & public sectorOpen
data.gov.uk
The UK Government's central catalogue of open public sector data, run by the Government Digital Service. More than 47,000 datasets cover health, housing, transport, demographics, and the environment, giving RAG systems a trusted base of UK facts, figures, and official records.
ukgovernmentopen-data - Data platforms & marketplacesOpen
Datahub.io
A platform for publishing and finding open data packages: datasets bundled with consistent metadata in standardised formats. It hosts curated collections of widely used reference data, from country codes to exchange rates, ready to drop into a pipeline.
data-platformopen-datadata-packages - Curated lists & meta-resourcesOpen
DataKind UK Open Data Sets
A curated list of UK-focused open datasets from DataKind UK, covering government, health, crime, housing, and social data. A quick way into British public data when you are building a RAG system with a UK focus.
directoryukopen-data - Encyclopaedic & general knowledgeOpen
DBpedia
A knowledge graph built by pulling the structured parts of Wikipedia, mainly the infoboxes, into machine-readable data. It holds billions of facts about people, places, organisations, and more as RDF triples, small subject, predicate, object statements, which you can search with the SPARQL query language.
knowledge-graphstructuredwikipedia - Web corporaOpen
DCLM-Baseline (DataComp-LM)
A filtered English web dataset from the DataComp-LM benchmark project, produced by running model-based quality filtering over Common Crawl. Built to show which data-curation choices most improve language model training.
web-crawlenglishpretraining - Code & technical documentationOpen
DevDocs
An open-source aggregator that pulls API documentation for major programming languages, frameworks, and tools into one searchable place. A tidy RAG source when you want up-to-date developer reference material without scraping dozens of separate documentation sites yourself.
api-documentationdeveloperaggregator - Web corporaOpen
Dolma
A 3T token open corpus from Allen AI (its name stands for "Data for Open Language Models' Appetite") combining web text, scientific papers, code, public-domain books, Reddit posts, and Wikipedia. Built to train the OLMo models.
web-crawlpretrainingmulti-source - Biomedical & healthLimited
DrugBank
A freely accessible database that combines detailed drug data with drug target information, covering more than 15,000 drug entries. It links chemistry, pharmacology, and biology in one place, which makes it a strong knowledge source for medical and pharmaceutical RAG.
drugspharmacologydrug-targets - Legal & regulatoryOpen
EDGAR (SEC Filings)
The US Securities and Exchange Commission's repository of corporate filings, including annual reports (10-K), quarterly reports (10-Q), current reports (8-K), and proxy statements. You can search the full text or pull filings in bulk through an API, which makes it a rich source for financial and regulatory RAG.
usfinancial-filingssec - Legal & regulatoryOpen
EUR-Lex
The official portal for European Union law. It provides EU treaties, legislation, case law, and legislative proposals in 24 languages, with bulk download available. A dependable source for multilingual legal and regulatory RAG across the EU.
eulegislationcase-law - Web corporaOpen
FineWeb / FineWeb-Edu
A 15T token English web corpus distilled from Common Crawl by HuggingFace, filtered and deduplicated for language model training. FineWeb-Edu is a subset filtered for educational content. One of the higher-quality open web datasets for pretraining and broad-coverage RAG.
web-crawlenglishpretraining - Web corporaOpen
FineWeb2
A multilingual extension of FineWeb from HuggingFace, built with an open curation pipeline that adapts filtering and deduplication across languages. Useful when you need clean web text beyond English for pretraining or multilingual RAG.
web-crawlmultilingualpretraining - RAG-specific & evaluationOpen
FlashRAG Benchmark Datasets
A research toolkit that bundles standardised, pre-processed benchmark datasets for RAG evaluation across many scenarios. Every dataset comes in a single unified JSONL format, so you can swap benchmarks in and out without rewriting your data loading each time.
evaluationbenchmarktoolkit - Legal & regulatoryOpen
Free Law Project / CourtListener
The largest open collection of US court opinions, oral argument recordings, and judicial financial disclosures, run by the nonprofit Free Law Project. Its RECAP Archive holds hundreds of millions of federal court docket entries, and journalists at the Wall Street Journal and ProPublica have used it for investigations. A solid base for US legal RAG.
legaluscase-law - Encyclopaedic & general knowledgeOpen
Freebase
A collaborative knowledge base once run by Google, now retired but still available as downloadable data dumps. It holds structured facts about millions of entities, and much of its content has since moved into Wikidata.
knowledge-basestructuredarchived - News, events & mediaOpen
GDELT
GDELT, the Global Database of Events, Language, and Tone, monitors print, broadcast, and web news in more than 100 languages from every country. It records events back to 1979, builds a Global Knowledge Graph, and adds sentiment and emotion analysis, all queryable on Google BigQuery.
newseventsmultilingual - Geospatial & mappingOpen
GeoNames
A geographical database of more than 11 million place names, each with coordinates, population, elevation, and administrative divisions, covering every country. It works well as a gazetteer for resolving place names to locations.
gazetteerplace-namescoordinates - Code & technical documentationOpen
GitHub Public Repositories (GH Archive / GHTorrent)
Two projects that capture public GitHub activity for analysis. GH Archive records the public GitHub event timeline as downloadable hourly archives, and GHTorrent offers a queryable mirror of GitHub metadata. Together they help you mine code, issues, pull requests, and documentation at scale.
githubevent-dataissues - Books & literatureOpen
Google Books Ngrams
Word and phrase frequency data drawn from Google's digitised book collection, spanning centuries of published text. It counts how often words and short phrases appear by year, which is useful for historical language analysis and time-aware RAG features.
ngramsword-frequencyhistorical - Data platforms & marketplacesOpen
Google Dataset Search
A search engine for datasets that indexes millions of them from thousands of repositories worldwide. It does not host data itself: it points you to wherever each dataset lives, which makes it a fast first stop when you are hunting for a source on a specific topic.
data-platformsearch-enginediscovery - Knowledge graphs & structured dataLimited
Google Knowledge Graph API
An API into Google's Knowledge Graph, the store of billions of facts about entities that powers the info panels you see beside search results. You send a name or query and get back matching entities with descriptions, types, and links. Free to use within rate limits.
apiknowledge-graphentities - Government & public sectorOpen
Hansard
The official word-for-word record of proceedings in the UK Parliament, both the Commons and the Lords. Full text searchable and available as XML and linked data going back centuries, it is a rich source of political and legislative debate for RAG systems.
ukparliamentgovernment - Data platforms & marketplacesOpen
Hugging Face Datasets
The largest open hub for machine learning datasets, with well over 100,000 datasets you can search, stream, and version. Many RAG-ready corpora live here, and its Python library lets you pull data straight into a pipeline in a few lines.
platformmachine-learningstreaming - News, events & mediaOpen
Internet Archive
A nonprofit digital library giving free access to millions of books, films, audio recordings, software, archived websites, and television news. Its Wayback Machine has saved more than 800 billion web pages, making it a deep well of historical and current text.
digital-librarynonprofitweb-archive - Data platforms & marketplacesOpen
Kaggle Datasets
A hub of user-contributed datasets spanning almost every domain, owned by Google. Alongside the data you get notebooks, competitions, and active discussion, which makes it a quick way to find something workable and see how others have already cleaned and used it.
data-platformcommunitynotebooks - Legal & regulatoryOpen
Legislation.gov.uk
The official UK government legislation database, holding all UK Acts of Parliament and statutory instruments. It is available as XML and linked data, so you can load the full statute book into a RAG system rather than scraping individual pages.
uklegislationstatutes - Books & literatureOpen
LibriVox
A volunteer project providing free public domain audiobooks. Volunteers record themselves reading works whose copyright has expired, which makes it a handy source of speech data for training speech-to-text models and building multimodal RAG.
public-domainaudiobooksaudio - Biomedical & healthLimited
MIMIC-III / MIMIC-IV
Freely accessible critical care databases with de-identified health data from more than 40,000 intensive care patients at Beth Israel Deaconess Medical Center. They include vital signs, lab results, medications, and clinical notes. Access needs credentialing and a data-use agreement.
clinicalcritical-careehr - Geospatial & mappingOpen
Natural Earth
A public domain map dataset built for cartography, offered at three scales: 1:10m, 1:50m, and 1:110m. It bundles cultural data (borders, cities, roads), physical data (coastlines, rivers, lakes), and raster imagery, all ready to drop into a map.
public-domainmapscartography - Knowledge graphs & structured dataOpen
NELL (Never-Ending Language Learner)
A machine learning system from Carnegie Mellon that has been reading the web since 2010 and building a knowledge base as it goes. It extracts beliefs, entities, and the relationships between them from text, and keeps refining them over time.
knowledge-graphmachine-learningweb-extraction - Curated lists & meta-resourcesOpen
NLP Datasets
An alphabetical list of free and public domain text datasets for natural language processing, covering corpora, dialogue, sentiment, and summarisation. Handy when you want text-heavy data to build or evaluate a RAG system.
awesome-listnlptext-corpora - Government & public sectorOpen
ONS (Office for National Statistics)
The UK's official statistics body. It publishes Census data, economic indicators, population estimates, labour market figures, and more, making it the authoritative source for UK statistics in a RAG system.
ukgovernmentstatistics - Books & literatureOpen
Open Library
An Internet Archive project building a web page for every book ever published. It holds more than 20 million catalogue records and lends many titles digitally, making it a rich source of book metadata for RAG systems.
booksmetadatacatalogue - RAG-specific & evaluationOpen
Open RAGBench (Vectara)
A RAG evaluation benchmark from Vectara built on 1,000 arXiv papers, with multimodal content extraction and 3,045 question-and-answer pairs across scientific domains. Designed to test retrieval and answer quality on real research documents rather than short, simplified passages.
evaluationbenchmarkarxiv - Academic & scientific literatureOpen
OpenAlex
A free, open catalogue of the world's scholarly works, authors, venues, institutions, and research topics. It succeeds Microsoft Academic Graph and holds hundreds of millions of works with rich metadata and citation links.
academicmetadatacitations - Encyclopaedic & general knowledgeOpen
OpenCyc
The open release of Cyc, one of the oldest attempts to hand-build common-sense knowledge for machines. It holds hundreds of thousands of concepts and millions of assertions about how the everyday world fits together. Now archived, but the data is still available.
knowledge-baseontologycommon-sense - Geospatial & mappingOpen
OpenStreetMap
A crowdsourced map of the world, built and maintained by a community of volunteers. It holds detailed geographic data on roads, buildings, land use, points of interest, and much more, which you can download as a full planet file or as smaller regional extracts.
crowdsourcedmapsgeographic - Web corporaOpen
OSCAR
A multilingual web corpus (its name stands for Open Super-large Crawled Aggregated coRpus) extracted from Common Crawl with per-document language classification. Covers more than 150 languages, useful for non-English and cross-lingual RAG.
web-crawlmultilingualpretraining - Geospatial & mappingOpen
Overture Maps Foundation
An open map data project run under the Linux Foundation by a group of industry members. It combines and harmonises data from OpenStreetMap, Microsoft, Meta, and others into consistent, interoperable map layers you can build on freely.
mapsopen-datainteroperable - Biomedical & healthOpen
PhysioNet
A repository of freely available medical research data, including physiological signals such as ECG and EEG, clinical databases, and related software. It hosts MIMIC and many other biomedical datasets, so it is a central hub for clinical RAG source material.
clinicalphysiological-signalsecg - Legal & regulatoryOpen
Pile of Law
A 256 GB dataset of open English-language legal and administrative text, covering court opinions, contracts, administrative rules, and legislative records. A ready starting point for building a legal RAG system without assembling sources yourself.
legalenglishus - RAG-specific & evaluationOpen
PleIAs RAG-Resources
A curated collection of datasets and resources assembled by PleIAs for building and evaluating RAG applications. A useful starting point when you want ready-made material to test retrieval pipelines without hunting for datasets one by one.
evaluationcollectioncurated - Books & literatureOpen
Project Gutenberg
A volunteer effort to digitise and archive public domain books, with more than 70,000 free ebooks. Mostly English, but it covers many languages. One of the oldest digital library projects, so it is a clean, permissively licensed source of full-text literature for RAG.
public-domainbooksliterature - Academic & scientific literatureOpen
PubMed / PubMed Central (PMC)
The US National Library of Medicine's database of biomedical and life sciences literature. PubMed indexes over 36 million citations and abstracts, and PubMed Central (PMC) adds free full-text access to a growing subset. You can download the data in bulk over FTP.
biomedicalhealthfull-text - RAG-specific & evaluationOpen
RAG-Mini-Wikipedia
A small evaluation set that pairs 918 question-and-answer pairs with a matching corpus of 3,200 Wikipedia passages. Built specifically for testing RAG systems, it is small enough to run quick evaluation loops while still covering realistic retrieval questions.
evaluationquestion-answeringwikipedia - Web corporaOpen
RedPajama
An open reproduction of the LLaMA training data from Together. V1 aggregates Wikipedia, books, arXiv, GitHub, Common Crawl, and Stack Exchange; V2 is a web-only corpus of more than 100T tokens with 46 quality signals per document.
web-crawlpretrainingmulti-source - Web corporaOpen
RefinedWeb
A filtered English web dataset from the Technology Innovation Institute, creators of the Falcon models. It showed that carefully cleaned web-only data can match curated multi-source corpora for language model training.
web-crawlenglishpretraining - Academic & scientific literatureOpen
Semantic Scholar / S2ORC
Allen AI's academic graph, covering hundreds of millions of papers linked by billions of citations. S2ORC, the Semantic Scholar Open Research Corpus, is the downloadable dataset with full text and parsed references, drawn from arXiv, PubMed, Crossref, publishers, and web crawlers.
academiccitationsfull-text - Web corporaOpen
SlimPajama
A cleaned and deduplicated version of RedPajama-V1 from Cerebras. It drops short documents and removes near-duplicate content, cutting 1.2T tokens down to a denser 627B tokens without losing coverage.
web-crawlpretrainingcleaned - Data platforms & marketplacesCommercial
Snowflake Marketplace
A data marketplace connecting more than 820 providers who offer over 3,400 live datasets, data services, and applications. Listings span financial, weather, demographic, and industry data, with a mix of free and paid options delivered straight into a Snowflake account.
data-platformmarketplacecommercial - Code & technical documentationOpen
Stack Exchange Data Dump
Periodic XML dumps of every site in the Stack Exchange network, including Stack Overflow. Each dump holds questions, answers, comments, votes, and user profiles. One of the richest question-and-answer datasets you can freely download, and a natural fit for building technical support and developer RAG systems.
question-answeringstack-overflowcommunity - Books & literatureOpen
Standardized Project Gutenberg Corpus (SPGC)
A research-ready version of the Project Gutenberg catalogue with consistent formatting, tidy metadata, and token counts for every book. It saves you the cleanup work, so you get uniform full-text literature ready to chunk and embed for RAG.
public-domainbooksliterature - Web corporaOpen
The Pile
An 825 GB curated English text dataset from EleutherAI, made of 22 sub-datasets spanning books, academic papers, code, web content, and more. Built as a diverse training corpus and widely used to train early open language models.
pretrainingmulti-sourceenglish - Code & technical documentationOpen
The Stack v1 / v2
A large source code dataset from the BigCode project, built from permissively licensed code on GitHub with duplicate and near-duplicate files removed. Version 2 draws from Software Heritage and covers 600+ programming languages, making it a strong base for training and retrieval in code-aware RAG systems.
source-codepermissive-licencepretraining - Curated lists & meta-resourcesLimited
UK Data Service
The UK's largest collection of economic, population, and social research data for teaching, learning, and public benefit. Many datasets need registration or an institutional login, so plan for an access step before you build with them.
uksocial-scienceresearch - Biomedical & healthOpen
UniProt
A resource for protein sequence and functional information, with more than 250 million protein sequences and rich annotations. It is a core reference for bioinformatics, so it grounds life sciences RAG systems in reliable protein facts.
proteinssequencesbioinformatics - Academic & scientific literatureOpen
Unpaywall
A database that tracks where scholarly articles can be read legally for free, covering over 30 million open-access papers. Maintained by the nonprofit OurResearch, it links each article's DOI to full-text versions hosted across repositories and journals.
open-accessdoimetadata - Encyclopaedic & general knowledgeOpen
Wikidata
A free, collaborative knowledge base that holds structured facts for Wikipedia and the other Wikimedia projects. More than 100M items, each one machine-readable with statements, relationships, and identifiers, so you can pull clean facts instead of parsing article prose. Released into the public domain under CC0.
knowledge-basestructuredmultilingual - Encyclopaedic & general knowledgeOpen
Wikipedia
Complete database dumps of Wikipedia and the other Wikimedia projects, in every language, refreshed roughly twice a month. The most widely used knowledge source for RAG systems, and an easy first corpus to build a retrieval pipeline on.
encyclopaediamultilingualgeneral-knowledge - Knowledge graphs & structured dataOpen
WordNet
A lexical database of English that groups nouns, verbs, adjectives, and adverbs into sets of synonyms called synsets, then links those sets by meaning. It maps how words relate, which sense means what, what is a kind of what, so software can work with meaning rather than just spelling.
lexical-databaseenglishsynonyms - Encyclopaedic & general knowledgeOpen
YAGO
A semantic knowledge base that combines facts from Wikipedia, WordNet, and GeoNames into a single, high-accuracy collection of statements about entities. It adds when and where each fact holds, so you get temporal and spatial detail alongside the plain relationships.
knowledge-basestructuredwikipedia - Data platforms & marketplacesOpen
Zenodo
A general-purpose open repository built by CERN where researchers can deposit datasets, software, reports, and any other digital output. Every upload gets a DOI, a permanent identifier that makes the work easy to cite and find again, which makes Zenodo a reliable long-term home for research data.
data-platformopen-accessresearch