Crossref
Crossref is the organisation most publishers use to register a DOI (Digital Object Identifier), the permanent link that always points to a given article or book even if its web address changes. Registering a work also means depositing its metadata, and Crossref makes that metadata openly available.
Each record can include titles, authors, affiliations, funding, licences, abstracts, and reference lists, which together map how the scholarly record fits together. You reach it through a free REST API or bulk metadata dumps, and the metadata is released under CC0, placing it in the public domain.
For RAG, Crossref is less about full text and more about clean, authoritative metadata: resolving citations, deduplicating a corpus by DOI, and enriching documents with reliable bibliographic detail.
Related sources
arXiv
An open-access preprint server for physics, mathematics, computer science, quantitative biology, statistics, and more, holding over 2.5 million papers. You can pull the full archive in bulk from Amazon S3, or harvest the metadata through OAI-PMH, a standard protocol for sharing records between repositories.
CORE
An aggregator of open-access research papers that harvests from thousands of repositories and journals worldwide. It holds over 300 million metadata records and more than 40 million full-text articles, all reachable through one search API.
OpenAlex
A free, open catalogue of the world's scholarly works, authors, venues, institutions, and research topics. It succeeds Microsoft Academic Graph and holds hundreds of millions of works with rich metadata and citation links.
PubMed / PubMed Central (PMC)
The US National Library of Medicine's database of biomedical and life sciences literature. PubMed indexes over 36 million citations and abstracts, and PubMed Central (PMC) adds free full-text access to a growing subset. You can download the data in bulk over FTP.