Unpaywall
Unpaywall answers a single question well: given a paper, is there a legal free-to-read copy, and where is it? It maps each work's DOI (its permanent identifier) to open-access versions it finds in repositories, on publisher sites, and in preprint archives.
The data comes from the same nonprofit behind OpenAlex, OurResearch, and is released under CC0, so you can reuse it freely. You can query it through a free API, download the full database snapshot, or use the browser extension that surfaces free copies as you browse.
For RAG, Unpaywall is the filter that turns a list of citations into a list of documents you can actually fetch and index. Pair it with a metadata source like Crossref or OpenAlex to find the papers, then use Unpaywall to locate the full text you are allowed to use.
Related sources
arXiv
An open-access preprint server for physics, mathematics, computer science, quantitative biology, statistics, and more, holding over 2.5 million papers. You can pull the full archive in bulk from Amazon S3, or harvest the metadata through OAI-PMH, a standard protocol for sharing records between repositories.
CORE
An aggregator of open-access research papers that harvests from thousands of repositories and journals worldwide. It holds over 300 million metadata records and more than 40 million full-text articles, all reachable through one search API.
Crossref
A nonprofit DOI registration agency that publishes metadata for over 150 million scholarly works, including journal articles, books, and conference proceedings. A DOI is the permanent identifier assigned to each work, and Crossref shares its metadata through a free API.
OpenAlex
A free, open catalogue of the world's scholarly works, authors, venues, institutions, and research topics. It succeeds Microsoft Academic Graph and holds hundreds of millions of works with rich metadata and citation links.