arXiv
arXiv has been the home of scientific preprints since 1991, letting researchers share work before, or instead of, formal journal publication. It covers physics, mathematics, computer science, quantitative biology, quantitative finance, statistics, electrical engineering, and economics, and most modern machine learning research lands here first.
For bulk work you have two routes. Full-text sources and PDFs are available as a requester-pays dataset on Amazon S3, meaning you cover the download costs. Metadata can be harvested for free through OAI-PMH, the Open Archives Initiative protocol that repositories use to expose their records for bulk collection.
Licensing varies by paper. Most are submitted under arXiv's non-exclusive distribution licence, while some carry a Creative Commons licence chosen by the author, so check the terms on each paper before redistributing text.
Related sources
CORE
An aggregator of open-access research papers that harvests from thousands of repositories and journals worldwide. It holds over 300 million metadata records and more than 40 million full-text articles, all reachable through one search API.
Crossref
A nonprofit DOI registration agency that publishes metadata for over 150 million scholarly works, including journal articles, books, and conference proceedings. A DOI is the permanent identifier assigned to each work, and Crossref shares its metadata through a free API.
OpenAlex
A free, open catalogue of the world's scholarly works, authors, venues, institutions, and research topics. It succeeds Microsoft Academic Graph and holds hundreds of millions of works with rich metadata and citation links.
PubMed / PubMed Central (PMC)
The US National Library of Medicine's database of biomedical and life sciences literature. PubMed indexes over 36 million citations and abstracts, and PubMed Central (PMC) adds free full-text access to a growing subset. You can download the data in bulk over FTP.