Internet Archive
The Internet Archive has been preserving digital and digitised culture since 1996. Its collections span scanned books and periodicals, audio and video, vintage software, and the Wayback Machine, which lets you view web pages as they looked at points in the past.
You can browse and download most items directly, and the Archive offers APIs and bulk access for larger projects, alongside command-line tools for fetching items and metadata. Licences vary by collection, so check the terms on each item before you reuse it, since some material is public domain and some is not.
For RAG, the Internet Archive is valuable for historical text, out-of-print books, and archived web pages you cannot find live any more. The Wayback Machine in particular is a strong source when you need a snapshot of how a page read at a specific date.
Related sources
CC-News
A subset of Common Crawl focused on news, containing millions of articles pulled from news websites around the world. It has fed several large language model training pipelines and gives you a ready news corpus without crawling sites yourself.
GDELT
GDELT, the Global Database of Events, Language, and Tone, monitors print, broadcast, and web news in more than 100 languages from every country. It records events back to 1979, builds a Global Knowledge Graph, and adds sentiment and emotion analysis, all queryable on Google BigQuery.