Skip to content
RAG Repo

Project Gutenberg

Project Gutenberg began in 1971 and is one of the oldest digital library projects in existence. Volunteers digitise, proofread, and archive books whose copyright has expired, giving you full-text works you can use without licensing worries. The collection is mostly English, with roughly 55,000 English-language titles, alongside thousands of books in other languages.

You can download individual books as plain text, EPUB, HTML, or Kindle files, or pull the whole catalogue in bulk. The plain-text editions are the easiest to chunk (split into passages) and embed for retrieval, though be aware that older files include a standard header and footer you will want to strip out first.

Copyright status is handled per work rather than across the whole collection, so a title that is public domain in the United States may still be under copyright elsewhere. Check the status of each work for your jurisdiction before you build on it.

public-domainbooksliteratureenglishmultilingualnonprofit

Related sources