Skip to content
RAG Repo

Cambridge Law Corpus

The Cambridge Law Corpus gathers over 250,000 UK court cases into a single dataset built for studying legal language and testing legal natural language processing, the task of getting software to read and reason over legal text. The bulk of the cases are recent, but the collection stretches back several centuries, giving it unusual historical depth.

Access is restricted. The corpus is released for research use rather than open download, so you will need to go through the University of Cambridge's process and agree to its terms before you can work with it. The accompanying paper on arXiv describes how the data was collected and annotated.

For RAG, it is most useful as a research and evaluation corpus for UK case law, rather than a source you can drop straight into a production system.

ukcase-lawlegalacademicrestricted-access

Related sources