CEFR Companion Volume Descriptors
The CEFR Companion Volume is the 2020 edition of the reference descriptors that accompany the Common European Framework of Reference for Languages, published by the Council of Europe's Language Policy Programme in Strasbourg (authors Brian North, Enrica Piccardo and Tim Goodier). It updates and extends the original 2001 descriptor set, adding a new Pre-A1 level, richer description at A1 and the C levels, and whole new scales for mediation, online interaction, and plurilingual and pluricultural competence, plus modality-inclusive wording for sign languages. Each descriptor is a short can-do statement calibrated to one of the six main levels, A1 to C2, grouped into Basic, Independent and Proficient User. It is a proficiency framework, not teaching content: it says what a learner should be able to do at each level, not how to teach it.
You access it free as a PDF from the Council of Europe site, with translations into many languages, and there is a web-based descriptors search tool plus a Bank of Supplementary Descriptors for browsing and filtering by skill and level. Paper copies are sold through the Council of Europe bookshop.
For a RAG pipeline the descriptors chunk cleanly, since each is self-contained and level-tagged, giving you a controlled vocabulary for levelling any language-learning corpus. For an AI language tutor the value is direct: the descriptors decide what a learner at a given level should be able to do, so they drive curriculum sequencing, the grading of practice items and reading material, and mastery modelling against a recognised international scale.
On licensing, read carefully before building a commercial product. The Council of Europe holds copyright, and its standard notice authorises reproduction of extracts for non-commercial purposes only, on condition that the source is quoted. Commercial reuse of the descriptor text needs separate permission. It plays the same standards-spine role for languages that Common Core, NGSS and CASE Network play for other subjects: pair it with teaching content such as OpenStax, Khan Academy, CK-12 or Siyavula, which supply lessons it can align and level.
Related sources
Achievement Standards Network
Machine-readable curriculum standards from US states, national bodies and other jurisdictions, modelled as an RDF graph of URI-addressable learning objectives with cross-jurisdiction alignments. Now run by D2L and free to use.
AGIEval
8,062 questions drawn from 20 official standardised exams (SAT, LSAT, GMAT, GRE, Gaokao, AMC/AIME and more) in English and Chinese, packaged as a human-centric benchmark for evaluating foundation models.
CASE Network (1EdTech)
A public registry of machine-readable learning-standard frameworks from all 50 US states and other issuing agencies, run by 1EdTech in the CASE JSON format. The digitally referenceable spine of what to teach, at which level and in what order, rather than the teaching content itself.
CEFR-J English Profiles (Open Language Profiles)
Open datasets from Tono Lab at TUFS mapping English vocabulary and grammar to fine-grained CEFR-J sub-levels, from A1 to C2. Grade words and structures by level and gauge text difficulty for language tutoring.