PhET Interactive Simulations
PhET Interactive Simulations is a collection of over 125 free interactive maths and science simulations, founded in 2002 by Nobel laureate Carl Wieman and run by the University of Colorado Boulder. Simulations span physics, chemistry, biology, earth science and mathematics, pitched from roughly grade 4 through to introductory university level, and are complemented by a searchable bank of teacher activities and lesson plans contributed by the PhET team and its user community. You access everything through the web portal: simulations run in the browser as HTML5 apps, activities download as documents, and the underlying simulation source code is published openly on GitHub (github.com/phetsims). There is no dataset dump or public API in the usual sense, so this is a resource you crawl and curate rather than bulk download. For RAG, the value sits in the activity and lesson bank: level-graded, subject-tagged teaching materials with learning goals, prerequisites and guiding questions that map cleanly onto a curriculum. For an AI tutor, that same structure supports curriculum sequencing and level-graded explanation, and the activities model tutoring pedagogy directly (guided enquiry, prompts that surface misconceptions, worked exploration steps), which is harder to extract from a plain textbook. The simulations themselves are interactive rather than text, so treat them as referenced artefacts, not retrievable passages. The licence position needs care. Teacher activities are CC BY (commercial use permitted with attribution). Simulations published before 29 March 2026 are CC BY 4.0, but simulations released after that date carry CC BY-NC 4.0, which prohibits commercial use, so anyone building a paid tutor must check each simulation's release date. Source code is separately licensed as MIT or GPLv3. Attribution is required throughout and there is no share-alike obligation. It complements the textbook-style corpora we list (openstax, ck-12, siyavula) and the standards we index (common-core, ngss) by adding interactive, activity-level pedagogy, sitting closer to khan-academy in intent than to raw content.
Related sources
Achievement Standards Network
Machine-readable curriculum standards from US states, national bodies and other jurisdictions, modelled as an RDF graph of URI-addressable learning objectives with cross-jurisdiction alignments. Now run by D2L and free to use.
AGIEval
8,062 questions drawn from 20 official standardised exams (SAT, LSAT, GMAT, GRE, Gaokao, AMC/AIME and more) in English and Chinese, packaged as a human-centric benchmark for evaluating foundation models.
CASE Network (1EdTech)
A public registry of machine-readable learning-standard frameworks from all 50 US states and other issuing agencies, run by 1EdTech in the CASE JSON format. The digitally referenceable spine of what to teach, at which level and in what order, rather than the teaching content itself.
CEFR Companion Volume Descriptors
The Council of Europe's 2020 CEFR Companion Volume descriptor set, the de facto international standard for levelling language proficiency from Pre-A1 to C2. A spine of can-do statements for sequencing and grading language teaching.