PSLC DataShop
DataShop is a repository and analysis service for learning interaction data, built and maintained by Carnegie Mellon University's LearnLab (the former Pittsburgh Science of Learning Center) and running since 2006. It specialises in fine grained logs of students working through intelligent tutoring systems, online courses, virtual labs, assessment systems and simulations, spanning mathematics, physics, chemistry, computer science and language learning (Chinese, French, English and Spanish) across school and university levels. Its distinguishing feature is that student actions are coded against knowledge components, the hypothesised skills a step requires, not merely marked correct or incorrect, which supports mastery modelling and the study of how practice maps to learning. Access is gated. You sign in with InCommon or Google single sign-on, then browse public (shareable) datasets, while private datasets need permission from the owning principal investigator. Data exports as tab delimited student-step and transaction files (usable in statistical software), with XML and CSV options and a web services API for programmatic pulls, so the site works as a web portal, a bulk export and an API in one. The licence position matters and is restrictive. The terms of use state the data is for research, not commercial, purposes, prohibit redistribution (collaborators must open their own accounts), and forbid any attempt to de-anonymise students. Per dataset terms and copyright owner permissions apply on top of the global terms, so reuse rights vary from set to set. For anyone building a commercial tutor this is prohibited territory as it stands: treat it as a research and evaluation resource, not a shippable content source. For an AI tutor this is a pedagogy and evaluation asset rather than a content corpus: it is where you study step level learner behaviour, calibrate mastery and knowledge tracing models, and benchmark tutoring policies against real student data. That complements the teaching content we list, such as openstax, ck-12, siyavula and khan-academy, the standards frameworks common-core and ngss, and the maths reasoning sets amps, megamath, stackmathqa, naturalproofs, case-network and openthoughts3: DataShop supplies the interaction evidence those content and reasoning sources do not.
Related sources
AWS Data Exchange
A marketplace for finding, subscribing to, and using third-party data inside the AWS cloud. It carries both free open datasets and paid commercial data products, so you can pull licensed data straight into your AWS workflows without setting up separate transfers.
AWS Open Data Registry
A registry of high-value datasets hosted on AWS and made publicly available, covering genomics, geospatial data, climate, satellite imagery, and more. Over 300 PB of data in total, free to access: you pay only for the compute you use to process it.
Datahub.io
A platform for publishing and finding open data packages: datasets bundled with consistent metadata in standardised formats. It hosts curated collections of widely used reference data, from country codes to exchange rates, ready to drop into a pipeline.
Google Dataset Search
A search engine for datasets that indexes millions of them from thousands of repositories worldwide. It does not host data itself: it points you to wherever each dataset lives, which makes it a fast first stop when you are hunting for a source on a specific topic.