Integral · an education lab
AI mediates a growing share of the world’s learning. For many, it is now the first place a question goes.
Learning with AI today leaves a lot to be desired. The fixes live at every level of the system. Some in the weights, some in the products built around them, some in distribution — kids who’d benefit most are served last. And the fourth layer barely exists at all: the infrastructure and policy for how humans and AI share a classroom.
AI cannot replace educators, but it should be integral to education. Our work spans every layer — weights, scaffolding, distribution, the institutions around them — and it begins with the instrument that makes the rest improvable: measurement.
The problem
Language models are excellent explainers. But can they teach?
Measured learning and felt learning routinely diverge.Deslauriers et al., PNAS 2019Measuring actual learning versus feeling of learning — students in active-learning classrooms learned more than students hearing a polished lecture, yet rated their own learning lower. Yet models are tuned against rubrics and rater preference. The education-specific version of sycophancy: models are optimized for the learner's experience rather than progress.
The labs have tuned for pedagogy. LearnLM shipped as a standalone model, but it was sunsetted as its improvements were fused into mainline Gemini. Other labs have shipped study modes as system prompts over their defaults, then removed them. Education-specific behavior does not survive at the labs because education is not the product.
A study mode is judged in the assistant's currency — engagement, satisfaction — and a mode built to slow you down loses in that court. The counter-metric that could defend it, measured learning, exists only as one-off deployment trials.The outcome-grounded results, in fullTutor CoPilot — AI assist for human tutors: ~4 percentage points of student mastery, ~9 for students of the weakest tutors, at about $20 per tutor.Rori — a WhatsApp math tutor in Ghana: ~0.37 SD.World Bank, 2025 — a GPT-4 after-school program in Nigeria: ~0.31 SD.Kestin et al., 2025 — an AI tutor in undergraduate physics at Harvard roughly doubled learning gains over active-learning classes.Each evaluates a single product, after the fact. Nothing published converts measured learning into a reusable training and evaluation signal. The ceiling is unexplored because no one can train toward it in the open.
Where we begin
-
Measure
Existing benchmarks score models against human judgement, not results.What they score againstMathTutorBench — a pedagogy reward model.TutorGym — ITS-anchored checks.MRBench — a rubric taxonomy. We build the first open benchmark of teaching anchored to measured student learning.
-
Collect
The data that doesn’t exist in the open. We create it two ways: instrumenting new sessions, and joining existing archives to corresponding outcomes.
-
Simulate
Training needs millions of episodes; real students supply thousands. We build a student simulator whose validity is the deliverable: it holds misconceptions, updates from instruction, sits a post-test the tutor never sees.
-
Train
RL against the signal the field has never had: learning gain on a held-out post-test. Randomized real-student rounds check that simulated gains predict measured ones; a miss retrains the simulator before training resumes.
-
Release
Everything ships open: evals, data, the simulator, the instrumented runtime, weights, code, and results — including negative ones.
Roadmap
-
Outcome-anchored evaluation spec, v0months 0–6
An evaluation suite for AI tutors anchored to measured student learning on externally normed assessments — not rater preference. Specification and rubric published for comment before any scores are.
-
Instrumented tutoring corpusFall 2026 →
Expert tutoring with instructional decisions recorded in situ, linked to longitudinal outcomes on the same students. Modern Bloom's elementary pilot (~80 students, grades 3–4, math and ELA, 18 tutors, NWEA MAP, three administrations) launches fall 2026. First joined records expected late 2026.
-
Student simulator validity program2027
A simulated learner grounded in public interaction data and calibrated against real tutoring with pre/post measures.
Who it’s for
Researchers & labs
Run our evaluations, train against our verifier, use our data, replicate our results. Everything ships with the documentation to check our work.
Builders
Point your own models at our eval suite, simulator, and instrumented runtime. The infrastructure is model-agnostic by design.
Schools & tutoring orgs
Partner with us on studies. You get rigorous measurement of what your students actually learn; the field gets the data it lacks.
Commitments
These are the rules we operate under:
- Studies are pre-registered. Outcomes are measured on externally normed assessments we don’t control.
- Semester-grain claims anchor only to those external assessments. Proximal measures we build are published in full and validated against them.
- Our own models enter our benchmark under held-out assessment banks administered by third parties. Integral never certifies itself.
- We publish negative results with the same prominence as positive ones.
- Consent to use data for research and training is obtained prospectively, at the point of collection.
who we are
Integral is a small research team spanning quantitative finance, AI research infrastructure, and the learning sciences. We are partnering with Modern Bloom, whose live tutoring study grounds this work in real students and real stakes.
Interested in working with us? Say hello.