Research Engineer
NYC/SF
About the role
We’re looking for a Research Engineer to work on building and testing AI systems for learning.
You’ll help shape the research questions, build the systems needed to investigate them, and carry experiments through to analysis and publication.
The initial focus will be teaching-system prototypes, evaluations, and studies with real learners. We’re interested in post-training where the evidence warrants it, but we don’t assume that better teaching necessarily requires changing model weights.
What you’ll work on
- Investigate teaching behavior. Identify where models help or hinder learning and develop testable hypotheses about what to change—for example, how they diagnose confusion, choose between explanations and hints, or adapt to a learner’s understanding.
- Build learning harnesses. Develop the prompts, interaction flows, learner-state tracking, practice tools, and other components that shape how a model teaches.
- Build experiment infrastructure. Create reliable pipelines for model comparisons, condition assignment, interaction logging, assessments, and analysis. Keep experiments reproducible and handle learner data carefully.
- Run studies with learners. Help design, implement, and analyze randomized controlled studies comparing models and teaching approaches. Measure immediate learning, delayed retention, and transfer alongside engagement and time spent.
- Develop evaluations and benchmarks. Build reusable tests of teaching behavior and investigate whether their scores predict real learning outcomes. Prepare code, methods, and appropriately shareable data for release.
What we’re looking for
- Strong software engineering skills, including proficiency in Python and the ability to build, debug, and maintain research software.
- Hands-on experience building with language models, such as applications, agents, evaluation pipelines, or training experiments.
- Good experimental judgment: you can formulate a testable question, choose useful measurements, identify confounds, and distinguish a promising result from an unreliable one.
- Comfort working with imperfect data and tracing unexpected results back through a system.
- Clear written communication and an interest in contributing to research, not only implementing it.
Helpful experience
You don’t need experience in all of these:
- LLM evaluations, benchmark development, or model-behavior analysis.
- Fine-tuning, preference optimization, or reinforcement learning.
- Human-subject experiments, causal inference, or educational assessment.
- Learning sciences, tutoring, or designing learning experiences.
- Open-source contributions or research writing.
Apply
Email hello@integraleducation.org with your CV, links to relevant work, and a short note about why this question interests you.
We’d particularly like to hear about something you built or investigated: what you were trying to achieve, what you contributed, and what you learned. This could be a research project, a product, an open-source contribution, or an independent experiment.