9th Grade Algebra: A Simulated Student Cohort
na8ve provides teachers and parents with individual and self-contained learning models for each student. The model adaptively scaffolds learning based on the child’s needs. Here, we describe what that model records, where it runs, what it decides, and how much of that reasoning a parent or teacher can actually see. The model described here is the one running in your na8ve learning studio, and the figures below are what it produced during a research run of a simulated 9th grade algebra cohort.
The na8ve studio has three parts:
- The first is a per-student memory:for every skill a student has touched, the model holds an estimate of how confident it is that the skill is genuinely learned, how long that confidence should hold before it fades, how effortful the student’s answers looked, and whether the student appears to be fatigued.
- The second is where that memory lives, which is a continuously running cognitive layer with a clock of its own, separate from the servers that render lesson pages, so it can decide something needs attention before anyone asks it a question.
- The third is a teacher-facing dashboard for student insights, which shows what the model currently believes about a student and why it believes it, and exists so that no decision about a child is made somewhere the adult responsible for that child cannot see, modify, and control it.
Our research cohort involved: three simulated student profiles; Maya, Theo, and Priya, who all worked ninth grade material across an algebra course and a physical science course. We wrote every question-answer pair simulated, including the failures, specifically to put the model in positions that ordinary homework rarely produces. The model’s responses are real and were measured during the research run.
The three are not three personalities. They are three positions a student can be in relative to a model that is scaffolding their learning, and the run is about what the model does from each of those positions.
- Maya and Theo arrive as the ordinary case. The model had watched them work but had formed no settled view about this particular skill yet, so its confidence sat at the neutral midpoint and what they did next is what decided. Both of them did well. What makes them worth naming separately is that they did well by measurably different amounts and still finished in exactly the same place, which is the part of the run that shows the model working correctly.
- Priya is the harder case, and the one the model keeps returning to. Here the model already held a settled and positive view about her on this skill, built from work it had watched her do earlier. Her profile carried two other things that shaped what happened next: a misconception in equation solving that it had already flagged, and a test anxiety reading high enough that the teaching approach it chose was built around reducing it. She then failed the same skill twice, and the two failures were not treated the same way.
Here is what the model held about each of them. Every figure below is one the model produced rather than one we assigned, and the three are authored profiles rather than children:
Maya
Began the skill at 0.50, the neutral starting point: the model held no settled view yet about whether she had learned forces and Newton's laws. Her work that day was scored 0.91, the strongest evidence of the three, and it moved her to 0.66.
Theo
Also began at 0.50, with the same absence of a prior. He scored 0.87 and landed on exactly the same 0.66 as Maya. The same starting point meeting the same class of result gives the same answer.
Priya
The only one the model already had a settled view about, at 0.66 before she began, built from work it had watched earlier. Her profile also carries a flagged misconception in equation solving and a test anxiety reading of 72, the highest of her readings, which is what the chosen teaching approach was built around. Two failures on the skill took her to 0.43 and then to 0.39.
The clearest of those results was a disagreement we discovered. On a single afternoon, all three students worked the same skill, forces and Newton’s laws. Maya and Theo answered well and were moved forward, both landing on precisely the same confidence figure of 0.66 because the same starting point meeting the same kind of result produced the same arithmetic every time. Priya failed that skill twice, on two separate attempts, and the model treated the two failures differently.
The first time, the model did not react. Its confidence that she understood forces and Newton’s laws was already fairly high, built from everything else it had watched her do, and one bad piece of work against a belief it had would be a model grading error, not a pattern. It moved her forward instead of automatically holding her back.
The second time she failed the same topic, its confidence in her had genuinely dropped, and this time it sent her back to re-teach it. On the same afternoon, two other students in the cohort, Maya and Theo, did the same topic and passed, and both were moved forward.
In the research run, three students, taught one topic, had two different results. This describes how na8ve models perceive and adapt differently from a general-purpose AI tutor: it keeps a real memory of each student, it uses that memory to decide what happens next, and it will change its own mind when the evidence changes, but not on a single bad afternoon.
Two further results are worth mentioning here.
In an earlier run, the model pulled a skill from an algebra course, slope and rate of change, back in front of two students partway through a physical science course several weeks later. This happened close to the point that it estimated memory was about to fade. Nothing in the physics course mentioned it, but in the hours around the run reported here, the layer holding all of this woke itself up 12 times to check for work, wrote 9 new observations into the model’s memories, read a student’s full record 6 times in order to write their next lesson, and finished 35 background jobs with nobody watching. These figures are detailed in our observation logs.
What none of this shows is that any of it teaches a child better. Different is not better, and the study that would settle the question needs real students, a real term, and a willingness to publish a negative result. It has not been run. What follows is an account of what the model does, how it decides, and what it shows a teacher while it works with students.
Adaptive learning trajectories

Dynamic Adaptive Pathways
Most course-building tools, ours included, work by planning everything up front. You describe what you want to teach, you get a full course back, and you teach it in order. That is the right tool when you already know the sequence of your curriculum: a syllabus you are following, a co-op where four families need to be on the same lesson in the same week, anything you intend to print.
It is the wrong tool when you genuinely do not know where a student is yet. So na8ve also offers a second way to start a course, and it begins with exactly one lesson.
In the na8ve Create a Course menu, you describe the course you want to teach. Behind the scenes, na8ve plans the whole outline. For adaptive learning, it only writes the first lesson: a short warm-up whose entire job is to find out where the student actually is, rather than where a curriculum written in advance assumes they are.
After the lesson is created, your student does the warm-up. Every other lesson in the outline stays as an outline of potential lessons, invisible to them until it has actually been written by the model. From there, each time your student finishes a piece of work, the model decides what they should do next and scaffolds pedagogical strategies: move forward, revisit something that has started to fade, or go back over something that did not land. Two students working through the same course can end up doing genuinely different courses.
What the learning model remembers about a student
The decision at the center of all of this is based on cognitive learning science. For every skill a student has touched, na8ve keeps a small, specific memory:
Confidence
How sure the model is that the skill is genuinely learned, not just answered correctly once. Updated by every piece of work, and it moves in different directions depending on which side of a threshold an answer lands.
Retention
How long that confidence is expected to hold before it fades, so a review can be scheduled right at the point it is about to be forgotten, which is when a review is worth the most.
Fluency
Whether an answer looked effortless, careful, or like a guess. Two students who both got the right answer did not necessarily do the same thing to get there.
Fatigue
Whether a student looks like they are running low, so the system can choose a shorter, gentler next step instead of piling on.
What makes this a memory and not a scoreboard is that it follows a student across courses. In an earlier run, two students in the cohort were sent back to review a skill called slope and rate of change, in the middle of an entirely different course on physical science. They had learned it weeks earlier in an algebra course. Nothing about the physics course mentioned it. The model brought it back because its own estimate of how long that memory would last said it was due, at almost exactly the point either of them would have started to forget it. That is the part I want you to sit with: the record is not about one class, it follows the child’s entire learning trajectory.
What students are presented with is a tutor slightly better tuned to how they work and learn, and lessons that occasionally bring them back something they learned a month ago, at the very moment they were about to lose it.
Inside the cognitive layer

The Cognitive Layer
Everything described so far, the memory that follows a student, the module built for them, the guide that introduces it, has to actually run somewhere.
Underneath the lesson pages and the Insights board sits what we call the cognitive layer: a place in the model built specifically to hold this kind of record: what a student has done, how confident the model is in each skill, when a memory is due to fade, and what teaching strategy is currently in use and why. The cognitive layer is a separate, continuously running process within the model, and it keeps working whether or not anyone has a browser open at that moment.
That distinction matters more than it sounds. A lesson page can only react to what just happened with the student’s engagement. The cognitive layer can decide something needs attention before anyone asks, because it has its own cognitive process rather than only responding when a request arrives.
Here is what that actually looked like, tracked in the hours around a research run. This is the layer’s own cognitive record:
Tracking the Cognitive Layer
The cognitive layer, measured
| What happened | Times |
|---|---|
| Its own clock woke it up to check for new work | 12 |
| A new observation written to a student's memory | 9 |
| A student's full record pulled to write their next lesson | 6 |
| An existing memory re-weighted by how recent and relevant it still is | 6 |
| A new answer compared numerically against everything already known about that student | 10 |
| Background jobs claimed and finished without anyone watching | 35 |
Two of those rows are the ones I would point a skeptical reader to. “A student’s full record pulled to write their next lesson” is the exact step behind what we’ve been researching: it is where Priya’s second failure and Maya’s and Theo’s passes each produced a different lesson, because each of those lessons were written after this step read the layer’s memory of that one specific student. And “an existing memory re-weighted by how recent and relevant it still is” is what keeps the Insights dashboard honest over time. A fact from January does not carry the same weight in March, and the layer downgrades it on its own cognition rather than waiting to be asked.
This is also, directly, where the Insights dashboard gets what it shows a teacher. Nothing on that page is computed the moment you open it. It is a window into the model that has already been written, by a process that has been running the whole time your student was working, whether or not a teacher ever opens the dashboard or not. The same record drives both things we’ve described: which module a student is handed next, and what a teacher sees when they ask why that strategy fired. All model decisions come from measured weights rather than prompted responses, and provide insights into its processes rather than taking a general-purpose AI black box on faith.
Cognitive insights, reading the learner's profile

Priya's actual Insights board, captured live
Everything above happens whether or not a teacher ever looks at it. But there is a dashboard for accessing a model’s decisions for full auditing and control.
In the dashboard, the student roster is alphabetical. In the dashboard, there is no rank, no composite score, and no way to compare one student against another, because that comparison is not something the model factors. Its focus is on the individual student. Open one student’s profile and you get a handful of readings on how they approach new material and how they respond when stuck, individual observations, and a strip of anything worth knowing right about the student such as fatigue or a fading skill.
Two rules govern the Insights page:
- The first is that a blank space means we have not seen this yet, never they cannot do it, and the model will not quietly fill a gap with an average-looking number to make the page look more complete than the evidence supports.
- The second is that nothing on the page ever touches a grade. Grading runs on a separate track, and every mark your student receives is still one a teacher approved, by hand.
Your student does not see their own insights. A child handed a running readout of their own cognition starts performing for it, and that is not what this is for. na8ve is built to keep teachers and parents in control and engaged in the student’s learning journey.
Why this only works with a human in charge
Stanford’s SCALE Initiative published a systematic review of the causal research on AI tools in education in 2026, and its central finding is uncomfortable for a lot of general-purpose AI tools. The findings mention that AI tutoring tools perform badly when they operate as the decision-maker, and they perform well when they sit behind a human who remains in charge.
That finding is the reason the Insights board exists at all, and the reason it is built the way it is. A model that decided what a student needed next and never told anyone why would be exactly the kind of tool the Stanford review found does not work. Every decision na8ve makes is meant to be seen, argued with, and if needed, overruled by the adult, and grading never leaves that adult’s hands regardless of what the model decides about the lesson itself.
The design decision that follows from that is a strict one: nothing here is allowed to replace the adult. It is only allowed to make the adult faster and better informed, which is why the guide that introduces an adaptive lesson exists, why the Insights dashboard explains its own reasoning rather than just stating a conclusion, and why a student never sees any of it unless shared by the adult.
The data behind every research run
The following represents the data from our research run, in full.
One skill, three students, four decisions
| Student | Skill | Evidence | Confidence, before → after | Decision |
|---|---|---|---|---|
| Priya (1st attempt) | Forces & Newton's Laws | 0.58 | 0.66 → 0.43 | Advance |
| Priya (2nd attempt) | Forces & Newton's Laws | 0.45 | 0.43 → 0.39 | Reteach |
| Maya | Forces & Newton's Laws | 0.91 | 0.50 → 0.66 | Advance |
| Theo | Forces & Newton's Laws | 0.87 | 0.50 → 0.66 | Advance |
Every one of the confidence figures is a small, checkable calculation, that allows us to look into a system that would otherwise be a black box. The confidence represents a fixed starting point, which moved a fraction of the way toward a target that depends on how the answer was classified. Maya’s and Theo’s numbers landed on exactly the same figure, 0.66, because the same starting confidence meeting the same kind of result produced the same answer every time. Priya’s first failure was read as a genuine wobble against a strong prior record and the model moved her on. Her second failure, on the same skill, is what took her confidence low enough, and made the pattern clear enough to send her back.
This doesn’t mean the model has an opinion about Priya, or that “reteach” is a judgment about her as a learner. It means two consecutive pieces of evidence pointed the same direction on one specific skill, that a single datapoint did not.
Getting a student started
In practice, this is what a parent or teacher actually does:
1. Describe the course
Same form as any other course: topic, grade level, what you want covered. Choose "Adaptive path" instead of "Everything" when you do not already know where your student is.
2. One lesson gets written
The outline is planned in full; only the first, warm-up lesson is actually written. Nothing else exists yet for your student to see.
3. Your student works through it
Same lesson experience as any other module: the tutor is there if they get stuck, and a short reflection at the end.
4. The next lesson gets written for them
A lesson plan and a short guide arrive with it, introducing what it is and why it turned up, since nobody chose this module by hand.
That fourth step is the one I think matters most to a parent evaluating this. A module nobody chose by hand needs something to introduce it, and the guide is that introduction. It opens automatically the first time your student reaches the lesson, once, and then it gets out of the way.
The future direction of na8ve development
Everything above is real: it happened on our own production model, and the numbers in that table are what it actually produced, not a simulation of what it might do.
What is not real is the cohort. Maya, Theo, and Priya are authored research profiles, not children, and every answer they submitted, including the failures, came from question-answer pairs written to provide the AI tutor with varying pedagogical strategies for adaptive teaching. Future studies would be needed to show whether our na8ve models help a real child learn faster than general-purpose AI tutors or well-planned courses written by hand. Future studies will follow the model’s teaching with real students.
The research run shows what the model does: it remembers a specific student across courses and weeks, it changes a lesson’s content and not just its topic based on its memory, it holds back a decision rather than relying on a single data point, and every one of its decisions is visible to an adult rather than hidden inside a black box. Whether that combination teaches a child better is the next thing worth finding out, and I will publish the findings once we do know.
Source: Stanford SCALE Initiative, “The Evidence Base on AI in K-12: A 2026 Review.”
The complete na8ve research archive, including the run data and the cognitive layer’s own operations record cited in this article, is published at na8ve.ai/research.
David Laurenvil is the developer of na8ve.ai. He previously served as Director of Education at the Fleet Science Center in San Diego, CA, and Executive Director of Kids MakeIt Institute, a 21st-century educational institution focused on exposing students to Science, Technology, Engineering, and Math (STEM) skills and careers.
