Research

na8ve.ai Education

What if every student had their own learning model, instead of a general AI tutor?

Better yet, what if parents and teachers had full authority and control of their child's individual AI, so they can see how and why their students learn and achieve?That's what we're building at na8ve.

An abstract illustration on a pale green ground: a stack of dark cubes rising from layered emerald plates, joined by a curved line to a badge carrying the na8ve mark, beside a glowing circle holding an open book with an arrow rising from its pages.

David Laurenvil18 min read

Something every parent already knows

Every parent who has helped with homework knows something that almost no gradebook knows. You can watch two children work the same problem, both arrive at the same correct answer, and understand within about four seconds that they are nowhere near each other. One of them guessed it. The other one ground it out, or half-remembered a rule from three weeks ago. You know this, but then you write down the same checkmark for both of them, because a checkmark is all grade requires.

I spent years leading STEM programs, and what bothered me the most is that we had no standard approach for serving each individual student’s learning needs. We had software, and the software was happy to tell us who was above or below grade level, but it couldn’t tell us the one thing every teacher in the building actually needed, which was how the students learned. Too often we leaned into inquiry-based learning and Socratic methods without knowing if that matched the student’s learning style. So we guessed, the way teachers have always guessed, and the kids who got the most out of us were the ones lucky enough to be sitting near an educator who happened to guess right about them. To address this, na8ve develops personal learning models that keep track of the thing we were guessing at, and I want to share what that looks like in practice using a week from our own research records, including the decision I would have gotten wrong.

A chat bot is one model, treating everyone the same

The way most schools are meeting AI right now is through a chat box they type into. It is the same model for your child as it is for mine, the same model for a fourth grader in Fresno as for a graduate student in Boston, and you have no control or access to any of the data or analysis on your child. Whatever it believes about how children learn, it believed your child learns the same, and it will keep believing that no matter how many times you tell it differently. That is what people mean when they talk about bias in these models, and the part that gets missed is that the bias does not arrive as an opinion. It arrives as a default setting. A shared model has to assume it’s talking to the average student in order to say anything at all about yours. Not to mention that every child who is not that student; the careful one, the one reading in a second language, the one whose hands do not keep up with their thinking, gets measured against the average student who has a different learning style.

na8ve is built the other way around. Your child gets a model of their own. It holds what it has actually observed and learned about them. It is theirs and nobody else’s. It does not average them into a cohort and then compare your child to another. It measures your child’s progress against yesterday’s learning gains. It never compares them to another child, because every child learns differently. That single design decision is what makes our models fair, and it’s also what makes the model something you can open, read, argue with, and overrule, which is the part I care about most: parental control of the AI model.

What your child's model keeps

For every child, and for every individual skill that child is building, the model focuses on four things: how confident it is that they have the skill, how long it thinks the knowledge will last before they forget it, whether the skill has become automatic or still costs them visible effort, and how fatigued they are while working on the problem. Notice that a grade answers none of those. A grade tells you what happened on one day, but what’s absent from a grade is what a child needs in order to develop and retain a skill. A grade can make a statement about how a child did on a specific lesson, but their personal model asks a different question. The model wants to know how that student learns, and it exists to accelerate learning for one specific child. Not engagement, not time in the app, not completion percentage, but the thing education was created to provide: learning.

  • Confidence
    Certainty of mastery
  • Retention
    Time until forgotten
  • Fluency
    Effort cost
  • Fatigue
    Current state

The reason it targets these four together is that any single one cannot answer the two questions a teacher has to answer every morning, which are whether this child understands a topic and whether they will still understand it three weeks from now. Those two questions are rarely merged into the context of student learning. What would it look like if they did merge?

What that looks like on a screen

This is where teacher and parent authority becomes real rather than rhetorical. The model is not a black box that hands a teacher a recommendation and asks to be trusted. It’s an open window to the student’s learning style. Below are two screenshots from our research cohort, taken in the same week, one for a student who is comfortably ahead and one for a student who is struggling. Nothing on these screens grades anything, issues a credential, or stops a child from doing anything they want to do.

Ada, comfortably ahead

The learning insight surface for Ada, a fast and accurate student: the model says move on, nothing is due, and the teacher has pinned Socratic questioning as the approach in use.

The line at the top is the whole recommendation in three words. For Ada it reads Move on. Socratic Questioning. Not tired, followed by a plain sentence saying she learns most like an explainer, which means she learns by putting things in her own words. Underneath, the next-step line explains itself: nothing is due, so new material is the best use of her time. Five skills tracked, none of them due, the nearest one eight days out.

The two charts show how she learns and where it sits. The eight-point chart on the left is not a personality test, and we will not put one in this model. There is no OCEAN score here, no Myers-Briggs letter, nothing that provides a psychological diagnosis. The eight points are instructional: explainer, builder, investigator, connector, reviser, persister, collaborator, self-navigator, and they describe approaches that have worked for this child, not traits she is stuck with. The five-point chart beside it groups what has been seen about knowledge, process, drive, self-direction, and working with others. Both charts carry the same footnote: a low axis means we have not seen it yet, not that they cannot do it. Absent is not zero. A model that fills in a neutral guess where it has no evidence has invented an observation about a child, and we would rather show you a dash and tell you we have not looked.

The teaching approach card is where you take over. Ada’s says Socratic Questioning, and next to it, pinned by you.A teacher looked at the recommendation, decided, and pinned it, and the model now works inside that decision instead of re-litigating it every week. Cleo’s card, in the next capture, has not been pinned, so it shows what the model chose, when it chose it, and what it chose from: worked examples, selected 22 hours ago, drawn from 11 sessions of evidence, because she has an open misconception about identifying equal shares and worked examples model the correct procedure rather than asking her to rediscover it. You can read that reasoning, disagree with it, and pin something else from the dropdown. The model does not get a vote on being overruled.

Cleo, two skills due now

The learning insight surface for Cleo, an inconsistent student: two skills due now, cognitive load at 90 and flagged as worth acting on, worked examples chosen from eleven sessions of evidence, and the full facet list showing how each axis was composed.

The skills card is the part parents ask for first. Four skills tracked for Cleo, two due now. Equivalent fractions and equal shares both sit at 0.49 with their retention falling, and both are marked due now. Adding fractions sits at 0.70, is marked fluent, and is not due for three days. Partitioning is at 0.69, fluent, due in five. Under all of it: pace is measured against this learner’s own usual pace on each skill, never against anyone else’s.

The strip across the top is the state of the child, not the state of the work. Cleo’s cognitive load reads 90 and is colored amber. Two cards over, in the facet list, the same reading appears as cognitive load 0.90 to 0.10, worth acting on, inverted before it feeds an axis. The inversion is arithmetic and not censorship. High load is a hard day of learning, so it doesn’t read as a strong score when the axes are composed, and it isn’t hidden from a teacher who needs to know that their child is at capacity. The composed axis uses the corrected value, and the facet list shows each one by name.

The last card is the audit trail. The axes were composeddraws every line from the individual things the model observed to the eight points on the chart, and the caption tells you what each line is: a declared weight times the facet’s polarity-corrected value. These weights are written down by the model and are provisional. The weights are the explanation that a teacher can read. What reaches the browser is scores, words, and the arithmetic that represent the underlying learning taking place.

One week, three children

In our research, we keep three simulated students on the platform so that we can watch how the model behaves:

  • Ada is fast and accurate, and she is there so we can check that a child who is genuinely ahead is allowed to move on instead of grinding through material she has already proven.
  • Ben is accurate and slow, and he matters more than the other two put together, for reasons I will get to.
  • Cleo is inconsistent, has more work submitted than any of the others, but knows the least.

Last week all three of them happened to work on the same skill, adding fractions, and all three got it right. Here is what the model did with three correct answers on one skill:

A line chart of three learning paths over time: a fast and straight dark charcoal line, a slow and steady emerald line that climbs in stages, and an erratic soft mint line that spikes and drops while trending upward.
Three shapes of progress. Fast and straight, slow and steady, and erratic all arrive somewhere; only one of them looks like progress on a chart.

Table 1 · One skill, three correct answers

StudentAnswerConfidence beforeConfidence afterGain
Ada0.96 in 17s0.880.91+0.03
Ben0.91 in 55s0.660.76+0.10
Cleo0.92 in 12s0.570.70+0.13

Cleo gained more than four times what Ada gained from an equally correct answer. This is not a curve, nobody is being handed credit they did not earn, and the number is not a grade and never becomes one. It is a confidence estimate, and a correct answer from a child we were genuinely unsure about tells us more than the tenth correct answer from a child we have already watched succeed nine times in a row. That is not the model being generous to Cleo. This mainly reflects how certain we were of the outcome beforehand. We wrote down 0.91, 0.75 and 0.69 before the run, but got back 0.91, 0.76 and 0.70, which means that Cleo had the largest gain in learning for the week.

The same answers, read the other way

Now the second chart, which reads those same three answers and disagrees about who benefited.

Table 2 · What the same answers were worth to memory

StudentRetention when they answeredExpected to lastNow expected to lastGrowth
Ada100%6.7 days8.2 days1.22x
Ben54%1.0 day3.2 days3.2x
Cleo54%1.0 day3.2 days3.2x

Ben and Cleo both answered at the edge of forgetting, while Ada answered something she had not begun to forget, so her recall proved nothing the scheduled review did not already know and earned almost nothing, while theirs tripled. There is a well-established finding in memory research sitting underneath this, which states that reviewing something right before you would have forgotten it is worth more in learning than reviewing something you still hold comfortably in working memory. A comfortable review feels productive, while in reality it teaches almost nothing. This is the piece I want a parent to understand, because it explains the thing about na8ve that can look wrong from the outside. The model will sometimes leave a skill alone for three weeks, but that is not neglect. It is waiting for the moment the review is actually worth something.

Skill retention estimate · Cleo

Before review1 day
After review3 days

Skill retention estimate · Ada

Before review7 days
After review8 days

In the simulations before this one, Ben came back to a skill at the point where his model put his chance of recalling it at 89%, just under the threshold it aims for. He got it right, and his estimate of how long that skill would last went from ten days to thirty-three on the strength of a single answer.

Ben's skill retention estimate

Before review10 days
After review33 days

Read the two charts together and you have the argument for the whole design. The first says the answer was worth most to Cleo. The second says the answer taught the model most about Ben and Cleo and almost nothing about Ada. One number cannot say both of those things at once, and a model built on one number has to pick which truth to throw away, which in practice means throwing away whichever one was harder to measure.

The decision I would have gotten wrong, the bias trap

Here is the part that convinced me the model worked better than my initial assumptions. Before every one of these test runs we write down what we expect to happen, and this time I predicted the model would move Cleo forward onto a new skill. She had just turned in her best work in seven weeks, and there is one topic in her course she has never touched at all, which makes closing that gap exactly what the model should have done. Every instinct I have from years of working with kids said move her up, she earned it, let her celebrate it.

But the model didn’t send her forward. It sent her backward instead, to two topics she had already covered. My assumptions were shattered, and the model was right. Those two older topics were quietly slipping away from her while she was busy succeeding at something else, both sitting at 49% confidence with retention falling. I had the priority order backwards: a review should outrank new coverage.

The model refused to march a struggling kid onto brand new grounds while the ground behind her was giving way, and it made that call in the same week that it moved Ben forward, because Ben’s memory of his older material was solid and hers was not. Two children, same skill, same week, both correct, but sent in opposite directions. Nobody programmed that outcome. It came from the model reading two different children’s actual memory states and reaching two different conclusions, which is exactly the judgment a good teacher would make. This judgment can’t be inferred by simply reading a gradebook. na8ve exists to accelerate learning, not grades, and the more the model understands the individual child, the less biased it is in serving that child’s learning needs.

How na8ve models work, by introducing fairness

Pace is read against your child, never against the class. The model watches pace and compares your child to their own previous pace. A slow, careful, and correct answer is read as mastery rather than as a warning sign. We built it that way deliberately, and Ben, our slow accurate student, exists in the test cohort specifically so that we notice if we ever break this commitment.

Our forthcoming research papers state this as the single most consequential fairness decision in the model, and it is worth fully understanding. Picture a child who reads at the 30th percentile for their age: a comprehension question the average student answers in forty-five seconds may take that child two minutes. Measured against the median, every reading answer they ever give comes back marked slow, no matter how strong their comprehension is. This is when general models start prescribing speed drills for a child whose comprehension was never the problem.

English learners get hit harder, because a student working in a second language needs the extra time to process the language itself, which means the delay is coming from translation and not from the mathematics in front of them. Any threshold calibrated on native speakers will read that as struggling with the subject for as long as the child is learning in a second language.

Children with dysgraphia and kids who are neurodivergent are hit the same way, since their delay lives in how they process motor and neural signals. In all three cases, the clock is measuring something that has nothing to do with what the child knows. This is why our comparison is always your child against their previous work. The model is more interested in the progression and acceleration of learning than it is in speed and grading.

  • The first week is the week the model knows the least.The model needs about five to seven lessons before it knows a new student's pace. It falls back on a sixty-second default until then, and we flag the profile as provisional during that stretch.

  • Long gaps between topics are usually on purpose.If your child has not seen something in three weeks, that is generally the model deciding the skill is secure and that coming back to review it too early would waste the session.

  • Going backward is not punishment.If your child is sent to review something they already passed, it is because the estimate of what they will still remember has dropped, not because they did badly. Cleo was sent backward in the week described above, and it was the best decision the model made all week.

  • A completion credential means coverage, not attendance.The model issues a credential when every skill the course set out to cover is secure, not when enough modules have been submitted. Ada's course closed and her credential issued on that basis in the run before this one.

  • You can see the model's reasoning, and you can overrule it.The teaching approach is yours to pin. The axes show the model's calculation. The facet list shows the measured value under its real name even where the composed axis corrects for it. Nothing on that surface grades your child or gates anything they are allowed to do. If you think the model has your child wrong, you can adjust the pin, because you know your child better than the model does, and sometimes the model will be wrong.

  • If your child seems stuck, please tell us.The most useful thing I learned this month is that a model failing a child and a child finding the work hard look identical from the outside unless something forces you to look closer. There will be failures we have not thought of, and you will see them before we do.

Who this is for

The parents, teachers, and children this model exists for are the ones who have historically been served last and understood least by general chat bots. With the progression of AI, the biases in large general models will keep growing if they are not addressed in their training and research. A single model answering every child the same way is a recipe for scaling biases onto a child it doesn’t understand. Alternatively, a model that belongs to one student and holds only what has actually been observed about them, one that shows an adult exactly how the student learns and how it reached its conclusion, one that hands over full authority and control to the adult, is a completely different kind of model. We’re excited to share that model with those who need it most, and to share the journey of na8ve, what we believe, and the evidence we discover while keeping parents and teachers in the loop of their child’s learning journey.

The complete data behind this article, including the predictions written before each run, the before and after snapshots, the full logs, and the defect record from the previous run, is published alongside it in the na8ve research archive.

David Laurenvil is the developer of na8ve.ai. He previously served as Director of Education at the Fleet Science Center in San Diego, CA, and Executive Director of Kids MakeIt Institute, a 21st-century educational institution focused on exposing students to Science, Technology, Engineering, and Math (STEM) skills and careers.