The Most Important Number in Our Build Data So Far: 55% of Kid Revisions Are Rejections
Early build data from Xyplor: in our first ~70 classified kid revisions, over half were 'no, fix this' rejections, not additions. Here's what that number does and doesn't tell us.
This is a build note, not a research paper. It's about a small, early data point from our own product telemetry — not an outcome claim, not a peer-reviewed study, and not something we'd stake a marketing headline on yet. We're publishing it because it's the kind of thing we'd want to see if we were evaluating someone else's AI-for-kids product.
TL;DR
We started tagging how kids revise their AI creations on Xyplor — specifically, whether a revision is a fix (rejecting what the AI produced and asking for something different), an add (building on what's there), or a new (starting over from scratch). In our first ~70 classified kid revisions, 55% were fixes — kids telling the AI "no, that's not right, do it differently" rather than layering on more.
That's a small sample and an early read. But it's the single number from our build data so far that we think says the most about what's actually happening when a kid uses Xyplor: they're not just accepting whatever the AI hands them. They're pushing back on it more often than not.
Why we started tracking this at all
When you watch a kid use a generative AI tool for the first time, there's a version of the experience that looks like magic and a version that looks like agency, and from the outside they can be hard to tell apart. A kid types "make a game where I fight dragons," a playable game appears in about a minute, and the kid says "cool" and moves on. Did they direct the AI, or did the AI just perform for them?
We built Xyplor around the idea that the reflection loop after a creation — showing the kid what just happened and giving them a concrete next move — is what turns "cool" into "wait, actually, make the dragons faster." (We've written about the mechanics of that loop in a previous post.) But a design principle is a hypothesis until you have data. So a few months ago we started classifying what kids actually type when they revise something they've already built.
What revisionType means, plainly
Every time a kid submits a follow-up prompt on an existing creation, we tag it into one of three buckets:
- fix — the kid is rejecting something the AI produced and asking for it to be corrected or changed. "The jump is broken," "that's not what I meant," "make the dragon faster, it's too slow right now." The AI's first attempt didn't match what the kid wanted, and the kid said so.
- add — the kid is building on top of what already exists without rejecting it. "Add a second level," "give the quiz five more questions," "put a scoreboard in the corner." The existing thing stays; something new gets layered on.
- new — the kid abandons the current direction and starts over with a different concept entirely. Not a fix, not an addition — a fresh start.
We're not claiming this taxonomy is the only reasonable way to slice revision behavior, or that our classification process is airtight. Some revisions are genuinely ambiguous — "make it better" could be read as a fix or an add depending on context — and a human or model doing the classifying has to make a judgment call. We're being explicit about that limitation because we'd rather you distrust a number we've flagged than trust one we haven't.
The number: 55% fixes, in the first ~70 classified revisions
Of the first roughly 70 kid revisions we classified, 55% were fixes. Not additions, not fresh starts — corrections. A kid looked at what the AI made, decided it wasn't right, and said so.
We think this is the most important number in our build data right now, for a specific reason: it's evidence of authorship, not just usage.
There's an easy, cynical read of any kid-facing generative AI product: the AI does the work, the kid clicks "yes," and the kid walks away having directed nothing. If that were what's happening on Xyplor, we'd expect revisions to skew heavily toward additions — kids happily accepting the first draft and just piling more content on top, the way you'd keep adding rooms to a house you never questioned the foundation of.
That's not what the early data shows. More than half the time, the kid's first move after seeing an AI-built creation is to reject part of it and ask for something different. That's a kid evaluating output critically and pushing back — which is closer to what we mean when we say we want kids to direct AI rather than simply consume what it produces.
What this number doesn't tell us
We want to be careful here, because it would be easy to overstate this.
- This is not a study. It's internal product telemetry from a small, early set of interactions, not a controlled research design. We have no comparison condition, no control group, and no claim about how this compares to revision behavior on any other platform.
- ~70 revisions is a small sample. We wouldn't trust a 55% figure from 70 data points to hold at 700 or 7,000 — and we're not asserting it will. It could move in either direction as more kids use the product and as we refine the classification process itself.
- Fix ≠ better learning outcome, measured. A high fix rate tells us kids are engaging critically with output. It doesn't, on its own, tell us anything about learning transfer, skill development, or long-term outcomes. We don't have that data, and we're not going to imply we do. [VERIFY: any future longitudinal claim about outcomes must be backed by actual measurement before publication]
- Classification involves judgment calls. As noted above, the fix/add/new boundary isn't always clean. We're working on tightening the definitions and, eventually, having multiple people classify the same revisions to check agreement — but we're not there yet with this batch.
We're publishing the number now, at N≈70, because we think transparency about early, imperfect data is more useful to parents and educators than waiting for a polished, larger dataset and presenting it as if it always looked clean. If this number changes meaningfully as we collect more revisions, we'll say so.
Why we think this matters beyond our own product
Even setting Xyplor aside, we think "fix vs. add vs. new" is a useful lens for anyone evaluating a kid-facing AI tool, whether it's ours or a competitor's.
If you're a parent or teacher watching a kid use any generative AI product, ask: when the AI gets something wrong, does the kid notice, and does the kid say so? A tool that produces impressive-looking output on the first try can accidentally train a kid toward passive acceptance — "the AI is always basically right, so why argue with it." A tool that surfaces the AI's limitations, and rewards a kid for pushing back, is doing something different: building the habit of evaluating AI output rather than deferring to it.
That's the design bet behind the reflection loop we've written about before, and it's part of why the 55% figure feels significant to us internally, even at this early sample size. It's a rough, early signal that the design is nudging kids toward the behavior we intended — critical evaluation, not passive consumption — rather than a promise that it's working at scale.
What we're doing next
A few concrete next steps on our end:
- Expanding the classified sample well past 70 revisions before drawing any firmer conclusions.
- Having more than one person (or process) classify the same revisions to check how consistent the fix/add/new labels are.
- Breaking the data down by age band, since we'd expect a 7-year-old's revision behavior to look different from a 15-year-old's, and lumping them together probably hides more than it reveals.
- Looking at how fix rate changes over a kid's first few weeks on the platform — whether kids get more or less willing to reject AI output as they get more comfortable directing it.
We'll follow up on this blog when we have a larger, more rigorously classified dataset. If the number holds up, we think it's a meaningful data point about what kid-directed AI creation actually looks like. If it doesn't hold up, we'll say that too.
If you're building in this space and tracking anything similar — revision types, correction rates, anything that gets at authorship versus passive use — we'd genuinely like to compare notes. Email us at partnerships@xyplor.com.