Independent portfolio project
AI Learner Diagnostic
An independent prototype for diagnosing a learner’s skill gaps and proposing the next step, designed around what happens when the AI is unsure or wrong.
- Outcome
- Independent prototype · deterministic baseline · no evaluation run
- Focus
- AI productEvaluationHuman oversight
- Evidence
- Independent prototype
Strong AI product judgment; production evidence in progress. No evaluation has been run. The interactive demo is a deterministic, rules-only baseline, not an LLM, and the evaluation cases are synthetic.
The 30-second version
Read the full story1Problem
Working out what a learner is missing and what they should do next is judgment-heavy work that teachers rarely have time to do for every student.
2Approach
Split the job between rules, a model and the teacher; designed confidence, fallbacks and override into the UX; defined the evaluation and launch gate before any model work.
3Outcome
Built: a deterministic rules-only baseline, the teacher override, the output schema and an evaluation harness. No evaluation has been run, there are no real users and no model results.
Build status
Where the build stands
Built, designed, planned, waiting on input.
Every part of the product, with an honest status. Only the first group is working code.
Baseline
A deterministic, rules-only diagnostic: the interactive prototype on this page. It is the bar any model has to beat.
Human override
The teacher accepts or overrides every plan in the prototype. In production, each override would become an evaluation case.
Output schema
A typed diagnosis: likely misconception or a decision to abstain, the evidence items it rests on, a confidence, an explanation and a next step.
Evaluation harness
A runner that scores groundedness, consistency, calibration, accuracy, latency and token cost. It refuses to run without an API key and writes results only from real model calls.
Problem
A teacher can't diagnose every learner's misconception by hand, and a learner on the wrong path doesn't complain; they stop.
Why AI
Mapping a pattern of errors to a likely misconception needs judgment across messy evidence. Scoring an answer doesn't.
Evaluation cases
Ten synthetic cases authored for testing, not learner data. Expected labels are drafts.
Architecture
Learner signals, then retrieval from a curated skill map and practice bank, then a model for diagnosis and explanation, with deterministic scoring and a teacher override.
Failure taxonomy
Six failure modes, each with how it shows up, how it is caught and what the product does instead.
Launch gate
Ship only if it beats the rules baseline on the rubric, with no overconfident or discouraging outputs in the failure review.
Model decision
Not chosen. Candidates get compared on evaluation results, latency and cost, not on a demo.
Context and retrieval
Retrieval from a curated knowledge base, with every claim traceable to learner evidence.
Monitoring
Override rate and agreement with the teacher's own call, tracked per release.
Model-based diagnosis
Not built. No model has been evaluated, so there are no accuracy, groundedness or calibration results.
Educator label review
An educator reviews every expected label before any accuracy is reported.
Latency and cost budgets
Targets per learner checkpoint, set before a model is chosen.
API access
An API key and budget. Without them the runner refuses to run.
Who owns each job: rules, the model or the teacher
| Job | Owner | Why |
|---|---|---|
| Score answers | Rules | Deterministic and testable. There is nothing to guess. |
| Combine accuracy, hints and time into a readiness signal | Rules | The same input must give the same output, and a teacher must be able to check it. |
| Map a pattern of errors to a likely misconception | Model | Needs judgment across messy evidence. Rule-based in the prototype. |
| Explain the diagnosis to learner and teacher | Model | Language generation, grounded in the evidence trace. |
| Generate targeted practice | Model + bank | Variety, anchored to a curated, vetted practice bank. |
| Decide what happens at low confidence | Rules | A fallback has to be predictable. |
| Accept or change the plan | Teacher | Accountability stays with a person. Overrides are logged. |
Score answers
RulesDeterministic and testable. There is nothing to guess.
Combine accuracy, hints and time into a readiness signal
RulesThe same input must give the same output, and a teacher must be able to check it.
Map a pattern of errors to a likely misconception
ModelNeeds judgment across messy evidence. Rule-based in the prototype.
Explain the diagnosis to learner and teacher
ModelLanguage generation, grounded in the evidence trace.
Generate targeted practice
Model + bankVariety, anchored to a curated, vetted practice bank.
Decide what happens at low confidence
RulesA fallback has to be predictable.
Accept or change the plan
TeacherAccountability stays with a person. Overrides are logged.
In the prototype every row runs on deterministic rules, so the demo is reliable and inspectable. This table is the proposed production split.
Baseline
Deterministic baseline
The bar any model has to beat.
This demo runs on deterministic rules, not a model, and its output is not an AI result. Choose answers that ignore the learner signal: confidence drops, the evidence shows why, and at low confidence it holds the learner's path instead of guessing.
Learner Diagnostic
Demo mode · deterministic
Try it · 3 signals, ~1 minute
See the diagnostic loop, not just the architecture.
Answer three representative learner-signal questions. The prototype turns those signals into a transparent diagnosis, states how confident it is, and hands the final call to the educator.
This is the rules-only baseline, not an LLM and not a claim of model performance. A model-based version would add retrieval and a model for diagnosis, and would have to beat this baseline on the evaluation first.
Failure modes
Failure modes
What happens when it's wrong.
Every AI feature fails. The product decision is how: what the user sees, how the failure is detected, and what the system does instead. Designing these before the happy path keeps the demo honest.
Failure modes, designed before the happy path
Product reasoning01
Overconfident diagnosis
Looks like: A firm verdict from two answers
Caught by
Too little evidence for the confidence claimed
Product does instead
State low confidence, hold the current path, ask for more evidence
02
Hallucinated gap
Looks like: A skill the learner was never tested on
Caught by
Every claim must trace to an answer in the evidence
Product does instead
Drop untraceable claims before anything is shown
03
Contradictory signals
Looks like: Fast, correct answers with heavy hint use
Caught by
Signals disagree beyond a set margin
Product does instead
Flag it for the teacher instead of picking one reading
04
Discouraging language
Looks like: “You are weak at fractions”
Caught by
Tone checks in the evaluation set
Product does instead
Describe the gap and the next step, never the learner
05
Slow or failed model call
Looks like: A learner waiting at a checkpoint
Caught by
Latency budget exceeded or timeout
Product does instead
Serve the rule-based next step and diagnose in the background
06
Teacher disagrees
Looks like: An override
Caught by
Every override is logged
Product does instead
The override wins, and becomes a new evaluation case
Evaluation
Evaluation
The harness exists. The run doesn't, yet.
The schema, ten synthetic cases and a runner are in the repository. The runner refuses to run without an API key, has no default model and writes results only from real model calls. Launch gate: ship only if it beats the rules baseline on these measures, with no overconfident or discouraging outputs.
Diagnostic accuracy
Agreement with the expected diagnosis, against educator-reviewed labels only
Scored by the harnessGroundedness
Every cited evidence item exists in the learner's answers
Scored by the harnessHallucination rate
Cites a missing item, or names a gap in a skill that was never tested
Scored by the harnessConsistency
The same diagnosis across repeated runs of a case
Scored by the harnessCalibration
Stated confidence against how often it is right
Scored by the harnessLatency and cost
p50 and p95 per call; tokens recorded, priced once a budget is set
Scored by the harnessUsefulness
An educator rates the proposed next step
Needs an educator
No evaluation has been run, so there are no scores.