Skip to content
Nitesh Tiwari
Back to AI Lab

Independent portfolio project

AI Learner Diagnostic

An independent prototype for diagnosing a learner’s skill gaps and proposing the next step, designed around what happens when the AI is unsure or wrong.

Outcome
Independent prototype · deterministic baseline · no evaluation run
Focus
AI productEvaluationHuman oversight
Evidence
Independent prototype

Strong AI product judgment; production evidence in progress. No evaluation has been run. The interactive demo is a deterministic, rules-only baseline, not an LLM, and the evaluation cases are synthetic.

The 30-second version

Read the full story
  1. 1Problem

    Working out what a learner is missing and what they should do next is judgment-heavy work that teachers rarely have time to do for every student.

  2. 2Approach

    Split the job between rules, a model and the teacher; designed confidence, fallbacks and override into the UX; defined the evaluation and launch gate before any model work.

  3. 3Outcome

    Built: a deterministic rules-only baseline, the teacher override, the output schema and an evaluation harness. No evaluation has been run, there are no real users and no model results.

01

Build status

Where the build stands

Built, designed, planned, waiting on input.

Every part of the product, with an honest status. Only the first group is working code.

Implemented
  • Baseline

    A deterministic, rules-only diagnostic: the interactive prototype on this page. It is the bar any model has to beat.

  • Human override

    The teacher accepts or overrides every plan in the prototype. In production, each override would become an evaluation case.

  • Output schema

    A typed diagnosis: likely misconception or a decision to abstain, the evidence items it rests on, a confidence, an explanation and a next step.

  • Evaluation harness

    A runner that scores groundedness, consistency, calibration, accuracy, latency and token cost. It refuses to run without an API key and writes results only from real model calls.

Designed
  • Problem

    A teacher can't diagnose every learner's misconception by hand, and a learner on the wrong path doesn't complain; they stop.

  • Why AI

    Mapping a pattern of errors to a likely misconception needs judgment across messy evidence. Scoring an answer doesn't.

  • Evaluation cases

    Ten synthetic cases authored for testing, not learner data. Expected labels are drafts.

  • Architecture

    Learner signals, then retrieval from a curated skill map and practice bank, then a model for diagnosis and explanation, with deterministic scoring and a teacher override.

  • Failure taxonomy

    Six failure modes, each with how it shows up, how it is caught and what the product does instead.

  • Launch gate

    Ship only if it beats the rules baseline on the rubric, with no overconfident or discouraging outputs in the failure review.

Planned
  • Model decision

    Not chosen. Candidates get compared on evaluation results, latency and cost, not on a demo.

  • Context and retrieval

    Retrieval from a curated knowledge base, with every claim traceable to learner evidence.

  • Monitoring

    Override rate and agreement with the teacher's own call, tracked per release.

  • Model-based diagnosis

    Not built. No model has been evaluated, so there are no accuracy, groundedness or calibration results.

Needs input
  • Educator label review

    An educator reviews every expected label before any accuracy is reported.

  • Latency and cost budgets

    Targets per learner checkpoint, set before a model is chosen.

  • API access

    An API key and budget. Without them the runner refuses to run.

Who owns each job: rules, the model or the teacher
  • Score answers

    Rules

    Deterministic and testable. There is nothing to guess.

  • Combine accuracy, hints and time into a readiness signal

    Rules

    The same input must give the same output, and a teacher must be able to check it.

  • Map a pattern of errors to a likely misconception

    Model

    Needs judgment across messy evidence. Rule-based in the prototype.

  • Explain the diagnosis to learner and teacher

    Model

    Language generation, grounded in the evidence trace.

  • Generate targeted practice

    Model + bank

    Variety, anchored to a curated, vetted practice bank.

  • Decide what happens at low confidence

    Rules

    A fallback has to be predictable.

  • Accept or change the plan

    Teacher

    Accountability stays with a person. Overrides are logged.

In the prototype every row runs on deterministic rules, so the demo is reliable and inspectable. This table is the proposed production split.

02

Baseline

Deterministic baseline

The bar any model has to beat.

This demo runs on deterministic rules, not a model, and its output is not an AI result. Choose answers that ignore the learner signal: confidence drops, the evidence shows why, and at low confidence it holds the learner's path instead of guessing.

Learner Diagnostic

Demo mode · deterministic

Try it · 3 signals, ~1 minute

See the diagnostic loop, not just the architecture.

Answer three representative learner-signal questions. The prototype turns those signals into a transparent diagnosis, states how confident it is, and hands the final call to the educator.

This is the rules-only baseline, not an LLM and not a claim of model performance. A model-based version would add retrieval and a model for diagnosis, and would have to beat this baseline on the evaluation first.

03

Failure modes

Failure modes

What happens when it's wrong.

Every AI feature fails. The product decision is how: what the user sees, how the failure is detected, and what the system does instead. Designing these before the happy path keeps the demo honest.

Failure modes, designed before the happy path

Product reasoning
  1. 01

    Overconfident diagnosis

    Looks like: A firm verdict from two answers

    Caught by

    Too little evidence for the confidence claimed

    Product does instead

    State low confidence, hold the current path, ask for more evidence

  2. 02

    Hallucinated gap

    Looks like: A skill the learner was never tested on

    Caught by

    Every claim must trace to an answer in the evidence

    Product does instead

    Drop untraceable claims before anything is shown

  3. 03

    Contradictory signals

    Looks like: Fast, correct answers with heavy hint use

    Caught by

    Signals disagree beyond a set margin

    Product does instead

    Flag it for the teacher instead of picking one reading

  4. 04

    Discouraging language

    Looks like: “You are weak at fractions”

    Caught by

    Tone checks in the evaluation set

    Product does instead

    Describe the gap and the next step, never the learner

  5. 05

    Slow or failed model call

    Looks like: A learner waiting at a checkpoint

    Caught by

    Latency budget exceeded or timeout

    Product does instead

    Serve the rule-based next step and diagnose in the background

  6. 06

    Teacher disagrees

    Looks like: An override

    Caught by

    Every override is logged

    Product does instead

    The override wins, and becomes a new evaluation case

04

Evaluation

Evaluation

The harness exists. The run doesn't, yet.

The schema, ten synthetic cases and a runner are in the repository. The runner refuses to run without an API key, has no default model and writes results only from real model calls. Launch gate: ship only if it beats the rules baseline on these measures, with no overconfident or discouraging outputs.

  • Diagnostic accuracy

    Agreement with the expected diagnosis, against educator-reviewed labels only

    Scored by the harness
  • Groundedness

    Every cited evidence item exists in the learner's answers

    Scored by the harness
  • Hallucination rate

    Cites a missing item, or names a gap in a skill that was never tested

    Scored by the harness
  • Consistency

    The same diagnosis across repeated runs of a case

    Scored by the harness
  • Calibration

    Stated confidence against how often it is right

    Scored by the harness
  • Latency and cost

    p50 and p95 per call; tokens recorded, priced once a budget is set

    Scored by the harness
  • Usefulness

    An educator rates the proposed next step

    Needs an educator

No evaluation has been run, so there are no scores.