Back to changelog

Quality gate: regression evals, judge calibration, and human review

The content quality gate now has a labeled regression suite, the grading judge is calibrated against human-labeled answers, and a human review console audits both quarantined content and a rolling sample of what the automated judge approved.

Wrong lessons are the one existential risk for a learning product - the gate that prevents them now has its own safety net.

  • A labeled regression suite runs every content format through the real quality gate: structurally broken lessons, thin quizzes, duplicate flashcards, and walkthroughs without checks must all be caught before any prompt or gate change ships.
  • Two poisoned-but-well-formed lessons - one with unsafe advice, one with confidently invented facts - verify that the AI judge itself catches what structure cannot. A bad artefact clearing the gate is the one metric that may never regress.
  • The short-answer grading judge is now calibrated against human-labeled answers - paraphrases that must pass, plausible wrong answers that must not. Calibration already caught and fixed a real flaw: the judge accepted keyword-stuffing answers that name every term without committing to one.
  • A new admin review console lets a human spot-check both quarantined content and a rolling sample of what the judge approved. Rejecting a lesson that is currently being served pulls it from learners immediately, and every verdict records who decided and why.
Continue learning