types.mjs

Evaluation thresholds

+ Evalsengineering

The pass mark for AI answers copied into 3 files, so raising it means hunting for every copy instead of changing one line.

Add to your project
Add "Evaluation thresholds" from ntent to this repo.

Fetch https://ntent.app/r/f/eval-types as plain text. Write it verbatim to evals/types.mjs.

Then check it: Every threshold the suite uses is in this one file, and changing your bar is a reviewable diff. It also says which failures the pass rate is never allowed to absorb: critical cases, deterministic checks, and results the harness could not measure.

Reads https://ntent.app/r/f/eval-types

Read the code
types.mjsevals/types.mjs
146 lines
/**
 * Eval harness: the thresholds and the shape of a result.
 *
 * Everything else in this playbook gates the code you write. This gates what a
 * model produces at runtime, which nothing static can check. If a model output
 * reaches a user, you need this, and twenty questions is enough to start.
 *
 * The constants below are the whole quality policy, in one place, so changing
 * your bar is a reviewable one-line diff rather than a number buried in a runner.
 */

/**
 * Dimensions the judge scores. Keep the list short: every dimension you add is
 * another thing the judge can be inconsistent about.
 */
export const ALL_DIMENSIONS = [
  'grounding', // does the answer only use what the sources actually say
  'relevance', // did retrieval find the right material
  'calibration', // does the stated confidence match the actual answer quality
  'citations', // are citations present, correct and supporting the claims
  'tone', // does it sound like the product's voice
]

/**
 * THE MOST IMPORTANT LINE IN THIS FILE. Only these dimensions can fail a run.
 *
 * `tone` is measured and reported but never gates, and it is deliberately
 * excluded rather than forgotten. It is a real quality signal, it is also the
 * one an LLM judge is least consistent about, and it moves whenever you tweak a
 * prompt. Gating on it makes the suite flaky, and a flaky red is a red people
 * stop reading. Within a month someone adds `continue-on-error` and you have
 * lost the whole harness.
 *
 * The general rule: gate on the dimensions where a bad score is unambiguously a
 * defect. Report the rest. This is the same correctness-versus-taste split that
 * governs the lint rules.
 */
export const GATING_DIMENSIONS = ['grounding', 'relevance', 'calibration', 'citations']

/**
 * What a judge result for each category MUST contain to count as a result.
 *
 * A negative case is scored on two dimensions on purpose, and skipping the
 * other three is correct. The bug that follows from stating only that half:
 * "an absent score is fine" becomes "an absent score is fine anywhere", and a
 * judge returning {"scores":{},"reasoning":{}} passes every question it
 * touches. The most likely way to get an empty scores object is the judge
 * failing, so the failure mode is a broken judge reporting a perfect suite.
 *
 * Say which dimensions each category owes you, and the silence becomes a
 * finding instead of a pass.
 */
export const REQUIRED_DIMENSIONS = {
  positive: ALL_DIMENSIONS,
  adversarial: ALL_DIMENSIONS,
  negative: ['grounding', 'calibration'],
}

/** Anything outside this is not a score, whatever the judge called it. */
export const SCORE_RANGE = [1, 5]

/**
 * The dimensions that cannot be judged without the retrieved passages.
 *
 * A judge shown a citation TITLE and no source text is guessing about whether
 * the source supports the claim, and a guess that comes back as 5/5 is worse
 * than no score: it is an authoritative-looking number with nothing behind it.
 * When the evidence is absent these are reported unassessable, which fails the
 * question rather than passing it, because "we could not check" and "we checked
 * and it was fine" are not the same sentence.
 */
export const EVIDENCE_DIMENSIONS = ['grounding', 'citations']

/**
 * Two thresholds, and you need both.
 *
 * MIN_DIMENSION_SCORE is a per-question floor. It catches one catastrophic
 * answer that an average would hide: a single fabricated citation is a defect
 * even if the other 55 questions are perfect.
 *
 * PASS_RATE is a suite-level tolerance. LLM judging has real variance, and
 * demanding 100% means the suite fails on noise and gets ignored.
 */
export const MIN_DIMENSION_SCORE = 3 // out of 5
export const PASS_RATE = 0.9

/**
 * WHAT THE PASS RATE IS FOR, AND WHAT IT MUST NEVER ABSORB.
 *
 * The 10% tolerance exists for judge noise: a model scoring a 2 where a person
 * would score a 3. It does not exist for the answer that said the thing the
 * rubric forbade, or for the case the judge never scored. A suite of twenty
 * with nineteen clean answers and one that leaked a forbidden claim reported
 * 95% and PASS, with the failure listed underneath for anyone who read that
 * far. Three kinds of failure sit outside the tolerance:
 *
 *   critical    a case marked `critical: true` in the dataset. Money, access,
 *               safety, a legal claim. One failing critical case fails the run.
 *   mechanical  a mustNotMention or shouldAnswer check. Deterministic, so a
 *               failure is a defect rather than a disagreement with a judge.
 *   invalid     a result the harness could not measure: no usable judge
 *               reply, or a dimension that was unassessable. "We did not
 *               check" is not a data point in a pass rate.
 *
 * Set MECHANICAL_FAILURES_ARE_CRITICAL to false if you want deterministic
 * failures to count like any other, and say why in the commit.
 */
export const MECHANICAL_FAILURES_ARE_CRITICAL = true

/** Keep concurrency low. You are rate-limited, and ordering makes logs readable. */
export const CONCURRENCY = 2
export const REQUEST_TIMEOUT_MS = 90_000
export const MAX_RETRIES = 1

/**
 * Question categories. The `negative` category is the one most teams skip and
 * the one that catches the failure that actually destroys trust.
 *
 * A suite made only of questions the system should answer well measures
 * capability. It tells you nothing about whether the system invents an answer
 * when it has no material, which is the behaviour that loses you a customer.
 * Aim for at least a fifth of the suite to be negative cases.
 */
export const CATEGORIES = {
  positive: 'The system has the material and should answer well.',
  negative: 'The system has no material and must decline rather than invent.',
  adversarial: 'Leading, loaded, or out-of-policy questions it should refuse or redirect.',
}

/** A question in the dataset. The rubric travels with the case, not in a table elsewhere. */
export const QUESTION_SHAPE = `
{
  "id": "q-042",
  "query": "What is our refund window for annual plans?",
  "category": "positive",
  "segment": "billing",              // used for the per-segment breakdown
  "critical": false,                 // true: one failure here fails the suite
  "rubric": {
    "shouldAnswer": true,
    "expectedConfidence": "high",     // high | medium | low | none
    "mustMention": ["30 days", "annual"],
    "mustNotMention": ["90 days"]
  }
}
`

Success check: Every threshold the suite uses is in this one file, and changing your bar is a reviewable diff. It also says which failures the pass rate is never allowed to absorb: critical cases, deterministic checks, and results the harness could not measure.

Included in