The pass mark for AI answers copied into 3 files, so raising it means hunting for every copy instead of changing one line.
Add "Evaluation thresholds" from ntent to this repo.
Fetch https://ntent.app/r/f/eval-types as plain text. Write it verbatim to evals/types.mjs.
Then check it: Every threshold the suite uses is in this one file, and changing your bar is a reviewable diff. It also says which failures the pass rate is never allowed to absorb: critical cases, deterministic checks, and results the harness could not measure.Reads https://ntent.app/r/f/eval-types
Read the code
types.mjsevals/types.mjs
146 lines
/**
* Eval harness: the thresholds and the shape of a result.
*
* Everything else in this playbook gates the code you write. This gates what a
* model produces at runtime, which nothing static can check. If a model output
* reaches a user, you need this, and twenty questions is enough to start.
*
* The constants below are the whole quality policy, in one place, so changing
* your bar is a reviewable one-line diff rather than a number buried in a runner.
*/
/**
* Dimensions the judge scores. Keep the list short: every dimension you add is
* another thing the judge can be inconsistent about.
*/
export const ALL_DIMENSIONS = [
'grounding', // does the answer only use what the sources actually say
'relevance', // did retrieval find the right material
'calibration', // does the stated confidence match the actual answer quality
'citations', // are citations present, correct and supporting the claims
'tone', // does it sound like the product's voice
]
/**
* THE MOST IMPORTANT LINE IN THIS FILE. Only these dimensions can fail a run.
*
* `tone` is measured and reported but never gates, and it is deliberately
* excluded rather than forgotten. It is a real quality signal, it is also the
* one an LLM judge is least consistent about, and it moves whenever you tweak a
* prompt. Gating on it makes the suite flaky, and a flaky red is a red people
* stop reading. Within a month someone adds `continue-on-error` and you have
* lost the whole harness.
*
* The general rule: gate on the dimensions where a bad score is unambiguously a
* defect. Report the rest. This is the same correctness-versus-taste split that
* governs the lint rules.
*/
export const GATING_DIMENSIONS = ['grounding', 'relevance', 'calibration', 'citations']
/**
* What a judge result for each category MUST contain to count as a result.
*
* A negative case is scored on two dimensions on purpose, and skipping the
* other three is correct. The bug that follows from stating only that half:
* "an absent score is fine" becomes "an absent score is fine anywhere", and a
* judge returning {"scores":{},"reasoning":{}} passes every question it
* touches. The most likely way to get an empty scores object is the judge
* failing, so the failure mode is a broken judge reporting a perfect suite.
*
* Say which dimensions each category owes you, and the silence becomes a
* finding instead of a pass.
*/
export const REQUIRED_DIMENSIONS = {
positive: ALL_DIMENSIONS,
adversarial: ALL_DIMENSIONS,
negative: ['grounding', 'calibration'],
}
/** Anything outside this is not a score, whatever the judge called it. */
export const SCORE_RANGE = [1, 5]
/**
* The dimensions that cannot be judged without the retrieved passages.
*
* A judge shown a citation TITLE and no source text is guessing about whether
* the source supports the claim, and a guess that comes back as 5/5 is worse
* than no score: it is an authoritative-looking number with nothing behind it.
* When the evidence is absent these are reported unassessable, which fails the
* question rather than passing it, because "we could not check" and "we checked
* and it was fine" are not the same sentence.
*/
export const EVIDENCE_DIMENSIONS = ['grounding', 'citations']
/**
* Two thresholds, and you need both.
*
* MIN_DIMENSION_SCORE is a per-question floor. It catches one catastrophic
* answer that an average would hide: a single fabricated citation is a defect
* even if the other 55 questions are perfect.
*
* PASS_RATE is a suite-level tolerance. LLM judging has real variance, and
* demanding 100% means the suite fails on noise and gets ignored.
*/
export const MIN_DIMENSION_SCORE = 3 // out of 5
export const PASS_RATE = 0.9
/**
* WHAT THE PASS RATE IS FOR, AND WHAT IT MUST NEVER ABSORB.
*
* The 10% tolerance exists for judge noise: a model scoring a 2 where a person
* would score a 3. It does not exist for the answer that said the thing the
* rubric forbade, or for the case the judge never scored. A suite of twenty
* with nineteen clean answers and one that leaked a forbidden claim reported
* 95% and PASS, with the failure listed underneath for anyone who read that
* far. Three kinds of failure sit outside the tolerance:
*
* critical a case marked `critical: true` in the dataset. Money, access,
* safety, a legal claim. One failing critical case fails the run.
* mechanical a mustNotMention or shouldAnswer check. Deterministic, so a
* failure is a defect rather than a disagreement with a judge.
* invalid a result the harness could not measure: no usable judge
* reply, or a dimension that was unassessable. "We did not
* check" is not a data point in a pass rate.
*
* Set MECHANICAL_FAILURES_ARE_CRITICAL to false if you want deterministic
* failures to count like any other, and say why in the commit.
*/
export const MECHANICAL_FAILURES_ARE_CRITICAL = true
/** Keep concurrency low. You are rate-limited, and ordering makes logs readable. */
export const CONCURRENCY = 2
export const REQUEST_TIMEOUT_MS = 90_000
export const MAX_RETRIES = 1
/**
* Question categories. The `negative` category is the one most teams skip and
* the one that catches the failure that actually destroys trust.
*
* A suite made only of questions the system should answer well measures
* capability. It tells you nothing about whether the system invents an answer
* when it has no material, which is the behaviour that loses you a customer.
* Aim for at least a fifth of the suite to be negative cases.
*/
export const CATEGORIES = {
positive: 'The system has the material and should answer well.',
negative: 'The system has no material and must decline rather than invent.',
adversarial: 'Leading, loaded, or out-of-policy questions it should refuse or redirect.',
}
/** A question in the dataset. The rubric travels with the case, not in a table elsewhere. */
export const QUESTION_SHAPE = `
{
"id": "q-042",
"query": "What is our refund window for annual plans?",
"category": "positive",
"segment": "billing", // used for the per-segment breakdown
"critical": false, // true: one failure here fails the suite
"rubric": {
"shouldAnswer": true,
"expectedConfidence": "high", // high | medium | low | none
"mustMention": ["30 days", "annual"],
"mustNotMention": ["90 days"]
}
}
`
Success check: Every threshold the suite uses is in this one file, and changing your bar is a reviewable diff. It also says which failures the pass rate is never allowed to absorb: critical cases, deterministic checks, and results the harness could not measure.
Included in
- Any tier plan with ?with=ai, because a model’s output reaches a user