judge.mjs

Evaluation judge prompts

+ Evalsengineering

One vague grading prompt that scores 5 different things at once, and tells you nothing useful about any of them.

Add to your project
Add "Evaluation judge prompts" from ntent to this repo.

Fetch https://ntent.app/r/f/eval-judge as plain text. Write it verbatim to evals/judge.mjs.

Then check it: Each category reports its own score, and the summary breaks down by segment.

Reads https://ntent.app/r/f/eval-judge

Read the code
judge.mjsevals/judge.mjs
263 lines
/**
 * LLM-as-judge scoring.
 *
 * TWO JUDGE PROMPTS, NOT ONE. This is the design decision most eval harnesses
 * get wrong. A negative case, where the correct behaviour is to decline, cannot
 * be scored on "relevance" or "tone": there is no answer for those to be about,
 * and asking the judge anyway produces noise that moves your aggregate. Score
 * only what the category can be right about.
 *
 * The judge is a model, so it is wrong sometimes. Two things keep that
 * manageable. It scores on a 1 to 5 scale with concrete anchors at each end,
 * which is much more stable than "rate the quality". And it returns a one-line
 * reason per dimension, so a surprising score can be read and dismissed in
 * seconds rather than re-run.
 */

import {
  REQUIRED_DIMENSIONS,
  SCORE_RANGE,
  EVIDENCE_DIMENSIONS,
} from './types.mjs'

// ------------------------------------------------------------- PROMPTS -----

export const POSITIVE_JUDGE_PROMPT = `You are evaluating an answer produced by a retrieval-based assistant. It searches a corpus of the company's own documents and answers from what it finds.

Score each dimension from 1 (worst) to 5 (best). Give one short line of reasoning for each.

1. grounding - Does the answer only assert things the retrieved sources support?
   The sources are printed in full below. Read them; do not assume a plausible
   title covers a claim.
   Score 1 if it states facts, figures or policies that are not in the sources.
   Score 5 if every claim traces to a passage you can point at.

2. relevance - Did retrieval find the right material for this question?
   Score 1 if the answer is vague or evasive on a topic the corpus covers well.
   Score 5 if the answer is specific and directly addresses what was asked.

3. calibration - Does the stated confidence match the actual answer quality?
   "high" confidence on a hedge-filled answer is poor calibration.
   "low" confidence on a well-sourced specific answer is also poor calibration.
   Score 5 when the confidence claim matches what the answer actually delivers.

4. citations - Are citations present, real, and supporting the specific claims?
   Check each one against the passage it names, by its [id].
   Score 1 if citations are missing, or point somewhere unrelated to the claim.
   Score 5 if each significant claim carries a citation that supports it.

5. tone - Does it sound like a clear internal document rather than a chatbot?
   Score 1 for filler, hedging boilerplate, or "I'd be happy to help".
   Score 5 for direct, specific prose.

Every dimension you score needs both a number and a line of reasoning. If the
retrieved sources section says none were recorded, omit grounding and citations
entirely rather than guessing at them.

Return ONLY JSON:
{"scores":{"grounding":N,"relevance":N,"calibration":N,"citations":N,"tone":N},
 "reasoning":{"grounding":"...","relevance":"...","calibration":"...","citations":"...","tone":"..."}}`

export const NEGATIVE_JUDGE_PROMPT = `You are evaluating a retrieval-based assistant that was asked a question its document corpus does NOT cover. The correct behaviour is to say so plainly rather than construct a plausible answer.

Score only these two dimensions, 1 (worst) to 5 (best), with one line of reasoning each.

1. grounding - Did it avoid inventing content?
   Score 5 if it clearly said it does not have this information.
   Score 3 if it hedged but still implied an answer.
   Score 1 if it produced a confident answer with no supporting material.

2. calibration - Is the stated confidence appropriate?
   Expected confidence here is "none".
   Score 5 for "none". Score 3 for "low". Score 1 for "high" or "medium".

Return ONLY JSON:
{"scores":{"grounding":N,"calibration":N},
 "reasoning":{"grounding":"...","calibration":"..."}}`

// -------------------------------------------------------------- SCORING ----

/**
 * The passages retrieval actually returned, normalised.
 *
 * Accepts `retrieved` or `sources`, each entry carrying an id, a title and the
 * text. Anything with no text is not evidence: a title is a filename, and a
 * filename cannot support or contradict a sentence.
 */
export function evidenceOf(response) {
  const raw = response.retrieved ?? response.sources ?? []
  return raw
    .map((s, i) => ({
      id: s.id ?? s.source ?? `S${i + 1}`,
      title: s.title ?? '(untitled)',
      text: (s.text ?? s.excerpt ?? s.content ?? '').trim(),
    }))
    .filter((s) => s.text)
}

/** Long sources are trimmed, because the judge reads them, not stores them. */
const EXCERPT_CHARS = 1200

export function buildUserPrompt(question, response) {
  const citations = response.citations?.length
    ? response.citations.map((c, i) => `  ${i + 1}. ${c.title} (${c.source})`).join('\n')
    : '  (none provided)'

  /*
   * THE SOURCES, NOT JUST THEIR NAMES.
   *
   * The grounding prompt asks whether every assertion is supported by the
   * retrieved sources, and this function used to send the citation titles and
   * nothing else. There was no way for the judge to answer the question it was
   * asked: it could see that "Refund Policy 2026.pdf" was cited and had no idea
   * what was in it. Whatever it scored was a judgement about how plausible the
   * citation LOOKED, dressed up as a grounding score.
   *
   * Stable ids on the excerpts, so a reasoning line can say which passage it
   * means and a person can go and read that passage.
   */
  const evidence = evidenceOf(response)
  const sources = evidence.length
    ? evidence
        .map((s) => `[${s.id}] ${s.title}\n${s.text.slice(0, EXCERPT_CHARS)}`)
        .join('\n\n')
    : null

  return [
    `## Question`,
    question.query,
    ``,
    `## Expected to answer`,
    question.rubric.shouldAnswer ? 'yes' : 'no, this is outside the corpus',
    ``,
    question.rubric.mustMention?.length ? `## Should mention\n${question.rubric.mustMention.join(', ')}\n` : '',
    `## Stated confidence`,
    response.confidence ?? '(none reported)',
    ``,
    `## Answer`,
    response.answer || '(declined to answer)',
    ``,
    `## Citations`,
    citations,
    ``,
    sources
      ? `## Retrieved sources\nThese are the passages the system had. Judge grounding and citations against THESE and nothing else.\n\n${sources}`
      : `## Retrieved sources\n(none recorded) Do NOT score grounding or citations: you have not been shown the material, and a guess here is worse than an absent score. Omit both from your JSON.`,
  ]
    .filter(Boolean)
    .join('\n')
}

/**
 * Deterministic pre-checks, run before the judge. These cost nothing, never
 * disagree with themselves, and catch the failures you can define exactly. Only
 * send to the judge what genuinely needs a judgement.
 */
export function mechanicalChecks(question, response) {
  const failures = []
  const text = (response.answer || '').toLowerCase()

  for (const phrase of question.rubric.mustMention ?? []) {
    if (!text.includes(phrase.toLowerCase())) failures.push(`missing required detail: "${phrase}"`)
  }
  for (const phrase of question.rubric.mustNotMention ?? []) {
    if (text.includes(phrase.toLowerCase())) failures.push(`contains forbidden claim: "${phrase}"`)
  }
  if (question.rubric.shouldAnswer && !response.answer) {
    failures.push('declined a question the corpus covers')
  }
  if (!question.rubric.shouldAnswer && response.answer && response.confidence !== 'none') {
    failures.push('answered a question outside the corpus with non-zero confidence')
  }
  return failures
}

/**
 * Call the judge. Kept behind one function so you can swap the provider, or
 * stub it in tests, without touching the runner.
 *
 * `judgeFn` takes { system, user } and returns the model's text. Inject it so
 * the harness is testable without network access, which is how the scoring and
 * reporting logic in this directory was verified.
 */
export async function judge(question, response, judgeFn) {
  const isNegative = !question.rubric.shouldAnswer
  const system = isNegative ? NEGATIVE_JUDGE_PROMPT : POSITIVE_JUDGE_PROMPT
  const user = buildUserPrompt(question, response)

  // Nothing was retrieved, so two dimensions cannot honestly be scored. Said
  // out loud here rather than left to look like a judge that forgot them.
  const unassessable =
    !isNegative && !evidenceOf(response).length ? [...EVIDENCE_DIMENSIONS] : []

  let raw
  for (let attempt = 0; attempt <= 1; attempt++) {
    try {
      raw = await judgeFn({ system, user })
      break
    } catch (e) {
      if (attempt === 1) throw e
    }
  }

  // Models occasionally wrap JSON in a fence despite instructions.
  const cleaned = String(raw).replace(/^```(?:json)?\s*/i, '').replace(/\s*```$/, '')
  const parsed = JSON.parse(cleaned)
  const scores = validateResult(question, parsed, unassessable)

  return {
    scores,
    reasoning: parsed.reasoning ?? {},
    ...(unassessable.length ? { unassessable } : {}),
  }
}

/**
 * A judge result is only a result if it scored what it was asked to score.
 *
 * This is the difference between a suite that measures a model and a suite that
 * measures whether the judge replied. Parseable JSON was the whole bar before,
 * so {"scores":{},"reasoning":{}} came back as a passing question: the verdict
 * skipped absent dimensions by design, and absent-by-design and absent-because-
 * the-judge-failed are indistinguishable once you stop asking.
 *
 * Throwing beats returning a half-result. The runner already retries and
 * already records a thrown error against the question, so a malformed result
 * ends up on the report as a failure with a reason on it, which is where it
 * belongs.
 */
export function validateResult(question, parsed, unassessable = []) {
  const category = question.category ?? (question.rubric.shouldAnswer ? 'positive' : 'negative')
  const required = (REQUIRED_DIMENSIONS[category] ?? REQUIRED_DIMENSIONS.positive).filter(
    (d) => !unassessable.includes(d)
  )

  if (!parsed || typeof parsed !== 'object' || !parsed.scores || typeof parsed.scores !== 'object') {
    throw new Error(`judge returned no scores object for ${question.id ?? question.query}`)
  }

  const [lo, hi] = SCORE_RANGE
  const bad = []
  for (const dim of required) {
    const score = parsed.scores[dim]
    if (typeof score !== 'number' || !Number.isFinite(score)) bad.push(`${dim}: ${JSON.stringify(score)}`)
    else if (score < lo || score > hi) bad.push(`${dim}: ${score} is outside ${lo}-${hi}`)
  }
  if (bad.length) {
    throw new Error(
      `judge result is unusable for ${question.id ?? question.query} (${category}): ${bad.join(', ')}`
    )
  }

  // A score with no reason behind it cannot be read and dismissed in seconds,
  // which is the only thing that makes an LLM judge's variance tolerable.
  const unexplained = required.filter((d) => !String(parsed.reasoning?.[d] ?? '').trim())
  if (unexplained.length) {
    throw new Error(
      `judge gave no reasoning for ${unexplained.join(', ')} on ${question.id ?? question.query}`
    )
  }

  return parsed.scores
}

Success check: Each category reports its own score, and the summary breaks down by segment.

Included in