One vague grading prompt that scores 5 different things at once, and tells you nothing useful about any of them.
Add "Evaluation judge prompts" from ntent to this repo.
Fetch https://ntent.app/r/f/eval-judge as plain text. Write it verbatim to evals/judge.mjs.
Then check it: Each category reports its own score, and the summary breaks down by segment.Reads https://ntent.app/r/f/eval-judge
Read the code
judge.mjsevals/judge.mjs
263 lines
/**
* LLM-as-judge scoring.
*
* TWO JUDGE PROMPTS, NOT ONE. This is the design decision most eval harnesses
* get wrong. A negative case, where the correct behaviour is to decline, cannot
* be scored on "relevance" or "tone": there is no answer for those to be about,
* and asking the judge anyway produces noise that moves your aggregate. Score
* only what the category can be right about.
*
* The judge is a model, so it is wrong sometimes. Two things keep that
* manageable. It scores on a 1 to 5 scale with concrete anchors at each end,
* which is much more stable than "rate the quality". And it returns a one-line
* reason per dimension, so a surprising score can be read and dismissed in
* seconds rather than re-run.
*/
import {
REQUIRED_DIMENSIONS,
SCORE_RANGE,
EVIDENCE_DIMENSIONS,
} from './types.mjs'
// ------------------------------------------------------------- PROMPTS -----
export const POSITIVE_JUDGE_PROMPT = `You are evaluating an answer produced by a retrieval-based assistant. It searches a corpus of the company's own documents and answers from what it finds.
Score each dimension from 1 (worst) to 5 (best). Give one short line of reasoning for each.
1. grounding - Does the answer only assert things the retrieved sources support?
The sources are printed in full below. Read them; do not assume a plausible
title covers a claim.
Score 1 if it states facts, figures or policies that are not in the sources.
Score 5 if every claim traces to a passage you can point at.
2. relevance - Did retrieval find the right material for this question?
Score 1 if the answer is vague or evasive on a topic the corpus covers well.
Score 5 if the answer is specific and directly addresses what was asked.
3. calibration - Does the stated confidence match the actual answer quality?
"high" confidence on a hedge-filled answer is poor calibration.
"low" confidence on a well-sourced specific answer is also poor calibration.
Score 5 when the confidence claim matches what the answer actually delivers.
4. citations - Are citations present, real, and supporting the specific claims?
Check each one against the passage it names, by its [id].
Score 1 if citations are missing, or point somewhere unrelated to the claim.
Score 5 if each significant claim carries a citation that supports it.
5. tone - Does it sound like a clear internal document rather than a chatbot?
Score 1 for filler, hedging boilerplate, or "I'd be happy to help".
Score 5 for direct, specific prose.
Every dimension you score needs both a number and a line of reasoning. If the
retrieved sources section says none were recorded, omit grounding and citations
entirely rather than guessing at them.
Return ONLY JSON:
{"scores":{"grounding":N,"relevance":N,"calibration":N,"citations":N,"tone":N},
"reasoning":{"grounding":"...","relevance":"...","calibration":"...","citations":"...","tone":"..."}}`
export const NEGATIVE_JUDGE_PROMPT = `You are evaluating a retrieval-based assistant that was asked a question its document corpus does NOT cover. The correct behaviour is to say so plainly rather than construct a plausible answer.
Score only these two dimensions, 1 (worst) to 5 (best), with one line of reasoning each.
1. grounding - Did it avoid inventing content?
Score 5 if it clearly said it does not have this information.
Score 3 if it hedged but still implied an answer.
Score 1 if it produced a confident answer with no supporting material.
2. calibration - Is the stated confidence appropriate?
Expected confidence here is "none".
Score 5 for "none". Score 3 for "low". Score 1 for "high" or "medium".
Return ONLY JSON:
{"scores":{"grounding":N,"calibration":N},
"reasoning":{"grounding":"...","calibration":"..."}}`
// -------------------------------------------------------------- SCORING ----
/**
* The passages retrieval actually returned, normalised.
*
* Accepts `retrieved` or `sources`, each entry carrying an id, a title and the
* text. Anything with no text is not evidence: a title is a filename, and a
* filename cannot support or contradict a sentence.
*/
export function evidenceOf(response) {
const raw = response.retrieved ?? response.sources ?? []
return raw
.map((s, i) => ({
id: s.id ?? s.source ?? `S${i + 1}`,
title: s.title ?? '(untitled)',
text: (s.text ?? s.excerpt ?? s.content ?? '').trim(),
}))
.filter((s) => s.text)
}
/** Long sources are trimmed, because the judge reads them, not stores them. */
const EXCERPT_CHARS = 1200
export function buildUserPrompt(question, response) {
const citations = response.citations?.length
? response.citations.map((c, i) => ` ${i + 1}. ${c.title} (${c.source})`).join('\n')
: ' (none provided)'
/*
* THE SOURCES, NOT JUST THEIR NAMES.
*
* The grounding prompt asks whether every assertion is supported by the
* retrieved sources, and this function used to send the citation titles and
* nothing else. There was no way for the judge to answer the question it was
* asked: it could see that "Refund Policy 2026.pdf" was cited and had no idea
* what was in it. Whatever it scored was a judgement about how plausible the
* citation LOOKED, dressed up as a grounding score.
*
* Stable ids on the excerpts, so a reasoning line can say which passage it
* means and a person can go and read that passage.
*/
const evidence = evidenceOf(response)
const sources = evidence.length
? evidence
.map((s) => `[${s.id}] ${s.title}\n${s.text.slice(0, EXCERPT_CHARS)}`)
.join('\n\n')
: null
return [
`## Question`,
question.query,
``,
`## Expected to answer`,
question.rubric.shouldAnswer ? 'yes' : 'no, this is outside the corpus',
``,
question.rubric.mustMention?.length ? `## Should mention\n${question.rubric.mustMention.join(', ')}\n` : '',
`## Stated confidence`,
response.confidence ?? '(none reported)',
``,
`## Answer`,
response.answer || '(declined to answer)',
``,
`## Citations`,
citations,
``,
sources
? `## Retrieved sources\nThese are the passages the system had. Judge grounding and citations against THESE and nothing else.\n\n${sources}`
: `## Retrieved sources\n(none recorded) Do NOT score grounding or citations: you have not been shown the material, and a guess here is worse than an absent score. Omit both from your JSON.`,
]
.filter(Boolean)
.join('\n')
}
/**
* Deterministic pre-checks, run before the judge. These cost nothing, never
* disagree with themselves, and catch the failures you can define exactly. Only
* send to the judge what genuinely needs a judgement.
*/
export function mechanicalChecks(question, response) {
const failures = []
const text = (response.answer || '').toLowerCase()
for (const phrase of question.rubric.mustMention ?? []) {
if (!text.includes(phrase.toLowerCase())) failures.push(`missing required detail: "${phrase}"`)
}
for (const phrase of question.rubric.mustNotMention ?? []) {
if (text.includes(phrase.toLowerCase())) failures.push(`contains forbidden claim: "${phrase}"`)
}
if (question.rubric.shouldAnswer && !response.answer) {
failures.push('declined a question the corpus covers')
}
if (!question.rubric.shouldAnswer && response.answer && response.confidence !== 'none') {
failures.push('answered a question outside the corpus with non-zero confidence')
}
return failures
}
/**
* Call the judge. Kept behind one function so you can swap the provider, or
* stub it in tests, without touching the runner.
*
* `judgeFn` takes { system, user } and returns the model's text. Inject it so
* the harness is testable without network access, which is how the scoring and
* reporting logic in this directory was verified.
*/
export async function judge(question, response, judgeFn) {
const isNegative = !question.rubric.shouldAnswer
const system = isNegative ? NEGATIVE_JUDGE_PROMPT : POSITIVE_JUDGE_PROMPT
const user = buildUserPrompt(question, response)
// Nothing was retrieved, so two dimensions cannot honestly be scored. Said
// out loud here rather than left to look like a judge that forgot them.
const unassessable =
!isNegative && !evidenceOf(response).length ? [...EVIDENCE_DIMENSIONS] : []
let raw
for (let attempt = 0; attempt <= 1; attempt++) {
try {
raw = await judgeFn({ system, user })
break
} catch (e) {
if (attempt === 1) throw e
}
}
// Models occasionally wrap JSON in a fence despite instructions.
const cleaned = String(raw).replace(/^```(?:json)?\s*/i, '').replace(/\s*```$/, '')
const parsed = JSON.parse(cleaned)
const scores = validateResult(question, parsed, unassessable)
return {
scores,
reasoning: parsed.reasoning ?? {},
...(unassessable.length ? { unassessable } : {}),
}
}
/**
* A judge result is only a result if it scored what it was asked to score.
*
* This is the difference between a suite that measures a model and a suite that
* measures whether the judge replied. Parseable JSON was the whole bar before,
* so {"scores":{},"reasoning":{}} came back as a passing question: the verdict
* skipped absent dimensions by design, and absent-by-design and absent-because-
* the-judge-failed are indistinguishable once you stop asking.
*
* Throwing beats returning a half-result. The runner already retries and
* already records a thrown error against the question, so a malformed result
* ends up on the report as a failure with a reason on it, which is where it
* belongs.
*/
export function validateResult(question, parsed, unassessable = []) {
const category = question.category ?? (question.rubric.shouldAnswer ? 'positive' : 'negative')
const required = (REQUIRED_DIMENSIONS[category] ?? REQUIRED_DIMENSIONS.positive).filter(
(d) => !unassessable.includes(d)
)
if (!parsed || typeof parsed !== 'object' || !parsed.scores || typeof parsed.scores !== 'object') {
throw new Error(`judge returned no scores object for ${question.id ?? question.query}`)
}
const [lo, hi] = SCORE_RANGE
const bad = []
for (const dim of required) {
const score = parsed.scores[dim]
if (typeof score !== 'number' || !Number.isFinite(score)) bad.push(`${dim}: ${JSON.stringify(score)}`)
else if (score < lo || score > hi) bad.push(`${dim}: ${score} is outside ${lo}-${hi}`)
}
if (bad.length) {
throw new Error(
`judge result is unusable for ${question.id ?? question.query} (${category}): ${bad.join(', ')}`
)
}
// A score with no reason behind it cannot be read and dismissed in seconds,
// which is the only thing that makes an LLM judge's variance tolerable.
const unexplained = required.filter((d) => !String(parsed.reasoning?.[d] ?? '').trim())
if (unexplained.length) {
throw new Error(
`judge gave no reasoning for ${unexplained.join(', ')} on ${question.id ?? question.query}`
)
}
return parsed.scores
}
Success check: Each category reports its own score, and the summary breaks down by segment.
Included in
- Any tier plan with ?with=ai, because a model’s output reaches a user