Tier 1 of 4

Prototypes & landing pages

A light setup for work that may be thrown away. You get one file of project rules, a few basic checks, and a place to note decisions.

1 personWeeksSomeone will ask: Nothing+ Evals: A model’s output reaches a user
Give this to your agent
Set this repo up to ntent tier 1.

Fetch https://ntent.app/r/tier/1.json?with=ai and follow the plan in it exactly. Write every file verbatim, and check each file's sha256 against its step before you write it. It also carries steps that follow from what the product is: a model’s output reaches a user. Those are as required as the tier's own.

Skip any step whose "needs" this repo does not satisfy, and do not install a framework just to satisfy one. Run each step's verify line before you call it done, write the manifest the plan describes, then tell me what you installed and what you skipped.

Reads https://ntent.app/r/tier/1.json?with=ai

Install plan10 steps

Follow the steps in order. Then 4 more that follow from what you said the product is. Each step ends with a quick check the agent must run before it says the step worked. At the end, it writes a short record of what it added, which version it came from, and what it skipped.

Tier 1

Prototypes & landing pages

6 steps · 01–06
  1. 01
    Shared project instructionsA person checks this

    Everyone inventing their own conventions, and every assistant inventing a different set again.

    writes AGENTS.md · point CLAUDE.md at it with a one-line @AGENTS.md import

    Then edit It is a template. Replace every <angle bracket>, delete the sections marked for tiers above yours, and cut any row of the enforcement table whose command this repo does not have.

    Verify Ask an assistant "what are the rules in this repo" and it answers from the file.

  2. 02

    2 rule files that disagree, because one tool reads CLAUDE.md and another reads AGENTS.md.

    Skipped unless the repo has AGENTS.md.

    Verify Run it twice: the second run says there is nothing to do. Then start a session and ask the assistant to name a rule that only exists in AGENTS.md. In Claude Code, /context lists CLAUDE.md under Memory files. Copy a line of AGENTS.md into CLAUDE.md and run it again: it names that line and leaves it where it is.

  3. 03
    Pre-commit checksBlocks the commit

    A mistake caught an hour later on the build server, after someone has already spent time reviewing it. This catches it on your machine, before the commit.

    writes scripts/git-hooks/pre-commit · make it executable

    Verify Stage a file that fails a check this tier installs and the commit is refused, naming the check and the fix. At tier 1 that is check-env: add a variable to the env schema and not to .env.example. The hardcoded-colour case needs the tier 2 lint rules.

  4. 04

    Store hook setup in git so every contributor and coding agent runs the same checks.

    merges into package.json · inside the "scripts" block

    Verify Run pnpm install, then `git config core.hooksPath` prints scripts/git-hooks. A fresh clone gets the same answer with no manual step. Leave .git/hooks alone: it may hold hooks somebody else installed.

  5. 05

    A setting the app needs that is missing from .env.example, so someone who downloads the repo cannot start it and has no idea why. Also the reverse, and code that reads process.env directly instead of through the schema.

    writes scripts/check-env.mjs

    Then edit Point CONFIG.schemaFile at this repo’s env schema, CONFIG.exampleFile at its example file, and CONFIG.sourceDirs at the directories it keeps source in, if they are not src/env.ts, .env.example and src. The pre-commit hook reads those three back out of the file, so moving them keeps the gate wired.

    Verify Add a variable to the schema, do not add it to .env.example, and the check names it.

  6. 06
    Decision logA person checks this

    The same decision being re-made in month 6 because nobody remembers the reason for the first one.

    writes DECISIONS.md

    Then edit The first entry is an example from a billing app: delete it. Then write the decisions this project has actually made, one entry each, and label any whose source is memory rather than a thread or a call as "(from memory)". A fresh project may have none yet, and a file with only the template entry in it is correct.

    Verify Every entry names a real decision, who agreed it and when. Nothing in the file is invented, and an entry reconstructed from memory says so.

Because

A model’s output reaches a user

4 steps · 07–10
  1. 07

    The pass mark for AI answers copied into 3 files, so raising it means hunting for every copy instead of changing one line.

    writes evals/types.mjs

    Verify Every threshold the suite uses is in this one file, and changing your bar is a reviewable diff. It also says which failures the pass rate is never allowed to absorb: critical cases, deterministic checks, and results the harness could not measure.

  2. 08

    One vague grading prompt that scores 5 different things at once, and tells you nothing useful about any of them.

    writes evals/judge.mjs

    Verify Each category reports its own score, and the summary breaks down by segment.

  3. 09

    A single pass rate that says something broke but not what. A grading result with no scores in it counting as a pass. One answer in 20 saying something it must never say, while the report still shows 95% and PASS.

    writes evals/report.mjs

    Skipped unless the repo has evals/types.mjs.

    Verify The report breaks down by segment and by category. Feed it a result whose gating dimensions are missing and the question fails rather than passes. Feed it nineteen passes and one case marked critical that failed, and the suite fails at 95%.

  4. 10

    An AI feature that got worse and nobody noticed, because the only test was somebody trying it once.

    writes evals/eval.test.mjs

    Skipped unless the repo has an AI feature that reaches users, evals/judge.mjs, evals/report.mjs, evals/types.mjs.

    Verify node --test evals/eval.test.mjs passes. That is the harness proving ITSELF, on stubbed judge responses: it is what makes the scoring trustworthy, not a measurement of your feature. Running your own questions is the integration step, and it is yours to write: give judge() a judgeFn that calls your provider, feed it real responses with their retrieved sources, and put the summary somewhere you keep.