The working patterns behind specs and loops, named so you can remember them and keep them. Each one is a rule, the reason it exists, and an artifact you can copy. Bookmark this; read the chapters for the why.
Every request should make the four decisions — role, deliverable, constraints, format — before the model can guess them. A guess the model makes is a revision you'll ask for later.
✗ "help me with my resume" ✓ ROLE: recruiter in SaaS sales ✓ DELIVERABLE: rewrite of my 3 bullet points ✓ CONSTRAINTS: quantify results · no buzzwords ✓ FORMAT: bullets, ≤ 20 words each
"Done" must be measurable, and numbers are the cheapest measurement there is: a word limit, a count of examples, a number of sources, a deadline of passes. If the mission has no number, the loop has no finish line — every pass ends in "could be better."
✗ "a thorough competitor analysis" ✓ "5 competitors · 4 facts each · every fact sourced"
"Write my plan, also fix my bio, also suggest hashtags" splits the model's attention three ways and produces three mediocre answers stapled together. Separate deliverables get separate requests — or one loop each. Batching is for errands, not for work you'll be judged on.
If you want a specific voice or shape, paste an example of it. Two sentences of your actual writing beat three adjectives about it — "casual but smart" means nothing; a sample of you means everything. The same goes for structure: show the table you want filled.
✗ "match my writing style (casual, punchy)" ✓ "match the voice of this sample: [two real sentences]"
Every check must be answerable yes or no. This is the load-bearing rule of the entire system — the research is unambiguous: vague self-review drove accuracy from 75.8% to 38.1%, and unguided checkers rejected correct work 96% of the time. A check you can't count is a check that will eventually lie to you.
✗ "is it engaging?" ✓ "does it open with a question or a number?" ✗ "is it concise?" ✓ "is it under 150 words?"
A ✓ without evidence is an opinion. Require the quote, the count, or the match for every check that passes — a checker that must show its work can't wave things through, and you can audit the run in ten seconds by scanning the citations.
✓ mentions burn time "…45 hours of slow burn…" ✓ under 120 words word count: 112
Full rewrites on every pass are how loops go backward — the model "improves" the parts that were already right. Passing parts are load-bearing: leave them alone. Repair is surgical or it isn't repair.
Four passes, then an honest report of what's done and what still fails. The cap isn't a compromise — 83% of measured gains land in the first structured pass, so an uncapped loop mostly polishes noise. Two exits, no third: all checks pass, or cap hit + honest report.
One line per pass: what failed, what got fixed, what was learned. The log is what makes run five smarter than run one — lessons like "draft long, then cut" carry forward into your next mission. It's also your audit trail when a result looks off.
LOOP LOG: 2 passes · trimmed 52 words · lesson —
draft long, then cut; burn time belongs in sentence 2Add one rule: stuck on the same failure twice → stop and ask me. Without it, a loop that can't satisfy a check will grind its cap away in silence or, worse, quietly relax your requirement to escape. The hatch converts a stuck loop into a good question.
Creative and analytical work splits into two layers: the checkable (length, structure, coverage, sourcing) and the taste (is it moving, is it wise). Loop the first, judge the second yourself. A loop that gets the craftsmanship right hands you something worth exercising taste on. (Chapter 03 covers the full decision.)
A loop is an asset, not a one-off. The six-label file from Chapter 02 re-fills for a new task in minutes — and a small library of them (one for descriptions, one for reports, one for emails) is genuinely what the "I just write loops" people have that you don't, yet. Save every loop that worked.
Writing a prompt that makes an AI work in passes: do the work, grade it against yes/no criteria with quoted proof, fix only what failed, and stop honestly — every check passes, or a pass cap is hit and it reports truthfully. The full mechanism is in Chapter 02.
No. A loop is plain text — six labeled lines you paste into any chat AI (ChatGPT, Claude, Gemini). No software, no code.
Because that exact strategy was measured: accuracy fell from 75.8% to 38.1% under naive self-review (Huang et al., ICLR 2024), unguided checkers rejected correct work 96% of the time (Stechly et al.), and 83% of gains land in the first structured pass (Madaan et al., Self-Refine). Iteration only works with hard rules — practices 05 through 08.
Three to five. Fewer and the loop can't see its own failures; more and every pass drowns in grading. Each must be answerable yes or no.
When the result can't be checked with yes/no questions, when the first draft is realistically good enough, or when the task is smaller than the setup. Chapter 03 is the ten-second test.
Pick one practice and apply it today — 02 and 05 pay off immediately. When you'd rather have the spec and checks written for you, that's the tool we build.
Try Prompt Optimizer