---
name: catchup-distill
description: "Distills a finished coding session — commits, decisions, architecture changes, and durable learnings — into structured JSON items and knowledge events, then submits them to the hosted Catchup MCP server, which renders them into a swipeable feed, narrated audio, and short video, entirely without a server-side LLM. Use this at the end of a work session: after a commit lands, a bug is fixed, a decision is made, or an incident is resolved. Do not use it mid-task, during read-only exploration, or when the session produced nothing another engineer would need to know about."
context: fork
version: "2026-08-20"
---

# Catchup Session Distiller

Runs once, at the end of a work session, in a forked context so none of this JSON scaffolding pollutes the coding session's window. On the other end is a server that renders mechanically and runs zero LLM — nothing downstream double-checks whether what you submit is true. That is what makes the rules below non-negotiable rather than suggestions.

## Non-negotiables

1. **Never claim anything about code you did not read this session.** If you're not sure, the correct output is "I don't have enough information" — not a guess dressed up as a fact.
2. **Quotes before claims.** Pull verbatim quotes from the code before you analyze anything. After you draft, go back and find a supporting quote for every claim; delete any claim you can't source.
3. **Every source needs a `path` and a `content_hash` pinned to the exact span you read.** Before you emit, re-read each cited span. If it doesn't say what the item claims, delete the claim, not the source. The support has to sit **inside** a region you cited — a true claim whose evidence lives outside every cited span is an under-citation, not a pass: add the source that actually supports it, per claim, not per card. Docs count as sources — when `CLAUDE.md` or a decisions file is where a number or a decision actually lives, cite that file exactly as you would code. (A real cold-repo grading pass found 3 of 11 otherwise-clean cards failing exactly this way: every fact true, one fact per card supported only by a span that was never cited.)
4. **Reason free-form first.** Do your thinking — reading, quoting, drafting, doubting — in plain prose. The JSON is the last thing you produce, not the medium you think in.
5. **Never write HTML.** Everything goes in the typed template fields — `title`, `bullets`, `narration`, `diagram`, `code`, `quiz` — or it doesn't go in at all.

## Craft

A developer audience rejects on style before it ever evaluates truth. A director ended a document review in about five seconds — *"I can't read this because of the AI dashes, write it in your own words"* — on a document about using AI more effectively, and the complaints that followed map directly onto what these cards do: mixed levels of abstraction, no headline, the path taken instead of the why. Our entire corpus is machine-written, for exactly this audience — run these five checks before you emit anything.

1. **Headline discipline.** Describe this card and the session's biggest change, each in one sentence. If the card isn't about the bigger one, re-rank or rewrite it — an `importance: 2` card from a session that shipped an architecture change is mis-ranked, not small.
2. **Why over path.** Write what is true now and why it changed. Send dead ends and pruned branches to the event's `raw` field, never to a bullet — nobody reading the card needs the branches you didn't take.
3. **One level of abstraction.** Pick one zoom level and hold it for the whole card. A card that opens on system architecture and closes on a variable rename is two cards, or one bad one — split it or cut half. This is what `difficulty` is for, and once the ranker reads it (10-04), a mixed card is ranked wrongly for everyone who sees it.
4. **The anti-tell pass.** Re-read your narration and bullets as a hostile reader hunting for tells — em-dash cadence, tricolons, "it's not X, it's Y", delve/leverage/robust, a sentence that restates the sentence before it — and rewrite anything you find. Run this gate before the other four, whatever order you read them in: the audience rejects on style before it evaluates truth, so check style first.
5. **Read the card back.** Pick any bullet and ask "why?" If you can't answer it from the spans you actually cited this session, cut the bullet. This is also your answer to the strongest cultural objection to generated docs — *"the author didn't read it so they can't answer the questions either"* — and here you genuinely can.

*TODO(09-03): gate 5 wants a real `contradicted` claim from the hallucination sample as its worked example, quoted next to the span that contradicts it, once that run is graded — see `docs/operations/hallucination-sample.md`. Left as a marker rather than filled with an invented failure: a fabricated example in a gate that exists to prevent fabrication would be the worst possible artifact here, and it would look identical to a real one to every future reader.*

## The writers room

Great cards are selected, not compressed. A card that tries to cover the whole session can only ever be a compressed changelog — the hollow opener and the summary ending are not writing failures, they are what a coverage mandate produces. Run these four passes instead. If your harness supports subagents, the skeptic and cold reader work best with fresh context; run them in order regardless.

**1. The miner.** Before drafting, list every point this session where reality pushed back: an approach that got reverted, a test that redirected the work, a constraint that forced an odd design, a number that moved more than expected. Pick one. The card is the story of that moment; everything else becomes a bullet or gets dropped.

**If nothing pushed back, you still have a story — a different shape.** Set up the change plainly, then bring in one true element from a different scale or subject (what it costs, what it makes newly possible, what it now resembles elsewhere in the system), and let the last beat connect that back so the setup means something it didn't mean at the start. Nothing opposes anything and it still turns, because the turn is a reframing rather than a reversal. Reach for this on the routine sessions — it is the honest shape for most of them, and "ignorance to knowledge" is a real turn even when no conflict exists. Only when neither shape fits does the card go bullets-only with no screenplay.

**2. The stakes pass — who benefits, who loses, why now.** One plain sentence, before you draft a single beat. Small and sharp beats large and vague: "the next person to add a lane loses an afternoon to this" is worth more than "this improves reliability." Then the *why now* half, which is the strongest test in this room: **is this true because of this change, or was it equally true a year ago?** A card that would have been just as accurate before the session is describing the system, not reporting it — that is a coverage-gap card at best, and it is the single most common way a session turns into a recap. Cards that fail this get re-mined, not rewritten.

**3. The screenwriter.** Fill the beat sheet. Beats are unequal on purpose — hook and turn short and hard, proof denser — and a three-beat card keeps 1, 4, 6.

| Beat | Job | What goes in it |
|---|---|---|
| 1 Gap | Open a loop | The true outcome, cause withheld. First eight words carry the anomaly: a number, a contradiction, a real question |
| 2 Stakes | Localise the cost | Units someone paid: ms, retries, lines, pages, rebuild minutes |
| 3 Obvious move | Install the expectation | What any competent dev would do — the tried-and-reverted approach, the tutorial answer |
| 4 Turn | Break it | Why the obvious move fails here: the constraint or discovery actually in the diff. This is the payload |
| 5 Receipt | Cash the claim | The number after, the test that pins it, the line quoted by name |
| 6 Residue | Consequence, never moral | What the change cost or traded away, or the true question it opened. State a fact that implies the theme; never state the theme |

Adjacent beats connect by "but" or "therefore", never "and then". If two neighbouring beats join with "and then" and lose nothing, the card is a list wearing a screenplay.

**4. The skeptic.** Run the five Craft checks above against the draft, as a hostile reader. The server re-checks the mechanical half and rejects what it can see; meaning is yours alone. Add one substitution test: **could any other project's session have produced this exact bullet?** If yes, it is borrowed traffic, not your session — cut it or make it specific.

**5. The cold reader.** Read the title and first beat alone, as a stranger to this repo. Would they stay past three seconds? Would the hook survive the project name being bleeped out? If not, open on the universal phenomenon and let the repo-specific detail enter at beat 2.

**Receipts over prose.** When unsure of your own writing — smaller model, unfamiliar domain, low confidence — quote reality instead of composing: the real error message, the real diff line, the oddly-specific number. A quoted receipt cannot be slop and cannot hallucinate, and the strongest short-form experiments are exactly this shape: the story told *as* its artifact — as the test log that redirected the work, as the error and the three wrong explanations it survived.

The one way this doctrine fails: **a receipt that just sits there.** Quoting is the floor, not the destination — a beat whose only job is delivering a fact has no reason to exist, and a card that quotes reality at every beat never actually says anything. Make the receipt the instrument of the turn: the line that flips the reader from not-knowing to knowing. If a beat is a receipt and nothing else, fold it into the beat it proves.

**Register: spoken, not written.** Read the narration aloud in your head; rewrite anything you would not say to a colleague at their desk. Lecture cadence, formula ("statement, explanation, example"), and smart-sounding connectives (furthermore, thus, moreover) read as fake-smart to exactly this audience — directness is the eloquence. Second person: "you shipped," not "the developer shipped" (d=0.30 retention, d=0.54 transfer). Narration runs 30-60 seconds spoken; one idea per item.

## Load session context

This runs before you see the rest of the file — you're not choosing to look, the harness already ran it.

Diff since the last commit:
!`git diff --stat HEAD~1`

Commit SHA, for reference:
!`git log -1 --format=%H`

This diff is one input, not the anchor — see the five evidence kinds in step 2 below. An empty or trivial diff does not mean nothing happened here: on a quiet week, the other four kinds (what's already in the corpus, structure, the code itself, reader signal) are exactly what carry the session.

## The process

0. **Read `project_state` in the brief before anything else — it decides which job you are doing.**

   | State | What it means | What you do |
   |---|---|---|
   | `needs_backfill: true` | The feed is still too thin for a first scroll — check `backfill_remaining` for how many are still owed | **Stop reading the diff.** Walk the repository instead: entry points, the main subsystems, the decisions visible in the git log. Produce at least `backfill_remaining` more evergreen items. A first scroll that is empty is the single most common reason someone never comes back. |
   | `quiet: true` | Nothing new for 5+ days | Don't wait for a commit. Mine: code untouched for months that a newcomer would trip over, patterns repeated across the codebase, decisions whose reasoning is no longer obvious. |
   | neither | Ordinary session | Continue with the steps below. |

   Backfill is the one time you should ignore what just happened in the session entirely. **A backfill often does not finish in one session** — the walkthrough is long and context runs out. That is expected and safe: `needs_backfill` stays true until the feed reaches its target, so the next session picks up where this one stopped. Check `backfill_remaining` rather than assuming you are starting from nothing, and call `catchup_check_coverage` first so you extend what is already covered instead of re-walking it.

   **Before you start walking the repo, tell the user roughly what it will cost.** A backfill runs on the user's own subscription, not ours — we run zero LLM server-side, which is exactly why this spend is otherwise invisible to everyone but them. `project_state.backfill_cost_estimate` is present whenever `needs_backfill` is true and carries `estimated_tokens` for the `items_remaining` still owed, plus `basis` naming where the number came from. State it in one line, in the user's own terms — "this walkthrough will run roughly N tokens on your own subscription for the M items still owed" — and never turn it into a dollar figure: we don't know their plan or rate, and a wrong price is worse than none.

   **After finishing a work stretch — whether or not the backfill is fully done — call `catchup_record_backfill_actual`** with how many items you produced and your own best-effort `input_tokens`/`output_tokens` for that stretch. Omit the token counts rather than guess if you genuinely can't estimate them; `items_produced` alone is still worth recording honestly. This is how the estimate improves over time instead of staying fixed at one initial measurement. **The first backfill card of a fresh project never opens on a definition.** Open on a question or a receipt instead. A question names the strange thing before explaining it — the first mining sitting's own candidate rewrites a card to open with *"Why does an airline chatbot keep four classifiers on disk?"* instead of leading with the answer. A receipt quotes reality instead of describing it — that same sitting's strongest cold-open candidate is a real passenger string, quoted exactly as the system saw it: *"جدة للمغرب امبارح"* (roughly, a Jeddah-to-Morocco itinerary). Both earn the next three seconds from a stranger who doesn't know yet why a definition would matter; a definition doesn't. This is a testable claim, not a taste call — `fast_skip` already measures instant dismissal on any card; watch its rate on a project's first-served card to find out whether it holds.

1. **Work the queue top-down.** Call `catchup_get_generation_brief` first. It returns `assignments` — a ranked queue, highest priority first, each with a `kind`, a `why`, and a `shape` telling you what to make. `budget.items_remaining_today` caps how many you can take; `budget.deferred` says how many of the queue are pushed to next session, so you never have to infer it. An assignment is not a quota: if the top item's honest output is nothing, report that (see worked example 5) and move on rather than manufacturing content to fill it. **Before that call, walk the repo's top-level directories and build `observed_areas`** — this is how the brief learns *structure* (step 2's evidence table, below). Skip exactly this set, the same `SKIP_DIRS` `scripts/mine.ts` uses: `.git`, `node_modules`, `.next`, `dist`, `build`, `coverage`, `.venv`, `venv`, `__pycache__`, `vendor`, `target`. `venv` — no leading dot — earned its place on that list: a Python virtualenv under exactly that name once put 15MB of vendored streamlit/plotly/numpy source at the top of a real repo's unexplained areas, ahead of every first-party file, during dogfood testing on backend-2. That was a real incident, not a hypothetical, and it is why skipping dot-directories alone isn't enough. For every directory that's left, note roughly how many source files and bytes live under it, and mark `flow: true` on anything sitting on a money-path flow — checkout, booking, payment, auth. That flag floors the area's urgency at 0.8 server-side regardless of size, so a small but critical flow doesn't lose to a large but peripheral one. Pass the result as `observed_areas: [{path, files, bytes, flow?}]` on the call.

   | Kind | Means | Do |
   |---|---|---|
   | `repair_flag` | A reader flagged an item | See the flag table under Submit, below |
   | `answer_question` | A reader asked, nobody's answered | `catchup_answer_question` with `target.question_id` |
   | `clarify_item` | Readers marked this card confused; nobody flagged it, so they told you it failed and not why | `catchup_repair_item` on `target.item_id` (no flag to resolve). Rewrite from what the card *assumes* — unexplained nouns, the identifier dropped from a span you already cite — not from a guess at their complaint |
   | `due_recall` | A topic is due for spaced retrieval | See Due-recall items, below |
   | `prerequisite_gap` | A live card assumes a concept no live card teaches | Write an intro card whose `teaches` names `target.concept` exactly — that's what clears the gate for every card that assumes it |
   | `coverage_gap` | Commits touch a topic with zero items | Write it, at the difficulty `shape` suggests |
   | `session` | Today already logged something durable, uncovered | Write it, or confirm today's diff already covers it |
   | `unexplained_area` | The area walk above found a directory no live card cites | Read `target.path`, then teach what's there, intro-first; `target.flow` flags a money-path directory |
   | `decision_point` | A topic tops out at intro; the code may hold a real trade-off | Read the code behind `target.topic`. A genuine non-obvious choice — write a `decision-card` at `working`+. No real choice — decline; that's a legitimate outcome, not a failure |
   | `weak_topic_depth` | Mastery is low and only one depth exists | Same topic, a genuinely different depth |
   | `stale_refresher` | Untouched 90+ days, project's gone quiet | Mine history for something worth refreshing |
   | `cross_project_pattern` | Same topic, different depth/decision across your own repos | Advisory — write only if the contrast genuinely teaches something |
   | `weekly_recap` | The week produced enough to summarise | Advisory — skip if there's nothing worth recapping |
   | `corpus_shape` | A form or craft deficit (too few quizzes, recurring craft findings) | Advisory — act where the material earns it; for craft, run the writers room below |
2. **Ask the session's question.** Every session, weigh five kinds of evidence and make one judgment: **what most advances this reader's understanding of this repo, given what they already have.**

   | Evidence | What it yields | Where it comes from |
   |---|---|---|
   | **What is already created** | the corpus: what is covered, at what depth, and what it depends on — so the next card extends rather than repeats | server, via the brief |
   | **Structure** | files and directories that exist and have never been explained; a subsystem that appeared | agent, filesystem |
   | **The code itself** | the key decision points, invariants and gotchas that no commit message describes | agent, reading source |
   | **Change** | what moved recently — one input, not the anchor | agent, working tree + session |
   | **Reader signal** | confusion, questions, flags, mastery — what this reader specifically failed to get | server, via the brief |

   The brief (step 1) already handed you two of these — the corpus and reader signal — and the area walk (also step 1) handed you a third, structure. The diff and commit loaded above are the fourth, change: real evidence, never the whole question. The fifth, the code itself, is what step 3's close read is for. Weigh what you have against the judgment above, not against the diff alone. No decision, no shipped change, no durable learning, nothing across any of the five — stop here and submit nothing. See worked example 5.
3. **Re-read the actual spans that changed**, not just the diff summary. Pull verbatim quotes as you go; this is the raw material for both `sources[]` and for every claim you're about to write.
4. **Draft in prose first.** What happened, why, what's non-obvious, what's worth testing someone on. No JSON yet.
5. **Split into items.** One idea per item. A session might produce zero items, one, or several — don't cram two ideas into one item's five bullets, and don't split one idea across two items.
6. **Check coverage before you commit to a topic.** Call `catchup_check_coverage` with the topics you're about to use. If something already exists at the same depth, either go deeper or drop the item. Do not restate what the feed already says.
7. **Check every claim against a quote.** If a sentence in your draft isn't backed by something you actually read this session, cut the sentence, not the rule.
8. **Build `sources[]`** for each item: `path`, the `region` you actually read, and its `content_hash` — the hashed shape (see The item, below), not the legacy `commit_sha` one.
9. **Decide on knowledge events, independently.** A decision, change, learning, architecture shift, or incident becomes an event whether or not it also became an item.
10. **Scrub secrets, then submit.** Both covered below.

## Payloads

### The item

```json
{
  "item_id": "client-generated-uuid",
  "template": "diff-walkthrough | decision-card | architecture-diagram | quiz-card | timeline | concept-pill",
  "title": "max 120 chars",
  "bullets": ["max 5, each max 200 chars"],
  "narration": "plain-text script for TTS, MAX 150 WORDS (~60s spoken)",
  "diagram": "mermaid source (optional)",
  "code": [{"lang": "ts", "snippet": "scrubbed source", "ref": "path:line"}],
  "quiz": {"question": "...", "options": ["2-4 options"], "answer": 1},
  "topics": ["1-8 topics"],
  "teaches": ["1-8 concept slugs, lowercase (optional)"],
  "assumes": ["0-4 concept slugs a reader needs first (optional)"],
  "difficulty": "intro | working | deep",
  "importance": 3,
  "sources": [{"path": "src/auth/session.ts", "region": [120, 168], "content_hash": "sha256:07818422d934755ce6b1592d61b0c5a0b83faa8b90167344eb39a6dd661af246"}],
  "generator": {"model": "your own model identifier", "skill_version": "this file's version, above (optional)"}
}
```

Lengths, enum values, and array bounds are enforced by the server's validator. Get the truth right; let the server catch a stray 201-character bullet.

- `item_id` — you generate this, client-side. A real UUID, not a placeholder string.
- `template` — pick the one shape that actually fits. Don't stretch a decision into a `quiz-card` just because you thought of a good question.
- `title`, `bullets`, `topics` — **shape discipline**, from the first mining sitting: the server's bounds (title 120 chars; 5 bullets at 200 chars each; 8 topics) are ceilings, not targets. The height gate rejected 8 of 11 first drafts written to those limits; what survived was a title **≤~54 chars**, **3–4 bullets, each ≤~80 chars**, and **≤3 topics** — every extra topic chip and every untrimmed bullet is less screen space for the idea itself. If the idea doesn't fit, split into a second item — never trim a bullet to make room. The bullet that gets cut to fit is always the last one, and the last one is where the caveat lives.
- `diagram`, `code`, and `quiz` are optional. Include them when they carry real information, not by default.
- `teaches`/`assumes` (both optional) — concept slugs, lowercase. `teaches` is what this card delivers: 1–8 of them, but keep it to the 1–3 concepts a reader actually takes away, not everything the card mentions (a slug is a short, reusable name — `rate-limiting`, not `the-rate-limiter-middleware-refactor`). `assumes` is what it presupposes: **at most 4** — a card assuming five concepts is a card teaching the wrong thing, not one that needs a longer list — and only what another card could plausibly `teach`; an intro card assumes nothing, because there's nothing earlier on the topic yet for it to presuppose. The ranker demotes any card whose `assumes` aren't met yet for that reader, so a concept named here that nothing `teaches` blocks every card that assumes it — exactly what a `prerequisite_gap` assignment (above) is asking you to fix.
- `importance` is your judgment call, 1 (minor) to 5 (changes how someone thinks about the system).
- `code[].snippet` gets scrubbed before it lands in this field — see Secret scrubbing below. The schema calls it "scrubbed source" for a reason.
- `sources` is never empty. Every item has at least one, in the **hashed shape**: `path`, the `region` you read (1-indexed, inclusive line numbers — omit only when the claim is about a whole file's shape, not a span), and `content_hash`, the sha256 hex digest of exactly that text — lines `region[0]`..`region[1]` joined by `\n`, no trailing newline (the file's raw text, unmodified, when `region` is omitted). That's exactly what `scripts/mine.ts`'s `hashRegion` computes; matching it is what lets a later session's staleness check agree with the hash you write today. The legacy `{kind, ref, commit_sha}` shape still validates — every card written before this version carries it — but this skill doesn't write it: hashed is what capture uses now, and unlike a commit pin, it works with no git history at all. The schema also carries an `excerpt` field (verbatim quoted text, for a reader with no repo access) — leave it unwritten. Whether this project may retain verbatim excerpts from a client repo is a founder-owned retention decision that hasn't landed; this skill doesn't populate the field until it does, even though a value there would pass validation.
- `generator` (optional) — your own model identifier and this file's `version` (see frontmatter) as `skill_version`. Not certain of your model string? Omit the whole field rather than guessing — a wrong label is worse than a null, because a null is honest and a guess silently corrupts a comparison.
- `assignment_id` (optional) — answering an assignment from `get_generation_brief`'s queue? Echo its `id` verbatim. Not answering one — an unprompted item, something worth saying that nothing asked for? Omit the field entirely; that's a legitimate, desirable item, not an error. Never attach the nearest plausible assignment id to make the record look tidy — that turns an honest omission into a false attribution, and every later question about which assignment kinds actually produce good content gets answered from contaminated data.

### The knowledge event

```json
{
  "event_id": "client-generated-uuid",
  "type": "decision|change|learning|architecture|incident",
  "summary": "one sentence, plain language",
  "why": "one sentence: the trigger or reasoning",
  "entities": ["files, modules, or concepts involved"],
  "code_ref": "path:line",
  "importance": 3,
  "raw": "longer verbatim note (optional)"
}
```

Events feed the durable record — decision history, per-topic mastery — independent of what gets rendered. A session can produce an event with no matching item (worth remembering, not worth 60 seconds of someone's attention) or an item with no matching event (a good explainer about something that isn't new). Most sessions worth submitting anything for produce at least one of each.

- `event_id` (optional) — a real UUID, generated the same way `item_id` is. Include it if you can; a call carrying it is safe to retry exactly like `send_content` (see Submit, below). Omitting it still works — the event still ingests — but a dropped connection on a call with no `event_id` cannot be told apart from a genuinely new one, so include it whenever you're able to.

## Quiz rules

These get their own section because they matter more than anything else about quizzes.

Well-built multiple-choice questions are unusually effective: later recall around d=1.30, against d=0.67 for cued recall, and they even lift recall of related material that was never directly tested. But that whole effect rides on the distractors. Non-competitive distractors — the obviously-wrong kind — produce zero benefit. And superficial distractors are the default failure mode of LLM-written quizzes.

> Every distractor must be something a reader who half-understood the session would actually believe. Ruling it out has to require recalling a real fact, not just noticing it's silly.

**Good — every option is true, only one answers the question asked**

Q: Why did this session move rate limiting from middleware to the queue consumer?
- Correct: Retries were counted as new requests, so a flaky connection burned a caller's quota.
- The queue consumer has access to the user's plan tier, and middleware did not. (True — but that was true before. Not why this changed.)
- Middleware couldn't read the parsed request body at that point in the chain. (A real class of problem. Not this one.)
- Rate limiting was removed entirely, not moved to a different layer. (Contradicts the premise, but reads plausibly to a skimmer.)

Every wrong option is a true statement about the system. Picking the right one takes having actually read the reasoning, not pattern-matching on vocabulary.

Note the *shape* as well as the content: all four options run to about the same length and the same level of detail. That is not incidental — see the checklist below.

**Bad — wrong on inspection, no recall required**

Q: Why did this session move rate limiting from middleware to the queue consumer?
- Correct: Retries were being counted as new requests.
- Because the developer prefers queues.
- Because middleware was deleted for no reason.
- Because rate limiting doesn't matter.

Three options are absurd without knowing anything about the change. The question ends up testing whether the reader can spot a silly sentence, not whether they read the item.

### Critique your own quiz before you send it

Writing the question and checking the question are different jobs, and doing them in one pass is how the failure above ships. After you've drafted a quiz, re-read it as a reader who did *not* attend the session and is trying to game it. Work down this list; if any answer is yes, rewrite before sending.

**Shape — can I pick the answer without knowing anything?**

1. Is the correct option noticeably longer or more detailed than the others? This is the most exploited flaw in multiple choice, and it is the one you will commit by default: you elaborate the true statement because it has to survive scrutiny, and stub the false ones because they don't. Give every option the same weight.
2. Do all the wrong options hedge one way and the right one the other — every distractor absolute (*always*, *never*, *only*, *must*) while the answer is measured? "Eliminate the absolutes" then wins outright.
3. Do two options say the same thing in different words?
4. Is there an "all of the above", a "none of these", a "both A and B"?
5. Is the answer in the same position as the last few quizzes you wrote? There is a real pull toward the first option. Move it.

**Meaning — does ruling each one out take a fact?**

6. For each distractor, name the specific thing a reader would have to know to eliminate it. If you can't name one, that distractor is decoration — replace it.
7. Would someone who half-read the session plausibly pick it? A distractor nobody would choose contributes nothing; the best ones are the misunderstanding you actually had before you traced the problem.
8. Is every option true as a statement about the world, with only one true *as an answer to the question asked*? That's the target the good example above hits.

The first five are checked on the server and a quiz that fails them is rejected with a repair hint — not because a machine can judge your distractors, but because these particular flaws are visible in the shape of the options alone. The last three are yours; nothing else in the system can see them. And a quiz is optional: if you can't build distractors that meet 6–8, send the item without one. A card with no quiz is worth more than a card with a quiz that teaches nothing.

## Due-recall items

When the queue includes a `due_recall` assignment, FSRS is asking for a genuine retrieval attempt on that topic, not a nudge to revisit old ground. A recall item only earns that name if it tests the same underlying fact from a different entry point than every item already listed in `seen_item_ids` — a consequence rather than the mechanism that produces it, a failure mode rather than the definition that names it. Restating `seen_item_ids` in new words is the re-exposure the brief's own `constraint` field exists to rule out, one layer above where the ranker already blocks it.

Some topics are small enough that no genuinely different entry point exists yet. When that's the case, the correct output is nothing — the same answer worked example 5 already gives for a session with nothing durable to report, and the same trade the quiz rules already make: a card with no quiz beats a quiz that teaches nothing, so a due topic left unanswered beats a recall item that only paraphrases the card it was supposed to replace.

## Screenplay and visual assets

An item's `screenplay` places named characters and objects into its video — beats of on-screen text, each naming which assets from this project's library appear and where. The server runs zero LLM and generates nothing per item: every asset a screenplay uses has to already exist, drawn once and reused across many items. Unlike a citation span, the server *can* verify a screenplay reference against its own database — so it does, at ingest, and that's what the rules below are shaped around.

**Call `catchup_list_assets` before writing a single beat.** It returns every character, object, and background available in this project right now, each with its poses and a short description. That call is the entire library. There is nothing wider to imagine your way into.

**Only ever reference a `name`/`pose` that call actually returned.** Never invent one, even a pose you're confident "should" exist for a character you can see is already there. The server checks every reference against this project's asset table before the item is stored, and a screenplay naming anything not on that exact list is rejected whole, with a repair hint pointing back at `catchup_list_assets`. Inventing one doesn't fail softly — it just spends a round-trip you didn't need to spend.

- Good: `catchup_list_assets` returned `fox-narrator` with poses `idle`, `talking`, `pointing`. The screenplay in the worked example below uses exactly that name and exactly those three poses, nothing else.
- Bad: that same library has no `terminal-icon` and no `excited` pose for `fox-narrator` — but both feel like they should exist, so a beat references them anyway. Both get the whole screenplay rejected. The library doesn't grow because a screenplay wished it would; it grows the way the next rule describes.

**`screenplay` is optional, exactly like `quiz`, `diagram`, and `code`.** If nothing currently in the library fits what this item needs, don't force it — send the item as text and narration only. A card with no screenplay beats a screenplay whose character doesn't actually match what's being explained: a mismatched visual competes with the words instead of reinforcing them, and this project would rather ship the words alone.

**You cannot create a new asset.** There is no tool call for it. Asset creation is `scripts/generate-asset.ts`, run by a human, on purpose: each new character or pose costs a real image-generation call, and the library is meant to grow slowly and deliberately — a handful of reusable characters and objects — not per item, and not from inside a session. If the topic needs a character the library doesn't have yet, that's not a gap to fill right now; it's the same case as any other missing asset. Drop the screenplay for this item.

### 5. screenplay — reusing an existing character

Session: explained why payment-webhook retries were being counted against the caller's rate limit (the same underlying fix as worked example 1 below). The project's asset library already has a narrator character, so this time the item ships a screenplay instead of narration-only video.

`catchup_list_assets` returned (abbreviated):
```json
[
  {
    "name": "fox-narrator",
    "kind": "character",
    "poses": ["idle", "talking", "pointing"],
    "description": "friendly orange fox mascot, transparent background"
  }
]
```

Item:
```json
{
  "item_id": "9e2f1c4a-7b3d-4a10-8f6e-2c1d9a3b5e70",
  "template": "diff-walkthrough",
  "title": "Retries were eating your own rate limit",
  "bullets": [
    "Webhook retries were hitting the rate limiter before dedup, so a flaky network could burn a user's whole quota on one event.",
    "The check moved from Express middleware to the queue consumer, which runs after the dedup step."
  ],
  "narration": "Your payment webhook handler had a quiet bug: every retry counted against the caller's rate limit, even though a retry isn't a new request. Moving the check to the queue consumer, after dedup, fixed it.",
  "screenplay": [
    {
      "text": "One flaky connection burned a caller's entire quota on a single event.",
      "assets": [{ "name": "fox-narrator", "pose": "idle", "position": "left" }]
    },
    {
      "text": "But a retry isn't a new request. The limiter was counting every attempt before dedup ever ran.",
      "assets": [{ "name": "fox-narrator", "pose": "talking", "position": "left" }]
    },
    {
      "text": "The check now lives in the queue consumer, after dedup. A burst of retries costs one unit, not one per attempt.",
      "assets": [{ "name": "fox-narrator", "pose": "pointing", "position": "center" }]
    }
  ],
  "code": [{"lang": "ts", "snippet": "if (isRetry(req)) return next();", "ref": "src/webhooks/consumer.ts:88"}],
  "topics": ["rate-limiting", "webhooks", "retries"],
  "difficulty": "working",
  "importance": 3,
  "sources": [{"path": "src/webhooks/consumer.ts", "region": [85, 90], "content_hash": "sha256:684c9829cea7f3883c9018113f29ece952521a5dc6d385dd4bde092a4424e1ec"}]
}
```

Three beats, one asset each, only the names and poses `catchup_list_assets` actually listed. Nothing here invents a character, a pose, or a fact the session didn't establish.

## Narration voice

Every card is narrated. By default it uses the project's voice, and
**omitting `voice` is the normal case** — a consistent narrator is what
makes a feed feel like one thing rather than a playlist.

Set `voice` only when the content genuinely justifies a different speaker:
a quoted incident report, or a card that contrasts with the one before it.
Call `catchup_list_voices` to see what is available — each voice carries a
measured phonation score and a sounds-like note.

**You cannot create, clone, or upload a voice** — a voice not in that list
is rejected at ingest rather than silently replaced.

## Secret scrubbing

Scrub before you send anything. The server scrubs again on the way in, but that's a backstop, not a plan — this is the first of two passes, and yours is the one that has to work.

Redact:
- AWS keys
- `gh*_` tokens
- `xox*-` tokens
- `sk-` keys
- Bearer tokens
- PEM blocks
- Any `key=value` where the key looks like `secret`, `token`, `password`, or `api_key`

If a snippet needs a redacted value to still make sense, replace it with a shape-preserving placeholder (`sk-***`, `AKIA***`) instead of deleting it outright.

## Submit

Call `catchup_send_content` once per item, and `catchup_send_event` once per knowledge event — independent calls, not a bundle. If the session produced neither, don't call anything; tell the user in one line that nothing durable happened.

If you're fulfilling an `answer_question` assignment, use `catchup_answer_question` with its `target.question_id` instead of `catchup_send_content`, so the question is marked answered rather than left in the queue.

If you're fulfilling a `repair_flag` assignment, repair it with `catchup_repair_item` using its `target.flag_id`. A flag is a reader telling you the card failed them, and each kind means something different:

| Flag | What went wrong | What to send back |
|---|---|---|
| `false` | The claim is wrong. The item is already hidden from the feed. | Re-read the cited span. If the claim was wrong, replace it with what the code actually does. If the claim was *right*, say so plainly and show the evidence — being right is worth explaining. |
| `unclear` | Right topic, wrong execution | Rewrite plainer. Do not add detail; unclear usually means too much, not too little. |
| `duplicate` | Already covered | Go deeper than the existing item or don't replace it at all. |
| `too_basic` / `too_deep` | Wrong depth for this reader | Same content, different `difficulty`. |

Repairing supersedes the original and closes the flag. Left alone, flags surface in every future brief forever.

**When a call comes back as an error, it is telling you how to fix it.** The server validates every payload and returns the exact field path and what was wrong — `sources.0.content_hash | must be sha256:<64 hex chars>`. Read it, correct that field, resubmit. Do not drop the item, do not work around the check by weakening a claim, and do not invent a SHA to satisfy the validator. If you cannot produce a real citation for a claim, the claim is the thing that was wrong.

Submitting the same `item_id` twice is safe — the server recognises it and does nothing. The same is true of `send_event` when you include `event_id`: a repeated call with the same `event_id` is recognised and inserts nothing. Retry freely if a call fails midway.

## Worked examples

### 1. diff-walkthrough — a retry bug

Session: fixed a bug where retry logic in the payment webhook handler counted each retry against the caller's rate limit. Moved the check from Express middleware to the queue consumer, after dedup.

Item:
```json
{
  "item_id": "b3f1c2a0-6e21-4b9a-9d3a-1e2f3a4b5c6d",
  "template": "diff-walkthrough",
  "title": "Retries were eating your own rate limit",
  "bullets": [
    "Webhook retries were hitting the rate limiter before dedup, so a flaky network could burn a user's whole quota on one event.",
    "The check moved from Express middleware to the queue consumer, which runs after the dedup step.",
    "isRetry(req) now short-circuits the limiter for anything already seen once."
  ],
  "narration": "Your payment webhook handler had a quiet bug: every retry counted against the caller's rate limit, even though a retry isn't a new request. You fixed it by moving the rate-limit check out of Express middleware and into the queue consumer, which runs after deduplication. Now a burst of retries for the same event costs one unit of quota, not one per attempt.",
  "code": [{"lang": "ts", "snippet": "if (isRetry(req)) return next();", "ref": "src/webhooks/consumer.ts:88"}],
  "topics": ["rate-limiting", "webhooks", "retries"],
  "difficulty": "working",
  "importance": 3,
  "sources": [{"path": "src/webhooks/consumer.ts", "region": [85, 90], "content_hash": "sha256:684c9829cea7f3883c9018113f29ece952521a5dc6d385dd4bde092a4424e1ec"}]
}
```

Event:
```json
{
  "type": "change",
  "summary": "Moved the webhook rate-limit check from Express middleware to the queue consumer, after dedup.",
  "why": "Retries were being counted as new requests and could exhaust a caller's rate limit on their own retries.",
  "entities": ["src/webhooks/consumer.ts", "rate limiter"],
  "code_ref": "src/webhooks/consumer.ts:88",
  "importance": 3
}
```

### 2. decision-card — polling to push

Session: replaced a polling-based sync job with a webhook-driven push model after diagnosing a 40-minute worst-case staleness window under backlog.

Item:
```json
{
  "item_id": "7a2d9e10-44f0-4c2b-8a11-2d3f9b7c5e60",
  "template": "decision-card",
  "title": "Sync went from polling to push",
  "bullets": [
    "The old sync job polled every 5 minutes, with a worst-case 40-minute staleness window under backlog.",
    "It is now driven by an incoming webhook, so updates land in seconds instead of minutes.",
    "The polling loop and its cron entry were deleted, not just disabled."
  ],
  "narration": "Sync used to poll on a five-minute cycle, which meant a worst-case staleness window of forty minutes whenever the queue backed up. You replaced it with a webhook: the upstream system now pushes on change, and the handler processes it within seconds. The old polling loop and its cron entry are gone, not just turned off, so there's no dead code path left to confuse the next person.",
  "topics": ["sync", "architecture", "webhooks"],
  "difficulty": "deep",
  "importance": 4,
  "sources": [{"path": "src/sync/scheduler.ts", "region": [8, 15], "content_hash": "sha256:d79de34c45b68a5529f1b6648f037d9453bf88d1f1bc14758e0e197ab2913a5b"}]
}
```

Event: same pairing as example 1 — `type: "decision"`, the 40-minute staleness window as the `why`, `src/sync/scheduler.ts:12` as the ref.

### 3. architecture-diagram — a queue in front of the workers

Session: added a queue between the API and the worker fleet so traffic spikes no longer hit workers directly.

Item:
```json
{
  "item_id": "c4e8a715-2b90-4f6d-9a3c-6f1d8e2b4a90",
  "template": "architecture-diagram",
  "title": "A queue now sits between the API and the workers",
  "bullets": [
    "Traffic spikes used to hit the worker fleet directly, causing timeouts under load.",
    "Requests now land in a queue; workers pull from it at their own pace.",
    "The API responds immediately with a job id instead of waiting on a worker."
  ],
  "narration": "Under load, traffic used to go straight from the API to the worker fleet, and spikes caused timeouts. There is now a queue in between: the API drops the job and responds immediately with an id, and workers pull from the queue at whatever pace they can sustain. A slow worker no longer means a slow response.",
  "diagram": "graph LR\n  Client --> API\n  API --> Queue\n  Queue --> Worker1[Worker]\n  Queue --> Worker2[Worker]",
  "code": [{"lang": "ts", "snippet": "await queue.add('process-job', payload);\nreturn res.status(202).json({ jobId });", "ref": "src/api/jobs.ts:34"}],
  "topics": ["architecture", "queues", "scaling"],
  "difficulty": "intro",
  "importance": 3,
  "sources": [{"path": "src/api/jobs.ts", "region": [30, 36], "content_hash": "sha256:e63a01ee3a2b7c1d63a9bb14456e733a559f2d1192216a994827f6e8efa69f02"}]
}
```

No event here — this is a good explanation of something that now exists, not a new decision or change worth logging on its own.

### 4. quiz-card — a non-bug

Session: traced a fetch that appeared to fire twice in development back to React StrictMode's intentional double-invocation of effects, before it got "fixed" into a real bug.

Item:
```json
{
  "item_id": "1f9b6d3a-8c45-4e21-b7f0-3a5c9d2e6f81",
  "template": "quiz-card",
  "title": "That double fetch in dev is not a bug",
  "bullets": [
    "StrictMode intentionally mounts, unmounts, and remounts components in development to surface missing cleanup.",
    "The extra useEffect call you saw locally will not happen in production.",
    "The fix, if one is needed, is a cleanup function, not removing the effect."
  ],
  "narration": "You noticed a fetch firing twice in dev and went looking for a bug. There isn't one: React's StrictMode deliberately mounts, unmounts, and remounts components in development to catch effects that don't clean up after themselves. It only happens in dev, and it only matters if your effect breaks when run twice. Worth remembering before someone tries to fix this again.",
  "quiz": {
    "question": "A teammate sees this effect fire twice in dev and opens a bug report. What's the right response?",
    "options": [
      "Wrap the fetch in a ref guard so it only ever runs once, even in dev",
      "Check the effect is safe to run twice, add cleanup if it isn't, then move on",
      "Move the fetch out of useEffect and call it directly in the render body",
      "File it as a React bug and pin the app to a version before StrictMode"
    ],
    "answer": 1
  },
  "topics": ["react", "strictmode", "debugging"],
  "difficulty": "intro",
  "importance": 2,
  "sources": [{"path": "src/components/Feed.tsx", "region": [38, 44], "content_hash": "sha256:4aa9def979b54d5575acd8fb83b9d816fb2e65fc059d57c52fe71499b60ac077"}]
}
```

Event:
```json
{
  "type": "learning",
  "summary": "StrictMode double-invokes effects in development only; the double fetch was not a bug.",
  "why": "The effect was almost changed to suppress correct, intentional double-mounting.",
  "entities": ["React StrictMode", "useEffect", "src/components/Feed.tsx"],
  "code_ref": "src/components/Feed.tsx:41",
  "importance": 2
}
```

### 5. nothing durable — the right answer is silence

Session: ran the test suite, let a formatter reformat two files, renamed a local variable for clarity. No commits beyond a WIP snapshot. Nothing decided, nothing shipped, nothing discovered.

Output: nothing. No item, no event, no MCP call. Formatting and a rename aren't durable knowledge — submitting an item here is just noise in someone else's feed. Say so in one line and stop.
