Finish the Line is a 60-second Friends quiz: three lines, two fakes, pick the one really said. The owner is about to author 156 more episodes — roughly 468 questions — and has a hard constraint: the pipeline stays LLM-free. This page walks one question from "pick an episode" to "it's live," and marks each step as either a deterministic program or a human judgement call.
The marking is the point of this page — that's what's worth commenting on. If a step you think is free actually needs judgement, or the reverse, that's exactly the kind of note this page is for.
This page is deliberately light while the rest of this design system is dark — it's a document meant to be annotated, and Hypothesis's highlighter is built for dark text on light paper. The other pages stay dark because they document a dark app as it actually ships.
Each row below carries a stable id, so a comment anchors to one specific step rather than drifting if the page changes later.
| # | Step | Marking |
|---|---|---|
| 1 | Pick an episode that's readyIt's a list — the next episode in the queue that has everything it needs. | Free |
| 2 | Pick which clip to useAn episode's script has hundreds of lines. Choosing the handful of seconds that will read well as a quiz clip — one clear speaker, not buried under overlapping dialogue, not already a line everyone can recite on sight — is a taste call. Nothing ranks clips by how good a quiz question they'll make. | Judgement |
| 3 | Fetch that clip's subtitlesA download, using yt-dlp. | Free |
| 4 | Find which lines are in the clipPlain text matching between the subtitles and the aired script. | Free |
| 5 | Pick which line makes a good questionSeveral lines from the clip are candidates. Deciding which single one is worth building a question around — a complete sentence, not a fragment, not something that gives itself away by how it's phrased — is a taste call a script can't apply. | Judgement |
| 6 | Prove the line is word-for-word in the aired scriptA validation pass (validate_episode.py) checks the chosen line matches the script exactly before anything is built on top of it. | Free |
| 7 | Generate candidate fakesnlpaug plus plain slicing produces a shortlist of near-miss variants of the real line — roughly ten per line. See the measured evidence below. | Free |
| 8 | Pick the two best fakesA person reads the ~10 generated candidates and keeps the two that are wrong
in a genuinely tempting way — close enough to make someone pause, not so close it's basically
a duplicate, not so far off it's obviously wrong. This is the step the measured evidence below
is about.
Invented example — not from the show
Original: "I already told the landlord about it." A plausible fake keeps the sentence's
shape: "I already called the landlord about it." A nonsense fake changes what happened:
"I already forgot the landlord's number."
|
Judgement |
| 9 | Write the cueA template — "Chandler says…" — filled in from who actually said the line. | Free |
| 10 | Write the reveal, differing words boldedA difflib diff between the real line and each fake decides which words get bolded — no one writes this by hand. | Free |
| 11 | Write the trap noteThe trap note explains, after the reveal, why the fake was tempting. A generated version can state the fact but reads flat — it needs the product's own dry voice, and voice isn't something a deterministic step produces. | Judgement |
| 12 | Write the scene setupThe setup line is quoted straight from the aired script — a lookup, not a composition. Whether it reads cleanly out of context sometimes needs a light polish; that smaller judgement never changes which line is quoted, which is why this step counts on the free side of the tally above. | Mostly free |
| 13 | Find the exact second for the clip linkMatches the chosen words to caption timings to produce a timestamped link. | Free |
| 14 | Validate the finished episodeAn automated check of the completed episode before it goes anywhere near publishing. | Free |
| 15 | Bundle and deployPackaging and shipping the finished episode. | Free |
Steps 2, 5, 8 and 11 are the ones that don't reduce to a program: which clip, which line, which two fakes, how the trap note reads. Every other step — 11 of 15 — already runs as a deterministic tool with no person in the loop.
These numbers are what make this page worth annotating rather than a diagram of intentions.
| Coverage | The fake generator ran over all 504 existing lines and kept 5,117 candidates — about 10 per line. |
| Agreement with a human | 47 lines got a machine candidate that exactly matched the fake a human had already written by hand. |
| Hand check | A 120-item manual check found 34% of candidates plausible and 66% nonsense. Most candidates are junk — the open question is whether two good ones per line reliably exist, and that is unmeasured. It's the next thing to settle. |
| A fixed bug | Fixing the tokenizer took garbled candidate output from 44.5% down to 0%. |
⛔ There is no third option for those four steps. Either a person decides, or a model does. LLM-free is achievable — but it means the owner makes roughly six decisions per episode, about 900 across the whole project, each one shaped "pick two from a list of about ten" rather than writing anything from a blank page.
If the hand-check ratio above holds at scale, some lines may not offer two plausible fakes in a shortlist of ten — which would turn a "pick two" decision into "write one yourself." That gap is unmeasured and worth flagging if it changes how many of the ~900 decisions are truly picks versus occasional writing.