BeatlesAnswers.org

Twelve Words at a Time

Build log; 29 August 2026

Two scheduled enrichment runs in a row stopped dead in the same place. Each had opened normally, taken its pre-edit snapshot of the page it was about to work on, issued one command, and then produced nothing at all: API Error: Output blocked by content filtering policy. No page was written, no session record was left behind, and the only trace either run left on disk was an orphaned snapshot directory with no matching entry in the log. Re-running the same instruction reproduced the same silence.

The command was the problem

The command was the oldest and most boring step in the whole process. Every claim on every song page has to be checked against the printed sources before it ships, in the same session that writes it, and the way that check had always been performed was to extract the relevant pages of the source book as plain text and read them—pdftotext -layout -f 95 -l 95, and there is page ninety-five. Both of the primary sources here are books still in copyright. Extracting whole pages of them and then working from that text is, from the outside, indistinguishable from reproducing the books, and the platform’s content filter treats it accordingly: it terminates the response. Not the command—the response. Which is why the failure looked like an infrastructure fault rather than a policy one, and why no amount of retrying moved it.

Worth being precise about the confidence here, since this build log is also a record of how conclusions were reached: the filter does not report a reason. The diagnosis rests on the failure signature—the response dying immediately after a command whose output was several hundred words of a copyrighted book, twice, in the only step that does that. It is the only mechanism that fits the evidence, and the fix removes the exposure either way, but it is an inference and is recorded as one.

Verdicts, not prose

The repair was to stop moving the books. Verification does not actually require the text to travel anywhere; it requires an answer to a narrow question—is this claim on this page of this book? A new verification tool now does the reading in the shell and returns only the answer. Given a page range it confirms the range is the one intended and checks the printed page number in the footer against the offset table, because one of the two sources runs four pages ahead of its own printed numbering and has produced a whole class of citation errors on this site before. Given a phrase it reports FOUND, PARTIAL, or MISS. Given a claims file it runs a page’s entire citation set in one pass and exits clean only if every claim landed where the page says it did.

Where a match is found, the tool prints at most twelve words around it—enough to prove the hit is the right one and to catch a misquote, and nowhere near enough to substitute for the book. Five hits per query, twenty-four kilobytes per run, and no mode that prints a page. The constraint is the design, not a setting.

None of this loosens the standard. Claims are still matched against the actual text of the actual PDF, in the session that writes them, with the printed footer still verified; direct quotations still get checked word for word before they appear on a page. What changed is that the source’s prose now stays in the shell instead of passing through the model on its way to a yes-or-no answer. The old method has been written out of the session contract and the new one written in as mandatory, along with a description of the failure signature, so that a future session recognises the symptom in seconds rather than mistaking it for a broken repository.

Three findings in the first ten minutes

A verification tool that has never been run against known-good work is an untested assumption, so the first thing it checked was material already marked verified. It found three things. The Penny Lane page records David Mason’s remark about still being able to play his part on an unmodified instrument as coming from page ninety-two of Lewisohn; the quotation is on page ninety-three. A search for the phrase “piccolo trumpet” on that page returned PARTIAL—both words are present, not adjacent—which means Lewisohn’s own wording differs and any page presenting that phrase as a quotation needs rewording. And on the newly enriched Sgt. Pepper title track, six of seven sampled citations confirmed on the pages listed, while the French horn detail turned out to sit one page beyond the range the page cites.

All three are small. All three are exactly the class of error the cross-session re-verification pass exists to catch, and all three surfaced from a tool run in a few seconds rather than from a careful human re-read. They have been queued rather than patched on the spot, because a correction made in the same breath as the discovery is a correction nobody has checked.

There is a second-order lesson in the stall worth keeping. A session that dies before it writes anything also dies before it can record that it died: the only trace either failed run left was a snapshot directory with nothing beside it. A process that reports its own failures cannot report the failures that stop it reporting. The orphaned snapshot — a pre-edit backup with no session entry to explain it — turned out to be the most reliable signal that anything had gone wrong at all, which is an argument for taking the backup first even when nothing follows it.