Activity log

2026-09

2026-09-14 update The translation captures removed from the memory bank, and pi's memory extension kept to the TUI

The user decided both questions the previous entry left open: the facts are not wanted, and the extension has no business running when pi is driven by a script. Neither change touched a page in this vault.

The bank cleanup went through Hindsight's REST API, since the memory MCP server has no delete. Deleting a document cascades to the facts extracted from it and to their links. The candidates were every document tagged project:ru-regen-0914 or project:ru-batch-0914, and each had to have a pi- id and no tag besides pi and its project, so nothing another writer had put under those tags could go with them; all 140 matched and were deleted, taking 2449 facts. The queue still held 102 retains carrying the same tags, which were cancelled, and nine more already in progress, which cannot be. A background job waited for those nine, the longest of them stuck since 07:15, and deleted the 10 documents and 234 facts they produced once they finished. Both tags now have no documents. One consolidation that began at 06:33, before any of this, still lists both tags among five it will refresh; it creates no documents, and it is shared with other tags, so it was left alone. In all, 150 documents and 2683 facts came out.

The fix is in ai-workspace, commit bafe67f, synced to this machine with ai-sync sync pi (doctor: clean). hindsight-memory.ts now runs its index, its recall and its capture only when ctx.mode is tui, the same gate herdr-agent-state.ts already used. A pi started with -p, --mode json or --mode rpc makes no request to the memory server unless HINDSIGHT_HEADLESS=1, while the memory_recall and memory_retain tools and /memory recall still work, because those run only when something asks for them. The Claude Code pi plugin was never affected; it already starts pi with --no-extensions. A new test drives one whole turn against a stubbed fetch and counts requests: none in print, json or rpc mode, four in the TUI. Five mutants each removed one piece of the gate, and the first run left one alive: taking the gate off capture turned nothing red, because the recall handler returned before it recorded the prompt, so capture had nothing to send either way. The prompt is now recorded before the mode check, each of the two paths has a gate of its own, and all five mutants fail the suite.

Measured after the sync: a one-word pi -p through the full profile takes 5.3 to 6.1 seconds, down from 13 to 14, and leaves no document in the bank. With only this extension loaded it costs 4.7 seconds against 2.5 for none, where it cost 15.9 against 1.2 before; the rest of the profile's extensions account for the remaining difference. translate-page.sh was not changed.

2026-09-14 update Russian translations — the backlog after both waves, 127 pages, and why pi takes 14 seconds to say "да"

The queue after the reading-list ingest was 125 pages: 9 stale translations and 117 missing ones, less deepseek-v4-roleplay-instruct, which stays in skip.txt. run-batch.sh took it in batches of 50 with two workers: 2026-09-14T1017 (50 pages, 722 seconds), 2026-09-14T1029 (37 pages) and 2026-09-14T1044 (38 pages, 571 seconds). The middle batch is short because the loop was stopped partway through when the run was called slow, and then resumed with the same settings; the two pages that were mid-attempt had written nothing and were picked up by the next batch. All 125 came back OK. One needed a second attempt, swr-bench-code-review, for a 161-character summary.

Two of the four pages the reading-list entry above names as silently out of date really were. mixture-of-experts and local-ai-is-not-opus had been translated at 09:55 from a revision already dated 2026-09-14, and the ingest added cross-references to both at 10:08 without moving the date, so warren translations read the old versions as current. Both were deleted and regenerated (2026-09-14T-xref-late, 2 pages, 35 seconds). The other two, qwen-gpt-reasoning-prefills and intelligence-vs-cost-linear, were new pages that got their first translation at 10:29 and 10:43, after the cross-references, and needed nothing. The date check has no way to see this case; comparing the commit times of the original and its translation does.

Coverage is 866 of 867, stale 0, and the one missing page is the skipped one. warren check reports no undeclared dangling links and warren lint no errors. The binary is the 1 September build again, six September commits behind its repository; a build of HEAD gives byte-identical translations and check output on this tree, so the gates stand, but the newer commits add the people/ exclusion to the scan, and the old binary run against the main checkout, where the private people/ directory lives, would index it.

The humanizer-ru scan over the 125: median 100, p10 88, minimum 56; 121 «чисто», 3 «правка», 1 «рерайт». The «рерайт» is interconnects, charged mostly for being a list of links (75% of its lines) and for noun density. Hard bans on 13 pages, 16 in all. Eleven are the English's own constructions — "one to three times a week", "6 to 18", "from about four months to about ten", "100 to 600+", "64 to 224 bytes", "from trivial to cross-package", "not only" in olmo and telnyx, "worth keeping in mind" in the reading list — and one is the scanner reading «воздержаться от вмешательства до тех пор» as a range. Four were added by the model, and none of them calls for a change to an English page. "Architecture as well as hardware" became «не только за счёт "железа", но и благодаря архитектуре» in qwen3-8-flash-next — the second time today that "as well as" came back as that frame. "To block threats instead of only watching them" became «не просто наблюдать, а блокировать» in lee-holloway-biography, against a prompt rule that already names "instead of just X". "Truly open is doing work in the name" became «играют важную роль в названии» in atom-project-american-truly-open-models, and «Об этом важно помнить» in interconnects has no English counterpart. They are candidates for system-prompt.md; the prompt was not changed. The sample was pigeon, llm-distillation and six-months-to-live-for-open-models: Latin-script terms (subagent, soft labels, logits, post-training), ё where it belongs, spaced hyphens, code untouched. No page carries the ё corruption fixed earlier today.

Why the run was slow. A page took 28 seconds against 13 to 16 on 1 September; the new pages are larger (8 KB against 4 to 5), but seconds per kilobyte had not changed and there were no retries, so the model was not the difference. A one-word prompt through pi -p took 13 to 14 seconds; with --no-extensions, 1.6. Loading the extensions one at a time put all of it on hindsight-memory.ts, the memory extension in pi's shared profile (15.9 seconds alone; every other extension 1.6 to 3.4). Every pi -p call is a new session to it, so on every page it builds the project memory index, runs a recall with the whole English page as the query (a four-second timeout, and the results go into the translating model's context), and after the answer waits up to ten seconds to retain the prompt and the translation in the shared bank bigbes. The extension is dated 11 September, after the fast batches. The retains are visible in the bank: 1990 facts tagged project:ru-regen-0914 and 383 tagged project:ru-batch-0914, retellings of the translated articles in Russian and lines like "Assistant produced a Russian translation of the note", against 83 for this vault's own memories — and the bank's own comment in that extension measures a recall slowing to 56 seconds while it digests a bulk import. Nothing was changed: passing --no-extensions from translate-page.sh and removing those documents from the bank are both waiting on the user.

2026-09-14 update Russian translations — 85 stale pages after the Telegram waves, and a ё rule that was corrupting words

The two Telegram waves added cross-references to 85 English pages and bumped their updated:, so the pre-push hook refused master with 85 stale translations. Only those 85 were regenerated: select-batch.sh orders missing and stale pages together by inbound links and would have pulled in new pages from the backlog, so translate-page.sh ran over an explicit list with the usual two workers. 1063 seconds, 12.5 seconds a page, all 85 OK. Five needed a second attempt, each for a defect the checker already knows: a transliteration twice ("дефолтный", "фичи"; "деплоить"), unparseable YAML in title: or summary: twice, and a 161-character summary once.

The humanizer-ru scan over the 85: median 97, p10 88, minimum 70; 82 «чисто», 3 «правка». Hard bans on seven pages, and none of them calls for a change to an English page. «От начала до конца» in reaction, not-understanding-your-codebase and zstd-lean-proof-automation is "top to bottom", "end-to-end" and "end to end" — idioms the scanner reads as false ranges. «От восьми до десяти часов» in agentic-coding-fatigue and «от радости до отвращения» in anti-llm-discourse are ranges the English has. «Стоит отметить отдельно» in agent-memory-components is "worth naming separately", and «не только вам, но и всем» in peril-of-laziness-lost is "not just you but everyone", the example audit.md already cites. Two constructions were the model's own: "Note that" became «Стоит отметить, что» in anti-llm-discourse, and "pays off mechanically as well as cognitively" became «не только когнитивно, но и механически» in peril-of-laziness-lost. Two sentences in 85 pages is not yet a pattern, so they are recorded here as candidates for system-prompt.md rather than added to it.

Pairing those findings with the English turned up a defect nothing was checking for: «свидётельства» in anti-llm-discourse. The ё rule in normalize-typography.sh, added on 1 September, replaced its closed list anywhere in a word, and a listed spelling inside a longer word is a different word — "идет" inside "свидетельство", "увидеть" and "сидеть", "еще" inside "вещей", "размещение", "смещения", "помещения". A scan of all of ru/ for those shapes found 24 corrupted pages. 19 came from this batch; the other five — bytecode-to-source-mapping, enshittification, golang-maps-swiss-tables, recall-to-judgment and redis-cost-of-ambition — had been regenerated after the rule landed and were already published. The checker could not see it, because the normaliser runs before the checker and the result is valid Russian-looking text with correct structure.

The substitution now needs a word start. macOS awk counts ё as two bytes and fails on substr inside a multibyte character, so the check could not be written there; the ё pass moved to perl with a Unicode lookbehind, and the dash and quote passes stay in awk. Endings are still allowed ("учета" becomes "учёта"), which means a prefixed form such as "придет" is now left without its ё — the cheaper mistake of the two. A test page with the corrupted words, inline code and a fenced block fails against the old script with exactly the words above and passes against the new one, and a mutant with the lookbehind removed fails it again. The 24 pages were deleted and regenerated rather than patched: 333 seconds, all 24 OK, one on a second attempt for "дефолтным", and the corpus-wide scan now finds no corrupted word.

Coverage is 750 of 820 (91.5%), stale 0; the 70 missing are the pages the two waves created, which is the next batch and does not block a push. warren check reports no undeclared dangling links and warren lint no errors.

2026-09-01 update Russian translations, batches 4-6 — 147 pages, a skip list

Three batches back to back on -medium, one commit each: 390, 381 and 364 seconds, 15 seconds a page — faster than before, because typography no longer costs a second call. 146 pages clean and one kept with a WARN (an "Инстансы" the blacklist did not stop), 7 retries in 147 pages; coverage is 297 of 751 (39.5%), warren check and warren lint exit 0.

The three FAILs were one page, three times: deepseek-v4-roleplay-instruct carries a Chinese system prompt in a code block with literal <think> tags, and the model strips the tags on every attempt — nine in all, because the most-linked page still missing is picked first in every batch. tools/translate/skip.txt now lists it with the reason and select-batch.sh leaves it out; a page goes there only after failing three times for a reason no prompt rule fixes.

What the retries were, now that the manifest keeps them: four of seven were code the model altered — an arrow in pseudocode, a # comment line, alignment — all caught by the code-block diff and clean on the second call; one was a wikilink the model invented (xz-utils-incident in maintainer-governance-ambiguity, which never links it), the one defect that would change the vault's graph if it got through; and two were transliterations from the blacklist that the model wrote anyway ("фичи", "перформанс"), kept as WARN. The sample from each batch read as the earlier ones did.

The push of these batches was the pre-push hook's first real catch, and a wrong one: it read the working tree, where an uncommitted edit had bumped enshittification's updated:, and refused a push of commits that did not carry that edit. The hook now exports the commit being pushed with git archive and checks that. The edit itself is right to be caught — once it is committed, the Russian enshittification is behind its original and has to be regenerated before that push.

2026-09-01 update Russian translations, batch 3 — back on -medium, and typography stops costing retries

-high had bought nothing visible for more retries, so the default model is -medium again. 50 pages, two workers, 521 seconds, 20 seconds a page; 48 landed on the first run and the two that failed landed on a second, so coverage is 150 of 751 (20.0%). warren check and warren lint exit 0.

The runner now keeps every attempt's verdict in the manifest, not just the last, and the first batch with that history settled a question: six of the seven retries were the same defect — an em dash or «ёлочки» carried over from the English into Russian prose on the first attempt and fixed on the second. That is a deterministic rule, not a judgment, so it is now applied rather than asked for: normalize-typography.sh runs over the prose of every result before the checker (code untouched), and a retry for typography — another model call, another slice of quota — no longer happens. habr-local-llm-quantization-deep-dive, which failed batch 2 on the punctuation of its Russian quotations, went through on the second attempt with the new prompt rule and would go through on the first now.

The normaliser cost one round of its own: a rule that squeezed " - " to " - " also squeezed the two-space indentation of a YAML list item, and both retried pages failed on "a frontmatter value changed" until it was removed. The other new failure was flue: the model collapsed the run of spaces that aligns a trailing comment in a code block, three times; the prompt now says whitespace in code is part of the code, and the page went through on the first attempt after that. thinking-mode-rule-erosion, the other page that had failed three times, was the same em-dash case and is clean now.

2026-09-01 update Russian translations, batch 2 — the next 50, on gemini-3.7-flash-high

Batch 1 cost about 6% of the antigravity quota for 51 pages, so the thinking level went up rather than down: from this batch the default model is lllm/antigravity/gemini-3.7-flash-high (antigravity carries the level in the model id; translator: in each translation records which). 50 pages, two workers, 513 seconds, 19 seconds a page — the same speed as -medium. 49 landed; coverage is 100 of 751 (13.3%), warren check and warren lint exit 0.

Two differences from batch 1. Seven pages needed a second or third attempt where batch 1 needed one — charlie-daemons and i-am-an-ai-hater three each, kv-cache-sizing, moe-cpu-offload, botctl, simonw-vibe-coding-agentic, claude-emotion-concepts two — all clean in the end; the manifest keeps only the final note, so what failed first is not recorded, which is worth changing. And one page failed outright: habr-local-llm-quantization-deep-dive is a summary of a Russian article and its English original already carries 24 lines of «ёлочки» and em dashes inside quoted passages; the model kept that punctuation three times over. The prompt now says quoted passages are re-punctuated like everything else, and the page stays in the queue.

The sample read as batch 1 did: "Демоны" (Daemons) от Charlie Labs, YAML-frontmatter'ом, Мозер with the argument's three levels laid out as in the English. Nothing in the fifty pages made the -high output visibly different from -medium; the quota cost per page is the number to compare next.

2026-09-01 update The translate skill and the coverage gate

What batch 1 taught is now written down where the next run will read it, and the question "is everything translated?" has a command. The translate skill in .claude/skills/ is the procedure — where things stand, run a batch, check it, commit and log, what to do before a push — and carries the list of what Gemini gets wrong with what was done about each: transliterations that only an explicit blacklist stopped, YAML broken by an unquoted colon or a quote inside a title until the prompt said "always double-quote", Russian summaries running past 160, an em dash inside a code comment that the model was right to keep and the checker wrong to reject, and the zsh status trap that hid ten results.

warren translations --src . --list (warren branch translations-cli) measures coverage of wiki/ and toolbox/: 751 pages, 51 current, 0 stale, 700 missing, 6.8% — and lists the missing most-linked first, which is the next batch's order. It exits non-zero on a stale translation. tools/translate/pre-push.sh is the vault's git pre-push hook: on a push to master it runs that and warren check, so an English edit cannot go out with a Russian version that says something else and a dangling link cannot go out undeclared; missing translations are the backlog and do not block. CLAUDE says so under Translations. Installed with ln -sf ../../tools/translate/pre-push.sh .git/hooks/pre-push.

2026-09-01 update Russian translations, batch 1 — the 50 most-linked pages

The first batch of the Russian version of the vault, by the rules in tools/translate/README.md: the 50 originals in wiki/ and toolbox/ with the most inbound links, one pi -p --no-tools call each to lllm/antigravity/gemini-3.7-flash-medium, every result through check-translation.sh, nothing edited by hand. Two workers, 564 seconds for the run, 18 seconds a page on average and 135 for the longest (watchlist, a 73-row table, which took two attempts). All 50 landed in ru/wiki/ and ru/toolbox/; with the pilot from the morning that is 51 translations. warren check and warren lint exit 0.

The batch taught the checker three things it did not know. A summary: with an unquoted colon, or a title: with a quote inside it, is invalid YAML — the page then serves with no properties at all — and the checker had compared key names only, so four such pages passed it and were caught by warren lint's translation-meta. It now parses the block with yq, and the prompt says to double-quote title: and summary: always. An em dash inside a code comment is copied verbatim from the original, as it should be, and the typography check had read the whole file, so hazmat failed three times on a line it was right to keep; the check now skips fenced and inline code. And translation-page.sh had a variable named status, which zsh treats as read-only — the script died after writing the file and before reporting it, ten pages went unlogged, and the same trap is already written up in the global CLAUDE.md. Those five pages were regenerated, all clean.

What the sample showed: register and terms hold — "Эндрю Келли (Andrew Kelley)", wikilink'и, Skill'ы, shim, and llm-wiki-pattern and zig read as written rather than translated. The transliteration blacklist added after the pilot removed the "раннер"/"шим"/"селф-хостед" class entirely: no WARN in the final set. Still open: summary: is often longer in Russian than the English it came from and sits near the 160 limit; the search UI does not yet show a hit's language.

2026-09-01 update The log split by month

wiki/log.md had reached 282 KB and 237 entries — past what a reader, an editor or a diff takes in at once, and past the Read limit of the tools that maintain it. The entries now live in one file per month, wiki/log/YYYY-MM.md, each with type: log: 2026-04 (123 entries), 2026-05 (60), 2026-07 (51), 2026-09 (this file). The split is mechanical — entries verbatim, in the order they stood in the old file — and log becomes the way in: the format, the header a new month file starts from, and the list of months. Nothing was renamed, so [[log]] still resolves.

CLAUDE says where an entry goes now (the month it is written in, never moved between files); the ingest, telegram-ingest and lint skills say the same. wiki.base excludes type: log from the Stale Pages view, since the month files would otherwise sit at the top of it forever.

warren needed almost nothing: it already split every type: log page into entries and grouped /log by month. What changed (warren branch log-months): the /log header now links to wiki/log.md by name rather than to whichever log page was scanned first, and the index test covers a month file contributing its entries alongside the old one.

2026-09-01 update Schema: ATX headings and fenced code only

A review of what warren's search chunker loses on the vault found two markdown forms it does not read: setext headings (Title underlined with === or ---) and code indented by four spaces. Measured over 1152 pages with fences and list continuations excluded, the vault had no setext heading at all — every candidate was a ```yaml fence followed by a YAML --- — and one indented block, the Beaver identity on mpc-beaver-triples, now fenced. Rather than teach the chunker either form, the schema forbids them in the authored sections: a paragraph followed directly by --- is a setext heading to CommonMark and Obsidian but a paragraph to warren, and an indented block cannot be told from a list continuation without a parser and splits at its first blank line when chunked. warren lint reports both as heading-setext and code-indented; sources/ is out of scope as always. CLAUDE.md carries the two conventions.

2026-09-01 update Russian translations, batches 8-16 — the rest of the vault, 403 pages

Nine batches back to back on -medium, one commit each: 359, 395, 364, 402, 353, 322, 331, 284 and 21 seconds, 14 seconds a page. 403 pages, 401 OK, two kept as WARN, one FAIL. Coverage is 750 of 751 (99.9%), stale 0, warren check reports 0 undeclared dangling links; warren lint's only error is the code-indented block in mpc-beaver-triples, an English page that carried it before any of this. The one page still missing is deepseek-v4-roleplay-instruct, which is in skip.txt.

Two of the three defects were the checker's, not the model's, and both were false positives on the same blacklist entry. ai-is-not-conscious-chiang was flagged for «шимпанзе» — \bшим[а-я]* has a boundary only at its start, so the ape matched the transliteration of "shim". freeink was flagged for «ШИМ», which is the Russian abbreviation for PWM and matched because the whole blacklist is searched case-insensitively. Both translations are right and neither was touched; check-translation.sh now checks «шим» apart from the rest, case-sensitively and bounded at both ends, and a mutant carrying «шимом» and «кастомного» still comes back WARN, so the check can still fire.

The one real FAIL was tweakcc, where the model translated the # comments inside a shell block («# Configure interactively» became «# Интерактивная настройка») three times over. The prompt already said comments inside code are left alone; what it lacked was the case stated as a case, so it now carries that example verbatim and says the rule covers #, //, --, ; and /* */ alike. The next batch is the evidence: tweakcc, which had failed three attempts out of three, came back clean on the first.

The same treatment did not work for deepseek-v4-roleplay-instruct. A rule naming its exact defect — a literal <think> inside a fenced block is a character of that block — changed nothing across three more attempts, twelve in all, so the rule was reverted rather than left in the prompt as decoration, and skip.txt records that it was tried.

The humanizer-ru scan over all 403: 367 «чисто», 36 «правка», none «рерайт»; median 95, p10 85, minimum 67. Hard bans on 22 pages against 25 of 297 last time. The classifier splits the 91 flagged sentences into 61 that the English original carries and 30 the model added — 0.074 model-added sentences a page against 0.138 in the pre-rules corpus, so the two rules from the audit are holding. What is left of the 30 is mostly the classifier's own doing («детектор ключевых объектов» is a salient-object detector, «ключевые слова» are keywords), and one genuine gap: "much of X" kept coming back as «значительная часть X» because the rule's replacements («много», «заметно») do not fit the partitive. The rule now names «большая часть X» for it. That change is unverified — the queue is empty, so no batch has run under it.

Reading the sample turned up something the checker does not look for at all: ё. freeink said «состоит из трех уровней». Across all 750 translations it was 16 occurrences on 8 pages — «еще», «трех», «объем», «надежность», «счет», «идет» — a rate of one page in a hundred, and deterministic on a closed list of words that have no е-spelling. So it went where the dashes went: normalize-typography.sh restores ё in prose, leaving code and inline code alone, and staying away from «все»/«всё» and «счет»/«счета», where a blind substitution would invent the wrong word. The eight pages were regenerated rather than patched, all eight clean on the first attempt, and the corpus now has none of those spellings left — including «счет», which the model wrote as «счёт» of its own accord the second time.

The sample read as the earlier batches did: zsh-glob-qualifiers, tcp-congestion-control and freeink — a concept, a concept and a toolbox page — with Latin-script terms declined (glob, bash, e-paper), the dash a spaced hyphen, and tables and code carried over unchanged.

2026-09-01 update Russian translations, batch 7 — 50 pages, one retry in the whole batch

50 pages on -medium, two workers, 390 seconds, 15.5 seconds a page. All 50 OK, no WARN, no FAIL; coverage is 347 of 751 (46.2%), stale 0, warren check reports 0 undeclared dangling links. warren lint exits 1 on one code-indented finding in mpc-beaver-triples — an English page, present before this batch and untouched by it — and 102 summary-long warnings, none of them from ru/.

One page needed more than one attempt, and it needed all three: pq-wireguard-revisited invented [[store-now-decrypt-later]], a wikilink its original does not carry, then wrote a summary of 161 characters, then came back clean. Both defects are ones the checker already catches, so nothing new goes into the prompt or into skip.txt from this batch — the first batch since the audit's two rules landed to run with a single retry in fifty pages.

The humanizer-ru scan over the fifty: median 96, p10 88, minimum 81 — better than the corpus calibration of 94 and 86. Two pages below 85, both scanner blind spots rather than prose: marius-blog (81) and sdl3-dos-port (84) are charged only for «номинальность» and «листикл», which audit.md says to ignore on technical text and on pages that are lists by design; their markers and hard_bans are empty. Two hard bans, both checked against the English and both faithful: «от одного до трёх срезов» in why-software-factories-fail is "one to three slices at a time", a real numeric range, and «работают от начала до конца» in zstd-lean-proof-automation is "end to end", an idiom the scanner reads as a false range. Neither is a defect in the English, so nothing was changed there either.

The sample was lazypi, skip-list and marius-blog — a toolbox page, a concept and an entity. Terms in Latin script are declined as the style asks (coding agent'а, skip list'ы, self-hosting'е, homelab'е), ё is everywhere it belongs, the dash is a spaced hyphen, and マリウス.com and its punycode survived the round trip intact.

2026-09-01 update Freshness tiers, a mechanical lint, and the disagreement rule

Read AnastasiyaW/knowledge-space — 851 LLM-written reference cards for agents across 26 domains, MkDocs, a 190-line Python lint — to see what it does that this vault does not. Its content is not ingestable: attribution is stripped on purpose (a pre-commit hook rejects author and course names), so the figures on its pages cannot be checked against anything. Three of its mechanisms were worth taking; the rest — llms.txt, difficulty badges, a checker for hardcoded article counts — either exists in warren already or solves a problem this vault does not have.

Created freshness-tiers: one page carrying a freshness: frontmatter mapping — type, section or tag to a maximum age — the way planned-pages carries planned:. A page's base tier comes from its type:, then its section, then default; a tag shortens a finite base and never revives never, so a concept tagged ai-agents ages in nine months while a summary with the same tag does not age at all. CLAUDE gained a Freshness section pointing at the page, a paragraph in the Ingest workflow on what to do when a new source disagrees with an existing page (work out which claim holds, update, and say on the page what changed and why — never skip because the topic has a page, never overwrite silently), and a first step in the Lint workflow that runs the machine before the reader. The lint skill got the same Step 0.

warren lint is the machine: a new subcommand (warren branch lint) reporting orphans by the catalog/log rule CLAUDE.md has stated since the catalog pages got type: catalog, pages missing from their section's index.md (a parent page's listing counts, for lectures), frontmatter missing a required key, duplicating one, with an unknown type: or a malformed date, or not parsing at all, summary: over the limit, and pages older than their tier. Schema defects gate; long summaries and stale pages warn. First run over this vault: 775 pages, 0 uncatalogued, 19 orphans, 5 pages whose frontmatter did not parse, 426 summaries over 100 characters, 0 stale.

The five unparseable headers were an unquoted colon in summary:fooling-yourself-with-ai, meerkat-introduction, neutrino-1-8b, scylladb-trie-index, zig-incremental-compilation. Quoted; until now warren had been serving those five with no properties at all, which a grep for the key could not show. The planned-pages link in CLAUDE.md was in backticks and so was that page's only inbound link without being one; unbackticked. 18 orphans remain for a later pass: agent-shell, boringbar, dhall, diy-cnc-machine, kokoro, lightroom-cc-on-linux, meetily, openscreen, scotty, slogbox, toolcraft, antonz-org, bitcoin-quantum-computing, cleberg-server-build-summary, go-transaction-linter, httpx2-blessed-fork, jep-540-json-api, since-you-arrived-vol4.

Two numbers to act on later. The 100-character summary limit is violated by 426 of 775 pages (106 over 160); the check warns rather than gates and the limit is a flag (--summary-max), so the rule needs a decision, not a cleanup. And 0 stale today is the vault's age, not its health: the oldest updated: is 2026-04-06, the claude-code tier starts firing in October, and warren lint --as-of 2027-06-01 already lists 208 pages, mostly the 9-month tiers.

2026-07

2026-07-29 update Video support — ytfetch, talks/, and the first talk page

Videos were the one source type the vault could not take. webfetch gets a player shell from YouTube and writes a file containing frontmatter and nothing else, so video links either died at triage or produced a stub that looked ingested. The batch earlier today recorded the failure honestly — the Andrew Kelley talk went into ingested-urls as "video with no fetchable transcript" — but the premise was wrong. The captions were always fetchable; nothing in the vault knew how to ask.

tools/ytfetch is a Go CLI alongside tools/webfetch, same output contract: markdown into sources/, kebab slug, --name override, never overwrites. It resolves a video with yt-dlp, prefers uploaded subtitles, falls back to automatic captions, and falls back again to a local whisper.cpp run over the downloaded audio. Output is prose with a timestamp on every paragraph that deep-links into the video, chapters as headings where the uploader provided them, and a captions: manual | auto | whisper field recording where the text came from.

Two things that were not obvious. YouTube's json3 caption format interleaves each spoken line with an aAppend event carrying only a newline — the on-screen scroll, not speech — and keeping those duplicates the entire transcript; ytfetch drops them, and the VTT path collapses the same duplication by dropping repeated lines. And YouTube publishes machine translations of its own ASR under every language tag, so asking for en on a Russian talk returns translationese with no marker; the -orig track is the real one and is preferred.

webfetch now refuses video hosts with a message pointing at ytfetch, instead of writing the stub. That guard is the actual fix for what happened this morning.

Schema. New talks/ folder and type: talk, parallel to books/ — conference talks, lectures, and recorded courses, with status: and a video_id: dedup key, since the same talk arrives as watch?v=, youtu.be/, and timestamped links and only the id is stable. A multi-lecture course is a parent page plus one page per lecture, because lectures are watched months apart and cover unrelated concepts. CLAUDE gains the Talk Pages section and a Video Ingest workflow; index is the fifth catalog page; the ingest and telegram-ingest skills route video URLs to the new ytfetch skill.

First talk page. kelley-dont-take-the-black-pill — Kelley's argument that software quality was sustained by programmers' accidental bargaining power rather than professional standards, and that despair about losing it is the trap. It sits against enshittification (he keeps one of Doctorow's four checks and drops the rest) and against llm-enshittification (same premises about scraping, opposite conclusion about whether to keep publishing free software). The page cites minute-level timestamps throughout, which is the point of the format — the transcript makes a moment as citable as a page number.

The transcript is machine-generated and the page says so rather than presenting the quotes as verified. talks/ sits outside wiki.base's folder filter, exactly as books/ does, so talk pages do not appear in the orphan and staleness dashboards yet.

Pages created: kelley-dont-take-the-black-pill, index. Pages updated: CLAUDE, index, zig, enshittification, ingested-urls. Tools: tools/ytfetch added, tools/webfetch video-host guard.

2026-07-22 update Declared gaps — `planned:` frontmatter and the check that honours it

Follow-up to the maintenance pass earlier today, which left six wikilinks pointing at pages that do not exist and argued each was a real gap rather than a defect. That argument lived only in a log entry, so warren check still failed on them and the vault had no way to tell a deliberate gap from a typo.

Now it does. planned-pages declares the six in a planned: frontmatter list and explains in a table what references each and why the page is worth writing. CLAUDE.md gains a "Deliberate Gaps" rule under Cross-References: a dangling wikilink is a defect unless some page declares its target, stubs are still forbidden, and dangling links inside sources/ are expected because those are clipped documents pointing into their own repositories.

warren was changed to match — internal/vault parses planned:, internal/linkcheck marks each dangling link declared or not, and warren check now fails on undeclared links outside sources/ rather than on any dangling link at all. It also warns when a declared target starts resolving, so the list gets pruned instead of decaying into a permanent allowlist. Eight tests cover the gate, the sources exemption, spelling variants of a declaration, and staleness.

The vault now passes warren check for the first time: 6,555 wikilinks, 6 dangling, 6 declared, 0 undeclared. Pages created: planned-pages. Pages updated: index (1 entry).

2026-05

2026-05-12 update Add books section

New books/ directory with index.md (Reading / Finished / Want to read / Abandoned). Added book page type to CLAUDE.md with frontmatter (author, year, isbn, status, rating) and content conventions. Linked from index.

2026-05-12 update Promote 15 Eaton speculative picks to individual book pages

The previously inline list under "Eaton bookclub — speculative future picks" is now 15 individual books/<slug>.md pages, all status: want-to-read, with publisher url: in frontmatter. Pages created: garbage-collection-handbook, designing-data-intensive-applications, high-performance-browser-networking, concurrency-works-of-leslie-lamport, fault-tolerant-design, replication-theory-and-practice, art-of-computer-systems-performance-analysis, virtual-machines-versatile-platforms, on-transactional-concurrency-control, feedback-control-for-computer-systems, hackers-delight, transaction-processing, memory-systems-cache-dram-disk, building-a-debugger, readings-in-database-systems. books/index.md speculative-list section converted to wikilinks. Cross-references added to existing wiki concepts where relevant (LSM, distributed-consensus, microvm, syscall-binary-rewriting, hold-on-to-your-hardware, etc.).

2026-04

2026-04-06 update Added Toolbox section

New page type toolbox for interesting projects to try later. Added wirez as first entry. Updated CLAUDE.md schema.

2026-04-06 update Inbox and URL tracking

Added inbox workflow: inbox/ for Obsidian web clips, inbox-processed/ for backup after ingest, inbox-assets/ for clip attachments. Created sources/ingested-urls.md to track URLs fetched from terminal that need manual backup. Updated CLAUDE.md and ingest skill.

2026-04-06 update Added parent field and updated date convention

Inspired by thinking-space's hierarchical metadata. Added optional parent: frontmatter field for nesting pages under broader topics. Added convention to bump updated: date when revising existing pages. Updated CLAUDE.md and ingest skill.

2026-04-06 update Wiki initialized

Created directory structure, schema (CLAUDE.md), index, and log. Ready for first ingest.