# Infra register — CI / build-pipeline debt

Operator-ruled 2026-08-02: infrastructure and CI-pipeline work items live HERE, not on the
`spt-bs-releases` board. The board carries product surface the operator triages; this register
carries what the gater triages. **Mandate: doyle sweeps this file at every milestone intake and
every release close, and composes ripe entries into waves/milestone riders.** An entry leaves
this file only by being built (link the lane) or being retired with a stated reason.

Entry format: status · origin · what/why · trigger condition (what makes it ripe) · size guess.

Last sweep: 2026-08-27, IO-PARSER (#22) INTAKE (wave map: `IO-PARSER-22-JIT.md` + #22 comment).
Rulings: IR-66 COMPOSED as the parallel hertz rider (the two attach-cell later-needle treatments,
attach.rs:561/:672 — the "rides next intake" note comes due here). IR-52 conditional rider
CARRIED FORWARD on this milestone's docs lane, same terms as WAX-SEAL (depth fix lands iff the
lane touches the docs-site CLI reference generator, else stays open). IR-67 COMPOSED as JIT gate
discipline (no wave battery leans on a `-p spt --bins` leg as integration coverage); the entry
stays open until its named construction fix. IR-55 stays armed with hertz; IR-6 LOCKSMITH
composition still unreported, stays; IR-2 trigger unevaluated this sweep (teardown just cleared
the unlanded-lane field — re-check next sweep). No entries retired; none filed.

Prior sweep: 2026-08-23, WAX-SEAL (#21) INTAKE (wave map: `WAX-SEAL-21-JIT.md` + #21 comment
5390092115). Rulings: IR-52 COMPOSED as a CONDITIONAL W4 rider on the wax-seal docs lane —
the milestone mints new nested `spt seal` CLI verbs, exactly the shallow-render class the entry
names; the depth fix lands iff that lane touches the docs-site CLI reference generator, else the
entry stays open here. IR-2 trigger NOT met (todlando's PR set parked unlanded for the next
golden chain). IR-55 stays ARMED-FOR-CAPTURE with hertz. IR-6 conditional composition from
LOCKSMITH still unreported — stays open. IR-56/57/58/59/60/61 are discipline/craft entries or
await their named triggers; IR-62 is hertz-class test/rig work, unscheduled. No entries retired;
none filed.

Prior sweep: 2026-08-19, CONCIERGE (#183) INTAKE (wave map: `CONCIERGE-183-JIT.md` + #183
comment 5347768797) — rows read from the LOCAL register stack (`4799031→ecd640c→8d4c224`;
origin/main lacks IR-49/50/51 until the chain lands — absence ≠ not-filed). Rulings: IR-50
COMPOSED — dispatched to hertz same day (sink-path-helper-FIRST sequencing per the entry;
`engine_room_bringup_e2e.rs` excluded, its panel sites ride the gated #199 lane), with the
engineroom.rs:145 misnomer as a rider on the same lane. IR-47 = candidate co-rider IF this
chain touches golden.yml/ci-notify, else holds. IR-48 holds (next xtask parity-cell touch).
IR-49 holds (next poolguard touch or first landed-lane takeover request). IR-51 holds —
gated on #199 attribution; a net-off fix must NOT precede attribution. IR-2 trigger NOT met
(queued unlanded lanes exist). No entries retired; none filed — the intake window's product
finding (rc terminal-Exit omission, hertz RCA: proven invariant violation, load excluded
92/92, H2 candidate unconfirmed) routed to the BOARD as releases#201, the correct venue for
product surface.

Prior sweep: 2026-08-18, KEYSTONE (#182) INTAKE — the prior sweep's four composition rulings
EXECUTED (wave map: `KEYSTONE-182-JIT.md`, base main @`8248bc3`): (1) IR-9 pin lane DISPATCHED to
hertz (W0 item 1; interim kitsubito-clippy-authoritative dies when it lands). (2) IR-21 remedy
(2) + IR-39 vs `f24e732` DISPOSED: the branch is a BEACHHEAD (golden.yml prebuild step = the
prebuild-made-a-rule arm, ONE fail-fast site, one ledger row) — hertz rebases it off its
abandoned parent `b483699` (the id-collision draft; content landed renumbered as IR-46 @`8248bc3`)
and lands it in W0; the class remainder (IR-39's SHARED precondition helper + 24-site sibling_bin
sweep, the 33 cross-package build-edge expressions) folds into the test-hygiene lane. (3)
Test-hygiene family DECIDED-ACTIVATED as a dedicated hertz lane inside the KEYSTONE window,
sequenced after W0 (members: IR-13/23/36/37/38 + IR-21 clarity half + IR-39 helper +
twohost.rs:394 doc comment). (4) First-execution-cells discipline restated in the JIT's
golden-head step. Board members #84/#85 ride hertz's W0; #166/#57/#185 are todlando's W1. IR-46
id-collision RULED this intake (renumber-and-land, parallel issuance not authored disagreement;
deployah landed it @`8248bc3`). No new entries filed; no entries retired.

Prior sweep: 2026-08-19, NAMEPLATE (#181) / v0.56.0 RELEASE CLOSE — shipped c91 @`60d74ea`
(tag == golden-tested sha, run 32209922535; one respin, both first-golden reds ruled rig defects).
Filed IR-42 (pool-claim writes / build enforces — BUILT in this same commit, AGENTS.md line),
IR-43 (knock NoReply past its 30s carrier bound), IR-44 (perch-sentinel comment overclaims
preservation), IR-45 (twohost rig home premise; SETTLED + RETIRED
2026-08-19 — pump paths resolve under the per-run temp root, see the entry; #189 corrected on the
board, comment 5337519936). Two cycle findings ruled RUNBOOK-homed and
landed in this commit rather than as entries: the first-execution-cells intake question
(RELEASE-RUNBOOK golden-head intake — name the never-executed cells before the run) and
deployah's sweep-vs-cascade mechanism (RELEASE-RUNBOOK board step — under golden CI the cascade
is DRIVEN via `state <mref> acceptance`, never swept). ⚠ LABELLED HOLE — CLOSED UNDERIVABLE
(2026-08-19): the close commune's batch list named "alchemy create-races"; its content did not
survive the author's context reset and was not recoverable from #181, the JIT records, or memory.
Deployah answered the query: he ran ZERO create ops at the cut (could not have witnessed a create
race), a fresh probe over the milestone window shows no duplicate mints (the 4-issues-in-2s batch
mint is batching, not duplication), and the only surviving trace is doyle's own pre-reset message
naming the item as already-known — a pointer, not a sighting. Item DROPPED; the hole stands as
the record. Deployah's sweep-vs-cascade ship-path trap (#181 comment 5337402639, runbook-homed
above) is explicitly NOT this item's content — do not fold it in. Composition: IR-9's pin lane
(`rust-toolchain.toml` @ 1.96.0) goes to hertz AT KEYSTONE #182 INTAKE per the 2026-08-05
ruling; IR-21 remedy (2) + IR-39's precondition helper compose with hertz's standing
fixture-prebuild-hardening branch `f24e732` — disposition at the same intake; the test-hygiene
family (IR-13/23/36/37/38 + IR-21's clarity half) stays the dedicated post-batch lane candidate,
decision at intake; IR-29's proving run + IR-30's instrument lanes ride the next golden batch.
IR-2's trigger explicitly NOT met (queued unlanded lanes exist: IR-29/IR-30 instruments,
four-arm refusal eprintln, f24e732). IR-14 hygiene movement: `.worktrees/nameplate-asm-2bd36f1`
reaped this sweep (+66.98 GB by FS delta 120.53→187.51; claim `gate-w7-courtesy` base `fd3dc5a`
in main = finished lane; zero inbound reparse points) — stale-lane audit itself still open.

Prior sweep: 2026-08-05, LOCKSMITH tranche-2 (#141) CLOSE — golden run 30971976024 green on all 9
jobs, main ff'd to `0a25b77`, v0.55.0. IR-40 filed (stale-resume-brief + early-informant class).
IR-9's decision trigger FIRED 2026-08-05: doyle read the runner-account versions off this run's
`test` legs and RULED — pin in-repo via `rust-toolchain.toml` @ 1.96.0, kitsubito's clippy leg
authoritative in the interim; pin lane to hertz at next intake (see the entry). IR-41 filed the
same night (queued main run superseded without a record; runbook step 3 corrected in the same
commit). IR-1/IR-4's golden-only steps (link probe both boxes, toolchain
print both legs) had their FIRST EXERCISE here, discharging the "unexercised until a golden run"
caveat at the CI-RIDER LANE STATE foot section. IR-35's re-measure rode the batch (`c65b838`).
Owlery-noun thin lane SCOPED and dispatched to hertz for the next batch (class A only, two sites;
class B on-disk rename is an explicit non-goal — see IR-40's kin discipline for why the boundary is
written into the brief rather than left to judgement).

---

## OPEN

### IR-1 — Quiet predicate needs a network axis (tailscale RTT probe)
- **Status:** open, LANE EXISTS UNLANDED — see [[CI-RIDER LANE STATE]] at the foot of this file;
  carried by `bf8c4a2` → `REQ-CI-LINK-HEALTH-PROBE`, mapping CONFIRMED by builder hertz 2026-08-03
  (by content: tailscale ping ×5, med/max RTT rows into the bench ledger, three arms
  success/NO-REPLY/UNAVAILABLE, always exit 0 — instrument, not gate). NOTE `bf8c4a2` is NOT
  single-purpose: it also corrects the free-space preflight floor read
  (`REQ-CI-FREE-SPACE-PREFLIGHT`, the ci-runner-has-no-warm-target shape) — no 1:1 commit→IR map
  for this commit · **Origin:** golden/bench-wiring red triage 2026-08-02 (ex releases#126)
- **What/why:** the shared-runner quiet predicate (zero non-terminal runs + no local
  cargo/rustc/nextest by parent chain) is process-shaped; both axes passed on a box whose only
  link was degrading (321s for a 1s checkout, bidirectional 10s QUIC dial timeouts). A tailscale
  RTT probe to the peer box before two-host rendezvous, carried in the bench ledger, would have
  called run 30771155390's red in seconds. Evidence: the arm-1 count table (PUMP_PEER_FAIL
  a 0→3→0, b 8→22→8 across green/red/rerun).
- **Permanent, not stopgap:** operator-confirmed 2026-08-02 that kitsubito cannot be provided
  ethernet — wifi-only indefinitely, so the link cannot be hardened and the predicate must see
  link health.
- **Ripe when:** next CI-touching wave, or the next network-shaped golden red — whichever first.
- **Size:** small (one probe step + ledger row + predicate doc).
- **Composed:** LOCKSMITH (#132) CI-rider cluster, hertz thin lane — 2026-08-03. GREENLIT with
  #132 and **DISPATCHED to hertz 2026-08-03** (direct brief; the register is the spec, there is no
  board issue). Leaves the register only when the lane lands or the entry is retired.
- **FIRST FIELD USE, and it DISCRIMINATED (2026-08-04, golden 30873007187 attempt 1):** the probe
  (landed via the rider lane, riding `4b37512`) read 4–7ms RTT healthy on both twohost legs
  minutes before both legs redded — REFUTING the degraded-link read for that red and steering
  triage to the real mechanism (the [[IR-29]] serve-window race) instead of a link chase. The
  instrument's first catch was a correct NEGATIVE — exactly the call run 30771155390 needed and
  could not make.

### IR-2 — Settle the warm-runner CARGO_INCREMENTAL delta
- **Status:** RETIRED 2026-08-28, measurement run, verdict NO FLIP — incremental stays ON.
  hertz paired goldens at main fed965f8: ON run 33167693879 GREEN; OFF run 33172022108 RED
  (Windows input_ack_deadlock SetupFailed at a 3.0s IPC read deadline — the box-load family,
  discarded with the leg). Disk: OFF target 7.49 GB vs ON 11.40 GB = 3.91 GB / 34.3% saved
  (smaller than the entry's 5.65 GB cold figure). Timing NOT causal-quality: sequential runs in
  one persistent workspace let OFF inherit ON's cache (order contamination dominates the
  apparent OFF speedups); Windows build-spt-bin +3.8%, Linux -5.1%. Retire reasoning: the
  DECISION is answerable now — OFF produced no green golden, the timing question needs 4+
  counterbalanced fresh-cache windows to answer cleanly, and the disk motivation has weakened
  (box at ~148 GB free under the teardown discipline; the LNK1318 pressure era predates it).
  Runs/artifacts preserved on the record. Reopen only if runner disk pressure returns as a
  recurring floor-class red.
- **Was:** open · **Origin:** #103/#108 bench-wiring lane 2026-08-02 (ex releases#127)
- **What/why:** the #103 measurement (−29.8% wall, −5.65 GB/target, n=3) is COLD-build only.
  Golden's runner `_work` target persists warm, where incremental is exactly what keeps it cheap;
  CARGO_INCREMENTAL=0 was applied only to the genuine cold build (n1-gate pinned old-broker
  cache) + local rig recipes (docs/GOLDEN-CI.md). Open question: does incremental still pay on
  the warm runner, weighed against 5.65 GB/target on a box with LNK1318 free-space history?
- **Method (hertz):** one full golden each way on a quiet box, outside a milestone, compared
  per-step from the bench ledger.
- **Ripe when:** a quiet between-milestone window with no queued lanes (the measurement burns
  two golden windows).
- **Size:** medium (two proving runs + verdict + possible leg flips).

### IR-3 — Daemon-level guard: broker-net wakeup rate bounded across endpoint churn
- **Status:** open · **Origin:** releases#125 remediation, todlando REQ call 1 (ex releases#128)
- **What/why:** the swarm-discovery GC-spin burned two cores for two weeks visible only in a
  process table — no suite assertion sees the class. Wanted: a daemon-level assertion that
  broker-net workers stay quiescent across repeated endpoint create/destroy churn.
- **Design constraint (pre-ruled):** assert on WAKEUP RATE / voluntary ctxt-switch delta over
  the churn window, NOT %CPU — CPU thresholds flake under CI load; the defect signature
  (~100 Hz per orphan loop) is load-independent. Mint the REQ at activation.
- **Near-product:** this is runtime-defect visibility, the most product-adjacent entry here —
  a candidate rider on any daemon-lifecycle milestone.
- **Ripe when:** the next milestone touching spt-net endpoint lifecycle or daemon supervision.
- **Size:** medium (churn harness + counter plumbing + flake-safe assertion).

### IR-4 — Lock-pin guard + lock-procedure rule + toolchain print (three riders, one lane)
- **Status:** open, LANE EXISTS UNLANDED — see [[CI-RIDER LANE STATE]] at the foot of this file;
  candidate commits `1275e47` + `49d4805` + `8ed006b`, mapping NOT yet confirmed by its builder ·
  **Origin:** releases#125 fix-lane intake hold (ex releases#129 + riders)
- **What/why, three parts that land together:**
  1. **xtask check leg:** assert Cargo.lock resolves swarm-discovery to git rev
     `89a2200d54a4e3cab2f46cc75ebff49a1fb07614` while the patch is load-bearing; the check's
     message states its own drop condition (upstream ships a post-PR#27 release AND iroh's pin
     reaches it). Without it, stanza removal or a routine iroh bump silently returns the lock to
     the spinning crate and nothing reds.
  2. **Procedure rule for lock-touching lanes** (docs): targeted `cargo update -p <crate>` only,
     never full re-resolve; count changed `[[package]]` blocks AND diff per-block edges (set-identical
     hid 8 windows-sys edge movers at acaaa4f); two resolutions disagreeing = toolchain drift —
     stop and compare against CI before shipping either lock; hand-edited lock acceptable iff
     `cargo check --workspace --locked` passes.
  3. **Toolchain-version print step in golden** (cargo/rustc versions, both OS legs): the acaaa4f
     comparison against CI was impossible because no run log prints a version. One-grep audit.
- **Ripe when:** next CI-touching wave; part 1 sooner if any iroh bump is proposed.
- **Size:** small-medium (one xtask leg, one docs section, one workflow step).
- **Composed:** LOCKSMITH (#132) CI-rider cluster, hertz thin lane — 2026-08-03. GREENLIT with
  #132 and **DISPATCHED to hertz 2026-08-03** (all three parts land together).

### IR-5 — Shared nextest summary parser
- **Status:** open · **Origin:** BAROMETER triage (standing recommendation, pre-register)
- **What/why:** two agents in one day wrote `[0-9]+ tests run` parsers that read "1 test run"
  (singular) as zero — a guard fed by a broken parser condemns valid rounds. One shared,
  singular-aware parser (single-source discriminant) for every consumer of nextest summaries.
- **Ripe-when REWORDED 2026-08-04 (todlando audit): the old trigger was unfireable as worded.** At
  `b7b00c3` the in-tree population of nextest-SUMMARY parsers is ZERO — the two scripts that read
  nextest output (`g6-curve.ps1:83-90` per-test lines, `g6-postbounce.ps1:44` display-only grep)
  neither parse counts nor carry the defect, so "next wave touching any gate script that reads
  nextest output" could fire on a non-defective script while the real population (agent-authored
  throwaway parsers, which never enter the tree) stays out of reach. New trigger: **the next time
  anyone — agent or lane — needs a nextest summary COUNT**, the shared parser is built FIRST and the
  need consumes it; rig briefs should name it so throwaways stop being authored.
- **Size:** small.
- **Composed:** LOCKSMITH (#132) CI-rider cluster, CONDITIONAL — lands iff the hertz thin lane
  touches gate scripts; otherwise stays open here. 2026-08-03. Carried in hertz's 2026-08-03
  dispatch brief as a **conditional** rider; hertz reports whether the condition fired. **Not yet
  known to be building** — an unreported condition leaves this entry open, not landed.

### IR-6 — Membership logging on subnet gates
- **Status:** open · **Origin:** BAROMETER triage (standing recommendation, pre-register)
- **What/why:** counts beside results, membership beside counts — gate logs that state a count
  without naming the population keep producing unreadable reds. Standardize membership
  enumeration in gate output.
- **Ripe when:** next wave touching gate scripts / CI legs that report counts.
- **Size:** small.
- **Composed:** LOCKSMITH (#132) CI-rider cluster, CONDITIONAL — lands iff the hertz thin lane
  touches gate scripts; otherwise stays open here. 2026-08-03. Carried in hertz's 2026-08-03
  dispatch brief as a **conditional** rider; hertz reports whether the condition fired. **Not yet
  known to be building** — an unreported condition leaves this entry open, not landed.

### IR-7 — Phase A rigs leak a daemon+brain pair on Windows (exe-lock kills notify relink)
- **Status:** open · **Origin:** BAROMETER post-publish triage (ex releases#124 — full mechanism on the closed issue)
- **What/why:** two Phase A rigs launch daemons that escape the job object via WMI-rung autostart
  (double-space unquoted cmdline fingerprint; SPT_HOME in the wrapper cmdline is the attribution
  key); the leaked pair holds target/debug/spt.exe and kills every golden job reaching the notify
  relink. CI reaps as tourniquet (fa6e597); in-job test launch is the fix.
- **Ripe when:** next wave touching the Phase A rigs or daemon autostart path.
- **Size:** medium.

### IR-8 — reap-census scoped_survivors=0 is blind to unreadable-path holders
- **Status:** BUILT 2026-08-18 on KEYSTONE #182 hygiene lane. · **Origin:** BAROMETER triage
  (ex releases#122)
- **What/why:** a zero that cannot see is not a zero — census scoping skips procs whose exe path is
  unreadable, so the survivors count can report clean while a holder lives. Needs a positive control
  / explicit unreadable bucket in the verdict line (unreadable_path count exists; the ZERO must
  refuse when it is nonzero).
- **Defining specimen (golden 30782259675, hfenduleam test job, 2026-08-03):**
  `CI-REAP summary: killed=5 kill_failed=1 scoped_survivors=0` — an admitted kill failure printed
  beside a zero-survivors claim on the same verdict line. The held image surfaced one step later:
  run-scoped tmp cleanup denied 5/5 attempts on `...\relshell\svcmock.exe` (2nd appearance of the
  svcmock hold; 1st @7a3c08c, pre-kill-auth). hertz's addendum: an image-held survivor also blocks
  WRITES to the exe path — the same class manufactures build/relink access-denied reds that mask as
  build problems, not just cleanup warnings. Not per-run: the same-sha green rerun's leg read
  `kill_failed=0 scoped_survivors=0` throughout (hertz, 30784469908) — intermittent sighting,
  second of its class, not a deterministic fixture property.
- **The Linux twin is strictly worse (todlando audit 2026-08-04, vs `b7b00c3`):** `reap-census.sh`
  has NO `kill_failed` anywhere (0 occurrences vs 2 in the `.ps1`) — its kill loop increments
  `killed` only in the success branch with no else, so a failed kill increments nothing and prints
  nothing. The specimen that made this class VISIBLE on Windows would be INVISIBLE on Linux.
  Population precision so this is not overclaimed: ESRCH is benign (already gone; the kill-time
[…2639ln elided…]
already refusing cannot be re-claimed by the ordinary route. The
  refusal text says so itself — meaning the bootstrap case is known, and a wrong-cwd claim is the
  ordinary way a lane lands in it.
- **The rule, as a must-do:** claim from the lane's own worktree so the identity the verb harvests
  is the lane's — `( cd <lane> && cargo run -p xtask -- pool-claim --pool "$PWD/target" … )`, or any
  form that makes cwd explicit. Then READ THE PRINTED RECORD BACK and confirm `owner_tree` and
  `branch` name the lane: the success line carries the record, so the check costs nothing.
- **Trigger condition:** ripe now — it is a doc/UX fix on a verb we run every lane. Ripest alongside
  any other `xtask` pool-verb touch.
- **Size guess:** small. Either make the verb REFUSE when cwd is not inside the tree containing
  `--pool` (loud, and it cannot be wrong), or print the harvested identity as a distinct confirmed
  line. The refusing form is preferred: it removes the reading step rather than adding one.
- **Built evidence, corrected at review:** `pool-claim` resolves the git worktree containing
  `--pool` from the pool's nearest existing ancestor and compares it with the cwd worktree before
  writing. An accidental cross-worktree call returns 2, names both worktrees and the remedy, and
  leaves the not-yet-existing pool absent. The sanctioned sequential takeover carries an explicit
  `--foreign-pool`, says that it is harvesting the arriving cwd identity, writes that identity, and
  returns 0. Both poolguard refusal remedies now print that flag. An executable xtask test drives
  both arms against two real temporary git repositories, pinning VERB-ADMITS-REMEDY rather than
  merely testing path equality. This remains identity-harvest validation only: the claim still
  writes without adjudicating admission, so IR-42's enforcement boundary remains in the build.

### IR-57 — an assembly pick's conflict resolution can silently DROP lines, and every gate run at assembly time is blind to it

- **Status:** OPEN, owned by **doyle** (gater craft, not a code defect — no lane to dispatch).
  Filed 2026-08-22 from doyle's own defect on the TURNKEY #212 assembly head.
- **What happened:** assembly head `5cce533d` **did not compile**, and was carried across a session
  boundary as a ready head with "treqs exit 0" attached to it. `crates/spt-store/src/access.rs` had
  an unclosed `mod tests`: the cherry-pick of the #206 lane commit `b2194aa9` conflicted against the
  #209 test block — both add functions to the same region, and the two sides share the trailing
  `);` / `}` / `}` — and the resolution deleted the markers without restoring the FIRST side's
  function closer. Exactly two lines lost: `        );` and `    }`.
- **Every instrument in the assembly path reported success:** `git cherry-pick` completed with no
  conflict remaining; `traceable-reqs check` returned 797/797, 0 findings, exit 0 — **it parses
  tags and never invokes the compiler**; and the lanes themselves were green and stayed provably
  clean (the #206 lane's brace balance is correct at all five of its commits). **Lane-green plus
  conflict-free is not a claim about the assembled head.** Same lesson as the #182 v2 assembly
  (`E0425`×4) by a different mechanism: that one was a semantic composition break between two
  correct hunks, this one is a fidelity LOSS during resolution.
- **Two must-dos, both cheap:**
  1. **Compile the assembled head before it is handed anywhere — including forward to your own next
     session.** A head is not ready on a treqs exit; it is ready on a build. Any sentence carrying a
     head sha to another agent must name the instrument that proved it.
  2. **Audit pick fidelity on CHANGED LINES for every pick on the chain**, not just the one that
     tripped you: hash `git show --format= <sha> | grep -E '^[+-]' | grep -vE '^(\+\+\+|---)'` for
     the lane source and for the pick, and compare. Whole-diff comparison is useless here — hunk
     headers and context legitimately drift once the head's copy of the file has moved.
- **Read COUNT and HASH together, because they fail in opposite directions.** Of 24 picks on this
  chain two did not match. `b2194aa9 → 4d6deed1` differed in COUNT (310 lane / 308 pick) — the real
  defect. `f27f15c6 → 5865cc98` had the SAME count and a different HASH — benign: a `CONTEXT.md`
  paragraph the head had already amended for #209, whose merged result correctly carries both lanes'
  sentences. A count check MISSES the first class entirely; a hash check FLAGS the second as if it
  were a defect. Neither alone classifies a pick.
- **Repair shape used, recorded because it is the cheap one:** reset to the commit before the bad
  pick, re-pick the lane source, resolve the same conflict correctly, then replay the remaining picks
  in chain order with the fidelity check on each. Verify the repair changed nothing else with a
  WHOLE-TREE diff against the old head — it must show exactly the restored lines and no other file.
  Where a later pick re-conflicts, check its pre-pick blobs against the old chain
  (`git rev-parse <old-pick>~1:<file>`); if they match, the old chain's post-pick blob is a
  transferable, already-reviewed resolution.
- **AND GATE THE REPAIR'S OWN SUBJECT.** The rebuilt head passed six legs — clippy `-D warnings`,
  sweep 660/660, treqs 803/803, xtask OK — while the restored lines live in `spt-store`, whose LIB
  tests `-p spt --bins` never runs. Six greens, and not one executed the function the repair
  restored; it was proven to COMPILE and assumed to PASS. Proven afterwards by name (4/4) and by the
  full `spt-store --lib` suite (519/519). **A repair's gate must include the suite that owns the
  repaired file**, which is not necessarily the suite the head gate runs.
- **Trigger condition:** ripe at the next assembly — this is procedure, and its cost is one command
  per pick.
- **Size guess:** small as a scripted check in `xtask` (fidelity audit over a pick range); zero as
  discipline, which is how it is being applied now.

### IR-58 — a `[[bin]]` grep is not a census: cargo AUTODISCOVERS `src/bin/*.rs` targets, and a hand-built bin list produced a confident false "this bin does not exist"

- **Status:** OPEN, unowned. Filed by doyle 2026-08-22; **its original filing was WRONG and is
  retracted in full below.** The entry is kept because the way it was wrong is the finding.
- **RETRACTED CLAIM, stated plainly so no reader acts on it:** this entry first said that
  `crates/spt/Cargo.toml:18-21` and `crates/spt/src/cli.rs:30355` forbid a bin-name collision with
  `xlate_choreo_fixture`, "a bin that does not exist", and that spt-daemon's third helper had been
  renamed to `summarizer_fixture`. **`xlate_choreo_fixture` EXISTS** — at
  `crates/spt-daemon/src/bin/xlate_choreo_fixture.rs`, verified at the assembly head. The collision
  rule's counterpart is real, both prose sites are CORRECT, and IR-21's original table row was
  correct too; it merely predated `summarizer_fixture`. Nothing in that prose needs fixing.
- **How the false claim was produced, which is the entry:** the census was built by grepping
  `^\[\[bin\]\]` across `Cargo.toml`s. **Cargo AUTODISCOVERS `src/bin/*.rs` as bin targets with no
  stanza at all**, so a stanza grep cannot see them, and it fails SILENTLY — it returns a clean,
  well-formed list that is simply short. From that list it read as established fact that a named bin
  had been renamed away, and that "fact" was then written into a register table, an IR entry, and a
  message to the builder. **A pattern-built population encodes the pattern's assumption; the count
  looks like a census and is a property of the pattern.**
  (Found by todlando, resolving one of doyle's figures against his own tree under the pointer rule.)
- **A CENSUS IS ALSO A PROPERTY OF A TREE.** Stanza counts differed legitimately across trees the
  same hour — 13 at the assembly head, 11 at `dc1c7532` — because `summarizer_fixture` exists at one
  and not the other. Two correct censuses of the same workspace disagree unless each names its sha.
- **The must-do, and it removes the list rather than lengthening it:** never hand-maintain a bin
  roster. Run `cargo build --workspace --bins` and let cargo enumerate its own targets — the same
  enumeration the test harness resolves against, so it cannot drift from what a fixture lookup
  expects. Stanza-derived lists, remembered pairs, and register tables are all the roster problem at
  different depths; only cargo's own enumeration is the population.
- **Kin:** IR-55's amended prebuild bullet (prebuild the census, not a remembered pair) and IR-42
  (a verb that writes without adjudicating). The shared shape is an instrument returning an orderly
  answer about a set it never fully saw.
- **Trigger condition:** ripe on any `xtask` check work — a duplicate-bin-name check over cargo's
  target enumeration (NOT over manifest stanzas) is a handful of lines and cannot go stale.
- **Size guess:** small. The prose fix originally proposed here is WITHDRAWN — there was nothing
  wrong with the prose.

### IR-59 — pool arithmetic: this box holds TWO cold pools, not four; and a build that exhausts the volume reds as a LINKER defect that names no disk

- **Status:** OPEN, unowned — it is arithmetic and discipline, not a code lane. Filed by doyle
  2026-08-22 from a live disk-floor abort during the TURNKEY #212 W4 gate; footprint figures measured
  by todlando the same hour.
- **SECOND FACE, measured 2026-08-23/24 (WAX-SEAL #21 W2 gate, doyle — the violator this time):**
  both predicted mechanisms fired at once, WITHOUT the floor guard because the gate legs ran as a
  plain script. (1) doyle minted a THIRD cold pool (gate rig) while todlando's lane pool was LIVE —
  the two-pool arithmetic, violated by the gater. (2) The lane pool had silently accumulated
  **104.8 GB** over repeated full-workspace sweeps (nothing bounds or even measures a pool's
  accumulation until the volume dies). The volume hit **3 MB free**; the reds wore exactly the
  entry's predicted costume — `LNK1318: Unexpected PDB error; LIMIT (12)` naming no disk — plus a
  second face worth keeping: `traceable-reqs check` died as a PANIC printing to stdout
  (`os error 112`), so a REGISTRY gate red can also be the disk's. Both gate verdicts VOIDED
  (clippy's full compile had exited 0 before space died — the code was never the question).
  Recovery measured: gate pool reap 9.6 GB, lane pool clean 100.2 GB, drive back to 109.6 GB free.
  **Remedies adopted at the incident:** the gate never mints its own pool while a builder lane is
  live on the box — it takes the lane pool over SEQUENTIALLY at pool-release (warm, loud, the
  releases#103 pattern); and gate/lane scripts get the disk-floor preflight the golden runner
  already carries (this entry's part 2, still unbuilt).
- **THIRD FACE, measured 2026-08-24 (WAX-SEAL #21 golden window, doyle):** two new mechanisms, both
  about the REMEDY rather than the red. (1) The wax-seal-w1 lane pool **re-accumulated to 152–169 GB
  (two meters, IR-46 caveat) within ONE DAY of the 08-23 clean** and rode the volume back down to
  4.92 GB free — a reap is an instant, not a state; nothing yet bounds a pool between reaps.
  (2) **The remedy itself arrived garbled through relay:** the golden push was held "waiting on the
  IR-59 reboot" when this entry prescribes no reboot anywhere — tens of releases have shipped without
  one; the operator refused the premise and the register's own recorded remedy (reap finished-lane
  pools) cleared it in ten minutes. A remedy relayed as a DIFFERENT remedy is the relay trap wearing
  infra clothes: the entry is the author — read it before adopting a hold. Recovery measured: five
  finished-lane pools reaped after ancestry classification (`git cherry` catches rebased lands that
  a plain `merge-base --is-ancestor` calls unlanded — gate-w4l1's 90.9 GB pool was reapable only by
  that meter), 4.92 → 260.88 GB free (~256 GB free-delta vs ~270–287 GB sum-of-lengths). Two main
  runs red inside the low-disk window (09:17Z/09:41Z) were re-run, not re-read, per this entry's
  must-do.
- **FOURTH FACE, measured 2026-08-24 same day (SIGNET #218 W2 gate, doyle — the gater's own pool
  this time):** the MAIN pool went **13.7 GB → 158.3 GB in ONE DAY** absorbing four full gate
  batteries across two lane tips, and dragged the volume to **18 GB free mid-gate** — so the
  entry's mechanism is not a lane-pool problem, it is EVERY pool under repeated full sweeps, the
  gate rig included. Two riders worth the ink: (1) the golden runner's **free-space floor step
  produced its first live catch** — a thin CI leg refused at 14:14 local over exactly this window
  with zero test signal, and the register's re-run-not-re-read rule resolved it in one command;
  (2) the reap-and-rebuild trade ("cleaning a live pool buys one window at rebuild price") was
  taken deliberately mid-gate and cost ~35 minutes cold — cheaper than one uninterpretable red.
  A red measured at 18 GB free in this gate got NO read at all, and the next sweep's red was a
  DIFFERENT test: low disk manufactures random victims before it manufactures link errors.
- **The event:** a gate's preflight refused with `GATE_ABORTED_DISK_FLOOR` at **4.75 GB free on a
  1863 GB volume**. Four cold pools plus the milestone head's pool were standing at once. The floor
  guard did its job — without it a full workspace build would have started with under 5 GB of
  headroom and produced a red belonging to nothing.
- **THE ARITHMETIC, and it corrects the intuition that worktrees are the cost.** Measured across the
  whole `.worktrees` tree: **45 worktrees = 0.79 GB total** (~18 MB each). **3 pools = 130.45 GB.**
  Worktrees are **0.6%** of the footprint; pools are **99.4%**. **One cold pool is worth roughly 2,500
  worktrees.** A rule keyed on worktree count would cost a day of `git worktree remove` work for under
  a gigabyte, risk removing a lane someone still wants, and leave the lever untouched.
- **THE CEILING, as arithmetic rather than hygiene:** at **45–75 GB per cold pool** on this workspace
  against a **40 GB floor**, this box supports **TWO live pools comfortably, THREE only if one is
  small**. The two rules that already exist need no replacement, only this number attached:
  - **Reap a pool when its lane is finished** → 45–75 GB back immediately.
  - **Do not run two builds at once** → not only the contention argument, but **75–150 GB of
    simultaneous allocation** against that floor.
- **A FULL VOLUME REDS AS A LINKER DEFECT.** `LINK: fatal error LNK1318: Unexpected PDB error`
  (variously `OK (0)` or `LIMIT (12)`), often beside `LNK4209: debugging information corrupt`. It
  reads as a corrupt artifact or a broken toolchain and is neither — it is the linker unable to grow
  a large `.pdb`. **Nothing in the failure text says "disk".** Do not key recognition on the
  parenthesised code; it varies.
- **THE DANGER WINDOW IS THE TAIL OF A COLD BUILD, not a steady state you can check once.** The
  2026-08-22 instance died at the link of `spt` — the last and largest artifact of a **74 GB** cold
  pool — while that same build was draining the volume out from under itself. A free-space reading
  taken before the build would have looked fine.
- **Positive tell, with its own refutation attached:** `traceable-reqs check` alone surviving a leg
  table is suggestive, because it is the only leg that never links. It is a POSITIVE tell ONLY — a
  full disk can red before any link step, so the signature's ABSENCE clears nothing.
- **THE FALSIFIER IS FREE SPACE AT THE TIME OF THE RUN**, not the shape of the leg table.
- **Must-do, both cheap:**
  1. **A free-space reading goes in the FIRST LINE of any build-failure report** — not in the
     controls, not as a follow-up. Two agents spent hours on a mechanism for this failure and neither
     ran `df`; the falsifier was one command away throughout.
  2. **Rig and gate logs should record free space at run start.** Today they do not, which makes "was
     the disk full" **unanswerable after the fact for every red already in hand**. Any red produced
     under ~1 GB free is UNINTERPRETABLE and must be **re-run, not re-read**.
- **Measurement caveat, observed twice in one hour:** sum-of-file-lengths and free-space delta
  disagree by several percent when anything else on the box is writing (98.07 GB of subtrees → 92.25
  GB reclaimed; 47.28 GB → 40.81 GB). Report BOTH, derive neither from the other, and treat any single
  free-space number as an instant rather than headroom — see IR-46.
- **FIFTH FACE, measured 2026-08-29 at the CONDUIT #236/v0.65.0 cut (deployah):**
  the finished milestone's sequential lane-sharing accumulated **140.51 GB**, while a freshly
  rebuilt pool measured **7.78 GB** — about **18× one build's working set** retained across lane
  hand-offs with no reap between them. This is the budgetable rate behind that cut's disk-floor
  red: sequential execution prevents concurrent ownership corruption, but it does not bound
  historical artifacts inside the shared pool. Treat takeover and reclamation as separate
  operations; a pool may be safe to reuse and still be too large to keep.
- **Composed for next intake (doyle, v0.65.0 close):** the still-open log-the-floor build half
  rides with [[IR-57]]'s scripted pick-fidelity audit as the next milestone's two tooling riders.
- **Trigger condition:** ripe now for the log-the-floor half (it rides any gate-script or CI touch);
  the arithmetic is discipline and applies immediately.
- **Size guess:** small — one line in each rig/gate preflight to record free space alongside the
  existing floor check.

### IR-60 — a wrapper that CAPTURES a leg's exit status ends with the capture, so the wrapper's own status is the capture's, not the leg's

- **Status:** OPEN, unowned — script-shape discipline, not a code lane. Filed by doyle 2026-08-22
  during the TURNKEY #212 W5 lane; mechanism measured by todlando, who reported it first as a
  possible harness misreport and then RETRACTED that reading himself on a minimal probe.
- **The reading that started it:** a background leg was summarised as `exit code 0` while the leg's
  own exit file read `100`. Reported as-is, that is a task runner lying about a red — the worst
  possible direction for a purposeful-red claim, since it would turn two reds that DID fire into two
  reds reported as not having fired.
- **The measurement, minimal and on this box:**

      ( exit 100 ); echo $? > probe.exit; FINAL=$?
      leg status captured to file : 100
      wrapper's own final status  : 0

  The wrapper's shape is `( cd <lane> && cargo nextest ... > leg.raw 2>&1 ); echo $? > leg.exit`.
  **The last command is the `echo`, which succeeds**, so the wrapper genuinely exits 0 while the leg
  genuinely exited 100. THE RUNNER TOLD THE TRUTH ABOUT THE WRAPPER. There is nothing to file against
  the task runner, and it would have been filed.
- **The class, and this is why it is worth an entry rather than a fix:** it is the truncation-pipe
  family — the meter measured something real, it just was not the thing about to be quoted. **Three
  shapes now share ONE trigger, and the trigger is not "is this a pipe":**
  1. a truncation pipe (`cmd | tail`, `cmd | head`) — `$?` measures the truncator, and `head` closing
     the pipe can MANUFACTURE a status (SIGPIPE → 101), so `PIPESTATUS` is not protection either;
  2. a flattened `$?` read after any intervening command;
  3. **a wrapper whose last command is the capture itself.**
  **The trigger in all three is the REPORTING** — every instance happened while trimming or capturing
  output in order to QUOTE it as evidence.
- **The practice that held, and the one to keep:** the leg's own file is the leg's verdict. Both reds
  were real, and the only reason that was knowable is that every leg wrote its exit to its own file
  rather than to the wrapper's status.
- **Remedy, in preference order:** (1) never read a wrapper's status as a leg's — read the leg's exit
  FILE, which the per-leg rule already requires; (2) if a wrapper's own status must be meaningful,
  end it by re-raising the captured status (`exit "$(cat leg.exit)"`) or set the status before the
  capture is the last word; (3) for anything that becomes evidence, redirect to a file and read the
  file.
- **Trigger condition:** ripe now — it rides the next gate-script touch, and the gate scripts in this
  milestone already write per-leg exit files, which is what makes this discipline rather than debt.
- **Size guess:** small — a convention line in the gate-script preamble; no product code.

### IR-61 — spacerun's test-module skip carries a SECOND literal tracker; the two answers can drift, and today only a cell stands between them

- **Status:** OPEN, unowned, accepted-with-mitigation. Filed by doyle 2026-08-22 at the TURNKEY #212
  head gate, as the named residual of the latch repair (the repair itself rides the milestone).
- **What the debt is.** `crates/xtask/src/spacerun.rs` now skips a column-0 `#[cfg(test)]` module by
  BRACE DEPTH and resumes after its close, rather than latching the scan to end-of-file. Counting
  braces safely means knowing when a brace is inside a string, a raw string or a char literal — and
  the scanner ALREADY tracks literals for its rendering pass. It does not reuse that tracking. There
  are now TWO answers in one file to "am I inside a literal", and **two answers to one question is a
  disagreement waiting for a reader who fixes one of them.**
- **Why it was accepted rather than refactored on the spot, stated so the trade is auditable:** the
  two trackers answer genuinely different questions — the rendering walk is PER LINE and forward
  from an opening quote ("where does this literal start and end"), while the skip needs "am I inside
  a literal right now" carried CONTINUOUSLY across lines. Unifying them is a real refactor of a check
  that is already gated, already carries five fixed defects, and sits on the tree being handed to
  golden. The narrow change with an OBSERVABLE failure mode beat the correct-shaped change with a
  wide blast radius, on that tree, on that day.
- **The mitigation, and its exact limit.** Three brace cells fail the moment the two trackers
  disagree about a raw string, a normal string or a char literal, plus a cell that reds when the skip
  stops consulting a literal tracker at all (`match lit` -> `match Lit::None`). **That is a cell
  standing in for a refactor.** It catches drift in the three forms it names and nothing else — a
  fourth literal form, or a change to only one tracker in a form neither cell covers, passes.
- **The general shape, which is why this is a register entry rather than a comment:** a check whose
  own correctness depends on a second implementation of a thing it already implements is one
  refactor away from the defect class it exists to prevent. This same file has already produced
  three instrument defects (an indented-marker latch, a column-0 latch, and a suppression that
  decided what got PARSED rather than what got REPORTED). The pattern is not carelessness; it is a
  scanner accumulating special cases.
- **Trigger condition:** the next SUBSTANTIVE change to spacerun's scanning — any new skip, any new
  literal form, any change to either tracker. At that point unify, rather than adding a third answer.
- **Size guess:** small to medium — one continuous literal-state walk serving both the render pass
  and the skip, with the existing corpus and brace cells as the regression net.

### IR-62 — an e2e daemon binds well-known ports and collides with the resident fleet on shared runners

Filed 2026-08-23 (doyle), from the #212 r2 golden's Windows Phase A red — one witnessed
instance, mechanism verified from the run's own capture, filed on the mechanism per the
register's standard.

**Mechanism.** `endpoint_autostart_e2e::saved_endpoint_replays_on_daemon_restart` failed its
own PRECONDITION ("daemon B must come up with a fresh brain") because daemon B's broker came up
degraded: `NODE_KEY_FAIL: identity unavailable` (net-less broker, no retry) and
`DOCS_SERVER_BIND_FAIL` port 5474 `os error 10048` — the docs port was held by a CO-RESIDENT
daemon. HFENDULEAM is live infra: the resident fleet's daemon legitimately holds well-known
ports, and any e2e that brings up a daemon with default port bindings is in a race with it BY
CONSTRUCTION. An isolated SPT_HOME isolates the store and broker socket, NOT globally-numbered
TCP ports.

**Classification history.** r2: red (35.4s, precondition panic). Run 1 and r3 on the same box:
green. Folded into r3 under a pre-registered predicate (reds twice → dedicated triage); it
greened, so this stays an environment-shaped intermittent, NOT a product defect and NOT
closed — the collision window is real and will re-fire under the right co-residence timing.

**Remedy direction (not built).** Test daemons should bind ephemeral/rig-scoped ports for
every advisory surface (the docs server is advisory — a bind failure should degrade the rig
loudly, not poison an unrelated cell's precondition), or the precondition should name the
port-collision cause distinctly so the red self-classifies. Either lane is test/CI work
(hertz's), triggered the next time this class fires anywhere.

**Kin.** IR-51 (e2e daemon net-enabled on real interfaces — hermeticity is disk-deep only);
the known not-ours red classes list in the #212 hand-off.

**BUILT 2026-08-24 (hertz, SIGNET #218 batch; PR #158).** The class fired its second witnessed
instance the same day — doyle's W2 gate rig, `engine_room_bringup_e2e` precondition panic with
the IR-50 panel naming `DOCS_SERVER_BIND_FAIL` port 5474 `os error 10048` — which armed this
entry's own trigger. Hertz repro'd deterministically (held the port; brain still reached
BRAIN_UP), pre-ranked three mechanisms before probing, and built the ephemeral-advisory form:
rig-only `SPT_TEST_EPHEMERAL_ADVISORY_PORTS=1` makes the daemon's docs listener bind port 0
(production config/`SPT_DOCS_PORT`/docs-url untouched; the flag matches the literal "1" only).
The golden workflow's test job sets it globally on both self-hosted legs, the standalone ER rig
sets it too. Scope note kept honest per his own ranking: the docs collision is retired; if a
co-diagnostic bind elsewhere still poisons a precondition, that is a residual face of THIS
entry, not a closed question.

### IR-63 — suite-mix sweeps leak job-escaped autostart daemons that LOCK the pool's spt.exe; and the same box conditions manufacture one random victim per sweep

- **Status:** OPEN, mitigated-by-rig-step — the reap is adopted discipline, not yet construction.
  Filed by doyle 2026-08-24 at the SIGNET #218 gates; first measured by todlando the same day
  (his W1 build: four reap rounds of 2, 8, 1, 1 processes, count-by-path stated each round).
- **Mechanism, two faces of one condition:** (1) e2e sweeps that exercise autostart leave
  job-escaped daemons running `target/debug/spt.exe`; the NEXT build or `xtask check` then dies
  on `os error 5` removing the exe — a rig red wearing a build defect's clothes. nextest's own
  `leaky` flags corroborate (4–8 per sweep measured). (2) The same co-residence (leaked daemons +
  resident fleet + runner traffic) manufactures ONE red per full sweep with a DIFFERENT victim
  each time: doyle's W2 gate measured FOUR distinct one-off victims in four same-sha sweeps
  (ER bringup, composite+bootstrap pair under CI load, resident_service, brain_resume), every
  one green alone and/or in a sibling sweep. The e2e-leaked-daemons rule holds: a different test
  dying each run on one sha is ONE environment cause, and hardening members never closes it.
- **Mitigation adopted (rig step, both todlando's lanes and doyle's gate runner):** reap
  `spt.exe` BY PATH (`*spt-core\target*`) before EVERY cargo invocation and after every sweep;
  report the count each time so a zero is a claim, not silence.
- **Remedy direction (unbuilt):** the durable form is construction, not discipline — either
  nextest wrap/xtask verb that performs the by-path reap as a pre-step on this box's rigs, or
  autostart e2es gain teardown that provably outlives job escape (kin IR-38's built wedge-namer,
  IR-62's ephemeral ports which retire one collision axis, IR-35's victim-rate ledger).
- **Trigger condition:** next rig/gate-script construction touch, or the first golden red that
  classifies to face (1).
- **Size guess:** small-medium — one wrapper seam plus adopting it in the gate/lane scripts.

### IR-64 — HFENDULEAM's disk floor is set by non-CI bulk: ~845 GB of operator payload leaves the golden box ~2 GB of slack against its own 32 GB preflight

- **Status:** OPEN · **Origin:** doyle, v0.63.0 close sweep 2026-08-26; mechanism first measured
  at the FIELD-SEAL golden window (deployah's IR-31 fifth addendum carries the POOL side of the
  same event by agreed division — this entry is the OTHER reservoir, deliberately not his).
- **Mechanism:** the 1.86 TB C: carries ~614 GB Steam + ~230 GB Downloads (operator payload, not
  CI state). With that floor fixed, normal lane traffic alone walks the box under the 32 GB
  golden preflight — a full-sweep pool WEIGHS ~90–112 GB steady state (IR-59's measurement) and
  five FIELD-SEAL lanes cost 148.48 GB (IR-31 fifth addendum), so ONE milestone's pools exceed
  the entire free margin. Reaping buys windows, not headroom: deployah measured free fall
  167 → 108 GB within hours of his reap, ~59 GB re-consumed by lane builds. The recurring shape:
  every milestone pays a reap-and-measure tax to rent space the box does not structurally have.
- **Why register-worthy rather than "clean up more":** agent-side discipline (IR-31's budgeting,
  pool reaps, teardown steps) is already adopted and still only rents windows — the reservoir
  that would durably move the floor is operator-owned bulk no agent may touch. Naming it here is
  the boundary: agents keep budgeting IN POOLS (not GB); moving Steam/Downloads (or adding a
  disk, or pinning golden to a box without operator payload) is an OPERATOR decision this entry
  exists to put in front of them once, with numbers, instead of re-deriving the floor each cut.
- **Ripe when:** operator rules on the bulk (move/expand/accept-the-tax), OR the first golden
  that dies at the 32 GB preflight despite adopted pool discipline (IR-59's LNK1180-class red).
- **Size:** zero code; one operator decision + at most a runbook line naming the chosen floor.

### IR-65 — kitsubito kernel-audit backpressure: a tailscale-snap AppArmor denial storm + no auditd turned the default audit backlog into a CI-wide stall amplifier (REMEDIATED AT BOX; re-check triggers named)

- **Status:** REMEDIATED-AT-BOX 2026-08-25 (doyle, operator-authorized root — ruling pinned
  releases#225 comment 5418800776); entry stays OPEN as the re-check record because the storm
  SOURCE persists and two named events can silently revert the fix.
- **Mechanism:** `snap.tailscale.tailscaled` (1.92.5) polls /proc and takes an AppArmor
  ptrace-read DENIED per poll — a continuous kernel-audit record storm scaling with process
  count (spawn-heavy serialized CI legs amplify their own storm). With NO auditd installed the
  records rode printk (kauditd throttling) and the 8192 kernel backlog overran
  (lost=1,537,743); the default `backlog_wait_time` 60000ms turns a full backlog into A 60s
  SLEEP INSIDE ANY AUDITED SYSCALL — the broker dispatch stall that stretched IR-30's
  Linux-face race window (see IR-30's Linux-face addendum; deaths clustered 59–63s ↔ this
  knob's value).
- **Remediation (verified at the EFFECTIVE layer, not the fragment):** auditd installed +
  active (storm consumed, backlog drains to 0, lost flat), backlog_limit 32768,
  backlog_wait_time 0 (full backlog may DROP, never STALL). ⚠ THE TRAP PAID FOR ONCE: the first
  knob write (`50-backlog.rules`) silently LOST the augenrules merge — apt's own
  `/etc/audit/rules.d/audit.rules` sorts LAST (digits before letters) and its `-b 8192` /
  `--backlog_wait_time 60000` won; a whole proof run executed under the defaults while the
  fragment grepped perfect. Values now live in the merge-WINNING file; verified `auditctl -s`.
- **Re-check triggers (the reason this entry stays):** (1) tailscale snap update — the plug
  landscape may change, the storm may stop or grow; (2) auditd package update/reinstall — may
  rewrite `rules.d/audit.rules` and re-lose the merge; (3) box reimage. On any of these: one
  `auditctl -s` (expect 32768/0) + one `journalctl -k | grep 'backlog limit'` (expect silence).
- **Ripe when:** a re-check trigger fires. **Size:** two read-only commands per check.

### IR-66 — first-chunk-needle test class: two latent members remain after the v0.63.0 fix (attach.rs:561, :672)

- **Status:** BUILT — hertz rider PR #161, landed ff at `d04b922d` (2026-08-27, IO-PARSER #22
  intake rider). Both members got the later-needle treatment (TICK39 delayed past burst on both
  OS arms; alt-screen entered before delayed ALT_VIEWPORT_MARKER). Gate: doyle — diff-scope
  review; Windows isolated worktree 3× TICK39 PASS + clippy + treqs; Linux CLEAN worktree at the
  PR sha on kitsubito 3× both cells PASS (independent of the builder's dirty-shared-tree proof,
  disclosure on record). · **Origin:** doyle tree-wide census at `dbe3daad`
  (releases#225 comment 5418844404 — the census that corrected my own resume.rs-scoped
  overclaim), after the class's first member was RCA'd and fixed in v0.63.0's head.
- **Mechanism (cite, not restate):** KNOWN-HAZARDS 6.9 same-session face, amended at
  `dbe3daad` — `spawn_session_pid`'s Spawned-wait consumes-and-discards a first output chunk
  that races the reply (~2ms window on a healthy Linux box; load-stretched under backpressure,
  see IR-30 Linux-face + IR-65). A test whose needle exists ONLY in the child's first chunk
  asserts winning a race the contract does not promise.
- **The two members:** `crates/spt-daemon/tests/attach.rs:561` (TICK0..39 instant burst then
  `cat`; needle TICK39 read directly after spawn — the whole burst can sit in the first
  chunks) and `:672` `cross_node_cold_attach_to_alt_screen_gets_clean_repaint` (unix-only;
  `printf '…ALT_VIEWPORT_MARKER'; cat` — needle in the first chunk). Failure shape differs from
  the fixed member: both children idle on `cat`, so a lost prefix is a HANG killed by nextest's
  backstop at exactly 240s (slow-timeout 60s × terminate-after 4) — not a natural-life death.
  Verified NOT in the class: daemon_e2e.rs:209 (needle = echo of post-spawn input), attach.rs:995
  + brain_swap.rs:170 (attach/replay reads, not spawn-waits).
- **Fix shape (ruled at the v0.63.0 close):** hertz lane, NEXT intake — the same one-line
  later-needle / robust-read treatment the resume seed got in `dbe3daad`; deliberately NOT
  landed mid-r4. Both cells passed r4 (0.010–0.058s) — the members are latent, not failing.
- **Ripe when:** next milestone intake (hertz test lane) or the first golden red at ~240s on an
  attach cell — either way the mechanism and fix are pre-derived, triage should cite this entry
  and skip the RCA. **Size:** two one-line test edits.

### IR-67 — wave-battery gap: a `-p spt --bins` leg compiles bin targets as harnesses and runs ZERO integration/e2e tests, so batteries that lean on it carry silent no-coverage legs

- **Status:** BUILT — remedy applied across the IO-PARSER #22 milestone (2026-08-27/28): every
  wave battery named explicit `-p spt --test <suite>` e2e legs beside the `--bins` unit leg
  (GW1..GW5 evidence), every builder evidence report carried a named real e2e per the dispatch
  template, and the one filter mistake (a `--test` name not matching its file) REFUSED loudly
  rather than silently skipping — the failure mode this entry exists to kill. Battery template =
  the dispatch text now carried forward in gate craft.
- **Was:** OPEN · **Origin:** doyle, banked at the SIGNET #218 gates 2026-08-25 ("gate
  battery gap learned"), filed at this sweep per the standing rule.
- **Mechanism:** `cargo test/nextest -p spt --bins` builds `[[bin]]` targets as TEST HARNESSES
  (unit tests inside bins only) — `tests/` integration and e2e suites of the crate NEVER run,
  and the leg's green output is indistinguishable from coverage. Kin of the two banked pool
  faces (`--bins` never emits fixture exes; fresh-pool fixture reds) but this face is about the
  BATTERY TEMPLATE: my SIGNET wave batteries carried a `-p spt --bins` leg believed to cover
  the crate.
- **Remedy:** at next intake, the wave-battery template gains an explicit `tests/` leg for the
  spt crate (nextest `-p spt --test <suites>` or unfiltered `-p spt` where pool history
  permits), and any battery doc that lists `--bins` as a coverage leg gets the one-line caveat.
- **Ripe when:** next milestone intake (the battery template is touched at every intake).
  **Size:** template lines only.

### IR-68 — PSYCHE_INGEST_FAIL cause class unproven: git index.lock did NOT reproduce the ingest failure it was blamed for

- **Status:** OPEN, investigation-shaped · **Origin:** todlando measurement at the W2 int-leg rig
  (IO-PARSER #22, PR #163 at `769fb0a9`, 2026-08-27), reported measure-first; filed by doyle.
- **The contradiction:** the six silent `PSYCHE_INGEST_FAIL:todlando` lines observed at the
  drop-dir probe (2026-08-25) were attributed to shared-checkout git `index.lock` contention. The
  W2 rig planted NON-EMPTY `index.lock` files in BOTH places git takes one (bare git-dir + every
  `<git_dir>/worktrees/<name>/`, mirroring `branchstore::sweep_stale_index_locks`, plant count
  asserted >= 2) — and the ingest COMMITTED ANYWAY (tier write `Written`; `route_slices`
  propagates `commit_live(...)?`, so a blocked checkpoint would have failed it). Non-empty was
  deliberate: KH 1.3 reaps only 0-byte locks, and `pulse_tick` runs no boot sweep. So on this
  store an `index.lock` does not block the checkpoint, and the field incident's mechanism is
  UNKNOWN, not merely unconfirmed. The rig's doc comment records the attempt; the shipped int leg
  fails the ingest at the drop READ instead (upstream of the store).
- **Strongest negative evidence, pinned to the executable rig:**
  `crates/spt-daemon/tests/commune_io_events_int.rs:95-107` documents that the lock mechanism was
  tried **first**, planted non-empty locks in both the bare git-dir and every linked-worktree
  location, asserted that the population was non-empty, and still observed `Written`. Its own
  sentence is the required scope boundary: “This leg does not claim to reproduce a mechanism it
  measured as inert.” The working injection is instead drop-file → directory replacement, which
  fails the upstream read deterministically and proves only the `COMMUNE_FAIL` event contract.
- **NOT claimed:** that the field incident was misattributed to ingest failure generally — the
  files did survive with failed-ingest log lines; what is unproven is the index.lock CAUSE.
- **Ripe when:** next PSYCHE_INGEST_FAIL sighting in the field (capture the failing store state
  before touching it), or a dedicated probe slot. W2's COMMUNE_FAIL event now pushes a NAMED
  reason on every failure, so the next occurrence self-reports its cause — read that first.
  **Size:** investigation; no product change until the mechanism is pinned.

### IR-69 — inherited stderr tokens can tear between format fragments; 26 test consumers parse that surface as structured truth

- **Status:** OPEN, product remedy lane owned by **todlando**; this entry is the exposure map and
  census discipline, not the implementation. · **Origin:** CONDUIT #236 r3 golden attempt 2,
  2026-08-29: `endpoint_autostart_e2e` missed its contiguous
  `ENDPOINT_AUTOSTART:gwauto` keystone although the matching fresh session id proved the replay
  happened. Attempt 3 was clean on the same discarded SHA after 139 GB was reaped; unequal load
  makes the pair corroborating, never a rate.
- **Mechanism, recovered from the literal torn bytes:** two daemon processes inherited one stderr
  pipe. `BRAIN_UP` and `ENDPOINT_AUTOSTART` are each one `eprintln!`, but `write_fmt` may reach the
  handle once per format fragment; the per-process stderr lock cannot serialize the other process.
  The observed interleave split the autostart token between `ENDPOINT_AUTOSTART:` and `gwauto`.
  This is fragment-level cross-process interleaving, not a failed replay and not a two-thread race.
- **Durable remedy, ruled product-side:** render a complete diagnostic line into one buffer and
  issue the complete rendered text, newline included, as **ONE handle call**. An OS short-write
  may retry the late tail so completeness wins over a silently truncated line; the deterministic
  property pins application-side fragmentation, not syscall count under that rare retry. It makes
  no claim that the OS write is atomic on Windows, and no CI re-fire against a 1-of-2 observation
  may stand in for it. Widening one test's grep is refused: it greens one consumer while every
  sibling token remains exposed.
- **Cut-SHA census** (todlando generator, independently hash-verified by hertz), measured from
  `4d6007ac4ee3d0991f4bc60b7e1a55825aad1aaf` via `git show`, not a dirty worktree:
  **487 emitter sites / 383 distinct tokens / 26 consumer test files**. Emitters by crate:
  spt-daemon 270, spt 195, spt-runtime 8, spt-live 5, spt-net 5, spt-store 4; by macro:
  `eprintln!` 485, `println!` 1, `eprint!` 1. `DRIVEN_BY` is the sole stdout row and therefore
  outside the stderr remedy. A consumer means the test both mentions `TOKEN:` and reads a
  stderr/log surface; bare-token matching inflated the population to 72 files by admitting prose.
  This is an **exposure map**, not a rate and not a claim that all 26 have torn.
- **Artifacts:** cut census SHA-256:
  sites `f10d6bfe5e0399f98945cf63bb5719a01a62e1251be84165c73030fa9bb224f5`,
  tokens `80b0d1835d9d46376b2f200b1c50be4ca0e185a135168444414255f9308ed92b`,
  consumers `249a2e33df43a30570ca82b5bd200d8b6e2bf07aa8250fcc8bace47fe23e3cbc`.
  Generator `81df82f01ef33a82307920d9804a2394b787ef2ffbe3b41d2d7dda560d9a62a9`
  asserts newline-terminated output, logical record counts, a known-positive
  `SUBSCRIBE_DECISION` sentinel, four-line macro lookback, and truncation only at a
  `#[cfg(test)]` that opens a module.
- **Census hazards paid for while constructing the population:** the provisional instrument moved
  `277/243 → 286 → 412 → 487` as three silent assumptions were found: same-line matching misses
  multi-line macros; truncating at the first `#[cfg(test)]` drops later shipping code; and a
  lookback must prefer the token line's own macro over a neighbouring arm. A file without a final
  newline also makes `wc -l` silently undercount logical records by one. Every future census must
  name its SHA, assert a known member, and count records inside the generator.
- **Separate anchored-matcher finding:** `IDLE`, `DISPATCH`, `BUSY`, and `BOUND` are common words
  matched as substrings in exposed consumers. One-write emission does not prevent an unrelated log
  line from satisfying them; those predicates need anchored token matching. Keep this separate from
  the tear remedy so neither defect is claimed to close the other.
- **Kin:** [[IR-50]] (the sink must be read before its absence means anything), [[IR-40]] (a
  confident signal that does not describe the source truth), [[IR-8]] (a plausible census whose
  blind spot returns a clean answer).
- **Ripe when:** todlando's product single-write lane lands; re-run the cut generator on that lane,
  pin the one-write property deterministically, and decide whether the four common-word consumers
  need a separate matcher lane. **Size:** product lane owned elsewhere; register follow-through
  small.

### IR-70 — a hand-built `IoBus` silently omitted later sinks while every gate stayed green

- **Status:** OPEN, remedy unscheduled; the v0.65.0 respin fixed the witnessed call site on PR #173.
  This entry owns the recurrence class, not that shipped correction. · **Origin:** releases#234,
  commit `c5459795167`, CONDUIT #236 respin.
- **Mechanism:** `publish_commune_io` constructed `IoBus` itself and registered only the sink it
  knew. `default_bus` already existed as the assembly point and its module contract explicitly says
  emitters must know only `IoBus::publish`, so adding a later sink costs one registration rather
  than an emitter sweep. The hand-built publisher recreated exactly the drift that contract warned
  against: `COMMUNE` and `COMMUNE_FAIL` reached the old sink but never the adapter log, leaving
  `spt api io-events` permanently short on two documented kinds.
- **Why the complete battery was green:** the e2e was vacuous on the broken kinds; units synthesized
  rows without traversing `publish_commune_io`; and traceability checked attached tags, not whether
  the tagged path exercised the real publisher. Three green instruments shared one blind seam.
  The respin routed the site through `crate::iobus::default_bus`, and its source comment now records
  the failure shape.
- **Durable remedy direction (todlando mechanism, not scheduled):** enforce the assembly point so
  the next hand-built bus is impossible rather than found — make production publishers obtain the
  composed bus through one construction API, and gate the real publish path for every documented
  kind. A search or comment is not enforcement; another emitter can satisfy both while rebuilding
  a partial sink list.
- **Kin:** [[IR-39]] (a green fixture never reached the missing dependency), [[IR-37]]
  (traceability tags prove coverage bookkeeping, not behavioral truth), [[IR-58]] (a hand-built
  population returns an orderly but incomplete answer).
- **Ripe when:** the next `IoBus` construction/API touch, or a new sink registration. **Size:**
  small-to-medium assembly-point enforcement plus a real-path population test.

### IR-71 — controller-seat release has no persisted breadcrumb, so tests cannot distinguish propagation lag from a missing detach

- **Status:** OPEN, observability gap; no remedy scheduled. · **Origin:** CONDUIT #236 r3
  `er_brief_once_per_session_e2e` intermittent, classified test-side and fixed on PR #174.
- **What/why:** killing rc1 reaps the controller process, but the broker is another process and
  notices the socket close later. On that detach edge it clears `driven_by` and `controlled` in
  `info.json`; rc2's preflight reads that file locally and can race the write, refuse normally, and
  exit 0 without dialing the broker. The test originally spawned rc2 immediately and had no
  observable precondition separating “release is propagating” from “release never happened.”
- **The missing surface:** broker lifecycle state records `session-detach was_controller=true`, but
  there is no always-on persisted breadcrumb that states the controller stamp was cleared, names
  the resulting `driven_by`/`controlled` state, or measures detach-to-persist latency. The repaired
  test therefore polls `info.json` itself, prints elapsed time (0.075s on the first local real run),
  and uses a named 60s timeout whose failure promotes the finding from timing to missing release.
  Its deliberately latched-seat unit proves the barrier can red.
- **Remedy direction:** add one structured release breadcrumb after the persisted clear, carrying
  session/connection identity and the resulting control state; if latency is carried, measure it
  from the detach edge. It must ride the persisted daemon sink and remain distinct from
  `session-detach`, which proves the connection event but not the file write.
- **Kin:** [[IR-50]] (the sink a test must actually read), [[IR-40]] (an event timestamp is not a
  completion signal), REQ-HAZARD-CONTROL-STAMP-CONVERGENCE.
- **Ripe when:** the next controller lifecycle or structured-breadcrumb touch, or a field timeout
  from PR #174's barrier. **Size:** small emitter plus one contract-level assertion.

### IR-72 — no product read verb surfaces which process holds a perch or controller seat

- **Status:** OPEN, observability gap; no remedy scheduled. · **Origin:** CONDUIT #236 r3
  two-host RCA and PR #174 (`70c1a303`), corrected by doyle after IR-18 and daemon-status were
  ruled out as different classes.
- **Source-verified boundary:** the holder identity already exists in durable/internal records
  (`InfoJson.pid`, `parent_pid`, and the broker's session pid), and product internals read it for
  liveness, teardown, and self-detection. No product **read verb** returns the process that holds a
  named perch/seat. PID-bearing teardown messages are failure outcomes, not an inspection surface;
  daemon status is process-wide, not perch- or seat-scoped. This also does not collapse into
  [[IR-18]], where an existing internal `read_pid` collapses absent and unreadable records.
- **Paid-for consequence 1 — rigs manufacture custody:** PR #174's two-host repair had to add
  `PidHolder` in `crates/spt/tests/twohost_cli.rs:401-437`, spawn a long-lived sibling process,
  write that pid directly into each synthetic perch, retain the child handle, and kill+wait it on
  drop. The rig can make a known holder; it cannot ask the product which holder the product sees.
- **Paid-for consequence 2 — ambiguity stayed latent until behavior flipped:** the r3 fixture
  seeded three perches with one test-process pid. Self-detection then enumerated a directory whose
  first matching row was a filesystem-order coin, so a green could name the wrong perch. The
  repaired resolver refuses an ambiguous candidate set instead of guessing, and the fixture's
  `assert_only_ancestor_candidate` now re-reads every `info.json` rig-side
  (`twohost_cli.rs:439-457`). Refusal is enforcement, not observability: it proves a tie but still
  offers no supported verb that names each holding process.
- **Remedy direction:** add one machine-readable, named-perch inspection surface that reports the
  custody fields the product actually used — holder pid and role/source (bind relay, stable harness
  parent, brokered controller/session) — without asking callers to parse private `info.json` or
  scan the process table. PID alone is recyclable; where a birth/image/session stamp exists, carry
  it so the output does not become a new bare-pid oracle.
- **Kin:** [[IR-18]] (read primitive loses error class, explicitly distinct), [[IR-71]] (release
  completion lacks a persisted breadcrumb), REQ-HAZARD-SELF-DETECT-TIE (refuse ambiguous ancestry
  rather than select by enumeration order).
- **Ripe when:** the next roster/status read-model change or another rig needs to reap a named
  holder. **Size:** small read-model/API addition, medium if seat and perch custody need separate
  typed variants.


Wall time: 0.18 seconds