# Infra register — CI / build-pipeline debt

Operator-ruled 2026-08-02: infrastructure and CI-pipeline work items live HERE, not on the
`spt-bs-releases` board. The board carries product surface the operator triages; this register
carries what the gater triages. **Mandate: doyle sweeps this file at every milestone intake and
every release close, and composes ripe entries into waves/milestone riders.** An entry leaves
this file only by being built (link the lane) or being retired with a stated reason.

Entry format: status · origin · what/why · trigger condition (what makes it ripe) · size guess.

Last sweep: 2026-08-19, CONCIERGE (#183) INTAKE (wave map: `CONCIERGE-183-JIT.md` + #183
comment 5347768797) — rows read from the LOCAL register stack (`4799031→ecd640c→8d4c224`;
origin/main lacks IR-49/50/51 until the chain lands — absence ≠ not-filed). Rulings: IR-50
COMPOSED — dispatched to hertz same day (sink-path-helper-FIRST sequencing per the entry;
`engine_room_bringup_e2e.rs` excluded, its panel sites ride the gated #199 lane), with the
engineroom.rs:145 misnomer as a rider on the same lane. IR-47 = candidate co-rider IF this
chain touches golden.yml/ci-notify, else holds. IR-48 holds (next xtask parity-cell touch).
IR-49 holds (next poolguard touch or first landed-lane takeover request). IR-51 holds —
gated on #199 attribution; a net-off fix must NOT precede attribution. IR-2 trigger NOT met
(queued unlanded lanes exist). No entries retired; none filed — the intake window's product
finding (rc terminal-Exit omission, hertz RCA: proven invariant violation, load excluded
92/92, H2 candidate unconfirmed) routed to the BOARD as releases#201, the correct venue for
product surface.

Prior sweep: 2026-08-18, KEYSTONE (#182) INTAKE — the prior sweep's four composition rulings
EXECUTED (wave map: `KEYSTONE-182-JIT.md`, base main @`8248bc3`): (1) IR-9 pin lane DISPATCHED to
hertz (W0 item 1; interim kitsubito-clippy-authoritative dies when it lands). (2) IR-21 remedy
(2) + IR-39 vs `f24e732` DISPOSED: the branch is a BEACHHEAD (golden.yml prebuild step = the
prebuild-made-a-rule arm, ONE fail-fast site, one ledger row) — hertz rebases it off its
abandoned parent `b483699` (the id-collision draft; content landed renumbered as IR-46 @`8248bc3`)
and lands it in W0; the class remainder (IR-39's SHARED precondition helper + 24-site sibling_bin
sweep, the 33 cross-package build-edge expressions) folds into the test-hygiene lane. (3)
Test-hygiene family DECIDED-ACTIVATED as a dedicated hertz lane inside the KEYSTONE window,
sequenced after W0 (members: IR-13/23/36/37/38 + IR-21 clarity half + IR-39 helper +
twohost.rs:394 doc comment). (4) First-execution-cells discipline restated in the JIT's
golden-head step. Board members #84/#85 ride hertz's W0; #166/#57/#185 are todlando's W1. IR-46
id-collision RULED this intake (renumber-and-land, parallel issuance not authored disagreement;
deployah landed it @`8248bc3`). No new entries filed; no entries retired.

Prior sweep: 2026-08-19, NAMEPLATE (#181) / v0.56.0 RELEASE CLOSE — shipped c91 @`60d74ea`
(tag == golden-tested sha, run 32209922535; one respin, both first-golden reds ruled rig defects).
Filed IR-42 (pool-claim writes / build enforces — BUILT in this same commit, AGENTS.md line),
IR-43 (knock NoReply past its 30s carrier bound), IR-44 (perch-sentinel comment overclaims
preservation), IR-45 (twohost rig home premise; SETTLED + RETIRED
2026-08-19 — pump paths resolve under the per-run temp root, see the entry; #189 corrected on the
board, comment 5337519936). Two cycle findings ruled RUNBOOK-homed and
landed in this commit rather than as entries: the first-execution-cells intake question
(RELEASE-RUNBOOK golden-head intake — name the never-executed cells before the run) and
deployah's sweep-vs-cascade mechanism (RELEASE-RUNBOOK board step — under golden CI the cascade
is DRIVEN via `state <mref> acceptance`, never swept). ⚠ LABELLED HOLE — CLOSED UNDERIVABLE
(2026-08-19): the close commune's batch list named "alchemy create-races"; its content did not
survive the author's context reset and was not recoverable from #181, the JIT records, or memory.
Deployah answered the query: he ran ZERO create ops at the cut (could not have witnessed a create
race), a fresh probe over the milestone window shows no duplicate mints (the 4-issues-in-2s batch
mint is batching, not duplication), and the only surviving trace is doyle's own pre-reset message
naming the item as already-known — a pointer, not a sighting. Item DROPPED; the hole stands as
the record. Deployah's sweep-vs-cascade ship-path trap (#181 comment 5337402639, runbook-homed
above) is explicitly NOT this item's content — do not fold it in. Composition: IR-9's pin lane
(`rust-toolchain.toml` @ 1.96.0) goes to hertz AT KEYSTONE #182 INTAKE per the 2026-08-05
ruling; IR-21 remedy (2) + IR-39's precondition helper compose with hertz's standing
fixture-prebuild-hardening branch `f24e732` — disposition at the same intake; the test-hygiene
family (IR-13/23/36/37/38 + IR-21's clarity half) stays the dedicated post-batch lane candidate,
decision at intake; IR-29's proving run + IR-30's instrument lanes ride the next golden batch.
IR-2's trigger explicitly NOT met (queued unlanded lanes exist: IR-29/IR-30 instruments,
four-arm refusal eprintln, f24e732). IR-14 hygiene movement: `.worktrees/nameplate-asm-2bd36f1`
reaped this sweep (+66.98 GB by FS delta 120.53→187.51; claim `gate-w7-courtesy` base `fd3dc5a`
in main = finished lane; zero inbound reparse points) — stale-lane audit itself still open.

Prior sweep: 2026-08-05, LOCKSMITH tranche-2 (#141) CLOSE — golden run 30971976024 green on all 9
jobs, main ff'd to `0a25b77`, v0.55.0. IR-40 filed (stale-resume-brief + early-informant class).
IR-9's decision trigger FIRED 2026-08-05: doyle read the runner-account versions off this run's
`test` legs and RULED — pin in-repo via `rust-toolchain.toml` @ 1.96.0, kitsubito's clippy leg
authoritative in the interim; pin lane to hertz at next intake (see the entry). IR-41 filed the
same night (queued main run superseded without a record; runbook step 3 corrected in the same
commit). IR-1/IR-4's golden-only steps (link probe both boxes, toolchain
print both legs) had their FIRST EXERCISE here, discharging the "unexercised until a golden run"
caveat at the CI-RIDER LANE STATE foot section. IR-35's re-measure rode the batch (`c65b838`).
Owlery-noun thin lane SCOPED and dispatched to hertz for the next batch (class A only, two sites;
class B on-disk rename is an explicit non-goal — see IR-40's kin discipline for why the boundary is
written into the brief rather than left to judgement).

---

## OPEN

### IR-1 — Quiet predicate needs a network axis (tailscale RTT probe)
- **Status:** open, LANE EXISTS UNLANDED — see [[CI-RIDER LANE STATE]] at the foot of this file;
  carried by `bf8c4a2` → `REQ-CI-LINK-HEALTH-PROBE`, mapping CONFIRMED by builder hertz 2026-08-03
  (by content: tailscale ping ×5, med/max RTT rows into the bench ledger, three arms
  success/NO-REPLY/UNAVAILABLE, always exit 0 — instrument, not gate). NOTE `bf8c4a2` is NOT
  single-purpose: it also corrects the free-space preflight floor read
  (`REQ-CI-FREE-SPACE-PREFLIGHT`, the ci-runner-has-no-warm-target shape) — no 1:1 commit→IR map
  for this commit · **Origin:** golden/bench-wiring red triage 2026-08-02 (ex releases#126)
- **What/why:** the shared-runner quiet predicate (zero non-terminal runs + no local
  cargo/rustc/nextest by parent chain) is process-shaped; both axes passed on a box whose only
  link was degrading (321s for a 1s checkout, bidirectional 10s QUIC dial timeouts). A tailscale
  RTT probe to the peer box before two-host rendezvous, carried in the bench ledger, would have
  called run 30771155390's red in seconds. Evidence: the arm-1 count table (PUMP_PEER_FAIL
  a 0→3→0, b 8→22→8 across green/red/rerun).
- **Permanent, not stopgap:** operator-confirmed 2026-08-02 that kitsubito cannot be provided
  ethernet — wifi-only indefinitely, so the link cannot be hardened and the predicate must see
  link health.
- **Ripe when:** next CI-touching wave, or the next network-shaped golden red — whichever first.
- **Size:** small (one probe step + ledger row + predicate doc).
- **Composed:** LOCKSMITH (#132) CI-rider cluster, hertz thin lane — 2026-08-03. GREENLIT with
  #132 and **DISPATCHED to hertz 2026-08-03** (direct brief; the register is the spec, there is no
  board issue). Leaves the register only when the lane lands or the entry is retired.
- **FIRST FIELD USE, and it DISCRIMINATED (2026-08-04, golden 30873007187 attempt 1):** the probe
  (landed via the rider lane, riding `4b37512`) read 4–7ms RTT healthy on both twohost legs
  minutes before both legs redded — REFUTING the degraded-link read for that red and steering
  triage to the real mechanism (the [[IR-29]] serve-window race) instead of a link chase. The
  instrument's first catch was a correct NEGATIVE — exactly the call run 30771155390 needed and
  could not make.

### IR-2 — Settle the warm-runner CARGO_INCREMENTAL delta
- **Status:** open · **Origin:** #103/#108 bench-wiring lane 2026-08-02 (ex releases#127)
- **What/why:** the #103 measurement (−29.8% wall, −5.65 GB/target, n=3) is COLD-build only.
  Golden's runner `_work` target persists warm, where incremental is exactly what keeps it cheap;
  CARGO_INCREMENTAL=0 was applied only to the genuine cold build (n1-gate pinned old-broker
  cache) + local rig recipes (docs/GOLDEN-CI.md). Open question: does incremental still pay on
  the warm runner, weighed against 5.65 GB/target on a box with LNK1318 free-space history?
- **Method (hertz):** one full golden each way on a quiet box, outside a milestone, compared
  per-step from the bench ledger.
- **Ripe when:** a quiet between-milestone window with no queued lanes (the measurement burns
  two golden windows).
- **Size:** medium (two proving runs + verdict + possible leg flips).

### IR-3 — Daemon-level guard: broker-net wakeup rate bounded across endpoint churn
- **Status:** open · **Origin:** releases#125 remediation, todlando REQ call 1 (ex releases#128)
- **What/why:** the swarm-discovery GC-spin burned two cores for two weeks visible only in a
  process table — no suite assertion sees the class. Wanted: a daemon-level assertion that
  broker-net workers stay quiescent across repeated endpoint create/destroy churn.
- **Design constraint (pre-ruled):** assert on WAKEUP RATE / voluntary ctxt-switch delta over
  the churn window, NOT %CPU — CPU thresholds flake under CI load; the defect signature
  (~100 Hz per orphan loop) is load-independent. Mint the REQ at activation.
- **Near-product:** this is runtime-defect visibility, the most product-adjacent entry here —
  a candidate rider on any daemon-lifecycle milestone.
- **Ripe when:** the next milestone touching spt-net endpoint lifecycle or daemon supervision.
- **Size:** medium (churn harness + counter plumbing + flake-safe assertion).

### IR-4 — Lock-pin guard + lock-procedure rule + toolchain print (three riders, one lane)
- **Status:** open, LANE EXISTS UNLANDED — see [[CI-RIDER LANE STATE]] at the foot of this file;
  candidate commits `1275e47` + `49d4805` + `8ed006b`, mapping NOT yet confirmed by its builder ·
  **Origin:** releases#125 fix-lane intake hold (ex releases#129 + riders)
- **What/why, three parts that land together:**
  1. **xtask check leg:** assert Cargo.lock resolves swarm-discovery to git rev
     `89a2200d54a4e3cab2f46cc75ebff49a1fb07614` while the patch is load-bearing; the check's
     message states its own drop condition (upstream ships a post-PR#27 release AND iroh's pin
     reaches it). Without it, stanza removal or a routine iroh bump silently returns the lock to
     the spinning crate and nothing reds.
  2. **Procedure rule for lock-touching lanes** (docs): targeted `cargo update -p <crate>` only,
     never full re-resolve; count changed `[[package]]` blocks AND diff per-block edges (set-identical
     hid 8 windows-sys edge movers at acaaa4f); two resolutions disagreeing = toolchain drift —
     stop and compare against CI before shipping either lock; hand-edited lock acceptable iff
     `cargo check --workspace --locked` passes.
  3. **Toolchain-version print step in golden** (cargo/rustc versions, both OS legs): the acaaa4f
     comparison against CI was impossible because no run log prints a version. One-grep audit.
- **Ripe when:** next CI-touching wave; part 1 sooner if any iroh bump is proposed.
- **Size:** small-medium (one xtask leg, one docs section, one workflow step).
- **Composed:** LOCKSMITH (#132) CI-rider cluster, hertz thin lane — 2026-08-03. GREENLIT with
  #132 and **DISPATCHED to hertz 2026-08-03** (all three parts land together).

### IR-5 — Shared nextest summary parser
- **Status:** open · **Origin:** BAROMETER triage (standing recommendation, pre-register)
- **What/why:** two agents in one day wrote `[0-9]+ tests run` parsers that read "1 test run"
  (singular) as zero — a guard fed by a broken parser condemns valid rounds. One shared,
  singular-aware parser (single-source discriminant) for every consumer of nextest summaries.
- **Ripe-when REWORDED 2026-08-04 (todlando audit): the old trigger was unfireable as worded.** At
  `b7b00c3` the in-tree population of nextest-SUMMARY parsers is ZERO — the two scripts that read
  nextest output (`g6-curve.ps1:83-90` per-test lines, `g6-postbounce.ps1:44` display-only grep)
  neither parse counts nor carry the defect, so "next wave touching any gate script that reads
  nextest output" could fire on a non-defective script while the real population (agent-authored
  throwaway parsers, which never enter the tree) stays out of reach. New trigger: **the next time
  anyone — agent or lane — needs a nextest summary COUNT**, the shared parser is built FIRST and the
  need consumes it; rig briefs should name it so throwaways stop being authored.
- **Size:** small.
- **Composed:** LOCKSMITH (#132) CI-rider cluster, CONDITIONAL — lands iff the hertz thin lane
  touches gate scripts; otherwise stays open here. 2026-08-03. Carried in hertz's 2026-08-03
  dispatch brief as a **conditional** rider; hertz reports whether the condition fired. **Not yet
  known to be building** — an unreported condition leaves this entry open, not landed.

### IR-6 — Membership logging on subnet gates
- **Status:** open · **Origin:** BAROMETER triage (standing recommendation, pre-register)
- **What/why:** counts beside results, membership beside counts — gate logs that state a count
  without naming the population keep producing unreadable reds. Standardize membership
  enumeration in gate output.
- **Ripe when:** next wave touching gate scripts / CI legs that report counts.
- **Size:** small.
- **Composed:** LOCKSMITH (#132) CI-rider cluster, CONDITIONAL — lands iff the hertz thin lane
  touches gate scripts; otherwise stays open here. 2026-08-03. Carried in hertz's 2026-08-03
  dispatch brief as a **conditional** rider; hertz reports whether the condition fired. **Not yet
  known to be building** — an unreported condition leaves this entry open, not landed.

### IR-7 — Phase A rigs leak a daemon+brain pair on Windows (exe-lock kills notify relink)
- **Status:** open · **Origin:** BAROMETER post-publish triage (ex releases#124 — full mechanism on the closed issue)
- **What/why:** two Phase A rigs launch daemons that escape the job object via WMI-rung autostart
  (double-space unquoted cmdline fingerprint; SPT_HOME in the wrapper cmdline is the attribution
  key); the leaked pair holds target/debug/spt.exe and kills every golden job reaching the notify
  relink. CI reaps as tourniquet (fa6e597); in-job test launch is the fix.
- **Ripe when:** next wave touching the Phase A rigs or daemon autostart path.
- **Size:** medium.

### IR-8 — reap-census scoped_survivors=0 is blind to unreadable-path holders
- **Status:** BUILT 2026-08-18 on KEYSTONE #182 hygiene lane. · **Origin:** BAROMETER triage
  (ex releases#122)
- **What/why:** a zero that cannot see is not a zero — census scoping skips procs whose exe path is
  unreadable, so the survivors count can report clean while a holder lives. Needs a positive control
  / explicit unreadable bucket in the verdict line (unreadable_path count exists; the ZERO must
  refuse when it is nonzero).
- **Defining specimen (golden 30782259675, hfenduleam test job, 2026-08-03):**
  `CI-REAP summary: killed=5 kill_failed=1 scoped_survivors=0` — an admitted kill failure printed
  beside a zero-survivors claim on the same verdict line. The held image surfaced one step later:
  run-scoped tmp cleanup denied 5/5 attempts on `...\relshell\svcmock.exe` (2nd appearance of the
  svcmock hold; 1st @7a3c08c, pre-kill-auth). hertz's addendum: an image-held survivor also blocks
  WRITES to the exe path — the same class manufactures build/relink access-denied reds that mask as
  build problems, not just cleanup warnings. Not per-run: the same-sha green rerun's leg read
  `kill_failed=0 scoped_survivors=0` throughout (hertz, 30784469908) — intermittent sighting,
  second of its class, not a deterministic fixture property.
- **The Linux twin is strictly worse (todlando audit 2026-08-04, vs `b7b00c3`):** `reap-census.sh`
  has NO `kill_failed` anywhere (0 occurrences vs 2 in the `.ps1`) — its kill loop increments
  `killed` only in the success branch with no else, so a failed kill increments nothing and prints
  nothing. The specimen that made this class VISIBLE on Windows would be INVISIBLE on Linux.
  Population precision so this is not overclaimed: ESRCH is benign (already gone; the kill-time
  re-resolve makes it the common case); the vanishing case is EPERM against another account's
  process. The `.sh` `scoped_survivors` DOES come from a post-reap census re-measure, so survivors
  are measured — the hole is failed kills and the unreadable bucket, not the survivor count.
  Windows precision from the same audit: `unreadable_path` IS on the CI-CENSUS line (:157) but the
  CI-REAP verdict line (:243, :246) still carries only killed/kill_failed/scoped_survivors — the
  zero still does not refuse, exactly this entry's ask.
- **Built evidence:** both verdict lines now carry `unreadable_path`; any nonzero unreadable
  family population renders `scoped_survivors=UNPROVEN` rather than a false zero. Linux also
  counts and reports failed kills. The shared predicate is mutation-pinned by
  `reap-census-selftest.sh`; the PowerShell implementation parses cleanly.
- **Ripe when:** next census/reap script wave (natural pair with IR-7's lane) — now BOTH platforms.
- **Size:** small.

### IR-9 — Golden boxes run different clippy versions
- **Status:** OPEN — RULED, awaiting its build lane. Measurement half LANDED AND EXERCISED:
  the toolchain print (`8ed006b` → `REQ-CI-TOOLCHAIN-VERSION-PRINT`) ran on both legs of golden
  run 30971976024 (`test` jobs 92198170694/92198170700 — the step is scoped to `test`, not
  n1-gate; grep token `TOOLCHAIN `). Runner-account facts, read by doyle 2026-08-05:
  hfenduleam `cargo/rustc 1.93.0` + `clippy 0.1.93`, kitsubito `cargo/rustc 1.96.0` +
  `clippy 0.1.96` — the interactive prior confirmed, and the load-bearing NEW fact is
  `stable (default)` on BOTH: neither box is pinned, so any box-side update re-opens the skew.
- **RULING (doyle, 2026-08-05, on the runner-account facts as the 2026-08-03 hold required):**
  (1) align-by-event REJECTED — with both boxes on unpinned stable, alignment decays silently;
  the skew is a mechanism and the fix must be one too. (2) declare-one-leg REJECTED as the
  terminal state — it repairs lint authority but leaves the legs resolving with different cargo
  versions, and the acaaa4f lock-attribution question this family started from is a resolver
  question. (3) **PIN IN-REPO: `rust-toolchain.toml`, `channel = "1.96.0"`** — both runner
  accounts already resolve their toolchain through rustup (witnessed by the `active=` line), so
  the pin self-applies with zero per-box maintenance, and the judge's version becomes a property
  of the TESTED SHA — the same object golden CI already guarantees. Bumps become reviewed lane
  commits, tested by the golden run they ride; the hfenduleam leg proves 1.96.0 on the pin
  lane
[…1739ln elided…]
s never the mechanism, it only widened the
  blind spot.** Remedy option (a) is the live one: a real placement rule in the checker (upstream
  experimplate), not a per-tag-nearest-item heuristic. Measurement caveat, carried at the
  measurer's own insistence: an intermediate probe printed `TAG_ADJACENT_AFTER_REVERT=0` and that
  was the PROBE wrong (`grep -B 1` above the fn lands on `#[test]`, not the tag above it), not a
  failed revert — the revert was verified directly (exactly one tag, above `#[test]`). The
  discriminator result is the record; the intermediate zero is not.
- **REMEDY DISPATCHED 2026-08-18 (doyle):** upstream lane opened in BigscreenVR/traceable-reqs
  (the actual upstream remote; the checkout at `~/Documents/projects/traceable-reqs` tracks it —
  "experimplate" above named the tool's origin project, not the repo) — checker-side placement
  rule, default-on in `check`, positive controls all four stages + separated-tag negatives
  including the single-tag displaced discriminator shape and the interposed-const shape.
  Requested by hertz as the IR-37 prerequisite of his test-hygiene family lane (KEYSTONE #182).
  spt-core consumes the upstream RELEASE only — no local shadow checker. Issue/branch/PR refs
  land here when the lane reports.

### IR-38 — CLASS: an e2e that ends with a live daemon-spawning binary wedges the NEXT build in its pool, and the diagnostic names the wrong lane
- **Status:** BUILT 2026-08-18 on KEYSTONE #182 hygiene lane. · **Origin:** todlando 2026-08-04,
  `shell_relink_force_e2e` on the #6 lane (filed by doyle; the finding and the fix are todlando's).
- **What/why:** the e2e ends with a LIVE binary by design; on Windows a live process holds its
  image open, so the NEXT cargo invocation in that target dir dies with
  `failed to remove file target\debug\spt.exe: Access is denied (os error 5)` — surfacing as
  exit 101 with NO test summary, because the failure names the BUILD and the FILE, never the test
  that leaked the holder. Actual holders measured: two lane-local `spt daemon run --detached` +
  `daemon brain` processes spawned under the test's temp SPT_HOME by a child CLI's
  `ensure_running`, still alive from an earlier run. **Reap BY IMAGE PATH** — the fleet's other
  eleven spt processes run from the installed binary; a name-based sweep takes your own daemon.
  The live shell-spawn population is eleven e2e files. Nine already end through product teardown
  or their authenticated home-scoped reap. The two intentionally live-at-success siblings now
  kill and prove the final resident gone AFTER all identity/wake assertions:
  `shell_relink_force_e2e` and `shell_stale_online_e2e`.
- **Kin:** [[IR-7]] (phase-A daemon+brain pair leak), the exe-lock reap-by-path rule, and the
  wrong-lane-diagnostic family ([[IR-21]]'s location-not-build class).
- **SECOND HOLDER CLASS, measured 2026-08-04 (doyle, own gate rig):** an **ORPHANED CONHOST with an
  inherited CWD** inside the test tree blocks `git worktree remove` with Permission-denied on a
  DIRECTORY handle — invisible to any exe-path scan (conhost runs from System32; the parent that
  spawned it was already dead). Found via `.github/ci/find-cwd-holders.ps1` (IR-11's instrument);
  killed by pid after parent-dead classification; 34.78 GB reclaimed. Mechanism chain: the
  windowless spawn path masks DETACHED_PROCESS off, the child owns a console, conhost inherits the
  child's cwd and can OUTLIVE it. Boundary (hertz's, correct): this find and his field-4 companion
  are the SAME MECHANISM FAMILY measured on DIFFERENT populations by different instruments — his
  capture measured a ppid-match COUNT (floor 2, companion identity UNMEASURED, his stated limit),
  and this conhost is the first FIELD identity evidence in the family, on THIS population. Do not
  cite his row as identity data it never carried. (His box-wide accretion count 77→80 is a third
  population again.) Teardown rule
  addendum: a disposal leg must GATE on the removal's exit code (a rig that logs "reclaimed 0.01
  GB" from a failed remove fabricates its own success) and sweep CWD HOLDERS, not only image
  locks — two populations, either alone is a clean zero on the other.
- **Ripe when:** the test-hygiene lane (IR-13/IR-23/IR-36 family) — sweep every e2e that ends
  with a live spawning binary, apply the same reap-after-assertions shape.
- **Size:** small per test; population unknown until swept.
- **Built evidence:** focused runs of both live-at-success tests pass back-to-back in one target
  pool, followed by a build invocation from that same pool; the second test's final kill is polled
  to proven process death rather than treated as fire-and-forget.

### IR-39 — CLASS: the missing-fixture-bin defect has TWO failure faces, and one of them impersonates a lifecycle defect of the subject under test
- **Status:** BUILT 2026-08-18 on KEYSTONE #182 hygiene lane. · **Origin:** hertz + todlando
  2026-08-04, same evening from opposite sides (filed by doyle; hertz relayed todlando's
  suggestion without endorsing scope).
- **What/why:** a rig whose narrowed build (`-p <pkg> --test <one>` / `-p spt --bin spt`) skips a
  cross-package fixture bin fails in one of two ways depending on the consuming site.
  Face 1 (cheap): `translate_proof_fixture` panics "must be built: <path>" — names the artifact
  and the recipe, costs one build. Face 2 (expensive): `sibling_bin("mock-session")` resolves to a
  nonexistent exe, nothing spawns, and the test reds at a PRECONDITION assert whose message
  ("warm bring-up spawned=false", left None / right Some("offline")) reads as a lifecycle defect
  of the subject — measured cost 186s + a triage inside a lent box window (hertz's F1 red arm,
  attempt 1). The build gap itself is the known cross-package row of the fixture-bin population;
  what this entry carries is the DIAGNOSTIC asymmetry.
- **Ask:** a fixture-bin existence precondition that NAMES the missing exe and its build recipe
  wherever a rig consumes one — `sibling_bin` (and the shared
  `crates/spt-term/tests/support/fixture_bin.rs` resolver already carries the recipe string as an
  argument, the shape to copy) refuses with the artifact + recipe instead of letting the consuming
  test fail downstream in its own vocabulary.
- **Kin:** [[IR-21]] (wrong-lane-diagnostic family: the failure names the wrong actor), the
  cargo-builds-package-bins-for-integration-tests rule (cross-package row is the only unguaranteed
  build), golden.yml's dedicated fixture prebuild step.
- **Ripe when:** the test-hygiene lane (IR-13/IR-23/IR-36/IR-38 family) — same sweep population;
  or standalone small if that lane keeps slipping.
- **Size:** small — one shared helper + a sweep of local `sibling_bin` resolver copies (24 known).
- **Addendum (doyle, 2026-08-04 late): third and fourth sightings inside 24 h.** Face 1 twice
  more the same evening — todlando's teach-lane leg set (his own translate_proof_fixture, caught
  by his prebuild), and doyle's cold gate rig for the same lane (first-pass spt-bins red at
  136/611, correctly withheld from the lane verdict as rig evidence, fixed by `cargo build -p spt
  --bins` + rerun 611/611). Four sightings in one day across three agents and both faces: the
  ripeness condition ("the test-hygiene lane keeps slipping") is under measurable pressure — the
  prebuild remedy is being re-derived per-rig per-agent, which is the recurring cost the shared
  precondition helper exists to delete.
- **Built evidence:** all 29 file-local `sibling_bin` resolvers now route through one shared
  precondition. A missing fixture names its path and the package-correct `cargo build -p <owner>
  --bin <name>` recovery command before product assertions or timeout paths run; a focused
  negative test pins that diagnostic.

### IR-40 — CLASS: NO envelope-side ordering signal on a resume brief tracks content recency, and a CI informant's per-job line is a claim about that STEP, not the run
- **Status:** open · **Origin:** three sightings in one night, 2026-08-04/05 — hertz, doyle,
  todlando (filed by doyle; the artifacts are hertz's and todlando's, the CI arm is doyle's with
  deployah's framing). Two arms of ONE family: a stale record that carries a confident ordering
  signal, and a status claim that arrives before the thing it describes.

- **ARM A — resume briefs. What/why:** a session can receive TWO start-of-session briefs, and the
  one routed LATER can carry the STALER project-context. Measured on hertz's pair (same node id):
  routed_at_ms `1785897407373` > `1785896642480`, spill-filename epoch `1785897764588` >
  `1785896658159`, and vector sequence `1242529/1242531` > `1241542/1241544` — **all three
  monotonic orderings rank the staler-content brief as newer.** todlando's pair is the NEGATIVE
  arm: his later brief (`…1785897765196`, routed `…577476`, vector `1242756/7`) genuinely WAS the
  fresher one, so the same heuristic happened to pick right.
- **The finding is UNCORRELATED, not inverted** — and that is the whole entry. An inverted signal
  is usable once you know to flip it; an uncorrelated one reads as reliable until the pair where it
  costs you. Neither agent could have known which pair he held without checking content against
  measurement. **Do not name routed_at_ms specifically** — naming the stamp invites the repair that
  fails ("the timestamp is unreliable, use the sequence number"). The vector is the most dangerous
  of the three: a stamp looks like a clock and invites suspicion, a sequence number looks like an
  ordering and does not.
- **Visible failure mode if a brief is obeyed as instructed** (hertz's, measured): re-open
  `releases#123` — work already ruled closed — and under-report his own claimed-pool count by one,
  the row sitting under a live gate. doyle's own sighting the same night: a brief describing main
  @`92b2482` with a teach lane "mid-flight" hours after every lane had gated.
- **PRE-REFUSED DISCRIMINATOR, recorded so it is not rediscovered as a finding:** file SIZE ranks
  the fresher brief correctly on BOTH pairs (hertz 16308 > 12649; todlando 14138 > 10418) — 2 of 2,
  the only signal surviving both samples. **Do not use it.** Size measures RICHNESS, not recency;
  the two correlated only because todlando's thin brief was a correction fragment, and hertz's
  STALER brief was a full 12649-byte dump. Two full dumps hours apart go flat or inverted. Handed
  over pre-refused by its own finder.
- **ARM B — CI informants (doyle, 2026-08-05, run 30971976024).** `CI-KITSUBITO` sent
  "CI SUCCESS … sha 0a25b77" while the run was measurably `status=in_progress` with
  `conclusion=""` — the EMPTY STRING, not `success`. The informant fires from a per-job `notify`
  step, and at send time `notify` itself had no conclusion. All eight substantive jobs were green,
  so the line was **right in the end but EARLY**. deployah's framing, kept verbatim in kind: *an
  informant line that is right in the end but early is worse than one that is simply wrong, because
  a wrong signal gets distrusted while an early-but-correct one gets promoted to the conclusion.*
  todlando's sentence is the one to keep at the point of use: **a per-job notify step's message is
  a claim about that STEP, not about the RUN.**
- **ARM B MECHANISM — hertz, source read 2026-08-05, zero box cost. It is NOT a race; it fires
  early on EVERY run, and the entry above understated it.** `notify` is a job INSIDE the run
  (`golden.yml:1138`), and a job cannot observe its own run's conclusion — so `conclusion: ""` is
  not a timing artifact, it is **the only value that can exist when the informant speaks. The
  informant has never once made a statement about a run conclusion.** What `ci-notify.sh` actually
  asserts is narrower and worth naming exactly: a verdict computed PURELY from six `needs` results
  (changes, traceability, test, n1-gate, twohost-a, twohost-b — the two matrices collapse to one
  result each) being neither failure nor cancelled. Nothing more.
  **The blind spot is exactly ONE job: `notify` itself, the 9th — it cannot appear in its own
  `needs` list. And that one job has a WITNESSED red path**, documented in the script's own header
  (2026-07-27, twice on kitsubito): verdict computed, then `spt send` HUNG ~5 minutes until the
  job's own 5-minute budget cancelled it, and that cancellation reddened the run. So *"informant
  says SUCCESS, run concludes FAILURE"* is a real observed sequence, not a hypothetical. The 30s
  `SEND_BOUND_SECS` added since bounds the hang but does not close the class — it makes the
  informant's own failure FAST rather than impossible.
  ⇒ The informant's SUCCESS covers 8 of 9 jobs and is blind to the 9th, and the 9th is the only one
  with an observed failure mode that reds the run. This is why the ff predicate is the run's
  conclusion field and never the informant line.
- **RECOVERY, and it is the same for both arms: never rank the claims — re-derive from the source
  object.** For briefs, hertz's four falsifiers are the method (main sha, claimed-pool count, lane
  tip, issue state) and they work because each is a content claim checkable in ONE command against
  a source NEITHER brief controls. For CI, read the RUN's own `conclusion` field, and prefer
  independent reads of the run object (`gh run watch --exit-status` polls the run, so its exit is a
  second measurement rather than an echo of `notify`). On the batch this was caught in, three
  sources agreed with the informant excluded from all three before main was fast-forwarded.
- **Standing rule this entry exists to enforce:** a resume brief's RULINGS are durable; its STATE
  is a HYPOTHESIS to re-derive before acting.
- **Kin:** [[IR-21]] and [[IR-39]] (the diagnostic names the wrong actor / the wrong subject),
  [[IR-32]] (a gate that cannot see its own blind spot), [[IR-34]] (an intermittent marker read as
  a signal).
- **Ripe when:** the next commune/psyche-tier touch for arm A (the routing layer is spt-core's own
  `spool`/psyche ingest, so a remedy is in-repo, not upstream); for arm B, the next `golden.yml`
  informant touch — the cheap fix is that the notify step names its own scope in the message body
  (STEP vs RUN) rather than emitting a bare "CI SUCCESS".
- **Size:** arm B small (message text + scope word in one workflow step). Arm A unsized —
  establishing whether ANY durable content-recency signal can ride the envelope is the measurement,
  and until it runs the answer is "re-derive, do not rank."

### IR-41 — A QUEUED main run is superseded without a record; `cancel-in-progress: false` protects only STARTED runs
- **Status:** open — mechanism CONFIRMED (doyle ruling 2026-08-05); the runbook sentence it
  falsified is already corrected (`docs/RELEASE-RUNBOOK.md` step 3, same commit as this entry)
  · **Origin:** deployah, v0.55.0 publish night, filed by him explicitly as UNPROVEN with the
  manual-cancel alternative not ruled out; confirmed by doyle on the run objects.
- **What/why:** `ci.yml` sets `cancel-in-progress: false` on main and the runbook read that as
  "main's concurrency policy never cancels this run; that preserves the record." FALSE for the
  queued phase: a concurrency group holds at most one running + one pending run, and a newer push
  REPLACES the pending run regardless of the flag — the flag governs only whether a STARTED run
  is cancelled. A superseded queued run leaves NO record: zero jobs, conclusion `cancelled`,
  nothing measured about its sha.
- **Evidence (all read off the run objects, not the informant):** run 30973909279 at `327f1f8`
  (push, main) — created 04:01:08Z, **zero jobs ever started**, cancelled 04:03:49Z, ONE SECOND
  after `e8805f7`'s run 30974041521 entered the group (created 04:03:48Z), while `0a25b77`'s run
  30973770039 was still in flight (done 04:07:59Z). One-running-one-pending, newer push landed,
  pending run died. The 1s coupling to an unrelated push is the discriminator against deployah's
  own manual-cancel alternative — a human cancel co-timed to the second with a push it had no
  view of is not a credible mechanism; supersession fires on exactly that trigger by design.
- **The hole this leaves:** under a rapid push sequence, an intermediate sha on MAIN can have no
  thin-CI verdict AT ALL, and the absence reads as nothing rather than as supersession. Anyone
  auditing "did sha X pass main CI" gets a hole where the runbook promised a record. Absence of
  a run is now a labelled state, not evidence about the sha.
- **Ripe when:** next ci.yml/runbook touch. Candidate remedies to weigh THEN, not now: accept and
  document (done — the runbook now carries the mechanism), or give main's group per-sha keys
  (`group: ci-${{ github.sha }}`) so runs never share a group — at the cost of concurrent main
  runs competing for the boxes, which is exactly what the group exists to prevent. The trade is
  real; do not fold it into a drive-by.
- **Size:** docs half DONE in this commit; ci.yml half small but load-bearing — needs its own
  lane if taken.
- **Kin:** [[IR-40]] (a record that is right-in-the-end-but-early vs a record that never comes to
  exist — both read as clean states unless labelled), [[cancelled-measurement-leaves-labelled-hole]].

### IR-42 — `pool-claim` writes a record; only the BUILD enforces — and AGENTS.md invited the misread
- **Status:** BUILT 2026-08-19 — the AGENTS.md corrective line lands in the SAME COMMIT as this
  entry (the register batch commit), which is the entry's stated exit condition. The optional
  code-side nicety (claim verb prints the incumbent record it overwrites, information only)
  remains unclaimed — a candidate rider on the IR-26 remedy lane, which already touches
  `pool_claim`'s read-before-write path, NOT a reason to keep this entry open ·
  **Origin:** doyle + todlando independently at source, 2026-08-19, during the #193 gate.
- **Mechanism (source-verified):** `xtask pool-claim` (crates/xtask/src/main.rs:2232-2294 at the
  read sha) parses its args, reads the lane's own git identity, builds a `PoolOwner` and calls
  `spt_poolguard::write_owner` UNCONDITIONALLY (:2280) — no `read_owner`, no verdict, no
  comparison against the incumbent claim anywhere in the function. Claiming is last-writer-wins
  by construction: two lanes can each "hold" a pool in sequence with only the last write
  surviving, and the displaced lane learns nothing at displacement time. All four enforcement
  arms — `Refuse` (SPT_POOL_FOREIGN), `Takeover` (loud), `Unproven`, `HatchOpen` — live in
  crates/spt-store/build.rs:31-113 and speak only at the next BUILD. (Same object [[IR-26]]'s
  CLAIM-OBSERVABILITY measurement saw from the displaced side; this entry is the CLAIMING side
  plus the doc that invited the misread.)
- **Measured consequence:** 2026-08-19, doyle's gate claim over todlando's unlanded
  redeem-echo-diag claim — no refusal, no notice (three sequential overwrites that session, one
  over a lane with an unlanded commit). todlando then predicted from AGENTS.md's refusal
  sentence that the gate's claim "will be REFUSED" — a false warning to a gater mid-run. The
  refusal sentence sat directly under the "Claim a pool at lane start" instruction, inviting the
  build-time semantics to be read onto the claim verb.
- **Corrective landed:** one AGENTS.md sentence — the claim WRITES a record and never refuses;
  predict refusals from builds, never from `pool-claim`.

### IR-43 — knock answer carrier: `NoReply` landed at 60.08s against a stated 30s carrier deadline
- **Status:** open, OBSERVATION — mechanism unmeasured, filed exactly as wide as the datum ·
  **Origin:** KNOCK-169 RCA layer 1 (todlando, read-only), golden 32191618042 @`93c1130`,
  twohost-b, 2026-08-18.
- **What/why:** B's `request_answer` yielded `NoReply` at 22:50:08.3605Z — 60.08s after A's
  death — while the carrier deadline is stated at 30s. Recorded in the RCA as a side observation
  with no claim; the cell's red itself was a CASCADE of A's earlier death (answerop.rs:73-81
  produces the courtesy BEFORE returning, so the reply was never sent — that half is settled and
  is NOT this entry). What this entry carries: a transport/liveness bound that reported at 2x its
  stated value. Candidate shapes, neither asserted: two sequential 30s legs each honoring its own
  bound (a composed wait wearing one bound's name), or a bound armed after a wait it does not
  cover (the [[IR-30]] nethost.rs permit-wait shape: `sem.acquire_owned()` outside the timeout).
- **Ask:** derive where the 30s is armed and what path yields a 60s `NoReply` — one code read
  from the emit site, before any instrument.
- **Ripe when:** next knock-carrier touch, or the next `NoReply`-shaped red — whichever first.
- **Size:** small (code read; a token only if the read forks).

### IR-44 — perch-sentinel comment overclaims: serialization is not preservation, and the comment teaches the false step
- **Status:** open · **Origin:** #188 RCA lock read (todlando flagged, doyle confirmed the lock
  read), 2026-08-18; the collapse-of-branches argument is on #188 and in the RCA record.
- **What/why:** `lock_perch_sentinel` (info.rs:778 at `b20e770`, the diag sha — re-derive the
  line on main before editing) takes a cross-process fs2 lock on a stable per-perch `.info.lock`
  sentinel, and its comment claims it serializes "ALL info.json writers so a whole-record write
  and a locked RMW can never lose each other's update (`REQ-HAZARD-INFO-RMW-LOST-UPDATE`)" —
  TRUE for interleaving, FALSE for preservation. `write_info_unlocked` is module-private with
  exactly three callers, all taking the sentinel first (bypass refuted by visibility), but every
  caller of the public `write_info` composes its record BEFORE the lock is taken (:842) — **the
  lock makes that write atomic; it cannot make it preserving.** Only `mutate_info` /
  `establish_locked` read under the hold and can preserve. Measured field case: census write #1
  (`twohost.rs:736`) clobbered `controlled` while correctly locked. A reader trusting the
  comment infers lost-update immunity the funnel does not provide — the #188 hunt burned a
  branch on exactly that inference before the collapse argument killed it.
- **Ask:** respell the comment — serializes writers; a composed `write_info` is atomic, never
  preserving; preservation requires `mutate_info`/`establish_locked` — and weigh whether
  `REQ-HAZARD-INFO-RMW-LOST-UPDATE`'s doc wording carries the same overclaim.
- **Ripe when:** next spt-store/info.rs touch. **Size:** tiny (comment + possibly one REQ doc
  line).

### IR-45 — twohost rig home premise, SETTLED: pump paths resolve under the per-run TEMP root (rig is disk-hermetic; the "live fleet roster" reading is retired)
- **Status:** RETIRED 2026-08-19 — ask answered by the one code read at main @`78a9a16`; residue
  re-homed (see Answer). · **Origin:** #188/#189 RCA readouts (todlando, doyle's 2a-2d), runs
  32097943571 + 32108557362 + 32111921251, 2026-08-18.
- **What/why, two measured halves in tension:** (1) B's pump dialed a THIRD node 155-159 times
  per run at kitsubito's OWN tailscale IP (`100.98.197.12`), different hex + different ephemeral
  port each run — a short-lived local endpoint re-minting identity between runs, present in both
  the golden and the diag run. Read at the time as: the rig uses the CANONICAL home
  (`canonical_pump_paths` → `perch::spt_home()`), so B's roster is kitsubito's live fleet
  roster — environmental coupling of a CI rig to host fleet state. (2) The ER perch measured in
  the chain runs resolves under a per-run TEMP root
  (`…/_temp/spt-test-tmp-<run>…/owlery/engine-room`), NOT the canonical home. Both measurements
  stand; they answer different artifacts (pump roster vs perches), and nothing on record settles
  which home the PUMP paths actually resolve. #189's third-node evidence and its
  cache-leg-never-heals reading rest on that premise (the tension is filed on the #188 record as
  an open question against #189).
- **Answer (2026-08-19, code read at `78a9a16`, unambiguous — no run token needed):** both role
  tests set `SPT_HOME` to a fresh per-run `TempDir` as their FIRST act (`twohost.rs:813-814` role
  B, `:1501-1502` role A), before any store touch. `canonical_pump_paths` resolves every path via
  `perch::spt_home()` (`twohost.rs:396-414`), which honors the `SPT_HOME` override first
  (`perch.rs:34-41`); the pump's roster source `presence::registry_snapshot_dir()` →
  `perch::identity_dir()` sits under the same root (`presence.rs:83-85`). The rig is
  DISK-HERMETIC per run; both prior measurements reconcile (ER perch observed under the temp
  root, pump paths resolve there too). The misleader was the helper's name + doc comment ("real
  homes, not test roots"), which describe production path LAYOUT, not the resolved ROOT.
- **Residue disposition:** (1) #189's third-node evidence was corrected on the board (comment
  5337519936) — the roster row arrived at RUNTIME into a temp-rooted store, ingress mechanism
  OPEN on #189, no longer explained-environmental. (2) KEYSTONE #182's hygiene lane repaired the
  `twohost.rs:394` comment to say exactly what the helper does: production path layout rooted in
  the process's fresh per-run `SPT_HOME`.

### IR-46 — the Windows disk-floor preflight asserts an INSTANT; workspace free space moves tens of GB inside the hour, so a green floor is not a claim about run headroom
- **Status:** open — mechanism CONFIRMED by a measurement series on hfenduleam; doyle ruled it
  register-shaped 2026-08-18 · **Origin:** deployah, NAMEPLATE #181 golden-head intake, on the
  run that the floor red-floored · **Filed as IR-46 after an id collision:** this row was authored
  as IR-42 and held unpushed under the v0.56.0 tag freeze, during which the release-close sweep
  (`78a9a16`) issued IR-42..45 to other findings. Parallel issuance, not authored disagreement —
  and noted here so a later reader chasing "IR-42" in this row's history does not hunt for a lost
  version of it.
- **What/why:** `golden.yml`'s Windows preflight computes `$freeBytes` once at job start and
  hard-fails under a `32GB` floor. As a fast-fail that is correct and it did its job — it refused
  before compiling rather than dying deep in a link step. The defect is what a PASS is then read
  to mean. The reading is a point sample of a quantity that moves by tens of GB unattended, so a
  green preflight licenses a run whose headroom was never measured. Nothing downstream re-checks.
- **Evidence — five readings of `C:` free on one box inside roughly one hour, 2026-08-18:**
  - `03:01:54Z` — `free_bytes=2666205184` (2.48 GiB) against `floor_bytes=34359738368`, the
    preflight's OWN reading; run `32093894524`, both `n1-gate` and `test` on Windows/hfenduleam
    failed at this same step, before any compilation.
  - `~03:2xZ` — ~2.50 GB, independent `Get-PSDrive` read, corroborating the preflight.
  - pre-reap — **33.70 GB**, immediately before deployah deleted anything.
  - post-reap-1 — 76.42 GB (42.72 GB reclaimed: `claude_skill_owl/target`, `golden-w1/target`).
  - pre-reap-2 — **73.81 GB**, i.e. 2.61 GB consumed unattended between the two reaps; then
    post-reap-2 129.70 GB (55.89 GB reclaimed: `.worktrees/nameplate-w1/target`).
- **LABELLED HOLE — do not let this harden:** roughly **31 GB was released between the preflight
  failure and the pre-reap reading by something OTHER than the reclaim**. Released-by-unknown.
  The runner cleaning its workspace after the job died is *plausible* and is NOT the recorded
  cause — it was never measured, and no one looked while it was happening. Recorded as a hole on
  purpose (doyle's explicit ask at filing): a plausible cause written down as the cause would make
  this row read as explained when the actual mechanism of the 31 GB is unknown.
- **The hazard this leaves:** a run can clear the floor at preflight and starve mid-build, and
  mid-build disk starvation does not present as disk — it surfaces as a link failure, a truncated
  artifact, or a rustc ICE, i.e. as a defect in the tree under test. The failure mode inverts the
  gate's purpose: the preflight's whole point is to keep box conditions from being read as lane
  reds, and a passing preflight actively argues the opposite. Note the polarity — this row is
  about the PASS, not the FAIL. The observed red was honest.
- **Ripe when:** next `golden.yml` touch. Remedies to weigh THEN, not as a drive-by: re-assert the
  floor at job END (turns a starvation into a labelled disk verdict instead of a fake lane red);
  or raise the floor to cover a full build's peak rather than its entry condition; or sample free
  space across the run and emit the minimum as a token. All three cost run time; none is obvious.
- **Size:** small in `golden.yml`, load-bearing in what a green run is taken to prove.
- **Kin:** [[IR-41]] and [[IR-40]] (a state that reads as a clean verdict unless it is labelled —
  here a PASS that is read as headroom it never measured), [[measure-the-box-before-the-instrument]],
  [[cancelled-measurement-leaves-labelled-hole]].

### IR-47 — ci-notify treats a missing co-author trailer as a quiet info line; silent attribution loss reads as "there was none"
- **Status:** open — patch AUTHORED and stashed unlanded since 2026-07-29
  (`.worktrees/_patches/ci-notify-missing-trailer.patch` + its `test-ci-notify.sh` harness in the
  sibling `.untracked` dir); surfaced by doyle's 2026-08-19 idle-queue verification of that stash ·
  **Origin:** the AGENTS.md trailer mandate's own hazard family (the space-spelling trailer is
  structurally invisible to git's tokenizer, so confident zeros already read as attribution loss
  once); this row is the NOTIFY-side twin.
- **What/why:** `.github/ci/ci-notify.sh` parses the head commit for the line-anchored
  `Co-authored by: <agent>` trailer to add a second notification recipient. When the trailer is
  absent or unparseable, the current script emits one stdout info line ("doyle only") and moves
  on — correct as a non-fatal outcome, wrong as a SILENT one: a lane that lost its attribution
  (typo'd trailer, squash that dropped the body, hyphenated spelling) notifies doyle alone on
  every run and nothing anywhere says an agent stopped being told about their own lane's verdicts.
  The stashed patch keeps absence non-fatal but promotes it to a `::warning` workflow annotation
  (visible on the run summary), and splits the legitimate quiet case (co-author IS doyle) from the
  loss case so the two stop sharing one message.
- **Evidence:** patch verified 2026-08-19 to still NOT be on main — the warning string is absent
  from `.github/ci/ci-notify.sh` at the current tree. Patch content re-read at verification; it
  applies against the script's post-`SEND_BOUND_SECS` shape.
- **Ripe when:** next `golden.yml`/ci-notify touch — natural co-rider with [[IR-46]]'s remedy
  window (same file family, same "weigh then, not drive-by" rule). The stashed test harness rides
  with it.
- **Size:** ~10 lines in one shell script plus its test.
- **Kin:** [[IR-40]] (an unlabelled state read as a clean verdict — here a quiet info line read as
  "no co-author existed"), the AGENTS.md `%(trailers:)` tokenizer mandate (the sibling silent-zero
  on the AUDIT side).

### IR-48 — the nextest.toml parity cell tolerates CRLF only by parser accident; a newline-sensitive arm added later reds fresh Windows checkouts alone
- **Status:** open — filed 2026-08-19 (doyle route, hertz verdict: REGISTER as preventive
  hardening, not current defect) · **Origin:** sibling sweep after the W3 gate CRLF finding —
  parent mechanism: `include_str!` embeds working-tree bytes verbatim, `core.autocrlf=true`
  smudges fresh checkouts CRLF, and the author's never-re-smudged tree stays LF — builder green,
  every fresh rig/golden checkout red.
- **What/why:** `crates/xtask/src/main.rs` `the_checked_in_config_is_in_parity` include_str!s
  `.config/nextest.toml` and feeds `phase_a_overrides_missing_from_ci_windows`. TODAY this is
  causally CRLF-tolerant — the parser is `str::lines()`+trim (`override_filters`, main.rs:434-440),
  and the 52/52 fresh-rig green at the hygiene gate (2026-08-19, this box) is therefore causal, not
  luck. The hazard is the MISSING CONTROL: nothing pins the tolerance, so a future
  newline-sensitive arm in the parity predicate reds only on fresh Windows checkouts — the
  builder-green/rig-red inversion, deterministic but reading as flake.
- **Remedy (hertz's, platform-independent):** feed the parity predicate an LF fixture AND the same
  fixture converted to CRLF, assert identical missing-sets (including one unmirrored negative), so
  a newline-sensitive arm reds on EVERY host rather than only fresh Windows checkouts.
- **Ripe when:** next xtask parity-cell touch; explicitly OUTSIDE the family-B lane (hertz
  scoping). **Size:** one fixture-pair cell.
- **Kin:** the W3 gate finding this swept out of (skeleton cells, fixed at the fixture edge in
  lane commit e99633f); [[IR-40]] (an untested tolerance read as a guarantee).

### IR-49 — poolguard's landed-lane predicate is ancestry-only, so a cherry-picked member lane can never read Settled and its pool can never be taken over
- **Status:** open — filed 2026-08-19 (todlando measurement at the keystone-w1/fix-197 reap,
  doyle ruling; doyle re-verified the predicate at source before filing) · **Origin:** todlando's
  post-landing landed-check used the guard's own predicate and it contradicted the (correct)
  landed fact — the checker was wrong with the guard, which is exactly how the guard will be
  wrong alone.
- **What/why:** `git_lane_state`'s final arm answers Settled/InFlight purely by
  `merge-base --is-ancestor <lane tip> <integration head>`
  (`crates/spt-poolguard/src/lib.rs:424-436`, probe at `:474-479`). Under ADR-0050 golden CI
  (thin member lanes cherry-picked onto a stage branch, ff-only main) a MEMBER lane's own branch
  tip never becomes an ancestor of main. Measured at origin/main `9ea595c`:
  `build/keystone-182-w2-sealed-multi` @`ba2d998` and `fix/197-absent-code-uncounted` @`5eb6a3b`
  both read NOT-ancestor while `git cherry -v` reports every lane commit `-` (already upstream).
  So `LaneState::Settled` is unreachable for any cherry-pick-assembled member lane whose branch
  still exists, the takeover arm can never fire, and the pool refuses forever — including with a
  dead holder. That contradicts the AGENTS.md mandate ("a merged branch is taken over") because
  the guard's notion of merged is ancestry-only and the golden model structurally never produces
  ancestry for member branches. The branch-vanished (`:391`) and re-pointed (`:399`) arms still
  fire, and the STAGE lane works (its tip IS an ancestor) — the hole is landed-member-branch-
  still-exists only. It did not bite at the 2026-08-19 reaps solely because deleting a pool
  deletes its claim; it bites the first time someone wants to TAKE OVER a landed lane's pool
  rather than reap it.
- **Remedy (todlando's candidate, doyle-endorsed direction):** when ancestry answers false, fall
  back to patch-id containment (`git cherry` / `rev-list --cherry-mark` against the claimed
  base) before concluding InFlight. **Known residual to state in the fix's docs/tests:** a
  conflict-adjusted pick changes the patch-id, so such a lane still reads InFlight; its release
  path is branch deletion (the vanished arm), which is at least reachable. Tests want a
  cherry-picked-landed fixture plus a conflict-adjusted negative.
- **Ripe when:** next poolguard touch, or the first landed-lane pool takeover request.
- **Size:** one fallback arm in `git_lane_state` + two fixtures.
- **Kin:** [[IR-26]] (dead-holder takeover AUTHORIZED when it shouldn't be — this is the mirror:
  takeover UNREACHABLE when it should fire), [[IR-42]] (claim writes, build enforces), the
  AGENTS.md pool mandate this contradicts.

### IR-50 — e2e failure panels read the PRE-REDIRECT stderr capture; every daemon-side diagnostic lands in a sink the rig deletes unread
- **Status:** BUILT 2026-08-19 — lane `fix/ir50-stderr-sink-census` @`1a2f6a2`+`2ac9543`
  (hertz; base `0c86b2d`; rides the CONCIERGE #183 chain; these shas are the 2026-08-19
  message-only reword of `cb4cd9d`+`7336e03` — trees proved identical by tree-id, so every
  measurement below carries): exported
  `spt_daemon::stderrlog::sink_path(home)` with `stderr_log_path` routed through it +
  equality unit; `daemon_stderr_panel` in tests/common renders BOTH channels and labels a
  read failure with channel + exact looked-at path + OS error (absent-sink regression pins
  it — the IR-40 unlabelled-absence class was caught at gate and fixed in `2ac9543`);
  censused every daemon-run captured-stderr panel outside `engine_room_bringup_e2e` (rides
  #199 lane) and `er_briefing_*` (rides #164 lane); engineroom.rs:145 misnomer rider landed
  in the same lane. Gate figure 781/781 at `2ac9543` (measured at the pre-reword tree,
  which is byte-identical; first-report 765 was transcription,
  corrected with raw line + JSON provenance). CLASS FOLD (doyle ruling, no fresh row):
  `er_briefing_presented_e2e` hit this exact class during its own authoring (blank panels
  on a real loud line) — fixed locally in `72efb6a` reading both channels; RESIDUAL: that
  cell re-spells the sink path and should adopt `sink_path()` once both lanes land
  (hertz's base predates the cell, so adoption is a post-land follow-up). ·
  **Originally filed:** 2026-08-19 (todlando RCA inside the #199 investigation; the immediate
  four-panel fix in `engine_room_bringup_e2e.rs` is his and rides the #199 lane — THIS entry is
  the class census, hertz-class) · **Origin:** the #199 evidence void — 40 rca-v2 runs held zero
  daemon-side evidence for either erhost face; `ENGINE_ROOM_SPAWN_FAIL` (broker.rs:5658) was
  written to a log the rig destroyed at teardown. Absence discipline honoured: the SUCCESS-path
  line (`ENGINE_ROOM_BROUGHT_UP`) was also absent from all 40, proving the channel dead rather
  than the event absent.
- **What/why:** product `stderrlog::install` (broker at cli.rs:7633, brain at brainproc.rs:200)
  repoints the process STD_ERROR_HANDLE within its first statements (`redirect_stderr_to`,
  stderrlog.rs:101 — Windows `SetStdHandle` + `mem::forget`), so a rig's `.stderr(File)` capture
  holds only the pre-redirect window; from that line on every diagnostic lands in
  `SPT_HOME/logs/daemon.stderr.log` (stderrlog.rs:34/42) inside the rig's temp home, destroyed
  unread at teardown. Panels are also mislabelled ("brain stderr" holding the broker's first
  lines). REQ-DAEMON-STDERR-PERSIST fixed this blindness PRODUCT-side ("the incident-night
  RCA-blind gap", cli.rs:7626-7631); no rig was ever taught to read the sink, so the fix does
  not reach any harness. Class: every e2e rig that spawns `spt daemon run` and prints a
  captured-stderr panel has the same void.
- **Remedy:** census all such rigs; failure panels additionally dump
  `SPT_HOME/logs/daemon.stderr.log` with the path read FROM `spt_daemon::stderrlog` (never
  re-spelled), and the labels corrected ("broker stderr (pre-redirect)" / "daemon stderr sink").
  `engine_room_bringup_e2e.rs`'s panel sites land in the #199 lane (todlando, scoped there —
  measured shape: FIVE producers incl. the :167 precondition panic, SEVEN print sites, seven
  relabels; the early "four" was an estimate the file refuted). ⚠ Rigs can only import
  `STDERR_LOG_BASENAME`; the `"logs"` dir segment has no exported const, so every rig
  re-spells it — the census lane should add an exported path helper (e.g.
  `stderrlog::sink_path(home)`) to spt-daemon FIRST, then consume it everywhere. The census
  and remainder are hertz-class.
- **Ripe when:** hertz's queue after current items, or the next e2e red printing an empty
  stderr panel.
- **Size:** census + mechanical panel edits per rig.
- **Kin:** [[IR-40]] (an absent signal read as a clean verdict), REQ-DAEMON-STDERR-PERSIST (the
  product half of the same incident), [[IR-8]] (a censused zero that cannot see its blind spot).

### IR-51 — engine-room e2e daemon runs NET-ENABLED on real interfaces; rig hermeticity is disk-deep only
- **Status:** open — filed 2026-08-19 (todlando's x40 run-4 sink read, the first live read any
  rig has had of the daemon sink; doyle ruled own-id rather than riding #199, at the finder's
  own request) · **Origin:** unasked-for find inside the #199 instrument's first catch.
- **What/why:** the rig's temp `SPT_HOME` buys DISK hermeticity ([[IR-45]]'s settlement — which
  explicitly covered paths, not network) but nothing turns the net off: the run-4 sink shows
  the rig's own daemon with `BRAIN_NET_CONSUMERS_UP` (dispatcher + peer pump started), three
  `NET_FAMILY_GATE: binding IPv4-only`, and three `PAIR_MEET_UP:erhome` carrying the box's REAL
  Tailscale (100.68.35.65) and LAN (192.168.1.81) addresses — live peer work during an e2e —
  plus an `INBOUND_REACHABILITY` warning whose firewall rule admits the CI-runner's binary path
  (`C:\actions-runner\_work\spt-bs-core\spt-bs-core\target\debug\spt.exe`) while this daemon
  ran from a dev tree. Consequences: test behavior conditioned on shared-box network state;
  rate-shaped flakes; e2e runs radiating real traffic. SAME SINK, RECORDED NOT DIAGNOSED: conn
  churn at ~150 ms cadence (327 write-start / 324 transport-close, conn ids past 300 in 58 s) —
  unattributed. ⚠ CANDIDATE INTERACTION, explicitly NOT established (finder's own framing
  honoured): the run-4 face (30 s bound blown, zero session events in the window, launch Ok, no
  error anywhere) matches the SHAPE of the filed offline-but-resolvable-peer pump-stall
  mechanism (dial does not fast-fail). Discrimination open; nothing here is a #199 attribution.
- **Remedy:** a rig-honoured net-off switch (peer pump + discovery disabled under test unless
  the test is ABOUT them) or loopback-only binding; census which e2e rigs start net consumers.
  Sequencing rule: do NOT fix hermeticity into #199's face before the face is attributed — a
  green bought by turning the net off would bury the mechanism unread.
- **Ripe when:** #199 attribution answers whether net state is causal; else the next
  hermeticity pass.
- **Size:** product switch + rig adoption + census.
- **Kin:** [[IR-45]] (the disk half), the subnet-pump dial-does-not-fast-fail mechanism (filed,
  pump-stall RCA), [[IR-38]] (a live daemon's side effects outliving the test's intent),
  releases#125 gc-spin (a conn-churn shape candidate).

### IR-52 — docs-site CLI reference renders SHALLOW; help text below the rendered depth is outside the drift gate, unit-held per-REQ or held by nothing
- **Status:** open — filed 2026-08-19 (doyle, W3 #160 gate prep; second measured instance of a
  class REQ-CLI-SURFACE-SECTION-SITED's own title already named for its sites) · **Origin:**
  todlando's #160 build report, fact flagged for the gate rather than folded silently.
- **What/why:** the generated `docs-site/src/cli/reference.md` renders `spt endpoint monic`
  but does not descend to `monic add` / `monic update`, so the #160 strike of the maintained
  kind enumeration at those two flag docs never appeared in the published reference and the
  docs-drift gate CANNOT see the strike — the REQ-CLI-MONIC-TRIGGER-SECTION rendered-help
  unit is the only hold. General mechanism: any help text below the generator's descent
  depth is invisible to the drift gate; a deep help edit (or regression) ships unpublished
  and ungated unless some REQ's unit happens to pin it. First measured instance:
  REQ-CLI-SURFACE-SECTION-SITED names "the three-deep sites the docs-site drift gate cannot
  reach" and closes them with its own pinned walk — per-REQ self-defense, not a gate.
  Consequence beyond drift: adapter builders build BLIND from the public docs (DRI
  protocol), so a deep-help divergence hands every adapter a thinner contract than the
  operator's terminal shows.
- **Remedy:** either xtask reference-gen descends the full command tree (deep help becomes
  published surface the drift gate already covers), or a generic rendered-help walk over all
  sites at all depths joins the drift gate; census which REQ units currently self-hold deep
  sites (known: SURFACE-SECTION-SITED walk, MONIC-TRIGGER-SECTION sited-help unit).
- **Ripe when:** next xtask docs-gen touch, or the first adapter filing traceable to a
  deep-help divergence.
- **Size:** xtask render change + regen + census of self-holding units.
- **Kin:** the REQ-CLI-SURFACE-SECTION-SITED pinned walk (the per-REQ closure shape this
  entry generalizes), [[IR-47]] only if its golden.yml/ci-notify surface rides the same
  docs-gate window.

### CI-RIDER LANE STATE — LANDED 2026-08-04 (post-v0.53.0 merge queue); history below kept for its mechanisms
- **RESOLVED:** the lane rebased clean onto the post-tag queue and landed ff-only as
  `2b33a47`/`2a7b016`/`d63f4ce`/`19d7f79` (content byte-identical to `1275e47..bf8c4a2` by lane-diff
  blob hash). IR-1 and IR-4 are BUILT AND LANDED; IR-9's measurement half is landed with its trigger
  now armed — the toolchain print reports runner-account versions on the NEXT GOLDEN RUN, which is
  when doyle's 2026-08-03 ruling (decide from runner-account versions, never the interactive prior)
  becomes executable. The two golden-only steps (link probe both boxes, toolchain print both legs)
  remain unexercised until that run — the landing does not change that caveat.
- **As recorded pre-landing (2026-08-04, doyle):** `ci/locksmith-riders` was at `bf8c4a2` with FOUR
  commits not contained in `origin/main` and not present in LOCKSMITH's golden head `b7b00c3`. The
  worktree `.worktrees/hertz-ci-riders` was clean, so the work existed and was simply unlanded:
  `1275e47` (xtask: assert a load-bearing patch pin is still in force, both ways it lapses) ·
  `49d4805` (docs/ci: a lock-touching lane reads edges, not just the package set) ·
  `8ed006b` (ci/golden: print the toolchain that judged the run, both legs) ·
  `bf8c4a2` (ci/golden: measure the link before rendezvous, read the floor after the checkout that
  clears it — the floor read AFTER its own reclaim is [[ci-runner-has-no-warm-target]]'s shape).
- **Why it is recorded rather than quietly re-dispatched:** [[IR-1]]/[[IR-4]]/[[IR-9]] were carried
  as DISPATCHED, which is now false in BOTH directions — the work is further along than dispatched,
  and it also did not ship. The cause was the same agent outage that cost `#123` its build, and the
  operator's standing rule from that drop applies here too: an outage must not be able to remove
  work from a batch silently. Board requests got a drop comment; register entries get this.
- **MAPPING CONFIRMED BY ITS BUILDER 2026-08-04, by REQ id rather than by recollection** — hertz
  noted first that its own context had been cleared between building the lane and answering, and
  declined to testify from memory about the original instrument. Everything here is either measured
  that day or read off the commits:
  `1275e47` → `REQ-CI-LOAD-BEARING-PATCH-PIN` ([[IR-4]] part 1, the mechanical pin guard in xtask) ·
  `49d4805` → `REQ-LOCK-TOUCHING-LANE-PROCEDURE` ([[IR-4]] part 2, procedure in docs/GOLDEN-CI.md) ·
  `8ed006b` → `REQ-CI-TOOLCHAIN-VERSION-PRINT` ([[IR-4]] part 3, carrying [[IR-9]]) ·
  `bf8c4a2` → `REQ-CI-LINK-HEALTH-PROBE` ([[IR-1]], the network axis).
- **CORRECTION TO THE GATER'S EARLIER WORDING, from the builder:** `8ed006b` does NOT discharge
  [[IR-9]] — it discharges IR-9's MEASUREMENT half only. IR-9's other half (align the two boxes, or
  declare one authoritative clippy leg) is held by doyle's own 2026-08-03 ruling: decide once the
  step reports RUNNER-ACCOUNT versions, never on the interactive-account prior. IR-9 therefore stays
  OPEN with its trigger now REACHABLE, which is a different state from "waiting".
- **Gate evidence, instrument NAMED:** `cargo nextest run -p xtask` in `.worktrees/hertz-ci-riders`
  on hfenduleam, box quiet with no runs in flight — 34 run, 34 passed, 0 skipped. Nextest rather
  than bare `cargo test`, per [[IR-25]]. ⚠ **What that does NOT cover, stated by the builder rather
  than discovered later:** the two golden steps the lane ADDS (link probe on both boxes, toolchain
  print on both legs) cannot be exercised outside a golden run. The probe's three arms
  (five-sample success, no-reply, absent-CLI) were exercised on the real boxes at commit time, and
  that is reported as a CLAIM RECORDED AT COMMIT TIME, not as a re-verification — the kitsubito arm
  could not be re-checked from here. This is why the lane wants a golden rather than a quiet ff.
- **[[IR-5]]/[[IR-6]] CONDITIONAL RIDERS: CONDITION NOT MET — the answer is NO and it closes the
  question for this lane.** Measured file set `6d291e0..bf8c4a2`: `.gitattributes`,
  `.github/bench/link-probe.ps1`, `.github/bench/link-probe.sh`, `.github/workflows/golden.yml`,
  `crates/xtask/src/main.rs`, `crates/xtask/tests/bench_row_parity.rs`, `docs/GOLDEN-CI.md`,
  `traceable-reqs.toml`. Nothing under `.github/ci/`, where the nextest-summary readers and
  count-reporting scripts live. IR-5: no consumer of nextest output touched. IR-6: no EXISTING gate
  output changed — though the NEW output was voluntarily written to IR-6's rule (the LINK line
  prints `samples=N/M` and the individual rtts, membership beside the count, in both shells). That
  SHRINKS IR-6's future scope by two lines; it does not discharge it.
- **Three mechanisms lifted from this lane's REQ titles, worth more than the lane:**
  (a) **An instrument must not be able to RED the run, and each shell breaks that differently** —
  bash: `| head` under `pipefail` SIGPIPEs the producer to 141, so a print step reds (trim with
  parameter expansion instead); pwsh under GitHub's wrapper: a MISSING COMMAND is a terminating
  error, so rustup's absence must be TESTED with `Get-Command`, never caught. One requirement, two
  constructions, and neither is portable reasoning about the other.
  (b) **A negative-control mutation must be COUNTED before it is trusted** — the absent-CLI arm was
  first "measured" against a tree where the mutation had not taken, silently re-measuring the
  unmutated arm and reading as a pass.
  (c) **Jitter, not the median, is the signal on a link that mostly works** — the motivating red's
  link ran 10..83ms and 3..75ms while IDLE, so a median-only probe reads healthy straight through
  the failure. Both med and max rows ride the ledger.
- **Disposition:** composes onto the next golden batch as a thin lane, NOT slipped onto main. It
  changes the golden pipeline itself, so it wants a golden rather than a quiet ff — and while
  v0.53.0 is untagged, any push to `main` moves the runbook's bare `git tag` off the tested sha.


Wall time: 0.36 seconds