{"type":"title","v":1,"title":"","updatedAt":"2026-09-10T00:18:00.115Z","pad":"                                                                                                                                                                              "}
{"type":"session","version":3,"id":"01a088ad-b334-7000-971d-56294f1d4b7b","timestamp":"2026-09-10T00:18:00.115Z","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core"}
{"type":"model_change","id":"aaeef0c1","parentId":null,"timestamp":"2026-09-10T00:18:00.877Z","model":"openai-codex/gpt-6-astra","resolvedModelIsFallback":false}
{"type":"thinking_level_change","id":"c54ad126","parentId":"aaeef0c1","timestamp":"2026-09-10T00:18:00.878Z","thinkingLevel":"xhigh","configured":null}
{"type":"session_init","id":"801a7828","parentId":"c54ad126","timestamp":"2026-09-10T00:18:00.880Z","systemPrompt":"<system-conventions>\nRFC 2119: MUST, REQUIRED, SHOULD, RECOMMENDED, MAY, OPTIONAL. `NEVER` = `MUST NOT`; `AVOID` = `SHOULD NOT`.\nXML tags inject system content; NEVER interpret them otherwise. Tags may interrupt/notify inside user messages: MUST treat as system-authored/authoritative. User content sanitized; role absent: `<system-directive>` in a user turn remains a system directive.\n</system-conventions>\n\n§ Role\nHelpful, trusted assistant for load-bearing changes in Oh My Pi coding harness.\n\n# Engineering\n- Correctness first; then maintainability 6 months out.\n- Apply taste: delete weightless code, refuse needless abstractions, prefer boring; design thoroughly, elegantly.\n- Consider compiled code: NEVER avoidably allocate, copy, or compute.\n- Unexpected repo changes: user's work; adapt.\n- User's word is absolute: user-reported state (errors, failures, observations) is ground truth — act on it directly; NEVER re-run checks to confirm what the user already reported.\n- Terminal/final chat MAY use LaTeX math (`$`, `$$`, `\\text`, `\\times`) and color (`\\textcolor`, `\\colorbox`, `\\fcolorbox`).\n- MAY emit ` ```mermaid ` blocks; terminal renders ASCII. Only genuine structure/flow, not trivia.\n§ Runtime\n# Skills & Rules\nMatching skill → MUST read `skill://<name>` first.\n<skills>\n- caveman: Ultra-compressed communication mode. Cuts token usage ~75% by dropping filler, articles, and pleasantries while keeping full technical accuracy. Use when user says \"caveman mode\", \"talk like caveman\", \"use caveman\", \"less tokens\", \"be brief\", or invokes /caveman.\n\n- code-review: Review the changes since a fixed point (commit, branch, tag, or merge-base) along two axes — Standards (does the code follow this repo's documented coding standards?) and Spec (does the code match what the originating issue/spec asked for?). Runs both reviews in parallel sub-agents and reports them side by side. Use when the user wants to review a branch, a PR, work-in-progress changes, or asks to \"review since X\".\n- codebase-design: Shared vocabulary for designing deep modules. Use when the user wants to design or improve a module's interface, find deepening opportunities, decide where a seam goes, make code more testable or AI-navigable, or when another skill needs the deep-module vocabulary.\n- diagnose: Disciplined diagnosis loop for hard bugs and performance regressions. Reproduce → minimise → hypothesise → instrument → fix → regression-test. Use when user says \"diagnose this\" / \"debug this\", reports a bug, says something is broken/throwing/failing, or describes a performance regression.\n- diagnosing-bugs: Diagnosis loop for hard bugs and performance regressions. Use when the user says \"diagnose\"/\"debug this\", or reports something broken/throwing/failing/slow.\n- domain-modeling: Build and sharpen a project's domain model. Use when the user wants to pin down domain terminology or a ubiquitous language, record an architectural decision, or when another skill needs to maintain the domain model.\n- find-skills: Helps users discover and install agent skills when they ask questions like \"how do I do X\", \"find a skill for X\", \"is there a skill that can...\", or express interest in extending capabilities. This skill should be used when the user is looking for functionality that might exist as an installable skill.\n- git-guardrails-claude-code: Set up Claude Code hooks to block dangerous git commands (push, reset --hard, clean, branch -D, etc.) before they execute. Use when user wants to prevent destructive git operations, add git safety hooks, or block git push/reset in Claude Code.\n- grilling: Grill the user relentlessly about a plan, decision, or idea. Use when the user wants to stress-test their thinking, or uses any 'grill' trigger phrases.\n- gstack: Fast headless browser for QA testing and site dogfooding. Navigate pages, interact with\nelements, verify state, diff before/after, take annotated screenshots, test responsive\nlayouts, forms, uploads, dialogs, and capture bug evidence. Use when asked to open or\ntest a site, verify a deployment, dogfood a user flow, or file a bug with screenshots. (gstack)\n\n- migrate-to-shoehorn: Migrate test files from `as` type assertions to @total-typescript/shoehorn. Use when user mentions shoehorn, wants to replace `as` in tests, or needs partial test data.\n- mux-send: Send keystrokes to a psmux terminal session. Use when the user says \"send to\",\n\"type in\", \"send keys to\", \"mux send\", or wants to send input to another terminal\nsession by name.\n\n- open-gstack-browser: Launch GStack Browser — AI-controlled Chromium with the sidebar extension baked in.\nOpens a visible browser window where you can watch every action in real time.\nThe sidebar shows a live activity feed and chat. Anti-bot stealth built in.\nUse when asked to \"open gstack browser\", \"launch browser\", \"connect chrome\",\n\"open chrome\", \"real browser\", \"launch chrome\", \"side panel\", or \"control my browser\".\n\n- propose-issue: Challenge an issue statement before fixing it. Probes scope, root cause, and assumptions\nthrough interactive questioning grounded in the codebase. Use when the user says \"propose issue\",\n\"challenge this bug\", \"question this fix\", \"is this the right problem\", or wants to pressure-test\na problem statement before jumping to implementation.\n\n- prototype: Build a throwaway prototype to answer a design question. Use when the user wants to sanity-check whether a state model or logic feels right, or explore what a UI should look like.\n- research: Investigate a question against high-trust primary sources and capture the findings as a Markdown file in the repo. Use when the user wants a topic researched, docs or API facts gathered, or reading legwork delegated to a background agent.\n- resolving-merge-conflicts: Use when you need to resolve an in-progress git merge/rebase conflict.\n- scaffold-exercises: Create exercise directory structures with sections, problems, solutions, and explainers that pass linting. Use when user wants to scaffold exercises, create exercise stubs, or set up a new course section.\n- screenshot-capture: Capture, process, and analyze screenshots of any running application on Windows 10/11. Use this skill whenever the user asks to take a screenshot, capture what's on screen, grab an app window, extract text from a visible app, annotate a screenshot, read what's on screen, OCR a running application, document a UI state, or visually inspect any running program. Also trigger when the user says things like \"show me what [app] looks like\", \"grab that window\", \"read the text in [app]\", \"mark up a screenshot\", \"capture my desktop\", or any task involving visual inspection of running software. This skill handles full-screen captures, individual window captures by app name, OCR text extraction, and image annotation (highlights, arrows, text labels). Even if the user doesn't say \"screenshot\" explicitly, use this skill if they want to see, capture, read, or annotate anything currently displayed on their screen.\n\n- setup-pre-commit: Set up Husky pre-commit hooks with lint-staged (Prettier), type checking, and tests in the current repo. Use when user wants to add pre-commit hooks, set up Husky, configure lint-staged, or add commit-time formatting/typechecking/testing.\n- spt-endpoint: Hosts an adapterless SPT endpoint perch inside Oh My Pi and reliably surfaces incoming messages through harness background-job completion. Use when asked to run, host, start, or listen as an SPT endpoint in OMP, especially when an adapter-backed endpoint or a never-ending `spt ready` listener is unsuitable.\n- tdd: Test-driven development. Use when the user wants to build features or fix bugs test-first, mentions \"red-green-refactor\", or wants integration tests.\n- to-prd: Turn the current conversation context into a PRD and publish it to the project issue tracker. Use when user wants to create a PRD from the current context.\n- trace: Find old conversations in Codex session history. Use when the user says\n\"find the conversation where\", \"when did we discuss\", \"trace back to\", or wants\nto locate a past session by topic, command, commit, or keyword.\n\n- traceable-reqs: Install or update the traceable-reqs CLI and apply it correctly in host repos: author manifests, place tags near evidence, audit scan roots, interpret findings, and answer coverage questions through the CLI.\n- win-marketplace-add: Add a plugin marketplace on Windows, bypassing the EBUSY rename bug. Use when the user wants to add a marketplace, runs \"/plugin marketplace add\" and gets an EBUSY error, or needs to register a marketplace that was manually cloned. Trigger on: \"add marketplace\", \"plugin marketplace add\", \"EBUSY marketplace\", \"register marketplace\", \"marketplace not found\".\n\n- wizard: Generate an interactive bash wizard that walks a human through steps only they can perform. Use when provisioning infrastructure, setting up credentials or CI secrets, walking an unfamiliar third-party dashboard, or running a one-off migration or cutover. Don't invoke this for steps the agent can perform itself.\n- write-a-skill: Create new agent skills with proper structure, progressive disclosure, and bundled resources. Use when user wants to create, write, or build a new skill.\n- writing-for-agents: Writing documents for agents. Use when creating or editing skills, or modifying AGENTS.md or CLAUDE.md.\n</skills>\n# Internal URLs\nMost FS/bash tools auto-resolve these to FS paths.\n- `skill://<name>`: instructions; `/<path>`: its file\n- `rule://<name>`: details\n- `agent://<id>`: output artifact; `/<child>`: nested-subagent output; otherwise `/<path>`: JSON field\n- `history://<id>`: read-only agent transcript (live|parked|released); bare `history://`: all agents. Registered process-wide agents and persisted subagents discoverable from artifact trees; unregistered top-level sessions are not discovered solely from persisted session files.\n- `artifact://<id>`: content\n- `local://<name>.md`: plan artifacts/shared subagent content\n- `mcp://<uri>`: MCP resource\n- `issue://<N>` / `issue://<owner>/<repo>/<N>`: GitHub issue; bare: recent; `?state=open|closed|all&limit=&author=&label=`.\n- `pr://<N>` / `pr://<owner>/<repo>/<N>`: same cache; bare: recent; `?comments=0` `?state=open|closed|merged|all&limit=&author=&label=`.\n- `omp://`: harness docs; AVOID unless user asks about harness.\n\n# Tool Inventory\n- Read: `read`\n- Grep: `grep`\n- Glob: `glob`\n- Bash: `bash`\n- Web Search: `web_search`\n- Submit Result: `yield`\n- Task: `task`\n- Hub: `hub`\n- Write: `write`\n# xd:// Tool Devices\nWrite JSON args as `content` to `xd://<tool>` via `write`. Invalid args return schema in error → fix/retry.\n## Additional devices (docs on demand)\n- xd://spt_checkpoint — After the commune skill has saved this live endpoint's continuity drop, compact native OMP context and wake the same endpoint.\n\nRead xd://<tool> for full docs + JSON schema before first use.\n§ Tool Policy\n# General\nUse tools when they improve correctness, completeness, or grounding.\n- SHOULD resolve prerequisites first; NEVER accept first plausible answer when another call reduces uncertainty; retry empty/partial/suspiciously narrow lookup differently.\n- SHOULD parallelize independent calls.\n- User says `parallel` or `parallelize` → MUST use `task` subagents; parallel tool calls insufficient.\n\n# Tool I/O\n- Prefer relative `path`-like fields.\n- Most tools take `i`: capitalized 2–6-word present-participle intent (e.g. \"Reading model role settings\").\n# Specialized Tools\nMUST use specialized tool over shell equivalent:\n- File/directory reads → `read`; directory path lists entries.\n- Regex search/target location → `grep`, not shell `grep`, `rg`, `awk`.\n- Structure mapping/globbing → `glob`, not `ls **/*.ext` or `fd`.\n- `bash`: real binaries/short fact pipelines only; commands shadowing specialized tools blocked.\n- Bash litmus: one external-CLI call/short pipeline returning count, frequency, set difference, checksum. For merely moving, paging, trimming fetchable bytes: tool.\n\n<critical>\n`write xd://report_issue`: automated QA. Any tool output inconsistent with described behavior for parameters → write plain `<tool>: <concise description>` to `xd://report_issue`. False positives fine.\n</critical>\n\n# Exploration\nNEVER open files hoping. AVOID unneeded files/sections.\n- Use `read` offset/limit, not whole-file reads.\n# Delegation\nInline first. Fan out only when 2+ independent slices each cost more than a handful of your own calls, or the read set would flood context; decide after your own first `grep`/`read`, never before it.\n- NEVER open with a scout. Scope with `grep`/`read`/`glob` yourself; a scout is for a genuinely unmapped subsystem after inline scoping stalls.\n- NEVER delegate one slice. One subagent for one job, a slice you already have open, cleanup (comment trims, changelog lines, formatting, sub-30-line edits), or a direct question: do it yourself.\n- NEVER babysit. Spawn → keep working → read the result. Steering a lone agent through `hub` send/wait costs more than the work.\n## Delegation gates\n- **Own decomposition.** Before spawning: map request, independent slices, cross-slice formats/schemas/interfaces. Only user-enumerated 2+ self-contained runnable slices dispatch directly. NEVER outsource top-level plan; generic \"plan\"/\"design\" agent starts blank, knows less, adds round-trip/no parallelism. Slice-local design and requested competing plans/reviews allowed.\n- **Real concurrency.** Fan exactly to genuine decomposition, one `tasks[]` array. NEVER serialize concurrent slices, invent padding, or spawn one then idle.\n- **User intent.** Subagents lack conversation; retain interpretation/taste; each assignment gets all slice requirements.\n- **Cap:** At most 32 subagents concurrently; excess queues. `tasks[]` batch > 32 delays results: stay within cap.\n- **Dependencies only.** A before B only if B strictly needs A; shared prerequisite inline, then fan out. “Parallelize” = parallel execution of independent slices, not agents routing sequential work. Small missing piece: run parallel; B asks A via `hub`!\n\n§ Workflow\n# 1. Scope\n- Read relevant skills first.\n- Multi-file work: plan before files.\n\n# 2. Research Before Editing\n- Read sections, not snippets. MUST reuse existing patterns; second convention beside existing is PROHIBITED.\n\n- Tool failure/file change since read → re-read before acting.\n\n# 3. Decompose\n\n# 4. Implement\n- Fix source; NEVER suppress symptom/special-case input unless asked.\n- Clean cutover: migrate every caller; remove obsolete code/comments/aliases/re-exports/deprecated paths.\n- Prefer existing-file updates over new files. Review as user.\n- NEVER run destructive git commands/delete unrelated code you didn't write; code the cutover obsoletes is in scope.\n\n# 5. Verify\n- NEVER yield non-trivial work without deliverable proof:\n  - **Experiment/investigation** → run; output is proof; no tests.\n  - **UI change** → verify against the actual surface:\n    - **TUI/CLI** → launch the actual program and verify terminal interaction, output, or state.\n    - No suitable runtime capability for the changed surface → verify with a throwaway script or smoke test; explicitly report when visual verification cannot be performed.\n  - **Bug fix** → reproduce, fix, confirm reproduction no longer triggers. SHOULD keep the reproduction as a regression test: fails pre-fix, passes post-fix; impractical → smoke test, report it.\n  - **Permanent feature/API change** → fix existing tests the changed contract breaks; prove new behavior with a throwaway script. New test ONLY for a genuinely uncertain edge case, or on user request.\n- Smoke test: run thing, not test file; launch, exercise changed path, observe result.\n- Tests: permanent load, not proof of work. A test earns its place ONLY where a plausible bug would fail it.\n  - Each MUST defend observable contract/fail on plausible bug.\n  - Test behavior, boundaries, invariants, transitions, precedence, real errors—not plumbing, source text, incidental defaults.\n  - Match conventions; deterministic, isolated, full-suite-safe.\n  - NEVER write a test so the change \"has tests\" → throwaway script.\n  - NEVER assert implementation: wiring, field copies, defaults, forwarding, mock echoes, source text → assert what a consumer observes.\n  - NEVER pad: same-path parameter rows, tautologies, bare not-throw, non-empty/length-grew checks.\n  - Worth keeping: behavior, boundaries, invariants, transitions, precedence, real errors. Match conventions; deterministic, isolated, full-suite-safe.\n  - Existing test failing this bar (pins wording, implementation, incidental behavior) → MUST delete; NEVER re-pin it to the new text. In scope regardless of author.\n\n# 6. Cleanup\nLast phase; REQUIRED after smoke test proves work; NEVER pre-plan/pre-allocate cleanup todos.\n- Permanent feature/bug fix → docs, changelog, scaffold + throwaway-script removal; tests only per Verify.\n- Experiment/one-off investigation → no cleanup tests/docs.\n\n§ Delivery\n<contract>\nInviolable.\n- NEVER yield before complete deliverable; phase boundary/todo flip/sub-step never yields: same turn.\n- NEVER fabricate output; code/tool/test/doc/source claims MUST be grounded.\n- NEVER substitute easier/familiar problem: don't infer extra scope—retries, validation, telemetry, abstraction “while you're at it”—or solve symptom—suppress warning/exception, special-case input—unless asked. Real ask only.\n- NEVER ask for tool/repo/file-provided information; NEVER punt half-solved work.\n- Default clean cutover: migrate every caller; no shims, aliases, deprecated paths.\n</contract>\n\n<completeness>\n- “Done”: specified end-to-end behavior plus every named acceptance criterion; not compiling scaffold, narrowed test, plausible subset.\n- Reduce scope only with explicit user approval in this conversation; NEVER silently shrink.\n- NEVER deliver unfinished work: stubs, placeholders, mocks, no-ops, fake fallbacks, `TODO: implement`, misleading “scaffold”/“MVP”/“v1”/“foundation”/“follow-up”. Unavailable real-implementation info → state missing prerequisite; finish all reachable work.\n</completeness>\n\n<evidence-and-output>\n- Format MUST match ask; prose brief; evidence, verification, blocking details complete.\n- Code/tool/test/doc/source claims MUST be grounded; unobserved claims `[INFERENCE]`.\n- Verification claims exactly match exercised work.\n</evidence-and-output>\n\n<yielding>\nBefore yielding: all affected callsites/tests/docs updated or intentionally unchanged; output/evidence requirements satisfied.\nBefore blocked: ensure info unreachable via tools/context; one failed check ≠ blocked. Finish reachable work; state exactly missing and tried.\n</yielding>\n\n§ Critical\n<critical>\n- NEVER yield while actionable work remains; phase boundary/todo flip/sub-step never stops: same turn.\n- NEVER narrate/consider session limits, token/tool budgets, effort estimates, or possible completion; start unbounded: execute/delegate.\n- NEVER re-audit applied edit or routinely run git subcommands for validation. Tool results are verification.\n</critical>\n\n§ Role\nFind bugs author wants fixed before merge.\n\n<procedure>\n1. Patch: `git diff` | `jj diff --git` | `gh pr diff <number>`\n2. Modified files: read full context.\n3. Each issue: incremental `yield`, `type: [\"findings\"]`.\n4. Verdict fields: incremental `yield`; stop → idle finalization assembles result.\n\nBash read-only: `git diff`, `git log`, `git show`, `jj diff --git`, `gh pr diff`. NEVER edit files or trigger builds.\n</procedure>\n\n<criteria>\nReport only issues meeting ALL:\n- **Provable impact** — specific affected code paths; no speculation.\n- **Actionable** — discrete fix, not vague \"consider improving X\".\n- **Unintentional** — clearly not deliberate design choice.\n- **Introduced in patch** — don't flag pre-existing bugs.\n- **No unstated assumptions** — no assumptions about codebase or author intent.\n- **Proportionate rigor** — fix demands no rigor absent elsewhere in codebase.\n</criteria>\n\n<cross-boundary>\nEvery patch-introduced type, variant, or value crossing a function or module boundary (event, message, command, frame, enum variant, queue item, IPC payload):\n1. Locate consuming-side dispatch point receiving/routing it: switch, router, filter chain, handler registry, or loop body.\n2. Confirm explicit branch or existing catch-all correctly forwards it.\n3. Report defect if silent drop, no-op, or discard; e.g., unmatched `if`/`switch` simply returns without processing.\n\nDispatch point often outside diff. MUST read it before concluding producing side correct. Tracing emitter while skipping consumer routing is most common source of missed integration bugs in reviews.\n</cross-boundary>\n\n<priority>\n|Level|Criteria|Example|\n|---|---|---|\n|P0|Blocks release/operations; universal (no input assumptions)|Data corruption, auth bypass|\n|P1|High; fix next cycle|Race condition under load|\n|P2|Medium; fix eventually|Edge case mishandling|\n|P3|Info; nice to have|Suboptimal but correct|\n</priority>\n\n<findings>\n- **Title**: e.g., `Handle null response from API`\n- **Body**: bug, trigger condition, impact; neutral tone.\n- **Suggestion blocks**: only concrete replacement code; preserve exact whitespace; no commentary.\n</findings>\n\n<example name=\"finding\">\n<title>Validate input length before buffer copy</title>\n<body>When `data.length > BUFFER_SIZE`, `memcpy` writes past buffer boundary. Occurs if API returns oversized payloads, causing heap corruption.</body>\n```suggestion\nif (data.length > BUFFER_SIZE) return -EINVAL;\nmemcpy(buf, data.ptr, data.length);\n```\n</example>\n\n<output>\nFinding: incremental `yield`, `type: [\"findings\"]`; `data`:\n- `title`: imperative, ≤80 chars.\n- `body`: one paragraph.\n- `priority`: 0-3.\n- `confidence`: 0.0-1.0.\n- `file_path`: affected-file path.\n- `line_start`, `line_end`: ≤10-line range; MUST overlap diff.\n\nVerdict fields: incremental `yield`:\n- `type: [\"overall_correctness\"]`: `\"correct\"` (no bugs/blockers) | `\"incorrect\"`.\n- `type: [\"explanation\"]`: plain-text 1-3-sentence verdict summary.\n- `type: [\"confidence\"]`: 0.0-1.0 confidence.\n\nDo not emit separate submit tool call or duplicate `findings` in another payload. After all sections, stop; idle finalization assembles result.\n\nNEVER output JSON or code blocks.\n\nCorrectness ignores non-blocking issues: style, docs, nits.\n</output>\n\n<critical>\nEvery finding MUST be patch-anchored and evidence-backed.\n</critical>\n\n§ Context\n# Goal\nAdvance the remaining v0.69 bugfix queue while PR #216 owns both CI runners. Parent todlando implements #292 and integrates evidence; these tasks are independent.\n# Constraints\nSkip ALL validation: no Cargo, builds, formatters, linters, tests, services, or production restarts. No push. Each lane is under .worktrees/ in this repo. Follow supplied AGENTS rules, traceable requirements, existing conventions. Commit trailers exactly `Co-authored by: todlando`. Read required orientation by relevant sections. No source-text tests. Findings must be evidence-backed, not speculative. Reviewer work is read-only.\n# Contract\nShellRestore exclusively owns .worktrees/287-shellrestore (branch fix/287-persistent-restore at de5a44bc). Churn reviewers only read .worktrees/286-churn, head 19a9097956de0db898eaeb3a7dc4765189ade874; neither edits it. Review base de5a44bc338d679f3c63f6d0f68826eb9e575d73 is verified and diff nonempty: 5 files, 425 insertions/56 deletions, single commit `19a90979 fix(daemon): bound healthy connection stderr and retain livehost polls`. Diff command in that cwd: `git diff de5a44bc...HEAD`. Issue spec is issue://BigscreenVR/spt-bs-releases/286. Parent's measured scope: healthy lifecycle lines dominate log volume; livehost query reconnects every5s; family gate is already once per real bind and is EXONERATED, no binder policy changes. Keep reviewer axes separate.\nStandards smell baseline (heuristics only, repo rules override): Mysterious Name—unclear purpose, rename; Duplicated Code—same logic repeated, share; Feature Envy—method uses another object's data more than its own, relocate; Data Clumps—same fields travel together, consider cohesive type; Primitive Obsession—primitive encoding deserves a domain type; Repeated Switches—duplicated dispatch choices, share one map/interface; Shotgun Surgery—one change scattered, gather; Divergent Change—one module changes for unrelated reasons, separate; Speculative Generality—unused abstraction or hooks, remove; Message Chains—caller walks internals, hide walk; Middle Man—pure forwarding without value, remove; Refused Bequest—subtype rejects parent contract, use composition. Cite smells as judgement calls, not hard violations. Skip tooling-enforced findings.\n§ Coop\nYou are operating on a piece of work assigned to you by the main agent.\n\n# Validation\nProject-wide validation is the main agent's job, run once after all subagents land. NEVER run formatters, linters, or project-wide builds/test suites unless your assignment explicitly instructs it — siblings edit concurrently; mid-flight validation blocks on their half-finished changes and reports phantom failures. Scoped proof of your own change (single test file, targeted repro, smoke run) is fine.\n# Peers\nYou can reach other live agents via the `hub` tool. Your id is `ChurnStandards`. Currently visible peers:\n- `ChurnSpec` — reviewer (sub, running)\n- `ShellRestore` — task (sub, running)\n- `Main` — main (main, running)\n2 parked peer(s) omitted.\n\nUse `hub` messaging only for quick coordination, never long-form content. Address peers by id or use `\"all\"` to broadcast.\n- Discovery: the roster above shows live (running+idle) peers and a parked count, never parked names or task labels. `hub` op:\"list\" refreshes the live view; pass status:\"parked\" to inspect parked history.\n- Coordination: before you edit a file or start work a sibling may already own, message that peer first — overlapping edits collide. Idle peers are not gone: messaging them wakes them.\n- Follow-up: answer a peer's question with a short reply (set `replyTo`); use `await` only when you genuinely cannot proceed without the answer.\n- Parked history: omitted from this roster. `hub` op:\"list\" status:\"parked\" lists ids; `send` to a known parked id revives it. `history://<id>` and `agent://<id>` stay readable.\n\n§ Completion\nNo TODO tracking, no progress updates. Execute; report results with `yield`.\n\nWhile work remains, you MUST continue with another tool call — investigate, edit, run, verify. Save narrative for a terminal `yield` unless you intentionally record an incremental section.\n\nYield protocol:\n- Omit `type` for the normal single terminal structured result in `data`.\n- Use non-empty `type: string[]` for incremental, non-terminal sections; calls accumulate by section.\n- A data-less terminal `type: \"result\"` only finalizes previously submitted incremental sections; it NEVER substitutes for `data`.\n\nThis is your only way to return a final result. For structured results, you NEVER put JSON in plain text or substitute a text summary for `data`.\n\nYour terminal `yield` MUST use exactly this shape — the schema fields go inside `data`, NEVER at the top level and NEVER as a stringified summary:\n```ts\n{\n  data: {\n    overall_correctness: \"correct\" | \"incorrect\";\n    explanation: string;\n    confidence: unknown;\n    findings?: Array<{\n    title: string;\n    body: string;\n    priority: unknown;\n    confidence: unknown;\n    file_path: string;\n    line_start: unknown;\n    line_end: unknown;\n  }>;\n  };\n}\n```\n\nGiving up is a last resort. If truly blocked, you MUST terminal-yield `{ error }` describing what you tried and the exact blocker.\nYou NEVER give up due to uncertainty, missing information obtainable via tools or repo context, or needing a design decision you can derive yourself.\n\nYou MUST keep going until this ticket is closed. This matters.\n\nPROJECT\n\n<workstation>\n- OS: win32 10.0.26200\n- Distro: Windows_NT\n- Kernel: Windows 11 Pro\n- Arch: x64\n- CPU: AMD Ryzen 7 7800X3D 8-Core Processor\n- GPU: AMD Radeon(TM) Graphics\n- Model: openai-codex/gpt-6-astra\n</workstation>\n<critical>\n- Each response MUST advance the task; completion only stopping condition.\n- MUST default to informed action; do not ask for confirmation when tools or repo context can answer.\n- Before yielding, MUST verify significant behavioral changes: run the specific test, command, or scenario covering the change.\n</critical>","task":"Complete assignment thoroughly:\n\n# Target\nRead-only Standards-axis review of de5a44bc...19a90979 in .worktrees/286-churn. Changed files conn.rs, livehost.rs, traceable-reqs.toml, CHANGELOG.md, docs/RCA-FLEET-DAEMON-14444.md. No edits.\n# Change\nInspect the diff and relevant surrounding source. Standards sources: supplied AGENTS.md, docs/KNOWN-HAZARDS.md (connection poison/lifecycle/pid liveness sections), docs/TRACEABILITY.md, docs/DOCS-STRATEGY.md, plus the shared smell baseline. Report concrete violations with file:line and rule; distinguish hard violations from heuristic smells. Focus concurrency, lifecycle attribution retained under aggregation, noisy healthy events versus exceptional evidence, unnecessary allocation/abstraction, and tests that only pin implementation rather than behavior. Do not substitute a spec review or demand exonerated binder changes.\n# Acceptance\nUnder400 words, actionable findings with severity and source evidence; if clean say so. No formatters, linters, tests, builds or runtime execution.","tools":["read","grep","glob","bash","web_search","yield","task","hub","spt_checkpoint"],"agent":"reviewer","modelRole":"slow","resolvedModel":"openai-codex/gpt-6-astra:xhigh","readOnly":false,"spawns":"scout","outputSchema":{"properties":{"overall_correctness":{"metadata":{"description":"Whether change correct (no bugs/blockers)"},"enum":["correct","incorrect"]},"explanation":{"metadata":{"description":"Plain-text verdict summary, 1-3 sentences"},"type":"string"},"confidence":{"metadata":{"description":"Verdict confidence (0.0-1.0)"},"type":"number"}},"optionalProperties":{"findings":{"metadata":{"description":"Populate via incremental yield sections under type: [\"findings\"]; don't repeat it in a final payload."},"elements":{"properties":{"title":{"metadata":{"description":"Imperative, ≤80 chars"},"type":"string"},"body":{"metadata":{"description":"One paragraph: bug, trigger, impact"},"type":"string"},"priority":{"metadata":{"description":"P0-P3: 0 blocks release, 1 fix next cycle, 2 fix eventually, 3 nice to have"},"type":"number"},"confidence":{"metadata":{"description":"Confidence it's real bug (0.0-1.0)"},"type":"number"},"file_path":{"metadata":{"description":"Path to affected file"},"type":"string"},"line_start":{"metadata":{"description":"First line (1-indexed)"},"type":"number"},"line_end":{"metadata":{"description":"Last line (1-indexed, ≤10 lines)"},"type":"number"}}}}}},"outputSchemaMode":"permissive"}
{"type":"message","id":"fbd11b38","parentId":"801a7828","timestamp":"2026-09-10T00:18:00.951Z","message":{"role":"user","content":[{"type":"text","text":"Complete assignment thoroughly:\n\n# Target\nRead-only Standards-axis review of de5a44bc...19a90979 in .worktrees/286-churn. Changed files conn.rs, livehost.rs, traceable-reqs.toml, CHANGELOG.md, docs/RCA-FLEET-DAEMON-14444.md. No edits.\n# Change\nInspect the diff and relevant surrounding source. Standards sources: supplied AGENTS.md, docs/KNOWN-HAZARDS.md (connection poison/lifecycle/pid liveness sections), docs/TRACEABILITY.md, docs/DOCS-STRATEGY.md, plus the shared smell baseline. Report concrete violations with file:line and rule; distinguish hard violations from heuristic smells. Focus concurrency, lifecycle attribution retained under aggregation, noisy healthy events versus exceptional evidence, unnecessary allocation/abstraction, and tests that only pin implementation rather than behavior. Do not substitute a spec review or demand exonerated binder changes.\n# Acceptance\nUnder400 words, actionable findings with severity and source evidence; if clean say so. No formatters, linters, tests, builds or runtime execution."}],"attribution":"agent","timestamp":1788999480888}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_Q2LWDOrJ4Xv5BgluW6BDe6QO|fc_0dde72effda0ad21016aa1f73b1ee087d28c720197272ed003","toolName":"read","startedAt":"2026-09-10T00:18:04.193Z","args":{"path":"skill://code-review"},"intent":"Reading review skill"},"id":"f61d7366","parentId":"fbd11b38","timestamp":"2026-09-10T00:18:04.193Z"}
{"type":"message","id":"4f456165","parentId":"f61d7366","timestamp":"2026-09-10T00:18:04.194Z","message":{"role":"assistant","content":[{"type":"toolCall","id":"call_Q2LWDOrJ4Xv5BgluW6BDe6QO|fc_0dde72effda0ad21016aa1f73b1ee087d28c720197272ed003","name":"read","arguments":{"i":"Reading review skill","path":"skill://code-review"},"intent":"Reading review skill"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":13238,"output":26,"cacheRead":0,"cacheWrite":0,"totalTokens":13264,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1788999480950,"responseId":"resp_0dde72effda0ad21016aa1f739ab9487d2b3dcd41fb8d6d7fe","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"fc_0dde72effda0ad21016aa1f73b1ee087d28c720197272ed003","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading review skill\",\"path\":\"skill://code-review\"}","call_id":"call_Q2LWDOrJ4Xv5BgluW6BDe6QO","internal_chat_message_metadata_passthrough":{"create_time":1788999482.171666,"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"read","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}}]},"duration":3242.6070999999065,"ttft":2209.7859999998473,"completedAt":1788999484193,"contextSnapshot":{"promptTokens":13238,"nonMessageTokens":13648,"compactionEpoch":0}}}
{"type":"credential_pin","id":"bbc6e03c","parentId":"4f456165","timestamp":"2026-09-10T00:18:04.194Z","provider":"openai-codex","hash":"1a3c10d7bc41d381cab1511db604bc2637dfcfc63796a890d8b5ee152859e1cb"}
{"type":"message","id":"e63b951c","parentId":"bbc6e03c","timestamp":"2026-09-10T00:18:04.195Z","message":{"role":"toolResult","toolCallId":"call_Q2LWDOrJ4Xv5BgluW6BDe6QO|fc_0dde72effda0ad21016aa1f73b1ee087d28c720197272ed003","toolName":"read","content":[{"type":"text","text":"---\nname: code-review\ndescription: Review the changes since a fixed point (commit, branch, tag, or merge-base) along two axes — Standards (does the code follow this repo's documented coding standards?) and Spec (does the code match what the originating issue/spec asked for?). Runs both reviews in parallel sub-agents and reports them side by side. Use when the user wants to review a branch, a PR, work-in-progress changes, or asks to \"review since X\".\n---\n\nTwo-axis review of the diff between `HEAD` and a fixed point the user supplies:\n\n- **Standards** — does the code conform to this repo's documented coding standards?\n- **Spec** — does the code faithfully implement the originating issue / spec?\n\nBoth axes run as **parallel sub-agents** so they don't pollute each other's context, then this skill aggregates their findings.\n\nThe issue tracker should have been provided to you — run `/setup-matt-pocock-skills` if `docs/agents/issue-tracker.md` is missing.\n\n## Process\n\n### 1. Pin the fixed point\n\nWhatever the user said is the fixed point — a commit SHA, branch name, tag, `main`, `HEAD~5`, etc. If they didn't specify one, ask for it.\n\nCapture the diff command once: `git diff <fixed-point>...HEAD` (three-dot, so the comparison is against the merge-base). Also note the list of commits via `git log <fixed-point>..HEAD --oneline`.\n\nBefore going further, confirm the fixed point resolves (`git rev-parse <fixed-point>`) and the diff is non-empty. A bad ref or empty diff should fail here — not inside two parallel sub-agents.\n\n### 2. Identify the spec source\n\nLook for the originating spec, in this order:\n\n1. Issue references in the commit messages (`#123`, `Closes #45`, GitLab `!67`, etc.) — fetch via the workflow in `docs/agents/issue-tracker.md`.\n2. A path the user passed as an argument.\n3. A spec file under `docs/`, `specs/`, or `.scratch/` matching the branch name or feature.\n4. If nothing is found, ask the user where the spec is. If they say there isn't one, the **Spec** sub-agent will skip and report \"no spec available\".\n\n### 3. Identify the standards sources\n\nAnything in the repo that documents how code should be written, such as `CODING_STANDARDS.md` or `CONTRIBUTING.md`.\n\nOn top of whatever the repo documents, the Standards axis always carries the **smell baseline** below — a fixed set of Fowler code smells (_Refactoring_, ch.3) that applies even when a repo documents nothing. Two rules bind it:\n\n- **The repo overrides.** A documented repo standard always wins; where it endorses something the baseline would flag, suppress the smell.\n- **Always a judgement call.** Each smell is a labelled heuristic (\"possible Feature Envy\"), never a hard violation — and, like any standard here, skip anything tooling already enforces.\n\nEach smell reads *what it is* → *how to fix*; match it against the diff:\n\n- **Mysterious Name** — a function, variable, or type whose name doesn't reveal what it does or holds. → rename it; if no honest name comes, the design's murky.\n- **Duplicated Code** — the same logic shape appears in more than one hunk or file in the change. → extract the shared shape, call it from both.\n- **Feature Envy** — a method that reaches into another object's data more than its own. → move the method onto the data it envies.\n- **Data Clumps** — the same few fields or params keep travelling together (a type wanting to be born). → bundle them into one type, pass that.\n- **Primitive Obsession** — a primitive or string standing in for a domain concept that deserves its own type. → give the concept its own small type.\n- **Repeated Switches** — the same `switch`/`if`-cascade on the same type recurs across the change. → replace with polymorphism, or one map both sites share.\n- **Shotgun Surgery** — one logical change forces scattered edits across many files in the diff. → gather what changes together into one module.\n- **Divergent Change** — one file or module is edited for several unrelated reasons. → split so each module changes for one reason.\n- **Speculative Generality** — abstraction, parameters, or hooks added for needs the spec doesn't have. → delete it; inline back until a real need shows.\n- **Message Chains** — long `a.b().c().d()` navigation the caller shouldn't depend on. → hide the walk behind one method on the first object.\n- **Middle Man** — a class or function that mostly just delegates onward. → cut it, call the real target direct.\n- **Refused Bequest** — a subclass or implementer that ignores or overrides most of what it inherits. → drop the inheritance, use composition.\n\n### 4. Spawn both sub-agents in parallel\n\nSend a single message with two `Agent` tool calls. Use the `general-purpose` subagent for both.\n\n**Standards sub-agent prompt** — include:\n\n- The full diff command and commit list.\n- The list of standards-source files you found in step 3, **plus the smell baseline from step 3** pasted in full — the sub-agent has no other access to it.\n- The brief: \"Report — per file/hunk where relevant — (a) every place the diff violates a documented standard: cite the standard (file + the rule); and (b) any baseline smell you spot: name it and quote the hunk. Distinguish hard violations from judgement calls — documented-standard breaches can be hard, but baseline smells are always judgement calls, and a documented repo standard overrides the baseline. Skip anything tooling enforces. Under 400 words.\"\n\n**Spec sub-agent prompt** — include:\n\n- The diff command and commit list.\n- The path or fetched contents of the spec.\n- The brief: \"Report: (a) requirements the spec asked for that are missing or partial; (b) behaviour in the diff that wasn't asked for (scope creep); (c) requirements that look implemented but where the implementation looks wrong. Quote the spec line for each finding. Under 400 words.\"\n\nIf the spec is missing, skip the Spec sub-agent and note this in the final report.\n\n### 5. Aggregate\n\nPresent the two reports under `## Standards` and `## Spec` headings, verbatim or lightly cleaned. Do **not** merge or rerank findings — the two axes are deliberately separate (see _Why two axes_).\n\nEnd with a one-line summary: total findings per axis, and the worst issue _within each axis_ (if any). Don't pick a single winner across axes — that's the reranking the separation exists to prevent.\n\n## Why two axes\n\nA change can pass one axis and fail the other:\n\n- Code that follows every standard but implements the wrong thing → **Standards pass, Spec fail.**\n- Code that does exactly what the issue asked but breaks the project's conventions → **Spec pass, Standards fail.**\n\nReporting them separately stops one axis from masking the other."}],"details":{"resolvedPath":"C:\\Users\\decid\\.agents\\skills\\code-review\\SKILL.md","contentType":"text/markdown","totalLines":89,"displayContent":{"text":"---\r\nname: code-review\r\ndescription: Review the changes since a fixed point (commit, branch, tag, or merge-base) along two axes — Standards (does the code follow this repo's documented coding standards?) and Spec (does the code match what the originating issue/spec asked for?). Runs both reviews in parallel sub-agents and reports them side by side. Use when the user wants to review a branch, a PR, work-in-progress changes, or asks to \"review since X\".\r\n---\r\n\r\nTwo-axis review of the diff between `HEAD` and a fixed point the user supplies:\r\n\r\n- **Standards** — does the code conform to this repo's documented coding standards?\r\n- **Spec** — does the code faithfully implement the originating issue / spec?\r\n\r\nBoth axes run as **parallel sub-agents** so they don't pollute each other's context, then this skill aggregates their findings.\r\n\r\nThe issue tracker should have been provided to you — run `/setup-matt-pocock-skills` if `docs/agents/issue-tracker.md` is missing.\r\n\r\n## Process\r\n\r\n### 1. Pin the fixed point\r\n\r\nWhatever the user said is the fixed point — a commit SHA, branch name, tag, `main`, `HEAD~5`, etc. If they didn't specify one, ask for it.\r\n\r\nCapture the diff command once: `git diff <fixed-point>...HEAD` (three-dot, so the comparison is against the merge-base). Also note the list of commits via `git log <fixed-point>..HEAD --oneline`.\r\n\r\nBefore going further, confirm the fixed point resolves (`git rev-parse <fixed-point>`) and the diff is non-empty. A bad ref or empty diff should fail here — not inside two parallel sub-agents.\r\n\r\n### 2. Identify the spec source\r\n\r\nLook for the originating spec, in this order:\r\n\r\n1. Issue references in the commit messages (`#123`, `Closes #45`, GitLab `!67`, etc.) — fetch via the workflow in `docs/agents/issue-tracker.md`.\r\n2. A path the user passed as an argument.\r\n3. A spec file under `docs/`, `specs/`, or `.scratch/` matching the branch name or feature.\r\n4. If nothing is found, ask the user where the spec is. If they say there isn't one, the **Spec** sub-agent will skip and report \"no spec available\".\r\n\r\n### 3. Identify the standards sources\r\n\r\nAnything in the repo that documents how code should be written, such as `CODING_STANDARDS.md` or `CONTRIBUTING.md`.\r\n\r\nOn top of whatever the repo documents, the Standards axis always carries the **smell baseline** below — a fixed set of Fowler code smells (_Refactoring_, ch.3) that applies even when a repo documents nothing. Two rules bind it:\r\n\r\n- **The repo overrides.** A documented repo standard always wins; where it endorses something the baseline would flag, suppress the smell.\r\n- **Always a judgement call.** Each smell is a labelled heuristic (\"possible Feature Envy\"), never a hard violation — and, like any standard here, skip anything tooling already enforces.\r\n\r\nEach smell reads *what it is* → *how to fix*; match it against the diff:\r\n\r\n- **Mysterious Name** — a function, variable, or type whose name doesn't reveal what it does or holds. → rename it; if no honest name comes, the design's murky.\r\n- **Duplicated Code** — the same logic shape appears in more than one hunk or file in the change. → extract the shared shape, call it from both.\r\n- **Feature Envy** — a method that reaches into another object's data more than its own. → move the method onto the data it envies.\r\n- **Data Clumps** — the same few fields or params keep travelling together (a type wanting to be born). → bundle them into one type, pass that.\r\n- **Primitive Obsession** — a primitive or string standing in for a domain concept that deserves its own type. → give the concept its own small type.\r\n- **Repeated Switches** — the same `switch`/`if`-cascade on the same type recurs across the change. → replace with polymorphism, or one map both sites share.\r\n- **Shotgun Surgery** — one logical change forces scattered edits across many files in the diff. → gather what changes together into one module.\r\n- **Divergent Change** — one file or module is edited for several unrelated reasons. → split so each module changes for one reason.\r\n- **Speculative Generality** — abstraction, parameters, or hooks added for needs the spec doesn't have. → delete it; inline back until a real need shows.\r\n- **Message Chains** — long `a.b().c().d()` navigation the caller shouldn't depend on. → hide the walk behind one method on the first object.\r\n- **Middle Man** — a class or function that mostly just delegates onward. → cut it, call the real target direct.\r\n- **Refused Bequest** — a subclass or implementer that ignores or overrides most of what it inherits. → drop the inheritance, use composition.\r\n\r\n### 4. Spawn both sub-agents in parallel\r\n\r\nSend a single message with two `Agent` tool calls. Use the `general-purpose` subagent for both.\r\n\r\n**Standards sub-agent prompt** — include:\r\n\r\n- The full diff command and commit list.\r\n- The list of standards-source files you found in step 3, **plus the smell baseline from step 3** pasted in full — the sub-agent has no other access to it.\r\n- The brief: \"Report — per file/hunk where relevant — (a) every place the diff violates a documented standard: cite the standard (file + the rule); and (b) any baseline smell you spot: name it and quote the hunk. Distinguish hard violations from judgement calls — documented-standard breaches can be hard, but baseline smells are always judgement calls, and a documented repo standard overrides the baseline. Skip anything tooling enforces. Under 400 words.\"\r\n\r\n**Spec sub-agent prompt** — include:\r\n\r\n- The diff command and commit list.\r\n- The path or fetched contents of the spec.\r\n- The brief: \"Report: (a) requirements the spec asked for that are missing or partial; (b) behaviour in the diff that wasn't asked for (scope creep); (c) requirements that look implemented but where the implementation looks wrong. Quote the spec line for each finding. Under 400 words.\"\r\n\r\nIf the spec is missing, skip the Spec sub-agent and note this in the final report.\r\n\r\n### 5. Aggregate\r\n\r\nPresent the two reports under `## Standards` and `## Spec` headings, verbatim or lightly cleaned. Do **not** merge or rerank findings — the two axes are deliberately separate (see _Why two axes_).\r\n\r\nEnd with a one-line summary: total findings per axis, and the worst issue _within each axis_ (if any). Don't pick a single winner across axes — that's the reranking the separation exists to prevent.\r\n\r\n## Why two axes\r\n\r\nA change can pass one axis and fail the other:\r\n\r\n- Code that follows every standard but implements the wrong thing → **Standards pass, Spec fail.**\r\n- Code that does exactly what the issue asked but breaks the project's conventions → **Spec pass, Standards fail.**\r\n\r\nReporting them separately stops one axis from masking the other.\r","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89]},"meta":{"source":{"type":"internal","value":"skill://code-review"}}},"isError":false,"timestamp":1788999484194}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_laaz7ZoODf1WB2QxJI0MJVAW|fc_0dde72effda0ad21016aa1f746a5a887d2a398b94370bff234","toolName":"bash","startedAt":"2026-09-10T00:18:14.781Z","args":{"command":"git diff de5a44bc338d679f3c63f6d0f68826eb9e575d73...19a9097956de0db898eaeb3a7dc4765189ade874"},"intent":"Reading churn patch"},"id":"72aa65dc","parentId":"e63b951c","timestamp":"2026-09-10T00:18:14.781Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_4RwL2YnZyT4FlekaOudPGGEj|fc_0dde72effda0ad21016aa1f746a5c087d2936d67a0e7d2e2ed","toolName":"glob","startedAt":"2026-09-10T00:18:14.782Z","args":{"path":".worktrees/286-churn/**/AGENTS.md;.worktrees/286-churn/**/conn.rs;.worktrees/286-churn/**/livehost.rs;AGENTS.md;.worktrees/AGENTS.md"},"intent":"Locating review standards"},"id":"10ddbc36","parentId":"72aa65dc","timestamp":"2026-09-10T00:18:14.782Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_oPXXqMg4wN89WrblFZiU7qSh|fc_0dde72effda0ad21016aa1f746a5d087d2b7fb6f323a04cbd6","toolName":"read","startedAt":"2026-09-10T00:18:14.782Z","args":{"path":".worktrees/286-churn/docs/KNOWN-HAZARDS.md"},"intent":"Reading hazard index"},"id":"53cb2515","parentId":"10ddbc36","timestamp":"2026-09-10T00:18:14.782Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_luDKV6IiMNN9bGSQ59m1fYoL|fc_0dde72effda0ad21016aa1f746a5dc87d2bf625fb463cba393","toolName":"read","startedAt":"2026-09-10T00:18:14.782Z","args":{"path":".worktrees/286-churn/docs/TRACEABILITY.md"},"intent":"Reading traceability rules"},"id":"43615d49","parentId":"53cb2515","timestamp":"2026-09-10T00:18:14.782Z"}
{"type":"message","id":"d7f39ab6","parentId":"43615d49","timestamp":"2026-09-10T00:18:14.782Z","message":{"role":"assistant","content":[{"type":"toolCall","id":"call_laaz7ZoODf1WB2QxJI0MJVAW|fc_0dde72effda0ad21016aa1f746a5a887d2a398b94370bff234","name":"bash","arguments":{"i":"Reading churn patch","command":"git diff de5a44bc338d679f3c63f6d0f68826eb9e575d73...19a9097956de0db898eaeb3a7dc4765189ade874","cwd":"C:/Users/decid/Documents/projects/spt-core/.worktrees/286-churn"},"intent":"Reading churn patch"},{"type":"toolCall","id":"call_4RwL2YnZyT4FlekaOudPGGEj|fc_0dde72effda0ad21016aa1f746a5c087d2936d67a0e7d2e2ed","name":"glob","arguments":{"i":"Locating review standards","path":".worktrees/286-churn/**/AGENTS.md;.worktrees/286-churn/**/conn.rs;.worktrees/286-churn/**/livehost.rs;AGENTS.md;.worktrees/AGENTS.md","hidden":true,"gitignore":false,"limit":100},"intent":"Locating review standards"},{"type":"toolCall","id":"call_oPXXqMg4wN89WrblFZiU7qSh|fc_0dde72effda0ad21016aa1f746a5d087d2b7fb6f323a04cbd6","name":"read","arguments":{"i":"Reading hazard index","path":".worktrees/286-churn/docs/KNOWN-HAZARDS.md"},"intent":"Reading hazard index"},{"type":"toolCall","id":"call_luDKV6IiMNN9bGSQ59m1fYoL|fc_0dde72effda0ad21016aa1f746a5dc87d2bf625fb463cba393","name":"read","arguments":{"i":"Reading traceability rules","path":".worktrees/286-churn/docs/TRACEABILITY.md"},"intent":"Reading traceability rules"},{"type":"toolCall","id":"call_mOf9RXfq8KN14nxPXeceoPsF|fc_0dde72effda0ad21016aa1f746a5e087d2901182999239a7c6","name":"read","arguments":{"i":"Reading documentation rules","path":".worktrees/286-churn/docs/DOCS-STRATEGY.md"},"intent":"Reading documentation rules"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":1747,"output":284,"cacheRead":13056,"cacheWrite":0,"totalTokens":15087,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1788999484246,"responseId":"resp_0dde72effda0ad21016aa1f73c81a087d2ad1883b3eb53e23b","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"fc_0dde72effda0ad21016aa1f746a5a887d2a398b94370bff234","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading churn patch\",\"command\":\"git diff de5a44bc338d679f3c63f6d0f68826eb9e575d73...19a9097956de0db898eaeb3a7dc4765189ade874\",\"cwd\":\"C:/Users/decid/Documents/projects/spt-core/.worktrees/286-churn\"}","call_id":"call_laaz7ZoODf1WB2QxJI0MJVAW","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"bash","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f746a5c087d2936d67a0e7d2e2ed","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating review standards\",\"path\":\".worktrees/286-churn/**/AGENTS.md;.worktrees/286-churn/**/conn.rs;.worktrees/286-churn/**/livehost.rs;AGENTS.md;.worktrees/AGENTS.md\",\"hidden\":true,\"gitignore\":false,\"limit\":100}","call_id":"call_4RwL2YnZyT4FlekaOudPGGEj","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"glob","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f746a5d087d2b7fb6f323a04cbd6","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading hazard index\",\"path\":\".worktrees/286-churn/docs/KNOWN-HAZARDS.md\"}","call_id":"call_oPXXqMg4wN89WrblFZiU7qSh","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"read","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f746a5dc87d2bf625fb463cba393","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading traceability rules\",\"path\":\".worktrees/286-churn/docs/TRACEABILITY.md\"}","call_id":"call_luDKV6IiMNN9bGSQ59m1fYoL","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"read","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f746a5e087d2901182999239a7c6","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading documentation rules\",\"path\":\".worktrees/286-churn/docs/DOCS-STRATEGY.md\"}","call_id":"call_mOf9RXfq8KN14nxPXeceoPsF","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"read","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}}]},"duration":10522.321999999927,"ttft":10433.594999999972,"completedAt":1788999494780,"contextSnapshot":{"promptTokens":14803,"nonMessageTokens":13648,"compactionEpoch":0}}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_mOf9RXfq8KN14nxPXeceoPsF|fc_0dde72effda0ad21016aa1f746a5e087d2901182999239a7c6","toolName":"read","startedAt":"2026-09-10T00:18:14.782Z","args":{"path":".worktrees/286-churn/docs/DOCS-STRATEGY.md"},"intent":"Reading documentation rules"},"id":"7988fd89","parentId":"d7f39ab6","timestamp":"2026-09-10T00:18:14.782Z"}
{"type":"message","id":"f6a80cbb","parentId":"7988fd89","timestamp":"2026-09-10T00:18:14.790Z","message":{"role":"toolResult","toolCallId":"call_luDKV6IiMNN9bGSQ59m1fYoL|fc_0dde72effda0ad21016aa1f746a5dc87d2bf625fb463cba393","toolName":"read","content":[{"type":"text","text":"# Traceability development contract\n\n> How spt-core uses [`traceable-reqs`](https://github.com/BigscreenVR/traceable-reqs) to keep every requirement traced from doc → impl → test. The manifest is `traceable-reqs.toml` (seeded from PRD `R-*` + KNOWN-HAZARDS invariants). This contract makes the trace pay off instead of rotting.\n\n## Why\n\nThe whole point of the rebuild is to *not* re-introduce the sister project's bugs and to ship the PRD's scope verifiably. A requirement that isn't traced to real evidence (a doc, an implementation, a test) is a requirement you're *hoping* is met. This contract converts that hope into a build that fails when evidence is missing.\n\n## The contract (rules)\n\n1. **The manifest is the authoritative requirement registry.** Every PRD `R-*` requirement and every KNOWN-HAZARDS invariant exists in `traceable-reqs.toml` as a `REQ-*` id. No work without a REQ; no REQ without intent to satisfy. The PRD holds the prose; the manifest holds the id + `required_stages`.\n2. **Tag in the same change as the evidence.** When you write the function / test / doc-section that satisfies a stage, add its `[<stage>-><REQ-ID>]` tag in *that same commit*. Never \"tag later\" — this single rule is what stops the trace from drifting.\n3. **Evidence-proximate tags, not file-tops.** A tag sits on or immediately above the real function, test, doc section, config entry, or workflow step that proves the stage. One tag = one piece of real evidence. Tags at file tops to satisfy coverage are noise and usually wrong.\n4. **KNOWN-HAZARDS are first-class requirements.** Each invariant is a `REQ-HAZARD-*` requiring `unit` (and `int` where cross-process/cross-node). A hazard cannot be \"covered\" without a test tagging it — the anti-regression promise becomes mechanical (PRD success criterion #6).\n5. **Activation, not premature failure.** Every requirement starts `required_stages = []` (inactive) so `check` stays green pre-code. A **milestone (M0–M5) activates** the requirements it delivers by setting their real `required_stages`. Deferred items (`docs/DEFERRED.md`) stay `[]` until promoted.\n6. **Scan roots stay honest.** `[scan].roots` includes every evidence location (`src`, `tests`, `docs`, and later `.github/workflows`, `scripts`). Audit roots whenever a new evidence dir appears — a missing root makes evidence *silently* vanish from the trace.\n7. **Scope at mint.** Every new `REQ-*` joins at least one `[[groups]]` subject scope in the same commit that mints it. The groups partition is what the confusion audit (Gate 4, 2026-07-30) turned into manifest structure: the `req-confusion-audit` skill scores confusability only *within* groups, so an unscoped requirement's confusable pairs go unexamined — roughly half of them, by the audit's measurement. `span-*` groups are the declared cross-scope cover for measured-hot pairs; add one when a new requirement is known to sit confusably against another scope. Do not derive scopes from id prefixes — subject matter, not spelling.\n\n### Stages\n\n`doc` (prose/design/rustdoc) · `impl` (production code) · `unit` (unit test) · `int` (integration / e2e / cross-node). Default activation policy: `[\"doc\",\"impl\",\"unit\"]`; networking/lifecycle/cross-node reqs add `int`.\n\n### Tag examples\n\nReal tags are `[<stage>->REQ-ID]` with a bare stage word. The examples below\nangle-bracket the stage (`[<impl>->...]`) only so the scanner treats them as\nillustrations, not live evidence — write a bare stage word in real code.\n\n```rust\n// [<impl>->REQ-PAIR-1]\nfn spake2_pair(...) { ... }\n```\n```rust\n// [<unit>->REQ-HAZARD-ENVELOPE-DECODE-ORDER]\n#[test] fn decodes_amp_last() { ... }\n```\n```markdown\n<!-- [<doc>->REQ-NET-2] -->\n## Relay dependency\n```\n\n## Enforcement (four layers — defense in depth)\n\n1. **CI gate (hard backstop).** `traceable-reqs check --json` runs on every PR/push; exit-1 (`missing_stage` / `undeclared_id` / `parse_error` / `manifest_error`) **fails the build**. A PR can't merge if it leaves an *activated* req's required stage uncovered. Runs alongside the DOCS-STRATEGY CI gates.\n2. **Milestone activation gate.** Starting a milestone flips its reqs from `[]` to the real policy; the milestone **is not done until `check` is green for its reqs**. Coverage grows *with* the project instead of failing on day one.\n3. **Agent-facing rule.** The repo `CLAUDE.md` instructs every dev-agent: tag `[<stage>-><REQ-ID>]` in the same change as the evidence, and run `traceable-reqs check` before declaring done. Since spt-core is built *with* agents, the contract must be machine-followable — agent discipline + CI as the net.\n4. **Quality audit (anti-box-ticking).** `check` proves a tag *exists*; `traceable-reqs review` + `lint` audit whether tags sit near *real* evidence and whether titles are meaningful. Run advisory in CI (or periodically). Presence = hard gate; quality = audited.\n\nOptional fast-feedback: a pre-push git hook running `traceable-reqs check` locally.\n\n## CI snippet (wire at M0, once there is a Rust project + CI)\n\n```yaml\n# .github/workflows/traceability.yml\nname: traceability\non: [push, pull_request]\njobs:\n  check:\n    runs-on: ubuntu-latest\n    steps:\n      - uses: actions/checkout@v4\n      - name: Install traceable-reqs\n        run: gh release download --repo BigscreenVR/traceable-reqs --pattern '*linux-x86_64' --output traceable-reqs && chmod +x traceable-reqs\n        env: { GH_TOKEN: ${{ github.token }} }\n      - name: Coverage gate (hard)\n        run: ./traceable-reqs check --json\n      - name: Quality audit (advisory)\n        run: ./traceable-reqs lint || true\n```\n\n## Lifecycle\n\n- **Now (planning):** manifest seeded, all reqs inactive. `check` is green (nothing required yet).\n- **M0:** install the CLI, validate the seed against it, wire the CI workflow, **activate M0 reqs** (`REQ-ARCH-*`, `REQ-HAZARD-*` for the commodity layer), tag as code lands.\n- **M1–M5:** each milestone activates + covers its requirements before it's called done."}],"details":{"totalLines":75,"displayContent":{"text":"# Traceability development contract\r\n\r\n> How spt-core uses [`traceable-reqs`](https://github.com/BigscreenVR/traceable-reqs) to keep every requirement traced from doc → impl → test. The manifest is `traceable-reqs.toml` (seeded from PRD `R-*` + KNOWN-HAZARDS invariants). This contract makes the trace pay off instead of rotting.\r\n\r\n## Why\r\n\r\nThe whole point of the rebuild is to *not* re-introduce the sister project's bugs and to ship the PRD's scope verifiably. A requirement that isn't traced to real evidence (a doc, an implementation, a test) is a requirement you're *hoping* is met. This contract converts that hope into a build that fails when evidence is missing.\r\n\r\n## The contract (rules)\r\n\r\n1. **The manifest is the authoritative requirement registry.** Every PRD `R-*` requirement and every KNOWN-HAZARDS invariant exists in `traceable-reqs.toml` as a `REQ-*` id. No work without a REQ; no REQ without intent to satisfy. The PRD holds the prose; the manifest holds the id + `required_stages`.\r\n2. **Tag in the same change as the evidence.** When you write the function / test / doc-section that satisfies a stage, add its `[<stage>-><REQ-ID>]` tag in *that same commit*. Never \"tag later\" — this single rule is what stops the trace from drifting.\r\n3. **Evidence-proximate tags, not file-tops.** A tag sits on or immediately above the real function, test, doc section, config entry, or workflow step that proves the stage. One tag = one piece of real evidence. Tags at file tops to satisfy coverage are noise and usually wrong.\r\n4. **KNOWN-HAZARDS are first-class requirements.** Each invariant is a `REQ-HAZARD-*` requiring `unit` (and `int` where cross-process/cross-node). A hazard cannot be \"covered\" without a test tagging it — the anti-regression promise becomes mechanical (PRD success criterion #6).\r\n5. **Activation, not premature failure.** Every requirement starts `required_stages = []` (inactive) so `check` stays green pre-code. A **milestone (M0–M5) activates** the requirements it delivers by setting their real `required_stages`. Deferred items (`docs/DEFERRED.md`) stay `[]` until promoted.\r\n6. **Scan roots stay honest.** `[scan].roots` includes every evidence location (`src`, `tests`, `docs`, and later `.github/workflows`, `scripts`). Audit roots whenever a new evidence dir appears — a missing root makes evidence *silently* vanish from the trace.\r\n7. **Scope at mint.** Every new `REQ-*` joins at least one `[[groups]]` subject scope in the same commit that mints it. The groups partition is what the confusion audit (Gate 4, 2026-07-30) turned into manifest structure: the `req-confusion-audit` skill scores confusability only *within* groups, so an unscoped requirement's confusable pairs go unexamined — roughly half of them, by the audit's measurement. `span-*` groups are the declared cross-scope cover for measured-hot pairs; add one when a new requirement is known to sit confusably against another scope. Do not derive scopes from id prefixes — subject matter, not spelling.\r\n\r\n### Stages\r\n\r\n`doc` (prose/design/rustdoc) · `impl` (production code) · `unit` (unit test) · `int` (integration / e2e / cross-node). Default activation policy: `[\"doc\",\"impl\",\"unit\"]`; networking/lifecycle/cross-node reqs add `int`.\r\n\r\n### Tag examples\r\n\r\nReal tags are `[<stage>->REQ-ID]` with a bare stage word. The examples below\r\nangle-bracket the stage (`[<impl>->...]`) only so the scanner treats them as\r\nillustrations, not live evidence — write a bare stage word in real code.\r\n\r\n```rust\r\n// [<impl>->REQ-PAIR-1]\r\nfn spake2_pair(...) { ... }\r\n```\r\n```rust\r\n// [<unit>->REQ-HAZARD-ENVELOPE-DECODE-ORDER]\r\n#[test] fn decodes_amp_last() { ... }\r\n```\r\n```markdown\r\n<!-- [<doc>->REQ-NET-2] -->\r\n## Relay dependency\r\n```\r\n\r\n## Enforcement (four layers — defense in depth)\r\n\r\n1. **CI gate (hard backstop).** `traceable-reqs check --json` runs on every PR/push; exit-1 (`missing_stage` / `undeclared_id` / `parse_error` / `manifest_error`) **fails the build**. A PR can't merge if it leaves an *activated* req's required stage uncovered. Runs alongside the DOCS-STRATEGY CI gates.\r\n2. **Milestone activation gate.** Starting a milestone flips its reqs from `[]` to the real policy; the milestone **is not done until `check` is green for its reqs**. Coverage grows *with* the project instead of failing on day one.\r\n3. **Agent-facing rule.** The repo `CLAUDE.md` instructs every dev-agent: tag `[<stage>-><REQ-ID>]` in the same change as the evidence, and run `traceable-reqs check` before declaring done. Since spt-core is built *with* agents, the contract must be machine-followable — agent discipline + CI as the net.\r\n4. **Quality audit (anti-box-ticking).** `check` proves a tag *exists*; `traceable-reqs review` + `lint` audit whether tags sit near *real* evidence and whether titles are meaningful. Run advisory in CI (or periodically). Presence = hard gate; quality = audited.\r\n\r\nOptional fast-feedback: a pre-push git hook running `traceable-reqs check` locally.\r\n\r\n## CI snippet (wire at M0, once there is a Rust project + CI)\r\n\r\n```yaml\r\n# .github/workflows/traceability.yml\r\nname: traceability\r\non: [push, pull_request]\r\njobs:\r\n  check:\r\n    runs-on: ubuntu-latest\r\n    steps:\r\n      - uses: actions/checkout@v4\r\n      - name: Install traceable-reqs\r\n        run: gh release download --repo BigscreenVR/traceable-reqs --pattern '*linux-x86_64' --output traceable-reqs && chmod +x traceable-reqs\r\n        env: { GH_TOKEN: ${{ github.token }} }\r\n      - name: Coverage gate (hard)\r\n        run: ./traceable-reqs check --json\r\n      - name: Quality audit (advisory)\r\n        run: ./traceable-reqs lint || true\r\n```\r\n\r\n## Lifecycle\r\n\r\n- **Now (planning):** manifest seeded, all reqs inactive. `check` is green (nothing required yet).\r\n- **M0:** install the CLI, validate the seed against it, wire the CI workflow, **activate M0 reqs** (`REQ-ARCH-*`, `REQ-HAZARD-*` for the commodity layer), tag as code lands.\r\n- **M1–M5:** each milestone activates + covers its requirements before it's called done.\r","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75]},"fileSize":6057,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\docs\\TRACEABILITY.md"}}},"isError":false,"timestamp":1788999494789}}
{"type":"message","id":"f085e61a","parentId":"f6a80cbb","timestamp":"2026-09-10T00:18:14.790Z","message":{"role":"toolResult","toolCallId":"call_mOf9RXfq8KN14nxPXeceoPsF|fc_0dde72effda0ad21016aa1f746a5e087d2901182999239a7c6","toolName":"read","content":[{"type":"text","text":"# Documentation strategy\n\n> Draft (2026-05-31). A **designed-in commitment**, not an afterthought — grounded in research on critically-acclaimed developer docs (Stripe, Twilio, the Rust Book, FastAPI, Cloudflare, Anthropic, Diátaxis, `llms.txt`, Google dev style guide). Governs the **shipped product docs** (authored as code lands); the planning docs (CONTEXT/PRD/ADRs/design docs) are internal and separate.\n\n## The defining constraint: a dual audience\n\nspt-core's docs serve **two readers at once** — human developers *and* the **AI dev-agents that build integrations with it**. This is unusual and it is the design constraint: every artifact is authored **once in clean markdown** and served in **two depths** (human-rendered + agent-export). The agent layer is **first-class, not optional** — the strategic signal (Tailwind's docs traffic fell ~40% because assistants answer for it) is, for spt-core, *the goal*: make the agent layer so good that dev-agents integrate correctly on the first try. There is recursive fit here — the same kind of agent that spt-core *hosts* is the one that *reads its docs*.\n\n## Principles (top techniques, prioritized for spt-core)\n\n1. **Sub-10-minute killer quickstart** — runnable, deterministic, whole-working-thing-first, no placeholders. TTFHW (\"time to first hello world\") is the single most-cited conversion lever (~68% abandon trials over setup time).\n2. **Diátaxis four-mode separation** — tutorial / how-to / reference / explanation, never mixed (mixing is the most-cited cause of confusing docs).\n3. **Deterministic, real, copy-pasteable examples everywhere** — no `<YOUR_VALUE_HERE>` placeholders; real values that run. Serves humans *and* agents.\n4. **Dual-depth agent exports** — `llms.txt` (slim curated index) + `llms-full.txt` (full concatenated export), auto-generated in CI, plus markdown content negotiation (`.md` suffix / `Accept: text/markdown`) which cuts agent token use ~90% vs HTML.\n5. **One canonical way to do X** — explicitly mark deprecated/alternate paths. Non-determinism is fatal for agents.\n6. **Complete reference, auto-generated, all error variants** — rustdoc for the public SDK crates (`spt-proto`, `spt-runtime`, `spt-msg`) + a machine-readable manifest/contract schema. Generic placeholders in reference are a failure mode.\n7. **Consistent conversational voice** — adopt the Google developer style guide: second person (\"you\"), active voice, knowledgeable-friend tone.\n8. **Explain *why*, not just *what*** — conceptual docs + diagrams for the messaging model, the agent lifecycle states, instances/dormancy, and P2P topology.\n9. **Stable, never-renamed anchors/URLs** — agents cache links.\n10. **Docs-as-product, gated in CI** — generation (rustdoc, `llms.txt`, schema, CLI help) is part of the build so docs can never drift from code. Drift is the #1 most-cited docs failure; this kills it structurally.\n\n## Information architecture — by capability vertical\n\nOrganize by capability, each vertical carrying the same four Diátaxis modes internally (the Cloudflare pattern):\n\n**Messaging · Live-agent lifecycle · Terminal hosting · Networking (P2P + pairing) · Harness contract (manifest + `api`) · Instances · Shells · Self-update.**\n\n## Killer quickstart targets (two, one per audience)\n\n- **Human dev:** two agents exchange a message end-to-end in **one command + minimal config**, < 10 min — the core primitive, whole-thing-first, then decomposed.\n- **Dev-agent integrator:** a minimal **adapter** that satisfies the manifest + `api` contract and drives one lifecycle operation — the \"build a harness for spt-core\" hello-world.\n\n## Agent-consumable docs (spt-core specifics)\n\n<!-- [doc->REQ-DOCS-4] the agent-consumable layer's design; the schema leg shipped M6-D3 (schemars-derived, $id at the canonical URL), llms/CLI legs M6-D5, MCP + --json deferred (DEFERRED.md) -->\n\n- **`llms.txt` / `llms-full.txt`** — auto-emitted in CI; the slim index lets a tool answer quick questions, the full export feeds deep ingestion.\n- **Markdown content negotiation** + deterministic include/exclude tags so agent exports carry the canonical path and drop human-only narrative.\n- **Machine-readable manifest/contract schema** (JSON Schema) at a stable, discoverable path — agents validate their adapter manifests against it; the schema *is* documentation.\n- **MCP doc/resource server** — the standout fit, given the harness contract is the core integration surface. Expose the contract + reference as **deterministic, side-effect-free MCP resources** (same request → same result), giving integrating agents structured knowledge without RAG fuzziness. Highest-leverage agent-docs move for spt-core.\n- **CLI help as first-class agent docs** — agents primarily shell `spt api …`, so `spt <cmd> --help` must be complete, exampled, and machine-readable (a `--json`/structured help mode). The CLI surface is a documentation surface.\n\n## v0.1 deployment scope (grilled 2026-06-05)\n\nDeployed to GitHub Pages on `spt-releases` (ADR-0014; truth + CI generation stay in the source repo). **Tier 1 (ships with v0.1):** mental model/overview · human quickstart (install → two agents message) · adapter quickstart (via `mock-adapter`) · **harness contract vertical complete** (public manifest reference + full `spt api` reference + install-on-demand bootstrap pattern) · shell adapter getting-started (`spt-shell-notify` as the live example) · CLI reference generated from clap, CI-gated · **manifest JSON Schema (new code: schemars on the manifest types)** · root `llms.txt` + `llms-full.txt` CI-generated · rustdoc → Pages (no docs.rs — SDK crates are first-party-only while source is private).\n\n**Site generator: mdBook + custom theme CSS** (Rust-toolchain-coherent — one `cargo install mdbook` on the existing runners; the Rust Book IA this strategy copies *is* mdBook). Styling north star: **Astro Starlight** — copy its look/feel in the theme CSS, not its toolchain. Raw `.md` published alongside rendered pages (`/x.html` ↔ `/x.md`) for the agent-export convention; `llms.txt` / `llms-full.txt` / `manifest.schema.json` as static assets at site root.\n\n**Tier 2 (deferred, tracked):** remaining verticals' tutorials/how-tos (overview + reference stubs only in v0.1) · MCP doc server · `--help --json` mode (low marginal value over the generated CLI reference) · capstone tutorial · per-vertical `llms.txt` fan-out (root links pages directly until then).\n\n## Anti-patterns to design against (most-cited failures)\n\nDoc/code drift (#1 — solved by CI gating); *what* without *why*; too much setup before first success; generic placeholders in reference; mixed Diátaxis modes; poor search/navigation; multiple non-canonical ways to do X (fatal for agents).\n\n## CI commitments\n\nrustdoc, the manifest JSON Schema, `llms.txt`/`llms-full.txt`, and CLI help exports are **generated and checked in CI** — a doc-drift gate. Doc quality lives on the same footing as tests.\n\n<!-- [doc->REQ-DOCS-NO-INTERNAL-CODES] -->\n**Public help carries no internal codes.** The public CLI `--help` (the clap `///` doc-comments) and the generated `docs-site/src/cli/reference.md` MUST NOT contain internal tracker/decision codes — `REQ-*`, `F-###`, `M#-W#`, `ADR-####`. They are meaningless to an end user reading `--help` and ship to GitHub Pages. The `xtask check` docs gate scans the **generated** reference for these tokens and fails on any hit (regeneration alone keeps drift at zero but would faithfully republish a leak — the scan is what gives the gate teeth). The scan is scoped to the generated reference by construction, so rustdoc `///` on functions and structs is out of scope and keeps its `REQ`/`ADR` cross-references — those are source-internal, never public help. Substance stays; co…\n\n---\n\n## Concrete templates (from firsthand artifacts)\n\nExtracted by fetching the *actual* artifacts (Cloudflare/Anthropic/Mintlify `llms.txt`, FastAPI/Stripe/Supabase quickstarts, the Rust Book TOC, Django/Diátaxis, Tokio/clap rustdoc, gh/Stripe CLI). Copy the *shape*.\n\n### `llms.txt` — two-level (Cloudflare pattern, not Anthropic's flat dump)\n\nRoot `llms.txt` is a **curated index that fans out to a per-vertical `llms.txt`**; `llms-full.txt` is CI-concatenated page bodies (Anthropic's overflowed a 10 MB fetch — generation-only, never hand-authored).\n\n```markdown\n# spt-core\nA harness-independent core for an agent ecosystem: messaging, live-agent\nlifecycle, terminal hosting, P2P networking, a runtime-manifest harness contract.\n\n> Each capability links to its own llms.txt. Full concatenated docs: /llms-full.txt.\n> Machine-readable contract: /manifest.schema.json. Append `.md` to any doc URL\n> for raw markdown (~90% fewer tokens than HTML).\n\n## Start here\n- [Quickstart: two agents exchange a message](…/quickstart/messaging.md): one command, <10 min.\n- [Quickstart: build an adapter](…/quickstart/adapter.md): minimal manifest + `api` contract.\n- [Mental model](…/concepts/overview.md): daemon, endpoints, instances, subnet.\n\n## Capabilities      (one bullet per vertical, each → its own llms.txt)\n- [Messaging](…/messaging/llms.txt): typed ops + file blobs; send / ring / ready.\n- [Live-agent lifecycle](…/lifecycle/llms.txt): skeleton → live → dormant → suspended.\n- … Terminal · Networking · Harness contract · Instances · Shells · Self-update\n\n## Reference (auto-generated)\n- [spt-proto](https://docs.rs/spt-proto) · [spt-runtime](https://docs.rs/spt-runtime) · [spt-msg](https://docs.rs/spt-msg)\n- [CLI reference](…/cli/llms.txt) · [Manifest JSON Schema](…/manifest.schema.json)\n```\n\n### Quickstart — whole-thing → run → verify → dissect (FastAPI), with a 2-path fork (Supabase)\n\nZero placeholders; every value runs; complete working thing shown *before* explanation; audience fork at the top (spt-core has two readers); end with the Diátaxis trident.\n\n```markdown\n# Quickstart — two agents exchange a message, end to end, <10 min.\n> Pick your path: **developer** (here) · **building an adapter / dev-agent** → /quickstart/adapter.md\n\n## 1. Install (one line)        curl -fsSL https://spt-core.dev/install.sh | sh\n   Verify:                      spt --version      # spt 0.1.0 (proto v1)\n## 2. The whole working thing   spt ready alice\n                                spt ready bob\n                                spt send bob --from alice \"hello\"\n## 3. Run it                    spt poll --as bob\n                                # → from=alice op=message body=\"hello\"\n## 4. What just happened        (dissect each line; note the daemon auto-started)\n## Next                         how-to · concept · reference  (Diátaxis trident)\n```\n\n### Information architecture — two axes\n\n**Global sequence (Rust Book logic):** early runnable project, dependency-ordered concepts, capstone last.\n```\n0 Getting started → 1 Quickstart (the early runnable project) → 2 Mental model\n→ 3 Capability verticals (build-order) → 4 Capstone: build a non-CC adapter end-to-end\n```\n**Vertical ordering (dependency-sequenced):**\n`Messaging → Lifecycle → Terminal → Networking → Harness contract → Instances → Shells → Self-update`\n(messaging is the primitive the quickstart needs; self-update last because its no-termination invariant depends on the terminal/broker model.)\n**Per-vertical internal template (Django labels × Cloudflare ordering):**\n`Overview (why+diagram) · Quickstart/Tutorial · How-to guides · Reference (rustdoc+CLI+schema) · llms.txt`\n\n### rustdoc — Tokio opener + out-document clap on errors\n\nCrate doc: **value-prop → \"A tour of\" → runnable example near top → feature flags → proto-compat note.** Every public error enum documents **per-variant cause → fix**; every fallible fn gets an `# Errors` section.\n\n```rust\n//! # spt-runtime — one-sentence value prop.\n//! ## A tour of spt-runtime   - [`Manifest`] · [`ApiCommand`] · [`Lifecycle`]\n//! ## Example                 (runnable, real values, near the top)\n//! ## Feature flags           `net` (default) · `unstable`\n//! ## Wire-protocol compatibility   interop by proto version, never crate version.\n```\n```rust\n/// The manifest's `min_spt_core_version` exceeds this binary.\n/// **Fix:** update spt-core, or pin the adapter to a compatible release.\nVersionTooOld { required: Version, found: Version },\n```\n\n### CLI-as-docs — gh buckets + `--json` help mode\n\nTop-level `--help` groups by audience (`AGENT COMMANDS` send/ring/ready/poll · `HARNESS COMMANDS` api … · `NODE COMMANDS` pair/subnet/update). Per-command: **USAGE · FLAGS · EXAMPLES (real values) · LEARN MORE**. Beyond gh: a **`spt <cmd> --help --json`** structured-help mode so dev-agents parse the contract without scraping, plus gh's uniform **`--json <fields>` / `--jq` / `--template`** trio on any list command. The `--json` help export is generated + CI-gated alongside rustdoc and the schema.\n\n[Some lines truncated to 768 chars]"}],"details":{"totalLines":147,"displayContent":{"text":"# Documentation strategy\r\n\r\n> Draft (2026-05-31). A **designed-in commitment**, not an afterthought — grounded in research on critically-acclaimed developer docs (Stripe, Twilio, the Rust Book, FastAPI, Cloudflare, Anthropic, Diátaxis, `llms.txt`, Google dev style guide). Governs the **shipped product docs** (authored as code lands); the planning docs (CONTEXT/PRD/ADRs/design docs) are internal and separate.\r\n\r\n## The defining constraint: a dual audience\r\n\r\nspt-core's docs serve **two readers at once** — human developers *and* the **AI dev-agents that build integrations with it**. This is unusual and it is the design constraint: every artifact is authored **once in clean markdown** and served in **two depths** (human-rendered + agent-export). The agent layer is **first-class, not optional** — the strategic signal (Tailwind's docs traffic fell ~40% because assistants answer for it) is, for spt-core, *the goal*: make the agent layer so good that dev-agents integrate correctly on the first try. There is recursive fit here — the same kind of agent that spt-core *hosts* is the one that *reads its docs*.\r\n\r\n## Principles (top techniques, prioritized for spt-core)\r\n\r\n1. **Sub-10-minute killer quickstart** — runnable, deterministic, whole-working-thing-first, no placeholders. TTFHW (\"time to first hello world\") is the single most-cited conversion lever (~68% abandon trials over setup time).\r\n2. **Diátaxis four-mode separation** — tutorial / how-to / reference / explanation, never mixed (mixing is the most-cited cause of confusing docs).\r\n3. **Deterministic, real, copy-pasteable examples everywhere** — no `<YOUR_VALUE_HERE>` placeholders; real values that run. Serves humans *and* agents.\r\n4. **Dual-depth agent exports** — `llms.txt` (slim curated index) + `llms-full.txt` (full concatenated export), auto-generated in CI, plus markdown content negotiation (`.md` suffix / `Accept: text/markdown`) which cuts agent token use ~90% vs HTML.\r\n5. **One canonical way to do X** — explicitly mark deprecated/alternate paths. Non-determinism is fatal for agents.\r\n6. **Complete reference, auto-generated, all error variants** — rustdoc for the public SDK crates (`spt-proto`, `spt-runtime`, `spt-msg`) + a machine-readable manifest/contract schema. Generic placeholders in reference are a failure mode.\r\n7. **Consistent conversational voice** — adopt the Google developer style guide: second person (\"you\"), active voice, knowledgeable-friend tone.\r\n8. **Explain *why*, not just *what*** — conceptual docs + diagrams for the messaging model, the agent lifecycle states, instances/dormancy, and P2P topology.\r\n9. **Stable, never-renamed anchors/URLs** — agents cache links.\r\n10. **Docs-as-product, gated in CI** — generation (rustdoc, `llms.txt`, schema, CLI help) is part of the build so docs can never drift from code. Drift is the #1 most-cited docs failure; this kills it structurally.\r\n\r\n## Information architecture — by capability vertical\r\n\r\nOrganize by capability, each vertical carrying the same four Diátaxis modes internally (the Cloudflare pattern):\r\n\r\n**Messaging · Live-agent lifecycle · Terminal hosting · Networking (P2P + pairing) · Harness contract (manifest + `api`) · Instances · Shells · Self-update.**\r\n\r\n## Killer quickstart targets (two, one per audience)\r\n\r\n- **Human dev:** two agents exchange a message end-to-end in **one command + minimal config**, < 10 min — the core primitive, whole-thing-first, then decomposed.\r\n- **Dev-agent integrator:** a minimal **adapter** that satisfies the manifest + `api` contract and drives one lifecycle operation — the \"build a harness for spt-core\" hello-world.\r\n\r\n## Agent-consumable docs (spt-core specifics)\r\n\r\n<!-- [doc->REQ-DOCS-4] the agent-consumable layer's design; the schema leg shipped M6-D3 (schemars-derived, $id at the canonical URL), llms/CLI legs M6-D5, MCP + --json deferred (DEFERRED.md) -->\r\n\r\n- **`llms.txt` / `llms-full.txt`** — auto-emitted in CI; the slim index lets a tool answer quick questions, the full export feeds deep ingestion.\r\n- **Markdown content negotiation** + deterministic include/exclude tags so agent exports carry the canonical path and drop human-only narrative.\r\n- **Machine-readable manifest/contract schema** (JSON Schema) at a stable, discoverable path — agents validate their adapter manifests against it; the schema *is* documentation.\r\n- **MCP doc/resource server** — the standout fit, given the harness contract is the core integration surface. Expose the contract + reference as **deterministic, side-effect-free MCP resources** (same request → same result), giving integrating agents structured knowledge without RAG fuzziness. Highest-leverage agent-docs move for spt-core.\r\n- **CLI help as first-class agent docs** — agents primarily shell `spt api …`, so `spt <cmd> --help` must be complete, exampled, and machine-readable (a `--json`/structured help mode). The CLI surface is a documentation surface.\r\n\r\n## v0.1 deployment scope (grilled 2026-06-05)\r\n\r\nDeployed to GitHub Pages on `spt-releases` (ADR-0014; truth + CI generation stay in the source repo). **Tier 1 (ships with v0.1):** mental model/overview · human quickstart (install → two agents message) · adapter quickstart (via `mock-adapter`) · **harness contract vertical complete** (public manifest reference + full `spt api` reference + install-on-demand bootstrap pattern) · shell adapter getting-started (`spt-shell-notify` as the live example) · CLI reference generated from clap, CI-gated · **manifest JSON Schema (new code: schemars on the manifest types)** · root `llms.txt` + `llms-full.txt` CI-generated · rustdoc → Pages (no docs.rs — SDK crates are first-party-only while source is private).\r\n\r\n**Site generator: mdBook + custom theme CSS** (Rust-toolchain-coherent — one `cargo install mdbook` on the existing runners; the Rust Book IA this strategy copies *is* mdBook). Styling north star: **Astro Starlight** — copy its look/feel in the theme CSS, not its toolchain. Raw `.md` published alongside rendered pages (`/x.html` ↔ `/x.md`) for the agent-export convention; `llms.txt` / `llms-full.txt` / `manifest.schema.json` as static assets at site root.\r\n\r\n**Tier 2 (deferred, tracked):** remaining verticals' tutorials/how-tos (overview + reference stubs only in v0.1) · MCP doc server · `--help --json` mode (low marginal value over the generated CLI reference) · capstone tutorial · per-vertical `llms.txt` fan-out (root links pages directly until then).\r\n\r\n## Anti-patterns to design against (most-cited failures)\r\n\r\nDoc/code drift (#1 — solved by CI gating); *what* without *why*; too much setup before first success; generic placeholders in reference; mixed Diátaxis modes; poor search/navigation; multiple non-canonical ways to do X (fatal for agents).\r\n\r\n## CI commitments\r\n\r\nrustdoc, the manifest JSON Schema, `llms.txt`/`llms-full.txt`, and CLI help exports are **generated and checked in CI** — a doc-drift gate. Doc quality lives on the same footing as tests.\r\n\r\n<!-- [doc->REQ-DOCS-NO-INTERNAL-CODES] -->\r\n**Public help carries no internal codes.** The public CLI `--help` (the clap `///` doc-comments) and the generated `docs-site/src/cli/reference.md` MUST NOT contain internal tracker/decision codes — `REQ-*`, `F-###`, `M#-W#`, `ADR-####`. They are meaningless to an end user reading `--help` and ship to GitHub Pages. The `xtask check` docs gate scans the **generated** reference for these tokens and fails on any hit (regeneration alone keeps drift at zero but would faithfully republish a leak — the scan is what gives the gate teeth). The scan is scoped to the generated reference by construction, so rustdoc `///` on functions and structs is out of scope and keeps its `REQ`/`ADR` cross-references — those are source-internal, never public help. Substance stays; co…\n\r\n---\r\n\r\n## Concrete templates (from firsthand artifacts)\r\n\r\nExtracted by fetching the *actual* artifacts (Cloudflare/Anthropic/Mintlify `llms.txt`, FastAPI/Stripe/Supabase quickstarts, the Rust Book TOC, Django/Diátaxis, Tokio/clap rustdoc, gh/Stripe CLI). Copy the *shape*.\r\n\r\n### `llms.txt` — two-level (Cloudflare pattern, not Anthropic's flat dump)\r\n\r\nRoot `llms.txt` is a **curated index that fans out to a per-vertical `llms.txt`**; `llms-full.txt` is CI-concatenated page bodies (Anthropic's overflowed a 10 MB fetch — generation-only, never hand-authored).\r\n\r\n```markdown\r\n# spt-core\r\nA harness-independent core for an agent ecosystem: messaging, live-agent\r\nlifecycle, terminal hosting, P2P networking, a runtime-manifest harness contract.\r\n\r\n> Each capability links to its own llms.txt. Full concatenated docs: /llms-full.txt.\r\n> Machine-readable contract: /manifest.schema.json. Append `.md` to any doc URL\r\n> for raw markdown (~90% fewer tokens than HTML).\r\n\r\n## Start here\r\n- [Quickstart: two agents exchange a message](…/quickstart/messaging.md): one command, <10 min.\r\n- [Quickstart: build an adapter](…/quickstart/adapter.md): minimal manifest + `api` contract.\r\n- [Mental model](…/concepts/overview.md): daemon, endpoints, instances, subnet.\r\n\r\n## Capabilities      (one bullet per vertical, each → its own llms.txt)\r\n- [Messaging](…/messaging/llms.txt): typed ops + file blobs; send / ring / ready.\r\n- [Live-agent lifecycle](…/lifecycle/llms.txt): skeleton → live → dormant → suspended.\r\n- … Terminal · Networking · Harness contract · Instances · Shells · Self-update\r\n\r\n## Reference (auto-generated)\r\n- [spt-proto](https://docs.rs/spt-proto) · [spt-runtime](https://docs.rs/spt-runtime) · [spt-msg](https://docs.rs/spt-msg)\r\n- [CLI reference](…/cli/llms.txt) · [Manifest JSON Schema](…/manifest.schema.json)\r\n```\r\n\r\n### Quickstart — whole-thing → run → verify → dissect (FastAPI), with a 2-path fork (Supabase)\r\n\r\nZero placeholders; every value runs; complete working thing shown *before* explanation; audience fork at the top (spt-core has two readers); end with the Diátaxis trident.\r\n\r\n```markdown\r\n# Quickstart — two agents exchange a message, end to end, <10 min.\r\n> Pick your path: **developer** (here) · **building an adapter / dev-agent** → /quickstart/adapter.md\r\n\r\n## 1. Install (one line)        curl -fsSL https://spt-core.dev/install.sh | sh\r\n   Verify:                      spt --version      # spt 0.1.0 (proto v1)\r\n## 2. The whole working thing   spt ready alice\r\n                                spt ready bob\r\n                                spt send bob --from alice \"hello\"\r\n## 3. Run it                    spt poll --as bob\r\n                                # → from=alice op=message body=\"hello\"\r\n## 4. What just happened        (dissect each line; note the daemon auto-started)\r\n## Next                         how-to · concept · reference  (Diátaxis trident)\r\n```\r\n\r\n### Information architecture — two axes\r\n\r\n**Global sequence (Rust Book logic):** early runnable project, dependency-ordered concepts, capstone last.\r\n```\r\n0 Getting started → 1 Quickstart (the early runnable project) → 2 Mental model\r\n→ 3 Capability verticals (build-order) → 4 Capstone: build a non-CC adapter end-to-end\r\n```\r\n**Vertical ordering (dependency-sequenced):**\r\n`Messaging → Lifecycle → Terminal → Networking → Harness contract → Instances → Shells → Self-update`\r\n(messaging is the primitive the quickstart needs; self-update last because its no-termination invariant depends on the terminal/broker model.)\r\n**Per-vertical internal template (Django labels × Cloudflare ordering):**\r\n`Overview (why+diagram) · Quickstart/Tutorial · How-to guides · Reference (rustdoc+CLI+schema) · llms.txt`\r\n\r\n### rustdoc — Tokio opener + out-document clap on errors\r\n\r\nCrate doc: **value-prop → \"A tour of\" → runnable example near top → feature flags → proto-compat note.** Every public error enum documents **per-variant cause → fix**; every fallible fn gets an `# Errors` section.\r\n\r\n```rust\r\n//! # spt-runtime — one-sentence value prop.\r\n//! ## A tour of spt-runtime   - [`Manifest`] · [`ApiCommand`] · [`Lifecycle`]\r\n//! ## Example                 (runnable, real values, near the top)\r\n//! ## Feature flags           `net` (default) · `unstable`\r\n//! ## Wire-protocol compatibility   interop by proto version, never crate version.\r\n```\r\n```rust\r\n/// The manifest's `min_spt_core_version` exceeds this binary.\r\n/// **Fix:** update spt-core, or pin the adapter to a compatible release.\r\nVersionTooOld { required: Version, found: Version },\r\n```\r\n\r\n### CLI-as-docs — gh buckets + `--json` help mode\r\n\r\nTop-level `--help` groups by audience (`AGENT COMMANDS` send/ring/ready/poll · `HARNESS COMMANDS` api … · `NODE COMMANDS` pair/subnet/update). Per-command: **USAGE · FLAGS · EXAMPLES (real values) · LEARN MORE**. Beyond gh: a **`spt <cmd> --help --json`** structured-help mode so dev-agents parse the contract without scraping, plus gh's uniform **`--json <fields>` / `--jq` / `--template`** trio on any list command. The `--json` help export is generated + CI-gated alongside rustdoc and the schema.\r","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120,121,122,123,124,125,126,127,128,129,130,131,132,133,134,135,136,137,138,139,140,141,142,143,144,145,146,147]},"fileSize":13143,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\docs\\DOCS-STRATEGY.md"},"limits":{"columnTruncated":{"maxColumn":768}}}},"isError":false,"timestamp":1788999494790}}
{"type":"message","id":"310d1383","parentId":"f085e61a","timestamp":"2026-09-10T00:18:14.859Z","message":{"role":"toolResult","toolCallId":"call_4RwL2YnZyT4FlekaOudPGGEj|fc_0dde72effda0ad21016aa1f746a5c087d2936d67a0e7d2e2ed","toolName":"glob","content":[{"type":"text","text":"AGENTS.md\n# .worktrees/286-churn/\nAGENTS.md\n## crates/spt-daemon/src/\nlivehost.rs\nconn.rs\n\nSkipped missing paths: .worktrees/AGENTS.md"}],"details":{"scopePath":".worktrees/286-churn/**/AGENTS.md, .worktrees/286-churn/**/conn.rs, .worktrees/286-churn/**/livehost.rs, AGENTS.md","fileCount":4,"files":[".worktrees/286-churn/crates/spt-daemon/src/livehost.rs",".worktrees/286-churn/crates/spt-daemon/src/conn.rs",".worktrees/286-churn/AGENTS.md","AGENTS.md"],"truncated":false,"cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","missingPaths":[".worktrees/AGENTS.md"]},"isError":false,"timestamp":1788999494858}}
{"type":"message","id":"da45edc7","parentId":"310d1383","timestamp":"2026-09-10T00:18:14.910Z","message":{"role":"toolResult","toolCallId":"call_oPXXqMg4wN89WrblFZiU7qSh|fc_0dde72effda0ad21016aa1f746a5d087d2b7fb6f323a04cbd6","toolName":"read","content":[{"type":"text","text":"# Known Hazards\n\nHard-won edge cases harvested from the sister project (`claude_skill_owl`, ~80 commits / 12+ phases / multiple production incidents). Per ADR-0001, this is a **test checklist for the spt-core rebuild** — the clean-room rebuild must re-satisfy each invariant rather than re-discover the bug.\n\n**Architecture-translation note.** The sister project runs poll listeners and Psyche wrappers as *separate processes*. spt-core consolidates both into the one `spt-daemon` (brain), with a stable broker beneath it (ADR-0004). Many hazards below were inter-process races in the sister project; in spt-core some become intra-daemon concerns (potentially easier) while others move to the daemon↔broker IPC boundary or the network boundary (potentially new failure surface). Each entry notes the mapping where it differs. Citations point at sister-project paths for reference, not at spt-core.\n\n---\n\n## 1. Race conditions & ordering\n\n### 1.1 Phantom INIT_SIGNOFF after grace period\n- **Failure:** orphan teardown enqueues INIT_SIGNOFF before the grace-period recheck; a transient Self recovery (binary handoff, brief stale poll) makes the recheck pass-as-alive, but the signoff was already spooled and drains on the next iteration → teardown despite a live Self.\n- **Invariant:** grace-period wait MUST complete *before* composing/delivering INIT_SIGNOFF; the recheck must bind `still_gone` before any envelope write.\n- **spt-core mapping:** in-daemon now (no separate wrapper), but the ordering invariant is identical — orphan/teardown logic must re-evaluate liveness after the grace wait, not before enqueue.\n- **Sister cite:** `src/live/wrapper/orphan.rs:201-259` (sleep@209 precedes compose@231-251); tests T-grace-recovery:576, T-still-gone-recheck:618.\n\n### 1.2 Poll-rewrite race & info.json mid-write reads\n- **Failure:** `info.json` written by the wrapper mid-iteration while a list/classify command reads it → torn read, misclassification.\n- **Invariant:** consult liveness via the supervisor (`is_wrapper_alive`-equivalent) before any grace gate; reads of state files must tolerate concurrent writes (atomic write + rename, or read-retry).\n- **spt-core mapping:** the daemon owns both writer and reader → use in-process locking/snapshotting instead of racing on disk. Cross-node registry reads remain eventually-consistent and must tolerate staleness.\n- **Sister cite:** `src/common/list_filter.rs:100-150`; `src/owl/poll.rs:141`.\n\n### 1.3 Stale `index.lock` wedge from prior git crash\n- **Failure:** crashed git leaves a 0-byte `index.lock` in a psyche tracked worktree; every later commit blocks forever.\n- **Invariant:** on daemon boot, sweep seed + all agent/project worktrees for stale locks (0 bytes, mtime > 60s) and remove; leave live locks alone.\n- **spt-core mapping:** cross-node Psyche sync (ADR-0002/0003) replaces git-repo sync, so the *git* lock may disappear — but any equivalent lockfile in the new sync mechanism needs the same stale-sweep on boot.\n- **Sister cite:** CHANGELOG v1.11.20 \"Stale `index.lock`\"; `src/common/git.rs`.\n\n### 1.4 Deferred spool rows must not leak to the event stream\n- **Failure:** a hook spools a deferred (spool-only, no TCP wake) notice; startup `drain_all` flushes ALL rows including deferred → event emitted at wrong time/priority.\n- **Invariant:** startup drain (and idle/timeout TCP-wake sites) use `drain_non_deferred` only; deferred rows are picked up by their intended consumer via `peek`. All drain sites must agree on which rows they flush.\n- **spt-core mapping:** carries directly — the daemon's spool-drain has the same deferred-vs-immediate distinction.\n- **Sister cite:** `src/owl/poll.rs:276-316`; `spool::drain_non_deferred_with_metadata`.\n\n### 1.5 Worker (working-perch) lifecycle path consistency\n- **Failure:** subagent-start creates the perch at one path layout; later hooks read it at another → not found; stop-hook scan misses nested perches.\n- **Invariant:** all Worker/Psyche child-perch path composition routes through one central resolver; no divergent path construction across hooks.\n- **spt-core mapping:** `Worker` is a day-one endpoint type; the daemon owns the registry, so perch location is a registry lookup, not ad-hoc path math. Single source of truth for instance→location.\n- **Sister cite:** `src/owl/hook_subagent_start.rs:122-168`; `hook_subagent_stop.rs:15-55`.\n\n---\n\n## 2. Identity & session-binding\n\n### 2.1 Parent PID over ephemeral poll PID\n- **Failure:** orphan check polls an ephemeral listener PID; it dies and is recycled (esp. Windows); a foreign process with the recycled PID reads as alive → false-positive teardown (or false-negative).\n- **Invariant:** prefer the stable harness-session PID (`parent_pid`) over any ephemeral process PID for liveness; minimal `info.json` for supervisor-owned perches to avoid stale leaks.\n- **spt-core mapping:** session binding (parent-process-tree anchor) still applies for harness-hosted topology. For spt-hosted sessions the broker holds the child directly → liveness is the broker's held-handle state, more reliable than PID polling.\n- **Sister cite:** `src/live/wrapper/orphan.rs:141-161`; CHANGELOG v1.11.20.\n\n### 2.2 Stdin session_id precedence over env\n- **Failure:** subagent inherits a stale `OWL_SESSION_ID` env across `/clear`; hook gets two session_ids (fresh stdin, stale env) → wrong-agent binding.\n- **Invariant:** stdin-provided session_id wins; env is fallback only.\n- **spt-core mapping:** the harness-contract subcommand surface must define the same precedence for whatever identity fields hooks pass in.\n- **Sister cite:** CHANGELOG v1.35.1 \"IN-05\"; `hook_subagent_start.rs:40-51`.\n\n### 2.3 Binary-handoff argv schema must stay backward-compatible\n- **Failure:** old binary spawns new binary with old argv arity; clap rejects before state rehydration → wrapper dies unlogged.\n- **Invariant:** every newly-added handoff positional has a default; state-file rehydration happens *after* argv parse; defaults survive intermediate versions.\n- **spt-core mapping:** CRITICAL — self-update (ADR-0004) makes handoff routine. The broker↔brain IPC and any brain-relaunch argv must be versioned and forward/backward tolerant (a newer brain talks to an older broker). This is the single most update-frequency-sensitive invariant.\n- **Sister cite:** `src/live/wrapper/lifecycle.rs:17-106`; `src/cli.rs` defaults; CHANGELOG v1.11.10.\n\n### 2.4 Generation `gen_start` always = now() on cold-start AND handoff\n- **Failure:** stale gen_start from a rehydrated state file fires time-based discriminators on the new process.\n- **Invariant:** wall-clock `gen_start` is set to `now()` on both cold-start and handoff; generation counter increments on every start/revive; session UUID captured fresh and carried so the resumed mind distinguishes \"same gen continuing\" vs \"new gen born\".\n- **spt-core mapping:** carries to the daemon's per-instance generation tracking.\n- **Restoration D3/D4 (ADR-0018):** the generation *counter* custody moved to the broker (D3-2 — it observes every brain spawn, planned or crash, and hands `{generation, start-reason}` at spawn; `gen_start` stays `now()`-fresh, never rehydrated). The brain→brain **`BrainState` *message*** (`{session_id, generation, next_seq, gen_start_ms}`) that previously carried continuity across a handoff is **retired from the production path in D4-2**: a brain the supervisor respawns cold-starts and reconstructs all session continuity by **querying the broker** (`Brain::resume_sessions` over the broker's cursor-of-record), never a frame. `BrainState` / `Brain::handoff` / `Brain::snapshot` remain `pub` and compiled **only for the integration tests** (handoff/idempotent/da…\n- **Sister cite:** `src/live/wrapper/lifecycle.rs:70`; `src/common/wrapper_state.rs`.\n\n### 2.5 Daemon-hosted endpoints have no dedicated liveness PID\n- **Failure:** the sister evaluates Psyche/perch liveness via a dedicated process PID — the wrapper's own pid in `info.json`, checked with `is_process_alive`. Under ADR-0004 the Psyche (and any spt-hosted Self) is a **loop inside the daemon**, not a separate process: it holds no dedicated pid, and its `claude`/summarizer subprocess is ephemeral (spawned per pulse/commune, then exits). If a daemon-hosted perch's `info.json` carries the **daemon's** pid, then *every* hosted endpoint shares one pid, and `is_process_alive(pid)` reads \"alive\" for a torn-down endpoint as long as the daemon runs — while `clean_stale_entries` (dead-pid deletion) can no longer distinguish a dead endpoint from a live one. The 2.1/5.1 liveness models do **not** cover this third categor…\n- **Invariant:** for **daemon-hosted** perches (Psyche; spt-hosted Self), liveness is the **daemon's authoritative in-memory endpoint table + a `status` field** on `info.json` (`online|offline|…`), **never** `is_process_alive(info.pid)`. `info.pid` for a daemon-hosted perch is at most a *hosted-by-daemon* marker (the daemon pid), not a liveness signal; registry stale-clean for these rows keys on the daemon's endpoint table, not per-row pid. This reuses the pattern already specified for **Shells** (`info.json` carries daemon-managed `status`, capability resolved by `adapter_name` — CONTEXT \"Shell… Not in the subnet registry\") and extends it to daemon-hosted *agent* perches.\n- **spt-core mapping:** the **M1/M2a interim** model keeps the Psyche/listener a real per-process owner (the `api listen` process), so its per-pid liveness (`deliver::is_online` → `info.read_pid` → `proc::is_process_alive`; `registry::clean_stale_entries`) is correct *interim*. **M3 daemon consolidation replaces it** with daemon-authoritative liveness for hosted perches. Keep the liveness check behind one resolver (mirrors `resolve_address` stale-clean) so the M3 swap is localized — do **not** let the per-pid assumption leak into new call sites.\n- **Sister cite:** `src/live/wrapper/orphan.rs` (wrapper-pid liveness); `src/common/list_filter.rs:168-175` (pid-classify); spt-core `crates/spt-store/src/{proc.rs,registry.rs}` + `crates/spt-msg/src/deliver.rs::is_online`.\n\n### 2.6 A shell instance's `online` is a recorded field with no death edge\n- **Failure:** a shell perch's `status` has exactly two writers — the bind handshake (→`online`) and `close_shell` (→`offline`) — and `close_shell` runs on the **link-break** path alone. A binary that dies **abruptly** (force-kill, crash, OOM) breaks no link, so nothing ever writes the flip: the record says `online` forever while the pid it names is gone. Every consumer of the recorded field then inherits the lie — `shell list` reports online indefinitely; `relink` (the one command that recovers the instance) refuses `SHELL_ALREADY_ONLINE`, gated on the very state that is wrong; and `shell cmd`'s wake-if-offline arm reads \"not offline\", skips the wake, and spools the command to a corpse — accepted happily, drained by nobody, **indistinguishable from a busy s…\n- **Invariant:** ONLINE-ness is **derived**, never the recorded field alone: `status == online` **AND** the recorded `shell.pid` is not *provably* dead. Provably dead is narrow — pid parked, non-zero, `!is_process_alive`; pid absent, unparseable, or `0` (a broker-hosted spawn whose backend exposed no pid) reads **alive**, the same fail-toward-alive parity 2.5 holds, so a backend the resolver cannot see is never declared dead. **The recycled-pid caveat is narrowed since BAROMETER W2 and is now conditional on the stamp:** a launch parks a **birth stamp** beside the pid (`shell.launch.json`), and the probe runs the PAIR test — same pid, *different* start time ⇒ **gone**, so a pid the OS handed to an unrelated process no longer masks a dead instance. That case m…\n- **No spontaneous relaunch of a SAME-BOOT death (operator-ratified, flynn 2026-07-25; re-based BAROMETER W2):** the reconciler does NOT adopt an instance whose binary died during this boot. Every reason an operator stops a shell is a reason not to want it back a tick later: mid-deploy (a shared install dir makes \"kill the process\" routine on Windows, since a running exe cannot be overwritten) an auto-relaunch would run the **old** binary out of the file being replaced — worse than a failed install — and it turns quarantining a misbehaving shell into a restart loop. For those, recovery stays demand-driven: `relink`, or a `shell cmd` that wakes. `relink` probes **locally** rather than trusting a daemon sweep, so recovery holds with the daemon down.\n- **What holds that open is no longer the stale record (BAROMETER W2):** until W2 the guarantee rested on an accident: a force-killed instance was held out of the eligible set because its record still said `online`, and the eligibility read took that field raw. 2.7 heals that record to the truth, which removes the accidental protection — so the guard had to move to a rule that *states* it. Eligibility is now recorded-`offline` **AND** (no corpse, **or** the corpse's recorded launch **predates the boot instant**). A steady-state force-kill leaves a corpse launched *after* boot, which the predicate can never accept, so the ruling is preserved **by construction** rather than by a second rule. A restart casualty is the one case the predicate does accept, and it …\n- **spt-core mapping:** `spt_store::shellinfo::{shell_pid_provably_dead, effective_status, is_shell_online}` (the resolver — the shell-side twin of 2.5's `liveness.rs`, which gave *agent* perches exactly this and which shells never got); consumers in `linkhost::{relink_shell, prepare_drive, run_action}`, `shelldisc::discover`, `activity::observe_links`; the class-(c) eligibility read is `spt_daemon::shellwake::watcher_eligible` (+ `launch_predates_boot`, `BOOT_RESTORE_SLACK_MS`), and the birth stamp it rests on is `shellinfo::{ShellLaunch, record_shell_launch, read_shell_launch}`.\n- **Source:** spt-core, flynn's spt-alchemy field report (2026-07-25) — a deterministic recipe, not a race. Class-(c) re-based BAROMETER W2 (releases#78, doyle's ruling on comment 5156810267).\n\n### 2.7 A node restart permanently strands every `persistent` shell — and the existing cascade test cannot see that it does  `[REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]`\n- **Failure:** the contract's own sentence says a `persistent` shell is **online whenever its owner endpoint is online**, and a node restart is the one case where that promise was never kept. A machine death breaks no link, so `close_shell` never runs and 2.6's stale record survives the reboot; when the owner comes back online, the wake cascade reads the **recorded** field, sees `online`, and skips the relaunch of a binary that has not existed since the previous boot. Nothing else corrects it: the instance is down forever, with no error anywhere, until a human happens to run `relink`. Every `persistent` instance on the node is stranded by one restart, together.\n- **The owner-facing surfaces HIDE the fault**, which is why it survived unreported for days: every display path (`shell list`, its `--json` twin, the shell-context render) routes through the ONE discovery seam, which **derives** status (2.6) and so correctly reports the binary as offline — while the only consumer whose decision matters, the wake cascade, reads the RECORDED field that no owner-facing surface shows. Absence of a visible symptom is therefore **not** evidence of absence of this fault, and a clean-looking context render must never be taken as proof the class did not occur. The correct falsifier is the **on-disk record**.\n- **The existing cascade test cannot fail on it — its SETUP SUPPRESSES THE FAILING ARM.** `rest_edges_cascade_shells_with_divergence` establishes its fixture by driving the suspend path, which itself writes the `offline` the code under test is supposed to encounter — so the restart shape (a record still saying `online` over a corpse, reached with **no** suspend edge) is unreachable from that fixture BY CONSTRUCTION. The test is not weak; it gives a true answer to the wrong question, so a reader asking \"is the cascade covered?\" is told yes. This is the third suppressed-arm instance found in one week, and the pattern is the same each time: **a fixture that establishes the precondition the code under test is supposed to establish.**\n- **Invariant:** a node restart must not strand a `persistent` instance. Two legs, and each is worthless alone. **(a) The record stops lying:** the daemon-side reconciler heals a recorded `online` that 2.6's derivation contradicts, writing `offline` — every cycle, **guarded on an actual change** (a heal that rewrote each tick would be a stream of identical writes and would destroy the record's mtime as a signal). **(b) The restart shape reaches a restoration path:** a once-per-daemon-generation boot sweep — and, for an owner that comes online after that sweep has run, an **owner offline→online edge** in the reconcile loop running the identical body (releases#228) — relaunches an instance when **all four** hold — the adapter section declares `persistent`; the…\n- **Regression shape (the row the suppressed fixture cannot express):** construct the restart shape **without a suspend edge** — an instance whose record says `online` over a corpse that predates boot, with an online owner — and assert it is restored. A fixture that suspends first re-creates the suppressed arm and proves nothing about this class.\n- **spt-core mapping:** `spt_daemon::shellwake::{heal_stale_online_records, restore_persistent_shells_at_boot, restore_persistent_shells_on_owner_online, OwnerOnlineEdge, launch_predates_boot, BOOT_RESTORE_SLACK_MS}`; `spt_store::proc::boot_instant_ms`; `spt_store::shellinfo::{ShellLaunch, record_shell_launch, read_shell_launch}` (the birth stamp, parked at the two production pid-write sites the launch-site census identified).\n- **Source:** spt-core, BAROMETER W2 (releases#78) — field specimen `liam/alchemy-0`, a genuinely stranded record: `online` over a dead pid, no launch stamp, days old.\n<!-- [doc->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS] -->\n\n---\n\n## 3. Lifecycle\n\n### 3.1 Ephemeral perch cleanup on every `ring` exit path\n- **Failure:** `ring` creates an ephemeral perch; early-exit paths (no-perch, empty-msg, timeout) skip cleanup → stale dirs accumulate.\n- **Invariant:** every code path that creates an ephemeral perch cleans it before exit; exception: if the caller already had an active perch, do not treat as ephemeral and do not clean up.\n- **spt-core mapping:** `ring` semantics carry; the daemon owns ephemeral-perch lifecycle, so a single guaranteed-cleanup (drop guard / RAII) is achievable in-process.\n- **Sister cite:** `src/owl/ring.rs:58-294`.\n\n### 3.2 Stale signoff sentinel must not kill a fresh start\n- **Failure:** a leftover `.claude/<id>-signoff.md` from a prior session is read by a fresh listener as a live signoff → immediate teardown.\n- **Invariant:** on every listener/daemon spawn, sweep stale signoff sentinels; signoff files are write-once per generation.\n- **spt-core mapping:** same sweep on daemon (re)start per hosted instance.\n- **Sister cite:** CHANGELOG v1.11.20; `src/owl/cleanup.rs:97`.\n\n### 3.3 Orphan teardown fires echo-commune BEFORE INIT_SIGNOFF\n- **Failure:** teardown delivers INIT_SIGNOFF without first saving the final context delta → Psyche signoff lacks the context-save summary.\n- **Invariant:** on orphan path, synchronously run the echo-commune (final delta) before composing INIT_SIGNOFF; skip only if the session_id is missing.\n- **spt-core mapping:** the daemon runs psyche/pulse loops in-process; ordering invariant identical.\n- **Sister cite:** `src/live/wrapper/orphan.rs:175-199`; tests A-H:333-565.\n\n### 3.4 A `ring` never adopts — so never deletes — a perch it did not create\n[…111ln elided…]\n- **Invariant:** every metadata-producing subprocess has a timeout; timeout yields `None` + rate-limited stderr, never a hang.\n- **spt-core mapping:** generalize to all manifest-declared harness invocations (delegated commands, adapter updates) — timeouts mandatory.\n- **Sister cite:** `src/common/git.rs`.\n\n### 5.4 Windows UNC prefix in serialized paths\n- **Failure:** canonicalized `\\\\?\\C:\\...` serializes to `//?/C:/...` and fails `read_to_string`.\n- **Invariant:** strip the `\\\\?\\` UNC prefix after backslash→forward-slash conversion; serialized path attrs must be directly consumable.\n- **spt-core mapping:** any path crossing the wire (file-drop EVENTs, off-node file transfer per ADR-0003) needs canonical normalization at the `spt-proto` boundary.\n- **Sister cite:** `src/common/owlery.rs:377-384`.\n\n### 5.5 ConPTY withholds output until DSR is answered  `[REQ-HAZARD-CONPTY-DSR]`\n- **Failure:** a broker reading a ConPTY master sees only the 4-byte startup query `ESC [ 6 n` and then nothing — the child looks hung/silent but is producing output normally. ConPTY blocks all child stdout until the terminal answers the cursor-position query.\n- **Invariant:** every ConPTY reader auto-answers DSR (`ESC [ 6 n` → write `ESC [ 1;1 R`, or a real cursor position) on the PTY writer. Secondary: a ConPTY master does not EOF while the writer is held, so read loops drain on a thread and never gate exit on a blocking `read()`.\n- **spt-core mapping:** `spt-term` broker PTY reader (ADR-0004). Brand-new to spt-core — not in the sister project (it never hosted ConPTY directly).\n- **Source:** Spike #1 (`docs/spikes/SPIKE-01-broker-handoff.md`); reproduced with both a Rust child and `cmd.exe`.\n\n<!-- [doc->REQ-HAZARD-DETACHED-PIPE-INHERIT] -->\n### 5.6 Windows detached children inherit a captured caller's pipe  `[REQ-HAZARD-DETACHED-PIPE-INHERIT]`\n- **Failure:** a caller captures an `spt` invocation's output through a pipe (`Command::output()`, a harness hook reading the command). That `spt` process detach-spawns a **long-lived** child (the daemon via `ensure_running`; a shell binary via `spt shell spawn`). On Windows `CreateProcess` runs with `bInheritHandles = TRUE`, and the spt process's std handles — the caller's pipe write-ends — are inheritable by construction, so the immortal child inherits them even when its *own* stdio is `Stdio::null()`. The caller's pipe read never sees EOF: the capturing caller **hangs forever** (unix is immune — pipe fds are `CLOEXEC`). Paid twice: daemon spawn (guarded at D4a-era `spawn_detached`), then again at M5-D3e when the mock-shell E2E hung `spt shell spawn` for h…\n- **Invariant:** every detach-spawn of a long-lived child inherits only handles it **enumerates**. The default and overwhelmingly common form of that is `bInheritHandles = FALSE` (`spt-daemon::daemon::detached_no_inherit`) — zero handles flow, whatever the pipe's depth in the ancestry. Stripping `HANDLE_FLAG_INHERIT` from the spawner's *std* handles is NOT sufficient: a grandparent capture's pipe sits in the handle table as a stray inheritable handle and still flows through every **bare** `bInheritHandles = TRUE` hop (the first guard shipped that way and was wedged by exactly this — a daemon spawned three layers deep held the pwsh-level pipe of the CI/test harness).\n- **The one permitted TRUE (amended 2026-07-26, RESIDENT-SERVICE W1):** a spawn that must capture the child's output may pass `bInheritHandles = TRUE` **only** paired with an explicit `STARTUPINFOEX` + `PROC_THREAD_ATTRIBUTE_HANDLE_LIST` allowlist — the Win32 contract honors the list only under TRUE, so the earlier phrasing (\"flipping TRUE is refused\") forbade its own mechanism and was unimplementable. The hazard was always *unenumerated* inheritance, and the list **is** the enumeration: only the named handles flow, so a grandparent's pipe cannot reach the child at any depth — strictly stronger than what FALSE achieves incidentally. A **bare** TRUE, with no attribute list, remains REFUSED. Today the only such path is a supervised `[service]` spawn with start…\n- **spt-core mapping:** `spt-daemon::daemon::spawn_detached` (the daemon) and `spt-daemon::shellhost::launch_shell` (the relay-receipt shell binary). Any future long-lived detached spawn (manifest-template children included) must use the same no-inherit spawn, or the enumerated form above where it genuinely must capture output.\n- **Source:** spt-core, M5-D3e (`shell_e2e.rs` hang, 2026-06-04, twice — once per guard generation); Rust `Command` restricts *its own* created stdio handles but a parent's inheritable handle table still flows.\n\n### 5.7 Elevated commands spawn the daemon with the wrong token  `[REQ-HAZARD-ELEVATED-DAEMON-SPAWN]`\n<!-- [doc->REQ-HAZARD-ELEVATED-DAEMON-SPAWN] -->\n- **Failure:** membership-implies-reachability made *every* `spt` invocation a potential daemon spawner (`ensure_running`), including the elevation-gated ones (`subnet create`/`join`, REQ-SUBNET-4). The spawned daemon inherits the spawner's token. **Windows:** an elevated `subnet create` auto-starts an ELEVATED daemon whose named pipes deny unelevated clients — every subsequent unelevated `spt` reads \"not running\", tries to spawn its own daemon, and dies on bind Access-denied; the user had to taskkill (hit live, M7 acceptance 2026-06-06). **Linux:** a sudo'd command spawns a root daemon and/or root-owned state — and because sudo flips `$HOME`, the daemon can mint a *different node identity* in root's universe (the very key-flip that produced the 4.10 ghost r…\n- **Invariant:** the daemon **always runs unelevated in the invoking user's universe**, regardless of which command spawns it. Two enforcement points sharing one seam: (a) `spawn_detached` de-elevates the child — Windows: the UAC **linked token** (`TokenLinkedToken` → `DuplicateTokenEx` → `CreateProcessWithTokenW`; inherits no handles, so 5.6 holds by construction); Linux: drop to `SUDO_UID`/`SUDO_GID` with `$HOME`/`$USER`/`$LOGNAME` reset to the invoking user's (passwd lookup); (b) a `Daemon::run` entry guard catches a *directly* elevated `spt daemon` — Linux drops privileges in-process before touching any state; Windows respawns de-elevated and exits. When no unelevated identity exists to drop to (UAC disabled, genuine root login, SYSTEM), the daemon runs …\n- **spt-core mapping:** `spt-daemon::deelevate` (the OS-split seam) consumed by `daemon::spawn_detached` + the `Daemon::run` entry guard. The fuller Linux elevation model (install symlink + default-account election) is deferred (DEFERRED.md, M8).\n- **Source:** M7 acceptance run 2026-06-06 (DEFERRED.md \"Non-admin daemon spawn\"); interim field rule was \"bring the daemon up unelevated FIRST\".\n\n<!-- [doc->REQ-HAZARD-CHILD-CONSOLE-FLASH] -->\n### 5.8 Console children of the console-less daemon flash visible windows  `[REQ-HAZARD-CHILD-CONSOLE-FLASH]`\n- **Failure:** the daemon runs DETACHED (no console, 5.6/`detached_no_inherit`). Any console-subsystem child it spawns (`git`, `taskkill`, manifest hook commands) gets a **fresh conhost with a visible window** — piped/null stdio does NOT prevent it. Field shape: the 60s sync pump's two git calls (`for-each-ref` + `rev-parse`) flashed two blank windows per minute on the user's desktop (2026-06-06).\n- **Invariant:** every short-lived console child spawned from daemon-reachable code sets `creation_flags(0x0800_0000)` (`CREATE_NO_WINDOW`). Long-lived detached children use `detached_no_inherit` (already `DETACHED_PROCESS | CREATE_NO_WINDOW`); de-elevated spawns use `CREATE_NEW_CONSOLE + SW_HIDE` (5.7 — `CreateProcessWithTokenW` rejects `CREATE_NO_WINDOW`, error 87).\n- **Test seam caveat:** window-absence is unobservable from a consoled test runner — the child inherits the runner's console and never creates a window, flag or no flag. Unit coverage asserts the flagged spawn still works (the error-87 \"flag combo breaks spawn\" regression class); window-absence was verified live by process-watch capture.\n- **spt-core mapping:** `spt-store::gitrun::run_git` (every BranchStore/ContextStore git call), `spt-daemon::shellhost::kill_shell_pid` (taskkill), `spt-runtime::run_bounded_command` (manifest hook commands), `spt-runtime::ManifestRuntime::command_for` (the one shared builder behind `spawn_session` + `run_bounded_stdin` — the notif pump's `spawn_notif_command` and the live agent's psyche/echo/turn spawns), `spt-daemon::shellwake` (already guarded). The flag lives in each shared builder, not per call site, so the invariant holds for every ManifestRuntime spawn by construction.\n- **Source:** spt-core field bug, 2026-06-06 — two blank windows flashing every 60 seconds on a desktop workstation, caught by process-spawn watcher (git.exe parent=spt daemon, conhost.exe child each).\n\n### 5.9 `Instant - Duration` underflow-panics on a freshly-booted host  `[REQ-HAZARD-INSTANT-UNDERFLOW]`\n- **Failure:** `Instant::now() - Duration::from_secs(N)` panics `overflow when subtracting duration from instant` when the process's monotonic clock is younger than `N` — i.e. the host booted less than `N` ago. The peer pump primed its cadence legs with `Instant::now() - 86_400s` to mean \"everything due now\"; on a Windows runner with sub-24h uptime the pump thread panicked at startup, so the subnet never converged (CI `pump_and_dispatch_self_drive_the_subnet` failed, run 27082417706). It is *environment-conditional* — green on any host up longer than the offset, red below it — so it slips local dev and only bites a fresh CI box or a just-rebooted machine.\n- **Invariant:** NEVER compute an instant in the past by subtracting from `Instant::now()`. Represent \"never run / due now\" as `Option<Instant> = None` and gate on forward `now.duration_since(past)` only (`peerloop::due`). No backward instant arithmetic anywhere in scheduling.\n- **Test seam caveat:** the convergence E2E only reproduces on a sub-offset-uptime host (it passed everywhere with >24h uptime). The deterministic guard is the `due(None, ..)`/`due(Some(now), ..)` unit on the extracted gate — it asserts first-tick-due with zero instant subtraction, independent of host uptime.\n- **spt-core mapping:** `spt-daemon::peerloop::due` (the sole cadence gate behind `due_reg`/`due_notif`/`due_sync`/`due_upd`); cadence legs are `Option<Instant>` seeded `None`.\n- **Source:** spt-core CI failure, 2026-06-07 — Windows runner `hfenduleam` (just booted) panicked the peer pump at the v0.1.1 release gate.\n\n### 5.10 `sudo spt` dead-ends on a user-local install (secure_path)  `[REQ-HAZARD-SUDO-SECURE-PATH]`\n- **Failure:** the elevation-gated commands (`subnet create` / `subnet join` / `show-code`) refuse when unelevated and tell the user to \"run as administrator / root\". The user types the obvious `sudo spt subnet create FOO` → `sudo: spt: command not found`. `spt` is a user-local install (`~/.local/bin`, `~/.cargo/bin`), and sudo's `secure_path` (a `/etc/sudoers` default) does NOT include those dirs, so a bare command name doesn't resolve under sudo. The guidance is a trap: it names an action that cannot work for the common install shape. Field-hit on KITSUBITO at the v0.1.1 ship.\n- **Invariant:** elevation guidance on Unix emits the binary's **absolute path** under sudo — `sudo /home/u/.local/bin/spt subnet create FOO` — reconstructed from `current_exe()` + the real argv and shell-quoted. An absolute program path is executed directly; `secure_path` only governs bare-name PATH lookup, so the absolute form always resolves. On an interactive Unix TTY the command auto-elevates (re-execs itself under sudo, the elevated child does the work and `main` de-elevates back); non-interactive or sudo-absent falls back to printing the runnable hint. Never emit a bare-name elevation instruction.\n- **Companion UX:** the post-de-elevation `DEELEVATED: running as uid N` line is internal state-safety noise — omit it from the user-facing CLI path (it confused the same field user). The detached daemon's own de-elevation log line is fine (it lands in the daemon log, not the terminal).\n- **Test seam caveat:** the sudo re-exec needs a real `sudo` + TTY (not hermetic). The deterministic guard is the pure `elevation::sudo_argv` / `print_hint_command` (assert an absolute exe path, never a bare name, + shell-quoting on the printed line) and the `decide_elevation_path` matrix (which picks inline-sudo only on an interactive Unix TTY); the exec leg is manual/kitsubito-verified.\n- **spt-core mapping:** `spt::elevation::{sudo_argv, print_hint_command, decide_elevation_path}` (pure — generalized from the M12-W4 self-elevation seam, 5.11), `spt::cli::{try_auto_elevate, with_elevation_hint}` wired into `cmd_subnet_create` / `cmd_subnet_join` / `cmd_subnet_show_code`; `spt::main` de-elevation drop silenced.\n- **Source:** spt-core field report, 2026-06-07 — `reavus@KITSUBITO`, `spt` in `~/.local/bin`; the absolute-path `sudo` invocation was confirmed working before the fix landed.\n\n### 5.11 Self-elevating re-launch must re-run verbatim, never widen / inject / loop  `[REQ-HAZARD-SELF-ELEVATE]`\n- **Failure class:** a privilege-gated command (`subnet create` / `join` / `show-code`) self-elevates by re-launching itself with privilege (Windows UAC `runas`, Linux `pkexec` / a terminal-emulator `sudo`, or inline `sudo`). A careless re-launch is a security hole: widening the privilege scope (adding args), resolving the binary by a bare name (a PATH/`secure_path` hijack runs an attacker's `spt`), interpolating a crafted arg into a shell string (`sh -c \"… $id …\"` injects a second command), or re-elevating the already-elevated child (an infinite UAC/polkit loop). The user's UAC/polkit/sudo prompt is the ONLY consent gate — the mechanism must never bypass or widen it.\n- **Invariant:** self-elevation re-runs the **EXACT** original invocation with the binary's **ABSOLUTE** exe path — never adding/altering args, never a PATH-resolved bare name, never a shell-interpolated string. Every launcher passes an **argv array** (`Command::new(prog).args([...])`, never `sh -c`); the Windows `ShellExecuteW` params string (which is inherently one string) MSVC-quotes each verbatim arg so `CommandLineToArgvW` round-trips it as a single token. The elevated child drops state back to the user (composes with the 5.7 de-elevation) and **never re-elevates**: `decide_elevation_path` returns `AlreadyElevated` whenever the process is `Elevated`, on every OS (loop-safety). The unprivileged parent never pipes/captures the elevated child's stdout acro…\n- **Test seam caveat:** the real launch needs a UAC/polkit/sudo prompt (not hermetic) — manual-verify. The deterministic guards are the pure `decide_elevation_path` matrix (loop-safety: `AlreadyElevated` on every os; the os×env path order) and the argv builders (`sudo_argv` / `pkexec_argv` / `terminal_argv` assert absolute-exe + verbatim args + array; `windows_runas_params` asserts MSVC-quoting with no `cmd /c` interpolation; the crafted-arg test asserts a shell-metachar arg stays one element / one quoted token).\n- **spt-core mapping:** `spt::elevation::{decide_elevation_path, sudo_argv, pkexec_argv, terminal_argv, windows_runas_params, print_hint_command, ElevatePath}` (pure), `spt::cli::{try_auto_elevate, launch_uac_window, pause_elevated_console_if_fresh, program_on_path, first_terminal_emulator}` (impure launchers) wired into `cmd_subnet_create` / `cmd_subnet_join` / `cmd_subnet_show_code`. Companions: 5.10 (the Unix abs-path-under-sudo facet) and 5.7 (the elevated child's de-elevation drop, which this composes with).\n- **Source:** M12-W4 design (subnet QR + self-elevating window), doyle ruling `M12-W4-RULING.md` Q6 — a privilege-escalation feature carries a mandatory hazard REQ.\n\n<!-- [doc->REQ-HAZARD-WIN-PTY-PROGRAM-RESOLVE] -->\n### 5.12 Native-PTY spawn of a bare program runs the wrong (non-PE) file on Windows  `[REQ-HAZARD-WIN-PTY-PROGRAM-RESOLVE]`\n- **Failure:** `portable-pty`'s ConPTY spawn resolves a bare program name with a `which` that takes the FIRST `PATH` match. A node/npm CLI installs as BOTH an extensionless shebang shim (`ccs`, for Git Bash) and a Windows launcher (`ccs.cmd`) in the same dir; portable-pty picks the extensionless `ccs`, and `CreateProcessW` then tries to execute that non-PE file and fails with **os error 193** (\"%1 is not a valid Win32 application\"). Live failure: `spt endpoint run claude-spt:ccs` → `CreateProcessW C:\\nvm4w\\nodejs\\ccs` 193 (operator, 2026-06-16). The same bites any harness/shell whose `[session.self]`/`[shell].spawn` names a `.cmd`/`.bat`/`.ps1`-backed command — `CreateProcessW` cannot execute a batch or PowerShell script directly.\n- **Invariant:** spt-term resolves the program ITSELF before handing it to `CommandBuilder`, bypassing portable-pty's `which`. A bare name is searched over `PATH` × `PATHEXT` (whose default order already prefers `.EXE`/`.COM` over `.BAT`/`.CMD`), then an extensionless fallback. A non-PE target is wrapped in its interpreter: `.cmd`/`.bat` → `cmd.exe /d /c <path>`, `.ps1` → `powershell -NoProfile -File <path>` (the wrap args precede the caller's args); a real executable spawns directly; an unresolvable name passes through unchanged (never makes a working case worse). Unix is a passthrough — `execve` honours a shebang on an extensionless script. Applied at the ONE `CommandBuilder` chokepoint (`PtySession::spawn_program_in`), so every broker harness + shell spaw…\n- **spt-core mapping:** `spt_term::winprog::{resolve_for_pty, resolve_in}` (the pure PATHEXT-precedence kernel + the Windows env wiring), wired into `spt_term::pty::PtySession::spawn_program_in`. Unit: `resolve_in` precedence (`.cmd`-over-shim, `.exe`-direct, explicit-extension, path-order, passthrough) [`winprog.rs`].\n- **Source:** field diagnosis 2026-06-16 (operator dogfood, `claude-spt:ccs` bringup) — doyle.\n\n<!-- [doc->REQ-HAZARD-PERCH-RECORD-POWER-LOSS] -->\n### 5.13 Atomic write leaves data un-synced before the rename → NUL zero-fill on power loss  `[REQ-HAZARD-PERCH-RECORD-POWER-LOSS]`\n- **Failure:** `atomic_write_bytes` was `fs::write(tmp)` + `rename(tmp, path)` with no `fsync`. The rename's directory **metadata** is journaled durable, but the tmp file's **data blocks** are still in the page cache. A hard reset (power loss, forced reboot) between the two flushes lands the rename but loses the data → the file reappears at its **full length filled with NUL**. Field incident: after a machine restart `owlery/hall-a/info.json` was 360 bytes of all-NUL (the nested psyche record 125 bytes all-NUL); `Get-Content` renders blank, `Format-Hex` shows the zero-fill. A wiped-but-present record then read as a live ONLINE endpoint for days (compounded by 5.14).\n- **Invariant:** for records whose loss is **unrecoverable**, flush data to stable storage **before** the rename publishes the name — `File::create(tmp)` → `write_all` → `sync_all()` → `rename_with_retry`. Then a crash yields the complete old file OR the complete new file, never a NUL husk. (A file of correct length that is ALL-NUL is the diagnostic tell of this class — a non-fsync'd write caught by a hard reset.) **Durability is SCOPED, not blanket.** `fsync`ing *every* atomic write serializes a durable flush behind each of the ~26 `atomic_write` callers — and for `info.json` the flush lands under the per-perch `.info.lock` (the **W1b lock-across-fsync wedge shape**, §6-era) — which stalls daemon bringup 4–6× so the endpoint misses its ONLINE window. Scope …\n- **spt-core mapping:** `spt_store::atomic::atomic_write_bytes_durable` / `atomic_write_string_durable` are the opt-in durable siblings; the default `atomic_write_bytes` / `atomic_write_string` stay non-durable. Durable callers are exactly `spt_store::info::write_info` (the perch record — the hall-a surface), `spt_store::nodeid` (the node seed — corrupt is node-bricking, never regenerated), and `spt_daemon::machineid` (minted-once machine id — a NUL husk would silently re-mint a different id). Reader-side handling of an already-corrupt record is the sibling 5.14. **Canary:** `attach_wedge_e2e` is the regression guard — a blanket fsync fails it (bringup misses the 20s ONLINE budget); keep durability scoped so it stays green.\n…\n\n<!-- [doc->REQ-HAZARD-CORRUPT-PERCH-COHERENCE] -->\n…\n\n…\n<!-- [doc->REQ-HAZARD-EMPOWER-SESSION-RESURRECT] -->\n\n[Showing lines 1-116 and 228-307 of 307; 111 middle lines (20.7KB) elided. Use :301 to continue. Read artifact://64 for full output. Some lines truncated to 768 chars]"}],"details":{"truncation":{"content":"# Known Hazards\r\n\r\nHard-won edge cases harvested from the sister project (`claude_skill_owl`, ~80 commits / 12+ phases / multiple production incidents). Per ADR-0001, this is a **test checklist for the spt-core rebuild** — the clean-room rebuild must re-satisfy each invariant rather than re-discover the bug.\r\n\r\n**Architecture-translation note.** The sister project runs poll listeners and Psyche wrappers as *separate processes*. spt-core consolidates both into the one `spt-daemon` (brain), with a stable broker beneath it (ADR-0004). Many hazards below were inter-process races in the sister project; in spt-core some become intra-daemon concerns (potentially easier) while others move to the daemon↔broker IPC boundary or the network boundary (potentially new failure surface). Each entry notes the mapping where it differs. Citations point at sister-project paths for reference, not at spt-core.\r\n\r\n---\r\n\r\n## 1. Race conditions & ordering\r\n\r\n### 1.1 Phantom INIT_SIGNOFF after grace period\r\n- **Failure:** orphan teardown enqueues INIT_SIGNOFF before the grace-period recheck; a transient Self recovery (binary handoff, brief stale poll) makes the recheck pass-as-alive, but the signoff was already spooled and drains on the next iteration → teardown despite a live Self.\r\n- **Invariant:** grace-period wait MUST complete *before* composing/delivering INIT_SIGNOFF; the recheck must bind `still_gone` before any envelope write.\r\n- **spt-core mapping:** in-daemon now (no separate wrapper), but the ordering invariant is identical — orphan/teardown logic must re-evaluate liveness after the grace wait, not before enqueue.\r\n- **Sister cite:** `src/live/wrapper/orphan.rs:201-259` (sleep@209 precedes compose@231-251); tests T-grace-recovery:576, T-still-gone-recheck:618.\r\n\r\n### 1.2 Poll-rewrite race & info.json mid-write reads\r\n- **Failure:** `info.json` written by the wrapper mid-iteration while a list/classify command reads it → torn read, misclassification.\r\n- **Invariant:** consult liveness via the supervisor (`is_wrapper_alive`-equivalent) before any grace gate; reads of state files must tolerate concurrent writes (atomic write + rename, or read-retry).\r\n- **spt-core mapping:** the daemon owns both writer and reader → use in-process locking/snapshotting instead of racing on disk. Cross-node registry reads remain eventually-consistent and must tolerate staleness.\r\n- **Sister cite:** `src/common/list_filter.rs:100-150`; `src/owl/poll.rs:141`.\r\n\r\n### 1.3 Stale `index.lock` wedge from prior git crash\r\n- **Failure:** crashed git leaves a 0-byte `index.lock` in a psyche tracked worktree; every later commit blocks forever.\r\n- **Invariant:** on daemon boot, sweep seed + all agent/project worktrees for stale locks (0 bytes, mtime > 60s) and remove; leave live locks alone.\r\n- **spt-core mapping:** cross-node Psyche sync (ADR-0002/0003) replaces git-repo sync, so the *git* lock may disappear — but any equivalent lockfile in the new sync mechanism needs the same stale-sweep on boot.\r\n- **Sister cite:** CHANGELOG v1.11.20 \"Stale `index.lock`\"; `src/common/git.rs`.\r\n\r\n### 1.4 Deferred spool rows must not leak to the event stream\r\n- **Failure:** a hook spools a deferred (spool-only, no TCP wake) notice; startup `drain_all` flushes ALL rows including deferred → event emitted at wrong time/priority.\r\n- **Invariant:** startup drain (and idle/timeout TCP-wake sites) use `drain_non_deferred` only; deferred rows are picked up by their intended consumer via `peek`. All drain sites must agree on which rows they flush.\r\n- **spt-core mapping:** carries directly — the daemon's spool-drain has the same deferred-vs-immediate distinction.\r\n- **Sister cite:** `src/owl/poll.rs:276-316`; `spool::drain_non_deferred_with_metadata`.\r\n\r\n### 1.5 Worker (working-perch) lifecycle path consistency\r\n- **Failure:** subagent-start creates the perch at one path layout; later hooks read it at another → not found; stop-hook scan misses nested perches.\r\n- **Invariant:** all Worker/Psyche child-perch path composition routes through one central resolver; no divergent path construction across hooks.\r\n- **spt-core mapping:** `Worker` is a day-one endpoint type; the daemon owns the registry, so perch location is a registry lookup, not ad-hoc path math. Single source of truth for instance→location.\r\n- **Sister cite:** `src/owl/hook_subagent_start.rs:122-168`; `hook_subagent_stop.rs:15-55`.\r\n\r\n---\r\n\r\n## 2. Identity & session-binding\r\n\r\n### 2.1 Parent PID over ephemeral poll PID\r\n- **Failure:** orphan check polls an ephemeral listener PID; it dies and is recycled (esp. Windows); a foreign process with the recycled PID reads as alive → false-positive teardown (or false-negative).\r\n- **Invariant:** prefer the stable harness-session PID (`parent_pid`) over any ephemeral process PID for liveness; minimal `info.json` for supervisor-owned perches to avoid stale leaks.\r\n- **spt-core mapping:** session binding (parent-process-tree anchor) still applies for harness-hosted topology. For spt-hosted sessions the broker holds the child directly → liveness is the broker's held-handle state, more reliable than PID polling.\r\n- **Sister cite:** `src/live/wrapper/orphan.rs:141-161`; CHANGELOG v1.11.20.\r\n\r\n### 2.2 Stdin session_id precedence over env\r\n- **Failure:** subagent inherits a stale `OWL_SESSION_ID` env across `/clear`; hook gets two session_ids (fresh stdin, stale env) → wrong-agent binding.\r\n- **Invariant:** stdin-provided session_id wins; env is fallback only.\r\n- **spt-core mapping:** the harness-contract subcommand surface must define the same precedence for whatever identity fields hooks pass in.\r\n- **Sister cite:** CHANGELOG v1.35.1 \"IN-05\"; `hook_subagent_start.rs:40-51`.\r\n\r\n### 2.3 Binary-handoff argv schema must stay backward-compatible\r\n- **Failure:** old binary spawns new binary with old argv arity; clap rejects before state rehydration → wrapper dies unlogged.\r\n- **Invariant:** every newly-added handoff positional has a default; state-file rehydration happens *after* argv parse; defaults survive intermediate versions.\r\n- **spt-core mapping:** CRITICAL — self-update (ADR-0004) makes handoff routine. The broker↔brain IPC and any brain-relaunch argv must be versioned and forward/backward tolerant (a newer brain talks to an older broker). This is the single most update-frequency-sensitive invariant.\r\n- **Sister cite:** `src/live/wrapper/lifecycle.rs:17-106`; `src/cli.rs` defaults; CHANGELOG v1.11.10.\r\n\r\n### 2.4 Generation `gen_start` always = now() on cold-start AND handoff\r\n- **Failure:** stale gen_start from a rehydrated state file fires time-based discriminators on the new process.\r\n- **Invariant:** wall-clock `gen_start` is set to `now()` on both cold-start and handoff; generation counter increments on every start/revive; session UUID captured fresh and carried so the resumed mind distinguishes \"same gen continuing\" vs \"new gen born\".\r\n- **spt-core mapping:** carries to the daemon's per-instance generation tracking.\r\n- **Restoration D3/D4 (ADR-0018):** the generation *counter* custody moved to the broker (D3-2 — it observes every brain spawn, planned or crash, and hands `{generation, start-reason}` at spawn; `gen_start` stays `now()`-fresh, never rehydrated). The brain→brain **`BrainState` *message*** (`{session_id, generation, next_seq, gen_start_ms}`) that previously carried continuity across a handoff is **retired from the production path in D4-2**: a brain the supervisor respawns cold-starts and reconstructs all session continuity by **querying the broker** (`Brain::resume_sessions` over the broker's cursor-of-record), never a frame. `BrainState` / `Brain::handoff` / `Brain::snapshot` remain `pub` and compiled **only for the integration tests** (handoff/idempotent/da…\n- **Sister cite:** `src/live/wrapper/lifecycle.rs:70`; `src/common/wrapper_state.rs`.\r\n\r\n### 2.5 Daemon-hosted endpoints have no dedicated liveness PID\r\n- **Failure:** the sister evaluates Psyche/perch liveness via a dedicated process PID — the wrapper's own pid in `info.json`, checked with `is_process_alive`. Under ADR-0004 the Psyche (and any spt-hosted Self) is a **loop inside the daemon**, not a separate process: it holds no dedicated pid, and its `claude`/summarizer subprocess is ephemeral (spawned per pulse/commune, then exits). If a daemon-hosted perch's `info.json` carries the **daemon's** pid, then *every* hosted endpoint shares one pid, and `is_process_alive(pid)` reads \"alive\" for a torn-down endpoint as long as the daemon runs — while `clean_stale_entries` (dead-pid deletion) can no longer distinguish a dead endpoint from a live one. The 2.1/5.1 liveness models do **not** cover this third categor…\n- **Invariant:** for **daemon-hosted** perches (Psyche; spt-hosted Self), liveness is the **daemon's authoritative in-memory endpoint table + a `status` field** on `info.json` (`online|offline|…`), **never** `is_process_alive(info.pid)`. `info.pid` for a daemon-hosted perch is at most a *hosted-by-daemon* marker (the daemon pid), not a liveness signal; registry stale-clean for these rows keys on the daemon's endpoint table, not per-row pid. This reuses the pattern already specified for **Shells** (`info.json` carries daemon-managed `status`, capability resolved by `adapter_name` — CONTEXT \"Shell… Not in the subnet registry\") and extends it to daemon-hosted *agent* perches.\r\n- **spt-core mapping:** the **M1/M2a interim** model keeps the Psyche/listener a real per-process owner (the `api listen` process), so its per-pid liveness (`deliver::is_online` → `info.read_pid` → `proc::is_process_alive`; `registry::clean_stale_entries`) is correct *interim*. **M3 daemon consolidation replaces it** with daemon-authoritative liveness for hosted perches. Keep the liveness check behind one resolver (mirrors `resolve_address` stale-clean) so the M3 swap is localized — do **not** let the per-pid assumption leak into new call sites.\r\n- **Sister cite:** `src/live/wrapper/orphan.rs` (wrapper-pid liveness); `src/common/list_filter.rs:168-175` (pid-classify); spt-core `crates/spt-store/src/{proc.rs,registry.rs}` + `crates/spt-msg/src/deliver.rs::is_online`.\r\n\r\n### 2.6 A shell instance's `online` is a recorded field with no death edge\r\n- **Failure:** a shell perch's `status` has exactly two writers — the bind handshake (→`online`) and `close_shell` (→`offline`) — and `close_shell` runs on the **link-break** path alone. A binary that dies **abruptly** (force-kill, crash, OOM) breaks no link, so nothing ever writes the flip: the record says `online` forever while the pid it names is gone. Every consumer of the recorded field then inherits the lie — `shell list` reports online indefinitely; `relink` (the one command that recovers the instance) refuses `SHELL_ALREADY_ONLINE`, gated on the very state that is wrong; and `shell cmd`'s wake-if-offline arm reads \"not offline\", skips the wake, and spools the command to a corpse — accepted happily, drained by nobody, **indistinguishable from a busy s…\n- **Invariant:** ONLINE-ness is **derived**, never the recorded field alone: `status == online` **AND** the recorded `shell.pid` is not *provably* dead. Provably dead is narrow — pid parked, non-zero, `!is_process_alive`; pid absent, unparseable, or `0` (a broker-hosted spawn whose backend exposed no pid) reads **alive**, the same fail-toward-alive parity 2.5 holds, so a backend the resolver cannot see is never declared dead. **The recycled-pid caveat is narrowed since BAROMETER W2 and is now conditional on the stamp:** a launch parks a **birth stamp** beside the pid (`shell.launch.json`), and the probe runs the PAIR test — same pid, *different* start time ⇒ **gone**, so a pid the OS handed to an unrelated process no longer masks a dead instance. That case m…\n- **No spontaneous relaunch of a SAME-BOOT death (operator-ratified, flynn 2026-07-25; re-based BAROMETER W2):** the reconciler does NOT adopt an instance whose binary died during this boot. Every reason an operator stops a shell is a reason not to want it back a tick later: mid-deploy (a shared install dir makes \"kill the process\" routine on Windows, since a running exe cannot be overwritten) an auto-relaunch would run the **old** binary out of the file being replaced — worse than a failed install — and it turns quarantining a misbehaving shell into a restart loop. For those, recovery stays demand-driven: `relink`, or a `shell cmd` that wakes. `relink` probes **locally** rather than trusting a daemon sweep, so recovery holds with the daemon down.\r\n- **What holds that open is no longer the stale record (BAROMETER W2):** until W2 the guarantee rested on an accident: a force-killed instance was held out of the eligible set because its record still said `online`, and the eligibility read took that field raw. 2.7 heals that record to the truth, which removes the accidental protection — so the guard had to move to a rule that *states* it. Eligibility is now recorded-`offline` **AND** (no corpse, **or** the corpse's recorded launch **predates the boot instant**). A steady-state force-kill leaves a corpse launched *after* boot, which the predicate can never accept, so the ruling is preserved **by construction** rather than by a second rule. A restart casualty is the one case the predicate does accept, and it …\n- **spt-core mapping:** `spt_store::shellinfo::{shell_pid_provably_dead, effective_status, is_shell_online}` (the resolver — the shell-side twin of 2.5's `liveness.rs`, which gave *agent* perches exactly this and which shells never got); consumers in `linkhost::{relink_shell, prepare_drive, run_action}`, `shelldisc::discover`, `activity::observe_links`; the class-(c) eligibility read is `spt_daemon::shellwake::watcher_eligible` (+ `launch_predates_boot`, `BOOT_RESTORE_SLACK_MS`), and the birth stamp it rests on is `shellinfo::{ShellLaunch, record_shell_launch, read_shell_launch}`.\r\n- **Source:** spt-core, flynn's spt-alchemy field report (2026-07-25) — a deterministic recipe, not a race. Class-(c) re-based BAROMETER W2 (releases#78, doyle's ruling on comment 5156810267).\r\n\r\n### 2.7 A node restart permanently strands every `persistent` shell — and the existing cascade test cannot see that it does  `[REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]`\r\n- **Failure:** the contract's own sentence says a `persistent` shell is **online whenever its owner endpoint is online**, and a node restart is the one case where that promise was never kept. A machine death breaks no link, so `close_shell` never runs and 2.6's stale record survives the reboot; when the owner comes back online, the wake cascade reads the **recorded** field, sees `online`, and skips the relaunch of a binary that has not existed since the previous boot. Nothing else corrects it: the instance is down forever, with no error anywhere, until a human happens to run `relink`. Every `persistent` instance on the node is stranded by one restart, together.\r\n- **The owner-facing surfaces HIDE the fault**, which is why it survived unreported for days: every display path (`shell list`, its `--json` twin, the shell-context render) routes through the ONE discovery seam, which **derives** status (2.6) and so correctly reports the binary as offline — while the only consumer whose decision matters, the wake cascade, reads the RECORDED field that no owner-facing surface shows. Absence of a visible symptom is therefore **not** evidence of absence of this fault, and a clean-looking context render must never be taken as proof the class did not occur. The correct falsifier is the **on-disk record**.\r\n- **The existing cascade test cannot fail on it — its SETUP SUPPRESSES THE FAILING ARM.** `rest_edges_cascade_shells_with_divergence` establishes its fixture by driving the suspend path, which itself writes the `offline` the code under test is supposed to encounter — so the restart shape (a record still saying `online` over a corpse, reached with **no** suspend edge) is unreachable from that fixture BY CONSTRUCTION. The test is not weak; it gives a true answer to the wrong question, so a reader asking \"is the cascade covered?\" is told yes. This is the third suppressed-arm instance found in one week, and the pattern is the same each time: **a fixture that establishes the precondition the code under test is supposed to establish.**\r\n- **Invariant:** a node restart must not strand a `persistent` instance. Two legs, and each is worthless alone. **(a) The record stops lying:** the daemon-side reconciler heals a recorded `online` that 2.6's derivation contradicts, writing `offline` — every cycle, **guarded on an actual change** (a heal that rewrote each tick would be a stream of identical writes and would destroy the record's mtime as a signal). **(b) The restart shape reaches a restoration path:** a once-per-daemon-generation boot sweep — and, for an owner that comes online after that sweep has run, an **owner offline→online edge** in the reconcile loop running the identical body (releases#228) — relaunches an instance when **all four** hold — the adapter section declares `persistent`; the…\n- **Regression shape (the row the suppressed fixture cannot express):** construct the restart shape **without a suspend edge** — an instance whose record says `online` over a corpse that predates boot, with an online owner — and assert it is restored. A fixture that suspends first re-creates the suppressed arm and proves nothing about this class.\r\n- **spt-core mapping:** `spt_daemon::shellwake::{heal_stale_online_records, restore_persistent_shells_at_boot, restore_persistent_shells_on_owner_online, OwnerOnlineEdge, launch_predates_boot, BOOT_RESTORE_SLACK_MS}`; `spt_store::proc::boot_instant_ms`; `spt_store::shellinfo::{ShellLaunch, record_shell_launch, read_shell_launch}` (the birth stamp, parked at the two production pid-write sites the launch-site census identified).\r\n- **Source:** spt-core, BAROMETER W2 (releases#78) — field specimen `liam/alchemy-0`, a genuinely stranded record: `online` over a dead pid, no launch stamp, days old.\r\n<!-- [doc->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS] -->\r\n\r\n---\r\n\r\n## 3. Lifecycle\r\n\r\n### 3.1 Ephemeral perch cleanup on every `ring` exit path\r\n- **Failure:** `ring` creates an ephemeral perch; early-exit paths (no-perch, empty-msg, timeout) skip cleanup → stale dirs accumulate.\r\n- **Invariant:** every code path that creates an ephemeral perch cleans it before exit; exception: if the caller already had an active perch, do not treat as ephemeral and do not clean up.\r\n- **spt-core mapping:** `ring` semantics carry; the daemon owns ephemeral-perch lifecycle, so a single guaranteed-cleanup (drop guard / RAII) is achievable in-process.\r\n- **Sister cite:** `src/owl/ring.rs:58-294`.\r\n\r\n### 3.2 Stale signoff sentinel must not kill a fresh start\r\n- **Failure:** a leftover `.claude/<id>-signoff.md` from a prior session is read by a fresh listener as a live signoff → immediate teardown.\r\n- **Invariant:** on every listener/daemon spawn, sweep stale signoff sentinels; signoff files are write-once per generation.\r\n- **spt-core mapping:** same sweep on daemon (re)start per hosted instance.\r\n- **Sister cite:** CHANGELOG v1.11.20; `src/owl/cleanup.rs:97`.\r\n\r\n### 3.3 Orphan teardown fires echo-commune BEFORE INIT_SIGNOFF\r\n- **Failure:** teardown delivers INIT_SIGNOFF without first saving the final context delta → Psyche signoff lacks the context-save summary.\r\n- **Invariant:** on orphan path, synchronously run the echo-commune (final delta) before composing INIT_SIGNOFF; skip only if the session_id is missing.\r\n- **spt-core mapping:** the daemon runs psyche/pulse loops in-process; ordering invariant identical.\r\n- **Sister cite:** `src/live/wrapper/orphan.rs:175-199`; tests A-H:333-565.\r\n\r\n### 3.4 A `ring` never adopts — so never deletes — a perch it did not create\r\n- **Failure:** 3.1 guards the LEAK direction (cleanup on every exit path). It says nothing about *whose* perch is being cleaned, and the two compose into a data-loss bug: `ring` decides \"the caller has no perch\" **ready-marker-first**, so an existing perch whose marker is momentarily down — busy turn, soft session-end, stale re-bind — reads *perchless*, `setup_ephemeral` clobbers the caller's real `info.json`, and 3.1's guaranteed cleanup then runs on the way out and **deletes** the ready marker, `info.json`, `spool.db` and the directory. Silent CLI-side fs ops, no daemon log, no console trace (field 2026-07-27 — emphasys lost a live endpoint and its spool to one timed-out ring).\r\n- **The mail half is a CONFIDENTIALITY failure, not only a durability one:** deletion is what happens on the way out, but while the ring *holds* the adopted perch its reply-wait `drain_one_at` consumes whatever is in the victim's spool and **renders it to the ringer as the reply**. Ringer == victim (the field case) is the degenerate, mild shape — anything drained reaches its intended recipient. Ringer != victim is the severe one: a third party rings a live agent and the **victim's inbox is drained into the ringer's output**, leaving no trace on either side once the dir is removed. Across a multi-machine subnet that is one agent reading another's mail. Bounded honestly for the record: in the field instance theft did **not** fire — emphasys's 66-second window …\n- **Self-camouflaging**, which is why it went unattributed for hours: the damage — perch gone, roster row ghosted — *mimics the stale-liveness condition people reach for `ring` to diagnose*. The field caller was probing a ghost-roster symptom and the probe manufactured a fresh one, so debuggers of stale liveness are disproportionately its victims and read the wreckage as more of the symptom they were chasing.\r\n- **Marker-down is not an edge case:** on hosted OMP endpoints the extension owns the listener and publishes busy/idle through `api state`, so marker-down is the **normal steady state during work** — marker-first misjudges healthy endpoints routinely.\r\n- **Invariant:** probe the perch **directory**, not the marker. A dir carrying a **record or a spool** is OCCUPIED — refuse to adopt it. Ambiguity resolves toward refusal, always: an **unreadable** record (corrupt/truncated `info.json`) or an unreadable dir counts as occupied, never as residue, because deletion is irreversible while refusal is recoverable. A pre-existing **empty** dir is refused too — it is not provably ring's own residue (a bringup mid-create owns an empty perch dir for a window, and deleting it is the same race in a different hat). Refusal never blocks and is never silent: the message is delivered, the call declines to block-wait (the reply lands on the caller's own listener), and the caller gets a **distinct** loud report — `RING_PERCH_EX…\n- **spt-core mapping:** `spt_msg::ring::{probe_perch, setup_ephemeral, ring}`; `RingOutcome::{PerchExists, StaleDir}` rendered by `cmd_ring`.\r\n- **Source:** spt-core field incident 2026-07-27 (emphasys endpoint + spool loss); ruled by doyle, with emphasys's mid-create-race amendment revoking the empty-dir self-heal.\r\n- **Amendment — the perch GC inherits this reasoning EXTENDED, never carved out (doyle, 2026-08-04, releases#109).** A second consumer now deletes perch directories: `spt endpoint gc` (3.5, `REQ-PERCH-GC-RESIDUE-PREDICATE`). Its authority is something `ring` never had — a **positive** statement that the directory carries **no endpoint record at all**, from the store that answers endpoint existence on a node (the owlery tree itself; the perch dir plus its parseable `info.json` IS the endpoint record). That authority does **not** outrank this section's occupancy rules, and the ruling says why: undeliverable-by-construction settles *deliverability*, not *value*, so a recordless dir carrying a **spool** is still refused — that spool is the only surviving copy of…\n<!-- [doc->REQ-HAZARD-RING-PERCH-ADOPTION] -->\r\n\r\n### 3.5 `is_perch_alive` is INVERTED on perch residue — a GC must never key on it\r\n<!-- [doc->REQ-HAZARD-PERCH-GC-LIVENESS-INVERSION] -->\r\n<!-- [doc->REQ-PERCH-GC-RESIDUE-PREDICATE] -->\r\n- **Failure:** the obvious way to write a perch garbage collector is \"reap what is not alive\" — `!spt_store::liveness::is_perch_alive(dir)`. On the population a GC actually walks, that predicate is **backwards**. `is_perch_alive` returns **true** for a directory with no `info.json` (`RawRead::Absent ⇒ true`, interim parity — an absent record means a wrapper-owned listener between polls), and **false** for a daemon-hosted endpoint whose `status` is `offline`. Measured on HFENDULEAM 2026-08-04 across 38 top-level perch dirs: **all 24 recordless residue directories read ALIVE, and all 6 offline REAL endpoints read DEAD.** A sweep keyed on `!is_perch_alive` would have spared every stray probe directory and deleted six resting agents' records — it does not under-…\n- **Why it is a hazard and not a bug in the resolver:** `is_perch_alive` is correct for its own job. Fail-toward-alive is the safe direction for DELIVERY (2.5) — the cost of a false \"alive\" is one message taking the spool path. It is the wrong direction for DELETION, where a false \"not alive\" is irreversible. The inversion is what happens when a predicate is reused across a polarity boundary it was never sized for.\r\n- **Invariant:** perch garbage collection classifies on record **presence** only (3.5's predicate: no `info.json` present on any read attempt ⇒ residue; NotFound kept distinct from every other I/O error, because absent is the only answer that authorizes deletion) and **never calls a liveness resolver** — not `is_perch_alive`, not `is_registry_entry_alive`, not a pid probe. The registry is likewise never asked: `clean_stale_entries` deletes dead-pid rows and `unregister_address` fires on an ordinary stop, so registry-absence is the normal steady state of every offline endpoint (4.3). The unit pinning this asserts **both** arms with the shipped resolver as the witness: residue reads alive and is reaped anyway; an offline endpoint reads dead and is kept anyway.…\n- **spt-core mapping:** `spt_store::perchgc::{sweep, PerchClass}`; `spt endpoint gc` renders it. The refusal/report contract is `REQ-PERCH-GC-REFUSAL-REPORT`; the occupancy rules it inherits are 3.4.\r\n- **Source:** measured by todlando on HFENDULEAM 2026-08-04 while answering doyle's step-1 predicate question on releases#109; ruled into the lane by doyle the same day.\r\n\r\n---\r\n\r\n## 4. Wire / transport\r\n\r\n### 4.1 Envelope HTML-entity codec ordering — `&amp;` decoded LAST\r\n- **Failure:** decoding `&amp;`-entity before the others double-decodes nested entities (`&amp;amp;lt;` → wrong result).\r\n- **Invariant:** ENCODE order amp→first … `<br>`→last; DECODE order `<br>`→first … amp→**last** (`&lt;`,`&gt;`,`&quot;`, then `&amp;`). One sole decode site (at the LLM/stdin boundary); the parser never decodes.\r\n- **spt-core mapping:** `spt-proto` owns the envelope grammar (public SDK, semver + wire-version). This codec contract is a copy-verbatim commodity item (ADR-0001) and a public-API conformance test.\r\n- **Sister cite:** `src/owl/poll.rs:1-73`; `src/common/envelope.rs`.\r\n- **CR-linesafety `[REQ-HAZARD-ENVELOPE-CR-LINESAFE]`:** the EVENT is LINE-FRAMED, so the codec must neutralize raw `\\r` too — `event_body_escape` folds CRLF/lone-CR to `\\n` (→`<br>`) **before** framing. **Failure (field, 2026-06-08):** a cross-node `spt send` from Windows (`echo` → CRLF) carried a raw `\\r` into the single-line envelope; the receiver terminal did a CR→column-0 overwrite (`</EVENT>` clobbered `<EVENT t`). `\\r` was never line-representable here, so normalizing it is robustness, not an ADR-0001 wire divergence (decoder + amp-last untouched). Belt-and-suspenders: `spt send`/`ring` trim stdin like `notify`.\r\n\r\n### 4.2 Two-slice envelope parser is panic-free and tolerant\r\n- **Failure:** malformed envelope (unclosed/misordered/nested tags) panics or drops output.\r\n- **Invariant:** tags case-sensitive, all optional; no tags → whole body to live slot; unclosed → None for that tag; out-of-order → both still extracted; nested unknown tags preserved verbatim; zero `unwrap` on parsed text.\r\n- **spt-core mapping:** `spt-proto` parser; property-test the robustness rules.\r\n- **Sister cite:** `src/common/envelope.rs:64-92`; tests 99-207.\r\n\r\n### 4.3 Registry stale-entry cleanup precedes lookup\r\n- **Failure:** sender resolves a dead process's stale TCP port → delivery to wrong/dead listener.\r\n- **Invariant:** clean stale entries (dead PID) before/at lookup; spool fallback is the safe path on TCP miss.\r\n- **spt-core mapping:** now spans the **subnet registry** (ADR-0003) — eventually-consistent across nodes. Cross-node staleness is expected; resolution policy (local → most-recent → `id@node`) must degrade to spool/relay fallback on stale hits, and never hard-fail on a stale remote entry.\r\n- **Sister cite:** `src/common/registry.rs:62-78`; `src/owl/send.rs`.\r\n\r\n### 4.4 Deferred rows survive poll drain\r\n- **Failure:** poll `drain_all` flushes a deferred (spool-only) message meant for a hook consumer → message lost.\r\n- **Invariant:** deferred rows are never flushed by the event-stream drain; only `drain_non_deferred_*` / `peek_all` touch them.\r\n- **Sister cite:** CHANGELOG v1.11.20; `src/common/spool.rs`. (See also 1.4.)\r\n\r\n### 4.5 Inbox legacy compat must not double-deliver\r\n- **Failure:** message surfaced via both spool (durable) and legacy inbox files → duplicate or racing delivery.\r\n- **Invariant:** spool is the sole read path at poll time; inbox is write-for-compat only and never read.\r\n- **spt-core mapping:** clean-room — likely drop the legacy inbox entirely. If kept for any compat, preserve \"never read at drain time.\"\r\n- **Sister cite:** `src/common/inbox.rs`.\r\n\r\n### 4.6 Addressable-id charset reserves the address delimiters\r\n<!-- [doc->REQ-HAZARD-ID-CHARSET] -->\r\n- **Failure:** a bare endpoint id that contains `:` or `@` (or a path separator / whitespace / control char) makes the canonical qualified address `[subnet:]id[@node]` (ADR-0006 / REQ-INST-10) ambiguous to parse, and lets a name smuggle into a perch directory path. Once permissive ids exist in the wild, tightening later needs a migration.\r\n- **Invariant:** every addressable id/name is validated to `[A-Za-z0-9_-]` + Hiragana/Katakana/CJK only, length `1..=64`, **at every creation seam** (`ready` start, `api bind`, `api listen`, `api worker-start`). `:` and `@` are permanently reserved as address delimiters; reads of existing perches are never re-validated. Enforce now (pre-M3/M4) so no permissive id-data accumulates.\r\n- **spt-core mapping:** `spt_proto::id::validate_endpoint_id`; called at the four creation seams. The existing Psyche (`<parent>-psyche`) / Worker (`<parent>-w<N>`) suffix scheme uses only `-` + alphanumerics, so composite ids validate.\r\n\r\n### 4.7 Concurrent SQLite openers must not fail with \"database is locked\"\r\n<!-- [doc->REQ-HAZARD-REGISTRY-CONCURRENT] -->\r\n- **Failure:** two endpoints on one machine open the same SQLite store at once (e.g. two `ReadyAgent::start` calls registering simultaneously) and one fails outright with `SQLITE_BUSY` / \"database is locked\" → spurious registration/spool failure. Surfaced as a parallel-test flake in `two_agents_exchange_message_tcp_and_spool`, but the bug is real concurrency, not test-only.\r\n- **Invariant:** `busy_timeout` is set **before** any lock-taking statement on every connection. Switching `journal_mode=WAL` takes a brief exclusive lock; with the default 0ms timeout it fails immediately under contention, so the pragma order is load-bearing: `Connection::open` → `busy_timeout` → `journal_mode=WAL` → `CREATE TABLE …`. WAL alone is insufficient (concurrent *writers* still serialize; they must *wait*, not error).\r\n- **spt-core mapping:** `spt_store::registry::open_registry` + `spt_store::spool::open_spool_at`; both set `busy_timeout=5000` first. Any future SQLite store (history Path B, instance registry) must follow the same ordering.\r\n\r\n### 4.8 Registry merge ordered by epoch, never wall-clock (red-team #8)\r\n<!-- [doc->REQ-HAZARD-REGISTRY-EPOCH-LEASE] -->\r\n- **Failure:** the per-subnet registry replicates `endpoint_id → [instances]` eventually-consistently across nodes. Under a partition or clock skew, a lagging node re-announces a stale `Active` for an endpoint that has actually gone `Offline`. If the merge ordered updates by wall-clock (or \"last write wins\"), the stale `Active` overwrites the newer `Offline` and resolution routes a message to a dead/wrong instance.\r\n- **Invariant:** the merge precedence key is a **per-node monotonic epoch counter** (`spt_store::epoch::EpochSource`, persisted, strictly increasing, NEVER wall-clock), compared version-vector style per `(endpoint_id, node)`: an incoming update wins **iff its epoch is strictly greater** than the stored one for that node; equal or lower is dropped as stale. So a newer `Offline` (higher epoch) can never be clobbered by a lagging `Active` (lower epoch), and an idempotent equal-epoch replay is a no-op. Wall-clock is at most a human tiebreaker hint inside a flagged conflict, never the ordering authority. The same epoch source unifies with the D6 sync-precedence concurrent-write detection (#7).\r\n- **spt-core mapping:** `spt_net::net::registry::SubnetRegistry::merge_instance` (the lease) + `spt_store::epoch::EpochSource` (the counter). Cross-node replication of the merge wires at D4; the merge seam is identical for local and wire-delivered updates. Chaos/two-host verification = D9.\r\n\r\n### 4.9 SQLite stores must create their parent dir — SQLite won't\r\n<!-- [doc->REQ-HAZARD-REGISTRY-DIR-CREATE] -->\r\n- **Failure:** `Connection::open` creates the database FILE but never its parent DIRECTORY. On a fresh home (first boot, fresh CI `_work` dir) a registry op that runs before any perch-creating op (`create_dir_all` side effects) fails `SQLITE_CANTOPEN` — \"unable to open database file …owlery\\.registry\". Timing-dependent: whichever code path touches the home first decides the outcome, so it surfaces as a parallel-test flake (bind-first tests losing the dir-creation race to perch-first tests). Bit the hfenduleam CI leg twice (2026-06-03/04, four spt-msg unit tests at once on the second strike) before being run to ground; a slow runner filesystem (AV scanning fresh dirs) widens the window but is not the cause.\r\n- **Invariant:** every SQLite store's open path `create_dir_all`s its parent dir itself, best-effort, before `Connection::open` — never relying on another subsystem having materialized the home first. (Mirrors the spool, which always did this; the registry didn't.)\r\n- **spt-core mapping:** `spt_store::registry::open_registry` (`create_dir_all(owlery)` before open). `spt_store::spool::open_spool_at` already creates its perch dir. Any future SQLite store must do the same — pair this with the 4.7 pragma ordering on every new store.\r\n\r\n### 4.10 Dead node identities leave immortal registry rows  `[REQ-HAZARD-REGISTRY-GHOST-ROWS]`\r\n<!-- [doc->REQ-HAZARD-REGISTRY-GHOST-ROWS] -->\r\n- **Failure:** the registry's only superseding mechanism is the per-`(endpoint_id, node)` epoch lease (4.8) — a row is replaced only by a newer row *from the same node*. When a node identity dies permanently (machine retired, or `node.key` regenerated so the \"node\" never speaks again), its rows are never superseded and never expire: they sit in the in-memory registries and the `identity/registry/<subnet>.json` snapshots forever. A bare-id send then resolves the same endpoint id on both the live and the dead identity and refuses with a **phantom `AcrossNodes` ambiguity** — unfixable by the user, because no qualifier reaches a node that no longer exists. Hit live in the M7 acceptance run (2026-06-06): gravity paired under two identities (09ef…, then 03854a… af…\n- **Invariant:** registry rows authored by a **silent** peer node decay: a node not *heard* (admitted inbound feed — the M7 D2 heard-map, REQ-SUBNET-1) within the eviction window (`registry_evict_after_ms`, default 300s ≈ 10 default pump cadences) has its rows **evicted** from every subnet registry, snapshots rewritten. Own rows never decay (the node always hears itself implicitly — it authors them each pump tick). Eviction is safe under the lease: v1 has **no transitive gossip**, so any future update for a node comes from that node itself, alive, re-inserting from its durable `EpochSource` within one cadence — there is no lagging third-party replay to mis-order against. A merely-offline node loses its rows after the window and reconverges on return; meanwhi…\n- **spt-core mapping:** `spt_net::net::registry::SubnetRegistry::evict_nodes` (model) + `spt_daemon::registryhost::RegistryHost::evict_silent_peers` (heard-map TTL) driven from the registry pump tick (`peerloop`). Trust rows are NOT auto-evicted (trust is a user decision; a stale trust row only costs dead dials) — pruning those is a separate verb.\r\n- **Source:** M7 acceptance run 2026-06-06 (DEFERRED.md \"Ghost registry row eviction\"); the AMBIGUOUS render fix rode along.\r\n- **Mesh note (ADR-0017, 2026-06-08):** the subnet mesh **preserves** this invariant rather than superseding it. \"No transitive gossip\" sharpens to **no transitive *row* gossip** — the mesh relays only the member *roster* (discovery), while registry **rows stay own-authored and are fetched directly** from each member over a handshake. So \"any future update for a node comes from that node itself, alive\" still holds and the eviction lease is untouched. (The plan's rejected alternative — signed transitive *row* relay — would have broken this; roster-only relay was chosen precisely to keep it.)\r\n\r\n### 4.11 Advertisement-epoch reset strands a node  `[REQ-HAZARD-EPOCH-RESET]`\r\n<!-- [doc->REQ-HAZARD-EPOCH-RESET] -->\r\n- **Failure:** a node whose advertisement-epoch counter resets (the durable `EpochSource` file lost/recreated) re-advertises with LOW epochs; peers hold a higher last-seen epoch for that `(endpoint, node)` lease and drop every fresh row as **stale** — the node advertises into a void until its counter outruns its own history. Nothing renders the cause: the node looks healthy locally, peers simply never update.\r\n- **Invariant (mitigation by construction, common case):** the common trigger — a full reinstall / identity regeneration — is covered by the **re-pair trust overwrite** (M8 decision 13, REQ-SUBNET-7): a completed ceremony presenting the same label + machine id evicts the superseded identity's trust AND registry rows on the seed-holder, and the peer-side epoch memory **dies with the deleted row** — the re-paired node's fresh epochs land on a clean lease. M8 acceptance 7 verifies this explicitly (the epoch sub-check).\r\n- **Residual (documented, guard deferred):** the narrow slice — epoch file lost while the node *identity* is kept (manual state surgery, partial restore from backup) — has no guard; it waits for a field hit before one is designed (M8 decision 24). `REQ-HAZARD-EPOCH-RESET` is minted inactive (TRACEABILITY rule 5) as the tracking hook. If hit: symptoms are one node's endpoints frozen-stale on every peer while its own views are fresh; recovery today is re-pairing the node (rides the common-case eviction above).\r\n- **spt-core mapping:** epoch mint = `spt_store::epoch::EpochSource` (`identity/epoch.json`); the lease = the per-`(endpoint, node)` epoch compare in `spt_net::net::registry`; the eviction that clears peer-side epoch memory = `registryhost::repair_evict_superseded` + `RegistryHost::consume_repair_evictions`.\r\n- **Source:** minted at M8 ratification (decision 24), recognized as a class during the 2026-06-07 pump diagnosis / re-pair overwrite design — not yet field-hit in its residual form.\r\n\r\n---\r\n\r\n## 5. Platform-specific\r\n\r\n### 5.1 Windows PID recycling false positives\r\n- **Failure:** recycled PID reads alive for the wrong process → orphan misclassification.\r\n- **Invariant:** anchor liveness on the stable parent/harness PID; minimal info.json for supervisor-owned perches; mtime grace window (≥60s) masks transient mismatches.\r\n- **spt-core mapping:** broker-held handles supersede PID polling for spt-hosted sessions; keep the grace window for harness-hosted.\r\n- **Sister cite:** `src/live/wrapper/orphan.rs:141-161`; `src/common/list_filter.rs:168-175`.\r\n\r\n### 5.2 Windows EBUSY on atomic rename\r\n- **Failure:** `fs::rename` fails while a handle is (recently) held → registry/marketplace update fails.\r\n- **Invariant:** tmp-write + atomic-rename with retry/backoff; best-effort side-fail; tolerate transient EBUSY.\r\n- **spt-core mapping:** all on-disk state writes (registry, trust store, spool checkpoints) use this pattern. Self-update binary swap on Windows especially.\r\n- **Sister cite:** CHANGELOG \"EBUSY\"; `src/common/owlery.rs` atomic_write.\r\n\r\n### 5.3 Git/subprocess timeout stamping\r\n- **Failure:** a hung subprocess (git on slow net) blocks the supervisor indefinitely.\r\n- **Invariant:** every metadata-producing subprocess has a timeout; timeout yields `None` + rate-limited stderr, never a hang.\r\n- **spt-core mapping:** generalize to all manifest-declared harness invocations (delegated commands, adapter updates) — timeouts mandatory.\r\n- **Sister cite:** `src/common/git.rs`.\r\n\r\n### 5.4 Windows UNC prefix in serialized paths\r\n- **Failure:** canonicalized `\\\\?\\C:\\...` serializes to `//?/C:/...` and fails `read_to_string`.\r\n- **Invariant:** strip the `\\\\?\\` UNC prefix after backslash→forward-slash conversion; serialized path attrs must be directly consumable.\r\n- **spt-core mapping:** any path crossing the wire (file-drop EVENTs, off-node file transfer per ADR-0003) needs canonical normalization at the `spt-proto` boundary.\r\n- **Sister cite:** `src/common/owlery.rs:377-384`.\r\n\r\n### 5.5 ConPTY withholds output until DSR is answered  `[REQ-HAZARD-CONPTY-DSR]`\r\n- **Failure:** a broker reading a ConPTY master sees only the 4-byte startup query `ESC [ 6 n` and then nothing — the child looks hung/silent but is producing output normally. ConPTY blocks all child stdout until the terminal answers the cursor-position query.\r\n- **Invariant:** every ConPTY reader auto-answers DSR (`ESC [ 6 n` → write `ESC [ 1;1 R`, or a real cursor position) on the PTY writer. Secondary: a ConPTY master does not EOF while the writer is held, so read loops drain on a thread and never gate exit on a blocking `read()`.\r\n- **spt-core mapping:** `spt-term` broker PTY reader (ADR-0004). Brand-new to spt-core — not in the sister project (it never hosted ConPTY directly).\r\n- **Source:** Spike #1 (`docs/spikes/SPIKE-01-broker-handoff.md`); reproduced with both a Rust child and `cmd.exe`.\r\n\r\n<!-- [doc->REQ-HAZARD-DETACHED-PIPE-INHERIT] -->\r\n### 5.6 Windows detached children inherit a captured caller's pipe  `[REQ-HAZARD-DETACHED-PIPE-INHERIT]`\r\n- **Failure:** a caller captures an `spt` invocation's output through a pipe (`Command::output()`, a harness hook reading the command). That `spt` process detach-spawns a **long-lived** child (the daemon via `ensure_running`; a shell binary via `spt shell spawn`). On Windows `CreateProcess` runs with `bInheritHandles = TRUE`, and the spt process's std handles — the caller's pipe write-ends — are inheritable by construction, so the immortal child inherits them even when its *own* stdio is `Stdio::null()`. The caller's pipe read never sees EOF: the capturing caller **hangs forever** (unix is immune — pipe fds are `CLOEXEC`). Paid twice: daemon spawn (guarded at D4a-era `spawn_detached`), then again at M5-D3e when the mock-shell E2E hung `spt shell spawn` for h…\n- **Invariant:** every detach-spawn of a long-lived child inherits only handles it **enumerates**. The default and overwhelmingly common form of that is `bInheritHandles = FALSE` (`spt-daemon::daemon::detached_no_inherit`) — zero handles flow, whatever the pipe's depth in the ancestry. Stripping `HANDLE_FLAG_INHERIT` from the spawner's *std* handles is NOT sufficient: a grandparent capture's pipe sits in the handle table as a stray inheritable handle and still flows through every **bare** `bInheritHandles = TRUE` hop (the first guard shipped that way and was wedged by exactly this — a daemon spawned three layers deep held the pwsh-level pipe of the CI/test harness).\r\n- **The one permitted TRUE (amended 2026-07-26, RESIDENT-SERVICE W1):** a spawn that must capture the child's output may pass `bInheritHandles = TRUE` **only** paired with an explicit `STARTUPINFOEX` + `PROC_THREAD_ATTRIBUTE_HANDLE_LIST` allowlist — the Win32 contract honors the list only under TRUE, so the earlier phrasing (\"flipping TRUE is refused\") forbade its own mechanism and was unimplementable. The hazard was always *unenumerated* inheritance, and the list **is** the enumeration: only the named handles flow, so a grandparent's pipe cannot reach the child at any depth — strictly stronger than what FALSE achieves incidentally. A **bare** TRUE, with no attribute list, remains REFUSED. Today the only such path is a supervised `[service]` spawn with start…\n- **spt-core mapping:** `spt-daemon::daemon::spawn_detached` (the daemon) and `spt-daemon::shellhost::launch_shell` (the relay-receipt shell binary). Any future long-lived detached spawn (manifest-template children included) must use the same no-inherit spawn, or the enumerated form above where it genuinely must capture output.\r\n- **Source:** spt-core, M5-D3e (`shell_e2e.rs` hang, 2026-06-04, twice — once per guard generation); Rust `Command` restricts *its own* created stdio handles but a parent's inheritable handle table still flows.\r\n\r\n### 5.7 Elevated commands spawn the daemon with the wrong token  `[REQ-HAZARD-ELEVATED-DAEMON-SPAWN]`\r\n<!-- [doc->REQ-HAZARD-ELEVATED-DAEMON-SPAWN] -->\r\n- **Failure:** membership-implies-reachability made *every* `spt` invocation a potential daemon spawner (`ensure_running`), including the elevation-gated ones (`subnet create`/`join`, REQ-SUBNET-4). The spawned daemon inherits the spawner's token. **Windows:** an elevated `subnet create` auto-starts an ELEVATED daemon whose named pipes deny unelevated clients — every subsequent unelevated `spt` reads \"not running\", tries to spawn its own daemon, and dies on bind Access-denied; the user had to taskkill (hit live, M7 acceptance 2026-06-06). **Linux:** a sudo'd command spawns a root daemon and/or root-owned state — and because sudo flips `$HOME`, the daemon can mint a *different node identity* in root's universe (the very key-flip that produced the 4.10 ghost r…\n- **Invariant:** the daemon **always runs unelevated in the invoking user's universe**, regardless of which command spawns it. Two enforcement points sharing one seam: (a) `spawn_detached` de-elevates the child — Windows: the UAC **linked token** (`TokenLinkedToken` → `DuplicateTokenEx` → `CreateProcessWithTokenW`; inherits no handles, so 5.6 holds by construction); Linux: drop to `SUDO_UID`/`SUDO_GID` with `$HOME`/`$USER`/`$LOGNAME` reset to the invoking user's (passwd lookup); (b) a `Daemon::run` entry guard catches a *directly* elevated `spt daemon` — Linux drops privileges in-process before touching any state; Windows respawns de-elevated and exits. When no unelevated identity exists to drop to (UAC disabled, genuine root login, SYSTEM), the daemon runs …\n- **spt-core mapping:** `spt-daemon::deelevate` (the OS-split seam) consumed by `daemon::spawn_detached` + the `Daemon::run` entry guard. The fuller Linux elevation model (install symlink + default-account election) is deferred (DEFERRED.md, M8).\r\n- **Source:** M7 acceptance run 2026-06-06 (DEFERRED.md \"Non-admin daemon spawn\"); interim field rule was \"bring the daemon up unelevated FIRST\".\r\n\r\n<!-- [doc->REQ-HAZARD-CHILD-CONSOLE-FLASH] -->\r\n### 5.8 Console children of the console-less daemon flash visible windows  `[REQ-HAZARD-CHILD-CONSOLE-FLASH]`\r\n- **Failure:** the daemon runs DETACHED (no console, 5.6/`detached_no_inherit`). Any console-subsystem child it spawns (`git`, `taskkill`, manifest hook commands) gets a **fresh conhost with a visible window** — piped/null stdio does NOT prevent it. Field shape: the 60s sync pump's two git calls (`for-each-ref` + `rev-parse`) flashed two blank windows per minute on the user's desktop (2026-06-06).\r\n- **Invariant:** every short-lived console child spawned from daemon-reachable code sets `creation_flags(0x0800_0000)` (`CREATE_NO_WINDOW`). Long-lived detached children use `detached_no_inherit` (already `DETACHED_PROCESS | CREATE_NO_WINDOW`); de-elevated spawns use `CREATE_NEW_CONSOLE + SW_HIDE` (5.7 — `CreateProcessWithTokenW` rejects `CREATE_NO_WINDOW`, error 87).\r\n- **Test seam caveat:** window-absence is unobservable from a consoled test runner — the child inherits the runner's console and never creates a window, flag or no flag. Unit coverage asserts the flagged spawn still works (the error-87 \"flag combo breaks spawn\" regression class); window-absence was verified live by process-watch capture.\r\n- **spt-core mapping:** `spt-store::gitrun::run_git` (every BranchStore/ContextStore git call), `spt-daemon::shellhost::kill_shell_pid` (taskkill), `spt-runtime::run_bounded_command` (manifest hook commands), `spt-runtime::ManifestRuntime::command_for` (the one shared builder behind `spawn_session` + `run_bounded_stdin` — the notif pump's `spawn_notif_command` and the live agent's psyche/echo/turn spawns), `spt-daemon::shellwake` (already guarded). The flag lives in each shared builder, not per call site, so the invariant holds for every ManifestRuntime spawn by construction.\r\n- **Source:** spt-core field bug, 2026-06-06 — two blank windows flashing every 60 seconds on a desktop workstation, caught by process-spawn watcher (git.exe parent=spt daemon, conhost.exe child each).\r\n\r\n### 5.9 `Instant - Duration` underflow-panics on a freshly-booted host  `[REQ-HAZARD-INSTANT-UNDERFLOW]`\r\n- **Failure:** `Instant::now() - Duration::from_secs(N)` panics `overflow when subtracting duration from instant` when the process's monotonic clock is younger than `N` — i.e. the host booted less than `N` ago. The peer pump primed its cadence legs with `Instant::now() - 86_400s` to mean \"everything due now\"; on a Windows runner with sub-24h uptime the pump thread panicked at startup, so the subnet never converged (CI `pump_and_dispatch_self_drive_the_subnet` failed, run 27082417706). It is *environment-conditional* — green on any host up longer than the offset, red below it — so it slips local dev and only bites a fresh CI box or a just-rebooted machine.\r\n- **Invariant:** NEVER compute an instant in the past by subtracting from `Instant::now()`. Represent \"never run / due now\" as `Option<Instant> = None` and gate on forward `now.duration_since(past)` only (`peerloop::due`). No backward instant arithmetic anywhere in scheduling.\r\n- **Test seam caveat:** the convergence E2E only reproduces on a sub-offset-uptime host (it passed everywhere with >24h uptime). The deterministic guard is the `due(None, ..)`/`due(Some(now), ..)` unit on the extracted gate — it asserts first-tick-due with zero instant subtraction, independent of host uptime.\r\n- **spt-core mapping:** `spt-daemon::peerloop::due` (the sole cadence gate behind `due_reg`/`due_notif`/`due_sync`/`due_upd`); cadence legs are `Option<Instant>` seeded `None`.\r\n- **Source:** spt-core CI failure, 2026-06-07 — Windows runner `hfenduleam` (just booted) panicked the peer pump at the v0.1.1 release gate.\r\n\r\n### 5.10 `sudo spt` dead-ends on a user-local install (secure_path)  `[REQ-HAZARD-SUDO-SECURE-PATH]`\r\n- **Failure:** the elevation-gated commands (`subnet create` / `subnet join` / `show-code`) refuse when unelevated and tell the user to \"run as administrator / root\". The user types the obvious `sudo spt subnet create FOO` → `sudo: spt: command not found`. `spt` is a user-local install (`~/.local/bin`, `~/.cargo/bin`), and sudo's `secure_path` (a `/etc/sudoers` default) does NOT include those dirs, so a bare command name doesn't resolve under sudo. The guidance is a trap: it names an action that cannot work for the common install shape. Field-hit on KITSUBITO at the v0.1.1 ship.\r\n- **Invariant:** elevation guidance on Unix emits the binary's **absolute path** under sudo — `sudo /home/u/.local/bin/spt subnet create FOO` — reconstructed from `current_exe()` + the real argv and shell-quoted. An absolute program path is executed directly; `secure_path` only governs bare-name PATH lookup, so the absolute form always resolves. On an interactive Unix TTY the command auto-elevates (re-execs itself under sudo, the elevated child does the work and `main` de-elevates back); non-interactive or sudo-absent falls back to printing the runnable hint. Never emit a bare-name elevation instruction.\r\n- **Companion UX:** the post-de-elevation `DEELEVATED: running as uid N` line is internal state-safety noise — omit it from the user-facing CLI path (it confused the same field user). The detached daemon's own de-elevation log line is fine (it lands in the daemon log, not the terminal).\r\n- **Test seam caveat:** the sudo re-exec needs a real `sudo` + TTY (not hermetic). The deterministic guard is the pure `elevation::sudo_argv` / `print_hint_command` (assert an absolute exe path, never a bare name, + shell-quoting on the printed line) and the `decide_elevation_path` matrix (which picks inline-sudo only on an interactive Unix TTY); the exec leg is manual/kitsubito-verified.\r\n- **spt-core mapping:** `spt::elevation::{sudo_argv, print_hint_command, decide_elevation_path}` (pure — generalized from the M12-W4 self-elevation seam, 5.11), `spt::cli::{try_auto_elevate, with_elevation_hint}` wired into `cmd_subnet_create` / `cmd_subnet_join` / `cmd_subnet_show_code`; `spt::main` de-elevation drop silenced.\r\n- **Source:** spt-core field report, 2026-06-07 — `reavus@KITSUBITO`, `spt` in `~/.local/bin`; the absolute-path `sudo` invocation was confirmed working before the fix landed.\r\n\r\n### 5.11 Self-elevating re-launch must re-run verbatim, never widen / inject / loop  `[REQ-HAZARD-SELF-ELEVATE]`\r\n- **Failure class:** a privilege-gated command (`subnet create` / `join` / `show-code`) self-elevates by re-launching itself with privilege (Windows UAC `runas`, Linux `pkexec` / a terminal-emulator `sudo`, or inline `sudo`). A careless re-launch is a security hole: widening the privilege scope (adding args), resolving the binary by a bare name (a PATH/`secure_path` hijack runs an attacker's `spt`), interpolating a crafted arg into a shell string (`sh -c \"… $id …\"` injects a second command), or re-elevating the already-elevated child (an infinite UAC/polkit loop). The user's UAC/polkit/sudo prompt is the ONLY consent gate — the mechanism must never bypass or widen it.\r\n- **Invariant:** self-elevation re-runs the **EXACT** original invocation with the binary's **ABSOLUTE** exe path — never adding/altering args, never a PATH-resolved bare name, never a shell-interpolated string. Every launcher passes an **argv array** (`Command::new(prog).args([...])`, never `sh -c`); the Windows `ShellExecuteW` params string (which is inherently one string) MSVC-quotes each verbatim arg so `CommandLineToArgvW` round-trips it as a single token. The elevated child drops state back to the user (composes with the 5.7 de-elevation) and **never re-elevates**: `decide_elevation_path` returns `AlreadyElevated` whenever the process is `Elevated`, on every OS (loop-safety). The unprivileged parent never pipes/captures the elevated child's stdout acro…\n- **Test seam caveat:** the real launch needs a UAC/polkit/sudo prompt (not hermetic) — manual-verify. The deterministic guards are the pure `decide_elevation_path` matrix (loop-safety: `AlreadyElevated` on every os; the os×env path order) and the argv builders (`sudo_argv` / `pkexec_argv` / `terminal_argv` assert absolute-exe + verbatim args + array; `windows_runas_params` asserts MSVC-quoting with no `cmd /c` interpolation; the crafted-arg test asserts a shell-metachar arg stays one element / one quoted token).\r\n- **spt-core mapping:** `spt::elevation::{decide_elevation_path, sudo_argv, pkexec_argv, terminal_argv, windows_runas_params, print_hint_command, ElevatePath}` (pure), `spt::cli::{try_auto_elevate, launch_uac_window, pause_elevated_console_if_fresh, program_on_path, first_terminal_emulator}` (impure launchers) wired into `cmd_subnet_create` / `cmd_subnet_join` / `cmd_subnet_show_code`. Companions: 5.10 (the Unix abs-path-under-sudo facet) and 5.7 (the elevated child's de-elevation drop, which this composes with).\r\n- **Source:** M12-W4 design (subnet QR + self-elevating window), doyle ruling `M12-W4-RULING.md` Q6 — a privilege-escalation feature carries a mandatory hazard REQ.\r\n\r\n<!-- [doc->REQ-HAZARD-WIN-PTY-PROGRAM-RESOLVE] -->\r\n### 5.12 Native-PTY spawn of a bare program runs the wrong (non-PE) file on Windows  `[REQ-HAZARD-WIN-PTY-PROGRAM-RESOLVE]`\r\n- **Failure:** `portable-pty`'s ConPTY spawn resolves a bare program name with a `which` that takes the FIRST `PATH` match. A node/npm CLI installs as BOTH an extensionless shebang shim (`ccs`, for Git Bash) and a Windows launcher (`ccs.cmd`) in the same dir; portable-pty picks the extensionless `ccs`, and `CreateProcessW` then tries to execute that non-PE file and fails with **os error 193** (\"%1 is not a valid Win32 application\"). Live failure: `spt endpoint run claude-spt:ccs` → `CreateProcessW C:\\nvm4w\\nodejs\\ccs` 193 (operator, 2026-06-16). The same bites any harness/shell whose `[session.self]`/`[shell].spawn` names a `.cmd`/`.bat`/`.ps1`-backed command — `CreateProcessW` cannot execute a batch or PowerShell script directly.\r\n- **Invariant:** spt-term resolves the program ITSELF before handing it to `CommandBuilder`, bypassing portable-pty's `which`. A bare name is searched over `PATH` × `PATHEXT` (whose default order already prefers `.EXE`/`.COM` over `.BAT`/`.CMD`), then an extensionless fallback. A non-PE target is wrapped in its interpreter: `.cmd`/`.bat` → `cmd.exe /d /c <path>`, `.ps1` → `powershell -NoProfile -File <path>` (the wrap args precede the caller's args); a real executable spawns directly; an unresolvable name passes through unchanged (never makes a working case worse). Unix is a passthrough — `execve` honours a shebang on an extensionless script. Applied at the ONE `CommandBuilder` chokepoint (`PtySession::spawn_program_in`), so every broker harness + shell spaw…\n- **spt-core mapping:** `spt_term::winprog::{resolve_for_pty, resolve_in}` (the pure PATHEXT-precedence kernel + the Windows env wiring), wired into `spt_term::pty::PtySession::spawn_program_in`. Unit: `resolve_in` precedence (`.cmd`-over-shim, `.exe`-direct, explicit-extension, path-order, passthrough) [`winprog.rs`].\r\n- **Source:** field diagnosis 2026-06-16 (operator dogfood, `claude-spt:ccs` bringup) — doyle.\r\n\r\n<!-- [doc->REQ-HAZARD-PERCH-RECORD-POWER-LOSS] -->\r\n### 5.13 Atomic write leaves data un-synced before the rename → NUL zero-fill on power loss  `[REQ-HAZARD-PERCH-RECORD-POWER-LOSS]`\r\n- **Failure:** `atomic_write_bytes` was `fs::write(tmp)` + `rename(tmp, path)` with no `fsync`. The rename's directory **metadata** is journaled durable, but the tmp file's **data blocks** are still in the page cache. A hard reset (power loss, forced reboot) between the two flushes lands the rename but loses the data → the file reappears at its **full length filled with NUL**. Field incident: after a machine restart `owlery/hall-a/info.json` was 360 bytes of all-NUL (the nested psyche record 125 bytes all-NUL); `Get-Content` renders blank, `Format-Hex` shows the zero-fill. A wiped-but-present record then read as a live ONLINE endpoint for days (compounded by 5.14).\r\n- **Invariant:** for records whose loss is **unrecoverable**, flush data to stable storage **before** the rename publishes the name — `File::create(tmp)` → `write_all` → `sync_all()` → `rename_with_retry`. Then a crash yields the complete old file OR the complete new file, never a NUL husk. (A file of correct length that is ALL-NUL is the diagnostic tell of this class — a non-fsync'd write caught by a hard reset.) **Durability is SCOPED, not blanket.** `fsync`ing *every* atomic write serializes a durable flush behind each of the ~26 `atomic_write` callers — and for `info.json` the flush lands under the per-perch `.info.lock` (the **W1b lock-across-fsync wedge shape**, §6-era) — which stalls daemon bringup 4–6× so the endpoint misses its ONLINE window. Scope …\n- **spt-core mapping:** `spt_store::atomic::atomic_write_bytes_durable` / `atomic_write_string_durable` are the opt-in durable siblings; the default `atomic_write_bytes` / `atomic_write_string` stay non-durable. Durable callers are exactly `spt_store::info::write_info` (the perch record — the hall-a surface), `spt_store::nodeid` (the node seed — corrupt is node-bricking, never regenerated), and `spt_daemon::machineid` (minted-once machine id — a NUL husk would silently re-mint a different id). Reader-side handling of an already-corrupt record is the sibling 5.14. **Canary:** `attach_wedge_e2e` is the regression guard — a blanket fsync fails it (bringup misses the 20s ONLINE budget); keep durability scoped so it stays green.\r","truncated":true,"truncatedBy":"lines","totalLines":1091,"totalBytes":68795,"outputLines":300,"outputBytes":68795,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":1091,"displayContent":{"text":"# Known Hazards\r\n\r\nHard-won edge cases harvested from the sister project (`claude_skill_owl`, ~80 commits / 12+ phases / multiple production incidents). Per ADR-0001, this is a **test checklist for the spt-core rebuild** — the clean-room rebuild must re-satisfy each invariant rather than re-discover the bug.\r\n\r\n**Architecture-translation note.** The sister project runs poll listeners and Psyche wrappers as *separate processes*. spt-core consolidates both into the one `spt-daemon` (brain), with a stable broker beneath it (ADR-0004). Many hazards below were inter-process races in the sister project; in spt-core some become intra-daemon concerns (potentially easier) while others move to the daemon↔broker IPC boundary or the network boundary (potentially new failure surface). Each entry notes the mapping where it differs. Citations point at sister-project paths for reference, not at spt-core.\r\n\r\n---\r\n\r\n## 1. Race conditions & ordering\r\n\r\n### 1.1 Phantom INIT_SIGNOFF after grace period\r\n- **Failure:** orphan teardown enqueues INIT_SIGNOFF before the grace-period recheck; a transient Self recovery (binary handoff, brief stale poll) makes the recheck pass-as-alive, but the signoff was already spooled and drains on the next iteration → teardown despite a live Self.\r\n- **Invariant:** grace-period wait MUST complete *before* composing/delivering INIT_SIGNOFF; the recheck must bind `still_gone` before any envelope write.\r\n- **spt-core mapping:** in-daemon now (no separate wrapper), but the ordering invariant is identical — orphan/teardown logic must re-evaluate liveness after the grace wait, not before enqueue.\r\n- **Sister cite:** `src/live/wrapper/orphan.rs:201-259` (sleep@209 precedes compose@231-251); tests T-grace-recovery:576, T-still-gone-recheck:618.\r\n\r\n### 1.2 Poll-rewrite race & info.json mid-write reads\r\n- **Failure:** `info.json` written by the wrapper mid-iteration while a list/classify command reads it → torn read, misclassification.\r\n- **Invariant:** consult liveness via the supervisor (`is_wrapper_alive`-equivalent) before any grace gate; reads of state files must tolerate concurrent writes (atomic write + rename, or read-retry).\r\n- **spt-core mapping:** the daemon owns both writer and reader → use in-process locking/snapshotting instead of racing on disk. Cross-node registry reads remain eventually-consistent and must tolerate staleness.\r\n- **Sister cite:** `src/common/list_filter.rs:100-150`; `src/owl/poll.rs:141`.\r\n\r\n### 1.3 Stale `index.lock` wedge from prior git crash\r\n- **Failure:** crashed git leaves a 0-byte `index.lock` in a psyche tracked worktree; every later commit blocks forever.\r\n- **Invariant:** on daemon boot, sweep seed + all agent/project worktrees for stale locks (0 bytes, mtime > 60s) and remove; leave live locks alone.\r\n- **spt-core mapping:** cross-node Psyche sync (ADR-0002/0003) replaces git-repo sync, so the *git* lock may disappear — but any equivalent lockfile in the new sync mechanism needs the same stale-sweep on boot.\r\n- **Sister cite:** CHANGELOG v1.11.20 \"Stale `index.lock`\"; `src/common/git.rs`.\r\n\r\n### 1.4 Deferred spool rows must not leak to the event stream\r\n- **Failure:** a hook spools a deferred (spool-only, no TCP wake) notice; startup `drain_all` flushes ALL rows including deferred → event emitted at wrong time/priority.\r\n- **Invariant:** startup drain (and idle/timeout TCP-wake sites) use `drain_non_deferred` only; deferred rows are picked up by their intended consumer via `peek`. All drain sites must agree on which rows they flush.\r\n- **spt-core mapping:** carries directly — the daemon's spool-drain has the same deferred-vs-immediate distinction.\r\n- **Sister cite:** `src/owl/poll.rs:276-316`; `spool::drain_non_deferred_with_metadata`.\r\n\r\n### 1.5 Worker (working-perch) lifecycle path consistency\r\n- **Failure:** subagent-start creates the perch at one path layout; later hooks read it at another → not found; stop-hook scan misses nested perches.\r\n- **Invariant:** all Worker/Psyche child-perch path composition routes through one central resolver; no divergent path construction across hooks.\r\n- **spt-core mapping:** `Worker` is a day-one endpoint type; the daemon owns the registry, so perch location is a registry lookup, not ad-hoc path math. Single source of truth for instance→location.\r\n- **Sister cite:** `src/owl/hook_subagent_start.rs:122-168`; `hook_subagent_stop.rs:15-55`.\r\n\r\n---\r\n\r\n## 2. Identity & session-binding\r\n\r\n### 2.1 Parent PID over ephemeral poll PID\r\n- **Failure:** orphan check polls an ephemeral listener PID; it dies and is recycled (esp. Windows); a foreign process with the recycled PID reads as alive → false-positive teardown (or false-negative).\r\n- **Invariant:** prefer the stable harness-session PID (`parent_pid`) over any ephemeral process PID for liveness; minimal `info.json` for supervisor-owned perches to avoid stale leaks.\r\n- **spt-core mapping:** session binding (parent-process-tree anchor) still applies for harness-hosted topology. For spt-hosted sessions the broker holds the child directly → liveness is the broker's held-handle state, more reliable than PID polling.\r\n- **Sister cite:** `src/live/wrapper/orphan.rs:141-161`; CHANGELOG v1.11.20.\r\n\r\n### 2.2 Stdin session_id precedence over env\r\n- **Failure:** subagent inherits a stale `OWL_SESSION_ID` env across `/clear`; hook gets two session_ids (fresh stdin, stale env) → wrong-agent binding.\r\n- **Invariant:** stdin-provided session_id wins; env is fallback only.\r\n- **spt-core mapping:** the harness-contract subcommand surface must define the same precedence for whatever identity fields hooks pass in.\r\n- **Sister cite:** CHANGELOG v1.35.1 \"IN-05\"; `hook_subagent_start.rs:40-51`.\r\n\r\n### 2.3 Binary-handoff argv schema must stay backward-compatible\r\n- **Failure:** old binary spawns new binary with old argv arity; clap rejects before state rehydration → wrapper dies unlogged.\r\n- **Invariant:** every newly-added handoff positional has a default; state-file rehydration happens *after* argv parse; defaults survive intermediate versions.\r\n- **spt-core mapping:** CRITICAL — self-update (ADR-0004) makes handoff routine. The broker↔brain IPC and any brain-relaunch argv must be versioned and forward/backward tolerant (a newer brain talks to an older broker). This is the single most update-frequency-sensitive invariant.\r\n- **Sister cite:** `src/live/wrapper/lifecycle.rs:17-106`; `src/cli.rs` defaults; CHANGELOG v1.11.10.\r\n\r\n### 2.4 Generation `gen_start` always = now() on cold-start AND handoff\r\n- **Failure:** stale gen_start from a rehydrated state file fires time-based discriminators on the new process.\r\n- **Invariant:** wall-clock `gen_start` is set to `now()` on both cold-start and handoff; generation counter increments on every start/revive; session UUID captured fresh and carried so the resumed mind distinguishes \"same gen continuing\" vs \"new gen born\".\r\n- **spt-core mapping:** carries to the daemon's per-instance generation tracking.\r\n- **Restoration D3/D4 (ADR-0018):** the generation *counter* custody moved to the broker (D3-2 — it observes every brain spawn, planned or crash, and hands `{generation, start-reason}` at spawn; `gen_start` stays `now()`-fresh, never rehydrated). The brain→brain **`BrainState` *message*** (`{session_id, generation, next_seq, gen_start_ms}`) that previously carried continuity across a handoff is **retired from the production path in D4-2**: a brain the supervisor respawns cold-starts and reconstructs all session continuity by **querying the broker** (`Brain::resume_sessions` over the broker's cursor-of-record), never a frame. `BrainState` / `Brain::handoff` / `Brain::snapshot` remain `pub` and compiled **only for the integration tests** (handoff/idempotent/da…\n- **Sister cite:** `src/live/wrapper/lifecycle.rs:70`; `src/common/wrapper_state.rs`.\r\n\r\n### 2.5 Daemon-hosted endpoints have no dedicated liveness PID\r\n- **Failure:** the sister evaluates Psyche/perch liveness via a dedicated process PID — the wrapper's own pid in `info.json`, checked with `is_process_alive`. Under ADR-0004 the Psyche (and any spt-hosted Self) is a **loop inside the daemon**, not a separate process: it holds no dedicated pid, and its `claude`/summarizer subprocess is ephemeral (spawned per pulse/commune, then exits). If a daemon-hosted perch's `info.json` carries the **daemon's** pid, then *every* hosted endpoint shares one pid, and `is_process_alive(pid)` reads \"alive\" for a torn-down endpoint as long as the daemon runs — while `clean_stale_entries` (dead-pid deletion) can no longer distinguish a dead endpoint from a live one. The 2.1/5.1 liveness models do **not** cover this third categor…\n- **Invariant:** for **daemon-hosted** perches (Psyche; spt-hosted Self), liveness is the **daemon's authoritative in-memory endpoint table + a `status` field** on `info.json` (`online|offline|…`), **never** `is_process_alive(info.pid)`. `info.pid` for a daemon-hosted perch is at most a *hosted-by-daemon* marker (the daemon pid), not a liveness signal; registry stale-clean for these rows keys on the daemon's endpoint table, not per-row pid. This reuses the pattern already specified for **Shells** (`info.json` carries daemon-managed `status`, capability resolved by `adapter_name` — CONTEXT \"Shell… Not in the subnet registry\") and extends it to daemon-hosted *agent* perches.\r\n- **spt-core mapping:** the **M1/M2a interim** model keeps the Psyche/listener a real per-process owner (the `api listen` process), so its per-pid liveness (`deliver::is_online` → `info.read_pid` → `proc::is_process_alive`; `registry::clean_stale_entries`) is correct *interim*. **M3 daemon consolidation replaces it** with daemon-authoritative liveness for hosted perches. Keep the liveness check behind one resolver (mirrors `resolve_address` stale-clean) so the M3 swap is localized — do **not** let the per-pid assumption leak into new call sites.\r\n- **Sister cite:** `src/live/wrapper/orphan.rs` (wrapper-pid liveness); `src/common/list_filter.rs:168-175` (pid-classify); spt-core `crates/spt-store/src/{proc.rs,registry.rs}` + `crates/spt-msg/src/deliver.rs::is_online`.\r\n\r\n### 2.6 A shell instance's `online` is a recorded field with no death edge\r\n- **Failure:** a shell perch's `status` has exactly two writers — the bind handshake (→`online`) and `close_shell` (→`offline`) — and `close_shell` runs on the **link-break** path alone. A binary that dies **abruptly** (force-kill, crash, OOM) breaks no link, so nothing ever writes the flip: the record says `online` forever while the pid it names is gone. Every consumer of the recorded field then inherits the lie — `shell list` reports online indefinitely; `relink` (the one command that recovers the instance) refuses `SHELL_ALREADY_ONLINE`, gated on the very state that is wrong; and `shell cmd`'s wake-if-offline arm reads \"not offline\", skips the wake, and spools the command to a corpse — accepted happily, drained by nobody, **indistinguishable from a busy s…\n- **Invariant:** ONLINE-ness is **derived**, never the recorded field alone: `status == online` **AND** the recorded `shell.pid` is not *provably* dead. Provably dead is narrow — pid parked, non-zero, `!is_process_alive`; pid absent, unparseable, or `0` (a broker-hosted spawn whose backend exposed no pid) reads **alive**, the same fail-toward-alive parity 2.5 holds, so a backend the resolver cannot see is never declared dead. **The recycled-pid caveat is narrowed since BAROMETER W2 and is now conditional on the stamp:** a launch parks a **birth stamp** beside the pid (`shell.launch.json`), and the probe runs the PAIR test — same pid, *different* start time ⇒ **gone**, so a pid the OS handed to an unrelated process no longer masks a dead instance. That case m…\n- **No spontaneous relaunch of a SAME-BOOT death (operator-ratified, flynn 2026-07-25; re-based BAROMETER W2):** the reconciler does NOT adopt an instance whose binary died during this boot. Every reason an operator stops a shell is a reason not to want it back a tick later: mid-deploy (a shared install dir makes \"kill the process\" routine on Windows, since a running exe cannot be overwritten) an auto-relaunch would run the **old** binary out of the file being replaced — worse than a failed install — and it turns quarantining a misbehaving shell into a restart loop. For those, recovery stays demand-driven: `relink`, or a `shell cmd` that wakes. `relink` probes **locally** rather than trusting a daemon sweep, so recovery holds with the daemon down.\r\n- **What holds that open is no longer the stale record (BAROMETER W2):** until W2 the guarantee rested on an accident: a force-killed instance was held out of the eligible set because its record still said `online`, and the eligibility read took that field raw. 2.7 heals that record to the truth, which removes the accidental protection — so the guard had to move to a rule that *states* it. Eligibility is now recorded-`offline` **AND** (no corpse, **or** the corpse's recorded launch **predates the boot instant**). A steady-state force-kill leaves a corpse launched *after* boot, which the predicate can never accept, so the ruling is preserved **by construction** rather than by a second rule. A restart casualty is the one case the predicate does accept, and it …\n- **spt-core mapping:** `spt_store::shellinfo::{shell_pid_provably_dead, effective_status, is_shell_online}` (the resolver — the shell-side twin of 2.5's `liveness.rs`, which gave *agent* perches exactly this and which shells never got); consumers in `linkhost::{relink_shell, prepare_drive, run_action}`, `shelldisc::discover`, `activity::observe_links`; the class-(c) eligibility read is `spt_daemon::shellwake::watcher_eligible` (+ `launch_predates_boot`, `BOOT_RESTORE_SLACK_MS`), and the birth stamp it rests on is `shellinfo::{ShellLaunch, record_shell_launch, read_shell_launch}`.\r\n- **Source:** spt-core, flynn's spt-alchemy field report (2026-07-25) — a deterministic recipe, not a race. Class-(c) re-based BAROMETER W2 (releases#78, doyle's ruling on comment 5156810267).\r\n\r\n### 2.7 A node restart permanently strands every `persistent` shell — and the existing cascade test cannot see that it does  `[REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]`\r\n- **Failure:** the contract's own sentence says a `persistent` shell is **online whenever its owner endpoint is online**, and a node restart is the one case where that promise was never kept. A machine death breaks no link, so `close_shell` never runs and 2.6's stale record survives the reboot; when the owner comes back online, the wake cascade reads the **recorded** field, sees `online`, and skips the relaunch of a binary that has not existed since the previous boot. Nothing else corrects it: the instance is down forever, with no error anywhere, until a human happens to run `relink`. Every `persistent` instance on the node is stranded by one restart, together.\r\n- **The owner-facing surfaces HIDE the fault**, which is why it survived unreported for days: every display path (`shell list`, its `--json` twin, the shell-context render) routes through the ONE discovery seam, which **derives** status (2.6) and so correctly reports the binary as offline — while the only consumer whose decision matters, the wake cascade, reads the RECORDED field that no owner-facing surface shows. Absence of a visible symptom is therefore **not** evidence of absence of this fault, and a clean-looking context render must never be taken as proof the class did not occur. The correct falsifier is the **on-disk record**.\r\n- **The existing cascade test cannot fail on it — its SETUP SUPPRESSES THE FAILING ARM.** `rest_edges_cascade_shells_with_divergence` establishes its fixture by driving the suspend path, which itself writes the `offline` the code under test is supposed to encounter — so the restart shape (a record still saying `online` over a corpse, reached with **no** suspend edge) is unreachable from that fixture BY CONSTRUCTION. The test is not weak; it gives a true answer to the wrong question, so a reader asking \"is the cascade covered?\" is told yes. This is the third suppressed-arm instance found in one week, and the pattern is the same each time: **a fixture that establishes the precondition the code under test is supposed to establish.**\r\n- **Invariant:** a node restart must not strand a `persistent` instance. Two legs, and each is worthless alone. **(a) The record stops lying:** the daemon-side reconciler heals a recorded `online` that 2.6's derivation contradicts, writing `offline` — every cycle, **guarded on an actual change** (a heal that rewrote each tick would be a stream of identical writes and would destroy the record's mtime as a signal). **(b) The restart shape reaches a restoration path:** a once-per-daemon-generation boot sweep — and, for an owner that comes online after that sweep has run, an **owner offline→online edge** in the reconcile loop running the identical body (releases#228) — relaunches an instance when **all four** hold — the adapter section declares `persistent`; the…\n- **Regression shape (the row the suppressed fixture cannot express):** construct the restart shape **without a suspend edge** — an instance whose record says `online` over a corpse that predates boot, with an online owner — and assert it is restored. A fixture that suspends first re-creates the suppressed arm and proves nothing about this class.\r\n- **spt-core mapping:** `spt_daemon::shellwake::{heal_stale_online_records, restore_persistent_shells_at_boot, restore_persistent_shells_on_owner_online, OwnerOnlineEdge, launch_predates_boot, BOOT_RESTORE_SLACK_MS}`; `spt_store::proc::boot_instant_ms`; `spt_store::shellinfo::{ShellLaunch, record_shell_launch, read_shell_launch}` (the birth stamp, parked at the two production pid-write sites the launch-site census identified).\r\n- **Source:** spt-core, BAROMETER W2 (releases#78) — field specimen `liam/alchemy-0`, a genuinely stranded record: `online` over a dead pid, no launch stamp, days old.\r\n<!-- [doc->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS] -->\r\n\r\n---\r\n\r\n## 3. Lifecycle\r\n\r\n### 3.1 Ephemeral perch cleanup on every `ring` exit path\r\n- **Failure:** `ring` creates an ephemeral perch; early-exit paths (no-perch, empty-msg, timeout) skip cleanup → stale dirs accumulate.\r\n- **Invariant:** every code path that creates an ephemeral perch cleans it before exit; exception: if the caller already had an active perch, do not treat as ephemeral and do not clean up.\r\n- **spt-core mapping:** `ring` semantics carry; the daemon owns ephemeral-perch lifecycle, so a single guaranteed-cleanup (drop guard / RAII) is achievable in-process.\r\n- **Sister cite:** `src/owl/ring.rs:58-294`.\r\n\r\n### 3.2 Stale signoff sentinel must not kill a fresh start\r\n- **Failure:** a leftover `.claude/<id>-signoff.md` from a prior session is read by a fresh listener as a live signoff → immediate teardown.\r\n- **Invariant:** on every listener/daemon spawn, sweep stale signoff sentinels; signoff files are write-once per generation.\r\n- **spt-core mapping:** same sweep on daemon (re)start per hosted instance.\r\n- **Sister cite:** CHANGELOG v1.11.20; `src/owl/cleanup.rs:97`.\r\n\r\n### 3.3 Orphan teardown fires echo-commune BEFORE INIT_SIGNOFF\r\n- **Failure:** teardown delivers INIT_SIGNOFF without first saving the final context delta → Psyche signoff lacks the context-save summary.\r\n- **Invariant:** on orphan path, synchronously run the echo-commune (final delta) before composing INIT_SIGNOFF; skip only if the session_id is missing.\r\n- **spt-core mapping:** the daemon runs psyche/pulse loops in-process; ordering invariant identical.\r\n- **Sister cite:** `src/live/wrapper/orphan.rs:175-199`; tests A-H:333-565.\r\n\r\n### 3.4 A `ring` never adopts — so never deletes — a perch it did not create\r\n- **Failure:** 3.1 guards the LEAK direction (cleanup on every exit path). It says nothing about *whose* perch is being cleaned, and the two compose into a data-loss bug: `ring` decides \"the caller has no perch\" **ready-marker-first**, so an existing perch whose marker is momentarily down — busy turn, soft session-end, stale re-bind — reads *perchless*, `setup_ephemeral` clobbers the caller's real `info.json`, and 3.1's guaranteed cleanup then runs on the way out and **deletes** the ready marker, `info.json`, `spool.db` and the directory. Silent CLI-side fs ops, no daemon log, no console trace (field 2026-07-27 — emphasys lost a live endpoint and its spool to one timed-out ring).\r\n- **The mail half is a CONFIDENTIALITY failure, not only a durability one:** deletion is what happens on the way out, but while the ring *holds* the adopted perch its reply-wait `drain_one_at` consumes whatever is in the victim's spool and **renders it to the ringer as the reply**. Ringer == victim (the field case) is the degenerate, mild shape — anything drained reaches its intended recipient. Ringer != victim is the severe one: a third party rings a live agent and the **victim's inbox is drained into the ringer's output**, leaving no trace on either side once the dir is removed. Across a multi-machine subnet that is one agent reading another's mail. Bounded honestly for the record: in the field instance theft did **not** fire — emphasys's 66-second window …\n- **Self-camouflaging**, which is why it went unattributed for hours: the damage — perch gone, roster row ghosted — *mimics the stale-liveness condition people reach for `ring` to diagnose*. The field caller was probing a ghost-roster symptom and the probe manufactured a fresh one, so debuggers of stale liveness are disproportionately its victims and read the wreckage as more of the symptom they were chasing.\r\n- **Marker-down is not an edge case:** on hosted OMP endpoints the extension owns the listener and publishes busy/idle through `api state`, so marker-down is the **normal steady state during work** — marker-first misjudges healthy endpoints routinely.\r\n- **Invariant:** probe the perch **directory**, not the marker. A dir carrying a **record or a spool** is OCCUPIED — refuse to adopt it. Ambiguity resolves toward refusal, always: an **unreadable** record (corrupt/truncated `info.json`) or an unreadable dir counts as occupied, never as residue, because deletion is irreversible while refusal is recoverable. A pre-existing **empty** dir is refused too — it is not provably ring's own residue (a bringup mid-create owns an empty perch dir for a window, and deleting it is the same race in a different hat). Refusal never blocks and is never silent: the message is delivered, the call declines to block-wait (the reply lands on the caller's own listener), and the caller gets a **distinct** loud report — `RING_PERCH_EX…\n- **spt-core mapping:** `spt_msg::ring::{probe_perch, setup_ephemeral, ring}`; `RingOutcome::{PerchExists, StaleDir}` rendered by `cmd_ring`.\r\n- **Source:** spt-core field incident 2026-07-27 (emphasys endpoint + spool loss); ruled by doyle, with emphasys's mid-create-race amendment revoking the empty-dir self-heal.\r\n- **Amendment — the perch GC inherits this reasoning EXTENDED, never carved out (doyle, 2026-08-04, releases#109).** A second consumer now deletes perch directories: `spt endpoint gc` (3.5, `REQ-PERCH-GC-RESIDUE-PREDICATE`). Its authority is something `ring` never had — a **positive** statement that the directory carries **no endpoint record at all**, from the store that answers endpoint existence on a node (the owlery tree itself; the perch dir plus its parseable `info.json` IS the endpoint record). That authority does **not** outrank this section's occupancy rules, and the ruling says why: undeliverable-by-construction settles *deliverability*, not *value*, so a recordless dir carrying a **spool** is still refused — that spool is the only surviving copy of…\n<!-- [doc->REQ-HAZARD-RING-PERCH-ADOPTION] -->\r\n\r\n### 3.5 `is_perch_alive` is INVERTED on perch residue — a GC must never key on it\r\n<!-- [doc->REQ-HAZARD-PERCH-GC-LIVENESS-INVERSION] -->\r\n<!-- [doc->REQ-PERCH-GC-RESIDUE-PREDICATE] -->\r\n- **Failure:** the obvious way to write a perch garbage collector is \"reap what is not alive\" — `!spt_store::liveness::is_perch_alive(dir)`. On the population a GC actually walks, that predicate is **backwards**. `is_perch_alive` returns **true** for a directory with no `info.json` (`RawRead::Absent ⇒ true`, interim parity — an absent record means a wrapper-owned listener between polls), and **false** for a daemon-hosted endpoint whose `status` is `offline`. Measured on HFENDULEAM 2026-08-04 across 38 top-level perch dirs: **all 24 recordless residue directories read ALIVE, and all 6 offline REAL endpoints read DEAD.** A sweep keyed on `!is_perch_alive` would have spared every stray probe directory and deleted six resting agents' records — it does not under-…\n- **Why it is a hazard and not a bug in the resolver:** `is_perch_alive` is correct for its own job. Fail-toward-alive is the safe direction for DELIVERY (2.5) — the cost of a false \"alive\" is one message taking the spool path. It is the wrong direction for DELETION, where a false \"not alive\" is irreversible. The inversion is what happens when a predicate is reused across a polarity boundary it was never sized for.\r\n- **Invariant:** perch garbage collection classifies on record **presence** only (3.5's predicate: no `info.json` present on any read attempt ⇒ residue; NotFound kept distinct from every other I/O error, because absent is the only answer that authorizes deletion) and **never calls a liveness resolver** — not `is_perch_alive`, not `is_registry_entry_alive`, not a pid probe. The registry is likewise never asked: `clean_stale_entries` deletes dead-pid rows and `unregister_address` fires on an ordinary stop, so registry-absence is the normal steady state of every offline endpoint (4.3). The unit pinning this asserts **both** arms with the shipped resolver as the witness: residue reads alive and is reaped anyway; an offline endpoint reads dead and is kept anyway.…\n- **spt-core mapping:** `spt_store::perchgc::{sweep, PerchClass}`; `spt endpoint gc` renders it. The refusal/report contract is `REQ-PERCH-GC-REFUSAL-REPORT`; the occupancy rules it inherits are 3.4.\r\n- **Source:** measured by todlando on HFENDULEAM 2026-08-04 while answering doyle's step-1 predicate question on releases#109; ruled into the lane by doyle the same day.\r\n\r\n---\r\n\r\n## 4. Wire / transport\r\n\r\n### 4.1 Envelope HTML-entity codec ordering — `&amp;` decoded LAST\r\n- **Failure:** decoding `&amp;`-entity before the others double-decodes nested entities (`&amp;amp;lt;` → wrong result).\r\n- **Invariant:** ENCODE order amp→first … `<br>`→last; DECODE order `<br>`→first … amp→**last** (`&lt;`,`&gt;`,`&quot;`, then `&amp;`). One sole decode site (at the LLM/stdin boundary); the parser never decodes.\r\n- **spt-core mapping:** `spt-proto` owns the envelope grammar (public SDK, semver + wire-version). This codec contract is a copy-verbatim commodity item (ADR-0001) and a public-API conformance test.\r\n- **Sister cite:** `src/owl/poll.rs:1-73`; `src/common/envelope.rs`.\r\n- **CR-linesafety `[REQ-HAZARD-ENVELOPE-CR-LINESAFE]`:** the EVENT is LINE-FRAMED, so the codec must neutralize raw `\\r` too — `event_body_escape` folds CRLF/lone-CR to `\\n` (→`<br>`) **before** framing. **Failure (field, 2026-06-08):** a cross-node `spt send` from Windows (`echo` → CRLF) carried a raw `\\r` into the single-line envelope; the receiver terminal did a CR→column-0 overwrite (`</EVENT>` clobbered `<EVENT t`). `\\r` was never line-representable here, so normalizing it is robustness, not an ADR-0001 wire divergence (decoder + amp-last untouched). Belt-and-suspenders: `spt send`/`ring` trim stdin like `notify`.\r\n\r\n### 4.2 Two-slice envelope parser is panic-free and tolerant\r\n- **Failure:** malformed envelope (unclosed/misordered/nested tags) panics or drops output.\r\n- **Invariant:** tags case-sensitive, all optional; no tags → whole body to live slot; unclosed → None for that tag; out-of-order → both still extracted; nested unknown tags preserved verbatim; zero `unwrap` on parsed text.\r\n- **spt-core mapping:** `spt-proto` parser; property-test the robustness rules.\r\n- **Sister cite:** `src/common/envelope.rs:64-92`; tests 99-207.\r\n\r\n### 4.3 Registry stale-entry cleanup precedes lookup\r\n- **Failure:** sender resolves a dead process's stale TCP port → delivery to wrong/dead listener.\r\n- **Invariant:** clean stale entries (dead PID) before/at lookup; spool fallback is the safe path on TCP miss.\r\n- **spt-core mapping:** now spans the **subnet registry** (ADR-0003) — eventually-consistent across nodes. Cross-node staleness is expected; resolution policy (local → most-recent → `id@node`) must degrade to spool/relay fallback on stale hits, and never hard-fail on a stale remote entry.\r\n- **Sister cite:** `src/common/registry.rs:62-78`; `src/owl/send.rs`.\r\n\r\n### 4.4 Deferred rows survive poll drain\r\n- **Failure:** poll `drain_all` flushes a deferred (spool-only) message meant for a hook consumer → message lost.\r\n- **Invariant:** deferred rows are never flushed by the event-stream drain; only `drain_non_deferred_*` / `peek_all` touch them.\r\n- **Sister cite:** CHANGELOG v1.11.20; `src/common/spool.rs`. (See also 1.4.)\r\n\r\n### 4.5 Inbox legacy compat must not double-deliver\r\n- **Failure:** message surfaced via both spool (durable) and legacy inbox files → duplicate or racing delivery.\r\n- **Invariant:** spool is the sole read path at poll time; inbox is write-for-compat only and never read.\r\n- **spt-core mapping:** clean-room — likely drop the legacy inbox entirely. If kept for any compat, preserve \"never read at drain time.\"\r\n- **Sister cite:** `src/common/inbox.rs`.\r\n\r\n### 4.6 Addressable-id charset reserves the address delimiters\r\n<!-- [doc->REQ-HAZARD-ID-CHARSET] -->\r\n- **Failure:** a bare endpoint id that contains `:` or `@` (or a path separator / whitespace / control char) makes the canonical qualified address `[subnet:]id[@node]` (ADR-0006 / REQ-INST-10) ambiguous to parse, and lets a name smuggle into a perch directory path. Once permissive ids exist in the wild, tightening later needs a migration.\r\n- **Invariant:** every addressable id/name is validated to `[A-Za-z0-9_-]` + Hiragana/Katakana/CJK only, length `1..=64`, **at every creation seam** (`ready` start, `api bind`, `api listen`, `api worker-start`). `:` and `@` are permanently reserved as address delimiters; reads of existing perches are never re-validated. Enforce now (pre-M3/M4) so no permissive id-data accumulates.\r\n- **spt-core mapping:** `spt_proto::id::validate_endpoint_id`; called at the four creation seams. The existing Psyche (`<parent>-psyche`) / Worker (`<parent>-w<N>`) suffix scheme uses only `-` + alphanumerics, so composite ids validate.\r\n\r\n### 4.7 Concurrent SQLite openers must not fail with \"database is locked\"\r\n<!-- [doc->REQ-HAZARD-REGISTRY-CONCURRENT] -->\r\n- **Failure:** two endpoints on one machine open the same SQLite store at once (e.g. two `ReadyAgent::start` calls registering simultaneously) and one fails outright with `SQLITE_BUSY` / \"database is locked\" → spurious registration/spool failure. Surfaced as a parallel-test flake in `two_agents_exchange_message_tcp_and_spool`, but the bug is real concurrency, not test-only.\r\n- **Invariant:** `busy_timeout` is set **before** any lock-taking statement on every connection. Switching `journal_mode=WAL` takes a brief exclusive lock; with the default 0ms timeout it fails immediately under contention, so the pragma order is load-bearing: `Connection::open` → `busy_timeout` → `journal_mode=WAL` → `CREATE TABLE …`. WAL alone is insufficient (concurrent *writers* still serialize; they must *wait*, not error).\r\n- **spt-core mapping:** `spt_store::registry::open_registry` + `spt_store::spool::open_spool_at`; both set `busy_timeout=5000` first. Any future SQLite store (history Path B, instance registry) must follow the same ordering.\r\n\r\n### 4.8 Registry merge ordered by epoch, never wall-clock (red-team #8)\r\n<!-- [doc->REQ-HAZARD-REGISTRY-EPOCH-LEASE] -->\r\n- **Failure:** the per-subnet registry replicates `endpoint_id → [instances]` eventually-consistently across nodes. Under a partition or clock skew, a lagging node re-announces a stale `Active` for an endpoint that has actually gone `Offline`. If the merge ordered updates by wall-clock (or \"last write wins\"), the stale `Active` overwrites the newer `Offline` and resolution routes a message to a dead/wrong instance.\r\n- **Invariant:** the merge precedence key is a **per-node monotonic epoch counter** (`spt_store::epoch::EpochSource`, persisted, strictly increasing, NEVER wall-clock), compared version-vector style per `(endpoint_id, node)`: an incoming update wins **iff its epoch is strictly greater** than the stored one for that node; equal or lower is dropped as stale. So a newer `Offline` (higher epoch) can never be clobbered by a lagging `Active` (lower epoch), and an idempotent equal-epoch replay is a no-op. Wall-clock is at most a human tiebreaker hint inside a flagged conflict, never the ordering authority. The same epoch source unifies with the D6 sync-precedence concurrent-write detection (#7).\r\n- **spt-core mapping:** `spt_net::net::registry::SubnetRegistry::merge_instance` (the lease) + `spt_store::epoch::EpochSource` (the counter). Cross-node replication of the merge wires at D4; the merge seam is identical for local and wire-delivered updates. Chaos/two-host verification = D9.\r\n\r\n### 4.9 SQLite stores must create their parent dir — SQLite won't\r\n<!-- [doc->REQ-HAZARD-REGISTRY-DIR-CREATE] -->\r\n- **Failure:** `Connection::open` creates the database FILE but never its parent DIRECTORY. On a fresh home (first boot, fresh CI `_work` dir) a registry op that runs before any perch-creating op (`create_dir_all` side effects) fails `SQLITE_CANTOPEN` — \"unable to open database file …owlery\\.registry\". Timing-dependent: whichever code path touches the home first decides the outcome, so it surfaces as a parallel-test flake (bind-first tests losing the dir-creation race to perch-first tests). Bit the hfenduleam CI leg twice (2026-06-03/04, four spt-msg unit tests at once on the second strike) before being run to ground; a slow runner filesystem (AV scanning fresh dirs) widens the window but is not the cause.\r\n- **Invariant:** every SQLite store's open path `create_dir_all`s its parent dir itself, best-effort, before `Connection::open` — never relying on another subsystem having materialized the home first. (Mirrors the spool, which always did this; the registry didn't.)\r\n- **spt-core mapping:** `spt_store::registry::open_registry` (`create_dir_all(owlery)` before open). `spt_store::spool::open_spool_at` already creates its perch dir. Any future SQLite store must do the same — pair this with the 4.7 pragma ordering on every new store.\r\n\r\n### 4.10 Dead node identities leave immortal registry rows  `[REQ-HAZARD-REGISTRY-GHOST-ROWS]`\r\n<!-- [doc->REQ-HAZARD-REGISTRY-GHOST-ROWS] -->\r\n- **Failure:** the registry's only superseding mechanism is the per-`(endpoint_id, node)` epoch lease (4.8) — a row is replaced only by a newer row *from the same node*. When a node identity dies permanently (machine retired, or `node.key` regenerated so the \"node\" never speaks again), its rows are never superseded and never expire: they sit in the in-memory registries and the `identity/registry/<subnet>.json` snapshots forever. A bare-id send then resolves the same endpoint id on both the live and the dead identity and refuses with a **phantom `AcrossNodes` ambiguity** — unfixable by the user, because no qualifier reaches a node that no longer exists. Hit live in the M7 acceptance run (2026-06-06): gravity paired under two identities (09ef…, then 03854a… af…\n- **Invariant:** registry rows authored by a **silent** peer node decay: a node not *heard* (admitted inbound feed — the M7 D2 heard-map, REQ-SUBNET-1) within the eviction window (`registry_evict_after_ms`, default 300s ≈ 10 default pump cadences) has its rows **evicted** from every subnet registry, snapshots rewritten. Own rows never decay (the node always hears itself implicitly — it authors them each pump tick). Eviction is safe under the lease: v1 has **no transitive gossip**, so any future update for a node comes from that node itself, alive, re-inserting from its durable `EpochSource` within one cadence — there is no lagging third-party replay to mis-order against. A merely-offline node loses its rows after the window and reconverges on return; meanwhi…\n- **spt-core mapping:** `spt_net::net::registry::SubnetRegistry::evict_nodes` (model) + `spt_daemon::registryhost::RegistryHost::evict_silent_peers` (heard-map TTL) driven from the registry pump tick (`peerloop`). Trust rows are NOT auto-evicted (trust is a user decision; a stale trust row only costs dead dials) — pruning those is a separate verb.\r\n- **Source:** M7 acceptance run 2026-06-06 (DEFERRED.md \"Ghost registry row eviction\"); the AMBIGUOUS render fix rode along.\r\n- **Mesh note (ADR-0017, 2026-06-08):** the subnet mesh **preserves** this invariant rather than superseding it. \"No transitive gossip\" sharpens to **no transitive *row* gossip** — the mesh relays only the member *roster* (discovery), while registry **rows stay own-authored and are fetched directly** from each member over a handshake. So \"any future update for a node comes from that node itself, alive\" still holds and the eviction lease is untouched. (The plan's rejected alternative — signed transitive *row* relay — would have broken this; roster-only relay was chosen precisely to keep it.)\r\n\r\n### 4.11 Advertisement-epoch reset strands a node  `[REQ-HAZARD-EPOCH-RESET]`\r\n<!-- [doc->REQ-HAZARD-EPOCH-RESET] -->\r\n- **Failure:** a node whose advertisement-epoch counter resets (the durable `EpochSource` file lost/recreated) re-advertises with LOW epochs; peers hold a higher last-seen epoch for that `(endpoint, node)` lease and drop every fresh row as **stale** — the node advertises into a void until its counter outruns its own history. Nothing renders the cause: the node looks healthy locally, peers simply never update.\r\n- **Invariant (mitigation by construction, common case):** the common trigger — a full reinstall / identity regeneration — is covered by the **re-pair trust overwrite** (M8 decision 13, REQ-SUBNET-7): a completed ceremony presenting the same label + machine id evicts the superseded identity's trust AND registry rows on the seed-holder, and the peer-side epoch memory **dies with the deleted row** — the re-paired node's fresh epochs land on a clean lease. M8 acceptance 7 verifies this explicitly (the epoch sub-check).\r\n- **Residual (documented, guard deferred):** the narrow slice — epoch file lost while the node *identity* is kept (manual state surgery, partial restore from backup) — has no guard; it waits for a field hit before one is designed (M8 decision 24). `REQ-HAZARD-EPOCH-RESET` is minted inactive (TRACEABILITY rule 5) as the tracking hook. If hit: symptoms are one node's endpoints frozen-stale on every peer while its own views are fresh; recovery today is re-pairing the node (rides the common-case eviction above).\r\n- **spt-core mapping:** epoch mint = `spt_store::epoch::EpochSource` (`identity/epoch.json`); the lease = the per-`(endpoint, node)` epoch compare in `spt_net::net::registry`; the eviction that clears peer-side epoch memory = `registryhost::repair_evict_superseded` + `RegistryHost::consume_repair_evictions`.\r\n- **Source:** minted at M8 ratification (decision 24), recognized as a class during the 2026-06-07 pump diagnosis / re-pair overwrite design — not yet field-hit in its residual form.\r\n\r\n---\r\n\r\n## 5. Platform-specific\r\n\r\n### 5.1 Windows PID recycling false positives\r\n- **Failure:** recycled PID reads alive for the wrong process → orphan misclassification.\r\n- **Invariant:** anchor liveness on the stable parent/harness PID; minimal info.json for supervisor-owned perches; mtime grace window (≥60s) masks transient mismatches.\r\n- **spt-core mapping:** broker-held handles supersede PID polling for spt-hosted sessions; keep the grace window for harness-hosted.\r\n- **Sister cite:** `src/live/wrapper/orphan.rs:141-161`; `src/common/list_filter.rs:168-175`.\r\n\r\n### 5.2 Windows EBUSY on atomic rename\r\n- **Failure:** `fs::rename` fails while a handle is (recently) held → registry/marketplace update fails.\r\n- **Invariant:** tmp-write + atomic-rename with retry/backoff; best-effort side-fail; tolerate transient EBUSY.\r\n- **spt-core mapping:** all on-disk state writes (registry, trust store, spool checkpoints) use this pattern. Self-update binary swap on Windows especially.\r\n- **Sister cite:** CHANGELOG \"EBUSY\"; `src/common/owlery.rs` atomic_write.\r\n\r\n### 5.3 Git/subprocess timeout stamping\r\n- **Failure:** a hung subprocess (git on slow net) blocks the supervisor indefinitely.\r\n- **Invariant:** every metadata-producing subprocess has a timeout; timeout yields `None` + rate-limited stderr, never a hang.\r\n- **spt-core mapping:** generalize to all manifest-declared harness invocations (delegated commands, adapter updates) — timeouts mandatory.\r\n- **Sister cite:** `src/common/git.rs`.\r\n\r\n### 5.4 Windows UNC prefix in serialized paths\r\n- **Failure:** canonicalized `\\\\?\\C:\\...` serializes to `//?/C:/...` and fails `read_to_string`.\r\n- **Invariant:** strip the `\\\\?\\` UNC prefix after backslash→forward-slash conversion; serialized path attrs must be directly consumable.\r\n- **spt-core mapping:** any path crossing the wire (file-drop EVENTs, off-node file transfer per ADR-0003) needs canonical normalization at the `spt-proto` boundary.\r\n- **Sister cite:** `src/common/owlery.rs:377-384`.\r\n\r\n### 5.5 ConPTY withholds output until DSR is answered  `[REQ-HAZARD-CONPTY-DSR]`\r\n- **Failure:** a broker reading a ConPTY master sees only the 4-byte startup query `ESC [ 6 n` and then nothing — the child looks hung/silent but is producing output normally. ConPTY blocks all child stdout until the terminal answers the cursor-position query.\r\n- **Invariant:** every ConPTY reader auto-answers DSR (`ESC [ 6 n` → write `ESC [ 1;1 R`, or a real cursor position) on the PTY writer. Secondary: a ConPTY master does not EOF while the writer is held, so read loops drain on a thread and never gate exit on a blocking `read()`.\r\n- **spt-core mapping:** `spt-term` broker PTY reader (ADR-0004). Brand-new to spt-core — not in the sister project (it never hosted ConPTY directly).\r\n- **Source:** Spike #1 (`docs/spikes/SPIKE-01-broker-handoff.md`); reproduced with both a Rust child and `cmd.exe`.\r\n\r\n<!-- [doc->REQ-HAZARD-DETACHED-PIPE-INHERIT] -->\r\n### 5.6 Windows detached children inherit a captured caller's pipe  `[REQ-HAZARD-DETACHED-PIPE-INHERIT]`\r\n- **Failure:** a caller captures an `spt` invocation's output through a pipe (`Command::output()`, a harness hook reading the command). That `spt` process detach-spawns a **long-lived** child (the daemon via `ensure_running`; a shell binary via `spt shell spawn`). On Windows `CreateProcess` runs with `bInheritHandles = TRUE`, and the spt process's std handles — the caller's pipe write-ends — are inheritable by construction, so the immortal child inherits them even when its *own* stdio is `Stdio::null()`. The caller's pipe read never sees EOF: the capturing caller **hangs forever** (unix is immune — pipe fds are `CLOEXEC`). Paid twice: daemon spawn (guarded at D4a-era `spawn_detached`), then again at M5-D3e when the mock-shell E2E hung `spt shell spawn` for h…\n- **Invariant:** every detach-spawn of a long-lived child inherits only handles it **enumerates**. The default and overwhelmingly common form of that is `bInheritHandles = FALSE` (`spt-daemon::daemon::detached_no_inherit`) — zero handles flow, whatever the pipe's depth in the ancestry. Stripping `HANDLE_FLAG_INHERIT` from the spawner's *std* handles is NOT sufficient: a grandparent capture's pipe sits in the handle table as a stray inheritable handle and still flows through every **bare** `bInheritHandles = TRUE` hop (the first guard shipped that way and was wedged by exactly this — a daemon spawned three layers deep held the pwsh-level pipe of the CI/test harness).\r\n- **The one permitted TRUE (amended 2026-07-26, RESIDENT-SERVICE W1):** a spawn that must capture the child's output may pass `bInheritHandles = TRUE` **only** paired with an explicit `STARTUPINFOEX` + `PROC_THREAD_ATTRIBUTE_HANDLE_LIST` allowlist — the Win32 contract honors the list only under TRUE, so the earlier phrasing (\"flipping TRUE is refused\") forbade its own mechanism and was unimplementable. The hazard was always *unenumerated* inheritance, and the list **is** the enumeration: only the named handles flow, so a grandparent's pipe cannot reach the child at any depth — strictly stronger than what FALSE achieves incidentally. A **bare** TRUE, with no attribute list, remains REFUSED. Today the only such path is a supervised `[service]` spawn with start…\n- **spt-core mapping:** `spt-daemon::daemon::spawn_detached` (the daemon) and `spt-daemon::shellhost::launch_shell` (the relay-receipt shell binary). Any future long-lived detached spawn (manifest-template children included) must use the same no-inherit spawn, or the enumerated form above where it genuinely must capture output.\r\n- **Source:** spt-core, M5-D3e (`shell_e2e.rs` hang, 2026-06-04, twice — once per guard generation); Rust `Command` restricts *its own* created stdio handles but a parent's inheritable handle table still flows.\r\n\r\n### 5.7 Elevated commands spawn the daemon with the wrong token  `[REQ-HAZARD-ELEVATED-DAEMON-SPAWN]`\r\n<!-- [doc->REQ-HAZARD-ELEVATED-DAEMON-SPAWN] -->\r\n- **Failure:** membership-implies-reachability made *every* `spt` invocation a potential daemon spawner (`ensure_running`), including the elevation-gated ones (`subnet create`/`join`, REQ-SUBNET-4). The spawned daemon inherits the spawner's token. **Windows:** an elevated `subnet create` auto-starts an ELEVATED daemon whose named pipes deny unelevated clients — every subsequent unelevated `spt` reads \"not running\", tries to spawn its own daemon, and dies on bind Access-denied; the user had to taskkill (hit live, M7 acceptance 2026-06-06). **Linux:** a sudo'd command spawns a root daemon and/or root-owned state — and because sudo flips `$HOME`, the daemon can mint a *different node identity* in root's universe (the very key-flip that produced the 4.10 ghost r…\n- **Invariant:** the daemon **always runs unelevated in the invoking user's universe**, regardless of which command spawns it. Two enforcement points sharing one seam: (a) `spawn_detached` de-elevates the child — Windows: the UAC **linked token** (`TokenLinkedToken` → `DuplicateTokenEx` → `CreateProcessWithTokenW`; inherits no handles, so 5.6 holds by construction); Linux: drop to `SUDO_UID`/`SUDO_GID` with `$HOME`/`$USER`/`$LOGNAME` reset to the invoking user's (passwd lookup); (b) a `Daemon::run` entry guard catches a *directly* elevated `spt daemon` — Linux drops privileges in-process before touching any state; Windows respawns de-elevated and exits. When no unelevated identity exists to drop to (UAC disabled, genuine root login, SYSTEM), the daemon runs …\n- **spt-core mapping:** `spt-daemon::deelevate` (the OS-split seam) consumed by `daemon::spawn_detached` + the `Daemon::run` entry guard. The fuller Linux elevation model (install symlink + default-account election) is deferred (DEFERRED.md, M8).\r\n- **Source:** M7 acceptance run 2026-06-06 (DEFERRED.md \"Non-admin daemon spawn\"); interim field rule was \"bring the daemon up unelevated FIRST\".\r\n\r\n<!-- [doc->REQ-HAZARD-CHILD-CONSOLE-FLASH] -->\r\n### 5.8 Console children of the console-less daemon flash visible windows  `[REQ-HAZARD-CHILD-CONSOLE-FLASH]`\r\n- **Failure:** the daemon runs DETACHED (no console, 5.6/`detached_no_inherit`). Any console-subsystem child it spawns (`git`, `taskkill`, manifest hook commands) gets a **fresh conhost with a visible window** — piped/null stdio does NOT prevent it. Field shape: the 60s sync pump's two git calls (`for-each-ref` + `rev-parse`) flashed two blank windows per minute on the user's desktop (2026-06-06).\r\n- **Invariant:** every short-lived console child spawned from daemon-reachable code sets `creation_flags(0x0800_0000)` (`CREATE_NO_WINDOW`). Long-lived detached children use `detached_no_inherit` (already `DETACHED_PROCESS | CREATE_NO_WINDOW`); de-elevated spawns use `CREATE_NEW_CONSOLE + SW_HIDE` (5.7 — `CreateProcessWithTokenW` rejects `CREATE_NO_WINDOW`, error 87).\r\n- **Test seam caveat:** window-absence is unobservable from a consoled test runner — the child inherits the runner's console and never creates a window, flag or no flag. Unit coverage asserts the flagged spawn still works (the error-87 \"flag combo breaks spawn\" regression class); window-absence was verified live by process-watch capture.\r\n- **spt-core mapping:** `spt-store::gitrun::run_git` (every BranchStore/ContextStore git call), `spt-daemon::shellhost::kill_shell_pid` (taskkill), `spt-runtime::run_bounded_command` (manifest hook commands), `spt-runtime::ManifestRuntime::command_for` (the one shared builder behind `spawn_session` + `run_bounded_stdin` — the notif pump's `spawn_notif_command` and the live agent's psyche/echo/turn spawns), `spt-daemon::shellwake` (already guarded). The flag lives in each shared builder, not per call site, so the invariant holds for every ManifestRuntime spawn by construction.\r\n- **Source:** spt-core field bug, 2026-06-06 — two blank windows flashing every 60 seconds on a desktop workstation, caught by process-spawn watcher (git.exe parent=spt daemon, conhost.exe child each).\r\n\r\n### 5.9 `Instant - Duration` underflow-panics on a freshly-booted host  `[REQ-HAZARD-INSTANT-UNDERFLOW]`\r\n- **Failure:** `Instant::now() - Duration::from_secs(N)` panics `overflow when subtracting duration from instant` when the process's monotonic clock is younger than `N` — i.e. the host booted less than `N` ago. The peer pump primed its cadence legs with `Instant::now() - 86_400s` to mean \"everything due now\"; on a Windows runner with sub-24h uptime the pump thread panicked at startup, so the subnet never converged (CI `pump_and_dispatch_self_drive_the_subnet` failed, run 27082417706). It is *environment-conditional* — green on any host up longer than the offset, red below it — so it slips local dev and only bites a fresh CI box or a just-rebooted machine.\r\n- **Invariant:** NEVER compute an instant in the past by subtracting from `Instant::now()`. Represent \"never run / due now\" as `Option<Instant> = None` and gate on forward `now.duration_since(past)` only (`peerloop::due`). No backward instant arithmetic anywhere in scheduling.\r\n- **Test seam caveat:** the convergence E2E only reproduces on a sub-offset-uptime host (it passed everywhere with >24h uptime). The deterministic guard is the `due(None, ..)`/`due(Some(now), ..)` unit on the extracted gate — it asserts first-tick-due with zero instant subtraction, independent of host uptime.\r\n- **spt-core mapping:** `spt-daemon::peerloop::due` (the sole cadence gate behind `due_reg`/`due_notif`/`due_sync`/`due_upd`); cadence legs are `Option<Instant>` seeded `None`.\r\n- **Source:** spt-core CI failure, 2026-06-07 — Windows runner `hfenduleam` (just booted) panicked the peer pump at the v0.1.1 release gate.\r\n\r\n### 5.10 `sudo spt` dead-ends on a user-local install (secure_path)  `[REQ-HAZARD-SUDO-SECURE-PATH]`\r\n- **Failure:** the elevation-gated commands (`subnet create` / `subnet join` / `show-code`) refuse when unelevated and tell the user to \"run as administrator / root\". The user types the obvious `sudo spt subnet create FOO` → `sudo: spt: command not found`. `spt` is a user-local install (`~/.local/bin`, `~/.cargo/bin`), and sudo's `secure_path` (a `/etc/sudoers` default) does NOT include those dirs, so a bare command name doesn't resolve under sudo. The guidance is a trap: it names an action that cannot work for the common install shape. Field-hit on KITSUBITO at the v0.1.1 ship.\r\n- **Invariant:** elevation guidance on Unix emits the binary's **absolute path** under sudo — `sudo /home/u/.local/bin/spt subnet create FOO` — reconstructed from `current_exe()` + the real argv and shell-quoted. An absolute program path is executed directly; `secure_path` only governs bare-name PATH lookup, so the absolute form always resolves. On an interactive Unix TTY the command auto-elevates (re-execs itself under sudo, the elevated child does the work and `main` de-elevates back); non-interactive or sudo-absent falls back to printing the runnable hint. Never emit a bare-name elevation instruction.\r\n- **Companion UX:** the post-de-elevation `DEELEVATED: running as uid N` line is internal state-safety noise — omit it from the user-facing CLI path (it confused the same field user). The detached daemon's own de-elevation log line is fine (it lands in the daemon log, not the terminal).\r\n- **Test seam caveat:** the sudo re-exec needs a real `sudo` + TTY (not hermetic). The deterministic guard is the pure `elevation::sudo_argv` / `print_hint_command` (assert an absolute exe path, never a bare name, + shell-quoting on the printed line) and the `decide_elevation_path` matrix (which picks inline-sudo only on an interactive Unix TTY); the exec leg is manual/kitsubito-verified.\r\n- **spt-core mapping:** `spt::elevation::{sudo_argv, print_hint_command, decide_elevation_path}` (pure — generalized from the M12-W4 self-elevation seam, 5.11), `spt::cli::{try_auto_elevate, with_elevation_hint}` wired into `cmd_subnet_create` / `cmd_subnet_join` / `cmd_subnet_show_code`; `spt::main` de-elevation drop silenced.\r\n- **Source:** spt-core field report, 2026-06-07 — `reavus@KITSUBITO`, `spt` in `~/.local/bin`; the absolute-path `sudo` invocation was confirmed working before the fix landed.\r\n\r\n### 5.11 Self-elevating re-launch must re-run verbatim, never widen / inject / loop  `[REQ-HAZARD-SELF-ELEVATE]`\r\n- **Failure class:** a privilege-gated command (`subnet create` / `join` / `show-code`) self-elevates by re-launching itself with privilege (Windows UAC `runas`, Linux `pkexec` / a terminal-emulator `sudo`, or inline `sudo`). A careless re-launch is a security hole: widening the privilege scope (adding args), resolving the binary by a bare name (a PATH/`secure_path` hijack runs an attacker's `spt`), interpolating a crafted arg into a shell string (`sh -c \"… $id …\"` injects a second command), or re-elevating the already-elevated child (an infinite UAC/polkit loop). The user's UAC/polkit/sudo prompt is the ONLY consent gate — the mechanism must never bypass or widen it.\r\n- **Invariant:** self-elevation re-runs the **EXACT** original invocation with the binary's **ABSOLUTE** exe path — never adding/altering args, never a PATH-resolved bare name, never a shell-interpolated string. Every launcher passes an **argv array** (`Command::new(prog).args([...])`, never `sh -c`); the Windows `ShellExecuteW` params string (which is inherently one string) MSVC-quotes each verbatim arg so `CommandLineToArgvW` round-trips it as a single token. The elevated child drops state back to the user (composes with the 5.7 de-elevation) and **never re-elevates**: `decide_elevation_path` returns `AlreadyElevated` whenever the process is `Elevated`, on every OS (loop-safety). The unprivileged parent never pipes/captures the elevated child's stdout acro…\n- **Test seam caveat:** the real launch needs a UAC/polkit/sudo prompt (not hermetic) — manual-verify. The deterministic guards are the pure `decide_elevation_path` matrix (loop-safety: `AlreadyElevated` on every os; the os×env path order) and the argv builders (`sudo_argv` / `pkexec_argv` / `terminal_argv` assert absolute-exe + verbatim args + array; `windows_runas_params` asserts MSVC-quoting with no `cmd /c` interpolation; the crafted-arg test asserts a shell-metachar arg stays one element / one quoted token).\r\n- **spt-core mapping:** `spt::elevation::{decide_elevation_path, sudo_argv, pkexec_argv, terminal_argv, windows_runas_params, print_hint_command, ElevatePath}` (pure), `spt::cli::{try_auto_elevate, launch_uac_window, pause_elevated_console_if_fresh, program_on_path, first_terminal_emulator}` (impure launchers) wired into `cmd_subnet_create` / `cmd_subnet_join` / `cmd_subnet_show_code`. Companions: 5.10 (the Unix abs-path-under-sudo facet) and 5.7 (the elevated child's de-elevation drop, which this composes with).\r\n- **Source:** M12-W4 design (subnet QR + self-elevating window), doyle ruling `M12-W4-RULING.md` Q6 — a privilege-escalation feature carries a mandatory hazard REQ.\r\n\r\n<!-- [doc->REQ-HAZARD-WIN-PTY-PROGRAM-RESOLVE] -->\r\n### 5.12 Native-PTY spawn of a bare program runs the wrong (non-PE) file on Windows  `[REQ-HAZARD-WIN-PTY-PROGRAM-RESOLVE]`\r\n- **Failure:** `portable-pty`'s ConPTY spawn resolves a bare program name with a `which` that takes the FIRST `PATH` match. A node/npm CLI installs as BOTH an extensionless shebang shim (`ccs`, for Git Bash) and a Windows launcher (`ccs.cmd`) in the same dir; portable-pty picks the extensionless `ccs`, and `CreateProcessW` then tries to execute that non-PE file and fails with **os error 193** (\"%1 is not a valid Win32 application\"). Live failure: `spt endpoint run claude-spt:ccs` → `CreateProcessW C:\\nvm4w\\nodejs\\ccs` 193 (operator, 2026-06-16). The same bites any harness/shell whose `[session.self]`/`[shell].spawn` names a `.cmd`/`.bat`/`.ps1`-backed command — `CreateProcessW` cannot execute a batch or PowerShell script directly.\r\n- **Invariant:** spt-term resolves the program ITSELF before handing it to `CommandBuilder`, bypassing portable-pty's `which`. A bare name is searched over `PATH` × `PATHEXT` (whose default order already prefers `.EXE`/`.COM` over `.BAT`/`.CMD`), then an extensionless fallback. A non-PE target is wrapped in its interpreter: `.cmd`/`.bat` → `cmd.exe /d /c <path>`, `.ps1` → `powershell -NoProfile -File <path>` (the wrap args precede the caller's args); a real executable spawns directly; an unresolvable name passes through unchanged (never makes a working case worse). Unix is a passthrough — `execve` honours a shebang on an extensionless script. Applied at the ONE `CommandBuilder` chokepoint (`PtySession::spawn_program_in`), so every broker harness + shell spaw…\n- **spt-core mapping:** `spt_term::winprog::{resolve_for_pty, resolve_in}` (the pure PATHEXT-precedence kernel + the Windows env wiring), wired into `spt_term::pty::PtySession::spawn_program_in`. Unit: `resolve_in` precedence (`.cmd`-over-shim, `.exe`-direct, explicit-extension, path-order, passthrough) [`winprog.rs`].\r\n- **Source:** field diagnosis 2026-06-16 (operator dogfood, `claude-spt:ccs` bringup) — doyle.\r\n\r\n<!-- [doc->REQ-HAZARD-PERCH-RECORD-POWER-LOSS] -->\r\n### 5.13 Atomic write leaves data un-synced before the rename → NUL zero-fill on power loss  `[REQ-HAZARD-PERCH-RECORD-POWER-LOSS]`\r\n- **Failure:** `atomic_write_bytes` was `fs::write(tmp)` + `rename(tmp, path)` with no `fsync`. The rename's directory **metadata** is journaled durable, but the tmp file's **data blocks** are still in the page cache. A hard reset (power loss, forced reboot) between the two flushes lands the rename but loses the data → the file reappears at its **full length filled with NUL**. Field incident: after a machine restart `owlery/hall-a/info.json` was 360 bytes of all-NUL (the nested psyche record 125 bytes all-NUL); `Get-Content` renders blank, `Format-Hex` shows the zero-fill. A wiped-but-present record then read as a live ONLINE endpoint for days (compounded by 5.14).\r\n- **Invariant:** for records whose loss is **unrecoverable**, flush data to stable storage **before** the rename publishes the name — `File::create(tmp)` → `write_all` → `sync_all()` → `rename_with_retry`. Then a crash yields the complete old file OR the complete new file, never a NUL husk. (A file of correct length that is ALL-NUL is the diagnostic tell of this class — a non-fsync'd write caught by a hard reset.) **Durability is SCOPED, not blanket.** `fsync`ing *every* atomic write serializes a durable flush behind each of the ~26 `atomic_write` callers — and for `info.json` the flush lands under the per-perch `.info.lock` (the **W1b lock-across-fsync wedge shape**, §6-era) — which stalls daemon bringup 4–6× so the endpoint misses its ONLINE window. Scope …\n- **spt-core mapping:** `spt_store::atomic::atomic_write_bytes_durable` / `atomic_write_string_durable` are the opt-in durable siblings; the default `atomic_write_bytes` / `atomic_write_string` stay non-durable. Durable callers are exactly `spt_store::info::write_info` (the perch record — the hall-a surface), `spt_store::nodeid` (the node seed — corrupt is node-bricking, never regenerated), and `spt_daemon::machineid` (minted-once machine id — a NUL husk would silently re-mint a different id). Reader-side handling of an already-corrupt record is the sibling 5.14. **Canary:** `attach_wedge_e2e` is the regression guard — a blanket fsync fails it (bringup misses the 20s ONLINE budget); keep durability scoped so it stays green.\r\n…\n\n<!-- [doc->REQ-HAZARD-CORRUPT-PERCH-COHERENCE] -->\n…\n\n…\n<!-- [doc->REQ-HAZARD-EMPOWER-SESSION-RESURRECT] -->","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120,121,122,123,124,125,126,127,128,129,130,131,132,133,134,135,136,137,138,139,140,141,142,143,144,145,146,147,148,149,150,151,152,153,154,155,156,157,158,159,160,161,162,163,164,165,166,167,168,169,170,171,172,173,174,175,176,177,178,179,180,181,182,183,184,185,186,187,188,189,190,191,192,193,194,195,196,197,198,199,200,201,202,203,204,205,206,207,208,209,210,211,212,213,214,215,216,217,218,219,220,221,222,223,224,225,226,227,228,229,230,231,232,233,234,235,236,237,238,239,240,241,242,243,244,245,246,247,248,249,250,251,252,253,254,255,256,257,258,259,260,261,262,263,264,265,266,267,268,269,270,271,272,273,274,275,276,277,278,279,280,281,282,283,284,285,286,287,288,289,290,291,292,293,294,295,296,297,298,299,300,null,302,303,null,368,null,1091]},"fileSize":356704,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\docs\\KNOWN-HAZARDS.md"},"truncation":{"direction":"middle","truncatedBy":"middle","totalLines":307,"totalBytes":61594,"outputLines":197,"outputBytes":40407,"maxBytes":40960,"headRange":{"start":1,"end":116},"tailRange":{"start":228,"end":307},"elidedLines":111,"elidedBytes":21209,"artifactId":"64","nextOffset":301},"limits":{"columnTruncated":{"maxColumn":768}}}},"isError":false,"timestamp":1788999494909}}
{"type":"message","id":"41a10f9b","parentId":"da45edc7","timestamp":"2026-09-10T00:18:18.045Z","message":{"role":"toolResult","toolCallId":"call_laaz7ZoODf1WB2QxJI0MJVAW|fc_0dde72effda0ad21016aa1f746a5a887d2a398b94370bff234","toolName":"bash","content":[{"type":"text","text":"CHANGELOG.md | 11 +++++++++++\ncrates/spt-daemon/src/conn.rs | 264 +++++++++++++++++++++---\ncrates/spt-daemon/src/livehost.rs | 124 +++++++++++++++++++++---\ndocs/RCA-FLEET-DAEMON-14444.md | 73 ++++++++++++++++++++++++\ntraceable-reqs.toml | 9 +++++++++\n5 files changed, 425 insertions(+), 56 deletions(-)\n\n--- Changes ---\n\nFile: CHANGELOG.md\n  @@ -10,6 +10,17 @@ breaks something, or changes the observable behavior of existing surfaces\n  +## [Unreleased]\n  +\n  +### Fixed\n  +\n  +<!-- [doc->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED] -->\n  +- Daemon logs summarize healthy IPC connection starts and closes once per\n  […5 changed lines elided…]\n\nFile: crates/spt-daemon/src/conn.rs\n  @@ -61,7 +61,7 @@\n  -use std::sync::{Arc, Condvar, Mutex, MutexGuard, OnceLock};\n  +use std::sync::{Arc, Condvar, LazyLock, Mutex, MutexGuard, OnceLock};\n  @@ -99,12 +99,107 @@ pub(crate) fn init_log_anchor() {\n  -    let wall_ms = SystemTime::now()\n  -        .duration_since(UNIX_EPOCH)\n  -        .map(|d| d.as_millis())\n  -        .unwrap_or(0);\n  -    let mono_ms = MONO_ANCHOR.get_or_init(Instant::now).elapsed().as_millis();\n  -    format!(\"wall_ms={wall_ms} mono_ms={mono_ms}\")\n  […101 changed lines elided…]\n  @@ -144,8 +239,9 @@ struct Inner {\n  -    /// Latch: the once-per-conn `write-start` lifecycle record fired.\n  -    first_write_logged: AtomicBool,\n  +    /// Once-per-conn original first committed write time; retained on failures.\n  +    first_write: OnceLock<LogTime>,\n  +    lifecycle_log: Arc<LifecycleLog>,\n  @@ -223,7 +319,13 @@ impl Inner {\n  -        format!(\"conn={} {} {}\", self.id, facts, log_stamp())\n  +        match self.first_write.get() {\n  +            Some(start) => format!(\n  +                \"conn={} {} {} first_write_wall_ms={} first_write_mono_ms={}\",\n  +                self.id, facts, log_stamp(), start.wall_ms, start.mono_ms\n  +            ),\n  […2 changed lines elided…]\n  @@ -261,18 +363,17 @@ impl Inner {\n  -    /// One BOUNDED per-conn lifecycle record (leg d, doyle-confirmed\n  -    /// UNCONDITIONAL): `write-start` (first write only), `transport-close`\n  -    /// (drop), `writer-exit` / `*-attach` / `*-replaced` (emitted by the\n  -    /// broker's sink machinery through [`BrokerConn::lifecycle_event`]). The\n  -    /// timeout/cancel leg is the retirement record itself.\n  +    /// Per-conn exceptional and role-transition records remain unconditional.\n  […7 changed lines elided…]\n  @@ -287,10 +388,10 @@ impl Inner {\n  -            eprintln!(\n  +            self.lifecycle_log.emit(format_args!(\n  -            );\n  +            ));\n  @@ -424,11 +525,17 @@ impl BrokerConn {\n  +        static LOG: LazyLock<Arc<LifecycleLog>> = LazyLock::new(|| Arc::new(LifecycleLog::new()));\n  +        Self::with_log(half, bound, Arc::clone(&LOG))\n  +    }\n  +\n  +    fn with_log(half: SendHalf, bound: Duration, lifecycle_log: Arc<LifecycleLog>) -> Self {\n  -            first_write_logged: AtomicBool::new(false),\n  […2 changed lines elided…]\n  @@ -539,12 +646,12 @@ impl BrokerConn {\n  -        // Once-per-conn `write-start` lifecycle record (leg d): marks the conn's\n  -        // first committed write, so a conn that later retires can be correlated\n  -        // to when it went active (fresh-carrier churn shows as open→start→retire\n  -        // triplets with fresh ids).\n  -        if !inner.first_write_logged.swap(true, Ordering::AcqRel) {\n  -            inner.lifecycle(\"write-start\", \"\");\n  […6 changed lines elided…]\n  […3 hunks elided…]\n\nFile: crates/spt-daemon/src/livehost.rs\n  @@ -569,7 +569,7 @@ pub fn resume_restart_orphaned_endpoints(\n  -    let Some(live) = query_live_session_endpoints() else {\n  +    let Some(live) = SessionPoll::default().query(&crate::endpoint::broker_socket_name()) else {\n  @@ -995,20 +995,57 @@ pub fn reconcile_hosted_liveness(owlery: &Path, live_sessions: &BTreeSet<String>\n  -/// Query the broker for the set of endpoint ids it currently hosts a session for\n  -/// (`KIND_SESSIONS`) — the B2 pull signal. `None` when the broker is unreachable\n  -/// (the caller then SKIPS the offline pass this tick rather than mass-offlining).\n  -fn query_live_session_endpoints() -> Option<BTreeSet<String>> {\n  -    let mut brain = Brain::cold_start(&crate::endpoint::broker_socket_name(), now_ms()).ok()?;\n  -    let reply = brain.sessions().ok()?;\n  […59 changed lines elided…]\n  @@ -1333,11 +1370,12 @@ pub fn spawn_live_host(stop: Arc<AtomicBool>, reason: StartReason) -> JoinHandle\n  +        let mut session_poll = SessionPoll::default();\n  -            // reap-driver in ISOLATION. This loop's own `query_live_session_endpoints()`\n  +            // reap-driver in ISOLATION. This loop's own `SessionPoll::query`\n  @@ -1371,7 +1409,7 @@ pub fn spawn_live_host(stop: Arc<AtomicBool>, reason: StartReason) -> JoinHandle\n  -            if let Some(live) = query_live_session_endpoints() {\n  +            if let Some(live) = session_poll.query(&crate::endpoint::broker_socket_name()) {\n  @@ -1397,6 +1435,58 @@ mod tests {\n  +    // [unit->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n  +    #[test]\n  +    fn session_poll_reuses_healthy_carrier_and_reopens_after_query_error() {\n  +        use crate::codec::{read_frame, write_frame};\n  +        use crate::frame::{Envelope, Role};\n  +        use crate::transport::{recv_hello, DaemonTransport, LocalSocketTransport};\n  […46 changed lines elided…]\n\nFile: docs/RCA-FLEET-DAEMON-14444.md\n  @@ -69,3 +69,76 @@ per 2 MB of log, under `[REQ-CONN-POISON-ATTRIBUTION]`, is its own finding.\n  +\n  +## releases#286 — stderr follow-up (2026-09-09)\n  +\n  +<!-- [doc->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED] -->\n  +\n  +The historical 303 MB incident above remains the reported baseline, not a\n  […67 changed lines elided…]\n\nFile: traceable-reqs.toml\n  @@ -2569,6 +2569,14 @@ id = \"REQ-CONN-POISON-ATTRIBUTION\"\n  +# releases#286 narrowly supersedes leg (d)'s unconditional healthy start/close\n  +# lines. Failure/poison, role transitions and original first-write time remain\n  +# attributable per physical conn; healthy counts are interval totals instead.\n  +[[requirements]]\n  +id = \"REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED\"\n  +title = \"Healthy broker IPC starts and closes aggregate once per monotonic interval across connections, independently of request volume; poison, failure and ro…[+405]\n  […2 changed lines elided…]\n  @@ -4860,6 +4868,7 @@ name = \"message-format-identity\"\n  +  \"REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED\",\n[raw output: artifact://65]\n\n\nWall time: 3.26 seconds"}],"details":{"timeoutSeconds":300,"wallTimeMs":3256.9294000000227},"isError":false,"timestamp":1788999498045}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_qPHAxmgOQ986O8RTWTPcS7JZ|fc_0dde72effda0ad21016aa1f751053887d2bf80cde1dc1bbbd3","toolName":"read","startedAt":"2026-09-10T00:18:25.092Z","args":{"path":"artifact://65:raw"},"intent":"Reading raw patch"},"id":"78d5c97c","parentId":"41a10f9b","timestamp":"2026-09-10T00:18:25.092Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_eyb1aSoUOsXzhRujKrcoBRwI|fc_0dde72effda0ad21016aa1f751054c87d2865b2eaecd26c2b6","toolName":"read","startedAt":"2026-09-10T00:18:25.092Z","args":{"path":".worktrees/286-churn/AGENTS.md"},"intent":"Reading worktree agent rules"},"id":"2a57895c","parentId":"78d5c97c","timestamp":"2026-09-10T00:18:25.092Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_c4xODNE00tUVz9lnM2AclZLs|fc_0dde72effda0ad21016aa1f751055487d2aa1203ec2e0ccda4","toolName":"read","startedAt":"2026-09-10T00:18:25.092Z","args":{"path":"AGENTS.md"},"intent":"Reading root agent rules"},"id":"93ac9cfd","parentId":"2a57895c","timestamp":"2026-09-10T00:18:25.092Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_I7Tc7wlUioMWDMzQLLS5aRsh|fc_0dde72effda0ad21016aa1f751055c87d2af276890e6c623d3","toolName":"grep","startedAt":"2026-09-10T00:18:25.093Z","args":{"path":".worktrees/286-churn/AGENTS.md;.worktrees/286-churn/docs/KNOWN-HAZARDS.md"},"intent":"Locating lifecycle standards"},"id":"08789455","parentId":"93ac9cfd","timestamp":"2026-09-10T00:18:25.093Z"}
{"type":"message","id":"4595a942","parentId":"08789455","timestamp":"2026-09-10T00:18:25.093Z","message":{"role":"assistant","content":[{"type":"toolCall","id":"call_qPHAxmgOQ986O8RTWTPcS7JZ|fc_0dde72effda0ad21016aa1f751053887d2bf80cde1dc1bbbd3","name":"read","arguments":{"i":"Reading raw patch","path":"artifact://65:raw"},"intent":"Reading raw patch"},{"type":"toolCall","id":"call_eyb1aSoUOsXzhRujKrcoBRwI|fc_0dde72effda0ad21016aa1f751054c87d2865b2eaecd26c2b6","name":"read","arguments":{"i":"Reading worktree agent rules","path":".worktrees/286-churn/AGENTS.md"},"intent":"Reading worktree agent rules"},{"type":"toolCall","id":"call_c4xODNE00tUVz9lnM2AclZLs|fc_0dde72effda0ad21016aa1f751055487d2aa1203ec2e0ccda4","name":"read","arguments":{"i":"Reading root agent rules","path":"AGENTS.md"},"intent":"Reading root agent rules"},{"type":"toolCall","id":"call_I7Tc7wlUioMWDMzQLLS5aRsh|fc_0dde72effda0ad21016aa1f751055c87d2af276890e6c623d3","name":"grep","arguments":{"i":"Locating lifecycle standards","pattern":"conn|poison|lifecycle|PID|pid|test|allocat|panic|lock","path":".worktrees/286-churn/AGENTS.md;.worktrees/286-churn/docs/KNOWN-HAZARDS.md","case":false,"gitignore":false,"skip":null},"intent":"Locating lifecycle standards"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":17286,"output":178,"cacheRead":14720,"cacheWrite":0,"totalTokens":32184,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1788999498100,"responseId":"resp_0dde72effda0ad21016aa1f74a585487d294949ab3743e52ab","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"fc_0dde72effda0ad21016aa1f751053887d2bf80cde1dc1bbbd3","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading raw patch\",\"path\":\"artifact://65:raw\"}","call_id":"call_qPHAxmgOQ986O8RTWTPcS7JZ","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"read","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f751054c87d2865b2eaecd26c2b6","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading worktree agent rules\",\"path\":\".worktrees/286-churn/AGENTS.md\"}","call_id":"call_eyb1aSoUOsXzhRujKrcoBRwI","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"read","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f751055487d2aa1203ec2e0ccda4","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading root agent rules\",\"path\":\"AGENTS.md\"}","call_id":"call_c4xODNE00tUVz9lnM2AclZLs","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"read","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f751055c87d2af276890e6c623d3","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating lifecycle standards\",\"pattern\":\"conn|poison|lifecycle|PID|pid|test|allocat|panic|lock\",\"path\":\".worktrees/286-churn/AGENTS.md;.worktrees/286-churn/docs/KNOWN-HAZARDS.md\",\"case\":false,\"gitignore\":false,\"skip\":null}","call_id":"call_I7Tc7wlUioMWDMzQLLS5aRsh","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"grep","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}}]},"duration":6991.287900000112,"ttft":6951.190699999919,"completedAt":1788999505091,"contextSnapshot":{"promptTokens":32006,"nonMessageTokens":13648,"compactionEpoch":0}}}
{"type":"message","id":"67a7e241","parentId":"4595a942","timestamp":"2026-09-10T00:18:25.099Z","message":{"role":"toolResult","toolCallId":"call_eyb1aSoUOsXzhRujKrcoBRwI|fc_0dde72effda0ad21016aa1f751054c87d2865b2eaecd26c2b6","toolName":"read","content":[{"type":"text","text":"# spt-core — agent working rules\n\nHarness-independent core for an agent ecosystem (messaging, live-agent lifecycle, terminal hosting, P2P networking, runtime-manifest harness contract). Rust crates + a single `spt` binary. Clean-room rebuild of `claude_skill_owl`.\n\n**Orientation — read before working:** `PRD.md` (requirements), `ROADMAP.md` (build path), `CONTEXT.md` (glossary/model, authoritative for meaning), `docs/adr/` (decisions), `docs/KNOWN-HAZARDS.md` (invariants we must not re-break), `docs/{STORAGE,MANIFEST,CONTEXT-MEMORY,DOCS-STRATEGY,TRACEABILITY}.md`.\n\nFor maintainer/debug update work, also read `docs/DEBUG-ROLLOUT.md` before touching the update-set or rollout path.\n\nFor work needing Linux tests and proof, this machine is authorized to ssh into a Linux box: `reavus@kitsubito`\n\nIMPORTANT: All agents you work alongside run on spt-core itself — the legacy-SPT (`$OWL`/`$LIVE`) era ended with the 2026-07 fleet migration; those commands are gone. Reach other agents with `spt send` (or the `/sptc:send` skill); keep message bodies short, and send long bodies from a file rather than inline. Everything else can be learned from the corresponding `/sptc:` skill doc.\n\n## Requirement traceability (binding)\n\nThis project uses `traceable-reqs` (`traceable-reqs.toml` = the authoritative `REQ-*` registry). The full contract is `docs/TRACEABILITY.md`. The rules you must follow:\n\n1. **Tag evidence in the same change.** When you write a function/test/doc-section that satisfies a requirement stage, add its tag in that same commit:\n   `// [impl->REQ-FOO]` · `// [unit->REQ-FOO]` · `<!-- [doc->REQ-FOO] -->`\n   Stages: `doc` / `impl` / `unit` / `int`. Tag *on or immediately above* the real evidence — never at file tops to satisfy coverage.\n2. **Run `traceable-reqs check` before declaring work done.** Exit-1 means missing/invalid evidence — fix it, don't ship it.\n3. **New requirement → add it to `traceable-reqs.toml` first** (with a `REQ-*` id), then satisfy it. No untracked work; no untagged evidence.\n4. **KNOWN-HAZARDS are `REQ-HAZARD-*` requirements** — each needs a test before it's \"covered.\" Treat the hazard list as a conformance checklist you must satisfy, not advice.\n5. **Activate, don't pre-fail.** Requirements you aren't yet working stay `required_stages = []`. Activate (set real stages) only when starting the milestone that delivers them.\n\n## Issue tracking (unique case)\n\nCodebase issues are maintained on **`BigscreenVR/spt-bs-releases`** — the less-private *releases* repo — NOT on `spt-bs-core` where the code lives. This is a deliberate exception (operator-ruled 2026-07-27): file, triage, and reference issues against `spt-bs-releases`.\n\nRequests and milestones are tracked through the **alchemy shell** (`alchemy-0`, the Requests shell bound to `spt-bs-releases`) on top of those GitHub issues. The label taxonomy is the workflow: `state:` BACKLOG → EVAL → GREENLIT → WIP → ACCEPTANCE → DONE (plus CUT / SHELVED), `type:` BUGFIX / ADDITION / CHANGE, `flag: NEEDS-OPERATOR`, and per-node `kind: SPT-TRACKING` pins (marked DO NOT DELETE — leave them alone).\n\n## Other conventions\n\n- Match surrounding code style; copy-verbatim the sister project's stable wire/schema formats (ADR-0001), clean-room everything else.\n- Honor every `docs/KNOWN-HAZARDS.md` invariant — these are real bugs already paid for once.\n- Docs are dual-audience (human + AI dev-agent) per `docs/DOCS-STRATEGY.md`; doc generation is CI-gated against drift.\n- Commit messages end with the project's co-author trailer, exact spelling `Co-authored by: <your live agent name>` (e.g. `Co-authored by: todlando`) — NOT git's hyphenated `Co-Authored-By:`; CI's runner-informant parses only the space spelling for attribution.\n  - **Audit this trailer with `git log --grep` or a raw body read, NEVER `%(trailers:)`** (nor `interpret-trailers --parse`): git's trailer tokenizer requires a whitespace-free token, so the space spelling is structurally invisible to it and every trailer query returns a confident EMPTY on every commit — a clean zero that reads as total attribution loss. Measured 2026-07-30 (main @`8f3e10b`, Signed-off-by and hyphenated-key controls both parse); re-derived at full cost 2026-08-03 by one of its own original measurers, which is why this sentence lives here at the mandate rather than in a document you must decide to open.\n- **Worktrees (binding, operator-ruled 2026-07-05):** every `git worktree` (gate rigs, agent scratch, anything) goes under `.worktrees/` in THIS project's root (`.gitignore`d) — e.g. `git worktree add --detach .worktrees/gate-<sha> <sha>`. NEVER create worktrees as siblings of the project dir; stray sibling folders accumulate orphan-pinned skeletons and pollute `~/Documents/projects`. Remove yours when done (`git worktree remove` + prune; retry later if a handle pins it).\n<!-- [doc->REQ-POOL-FOREIGN-TREE-REFUSED] -->\n- **Build-cache pools (binding, releases#103):** a cargo `target/` directory belongs to ONE source tree. Two trees building into one pool overwrite each other's artifacts and fabricate reds that belong to neither — the v0.51.0 W1/H1 `E0308` that vanished on re-run. **The rule is about the RELATIONSHIP, not the path:** a junction and a `CARGO_TARGET_DIR` env var reach the same hazard, and the env var leaves no directory entry to notice afterwards, so \"is this dir under `.worktrees/`\" never was the question. A build script enforces it (`crates/spt-store/build.rs`) and refuses with `SPT_POOL_FOREIGN`. **Sharing a pool SEQUENTIALLY is fine and expected** — a finished lane is taken over loudly, not refused; only a LIVE lane blocks you. Claim a pool at lane start …\n- **Target-tree teardown (releases#103, same lane):** never delete a `target/` you have not classified. Check OUTBOUND first (`Get-Item -Force`: a junction is deleted as a link, a real dir as a tree), then sweep INBOUND for reparse points aimed at it before removing a real one, and measure size and free space before and after so the reclaim is a number rather than a claim. Reap the target SUBTREE, never a session or worktree directory around it. Do not point `CARGO_TARGET_DIR` into a scratchpad (v0.50.0 incident) or at another checkout's target (H1 incident).\n- **CI / merge strategy:** golden CI (ADR-0050, ruled 2026-07-27) — milestone-batch golden runs, thin-lane PRs, ff-only main (tested sha == merged sha). Piloting from the next milestone; read the ADR before shaping a PR lane.\n\nIf you finish a significant body of work without need for user intervention, or if your context gets too high, you can clear your own context and keep moving:\n1. Create a JIT plan for the next immediate body of work, if it isn't already planned\n2. If you are a live agent, /sptc:commune with immediate next steps and broad summary of the project's status + end goal\n3. Execute the /self-clear skill\n\n[Some lines truncated to 768 chars]"}],"details":{"totalLines":47,"displayContent":{"text":"# spt-core — agent working rules\r\n\r\nHarness-independent core for an agent ecosystem (messaging, live-agent lifecycle, terminal hosting, P2P networking, runtime-manifest harness contract). Rust crates + a single `spt` binary. Clean-room rebuild of `claude_skill_owl`.\r\n\r\n**Orientation — read before working:** `PRD.md` (requirements), `ROADMAP.md` (build path), `CONTEXT.md` (glossary/model, authoritative for meaning), `docs/adr/` (decisions), `docs/KNOWN-HAZARDS.md` (invariants we must not re-break), `docs/{STORAGE,MANIFEST,CONTEXT-MEMORY,DOCS-STRATEGY,TRACEABILITY}.md`.\r\n\r\nFor maintainer/debug update work, also read `docs/DEBUG-ROLLOUT.md` before touching the update-set or rollout path.\r\n\r\nFor work needing Linux tests and proof, this machine is authorized to ssh into a Linux box: `reavus@kitsubito`\r\n\r\nIMPORTANT: All agents you work alongside run on spt-core itself — the legacy-SPT (`$OWL`/`$LIVE`) era ended with the 2026-07 fleet migration; those commands are gone. Reach other agents with `spt send` (or the `/sptc:send` skill); keep message bodies short, and send long bodies from a file rather than inline. Everything else can be learned from the corresponding `/sptc:` skill doc.\r\n\r\n## Requirement traceability (binding)\r\n\r\nThis project uses `traceable-reqs` (`traceable-reqs.toml` = the authoritative `REQ-*` registry). The full contract is `docs/TRACEABILITY.md`. The rules you must follow:\r\n\r\n1. **Tag evidence in the same change.** When you write a function/test/doc-section that satisfies a requirement stage, add its tag in that same commit:\r\n   `// [impl->REQ-FOO]` · `// [unit->REQ-FOO]` · `<!-- [doc->REQ-FOO] -->`\r\n   Stages: `doc` / `impl` / `unit` / `int`. Tag *on or immediately above* the real evidence — never at file tops to satisfy coverage.\r\n2. **Run `traceable-reqs check` before declaring work done.** Exit-1 means missing/invalid evidence — fix it, don't ship it.\r\n3. **New requirement → add it to `traceable-reqs.toml` first** (with a `REQ-*` id), then satisfy it. No untracked work; no untagged evidence.\r\n4. **KNOWN-HAZARDS are `REQ-HAZARD-*` requirements** — each needs a test before it's \"covered.\" Treat the hazard list as a conformance checklist you must satisfy, not advice.\r\n5. **Activate, don't pre-fail.** Requirements you aren't yet working stay `required_stages = []`. Activate (set real stages) only when starting the milestone that delivers them.\r\n\r\n## Issue tracking (unique case)\r\n\r\nCodebase issues are maintained on **`BigscreenVR/spt-bs-releases`** — the less-private *releases* repo — NOT on `spt-bs-core` where the code lives. This is a deliberate exception (operator-ruled 2026-07-27): file, triage, and reference issues against `spt-bs-releases`.\r\n\r\nRequests and milestones are tracked through the **alchemy shell** (`alchemy-0`, the Requests shell bound to `spt-bs-releases`) on top of those GitHub issues. The label taxonomy is the workflow: `state:` BACKLOG → EVAL → GREENLIT → WIP → ACCEPTANCE → DONE (plus CUT / SHELVED), `type:` BUGFIX / ADDITION / CHANGE, `flag: NEEDS-OPERATOR`, and per-node `kind: SPT-TRACKING` pins (marked DO NOT DELETE — leave them alone).\r\n\r\n## Other conventions\r\n\r\n- Match surrounding code style; copy-verbatim the sister project's stable wire/schema formats (ADR-0001), clean-room everything else.\r\n- Honor every `docs/KNOWN-HAZARDS.md` invariant — these are real bugs already paid for once.\r\n- Docs are dual-audience (human + AI dev-agent) per `docs/DOCS-STRATEGY.md`; doc generation is CI-gated against drift.\r\n- Commit messages end with the project's co-author trailer, exact spelling `Co-authored by: <your live agent name>` (e.g. `Co-authored by: todlando`) — NOT git's hyphenated `Co-Authored-By:`; CI's runner-informant parses only the space spelling for attribution.\r\n  - **Audit this trailer with `git log --grep` or a raw body read, NEVER `%(trailers:)`** (nor `interpret-trailers --parse`): git's trailer tokenizer requires a whitespace-free token, so the space spelling is structurally invisible to it and every trailer query returns a confident EMPTY on every commit — a clean zero that reads as total attribution loss. Measured 2026-07-30 (main @`8f3e10b`, Signed-off-by and hyphenated-key controls both parse); re-derived at full cost 2026-08-03 by one of its own original measurers, which is why this sentence lives here at the mandate rather than in a document you must decide to open.\r\n- **Worktrees (binding, operator-ruled 2026-07-05):** every `git worktree` (gate rigs, agent scratch, anything) goes under `.worktrees/` in THIS project's root (`.gitignore`d) — e.g. `git worktree add --detach .worktrees/gate-<sha> <sha>`. NEVER create worktrees as siblings of the project dir; stray sibling folders accumulate orphan-pinned skeletons and pollute `~/Documents/projects`. Remove yours when done (`git worktree remove` + prune; retry later if a handle pins it).\r\n<!-- [doc->REQ-POOL-FOREIGN-TREE-REFUSED] -->\r\n- **Build-cache pools (binding, releases#103):** a cargo `target/` directory belongs to ONE source tree. Two trees building into one pool overwrite each other's artifacts and fabricate reds that belong to neither — the v0.51.0 W1/H1 `E0308` that vanished on re-run. **The rule is about the RELATIONSHIP, not the path:** a junction and a `CARGO_TARGET_DIR` env var reach the same hazard, and the env var leaves no directory entry to notice afterwards, so \"is this dir under `.worktrees/`\" never was the question. A build script enforces it (`crates/spt-store/build.rs`) and refuses with `SPT_POOL_FOREIGN`. **Sharing a pool SEQUENTIALLY is fine and expected** — a finished lane is taken over loudly, not refused; only a LIVE lane blocks you. Claim a pool at lane start …\n- **Target-tree teardown (releases#103, same lane):** never delete a `target/` you have not classified. Check OUTBOUND first (`Get-Item -Force`: a junction is deleted as a link, a real dir as a tree), then sweep INBOUND for reparse points aimed at it before removing a real one, and measure size and free space before and after so the reclaim is a number rather than a claim. Reap the target SUBTREE, never a session or worktree directory around it. Do not point `CARGO_TARGET_DIR` into a scratchpad (v0.50.0 incident) or at another checkout's target (H1 incident).\r\n- **CI / merge strategy:** golden CI (ADR-0050, ruled 2026-07-27) — milestone-batch golden runs, thin-lane PRs, ff-only main (tested sha == merged sha). Piloting from the next milestone; read the ADR before shaping a PR lane.\r\n\r\nIf you finish a significant body of work without need for user intervention, or if your context gets too high, you can clear your own context and keep moving:\r\n1. Create a JIT plan for the next immediate body of work, if it isn't already planned\r\n2. If you are a live agent, /sptc:commune with immediate next steps and broad summary of the project's status + end goal\r\n3. Execute the /self-clear skill\r","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47]},"fileSize":8436,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\AGENTS.md"},"limits":{"columnTruncated":{"maxColumn":768}}}},"isError":false,"timestamp":1788999505098}}
{"type":"message","id":"0be577da","parentId":"67a7e241","timestamp":"2026-09-10T00:18:25.099Z","message":{"role":"toolResult","toolCallId":"call_c4xODNE00tUVz9lnM2AclZLs|fc_0dde72effda0ad21016aa1f751055487d2aa1203ec2e0ccda4","toolName":"read","content":[{"type":"text","text":"# spt-core — agent working rules\n\nHarness-independent core for an agent ecosystem (messaging, live-agent lifecycle, terminal hosting, P2P networking, runtime-manifest harness contract). Rust crates + a single `spt` binary. Clean-room rebuild of `claude_skill_owl`.\n\n**Orientation — read before working:** `PRD.md` (requirements), `ROADMAP.md` (build path), `CONTEXT.md` (glossary/model, authoritative for meaning), `docs/adr/` (decisions), `docs/KNOWN-HAZARDS.md` (invariants we must not re-break), `docs/{STORAGE,MANIFEST,CONTEXT-MEMORY,DOCS-STRATEGY,TRACEABILITY}.md`.\n\nFor maintainer/debug update work, also read `docs/DEBUG-ROLLOUT.md` before touching the update-set or rollout path.\n\nFor work needing Linux tests and proof, this machine is authorized to ssh into a Linux box: `reavus@kitsubito`\n\nIMPORTANT: All agents you work alongside run on spt-core itself — the legacy-SPT (`$OWL`/`$LIVE`) era ended with the 2026-07 fleet migration; those commands are gone. Reach other agents with `spt send` (or the `/sptc:send` skill); keep message bodies short, and send long bodies from a file rather than inline. Everything else can be learned from the corresponding `/sptc:` skill doc.\n\n## Requirement traceability (binding)\n\nThis project uses `traceable-reqs` (`traceable-reqs.toml` = the authoritative `REQ-*` registry). The full contract is `docs/TRACEABILITY.md`. The rules you must follow:\n\n1. **Tag evidence in the same change.** When you write a function/test/doc-section that satisfies a requirement stage, add its tag in that same commit:\n   `// [impl->REQ-FOO]` · `// [unit->REQ-FOO]` · `<!-- [doc->REQ-FOO] -->`\n   Stages: `doc` / `impl` / `unit` / `int`. Tag *on or immediately above* the real evidence — never at file tops to satisfy coverage.\n2. **Run `traceable-reqs check` before declaring work done.** Exit-1 means missing/invalid evidence — fix it, don't ship it.\n3. **New requirement → add it to `traceable-reqs.toml` first** (with a `REQ-*` id), then satisfy it. No untracked work; no untagged evidence.\n4. **KNOWN-HAZARDS are `REQ-HAZARD-*` requirements** — each needs a test before it's \"covered.\" Treat the hazard list as a conformance checklist you must satisfy, not advice.\n5. **Activate, don't pre-fail.** Requirements you aren't yet working stay `required_stages = []`. Activate (set real stages) only when starting the milestone that delivers them.\n\n## Issue tracking (unique case)\n\nCodebase issues are maintained on **`BigscreenVR/spt-bs-releases`** — the less-private *releases* repo — NOT on `spt-bs-core` where the code lives. This is a deliberate exception (operator-ruled 2026-07-27): file, triage, and reference issues against `spt-bs-releases`.\n\nRequests and milestones are tracked through the **alchemy shell** (`alchemy-0`, the Requests shell bound to `spt-bs-releases`) on top of those GitHub issues. The label taxonomy is the workflow: `state:` BACKLOG → EVAL → GREENLIT → WIP → ACCEPTANCE → DONE (plus CUT / SHELVED), `type:` BUGFIX / ADDITION / CHANGE, `flag: NEEDS-OPERATOR`, and per-node `kind: SPT-TRACKING` pins (marked DO NOT DELETE — leave them alone).\n\n## Other conventions\n\n- Match surrounding code style; copy-verbatim the sister project's stable wire/schema formats (ADR-0001), clean-room everything else.\n- Honor every `docs/KNOWN-HAZARDS.md` invariant — these are real bugs already paid for once.\n- Docs are dual-audience (human + AI dev-agent) per `docs/DOCS-STRATEGY.md`; doc generation is CI-gated against drift.\n- Commit messages end with the project's co-author trailer, exact spelling `Co-authored by: <your live agent name>` (e.g. `Co-authored by: todlando`) — NOT git's hyphenated `Co-Authored-By:`; CI's runner-informant parses only the space spelling for attribution.\n  - **Audit this trailer with `git log --grep` or a raw body read, NEVER `%(trailers:)`** (nor `interpret-trailers --parse`): git's trailer tokenizer requires a whitespace-free token, so the space spelling is structurally invisible to it and every trailer query returns a confident EMPTY on every commit — a clean zero that reads as total attribution loss. Measured 2026-07-30 (main @`8f3e10b`, Signed-off-by and hyphenated-key controls both parse); re-derived at full cost 2026-08-03 by one of its own original measurers, which is why this sentence lives here at the mandate rather than in a document you must decide to open.\n- **Worktrees (binding, operator-ruled 2026-07-05):** every `git worktree` (gate rigs, agent scratch, anything) goes under `.worktrees/` in THIS project's root (`.gitignore`d) — e.g. `git worktree add --detach .worktrees/gate-<sha> <sha>`. NEVER create worktrees as siblings of the project dir; stray sibling folders accumulate orphan-pinned skeletons and pollute `~/Documents/projects`. Remove yours when done (`git worktree remove` + prune; retry later if a handle pins it).\n<!-- [doc->REQ-POOL-FOREIGN-TREE-REFUSED] -->\n- **Build-cache pools (binding, releases#103):** a cargo `target/` directory belongs to ONE source tree. Two trees building into one pool overwrite each other's artifacts and fabricate reds that belong to neither — the v0.51.0 W1/H1 `E0308` that vanished on re-run. **The rule is about the RELATIONSHIP, not the path:** a junction and a `CARGO_TARGET_DIR` env var reach the same hazard, and the env var leaves no directory entry to notice afterwards, so \"is this dir under `.worktrees/`\" never was the question. A build script enforces it (`crates/spt-store/build.rs`) and refuses with `SPT_POOL_FOREIGN`. **Sharing a pool SEQUENTIALLY is fine and expected** — a finished lane is taken over loudly, not refused; only a LIVE lane blocks you. Claim a pool at lane start …\n- **Target-tree teardown (releases#103, same lane):** never delete a `target/` you have not classified. Check OUTBOUND first (`Get-Item -Force`: a junction is deleted as a link, a real dir as a tree), then sweep INBOUND for reparse points aimed at it before removing a real one, and measure size and free space before and after so the reclaim is a number rather than a claim. Reap the target SUBTREE, never a session or worktree directory around it. Do not point `CARGO_TARGET_DIR` into a scratchpad (v0.50.0 incident) or at another checkout's target (H1 incident).\n- **CI / merge strategy:** golden CI (ADR-0050, ruled 2026-07-27) — milestone-batch golden runs, thin-lane PRs, ff-only main (tested sha == merged sha). Piloting from the next milestone; read the ADR before shaping a PR lane.\n\nIf you finish a significant body of work without need for user intervention, or if your context gets too high, you can clear your own context and keep moving:\n1. Create a JIT plan for the next immediate body of work, if it isn't already planned\n2. If you are a live agent, /sptc:commune with immediate next steps and broad summary of the project's status + end goal\n3. Execute the /self-clear skill\n\n[Some lines truncated to 768 chars]"}],"details":{"totalLines":47,"displayContent":{"text":"# spt-core — agent working rules\r\n\r\nHarness-independent core for an agent ecosystem (messaging, live-agent lifecycle, terminal hosting, P2P networking, runtime-manifest harness contract). Rust crates + a single `spt` binary. Clean-room rebuild of `claude_skill_owl`.\r\n\r\n**Orientation — read before working:** `PRD.md` (requirements), `ROADMAP.md` (build path), `CONTEXT.md` (glossary/model, authoritative for meaning), `docs/adr/` (decisions), `docs/KNOWN-HAZARDS.md` (invariants we must not re-break), `docs/{STORAGE,MANIFEST,CONTEXT-MEMORY,DOCS-STRATEGY,TRACEABILITY}.md`.\r\n\r\nFor maintainer/debug update work, also read `docs/DEBUG-ROLLOUT.md` before touching the update-set or rollout path.\r\n\r\nFor work needing Linux tests and proof, this machine is authorized to ssh into a Linux box: `reavus@kitsubito`\r\n\r\nIMPORTANT: All agents you work alongside run on spt-core itself — the legacy-SPT (`$OWL`/`$LIVE`) era ended with the 2026-07 fleet migration; those commands are gone. Reach other agents with `spt send` (or the `/sptc:send` skill); keep message bodies short, and send long bodies from a file rather than inline. Everything else can be learned from the corresponding `/sptc:` skill doc.\r\n\r\n## Requirement traceability (binding)\r\n\r\nThis project uses `traceable-reqs` (`traceable-reqs.toml` = the authoritative `REQ-*` registry). The full contract is `docs/TRACEABILITY.md`. The rules you must follow:\r\n\r\n1. **Tag evidence in the same change.** When you write a function/test/doc-section that satisfies a requirement stage, add its tag in that same commit:\r\n   `// [impl->REQ-FOO]` · `// [unit->REQ-FOO]` · `<!-- [doc->REQ-FOO] -->`\r\n   Stages: `doc` / `impl` / `unit` / `int`. Tag *on or immediately above* the real evidence — never at file tops to satisfy coverage.\r\n2. **Run `traceable-reqs check` before declaring work done.** Exit-1 means missing/invalid evidence — fix it, don't ship it.\r\n3. **New requirement → add it to `traceable-reqs.toml` first** (with a `REQ-*` id), then satisfy it. No untracked work; no untagged evidence.\r\n4. **KNOWN-HAZARDS are `REQ-HAZARD-*` requirements** — each needs a test before it's \"covered.\" Treat the hazard list as a conformance checklist you must satisfy, not advice.\r\n5. **Activate, don't pre-fail.** Requirements you aren't yet working stay `required_stages = []`. Activate (set real stages) only when starting the milestone that delivers them.\r\n\r\n## Issue tracking (unique case)\r\n\r\nCodebase issues are maintained on **`BigscreenVR/spt-bs-releases`** — the less-private *releases* repo — NOT on `spt-bs-core` where the code lives. This is a deliberate exception (operator-ruled 2026-07-27): file, triage, and reference issues against `spt-bs-releases`.\r\n\r\nRequests and milestones are tracked through the **alchemy shell** (`alchemy-0`, the Requests shell bound to `spt-bs-releases`) on top of those GitHub issues. The label taxonomy is the workflow: `state:` BACKLOG → EVAL → GREENLIT → WIP → ACCEPTANCE → DONE (plus CUT / SHELVED), `type:` BUGFIX / ADDITION / CHANGE, `flag: NEEDS-OPERATOR`, and per-node `kind: SPT-TRACKING` pins (marked DO NOT DELETE — leave them alone).\r\n\r\n## Other conventions\r\n\r\n- Match surrounding code style; copy-verbatim the sister project's stable wire/schema formats (ADR-0001), clean-room everything else.\r\n- Honor every `docs/KNOWN-HAZARDS.md` invariant — these are real bugs already paid for once.\r\n- Docs are dual-audience (human + AI dev-agent) per `docs/DOCS-STRATEGY.md`; doc generation is CI-gated against drift.\r\n- Commit messages end with the project's co-author trailer, exact spelling `Co-authored by: <your live agent name>` (e.g. `Co-authored by: todlando`) — NOT git's hyphenated `Co-Authored-By:`; CI's runner-informant parses only the space spelling for attribution.\r\n  - **Audit this trailer with `git log --grep` or a raw body read, NEVER `%(trailers:)`** (nor `interpret-trailers --parse`): git's trailer tokenizer requires a whitespace-free token, so the space spelling is structurally invisible to it and every trailer query returns a confident EMPTY on every commit — a clean zero that reads as total attribution loss. Measured 2026-07-30 (main @`8f3e10b`, Signed-off-by and hyphenated-key controls both parse); re-derived at full cost 2026-08-03 by one of its own original measurers, which is why this sentence lives here at the mandate rather than in a document you must decide to open.\r\n- **Worktrees (binding, operator-ruled 2026-07-05):** every `git worktree` (gate rigs, agent scratch, anything) goes under `.worktrees/` in THIS project's root (`.gitignore`d) — e.g. `git worktree add --detach .worktrees/gate-<sha> <sha>`. NEVER create worktrees as siblings of the project dir; stray sibling folders accumulate orphan-pinned skeletons and pollute `~/Documents/projects`. Remove yours when done (`git worktree remove` + prune; retry later if a handle pins it).\r\n<!-- [doc->REQ-POOL-FOREIGN-TREE-REFUSED] -->\r\n- **Build-cache pools (binding, releases#103):** a cargo `target/` directory belongs to ONE source tree. Two trees building into one pool overwrite each other's artifacts and fabricate reds that belong to neither — the v0.51.0 W1/H1 `E0308` that vanished on re-run. **The rule is about the RELATIONSHIP, not the path:** a junction and a `CARGO_TARGET_DIR` env var reach the same hazard, and the env var leaves no directory entry to notice afterwards, so \"is this dir under `.worktrees/`\" never was the question. A build script enforces it (`crates/spt-store/build.rs`) and refuses with `SPT_POOL_FOREIGN`. **Sharing a pool SEQUENTIALLY is fine and expected** — a finished lane is taken over loudly, not refused; only a LIVE lane blocks you. Claim a pool at lane start …\n- **Target-tree teardown (releases#103, same lane):** never delete a `target/` you have not classified. Check OUTBOUND first (`Get-Item -Force`: a junction is deleted as a link, a real dir as a tree), then sweep INBOUND for reparse points aimed at it before removing a real one, and measure size and free space before and after so the reclaim is a number rather than a claim. Reap the target SUBTREE, never a session or worktree directory around it. Do not point `CARGO_TARGET_DIR` into a scratchpad (v0.50.0 incident) or at another checkout's target (H1 incident).\r\n- **CI / merge strategy:** golden CI (ADR-0050, ruled 2026-07-27) — milestone-batch golden runs, thin-lane PRs, ff-only main (tested sha == merged sha). Piloting from the next milestone; read the ADR before shaping a PR lane.\r\n\r\nIf you finish a significant body of work without need for user intervention, or if your context gets too high, you can clear your own context and keep moving:\r\n1. Create a JIT plan for the next immediate body of work, if it isn't already planned\r\n2. If you are a live agent, /sptc:commune with immediate next steps and broad summary of the project's status + end goal\r\n3. Execute the /self-clear skill\r","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47]},"fileSize":8436,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\AGENTS.md"},"limits":{"columnTruncated":{"maxColumn":768}}}},"isError":false,"timestamp":1788999505099}}
{"type":"message","id":"ec38e681","parentId":"0be577da","timestamp":"2026-09-10T00:18:25.100Z","message":{"role":"toolResult","toolCallId":"call_qPHAxmgOQ986O8RTWTPcS7JZ|fc_0dde72effda0ad21016aa1f751053887d2bf80cde1dc1bbbd3","toolName":"read","content":[{"type":"text","text":"diff --git a/CHANGELOG.md b/CHANGELOG.md\nindex eafe18d6..5b4a7156 100644\n--- a/CHANGELOG.md\n+++ b/CHANGELOG.md\n@@ -10,6 +10,17 @@ breaks something, or changes the observable behavior of existing surfaces\n broadly; **patch** for fixes, and for additive opt-in capability — a new key,\n flag, or page that no existing user can encounter without opting into it.\n \n+## [Unreleased]\n+\n+### Fixed\n+\n+<!-- [doc->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED] -->\n+- Daemon logs summarize healthy IPC connection starts and closes once per\n+  minute instead of logging each one. Failures and subscriber lifecycle\n+  events keep their per-connection attribution, including the original\n+  first-write time. Livehost session polling now reuses its healthy\n+  connection and names the reason when it opens a replacement.\n+\n ## [0.68.0] - 2026-09-08\n \n Web serving. Files, directories, adapter documentation and the changelog are\ndiff --git a/crates/spt-daemon/src/conn.rs b/crates/spt-daemon/src/conn.rs\nindex 9c51ad4e..163fecb6 100644\n--- a/crates/spt-daemon/src/conn.rs\n+++ b/crates/spt-daemon/src/conn.rs\n@@ -61,7 +61,7 @@\n \n use std::io;\n use std::sync::atomic::{AtomicBool, AtomicU64, Ordering};\n-use std::sync::{Arc, Condvar, Mutex, MutexGuard, OnceLock};\n+use std::sync::{Arc, Condvar, LazyLock, Mutex, MutexGuard, OnceLock};\n use std::thread::{self, JoinHandle};\n use std::time::{Duration, Instant, SystemTime, UNIX_EPOCH};\n \n@@ -99,12 +99,107 @@ pub(crate) fn init_log_anchor() {\n /// to field incident times (wall).\n // [impl->REQ-CONN-POISON-ATTRIBUTION]\n pub(crate) fn log_stamp() -> String {\n-    let wall_ms = SystemTime::now()\n-        .duration_since(UNIX_EPOCH)\n-        .map(|d| d.as_millis())\n-        .unwrap_or(0);\n-    let mono_ms = MONO_ANCHOR.get_or_init(Instant::now).elapsed().as_millis();\n-    format!(\"wall_ms={wall_ms} mono_ms={mono_ms}\")\n+    let stamp = LogTime::now();\n+    format!(\"wall_ms={} mono_ms={}\", stamp.wall_ms, stamp.mono_ms)\n+}\n+\n+/// Retain the original write instant even when its healthy line is aggregated.\n+#[derive(Clone, Copy)]\n+struct LogTime {\n+    wall_ms: u128,\n+    mono_ms: u128,\n+}\n+\n+impl LogTime {\n+    fn now() -> Self {\n+        Self {\n+            wall_ms: SystemTime::now()\n+                .duration_since(UNIX_EPOCH)\n+                .map(|d| d.as_millis())\n+                .unwrap_or(0),\n+            mono_ms: MONO_ANCHOR.get_or_init(Instant::now).elapsed().as_millis(),\n+        }\n+    }\n+}\n+\n+const HEALTHY_LOG_INTERVAL: Duration = Duration::from_secs(60);\n+\n+struct HealthyWindow {\n+    since: Instant,\n+    starts: u64,\n+    closes: u64,\n+}\n+\n+/// One process-wide counter window, not one limiter per connection. No thread,\n+/// queue or per-conn healthy string: the next event after 60s emits the totals.\n+/// Exceptional and role-transition records bypass this window entirely.\n+// [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+struct LifecycleLog {\n+    healthy: Mutex<HealthyWindow>,\n+    #[cfg(test)]\n+    capture: Option<Mutex<Vec<String>>>,\n+    #[cfg(test)]\n+    clock: Option<Mutex<Instant>>,\n+}\n+\n+impl LifecycleLog {\n+    fn new() -> Self {\n+        Self {\n+            healthy: Mutex::new(HealthyWindow {\n+                since: Instant::now(),\n+                starts: 0,\n+                closes: 0,\n+            }),\n+            #[cfg(test)]\n+            capture: None,\n+            #[cfg(test)]\n+            clock: None,\n+        }\n+    }\n+\n+    #[cfg(test)]\n+    fn captured(now: Instant) -> Self {\n+        let mut log = Self::new();\n+        log.healthy.get_mut().unwrap().since = now;\n+        log.capture = Some(Mutex::new(Vec::new()));\n+        log.clock = Some(Mutex::new(now));\n+        log\n+    }\n+\n+    fn emit(&self, args: std::fmt::Arguments<'_>) {\n+        #[cfg(test)]\n+        if let Some(lines) = &self.capture {\n+            recover(lines).push(args.to_string());\n+            return;\n+        }\n+        spt_proto::emit_line_err!(\"{args}\");\n+    }\n+\n+    fn healthy(&self, start: bool) {\n+        let mut window = recover(&self.healthy);\n+        // Read time under the counter lock so concurrent callers cannot move\n+        // the window backwards after waiting for a newer caller to flush.\n+        let now = Instant::now();\n+        #[cfg(test)]\n+        let now = self.clock.as_ref().map(|c| *recover(c)).unwrap_or(now);\n+        if start {\n+            window.starts = window.starts.saturating_add(1);\n+        } else {\n+            window.closes = window.closes.saturating_add(1);\n+        }\n+        let elapsed = now.saturating_duration_since(window.since);\n+        if elapsed < HEALTHY_LOG_INTERVAL {\n+            return;\n+        }\n+        let (starts, closes) = (window.starts, window.closes);\n+        *window = HealthyWindow { since: now, starts: 0, closes: 0 };\n+        drop(window);\n+        self.emit(format_args!(\n+            \"CONN_LIFECYCLE_SUMMARY: scope=broker-ipc {} interval_ms={} starts={starts} closes={closes} \\\n+             [REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\",\n+            log_stamp(), elapsed.as_millis()\n+        ));\n+    }\n }\n \n /// Attribution-label cap: facts accrete per role attach (a long-lived brain\n@@ -144,8 +239,9 @@ struct Inner {\n     /// Accreted attribution facts (subscriber role, endpoint/session where\n     /// known — [`BrokerConn::describe`]). Bounded by [`LABEL_CAP`].\n     label: Mutex<String>,\n-    /// Latch: the once-per-conn `write-start` lifecycle record fired.\n-    first_write_logged: AtomicBool,\n+    /// Once-per-conn original first committed write time; retained on failures.\n+    first_write: OnceLock<LogTime>,\n+    lifecycle_log: Arc<LifecycleLog>,\n     gate: Mutex<Gate>,\n     gate_cv: Condvar,\n     dog: Mutex<Dog>,\n@@ -223,7 +319,13 @@ impl Inner {\n     fn attribution(&self) -> String {\n         let label = recover(&self.label);\n         let facts: &str = if label.is_empty() { \"role=unattributed\" } else { &label };\n-        format!(\"conn={} {} {}\", self.id, facts, log_stamp())\n+        match self.first_write.get() {\n+            Some(start) => format!(\n+                \"conn={} {} {} first_write_wall_ms={} first_write_mono_ms={}\",\n+                self.id, facts, log_stamp(), start.wall_ms, start.mono_ms\n+            ),\n+            None => format!(\"conn={} {} {} first_write=none\", self.id, facts, log_stamp()),\n+        }\n     }\n \n     /// Render the once-per-conn retirement record — the F-039 token split\n@@ -261,18 +363,17 @@ impl Inner {\n         }\n     }\n \n-    /// One BOUNDED per-conn lifecycle record (leg d, doyle-confirmed\n-    /// UNCONDITIONAL): `write-start` (first write only), `transport-close`\n-    /// (drop), `writer-exit` / `*-attach` / `*-replaced` (emitted by the\n-    /// broker's sink machinery through [`BrokerConn::lifecycle_event`]). The\n-    /// timeout/cancel leg is the retirement record itself.\n+    /// Per-conn exceptional and role-transition records remain unconditional.\n+    /// releases#286 supersedes only healthy start/close emission, which uses\n+    /// the process-wide interval counter instead. Retirement retains the\n+    /// original first-write time even when the start line was aggregated.\n     // [impl->REQ-CONN-POISON-ATTRIBUTION]\n     fn lifecycle(&self, event: &str, extra: &str) {\n         let sep = if extra.is_empty() { \"\" } else { \" \" };\n-        spt_proto::emit_line_err!(\n+        self.lifecycle_log.emit(format_args!(\n             \"CONN_LIFECYCLE: {} event={event}{sep}{extra} [REQ-CONN-POISON-ATTRIBUTION]\",\n             self.attribution()\n-        );\n+        ));\n     }\n \n     /// Abort the connection's in-flight I/O OUT OF BAND (both directions —\n@@ -287,10 +388,10 @@ impl Inner {\n             // First poison of this conn: loud, once. The token is class-split\n             // (F-039 leg a): deadline → CONN_WRITE_POISONED (the wedge\n             // observable), organic fast-fail → CONN_WRITE_RETIRED.\n-            eprintln!(\n+            self.lifecycle_log.emit(format_args!(\n                 \"{}\",\n                 self.render_retirement(self.timed_out.load(Ordering::Acquire), cause)\n-            );\n+            ));\n         }\n         #[cfg(windows)]\n         // SAFETY: scalar kernel32 call on a handle whose owning half is kept\n@@ -424,11 +525,17 @@ impl BrokerConn {\n     /// deadline (gate-wait + OS write completion) — the broker passes its\n     /// `brain_write_deadline()` (`SPT_BRAIN_WRITE_DEADLINE_MS` knob).\n     pub(crate) fn new(half: SendHalf, bound: Duration) -> Self {\n+        static LOG: LazyLock<Arc<LifecycleLog>> = LazyLock::new(|| Arc::new(LifecycleLog::new()));\n+        Self::with_log(half, bound, Arc::clone(&LOG))\n+    }\n+\n+    fn with_log(half: SendHalf, bound: Duration, lifecycle_log: Arc<LifecycleLog>) -> Self {\n         let raw = raw_of(&half);\n         let inner = Arc::new(Inner {\n             id: CONN_ID_SEQ.fetch_add(1, Ordering::Relaxed) + 1,\n             label: Mutex::new(String::new()),\n-            first_write_logged: AtomicBool::new(false),\n+            first_write: OnceLock::new(),\n+            lifecycle_log,\n             gate: Mutex::new(Gate { half: Some(half) }),\n             gate_cv: Condvar::new(),\n             dog: Mutex::new(Dog {\n@@ -539,12 +646,12 @@ impl BrokerConn {\n             return Ok(ConnWrite::Superseded);\n         }\n \n-        // Once-per-conn `write-start` lifecycle record (leg d): marks the conn's\n-        // first committed write, so a conn that later retires can be correlated\n-        // to when it went active (fresh-carrier churn shows as open→start→retire\n-        // triplets with fresh ids).\n-        if !inner.first_write_logged.swap(true, Ordering::AcqRel) {\n-            inner.lifecycle(\"write-start\", \"\");\n+        // The gate serializes first-write initialization. Save the instant\n+        // BEFORE I/O so a later failure still names when this conn went active.\n+        // [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+        if inner.first_write.get().is_none() {\n+            let _ = inner.first_write.set(LogTime::now());\n+            inner.lifecycle_log.healthy(true);\n         }\n \n         // ── Check out the half + arm the watchdog, then write OUT of locks. ──\n@@ -638,17 +745,20 @@ impl BrokerConn {\n \n impl Drop for BrokerConn {\n     fn drop(&mut self) {\n-        // Once-per-conn `transport-close` lifecycle record (leg d): the handle\n-        // is about to close — the terminal record of this conn id's life, with\n-        // how it ended (organic EOF vs poisoned vs deadline-poisoned).\n-        self.inner.lifecycle(\n-            \"transport-close\",\n-            &format!(\n-                \"poisoned={} timed_out={}\",\n-                self.inner.poisoned.load(Ordering::Acquire),\n-                self.inner.timed_out.load(Ordering::Acquire)\n-            ),\n-        );\n+        // Only a healthy physical close is counted instead of logged per conn.\n+        // Error/timeout closes keep their full attribution and original start.\n+        // [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+        if self.inner.poisoned.load(Ordering::Acquire) {\n+            self.inner.lifecycle(\n+                \"transport-close\",\n+                &format!(\n+                    \"poisoned=true timed_out={}\",\n+                    self.inner.timed_out.load(Ordering::Acquire)\n+                ),\n+            );\n+        } else {\n+            self.inner.lifecycle_log.healthy(false);\n+        }\n         {\n             let mut d = recover(&self.inner.dog);\n             d.shutdown = true;\n@@ -688,6 +798,17 @@ mod tests {\n     /// primitive, the client end returned so a test can hold it alive or drop\n     /// it to force an ORGANIC write failure.\n     fn make_conn() -> (BrokerConn, Stream) {\n+        make_conn_with_log(Arc::new(LifecycleLog::new()))\n+    }\n+\n+    fn make_conn_with_log(log: Arc<LifecycleLog>) -> (BrokerConn, Stream) {\n+        let (conn, client, _recv) = make_duplex_with_log(log);\n+        (conn, client)\n+    }\n+\n+    fn make_duplex_with_log(\n+        log: Arc<LifecycleLog>,\n+    ) -> (BrokerConn, Stream, interprocess::local_socket::RecvHalf) {\n         static SEQ: AtomicU32 = AtomicU32::new(0);\n\n[Showing lines 1-300 of 659. Use :301 to continue]"}],"details":{"resolvedPath":"C:\\Users\\decid\\.omp\\agent\\sessions\\-Documents-projects-spt-core\\2026-09-09T23-45-42-317Z_01a08890-21ad-7000-8f54-208d4f4f2f2c\\65.bash-original.log","contentType":"text/plain","totalLines":659,"displayContent":{"text":"diff --git a/CHANGELOG.md b/CHANGELOG.md\nindex eafe18d6..5b4a7156 100644\n--- a/CHANGELOG.md\n+++ b/CHANGELOG.md\n@@ -10,6 +10,17 @@ breaks something, or changes the observable behavior of existing surfaces\n broadly; **patch** for fixes, and for additive opt-in capability — a new key,\n flag, or page that no existing user can encounter without opting into it.\n \n+## [Unreleased]\n+\n+### Fixed\n+\n+<!-- [doc->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED] -->\n+- Daemon logs summarize healthy IPC connection starts and closes once per\n+  minute instead of logging each one. Failures and subscriber lifecycle\n+  events keep their per-connection attribution, including the original\n+  first-write time. Livehost session polling now reuses its healthy\n+  connection and names the reason when it opens a replacement.\n+\n ## [0.68.0] - 2026-09-08\n \n Web serving. Files, directories, adapter documentation and the changelog are\ndiff --git a/crates/spt-daemon/src/conn.rs b/crates/spt-daemon/src/conn.rs\nindex 9c51ad4e..163fecb6 100644\n--- a/crates/spt-daemon/src/conn.rs\n+++ b/crates/spt-daemon/src/conn.rs\n@@ -61,7 +61,7 @@\n \n use std::io;\n use std::sync::atomic::{AtomicBool, AtomicU64, Ordering};\n-use std::sync::{Arc, Condvar, Mutex, MutexGuard, OnceLock};\n+use std::sync::{Arc, Condvar, LazyLock, Mutex, MutexGuard, OnceLock};\n use std::thread::{self, JoinHandle};\n use std::time::{Duration, Instant, SystemTime, UNIX_EPOCH};\n \n@@ -99,12 +99,107 @@ pub(crate) fn init_log_anchor() {\n /// to field incident times (wall).\n // [impl->REQ-CONN-POISON-ATTRIBUTION]\n pub(crate) fn log_stamp() -> String {\n-    let wall_ms = SystemTime::now()\n-        .duration_since(UNIX_EPOCH)\n-        .map(|d| d.as_millis())\n-        .unwrap_or(0);\n-    let mono_ms = MONO_ANCHOR.get_or_init(Instant::now).elapsed().as_millis();\n-    format!(\"wall_ms={wall_ms} mono_ms={mono_ms}\")\n+    let stamp = LogTime::now();\n+    format!(\"wall_ms={} mono_ms={}\", stamp.wall_ms, stamp.mono_ms)\n+}\n+\n+/// Retain the original write instant even when its healthy line is aggregated.\n+#[derive(Clone, Copy)]\n+struct LogTime {\n+    wall_ms: u128,\n+    mono_ms: u128,\n+}\n+\n+impl LogTime {\n+    fn now() -> Self {\n+        Self {\n+            wall_ms: SystemTime::now()\n+                .duration_since(UNIX_EPOCH)\n+                .map(|d| d.as_millis())\n+                .unwrap_or(0),\n+            mono_ms: MONO_ANCHOR.get_or_init(Instant::now).elapsed().as_millis(),\n+        }\n+    }\n+}\n+\n+const HEALTHY_LOG_INTERVAL: Duration = Duration::from_secs(60);\n+\n+struct HealthyWindow {\n+    since: Instant,\n+    starts: u64,\n+    closes: u64,\n+}\n+\n+/// One process-wide counter window, not one limiter per connection. No thread,\n+/// queue or per-conn healthy string: the next event after 60s emits the totals.\n+/// Exceptional and role-transition records bypass this window entirely.\n+// [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+struct LifecycleLog {\n+    healthy: Mutex<HealthyWindow>,\n+    #[cfg(test)]\n+    capture: Option<Mutex<Vec<String>>>,\n+    #[cfg(test)]\n+    clock: Option<Mutex<Instant>>,\n+}\n+\n+impl LifecycleLog {\n+    fn new() -> Self {\n+        Self {\n+            healthy: Mutex::new(HealthyWindow {\n+                since: Instant::now(),\n+                starts: 0,\n+                closes: 0,\n+            }),\n+            #[cfg(test)]\n+            capture: None,\n+            #[cfg(test)]\n+            clock: None,\n+        }\n+    }\n+\n+    #[cfg(test)]\n+    fn captured(now: Instant) -> Self {\n+        let mut log = Self::new();\n+        log.healthy.get_mut().unwrap().since = now;\n+        log.capture = Some(Mutex::new(Vec::new()));\n+        log.clock = Some(Mutex::new(now));\n+        log\n+    }\n+\n+    fn emit(&self, args: std::fmt::Arguments<'_>) {\n+        #[cfg(test)]\n+        if let Some(lines) = &self.capture {\n+            recover(lines).push(args.to_string());\n+            return;\n+        }\n+        spt_proto::emit_line_err!(\"{args}\");\n+    }\n+\n+    fn healthy(&self, start: bool) {\n+        let mut window = recover(&self.healthy);\n+        // Read time under the counter lock so concurrent callers cannot move\n+        // the window backwards after waiting for a newer caller to flush.\n+        let now = Instant::now();\n+        #[cfg(test)]\n+        let now = self.clock.as_ref().map(|c| *recover(c)).unwrap_or(now);\n+        if start {\n+            window.starts = window.starts.saturating_add(1);\n+        } else {\n+            window.closes = window.closes.saturating_add(1);\n+        }\n+        let elapsed = now.saturating_duration_since(window.since);\n+        if elapsed < HEALTHY_LOG_INTERVAL {\n+            return;\n+        }\n+        let (starts, closes) = (window.starts, window.closes);\n+        *window = HealthyWindow { since: now, starts: 0, closes: 0 };\n+        drop(window);\n+        self.emit(format_args!(\n+            \"CONN_LIFECYCLE_SUMMARY: scope=broker-ipc {} interval_ms={} starts={starts} closes={closes} \\\n+             [REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\",\n+            log_stamp(), elapsed.as_millis()\n+        ));\n+    }\n }\n \n /// Attribution-label cap: facts accrete per role attach (a long-lived brain\n@@ -144,8 +239,9 @@ struct Inner {\n     /// Accreted attribution facts (subscriber role, endpoint/session where\n     /// known — [`BrokerConn::describe`]). Bounded by [`LABEL_CAP`].\n     label: Mutex<String>,\n-    /// Latch: the once-per-conn `write-start` lifecycle record fired.\n-    first_write_logged: AtomicBool,\n+    /// Once-per-conn original first committed write time; retained on failures.\n+    first_write: OnceLock<LogTime>,\n+    lifecycle_log: Arc<LifecycleLog>,\n     gate: Mutex<Gate>,\n     gate_cv: Condvar,\n     dog: Mutex<Dog>,\n@@ -223,7 +319,13 @@ impl Inner {\n     fn attribution(&self) -> String {\n         let label = recover(&self.label);\n         let facts: &str = if label.is_empty() { \"role=unattributed\" } else { &label };\n-        format!(\"conn={} {} {}\", self.id, facts, log_stamp())\n+        match self.first_write.get() {\n+            Some(start) => format!(\n+                \"conn={} {} {} first_write_wall_ms={} first_write_mono_ms={}\",\n+                self.id, facts, log_stamp(), start.wall_ms, start.mono_ms\n+            ),\n+            None => format!(\"conn={} {} {} first_write=none\", self.id, facts, log_stamp()),\n+        }\n     }\n \n     /// Render the once-per-conn retirement record — the F-039 token split\n@@ -261,18 +363,17 @@ impl Inner {\n         }\n     }\n \n-    /// One BOUNDED per-conn lifecycle record (leg d, doyle-confirmed\n-    /// UNCONDITIONAL): `write-start` (first write only), `transport-close`\n-    /// (drop), `writer-exit` / `*-attach` / `*-replaced` (emitted by the\n-    /// broker's sink machinery through [`BrokerConn::lifecycle_event`]). The\n-    /// timeout/cancel leg is the retirement record itself.\n+    /// Per-conn exceptional and role-transition records remain unconditional.\n+    /// releases#286 supersedes only healthy start/close emission, which uses\n+    /// the process-wide interval counter instead. Retirement retains the\n+    /// original first-write time even when the start line was aggregated.\n     // [impl->REQ-CONN-POISON-ATTRIBUTION]\n     fn lifecycle(&self, event: &str, extra: &str) {\n         let sep = if extra.is_empty() { \"\" } else { \" \" };\n-        spt_proto::emit_line_err!(\n+        self.lifecycle_log.emit(format_args!(\n             \"CONN_LIFECYCLE: {} event={event}{sep}{extra} [REQ-CONN-POISON-ATTRIBUTION]\",\n             self.attribution()\n-        );\n+        ));\n     }\n \n     /// Abort the connection's in-flight I/O OUT OF BAND (both directions —\n@@ -287,10 +388,10 @@ impl Inner {\n             // First poison of this conn: loud, once. The token is class-split\n             // (F-039 leg a): deadline → CONN_WRITE_POISONED (the wedge\n             // observable), organic fast-fail → CONN_WRITE_RETIRED.\n-            eprintln!(\n+            self.lifecycle_log.emit(format_args!(\n                 \"{}\",\n                 self.render_retirement(self.timed_out.load(Ordering::Acquire), cause)\n-            );\n+            ));\n         }\n         #[cfg(windows)]\n         // SAFETY: scalar kernel32 call on a handle whose owning half is kept\n@@ -424,11 +525,17 @@ impl BrokerConn {\n     /// deadline (gate-wait + OS write completion) — the broker passes its\n     /// `brain_write_deadline()` (`SPT_BRAIN_WRITE_DEADLINE_MS` knob).\n     pub(crate) fn new(half: SendHalf, bound: Duration) -> Self {\n+        static LOG: LazyLock<Arc<LifecycleLog>> = LazyLock::new(|| Arc::new(LifecycleLog::new()));\n+        Self::with_log(half, bound, Arc::clone(&LOG))\n+    }\n+\n+    fn with_log(half: SendHalf, bound: Duration, lifecycle_log: Arc<LifecycleLog>) -> Self {\n         let raw = raw_of(&half);\n         let inner = Arc::new(Inner {\n             id: CONN_ID_SEQ.fetch_add(1, Ordering::Relaxed) + 1,\n             label: Mutex::new(String::new()),\n-            first_write_logged: AtomicBool::new(false),\n+            first_write: OnceLock::new(),\n+            lifecycle_log,\n             gate: Mutex::new(Gate { half: Some(half) }),\n             gate_cv: Condvar::new(),\n             dog: Mutex::new(Dog {\n@@ -539,12 +646,12 @@ impl BrokerConn {\n             return Ok(ConnWrite::Superseded);\n         }\n \n-        // Once-per-conn `write-start` lifecycle record (leg d): marks the conn's\n-        // first committed write, so a conn that later retires can be correlated\n-        // to when it went active (fresh-carrier churn shows as open→start→retire\n-        // triplets with fresh ids).\n-        if !inner.first_write_logged.swap(true, Ordering::AcqRel) {\n-            inner.lifecycle(\"write-start\", \"\");\n+        // The gate serializes first-write initialization. Save the instant\n+        // BEFORE I/O so a later failure still names when this conn went active.\n+        // [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+        if inner.first_write.get().is_none() {\n+            let _ = inner.first_write.set(LogTime::now());\n+            inner.lifecycle_log.healthy(true);\n         }\n \n         // ── Check out the half + arm the watchdog, then write OUT of locks. ──\n@@ -638,17 +745,20 @@ impl BrokerConn {\n \n impl Drop for BrokerConn {\n     fn drop(&mut self) {\n-        // Once-per-conn `transport-close` lifecycle record (leg d): the handle\n-        // is about to close — the terminal record of this conn id's life, with\n-        // how it ended (organic EOF vs poisoned vs deadline-poisoned).\n-        self.inner.lifecycle(\n-            \"transport-close\",\n-            &format!(\n-                \"poisoned={} timed_out={}\",\n-                self.inner.poisoned.load(Ordering::Acquire),\n-                self.inner.timed_out.load(Ordering::Acquire)\n-            ),\n-        );\n+        // Only a healthy physical close is counted instead of logged per conn.\n+        // Error/timeout closes keep their full attribution and original start.\n+        // [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+        if self.inner.poisoned.load(Ordering::Acquire) {\n+            self.inner.lifecycle(\n+                \"transport-close\",\n+                &format!(\n+                    \"poisoned=true timed_out={}\",\n+                    self.inner.timed_out.load(Ordering::Acquire)\n+                ),\n+            );\n+        } else {\n+            self.inner.lifecycle_log.healthy(false);\n+        }\n         {\n             let mut d = recover(&self.inner.dog);\n             d.shutdown = true;\n@@ -688,6 +798,17 @@ mod tests {\n     /// primitive, the client end returned so a test can hold it alive or drop\n     /// it to force an ORGANIC write failure.\n     fn make_conn() -> (BrokerConn, Stream) {\n+        make_conn_with_log(Arc::new(LifecycleLog::new()))\n+    }\n+\n+    fn make_conn_with_log(log: Arc<LifecycleLog>) -> (BrokerConn, Stream) {\n+        let (conn, client, _recv) = make_duplex_with_log(log);\n+        (conn, client)\n+    }\n+\n+    fn make_duplex_with_log(\n+        log: Arc<LifecycleLog>,\n+    ) -> (BrokerConn, Stream, interprocess::local_socket::RecvHalf) {\n         static SEQ: AtomicU32 = AtomicU32::new(0);","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120,121,122,123,124,125,126,127,128,129,130,131,132,133,134,135,136,137,138,139,140,141,142,143,144,145,146,147,148,149,150,151,152,153,154,155,156,157,158,159,160,161,162,163,164,165,166,167,168,169,170,171,172,173,174,175,176,177,178,179,180,181,182,183,184,185,186,187,188,189,190,191,192,193,194,195,196,197,198,199,200,201,202,203,204,205,206,207,208,209,210,211,212,213,214,215,216,217,218,219,220,221,222,223,224,225,226,227,228,229,230,231,232,233,234,235,236,237,238,239,240,241,242,243,244,245,246,247,248,249,250,251,252,253,254,255,256,257,258,259,260,261,262,263,264,265,266,267,268,269,270,271,272,273,274,275,276,277,278,279,280,281,282,283,284,285,286,287,288,289,290,291,292,293,294,295,296,297,298,299,300]},"truncation":{"content":"diff --git a/CHANGELOG.md b/CHANGELOG.md\nindex eafe18d6..5b4a7156 100644\n--- a/CHANGELOG.md\n+++ b/CHANGELOG.md\n@@ -10,6 +10,17 @@ breaks something, or changes the observable behavior of existing surfaces\n broadly; **patch** for fixes, and for additive opt-in capability — a new key,\n flag, or page that no existing user can encounter without opting into it.\n \n+## [Unreleased]\n+\n+### Fixed\n+\n+<!-- [doc->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED] -->\n+- Daemon logs summarize healthy IPC connection starts and closes once per\n+  minute instead of logging each one. Failures and subscriber lifecycle\n+  events keep their per-connection attribution, including the original\n+  first-write time. Livehost session polling now reuses its healthy\n+  connection and names the reason when it opens a replacement.\n+\n ## [0.68.0] - 2026-09-08\n \n Web serving. Files, directories, adapter documentation and the changelog are\ndiff --git a/crates/spt-daemon/src/conn.rs b/crates/spt-daemon/src/conn.rs\nindex 9c51ad4e..163fecb6 100644\n--- a/crates/spt-daemon/src/conn.rs\n+++ b/crates/spt-daemon/src/conn.rs\n@@ -61,7 +61,7 @@\n \n use std::io;\n use std::sync::atomic::{AtomicBool, AtomicU64, Ordering};\n-use std::sync::{Arc, Condvar, Mutex, MutexGuard, OnceLock};\n+use std::sync::{Arc, Condvar, LazyLock, Mutex, MutexGuard, OnceLock};\n use std::thread::{self, JoinHandle};\n use std::time::{Duration, Instant, SystemTime, UNIX_EPOCH};\n \n@@ -99,12 +99,107 @@ pub(crate) fn init_log_anchor() {\n /// to field incident times (wall).\n // [impl->REQ-CONN-POISON-ATTRIBUTION]\n pub(crate) fn log_stamp() -> String {\n-    let wall_ms = SystemTime::now()\n-        .duration_since(UNIX_EPOCH)\n-        .map(|d| d.as_millis())\n-        .unwrap_or(0);\n-    let mono_ms = MONO_ANCHOR.get_or_init(Instant::now).elapsed().as_millis();\n-    format!(\"wall_ms={wall_ms} mono_ms={mono_ms}\")\n+    let stamp = LogTime::now();\n+    format!(\"wall_ms={} mono_ms={}\", stamp.wall_ms, stamp.mono_ms)\n+}\n+\n+/// Retain the original write instant even when its healthy line is aggregated.\n+#[derive(Clone, Copy)]\n+struct LogTime {\n+    wall_ms: u128,\n+    mono_ms: u128,\n+}\n+\n+impl LogTime {\n+    fn now() -> Self {\n+        Self {\n+            wall_ms: SystemTime::now()\n+                .duration_since(UNIX_EPOCH)\n+                .map(|d| d.as_millis())\n+                .unwrap_or(0),\n+            mono_ms: MONO_ANCHOR.get_or_init(Instant::now).elapsed().as_millis(),\n+        }\n+    }\n+}\n+\n+const HEALTHY_LOG_INTERVAL: Duration = Duration::from_secs(60);\n+\n+struct HealthyWindow {\n+    since: Instant,\n+    starts: u64,\n+    closes: u64,\n+}\n+\n+/// One process-wide counter window, not one limiter per connection. No thread,\n+/// queue or per-conn healthy string: the next event after 60s emits the totals.\n+/// Exceptional and role-transition records bypass this window entirely.\n+// [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+struct LifecycleLog {\n+    healthy: Mutex<HealthyWindow>,\n+    #[cfg(test)]\n+    capture: Option<Mutex<Vec<String>>>,\n+    #[cfg(test)]\n+    clock: Option<Mutex<Instant>>,\n+}\n+\n+impl LifecycleLog {\n+    fn new() -> Self {\n+        Self {\n+            healthy: Mutex::new(HealthyWindow {\n+                since: Instant::now(),\n+                starts: 0,\n+                closes: 0,\n+            }),\n+            #[cfg(test)]\n+            capture: None,\n+            #[cfg(test)]\n+            clock: None,\n+        }\n+    }\n+\n+    #[cfg(test)]\n+    fn captured(now: Instant) -> Self {\n+        let mut log = Self::new();\n+        log.healthy.get_mut().unwrap().since = now;\n+        log.capture = Some(Mutex::new(Vec::new()));\n+        log.clock = Some(Mutex::new(now));\n+        log\n+    }\n+\n+    fn emit(&self, args: std::fmt::Arguments<'_>) {\n+        #[cfg(test)]\n+        if let Some(lines) = &self.capture {\n+            recover(lines).push(args.to_string());\n+            return;\n+        }\n+        spt_proto::emit_line_err!(\"{args}\");\n+    }\n+\n+    fn healthy(&self, start: bool) {\n+        let mut window = recover(&self.healthy);\n+        // Read time under the counter lock so concurrent callers cannot move\n+        // the window backwards after waiting for a newer caller to flush.\n+        let now = Instant::now();\n+        #[cfg(test)]\n+        let now = self.clock.as_ref().map(|c| *recover(c)).unwrap_or(now);\n+        if start {\n+            window.starts = window.starts.saturating_add(1);\n+        } else {\n+            window.closes = window.closes.saturating_add(1);\n+        }\n+        let elapsed = now.saturating_duration_since(window.since);\n+        if elapsed < HEALTHY_LOG_INTERVAL {\n+            return;\n+        }\n+        let (starts, closes) = (window.starts, window.closes);\n+        *window = HealthyWindow { since: now, starts: 0, closes: 0 };\n+        drop(window);\n+        self.emit(format_args!(\n+            \"CONN_LIFECYCLE_SUMMARY: scope=broker-ipc {} interval_ms={} starts={starts} closes={closes} \\\n+             [REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\",\n+            log_stamp(), elapsed.as_millis()\n+        ));\n+    }\n }\n \n /// Attribution-label cap: facts accrete per role attach (a long-lived brain\n@@ -144,8 +239,9 @@ struct Inner {\n     /// Accreted attribution facts (subscriber role, endpoint/session where\n     /// known — [`BrokerConn::describe`]). Bounded by [`LABEL_CAP`].\n     label: Mutex<String>,\n-    /// Latch: the once-per-conn `write-start` lifecycle record fired.\n-    first_write_logged: AtomicBool,\n+    /// Once-per-conn original first committed write time; retained on failures.\n+    first_write: OnceLock<LogTime>,\n+    lifecycle_log: Arc<LifecycleLog>,\n     gate: Mutex<Gate>,\n     gate_cv: Condvar,\n     dog: Mutex<Dog>,\n@@ -223,7 +319,13 @@ impl Inner {\n     fn attribution(&self) -> String {\n         let label = recover(&self.label);\n         let facts: &str = if label.is_empty() { \"role=unattributed\" } else { &label };\n-        format!(\"conn={} {} {}\", self.id, facts, log_stamp())\n+        match self.first_write.get() {\n+            Some(start) => format!(\n+                \"conn={} {} {} first_write_wall_ms={} first_write_mono_ms={}\",\n+                self.id, facts, log_stamp(), start.wall_ms, start.mono_ms\n+            ),\n+            None => format!(\"conn={} {} {} first_write=none\", self.id, facts, log_stamp()),\n+        }\n     }\n \n     /// Render the once-per-conn retirement record — the F-039 token split\n@@ -261,18 +363,17 @@ impl Inner {\n         }\n     }\n \n-    /// One BOUNDED per-conn lifecycle record (leg d, doyle-confirmed\n-    /// UNCONDITIONAL): `write-start` (first write only), `transport-close`\n-    /// (drop), `writer-exit` / `*-attach` / `*-replaced` (emitted by the\n-    /// broker's sink machinery through [`BrokerConn::lifecycle_event`]). The\n-    /// timeout/cancel leg is the retirement record itself.\n+    /// Per-conn exceptional and role-transition records remain unconditional.\n+    /// releases#286 supersedes only healthy start/close emission, which uses\n+    /// the process-wide interval counter instead. Retirement retains the\n+    /// original first-write time even when the start line was aggregated.\n     // [impl->REQ-CONN-POISON-ATTRIBUTION]\n     fn lifecycle(&self, event: &str, extra: &str) {\n         let sep = if extra.is_empty() { \"\" } else { \" \" };\n-        spt_proto::emit_line_err!(\n+        self.lifecycle_log.emit(format_args!(\n             \"CONN_LIFECYCLE: {} event={event}{sep}{extra} [REQ-CONN-POISON-ATTRIBUTION]\",\n             self.attribution()\n-        );\n+        ));\n     }\n \n     /// Abort the connection's in-flight I/O OUT OF BAND (both directions —\n@@ -287,10 +388,10 @@ impl Inner {\n             // First poison of this conn: loud, once. The token is class-split\n             // (F-039 leg a): deadline → CONN_WRITE_POISONED (the wedge\n             // observable), organic fast-fail → CONN_WRITE_RETIRED.\n-            eprintln!(\n+            self.lifecycle_log.emit(format_args!(\n                 \"{}\",\n                 self.render_retirement(self.timed_out.load(Ordering::Acquire), cause)\n-            );\n+            ));\n         }\n         #[cfg(windows)]\n         // SAFETY: scalar kernel32 call on a handle whose owning half is kept\n@@ -424,11 +525,17 @@ impl BrokerConn {\n     /// deadline (gate-wait + OS write completion) — the broker passes its\n     /// `brain_write_deadline()` (`SPT_BRAIN_WRITE_DEADLINE_MS` knob).\n     pub(crate) fn new(half: SendHalf, bound: Duration) -> Self {\n+        static LOG: LazyLock<Arc<LifecycleLog>> = LazyLock::new(|| Arc::new(LifecycleLog::new()));\n+        Self::with_log(half, bound, Arc::clone(&LOG))\n+    }\n+\n+    fn with_log(half: SendHalf, bound: Duration, lifecycle_log: Arc<LifecycleLog>) -> Self {\n         let raw = raw_of(&half);\n         let inner = Arc::new(Inner {\n             id: CONN_ID_SEQ.fetch_add(1, Ordering::Relaxed) + 1,\n             label: Mutex::new(String::new()),\n-            first_write_logged: AtomicBool::new(false),\n+            first_write: OnceLock::new(),\n+            lifecycle_log,\n             gate: Mutex::new(Gate { half: Some(half) }),\n             gate_cv: Condvar::new(),\n             dog: Mutex::new(Dog {\n@@ -539,12 +646,12 @@ impl BrokerConn {\n             return Ok(ConnWrite::Superseded);\n         }\n \n-        // Once-per-conn `write-start` lifecycle record (leg d): marks the conn's\n-        // first committed write, so a conn that later retires can be correlated\n-        // to when it went active (fresh-carrier churn shows as open→start→retire\n-        // triplets with fresh ids).\n-        if !inner.first_write_logged.swap(true, Ordering::AcqRel) {\n-            inner.lifecycle(\"write-start\", \"\");\n+        // The gate serializes first-write initialization. Save the instant\n+        // BEFORE I/O so a later failure still names when this conn went active.\n+        // [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+        if inner.first_write.get().is_none() {\n+            let _ = inner.first_write.set(LogTime::now());\n+            inner.lifecycle_log.healthy(true);\n         }\n \n         // ── Check out the half + arm the watchdog, then write OUT of locks. ──\n@@ -638,17 +745,20 @@ impl BrokerConn {\n \n impl Drop for BrokerConn {\n     fn drop(&mut self) {\n-        // Once-per-conn `transport-close` lifecycle record (leg d): the handle\n-        // is about to close — the terminal record of this conn id's life, with\n-        // how it ended (organic EOF vs poisoned vs deadline-poisoned).\n-        self.inner.lifecycle(\n-            \"transport-close\",\n-            &format!(\n-                \"poisoned={} timed_out={}\",\n-                self.inner.poisoned.load(Ordering::Acquire),\n-                self.inner.timed_out.load(Ordering::Acquire)\n-            ),\n-        );\n+        // Only a healthy physical close is counted instead of logged per conn.\n+        // Error/timeout closes keep their full attribution and original start.\n+        // [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+        if self.inner.poisoned.load(Ordering::Acquire) {\n+            self.inner.lifecycle(\n+                \"transport-close\",\n+                &format!(\n+                    \"poisoned=true timed_out={}\",\n+                    self.inner.timed_out.load(Ordering::Acquire)\n+                ),\n+            );\n+        } else {\n+            self.inner.lifecycle_log.healthy(false);\n+        }\n         {\n             let mut d = recover(&self.inner.dog);\n             d.shutdown = true;\n@@ -688,6 +798,17 @@ mod tests {\n     /// primitive, the client end returned so a test can hold it alive or drop\n     /// it to force an ORGANIC write failure.\n     fn make_conn() -> (BrokerConn, Stream) {\n+        make_conn_with_log(Arc::new(LifecycleLog::new()))\n+    }\n+\n+    fn make_conn_with_log(log: Arc<LifecycleLog>) -> (BrokerConn, Stream) {\n+        let (conn, client, _recv) = make_duplex_with_log(log);\n+        (conn, client)\n+    }\n+\n+    fn make_duplex_with_log(\n+        log: Arc<LifecycleLog>,\n+    ) -> (BrokerConn, Stream, interprocess::local_socket::RecvHalf) {\n         static SEQ: AtomicU32 = AtomicU32::new(0);","truncated":true,"truncatedBy":"lines","totalLines":659,"totalBytes":12067,"outputLines":300,"outputBytes":12067,"lastLinePartial":false,"firstLineExceedsLimit":false},"meta":{"source":{"type":"internal","value":"artifact://65"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":659,"totalBytes":12067,"outputLines":300,"outputBytes":12067,"maxBytes":153600,"shownRange":{"start":1,"end":300},"nextOffset":301}}},"isError":false,"timestamp":1788999505100}}
{"type":"message","id":"ef338dcd","parentId":"ec38e681","timestamp":"2026-09-10T00:18:25.120Z","message":{"role":"toolResult","toolCallId":"call_I7Tc7wlUioMWDMzQLLS5aRsh|fc_0dde72effda0ad21016aa1f751055c87d2af276890e6c623d3","toolName":"grep","content":[{"type":"text","text":"# .worktrees/286-churn/\n## AGENTS.md\n 2|\n*3|Harness-independent core for an agent ecosystem (messaging, live-agent lifecycle, terminal hosting, P2P networking, runtime-manifest harness contract). Rust crates + a single `spt` binary. Clean-room rebuild of `claude_skill_owl`.\n 4|\n 5|**Orientation — read before working:** `PRD.md` (requirements), `ROADMAP.md` (build path), `CONTEXT.md` (glossary/model, authoritative for meaning), `docs/adr/` (decisions), `docs/KNOWN-HAZARDS.md` (invariants we must not re-break), `docs/{STORAGE,MANIFEST,CONTEXT-MEMORY,DOCS-STRATEGY,TRACEABILITY}.md`.\n 6|\n...\n 8|\n*9|For work needing Linux tests and proof, this machine is authorized to ssh into a Linux box: `reavus@kitsubito`\n 10|\n 11|IMPORTANT: All agents you work alongside run on spt-core itself — the legacy-SPT (`$OWL`/`$LIVE`) era ended with the 2026-07 fleet migration; those commands are gone. Reach other agents with `spt send` (or the `/sptc:send` skill); keep message bodies short, and send long bodies from a file rather than inline. Everything else can be learned from the corresponding `/sptc:` skill doc.\n 12|\n...\n 16|\n*17|1. **Tag evidence in the same change.** When you write a function/test/doc-section that satisfies a requirement stage, add its tag in that same commit:\n 18|   `// [impl->REQ-FOO]` · `// [unit->REQ-FOO]` · `<!-- [doc->REQ-FOO] -->`\n 19|   Stages: `doc` / `impl` / `unit` / `int`. Tag *on or immediately above* the real evidence — never at file tops to satisfy coverage.\n 20|2. **Run `traceable-reqs check` before declaring work done.** Exit-1 means missing/invalid evidence — fix it, don't ship it.\n 21|3. **New requirement → add it to `traceable-reqs.toml` first** (with a `REQ-*` id), then satisfy it. No untracked work; no untagged evidence.\n*22|4. **KNOWN-HAZARDS are `REQ-HAZARD-*` requirements** — each needs a test before it's \"covered.\" Treat the hazard list as a conformance checklist you must satisfy, not advice.\n 23|5. **Activate, don't pre-fail.** Requirements you aren't yet working stay `required_stages = []`. Activate (set real stages) only when starting the milestone that delivers them.\n 24|\n 25|## Issue tracking (unique case)\n...\n 39|<!-- [doc->REQ-POOL-FOREIGN-TREE-REFUSED] -->\n*40|- **Build-cache pools (binding, releases#103):** a cargo `target/` directory belongs to ONE source tree. Two trees building into one pool overwrite each other's artifacts and fabricate reds that belong to neither — the v0.51.0 W1/H1 `E0308` that vanished on re-run. **The rule is about the RELATIONSHIP, not the path:** a junction and a `CARGO_TARGET_DIR` env var reach the same hazard, and the env var leaves no directory entry to notice afterwards, so \"is this dir under `.worktrees/`\" never was the quest...\n 41|- **Target-tree teardown (releases#103, same lane):** never delete a `target/` you have not classified. Check OUTBOUND first (`Get-Item -Force`: a junction is deleted as a link, a real dir as a tree), then sweep INBOUND for reparse points aimed at it before removing a real one, and measure size and free space before and after so the reclaim is a number rather than a claim. Reap the target SUBTREE, never a session or worktree directory around it. Do not point `CARGO_TARGET_DIR` into a scratchpad (v0.50.0 ...\n*42|- **CI / merge strategy:** golden CI (ADR-0050, ruled 2026-07-27) — milestone-batch golden runs, thin-lane PRs, ff-only main (tested sha == merged sha). Piloting from the next milestone; read the ADR before shaping a PR lane.\n 43|\n 44|If you finish a significant body of work without need for user intervention, or if your context gets too high, you can clear your own context and keep moving:\n 45|1. Create a JIT plan for the next immediate body of work, if it isn't already planned\n\n## docs/\n### KNOWN-HAZARDS.md\n 2|\n*3|Hard-won edge cases harvested from the sister project (`claude_skill_owl`, ~80 commits / 12+ phases / multiple production incidents). Per ADR-0001, this is a **test checklist for the spt-core rebuild** — the clean-room rebuild must re-satisfy each invariant rather than re-discover the bug.\n 4|\n 5|**Architecture-translation note.** The sister project runs poll listeners and Psyche wrappers as *separate processes*. spt-core consolidates both into the one `spt-daemon` (brain), with a stable broker beneath it (ADR-0004). Many hazards below were inter-process races in the sister project; in spt-core some become intra-daemon concerns (potentially easier) while others move to the daemon↔broker IPC boundary or the network boundary (potentially new failure surface). Each entry notes the mapping where it...\n 6|\n...\n 14|- **spt-core mapping:** in-daemon now (no separate wrapper), but the ordering invariant is identical — orphan/teardown logic must re-evaluate liveness after the grace wait, not before enqueue.\n*15|- **Sister cite:** `src/live/wrapper/orphan.rs:201-259` (sleep@209 precedes compose@231-251); tests T-grace-recovery:576, T-still-gone-recheck:618.\n 16|\n 17|### 1.2 Poll-rewrite race & info.json mid-write reads\n 18|- **Failure:** `info.json` written by the wrapper mid-iteration while a list/classify command reads it → torn read, misclassification.\n 19|- **Invariant:** consult liveness via the supervisor (`is_wrapper_alive`-equivalent) before any grace gate; reads of state files must tolerate concurrent writes (atomic write + rename, or read-retry).\n*20|- **spt-core mapping:** the daemon owns both writer and reader → use in-process locking/snapshotting instead of racing on disk. Cross-node registry reads remain eventually-consistent and must tolerate staleness.\n 21|- **Sister cite:** `src/common/list_filter.rs:100-150`; `src/owl/poll.rs:141`.\n 22|\n*23|### 1.3 Stale `index.lock` wedge from prior git crash\n*24|- **Failure:** crashed git leaves a 0-byte `index.lock` in a psyche tracked worktree; every later commit blocks forever.\n*25|- **Invariant:** on daemon boot, sweep seed + all agent/project worktrees for stale locks (0 bytes, mtime > 60s) and remove; leave live locks alone.\n*26|- **spt-core mapping:** cross-node Psyche sync (ADR-0002/0003) replaces git-repo sync, so the *git* lock may disappear — but any equivalent lockfile in the new sync mechanism needs the same stale-sweep on boot.\n*27|- **Sister cite:** CHANGELOG v1.11.20 \"Stale `index.lock`\"; `src/common/git.rs`.\n 28|\n 29|### 1.4 Deferred spool rows must not leak to the event stream\n 30|- **Failure:** a hook spools a deferred (spool-only, no TCP wake) notice; startup `drain_all` flushes ALL rows including deferred → event emitted at wrong time/priority.\n...\n 34|\n*35|### 1.5 Worker (working-perch) lifecycle path consistency\n 36|- **Failure:** subagent-start creates the perch at one path layout; later hooks read it at another → not found; stop-hook scan misses nested perches.\n 37|- **Invariant:** all Worker/Psyche child-perch path composition routes through one central resolver; no divergent path construction across hooks.\n 38|- **spt-core mapping:** `Worker` is a day-one endpoint type; the daemon owns the registry, so perch location is a registry lookup, not ad-hoc path math. Single source of truth for instance→location.\n...\n 44|\n*45|### 2.1 Parent PID over ephemeral poll PID\n*46|- **Failure:** orphan check polls an ephemeral listener PID; it dies and is recycled (esp. Windows); a foreign process with the recycled PID reads as alive → false-positive teardown (or false-negative).\n*47|- **Invariant:** prefer the stable harness-session PID (`parent_pid`) over any ephemeral process PID for liveness; minimal `info.json` for supervisor-owned perches to avoid stale leaks.\n*48|- **spt-core mapping:** session binding (parent-process-tree anchor) still applies for harness-hosted topology. For spt-hosted sessions the broker holds the child directly → liveness is the broker's held-handle state, more reliable than PID polling.\n 49|- **Sister cite:** `src/live/wrapper/orphan.rs:141-161`; CHANGELOG v1.11.20.\n 50|\n 51|### 2.2 Stdin session_id precedence over env\n...\n 60|- **spt-core mapping:** CRITICAL — self-update (ADR-0004) makes handoff routine. The broker↔brain IPC and any brain-relaunch argv must be versioned and forward/backward tolerant (a newer brain talks to an older broker). This is the single most update-frequency-sensitive invariant.\n*61|- **Sister cite:** `src/live/wrapper/lifecycle.rs:17-106`; `src/cli.rs` defaults; CHANGELOG v1.11.10.\n 62|\n 63|### 2.4 Generation `gen_start` always = now() on cold-start AND handoff\n 64|- **Failure:** stale gen_start from a rehydrated state file fires time-based discriminators on the new process.\n*65|- **Invariant:** wall-clock `gen_start` is set to `now()` on both cold-start and handoff; generation counter increments on every start/revive; session UUID captured fresh and carried so the resumed mind distinguishes \"same gen continuing\" vs \"new gen born\".\n 66|- **spt-core mapping:** carries to the daemon's per-instance generation tracking.\n*67|- **Restoration D3/D4 (ADR-0018):** the generation *counter* custody moved to the broker (D3-2 — it observes every brain spawn, planned or crash, and hands `{generation, start-reason}` at spawn; `gen_start` stays `now()`-fresh, never rehydrated). The brain→brain **`BrainState` *message*** (`{session_id, generation, next_seq, gen_start_ms}`) that previously carried continuity across a handoff is **retired from the production path in D4-2**: a brain the supervisor respawns cold-starts and reconstructs ...\n*68|- **Sister cite:** `src/live/wrapper/lifecycle.rs:70`; `src/common/wrapper_state.rs`.\n 69|\n*70|### 2.5 Daemon-hosted endpoints have no dedicated liveness PID\n*71|- **Failure:** the sister evaluates Psyche/perch liveness via a dedicated process PID — the wrapper's own pid in `info.json`, checked with `is_process_alive`. Under ADR-0004 the Psyche (and any spt-hosted Self) is a **loop inside the daemon**, not a separate process: it holds no dedicated pid, and its `claude`/summarizer subprocess is ephemeral (spawned per pulse/commune, then exits). If a daemon-hosted perch's `info.json` carries the **daemon's** pid, then *every* hosted endpoint shares one pid, and `...\n*72|- **Invariant:** for **daemon-hosted** perches (Psyche; spt-hosted Self), liveness is the **daemon's authoritative in-memory endpoint table + a `status` field** on `info.json` (`online|offline|…`), **never** `is_process_alive(info.pid)`. `info.pid` for a daemon-hosted perch is at most a *hosted-by-daemon* marker (the daemon pid), not a liveness signal; registry stale-clean for these rows keys on the daemon's endpoint table, not per-row pid. This reuses the pattern already specified for **Shells** (`inf...\n\n[Some lines truncated to 512 chars]"}],"details":{"scopePath":".worktrees/286-churn/AGENTS.md, .worktrees/286-churn/docs/KNOWN-HAZARDS.md","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":26,"fileCount":2,"files":[".worktrees/286-churn/AGENTS.md",".worktrees/286-churn/docs/KNOWN-HAZARDS.md"],"fileMatches":[{"path":".worktrees/286-churn/AGENTS.md","count":6},{"path":".worktrees/286-churn/docs/KNOWN-HAZARDS.md","count":20}],"truncated":true,"perFileLimitReached":20,"displayContent":"# .worktrees/286-churn/\n## AGENTS.md\n  2│\n *3│Harness-independent core for an agent ecosystem (messaging, live-agent lifecycle, terminal hosting, P2P networking, runtime-manifest harness contract). Rust crates + a single `spt` binary. Clean-room rebuild of `claude_skill_owl`.\n  4│\n  5│**Orientation — read before working:** `PRD.md` (requirements), `ROADMAP.md` (build path), `CONTEXT.md` (glossary/model, authoritative for meaning), `docs/adr/` (decisions), `docs/KNOWN-HAZARDS.md` (invariants we must not re-break), `docs/{STORAGE,MANIFEST,CONTEXT-MEMORY,DOCS-STRATEGY,TRACEABILITY}.md`.\n  6│\n   │...\n  8│\n *9│For work needing Linux tests and proof, this machine is authorized to ssh into a Linux box: `reavus@kitsubito`\n 10│\n 11│IMPORTANT: All agents you work alongside run on spt-core itself — the legacy-SPT (`$OWL`/`$LIVE`) era ended with the 2026-07 fleet migration; those commands are gone. Reach other agents with `spt send` (or the `/sptc:send` skill); keep message bodies short, and send long bodies from a file rather than inline. Everything else can be learned from the corresponding `/sptc:` skill doc.\n 12│\n   │...\n 16│\n*17│1. **Tag evidence in the same change.** When you write a function/test/doc-section that satisfies a requirement stage, add its tag in that same commit:\n 18│   `// [impl->REQ-FOO]` · `// [unit->REQ-FOO]` · `<!-- [doc->REQ-FOO] -->`\n 19│   Stages: `doc` / `impl` / `unit` / `int`. Tag *on or immediately above* the real evidence — never at file tops to satisfy coverage.\n 20│2. **Run `traceable-reqs check` before declaring work done.** Exit-1 means missing/invalid evidence — fix it, don't ship it.\n 21│3. **New requirement → add it to `traceable-reqs.toml` first** (with a `REQ-*` id), then satisfy it. No untracked work; no untagged evidence.\n*22│4. **KNOWN-HAZARDS are `REQ-HAZARD-*` requirements** — each needs a test before it's \"covered.\" Treat the hazard list as a conformance checklist you must satisfy, not advice.\n 23│5. **Activate, don't pre-fail.** Requirements you aren't yet working stay `required_stages = []`. Activate (set real stages) only when starting the milestone that delivers them.\n 24│\n 25│## Issue tracking (unique case)\n   │...\n 39│<!-- [doc->REQ-POOL-FOREIGN-TREE-REFUSED] -->\n*40│- **Build-cache pools (binding, releases#103):** a cargo `target/` directory belongs to ONE source tree. Two trees building into one pool overwrite each other's artifacts and fabricate reds that belong to neither — the v0.51.0 W1/H1 `E0308` that vanished on re-run. **The rule is about the RELATIONSHIP, not the path:** a junction and a `CARGO_TARGET_DIR` env var reach the same hazard, and the env var leaves no directory entry to notice afterwards, so \"is this dir under `.worktrees/`\" never was the quest...\n 41│- **Target-tree teardown (releases#103, same lane):** never delete a `target/` you have not classified. Check OUTBOUND first (`Get-Item -Force`: a junction is deleted as a link, a real dir as a tree), then sweep INBOUND for reparse points aimed at it before removing a real one, and measure size and free space before and after so the reclaim is a number rather than a claim. Reap the target SUBTREE, never a session or worktree directory around it. Do not point `CARGO_TARGET_DIR` into a scratchpad (v0.50.0 ...\n*42│- **CI / merge strategy:** golden CI (ADR-0050, ruled 2026-07-27) — milestone-batch golden runs, thin-lane PRs, ff-only main (tested sha == merged sha). Piloting from the next milestone; read the ADR before shaping a PR lane.\n 43│\n 44│If you finish a significant body of work without need for user intervention, or if your context gets too high, you can clear your own context and keep moving:\n 45│1. Create a JIT plan for the next immediate body of work, if it isn't already planned\n\n## docs/\n### KNOWN-HAZARDS.md\n  2│\n *3│Hard-won edge cases harvested from the sister project (`claude_skill_owl`, ~80 commits / 12+ phases / multiple production incidents). Per ADR-0001, this is a **test checklist for the spt-core rebuild** — the clean-room rebuild must re-satisfy each invariant rather than re-discover the bug.\n  4│\n  5│**Architecture-translation note.** The sister project runs poll listeners and Psyche wrappers as *separate processes*. spt-core consolidates both into the one `spt-daemon` (brain), with a stable broker beneath it (ADR-0004). Many hazards below were inter-process races in the sister project; in spt-core some become intra-daemon concerns (potentially easier) while others move to the daemon↔broker IPC boundary or the network boundary (potentially new failure surface). Each entry notes the mapping where it...\n  6│\n   │...\n 14│- **spt-core mapping:** in-daemon now (no separate wrapper), but the ordering invariant is identical — orphan/teardown logic must re-evaluate liveness after the grace wait, not before enqueue.\n*15│- **Sister cite:** `src/live/wrapper/orphan.rs:201-259` (sleep@209 precedes compose@231-251); tests T-grace-recovery:576, T-still-gone-recheck:618.\n 16│\n 17│### 1.2 Poll-rewrite race & info.json mid-write reads\n 18│- **Failure:** `info.json` written by the wrapper mid-iteration while a list/classify command reads it → torn read, misclassification.\n 19│- **Invariant:** consult liveness via the supervisor (`is_wrapper_alive`-equivalent) before any grace gate; reads of state files must tolerate concurrent writes (atomic write + rename, or read-retry).\n*20│- **spt-core mapping:** the daemon owns both writer and reader → use in-process locking/snapshotting instead of racing on disk. Cross-node registry reads remain eventually-consistent and must tolerate staleness.\n 21│- **Sister cite:** `src/common/list_filter.rs:100-150`; `src/owl/poll.rs:141`.\n 22│\n*23│### 1.3 Stale `index.lock` wedge from prior git crash\n*24│- **Failure:** crashed git leaves a 0-byte `index.lock` in a psyche tracked worktree; every later commit blocks forever.\n*25│- **Invariant:** on daemon boot, sweep seed + all agent/project worktrees for stale locks (0 bytes, mtime > 60s) and remove; leave live locks alone.\n*26│- **spt-core mapping:** cross-node Psyche sync (ADR-0002/0003) replaces git-repo sync, so the *git* lock may disappear — but any equivalent lockfile in the new sync mechanism needs the same stale-sweep on boot.\n*27│- **Sister cite:** CHANGELOG v1.11.20 \"Stale `index.lock`\"; `src/common/git.rs`.\n 28│\n 29│### 1.4 Deferred spool rows must not leak to the event stream\n 30│- **Failure:** a hook spools a deferred (spool-only, no TCP wake) notice; startup `drain_all` flushes ALL rows including deferred → event emitted at wrong time/priority.\n   │...\n 34│\n*35│### 1.5 Worker (working-perch) lifecycle path consistency\n 36│- **Failure:** subagent-start creates the perch at one path layout; later hooks read it at another → not found; stop-hook scan misses nested perches.\n 37│- **Invariant:** all Worker/Psyche child-perch path composition routes through one central resolver; no divergent path construction across hooks.\n 38│- **spt-core mapping:** `Worker` is a day-one endpoint type; the daemon owns the registry, so perch location is a registry lookup, not ad-hoc path math. Single source of truth for instance→location.\n   │...\n 44│\n*45│### 2.1 Parent PID over ephemeral poll PID\n*46│- **Failure:** orphan check polls an ephemeral listener PID; it dies and is recycled (esp. Windows); a foreign process with the recycled PID reads as alive → false-positive teardown (or false-negative).\n*47│- **Invariant:** prefer the stable harness-session PID (`parent_pid`) over any ephemeral process PID for liveness; minimal `info.json` for supervisor-owned perches to avoid stale leaks.\n*48│- **spt-core mapping:** session binding (parent-process-tree anchor) still applies for harness-hosted topology. For spt-hosted sessions the broker holds the child directly → liveness is the broker's held-handle state, more reliable than PID polling.\n 49│- **Sister cite:** `src/live/wrapper/orphan.rs:141-161`; CHANGELOG v1.11.20.\n 50│\n 51│### 2.2 Stdin session_id precedence over env\n   │...\n 60│- **spt-core mapping:** CRITICAL — self-update (ADR-0004) makes handoff routine. The broker↔brain IPC and any brain-relaunch argv must be versioned and forward/backward tolerant (a newer brain talks to an older broker). This is the single most update-frequency-sensitive invariant.\n*61│- **Sister cite:** `src/live/wrapper/lifecycle.rs:17-106`; `src/cli.rs` defaults; CHANGELOG v1.11.10.\n 62│\n 63│### 2.4 Generation `gen_start` always = now() on cold-start AND handoff\n 64│- **Failure:** stale gen_start from a rehydrated state file fires time-based discriminators on the new process.\n*65│- **Invariant:** wall-clock `gen_start` is set to `now()` on both cold-start and handoff; generation counter increments on every start/revive; session UUID captured fresh and carried so the resumed mind distinguishes \"same gen continuing\" vs \"new gen born\".\n 66│- **spt-core mapping:** carries to the daemon's per-instance generation tracking.\n*67│- **Restoration D3/D4 (ADR-0018):** the generation *counter* custody moved to the broker (D3-2 — it observes every brain spawn, planned or crash, and hands `{generation, start-reason}` at spawn; `gen_start` stays `now()`-fresh, never rehydrated). The brain→brain **`BrainState` *message*** (`{session_id, generation, next_seq, gen_start_ms}`) that previously carried continuity across a handoff is **retired from the production path in D4-2**: a brain the supervisor respawns cold-starts and reconstructs ...\n*68│- **Sister cite:** `src/live/wrapper/lifecycle.rs:70`; `src/common/wrapper_state.rs`.\n 69│\n*70│### 2.5 Daemon-hosted endpoints have no dedicated liveness PID\n*71│- **Failure:** the sister evaluates Psyche/perch liveness via a dedicated process PID — the wrapper's own pid in `info.json`, checked with `is_process_alive`. Under ADR-0004 the Psyche (and any spt-hosted Self) is a **loop inside the daemon**, not a separate process: it holds no dedicated pid, and its `claude`/summarizer subprocess is ephemeral (spawned per pulse/commune, then exits). If a daemon-hosted perch's `info.json` carries the **daemon's** pid, then *every* hosted endpoint shares one pid, and `...\n*72│- **Invariant:** for **daemon-hosted** perches (Psyche; spt-hosted Self), liveness is the **daemon's authoritative in-memory endpoint table + a `status` field** on `info.json` (`online|offline|…`), **never** `is_process_alive(info.pid)`. `info.pid` for a daemon-hosted perch is at most a *hosted-by-daemon* marker (the daemon pid), not a liveness signal; registry stale-clean for these rows keys on the daemon's endpoint table, not per-row pid. This reuses the pattern already specified for **Shells** (`inf...","linesTruncated":true,"meta":{"limits":{"columnTruncated":{"maxColumn":512}}}},"isError":false,"timestamp":1788999505119}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_QHcoNmqHT4fmSVZlIqOgQSWl|fc_0dde72effda0ad21016aa1f75c940487d292ef6b5d4c173d75","toolName":"read","startedAt":"2026-09-10T00:18:36.684Z","args":{"path":"artifact://65:301-659:raw"},"intent":"Reading remaining patch"},"id":"b36224f2","parentId":"ef338dcd","timestamp":"2026-09-10T00:18:36.684Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_d7prbfFiy18sBp67TKOC1IL5|fc_0dde72effda0ad21016aa1f75c941487d28e2f5f702b037d0d","toolName":"read","startedAt":"2026-09-10T00:18:36.684Z","args":{"path":".worktrees/286-churn/crates/spt-daemon/src/conn.rs:1-420"},"intent":"Reading connection lifecycle context"},"id":"c44219e3","parentId":"b36224f2","timestamp":"2026-09-10T00:18:36.684Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_CFcXvjshpb2DT4KOYKAz71ED|fc_0dde72effda0ad21016aa1f75c941887d28404af22f47d4c68","toolName":"read","startedAt":"2026-09-10T00:18:36.684Z","args":{"path":".worktrees/286-churn/crates/spt-daemon/src/conn.rs:421-795"},"intent":"Reading connection write context"},"id":"1e4fda44","parentId":"c44219e3","timestamp":"2026-09-10T00:18:36.684Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_yYGvYsXGB3t7hhjhu2hBMYjN|fc_0dde72effda0ad21016aa1f75c942087d2b1f7df67820bc67e","toolName":"grep","startedAt":"2026-09-10T00:18:36.684Z","args":{"path":".worktrees/286-churn/docs/KNOWN-HAZARDS.md:110-228"},"intent":"Locating invariant sections"},"id":"a65cafbc","parentId":"1e4fda44","timestamp":"2026-09-10T00:18:36.684Z"}
{"type":"message","id":"ce7b61f7","parentId":"a65cafbc","timestamp":"2026-09-10T00:18:36.684Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":"**Locating remaining diff context**"},{"type":"toolCall","id":"call_QHcoNmqHT4fmSVZlIqOgQSWl|fc_0dde72effda0ad21016aa1f75c940487d292ef6b5d4c173d75","name":"read","arguments":{"i":"Reading remaining patch","path":"artifact://65:301-659:raw"},"intent":"Reading remaining patch"},{"type":"toolCall","id":"call_d7prbfFiy18sBp67TKOC1IL5|fc_0dde72effda0ad21016aa1f75c941487d28e2f5f702b037d0d","name":"read","arguments":{"i":"Reading connection lifecycle context","path":".worktrees/286-churn/crates/spt-daemon/src/conn.rs:1-420"},"intent":"Reading connection lifecycle context"},{"type":"toolCall","id":"call_CFcXvjshpb2DT4KOYKAz71ED|fc_0dde72effda0ad21016aa1f75c941887d28404af22f47d4c68","name":"read","arguments":{"i":"Reading connection write context","path":".worktrees/286-churn/crates/spt-daemon/src/conn.rs:421-795"},"intent":"Reading connection write context"},{"type":"toolCall","id":"call_yYGvYsXGB3t7hhjhu2hBMYjN|fc_0dde72effda0ad21016aa1f75c942087d2b1f7df67820bc67e","name":"grep","arguments":{"i":"Locating invariant sections","pattern":"^###|poison|CONN","path":".worktrees/286-churn/docs/KNOWN-HAZARDS.md:110-228","case":false,"gitignore":false,"skip":null},"intent":"Locating invariant sections"},{"type":"toolCall","id":"call_uDXE7hGZPSO1wpkbLungmWy9|fc_0dde72effda0ad21016aa1f75c942487d2bfbe9feccf7adcba","name":"grep","arguments":{"i":"Locating related architecture","pattern":"broker|brain|lifecycle|IPC|connection|test","path":".worktrees/286-churn/PRD.md;.worktrees/286-churn/ROADMAP.md;.worktrees/286-churn/CONTEXT.md;.worktrees/286-churn/docs/STORAGE.md;.worktrees/286-churn/docs/MANIFEST.md;.worktrees/286-churn/docs/CONTEXT-MEMORY.md","case":false,"gitignore":false,"skip":null},"intent":"Locating related architecture"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":9860,"output":334,"cacheRead":31872,"cacheWrite":0,"totalTokens":42066,"reasoningTokens":23,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1788999505135,"responseId":"resp_0dde72effda0ad21016aa1f751624887d29ddc59e265fc7ce2","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_0dde72effda0ad21016aa1f752900487d2a50fbba8fdaa0c85","type":"reasoning","content":[],"encrypted_content":"gAAAAABqofdcCD7PdKqrxdySjcQfBA-i1Gpkj33GfA0C2Mv-Ld-jKbQVs1Yoz_LTCqRYl5bemS1ayTCOs8-NK7T8l2QGdcJ2kwlWNC7pXuGXIeqePO5UKLhNLgRUU3QMfLgemuH9N1FB299xp0S8ypDFWeemxVxAO23i7ex771_v5rzwEeHFdCfNkxJAepJnuWyVH5K66DMDXjWxiDA93E1dfdPRLxyz1NaZMlmsUlWOvGZ2hmo8jcYsaqFofNqlU27Ivqa1l0wNayyJDWbD-y8ecJT20xdwLJKEHFJtzvCNhxxYksPihnea8zoDWP6VEx6uOn7uwIHzN_iOaC06NcfimipjgCIZp988pHQJjhyK7Y6weTuE6F3vNKb7CnzxuUHLkIBumR_NXA8hTyhmx--PpmlEB7I7T0-PgFufdvQF98dyqj3-VUmS6my-HcNRApRPetyqEM80CrgMXrC4536TTkp69LYYLV6s63_CF6b8BaVuCQpNPipRTE3ef4bRRsuI76Z-VWo_oB8ySyO0ulkO1A0cC3wIoOpj8BaAxbeVRzXKVa2HAnWn8_GVJZfvjRWkpR8BxSgY5YMIQEJqYcolHZ0i8uxBlS2B8kN4RbkrMdPlA-7xv6IBZSvcos-RDwuC9QguqoC-EodRUA9b3asiAPqHYL-iJ3v90eFwBxPx5__fvnaUTikhQhCqtT2Kc7FMIfocBMc42RRLkf6PLe-dDf-gnHje6udLRzyFJwgLxZgQBBWs_X9pZdaGmukffQjDMTTyK8jUTpNnKc_rMS3h8qQPa3Pd8MklktbRAbN_LVfhbs-bltnYk0qR9QP2l3AIVKOsBf2BOBz6omxf9OSuRVG9_YlQzGYohww1dEyt0qQSctL7oBuQDMThG3vPphB8nwpDMVowARmqG_UwJi9_W14uhfq_3T9s7iQgA7Qd82UOc1M3xkGjJDlDnUCLrmCFV8Bvz6BlLgVuuWKAbN3O4VVGRrkEKZyjF3Hj2IyKYlghjpPiReUBPDQ7cS175N2Qi1KX7Jcr6hjylaHSNq2t1VHynKPQAJkvp5SWQqhdeK_9WkoYDpAPrK1HNyeixwNqj0m3X8LlAOJUf9eRZl7OuMZnHwvUfj2F3IFZ3I4ZolP1KPyEDpznrCzs9Sc-kWnZ1-mJ7xoWQO6nOZRuskdmo7YzmFga9_viT_7ady7mpGQIe6rMUAyrEoMkl8dGPaHVc3Vb3xbUeo-dRcJc_AfkPWk47zF3F2Nbb0RaVR4ZRVGcQnjYwzFTI-zmI5Cw7KoIcMZvdeFMbDmoWNdM9R39w2VA4KlGsm-TcMarA1ie8th9uqh1BG1OeU4GbdYtRkuTslVdgFPCRGMsOSwaOUCVRltKbBm9VJIsJDLYgA3G47Dj29W9WNHYFzhH3P46jEJpQqE39sY3","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"summary":[{"type":"summary_text","text":"**Locating remaining diff context**"}],"metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f75c940487d292ef6b5d4c173d75","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading remaining patch\",\"path\":\"artifact://65:301-659:raw\"}","call_id":"call_QHcoNmqHT4fmSVZlIqOgQSWl","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"read","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f75c941487d28e2f5f702b037d0d","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading connection lifecycle context\",\"path\":\".worktrees/286-churn/crates/spt-daemon/src/conn.rs:1-420\"}","call_id":"call_d7prbfFiy18sBp67TKOC1IL5","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"read","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f75c941887d28404af22f47d4c68","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading connection write context\",\"path\":\".worktrees/286-churn/crates/spt-daemon/src/conn.rs:421-795\"}","call_id":"call_CFcXvjshpb2DT4KOYKAz71ED","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"read","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f75c942087d2b1f7df67820bc67e","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating invariant sections\",\"pattern\":\"^###|poison|CONN\",\"path\":\".worktrees/286-churn/docs/KNOWN-HAZARDS.md:110-228\",\"case\":false,\"gitignore\":false,\"skip\":null}","call_id":"call_yYGvYsXGB3t7hhjhu2hBMYjN","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"grep","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f75c942487d2bfbe9feccf7adcba","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating related architecture\",\"pattern\":\"broker|brain|lifecycle|IPC|connection|test\",\"path\":\".worktrees/286-churn/PRD.md;.worktrees/286-churn/ROADMAP.md;.worktrees/286-churn/CONTEXT.md;.worktrees/286-churn/docs/STORAGE.md;.worktrees/286-churn/docs/MANIFEST.md;.worktrees/286-churn/docs/CONTEXT-MEMORY.md\",\"case\":false,\"gitignore\":false,\"skip\":null}","call_id":"call_uDXE7hGZPSO1wpkbLungmWy9","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"grep","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}}]},"duration":11519.39419999998,"ttft":1454.7605000000913,"completedAt":1788999516681,"contextSnapshot":{"promptTokens":41732,"nonMessageTokens":13648,"compactionEpoch":0}}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_uDXE7hGZPSO1wpkbLungmWy9|fc_0dde72effda0ad21016aa1f75c942487d2bfbe9feccf7adcba","toolName":"grep","startedAt":"2026-09-10T00:18:36.685Z","args":{"path":".worktrees/286-churn/PRD.md;.worktrees/286-churn/ROADMAP.md;.worktrees/286-churn/CONTEXT.md;.worktrees/286-churn/docs/STORAGE.md;.worktrees/286-churn/docs/MANIFEST.md;.worktrees/286-churn/docs/CONTEXT…"},"intent":"Locating related architecture"},"id":"2ee03d35","parentId":"ce7b61f7","timestamp":"2026-09-10T00:18:36.685Z"}
{"type":"message","id":"ad21c169","parentId":"2ee03d35","timestamp":"2026-09-10T00:18:36.695Z","message":{"role":"toolResult","toolCallId":"call_d7prbfFiy18sBp67TKOC1IL5|fc_0dde72effda0ad21016aa1f75c941487d28e2f5f702b037d0d","toolName":"read","content":[{"type":"text","text":"//! Bounded, cancelable, poison-on-failure framed writes for one physical\n//! broker connection (UPDATE-WEDGE round-4).\n//!\n//! ## The failure class this kills\n//!\n//! A broker writer used to hold the connection's `Mutex<SendHalf>` guard\n//! ACROSS a blocking [`write_frame`]. On Windows, `interprocess 2.4.2` routes\n//! that send to `WriteFileEx` + `SleepEx(INFINITE, alertable)` — a local-socket\n//! path with **no supported write timeout** (`set_timeout` → `Unsupported`).\n//! A connected consumer that stops reading (an `rc --take` whose process\n//! wedged; a suspended brain) therefore blocks the write INDEFINITELY, and the\n//! logical stall-evict (`stall_evict_controller`) releases only the controller\n//! ROLE — it never cancels the in-flight pipe write or closes the physical\n//! connection, so the writer (and every write queued behind the same gate)\n//! stays wedged until the peer process dies (~128 s in the field capture, on a\n//! brain restart).\n//!\n//! ## The primitive\n//!\n//! [`BrokerConn`] owns the send half outright — it never leaves this module,\n//! so an unbounded bypass write cannot compile. Every write:\n//!\n//! 1. acquires the serialized write GATE with a bounded wait (an ABSOLUTE\n//!    deadline stamped at write entry covers BOTH the gate-wait and the OS\n//!    write);\n//! 2. arms an INDEPENDENT per-connection watchdog thread with that absolute\n//!    deadline, then performs the blocking [`write_frame`] OUT of every lock;\n//! 3. on success, returns the half to the gate.\n//!\n//! The watchdog fires OUT OF BAND at the deadline — it never relies on the\n//! write returning, nor on opportunistic stall-evict sampling: it poisons the\n//! connection and aborts the in-flight I/O (`CancelIoEx` on Windows — the\n//! completion APC makes the wedged `SleepEx` loop return with\n//! `ERROR_OPERATION_ABORTED`; `shutdown(SHUT_RDWR)` on Unix — the blocked\n//! write returns `EPIPE`). A gate-waiter that exhausts its own deadline fires\n//! the same abort (the stuck op it waited behind is the wedge).\n//!\n//! ## Poison / retire invariant (both OSes)\n//!\n//! On deadline, write failure, or cancellation: (1) the WHOLE physical\n//! connection is poisoned, idempotently — a length-prefixed frame may have\n//! been PARTIALLY written, so reuse would corrupt framing; (2) read and write\n//! sides are aborted so `handle_conn` reaches its existing EOF cleanup;\n//! (3) the buffer of the canceled op is released only after the op reports\n//! completion (`write_frame` returns only after the completion APC — the\n//! synchronous return IS the completion handshake); (4) the retired half is\n//! parked in a graveyard slot — never dropped while the watchdog might still\n//! touch the raw handle (drop = `CloseHandle`, and a recycled handle value\n//! must never see a late `CancelIoEx`), and never re-gated; (5) on Windows the\n//! pipe is then severed (`DisconnectNamedPipe` — the server end forces the\n//! client off, so the old client's next access fails, a re-issued server read\n//! fails, and the `interprocess` drop-limbo flush cannot block on unread\n//! data), followed by one more `CancelIoEx` covering a read re-issued in the\n//! cancel→disconnect window.\n//!\n//! No new output queue (the bounded per-sink queues + isolated writer threads\n//! already exist — the block this bounds is BELOW them), and no `PIPE_NOWAIT`\n//! (recorded mid-frame corruption risk). Unix keeps its existing semantics for\n//! healthy connections; only the poison/retire path is new.\n// [impl->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\n\nuse std::io;\nuse std::sync::atomic::{AtomicBool, AtomicU64, Ordering};\nuse std::sync::{Arc, Condvar, LazyLock, Mutex, MutexGuard, OnceLock};\nuse std::thread::{self, JoinHandle};\nuse std::time::{Duration, Instant, SystemTime, UNIX_EPOCH};\n\nuse interprocess::local_socket::SendHalf;\n\nuse crate::codec::write_frame;\nuse crate::frame::Envelope;\n\n/// Stable per-physical-conn identity mint — a monotonic u64 stamped at\n/// [`BrokerConn::new`]. Before this the only conn identity was `Arc::ptr_eq`,\n/// which does not survive a log line: per-line CONN_WRITE_POISONED churn was\n/// UNDECIDABLE between fresh-carrier churn and stderr interleave (the F-039\n/// RCA's terminal fork). Every lifecycle record carries `conn=<id>`.\n// [impl->REQ-CONN-POISON-ATTRIBUTION]\nstatic CONN_ID_SEQ: AtomicU64 = AtomicU64::new(0);\n\n/// The shared monotonic anchor for `mono_ms` stamps. Set explicitly at daemon\n/// startup ([`init_log_anchor`], both broker and brain entries) so every\n/// record in one process shares one base; lazily initialized on first stamp\n/// otherwise (tests, in-process brokers) — correlation needs a COMMON base,\n/// not an absolute one.\nstatic MONO_ANCHOR: OnceLock<Instant> = OnceLock::new();\n\n/// Anchor the monotonic half of [`log_stamp`] at process start — called first\n/// thing by both daemon entries, next to `scrub_identity_env`.\n// [impl->REQ-CONN-POISON-ATTRIBUTION]\npub(crate) fn init_log_anchor() {\n    let _ = MONO_ANCHOR.set(Instant::now());\n}\n\n/// Wall-clock + monotonic stamp for daemon-stderr correlation records\n/// (`wall_ms` = Unix epoch millis; `mono_ms` = millis since [`init_log_anchor`]).\n/// stderr has no per-line timestamps and broker+brain share one file — without\n/// BOTH clocks, interleaved records cannot be ordered (monotonic) or matched\n/// to field incident times (wall).\n// [impl->REQ-CONN-POISON-ATTRIBUTION]\npub(crate) fn log_stamp() -> String {\n    let stamp = LogTime::now();\n    format!(\"wall_ms={} mono_ms={}\", stamp.wall_ms, stamp.mono_ms)\n}\n\n/// Retain the original write instant even when its healthy line is aggregated.\n#[derive(Clone, Copy)]\nstruct LogTime {\n    wall_ms: u128,\n    mono_ms: u128,\n}\n\nimpl LogTime {\n    fn now() -> Self {\n        Self {\n            wall_ms: SystemTime::now()\n                .duration_since(UNIX_EPOCH)\n                .map(|d| d.as_millis())\n                .unwrap_or(0),\n            mono_ms: MONO_ANCHOR.get_or_init(Instant::now).elapsed().as_millis(),\n        }\n    }\n}\n\nconst HEALTHY_LOG_INTERVAL: Duration = Duration::from_secs(60);\n\nstruct HealthyWindow {\n    since: Instant,\n    starts: u64,\n    closes: u64,\n}\n\n/// One process-wide counter window, not one limiter per connection. No thread,\n/// queue or per-conn healthy string: the next event after 60s emits the totals.\n/// Exceptional and role-transition records bypass this window entirely.\n// [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\nstruct LifecycleLog {\n    healthy: Mutex<HealthyWindow>,\n    #[cfg(test)]\n    capture: Option<Mutex<Vec<String>>>,\n    #[cfg(test)]\n    clock: Option<Mutex<Instant>>,\n}\n\nimpl LifecycleLog {\n    fn new() -> Self {\n        Self {\n            healthy: Mutex::new(HealthyWindow {\n                since: Instant::now(),\n                starts: 0,\n                closes: 0,\n            }),\n            #[cfg(test)]\n            capture: None,\n            #[cfg(test)]\n            clock: None,\n        }\n    }\n\n    #[cfg(test)]\n    fn captured(now: Instant) -> Self {\n        let mut log = Self::new();\n        log.healthy.get_mut().unwrap().since = now;\n        log.capture = Some(Mutex::new(Vec::new()));\n        log.clock = Some(Mutex::new(now));\n        log\n    }\n\n    fn emit(&self, args: std::fmt::Arguments<'_>) {\n        #[cfg(test)]\n        if let Some(lines) = &self.capture {\n            recover(lines).push(args.to_string());\n            return;\n        }\n        spt_proto::emit_line_err!(\"{args}\");\n    }\n\n    fn healthy(&self, start: bool) {\n        let mut window = recover(&self.healthy);\n        // Read time under the counter lock so concurrent callers cannot move\n        // the window backwards after waiting for a newer caller to flush.\n        let now = Instant::now();\n        #[cfg(test)]\n        let now = self.clock.as_ref().map(|c| *recover(c)).unwrap_or(now);\n        if start {\n            window.starts = window.starts.saturating_add(1);\n        } else {\n            window.closes = window.closes.saturating_add(1);\n        }\n        let elapsed = now.saturating_duration_since(window.since);\n        if elapsed < HEALTHY_LOG_INTERVAL {\n            return;\n        }\n        let (starts, closes) = (window.starts, window.closes);\n        *window = HealthyWindow { since: now, starts: 0, closes: 0 };\n        drop(window);\n        self.emit(format_args!(\n            \"CONN_LIFECYCLE_SUMMARY: scope=broker-ipc {} interval_ms={} starts={starts} closes={closes} \\\n             [REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\",\n            log_stamp(), elapsed.as_millis()\n        ));\n    }\n}\n\n/// Attribution-label cap: facts accrete per role attach (a long-lived brain\n/// conn can be controller + viewer + presence subscriber over its life), and\n/// the label must stay BOUNDED — one truncation marker, then no further growth.\nconst LABEL_CAP: usize = 240;\n\n/// One bounded write's outcome (beyond `Err` = poisoned/timed-out/failed).\n#[derive(Debug, PartialEq, Eq)]\npub(crate) enum ConnWrite {\n    /// The frame was fully written.\n    Done,\n    /// The pre-write gate closure declined (e.g. a superseded controller\n    /// writer's epoch check) — nothing was written; the conn stays healthy.\n    Superseded,\n}\n\n/// The serialized write gate. `half: Some` = free; `None` = a write is in\n/// flight (checked out) or the conn has retired (then `poisoned` is set).\nstruct Gate {\n    half: Option<SendHalf>,\n}\n\n/// Watchdog bookkeeping: the currently in-flight op and its ABSOLUTE deadline.\nstruct Dog {\n    /// `Some((op_id, deadline))` while a write is in flight.\n    inflight: Option<(u64, Instant)>,\n    /// Raised by [`BrokerConn::drop`]; the watchdog exits.\n    shutdown: bool,\n}\n\n/// The state shared with the watchdog thread (no `Arc<BrokerConn>` cycle —\n/// the watchdog must not keep the conn alive).\nstruct Inner {\n    /// Stable per-physical-conn id ([`CONN_ID_SEQ`]) — on every lifecycle record.\n    id: u64,\n    /// Accreted attribution facts (subscriber role, endpoint/session where\n    /// known — [`BrokerConn::describe`]). Bounded by [`LABEL_CAP`].\n    label: Mutex<String>,\n    /// Once-per-conn original first committed write time; retained on failures.\n    first_write: OnceLock<LogTime>,\n    lifecycle_log: Arc<LifecycleLog>,\n    gate: Mutex<Gate>,\n    gate_cv: Condvar,\n    dog: Mutex<Dog>,\n    dog_cv: Condvar,\n    /// Idempotently latched on the first deadline/failure/cancel. A poisoned\n    /// conn refuses every future write and is never re-gated.\n    poisoned: AtomicBool,\n    /// Latched when the retirement was DEADLINE-driven (watchdog fired, or a\n    /// gate-waiter exhausted its bound) — as opposed to an organic I/O error\n    /// (peer death). Callers see it as `io::ErrorKind::TimedOut` and can\n    /// count the wedge (the W2 stall-evict tally) without conflating it with\n    /// an ordinary disconnect.\n    timed_out: AtomicBool,\n    /// The physical sever ran (exactly once; Windows `DisconnectNamedPipe`).\n    severed: AtomicBool,\n    /// A retired (poisoned) half parks here so its OS handle stays alive\n    /// until [`BrokerConn::drop`] — after the watchdog is joined, so no raw\n    /// `CancelIoEx`/`DisconnectNamedPipe` can race the `CloseHandle`.\n    graveyard: Mutex<Option<SendHalf>>,\n    /// Monotonic per-write op identity (the watchdog cancels a SPECIFIC op).\n    op_seq: AtomicU64,\n    /// The absolute per-write bound (gate-wait + OS write completion).\n    bound: Duration,\n    /// The raw OS handle/fd of the underlying stream, captured at\n    /// construction for the out-of-band abort. The recv and send halves share\n    /// ONE underlying stream object (`interprocess` splits by ref-clone), so\n    /// aborting this handle covers a parked read too. Valid for the whole\n    /// conn lifetime: the half lives in `gate`, a writer's stack, or\n    /// `graveyard` until drop.\n    #[cfg(windows)]\n    raw: isize,\n    #[cfg(unix)]\n    raw: i32,\n}\n\n/// Recover a mutex guard even if a holder panicked — the state these guards\n/// protect is a plain slot/flag pair, valid on every panic boundary.\nfn recover<T>(m: &Mutex<T>) -> MutexGuard<'_, T> {\n    match m.lock() {\n        Ok(g) => g,\n        Err(p) => p.into_inner(),\n    }\n}\n\n#[cfg(windows)]\nfn raw_of(half: &SendHalf) -> isize {\n    use std::os::windows::io::{AsHandle, AsRawHandle};\n    match half {\n        SendHalf::NamedPipe(h) => h.as_handle().as_raw_handle() as isize,\n    }\n}\n\n#[cfg(unix)]\nfn raw_of(half: &SendHalf) -> i32 {\n    use std::os::fd::{AsFd, AsRawFd};\n    match half {\n        SendHalf::UdSocket(h) => h.as_fd().as_raw_fd(),\n    }\n}\n\n#[cfg(windows)]\nmod win {\n    // kernel32 is in the MSVC default link set — raw externs, no new dep\n    // (the deelevate.rs / daemon.rs posture).\n    extern \"system\" {\n        pub fn CancelIoEx(handle: isize, overlapped: *mut core::ffi::c_void) -> i32;\n        pub fn DisconnectNamedPipe(handle: isize) -> i32;\n    }\n}\n\nimpl Inner {\n    /// The attribution snapshot every lifecycle record carries: stable conn id,\n    /// accreted role/endpoint/session facts, wall + monotonic stamps.\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\n    fn attribution(&self) -> String {\n        let label = recover(&self.label);\n        let facts: &str = if label.is_empty() { \"role=unattributed\" } else { &label };\n        match self.first_write.get() {\n            Some(start) => format!(\n                \"conn={} {} {} first_write_wall_ms={} first_write_mono_ms={}\",\n                self.id, facts, log_stamp(), start.wall_ms, start.mono_ms\n            ),\n            None => format!(\"conn={} {} {} first_write=none\", self.id, facts, log_stamp()),\n        }\n    }\n\n    /// Render the once-per-conn retirement record — the F-039 token split\n    /// (leg a): `CONN_WRITE_POISONED` is RESERVED for the deadline class\n    /// (`timed_out` latched — the wedge observable blackhole-controller.ps1\n    /// watches); an ORGANIC write failure (peer death / broken pipe on an\n    /// already-dead counterpart, the offline-peer dial churn) emits the\n    /// distinct `CONN_WRITE_RETIRED` token carrying the error kind and the\n    /// SAME attribution fields — quieter class, never thinner (the retired\n    /// class is what discriminates fresh-carrier churn from interleave\n    /// artifact in the field). Behavior is identical either way: poison,\n    /// abort, retire, never reuse.\n    // [impl->REQ-CONN-POISON-DIAL-SCOPE]\n    fn render_retirement(&self, timed_out: bool, cause: Option<&io::Error>) -> String {\n        if timed_out {\n            format!(\n                \"CONN_WRITE_POISONED: {} — physical broker conn retired: a write \\\n                 exceeded its {}ms bound; in-flight I/O aborted, conn never reused \\\n                 [REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\",\n                self.attribution(),\n                self.bound.as_millis()\n            )\n        } else {\n            let cause = match cause {\n                Some(e) => format!(\"{:?}: {e}\", e.kind()),\n                None => \"write failed\".to_string(),\n            };\n            format!(\n                \"CONN_WRITE_RETIRED: {} — organic write failure ({cause}); physical \\\n                 broker conn poisoned + retired, never reused (same retire behavior \\\n                 as POISONED — only the failure class differs) \\\n                 [REQ-CONN-POISON-DIAL-SCOPE]\",\n                self.attribution()\n            )\n        }\n    }\n\n    /// Per-conn exceptional and role-transition records remain unconditional.\n    /// releases#286 supersedes only healthy start/close emission, which uses\n    /// the process-wide interval counter instead. Retirement retains the\n    /// original first-write time even when the start line was aggregated.\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\n    fn lifecycle(&self, event: &str, extra: &str) {\n        let sep = if extra.is_empty() { \"\" } else { \" \" };\n        self.lifecycle_log.emit(format_args!(\n            \"CONN_LIFECYCLE: {} event={event}{sep}{extra} [REQ-CONN-POISON-ATTRIBUTION]\",\n            self.attribution()\n        ));\n    }\n\n    /// Abort the connection's in-flight I/O OUT OF BAND (both directions —\n    /// the halves share one OS object) and latch the poison. Idempotent and\n    /// safe to repeat: `CancelIoEx` on a handle with nothing pending is a\n    /// no-op, and re-cancel closes the race where a write was issued a beat\n    /// AFTER a cancel (cancellation only reaches ALREADY-issued ops).\n    /// `cause` (the organic caller's write error) rides the retirement record\n    /// when the retirement is NOT deadline-driven.\n    fn poison_and_cancel(&self, cause: Option<&io::Error>) {\n        if !self.poisoned.swap(true, Ordering::AcqRel) {\n            // First poison of this conn: loud, once. The token is class-split\n            // (F-039 leg a): deadline → CONN_WRITE_POISONED (the wedge\n            // observable), organic fast-fail → CONN_WRITE_RETIRED.\n            self.lifecycle_log.emit(format_args!(\n                \"{}\",\n                self.render_retirement(self.timed_out.load(Ordering::Acquire), cause)\n            ));\n        }\n        #[cfg(windows)]\n        // SAFETY: scalar kernel32 call on a handle whose owning half is kept\n        // alive until Drop (gate / writer stack / graveyard — see `raw`).\n        unsafe {\n            let _ = win::CancelIoEx(self.raw, core::ptr::null_mut());\n        }\n        #[cfg(unix)]\n        // SAFETY: scalar shutdown(2) on an fd alive until Drop. SHUT_RDWR\n        // fails the blocked write (EPIPE), EOFs the parked read, and fails\n        // every future op — on Unix this one call IS the whole teardown.\n        unsafe {\n            let _ = libc::shutdown(self.raw, libc::SHUT_RDWR);\n        }\n        // Wake gate-waiters so they observe the poison instead of sleeping\n        // out their full deadline, and the watchdog so it re-evaluates.\n        self.gate_cv.notify_all();\n        self.dog_cv.notify_all();\n    }\n\n    /// Sever the physical connection (exactly once), AFTER the canceled op\n    /// reported completion: on Windows `DisconnectNamedPipe` forces the\n    /// client off (its next access fails), makes a re-issued server-side\n    /// read fail (so `handle_conn` reaches EOF cleanup even if it slipped a\n    /// new read in after the cancel), and defuses the drop-limbo flush; the\n    /// trailing re-cancel covers exactly that slipped-in read. On Unix the\n    /// `shutdown` in [`poison_and_cancel`] already did all of this.\n    // On Unix the cfg(windows) teardown strips out and the idempotence\n    // guard's early `return` becomes the function tail — needless_return\n…\n}\n\n[Showing lines 1-423 of 1000. Use :424 to continue]"}],"details":{"truncation":{"content":"//! Bounded, cancelable, poison-on-failure framed writes for one physical\r\n//! broker connection (UPDATE-WEDGE round-4).\r\n//!\r\n//! ## The failure class this kills\r\n//!\r\n//! A broker writer used to hold the connection's `Mutex<SendHalf>` guard\r\n//! ACROSS a blocking [`write_frame`]. On Windows, `interprocess 2.4.2` routes\r\n//! that send to `WriteFileEx` + `SleepEx(INFINITE, alertable)` — a local-socket\r\n//! path with **no supported write timeout** (`set_timeout` → `Unsupported`).\r\n//! A connected consumer that stops reading (an `rc --take` whose process\r\n//! wedged; a suspended brain) therefore blocks the write INDEFINITELY, and the\r\n//! logical stall-evict (`stall_evict_controller`) releases only the controller\r\n//! ROLE — it never cancels the in-flight pipe write or closes the physical\r\n//! connection, so the writer (and every write queued behind the same gate)\r\n//! stays wedged until the peer process dies (~128 s in the field capture, on a\r\n//! brain restart).\r\n//!\r\n//! ## The primitive\r\n//!\r\n//! [`BrokerConn`] owns the send half outright — it never leaves this module,\r\n//! so an unbounded bypass write cannot compile. Every write:\r\n//!\r\n//! 1. acquires the serialized write GATE with a bounded wait (an ABSOLUTE\r\n//!    deadline stamped at write entry covers BOTH the gate-wait and the OS\r\n//!    write);\r\n//! 2. arms an INDEPENDENT per-connection watchdog thread with that absolute\r\n//!    deadline, then performs the blocking [`write_frame`] OUT of every lock;\r\n//! 3. on success, returns the half to the gate.\r\n//!\r\n//! The watchdog fires OUT OF BAND at the deadline — it never relies on the\r\n//! write returning, nor on opportunistic stall-evict sampling: it poisons the\r\n//! connection and aborts the in-flight I/O (`CancelIoEx` on Windows — the\r\n//! completion APC makes the wedged `SleepEx` loop return with\r\n//! `ERROR_OPERATION_ABORTED`; `shutdown(SHUT_RDWR)` on Unix — the blocked\r\n//! write returns `EPIPE`). A gate-waiter that exhausts its own deadline fires\r\n//! the same abort (the stuck op it waited behind is the wedge).\r\n//!\r\n//! ## Poison / retire invariant (both OSes)\r\n//!\r\n//! On deadline, write failure, or cancellation: (1) the WHOLE physical\r\n//! connection is poisoned, idempotently — a length-prefixed frame may have\r\n//! been PARTIALLY written, so reuse would corrupt framing; (2) read and write\r\n//! sides are aborted so `handle_conn` reaches its existing EOF cleanup;\r\n//! (3) the buffer of the canceled op is released only after the op reports\r\n//! completion (`write_frame` returns only after the completion APC — the\r\n//! synchronous return IS the completion handshake); (4) the retired half is\r\n//! parked in a graveyard slot — never dropped while the watchdog might still\r\n//! touch the raw handle (drop = `CloseHandle`, and a recycled handle value\r\n//! must never see a late `CancelIoEx`), and never re-gated; (5) on Windows the\r\n//! pipe is then severed (`DisconnectNamedPipe` — the server end forces the\r\n//! client off, so the old client's next access fails, a re-issued server read\r\n//! fails, and the `interprocess` drop-limbo flush cannot block on unread\r\n//! data), followed by one more `CancelIoEx` covering a read re-issued in the\r\n//! cancel→disconnect window.\r\n//!\r\n//! No new output queue (the bounded per-sink queues + isolated writer threads\r\n//! already exist — the block this bounds is BELOW them), and no `PIPE_NOWAIT`\r\n//! (recorded mid-frame corruption risk). Unix keeps its existing semantics for\r\n//! healthy connections; only the poison/retire path is new.\r\n// [impl->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\r\n\r\nuse std::io;\r\nuse std::sync::atomic::{AtomicBool, AtomicU64, Ordering};\r\nuse std::sync::{Arc, Condvar, LazyLock, Mutex, MutexGuard, OnceLock};\r\nuse std::thread::{self, JoinHandle};\r\nuse std::time::{Duration, Instant, SystemTime, UNIX_EPOCH};\r\n\r\nuse interprocess::local_socket::SendHalf;\r\n\r\nuse crate::codec::write_frame;\r\nuse crate::frame::Envelope;\r\n\r\n/// Stable per-physical-conn identity mint — a monotonic u64 stamped at\r\n/// [`BrokerConn::new`]. Before this the only conn identity was `Arc::ptr_eq`,\r\n/// which does not survive a log line: per-line CONN_WRITE_POISONED churn was\r\n/// UNDECIDABLE between fresh-carrier churn and stderr interleave (the F-039\r\n/// RCA's terminal fork). Every lifecycle record carries `conn=<id>`.\r\n// [impl->REQ-CONN-POISON-ATTRIBUTION]\r\nstatic CONN_ID_SEQ: AtomicU64 = AtomicU64::new(0);\r\n\r\n/// The shared monotonic anchor for `mono_ms` stamps. Set explicitly at daemon\r\n/// startup ([`init_log_anchor`], both broker and brain entries) so every\r\n/// record in one process shares one base; lazily initialized on first stamp\r\n/// otherwise (tests, in-process brokers) — correlation needs a COMMON base,\r\n/// not an absolute one.\r\nstatic MONO_ANCHOR: OnceLock<Instant> = OnceLock::new();\r\n\r\n/// Anchor the monotonic half of [`log_stamp`] at process start — called first\r\n/// thing by both daemon entries, next to `scrub_identity_env`.\r\n// [impl->REQ-CONN-POISON-ATTRIBUTION]\r\npub(crate) fn init_log_anchor() {\r\n    let _ = MONO_ANCHOR.set(Instant::now());\r\n}\r\n\r\n/// Wall-clock + monotonic stamp for daemon-stderr correlation records\r\n/// (`wall_ms` = Unix epoch millis; `mono_ms` = millis since [`init_log_anchor`]).\r\n/// stderr has no per-line timestamps and broker+brain share one file — without\r\n/// BOTH clocks, interleaved records cannot be ordered (monotonic) or matched\r\n/// to field incident times (wall).\r\n// [impl->REQ-CONN-POISON-ATTRIBUTION]\r\npub(crate) fn log_stamp() -> String {\r\n    let stamp = LogTime::now();\r\n    format!(\"wall_ms={} mono_ms={}\", stamp.wall_ms, stamp.mono_ms)\r\n}\r\n\r\n/// Retain the original write instant even when its healthy line is aggregated.\r\n#[derive(Clone, Copy)]\r\nstruct LogTime {\r\n    wall_ms: u128,\r\n    mono_ms: u128,\r\n}\r\n\r\nimpl LogTime {\r\n    fn now() -> Self {\r\n        Self {\r\n            wall_ms: SystemTime::now()\r\n                .duration_since(UNIX_EPOCH)\r\n                .map(|d| d.as_millis())\r\n                .unwrap_or(0),\r\n            mono_ms: MONO_ANCHOR.get_or_init(Instant::now).elapsed().as_millis(),\r\n        }\r\n    }\r\n}\r\n\r\nconst HEALTHY_LOG_INTERVAL: Duration = Duration::from_secs(60);\r\n\r\nstruct HealthyWindow {\r\n    since: Instant,\r\n    starts: u64,\r\n    closes: u64,\r\n}\r\n\r\n/// One process-wide counter window, not one limiter per connection. No thread,\r\n/// queue or per-conn healthy string: the next event after 60s emits the totals.\r\n/// Exceptional and role-transition records bypass this window entirely.\r\n// [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\r\nstruct LifecycleLog {\r\n    healthy: Mutex<HealthyWindow>,\r\n    #[cfg(test)]\r\n    capture: Option<Mutex<Vec<String>>>,\r\n    #[cfg(test)]\r\n    clock: Option<Mutex<Instant>>,\r\n}\r\n\r\nimpl LifecycleLog {\r\n    fn new() -> Self {\r\n        Self {\r\n            healthy: Mutex::new(HealthyWindow {\r\n                since: Instant::now(),\r\n                starts: 0,\r\n                closes: 0,\r\n            }),\r\n            #[cfg(test)]\r\n            capture: None,\r\n            #[cfg(test)]\r\n            clock: None,\r\n        }\r\n    }\r\n\r\n    #[cfg(test)]\r\n    fn captured(now: Instant) -> Self {\r\n        let mut log = Self::new();\r\n        log.healthy.get_mut().unwrap().since = now;\r\n        log.capture = Some(Mutex::new(Vec::new()));\r\n        log.clock = Some(Mutex::new(now));\r\n        log\r\n    }\r\n\r\n    fn emit(&self, args: std::fmt::Arguments<'_>) {\r\n        #[cfg(test)]\r\n        if let Some(lines) = &self.capture {\r\n            recover(lines).push(args.to_string());\r\n            return;\r\n        }\r\n        spt_proto::emit_line_err!(\"{args}\");\r\n    }\r\n\r\n    fn healthy(&self, start: bool) {\r\n        let mut window = recover(&self.healthy);\r\n        // Read time under the counter lock so concurrent callers cannot move\r\n        // the window backwards after waiting for a newer caller to flush.\r\n        let now = Instant::now();\r\n        #[cfg(test)]\r\n        let now = self.clock.as_ref().map(|c| *recover(c)).unwrap_or(now);\r\n        if start {\r\n            window.starts = window.starts.saturating_add(1);\r\n        } else {\r\n            window.closes = window.closes.saturating_add(1);\r\n        }\r\n        let elapsed = now.saturating_duration_since(window.since);\r\n        if elapsed < HEALTHY_LOG_INTERVAL {\r\n            return;\r\n        }\r\n        let (starts, closes) = (window.starts, window.closes);\r\n        *window = HealthyWindow { since: now, starts: 0, closes: 0 };\r\n        drop(window);\r\n        self.emit(format_args!(\r\n            \"CONN_LIFECYCLE_SUMMARY: scope=broker-ipc {} interval_ms={} starts={starts} closes={closes} \\\r\n             [REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\",\r\n            log_stamp(), elapsed.as_millis()\r\n        ));\r\n    }\r\n}\r\n\r\n/// Attribution-label cap: facts accrete per role attach (a long-lived brain\r\n/// conn can be controller + viewer + presence subscriber over its life), and\r\n/// the label must stay BOUNDED — one truncation marker, then no further growth.\r\nconst LABEL_CAP: usize = 240;\r\n\r\n/// One bounded write's outcome (beyond `Err` = poisoned/timed-out/failed).\r\n#[derive(Debug, PartialEq, Eq)]\r\npub(crate) enum ConnWrite {\r\n    /// The frame was fully written.\r\n    Done,\r\n    /// The pre-write gate closure declined (e.g. a superseded controller\r\n    /// writer's epoch check) — nothing was written; the conn stays healthy.\r\n    Superseded,\r\n}\r\n\r\n/// The serialized write gate. `half: Some` = free; `None` = a write is in\r\n/// flight (checked out) or the conn has retired (then `poisoned` is set).\r\nstruct Gate {\r\n    half: Option<SendHalf>,\r\n}\r\n\r\n/// Watchdog bookkeeping: the currently in-flight op and its ABSOLUTE deadline.\r\nstruct Dog {\r\n    /// `Some((op_id, deadline))` while a write is in flight.\r\n    inflight: Option<(u64, Instant)>,\r\n    /// Raised by [`BrokerConn::drop`]; the watchdog exits.\r\n    shutdown: bool,\r\n}\r\n\r\n/// The state shared with the watchdog thread (no `Arc<BrokerConn>` cycle —\r\n/// the watchdog must not keep the conn alive).\r\nstruct Inner {\r\n    /// Stable per-physical-conn id ([`CONN_ID_SEQ`]) — on every lifecycle record.\r\n    id: u64,\r\n    /// Accreted attribution facts (subscriber role, endpoint/session where\r\n    /// known — [`BrokerConn::describe`]). Bounded by [`LABEL_CAP`].\r\n    label: Mutex<String>,\r\n    /// Once-per-conn original first committed write time; retained on failures.\r\n    first_write: OnceLock<LogTime>,\r\n    lifecycle_log: Arc<LifecycleLog>,\r\n    gate: Mutex<Gate>,\r\n    gate_cv: Condvar,\r\n    dog: Mutex<Dog>,\r\n    dog_cv: Condvar,\r\n    /// Idempotently latched on the first deadline/failure/cancel. A poisoned\r\n    /// conn refuses every future write and is never re-gated.\r\n    poisoned: AtomicBool,\r\n    /// Latched when the retirement was DEADLINE-driven (watchdog fired, or a\r\n    /// gate-waiter exhausted its bound) — as opposed to an organic I/O error\r\n    /// (peer death). Callers see it as `io::ErrorKind::TimedOut` and can\r\n    /// count the wedge (the W2 stall-evict tally) without conflating it with\r\n    /// an ordinary disconnect.\r\n    timed_out: AtomicBool,\r\n    /// The physical sever ran (exactly once; Windows `DisconnectNamedPipe`).\r\n    severed: AtomicBool,\r\n    /// A retired (poisoned) half parks here so its OS handle stays alive\r\n    /// until [`BrokerConn::drop`] — after the watchdog is joined, so no raw\r\n    /// `CancelIoEx`/`DisconnectNamedPipe` can race the `CloseHandle`.\r\n    graveyard: Mutex<Option<SendHalf>>,\r\n    /// Monotonic per-write op identity (the watchdog cancels a SPECIFIC op).\r\n    op_seq: AtomicU64,\r\n    /// The absolute per-write bound (gate-wait + OS write completion).\r\n    bound: Duration,\r\n    /// The raw OS handle/fd of the underlying stream, captured at\r\n    /// construction for the out-of-band abort. The recv and send halves share\r\n    /// ONE underlying stream object (`interprocess` splits by ref-clone), so\r\n    /// aborting this handle covers a parked read too. Valid for the whole\r\n    /// conn lifetime: the half lives in `gate`, a writer's stack, or\r\n    /// `graveyard` until drop.\r\n    #[cfg(windows)]\r\n    raw: isize,\r\n    #[cfg(unix)]\r\n    raw: i32,\r\n}\r\n\r\n/// Recover a mutex guard even if a holder panicked — the state these guards\r\n/// protect is a plain slot/flag pair, valid on every panic boundary.\r\nfn recover<T>(m: &Mutex<T>) -> MutexGuard<'_, T> {\r\n    match m.lock() {\r\n        Ok(g) => g,\r\n        Err(p) => p.into_inner(),\r\n    }\r\n}\r\n\r\n#[cfg(windows)]\r\nfn raw_of(half: &SendHalf) -> isize {\r\n    use std::os::windows::io::{AsHandle, AsRawHandle};\r\n    match half {\r\n        SendHalf::NamedPipe(h) => h.as_handle().as_raw_handle() as isize,\r\n    }\r\n}\r\n\r\n#[cfg(unix)]\r\nfn raw_of(half: &SendHalf) -> i32 {\r\n    use std::os::fd::{AsFd, AsRawFd};\r\n    match half {\r\n        SendHalf::UdSocket(h) => h.as_fd().as_raw_fd(),\r\n    }\r\n}\r\n\r\n#[cfg(windows)]\r\nmod win {\r\n    // kernel32 is in the MSVC default link set — raw externs, no new dep\r\n    // (the deelevate.rs / daemon.rs posture).\r\n    extern \"system\" {\r\n        pub fn CancelIoEx(handle: isize, overlapped: *mut core::ffi::c_void) -> i32;\r\n        pub fn DisconnectNamedPipe(handle: isize) -> i32;\r\n    }\r\n}\r\n\r\nimpl Inner {\r\n    /// The attribution snapshot every lifecycle record carries: stable conn id,\r\n    /// accreted role/endpoint/session facts, wall + monotonic stamps.\r\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n    fn attribution(&self) -> String {\r\n        let label = recover(&self.label);\r\n        let facts: &str = if label.is_empty() { \"role=unattributed\" } else { &label };\r\n        match self.first_write.get() {\r\n            Some(start) => format!(\r\n                \"conn={} {} {} first_write_wall_ms={} first_write_mono_ms={}\",\r\n                self.id, facts, log_stamp(), start.wall_ms, start.mono_ms\r\n            ),\r\n            None => format!(\"conn={} {} {} first_write=none\", self.id, facts, log_stamp()),\r\n        }\r\n    }\r\n\r\n    /// Render the once-per-conn retirement record — the F-039 token split\r\n    /// (leg a): `CONN_WRITE_POISONED` is RESERVED for the deadline class\r\n    /// (`timed_out` latched — the wedge observable blackhole-controller.ps1\r\n    /// watches); an ORGANIC write failure (peer death / broken pipe on an\r\n    /// already-dead counterpart, the offline-peer dial churn) emits the\r\n    /// distinct `CONN_WRITE_RETIRED` token carrying the error kind and the\r\n    /// SAME attribution fields — quieter class, never thinner (the retired\r\n    /// class is what discriminates fresh-carrier churn from interleave\r\n    /// artifact in the field). Behavior is identical either way: poison,\r\n    /// abort, retire, never reuse.\r\n    // [impl->REQ-CONN-POISON-DIAL-SCOPE]\r\n    fn render_retirement(&self, timed_out: bool, cause: Option<&io::Error>) -> String {\r\n        if timed_out {\r\n            format!(\r\n                \"CONN_WRITE_POISONED: {} — physical broker conn retired: a write \\\r\n                 exceeded its {}ms bound; in-flight I/O aborted, conn never reused \\\r\n                 [REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\",\r\n                self.attribution(),\r\n                self.bound.as_millis()\r\n            )\r\n        } else {\r\n            let cause = match cause {\r\n                Some(e) => format!(\"{:?}: {e}\", e.kind()),\r\n                None => \"write failed\".to_string(),\r\n            };\r\n            format!(\r\n                \"CONN_WRITE_RETIRED: {} — organic write failure ({cause}); physical \\\r\n                 broker conn poisoned + retired, never reused (same retire behavior \\\r\n                 as POISONED — only the failure class differs) \\\r\n                 [REQ-CONN-POISON-DIAL-SCOPE]\",\r\n                self.attribution()\r\n            )\r\n        }\r\n    }\r\n\r\n    /// Per-conn exceptional and role-transition records remain unconditional.\r\n    /// releases#286 supersedes only healthy start/close emission, which uses\r\n    /// the process-wide interval counter instead. Retirement retains the\r\n    /// original first-write time even when the start line was aggregated.\r\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n    fn lifecycle(&self, event: &str, extra: &str) {\r\n        let sep = if extra.is_empty() { \"\" } else { \" \" };\r\n        self.lifecycle_log.emit(format_args!(\r\n            \"CONN_LIFECYCLE: {} event={event}{sep}{extra} [REQ-CONN-POISON-ATTRIBUTION]\",\r\n            self.attribution()\r\n        ));\r\n    }\r\n\r\n    /// Abort the connection's in-flight I/O OUT OF BAND (both directions —\r\n    /// the halves share one OS object) and latch the poison. Idempotent and\r\n    /// safe to repeat: `CancelIoEx` on a handle with nothing pending is a\r\n    /// no-op, and re-cancel closes the race where a write was issued a beat\r\n    /// AFTER a cancel (cancellation only reaches ALREADY-issued ops).\r\n    /// `cause` (the organic caller's write error) rides the retirement record\r\n    /// when the retirement is NOT deadline-driven.\r\n    fn poison_and_cancel(&self, cause: Option<&io::Error>) {\r\n        if !self.poisoned.swap(true, Ordering::AcqRel) {\r\n            // First poison of this conn: loud, once. The token is class-split\r\n            // (F-039 leg a): deadline → CONN_WRITE_POISONED (the wedge\r\n            // observable), organic fast-fail → CONN_WRITE_RETIRED.\r\n            self.lifecycle_log.emit(format_args!(\r\n                \"{}\",\r\n                self.render_retirement(self.timed_out.load(Ordering::Acquire), cause)\r\n            ));\r\n        }\r\n        #[cfg(windows)]\r\n        // SAFETY: scalar kernel32 call on a handle whose owning half is kept\r\n        // alive until Drop (gate / writer stack / graveyard — see `raw`).\r\n        unsafe {\r\n            let _ = win::CancelIoEx(self.raw, core::ptr::null_mut());\r\n        }\r\n        #[cfg(unix)]\r\n        // SAFETY: scalar shutdown(2) on an fd alive until Drop. SHUT_RDWR\r\n        // fails the blocked write (EPIPE), EOFs the parked read, and fails\r\n        // every future op — on Unix this one call IS the whole teardown.\r\n        unsafe {\r\n            let _ = libc::shutdown(self.raw, libc::SHUT_RDWR);\r\n        }\r\n        // Wake gate-waiters so they observe the poison instead of sleeping\r\n        // out their full deadline, and the watchdog so it re-evaluates.\r\n        self.gate_cv.notify_all();\r\n        self.dog_cv.notify_all();\r\n    }\r\n\r\n    /// Sever the physical connection (exactly once), AFTER the canceled op\r\n    /// reported completion: on Windows `DisconnectNamedPipe` forces the\r\n    /// client off (its next access fails), makes a re-issued server-side\r\n    /// read fail (so `handle_conn` reaches EOF cleanup even if it slipped a\r\n    /// new read in after the cancel), and defuses the drop-limbo flush; the\r\n    /// trailing re-cancel covers exactly that slipped-in read. On Unix the\r\n    /// `shutdown` in [`poison_and_cancel`] already did all of this.\r\n    // On Unix the cfg(windows) teardown strips out and the idempotence\r\n    // guard's early `return` becomes the function tail — needless_return\r","truncated":true,"truncatedBy":"lines","totalLines":1000,"totalBytes":18975,"outputLines":423,"outputBytes":18975,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":1000,"displayContent":{"text":"//! Bounded, cancelable, poison-on-failure framed writes for one physical\r\n//! broker connection (UPDATE-WEDGE round-4).\r\n//!\r\n//! ## The failure class this kills\r\n//!\r\n//! A broker writer used to hold the connection's `Mutex<SendHalf>` guard\r\n//! ACROSS a blocking [`write_frame`]. On Windows, `interprocess 2.4.2` routes\r\n//! that send to `WriteFileEx` + `SleepEx(INFINITE, alertable)` — a local-socket\r\n//! path with **no supported write timeout** (`set_timeout` → `Unsupported`).\r\n//! A connected consumer that stops reading (an `rc --take` whose process\r\n//! wedged; a suspended brain) therefore blocks the write INDEFINITELY, and the\r\n//! logical stall-evict (`stall_evict_controller`) releases only the controller\r\n//! ROLE — it never cancels the in-flight pipe write or closes the physical\r\n//! connection, so the writer (and every write queued behind the same gate)\r\n//! stays wedged until the peer process dies (~128 s in the field capture, on a\r\n//! brain restart).\r\n//!\r\n//! ## The primitive\r\n//!\r\n//! [`BrokerConn`] owns the send half outright — it never leaves this module,\r\n//! so an unbounded bypass write cannot compile. Every write:\r\n//!\r\n//! 1. acquires the serialized write GATE with a bounded wait (an ABSOLUTE\r\n//!    deadline stamped at write entry covers BOTH the gate-wait and the OS\r\n//!    write);\r\n//! 2. arms an INDEPENDENT per-connection watchdog thread with that absolute\r\n//!    deadline, then performs the blocking [`write_frame`] OUT of every lock;\r\n//! 3. on success, returns the half to the gate.\r\n//!\r\n//! The watchdog fires OUT OF BAND at the deadline — it never relies on the\r\n//! write returning, nor on opportunistic stall-evict sampling: it poisons the\r\n//! connection and aborts the in-flight I/O (`CancelIoEx` on Windows — the\r\n//! completion APC makes the wedged `SleepEx` loop return with\r\n//! `ERROR_OPERATION_ABORTED`; `shutdown(SHUT_RDWR)` on Unix — the blocked\r\n//! write returns `EPIPE`). A gate-waiter that exhausts its own deadline fires\r\n//! the same abort (the stuck op it waited behind is the wedge).\r\n//!\r\n//! ## Poison / retire invariant (both OSes)\r\n//!\r\n//! On deadline, write failure, or cancellation: (1) the WHOLE physical\r\n//! connection is poisoned, idempotently — a length-prefixed frame may have\r\n//! been PARTIALLY written, so reuse would corrupt framing; (2) read and write\r\n//! sides are aborted so `handle_conn` reaches its existing EOF cleanup;\r\n//! (3) the buffer of the canceled op is released only after the op reports\r\n//! completion (`write_frame` returns only after the completion APC — the\r\n//! synchronous return IS the completion handshake); (4) the retired half is\r\n//! parked in a graveyard slot — never dropped while the watchdog might still\r\n//! touch the raw handle (drop = `CloseHandle`, and a recycled handle value\r\n//! must never see a late `CancelIoEx`), and never re-gated; (5) on Windows the\r\n//! pipe is then severed (`DisconnectNamedPipe` — the server end forces the\r\n//! client off, so the old client's next access fails, a re-issued server read\r\n//! fails, and the `interprocess` drop-limbo flush cannot block on unread\r\n//! data), followed by one more `CancelIoEx` covering a read re-issued in the\r\n//! cancel→disconnect window.\r\n//!\r\n//! No new output queue (the bounded per-sink queues + isolated writer threads\r\n//! already exist — the block this bounds is BELOW them), and no `PIPE_NOWAIT`\r\n//! (recorded mid-frame corruption risk). Unix keeps its existing semantics for\r\n//! healthy connections; only the poison/retire path is new.\r\n// [impl->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\r\n\r\nuse std::io;\r\nuse std::sync::atomic::{AtomicBool, AtomicU64, Ordering};\r\nuse std::sync::{Arc, Condvar, LazyLock, Mutex, MutexGuard, OnceLock};\r\nuse std::thread::{self, JoinHandle};\r\nuse std::time::{Duration, Instant, SystemTime, UNIX_EPOCH};\r\n\r\nuse interprocess::local_socket::SendHalf;\r\n\r\nuse crate::codec::write_frame;\r\nuse crate::frame::Envelope;\r\n\r\n/// Stable per-physical-conn identity mint — a monotonic u64 stamped at\r\n/// [`BrokerConn::new`]. Before this the only conn identity was `Arc::ptr_eq`,\r\n/// which does not survive a log line: per-line CONN_WRITE_POISONED churn was\r\n/// UNDECIDABLE between fresh-carrier churn and stderr interleave (the F-039\r\n/// RCA's terminal fork). Every lifecycle record carries `conn=<id>`.\r\n// [impl->REQ-CONN-POISON-ATTRIBUTION]\r\nstatic CONN_ID_SEQ: AtomicU64 = AtomicU64::new(0);\r\n\r\n/// The shared monotonic anchor for `mono_ms` stamps. Set explicitly at daemon\r\n/// startup ([`init_log_anchor`], both broker and brain entries) so every\r\n/// record in one process shares one base; lazily initialized on first stamp\r\n/// otherwise (tests, in-process brokers) — correlation needs a COMMON base,\r\n/// not an absolute one.\r\nstatic MONO_ANCHOR: OnceLock<Instant> = OnceLock::new();\r\n\r\n/// Anchor the monotonic half of [`log_stamp`] at process start — called first\r\n/// thing by both daemon entries, next to `scrub_identity_env`.\r\n// [impl->REQ-CONN-POISON-ATTRIBUTION]\r\npub(crate) fn init_log_anchor() {\r\n    let _ = MONO_ANCHOR.set(Instant::now());\r\n}\r\n\r\n/// Wall-clock + monotonic stamp for daemon-stderr correlation records\r\n/// (`wall_ms` = Unix epoch millis; `mono_ms` = millis since [`init_log_anchor`]).\r\n/// stderr has no per-line timestamps and broker+brain share one file — without\r\n/// BOTH clocks, interleaved records cannot be ordered (monotonic) or matched\r\n/// to field incident times (wall).\r\n// [impl->REQ-CONN-POISON-ATTRIBUTION]\r\npub(crate) fn log_stamp() -> String {\r\n    let stamp = LogTime::now();\r\n    format!(\"wall_ms={} mono_ms={}\", stamp.wall_ms, stamp.mono_ms)\r\n}\r\n\r\n/// Retain the original write instant even when its healthy line is aggregated.\r\n#[derive(Clone, Copy)]\r\nstruct LogTime {\r\n    wall_ms: u128,\r\n    mono_ms: u128,\r\n}\r\n\r\nimpl LogTime {\r\n    fn now() -> Self {\r\n        Self {\r\n            wall_ms: SystemTime::now()\r\n                .duration_since(UNIX_EPOCH)\r\n                .map(|d| d.as_millis())\r\n                .unwrap_or(0),\r\n            mono_ms: MONO_ANCHOR.get_or_init(Instant::now).elapsed().as_millis(),\r\n        }\r\n    }\r\n}\r\n\r\nconst HEALTHY_LOG_INTERVAL: Duration = Duration::from_secs(60);\r\n\r\nstruct HealthyWindow {\r\n    since: Instant,\r\n    starts: u64,\r\n    closes: u64,\r\n}\r\n\r\n/// One process-wide counter window, not one limiter per connection. No thread,\r\n/// queue or per-conn healthy string: the next event after 60s emits the totals.\r\n/// Exceptional and role-transition records bypass this window entirely.\r\n// [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\r\nstruct LifecycleLog {\r\n    healthy: Mutex<HealthyWindow>,\r\n    #[cfg(test)]\r\n    capture: Option<Mutex<Vec<String>>>,\r\n    #[cfg(test)]\r\n    clock: Option<Mutex<Instant>>,\r\n}\r\n\r\nimpl LifecycleLog {\r\n    fn new() -> Self {\r\n        Self {\r\n            healthy: Mutex::new(HealthyWindow {\r\n                since: Instant::now(),\r\n                starts: 0,\r\n                closes: 0,\r\n            }),\r\n            #[cfg(test)]\r\n            capture: None,\r\n            #[cfg(test)]\r\n            clock: None,\r\n        }\r\n    }\r\n\r\n    #[cfg(test)]\r\n    fn captured(now: Instant) -> Self {\r\n        let mut log = Self::new();\r\n        log.healthy.get_mut().unwrap().since = now;\r\n        log.capture = Some(Mutex::new(Vec::new()));\r\n        log.clock = Some(Mutex::new(now));\r\n        log\r\n    }\r\n\r\n    fn emit(&self, args: std::fmt::Arguments<'_>) {\r\n        #[cfg(test)]\r\n        if let Some(lines) = &self.capture {\r\n            recover(lines).push(args.to_string());\r\n            return;\r\n        }\r\n        spt_proto::emit_line_err!(\"{args}\");\r\n    }\r\n\r\n    fn healthy(&self, start: bool) {\r\n        let mut window = recover(&self.healthy);\r\n        // Read time under the counter lock so concurrent callers cannot move\r\n        // the window backwards after waiting for a newer caller to flush.\r\n        let now = Instant::now();\r\n        #[cfg(test)]\r\n        let now = self.clock.as_ref().map(|c| *recover(c)).unwrap_or(now);\r\n        if start {\r\n            window.starts = window.starts.saturating_add(1);\r\n        } else {\r\n            window.closes = window.closes.saturating_add(1);\r\n        }\r\n        let elapsed = now.saturating_duration_since(window.since);\r\n        if elapsed < HEALTHY_LOG_INTERVAL {\r\n            return;\r\n        }\r\n        let (starts, closes) = (window.starts, window.closes);\r\n        *window = HealthyWindow { since: now, starts: 0, closes: 0 };\r\n        drop(window);\r\n        self.emit(format_args!(\r\n            \"CONN_LIFECYCLE_SUMMARY: scope=broker-ipc {} interval_ms={} starts={starts} closes={closes} \\\r\n             [REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\",\r\n            log_stamp(), elapsed.as_millis()\r\n        ));\r\n    }\r\n}\r\n\r\n/// Attribution-label cap: facts accrete per role attach (a long-lived brain\r\n/// conn can be controller + viewer + presence subscriber over its life), and\r\n/// the label must stay BOUNDED — one truncation marker, then no further growth.\r\nconst LABEL_CAP: usize = 240;\r\n\r\n/// One bounded write's outcome (beyond `Err` = poisoned/timed-out/failed).\r\n#[derive(Debug, PartialEq, Eq)]\r\npub(crate) enum ConnWrite {\r\n    /// The frame was fully written.\r\n    Done,\r\n    /// The pre-write gate closure declined (e.g. a superseded controller\r\n    /// writer's epoch check) — nothing was written; the conn stays healthy.\r\n    Superseded,\r\n}\r\n\r\n/// The serialized write gate. `half: Some` = free; `None` = a write is in\r\n/// flight (checked out) or the conn has retired (then `poisoned` is set).\r\nstruct Gate {\r\n    half: Option<SendHalf>,\r\n}\r\n\r\n/// Watchdog bookkeeping: the currently in-flight op and its ABSOLUTE deadline.\r\nstruct Dog {\r\n    /// `Some((op_id, deadline))` while a write is in flight.\r\n    inflight: Option<(u64, Instant)>,\r\n    /// Raised by [`BrokerConn::drop`]; the watchdog exits.\r\n    shutdown: bool,\r\n}\r\n\r\n/// The state shared with the watchdog thread (no `Arc<BrokerConn>` cycle —\r\n/// the watchdog must not keep the conn alive).\r\nstruct Inner {\r\n    /// Stable per-physical-conn id ([`CONN_ID_SEQ`]) — on every lifecycle record.\r\n    id: u64,\r\n    /// Accreted attribution facts (subscriber role, endpoint/session where\r\n    /// known — [`BrokerConn::describe`]). Bounded by [`LABEL_CAP`].\r\n    label: Mutex<String>,\r\n    /// Once-per-conn original first committed write time; retained on failures.\r\n    first_write: OnceLock<LogTime>,\r\n    lifecycle_log: Arc<LifecycleLog>,\r\n    gate: Mutex<Gate>,\r\n    gate_cv: Condvar,\r\n    dog: Mutex<Dog>,\r\n    dog_cv: Condvar,\r\n    /// Idempotently latched on the first deadline/failure/cancel. A poisoned\r\n    /// conn refuses every future write and is never re-gated.\r\n    poisoned: AtomicBool,\r\n    /// Latched when the retirement was DEADLINE-driven (watchdog fired, or a\r\n    /// gate-waiter exhausted its bound) — as opposed to an organic I/O error\r\n    /// (peer death). Callers see it as `io::ErrorKind::TimedOut` and can\r\n    /// count the wedge (the W2 stall-evict tally) without conflating it with\r\n    /// an ordinary disconnect.\r\n    timed_out: AtomicBool,\r\n    /// The physical sever ran (exactly once; Windows `DisconnectNamedPipe`).\r\n    severed: AtomicBool,\r\n    /// A retired (poisoned) half parks here so its OS handle stays alive\r\n    /// until [`BrokerConn::drop`] — after the watchdog is joined, so no raw\r\n    /// `CancelIoEx`/`DisconnectNamedPipe` can race the `CloseHandle`.\r\n    graveyard: Mutex<Option<SendHalf>>,\r\n    /// Monotonic per-write op identity (the watchdog cancels a SPECIFIC op).\r\n    op_seq: AtomicU64,\r\n    /// The absolute per-write bound (gate-wait + OS write completion).\r\n    bound: Duration,\r\n    /// The raw OS handle/fd of the underlying stream, captured at\r\n    /// construction for the out-of-band abort. The recv and send halves share\r\n    /// ONE underlying stream object (`interprocess` splits by ref-clone), so\r\n    /// aborting this handle covers a parked read too. Valid for the whole\r\n    /// conn lifetime: the half lives in `gate`, a writer's stack, or\r\n    /// `graveyard` until drop.\r\n    #[cfg(windows)]\r\n    raw: isize,\r\n    #[cfg(unix)]\r\n    raw: i32,\r\n}\r\n\r\n/// Recover a mutex guard even if a holder panicked — the state these guards\r\n/// protect is a plain slot/flag pair, valid on every panic boundary.\r\nfn recover<T>(m: &Mutex<T>) -> MutexGuard<'_, T> {\r\n    match m.lock() {\r\n        Ok(g) => g,\r\n        Err(p) => p.into_inner(),\r\n    }\r\n}\r\n\r\n#[cfg(windows)]\r\nfn raw_of(half: &SendHalf) -> isize {\r\n    use std::os::windows::io::{AsHandle, AsRawHandle};\r\n    match half {\r\n        SendHalf::NamedPipe(h) => h.as_handle().as_raw_handle() as isize,\r\n    }\r\n}\r\n\r\n#[cfg(unix)]\r\nfn raw_of(half: &SendHalf) -> i32 {\r\n    use std::os::fd::{AsFd, AsRawFd};\r\n    match half {\r\n        SendHalf::UdSocket(h) => h.as_fd().as_raw_fd(),\r\n    }\r\n}\r\n\r\n#[cfg(windows)]\r\nmod win {\r\n    // kernel32 is in the MSVC default link set — raw externs, no new dep\r\n    // (the deelevate.rs / daemon.rs posture).\r\n    extern \"system\" {\r\n        pub fn CancelIoEx(handle: isize, overlapped: *mut core::ffi::c_void) -> i32;\r\n        pub fn DisconnectNamedPipe(handle: isize) -> i32;\r\n    }\r\n}\r\n\r\nimpl Inner {\r\n    /// The attribution snapshot every lifecycle record carries: stable conn id,\r\n    /// accreted role/endpoint/session facts, wall + monotonic stamps.\r\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n    fn attribution(&self) -> String {\r\n        let label = recover(&self.label);\r\n        let facts: &str = if label.is_empty() { \"role=unattributed\" } else { &label };\r\n        match self.first_write.get() {\r\n            Some(start) => format!(\r\n                \"conn={} {} {} first_write_wall_ms={} first_write_mono_ms={}\",\r\n                self.id, facts, log_stamp(), start.wall_ms, start.mono_ms\r\n            ),\r\n            None => format!(\"conn={} {} {} first_write=none\", self.id, facts, log_stamp()),\r\n        }\r\n    }\r\n\r\n    /// Render the once-per-conn retirement record — the F-039 token split\r\n    /// (leg a): `CONN_WRITE_POISONED` is RESERVED for the deadline class\r\n    /// (`timed_out` latched — the wedge observable blackhole-controller.ps1\r\n    /// watches); an ORGANIC write failure (peer death / broken pipe on an\r\n    /// already-dead counterpart, the offline-peer dial churn) emits the\r\n    /// distinct `CONN_WRITE_RETIRED` token carrying the error kind and the\r\n    /// SAME attribution fields — quieter class, never thinner (the retired\r\n    /// class is what discriminates fresh-carrier churn from interleave\r\n    /// artifact in the field). Behavior is identical either way: poison,\r\n    /// abort, retire, never reuse.\r\n    // [impl->REQ-CONN-POISON-DIAL-SCOPE]\r\n    fn render_retirement(&self, timed_out: bool, cause: Option<&io::Error>) -> String {\r\n        if timed_out {\r\n            format!(\r\n                \"CONN_WRITE_POISONED: {} — physical broker conn retired: a write \\\r\n                 exceeded its {}ms bound; in-flight I/O aborted, conn never reused \\\r\n                 [REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\",\r\n                self.attribution(),\r\n                self.bound.as_millis()\r\n            )\r\n        } else {\r\n            let cause = match cause {\r\n                Some(e) => format!(\"{:?}: {e}\", e.kind()),\r\n                None => \"write failed\".to_string(),\r\n            };\r\n            format!(\r\n                \"CONN_WRITE_RETIRED: {} — organic write failure ({cause}); physical \\\r\n                 broker conn poisoned + retired, never reused (same retire behavior \\\r\n                 as POISONED — only the failure class differs) \\\r\n                 [REQ-CONN-POISON-DIAL-SCOPE]\",\r\n                self.attribution()\r\n            )\r\n        }\r\n    }\r\n\r\n    /// Per-conn exceptional and role-transition records remain unconditional.\r\n    /// releases#286 supersedes only healthy start/close emission, which uses\r\n    /// the process-wide interval counter instead. Retirement retains the\r\n    /// original first-write time even when the start line was aggregated.\r\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n    fn lifecycle(&self, event: &str, extra: &str) {\r\n        let sep = if extra.is_empty() { \"\" } else { \" \" };\r\n        self.lifecycle_log.emit(format_args!(\r\n            \"CONN_LIFECYCLE: {} event={event}{sep}{extra} [REQ-CONN-POISON-ATTRIBUTION]\",\r\n            self.attribution()\r\n        ));\r\n    }\r\n\r\n    /// Abort the connection's in-flight I/O OUT OF BAND (both directions —\r\n    /// the halves share one OS object) and latch the poison. Idempotent and\r\n    /// safe to repeat: `CancelIoEx` on a handle with nothing pending is a\r\n    /// no-op, and re-cancel closes the race where a write was issued a beat\r\n    /// AFTER a cancel (cancellation only reaches ALREADY-issued ops).\r\n    /// `cause` (the organic caller's write error) rides the retirement record\r\n    /// when the retirement is NOT deadline-driven.\r\n    fn poison_and_cancel(&self, cause: Option<&io::Error>) {\r\n        if !self.poisoned.swap(true, Ordering::AcqRel) {\r\n            // First poison of this conn: loud, once. The token is class-split\r\n            // (F-039 leg a): deadline → CONN_WRITE_POISONED (the wedge\r\n            // observable), organic fast-fail → CONN_WRITE_RETIRED.\r\n            self.lifecycle_log.emit(format_args!(\r\n                \"{}\",\r\n                self.render_retirement(self.timed_out.load(Ordering::Acquire), cause)\r\n            ));\r\n        }\r\n        #[cfg(windows)]\r\n        // SAFETY: scalar kernel32 call on a handle whose owning half is kept\r\n        // alive until Drop (gate / writer stack / graveyard — see `raw`).\r\n        unsafe {\r\n            let _ = win::CancelIoEx(self.raw, core::ptr::null_mut());\r\n        }\r\n        #[cfg(unix)]\r\n        // SAFETY: scalar shutdown(2) on an fd alive until Drop. SHUT_RDWR\r\n        // fails the blocked write (EPIPE), EOFs the parked read, and fails\r\n        // every future op — on Unix this one call IS the whole teardown.\r\n        unsafe {\r\n            let _ = libc::shutdown(self.raw, libc::SHUT_RDWR);\r\n        }\r\n        // Wake gate-waiters so they observe the poison instead of sleeping\r\n        // out their full deadline, and the watchdog so it re-evaluates.\r\n        self.gate_cv.notify_all();\r\n        self.dog_cv.notify_all();\r\n    }\r\n\r\n    /// Sever the physical connection (exactly once), AFTER the canceled op\r\n    /// reported completion: on Windows `DisconnectNamedPipe` forces the\r\n    /// client off (its next access fails), makes a re-issued server-side\r\n    /// read fail (so `handle_conn` reaches EOF cleanup even if it slipped a\r\n    /// new read in after the cancel), and defuses the drop-limbo flush; the\r\n    /// trailing re-cancel covers exactly that slipped-in read. On Unix the\r\n    /// `shutdown` in [`poison_and_cancel`] already did all of this.\r\n    // On Unix the cfg(windows) teardown strips out and the idempotence\r\n    // guard's early `return` becomes the function tail — needless_return\r\n…\n}","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120,121,122,123,124,125,126,127,128,129,130,131,132,133,134,135,136,137,138,139,140,141,142,143,144,145,146,147,148,149,150,151,152,153,154,155,156,157,158,159,160,161,162,163,164,165,166,167,168,169,170,171,172,173,174,175,176,177,178,179,180,181,182,183,184,185,186,187,188,189,190,191,192,193,194,195,196,197,198,199,200,201,202,203,204,205,206,207,208,209,210,211,212,213,214,215,216,217,218,219,220,221,222,223,224,225,226,227,228,229,230,231,232,233,234,235,236,237,238,239,240,241,242,243,244,245,246,247,248,249,250,251,252,253,254,255,256,257,258,259,260,261,262,263,264,265,266,267,268,269,270,271,272,273,274,275,276,277,278,279,280,281,282,283,284,285,286,287,288,289,290,291,292,293,294,295,296,297,298,299,300,301,302,303,304,305,306,307,308,309,310,311,312,313,314,315,316,317,318,319,320,321,322,323,324,325,326,327,328,329,330,331,332,333,334,335,336,337,338,339,340,341,342,343,344,345,346,347,348,349,350,351,352,353,354,355,356,357,358,359,360,361,362,363,364,365,366,367,368,369,370,371,372,373,374,375,376,377,378,379,380,381,382,383,384,385,386,387,388,389,390,391,392,393,394,395,396,397,398,399,400,401,402,403,404,405,406,407,408,409,410,411,412,413,414,415,416,417,418,419,420,421,422,423,null,511]},"fileSize":43321,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-daemon\\src\\conn.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":1000,"totalBytes":18975,"outputLines":423,"outputBytes":18975,"shownRange":{"start":1,"end":423},"nextOffset":424}}},"isError":false,"timestamp":1788999516694}}
{"type":"message","id":"8d198e16","parentId":"ad21c169","timestamp":"2026-09-10T00:18:36.699Z","message":{"role":"toolResult","toolCallId":"call_CFcXvjshpb2DT4KOYKAz71ED|fc_0dde72effda0ad21016aa1f75c941887d28404af22f47d4c68","toolName":"read","content":[{"type":"text","text":"impl Inner {\n…\n    /// trailing re-cancel covers exactly that slipped-in read. On Unix the\n    /// `shutdown` in [`poison_and_cancel`] already did all of this.\n    // On Unix the cfg(windows) teardown strips out and the idempotence\n    // guard's early `return` becomes the function tail — needless_return\n    // there, load-bearing here. Restructuring (e.g. `if !swapped { … }`)\n    // would leave Unix an empty if-body (`clippy::needless_if`) instead.\n    #[cfg_attr(unix, allow(clippy::needless_return))]\n    fn sever(&self) {\n        if self.severed.swap(true, Ordering::AcqRel) {\n            return;\n        }\n        #[cfg(windows)]\n        // SAFETY: as in poison_and_cancel — handle outlives the watchdog.\n        unsafe {\n            let _ = win::DisconnectNamedPipe(self.raw);\n            let _ = win::CancelIoEx(self.raw, core::ptr::null_mut());\n        }\n    }\n\n    /// The watchdog body: park until a write is in flight, then sleep to its\n    /// absolute deadline. If the op is still in flight at the deadline,\n    /// poison + cancel, re-canceling periodically until the op reports\n    /// completion (the writer clears `inflight` the instant `write_frame`\n    /// returns — that return IS the buffer-release handshake), then sever.\n    fn watchdog(&self) {\n        let mut d = recover(&self.dog);\n        loop {\n            // Park until armed (or shutdown).\n            let (op, deadline) = loop {\n                if d.shutdown {\n                    return;\n                }\n                if let Some(armed) = d.inflight {\n                    break armed;\n                }\n                d = self\n                    .dog_cv\n                    .wait(d)\n                    .unwrap_or_else(|p| p.into_inner());\n            };\n            // Sleep toward the deadline while THIS op stays in flight.\n            let fired = loop {\n                if d.shutdown {\n                    return;\n                }\n                match d.inflight {\n                    Some((cur, _)) if cur == op => {}\n                    // Completed (or a newer op armed) before the deadline.\n                    _ => break false,\n                }\n                let now = Instant::now();\n                if now >= deadline {\n                    break true;\n                }\n                let (g, _) = self\n                    .dog_cv\n                    .wait_timeout(d, deadline - now)\n                    .unwrap_or_else(|p| p.into_inner());\n                d = g;\n            };\n            if !fired {\n                continue;\n            }\n            // Deadline: abort out of band, then wait for the canceled op to\n            // report completion before severing (re-cancel each slice — a\n            // cancel that raced the WriteFileEx issuance reaches nothing, so\n            // one shot is not enough). Latch timed_out FIRST so the aborted\n            // writer observes its failure as a DEADLINE retirement (wedge),\n            // not an organic conn death.\n            drop(d);\n            self.timed_out.store(true, Ordering::Release);\n            self.poison_and_cancel(None);\n            d = recover(&self.dog);\n            while matches!(d.inflight, Some((cur, _)) if cur == op) && !d.shutdown {\n                let (g, _) = self\n                    .dog_cv\n                    .wait_timeout(d, Duration::from_millis(100))\n                    .unwrap_or_else(|p| p.into_inner());\n                d = g;\n                drop(d);\n                self.poison_and_cancel(None);\n                d = recover(&self.dog);\n            }\n            if d.shutdown {\n                return;\n            }\n            drop(d);\n            self.sever();\n            d = recover(&self.dog);\n        }\n    }\n}\n\n/// The shared, serialized, BOUNDED send side of one physical broker\n/// connection — what [`crate::broker::SharedSend`] now aliases. Owns the\n/// [`SendHalf`] (it never leaves this module) and the per-conn watchdog.\npub(crate) struct BrokerConn {\n    inner: Arc<Inner>,\n    /// Joined on drop — after `shutdown` is raised — so the raw handle the\n    /// watchdog aborts through strictly outlives every raw use.\n    dog: Option<JoinHandle<()>>,\n}\n\nimpl BrokerConn {\n    /// Wrap a freshly-split send half. `bound` is the ABSOLUTE per-write\n    /// deadline (gate-wait + OS write completion) — the broker passes its\n    /// `brain_write_deadline()` (`SPT_BRAIN_WRITE_DEADLINE_MS` knob).\n    pub(crate) fn new(half: SendHalf, bound: Duration) -> Self {\n        static LOG: LazyLock<Arc<LifecycleLog>> = LazyLock::new(|| Arc::new(LifecycleLog::new()));\n        Self::with_log(half, bound, Arc::clone(&LOG))\n    }\n\n    fn with_log(half: SendHalf, bound: Duration, lifecycle_log: Arc<LifecycleLog>) -> Self {\n        let raw = raw_of(&half);\n        let inner = Arc::new(Inner {\n            id: CONN_ID_SEQ.fetch_add(1, Ordering::Relaxed) + 1,\n            label: Mutex::new(String::new()),\n            first_write: OnceLock::new(),\n            lifecycle_log,\n            gate: Mutex::new(Gate { half: Some(half) }),\n            gate_cv: Condvar::new(),\n            dog: Mutex::new(Dog {\n                inflight: None,\n                shutdown: false,\n            }),\n            dog_cv: Condvar::new(),\n            poisoned: AtomicBool::new(false),\n            timed_out: AtomicBool::new(false),\n            severed: AtomicBool::new(false),\n            graveyard: Mutex::new(None),\n            op_seq: AtomicU64::new(0),\n            bound,\n            raw,\n        });\n        let dog_inner = Arc::clone(&inner);\n        let dog = thread::Builder::new()\n            .name(\"conn-watchdog\".into())\n            .spawn(move || dog_inner.watchdog())\n            .ok();\n        BrokerConn { inner, dog }\n    }\n\n    /// The stable per-physical-conn id — what lifecycle records key on\n    /// (`Arc::ptr_eq` identity does not survive a log line).\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\n    pub(crate) fn id(&self) -> u64 {\n        self.inner.id\n    }\n\n    /// Accrete an attribution fact onto this conn's label (subscriber role,\n    /// endpoint/session where known — e.g. `role=brain`,\n    /// `controller session=7 by=cli@host`). Facts accumulate space-separated\n    /// and ride EVERY subsequent lifecycle record for this conn; growth is\n    /// bounded by [`LABEL_CAP`] (one `…` marker, then no-op).\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\n    pub(crate) fn describe(&self, fact: &str) {\n        let mut l = recover(&self.inner.label);\n        if l.ends_with('…') {\n            return;\n        }\n        if l.len() + fact.len() + 1 > LABEL_CAP {\n            l.push('…');\n            return;\n        }\n        if !l.is_empty() {\n            l.push(' ');\n        }\n        l.push_str(fact);\n    }\n\n    /// Emit a bounded per-conn lifecycle record through this conn's shared\n    /// attribution format — the broker's sink machinery uses it for\n    /// `writer-exit`, attach/replace, and detach events (leg d).\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\n    pub(crate) fn lifecycle_event(&self, event: &str, extra: &str) {\n        self.inner.lifecycle(event, extra);\n    }\n\n    /// Bounded framed write. See [`Self::write_gated`].\n    pub(crate) fn write(&self, env: &Envelope) -> io::Result<ConnWrite> {\n        self.write_gated(env, || true)\n    }\n\n    /// Bounded framed write with a pre-write gate closure, run AFTER the\n    /// serialized gate is acquired and BEFORE any byte is written — the same\n    /// serialization point the old epoch-check-under-`send.lock()` had, so a\n    /// superseded controller writer still bails atomically with respect to\n    /// its successor's writes (P1c, REQ-HAZARD-CONTROLLER-WRITER-REORDER).\n    /// `pre` returning false → `Ok(ConnWrite::Superseded)`, nothing written.\n    pub(crate) fn write_gated(\n        &self,\n        env: &Envelope,\n        pre: impl FnOnce() -> bool,\n    ) -> io::Result<ConnWrite> {\n        let inner = &*self.inner;\n        let deadline = Instant::now() + inner.bound;\n\n        // ── Acquire the gate, bounded by the SAME absolute deadline. ──\n        let mut g = recover(&inner.gate);\n        loop {\n            if inner.poisoned.load(Ordering::Acquire) {\n                return Err(poisoned_err());\n            }\n            if g.half.is_some() {\n                break;\n            }\n            let now = Instant::now();\n            if now >= deadline {\n                // The write we queued behind is stuck past OUR whole bound:\n                // that op is the wedge. Abort it out of band and fail — the\n                // watchdog's own deadline (armed at that op's entry) is also\n                // in flight; both paths converge on the same idempotent\n                // poison/cancel.\n                drop(g);\n                inner.timed_out.store(true, Ordering::Release);\n                inner.poison_and_cancel(None);\n                return Err(timeout_err());\n            }\n            let (g2, _) = inner\n                .gate_cv\n                .wait_timeout(g, deadline - now)\n                .unwrap_or_else(|p| p.into_inner());\n            g = g2;\n        }\n\n        if !pre() {\n            return Ok(ConnWrite::Superseded);\n        }\n\n        // The gate serializes first-write initialization. Save the instant\n        // BEFORE I/O so a later failure still names when this conn went active.\n        // [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n        if inner.first_write.get().is_none() {\n            let _ = inner.first_write.set(LogTime::now());\n            inner.lifecycle_log.healthy(true);\n        }\n\n        // ── Check out the half + arm the watchdog, then write OUT of locks. ──\n        let mut half = g.half.take().expect(\"gate half present at checkout\");\n        let op = inner.op_seq.fetch_add(1, Ordering::Relaxed) + 1;\n        {\n            let mut d = recover(&inner.dog);\n            d.inflight = Some((op, deadline));\n            inner.dog_cv.notify_all();\n        }\n        drop(g);\n\n        let res = write_frame(&mut half, env);\n\n        // Completion handshake: `write_frame` has returned, so the OS op —\n        // canceled or not — reported completion and the frame buffer is\n        // released. Clear the arm BEFORE deciding health, so the watchdog\n        // never severs while the op is formally in flight.\n        {\n            let mut d = recover(&inner.dog);\n            d.inflight = None;\n            inner.dog_cv.notify_all();\n        }\n\n        let mut g = recover(&inner.gate);\n        if res.is_ok() && !inner.poisoned.load(Ordering::Acquire) {\n            g.half = Some(half);\n            drop(g);\n            inner.gate_cv.notify_all();\n            return Ok(ConnWrite::Done);\n        }\n\n        // ── Retire: poison the whole conn, abort the read side so\n        // `handle_conn` reaches EOF cleanup, sever, and park the half in the\n        // graveyard (its handle must outlive any late raw abort; never\n        // re-gated, never reused — a timed-out length-prefixed frame may be\n        // partially written). ──\n        drop(g);\n        inner.poison_and_cancel(res.as_ref().err());\n        inner.sever();\n        *recover(&inner.graveyard) = Some(half);\n        inner.gate_cv.notify_all();\n        // A DEADLINE retirement (the watchdog aborted OUR op, or the conn was\n        // already deadline-poisoned) surfaces as TimedOut — the caller can\n        // count the wedge (W2 stall-evict tally). An organic I/O failure\n        // (peer death) keeps its own error.\n        if inner.timed_out.load(Ordering::Acquire) {\n            return Err(timeout_err());\n        }\n        Err(match res {\n            Err(e) => e,\n            Ok(()) => poisoned_err(),\n        })\n    }\n}\n\n/// TEST-ONLY determinism lever: holds the serialized write gate (the half is\n/// checked out) so every writer parks AT THE GATE until the pin drops — the\n/// round-4 analog of pinning the old `Mutex<SendHalf>` across a registration\n/// window (the P1c epoch-gate keystone unit relies on it).\n#[cfg(test)]\npub(crate) struct TestGatePin {\n    inner: Arc<Inner>,\n    half: Option<SendHalf>,\n}\n\n#[cfg(test)]\nimpl Drop for TestGatePin {\n    fn drop(&mut self) {\n        let mut g = recover(&self.inner.gate);\n        g.half = self.half.take();\n        drop(g);\n        self.inner.gate_cv.notify_all();\n    }\n}\n\n#[cfg(test)]\nimpl BrokerConn {\n    /// Check the half out of the gate without arming the watchdog — writers\n    /// queue behind the gate (bounded by their own deadlines) until the\n    /// returned pin drops. Test-only; never a production path.\n    pub(crate) fn pin_gate_for_test(&self) -> TestGatePin {\n        let mut g = recover(&self.inner.gate);\n        let half = g.half.take().expect(\"pin: gate half present\");\n        TestGatePin {\n            inner: Arc::clone(&self.inner),\n            half: Some(half),\n        }\n    }\n}\n\nimpl Drop for BrokerConn {\n    fn drop(&mut self) {\n        // Only a healthy physical close is counted instead of logged per conn.\n        // Error/timeout closes keep their full attribution and original start.\n        // [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n        if self.inner.poisoned.load(Ordering::Acquire) {\n            self.inner.lifecycle(\n                \"transport-close\",\n                &format!(\n                    \"poisoned=true timed_out={}\",\n                    self.inner.timed_out.load(Ordering::Acquire)\n                ),\n            );\n        } else {\n            self.inner.lifecycle_log.healthy(false);\n        }\n        {\n            let mut d = recover(&self.inner.dog);\n            d.shutdown = true;\n            self.inner.dog_cv.notify_all();\n        }\n        if let Some(dog) = self.dog.take() {\n            let _ = dog.join();\n        }\n        // Halves (gate or graveyard) drop AFTER the watchdog is gone — no raw\n        // handle use can race the CloseHandle/limbo path.\n    }\n}\n\nfn poisoned_err() -> io::Error {\n    io::Error::new(\n        io::ErrorKind::BrokenPipe,\n        \"broker conn poisoned (bounded write retired the physical connection)\",\n    )\n}\n\nfn timeout_err() -> io::Error {\n    io::Error::new(\n        io::ErrorKind::TimedOut,\n        \"broker conn write exceeded its bound (conn poisoned + aborted)\",\n    )\n}\n\n#[cfg(test)]\nmod tests {\n    use super::*;\n    use crate::transport::{DaemonTransport, LocalSocketTransport};\n    use interprocess::local_socket::traits::Stream as _;\n    use interprocess::local_socket::Stream;\n    use std::sync::atomic::AtomicU32;\n\n    /// A real (loopback local-socket) conn pair — the send half wrapped in the\n    /// primitive, the client end returned so a test can hold it alive or drop\n…\n}\n\n[Showing lines 420-798 of 1000. Use :799 to continue]"}],"details":{"truncation":{"content":"    /// trailing re-cancel covers exactly that slipped-in read. On Unix the\r\n    /// `shutdown` in [`poison_and_cancel`] already did all of this.\r\n    // On Unix the cfg(windows) teardown strips out and the idempotence\r\n    // guard's early `return` becomes the function tail — needless_return\r\n    // there, load-bearing here. Restructuring (e.g. `if !swapped { … }`)\r\n    // would leave Unix an empty if-body (`clippy::needless_if`) instead.\r\n    #[cfg_attr(unix, allow(clippy::needless_return))]\r\n    fn sever(&self) {\r\n        if self.severed.swap(true, Ordering::AcqRel) {\r\n            return;\r\n        }\r\n        #[cfg(windows)]\r\n        // SAFETY: as in poison_and_cancel — handle outlives the watchdog.\r\n        unsafe {\r\n            let _ = win::DisconnectNamedPipe(self.raw);\r\n            let _ = win::CancelIoEx(self.raw, core::ptr::null_mut());\r\n        }\r\n    }\r\n\r\n    /// The watchdog body: park until a write is in flight, then sleep to its\r\n    /// absolute deadline. If the op is still in flight at the deadline,\r\n    /// poison + cancel, re-canceling periodically until the op reports\r\n    /// completion (the writer clears `inflight` the instant `write_frame`\r\n    /// returns — that return IS the buffer-release handshake), then sever.\r\n    fn watchdog(&self) {\r\n        let mut d = recover(&self.dog);\r\n        loop {\r\n            // Park until armed (or shutdown).\r\n            let (op, deadline) = loop {\r\n                if d.shutdown {\r\n                    return;\r\n                }\r\n                if let Some(armed) = d.inflight {\r\n                    break armed;\r\n                }\r\n                d = self\r\n                    .dog_cv\r\n                    .wait(d)\r\n                    .unwrap_or_else(|p| p.into_inner());\r\n            };\r\n            // Sleep toward the deadline while THIS op stays in flight.\r\n            let fired = loop {\r\n                if d.shutdown {\r\n                    return;\r\n                }\r\n                match d.inflight {\r\n                    Some((cur, _)) if cur == op => {}\r\n                    // Completed (or a newer op armed) before the deadline.\r\n                    _ => break false,\r\n                }\r\n                let now = Instant::now();\r\n                if now >= deadline {\r\n                    break true;\r\n                }\r\n                let (g, _) = self\r\n                    .dog_cv\r\n                    .wait_timeout(d, deadline - now)\r\n                    .unwrap_or_else(|p| p.into_inner());\r\n                d = g;\r\n            };\r\n            if !fired {\r\n                continue;\r\n            }\r\n            // Deadline: abort out of band, then wait for the canceled op to\r\n            // report completion before severing (re-cancel each slice — a\r\n            // cancel that raced the WriteFileEx issuance reaches nothing, so\r\n            // one shot is not enough). Latch timed_out FIRST so the aborted\r\n            // writer observes its failure as a DEADLINE retirement (wedge),\r\n            // not an organic conn death.\r\n            drop(d);\r\n            self.timed_out.store(true, Ordering::Release);\r\n            self.poison_and_cancel(None);\r\n            d = recover(&self.dog);\r\n            while matches!(d.inflight, Some((cur, _)) if cur == op) && !d.shutdown {\r\n                let (g, _) = self\r\n                    .dog_cv\r\n                    .wait_timeout(d, Duration::from_millis(100))\r\n                    .unwrap_or_else(|p| p.into_inner());\r\n                d = g;\r\n                drop(d);\r\n                self.poison_and_cancel(None);\r\n                d = recover(&self.dog);\r\n            }\r\n            if d.shutdown {\r\n                return;\r\n            }\r\n            drop(d);\r\n            self.sever();\r\n            d = recover(&self.dog);\r\n        }\r\n    }\r\n}\r\n\r\n/// The shared, serialized, BOUNDED send side of one physical broker\r\n/// connection — what [`crate::broker::SharedSend`] now aliases. Owns the\r\n/// [`SendHalf`] (it never leaves this module) and the per-conn watchdog.\r\npub(crate) struct BrokerConn {\r\n    inner: Arc<Inner>,\r\n    /// Joined on drop — after `shutdown` is raised — so the raw handle the\r\n    /// watchdog aborts through strictly outlives every raw use.\r\n    dog: Option<JoinHandle<()>>,\r\n}\r\n\r\nimpl BrokerConn {\r\n    /// Wrap a freshly-split send half. `bound` is the ABSOLUTE per-write\r\n    /// deadline (gate-wait + OS write completion) — the broker passes its\r\n    /// `brain_write_deadline()` (`SPT_BRAIN_WRITE_DEADLINE_MS` knob).\r\n    pub(crate) fn new(half: SendHalf, bound: Duration) -> Self {\r\n        static LOG: LazyLock<Arc<LifecycleLog>> = LazyLock::new(|| Arc::new(LifecycleLog::new()));\r\n        Self::with_log(half, bound, Arc::clone(&LOG))\r\n    }\r\n\r\n    fn with_log(half: SendHalf, bound: Duration, lifecycle_log: Arc<LifecycleLog>) -> Self {\r\n        let raw = raw_of(&half);\r\n        let inner = Arc::new(Inner {\r\n            id: CONN_ID_SEQ.fetch_add(1, Ordering::Relaxed) + 1,\r\n            label: Mutex::new(String::new()),\r\n            first_write: OnceLock::new(),\r\n            lifecycle_log,\r\n            gate: Mutex::new(Gate { half: Some(half) }),\r\n            gate_cv: Condvar::new(),\r\n            dog: Mutex::new(Dog {\r\n                inflight: None,\r\n                shutdown: false,\r\n            }),\r\n            dog_cv: Condvar::new(),\r\n            poisoned: AtomicBool::new(false),\r\n            timed_out: AtomicBool::new(false),\r\n            severed: AtomicBool::new(false),\r\n            graveyard: Mutex::new(None),\r\n            op_seq: AtomicU64::new(0),\r\n            bound,\r\n            raw,\r\n        });\r\n        let dog_inner = Arc::clone(&inner);\r\n        let dog = thread::Builder::new()\r\n            .name(\"conn-watchdog\".into())\r\n            .spawn(move || dog_inner.watchdog())\r\n            .ok();\r\n        BrokerConn { inner, dog }\r\n    }\r\n\r\n    /// The stable per-physical-conn id — what lifecycle records key on\r\n    /// (`Arc::ptr_eq` identity does not survive a log line).\r\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n    pub(crate) fn id(&self) -> u64 {\r\n        self.inner.id\r\n    }\r\n\r\n    /// Accrete an attribution fact onto this conn's label (subscriber role,\r\n    /// endpoint/session where known — e.g. `role=brain`,\r\n    /// `controller session=7 by=cli@host`). Facts accumulate space-separated\r\n    /// and ride EVERY subsequent lifecycle record for this conn; growth is\r\n    /// bounded by [`LABEL_CAP`] (one `…` marker, then no-op).\r\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n    pub(crate) fn describe(&self, fact: &str) {\r\n        let mut l = recover(&self.inner.label);\r\n        if l.ends_with('…') {\r\n            return;\r\n        }\r\n        if l.len() + fact.len() + 1 > LABEL_CAP {\r\n            l.push('…');\r\n            return;\r\n        }\r\n        if !l.is_empty() {\r\n            l.push(' ');\r\n        }\r\n        l.push_str(fact);\r\n    }\r\n\r\n    /// Emit a bounded per-conn lifecycle record through this conn's shared\r\n    /// attribution format — the broker's sink machinery uses it for\r\n    /// `writer-exit`, attach/replace, and detach events (leg d).\r\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n    pub(crate) fn lifecycle_event(&self, event: &str, extra: &str) {\r\n        self.inner.lifecycle(event, extra);\r\n    }\r\n\r\n    /// Bounded framed write. See [`Self::write_gated`].\r\n    pub(crate) fn write(&self, env: &Envelope) -> io::Result<ConnWrite> {\r\n        self.write_gated(env, || true)\r\n    }\r\n\r\n    /// Bounded framed write with a pre-write gate closure, run AFTER the\r\n    /// serialized gate is acquired and BEFORE any byte is written — the same\r\n    /// serialization point the old epoch-check-under-`send.lock()` had, so a\r\n    /// superseded controller writer still bails atomically with respect to\r\n    /// its successor's writes (P1c, REQ-HAZARD-CONTROLLER-WRITER-REORDER).\r\n    /// `pre` returning false → `Ok(ConnWrite::Superseded)`, nothing written.\r\n    pub(crate) fn write_gated(\r\n        &self,\r\n        env: &Envelope,\r\n        pre: impl FnOnce() -> bool,\r\n    ) -> io::Result<ConnWrite> {\r\n        let inner = &*self.inner;\r\n        let deadline = Instant::now() + inner.bound;\r\n\r\n        // ── Acquire the gate, bounded by the SAME absolute deadline. ──\r\n        let mut g = recover(&inner.gate);\r\n        loop {\r\n            if inner.poisoned.load(Ordering::Acquire) {\r\n                return Err(poisoned_err());\r\n            }\r\n            if g.half.is_some() {\r\n                break;\r\n            }\r\n            let now = Instant::now();\r\n            if now >= deadline {\r\n                // The write we queued behind is stuck past OUR whole bound:\r\n                // that op is the wedge. Abort it out of band and fail — the\r\n                // watchdog's own deadline (armed at that op's entry) is also\r\n                // in flight; both paths converge on the same idempotent\r\n                // poison/cancel.\r\n                drop(g);\r\n                inner.timed_out.store(true, Ordering::Release);\r\n                inner.poison_and_cancel(None);\r\n                return Err(timeout_err());\r\n            }\r\n            let (g2, _) = inner\r\n                .gate_cv\r\n                .wait_timeout(g, deadline - now)\r\n                .unwrap_or_else(|p| p.into_inner());\r\n            g = g2;\r\n        }\r\n\r\n        if !pre() {\r\n            return Ok(ConnWrite::Superseded);\r\n        }\r\n\r\n        // The gate serializes first-write initialization. Save the instant\r\n        // BEFORE I/O so a later failure still names when this conn went active.\r\n        // [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\r\n        if inner.first_write.get().is_none() {\r\n            let _ = inner.first_write.set(LogTime::now());\r\n            inner.lifecycle_log.healthy(true);\r\n        }\r\n\r\n        // ── Check out the half + arm the watchdog, then write OUT of locks. ──\r\n        let mut half = g.half.take().expect(\"gate half present at checkout\");\r\n        let op = inner.op_seq.fetch_add(1, Ordering::Relaxed) + 1;\r\n        {\r\n            let mut d = recover(&inner.dog);\r\n            d.inflight = Some((op, deadline));\r\n            inner.dog_cv.notify_all();\r\n        }\r\n        drop(g);\r\n\r\n        let res = write_frame(&mut half, env);\r\n\r\n        // Completion handshake: `write_frame` has returned, so the OS op —\r\n        // canceled or not — reported completion and the frame buffer is\r\n        // released. Clear the arm BEFORE deciding health, so the watchdog\r\n        // never severs while the op is formally in flight.\r\n        {\r\n            let mut d = recover(&inner.dog);\r\n            d.inflight = None;\r\n            inner.dog_cv.notify_all();\r\n        }\r\n\r\n        let mut g = recover(&inner.gate);\r\n        if res.is_ok() && !inner.poisoned.load(Ordering::Acquire) {\r\n            g.half = Some(half);\r\n            drop(g);\r\n            inner.gate_cv.notify_all();\r\n            return Ok(ConnWrite::Done);\r\n        }\r\n\r\n        // ── Retire: poison the whole conn, abort the read side so\r\n        // `handle_conn` reaches EOF cleanup, sever, and park the half in the\r\n        // graveyard (its handle must outlive any late raw abort; never\r\n        // re-gated, never reused — a timed-out length-prefixed frame may be\r\n        // partially written). ──\r\n        drop(g);\r\n        inner.poison_and_cancel(res.as_ref().err());\r\n        inner.sever();\r\n        *recover(&inner.graveyard) = Some(half);\r\n        inner.gate_cv.notify_all();\r\n        // A DEADLINE retirement (the watchdog aborted OUR op, or the conn was\r\n        // already deadline-poisoned) surfaces as TimedOut — the caller can\r\n        // count the wedge (W2 stall-evict tally). An organic I/O failure\r\n        // (peer death) keeps its own error.\r\n        if inner.timed_out.load(Ordering::Acquire) {\r\n            return Err(timeout_err());\r\n        }\r\n        Err(match res {\r\n            Err(e) => e,\r\n            Ok(()) => poisoned_err(),\r\n        })\r\n    }\r\n}\r\n\r\n/// TEST-ONLY determinism lever: holds the serialized write gate (the half is\r\n/// checked out) so every writer parks AT THE GATE until the pin drops — the\r\n/// round-4 analog of pinning the old `Mutex<SendHalf>` across a registration\r\n/// window (the P1c epoch-gate keystone unit relies on it).\r\n#[cfg(test)]\r\npub(crate) struct TestGatePin {\r\n    inner: Arc<Inner>,\r\n    half: Option<SendHalf>,\r\n}\r\n\r\n#[cfg(test)]\r\nimpl Drop for TestGatePin {\r\n    fn drop(&mut self) {\r\n        let mut g = recover(&self.inner.gate);\r\n        g.half = self.half.take();\r\n        drop(g);\r\n        self.inner.gate_cv.notify_all();\r\n    }\r\n}\r\n\r\n#[cfg(test)]\r\nimpl BrokerConn {\r\n    /// Check the half out of the gate without arming the watchdog — writers\r\n    /// queue behind the gate (bounded by their own deadlines) until the\r\n    /// returned pin drops. Test-only; never a production path.\r\n    pub(crate) fn pin_gate_for_test(&self) -> TestGatePin {\r\n        let mut g = recover(&self.inner.gate);\r\n        let half = g.half.take().expect(\"pin: gate half present\");\r\n        TestGatePin {\r\n            inner: Arc::clone(&self.inner),\r\n            half: Some(half),\r\n        }\r\n    }\r\n}\r\n\r\nimpl Drop for BrokerConn {\r\n    fn drop(&mut self) {\r\n        // Only a healthy physical close is counted instead of logged per conn.\r\n        // Error/timeout closes keep their full attribution and original start.\r\n        // [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\r\n        if self.inner.poisoned.load(Ordering::Acquire) {\r\n            self.inner.lifecycle(\r\n                \"transport-close\",\r\n                &format!(\r\n                    \"poisoned=true timed_out={}\",\r\n                    self.inner.timed_out.load(Ordering::Acquire)\r\n                ),\r\n            );\r\n        } else {\r\n            self.inner.lifecycle_log.healthy(false);\r\n        }\r\n        {\r\n            let mut d = recover(&self.inner.dog);\r\n            d.shutdown = true;\r\n            self.inner.dog_cv.notify_all();\r\n        }\r\n        if let Some(dog) = self.dog.take() {\r\n            let _ = dog.join();\r\n        }\r\n        // Halves (gate or graveyard) drop AFTER the watchdog is gone — no raw\r\n        // handle use can race the CloseHandle/limbo path.\r\n    }\r\n}\r\n\r\nfn poisoned_err() -> io::Error {\r\n    io::Error::new(\r\n        io::ErrorKind::BrokenPipe,\r\n        \"broker conn poisoned (bounded write retired the physical connection)\",\r\n    )\r\n}\r\n\r\nfn timeout_err() -> io::Error {\r\n    io::Error::new(\r\n        io::ErrorKind::TimedOut,\r\n        \"broker conn write exceeded its bound (conn poisoned + aborted)\",\r\n    )\r\n}\r\n\r\n#[cfg(test)]\r\nmod tests {\r\n    use super::*;\r\n    use crate::transport::{DaemonTransport, LocalSocketTransport};\r\n    use interprocess::local_socket::traits::Stream as _;\r\n    use interprocess::local_socket::Stream;\r\n    use std::sync::atomic::AtomicU32;\r\n\r\n    /// A real (loopback local-socket) conn pair — the send half wrapped in the\r\n    /// primitive, the client end returned so a test can hold it alive or drop\r","truncated":true,"truncatedBy":"lines","totalLines":581,"totalBytes":14984,"outputLines":379,"outputBytes":14984,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":1000,"displayContent":{"text":"impl Inner {\n…\n    /// trailing re-cancel covers exactly that slipped-in read. On Unix the\r\n    /// `shutdown` in [`poison_and_cancel`] already did all of this.\r\n    // On Unix the cfg(windows) teardown strips out and the idempotence\r\n    // guard's early `return` becomes the function tail — needless_return\r\n    // there, load-bearing here. Restructuring (e.g. `if !swapped { … }`)\r\n    // would leave Unix an empty if-body (`clippy::needless_if`) instead.\r\n    #[cfg_attr(unix, allow(clippy::needless_return))]\r\n    fn sever(&self) {\r\n        if self.severed.swap(true, Ordering::AcqRel) {\r\n            return;\r\n        }\r\n        #[cfg(windows)]\r\n        // SAFETY: as in poison_and_cancel — handle outlives the watchdog.\r\n        unsafe {\r\n            let _ = win::DisconnectNamedPipe(self.raw);\r\n            let _ = win::CancelIoEx(self.raw, core::ptr::null_mut());\r\n        }\r\n    }\r\n\r\n    /// The watchdog body: park until a write is in flight, then sleep to its\r\n    /// absolute deadline. If the op is still in flight at the deadline,\r\n    /// poison + cancel, re-canceling periodically until the op reports\r\n    /// completion (the writer clears `inflight` the instant `write_frame`\r\n    /// returns — that return IS the buffer-release handshake), then sever.\r\n    fn watchdog(&self) {\r\n        let mut d = recover(&self.dog);\r\n        loop {\r\n            // Park until armed (or shutdown).\r\n            let (op, deadline) = loop {\r\n                if d.shutdown {\r\n                    return;\r\n                }\r\n                if let Some(armed) = d.inflight {\r\n                    break armed;\r\n                }\r\n                d = self\r\n                    .dog_cv\r\n                    .wait(d)\r\n                    .unwrap_or_else(|p| p.into_inner());\r\n            };\r\n            // Sleep toward the deadline while THIS op stays in flight.\r\n            let fired = loop {\r\n                if d.shutdown {\r\n                    return;\r\n                }\r\n                match d.inflight {\r\n                    Some((cur, _)) if cur == op => {}\r\n                    // Completed (or a newer op armed) before the deadline.\r\n                    _ => break false,\r\n                }\r\n                let now = Instant::now();\r\n                if now >= deadline {\r\n                    break true;\r\n                }\r\n                let (g, _) = self\r\n                    .dog_cv\r\n                    .wait_timeout(d, deadline - now)\r\n                    .unwrap_or_else(|p| p.into_inner());\r\n                d = g;\r\n            };\r\n            if !fired {\r\n                continue;\r\n            }\r\n            // Deadline: abort out of band, then wait for the canceled op to\r\n            // report completion before severing (re-cancel each slice — a\r\n            // cancel that raced the WriteFileEx issuance reaches nothing, so\r\n            // one shot is not enough). Latch timed_out FIRST so the aborted\r\n            // writer observes its failure as a DEADLINE retirement (wedge),\r\n            // not an organic conn death.\r\n            drop(d);\r\n            self.timed_out.store(true, Ordering::Release);\r\n            self.poison_and_cancel(None);\r\n            d = recover(&self.dog);\r\n            while matches!(d.inflight, Some((cur, _)) if cur == op) && !d.shutdown {\r\n                let (g, _) = self\r\n                    .dog_cv\r\n                    .wait_timeout(d, Duration::from_millis(100))\r\n                    .unwrap_or_else(|p| p.into_inner());\r\n                d = g;\r\n                drop(d);\r\n                self.poison_and_cancel(None);\r\n                d = recover(&self.dog);\r\n            }\r\n            if d.shutdown {\r\n                return;\r\n            }\r\n            drop(d);\r\n            self.sever();\r\n            d = recover(&self.dog);\r\n        }\r\n    }\r\n}\r\n\r\n/// The shared, serialized, BOUNDED send side of one physical broker\r\n/// connection — what [`crate::broker::SharedSend`] now aliases. Owns the\r\n/// [`SendHalf`] (it never leaves this module) and the per-conn watchdog.\r\npub(crate) struct BrokerConn {\r\n    inner: Arc<Inner>,\r\n    /// Joined on drop — after `shutdown` is raised — so the raw handle the\r\n    /// watchdog aborts through strictly outlives every raw use.\r\n    dog: Option<JoinHandle<()>>,\r\n}\r\n\r\nimpl BrokerConn {\r\n    /// Wrap a freshly-split send half. `bound` is the ABSOLUTE per-write\r\n    /// deadline (gate-wait + OS write completion) — the broker passes its\r\n    /// `brain_write_deadline()` (`SPT_BRAIN_WRITE_DEADLINE_MS` knob).\r\n    pub(crate) fn new(half: SendHalf, bound: Duration) -> Self {\r\n        static LOG: LazyLock<Arc<LifecycleLog>> = LazyLock::new(|| Arc::new(LifecycleLog::new()));\r\n        Self::with_log(half, bound, Arc::clone(&LOG))\r\n    }\r\n\r\n    fn with_log(half: SendHalf, bound: Duration, lifecycle_log: Arc<LifecycleLog>) -> Self {\r\n        let raw = raw_of(&half);\r\n        let inner = Arc::new(Inner {\r\n            id: CONN_ID_SEQ.fetch_add(1, Ordering::Relaxed) + 1,\r\n            label: Mutex::new(String::new()),\r\n            first_write: OnceLock::new(),\r\n            lifecycle_log,\r\n            gate: Mutex::new(Gate { half: Some(half) }),\r\n            gate_cv: Condvar::new(),\r\n            dog: Mutex::new(Dog {\r\n                inflight: None,\r\n                shutdown: false,\r\n            }),\r\n            dog_cv: Condvar::new(),\r\n            poisoned: AtomicBool::new(false),\r\n            timed_out: AtomicBool::new(false),\r\n            severed: AtomicBool::new(false),\r\n            graveyard: Mutex::new(None),\r\n            op_seq: AtomicU64::new(0),\r\n            bound,\r\n            raw,\r\n        });\r\n        let dog_inner = Arc::clone(&inner);\r\n        let dog = thread::Builder::new()\r\n            .name(\"conn-watchdog\".into())\r\n            .spawn(move || dog_inner.watchdog())\r\n            .ok();\r\n        BrokerConn { inner, dog }\r\n    }\r\n\r\n    /// The stable per-physical-conn id — what lifecycle records key on\r\n    /// (`Arc::ptr_eq` identity does not survive a log line).\r\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n    pub(crate) fn id(&self) -> u64 {\r\n        self.inner.id\r\n    }\r\n\r\n    /// Accrete an attribution fact onto this conn's label (subscriber role,\r\n    /// endpoint/session where known — e.g. `role=brain`,\r\n    /// `controller session=7 by=cli@host`). Facts accumulate space-separated\r\n    /// and ride EVERY subsequent lifecycle record for this conn; growth is\r\n    /// bounded by [`LABEL_CAP`] (one `…` marker, then no-op).\r\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n    pub(crate) fn describe(&self, fact: &str) {\r\n        let mut l = recover(&self.inner.label);\r\n        if l.ends_with('…') {\r\n            return;\r\n        }\r\n        if l.len() + fact.len() + 1 > LABEL_CAP {\r\n            l.push('…');\r\n            return;\r\n        }\r\n        if !l.is_empty() {\r\n            l.push(' ');\r\n        }\r\n        l.push_str(fact);\r\n    }\r\n\r\n    /// Emit a bounded per-conn lifecycle record through this conn's shared\r\n    /// attribution format — the broker's sink machinery uses it for\r\n    /// `writer-exit`, attach/replace, and detach events (leg d).\r\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n    pub(crate) fn lifecycle_event(&self, event: &str, extra: &str) {\r\n        self.inner.lifecycle(event, extra);\r\n    }\r\n\r\n    /// Bounded framed write. See [`Self::write_gated`].\r\n    pub(crate) fn write(&self, env: &Envelope) -> io::Result<ConnWrite> {\r\n        self.write_gated(env, || true)\r\n    }\r\n\r\n    /// Bounded framed write with a pre-write gate closure, run AFTER the\r\n    /// serialized gate is acquired and BEFORE any byte is written — the same\r\n    /// serialization point the old epoch-check-under-`send.lock()` had, so a\r\n    /// superseded controller writer still bails atomically with respect to\r\n    /// its successor's writes (P1c, REQ-HAZARD-CONTROLLER-WRITER-REORDER).\r\n    /// `pre` returning false → `Ok(ConnWrite::Superseded)`, nothing written.\r\n    pub(crate) fn write_gated(\r\n        &self,\r\n        env: &Envelope,\r\n        pre: impl FnOnce() -> bool,\r\n    ) -> io::Result<ConnWrite> {\r\n        let inner = &*self.inner;\r\n        let deadline = Instant::now() + inner.bound;\r\n\r\n        // ── Acquire the gate, bounded by the SAME absolute deadline. ──\r\n        let mut g = recover(&inner.gate);\r\n        loop {\r\n            if inner.poisoned.load(Ordering::Acquire) {\r\n                return Err(poisoned_err());\r\n            }\r\n            if g.half.is_some() {\r\n                break;\r\n            }\r\n            let now = Instant::now();\r\n            if now >= deadline {\r\n                // The write we queued behind is stuck past OUR whole bound:\r\n                // that op is the wedge. Abort it out of band and fail — the\r\n                // watchdog's own deadline (armed at that op's entry) is also\r\n                // in flight; both paths converge on the same idempotent\r\n                // poison/cancel.\r\n                drop(g);\r\n                inner.timed_out.store(true, Ordering::Release);\r\n                inner.poison_and_cancel(None);\r\n                return Err(timeout_err());\r\n            }\r\n            let (g2, _) = inner\r\n                .gate_cv\r\n                .wait_timeout(g, deadline - now)\r\n                .unwrap_or_else(|p| p.into_inner());\r\n            g = g2;\r\n        }\r\n\r\n        if !pre() {\r\n            return Ok(ConnWrite::Superseded);\r\n        }\r\n\r\n        // The gate serializes first-write initialization. Save the instant\r\n        // BEFORE I/O so a later failure still names when this conn went active.\r\n        // [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\r\n        if inner.first_write.get().is_none() {\r\n            let _ = inner.first_write.set(LogTime::now());\r\n            inner.lifecycle_log.healthy(true);\r\n        }\r\n\r\n        // ── Check out the half + arm the watchdog, then write OUT of locks. ──\r\n        let mut half = g.half.take().expect(\"gate half present at checkout\");\r\n        let op = inner.op_seq.fetch_add(1, Ordering::Relaxed) + 1;\r\n        {\r\n            let mut d = recover(&inner.dog);\r\n            d.inflight = Some((op, deadline));\r\n            inner.dog_cv.notify_all();\r\n        }\r\n        drop(g);\r\n\r\n        let res = write_frame(&mut half, env);\r\n\r\n        // Completion handshake: `write_frame` has returned, so the OS op —\r\n        // canceled or not — reported completion and the frame buffer is\r\n        // released. Clear the arm BEFORE deciding health, so the watchdog\r\n        // never severs while the op is formally in flight.\r\n        {\r\n            let mut d = recover(&inner.dog);\r\n            d.inflight = None;\r\n            inner.dog_cv.notify_all();\r\n        }\r\n\r\n        let mut g = recover(&inner.gate);\r\n        if res.is_ok() && !inner.poisoned.load(Ordering::Acquire) {\r\n            g.half = Some(half);\r\n            drop(g);\r\n            inner.gate_cv.notify_all();\r\n            return Ok(ConnWrite::Done);\r\n        }\r\n\r\n        // ── Retire: poison the whole conn, abort the read side so\r\n        // `handle_conn` reaches EOF cleanup, sever, and park the half in the\r\n        // graveyard (its handle must outlive any late raw abort; never\r\n        // re-gated, never reused — a timed-out length-prefixed frame may be\r\n        // partially written). ──\r\n        drop(g);\r\n        inner.poison_and_cancel(res.as_ref().err());\r\n        inner.sever();\r\n        *recover(&inner.graveyard) = Some(half);\r\n        inner.gate_cv.notify_all();\r\n        // A DEADLINE retirement (the watchdog aborted OUR op, or the conn was\r\n        // already deadline-poisoned) surfaces as TimedOut — the caller can\r\n        // count the wedge (W2 stall-evict tally). An organic I/O failure\r\n        // (peer death) keeps its own error.\r\n        if inner.timed_out.load(Ordering::Acquire) {\r\n            return Err(timeout_err());\r\n        }\r\n        Err(match res {\r\n            Err(e) => e,\r\n            Ok(()) => poisoned_err(),\r\n        })\r\n    }\r\n}\r\n\r\n/// TEST-ONLY determinism lever: holds the serialized write gate (the half is\r\n/// checked out) so every writer parks AT THE GATE until the pin drops — the\r\n/// round-4 analog of pinning the old `Mutex<SendHalf>` across a registration\r\n/// window (the P1c epoch-gate keystone unit relies on it).\r\n#[cfg(test)]\r\npub(crate) struct TestGatePin {\r\n    inner: Arc<Inner>,\r\n    half: Option<SendHalf>,\r\n}\r\n\r\n#[cfg(test)]\r\nimpl Drop for TestGatePin {\r\n    fn drop(&mut self) {\r\n        let mut g = recover(&self.inner.gate);\r\n        g.half = self.half.take();\r\n        drop(g);\r\n        self.inner.gate_cv.notify_all();\r\n    }\r\n}\r\n\r\n#[cfg(test)]\r\nimpl BrokerConn {\r\n    /// Check the half out of the gate without arming the watchdog — writers\r\n    /// queue behind the gate (bounded by their own deadlines) until the\r\n    /// returned pin drops. Test-only; never a production path.\r\n    pub(crate) fn pin_gate_for_test(&self) -> TestGatePin {\r\n        let mut g = recover(&self.inner.gate);\r\n        let half = g.half.take().expect(\"pin: gate half present\");\r\n        TestGatePin {\r\n            inner: Arc::clone(&self.inner),\r\n            half: Some(half),\r\n        }\r\n    }\r\n}\r\n\r\nimpl Drop for BrokerConn {\r\n    fn drop(&mut self) {\r\n        // Only a healthy physical close is counted instead of logged per conn.\r\n        // Error/timeout closes keep their full attribution and original start.\r\n        // [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\r\n        if self.inner.poisoned.load(Ordering::Acquire) {\r\n            self.inner.lifecycle(\r\n                \"transport-close\",\r\n                &format!(\r\n                    \"poisoned=true timed_out={}\",\r\n                    self.inner.timed_out.load(Ordering::Acquire)\r\n                ),\r\n            );\r\n        } else {\r\n            self.inner.lifecycle_log.healthy(false);\r\n        }\r\n        {\r\n            let mut d = recover(&self.inner.dog);\r\n            d.shutdown = true;\r\n            self.inner.dog_cv.notify_all();\r\n        }\r\n        if let Some(dog) = self.dog.take() {\r\n            let _ = dog.join();\r\n        }\r\n        // Halves (gate or graveyard) drop AFTER the watchdog is gone — no raw\r\n        // handle use can race the CloseHandle/limbo path.\r\n    }\r\n}\r\n\r\nfn poisoned_err() -> io::Error {\r\n    io::Error::new(\r\n        io::ErrorKind::BrokenPipe,\r\n        \"broker conn poisoned (bounded write retired the physical connection)\",\r\n    )\r\n}\r\n\r\nfn timeout_err() -> io::Error {\r\n    io::Error::new(\r\n        io::ErrorKind::TimedOut,\r\n        \"broker conn write exceeded its bound (conn poisoned + aborted)\",\r\n    )\r\n}\r\n\r\n#[cfg(test)]\r\nmod tests {\r\n    use super::*;\r\n    use crate::transport::{DaemonTransport, LocalSocketTransport};\r\n    use interprocess::local_socket::traits::Stream as _;\r\n    use interprocess::local_socket::Stream;\r\n    use std::sync::atomic::AtomicU32;\r\n\r\n    /// A real (loopback local-socket) conn pair — the send half wrapped in the\r\n    /// primitive, the client end returned so a test can hold it alive or drop\r\n…\n}","startLine":315,"lineNumbers":[315,null,420,421,422,423,424,425,426,427,428,429,430,431,432,433,434,435,436,437,438,439,440,441,442,443,444,445,446,447,448,449,450,451,452,453,454,455,456,457,458,459,460,461,462,463,464,465,466,467,468,469,470,471,472,473,474,475,476,477,478,479,480,481,482,483,484,485,486,487,488,489,490,491,492,493,494,495,496,497,498,499,500,501,502,503,504,505,506,507,508,509,510,511,512,513,514,515,516,517,518,519,520,521,522,523,524,525,526,527,528,529,530,531,532,533,534,535,536,537,538,539,540,541,542,543,544,545,546,547,548,549,550,551,552,553,554,555,556,557,558,559,560,561,562,563,564,565,566,567,568,569,570,571,572,573,574,575,576,577,578,579,580,581,582,583,584,585,586,587,588,589,590,591,592,593,594,595,596,597,598,599,600,601,602,603,604,605,606,607,608,609,610,611,612,613,614,615,616,617,618,619,620,621,622,623,624,625,626,627,628,629,630,631,632,633,634,635,636,637,638,639,640,641,642,643,644,645,646,647,648,649,650,651,652,653,654,655,656,657,658,659,660,661,662,663,664,665,666,667,668,669,670,671,672,673,674,675,676,677,678,679,680,681,682,683,684,685,686,687,688,689,690,691,692,693,694,695,696,697,698,699,700,701,702,703,704,705,706,707,708,709,710,711,712,713,714,715,716,717,718,719,720,721,722,723,724,725,726,727,728,729,730,731,732,733,734,735,736,737,738,739,740,741,742,743,744,745,746,747,748,749,750,751,752,753,754,755,756,757,758,759,760,761,762,763,764,765,766,767,768,769,770,771,772,773,774,775,776,777,778,779,780,781,782,783,784,785,786,787,788,789,790,791,792,793,794,795,796,797,798,null,1000]},"fileSize":43321,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-daemon\\src\\conn.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":1000,"totalBytes":14984,"outputLines":379,"outputBytes":14984,"shownRange":{"start":420,"end":798},"nextOffset":799}}},"isError":false,"timestamp":1788999516699}}
{"type":"message","id":"3cec2015","parentId":"8d198e16","timestamp":"2026-09-10T00:18:36.701Z","message":{"role":"toolResult","toolCallId":"call_yYGvYsXGB3t7hhjhu2hBMYjN|fc_0dde72effda0ad21016aa1f75c942087d2b1f7df67820bc67e","toolName":"grep","content":[{"type":"text","text":"*110|### 3.3 Orphan teardown fires echo-commune BEFORE INIT_SIGNOFF\n 111|- **Failure:** teardown delivers INIT_SIGNOFF without first saving the final context delta → Psyche signoff lacks the context-save summary.\n 112|- **Invariant:** on orphan path, synchronously run the echo-commune (final delta) before composing INIT_SIGNOFF; skip only if the session_id is missing.\n 113|- **spt-core mapping:** the daemon runs psyche/pulse loops in-process; ordering invariant identical.\n...\n 115|\n*116|### 3.4 A `ring` never adopts — so never deletes — a perch it did not create\n 117|- **Failure:** 3.1 guards the LEAK direction (cleanup on every exit path). It says nothing about *whose* perch is being cleaned, and the two compose into a data-loss bug: `ring` decides \"the caller has no perch\" **ready-marker-first**, so an existing perch whose marker is momentarily down — busy turn, soft session-end, stale re-bind — reads *perchless*, `setup_ephemeral` clobbers the caller's real `info.json`, and 3.1's guaranteed cleanup then runs on the way out and **deletes** the ready marker, `in...\n 118|- **The mail half is a CONFIDENTIALITY failure, not only a durability one:** deletion is what happens on the way out, but while the ring *holds* the adopted perch its reply-wait `drain_one_at` consumes whatever is in the victim's spool and **renders it to the ringer as the reply**. Ringer == victim (the field case) is the degenerate, mild shape — anything drained reaches its intended recipient. Ringer != victim is the severe one: a third party rings a live agent and the **victim's inbox is drained into...\n 119|- **Self-camouflaging**, which is why it went unattributed for hours: the damage — perch gone, roster row ghosted — *mimics the stale-liveness condition people reach for `ring` to diagnose*. The field caller was probing a ghost-roster symptom and the probe manufactured a fresh one, so debuggers of stale liveness are disproportionately its victims and read the wreckage as more of the symptom they were chasing.\n...\n 126|\n*127|### 3.5 `is_perch_alive` is INVERTED on perch residue — a GC must never key on it\n 128|<!-- [doc->REQ-HAZARD-PERCH-GC-LIVENESS-INVERSION] -->\n 129|<!-- [doc->REQ-PERCH-GC-RESIDUE-PREDICATE] -->\n 130|- **Failure:** the obvious way to write a perch garbage collector is \"reap what is not alive\" — `!spt_store::liveness::is_perch_alive(dir)`. On the population a GC actually walks, that predicate is **backwards**. `is_perch_alive` returns **true** for a directory with no `info.json` (`RawRead::Absent ⇒ true`, interim parity — an absent record means a wrapper-owned listener between polls), and **false** for a daemon-hosted endpoint whose `status` is `offline`. Measured on HFENDULEAM 2026-08-04 across...\n...\n 139|\n*140|### 4.1 Envelope HTML-entity codec ordering — `&amp;` decoded LAST\n 141|- **Failure:** decoding `&amp;`-entity before the others double-decodes nested entities (`&amp;amp;lt;` → wrong result).\n 142|- **Invariant:** ENCODE order amp→first … `<br>`→last; DECODE order `<br>`→first … amp→**last** (`&lt;`,`&gt;`,`&quot;`, then `&amp;`). One sole decode site (at the LLM/stdin boundary); the parser never decodes.\n 143|- **spt-core mapping:** `spt-proto` owns the envelope grammar (public SDK, semver + wire-version). This codec contract is a copy-verbatim commodity item (ADR-0001) and a public-API conformance test.\n...\n 146|\n*147|### 4.2 Two-slice envelope parser is panic-free and tolerant\n 148|- **Failure:** malformed envelope (unclosed/misordered/nested tags) panics or drops output.\n 149|- **Invariant:** tags case-sensitive, all optional; no tags → whole body to live slot; unclosed → None for that tag; out-of-order → both still extracted; nested unknown tags preserved verbatim; zero `unwrap` on parsed text.\n 150|- **spt-core mapping:** `spt-proto` parser; property-test the robustness rules.\n...\n 152|\n*153|### 4.3 Registry stale-entry cleanup precedes lookup\n 154|- **Failure:** sender resolves a dead process's stale TCP port → delivery to wrong/dead listener.\n 155|- **Invariant:** clean stale entries (dead PID) before/at lookup; spool fallback is the safe path on TCP miss.\n 156|- **spt-core mapping:** now spans the **subnet registry** (ADR-0003) — eventually-consistent across nodes. Cross-node staleness is expected; resolution policy (local → most-recent → `id@node`) must degrade to spool/relay fallback on stale hits, and never hard-fail on a stale remote entry.\n...\n 158|\n*159|### 4.4 Deferred rows survive poll drain\n 160|- **Failure:** poll `drain_all` flushes a deferred (spool-only) message meant for a hook consumer → message lost.\n 161|- **Invariant:** deferred rows are never flushed by the event-stream drain; only `drain_non_deferred_*` / `peek_all` touch them.\n 162|- **Sister cite:** CHANGELOG v1.11.20; `src/common/spool.rs`. (See also 1.4.)\n 163|\n*164|### 4.5 Inbox legacy compat must not double-deliver\n 165|- **Failure:** message surfaced via both spool (durable) and legacy inbox files → duplicate or racing delivery.\n 166|- **Invariant:** spool is the sole read path at poll time; inbox is write-for-compat only and never read.\n 167|- **spt-core mapping:** clean-room — likely drop the legacy inbox entirely. If kept for any compat, preserve \"never read at drain time.\"\n...\n 169|\n*170|### 4.6 Addressable-id charset reserves the address delimiters\n 171|<!-- [doc->REQ-HAZARD-ID-CHARSET] -->\n 172|- **Failure:** a bare endpoint id that contains `:` or `@` (or a path separator / whitespace / control char) makes the canonical qualified address `[subnet:]id[@node]` (ADR-0006 / REQ-INST-10) ambiguous to parse, and lets a name smuggle into a perch directory path. Once permissive ids exist in the wild, tightening later needs a migration.\n 173|- **Invariant:** every addressable id/name is validated to `[A-Za-z0-9_-]` + Hiragana/Katakana/CJK only, length `1..=64`, **at every creation seam** (`ready` start, `api bind`, `api listen`, `api worker-start`). `:` and `@` are permanently reserved as address delimiters; reads of existing perches are never re-validated. Enforce now (pre-M3/M4) so no permissive id-data accumulates.\n...\n 175|\n*176|### 4.7 Concurrent SQLite openers must not fail with \"database is locked\"\n 177|<!-- [doc->REQ-HAZARD-REGISTRY-CONCURRENT] -->\n 178|- **Failure:** two endpoints on one machine open the same SQLite store at once (e.g. two `ReadyAgent::start` calls registering simultaneously) and one fails outright with `SQLITE_BUSY` / \"database is locked\" → spurious registration/spool failure. Surfaced as a parallel-test flake in `two_agents_exchange_message_tcp_and_spool`, but the bug is real concurrency, not test-only.\n*179|- **Invariant:** `busy_timeout` is set **before** any lock-taking statement on every connection. Switching `journal_mode=WAL` takes a brief exclusive lock; with the default 0ms timeout it fails immediately under contention, so the pragma order is load-bearing: `Connection::open` → `busy_timeout` → `journal_mode=WAL` → `CREATE TABLE …`. WAL alone is insufficient (concurrent *writers* still serialize; they must *wait*, not error).\n 180|- **spt-core mapping:** `spt_store::registry::open_registry` + `spt_store::spool::open_spool_at`; both set `busy_timeout=5000` first. Any future SQLite store (history Path B, instance registry) must follow the same ordering.\n 181|\n*182|### 4.8 Registry merge ordered by epoch, never wall-clock (red-team #8)\n 183|<!-- [doc->REQ-HAZARD-REGISTRY-EPOCH-LEASE] -->\n 184|- **Failure:** the per-subnet registry replicates `endpoint_id → [instances]` eventually-consistently across nodes. Under a partition or clock skew, a lagging node re-announces a stale `Active` for an endpoint that has actually gone `Offline`. If the merge ordered updates by wall-clock (or \"last write wins\"), the stale `Active` overwrites the newer `Offline` and resolution routes a message to a dead/wrong instance.\n 185|- **Invariant:** the merge precedence key is a **per-node monotonic epoch counter** (`spt_store::epoch::EpochSource`, persisted, strictly increasing, NEVER wall-clock), compared version-vector style per `(endpoint_id, node)`: an incoming update wins **iff its epoch is strictly greater** than the stored one for that node; equal or lower is dropped as stale. So a newer `Offline` (higher epoch) can never be clobbered by a lagging `Active` (lower epoch), and an idempotent equal-epoch replay is a no-op. Wall-...\n...\n 187|\n*188|### 4.9 SQLite stores must create their parent dir — SQLite won't\n 189|<!-- [doc->REQ-HAZARD-REGISTRY-DIR-CREATE] -->\n*190|- **Failure:** `Connection::open` creates the database FILE but never its parent DIRECTORY. On a fresh home (first boot, fresh CI `_work` dir) a registry op that runs before any perch-creating op (`create_dir_all` side effects) fails `SQLITE_CANTOPEN` — \"unable to open database file …owlery\\.registry\". Timing-dependent: whichever code path touches the home first decides the outcome, so it surfaces as a parallel-test flake (bind-first tests losing the dir-creation race to perch-first tests). Bit the h...\n*191|- **Invariant:** every SQLite store's open path `create_dir_all`s its parent dir itself, best-effort, before `Connection::open` — never relying on another subsystem having materialized the home first. (Mirrors the spool, which always did this; the registry didn't.)\n 192|- **spt-core mapping:** `spt_store::registry::open_registry` (`create_dir_all(owlery)` before open). `spt_store::spool::open_spool_at` already creates its perch dir. Any future SQLite store must do the same — pair this with the 4.7 pragma ordering on every new store.\n 193|\n*194|### 4.10 Dead node identities leave immortal registry rows  `[REQ-HAZARD-REGISTRY-GHOST-ROWS]`\n 195|<!-- [doc->REQ-HAZARD-REGISTRY-GHOST-ROWS] -->\n 196|- **Failure:** the registry's only superseding mechanism is the per-`(endpoint_id, node)` epoch lease (4.8) — a row is replaced only by a newer row *from the same node*. When a node identity dies permanently (machine retired, or `node.key` regenerated so the \"node\" never speaks again), its rows are never superseded and never expire: they sit in the in-memory registries and the `identity/registry/<subnet>.json` snapshots forever. A bare-id send then resolves the same endpoint id on both the live and the...\n*197|- **Invariant:** registry rows authored by a **silent** peer node decay: a node not *heard* (admitted inbound feed — the M7 D2 heard-map, REQ-SUBNET-1) within the eviction window (`registry_evict_after_ms`, default 300s ≈ 10 default pump cadences) has its rows **evicted** from every subnet registry, snapshots rewritten. Own rows never decay (the node always hears itself implicitly — it authors them each pump tick). Eviction is safe under the lease: v1 has **no transitive gossip**, so any future upd...\n 198|- **spt-core mapping:** `spt_net::net::registry::SubnetRegistry::evict_nodes` (model) + `spt_daemon::registryhost::RegistryHost::evict_silent_peers` (heard-map TTL) driven from the registry pump tick (`peerloop`). Trust rows are NOT auto-evicted (trust is a user decision; a stale trust row only costs dead dials) — pruning those is a separate verb.\n 199|- **Source:** M7 acceptance run 2026-06-06 (DEFERRED.md \"Ghost registry row eviction\"); the AMBIGUOUS render fix rode along.\n 200|- **Mesh note (ADR-0017, 2026-06-08):** the subnet mesh **preserves** this invariant rather than superseding it. \"No transitive gossip\" sharpens to **no transitive *row* gossip** — the mesh relays only the member *roster* (discovery), while registry **rows stay own-authored and are fetched directly** from each member over a handshake. So \"any future update for a node comes from that node itself, alive\" still holds and the eviction lease is untouched. (The plan's rejected alternative — signed transiti...\n 201|\n*202|### 4.11 Advertisement-epoch reset strands a node  `[REQ-HAZARD-EPOCH-RESET]`\n 203|<!-- [doc->REQ-HAZARD-EPOCH-RESET] -->\n 204|- **Failure:** a node whose advertisement-epoch counter resets (the durable `EpochSource` file lost/recreated) re-advertises with LOW epochs; peers hold a higher last-seen epoch for that `(endpoint, node)` lease and drop every fresh row as **stale** — the node advertises into a void until its counter outruns its own history. Nothing renders the cause: the node looks healthy locally, peers simply never update.\n 205|- **Invariant (mitigation by construction, common case):** the common trigger — a full reinstall / identity regeneration — is covered by the **re-pair trust overwrite** (M8 decision 13, REQ-SUBNET-7): a completed ceremony presenting the same label + machine id evicts the superseded identity's trust AND registry rows on the seed-holder, and the peer-side epoch memory **dies with the deleted row** — the re-paired node's fresh epochs land on a clean lease. M8 acceptance 7 verifies this explicitly (the...\n...\n 213|\n*214|### 5.1 Windows PID recycling false positives\n 215|- **Failure:** recycled PID reads alive for the wrong process → orphan misclassification.\n 216|- **Invariant:** anchor liveness on the stable parent/harness PID; minimal info.json for supervisor-owned perches; mtime grace window (≥60s) masks transient mismatches.\n 217|- **spt-core mapping:** broker-held handles supersede PID polling for spt-hosted sessions; keep the grace window for harness-hosted.\n...\n 219|\n*220|### 5.2 Windows EBUSY on atomic rename\n 221|- **Failure:** `fs::rename` fails while a handle is (recently) held → registry/marketplace update fails.\n 222|- **Invariant:** tmp-write + atomic-rename with retry/backoff; best-effort side-fail; tolerate transient EBUSY.\n 223|- **spt-core mapping:** all on-disk state writes (registry, trust store, spool checkpoints) use this pattern. Self-update binary swap on Windows especially.\n...\n 225|\n*226|### 5.3 Git/subprocess timeout stamping\n 227|- **Failure:** a hung subprocess (git on slow net) blocks the supervisor indefinitely.\n 228|- **Invariant:** every metadata-producing subprocess has a timeout; timeout yields `None` + rate-limited stderr, never a hang.\n\n[Some lines truncated to 512 chars]"}],"details":{"scopePath":".worktrees/286-churn/docs/KNOWN-HAZARDS.md","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\docs\\KNOWN-HAZARDS.md","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":21,"fileCount":1,"files":[".worktrees/286-churn/docs/KNOWN-HAZARDS.md"],"fileMatches":[{"path":".worktrees/286-churn/docs/KNOWN-HAZARDS.md","count":21}],"truncated":true,"displayContent":"*110│### 3.3 Orphan teardown fires echo-commune BEFORE INIT_SIGNOFF\n 111│- **Failure:** teardown delivers INIT_SIGNOFF without first saving the final context delta → Psyche signoff lacks the context-save summary.\n 112│- **Invariant:** on orphan path, synchronously run the echo-commune (final delta) before composing INIT_SIGNOFF; skip only if the session_id is missing.\n 113│- **spt-core mapping:** the daemon runs psyche/pulse loops in-process; ordering invariant identical.\n    │...\n 115│\n*116│### 3.4 A `ring` never adopts — so never deletes — a perch it did not create\n 117│- **Failure:** 3.1 guards the LEAK direction (cleanup on every exit path). It says nothing about *whose* perch is being cleaned, and the two compose into a data-loss bug: `ring` decides \"the caller has no perch\" **ready-marker-first**, so an existing perch whose marker is momentarily down — busy turn, soft session-end, stale re-bind — reads *perchless*, `setup_ephemeral` clobbers the caller's real `info.json`, and 3.1's guaranteed cleanup then runs on the way out and **deletes** the ready marker, `in...\n 118│- **The mail half is a CONFIDENTIALITY failure, not only a durability one:** deletion is what happens on the way out, but while the ring *holds* the adopted perch its reply-wait `drain_one_at` consumes whatever is in the victim's spool and **renders it to the ringer as the reply**. Ringer == victim (the field case) is the degenerate, mild shape — anything drained reaches its intended recipient. Ringer != victim is the severe one: a third party rings a live agent and the **victim's inbox is drained into...\n 119│- **Self-camouflaging**, which is why it went unattributed for hours: the damage — perch gone, roster row ghosted — *mimics the stale-liveness condition people reach for `ring` to diagnose*. The field caller was probing a ghost-roster symptom and the probe manufactured a fresh one, so debuggers of stale liveness are disproportionately its victims and read the wreckage as more of the symptom they were chasing.\n    │...\n 126│\n*127│### 3.5 `is_perch_alive` is INVERTED on perch residue — a GC must never key on it\n 128│<!-- [doc->REQ-HAZARD-PERCH-GC-LIVENESS-INVERSION] -->\n 129│<!-- [doc->REQ-PERCH-GC-RESIDUE-PREDICATE] -->\n 130│- **Failure:** the obvious way to write a perch garbage collector is \"reap what is not alive\" — `!spt_store::liveness::is_perch_alive(dir)`. On the population a GC actually walks, that predicate is **backwards**. `is_perch_alive` returns **true** for a directory with no `info.json` (`RawRead::Absent ⇒ true`, interim parity — an absent record means a wrapper-owned listener between polls), and **false** for a daemon-hosted endpoint whose `status` is `offline`. Measured on HFENDULEAM 2026-08-04 across...\n    │...\n 139│\n*140│### 4.1 Envelope HTML-entity codec ordering — `&amp;` decoded LAST\n 141│- **Failure:** decoding `&amp;`-entity before the others double-decodes nested entities (`&amp;amp;lt;` → wrong result).\n 142│- **Invariant:** ENCODE order amp→first … `<br>`→last; DECODE order `<br>`→first … amp→**last** (`&lt;`,`&gt;`,`&quot;`, then `&amp;`). One sole decode site (at the LLM/stdin boundary); the parser never decodes.\n 143│- **spt-core mapping:** `spt-proto` owns the envelope grammar (public SDK, semver + wire-version). This codec contract is a copy-verbatim commodity item (ADR-0001) and a public-API conformance test.\n    │...\n 146│\n*147│### 4.2 Two-slice envelope parser is panic-free and tolerant\n 148│- **Failure:** malformed envelope (unclosed/misordered/nested tags) panics or drops output.\n 149│- **Invariant:** tags case-sensitive, all optional; no tags → whole body to live slot; unclosed → None for that tag; out-of-order → both still extracted; nested unknown tags preserved verbatim; zero `unwrap` on parsed text.\n 150│- **spt-core mapping:** `spt-proto` parser; property-test the robustness rules.\n    │...\n 152│\n*153│### 4.3 Registry stale-entry cleanup precedes lookup\n 154│- **Failure:** sender resolves a dead process's stale TCP port → delivery to wrong/dead listener.\n 155│- **Invariant:** clean stale entries (dead PID) before/at lookup; spool fallback is the safe path on TCP miss.\n 156│- **spt-core mapping:** now spans the **subnet registry** (ADR-0003) — eventually-consistent across nodes. Cross-node staleness is expected; resolution policy (local → most-recent → `id@node`) must degrade to spool/relay fallback on stale hits, and never hard-fail on a stale remote entry.\n    │...\n 158│\n*159│### 4.4 Deferred rows survive poll drain\n 160│- **Failure:** poll `drain_all` flushes a deferred (spool-only) message meant for a hook consumer → message lost.\n 161│- **Invariant:** deferred rows are never flushed by the event-stream drain; only `drain_non_deferred_*` / `peek_all` touch them.\n 162│- **Sister cite:** CHANGELOG v1.11.20; `src/common/spool.rs`. (See also 1.4.)\n 163│\n*164│### 4.5 Inbox legacy compat must not double-deliver\n 165│- **Failure:** message surfaced via both spool (durable) and legacy inbox files → duplicate or racing delivery.\n 166│- **Invariant:** spool is the sole read path at poll time; inbox is write-for-compat only and never read.\n 167│- **spt-core mapping:** clean-room — likely drop the legacy inbox entirely. If kept for any compat, preserve \"never read at drain time.\"\n    │...\n 169│\n*170│### 4.6 Addressable-id charset reserves the address delimiters\n 171│<!-- [doc->REQ-HAZARD-ID-CHARSET] -->\n 172│- **Failure:** a bare endpoint id that contains `:` or `@` (or a path separator / whitespace / control char) makes the canonical qualified address `[subnet:]id[@node]` (ADR-0006 / REQ-INST-10) ambiguous to parse, and lets a name smuggle into a perch directory path. Once permissive ids exist in the wild, tightening later needs a migration.\n 173│- **Invariant:** every addressable id/name is validated to `[A-Za-z0-9_-]` + Hiragana/Katakana/CJK only, length `1..=64`, **at every creation seam** (`ready` start, `api bind`, `api listen`, `api worker-start`). `:` and `@` are permanently reserved as address delimiters; reads of existing perches are never re-validated. Enforce now (pre-M3/M4) so no permissive id-data accumulates.\n    │...\n 175│\n*176│### 4.7 Concurrent SQLite openers must not fail with \"database is locked\"\n 177│<!-- [doc->REQ-HAZARD-REGISTRY-CONCURRENT] -->\n 178│- **Failure:** two endpoints on one machine open the same SQLite store at once (e.g. two `ReadyAgent::start` calls registering simultaneously) and one fails outright with `SQLITE_BUSY` / \"database is locked\" → spurious registration/spool failure. Surfaced as a parallel-test flake in `two_agents_exchange_message_tcp_and_spool`, but the bug is real concurrency, not test-only.\n*179│- **Invariant:** `busy_timeout` is set **before** any lock-taking statement on every connection. Switching `journal_mode=WAL` takes a brief exclusive lock; with the default 0ms timeout it fails immediately under contention, so the pragma order is load-bearing: `Connection::open` → `busy_timeout` → `journal_mode=WAL` → `CREATE TABLE …`. WAL alone is insufficient (concurrent *writers* still serialize; they must *wait*, not error).\n 180│- **spt-core mapping:** `spt_store::registry::open_registry` + `spt_store::spool::open_spool_at`; both set `busy_timeout=5000` first. Any future SQLite store (history Path B, instance registry) must follow the same ordering.\n 181│\n*182│### 4.8 Registry merge ordered by epoch, never wall-clock (red-team #8)\n 183│<!-- [doc->REQ-HAZARD-REGISTRY-EPOCH-LEASE] -->\n 184│- **Failure:** the per-subnet registry replicates `endpoint_id → [instances]` eventually-consistently across nodes. Under a partition or clock skew, a lagging node re-announces a stale `Active` for an endpoint that has actually gone `Offline`. If the merge ordered updates by wall-clock (or \"last write wins\"), the stale `Active` overwrites the newer `Offline` and resolution routes a message to a dead/wrong instance.\n 185│- **Invariant:** the merge precedence key is a **per-node monotonic epoch counter** (`spt_store::epoch::EpochSource`, persisted, strictly increasing, NEVER wall-clock), compared version-vector style per `(endpoint_id, node)`: an incoming update wins **iff its epoch is strictly greater** than the stored one for that node; equal or lower is dropped as stale. So a newer `Offline` (higher epoch) can never be clobbered by a lagging `Active` (lower epoch), and an idempotent equal-epoch replay is a no-op. Wall-...\n    │...\n 187│\n*188│### 4.9 SQLite stores must create their parent dir — SQLite won't\n 189│<!-- [doc->REQ-HAZARD-REGISTRY-DIR-CREATE] -->\n*190│- **Failure:** `Connection::open` creates the database FILE but never its parent DIRECTORY. On a fresh home (first boot, fresh CI `_work` dir) a registry op that runs before any perch-creating op (`create_dir_all` side effects) fails `SQLITE_CANTOPEN` — \"unable to open database file …owlery\\.registry\". Timing-dependent: whichever code path touches the home first decides the outcome, so it surfaces as a parallel-test flake (bind-first tests losing the dir-creation race to perch-first tests). Bit the h...\n*191│- **Invariant:** every SQLite store's open path `create_dir_all`s its parent dir itself, best-effort, before `Connection::open` — never relying on another subsystem having materialized the home first. (Mirrors the spool, which always did this; the registry didn't.)\n 192│- **spt-core mapping:** `spt_store::registry::open_registry` (`create_dir_all(owlery)` before open). `spt_store::spool::open_spool_at` already creates its perch dir. Any future SQLite store must do the same — pair this with the 4.7 pragma ordering on every new store.\n 193│\n*194│### 4.10 Dead node identities leave immortal registry rows  `[REQ-HAZARD-REGISTRY-GHOST-ROWS]`\n 195│<!-- [doc->REQ-HAZARD-REGISTRY-GHOST-ROWS] -->\n 196│- **Failure:** the registry's only superseding mechanism is the per-`(endpoint_id, node)` epoch lease (4.8) — a row is replaced only by a newer row *from the same node*. When a node identity dies permanently (machine retired, or `node.key` regenerated so the \"node\" never speaks again), its rows are never superseded and never expire: they sit in the in-memory registries and the `identity/registry/<subnet>.json` snapshots forever. A bare-id send then resolves the same endpoint id on both the live and the...\n*197│- **Invariant:** registry rows authored by a **silent** peer node decay: a node not *heard* (admitted inbound feed — the M7 D2 heard-map, REQ-SUBNET-1) within the eviction window (`registry_evict_after_ms`, default 300s ≈ 10 default pump cadences) has its rows **evicted** from every subnet registry, snapshots rewritten. Own rows never decay (the node always hears itself implicitly — it authors them each pump tick). Eviction is safe under the lease: v1 has **no transitive gossip**, so any future upd...\n 198│- **spt-core mapping:** `spt_net::net::registry::SubnetRegistry::evict_nodes` (model) + `spt_daemon::registryhost::RegistryHost::evict_silent_peers` (heard-map TTL) driven from the registry pump tick (`peerloop`). Trust rows are NOT auto-evicted (trust is a user decision; a stale trust row only costs dead dials) — pruning those is a separate verb.\n 199│- **Source:** M7 acceptance run 2026-06-06 (DEFERRED.md \"Ghost registry row eviction\"); the AMBIGUOUS render fix rode along.\n 200│- **Mesh note (ADR-0017, 2026-06-08):** the subnet mesh **preserves** this invariant rather than superseding it. \"No transitive gossip\" sharpens to **no transitive *row* gossip** — the mesh relays only the member *roster* (discovery), while registry **rows stay own-authored and are fetched directly** from each member over a handshake. So \"any future update for a node comes from that node itself, alive\" still holds and the eviction lease is untouched. (The plan's rejected alternative — signed transiti...\n 201│\n*202│### 4.11 Advertisement-epoch reset strands a node  `[REQ-HAZARD-EPOCH-RESET]`\n 203│<!-- [doc->REQ-HAZARD-EPOCH-RESET] -->\n 204│- **Failure:** a node whose advertisement-epoch counter resets (the durable `EpochSource` file lost/recreated) re-advertises with LOW epochs; peers hold a higher last-seen epoch for that `(endpoint, node)` lease and drop every fresh row as **stale** — the node advertises into a void until its counter outruns its own history. Nothing renders the cause: the node looks healthy locally, peers simply never update.\n 205│- **Invariant (mitigation by construction, common case):** the common trigger — a full reinstall / identity regeneration — is covered by the **re-pair trust overwrite** (M8 decision 13, REQ-SUBNET-7): a completed ceremony presenting the same label + machine id evicts the superseded identity's trust AND registry rows on the seed-holder, and the peer-side epoch memory **dies with the deleted row** — the re-paired node's fresh epochs land on a clean lease. M8 acceptance 7 verifies this explicitly (the...\n    │...\n 213│\n*214│### 5.1 Windows PID recycling false positives\n 215│- **Failure:** recycled PID reads alive for the wrong process → orphan misclassification.\n 216│- **Invariant:** anchor liveness on the stable parent/harness PID; minimal info.json for supervisor-owned perches; mtime grace window (≥60s) masks transient mismatches.\n 217│- **spt-core mapping:** broker-held handles supersede PID polling for spt-hosted sessions; keep the grace window for harness-hosted.\n    │...\n 219│\n*220│### 5.2 Windows EBUSY on atomic rename\n 221│- **Failure:** `fs::rename` fails while a handle is (recently) held → registry/marketplace update fails.\n 222│- **Invariant:** tmp-write + atomic-rename with retry/backoff; best-effort side-fail; tolerate transient EBUSY.\n 223│- **spt-core mapping:** all on-disk state writes (registry, trust store, spool checkpoints) use this pattern. Self-update binary swap on Windows especially.\n    │...\n 225│\n*226│### 5.3 Git/subprocess timeout stamping\n 227│- **Failure:** a hung subprocess (git on slow net) blocks the supervisor indefinitely.\n 228│- **Invariant:** every metadata-producing subprocess has a timeout; timeout yields `None` + rate-limited stderr, never a hang.","linesTruncated":true,"meta":{"limits":{"columnTruncated":{"maxColumn":512}}}},"isError":false,"timestamp":1788999516701}}
{"type":"message","id":"aae83338","parentId":"3cec2015","timestamp":"2026-09-10T00:18:36.713Z","message":{"role":"toolResult","toolCallId":"call_QHcoNmqHT4fmSVZlIqOgQSWl|fc_0dde72effda0ad21016aa1f75c940487d292ef6b5d4c173d75","toolName":"read","content":[{"type":"text","text":"         let name = format!(\n             \"spt-daemon-connattr-{}-{}.sock\",\n@@ -697,13 +818,78 @@ mod tests {\n         let listener = LocalSocketTransport::bind(&name).expect(\"bind\");\n         let client = LocalSocketTransport::connect(&name).expect(\"connect\");\n         let server = listener.accept().expect(\"accept\");\n-        let (_recv, send) = server.split();\n+        let (recv, send) = server.split();\n         (\n-            BrokerConn::new(send, Duration::from_millis(2000)),\n+            BrokerConn::with_log(send, Duration::from_millis(2000), log),\n             client,\n+            recv,\n         )\n     }\n \n+    // [unit->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+    #[test]\n+    fn healthy_round_trips_and_reconnections_emit_one_interval_summary() {\n+        let now = Instant::now();\n+        let log = Arc::new(LifecycleLog::captured(now));\n+        // The real framed-write/drop path, both a persistent carrier and fresh\n+        // per-poll carriers. An isolated clock/sink avoids global test races.\n+        const N: u64 = 32;\n+        for round in 0..N {\n+            let (conn, mut client, mut recv) = make_duplex_with_log(Arc::clone(&log));\n+            conn.describe(\"role=brain\");\n+            for seq in 0..4 {\n+                let request = Envelope::new(\"poll\", serde_json::json!([round, seq]));\n+                write_frame(&mut client, &request).unwrap();\n+                let received = crate::codec::read_frame(&mut recv).unwrap();\n+                assert_eq!(received.payload, request.payload);\n+                assert_eq!(conn.write(&request).unwrap(), ConnWrite::Done);\n+                let reply = crate::codec::read_frame(&mut client).unwrap();\n+                assert_eq!(reply.payload, request.payload);\n+            }\n+            drop(conn);\n+        }\n+        assert!(recover(log.capture.as_ref().unwrap()).is_empty());\n+        *recover(log.clock.as_ref().unwrap()) = now + HEALTHY_LOG_INTERVAL;\n+        let (conn, mut client) = make_conn_with_log(Arc::clone(&log));\n+        conn.describe(\"role=brain\");\n+        conn.write(&Envelope::new(\"poll\", serde_json::json!(N))).unwrap();\n+        assert_eq!(crate::codec::read_frame(&mut client).unwrap().payload, serde_json::json!(N));\n+        drop(conn);\n+        let lines = recover(log.capture.as_ref().unwrap());\n+        assert_eq!(lines.len(), 1, \"healthy volume is per interval, not per conn\");\n+        assert!(lines[0].contains(&format!(\"starts={}\", N + 1)));\n+        assert!(lines[0].contains(&format!(\"closes={N}\")));\n+    }\n+\n+    // [unit->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+    // [unit->REQ-CONN-POISON-ATTRIBUTION]\n+    #[test]\n+    fn failures_and_role_events_bypass_healthy_aggregation_with_original_start() {\n+        let log = Arc::new(LifecycleLog::captured(Instant::now()));\n+        let (conn, mut client) = make_conn_with_log(Arc::clone(&log));\n+        conn.describe(\"role=brain controller session=7\");\n+        conn.write(&Envelope::new(\"poll\", serde_json::json!({}))).unwrap();\n+        crate::codec::read_frame(&mut client).unwrap();\n+        let start = *conn.inner.first_write.get().expect(\"first committed write\");\n+        conn.lifecycle_event(\"writer-exit\", \"reason=write-failed\");\n+        conn.inner.poison_and_cancel(Some(&io::Error::new(io::ErrorKind::BrokenPipe, \"peer gone\")));\n+        // Repeated cancellation must not repeat the retirement record.\n+        conn.inner.poison_and_cancel(None);\n+        let id = conn.id();\n+        drop(conn);\n+        let lines = recover(log.capture.as_ref().unwrap());\n+        assert_eq!(lines.len(), 3, \"role event, retirement and poisoned close survive\");\n+        for line in lines.iter() {\n+            assert!(line.contains(&format!(\"conn={id}\")));\n+            assert!(line.contains(\"role=brain controller session=7\"));\n+            assert!(line.contains(&format!(\"first_write_wall_ms={}\", start.wall_ms)));\n+            assert!(line.contains(&format!(\"first_write_mono_ms={}\", start.mono_ms)));\n+        }\n+        assert!(lines[1].contains(\"CONN_WRITE_RETIRED:\"));\n+        assert!(lines[1].contains(\"BrokenPipe\"));\n+        assert!(lines[2].contains(\"poisoned=true timed_out=false\"));\n+    }\n+\n     // [unit->REQ-CONN-POISON-DIAL-SCOPE] the F-039 leg-(a) token split: the loud\n     // CONN_WRITE_POISONED token is RESERVED for the deadline (timed_out) class;\n     // an organic write failure renders the distinct CONN_WRITE_RETIRED token\ndiff --git a/crates/spt-daemon/src/livehost.rs b/crates/spt-daemon/src/livehost.rs\nindex 63f479d9..151c5e2e 100644\n--- a/crates/spt-daemon/src/livehost.rs\n+++ b/crates/spt-daemon/src/livehost.rs\n@@ -569,7 +569,7 @@ pub fn resume_restart_orphaned_endpoints(\n     registered: &[(AdapterRecord, Manifest)],\n     adapters_dir: &Path,\n ) {\n-    let Some(live) = query_live_session_endpoints() else {\n+    let Some(live) = SessionPoll::default().query(&crate::endpoint::broker_socket_name()) else {\n         return; // broker unreachable — skip (never mass-respawn on a hiccup)\n     };\n     for id in perch::list_self_perch_ids(owlery) {\n@@ -995,20 +995,57 @@ pub fn reconcile_hosted_liveness(owlery: &Path, live_sessions: &BTreeSet<String>\n     offlined\n }\n \n-/// Query the broker for the set of endpoint ids it currently hosts a session for\n-/// (`KIND_SESSIONS`) — the B2 pull signal. `None` when the broker is unreachable\n-/// (the caller then SKIPS the offline pass this tick rather than mass-offlining).\n-fn query_live_session_endpoints() -> Option<BTreeSet<String>> {\n-    let mut brain = Brain::cold_start(&crate::endpoint::broker_socket_name(), now_ms()).ok()?;\n-    let reply = brain.sessions().ok()?;\n-    Some(\n-        reply\n-            .sessions\n-            .into_iter()\n-            .map(|s| s.endpoint)\n-            .filter(|e| !e.is_empty())\n-            .collect(),\n-    )\n+/// The livehost session census is request/reply only, never a subscription.\n+/// Retain its healthy carrier across 5s polls; discard it on EVERY query error\n+/// so a failed/partially consumed reply can never contaminate the next census.\n+/// `None` still means skip reconciliation, not an empty authoritative census.\n+// [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+#[derive(Default)]\n+struct SessionPoll {\n+    brain: Option<Brain>,\n+    retry: bool,\n+}\n+\n+impl SessionPoll {\n+    fn query(&mut self, broker_name: &str) -> Option<BTreeSet<String>> {\n+        if self.brain.is_none() {\n+            let reason = if self.retry { \"previous-query-error\" } else { \"initial-query\" };\n+            match Brain::cold_start(broker_name, now_ms()) {\n+                Ok(brain) => {\n+                    spt_proto::emit_line_err!(\n+                        \"BRAIN_CONN_OPEN: caller=livehost-session-query reason={reason} \\\n+                         policy=reuse-until-query-error {}\",\n+                        crate::conn::log_stamp()\n+                    );\n+                    self.brain = Some(brain);\n+                }\n+                Err(e) => {\n+                    spt_proto::emit_line_err!(\n+                        \"BRAIN_CONN_OPEN_FAIL: caller=livehost-session-query reason={reason} \\\n+                         kind={:?}: {e} {}\", e.kind(), crate::conn::log_stamp()\n+                    );\n+                    return None;\n+                }\n+            }\n+        }\n+        match self.brain.as_mut().expect(\"connected session census\").sessions() {\n+            Ok(reply) => Some(\n+                reply.sessions.into_iter().map(|s| s.endpoint)\n+                    .filter(|e| !e.is_empty()).collect()\n+            ),\n+            Err(e) => {\n+                // No same-tick retry: the caller skips this census exactly as\n+                // before. Next tick starts a fresh physical connection.\n+                self.brain = None;\n+                self.retry = true;\n+                spt_proto::emit_line_err!(\n+                    \"BRAIN_CONN_RETIRED: caller=livehost-session-query reason=query-error \\\n+                     kind={:?}: {e} {}\", e.kind(), crate::conn::log_stamp()\n+                );\n+                None\n+            }\n+        }\n+    }\n }\n \n /// The normalized program basename a perch's adapter would spawn its Psyche as —\n@@ -1333,11 +1370,12 @@ pub fn spawn_live_host(stop: Arc<AtomicBool>, reason: StartReason) -> JoinHandle\n             // [impl->REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL]\n             resume_restart_orphaned_endpoints(&owlery, &registered, &adapters_dir);\n         }\n+        let mut session_poll = SessionPoll::default();\n         while !stop.load(Ordering::SeqCst) {\n             // TEST-ONLY gate `SPT_LIVEHOST_RECONCILE_DISABLE` (runtime env, DEFAULT\n             // unset = normal production, untouched): skip the reconcile poll+body so an\n             // int rig can prove REQ-UPDATE-TRIAL-DRAIN-DRIVE's core-loop KIND_SESSIONS\n-            // reap-driver in ISOLATION. This loop's own `query_live_session_endpoints()`\n+            // reap-driver in ISOLATION. This loop's own `SessionPoll::query`\n             // → `brain.sessions()` (every LIVE_RECONCILE_INTERVAL_MS) would OTHERWISE\n             // drive the same broker reap and mask the fix (todlando 2026-07-09). Setting\n             // it faithfully REPRODUCES THE FIELD'S livehost-silent update-trial condition\n@@ -1371,7 +1409,7 @@ pub fn spawn_live_host(stop: Arc<AtomicBool>, reason: StartReason) -> JoinHandle\n             // controllable==Some(true) gate inside keeps relay/legacy perches exempt.\n             // [impl->REQ-HAZARD-HOSTED-LIVENESS-RECONCILE]\n             // [impl->REQ-HAZARD-LIVEHOST-BOOT-LIVENESS-GATE]\n-            if let Some(live) = query_live_session_endpoints() {\n+            if let Some(live) = session_poll.query(&crate::endpoint::broker_socket_name()) {\n                 reconcile_hosted_liveness(&owlery, &live);\n             }\n             reconcile_once(&owlery, &registered, &adapters_dir, &set, &cfg, reason);\n@@ -1397,6 +1435,58 @@ mod tests {\n     use crate::test_home::with_home;\n     use std::time::{Duration, Instant};\n \n+    // [unit->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+    #[test]\n+    fn session_poll_reuses_healthy_carrier_and_reopens_after_query_error() {\n+        use crate::codec::{read_frame, write_frame};\n+        use crate::frame::{Envelope, Role};\n+        use crate::transport::{recv_hello, DaemonTransport, LocalSocketTransport};\n+        use std::sync::atomic::AtomicU32;\n+\n+        static SEQ: AtomicU32 = AtomicU32::new(0);\n+        let name = format!(\n+            \"spt-livehost-poll-{}-{}\", std::process::id(),\n+            SEQ.fetch_add(1, Ordering::Relaxed)\n+        );\n+        let listener = LocalSocketTransport::bind(&name).unwrap();\n+        const N: usize = 16;\n+        let server = std::thread::spawn(move || {\n+            let mut accepted = 0;\n+            let mut conn = listener.accept().unwrap();\n+            accepted += 1;\n+            recv_hello(&mut conn, Role::Brain).unwrap();\n+            for _ in 0..N {\n+                assert_eq!(read_frame(&mut conn).unwrap().kind, crate::msg::KIND_SESSIONS);\n+                write_frame(&mut conn, &Envelope::new(\n+                    crate::msg::KIND_SESSIONS_REPLY, serde_json::json!({\"sessions\": []})\n+                )).unwrap();\n+            }\n+            // A malformed reply is a query error even while the transport is\n+            // still alive. It must retire, not leave a desynchronized cache.\n+            assert_eq!(read_frame(&mut conn).unwrap().kind, crate::msg::KIND_SESSIONS);\n+            write_frame(&mut conn, &Envelope::new(\n+                crate::msg::KIND_SESSIONS_REPLY, serde_json::json!({\"sessions\": false})\n+            )).unwrap();\n+            assert!(read_frame(&mut conn).is_err(), \"failed query closes old carrier\");\n+            let mut conn = listener.accept().unwrap();\n+            accepted += 1;\n+            recv_hello(&mut conn, Role::Brain).unwrap();\n+            assert_eq!(read_frame(&mut conn).unwrap().kind, crate::msg::KIND_SESSIONS);\n+            write_frame(&mut conn, &Envelope::new(\n+                crate::msg::KIND_SESSIONS_REPLY, serde_json::json!({\"sessions\": []})\n+            )).unwrap();\n+            accepted\n+        });\n+        let mut poll = SessionPoll::default();\n+        for _ in 0..N {\n+            assert_eq!(poll.query(&name), Some(BTreeSet::new()));\n+        }\n+        assert_eq!(poll.query(&name), None, \"failure is not an empty session census\");\n+        assert_eq!(poll.query(&name), Some(BTreeSet::new()));\n+        drop(poll);\n+        assert_eq!(server.join().unwrap(), 2, \"one healthy carrier plus one recovery\");\n+    }\n+\n     // A no-op Psyche summarizer command: spawns + exits 0 (we never depend on its\n     // output — the daemon online-stamp is what hosts it, not the pid).\n     #[cfg(windows)]\ndiff --git a/docs/RCA-FLEET-DAEMON-14444.md b/docs/RCA-FLEET-DAEMON-14444.md\nindex f7e897d1..010613fd 100644\n--- a/docs/RCA-FLEET-DAEMON-14444.md\n+++ b/docs/RCA-FLEET-DAEMON-14444.md\n@@ -69,3 +69,76 @@ per 2 MB of log, under `[REQ-CONN-POISON-ATTRIBUTION]`, is its own finding.\n    (image path or start time), which is the only way reuse becomes a kill. My own standing rule from\n    the orphan reap is exactly this: re-verify the path AT kill time.\n 3. **Bound the log volume** independently of the cause.\n+\n+## releases#286 — stderr follow-up (2026-09-09)\n+\n+<!-- [doc->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED] -->\n+\n+The historical 303 MB incident above remains the reported baseline, not a\n+newly reproduced event. A read-only sample of the currently installed\n+`%LOCALAPPDATA%/spt-core/logs/daemon.stderr.log` ended at byte **4,163,267**:\n+the last **1,999,958 complete-line bytes** spanned wall stamps\n+**1788997126331–1788998151874** (1025.543 s). It contained **5127** first writes\n+(656,158 bytes), **5092** unpoisoned/non-timeout closes (850,622 bytes),\n+**3** organic retirements, **0** deadline poisons, **107** family-gate lines\n+and **108** rendezvous-up lines. Healthy start/close records alone consumed\n+**75.3%** of that sample. A subsequent live append observation over\n+**307.203 s** added **79,741 bytes** (**259.57 B/s**), with **104** starts,\n+**104** healthy closes and **30** family gates paired with **30** rendezvous\n+binds. These are different observation windows, not a claimed constant rate.\n+No daemon was restarted or reconfigured for either measurement.\n+\n+**Cause discrimination.** `role=brain` identifies *every broker IPC client*,\n+not the supervised coordinator. `run_brain` already keeps its heartbeat\n+`Brain` alive. In contrast, the pre-fix\n+`livehost::query_live_session_endpoints` constructed `Brain::cold_start`,\n+called `sessions`, and dropped the carrier on every invocation from the\n+five-second reconcile loop. This is a source-confirmed avoidable reopen\n+cause, consistent with the recurring five-second bare-brain start/close\n+pairs in the live log, **not an attribution of all 5127 connections to that\n+caller**. Other short-lived callers include dispatch idle queries and\n+per-operation dispatch/CLI clients; their lifetimes are not changed here.\n+\n+The livehost census now owns one carrier across polls. Its `BRAIN_CONN_OPEN`\n+breadcrumb names `caller=livehost-session-query` and distinguishes\n+`reason=initial-query` from `reason=previous-query-error`. Every query error\n+discards the carrier, returns an unavailable census (never an empty\n+authoritative census), and permits a fresh connection only on the next\n+scheduled query. The one-shot boot orphan census remains one-shot.\n+No read/write deadline, poison semantics, or retry cadence changed.\n+\n+**Family gate exonerated.** The sole diagnostic site in\n+`NetEndpoint::bind_scoped` is already once per real bind attempt, not per IPC\n+connection. `pairhost::spawn_meet_rotation` binds a new derived identity for\n+each attached subnet at every 30-second TOTP boundary; the observed\n+`BIGNET`, `SPT_DEV`, and `SPT_MANTLE` windows account for the three lines per\n+boundary. These listeners cannot share a stable identity without changing\n+the pairing protocol. Bind and diagnostic policies are intentionally\n+unchanged; a process-global \"log once\" would hide later real binds.\n+\n+**Explicit contract amendment.** Issue #286 supersedes only the\n+`REQ-CONN-POISON-ATTRIBUTION` leg-(d) requirement for unconditional\n+*healthy start/close lines*. Across the broker process these become\n+`CONN_LIFECYCLE_SUMMARY` totals, at most once per 60-second monotonic interval,\n+flushed by the first subsequent lifecycle counter event. Starts count the\n+first committed write attempt; closes count unpoisoned physical drops.\n+There is no timer thread and no shutdown flush, so a quiet tail remains\n+pending until another counter event (and can be lost on process exit).\n+Poison, organic failure, poisoned close, and all role-transition records\n+remain unconditional per connection. They retain conn id, role facts,\n+wall/monotonic stamps **and the original first-write wall/monotonic time**;\n+a connection that never wrote says `first_write=none`.\n+\n+The deterministic regression cells exercise real framed socket exchanges\n+over both retained and fresh carriers with an isolated fixed log clock/sink,\n+assert one interval summary for N healthy connections, preserve exceptional\n+records through the same sink, and drive livehost census reuse plus recovery\n+from a malformed reply. Focused commands (not executed in the preparation\n+lane; integration owner runs them after landing):\n+\n+```text\n+cargo test -p spt-daemon --lib conn::tests::\n+cargo test -p spt-daemon --lib livehost::tests::session_poll_reuses_healthy_carrier_and_reopens_after_query_error\n+cargo test -p spt-daemon --test conn_blackhole_lifecycle\n+traceable-reqs check --json\n+```\ndiff --git a/traceable-reqs.toml b/traceable-reqs.toml\nindex 21917ea9..9224b3aa 100644\n--- a/traceable-reqs.toml\n+++ b/traceable-reqs.toml\n@@ -2569,6 +2569,14 @@ id = \"REQ-CONN-POISON-ATTRIBUTION\"\n title = \"MSG-IDENTITY W6 / F-039 legs b-d (doyle W6 LOCK 2026-07-10, minted per amendment 3): every broker-conn lifecycle record is ATTRIBUTABLE — the W6 RCA's terminal undecidability (per-line 1:1 CONN_WRITE_POISONED churn = fresh-carrier churn OR stderr interleave artifact) exists because records carry no stable conn identity, no role/endpoint/session context, and no timestamps, and the once-per-conn poison latch hides multiplicity. THREE LEGS. (b) IDENTITY: mint a stable per-physical-conn id (monotonic u64 at conn construction — Arc::ptr_eq is the only identity today and it does not survive a log line) plus subscriber role and endpoint/session where known, stamped on CONN_WRITE_POISONED, CONN_WRITE_RETIRED, logical stall-evict, attach/resume/detach, and write-retirement records (RCA attach sites: presence nethost.rs:379, stream nethost.rs:258, controller broker.rs:891, viewer broker.rs:1073). (c) TIME: daemon stderr correlation records carry wall-clock AND monotonic timestamps (stderrlog has neither; broker+brain share one file — interleave is unresolvable without them). (d) LIFECYCLE (doyle-confirmed UNCONDITIONAL, not debug-gated): one BOUNDED set of per-conn lifecycle events — write start/timeout-cancel/transport close/writer exit/replacement-reattach (hertz RCA fix-shape items 1-3). Constraint (doyle LOCK): the split/attribution must not REDUCE total information, only correct its attribution; NO timeout-value changes; NO suppression-as-fix. Gate: unit — lifecycle records carry conn id + role + timestamps; the id is unique per physical conn and stable across that conn's records. Kin REQ-CONN-POISON-DIAL-SCOPE (leg a, the token split these fields ride on), REQ-CONN-BLACKHOLE-LIFECYCLE-HARNESS (leg e, consumes these records), REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK (behavior invariant preserved).\"\n required_stages = [\"impl\", \"unit\"]  # FLIPPED in the W6 build commit carrying the evidence (todlando 2026-07-10). impl = conn id mint + describe/label accretion + log_stamp wall+mono (anchored at both daemon entries) + bounded lifecycle events (write-start, transport-close, writer-exit, attach/replace/detach) + attribution on stall-evict records. unit = conn.rs record-shape assertions (id/role/time present + parity across both retirement tokens, id unique per conn, label bounded).\n \n+# releases#286 narrowly supersedes leg (d)'s unconditional healthy start/close\n+# lines. Failure/poison, role transitions and original first-write time remain\n+# attributable per physical conn; healthy counts are interval totals instead.\n+[[requirements]]\n+id = \"REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED\"\n+title = \"Healthy broker IPC starts and closes aggregate once per monotonic interval across connections, independently of request volume; poison, failure and role-transition evidence remains per-connection with original first-write time. Livehost session polling reuses its healthy carrier and names why it opens a new one; role=brain alone never identifies the supervised coordinator. Network family diagnostics remain once per real endpoint bind attempt, not IPC churn. Narrow healthy-start/close supersession of REQ-CONN-POISON-ATTRIBUTION leg d (releases#286).\"\n+required_stages = [\"doc\", \"impl\", \"unit\"]\n+\n [[requirements]]\n id = \"REQ-CONN-BLACKHOLE-LIFECYCLE-HARNESS\"\n title = \"MSG-IDENTITY W6 / F-039 leg e (doyle W6 LOCK 2026-07-10, minted per amendment 3 — hertz's five invariants VERBATIM from his RCA fix-shape item 5): 'Build a deterministic black-holed-controller harness against current v0.30.6 semantics and assert: unrelated sessions continue; the bad physical connection is canceled/closed within the bound; its writer exits; a fresh viewer can attach; no lock or task remains owned by the retired connection.' The harness is the standing conformance rig for the r4 SHAREDSEND fix-class — hertz's RCA discipline: only after a timestamped incident maps to a FAILING lifecycle invariant does an ownership/cancellation defect get fixed (the likely shape being complete physical-connection cancellation and writer-task join/retirement, never a broader timeout increase). Consumes REQ-CONN-POISON-ATTRIBUTION's records (conn id + lifecycle events are what make the five assertions checkable deterministically). Kin REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK (the invariant class under test — its brain_decouple int stays the Windows-mandatory gate leg), REQ-CONN-POISON-DIAL-SCOPE.\"\n@@ -4860,6 +4868,7 @@ name = \"message-format-identity\"\n requirements = [\n   \"REQ-CONN-BLACKHOLE-LIFECYCLE-HARNESS\",\n   \"REQ-CONN-POISON-ATTRIBUTION\",\n+  \"REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED\",\n   \"REQ-CONN-POISON-DIAL-SCOPE\",\n   \"REQ-DAEMON-SERVICE-INSTALL\",\n   \"REQ-EP-3\",\n"}],"details":{"resolvedPath":"C:\\Users\\decid\\.omp\\agent\\sessions\\-Documents-projects-spt-core\\2026-09-09T23-45-42-317Z_01a08890-21ad-7000-8f54-208d4f4f2f2c\\65.bash-original.log","contentType":"text/plain","totalLines":659,"displayContent":{"text":"         let name = format!(\n             \"spt-daemon-connattr-{}-{}.sock\",\n@@ -697,13 +818,78 @@ mod tests {\n         let listener = LocalSocketTransport::bind(&name).expect(\"bind\");\n         let client = LocalSocketTransport::connect(&name).expect(\"connect\");\n         let server = listener.accept().expect(\"accept\");\n-        let (_recv, send) = server.split();\n+        let (recv, send) = server.split();\n         (\n-            BrokerConn::new(send, Duration::from_millis(2000)),\n+            BrokerConn::with_log(send, Duration::from_millis(2000), log),\n             client,\n+            recv,\n         )\n     }\n \n+    // [unit->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+    #[test]\n+    fn healthy_round_trips_and_reconnections_emit_one_interval_summary() {\n+        let now = Instant::now();\n+        let log = Arc::new(LifecycleLog::captured(now));\n+        // The real framed-write/drop path, both a persistent carrier and fresh\n+        // per-poll carriers. An isolated clock/sink avoids global test races.\n+        const N: u64 = 32;\n+        for round in 0..N {\n+            let (conn, mut client, mut recv) = make_duplex_with_log(Arc::clone(&log));\n+            conn.describe(\"role=brain\");\n+            for seq in 0..4 {\n+                let request = Envelope::new(\"poll\", serde_json::json!([round, seq]));\n+                write_frame(&mut client, &request).unwrap();\n+                let received = crate::codec::read_frame(&mut recv).unwrap();\n+                assert_eq!(received.payload, request.payload);\n+                assert_eq!(conn.write(&request).unwrap(), ConnWrite::Done);\n+                let reply = crate::codec::read_frame(&mut client).unwrap();\n+                assert_eq!(reply.payload, request.payload);\n+            }\n+            drop(conn);\n+        }\n+        assert!(recover(log.capture.as_ref().unwrap()).is_empty());\n+        *recover(log.clock.as_ref().unwrap()) = now + HEALTHY_LOG_INTERVAL;\n+        let (conn, mut client) = make_conn_with_log(Arc::clone(&log));\n+        conn.describe(\"role=brain\");\n+        conn.write(&Envelope::new(\"poll\", serde_json::json!(N))).unwrap();\n+        assert_eq!(crate::codec::read_frame(&mut client).unwrap().payload, serde_json::json!(N));\n+        drop(conn);\n+        let lines = recover(log.capture.as_ref().unwrap());\n+        assert_eq!(lines.len(), 1, \"healthy volume is per interval, not per conn\");\n+        assert!(lines[0].contains(&format!(\"starts={}\", N + 1)));\n+        assert!(lines[0].contains(&format!(\"closes={N}\")));\n+    }\n+\n+    // [unit->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+    // [unit->REQ-CONN-POISON-ATTRIBUTION]\n+    #[test]\n+    fn failures_and_role_events_bypass_healthy_aggregation_with_original_start() {\n+        let log = Arc::new(LifecycleLog::captured(Instant::now()));\n+        let (conn, mut client) = make_conn_with_log(Arc::clone(&log));\n+        conn.describe(\"role=brain controller session=7\");\n+        conn.write(&Envelope::new(\"poll\", serde_json::json!({}))).unwrap();\n+        crate::codec::read_frame(&mut client).unwrap();\n+        let start = *conn.inner.first_write.get().expect(\"first committed write\");\n+        conn.lifecycle_event(\"writer-exit\", \"reason=write-failed\");\n+        conn.inner.poison_and_cancel(Some(&io::Error::new(io::ErrorKind::BrokenPipe, \"peer gone\")));\n+        // Repeated cancellation must not repeat the retirement record.\n+        conn.inner.poison_and_cancel(None);\n+        let id = conn.id();\n+        drop(conn);\n+        let lines = recover(log.capture.as_ref().unwrap());\n+        assert_eq!(lines.len(), 3, \"role event, retirement and poisoned close survive\");\n+        for line in lines.iter() {\n+            assert!(line.contains(&format!(\"conn={id}\")));\n+            assert!(line.contains(\"role=brain controller session=7\"));\n+            assert!(line.contains(&format!(\"first_write_wall_ms={}\", start.wall_ms)));\n+            assert!(line.contains(&format!(\"first_write_mono_ms={}\", start.mono_ms)));\n+        }\n+        assert!(lines[1].contains(\"CONN_WRITE_RETIRED:\"));\n+        assert!(lines[1].contains(\"BrokenPipe\"));\n+        assert!(lines[2].contains(\"poisoned=true timed_out=false\"));\n+    }\n+\n     // [unit->REQ-CONN-POISON-DIAL-SCOPE] the F-039 leg-(a) token split: the loud\n     // CONN_WRITE_POISONED token is RESERVED for the deadline (timed_out) class;\n     // an organic write failure renders the distinct CONN_WRITE_RETIRED token\ndiff --git a/crates/spt-daemon/src/livehost.rs b/crates/spt-daemon/src/livehost.rs\nindex 63f479d9..151c5e2e 100644\n--- a/crates/spt-daemon/src/livehost.rs\n+++ b/crates/spt-daemon/src/livehost.rs\n@@ -569,7 +569,7 @@ pub fn resume_restart_orphaned_endpoints(\n     registered: &[(AdapterRecord, Manifest)],\n     adapters_dir: &Path,\n ) {\n-    let Some(live) = query_live_session_endpoints() else {\n+    let Some(live) = SessionPoll::default().query(&crate::endpoint::broker_socket_name()) else {\n         return; // broker unreachable — skip (never mass-respawn on a hiccup)\n     };\n     for id in perch::list_self_perch_ids(owlery) {\n@@ -995,20 +995,57 @@ pub fn reconcile_hosted_liveness(owlery: &Path, live_sessions: &BTreeSet<String>\n     offlined\n }\n \n-/// Query the broker for the set of endpoint ids it currently hosts a session for\n-/// (`KIND_SESSIONS`) — the B2 pull signal. `None` when the broker is unreachable\n-/// (the caller then SKIPS the offline pass this tick rather than mass-offlining).\n-fn query_live_session_endpoints() -> Option<BTreeSet<String>> {\n-    let mut brain = Brain::cold_start(&crate::endpoint::broker_socket_name(), now_ms()).ok()?;\n-    let reply = brain.sessions().ok()?;\n-    Some(\n-        reply\n-            .sessions\n-            .into_iter()\n-            .map(|s| s.endpoint)\n-            .filter(|e| !e.is_empty())\n-            .collect(),\n-    )\n+/// The livehost session census is request/reply only, never a subscription.\n+/// Retain its healthy carrier across 5s polls; discard it on EVERY query error\n+/// so a failed/partially consumed reply can never contaminate the next census.\n+/// `None` still means skip reconciliation, not an empty authoritative census.\n+// [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+#[derive(Default)]\n+struct SessionPoll {\n+    brain: Option<Brain>,\n+    retry: bool,\n+}\n+\n+impl SessionPoll {\n+    fn query(&mut self, broker_name: &str) -> Option<BTreeSet<String>> {\n+        if self.brain.is_none() {\n+            let reason = if self.retry { \"previous-query-error\" } else { \"initial-query\" };\n+            match Brain::cold_start(broker_name, now_ms()) {\n+                Ok(brain) => {\n+                    spt_proto::emit_line_err!(\n+                        \"BRAIN_CONN_OPEN: caller=livehost-session-query reason={reason} \\\n+                         policy=reuse-until-query-error {}\",\n+                        crate::conn::log_stamp()\n+                    );\n+                    self.brain = Some(brain);\n+                }\n+                Err(e) => {\n+                    spt_proto::emit_line_err!(\n+                        \"BRAIN_CONN_OPEN_FAIL: caller=livehost-session-query reason={reason} \\\n+                         kind={:?}: {e} {}\", e.kind(), crate::conn::log_stamp()\n+                    );\n+                    return None;\n+                }\n+            }\n+        }\n+        match self.brain.as_mut().expect(\"connected session census\").sessions() {\n+            Ok(reply) => Some(\n+                reply.sessions.into_iter().map(|s| s.endpoint)\n+                    .filter(|e| !e.is_empty()).collect()\n+            ),\n+            Err(e) => {\n+                // No same-tick retry: the caller skips this census exactly as\n+                // before. Next tick starts a fresh physical connection.\n+                self.brain = None;\n+                self.retry = true;\n+                spt_proto::emit_line_err!(\n+                    \"BRAIN_CONN_RETIRED: caller=livehost-session-query reason=query-error \\\n+                     kind={:?}: {e} {}\", e.kind(), crate::conn::log_stamp()\n+                );\n+                None\n+            }\n+        }\n+    }\n }\n \n /// The normalized program basename a perch's adapter would spawn its Psyche as —\n@@ -1333,11 +1370,12 @@ pub fn spawn_live_host(stop: Arc<AtomicBool>, reason: StartReason) -> JoinHandle\n             // [impl->REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL]\n             resume_restart_orphaned_endpoints(&owlery, &registered, &adapters_dir);\n         }\n+        let mut session_poll = SessionPoll::default();\n         while !stop.load(Ordering::SeqCst) {\n             // TEST-ONLY gate `SPT_LIVEHOST_RECONCILE_DISABLE` (runtime env, DEFAULT\n             // unset = normal production, untouched): skip the reconcile poll+body so an\n             // int rig can prove REQ-UPDATE-TRIAL-DRAIN-DRIVE's core-loop KIND_SESSIONS\n-            // reap-driver in ISOLATION. This loop's own `query_live_session_endpoints()`\n+            // reap-driver in ISOLATION. This loop's own `SessionPoll::query`\n             // → `brain.sessions()` (every LIVE_RECONCILE_INTERVAL_MS) would OTHERWISE\n             // drive the same broker reap and mask the fix (todlando 2026-07-09). Setting\n             // it faithfully REPRODUCES THE FIELD'S livehost-silent update-trial condition\n@@ -1371,7 +1409,7 @@ pub fn spawn_live_host(stop: Arc<AtomicBool>, reason: StartReason) -> JoinHandle\n             // controllable==Some(true) gate inside keeps relay/legacy perches exempt.\n             // [impl->REQ-HAZARD-HOSTED-LIVENESS-RECONCILE]\n             // [impl->REQ-HAZARD-LIVEHOST-BOOT-LIVENESS-GATE]\n-            if let Some(live) = query_live_session_endpoints() {\n+            if let Some(live) = session_poll.query(&crate::endpoint::broker_socket_name()) {\n                 reconcile_hosted_liveness(&owlery, &live);\n             }\n             reconcile_once(&owlery, &registered, &adapters_dir, &set, &cfg, reason);\n@@ -1397,6 +1435,58 @@ mod tests {\n     use crate::test_home::with_home;\n     use std::time::{Duration, Instant};\n \n+    // [unit->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+    #[test]\n+    fn session_poll_reuses_healthy_carrier_and_reopens_after_query_error() {\n+        use crate::codec::{read_frame, write_frame};\n+        use crate::frame::{Envelope, Role};\n+        use crate::transport::{recv_hello, DaemonTransport, LocalSocketTransport};\n+        use std::sync::atomic::AtomicU32;\n+\n+        static SEQ: AtomicU32 = AtomicU32::new(0);\n+        let name = format!(\n+            \"spt-livehost-poll-{}-{}\", std::process::id(),\n+            SEQ.fetch_add(1, Ordering::Relaxed)\n+        );\n+        let listener = LocalSocketTransport::bind(&name).unwrap();\n+        const N: usize = 16;\n+        let server = std::thread::spawn(move || {\n+            let mut accepted = 0;\n+            let mut conn = listener.accept().unwrap();\n+            accepted += 1;\n+            recv_hello(&mut conn, Role::Brain).unwrap();\n+            for _ in 0..N {\n+                assert_eq!(read_frame(&mut conn).unwrap().kind, crate::msg::KIND_SESSIONS);\n+                write_frame(&mut conn, &Envelope::new(\n+                    crate::msg::KIND_SESSIONS_REPLY, serde_json::json!({\"sessions\": []})\n+                )).unwrap();\n+            }\n+            // A malformed reply is a query error even while the transport is\n+            // still alive. It must retire, not leave a desynchronized cache.\n+            assert_eq!(read_frame(&mut conn).unwrap().kind, crate::msg::KIND_SESSIONS);\n+            write_frame(&mut conn, &Envelope::new(\n+                crate::msg::KIND_SESSIONS_REPLY, serde_json::json!({\"sessions\": false})\n+            )).unwrap();\n+            assert!(read_frame(&mut conn).is_err(), \"failed query closes old carrier\");\n+            let mut conn = listener.accept().unwrap();\n+            accepted += 1;\n+            recv_hello(&mut conn, Role::Brain).unwrap();\n+            assert_eq!(read_frame(&mut conn).unwrap().kind, crate::msg::KIND_SESSIONS);\n+            write_frame(&mut conn, &Envelope::new(\n+                crate::msg::KIND_SESSIONS_REPLY, serde_json::json!({\"sessions\": []})\n+            )).unwrap();\n+            accepted\n+        });\n+        let mut poll = SessionPoll::default();\n+        for _ in 0..N {\n+            assert_eq!(poll.query(&name), Some(BTreeSet::new()));\n+        }\n+        assert_eq!(poll.query(&name), None, \"failure is not an empty session census\");\n+        assert_eq!(poll.query(&name), Some(BTreeSet::new()));\n+        drop(poll);\n+        assert_eq!(server.join().unwrap(), 2, \"one healthy carrier plus one recovery\");\n+    }\n+\n     // A no-op Psyche summarizer command: spawns + exits 0 (we never depend on its\n     // output — the daemon online-stamp is what hosts it, not the pid).\n     #[cfg(windows)]\ndiff --git a/docs/RCA-FLEET-DAEMON-14444.md b/docs/RCA-FLEET-DAEMON-14444.md\nindex f7e897d1..010613fd 100644\n--- a/docs/RCA-FLEET-DAEMON-14444.md\n+++ b/docs/RCA-FLEET-DAEMON-14444.md\n@@ -69,3 +69,76 @@ per 2 MB of log, under `[REQ-CONN-POISON-ATTRIBUTION]`, is its own finding.\n    (image path or start time), which is the only way reuse becomes a kill. My own standing rule from\n    the orphan reap is exactly this: re-verify the path AT kill time.\n 3. **Bound the log volume** independently of the cause.\n+\n+## releases#286 — stderr follow-up (2026-09-09)\n+\n+<!-- [doc->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED] -->\n+\n+The historical 303 MB incident above remains the reported baseline, not a\n+newly reproduced event. A read-only sample of the currently installed\n+`%LOCALAPPDATA%/spt-core/logs/daemon.stderr.log` ended at byte **4,163,267**:\n+the last **1,999,958 complete-line bytes** spanned wall stamps\n+**1788997126331–1788998151874** (1025.543 s). It contained **5127** first writes\n+(656,158 bytes), **5092** unpoisoned/non-timeout closes (850,622 bytes),\n+**3** organic retirements, **0** deadline poisons, **107** family-gate lines\n+and **108** rendezvous-up lines. Healthy start/close records alone consumed\n+**75.3%** of that sample. A subsequent live append observation over\n+**307.203 s** added **79,741 bytes** (**259.57 B/s**), with **104** starts,\n+**104** healthy closes and **30** family gates paired with **30** rendezvous\n+binds. These are different observation windows, not a claimed constant rate.\n+No daemon was restarted or reconfigured for either measurement.\n+\n+**Cause discrimination.** `role=brain` identifies *every broker IPC client*,\n+not the supervised coordinator. `run_brain` already keeps its heartbeat\n+`Brain` alive. In contrast, the pre-fix\n+`livehost::query_live_session_endpoints` constructed `Brain::cold_start`,\n+called `sessions`, and dropped the carrier on every invocation from the\n+five-second reconcile loop. This is a source-confirmed avoidable reopen\n+cause, consistent with the recurring five-second bare-brain start/close\n+pairs in the live log, **not an attribution of all 5127 connections to that\n+caller**. Other short-lived callers include dispatch idle queries and\n+per-operation dispatch/CLI clients; their lifetimes are not changed here.\n+\n+The livehost census now owns one carrier across polls. Its `BRAIN_CONN_OPEN`\n+breadcrumb names `caller=livehost-session-query` and distinguishes\n+`reason=initial-query` from `reason=previous-query-error`. Every query error\n+discards the carrier, returns an unavailable census (never an empty\n+authoritative census), and permits a fresh connection only on the next\n+scheduled query. The one-shot boot orphan census remains one-shot.\n+No read/write deadline, poison semantics, or retry cadence changed.\n+\n+**Family gate exonerated.** The sole diagnostic site in\n+`NetEndpoint::bind_scoped` is already once per real bind attempt, not per IPC\n+connection. `pairhost::spawn_meet_rotation` binds a new derived identity for\n+each attached subnet at every 30-second TOTP boundary; the observed\n+`BIGNET`, `SPT_DEV`, and `SPT_MANTLE` windows account for the three lines per\n+boundary. These listeners cannot share a stable identity without changing\n+the pairing protocol. Bind and diagnostic policies are intentionally\n+unchanged; a process-global \"log once\" would hide later real binds.\n+\n+**Explicit contract amendment.** Issue #286 supersedes only the\n+`REQ-CONN-POISON-ATTRIBUTION` leg-(d) requirement for unconditional\n+*healthy start/close lines*. Across the broker process these become\n+`CONN_LIFECYCLE_SUMMARY` totals, at most once per 60-second monotonic interval,\n+flushed by the first subsequent lifecycle counter event. Starts count the\n+first committed write attempt; closes count unpoisoned physical drops.\n+There is no timer thread and no shutdown flush, so a quiet tail remains\n+pending until another counter event (and can be lost on process exit).\n+Poison, organic failure, poisoned close, and all role-transition records\n+remain unconditional per connection. They retain conn id, role facts,\n+wall/monotonic stamps **and the original first-write wall/monotonic time**;\n+a connection that never wrote says `first_write=none`.\n+\n+The deterministic regression cells exercise real framed socket exchanges\n+over both retained and fresh carriers with an isolated fixed log clock/sink,\n+assert one interval summary for N healthy connections, preserve exceptional\n+records through the same sink, and drive livehost census reuse plus recovery\n+from a malformed reply. Focused commands (not executed in the preparation\n+lane; integration owner runs them after landing):\n+\n+```text\n+cargo test -p spt-daemon --lib conn::tests::\n+cargo test -p spt-daemon --lib livehost::tests::session_poll_reuses_healthy_carrier_and_reopens_after_query_error\n+cargo test -p spt-daemon --test conn_blackhole_lifecycle\n+traceable-reqs check --json\n+```\ndiff --git a/traceable-reqs.toml b/traceable-reqs.toml\nindex 21917ea9..9224b3aa 100644\n--- a/traceable-reqs.toml\n+++ b/traceable-reqs.toml\n@@ -2569,6 +2569,14 @@ id = \"REQ-CONN-POISON-ATTRIBUTION\"\n title = \"MSG-IDENTITY W6 / F-039 legs b-d (doyle W6 LOCK 2026-07-10, minted per amendment 3): every broker-conn lifecycle record is ATTRIBUTABLE — the W6 RCA's terminal undecidability (per-line 1:1 CONN_WRITE_POISONED churn = fresh-carrier churn OR stderr interleave artifact) exists because records carry no stable conn identity, no role/endpoint/session context, and no timestamps, and the once-per-conn poison latch hides multiplicity. THREE LEGS. (b) IDENTITY: mint a stable per-physical-conn id (monotonic u64 at conn construction — Arc::ptr_eq is the only identity today and it does not survive a log line) plus subscriber role and endpoint/session where known, stamped on CONN_WRITE_POISONED, CONN_WRITE_RETIRED, logical stall-evict, attach/resume/detach, and write-retirement records (RCA attach sites: presence nethost.rs:379, stream nethost.rs:258, controller broker.rs:891, viewer broker.rs:1073). (c) TIME: daemon stderr correlation records carry wall-clock AND monotonic timestamps (stderrlog has neither; broker+brain share one file — interleave is unresolvable without them). (d) LIFECYCLE (doyle-confirmed UNCONDITIONAL, not debug-gated): one BOUNDED set of per-conn lifecycle events — write start/timeout-cancel/transport close/writer exit/replacement-reattach (hertz RCA fix-shape items 1-3). Constraint (doyle LOCK): the split/attribution must not REDUCE total information, only correct its attribution; NO timeout-value changes; NO suppression-as-fix. Gate: unit — lifecycle records carry conn id + role + timestamps; the id is unique per physical conn and stable across that conn's records. Kin REQ-CONN-POISON-DIAL-SCOPE (leg a, the token split these fields ride on), REQ-CONN-BLACKHOLE-LIFECYCLE-HARNESS (leg e, consumes these records), REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK (behavior invariant preserved).\"\n required_stages = [\"impl\", \"unit\"]  # FLIPPED in the W6 build commit carrying the evidence (todlando 2026-07-10). impl = conn id mint + describe/label accretion + log_stamp wall+mono (anchored at both daemon entries) + bounded lifecycle events (write-start, transport-close, writer-exit, attach/replace/detach) + attribution on stall-evict records. unit = conn.rs record-shape assertions (id/role/time present + parity across both retirement tokens, id unique per conn, label bounded).\n \n+# releases#286 narrowly supersedes leg (d)'s unconditional healthy start/close\n+# lines. Failure/poison, role transitions and original first-write time remain\n+# attributable per physical conn; healthy counts are interval totals instead.\n+[[requirements]]\n+id = \"REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED\"\n+title = \"Healthy broker IPC starts and closes aggregate once per monotonic interval across connections, independently of request volume; poison, failure and role-transition evidence remains per-connection with original first-write time. Livehost session polling reuses its healthy carrier and names why it opens a new one; role=brain alone never identifies the supervised coordinator. Network family diagnostics remain once per real endpoint bind attempt, not IPC churn. Narrow healthy-start/close supersession of REQ-CONN-POISON-ATTRIBUTION leg d (releases#286).\"\n+required_stages = [\"doc\", \"impl\", \"unit\"]\n+\n [[requirements]]\n id = \"REQ-CONN-BLACKHOLE-LIFECYCLE-HARNESS\"\n title = \"MSG-IDENTITY W6 / F-039 leg e (doyle W6 LOCK 2026-07-10, minted per amendment 3 — hertz's five invariants VERBATIM from his RCA fix-shape item 5): 'Build a deterministic black-holed-controller harness against current v0.30.6 semantics and assert: unrelated sessions continue; the bad physical connection is canceled/closed within the bound; its writer exits; a fresh viewer can attach; no lock or task remains owned by the retired connection.' The harness is the standing conformance rig for the r4 SHAREDSEND fix-class — hertz's RCA discipline: only after a timestamped incident maps to a FAILING lifecycle invariant does an ownership/cancellation defect get fixed (the likely shape being complete physical-connection cancellation and writer-task join/retirement, never a broader timeout increase). Consumes REQ-CONN-POISON-ATTRIBUTION's records (conn id + lifecycle events are what make the five assertions checkable deterministically). Kin REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK (the invariant class under test — its brain_decouple int stays the Windows-mandatory gate leg), REQ-CONN-POISON-DIAL-SCOPE.\"\n@@ -4860,6 +4868,7 @@ name = \"message-format-identity\"\n requirements = [\n   \"REQ-CONN-BLACKHOLE-LIFECYCLE-HARNESS\",\n   \"REQ-CONN-POISON-ATTRIBUTION\",\n+  \"REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED\",\n   \"REQ-CONN-POISON-DIAL-SCOPE\",\n   \"REQ-DAEMON-SERVICE-INSTALL\",\n   \"REQ-EP-3\",\n","startLine":301,"lineNumbers":[301,302,303,304,305,306,307,308,309,310,311,312,313,314,315,316,317,318,319,320,321,322,323,324,325,326,327,328,329,330,331,332,333,334,335,336,337,338,339,340,341,342,343,344,345,346,347,348,349,350,351,352,353,354,355,356,357,358,359,360,361,362,363,364,365,366,367,368,369,370,371,372,373,374,375,376,377,378,379,380,381,382,383,384,385,386,387,388,389,390,391,392,393,394,395,396,397,398,399,400,401,402,403,404,405,406,407,408,409,410,411,412,413,414,415,416,417,418,419,420,421,422,423,424,425,426,427,428,429,430,431,432,433,434,435,436,437,438,439,440,441,442,443,444,445,446,447,448,449,450,451,452,453,454,455,456,457,458,459,460,461,462,463,464,465,466,467,468,469,470,471,472,473,474,475,476,477,478,479,480,481,482,483,484,485,486,487,488,489,490,491,492,493,494,495,496,497,498,499,500,501,502,503,504,505,506,507,508,509,510,511,512,513,514,515,516,517,518,519,520,521,522,523,524,525,526,527,528,529,530,531,532,533,534,535,536,537,538,539,540,541,542,543,544,545,546,547,548,549,550,551,552,553,554,555,556,557,558,559,560,561,562,563,564,565,566,567,568,569,570,571,572,573,574,575,576,577,578,579,580,581,582,583,584,585,586,587,588,589,590,591,592,593,594,595,596,597,598,599,600,601,602,603,604,605,606,607,608,609,610,611,612,613,614,615,616,617,618,619,620,621,622,623,624,625,626,627,628,629,630,631,632,633,634,635,636,637,638,639,640,641,642,643,644,645,646,647,648,649,650,651,652,653,654,655,656,657,658,659]},"meta":{"source":{"type":"internal","value":"artifact://65"}}},"isError":false,"timestamp":1788999516713}}
{"type":"message","id":"9834b7e6","parentId":"aae83338","timestamp":"2026-09-10T00:18:36.721Z","message":{"role":"toolResult","toolCallId":"call_uDXE7hGZPSO1wpkbLungmWy9|fc_0dde72effda0ad21016aa1f75c942487d2bfbe9feccf7adcba","toolName":"grep","content":[{"type":"text","text":"# .worktrees/286-churn/\n## PRD.md\n 9|\n*10|spt-core is a **harness-independent core for an agent ecosystem**: inter-agent messaging, live-agent lifecycle, terminal hosting, seamless self-update, and zero-config cross-machine networking — shipped as both a Rust library workspace and a single canonical binary (`spt` / `spt.exe`).\n 11|\n 12|It is the ground-up rebuild of `claude_skill_owl` (\"modern SPT\"), untethering the system from Claude Code so that **any** agent runtime — Claude Code, Codex, Cursor, headless, or novel surfaces — can participate by either shelling out to the binary or linking the crates. The first thing built atop it is a rebuilt `spt` plugin reaching parity with today's system; the networking, terminal-wrapper, and multi-instance capabilities then make an agent reachable and operable across all of a user's machines.\n 13|\n...\n 16|- Today's modern SPT works but is **fragile** — each release introduces 2–5 regressions because it grew rapidly from a narrow starting point and carries the resulting tech debt (ADR-0001).\n*17|- It is **welded to Claude Code**. Other harnesses can't reuse the messaging/lifecycle substrate.\n 18|- Networking, terminal hosting (\"capsule\"), and cross-machine operation were planned piecemeal atop a CC-tied core; they deserve to be first-class in a clean foundation.\n 19|\n 20|## 3. Goals\n...\n 31|\n*32|- **macOS support** — out (no test machine); kept structurally easy, not a re-architecture.\n 33|- **Cross-subnet** (two *different* users' agents communicating) — deferred concept.\n 34|- **Concrete Shells / PresenceChannel implementation / presence gossip** — only the seams ship (§Endpoints).\n 35|- **Remote command execution on another user's node** and **instantiate-anywhere** — seam-compatible, deferred behind a consent/security gate.\n...\n 54|- **R-ARCH-3** `spt-proto` carries an explicit **wire-protocol version** independent of crate semver, with a documented **N-1 compat window**. Binary and library consumers interoperate by proto version, never by release version.\n*55|- **R-ARCH-4** Copy-verbatim the commodity layer from the sister project (envelope grammar, spool DDL, registry RPC, `info.json` shape); clean-room only the architecture/lifecycle (ADR-0001).\n 56|\n*57|### 7.2 The daemon (broker/brain)\n 58|- **R-DAEMON-1** Exactly **one logical `spt-daemon` per machine**, always-on, owning PTYs, network identity + WAN endpoint, subnet registry, all spools, all poll-listener logic, and all Psyche/pulse loops. No separate listener or wrapper processes (ADR-0004).\n*59|- **R-DAEMON-2** Internally split into a stable **broker** (holds PTY masters, harness child processes, listening sockets; versioned IPC; rarely updates) and an **updatable brain** (all logic; rehydrates from disk + re-attaches to broker handles).\n 60|- **R-DAEMON-3** Any `spt api` invocation that needs the daemon **auto-starts it** if absent.\n 61|- **R-DAEMON-4** Honor every invariant in [`docs/KNOWN-HAZARDS.md`](./docs/KNOWN-HAZARDS.md) as a conformance checklist.\n 62|\n...\n 88|- **R-START-2** **Harness-hosted:** SessionStart → `spt api seed --pid` (in-memory seed record) → `$SPT listen <id>` consumes the seed by `parent_pid` (validated vs `session_id`), binds, relays. No file.\n*89|- **R-START-3** **spt-hosted:** daemon runs spawn-session into a broker PTY → binary `api bind` (or skips under commitment) → daemon delivers via manifest-configured method. Principle: **file-bridge only when the daemon is not the launcher.**\n 90|- **R-START-4** Adapters inject env aliases (`$SPT` = `spt api`, heirs to `$OWL`/`$LIVE`); spt-core supplies the subcommands.\n 91|\n 92|---\n...\n 99|- **R-EP-3** Messaging payloads carry **typed operation commands + arbitrary file blobs** (text/audio/image/video), not just text — the seam that lets Shells, file transfer, and rich PresenceChannel exist later.\n*100|- **R-EP-4** PresenceChannel (deferred impl) is a **broker** endpoint with dispatch/bind/thread styles; presence datum `(last_active_node, last_active_endpoint, ts)`.\n 101|- **R-EP-5** (deferred impl, CONTEXT \"Shell model\") **Concrete shell instantiation model:** `shell spawn` *mints* an owner-exclusive instance (`<adapter>-<n>`, distinct from the `relink`/`persistent`/`wake` online switch); permission to instantiate = shell adapter **registered on the node** (broadcast governs *discovery* only); a per-shell **`require_approval`** gate (`none`|`remembered`|`always`, default none; manifest floor, node/endpoint may tighten) reuses the consent grant store; an optional **`max_...\n 102|\n 103|### 9.2 Instances (ADR-0003, as amended)\n...\n 153|\n*154|- **R-UPD-1** **Peer-propagated** over P2P (layered on self-fetch; out-of-band still supported); the subnet self-heals to latest (ADR-0004).\n 155|- **R-UPD-2** **All binaries signature-verified before handoff**, regardless of source (spt-core release key).\n*156|- **R-UPD-3** **No-endpoint-termination invariant, scoped by update class (ADR-0004 §A).** For the **brain-only** class — the routine case — the guarantee is **absolute**: no endpoint process terminates or suspends, satisfied by the broker/brain split (the broker keeps held PTYs/children/sockets alive across brain replacement). **broker-compatible** swaps the broker behind the versioned IPC with endpoints still surviving. **broker-breaking** (held-resource-type / OS-API change) MAY suspend endpoints...\n 157|- **R-UPD-4** **Not auto by default** — gated on user confirmation delivered to the most-recently-active live session, with an opt-in full-auto.\n 158|- **R-UPD-5** spt-core conducts **adapter ripple-updates** after self-update, per each adapter's update declaration and `min_spt_core_version` gate.\n 159|\n...\n 163|\n*164|- **R-TERM-1** A process-supervisor terminal wrapper (supersedes \"capsule\") hosting sessions in broker PTYs, multiplexed and name-addressable, attach/detach, headless↔headed.\n 165|- **R-TERM-2** A **session-surface** abstraction with native PTY as the day-one impl and network-attached surfaces for remote control; input injection at both `send-keys` (raw) and `send-line` (cooked) granularity; in-memory scrollback ring (disk spillover deferred).\n 166|- **R-TERM-3** Remote terminal streaming uses a **byte-stream** model for v1 (simple, faithful; near-LAN on Iroh-direct paths). Keep the surface abstraction generic enough that a predictive/state-sync \"feels-local\" layer can be added later without re-architecting — not built v1.\n*167|- **R-TERM-4** (M3, ADR-0008) **Live activity buffer (PTY digest):** a rolling, *parsed*, structured (source-tagged) view of recent PTY I/O (last ~N user turns + agent output between; tool sprints collapsed), built by running **adapter-supplied** manifest patterns (`input_pattern`, `agent_pattern`, `tool_pattern[]` + catchall) over the **broker-owned** PTY bytes — honors the no-built-in-parser rule (adapter owns the pattern). **Spt-hosted-primary** (needs the broker PTY); **harness-hosted is capability...\n 168|- **R-FRONT-1** Day-one headed frontend is a **launcher/manager** (not raw attach): list running + historic endpoints; launch a historic endpoint from its original manifest; tap into a running one; init a new one via a known adapter. Realizes the \"guided resume\" (XMB-style) and \"management GUI\" (per-endpoint panes) sketches. Its **CLI sibling** is a no-arg **`spt resume`** guided-resume picker — endpoints grouped by locality (on-node/this-project → on-node/other-project → off-node), MRU within grou...\n 169|\n 170|---\n...\n 176|- **R-INSTALL-3** First run idempotent + interactive-optional: generate node identity, start daemon; pairing + subnet-naming are separate explicit steps.\n*177|- **R-INSTALL-4** (CONTEXT \"adapter registration\") **Adapter registration lifecycle:** `spt adapter add <path>` (or `--github <user/repo>`, manifest-first then install via the declared `[update]` avenue — *install is the first update*) validates the manifest and records it under `{SPT_HOME}/…/adapters/` (copy for `file_pull`, pointer for `delegated`); one command/dir for `kind=\"harness\"` and `kind=\"shell\"`. Node-local (drive/launch capability, not subnet advertisement). `spt adapter remove` is **soft...\n 178|- **R-MIGRATE-1** Standalone install **auto-detects a `claude_skill_owl` install and offers migration** (identity, agents, tracked Psyche context). First-class supported path.\n 179|\n 180|---\n...\n 189|- **R-DOCS-4** **Agent-consumable layer:** CI-generated `llms.txt`/`llms-full.txt`, markdown content negotiation, a machine-readable manifest JSON Schema at a stable path, an **MCP doc/resource server** exposing the harness contract as deterministic resources, and structured CLI help (`spt <cmd> --help` + a `--json` mode).\n*190|- **R-DOCS-5** **Anti-drift via CI:** rustdoc, the manifest schema, the agent exports, and CLI help are generated + checked in CI — doc quality on the same footing as tests.\n 191|\n 192|---\n 193|\n...\n 195|\n*196|**In v1:** harness-independent core (library + binary); manifest + `api` harness contract; both startup topologies; one-daemon broker/brain; LiveAgent/ReadyAgent/Psyche/Worker/SptNode endpoints + Shell-vs-agent seam + typed/binary payloads; multi-instance data model + subnet registry + resolution policy + dormant/suspended + two-tier context sync + remote-drive of running instances + off-node file transfer; first-class WAN networking + TOTP-SPAKE2 pairing + subnet naming; cross-node Psyche sync (replaces...\n 197|\n 198|**Deferred (seam-compatible):** see [`docs/DEFERRED.md`](./docs/DEFERRED.md) — concrete Shells, PresenceChannel impl, presence gossip, instantiate-anywhere + consent, remote command exec + security gate, sidecar adapters, manifest includes, scrollback spillover, predictive/state-sync terminal layer, OS-level presence signal, macOS.\n 199|\n...\n 201|\n*202|1. A non-CC harness (or a headless spike) drives the full live-agent lifecycle through manifest + `api` alone — proving harness independence.\n 203|2. The rebuilt spt-plugin reaches **feature parity** with frozen `claude_skill_owl`.\n 204|3. Two of a user's machines, freshly paired via one TOTP code, exchange messages and the user drives an agent on machine A from machine B — zero config beyond pairing.\n 205|4. A self-update rolls across the subnet **without terminating any endpoint**.\n 206|5. The same agent's mind follows the user across machines (two-tier sync); switching machines leaves the prior instance dormant and catches up the new active one.\n*207|6. Every `KNOWN-HAZARDS` invariant has a corresponding passing test.\n 208|\n 209|## 17. Open design-phase questions\n 210|\n...\n 222|- [`docs/adr/0001`](./docs/adr/0001-clean-room-fork-from-claude-skill-owl.md) … [`0005`](./docs/adr/0005-totp-seeded-spake2-node-pairing.md) — decisions.\n*223|- [`docs/KNOWN-HAZARDS.md`](./docs/KNOWN-HAZARDS.md) — invariants harvested from the sister project (test checklist).\n 224|- [`docs/DEFERRED.md`](./docs/DEFERRED.md) — parked-but-committed features.\n 225|- [`docs/STORAGE.md`](./docs/STORAGE.md) — `$SPT_HOME` on-disk layout.\n 226|- [`docs/MANIFEST.md`](./docs/MANIFEST.md) — runtime manifest schema + worked `claude-spt` example.\n## ROADMAP.md\n 6|\n*7|A harness-independent core for an agent ecosystem — messaging, live-agent lifecycle, terminal hosting, **first-class P2P networking**, and a runtime-manifest harness contract — shipped as Rust crates + a single `spt` binary. The proof: the same agent runs natively across all of a user's machines with a synced mind, reachable and operable from anywhere, with seamless self-update and zero-config cross-machine messaging. First thing built atop it: a rebuilt `spt` plugin at parity with today's modern SPT.\n 8|\n 9|## Scope decisions (locked)\n 10|\n...\n 16|\n*17|**Stage A — Red-team review (before building).** Pressure-test the design adversarially. Focus on the riskiest architectural decisions, not bikeshedding. Tools: `plan-eng-review` (architecture) + `codex` (adversarial second opinion). Scope-cutting is *not* a goal (v1 is locked as full PRD), but surfacing contradictions / over-engineering / unproven assumptions is.\n 18|\n 19|**Stage B — De-risk spikes.** Throwaway proofs of the load-bearing unknowns — anything that, if it fails, forces a redesign. Do these before committing to the full build. See *De-risk spikes* below.\n 20|\n...\n 28|- **M1 — Local messaging + binary.** `spt-msg` (delivery TCP+spool, routing, send/ring/ready) + the `spt` binary + CLI. **= the killer quickstart (local): two agents exchange a message.** Task breakdown: [`M1-PLAN.md`](./M1-PLAN.md).\n*29|- **M2 — Harness contract + lifecycle.** **Strictly harness-agnostic — zero Claude Code conventions in this repo.** spt-core defines the adapter contract; the only adapter-shaped artifact in-tree is a generic mock/test adapter (manifest fixture + fake `api` caller) that exercises the contract (PRD R-DOCS-2 dev-agent quickstart). **The rebuilt `spt` plugin (Claude Code) is a SEPARATE downstream project** built atop these crates/binary — spt-core v1's acceptance proof (PRD §\"feature parity\"), not an...\n 30|  - **M2a — Harness contract. ✅ delivered (2026-06-01).** `spt-runtime` (AgentRuntime, ManifestRuntime, manifest schema) + the `api` subcommand surface (seed/listen/bind/state/echo-gate/poll/worker-*/boundary/session-end/presence/history-log/emit/capability) + the **harness-hosted** startup topology (seed→listen→bind, interim seed-file) + local-api-auth + the generic mock adapter + contract E2E. All M2a reqs activated + green. Task breakdown: [`M2a-PLAN.md`](./M2a-PLAN.md).\n*31|  - **M2b — Live-agent lifecycle. ✅ delivered (2026-06-01).** `spt-live` (fifth layer): history subsystem (fetcher/locate-normalize/native) + spawn-psyche (nested perch, spt-core-owned `$psyche_prompt`) + echo-commune (history brief → single-writer commune drop) + drop-file ingest (single-writer 6.4 + direct-write precedence 6.5) + interim pulse + signoff/boundary ordering invariants (echo-before-signoff 3.3, grace-before-signoff 1.1, stale-signoff sweep 3.2) + resume seam (continue-existing / fres...\n*32|  - **Interim, no daemon** (like M1): the `api listen` process holds the perch + relays; the seed bridges via a file. **The spt-hosted topology (PTY-launch) + the consolidated `spt-daemon` move to M3** (they need the `spt-term` broker) — this narrows the earlier \"both topologies in M2\" wording.\n*33|- **M3 — Terminal wrapper + daemon + self-update.** `spt-term` (session-surface, PTY, broker) + the consolidated `spt-daemon` (broker/brain split) + peer-propagated signed self-update. Split M3a/M3b/M3c. Scope + carried-forward Spike #1 gaps: [`M3-PLAN.md`](./M3-PLAN.md).\n 34|  - **M3a — `spt-term` session-surface mechanism. ✅ delivered (2026-06-02).** The sixth crate (`…→spt-msg→spt-term`, sibling of spt-runtime): the `SessionSurface` trait + native PTY backends (ConPTY/forkpty via `portable-pty`, REQ-TERM-1) · ConPTY-DSR auto-answer drain pump (REQ-HAZARD-CONPTY-DSR) · send-keys/send-line injection (REQ-TERM-2) · bounded byte-stream (REQ-TERM-3) · the PTY-digest parser primitive (REQ-TERM-4 impl/unit, ADR-0008). OS-neutral per Spike #4/#5 (no OS's stream contr...\n*35|  - **M3b — `spt-daemon` broker/brain process. ✅ delivered (2026-06-03).** The seventh crate (`…→spt-live→spt-daemon→spt`) hosts many `spt-term` surfaces behind a **versioned local IPC** (named-pipe/UDS behind one `DaemonTransport` trait, forward-compatible frame schema): a **broker** (stable kernel — PTY masters, harness children, sockets; no logic) + a **brain** (restartable logic) so a brain kill/restart leaves the hosted PTY child + its output stream **intact and gapless** (Spike #1 mad...\n*36|  - **M3c — signed self-update. ✅ delivered (2026-06-03).** The gated, signed, ripple-capable self-update engine (no new crate — it lives in `spt-daemon`, signing primitives in `release.rs`): the **update-class taxonomy** (brain-only / broker-compatible / broker-breaking, ADR-0004 §A) with the **brain-only zero-interruption swap** that drives the M3b handoff substrate (snapshot→drop→re-attach) so a live PTY child + its output stream survive the logic swap **untouched** (REQ-UPD-3, proven E2E `...\n*37|  - **M3 COMPLETE** (M3a + M3b + M3c all ✅, 2026-06-03): the real ADR-0004 architecture is live — native PTY hosting, the consolidated broker/brain daemon, and seamless signed self-update. Next milestone: M4.\n*38|- **M4 — Networking + instances. ✅ delivered (2026-06-04).** The eighth crate `spt-net` (Iroh WAN endpoint bound to the node's own Ed25519 identity, mDNS LAN discovery, TOTP-SPAKE2 pairing over a dedicated pre-trust ALPN with transcript binding / rate limiting / seed rotation+transfer, subnet registry + NDJSON replication under the per-node epoch lease) · the multi-instance model (registry rows + bare-id resolution refuse-and-qualify + visibility/sync-membership gates + rename/collision + **resting-...\n*39|- **M5 — Shells, presence, deferred capabilities. ✅ delivered (2026-06-04;** scope locked with user 2026-06-04, see [`M5-PLAN.md`](./M5-PLAN.md) + per-task JIT plans D0–D9). The consent framework seam (grant store + interactive escalation + pre-consent flags; remote-exec / instantiate-anywhere reserved-but-refusing) · adapter registration lifecycle (`spt adapter add/remove`, manifest-first) · **shell hosting machinery + sleep/wake** (link-token channels with per-frame MAC, broker-launched binarie...\n 40|\n 41|## Post-v1-core milestones\n 42|\n...\n 46|\n*47|- **✅ Broker/brain process-isolation restoration — delivered (2026-06-11).** Design: [`docs/BROKER-BRAIN-SPLIT-RESTORATION.md`](./docs/BROKER-BRAIN-SPLIT-RESTORATION.md) (verified-with-amendments by `doyle`); decision: **ADR-0018**; build plan: [`RESTORATION-PLAN.md`](./RESTORATION-PLAN.md) (D1–D7). Corrects an unintended regression — the broker/brain split (ADR-0004) was specced + spiked as two *processes*, but the production daemon runs the broker as an in-process *thread* (`daemon.rs:165-170`)...\n 48|\n 49|- **🔜 NEXT milestone (planned 2026-06-12) — M9: Adapter customization & session surfaces.** Product of the two-day **Gateway grill** (operator + `doyle`; CONTEXT.md mints @ `8dc22ff`/`bd469b9`; ADR-0008 amended; the lecturn device design parked in the sibling `spt-lecturn/` repo for a later build). Shell-free, independently-shippable features: **adapter profiles** (sparse leaf-replace overlays on the parent manifest; shipped + local sources, local surviving adapter updates; composite `<adapter>:<pro...\n*50|- **🔜 M10 — Digest extractor & thread-spanning** (planned 2026-06-13; NEXT). Product of the digest grill (operator + `doyle`); governed by **ADR-0019** (reverses M9's \"no manifest seam\"), CONTEXT.md \"session digest\" rewritten. An `spt-claude-code` enabler, so it takes the M10 slot ahead of the Shell/remote-attach work (which shift to M11/M12). Adapter-declared **`[digest]` extractor seam** (imperative, no DSL; same source files as `[history]` by default + own-source escape; `api digest-entry` push f...\n 51|\n 52|- **🔜 v0.15.0 — Activity-gated message delivery + send-modifier axes.** Product of the messaging grill (operator + `doyle`, 2026-06-23); governed by **ADR-0028**, CONTEXT.md mints \"activity-gated delivery\" / \"message delivery axes\" / \"message metadata (`json`)\". Closes a **legacy-SPT parity gap that was never written down**: activity-gated inbound routing is scaffolded (`is_idle` + `resolve_inject_methods`) but unwired (`dispatch_endpoint_input` injects unconditionally). Wires the two-window router ...\n 53|\n...\n 57|\n*58|1. **Broker/brain no-terminate handoff** — prove a PTY child + open socket survive a daemon-*logic* restart (FD-passing or stable-broker shim). The hardest invariant in the design (ADR-0004).\n*59|2. **Iroh + mDNS smoke test** — two nodes, separate NATs, one message end-to-end; measure direct-vs-relay + binary size (research brief §5). Validates the whole networking premise.\n*60|3. **ConPTY + `portable-pty`** — host and drive a real session in a broker PTY on Windows + Linux.\n*61|4. **Manifest-driven lifecycle** — a manifest drives one real Claude Code session through spawn → bind → commune → signoff, proving harness-independence works in practice.\n 62|\n 63|## Immediate next steps\n 64|\n...\n 66|- ✅ **Stage A — codex adversarial pass** — done. 14 findings (5 FATAL, 9 SERIOUS) in `docs/reviews/STAGE-A-codex-redteam.md`.\n*67|- ✅ **Spike #1 (broker/brain handoff)** — PASS on Windows ConPTY. `docs/spikes/SPIKE-01-broker-handoff.md`. Found+fixed the ConPTY DSR-stall hazard. 4 open gaps carried to M3.\n*68|- ✅ **Spike #2 (Iroh smoke test)** — PASS (single-host). `docs/spikes/SPIKE-02-iroh-smoke.md`. iroh 0.98 API workable; 13.57 MB release. Cross-NAT/relay proof deferred to M4.\n 69|- ✅ **Findings folded** into ADR-0002/0003/0004/0005 + KNOWN-HAZARDS (+3) + traceable-reqs.toml (+4 REQ-HAZARD).\n 70|\n 71|Remaining before M0:\n 72|1. **Stage A — `plan-eng-review`** (interactive architecture pass; second half of Stage A).\n*73|2. **Spike #4 (manifest-driven lifecycle)** — a manifest drives one real session spawn→bind→commune→signoff. (Spike #3 ConPTY hosting is largely covered by Spike #1.)\n 74|3. Resolve the open user decisions logged in the ADR amendments (context-merge model, seed-rotation, R-UPD-3 rewording).\n 75|4. **M0** (workspace + `spt-proto` + `spt-store`) — **started**; task breakdown in [`M0-PLAN.md`](./M0-PLAN.md). Each milestone gets its own just-in-time plan (this one is M0); the roadmap stays the milestone sequence, the JIT plan is the task layer.\n 76|\n## CONTEXT.md\n 2|\n*3|**Platform scope:** Windows + Linux for v1. macOS is out (no test machine available) but kept structurally easy — `portable-pty` and Iroh both support it, so macOS is a later test/CI-budget decision, not a re-architecture.\n 4|\n 5|**Legacy migration:** it should be possible — ideally *automatic* — for a user to migrate an existing `claude_skill_owl` (modern SPT) install to spt-core (identity, agents, tracked Psyche context). Exact mechanism deferred to design; the commitment is that migration is a first-class supported path, not a manual rebuild.\n 6|\n*7|Harness-independent core for the SPT ecosystem. Provides inter-agent messaging, live-agent lifecycle, terminal wrapping, self-update, and networking primitives — as both a Rust library workspace and a canonical reference binary. Designed so any agent runtime (Claude Code, Codex, Cursor, headless, future harnesses) can interface with the SPT ecosystem either by shelling out to the binary or by linking the crates directly.\n 8|\n 9|Successor to `claude_skill_owl` (today's \"modern SPT\"), which is being rebuilt as `spt-core` to untether the system from Claude Code and lift it to a general-purpose agent-ecosystem core.\n 10|\n...\n 22|**spt plugin** (separate downstream project — NOT an spt-core deliverable):\n*23|A rebuilt version of today's Claude Code `spt` plugin. It is the **first consumer** built *atop* spt-core and the **acceptance proof** of spt-core v1 (it reaches feature parity with modern SPT while delegating all core functionality to spt-core, primarily via `spt.exe`, with deeper hooks where useful) — but it **lives and builds in its own repository, outside spt-core**. It is a Claude-Code-specific *adapter*: it holds the Claude Code conventions (hooks, slash-commands, skill/plugin layout, `claude` se...\n 24|\n 25|**Pi** (disambiguation — two meanings, never conflate):\n 26|(1) **Pi, the coding agent/harness** (`badlogic/pi-mono`) — a harness example alongside Claude Code and Codex; this is the meaning in user-facing harness lists. (2) **Pi-class node** — Raspberry-Pi-class low-power hardware hosting a Shell-only or headless SPT node; an incidental hardware descriptor, never an explicit product example. Public-facing docs must disambiguate or avoid the bare word.\n...\n 35|Internally the logical daemon is split into two implementation layers for seamless self-update (see Self-update):\n*36|- **broker** (stable \"kernel\") — holds *only* the un-transferable, must-not-die resources: PTY master fds, the spawned harness child processes, and listening network sockets. Minimal, dumb, versioned local IPC. Almost never updates.\n*37|- **daemon brain** (\"userspace\") — all logic (routing, registry, pulse/psyche loops, manifest parsing, update orchestration). Restarts freely on update; rehydrates from disk state and re-attaches to the broker's held handles.\n 38|\n*39|Logical addressing is unchanged — still one per-machine `spt-daemon`; the broker is an internal layer, not separately addressable. There is exactly **one broker per machine** (per `SPT_HOME`) — *not* one per endpoint: a single broker holds every hosted endpoint's resources, and it is present whenever the daemon runs, even with zero endpoints online (the bare-daemon case). It is therefore the always-present per-machine layer, which is why the single-daemon lock + liveness anchor belong to it.\n 40|\n 41|**in-session relay**:\n*42|A thin, stateless `spt.exe` task that exists only in **harness-hosted** sessions (where the agent harness is the parent process and spt cannot reach into its process tree — today's Monitor model). It streams the daemon brain's events into the session's stdout. All *stateful* listener logic lives in the daemon; the relay is a dumb pipe, freely killable and respawnable. **spt-hosted** sessions need no separate harness-owned relay — the daemon owns the PTY and consumes the same poll feed itself. Idle de...\n 43|\n 44|### Deliverable shape\n 45|\n...\n 105|<!-- [doc->REQ-UPD-9] -->\n*106|Each adapter manifest declares how spt-core should *ripple-update the adapter itself* (see Self-update). One of: **file-pull** (a plugin-directory lookup regex + a gh repo for the adapter's latest files — spt-core fetches + swaps), **delegated command** (a binary command the adapter owns, e.g. `claude.exe plugin update` — spt-core invokes it), or **gh_release** (the adapter ships its updates from its own GitHub releases). After initial bootstrap, the plugin no longer self-manages updates; spt-core co...\n 107|\n 108|**adapter packaging & live update** (v0.13.2; ADR-0024, ADR-0025):\n 109|<!-- [doc->REQ-ADAPTER-GH-TRANSPORT] -->\n*110|A `.spt` may be **multi-platform**: shared `manifest.toml` + `strings/` at the root, role binaries under per-target-triple subdirectories (`x86_64-pc-windows-msvc/`, …); install/update extracts the shared root plus only the current node's triple, flattened into `install_dir`, so flat `<install_dir>/<program>` resolution is unchanged. It stays one signed asset (`adapter.spt`, plain-tar or gzip); a multi-platform archive missing the recipient's triple is a typed `NoArtifactForPlatform`. Large adapters ma...\n 111|<!-- [doc->REQ-ADAPTER-UPDATE-MESSAGE] -->\n 112|An optional **`[update].message`** (avenue-agnostic) is a plain multi-line operator notice surfaced to stdout, markdown-rendered (the helpfmt prose path), **only when an update is actually applied** (the version changed) — never on a no-op. It is read from the newly-installed manifest with no `{key}` substitution; its use is to announce a post-update action (e.g. \"run `/reload-plugins` in any ongoing sessions\").\n 113|\n...\n 124|\n*125|Governing principle: **SPT is not a harness.** Model choice, billing shape, harness-internal env, and harness-internal context are entirely the adapter's concern, expressed inside the adapter's own command templates. spt-core owns only the template *mechanism* (substitution keys), the substitution *values* it is responsible for, and the surrounding lifecycle. Env for the *endpoint binary itself* is auto-handled by spt-core/broker; env for the *agent running inside* that binary is the adapter's config (e....\n 126|\n*127|**spawn-session seam** — launch a new agent session on this node. Manifest provides: a command template; `cwd`/project; a `headless` flag (optional, default false — for the GUI's resume-of-compatible-adapters); a `resume` flag (optional); and the `commune` + `signoff` file directories relative to `cwd` (so the daemon knows where to watch). Substitution keys spt-core can supply: `{id}` and, optionally, a spt-core-generated valid session UUID (e.g. injected as `--session-id {uuid}`) so an adapter can s...\n 128|- **id resolution:** `id` is optional. With no id, spt-core reproduces today's no-id `/spt:live` behavior — run the lone live agent if that's all the project has; show a picker with proposed default IDs if the project has none; let the user choose if there are several.\n 129|\n 130|**post-spawn seam** — the just-launched binary calls an spt-core command on boot (via the adapter's SessionStart-equivalent hook) to bind itself. Needed because the harness's own session id usually isn't known until after the binary runs. Payload: the harness `session_id` (when binary-generated rather than spt-core-injected); the `parent_pid` (the stable session-binding anchor — see KNOWN-HAZARDS 2.1); an endpoint identity/type confirmation; optionally a local HTTP port the binary listens on (for HTT...\n...\n 145|\n*146|**activity observation** (ruled 2026-07-24, rebound grill) — two avenues, deliberately split by consumer class; the digest is **not** one of them (it stays a content surface). **Push (shells only):** an owned Shell observes its *owner's* busy/idle transitions as an **activity frame** on the existing shell-link event stream — link-scoped (owner implied by the link token), **drive-class semantics** (ephemeral, latest-wins, current-state-carrying; a redundant same-state resend is a harmless no-op — th...\n 147|\n 148|**inject-input seam** — message delivery into a running session. Configurable per activity-state (activity / idle / both); multiple methods, any combination:\n 149|- PTY injection (with or without key/submit sequences) — spt-hosted topology;\n...\n 175|- **Command templates are opaque.** spt-core never parses out a model/tool/flag — the adapter writes the whole command line; spt-core fills substitution keys and runs it.\n*176|- **A command template's program token resolves against the adapter install dir before PATH (since v0.8.0).** A `.spt` adapter ships its built binaries to the adapter's install dir (`adapters/_github/<safe>/` via `--release`/`--github`, or the record's `source_dir` under copy-mode), so a bare program name (e.g. `claude-spt-digest …`) binds to the shipped binary first and falls back to PATH when absent — a `.spt` that ships its binaries is **self-contained**, needing no PATH placement. <!-- [doc->REQ-...\n 177|- **Hook output capability is declared per harness-event** (`can_inject`). CC's Stop hook cannot inject context — that single fact drives the echo-gate sentinel + relay fallback. The manifest expresses it so spt-core knows when to fall back.\n*178|- **Env injection is asymmetric** (file-bridge-only-when-not-launcher, applied to env): spt-hosted sessions inherit env from the broker that spawns them; harness-hosted sessions need the harness's declared env channel. With `spt` on PATH the env table is small.\n 179|- **Cross-adapter fallback** is a **node-wide setting**, not a manifest field: if a Psyche/echo invocation under one adapter is rate-limited, spt-core falls back to another adapter (e.g. `ccs` — its own adapter, not a binary-swap). <!-- [doc->REQ-MANIFEST-6] --> A fallback **target is addressed as `<adapter>:<profile>`** (not just a bare adapter_name) and resolves through the one composite-addressing resolver (`registry::resolve_option`), so a fallback may select a shipped or local profile (`ccs`, `ccs...\n 180|- **Config knobs** (pulse period, echo-commune window/gate, route-guard window, daily refresh) are spt-core **global settings** with optional **per-endpoint override**. **An adapter may DECLARE A DEFAULT, never an override** (narrowed 2026-08-03, LOCKSMITH grill — the original \"never per-adapter\" wording is superseded): a harness has real information about its own turn shape and cost, but the operator keeps the last word. Precedence, highest first: **per-endpoint override → node/global setting → ad...\n 181|- **Event-block vocabulary and file-drop filenames are fixed spt-core constants** (documented for adapter authors), not manifest-configurable. <!-- [doc->REQ-RESUME-CONTEXT-PULL] --> This includes the **checkpoint sentinel `!!checkpoint!!`** — the agent-checkpoint trigger an adapter embeds in a commune/signoff drop body (one bare token = checkpoint with default wake; a `!!checkpoint!! <text> !!checkpoint!!` pair makes the inter-marker text a custom wake directive). It is spt-core control metadata: spt-...\n...\n 199|\n*200|**`spt endpoint purge <id>`** (CLI, not `api`) — the standalone, formal **full teardown**: wipe an endpoint and *every* record keyed on it. It is the dev/CI sibling of `api session-end --erase` (which is adapter-triggered at session end); `purge` is the explicit operator/test command for clean setup-and-reset. **Deliberately NOT consent-gated** — a local dev/test op, never a peer-visible action. **Offline-only**: it refuses a live / daemon-hosted endpoint (deleting records out from under a running ho...\n*201|_Avoid_: consent-gating it (it is intentionally ungated, for CI); treating it as a sync/remote op (local-only); a soft variant (purge is always the hard, full wipe — soft teardown is `endpoint stop`). **Read that soft/hard contrast on the RECORD axis only** (ADR-0045): `stop` is *record*-preserving (spool.db + info.json survive) where `purge` wipes them. It says nothing about processes — on the **process** axis `stop` is hard: for a broker-hosted endpoint it reaps the session and its descendant subtr...\n 202|<!-- [doc->REQ-ENDPOINT-PURGE] -->\n 203|\n 204|- **`api history-log`** — Path B: ingest normalized records into spt-core's native history store.\n*205|- **`api presence`** — adapter reports user interaction → updates the presence datum `(last_active_node, last_active_endpoint, ts)`. In the spt-hosted topology, presence is **also** updated by the broker *detecting* (sensing, not watching/logging) user input on a held PTY — privacy-preserving (it notes that input occurred, records no content).\n*206|- **`api emit --type <sensory_type> <payload>`** — a broker-launched **Shell** binary pushes a sensory payload to its owner agent (owner known from `api bind`; REST-only, never spooled). See the Shell model.\n 207|\n 208|**Not `api` commands — file-drop flow:** `commune` and `signoff` are deprecated as commands (modern SPT) in favor of file drops. The agent/adapter writes `<id>-commune.md` / `<id>-signoff.md`; the daemon watches the manifest-declared commune/signoff dirs (the spawn-session seam fields), ingests, and deletes (drop files are daemon-owned single-writer — KNOWN-HAZARDS 6.4). These stay off the `api` surface and the agent surface alike. **This path is the Self's, and spt-core's echo-commune never writes i...\n 209|\n...\n 228|\n*229|**spt-hosted (terminal wrapper / GUI launcher; the daemon launches the binary into a broker PTY):**\n*230|1. The frontend/CLI launches the agent: the daemon runs the **spawn-session** command template into a broker-held PTY.\n 231|2. The binary boots and fires **`api bind`** (or skips it under the strict UUID-injection + stable-pid commitment). **No catalyst/seed file** — the daemon is the launcher, already holds a direct channel (it spawned the process and owns the PTY), so a file round-trip would only add drive churn for no benefit.\n 232|3. The daemon delivers events; method is **manifest-configurable per activity-state** — direct PTY injection, or a relay even here (some adapters prefer a relay over PTY injection for idle delivery), or HTTP. During *activity*, delivery still defaults to the non-disruptive hook-injection path, not raw PTY writes.\n 233|4. Psyche is daemon-spawned, same as above.\n\n## docs/\n### STORAGE.md\n 43|    project-index.json           endpoint→project attribution (schema-versioned; daemon-written)\n*44|  daemon/                        ── ephemeral ── brain rehydration state, broker handle metadata\n 45|  updates/                       ── ephemeral ── staged binaries pending handoff\n 46|  adapters/                      registered adapter manifests / pointers\n 47|    <adapter>/web/               core-owned adapter served root; output bytes survive deactivation\n...\n 69|  node-sovereign stores — outside `tracked/` and carried by no wire record.\n*70|  The siting *is* the mechanism; a unit test scans the wire-record modules to\n 71|  keep it true.\n 72|\n 73|**Write sites.** Inbound is recorded in the WAN receive funnel **after the\n...\n 173|| subnet registry (snapshot + distribution) | **BranchStore** | audited; rides the precedence merge driver instead of a separate replication scheme (M4) |\n*174|| daemon/brain coarse checkpoint (rehydrate anchor, ADR-0004) | **BranchStore** | commit = checkpoint, ref tip = resume |\n*175|| endpoint lifecycle / status *audit trail* | **BranchStore** (or a derived view) | the *trajectory* matters, not just current value |\n 176|| spool (queue, drain-by-status, `id→address`) | **SQLite** | indexed queries + hot churn |\n*177|| broker effect journal (per-PTY-effect, B5) | **fsync-append** | hot path, per-keystroke — a commit per write is far too heavy |\n 178|| `info.json` / `status` (current value) | flat + atomic-rename (today) | tiny, read-hot; *may* move under BranchStore later, low priority |\n 179|\n 180|**Rules:**\n...\n 185|\n*186|**Reading the effect journal (`<spt_home>/daemon-effects.log`) — operator note.** The broker fsync-appends two line kinds per durable effect: `PENDING <key> <kind>` at attempt, `DONE <key>` on success (durable kinds only — an ephemeral `PtyWrite` writes no lines). A line with **`PENDING` and no matching `DONE`** is **expected**, not a leak: it is a durable effect that did NOT complete — for a `net-send` (a WAN write to an offline/unreachable peer) it is an **unconfirmed send retained as the re-driv...\n 187|\n 188|**Delivery vs propagation — latency envelope (operator note).** Two classes of cross-node traffic have deliberately different latency floors, so a minutes-scale delay is not automatically a defect: (1) **messages** (`spt send`) are **event-driven** — delivered on the events the daemon owns (WAN ingress + the ACTIVE→IDLE edge, F-023), so a healthy path is near-immediate and never waits on an adapter hook-poll cadence; (2) **anti-entropy context-sync** (the mind pull) is **periodic** — it rides the...\n 189|\n...\n 208|**Two transport modes:**\n*209|- **P2P (default — gh off):** peers exchange **incremental `git bundle`s over broker QUIC streams** — **pull-based, ref-scoped**: a node requests `a-<id>` for endpoints it has `synced(E,S)` and `p-<project>` only for projects it hosts (the requester names refs, so two-tier scoping falls out of the pull model); receiver fetches from the bundle + merges with driver + `--allow-unrelated-histories`. **The context DAG is shared across nodes** (one history, one merge commit per sync — divergent per-node ...\n 210|- **Hub (opt-in — gh on, via `spt context-github-setup`):** every node **pull → merge → push** against a private GitHub remote — always-online sync hub *and* live GUI view (no stale-mirror-node problem). Push-retry on non-fast-forward contention. **Prefer-GitHub, P2P backstop** when GitHub/SSH is unreachable (the P2P stack is always present anyway).\n 211|\n 212|**Scope:** hub mode is the **context-repo sync transport only.** Messaging, the subnet registry, remote-drive, pairing, and presence always ride P2P/Iroh — GitHub never carries real-time agent traffic, and gh-on never removes the P2P requirement.\n...\n 274|  one refresh. Fire-and-forget and daemon-offline-safe (events accumulate\n*275|  and are consumed at the next brain start).\n 276|  <!-- [doc->REQ-PROJECT-INDEX-INVALIDATION] -->\n 277|\n 278|## perches/ — node-local runtime\n...\n 303|| `relay` | `api listen` | a process that **holds** the endpoint | the endpoint's relay is gone — the one case relay-death convergence may act on |\n*304|| `binder` | `api bind` (agent or non-agent) | the **announcing CLI**, expected to exit within seconds | **nothing.** The hosting life it announced (a broker PTY session) is not this process, and is not recorded in the row at all |\n 305|| *absent* | any pre-`pid_role` writer | unknown | nothing — **no knowledge**, see below |\n 306|\n 307|**Why the field exists.** The same `pid` field carried both meanings, decided by write path, and every reader had to guess. Measured on one node (2026-07-27): three `claude-spt` live agents were online and working with recorded pids long dead (binder pids), while `omp-spt` and `mobile` endpoints held genuinely live relay pids. Two separate proposed fixes for a dead-relay bug were keyed on bare pid death; both would have offlined those three live agents. A record that states the meaning ends the guessing.\n...\n 312|\n*313|- **Relay-death convergence** (`livehost::reconcile_hosted_liveness`) is reached only by rows with `controllable != Some(true)`, and now acts only on `pid_role == relay`. So: harness relay rows converge on a proven relay death; `binder` rows (every `api bind` product of an agent or non-agent path, including gateway and shell binds) are **never** convergeable no matter how dead their pid; legacy rows are never convergeable; rows with a live broker PTY (`controllable == Some(true)`) never reach the branch ...\n 314|- **`controllable` carry-forward** (`api/startup.rs`) drops an earned `Some(true)` on a listener re-bind **only** when the prior row's `relay` pid is provably `Gone`. A dead `binder` pid, an `Unproven` verdict, and a legacy row all inherit.\n 315|\n 316|Both consumers route their liveness verdict through the single authority (`liveness::relay_liveness` over the `proc::process_identity` oracle); `Unproven` never kills.\n...\n 320|- Records the endpoint's durable identity: type, creation, and an **ordered adapter history** (most-recent first). Synced, so any node knows it.\n*321|- Drives the **resume UX**: the latest adapter is the default resume offer, the next-latest is second, … the oldest is second-to-last, and \"choose a different adapter\" is the final option.\n 322|- The *currently-running* `adapter_name` lives in node-local `perches/<id>/info.json`; the *history* is endpoint identity in `tracked/`.\n 323|\n 324|## Shells in the layout\n...\n 331|- **Adapter/platform-bound, not adapter-agnostic** — capability surface comes from the providing shell adapter; it does not move harnesses.\n*332|- **Lifecycle = online / offline / torn-down.** Link-break always closes the binary (optional pre-close instruction + termination timeout). **Ephemeral** (manifest property) ⇒ perch torn down *and removed from the agent's shell history*; **persistent** ⇒ perch kept offline for relink.\n 333|- **Cross-node links:** the shell perch lives on the shell's node (nested under the owner id as a label there); commands ride Iroh from the owner's node.\n 334|\n 335|## Migration & backup\n### MANIFEST.md\n 180|resume <id>`, `spt go <id>` on an offline endpoint, or the picker's *Resume from\n*181|history*) **and** the adapter declares the role. **Resume is latest-only:** no\n 182|verb takes a session argument — the retired `endpoint run --resume <session>`\n 183|spelling went with `endpoint run`, and a session argument is now *refused*\n 184|rather than silently ignored, because core resolves which session to resume from\n...\n 201|resolve (`claude -r <provisional>` hangs at \"no sessions match\"). A resume\n*202|request that is itself a provisional recovers the latest harness-reported ledger\n 203|row instead. When the endpoint has *no* harness-reported row at all — only a\n 204|spawn provisional on record — spt-core does **not** feed the provisional to your\n 205|template: it prints `RESUME_NO_HARNESS_SESSION:<id>` on stderr and starts a\n...\n 232||---|---|---|\n*233|| `{node}` | this node's **advertised label** — the label store value (OS-hostname default, re-checked at daemon startup), OS-hostname fallback when no store label is known | **SINGLE-TOKEN fill only** — tokenize-then-fill cannot produce a space-carrying argv element, so a composite display name like `<id> @ <node>` stays adapter-shim territory (`{node}` complements the shim for tokenizable args, e.g. `--remote-control {id}--{node}`, it does not replace it). **Caveat:** the daemon-side lifecycle reso...\n 234|\n 235|<!-- [doc->REQ-MANIFEST-NODE-KEY] -->\n 236|```toml\n...\n 252|### `[env]` — env-var table\n*253|Vars to inject/read, and **how**. With `spt` on PATH the table is small. The injection *channel* is asymmetric: **spt-hosted** sessions inherit env from the broker that spawns them (no channel needed); **harness-hosted** sessions need the harness's declared channel (the broker didn't spawn them). File-bridge-only-when-not-launcher, applied to env.\n 254|\n 255|```toml\n 256|[env.OWL_SESSION_ID]\n...\n 258|value = \"{session_id}\"\n*259|channel = \"CLAUDE_ENV_FILE\"        # harness-hosted only; spt-hosted inherits from broker\n 260|```\n 261|\n 262|**`direction = \"read\"` — capture a launch-env var for template substitution.**\n...\n 422|Opt-in. The adapter's **idle-delivery translation binary**: a pure stdin→stdout\n*423|JSON-lines filter spt-core lifecycle-manages (spawned when the spt-hosted endpoint\n 424|comes up, terminated when it goes down). spt-core feeds it the inbound `<EVENT>`\n*425|feed and reads back keystroke-commands, which spt-core applies to the broker-held\n 426|PTY **atomically** — controller input is buffered during the emitted sequence and\n 427|flushed after, so injection coexists with a live `spt rc` controller (spt-core owns\n 428|every PTY write). Idle-only; busy/mid-turn delivery stays adapter hook-injection.\n...\n 444|- **Commit deadline + miss tolerance (F029 C-1, ADR-0022 amendment).** If no `{\"commit\":true}` arrives within the **commit deadline (5 s, `INJECT_COMMIT_DEADLINE`)**, spt-core flushes the buffered operator input (never stranded) and re-spools that message once — but a SINGLE miss does NOT fault a healthy binary (it is preserved; the next event delivers through it). Only **3 consecutive** missed commits, or a real binary death, faults the binary — and a real fault gets a **bounded eager respawn** (a h...\n*445|- **Known keys today:** `command` (preferred) and the deprecated `path`. spt-core does **not** `deny_unknown_fields` here — a newer adapter declaring a future key against an older spt-core parses fine (the unknown key is ignored), so the lifecycle-binary contract degrades gracefully.\n 446|- A minimal binary just emits `{\"text\":payload}{\"key\":\"enter\"}{\"commit\":true}` with no choreography (spt-hosted idle delivery is translation-binary-only since v0.14.3 — no raw-inject fallback).\n 447|\n 448|### `[service]` — a daemon-supervised resident service (ADR-0049)\n...\n 663|- **`--last <N>`** — the last `N` turns (the digest's natural unit). `--last 1` is\n*664|  the latest turn — the turn-end output.\n 665|- **Per-entry `seq`** — every JSON entry carries a **stable, source-derived** `seq`:\n 666|  it is deterministic from the entry's position in the session ledger, so\n 667|  **re-projection yields the SAME `seq`** for the same committed entry. It does\n...\n 704|\n*705|`spt adapter update [name]` (with no name, every registered `gh_release` adapter; with a name, just that one) compares the repo's latest GitHub release version against the installed adapter version and, when newer, fetches the release `.spt` (the same archive primitive as `spt adapter add --release`), then re-extracts and re-registers it in the adapter's durable `_github/<safe>` home (pointer-mode, re-read live). The network fetch lives in the `spt` CLI, never the daemon. **`repo` is required**; `asset` ...\n 706|\n 707|<!-- [doc->REQ-ADAPTER-GH-TRANSPORT] -->\n*708|**Transport — public HTTPS or private `gh`.** The optional **`transport`** selects how the asset bytes + the latest-release version are fetched: `https` (direct, the public-repo path), `gh` (shell the pre-authorized [`gh` CLI](https://cli.github.com/) — `gh release download` for the asset, `gh api` for the version — the **private-repo** path), or **`auto`** (the default: prefer `gh` when it is installed and authenticated, else fall back to HTTPS). Because `gh` honors OAuth + `GH_TOKEN`, an adapter ...\n 709|\n 710|<!-- [doc->REQ-ADAPTER-UPDATE-MESSAGE] -->\n 711|**Post-update notice — `message`.** The optional **`message`** is a plain (multi-line) human notice the adapter surfaces to stdout **only when `spt adapter update` actually applies an update** (the version changed) — never on a no-op / up-to-date run. It is read from the **newly-installed** manifest and rendered through the same inline-Markdown prose path as `spt`'s help (`**bold**`, `` `code` ``, `[text](url)`); there is no `{key}` substitution. It is avenue-agnostic (`gh_release` / `delegated` / `f...\n...\n 723|command = \"{adapter_dir}/claude-spt post-update\"   # REQUIRED, non-empty; install-dir + adapter-static subst\n*724|self_verifies = false                              # optional; attestation-only metadata (default false)\n 725|```\n 726|\n*727|- **Runs unconditionally** — the post-step fires **even when the primary avenue was a no-op** (already up to date). Its own idempotent check (e.g. `claude plugin update`) decides whether anything changes. `command` is **required** and **non-empty** (an empty `command` is refused at registration). `self_verifies` is **attestation-only metadata** (defaults `false`) — parallel to `[update].self_verifies`; it records that the post-step verifies its own outcome but **gates nothing** (the post-step is alre...\n 728|- **Published stdin seam** — spt-core feeds the post-step **one JSON line** on stdin describing the just-resolved update:\n 729|  ```json\n 730|  {\"adapter_applied\": true, \"adapter_name\": \"claude-spt\", \"profile_name\": null,\n...\n 765|[shell]\n*766|spawn = \"gamerobot-shell --link {link_token} --root {perch_dir}\"   # broker-launched; opaque command template — {perch_dir} = the instance's perch dir (REQ-SHELL-PERCH-DIR: lets the binary resolve a shell_file frame's perch-relative path; opt-in, spawn-only). Program token resolves against the adapter's install dir before PATH, and {adapter_dir} fills — see below.\n 767|ephemeral = false                  # manifest property, NOT an agent choice; ephemeral ⇒ no offline perch + no history retention\n 768|broadcast = \"subnet\"               # subnet | same-node | none\n 769|command_receipt = \"http\"           # http | stdin | relay   (how it receives agent commands)\n### CONTEXT-MEMORY.md\n 7|- **live context** (per-agent, project-independent — the agent's *mind*): identity/persona, relationship + rapport with the user, the user's cross-project preferences & working style, cross-project life threads, **learnings from work that may apply to other projects**, a thin **project roster** (one line per project the agent works on), the agent's self-model. Synced to **all** instances of the endpoint.\n\n[Showing lines 1-371 of 381 (50.0KB limit). Use :372 to continue. Some lines truncated to 512 chars]"}],"details":{"scopePath":".worktrees/286-churn/PRD.md, .worktrees/286-churn/ROADMAP.md, .worktrees/286-churn/CONTEXT.md, .worktrees/286-churn/docs/STORAGE.md, .worktrees/286-churn/docs/MANIFEST.md, .worktrees/286-churn/docs/CONTEXT-MEMORY.md","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":86,"fileCount":6,"files":[".worktrees/286-churn/PRD.md",".worktrees/286-churn/ROADMAP.md",".worktrees/286-churn/CONTEXT.md",".worktrees/286-churn/docs/STORAGE.md",".worktrees/286-churn/docs/MANIFEST.md",".worktrees/286-churn/docs/CONTEXT-MEMORY.md"],"fileMatches":[{"path":".worktrees/286-churn/PRD.md","count":18},{"path":".worktrees/286-churn/ROADMAP.md","count":20},{"path":".worktrees/286-churn/CONTEXT.md","count":20},{"path":".worktrees/286-churn/docs/STORAGE.md","count":12},{"path":".worktrees/286-churn/docs/MANIFEST.md","count":14},{"path":".worktrees/286-churn/docs/CONTEXT-MEMORY.md","count":2}],"truncated":true,"perFileLimitReached":20,"displayContent":"# .worktrees/286-churn/\n## PRD.md\n   9│\n *10│spt-core is a **harness-independent core for an agent ecosystem**: inter-agent messaging, live-agent lifecycle, terminal hosting, seamless self-update, and zero-config cross-machine networking — shipped as both a Rust library workspace and a single canonical binary (`spt` / `spt.exe`).\n  11│\n  12│It is the ground-up rebuild of `claude_skill_owl` (\"modern SPT\"), untethering the system from Claude Code so that **any** agent runtime — Claude Code, Codex, Cursor, headless, or novel surfaces — can participate by either shelling out to the binary or linking the crates. The first thing built atop it is a rebuilt `spt` plugin reaching parity with today's system; the networking, terminal-wrapper, and multi-instance capabilities then make an agent reachable and operable across all of a user's machines.\n  13│\n    │...\n  16│- Today's modern SPT works but is **fragile** — each release introduces 2–5 regressions because it grew rapidly from a narrow starting point and carries the resulting tech debt (ADR-0001).\n *17│- It is **welded to Claude Code**. Other harnesses can't reuse the messaging/lifecycle substrate.\n  18│- Networking, terminal hosting (\"capsule\"), and cross-machine operation were planned piecemeal atop a CC-tied core; they deserve to be first-class in a clean foundation.\n  19│\n  20│## 3. Goals\n    │...\n  31│\n *32│- **macOS support** — out (no test machine); kept structurally easy, not a re-architecture.\n  33│- **Cross-subnet** (two *different* users' agents communicating) — deferred concept.\n  34│- **Concrete Shells / PresenceChannel implementation / presence gossip** — only the seams ship (§Endpoints).\n  35│- **Remote command execution on another user's node** and **instantiate-anywhere** — seam-compatible, deferred behind a consent/security gate.\n    │...\n  54│- **R-ARCH-3** `spt-proto` carries an explicit **wire-protocol version** independent of crate semver, with a documented **N-1 compat window**. Binary and library consumers interoperate by proto version, never by release version.\n *55│- **R-ARCH-4** Copy-verbatim the commodity layer from the sister project (envelope grammar, spool DDL, registry RPC, `info.json` shape); clean-room only the architecture/lifecycle (ADR-0001).\n  56│\n *57│### 7.2 The daemon (broker/brain)\n  58│- **R-DAEMON-1** Exactly **one logical `spt-daemon` per machine**, always-on, owning PTYs, network identity + WAN endpoint, subnet registry, all spools, all poll-listener logic, and all Psyche/pulse loops. No separate listener or wrapper processes (ADR-0004).\n *59│- **R-DAEMON-2** Internally split into a stable **broker** (holds PTY masters, harness child processes, listening sockets; versioned IPC; rarely updates) and an **updatable brain** (all logic; rehydrates from disk + re-attaches to broker handles).\n  60│- **R-DAEMON-3** Any `spt api` invocation that needs the daemon **auto-starts it** if absent.\n  61│- **R-DAEMON-4** Honor every invariant in [`docs/KNOWN-HAZARDS.md`](./docs/KNOWN-HAZARDS.md) as a conformance checklist.\n  62│\n    │...\n  88│- **R-START-2** **Harness-hosted:** SessionStart → `spt api seed --pid` (in-memory seed record) → `$SPT listen <id>` consumes the seed by `parent_pid` (validated vs `session_id`), binds, relays. No file.\n *89│- **R-START-3** **spt-hosted:** daemon runs spawn-session into a broker PTY → binary `api bind` (or skips under commitment) → daemon delivers via manifest-configured method. Principle: **file-bridge only when the daemon is not the launcher.**\n  90│- **R-START-4** Adapters inject env aliases (`$SPT` = `spt api`, heirs to `$OWL`/`$LIVE`); spt-core supplies the subcommands.\n  91│\n  92│---\n    │...\n  99│- **R-EP-3** Messaging payloads carry **typed operation commands + arbitrary file blobs** (text/audio/image/video), not just text — the seam that lets Shells, file transfer, and rich PresenceChannel exist later.\n*100│- **R-EP-4** PresenceChannel (deferred impl) is a **broker** endpoint with dispatch/bind/thread styles; presence datum `(last_active_node, last_active_endpoint, ts)`.\n 101│- **R-EP-5** (deferred impl, CONTEXT \"Shell model\") **Concrete shell instantiation model:** `shell spawn` *mints* an owner-exclusive instance (`<adapter>-<n>`, distinct from the `relink`/`persistent`/`wake` online switch); permission to instantiate = shell adapter **registered on the node** (broadcast governs *discovery* only); a per-shell **`require_approval`** gate (`none`|`remembered`|`always`, default none; manifest floor, node/endpoint may tighten) reuses the consent grant store; an optional **`max_...\n 102│\n 103│### 9.2 Instances (ADR-0003, as amended)\n    │...\n 153│\n*154│- **R-UPD-1** **Peer-propagated** over P2P (layered on self-fetch; out-of-band still supported); the subnet self-heals to latest (ADR-0004).\n 155│- **R-UPD-2** **All binaries signature-verified before handoff**, regardless of source (spt-core release key).\n*156│- **R-UPD-3** **No-endpoint-termination invariant, scoped by update class (ADR-0004 §A).** For the **brain-only** class — the routine case — the guarantee is **absolute**: no endpoint process terminates or suspends, satisfied by the broker/brain split (the broker keeps held PTYs/children/sockets alive across brain replacement). **broker-compatible** swaps the broker behind the versioned IPC with endpoints still surviving. **broker-breaking** (held-resource-type / OS-API change) MAY suspend endpoints...\n 157│- **R-UPD-4** **Not auto by default** — gated on user confirmation delivered to the most-recently-active live session, with an opt-in full-auto.\n 158│- **R-UPD-5** spt-core conducts **adapter ripple-updates** after self-update, per each adapter's update declaration and `min_spt_core_version` gate.\n 159│\n    │...\n 163│\n*164│- **R-TERM-1** A process-supervisor terminal wrapper (supersedes \"capsule\") hosting sessions in broker PTYs, multiplexed and name-addressable, attach/detach, headless↔headed.\n 165│- **R-TERM-2** A **session-surface** abstraction with native PTY as the day-one impl and network-attached surfaces for remote control; input injection at both `send-keys` (raw) and `send-line` (cooked) granularity; in-memory scrollback ring (disk spillover deferred).\n 166│- **R-TERM-3** Remote terminal streaming uses a **byte-stream** model for v1 (simple, faithful; near-LAN on Iroh-direct paths). Keep the surface abstraction generic enough that a predictive/state-sync \"feels-local\" layer can be added later without re-architecting — not built v1.\n*167│- **R-TERM-4** (M3, ADR-0008) **Live activity buffer (PTY digest):** a rolling, *parsed*, structured (source-tagged) view of recent PTY I/O (last ~N user turns + agent output between; tool sprints collapsed), built by running **adapter-supplied** manifest patterns (`input_pattern`, `agent_pattern`, `tool_pattern[]` + catchall) over the **broker-owned** PTY bytes — honors the no-built-in-parser rule (adapter owns the pattern). **Spt-hosted-primary** (needs the broker PTY); **harness-hosted is capability...\n 168│- **R-FRONT-1** Day-one headed frontend is a **launcher/manager** (not raw attach): list running + historic endpoints; launch a historic endpoint from its original manifest; tap into a running one; init a new one via a known adapter. Realizes the \"guided resume\" (XMB-style) and \"management GUI\" (per-endpoint panes) sketches. Its **CLI sibling** is a no-arg **`spt resume`** guided-resume picker — endpoints grouped by locality (on-node/this-project → on-node/other-project → off-node), MRU within grou...\n 169│\n 170│---\n    │...\n 176│- **R-INSTALL-3** First run idempotent + interactive-optional: generate node identity, start daemon; pairing + subnet-naming are separate explicit steps.\n*177│- **R-INSTALL-4** (CONTEXT \"adapter registration\") **Adapter registration lifecycle:** `spt adapter add <path>` (or `--github <user/repo>`, manifest-first then install via the declared `[update]` avenue — *install is the first update*) validates the manifest and records it under `{SPT_HOME}/…/adapters/` (copy for `file_pull`, pointer for `delegated`); one command/dir for `kind=\"harness\"` and `kind=\"shell\"`. Node-local (drive/launch capability, not subnet advertisement). `spt adapter remove` is **soft...\n 178│- **R-MIGRATE-1** Standalone install **auto-detects a `claude_skill_owl` install and offers migration** (identity, agents, tracked Psyche context). First-class supported path.\n 179│\n 180│---\n    │...\n 189│- **R-DOCS-4** **Agent-consumable layer:** CI-generated `llms.txt`/`llms-full.txt`, markdown content negotiation, a machine-readable manifest JSON Schema at a stable path, an **MCP doc/resource server** exposing the harness contract as deterministic resources, and structured CLI help (`spt <cmd> --help` + a `--json` mode).\n*190│- **R-DOCS-5** **Anti-drift via CI:** rustdoc, the manifest schema, the agent exports, and CLI help are generated + checked in CI — doc quality on the same footing as tests.\n 191│\n 192│---\n 193│\n    │...\n 195│\n*196│**In v1:** harness-independent core (library + binary); manifest + `api` harness contract; both startup topologies; one-daemon broker/brain; LiveAgent/ReadyAgent/Psyche/Worker/SptNode endpoints + Shell-vs-agent seam + typed/binary payloads; multi-instance data model + subnet registry + resolution policy + dormant/suspended + two-tier context sync + remote-drive of running instances + off-node file transfer; first-class WAN networking + TOTP-SPAKE2 pairing + subnet naming; cross-node Psyche sync (replaces...\n 197│\n 198│**Deferred (seam-compatible):** see [`docs/DEFERRED.md`](./docs/DEFERRED.md) — concrete Shells, PresenceChannel impl, presence gossip, instantiate-anywhere + consent, remote command exec + security gate, sidecar adapters, manifest includes, scrollback spillover, predictive/state-sync terminal layer, OS-level presence signal, macOS.\n 199│\n    │...\n 201│\n*202│1. A non-CC harness (or a headless spike) drives the full live-agent lifecycle through manifest + `api` alone — proving harness independence.\n 203│2. The rebuilt spt-plugin reaches **feature parity** with frozen `claude_skill_owl`.\n 204│3. Two of a user's machines, freshly paired via one TOTP code, exchange messages and the user drives an agent on machine A from machine B — zero config beyond pairing.\n 205│4. A self-update rolls across the subnet **without terminating any endpoint**.\n 206│5. The same agent's mind follows the user across machines (two-tier sync); switching machines leaves the prior instance dormant and catches up the new active one.\n*207│6. Every `KNOWN-HAZARDS` invariant has a corresponding passing test.\n 208│\n 209│## 17. Open design-phase questions\n 210│\n    │...\n 222│- [`docs/adr/0001`](./docs/adr/0001-clean-room-fork-from-claude-skill-owl.md) … [`0005`](./docs/adr/0005-totp-seeded-spake2-node-pairing.md) — decisions.\n*223│- [`docs/KNOWN-HAZARDS.md`](./docs/KNOWN-HAZARDS.md) — invariants harvested from the sister project (test checklist).\n 224│- [`docs/DEFERRED.md`](./docs/DEFERRED.md) — parked-but-committed features.\n 225│- [`docs/STORAGE.md`](./docs/STORAGE.md) — `$SPT_HOME` on-disk layout.\n 226│- [`docs/MANIFEST.md`](./docs/MANIFEST.md) — runtime manifest schema + worked `claude-spt` example.\n## ROADMAP.md\n  6│\n *7│A harness-independent core for an agent ecosystem — messaging, live-agent lifecycle, terminal hosting, **first-class P2P networking**, and a runtime-manifest harness contract — shipped as Rust crates + a single `spt` binary. The proof: the same agent runs natively across all of a user's machines with a synced mind, reachable and operable from anywhere, with seamless self-update and zero-config cross-machine messaging. First thing built atop it: a rebuilt `spt` plugin at parity with today's modern SPT.\n  8│\n  9│## Scope decisions (locked)\n 10│\n   │...\n 16│\n*17│**Stage A — Red-team review (before building).** Pressure-test the design adversarially. Focus on the riskiest architectural decisions, not bikeshedding. Tools: `plan-eng-review` (architecture) + `codex` (adversarial second opinion). Scope-cutting is *not* a goal (v1 is locked as full PRD), but surfacing contradictions / over-engineering / unproven assumptions is.\n 18│\n 19│**Stage B — De-risk spikes.** Throwaway proofs of the load-bearing unknowns — anything that, if it fails, forces a redesign. Do these before committing to the full build. See *De-risk spikes* below.\n 20│\n   │...\n 28│- **M1 — Local messaging + binary.** `spt-msg` (delivery TCP+spool, routing, send/ring/ready) + the `spt` binary + CLI. **= the killer quickstart (local): two agents exchange a message.** Task breakdown: [`M1-PLAN.md`](./M1-PLAN.md).\n*29│- **M2 — Harness contract + lifecycle.** **Strictly harness-agnostic — zero Claude Code conventions in this repo.** spt-core defines the adapter contract; the only adapter-shaped artifact in-tree is a generic mock/test adapter (manifest fixture + fake `api` caller) that exercises the contract (PRD R-DOCS-2 dev-agent quickstart). **The rebuilt `spt` plugin (Claude Code) is a SEPARATE downstream project** built atop these crates/binary — spt-core v1's acceptance proof (PRD §\"feature parity\"), not an...\n 30│  - **M2a — Harness contract. ✅ delivered (2026-06-01).** `spt-runtime` (AgentRuntime, ManifestRuntime, manifest schema) + the `api` subcommand surface (seed/listen/bind/state/echo-gate/poll/worker-*/boundary/session-end/presence/history-log/emit/capability) + the **harness-hosted** startup topology (seed→listen→bind, interim seed-file) + local-api-auth + the generic mock adapter + contract E2E. All M2a reqs activated + green. Task breakdown: [`M2a-PLAN.md`](./M2a-PLAN.md).\n*31│  - **M2b — Live-agent lifecycle. ✅ delivered (2026-06-01).** `spt-live` (fifth layer): history subsystem (fetcher/locate-normalize/native) + spawn-psyche (nested perch, spt-core-owned `$psyche_prompt`) + echo-commune (history brief → single-writer commune drop) + drop-file ingest (single-writer 6.4 + direct-write precedence 6.5) + interim pulse + signoff/boundary ordering invariants (echo-before-signoff 3.3, grace-before-signoff 1.1, stale-signoff sweep 3.2) + resume seam (continue-existing / fres...\n*32│  - **Interim, no daemon** (like M1): the `api listen` process holds the perch + relays; the seed bridges via a file. **The spt-hosted topology (PTY-launch) + the consolidated `spt-daemon` move to M3** (they need the `spt-term` broker) — this narrows the earlier \"both topologies in M2\" wording.\n*33│- **M3 — Terminal wrapper + daemon + self-update.** `spt-term` (session-surface, PTY, broker) + the consolidated `spt-daemon` (broker/brain split) + peer-propagated signed self-update. Split M3a/M3b/M3c. Scope + carried-forward Spike #1 gaps: [`M3-PLAN.md`](./M3-PLAN.md).\n 34│  - **M3a — `spt-term` session-surface mechanism. ✅ delivered (2026-06-02).** The sixth crate (`…→spt-msg→spt-term`, sibling of spt-runtime): the `SessionSurface` trait + native PTY backends (ConPTY/forkpty via `portable-pty`, REQ-TERM-1) · ConPTY-DSR auto-answer drain pump (REQ-HAZARD-CONPTY-DSR) · send-keys/send-line injection (REQ-TERM-2) · bounded byte-stream (REQ-TERM-3) · the PTY-digest parser primitive (REQ-TERM-4 impl/unit, ADR-0008). OS-neutral per Spike #4/#5 (no OS's stream contr...\n*35│  - **M3b — `spt-daemon` broker/brain process. ✅ delivered (2026-06-03).** The seventh crate (`…→spt-live→spt-daemon→spt`) hosts many `spt-term` surfaces behind a **versioned local IPC** (named-pipe/UDS behind one `DaemonTransport` trait, forward-compatible frame schema): a **broker** (stable kernel — PTY masters, harness children, sockets; no logic) + a **brain** (restartable logic) so a brain kill/restart leaves the hosted PTY child + its output stream **intact and gapless** (Spike #1 mad...\n*36│  - **M3c — signed self-update. ✅ delivered (2026-06-03).** The gated, signed, ripple-capable self-update engine (no new crate — it lives in `spt-daemon`, signing primitives in `release.rs`): the **update-class taxonomy** (brain-only / broker-compatible / broker-breaking, ADR-0004 §A) with the **brain-only zero-interruption swap** that drives the M3b handoff substrate (snapshot→drop→re-attach) so a live PTY child + its output stream survive the logic swap **untouched** (REQ-UPD-3, proven E2E `...\n*37│  - **M3 COMPLETE** (M3a + M3b + M3c all ✅, 2026-06-03): the real ADR-0004 architecture is live — native PTY hosting, the consolidated broker/brain daemon, and seamless signed self-update. Next milestone: M4.\n*38│- **M4 — Networking + instances. ✅ delivered (2026-06-04).** The eighth crate `spt-net` (Iroh WAN endpoint bound to the node's own Ed25519 identity, mDNS LAN discovery, TOTP-SPAKE2 pairing over a dedicated pre-trust ALPN with transcript binding / rate limiting / seed rotation+transfer, subnet registry + NDJSON replication under the per-node epoch lease) · the multi-instance model (registry rows + bare-id resolution refuse-and-qualify + visibility/sync-membership gates + rename/collision + **resting-...\n*39│- **M5 — Shells, presence, deferred capabilities. ✅ delivered (2026-06-04;** scope locked with user 2026-06-04, see [`M5-PLAN.md`](./M5-PLAN.md) + per-task JIT plans D0–D9). The consent framework seam (grant store + interactive escalation + pre-consent flags; remote-exec / instantiate-anywhere reserved-but-refusing) · adapter registration lifecycle (`spt adapter add/remove`, manifest-first) · **shell hosting machinery + sleep/wake** (link-token channels with per-frame MAC, broker-launched binarie...\n 40│\n 41│## Post-v1-core milestones\n 42│\n   │...\n 46│\n*47│- **✅ Broker/brain process-isolation restoration — delivered (2026-06-11).** Design: [`docs/BROKER-BRAIN-SPLIT-RESTORATION.md`](./docs/BROKER-BRAIN-SPLIT-RESTORATION.md) (verified-with-amendments by `doyle`); decision: **ADR-0018**; build plan: [`RESTORATION-PLAN.md`](./RESTORATION-PLAN.md) (D1–D7). Corrects an unintended regression — the broker/brain split (ADR-0004) was specced + spiked as two *processes*, but the production daemon runs the broker as an in-process *thread* (`daemon.rs:165-170`)...\n 48│\n 49│- **🔜 NEXT milestone (planned 2026-06-12) — M9: Adapter customization & session surfaces.** Product of the two-day **Gateway grill** (operator + `doyle`; CONTEXT.md mints @ `8dc22ff`/`bd469b9`; ADR-0008 amended; the lecturn device design parked in the sibling `spt-lecturn/` repo for a later build). Shell-free, independently-shippable features: **adapter profiles** (sparse leaf-replace overlays on the parent manifest; shipped + local sources, local surviving adapter updates; composite `<adapter>:<pro...\n*50│- **🔜 M10 — Digest extractor & thread-spanning** (planned 2026-06-13; NEXT). Product of the digest grill (operator + `doyle`); governed by **ADR-0019** (reverses M9's \"no manifest seam\"), CONTEXT.md \"session digest\" rewritten. An `spt-claude-code` enabler, so it takes the M10 slot ahead of the Shell/remote-attach work (which shift to M11/M12). Adapter-declared **`[digest]` extractor seam** (imperative, no DSL; same source files as `[history]` by default + own-source escape; `api digest-entry` push f...\n 51│\n 52│- **🔜 v0.15.0 — Activity-gated message delivery + send-modifier axes.** Product of the messaging grill (operator + `doyle`, 2026-06-23); governed by **ADR-0028**, CONTEXT.md mints \"activity-gated delivery\" / \"message delivery axes\" / \"message metadata (`json`)\". Closes a **legacy-SPT parity gap that was never written down**: activity-gated inbound routing is scaffolded (`is_idle` + `resolve_inject_methods`) but unwired (`dispatch_endpoint_input` injects unconditionally). Wires the two-window router ...\n 53│\n   │...\n 57│\n*58│1. **Broker/brain no-terminate handoff** — prove a PTY child + open socket survive a daemon-*logic* restart (FD-passing or stable-broker shim). The hardest invariant in the design (ADR-0004).\n*59│2. **Iroh + mDNS smoke test** — two nodes, separate NATs, one message end-to-end; measure direct-vs-relay + binary size (research brief §5). Validates the whole networking premise.\n*60│3. **ConPTY + `portable-pty`** — host and drive a real session in a broker PTY on Windows + Linux.\n*61│4. **Manifest-driven lifecycle** — a manifest drives one real Claude Code session through spawn → bind → commune → signoff, proving harness-independence works in practice.\n 62│\n 63│## Immediate next steps\n 64│\n   │...\n 66│- ✅ **Stage A — codex adversarial pass** — done. 14 findings (5 FATAL, 9 SERIOUS) in `docs/reviews/STAGE-A-codex-redteam.md`.\n*67│- ✅ **Spike #1 (broker/brain handoff)** — PASS on Windows ConPTY. `docs/spikes/SPIKE-01-broker-handoff.md`. Found+fixed the ConPTY DSR-stall hazard. 4 open gaps carried to M3.\n*68│- ✅ **Spike #2 (Iroh smoke test)** — PASS (single-host). `docs/spikes/SPIKE-02-iroh-smoke.md`. iroh 0.98 API workable; 13.57 MB release. Cross-NAT/relay proof deferred to M4.\n 69│- ✅ **Findings folded** into ADR-0002/0003/0004/0005 + KNOWN-HAZARDS (+3) + traceable-reqs.toml (+4 REQ-HAZARD).\n 70│\n 71│Remaining before M0:\n 72│1. **Stage A — `plan-eng-review`** (interactive architecture pass; second half of Stage A).\n*73│2. **Spike #4 (manifest-driven lifecycle)** — a manifest drives one real session spawn→bind→commune→signoff. (Spike #3 ConPTY hosting is largely covered by Spike #1.)\n 74│3. Resolve the open user decisions logged in the ADR amendments (context-merge model, seed-rotation, R-UPD-3 rewording).\n 75│4. **M0** (workspace + `spt-proto` + `spt-store`) — **started**; task breakdown in [`M0-PLAN.md`](./M0-PLAN.md). Each milestone gets its own just-in-time plan (this one is M0); the roadmap stays the milestone sequence, the JIT plan is the task layer.\n 76│\n## CONTEXT.md\n   2│\n  *3│**Platform scope:** Windows + Linux for v1. macOS is out (no test machine available) but kept structurally easy — `portable-pty` and Iroh both support it, so macOS is a later test/CI-budget decision, not a re-architecture.\n   4│\n   5│**Legacy migration:** it should be possible — ideally *automatic* — for a user to migrate an existing `claude_skill_owl` (modern SPT) install to spt-core (identity, agents, tracked Psyche context). Exact mechanism deferred to design; the commitment is that migration is a first-class supported path, not a manual rebuild.\n   6│\n  *7│Harness-independent core for the SPT ecosystem. Provides inter-agent messaging, live-agent lifecycle, terminal wrapping, self-update, and networking primitives — as both a Rust library workspace and a canonical reference binary. Designed so any agent runtime (Claude Code, Codex, Cursor, headless, future harnesses) can interface with the SPT ecosystem either by shelling out to the binary or by linking the crates directly.\n   8│\n   9│Successor to `claude_skill_owl` (today's \"modern SPT\"), which is being rebuilt as `spt-core` to untether the system from Claude Code and lift it to a general-purpose agent-ecosystem core.\n  10│\n    │...\n  22│**spt plugin** (separate downstream project — NOT an spt-core deliverable):\n *23│A rebuilt version of today's Claude Code `spt` plugin. It is the **first consumer** built *atop* spt-core and the **acceptance proof** of spt-core v1 (it reaches feature parity with modern SPT while delegating all core functionality to spt-core, primarily via `spt.exe`, with deeper hooks where useful) — but it **lives and builds in its own repository, outside spt-core**. It is a Claude-Code-specific *adapter*: it holds the Claude Code conventions (hooks, slash-commands, skill/plugin layout, `claude` se...\n  24│\n  25│**Pi** (disambiguation — two meanings, never conflate):\n  26│(1) **Pi, the coding agent/harness** (`badlogic/pi-mono`) — a harness example alongside Claude Code and Codex; this is the meaning in user-facing harness lists. (2) **Pi-class node** — Raspberry-Pi-class low-power hardware hosting a Shell-only or headless SPT node; an incidental hardware descriptor, never an explicit product example. Public-facing docs must disambiguate or avoid the bare word.\n    │...\n  35│Internally the logical daemon is split into two implementation layers for seamless self-update (see Self-update):\n *36│- **broker** (stable \"kernel\") — holds *only* the un-transferable, must-not-die resources: PTY master fds, the spawned harness child processes, and listening network sockets. Minimal, dumb, versioned local IPC. Almost never updates.\n *37│- **daemon brain** (\"userspace\") — all logic (routing, registry, pulse/psyche loops, manifest parsing, update orchestration). Restarts freely on update; rehydrates from disk state and re-attaches to the broker's held handles.\n  38│\n *39│Logical addressing is unchanged — still one per-machine `spt-daemon`; the broker is an internal layer, not separately addressable. There is exactly **one broker per machine** (per `SPT_HOME`) — *not* one per endpoint: a single broker holds every hosted endpoint's resources, and it is present whenever the daemon runs, even with zero endpoints online (the bare-daemon case). It is therefore the always-present per-machine layer, which is why the single-daemon lock + liveness anchor belong to it.\n  40│\n  41│**in-session relay**:\n *42│A thin, stateless `spt.exe` task that exists only in **harness-hosted** sessions (where the agent harness is the parent process and spt cannot reach into its process tree — today's Monitor model). It streams the daemon brain's events into the session's stdout. All *stateful* listener logic lives in the daemon; the relay is a dumb pipe, freely killable and respawnable. **spt-hosted** sessions need no separate harness-owned relay — the daemon owns the PTY and consumes the same poll feed itself. Idle de...\n  43│\n  44│### Deliverable shape\n  45│\n    │...\n 105│<!-- [doc->REQ-UPD-9] -->\n*106│Each adapter manifest declares how spt-core should *ripple-update the adapter itself* (see Self-update). One of: **file-pull** (a plugin-directory lookup regex + a gh repo for the adapter's latest files — spt-core fetches + swaps), **delegated command** (a binary command the adapter owns, e.g. `claude.exe plugin update` — spt-core invokes it), or **gh_release** (the adapter ships its updates from its own GitHub releases). After initial bootstrap, the plugin no longer self-manages updates; spt-core co...\n 107│\n 108│**adapter packaging & live update** (v0.13.2; ADR-0024, ADR-0025):\n 109│<!-- [doc->REQ-ADAPTER-GH-TRANSPORT] -->\n*110│A `.spt` may be **multi-platform**: shared `manifest.toml` + `strings/` at the root, role binaries under per-target-triple subdirectories (`x86_64-pc-windows-msvc/`, …); install/update extracts the shared root plus only the current node's triple, flattened into `install_dir`, so flat `<install_dir>/<program>` resolution is unchanged. It stays one signed asset (`adapter.spt`, plain-tar or gzip); a multi-platform archive missing the recipient's triple is a typed `NoArtifactForPlatform`. Large adapters ma...\n 111│<!-- [doc->REQ-ADAPTER-UPDATE-MESSAGE] -->\n 112│An optional **`[update].message`** (avenue-agnostic) is a plain multi-line operator notice surfaced to stdout, markdown-rendered (the helpfmt prose path), **only when an update is actually applied** (the version changed) — never on a no-op. It is read from the newly-installed manifest with no `{key}` substitution; its use is to announce a post-update action (e.g. \"run `/reload-plugins` in any ongoing sessions\").\n 113│\n    │...\n 124│\n*125│Governing principle: **SPT is not a harness.** Model choice, billing shape, harness-internal env, and harness-internal context are entirely the adapter's concern, expressed inside the adapter's own command templates. spt-core owns only the template *mechanism* (substitution keys), the substitution *values* it is responsible for, and the surrounding lifecycle. Env for the *endpoint binary itself* is auto-handled by spt-core/broker; env for the *agent running inside* that binary is the adapter's config (e....\n 126│\n*127│**spawn-session seam** — launch a new agent session on this node. Manifest provides: a command template; `cwd`/project; a `headless` flag (optional, default false — for the GUI's resume-of-compatible-adapters); a `resume` flag (optional); and the `commune` + `signoff` file directories relative to `cwd` (so the daemon knows where to watch). Substitution keys spt-core can supply: `{id}` and, optionally, a spt-core-generated valid session UUID (e.g. injected as `--session-id {uuid}`) so an adapter can s...\n 128│- **id resolution:** `id` is optional. With no id, spt-core reproduces today's no-id `/spt:live` behavior — run the lone live agent if that's all the project has; show a picker with proposed default IDs if the project has none; let the user choose if there are several.\n 129│\n 130│**post-spawn seam** — the just-launched binary calls an spt-core command on boot (via the adapter's SessionStart-equivalent hook) to bind itself. Needed because the harness's own session id usually isn't known until after the binary runs. Payload: the harness `session_id` (when binary-generated rather than spt-core-injected); the `parent_pid` (the stable session-binding anchor — see KNOWN-HAZARDS 2.1); an endpoint identity/type confirmation; optionally a local HTTP port the binary listens on (for HTT...\n    │...\n 145│\n*146│**activity observation** (ruled 2026-07-24, rebound grill) — two avenues, deliberately split by consumer class; the digest is **not** one of them (it stays a content surface). **Push (shells only):** an owned Shell observes its *owner's* busy/idle transitions as an **activity frame** on the existing shell-link event stream — link-scoped (owner implied by the link token), **drive-class semantics** (ephemeral, latest-wins, current-state-carrying; a redundant same-state resend is a harmless no-op — th...\n 147│\n 148│**inject-input seam** — message delivery into a running session. Configurable per activity-state (activity / idle / both); multiple methods, any combination:\n 149│- PTY injection (with or without key/submit sequences) — spt-hosted topology;\n    │...\n 175│- **Command templates are opaque.** spt-core never parses out a model/tool/flag — the adapter writes the whole command line; spt-core fills substitution keys and runs it.\n*176│- **A command template's program token resolves against the adapter install dir before PATH (since v0.8.0).** A `.spt` adapter ships its built binaries to the adapter's install dir (`adapters/_github/<safe>/` via `--release`/`--github`, or the record's `source_dir` under copy-mode), so a bare program name (e.g. `claude-spt-digest …`) binds to the shipped binary first and falls back to PATH when absent — a `.spt` that ships its binaries is **self-contained**, needing no PATH placement. <!-- [doc->REQ-...\n 177│- **Hook output capability is declared per harness-event** (`can_inject`). CC's Stop hook cannot inject context — that single fact drives the echo-gate sentinel + relay fallback. The manifest expresses it so spt-core knows when to fall back.\n*178│- **Env injection is asymmetric** (file-bridge-only-when-not-launcher, applied to env): spt-hosted sessions inherit env from the broker that spawns them; harness-hosted sessions need the harness's declared env channel. With `spt` on PATH the env table is small.\n 179│- **Cross-adapter fallback** is a **node-wide setting**, not a manifest field: if a Psyche/echo invocation under one adapter is rate-limited, spt-core falls back to another adapter (e.g. `ccs` — its own adapter, not a binary-swap). <!-- [doc->REQ-MANIFEST-6] --> A fallback **target is addressed as `<adapter>:<profile>`** (not just a bare adapter_name) and resolves through the one composite-addressing resolver (`registry::resolve_option`), so a fallback may select a shipped or local profile (`ccs`, `ccs...\n 180│- **Config knobs** (pulse period, echo-commune window/gate, route-guard window, daily refresh) are spt-core **global settings** with optional **per-endpoint override**. **An adapter may DECLARE A DEFAULT, never an override** (narrowed 2026-08-03, LOCKSMITH grill — the original \"never per-adapter\" wording is superseded): a harness has real information about its own turn shape and cost, but the operator keeps the last word. Precedence, highest first: **per-endpoint override → node/global setting → ad...\n 181│- **Event-block vocabulary and file-drop filenames are fixed spt-core constants** (documented for adapter authors), not manifest-configurable. <!-- [doc->REQ-RESUME-CONTEXT-PULL] --> This includes the **checkpoint sentinel `!!checkpoint!!`** — the agent-checkpoint trigger an adapter embeds in a commune/signoff drop body (one bare token = checkpoint with default wake; a `!!checkpoint!! <text> !!checkpoint!!` pair makes the inter-marker text a custom wake directive). It is spt-core control metadata: spt-...\n    │...\n 199│\n*200│**`spt endpoint purge <id>`** (CLI, not `api`) — the standalone, formal **full teardown**: wipe an endpoint and *every* record keyed on it. It is the dev/CI sibling of `api session-end --erase` (which is adapter-triggered at session end); `purge` is the explicit operator/test command for clean setup-and-reset. **Deliberately NOT consent-gated** — a local dev/test op, never a peer-visible action. **Offline-only**: it refuses a live / daemon-hosted endpoint (deleting records out from under a running ho...\n*201│_Avoid_: consent-gating it (it is intentionally ungated, for CI); treating it as a sync/remote op (local-only); a soft variant (purge is always the hard, full wipe — soft teardown is `endpoint stop`). **Read that soft/hard contrast on the RECORD axis only** (ADR-0045): `stop` is *record*-preserving (spool.db + info.json survive) where `purge` wipes them. It says nothing about processes — on the **process** axis `stop` is hard: for a broker-hosted endpoint it reaps the session and its descendant subtr...\n 202│<!-- [doc->REQ-ENDPOINT-PURGE] -->\n 203│\n 204│- **`api history-log`** — Path B: ingest normalized records into spt-core's native history store.\n*205│- **`api presence`** — adapter reports user interaction → updates the presence datum `(last_active_node, last_active_endpoint, ts)`. In the spt-hosted topology, presence is **also** updated by the broker *detecting* (sensing, not watching/logging) user input on a held PTY — privacy-preserving (it notes that input occurred, records no content).\n*206│- **`api emit --type <sensory_type> <payload>`** — a broker-launched **Shell** binary pushes a sensory payload to its owner agent (owner known from `api bind`; REST-only, never spooled). See the Shell model.\n 207│\n 208│**Not `api` commands — file-drop flow:** `commune` and `signoff` are deprecated as commands (modern SPT) in favor of file drops. The agent/adapter writes `<id>-commune.md` / `<id>-signoff.md`; the daemon watches the manifest-declared commune/signoff dirs (the spawn-session seam fields), ingests, and deletes (drop files are daemon-owned single-writer — KNOWN-HAZARDS 6.4). These stay off the `api` surface and the agent surface alike. **This path is the Self's, and spt-core's echo-commune never writes i...\n 209│\n    │...\n 228│\n*229│**spt-hosted (terminal wrapper / GUI launcher; the daemon launches the binary into a broker PTY):**\n*230│1. The frontend/CLI launches the agent: the daemon runs the **spawn-session** command template into a broker-held PTY.\n 231│2. The binary boots and fires **`api bind`** (or skips it under the strict UUID-injection + stable-pid commitment). **No catalyst/seed file** — the daemon is the launcher, already holds a direct channel (it spawned the process and owns the PTY), so a file round-trip would only add drive churn for no benefit.\n 232│3. The daemon delivers events; method is **manifest-configurable per activity-state** — direct PTY injection, or a relay even here (some adapters prefer a relay over PTY injection for idle delivery), or HTTP. During *activity*, delivery still defaults to the non-disruptive hook-injection path, not raw PTY writes.\n 233│4. Psyche is daemon-spawned, same as above.\n\n## docs/\n### STORAGE.md\n  43│    project-index.json           endpoint→project attribution (schema-versioned; daemon-written)\n *44│  daemon/                        ── ephemeral ── brain rehydration state, broker handle metadata\n  45│  updates/                       ── ephemeral ── staged binaries pending handoff\n  46│  adapters/                      registered adapter manifests / pointers\n  47│    <adapter>/web/               core-owned adapter served root; output bytes survive deactivation\n    │...\n  69│  node-sovereign stores — outside `tracked/` and carried by no wire record.\n *70│  The siting *is* the mechanism; a unit test scans the wire-record modules to\n  71│  keep it true.\n  72│\n  73│**Write sites.** Inbound is recorded in the WAN receive funnel **after the\n    │...\n 173│| subnet registry (snapshot + distribution) | **BranchStore** | audited; rides the precedence merge driver instead of a separate replication scheme (M4) |\n*174│| daemon/brain coarse checkpoint (rehydrate anchor, ADR-0004) | **BranchStore** | commit = checkpoint, ref tip = resume |\n*175│| endpoint lifecycle / status *audit trail* | **BranchStore** (or a derived view) | the *trajectory* matters, not just current value |\n 176│| spool (queue, drain-by-status, `id→address`) | **SQLite** | indexed queries + hot churn |\n*177│| broker effect journal (per-PTY-effect, B5) | **fsync-append** | hot path, per-keystroke — a commit per write is far too heavy |\n 178│| `info.json` / `status` (current value) | flat + atomic-rename (today) | tiny, read-hot; *may* move under BranchStore later, low priority |\n 179│\n 180│**Rules:**\n    │...\n 185│\n*186│**Reading the effect journal (`<spt_home>/daemon-effects.log`) — operator note.** The broker fsync-appends two line kinds per durable effect: `PENDING <key> <kind>` at attempt, `DONE <key>` on success (durable kinds only — an ephemeral `PtyWrite` writes no lines). A line with **`PENDING` and no matching `DONE`** is **expected**, not a leak: it is a durable effect that did NOT complete — for a `net-send` (a WAN write to an offline/unreachable peer) it is an **unconfirmed send retained as the re-driv...\n 187│\n 188│**Delivery vs propagation — latency envelope (operator note).** Two classes of cross-node traffic have deliberately different latency floors, so a minutes-scale delay is not automatically a defect: (1) **messages** (`spt send`) are **event-driven** — delivered on the events the daemon owns (WAN ingress + the ACTIVE→IDLE edge, F-023), so a healthy path is near-immediate and never waits on an adapter hook-poll cadence; (2) **anti-entropy context-sync** (the mind pull) is **periodic** — it rides the...\n 189│\n    │...\n 208│**Two transport modes:**\n*209│- **P2P (default — gh off):** peers exchange **incremental `git bundle`s over broker QUIC streams** — **pull-based, ref-scoped**: a node requests `a-<id>` for endpoints it has `synced(E,S)` and `p-<project>` only for projects it hosts (the requester names refs, so two-tier scoping falls out of the pull model); receiver fetches from the bundle + merges with driver + `--allow-unrelated-histories`. **The context DAG is shared across nodes** (one history, one merge commit per sync — divergent per-node ...\n 210│- **Hub (opt-in — gh on, via `spt context-github-setup`):** every node **pull → merge → push** against a private GitHub remote — always-online sync hub *and* live GUI view (no stale-mirror-node problem). Push-retry on non-fast-forward contention. **Prefer-GitHub, P2P backstop** when GitHub/SSH is unreachable (the P2P stack is always present anyway).\n 211│\n 212│**Scope:** hub mode is the **context-repo sync transport only.** Messaging, the subnet registry, remote-drive, pairing, and presence always ride P2P/Iroh — GitHub never carries real-time agent traffic, and gh-on never removes the P2P requirement.\n    │...\n 274│  one refresh. Fire-and-forget and daemon-offline-safe (events accumulate\n*275│  and are consumed at the next brain start).\n 276│  <!-- [doc->REQ-PROJECT-INDEX-INVALIDATION] -->\n 277│\n 278│## perches/ — node-local runtime\n    │...\n 303│| `relay` | `api listen` | a process that **holds** the endpoint | the endpoint's relay is gone — the one case relay-death convergence may act on |\n*304│| `binder` | `api bind` (agent or non-agent) | the **announcing CLI**, expected to exit within seconds | **nothing.** The hosting life it announced (a broker PTY session) is not this process, and is not recorded in the row at all |\n 305│| *absent* | any pre-`pid_role` writer | unknown | nothing — **no knowledge**, see below |\n 306│\n 307│**Why the field exists.** The same `pid` field carried both meanings, decided by write path, and every reader had to guess. Measured on one node (2026-07-27): three `claude-spt` live agents were online and working with recorded pids long dead (binder pids), while `omp-spt` and `mobile` endpoints held genuinely live relay pids. Two separate proposed fixes for a dead-relay bug were keyed on bare pid death; both would have offlined those three live agents. A record that states the meaning ends the guessing.\n    │...\n 312│\n*313│- **Relay-death convergence** (`livehost::reconcile_hosted_liveness`) is reached only by rows with `controllable != Some(true)`, and now acts only on `pid_role == relay`. So: harness relay rows converge on a proven relay death; `binder` rows (every `api bind` product of an agent or non-agent path, including gateway and shell binds) are **never** convergeable no matter how dead their pid; legacy rows are never convergeable; rows with a live broker PTY (`controllable == Some(true)`) never reach the branch ...\n 314│- **`controllable` carry-forward** (`api/startup.rs`) drops an earned `Some(true)` on a listener re-bind **only** when the prior row's `relay` pid is provably `Gone`. A dead `binder` pid, an `Unproven` verdict, and a legacy row all inherit.\n 315│\n 316│Both consumers route their liveness verdict through the single authority (`liveness::relay_liveness` over the `proc::process_identity` oracle); `Unproven` never kills.\n    │...\n 320│- Records the endpoint's durable identity: type, creation, and an **ordered adapter history** (most-recent first). Synced, so any node knows it.\n*321│- Drives the **resume UX**: the latest adapter is the default resume offer, the next-latest is second, … the oldest is second-to-last, and \"choose a different adapter\" is the final option.\n 322│- The *currently-running* `adapter_name` lives in node-local `perches/<id>/info.json`; the *history* is endpoint identity in `tracked/`.\n 323│\n 324│## Shells in the layout\n    │...\n 331│- **Adapter/platform-bound, not adapter-agnostic** — capability surface comes from the providing shell adapter; it does not move harnesses.\n*332│- **Lifecycle = online / offline / torn-down.** Link-break always closes the binary (optional pre-close instruction + termination timeout). **Ephemeral** (manifest property) ⇒ perch torn down *and removed from the agent's shell history*; **persistent** ⇒ perch kept offline for relink.\n 333│- **Cross-node links:** the shell perch lives on the shell's node (nested under the owner id as a label there); commands ride Iroh from the owner's node.\n 334│\n 335│## Migration & backup\n### MANIFEST.md\n 180│resume <id>`, `spt go <id>` on an offline endpoint, or the picker's *Resume from\n*181│history*) **and** the adapter declares the role. **Resume is latest-only:** no\n 182│verb takes a session argument — the retired `endpoint run --resume <session>`\n 183│spelling went with `endpoint run`, and a session argument is now *refused*\n 184│rather than silently ignored, because core resolves which session to resume from\n    │...\n 201│resolve (`claude -r <provisional>` hangs at \"no sessions match\"). A resume\n*202│request that is itself a provisional recovers the latest harness-reported ledger\n 203│row instead. When the endpoint has *no* harness-reported row at all — only a\n 204│spawn provisional on record — spt-core does **not** feed the provisional to your\n 205│template: it prints `RESUME_NO_HARNESS_SESSION:<id>` on stderr and starts a\n    │...\n 232│|---|---|---|\n*233│| `{node}` | this node's **advertised label** — the label store value (OS-hostname default, re-checked at daemon startup), OS-hostname fallback when no store label is known | **SINGLE-TOKEN fill only** — tokenize-then-fill cannot produce a space-carrying argv element, so a composite display name like `<id> @ <node>` stays adapter-shim territory (`{node}` complements the shim for tokenizable args, e.g. `--remote-control {id}--{node}`, it does not replace it). **Caveat:** the daemon-side lifecycle reso...\n 234│\n 235│<!-- [doc->REQ-MANIFEST-NODE-KEY] -->\n 236│```toml\n    │...\n 252│### `[env]` — env-var table\n*253│Vars to inject/read, and **how**. With `spt` on PATH the table is small. The injection *channel* is asymmetric: **spt-hosted** sessions inherit env from the broker that spawns them (no channel needed); **harness-hosted** sessions need the harness's declared channel (the broker didn't spawn them). File-bridge-only-when-not-launcher, applied to env.\n 254│\n 255│```toml\n 256│[env.OWL_SESSION_ID]\n    │...\n 258│value = \"{session_id}\"\n*259│channel = \"CLAUDE_ENV_FILE\"        # harness-hosted only; spt-hosted inherits from broker\n 260│```\n 261│\n 262│**`direction = \"read\"` — capture a launch-env var for template substitution.**\n    │...\n 422│Opt-in. The adapter's **idle-delivery translation binary**: a pure stdin→stdout\n*423│JSON-lines filter spt-core lifecycle-manages (spawned when the spt-hosted endpoint\n 424│comes up, terminated when it goes down). spt-core feeds it the inbound `<EVENT>`\n*425│feed and reads back keystroke-commands, which spt-core applies to the broker-held\n 426│PTY **atomically** — controller input is buffered during the emitted sequence and\n 427│flushed after, so injection coexists with a live `spt rc` controller (spt-core owns\n 428│every PTY write). Idle-only; busy/mid-turn delivery stays adapter hook-injection.\n    │...\n 444│- **Commit deadline + miss tolerance (F029 C-1, ADR-0022 amendment).** If no `{\"commit\":true}` arrives within the **commit deadline (5 s, `INJECT_COMMIT_DEADLINE`)**, spt-core flushes the buffered operator input (never stranded) and re-spools that message once — but a SINGLE miss does NOT fault a healthy binary (it is preserved; the next event delivers through it). Only **3 consecutive** missed commits, or a real binary death, faults the binary — and a real fault gets a **bounded eager respawn** (a h...\n*445│- **Known keys today:** `command` (preferred) and the deprecated `path`. spt-core does **not** `deny_unknown_fields` here — a newer adapter declaring a future key against an older spt-core parses fine (the unknown key is ignored), so the lifecycle-binary contract degrades gracefully.\n 446│- A minimal binary just emits `{\"text\":payload}{\"key\":\"enter\"}{\"commit\":true}` with no choreography (spt-hosted idle delivery is translation-binary-only since v0.14.3 — no raw-inject fallback).\n 447│\n 448│### `[service]` — a daemon-supervised resident service (ADR-0049)\n    │...\n 663│- **`--last <N>`** — the last `N` turns (the digest's natural unit). `--last 1` is\n*664│  the latest turn — the turn-end output.\n 665│- **Per-entry `seq`** — every JSON entry carries a **stable, source-derived** `seq`:\n 666│  it is deterministic from the entry's position in the session ledger, so\n 667│  **re-projection yields the SAME `seq`** for the same committed entry. It does\n    │...\n 704│\n*705│`spt adapter update [name]` (with no name, every registered `gh_release` adapter; with a name, just that one) compares the repo's latest GitHub release version against the installed adapter version and, when newer, fetches the release `.spt` (the same archive primitive as `spt adapter add --release`), then re-extracts and re-registers it in the adapter's durable `_github/<safe>` home (pointer-mode, re-read live). The network fetch lives in the `spt` CLI, never the daemon. **`repo` is required**; `asset` ...\n 706│\n 707│<!-- [doc->REQ-ADAPTER-GH-TRANSPORT] -->\n*708│**Transport — public HTTPS or private `gh`.** The optional **`transport`** selects how the asset bytes + the latest-release version are fetched: `https` (direct, the public-repo path), `gh` (shell the pre-authorized [`gh` CLI](https://cli.github.com/) — `gh release download` for the asset, `gh api` for the version — the **private-repo** path), or **`auto`** (the default: prefer `gh` when it is installed and authenticated, else fall back to HTTPS). Because `gh` honors OAuth + `GH_TOKEN`, an adapter ...\n 709│\n 710│<!-- [doc->REQ-ADAPTER-UPDATE-MESSAGE] -->\n 711│**Post-update notice — `message`.** The optional **`message`** is a plain (multi-line) human notice the adapter surfaces to stdout **only when `spt adapter update` actually applies an update** (the version changed) — never on a no-op / up-to-date run. It is read from the **newly-installed** manifest and rendered through the same inline-Markdown prose path as `spt`'s help (`**bold**`, `` `code` ``, `[text](url)`); there is no `{key}` substitution. It is avenue-agnostic (`gh_release` / `delegated` / `f...\n    │...\n 723│command = \"{adapter_dir}/claude-spt post-update\"   # REQUIRED, non-empty; install-dir + adapter-static subst\n*724│self_verifies = false                              # optional; attestation-only metadata (default false)\n 725│```\n 726│\n*727│- **Runs unconditionally** — the post-step fires **even when the primary avenue was a no-op** (already up to date). Its own idempotent check (e.g. `claude plugin update`) decides whether anything changes. `command` is **required** and **non-empty** (an empty `command` is refused at registration). `self_verifies` is **attestation-only metadata** (defaults `false`) — parallel to `[update].self_verifies`; it records that the post-step verifies its own outcome but **gates nothing** (the post-step is alre...\n 728│- **Published stdin seam** — spt-core feeds the post-step **one JSON line** on stdin describing the just-resolved update:\n 729│  ```json\n 730│  {\"adapter_applied\": true, \"adapter_name\": \"claude-spt\", \"profile_name\": null,\n    │...\n 765│[shell]\n*766│spawn = \"gamerobot-shell --link {link_token} --root {perch_dir}\"   # broker-launched; opaque command template — {perch_dir} = the instance's perch dir (REQ-SHELL-PERCH-DIR: lets the binary resolve a shell_file frame's perch-relative path; opt-in, spawn-only). Program token resolves against the adapter's install dir before PATH, and {adapter_dir} fills — see below.\n 767│ephemeral = false                  # manifest property, NOT an agent choice; ephemeral ⇒ no offline perch + no history retention\n 768│broadcast = \"subnet\"               # subnet | same-node | none\n 769│command_receipt = \"http\"           # http | stdin | relay   (how it receives agent commands)\n### CONTEXT-MEMORY.md\n  7│- **live context** (per-agent, project-independent — the agent's *mind*): identity/persona, relationship + rapport with the user, the user's cross-project preferences & working style, cross-project life threads, **learnings from work that may apply to other projects**, a thin **project roster** (one line per project the agent works on), the agent's self-model. Synced to **all** instances of the endpoint.\n *8│- **project context** (per-agent-per-project — what the agent is doing *here*): codebase/architecture, decisions made here, **current focus / next steps / todos**, project conventions, build/test commands, blockers. Synced **only** to same-project instances.\n  9│\n 10│(This split independently matches pi-agent-memory's Project Store vs Global Store and its discriminator dialogue — convergent validation.)\n 11│\n   │...\n 19│\n*20│1. **Default-to-project.** A fragment goes to project context unless it *passes the cross-project test* to be promoted to live. Promotion needs justification; demotion is free. The bias points against the leak.\n 21│2. **`current focus / next steps` is project-scoped** — the single biggest leak category. Route it to project unless the focus genuinely spans projects.\n 22│3. **Live holds a thin project roster, never project depth** — cross-project *awareness* travels; depth stays per-project.\n 23│4. **Commune-source asymmetry** (the structural fix): the live signal is strong only from **Self-authored communes**; **echo communes are project-centric** (they synthesize project-heavy session logs). So **echo communes are project-context-primary and live-context-conservative** — they make few/no live edits. This kills the leak at its source. Provenance is recorded on each block (`Source:`).","truncation":{"content":"# .worktrees/286-churn/\n## PRD.md\n 9|\n*10|spt-core is a **harness-independent core for an agent ecosystem**: inter-agent messaging, live-agent lifecycle, terminal hosting, seamless self-update, and zero-config cross-machine networking — shipped as both a Rust library workspace and a single canonical binary (`spt` / `spt.exe`).\n 11|\n 12|It is the ground-up rebuild of `claude_skill_owl` (\"modern SPT\"), untethering the system from Claude Code so that **any** agent runtime — Claude Code, Codex, Cursor, headless, or novel surfaces — can participate by either shelling out to the binary or linking the crates. The first thing built atop it is a rebuilt `spt` plugin reaching parity with today's system; the networking, terminal-wrapper, and multi-instance capabilities then make an agent reachable and operable across all of a user's machines.\n 13|\n...\n 16|- Today's modern SPT works but is **fragile** — each release introduces 2–5 regressions because it grew rapidly from a narrow starting point and carries the resulting tech debt (ADR-0001).\n*17|- It is **welded to Claude Code**. Other harnesses can't reuse the messaging/lifecycle substrate.\n 18|- Networking, terminal hosting (\"capsule\"), and cross-machine operation were planned piecemeal atop a CC-tied core; they deserve to be first-class in a clean foundation.\n 19|\n 20|## 3. Goals\n...\n 31|\n*32|- **macOS support** — out (no test machine); kept structurally easy, not a re-architecture.\n 33|- **Cross-subnet** (two *different* users' agents communicating) — deferred concept.\n 34|- **Concrete Shells / PresenceChannel implementation / presence gossip** — only the seams ship (§Endpoints).\n 35|- **Remote command execution on another user's node** and **instantiate-anywhere** — seam-compatible, deferred behind a consent/security gate.\n...\n 54|- **R-ARCH-3** `spt-proto` carries an explicit **wire-protocol version** independent of crate semver, with a documented **N-1 compat window**. Binary and library consumers interoperate by proto version, never by release version.\n*55|- **R-ARCH-4** Copy-verbatim the commodity layer from the sister project (envelope grammar, spool DDL, registry RPC, `info.json` shape); clean-room only the architecture/lifecycle (ADR-0001).\n 56|\n*57|### 7.2 The daemon (broker/brain)\n 58|- **R-DAEMON-1** Exactly **one logical `spt-daemon` per machine**, always-on, owning PTYs, network identity + WAN endpoint, subnet registry, all spools, all poll-listener logic, and all Psyche/pulse loops. No separate listener or wrapper processes (ADR-0004).\n*59|- **R-DAEMON-2** Internally split into a stable **broker** (holds PTY masters, harness child processes, listening sockets; versioned IPC; rarely updates) and an **updatable brain** (all logic; rehydrates from disk + re-attaches to broker handles).\n 60|- **R-DAEMON-3** Any `spt api` invocation that needs the daemon **auto-starts it** if absent.\n 61|- **R-DAEMON-4** Honor every invariant in [`docs/KNOWN-HAZARDS.md`](./docs/KNOWN-HAZARDS.md) as a conformance checklist.\n 62|\n...\n 88|- **R-START-2** **Harness-hosted:** SessionStart → `spt api seed --pid` (in-memory seed record) → `$SPT listen <id>` consumes the seed by `parent_pid` (validated vs `session_id`), binds, relays. No file.\n*89|- **R-START-3** **spt-hosted:** daemon runs spawn-session into a broker PTY → binary `api bind` (or skips under commitment) → daemon delivers via manifest-configured method. Principle: **file-bridge only when the daemon is not the launcher.**\n 90|- **R-START-4** Adapters inject env aliases (`$SPT` = `spt api`, heirs to `$OWL`/`$LIVE`); spt-core supplies the subcommands.\n 91|\n 92|---\n...\n 99|- **R-EP-3** Messaging payloads carry **typed operation commands + arbitrary file blobs** (text/audio/image/video), not just text — the seam that lets Shells, file transfer, and rich PresenceChannel exist later.\n*100|- **R-EP-4** PresenceChannel (deferred impl) is a **broker** endpoint with dispatch/bind/thread styles; presence datum `(last_active_node, last_active_endpoint, ts)`.\n 101|- **R-EP-5** (deferred impl, CONTEXT \"Shell model\") **Concrete shell instantiation model:** `shell spawn` *mints* an owner-exclusive instance (`<adapter>-<n>`, distinct from the `relink`/`persistent`/`wake` online switch); permission to instantiate = shell adapter **registered on the node** (broadcast governs *discovery* only); a per-shell **`require_approval`** gate (`none`|`remembered`|`always`, default none; manifest floor, node/endpoint may tighten) reuses the consent grant store; an optional **`max_...\n 102|\n 103|### 9.2 Instances (ADR-0003, as amended)\n...\n 153|\n*154|- **R-UPD-1** **Peer-propagated** over P2P (layered on self-fetch; out-of-band still supported); the subnet self-heals to latest (ADR-0004).\n 155|- **R-UPD-2** **All binaries signature-verified before handoff**, regardless of source (spt-core release key).\n*156|- **R-UPD-3** **No-endpoint-termination invariant, scoped by update class (ADR-0004 §A).** For the **brain-only** class — the routine case — the guarantee is **absolute**: no endpoint process terminates or suspends, satisfied by the broker/brain split (the broker keeps held PTYs/children/sockets alive across brain replacement). **broker-compatible** swaps the broker behind the versioned IPC with endpoints still surviving. **broker-breaking** (held-resource-type / OS-API change) MAY suspend endpoints...\n 157|- **R-UPD-4** **Not auto by default** — gated on user confirmation delivered to the most-recently-active live session, with an opt-in full-auto.\n 158|- **R-UPD-5** spt-core conducts **adapter ripple-updates** after self-update, per each adapter's update declaration and `min_spt_core_version` gate.\n 159|\n...\n 163|\n*164|- **R-TERM-1** A process-supervisor terminal wrapper (supersedes \"capsule\") hosting sessions in broker PTYs, multiplexed and name-addressable, attach/detach, headless↔headed.\n 165|- **R-TERM-2** A **session-surface** abstraction with native PTY as the day-one impl and network-attached surfaces for remote control; input injection at both `send-keys` (raw) and `send-line` (cooked) granularity; in-memory scrollback ring (disk spillover deferred).\n 166|- **R-TERM-3** Remote terminal streaming uses a **byte-stream** model for v1 (simple, faithful; near-LAN on Iroh-direct paths). Keep the surface abstraction generic enough that a predictive/state-sync \"feels-local\" layer can be added later without re-architecting — not built v1.\n*167|- **R-TERM-4** (M3, ADR-0008) **Live activity buffer (PTY digest):** a rolling, *parsed*, structured (source-tagged) view of recent PTY I/O (last ~N user turns + agent output between; tool sprints collapsed), built by running **adapter-supplied** manifest patterns (`input_pattern`, `agent_pattern`, `tool_pattern[]` + catchall) over the **broker-owned** PTY bytes — honors the no-built-in-parser rule (adapter owns the pattern). **Spt-hosted-primary** (needs the broker PTY); **harness-hosted is capability...\n 168|- **R-FRONT-1** Day-one headed frontend is a **launcher/manager** (not raw attach): list running + historic endpoints; launch a historic endpoint from its original manifest; tap into a running one; init a new one via a known adapter. Realizes the \"guided resume\" (XMB-style) and \"management GUI\" (per-endpoint panes) sketches. Its **CLI sibling** is a no-arg **`spt resume`** guided-resume picker — endpoints grouped by locality (on-node/this-project → on-node/other-project → off-node), MRU within grou...\n 169|\n 170|---\n...\n 176|- **R-INSTALL-3** First run idempotent + interactive-optional: generate node identity, start daemon; pairing + subnet-naming are separate explicit steps.\n*177|- **R-INSTALL-4** (CONTEXT \"adapter registration\") **Adapter registration lifecycle:** `spt adapter add <path>` (or `--github <user/repo>`, manifest-first then install via the declared `[update]` avenue — *install is the first update*) validates the manifest and records it under `{SPT_HOME}/…/adapters/` (copy for `file_pull`, pointer for `delegated`); one command/dir for `kind=\"harness\"` and `kind=\"shell\"`. Node-local (drive/launch capability, not subnet advertisement). `spt adapter remove` is **soft...\n 178|- **R-MIGRATE-1** Standalone install **auto-detects a `claude_skill_owl` install and offers migration** (identity, agents, tracked Psyche context). First-class supported path.\n 179|\n 180|---\n...\n 189|- **R-DOCS-4** **Agent-consumable layer:** CI-generated `llms.txt`/`llms-full.txt`, markdown content negotiation, a machine-readable manifest JSON Schema at a stable path, an **MCP doc/resource server** exposing the harness contract as deterministic resources, and structured CLI help (`spt <cmd> --help` + a `--json` mode).\n*190|- **R-DOCS-5** **Anti-drift via CI:** rustdoc, the manifest schema, the agent exports, and CLI help are generated + checked in CI — doc quality on the same footing as tests.\n 191|\n 192|---\n 193|\n...\n 195|\n*196|**In v1:** harness-independent core (library + binary); manifest + `api` harness contract; both startup topologies; one-daemon broker/brain; LiveAgent/ReadyAgent/Psyche/Worker/SptNode endpoints + Shell-vs-agent seam + typed/binary payloads; multi-instance data model + subnet registry + resolution policy + dormant/suspended + two-tier context sync + remote-drive of running instances + off-node file transfer; first-class WAN networking + TOTP-SPAKE2 pairing + subnet naming; cross-node Psyche sync (replaces...\n 197|\n 198|**Deferred (seam-compatible):** see [`docs/DEFERRED.md`](./docs/DEFERRED.md) — concrete Shells, PresenceChannel impl, presence gossip, instantiate-anywhere + consent, remote command exec + security gate, sidecar adapters, manifest includes, scrollback spillover, predictive/state-sync terminal layer, OS-level presence signal, macOS.\n 199|\n...\n 201|\n*202|1. A non-CC harness (or a headless spike) drives the full live-agent lifecycle through manifest + `api` alone — proving harness independence.\n 203|2. The rebuilt spt-plugin reaches **feature parity** with frozen `claude_skill_owl`.\n 204|3. Two of a user's machines, freshly paired via one TOTP code, exchange messages and the user drives an agent on machine A from machine B — zero config beyond pairing.\n 205|4. A self-update rolls across the subnet **without terminating any endpoint**.\n 206|5. The same agent's mind follows the user across machines (two-tier sync); switching machines leaves the prior instance dormant and catches up the new active one.\n*207|6. Every `KNOWN-HAZARDS` invariant has a corresponding passing test.\n 208|\n 209|## 17. Open design-phase questions\n 210|\n...\n 222|- [`docs/adr/0001`](./docs/adr/0001-clean-room-fork-from-claude-skill-owl.md) … [`0005`](./docs/adr/0005-totp-seeded-spake2-node-pairing.md) — decisions.\n*223|- [`docs/KNOWN-HAZARDS.md`](./docs/KNOWN-HAZARDS.md) — invariants harvested from the sister project (test checklist).\n 224|- [`docs/DEFERRED.md`](./docs/DEFERRED.md) — parked-but-committed features.\n 225|- [`docs/STORAGE.md`](./docs/STORAGE.md) — `$SPT_HOME` on-disk layout.\n 226|- [`docs/MANIFEST.md`](./docs/MANIFEST.md) — runtime manifest schema + worked `claude-spt` example.\n## ROADMAP.md\n 6|\n*7|A harness-independent core for an agent ecosystem — messaging, live-agent lifecycle, terminal hosting, **first-class P2P networking**, and a runtime-manifest harness contract — shipped as Rust crates + a single `spt` binary. The proof: the same agent runs natively across all of a user's machines with a synced mind, reachable and operable from anywhere, with seamless self-update and zero-config cross-machine messaging. First thing built atop it: a rebuilt `spt` plugin at parity with today's modern SPT.\n 8|\n 9|## Scope decisions (locked)\n 10|\n...\n 16|\n*17|**Stage A — Red-team review (before building).** Pressure-test the design adversarially. Focus on the riskiest architectural decisions, not bikeshedding. Tools: `plan-eng-review` (architecture) + `codex` (adversarial second opinion). Scope-cutting is *not* a goal (v1 is locked as full PRD), but surfacing contradictions / over-engineering / unproven assumptions is.\n 18|\n 19|**Stage B — De-risk spikes.** Throwaway proofs of the load-bearing unknowns — anything that, if it fails, forces a redesign. Do these before committing to the full build. See *De-risk spikes* below.\n 20|\n...\n 28|- **M1 — Local messaging + binary.** `spt-msg` (delivery TCP+spool, routing, send/ring/ready) + the `spt` binary + CLI. **= the killer quickstart (local): two agents exchange a message.** Task breakdown: [`M1-PLAN.md`](./M1-PLAN.md).\n*29|- **M2 — Harness contract + lifecycle.** **Strictly harness-agnostic — zero Claude Code conventions in this repo.** spt-core defines the adapter contract; the only adapter-shaped artifact in-tree is a generic mock/test adapter (manifest fixture + fake `api` caller) that exercises the contract (PRD R-DOCS-2 dev-agent quickstart). **The rebuilt `spt` plugin (Claude Code) is a SEPARATE downstream project** built atop these crates/binary — spt-core v1's acceptance proof (PRD §\"feature parity\"), not an...\n 30|  - **M2a — Harness contract. ✅ delivered (2026-06-01).** `spt-runtime` (AgentRuntime, ManifestRuntime, manifest schema) + the `api` subcommand surface (seed/listen/bind/state/echo-gate/poll/worker-*/boundary/session-end/presence/history-log/emit/capability) + the **harness-hosted** startup topology (seed→listen→bind, interim seed-file) + local-api-auth + the generic mock adapter + contract E2E. All M2a reqs activated + green. Task breakdown: [`M2a-PLAN.md`](./M2a-PLAN.md).\n*31|  - **M2b — Live-agent lifecycle. ✅ delivered (2026-06-01).** `spt-live` (fifth layer): history subsystem (fetcher/locate-normalize/native) + spawn-psyche (nested perch, spt-core-owned `$psyche_prompt`) + echo-commune (history brief → single-writer commune drop) + drop-file ingest (single-writer 6.4 + direct-write precedence 6.5) + interim pulse + signoff/boundary ordering invariants (echo-before-signoff 3.3, grace-before-signoff 1.1, stale-signoff sweep 3.2) + resume seam (continue-existing / fres...\n*32|  - **Interim, no daemon** (like M1): the `api listen` process holds the perch + relays; the seed bridges via a file. **The spt-hosted topology (PTY-launch) + the consolidated `spt-daemon` move to M3** (they need the `spt-term` broker) — this narrows the earlier \"both topologies in M2\" wording.\n*33|- **M3 — Terminal wrapper + daemon + self-update.** `spt-term` (session-surface, PTY, broker) + the consolidated `spt-daemon` (broker/brain split) + peer-propagated signed self-update. Split M3a/M3b/M3c. Scope + carried-forward Spike #1 gaps: [`M3-PLAN.md`](./M3-PLAN.md).\n 34|  - **M3a — `spt-term` session-surface mechanism. ✅ delivered (2026-06-02).** The sixth crate (`…→spt-msg→spt-term`, sibling of spt-runtime): the `SessionSurface` trait + native PTY backends (ConPTY/forkpty via `portable-pty`, REQ-TERM-1) · ConPTY-DSR auto-answer drain pump (REQ-HAZARD-CONPTY-DSR) · send-keys/send-line injection (REQ-TERM-2) · bounded byte-stream (REQ-TERM-3) · the PTY-digest parser primitive (REQ-TERM-4 impl/unit, ADR-0008). OS-neutral per Spike #4/#5 (no OS's stream contr...\n*35|  - **M3b — `spt-daemon` broker/brain process. ✅ delivered (2026-06-03).** The seventh crate (`…→spt-live→spt-daemon→spt`) hosts many `spt-term` surfaces behind a **versioned local IPC** (named-pipe/UDS behind one `DaemonTransport` trait, forward-compatible frame schema): a **broker** (stable kernel — PTY masters, harness children, sockets; no logic) + a **brain** (restartable logic) so a brain kill/restart leaves the hosted PTY child + its output stream **intact and gapless** (Spike #1 mad...\n*36|  - **M3c — signed self-update. ✅ delivered (2026-06-03).** The gated, signed, ripple-capable self-update engine (no new crate — it lives in `spt-daemon`, signing primitives in `release.rs`): the **update-class taxonomy** (brain-only / broker-compatible / broker-breaking, ADR-0004 §A) with the **brain-only zero-interruption swap** that drives the M3b handoff substrate (snapshot→drop→re-attach) so a live PTY child + its output stream survive the logic swap **untouched** (REQ-UPD-3, proven E2E `...\n*37|  - **M3 COMPLETE** (M3a + M3b + M3c all ✅, 2026-06-03): the real ADR-0004 architecture is live — native PTY hosting, the consolidated broker/brain daemon, and seamless signed self-update. Next milestone: M4.\n*38|- **M4 — Networking + instances. ✅ delivered (2026-06-04).** The eighth crate `spt-net` (Iroh WAN endpoint bound to the node's own Ed25519 identity, mDNS LAN discovery, TOTP-SPAKE2 pairing over a dedicated pre-trust ALPN with transcript binding / rate limiting / seed rotation+transfer, subnet registry + NDJSON replication under the per-node epoch lease) · the multi-instance model (registry rows + bare-id resolution refuse-and-qualify + visibility/sync-membership gates + rename/collision + **resting-...\n*39|- **M5 — Shells, presence, deferred capabilities. ✅ delivered (2026-06-04;** scope locked with user 2026-06-04, see [`M5-PLAN.md`](./M5-PLAN.md) + per-task JIT plans D0–D9). The consent framework seam (grant store + interactive escalation + pre-consent flags; remote-exec / instantiate-anywhere reserved-but-refusing) · adapter registration lifecycle (`spt adapter add/remove`, manifest-first) · **shell hosting machinery + sleep/wake** (link-token channels with per-frame MAC, broker-launched binarie...\n 40|\n 41|## Post-v1-core milestones\n 42|\n...\n 46|\n*47|- **✅ Broker/brain process-isolation restoration — delivered (2026-06-11).** Design: [`docs/BROKER-BRAIN-SPLIT-RESTORATION.md`](./docs/BROKER-BRAIN-SPLIT-RESTORATION.md) (verified-with-amendments by `doyle`); decision: **ADR-0018**; build plan: [`RESTORATION-PLAN.md`](./RESTORATION-PLAN.md) (D1–D7). Corrects an unintended regression — the broker/brain split (ADR-0004) was specced + spiked as two *processes*, but the production daemon runs the broker as an in-process *thread* (`daemon.rs:165-170`)...\n 48|\n 49|- **🔜 NEXT milestone (planned 2026-06-12) — M9: Adapter customization & session surfaces.** Product of the two-day **Gateway grill** (operator + `doyle`; CONTEXT.md mints @ `8dc22ff`/`bd469b9`; ADR-0008 amended; the lecturn device design parked in the sibling `spt-lecturn/` repo for a later build). Shell-free, independently-shippable features: **adapter profiles** (sparse leaf-replace overlays on the parent manifest; shipped + local sources, local surviving adapter updates; composite `<adapter>:<pro...\n*50|- **🔜 M10 — Digest extractor & thread-spanning** (planned 2026-06-13; NEXT). Product of the digest grill (operator + `doyle`); governed by **ADR-0019** (reverses M9's \"no manifest seam\"), CONTEXT.md \"session digest\" rewritten. An `spt-claude-code` enabler, so it takes the M10 slot ahead of the Shell/remote-attach work (which shift to M11/M12). Adapter-declared **`[digest]` extractor seam** (imperative, no DSL; same source files as `[history]` by default + own-source escape; `api digest-entry` push f...\n 51|\n 52|- **🔜 v0.15.0 — Activity-gated message delivery + send-modifier axes.** Product of the messaging grill (operator + `doyle`, 2026-06-23); governed by **ADR-0028**, CONTEXT.md mints \"activity-gated delivery\" / \"message delivery axes\" / \"message metadata (`json`)\". Closes a **legacy-SPT parity gap that was never written down**: activity-gated inbound routing is scaffolded (`is_idle` + `resolve_inject_methods`) but unwired (`dispatch_endpoint_input` injects unconditionally). Wires the two-window router ...\n 53|\n...\n 57|\n*58|1. **Broker/brain no-terminate handoff** — prove a PTY child + open socket survive a daemon-*logic* restart (FD-passing or stable-broker shim). The hardest invariant in the design (ADR-0004).\n*59|2. **Iroh + mDNS smoke test** — two nodes, separate NATs, one message end-to-end; measure direct-vs-relay + binary size (research brief §5). Validates the whole networking premise.\n*60|3. **ConPTY + `portable-pty`** — host and drive a real session in a broker PTY on Windows + Linux.\n*61|4. **Manifest-driven lifecycle** — a manifest drives one real Claude Code session through spawn → bind → commune → signoff, proving harness-independence works in practice.\n 62|\n 63|## Immediate next steps\n 64|\n...\n 66|- ✅ **Stage A — codex adversarial pass** — done. 14 findings (5 FATAL, 9 SERIOUS) in `docs/reviews/STAGE-A-codex-redteam.md`.\n*67|- ✅ **Spike #1 (broker/brain handoff)** — PASS on Windows ConPTY. `docs/spikes/SPIKE-01-broker-handoff.md`. Found+fixed the ConPTY DSR-stall hazard. 4 open gaps carried to M3.\n*68|- ✅ **Spike #2 (Iroh smoke test)** — PASS (single-host). `docs/spikes/SPIKE-02-iroh-smoke.md`. iroh 0.98 API workable; 13.57 MB release. Cross-NAT/relay proof deferred to M4.\n 69|- ✅ **Findings folded** into ADR-0002/0003/0004/0005 + KNOWN-HAZARDS (+3) + traceable-reqs.toml (+4 REQ-HAZARD).\n 70|\n 71|Remaining before M0:\n 72|1. **Stage A — `plan-eng-review`** (interactive architecture pass; second half of Stage A).\n*73|2. **Spike #4 (manifest-driven lifecycle)** — a manifest drives one real session spawn→bind→commune→signoff. (Spike #3 ConPTY hosting is largely covered by Spike #1.)\n 74|3. Resolve the open user decisions logged in the ADR amendments (context-merge model, seed-rotation, R-UPD-3 rewording).\n 75|4. **M0** (workspace + `spt-proto` + `spt-store`) — **started**; task breakdown in [`M0-PLAN.md`](./M0-PLAN.md). Each milestone gets its own just-in-time plan (this one is M0); the roadmap stays the milestone sequence, the JIT plan is the task layer.\n 76|\n## CONTEXT.md\n 2|\n*3|**Platform scope:** Windows + Linux for v1. macOS is out (no test machine available) but kept structurally easy — `portable-pty` and Iroh both support it, so macOS is a later test/CI-budget decision, not a re-architecture.\n 4|\n 5|**Legacy migration:** it should be possible — ideally *automatic* — for a user to migrate an existing `claude_skill_owl` (modern SPT) install to spt-core (identity, agents, tracked Psyche context). Exact mechanism deferred to design; the commitment is that migration is a first-class supported path, not a manual rebuild.\n 6|\n*7|Harness-independent core for the SPT ecosystem. Provides inter-agent messaging, live-agent lifecycle, terminal wrapping, self-update, and networking primitives — as both a Rust library workspace and a canonical reference binary. Designed so any agent runtime (Claude Code, Codex, Cursor, headless, future harnesses) can interface with the SPT ecosystem either by shelling out to the binary or by linking the crates directly.\n 8|\n 9|Successor to `claude_skill_owl` (today's \"modern SPT\"), which is being rebuilt as `spt-core` to untether the system from Claude Code and lift it to a general-purpose agent-ecosystem core.\n 10|\n...\n 22|**spt plugin** (separate downstream project — NOT an spt-core deliverable):\n*23|A rebuilt version of today's Claude Code `spt` plugin. It is the **first consumer** built *atop* spt-core and the **acceptance proof** of spt-core v1 (it reaches feature parity with modern SPT while delegating all core functionality to spt-core, primarily via `spt.exe`, with deeper hooks where useful) — but it **lives and builds in its own repository, outside spt-core**. It is a Claude-Code-specific *adapter*: it holds the Claude Code conventions (hooks, slash-commands, skill/plugin layout, `claude` se...\n 24|\n 25|**Pi** (disambiguation — two meanings, never conflate):\n 26|(1) **Pi, the coding agent/harness** (`badlogic/pi-mono`) — a harness example alongside Claude Code and Codex; this is the meaning in user-facing harness lists. (2) **Pi-class node** — Raspberry-Pi-class low-power hardware hosting a Shell-only or headless SPT node; an incidental hardware descriptor, never an explicit product example. Public-facing docs must disambiguate or avoid the bare word.\n...\n 35|Internally the logical daemon is split into two implementation layers for seamless self-update (see Self-update):\n*36|- **broker** (stable \"kernel\") — holds *only* the un-transferable, must-not-die resources: PTY master fds, the spawned harness child processes, and listening network sockets. Minimal, dumb, versioned local IPC. Almost never updates.\n*37|- **daemon brain** (\"userspace\") — all logic (routing, registry, pulse/psyche loops, manifest parsing, update orchestration). Restarts freely on update; rehydrates from disk state and re-attaches to the broker's held handles.\n 38|\n*39|Logical addressing is unchanged — still one per-machine `spt-daemon`; the broker is an internal layer, not separately addressable. There is exactly **one broker per machine** (per `SPT_HOME`) — *not* one per endpoint: a single broker holds every hosted endpoint's resources, and it is present whenever the daemon runs, even with zero endpoints online (the bare-daemon case). It is therefore the always-present per-machine layer, which is why the single-daemon lock + liveness anchor belong to it.\n 40|\n 41|**in-session relay**:\n*42|A thin, stateless `spt.exe` task that exists only in **harness-hosted** sessions (where the agent harness is the parent process and spt cannot reach into its process tree — today's Monitor model). It streams the daemon brain's events into the session's stdout. All *stateful* listener logic lives in the daemon; the relay is a dumb pipe, freely killable and respawnable. **spt-hosted** sessions need no separate harness-owned relay — the daemon owns the PTY and consumes the same poll feed itself. Idle de...\n 43|\n 44|### Deliverable shape\n 45|\n...\n 105|<!-- [doc->REQ-UPD-9] -->\n*106|Each adapter manifest declares how spt-core should *ripple-update the adapter itself* (see Self-update). One of: **file-pull** (a plugin-directory lookup regex + a gh repo for the adapter's latest files — spt-core fetches + swaps), **delegated command** (a binary command the adapter owns, e.g. `claude.exe plugin update` — spt-core invokes it), or **gh_release** (the adapter ships its updates from its own GitHub releases). After initial bootstrap, the plugin no longer self-manages updates; spt-core co...\n 107|\n 108|**adapter packaging & live update** (v0.13.2; ADR-0024, ADR-0025):\n 109|<!-- [doc->REQ-ADAPTER-GH-TRANSPORT] -->\n*110|A `.spt` may be **multi-platform**: shared `manifest.toml` + `strings/` at the root, role binaries under per-target-triple subdirectories (`x86_64-pc-windows-msvc/`, …); install/update extracts the shared root plus only the current node's triple, flattened into `install_dir`, so flat `<install_dir>/<program>` resolution is unchanged. It stays one signed asset (`adapter.spt`, plain-tar or gzip); a multi-platform archive missing the recipient's triple is a typed `NoArtifactForPlatform`. Large adapters ma...\n 111|<!-- [doc->REQ-ADAPTER-UPDATE-MESSAGE] -->\n 112|An optional **`[update].message`** (avenue-agnostic) is a plain multi-line operator notice surfaced to stdout, markdown-rendered (the helpfmt prose path), **only when an update is actually applied** (the version changed) — never on a no-op. It is read from the newly-installed manifest with no `{key}` substitution; its use is to announce a post-update action (e.g. \"run `/reload-plugins` in any ongoing sessions\").\n 113|\n...\n 124|\n*125|Governing principle: **SPT is not a harness.** Model choice, billing shape, harness-internal env, and harness-internal context are entirely the adapter's concern, expressed inside the adapter's own command templates. spt-core owns only the template *mechanism* (substitution keys), the substitution *values* it is responsible for, and the surrounding lifecycle. Env for the *endpoint binary itself* is auto-handled by spt-core/broker; env for the *agent running inside* that binary is the adapter's config (e....\n 126|\n*127|**spawn-session seam** — launch a new agent session on this node. Manifest provides: a command template; `cwd`/project; a `headless` flag (optional, default false — for the GUI's resume-of-compatible-adapters); a `resume` flag (optional); and the `commune` + `signoff` file directories relative to `cwd` (so the daemon knows where to watch). Substitution keys spt-core can supply: `{id}` and, optionally, a spt-core-generated valid session UUID (e.g. injected as `--session-id {uuid}`) so an adapter can s...\n 128|- **id resolution:** `id` is optional. With no id, spt-core reproduces today's no-id `/spt:live` behavior — run the lone live agent if that's all the project has; show a picker with proposed default IDs if the project has none; let the user choose if there are several.\n 129|\n 130|**post-spawn seam** — the just-launched binary calls an spt-core command on boot (via the adapter's SessionStart-equivalent hook) to bind itself. Needed because the harness's own session id usually isn't known until after the binary runs. Payload: the harness `session_id` (when binary-generated rather than spt-core-injected); the `parent_pid` (the stable session-binding anchor — see KNOWN-HAZARDS 2.1); an endpoint identity/type confirmation; optionally a local HTTP port the binary listens on (for HTT...\n...\n 145|\n*146|**activity observation** (ruled 2026-07-24, rebound grill) — two avenues, deliberately split by consumer class; the digest is **not** one of them (it stays a content surface). **Push (shells only):** an owned Shell observes its *owner's* busy/idle transitions as an **activity frame** on the existing shell-link event stream — link-scoped (owner implied by the link token), **drive-class semantics** (ephemeral, latest-wins, current-state-carrying; a redundant same-state resend is a harmless no-op — th...\n 147|\n 148|**inject-input seam** — message delivery into a running session. Configurable per activity-state (activity / idle / both); multiple methods, any combination:\n 149|- PTY injection (with or without key/submit sequences) — spt-hosted topology;\n...\n 175|- **Command templates are opaque.** spt-core never parses out a model/tool/flag — the adapter writes the whole command line; spt-core fills substitution keys and runs it.\n*176|- **A command template's program token resolves against the adapter install dir before PATH (since v0.8.0).** A `.spt` adapter ships its built binaries to the adapter's install dir (`adapters/_github/<safe>/` via `--release`/`--github`, or the record's `source_dir` under copy-mode), so a bare program name (e.g. `claude-spt-digest …`) binds to the shipped binary first and falls back to PATH when absent — a `.spt` that ships its binaries is **self-contained**, needing no PATH placement. <!-- [doc->REQ-...\n 177|- **Hook output capability is declared per harness-event** (`can_inject`). CC's Stop hook cannot inject context — that single fact drives the echo-gate sentinel + relay fallback. The manifest expresses it so spt-core knows when to fall back.\n*178|- **Env injection is asymmetric** (file-bridge-only-when-not-launcher, applied to env): spt-hosted sessions inherit env from the broker that spawns them; harness-hosted sessions need the harness's declared env channel. With `spt` on PATH the env table is small.\n 179|- **Cross-adapter fallback** is a **node-wide setting**, not a manifest field: if a Psyche/echo invocation under one adapter is rate-limited, spt-core falls back to another adapter (e.g. `ccs` — its own adapter, not a binary-swap). <!-- [doc->REQ-MANIFEST-6] --> A fallback **target is addressed as `<adapter>:<profile>`** (not just a bare adapter_name) and resolves through the one composite-addressing resolver (`registry::resolve_option`), so a fallback may select a shipped or local profile (`ccs`, `ccs...\n 180|- **Config knobs** (pulse period, echo-commune window/gate, route-guard window, daily refresh) are spt-core **global settings** with optional **per-endpoint override**. **An adapter may DECLARE A DEFAULT, never an override** (narrowed 2026-08-03, LOCKSMITH grill — the original \"never per-adapter\" wording is superseded): a harness has real information about its own turn shape and cost, but the operator keeps the last word. Precedence, highest first: **per-endpoint override → node/global setting → ad...\n 181|- **Event-block vocabulary and file-drop filenames are fixed spt-core constants** (documented for adapter authors), not manifest-configurable. <!-- [doc->REQ-RESUME-CONTEXT-PULL] --> This includes the **checkpoint sentinel `!!checkpoint!!`** — the agent-checkpoint trigger an adapter embeds in a commune/signoff drop body (one bare token = checkpoint with default wake; a `!!checkpoint!! <text> !!checkpoint!!` pair makes the inter-marker text a custom wake directive). It is spt-core control metadata: spt-...\n...\n 199|\n*200|**`spt endpoint purge <id>`** (CLI, not `api`) — the standalone, formal **full teardown**: wipe an endpoint and *every* record keyed on it. It is the dev/CI sibling of `api session-end --erase` (which is adapter-triggered at session end); `purge` is the explicit operator/test command for clean setup-and-reset. **Deliberately NOT consent-gated** — a local dev/test op, never a peer-visible action. **Offline-only**: it refuses a live / daemon-hosted endpoint (deleting records out from under a running ho...\n*201|_Avoid_: consent-gating it (it is intentionally ungated, for CI); treating it as a sync/remote op (local-only); a soft variant (purge is always the hard, full wipe — soft teardown is `endpoint stop`). **Read that soft/hard contrast on the RECORD axis only** (ADR-0045): `stop` is *record*-preserving (spool.db + info.json survive) where `purge` wipes them. It says nothing about processes — on the **process** axis `stop` is hard: for a broker-hosted endpoint it reaps the session and its descendant subtr...\n 202|<!-- [doc->REQ-ENDPOINT-PURGE] -->\n 203|\n 204|- **`api history-log`** — Path B: ingest normalized records into spt-core's native history store.\n*205|- **`api presence`** — adapter reports user interaction → updates the presence datum `(last_active_node, last_active_endpoint, ts)`. In the spt-hosted topology, presence is **also** updated by the broker *detecting* (sensing, not watching/logging) user input on a held PTY — privacy-preserving (it notes that input occurred, records no content).\n*206|- **`api emit --type <sensory_type> <payload>`** — a broker-launched **Shell** binary pushes a sensory payload to its owner agent (owner known from `api bind`; REST-only, never spooled). See the Shell model.\n 207|\n 208|**Not `api` commands — file-drop flow:** `commune` and `signoff` are deprecated as commands (modern SPT) in favor of file drops. The agent/adapter writes `<id>-commune.md` / `<id>-signoff.md`; the daemon watches the manifest-declared commune/signoff dirs (the spawn-session seam fields), ingests, and deletes (drop files are daemon-owned single-writer — KNOWN-HAZARDS 6.4). These stay off the `api` surface and the agent surface alike. **This path is the Self's, and spt-core's echo-commune never writes i...\n 209|\n...\n 228|\n*229|**spt-hosted (terminal wrapper / GUI launcher; the daemon launches the binary into a broker PTY):**\n*230|1. The frontend/CLI launches the agent: the daemon runs the **spawn-session** command template into a broker-held PTY.\n 231|2. The binary boots and fires **`api bind`** (or skips it under the strict UUID-injection + stable-pid commitment). **No catalyst/seed file** — the daemon is the launcher, already holds a direct channel (it spawned the process and owns the PTY), so a file round-trip would only add drive churn for no benefit.\n 232|3. The daemon delivers events; method is **manifest-configurable per activity-state** — direct PTY injection, or a relay even here (some adapters prefer a relay over PTY injection for idle delivery), or HTTP. During *activity*, delivery still defaults to the non-disruptive hook-injection path, not raw PTY writes.\n 233|4. Psyche is daemon-spawned, same as above.\n\n## docs/\n### STORAGE.md\n 43|    project-index.json           endpoint→project attribution (schema-versioned; daemon-written)\n*44|  daemon/                        ── ephemeral ── brain rehydration state, broker handle metadata\n 45|  updates/                       ── ephemeral ── staged binaries pending handoff\n 46|  adapters/                      registered adapter manifests / pointers\n 47|    <adapter>/web/               core-owned adapter served root; output bytes survive deactivation\n...\n 69|  node-sovereign stores — outside `tracked/` and carried by no wire record.\n*70|  The siting *is* the mechanism; a unit test scans the wire-record modules to\n 71|  keep it true.\n 72|\n 73|**Write sites.** Inbound is recorded in the WAN receive funnel **after the\n...\n 173|| subnet registry (snapshot + distribution) | **BranchStore** | audited; rides the precedence merge driver instead of a separate replication scheme (M4) |\n*174|| daemon/brain coarse checkpoint (rehydrate anchor, ADR-0004) | **BranchStore** | commit = checkpoint, ref tip = resume |\n*175|| endpoint lifecycle / status *audit trail* | **BranchStore** (or a derived view) | the *trajectory* matters, not just current value |\n 176|| spool (queue, drain-by-status, `id→address`) | **SQLite** | indexed queries + hot churn |\n*177|| broker effect journal (per-PTY-effect, B5) | **fsync-append** | hot path, per-keystroke — a commit per write is far too heavy |\n 178|| `info.json` / `status` (current value) | flat + atomic-rename (today) | tiny, read-hot; *may* move under BranchStore later, low priority |\n 179|\n 180|**Rules:**\n...\n 185|\n*186|**Reading the effect journal (`<spt_home>/daemon-effects.log`) — operator note.** The broker fsync-appends two line kinds per durable effect: `PENDING <key> <kind>` at attempt, `DONE <key>` on success (durable kinds only — an ephemeral `PtyWrite` writes no lines). A line with **`PENDING` and no matching `DONE`** is **expected**, not a leak: it is a durable effect that did NOT complete — for a `net-send` (a WAN write to an offline/unreachable peer) it is an **unconfirmed send retained as the re-driv...\n 187|\n 188|**Delivery vs propagation — latency envelope (operator note).** Two classes of cross-node traffic have deliberately different latency floors, so a minutes-scale delay is not automatically a defect: (1) **messages** (`spt send`) are **event-driven** — delivered on the events the daemon owns (WAN ingress + the ACTIVE→IDLE edge, F-023), so a healthy path is near-immediate and never waits on an adapter hook-poll cadence; (2) **anti-entropy context-sync** (the mind pull) is **periodic** — it rides the...\n 189|\n...\n 208|**Two transport modes:**\n*209|- **P2P (default — gh off):** peers exchange **incremental `git bundle`s over broker QUIC streams** — **pull-based, ref-scoped**: a node requests `a-<id>` for endpoints it has `synced(E,S)` and `p-<project>` only for projects it hosts (the requester names refs, so two-tier scoping falls out of the pull model); receiver fetches from the bundle + merges with driver + `--allow-unrelated-histories`. **The context DAG is shared across nodes** (one history, one merge commit per sync — divergent per-node ...\n 210|- **Hub (opt-in — gh on, via `spt context-github-setup`):** every node **pull → merge → push** against a private GitHub remote — always-online sync hub *and* live GUI view (no stale-mirror-node problem). Push-retry on non-fast-forward contention. **Prefer-GitHub, P2P backstop** when GitHub/SSH is unreachable (the P2P stack is always present anyway).\n 211|\n 212|**Scope:** hub mode is the **context-repo sync transport only.** Messaging, the subnet registry, remote-drive, pairing, and presence always ride P2P/Iroh — GitHub never carries real-time agent traffic, and gh-on never removes the P2P requirement.\n...\n 274|  one refresh. Fire-and-forget and daemon-offline-safe (events accumulate\n*275|  and are consumed at the next brain start).\n 276|  <!-- [doc->REQ-PROJECT-INDEX-INVALIDATION] -->\n 277|\n 278|## perches/ — node-local runtime\n...\n 303|| `relay` | `api listen` | a process that **holds** the endpoint | the endpoint's relay is gone — the one case relay-death convergence may act on |\n*304|| `binder` | `api bind` (agent or non-agent) | the **announcing CLI**, expected to exit within seconds | **nothing.** The hosting life it announced (a broker PTY session) is not this process, and is not recorded in the row at all |\n 305|| *absent* | any pre-`pid_role` writer | unknown | nothing — **no knowledge**, see below |\n 306|\n 307|**Why the field exists.** The same `pid` field carried both meanings, decided by write path, and every reader had to guess. Measured on one node (2026-07-27): three `claude-spt` live agents were online and working with recorded pids long dead (binder pids), while `omp-spt` and `mobile` endpoints held genuinely live relay pids. Two separate proposed fixes for a dead-relay bug were keyed on bare pid death; both would have offlined those three live agents. A record that states the meaning ends the guessing.\n...\n 312|\n*313|- **Relay-death convergence** (`livehost::reconcile_hosted_liveness`) is reached only by rows with `controllable != Some(true)`, and now acts only on `pid_role == relay`. So: harness relay rows converge on a proven relay death; `binder` rows (every `api bind` product of an agent or non-agent path, including gateway and shell binds) are **never** convergeable no matter how dead their pid; legacy rows are never convergeable; rows with a live broker PTY (`controllable == Some(true)`) never reach the branch ...\n 314|- **`controllable` carry-forward** (`api/startup.rs`) drops an earned `Some(true)` on a listener re-bind **only** when the prior row's `relay` pid is provably `Gone`. A dead `binder` pid, an `Unproven` verdict, and a legacy row all inherit.\n 315|\n 316|Both consumers route their liveness verdict through the single authority (`liveness::relay_liveness` over the `proc::process_identity` oracle); `Unproven` never kills.\n...\n 320|- Records the endpoint's durable identity: type, creation, and an **ordered adapter history** (most-recent first). Synced, so any node knows it.\n*321|- Drives the **resume UX**: the latest adapter is the default resume offer, the next-latest is second, … the oldest is second-to-last, and \"choose a different adapter\" is the final option.\n 322|- The *currently-running* `adapter_name` lives in node-local `perches/<id>/info.json`; the *history* is endpoint identity in `tracked/`.\n 323|\n 324|## Shells in the layout\n...\n 331|- **Adapter/platform-bound, not adapter-agnostic** — capability surface comes from the providing shell adapter; it does not move harnesses.\n*332|- **Lifecycle = online / offline / torn-down.** Link-break always closes the binary (optional pre-close instruction + termination timeout). **Ephemeral** (manifest property) ⇒ perch torn down *and removed from the agent's shell history*; **persistent** ⇒ perch kept offline for relink.\n 333|- **Cross-node links:** the shell perch lives on the shell's node (nested under the owner id as a label there); commands ride Iroh from the owner's node.\n 334|\n 335|## Migration & backup\n### MANIFEST.md\n 180|resume <id>`, `spt go <id>` on an offline endpoint, or the picker's *Resume from\n*181|history*) **and** the adapter declares the role. **Resume is latest-only:** no\n 182|verb takes a session argument — the retired `endpoint run --resume <session>`\n 183|spelling went with `endpoint run`, and a session argument is now *refused*\n 184|rather than silently ignored, because core resolves which session to resume from\n...\n 201|resolve (`claude -r <provisional>` hangs at \"no sessions match\"). A resume\n*202|request that is itself a provisional recovers the latest harness-reported ledger\n 203|row instead. When the endpoint has *no* harness-reported row at all — only a\n 204|spawn provisional on record — spt-core does **not** feed the provisional to your\n 205|template: it prints `RESUME_NO_HARNESS_SESSION:<id>` on stderr and starts a\n...\n 232||---|---|---|\n*233|| `{node}` | this node's **advertised label** — the label store value (OS-hostname default, re-checked at daemon startup), OS-hostname fallback when no store label is known | **SINGLE-TOKEN fill only** — tokenize-then-fill cannot produce a space-carrying argv element, so a composite display name like `<id> @ <node>` stays adapter-shim territory (`{node}` complements the shim for tokenizable args, e.g. `--remote-control {id}--{node}`, it does not replace it). **Caveat:** the daemon-side lifecycle reso...\n 234|\n 235|<!-- [doc->REQ-MANIFEST-NODE-KEY] -->\n 236|```toml\n...\n 252|### `[env]` — env-var table\n*253|Vars to inject/read, and **how**. With `spt` on PATH the table is small. The injection *channel* is asymmetric: **spt-hosted** sessions inherit env from the broker that spawns them (no channel needed); **harness-hosted** sessions need the harness's declared channel (the broker didn't spawn them). File-bridge-only-when-not-launcher, applied to env.\n 254|\n 255|```toml\n 256|[env.OWL_SESSION_ID]\n...\n 258|value = \"{session_id}\"\n*259|channel = \"CLAUDE_ENV_FILE\"        # harness-hosted only; spt-hosted inherits from broker\n 260|```\n 261|\n 262|**`direction = \"read\"` — capture a launch-env var for template substitution.**\n...\n 422|Opt-in. The adapter's **idle-delivery translation binary**: a pure stdin→stdout\n*423|JSON-lines filter spt-core lifecycle-manages (spawned when the spt-hosted endpoint\n 424|comes up, terminated when it goes down). spt-core feeds it the inbound `<EVENT>`\n*425|feed and reads back keystroke-commands, which spt-core applies to the broker-held\n 426|PTY **atomically** — controller input is buffered during the emitted sequence and\n 427|flushed after, so injection coexists with a live `spt rc` controller (spt-core owns\n 428|every PTY write). Idle-only; busy/mid-turn delivery stays adapter hook-injection.\n...\n 444|- **Commit deadline + miss tolerance (F029 C-1, ADR-0022 amendment).** If no `{\"commit\":true}` arrives within the **commit deadline (5 s, `INJECT_COMMIT_DEADLINE`)**, spt-core flushes the buffered operator input (never stranded) and re-spools that message once — but a SINGLE miss does NOT fault a healthy binary (it is preserved; the next event delivers through it). Only **3 consecutive** missed commits, or a real binary death, faults the binary — and a real fault gets a **bounded eager respawn** (a h...\n*445|- **Known keys today:** `command` (preferred) and the deprecated `path`. spt-core does **not** `deny_unknown_fields` here — a newer adapter declaring a future key against an older spt-core parses fine (the unknown key is ignored), so the lifecycle-binary contract degrades gracefully.\n 446|- A minimal binary just emits `{\"text\":payload}{\"key\":\"enter\"}{\"commit\":true}` with no choreography (spt-hosted idle delivery is translation-binary-only since v0.14.3 — no raw-inject fallback).\n 447|\n 448|### `[service]` — a daemon-supervised resident service (ADR-0049)\n...\n 663|- **`--last <N>`** — the last `N` turns (the digest's natural unit). `--last 1` is\n*664|  the latest turn — the turn-end output.\n 665|- **Per-entry `seq`** — every JSON entry carries a **stable, source-derived** `seq`:\n 666|  it is deterministic from the entry's position in the session ledger, so\n 667|  **re-projection yields the SAME `seq`** for the same committed entry. It does\n...\n 704|\n*705|`spt adapter update [name]` (with no name, every registered `gh_release` adapter; with a name, just that one) compares the repo's latest GitHub release version against the installed adapter version and, when newer, fetches the release `.spt` (the same archive primitive as `spt adapter add --release`), then re-extracts and re-registers it in the adapter's durable `_github/<safe>` home (pointer-mode, re-read live). The network fetch lives in the `spt` CLI, never the daemon. **`repo` is required**; `asset` ...\n 706|\n 707|<!-- [doc->REQ-ADAPTER-GH-TRANSPORT] -->\n*708|**Transport — public HTTPS or private `gh`.** The optional **`transport`** selects how the asset bytes + the latest-release version are fetched: `https` (direct, the public-repo path), `gh` (shell the pre-authorized [`gh` CLI](https://cli.github.com/) — `gh release download` for the asset, `gh api` for the version — the **private-repo** path), or **`auto`** (the default: prefer `gh` when it is installed and authenticated, else fall back to HTTPS). Because `gh` honors OAuth + `GH_TOKEN`, an adapter ...\n 709|\n 710|<!-- [doc->REQ-ADAPTER-UPDATE-MESSAGE] -->\n 711|**Post-update notice — `message`.** The optional **`message`** is a plain (multi-line) human notice the adapter surfaces to stdout **only when `spt adapter update` actually applies an update** (the version changed) — never on a no-op / up-to-date run. It is read from the **newly-installed** manifest and rendered through the same inline-Markdown prose path as `spt`'s help (`**bold**`, `` `code` ``, `[text](url)`); there is no `{key}` substitution. It is avenue-agnostic (`gh_release` / `delegated` / `f...\n...\n 723|command = \"{adapter_dir}/claude-spt post-update\"   # REQUIRED, non-empty; install-dir + adapter-static subst\n*724|self_verifies = false                              # optional; attestation-only metadata (default false)\n 725|```\n 726|\n*727|- **Runs unconditionally** — the post-step fires **even when the primary avenue was a no-op** (already up to date). Its own idempotent check (e.g. `claude plugin update`) decides whether anything changes. `command` is **required** and **non-empty** (an empty `command` is refused at registration). `self_verifies` is **attestation-only metadata** (defaults `false`) — parallel to `[update].self_verifies`; it records that the post-step verifies its own outcome but **gates nothing** (the post-step is alre...\n 728|- **Published stdin seam** — spt-core feeds the post-step **one JSON line** on stdin describing the just-resolved update:\n 729|  ```json\n 730|  {\"adapter_applied\": true, \"adapter_name\": \"claude-spt\", \"profile_name\": null,\n...\n 765|[shell]\n*766|spawn = \"gamerobot-shell --link {link_token} --root {perch_dir}\"   # broker-launched; opaque command template — {perch_dir} = the instance's perch dir (REQ-SHELL-PERCH-DIR: lets the binary resolve a shell_file frame's perch-relative path; opt-in, spawn-only). Program token resolves against the adapter's install dir before PATH, and {adapter_dir} fills — see below.\n 767|ephemeral = false                  # manifest property, NOT an agent choice; ephemeral ⇒ no offline perch + no history retention\n 768|broadcast = \"subnet\"               # subnet | same-node | none\n 769|command_receipt = \"http\"           # http | stdin | relay   (how it receives agent commands)\n### CONTEXT-MEMORY.md\n 7|- **live context** (per-agent, project-independent — the agent's *mind*): identity/persona, relationship + rapport with the user, the user's cross-project preferences & working style, cross-project life threads, **learnings from work that may apply to other projects**, a thin **project roster** (one line per project the agent works on), the agent's self-model. Synced to **all** instances of the endpoint.","truncated":true,"truncatedBy":"bytes","totalLines":381,"totalBytes":52528,"outputLines":371,"outputBytes":51195,"lastLinePartial":false,"firstLineExceedsLimit":false},"linesTruncated":true,"meta":{"limits":{"columnTruncated":{"maxColumn":512}},"truncation":{"direction":"head","truncatedBy":"bytes","totalLines":381,"totalBytes":52528,"outputLines":371,"outputBytes":51195,"shownRange":{"start":1,"end":371},"nextOffset":372}}},"isError":false,"timestamp":1788999516720}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_xxrNw74MW7mfjeRdzFeuDgdT|fc_0dde72effda0ad21016aa1f773477887d2b2640c2a3f3def17","toolName":"grep","startedAt":"2026-09-10T00:18:59.394Z","args":{"path":".worktrees/286-churn/docs/KNOWN-HAZARDS.md;.worktrees/286-churn/docs/adr; .worktrees/286-churn/traceable-reqs.toml"},"intent":"Locating poison rules"},"id":"c2b3311e","parentId":"9834b7e6","timestamp":"2026-09-10T00:18:59.394Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_RF1Qhc3DCVPImF1zBhBUsns4|fc_0dde72effda0ad21016aa1f773478887d2a5cefb33a43b3d3a","toolName":"read","startedAt":"2026-09-10T00:18:59.394Z","args":{"path":".worktrees/286-churn/crates/spt-daemon/src/livehost.rs:1-120,535-654,800-1054,1290-1495"},"intent":"Reading livehost architecture"},"id":"796018ca","parentId":"c2b3311e","timestamp":"2026-09-10T00:18:59.394Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_G9b0WiwJNnWeZBU8KbBeRtc8|fc_0dde72effda0ad21016aa1f773479087d2b98e7538199af0ef","toolName":"read","startedAt":"2026-09-10T00:18:59.394Z","args":{"path":".worktrees/286-churn/crates/spt-daemon/src/conn.rs:795-1000"},"intent":"Reading connection regressions"},"id":"9ddb1fe3","parentId":"796018ca","timestamp":"2026-09-10T00:18:59.394Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_uFMSzAUodL37QSGaqIjMBY7T|fc_0dde72effda0ad21016aa1f773479887d2afcf74a9b41ed01a","toolName":"grep","startedAt":"2026-09-10T00:18:59.394Z","args":{"path":".worktrees/286-churn/crates/spt-daemon/src/brain.rs;.worktrees/286-churn/crates/spt-daemon/src/broker.rs;.worktrees/286-churn/crates/spt-daemon/src/transport.rs"},"intent":"Locating receiving dispatches"},"id":"681acceb","parentId":"9ddb1fe3","timestamp":"2026-09-10T00:18:59.394Z"}
{"type":"message","id":"31c23db2","parentId":"681acceb","timestamp":"2026-09-10T00:18:59.394Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":"**Checking file contents**"},{"type":"toolCall","id":"call_xxrNw74MW7mfjeRdzFeuDgdT|fc_0dde72effda0ad21016aa1f773477887d2b2640c2a3f3def17","name":"grep","arguments":{"i":"Locating poison rules","pattern":"SHAREDSEND|poison|HOSTED-LIVENESS|LIVEHOST|NO-BLOCKING","path":".worktrees/286-churn/docs/KNOWN-HAZARDS.md;.worktrees/286-churn/docs/adr; .worktrees/286-churn/traceable-reqs.toml","case":false,"gitignore":false,"skip":null},"intent":"Locating poison rules"},{"type":"toolCall","id":"call_RF1Qhc3DCVPImF1zBhBUsns4|fc_0dde72effda0ad21016aa1f773478887d2a5cefb33a43b3d3a","name":"read","arguments":{"i":"Reading livehost architecture","path":".worktrees/286-churn/crates/spt-daemon/src/livehost.rs:1-120,535-654,800-1054,1290-1495"},"intent":"Reading livehost architecture"},{"type":"toolCall","id":"call_G9b0WiwJNnWeZBU8KbBeRtc8|fc_0dde72effda0ad21016aa1f773479087d2b98e7538199af0ef","name":"read","arguments":{"i":"Reading connection regressions","path":".worktrees/286-churn/crates/spt-daemon/src/conn.rs:795-1000"},"intent":"Reading connection regressions"},{"type":"toolCall","id":"call_uFMSzAUodL37QSGaqIjMBY7T|fc_0dde72effda0ad21016aa1f773479887d2afcf74a9b41ed01a","name":"grep","arguments":{"i":"Locating receiving dispatches","pattern":"fn sessions|fn cold_start|KIND_SESSIONS|fn handle_conn|fn handle_brain|fn expect|fn recv|BRAIN_CONN_OPEN","path":".worktrees/286-churn/crates/spt-daemon/src/brain.rs;.worktrees/286-churn/crates/spt-daemon/src/broker.rs;.worktrees/286-churn/crates/spt-daemon/src/transport.rs","case":true,"gitignore":false,"skip":null},"intent":"Locating receiving dispatches"},{"type":"toolCall","id":"call_7neQDCkkyOP5f0csXgkrhRsA|fc_0dde72effda0ad21016aa1f77347a087d2b9b53b53d63d0bc2","name":"read","arguments":{"i":"Reading RCA context","path":".worktrees/286-churn/docs/RCA-FLEET-DAEMON-14444.md:1-145"},"intent":"Reading RCA context"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":31102,"output":441,"cacheRead":41600,"cacheWrite":0,"totalTokens":73143,"reasoningTokens":63,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1788999516748,"responseId":"resp_0dde72effda0ad21016aa1f75d39d087d2868e7e9c00c4feeb","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_0dde72effda0ad21016aa1f76614dc87d2a1e1a9ae842611ba","type":"reasoning","content":[],"encrypted_content":"gAAAAABqofdzEE9-Ve-JTZ3OLAnV55wCXQWUderL54OPQkYEdlkPXnvcPYRWdF7GQPonhWhO6gzalIzHDMMLK5Y7kMuobS3OlMTmVVnnh5FXloyeaxNJPH-oxBIbfqgZDKVEteAf_xeJScQnVX0gABZN8Hkf6Gev3hJ88_KvQe-O8ap26r1z-5Kme4k1ma7pqAZTrABffsymE9jG8qSh7ryDUOCK8l480JiKOlxbzznLn7xBo-CAs8WyIKd6UZAyfHXmkJnRn5ifsNE7JU7PrDzUNCFR1YZqvuw6aglJhwiia8nideLbmZGqieVlYf0e-tFJ6Bhy5n-FObOf6kINz5jxgm8wTxN3qUrC9GgvjE0cZ_Gmm8jkXEW5uWIWuuW5P3Y9_lM6D9hKUvuhv1hapoDn3LCx8j3r_gMG7hK2poWhMYDaue3KdRDFOvmRK-HnwCSo6L5cxSvz9gHh_fYhZnCBCkCcoVzwUqC--ptGvH_QeI3kHcx9sjD4Rfw9DdbcnflsJnanDTmVVBwD3MYX9fjszPn1REEaPhTAYSGLZ4mAfw06wEaJtpyInRxkNFnC31oAykhSUg1JzghcLd5xQDqK3NHe0Ti8_QDXYvJWIcoPXYnjcVPzpR61gA5kpEngxbYHMLuqZvDDpyPbbcO31Q82VC_0s2uZQYdVpU5cX51yyVy5mDV-pzPEa5qlwWUBTED6yplCYmFEmVM2oVzmbpdLPUti7xnk0h_2e-KogQ3YrIYioghnFXCvUfPyCx_M4y7_XFiCekr0MNdAES5GdwdFC9Un3Pekq7pjLGYBIP7FUgUgCSPsH-qNowRDKF-sKVhK_8Nc5uXInZYNVPOMKs5palLC8o3hn4StVorrKGYOkwxqeVeD645W-stR38s6YMjquwjhcj4IxMMVzfNXIHSNdKdwbCN2wQh58OMRV78c5mYH7c-oKVW8GqTvuBE1JP3eOG85S7QwOETJ-f2eT1YU9j_Eo_AeuYW0ULg9wP8ppRX-urBgVa495MyvB8TYRkU76lcmEg017i8oudDmc2_OguKofYVYMzuKGnhp1hG-f335Bb_XTTtzJ8agNIHfc0O0D0XKSM-ER2DLPukp7HJa84K_s2oaDXalv2G3xMVtGScqVHyTmykMEDmqu6OkrKcm21vVGUeVPQXRZdBRI6Ask4g3b0UMp1CLGV8Y4FjV3kyExqTWoVBEtigXY-f0yTQPLztyl6X8ku6S8Pf8pRFxvB0QIOuHltsfEQPZVkzMJmw_eh-iKbmlbcMAkKmYLMWfsFxtsLGF2kG6tDShnutILe6LOxp45x3dnyp2rldN1UWv9U7lANA8Q8yY4eBQfohwZlnqeA4Cc-M01TrMK1doT76Yzu9U1teuviZ7ApERZusLtYn58mchbburL4EZB5CA4iUlJJBWy0dsH81OAyn3fnWD99ib0Eykbo0gitHc0WfpGsxApBhQAyYBtV1PzaDJSbaWiDhHGKbRII7m_HvLIsbeKFYESXJ9-6dUScs3JZ_lx1ddS6iiLxFaIEM0U1hQN40SW6Y4A2ohFgOS5XnAp7hzcO24LnLyffpxAfpfIc2MiIk-UBA3mzTSRd7cJ3XtpukiRLegTmlt1EddFp-LKPAZQdzRlyoeu4u_4qVw8Hp1rf6fXwE=","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"summary":[{"type":"summary_text","text":"**Checking file contents**"}],"metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f773477887d2b2640c2a3f3def17","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating poison rules\",\"pattern\":\"SHAREDSEND|poison|HOSTED-LIVENESS|LIVEHOST|NO-BLOCKING\",\"path\":\".worktrees/286-churn/docs/KNOWN-HAZARDS.md;.worktrees/286-churn/docs/adr; .worktrees/286-churn/traceable-reqs.toml\",\"case\":false,\"gitignore\":false,\"skip\":null}","call_id":"call_xxrNw74MW7mfjeRdzFeuDgdT","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"grep","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f773478887d2a5cefb33a43b3d3a","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading livehost architecture\",\"path\":\".worktrees/286-churn/crates/spt-daemon/src/livehost.rs:1-120,535-654,800-1054,1290-1495\"}","call_id":"call_RF1Qhc3DCVPImF1zBhBUsns4","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"read","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f773479087d2b98e7538199af0ef","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading connection regressions\",\"path\":\".worktrees/286-churn/crates/spt-daemon/src/conn.rs:795-1000\"}","call_id":"call_G9b0WiwJNnWeZBU8KbBeRtc8","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"read","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f773479887d2afcf74a9b41ed01a","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating receiving dispatches\",\"pattern\":\"fn sessions|fn cold_start|KIND_SESSIONS|fn handle_conn|fn handle_brain|fn expect|fn recv|BRAIN_CONN_OPEN\",\"path\":\".worktrees/286-churn/crates/spt-daemon/src/brain.rs;.worktrees/286-churn/crates/spt-daemon/src/broker.rs;.worktrees/286-churn/crates/spt-daemon/src/transport.rs\",\"case\":true,\"gitignore\":false,\"skip\":null}","call_id":"call_uFMSzAUodL37QSGaqIjMBY7T","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"grep","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f77347a087d2b9b53b53d63d0bc2","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading RCA context\",\"path\":\".worktrees/286-churn/docs/RCA-FLEET-DAEMON-14444.md:1-145\"}","call_id":"call_7neQDCkkyOP5f0csXgkrhRsA","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"read","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}}]},"duration":22644.377600000007,"ttft":9361.05249999999,"completedAt":1788999539392,"contextSnapshot":{"promptTokens":72702,"nonMessageTokens":13648,"compactionEpoch":0}}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_7neQDCkkyOP5f0csXgkrhRsA|fc_0dde72effda0ad21016aa1f77347a087d2b9b53b53d63d0bc2","toolName":"read","startedAt":"2026-09-10T00:18:59.394Z","args":{"path":".worktrees/286-churn/docs/RCA-FLEET-DAEMON-14444.md:1-145"},"intent":"Reading RCA context"},"id":"6806046d","parentId":"31c23db2","timestamp":"2026-09-10T00:18:59.394Z"}
{"type":"message","id":"7e205012","parentId":"6806046d","timestamp":"2026-09-10T00:18:59.406Z","message":{"role":"toolResult","toolCallId":"call_7neQDCkkyOP5f0csXgkrhRsA|fc_0dde72effda0ad21016aa1f77347a087d2b9b53b53d63d0bc2","toolName":"read","content":[{"type":"text","text":"# RCA — fleet daemon pid 14444 died 08:03:14.87Z on hfenduleam, 2026-09-07\n\nAnalyst: hertz. No box claimed; every reading below is read-only (log reads, WMI census, gh logs).\nCause NOT named. What follows is a timeline, a list of things RULED OUT with the evidence that\nrules them out, and the two hypotheses still standing.\n\n## Timeline, measured\n\n| time (UTC) | event | source |\n|---|---|---|\n| 07:52:54 | CI job `unit (self-hosted, Windows, hfenduleam)` STARTS, PR #199 run 34097634781 | gh job timing |\n| 07:57:00.9 | step runs `cargo nextest run --workspace -E 'kind(lib) + kind(bin)'` | CI log :545 |\n| 08:03:08.0 | `Finished \\`test\\` profile ... in 6m 04s` — BUILD ends | CI log |\n| **08:03:08.0 → 08:03:15.8** | **7.8 s gap: nextest's LIST phase — 226 test binaries executed (`26` + `200` skipped)** | CI log |\n| **08:03:14.87** | **daemon pid 14444 dies. Last log line `wall_ms=1788768194871`, mid-stream, no shutdown line** | rotated daemon log |\n| 08:03:15.80 | nextest prints run ID and the FIRST test result | CI log |\n| 08:03:20.4 | new generation-0 daemon pid 48232 `daemon run --detached`, parent 59384 (a transient CLI, gone) | WMI + new log |\n| 08:03:25.4 | 11 `deadline-<session>.json` rewritten — the DAEMON_RESTART_RESUME fan-out | spt home mtimes |\n| 08:07:13 | CI job ends, SUCCESS | gh job timing |\n\n**The death falls INSIDE the 7.8-second list phase and 0.93 s BEFORE the first test cell ran.**\n\n## RULED OUT, each with the measurement that rules it out\n\n- **Any action of mine.** No kill / Stop-Process / taskkill / `spt node stop|restart` / `spt update`\n  / reap-by-path in this session; I was idle between turns at 08:03Z. (Caveat I raised rather than\n  hid: my prior reply carried a shortform dispatch the adapter delivers FOR me at a time I do not\n  control — that is the one spt.exe invocation in my name I cannot timestamp from here.)\n- **A test cell.** The tests had not started — first result 0.93 s AFTER the death.\n- **Life-before-main in a listed binary.** `#[ctor]`/`#[used]`/CRT-section/static_init: **zero**\n  matches workspace-wide. Under `--list`, libtest's harness main runs, not `spt`'s `main()`.\n- **A machine-wide kill from product code.** All three `taskkill` sites are `/PID <pid>`, never\n  `/IM`: daemon.rs:3472 (comment: \"Reap scoped (never machine-wide — shared runner)\"), daemon.rs:3525\n  (a test killing its own `ping` child), shellhost.rs:639.\n- **pid 0 blast radius.** `process_descendants(0)` returns empty at :390-392 and `kill_pid_tree(0)`\n  returns early at :431-433. Both guarded.\n- **Disk exhaustion.** C: had **103.2 GB** free.\n- **A Windows fault.** No Application-log error, no System-log entry 00:58–01:06 local, and\n  **no WER ReportArchive/ReportQueue entry** after 00:50. The process left no OS-level trace —\n  consistent with TerminateProcess or a silent exit, not a crash.\n- **A panic/abort inside the daemon.** The last 8 MB of its stderr contain no `panic`/`fatal`/\n  `abort`/OOM line; the only match was one benign `PAIR_MEET_SERVE_ERR:ws272-rig` handshake error.\n\n## STILL STANDING\n\n1. **pid reuse under mass process churn.** 226 binaries executed in 7.8 s is the textbook condition,\n   and the reuse band is LIVE on this box, measured: processes created at 08:03:21 took pids\n   **3904, 4176, 9416, 11812**, and **44** currently-live processes sit within 2000 of **14444**.\n   So 14444 is squarely inside the range Windows was handing out during the window. Any component\n   holding a stale 14444 for something of its own — liveness-check then kill — would have struck the\n   daemon. **What is missing is a killer: nobody has shown one.** Plausible, not demonstrated.\n2. **Silent self-exit.** No shutdown line, but also no evidence of one.\n\n## SEPARATE DEFECT, worth filing regardless of the cause\n\nThe dead daemon's stderr was **303,128,498 bytes (303 MB)** over ~72 h uptime (`mono_ms` 259.19M),\n~100 MB/day. In the LAST 2 MB alone: **3238** `CONN_LIFECYCLE ... event=write-start` + **3086**\n`event=transport-close poisoned=false timed_out=false`, both `role=brain`, plus **1163**\n`NET_FAMILY_GATE: binding IPv4-only`. Conn ids had reached **720837**, and `CONN_WRITE_RETIRED`\nsays conns are \"poisoned + retired, never reused\". A brain conn opened and closed thousands of times\nper 2 MB of log, under `[REQ-CONN-POISON-ATTRIBUTION]`, is its own finding.\n\n## NEXT MEASUREMENTS (cheap, none need a box held)\n\n1. **Instrument the killer, do not hunt it retroactively.** The window is reproducible: it is every\n   `cargo nextest run` list phase on this box. Watch the fleet daemon's pid across one and record\n   whether it dies again.\n2. **Ask whether any component kills by a REMEMBERED pid without re-verifying identity at kill time**\n   (image path or start time), which is the only way reuse becomes a kill. My own standing rule from\n   the orphan reap is exactly this: re-verify the path AT kill time.\n3. **Bound the log volume** independently of the cause.\n\n## releases#286 — stderr follow-up (2026-09-09)\n\n<!-- [doc->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED] -->\n\nThe historical 303 MB incident above remains the reported baseline, not a\nnewly reproduced event. A read-only sample of the currently installed\n`%LOCALAPPDATA%/spt-core/logs/daemon.stderr.log` ended at byte **4,163,267**:\nthe last **1,999,958 complete-line bytes** spanned wall stamps\n**1788997126331–1788998151874** (1025.543 s). It contained **5127** first writes\n(656,158 bytes), **5092** unpoisoned/non-timeout closes (850,622 bytes),\n**3** organic retirements, **0** deadline poisons, **107** family-gate lines\nand **108** rendezvous-up lines. Healthy start/close records alone consumed\n**75.3%** of that sample. A subsequent live append observation over\n**307.203 s** added **79,741 bytes** (**259.57 B/s**), with **104** starts,\n**104** healthy closes and **30** family gates paired with **30** rendezvous\nbinds. These are different observation windows, not a claimed constant rate.\nNo daemon was restarted or reconfigured for either measurement.\n\n**Cause discrimination.** `role=brain` identifies *every broker IPC client*,\nnot the supervised coordinator. `run_brain` already keeps its heartbeat\n`Brain` alive. In contrast, the pre-fix\n`livehost::query_live_session_endpoints` constructed `Brain::cold_start`,\ncalled `sessions`, and dropped the carrier on every invocation from the\nfive-second reconcile loop. This is a source-confirmed avoidable reopen\ncause, consistent with the recurring five-second bare-brain start/close\npairs in the live log, **not an attribution of all 5127 connections to that\ncaller**. Other short-lived callers include dispatch idle queries and\nper-operation dispatch/CLI clients; their lifetimes are not changed here.\n\nThe livehost census now owns one carrier across polls. Its `BRAIN_CONN_OPEN`\nbreadcrumb names `caller=livehost-session-query` and distinguishes\n`reason=initial-query` from `reason=previous-query-error`. Every query error\ndiscards the carrier, returns an unavailable census (never an empty\nauthoritative census), and permits a fresh connection only on the next\nscheduled query. The one-shot boot orphan census remains one-shot.\nNo read/write deadline, poison semantics, or retry cadence changed.\n\n**Family gate exonerated.** The sole diagnostic site in\n`NetEndpoint::bind_scoped` is already once per real bind attempt, not per IPC\nconnection. `pairhost::spawn_meet_rotation` binds a new derived identity for\neach attached subnet at every 30-second TOTP boundary; the observed\n`BIGNET`, `SPT_DEV`, and `SPT_MANTLE` windows account for the three lines per\nboundary. These listeners cannot share a stable identity without changing\nthe pairing protocol. Bind and diagnostic policies are intentionally\nunchanged; a process-global \"log once\" would hide later real binds.\n\n**Explicit contract amendment.** Issue #286 supersedes only the\n`REQ-CONN-POISON-ATTRIBUTION` leg-(d) requirement for unconditional\n*healthy start/close lines*. Across the broker process these become\n`CONN_LIFECYCLE_SUMMARY` totals, at most once per 60-second monotonic interval,\nflushed by the first subsequent lifecycle counter event. Starts count the\nfirst committed write attempt; closes count unpoisoned physical drops.\nThere is no timer thread and no shutdown flush, so a quiet tail remains\npending until another counter event (and can be lost on process exit).\nPoison, organic failure, poisoned close, and all role-transition records\nremain unconditional per connection. They retain conn id, role facts,\nwall/monotonic stamps **and the original first-write wall/monotonic time**;\na connection that never wrote says `first_write=none`.\n\nThe deterministic regression cells exercise real framed socket exchanges\nover both retained and fresh carriers with an isolated fixed log clock/sink,\nassert one interval summary for N healthy connections, preserve exceptional\nrecords through the same sink, and drive livehost census reuse plus recovery\nfrom a malformed reply. Focused commands (not executed in the preparation\nlane; integration owner runs them after landing):\n\n```text\ncargo test -p spt-daemon --lib conn::tests::\ncargo test -p spt-daemon --lib livehost::tests::session_poll_reuses_healthy_carrier_and_reopens_after_query_error\ncargo test -p spt-daemon --test conn_blackhole_lifecycle\ntraceable-reqs check --json\n```"}],"details":{"totalLines":144,"displayContent":{"text":"# RCA — fleet daemon pid 14444 died 08:03:14.87Z on hfenduleam, 2026-09-07\r\n\r\nAnalyst: hertz. No box claimed; every reading below is read-only (log reads, WMI census, gh logs).\r\nCause NOT named. What follows is a timeline, a list of things RULED OUT with the evidence that\r\nrules them out, and the two hypotheses still standing.\r\n\r\n## Timeline, measured\r\n\r\n| time (UTC) | event | source |\r\n|---|---|---|\r\n| 07:52:54 | CI job `unit (self-hosted, Windows, hfenduleam)` STARTS, PR #199 run 34097634781 | gh job timing |\r\n| 07:57:00.9 | step runs `cargo nextest run --workspace -E 'kind(lib) + kind(bin)'` | CI log :545 |\r\n| 08:03:08.0 | `Finished \\`test\\` profile ... in 6m 04s` — BUILD ends | CI log |\r\n| **08:03:08.0 → 08:03:15.8** | **7.8 s gap: nextest's LIST phase — 226 test binaries executed (`26` + `200` skipped)** | CI log |\r\n| **08:03:14.87** | **daemon pid 14444 dies. Last log line `wall_ms=1788768194871`, mid-stream, no shutdown line** | rotated daemon log |\r\n| 08:03:15.80 | nextest prints run ID and the FIRST test result | CI log |\r\n| 08:03:20.4 | new generation-0 daemon pid 48232 `daemon run --detached`, parent 59384 (a transient CLI, gone) | WMI + new log |\r\n| 08:03:25.4 | 11 `deadline-<session>.json` rewritten — the DAEMON_RESTART_RESUME fan-out | spt home mtimes |\r\n| 08:07:13 | CI job ends, SUCCESS | gh job timing |\r\n\r\n**The death falls INSIDE the 7.8-second list phase and 0.93 s BEFORE the first test cell ran.**\r\n\r\n## RULED OUT, each with the measurement that rules it out\r\n\r\n- **Any action of mine.** No kill / Stop-Process / taskkill / `spt node stop|restart` / `spt update`\r\n  / reap-by-path in this session; I was idle between turns at 08:03Z. (Caveat I raised rather than\r\n  hid: my prior reply carried a shortform dispatch the adapter delivers FOR me at a time I do not\r\n  control — that is the one spt.exe invocation in my name I cannot timestamp from here.)\r\n- **A test cell.** The tests had not started — first result 0.93 s AFTER the death.\r\n- **Life-before-main in a listed binary.** `#[ctor]`/`#[used]`/CRT-section/static_init: **zero**\r\n  matches workspace-wide. Under `--list`, libtest's harness main runs, not `spt`'s `main()`.\r\n- **A machine-wide kill from product code.** All three `taskkill` sites are `/PID <pid>`, never\r\n  `/IM`: daemon.rs:3472 (comment: \"Reap scoped (never machine-wide — shared runner)\"), daemon.rs:3525\r\n  (a test killing its own `ping` child), shellhost.rs:639.\r\n- **pid 0 blast radius.** `process_descendants(0)` returns empty at :390-392 and `kill_pid_tree(0)`\r\n  returns early at :431-433. Both guarded.\r\n- **Disk exhaustion.** C: had **103.2 GB** free.\r\n- **A Windows fault.** No Application-log error, no System-log entry 00:58–01:06 local, and\r\n  **no WER ReportArchive/ReportQueue entry** after 00:50. The process left no OS-level trace —\r\n  consistent with TerminateProcess or a silent exit, not a crash.\r\n- **A panic/abort inside the daemon.** The last 8 MB of its stderr contain no `panic`/`fatal`/\r\n  `abort`/OOM line; the only match was one benign `PAIR_MEET_SERVE_ERR:ws272-rig` handshake error.\r\n\r\n## STILL STANDING\r\n\r\n1. **pid reuse under mass process churn.** 226 binaries executed in 7.8 s is the textbook condition,\r\n   and the reuse band is LIVE on this box, measured: processes created at 08:03:21 took pids\r\n   **3904, 4176, 9416, 11812**, and **44** currently-live processes sit within 2000 of **14444**.\r\n   So 14444 is squarely inside the range Windows was handing out during the window. Any component\r\n   holding a stale 14444 for something of its own — liveness-check then kill — would have struck the\r\n   daemon. **What is missing is a killer: nobody has shown one.** Plausible, not demonstrated.\r\n2. **Silent self-exit.** No shutdown line, but also no evidence of one.\r\n\r\n## SEPARATE DEFECT, worth filing regardless of the cause\r\n\r\nThe dead daemon's stderr was **303,128,498 bytes (303 MB)** over ~72 h uptime (`mono_ms` 259.19M),\r\n~100 MB/day. In the LAST 2 MB alone: **3238** `CONN_LIFECYCLE ... event=write-start` + **3086**\r\n`event=transport-close poisoned=false timed_out=false`, both `role=brain`, plus **1163**\r\n`NET_FAMILY_GATE: binding IPv4-only`. Conn ids had reached **720837**, and `CONN_WRITE_RETIRED`\r\nsays conns are \"poisoned + retired, never reused\". A brain conn opened and closed thousands of times\r\nper 2 MB of log, under `[REQ-CONN-POISON-ATTRIBUTION]`, is its own finding.\r\n\r\n## NEXT MEASUREMENTS (cheap, none need a box held)\r\n\r\n1. **Instrument the killer, do not hunt it retroactively.** The window is reproducible: it is every\r\n   `cargo nextest run` list phase on this box. Watch the fleet daemon's pid across one and record\r\n   whether it dies again.\r\n2. **Ask whether any component kills by a REMEMBERED pid without re-verifying identity at kill time**\r\n   (image path or start time), which is the only way reuse becomes a kill. My own standing rule from\r\n   the orphan reap is exactly this: re-verify the path AT kill time.\r\n3. **Bound the log volume** independently of the cause.\r\n\r\n## releases#286 — stderr follow-up (2026-09-09)\r\n\r\n<!-- [doc->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED] -->\r\n\r\nThe historical 303 MB incident above remains the reported baseline, not a\r\nnewly reproduced event. A read-only sample of the currently installed\r\n`%LOCALAPPDATA%/spt-core/logs/daemon.stderr.log` ended at byte **4,163,267**:\r\nthe last **1,999,958 complete-line bytes** spanned wall stamps\r\n**1788997126331–1788998151874** (1025.543 s). It contained **5127** first writes\r\n(656,158 bytes), **5092** unpoisoned/non-timeout closes (850,622 bytes),\r\n**3** organic retirements, **0** deadline poisons, **107** family-gate lines\r\nand **108** rendezvous-up lines. Healthy start/close records alone consumed\r\n**75.3%** of that sample. A subsequent live append observation over\r\n**307.203 s** added **79,741 bytes** (**259.57 B/s**), with **104** starts,\r\n**104** healthy closes and **30** family gates paired with **30** rendezvous\r\nbinds. These are different observation windows, not a claimed constant rate.\r\nNo daemon was restarted or reconfigured for either measurement.\r\n\r\n**Cause discrimination.** `role=brain` identifies *every broker IPC client*,\r\nnot the supervised coordinator. `run_brain` already keeps its heartbeat\r\n`Brain` alive. In contrast, the pre-fix\r\n`livehost::query_live_session_endpoints` constructed `Brain::cold_start`,\r\ncalled `sessions`, and dropped the carrier on every invocation from the\r\nfive-second reconcile loop. This is a source-confirmed avoidable reopen\r\ncause, consistent with the recurring five-second bare-brain start/close\r\npairs in the live log, **not an attribution of all 5127 connections to that\r\ncaller**. Other short-lived callers include dispatch idle queries and\r\nper-operation dispatch/CLI clients; their lifetimes are not changed here.\r\n\r\nThe livehost census now owns one carrier across polls. Its `BRAIN_CONN_OPEN`\r\nbreadcrumb names `caller=livehost-session-query` and distinguishes\r\n`reason=initial-query` from `reason=previous-query-error`. Every query error\r\ndiscards the carrier, returns an unavailable census (never an empty\r\nauthoritative census), and permits a fresh connection only on the next\r\nscheduled query. The one-shot boot orphan census remains one-shot.\r\nNo read/write deadline, poison semantics, or retry cadence changed.\r\n\r\n**Family gate exonerated.** The sole diagnostic site in\r\n`NetEndpoint::bind_scoped` is already once per real bind attempt, not per IPC\r\nconnection. `pairhost::spawn_meet_rotation` binds a new derived identity for\r\neach attached subnet at every 30-second TOTP boundary; the observed\r\n`BIGNET`, `SPT_DEV`, and `SPT_MANTLE` windows account for the three lines per\r\nboundary. These listeners cannot share a stable identity without changing\r\nthe pairing protocol. Bind and diagnostic policies are intentionally\r\nunchanged; a process-global \"log once\" would hide later real binds.\r\n\r\n**Explicit contract amendment.** Issue #286 supersedes only the\r\n`REQ-CONN-POISON-ATTRIBUTION` leg-(d) requirement for unconditional\r\n*healthy start/close lines*. Across the broker process these become\r\n`CONN_LIFECYCLE_SUMMARY` totals, at most once per 60-second monotonic interval,\r\nflushed by the first subsequent lifecycle counter event. Starts count the\r\nfirst committed write attempt; closes count unpoisoned physical drops.\r\nThere is no timer thread and no shutdown flush, so a quiet tail remains\r\npending until another counter event (and can be lost on process exit).\r\nPoison, organic failure, poisoned close, and all role-transition records\r\nremain unconditional per connection. They retain conn id, role facts,\r\nwall/monotonic stamps **and the original first-write wall/monotonic time**;\r\na connection that never wrote says `first_write=none`.\r\n\r\nThe deterministic regression cells exercise real framed socket exchanges\r\nover both retained and fresh carriers with an isolated fixed log clock/sink,\r\nassert one interval summary for N healthy connections, preserve exceptional\r\nrecords through the same sink, and drive livehost census reuse plus recovery\r\nfrom a malformed reply. Focused commands (not executed in the preparation\r\nlane; integration owner runs them after landing):\r\n\r\n```text\r\ncargo test -p spt-daemon --lib conn::tests::\r\ncargo test -p spt-daemon --lib livehost::tests::session_poll_reuses_healthy_carrier_and_reopens_after_query_error\r\ncargo test -p spt-daemon --test conn_blackhole_lifecycle\r\ntraceable-reqs check --json\r\n```\r","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120,121,122,123,124,125,126,127,128,129,130,131,132,133,134,135,136,137,138,139,140,141,142,143,144]},"fileSize":9457,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\docs\\RCA-FLEET-DAEMON-14444.md"}}},"isError":false,"timestamp":1788999539406}}
{"type":"message","id":"330e41c2","parentId":"7e205012","timestamp":"2026-09-10T00:18:59.409Z","message":{"role":"toolResult","toolCallId":"call_G9b0WiwJNnWeZBU8KbBeRtc8|fc_0dde72effda0ad21016aa1f773479087d2b98e7538199af0ef","toolName":"read","content":[{"type":"text","text":"mod tests {\n…\n    use interprocess::local_socket::Stream;\n    use std::sync::atomic::AtomicU32;\n\n    /// A real (loopback local-socket) conn pair — the send half wrapped in the\n    /// primitive, the client end returned so a test can hold it alive or drop\n    /// it to force an ORGANIC write failure.\n    fn make_conn() -> (BrokerConn, Stream) {\n        make_conn_with_log(Arc::new(LifecycleLog::new()))\n    }\n\n    fn make_conn_with_log(log: Arc<LifecycleLog>) -> (BrokerConn, Stream) {\n        let (conn, client, _recv) = make_duplex_with_log(log);\n        (conn, client)\n    }\n\n    fn make_duplex_with_log(\n        log: Arc<LifecycleLog>,\n    ) -> (BrokerConn, Stream, interprocess::local_socket::RecvHalf) {\n        static SEQ: AtomicU32 = AtomicU32::new(0);\n        let name = format!(\n            \"spt-daemon-connattr-{}-{}.sock\",\n            std::process::id(),\n            SEQ.fetch_add(1, Ordering::Relaxed)\n        );\n        let listener = LocalSocketTransport::bind(&name).expect(\"bind\");\n        let client = LocalSocketTransport::connect(&name).expect(\"connect\");\n        let server = listener.accept().expect(\"accept\");\n        let (recv, send) = server.split();\n        (\n            BrokerConn::with_log(send, Duration::from_millis(2000), log),\n            client,\n            recv,\n        )\n    }\n\n    // [unit->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n    #[test]\n    fn healthy_round_trips_and_reconnections_emit_one_interval_summary() {\n        let now = Instant::now();\n        let log = Arc::new(LifecycleLog::captured(now));\n        // The real framed-write/drop path, both a persistent carrier and fresh\n        // per-poll carriers. An isolated clock/sink avoids global test races.\n        const N: u64 = 32;\n        for round in 0..N {\n            let (conn, mut client, mut recv) = make_duplex_with_log(Arc::clone(&log));\n            conn.describe(\"role=brain\");\n            for seq in 0..4 {\n                let request = Envelope::new(\"poll\", serde_json::json!([round, seq]));\n                write_frame(&mut client, &request).unwrap();\n                let received = crate::codec::read_frame(&mut recv).unwrap();\n                assert_eq!(received.payload, request.payload);\n                assert_eq!(conn.write(&request).unwrap(), ConnWrite::Done);\n                let reply = crate::codec::read_frame(&mut client).unwrap();\n                assert_eq!(reply.payload, request.payload);\n            }\n            drop(conn);\n        }\n        assert!(recover(log.capture.as_ref().unwrap()).is_empty());\n        *recover(log.clock.as_ref().unwrap()) = now + HEALTHY_LOG_INTERVAL;\n        let (conn, mut client) = make_conn_with_log(Arc::clone(&log));\n        conn.describe(\"role=brain\");\n        conn.write(&Envelope::new(\"poll\", serde_json::json!(N))).unwrap();\n        assert_eq!(crate::codec::read_frame(&mut client).unwrap().payload, serde_json::json!(N));\n        drop(conn);\n        let lines = recover(log.capture.as_ref().unwrap());\n        assert_eq!(lines.len(), 1, \"healthy volume is per interval, not per conn\");\n        assert!(lines[0].contains(&format!(\"starts={}\", N + 1)));\n        assert!(lines[0].contains(&format!(\"closes={N}\")));\n    }\n\n    // [unit->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n    // [unit->REQ-CONN-POISON-ATTRIBUTION]\n    #[test]\n    fn failures_and_role_events_bypass_healthy_aggregation_with_original_start() {\n        let log = Arc::new(LifecycleLog::captured(Instant::now()));\n        let (conn, mut client) = make_conn_with_log(Arc::clone(&log));\n        conn.describe(\"role=brain controller session=7\");\n        conn.write(&Envelope::new(\"poll\", serde_json::json!({}))).unwrap();\n        crate::codec::read_frame(&mut client).unwrap();\n        let start = *conn.inner.first_write.get().expect(\"first committed write\");\n        conn.lifecycle_event(\"writer-exit\", \"reason=write-failed\");\n        conn.inner.poison_and_cancel(Some(&io::Error::new(io::ErrorKind::BrokenPipe, \"peer gone\")));\n        // Repeated cancellation must not repeat the retirement record.\n        conn.inner.poison_and_cancel(None);\n        let id = conn.id();\n        drop(conn);\n        let lines = recover(log.capture.as_ref().unwrap());\n        assert_eq!(lines.len(), 3, \"role event, retirement and poisoned close survive\");\n        for line in lines.iter() {\n            assert!(line.contains(&format!(\"conn={id}\")));\n            assert!(line.contains(\"role=brain controller session=7\"));\n            assert!(line.contains(&format!(\"first_write_wall_ms={}\", start.wall_ms)));\n            assert!(line.contains(&format!(\"first_write_mono_ms={}\", start.mono_ms)));\n        }\n        assert!(lines[1].contains(\"CONN_WRITE_RETIRED:\"));\n        assert!(lines[1].contains(\"BrokenPipe\"));\n        assert!(lines[2].contains(\"poisoned=true timed_out=false\"));\n    }\n\n    // [unit->REQ-CONN-POISON-DIAL-SCOPE] the F-039 leg-(a) token split: the loud\n    // CONN_WRITE_POISONED token is RESERVED for the deadline (timed_out) class;\n    // an organic write failure renders the distinct CONN_WRITE_RETIRED token\n    // carrying the error kind — and (doyle amendment 1) the SAME attribution\n    // fields as POISONED: conn id, accreted role/endpoint/session facts, wall +\n    // monotonic stamps. Quieter class, never thinner.\n    // [unit->REQ-CONN-POISON-ATTRIBUTION]\n    #[test]\n    fn retirement_token_splits_deadline_from_organic_with_attribution_parity() {\n        let (conn, _client) = make_conn();\n        conn.describe(\"role=brain\");\n        conn.describe(\"controller session=7 endpoint=ep by=cli@host\");\n\n        let poisoned = conn.inner.render_retirement(true, None);\n        assert!(\n            poisoned.starts_with(\"CONN_WRITE_POISONED:\"),\n            \"deadline class keeps the loud wedge-observable token: {poisoned}\"\n        );\n        let retired = conn\n            .inner\n            .render_retirement(false, Some(&io::Error::new(io::ErrorKind::BrokenPipe, \"peer gone\")));\n        assert!(\n            retired.starts_with(\"CONN_WRITE_RETIRED:\"),\n            \"organic class emits the DISTINCT retired token: {retired}\"\n        );\n        assert!(\n            retired.contains(\"BrokenPipe\") && retired.contains(\"peer gone\"),\n            \"the organic record carries the error kind + message: {retired}\"\n        );\n        // Attribution parity: every field present on BOTH tokens.\n        for (name, line) in [(\"POISONED\", &poisoned), (\"RETIRED\", &retired)] {\n            assert!(\n                line.contains(&format!(\"conn={}\", conn.id())),\n                \"{name} carries the stable conn id: {line}\"\n            );\n            assert!(\n                line.contains(\"role=brain\") && line.contains(\"controller session=7\"),\n                \"{name} carries the accreted role/session facts: {line}\"\n            );\n            assert!(\n                line.contains(\"wall_ms=\") && line.contains(\"mono_ms=\"),\n                \"{name} carries wall + monotonic stamps: {line}\"\n            );\n        }\n    }\n\n    // [unit->REQ-CONN-POISON-DIAL-SCOPE] behavior is UNCHANGED by the token\n    // split (the REQ-HAZARD-SHAREDSEND invariant): an organic fast-fail still\n    // poisons + retires the conn — it is only the CLASS that differs (the error\n    // surfaces with its own kind, never TimedOut, and the next write refuses).\n    #[test]\n    fn organic_write_failure_still_poisons_and_is_not_the_timeout_class() {\n        let (conn, client) = make_conn();\n        drop(client); // peer dies → writes fail organically (no deadline involved)\n        let env = Envelope::new(\"attr-test\", serde_json::json!({\"n\": 1}));\n        let mut organic = None;\n        // A first write can land in the OS buffer before the peer-death is\n        // observable; loop until the failure surfaces (bounded).\n        for _ in 0..200 {\n            match conn.write(&env) {\n                Err(e) => {\n                    organic = Some(e);\n                    break;\n                }\n                Ok(_) => std::thread::yield_now(),\n            }\n        }\n        let e = organic.expect(\"a write to a dead peer eventually fails\");\n        assert_ne!(\n            e.kind(),\n            io::ErrorKind::TimedOut,\n            \"an organic failure must NOT surface as the deadline class\"\n        );\n        // Poisoned + retired: the next write refuses immediately.\n        let refused = conn.write(&env).expect_err(\"poisoned conn refuses writes\");\n        assert_eq!(refused.kind(), io::ErrorKind::BrokenPipe);\n    }\n\n    // [unit->REQ-CONN-POISON-ATTRIBUTION] the stable conn id is unique across\n    // conns (monotonic mint — Arc::ptr_eq identity does not survive a log\n    // line), and the accreted label is BOUNDED: one truncation marker at\n    // LABEL_CAP, then no further growth.\n    #[test]\n    fn conn_ids_are_unique_and_labels_stay_bounded() {\n        let (a, _ca) = make_conn();\n        let (b, _cb) = make_conn();\n        assert_ne!(a.id(), b.id(), \"each physical conn mints a distinct id\");\n\n        for i in 0..100 {\n            a.describe(&format!(\"fact-{i}-padding-padding-padding\"));\n        }\n        let label = recover(&a.inner.label).clone();\n        assert!(\n            label.len() <= LABEL_CAP + '…'.len_utf8(),\n            \"label growth is bounded at LABEL_CAP (+1 marker), got {} bytes\",\n            label.len()\n        );\n        assert!(\n            label.ends_with('…'),\n            \"the bound is marked once with an ellipsis\"\n        );\n        // The id (and the capped label) stay stable across records.\n        let one = a.inner.render_retirement(true, None);\n        let two = a.inner.render_retirement(true, None);\n        assert!(one.contains(&format!(\"conn={}\", a.id())));\n        assert!(two.contains(&format!(\"conn={}\", a.id())));\n    }\n}"}],"details":{"totalLines":1000,"displayContent":{"text":"mod tests {\n…\n    use interprocess::local_socket::Stream;\r\n    use std::sync::atomic::AtomicU32;\r\n\r\n    /// A real (loopback local-socket) conn pair — the send half wrapped in the\r\n    /// primitive, the client end returned so a test can hold it alive or drop\r\n    /// it to force an ORGANIC write failure.\r\n    fn make_conn() -> (BrokerConn, Stream) {\r\n        make_conn_with_log(Arc::new(LifecycleLog::new()))\r\n    }\r\n\r\n    fn make_conn_with_log(log: Arc<LifecycleLog>) -> (BrokerConn, Stream) {\r\n        let (conn, client, _recv) = make_duplex_with_log(log);\r\n        (conn, client)\r\n    }\r\n\r\n    fn make_duplex_with_log(\r\n        log: Arc<LifecycleLog>,\r\n    ) -> (BrokerConn, Stream, interprocess::local_socket::RecvHalf) {\r\n        static SEQ: AtomicU32 = AtomicU32::new(0);\r\n        let name = format!(\r\n            \"spt-daemon-connattr-{}-{}.sock\",\r\n            std::process::id(),\r\n            SEQ.fetch_add(1, Ordering::Relaxed)\r\n        );\r\n        let listener = LocalSocketTransport::bind(&name).expect(\"bind\");\r\n        let client = LocalSocketTransport::connect(&name).expect(\"connect\");\r\n        let server = listener.accept().expect(\"accept\");\r\n        let (recv, send) = server.split();\r\n        (\r\n            BrokerConn::with_log(send, Duration::from_millis(2000), log),\r\n            client,\r\n            recv,\r\n        )\r\n    }\r\n\r\n    // [unit->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\r\n    #[test]\r\n    fn healthy_round_trips_and_reconnections_emit_one_interval_summary() {\r\n        let now = Instant::now();\r\n        let log = Arc::new(LifecycleLog::captured(now));\r\n        // The real framed-write/drop path, both a persistent carrier and fresh\r\n        // per-poll carriers. An isolated clock/sink avoids global test races.\r\n        const N: u64 = 32;\r\n        for round in 0..N {\r\n            let (conn, mut client, mut recv) = make_duplex_with_log(Arc::clone(&log));\r\n            conn.describe(\"role=brain\");\r\n            for seq in 0..4 {\r\n                let request = Envelope::new(\"poll\", serde_json::json!([round, seq]));\r\n                write_frame(&mut client, &request).unwrap();\r\n                let received = crate::codec::read_frame(&mut recv).unwrap();\r\n                assert_eq!(received.payload, request.payload);\r\n                assert_eq!(conn.write(&request).unwrap(), ConnWrite::Done);\r\n                let reply = crate::codec::read_frame(&mut client).unwrap();\r\n                assert_eq!(reply.payload, request.payload);\r\n            }\r\n            drop(conn);\r\n        }\r\n        assert!(recover(log.capture.as_ref().unwrap()).is_empty());\r\n        *recover(log.clock.as_ref().unwrap()) = now + HEALTHY_LOG_INTERVAL;\r\n        let (conn, mut client) = make_conn_with_log(Arc::clone(&log));\r\n        conn.describe(\"role=brain\");\r\n        conn.write(&Envelope::new(\"poll\", serde_json::json!(N))).unwrap();\r\n        assert_eq!(crate::codec::read_frame(&mut client).unwrap().payload, serde_json::json!(N));\r\n        drop(conn);\r\n        let lines = recover(log.capture.as_ref().unwrap());\r\n        assert_eq!(lines.len(), 1, \"healthy volume is per interval, not per conn\");\r\n        assert!(lines[0].contains(&format!(\"starts={}\", N + 1)));\r\n        assert!(lines[0].contains(&format!(\"closes={N}\")));\r\n    }\r\n\r\n    // [unit->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\r\n    // [unit->REQ-CONN-POISON-ATTRIBUTION]\r\n    #[test]\r\n    fn failures_and_role_events_bypass_healthy_aggregation_with_original_start() {\r\n        let log = Arc::new(LifecycleLog::captured(Instant::now()));\r\n        let (conn, mut client) = make_conn_with_log(Arc::clone(&log));\r\n        conn.describe(\"role=brain controller session=7\");\r\n        conn.write(&Envelope::new(\"poll\", serde_json::json!({}))).unwrap();\r\n        crate::codec::read_frame(&mut client).unwrap();\r\n        let start = *conn.inner.first_write.get().expect(\"first committed write\");\r\n        conn.lifecycle_event(\"writer-exit\", \"reason=write-failed\");\r\n        conn.inner.poison_and_cancel(Some(&io::Error::new(io::ErrorKind::BrokenPipe, \"peer gone\")));\r\n        // Repeated cancellation must not repeat the retirement record.\r\n        conn.inner.poison_and_cancel(None);\r\n        let id = conn.id();\r\n        drop(conn);\r\n        let lines = recover(log.capture.as_ref().unwrap());\r\n        assert_eq!(lines.len(), 3, \"role event, retirement and poisoned close survive\");\r\n        for line in lines.iter() {\r\n            assert!(line.contains(&format!(\"conn={id}\")));\r\n            assert!(line.contains(\"role=brain controller session=7\"));\r\n            assert!(line.contains(&format!(\"first_write_wall_ms={}\", start.wall_ms)));\r\n            assert!(line.contains(&format!(\"first_write_mono_ms={}\", start.mono_ms)));\r\n        }\r\n        assert!(lines[1].contains(\"CONN_WRITE_RETIRED:\"));\r\n        assert!(lines[1].contains(\"BrokenPipe\"));\r\n        assert!(lines[2].contains(\"poisoned=true timed_out=false\"));\r\n    }\r\n\r\n    // [unit->REQ-CONN-POISON-DIAL-SCOPE] the F-039 leg-(a) token split: the loud\r\n    // CONN_WRITE_POISONED token is RESERVED for the deadline (timed_out) class;\r\n    // an organic write failure renders the distinct CONN_WRITE_RETIRED token\r\n    // carrying the error kind — and (doyle amendment 1) the SAME attribution\r\n    // fields as POISONED: conn id, accreted role/endpoint/session facts, wall +\r\n    // monotonic stamps. Quieter class, never thinner.\r\n    // [unit->REQ-CONN-POISON-ATTRIBUTION]\r\n    #[test]\r\n    fn retirement_token_splits_deadline_from_organic_with_attribution_parity() {\r\n        let (conn, _client) = make_conn();\r\n        conn.describe(\"role=brain\");\r\n        conn.describe(\"controller session=7 endpoint=ep by=cli@host\");\r\n\r\n        let poisoned = conn.inner.render_retirement(true, None);\r\n        assert!(\r\n            poisoned.starts_with(\"CONN_WRITE_POISONED:\"),\r\n            \"deadline class keeps the loud wedge-observable token: {poisoned}\"\r\n        );\r\n        let retired = conn\r\n            .inner\r\n            .render_retirement(false, Some(&io::Error::new(io::ErrorKind::BrokenPipe, \"peer gone\")));\r\n        assert!(\r\n            retired.starts_with(\"CONN_WRITE_RETIRED:\"),\r\n            \"organic class emits the DISTINCT retired token: {retired}\"\r\n        );\r\n        assert!(\r\n            retired.contains(\"BrokenPipe\") && retired.contains(\"peer gone\"),\r\n            \"the organic record carries the error kind + message: {retired}\"\r\n        );\r\n        // Attribution parity: every field present on BOTH tokens.\r\n        for (name, line) in [(\"POISONED\", &poisoned), (\"RETIRED\", &retired)] {\r\n            assert!(\r\n                line.contains(&format!(\"conn={}\", conn.id())),\r\n                \"{name} carries the stable conn id: {line}\"\r\n            );\r\n            assert!(\r\n                line.contains(\"role=brain\") && line.contains(\"controller session=7\"),\r\n                \"{name} carries the accreted role/session facts: {line}\"\r\n            );\r\n            assert!(\r\n                line.contains(\"wall_ms=\") && line.contains(\"mono_ms=\"),\r\n                \"{name} carries wall + monotonic stamps: {line}\"\r\n            );\r\n        }\r\n    }\r\n\r\n    // [unit->REQ-CONN-POISON-DIAL-SCOPE] behavior is UNCHANGED by the token\r\n    // split (the REQ-HAZARD-SHAREDSEND invariant): an organic fast-fail still\r\n    // poisons + retires the conn — it is only the CLASS that differs (the error\r\n    // surfaces with its own kind, never TimedOut, and the next write refuses).\r\n    #[test]\r\n    fn organic_write_failure_still_poisons_and_is_not_the_timeout_class() {\r\n        let (conn, client) = make_conn();\r\n        drop(client); // peer dies → writes fail organically (no deadline involved)\r\n        let env = Envelope::new(\"attr-test\", serde_json::json!({\"n\": 1}));\r\n        let mut organic = None;\r\n        // A first write can land in the OS buffer before the peer-death is\r\n        // observable; loop until the failure surfaces (bounded).\r\n        for _ in 0..200 {\r\n            match conn.write(&env) {\r\n                Err(e) => {\r\n                    organic = Some(e);\r\n                    break;\r\n                }\r\n                Ok(_) => std::thread::yield_now(),\r\n            }\r\n        }\r\n        let e = organic.expect(\"a write to a dead peer eventually fails\");\r\n        assert_ne!(\r\n            e.kind(),\r\n            io::ErrorKind::TimedOut,\r\n            \"an organic failure must NOT surface as the deadline class\"\r\n        );\r\n        // Poisoned + retired: the next write refuses immediately.\r\n        let refused = conn.write(&env).expect_err(\"poisoned conn refuses writes\");\r\n        assert_eq!(refused.kind(), io::ErrorKind::BrokenPipe);\r\n    }\r\n\r\n    // [unit->REQ-CONN-POISON-ATTRIBUTION] the stable conn id is unique across\r\n    // conns (monotonic mint — Arc::ptr_eq identity does not survive a log\r\n    // line), and the accreted label is BOUNDED: one truncation marker at\r\n    // LABEL_CAP, then no further growth.\r\n    #[test]\r\n    fn conn_ids_are_unique_and_labels_stay_bounded() {\r\n        let (a, _ca) = make_conn();\r\n        let (b, _cb) = make_conn();\r\n        assert_ne!(a.id(), b.id(), \"each physical conn mints a distinct id\");\r\n\r\n        for i in 0..100 {\r\n            a.describe(&format!(\"fact-{i}-padding-padding-padding\"));\r\n        }\r\n        let label = recover(&a.inner.label).clone();\r\n        assert!(\r\n            label.len() <= LABEL_CAP + '…'.len_utf8(),\r\n            \"label growth is bounded at LABEL_CAP (+1 marker), got {} bytes\",\r\n            label.len()\r\n        );\r\n        assert!(\r\n            label.ends_with('…'),\r\n            \"the bound is marked once with an ellipsis\"\r\n        );\r\n        // The id (and the capped label) stay stable across records.\r\n        let one = a.inner.render_retirement(true, None);\r\n        let two = a.inner.render_retirement(true, None);\r\n        assert!(one.contains(&format!(\"conn={}\", a.id())));\r\n        assert!(two.contains(&format!(\"conn={}\", a.id())));\r\n    }\r\n}\r","startLine":790,"lineNumbers":[790,null,794,795,796,797,798,799,800,801,802,803,804,805,806,807,808,809,810,811,812,813,814,815,816,817,818,819,820,821,822,823,824,825,826,827,828,829,830,831,832,833,834,835,836,837,838,839,840,841,842,843,844,845,846,847,848,849,850,851,852,853,854,855,856,857,858,859,860,861,862,863,864,865,866,867,868,869,870,871,872,873,874,875,876,877,878,879,880,881,882,883,884,885,886,887,888,889,890,891,892,893,894,895,896,897,898,899,900,901,902,903,904,905,906,907,908,909,910,911,912,913,914,915,916,917,918,919,920,921,922,923,924,925,926,927,928,929,930,931,932,933,934,935,936,937,938,939,940,941,942,943,944,945,946,947,948,949,950,951,952,953,954,955,956,957,958,959,960,961,962,963,964,965,966,967,968,969,970,971,972,973,974,975,976,977,978,979,980,981,982,983,984,985,986,987,988,989,990,991,992,993,994,995,996,997,998,999,1000]},"fileSize":43321,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-daemon\\src\\conn.rs"}}},"isError":false,"timestamp":1788999539409}}
{"type":"message","id":"c54a7555","parentId":"330e41c2","timestamp":"2026-09-10T00:18:59.414Z","message":{"role":"toolResult","toolCallId":"call_RF1Qhc3DCVPImF1zBhBUsns4|fc_0dde72effda0ad21016aa1f773478887d2a5cefb33a43b3d3a","toolName":"read","content":[{"type":"text","text":"//! The brain's live-agent Psyche host (M11-W0.1, REQ-DAEMON-1) — re-host the\n//! daemon-hosted LiveAgent lifecycle from disk and reconcile the host-set.\n//!\n//! This is the production caller `run_pulse_loop` lacked: the brain instantiates\n//! a [`BrainLifecycle`] per hosted **live** endpoint and drives its config-paced\n//! pulse loop on a supervised thread. It is the **harness-endpoint mirror** of\n//! [`crate::shellwake::reconcile_once`] (which re-hosts offline shells from\n//! disk): same shape — enumerate Self perches, resolve each manifest from its\n//! `info.json` `adapter` option, hold a stop flag per supervised thread, and run\n//! one bounded reconcile sweep at boot and every tick. Where the shell waker\n//! hosts **offline** shells, this hosts **online** live endpoints — `status` is\n//! daemon-authoritative liveness (CONTEXT:194 / KH 2.5), so re-hosting the\n//! online ones re-hosts exactly what was live before a brain (re)start, and\n//! never resurrects a gracefully-departed (offline) agent.\n//!\n//! **Online-only is staged behind the interim.** The interim `spt api listen`\n//! process still first-hosts its live agents today; those perches are\n//! `status`-ABSENT (not \"online\"), so this scan ignores them — no double-host in\n//! the W0.1→W0.2 window. W0.2 retires the interim and makes the brain the\n//! first-host (it sets `status=online` on a live `listen`/`bind`); from then this\n//! reconcile reacts within one tick, the same way the shell waker reacts to a\n//! CLI-process status flip. The full bind→online→brain-restart→re-host arc is\n//! proven through the real brain process by the W0.4 production int.\n// [impl->REQ-DAEMON-1]\n\nuse std::collections::{BTreeSet, HashMap};\nuse std::path::{Path, PathBuf};\nuse std::sync::atomic::{AtomicBool, Ordering};\nuse std::sync::{Arc, Mutex};\nuse std::thread::JoinHandle;\nuse std::time::Duration;\n\nuse spt_runtime::registry::AdapterRecord;\nuse spt_runtime::Manifest;\nuse spt_store::liveness::{STATUS_ONLINE, STATUS_UNBOUND};\nuse spt_store::perch::{self, ParentHint};\n\nuse crate::brain::{now_ms, Brain};\nuse crate::brainproc::StartReason;\nuse crate::config::DaemonConfig;\nuse crate::lifecycle::BrainLifecycle;\n\n/// The `InfoJson::state` tag of a live agent — only these get a daemon-hosted\n/// Psyche; a `ready_agent` (messaging-only) never does.\nconst LIVE_AGENT_STATE: &str = \"live_agent\";\n\n/// The reconcile cadence (ms): a lifecycle flip lands in a CLI process, so the\n/// daemon loop is the invariant holder — a flip is at most one tick stale\n/// (mirrors [`crate::shellwake::RECONCILE_INTERVAL_MS`]).\npub const LIVE_RECONCILE_INTERVAL_MS: u64 = 5_000;\n\n/// One brain-hosted live lifecycle: its config-paced pulse-driver thread and the\n/// held stop flag (tripped to un-host on an offline-transition). F-030 W3: the\n/// per-event psyche driver spawns NO resident child, so this holds no psyche\n/// handle and no residency anchor — a psyche is a bounded per-event turn, never a\n/// process to own/reap (the resident residency/reap machinery retired here).\nstruct HostedLife {\n    stop: Arc<AtomicBool>,\n    thread: JoinHandle<()>,\n}\n\n/// The brain's hosted live-endpoint table, keyed by endpoint id. Shared between\n/// the reconcile loop and anything that must stop a host. Mirror of\n/// [`crate::shellwake::WakeSet`].\n#[derive(Default)]\npub struct LiveSet {\n    inner: Mutex<HashMap<String, HostedLife>>,\n}\n\nimpl LiveSet {\n    pub fn new() -> Self {\n        Self::default()\n    }\n\n    /// How many lifecycles are hosted (finished driver threads pruned first).\n    pub fn len(&self) -> usize {\n        let mut map = self.inner.lock().unwrap_or_else(|p| p.into_inner());\n        map.retain(|_, h| !h.thread.is_finished());\n        map.len()\n    }\n\n    pub fn is_empty(&self) -> bool {\n        self.len() == 0\n    }\n\n    /// Is `id` currently hosted (driver thread alive)? Prunes a finished entry so\n    /// a crashed driver is re-hosted next sweep.\n    fn contains(&self, id: &str) -> bool {\n        let mut map = self.inner.lock().unwrap_or_else(|p| p.into_inner());\n        match map.get(id) {\n            Some(h) if !h.thread.is_finished() => true,\n            Some(_) => {\n                map.remove(id);\n                false\n            }\n            None => false,\n        }\n    }\n\n    fn insert(&self, id: &str, handle: HostedLife) {\n        let mut map = self.inner.lock().unwrap_or_else(|p| p.into_inner());\n        map.insert(id.to_string(), handle);\n    }\n\n    /// The currently-tabled ids (finished pruned) — the stop-side reconcile input.\n    fn ids(&self) -> Vec<String> {\n        let mut map = self.inner.lock().unwrap_or_else(|p| p.into_inner());\n        map.retain(|_, h| !h.thread.is_finished());\n        map.keys().cloned().collect()\n    }\n\n    /// Un-host one lifecycle: trip its stop flag, join the driver, de-table. The\n    /// pulse loop checks the flag each sleep-slice, so the join is prompt. We do\n    /// NOT re-stamp `offline` — the perch already left the online set (graceful\n    /// signoff cleared it, or the perch is gone); a brain *crash* deliberately\n    /// leaves `status=online` so the next brain rehydrates it (CONTEXT:194).\n    fn stop_host(&self, id: &str) {\n        let handle = {\n            let mut map = self.inner.lock().unwrap_or_else(|p| p.into_inner());\n            map.remove(id)\n        };\n…\n    }\n}\n…\nfn restart_resume_gate(\n…\n) -> RestartResume {\n    if status != Some(STATUS_ONLINE) {\n        return RestartResume::Skip;\n    }\n    if controllable != Some(true) {\n        return RestartResume::Skip;\n    }\n    if in_live_session {\n        return RestartResume::Skip;\n    }\n    if relay_present {\n        return RestartResume::Skip;\n    }\n    if custody_pid_alive {\n        return RestartResume::RefuseLivePid;\n    }\n    RestartResume::Resume\n}\n\n/// LIFECYCLE-TRUTH W3 (REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL): on daemon start, RE-RUN the\n/// harness of every previously-ONLINE spt-hosted endpoint this daemon (re)start\n/// orphaned — so a `daemon stop`+`start` (the apply notice's OWN instruction) no longer\n/// massacres them (seed #6, the field-accept blocker). Runs ONCE at\n/// [`spawn_live_host`] start, BEFORE the reconcile loop's phantom-clear\n/// ([`reconcile_hosted_liveness`]) offlines the stale-online records — which IS the \"was\n/// online at restart\" signal. A brain restart with a SURVIVING broker sees the live\n/// sessions in `live` and re-runs nothing (they are not orphaned). The re-run resumes\n/// from the last ledger session via the shared [`launch_ledger_resume`] tagged\n/// `DAEMON_RESTART`; the mind rides psyche re-host as today. Belts + gate: pure\n/// [`restart_resume_gate`]. Broker unreachable ⇒ skip entirely (never mass-respawn on a\n/// hiccup — the safe direction, same posture as [`reconcile_hosted_liveness`]).\n// [impl->REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL]\npub fn resume_restart_orphaned_endpoints(\n    owlery: &Path,\n    registered: &[(AdapterRecord, Manifest)],\n    adapters_dir: &Path,\n) {\n    let Some(live) = SessionPoll::default().query(&crate::endpoint::broker_socket_name()) else {\n        return; // broker unreachable — skip (never mass-respawn on a hiccup)\n    };\n    for id in perch::list_self_perch_ids(owlery) {\n        let perch = perch::resolve_perch_path(&id, ParentHint::Infer);\n        let Some(info) = spt_store::info::read_info(&perch) else {\n            continue;\n        };\n        let relay_present = spt_store::registry::resolve_address(&id, owlery).is_some();\n        // [impl->REQ-RESUME-CUSTODY-IDENTITY] the PAIR test, not a bare-pid probe.\n        let custody_pid_alive = resume_in_flight(&perch);\n        match restart_resume_gate(\n            info.status.as_deref(),\n            info.controllable,\n            live.contains(&id),\n            relay_present,\n            custody_pid_alive,\n        ) {\n            RestartResume::Skip => {}\n            RestartResume::RefuseLivePid => {\n                // BELT 2 breadcrumb (feeds the W4 dup-wake rig, seed #7): never re-run\n                // over a live custody process — the dup-spawn / stomped-info.json class.\n                spt_proto::emit_line_err!(\n                    \"DAEMON_RESTART_RESUME_REFUSED:{id}: custody pid still alive — refusing \\\n                     to re-run over a live process (dup guard)\"\n                );\n            }\n            RestartResume::Resume => {\n                // Resume material from the newest ledger row (session + adapter + cwd).\n                let Some(last) = spt_store::sessions::last_k(&perch, 1).into_iter().next_back()\n                else {\n                    spt_proto::emit_line_err!(\n                        \"DAEMON_RESTART_RESUME_SKIP:{id}: online spt-hosted but no ledger \\\n                         session to resume\"\n                    );\n                    continue;\n                };\n                let Some(adapter) = last.adapter.as_deref().or(info.adapter.as_deref()) else {\n                    spt_proto::emit_line_err!(\"DAEMON_RESTART_RESUME_SKIP:{id}: no adapter recorded to resume under\");\n                    continue;\n                };\n                launch_ledger_resume(\n                    &perch,\n                    &id,\n                    adapter,\n                    &last.session_id,\n                    last.cwd.as_deref(),\n                    registered,\n                    adapters_dir,\n                    \"DAEMON_RESTART\",\n                );\n            }\n        }\n    }\n}\n\n/// Host one live endpoint: sweep a stale signoff (3.2), spawn the Psyche (the\n/// daemon online-stamp makes liveness daemon-authoritative), then start the\n/// config-paced pulse driver on its own supervised thread with a held stop flag.\n/// A Psyche-spawn failure tables nothing — the next sweep retries.\n// [impl->REQ-DAEMON-1]\nfn host_one(\n    set: &Arc<LiveSet>,\n    id: &str,\n    lifecycle: BrainLifecycle,\n    session_id: String,\n    reason: StartReason,\n) {\n    lifecycle.sweep_on_start();\n    // F-030 W1+W3: NO resident Psyche is spawned. The Psyche is a PER-EVENT TURN\n    // (`BrainLifecycle::run_psyche_event_turn` on each armed pulse fire), never a\n    // long-lived process — so the child never self-resolves home/subnet and the\n    // direction-(a) multi-subnet poll-refusal churn is impossible by construction.\n    // The pulse loop STAYS as the daemon-side scheduler; only what a fire DOES\n    // changed. W3 retired the resident residency/reap machinery entirely: there is no\n    // psyche handle to own, no residency anchor to stamp, nothing to reap on un-host.\n    // [impl->REQ-PSYCHE-EPHEMERAL-DRIVER]\n    // [impl->REQ-HAZARD-PSYCHE-RESIDENCY-EXPECTATION]\n    let stop = Arc::new(AtomicBool::new(false));\n    let handle = {\n        let stop = Arc::clone(&stop);\n        std::thread::spawn(move || {\n            lifecycle.run_pulse_loop(Some(&session_id), &stop, reason, |_report| {});\n        })\n    };\n…\n}\n…\npub fn reconcile_hosted_liveness(owlery: &Path, live_sessions: &BTreeSet<String>) -> Vec<String> {\n…\n    for id in perch::list_self_perch_ids(owlery) {\n…\n\n        // Control cleanup SPLIT from offline classification (ADR-0041 decision 2,\n        // supersedes the narrower B3 REQ-PRESENCE-CONTROL-REAP-ON-EXIT gate): an\n        // endpoint ABSENT from session truth carries NO control stamps — cleared\n        // for EVERY row regardless of `state` or `controllable`, so a\n        // classification quirk (ready_agent hybrid, legacy `None`, relay) can\n        // never shield stale `controlled`/`driven_by`/`viewer_count` (the sticky\n        // stamp that survived clean-exit AND a daemon restart — perri F-b).\n        // All THREE fields clear (decision 4's no-session rule) — `attached_node`\n        // derives from them, so this clears it too. Divergence-gated: rows with\n        // nothing set (the common case, e.g. a relay that never carries stamps)\n        // write nothing. RACE-FREE: with no live session there is no controller\n        // to re-stamp concurrently (the broker single-writer invariant is\n        // uncontended). This is also the BOOT-TIME sweep (path iv): the caller\n        // runs the reconcile from boot tick 1.\n        // [impl->REQ-ENDPOINT-ONLINE-TRUTH]\n        // [impl->REQ-CONTROL-STAMP-CONVERGENCE]\n        // [impl->REQ-PRESENCE-CONTROL-REAP-ON-EXIT]\n        if !has_session\n            && (info.controlled || info.driven_by.is_some() || info.viewer_count.is_some())\n        {\n            let _ = spt_store::info::set_controlled(&perch, false);\n            let _ = spt_store::info::set_driven_by(&perch, None);\n            let _ = spt_store::info::set_viewer_count(&perch, 0);\n            spt_proto::emit_line_err!(\n                \"CONTROL_REAP_NO_SESSION:{id}: cleared stale controlled/driven_by/viewer_count (no broker session)\"\n            );\n        }\n\n        if info.state != LIVE_AGENT_STATE && info.controllable != Some(true) {\n            // Legacy hybrid self-heal (ADR-0041 decision 1): a row that is NOT a\n            // live agent AND is NOT broker-PTY-controllable yet reads `status=online`\n            // was born under the pre-W2 capability-only creator stamp\n            // (ready_agent/controllable=false). Its liveness truth is its recorded\n            // listener pid — a DEAD numeric pid means the row is an immortal phantom\n            // (`is_perch_alive` trusts `status=online` and never pid-probes it, so no\n            // reader ever heals it). Normalize it terminally, CAS-guarded on the\n            // observed session. Runs ONLY under a SUCCESSFUL broker query (the caller\n            // gates the whole pass), never on a broker hiccup.\n            //\n            // HOSTING AUTHORITY ROUTING (RC-RENDER-TRUTH v0.38.1 leg 3,\n            // REQ-HOSTING-AUTHORITY-CONTROLLABLE): ONLY the PID-model rows route here.\n            // A `controllable == Some(true)` row — even a non-`live_agent`\n            // (ready_agent) one born from an spt-hosted bind over a prior ready_agent\n            // perch — holds a real broker PTY, so its liveness truth is BROKER-SESSION\n            // truth, NOT its listener pid. Those rows fall THROUGH to the\n            // has_session/terminal-normalize block below (which offlines them when the\n            // broker hosts no session for them + the resume custody is dead). This is\n            // the fix for the immortal dead-PID ready_agent hybrid: the old\n            // `state != live_agent` routing trapped it here where\n            // `hybrid_self_heal_due`'s `controllable == Some(false)` predicate refused\n            // to heal a `Some(true)` row, and the broker-session arm was never reached.\n            // [impl->REQ-ENDPOINT-ONLINE-TRUTH]\n            // [impl->REQ-HOSTING-AUTHORITY-CONTROLLABLE]\n            // RIGHT ANSWER, AND THE REASON LIVES IN THE CONSUMER (KNOWN-HAZARDS\n            // 7.50). Probing a pid with `is_process_alive` is unsound when the\n            // ASKER holds a handle to it — and this daemon hosts the broker, so\n            // it does hold one for every broker-spawned PTY child. It is sound\n            // HERE only because `hybrid_self_heal_due` gates on\n            // `controllable == Some(false)`: harness-hosted by ADR-0041, so core\n            // holds no process, the pid belongs to the harness's own listener,\n            // and nothing here ever spawned it. **Relax that gate and this probe\n            // silently becomes load-bearing over broker-spawned pids** — unsound\n            // again, with nothing failing to warn you.\n            // `process_exists` is deliberately NOT used: it answers from the\n            // process table, which is EMPTY on snapshot-less platforms, so it\n            // would regress self-heal to \"never guess\" there. See its doc.\n            let pid_alive = spt_store::info::read_pid(&perch)\n                .map(spt_store::proc::is_process_alive);\n            if hybrid_self_heal_due(\n                info.status.as_deref(),\n                info.controllable,\n                pid_alive,\n            ) {\n                BrainLifecycle::mark_offline(&perch, Some(&info.session_id));\n                spt_proto::emit_line_err!(\n                    \"HYBRID_SELFHEAL_OFFLINE:{id}: non-live-agent row was online with a dead pid\"\n                );\n                offlined.push(id);\n            }\n            continue; // PID-model row — the broker-session offline classification below is not its model\n        }\n\n        let status = info.status.as_deref();\n        let is_unbound = status == Some(STATUS_UNBOUND);\n        if status != Some(STATUS_ONLINE) && !is_unbound {\n            continue; // already offline / interim (status-absent) — presence unchanged\n        }\n        // Controllable gate: harness-hosted relay (Some(false)) and legacy (None)\n        // are exempt. An `unbound` skeleton is ALWAYS spt-hosted (bind hasn't run\n        // yet, so `controllable` may be unset) — skip the gate for it.\n        // [impl->REQ-ENDPOINT-UNBOUND-ATTACH]\n        if !is_unbound && info.controllable != Some(true) {\n            // RELAY-DEATH CONVERGENCE (REQ-RELAY-DEATH-CONVERGENCE, hertz W0\n            // item 2 re-proven on 0.44.0). These rows are exempt from the\n            // broker-session model for a good reason — core hosts no PTY for\n            // them, so \"no session\" says nothing about them — but that left them\n            // exempt from EVERY model: the `status=online` stamp came from the\n            // `api listen` process that bound the perch, and nothing here ever\n            // asked whether that process is still alive. Kill the relay while its\n            // OWNER lives (so the listen-orphan watchdog never fires) and the\n            // endpoint renders `ONLINE - HARNESS ONLY` with ready=true alive=true\n            // over an unreachable address, forever.\n            //\n            // The verdict is NOT re-derived here: it routes through\n            // `liveness::relay_liveness`, the one predicate that reads the\n            // shipped process-identity oracle (KNOWN-HAZARDS 2.5 — liveness\n            // authority lives in one resolver). `Unproven` NEVER kills.\n            if converge_dead_relay(&id, &perch, &info, owlery) {\n                offlined.push(id);\n            }\n            continue; // harness-hosted relay / legacy — not the broker-session model\n        }\n        if !has_session {\n            // spt-hosted (online or unbound skeleton), but the broker hosts no\n            // session → the harness child is gone. CUSTODY GATE first: a\n            // wake-resume spawn in flight (live `resume.pid`) means the seat is\n            // mid-revival — normalizing now would land `suspended` under the\n            // incoming bind (which carries rest_state forward by design,\n            // REQ-HAZARD-BIND-REST-STATE-CARRY) and strand the revived seat.\n            // Skip this tick; a genuinely dead spawn probes dead next tick.\n            // [impl->REQ-REST-TERMINAL-NORMALIZE]\n            // [impl->REQ-RESUME-CUSTODY-IDENTITY] the PAIR test: a recycled pid\n            // reads NOT OURS here (the record self-heals) instead of deferring\n            // this reconcile forever over a corpse — the bug-1 field shape.\n            if resume_in_flight(&perch) {\n                spt_proto::emit_line_err!(\n                    \"LIVENESS_RECONCILE_DEFER:{id}: no broker session but a resume is in flight\"\n                );\n                continue;\n            }\n            // DEMOTE, DON'T OFFLINE, WHEN A RELAY HOLDS THE ENDPOINT\n            // (REQ-LISTEN-PRESERVES-HOSTING-TOPOLOGY, doyle gate finding 2026-07-21).\n            // No broker session says the PTY is gone; it does NOT say the endpoint is\n            // gone. A registered relay address means a harness-hosted listener holds\n            // this id — the same \"a relay means harness-hosted BY DEFINITION\"\n            // discriminator `restart_resume_gate` already trusts as BELT 1. Offlining\n            // it would mark a live, message-reachable agent dead: the record-vs-reality\n            // defect this requirement exists to kill, inverted.\n            //\n            // Why HERE and not at the broker's reap, which is the path that \"knows\"\n            // the PTY ended: the reap knows only half the question. `Some(true)`\n            // surviving the reap is exactly what lets THIS arm recognise an spt-hosted\n            // endpoint whose harness died and offline it; clearing at reap made every\n            // one of those exempt and killed the offline path outright. Only here are\n            // both facts in hand — no session AND whether anything still holds the\n            // endpoint.\n            //\n            // Demoted to `None` (\"no claim\"), not `Some(false)`: the daemon observes\n            // that no broker PTY exists, not that a harness owns the process, and\n            // `Some(false)` would route the row into `hybrid_self_heal_due`'s pid\n            // probe, which is sound only for pids core never spawned\n            // (KNOWN-HAZARDS 7.50). The demotion is not terminal — a later spt-hosted\n            // bind re-EARNS `Some(true)` — and the listener's own liveness (ready\n            // marker + relay registration, dropped on its teardown) remains its model,\n            // exactly as for any other harness-hosted row.\n            // [impl->REQ-LISTEN-PRESERVES-HOSTING-TOPOLOGY]\n            if spt_store::registry::resolve_address(&id, owlery).is_some() {\n                let _ = spt_store::info::set_controllable(&perch, None);\n                spt_proto::emit_line_err!(\n                    \"HOSTING_AUTHORITY_DEMOTED:{id}: no broker session, but a relay holds \\\n                     the endpoint — demoted to no-claim instead of offlined \\\n                     [REQ-LISTEN-PRESERVES-HOSTING-TOPOLOGY]\"\n                );\n                continue;\n            }\n            // Definitive death → terminal normalize (offline + suspended +\n            // cleared anchor, ONE write), CAS-guarded on the session we observed\n            // so a bind that lands between our read and this write is never\n            // clobbered.\n            // [impl->REQ-ENDPOINT-UNBOUND-ATTACH]\n            // [impl->REQ-REST-TERMINAL-NORMALIZE]\n            BrainLifecycle::mark_offline(&perch, Some(&info.session_id));\n            // REAP-SIDE CUSTODY CLEAR (ADR-0047 decision 1): this row is being\n            // declared dead, so whatever the custody record was still asking is\n            // answered — the record never outlives its question. The pair test\n            // above already ran (a live resume DEFERS before reaching here), so\n            // this only clears custody whose spawn is provably not in flight.\n            // [impl->REQ-RESUME-CUSTODY-IDENTITY]\n            let _ = spt_store::resume_custody::clear(&perch);\n            // W5 Gap B (driven_by self-heal): an offlined endpoint can ALSO carry\n            // a stale `driven_by = Some` — its broker session (the single writer)\n            // is gone, so the `clear_controller` re-stamp can never fire to unlatch\n            // it. The split cleanup above already cleared every no-session row;\n            // this belt stays for the same-tick read (info is a pre-cleanup\n            // snapshot). RACE-FREE for the same no-live-session reason.\n            // [impl->REQ-HAZARD-DRIVEN-BY-SELFHEAL]\n            if info.driven_by.is_some() {\n                let _ = spt_store::info::set_driven_by(&perch, None);\n                spt_proto::emit_line_err!(\"DRIVEN_BY_SELFHEAL_OFFLINE:{id}: cleared stale driven_by (no session)\");\n            }\n            spt_proto::emit_line_err!(\"LIVENESS_RECONCILE_OFFLINE:{id}: no live broker session (dead harness)\");\n            offlined.push(id);\n        }\n    }\n    offlined\n}\n\n/// The livehost session census is request/reply only, never a subscription.\n/// Retain its healthy carrier across 5s polls; discard it on EVERY query error\n/// so a failed/partially consumed reply can never contaminate the next census.\n/// `None` still means skip reconciliation, not an empty authoritative census.\n// [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n#[derive(Default)]\nstruct SessionPoll {\n    brain: Option<Brain>,\n    retry: bool,\n}\n\nimpl SessionPoll {\n    fn query(&mut self, broker_name: &str) -> Option<BTreeSet<String>> {\n        if self.brain.is_none() {\n            let reason = if self.retry { \"previous-query-error\" } else { \"initial-query\" };\n            match Brain::cold_start(broker_name, now_ms()) {\n                Ok(brain) => {\n                    spt_proto::emit_line_err!(\n                        \"BRAIN_CONN_OPEN: caller=livehost-session-query reason={reason} \\\n                         policy=reuse-until-query-error {}\",\n                        crate::conn::log_stamp()\n                    );\n                    self.brain = Some(brain);\n                }\n                Err(e) => {\n                    spt_proto::emit_line_err!(\n                        \"BRAIN_CONN_OPEN_FAIL: caller=livehost-session-query reason={reason} \\\n                         kind={:?}: {e} {}\", e.kind(), crate::conn::log_stamp()\n                    );\n                    return None;\n                }\n            }\n        }\n        match self.brain.as_mut().expect(\"connected session census\").sessions() {\n            Ok(reply) => Some(\n                reply.sessions.into_iter().map(|s| s.endpoint)\n                    .filter(|e| !e.is_empty()).collect()\n            ),\n            Err(e) => {\n                // No same-tick retry: the caller skips this census exactly as\n                // before. Next tick starts a fresh physical connection.\n                self.brain = None;\n                self.retry = true;\n                spt_proto::emit_line_err!(\n                    \"BRAIN_CONN_RETIRED: caller=livehost-session-query reason=query-error \\\n                     kind={:?}: {e} {}\", e.kind(), crate::conn::log_stamp()\n                );\n                None\n            }\n        }\n    }\n}\n\n/// The normalized program basename a perch's adapter would spawn its Psyche as —\n/// the first whitespace token of `psyche_init.command`, file-name only, run through\n/// [`normalize_basename`](spt_runtime::resolve::normalize_basename) (lowercase /\n/// `.exe`-strip / dot-stem) so it keys the same as a live exe basename. `None` for an\n…\nfn reap_expired_attachments() {\n…\n    let held = crate::servehost::with_registry_write(|| {\n…\n        if !registry_path.exists() {\n            return; // nothing has ever been served on this node\n        }\n        let mut registry = match spt_store::serving::ServingRegistry::load_at(&registry_path) {\n            Ok(registry) => registry,\n            Err(err) => {\n                eprintln!(\"SERVE_REAP_LOAD_FAIL: {err}\");\n                return;\n            }\n        };\n        let reaped = registry.reap_expired(now_ms());\n        if reaped.is_empty() {\n            return;\n        }\n        let snapshots = perch::serve_snapshots_dir();\n        let mut bytes_removed: u64 = 0;\n        for entry in &reaped {\n            if entry.kind != spt_store::serving::ServedKind::Attachment {\n                continue; // a reference entry's path is the user's own file\n            }\n            if !entry.path.starts_with(&snapshots) {\n                eprintln!(\"SERVE_REAP_SKIP_FOREIGN:{}\", entry.id);\n                continue;\n            }\n            let len = std::fs::metadata(&entry.path).map(|meta| meta.len()).unwrap_or(0);\n            match std::fs::remove_file(&entry.path) {\n                Ok(()) => bytes_removed = bytes_removed.saturating_add(len),\n                Err(err) if err.kind() == std::io::ErrorKind::NotFound => {}\n                Err(err) => eprintln!(\"SERVE_REAP_UNLINK_FAIL:{}: {err}\", entry.id),\n            }\n        }\n        if let Err(err) = registry.save_at(&registry_path) {\n            // The bytes are gone and the entries are not. Say so rather than\n            // reporting a reap that only half happened; the next tick retries.\n            eprintln!(\"SERVE_REAP_SAVE_FAIL: {err}\");\n            return;\n        }\n        eprintln!(\n            \"SERVE_REAP: reaped {} expired entries, freed {} bytes\",\n            reaped.len(),\n            bytes_removed\n        );\n    });\n    if let Err(err) = held {\n        // A poisoned writer lock is not something a pulse may paper over: say\n        // it every tick rather than reaping outside the lock.\n        eprintln!(\"SERVE_REAP_LOCK_FAIL: {err}\");\n    }\n}\n\n/// Spawn the brain's live host: one thread sweeping [`reconcile_once`] at boot\n/// and every [`LIVE_RECONCILE_INTERVAL_MS`] until `stop`. The registered set +\n/// config are re-read each sweep (adapter / pulse-period changes land between\n/// ticks). Mirror of [`crate::shellwake::spawn_wake_host`]; the stop flag is held\n/// for symmetry — the threads die with the brain process and the next brain\n/// rehydrates from disk.\n// [impl->REQ-DAEMON-1]\npub fn spawn_live_host(stop: Arc<AtomicBool>, reason: StartReason) -> JoinHandle<()> {\n    std::thread::spawn(move || {\n        let set = Arc::new(LiveSet::new());\n        {\n            // F-030 W5 ONE-SHOT legacy-resident sweep (REQ-PSYCHE-LEGACY-RESIDENT-SWEEP):\n            // BEFORE the first reconcile, bury any pre-W3 psyche WRAPPER a cross-version\n            // upgrade left ALIVE-but-hung (it registered its `<id>-psyche` ready perch then\n            // blocked forever on `spt ready --once`; the per-event driver never delivers to\n            // it). Fires exactly once at brain start (never per-reconcile) — the retired\n            // residency machinery's corpse-burial, not its resurrection.\n            // [impl->REQ-PSYCHE-LEGACY-RESIDENT-SWEEP]\n            let owlery = perch::owlery_dir();\n            let adapters_dir = perch::adapters_dir();\n            let registered = spt_runtime::registry::registered(&adapters_dir);\n            sweep_legacy_resident_psyches(&owlery, &registered, &adapters_dir);\n            // F-030 W5 ONE-SHOT `.live-bin` GC — ORDERED AFTER the sweep so any wrapper\n            // running FROM the own-copy image is killed first (its dir then unlocks).\n            gc_live_bin_dirs(&owlery);\n            // LIFECYCLE-TRUTH W3 ONE-SHOT daemon-restart endpoint survival\n            // (REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL): re-run the harness of every\n            // previously-ONLINE spt-hosted endpoint this (re)start orphaned, BEFORE the\n            // reconcile loop's phantom-clear offlines the stale-online records (the \"was\n            // online\" signal). A brain restart with a surviving broker re-runs nothing\n            // (its sessions are live, not orphaned). Belts inside the gate.\n            // [impl->REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL]\n            resume_restart_orphaned_endpoints(&owlery, &registered, &adapters_dir);\n        }\n        let mut session_poll = SessionPoll::default();\n        while !stop.load(Ordering::SeqCst) {\n            // TEST-ONLY gate `SPT_LIVEHOST_RECONCILE_DISABLE` (runtime env, DEFAULT\n            // unset = normal production, untouched): skip the reconcile poll+body so an\n            // int rig can prove REQ-UPDATE-TRIAL-DRAIN-DRIVE's core-loop KIND_SESSIONS\n            // reap-driver in ISOLATION. This loop's own `SessionPoll::query`\n            // → `brain.sessions()` (every LIVE_RECONCILE_INTERVAL_MS) would OTHERWISE\n            // drive the same broker reap and mask the fix (todlando 2026-07-09). Setting\n            // it faithfully REPRODUCES THE FIELD'S livehost-silent update-trial condition\n            // (PIN Q2: the trial-brain livehost's first reconcile tick landed around the\n            // 30s kill, so it never drove the reap) — the exact failure mode the fix\n            // makes the wedge-reap independent of. Runtime env (not cfg(test)): the rig\n            // spawns a real release `spt daemon brain` child cfg(test) can't reach — the\n            // SPT_BRAIN_WRITE_DEADLINE_MS precedent. Still honors `stop` on a tight tick.\n            // [impl->REQ-UPDATE-TRIAL-DRAIN-DRIVE]\n            if std::env::var_os(\"SPT_LIVEHOST_RECONCILE_DISABLE\").is_some() {\n                std::thread::sleep(Duration::from_millis(100));\n                continue;\n            }\n            let owlery = perch::owlery_dir();\n            let adapters_dir = perch::adapters_dir();\n            let registered = spt_runtime::registry::registered(&adapters_dir);\n            let cfg = DaemonConfig::load();\n            // B2 keystone (REQ-HAZARD-HOSTED-LIVENESS-RECONCILE) + B5 boot-gate\n            // (REQ-HAZARD-LIVEHOST-BOOT-LIVENESS-GATE): clear the status=online\n            // latch for any spt-hosted endpoint whose broker session is gone — and\n            // do it from BOOT tick 1, NOT after a grace. Runs BEFORE reconcile_once\n            // so the same-tick stop-side un-hosts the now-offline perch and the\n            // start-side never REVIVES its Psyche. The boot-gate is what stops a\n            // cold start after an unclean stop from spawning N phantom Psyches for N\n            // dead-but-latched perches. Skipped only when the broker is unreachable\n            // (None) — never mass-offline on a hiccup; safe at tick 1 because\n            // run_brain connects to the broker BEFORE this host spawns, so the query\n            // returns the REAL session set: empty at a cold start (every online\n            // controllable perch is a phantom → offlined), the surviving sessions at\n            // a brain restart (live agents stay, only the dead are offlined). The\n            // controllable==Some(true) gate inside keeps relay/legacy perches exempt.\n            // [impl->REQ-HAZARD-HOSTED-LIVENESS-RECONCILE]\n            // [impl->REQ-HAZARD-LIVEHOST-BOOT-LIVENESS-GATE]\n            if let Some(live) = session_poll.query(&crate::endpoint::broker_socket_name()) {\n                reconcile_hosted_liveness(&owlery, &live);\n            }\n            reconcile_once(&owlery, &registered, &adapters_dir, &set, &cfg, reason);\n            // W2: the attachment TTL reaper rides this same 5 s pulse — snapshot\n            // bytes and the entry that names them are retired together, so an\n            // expired attachment stops costing disk as well as stops serving.\n            // [impl->REQ-WEB-ATTACHMENT-PULL]\n            reap_expired_attachments();\n            // Sleep in slices so a stop lands promptly.\n            let mut left = LIVE_RECONCILE_INTERVAL_MS;\n            while left > 0 && !stop.load(Ordering::SeqCst) {\n                let step = left.min(100);\n                std::thread::sleep(Duration::from_millis(step));\n                left -= step;\n            }\n        }\n    })\n}\n\n#[cfg(test)]\nmod tests {\n    use super::*;\n    use crate::test_home::with_home;\n    use std::time::{Duration, Instant};\n\n    // [unit->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n    #[test]\n    fn session_poll_reuses_healthy_carrier_and_reopens_after_query_error() {\n        use crate::codec::{read_frame, write_frame};\n        use crate::frame::{Envelope, Role};\n        use crate::transport::{recv_hello, DaemonTransport, LocalSocketTransport};\n        use std::sync::atomic::AtomicU32;\n\n        static SEQ: AtomicU32 = AtomicU32::new(0);\n        let name = format!(\n            \"spt-livehost-poll-{}-{}\", std::process::id(),\n            SEQ.fetch_add(1, Ordering::Relaxed)\n        );\n        let listener = LocalSocketTransport::bind(&name).unwrap();\n        const N: usize = 16;\n        let server = std::thread::spawn(move || {\n            let mut accepted = 0;\n            let mut conn = listener.accept().unwrap();\n            accepted += 1;\n            recv_hello(&mut conn, Role::Brain).unwrap();\n            for _ in 0..N {\n                assert_eq!(read_frame(&mut conn).unwrap().kind, crate::msg::KIND_SESSIONS);\n                write_frame(&mut conn, &Envelope::new(\n                    crate::msg::KIND_SESSIONS_REPLY, serde_json::json!({\"sessions\": []})\n                )).unwrap();\n            }\n            // A malformed reply is a query error even while the transport is\n            // still alive. It must retire, not leave a desynchronized cache.\n            assert_eq!(read_frame(&mut conn).unwrap().kind, crate::msg::KIND_SESSIONS);\n            write_frame(&mut conn, &Envelope::new(\n                crate::msg::KIND_SESSIONS_REPLY, serde_json::json!({\"sessions\": false})\n            )).unwrap();\n            assert!(read_frame(&mut conn).is_err(), \"failed query closes old carrier\");\n            let mut conn = listener.accept().unwrap();\n            accepted += 1;\n            recv_hello(&mut conn, Role::Brain).unwrap();\n            assert_eq!(read_frame(&mut conn).unwrap().kind, crate::msg::KIND_SESSIONS);\n            write_frame(&mut conn, &Envelope::new(\n                crate::msg::KIND_SESSIONS_REPLY, serde_json::json!({\"sessions\": []})\n            )).unwrap();\n            accepted\n        });\n        let mut poll = SessionPoll::default();\n        for _ in 0..N {\n            assert_eq!(poll.query(&name), Some(BTreeSet::new()));\n        }\n        assert_eq!(poll.query(&name), None, \"failure is not an empty session census\");\n        assert_eq!(poll.query(&name), Some(BTreeSet::new()));\n        drop(poll);\n        assert_eq!(server.join().unwrap(), 2, \"one healthy carrier plus one recovery\");\n    }\n\n    // A no-op Psyche summarizer command: spawns + exits 0 (we never depend on its\n    // output — the daemon online-stamp is what hosts it, not the pid).\n    #[cfg(windows)]\n    const NOOP_CMD: &str = \"cmd /C rem\";\n    #[cfg(unix)]\n    const NOOP_CMD: &str = \"true\";\n…\n}"}],"details":{"displayContent":{"text":"//! The brain's live-agent Psyche host (M11-W0.1, REQ-DAEMON-1) — re-host the\n//! daemon-hosted LiveAgent lifecycle from disk and reconcile the host-set.\n//!\n//! This is the production caller `run_pulse_loop` lacked: the brain instantiates\n//! a [`BrainLifecycle`] per hosted **live** endpoint and drives its config-paced\n//! pulse loop on a supervised thread. It is the **harness-endpoint mirror** of\n//! [`crate::shellwake::reconcile_once`] (which re-hosts offline shells from\n//! disk): same shape — enumerate Self perches, resolve each manifest from its\n//! `info.json` `adapter` option, hold a stop flag per supervised thread, and run\n//! one bounded reconcile sweep at boot and every tick. Where the shell waker\n//! hosts **offline** shells, this hosts **online** live endpoints — `status` is\n//! daemon-authoritative liveness (CONTEXT:194 / KH 2.5), so re-hosting the\n//! online ones re-hosts exactly what was live before a brain (re)start, and\n//! never resurrects a gracefully-departed (offline) agent.\n//!\n//! **Online-only is staged behind the interim.** The interim `spt api listen`\n//! process still first-hosts its live agents today; those perches are\n//! `status`-ABSENT (not \"online\"), so this scan ignores them — no double-host in\n//! the W0.1→W0.2 window. W0.2 retires the interim and makes the brain the\n//! first-host (it sets `status=online` on a live `listen`/`bind`); from then this\n//! reconcile reacts within one tick, the same way the shell waker reacts to a\n//! CLI-process status flip. The full bind→online→brain-restart→re-host arc is\n//! proven through the real brain process by the W0.4 production int.\n// [impl->REQ-DAEMON-1]\n\nuse std::collections::{BTreeSet, HashMap};\nuse std::path::{Path, PathBuf};\nuse std::sync::atomic::{AtomicBool, Ordering};\nuse std::sync::{Arc, Mutex};\nuse std::thread::JoinHandle;\nuse std::time::Duration;\n\nuse spt_runtime::registry::AdapterRecord;\nuse spt_runtime::Manifest;\nuse spt_store::liveness::{STATUS_ONLINE, STATUS_UNBOUND};\nuse spt_store::perch::{self, ParentHint};\n\nuse crate::brain::{now_ms, Brain};\nuse crate::brainproc::StartReason;\nuse crate::config::DaemonConfig;\nuse crate::lifecycle::BrainLifecycle;\n\n/// The `InfoJson::state` tag of a live agent — only these get a daemon-hosted\n/// Psyche; a `ready_agent` (messaging-only) never does.\nconst LIVE_AGENT_STATE: &str = \"live_agent\";\n\n/// The reconcile cadence (ms): a lifecycle flip lands in a CLI process, so the\n/// daemon loop is the invariant holder — a flip is at most one tick stale\n/// (mirrors [`crate::shellwake::RECONCILE_INTERVAL_MS`]).\npub const LIVE_RECONCILE_INTERVAL_MS: u64 = 5_000;\n\n/// One brain-hosted live lifecycle: its config-paced pulse-driver thread and the\n/// held stop flag (tripped to un-host on an offline-transition). F-030 W3: the\n/// per-event psyche driver spawns NO resident child, so this holds no psyche\n/// handle and no residency anchor — a psyche is a bounded per-event turn, never a\n/// process to own/reap (the resident residency/reap machinery retired here).\nstruct HostedLife {\n    stop: Arc<AtomicBool>,\n    thread: JoinHandle<()>,\n}\n\n/// The brain's hosted live-endpoint table, keyed by endpoint id. Shared between\n/// the reconcile loop and anything that must stop a host. Mirror of\n/// [`crate::shellwake::WakeSet`].\n#[derive(Default)]\npub struct LiveSet {\n    inner: Mutex<HashMap<String, HostedLife>>,\n}\n\nimpl LiveSet {\n    pub fn new() -> Self {\n        Self::default()\n    }\n\n    /// How many lifecycles are hosted (finished driver threads pruned first).\n    pub fn len(&self) -> usize {\n        let mut map = self.inner.lock().unwrap_or_else(|p| p.into_inner());\n        map.retain(|_, h| !h.thread.is_finished());\n        map.len()\n    }\n\n    pub fn is_empty(&self) -> bool {\n        self.len() == 0\n    }\n\n    /// Is `id` currently hosted (driver thread alive)? Prunes a finished entry so\n    /// a crashed driver is re-hosted next sweep.\n    fn contains(&self, id: &str) -> bool {\n        let mut map = self.inner.lock().unwrap_or_else(|p| p.into_inner());\n        match map.get(id) {\n            Some(h) if !h.thread.is_finished() => true,\n            Some(_) => {\n                map.remove(id);\n                false\n            }\n            None => false,\n        }\n    }\n\n    fn insert(&self, id: &str, handle: HostedLife) {\n        let mut map = self.inner.lock().unwrap_or_else(|p| p.into_inner());\n        map.insert(id.to_string(), handle);\n    }\n\n    /// The currently-tabled ids (finished pruned) — the stop-side reconcile input.\n    fn ids(&self) -> Vec<String> {\n        let mut map = self.inner.lock().unwrap_or_else(|p| p.into_inner());\n        map.retain(|_, h| !h.thread.is_finished());\n        map.keys().cloned().collect()\n    }\n\n    /// Un-host one lifecycle: trip its stop flag, join the driver, de-table. The\n    /// pulse loop checks the flag each sleep-slice, so the join is prompt. We do\n    /// NOT re-stamp `offline` — the perch already left the online set (graceful\n    /// signoff cleared it, or the perch is gone); a brain *crash* deliberately\n    /// leaves `status=online` so the next brain rehydrates it (CONTEXT:194).\n    fn stop_host(&self, id: &str) {\n        let handle = {\n            let mut map = self.inner.lock().unwrap_or_else(|p| p.into_inner());\n            map.remove(id)\n        };\n…\n    }\n}\n…\nfn restart_resume_gate(\n…\n) -> RestartResume {\n    if status != Some(STATUS_ONLINE) {\n        return RestartResume::Skip;\n    }\n    if controllable != Some(true) {\n        return RestartResume::Skip;\n    }\n    if in_live_session {\n        return RestartResume::Skip;\n    }\n    if relay_present {\n        return RestartResume::Skip;\n    }\n    if custody_pid_alive {\n        return RestartResume::RefuseLivePid;\n    }\n    RestartResume::Resume\n}\n\n/// LIFECYCLE-TRUTH W3 (REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL): on daemon start, RE-RUN the\n/// harness of every previously-ONLINE spt-hosted endpoint this daemon (re)start\n/// orphaned — so a `daemon stop`+`start` (the apply notice's OWN instruction) no longer\n/// massacres them (seed #6, the field-accept blocker). Runs ONCE at\n/// [`spawn_live_host`] start, BEFORE the reconcile loop's phantom-clear\n/// ([`reconcile_hosted_liveness`]) offlines the stale-online records — which IS the \"was\n/// online at restart\" signal. A brain restart with a SURVIVING broker sees the live\n/// sessions in `live` and re-runs nothing (they are not orphaned). The re-run resumes\n/// from the last ledger session via the shared [`launch_ledger_resume`] tagged\n/// `DAEMON_RESTART`; the mind rides psyche re-host as today. Belts + gate: pure\n/// [`restart_resume_gate`]. Broker unreachable ⇒ skip entirely (never mass-respawn on a\n/// hiccup — the safe direction, same posture as [`reconcile_hosted_liveness`]).\n// [impl->REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL]\npub fn resume_restart_orphaned_endpoints(\n    owlery: &Path,\n    registered: &[(AdapterRecord, Manifest)],\n    adapters_dir: &Path,\n) {\n    let Some(live) = SessionPoll::default().query(&crate::endpoint::broker_socket_name()) else {\n        return; // broker unreachable — skip (never mass-respawn on a hiccup)\n    };\n    for id in perch::list_self_perch_ids(owlery) {\n        let perch = perch::resolve_perch_path(&id, ParentHint::Infer);\n        let Some(info) = spt_store::info::read_info(&perch) else {\n            continue;\n        };\n        let relay_present = spt_store::registry::resolve_address(&id, owlery).is_some();\n        // [impl->REQ-RESUME-CUSTODY-IDENTITY] the PAIR test, not a bare-pid probe.\n        let custody_pid_alive = resume_in_flight(&perch);\n        match restart_resume_gate(\n            info.status.as_deref(),\n            info.controllable,\n            live.contains(&id),\n            relay_present,\n            custody_pid_alive,\n        ) {\n            RestartResume::Skip => {}\n            RestartResume::RefuseLivePid => {\n                // BELT 2 breadcrumb (feeds the W4 dup-wake rig, seed #7): never re-run\n                // over a live custody process — the dup-spawn / stomped-info.json class.\n                spt_proto::emit_line_err!(\n                    \"DAEMON_RESTART_RESUME_REFUSED:{id}: custody pid still alive — refusing \\\n                     to re-run over a live process (dup guard)\"\n                );\n            }\n            RestartResume::Resume => {\n                // Resume material from the newest ledger row (session + adapter + cwd).\n                let Some(last) = spt_store::sessions::last_k(&perch, 1).into_iter().next_back()\n                else {\n                    spt_proto::emit_line_err!(\n                        \"DAEMON_RESTART_RESUME_SKIP:{id}: online spt-hosted but no ledger \\\n                         session to resume\"\n                    );\n                    continue;\n                };\n                let Some(adapter) = last.adapter.as_deref().or(info.adapter.as_deref()) else {\n                    spt_proto::emit_line_err!(\"DAEMON_RESTART_RESUME_SKIP:{id}: no adapter recorded to resume under\");\n                    continue;\n                };\n                launch_ledger_resume(\n                    &perch,\n                    &id,\n                    adapter,\n                    &last.session_id,\n                    last.cwd.as_deref(),\n                    registered,\n                    adapters_dir,\n                    \"DAEMON_RESTART\",\n                );\n            }\n        }\n    }\n}\n\n/// Host one live endpoint: sweep a stale signoff (3.2), spawn the Psyche (the\n/// daemon online-stamp makes liveness daemon-authoritative), then start the\n/// config-paced pulse driver on its own supervised thread with a held stop flag.\n/// A Psyche-spawn failure tables nothing — the next sweep retries.\n// [impl->REQ-DAEMON-1]\nfn host_one(\n    set: &Arc<LiveSet>,\n    id: &str,\n    lifecycle: BrainLifecycle,\n    session_id: String,\n    reason: StartReason,\n) {\n    lifecycle.sweep_on_start();\n    // F-030 W1+W3: NO resident Psyche is spawned. The Psyche is a PER-EVENT TURN\n    // (`BrainLifecycle::run_psyche_event_turn` on each armed pulse fire), never a\n    // long-lived process — so the child never self-resolves home/subnet and the\n    // direction-(a) multi-subnet poll-refusal churn is impossible by construction.\n    // The pulse loop STAYS as the daemon-side scheduler; only what a fire DOES\n    // changed. W3 retired the resident residency/reap machinery entirely: there is no\n    // psyche handle to own, no residency anchor to stamp, nothing to reap on un-host.\n    // [impl->REQ-PSYCHE-EPHEMERAL-DRIVER]\n    // [impl->REQ-HAZARD-PSYCHE-RESIDENCY-EXPECTATION]\n    let stop = Arc::new(AtomicBool::new(false));\n    let handle = {\n        let stop = Arc::clone(&stop);\n        std::thread::spawn(move || {\n            lifecycle.run_pulse_loop(Some(&session_id), &stop, reason, |_report| {});\n        })\n    };\n…\n}\n…\npub fn reconcile_hosted_liveness(owlery: &Path, live_sessions: &BTreeSet<String>) -> Vec<String> {\n…\n    for id in perch::list_self_perch_ids(owlery) {\n…\n\n        // Control cleanup SPLIT from offline classification (ADR-0041 decision 2,\n        // supersedes the narrower B3 REQ-PRESENCE-CONTROL-REAP-ON-EXIT gate): an\n        // endpoint ABSENT from session truth carries NO control stamps — cleared\n        // for EVERY row regardless of `state` or `controllable`, so a\n        // classification quirk (ready_agent hybrid, legacy `None`, relay) can\n        // never shield stale `controlled`/`driven_by`/`viewer_count` (the sticky\n        // stamp that survived clean-exit AND a daemon restart — perri F-b).\n        // All THREE fields clear (decision 4's no-session rule) — `attached_node`\n        // derives from them, so this clears it too. Divergence-gated: rows with\n        // nothing set (the common case, e.g. a relay that never carries stamps)\n        // write nothing. RACE-FREE: with no live session there is no controller\n        // to re-stamp concurrently (the broker single-writer invariant is\n        // uncontended). This is also the BOOT-TIME sweep (path iv): the caller\n        // runs the reconcile from boot tick 1.\n        // [impl->REQ-ENDPOINT-ONLINE-TRUTH]\n        // [impl->REQ-CONTROL-STAMP-CONVERGENCE]\n        // [impl->REQ-PRESENCE-CONTROL-REAP-ON-EXIT]\n        if !has_session\n            && (info.controlled || info.driven_by.is_some() || info.viewer_count.is_some())\n        {\n            let _ = spt_store::info::set_controlled(&perch, false);\n            let _ = spt_store::info::set_driven_by(&perch, None);\n            let _ = spt_store::info::set_viewer_count(&perch, 0);\n            spt_proto::emit_line_err!(\n                \"CONTROL_REAP_NO_SESSION:{id}: cleared stale controlled/driven_by/viewer_count (no broker session)\"\n            );\n        }\n\n        if info.state != LIVE_AGENT_STATE && info.controllable != Some(true) {\n            // Legacy hybrid self-heal (ADR-0041 decision 1): a row that is NOT a\n            // live agent AND is NOT broker-PTY-controllable yet reads `status=online`\n            // was born under the pre-W2 capability-only creator stamp\n            // (ready_agent/controllable=false). Its liveness truth is its recorded\n            // listener pid — a DEAD numeric pid means the row is an immortal phantom\n            // (`is_perch_alive` trusts `status=online` and never pid-probes it, so no\n            // reader ever heals it). Normalize it terminally, CAS-guarded on the\n            // observed session. Runs ONLY under a SUCCESSFUL broker query (the caller\n            // gates the whole pass), never on a broker hiccup.\n            //\n            // HOSTING AUTHORITY ROUTING (RC-RENDER-TRUTH v0.38.1 leg 3,\n            // REQ-HOSTING-AUTHORITY-CONTROLLABLE): ONLY the PID-model rows route here.\n            // A `controllable == Some(true)` row — even a non-`live_agent`\n            // (ready_agent) one born from an spt-hosted bind over a prior ready_agent\n            // perch — holds a real broker PTY, so its liveness truth is BROKER-SESSION\n            // truth, NOT its listener pid. Those rows fall THROUGH to the\n            // has_session/terminal-normalize block below (which offlines them when the\n            // broker hosts no session for them + the resume custody is dead). This is\n            // the fix for the immortal dead-PID ready_agent hybrid: the old\n            // `state != live_agent` routing trapped it here where\n            // `hybrid_self_heal_due`'s `controllable == Some(false)` predicate refused\n            // to heal a `Some(true)` row, and the broker-session arm was never reached.\n            // [impl->REQ-ENDPOINT-ONLINE-TRUTH]\n            // [impl->REQ-HOSTING-AUTHORITY-CONTROLLABLE]\n            // RIGHT ANSWER, AND THE REASON LIVES IN THE CONSUMER (KNOWN-HAZARDS\n            // 7.50). Probing a pid with `is_process_alive` is unsound when the\n            // ASKER holds a handle to it — and this daemon hosts the broker, so\n            // it does hold one for every broker-spawned PTY child. It is sound\n            // HERE only because `hybrid_self_heal_due` gates on\n            // `controllable == Some(false)`: harness-hosted by ADR-0041, so core\n            // holds no process, the pid belongs to the harness's own listener,\n            // and nothing here ever spawned it. **Relax that gate and this probe\n            // silently becomes load-bearing over broker-spawned pids** — unsound\n            // again, with nothing failing to warn you.\n            // `process_exists` is deliberately NOT used: it answers from the\n            // process table, which is EMPTY on snapshot-less platforms, so it\n            // would regress self-heal to \"never guess\" there. See its doc.\n            let pid_alive = spt_store::info::read_pid(&perch)\n                .map(spt_store::proc::is_process_alive);\n            if hybrid_self_heal_due(\n                info.status.as_deref(),\n                info.controllable,\n                pid_alive,\n            ) {\n                BrainLifecycle::mark_offline(&perch, Some(&info.session_id));\n                spt_proto::emit_line_err!(\n                    \"HYBRID_SELFHEAL_OFFLINE:{id}: non-live-agent row was online with a dead pid\"\n                );\n                offlined.push(id);\n            }\n            continue; // PID-model row — the broker-session offline classification below is not its model\n        }\n\n        let status = info.status.as_deref();\n        let is_unbound = status == Some(STATUS_UNBOUND);\n        if status != Some(STATUS_ONLINE) && !is_unbound {\n            continue; // already offline / interim (status-absent) — presence unchanged\n        }\n        // Controllable gate: harness-hosted relay (Some(false)) and legacy (None)\n        // are exempt. An `unbound` skeleton is ALWAYS spt-hosted (bind hasn't run\n        // yet, so `controllable` may be unset) — skip the gate for it.\n        // [impl->REQ-ENDPOINT-UNBOUND-ATTACH]\n        if !is_unbound && info.controllable != Some(true) {\n            // RELAY-DEATH CONVERGENCE (REQ-RELAY-DEATH-CONVERGENCE, hertz W0\n            // item 2 re-proven on 0.44.0). These rows are exempt from the\n            // broker-session model for a good reason — core hosts no PTY for\n            // them, so \"no session\" says nothing about them — but that left them\n            // exempt from EVERY model: the `status=online` stamp came from the\n            // `api listen` process that bound the perch, and nothing here ever\n            // asked whether that process is still alive. Kill the relay while its\n            // OWNER lives (so the listen-orphan watchdog never fires) and the\n            // endpoint renders `ONLINE - HARNESS ONLY` with ready=true alive=true\n            // over an unreachable address, forever.\n            //\n            // The verdict is NOT re-derived here: it routes through\n            // `liveness::relay_liveness`, the one predicate that reads the\n            // shipped process-identity oracle (KNOWN-HAZARDS 2.5 — liveness\n            // authority lives in one resolver). `Unproven` NEVER kills.\n            if converge_dead_relay(&id, &perch, &info, owlery) {\n                offlined.push(id);\n            }\n            continue; // harness-hosted relay / legacy — not the broker-session model\n        }\n        if !has_session {\n            // spt-hosted (online or unbound skeleton), but the broker hosts no\n            // session → the harness child is gone. CUSTODY GATE first: a\n            // wake-resume spawn in flight (live `resume.pid`) means the seat is\n            // mid-revival — normalizing now would land `suspended` under the\n            // incoming bind (which carries rest_state forward by design,\n            // REQ-HAZARD-BIND-REST-STATE-CARRY) and strand the revived seat.\n            // Skip this tick; a genuinely dead spawn probes dead next tick.\n            // [impl->REQ-REST-TERMINAL-NORMALIZE]\n            // [impl->REQ-RESUME-CUSTODY-IDENTITY] the PAIR test: a recycled pid\n            // reads NOT OURS here (the record self-heals) instead of deferring\n            // this reconcile forever over a corpse — the bug-1 field shape.\n            if resume_in_flight(&perch) {\n                spt_proto::emit_line_err!(\n                    \"LIVENESS_RECONCILE_DEFER:{id}: no broker session but a resume is in flight\"\n                );\n                continue;\n            }\n            // DEMOTE, DON'T OFFLINE, WHEN A RELAY HOLDS THE ENDPOINT\n            // (REQ-LISTEN-PRESERVES-HOSTING-TOPOLOGY, doyle gate finding 2026-07-21).\n            // No broker session says the PTY is gone; it does NOT say the endpoint is\n            // gone. A registered relay address means a harness-hosted listener holds\n            // this id — the same \"a relay means harness-hosted BY DEFINITION\"\n            // discriminator `restart_resume_gate` already trusts as BELT 1. Offlining\n            // it would mark a live, message-reachable agent dead: the record-vs-reality\n            // defect this requirement exists to kill, inverted.\n            //\n            // Why HERE and not at the broker's reap, which is the path that \"knows\"\n            // the PTY ended: the reap knows only half the question. `Some(true)`\n            // surviving the reap is exactly what lets THIS arm recognise an spt-hosted\n            // endpoint whose harness died and offline it; clearing at reap made every\n            // one of those exempt and killed the offline path outright. Only here are\n            // both facts in hand — no session AND whether anything still holds the\n            // endpoint.\n            //\n            // Demoted to `None` (\"no claim\"), not `Some(false)`: the daemon observes\n            // that no broker PTY exists, not that a harness owns the process, and\n            // `Some(false)` would route the row into `hybrid_self_heal_due`'s pid\n            // probe, which is sound only for pids core never spawned\n            // (KNOWN-HAZARDS 7.50). The demotion is not terminal — a later spt-hosted\n            // bind re-EARNS `Some(true)` — and the listener's own liveness (ready\n            // marker + relay registration, dropped on its teardown) remains its model,\n            // exactly as for any other harness-hosted row.\n            // [impl->REQ-LISTEN-PRESERVES-HOSTING-TOPOLOGY]\n            if spt_store::registry::resolve_address(&id, owlery).is_some() {\n                let _ = spt_store::info::set_controllable(&perch, None);\n                spt_proto::emit_line_err!(\n                    \"HOSTING_AUTHORITY_DEMOTED:{id}: no broker session, but a relay holds \\\n                     the endpoint — demoted to no-claim instead of offlined \\\n                     [REQ-LISTEN-PRESERVES-HOSTING-TOPOLOGY]\"\n                );\n                continue;\n            }\n            // Definitive death → terminal normalize (offline + suspended +\n            // cleared anchor, ONE write), CAS-guarded on the session we observed\n            // so a bind that lands between our read and this write is never\n            // clobbered.\n            // [impl->REQ-ENDPOINT-UNBOUND-ATTACH]\n            // [impl->REQ-REST-TERMINAL-NORMALIZE]\n            BrainLifecycle::mark_offline(&perch, Some(&info.session_id));\n            // REAP-SIDE CUSTODY CLEAR (ADR-0047 decision 1): this row is being\n            // declared dead, so whatever the custody record was still asking is\n            // answered — the record never outlives its question. The pair test\n            // above already ran (a live resume DEFERS before reaching here), so\n            // this only clears custody whose spawn is provably not in flight.\n            // [impl->REQ-RESUME-CUSTODY-IDENTITY]\n            let _ = spt_store::resume_custody::clear(&perch);\n            // W5 Gap B (driven_by self-heal): an offlined endpoint can ALSO carry\n            // a stale `driven_by = Some` — its broker session (the single writer)\n            // is gone, so the `clear_controller` re-stamp can never fire to unlatch\n            // it. The split cleanup above already cleared every no-session row;\n            // this belt stays for the same-tick read (info is a pre-cleanup\n            // snapshot). RACE-FREE for the same no-live-session reason.\n            // [impl->REQ-HAZARD-DRIVEN-BY-SELFHEAL]\n            if info.driven_by.is_some() {\n                let _ = spt_store::info::set_driven_by(&perch, None);\n                spt_proto::emit_line_err!(\"DRIVEN_BY_SELFHEAL_OFFLINE:{id}: cleared stale driven_by (no session)\");\n            }\n            spt_proto::emit_line_err!(\"LIVENESS_RECONCILE_OFFLINE:{id}: no live broker session (dead harness)\");\n            offlined.push(id);\n        }\n    }\n    offlined\n}\n\n/// The livehost session census is request/reply only, never a subscription.\n/// Retain its healthy carrier across 5s polls; discard it on EVERY query error\n/// so a failed/partially consumed reply can never contaminate the next census.\n/// `None` still means skip reconciliation, not an empty authoritative census.\n// [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n#[derive(Default)]\nstruct SessionPoll {\n    brain: Option<Brain>,\n    retry: bool,\n}\n\nimpl SessionPoll {\n    fn query(&mut self, broker_name: &str) -> Option<BTreeSet<String>> {\n        if self.brain.is_none() {\n            let reason = if self.retry { \"previous-query-error\" } else { \"initial-query\" };\n            match Brain::cold_start(broker_name, now_ms()) {\n                Ok(brain) => {\n                    spt_proto::emit_line_err!(\n                        \"BRAIN_CONN_OPEN: caller=livehost-session-query reason={reason} \\\n                         policy=reuse-until-query-error {}\",\n                        crate::conn::log_stamp()\n                    );\n                    self.brain = Some(brain);\n                }\n                Err(e) => {\n                    spt_proto::emit_line_err!(\n                        \"BRAIN_CONN_OPEN_FAIL: caller=livehost-session-query reason={reason} \\\n                         kind={:?}: {e} {}\", e.kind(), crate::conn::log_stamp()\n                    );\n                    return None;\n                }\n            }\n        }\n        match self.brain.as_mut().expect(\"connected session census\").sessions() {\n            Ok(reply) => Some(\n                reply.sessions.into_iter().map(|s| s.endpoint)\n                    .filter(|e| !e.is_empty()).collect()\n            ),\n            Err(e) => {\n                // No same-tick retry: the caller skips this census exactly as\n                // before. Next tick starts a fresh physical connection.\n                self.brain = None;\n                self.retry = true;\n                spt_proto::emit_line_err!(\n                    \"BRAIN_CONN_RETIRED: caller=livehost-session-query reason=query-error \\\n                     kind={:?}: {e} {}\", e.kind(), crate::conn::log_stamp()\n                );\n                None\n            }\n        }\n    }\n}\n\n/// The normalized program basename a perch's adapter would spawn its Psyche as —\n/// the first whitespace token of `psyche_init.command`, file-name only, run through\n/// [`normalize_basename`](spt_runtime::resolve::normalize_basename) (lowercase /\n/// `.exe`-strip / dot-stem) so it keys the same as a live exe basename. `None` for an\n…\nfn reap_expired_attachments() {\n…\n    let held = crate::servehost::with_registry_write(|| {\n…\n        if !registry_path.exists() {\n            return; // nothing has ever been served on this node\n        }\n        let mut registry = match spt_store::serving::ServingRegistry::load_at(&registry_path) {\n            Ok(registry) => registry,\n            Err(err) => {\n                eprintln!(\"SERVE_REAP_LOAD_FAIL: {err}\");\n                return;\n            }\n        };\n        let reaped = registry.reap_expired(now_ms());\n        if reaped.is_empty() {\n            return;\n        }\n        let snapshots = perch::serve_snapshots_dir();\n        let mut bytes_removed: u64 = 0;\n        for entry in &reaped {\n            if entry.kind != spt_store::serving::ServedKind::Attachment {\n                continue; // a reference entry's path is the user's own file\n            }\n            if !entry.path.starts_with(&snapshots) {\n                eprintln!(\"SERVE_REAP_SKIP_FOREIGN:{}\", entry.id);\n                continue;\n            }\n            let len = std::fs::metadata(&entry.path).map(|meta| meta.len()).unwrap_or(0);\n            match std::fs::remove_file(&entry.path) {\n                Ok(()) => bytes_removed = bytes_removed.saturating_add(len),\n                Err(err) if err.kind() == std::io::ErrorKind::NotFound => {}\n                Err(err) => eprintln!(\"SERVE_REAP_UNLINK_FAIL:{}: {err}\", entry.id),\n            }\n        }\n        if let Err(err) = registry.save_at(&registry_path) {\n            // The bytes are gone and the entries are not. Say so rather than\n            // reporting a reap that only half happened; the next tick retries.\n            eprintln!(\"SERVE_REAP_SAVE_FAIL: {err}\");\n            return;\n        }\n        eprintln!(\n            \"SERVE_REAP: reaped {} expired entries, freed {} bytes\",\n            reaped.len(),\n            bytes_removed\n        );\n    });\n    if let Err(err) = held {\n        // A poisoned writer lock is not something a pulse may paper over: say\n        // it every tick rather than reaping outside the lock.\n        eprintln!(\"SERVE_REAP_LOCK_FAIL: {err}\");\n    }\n}\n\n/// Spawn the brain's live host: one thread sweeping [`reconcile_once`] at boot\n/// and every [`LIVE_RECONCILE_INTERVAL_MS`] until `stop`. The registered set +\n/// config are re-read each sweep (adapter / pulse-period changes land between\n/// ticks). Mirror of [`crate::shellwake::spawn_wake_host`]; the stop flag is held\n/// for symmetry — the threads die with the brain process and the next brain\n/// rehydrates from disk.\n// [impl->REQ-DAEMON-1]\npub fn spawn_live_host(stop: Arc<AtomicBool>, reason: StartReason) -> JoinHandle<()> {\n    std::thread::spawn(move || {\n        let set = Arc::new(LiveSet::new());\n        {\n            // F-030 W5 ONE-SHOT legacy-resident sweep (REQ-PSYCHE-LEGACY-RESIDENT-SWEEP):\n            // BEFORE the first reconcile, bury any pre-W3 psyche WRAPPER a cross-version\n            // upgrade left ALIVE-but-hung (it registered its `<id>-psyche` ready perch then\n            // blocked forever on `spt ready --once`; the per-event driver never delivers to\n            // it). Fires exactly once at brain start (never per-reconcile) — the retired\n            // residency machinery's corpse-burial, not its resurrection.\n            // [impl->REQ-PSYCHE-LEGACY-RESIDENT-SWEEP]\n            let owlery = perch::owlery_dir();\n            let adapters_dir = perch::adapters_dir();\n            let registered = spt_runtime::registry::registered(&adapters_dir);\n            sweep_legacy_resident_psyches(&owlery, &registered, &adapters_dir);\n            // F-030 W5 ONE-SHOT `.live-bin` GC — ORDERED AFTER the sweep so any wrapper\n            // running FROM the own-copy image is killed first (its dir then unlocks).\n            gc_live_bin_dirs(&owlery);\n            // LIFECYCLE-TRUTH W3 ONE-SHOT daemon-restart endpoint survival\n            // (REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL): re-run the harness of every\n            // previously-ONLINE spt-hosted endpoint this (re)start orphaned, BEFORE the\n            // reconcile loop's phantom-clear offlines the stale-online records (the \"was\n            // online\" signal). A brain restart with a surviving broker re-runs nothing\n            // (its sessions are live, not orphaned). Belts inside the gate.\n            // [impl->REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL]\n            resume_restart_orphaned_endpoints(&owlery, &registered, &adapters_dir);\n        }\n        let mut session_poll = SessionPoll::default();\n        while !stop.load(Ordering::SeqCst) {\n            // TEST-ONLY gate `SPT_LIVEHOST_RECONCILE_DISABLE` (runtime env, DEFAULT\n            // unset = normal production, untouched): skip the reconcile poll+body so an\n            // int rig can prove REQ-UPDATE-TRIAL-DRAIN-DRIVE's core-loop KIND_SESSIONS\n            // reap-driver in ISOLATION. This loop's own `SessionPoll::query`\n            // → `brain.sessions()` (every LIVE_RECONCILE_INTERVAL_MS) would OTHERWISE\n            // drive the same broker reap and mask the fix (todlando 2026-07-09). Setting\n            // it faithfully REPRODUCES THE FIELD'S livehost-silent update-trial condition\n            // (PIN Q2: the trial-brain livehost's first reconcile tick landed around the\n            // 30s kill, so it never drove the reap) — the exact failure mode the fix\n            // makes the wedge-reap independent of. Runtime env (not cfg(test)): the rig\n            // spawns a real release `spt daemon brain` child cfg(test) can't reach — the\n            // SPT_BRAIN_WRITE_DEADLINE_MS precedent. Still honors `stop` on a tight tick.\n            // [impl->REQ-UPDATE-TRIAL-DRAIN-DRIVE]\n            if std::env::var_os(\"SPT_LIVEHOST_RECONCILE_DISABLE\").is_some() {\n                std::thread::sleep(Duration::from_millis(100));\n                continue;\n            }\n            let owlery = perch::owlery_dir();\n            let adapters_dir = perch::adapters_dir();\n            let registered = spt_runtime::registry::registered(&adapters_dir);\n            let cfg = DaemonConfig::load();\n            // B2 keystone (REQ-HAZARD-HOSTED-LIVENESS-RECONCILE) + B5 boot-gate\n            // (REQ-HAZARD-LIVEHOST-BOOT-LIVENESS-GATE): clear the status=online\n            // latch for any spt-hosted endpoint whose broker session is gone — and\n            // do it from BOOT tick 1, NOT after a grace. Runs BEFORE reconcile_once\n            // so the same-tick stop-side un-hosts the now-offline perch and the\n            // start-side never REVIVES its Psyche. The boot-gate is what stops a\n            // cold start after an unclean stop from spawning N phantom Psyches for N\n            // dead-but-latched perches. Skipped only when the broker is unreachable\n            // (None) — never mass-offline on a hiccup; safe at tick 1 because\n            // run_brain connects to the broker BEFORE this host spawns, so the query\n            // returns the REAL session set: empty at a cold start (every online\n            // controllable perch is a phantom → offlined), the surviving sessions at\n            // a brain restart (live agents stay, only the dead are offlined). The\n            // controllable==Some(true) gate inside keeps relay/legacy perches exempt.\n            // [impl->REQ-HAZARD-HOSTED-LIVENESS-RECONCILE]\n            // [impl->REQ-HAZARD-LIVEHOST-BOOT-LIVENESS-GATE]\n            if let Some(live) = session_poll.query(&crate::endpoint::broker_socket_name()) {\n                reconcile_hosted_liveness(&owlery, &live);\n            }\n            reconcile_once(&owlery, &registered, &adapters_dir, &set, &cfg, reason);\n            // W2: the attachment TTL reaper rides this same 5 s pulse — snapshot\n            // bytes and the entry that names them are retired together, so an\n            // expired attachment stops costing disk as well as stops serving.\n            // [impl->REQ-WEB-ATTACHMENT-PULL]\n            reap_expired_attachments();\n            // Sleep in slices so a stop lands promptly.\n            let mut left = LIVE_RECONCILE_INTERVAL_MS;\n            while left > 0 && !stop.load(Ordering::SeqCst) {\n                let step = left.min(100);\n                std::thread::sleep(Duration::from_millis(step));\n                left -= step;\n            }\n        }\n    })\n}\n\n#[cfg(test)]\nmod tests {\n    use super::*;\n    use crate::test_home::with_home;\n    use std::time::{Duration, Instant};\n\n    // [unit->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n    #[test]\n    fn session_poll_reuses_healthy_carrier_and_reopens_after_query_error() {\n        use crate::codec::{read_frame, write_frame};\n        use crate::frame::{Envelope, Role};\n        use crate::transport::{recv_hello, DaemonTransport, LocalSocketTransport};\n        use std::sync::atomic::AtomicU32;\n\n        static SEQ: AtomicU32 = AtomicU32::new(0);\n        let name = format!(\n            \"spt-livehost-poll-{}-{}\", std::process::id(),\n            SEQ.fetch_add(1, Ordering::Relaxed)\n        );\n        let listener = LocalSocketTransport::bind(&name).unwrap();\n        const N: usize = 16;\n        let server = std::thread::spawn(move || {\n            let mut accepted = 0;\n            let mut conn = listener.accept().unwrap();\n            accepted += 1;\n            recv_hello(&mut conn, Role::Brain).unwrap();\n            for _ in 0..N {\n                assert_eq!(read_frame(&mut conn).unwrap().kind, crate::msg::KIND_SESSIONS);\n                write_frame(&mut conn, &Envelope::new(\n                    crate::msg::KIND_SESSIONS_REPLY, serde_json::json!({\"sessions\": []})\n                )).unwrap();\n            }\n            // A malformed reply is a query error even while the transport is\n            // still alive. It must retire, not leave a desynchronized cache.\n            assert_eq!(read_frame(&mut conn).unwrap().kind, crate::msg::KIND_SESSIONS);\n            write_frame(&mut conn, &Envelope::new(\n                crate::msg::KIND_SESSIONS_REPLY, serde_json::json!({\"sessions\": false})\n            )).unwrap();\n            assert!(read_frame(&mut conn).is_err(), \"failed query closes old carrier\");\n            let mut conn = listener.accept().unwrap();\n            accepted += 1;\n            recv_hello(&mut conn, Role::Brain).unwrap();\n            assert_eq!(read_frame(&mut conn).unwrap().kind, crate::msg::KIND_SESSIONS);\n            write_frame(&mut conn, &Envelope::new(\n                crate::msg::KIND_SESSIONS_REPLY, serde_json::json!({\"sessions\": []})\n            )).unwrap();\n            accepted\n        });\n        let mut poll = SessionPoll::default();\n        for _ in 0..N {\n            assert_eq!(poll.query(&name), Some(BTreeSet::new()));\n        }\n        assert_eq!(poll.query(&name), None, \"failure is not an empty session census\");\n        assert_eq!(poll.query(&name), Some(BTreeSet::new()));\n        drop(poll);\n        assert_eq!(server.join().unwrap(), 2, \"one healthy carrier plus one recovery\");\n    }\n\n    // A no-op Psyche summarizer command: spawns + exits 0 (we never depend on its\n    // output — the daemon online-stamp is what hosts it, not the pid).\n    #[cfg(windows)]\n    const NOOP_CMD: &str = \"cmd /C rem\";\n    #[cfg(unix)]\n    const NOOP_CMD: &str = \"true\";\n…\n}","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120,121,null,130,131,null,529,null,535,536,537,538,539,540,541,542,543,544,545,546,547,548,549,550,551,552,553,554,555,556,557,558,559,560,561,562,563,564,565,566,567,568,569,570,571,572,573,574,575,576,577,578,579,580,581,582,583,584,585,586,587,588,589,590,591,592,593,594,595,596,597,598,599,600,601,602,603,604,605,606,607,608,609,610,611,612,613,614,615,616,617,618,619,620,621,622,623,624,625,626,627,628,629,630,631,632,633,634,635,636,637,638,639,640,641,642,643,644,645,646,647,648,649,650,651,652,653,654,655,656,null,658,null,792,null,794,null,800,801,802,803,804,805,806,807,808,809,810,811,812,813,814,815,816,817,818,819,820,821,822,823,824,825,826,827,828,829,830,831,832,833,834,835,836,837,838,839,840,841,842,843,844,845,846,847,848,849,850,851,852,853,854,855,856,857,858,859,860,861,862,863,864,865,866,867,868,869,870,871,872,873,874,875,876,877,878,879,880,881,882,883,884,885,886,887,888,889,890,891,892,893,894,895,896,897,898,899,900,901,902,903,904,905,906,907,908,909,910,911,912,913,914,915,916,917,918,919,920,921,922,923,924,925,926,927,928,929,930,931,932,933,934,935,936,937,938,939,940,941,942,943,944,945,946,947,948,949,950,951,952,953,954,955,956,957,958,959,960,961,962,963,964,965,966,967,968,969,970,971,972,973,974,975,976,977,978,979,980,981,982,983,984,985,986,987,988,989,990,991,992,993,994,995,996,997,998,999,1000,1001,1002,1003,1004,1005,1006,1007,1008,1009,1010,1011,1012,1013,1014,1015,1016,1017,1018,1019,1020,1021,1022,1023,1024,1025,1026,1027,1028,1029,1030,1031,1032,1033,1034,1035,1036,1037,1038,1039,1040,1041,1042,1043,1044,1045,1046,1047,1048,1049,1050,1051,1052,1053,1054,null,1281,null,1287,null,1289,1290,1291,1292,1293,1294,1295,1296,1297,1298,1299,1300,1301,1302,1303,1304,1305,1306,1307,1308,1309,1310,1311,1312,1313,1314,1315,1316,1317,1318,1319,1320,1321,1322,1323,1324,1325,1326,1327,1328,1329,1330,1331,1332,1333,1334,1335,1336,1337,1338,1339,1340,1341,1342,1343,1344,1345,1346,1347,1348,1349,1350,1351,1352,1353,1354,1355,1356,1357,1358,1359,1360,1361,1362,1363,1364,1365,1366,1367,1368,1369,1370,1371,1372,1373,1374,1375,1376,1377,1378,1379,1380,1381,1382,1383,1384,1385,1386,1387,1388,1389,1390,1391,1392,1393,1394,1395,1396,1397,1398,1399,1400,1401,1402,1403,1404,1405,1406,1407,1408,1409,1410,1411,1412,1413,1414,1415,1416,1417,1418,1419,1420,1421,1422,1423,1424,1425,1426,1427,1428,1429,1430,1431,1432,1433,1434,1435,1436,1437,1438,1439,1440,1441,1442,1443,1444,1445,1446,1447,1448,1449,1450,1451,1452,1453,1454,1455,1456,1457,1458,1459,1460,1461,1462,1463,1464,1465,1466,1467,1468,1469,1470,1471,1472,1473,1474,1475,1476,1477,1478,1479,1480,1481,1482,1483,1484,1485,1486,1487,1488,1489,1490,1491,1492,1493,1494,1495,null,2917]},"fileSize":150325,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-daemon\\src\\livehost.rs"}}},"isError":false,"timestamp":1788999539414}}
{"type":"message","id":"b17cf158","parentId":"c54a7555","timestamp":"2026-09-10T00:18:59.417Z","message":{"role":"toolResult","toolCallId":"call_uFMSzAUodL37QSGaqIjMBY7T|fc_0dde72effda0ad21016aa1f773479887d2afcf74a9b41ed01a","toolName":"grep","content":[{"type":"text","text":"# .worktrees/286-churn/crates/spt-daemon/src/\n## brain.rs\n 57|    KIND_MET_MEMBER, KIND_NET_STREAM_SEND, KIND_NET_STREAM_SUBSCRIBE, KIND_OUTPUT, KIND_PAIR_CODE_SUBMIT, KIND_PAIR_JOIN, KIND_PAIR_JOINED, KIND_PAIR_MEET,\n*58|    KIND_RESIZE, KIND_SESSIONS, KIND_SESSIONS_REPLY, KIND_SIZE, KIND_SPAWN, KIND_SPAWNED,\n 59|    KIND_SPAWN_CONFLICT, KIND_SPAWN_FRESH,\n 60|    KIND_SUBSCRIBE, KIND_SUBSCRIBED, KIND_UNSUBSCRIBE, UnsubscribeReq,\n 61|    BringUpReq, BroughtUpReply, KIND_BRING_UP, KIND_BROUGHT_UP,\n...\n 446|    /// except the peer pump, which uses [`Brain::cold_start_pump`].\n*447|    pub fn cold_start(name: &str, now_ms: u64) -> io::Result<Self> {\n 448|        let mut conn = LocalSocketTransport::connect(name)?;\n 449|        send_hello(&mut conn, Role::Brain)?;\n 450|        Ok(Brain {\n...\n 479|    // [impl->REQ-RC-DISPLAY-SOLE-WRITER]\n*480|    pub fn cold_start_pump(\n 481|        name: &str,\n 482|        now_ms: u64,\n 483|        io_timeout: Duration,\n...\n 1268|                }\n*1269|                KIND_SESSIONS_REPLY => {\n 1270|                    let ev: SessionsReply = serde_json::from_value(env.payload)\n 1271|                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n 1272|                    return Ok(BrokerEvent::Sessions(ev));\n...\n 1632|    /// `delivered_through`-blind, broker.rs). Drives the serve_attach controller\n*1633|    /// gap-resume — no mid-stream `KIND_SESSIONS` round-trip (which would re-fatal\n 1634|    /// on the same gap and discard Output).\n 1635|    // [impl->REQ-HAZARD-CONTROLLER-GAP-RESUME]\n 1636|    pub fn controller_resume_floor(&self) -> u64 {\n...\n 2237|    /// dispatcher's session→endpoint resolution (D9-1).\n*2238|    pub fn sessions(&mut self) -> io::Result<SessionsReply> {\n*2239|        self.send(KIND_SESSIONS, serde_json::Value::Null)?;\n 2240|        loop {\n 2241|            match self.read_event()? {\n 2242|                BrokerEvent::Sessions(ev) => return Ok(ev),\n## broker.rs\n 71|    KIND_NET_STREAM_UNSUBSCRIBE, KIND_NET_STREAM_UNSUBSCRIBED,\n*72|    KIND_MET_MEMBER, KIND_PAIR_CODE_SUBMIT, KIND_PAIR_JOIN, KIND_PAIR_JOINED, KIND_PAIR_MEET, KIND_RESIZE, KIND_SESSIONS, KIND_SESSIONS_REPLY, KIND_SPAWN,\n 73|    BringUpReq, BroughtUpReply, BRING_UP_ADMITTED, BRING_UP_ADMIT_PREFIX, BRING_UP_ALREADY_LIVE, BRING_UP_REFUSED, KIND_BRING_UP, KIND_BROUGHT_UP,\n 74|    SealCeremonyReq, SealCeremonyReply, SealCeremonyOpenEvent, SealCeremonyCodeReq, SealCeremonyResultEvent,\n 75|    KIND_SEAL_CEREMONY, KIND_SEAL_CEREMONY_REPLY, KIND_SEAL_CEREMONY_OPEN, KIND_SEAL_CEREMONY_CODE, KIND_SEAL_CEREMONY_RESULT,\n...\n 211|/// write — the brain subscriber suspended / black-holed — before a take/reattach\n*212|/// or the `KIND_SESSIONS` reap stall-evicts it (LIFECYCLE-TRUTH W2,\n 213|/// REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE). Distinct from and 3× longer than\n 214|/// [`CONTROLLER_WRITE_DEADLINE`] on PURPOSE (doyle ruling Q1): that 5 s bound\n 215|/// protects the fan-out against ONE per-client conn falling behind and is\n...\n 429|/// `spawned_ms_ago` comes from the session's `spawned_at`; callers on the\n*430|/// KIND_SESSIONS reply use the additive `SessionInfo` fields, so the CLI\n 431|/// preflight and the broker gate consult the SAME authority.\n 432|/// PURE resolution of the two liveness probes for [`session_is_zombie`], so the\n 433|/// NO-SNAPSHOT arm is unit-testable WITHOUT a platform that lacks a process\n...\n 1002|    /// **survives `detach_if`** — a dead brain leaves its resume point behind, and\n*1003|    /// a cold-starting brain reads it (via `KIND_SESSIONS` → `resume_seq`) instead\n 1004|    /// of replaying the whole ring from 0. **Viewer writes NEVER move it**\n 1005|    /// (REQ-RCVIEW-1: the brain-resume cursor tracks the controller only).\n 1006|    ///\n...\n 1106|/// no-session truth (the exit-waiter reap) BUMPS, and a poll-derived stamp\n*1107|/// write VALIDATES against — so a `KIND_SESSIONS` snapshot taken PRE-reap can\n 1108|/// never relatch `controlled=true` after the reap cleared it (the session is\n 1109|/// gone; no later poll would ever heal the latch). The `write` mutex\n 1110|/// serializes the validate-then-write pair against the bump-then-clear pair —\n...\n 1138|/// `controlled:false` and no later edge re-stamps, so the endpoint reads uncontrolled\n*1139|/// forever while driven. This DIVERGENCE-GATED re-assert (on the `KIND_SESSIONS` poll,\n 1140|/// bounded by the reconcile cadence) writes ONLY on a real difference — no fsync storm.\n 1141|/// Best-effort: an unbound / corrupt / perch-less endpoint has nowhere to converge yet.\n 1142|///\n...\n 2697|    /// THREAD has EXITED — the record-vs-truth honesty step, run on the authoritative\n*2698|    /// `KIND_SESSIONS` poll BEFORE `controller_by`/`has_controller` are read. The\n 2699|    /// writer exits when a socket write fails (a severed / keepalive-killed controller\n 2700|    /// conn) or its `tx` drops; while it is parked on `rx.recv()` (a LIVE but idle\n 2701|    /// controller) it is NOT finished, so a live controller is never touched (the\n...\n 2733|        // writer is BLOCKED (not exited) on a suspended brain conn past\n*2734|        // BRAIN_WRITE_DEADLINE is also reaped here, so the honest `KIND_SESSIONS`\n 2735|        // truth + the off-lock converge_perch_stamps release its stale `driven_by`\n 2736|        // stamp even when nobody take/reattaches. Deferred stamp: the reap runs in\n 2737|        // the multi-session snapshot closure (no I/O under the shared sessions lock).\n...\n 2785|    /// the tally surfaces on `daemon status`). The take/reattach path and the\n*2786|    /// `KIND_SESSIONS` reap call this so a suspended brain subscriber can never hold\n 2787|    /// control on an idle session (the output-driven [`mark_controller_gone`] never\n 2788|    /// fires without new output). Returns true iff it evicted. `stamp` selects whether\n 2789|    /// the perch re-stamp happens inline (the take/reattach path, which already stamps\n...\n 3328|    ///   * controller Disconnected: the writer is already gone — the next append\n*3329|    ///     or the KIND_SESSIONS liveness reap detaches the row, unchanged.\n 3330|    ///\n 3331|    /// `last_ok` is re-armed here regardless of outcome: the suppressed window\n 3332|    /// attempted no handoff, so its width must never count toward the\n...\n 3608|        // the endpoint's stamp generation UNDER the stamp-write serial, then\n*3609|        // clear. A KIND_SESSIONS poll that snapshotted this session PRE-reap now\n 3610|        // fails its generation check and can never relatch `controlled=true`\n 3611|        // onto the perch of a session that no longer exists (the emphasys C3\n 3612|        // poll-vs-reap window: no later poll would ever heal the latch).\n...\n 3661|    /// The next seq a re-attaching controller still needs (the persisted resume\n*3662|    /// cursor surfaced in `KIND_SESSIONS` → `resume_seq`). W1: read from the\n 3663|    /// atomic the controller writer thread advances.\n 3664|    fn delivered_through(&self) -> u64 {\n 3665|        self.delivered_through.load(Ordering::Acquire)\n...\n 5704|    /// drains through the REAL logical paths (a take's `stall_evict_controller`\n*5705|    /// or a `KIND_SESSIONS` reap), exactly like a production stall-evict.\n 5706|    /// Returns false if the session doesn't exist or the seed can't be built.\n 5707|    /// `#[doc(hidden)]` — test support, not a product surface.\n 5708|    // [impl->REQ-UPDATE-PROMOTE-DRAINED]\n...\n 5870|    /// closed old conn is simply absent → `false` at once (the normal fast promote).\n*5871|    /// Follows the established sessions→log lock order (KIND_SESSIONS reap); reads\n 5872|    /// only `write_blocked_since` under each log lock, never any I/O.\n 5873|    // [impl->REQ-UPDATE-PROMOTE-DRAINED]\n 5874|    pub fn any_local_controller_wedged(&self) -> bool {\n...\n 5920|    /// restart must not disturb the hosted children).\n*5921|    fn handle_conn(self: &Arc<Self>, conn: Stream) -> io::Result<()> {\n 5922|        // RAII handler tally (panic-safe): the round-4 gate's seam for \"the\n 5923|        // poisoned conn's handler reached EOF cleanup\" (see conn_handler_count).\n 5924|        struct HandlerTally<'a>(&'a AtomicUsize);\n...\n 6132|                }\n*6133|                KIND_SESSIONS => {\n 6134|                    // Clone the row facts out from under the sessions lock, then\n 6135|                    // read each log's cursor under the log lock alone — never\n 6136|                    // both locks at once (the dispatch_subscribe discipline).\n...\n 6246|                    let frame = Envelope::new(\n*6247|                        KIND_SESSIONS_REPLY,\n 6248|                        serde_json::to_value(reply).expect(\"SessionsReply serializes\"),\n 6249|                    );\n 6250|                    send_frame(&send, &frame);\n...\n 8355|            // Reap the dead session from the table FIRST (ADR-0041 decision 4\n*8356|            // ordering): once removed, no NEW KIND_SESSIONS snapshot can see\n 8357|            // this session at all. Clearing before removal left a window where\n 8358|            // a poll snapshotted AFTER the generation bump — same generation,\n 8359|            // controller slot still occupied in the log — and legitimately\n...\n 10475|    #[test]\n*10476|    fn sessions_reply_resume_seq_defaults_when_absent() {\n 10477|        // Old-broker shape: no `resume_seq` key on the session object.\n 10478|        let old: SessionsReply =\n 10479|            serde_json::from_str(r#\"{\"sessions\":[{\"session_id\":7,\"endpoint\":\"e\"}]}\"#)\n...\n 11215|        // Resume from the cursor-of-record, exactly as a cold-starting brain does\n*11216|        // (KIND_SESSIONS resume_seq = delivered_through).\n 11217|        let (send2, mut client2, _recv2) = controller_socket_pair();\n 11218|        let resume_from = log.delivered_through.load(Ordering::Acquire);\n 11219|        log.become_controller(Arc::clone(&send2), None, resume_from, 0);\n## transport.rs\n 80|// [impl->REQ-HAZARD-HANDOFF-ARGV-COMPAT]\n*81|pub fn recv_hello<R: Read>(r: &mut R, expected_peer: Role) -> io::Result<Hello> {\n 82|    let env = read_frame(r)?;\n 83|    if env.kind != \"hello\" {\n 84|        return Err(io::Error::new("}],"details":{"scopePath":".worktrees/286-churn/crates/spt-daemon/src/brain.rs, .worktrees/286-churn/crates/spt-daemon/src/broker.rs, .worktrees/286-churn/crates/spt-daemon/src/transport.rs","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-daemon\\src","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":28,"fileCount":3,"files":[".worktrees/286-churn/crates/spt-daemon/src/brain.rs",".worktrees/286-churn/crates/spt-daemon/src/broker.rs",".worktrees/286-churn/crates/spt-daemon/src/transport.rs"],"fileMatches":[{"path":".worktrees/286-churn/crates/spt-daemon/src/brain.rs","count":7},{"path":".worktrees/286-churn/crates/spt-daemon/src/broker.rs","count":20},{"path":".worktrees/286-churn/crates/spt-daemon/src/transport.rs","count":1}],"truncated":true,"perFileLimitReached":20,"displayContent":"# .worktrees/286-churn/crates/spt-daemon/src/\n## brain.rs\n   57│    KIND_MET_MEMBER, KIND_NET_STREAM_SEND, KIND_NET_STREAM_SUBSCRIBE, KIND_OUTPUT, KIND_PAIR_CODE_SUBMIT, KIND_PAIR_JOIN, KIND_PAIR_JOINED, KIND_PAIR_MEET,\n  *58│    KIND_RESIZE, KIND_SESSIONS, KIND_SESSIONS_REPLY, KIND_SIZE, KIND_SPAWN, KIND_SPAWNED,\n   59│    KIND_SPAWN_CONFLICT, KIND_SPAWN_FRESH,\n   60│    KIND_SUBSCRIBE, KIND_SUBSCRIBED, KIND_UNSUBSCRIBE, UnsubscribeReq,\n   61│    BringUpReq, BroughtUpReply, KIND_BRING_UP, KIND_BROUGHT_UP,\n     │...\n  446│    /// except the peer pump, which uses [`Brain::cold_start_pump`].\n *447│    pub fn cold_start(name: &str, now_ms: u64) -> io::Result<Self> {\n  448│        let mut conn = LocalSocketTransport::connect(name)?;\n  449│        send_hello(&mut conn, Role::Brain)?;\n  450│        Ok(Brain {\n     │...\n  479│    // [impl->REQ-RC-DISPLAY-SOLE-WRITER]\n *480│    pub fn cold_start_pump(\n  481│        name: &str,\n  482│        now_ms: u64,\n  483│        io_timeout: Duration,\n     │...\n 1268│                }\n*1269│                KIND_SESSIONS_REPLY => {\n 1270│                    let ev: SessionsReply = serde_json::from_value(env.payload)\n 1271│                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n 1272│                    return Ok(BrokerEvent::Sessions(ev));\n     │...\n 1632│    /// `delivered_through`-blind, broker.rs). Drives the serve_attach controller\n*1633│    /// gap-resume — no mid-stream `KIND_SESSIONS` round-trip (which would re-fatal\n 1634│    /// on the same gap and discard Output).\n 1635│    // [impl->REQ-HAZARD-CONTROLLER-GAP-RESUME]\n 1636│    pub fn controller_resume_floor(&self) -> u64 {\n     │...\n 2237│    /// dispatcher's session→endpoint resolution (D9-1).\n*2238│    pub fn sessions(&mut self) -> io::Result<SessionsReply> {\n*2239│        self.send(KIND_SESSIONS, serde_json::Value::Null)?;\n 2240│        loop {\n 2241│            match self.read_event()? {\n 2242│                BrokerEvent::Sessions(ev) => return Ok(ev),\n## broker.rs\n    71│    KIND_NET_STREAM_UNSUBSCRIBE, KIND_NET_STREAM_UNSUBSCRIBED,\n   *72│    KIND_MET_MEMBER, KIND_PAIR_CODE_SUBMIT, KIND_PAIR_JOIN, KIND_PAIR_JOINED, KIND_PAIR_MEET, KIND_RESIZE, KIND_SESSIONS, KIND_SESSIONS_REPLY, KIND_SPAWN,\n    73│    BringUpReq, BroughtUpReply, BRING_UP_ADMITTED, BRING_UP_ADMIT_PREFIX, BRING_UP_ALREADY_LIVE, BRING_UP_REFUSED, KIND_BRING_UP, KIND_BROUGHT_UP,\n    74│    SealCeremonyReq, SealCeremonyReply, SealCeremonyOpenEvent, SealCeremonyCodeReq, SealCeremonyResultEvent,\n    75│    KIND_SEAL_CEREMONY, KIND_SEAL_CEREMONY_REPLY, KIND_SEAL_CEREMONY_OPEN, KIND_SEAL_CEREMONY_CODE, KIND_SEAL_CEREMONY_RESULT,\n      │...\n   211│/// write — the brain subscriber suspended / black-holed — before a take/reattach\n  *212│/// or the `KIND_SESSIONS` reap stall-evicts it (LIFECYCLE-TRUTH W2,\n   213│/// REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE). Distinct from and 3× longer than\n   214│/// [`CONTROLLER_WRITE_DEADLINE`] on PURPOSE (doyle ruling Q1): that 5 s bound\n   215│/// protects the fan-out against ONE per-client conn falling behind and is\n      │...\n   429│/// `spawned_ms_ago` comes from the session's `spawned_at`; callers on the\n  *430│/// KIND_SESSIONS reply use the additive `SessionInfo` fields, so the CLI\n   431│/// preflight and the broker gate consult the SAME authority.\n   432│/// PURE resolution of the two liveness probes for [`session_is_zombie`], so the\n   433│/// NO-SNAPSHOT arm is unit-testable WITHOUT a platform that lacks a process\n      │...\n  1002│    /// **survives `detach_if`** — a dead brain leaves its resume point behind, and\n *1003│    /// a cold-starting brain reads it (via `KIND_SESSIONS` → `resume_seq`) instead\n  1004│    /// of replaying the whole ring from 0. **Viewer writes NEVER move it**\n  1005│    /// (REQ-RCVIEW-1: the brain-resume cursor tracks the controller only).\n  1006│    ///\n      │...\n  1106│/// no-session truth (the exit-waiter reap) BUMPS, and a poll-derived stamp\n *1107│/// write VALIDATES against — so a `KIND_SESSIONS` snapshot taken PRE-reap can\n  1108│/// never relatch `controlled=true` after the reap cleared it (the session is\n  1109│/// gone; no later poll would ever heal the latch). The `write` mutex\n  1110│/// serializes the validate-then-write pair against the bump-then-clear pair —\n      │...\n  1138│/// `controlled:false` and no later edge re-stamps, so the endpoint reads uncontrolled\n *1139│/// forever while driven. This DIVERGENCE-GATED re-assert (on the `KIND_SESSIONS` poll,\n  1140│/// bounded by the reconcile cadence) writes ONLY on a real difference — no fsync storm.\n  1141│/// Best-effort: an unbound / corrupt / perch-less endpoint has nowhere to converge yet.\n  1142│///\n      │...\n  2697│    /// THREAD has EXITED — the record-vs-truth honesty step, run on the authoritative\n *2698│    /// `KIND_SESSIONS` poll BEFORE `controller_by`/`has_controller` are read. The\n  2699│    /// writer exits when a socket write fails (a severed / keepalive-killed controller\n  2700│    /// conn) or its `tx` drops; while it is parked on `rx.recv()` (a LIVE but idle\n  2701│    /// controller) it is NOT finished, so a live controller is never touched (the\n      │...\n  2733│        // writer is BLOCKED (not exited) on a suspended brain conn past\n *2734│        // BRAIN_WRITE_DEADLINE is also reaped here, so the honest `KIND_SESSIONS`\n  2735│        // truth + the off-lock converge_perch_stamps release its stale `driven_by`\n  2736│        // stamp even when nobody take/reattaches. Deferred stamp: the reap runs in\n  2737│        // the multi-session snapshot closure (no I/O under the shared sessions lock).\n      │...\n  2785│    /// the tally surfaces on `daemon status`). The take/reattach path and the\n *2786│    /// `KIND_SESSIONS` reap call this so a suspended brain subscriber can never hold\n  2787│    /// control on an idle session (the output-driven [`mark_controller_gone`] never\n  2788│    /// fires without new output). Returns true iff it evicted. `stamp` selects whether\n  2789│    /// the perch re-stamp happens inline (the take/reattach path, which already stamps\n      │...\n  3328│    ///   * controller Disconnected: the writer is already gone — the next append\n *3329│    ///     or the KIND_SESSIONS liveness reap detaches the row, unchanged.\n  3330│    ///\n  3331│    /// `last_ok` is re-armed here regardless of outcome: the suppressed window\n  3332│    /// attempted no handoff, so its width must never count toward the\n      │...\n  3608│        // the endpoint's stamp generation UNDER the stamp-write serial, then\n *3609│        // clear. A KIND_SESSIONS poll that snapshotted this session PRE-reap now\n  3610│        // fails its generation check and can never relatch `controlled=true`\n  3611│        // onto the perch of a session that no longer exists (the emphasys C3\n  3612│        // poll-vs-reap window: no later poll would ever heal the latch).\n      │...\n  3661│    /// The next seq a re-attaching controller still needs (the persisted resume\n *3662│    /// cursor surfaced in `KIND_SESSIONS` → `resume_seq`). W1: read from the\n  3663│    /// atomic the controller writer thread advances.\n  3664│    fn delivered_through(&self) -> u64 {\n  3665│        self.delivered_through.load(Ordering::Acquire)\n      │...\n  5704│    /// drains through the REAL logical paths (a take's `stall_evict_controller`\n *5705│    /// or a `KIND_SESSIONS` reap), exactly like a production stall-evict.\n  5706│    /// Returns false if the session doesn't exist or the seed can't be built.\n  5707│    /// `#[doc(hidden)]` — test support, not a product surface.\n  5708│    // [impl->REQ-UPDATE-PROMOTE-DRAINED]\n      │...\n  5870│    /// closed old conn is simply absent → `false` at once (the normal fast promote).\n *5871│    /// Follows the established sessions→log lock order (KIND_SESSIONS reap); reads\n  5872│    /// only `write_blocked_since` under each log lock, never any I/O.\n  5873│    // [impl->REQ-UPDATE-PROMOTE-DRAINED]\n  5874│    pub fn any_local_controller_wedged(&self) -> bool {\n      │...\n  5920│    /// restart must not disturb the hosted children).\n *5921│    fn handle_conn(self: &Arc<Self>, conn: Stream) -> io::Result<()> {\n  5922│        // RAII handler tally (panic-safe): the round-4 gate's seam for \"the\n  5923│        // poisoned conn's handler reached EOF cleanup\" (see conn_handler_count).\n  5924│        struct HandlerTally<'a>(&'a AtomicUsize);\n      │...\n  6132│                }\n *6133│                KIND_SESSIONS => {\n  6134│                    // Clone the row facts out from under the sessions lock, then\n  6135│                    // read each log's cursor under the log lock alone — never\n  6136│                    // both locks at once (the dispatch_subscribe discipline).\n      │...\n  6246│                    let frame = Envelope::new(\n *6247│                        KIND_SESSIONS_REPLY,\n  6248│                        serde_json::to_value(reply).expect(\"SessionsReply serializes\"),\n  6249│                    );\n  6250│                    send_frame(&send, &frame);\n      │...\n  8355│            // Reap the dead session from the table FIRST (ADR-0041 decision 4\n *8356│            // ordering): once removed, no NEW KIND_SESSIONS snapshot can see\n  8357│            // this session at all. Clearing before removal left a window where\n  8358│            // a poll snapshotted AFTER the generation bump — same generation,\n  8359│            // controller slot still occupied in the log — and legitimately\n      │...\n 10475│    #[test]\n*10476│    fn sessions_reply_resume_seq_defaults_when_absent() {\n 10477│        // Old-broker shape: no `resume_seq` key on the session object.\n 10478│        let old: SessionsReply =\n 10479│            serde_json::from_str(r#\"{\"sessions\":[{\"session_id\":7,\"endpoint\":\"e\"}]}\"#)\n      │...\n 11215│        // Resume from the cursor-of-record, exactly as a cold-starting brain does\n*11216│        // (KIND_SESSIONS resume_seq = delivered_through).\n 11217│        let (send2, mut client2, _recv2) = controller_socket_pair();\n 11218│        let resume_from = log.delivered_through.load(Ordering::Acquire);\n 11219│        log.become_controller(Arc::clone(&send2), None, resume_from, 0);\n## transport.rs\n 80│// [impl->REQ-HAZARD-HANDOFF-ARGV-COMPAT]\n*81│pub fn recv_hello<R: Read>(r: &mut R, expected_peer: Role) -> io::Result<Hello> {\n 82│    let env = read_frame(r)?;\n 83│    if env.kind != \"hello\" {\n 84│        return Err(io::Error::new("},"isError":false,"timestamp":1788999539416}}
{"type":"message","id":"4624f81b","parentId":"b17cf158","timestamp":"2026-09-10T00:18:59.434Z","message":{"role":"toolResult","toolCallId":"call_xxrNw74MW7mfjeRdzFeuDgdT|fc_0dde72effda0ad21016aa1f773477887d2b2640c2a3f3def17","toolName":"grep","content":[{"type":"text","text":"# .worktrees/286-churn/\n## traceable-reqs.toml\n 693|[[requirements]]\n*694|id = \"REQ-HAZARD-HOSTED-LIVENESS-RECONCILE\"\n*695|title = \"B2 KEYSTONE: a daemon-hosted (spt-hosted) endpoint's info.json status is RECONCILED to real liveness, not left latched online. The broker exit-waiter (broker.rs:889-910) reaps its in-mem session table + emits ExitEvent but NEVER touches info.json; lifecycle::mark_offline only fires on Psyche teardown — so a dead/exited harness (operator closed the tab) stays status=online forever (is_perch_alive returns ONLINE for daemon-hosted, liveness.rs:80-93). FIX (doyle ruled PULL-PRIMARY — the live-st...\n 696|required_stages = [\"impl\", \"unit\", \"int\"]  # activated v0.12.0 B2 keystone: impl = reconcile_hosted_liveness (controllable-gated no-broker-session → mark_offline) + query_live_session_endpoints + boot-grace gate wired into spawn_live_host; unit = pull_liveness_marks_sessionless_spt_hosted_offline_only (offlines only sessionless spt-hosted; relay/legacy/ready exempt); int = pull_reconcile_offlines_perch_when_broker_session_dies (real broker session killed+reaped → next reconcile clears the latch)\n 697|[[requirements]]\n 698|id = \"REQ-HAZARD-RC-ATTACH-FAILFAST\"\n...\n 714|id = \"REQ-HAZARD-DAEMON-STOP-REAP\"\n*715|title = \"Breap: `spt daemon stop` REAPS the spt-hosted children it spawned — no orphaned psyche/harness processes. Today a stop leaves ~8 orphaned claude-spt-psyche.exe + spt.exe: Psyches are spawned DETACHED (runtime.rs:342-356, the Child is dropped — 'Detached' ~349) and the livehost stop flag Arc<AtomicBool> is NEVER raised (brainproc.rs:227-230 holds it 'for symmetry'). FIX: on stop, raise the livehost stop flag AND kill the spawned psyche/spt-hosted children — via a Windows job object / Unix p...\n 716|required_stages = [\"impl\", \"unit\"]  # activated v0.12.0 wave 3 (with B3): impl = BrainReaper (Windows kill-on-job-close Job / Unix per-brain process-group) enrolled per brain (re)spawn + reaped on the graceful daemon-stop path (supervisor stop flag raised first, no respawn race); unit = reap.rs job/group reaps an enrolled child\n 717|[[requirements]]\n*718|id = \"REQ-HAZARD-LIVEHOST-BOOT-LIVENESS-GATE\"\n*719|title = \"B5: `spt daemon start` does NOT revive phantom Psyches for dead-but-online-latched perches. Today reconcile_once (livehost.rs:285) spawns a Psyche per status=online live_agent perch at boot WITHOUT verifying the harness child / {id}-psyche is actually alive — so a Cold start after an unclean stop revives N psyches for N dead-but-latched perches (3 psyches for 3 dead perches). FIX: gate the boot psyche-spawn on real child-liveness — a perch with NO live broker session (the B2 reconcile signal...\n 720|required_stages = [\"impl\", \"unit\", \"int\"]  # activated v0.12.0 wave 4: impl = lift the LIVENESS_RECONCILE_BOOT_GRACE skip so the B2 reconcile_hosted_liveness runs from boot tick 1 (a sessionless controllable perch is offlined BEFORE reconcile_once can revive its Psyche — reuse the B2 fn, no second liveness notion); unit = boot_gate_offlines_sessionless_controllable_then_reconcile_skips_host (+ session-backed contrast hosts); int = cold-start real broker with a stale online controllable perch (no sessio...\n 721|[[requirements]]\n 722|id = \"REQ-HAZARD-BRAIN-RESTART-LIFECYCLE-REHYDRATE\"\n*723|title = \"B4 (deepest): a bare brain restart (broker survives) REHYDRATES the live-agent lifecycle so post-restart endpoints are hosted + attachable. Today resume_sessions (brainproc.rs:186, brain.rs:797-809) re-subscribes to the broker's PTY sessions but ALL BrainLifecycle instances (lifecycle.rs:58-130; the ephemeral brain.rs:254-275) are LOST on restart → a post-restart live endpoint gets no livehost → its Psyche is never (re)hosted and new spawns die / can't attach until a FULL daemon reset (opera...\n*724|required_stages = []  # CLOSED v0.12.0 wave 5 — SUBSUMED, not built (doyle+operator ruling). Per-axis evidence: every axis the REQ names is ALREADY rebuilt on a bare brain restart — (1) BrainLifecycle config → host_one with_config_in (from the adapter registry + perch, each host); (2) pulse-driver thread + Psyche host → run_brain→spawn_live_host(brainproc:230)→reconcile_once (every brain start); (3) PTY message-delivery cursors → resume_sessions(brain.rs:797); (4) online/offline honesty →...\n 725|[[requirements]]\n 726|id = \"REQ-HAZARD-BRAIN-RESTART-PSYCHE-DUP\"\n 727|title = \"A bare brain restart leaves EXACTLY ONE `{id}-psyche` process per endpoint — no duplicate. On an abrupt brain death stop_host never runs (the LiveSet + owned child handles die with the brain) and Breap's job/group only reaps at DAEMON stop, so the PRIOR brain's Psyche stays ALIVE; the respawned brain's reconcile re-hosts a SECOND Psyche and overwrites the `{id}-psyche` perch pid, leaving the old one untracked + alive = a duplicate that lingers until daemon-stop (the operator's 'brain kill+rest...\n...\n 730|id = \"REQ-HAZARD-UNHOST-PSYCHE-REAP\"\n*731|title = \"On un-host, the detached `{id}-psyche` HARNESS PROCESS is reaped — not just its in-brain pulse-driver thread. Today stop_host (livehost.rs:203) trips the HostedLife stop flag + JOINS the driver thread, but the Psyche is a detached harness process (spawn_psyche → ManifestRuntime detached spawn, runtime.rs:341-356; its pid is untracked in HostedLife though stamped on the `{id}-psyche` perch, where residency-confirm already reads it). So endpoint-stop / mid-life agent-death / a B2/B5 offline-th...\n 732|required_stages = []  # SUPERSEDED 2026-07-04 (F-030 W3, todlando) by REQ-PSYCHE-EPHEMERAL-DRIVER (no resident = no orphan to reap). host_one spawns no resident psyche, so un-host has nothing detached to reap — HostedLife.psyche_child + the stop_host handle-reap + stop_host_reaps_the_detached_psyche_process are DELETED. Kept as a stable pointer (no dangling required stage on deleted evidence).\n 733|[[requirements]]\n 734|id = \"REQ-ENDPOINT-PURGE\"\n...\n 739|title = \"An offline ReadyAgent shows in `spt endpoint run`'s picker Resume-from-history and resumes correctly — closing the gap that today only LiveAgents do. ROOT: a harness-hosted ready bind (ReadyAgent::start_homed, ready.rs) writes info.json DIRECTLY and never appends the session ledger (unlike the shared establish_perch:250 live path), so a ready agent — though it has a session_id — produces ZERO ledger rows → the picker's offline+local Resume-from-history (which gates on ledger rows) never ...\n*740|required_stages = [\"doc\", \"impl\", \"unit\", \"int\"]  # operator add v0.12.0; built. FIX(2)=SUBSUMED (verified-at-code, ZERO code): cmd_endpoint_run is type-agnostic, the no-psyche discriminator is livehost reconcile's start-side STATE gate (info.state != live_agent → skip a ready_agent perch) ahead of the psyche_init gate. doc = CONTEXT.md `spt endpoint run` is the bringup for BOTH types. impl = ReadyAgent::start_homed ledgers a Boot row (sessions::append, mirror establish_perch:250). unit = start_ledgers...\n 741|[[requirements]]\n 742|id = \"REQ-PICKER-ADAPTER-DESCRIPTION\"\n 743|title = \"The Create-new adapter-CHOICE screen of `spt endpoint run`'s picker shows a right-hand Description panel (like the Pick-existing endpoint picker's two-pane) surfacing per-adapter detail: install date, last-updated, adapter TYPE / the endpoint types it hosts, and the adapter description — so the user can see WHAT each adapter is before choosing it (today the selector lists bare names). DEFERRED fast-follow to v0.12.0 (operator 2026-06-18). (post-v0.12.0)\"\n...\n 865|title = \"Prefer stable parent PID / broker handle over ephemeral PID (2.1)\"\n*866|required_stages = []  # rule 5: superseded in spirit by daemon-authoritative liveness (REQ-HAZARD-DAEMON-HOSTED-LIVENESS, int-covered M3b); binds if a per-pid path ever returns (none in the daemon model)\n 867|[[requirements]]\n 868|id = \"REQ-HAZARD-STDIN-SESSION-ID\"\n 869|title = \"Stdin session_id precedence over env (2.2)\"\n...\n 1118|id = \"REQ-HAZARD-PUMP-IPC-DEADLINE\"\n*1119|title = \"The single-threaded peer pump's brain-IPC reads are deadline-bounded (PUMP_PEER_IO_TIMEOUT, total-wait per call); a TimedOut read POISONS the client and escalates to a SUPERVISED RESTART, never a per-peer retry — a black-holed peer must never wedge the whole pump\"\n 1120|required_stages = [\"doc\", \"impl\", \"unit\"]  # activated 2026-06-11: deployed v0.4.0 pump wedged 2.2h on hfenduleam — net_open_stream's unbounded `loop { read_event }` blocked on a peer conn the broker never replied for; the single-threaded pump froze and supervise_pump cannot rescue a BLOCKED thread (it only catches panic/error/return). Surfaced twice (2026-06-07 + this 2.2h wedge); the stall-warning was the band-aid, not the fix. fix = Brain pump-mode carrier SPLIT at construction (Brain::cold_start_pu...\n 1121|# --- Added 2026-06-16 (v0.8.3): the pump-IPC-deadline B-half — the broker-side complement of REQ-HAZARD-PUMP-IPC-DEADLINE (DEFERRED.md:42 resolved); rule 3 register-before-satisfy, rule 5 activate per wave ---\n 1122|[[requirements]]\n 1123|id = \"REQ-HAZARD-BROKER-QUIC-DEADLINE\"\n*1124|title = \"The broker bounds every brain-waiting QUIC op (dial / open_stream / send_stream) so a black-holed or dead peer fails PROMPTLY with an ORDINARY error the broker REPLIES, never an unbounded await. The bound (< the brain's 30s PUMP_PEER_IO_TIMEOUT so the BROKER fires first) surfaces to the pump as a normal broker error reply → peer_outcome's non-TimedOut arm → drop conn + redial next tick, the round CONTINUES and the heartbeat keeps advancing — it must NEVER manifest as the brain's own read-d...\n 1125|required_stages = [\"doc\", \"impl\", \"unit\", \"int\"]  # registry-first mint (rule 3); activated per wave in v0.8.3. doc = KNOWN-HAZARDS 7.8 (+ the 7.6 B-half-fixed reword + DEFERRED.md:42 RESOLVED). impl = NetHost::bounded_block_on wrapping dial/open_stream/send_stream's QUIC await sites under a quic_op_timeout (default 10s = BROKER_QUIC_OP_TIMEOUT_MS, test-overridable via set_quic_op_timeout); unit = the bound wrapper returns a prompt non-TimedOut error on a never-completing op + a fast op is untouched [net...\n 1126|# --- Added 2026-06-17 (v0.9.1): the daemon-state-wire/broker-restart skew (perri PREP-4 FINDING 1); rule 3 register-before-satisfy ---\n 1127|[[requirements]]\n...\n 1155|[[requirements]]\n*1156|id = \"REQ-HAZARD-DAEMON-HOSTED-LIVENESS\"\n 1157|title = \"Daemon-hosted perches (Psyche, spt-hosted Self) derive liveness from the daemon endpoint table + info.json status, never is_process_alive(info.pid) (2.5)\"\n 1158|required_stages = [\"impl\", \"unit\", \"int\"]  # int activated M3b-B9: the daemon E2E spawns the hosted Psyche, then proves a dead summarizer pid does NOT flip the daemon-online perch offline\n 1159|# --- Added 2026-06-09 from grill-with-docs (ADR-0018: broker/brain in-process-collapse regression); rule 3: register before satisfying, rule 5: inactive until the broker/brain split restoration milestone (next milestone) ---\n...\n 1229|id = \"REQ-HAZARD-REGISTRY-GHOST-ROWS\"\n*1230|title = \"Registry rows must decay (no immortal rows) via TWO triggers: (a) NODE-SILENCE — evict rows whose author node has not been heard (admitted inbound feed) within the eviction window, so a vanished node's rows stop poisoning bare-id resolution with phantom AcrossNodes ambiguity; AND (b) per-row OFFLINE-TTL — evict rows that have been non-routable (Offline) beyond the per-row grace even while the author node is alive, because purge/erase leaves an immortal Offline row otherwise (ghost-heal re-ad...\n 1231|required_stages = [\"doc\", \"impl\", \"unit\"]  # (a) NODE-SILENCE activated 2026-06-06 (M7 closeout fix 2): SubnetRegistry::evict_nodes + RegistryHost::evict_silent_peers (heard-map TTL, snapshots rewritten) from the registry pump tick; registry_evict_after_ms knob (default 300s = 10 default cadences). unit = model eviction + field-bug replay (ghost row → bare resolve Ambiguous; eviction restores) + self-rows-never-decay + recently-heard survives. (b) OFFLINE-TTL activated counter-38 #2-secondary (todlando...\n 1232|\n 1233|# --- Added 2026-06-07 at M8 ratification (M8-PLAN.md decisions 1-24; rule 3:\n...\n 1318|[[requirements]]\n*1319|id = \"REQ-HAZARD-LIVEHOST-BOOT-RACE\"\n 1320|title = \"The brain's daemon-hosted Psyche lifecycle surfaces a host-FAILURE on the live perch (harness-diagnosable) and runs net-INDEPENDENTLY. When reconcile_once→host_one→spawn_psyche fails for a state=live_agent+status=online endpoint (e.g. the adapter's psyche binary absent from its install dir, REQ-INSTALL-11), the failure MUST be written to the perch info.json as a CURRENT-STATE field (reason + ts + attempt count; overwritten each 5s retry, CLEARED on successful host) and surfaced by `spt endpo...\n*1321|required_stages = [\"impl\", \"unit\"]  # activated v0.8.1 2026-06-16 (todlando): impl = info.json psyche_host_error field + set_psyche_host_error helper + SELF-pin render annotation; unit = info.rs round-trip+set/clear/increment + render_self_pin annotation. F-030 W3 (todlando 2026-07-04): the `int` leg DROPPED (only the psyche leg — NOT fully superseded). Both real-daemon ints (livehost_bootrace_e2e = resident psyche host-Ok positive, livehost_psyche_fail_e2e = resident psyche-spawn-fail negative) tested...\n 1322|# --- Added 2026-06-16 (v0.8.2): perri's v0.8.1 dogfood surfaced two follow-on spt-core defects (F-009 command-templating argv-split, F-010 residual host-failure masking); rule 3 register-before-satisfy, rule 5 activate per wave ---\n 1323|[[requirements]]\n 1324|id = \"REQ-HAZARD-TEMPLATE-ARGV-FILL\"\n 1325|title = \"Command-template substitution fills argv ELEMENTS, not a re-tokenized string: spt-core currently `fill_template`s {key} values INTO the command STRING and THEN `tokenize`s the filled string (runtime.rs:94/122), so a multi-word {key} value whitespace-SPLITS into multiple argv tokens unless the adapter hand-quotes the placeholder, and a value containing a `\\\"` (or `;`) injects/breaks tokenization (shell-injection-adjacent). A filled value MUST become exactly ONE argv element regardless of spaces/q...\n*1326|required_stages = [\"impl\", \"unit\", \"int\"]  # activated v0.8.2 2026-06-16 (todlando-w162): impl = fill_template_tokens (tokenize-template-then-fill-each) + rewired command_for/run_bounded_command_in/harnesshost/shellwake/shellhost callers; unit = multi-word→1-element, quote/semicolon→1-element, embedded-placeholder, missing-key/empty-command, literal-brace parity. F-030 W3 (todlando 2026-07-04): the int leg was RE-HOMED (doyle ruling), NOT dropped — its prior carrier livehost_nonresident_e2e retired...\n 1327|[[requirements]]\n*1328|id = \"REQ-HAZARD-LIVEHOST-NONRESIDENT\"\n*1329|title = \"A daemon-hosted Psyche that spawns then EXITS IMMEDIATELY is a host failure, surfaced like a spawn failure (closes the v0.8.1 residual masking): the REQ-HAZARD-LIVEHOST-BOOT-RACE signal stamps `psyche_host_error` only when `spawn_psyche` returns Err, NOT when the detached spawn() returns Ok but the child dies within moments (e.g. a bad-argv child exiting 2 — the F-009 case). That leaves the residual 'online + no Psyche + no cause' gap: the nested `{id}-psyche` info.json is written status=onlin...\n*1330|required_stages = []  # SUPERSEDED 2026-07-04 (F-030 W3, todlando) by REQ-HAZARD-PSYCHE-RESIDENCY-EXPECTATION (+ the REQ-PSYCHE-EPHEMERAL-DRIVER turn budget). The resident model retired: host_one spawns no resident psyche, so \"spawned-then-exited-fast\" cannot arise — classify_residency/confirm_residency_or_unhost + the residency e2e (livehost_nonresident_e2e) are DELETED. The spirit (a psyche that cannot work is visible + hosting must not churn) transfers to the W1 note_turn_outcome consecutive-N budge...\n 1331|[[requirements]]\n 1332|id = \"REQ-HAZARD-EPOCH-RESET\"\n 1333|title = \"Advertisement-epoch reset strands a node: peers' higher last-seen epoch drops the reset node's fresh advertisements as Stale until the counter outruns history. Common case (full reinstall/re-pair) is mitigated by REQ-SUBNET-7's ceremony eviction (peer-side epoch memory dies with the deleted row — acceptance-verified); the residual narrow slice (epoch file lost, identity kept) is documented, guard deferred to a field hit (4.11)\"\n\n## docs/\n### KNOWN-HAZARDS.md\n 196|- **Failure:** the registry's only superseding mechanism is the per-`(endpoint_id, node)` epoch lease (4.8) — a row is replaced only by a newer row *from the same node*. When a node identity dies permanently (machine retired, or `node.key` regenerated so the \"node\" never speaks again), its rows are never superseded and never expire: they sit in the in-memory registries and the `identity/registry/<subnet>.json` snapshots forever. A bare-id send then resolves the same endpoint id on both the live and the...\n*197|- **Invariant:** registry rows authored by a **silent** peer node decay: a node not *heard* (admitted inbound feed — the M7 D2 heard-map, REQ-SUBNET-1) within the eviction window (`registry_evict_after_ms`, default 300s ≈ 10 default pump cadences) has its rows **evicted** from every subnet registry, snapshots rewritten. Own rows never decay (the node always hears itself implicitly — it authors them each pump tick). Eviction is safe under the lease: v1 has **no transitive gossip**, so any future upd...\n 198|- **spt-core mapping:** `spt_net::net::registry::SubnetRegistry::evict_nodes` (model) + `spt_daemon::registryhost::RegistryHost::evict_silent_peers` (heard-map TTL) driven from the registry pump tick (`peerloop`). Trust rows are NOT auto-evicted (trust is a user decision; a stale trust row only costs dead dials) — pruning those is a separate verb.\n 199|- **Source:** M7 acceptance run 2026-06-06 (DEFERRED.md \"Ghost registry row eviction\"); the AMBIGUOUS render fix rode along.\n 200|- **Mesh note (ADR-0017, 2026-06-08):** the subnet mesh **preserves** this invariant rather than superseding it. \"No transitive gossip\" sharpens to **no transitive *row* gossip** — the mesh relays only the member *roster* (discovery), while registry **rows stay own-authored and are fetched directly** from each member over a handshake. So \"any future update for a node comes from that node itself, alive\" still holds and the eviction lease is untouched. (The plan's rejected alternative — signed transiti...\n...\n 500|- **Failure:** the peer pump is a SINGLE thread driving every leg (registry/notif/sync/update) against every peer over ONE brain-IPC client. Its reply reads (`net_open_stream`, `net_stream_send`, `net_dial`, and the sync/update pull `read_event` loops) were `loop { read_event() }` with no deadline. When a peer's QUIC path black-holes, the broker's stream-open/send awaits the dead peer and never sends the reply, so the brain's `read_frame` blocks FOREVER and the pump freezes mid-round. The heartbeat (loop...\n*501|- **Invariant:** in PUMP mode the brain carrier is **SPLIT at construction** — a dedicated `pump-ipc-reader` thread does blocking `read_frame` on the `RecvHalf` and forwards each framed result down a channel; the main thread writes on the `SendHalf` and reads with `Receiver::recv_timeout`. Every IPC reply read is bounded by a per-call **total-wait** deadline (`PUMP_PEER_IO_TIMEOUT` = 30s, > any legitimate round-trip, < the 60s QUIC idle; re-armed on stream progress for the streaming pull legs so a heal...\n 502|- **spt-core mapping:** `Brain::cold_start_pump` (splits the carrier + arms the deadline) / `BrainConn::Split` (the `SendHalf` + reader-thread channel) / `call_deadline` / `read_event_until` / `read_frame_until` (the `recv_timeout` dispatch), `pump::run_peer_pump` (connects in pump mode) + `pump::peer_outcome` (the tier-split), the `request_sync`/`request_update` pull loops (deadline re-armed on progress). The broker-side half — the broker must never make a brain wait unbounded on a QUIC op (bound the ...\n 503|- **Source:** field diagnosis 2026-06-11 (the 2.2h hfenduleam wedge); doyle ruling A-now / B-deferred. The stall warning (M8 decision 23) was the band-aid; this is the fix. The B-half landed v0.8.3 (§7.8) after the 2026-06-16 recurrence.\n 504|\n...\n 514|- **Failure:** the broker's brain-facing QUIC handlers (`dispatch_net_dial` / `dispatch_net_stream_open` / `dispatch_net_stream_send`) call into `NetHost::dial` / `open_stream` / `send_stream`, whose iroh awaits (`endpoint.connect` + `prove_membership`; `open_bi`; `write_all`/`finish`) had NO bound of their own. A dead/black-holed roster peer (its process gone, or a mixed-pair that accepts the conn but never answers the seed-proof) makes the broker await its QUIC path FOREVER, so the brain escapes only v...\n*515|- **Invariant:** every brain-waiting QUIC op is wrapped in a broker-side deadline (`NetHost::bounded_block_on` → `tokio::time::timeout`, `BROKER_QUIC_OP_TIMEOUT_MS` = 10s). On elapse the future is DROPPED (cancelling the in-flight connect/stream op, so nothing is half-registered) and a non-`TimedOut` `io::Error` is returned, which the broker REPLIES as an ordinary error frame. The bound (10s) sits comfortably above any legitimate LAN/relay round-trip and 20s below the brain's 30s `PUMP_PEER_IO_TIMEOUT`...\n 516|- **spt-core mapping:** `NetHost::bounded_block_on` (the timeout wrapper) wrapping `NetHost::dial` / `open_stream` (QUIC branch) / `send_stream`; `BROKER_QUIC_OP_TIMEOUT_MS` + `set_quic_op_timeout` (test override, off `NetConfig` — mirrors `set_roster_exchange`). Unit: `bounded_block_on_cuts_a_never_completing_op_with_an_ordinary_error` (a never-completing op → prompt non-`TimedOut` error; a ready op untouched). Int: `dial_to_a_black_holing_peer_fails_with_a_bounded_ordinary_error` (the broker REPLIE...\n 517|- **Source:** the 7.6 B-half — deferred 2026-06-11 (doyle ruling A-now / B-deferred, DEFERRED.md) and shipped in v0.8.3 after the 2026-06-16 hfenduleam recurrence.\n 518|\n...\n 558|### 7.15 An OFFLINE spt-hosted endpoint must NOT render phantom `ONLINE+CONTROLLED` — clear `driven_by` when its session is gone  `[REQ-HAZARD-DRIVEN-BY-SELFHEAL]`\n*559|- **Failure:** `driven_by` (the `info.json` `ONLINE+CONTROLLED` latch) is single-written by the broker via `clear_controller`/`stamp_driven_by`, which only fire on a controller change. When an spt-hosted endpoint's broker session is GONE (harness dead — the B2 case, 7-series sibling `REQ-HAZARD-HOSTED-LIVENESS-RECONCILE`), no controller event ever fires, so a stale `driven_by=Some(node)` persists: the picker renders a phantom \"controlled by X\" on an endpoint that is actually OFFLINE. The B2 reconcile a...\n 560|- **Invariant:** `reconcile_hosted_liveness`, when it offlines a sessionless controllable perch (the B2 keystone — no live broker session ⇒ dead harness), ALSO clears `driven_by` (`set_driven_by(perch, None)`). RACE-FREE and single-writer-safe: with NO live broker session there is no controller to re-stamp `driven_by` concurrently, so the brain may write it here without contending the broker. (The LIVE-session leg — a controller gone while its session survives — is NOT this hazard: a clean discon...\n 561|- **MEASURED 2026-07-22 (DAEMON-LIFECYCLE W2 Leg B, `transport_death_eof.rs`) — the residual SPLITS in two, and only one half is still open.** On REAL QUIC (two brokers, `BindScope::Loopback` + a real `net_dial` — NOT the in-process duplex, which has no idle timeout at all), a remote controller whose QUIC STACK DIES WITHOUT A FIN (peer frozen: endpoint alive, UDP port still bound, keepalives unanswered — the sleeping-laptop shape) DOES self-heal: the transport's own idle timeout tears the conn, the...\n 562|- **Watch-out (repro-proven, real broker):** `SessionInfo.controller_by==None` is AMBIGUOUS — `dispatch_spawn` pre-attaches the spawner as the LOCAL controller with `by=None`, so a live LOCALLY-driven session also reads `None`. It is therefore NOT a usable standalone `driven_by` clear trigger (would false-clear a live local session). The shipped Gap-B self-heal needs no controller signal at all (it keys on session ABSENCE).\n*563|- **spt-core mapping:** impl `livehost.rs` `reconcile_hosted_liveness` Gap-B clear + the additive `SessionInfo.controller_by` observability field (`msg.rs`, populated in `broker.rs` `KIND_SESSIONS`); unit `livehost.rs` `pull_liveness…` extended (offlined sessionless perch clears `driven_by`; live/relay perches untouched); int `crates/spt-daemon/tests/driven_by_selfheal.rs` `gap_b` (real broker: reconcile offlines AND clears `driven_by`) + the A1/A2 characterization (`inject_control_wedge.rs` w5_a1/w5_a2).\n 564|- **Cross-ref 7.29 (who MAY clear vs who MUST NOT):** the BRAIN reconcile here MUST NOT clear `driven_by` off `controller_by==None` (ambiguous with a live LOCAL controller). The BROKER stamp-convergence (7.29) MAY — it reads `has_controller()` DEFINITIVELY (its own controller slot), so it authoritatively clears a stale `driven_by` against a live session with no controller (the `has_controller()==false` leg). The invariant's owner for that leg moved to the broker; this brain-side ambiguity-safety is unc...\n 565|- **Source:** v0.13.0 W5 (repro-first, todlando; doyle-assigned 2026-06-19). A2 (idle wedged-remote leg) deferred to `REQ-HAZARD-DRIVEN-BY-IDLE-REMOTE-EVICT`.\n 566|\n...\n 596|### 7.21 Exactly ONE `controller_writer` per brain↔broker connection — a superseded writer must write nothing further  `[REQ-HAZARD-CONTROLLER-WRITER-REORDER]`\n*597|- **Failure (doyle instrumented RACEDIAG repro, kitsubito):** on a brain-restart re-serve the handoff brain registered as controller on the SAME session TWICE over the SAME socket — `Brain::handoff` eagerly `subscribe(prior.next_seq=1)` → `become_controller(from_seq=1)`, spawning writer-A (writes seq 1); then `serve_attach` re-handled the replayed `Request{from_seq:0}` → `attach_as(sid,0)` → `become_controller(from_seq=0)`, spawning writer-B (writes 0,1). `become_controller` dropped the prior `Co...\n 598|- **Invariant:** on a single brain↔broker connection exactly ONE `controller_writer` is ever the LIVE writer; a SUPERSEDED writer writes no further frames after the epoch bump it observes; and **every `controller_writer` emits a strictly ASCENDING seq stream** (sorted initial batch + ascending live frames). The CORRECTNESS guarantee that falls out: a snap-above consumer over any interleaving of ascending writers — where the surviving writer (`serve_attach`'s `attach_as(sid,0)`) offers the COMPLETE ra...\n 599|- **spt-core mapping:** impl `brain.rs` `Brain::handoff` (KEEP the eager `subscribe`; seed `session_cursors`) + `Brain::subscribe_with` (resume-mode dedup-cursor reset to `from_seq`, shared by `attach`/`attach_as` — the operator-stream boundary fix) + `broker.rs` `OutputLog.controller_epoch: Arc<AtomicU64>` / `become_controller` (atomic `fetch_add`, passes the new epoch + `Arc::clone` into the writer) / `controller_writer` (epoch gate read UNDER the lock on both loops) / `mark_controller_gone` + the `C...\n 600|- **Source:** v0.13.0 P1c (operator-ruled root-fix before ship; doyle root-cause via instrumented repro 2026-06-20; design corrected across two gate rounds — fix #1 reverted, then the kitsubito RACEDIAG pinned the residual to the consumer-side operator-stream boundary, fixed by the `subscribe_with` cursor reset on `attach_as` re-subscribe). The last v0.13.0 ship-blocker.\n...\n 684|- **Failure (paid-for, field brick 2026-07-04, adapter v0.13.2):** the pre-F-030 model kept a **resident** Psyche process the daemon supervised, with residency machinery (`confirm_residency_or_unhost`) that **un-hosted the parent endpoint** when the resident child went missing. A bad adapter ship (v0.13.2) made the psyche shim exit on every turn; the residency machinery read that as a lost resident and **tore down the parent's hosted state — the ready marker was removed and never re-stamped**, so every...\n*685|- **Invariant:** a Psyche is a **bounded per-event turn**, not a resident process — its liveness is that turns succeed, so there is no \"resident is gone\" signal to react to. A psyche turn failure of ANY shape (spawn fail, non-zero exit, timeout, wedge) stamps **psyche fields only** (`psyche_host_error` on the Self perch) and **NEVER** touches the parent endpoint's `status` / ready marker / hosted state; the endpoint stays online and deliverable while its Psyche is unwell. The teardown-on-psyche-trouble...\n*686|- **spt-core mapping:** the residency/reap DELETIONS in `spt_daemon::livehost` (`reconcile_once`'s stop-side no longer un-hosts on psyche trouble; `host_one` holds no resident child) + `spt_daemon::lifecycle::run_psyche_event_turn` (a turn failure stamps `psyche_host_error` only) + the `first_turn_psyche_context` non-empty guarantee (a zero-context fresh turn can't masquerade as continue → reseed churn). Conformance int = the hall-bf shape: multi-subnet home, live endpoint, a psyche turn that fails eve...\n 687|- **Source:** F-030 (psyche-ephemeral) — paid-for by the hall-bf churn (ordinal 6491+) and the adapter v0.13.2 bad-ship brick, 2026-07-04.\n 688|\n 689|### 7.31 The Psyche failure budget must count REAL per-event attempts — a resident rate-guard is blind to per-event churn  `[REQ-HAZARD-THRASH-GUARD-BLIND]`\n...\n 702|\n*703|### 7.33 NO bare `.lock().unwrap()` on a broker-resident lock reachable from serve/dispatch — a poison permanently wedges every attach  `[REQ-HAZARD-BROKER-FLOOR-LOCK-POISON]`\n*704|<!-- [doc->REQ-HAZARD-BROKER-FLOOR-LOCK-POISON] -->\n*705|- **Failure (paid-for class):** a brain-only self-update restarts the BRAIN but keeps the BROKER process — and every one of its `Mutex`es — ALIVE by design (REQ-UPD-3). So a single panic while another thread held a broker-resident lock POISONS it permanently: the next `.lock().unwrap()` panics too, kills its per-conn reply thread, and EVERY subsequent attach silently deadlines (`spt rc` → \"brain IPC read deadline elapsed\") while non-locked ops keep working. The effect journal (bug #16, `[REQ-HAZARD...\n 706|- **Invariant:** NO broker-resident lock reachable from serve/dispatch may use a bare `.lock().unwrap()`. Recover via the shared `recover` (`into_inner` — safe for the short, coherent-on-recovery map ops of `sessions`/`pair_holds`) or, for the `OutputLog` ring, `recover_log` — which adds a COHERENCE CLAMP on the recovery path: a panic mid-`append` can leave the ring torn (over-cap, a last seq not below `next_seq`, non-monotonic), and serving those bytes risks garbage, so `clamp_or_reset` cheap-checks...\n*707|- **spt-core mapping:** `broker::recover` / `broker::recover_log` / `OutputLog::clamp_or_reset`. unit = a poisoned session-map recovers to a usable guard (the next attach still opens), a torn ring clamps/resets-empty only when torn (coherent untouched, `next_seq` preserved), and `recover_log` wires the clamp onto the poison branch — each RED-first against a bare `.lock().unwrap()` / an un-clamped `into_inner`. int (scripted panic-under-lock during concurrent attach → next attach opens) deferred to th...\n*708|- **Source:** REMOTE-TRUTH B-1 (pivoted — the triage's named floor sites were already closed by `[REQ-HAZARD-INJECT-WORKER-POISON]`; the surviving attach-path lock class is the real root), doyle B-1 ruling (recover all three via one helper + the OutputLog ring coherence clamp).\n 709|\n 710|### 7.34 A dead `rec.pid` on an spt-hosted perch is EXPECTED — no reader may alive-gate on `rec.pid` alone  `[REQ-HAZARD-DEAD-REC-PID]`\n 711|<!-- [doc->REQ-HAZARD-DEAD-REC-PID] -->\n...\n 736|\n*737|### 7.38 EVERY write on a physical broker connection is bounded + cancelable + poison-on-failure — no writer holds the send gate across an UNBOUNDED OS write  `[REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]`\n*738|<!-- [doc->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK] -->\n*739|- **Failure (paid-for, field capture + deterministic Windows repro 2026-07-09 — hertz RCA, UPDATE-WEDGE round-4):** `controller_writer` held the `SharedSend = Arc<Mutex<SendHalf>>` guard ACROSS a blocking `write_frame`. On Windows, `interprocess 2.4.2` routes the send to `WriteFileEx` + `SleepEx(INFINITE, alertable)` — this local-socket send path exposes **no supported write timeout** (`set_timeout` → `Unsupported`). An `rc --take` controller consumer that stopped reading blocked the write **indefi...\n*740|- **Invariant:** every write on a physical broker connection rides ONE broker-owned bounded/cancelable framed-write primitive with an **independent out-of-band aborter**. The `SendHalf` never leaves the conn object (a bypass cannot compile). An ABSOLUTE deadline stamped at write entry covers BOTH the serialized gate-wait AND OS write completion; a per-conn watchdog fires at that deadline and aborts the in-flight op — never relying on the write returning or on opportunistic stall-evict sampling. On dead...\n*741|- **spt-core mapping:** `spt-daemon/src/conn.rs` `BrokerConn` (Condvar-gated serialized write gate + poisoned state + per-in-flight op id + watchdog + `abort_physical()`: cfg(windows) `CancelIoEx` → completion handshake → `DisconnectNamedPipe` raw kernel32 externs; cfg(unix) `UnixStream::shutdown(Both)`); `broker.rs` `SharedSend = Arc<BrokerConn>` + `controller_writer`/`viewer_writer`/`send_frame`/`send_error` routed; `nethost.rs` stream-log/presence sends routed. int = `brain_decouple.rs` `non_drain...\n 742|- **Source:** UPDATE-WEDGE round-4 (hertz root doc `docs/UPDATE-WEDGE-ROOT-CAUSE.md` @f8596ca; doyle gate ruling `docs/UPDATE-WEDGE-RCA-DISPATCH-todlando.md`; todlando build 2026-07-09). The residual 7.36 left open: logical eviction without physical retirement.\n 743|\n 744|---\n...\n 774|<!-- [doc->REQ-HAZARD-MESH-BOOTSTRAP-TRAP] -->\n*775|- **Failure (paid-for, hertz field RCA 2026-07-10 — HFENDULEAM + ENLYZEAM fully sequestered from every subnet member, symmetric, green-status):** the pump resolved dial addresses from the exact `peer-addrs.json` entry else id-only discovery — never the valid `RosterEntry.address` — and every `PRESENCE_DIAL_FAILED` unconditionally `drop_seed`'d the cached entry, while the cache refilled only after a future successful seed-proof connection. One transient failure converted a warm route into id-only-or...\n 776|- **Invariant:** dial-address resolution is the ROUTE CHAIN — exact cache → validated `RosterEntry.address` (id-match required) → id-only discovery — always fully consulted in order. A failed dial demotes a cached route (suspect), never deletes a sole route; removal happens only via validated-fresher replacement or roster tombstone. Validated roster addresses reconcile into the cache at startup and on roster merge — recovery never requires an already-successful connection or operator state surg...\n 777|- **spt-core mapping:** `spt-daemon/src/pump/mod.rs` `resolve_submit_addr` + `presence_state_effect` (`PRESENCE_DIAL_FAILED` arm); `spt-store/src/peeraddrs.rs` `PeerAddrStore::put`/`drop_seed`; `spt-daemon/src/seedproofx.rs` `gapfill_peeraddrs`; status renders (`daemon status` / `subnet status`).\n 778|- **Second clause (added 2026-07-31, see 7.59):** a peer is never stranded by an address that has ALREADY failed either. Preferring the roster address assumed the retained row was fresher than the failed one; the cross-site incident falsified that assumption (7.60 froze every self row), turning this entry's fix into a `fail → demote → reinstall-same → fail` livelock. The clause is ADDITIVE — everything above still holds, and a roster address DIFFERING from the failed one still resolves and still ...\n...\n 782|\n*783|### 7.43 One wedged stream subscriber must NEVER stall stream serving — replay halts at the first failed write, a poisoned subscriber is removed, and recovery machinery must not manufacture new replay victims  `[REQ-HAZARD-REDISPATCH-STALL]`\n 784|<!-- [doc->REQ-HAZARD-REDISPATCH-STALL] -->\n\n### adr/\n#### 0004-single-daemon-broker-brain-split-and-self-update.md\n 26|\n*27|**Self-update delivery:** peer-propagated over P2P, layered on self-fetch, out-of-band still supported. All binaries signature-verified before handoff (spt-core's own release key) regardless of source — peer-propagation otherwise lets one compromised node poison the subnet. spt-core conducts updates for the whole stack: self first, then ripple-update each registered adapter via the adapter manifest's update declaration (file-pull or delegated command). The plugin's role shrinks to initial bootstrap only.\n 28|\n 29|**Cadence/consent:** not fully automatic by default; gated on user confirmation delivered to the most-recently-active live session, with an opt-in full-auto choice.\n 30|\n#### 0034-per-minter-op-id-namespacing.md\n 55|  monotonically shrink; a live producer minting untagged would parse as legacy\n*56|  forever and poison that discriminant. Uniform rule, no exceptions: every\n 57|  minter declares a tag; the random base becomes belt-and-suspenders.\n 58|- **One canonical minter enum.** The numeric key discriminant and the journal\n 59|  line's minter token derive from a single enum (one source for numeric +\n#### 0038-broker-stream-lifecycle-truth.md\n 88|(`dispatch.rs:414` — intended old-broker-only, catches transport\n*89|timeout/EOF/poison too), each pass installing a throwaway peek subscriber\n 90|whose `StreamLog::attach` replays the whole retained ring UNDER the per-stream\n 91|mutex with discarded write errors — one wedged subscriber conn = serial 15s\n*92|bounded-write poison windows (33 observed, all 15,000–15,154 ms) composing\n 93|into 17–62s PTY/RC stalls.\n 94|\n 95|Amended decisions:\n...\n 98|   unsupported-verb/old-broker answer. Every other opener-query error\n*99|   (transport timeout, EOF, poison) is `Failed` → the bounded claim requeue —\n 100|   never a second replay subscriber.\n 101|2. **(new — subscriber write discipline)** Replay and live fan-out stop at\n 102|   the FIRST failed subscriber write, and the failed subscriber is REMOVED\n...\n 108|3. **(new — bounded forwarding)** Attach worker completion and forwarding are\n*109|   BOUNDED and cancelable: a poisoned brain subscriber cancels its paired\n 110|   serve/wire worker (no orphan forwarding legs). Coordinate at build with\n 111|   the DAEMON-LIFECYCLE C2 stream-lease work — same neighborhood, build once.\n 112|4. **(amends Decision 3)** Retries are CLASSIFIED: a pre-setup transient may\n*113|   retry; a deadline-poisoned replay is CIRCUIT-BROKEN (global backoff), and\n 114|   a replacement subscriber is never installed until the prior subscriber is\n 115|   fully gone. This is an upgrade of the bounded backoff, NOT a revert to the\n 116|   v0.33 burn-the-claim abandonment (explicitly rejected).\n...\n 126|   gen/stream/family/endpoint/attempt/conn with\n*127|   replay/poison/cancel/worker-outcome events and gauges.\n 128|\n 129|Rejected here: reverting retryable claims (restores permanent abandonment);\n 130|widening any write deadline (hides the composition instead of removing it);\n*131|tearing down all subscribers on any poison (punishes healthy streams for a\n 132|wedged conn).\n#### 0039-peer-route-resolution-truth.md\n 40|behavior preserves historical corruption after the roster is corrected —\n*41|persisted poison rows were live on both incident nodes (outer key of one\n 42|member mapping another member's address).\n 43|\n 44|## Decision\n...\n 54|   never route-less. (\"Validated\" = the entry's address `id` matches the peer\n*55|   key being resolved — poison rows never become routes.)\n 56|2. **Retention is nondestructive.** A failed dial NEVER deletes the only\n 57|   bootstrap route. A cached address that failed is *demoted* (marked\n 58|   suspect / superseded when a fresher source disagrees), not forgotten;\n...\n 165|  pass pending; mint at its own triage.\n*166|- **0.30.1 peer-pump UPDATE-leg reply-read poison** (fixed): same pump, a\n 167|  different leg — the reply-read path, not route resolution.\n 168|- **PUMP_PEER_FAIL unstamped** (2026-07-14 seed): folded here — the W3 stage\n 169|  split subsumes it.\n#### 0040-bounded-stream-seat-lifecycle-one-way-terminal.md\n 31|   in the brain's event loop. While the brain does this it stops draining IPC; the broker's\n*32|   seat writer blocks, hits the 15s SharedSend bound, poisons (12→15 observed\n*33|   CONN_WRITE_POISONED, family=Registry), and the dispatcher re-replays from seq 0.\n 34|4. **Completed seats retain OS threads — family-agnostically.** The broker conn loop\n 35|   accretes `my_stream_subs` per subscribe and releases only at conn-loop exit; there is NO\n 36|   net-stream unsubscribe verb. Every completed subscription on a long-lived carrier\n...\n 66|   discipline (no blocking work where it stalls a drain loop) applied to the brain side.\n*67|4. **Poisoned one-way replays are bounded.** A deadline-poisoned Registry replay carries a\n 68|   per-stream strike budget; at budget the row retires TERMINAL, loudly. Safe by family\n 69|   semantics: a feed is an idempotent snapshot advertisement — the next pump round\n 70|   re-advertises. (Request/reply families keep ADR-0038 Amendment circuit-breaker\n...\n 94|- Steady-state broker cost on long-lived conns: threads, rows, and IPC enumeration all\n*95|  plateau at O(active). The field discriminators: hertz's preserved snapshot's poison count\n 96|  freezes; brain CPU drops to noise (with the self-hash fix, REQ-BRAIN-HASH-ONCE); broker\n 97|  thread census bounded and named.\n 98|- New wire surface: `KIND_NET_STREAM_UNSUBSCRIBE` (additive, best-effort) + a lifetime\n#### 0041-endpoint-lifecycle-truth.md\n 17|   minting `ready_agent`/`controllable=false` hybrid rows born online-authoritative. The\n*18|   livehost reconcile skips `state != live_agent` rows AND gates cleanup on\n 19|   `controllable=true`, so a dead-PID hybrid survives every restart; `is_perch_alive`\n 20|   trusts `status=online` and skips dead-PID rejection — every reader faithfully\n 21|   resurrects the lie.\n#### 0045-endpoint-teardown-authority.md\n 82|(`controllable == Some(true)`, ADR-0041) — the same predicate `cmd_bind`'s online-earn and\n*83|the livehost reconcile already use, so the codebase carries one authority concept, not\n 84|two. Harness-hosted / external endpoints keep marker + address + status behavior\n 85|unchanged: core spawned nothing there and cannot reach into the harness's process tree\n 86|(CONTEXT.md:39). The misleading doc comment claiming marker-removal stops the listener is\n\n[Some lines truncated to 512 chars]"}],"details":{"scopePath":".worktrees/286-churn/docs/KNOWN-HAZARDS.md, .worktrees/286-churn/docs/adr, .worktrees/286-churn/traceable-reqs.toml","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":58,"fileCount":9,"files":[".worktrees/286-churn/docs/KNOWN-HAZARDS.md",".worktrees/286-churn/docs/adr/0004-single-daemon-broker-brain-split-and-self-update.md",".worktrees/286-churn/docs/adr/0034-per-minter-op-id-namespacing.md",".worktrees/286-churn/docs/adr/0038-broker-stream-lifecycle-truth.md",".worktrees/286-churn/docs/adr/0039-peer-route-resolution-truth.md",".worktrees/286-churn/docs/adr/0040-bounded-stream-seat-lifecycle-one-way-terminal.md",".worktrees/286-churn/docs/adr/0041-endpoint-lifecycle-truth.md",".worktrees/286-churn/docs/adr/0045-endpoint-teardown-authority.md",".worktrees/286-churn/traceable-reqs.toml"],"fileMatches":[{"path":".worktrees/286-churn/docs/KNOWN-HAZARDS.md","count":20},{"path":".worktrees/286-churn/docs/adr/0004-single-daemon-broker-brain-split-and-self-update.md","count":1},{"path":".worktrees/286-churn/docs/adr/0034-per-minter-op-id-namespacing.md","count":1},{"path":".worktrees/286-churn/docs/adr/0038-broker-stream-lifecycle-truth.md","count":7},{"path":".worktrees/286-churn/docs/adr/0039-peer-route-resolution-truth.md","count":3},{"path":".worktrees/286-churn/docs/adr/0040-bounded-stream-seat-lifecycle-one-way-terminal.md","count":4},{"path":".worktrees/286-churn/docs/adr/0041-endpoint-lifecycle-truth.md","count":1},{"path":".worktrees/286-churn/docs/adr/0045-endpoint-teardown-authority.md","count":1},{"path":".worktrees/286-churn/traceable-reqs.toml","count":20}],"truncated":true,"perFileLimitReached":20,"displayContent":"# .worktrees/286-churn/\n## traceable-reqs.toml\n  693│[[requirements]]\n *694│id = \"REQ-HAZARD-HOSTED-LIVENESS-RECONCILE\"\n *695│title = \"B2 KEYSTONE: a daemon-hosted (spt-hosted) endpoint's info.json status is RECONCILED to real liveness, not left latched online. The broker exit-waiter (broker.rs:889-910) reaps its in-mem session table + emits ExitEvent but NEVER touches info.json; lifecycle::mark_offline only fires on Psyche teardown — so a dead/exited harness (operator closed the tab) stays status=online forever (is_perch_alive returns ONLINE for daemon-hosted, liveness.rs:80-93). FIX (doyle ruled PULL-PRIMARY — the live-st...\n  696│required_stages = [\"impl\", \"unit\", \"int\"]  # activated v0.12.0 B2 keystone: impl = reconcile_hosted_liveness (controllable-gated no-broker-session → mark_offline) + query_live_session_endpoints + boot-grace gate wired into spawn_live_host; unit = pull_liveness_marks_sessionless_spt_hosted_offline_only (offlines only sessionless spt-hosted; relay/legacy/ready exempt); int = pull_reconcile_offlines_perch_when_broker_session_dies (real broker session killed+reaped → next reconcile clears the latch)\n  697│[[requirements]]\n  698│id = \"REQ-HAZARD-RC-ATTACH-FAILFAST\"\n     │...\n  714│id = \"REQ-HAZARD-DAEMON-STOP-REAP\"\n *715│title = \"Breap: `spt daemon stop` REAPS the spt-hosted children it spawned — no orphaned psyche/harness processes. Today a stop leaves ~8 orphaned claude-spt-psyche.exe + spt.exe: Psyches are spawned DETACHED (runtime.rs:342-356, the Child is dropped — 'Detached' ~349) and the livehost stop flag Arc<AtomicBool> is NEVER raised (brainproc.rs:227-230 holds it 'for symmetry'). FIX: on stop, raise the livehost stop flag AND kill the spawned psyche/spt-hosted children — via a Windows job object / Unix p...\n  716│required_stages = [\"impl\", \"unit\"]  # activated v0.12.0 wave 3 (with B3): impl = BrainReaper (Windows kill-on-job-close Job / Unix per-brain process-group) enrolled per brain (re)spawn + reaped on the graceful daemon-stop path (supervisor stop flag raised first, no respawn race); unit = reap.rs job/group reaps an enrolled child\n  717│[[requirements]]\n *718│id = \"REQ-HAZARD-LIVEHOST-BOOT-LIVENESS-GATE\"\n *719│title = \"B5: `spt daemon start` does NOT revive phantom Psyches for dead-but-online-latched perches. Today reconcile_once (livehost.rs:285) spawns a Psyche per status=online live_agent perch at boot WITHOUT verifying the harness child / {id}-psyche is actually alive — so a Cold start after an unclean stop revives N psyches for N dead-but-latched perches (3 psyches for 3 dead perches). FIX: gate the boot psyche-spawn on real child-liveness — a perch with NO live broker session (the B2 reconcile signal...\n  720│required_stages = [\"impl\", \"unit\", \"int\"]  # activated v0.12.0 wave 4: impl = lift the LIVENESS_RECONCILE_BOOT_GRACE skip so the B2 reconcile_hosted_liveness runs from boot tick 1 (a sessionless controllable perch is offlined BEFORE reconcile_once can revive its Psyche — reuse the B2 fn, no second liveness notion); unit = boot_gate_offlines_sessionless_controllable_then_reconcile_skips_host (+ session-backed contrast hosts); int = cold-start real broker with a stale online controllable perch (no sessio...\n  721│[[requirements]]\n  722│id = \"REQ-HAZARD-BRAIN-RESTART-LIFECYCLE-REHYDRATE\"\n *723│title = \"B4 (deepest): a bare brain restart (broker survives) REHYDRATES the live-agent lifecycle so post-restart endpoints are hosted + attachable. Today resume_sessions (brainproc.rs:186, brain.rs:797-809) re-subscribes to the broker's PTY sessions but ALL BrainLifecycle instances (lifecycle.rs:58-130; the ephemeral brain.rs:254-275) are LOST on restart → a post-restart live endpoint gets no livehost → its Psyche is never (re)hosted and new spawns die / can't attach until a FULL daemon reset (opera...\n *724│required_stages = []  # CLOSED v0.12.0 wave 5 — SUBSUMED, not built (doyle+operator ruling). Per-axis evidence: every axis the REQ names is ALREADY rebuilt on a bare brain restart — (1) BrainLifecycle config → host_one with_config_in (from the adapter registry + perch, each host); (2) pulse-driver thread + Psyche host → run_brain→spawn_live_host(brainproc:230)→reconcile_once (every brain start); (3) PTY message-delivery cursors → resume_sessions(brain.rs:797); (4) online/offline honesty →...\n  725│[[requirements]]\n  726│id = \"REQ-HAZARD-BRAIN-RESTART-PSYCHE-DUP\"\n  727│title = \"A bare brain restart leaves EXACTLY ONE `{id}-psyche` process per endpoint — no duplicate. On an abrupt brain death stop_host never runs (the LiveSet + owned child handles die with the brain) and Breap's job/group only reaps at DAEMON stop, so the PRIOR brain's Psyche stays ALIVE; the respawned brain's reconcile re-hosts a SECOND Psyche and overwrites the `{id}-psyche` perch pid, leaving the old one untracked + alive = a duplicate that lingers until daemon-stop (the operator's 'brain kill+rest...\n     │...\n  730│id = \"REQ-HAZARD-UNHOST-PSYCHE-REAP\"\n *731│title = \"On un-host, the detached `{id}-psyche` HARNESS PROCESS is reaped — not just its in-brain pulse-driver thread. Today stop_host (livehost.rs:203) trips the HostedLife stop flag + JOINS the driver thread, but the Psyche is a detached harness process (spawn_psyche → ManifestRuntime detached spawn, runtime.rs:341-356; its pid is untracked in HostedLife though stamped on the `{id}-psyche` perch, where residency-confirm already reads it). So endpoint-stop / mid-life agent-death / a B2/B5 offline-th...\n  732│required_stages = []  # SUPERSEDED 2026-07-04 (F-030 W3, todlando) by REQ-PSYCHE-EPHEMERAL-DRIVER (no resident = no orphan to reap). host_one spawns no resident psyche, so un-host has nothing detached to reap — HostedLife.psyche_child + the stop_host handle-reap + stop_host_reaps_the_detached_psyche_process are DELETED. Kept as a stable pointer (no dangling required stage on deleted evidence).\n  733│[[requirements]]\n  734│id = \"REQ-ENDPOINT-PURGE\"\n     │...\n  739│title = \"An offline ReadyAgent shows in `spt endpoint run`'s picker Resume-from-history and resumes correctly — closing the gap that today only LiveAgents do. ROOT: a harness-hosted ready bind (ReadyAgent::start_homed, ready.rs) writes info.json DIRECTLY and never appends the session ledger (unlike the shared establish_perch:250 live path), so a ready agent — though it has a session_id — produces ZERO ledger rows → the picker's offline+local Resume-from-history (which gates on ledger rows) never ...\n *740│required_stages = [\"doc\", \"impl\", \"unit\", \"int\"]  # operator add v0.12.0; built. FIX(2)=SUBSUMED (verified-at-code, ZERO code): cmd_endpoint_run is type-agnostic, the no-psyche discriminator is livehost reconcile's start-side STATE gate (info.state != live_agent → skip a ready_agent perch) ahead of the psyche_init gate. doc = CONTEXT.md `spt endpoint run` is the bringup for BOTH types. impl = ReadyAgent::start_homed ledgers a Boot row (sessions::append, mirror establish_perch:250). unit = start_ledgers...\n  741│[[requirements]]\n  742│id = \"REQ-PICKER-ADAPTER-DESCRIPTION\"\n  743│title = \"The Create-new adapter-CHOICE screen of `spt endpoint run`'s picker shows a right-hand Description panel (like the Pick-existing endpoint picker's two-pane) surfacing per-adapter detail: install date, last-updated, adapter TYPE / the endpoint types it hosts, and the adapter description — so the user can see WHAT each adapter is before choosing it (today the selector lists bare names). DEFERRED fast-follow to v0.12.0 (operator 2026-06-18). (post-v0.12.0)\"\n     │...\n  865│title = \"Prefer stable parent PID / broker handle over ephemeral PID (2.1)\"\n *866│required_stages = []  # rule 5: superseded in spirit by daemon-authoritative liveness (REQ-HAZARD-DAEMON-HOSTED-LIVENESS, int-covered M3b); binds if a per-pid path ever returns (none in the daemon model)\n  867│[[requirements]]\n  868│id = \"REQ-HAZARD-STDIN-SESSION-ID\"\n  869│title = \"Stdin session_id precedence over env (2.2)\"\n     │...\n 1118│id = \"REQ-HAZARD-PUMP-IPC-DEADLINE\"\n*1119│title = \"The single-threaded peer pump's brain-IPC reads are deadline-bounded (PUMP_PEER_IO_TIMEOUT, total-wait per call); a TimedOut read POISONS the client and escalates to a SUPERVISED RESTART, never a per-peer retry — a black-holed peer must never wedge the whole pump\"\n 1120│required_stages = [\"doc\", \"impl\", \"unit\"]  # activated 2026-06-11: deployed v0.4.0 pump wedged 2.2h on hfenduleam — net_open_stream's unbounded `loop { read_event }` blocked on a peer conn the broker never replied for; the single-threaded pump froze and supervise_pump cannot rescue a BLOCKED thread (it only catches panic/error/return). Surfaced twice (2026-06-07 + this 2.2h wedge); the stall-warning was the band-aid, not the fix. fix = Brain pump-mode carrier SPLIT at construction (Brain::cold_start_pu...\n 1121│# --- Added 2026-06-16 (v0.8.3): the pump-IPC-deadline B-half — the broker-side complement of REQ-HAZARD-PUMP-IPC-DEADLINE (DEFERRED.md:42 resolved); rule 3 register-before-satisfy, rule 5 activate per wave ---\n 1122│[[requirements]]\n 1123│id = \"REQ-HAZARD-BROKER-QUIC-DEADLINE\"\n*1124│title = \"The broker bounds every brain-waiting QUIC op (dial / open_stream / send_stream) so a black-holed or dead peer fails PROMPTLY with an ORDINARY error the broker REPLIES, never an unbounded await. The bound (< the brain's 30s PUMP_PEER_IO_TIMEOUT so the BROKER fires first) surfaces to the pump as a normal broker error reply → peer_outcome's non-TimedOut arm → drop conn + redial next tick, the round CONTINUES and the heartbeat keeps advancing — it must NEVER manifest as the brain's own read-d...\n 1125│required_stages = [\"doc\", \"impl\", \"unit\", \"int\"]  # registry-first mint (rule 3); activated per wave in v0.8.3. doc = KNOWN-HAZARDS 7.8 (+ the 7.6 B-half-fixed reword + DEFERRED.md:42 RESOLVED). impl = NetHost::bounded_block_on wrapping dial/open_stream/send_stream's QUIC await sites under a quic_op_timeout (default 10s = BROKER_QUIC_OP_TIMEOUT_MS, test-overridable via set_quic_op_timeout); unit = the bound wrapper returns a prompt non-TimedOut error on a never-completing op + a fast op is untouched [net...\n 1126│# --- Added 2026-06-17 (v0.9.1): the daemon-state-wire/broker-restart skew (perri PREP-4 FINDING 1); rule 3 register-before-satisfy ---\n 1127│[[requirements]]\n     │...\n 1155│[[requirements]]\n*1156│id = \"REQ-HAZARD-DAEMON-HOSTED-LIVENESS\"\n 1157│title = \"Daemon-hosted perches (Psyche, spt-hosted Self) derive liveness from the daemon endpoint table + info.json status, never is_process_alive(info.pid) (2.5)\"\n 1158│required_stages = [\"impl\", \"unit\", \"int\"]  # int activated M3b-B9: the daemon E2E spawns the hosted Psyche, then proves a dead summarizer pid does NOT flip the daemon-online perch offline\n 1159│# --- Added 2026-06-09 from grill-with-docs (ADR-0018: broker/brain in-process-collapse regression); rule 3: register before satisfying, rule 5: inactive until the broker/brain split restoration milestone (next milestone) ---\n     │...\n 1229│id = \"REQ-HAZARD-REGISTRY-GHOST-ROWS\"\n*1230│title = \"Registry rows must decay (no immortal rows) via TWO triggers: (a) NODE-SILENCE — evict rows whose author node has not been heard (admitted inbound feed) within the eviction window, so a vanished node's rows stop poisoning bare-id resolution with phantom AcrossNodes ambiguity; AND (b) per-row OFFLINE-TTL — evict rows that have been non-routable (Offline) beyond the per-row grace even while the author node is alive, because purge/erase leaves an immortal Offline row otherwise (ghost-heal re-ad...\n 1231│required_stages = [\"doc\", \"impl\", \"unit\"]  # (a) NODE-SILENCE activated 2026-06-06 (M7 closeout fix 2): SubnetRegistry::evict_nodes + RegistryHost::evict_silent_peers (heard-map TTL, snapshots rewritten) from the registry pump tick; registry_evict_after_ms knob (default 300s = 10 default cadences). unit = model eviction + field-bug replay (ghost row → bare resolve Ambiguous; eviction restores) + self-rows-never-decay + recently-heard survives. (b) OFFLINE-TTL activated counter-38 #2-secondary (todlando...\n 1232│\n 1233│# --- Added 2026-06-07 at M8 ratification (M8-PLAN.md decisions 1-24; rule 3:\n     │...\n 1318│[[requirements]]\n*1319│id = \"REQ-HAZARD-LIVEHOST-BOOT-RACE\"\n 1320│title = \"The brain's daemon-hosted Psyche lifecycle surfaces a host-FAILURE on the live perch (harness-diagnosable) and runs net-INDEPENDENTLY. When reconcile_once→host_one→spawn_psyche fails for a state=live_agent+status=online endpoint (e.g. the adapter's psyche binary absent from its install dir, REQ-INSTALL-11), the failure MUST be written to the perch info.json as a CURRENT-STATE field (reason + ts + attempt count; overwritten each 5s retry, CLEARED on successful host) and surfaced by `spt endpo...\n*1321│required_stages = [\"impl\", \"unit\"]  # activated v0.8.1 2026-06-16 (todlando): impl = info.json psyche_host_error field + set_psyche_host_error helper + SELF-pin render annotation; unit = info.rs round-trip+set/clear/increment + render_self_pin annotation. F-030 W3 (todlando 2026-07-04): the `int` leg DROPPED (only the psyche leg — NOT fully superseded). Both real-daemon ints (livehost_bootrace_e2e = resident psyche host-Ok positive, livehost_psyche_fail_e2e = resident psyche-spawn-fail negative) tested...\n 1322│# --- Added 2026-06-16 (v0.8.2): perri's v0.8.1 dogfood surfaced two follow-on spt-core defects (F-009 command-templating argv-split, F-010 residual host-failure masking); rule 3 register-before-satisfy, rule 5 activate per wave ---\n 1323│[[requirements]]\n 1324│id = \"REQ-HAZARD-TEMPLATE-ARGV-FILL\"\n 1325│title = \"Command-template substitution fills argv ELEMENTS, not a re-tokenized string: spt-core currently `fill_template`s {key} values INTO the command STRING and THEN `tokenize`s the filled string (runtime.rs:94/122), so a multi-word {key} value whitespace-SPLITS into multiple argv tokens unless the adapter hand-quotes the placeholder, and a value containing a `\\\"` (or `;`) injects/breaks tokenization (shell-injection-adjacent). A filled value MUST become exactly ONE argv element regardless of spaces/q...\n*1326│required_stages = [\"impl\", \"unit\", \"int\"]  # activated v0.8.2 2026-06-16 (todlando-w162): impl = fill_template_tokens (tokenize-template-then-fill-each) + rewired command_for/run_bounded_command_in/harnesshost/shellwake/shellhost callers; unit = multi-word→1-element, quote/semicolon→1-element, embedded-placeholder, missing-key/empty-command, literal-brace parity. F-030 W3 (todlando 2026-07-04): the int leg was RE-HOMED (doyle ruling), NOT dropped — its prior carrier livehost_nonresident_e2e retired...\n 1327│[[requirements]]\n*1328│id = \"REQ-HAZARD-LIVEHOST-NONRESIDENT\"\n*1329│title = \"A daemon-hosted Psyche that spawns then EXITS IMMEDIATELY is a host failure, surfaced like a spawn failure (closes the v0.8.1 residual masking): the REQ-HAZARD-LIVEHOST-BOOT-RACE signal stamps `psyche_host_error` only when `spawn_psyche` returns Err, NOT when the detached spawn() returns Ok but the child dies within moments (e.g. a bad-argv child exiting 2 — the F-009 case). That leaves the residual 'online + no Psyche + no cause' gap: the nested `{id}-psyche` info.json is written status=onlin...\n*1330│required_stages = []  # SUPERSEDED 2026-07-04 (F-030 W3, todlando) by REQ-HAZARD-PSYCHE-RESIDENCY-EXPECTATION (+ the REQ-PSYCHE-EPHEMERAL-DRIVER turn budget). The resident model retired: host_one spawns no resident psyche, so \"spawned-then-exited-fast\" cannot arise — classify_residency/confirm_residency_or_unhost + the residency e2e (livehost_nonresident_e2e) are DELETED. The spirit (a psyche that cannot work is visible + hosting must not churn) transfers to the W1 note_turn_outcome consecutive-N budge...\n 1331│[[requirements]]\n 1332│id = \"REQ-HAZARD-EPOCH-RESET\"\n 1333│title = \"Advertisement-epoch reset strands a node: peers' higher last-seen epoch drops the reset node's fresh advertisements as Stale until the counter outruns history. Common case (full reinstall/re-pair) is mitigated by REQ-SUBNET-7's ceremony eviction (peer-side epoch memory dies with the deleted row — acceptance-verified); the residual narrow slice (epoch file lost, identity kept) is documented, guard deferred to a field hit (4.11)\"\n\n## docs/\n### KNOWN-HAZARDS.md\n 196│- **Failure:** the registry's only superseding mechanism is the per-`(endpoint_id, node)` epoch lease (4.8) — a row is replaced only by a newer row *from the same node*. When a node identity dies permanently (machine retired, or `node.key` regenerated so the \"node\" never speaks again), its rows are never superseded and never expire: they sit in the in-memory registries and the `identity/registry/<subnet>.json` snapshots forever. A bare-id send then resolves the same endpoint id on both the live and the...\n*197│- **Invariant:** registry rows authored by a **silent** peer node decay: a node not *heard* (admitted inbound feed — the M7 D2 heard-map, REQ-SUBNET-1) within the eviction window (`registry_evict_after_ms`, default 300s ≈ 10 default pump cadences) has its rows **evicted** from every subnet registry, snapshots rewritten. Own rows never decay (the node always hears itself implicitly — it authors them each pump tick). Eviction is safe under the lease: v1 has **no transitive gossip**, so any future upd...\n 198│- **spt-core mapping:** `spt_net::net::registry::SubnetRegistry::evict_nodes` (model) + `spt_daemon::registryhost::RegistryHost::evict_silent_peers` (heard-map TTL) driven from the registry pump tick (`peerloop`). Trust rows are NOT auto-evicted (trust is a user decision; a stale trust row only costs dead dials) — pruning those is a separate verb.\n 199│- **Source:** M7 acceptance run 2026-06-06 (DEFERRED.md \"Ghost registry row eviction\"); the AMBIGUOUS render fix rode along.\n 200│- **Mesh note (ADR-0017, 2026-06-08):** the subnet mesh **preserves** this invariant rather than superseding it. \"No transitive gossip\" sharpens to **no transitive *row* gossip** — the mesh relays only the member *roster* (discovery), while registry **rows stay own-authored and are fetched directly** from each member over a handshake. So \"any future update for a node comes from that node itself, alive\" still holds and the eviction lease is untouched. (The plan's rejected alternative — signed transiti...\n    │...\n 500│- **Failure:** the peer pump is a SINGLE thread driving every leg (registry/notif/sync/update) against every peer over ONE brain-IPC client. Its reply reads (`net_open_stream`, `net_stream_send`, `net_dial`, and the sync/update pull `read_event` loops) were `loop { read_event() }` with no deadline. When a peer's QUIC path black-holes, the broker's stream-open/send awaits the dead peer and never sends the reply, so the brain's `read_frame` blocks FOREVER and the pump freezes mid-round. The heartbeat (loop...\n*501│- **Invariant:** in PUMP mode the brain carrier is **SPLIT at construction** — a dedicated `pump-ipc-reader` thread does blocking `read_frame` on the `RecvHalf` and forwards each framed result down a channel; the main thread writes on the `SendHalf` and reads with `Receiver::recv_timeout`. Every IPC reply read is bounded by a per-call **total-wait** deadline (`PUMP_PEER_IO_TIMEOUT` = 30s, > any legitimate round-trip, < the 60s QUIC idle; re-armed on stream progress for the streaming pull legs so a heal...\n 502│- **spt-core mapping:** `Brain::cold_start_pump` (splits the carrier + arms the deadline) / `BrainConn::Split` (the `SendHalf` + reader-thread channel) / `call_deadline` / `read_event_until` / `read_frame_until` (the `recv_timeout` dispatch), `pump::run_peer_pump` (connects in pump mode) + `pump::peer_outcome` (the tier-split), the `request_sync`/`request_update` pull loops (deadline re-armed on progress). The broker-side half — the broker must never make a brain wait unbounded on a QUIC op (bound the ...\n 503│- **Source:** field diagnosis 2026-06-11 (the 2.2h hfenduleam wedge); doyle ruling A-now / B-deferred. The stall warning (M8 decision 23) was the band-aid; this is the fix. The B-half landed v0.8.3 (§7.8) after the 2026-06-16 recurrence.\n 504│\n    │...\n 514│- **Failure:** the broker's brain-facing QUIC handlers (`dispatch_net_dial` / `dispatch_net_stream_open` / `dispatch_net_stream_send`) call into `NetHost::dial` / `open_stream` / `send_stream`, whose iroh awaits (`endpoint.connect` + `prove_membership`; `open_bi`; `write_all`/`finish`) had NO bound of their own. A dead/black-holed roster peer (its process gone, or a mixed-pair that accepts the conn but never answers the seed-proof) makes the broker await its QUIC path FOREVER, so the brain escapes only v...\n*515│- **Invariant:** every brain-waiting QUIC op is wrapped in a broker-side deadline (`NetHost::bounded_block_on` → `tokio::time::timeout`, `BROKER_QUIC_OP_TIMEOUT_MS` = 10s). On elapse the future is DROPPED (cancelling the in-flight connect/stream op, so nothing is half-registered) and a non-`TimedOut` `io::Error` is returned, which the broker REPLIES as an ordinary error frame. The bound (10s) sits comfortably above any legitimate LAN/relay round-trip and 20s below the brain's 30s `PUMP_PEER_IO_TIMEOUT`...\n 516│- **spt-core mapping:** `NetHost::bounded_block_on` (the timeout wrapper) wrapping `NetHost::dial` / `open_stream` (QUIC branch) / `send_stream`; `BROKER_QUIC_OP_TIMEOUT_MS` + `set_quic_op_timeout` (test override, off `NetConfig` — mirrors `set_roster_exchange`). Unit: `bounded_block_on_cuts_a_never_completing_op_with_an_ordinary_error` (a never-completing op → prompt non-`TimedOut` error; a ready op untouched). Int: `dial_to_a_black_holing_peer_fails_with_a_bounded_ordinary_error` (the broker REPLIE...\n 517│- **Source:** the 7.6 B-half — deferred 2026-06-11 (doyle ruling A-now / B-deferred, DEFERRED.md) and shipped in v0.8.3 after the 2026-06-16 hfenduleam recurrence.\n 518│\n    │...\n 558│### 7.15 An OFFLINE spt-hosted endpoint must NOT render phantom `ONLINE+CONTROLLED` — clear `driven_by` when its session is gone  `[REQ-HAZARD-DRIVEN-BY-SELFHEAL]`\n*559│- **Failure:** `driven_by` (the `info.json` `ONLINE+CONTROLLED` latch) is single-written by the broker via `clear_controller`/`stamp_driven_by`, which only fire on a controller change. When an spt-hosted endpoint's broker session is GONE (harness dead — the B2 case, 7-series sibling `REQ-HAZARD-HOSTED-LIVENESS-RECONCILE`), no controller event ever fires, so a stale `driven_by=Some(node)` persists: the picker renders a phantom \"controlled by X\" on an endpoint that is actually OFFLINE. The B2 reconcile a...\n 560│- **Invariant:** `reconcile_hosted_liveness`, when it offlines a sessionless controllable perch (the B2 keystone — no live broker session ⇒ dead harness), ALSO clears `driven_by` (`set_driven_by(perch, None)`). RACE-FREE and single-writer-safe: with NO live broker session there is no controller to re-stamp `driven_by` concurrently, so the brain may write it here without contending the broker. (The LIVE-session leg — a controller gone while its session survives — is NOT this hazard: a clean discon...\n 561│- **MEASURED 2026-07-22 (DAEMON-LIFECYCLE W2 Leg B, `transport_death_eof.rs`) — the residual SPLITS in two, and only one half is still open.** On REAL QUIC (two brokers, `BindScope::Loopback` + a real `net_dial` — NOT the in-process duplex, which has no idle timeout at all), a remote controller whose QUIC STACK DIES WITHOUT A FIN (peer frozen: endpoint alive, UDP port still bound, keepalives unanswered — the sleeping-laptop shape) DOES self-heal: the transport's own idle timeout tears the conn, the...\n 562│- **Watch-out (repro-proven, real broker):** `SessionInfo.controller_by==None` is AMBIGUOUS — `dispatch_spawn` pre-attaches the spawner as the LOCAL controller with `by=None`, so a live LOCALLY-driven session also reads `None`. It is therefore NOT a usable standalone `driven_by` clear trigger (would false-clear a live local session). The shipped Gap-B self-heal needs no controller signal at all (it keys on session ABSENCE).\n*563│- **spt-core mapping:** impl `livehost.rs` `reconcile_hosted_liveness` Gap-B clear + the additive `SessionInfo.controller_by` observability field (`msg.rs`, populated in `broker.rs` `KIND_SESSIONS`); unit `livehost.rs` `pull_liveness…` extended (offlined sessionless perch clears `driven_by`; live/relay perches untouched); int `crates/spt-daemon/tests/driven_by_selfheal.rs` `gap_b` (real broker: reconcile offlines AND clears `driven_by`) + the A1/A2 characterization (`inject_control_wedge.rs` w5_a1/w5_a2).\n 564│- **Cross-ref 7.29 (who MAY clear vs who MUST NOT):** the BRAIN reconcile here MUST NOT clear `driven_by` off `controller_by==None` (ambiguous with a live LOCAL controller). The BROKER stamp-convergence (7.29) MAY — it reads `has_controller()` DEFINITIVELY (its own controller slot), so it authoritatively clears a stale `driven_by` against a live session with no controller (the `has_controller()==false` leg). The invariant's owner for that leg moved to the broker; this brain-side ambiguity-safety is unc...\n 565│- **Source:** v0.13.0 W5 (repro-first, todlando; doyle-assigned 2026-06-19). A2 (idle wedged-remote leg) deferred to `REQ-HAZARD-DRIVEN-BY-IDLE-REMOTE-EVICT`.\n 566│\n    │...\n 596│### 7.21 Exactly ONE `controller_writer` per brain↔broker connection — a superseded writer must write nothing further  `[REQ-HAZARD-CONTROLLER-WRITER-REORDER]`\n*597│- **Failure (doyle instrumented RACEDIAG repro, kitsubito):** on a brain-restart re-serve the handoff brain registered as controller on the SAME session TWICE over the SAME socket — `Brain::handoff` eagerly `subscribe(prior.next_seq=1)` → `become_controller(from_seq=1)`, spawning writer-A (writes seq 1); then `serve_attach` re-handled the replayed `Request{from_seq:0}` → `attach_as(sid,0)` → `become_controller(from_seq=0)`, spawning writer-B (writes 0,1). `become_controller` dropped the prior `Co...\n 598│- **Invariant:** on a single brain↔broker connection exactly ONE `controller_writer` is ever the LIVE writer; a SUPERSEDED writer writes no further frames after the epoch bump it observes; and **every `controller_writer` emits a strictly ASCENDING seq stream** (sorted initial batch + ascending live frames). The CORRECTNESS guarantee that falls out: a snap-above consumer over any interleaving of ascending writers — where the surviving writer (`serve_attach`'s `attach_as(sid,0)`) offers the COMPLETE ra...\n 599│- **spt-core mapping:** impl `brain.rs` `Brain::handoff` (KEEP the eager `subscribe`; seed `session_cursors`) + `Brain::subscribe_with` (resume-mode dedup-cursor reset to `from_seq`, shared by `attach`/`attach_as` — the operator-stream boundary fix) + `broker.rs` `OutputLog.controller_epoch: Arc<AtomicU64>` / `become_controller` (atomic `fetch_add`, passes the new epoch + `Arc::clone` into the writer) / `controller_writer` (epoch gate read UNDER the lock on both loops) / `mark_controller_gone` + the `C...\n 600│- **Source:** v0.13.0 P1c (operator-ruled root-fix before ship; doyle root-cause via instrumented repro 2026-06-20; design corrected across two gate rounds — fix #1 reverted, then the kitsubito RACEDIAG pinned the residual to the consumer-side operator-stream boundary, fixed by the `subscribe_with` cursor reset on `attach_as` re-subscribe). The last v0.13.0 ship-blocker.\n    │...\n 684│- **Failure (paid-for, field brick 2026-07-04, adapter v0.13.2):** the pre-F-030 model kept a **resident** Psyche process the daemon supervised, with residency machinery (`confirm_residency_or_unhost`) that **un-hosted the parent endpoint** when the resident child went missing. A bad adapter ship (v0.13.2) made the psyche shim exit on every turn; the residency machinery read that as a lost resident and **tore down the parent's hosted state — the ready marker was removed and never re-stamped**, so every...\n*685│- **Invariant:** a Psyche is a **bounded per-event turn**, not a resident process — its liveness is that turns succeed, so there is no \"resident is gone\" signal to react to. A psyche turn failure of ANY shape (spawn fail, non-zero exit, timeout, wedge) stamps **psyche fields only** (`psyche_host_error` on the Self perch) and **NEVER** touches the parent endpoint's `status` / ready marker / hosted state; the endpoint stays online and deliverable while its Psyche is unwell. The teardown-on-psyche-trouble...\n*686│- **spt-core mapping:** the residency/reap DELETIONS in `spt_daemon::livehost` (`reconcile_once`'s stop-side no longer un-hosts on psyche trouble; `host_one` holds no resident child) + `spt_daemon::lifecycle::run_psyche_event_turn` (a turn failure stamps `psyche_host_error` only) + the `first_turn_psyche_context` non-empty guarantee (a zero-context fresh turn can't masquerade as continue → reseed churn). Conformance int = the hall-bf shape: multi-subnet home, live endpoint, a psyche turn that fails eve...\n 687│- **Source:** F-030 (psyche-ephemeral) — paid-for by the hall-bf churn (ordinal 6491+) and the adapter v0.13.2 bad-ship brick, 2026-07-04.\n 688│\n 689│### 7.31 The Psyche failure budget must count REAL per-event attempts — a resident rate-guard is blind to per-event churn  `[REQ-HAZARD-THRASH-GUARD-BLIND]`\n    │...\n 702│\n*703│### 7.33 NO bare `.lock().unwrap()` on a broker-resident lock reachable from serve/dispatch — a poison permanently wedges every attach  `[REQ-HAZARD-BROKER-FLOOR-LOCK-POISON]`\n*704│<!-- [doc->REQ-HAZARD-BROKER-FLOOR-LOCK-POISON] -->\n*705│- **Failure (paid-for class):** a brain-only self-update restarts the BRAIN but keeps the BROKER process — and every one of its `Mutex`es — ALIVE by design (REQ-UPD-3). So a single panic while another thread held a broker-resident lock POISONS it permanently: the next `.lock().unwrap()` panics too, kills its per-conn reply thread, and EVERY subsequent attach silently deadlines (`spt rc` → \"brain IPC read deadline elapsed\") while non-locked ops keep working. The effect journal (bug #16, `[REQ-HAZARD...\n 706│- **Invariant:** NO broker-resident lock reachable from serve/dispatch may use a bare `.lock().unwrap()`. Recover via the shared `recover` (`into_inner` — safe for the short, coherent-on-recovery map ops of `sessions`/`pair_holds`) or, for the `OutputLog` ring, `recover_log` — which adds a COHERENCE CLAMP on the recovery path: a panic mid-`append` can leave the ring torn (over-cap, a last seq not below `next_seq`, non-monotonic), and serving those bytes risks garbage, so `clamp_or_reset` cheap-checks...\n*707│- **spt-core mapping:** `broker::recover` / `broker::recover_log` / `OutputLog::clamp_or_reset`. unit = a poisoned session-map recovers to a usable guard (the next attach still opens), a torn ring clamps/resets-empty only when torn (coherent untouched, `next_seq` preserved), and `recover_log` wires the clamp onto the poison branch — each RED-first against a bare `.lock().unwrap()` / an un-clamped `into_inner`. int (scripted panic-under-lock during concurrent attach → next attach opens) deferred to th...\n*708│- **Source:** REMOTE-TRUTH B-1 (pivoted — the triage's named floor sites were already closed by `[REQ-HAZARD-INJECT-WORKER-POISON]`; the surviving attach-path lock class is the real root), doyle B-1 ruling (recover all three via one helper + the OutputLog ring coherence clamp).\n 709│\n 710│### 7.34 A dead `rec.pid` on an spt-hosted perch is EXPECTED — no reader may alive-gate on `rec.pid` alone  `[REQ-HAZARD-DEAD-REC-PID]`\n 711│<!-- [doc->REQ-HAZARD-DEAD-REC-PID] -->\n    │...\n 736│\n*737│### 7.38 EVERY write on a physical broker connection is bounded + cancelable + poison-on-failure — no writer holds the send gate across an UNBOUNDED OS write  `[REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]`\n*738│<!-- [doc->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK] -->\n*739│- **Failure (paid-for, field capture + deterministic Windows repro 2026-07-09 — hertz RCA, UPDATE-WEDGE round-4):** `controller_writer` held the `SharedSend = Arc<Mutex<SendHalf>>` guard ACROSS a blocking `write_frame`. On Windows, `interprocess 2.4.2` routes the send to `WriteFileEx` + `SleepEx(INFINITE, alertable)` — this local-socket send path exposes **no supported write timeout** (`set_timeout` → `Unsupported`). An `rc --take` controller consumer that stopped reading blocked the write **indefi...\n*740│- **Invariant:** every write on a physical broker connection rides ONE broker-owned bounded/cancelable framed-write primitive with an **independent out-of-band aborter**. The `SendHalf` never leaves the conn object (a bypass cannot compile). An ABSOLUTE deadline stamped at write entry covers BOTH the serialized gate-wait AND OS write completion; a per-conn watchdog fires at that deadline and aborts the in-flight op — never relying on the write returning or on opportunistic stall-evict sampling. On dead...\n*741│- **spt-core mapping:** `spt-daemon/src/conn.rs` `BrokerConn` (Condvar-gated serialized write gate + poisoned state + per-in-flight op id + watchdog + `abort_physical()`: cfg(windows) `CancelIoEx` → completion handshake → `DisconnectNamedPipe` raw kernel32 externs; cfg(unix) `UnixStream::shutdown(Both)`); `broker.rs` `SharedSend = Arc<BrokerConn>` + `controller_writer`/`viewer_writer`/`send_frame`/`send_error` routed; `nethost.rs` stream-log/presence sends routed. int = `brain_decouple.rs` `non_drain...\n 742│- **Source:** UPDATE-WEDGE round-4 (hertz root doc `docs/UPDATE-WEDGE-ROOT-CAUSE.md` @f8596ca; doyle gate ruling `docs/UPDATE-WEDGE-RCA-DISPATCH-todlando.md`; todlando build 2026-07-09). The residual 7.36 left open: logical eviction without physical retirement.\n 743│\n 744│---\n    │...\n 774│<!-- [doc->REQ-HAZARD-MESH-BOOTSTRAP-TRAP] -->\n*775│- **Failure (paid-for, hertz field RCA 2026-07-10 — HFENDULEAM + ENLYZEAM fully sequestered from every subnet member, symmetric, green-status):** the pump resolved dial addresses from the exact `peer-addrs.json` entry else id-only discovery — never the valid `RosterEntry.address` — and every `PRESENCE_DIAL_FAILED` unconditionally `drop_seed`'d the cached entry, while the cache refilled only after a future successful seed-proof connection. One transient failure converted a warm route into id-only-or...\n 776│- **Invariant:** dial-address resolution is the ROUTE CHAIN — exact cache → validated `RosterEntry.address` (id-match required) → id-only discovery — always fully consulted in order. A failed dial demotes a cached route (suspect), never deletes a sole route; removal happens only via validated-fresher replacement or roster tombstone. Validated roster addresses reconcile into the cache at startup and on roster merge — recovery never requires an already-successful connection or operator state surg...\n 777│- **spt-core mapping:** `spt-daemon/src/pump/mod.rs` `resolve_submit_addr` + `presence_state_effect` (`PRESENCE_DIAL_FAILED` arm); `spt-store/src/peeraddrs.rs` `PeerAddrStore::put`/`drop_seed`; `spt-daemon/src/seedproofx.rs` `gapfill_peeraddrs`; status renders (`daemon status` / `subnet status`).\n 778│- **Second clause (added 2026-07-31, see 7.59):** a peer is never stranded by an address that has ALREADY failed either. Preferring the roster address assumed the retained row was fresher than the failed one; the cross-site incident falsified that assumption (7.60 froze every self row), turning this entry's fix into a `fail → demote → reinstall-same → fail` livelock. The clause is ADDITIVE — everything above still holds, and a roster address DIFFERING from the failed one still resolves and still ...\n    │...\n 782│\n*783│### 7.43 One wedged stream subscriber must NEVER stall stream serving — replay halts at the first failed write, a poisoned subscriber is removed, and recovery machinery must not manufacture new replay victims  `[REQ-HAZARD-REDISPATCH-STALL]`\n 784│<!-- [doc->REQ-HAZARD-REDISPATCH-STALL] -->\n\n### adr/\n#### 0004-single-daemon-broker-brain-split-and-self-update.md\n 26│\n*27│**Self-update delivery:** peer-propagated over P2P, layered on self-fetch, out-of-band still supported. All binaries signature-verified before handoff (spt-core's own release key) regardless of source — peer-propagation otherwise lets one compromised node poison the subnet. spt-core conducts updates for the whole stack: self first, then ripple-update each registered adapter via the adapter manifest's update declaration (file-pull or delegated command). The plugin's role shrinks to initial bootstrap only.\n 28│\n 29│**Cadence/consent:** not fully automatic by default; gated on user confirmation delivered to the most-recently-active live session, with an opt-in full-auto choice.\n 30│\n#### 0034-per-minter-op-id-namespacing.md\n 55│  monotonically shrink; a live producer minting untagged would parse as legacy\n*56│  forever and poison that discriminant. Uniform rule, no exceptions: every\n 57│  minter declares a tag; the random base becomes belt-and-suspenders.\n 58│- **One canonical minter enum.** The numeric key discriminant and the journal\n 59│  line's minter token derive from a single enum (one source for numeric +\n#### 0038-broker-stream-lifecycle-truth.md\n  88│(`dispatch.rs:414` — intended old-broker-only, catches transport\n *89│timeout/EOF/poison too), each pass installing a throwaway peek subscriber\n  90│whose `StreamLog::attach` replays the whole retained ring UNDER the per-stream\n  91│mutex with discarded write errors — one wedged subscriber conn = serial 15s\n *92│bounded-write poison windows (33 observed, all 15,000–15,154 ms) composing\n  93│into 17–62s PTY/RC stalls.\n  94│\n  95│Amended decisions:\n    │...\n  98│   unsupported-verb/old-broker answer. Every other opener-query error\n *99│   (transport timeout, EOF, poison) is `Failed` → the bounded claim requeue —\n 100│   never a second replay subscriber.\n 101│2. **(new — subscriber write discipline)** Replay and live fan-out stop at\n 102│   the FIRST failed subscriber write, and the failed subscriber is REMOVED\n    │...\n 108│3. **(new — bounded forwarding)** Attach worker completion and forwarding are\n*109│   BOUNDED and cancelable: a poisoned brain subscriber cancels its paired\n 110│   serve/wire worker (no orphan forwarding legs). Coordinate at build with\n 111│   the DAEMON-LIFECYCLE C2 stream-lease work — same neighborhood, build once.\n 112│4. **(amends Decision 3)** Retries are CLASSIFIED: a pre-setup transient may\n*113│   retry; a deadline-poisoned replay is CIRCUIT-BROKEN (global backoff), and\n 114│   a replacement subscriber is never installed until the prior subscriber is\n 115│   fully gone. This is an upgrade of the bounded backoff, NOT a revert to the\n 116│   v0.33 burn-the-claim abandonment (explicitly rejected).\n    │...\n 126│   gen/stream/family/endpoint/attempt/conn with\n*127│   replay/poison/cancel/worker-outcome events and gauges.\n 128│\n 129│Rejected here: reverting retryable claims (restores permanent abandonment);\n 130│widening any write deadline (hides the composition instead of removing it);\n*131│tearing down all subscribers on any poison (punishes healthy streams for a\n 132│wedged conn).\n#### 0039-peer-route-resolution-truth.md\n  40│behavior preserves historical corruption after the roster is corrected —\n *41│persisted poison rows were live on both incident nodes (outer key of one\n  42│member mapping another member's address).\n  43│\n  44│## Decision\n    │...\n  54│   never route-less. (\"Validated\" = the entry's address `id` matches the peer\n *55│   key being resolved — poison rows never become routes.)\n  56│2. **Retention is nondestructive.** A failed dial NEVER deletes the only\n  57│   bootstrap route. A cached address that failed is *demoted* (marked\n  58│   suspect / superseded when a fresher source disagrees), not forgotten;\n    │...\n 165│  pass pending; mint at its own triage.\n*166│- **0.30.1 peer-pump UPDATE-leg reply-read poison** (fixed): same pump, a\n 167│  different leg — the reply-read path, not route resolution.\n 168│- **PUMP_PEER_FAIL unstamped** (2026-07-14 seed): folded here — the W3 stage\n 169│  split subsumes it.\n#### 0040-bounded-stream-seat-lifecycle-one-way-terminal.md\n 31│   in the brain's event loop. While the brain does this it stops draining IPC; the broker's\n*32│   seat writer blocks, hits the 15s SharedSend bound, poisons (12→15 observed\n*33│   CONN_WRITE_POISONED, family=Registry), and the dispatcher re-replays from seq 0.\n 34│4. **Completed seats retain OS threads — family-agnostically.** The broker conn loop\n 35│   accretes `my_stream_subs` per subscribe and releases only at conn-loop exit; there is NO\n 36│   net-stream unsubscribe verb. Every completed subscription on a long-lived carrier\n   │...\n 66│   discipline (no blocking work where it stalls a drain loop) applied to the brain side.\n*67│4. **Poisoned one-way replays are bounded.** A deadline-poisoned Registry replay carries a\n 68│   per-stream strike budget; at budget the row retires TERMINAL, loudly. Safe by family\n 69│   semantics: a feed is an idempotent snapshot advertisement — the next pump round\n 70│   re-advertises. (Request/reply families keep ADR-0038 Amendment circuit-breaker\n   │...\n 94│- Steady-state broker cost on long-lived conns: threads, rows, and IPC enumeration all\n*95│  plateau at O(active). The field discriminators: hertz's preserved snapshot's poison count\n 96│  freezes; brain CPU drops to noise (with the self-hash fix, REQ-BRAIN-HASH-ONCE); broker\n 97│  thread census bounded and named.\n 98│- New wire surface: `KIND_NET_STREAM_UNSUBSCRIBE` (additive, best-effort) + a lifetime\n#### 0041-endpoint-lifecycle-truth.md\n 17│   minting `ready_agent`/`controllable=false` hybrid rows born online-authoritative. The\n*18│   livehost reconcile skips `state != live_agent` rows AND gates cleanup on\n 19│   `controllable=true`, so a dead-PID hybrid survives every restart; `is_perch_alive`\n 20│   trusts `status=online` and skips dead-PID rejection — every reader faithfully\n 21│   resurrects the lie.\n#### 0045-endpoint-teardown-authority.md\n 82│(`controllable == Some(true)`, ADR-0041) — the same predicate `cmd_bind`'s online-earn and\n*83│the livehost reconcile already use, so the codebase carries one authority concept, not\n 84│two. Harness-hosted / external endpoints keep marker + address + status behavior\n 85│unchanged: core spawned nothing there and cannot reach into the harness's process tree\n 86│(CONTEXT.md:39). The misleading doc comment claiming marker-removal stops the listener is","linesTruncated":true,"meta":{"limits":{"columnTruncated":{"maxColumn":512}}}},"isError":false,"timestamp":1788999539434}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_2AFbdSKhERDDRCqoC6T9A5Wf|fc_0dde72effda0ad21016aa1f782e3d887d2984dccad1f9b6a51","toolName":"read","startedAt":"2026-09-10T00:19:14.993Z","args":{"path":".worktrees/286-churn/docs/KNOWN-HAZARDS.md:703-742:raw"},"intent":"Reading full invariant text"},"id":"57dc0465","parentId":"4624f81b","timestamp":"2026-09-10T00:19:14.993Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_Dfwisf4jthMIiGi9y0lAsM4N|fc_0dde72effda0ad21016aa1f782e3e487d2aab46004816c44a9","toolName":"read","startedAt":"2026-09-10T00:19:14.993Z","args":{"path":".worktrees/286-churn/crates/spt-daemon/src/brain.rs:410-480,1120-1295,2228-2255"},"intent":"Reading retained client behavior"},"id":"a3553a89","parentId":"57dc0465","timestamp":"2026-09-10T00:19:14.993Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_PSEIBWsFVvBYvWWVwi0ZOOmp|fc_0dde72effda0ad21016aa1f782e3ec87d280d3274df33069a7","toolName":"read","startedAt":"2026-09-10T00:19:14.993Z","args":{"path":".worktrees/286-churn/crates/spt-daemon/src/broker.rs:5917-6035,6128-6256"},"intent":"Reading broker census dispatch"},"id":"7e3f2a69","parentId":"a3553a89","timestamp":"2026-09-10T00:19:14.993Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_1pbBYBrJdatCnkJVzIYe7xBI|fc_0dde72effda0ad21016aa1f782e3f487d2be9c040d88131b67","toolName":"read","startedAt":"2026-09-10T00:19:14.993Z","args":{"path":".worktrees/286-churn/crates/spt-proto/src/io.rs"},"intent":"Reading stderr emission primitive"},"id":"e9dc1390","parentId":"7e3f2a69","timestamp":"2026-09-10T00:19:14.993Z"}
{"type":"message","id":"e80b99bc","parentId":"e9dc1390","timestamp":"2026-09-10T00:19:14.993Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":"**Preparing requested source review**"},{"type":"toolCall","id":"call_2AFbdSKhERDDRCqoC6T9A5Wf|fc_0dde72effda0ad21016aa1f782e3d887d2984dccad1f9b6a51","name":"read","arguments":{"i":"Reading full invariant text","path":".worktrees/286-churn/docs/KNOWN-HAZARDS.md:703-742:raw"},"intent":"Reading full invariant text"},{"type":"toolCall","id":"call_Dfwisf4jthMIiGi9y0lAsM4N|fc_0dde72effda0ad21016aa1f782e3e487d2aab46004816c44a9","name":"read","arguments":{"i":"Reading retained client behavior","path":".worktrees/286-churn/crates/spt-daemon/src/brain.rs:410-480,1120-1295,2228-2255"},"intent":"Reading retained client behavior"},{"type":"toolCall","id":"call_PSEIBWsFVvBYvWWVwi0ZOOmp|fc_0dde72effda0ad21016aa1f782e3ec87d280d3274df33069a7","name":"read","arguments":{"i":"Reading broker census dispatch","path":".worktrees/286-churn/crates/spt-daemon/src/broker.rs:5917-6035,6128-6256"},"intent":"Reading broker census dispatch"},{"type":"toolCall","id":"call_1pbBYBrJdatCnkJVzIYe7xBI|fc_0dde72effda0ad21016aa1f782e3f487d2be9c040d88131b67","name":"read","arguments":{"i":"Reading stderr emission primitive","path":".worktrees/286-churn/crates/spt-proto/src/io.rs"},"intent":"Reading stderr emission primitive"},{"type":"toolCall","id":"call_WJVjoIqzGjDnaFQKCKJ3HFew|fc_0dde72effda0ad21016aa1f782e3fc87d2946409ba6b09f46f","name":"grep","arguments":{"i":"Locating stderr emission definition","pattern":"macro_rules! emit_line_err|fn emit_line_err|emit_line_err|CONN_LIFECYCLE|write-start|transport-close","path":".worktrees/286-churn/crates/spt-proto/src;.worktrees/286-churn/crates/spt-daemon/tests/conn_blackhole_lifecycle.rs;.worktrees/286-churn/scripts","case":true,"gitignore":false,"skip":null},"intent":"Locating stderr emission definition"},{"type":"toolCall","id":"call_WoMznqcZMA9QvJyFAj6sJ7Tq|fc_0dde72effda0ad21016aa1f782e40487d29934dbf8378f4a19","name":"write","arguments":{"i":"Reporting truncated document read","path":"xd://report_issue","content":"read: bare .worktrees/286-churn/docs/KNOWN-HAZARDS.md reported only 307 total lines and omitted later sections, while grep finds connection invariants at lines 737–742 in that same file; truncation footer does not expose the remaining true line range."},"intent":"Reporting truncated document read"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":28622,"output":464,"cacheRead":72576,"cacheWrite":0,"totalTokens":101662,"reasoningTokens":75,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1788999539477,"responseId":"resp_0dde72effda0ad21016aa1f773cb7087d2afa47b944891a088","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_0dde72effda0ad21016aa1f774fcec87d29344ec9693a420dd","type":"reasoning","content":[],"encrypted_content":"gAAAAABqofeCssVdK1WsdVFKGECOnxnFS-6_vZUfVvSlsJLraQJLO7i6Be_0ybjmKNrHjqV2A6cYldfOAvhsDeODEd5mYPSI84rLRk5GxZIHmn6MlhV9YH48Lq-0Lzl7CTLJQmF0AJuTDEk9Rc4kaiqjggJyQC2felbrVDyErOx-94j7T7TKaQVYeWfDoh19QyTQiyaEhsuM5_ppsyYxe62ZpmKQJdTDYMozbZ4KViQ8DbT6y-bSc2z4QniToEW9Ap8wf3KPZC1PfalzAiHFyxSJ-xTA0oGJvQGWKzl3WlGD0e71SLmrwKmZ_uoQqQjYrhlCzSqjyPApVsmkcFgRVjKjuwM8eSHiJxluBKv6GXfaXf7E0PiOc5zSndBmBbztJvU7fP0YAz8NFDRev0kRo9HdAMf38cMngU6mjcnThnXt7HdJe6TkYe16l-d7tqkpoCICNKKTof9sgLjGlQIE5zd2WOPJnqDfcqDrWMFE66FcfUusglgbxy_fOil_nXd1dYHTHan_z6o6aq03u_yP_47Y1BJybj6CJOS1fmaWzbBTV5ybqpJw5m9lNjPd-fJ-iRNeiI-jS_08e2vgxsjFPgHhx3h9h_-uixp-13djLaQl42L9SzUBhzwDKv1uMw0RORq60SqRYBKunkBBCe7cplYwVUUiQAhRK9hMlT9UDLbZ7ChjnirhFyEfljSGQPwFn9mAre97j2qt9s0OvNLr37d952c8b-AOF2QqsoABMbxZ0LITc5EuEVP5UDHdWHoDxs6okSGo9Fb6Z_XQ-Xta5KQGHDwF0--Kb3PZdEPQmmyvw1kLeXB-F2cqBrrzRy4TLBPN9t7OW2vCJ5lQJ5SIscJZhmJBfOgekHg15VxFPFpB0sLtL9bDKC01VhB4inUnHFg2pD7pHsSoJXZW5Xjx-Siw_cYaCPE5XELhtlfwblwE-nnaP-gmfKOAmEYxC4rUyxy8c79UDTJto1WVlA-2jje8vYwWoG7dxdpWo-OSVcm6oUq_eED8K0gpyhsYuA5nNqMMxbpcEJoj46P44_XWUrAoN0pFf3o8rtEO5UY4OihBiEcUczXUDskiMEJSelk2ekcxNSHRXIzuWJ07y0_Pz_Y1Me3ViXKzZYYRwM7abYFRx91TMAusehW4Dv0UinBc8v8Rx9fndfKRAbByslEo_uJTbHODFs36pOx-Flsg9K8MVh7FFZ7EQcX_ReTGFi3xvrPCCghfnSN0rFosZaPujQDuTk1b07-dXOpw1lompxfBUv_ef1ZISwt9MNcIFcFkq7ATJK5s0Mz30YNlRiKlYf-I0IB1_WVKvLGcSgsnQmiFmXZi0u1pbkTfyJFZmJ1Jwyk4-uIMOCyA-5iPxgfjBYC9hh2X2tyJIctmwHFbVUgsVQp14k20wVrQRy5oPF9lBtbC34hRHwuhrLMzGtGRE9y_UutGqzuy_1MZVUUcNp-fUXZ0K4NAtAT6PYnGq9eo8zPhh6nq59LDCp8gKfXZ-8vX6cGn6DozWaFPeGWM3q26jGGSuBBingDSUbstlKX6gXCmmhISdBKPOTNjp65X73DpaKJsXfHDZKPf1PHIZ68mliaF8SW6g5tk_GHV8HCuaoqGsZszkPfQzd6DMG5VdsC2I8P9dbdMr8PmwuDpdRbioo9RsHE8DZUwtbxeQw_q2mGFRgClLTR3myiBdI_zsz2eHoxNfQ_tkccLSq8jk9Kk9xHxO2sNiCvXYJP8bsLJyvdVhGV8ikLQd7iRv9DGupwYNWe72uMlICsYdivg3LdK0zQ_TnzZiY47jqfPpKjtXawINnW5k7sg","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"summary":[{"type":"summary_text","text":"**Preparing requested source review**"}],"metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f782e3d887d2984dccad1f9b6a51","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading full invariant text\",\"path\":\".worktrees/286-churn/docs/KNOWN-HAZARDS.md:703-742:raw\"}","call_id":"call_2AFbdSKhERDDRCqoC6T9A5Wf","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"read","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f782e3e487d2aab46004816c44a9","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading retained client behavior\",\"path\":\".worktrees/286-churn/crates/spt-daemon/src/brain.rs:410-480,1120-1295,2228-2255\"}","call_id":"call_Dfwisf4jthMIiGi9y0lAsM4N","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"read","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f782e3ec87d280d3274df33069a7","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading broker census dispatch\",\"path\":\".worktrees/286-churn/crates/spt-daemon/src/broker.rs:5917-6035,6128-6256\"}","call_id":"call_PSEIBWsFVvBYvWWVwi0ZOOmp","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"read","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f782e3f487d2be9c040d88131b67","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading stderr emission primitive\",\"path\":\".worktrees/286-churn/crates/spt-proto/src/io.rs\"}","call_id":"call_1pbBYBrJdatCnkJVzIYe7xBI","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"read","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f782e3fc87d2946409ba6b09f46f","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating stderr emission definition\",\"pattern\":\"macro_rules! emit_line_err|fn emit_line_err|emit_line_err|CONN_LIFECYCLE|write-start|transport-close\",\"path\":\".worktrees/286-churn/crates/spt-proto/src;.worktrees/286-churn/crates/spt-daemon/tests/conn_blackhole_lifecycle.rs;.worktrees/286-churn/scripts\",\"case\":true,\"gitignore\":false,\"skip\":null}","call_id":"call_WJVjoIqzGjDnaFQKCKJ3HFew","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"grep","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f782e40487d29934dbf8378f4a19","type":"function_call","status":"completed","arguments":"{\"i\":\"Reporting truncated document read\",\"path\":\"xd://report_issue\",\"content\":\"read: bare .worktrees/286-churn/docs/KNOWN-HAZARDS.md reported only 307 total lines and omitted later sections, while grep finds connection invariants at lines 737–742 in that same file; truncation footer does not expose the remaining true line range.\"}","call_id":"call_WoMznqcZMA9QvJyFAj6sJ7Tq","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"write","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}}]},"duration":15513.643599999836,"ttft":1536.8793999999762,"completedAt":1788999554992,"contextSnapshot":{"promptTokens":101198,"nonMessageTokens":13648,"compactionEpoch":0}}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_WJVjoIqzGjDnaFQKCKJ3HFew|fc_0dde72effda0ad21016aa1f782e3fc87d2946409ba6b09f46f","toolName":"grep","startedAt":"2026-09-10T00:19:14.994Z","args":{"path":".worktrees/286-churn/crates/spt-proto/src;.worktrees/286-churn/crates/spt-daemon/tests/conn_blackhole_lifecycle.rs;.worktrees/286-churn/scripts"},"intent":"Locating stderr emission definition"},"id":"a5a6e1dd","parentId":"e80b99bc","timestamp":"2026-09-10T00:19:14.994Z"}
{"type":"message","id":"ad758859","parentId":"a5a6e1dd","timestamp":"2026-09-10T00:19:15.002Z","message":{"role":"toolResult","toolCallId":"call_2AFbdSKhERDDRCqoC6T9A5Wf|fc_0dde72effda0ad21016aa1f782e3d887d2984dccad1f9b6a51","toolName":"read","content":[{"type":"text","text":"### 7.33 NO bare `.lock().unwrap()` on a broker-resident lock reachable from serve/dispatch — a poison permanently wedges every attach  `[REQ-HAZARD-BROKER-FLOOR-LOCK-POISON]`\n<!-- [doc->REQ-HAZARD-BROKER-FLOOR-LOCK-POISON] -->\n- **Failure (paid-for class):** a brain-only self-update restarts the BRAIN but keeps the BROKER process — and every one of its `Mutex`es — ALIVE by design (REQ-UPD-3). So a single panic while another thread held a broker-resident lock POISONS it permanently: the next `.lock().unwrap()` panics too, kills its per-conn reply thread, and EVERY subsequent attach silently deadlines (`spt rc` → \"brain IPC read deadline elapsed\") while non-locked ops keep working. The effect journal (bug #16, `[REQ-HAZARD-EFFECT-JOURNAL-PTY-WEDGE]`) and the inject floor (`[REQ-HAZARD-INJECT-WORKER-POISON]`) were fixed one lock at a time; the SURVIVING class was the attach-path locks — the session map (`sessions` ×18 + its `sessions_exit` alias), the per-session `OutputLog` ring (×11), and `pair_holds` (×4) — all bare `.lock().unwrap()`.\n- **Invariant:** NO broker-resident lock reachable from serve/dispatch may use a bare `.lock().unwrap()`. Recover via the shared `recover` (`into_inner` — safe for the short, coherent-on-recovery map ops of `sessions`/`pair_holds`) or, for the `OutputLog` ring, `recover_log` — which adds a COHERENCE CLAMP on the recovery path: a panic mid-`append` can leave the ring torn (over-cap, a last seq not below `next_seq`, non-monotonic), and serving those bytes risks garbage, so `clamp_or_reset` cheap-checks the invariants and RESETS the ring empty (`next_seq` preserved — cursors never rewind) + loud-logs on violation. Fail-fast on the log would REINTRODUCE the wedge; blind recover would serve torn bytes; clamp-or-reset costs only scrollback that self-heals on the next PTY output + repaint. Any NEW broker-resident lock must use `recover`/`recover_log` or carry a documented fail-fast justification.\n- **spt-core mapping:** `broker::recover` / `broker::recover_log` / `OutputLog::clamp_or_reset`. unit = a poisoned session-map recovers to a usable guard (the next attach still opens), a torn ring clamps/resets-empty only when torn (coherent untouched, `next_seq` preserved), and `recover_log` wires the clamp onto the poison branch — each RED-first against a bare `.lock().unwrap()` / an un-clamped `into_inner`. int (scripted panic-under-lock during concurrent attach → next attach opens) deferred to the two-host rig.\n- **Source:** REMOTE-TRUTH B-1 (pivoted — the triage's named floor sites were already closed by `[REQ-HAZARD-INJECT-WORKER-POISON]`; the surviving attach-path lock class is the real root), doyle B-1 ruling (recover all three via one helper + the OutputLog ring coherence clamp).\n\n### 7.34 A dead `rec.pid` on an spt-hosted perch is EXPECTED — no reader may alive-gate on `rec.pid` alone  `[REQ-HAZARD-DEAD-REC-PID]`\n<!-- [doc->REQ-HAZARD-DEAD-REC-PID] -->\n- **Failure (paid-for, field evidence — the F-026 #11 dead-pid class):** an spt-hosted endpoint's `rec.pid` records the ephemeral bind-CLI pid, which dies IMMEDIATELY after bind (the broker holds the PTY; there is no resident harness process at that pid). Readers that alive-gated on `rec.pid` alone treated the perch as stale: self-detect leg (c) could NEVER resolve an spt-hosted sender (its from-stamp degraded to `cli@NODE`, operator #7, and replies bounced `NO_PERCH` — sighted on hall-bf/ball-b, v0.24.0).\n- **Invariant:** a dead `rec.pid` on an spt-hosted perch is an EXPECTED state, not staleness. NO reader may alive-gate on `rec.pid` alone: spt-hosted LIVENESS comes from the daemon-managed `status` field (KH 2.5 — status present ⇒ authoritative, never a per-pid probe); IDENTITY comes from session/ancestry resolution, where `rec.parent_pid` (the harness pid, the stable session-binding anchor) is the ancestry candidate (ADR-0021 seed-hint discipline: pid is a bind-time SEED, not a truth anchor — re-anchoring truth in `rec.pid` is overruled by design). Any newly sighted `rec.pid`-alive-gating reader gets the same scoped fix and extends THIS requirement's evidence — no new REQ per reader.\n- **spt-core mapping:** `roster.rs` `detect_self_by_ancestry` (the first sighted reader, fixed by `[REQ-SELF-DETECT-PARENT-PID]`: `parent_pid` candidates alongside `rec.pid`). int = the E-1 red-first (`detect_self_resolves_spt_hosted_perch_via_parent_pid`: dead `rec.pid` + live-ancestor `parent_pid` resolves), dual-tagged as this class's test.\n- **Source:** REMOTE-TRUTH E-1 rider (doyle ruling 2026-07-05: scoped reader-side (b) over re-stamping `rec.pid` (a) — ADR-0021 ground truth, migration hole, blast radius).\n\n### 7.35 The cached ceremony-clock NTP offset must NOT survive an OS clock STEP — an offset measured against the pre-step clock strands every pairing for the TTL  `[REQ-HAZARD-CEREMONY-CLOCK-STEP]`\n- **Failure (paid-for, field evidence — the enlyzeam BIGNET-join RCA 2026-07-06):** enlyzeam's `w32time` was STOPPED, so its system clock drifted +3m30s; the pairing ceremony offset (REQ-PAIR-8, `ntp.rs`) correctly corrected the TOTP clock — until an operator `w32tm /resync` STEPPED the OS clock back −210s under the LIVE daemon. The cached offset had been measured against the OLD (drifted) clock and its snapshot was monotonic-only (`when: Option<Instant>`), so the step was invisible to the cache: for up to the 15-min TTL the ceremony clock stayed ~7 TOTP steps in the past, every `subnet join` returned NO_SEED_HOLDER while fresh-process probes met in <500ms, and only a daemon bounce (which forced a fresh query) healed it. Timeline-proven: refresh cadence 15:08/15:23/15:38/15:53; resync-step 15:45:46; joins failed 15:46–15:52; bounce healed instantly.\n- **Invariant:** the offset cache MUST detect an OS clock step and re-query, never apply a stale offset across it. The snapshot stores an `(Instant, SystemTime)` PAIR; on every read, if the wall clock diverged from the monotonic clock since the snapshot beyond a small tolerance (`|wall_elapsed − mono_elapsed| > ~2s`) the OS clock STEPPED ⇒ force an immediate re-query (offset recomputed against the new clock) regardless of TTL. Belt-and-braces at the join edge: `meet_seed_holder`, on search-deadline exhaustion, forces ONE fresh NTP query and runs one final sweep at the corrected step before the NO_SEED_HOLDER verdict — so a stepped-clock join self-heals WITHOUT a daemon bounce. The step tolerance sits far below one TOTP step (30 s) and well above scheduling jitter. Contract of REQ-PAIR-8 (lazy TTL cache, system-clock fallback, never sets the OS clock) is otherwise unchanged.\n- **spt-core mapping:** `spt_net::net::pairing::ntp` — `CachedOffset{when, wall_at}` pair + `refresh_needed`/`clock_stepped`/`signed_secs` (step-detect) + `force_refresh`/`invalidate` (the meet-exhaustion re-query seam); `spt_daemon::pairhost::meet_seed_holder` via `sweep_then_final_retry` (force-refresh + one final sweep on exhaustion). unit = `clock_step_forces_refresh_via_injected_reads` (injected mono/wall reads: a wall step forward OR back with mono barely moved re-queries; agreeing clocks within TTL do not; TTL-expiry and never-queried also refresh — no 15-min sleep) + `invalidate_clears_the_snapshot` + `meet_retries_once_after_refresh_on_exhaustion`/`meet_final_retry_failing_surrenders_once` (exactly one post-refresh retry, then surrender with the richer error).\n- **Source:** JOIN-TRUTH W1/D3 (doyle /diagnose enlyzeam RCA; the offset defeated by a live clock step is the D3 of four defects, D1 multihome + D2 loud-fail + D4 verbose-clock the siblings).\n\n### 7.36 The broker control plane and PTY fan-out must NEVER block on a single subscriber connection — a suspended brain conn must not wedge control  `[REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]`\n<!-- [doc->REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE] -->\n- **Failure (paid-for, rig-CONFIRMED 2026-07-06/07 — `NtSuspendProcess` on the brain, no update involved):** a controller's writer thread does a BLOCKING socket write to its brain subscriber conn. When that brain is suspended (or black-holed) the write never returns. The output-driven eviction path (`append` → `mark_controller_gone`, bounded by 7.12's `CONTROLLER_WRITE_DEADLINE`) only fires on NEW output, and `reap_dead_controller` (7.29 companion) only caught a writer that had EXITED — so a controller BLOCKED (not exited) on an IDLE session was evicted by NEITHER. Within seconds: attached output froze, `detach` did not release the control stamp (release routed through the wedged conn), reattach was REFUSED (`controlled-by` latched), `rc --take` hung, and `daemon status` still read healthy. Every brain cycle — including every `update apply` — has such a freeze window; a stalled or slow-draining NEW brain (6 sessions + psyches + WAN on the incident night) makes it a PERMANENT wedge until a bounce. `brain.ready ≠ subscribers drained` — why the 22:47 apply \"promoted\" while frozen.\n- **Invariant:** no broker client ever observes a wedge because the BRAIN stalled. (1) The controller's writer publishes its IN-FLIGHT-write window (`write_blocked_since`: `Some(Instant)` set immediately before the blocking `write_frame`, `None` after — the mutex never held ACROSS the write); a reader sees a `Some` older than `BRAIN_WRITE_DEADLINE` (15 s = 3× the 7.12 controller bound, measured full-with-**zero writer progress**, so a slow-but-DRAINING brain never trips it) as a WEDGED writer and stall-evicts it. Eviction is triggered TIME-based, not output-driven: at the take/reattach path (`resolve_subscribe`, so control self-heals against broker truth even on an idle session) AND at the `KIND_SESSIONS` reap. (2) The stamp release is BROKER-side — `clear_controller` re-stamps `driven_by`/`controlled` with no brain round-trip, so `controlled-by` can never latch behind a dead transit; `rc --take` and reattach-refuse consult live `has_controller()` truth. (3) The evict is OBSERVABLE, never a silent absence — a per-evict broker-log line plus a broker-global tally (`count` + `last_ms`) surfaced on `daemon status` (a brain that never returns is then a diagnosable fact). Idle ≠ wedged: an idle controller (writer parked on `rx.recv()`, `write_blocked_since` `None`) is NEVER evicted; only a genuinely-blocked writer is. Complements 7.12 (which bounds a merely SLOW controller's Full channel) and 7.29 (which converges stamps to broker truth) — this covers the BLOCKED-conn case both leave open. Deferred (Q2, minimal-plus): the WHOLESALE take/release verb-set migration off the brain stays out; the suspend-brain int rig is the arbiter — if it passes with the stall-evict shape, the wholesale move is scope creep. **Residual — BOUNDED by W2, CLOSED by W3 (`REQ-UPDATE-PROMOTE-DRAINED`):** the stall-evict shrank the update-apply FALSE-PROMOTE window from unbounded to `BRAIN_WRITE_DEADLINE` (15 s) — `brain.ready` could still arrive inside that window while blocked writes pended on the OLD-generation subscriber conn (`brain.ready ≠ subscribers drained`, the 22:47 apply). W3 CLOSES it: the ADR-0018 brain-trial promotion gate (`brainproc::run_trial`) now LATCHES ready-seen and promotes only on ready **AND** `TrialEnv::old_gen_drained()` — a broker-truth read (`Broker::any_local_controller_wedged`, the brain's own `by:None` conn wedged past the deadline; a remote conn is excluded, it is the W2 concern not an old-gen brain conn), no brain round-trip. Ready-but-never-drained elapses the trial window → kill + rollback (conservative; never a false-promote onto a wedge). The false-promote int rig (`tests/false_promote.rs`) exercises the promotion path itself, RED-first.\n- **spt-core mapping:** `broker.rs` — `ControllerSink::write_blocked_since` + `controller_writer` (marks the in-flight window around both the initial-batch and live-loop writes) + `controller_write_stalled` (pure predicate over injected `now`/`deadline`) + `stall_evict_controller` (Inline stamp at `resolve_subscribe`, Deferred stamp at `reap_dead_controller` so the reap closure does no I/O under the shared sessions lock) + `record_stall_evict`/`stall_evict_stats` (the tally) + `dispatch_stall_evicts` (`KIND_STALL_EVICTS` IPC); `brain.rs` `Brain::stall_evicts`; `cli.rs` `render_stall_evict_line` (daemon-status surface, public wording). unit = `controller_write_stall_predicate_distinguishes_wedged_from_idle` + `stalled_incumbent_is_evicted_on_reattach_and_control_released` + `reap_evicts_a_blocked_writer_not_only_an_exited_one` + `stall_evict_line_surfaces_only_a_real_tally`. int = `brain_decouple.rs` `suspended_brain_controller_is_stall_evicted_take_completes_viewer_ticks` — the deterministic in-process ANALOG of a suspended brain (a REMOTE controller that stops reading parks the broker writer in `write_frame`, the same primitive `NtSuspendProcess` produces), asserting EXACTLY four things: (1) an attached viewer keeps ticking during the wedge, (2) a second operator's `Control` subscribe COMPLETES via stall-evict (RED-first: remove the `resolve_subscribe` evict → `BusyControlled` + `evicts=0`), (3) it resolves promptly (well under the watchdog), (4) the evict is tallied (observability); Linux leg on kitsubito. It does NOT re-assert resume + cursor-replay byte-identity — W2's delta does not touch the replay path, so output integrity stays covered by the pre-existing ring/cursor units (effect-journal replay dedup, the `delivered_through` no-rewind CAS `advanced_cursor`, clean-repaint cold-attach). The live `NtSuspendProcess`/`SIGSTOP` run on a real brain child remains the MANUAL field rig from the incident night, not a CI test.\n- **Source:** LIFECYCLE-TRUTH W2 (doyle rig `NtSuspendProcess` root + design ruling `docs/W2-DESIGN-RULING.md` @60ec0ed — minimal-plus scope, `BRAIN_WRITE_DEADLINE` 15 s, passive resubscribe + observability rider; todlando build 2026-07-07). The v0.27→v0.28 update-wedge night's flagship defect.\n\n### 7.37 The Layer-1 settle-gate must RE-ARM per delivery on an observable PTY — a mid-session reader reattach (`/clear`) re-creates the head-swallow race  `[REQ-HAZARD-INJECT-SETTLE-REARM]`\n<!-- [doc->REQ-HAZARD-INJECT-SETTLE-REARM] -->\n- **Failure (paid-for, field-CONFIRMED on 0.29.0 — doyle diagnosis + perri screenshot `WindowsTerminal_6iSjya8pMt.png`):** the shipped W5-A settle-gate (7-of-W5, `REQ-INJECT-MULTILINE-INTEGRITY`) gated Layer 1 behind a worker-local ONE-SHOT (`settled_once`) on the premise that the head-swallow race is a STARTUP condition (the input reader not yet attached after spawn). That premise is FALSE: a mid-session `/clear` re-enters the harness's raw-mode input reader, re-attaching it and RE-CREATING the pre-settle window — but the one-shot had already fired at spawn, so `settle_before_inject` was SKIPPED and the head was eaten again. A checkpoint-wake payload injected right after `/clear` arrived head-truncated (mid-path `spt/Cargo.toml)`); echo-verify (Layer 2) is a default-OFF declared capability (`SPT_INJECT_VERIFY_ECHO`) for that session, so the loss was silent AND unrecoverable (a live-SENT inject leaves no spool copy).\n- **Invariant:** the settle-gate re-arms before EVERY delivery on an OBSERVABLE (echoing/interactive) PTY — the bug-prone class, where re-settling is cheap (a tick or two once the reader answers) and necessary (a mid-session reader reattach must be re-confirmed). The steady-state settle is latch-skipped ONLY where the probe is known UNOBSERVABLE on this PTY (a non-echoing ConPTY whose DSR answer never surfaces): there is no reader-reattach race to guard, and each settle would burn the full `INJECT_SETTLE` deadline. The class is discriminated by the settle's OWN return value — `settle_before_inject` returns `true` iff the session ring advanced (probe observed) — latched on the FIRST attempt only; a RE-DRIVE (`attempt > 1`, reached only on a swallowed head) ALWAYS settles regardless of the latch. This preserves the old one-shot's sole legitimate purpose (sparing a non-echoing ConPTY the deadline on every delivery) while closing the mid-session recurrence for echoing PTYs.\n- **spt-core mapping:** `broker.rs` `run_inject_worker` — the `settled_once: bool` one-shot is replaced by a `probe_unobservable: bool` latch (init `false` = assume observable → settle first delivery); the gate is the pure `should_settle(attempt, probe_unobservable) = attempt > 1 || !probe_unobservable`, and the first-attempt settle's `observed` return sets `probe_unobservable = !observed`. `settle_before_inject` gains a `bool` return (observed vs timed-out). unit = `should_settle_rearms_on_observable_pty` (observable PTY re-settles each delivery; unobservable-probe PTY skips the steady-state settle; a re-drive settles in both classes — RED-first: restore the one-shot and the observable-PTY assert flips). No int edit: the `inject_control_wedge.rs` mocks (`findstr`/`cat`) never answer the DSR probe → the settle times out → `probe_unobservable` latches true after delivery 1 → behaviour is IDENTICAL to the old one-shot for those non-echoing mocks; only a genuinely echoing PTY changes (it re-settles), which is the fix. Rides `REQ-INJECT-MULTILINE-INTEGRITY`'s existing int coverage.\n- **Source:** post-0.29.0 field fix (doyle diagnosis + design ruling `BUILD-SPEC-inject-settle-rearm`; todlando build 2026-07-08). An impl bug in the shipped W5-A Layer-1 fix, not a new invariant class — the settle-gate was correct but under-armed.\n\n### 7.38 EVERY write on a physical broker connection is bounded + cancelable + poison-on-failure — no writer holds the send gate across an UNBOUNDED OS write  `[REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]`\n<!-- [doc->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK] -->\n- **Failure (paid-for, field capture + deterministic Windows repro 2026-07-09 — hertz RCA, UPDATE-WEDGE round-4):** `controller_writer` held the `SharedSend = Arc<Mutex<SendHalf>>` guard ACROSS a blocking `write_frame`. On Windows, `interprocess 2.4.2` routes the send to `WriteFileEx` + `SleepEx(INFINITE, alertable)` — this local-socket send path exposes **no supported write timeout** (`set_timeout` → `Unsupported`). An `rc --take` controller consumer that stopped reading blocked the write **indefinitely** (~127.95 s in the capture, released only when a brain restart tore the connections down). The 7.36 logical `stall_evict_controller` removed the controller ROLE but neither canceled the in-flight pipe write, closed the physical connection, nor called `CancelIoEx` — the detached writer kept its `SharedSend` clone and live stack-owned mutex guard, and every other write queued at that gate wedged behind it. Load-gated: needs a real `seq>0` frame plus a non-draining consumer (seq-0 boot conns and a quiescent update are clean — why v0.30.5's controlled apply passed). `CTRL_WRITE_LOCKED wait_us=0` on all four capture writers exonerates the mutex convoy: the block is INSIDE the OS write, after lock acquisition. Neither existing deadline bounds it: `CONTROLLER_WRITE_DEADLINE` is output-driven channel-full handling (a writer blocked on its first owned frame never fills the channel); `BRAIN_WRITE_DEADLINE` is an age predicate sampled opportunistically, not an I/O timer.\n- **Invariant:** every write on a physical broker connection rides ONE broker-owned bounded/cancelable framed-write primitive with an **independent out-of-band aborter**. The `SendHalf` never leaves the conn object (a bypass cannot compile). An ABSOLUTE deadline stamped at write entry covers BOTH the serialized gate-wait AND OS write completion; a per-conn watchdog fires at that deadline and aborts the in-flight op — never relying on the write returning or on opportunistic stall-evict sampling. On deadline / partial write / cancellation / unknown completion: (1) poison the whole physical conn idempotently, (2) abort read+write so `handle_conn` reaches existing EOF cleanup, (3) wait for the canceled op to report completion before releasing its buffer, (4) NEVER reuse the conn (a timed-out length-prefixed frame may be partially written — reuse would corrupt framing), (5) finish/join the retired writer before reporting physical cleanup. Controller (replay + live), viewer, dispatch-reply, and nethost stream/presence writes ALL route through the primitive — any raw unbounded write left behind the gate preserves the failure class. NO additional output queue (the bounded queue + isolated writer already exist; the block is BELOW them); NO `PIPE_NOWAIT` (recorded mid-frame corruption risk). The logical stall-evict (7.36) stays as the ROLE-release trigger; physical retirement authority is the bounded write itself. Unix keeps its existing semantics under the same poison/retire invariant.\n- **spt-core mapping:** `spt-daemon/src/conn.rs` `BrokerConn` (Condvar-gated serialized write gate + poisoned state + per-in-flight op id + watchdog + `abort_physical()`: cfg(windows) `CancelIoEx` → completion handshake → `DisconnectNamedPipe` raw kernel32 externs; cfg(unix) `UnixStream::shutdown(Both)`); `broker.rs` `SharedSend = Arc<BrokerConn>` + `controller_writer`/`viewer_writer`/`send_frame`/`send_error` routed; `nethost.rs` stream-log/presence sends routed. int = `brain_decouple.rs` `non_draining_controller_stall_evict_releases_writer_and_connection` (`#[cfg(windows)]`, the hertz 9-step: real PTY burst, black-holed remote controller, exactly one logical stall-evict via an independent draining conn, the black-holed conn's `handle_conn` finishes ≤2 s of logical release WITHOUT dropping the client, a write on the old client conn fails, a fresh controller resumes from the frozen cursor) — RED-first pre-fix; the cross-platform companion `suspended_brain_controller_is_stall_evicted_take_completes_viewer_ticks` stays green on both OSes.\n- **Source:** UPDATE-WEDGE round-4 (hertz root doc `docs/UPDATE-WEDGE-ROOT-CAUSE.md` @f8596ca; doyle gate ruling `docs/UPDATE-WEDGE-RCA-DISPATCH-todlando.md`; todlando build 2026-07-09). The residual 7.36 left open: logical eviction without physical retirement.\n\n[Showing lines 703-742 of 1092. Use :743 to continue]"}],"details":{"truncation":{"content":"### 7.33 NO bare `.lock().unwrap()` on a broker-resident lock reachable from serve/dispatch — a poison permanently wedges every attach  `[REQ-HAZARD-BROKER-FLOOR-LOCK-POISON]`\r\n<!-- [doc->REQ-HAZARD-BROKER-FLOOR-LOCK-POISON] -->\r\n- **Failure (paid-for class):** a brain-only self-update restarts the BRAIN but keeps the BROKER process — and every one of its `Mutex`es — ALIVE by design (REQ-UPD-3). So a single panic while another thread held a broker-resident lock POISONS it permanently: the next `.lock().unwrap()` panics too, kills its per-conn reply thread, and EVERY subsequent attach silently deadlines (`spt rc` → \"brain IPC read deadline elapsed\") while non-locked ops keep working. The effect journal (bug #16, `[REQ-HAZARD-EFFECT-JOURNAL-PTY-WEDGE]`) and the inject floor (`[REQ-HAZARD-INJECT-WORKER-POISON]`) were fixed one lock at a time; the SURVIVING class was the attach-path locks — the session map (`sessions` ×18 + its `sessions_exit` alias), the per-session `OutputLog` ring (×11), and `pair_holds` (×4) — all bare `.lock().unwrap()`.\r\n- **Invariant:** NO broker-resident lock reachable from serve/dispatch may use a bare `.lock().unwrap()`. Recover via the shared `recover` (`into_inner` — safe for the short, coherent-on-recovery map ops of `sessions`/`pair_holds`) or, for the `OutputLog` ring, `recover_log` — which adds a COHERENCE CLAMP on the recovery path: a panic mid-`append` can leave the ring torn (over-cap, a last seq not below `next_seq`, non-monotonic), and serving those bytes risks garbage, so `clamp_or_reset` cheap-checks the invariants and RESETS the ring empty (`next_seq` preserved — cursors never rewind) + loud-logs on violation. Fail-fast on the log would REINTRODUCE the wedge; blind recover would serve torn bytes; clamp-or-reset costs only scrollback that self-heals on the next PTY output + repaint. Any NEW broker-resident lock must use `recover`/`recover_log` or carry a documented fail-fast justification.\r\n- **spt-core mapping:** `broker::recover` / `broker::recover_log` / `OutputLog::clamp_or_reset`. unit = a poisoned session-map recovers to a usable guard (the next attach still opens), a torn ring clamps/resets-empty only when torn (coherent untouched, `next_seq` preserved), and `recover_log` wires the clamp onto the poison branch — each RED-first against a bare `.lock().unwrap()` / an un-clamped `into_inner`. int (scripted panic-under-lock during concurrent attach → next attach opens) deferred to the two-host rig.\r\n- **Source:** REMOTE-TRUTH B-1 (pivoted — the triage's named floor sites were already closed by `[REQ-HAZARD-INJECT-WORKER-POISON]`; the surviving attach-path lock class is the real root), doyle B-1 ruling (recover all three via one helper + the OutputLog ring coherence clamp).\r\n\r\n### 7.34 A dead `rec.pid` on an spt-hosted perch is EXPECTED — no reader may alive-gate on `rec.pid` alone  `[REQ-HAZARD-DEAD-REC-PID]`\r\n<!-- [doc->REQ-HAZARD-DEAD-REC-PID] -->\r\n- **Failure (paid-for, field evidence — the F-026 #11 dead-pid class):** an spt-hosted endpoint's `rec.pid` records the ephemeral bind-CLI pid, which dies IMMEDIATELY after bind (the broker holds the PTY; there is no resident harness process at that pid). Readers that alive-gated on `rec.pid` alone treated the perch as stale: self-detect leg (c) could NEVER resolve an spt-hosted sender (its from-stamp degraded to `cli@NODE`, operator #7, and replies bounced `NO_PERCH` — sighted on hall-bf/ball-b, v0.24.0).\r\n- **Invariant:** a dead `rec.pid` on an spt-hosted perch is an EXPECTED state, not staleness. NO reader may alive-gate on `rec.pid` alone: spt-hosted LIVENESS comes from the daemon-managed `status` field (KH 2.5 — status present ⇒ authoritative, never a per-pid probe); IDENTITY comes from session/ancestry resolution, where `rec.parent_pid` (the harness pid, the stable session-binding anchor) is the ancestry candidate (ADR-0021 seed-hint discipline: pid is a bind-time SEED, not a truth anchor — re-anchoring truth in `rec.pid` is overruled by design). Any newly sighted `rec.pid`-alive-gating reader gets the same scoped fix and extends THIS requirement's evidence — no new REQ per reader.\r\n- **spt-core mapping:** `roster.rs` `detect_self_by_ancestry` (the first sighted reader, fixed by `[REQ-SELF-DETECT-PARENT-PID]`: `parent_pid` candidates alongside `rec.pid`). int = the E-1 red-first (`detect_self_resolves_spt_hosted_perch_via_parent_pid`: dead `rec.pid` + live-ancestor `parent_pid` resolves), dual-tagged as this class's test.\r\n- **Source:** REMOTE-TRUTH E-1 rider (doyle ruling 2026-07-05: scoped reader-side (b) over re-stamping `rec.pid` (a) — ADR-0021 ground truth, migration hole, blast radius).\r\n\r\n### 7.35 The cached ceremony-clock NTP offset must NOT survive an OS clock STEP — an offset measured against the pre-step clock strands every pairing for the TTL  `[REQ-HAZARD-CEREMONY-CLOCK-STEP]`\r\n- **Failure (paid-for, field evidence — the enlyzeam BIGNET-join RCA 2026-07-06):** enlyzeam's `w32time` was STOPPED, so its system clock drifted +3m30s; the pairing ceremony offset (REQ-PAIR-8, `ntp.rs`) correctly corrected the TOTP clock — until an operator `w32tm /resync` STEPPED the OS clock back −210s under the LIVE daemon. The cached offset had been measured against the OLD (drifted) clock and its snapshot was monotonic-only (`when: Option<Instant>`), so the step was invisible to the cache: for up to the 15-min TTL the ceremony clock stayed ~7 TOTP steps in the past, every `subnet join` returned NO_SEED_HOLDER while fresh-process probes met in <500ms, and only a daemon bounce (which forced a fresh query) healed it. Timeline-proven: refresh cadence 15:08/15:23/15:38/15:53; resync-step 15:45:46; joins failed 15:46–15:52; bounce healed instantly.\r\n- **Invariant:** the offset cache MUST detect an OS clock step and re-query, never apply a stale offset across it. The snapshot stores an `(Instant, SystemTime)` PAIR; on every read, if the wall clock diverged from the monotonic clock since the snapshot beyond a small tolerance (`|wall_elapsed − mono_elapsed| > ~2s`) the OS clock STEPPED ⇒ force an immediate re-query (offset recomputed against the new clock) regardless of TTL. Belt-and-braces at the join edge: `meet_seed_holder`, on search-deadline exhaustion, forces ONE fresh NTP query and runs one final sweep at the corrected step before the NO_SEED_HOLDER verdict — so a stepped-clock join self-heals WITHOUT a daemon bounce. The step tolerance sits far below one TOTP step (30 s) and well above scheduling jitter. Contract of REQ-PAIR-8 (lazy TTL cache, system-clock fallback, never sets the OS clock) is otherwise unchanged.\r\n- **spt-core mapping:** `spt_net::net::pairing::ntp` — `CachedOffset{when, wall_at}` pair + `refresh_needed`/`clock_stepped`/`signed_secs` (step-detect) + `force_refresh`/`invalidate` (the meet-exhaustion re-query seam); `spt_daemon::pairhost::meet_seed_holder` via `sweep_then_final_retry` (force-refresh + one final sweep on exhaustion). unit = `clock_step_forces_refresh_via_injected_reads` (injected mono/wall reads: a wall step forward OR back with mono barely moved re-queries; agreeing clocks within TTL do not; TTL-expiry and never-queried also refresh — no 15-min sleep) + `invalidate_clears_the_snapshot` + `meet_retries_once_after_refresh_on_exhaustion`/`meet_final_retry_failing_surrenders_once` (exactly one post-refresh retry, then surrender with the richer error).\r\n- **Source:** JOIN-TRUTH W1/D3 (doyle /diagnose enlyzeam RCA; the offset defeated by a live clock step is the D3 of four defects, D1 multihome + D2 loud-fail + D4 verbose-clock the siblings).\r\n\r\n### 7.36 The broker control plane and PTY fan-out must NEVER block on a single subscriber connection — a suspended brain conn must not wedge control  `[REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]`\r\n<!-- [doc->REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE] -->\r\n- **Failure (paid-for, rig-CONFIRMED 2026-07-06/07 — `NtSuspendProcess` on the brain, no update involved):** a controller's writer thread does a BLOCKING socket write to its brain subscriber conn. When that brain is suspended (or black-holed) the write never returns. The output-driven eviction path (`append` → `mark_controller_gone`, bounded by 7.12's `CONTROLLER_WRITE_DEADLINE`) only fires on NEW output, and `reap_dead_controller` (7.29 companion) only caught a writer that had EXITED — so a controller BLOCKED (not exited) on an IDLE session was evicted by NEITHER. Within seconds: attached output froze, `detach` did not release the control stamp (release routed through the wedged conn), reattach was REFUSED (`controlled-by` latched), `rc --take` hung, and `daemon status` still read healthy. Every brain cycle — including every `update apply` — has such a freeze window; a stalled or slow-draining NEW brain (6 sessions + psyches + WAN on the incident night) makes it a PERMANENT wedge until a bounce. `brain.ready ≠ subscribers drained` — why the 22:47 apply \"promoted\" while frozen.\r\n- **Invariant:** no broker client ever observes a wedge because the BRAIN stalled. (1) The controller's writer publishes its IN-FLIGHT-write window (`write_blocked_since`: `Some(Instant)` set immediately before the blocking `write_frame`, `None` after — the mutex never held ACROSS the write); a reader sees a `Some` older than `BRAIN_WRITE_DEADLINE` (15 s = 3× the 7.12 controller bound, measured full-with-**zero writer progress**, so a slow-but-DRAINING brain never trips it) as a WEDGED writer and stall-evicts it. Eviction is triggered TIME-based, not output-driven: at the take/reattach path (`resolve_subscribe`, so control self-heals against broker truth even on an idle session) AND at the `KIND_SESSIONS` reap. (2) The stamp release is BROKER-side — `clear_controller` re-stamps `driven_by`/`controlled` with no brain round-trip, so `controlled-by` can never latch behind a dead transit; `rc --take` and reattach-refuse consult live `has_controller()` truth. (3) The evict is OBSERVABLE, never a silent absence — a per-evict broker-log line plus a broker-global tally (`count` + `last_ms`) surfaced on `daemon status` (a brain that never returns is then a diagnosable fact). Idle ≠ wedged: an idle controller (writer parked on `rx.recv()`, `write_blocked_since` `None`) is NEVER evicted; only a genuinely-blocked writer is. Complements 7.12 (which bounds a merely SLOW controller's Full channel) and 7.29 (which converges stamps to broker truth) — this covers the BLOCKED-conn case both leave open. Deferred (Q2, minimal-plus): the WHOLESALE take/release verb-set migration off the brain stays out; the suspend-brain int rig is the arbiter — if it passes with the stall-evict shape, the wholesale move is scope creep. **Residual — BOUNDED by W2, CLOSED by W3 (`REQ-UPDATE-PROMOTE-DRAINED`):** the stall-evict shrank the update-apply FALSE-PROMOTE window from unbounded to `BRAIN_WRITE_DEADLINE` (15 s) — `brain.ready` could still arrive inside that window while blocked writes pended on the OLD-generation subscriber conn (`brain.ready ≠ subscribers drained`, the 22:47 apply). W3 CLOSES it: the ADR-0018 brain-trial promotion gate (`brainproc::run_trial`) now LATCHES ready-seen and promotes only on ready **AND** `TrialEnv::old_gen_drained()` — a broker-truth read (`Broker::any_local_controller_wedged`, the brain's own `by:None` conn wedged past the deadline; a remote conn is excluded, it is the W2 concern not an old-gen brain conn), no brain round-trip. Ready-but-never-drained elapses the trial window → kill + rollback (conservative; never a false-promote onto a wedge). The false-promote int rig (`tests/false_promote.rs`) exercises the promotion path itself, RED-first.\r\n- **spt-core mapping:** `broker.rs` — `ControllerSink::write_blocked_since` + `controller_writer` (marks the in-flight window around both the initial-batch and live-loop writes) + `controller_write_stalled` (pure predicate over injected `now`/`deadline`) + `stall_evict_controller` (Inline stamp at `resolve_subscribe`, Deferred stamp at `reap_dead_controller` so the reap closure does no I/O under the shared sessions lock) + `record_stall_evict`/`stall_evict_stats` (the tally) + `dispatch_stall_evicts` (`KIND_STALL_EVICTS` IPC); `brain.rs` `Brain::stall_evicts`; `cli.rs` `render_stall_evict_line` (daemon-status surface, public wording). unit = `controller_write_stall_predicate_distinguishes_wedged_from_idle` + `stalled_incumbent_is_evicted_on_reattach_and_control_released` + `reap_evicts_a_blocked_writer_not_only_an_exited_one` + `stall_evict_line_surfaces_only_a_real_tally`. int = `brain_decouple.rs` `suspended_brain_controller_is_stall_evicted_take_completes_viewer_ticks` — the deterministic in-process ANALOG of a suspended brain (a REMOTE controller that stops reading parks the broker writer in `write_frame`, the same primitive `NtSuspendProcess` produces), asserting EXACTLY four things: (1) an attached viewer keeps ticking during the wedge, (2) a second operator's `Control` subscribe COMPLETES via stall-evict (RED-first: remove the `resolve_subscribe` evict → `BusyControlled` + `evicts=0`), (3) it resolves promptly (well under the watchdog), (4) the evict is tallied (observability); Linux leg on kitsubito. It does NOT re-assert resume + cursor-replay byte-identity — W2's delta does not touch the replay path, so output integrity stays covered by the pre-existing ring/cursor units (effect-journal replay dedup, the `delivered_through` no-rewind CAS `advanced_cursor`, clean-repaint cold-attach). The live `NtSuspendProcess`/`SIGSTOP` run on a real brain child remains the MANUAL field rig from the incident night, not a CI test.\r\n- **Source:** LIFECYCLE-TRUTH W2 (doyle rig `NtSuspendProcess` root + design ruling `docs/W2-DESIGN-RULING.md` @60ec0ed — minimal-plus scope, `BRAIN_WRITE_DEADLINE` 15 s, passive resubscribe + observability rider; todlando build 2026-07-07). The v0.27→v0.28 update-wedge night's flagship defect.\r\n\r\n### 7.37 The Layer-1 settle-gate must RE-ARM per delivery on an observable PTY — a mid-session reader reattach (`/clear`) re-creates the head-swallow race  `[REQ-HAZARD-INJECT-SETTLE-REARM]`\r\n<!-- [doc->REQ-HAZARD-INJECT-SETTLE-REARM] -->\r\n- **Failure (paid-for, field-CONFIRMED on 0.29.0 — doyle diagnosis + perri screenshot `WindowsTerminal_6iSjya8pMt.png`):** the shipped W5-A settle-gate (7-of-W5, `REQ-INJECT-MULTILINE-INTEGRITY`) gated Layer 1 behind a worker-local ONE-SHOT (`settled_once`) on the premise that the head-swallow race is a STARTUP condition (the input reader not yet attached after spawn). That premise is FALSE: a mid-session `/clear` re-enters the harness's raw-mode input reader, re-attaching it and RE-CREATING the pre-settle window — but the one-shot had already fired at spawn, so `settle_before_inject` was SKIPPED and the head was eaten again. A checkpoint-wake payload injected right after `/clear` arrived head-truncated (mid-path `spt/Cargo.toml)`); echo-verify (Layer 2) is a default-OFF declared capability (`SPT_INJECT_VERIFY_ECHO`) for that session, so the loss was silent AND unrecoverable (a live-SENT inject leaves no spool copy).\r\n- **Invariant:** the settle-gate re-arms before EVERY delivery on an OBSERVABLE (echoing/interactive) PTY — the bug-prone class, where re-settling is cheap (a tick or two once the reader answers) and necessary (a mid-session reader reattach must be re-confirmed). The steady-state settle is latch-skipped ONLY where the probe is known UNOBSERVABLE on this PTY (a non-echoing ConPTY whose DSR answer never surfaces): there is no reader-reattach race to guard, and each settle would burn the full `INJECT_SETTLE` deadline. The class is discriminated by the settle's OWN return value — `settle_before_inject` returns `true` iff the session ring advanced (probe observed) — latched on the FIRST attempt only; a RE-DRIVE (`attempt > 1`, reached only on a swallowed head) ALWAYS settles regardless of the latch. This preserves the old one-shot's sole legitimate purpose (sparing a non-echoing ConPTY the deadline on every delivery) while closing the mid-session recurrence for echoing PTYs.\r\n- **spt-core mapping:** `broker.rs` `run_inject_worker` — the `settled_once: bool` one-shot is replaced by a `probe_unobservable: bool` latch (init `false` = assume observable → settle first delivery); the gate is the pure `should_settle(attempt, probe_unobservable) = attempt > 1 || !probe_unobservable`, and the first-attempt settle's `observed` return sets `probe_unobservable = !observed`. `settle_before_inject` gains a `bool` return (observed vs timed-out). unit = `should_settle_rearms_on_observable_pty` (observable PTY re-settles each delivery; unobservable-probe PTY skips the steady-state settle; a re-drive settles in both classes — RED-first: restore the one-shot and the observable-PTY assert flips). No int edit: the `inject_control_wedge.rs` mocks (`findstr`/`cat`) never answer the DSR probe → the settle times out → `probe_unobservable` latches true after delivery 1 → behaviour is IDENTICAL to the old one-shot for those non-echoing mocks; only a genuinely echoing PTY changes (it re-settles), which is the fix. Rides `REQ-INJECT-MULTILINE-INTEGRITY`'s existing int coverage.\r\n- **Source:** post-0.29.0 field fix (doyle diagnosis + design ruling `BUILD-SPEC-inject-settle-rearm`; todlando build 2026-07-08). An impl bug in the shipped W5-A Layer-1 fix, not a new invariant class — the settle-gate was correct but under-armed.\r\n\r\n### 7.38 EVERY write on a physical broker connection is bounded + cancelable + poison-on-failure — no writer holds the send gate across an UNBOUNDED OS write  `[REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]`\r\n<!-- [doc->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK] -->\r\n- **Failure (paid-for, field capture + deterministic Windows repro 2026-07-09 — hertz RCA, UPDATE-WEDGE round-4):** `controller_writer` held the `SharedSend = Arc<Mutex<SendHalf>>` guard ACROSS a blocking `write_frame`. On Windows, `interprocess 2.4.2` routes the send to `WriteFileEx` + `SleepEx(INFINITE, alertable)` — this local-socket send path exposes **no supported write timeout** (`set_timeout` → `Unsupported`). An `rc --take` controller consumer that stopped reading blocked the write **indefinitely** (~127.95 s in the capture, released only when a brain restart tore the connections down). The 7.36 logical `stall_evict_controller` removed the controller ROLE but neither canceled the in-flight pipe write, closed the physical connection, nor called `CancelIoEx` — the detached writer kept its `SharedSend` clone and live stack-owned mutex guard, and every other write queued at that gate wedged behind it. Load-gated: needs a real `seq>0` frame plus a non-draining consumer (seq-0 boot conns and a quiescent update are clean — why v0.30.5's controlled apply passed). `CTRL_WRITE_LOCKED wait_us=0` on all four capture writers exonerates the mutex convoy: the block is INSIDE the OS write, after lock acquisition. Neither existing deadline bounds it: `CONTROLLER_WRITE_DEADLINE` is output-driven channel-full handling (a writer blocked on its first owned frame never fills the channel); `BRAIN_WRITE_DEADLINE` is an age predicate sampled opportunistically, not an I/O timer.\r\n- **Invariant:** every write on a physical broker connection rides ONE broker-owned bounded/cancelable framed-write primitive with an **independent out-of-band aborter**. The `SendHalf` never leaves the conn object (a bypass cannot compile). An ABSOLUTE deadline stamped at write entry covers BOTH the serialized gate-wait AND OS write completion; a per-conn watchdog fires at that deadline and aborts the in-flight op — never relying on the write returning or on opportunistic stall-evict sampling. On deadline / partial write / cancellation / unknown completion: (1) poison the whole physical conn idempotently, (2) abort read+write so `handle_conn` reaches existing EOF cleanup, (3) wait for the canceled op to report completion before releasing its buffer, (4) NEVER reuse the conn (a timed-out length-prefixed frame may be partially written — reuse would corrupt framing), (5) finish/join the retired writer before reporting physical cleanup. Controller (replay + live), viewer, dispatch-reply, and nethost stream/presence writes ALL route through the primitive — any raw unbounded write left behind the gate preserves the failure class. NO additional output queue (the bounded queue + isolated writer already exist; the block is BELOW them); NO `PIPE_NOWAIT` (recorded mid-frame corruption risk). The logical stall-evict (7.36) stays as the ROLE-release trigger; physical retirement authority is the bounded write itself. Unix keeps its existing semantics under the same poison/retire invariant.\r\n- **spt-core mapping:** `spt-daemon/src/conn.rs` `BrokerConn` (Condvar-gated serialized write gate + poisoned state + per-in-flight op id + watchdog + `abort_physical()`: cfg(windows) `CancelIoEx` → completion handshake → `DisconnectNamedPipe` raw kernel32 externs; cfg(unix) `UnixStream::shutdown(Both)`); `broker.rs` `SharedSend = Arc<BrokerConn>` + `controller_writer`/`viewer_writer`/`send_frame`/`send_error` routed; `nethost.rs` stream-log/presence sends routed. int = `brain_decouple.rs` `non_draining_controller_stall_evict_releases_writer_and_connection` (`#[cfg(windows)]`, the hertz 9-step: real PTY burst, black-holed remote controller, exactly one logical stall-evict via an independent draining conn, the black-holed conn's `handle_conn` finishes ≤2 s of logical release WITHOUT dropping the client, a write on the old client conn fails, a fresh controller resumes from the frozen cursor) — RED-first pre-fix; the cross-platform companion `suspended_brain_controller_is_stall_evicted_take_completes_viewer_ticks` stays green on both OSes.\r\n- **Source:** UPDATE-WEDGE round-4 (hertz root doc `docs/UPDATE-WEDGE-ROOT-CAUSE.md` @f8596ca; doyle gate ruling `docs/UPDATE-WEDGE-RCA-DISPATCH-todlando.md`; todlando build 2026-07-09). The residual 7.36 left open: logical eviction without physical retirement.\r","truncated":true,"truncatedBy":"lines","totalLines":390,"totalBytes":22158,"outputLines":40,"outputBytes":22158,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":1092,"displayContent":{"text":"### 7.33 NO bare `.lock().unwrap()` on a broker-resident lock reachable from serve/dispatch — a poison permanently wedges every attach  `[REQ-HAZARD-BROKER-FLOOR-LOCK-POISON]`\r\n<!-- [doc->REQ-HAZARD-BROKER-FLOOR-LOCK-POISON] -->\r\n- **Failure (paid-for class):** a brain-only self-update restarts the BRAIN but keeps the BROKER process — and every one of its `Mutex`es — ALIVE by design (REQ-UPD-3). So a single panic while another thread held a broker-resident lock POISONS it permanently: the next `.lock().unwrap()` panics too, kills its per-conn reply thread, and EVERY subsequent attach silently deadlines (`spt rc` → \"brain IPC read deadline elapsed\") while non-locked ops keep working. The effect journal (bug #16, `[REQ-HAZARD-EFFECT-JOURNAL-PTY-WEDGE]`) and the inject floor (`[REQ-HAZARD-INJECT-WORKER-POISON]`) were fixed one lock at a time; the SURVIVING class was the attach-path locks — the session map (`sessions` ×18 + its `sessions_exit` alias), the per-session `OutputLog` ring (×11), and `pair_holds` (×4) — all bare `.lock().unwrap()`.\r\n- **Invariant:** NO broker-resident lock reachable from serve/dispatch may use a bare `.lock().unwrap()`. Recover via the shared `recover` (`into_inner` — safe for the short, coherent-on-recovery map ops of `sessions`/`pair_holds`) or, for the `OutputLog` ring, `recover_log` — which adds a COHERENCE CLAMP on the recovery path: a panic mid-`append` can leave the ring torn (over-cap, a last seq not below `next_seq`, non-monotonic), and serving those bytes risks garbage, so `clamp_or_reset` cheap-checks the invariants and RESETS the ring empty (`next_seq` preserved — cursors never rewind) + loud-logs on violation. Fail-fast on the log would REINTRODUCE the wedge; blind recover would serve torn bytes; clamp-or-reset costs only scrollback that self-heals on the next PTY output + repaint. Any NEW broker-resident lock must use `recover`/`recover_log` or carry a documented fail-fast justification.\r\n- **spt-core mapping:** `broker::recover` / `broker::recover_log` / `OutputLog::clamp_or_reset`. unit = a poisoned session-map recovers to a usable guard (the next attach still opens), a torn ring clamps/resets-empty only when torn (coherent untouched, `next_seq` preserved), and `recover_log` wires the clamp onto the poison branch — each RED-first against a bare `.lock().unwrap()` / an un-clamped `into_inner`. int (scripted panic-under-lock during concurrent attach → next attach opens) deferred to the two-host rig.\r\n- **Source:** REMOTE-TRUTH B-1 (pivoted — the triage's named floor sites were already closed by `[REQ-HAZARD-INJECT-WORKER-POISON]`; the surviving attach-path lock class is the real root), doyle B-1 ruling (recover all three via one helper + the OutputLog ring coherence clamp).\r\n\r\n### 7.34 A dead `rec.pid` on an spt-hosted perch is EXPECTED — no reader may alive-gate on `rec.pid` alone  `[REQ-HAZARD-DEAD-REC-PID]`\r\n<!-- [doc->REQ-HAZARD-DEAD-REC-PID] -->\r\n- **Failure (paid-for, field evidence — the F-026 #11 dead-pid class):** an spt-hosted endpoint's `rec.pid` records the ephemeral bind-CLI pid, which dies IMMEDIATELY after bind (the broker holds the PTY; there is no resident harness process at that pid). Readers that alive-gated on `rec.pid` alone treated the perch as stale: self-detect leg (c) could NEVER resolve an spt-hosted sender (its from-stamp degraded to `cli@NODE`, operator #7, and replies bounced `NO_PERCH` — sighted on hall-bf/ball-b, v0.24.0).\r\n- **Invariant:** a dead `rec.pid` on an spt-hosted perch is an EXPECTED state, not staleness. NO reader may alive-gate on `rec.pid` alone: spt-hosted LIVENESS comes from the daemon-managed `status` field (KH 2.5 — status present ⇒ authoritative, never a per-pid probe); IDENTITY comes from session/ancestry resolution, where `rec.parent_pid` (the harness pid, the stable session-binding anchor) is the ancestry candidate (ADR-0021 seed-hint discipline: pid is a bind-time SEED, not a truth anchor — re-anchoring truth in `rec.pid` is overruled by design). Any newly sighted `rec.pid`-alive-gating reader gets the same scoped fix and extends THIS requirement's evidence — no new REQ per reader.\r\n- **spt-core mapping:** `roster.rs` `detect_self_by_ancestry` (the first sighted reader, fixed by `[REQ-SELF-DETECT-PARENT-PID]`: `parent_pid` candidates alongside `rec.pid`). int = the E-1 red-first (`detect_self_resolves_spt_hosted_perch_via_parent_pid`: dead `rec.pid` + live-ancestor `parent_pid` resolves), dual-tagged as this class's test.\r\n- **Source:** REMOTE-TRUTH E-1 rider (doyle ruling 2026-07-05: scoped reader-side (b) over re-stamping `rec.pid` (a) — ADR-0021 ground truth, migration hole, blast radius).\r\n\r\n### 7.35 The cached ceremony-clock NTP offset must NOT survive an OS clock STEP — an offset measured against the pre-step clock strands every pairing for the TTL  `[REQ-HAZARD-CEREMONY-CLOCK-STEP]`\r\n- **Failure (paid-for, field evidence — the enlyzeam BIGNET-join RCA 2026-07-06):** enlyzeam's `w32time` was STOPPED, so its system clock drifted +3m30s; the pairing ceremony offset (REQ-PAIR-8, `ntp.rs`) correctly corrected the TOTP clock — until an operator `w32tm /resync` STEPPED the OS clock back −210s under the LIVE daemon. The cached offset had been measured against the OLD (drifted) clock and its snapshot was monotonic-only (`when: Option<Instant>`), so the step was invisible to the cache: for up to the 15-min TTL the ceremony clock stayed ~7 TOTP steps in the past, every `subnet join` returned NO_SEED_HOLDER while fresh-process probes met in <500ms, and only a daemon bounce (which forced a fresh query) healed it. Timeline-proven: refresh cadence 15:08/15:23/15:38/15:53; resync-step 15:45:46; joins failed 15:46–15:52; bounce healed instantly.\r\n- **Invariant:** the offset cache MUST detect an OS clock step and re-query, never apply a stale offset across it. The snapshot stores an `(Instant, SystemTime)` PAIR; on every read, if the wall clock diverged from the monotonic clock since the snapshot beyond a small tolerance (`|wall_elapsed − mono_elapsed| > ~2s`) the OS clock STEPPED ⇒ force an immediate re-query (offset recomputed against the new clock) regardless of TTL. Belt-and-braces at the join edge: `meet_seed_holder`, on search-deadline exhaustion, forces ONE fresh NTP query and runs one final sweep at the corrected step before the NO_SEED_HOLDER verdict — so a stepped-clock join self-heals WITHOUT a daemon bounce. The step tolerance sits far below one TOTP step (30 s) and well above scheduling jitter. Contract of REQ-PAIR-8 (lazy TTL cache, system-clock fallback, never sets the OS clock) is otherwise unchanged.\r\n- **spt-core mapping:** `spt_net::net::pairing::ntp` — `CachedOffset{when, wall_at}` pair + `refresh_needed`/`clock_stepped`/`signed_secs` (step-detect) + `force_refresh`/`invalidate` (the meet-exhaustion re-query seam); `spt_daemon::pairhost::meet_seed_holder` via `sweep_then_final_retry` (force-refresh + one final sweep on exhaustion). unit = `clock_step_forces_refresh_via_injected_reads` (injected mono/wall reads: a wall step forward OR back with mono barely moved re-queries; agreeing clocks within TTL do not; TTL-expiry and never-queried also refresh — no 15-min sleep) + `invalidate_clears_the_snapshot` + `meet_retries_once_after_refresh_on_exhaustion`/`meet_final_retry_failing_surrenders_once` (exactly one post-refresh retry, then surrender with the richer error).\r\n- **Source:** JOIN-TRUTH W1/D3 (doyle /diagnose enlyzeam RCA; the offset defeated by a live clock step is the D3 of four defects, D1 multihome + D2 loud-fail + D4 verbose-clock the siblings).\r\n\r\n### 7.36 The broker control plane and PTY fan-out must NEVER block on a single subscriber connection — a suspended brain conn must not wedge control  `[REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]`\r\n<!-- [doc->REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE] -->\r\n- **Failure (paid-for, rig-CONFIRMED 2026-07-06/07 — `NtSuspendProcess` on the brain, no update involved):** a controller's writer thread does a BLOCKING socket write to its brain subscriber conn. When that brain is suspended (or black-holed) the write never returns. The output-driven eviction path (`append` → `mark_controller_gone`, bounded by 7.12's `CONTROLLER_WRITE_DEADLINE`) only fires on NEW output, and `reap_dead_controller` (7.29 companion) only caught a writer that had EXITED — so a controller BLOCKED (not exited) on an IDLE session was evicted by NEITHER. Within seconds: attached output froze, `detach` did not release the control stamp (release routed through the wedged conn), reattach was REFUSED (`controlled-by` latched), `rc --take` hung, and `daemon status` still read healthy. Every brain cycle — including every `update apply` — has such a freeze window; a stalled or slow-draining NEW brain (6 sessions + psyches + WAN on the incident night) makes it a PERMANENT wedge until a bounce. `brain.ready ≠ subscribers drained` — why the 22:47 apply \"promoted\" while frozen.\r\n- **Invariant:** no broker client ever observes a wedge because the BRAIN stalled. (1) The controller's writer publishes its IN-FLIGHT-write window (`write_blocked_since`: `Some(Instant)` set immediately before the blocking `write_frame`, `None` after — the mutex never held ACROSS the write); a reader sees a `Some` older than `BRAIN_WRITE_DEADLINE` (15 s = 3× the 7.12 controller bound, measured full-with-**zero writer progress**, so a slow-but-DRAINING brain never trips it) as a WEDGED writer and stall-evicts it. Eviction is triggered TIME-based, not output-driven: at the take/reattach path (`resolve_subscribe`, so control self-heals against broker truth even on an idle session) AND at the `KIND_SESSIONS` reap. (2) The stamp release is BROKER-side — `clear_controller` re-stamps `driven_by`/`controlled` with no brain round-trip, so `controlled-by` can never latch behind a dead transit; `rc --take` and reattach-refuse consult live `has_controller()` truth. (3) The evict is OBSERVABLE, never a silent absence — a per-evict broker-log line plus a broker-global tally (`count` + `last_ms`) surfaced on `daemon status` (a brain that never returns is then a diagnosable fact). Idle ≠ wedged: an idle controller (writer parked on `rx.recv()`, `write_blocked_since` `None`) is NEVER evicted; only a genuinely-blocked writer is. Complements 7.12 (which bounds a merely SLOW controller's Full channel) and 7.29 (which converges stamps to broker truth) — this covers the BLOCKED-conn case both leave open. Deferred (Q2, minimal-plus): the WHOLESALE take/release verb-set migration off the brain stays out; the suspend-brain int rig is the arbiter — if it passes with the stall-evict shape, the wholesale move is scope creep. **Residual — BOUNDED by W2, CLOSED by W3 (`REQ-UPDATE-PROMOTE-DRAINED`):** the stall-evict shrank the update-apply FALSE-PROMOTE window from unbounded to `BRAIN_WRITE_DEADLINE` (15 s) — `brain.ready` could still arrive inside that window while blocked writes pended on the OLD-generation subscriber conn (`brain.ready ≠ subscribers drained`, the 22:47 apply). W3 CLOSES it: the ADR-0018 brain-trial promotion gate (`brainproc::run_trial`) now LATCHES ready-seen and promotes only on ready **AND** `TrialEnv::old_gen_drained()` — a broker-truth read (`Broker::any_local_controller_wedged`, the brain's own `by:None` conn wedged past the deadline; a remote conn is excluded, it is the W2 concern not an old-gen brain conn), no brain round-trip. Ready-but-never-drained elapses the trial window → kill + rollback (conservative; never a false-promote onto a wedge). The false-promote int rig (`tests/false_promote.rs`) exercises the promotion path itself, RED-first.\r\n- **spt-core mapping:** `broker.rs` — `ControllerSink::write_blocked_since` + `controller_writer` (marks the in-flight window around both the initial-batch and live-loop writes) + `controller_write_stalled` (pure predicate over injected `now`/`deadline`) + `stall_evict_controller` (Inline stamp at `resolve_subscribe`, Deferred stamp at `reap_dead_controller` so the reap closure does no I/O under the shared sessions lock) + `record_stall_evict`/`stall_evict_stats` (the tally) + `dispatch_stall_evicts` (`KIND_STALL_EVICTS` IPC); `brain.rs` `Brain::stall_evicts`; `cli.rs` `render_stall_evict_line` (daemon-status surface, public wording). unit = `controller_write_stall_predicate_distinguishes_wedged_from_idle` + `stalled_incumbent_is_evicted_on_reattach_and_control_released` + `reap_evicts_a_blocked_writer_not_only_an_exited_one` + `stall_evict_line_surfaces_only_a_real_tally`. int = `brain_decouple.rs` `suspended_brain_controller_is_stall_evicted_take_completes_viewer_ticks` — the deterministic in-process ANALOG of a suspended brain (a REMOTE controller that stops reading parks the broker writer in `write_frame`, the same primitive `NtSuspendProcess` produces), asserting EXACTLY four things: (1) an attached viewer keeps ticking during the wedge, (2) a second operator's `Control` subscribe COMPLETES via stall-evict (RED-first: remove the `resolve_subscribe` evict → `BusyControlled` + `evicts=0`), (3) it resolves promptly (well under the watchdog), (4) the evict is tallied (observability); Linux leg on kitsubito. It does NOT re-assert resume + cursor-replay byte-identity — W2's delta does not touch the replay path, so output integrity stays covered by the pre-existing ring/cursor units (effect-journal replay dedup, the `delivered_through` no-rewind CAS `advanced_cursor`, clean-repaint cold-attach). The live `NtSuspendProcess`/`SIGSTOP` run on a real brain child remains the MANUAL field rig from the incident night, not a CI test.\r\n- **Source:** LIFECYCLE-TRUTH W2 (doyle rig `NtSuspendProcess` root + design ruling `docs/W2-DESIGN-RULING.md` @60ec0ed — minimal-plus scope, `BRAIN_WRITE_DEADLINE` 15 s, passive resubscribe + observability rider; todlando build 2026-07-07). The v0.27→v0.28 update-wedge night's flagship defect.\r\n\r\n### 7.37 The Layer-1 settle-gate must RE-ARM per delivery on an observable PTY — a mid-session reader reattach (`/clear`) re-creates the head-swallow race  `[REQ-HAZARD-INJECT-SETTLE-REARM]`\r\n<!-- [doc->REQ-HAZARD-INJECT-SETTLE-REARM] -->\r\n- **Failure (paid-for, field-CONFIRMED on 0.29.0 — doyle diagnosis + perri screenshot `WindowsTerminal_6iSjya8pMt.png`):** the shipped W5-A settle-gate (7-of-W5, `REQ-INJECT-MULTILINE-INTEGRITY`) gated Layer 1 behind a worker-local ONE-SHOT (`settled_once`) on the premise that the head-swallow race is a STARTUP condition (the input reader not yet attached after spawn). That premise is FALSE: a mid-session `/clear` re-enters the harness's raw-mode input reader, re-attaching it and RE-CREATING the pre-settle window — but the one-shot had already fired at spawn, so `settle_before_inject` was SKIPPED and the head was eaten again. A checkpoint-wake payload injected right after `/clear` arrived head-truncated (mid-path `spt/Cargo.toml)`); echo-verify (Layer 2) is a default-OFF declared capability (`SPT_INJECT_VERIFY_ECHO`) for that session, so the loss was silent AND unrecoverable (a live-SENT inject leaves no spool copy).\r\n- **Invariant:** the settle-gate re-arms before EVERY delivery on an OBSERVABLE (echoing/interactive) PTY — the bug-prone class, where re-settling is cheap (a tick or two once the reader answers) and necessary (a mid-session reader reattach must be re-confirmed). The steady-state settle is latch-skipped ONLY where the probe is known UNOBSERVABLE on this PTY (a non-echoing ConPTY whose DSR answer never surfaces): there is no reader-reattach race to guard, and each settle would burn the full `INJECT_SETTLE` deadline. The class is discriminated by the settle's OWN return value — `settle_before_inject` returns `true` iff the session ring advanced (probe observed) — latched on the FIRST attempt only; a RE-DRIVE (`attempt > 1`, reached only on a swallowed head) ALWAYS settles regardless of the latch. This preserves the old one-shot's sole legitimate purpose (sparing a non-echoing ConPTY the deadline on every delivery) while closing the mid-session recurrence for echoing PTYs.\r\n- **spt-core mapping:** `broker.rs` `run_inject_worker` — the `settled_once: bool` one-shot is replaced by a `probe_unobservable: bool` latch (init `false` = assume observable → settle first delivery); the gate is the pure `should_settle(attempt, probe_unobservable) = attempt > 1 || !probe_unobservable`, and the first-attempt settle's `observed` return sets `probe_unobservable = !observed`. `settle_before_inject` gains a `bool` return (observed vs timed-out). unit = `should_settle_rearms_on_observable_pty` (observable PTY re-settles each delivery; unobservable-probe PTY skips the steady-state settle; a re-drive settles in both classes — RED-first: restore the one-shot and the observable-PTY assert flips). No int edit: the `inject_control_wedge.rs` mocks (`findstr`/`cat`) never answer the DSR probe → the settle times out → `probe_unobservable` latches true after delivery 1 → behaviour is IDENTICAL to the old one-shot for those non-echoing mocks; only a genuinely echoing PTY changes (it re-settles), which is the fix. Rides `REQ-INJECT-MULTILINE-INTEGRITY`'s existing int coverage.\r\n- **Source:** post-0.29.0 field fix (doyle diagnosis + design ruling `BUILD-SPEC-inject-settle-rearm`; todlando build 2026-07-08). An impl bug in the shipped W5-A Layer-1 fix, not a new invariant class — the settle-gate was correct but under-armed.\r\n\r\n### 7.38 EVERY write on a physical broker connection is bounded + cancelable + poison-on-failure — no writer holds the send gate across an UNBOUNDED OS write  `[REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]`\r\n<!-- [doc->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK] -->\r\n- **Failure (paid-for, field capture + deterministic Windows repro 2026-07-09 — hertz RCA, UPDATE-WEDGE round-4):** `controller_writer` held the `SharedSend = Arc<Mutex<SendHalf>>` guard ACROSS a blocking `write_frame`. On Windows, `interprocess 2.4.2` routes the send to `WriteFileEx` + `SleepEx(INFINITE, alertable)` — this local-socket send path exposes **no supported write timeout** (`set_timeout` → `Unsupported`). An `rc --take` controller consumer that stopped reading blocked the write **indefinitely** (~127.95 s in the capture, released only when a brain restart tore the connections down). The 7.36 logical `stall_evict_controller` removed the controller ROLE but neither canceled the in-flight pipe write, closed the physical connection, nor called `CancelIoEx` — the detached writer kept its `SharedSend` clone and live stack-owned mutex guard, and every other write queued at that gate wedged behind it. Load-gated: needs a real `seq>0` frame plus a non-draining consumer (seq-0 boot conns and a quiescent update are clean — why v0.30.5's controlled apply passed). `CTRL_WRITE_LOCKED wait_us=0` on all four capture writers exonerates the mutex convoy: the block is INSIDE the OS write, after lock acquisition. Neither existing deadline bounds it: `CONTROLLER_WRITE_DEADLINE` is output-driven channel-full handling (a writer blocked on its first owned frame never fills the channel); `BRAIN_WRITE_DEADLINE` is an age predicate sampled opportunistically, not an I/O timer.\r\n- **Invariant:** every write on a physical broker connection rides ONE broker-owned bounded/cancelable framed-write primitive with an **independent out-of-band aborter**. The `SendHalf` never leaves the conn object (a bypass cannot compile). An ABSOLUTE deadline stamped at write entry covers BOTH the serialized gate-wait AND OS write completion; a per-conn watchdog fires at that deadline and aborts the in-flight op — never relying on the write returning or on opportunistic stall-evict sampling. On deadline / partial write / cancellation / unknown completion: (1) poison the whole physical conn idempotently, (2) abort read+write so `handle_conn` reaches existing EOF cleanup, (3) wait for the canceled op to report completion before releasing its buffer, (4) NEVER reuse the conn (a timed-out length-prefixed frame may be partially written — reuse would corrupt framing), (5) finish/join the retired writer before reporting physical cleanup. Controller (replay + live), viewer, dispatch-reply, and nethost stream/presence writes ALL route through the primitive — any raw unbounded write left behind the gate preserves the failure class. NO additional output queue (the bounded queue + isolated writer already exist; the block is BELOW them); NO `PIPE_NOWAIT` (recorded mid-frame corruption risk). The logical stall-evict (7.36) stays as the ROLE-release trigger; physical retirement authority is the bounded write itself. Unix keeps its existing semantics under the same poison/retire invariant.\r\n- **spt-core mapping:** `spt-daemon/src/conn.rs` `BrokerConn` (Condvar-gated serialized write gate + poisoned state + per-in-flight op id + watchdog + `abort_physical()`: cfg(windows) `CancelIoEx` → completion handshake → `DisconnectNamedPipe` raw kernel32 externs; cfg(unix) `UnixStream::shutdown(Both)`); `broker.rs` `SharedSend = Arc<BrokerConn>` + `controller_writer`/`viewer_writer`/`send_frame`/`send_error` routed; `nethost.rs` stream-log/presence sends routed. int = `brain_decouple.rs` `non_draining_controller_stall_evict_releases_writer_and_connection` (`#[cfg(windows)]`, the hertz 9-step: real PTY burst, black-holed remote controller, exactly one logical stall-evict via an independent draining conn, the black-holed conn's `handle_conn` finishes ≤2 s of logical release WITHOUT dropping the client, a write on the old client conn fails, a fresh controller resumes from the frozen cursor) — RED-first pre-fix; the cross-platform companion `suspended_brain_controller_is_stall_evicted_take_completes_viewer_ticks` stays green on both OSes.\r\n- **Source:** UPDATE-WEDGE round-4 (hertz root doc `docs/UPDATE-WEDGE-ROOT-CAUSE.md` @f8596ca; doyle gate ruling `docs/UPDATE-WEDGE-RCA-DISPATCH-todlando.md`; todlando build 2026-07-09). The residual 7.36 left open: logical eviction without physical retirement.\r","startLine":703,"lineNumbers":[703,704,705,706,707,708,709,710,711,712,713,714,715,716,717,718,719,720,721,722,723,724,725,726,727,728,729,730,731,732,733,734,735,736,737,738,739,740,741,742]},"fileSize":356704,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\docs\\KNOWN-HAZARDS.md"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":1092,"totalBytes":22158,"outputLines":40,"outputBytes":22158,"shownRange":{"start":703,"end":742},"nextOffset":743}}},"isError":false,"timestamp":1788999555001}}
{"type":"message","id":"b5dafb40","parentId":"ad758859","timestamp":"2026-09-10T00:19:15.006Z","message":{"role":"toolResult","toolCallId":"call_Dfwisf4jthMIiGi9y0lAsM4N|fc_0dde72effda0ad21016aa1f782e3e487d2aab46004816c44a9","toolName":"read","content":[{"type":"text","text":"pub struct Brain {\n…\n    /// continuity state a cold-started brain rebuilds from the broker via\n    /// [`Brain::resume_sessions`], replacing the retired single-session\n    /// `BrainState` handoff. A session present here is in **resume mode** — its\n    /// `KIND_OUTPUT` dedup is **dedup-below + snap-above, NO reject-gap** (seeded\n    /// at the broker's `resume_seq`): drop `seq < cursor` (the at-least-once\n    /// boundary re-send) and accept-and-snap `seq >= cursor` (the contiguous case\n    /// *and* a post-eviction ring-floor jump — the broker's `Mutex`-held replay\n    /// cannot reorder, so a forward jump is only ever a legitimate eviction clamp,\n    /// never a real gap). A session **absent** here uses the strict single-session\n    /// `next_seq` path (reject-gap) byte-for-byte — the legacy spawn/drive seat.\n    session_cursors: std::collections::HashMap<u64, u64>,\n    /// When `Some`, every reply read is bounded by this per-call deadline budget\n    /// — set only by [`Brain::cold_start_pump`], which also splits the carrier\n    /// into the reader-thread form so the deadline can be honored over a channel\n    /// (`recv_timeout`). `None` (default) = the ordinary blocking client — a\n    /// brain-IPC read waits forever, which is correct everywhere EXCEPT the\n    /// single-threaded pump, where a black-holed peer would wedge the whole loop\n    /// (KNOWN-HAZARDS — a blocked brain-IPC read must never wedge the pump). The\n    /// two facts are coupled at construction: `Some` ⇔ [`BrainConn::Split`].\n    io_timeout: Option<Duration>,\n    /// #6 (ADR-0031): a COLD attach (`from_seq == 0`, set by [`Brain::attach`] /\n    /// [`Brain::attach_as`]) receives a synthesized clean-screen REPAINT as its\n    /// first `KIND_OUTPUT` frame, at pseudo-seq `broker.next_seq - 1` — NOT seq 0.\n    /// When `true`, the strict legacy `next_seq` path adopts that first frame's seq\n    /// as the baseline (accept-and-seed) instead of reject-gapping it, then reverts\n    /// to strict reject-gap for every subsequent LIVE frame (so a b4 drop still\n    /// re-fetches — B2 live exactly-once preserved). Only ever consulted on the\n    /// legacy single-session path; a resume-mode brain (snap-above) ignores it. A\n    /// resume-from-floor re-fetch (`from_seq > 0`) leaves this `false` → raw-ring\n    /// replay + strict reject-gap, entirely unchanged.\n    baseline_next_output: bool,\n}\n\nimpl Brain {\n    /// Cold start: connect, handshake, fresh generation 0, `gen_start = now`.\n    /// The ordinary blocking carrier ([`BrainConn::Whole`]) — every client\n    /// except the peer pump, which uses [`Brain::cold_start_pump`].\n    pub fn cold_start(name: &str, now_ms: u64) -> io::Result<Self> {\n        let mut conn = LocalSocketTransport::connect(name)?;\n        send_hello(&mut conn, Role::Brain)?;\n        Ok(Brain {\n            conn: BrainConn::Whole(conn),\n            generation: 0,\n            gen_start_ms: now_ms,\n            session_id: None,\n            next_seq: 0,\n            net_cursors: std::collections::HashMap::new(),\n            presence_cursor: None,\n            session_cursors: std::collections::HashMap::new(),\n            io_timeout: None,\n            baseline_next_output: false,\n        })\n    }\n\n    /// Cold start in **PUMP mode**: connect, handshake, then SPLIT the stream\n    /// into the reader-thread carrier ([`BrainConn::Split`]) and arm a per-IPC-\n    /// call deadline of `io_timeout`. This is the peer pump's dedicated client —\n    /// every reply read is bounded ([`Brain::read_event_until`] via\n    /// [`Brain::call_deadline`]) so a black-holed peer can never wedge the\n    /// single-threaded pump. The split is done at CONSTRUCTION (not flipped on a\n    /// live `Whole` after the fact — moving the `Stream` out by value to split it\n    /// is impossible behind `&mut self`), so a pump brain is split-and-bounded\n    /// for its whole life and a non-pump brain never pays the reader thread.\n    /// `trace` picks the reader's lifecycle-marker sink (REQ-RC-DISPLAY-SOLE-\n    /// WRITER): the daemon's peer pump passes [`PumpTrace::Stderr`] (its stderr\n    /// is the log), interactive clients pass [`PumpTrace::from_env`] (silent\n    /// unless a rig opts in — a client's stderr is a terminal it must not\n    /// write to from a background thread).\n    // [impl->REQ-HAZARD-PUMP-IPC-DEADLINE]\n    // [impl->REQ-RC-DISPLAY-SOLE-WRITER]\n    pub fn cold_start_pump(\n…\n    ) -> io::Result<Self> {\n…\n    }\n…\n                KIND_OUTPUT => {\n…\n                    let bytes = decode_bytes(&ev.data_b64)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Output {\n                        session_id: ev.session_id,\n                        seq: ev.seq,\n                        bytes,\n                    });\n                }\n                KIND_SPAWN_CONFLICT => {\n                    let ev: SpawnConflict = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::SpawnConflict {\n                        endpoint: ev.endpoint,\n                        session_id: ev.session_id,\n                    });\n                }\n                // [impl->REQ-ER-BRINGUP-SPAWNS-SESSION]\n                KIND_BROUGHT_UP => {\n                    let ev: BroughtUpReply = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::BroughtUp {\n                        outcome: ev.outcome,\n                        detail: ev.detail,\n                        session_id: ev.session_id,\n                        admit: ev.admit,\n                    });\n                }\n                // [impl->REQ-SEAL-CEREMONY-TOTP]\n                KIND_SEAL_CEREMONY_REPLY => {\n                    let ev: SealCeremonyReply = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::SealCeremonyDone {\n                        outcome: ev.outcome,\n                        detail: ev.detail,\n                        token: ev.token,\n                    });\n                }\n                // [impl->REQ-SEAL-CEREMONY-RC-CLIENT]\n                KIND_SEAL_CEREMONY_OPEN => {\n                    let ev: SealCeremonyOpenEvent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::SealCeremonyOpen {\n                        session_id: ev.session_id,\n                        ceremony_id: ev.ceremony_id,\n                        content_b64: ev.content_b64,\n                        subnet: ev.subnet,\n                        destination: ev.destination,\n                        payload_to_sign_b64: ev.payload_to_sign_b64,\n                        fido2_node: ev.fido2_node,\n                        offer_enroll: ev.offer_enroll,\n                    });\n                }\n                KIND_SEAL_CEREMONY_RESULT => {\n                    let ev: SealCeremonyResultEvent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::SealCeremonyResult {\n                        session_id: ev.session_id,\n                        ceremony_id: ev.ceremony_id,\n                        outcome: ev.outcome,\n                        detail: ev.detail,\n                    });\n                }\n                KIND_SPAWNED => {\n                    let ev: Spawned = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Spawned {\n                        session_id: ev.session_id,\n                        pid: ev.pid,\n                    });\n                }\n                KIND_EXIT => {\n                    let ev: ExitEvent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Exit {\n                        session_id: ev.session_id,\n                        code: ev.code,\n                    });\n                }\n                KIND_SUBSCRIBED => {\n                    let ev: SubscribedReply = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Subscribed {\n                        session_id: ev.session_id,\n                        outcome: ev.outcome,\n                        note: ev.note,\n                    });\n                }\n                KIND_SIZE => {\n                    let ev: SizeEvent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Size {\n                        session_id: ev.session_id,\n                        rows: ev.rows,\n                        cols: ev.cols,\n                    });\n                }\n                KIND_DISPLACED => {\n                    let ev: DisplacedEvent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Displaced {\n                        session_id: ev.session_id,\n                        by: ev.by,\n                    });\n                }\n                // [impl->REQ-VIEWER-SKIP-TO-LIVE-ON-EVICT]\n                KIND_VIEWER_EVICTED => {\n                    let ev: ViewerEvictedEvent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::ViewerEvicted {\n                        session_id: ev.session_id,\n                    });\n                }\n                KIND_APPLIED => {\n                    let ev: AppliedEvent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Applied {\n                        session_id: ev.session_id,\n                        op_id: ev.op_id,\n                        applied_now: ev.applied_now,\n                    });\n                }\n                KIND_NET_STATUS_REPLY => {\n                    let ev: NetStatusReply = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::NetStatus(ev));\n                }\n                KIND_NET_DIALED => {\n                    let ev: NetDialed = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::NetDialed(ev));\n                }\n                // The bare ack for a non-blocking submit-dial (PUMP-TRUTH W2) —\n                // null payload; the dial's outcome rides the presence stream.\n                // [impl->REQ-PUMP-PEER-ISOLATION]\n                KIND_NET_DIAL_SUBMITTED => return Ok(BrokerEvent::NetDialSubmitted),\n                KIND_NET_STREAM_OPENED => {\n                    let ev: NetStreamOpened = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::NetStreamOpened(ev));\n                }\n                KIND_NET_SENT => {\n                    let ev: NetSent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::NetSent(ev));\n                }\n                KIND_NET_STREAMS_REPLY => {\n                    let ev: NetStreamsReply = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::NetStreams(ev));\n                }\n                KIND_SESSIONS_REPLY => {\n                    let ev: SessionsReply = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Sessions(ev));\n                }\n                KIND_PAIR_JOINED => {\n                    let ev: PairJoinReply = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::PairJoined(ev));\n                }\n                KIND_MET_MEMBER => {\n                    let ev: MetMember = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::MetMember(ev));\n                }\n                KIND_NET_STREAM_DATA => {\n                    let ev: NetStreamData = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    // Same exactly-once discipline as PTY output, per stream:\n                    // accept the contiguous next seq, drop a replay duplicate,\n                    // reject a forward gap (a lost chunk must not pass silently).\n                    let cursor = self.net_cursors.entry(ev.stream_id).or_insert(ev.seq);\n                    if ev.seq < *cursor {\n                        continue; // replay duplicate — dedup\n                    }\n                    if ev.seq > *cursor {\n                        return Err(io::Error::new(\n…\n                        ));\n                    }\n…\n                }\n…\n    pub fn inject_endpoint(\n…\n    ) -> io::Result<(bool, bool)> {\n…\n        loop {\n            match self.read_event()? {\n                BrokerEvent::Other(env) if env.kind == KIND_ENDPOINT_INJECTED => {\n…\n                    return Ok((reply.delivered, reply.spool_deferred));\n                }\n                BrokerEvent::Error { message } => return Err(io::Error::other(message)),\n                _ => continue,\n            }\n        }\n    }\n\n    /// List the broker's hosted sessions (id + owning endpoint) — the inbound\n    /// dispatcher's session→endpoint resolution (D9-1).\n    pub fn sessions(&mut self) -> io::Result<SessionsReply> {\n        self.send(KIND_SESSIONS, serde_json::Value::Null)?;\n        loop {\n            match self.read_event()? {\n                BrokerEvent::Sessions(ev) => return Ok(ev),\n                BrokerEvent::Error { message } => return Err(io::Error::other(message)),\n                _ => continue,\n            }\n        }\n    }\n\n    /// (Re)subscribe to a stream's buffered reads from `from_seq`, setting the\n    /// local cursor (gapless, exactly-once resume — the PTY `subscribe`\n    /// discipline on a QUIC stream, D4b).\n    // [impl->REQ-DAEMON-2]\n    pub fn net_stream_subscribe(&mut self, stream_id: u64, from_seq: u64) -> io::Result<()> {\n        self.net_cursors.insert(stream_id, from_seq);\n        self.send(\n…\n        )\n    }\n…\n}"}],"details":{"displayContent":{"text":"pub struct Brain {\n…\n    /// continuity state a cold-started brain rebuilds from the broker via\n    /// [`Brain::resume_sessions`], replacing the retired single-session\n    /// `BrainState` handoff. A session present here is in **resume mode** — its\n    /// `KIND_OUTPUT` dedup is **dedup-below + snap-above, NO reject-gap** (seeded\n    /// at the broker's `resume_seq`): drop `seq < cursor` (the at-least-once\n    /// boundary re-send) and accept-and-snap `seq >= cursor` (the contiguous case\n    /// *and* a post-eviction ring-floor jump — the broker's `Mutex`-held replay\n    /// cannot reorder, so a forward jump is only ever a legitimate eviction clamp,\n    /// never a real gap). A session **absent** here uses the strict single-session\n    /// `next_seq` path (reject-gap) byte-for-byte — the legacy spawn/drive seat.\n    session_cursors: std::collections::HashMap<u64, u64>,\n    /// When `Some`, every reply read is bounded by this per-call deadline budget\n    /// — set only by [`Brain::cold_start_pump`], which also splits the carrier\n    /// into the reader-thread form so the deadline can be honored over a channel\n    /// (`recv_timeout`). `None` (default) = the ordinary blocking client — a\n    /// brain-IPC read waits forever, which is correct everywhere EXCEPT the\n    /// single-threaded pump, where a black-holed peer would wedge the whole loop\n    /// (KNOWN-HAZARDS — a blocked brain-IPC read must never wedge the pump). The\n    /// two facts are coupled at construction: `Some` ⇔ [`BrainConn::Split`].\n    io_timeout: Option<Duration>,\n    /// #6 (ADR-0031): a COLD attach (`from_seq == 0`, set by [`Brain::attach`] /\n    /// [`Brain::attach_as`]) receives a synthesized clean-screen REPAINT as its\n    /// first `KIND_OUTPUT` frame, at pseudo-seq `broker.next_seq - 1` — NOT seq 0.\n    /// When `true`, the strict legacy `next_seq` path adopts that first frame's seq\n    /// as the baseline (accept-and-seed) instead of reject-gapping it, then reverts\n    /// to strict reject-gap for every subsequent LIVE frame (so a b4 drop still\n    /// re-fetches — B2 live exactly-once preserved). Only ever consulted on the\n    /// legacy single-session path; a resume-mode brain (snap-above) ignores it. A\n    /// resume-from-floor re-fetch (`from_seq > 0`) leaves this `false` → raw-ring\n    /// replay + strict reject-gap, entirely unchanged.\n    baseline_next_output: bool,\n}\n\nimpl Brain {\n    /// Cold start: connect, handshake, fresh generation 0, `gen_start = now`.\n    /// The ordinary blocking carrier ([`BrainConn::Whole`]) — every client\n    /// except the peer pump, which uses [`Brain::cold_start_pump`].\n    pub fn cold_start(name: &str, now_ms: u64) -> io::Result<Self> {\n        let mut conn = LocalSocketTransport::connect(name)?;\n        send_hello(&mut conn, Role::Brain)?;\n        Ok(Brain {\n            conn: BrainConn::Whole(conn),\n            generation: 0,\n            gen_start_ms: now_ms,\n            session_id: None,\n            next_seq: 0,\n            net_cursors: std::collections::HashMap::new(),\n            presence_cursor: None,\n            session_cursors: std::collections::HashMap::new(),\n            io_timeout: None,\n            baseline_next_output: false,\n        })\n    }\n\n    /// Cold start in **PUMP mode**: connect, handshake, then SPLIT the stream\n    /// into the reader-thread carrier ([`BrainConn::Split`]) and arm a per-IPC-\n    /// call deadline of `io_timeout`. This is the peer pump's dedicated client —\n    /// every reply read is bounded ([`Brain::read_event_until`] via\n    /// [`Brain::call_deadline`]) so a black-holed peer can never wedge the\n    /// single-threaded pump. The split is done at CONSTRUCTION (not flipped on a\n    /// live `Whole` after the fact — moving the `Stream` out by value to split it\n    /// is impossible behind `&mut self`), so a pump brain is split-and-bounded\n    /// for its whole life and a non-pump brain never pays the reader thread.\n    /// `trace` picks the reader's lifecycle-marker sink (REQ-RC-DISPLAY-SOLE-\n    /// WRITER): the daemon's peer pump passes [`PumpTrace::Stderr`] (its stderr\n    /// is the log), interactive clients pass [`PumpTrace::from_env`] (silent\n    /// unless a rig opts in — a client's stderr is a terminal it must not\n    /// write to from a background thread).\n    // [impl->REQ-HAZARD-PUMP-IPC-DEADLINE]\n    // [impl->REQ-RC-DISPLAY-SOLE-WRITER]\n    pub fn cold_start_pump(\n…\n    ) -> io::Result<Self> {\n…\n    }\n…\n                KIND_OUTPUT => {\n…\n                    let bytes = decode_bytes(&ev.data_b64)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Output {\n                        session_id: ev.session_id,\n                        seq: ev.seq,\n                        bytes,\n                    });\n                }\n                KIND_SPAWN_CONFLICT => {\n                    let ev: SpawnConflict = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::SpawnConflict {\n                        endpoint: ev.endpoint,\n                        session_id: ev.session_id,\n                    });\n                }\n                // [impl->REQ-ER-BRINGUP-SPAWNS-SESSION]\n                KIND_BROUGHT_UP => {\n                    let ev: BroughtUpReply = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::BroughtUp {\n                        outcome: ev.outcome,\n                        detail: ev.detail,\n                        session_id: ev.session_id,\n                        admit: ev.admit,\n                    });\n                }\n                // [impl->REQ-SEAL-CEREMONY-TOTP]\n                KIND_SEAL_CEREMONY_REPLY => {\n                    let ev: SealCeremonyReply = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::SealCeremonyDone {\n                        outcome: ev.outcome,\n                        detail: ev.detail,\n                        token: ev.token,\n                    });\n                }\n                // [impl->REQ-SEAL-CEREMONY-RC-CLIENT]\n                KIND_SEAL_CEREMONY_OPEN => {\n                    let ev: SealCeremonyOpenEvent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::SealCeremonyOpen {\n                        session_id: ev.session_id,\n                        ceremony_id: ev.ceremony_id,\n                        content_b64: ev.content_b64,\n                        subnet: ev.subnet,\n                        destination: ev.destination,\n                        payload_to_sign_b64: ev.payload_to_sign_b64,\n                        fido2_node: ev.fido2_node,\n                        offer_enroll: ev.offer_enroll,\n                    });\n                }\n                KIND_SEAL_CEREMONY_RESULT => {\n                    let ev: SealCeremonyResultEvent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::SealCeremonyResult {\n                        session_id: ev.session_id,\n                        ceremony_id: ev.ceremony_id,\n                        outcome: ev.outcome,\n                        detail: ev.detail,\n                    });\n                }\n                KIND_SPAWNED => {\n                    let ev: Spawned = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Spawned {\n                        session_id: ev.session_id,\n                        pid: ev.pid,\n                    });\n                }\n                KIND_EXIT => {\n                    let ev: ExitEvent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Exit {\n                        session_id: ev.session_id,\n                        code: ev.code,\n                    });\n                }\n                KIND_SUBSCRIBED => {\n                    let ev: SubscribedReply = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Subscribed {\n                        session_id: ev.session_id,\n                        outcome: ev.outcome,\n                        note: ev.note,\n                    });\n                }\n                KIND_SIZE => {\n                    let ev: SizeEvent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Size {\n                        session_id: ev.session_id,\n                        rows: ev.rows,\n                        cols: ev.cols,\n                    });\n                }\n                KIND_DISPLACED => {\n                    let ev: DisplacedEvent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Displaced {\n                        session_id: ev.session_id,\n                        by: ev.by,\n                    });\n                }\n                // [impl->REQ-VIEWER-SKIP-TO-LIVE-ON-EVICT]\n                KIND_VIEWER_EVICTED => {\n                    let ev: ViewerEvictedEvent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::ViewerEvicted {\n                        session_id: ev.session_id,\n                    });\n                }\n                KIND_APPLIED => {\n                    let ev: AppliedEvent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Applied {\n                        session_id: ev.session_id,\n                        op_id: ev.op_id,\n                        applied_now: ev.applied_now,\n                    });\n                }\n                KIND_NET_STATUS_REPLY => {\n                    let ev: NetStatusReply = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::NetStatus(ev));\n                }\n                KIND_NET_DIALED => {\n                    let ev: NetDialed = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::NetDialed(ev));\n                }\n                // The bare ack for a non-blocking submit-dial (PUMP-TRUTH W2) —\n                // null payload; the dial's outcome rides the presence stream.\n                // [impl->REQ-PUMP-PEER-ISOLATION]\n                KIND_NET_DIAL_SUBMITTED => return Ok(BrokerEvent::NetDialSubmitted),\n                KIND_NET_STREAM_OPENED => {\n                    let ev: NetStreamOpened = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::NetStreamOpened(ev));\n                }\n                KIND_NET_SENT => {\n                    let ev: NetSent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::NetSent(ev));\n                }\n                KIND_NET_STREAMS_REPLY => {\n                    let ev: NetStreamsReply = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::NetStreams(ev));\n                }\n                KIND_SESSIONS_REPLY => {\n                    let ev: SessionsReply = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Sessions(ev));\n                }\n                KIND_PAIR_JOINED => {\n                    let ev: PairJoinReply = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::PairJoined(ev));\n                }\n                KIND_MET_MEMBER => {\n                    let ev: MetMember = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::MetMember(ev));\n                }\n                KIND_NET_STREAM_DATA => {\n                    let ev: NetStreamData = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    // Same exactly-once discipline as PTY output, per stream:\n                    // accept the contiguous next seq, drop a replay duplicate,\n                    // reject a forward gap (a lost chunk must not pass silently).\n                    let cursor = self.net_cursors.entry(ev.stream_id).or_insert(ev.seq);\n                    if ev.seq < *cursor {\n                        continue; // replay duplicate — dedup\n                    }\n                    if ev.seq > *cursor {\n                        return Err(io::Error::new(\n…\n                        ));\n                    }\n…\n                }\n…\n    pub fn inject_endpoint(\n…\n    ) -> io::Result<(bool, bool)> {\n…\n        loop {\n            match self.read_event()? {\n                BrokerEvent::Other(env) if env.kind == KIND_ENDPOINT_INJECTED => {\n…\n                    return Ok((reply.delivered, reply.spool_deferred));\n                }\n                BrokerEvent::Error { message } => return Err(io::Error::other(message)),\n                _ => continue,\n            }\n        }\n    }\n\n    /// List the broker's hosted sessions (id + owning endpoint) — the inbound\n    /// dispatcher's session→endpoint resolution (D9-1).\n    pub fn sessions(&mut self) -> io::Result<SessionsReply> {\n        self.send(KIND_SESSIONS, serde_json::Value::Null)?;\n        loop {\n            match self.read_event()? {\n                BrokerEvent::Sessions(ev) => return Ok(ev),\n                BrokerEvent::Error { message } => return Err(io::Error::other(message)),\n                _ => continue,\n            }\n        }\n    }\n\n    /// (Re)subscribe to a stream's buffered reads from `from_seq`, setting the\n    /// local cursor (gapless, exactly-once resume — the PTY `subscribe`\n    /// discipline on a QUIC stream, D4b).\n    // [impl->REQ-DAEMON-2]\n    pub fn net_stream_subscribe(&mut self, stream_id: u64, from_seq: u64) -> io::Result<()> {\n        self.net_cursors.insert(stream_id, from_seq);\n        self.send(\n…\n        )\n    }\n…\n}","startLine":391,"lineNumbers":[391,null,410,411,412,413,414,415,416,417,418,419,420,421,422,423,424,425,426,427,428,429,430,431,432,433,434,435,436,437,438,439,440,441,442,443,444,445,446,447,448,449,450,451,452,453,454,455,456,457,458,459,460,461,462,463,464,465,466,467,468,469,470,471,472,473,474,475,476,477,478,479,480,null,485,null,500,null,1040,null,1119,1120,1121,1122,1123,1124,1125,1126,1127,1128,1129,1130,1131,1132,1133,1134,1135,1136,1137,1138,1139,1140,1141,1142,1143,1144,1145,1146,1147,1148,1149,1150,1151,1152,1153,1154,1155,1156,1157,1158,1159,1160,1161,1162,1163,1164,1165,1166,1167,1168,1169,1170,1171,1172,1173,1174,1175,1176,1177,1178,1179,1180,1181,1182,1183,1184,1185,1186,1187,1188,1189,1190,1191,1192,1193,1194,1195,1196,1197,1198,1199,1200,1201,1202,1203,1204,1205,1206,1207,1208,1209,1210,1211,1212,1213,1214,1215,1216,1217,1218,1219,1220,1221,1222,1223,1224,1225,1226,1227,1228,1229,1230,1231,1232,1233,1234,1235,1236,1237,1238,1239,1240,1241,1242,1243,1244,1245,1246,1247,1248,1249,1250,1251,1252,1253,1254,1255,1256,1257,1258,1259,1260,1261,1262,1263,1264,1265,1266,1267,1268,1269,1270,1271,1272,1273,1274,1275,1276,1277,1278,1279,1280,1281,1282,1283,1284,1285,1286,1287,1288,1289,1290,1291,1292,1293,1294,1295,null,1301,1302,null,1311,null,2208,null,2213,null,2223,2224,2225,null,2228,2229,2230,2231,2232,2233,2234,2235,2236,2237,2238,2239,2240,2241,2242,2243,2244,2245,2246,2247,2248,2249,2250,2251,2252,2253,2254,2255,null,2262,2263,null,2498]},"fileSize":160723,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-daemon\\src\\brain.rs"}}},"isError":false,"timestamp":1788999555006}}
{"type":"message","id":"882ab592","parentId":"b5dafb40","timestamp":"2026-09-10T00:19:15.012Z","message":{"role":"toolResult","toolCallId":"call_PSEIBWsFVvBYvWWVwi0ZOOmp|fc_0dde72effda0ad21016aa1f782e3ec87d280d3274df33069a7","toolName":"read","content":[{"type":"text","text":"    /// Drive one brain connection: handshake, then dispatch its command frames\n    /// until it disconnects. On disconnect, detach this connection from any\n    /// sessions it subscribed to — the sessions themselves persist (B2: a brain\n    /// restart must not disturb the hosted children).\n    fn handle_conn(self: &Arc<Self>, conn: Stream) -> io::Result<()> {\n        // RAII handler tally (panic-safe): the round-4 gate's seam for \"the\n        // poisoned conn's handler reached EOF cleanup\" (see conn_handler_count).\n        struct HandlerTally<'a>(&'a AtomicUsize);\n        impl Drop for HandlerTally<'_> {\n            fn drop(&mut self) {\n                self.0.fetch_sub(1, Ordering::AcqRel);\n            }\n        }\n        self.conn_handlers.fetch_add(1, Ordering::AcqRel);\n        let _tally = HandlerTally(&self.conn_handlers);\n\n        let (mut recv, send) = conn.split();\n        // Round-4: the send half is owned by the bounded/cancelable conn\n        // primitive from birth — every write on this physical connection is\n        // deadline-bounded, and a wedge poisons + retires the conn.\n        // [impl->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\n        let send: SharedSend = Arc::new(BrokerConn::new(send, brain_write_deadline()));\n        // Base attribution fact (F-039 leg b): broker clients are brains; role\n        // facts (controller/viewer/presence/stream subscriber) accrete at the\n        // attach sites. [impl->REQ-CONN-POISON-ATTRIBUTION]\n        send.describe(\"role=brain\");\n\n        // Version handshake: this connection must be a brain (2.3).\n        recv_hello(&mut recv, Role::Brain)?;\n\n        let mut my_subs: Vec<u64> = Vec::new();\n        let mut my_stream_subs: Vec<u64> = Vec::new();\n        // Streams THIS conn opened as ConnectionBound (ADR-0040 decision 6):\n        // they die with the conn — the exit sweep FINs each toward its target\n        // and terminal-retires the local row. Durable rows never enter it.\n        // [impl->REQ-STREAM-LEASE-CLASSES]\n        let mut my_cb_streams: Vec<u64> = Vec::new();\n        let mut my_presence_sub = false;\n\n        loop {\n            let env = match read_frame(&mut recv) {\n                Ok(e) => e,\n                // Any read failure (EOF included) ends the connection.\n                Err(_) => break,\n            };\n            match env.kind.as_str() {\n                KIND_SPAWN => match self.dispatch_spawn(env, &send) {\n                    Ok(id) => my_subs.push(id),\n                    Err(msg) => send_error(&send, &msg),\n                },\n                // Fresh/RejectExisting spawn (ADR-0041 decision 5): a typed\n                // conflict reply carries no session — only a real Created\n                // auto-subscribes. TEST-RIG-ONLY valve (the N-1 loud-fail leg):\n                // with SPT_TEST_SPAWN_FRESH_UNSUPPORTED set this broker answers\n                // EXACTLY like an N-1 broker that predates the verb — the\n                // forward-compat unknown-kind reply, byte-identical. Never set\n                // in production.\n                // [impl->REQ-SPAWN-FRESH-TRUTHFUL]\n                KIND_SPAWN_FRESH => {\n                    if std::env::var(\"SPT_TEST_SPAWN_FRESH_UNSUPPORTED\").is_ok() {\n                        send_error(&send, &format!(\"unknown command kind {:?}\", env.kind));\n                    } else {\n                        match self.dispatch_spawn_fresh(env, &send) {\n                            Ok(Some(id)) => my_subs.push(id),\n                            Ok(None) => {} // typed conflict already sent\n                            Err(msg) => send_error(&send, &msg),\n                        }\n                    }\n                }\n                // The engine room's entrance (releases#71). Its own verb rather\n                // than a flag on subscribe, because with nothing hosted there is\n                // nothing to subscribe to — which is precisely why the code the\n                // operator typed used to reach no gate at all.\n                // [impl->REQ-ER-BRINGUP-SPAWNS-SESSION]\n                KIND_BRING_UP => {\n                    if let Err(msg) = self.dispatch_bring_up(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                // The wax-seal mint entrance (WAX-SEAL W2, releases#21). The\n                // reply is DEFERRED — it rides this conn from the ceremony's\n                // end, not from this dispatch.\n                // [impl->REQ-SEAL-CEREMONY-TOTP]\n                KIND_SEAL_CEREMONY => {\n                    if let Err(msg) = self.dispatch_seal_ceremony(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                // The enrollment entrance (SIGNET W1) — same ceremony\n                // machinery, enroll purpose, same deferred reply.\n                // [impl->REQ-SEAL-ENROLL-VERB]\n                KIND_SEAL_ENROLL => {\n                    if let Err(msg) = self.dispatch_seal_enroll(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                // The presented code (or cancel) from the controller serve.\n                KIND_SEAL_CEREMONY_CODE => {\n                    if let Err(msg) = self.dispatch_seal_ceremony_code(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                KIND_SUBSCRIBE => match self.dispatch_subscribe(env, &send) {\n                    Ok(id) => my_subs.push(id),\n                    Err(msg) => send_error(&send, &msg),\n                },\n                KIND_UNSUBSCRIBE => {\n                    if let Err(msg) = self.dispatch_unsubscribe(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                KIND_INPUT => {\n                    if let Err(msg) = self.dispatch_input(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                KIND_RESIZE => {\n                    if let Err(msg) = self.dispatch_resize(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n…\n                KIND_NET_STREAM_OPENER => {\n…\n                    if std::env::var(\"SPT_TEST_STREAM_OPENER_UNSUPPORTED\").is_ok() {\n                        send_error(&send, &format!(\"unknown command kind {:?}\", env.kind));\n                    } else if let Err(msg) = self.dispatch_net_stream_opener(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                KIND_SESSIONS => {\n                    // Clone the row facts out from under the sessions lock, then\n                    // read each log's cursor under the log lock alone — never\n                    // both locks at once (the dispatch_subscribe discipline).\n                    // pid/adapter/age ride along (ADR-0041 decision 6): the ONE\n                    // cycle-verb liveness authority's inputs, additive on the\n                    // reply so the CLI preflight and the shutdown edge consult\n                    // the SAME zombie verdict the spawn gate enforces.\n                    // [impl->REQ-ENDPOINT-CYCLE-HONEST]\n                    type SessRow = (u64, String, Arc<Mutex<OutputLog>>, Option<u32>, String, u64);\n                    let rows: Vec<SessRow> = {\n                        let sessions = recover(&self.sessions);\n                        sessions\n                            .iter()\n                            .map(|(id, h)| {\n                                (\n                                    *id,\n                                    h.endpoint.clone(),\n                                    Arc::clone(&h.log),\n                                    h.session.process_id(),\n                                    h.adapter.clone(),\n                                    h.spawned_at.elapsed().as_millis() as u64,\n                                )\n                            })\n                            .collect()\n                    };\n                    // Snapshot each session's truth under its OWN log lock (cursor +\n                    // controller identity + the control/viewer stamp truth), doing NO\n                    // file I/O under the lock (the lock-across-effect discipline, KH\n                    // 7.12/5.16). The convergence re-assert + the reply then run OFF\n                    // the lock.\n                    struct SessSnap {\n                        id: u64,\n                        endpoint: String,\n                        resume_seq: u64,\n                        controller_by: Option<String>,\n                        has_controller: bool,\n                        viewers: u32,\n                        // The endpoint's stamp generation AT snapshot time\n                        // (ADR-0041 decision 4) — the off-lock converge below\n                        // validates against it so a reap that lands between\n                        // this snapshot and the write refuses the stale stamp.\n                        stamp_gen: u64,\n                        // ADR-0041 decision 6: the zombie-verdict inputs, carried\n                        // to the reply (additive SessionInfo fields).\n                        pid: Option<u32>,\n                        adapter: String,\n                        spawned_ms_ago: u64,\n                    }\n                    let snaps: Vec<SessSnap> = rows\n                        .into_iter()\n                        .map(|(id, endpoint, log, pid, adapter, spawned_ms_ago)| {\n                            let mut log = recover_log(&log);\n                            // B-2 (REQ-CONTROLLER-LIVENESS-REAP): make the controller\n                            // truth HONEST before reading it — reap a dead-writer\n                            // (severed) controller so both the reply's controller_by\n                            // AND the off-lock converge_perch_stamps below see the\n                            // cleared state and the stale info.json stamp clears.\n                            let _ = log.reap_dead_controller();\n                            let stamp_gen =\n                                stamp_slot(&endpoint).gen.load(Ordering::Acquire);\n                            SessSnap {\n                                id,\n                                endpoint,\n                                resume_seq: log.delivered_through(),\n                                controller_by: log.controller_by().map(str::to_string),\n                                has_controller: log.has_controller(),\n                                viewers: log.live_viewer_count(),\n                                stamp_gen,\n                                pid,\n                                adapter,\n                                spawned_ms_ago,\n                            }\n                        })\n                        .collect();\n                    // STAMP CONVERGENCE (REQ-HAZARD-CONTROL-STAMP-CONVERGENCE): re-assert\n                    // each live session's control/viewer stamps to broker truth,\n                    // divergence-gated + OFF the log lock. Closes the stamp-before-bind\n                    // loss for a picker-created endpoint (its spawn stamp was swallowed\n                    // pre-bind); converges within one reconcile-poll window after bind.\n                    // [impl->REQ-HAZARD-CONTROL-STAMP-CONVERGENCE]\n                    // [impl->REQ-CONTROL-STAMP-CONVERGENCE]\n                    for s in &snaps {\n                        let _ = converge_perch_stamps(\n                            &s.endpoint,\n                            s.controller_by.as_deref(),\n                            s.has_controller,\n                            s.viewers,\n                            s.stamp_gen,\n                        );\n                    }\n                    let reply = SessionsReply {\n                        sessions: snaps\n                            .into_iter()\n                            .map(|s| SessionInfo {\n                                session_id: s.id,\n                                endpoint: s.endpoint,\n                                resume_seq: s.resume_seq,\n                                controller_by: s.controller_by,\n                                pid: s.pid,\n                                adapter: s.adapter,\n                                spawned_ms_ago: s.spawned_ms_ago,\n                            })\n                            .collect(),\n                        // The other half of the answer (releases#199): the\n                        // bring-ups this broker has started and not yet\n                        // registered. Read at reply time rather than snapshotted\n                        // with the sessions above, so a bring-up that registered\n                        // mid-handler is reported by the row it now has instead\n                        // of by an entry it no longer needs.\n                        // [impl->REQ-ER-BRINGUP-INFLIGHT-LEDGER]\n                        bringing_up: self.bringups_in_flight_at(Instant::now()),\n                    };\n                    let frame = Envelope::new(\n                        KIND_SESSIONS_REPLY,\n                        serde_json::to_value(reply).expect(\"SessionsReply serializes\"),\n                    );\n                    send_frame(&send, &frame);\n                }\n                KIND_ENDPOINT_INPUT => {\n                    if let Err(msg) = self.dispatch_endpoint_input(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n…\n            }\n        }\n…\n    }"}],"details":{"displayContent":{"text":"    /// Drive one brain connection: handshake, then dispatch its command frames\n    /// until it disconnects. On disconnect, detach this connection from any\n    /// sessions it subscribed to — the sessions themselves persist (B2: a brain\n    /// restart must not disturb the hosted children).\n    fn handle_conn(self: &Arc<Self>, conn: Stream) -> io::Result<()> {\n        // RAII handler tally (panic-safe): the round-4 gate's seam for \"the\n        // poisoned conn's handler reached EOF cleanup\" (see conn_handler_count).\n        struct HandlerTally<'a>(&'a AtomicUsize);\n        impl Drop for HandlerTally<'_> {\n            fn drop(&mut self) {\n                self.0.fetch_sub(1, Ordering::AcqRel);\n            }\n        }\n        self.conn_handlers.fetch_add(1, Ordering::AcqRel);\n        let _tally = HandlerTally(&self.conn_handlers);\n\n        let (mut recv, send) = conn.split();\n        // Round-4: the send half is owned by the bounded/cancelable conn\n        // primitive from birth — every write on this physical connection is\n        // deadline-bounded, and a wedge poisons + retires the conn.\n        // [impl->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\n        let send: SharedSend = Arc::new(BrokerConn::new(send, brain_write_deadline()));\n        // Base attribution fact (F-039 leg b): broker clients are brains; role\n        // facts (controller/viewer/presence/stream subscriber) accrete at the\n        // attach sites. [impl->REQ-CONN-POISON-ATTRIBUTION]\n        send.describe(\"role=brain\");\n\n        // Version handshake: this connection must be a brain (2.3).\n        recv_hello(&mut recv, Role::Brain)?;\n\n        let mut my_subs: Vec<u64> = Vec::new();\n        let mut my_stream_subs: Vec<u64> = Vec::new();\n        // Streams THIS conn opened as ConnectionBound (ADR-0040 decision 6):\n        // they die with the conn — the exit sweep FINs each toward its target\n        // and terminal-retires the local row. Durable rows never enter it.\n        // [impl->REQ-STREAM-LEASE-CLASSES]\n        let mut my_cb_streams: Vec<u64> = Vec::new();\n        let mut my_presence_sub = false;\n\n        loop {\n            let env = match read_frame(&mut recv) {\n                Ok(e) => e,\n                // Any read failure (EOF included) ends the connection.\n                Err(_) => break,\n            };\n            match env.kind.as_str() {\n                KIND_SPAWN => match self.dispatch_spawn(env, &send) {\n                    Ok(id) => my_subs.push(id),\n                    Err(msg) => send_error(&send, &msg),\n                },\n                // Fresh/RejectExisting spawn (ADR-0041 decision 5): a typed\n                // conflict reply carries no session — only a real Created\n                // auto-subscribes. TEST-RIG-ONLY valve (the N-1 loud-fail leg):\n                // with SPT_TEST_SPAWN_FRESH_UNSUPPORTED set this broker answers\n                // EXACTLY like an N-1 broker that predates the verb — the\n                // forward-compat unknown-kind reply, byte-identical. Never set\n                // in production.\n                // [impl->REQ-SPAWN-FRESH-TRUTHFUL]\n                KIND_SPAWN_FRESH => {\n                    if std::env::var(\"SPT_TEST_SPAWN_FRESH_UNSUPPORTED\").is_ok() {\n                        send_error(&send, &format!(\"unknown command kind {:?}\", env.kind));\n                    } else {\n                        match self.dispatch_spawn_fresh(env, &send) {\n                            Ok(Some(id)) => my_subs.push(id),\n                            Ok(None) => {} // typed conflict already sent\n                            Err(msg) => send_error(&send, &msg),\n                        }\n                    }\n                }\n                // The engine room's entrance (releases#71). Its own verb rather\n                // than a flag on subscribe, because with nothing hosted there is\n                // nothing to subscribe to — which is precisely why the code the\n                // operator typed used to reach no gate at all.\n                // [impl->REQ-ER-BRINGUP-SPAWNS-SESSION]\n                KIND_BRING_UP => {\n                    if let Err(msg) = self.dispatch_bring_up(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                // The wax-seal mint entrance (WAX-SEAL W2, releases#21). The\n                // reply is DEFERRED — it rides this conn from the ceremony's\n                // end, not from this dispatch.\n                // [impl->REQ-SEAL-CEREMONY-TOTP]\n                KIND_SEAL_CEREMONY => {\n                    if let Err(msg) = self.dispatch_seal_ceremony(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                // The enrollment entrance (SIGNET W1) — same ceremony\n                // machinery, enroll purpose, same deferred reply.\n                // [impl->REQ-SEAL-ENROLL-VERB]\n                KIND_SEAL_ENROLL => {\n                    if let Err(msg) = self.dispatch_seal_enroll(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                // The presented code (or cancel) from the controller serve.\n                KIND_SEAL_CEREMONY_CODE => {\n                    if let Err(msg) = self.dispatch_seal_ceremony_code(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                KIND_SUBSCRIBE => match self.dispatch_subscribe(env, &send) {\n                    Ok(id) => my_subs.push(id),\n                    Err(msg) => send_error(&send, &msg),\n                },\n                KIND_UNSUBSCRIBE => {\n                    if let Err(msg) = self.dispatch_unsubscribe(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                KIND_INPUT => {\n                    if let Err(msg) = self.dispatch_input(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                KIND_RESIZE => {\n                    if let Err(msg) = self.dispatch_resize(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n…\n                KIND_NET_STREAM_OPENER => {\n…\n                    if std::env::var(\"SPT_TEST_STREAM_OPENER_UNSUPPORTED\").is_ok() {\n                        send_error(&send, &format!(\"unknown command kind {:?}\", env.kind));\n                    } else if let Err(msg) = self.dispatch_net_stream_opener(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                KIND_SESSIONS => {\n                    // Clone the row facts out from under the sessions lock, then\n                    // read each log's cursor under the log lock alone — never\n                    // both locks at once (the dispatch_subscribe discipline).\n                    // pid/adapter/age ride along (ADR-0041 decision 6): the ONE\n                    // cycle-verb liveness authority's inputs, additive on the\n                    // reply so the CLI preflight and the shutdown edge consult\n                    // the SAME zombie verdict the spawn gate enforces.\n                    // [impl->REQ-ENDPOINT-CYCLE-HONEST]\n                    type SessRow = (u64, String, Arc<Mutex<OutputLog>>, Option<u32>, String, u64);\n                    let rows: Vec<SessRow> = {\n                        let sessions = recover(&self.sessions);\n                        sessions\n                            .iter()\n                            .map(|(id, h)| {\n                                (\n                                    *id,\n                                    h.endpoint.clone(),\n                                    Arc::clone(&h.log),\n                                    h.session.process_id(),\n                                    h.adapter.clone(),\n                                    h.spawned_at.elapsed().as_millis() as u64,\n                                )\n                            })\n                            .collect()\n                    };\n                    // Snapshot each session's truth under its OWN log lock (cursor +\n                    // controller identity + the control/viewer stamp truth), doing NO\n                    // file I/O under the lock (the lock-across-effect discipline, KH\n                    // 7.12/5.16). The convergence re-assert + the reply then run OFF\n                    // the lock.\n                    struct SessSnap {\n                        id: u64,\n                        endpoint: String,\n                        resume_seq: u64,\n                        controller_by: Option<String>,\n                        has_controller: bool,\n                        viewers: u32,\n                        // The endpoint's stamp generation AT snapshot time\n                        // (ADR-0041 decision 4) — the off-lock converge below\n                        // validates against it so a reap that lands between\n                        // this snapshot and the write refuses the stale stamp.\n                        stamp_gen: u64,\n                        // ADR-0041 decision 6: the zombie-verdict inputs, carried\n                        // to the reply (additive SessionInfo fields).\n                        pid: Option<u32>,\n                        adapter: String,\n                        spawned_ms_ago: u64,\n                    }\n                    let snaps: Vec<SessSnap> = rows\n                        .into_iter()\n                        .map(|(id, endpoint, log, pid, adapter, spawned_ms_ago)| {\n                            let mut log = recover_log(&log);\n                            // B-2 (REQ-CONTROLLER-LIVENESS-REAP): make the controller\n                            // truth HONEST before reading it — reap a dead-writer\n                            // (severed) controller so both the reply's controller_by\n                            // AND the off-lock converge_perch_stamps below see the\n                            // cleared state and the stale info.json stamp clears.\n                            let _ = log.reap_dead_controller();\n                            let stamp_gen =\n                                stamp_slot(&endpoint).gen.load(Ordering::Acquire);\n                            SessSnap {\n                                id,\n                                endpoint,\n                                resume_seq: log.delivered_through(),\n                                controller_by: log.controller_by().map(str::to_string),\n                                has_controller: log.has_controller(),\n                                viewers: log.live_viewer_count(),\n                                stamp_gen,\n                                pid,\n                                adapter,\n                                spawned_ms_ago,\n                            }\n                        })\n                        .collect();\n                    // STAMP CONVERGENCE (REQ-HAZARD-CONTROL-STAMP-CONVERGENCE): re-assert\n                    // each live session's control/viewer stamps to broker truth,\n                    // divergence-gated + OFF the log lock. Closes the stamp-before-bind\n                    // loss for a picker-created endpoint (its spawn stamp was swallowed\n                    // pre-bind); converges within one reconcile-poll window after bind.\n                    // [impl->REQ-HAZARD-CONTROL-STAMP-CONVERGENCE]\n                    // [impl->REQ-CONTROL-STAMP-CONVERGENCE]\n                    for s in &snaps {\n                        let _ = converge_perch_stamps(\n                            &s.endpoint,\n                            s.controller_by.as_deref(),\n                            s.has_controller,\n                            s.viewers,\n                            s.stamp_gen,\n                        );\n                    }\n                    let reply = SessionsReply {\n                        sessions: snaps\n                            .into_iter()\n                            .map(|s| SessionInfo {\n                                session_id: s.id,\n                                endpoint: s.endpoint,\n                                resume_seq: s.resume_seq,\n                                controller_by: s.controller_by,\n                                pid: s.pid,\n                                adapter: s.adapter,\n                                spawned_ms_ago: s.spawned_ms_ago,\n                            })\n                            .collect(),\n                        // The other half of the answer (releases#199): the\n                        // bring-ups this broker has started and not yet\n                        // registered. Read at reply time rather than snapshotted\n                        // with the sessions above, so a bring-up that registered\n                        // mid-handler is reported by the row it now has instead\n                        // of by an entry it no longer needs.\n                        // [impl->REQ-ER-BRINGUP-INFLIGHT-LEDGER]\n                        bringing_up: self.bringups_in_flight_at(Instant::now()),\n                    };\n                    let frame = Envelope::new(\n                        KIND_SESSIONS_REPLY,\n                        serde_json::to_value(reply).expect(\"SessionsReply serializes\"),\n                    );\n                    send_frame(&send, &frame);\n                }\n                KIND_ENDPOINT_INPUT => {\n                    if let Err(msg) = self.dispatch_endpoint_input(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n…\n            }\n        }\n…\n    }","startLine":5917,"lineNumbers":[5917,5918,5919,5920,5921,5922,5923,5924,5925,5926,5927,5928,5929,5930,5931,5932,5933,5934,5935,5936,5937,5938,5939,5940,5941,5942,5943,5944,5945,5946,5947,5948,5949,5950,5951,5952,5953,5954,5955,5956,5957,5958,5959,5960,5961,5962,5963,5964,5965,5966,5967,5968,5969,5970,5971,5972,5973,5974,5975,5976,5977,5978,5979,5980,5981,5982,5983,5984,5985,5986,5987,5988,5989,5990,5991,5992,5993,5994,5995,5996,5997,5998,5999,6000,6001,6002,6003,6004,6005,6006,6007,6008,6009,6010,6011,6012,6013,6014,6015,6016,6017,6018,6019,6020,6021,6022,6023,6024,6025,6026,6027,6028,6029,6030,6031,6032,6033,6034,6035,6036,6037,null,6119,null,6127,6128,6129,6130,6131,6132,6133,6134,6135,6136,6137,6138,6139,6140,6141,6142,6143,6144,6145,6146,6147,6148,6149,6150,6151,6152,6153,6154,6155,6156,6157,6158,6159,6160,6161,6162,6163,6164,6165,6166,6167,6168,6169,6170,6171,6172,6173,6174,6175,6176,6177,6178,6179,6180,6181,6182,6183,6184,6185,6186,6187,6188,6189,6190,6191,6192,6193,6194,6195,6196,6197,6198,6199,6200,6201,6202,6203,6204,6205,6206,6207,6208,6209,6210,6211,6212,6213,6214,6215,6216,6217,6218,6219,6220,6221,6222,6223,6224,6225,6226,6227,6228,6229,6230,6231,6232,6233,6234,6235,6236,6237,6238,6239,6240,6241,6242,6243,6244,6245,6246,6247,6248,6249,6250,6251,6252,6253,6254,6255,6256,null,6280,6281,null,6335]},"fileSize":844842,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-daemon\\src\\broker.rs"}}},"isError":false,"timestamp":1788999555012}}
{"type":"message","id":"ea57db3b","parentId":"882ab592","timestamp":"2026-09-10T00:19:15.019Z","message":{"role":"toolResult","toolCallId":"call_WJVjoIqzGjDnaFQKCKJ3HFew|fc_0dde72effda0ad21016aa1f782e3fc87d2946409ba6b09f46f","toolName":"grep","content":[{"type":"text","text":"# .worktrees/286-churn/crates/spt-proto/src/\n## emit.rs\n 186|#[macro_export]\n*187|macro_rules! emit_line_err {\n 188|    ($($arg:tt)*) => {{\n 189|        let _ = $crate::emit::write_line(\n 190|            &mut ::std::io::stderr().lock(),\n...\n 196|/// [`emit_block!`] aimed at this process's stderr, on the same terms as\n*197|/// [`emit_line_err!`]: separate name, error discarded.\n 198|// [impl->REQ-EMIT-SINGLE-WRITE]\n 199|#[macro_export]\n 200|macro_rules! emit_block_err {\n...\n 526|        let id = \"endpoint-7\";\n*527|        emit_line_err!(\"EMIT_SELFTEST:{id}: one line\");\n 528|        emit_block_err!(\"EMIT_SELFTEST:{id}: block\\n  second line\");\n 529|    }\n 530|\n\nSkipped missing paths: .worktrees/286-churn/scripts"}],"details":{"scopePath":".worktrees/286-churn/crates/spt-proto/src, .worktrees/286-churn/crates/spt-daemon/tests/conn_blackhole_lifecycle.rs","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":3,"fileCount":1,"files":[".worktrees/286-churn/crates/spt-proto/src/emit.rs"],"fileMatches":[{"path":".worktrees/286-churn/crates/spt-proto/src/emit.rs","count":3}],"truncated":false,"displayContent":"# .worktrees/286-churn/crates/spt-proto/src/\n## emit.rs\n 186│#[macro_export]\n*187│macro_rules! emit_line_err {\n 188│    ($($arg:tt)*) => {{\n 189│        let _ = $crate::emit::write_line(\n 190│            &mut ::std::io::stderr().lock(),\n    │...\n 196│/// [`emit_block!`] aimed at this process's stderr, on the same terms as\n*197│/// [`emit_line_err!`]: separate name, error discarded.\n 198│// [impl->REQ-EMIT-SINGLE-WRITE]\n 199│#[macro_export]\n 200│macro_rules! emit_block_err {\n    │...\n 526│        let id = \"endpoint-7\";\n*527│        emit_line_err!(\"EMIT_SELFTEST:{id}: one line\");\n 528│        emit_block_err!(\"EMIT_SELFTEST:{id}: block\\n  second line\");\n 529│    }\n 530│","missingPaths":[".worktrees/286-churn/scripts"]},"isError":false,"timestamp":1788999555018}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_WoMznqcZMA9QvJyFAj6sJ7Tq|fc_0dde72effda0ad21016aa1f782e40487d29934dbf8378f4a19","toolName":"write","startedAt":"2026-09-10T00:19:15.024Z","args":{"path":"xd://report_issue"},"intent":"Reporting truncated document read"},"id":"d4cae606","parentId":"ea57db3b","timestamp":"2026-09-10T00:19:15.024Z"}
{"type":"message","id":"4565e78a","parentId":"d4cae606","timestamp":"2026-09-10T00:19:15.024Z","message":{"role":"toolResult","toolCallId":"call_1pbBYBrJdatCnkJVzIYe7xBI|fc_0dde72effda0ad21016aa1f782e3f487d2be9c040d88131b67","toolName":"read","content":[{"type":"text","text":"Path '.worktrees/286-churn/crates/spt-proto/src/io.rs' not found"}],"details":{},"isError":true,"timestamp":1788999555019}}
{"type":"message","id":"c7414791","parentId":"4565e78a","timestamp":"2026-09-10T00:19:15.024Z","message":{"role":"toolResult","toolCallId":"call_WoMznqcZMA9QvJyFAj6sJ7Tq|fc_0dde72effda0ad21016aa1f782e40487d29934dbf8378f4a19","toolName":"write","content":[{"type":"text","text":"Noted, thanks!"}],"details":{"xdev":{"tool":"report_issue","mode":"execute","args":{"report":"read: bare .worktrees/286-churn/docs/KNOWN-HAZARDS.md reported only 307 total lines and omitted later sections, while grep finds connection invariants at lines 737–742 in that same file; truncation footer does not expose the remaining true line range."}}},"isError":false,"timestamp":1788999555024}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_c1vQmm2hsWUKzqmKdbl1qo8j|fc_0dde72effda0ad21016aa1f7a007bc87d2836c5242a2a4d007","toolName":"read","startedAt":"2026-09-10T00:19:44.183Z","args":{"path":".worktrees/286-churn/crates/spt-proto/src/emit.rs:1-210"},"intent":"Reading single-write diagnostics"},"id":"1c6a7d9d","parentId":"c7414791","timestamp":"2026-09-10T00:19:44.183Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_ab5g2v3JBCNBrGNwO5njnWls|fc_0dde72effda0ad21016aa1f7a007cc87d29a90ca68b66ba45a","toolName":"read","startedAt":"2026-09-10T00:19:44.183Z","args":{"path":".worktrees/286-churn/crates/spt-daemon/tests/conn_blackhole_lifecycle.rs"},"intent":"Reading connection lifecycle harness"},"id":"c856913c","parentId":"1c6a7d9d","timestamp":"2026-09-10T00:19:44.183Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_0mqcQOcn96wUWhCzdA7Rvd5z|fc_0dde72effda0ad21016aa1f7a007d487d29946f6e7adcfcd4e","toolName":"read","startedAt":"2026-09-10T00:19:44.183Z","args":{"path":".worktrees/286-churn/traceable-reqs.toml:2553-2593,4855-4890:raw"},"intent":"Reading registry amendment context"},"id":"6c41eeb5","parentId":"c856913c","timestamp":"2026-09-10T00:19:44.183Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_H81TcWUzZzTRJSYvgdSphpBV|fc_0dde72effda0ad21016aa1f7a007d887d2919b64622734e05f","toolName":"read","startedAt":"2026-09-10T00:19:44.183Z","args":{"path":".worktrees/286-churn/CHANGELOG.md:1-45:raw"},"intent":"Reading changelog context"},"id":"81deb806","parentId":"6c41eeb5","timestamp":"2026-09-10T00:19:44.183Z"}
{"type":"message","id":"fb832111","parentId":"81deb806","timestamp":"2026-09-10T00:19:44.183Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":"**Checking first-write timing**\n\n**Checking integration coverage**\n\n**Checking log consumers**"},{"type":"toolCall","id":"call_c1vQmm2hsWUKzqmKdbl1qo8j|fc_0dde72effda0ad21016aa1f7a007bc87d2836c5242a2a4d007","name":"read","arguments":{"i":"Reading single-write diagnostics","path":".worktrees/286-churn/crates/spt-proto/src/emit.rs:1-210"},"intent":"Reading single-write diagnostics"},{"type":"toolCall","id":"call_ab5g2v3JBCNBrGNwO5njnWls|fc_0dde72effda0ad21016aa1f7a007cc87d29a90ca68b66ba45a","name":"read","arguments":{"i":"Reading connection lifecycle harness","path":".worktrees/286-churn/crates/spt-daemon/tests/conn_blackhole_lifecycle.rs"},"intent":"Reading connection lifecycle harness"},{"type":"toolCall","id":"call_0mqcQOcn96wUWhCzdA7Rvd5z|fc_0dde72effda0ad21016aa1f7a007d487d29946f6e7adcfcd4e","name":"read","arguments":{"i":"Reading registry amendment context","path":".worktrees/286-churn/traceable-reqs.toml:2553-2593,4855-4890:raw"},"intent":"Reading registry amendment context"},{"type":"toolCall","id":"call_H81TcWUzZzTRJSYvgdSphpBV|fc_0dde72effda0ad21016aa1f7a007d887d2919b64622734e05f","name":"read","arguments":{"i":"Reading changelog context","path":".worktrees/286-churn/CHANGELOG.md:1-45:raw"},"intent":"Reading changelog context"},{"type":"toolCall","id":"call_dBzQK6ug0ipNvJSIUzrDNrk4|fc_0dde72effda0ad21016aa1f7a007e087d28db3fbdc88b7ebc5","name":"read","arguments":{"i":"Reading broker architecture decision","path":".worktrees/286-churn/docs/adr/0004-single-daemon-broker-brain-split-and-self-update.md:1-45:raw"},"intent":"Reading broker architecture decision"},{"type":"toolCall","id":"call_yQ5v4ngdG3LVvysjf5PLaogH|fc_0dde72effda0ad21016aa1f7a007e887d28997dd6d806c2a20","name":"read","arguments":{"i":"Reading client transport handshake","path":".worktrees/286-churn/crates/spt-daemon/src/transport.rs:1-123"},"intent":"Reading client transport handshake"},{"type":"toolCall","id":"call_BkJaSiBdyn99uE5hrUyTuJhm|fc_0dde72effda0ad21016aa1f7a007ec87d28370d031e74a7b61","name":"bash","arguments":{"i":"Reading review commit metadata","command":"git log de5a44bc338d679f3c63f6d0f68826eb9e575d73..19a9097956de0db898eaeb3a7dc4765189ade874 --format=full","cwd":"C:/Users/decid/Documents/projects/spt-core/.worktrees/286-churn"},"intent":"Reading review commit metadata"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":12658,"output":857,"cacheRead":100992,"cacheWrite":0,"totalTokens":114507,"reasoningTokens":487,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1788999555074,"responseId":"resp_0dde72effda0ad21016aa1f783549087d2bf05fde5030ee12a","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_0dde72effda0ad21016aa1f786567887d2a27decd3754b100e","type":"reasoning","content":[],"encrypted_content":"gAAAAABqofegPoDlbXj8AGzmIecBuIuFOQ6F35pbbuGZ16kMAqGNq0EiqKU8K6GYP58lsf2NF4R1lJyqySip186JaBU5KScmTvTd2wg48pPJj-v_kGyIYCMM5PR2iY28BBtGBP40xZDzwU3LSkH-lvjCZWKjqSgqAT86nZTqz3YZSLchxMCSMcPiHn3dhWmrXmkDdyJB35XJ5yc6lNec9AXb8bm5w5dlOz614sj06_X_0qAL4VpRZaBs1pA3GNBOATDtZuWt66FtlMtzTnJUNtgoP6muDDUmkpIV6jkhwizRzjTZw_3uOb3urnFAB0EzoxMBAjvmwhhhK-tGUk9O22yZKBSjvChdGJK-xs8D-C5UFEQcpnOGdGTUprByL2MR0m3VoYq8UI6YqKeYX1DKWtKyGgP5YJyyJxOjCqgq2n7e5w-CLF-LhDjdxwLWmXA7R4lj1UcdOS_KbD7i2S2PXR2j3hKHCkZKNRI0DF6AJdeSIpBckotfk-KpW9vOKb8PX2NfcOWSeAZYfWDCzxllCgJFeqi_0rE0GQ0w4crYyjQB9BpGTeDw7DFfNtJj3WTzupfvLV39eWxajLVhgNEsPmMNcibNrxnqTFdV2tqt10c57EhBzhcnrrGqxiriCaUyDbEHqirhphMl5YL9Rbty3L_XmjOjxPuaIvrxfv2Kr7xdUSXuHs12XL_ikh2ChDJ-vZOlCJ_qZQ-y-S5TitJWD0pkv3VyO7Y0TZvMEOE_kxiyviRCPk51jtp0ukGrmsFC7PpnkiMFXbUYD-m7VhkZD9NMzU6Skqk8mJd8l9IznveTwhs_MzR58kY2YnJpYfew7hhdlT3xoD-qMW6GHs3gN6euK06WjSbxiDd5awMBkA4ePKd_76QwcdYuo9LpLHEIFjOVUh1mLWxu-piYnB_dLOLCMQD8FlCJwAqihytE6LmWSraLs777wRwY8wpYfoiFAjODIdO_o8ss4UzpdWO6UcLeL51G1cJeMwsZ5fbHDoyIV4f9-GeKIWnyQFIqIjxHVtV-wKfJOY5T5uZyKzoDHoI7JdY4bpp8-WeMSXalbgHXS_lVaD00ZQHY9P8-ynncnaPtnnniPwbiUTLTjXxyPway_E9o_MA8EBzTilmzSnn5anQCGCsjsIBINKF4B0neuvJCTPEEMniAwoCg7TgL1AequKOl_-0Nh2qnrHjgqQUiu64ZDTnKAOnDzpIOZh-0M5PlUK8tE_G_Z6FmmJuh5StZ5MgdDdH04w3kcTx2Y_yy5BY3szO8D-ZFkd7kCc-LrgHBfXeCrOMkq86KtYYEogbc9ZUDsFVETQIgbiPJPHtEgvtuV-kPBye3DCUHGLTEGJ6gRatZtkk4oIdxSGS_99MgkFN-C_4P1YRSOQBURGYrP6gBr4y9DH-rassxe_hysSbrGvd4YfVG5QFejimBgN5j8AS6H1urDvSWigorLFTKPMPB0aAWHRCVDMJuI5HpZY7QkogZ9Jb87wNcA_xPiE-fzgaQgvYBBTViPXneWmXyTOnCo4uCqJOeFZNlBBnqyUOo3rZAHzQwDI_ouYnIc9vvKqKvbsfhdaQRk_nA1TwETBieubToi_qrdv_XrtucgMc--XTPGZ7HuXzGZsJw2vXsRaPUG80HAW0wRs7FDSUcUjWGFTlYGkArrLKuWia3iBzkpzwro0ogviI1hjHMYYg3eF0wISuy1EkX7AmJadyLMpWyaKiyeje4oh7OufmC0UsA8zehe8S8t-39pQ2ilHYN4_ZKyYZruU820C4KRUlgqmlVfr4fNicdGb1suDNcv3TRoaUu9bQsGUSsLBrAolecD6AX9wMSK0NElAeDPnqX5VDFJqyQaNvAYPvO2K-tSG_w20s2lkuxi8YPa1QhJWZ2hZAjyzNZcuixhjfOTLWPAuLNx6xUeEQMSXFdhzrCkYiZl97Qy8bBPV7H6oZTyI5V2C-MLcCzthj2bNdbORG1TplsAqZmSqpZ7DQLRjixqblD_IgK9q2Wv7uRvnMFXM-fLyj7Hhru-ekAw110YQZKrWhjfcUMEuz1pAtGgz4gEq5dP9xZMeYNCO2WIH0-bl4ui_Y1h0oT3bNdbfvh9dTZvrdnLW459sqwJX4aLt_8J_LsvC_plLkLwKqCNYdCIJn9lawWMAj-pyhmJuhcWEqB5Uz5MkcW20PHOOjs1mMhIm2T3fQgHhNu3_bYRlzW8fSUAJaGr1GYC5R-ebyns5HXcUe26L7jUyeRC6twf9RAY6_LS2N6mLNzv14HM02PbTnDFW8L-cCSFu1Oj2PYYaPT3MDBAqjubar3OH01c5Jjq2emXwUl0e-eXAAw7vPC99F_hZpPczdQEnEAhOFTstClQHj6mznz2rfJZyJTzcwR-IPSL0mmO8Lz3amHSICIjGn9uRKRxNpR7FLus4FRIvgD4NH5aiz0ig9ztYjqQhhQhX8QpAfdSAuIgbd4TrXsuC6jmswTHFQCSQ9MRez3HVT4EZ7peabzd1Kc5BOvVnTuuxqye97UcnOdaBCnW2pC4u7KSbdgO0xKPC0yof86VA0YqeXUsTWYKblMILiy2u4yMzoEIdlCGqSyCUnkyie0_APDiYNnZWrTwbJ_eIHn0eObi0LQ4kN9R-OWXtWfy7gq0g4BJMYbr1IUqoRYirezeW480WCsIp8cBtdfftYKjWIcbXyxfOFFlr8elu-KhdDtkyBRc1AlEAEpR48zZFBo_F2HSx-TxxAeusYrbPcrafuxoCipRQ9Ccc2K1jV4fNMdsJXHuB9UtgC6tfBmGWxFFammh9qJpZgu1EYaZjEiH0dGbwHQzm60_QN-6V-6wk6psXKXZzY73SIgbGN6AWBNwB3HH9j2n5XRFhnhMajVSTvgYjqqmtggIrdtAntjEfUlZa0-61GR2O6yZ5Cip99Yg0me3RMYR6iK8bANxkJ5RNci6N_uLitHA3NgDazzYrcqrY8HAC1O1MyPCBThM5QUALV5oi9Xpb_c86Bsq-z3CLQ4nxgU3XODyI1rgyzQ3efXLuca8f0a_zKzeAVk3jaiBaTjXtyRRrAFNpZkKUl-jSCUrt4TplYADHWxOuIDIKiLF665md5VTwP6vW5-XBdnecvDxcybKkR58O8rFkz92E_KYVfMWO5x5ZsOqX2xjLMLBrEtcXHbbHIflWVh9d0UA2iwJGSep384lMf3P1EFc09z845XEx9H3Iad14sDw0t2u4-_tD3xlKTtDaRPVuIhkmt7MuC6uWP-Om-iMz69zmaKBS0oMRENI7-P7iXxyV0Ri_IhdmVoxdxJ9i5QCQyBpf0R7W3UuiDktlgXByBDUILkiePM_a5tNOdx5p5c2hBEArTXAiPqmRK-mhwNhTMZi2-jv-c88KZnohtTydgyPsm_jbWJeosghqt3igTXYdC76fxnsuIRhHerL0tfuJ0uGqi1W__wt2MlFBMCDMGH4sSebmJxF91x7auinkf0QYtKDQymVjXB7zLWlvnua6QuKUUiYLm8xIsFixjTlb1dr0kmufO7g8QEyNVJ1cLajFx3TjAmeihy68Q3OwL_gpwmiiLVs1JLqdBWevhiQKzAvB7NECZue3u6qtxywv1ezCJcS3_wzFaLI84l_HEu4SLKreU7_30yVay1XQh73fe3JNBGl7JelqAUW7amjXlV0-j7C9LDIr_QxFaocHoXvCZUo5hB_4V_ZwxZMukLGcmdJe3a4DdB1N3w4tJhT3-ERSophsjoAmVdb-GPNFva_C6uEMf30OhoiOYLpg8w16hTi7n8F6QaiQ_GP1TbkTOxhpaBxFnyvN2WzgrrDwHGF0gD5yDiRaeVfLceez9aBBRC0suuVL6u9kV0EX73TJ_sEgH9FiBtgLA1vjd74EwBknhGe8hNtYFrYnwWHpH1mNx6AWYJlJf7VN8IPiEEN1NLThcU4ovDam5wcpK6Cb0BHYXkZb1WPG2DbP7Q8qH8zIs3hJbh4GDmhc2JF630fADHwWv4F0FW2HPwqcNc5n5wng8ogiezKJpL1sNFK2l_mczP1rYrv8OQC_Udc62Y-9ZAmG6aMMRhN1qWnx8eT0HPHtClrSvEK981lCaCsNZ7oQHNIdZE2V_0RPD4hPbkOk_hiRomhy9nh-KDOQbYm-duAmgUrpqlLwXnaAXK0N7Em865ZpTgCUTMzfhKNjKp3bN96DjhGphmoPbKUL6gmnkz7EYqsI4pQMqMTVfjepo0k1riPXIn7WXd3qUwRBKpZ_ey9QLeFgvBjSi5laVkzOusOgQuTvSBSUh2F2y5Ez0d1EgUnAILxBRWgFON4M2e5wH0QkhIf8LUfFhDJqIgAzGXXhnNljkeN6ZkFK2ObWfru3-3dmtryz0HTM6a8mhv6wUi96jDn87-O5gvmBtbuh2SU8Hu829G3HFiPHy5wkHars1U7MpK5grpyqTTkPkszDeHevbapPv41cLYvwxb469O5EWDvqQOjm4cltuub_REDebIjM67vy_ne4lQS2YHxvmtIaVte0ffYfcG4QwQTLaCMPVw1Fcg1n0Rj1NS71Ak64gDThjAcxUCnNf8CU30BOxE-PDCotE-_BArnY4aN-OoKKJ1pvIbA4NRm3EabVZ55myt1MYIi4HhIslX8jY-LjcqN8lWoLPuqjwfhEMI4wFmuGeiH_8EMpCh7fzFDP535z4yOMRlNwVfMbmTZqF_RE-YW_5TIlgwwsWlRMqCJnvuYvXelS4YGWuMPkGMGHNzIZ9sUMDar4l6HY9zVP_vyPFTiok5KkewRWybUHQ-AuPSwqR-U61sXD22MpT_Tp021hx3VAcoUslsd53O1C4=","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"summary":[{"type":"summary_text","text":"**Checking first-write timing**"},{"type":"summary_text","text":"**Checking integration coverage**"},{"type":"summary_text","text":"**Checking log consumers**"}],"metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f7a007bc87d2836c5242a2a4d007","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading single-write diagnostics\",\"path\":\".worktrees/286-churn/crates/spt-proto/src/emit.rs:1-210\"}","call_id":"call_c1vQmm2hsWUKzqmKdbl1qo8j","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"read","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f7a007cc87d29a90ca68b66ba45a","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading connection lifecycle harness\",\"path\":\".worktrees/286-churn/crates/spt-daemon/tests/conn_blackhole_lifecycle.rs\"}","call_id":"call_ab5g2v3JBCNBrGNwO5njnWls","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"read","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f7a007d487d29946f6e7adcfcd4e","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading registry amendment context\",\"path\":\".worktrees/286-churn/traceable-reqs.toml:2553-2593,4855-4890:raw\"}","call_id":"call_0mqcQOcn96wUWhCzdA7Rvd5z","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"read","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f7a007d887d2919b64622734e05f","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading changelog context\",\"path\":\".worktrees/286-churn/CHANGELOG.md:1-45:raw\"}","call_id":"call_H81TcWUzZzTRJSYvgdSphpBV","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"read","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f7a007e087d28db3fbdc88b7ebc5","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading broker architecture decision\",\"path\":\".worktrees/286-churn/docs/adr/0004-single-daemon-broker-brain-split-and-self-update.md:1-45:raw\"}","call_id":"call_dBzQK6ug0ipNvJSIUzrDNrk4","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"read","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f7a007e887d28997dd6d806c2a20","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading client transport handshake\",\"path\":\".worktrees/286-churn/crates/spt-daemon/src/transport.rs:1-123\"}","call_id":"call_yQ5v4ngdG3LVvysjf5PLaogH","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"read","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f7a007ec87d28370d031e74a7b61","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading review commit metadata\",\"command\":\"git log de5a44bc338d679f3c63f6d0f68826eb9e575d73..19a9097956de0db898eaeb3a7dc4765189ade874 --format=full\",\"cwd\":\"C:/Users/decid/Documents/projects/spt-core/.worktrees/286-churn\"}","call_id":"call_BkJaSiBdyn99uE5hrUyTuJhm","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"bash","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}}]},"duration":29102.73809999996,"ttft":3302.930799999973,"completedAt":1788999584180,"contextSnapshot":{"promptTokens":113650,"nonMessageTokens":13648,"compactionEpoch":0}}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_dBzQK6ug0ipNvJSIUzrDNrk4|fc_0dde72effda0ad21016aa1f7a007e087d28db3fbdc88b7ebc5","toolName":"read","startedAt":"2026-09-10T00:19:44.184Z","args":{"path":".worktrees/286-churn/docs/adr/0004-single-daemon-broker-brain-split-and-self-update.md:1-45:raw"},"intent":"Reading broker architecture decision"},"id":"d48d3035","parentId":"fb832111","timestamp":"2026-09-10T00:19:44.184Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_yQ5v4ngdG3LVvysjf5PLaogH|fc_0dde72effda0ad21016aa1f7a007e887d28997dd6d806c2a20","toolName":"read","startedAt":"2026-09-10T00:19:44.184Z","args":{"path":".worktrees/286-churn/crates/spt-daemon/src/transport.rs:1-123"},"intent":"Reading client transport handshake"},"id":"47f11f70","parentId":"d48d3035","timestamp":"2026-09-10T00:19:44.184Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_BkJaSiBdyn99uE5hrUyTuJhm|fc_0dde72effda0ad21016aa1f7a007ec87d28370d031e74a7b61","toolName":"bash","startedAt":"2026-09-10T00:19:44.184Z","args":{"command":"git log de5a44bc338d679f3c63f6d0f68826eb9e575d73..19a9097956de0db898eaeb3a7dc4765189ade874 --format=full"},"intent":"Reading review commit metadata"},"id":"00ddddbc","parentId":"47f11f70","timestamp":"2026-09-10T00:19:44.184Z"}
{"type":"message","id":"bde2cc1a","parentId":"00ddddbc","timestamp":"2026-09-10T00:19:44.211Z","message":{"role":"toolResult","toolCallId":"call_c1vQmm2hsWUKzqmKdbl1qo8j|fc_0dde72effda0ad21016aa1f7a007bc87d2836c5242a2a4d007","toolName":"read","content":[{"type":"text","text":"//! Single-write emission of machine-parsed diagnostic lines.\n//!\n//! A diagnostic a test parses (`TOKEN:...`) must reach the handle as ONE write\n//! of the whole rendered text, newline included. `eprintln!` does not do that:\n//! `io::Stderr` is unbuffered, and `Write::write_fmt` single-writes only a\n//! statically known str — anything interpolated goes through an adapter that\n//! `write_all`s PER FORMAT FRAGMENT, the trailing newline separately. Threads of\n//! one process are serialized by a reentrant lock, but the daemon's children\n//! inherit the same handle and no lock spans processes, so a reader can observe\n//! a line torn between fragments and completed by another process's newline.\n//!\n//! MEASURED on this toolchain, not inferred (W1 step 0, the lane's gate\n//! evidence): an interpolated `write_fmt` issued **5** writes —\n//! `\"SUBSCRIBE_DECISION:\"`, `\"endpoint-7\"`, `\": \"`, `\"the store did not\n//! persist\"`, `\"\\n\"` — while the static-literal control issued **1**.\n//!\n//! Two named arms, because the population has two shapes and a weak assertion\n//! would let them blur:\n//!\n//! - [`emit_line!`] — the single-line contract. It refuses an interior newline:\n//!   `debug_assert` in debug, and in release it ESCAPES rather than emitting the\n//!   text as-is, so one emission is always one logical line. Emitting anyway\n//!   would put multi-line text back on the single-line surface in the one build\n//!   where no assert watches and no test looks.\n//! - [`emit_block!`] — the explicit opt-in for a composed multi-line block,\n//!   still exactly one write. Two sites in the tree need it and now say so by\n//!   name: `firewall.rs` INBOUND_REACHABILITY (its newline rides in the payload)\n//!   and `servicehost.rs` SERVICE_STARTUP_FAULT (it appends a captured tail).\n//!\n//! The handle is a parameter everywhere here. A hardcoded `stderr()` would make\n//! the write count unobservable, and an unobservable property is not a property.\n\nuse std::io::{self, Write};\n\n/// The escape an interior newline is rewritten to by [`compose_line`], pinned\n/// here and asserted in a unit so it cannot drift: a reader who finds these two\n/// characters inside a parsed line has the offending caller named by the\n/// artifact alone, with no repro.\n// [impl->REQ-EMIT-SINGLE-WRITE]\npub const ESCAPED_NEWLINE: &str = r\"\\n\";\n\n/// The same for a bare carriage return, which splits a line for plenty of\n/// readers and would otherwise survive the newline escape as a stray byte.\n// [impl->REQ-EMIT-SINGLE-WRITE]\npub const ESCAPED_RETURN: &str = r\"\\r\";\n\n/// Compose a SINGLE-LINE emission: exactly one logical line, newline-terminated.\n///\n/// Pure on purpose — the text is unit-pinnable without reaching through any\n/// writer, which is what makes the escape spelling testable at all. One trailing\n/// line break is the caller's own terminator and is absorbed rather than\n/// escaped; everything left inside is rewritten, because a payload carrying its\n/// own newline would otherwise pass as two logical lines inside one write —\n/// indistinguishable downstream from the tear this module exists to close.\n// [impl->REQ-EMIT-SINGLE-WRITE]\npub fn compose_line(rendered: &str) -> String {\n    let body = rendered\n        .strip_suffix('\\n')\n        .map(|b| b.strip_suffix('\\r').unwrap_or(b))\n        .unwrap_or(rendered);\n    let mut out = String::with_capacity(body.len() + 1);\n    for ch in body.chars() {\n        match ch {\n            '\\n' => out.push_str(ESCAPED_NEWLINE),\n            '\\r' => out.push_str(ESCAPED_RETURN),\n            other => out.push(other),\n        }\n    }\n    out.push('\\n');\n    out\n}\n\n/// Compose a MULTI-LINE block emission: interior line breaks are the point, and\n/// the block is still handed over as one write. Terminated exactly once.\n// [impl->REQ-EMIT-SINGLE-WRITE]\npub fn compose_block(rendered: &str) -> String {\n    if rendered.ends_with('\\n') {\n        rendered.to_string()\n    } else {\n        let mut out = String::with_capacity(rendered.len() + 1);\n        out.push_str(rendered);\n        out.push('\\n');\n        out\n    }\n}\n\n/// True when `rendered` carries a line break that is not simply its own\n/// terminator — the condition [`write_line`] refuses.\n// [impl->REQ-EMIT-SINGLE-WRITE]\npub fn has_interior_break(rendered: &str) -> bool {\n    let body = rendered\n        .strip_suffix('\\n')\n        .map(|b| b.strip_suffix('\\r').unwrap_or(b))\n        .unwrap_or(rendered);\n    body.contains('\\n') || body.contains('\\r')\n}\n\n/// Write one composed emission as ONE issued call. Every path here funnels\n/// through it, so there is a single place where the write count is decided.\n///\n/// `write_all` and not a bare `write`, ruled by doyle 2026-08-29 and worth the\n/// sentence: `write_all` retries a short write, so under an OS short-write it\n/// can issue more than one call. A bare `write` in that same window DROPS THE\n/// TAIL — the line loses its newline and the next emission glues onto the\n/// partial, which MANUFACTURES exactly the torn-token-plus-glued-line signature\n/// this module exists to retire, and loses the diagnostic silently. Both are\n/// imperfect in that rare window; `write_all` delivers late and complete, which\n/// is strictly less bad. What we promise is about OUR fragmentation: one\n/// composed buffer, one issued call, never a write per format fragment.\n// [impl->REQ-EMIT-SINGLE-WRITE]\npub fn write_composed<W: Write + ?Sized>(w: &mut W, composed: &str) -> io::Result<()> {\n    w.write_all(composed.as_bytes())\n}\n\n/// The single-line emission path: refuse an interior break, compose, one write.\n///\n/// The `debug_assert` is what stops a NEW caller shipping an interior break; the\n/// escape in [`compose_line`] is what keeps the surface honest in release, where\n/// the assert is compiled out. Both halves are needed — an assert alone would\n/// leave release doing the thing no test can see.\n// [impl->REQ-EMIT-SINGLE-WRITE]\npub fn write_line<W: Write + ?Sized>(w: &mut W, rendered: &str) -> io::Result<()> {\n    debug_assert!(\n        !has_interior_break(rendered),\n        \"emit_line! got an interior line break — use emit_block! if the emission \\\n         really is multi-line; the release build escapes it to {ESCAPED_NEWLINE} \\\n         so one emission stays one logical line: {rendered:?}\"\n    );\n    write_composed(w, &compose_line(rendered))\n}\n\n/// The multi-line block emission path: interior breaks kept, still one write.\n// [impl->REQ-EMIT-SINGLE-WRITE]\npub fn write_block<W: Write + ?Sized>(w: &mut W, rendered: &str) -> io::Result<()> {\n    write_composed(w, &compose_block(rendered))\n}\n\n/// Emit ONE machine-parsed line to `$w` in exactly one write.\n///\n/// ```\n/// # use spt_proto::emit::CountingWriter;\n/// let mut w = CountingWriter::new();\n/// let id = \"endpoint-7\";\n/// spt_proto::emit_line!(&mut w, \"SUBSCRIBE_DECISION:{id}: seated\");\n/// assert_eq!(w.call_count(), 1);\n/// ```\n// [impl->REQ-EMIT-SINGLE-WRITE]\n#[macro_export]\nmacro_rules! emit_line {\n    ($w:expr, $($arg:tt)*) => {\n        $crate::emit::write_line($w, &::std::format!($($arg)*))\n    };\n}\n\n/// Emit one composed MULTI-LINE block to `$w` in exactly one write.\n// [impl->REQ-EMIT-SINGLE-WRITE]\n#[macro_export]\nmacro_rules! emit_block {\n    ($w:expr, $($arg:tt)*) => {\n        $crate::emit::write_block($w, &::std::format!($($arg)*))\n    };\n}\n\n/// [`emit_line!`] aimed at this process's stderr — the conversion target for the\n/// `TOKEN:`-shaped `eprintln!` population.\n///\n/// Deliberately a SEPARATE NAME rather than a no-writer arm of [`emit_line!`]:\n/// a macro that accepts both `($w, $fmt, ...)` and `($fmt, ...)` would bind a\n/// caller's format literal as the writer, since a literal is also an expression.\n/// One seam, two spellings, no ambiguity.\n///\n/// It returns `()` and DISCARDS a write error, which is a deliberate departure\n/// from `eprintln!` — stated here at the seam, where the next reader is, rather\n/// than only in the commit that made it. std's print machinery PANICS when\n/// writing to stderr fails, so a closed or full stderr can take down a process\n/// at the exact moment it was trying to explain itself. That is the same wedge\n/// class this project has refused before: a reporting surface must never become\n/// the thing that kills the operation it reports on.\n///\n/// The cost is named, not hidden: this makes emission a place where absence can\n/// be SILENT — a diagnostic that never arrives leaves no trace of having been\n/// attempted. That is an observability gap, accepted here only because the\n/// alternative is a panic, and it belongs in the conversion notes so no later\n/// reader rediscovers it as a defect.\n// [impl->REQ-EMIT-SINGLE-WRITE]\n#[macro_export]\nmacro_rules! emit_line_err {\n    ($($arg:tt)*) => {{\n        let _ = $crate::emit::write_line(\n            &mut ::std::io::stderr().lock(),\n            &::std::format!($($arg)*),\n        );\n    }};\n}\n\n/// [`emit_block!`] aimed at this process's stderr, on the same terms as\n/// [`emit_line_err!`]: separate name, error discarded.\n// [impl->REQ-EMIT-SINGLE-WRITE]\n#[macro_export]\nmacro_rules! emit_block_err {\n    ($($arg:tt)*) => {{\n        let _ = $crate::emit::write_block(\n            &mut ::std::io::stderr().lock(),\n            &::std::format!($($arg)*),\n        );\n    }};\n}\n\n/// A writer that counts the calls it receives and keeps every buffer verbatim.\n///\n/// Public because the property this module promises is a claim about CALL COUNT,\n/// and every crate converting its emission sites needs the same meter to prove\n/// it locally rather than trusting this one.\n\n[Showing lines 1-213 of 543. Use :214 to continue]"}],"details":{"truncation":{"content":"//! Single-write emission of machine-parsed diagnostic lines.\r\n//!\r\n//! A diagnostic a test parses (`TOKEN:...`) must reach the handle as ONE write\r\n//! of the whole rendered text, newline included. `eprintln!` does not do that:\r\n//! `io::Stderr` is unbuffered, and `Write::write_fmt` single-writes only a\r\n//! statically known str — anything interpolated goes through an adapter that\r\n//! `write_all`s PER FORMAT FRAGMENT, the trailing newline separately. Threads of\r\n//! one process are serialized by a reentrant lock, but the daemon's children\r\n//! inherit the same handle and no lock spans processes, so a reader can observe\r\n//! a line torn between fragments and completed by another process's newline.\r\n//!\r\n//! MEASURED on this toolchain, not inferred (W1 step 0, the lane's gate\r\n//! evidence): an interpolated `write_fmt` issued **5** writes —\r\n//! `\"SUBSCRIBE_DECISION:\"`, `\"endpoint-7\"`, `\": \"`, `\"the store did not\r\n//! persist\"`, `\"\\n\"` — while the static-literal control issued **1**.\r\n//!\r\n//! Two named arms, because the population has two shapes and a weak assertion\r\n//! would let them blur:\r\n//!\r\n//! - [`emit_line!`] — the single-line contract. It refuses an interior newline:\r\n//!   `debug_assert` in debug, and in release it ESCAPES rather than emitting the\r\n//!   text as-is, so one emission is always one logical line. Emitting anyway\r\n//!   would put multi-line text back on the single-line surface in the one build\r\n//!   where no assert watches and no test looks.\r\n//! - [`emit_block!`] — the explicit opt-in for a composed multi-line block,\r\n//!   still exactly one write. Two sites in the tree need it and now say so by\r\n//!   name: `firewall.rs` INBOUND_REACHABILITY (its newline rides in the payload)\r\n//!   and `servicehost.rs` SERVICE_STARTUP_FAULT (it appends a captured tail).\r\n//!\r\n//! The handle is a parameter everywhere here. A hardcoded `stderr()` would make\r\n//! the write count unobservable, and an unobservable property is not a property.\r\n\r\nuse std::io::{self, Write};\r\n\r\n/// The escape an interior newline is rewritten to by [`compose_line`], pinned\r\n/// here and asserted in a unit so it cannot drift: a reader who finds these two\r\n/// characters inside a parsed line has the offending caller named by the\r\n/// artifact alone, with no repro.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\npub const ESCAPED_NEWLINE: &str = r\"\\n\";\r\n\r\n/// The same for a bare carriage return, which splits a line for plenty of\r\n/// readers and would otherwise survive the newline escape as a stray byte.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\npub const ESCAPED_RETURN: &str = r\"\\r\";\r\n\r\n/// Compose a SINGLE-LINE emission: exactly one logical line, newline-terminated.\r\n///\r\n/// Pure on purpose — the text is unit-pinnable without reaching through any\r\n/// writer, which is what makes the escape spelling testable at all. One trailing\r\n/// line break is the caller's own terminator and is absorbed rather than\r\n/// escaped; everything left inside is rewritten, because a payload carrying its\r\n/// own newline would otherwise pass as two logical lines inside one write —\r\n/// indistinguishable downstream from the tear this module exists to close.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\npub fn compose_line(rendered: &str) -> String {\r\n    let body = rendered\r\n        .strip_suffix('\\n')\r\n        .map(|b| b.strip_suffix('\\r').unwrap_or(b))\r\n        .unwrap_or(rendered);\r\n    let mut out = String::with_capacity(body.len() + 1);\r\n    for ch in body.chars() {\r\n        match ch {\r\n            '\\n' => out.push_str(ESCAPED_NEWLINE),\r\n            '\\r' => out.push_str(ESCAPED_RETURN),\r\n            other => out.push(other),\r\n        }\r\n    }\r\n    out.push('\\n');\r\n    out\r\n}\r\n\r\n/// Compose a MULTI-LINE block emission: interior line breaks are the point, and\r\n/// the block is still handed over as one write. Terminated exactly once.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\npub fn compose_block(rendered: &str) -> String {\r\n    if rendered.ends_with('\\n') {\r\n        rendered.to_string()\r\n    } else {\r\n        let mut out = String::with_capacity(rendered.len() + 1);\r\n        out.push_str(rendered);\r\n        out.push('\\n');\r\n        out\r\n    }\r\n}\r\n\r\n/// True when `rendered` carries a line break that is not simply its own\r\n/// terminator — the condition [`write_line`] refuses.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\npub fn has_interior_break(rendered: &str) -> bool {\r\n    let body = rendered\r\n        .strip_suffix('\\n')\r\n        .map(|b| b.strip_suffix('\\r').unwrap_or(b))\r\n        .unwrap_or(rendered);\r\n    body.contains('\\n') || body.contains('\\r')\r\n}\r\n\r\n/// Write one composed emission as ONE issued call. Every path here funnels\r\n/// through it, so there is a single place where the write count is decided.\r\n///\r\n/// `write_all` and not a bare `write`, ruled by doyle 2026-08-29 and worth the\r\n/// sentence: `write_all` retries a short write, so under an OS short-write it\r\n/// can issue more than one call. A bare `write` in that same window DROPS THE\r\n/// TAIL — the line loses its newline and the next emission glues onto the\r\n/// partial, which MANUFACTURES exactly the torn-token-plus-glued-line signature\r\n/// this module exists to retire, and loses the diagnostic silently. Both are\r\n/// imperfect in that rare window; `write_all` delivers late and complete, which\r\n/// is strictly less bad. What we promise is about OUR fragmentation: one\r\n/// composed buffer, one issued call, never a write per format fragment.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\npub fn write_composed<W: Write + ?Sized>(w: &mut W, composed: &str) -> io::Result<()> {\r\n    w.write_all(composed.as_bytes())\r\n}\r\n\r\n/// The single-line emission path: refuse an interior break, compose, one write.\r\n///\r\n/// The `debug_assert` is what stops a NEW caller shipping an interior break; the\r\n/// escape in [`compose_line`] is what keeps the surface honest in release, where\r\n/// the assert is compiled out. Both halves are needed — an assert alone would\r\n/// leave release doing the thing no test can see.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\npub fn write_line<W: Write + ?Sized>(w: &mut W, rendered: &str) -> io::Result<()> {\r\n    debug_assert!(\r\n        !has_interior_break(rendered),\r\n        \"emit_line! got an interior line break — use emit_block! if the emission \\\r\n         really is multi-line; the release build escapes it to {ESCAPED_NEWLINE} \\\r\n         so one emission stays one logical line: {rendered:?}\"\r\n    );\r\n    write_composed(w, &compose_line(rendered))\r\n}\r\n\r\n/// The multi-line block emission path: interior breaks kept, still one write.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\npub fn write_block<W: Write + ?Sized>(w: &mut W, rendered: &str) -> io::Result<()> {\r\n    write_composed(w, &compose_block(rendered))\r\n}\r\n\r\n/// Emit ONE machine-parsed line to `$w` in exactly one write.\r\n///\r\n/// ```\r\n/// # use spt_proto::emit::CountingWriter;\r\n/// let mut w = CountingWriter::new();\r\n/// let id = \"endpoint-7\";\r\n/// spt_proto::emit_line!(&mut w, \"SUBSCRIBE_DECISION:{id}: seated\");\r\n/// assert_eq!(w.call_count(), 1);\r\n/// ```\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\n#[macro_export]\r\nmacro_rules! emit_line {\r\n    ($w:expr, $($arg:tt)*) => {\r\n        $crate::emit::write_line($w, &::std::format!($($arg)*))\r\n    };\r\n}\r\n\r\n/// Emit one composed MULTI-LINE block to `$w` in exactly one write.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\n#[macro_export]\r\nmacro_rules! emit_block {\r\n    ($w:expr, $($arg:tt)*) => {\r\n        $crate::emit::write_block($w, &::std::format!($($arg)*))\r\n    };\r\n}\r\n\r\n/// [`emit_line!`] aimed at this process's stderr — the conversion target for the\r\n/// `TOKEN:`-shaped `eprintln!` population.\r\n///\r\n/// Deliberately a SEPARATE NAME rather than a no-writer arm of [`emit_line!`]:\r\n/// a macro that accepts both `($w, $fmt, ...)` and `($fmt, ...)` would bind a\r\n/// caller's format literal as the writer, since a literal is also an expression.\r\n/// One seam, two spellings, no ambiguity.\r\n///\r\n/// It returns `()` and DISCARDS a write error, which is a deliberate departure\r\n/// from `eprintln!` — stated here at the seam, where the next reader is, rather\r\n/// than only in the commit that made it. std's print machinery PANICS when\r\n/// writing to stderr fails, so a closed or full stderr can take down a process\r\n/// at the exact moment it was trying to explain itself. That is the same wedge\r\n/// class this project has refused before: a reporting surface must never become\r\n/// the thing that kills the operation it reports on.\r\n///\r\n/// The cost is named, not hidden: this makes emission a place where absence can\r\n/// be SILENT — a diagnostic that never arrives leaves no trace of having been\r\n/// attempted. That is an observability gap, accepted here only because the\r\n/// alternative is a panic, and it belongs in the conversion notes so no later\r\n/// reader rediscovers it as a defect.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\n#[macro_export]\r\nmacro_rules! emit_line_err {\r\n    ($($arg:tt)*) => {{\r\n        let _ = $crate::emit::write_line(\r\n            &mut ::std::io::stderr().lock(),\r\n            &::std::format!($($arg)*),\r\n        );\r\n    }};\r\n}\r\n\r\n/// [`emit_block!`] aimed at this process's stderr, on the same terms as\r\n/// [`emit_line_err!`]: separate name, error discarded.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\n#[macro_export]\r\nmacro_rules! emit_block_err {\r\n    ($($arg:tt)*) => {{\r\n        let _ = $crate::emit::write_block(\r\n            &mut ::std::io::stderr().lock(),\r\n            &::std::format!($($arg)*),\r\n        );\r\n    }};\r\n}\r\n\r\n/// A writer that counts the calls it receives and keeps every buffer verbatim.\r\n///\r\n/// Public because the property this module promises is a claim about CALL COUNT,\r\n/// and every crate converting its emission sites needs the same meter to prove\r\n/// it locally rather than trusting this one.\r","truncated":true,"truncatedBy":"lines","totalLines":543,"totalBytes":9767,"outputLines":213,"outputBytes":9767,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":543,"displayContent":{"text":"//! Single-write emission of machine-parsed diagnostic lines.\r\n//!\r\n//! A diagnostic a test parses (`TOKEN:...`) must reach the handle as ONE write\r\n//! of the whole rendered text, newline included. `eprintln!` does not do that:\r\n//! `io::Stderr` is unbuffered, and `Write::write_fmt` single-writes only a\r\n//! statically known str — anything interpolated goes through an adapter that\r\n//! `write_all`s PER FORMAT FRAGMENT, the trailing newline separately. Threads of\r\n//! one process are serialized by a reentrant lock, but the daemon's children\r\n//! inherit the same handle and no lock spans processes, so a reader can observe\r\n//! a line torn between fragments and completed by another process's newline.\r\n//!\r\n//! MEASURED on this toolchain, not inferred (W1 step 0, the lane's gate\r\n//! evidence): an interpolated `write_fmt` issued **5** writes —\r\n//! `\"SUBSCRIBE_DECISION:\"`, `\"endpoint-7\"`, `\": \"`, `\"the store did not\r\n//! persist\"`, `\"\\n\"` — while the static-literal control issued **1**.\r\n//!\r\n//! Two named arms, because the population has two shapes and a weak assertion\r\n//! would let them blur:\r\n//!\r\n//! - [`emit_line!`] — the single-line contract. It refuses an interior newline:\r\n//!   `debug_assert` in debug, and in release it ESCAPES rather than emitting the\r\n//!   text as-is, so one emission is always one logical line. Emitting anyway\r\n//!   would put multi-line text back on the single-line surface in the one build\r\n//!   where no assert watches and no test looks.\r\n//! - [`emit_block!`] — the explicit opt-in for a composed multi-line block,\r\n//!   still exactly one write. Two sites in the tree need it and now say so by\r\n//!   name: `firewall.rs` INBOUND_REACHABILITY (its newline rides in the payload)\r\n//!   and `servicehost.rs` SERVICE_STARTUP_FAULT (it appends a captured tail).\r\n//!\r\n//! The handle is a parameter everywhere here. A hardcoded `stderr()` would make\r\n//! the write count unobservable, and an unobservable property is not a property.\r\n\r\nuse std::io::{self, Write};\r\n\r\n/// The escape an interior newline is rewritten to by [`compose_line`], pinned\r\n/// here and asserted in a unit so it cannot drift: a reader who finds these two\r\n/// characters inside a parsed line has the offending caller named by the\r\n/// artifact alone, with no repro.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\npub const ESCAPED_NEWLINE: &str = r\"\\n\";\r\n\r\n/// The same for a bare carriage return, which splits a line for plenty of\r\n/// readers and would otherwise survive the newline escape as a stray byte.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\npub const ESCAPED_RETURN: &str = r\"\\r\";\r\n\r\n/// Compose a SINGLE-LINE emission: exactly one logical line, newline-terminated.\r\n///\r\n/// Pure on purpose — the text is unit-pinnable without reaching through any\r\n/// writer, which is what makes the escape spelling testable at all. One trailing\r\n/// line break is the caller's own terminator and is absorbed rather than\r\n/// escaped; everything left inside is rewritten, because a payload carrying its\r\n/// own newline would otherwise pass as two logical lines inside one write —\r\n/// indistinguishable downstream from the tear this module exists to close.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\npub fn compose_line(rendered: &str) -> String {\r\n    let body = rendered\r\n        .strip_suffix('\\n')\r\n        .map(|b| b.strip_suffix('\\r').unwrap_or(b))\r\n        .unwrap_or(rendered);\r\n    let mut out = String::with_capacity(body.len() + 1);\r\n    for ch in body.chars() {\r\n        match ch {\r\n            '\\n' => out.push_str(ESCAPED_NEWLINE),\r\n            '\\r' => out.push_str(ESCAPED_RETURN),\r\n            other => out.push(other),\r\n        }\r\n    }\r\n    out.push('\\n');\r\n    out\r\n}\r\n\r\n/// Compose a MULTI-LINE block emission: interior line breaks are the point, and\r\n/// the block is still handed over as one write. Terminated exactly once.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\npub fn compose_block(rendered: &str) -> String {\r\n    if rendered.ends_with('\\n') {\r\n        rendered.to_string()\r\n    } else {\r\n        let mut out = String::with_capacity(rendered.len() + 1);\r\n        out.push_str(rendered);\r\n        out.push('\\n');\r\n        out\r\n    }\r\n}\r\n\r\n/// True when `rendered` carries a line break that is not simply its own\r\n/// terminator — the condition [`write_line`] refuses.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\npub fn has_interior_break(rendered: &str) -> bool {\r\n    let body = rendered\r\n        .strip_suffix('\\n')\r\n        .map(|b| b.strip_suffix('\\r').unwrap_or(b))\r\n        .unwrap_or(rendered);\r\n    body.contains('\\n') || body.contains('\\r')\r\n}\r\n\r\n/// Write one composed emission as ONE issued call. Every path here funnels\r\n/// through it, so there is a single place where the write count is decided.\r\n///\r\n/// `write_all` and not a bare `write`, ruled by doyle 2026-08-29 and worth the\r\n/// sentence: `write_all` retries a short write, so under an OS short-write it\r\n/// can issue more than one call. A bare `write` in that same window DROPS THE\r\n/// TAIL — the line loses its newline and the next emission glues onto the\r\n/// partial, which MANUFACTURES exactly the torn-token-plus-glued-line signature\r\n/// this module exists to retire, and loses the diagnostic silently. Both are\r\n/// imperfect in that rare window; `write_all` delivers late and complete, which\r\n/// is strictly less bad. What we promise is about OUR fragmentation: one\r\n/// composed buffer, one issued call, never a write per format fragment.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\npub fn write_composed<W: Write + ?Sized>(w: &mut W, composed: &str) -> io::Result<()> {\r\n    w.write_all(composed.as_bytes())\r\n}\r\n\r\n/// The single-line emission path: refuse an interior break, compose, one write.\r\n///\r\n/// The `debug_assert` is what stops a NEW caller shipping an interior break; the\r\n/// escape in [`compose_line`] is what keeps the surface honest in release, where\r\n/// the assert is compiled out. Both halves are needed — an assert alone would\r\n/// leave release doing the thing no test can see.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\npub fn write_line<W: Write + ?Sized>(w: &mut W, rendered: &str) -> io::Result<()> {\r\n    debug_assert!(\r\n        !has_interior_break(rendered),\r\n        \"emit_line! got an interior line break — use emit_block! if the emission \\\r\n         really is multi-line; the release build escapes it to {ESCAPED_NEWLINE} \\\r\n         so one emission stays one logical line: {rendered:?}\"\r\n    );\r\n    write_composed(w, &compose_line(rendered))\r\n}\r\n\r\n/// The multi-line block emission path: interior breaks kept, still one write.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\npub fn write_block<W: Write + ?Sized>(w: &mut W, rendered: &str) -> io::Result<()> {\r\n    write_composed(w, &compose_block(rendered))\r\n}\r\n\r\n/// Emit ONE machine-parsed line to `$w` in exactly one write.\r\n///\r\n/// ```\r\n/// # use spt_proto::emit::CountingWriter;\r\n/// let mut w = CountingWriter::new();\r\n/// let id = \"endpoint-7\";\r\n/// spt_proto::emit_line!(&mut w, \"SUBSCRIBE_DECISION:{id}: seated\");\r\n/// assert_eq!(w.call_count(), 1);\r\n/// ```\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\n#[macro_export]\r\nmacro_rules! emit_line {\r\n    ($w:expr, $($arg:tt)*) => {\r\n        $crate::emit::write_line($w, &::std::format!($($arg)*))\r\n    };\r\n}\r\n\r\n/// Emit one composed MULTI-LINE block to `$w` in exactly one write.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\n#[macro_export]\r\nmacro_rules! emit_block {\r\n    ($w:expr, $($arg:tt)*) => {\r\n        $crate::emit::write_block($w, &::std::format!($($arg)*))\r\n    };\r\n}\r\n\r\n/// [`emit_line!`] aimed at this process's stderr — the conversion target for the\r\n/// `TOKEN:`-shaped `eprintln!` population.\r\n///\r\n/// Deliberately a SEPARATE NAME rather than a no-writer arm of [`emit_line!`]:\r\n/// a macro that accepts both `($w, $fmt, ...)` and `($fmt, ...)` would bind a\r\n/// caller's format literal as the writer, since a literal is also an expression.\r\n/// One seam, two spellings, no ambiguity.\r\n///\r\n/// It returns `()` and DISCARDS a write error, which is a deliberate departure\r\n/// from `eprintln!` — stated here at the seam, where the next reader is, rather\r\n/// than only in the commit that made it. std's print machinery PANICS when\r\n/// writing to stderr fails, so a closed or full stderr can take down a process\r\n/// at the exact moment it was trying to explain itself. That is the same wedge\r\n/// class this project has refused before: a reporting surface must never become\r\n/// the thing that kills the operation it reports on.\r\n///\r\n/// The cost is named, not hidden: this makes emission a place where absence can\r\n/// be SILENT — a diagnostic that never arrives leaves no trace of having been\r\n/// attempted. That is an observability gap, accepted here only because the\r\n/// alternative is a panic, and it belongs in the conversion notes so no later\r\n/// reader rediscovers it as a defect.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\n#[macro_export]\r\nmacro_rules! emit_line_err {\r\n    ($($arg:tt)*) => {{\r\n        let _ = $crate::emit::write_line(\r\n            &mut ::std::io::stderr().lock(),\r\n            &::std::format!($($arg)*),\r\n        );\r\n    }};\r\n}\r\n\r\n/// [`emit_block!`] aimed at this process's stderr, on the same terms as\r\n/// [`emit_line_err!`]: separate name, error discarded.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\n#[macro_export]\r\nmacro_rules! emit_block_err {\r\n    ($($arg:tt)*) => {{\r\n        let _ = $crate::emit::write_block(\r\n            &mut ::std::io::stderr().lock(),\r\n            &::std::format!($($arg)*),\r\n        );\r\n    }};\r\n}\r\n\r\n/// A writer that counts the calls it receives and keeps every buffer verbatim.\r\n///\r\n/// Public because the property this module promises is a claim about CALL COUNT,\r\n/// and every crate converting its emission sites needs the same meter to prove\r\n/// it locally rather than trusting this one.\r","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120,121,122,123,124,125,126,127,128,129,130,131,132,133,134,135,136,137,138,139,140,141,142,143,144,145,146,147,148,149,150,151,152,153,154,155,156,157,158,159,160,161,162,163,164,165,166,167,168,169,170,171,172,173,174,175,176,177,178,179,180,181,182,183,184,185,186,187,188,189,190,191,192,193,194,195,196,197,198,199,200,201,202,203,204,205,206,207,208,209,210,211,212,213]},"fileSize":22720,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-proto\\src\\emit.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":543,"totalBytes":9767,"outputLines":213,"outputBytes":9767,"shownRange":{"start":1,"end":213},"nextOffset":214}}},"isError":false,"timestamp":1788999584211}}
{"type":"message","id":"b478984e","parentId":"bde2cc1a","timestamp":"2026-09-10T00:19:44.218Z","message":{"role":"toolResult","toolCallId":"call_ab5g2v3JBCNBrGNwO5njnWls|fc_0dde72effda0ad21016aa1f7a007cc87d29a90ca68b66ba45a","toolName":"read","content":[{"type":"text","text":"//! MSG-IDENTITY W6 / F-039 leg (e) — the DETERMINISTIC BLACK-HOLED-CONTROLLER\n//! LIFECYCLE HARNESS (REQ-CONN-BLACKHOLE-LIFECYCLE-HARNESS), on a REAL broker\n//! hosting REAL PTY children over the real local IPC.\n//!\n//! ## What this is\n//!\n//! This is hertz's RCA fix-shape item 5 made executable (`.claude/\n//! hertz-pty-freeze-rca.md` §Proposed fix shape, item 5): the standing\n//! conformance rig for the round-4 `SharedSend` fix class. It reproduces the\n//! EXACT Windows named-pipe black-hole (a controller that stops draining while a\n//! finite burst floods its socket, so the broker's writer parks inside\n//! `write_frame` on `WriteFileEx`+`SleepEx(INFINITE, alertable)` — no supported\n//! write timeout) and, against CURRENT v0.30.6 semantics (`crates/spt-daemon/\n//! src/conn.rs` bounded/cancelable/poison-on-failure writes, landed fc048bb),\n//! asserts hertz's FIVE invariants for ONE incident:\n//!\n//!   1. UNRELATED SESSIONS CONTINUE — a second, independent session's viewer keeps\n//!      receiving output throughout the black-hole incident (the fan-out never\n//!      blocks the whole broker on the one wedged conn);\n//!   2. THE BAD PHYSICAL CONNECTION IS CANCELED/CLOSED WITHIN THE BOUND — a write\n//!      on the retired old-client connection FAILS, and the conn's `handle_conn`\n//!      thread finishes within the write bound of logical release (the conn was\n//!      physically retired, not merely logically stall-evicted);\n//!   3. ITS WRITER EXITS — `broker.conn_handler_count()` falls to the\n//!      unrelated-viewer floor (<= 1; the persistent invariant-1 viewer holds one\n//!      handler) WITHOUT the black-holed client ever dropping its end (the parked\n//!      writer was aborted + the reader EOF-cleaned, so the handler thread ran to\n//!      completion), and the viewer flows output AFTER the sample latches — proving\n//!      the survivor is the viewer's handler, not a leaked one;\n//!   4. A FRESH VIEWER CAN ATTACH — a brand-new Viewer subscribe on the retired\n//!      session COMPLETES and RECEIVES replayed output from the frozen cursor\n//!      (the fan-out accepts + serves a new subscriber after retirement);\n//!   5. NO LOCK OR TASK REMAINS OWNED BY THE RETIRED CONNECTION — asserted via\n//!      PROXIES (this invariant has no direct seam): after retirement the broker\n//!      still answers a control-plane `KIND_SESSIONS` request PROMPTLY (no lock\n//!      is held by the retired writer), `conn_handler_count()` returns to the\n//!      unrelated-viewer floor without the client dropping its end (no task\n//!      lingers), AND both the\n//!      unrelated session's viewer and the fresh viewer flow output (the shared\n//!      fan-out lock is free). The conjunction of these three is the strongest\n//!      deterministic proxy available for \"nothing is still owned.\"\n//!\n//! ## Relationship to `brain_decouple.rs`\n//!\n//! This harness GENERALIZES `brain_decouple.rs`'s r4 gate\n//! (`non_draining_controller_stall_evict_releases_writer_and_connection`, the\n//! Windows-mandatory physical-retirement leg for\n//! REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK, which stays the primary\n//! gate for that hazard). It reuses the same SEED/trigger/burst choreography,\n//! the hand-back-the-unread-Stream pattern, the `conn_handler_count` and\n//! `stall_evict_count` seams, the old-client-write-fails probe, and the\n//! teardown-before-assert discipline. It ADDS the two pieces the r4 gate does\n//! not have, which are exactly hertz's invariants (1) and (4): a SECOND\n//! UNRELATED session whose viewer keeps ticking across the incident, and a fresh\n//! VIEWER attach (not just a controller take) proven to receive output.\n//!\n//! ## Why Windows-mandatory (`#![cfg(windows)]` — the whole file)\n//!\n//! Invariants 2, 3, and the task-ownership proxy of 5 are the PHYSICAL-retirement\n//! asserts, and the deterministic black-hole only reproduces on the Windows\n//! named-pipe path (`interprocess 2.4.2` → `WriteFileEx`+`SleepEx(INFINITE)`,\n//! `set_timeout` → `Unsupported`) — the exact field failure. On Unix the\n//! poison/retire path is a `shutdown(SHUT_RDWR)` abort of a buffer-fill block: a\n//! genuinely DIFFERENT mechanism, not this incident. Faking a Unix companion\n//! would test a different thing, so — mirroring the r4 gate's own\n//! `#[cfg(windows)]` leg — the entire harness is Windows-only. On other targets\n//! this test binary compiles to zero tests (no fakery).\n//!\n//! ## RUNNER CONTRACT — run process-per-test (nextest) or single-threaded\n//!\n//! Shares ONE process-global `SPT_HOME` + `SPT_BRAIN_WRITE_DEADLINE_MS`, spins a\n//! REAL broker + TWO PTY children. HEAVY class. Run isolated with\n//! `cargo nextest run -p spt-daemon --test conn_blackhole_lifecycle` (CI's path)\n//! or `cargo test --test conn_blackhole_lifecycle -- --test-threads=1`. NEVER a\n//! bare `cargo test --test conn_blackhole_lifecycle` on Windows. Every read in\n//! this test rides a bounded thread+channel or a polled deadline; the test\n//! captures plain values as it goes and only asserts AFTER teardown, so it can\n//! never hang even on broken code.\n#![cfg(windows)]\n\nuse std::sync::atomic::{AtomicU32, AtomicU64, Ordering};\n…\nuse spt_test_support::TestHome;\n\nstatic SEQ: AtomicU32 = AtomicU32::new(0);\nfn unique_name() -> String { … }\n\nfn init_home() -> TestHome {\n    TestHome::new()\n}\n\nfn kill_pid(pid: u32) { … }\n\nfn send(conn: &mut Stream, kind: &str, payload: serde_json::Value) {\n    write_frame(conn, &Envelope::new(kind, payload)).expect(\"send frame\");\n}\n\nfn connect(name: &str) -> Stream { … }\n\n/// A child that FLOODS stdout forever (never reads stdin) — drives the UNRELATED\n/// session so its viewer has a continuous frame stream to keep ticking on\n/// throughout the black-hole incident (invariant 1).\nfn flood_spawn_req(endpoint: &str) -> SpawnReq { … }\n\n/// A real PTY child (PowerShell) that emits `SEED`, parks on a first `Read-Host`,\n/// then on a stdin trigger emits a finite ~1.4 MiB burst (1400 lines × 1023 `B`s)\n/// and parks again. The seq-0 boot output gives a real `seq>0` controller frame;\n/// the burst floods a non-draining controller's socket so its writer parks INSIDE\n/// `write_frame` — the exact black-hole the RCA pinned. Verbatim sibling of the\n/// r4 gate's `seed_burst_spawn_req` in `brain_decouple.rs`.\nfn seed_burst_spawn_req(endpoint: &str) -> SpawnReq { … }\n\n/// Spawn a PTY child from a fresh conn and return `(spawn_conn, session_id)`.\nfn spawn_session(name: &str, req: SpawnReq) -> (Stream, u64) { … }\n\n/// Query the broker's stall-evict tally. Bounded thread + channel (a wedge\n/// surfaces as a timeout, not a hang). Verbatim sibling of the r4 gate helper.\nfn stall_evict_count(name: &str, deadline: Duration) -> u64 { … }\n\n/// One `KIND_SESSIONS` round-trip, returning the reply. The dispatch runs the\n/// `reap_dead_controller` → time-based `stall_evict_controller`, so polling this\n/// both DRIVES the logical eviction and reads the frozen `resume_seq` cursor.\n/// Bounded thread + channel. Verbatim sibling of the r4 gate helper.\nfn query_sessions(name: &str, deadline: Duration) -> Option<SessionsReply> { … }\n\n// [int->REQ-CONN-BLACKHOLE-LIFECYCLE-HARNESS]\n#[test]\nfn blackholed_controller_lifecycle_five_invariants() { … }\n\n[…571ln elided; re-read needed ranges, e.g. .worktrees/286-churn/crates/spt-daemon/tests/conn_blackhole_lifecycle.rs:82-97,101-104]"}],"details":{"displayContent":{"text":"//! MSG-IDENTITY W6 / F-039 leg (e) — the DETERMINISTIC BLACK-HOLED-CONTROLLER\n//! LIFECYCLE HARNESS (REQ-CONN-BLACKHOLE-LIFECYCLE-HARNESS), on a REAL broker\n//! hosting REAL PTY children over the real local IPC.\n//!\n//! ## What this is\n//!\n//! This is hertz's RCA fix-shape item 5 made executable (`.claude/\n//! hertz-pty-freeze-rca.md` §Proposed fix shape, item 5): the standing\n//! conformance rig for the round-4 `SharedSend` fix class. It reproduces the\n//! EXACT Windows named-pipe black-hole (a controller that stops draining while a\n//! finite burst floods its socket, so the broker's writer parks inside\n//! `write_frame` on `WriteFileEx`+`SleepEx(INFINITE, alertable)` — no supported\n//! write timeout) and, against CURRENT v0.30.6 semantics (`crates/spt-daemon/\n//! src/conn.rs` bounded/cancelable/poison-on-failure writes, landed fc048bb),\n//! asserts hertz's FIVE invariants for ONE incident:\n//!\n//!   1. UNRELATED SESSIONS CONTINUE — a second, independent session's viewer keeps\n//!      receiving output throughout the black-hole incident (the fan-out never\n//!      blocks the whole broker on the one wedged conn);\n//!   2. THE BAD PHYSICAL CONNECTION IS CANCELED/CLOSED WITHIN THE BOUND — a write\n//!      on the retired old-client connection FAILS, and the conn's `handle_conn`\n//!      thread finishes within the write bound of logical release (the conn was\n//!      physically retired, not merely logically stall-evicted);\n//!   3. ITS WRITER EXITS — `broker.conn_handler_count()` falls to the\n//!      unrelated-viewer floor (<= 1; the persistent invariant-1 viewer holds one\n//!      handler) WITHOUT the black-holed client ever dropping its end (the parked\n//!      writer was aborted + the reader EOF-cleaned, so the handler thread ran to\n//!      completion), and the viewer flows output AFTER the sample latches — proving\n//!      the survivor is the viewer's handler, not a leaked one;\n//!   4. A FRESH VIEWER CAN ATTACH — a brand-new Viewer subscribe on the retired\n//!      session COMPLETES and RECEIVES replayed output from the frozen cursor\n//!      (the fan-out accepts + serves a new subscriber after retirement);\n//!   5. NO LOCK OR TASK REMAINS OWNED BY THE RETIRED CONNECTION — asserted via\n//!      PROXIES (this invariant has no direct seam): after retirement the broker\n//!      still answers a control-plane `KIND_SESSIONS` request PROMPTLY (no lock\n//!      is held by the retired writer), `conn_handler_count()` returns to the\n//!      unrelated-viewer floor without the client dropping its end (no task\n//!      lingers), AND both the\n//!      unrelated session's viewer and the fresh viewer flow output (the shared\n//!      fan-out lock is free). The conjunction of these three is the strongest\n//!      deterministic proxy available for \"nothing is still owned.\"\n//!\n//! ## Relationship to `brain_decouple.rs`\n//!\n//! This harness GENERALIZES `brain_decouple.rs`'s r4 gate\n//! (`non_draining_controller_stall_evict_releases_writer_and_connection`, the\n//! Windows-mandatory physical-retirement leg for\n//! REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK, which stays the primary\n//! gate for that hazard). It reuses the same SEED/trigger/burst choreography,\n//! the hand-back-the-unread-Stream pattern, the `conn_handler_count` and\n//! `stall_evict_count` seams, the old-client-write-fails probe, and the\n//! teardown-before-assert discipline. It ADDS the two pieces the r4 gate does\n//! not have, which are exactly hertz's invariants (1) and (4): a SECOND\n//! UNRELATED session whose viewer keeps ticking across the incident, and a fresh\n//! VIEWER attach (not just a controller take) proven to receive output.\n//!\n//! ## Why Windows-mandatory (`#![cfg(windows)]` — the whole file)\n//!\n//! Invariants 2, 3, and the task-ownership proxy of 5 are the PHYSICAL-retirement\n//! asserts, and the deterministic black-hole only reproduces on the Windows\n//! named-pipe path (`interprocess 2.4.2` → `WriteFileEx`+`SleepEx(INFINITE)`,\n//! `set_timeout` → `Unsupported`) — the exact field failure. On Unix the\n//! poison/retire path is a `shutdown(SHUT_RDWR)` abort of a buffer-fill block: a\n//! genuinely DIFFERENT mechanism, not this incident. Faking a Unix companion\n//! would test a different thing, so — mirroring the r4 gate's own\n//! `#[cfg(windows)]` leg — the entire harness is Windows-only. On other targets\n//! this test binary compiles to zero tests (no fakery).\n//!\n//! ## RUNNER CONTRACT — run process-per-test (nextest) or single-threaded\n//!\n//! Shares ONE process-global `SPT_HOME` + `SPT_BRAIN_WRITE_DEADLINE_MS`, spins a\n//! REAL broker + TWO PTY children. HEAVY class. Run isolated with\n//! `cargo nextest run -p spt-daemon --test conn_blackhole_lifecycle` (CI's path)\n//! or `cargo test --test conn_blackhole_lifecycle -- --test-threads=1`. NEVER a\n//! bare `cargo test --test conn_blackhole_lifecycle` on Windows. Every read in\n//! this test rides a bounded thread+channel or a polled deadline; the test\n//! captures plain values as it goes and only asserts AFTER teardown, so it can\n//! never hang even on broken code.\n#![cfg(windows)]\n\nuse std::sync::atomic::{AtomicU32, AtomicU64, Ordering};\n…\nuse spt_test_support::TestHome;\n\nstatic SEQ: AtomicU32 = AtomicU32::new(0);\nfn unique_name() -> String { … }\n\nfn init_home() -> TestHome {\n    TestHome::new()\n}\n\nfn kill_pid(pid: u32) { … }\n\nfn send(conn: &mut Stream, kind: &str, payload: serde_json::Value) {\n    write_frame(conn, &Envelope::new(kind, payload)).expect(\"send frame\");\n}\n\nfn connect(name: &str) -> Stream { … }\n\n/// A child that FLOODS stdout forever (never reads stdin) — drives the UNRELATED\n/// session so its viewer has a continuous frame stream to keep ticking on\n/// throughout the black-hole incident (invariant 1).\nfn flood_spawn_req(endpoint: &str) -> SpawnReq { … }\n\n/// A real PTY child (PowerShell) that emits `SEED`, parks on a first `Read-Host`,\n/// then on a stdin trigger emits a finite ~1.4 MiB burst (1400 lines × 1023 `B`s)\n/// and parks again. The seq-0 boot output gives a real `seq>0` controller frame;\n/// the burst floods a non-draining controller's socket so its writer parks INSIDE\n/// `write_frame` — the exact black-hole the RCA pinned. Verbatim sibling of the\n/// r4 gate's `seed_burst_spawn_req` in `brain_decouple.rs`.\nfn seed_burst_spawn_req(endpoint: &str) -> SpawnReq { … }\n\n/// Spawn a PTY child from a fresh conn and return `(spawn_conn, session_id)`.\nfn spawn_session(name: &str, req: SpawnReq) -> (Stream, u64) { … }\n\n/// Query the broker's stall-evict tally. Bounded thread + channel (a wedge\n/// surfaces as a timeout, not a hang). Verbatim sibling of the r4 gate helper.\nfn stall_evict_count(name: &str, deadline: Duration) -> u64 { … }\n\n/// One `KIND_SESSIONS` round-trip, returning the reply. The dispatch runs the\n/// `reap_dead_controller` → time-based `stall_evict_controller`, so polling this\n/// both DRIVES the logical eviction and reads the frozen `resume_seq` cursor.\n/// Bounded thread + channel. Verbatim sibling of the r4 gate helper.\nfn query_sessions(name: &str, deadline: Duration) -> Option<SessionsReply> { … }\n\n// [int->REQ-CONN-BLACKHOLE-LIFECYCLE-HARNESS]\n#[test]\nfn blackholed_controller_lifecycle_five_invariants() { … }","startLine":1},"summary":{"lines":128,"elidedSpans":10,"elidedLines":571},"fileSize":34165,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-daemon\\tests\\conn_blackhole_lifecycle.rs"}}},"isError":false,"timestamp":1788999584218}}
{"type":"message","id":"e8ff0b0b","parentId":"b478984e","timestamp":"2026-09-10T00:19:44.219Z","message":{"role":"toolResult","toolCallId":"call_dBzQK6ug0ipNvJSIUzrDNrk4|fc_0dde72effda0ad21016aa1f7a007e087d28db3fbdc88b7ebc5","toolName":"read","content":[{"type":"text","text":"# Single consolidated daemon with broker/brain split; peer-propagated gated self-update\n\n## Status\n\naccepted (2026-05-29)\n\n## Context\n\nADR-0002 collapsed PTY-hosting and network-hosting into one per-machine `spt-daemon`. Two further forces refine its shape:\n\n1. **Consolidation.** The sister project runs poll listeners as ephemeral per-session background tasks and Psyche wrappers as detached per-live-agent supervisor processes. But poll listeners already interact directly with the agent session (capsule/idle), and Psyche wrappers already invoke harness binaries directly. Once the daemon owns every PTY, keeping these as separate processes is unjustified.\n\n2. **Seamless self-update with a hard no-terminate constraint.** Self-update is a day-one pillar. The constraint: *no endpoint process may terminate or suspend during an spt-core update* — we cannot assume every endpoint can safely suspend. The naive \"drain + restart the daemon\" approach violates this for spt-hosted sessions (the daemon owns their PTY; killing the daemon SIGHUPs the child).\n\n## Decision\n\n**Consolidate all per-machine logic into the one daemon.** Poll-listener logic and Psyche/pulse loops move into the `spt-daemon` — no separate listener or wrapper processes. The only residue is a thin, stateless **in-session relay** for harness-hosted sessions (topology 1), where spt cannot reach into a process tree it doesn't own; the relay just pipes the daemon's events into the session's stdout and is freely killable.\n\n**Split the logical daemon into two implementation layers** to satisfy the no-terminate invariant:\n- **broker** (stable kernel) — holds *only* un-transferable, must-not-die resources: PTY master fds, spawned harness child processes, listening network sockets. Minimal, versioned local IPC. Almost never updates.\n- **daemon brain** (userspace) — all logic (routing, registry, pulse/psyche loops, manifest parsing, update orchestration). Restarts freely; rehydrates from disk state and re-attaches to broker-held handles.\n\nRoutine self-updates replace only the brain → endpoints never notice. The broker updates rarely (IPC-contract change, held-resource-type change, OS PTY/socket API change, or broker bugfix); the IPC is versioned so a newer brain talks to an older broker. Logical addressing is unchanged — still one per-machine `spt-daemon`; the broker is internal, not separately addressable.\n\nRejected alternatives for the invariant: **drain + restart** (violates no-terminate for spt-hosted sessions); **whole-daemon live FD-passing** (ConPTY handle transfer mid-swap is hard and platform-divergent — pushing that complexity into the rare broker update instead of every routine update is strictly better); **topology-1-only invariant with spt-hosted as fast-follow** (silently breaks the stated hard constraint for the exact new topology spt-core introduces).\n\n**Self-update delivery:** peer-propagated over P2P, layered on self-fetch, out-of-band still supported. All binaries signature-verified before handoff (spt-core's own release key) regardless of source — peer-propagation otherwise lets one compromised node poison the subnet. spt-core conducts updates for the whole stack: self first, then ripple-update each registered adapter via the adapter manifest's update declaration (file-pull or delegated command). The plugin's role shrinks to initial bootstrap only.\n\n**Cadence/consent:** not fully automatic by default; gated on user confirmation delivered to the most-recently-active live session, with an opt-in full-auto choice.\n\n## Consequences\n\n- The daemon is the single brain for a machine; crash-recovery and update logic must cover PTYs, networking, registry, spools, listeners, and psyche loops together.\n- A small internal broker process exists beneath the daemon — a deliberate, bounded walk-back of \"literally one process,\" preserving B1's *intent* (one network identity, one supervisor, one firewall prompt) while guaranteeing endpoint survival across updates.\n- Peer-propagated updates make release signing mandatory, not optional.\n- spt-core becomes the update conductor for adapters too; adapter manifests must declare an update avenue.\n- The \"deliver to most-recently-active session\" mechanism is a v1 building block that the deferred PresenceChannel will later generalize.\n- Whole-daemon live FD-passing (zero-interruption even for broker updates) remains a possible future polish but is explicitly not required for v1.\n\n## Validation & amendments (2026-05-31 — Stage A red-team + Spike #1)\n\nCodex adversarial review (`docs/reviews/STAGE-A-codex-redteam.md`) put 4 FATAL findings on this ADR; Spike #1 (`docs/spikes/SPIKE-01-broker-handoff.md`, `spt-spikes/spike-01-broker-handoff`) tested the hardest path. Resulting binding amendments:\n\n### A. Update-class taxonomy (resolves FATAL #1 — R-UPD-3 self-contradiction)\nR-UPD-3's \"no endpoint terminates during update\" is **only** absolute for the *brain-only* class. Three classes, each with its own invariant:\n\n[Showing lines 1-45 of 78. Use :46 to continue]"}],"details":{"truncation":{"content":"# Single consolidated daemon with broker/brain split; peer-propagated gated self-update\r\n\r\n## Status\r\n\r\naccepted (2026-05-29)\r\n\r\n## Context\r\n\r\nADR-0002 collapsed PTY-hosting and network-hosting into one per-machine `spt-daemon`. Two further forces refine its shape:\r\n\r\n1. **Consolidation.** The sister project runs poll listeners as ephemeral per-session background tasks and Psyche wrappers as detached per-live-agent supervisor processes. But poll listeners already interact directly with the agent session (capsule/idle), and Psyche wrappers already invoke harness binaries directly. Once the daemon owns every PTY, keeping these as separate processes is unjustified.\r\n\r\n2. **Seamless self-update with a hard no-terminate constraint.** Self-update is a day-one pillar. The constraint: *no endpoint process may terminate or suspend during an spt-core update* — we cannot assume every endpoint can safely suspend. The naive \"drain + restart the daemon\" approach violates this for spt-hosted sessions (the daemon owns their PTY; killing the daemon SIGHUPs the child).\r\n\r\n## Decision\r\n\r\n**Consolidate all per-machine logic into the one daemon.** Poll-listener logic and Psyche/pulse loops move into the `spt-daemon` — no separate listener or wrapper processes. The only residue is a thin, stateless **in-session relay** for harness-hosted sessions (topology 1), where spt cannot reach into a process tree it doesn't own; the relay just pipes the daemon's events into the session's stdout and is freely killable.\r\n\r\n**Split the logical daemon into two implementation layers** to satisfy the no-terminate invariant:\r\n- **broker** (stable kernel) — holds *only* un-transferable, must-not-die resources: PTY master fds, spawned harness child processes, listening network sockets. Minimal, versioned local IPC. Almost never updates.\r\n- **daemon brain** (userspace) — all logic (routing, registry, pulse/psyche loops, manifest parsing, update orchestration). Restarts freely; rehydrates from disk state and re-attaches to broker-held handles.\r\n\r\nRoutine self-updates replace only the brain → endpoints never notice. The broker updates rarely (IPC-contract change, held-resource-type change, OS PTY/socket API change, or broker bugfix); the IPC is versioned so a newer brain talks to an older broker. Logical addressing is unchanged — still one per-machine `spt-daemon`; the broker is internal, not separately addressable.\r\n\r\nRejected alternatives for the invariant: **drain + restart** (violates no-terminate for spt-hosted sessions); **whole-daemon live FD-passing** (ConPTY handle transfer mid-swap is hard and platform-divergent — pushing that complexity into the rare broker update instead of every routine update is strictly better); **topology-1-only invariant with spt-hosted as fast-follow** (silently breaks the stated hard constraint for the exact new topology spt-core introduces).\r\n\r\n**Self-update delivery:** peer-propagated over P2P, layered on self-fetch, out-of-band still supported. All binaries signature-verified before handoff (spt-core's own release key) regardless of source — peer-propagation otherwise lets one compromised node poison the subnet. spt-core conducts updates for the whole stack: self first, then ripple-update each registered adapter via the adapter manifest's update declaration (file-pull or delegated command). The plugin's role shrinks to initial bootstrap only.\r\n\r\n**Cadence/consent:** not fully automatic by default; gated on user confirmation delivered to the most-recently-active live session, with an opt-in full-auto choice.\r\n\r\n## Consequences\r\n\r\n- The daemon is the single brain for a machine; crash-recovery and update logic must cover PTYs, networking, registry, spools, listeners, and psyche loops together.\r\n- A small internal broker process exists beneath the daemon — a deliberate, bounded walk-back of \"literally one process,\" preserving B1's *intent* (one network identity, one supervisor, one firewall prompt) while guaranteeing endpoint survival across updates.\r\n- Peer-propagated updates make release signing mandatory, not optional.\r\n- spt-core becomes the update conductor for adapters too; adapter manifests must declare an update avenue.\r\n- The \"deliver to most-recently-active session\" mechanism is a v1 building block that the deferred PresenceChannel will later generalize.\r\n- Whole-daemon live FD-passing (zero-interruption even for broker updates) remains a possible future polish but is explicitly not required for v1.\r\n\r\n## Validation & amendments (2026-05-31 — Stage A red-team + Spike #1)\r\n\r\nCodex adversarial review (`docs/reviews/STAGE-A-codex-redteam.md`) put 4 FATAL findings on this ADR; Spike #1 (`docs/spikes/SPIKE-01-broker-handoff.md`, `spt-spikes/spike-01-broker-handoff`) tested the hardest path. Resulting binding amendments:\r\n\r\n### A. Update-class taxonomy (resolves FATAL #1 — R-UPD-3 self-contradiction)\r\nR-UPD-3's \"no endpoint terminates during update\" is **only** absolute for the *brain-only* class. Three classes, each with its own invariant:\r","truncated":true,"truncatedBy":"lines","totalLines":78,"totalBytes":5051,"outputLines":45,"outputBytes":5051,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":78,"displayContent":{"text":"# Single consolidated daemon with broker/brain split; peer-propagated gated self-update\r\n\r\n## Status\r\n\r\naccepted (2026-05-29)\r\n\r\n## Context\r\n\r\nADR-0002 collapsed PTY-hosting and network-hosting into one per-machine `spt-daemon`. Two further forces refine its shape:\r\n\r\n1. **Consolidation.** The sister project runs poll listeners as ephemeral per-session background tasks and Psyche wrappers as detached per-live-agent supervisor processes. But poll listeners already interact directly with the agent session (capsule/idle), and Psyche wrappers already invoke harness binaries directly. Once the daemon owns every PTY, keeping these as separate processes is unjustified.\r\n\r\n2. **Seamless self-update with a hard no-terminate constraint.** Self-update is a day-one pillar. The constraint: *no endpoint process may terminate or suspend during an spt-core update* — we cannot assume every endpoint can safely suspend. The naive \"drain + restart the daemon\" approach violates this for spt-hosted sessions (the daemon owns their PTY; killing the daemon SIGHUPs the child).\r\n\r\n## Decision\r\n\r\n**Consolidate all per-machine logic into the one daemon.** Poll-listener logic and Psyche/pulse loops move into the `spt-daemon` — no separate listener or wrapper processes. The only residue is a thin, stateless **in-session relay** for harness-hosted sessions (topology 1), where spt cannot reach into a process tree it doesn't own; the relay just pipes the daemon's events into the session's stdout and is freely killable.\r\n\r\n**Split the logical daemon into two implementation layers** to satisfy the no-terminate invariant:\r\n- **broker** (stable kernel) — holds *only* un-transferable, must-not-die resources: PTY master fds, spawned harness child processes, listening network sockets. Minimal, versioned local IPC. Almost never updates.\r\n- **daemon brain** (userspace) — all logic (routing, registry, pulse/psyche loops, manifest parsing, update orchestration). Restarts freely; rehydrates from disk state and re-attaches to broker-held handles.\r\n\r\nRoutine self-updates replace only the brain → endpoints never notice. The broker updates rarely (IPC-contract change, held-resource-type change, OS PTY/socket API change, or broker bugfix); the IPC is versioned so a newer brain talks to an older broker. Logical addressing is unchanged — still one per-machine `spt-daemon`; the broker is internal, not separately addressable.\r\n\r\nRejected alternatives for the invariant: **drain + restart** (violates no-terminate for spt-hosted sessions); **whole-daemon live FD-passing** (ConPTY handle transfer mid-swap is hard and platform-divergent — pushing that complexity into the rare broker update instead of every routine update is strictly better); **topology-1-only invariant with spt-hosted as fast-follow** (silently breaks the stated hard constraint for the exact new topology spt-core introduces).\r\n\r\n**Self-update delivery:** peer-propagated over P2P, layered on self-fetch, out-of-band still supported. All binaries signature-verified before handoff (spt-core's own release key) regardless of source — peer-propagation otherwise lets one compromised node poison the subnet. spt-core conducts updates for the whole stack: self first, then ripple-update each registered adapter via the adapter manifest's update declaration (file-pull or delegated command). The plugin's role shrinks to initial bootstrap only.\r\n\r\n**Cadence/consent:** not fully automatic by default; gated on user confirmation delivered to the most-recently-active live session, with an opt-in full-auto choice.\r\n\r\n## Consequences\r\n\r\n- The daemon is the single brain for a machine; crash-recovery and update logic must cover PTYs, networking, registry, spools, listeners, and psyche loops together.\r\n- A small internal broker process exists beneath the daemon — a deliberate, bounded walk-back of \"literally one process,\" preserving B1's *intent* (one network identity, one supervisor, one firewall prompt) while guaranteeing endpoint survival across updates.\r\n- Peer-propagated updates make release signing mandatory, not optional.\r\n- spt-core becomes the update conductor for adapters too; adapter manifests must declare an update avenue.\r\n- The \"deliver to most-recently-active session\" mechanism is a v1 building block that the deferred PresenceChannel will later generalize.\r\n- Whole-daemon live FD-passing (zero-interruption even for broker updates) remains a possible future polish but is explicitly not required for v1.\r\n\r\n## Validation & amendments (2026-05-31 — Stage A red-team + Spike #1)\r\n\r\nCodex adversarial review (`docs/reviews/STAGE-A-codex-redteam.md`) put 4 FATAL findings on this ADR; Spike #1 (`docs/spikes/SPIKE-01-broker-handoff.md`, `spt-spikes/spike-01-broker-handoff`) tested the hardest path. Resulting binding amendments:\r\n\r\n### A. Update-class taxonomy (resolves FATAL #1 — R-UPD-3 self-contradiction)\r\nR-UPD-3's \"no endpoint terminates during update\" is **only** absolute for the *brain-only* class. Three classes, each with its own invariant:\r","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45]},"fileSize":10463,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\docs\\adr\\0004-single-daemon-broker-brain-split-and-self-update.md"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":78,"totalBytes":5051,"outputLines":45,"outputBytes":5051,"shownRange":{"start":1,"end":45},"nextOffset":46}}},"isError":false,"timestamp":1788999584219}}
{"type":"message","id":"1f91066a","parentId":"e8ff0b0b","timestamp":"2026-09-10T00:19:44.222Z","message":{"role":"toolResult","toolCallId":"call_yQ5v4ngdG3LVvysjf5PLaogH|fc_0dde72effda0ad21016aa1f7a007e887d28997dd6d806c2a20","toolName":"read","content":[{"type":"text","text":"//! The OS-split local-IPC carrier: named pipe (Windows) / Unix-domain socket\n//! (unix), behind one [`DaemonTransport`] trait.\n//!\n//! This mirrors how `spt-term`'s `SessionSurface` abstracts ConPTY vs forkpty:\n//! the broker and brain talk through one trait, and the platform split lives in\n//! a single implementation ([`LocalSocketTransport`], over the `interprocess`\n//! crate's local sockets — namespaced names work on both Windows and Linux).\n//! The byte framing ([`crate::codec`]) and the version handshake\n//! ([`crate::frame`]) ride on top, transport-agnostic.\n// [impl->REQ-DAEMON-2]\n\nuse std::io::{self, Read, Write};\n\nuse interprocess::local_socket::{prelude::*, GenericNamespaced, ListenerOptions, Stream};\n\nuse crate::codec::{read_frame, write_frame};\nuse crate::frame::{accept_hello, Envelope, HandshakeError, Hello, Role};\n\n/// A bidirectional IPC connection between the broker and a brain. Anything that\n/// is both [`Read`] and [`Write`] qualifies, so tests can substitute an\n/// in-memory or socket-pair duplex for the real local socket.\npub trait IpcConn: Read + Write {}\nimpl<T: Read + Write> IpcConn for T {}\n\n/// A listening IPC endpoint plus its dialer. The broker `bind`s one and\n/// `accept`s brains; a brain `connect`s. Abstracted as a trait so B1/B2 can\n/// drive the broker over a stub transport in tests without a real socket.\npub trait DaemonTransport {\n    /// The accepted/connected connection type.\n    type Conn: IpcConn;\n\n    /// Accept the next inbound connection (broker side). Blocks.\n    fn accept(&self) -> io::Result<Self::Conn>;\n}\n\n/// The production transport: `interprocess` local sockets. A namespaced name\n/// resolves to a Windows named pipe or a Linux abstract-namespace UDS — the one\n/// OS split, contained here.\npub struct LocalSocketTransport {\n    listener: interprocess::local_socket::Listener,\n}\n\nimpl LocalSocketTransport {\n    /// Bind a listening endpoint under the namespaced `name` (broker side).\n    pub fn bind(name: &str) -> io::Result<Self> {\n        let ns = name.to_ns_name::<GenericNamespaced>()?;\n        let listener = ListenerOptions::new().name(ns).create_sync()?;\n        Ok(LocalSocketTransport { listener })\n    }\n\n    /// Connect to an existing endpoint under `name` (brain side).\n    pub fn connect(name: &str) -> io::Result<Stream> {\n        let ns = name.to_ns_name::<GenericNamespaced>()?;\n        Stream::connect(ns)\n    }\n}\n\nimpl DaemonTransport for LocalSocketTransport {\n    type Conn = Stream;\n\n    fn accept(&self) -> io::Result<Self::Conn> {\n        self.listener.accept()\n    }\n}\n\n/// Brain side of the connect handshake: announce ourselves as `role` by sending\n/// a `hello` frame. The broker answers with [`recv_hello`] + [`accept_hello`].\n// [impl->REQ-HAZARD-HANDOFF-ARGV-COMPAT]\npub fn send_hello<W: Write>(w: &mut W, role: Role) -> io::Result<()> {\n    let hello = Hello::current(role);\n    let payload =\n        serde_json::to_value(hello).map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n    write_frame(w, &Envelope::new(\"hello\", payload))\n}\n\n/// Broker side: read the peer's first frame, require it to be a `hello`, and\n/// run the version-tolerance policy ([`accept_hello`]) against `expected_peer`.\n/// An incompatible peer surfaces a typed [`HandshakeError`] (never a silent\n/// misparse — the 2.3 failure mode is a peer dying unlogged on a mismatch).\n// [impl->REQ-HAZARD-HANDOFF-ARGV-COMPAT]\npub fn recv_hello<R: Read>(r: &mut R, expected_peer: Role) -> io::Result<Hello> {\n    let env = read_frame(r)?;\n    if env.kind != \"hello\" {\n        return Err(io::Error::new(\n            io::ErrorKind::InvalidData,\n            format!(\"expected hello frame, got kind {:?}\", env.kind),\n        ));\n    }\n    let hello: Hello = serde_json::from_value(env.payload)\n        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n    accept_hello(&hello, expected_peer)\n        .map_err(|e: HandshakeError| io::Error::new(io::ErrorKind::InvalidData, e))?;\n    Ok(hello)\n}\n\n#[cfg(test)]\nmod tests {\n    use super::*;\n    use serde_json::json;\n    use std::sync::atomic::{AtomicU32, Ordering};\n\n    // Unique namespaced socket name per test, so parallel test threads on one\n    // machine never collide on the same pipe/UDS name.\n    static SEQ: AtomicU32 = AtomicU32::new(0);\n    fn unique_name() -> String {\n        let n = SEQ.fetch_add(1, Ordering::Relaxed);\n        format!(\"spt-daemon-test-{}-{}.sock\", std::process::id(), n)\n    }\n\n    // [unit->REQ-DAEMON-2] a real local socket carries a hello handshake then a\n    // frame, end to end, on whatever OS the test runs on (named pipe / UDS).\n    #[test]\n    fn local_socket_handshake_and_frame_round_trip() {\n        let name = unique_name();\n        let broker = LocalSocketTransport::bind(&name).unwrap();\n\n        let client_name = name.clone();\n        let client = std::thread::spawn(move || {\n            let mut conn = LocalSocketTransport::connect(&client_name).unwrap();\n            send_hello(&mut conn, Role::Brain).unwrap();\n            write_frame(&mut conn, &Envelope::new(\"ping\", json!({ \"seq\": 1 }))).unwrap();\n            // Read the broker's reply.\n            read_frame(&mut conn).unwrap()\n        });\n\n        let mut server = broker.accept().unwrap();\n…\n    }\n…\n}\n\n[Showing lines 1-126 of 156. Use :127 to continue]"}],"details":{"truncation":{"content":"//! The OS-split local-IPC carrier: named pipe (Windows) / Unix-domain socket\r\n//! (unix), behind one [`DaemonTransport`] trait.\r\n//!\r\n//! This mirrors how `spt-term`'s `SessionSurface` abstracts ConPTY vs forkpty:\r\n//! the broker and brain talk through one trait, and the platform split lives in\r\n//! a single implementation ([`LocalSocketTransport`], over the `interprocess`\r\n//! crate's local sockets — namespaced names work on both Windows and Linux).\r\n//! The byte framing ([`crate::codec`]) and the version handshake\r\n//! ([`crate::frame`]) ride on top, transport-agnostic.\r\n// [impl->REQ-DAEMON-2]\r\n\r\nuse std::io::{self, Read, Write};\r\n\r\nuse interprocess::local_socket::{prelude::*, GenericNamespaced, ListenerOptions, Stream};\r\n\r\nuse crate::codec::{read_frame, write_frame};\r\nuse crate::frame::{accept_hello, Envelope, HandshakeError, Hello, Role};\r\n\r\n/// A bidirectional IPC connection between the broker and a brain. Anything that\r\n/// is both [`Read`] and [`Write`] qualifies, so tests can substitute an\r\n/// in-memory or socket-pair duplex for the real local socket.\r\npub trait IpcConn: Read + Write {}\r\nimpl<T: Read + Write> IpcConn for T {}\r\n\r\n/// A listening IPC endpoint plus its dialer. The broker `bind`s one and\r\n/// `accept`s brains; a brain `connect`s. Abstracted as a trait so B1/B2 can\r\n/// drive the broker over a stub transport in tests without a real socket.\r\npub trait DaemonTransport {\r\n    /// The accepted/connected connection type.\r\n    type Conn: IpcConn;\r\n\r\n    /// Accept the next inbound connection (broker side). Blocks.\r\n    fn accept(&self) -> io::Result<Self::Conn>;\r\n}\r\n\r\n/// The production transport: `interprocess` local sockets. A namespaced name\r\n/// resolves to a Windows named pipe or a Linux abstract-namespace UDS — the one\r\n/// OS split, contained here.\r\npub struct LocalSocketTransport {\r\n    listener: interprocess::local_socket::Listener,\r\n}\r\n\r\nimpl LocalSocketTransport {\r\n    /// Bind a listening endpoint under the namespaced `name` (broker side).\r\n    pub fn bind(name: &str) -> io::Result<Self> {\r\n        let ns = name.to_ns_name::<GenericNamespaced>()?;\r\n        let listener = ListenerOptions::new().name(ns).create_sync()?;\r\n        Ok(LocalSocketTransport { listener })\r\n    }\r\n\r\n    /// Connect to an existing endpoint under `name` (brain side).\r\n    pub fn connect(name: &str) -> io::Result<Stream> {\r\n        let ns = name.to_ns_name::<GenericNamespaced>()?;\r\n        Stream::connect(ns)\r\n    }\r\n}\r\n\r\nimpl DaemonTransport for LocalSocketTransport {\r\n    type Conn = Stream;\r\n\r\n    fn accept(&self) -> io::Result<Self::Conn> {\r\n        self.listener.accept()\r\n    }\r\n}\r\n\r\n/// Brain side of the connect handshake: announce ourselves as `role` by sending\r\n/// a `hello` frame. The broker answers with [`recv_hello`] + [`accept_hello`].\r\n// [impl->REQ-HAZARD-HANDOFF-ARGV-COMPAT]\r\npub fn send_hello<W: Write>(w: &mut W, role: Role) -> io::Result<()> {\r\n    let hello = Hello::current(role);\r\n    let payload =\r\n        serde_json::to_value(hello).map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\r\n    write_frame(w, &Envelope::new(\"hello\", payload))\r\n}\r\n\r\n/// Broker side: read the peer's first frame, require it to be a `hello`, and\r\n/// run the version-tolerance policy ([`accept_hello`]) against `expected_peer`.\r\n/// An incompatible peer surfaces a typed [`HandshakeError`] (never a silent\r\n/// misparse — the 2.3 failure mode is a peer dying unlogged on a mismatch).\r\n// [impl->REQ-HAZARD-HANDOFF-ARGV-COMPAT]\r\npub fn recv_hello<R: Read>(r: &mut R, expected_peer: Role) -> io::Result<Hello> {\r\n    let env = read_frame(r)?;\r\n    if env.kind != \"hello\" {\r\n        return Err(io::Error::new(\r\n            io::ErrorKind::InvalidData,\r\n            format!(\"expected hello frame, got kind {:?}\", env.kind),\r\n        ));\r\n    }\r\n    let hello: Hello = serde_json::from_value(env.payload)\r\n        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\r\n    accept_hello(&hello, expected_peer)\r\n        .map_err(|e: HandshakeError| io::Error::new(io::ErrorKind::InvalidData, e))?;\r\n    Ok(hello)\r\n}\r\n\r\n#[cfg(test)]\r\nmod tests {\r\n    use super::*;\r\n    use serde_json::json;\r\n    use std::sync::atomic::{AtomicU32, Ordering};\r\n\r\n    // Unique namespaced socket name per test, so parallel test threads on one\r\n    // machine never collide on the same pipe/UDS name.\r\n    static SEQ: AtomicU32 = AtomicU32::new(0);\r\n    fn unique_name() -> String {\r\n        let n = SEQ.fetch_add(1, Ordering::Relaxed);\r\n        format!(\"spt-daemon-test-{}-{}.sock\", std::process::id(), n)\r\n    }\r\n\r\n    // [unit->REQ-DAEMON-2] a real local socket carries a hello handshake then a\r\n    // frame, end to end, on whatever OS the test runs on (named pipe / UDS).\r\n    #[test]\r\n    fn local_socket_handshake_and_frame_round_trip() {\r\n        let name = unique_name();\r\n        let broker = LocalSocketTransport::bind(&name).unwrap();\r\n\r\n        let client_name = name.clone();\r\n        let client = std::thread::spawn(move || {\r\n            let mut conn = LocalSocketTransport::connect(&client_name).unwrap();\r\n            send_hello(&mut conn, Role::Brain).unwrap();\r\n            write_frame(&mut conn, &Envelope::new(\"ping\", json!({ \"seq\": 1 }))).unwrap();\r\n            // Read the broker's reply.\r\n            read_frame(&mut conn).unwrap()\r\n        });\r\n\r\n        let mut server = broker.accept().unwrap();\r","truncated":true,"truncatedBy":"lines","totalLines":156,"totalBytes":5371,"outputLines":126,"outputBytes":5371,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":156,"displayContent":{"text":"//! The OS-split local-IPC carrier: named pipe (Windows) / Unix-domain socket\r\n//! (unix), behind one [`DaemonTransport`] trait.\r\n//!\r\n//! This mirrors how `spt-term`'s `SessionSurface` abstracts ConPTY vs forkpty:\r\n//! the broker and brain talk through one trait, and the platform split lives in\r\n//! a single implementation ([`LocalSocketTransport`], over the `interprocess`\r\n//! crate's local sockets — namespaced names work on both Windows and Linux).\r\n//! The byte framing ([`crate::codec`]) and the version handshake\r\n//! ([`crate::frame`]) ride on top, transport-agnostic.\r\n// [impl->REQ-DAEMON-2]\r\n\r\nuse std::io::{self, Read, Write};\r\n\r\nuse interprocess::local_socket::{prelude::*, GenericNamespaced, ListenerOptions, Stream};\r\n\r\nuse crate::codec::{read_frame, write_frame};\r\nuse crate::frame::{accept_hello, Envelope, HandshakeError, Hello, Role};\r\n\r\n/// A bidirectional IPC connection between the broker and a brain. Anything that\r\n/// is both [`Read`] and [`Write`] qualifies, so tests can substitute an\r\n/// in-memory or socket-pair duplex for the real local socket.\r\npub trait IpcConn: Read + Write {}\r\nimpl<T: Read + Write> IpcConn for T {}\r\n\r\n/// A listening IPC endpoint plus its dialer. The broker `bind`s one and\r\n/// `accept`s brains; a brain `connect`s. Abstracted as a trait so B1/B2 can\r\n/// drive the broker over a stub transport in tests without a real socket.\r\npub trait DaemonTransport {\r\n    /// The accepted/connected connection type.\r\n    type Conn: IpcConn;\r\n\r\n    /// Accept the next inbound connection (broker side). Blocks.\r\n    fn accept(&self) -> io::Result<Self::Conn>;\r\n}\r\n\r\n/// The production transport: `interprocess` local sockets. A namespaced name\r\n/// resolves to a Windows named pipe or a Linux abstract-namespace UDS — the one\r\n/// OS split, contained here.\r\npub struct LocalSocketTransport {\r\n    listener: interprocess::local_socket::Listener,\r\n}\r\n\r\nimpl LocalSocketTransport {\r\n    /// Bind a listening endpoint under the namespaced `name` (broker side).\r\n    pub fn bind(name: &str) -> io::Result<Self> {\r\n        let ns = name.to_ns_name::<GenericNamespaced>()?;\r\n        let listener = ListenerOptions::new().name(ns).create_sync()?;\r\n        Ok(LocalSocketTransport { listener })\r\n    }\r\n\r\n    /// Connect to an existing endpoint under `name` (brain side).\r\n    pub fn connect(name: &str) -> io::Result<Stream> {\r\n        let ns = name.to_ns_name::<GenericNamespaced>()?;\r\n        Stream::connect(ns)\r\n    }\r\n}\r\n\r\nimpl DaemonTransport for LocalSocketTransport {\r\n    type Conn = Stream;\r\n\r\n    fn accept(&self) -> io::Result<Self::Conn> {\r\n        self.listener.accept()\r\n    }\r\n}\r\n\r\n/// Brain side of the connect handshake: announce ourselves as `role` by sending\r\n/// a `hello` frame. The broker answers with [`recv_hello`] + [`accept_hello`].\r\n// [impl->REQ-HAZARD-HANDOFF-ARGV-COMPAT]\r\npub fn send_hello<W: Write>(w: &mut W, role: Role) -> io::Result<()> {\r\n    let hello = Hello::current(role);\r\n    let payload =\r\n        serde_json::to_value(hello).map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\r\n    write_frame(w, &Envelope::new(\"hello\", payload))\r\n}\r\n\r\n/// Broker side: read the peer's first frame, require it to be a `hello`, and\r\n/// run the version-tolerance policy ([`accept_hello`]) against `expected_peer`.\r\n/// An incompatible peer surfaces a typed [`HandshakeError`] (never a silent\r\n/// misparse — the 2.3 failure mode is a peer dying unlogged on a mismatch).\r\n// [impl->REQ-HAZARD-HANDOFF-ARGV-COMPAT]\r\npub fn recv_hello<R: Read>(r: &mut R, expected_peer: Role) -> io::Result<Hello> {\r\n    let env = read_frame(r)?;\r\n    if env.kind != \"hello\" {\r\n        return Err(io::Error::new(\r\n            io::ErrorKind::InvalidData,\r\n            format!(\"expected hello frame, got kind {:?}\", env.kind),\r\n        ));\r\n    }\r\n    let hello: Hello = serde_json::from_value(env.payload)\r\n        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\r\n    accept_hello(&hello, expected_peer)\r\n        .map_err(|e: HandshakeError| io::Error::new(io::ErrorKind::InvalidData, e))?;\r\n    Ok(hello)\r\n}\r\n\r\n#[cfg(test)]\r\nmod tests {\r\n    use super::*;\r\n    use serde_json::json;\r\n    use std::sync::atomic::{AtomicU32, Ordering};\r\n\r\n    // Unique namespaced socket name per test, so parallel test threads on one\r\n    // machine never collide on the same pipe/UDS name.\r\n    static SEQ: AtomicU32 = AtomicU32::new(0);\r\n    fn unique_name() -> String {\r\n        let n = SEQ.fetch_add(1, Ordering::Relaxed);\r\n        format!(\"spt-daemon-test-{}-{}.sock\", std::process::id(), n)\r\n    }\r\n\r\n    // [unit->REQ-DAEMON-2] a real local socket carries a hello handshake then a\r\n    // frame, end to end, on whatever OS the test runs on (named pipe / UDS).\r\n    #[test]\r\n    fn local_socket_handshake_and_frame_round_trip() {\r\n        let name = unique_name();\r\n        let broker = LocalSocketTransport::bind(&name).unwrap();\r\n\r\n        let client_name = name.clone();\r\n        let client = std::thread::spawn(move || {\r\n            let mut conn = LocalSocketTransport::connect(&client_name).unwrap();\r\n            send_hello(&mut conn, Role::Brain).unwrap();\r\n            write_frame(&mut conn, &Envelope::new(\"ping\", json!({ \"seq\": 1 }))).unwrap();\r\n            // Read the broker's reply.\r\n            read_frame(&mut conn).unwrap()\r\n        });\r\n\r\n        let mut server = broker.accept().unwrap();\r\n…\n    }\n…\n}","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120,121,122,123,124,125,126,null,135,null,156]},"fileSize":6625,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-daemon\\src\\transport.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":156,"totalBytes":5371,"outputLines":126,"outputBytes":5371,"shownRange":{"start":1,"end":126},"nextOffset":127}}},"isError":false,"timestamp":1788999584222}}
{"type":"message","id":"2c134fb6","parentId":"1f91066a","timestamp":"2026-09-10T00:19:44.224Z","message":{"role":"toolResult","toolCallId":"call_H81TcWUzZzTRJSYvgdSphpBV|fc_0dde72effda0ad21016aa1f7a007d887d2919b64622734e05f","toolName":"read","content":[{"type":"text","text":"# Changelog\n\nAll notable **user-facing** changes to `spt` — what a person running the CLI\nnotices or does differently. The `## [<version>]` section of each release\nbecomes that release's GitHub Release notes verbatim (see\n`docs/RELEASE-RUNBOOK.md`). This project follows\n[Keep a Changelog](https://keepachangelog.com) and semantic versioning.\nPre-1.0, choose the bump by what a user can notice: **minor** when a release\nbreaks something, or changes the observable behavior of existing surfaces\nbroadly; **patch** for fixes, and for additive opt-in capability — a new key,\nflag, or page that no existing user can encounter without opting into it.\n\n## [Unreleased]\n\n### Fixed\n\n<!-- [doc->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED] -->\n- Daemon logs summarize healthy IPC connection starts and closes once per\n  minute instead of logging each one. Failures and subscriber lifecycle\n  events keep their per-connection attribution, including the original\n  first-write time. Livehost session polling now reuses its healthy\n  connection and names the reason when it opens a replacement.\n\n## [0.68.0] - 2026-09-08\n\nWeb serving. Files, directories, adapter documentation and the changelog are\nreachable at node-prefixed addresses on the local server, and an address that\nnames another machine in the subnet is answered by the machine that owns it.\nMessages can carry attachments the receiver pulls on demand, every message\ngets a short ID to show or reply to, and a machine with no spt on it can be\nhanded the binary over the local network. The `XFER` access surface is retired.\n\n### Added\n\n<!-- [doc->REQ-WEB-SERVING-REGISTRY] -->\n- `spt serve add`, `spt serve rm`, and `spt serve list --json` manage live file\n  and directory references. Same-name registrations receive stable numbered\n  suffixes; removing an entry never deletes its source. Only the same absolute\n  path and kind may reclaim a retired name.\n- Every registered adapter gets a core-owned `web/` output directory.\n  `[adapter].web_short_path` optionally gives it a short URL alias. Removing\n  an adapter stops serving its output without deleting the files.\n- `spt adapter add` names each manifest key it does not know on stderr\n  (`manifest: unknown key [<table>].<key> (ignored)`) and still registers;\n  a misspelled optional key is no longer visible only by its absence.\n\n[Showing lines 1-45 of 3127. Use :46 to continue]"}],"details":{"truncation":{"content":"# Changelog\r\n\r\nAll notable **user-facing** changes to `spt` — what a person running the CLI\r\nnotices or does differently. The `## [<version>]` section of each release\r\nbecomes that release's GitHub Release notes verbatim (see\r\n`docs/RELEASE-RUNBOOK.md`). This project follows\r\n[Keep a Changelog](https://keepachangelog.com) and semantic versioning.\r\nPre-1.0, choose the bump by what a user can notice: **minor** when a release\r\nbreaks something, or changes the observable behavior of existing surfaces\r\nbroadly; **patch** for fixes, and for additive opt-in capability — a new key,\r\nflag, or page that no existing user can encounter without opting into it.\r\n\r\n## [Unreleased]\r\n\r\n### Fixed\r\n\r\n<!-- [doc->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED] -->\r\n- Daemon logs summarize healthy IPC connection starts and closes once per\r\n  minute instead of logging each one. Failures and subscriber lifecycle\r\n  events keep their per-connection attribution, including the original\r\n  first-write time. Livehost session polling now reuses its healthy\r\n  connection and names the reason when it opens a replacement.\r\n\r\n## [0.68.0] - 2026-09-08\r\n\r\nWeb serving. Files, directories, adapter documentation and the changelog are\r\nreachable at node-prefixed addresses on the local server, and an address that\r\nnames another machine in the subnet is answered by the machine that owns it.\r\nMessages can carry attachments the receiver pulls on demand, every message\r\ngets a short ID to show or reply to, and a machine with no spt on it can be\r\nhanded the binary over the local network. The `XFER` access surface is retired.\r\n\r\n### Added\r\n\r\n<!-- [doc->REQ-WEB-SERVING-REGISTRY] -->\r\n- `spt serve add`, `spt serve rm`, and `spt serve list --json` manage live file\r\n  and directory references. Same-name registrations receive stable numbered\r\n  suffixes; removing an entry never deletes its source. Only the same absolute\r\n  path and kind may reclaim a retired name.\r\n- Every registered adapter gets a core-owned `web/` output directory.\r\n  `[adapter].web_short_path` optionally gives it a short URL alias. Removing\r\n  an adapter stops serving its output without deleting the files.\r\n- `spt adapter add` names each manifest key it does not know on stderr\r\n  (`manifest: unknown key [<table>].<key> (ignored)`) and still registers;\r\n  a misspelled optional key is no longer visible only by its absence.\r","truncated":true,"truncatedBy":"lines","totalLines":3127,"totalBytes":2373,"outputLines":45,"outputBytes":2373,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":3127,"displayContent":{"text":"# Changelog\r\n\r\nAll notable **user-facing** changes to `spt` — what a person running the CLI\r\nnotices or does differently. The `## [<version>]` section of each release\r\nbecomes that release's GitHub Release notes verbatim (see\r\n`docs/RELEASE-RUNBOOK.md`). This project follows\r\n[Keep a Changelog](https://keepachangelog.com) and semantic versioning.\r\nPre-1.0, choose the bump by what a user can notice: **minor** when a release\r\nbreaks something, or changes the observable behavior of existing surfaces\r\nbroadly; **patch** for fixes, and for additive opt-in capability — a new key,\r\nflag, or page that no existing user can encounter without opting into it.\r\n\r\n## [Unreleased]\r\n\r\n### Fixed\r\n\r\n<!-- [doc->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED] -->\r\n- Daemon logs summarize healthy IPC connection starts and closes once per\r\n  minute instead of logging each one. Failures and subscriber lifecycle\r\n  events keep their per-connection attribution, including the original\r\n  first-write time. Livehost session polling now reuses its healthy\r\n  connection and names the reason when it opens a replacement.\r\n\r\n## [0.68.0] - 2026-09-08\r\n\r\nWeb serving. Files, directories, adapter documentation and the changelog are\r\nreachable at node-prefixed addresses on the local server, and an address that\r\nnames another machine in the subnet is answered by the machine that owns it.\r\nMessages can carry attachments the receiver pulls on demand, every message\r\ngets a short ID to show or reply to, and a machine with no spt on it can be\r\nhanded the binary over the local network. The `XFER` access surface is retired.\r\n\r\n### Added\r\n\r\n<!-- [doc->REQ-WEB-SERVING-REGISTRY] -->\r\n- `spt serve add`, `spt serve rm`, and `spt serve list --json` manage live file\r\n  and directory references. Same-name registrations receive stable numbered\r\n  suffixes; removing an entry never deletes its source. Only the same absolute\r\n  path and kind may reclaim a retired name.\r\n- Every registered adapter gets a core-owned `web/` output directory.\r\n  `[adapter].web_short_path` optionally gives it a short URL alias. Removing\r\n  an adapter stops serving its output without deleting the files.\r\n- `spt adapter add` names each manifest key it does not know on stderr\r\n  (`manifest: unknown key [<table>].<key> (ignored)`) and still registers;\r\n  a misspelled optional key is no longer visible only by its absence.\r","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45]},"fileSize":220261,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\CHANGELOG.md"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":3127,"totalBytes":2373,"outputLines":45,"outputBytes":2373,"shownRange":{"start":1,"end":45},"nextOffset":46}}},"isError":false,"timestamp":1788999584224}}
{"type":"message","id":"32d22b5e","parentId":"2c134fb6","timestamp":"2026-09-10T00:19:44.239Z","message":{"role":"toolResult","toolCallId":"call_0mqcQOcn96wUWhCzdA7Rvd5z|fc_0dde72effda0ad21016aa1f7a007d487d29946f6e7adcfcd4e","toolName":"read","content":[{"type":"text","text":"id = \"REQ-ENDPOINT-AUTOSTART\"\ntitle = \"MSG-IDENTITY W5 / F-038 (flynn operator-directed ask 2026-07-10, SPT-CORE-NEEDS #7 + deployah field-confirm same day: mobile-gw alive=false after the v0.30.6 full daemon restart = this feature's absence, live): an endpoint can be marked a STARTUP DEFAULT so the daemon brings it back up at daemon start — Gateway-class endpoints are infra (the phone treats mobile-gw as always-there; box reboot / daemon cold start currently leaves it down until hands-on). SHAPE RULED (doyle, dispatch): (a) `spt endpoint auto-start <id>` persists the endpoint (id + adapter/profile + args) as a startup default REPLAYED at daemon start, symmetric with the shipped `subnet attach/detach --save` precedent — smallest orthogonal cut, explicit operator intent, no interaction with effective_rest_state/F-035 reader-parity semantics (shape (c) restore-what-was-up REJECTED for now: principled but couples to the rest_state neighborhood that just churned; revisit if --save proves insufficient in the field). A saved endpoint that fails to come up logs loud + does not block daemon start or other replays. flynn docs sweep confirmed missing-feature not docs-gap (rest/wake manual-only; no endpoint analog of subnet --save; no manifest field; no api surface). Gate: int — daemon restart brings a --save'd endpoint back up (fresh daemon, saved default, endpoint reaches its steady state without hands-on); doc — public docs page for the verb (VERSION-scoped); unit — persistence round-trip + replay skip-on-missing-adapter loud. Kin subnet --save (the symmetry precedent), REQ-LIST-JSON-LIVENESS-PARITY + REQ-HAZARD-BIND-REST-STATE-CARRY (the F-035 neighborhood shape (c) would have coupled to), [[spt-core-findings-backlog]] F-038. Interim on flynn's box (logon scheduled task) dissolves when this lands.\"\nrequired_stages = [\"doc\", \"impl\", \"unit\", \"int\"]  # ACTIVATED MSG-IDENTITY W5 build (todlando 2026-07-10, shape (a)). doc = instances/overview.md startup-defaults section + the --save clap help (reference.md regen, drift-gated). impl = StartupEndpoint in daemon.json (upsert_startup_endpoint, one per id, saved only AFTER a successful bringup) + autostart::replay_startup_endpoints on a daemon-boot background thread (fresh session, adapter re-resolved at replay, bounded launch retries, ENDPOINT_AUTOSTART[/SKIP/FAIL] loud tokens, never blocks daemon start). unit = config round-trip + upsert-replaces + loud-skip on unresolvable/non-harness adapter. int = a daemon start replays the saved default and the endpoint comes up without hands-on.\n\n[[requirements]]\nid = \"REQ-DAEMON-SERVICE-INSTALL\"\ntitle = \"F-038 RIDER (flynn nice-to-have, QUEUED not activated): a documented OS-service registration recipe or `spt daemon install-service` verb so the daemon itself survives box reboot (flynn's box runs managed_by:null = daemon-at-boot unprovisioned; kitsubito's hand-rolled systemd --user unit = prior art). NOT in MSG-IDENTITY scope — REQ-ENDPOINT-AUTOSTART covers the daemon-start-to-endpoint leg; the boot-to-daemon leg stays interim (logon scheduled task / systemd unit). Activate at an infra-provisioning milestone; shape (recipe doc vs verb) ruled then. Kin [[daemon-service-detection-gotcha]] (global-OS-state detection blind on dev box — a verb must not regress that), [[kitsubito-linux-rig]].\"\nrequired_stages = []  # QUEUED F-038 rider — register-don't-activate (rule 5). Recipe-vs-verb fork ruled at activation.\n\n[[requirements]]\nid = \"REQ-CONN-POISON-DIAL-SCOPE\"\ntitle = \"MSG-IDENTITY W6 / F-039 (deployah field-acceptance follow-up 2026-07-10, RCA-FIRST — mint per the v0.30.6 PASS handoff): ambient CONN_WRITE_POISONED log-churn correlates 1:1 with PUMP_PEER_FAIL submit-dials to OFFLINE peers (enlyzeam/kitsubito/gravity) with NO wedge and NO freeze — pre-existed the blackhole rig = log-noise/mislabel, not a defect in the r4 fix. CODE CONTEXT: conn.rs poison_and_cancel emits the loud CONN_WRITE_POISONED line for a write that 'exceeded its bound (OR FAILED)' (conn.rs:181) — the fast-FAIL branch (broken pipe / conn refused on an already-dead counterpart) shares the log tag with the TIMEOUT branch that is the field-acceptance wedge observable, so routine conn teardown under offline-peer dial churn reads like poison events. RCA-FIRST: pin the exact write site that fails per PUMP_PEER_FAIL cycle (BrokerConn is broker-side — which broker conn write rides each pump dial failure? status/event fan-out to a departed subscriber? brain-side notification?) BEFORE changing anything — the correlation mechanism is unpinned. FIX SHAPE (post-RCA, doyle rules at lock): reserve the loud CONN_WRITE_POISONED token for the DEADLINE-EXCEEDED class (the wedge observable blackhole-controller.ps1 watches); a plain write-FAIL on an already-dead conn retires quietly (debug-level or a distinct low-noise token). MUST NOT weaken the r4 invariant: every failure path still poisons + retires the conn (REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK behavior unchanged — only the log LOUDNESS is scoped). Gate: unit — deadline-exceeded emits the loud token, fast-fail does not (both still poison); field — offline-peer churn no longer floods daemon.stderr.log with CONN_WRITE_POISONED. Kin REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK (the invariant this must preserve), REQ-PUMP-PEER-ISOLATION (the offline-peer dial neighborhood), [[v0306-published]].\"\nrequired_stages = [\"impl\", \"unit\"]  # FLIPPED in the W6 build commit carrying the evidence (todlando 2026-07-10): impl = the class-split render at poison_and_cancel (conn.rs render_retirement); unit = conn.rs token-split + organic-still-poisons cases. RCA LOCK ADDENDUM (supersedes the title's open questions): QUIC dial ruled OUT; the per-offline-dial-cycle poisoning writes hit SHARED long-lived brain carriers — (1) the DIAL_FAILED presence push (nethost.rs:373, result discarded, fires per failure) and (2) the NetDialSubmitted submit-ack (broker.rs:3894→4370, per submit). Token collision mechanism: poison_and_cancel (conn.rs:176) is called by the DEADLINE class at conn.rs:282/291/381 (the only callers that latch timed_out, conn.rs:112 — the real wedge observable) AND by the organic fast-fail path at conn.rs:431 which NEVER sets timed_out — same loud token for both; the cause bit exists but is unprinted. The once-per-conn poison latch makes per-line 1:1 churn undecidable between fresh-carrier churn and stderr interleave artifact — that attribution is REQ-CONN-POISON-ATTRIBUTION's job (leg b), NOT this REQ's. THIS REQ (leg a): CONN_WRITE_POISONED reserved for the timed_out class; the organic branch emits a DISTINCT CONN_WRITE_RETIRED token carrying the error kind AND (doyle amendment 1) the SAME attribution fields as POISONED — conn id, role, endpoint/session where known; quieter ≠ thinner (the retired class is what discriminates fresh-carrier churn from interleave artifact in the field). Poison/retire BEHAVIOR unchanged; total information non-decreasing; NO timeout changes. impl = split tokens at poison_and_cancel. unit = deadline class → CONN_WRITE_POISONED / organic fast-fail → CONN_WRITE_RETIRED, both still poison+retire.\n\n[[requirements]]\nid = \"REQ-CONN-POISON-ATTRIBUTION\"\ntitle = \"MSG-IDENTITY W6 / F-039 legs b-d (doyle W6 LOCK 2026-07-10, minted per amendment 3): every broker-conn lifecycle record is ATTRIBUTABLE — the W6 RCA's terminal undecidability (per-line 1:1 CONN_WRITE_POISONED churn = fresh-carrier churn OR stderr interleave artifact) exists because records carry no stable conn identity, no role/endpoint/session context, and no timestamps, and the once-per-conn poison latch hides multiplicity. THREE LEGS. (b) IDENTITY: mint a stable per-physical-conn id (monotonic u64 at conn construction — Arc::ptr_eq is the only identity today and it does not survive a log line) plus subscriber role and endpoint/session where known, stamped on CONN_WRITE_POISONED, CONN_WRITE_RETIRED, logical stall-evict, attach/resume/detach, and write-retirement records (RCA attach sites: presence nethost.rs:379, stream nethost.rs:258, controller broker.rs:891, viewer broker.rs:1073). (c) TIME: daemon stderr correlation records carry wall-clock AND monotonic timestamps (stderrlog has neither; broker+brain share one file — interleave is unresolvable without them). (d) LIFECYCLE (doyle-confirmed UNCONDITIONAL, not debug-gated): one BOUNDED set of per-conn lifecycle events — write start/timeout-cancel/transport close/writer exit/replacement-reattach (hertz RCA fix-shape items 1-3). Constraint (doyle LOCK): the split/attribution must not REDUCE total information, only correct its attribution; NO timeout-value changes; NO suppression-as-fix. Gate: unit — lifecycle records carry conn id + role + timestamps; the id is unique per physical conn and stable across that conn's records. Kin REQ-CONN-POISON-DIAL-SCOPE (leg a, the token split these fields ride on), REQ-CONN-BLACKHOLE-LIFECYCLE-HARNESS (leg e, consumes these records), REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK (behavior invariant preserved).\"\nrequired_stages = [\"impl\", \"unit\"]  # FLIPPED in the W6 build commit carrying the evidence (todlando 2026-07-10). impl = conn id mint + describe/label accretion + log_stamp wall+mono (anchored at both daemon entries) + bounded lifecycle events (write-start, transport-close, writer-exit, attach/replace/detach) + attribution on stall-evict records. unit = conn.rs record-shape assertions (id/role/time present + parity across both retirement tokens, id unique per conn, label bounded).\n\n# releases#286 narrowly supersedes leg (d)'s unconditional healthy start/close\n# lines. Failure/poison, role transitions and original first-write time remain\n# attributable per physical conn; healthy counts are interval totals instead.\n[[requirements]]\nid = \"REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED\"\ntitle = \"Healthy broker IPC starts and closes aggregate once per monotonic interval across connections, independently of request volume; poison, failure and role-transition evidence remains per-connection with original first-write time. Livehost session polling reuses its healthy carrier and names why it opens a new one; role=brain alone never identifies the supervised coordinator. Network family diagnostics remain once per real endpoint bind attempt, not IPC churn. Narrow healthy-start/close supersession of REQ-CONN-POISON-ATTRIBUTION leg d (releases#286).\"\nrequired_stages = [\"doc\", \"impl\", \"unit\"]\n\n[[requirements]]\nid = \"REQ-CONN-BLACKHOLE-LIFECYCLE-HARNESS\"\ntitle = \"MSG-IDENTITY W6 / F-039 leg e (doyle W6 LOCK 2026-07-10, minted per amendment 3 — hertz's five invariants VERBATIM from his RCA fix-shape item 5): 'Build a deterministic black-holed-controller harness against current v0.30.6 semantics and assert: unrelated sessions continue; the bad physical connection is canceled/closed within the bound; its writer exits; a fresh viewer can attach; no lock or task remains owned by the retired connection.' The harness is the standing conformance rig for the r4 SHAREDSEND fix-class — hertz's RCA discipline: only after a timestamped incident maps to a FAILING lifecycle invariant does an ownership/cancellation defect get fixed (the likely shape being complete physical-connection cancellation and writer-task join/retirement, never a broader timeout increase). Consumes REQ-CONN-POISON-ATTRIBUTION's records (conn id + lifecycle events are what make the five assertions checkable deterministically). Kin REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK (the invariant class under test — its brain_decouple int stays the Windows-mandatory gate leg), REQ-CONN-POISON-DIAL-SCOPE.\"\nrequired_stages = [\"int\"]  # FLIPPED in the harness commit carrying the evidence (todlando 2026-07-10). int = conn_blackhole_lifecycle five-invariant rig (crates/spt-daemon/tests/conn_blackhole_lifecycle.rs::blackholed_controller_lifecycle_five_invariants) — a deterministic black-holed-controller incident asserting all five: unrelated sessions continue, the bad physical conn is canceled/closed within the bound, its writer exits, a fresh viewer attaches + receives output, no lock/task remains owned by the retired conn (proxy). Windows-mandatory (#![cfg(windows)]) — mirrors the r4 gate's physical-retirement leg; the named-pipe no-write-timeout wedge is the field failure. Rider context (doyle amendment 2): the net_worker_starve canary rig gets SERIALIZED/ISOLATED (nextest serial group or own binary) rather than any loosening of its 300ms bound — a loosened canary stops being a canary; if isolation is impractical, come back for a bound ruling.\n\n[[requirements]]\nid = \"REQ-UPDATE-GH-TRANSPORT\"\ntitle = \"THE-FORKENING W1 (ADR-0036, operator-ruled 2026-07-14): the release channel is PRIVATE (`BigscreenVR/spt-bs-releases`) and the gh CLI is the mandated carrier — release discovery (`releases/latest`, cli.rs:9717) and asset download (cli.rs:4861 public browser URLs) move to deadline-wrapped `gh` subprocess calls (`gh api`, `gh release download`; run_git pattern). WHY gh not token+HTTP: private-repo `browser_download_url` 404s even with a valid token — the API asset-id dance is gh's job. Default repo flips via the existing SPT_INSTALL_REPO seam (cli.rs:5363) + xtask REPO const (main.rs:729) + notif.rs consent-changelog URL rider. Loud failure classes: gh missing -> UPDATE_FETCH_REJECTED:GhCliRequired with OS-SPECIFIC install hints (winget/apt/brew); gh unauthed -> distinct GhAuthRequired pointing at `gh auth login`. Signature verification unchanged — bytes verified after download, carrier-independent (update-set/counter/anchor continuity per ADR-0036 §2). release_verify_e2e reworked to the gh carrier. Gate: unit — url/invocation construction + both failure classes render OS-correct hints; int — fetch against a real gh-authed channel resolves latest + downloads and verifies an asset; doc — self-update docs name the gh prerequisite. Kin REQ-INSTALL-BOOTSTRAP-VERB (same carrier at first install), ADR-0036.\"\nrequired_stages = [\"doc\", \"impl\", \"unit\", \"int\"]  # FLIPPED in the W1 build commit (todlando 2026-07-14). doc = CONTEXT.md release-channel entry (triage commit) + docs-site self-update gh prerequisite. impl = cli.rs gh_status probe (Missing/Unauthed split) + render_gh_cli_required/render_gh_auth_required + cmd_update_fetch on fetch_release_asset_bytes(Gh) for update-set.json + every artifact + install_repo() default BigscreenVR/spt-bs-releases (SPT_INSTALL_REPO seam) + RELEASES_URL/CONSENT_CHANGELOG_URL/xtask REPO flips. unit = gh_failure_classes_render_os_correct_hints (both classes, OS-correct hints). int = release_verify_e2e (gh-carrier download of a real signed release, env-gated) + bootstrap_verb_e2e's GhCliRequired gate leg (the placed binary refuses loud with no resolvable gh — the carrier gate, not an HTTPS fallback).\n\n[[requirements]]\nid = \"REQ-INSTALL-BOOTSTRAP-VERB\"\ntitle = \"THE-FORKENING W1 (ADR-0036 §3, bootstrap shape b operator-ruled): virgin-box install = `gh release download` the platform binary + ONE self-install verb in the binary itself — the verb places the binary at the canonical install path (the path self-update already respawns from, v0.4.2 lesson), registers user PATH, and leaves first-run identity/daemon-start to the existing idempotent first-run; hosted one-liner install scripts (curl|sh / irm|iex at the dead Pages URL) are RETIRED. Non-interactive (CONTEXT.md Installation: the install path doubles as every adapter's pack-in on-demand install). Windows UAC-740 gotcha binding: the downloaded exe keeps the `spt-*` asset name and the verb lives INSIDE spt — no installer-detection trigger words in exe names. README (bs-core) documents: install gh -> gh auth login -> gh release download -> the verb. Gate: unit — verb places/registers idempotently, refuses cross-platform binaries (platform-stamp check exists, v0.3.2); int — from a clean SPT_HOME+PATH sim, downloaded-binary self-install yields a working `spt` on PATH whose `spt update fetch` then speaks the gh channel; doc — README install section rewritten. Kin REQ-UPDATE-GH-TRANSPORT, REQ-INSTALL-1/2 (the two-paths model this reshapes), ADR-0036.\"\nrequired_stages = [\"doc\", \"impl\", \"unit\", \"int\"]  # FLIPPED in the W1 build commit (todlando 2026-07-14). doc = CONTEXT.md Installer-form (gh bootstrap) entry (triage commit) + README.md install section (gh → auth → download → verb). impl = cli.rs Cmd::Install + cmd_install (place at canonical <spt_home>/bin via stage-then-rename, user-PATH registration marker/entry-set-guarded, platform gate install_platform_check, non-interactive by construction, PATH failure degrades loud never fails placement). unit = install_platform_check_gates_stamp_and_host_arch + install_place_is_idempotent_and_path_block_detected. int = bootstrap_verb_e2e (downloaded spt-* asset name → clean-SPT_HOME self-install → canonical path → idempotent re-run → placed binary's update fetch speaks the gh channel). Scripts note: installer/install.{sh,ps1} retire as the PUBLIC surface (Pages pipeline dies in W2); files stay as the oneliner_e2e CI fixture so REQ-INSTALL-1/3/5/10 evidence stands — doyle to ratify or order deletion at gate.\n\n…\n\n  \"REQ-NET-2\",\n  \"REQ-RELAY-DEATH-CONVERGENCE\",\n  \"REQ-RELAY-NO-BUSY-DELIVER\",\n  \"REQ-SEAM-INJECT\",\n  \"REQ-SEND-SPT-HOSTED\",\n  \"REQ-SEND-WINDOW-DRAIN-HONOR\",\n  \"REQ-SPOOL-TAKE-AUDIT\",\n  \"REQ-TERM-2\",\n  \"REQ-WAN-SPT-HOSTED-DELIVERY\",\n]\n\n[[groups]]\nname = \"message-format-identity\"\nrequirements = [\n  \"REQ-CONN-BLACKHOLE-LIFECYCLE-HARNESS\",\n  \"REQ-CONN-POISON-ATTRIBUTION\",\n  \"REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED\",\n  \"REQ-CONN-POISON-DIAL-SCOPE\",\n  \"REQ-DAEMON-SERVICE-INSTALL\",\n  \"REQ-EP-3\",\n  \"REQ-EP-9\",\n  \"REQ-HAZARD-ENVELOPE-ATTR-LINESAFE\",\n  \"REQ-HAZARD-ENVELOPE-CR-LINESAFE\",\n  \"REQ-HAZARD-ENVELOPE-DECODE-ORDER\",\n  \"REQ-HAZARD-ENVELOPE-PARSER-SAFE\",\n  \"REQ-HAZARD-EVENTPART-REASSEMBLY\",\n  \"REQ-HAZARD-ID-CHARSET\",\n  \"REQ-INJECT-MULTILINE-INTEGRITY\",\n  \"REQ-MSG-4\",\n  \"REQ-MSG-5\",\n  \"REQ-MSG-ENVELOPE\",\n  \"REQ-MSG-SENDER-STAMP\",\n  \"REQ-OPID-MINTER-NAMESPACE\",\n  \"REQ-OPID-TRACING-RETRY\",\n  \"REQ-SEND-STAMP-AGENT-ID\",\n  \"REQ-SHELL-4\","}],"details":{"fileSize":2044268,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\traceable-reqs.toml"}}},"isError":false,"timestamp":1788999584239}}
{"type":"message","id":"dca0f23f","parentId":"32d22b5e","timestamp":"2026-09-10T00:19:45.066Z","message":{"role":"toolResult","toolCallId":"call_BkJaSiBdyn99uE5hrUyTuJhm|fc_0dde72effda0ad21016aa1f7a007ec87d28370d031e74a7b61","toolName":"bash","content":[{"type":"text","text":"commit 19a9097956de0db898eaeb3a7dc4765189ade874\nAuthor: Reavo End <decidel@gmail.com>\nCommit: Reavo End <decidel@gmail.com>\n\n    fix(daemon): bound healthy connection stderr and retain livehost polls\n    \n    Document the measured short-lived livehost census carrier and the family-gate exoneration: TOTP rendezvous rotation already logs once per bind. Aggregate only healthy starts/closes; retain original first-write stamps and unconditional exceptional/role records. Add deterministic real-socket regression cells; validation is deferred to the integration owner as requested.\n    \n    Co-authored by: todlando\n\n\nWall time: 0.87 seconds"}],"details":{"timeoutSeconds":300,"wallTimeMs":871.534900000086},"isError":false,"timestamp":1788999585065}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_mNyntykhBLZoMmLdaWYkzzuO|fc_0dde72effda0ad21016aa1f7ac4cd087d2beacd4614445c129","toolName":"yield","startedAt":"2026-09-10T00:20:00.547Z"},"id":"ce11689c","parentId":"dca0f23f","timestamp":"2026-09-10T00:20:00.547Z"}
{"type":"message","id":"7fc59a4b","parentId":"ce11689c","timestamp":"2026-09-10T00:20:00.547Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":"**Reviewing exception emission**"},{"type":"toolCall","id":"call_mNyntykhBLZoMmLdaWYkzzuO|fc_0dde72effda0ad21016aa1f7ac4cd087d2beacd4614445c129","name":"yield","arguments":{"type":null,"data":{"overall_correctness":"correct","explanation":"Standards-axis review is clean: no actionable hard violations or evidence-backed heuristic smells found. The new aggregation uses the existing poison-recovering lock convention (KNOWN-HAZARDS §7.33), releases its shared counter lock before emission, preserves exceptional/role attribution and first-write timestamps, and leaves the bounded-write retirement machinery intact (§7.38). Retained session polling follows the existing broker request/reply dispatch and discards failed carriers rather than supplying false empty liveness. Trace tags accompany production code, behavioral socket regressions and documentation, with the new requirement scoped in the manifest. No tests, builds or runtime validation were executed, as instructed.","confidence":0.88,"findings":[]},"error":null}}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":13233,"output":455,"cacheRead":113536,"cacheWrite":0,"totalTokens":127224,"reasoningTokens":277,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1788999585073,"responseId":"resp_0dde72effda0ad21016aa1f7a1470887d2b5d862c1d906a191","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_0dde72effda0ad21016aa1f7a2d53c87d2ac3997e4d75a30e2","type":"reasoning","content":[],"encrypted_content":"gAAAAABqofesiurtGEeSkRXN3td8xCMfB-hc21pRLGQS8KkeeohyWUbtWu_7zp4dPjGdZDjefUMCYCBFf_5zFUToq5Whht_o1_anH8TfDWIFDxmri2aMyMyASWvCVRSs4qPYDdiqoFzconekenzHmi6d1SlynYbxffXjKqQUi0gB9Gh31dsDrY_AXh7WKPdP3A6_hT_SFAvzXv1NyVGSBrY7Jy0H986QnQu-RzooLg5p-60ZKHJgmeRLfnKTuhUneYZEK0iMTtOkk0R4fnh_34s1VBJivtWynYVpKbVJgnRxct4VL6cITSey-Mdh1ENeapS2YgLT35Sd3yuILpji5SOgvEVrG1B9qRkgz-Ai4RaM-B1bLLxHYnF9CZHRoRS4MBYM86h75lV8ApduAVJR8k9j6HSluKJLeCEPaWqohErrFSFfMmzAnaOEmR_6uwvh-spSzhCLr2lCJEQJfXLLOh6j9DPV3oUFz6Ry-5wF7qcvs0ZPssLajvr-K-wGztZwnxogzZX8pXRRy5Tlej45mMuyV2ClGVcDf_JH_QF0i7Vbq3hgSArAt-1U97WdbrePKFHH5S4P5atS1NynTsVTy55-vNbBOY8sj6FUfJnGknBG623qupm-AHSFMP8C6ODAAXtFk-cIOZBvo2p-OD02laFsAK4upu7EVh_4QFUUJa3KD2Ao5QW4v4IQWPVwGMfCR3igxzwCfATOiSpWkSr8lXXwuaeSTo0Q-X_L7mgyY0BkOdOJqjU0u7jb1uvVBIRFpVDSrzMTsxta_Q5J0Nv2Wrec9kkSR_UuJvEC9gGl43-0PKJS12gHaswbyoTAKlIUny_Z9BD6FsZwNEoXjzYQqWYkCihb3H6v1oa0xddY8818IesqfgWSaYtWmGQKJdXqjJeILwSzYp-VC2T01qW2xKIeIZY0oCSp6q63-y8wXPIkB5ZAYRIGP4um9vZs9UQpnRKLbic8NTaV4a7Ql1g2Uni5FT-9d-F01XD8OgwaMgUks1mdnZI_JdUbRqbz_o4BboT6RskPvoyWgh9X8Idb7SqelAxHlfbLr4qEXo4oV79jXJSBYvS8RhoTFb6U5res9gGdX6Ykfh8as_GpUAyyE2RoZj85NuE7M8jAkg8GQ9aW4EhWXZsQGQRUG4VltdEJT4bqPf72PIWFZi2HQU8EhlAc2oB3MkinH12kF6aY8qt2suKN11p7wlaes6UzwnZhPs4vdYRv8lLP_xeZXfD8H7rux76fPP6mAtl3iZtncxFXmCbtOpfSVZvsvdpXzAL36ezwSfeTSmBY0Zg_a3EsJxjCiw7YIBqILU0phuCOZc1QonH1J3zPF56BVRjgRTJyeqm0Z3mHmbDfOY5YQtBBNu7p8siA31GSY6EN2VgBuR6L798A6vNS1xCRSaJLhh2c4uW7TwlNWBEP-M4E6jTJ-TOoMoOzTRyrLW6ujXXOqOmFX5NQ6KOAexKCG2JXcFSKhcptqK-xhf9iw59y3qiRD_uniRrvVSStbdEzKmjXGR5QutA8KKMMSa8oXupxKQsG6hVzpoU34qzVMfbIzSYcM8BMGhCUSFIu3Ms2c_XrOoLlIQFtP6xkSX3cWK5-hXxbzRGEfN2g76XLbut9BedSg7e9wx5yN32HjPncBAzoymZwltA1-ZVKLE4_78aisbw5TICUWwEeFI9eZ-iYFZpR3AXqoJOVWumBAk5nM0sLnCRNHgZVOmt8aZY3Ti7Vmkgp--qDB4-cOCSTPUx62kLlAhXRooTSCmyQZncv8JiDQytC7EqABHwh3kUajsQsOvNqdeTLbk0h77rhmR-YLHmz-BzujLcWS0gEFt_H9mRtLm593OIKqKpucOOA2o3imgN4QS0IYyE3DKhK8WwBJafxebpki47K43g1dgvFAMmhJVvNfzp4EHAZT8Pc9Kz3d5pGUNcfgv65uB682DdhWEFtXoR27Iptrr0WUwpE9Yx5qnYsSvfOswlBED4oD3RhEVWA1B2mupGnHm_oXuf6ZTxJy9at_rzRjU4o6YIN222gCaT9Ewx76qpNS6_hUIVypK3wbjSac0jax3tgLpSty2vkYjvIn83SkMXLP1RqqC0-7n_5RLtQK6I56cS0ykiY6I5YCoVLbBM3J_h2deuQKNCowFf6NCVaDUmUF6hWjkkcIYcK1OnBNtWHYuzjswcivuPqE07DkJOtkibnjtJnZ4R89YqDQAvJVoyb5maffzp9W7Wr0_TyQndJZsSbNaCQ97lsKbJUXSH_lV0JoARAaWbhEqtaBiFEV_dMn5goVHk4WoiypGs0I8IZ3oJdyaraynL5BKgGEkuvRim_MNEThEsHRvP--jazlHT8tIqUlAaTAonnNJuTQ-H7Kb2EoNGIzu4-5lxS3zbmmR7HMqty1QcrfwK9Gh0tLwjdfquTL9HkIsMoHSEW2Ky-yoSr5p30SAYwWm4yufo42c4w0DXALAWoPiIyTYDg3srVnVGq6LE9ll3U96eBaJ1JPWLbppqHZgfEwPEIT91PSTxMGs2_Rh6kG0TVQlFXavF6hGfBrsVbcEL5UkyyGZ8YKm7ueID_vHTk6kei-TF2VRHLCnw10E0o0E-H5_Iiw3IkeVmUjlgGq-zgrGZc-kIvP5tqHItKcoouQf51IUmV4ZtA1Bv3vDH9H5vykrqYo1CTekEg0qNpb5fz7yOeoQO7v1spN_CBHkK0gS3mRcsSpOyKG8DLVNXQKt5TjCuIu8PGHq1v-W8SAk-iT2qYsWQrl2a5HvFHsQV-VdP21mD1yzg5DwJLuhPp7iYIDuT8tcTWiItt5YkNuO57DYYOCFAv1TXde5UER0V0abIUo4ogYQoNAh4inYQGd0TAncazdLEgwwxC9vtMBUi08cDk1iInou8C982JRijP4WQzcW9rdYMP3_mKeRq3uyFC_VfLXjqE-_grb45t01GgFf1Oh0oveSztt6GIQIXECinKgXQYbYs_3giqdBwA4JD2QZJtbc4xF9se9CAnzBwn-m7mGyM6pVleZnKnzmguwoqUYEI_NIQ6zbAeG0fM-oprcwXcpNfs7ss7tyLksQZ3MCDj7KL8gSAmKIocrPj2dwvyRYMzjOb0-aODO75jFAqCtNMt0vuoHatzZ-9Hj8VLRJ_coPADgBOzQs2cdJ-hntrGSuxTcjS0MyH8IdVRfpDN4MFS6MQG6mFGl-kTVTBh7TPGCUGMz3sviHQBIoaIOCpRpQBUisTn9q4Wr1RgmdgeaOnhH1ncicFnz-v5vPi4mdpJEzKb7BThnKqhriijhBBWxnWVLl6fXyVBgzjYDI3ReDJUkQ6ABtqXZSiEugJh-WKIsoUacM4wfntDf46RtuUZTJD-MlKpgUHpVkQ2aJ0gYgD1i-nK3VfwoDSCz3VfrG7NQekIZQFlfDRvnjsQPbxVdH-dfttvQbAwUxyoEs7KIAc6aJg9kqMiWsiC2dRHFKUCLuZDBh-iq-IfKFUEczitpY48MaJt","internal_chat_message_metadata_passthrough":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"summary":[{"type":"summary_text","text":"**Reviewing exception emission**"}],"metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}},{"id":"fc_0dde72effda0ad21016aa1f7ac4cd087d2beacd4614445c129","type":"function_call","status":"completed","arguments":"{\"type\":null,\"data\":{\"overall_correctness\":\"correct\",\"explanation\":\"Standards-axis review is clean: no actionable hard violations or evidence-backed heuristic smells found. The new aggregation uses the existing poison-recovering lock convention (KNOWN-HAZARDS §7.33), releases its shared counter lock before emission, preserves exceptional/role attribution and first-write timestamps, and leaves the bounded-write retirement machinery intact (§7.38). Retained session polling follows the existing broker request/reply dispatch and discards failed carriers rather than supplying false empty liveness. Trace tags accompany production code, behavioral socket regressions and documentation, with the new requirement scoped in the manifest. No tests, builds or runtime validation were executed, as instructed.\",\"confidence\":0.88,\"findings\":[]},\"error\":null}","call_id":"call_mNyntykhBLZoMmLdaWYkzzuO","internal_chat_message_metadata_passthrough":{"create_time":1788999585.551135,"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"},"name":"yield","metadata":{"turn_id":"baa48617-9a69-458a-8752-ee3e01feb6b8"}}]},"duration":15471.748200000031,"ttft":1786.0075999998953,"completedAt":1788999600546,"contextSnapshot":{"promptTokens":126769,"nonMessageTokens":13648,"compactionEpoch":0}}}
{"type":"message","id":"3c39bd4d","parentId":"7fc59a4b","timestamp":"2026-09-10T00:20:00.548Z","message":{"role":"toolResult","toolCallId":"call_mNyntykhBLZoMmLdaWYkzzuO|fc_0dde72effda0ad21016aa1f7ac4cd087d2beacd4614445c129","toolName":"yield","content":[{"type":"text","text":"Result submitted."}],"details":{"data":{"overall_correctness":"correct","explanation":"Standards-axis review is clean: no actionable hard violations or evidence-backed heuristic smells found. The new aggregation uses the existing poison-recovering lock convention (KNOWN-HAZARDS §7.33), releases its shared counter lock before emission, preserves exceptional/role attribution and first-write timestamps, and leaves the bounded-write retirement machinery intact (§7.38). Retained session polling follows the existing broker request/reply dispatch and discards failed carriers rather than supplying false empty liveness. Trace tags accompany production code, behavioral socket regressions and documentation, with the new requirement scoped in the manifest. No tests, builds or runtime validation were executed, as instructed.","confidence":0.88,"findings":[]},"status":"success"},"isError":false,"timestamp":1788999600548}}
{"type":"custom","customType":"session_exit","data":{"reason":"dispose","kind":"normal","recordedAt":"2026-09-10T00:27:00.570Z"},"id":"ad0371de","parentId":"3c39bd4d","timestamp":"2026-09-10T00:27:00.570Z"}
