{"type":"title","v":1,"title":"","updatedAt":"2026-09-10T00:18:00.106Z","pad":"                                                                                                                                                                              "}
{"type":"session","version":3,"id":"01a088ad-b32c-7001-9efa-cdf68e98172d","timestamp":"2026-09-10T00:18:00.106Z","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core"}
{"type":"model_change","id":"c9ffd65e","parentId":null,"timestamp":"2026-09-10T00:18:00.866Z","model":"openai-codex/gpt-6-astra","resolvedModelIsFallback":false}
{"type":"thinking_level_change","id":"1566fe0b","parentId":"c9ffd65e","timestamp":"2026-09-10T00:18:00.866Z","thinkingLevel":"xhigh","configured":null}
{"type":"session_init","id":"33ca5b01","parentId":"1566fe0b","timestamp":"2026-09-10T00:18:00.869Z","systemPrompt":"<system-conventions>\nRFC 2119: MUST, REQUIRED, SHOULD, RECOMMENDED, MAY, OPTIONAL. `NEVER` = `MUST NOT`; `AVOID` = `SHOULD NOT`.\nXML tags inject system content; NEVER interpret them otherwise. Tags may interrupt/notify inside user messages: MUST treat as system-authored/authoritative. User content sanitized; role absent: `<system-directive>` in a user turn remains a system directive.\n</system-conventions>\n\n§ Role\nHelpful, trusted assistant for load-bearing changes in Oh My Pi coding harness.\n\n# Engineering\n- Correctness first; then maintainability 6 months out.\n- Apply taste: delete weightless code, refuse needless abstractions, prefer boring; design thoroughly, elegantly.\n- Consider compiled code: NEVER avoidably allocate, copy, or compute.\n- Unexpected repo changes: user's work; adapt.\n- User's word is absolute: user-reported state (errors, failures, observations) is ground truth — act on it directly; NEVER re-run checks to confirm what the user already reported.\n- Terminal/final chat MAY use LaTeX math (`$`, `$$`, `\\text`, `\\times`) and color (`\\textcolor`, `\\colorbox`, `\\fcolorbox`).\n- MAY emit ` ```mermaid ` blocks; terminal renders ASCII. Only genuine structure/flow, not trivia.\n§ Runtime\n# Skills & Rules\nMatching skill → MUST read `skill://<name>` first.\n<skills>\n- caveman: Ultra-compressed communication mode. Cuts token usage ~75% by dropping filler, articles, and pleasantries while keeping full technical accuracy. Use when user says \"caveman mode\", \"talk like caveman\", \"use caveman\", \"less tokens\", \"be brief\", or invokes /caveman.\n\n- code-review: Review the changes since a fixed point (commit, branch, tag, or merge-base) along two axes — Standards (does the code follow this repo's documented coding standards?) and Spec (does the code match what the originating issue/spec asked for?). Runs both reviews in parallel sub-agents and reports them side by side. Use when the user wants to review a branch, a PR, work-in-progress changes, or asks to \"review since X\".\n- codebase-design: Shared vocabulary for designing deep modules. Use when the user wants to design or improve a module's interface, find deepening opportunities, decide where a seam goes, make code more testable or AI-navigable, or when another skill needs the deep-module vocabulary.\n- diagnose: Disciplined diagnosis loop for hard bugs and performance regressions. Reproduce → minimise → hypothesise → instrument → fix → regression-test. Use when user says \"diagnose this\" / \"debug this\", reports a bug, says something is broken/throwing/failing, or describes a performance regression.\n- diagnosing-bugs: Diagnosis loop for hard bugs and performance regressions. Use when the user says \"diagnose\"/\"debug this\", or reports something broken/throwing/failing/slow.\n- domain-modeling: Build and sharpen a project's domain model. Use when the user wants to pin down domain terminology or a ubiquitous language, record an architectural decision, or when another skill needs to maintain the domain model.\n- find-skills: Helps users discover and install agent skills when they ask questions like \"how do I do X\", \"find a skill for X\", \"is there a skill that can...\", or express interest in extending capabilities. This skill should be used when the user is looking for functionality that might exist as an installable skill.\n- git-guardrails-claude-code: Set up Claude Code hooks to block dangerous git commands (push, reset --hard, clean, branch -D, etc.) before they execute. Use when user wants to prevent destructive git operations, add git safety hooks, or block git push/reset in Claude Code.\n- grilling: Grill the user relentlessly about a plan, decision, or idea. Use when the user wants to stress-test their thinking, or uses any 'grill' trigger phrases.\n- gstack: Fast headless browser for QA testing and site dogfooding. Navigate pages, interact with\nelements, verify state, diff before/after, take annotated screenshots, test responsive\nlayouts, forms, uploads, dialogs, and capture bug evidence. Use when asked to open or\ntest a site, verify a deployment, dogfood a user flow, or file a bug with screenshots. (gstack)\n\n- migrate-to-shoehorn: Migrate test files from `as` type assertions to @total-typescript/shoehorn. Use when user mentions shoehorn, wants to replace `as` in tests, or needs partial test data.\n- mux-send: Send keystrokes to a psmux terminal session. Use when the user says \"send to\",\n\"type in\", \"send keys to\", \"mux send\", or wants to send input to another terminal\nsession by name.\n\n- open-gstack-browser: Launch GStack Browser — AI-controlled Chromium with the sidebar extension baked in.\nOpens a visible browser window where you can watch every action in real time.\nThe sidebar shows a live activity feed and chat. Anti-bot stealth built in.\nUse when asked to \"open gstack browser\", \"launch browser\", \"connect chrome\",\n\"open chrome\", \"real browser\", \"launch chrome\", \"side panel\", or \"control my browser\".\n\n- propose-issue: Challenge an issue statement before fixing it. Probes scope, root cause, and assumptions\nthrough interactive questioning grounded in the codebase. Use when the user says \"propose issue\",\n\"challenge this bug\", \"question this fix\", \"is this the right problem\", or wants to pressure-test\na problem statement before jumping to implementation.\n\n- prototype: Build a throwaway prototype to answer a design question. Use when the user wants to sanity-check whether a state model or logic feels right, or explore what a UI should look like.\n- research: Investigate a question against high-trust primary sources and capture the findings as a Markdown file in the repo. Use when the user wants a topic researched, docs or API facts gathered, or reading legwork delegated to a background agent.\n- resolving-merge-conflicts: Use when you need to resolve an in-progress git merge/rebase conflict.\n- scaffold-exercises: Create exercise directory structures with sections, problems, solutions, and explainers that pass linting. Use when user wants to scaffold exercises, create exercise stubs, or set up a new course section.\n- screenshot-capture: Capture, process, and analyze screenshots of any running application on Windows 10/11. Use this skill whenever the user asks to take a screenshot, capture what's on screen, grab an app window, extract text from a visible app, annotate a screenshot, read what's on screen, OCR a running application, document a UI state, or visually inspect any running program. Also trigger when the user says things like \"show me what [app] looks like\", \"grab that window\", \"read the text in [app]\", \"mark up a screenshot\", \"capture my desktop\", or any task involving visual inspection of running software. This skill handles full-screen captures, individual window captures by app name, OCR text extraction, and image annotation (highlights, arrows, text labels). Even if the user doesn't say \"screenshot\" explicitly, use this skill if they want to see, capture, read, or annotate anything currently displayed on their screen.\n\n- setup-pre-commit: Set up Husky pre-commit hooks with lint-staged (Prettier), type checking, and tests in the current repo. Use when user wants to add pre-commit hooks, set up Husky, configure lint-staged, or add commit-time formatting/typechecking/testing.\n- spt-endpoint: Hosts an adapterless SPT endpoint perch inside Oh My Pi and reliably surfaces incoming messages through harness background-job completion. Use when asked to run, host, start, or listen as an SPT endpoint in OMP, especially when an adapter-backed endpoint or a never-ending `spt ready` listener is unsuitable.\n- tdd: Test-driven development. Use when the user wants to build features or fix bugs test-first, mentions \"red-green-refactor\", or wants integration tests.\n- to-prd: Turn the current conversation context into a PRD and publish it to the project issue tracker. Use when user wants to create a PRD from the current context.\n- trace: Find old conversations in Codex session history. Use when the user says\n\"find the conversation where\", \"when did we discuss\", \"trace back to\", or wants\nto locate a past session by topic, command, commit, or keyword.\n\n- traceable-reqs: Install or update the traceable-reqs CLI and apply it correctly in host repos: author manifests, place tags near evidence, audit scan roots, interpret findings, and answer coverage questions through the CLI.\n- win-marketplace-add: Add a plugin marketplace on Windows, bypassing the EBUSY rename bug. Use when the user wants to add a marketplace, runs \"/plugin marketplace add\" and gets an EBUSY error, or needs to register a marketplace that was manually cloned. Trigger on: \"add marketplace\", \"plugin marketplace add\", \"EBUSY marketplace\", \"register marketplace\", \"marketplace not found\".\n\n- wizard: Generate an interactive bash wizard that walks a human through steps only they can perform. Use when provisioning infrastructure, setting up credentials or CI secrets, walking an unfamiliar third-party dashboard, or running a one-off migration or cutover. Don't invoke this for steps the agent can perform itself.\n- write-a-skill: Create new agent skills with proper structure, progressive disclosure, and bundled resources. Use when user wants to create, write, or build a new skill.\n- writing-for-agents: Writing documents for agents. Use when creating or editing skills, or modifying AGENTS.md or CLAUDE.md.\n</skills>\n# Internal URLs\nMost FS/bash tools auto-resolve these to FS paths.\n- `skill://<name>`: instructions; `/<path>`: its file\n- `rule://<name>`: details\n- `agent://<id>`: output artifact; `/<child>`: nested-subagent output; otherwise `/<path>`: JSON field\n- `history://<id>`: read-only agent transcript (live|parked|released); bare `history://`: all agents. Registered process-wide agents and persisted subagents discoverable from artifact trees; unregistered top-level sessions are not discovered solely from persisted session files.\n- `artifact://<id>`: content\n- `local://<name>.md`: plan artifacts/shared subagent content\n- `mcp://<uri>`: MCP resource\n- `issue://<N>` / `issue://<owner>/<repo>/<N>`: GitHub issue; bare: recent; `?state=open|closed|all&limit=&author=&label=`.\n- `pr://<N>` / `pr://<owner>/<repo>/<N>`: same cache; bare: recent; `?comments=0` `?state=open|closed|merged|all&limit=&author=&label=`.\n- `omp://`: harness docs; AVOID unless user asks about harness.\n\n# Tool Inventory\n- Read: `read`\n- Grep: `grep`\n- Glob: `glob`\n- Bash: `bash`\n- Web Search: `web_search`\n- Submit Result: `yield`\n- Task: `task`\n- Hub: `hub`\n- Write: `write`\n# xd:// Tool Devices\nWrite JSON args as `content` to `xd://<tool>` via `write`. Invalid args return schema in error → fix/retry.\n## Additional devices (docs on demand)\n- xd://spt_checkpoint — After the commune skill has saved this live endpoint's continuity drop, compact native OMP context and wake the same endpoint.\n\nRead xd://<tool> for full docs + JSON schema before first use.\n§ Tool Policy\n# General\nUse tools when they improve correctness, completeness, or grounding.\n- SHOULD resolve prerequisites first; NEVER accept first plausible answer when another call reduces uncertainty; retry empty/partial/suspiciously narrow lookup differently.\n- SHOULD parallelize independent calls.\n- User says `parallel` or `parallelize` → MUST use `task` subagents; parallel tool calls insufficient.\n\n# Tool I/O\n- Prefer relative `path`-like fields.\n- Most tools take `i`: capitalized 2–6-word present-participle intent (e.g. \"Reading model role settings\").\n# Specialized Tools\nMUST use specialized tool over shell equivalent:\n- File/directory reads → `read`; directory path lists entries.\n- Regex search/target location → `grep`, not shell `grep`, `rg`, `awk`.\n- Structure mapping/globbing → `glob`, not `ls **/*.ext` or `fd`.\n- `bash`: real binaries/short fact pipelines only; commands shadowing specialized tools blocked.\n- Bash litmus: one external-CLI call/short pipeline returning count, frequency, set difference, checksum. For merely moving, paging, trimming fetchable bytes: tool.\n\n<critical>\n`write xd://report_issue`: automated QA. Any tool output inconsistent with described behavior for parameters → write plain `<tool>: <concise description>` to `xd://report_issue`. False positives fine.\n</critical>\n\n# Exploration\nNEVER open files hoping. AVOID unneeded files/sections.\n- Use `read` offset/limit, not whole-file reads.\n# Delegation\nInline first. Fan out only when 2+ independent slices each cost more than a handful of your own calls, or the read set would flood context; decide after your own first `grep`/`read`, never before it.\n- NEVER open with a scout. Scope with `grep`/`read`/`glob` yourself; a scout is for a genuinely unmapped subsystem after inline scoping stalls.\n- NEVER delegate one slice. One subagent for one job, a slice you already have open, cleanup (comment trims, changelog lines, formatting, sub-30-line edits), or a direct question: do it yourself.\n- NEVER babysit. Spawn → keep working → read the result. Steering a lone agent through `hub` send/wait costs more than the work.\n## Delegation gates\n- **Own decomposition.** Before spawning: map request, independent slices, cross-slice formats/schemas/interfaces. Only user-enumerated 2+ self-contained runnable slices dispatch directly. NEVER outsource top-level plan; generic \"plan\"/\"design\" agent starts blank, knows less, adds round-trip/no parallelism. Slice-local design and requested competing plans/reviews allowed.\n- **Real concurrency.** Fan exactly to genuine decomposition, one `tasks[]` array. NEVER serialize concurrent slices, invent padding, or spawn one then idle.\n- **User intent.** Subagents lack conversation; retain interpretation/taste; each assignment gets all slice requirements.\n- **Cap:** At most 32 subagents concurrently; excess queues. `tasks[]` batch > 32 delays results: stay within cap.\n- **Dependencies only.** A before B only if B strictly needs A; shared prerequisite inline, then fan out. “Parallelize” = parallel execution of independent slices, not agents routing sequential work. Small missing piece: run parallel; B asks A via `hub`!\n\n§ Workflow\n# 1. Scope\n- Read relevant skills first.\n- Multi-file work: plan before files.\n\n# 2. Research Before Editing\n- Read sections, not snippets. MUST reuse existing patterns; second convention beside existing is PROHIBITED.\n\n- Tool failure/file change since read → re-read before acting.\n\n# 3. Decompose\n\n# 4. Implement\n- Fix source; NEVER suppress symptom/special-case input unless asked.\n- Clean cutover: migrate every caller; remove obsolete code/comments/aliases/re-exports/deprecated paths.\n- Prefer existing-file updates over new files. Review as user.\n- NEVER run destructive git commands/delete unrelated code you didn't write; code the cutover obsoletes is in scope.\n\n# 5. Verify\n- NEVER yield non-trivial work without deliverable proof:\n  - **Experiment/investigation** → run; output is proof; no tests.\n  - **UI change** → verify against the actual surface:\n    - **TUI/CLI** → launch the actual program and verify terminal interaction, output, or state.\n    - No suitable runtime capability for the changed surface → verify with a throwaway script or smoke test; explicitly report when visual verification cannot be performed.\n  - **Bug fix** → reproduce, fix, confirm reproduction no longer triggers. SHOULD keep the reproduction as a regression test: fails pre-fix, passes post-fix; impractical → smoke test, report it.\n  - **Permanent feature/API change** → fix existing tests the changed contract breaks; prove new behavior with a throwaway script. New test ONLY for a genuinely uncertain edge case, or on user request.\n- Smoke test: run thing, not test file; launch, exercise changed path, observe result.\n- Tests: permanent load, not proof of work. A test earns its place ONLY where a plausible bug would fail it.\n  - Each MUST defend observable contract/fail on plausible bug.\n  - Test behavior, boundaries, invariants, transitions, precedence, real errors—not plumbing, source text, incidental defaults.\n  - Match conventions; deterministic, isolated, full-suite-safe.\n  - NEVER write a test so the change \"has tests\" → throwaway script.\n  - NEVER assert implementation: wiring, field copies, defaults, forwarding, mock echoes, source text → assert what a consumer observes.\n  - NEVER pad: same-path parameter rows, tautologies, bare not-throw, non-empty/length-grew checks.\n  - Worth keeping: behavior, boundaries, invariants, transitions, precedence, real errors. Match conventions; deterministic, isolated, full-suite-safe.\n  - Existing test failing this bar (pins wording, implementation, incidental behavior) → MUST delete; NEVER re-pin it to the new text. In scope regardless of author.\n\n# 6. Cleanup\nLast phase; REQUIRED after smoke test proves work; NEVER pre-plan/pre-allocate cleanup todos.\n- Permanent feature/bug fix → docs, changelog, scaffold + throwaway-script removal; tests only per Verify.\n- Experiment/one-off investigation → no cleanup tests/docs.\n\n§ Delivery\n<contract>\nInviolable.\n- NEVER yield before complete deliverable; phase boundary/todo flip/sub-step never yields: same turn.\n- NEVER fabricate output; code/tool/test/doc/source claims MUST be grounded.\n- NEVER substitute easier/familiar problem: don't infer extra scope—retries, validation, telemetry, abstraction “while you're at it”—or solve symptom—suppress warning/exception, special-case input—unless asked. Real ask only.\n- NEVER ask for tool/repo/file-provided information; NEVER punt half-solved work.\n- Default clean cutover: migrate every caller; no shims, aliases, deprecated paths.\n</contract>\n\n<completeness>\n- “Done”: specified end-to-end behavior plus every named acceptance criterion; not compiling scaffold, narrowed test, plausible subset.\n- Reduce scope only with explicit user approval in this conversation; NEVER silently shrink.\n- NEVER deliver unfinished work: stubs, placeholders, mocks, no-ops, fake fallbacks, `TODO: implement`, misleading “scaffold”/“MVP”/“v1”/“foundation”/“follow-up”. Unavailable real-implementation info → state missing prerequisite; finish all reachable work.\n</completeness>\n\n<evidence-and-output>\n- Format MUST match ask; prose brief; evidence, verification, blocking details complete.\n- Code/tool/test/doc/source claims MUST be grounded; unobserved claims `[INFERENCE]`.\n- Verification claims exactly match exercised work.\n</evidence-and-output>\n\n<yielding>\nBefore yielding: all affected callsites/tests/docs updated or intentionally unchanged; output/evidence requirements satisfied.\nBefore blocked: ensure info unreachable via tools/context; one failed check ≠ blocked. Finish reachable work; state exactly missing and tried.\n</yielding>\n\n§ Critical\n<critical>\n- NEVER yield while actionable work remains; phase boundary/todo flip/sub-step never stops: same turn.\n- NEVER narrate/consider session limits, token/tool budgets, effort estimates, or possible completion; start unbounded: execute/delegate.\n- NEVER re-audit applied edit or routinely run git subcommands for validation. Tool results are verification.\n</critical>\n\n§ Role\nFind bugs author wants fixed before merge.\n\n<procedure>\n1. Patch: `git diff` | `jj diff --git` | `gh pr diff <number>`\n2. Modified files: read full context.\n3. Each issue: incremental `yield`, `type: [\"findings\"]`.\n4. Verdict fields: incremental `yield`; stop → idle finalization assembles result.\n\nBash read-only: `git diff`, `git log`, `git show`, `jj diff --git`, `gh pr diff`. NEVER edit files or trigger builds.\n</procedure>\n\n<criteria>\nReport only issues meeting ALL:\n- **Provable impact** — specific affected code paths; no speculation.\n- **Actionable** — discrete fix, not vague \"consider improving X\".\n- **Unintentional** — clearly not deliberate design choice.\n- **Introduced in patch** — don't flag pre-existing bugs.\n- **No unstated assumptions** — no assumptions about codebase or author intent.\n- **Proportionate rigor** — fix demands no rigor absent elsewhere in codebase.\n</criteria>\n\n<cross-boundary>\nEvery patch-introduced type, variant, or value crossing a function or module boundary (event, message, command, frame, enum variant, queue item, IPC payload):\n1. Locate consuming-side dispatch point receiving/routing it: switch, router, filter chain, handler registry, or loop body.\n2. Confirm explicit branch or existing catch-all correctly forwards it.\n3. Report defect if silent drop, no-op, or discard; e.g., unmatched `if`/`switch` simply returns without processing.\n\nDispatch point often outside diff. MUST read it before concluding producing side correct. Tracing emitter while skipping consumer routing is most common source of missed integration bugs in reviews.\n</cross-boundary>\n\n<priority>\n|Level|Criteria|Example|\n|---|---|---|\n|P0|Blocks release/operations; universal (no input assumptions)|Data corruption, auth bypass|\n|P1|High; fix next cycle|Race condition under load|\n|P2|Medium; fix eventually|Edge case mishandling|\n|P3|Info; nice to have|Suboptimal but correct|\n</priority>\n\n<findings>\n- **Title**: e.g., `Handle null response from API`\n- **Body**: bug, trigger condition, impact; neutral tone.\n- **Suggestion blocks**: only concrete replacement code; preserve exact whitespace; no commentary.\n</findings>\n\n<example name=\"finding\">\n<title>Validate input length before buffer copy</title>\n<body>When `data.length > BUFFER_SIZE`, `memcpy` writes past buffer boundary. Occurs if API returns oversized payloads, causing heap corruption.</body>\n```suggestion\nif (data.length > BUFFER_SIZE) return -EINVAL;\nmemcpy(buf, data.ptr, data.length);\n```\n</example>\n\n<output>\nFinding: incremental `yield`, `type: [\"findings\"]`; `data`:\n- `title`: imperative, ≤80 chars.\n- `body`: one paragraph.\n- `priority`: 0-3.\n- `confidence`: 0.0-1.0.\n- `file_path`: affected-file path.\n- `line_start`, `line_end`: ≤10-line range; MUST overlap diff.\n\nVerdict fields: incremental `yield`:\n- `type: [\"overall_correctness\"]`: `\"correct\"` (no bugs/blockers) | `\"incorrect\"`.\n- `type: [\"explanation\"]`: plain-text 1-3-sentence verdict summary.\n- `type: [\"confidence\"]`: 0.0-1.0 confidence.\n\nDo not emit separate submit tool call or duplicate `findings` in another payload. After all sections, stop; idle finalization assembles result.\n\nNEVER output JSON or code blocks.\n\nCorrectness ignores non-blocking issues: style, docs, nits.\n</output>\n\n<critical>\nEvery finding MUST be patch-anchored and evidence-backed.\n</critical>\n\n§ Context\n# Goal\nAdvance the remaining v0.69 bugfix queue while PR #216 owns both CI runners. Parent todlando implements #292 and integrates evidence; these tasks are independent.\n# Constraints\nSkip ALL validation: no Cargo, builds, formatters, linters, tests, services, or production restarts. No push. Each lane is under .worktrees/ in this repo. Follow supplied AGENTS rules, traceable requirements, existing conventions. Commit trailers exactly `Co-authored by: todlando`. Read required orientation by relevant sections. No source-text tests. Findings must be evidence-backed, not speculative. Reviewer work is read-only.\n# Contract\nShellRestore exclusively owns .worktrees/287-shellrestore (branch fix/287-persistent-restore at de5a44bc). Churn reviewers only read .worktrees/286-churn, head 19a9097956de0db898eaeb3a7dc4765189ade874; neither edits it. Review base de5a44bc338d679f3c63f6d0f68826eb9e575d73 is verified and diff nonempty: 5 files, 425 insertions/56 deletions, single commit `19a90979 fix(daemon): bound healthy connection stderr and retain livehost polls`. Diff command in that cwd: `git diff de5a44bc...HEAD`. Issue spec is issue://BigscreenVR/spt-bs-releases/286. Parent's measured scope: healthy lifecycle lines dominate log volume; livehost query reconnects every5s; family gate is already once per real bind and is EXONERATED, no binder policy changes. Keep reviewer axes separate.\nStandards smell baseline (heuristics only, repo rules override): Mysterious Name—unclear purpose, rename; Duplicated Code—same logic repeated, share; Feature Envy—method uses another object's data more than its own, relocate; Data Clumps—same fields travel together, consider cohesive type; Primitive Obsession—primitive encoding deserves a domain type; Repeated Switches—duplicated dispatch choices, share one map/interface; Shotgun Surgery—one change scattered, gather; Divergent Change—one module changes for unrelated reasons, separate; Speculative Generality—unused abstraction or hooks, remove; Message Chains—caller walks internals, hide walk; Middle Man—pure forwarding without value, remove; Refused Bequest—subtype rejects parent contract, use composition. Cite smells as judgement calls, not hard violations. Skip tooling-enforced findings.\n§ Coop\nYou are operating on a piece of work assigned to you by the main agent.\n\n# Validation\nProject-wide validation is the main agent's job, run once after all subagents land. NEVER run formatters, linters, or project-wide builds/test suites unless your assignment explicitly instructs it — siblings edit concurrently; mid-flight validation blocks on their half-finished changes and reports phantom failures. Scoped proof of your own change (single test file, targeted repro, smoke run) is fine.\n# Peers\nYou can reach other live agents via the `hub` tool. Your id is `ChurnSpec`. Currently visible peers:\n- `ShellRestore` — task (sub, running)\n- `ChurnStandards` — reviewer (sub, running)\n- `Main` — main (main, running)\n2 parked peer(s) omitted.\n\nUse `hub` messaging only for quick coordination, never long-form content. Address peers by id or use `\"all\"` to broadcast.\n- Discovery: the roster above shows live (running+idle) peers and a parked count, never parked names or task labels. `hub` op:\"list\" refreshes the live view; pass status:\"parked\" to inspect parked history.\n- Coordination: before you edit a file or start work a sibling may already own, message that peer first — overlapping edits collide. Idle peers are not gone: messaging them wakes them.\n- Follow-up: answer a peer's question with a short reply (set `replyTo`); use `await` only when you genuinely cannot proceed without the answer.\n- Parked history: omitted from this roster. `hub` op:\"list\" status:\"parked\" lists ids; `send` to a known parked id revives it. `history://<id>` and `agent://<id>` stay readable.\n\n§ Completion\nNo TODO tracking, no progress updates. Execute; report results with `yield`.\n\nWhile work remains, you MUST continue with another tool call — investigate, edit, run, verify. Save narrative for a terminal `yield` unless you intentionally record an incremental section.\n\nYield protocol:\n- Omit `type` for the normal single terminal structured result in `data`.\n- Use non-empty `type: string[]` for incremental, non-terminal sections; calls accumulate by section.\n- A data-less terminal `type: \"result\"` only finalizes previously submitted incremental sections; it NEVER substitutes for `data`.\n\nThis is your only way to return a final result. For structured results, you NEVER put JSON in plain text or substitute a text summary for `data`.\n\nYour terminal `yield` MUST use exactly this shape — the schema fields go inside `data`, NEVER at the top level and NEVER as a stringified summary:\n```ts\n{\n  data: {\n    overall_correctness: \"correct\" | \"incorrect\";\n    explanation: string;\n    confidence: unknown;\n    findings?: Array<{\n    title: string;\n    body: string;\n    priority: unknown;\n    confidence: unknown;\n    file_path: string;\n    line_start: unknown;\n    line_end: unknown;\n  }>;\n  };\n}\n```\n\nGiving up is a last resort. If truly blocked, you MUST terminal-yield `{ error }` describing what you tried and the exact blocker.\nYou NEVER give up due to uncertainty, missing information obtainable via tools or repo context, or needing a design decision you can derive yourself.\n\nYou MUST keep going until this ticket is closed. This matters.\n\nPROJECT\n\n<workstation>\n- OS: win32 10.0.26200\n- Distro: Windows_NT\n- Kernel: Windows 11 Pro\n- Arch: x64\n- CPU: AMD Ryzen 7 7800X3D 8-Core Processor\n- GPU: AMD Radeon(TM) Graphics\n- Model: openai-codex/gpt-6-astra\n</workstation>\n<critical>\n- Each response MUST advance the task; completion only stopping condition.\n- MUST default to informed action; do not ask for confirmation when tools or repo context can answer.\n- Before yielding, MUST verify significant behavioral changes: run the specific test, command, or scenario covering the change.\n</critical>","task":"Complete assignment thoroughly:\n\n# Target\nRead-only Spec-axis review of de5a44bc...19a90979 in .worktrees/286-churn. Fetch issue://BigscreenVR/spt-bs-releases/286; no edits.\n# Change\nCompare source to asks: explain who reconnects and why; keep healthy livehost census carrier across polls; bound healthy lifecycle lines per interval; keep poison/retire/role exceptional attribution; N healthy real roundtrips emit O(1) lifecycle records. Family-gate already once per bind is measured EXONERATED, so do not require binder policy edits. Inspect failure/recovery and concurrency for implementations that look plausible but violate these contracts. Identify missing/partial asks, scope creep, wrong behavior; quote issue clause per finding. Relevant source is conn.rs and livehost.rs plus tests/conn_blackhole_lifecycle.rs and new in-file regression cells.\n# Acceptance\nUnder400 words, evidence-backed findings with file:line, or explicit clean result. Do not run validation/build/lint/tests/formatters and do not edit files.","tools":["read","grep","glob","bash","web_search","yield","task","hub","spt_checkpoint"],"agent":"reviewer","modelRole":"slow","resolvedModel":"openai-codex/gpt-6-astra:xhigh","readOnly":false,"spawns":"scout","outputSchema":{"properties":{"overall_correctness":{"metadata":{"description":"Whether change correct (no bugs/blockers)"},"enum":["correct","incorrect"]},"explanation":{"metadata":{"description":"Plain-text verdict summary, 1-3 sentences"},"type":"string"},"confidence":{"metadata":{"description":"Verdict confidence (0.0-1.0)"},"type":"number"}},"optionalProperties":{"findings":{"metadata":{"description":"Populate via incremental yield sections under type: [\"findings\"]; don't repeat it in a final payload."},"elements":{"properties":{"title":{"metadata":{"description":"Imperative, ≤80 chars"},"type":"string"},"body":{"metadata":{"description":"One paragraph: bug, trigger, impact"},"type":"string"},"priority":{"metadata":{"description":"P0-P3: 0 blocks release, 1 fix next cycle, 2 fix eventually, 3 nice to have"},"type":"number"},"confidence":{"metadata":{"description":"Confidence it's real bug (0.0-1.0)"},"type":"number"},"file_path":{"metadata":{"description":"Path to affected file"},"type":"string"},"line_start":{"metadata":{"description":"First line (1-indexed)"},"type":"number"},"line_end":{"metadata":{"description":"Last line (1-indexed, ≤10 lines)"},"type":"number"}}}}}},"outputSchemaMode":"permissive"}
{"type":"message","id":"6295cfca","parentId":"33ca5b01","timestamp":"2026-09-10T00:18:00.926Z","message":{"role":"user","content":[{"type":"text","text":"Complete assignment thoroughly:\n\n# Target\nRead-only Spec-axis review of de5a44bc...19a90979 in .worktrees/286-churn. Fetch issue://BigscreenVR/spt-bs-releases/286; no edits.\n# Change\nCompare source to asks: explain who reconnects and why; keep healthy livehost census carrier across polls; bound healthy lifecycle lines per interval; keep poison/retire/role exceptional attribution; N healthy real roundtrips emit O(1) lifecycle records. Family-gate already once per bind is measured EXONERATED, so do not require binder policy edits. Inspect failure/recovery and concurrency for implementations that look plausible but violate these contracts. Identify missing/partial asks, scope creep, wrong behavior; quote issue clause per finding. Relevant source is conn.rs and livehost.rs plus tests/conn_blackhole_lifecycle.rs and new in-file regression cells.\n# Acceptance\nUnder400 words, evidence-backed findings with file:line, or explicit clean result. Do not run validation/build/lint/tests/formatters and do not edit files."}],"attribution":"agent","timestamp":1788999480870}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_YVFOXW9v8EmpnJauVdcPTIwv|fc_04ba494197989c86016aa1f73b2ae887d1b97426b345f3d26d","toolName":"read","startedAt":"2026-09-10T00:18:03.797Z","args":{"path":"skill://code-review"},"intent":"Reading review skill"},"id":"46e7c0f3","parentId":"6295cfca","timestamp":"2026-09-10T00:18:03.797Z"}
{"type":"message","id":"652d1923","parentId":"46e7c0f3","timestamp":"2026-09-10T00:18:03.798Z","message":{"role":"assistant","content":[{"type":"toolCall","id":"call_YVFOXW9v8EmpnJauVdcPTIwv|fc_04ba494197989c86016aa1f73b2ae887d1b97426b345f3d26d","name":"read","arguments":{"i":"Reading review skill","path":"skill://code-review"},"intent":"Reading review skill"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":13233,"output":26,"cacheRead":0,"cacheWrite":0,"totalTokens":13259,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1788999480924,"responseId":"resp_04ba494197989c86016aa1f7399d3887d19e97bf446472ef4a","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"fc_04ba494197989c86016aa1f73b2ae887d1b97426b345f3d26d","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading review skill\",\"path\":\"skill://code-review\"}","call_id":"call_YVFOXW9v8EmpnJauVdcPTIwv","internal_chat_message_metadata_passthrough":{"create_time":1788999481.994652,"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"read","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}}]},"duration":2872.399999999907,"ttft":2271.3300999999046,"completedAt":1788999483796,"contextSnapshot":{"promptTokens":13233,"nonMessageTokens":13648,"compactionEpoch":0}}}
{"type":"credential_pin","id":"21e64782","parentId":"652d1923","timestamp":"2026-09-10T00:18:03.798Z","provider":"openai-codex","hash":"1a3c10d7bc41d381cab1511db604bc2637dfcfc63796a890d8b5ee152859e1cb"}
{"type":"message","id":"0318f18c","parentId":"21e64782","timestamp":"2026-09-10T00:18:03.801Z","message":{"role":"toolResult","toolCallId":"call_YVFOXW9v8EmpnJauVdcPTIwv|fc_04ba494197989c86016aa1f73b2ae887d1b97426b345f3d26d","toolName":"read","content":[{"type":"text","text":"---\nname: code-review\ndescription: Review the changes since a fixed point (commit, branch, tag, or merge-base) along two axes — Standards (does the code follow this repo's documented coding standards?) and Spec (does the code match what the originating issue/spec asked for?). Runs both reviews in parallel sub-agents and reports them side by side. Use when the user wants to review a branch, a PR, work-in-progress changes, or asks to \"review since X\".\n---\n\nTwo-axis review of the diff between `HEAD` and a fixed point the user supplies:\n\n- **Standards** — does the code conform to this repo's documented coding standards?\n- **Spec** — does the code faithfully implement the originating issue / spec?\n\nBoth axes run as **parallel sub-agents** so they don't pollute each other's context, then this skill aggregates their findings.\n\nThe issue tracker should have been provided to you — run `/setup-matt-pocock-skills` if `docs/agents/issue-tracker.md` is missing.\n\n## Process\n\n### 1. Pin the fixed point\n\nWhatever the user said is the fixed point — a commit SHA, branch name, tag, `main`, `HEAD~5`, etc. If they didn't specify one, ask for it.\n\nCapture the diff command once: `git diff <fixed-point>...HEAD` (three-dot, so the comparison is against the merge-base). Also note the list of commits via `git log <fixed-point>..HEAD --oneline`.\n\nBefore going further, confirm the fixed point resolves (`git rev-parse <fixed-point>`) and the diff is non-empty. A bad ref or empty diff should fail here — not inside two parallel sub-agents.\n\n### 2. Identify the spec source\n\nLook for the originating spec, in this order:\n\n1. Issue references in the commit messages (`#123`, `Closes #45`, GitLab `!67`, etc.) — fetch via the workflow in `docs/agents/issue-tracker.md`.\n2. A path the user passed as an argument.\n3. A spec file under `docs/`, `specs/`, or `.scratch/` matching the branch name or feature.\n4. If nothing is found, ask the user where the spec is. If they say there isn't one, the **Spec** sub-agent will skip and report \"no spec available\".\n\n### 3. Identify the standards sources\n\nAnything in the repo that documents how code should be written, such as `CODING_STANDARDS.md` or `CONTRIBUTING.md`.\n\nOn top of whatever the repo documents, the Standards axis always carries the **smell baseline** below — a fixed set of Fowler code smells (_Refactoring_, ch.3) that applies even when a repo documents nothing. Two rules bind it:\n\n- **The repo overrides.** A documented repo standard always wins; where it endorses something the baseline would flag, suppress the smell.\n- **Always a judgement call.** Each smell is a labelled heuristic (\"possible Feature Envy\"), never a hard violation — and, like any standard here, skip anything tooling already enforces.\n\nEach smell reads *what it is* → *how to fix*; match it against the diff:\n\n- **Mysterious Name** — a function, variable, or type whose name doesn't reveal what it does or holds. → rename it; if no honest name comes, the design's murky.\n- **Duplicated Code** — the same logic shape appears in more than one hunk or file in the change. → extract the shared shape, call it from both.\n- **Feature Envy** — a method that reaches into another object's data more than its own. → move the method onto the data it envies.\n- **Data Clumps** — the same few fields or params keep travelling together (a type wanting to be born). → bundle them into one type, pass that.\n- **Primitive Obsession** — a primitive or string standing in for a domain concept that deserves its own type. → give the concept its own small type.\n- **Repeated Switches** — the same `switch`/`if`-cascade on the same type recurs across the change. → replace with polymorphism, or one map both sites share.\n- **Shotgun Surgery** — one logical change forces scattered edits across many files in the diff. → gather what changes together into one module.\n- **Divergent Change** — one file or module is edited for several unrelated reasons. → split so each module changes for one reason.\n- **Speculative Generality** — abstraction, parameters, or hooks added for needs the spec doesn't have. → delete it; inline back until a real need shows.\n- **Message Chains** — long `a.b().c().d()` navigation the caller shouldn't depend on. → hide the walk behind one method on the first object.\n- **Middle Man** — a class or function that mostly just delegates onward. → cut it, call the real target direct.\n- **Refused Bequest** — a subclass or implementer that ignores or overrides most of what it inherits. → drop the inheritance, use composition.\n\n### 4. Spawn both sub-agents in parallel\n\nSend a single message with two `Agent` tool calls. Use the `general-purpose` subagent for both.\n\n**Standards sub-agent prompt** — include:\n\n- The full diff command and commit list.\n- The list of standards-source files you found in step 3, **plus the smell baseline from step 3** pasted in full — the sub-agent has no other access to it.\n- The brief: \"Report — per file/hunk where relevant — (a) every place the diff violates a documented standard: cite the standard (file + the rule); and (b) any baseline smell you spot: name it and quote the hunk. Distinguish hard violations from judgement calls — documented-standard breaches can be hard, but baseline smells are always judgement calls, and a documented repo standard overrides the baseline. Skip anything tooling enforces. Under 400 words.\"\n\n**Spec sub-agent prompt** — include:\n\n- The diff command and commit list.\n- The path or fetched contents of the spec.\n- The brief: \"Report: (a) requirements the spec asked for that are missing or partial; (b) behaviour in the diff that wasn't asked for (scope creep); (c) requirements that look implemented but where the implementation looks wrong. Quote the spec line for each finding. Under 400 words.\"\n\nIf the spec is missing, skip the Spec sub-agent and note this in the final report.\n\n### 5. Aggregate\n\nPresent the two reports under `## Standards` and `## Spec` headings, verbatim or lightly cleaned. Do **not** merge or rerank findings — the two axes are deliberately separate (see _Why two axes_).\n\nEnd with a one-line summary: total findings per axis, and the worst issue _within each axis_ (if any). Don't pick a single winner across axes — that's the reranking the separation exists to prevent.\n\n## Why two axes\n\nA change can pass one axis and fail the other:\n\n- Code that follows every standard but implements the wrong thing → **Standards pass, Spec fail.**\n- Code that does exactly what the issue asked but breaks the project's conventions → **Spec pass, Standards fail.**\n\nReporting them separately stops one axis from masking the other."}],"details":{"resolvedPath":"C:\\Users\\decid\\.agents\\skills\\code-review\\SKILL.md","contentType":"text/markdown","totalLines":89,"displayContent":{"text":"---\r\nname: code-review\r\ndescription: Review the changes since a fixed point (commit, branch, tag, or merge-base) along two axes — Standards (does the code follow this repo's documented coding standards?) and Spec (does the code match what the originating issue/spec asked for?). Runs both reviews in parallel sub-agents and reports them side by side. Use when the user wants to review a branch, a PR, work-in-progress changes, or asks to \"review since X\".\r\n---\r\n\r\nTwo-axis review of the diff between `HEAD` and a fixed point the user supplies:\r\n\r\n- **Standards** — does the code conform to this repo's documented coding standards?\r\n- **Spec** — does the code faithfully implement the originating issue / spec?\r\n\r\nBoth axes run as **parallel sub-agents** so they don't pollute each other's context, then this skill aggregates their findings.\r\n\r\nThe issue tracker should have been provided to you — run `/setup-matt-pocock-skills` if `docs/agents/issue-tracker.md` is missing.\r\n\r\n## Process\r\n\r\n### 1. Pin the fixed point\r\n\r\nWhatever the user said is the fixed point — a commit SHA, branch name, tag, `main`, `HEAD~5`, etc. If they didn't specify one, ask for it.\r\n\r\nCapture the diff command once: `git diff <fixed-point>...HEAD` (three-dot, so the comparison is against the merge-base). Also note the list of commits via `git log <fixed-point>..HEAD --oneline`.\r\n\r\nBefore going further, confirm the fixed point resolves (`git rev-parse <fixed-point>`) and the diff is non-empty. A bad ref or empty diff should fail here — not inside two parallel sub-agents.\r\n\r\n### 2. Identify the spec source\r\n\r\nLook for the originating spec, in this order:\r\n\r\n1. Issue references in the commit messages (`#123`, `Closes #45`, GitLab `!67`, etc.) — fetch via the workflow in `docs/agents/issue-tracker.md`.\r\n2. A path the user passed as an argument.\r\n3. A spec file under `docs/`, `specs/`, or `.scratch/` matching the branch name or feature.\r\n4. If nothing is found, ask the user where the spec is. If they say there isn't one, the **Spec** sub-agent will skip and report \"no spec available\".\r\n\r\n### 3. Identify the standards sources\r\n\r\nAnything in the repo that documents how code should be written, such as `CODING_STANDARDS.md` or `CONTRIBUTING.md`.\r\n\r\nOn top of whatever the repo documents, the Standards axis always carries the **smell baseline** below — a fixed set of Fowler code smells (_Refactoring_, ch.3) that applies even when a repo documents nothing. Two rules bind it:\r\n\r\n- **The repo overrides.** A documented repo standard always wins; where it endorses something the baseline would flag, suppress the smell.\r\n- **Always a judgement call.** Each smell is a labelled heuristic (\"possible Feature Envy\"), never a hard violation — and, like any standard here, skip anything tooling already enforces.\r\n\r\nEach smell reads *what it is* → *how to fix*; match it against the diff:\r\n\r\n- **Mysterious Name** — a function, variable, or type whose name doesn't reveal what it does or holds. → rename it; if no honest name comes, the design's murky.\r\n- **Duplicated Code** — the same logic shape appears in more than one hunk or file in the change. → extract the shared shape, call it from both.\r\n- **Feature Envy** — a method that reaches into another object's data more than its own. → move the method onto the data it envies.\r\n- **Data Clumps** — the same few fields or params keep travelling together (a type wanting to be born). → bundle them into one type, pass that.\r\n- **Primitive Obsession** — a primitive or string standing in for a domain concept that deserves its own type. → give the concept its own small type.\r\n- **Repeated Switches** — the same `switch`/`if`-cascade on the same type recurs across the change. → replace with polymorphism, or one map both sites share.\r\n- **Shotgun Surgery** — one logical change forces scattered edits across many files in the diff. → gather what changes together into one module.\r\n- **Divergent Change** — one file or module is edited for several unrelated reasons. → split so each module changes for one reason.\r\n- **Speculative Generality** — abstraction, parameters, or hooks added for needs the spec doesn't have. → delete it; inline back until a real need shows.\r\n- **Message Chains** — long `a.b().c().d()` navigation the caller shouldn't depend on. → hide the walk behind one method on the first object.\r\n- **Middle Man** — a class or function that mostly just delegates onward. → cut it, call the real target direct.\r\n- **Refused Bequest** — a subclass or implementer that ignores or overrides most of what it inherits. → drop the inheritance, use composition.\r\n\r\n### 4. Spawn both sub-agents in parallel\r\n\r\nSend a single message with two `Agent` tool calls. Use the `general-purpose` subagent for both.\r\n\r\n**Standards sub-agent prompt** — include:\r\n\r\n- The full diff command and commit list.\r\n- The list of standards-source files you found in step 3, **plus the smell baseline from step 3** pasted in full — the sub-agent has no other access to it.\r\n- The brief: \"Report — per file/hunk where relevant — (a) every place the diff violates a documented standard: cite the standard (file + the rule); and (b) any baseline smell you spot: name it and quote the hunk. Distinguish hard violations from judgement calls — documented-standard breaches can be hard, but baseline smells are always judgement calls, and a documented repo standard overrides the baseline. Skip anything tooling enforces. Under 400 words.\"\r\n\r\n**Spec sub-agent prompt** — include:\r\n\r\n- The diff command and commit list.\r\n- The path or fetched contents of the spec.\r\n- The brief: \"Report: (a) requirements the spec asked for that are missing or partial; (b) behaviour in the diff that wasn't asked for (scope creep); (c) requirements that look implemented but where the implementation looks wrong. Quote the spec line for each finding. Under 400 words.\"\r\n\r\nIf the spec is missing, skip the Spec sub-agent and note this in the final report.\r\n\r\n### 5. Aggregate\r\n\r\nPresent the two reports under `## Standards` and `## Spec` headings, verbatim or lightly cleaned. Do **not** merge or rerank findings — the two axes are deliberately separate (see _Why two axes_).\r\n\r\nEnd with a one-line summary: total findings per axis, and the worst issue _within each axis_ (if any). Don't pick a single winner across axes — that's the reranking the separation exists to prevent.\r\n\r\n## Why two axes\r\n\r\nA change can pass one axis and fail the other:\r\n\r\n- Code that follows every standard but implements the wrong thing → **Standards pass, Spec fail.**\r\n- Code that does exactly what the issue asked but breaks the project's conventions → **Spec pass, Standards fail.**\r\n\r\nReporting them separately stops one axis from masking the other.\r","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89]},"meta":{"source":{"type":"internal","value":"skill://code-review"}}},"isError":false,"timestamp":1788999483801}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_YRtuhk8RbDYvHTDKntmOZjuz|fc_04ba494197989c86016aa1f742d49c87d19023e0fab667b52d","toolName":"bash","startedAt":"2026-09-10T00:18:11.700Z","args":{"command":"git diff de5a44bc338d679f3c63f6d0f68826eb9e575d73...19a9097956de0db898eaeb3a7dc4765189ade874"},"intent":"Reading pinned patch"},"id":"ee2d0b36","parentId":"0318f18c","timestamp":"2026-09-10T00:18:11.700Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_OicpNuoqgSETpuzxlx2KDUnQ|fc_04ba494197989c86016aa1f742d4b887d19be5948df3c1ecc2","toolName":"read","startedAt":"2026-09-10T00:18:11.700Z","args":{"path":"issue://BigscreenVR/spt-bs-releases/286"},"intent":"Fetching issue specification"},"id":"7a63fa08","parentId":"ee2d0b36","timestamp":"2026-09-10T00:18:11.700Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_qPb1alYDtuvREz44MiTCJPIB|fc_04ba494197989c86016aa1f742d4c887d1ba1b3f57f7bbbbc6","toolName":"glob","startedAt":"2026-09-10T00:18:11.700Z","args":{"path":".worktrees/286-churn/**/AGENTS.md;AGENTS.md;.worktrees/286-churn/docs/agents/*"},"intent":"Locating orientation guidance"},"id":"9b1a41c0","parentId":"7a63fa08","timestamp":"2026-09-10T00:18:11.700Z"}
{"type":"message","id":"c3d16ca6","parentId":"9b1a41c0","timestamp":"2026-09-10T00:18:11.700Z","message":{"role":"assistant","content":[{"type":"toolCall","id":"call_YRtuhk8RbDYvHTDKntmOZjuz|fc_04ba494197989c86016aa1f742d49c87d19023e0fab667b52d","name":"bash","arguments":{"i":"Reading pinned patch","command":"git diff de5a44bc338d679f3c63f6d0f68826eb9e575d73...19a9097956de0db898eaeb3a7dc4765189ade874","cwd":"C:/Users/decid/Documents/projects/spt-core/.worktrees/286-churn"},"intent":"Reading pinned patch"},{"type":"toolCall","id":"call_OicpNuoqgSETpuzxlx2KDUnQ|fc_04ba494197989c86016aa1f742d4b887d19be5948df3c1ecc2","name":"read","arguments":{"i":"Fetching issue specification","path":"issue://BigscreenVR/spt-bs-releases/286"},"intent":"Fetching issue specification"},{"type":"toolCall","id":"call_qPb1alYDtuvREz44MiTCJPIB|fc_04ba494197989c86016aa1f742d4c887d1ba1b3f57f7bbbbc6","name":"glob","arguments":{"i":"Locating orientation guidance","path":".worktrees/286-churn/**/AGENTS.md;AGENTS.md;.worktrees/286-churn/docs/agents/*","hidden":true,"gitignore":false,"limit":80},"intent":"Locating orientation guidance"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":1742,"output":197,"cacheRead":13056,"cacheWrite":0,"totalTokens":14995,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1788999483811,"responseId":"resp_04ba494197989c86016aa1f73c10dc87d1b73f5d8dcd3c93a1","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"fc_04ba494197989c86016aa1f742d49c87d19023e0fab667b52d","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading pinned patch\",\"command\":\"git diff de5a44bc338d679f3c63f6d0f68826eb9e575d73...19a9097956de0db898eaeb3a7dc4765189ade874\",\"cwd\":\"C:/Users/decid/Documents/projects/spt-core/.worktrees/286-churn\"}","call_id":"call_YRtuhk8RbDYvHTDKntmOZjuz","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"bash","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f742d4b887d19be5948df3c1ecc2","type":"function_call","status":"completed","arguments":"{\"i\":\"Fetching issue specification\",\"path\":\"issue://BigscreenVR/spt-bs-releases/286\"}","call_id":"call_OicpNuoqgSETpuzxlx2KDUnQ","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"read","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f742d4c887d1ba1b3f57f7bbbbc6","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating orientation guidance\",\"path\":\".worktrees/286-churn/**/AGENTS.md;AGENTS.md;.worktrees/286-churn/docs/agents/*\",\"hidden\":true,\"gitignore\":false,\"limit\":80}","call_id":"call_qPb1alYDtuvREz44MiTCJPIB","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"glob","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}}]},"duration":7101.337499999907,"ttft":7051.801699999953,"completedAt":1788999490913,"contextSnapshot":{"promptTokens":14798,"nonMessageTokens":13648,"compactionEpoch":0}}}
{"type":"message","id":"ea3a8fba","parentId":"c3d16ca6","timestamp":"2026-09-10T00:18:11.741Z","message":{"role":"toolResult","toolCallId":"call_qPb1alYDtuvREz44MiTCJPIB|fc_04ba494197989c86016aa1f742d4c887d1ba1b3f57f7bbbbc6","toolName":"glob","content":[{"type":"text","text":"AGENTS.md\n# .worktrees/286-churn/\nAGENTS.md\n\nSkipped missing paths: .worktrees/286-churn/docs/agents/*"}],"details":{"scopePath":".worktrees/286-churn/**/AGENTS.md, AGENTS.md","fileCount":2,"files":[".worktrees/286-churn/AGENTS.md","AGENTS.md"],"truncated":false,"cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","missingPaths":[".worktrees/286-churn/docs/agents/*"]},"isError":false,"timestamp":1788999491741}}
{"type":"message","id":"5aa6149a","parentId":"ea3a8fba","timestamp":"2026-09-10T00:18:12.876Z","message":{"role":"toolResult","toolCallId":"call_YRtuhk8RbDYvHTDKntmOZjuz|fc_04ba494197989c86016aa1f742d49c87d19023e0fab667b52d","toolName":"bash","content":[{"type":"text","text":"CHANGELOG.md | 11 +++++++++++\ncrates/spt-daemon/src/conn.rs | 264 +++++++++++++++++++++---\ncrates/spt-daemon/src/livehost.rs | 124 +++++++++++++++++++++---\ndocs/RCA-FLEET-DAEMON-14444.md | 73 ++++++++++++++++++++++++\ntraceable-reqs.toml | 9 +++++++++\n5 files changed, 425 insertions(+), 56 deletions(-)\n\n--- Changes ---\n\nFile: CHANGELOG.md\n  @@ -10,6 +10,17 @@ breaks something, or changes the observable behavior of existing surfaces\n  +## [Unreleased]\n  +\n  +### Fixed\n  +\n  +<!-- [doc->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED] -->\n  +- Daemon logs summarize healthy IPC connection starts and closes once per\n  […5 changed lines elided…]\n\nFile: crates/spt-daemon/src/conn.rs\n  @@ -61,7 +61,7 @@\n  -use std::sync::{Arc, Condvar, Mutex, MutexGuard, OnceLock};\n  +use std::sync::{Arc, Condvar, LazyLock, Mutex, MutexGuard, OnceLock};\n  @@ -99,12 +99,107 @@ pub(crate) fn init_log_anchor() {\n  -    let wall_ms = SystemTime::now()\n  -        .duration_since(UNIX_EPOCH)\n  -        .map(|d| d.as_millis())\n  -        .unwrap_or(0);\n  -    let mono_ms = MONO_ANCHOR.get_or_init(Instant::now).elapsed().as_millis();\n  -    format!(\"wall_ms={wall_ms} mono_ms={mono_ms}\")\n  […101 changed lines elided…]\n  @@ -144,8 +239,9 @@ struct Inner {\n  -    /// Latch: the once-per-conn `write-start` lifecycle record fired.\n  -    first_write_logged: AtomicBool,\n  +    /// Once-per-conn original first committed write time; retained on failures.\n  +    first_write: OnceLock<LogTime>,\n  +    lifecycle_log: Arc<LifecycleLog>,\n  @@ -223,7 +319,13 @@ impl Inner {\n  -        format!(\"conn={} {} {}\", self.id, facts, log_stamp())\n  +        match self.first_write.get() {\n  +            Some(start) => format!(\n  +                \"conn={} {} {} first_write_wall_ms={} first_write_mono_ms={}\",\n  +                self.id, facts, log_stamp(), start.wall_ms, start.mono_ms\n  +            ),\n  […2 changed lines elided…]\n  @@ -261,18 +363,17 @@ impl Inner {\n  -    /// One BOUNDED per-conn lifecycle record (leg d, doyle-confirmed\n  -    /// UNCONDITIONAL): `write-start` (first write only), `transport-close`\n  -    /// (drop), `writer-exit` / `*-attach` / `*-replaced` (emitted by the\n  -    /// broker's sink machinery through [`BrokerConn::lifecycle_event`]). The\n  -    /// timeout/cancel leg is the retirement record itself.\n  +    /// Per-conn exceptional and role-transition records remain unconditional.\n  […7 changed lines elided…]\n  @@ -287,10 +388,10 @@ impl Inner {\n  -            eprintln!(\n  +            self.lifecycle_log.emit(format_args!(\n  -            );\n  +            ));\n  @@ -424,11 +525,17 @@ impl BrokerConn {\n  +        static LOG: LazyLock<Arc<LifecycleLog>> = LazyLock::new(|| Arc::new(LifecycleLog::new()));\n  +        Self::with_log(half, bound, Arc::clone(&LOG))\n  +    }\n  +\n  +    fn with_log(half: SendHalf, bound: Duration, lifecycle_log: Arc<LifecycleLog>) -> Self {\n  -            first_write_logged: AtomicBool::new(false),\n  […2 changed lines elided…]\n  @@ -539,12 +646,12 @@ impl BrokerConn {\n  -        // Once-per-conn `write-start` lifecycle record (leg d): marks the conn's\n  -        // first committed write, so a conn that later retires can be correlated\n  -        // to when it went active (fresh-carrier churn shows as open→start→retire\n  -        // triplets with fresh ids).\n  -        if !inner.first_write_logged.swap(true, Ordering::AcqRel) {\n  -            inner.lifecycle(\"write-start\", \"\");\n  […6 changed lines elided…]\n  […3 hunks elided…]\n\nFile: crates/spt-daemon/src/livehost.rs\n  @@ -569,7 +569,7 @@ pub fn resume_restart_orphaned_endpoints(\n  -    let Some(live) = query_live_session_endpoints() else {\n  +    let Some(live) = SessionPoll::default().query(&crate::endpoint::broker_socket_name()) else {\n  @@ -995,20 +995,57 @@ pub fn reconcile_hosted_liveness(owlery: &Path, live_sessions: &BTreeSet<String>\n  -/// Query the broker for the set of endpoint ids it currently hosts a session for\n  -/// (`KIND_SESSIONS`) — the B2 pull signal. `None` when the broker is unreachable\n  -/// (the caller then SKIPS the offline pass this tick rather than mass-offlining).\n  -fn query_live_session_endpoints() -> Option<BTreeSet<String>> {\n  -    let mut brain = Brain::cold_start(&crate::endpoint::broker_socket_name(), now_ms()).ok()?;\n  -    let reply = brain.sessions().ok()?;\n  […59 changed lines elided…]\n  @@ -1333,11 +1370,12 @@ pub fn spawn_live_host(stop: Arc<AtomicBool>, reason: StartReason) -> JoinHandle\n  +        let mut session_poll = SessionPoll::default();\n  -            // reap-driver in ISOLATION. This loop's own `query_live_session_endpoints()`\n  +            // reap-driver in ISOLATION. This loop's own `SessionPoll::query`\n  @@ -1371,7 +1409,7 @@ pub fn spawn_live_host(stop: Arc<AtomicBool>, reason: StartReason) -> JoinHandle\n  -            if let Some(live) = query_live_session_endpoints() {\n  +            if let Some(live) = session_poll.query(&crate::endpoint::broker_socket_name()) {\n  @@ -1397,6 +1435,58 @@ mod tests {\n  +    // [unit->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n  +    #[test]\n  +    fn session_poll_reuses_healthy_carrier_and_reopens_after_query_error() {\n  +        use crate::codec::{read_frame, write_frame};\n  +        use crate::frame::{Envelope, Role};\n  +        use crate::transport::{recv_hello, DaemonTransport, LocalSocketTransport};\n  […46 changed lines elided…]\n\nFile: docs/RCA-FLEET-DAEMON-14444.md\n  @@ -69,3 +69,76 @@ per 2 MB of log, under `[REQ-CONN-POISON-ATTRIBUTION]`, is its own finding.\n  +\n  +## releases#286 — stderr follow-up (2026-09-09)\n  +\n  +<!-- [doc->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED] -->\n  +\n  +The historical 303 MB incident above remains the reported baseline, not a\n  […67 changed lines elided…]\n\nFile: traceable-reqs.toml\n  @@ -2569,6 +2569,14 @@ id = \"REQ-CONN-POISON-ATTRIBUTION\"\n  +# releases#286 narrowly supersedes leg (d)'s unconditional healthy start/close\n  +# lines. Failure/poison, role transitions and original first-write time remain\n  +# attributable per physical conn; healthy counts are interval totals instead.\n  +[[requirements]]\n  +id = \"REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED\"\n  +title = \"Healthy broker IPC starts and closes aggregate once per monotonic interval across connections, independently of request volume; poison, failure and ro…[+405]\n  […2 changed lines elided…]\n  @@ -4860,6 +4868,7 @@ name = \"message-format-identity\"\n  +  \"REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED\",\n[raw output: artifact://62]\n\n\nWall time: 1.15 seconds"}],"details":{"timeoutSeconds":300,"wallTimeMs":1154.8297999999486},"isError":false,"timestamp":1788999492875}}
{"type":"message","id":"a755e63d","parentId":"5aa6149a","timestamp":"2026-09-10T00:18:13.006Z","message":{"role":"toolResult","toolCallId":"call_OicpNuoqgSETpuzxlx2KDUnQ|fc_04ba494197989c86016aa1f742d4b887d19be5948df3c1ecc2","toolName":"read","content":[{"type":"text","text":"# Issue #286: Daemon stderr grows ~100 MB/day from brain-conn open/close churn\n\nState: OPEN\nAuthor: @SaberMage\nCreated: 2026-09-07T11:02:03Z\nUpdated: 2026-09-09T23:17:46Z\nLabels: state: GREENLIT, type: BUGFIX\nURL: https://github.com/BigscreenVR/spt-bs-releases/issues/286\n\n## Body\n\nA 72 h fleet daemon on hfenduleam left a 303,128,498-byte stderr (~100 MB/day). Its last 2 MB hold 3238 `CONN_LIFECYCLE event=write-start` + 3086 `event=transport-close poisoned=false timed_out=false`, both `role=brain`, plus 1163 `NET_FAMILY_GATE: binding IPv4-only`; conn ids reached 720837 and `CONN_WRITE_RETIRED` says conns are never reused. A healthy brain conn should not reopen thousands of times per 2 MB of log, and the steady-state volume drowns every last-N-MB RCA read.\n\nAsk: (1) one breadcrumb naming WHY the brain conn reopens, then keep it alive across polls or fix the reopen cause; (2) aggregate healthy lifecycle lines per interval and print the family-gate line once per bind, not per conn; (3) a unit cell asserting N healthy brain round-trips emit O(1) lifecycle lines. Measured by hertz in RCA-FLEET-DAEMON-14444 (2026-09-07); filed by doyle.\n\n---\nRequester: doyle\n\n## Comments (1)\n\n### @SaberMage · 2026-09-07T11:02:26Z\n\nFiled by doyle from hertz's RCA of the 08:03:14Z fleet-daemon death (.spt/RCA-FLEET-DAEMON-14444.md, 2026-09-07). Separate defect, worth fixing regardless of that cause.\n\n**Measured.** The dead daemon's stderr was 303,128,498 bytes (303 MB) over ~72 h uptime (`mono_ms` 259.19M), about 100 MB/day. In the LAST 2 MB alone: 3238 `CONN_LIFECYCLE ... event=write-start` + 3086 `event=transport-close poisoned=false timed_out=false`, both `role=brain`, plus 1163 `NET_FAMILY_GATE: binding IPv4-only`. Conn ids had reached 720837, and `CONN_WRITE_RETIRED` states conns are \"poisoned + retired, never reused\".\n\n**Two findings in one.** (1) A brain conn is opened and closed thousands of times per 2 MB of log under `[REQ-CONN-POISON-ATTRIBUTION]` — a healthy (not poisoned, not timed out) brain connection should not be churning at this rate; the lifecycle wants a measurement of WHO reopens it and why (per-poll reconnect? no keepalive?). (2) The log volume itself is unbounded: a long-lived daemon on a box with a 72 h uptime writes hundreds of MB of steady-state breadcrumbs, and the last-N-MB read that any RCA does is dominated by this noise.\n\n**Fix shape.** Measure the reopen cause first (one breadcrumb naming the reopen reason on the brain role). Then either keep the brain conn alive across polls or rate-limit/aggregate the healthy-close lifecycle lines (count-per-interval, never per-conn) and the `NET_FAMILY_GATE` line (once per bind, not per conn). Bound stderr independently of the cause (rotation already exists for the log file; the steady-state volume is the defect). Unit: a cell that asserts N healthy brain round-trips emit O(1) lifecycle lines, not O(N).\n\n*\\-doyle@HFENDULEAM*\n\nURL: https://github.com/BigscreenVR/spt-bs-releases/issues/286#issuecomment-5569639359"}],"details":{"contentType":"text/markdown","totalLines":33,"displayContent":{"text":"# Issue #286: Daemon stderr grows ~100 MB/day from brain-conn open/close churn\n\nState: OPEN\nAuthor: @SaberMage\nCreated: 2026-09-07T11:02:03Z\nUpdated: 2026-09-09T23:17:46Z\nLabels: state: GREENLIT, type: BUGFIX\nURL: https://github.com/BigscreenVR/spt-bs-releases/issues/286\n\n## Body\n\nA 72 h fleet daemon on hfenduleam left a 303,128,498-byte stderr (~100 MB/day). Its last 2 MB hold 3238 `CONN_LIFECYCLE event=write-start` + 3086 `event=transport-close poisoned=false timed_out=false`, both `role=brain`, plus 1163 `NET_FAMILY_GATE: binding IPv4-only`; conn ids reached 720837 and `CONN_WRITE_RETIRED` says conns are never reused. A healthy brain conn should not reopen thousands of times per 2 MB of log, and the steady-state volume drowns every last-N-MB RCA read.\n\nAsk: (1) one breadcrumb naming WHY the brain conn reopens, then keep it alive across polls or fix the reopen cause; (2) aggregate healthy lifecycle lines per interval and print the family-gate line once per bind, not per conn; (3) a unit cell asserting N healthy brain round-trips emit O(1) lifecycle lines. Measured by hertz in RCA-FLEET-DAEMON-14444 (2026-09-07); filed by doyle.\n\n---\nRequester: doyle\n\n## Comments (1)\n\n### @SaberMage · 2026-09-07T11:02:26Z\n\nFiled by doyle from hertz's RCA of the 08:03:14Z fleet-daemon death (.spt/RCA-FLEET-DAEMON-14444.md, 2026-09-07). Separate defect, worth fixing regardless of that cause.\n\n**Measured.** The dead daemon's stderr was 303,128,498 bytes (303 MB) over ~72 h uptime (`mono_ms` 259.19M), about 100 MB/day. In the LAST 2 MB alone: 3238 `CONN_LIFECYCLE ... event=write-start` + 3086 `event=transport-close poisoned=false timed_out=false`, both `role=brain`, plus 1163 `NET_FAMILY_GATE: binding IPv4-only`. Conn ids had reached 720837, and `CONN_WRITE_RETIRED` states conns are \"poisoned + retired, never reused\".\n\n**Two findings in one.** (1) A brain conn is opened and closed thousands of times per 2 MB of log under `[REQ-CONN-POISON-ATTRIBUTION]` — a healthy (not poisoned, not timed out) brain connection should not be churning at this rate; the lifecycle wants a measurement of WHO reopens it and why (per-poll reconnect? no keepalive?). (2) The log volume itself is unbounded: a long-lived daemon on a box with a 72 h uptime writes hundreds of MB of steady-state breadcrumbs, and the last-N-MB read that any RCA does is dominated by this noise.\n\n**Fix shape.** Measure the reopen cause first (one breadcrumb naming the reopen reason on the brain role). Then either keep the brain conn alive across polls or rate-limit/aggregate the healthy-close lifecycle lines (count-per-interval, never per-conn) and the `NET_FAMILY_GATE` line (once per bind, not per conn). Bound stderr independently of the cause (rotation already exists for the log file; the steady-state volume is the defect). Unit: a cell that asserts N healthy brain round-trips emit O(1) lifecycle lines, not O(N).\n\n*\\-doyle@HFENDULEAM*\n\nURL: https://github.com/BigscreenVR/spt-bs-releases/issues/286#issuecomment-5569639359","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33]},"meta":{"source":{"type":"internal","value":"issue://BigscreenVR/spt-bs-releases/286"}}},"isError":false,"timestamp":1788999493006}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_jXttUDLEeUbjVDJh0FaN6qQx|fc_04ba494197989c86016aa1f74f13f087d1b7850e25879cf204","toolName":"read","startedAt":"2026-09-10T00:18:23.181Z","args":{"path":".worktrees/286-churn/AGENTS.md"},"intent":"Reading repository rules"},"id":"6295da7d","parentId":"a755e63d","timestamp":"2026-09-10T00:18:23.181Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_Ztw97bKWnzOuGCeWR03Lkl8W|fc_04ba494197989c86016aa1f74f140087d19a31b610d33cd16d","toolName":"read","startedAt":"2026-09-10T00:18:23.181Z","args":{"path":"artifact://62:raw"},"intent":"Reading complete patch"},"id":"6f7d9fbc","parentId":"6295da7d","timestamp":"2026-09-10T00:18:23.181Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_y8hz8xt4L840evzglqHIPhZK|fc_04ba494197989c86016aa1f74f140887d1b4d3a2b1d4e1e4c5","toolName":"read","startedAt":"2026-09-10T00:18:23.181Z","args":{"path":".worktrees/286-churn/crates/spt-daemon/src/conn.rs"},"intent":"Mapping connection module"},"id":"6b3d113d","parentId":"6f7d9fbc","timestamp":"2026-09-10T00:18:23.181Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_sciRWA6d1VDbXbF75C6V7j2a|fc_04ba494197989c86016aa1f74f141087d188f8edf4b4209b97","toolName":"read","startedAt":"2026-09-10T00:18:23.181Z","args":{"path":".worktrees/286-churn/crates/spt-daemon/src/livehost.rs"},"intent":"Mapping livehost module"},"id":"1de5e763","parentId":"6b3d113d","timestamp":"2026-09-10T00:18:23.181Z"}
{"type":"message","id":"cbe66330","parentId":"1de5e763","timestamp":"2026-09-10T00:18:23.181Z","message":{"role":"assistant","content":[{"type":"toolCall","id":"call_jXttUDLEeUbjVDJh0FaN6qQx|fc_04ba494197989c86016aa1f74f13f087d1b7850e25879cf204","name":"read","arguments":{"i":"Reading repository rules","path":".worktrees/286-churn/AGENTS.md"},"intent":"Reading repository rules"},{"type":"toolCall","id":"call_Ztw97bKWnzOuGCeWR03Lkl8W|fc_04ba494197989c86016aa1f74f140087d19a31b610d33cd16d","name":"read","arguments":{"i":"Reading complete patch","path":"artifact://62:raw"},"intent":"Reading complete patch"},{"type":"toolCall","id":"call_y8hz8xt4L840evzglqHIPhZK|fc_04ba494197989c86016aa1f74f140887d1b4d3a2b1d4e1e4c5","name":"read","arguments":{"i":"Mapping connection module","path":".worktrees/286-churn/crates/spt-daemon/src/conn.rs"},"intent":"Mapping connection module"},{"type":"toolCall","id":"call_sciRWA6d1VDbXbF75C6V7j2a|fc_04ba494197989c86016aa1f74f141087d188f8edf4b4209b97","name":"read","arguments":{"i":"Mapping livehost module","path":".worktrees/286-churn/crates/spt-daemon/src/livehost.rs"},"intent":"Mapping livehost module"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":3174,"output":143,"cacheRead":14592,"cacheWrite":0,"totalTokens":17909,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1788999493054,"responseId":"resp_04ba494197989c86016aa1f74541e087d1b145a1575d6a20c1","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"fc_04ba494197989c86016aa1f74f13f087d1b7850e25879cf204","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading repository rules\",\"path\":\".worktrees/286-churn/AGENTS.md\"}","call_id":"call_jXttUDLEeUbjVDJh0FaN6qQx","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"read","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f74f140087d19a31b610d33cd16d","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading complete patch\",\"path\":\"artifact://62:raw\"}","call_id":"call_Ztw97bKWnzOuGCeWR03Lkl8W","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"read","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f74f140887d1b4d3a2b1d4e1e4c5","type":"function_call","status":"completed","arguments":"{\"i\":\"Mapping connection module\",\"path\":\".worktrees/286-churn/crates/spt-daemon/src/conn.rs\"}","call_id":"call_y8hz8xt4L840evzglqHIPhZK","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"read","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f74f141087d188f8edf4b4209b97","type":"function_call","status":"completed","arguments":"{\"i\":\"Mapping livehost module\",\"path\":\".worktrees/286-churn/crates/spt-daemon/src/livehost.rs\"}","call_id":"call_sciRWA6d1VDbXbF75C6V7j2a","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"read","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}}]},"duration":10124.544999999925,"ttft":10072.242300000042,"completedAt":1788999503179,"contextSnapshot":{"promptTokens":17766,"nonMessageTokens":13648,"compactionEpoch":0}}}
{"type":"message","id":"89426a0f","parentId":"cbe66330","timestamp":"2026-09-10T00:18:23.234Z","message":{"role":"toolResult","toolCallId":"call_y8hz8xt4L840evzglqHIPhZK|fc_04ba494197989c86016aa1f74f140887d1b4d3a2b1d4e1e4c5","toolName":"read","content":[{"type":"text","text":"//! Bounded, cancelable, poison-on-failure framed writes for one physical\n//! broker connection (UPDATE-WEDGE round-4).\n//!\n//! ## The failure class this kills\n//!\n//! A broker writer used to hold the connection's `Mutex<SendHalf>` guard\n//! ACROSS a blocking [`write_frame`]. On Windows, `interprocess 2.4.2` routes\n//! that send to `WriteFileEx` + `SleepEx(INFINITE, alertable)` — a local-socket\n//! path with **no supported write timeout** (`set_timeout` → `Unsupported`).\n//! A connected consumer that stops reading (an `rc --take` whose process\n//! wedged; a suspended brain) therefore blocks the write INDEFINITELY, and the\n//! logical stall-evict (`stall_evict_controller`) releases only the controller\n//! ROLE — it never cancels the in-flight pipe write or closes the physical\n//! connection, so the writer (and every write queued behind the same gate)\n//! stays wedged until the peer process dies (~128 s in the field capture, on a\n//! brain restart).\n//!\n//! ## The primitive\n//!\n//! [`BrokerConn`] owns the send half outright — it never leaves this module,\n//! so an unbounded bypass write cannot compile. Every write:\n//!\n//! 1. acquires the serialized write GATE with a bounded wait (an ABSOLUTE\n//!    deadline stamped at write entry covers BOTH the gate-wait and the OS\n//!    write);\n//! 2. arms an INDEPENDENT per-connection watchdog thread with that absolute\n//!    deadline, then performs the blocking [`write_frame`] OUT of every lock;\n//! 3. on success, returns the half to the gate.\n//!\n//! The watchdog fires OUT OF BAND at the deadline — it never relies on the\n//! write returning, nor on opportunistic stall-evict sampling: it poisons the\n//! connection and aborts the in-flight I/O (`CancelIoEx` on Windows — the\n//! completion APC makes the wedged `SleepEx` loop return with\n//! `ERROR_OPERATION_ABORTED`; `shutdown(SHUT_RDWR)` on Unix — the blocked\n//! write returns `EPIPE`). A gate-waiter that exhausts its own deadline fires\n//! the same abort (the stuck op it waited behind is the wedge).\n//!\n//! ## Poison / retire invariant (both OSes)\n//!\n//! On deadline, write failure, or cancellation: (1) the WHOLE physical\n//! connection is poisoned, idempotently — a length-prefixed frame may have\n//! been PARTIALLY written, so reuse would corrupt framing; (2) read and write\n//! sides are aborted so `handle_conn` reaches its existing EOF cleanup;\n//! (3) the buffer of the canceled op is released only after the op reports\n//! completion (`write_frame` returns only after the completion APC — the\n//! synchronous return IS the completion handshake); (4) the retired half is\n//! parked in a graveyard slot — never dropped while the watchdog might still\n//! touch the raw handle (drop = `CloseHandle`, and a recycled handle value\n//! must never see a late `CancelIoEx`), and never re-gated; (5) on Windows the\n//! pipe is then severed (`DisconnectNamedPipe` — the server end forces the\n//! client off, so the old client's next access fails, a re-issued server read\n//! fails, and the `interprocess` drop-limbo flush cannot block on unread\n//! data), followed by one more `CancelIoEx` covering a read re-issued in the\n//! cancel→disconnect window.\n//!\n//! No new output queue (the bounded per-sink queues + isolated writer threads\n//! already exist — the block this bounds is BELOW them), and no `PIPE_NOWAIT`\n//! (recorded mid-frame corruption risk). Unix keeps its existing semantics for\n//! healthy connections; only the poison/retire path is new.\n// [impl->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\n\nuse std::io;\n…\nuse crate::frame::Envelope;\n\n/// Stable per-physical-conn identity mint — a monotonic u64 stamped at\n/// [`BrokerConn::new`]. Before this the only conn identity was `Arc::ptr_eq`,\n/// which does not survive a log line: per-line CONN_WRITE_POISONED churn was\n/// UNDECIDABLE between fresh-carrier churn and stderr interleave (the F-039\n/// RCA's terminal fork). Every lifecycle record carries `conn=<id>`.\n// [impl->REQ-CONN-POISON-ATTRIBUTION]\nstatic CONN_ID_SEQ: AtomicU64 = AtomicU64::new(0);\n\n/// The shared monotonic anchor for `mono_ms` stamps. Set explicitly at daemon\n/// startup ([`init_log_anchor`], both broker and brain entries) so every\n/// record in one process shares one base; lazily initialized on first stamp\n/// otherwise (tests, in-process brokers) — correlation needs a COMMON base,\n/// not an absolute one.\nstatic MONO_ANCHOR: OnceLock<Instant> = OnceLock::new();\n\n/// Anchor the monotonic half of [`log_stamp`] at process start — called first\n/// thing by both daemon entries, next to `scrub_identity_env`.\n// [impl->REQ-CONN-POISON-ATTRIBUTION]\npub(crate) fn init_log_anchor() {\n    let _ = MONO_ANCHOR.set(Instant::now());\n}\n\n/// Wall-clock + monotonic stamp for daemon-stderr correlation records\n/// (`wall_ms` = Unix epoch millis; `mono_ms` = millis since [`init_log_anchor`]).\n/// stderr has no per-line timestamps and broker+brain share one file — without\n/// BOTH clocks, interleaved records cannot be ordered (monotonic) or matched\n/// to field incident times (wall).\n// [impl->REQ-CONN-POISON-ATTRIBUTION]\npub(crate) fn log_stamp() -> String { … }\n\n/// Retain the original write instant even when its healthy line is aggregated.\n#[derive(Clone, Copy)]\nstruct LogTime { … }\n\nimpl LogTime { … }\n\nconst HEALTHY_LOG_INTERVAL: Duration = Duration::from_secs(60);\n\nstruct HealthyWindow { … }\n\n/// One process-wide counter window, not one limiter per connection. No thread,\n/// queue or per-conn healthy string: the next event after 60s emits the totals.\n/// Exceptional and role-transition records bypass this window entirely.\n// [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\nstruct LifecycleLog { … }\n\nimpl LifecycleLog { … }\n\n/// Attribution-label cap: facts accrete per role attach (a long-lived brain\n/// conn can be controller + viewer + presence subscriber over its life), and\n/// the label must stay BOUNDED — one truncation marker, then no further growth.\nconst LABEL_CAP: usize = 240;\n\n/// One bounded write's outcome (beyond `Err` = poisoned/timed-out/failed).\n#[derive(Debug, PartialEq, Eq)]\npub(crate) enum ConnWrite { … }\n\n/// The serialized write gate. `half: Some` = free; `None` = a write is in\n/// flight (checked out) or the conn has retired (then `poisoned` is set).\nstruct Gate {\n    half: Option<SendHalf>,\n}\n\n/// Watchdog bookkeeping: the currently in-flight op and its ABSOLUTE deadline.\nstruct Dog { … }\n\n/// The state shared with the watchdog thread (no `Arc<BrokerConn>` cycle —\n/// the watchdog must not keep the conn alive).\nstruct Inner { … }\n\n/// Recover a mutex guard even if a holder panicked — the state these guards\n/// protect is a plain slot/flag pair, valid on every panic boundary.\nfn recover<T>(m: &Mutex<T>) -> MutexGuard<'_, T> { … }\n\n#[cfg(windows)]\nfn raw_of(half: &SendHalf) -> isize { … }\n\n#[cfg(unix)]\nfn raw_of(half: &SendHalf) -> i32 { … }\n\n#[cfg(windows)]\nmod win { … }\n\nimpl Inner { … }\n\n/// The shared, serialized, BOUNDED send side of one physical broker\n/// connection — what [`crate::broker::SharedSend`] now aliases. Owns the\n/// [`SendHalf`] (it never leaves this module) and the per-conn watchdog.\npub(crate) struct BrokerConn { … }\n\nimpl BrokerConn { … }\n\n/// TEST-ONLY determinism lever: holds the serialized write gate (the half is\n/// checked out) so every writer parks AT THE GATE until the pin drops — the\n/// round-4 analog of pinning the old `Mutex<SendHalf>` across a registration\n/// window (the P1c epoch-gate keystone unit relies on it).\n#[cfg(test)]\npub(crate) struct TestGatePin { … }\n\n#[cfg(test)]\nimpl Drop for TestGatePin { … }\n\n#[cfg(test)]\nimpl BrokerConn { … }\n\nimpl Drop for BrokerConn { … }\n\nfn poisoned_err() -> io::Error { … }\n\nfn timeout_err() -> io::Error { … }\n\n#[cfg(test)]\nmod tests { … }\n\n[…800ln elided; re-read needed ranges, e.g. .worktrees/286-churn/crates/spt-daemon/src/conn.rs:63-70,101-104]"}],"details":{"displayContent":{"text":"//! Bounded, cancelable, poison-on-failure framed writes for one physical\n//! broker connection (UPDATE-WEDGE round-4).\n//!\n//! ## The failure class this kills\n//!\n//! A broker writer used to hold the connection's `Mutex<SendHalf>` guard\n//! ACROSS a blocking [`write_frame`]. On Windows, `interprocess 2.4.2` routes\n//! that send to `WriteFileEx` + `SleepEx(INFINITE, alertable)` — a local-socket\n//! path with **no supported write timeout** (`set_timeout` → `Unsupported`).\n//! A connected consumer that stops reading (an `rc --take` whose process\n//! wedged; a suspended brain) therefore blocks the write INDEFINITELY, and the\n//! logical stall-evict (`stall_evict_controller`) releases only the controller\n//! ROLE — it never cancels the in-flight pipe write or closes the physical\n//! connection, so the writer (and every write queued behind the same gate)\n//! stays wedged until the peer process dies (~128 s in the field capture, on a\n//! brain restart).\n//!\n//! ## The primitive\n//!\n//! [`BrokerConn`] owns the send half outright — it never leaves this module,\n//! so an unbounded bypass write cannot compile. Every write:\n//!\n//! 1. acquires the serialized write GATE with a bounded wait (an ABSOLUTE\n//!    deadline stamped at write entry covers BOTH the gate-wait and the OS\n//!    write);\n//! 2. arms an INDEPENDENT per-connection watchdog thread with that absolute\n//!    deadline, then performs the blocking [`write_frame`] OUT of every lock;\n//! 3. on success, returns the half to the gate.\n//!\n//! The watchdog fires OUT OF BAND at the deadline — it never relies on the\n//! write returning, nor on opportunistic stall-evict sampling: it poisons the\n//! connection and aborts the in-flight I/O (`CancelIoEx` on Windows — the\n//! completion APC makes the wedged `SleepEx` loop return with\n//! `ERROR_OPERATION_ABORTED`; `shutdown(SHUT_RDWR)` on Unix — the blocked\n//! write returns `EPIPE`). A gate-waiter that exhausts its own deadline fires\n//! the same abort (the stuck op it waited behind is the wedge).\n//!\n//! ## Poison / retire invariant (both OSes)\n//!\n//! On deadline, write failure, or cancellation: (1) the WHOLE physical\n//! connection is poisoned, idempotently — a length-prefixed frame may have\n//! been PARTIALLY written, so reuse would corrupt framing; (2) read and write\n//! sides are aborted so `handle_conn` reaches its existing EOF cleanup;\n//! (3) the buffer of the canceled op is released only after the op reports\n//! completion (`write_frame` returns only after the completion APC — the\n//! synchronous return IS the completion handshake); (4) the retired half is\n//! parked in a graveyard slot — never dropped while the watchdog might still\n//! touch the raw handle (drop = `CloseHandle`, and a recycled handle value\n//! must never see a late `CancelIoEx`), and never re-gated; (5) on Windows the\n//! pipe is then severed (`DisconnectNamedPipe` — the server end forces the\n//! client off, so the old client's next access fails, a re-issued server read\n//! fails, and the `interprocess` drop-limbo flush cannot block on unread\n//! data), followed by one more `CancelIoEx` covering a read re-issued in the\n//! cancel→disconnect window.\n//!\n//! No new output queue (the bounded per-sink queues + isolated writer threads\n//! already exist — the block this bounds is BELOW them), and no `PIPE_NOWAIT`\n//! (recorded mid-frame corruption risk). Unix keeps its existing semantics for\n//! healthy connections; only the poison/retire path is new.\n// [impl->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\n\nuse std::io;\n…\nuse crate::frame::Envelope;\n\n/// Stable per-physical-conn identity mint — a monotonic u64 stamped at\n/// [`BrokerConn::new`]. Before this the only conn identity was `Arc::ptr_eq`,\n/// which does not survive a log line: per-line CONN_WRITE_POISONED churn was\n/// UNDECIDABLE between fresh-carrier churn and stderr interleave (the F-039\n/// RCA's terminal fork). Every lifecycle record carries `conn=<id>`.\n// [impl->REQ-CONN-POISON-ATTRIBUTION]\nstatic CONN_ID_SEQ: AtomicU64 = AtomicU64::new(0);\n\n/// The shared monotonic anchor for `mono_ms` stamps. Set explicitly at daemon\n/// startup ([`init_log_anchor`], both broker and brain entries) so every\n/// record in one process shares one base; lazily initialized on first stamp\n/// otherwise (tests, in-process brokers) — correlation needs a COMMON base,\n/// not an absolute one.\nstatic MONO_ANCHOR: OnceLock<Instant> = OnceLock::new();\n\n/// Anchor the monotonic half of [`log_stamp`] at process start — called first\n/// thing by both daemon entries, next to `scrub_identity_env`.\n// [impl->REQ-CONN-POISON-ATTRIBUTION]\npub(crate) fn init_log_anchor() {\n    let _ = MONO_ANCHOR.set(Instant::now());\n}\n\n/// Wall-clock + monotonic stamp for daemon-stderr correlation records\n/// (`wall_ms` = Unix epoch millis; `mono_ms` = millis since [`init_log_anchor`]).\n/// stderr has no per-line timestamps and broker+brain share one file — without\n/// BOTH clocks, interleaved records cannot be ordered (monotonic) or matched\n/// to field incident times (wall).\n// [impl->REQ-CONN-POISON-ATTRIBUTION]\npub(crate) fn log_stamp() -> String { … }\n\n/// Retain the original write instant even when its healthy line is aggregated.\n#[derive(Clone, Copy)]\nstruct LogTime { … }\n\nimpl LogTime { … }\n\nconst HEALTHY_LOG_INTERVAL: Duration = Duration::from_secs(60);\n\nstruct HealthyWindow { … }\n\n/// One process-wide counter window, not one limiter per connection. No thread,\n/// queue or per-conn healthy string: the next event after 60s emits the totals.\n/// Exceptional and role-transition records bypass this window entirely.\n// [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\nstruct LifecycleLog { … }\n\nimpl LifecycleLog { … }\n\n/// Attribution-label cap: facts accrete per role attach (a long-lived brain\n/// conn can be controller + viewer + presence subscriber over its life), and\n/// the label must stay BOUNDED — one truncation marker, then no further growth.\nconst LABEL_CAP: usize = 240;\n\n/// One bounded write's outcome (beyond `Err` = poisoned/timed-out/failed).\n#[derive(Debug, PartialEq, Eq)]\npub(crate) enum ConnWrite { … }\n\n/// The serialized write gate. `half: Some` = free; `None` = a write is in\n/// flight (checked out) or the conn has retired (then `poisoned` is set).\nstruct Gate {\n    half: Option<SendHalf>,\n}\n\n/// Watchdog bookkeeping: the currently in-flight op and its ABSOLUTE deadline.\nstruct Dog { … }\n\n/// The state shared with the watchdog thread (no `Arc<BrokerConn>` cycle —\n/// the watchdog must not keep the conn alive).\nstruct Inner { … }\n\n/// Recover a mutex guard even if a holder panicked — the state these guards\n/// protect is a plain slot/flag pair, valid on every panic boundary.\nfn recover<T>(m: &Mutex<T>) -> MutexGuard<'_, T> { … }\n\n#[cfg(windows)]\nfn raw_of(half: &SendHalf) -> isize { … }\n\n#[cfg(unix)]\nfn raw_of(half: &SendHalf) -> i32 { … }\n\n#[cfg(windows)]\nmod win { … }\n\nimpl Inner { … }\n\n/// The shared, serialized, BOUNDED send side of one physical broker\n/// connection — what [`crate::broker::SharedSend`] now aliases. Owns the\n/// [`SendHalf`] (it never leaves this module) and the per-conn watchdog.\npub(crate) struct BrokerConn { … }\n\nimpl BrokerConn { … }\n\n/// TEST-ONLY determinism lever: holds the serialized write gate (the half is\n/// checked out) so every writer parks AT THE GATE until the pin drops — the\n/// round-4 analog of pinning the old `Mutex<SendHalf>` across a registration\n/// window (the P1c epoch-gate keystone unit relies on it).\n#[cfg(test)]\npub(crate) struct TestGatePin { … }\n\n#[cfg(test)]\nimpl Drop for TestGatePin { … }\n\n#[cfg(test)]\nimpl BrokerConn { … }\n\nimpl Drop for BrokerConn { … }\n\nfn poisoned_err() -> io::Error { … }\n\nfn timeout_err() -> io::Error { … }\n\n#[cfg(test)]\nmod tests { … }","startLine":1},"summary":{"lines":178,"elidedSpans":24,"elidedLines":800},"fileSize":43321,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-daemon\\src\\conn.rs"}}},"isError":false,"timestamp":1788999503234}}
{"type":"message","id":"5ddc7f24","parentId":"89426a0f","timestamp":"2026-09-10T00:18:23.250Z","message":{"role":"toolResult","toolCallId":"call_sciRWA6d1VDbXbF75C6V7j2a|fc_04ba494197989c86016aa1f74f141087d188f8edf4b4209b97","toolName":"read","content":[{"type":"text","text":"//! The brain's live-agent Psyche host (M11-W0.1, REQ-DAEMON-1) — re-host the\n//! daemon-hosted LiveAgent lifecycle from disk and reconcile the host-set.\n//!\n//! This is the production caller `run_pulse_loop` lacked: the brain instantiates\n//! a [`BrainLifecycle`] per hosted **live** endpoint and drives its config-paced\n//! pulse loop on a supervised thread. It is the **harness-endpoint mirror** of\n//! [`crate::shellwake::reconcile_once`] (which re-hosts offline shells from\n//! disk): same shape — enumerate Self perches, resolve each manifest from its\n//! `info.json` `adapter` option, hold a stop flag per supervised thread, and run\n//! one bounded reconcile sweep at boot and every tick. Where the shell waker\n//! hosts **offline** shells, this hosts **online** live endpoints — `status` is\n//! daemon-authoritative liveness (CONTEXT:194 / KH 2.5), so re-hosting the\n//! online ones re-hosts exactly what was live before a brain (re)start, and\n//! never resurrects a gracefully-departed (offline) agent.\n//!\n//! **Online-only is staged behind the interim.** The interim `spt api listen`\n//! process still first-hosts its live agents today; those perches are\n//! `status`-ABSENT (not \"online\"), so this scan ignores them — no double-host in\n//! the W0.1→W0.2 window. W0.2 retires the interim and makes the brain the\n//! first-host (it sets `status=online` on a live `listen`/`bind`); from then this\n//! reconcile reacts within one tick, the same way the shell waker reacts to a\n//! CLI-process status flip. The full bind→online→brain-restart→re-host arc is\n//! proven through the real brain process by the W0.4 production int.\n// [impl->REQ-DAEMON-1]\n\nuse std::collections::{BTreeSet, HashMap};\n…\nuse crate::lifecycle::BrainLifecycle;\n\n/// The `InfoJson::state` tag of a live agent — only these get a daemon-hosted\n/// Psyche; a `ready_agent` (messaging-only) never does.\nconst LIVE_AGENT_STATE: &str = \"live_agent\";\n\n/// The reconcile cadence (ms): a lifecycle flip lands in a CLI process, so the\n/// daemon loop is the invariant holder — a flip is at most one tick stale\n/// (mirrors [`crate::shellwake::RECONCILE_INTERVAL_MS`]).\npub const LIVE_RECONCILE_INTERVAL_MS: u64 = 5_000;\n\n/// One brain-hosted live lifecycle: its config-paced pulse-driver thread and the\n/// held stop flag (tripped to un-host on an offline-transition). F-030 W3: the\n/// per-event psyche driver spawns NO resident child, so this holds no psyche\n/// handle and no residency anchor — a psyche is a bounded per-event turn, never a\n/// process to own/reap (the resident residency/reap machinery retired here).\nstruct HostedLife { … }\n\n/// The brain's hosted live-endpoint table, keyed by endpoint id. Shared between\n/// the reconcile loop and anything that must stop a host. Mirror of\n/// [`crate::shellwake::WakeSet`].\n#[derive(Default)]\npub struct LiveSet {\n    inner: Mutex<HashMap<String, HostedLife>>,\n}\n\nimpl LiveSet { … }\n\n/// One reconcile sweep — the brain's live host-set invariant holder, run at boot\n/// and every [`LIVE_RECONCILE_INTERVAL_MS`]. The complete set-diff against the\n/// on-disk **online** live endpoints (mirror of [`crate::shellwake::reconcile_once`]):\n///\n/// - **un-host** (tabled ∧ ¬online) — stop the driver + de-table (exit guard: a\n///   signed-off / removed endpoint must not keep pulsing — thread + dead-agent leak).\n/// - **host** (online ∧ ¬tabled) — spawn the Psyche + start the driver.\n/// - **leave-alone** (online ∧ tabled).\n///\n/// `status` is daemon-authoritative (CONTEXT:194); interim `listen` perches are\n/// `status`-ABSENT so they are never hosted here (no double-host) until W0.2.\n// [impl->REQ-DAEMON-1]\npub fn reconcile_once(\n    owlery: &Path,\n    registered: &[(AdapterRecord, Manifest)],\n    adapters_dir: &Path,\n    set: &Arc<LiveSet>,\n    cfg: &DaemonConfig,\n    reason: StartReason,\n) { … }\n\n/// Reap every top-level parent's REAPABLE nested workers (W-3, REQ-WORKER-REAP) —\n/// the reconcile-cadence GC. Parent liveness is the ONE 2.5 resolver\n/// ([`is_perch_alive`]), NEVER the worker's dead-by-design pid; the TTL is the\n/// clamped `worker_reap_ttl` knob. Best-effort + idempotent: a reap this pass missed\n/// (transient false-dead within the orphan grace) reaps on the next.\n// [impl->REQ-WORKER-REAP]\nfn reap_orphaned_workers(owlery: &Path, cfg: &DaemonConfig) { … }\n\n/// Whether an in-flight wake-resume spawn holds custody of this perch (A-2).\n/// **Custody-only** — the F-030 nested-record discipline: it guards the boot\n/// window against a double-launch and enables a crashed-pre-bind retry, and is\n/// NEVER a liveness input (neither `liveness` nor the A-1 effective-state\n/// derivation reads it; `status=online` is what closes the resume gate).\n///\n/// The record is an IDENTITY PAIR (pid + process start time), not a bare pid\n/// (ADR-0047 decision 1 / KNOWN-HAZARDS 7.51): a pid recycled onto an unrelated\n/// process reads NOT OURS and the discovering reader deletes the record, so a\n/// dead spawn can no longer defer the reconcile forever (the v0.39.4\n/// false-ONLINE field bug). The self-heal lives in the store call — every\n/// consumer here is one honest question.\n// [impl->REQ-RESUME-CUSTODY-IDENTITY]\nfn resume_in_flight(perch: &Path) -> bool {\n    spt_store::resume_custody::resume_in_flight(perch)\n}\n\n/// The pure wake-resume decision for one not-online live-agent perch (A-2). Pure\n/// over the observable inputs so the whole table is a unit test; the caller does\n/// the IO (spawn / pid record / host_error stamp).\n#[derive(Debug, PartialEq, Eq)]\nenum ResumeAction { … }\n\n// [impl->REQ-WAKE-RESUME-LEG]\nfn decide_resume(\n    rest_state: Option<crate::resting::RestState>,\n    resume_pid_alive: bool,\n    last_sid: Option<&str>,\n    recorded_adapter: Option<&str>,\n    adapter_registered: bool,\n) -> ResumeAction { … }\n\n/// A-2 wake-resume leg (REQ-WAKE-RESUME-LEG, ADR-0033): a live-agent perch whose\n/// rest INTENT is Active but whose harness session is cold (status ≠ online) gets\n/// its harness RESUMED from the last ledger session, under the RECORDED adapter\n/// (D-2), mirroring `shellwake::resolve_wake`. This is the daemon LIFT of the thin\n/// `spt wake` edge (ADR-0033: the CLI writes intent, the daemon does the work). It\n/// NEVER stamps `status` — the resumed harness self-binds → online, and the next\n/// reconcile tick hosts its pulse driver. Psyche is untouched (per-event, W3).\n// [impl->REQ-WAKE-RESUME-LEG]\nfn resume_woken_endpoint(\n    perch: &Path,\n    id: &str,\n    info: &spt_store::info::InfoJson,\n    registered: &[(AdapterRecord, Manifest)],\n    adapters_dir: &Path,\n) { … }\n\n/// Re-run an endpoint's harness by resuming its LAST LEDGER session via the adapter's\n/// `[session.resume]` (`launch_harness_brokered_in`, `is_resume=true`) — the shared\n/// action behind BOTH the wake-resume leg ([`resume_woken_endpoint`]) and the\n/// daemon-restart survival pass ([`resume_restart_orphaned_endpoints`]). `reason_tag`\n/// names the trigger in the log breadcrumb (`WAKE` / `DAEMON_RESTART`), scoped so an\n/// operator can tell a woken resume from a restart-survival re-run. Records the custody\n/// guard pid + clears any stale `host_error` on success; stamps `host_error` on\n/// failure. `status` is NEVER stamped here — the harness self-binds → online (liveness\n/// stays daemon-authoritative, the D9-2 posture the resume leg already keeps).\n// [impl->REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL]\n#[allow(clippy::too_many_arguments)] // resume material + the two registry lookups; a struct would obscure the call\nfn launch_ledger_resume(\n    perch: &Path,\n    id: &str,\n    adapter: &str,\n    session_id: &str,\n    cwd: Option<&str>,\n    registered: &[(AdapterRecord, Manifest)],\n    adapters_dir: &Path,\n    reason_tag: &str,\n) { … }\n\n/// The daemon-restart re-run decision (REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL, refined\n/// by RC-RENDER-TRUTH v0.38.1 leg 3 REQ-HOSTING-AUTHORITY-CONTROLLABLE). A pure\n/// function over a candidate perch's facts, so the full belt matrix is unit-tested\n/// without a broker. Only the daemon-restart-ORPHANED, spt-hosted, DEAD set re-runs.\n///\n/// HOSTING AUTHORITY = `controllable`, NOT `state` (leg 3): `controllable == Some(true)`\n/// is the source-definitive broker-PTY authority. The prior `state == live_agent`\n/// rejection was WRONG — an spt-hosted bind over a prior `ready_agent` perch preserves\n/// `state = ready_agent` (REQ-EP-6 prior-type preserve, intentional) while holding a\n/// real broker PTY (`controllable == Some(true)`); gating on `state` left that orphan\n/// unresumed forever. endpoint_survival tables BOTH `live_agent` and `ready_agent`\n/// broker-owned rows — both resume their PTY here; only `live_agent` also gets a Psyche\n/// (state-gated separately at the reconcile pulse-host, `ready` stays no-Psyche). The\n/// belts (in order):\n/// - `status != online` → Skip (wasn't online at the fresh-broker restart — a signed-off\n///   / already-offline endpoint is NOT resurrected).\n/// - `controllable != Some(true)` → Skip (the spt-hosted discriminator, source-definitive:\n///   `Some(true)` = a broker-PTY spt-hosted session ONLY; a user-terminal harness-hosted\n///   agent is `Some(false)`, a legacy/adapterless record `None` — both fall OUT, the safe\n///   direction, no dup risk).\n/// - `in_live_session` → Skip (the broker already hosts a session for it — a surviving\n///   brain-restart session, not a daemon-restart orphan).\n/// - `relay_present` → Skip (BELT 1: a relay address means harness-hosted by definition —\n///   the same `is_spt_hosted_no_relay` discriminator the force-native path uses).\n/// - `custody_pid_alive` → RefuseLivePid (BELT 2: the dup-spawn guard — never re-run over\n///   a live custody process; refuse loudly, the breadcrumb feeds the W4 dup-wake rig).\n// [impl->REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL] [impl->REQ-HOSTING-AUTHORITY-CONTROLLABLE]\n#[derive(Debug, PartialEq, Eq)]\nenum RestartResume { … }\n\n// [impl->REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL] [impl->REQ-HOSTING-AUTHORITY-CONTROLLABLE]\nfn restart_resume_gate(\n    status: Option<&str>,\n    controllable: Option<bool>,\n    in_live_session: bool,\n    relay_present: bool,\n    custody_pid_alive: bool,\n) -> RestartResume { … }\n\n/// LIFECYCLE-TRUTH W3 (REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL): on daemon start, RE-RUN the\n/// harness of every previously-ONLINE spt-hosted endpoint this daemon (re)start\n/// orphaned — so a `daemon stop`+`start` (the apply notice's OWN instruction) no longer\n/// massacres them (seed #6, the field-accept blocker). Runs ONCE at\n/// [`spawn_live_host`] start, BEFORE the reconcile loop's phantom-clear\n/// ([`reconcile_hosted_liveness`]) offlines the stale-online records — which IS the \"was\n/// online at restart\" signal. A brain restart with a SURVIVING broker sees the live\n/// sessions in `live` and re-runs nothing (they are not orphaned). The re-run resumes\n/// from the last ledger session via the shared [`launch_ledger_resume`] tagged\n/// `DAEMON_RESTART`; the mind rides psyche re-host as today. Belts + gate: pure\n/// [`restart_resume_gate`]. Broker unreachable ⇒ skip entirely (never mass-respawn on a\n/// hiccup — the safe direction, same posture as [`reconcile_hosted_liveness`]).\n// [impl->REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL]\npub fn resume_restart_orphaned_endpoints(\n    owlery: &Path,\n    registered: &[(AdapterRecord, Manifest)],\n    adapters_dir: &Path,\n) { … }\n\n/// Host one live endpoint: sweep a stale signoff (3.2), spawn the Psyche (the\n/// daemon online-stamp makes liveness daemon-authoritative), then start the\n/// config-paced pulse driver on its own supervised thread with a held stop flag.\n/// A Psyche-spawn failure tables nothing — the next sweep retries.\n// [impl->REQ-DAEMON-1]\nfn host_one(\n    set: &Arc<LiveSet>,\n    id: &str,\n    lifecycle: BrainLifecycle,\n    session_id: String,\n    reason: StartReason,\n) { … }\n\n/// B2 KEYSTONE — PULL liveness reconcile (REQ-HAZARD-HOSTED-LIVENESS-RECONCILE):\n/// mark OFFLINE any spt-hosted (controllable) `status=online` `live_agent` perch\n/// whose broker session is GONE — the correction for the one-way `status=online`\n/// latch (the broker exit-waiter reaps its session but never clears `info.json`).\n/// This is the LIVE-STATUS analog of the v0.11.0 roster ghost-heal\n/// (`advertise_local` → Offline): same crash-robust, self-healing-next-tick shape.\n///\n/// `live_sessions` is the set of endpoint ids the broker currently hosts a PTY\n/// session for ([`crate::msg::SessionsReply`] via `KIND_SESSIONS`). A perch that\n/// is online + controllable but ABSENT from it has a dead/exited harness → mark it\n/// offline; the existing [`reconcile_once`] stop-side then un-hosts its driver and\n/// the start-side stops re-hosting it.\n///\n/// **Controllable gate (load-bearing):** only `controllable == Some(true)`\n/// (spt-hosted, broker-PTY) perches are reconciled here. A HARNESS-HOSTED relay\n/// live agent is legitimately online with NO broker session (its liveness is the\n/// relay + ready marker), so it is NEVER marked offline by this pass. A\n/// `controllable == None` perch (legacy / dead-psyche like wall-a) is also exempt\n/// here — its liveness/phantom class is not the broker-session arm's concern.\n///\n/// **What now carries which value (REQ-LISTEN-PRESERVES-HOSTING-TOPOLOGY,\n/// 2026-07-21).** `api listen` no longer stamps `Some(false)` — it holds listener\n/// custody and asserts nothing about hosting topology — so a harness-hosted\n/// listener reaches this pass as `None` (fresh) or as whatever an earlier bind\n/// earned. The authority deliberately SURVIVES the broker's reap — retiring it\n/// there reads right (that is the path that knows the PTY ended) but was tried\n/// and falsified: the reap knows only half the question, and clearing there made\n/// every spt-hosted-endpoint-whose-harness-died exempt from this arm, so a\n/// killed harness never went offline (`broker::stamp_reaped` and its unit\n/// `a_reap_clears_the_control_stamps_but_never_the_hosting_authority` pin the\n/// survival). The authority is retired HERE instead, by the relay-holds demotion\n/// below, where both facts are in hand. `Some(false)` survives only on\n/// records written by the pre-fix listen path. Both non-`Some(true)` values are\n/// exempt here, so this gate is unchanged in effect — but do not read\n/// `Some(false)` as \"listen-born\" any more. (The former residency-confirm arm that\n/// once handled it was retired with the resident-Psyche model in F-030; a\n/// genuinely stranded pre-upgrade resident is now reaped by the one-shot\n/// legacy-resident sweep at brain start.)\n///\n/// The caller passes `live_sessions` ONLY when the broker is reachable (it runs\n/// this from boot tick 1 — the B5 boot-gate — so a phantom is never revived, and\n/// SKIPS the pass entirely when the broker is unreachable so a transient hiccup\n/// never mass-offlines). Returns the ids it marked offline.\n// [impl->REQ-HAZARD-HOSTED-LIVENESS-RECONCILE]\n/// The legacy-hybrid self-heal decision (ADR-0041 decision 1), pure over its\n/// observables so the whole table is a unit test. A NON-live-agent row is due\n/// for a terminal heal exactly when it reads `status=online` (the immortal\n/// stamp no reader pid-probes), was born by the listen path\n/// (`controllable == Some(false)` — a gateway `None` or daemon-hosted\n/// `Some(true)` row is somebody else's liveness model), and its recorded\n/// NUMERIC pid probes DEAD (`Some(false)`). A BUSY/absent pid (`None`) is never\n/// guessed at — conservative, next tick retries.\n// [impl->REQ-ENDPOINT-ONLINE-TRUTH]\nfn hybrid_self_heal_due(\n    status: Option<&str>,\n    controllable: Option<bool>,\n    pid_alive: Option<bool>,\n) -> bool {\n    status == Some(STATUS_ONLINE) && controllable == Some(false) && pid_alive == Some(false)\n}\n\n/// Converge ONE harness-hosted endpoint whose relay process is provably gone\n/// (REQ-RELAY-DEATH-CONVERGENCE). Returns whether the endpoint was converged.\n///\n/// The verdict comes from [`spt_store::liveness::relay_liveness`] — the single\n/// relay-liveness authority, reading the process-identity oracle. `Held` and\n/// `Unproven` change nothing (a heal may be missed; it is never mis-fired).\n///\n/// On `Gone` the endpoint's whole projection is retired in the order that keeps a\n/// racing re-bind safe: the RECORD first, under\n/// [`converge_dead_relay`](spt_store::info::converge_dead_relay)'s session+pid\n/// guard, and the ready marker + registered address only once that write actually\n/// applied. `endpoint stop` can afford the opposite order (marker first, as a\n/// stop-accepting gate) because the operator's command IS the death observation;\n/// here the observation is a snapshot, so a re-bound endpoint must be able to keep\n/// its marker and address. What lands is exactly the projection a clean\n/// `endpoint stop` leaves: status offline, rest suspended, no ready marker, no\n/// registered address — so `ready`/`alive`/`address` all converge together instead\n/// of one reader disagreeing with another.\n// [impl->REQ-RELAY-DEATH-CONVERGENCE]\nfn converge_dead_relay(\n    id: &str,\n    perch_path: &Path,\n    info: &spt_store::info::InfoJson,\n    owlery: &Path,\n) -> bool { … }\n\npub fn reconcile_hosted_liveness(owlery: &Path, live_sessions: &BTreeSet<String>) -> Vec<String> { … }\n\n/// The livehost session census is request/reply only, never a subscription.\n/// Retain its healthy carrier across 5s polls; discard it on EVERY query error\n/// so a failed/partially consumed reply can never contaminate the next census.\n/// `None` still means skip reconciliation, not an empty authoritative census.\n// [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n#[derive(Default)]\nstruct SessionPoll { … }\n\nimpl SessionPoll { … }\n\n/// The normalized program basename a perch's adapter would spawn its Psyche as —\n/// the first whitespace token of `psyche_init.command`, file-name only, run through\n/// [`normalize_basename`](spt_runtime::resolve::normalize_basename) (lowercase /\n/// `.exe`-strip / dot-stem) so it keys the same as a live exe basename. `None` for an\n/// empty command. Adapter-AGNOSTIC: the expected program is the MANIFEST-declared psyche\n/// command, NEVER a hardcoded adapter name (glue-model boundary).\n// [impl->REQ-PSYCHE-LEGACY-RESIDENT-SWEEP]\nfn psyche_program_norm(command: &str) -> Option<String> { … }\n\n/// FAIL-SAFE-DECLINE, positive-match-only guard for a stranded legacy-resident\n/// `<id>-psyche` process (F-030 W5, REQ-PSYCHE-LEGACY-RESIDENT-SWEEP, pin 2). Kill IFF\n/// ALL THREE hold: (a) `pid` is alive, (b) its exe basename == the adapter's\n/// MANIFEST-declared psyche program (`expected_prog_norm`, via `normalize_basename` —\n/// never a hardcoded adapter name), AND (c) its command line contains the FULL\n/// `<id>-psyche` id (baked into the wrapper argv via `{id}`; a sibling\n/// `agent-other-psyche` never contains `agent-psyche`). ANY None — dead pid, unreadable\n/// basename, basename mismatch (a same-basename SIBLING), or unreadable cmdline — returns\n/// `false`, i.e. DECLINE to reap: a missed orphan is a visible idle process (cheap), a\n/// wrong-kill is catastrophic (infra never-kill inside the sweep). Pure over the process\n/// signals so the reap/spare/decline verdict is unit-testable.\n// [impl->REQ-PSYCHE-LEGACY-RESIDENT-SWEEP]\nfn legacy_psyche_should_reap(psyche_id: &str, expected_prog_norm: &str, pid: u32) -> bool { … }\n\n/// F-030 W5 ONE-SHOT legacy-resident sweep (REQ-PSYCHE-LEGACY-RESIDENT-SWEEP) — bury the\n/// residency era's corpse, NOT resurrect its machinery. A box upgraded across F-030 W4\n/// can carry a pre-W3 v0.13.x psyche WRAPPER still ALIVE: it registered its `<id>-psyche`\n/// READY perch then BLOCKED FOREVER on `spt ready <id>-psyche --once` (W4's nested-ready\n/// fix converted the old accidental READY_FAIL self-reap into a silent hang — KH 2.6\n/// invisible-loop, one level up). The per-event driver never delivers to a psyche perch,\n/// so it lingers idle, no exit, no fault. This sweep is its INTENDED replacement.\n///\n/// Runs ONCE at brain start (pin 1), BEFORE the first reconcile — so no per-event turn is\n/// in flight yet, and any `<id>-psyche` psyche-program process alive at this instant is\n/// unambiguously a stranded pre-W3 resident (fratricide closed by TIMING: the F-030 shim\n/// is daemon-spawned PER EVENT and its only alive-window is an in-flight resume turn,\n/// which cannot exist before the first reconcile). For each self-perch live-agent id it\n/// reads the pid the wrapper's ready registration recorded in the nested `<id>-psyche`\n/// info.json — a pure F-030 psyche writes NO info.json (custody is a dedicated\n/// `psyche-custody.json`, FORK-2), so `read_pid` is `None` for it and the sweep declines,\n/// making the info.json-with-pid the exact discriminator of a legacy wrapper.\n// [impl->REQ-PSYCHE-LEGACY-RESIDENT-SWEEP]\npub fn sweep_legacy_resident_psyches(\n    owlery: &Path,\n    registered: &[(AdapterRecord, Manifest)],\n    adapters_dir: &Path,\n) { … }\n\n/// Sweep ONE endpoint's stranded legacy-resident psyche (see\n/// [`sweep_legacy_resident_psyches`]). Positive-match-gated by\n/// [`legacy_psyche_should_reap`]; on a match, tree-kill the wrapper (so any harness\n/// children it spawned die with it, not orphaned) THEN clear the stale `<id>-psyche` ready\n/// registration (pin 3 residue: the info.json + `ready` marker the wrapper wrote, else a\n/// dead-pid phantom ready perch lingers — the REMOTE-TRUTH presence-lie class). Loud,\n/// non-fatal.\n// [impl->REQ-PSYCHE-LEGACY-RESIDENT-SWEEP]\nfn sweep_legacy_resident_psyche_for(\n    id: &str,\n    registered: &[(AdapterRecord, Manifest)],\n    adapters_dir: &Path,\n) { … }\n\n/// F-030 W5 ONE-SHOT brain-start GC of stranded `.live-bin` own-copy dirs (design\n/// §W5, [`REQ-PSYCHE-LEGACY-RESIDENT-SWEEP`]). The retired resident spawn kept a\n/// per-endpoint `<perch>/.live-bin/<prog>` PRIVATE COPY of the psyche binary; the\n/// ephemeral per-event turn needs none (it runs from the install-dir/PATH snapshot,\n/// image-lock exposure now turn-bounded), so a cross-version upgrade leaves a multi-MB\n/// `.live-bin` per pre-W5 endpoint forever. Runs ONCE at brain start AFTER\n/// [`sweep_legacy_resident_psyches`] (which kills any wrapper running FROM the image →\n/// the dir then unlocks). A still-image-locked dir → F-028 C1 DISPLACE (a locked exe\n/// can be MOVED even when it can't be unlinked): rename the dir aside to a fresh\n/// `.live-bin.old-N` sibling + best-effort remove; a still-locked displaced dir stays\n/// uniquely-named for a later boot's GC. NEVER delete-fail-abort. Also sweeps stale\n/// `.live-bin.old*` litter a prior locked displace left behind.\n// [impl->REQ-PSYCHE-LEGACY-RESIDENT-SWEEP]\npub fn gc_live_bin_dirs(owlery: &Path) {\n    gc_live_bin_dirs_with(owlery, &|d| std::fs::remove_dir_all(d));\n}\n\n/// [`gc_live_bin_dirs`] over an injected `remove` (the crc_swap injectable-seam\n/// pattern) so the image-locked DISPLACE path is unit-testable deterministically\n/// without a real OS lock: a unit feeds a `remove` that fails for `.live-bin` and\n/// asserts the dir is DISPLACED (renamed aside), never errored/aborted.\nfn gc_live_bin_dirs_with(owlery: &Path, remove: &dyn Fn(&Path) -> std::io::Result<()>) { … }\n\n/// A fresh, non-colliding `.live-bin.old[-N]` sibling of a `.live-bin` dir (mirrors\n/// crc_swap's `fresh_old_path` so a still-mapped prior `.old` is never renamed OVER).\nfn fresh_live_bin_old(live_bin: &Path) -> PathBuf { … }\n\n/// Delete every EXPIRED attachment snapshot and its registry entry together, on\n/// the live host's existing tick.\n///\n/// Serve-time masking already refuses an expired entry, so this is not what\n/// stops an expired link from working — it is what stops\n/// `$SPT_HOME/serve/snapshots/` from growing without bound, which masking alone\n/// never does.\n///\n/// **The order is deliberate: bytes first, then the entry.** Neither step is\n/// atomic with the other, so the question is which crash heals. Deleting the\n/// file and then failing to save leaves an expired entry naming a file that is\n/// gone: the next tick finds it expired again and `remove_file` is idempotent,\n/// so the pair completes on its own. Saving first and then failing to delete\n/// orphans the bytes with no entry left to name them, and no later tick can\n/// repair that — it is exactly the unbounded growth this exists to prevent.\n///\n/// **Only a snapshot is ever deleted.** `ServedKind::Attachment` marks an entry\n/// whose `path` this node minted into its own snapshot store; a `File` or `Dir`\n/// entry's `path` is the USER'S OWN file, which a reaper must never touch. The\n/// containment check is belt-and-braces on that same point: a hand-edited or\n/// corrupted registry must not be able to aim a delete at an arbitrary path.\n// [impl->REQ-WEB-ATTACHMENT-PULL]\nfn reap_expired_attachments() { … }\n\n/// Spawn the brain's live host: one thread sweeping [`reconcile_once`] at boot\n/// and every [`LIVE_RECONCILE_INTERVAL_MS`] until `stop`. The registered set +\n/// config are re-read each sweep (adapter / pulse-period changes land between\n/// ticks). Mirror of [`crate::shellwake::spawn_wake_host`]; the stop flag is held\n/// for symmetry — the threads die with the brain process and the next brain\n/// rehydrates from disk.\n// [impl->REQ-DAEMON-1]\npub fn spawn_live_host(stop: Arc<AtomicBool>, reason: StartReason) -> JoinHandle<()> { … }\n\n#[cfg(test)]\nmod tests { … }\n\n[…2446ln elided; re-read needed ranges, e.g. .worktrees/286-churn/crates/spt-daemon/src/livehost.rs:27-40,57-60]"}],"details":{"displayContent":{"text":"//! The brain's live-agent Psyche host (M11-W0.1, REQ-DAEMON-1) — re-host the\n//! daemon-hosted LiveAgent lifecycle from disk and reconcile the host-set.\n//!\n//! This is the production caller `run_pulse_loop` lacked: the brain instantiates\n//! a [`BrainLifecycle`] per hosted **live** endpoint and drives its config-paced\n//! pulse loop on a supervised thread. It is the **harness-endpoint mirror** of\n//! [`crate::shellwake::reconcile_once`] (which re-hosts offline shells from\n//! disk): same shape — enumerate Self perches, resolve each manifest from its\n//! `info.json` `adapter` option, hold a stop flag per supervised thread, and run\n//! one bounded reconcile sweep at boot and every tick. Where the shell waker\n//! hosts **offline** shells, this hosts **online** live endpoints — `status` is\n//! daemon-authoritative liveness (CONTEXT:194 / KH 2.5), so re-hosting the\n//! online ones re-hosts exactly what was live before a brain (re)start, and\n//! never resurrects a gracefully-departed (offline) agent.\n//!\n//! **Online-only is staged behind the interim.** The interim `spt api listen`\n//! process still first-hosts its live agents today; those perches are\n//! `status`-ABSENT (not \"online\"), so this scan ignores them — no double-host in\n//! the W0.1→W0.2 window. W0.2 retires the interim and makes the brain the\n//! first-host (it sets `status=online` on a live `listen`/`bind`); from then this\n//! reconcile reacts within one tick, the same way the shell waker reacts to a\n//! CLI-process status flip. The full bind→online→brain-restart→re-host arc is\n//! proven through the real brain process by the W0.4 production int.\n// [impl->REQ-DAEMON-1]\n\nuse std::collections::{BTreeSet, HashMap};\n…\nuse crate::lifecycle::BrainLifecycle;\n\n/// The `InfoJson::state` tag of a live agent — only these get a daemon-hosted\n/// Psyche; a `ready_agent` (messaging-only) never does.\nconst LIVE_AGENT_STATE: &str = \"live_agent\";\n\n/// The reconcile cadence (ms): a lifecycle flip lands in a CLI process, so the\n/// daemon loop is the invariant holder — a flip is at most one tick stale\n/// (mirrors [`crate::shellwake::RECONCILE_INTERVAL_MS`]).\npub const LIVE_RECONCILE_INTERVAL_MS: u64 = 5_000;\n\n/// One brain-hosted live lifecycle: its config-paced pulse-driver thread and the\n/// held stop flag (tripped to un-host on an offline-transition). F-030 W3: the\n/// per-event psyche driver spawns NO resident child, so this holds no psyche\n/// handle and no residency anchor — a psyche is a bounded per-event turn, never a\n/// process to own/reap (the resident residency/reap machinery retired here).\nstruct HostedLife { … }\n\n/// The brain's hosted live-endpoint table, keyed by endpoint id. Shared between\n/// the reconcile loop and anything that must stop a host. Mirror of\n/// [`crate::shellwake::WakeSet`].\n#[derive(Default)]\npub struct LiveSet {\n    inner: Mutex<HashMap<String, HostedLife>>,\n}\n\nimpl LiveSet { … }\n\n/// One reconcile sweep — the brain's live host-set invariant holder, run at boot\n/// and every [`LIVE_RECONCILE_INTERVAL_MS`]. The complete set-diff against the\n/// on-disk **online** live endpoints (mirror of [`crate::shellwake::reconcile_once`]):\n///\n/// - **un-host** (tabled ∧ ¬online) — stop the driver + de-table (exit guard: a\n///   signed-off / removed endpoint must not keep pulsing — thread + dead-agent leak).\n/// - **host** (online ∧ ¬tabled) — spawn the Psyche + start the driver.\n/// - **leave-alone** (online ∧ tabled).\n///\n/// `status` is daemon-authoritative (CONTEXT:194); interim `listen` perches are\n/// `status`-ABSENT so they are never hosted here (no double-host) until W0.2.\n// [impl->REQ-DAEMON-1]\npub fn reconcile_once(\n    owlery: &Path,\n    registered: &[(AdapterRecord, Manifest)],\n    adapters_dir: &Path,\n    set: &Arc<LiveSet>,\n    cfg: &DaemonConfig,\n    reason: StartReason,\n) { … }\n\n/// Reap every top-level parent's REAPABLE nested workers (W-3, REQ-WORKER-REAP) —\n/// the reconcile-cadence GC. Parent liveness is the ONE 2.5 resolver\n/// ([`is_perch_alive`]), NEVER the worker's dead-by-design pid; the TTL is the\n/// clamped `worker_reap_ttl` knob. Best-effort + idempotent: a reap this pass missed\n/// (transient false-dead within the orphan grace) reaps on the next.\n// [impl->REQ-WORKER-REAP]\nfn reap_orphaned_workers(owlery: &Path, cfg: &DaemonConfig) { … }\n\n/// Whether an in-flight wake-resume spawn holds custody of this perch (A-2).\n/// **Custody-only** — the F-030 nested-record discipline: it guards the boot\n/// window against a double-launch and enables a crashed-pre-bind retry, and is\n/// NEVER a liveness input (neither `liveness` nor the A-1 effective-state\n/// derivation reads it; `status=online` is what closes the resume gate).\n///\n/// The record is an IDENTITY PAIR (pid + process start time), not a bare pid\n/// (ADR-0047 decision 1 / KNOWN-HAZARDS 7.51): a pid recycled onto an unrelated\n/// process reads NOT OURS and the discovering reader deletes the record, so a\n/// dead spawn can no longer defer the reconcile forever (the v0.39.4\n/// false-ONLINE field bug). The self-heal lives in the store call — every\n/// consumer here is one honest question.\n// [impl->REQ-RESUME-CUSTODY-IDENTITY]\nfn resume_in_flight(perch: &Path) -> bool {\n    spt_store::resume_custody::resume_in_flight(perch)\n}\n\n/// The pure wake-resume decision for one not-online live-agent perch (A-2). Pure\n/// over the observable inputs so the whole table is a unit test; the caller does\n/// the IO (spawn / pid record / host_error stamp).\n#[derive(Debug, PartialEq, Eq)]\nenum ResumeAction { … }\n\n// [impl->REQ-WAKE-RESUME-LEG]\nfn decide_resume(\n    rest_state: Option<crate::resting::RestState>,\n    resume_pid_alive: bool,\n    last_sid: Option<&str>,\n    recorded_adapter: Option<&str>,\n    adapter_registered: bool,\n) -> ResumeAction { … }\n\n/// A-2 wake-resume leg (REQ-WAKE-RESUME-LEG, ADR-0033): a live-agent perch whose\n/// rest INTENT is Active but whose harness session is cold (status ≠ online) gets\n/// its harness RESUMED from the last ledger session, under the RECORDED adapter\n/// (D-2), mirroring `shellwake::resolve_wake`. This is the daemon LIFT of the thin\n/// `spt wake` edge (ADR-0033: the CLI writes intent, the daemon does the work). It\n/// NEVER stamps `status` — the resumed harness self-binds → online, and the next\n/// reconcile tick hosts its pulse driver. Psyche is untouched (per-event, W3).\n// [impl->REQ-WAKE-RESUME-LEG]\nfn resume_woken_endpoint(\n    perch: &Path,\n    id: &str,\n    info: &spt_store::info::InfoJson,\n    registered: &[(AdapterRecord, Manifest)],\n    adapters_dir: &Path,\n) { … }\n\n/// Re-run an endpoint's harness by resuming its LAST LEDGER session via the adapter's\n/// `[session.resume]` (`launch_harness_brokered_in`, `is_resume=true`) — the shared\n/// action behind BOTH the wake-resume leg ([`resume_woken_endpoint`]) and the\n/// daemon-restart survival pass ([`resume_restart_orphaned_endpoints`]). `reason_tag`\n/// names the trigger in the log breadcrumb (`WAKE` / `DAEMON_RESTART`), scoped so an\n/// operator can tell a woken resume from a restart-survival re-run. Records the custody\n/// guard pid + clears any stale `host_error` on success; stamps `host_error` on\n/// failure. `status` is NEVER stamped here — the harness self-binds → online (liveness\n/// stays daemon-authoritative, the D9-2 posture the resume leg already keeps).\n// [impl->REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL]\n#[allow(clippy::too_many_arguments)] // resume material + the two registry lookups; a struct would obscure the call\nfn launch_ledger_resume(\n    perch: &Path,\n    id: &str,\n    adapter: &str,\n    session_id: &str,\n    cwd: Option<&str>,\n    registered: &[(AdapterRecord, Manifest)],\n    adapters_dir: &Path,\n    reason_tag: &str,\n) { … }\n\n/// The daemon-restart re-run decision (REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL, refined\n/// by RC-RENDER-TRUTH v0.38.1 leg 3 REQ-HOSTING-AUTHORITY-CONTROLLABLE). A pure\n/// function over a candidate perch's facts, so the full belt matrix is unit-tested\n/// without a broker. Only the daemon-restart-ORPHANED, spt-hosted, DEAD set re-runs.\n///\n/// HOSTING AUTHORITY = `controllable`, NOT `state` (leg 3): `controllable == Some(true)`\n/// is the source-definitive broker-PTY authority. The prior `state == live_agent`\n/// rejection was WRONG — an spt-hosted bind over a prior `ready_agent` perch preserves\n/// `state = ready_agent` (REQ-EP-6 prior-type preserve, intentional) while holding a\n/// real broker PTY (`controllable == Some(true)`); gating on `state` left that orphan\n/// unresumed forever. endpoint_survival tables BOTH `live_agent` and `ready_agent`\n/// broker-owned rows — both resume their PTY here; only `live_agent` also gets a Psyche\n/// (state-gated separately at the reconcile pulse-host, `ready` stays no-Psyche). The\n/// belts (in order):\n/// - `status != online` → Skip (wasn't online at the fresh-broker restart — a signed-off\n///   / already-offline endpoint is NOT resurrected).\n/// - `controllable != Some(true)` → Skip (the spt-hosted discriminator, source-definitive:\n///   `Some(true)` = a broker-PTY spt-hosted session ONLY; a user-terminal harness-hosted\n///   agent is `Some(false)`, a legacy/adapterless record `None` — both fall OUT, the safe\n///   direction, no dup risk).\n/// - `in_live_session` → Skip (the broker already hosts a session for it — a surviving\n///   brain-restart session, not a daemon-restart orphan).\n/// - `relay_present` → Skip (BELT 1: a relay address means harness-hosted by definition —\n///   the same `is_spt_hosted_no_relay` discriminator the force-native path uses).\n/// - `custody_pid_alive` → RefuseLivePid (BELT 2: the dup-spawn guard — never re-run over\n///   a live custody process; refuse loudly, the breadcrumb feeds the W4 dup-wake rig).\n// [impl->REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL] [impl->REQ-HOSTING-AUTHORITY-CONTROLLABLE]\n#[derive(Debug, PartialEq, Eq)]\nenum RestartResume { … }\n\n// [impl->REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL] [impl->REQ-HOSTING-AUTHORITY-CONTROLLABLE]\nfn restart_resume_gate(\n    status: Option<&str>,\n    controllable: Option<bool>,\n    in_live_session: bool,\n    relay_present: bool,\n    custody_pid_alive: bool,\n) -> RestartResume { … }\n\n/// LIFECYCLE-TRUTH W3 (REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL): on daemon start, RE-RUN the\n/// harness of every previously-ONLINE spt-hosted endpoint this daemon (re)start\n/// orphaned — so a `daemon stop`+`start` (the apply notice's OWN instruction) no longer\n/// massacres them (seed #6, the field-accept blocker). Runs ONCE at\n/// [`spawn_live_host`] start, BEFORE the reconcile loop's phantom-clear\n/// ([`reconcile_hosted_liveness`]) offlines the stale-online records — which IS the \"was\n/// online at restart\" signal. A brain restart with a SURVIVING broker sees the live\n/// sessions in `live` and re-runs nothing (they are not orphaned). The re-run resumes\n/// from the last ledger session via the shared [`launch_ledger_resume`] tagged\n/// `DAEMON_RESTART`; the mind rides psyche re-host as today. Belts + gate: pure\n/// [`restart_resume_gate`]. Broker unreachable ⇒ skip entirely (never mass-respawn on a\n/// hiccup — the safe direction, same posture as [`reconcile_hosted_liveness`]).\n// [impl->REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL]\npub fn resume_restart_orphaned_endpoints(\n    owlery: &Path,\n    registered: &[(AdapterRecord, Manifest)],\n    adapters_dir: &Path,\n) { … }\n\n/// Host one live endpoint: sweep a stale signoff (3.2), spawn the Psyche (the\n/// daemon online-stamp makes liveness daemon-authoritative), then start the\n/// config-paced pulse driver on its own supervised thread with a held stop flag.\n/// A Psyche-spawn failure tables nothing — the next sweep retries.\n// [impl->REQ-DAEMON-1]\nfn host_one(\n    set: &Arc<LiveSet>,\n    id: &str,\n    lifecycle: BrainLifecycle,\n    session_id: String,\n    reason: StartReason,\n) { … }\n\n/// B2 KEYSTONE — PULL liveness reconcile (REQ-HAZARD-HOSTED-LIVENESS-RECONCILE):\n/// mark OFFLINE any spt-hosted (controllable) `status=online` `live_agent` perch\n/// whose broker session is GONE — the correction for the one-way `status=online`\n/// latch (the broker exit-waiter reaps its session but never clears `info.json`).\n/// This is the LIVE-STATUS analog of the v0.11.0 roster ghost-heal\n/// (`advertise_local` → Offline): same crash-robust, self-healing-next-tick shape.\n///\n/// `live_sessions` is the set of endpoint ids the broker currently hosts a PTY\n/// session for ([`crate::msg::SessionsReply`] via `KIND_SESSIONS`). A perch that\n/// is online + controllable but ABSENT from it has a dead/exited harness → mark it\n/// offline; the existing [`reconcile_once`] stop-side then un-hosts its driver and\n/// the start-side stops re-hosting it.\n///\n/// **Controllable gate (load-bearing):** only `controllable == Some(true)`\n/// (spt-hosted, broker-PTY) perches are reconciled here. A HARNESS-HOSTED relay\n/// live agent is legitimately online with NO broker session (its liveness is the\n/// relay + ready marker), so it is NEVER marked offline by this pass. A\n/// `controllable == None` perch (legacy / dead-psyche like wall-a) is also exempt\n/// here — its liveness/phantom class is not the broker-session arm's concern.\n///\n/// **What now carries which value (REQ-LISTEN-PRESERVES-HOSTING-TOPOLOGY,\n/// 2026-07-21).** `api listen` no longer stamps `Some(false)` — it holds listener\n/// custody and asserts nothing about hosting topology — so a harness-hosted\n/// listener reaches this pass as `None` (fresh) or as whatever an earlier bind\n/// earned. The authority deliberately SURVIVES the broker's reap — retiring it\n/// there reads right (that is the path that knows the PTY ended) but was tried\n/// and falsified: the reap knows only half the question, and clearing there made\n/// every spt-hosted-endpoint-whose-harness-died exempt from this arm, so a\n/// killed harness never went offline (`broker::stamp_reaped` and its unit\n/// `a_reap_clears_the_control_stamps_but_never_the_hosting_authority` pin the\n/// survival). The authority is retired HERE instead, by the relay-holds demotion\n/// below, where both facts are in hand. `Some(false)` survives only on\n/// records written by the pre-fix listen path. Both non-`Some(true)` values are\n/// exempt here, so this gate is unchanged in effect — but do not read\n/// `Some(false)` as \"listen-born\" any more. (The former residency-confirm arm that\n/// once handled it was retired with the resident-Psyche model in F-030; a\n/// genuinely stranded pre-upgrade resident is now reaped by the one-shot\n/// legacy-resident sweep at brain start.)\n///\n/// The caller passes `live_sessions` ONLY when the broker is reachable (it runs\n/// this from boot tick 1 — the B5 boot-gate — so a phantom is never revived, and\n/// SKIPS the pass entirely when the broker is unreachable so a transient hiccup\n/// never mass-offlines). Returns the ids it marked offline.\n// [impl->REQ-HAZARD-HOSTED-LIVENESS-RECONCILE]\n/// The legacy-hybrid self-heal decision (ADR-0041 decision 1), pure over its\n/// observables so the whole table is a unit test. A NON-live-agent row is due\n/// for a terminal heal exactly when it reads `status=online` (the immortal\n/// stamp no reader pid-probes), was born by the listen path\n/// (`controllable == Some(false)` — a gateway `None` or daemon-hosted\n/// `Some(true)` row is somebody else's liveness model), and its recorded\n/// NUMERIC pid probes DEAD (`Some(false)`). A BUSY/absent pid (`None`) is never\n/// guessed at — conservative, next tick retries.\n// [impl->REQ-ENDPOINT-ONLINE-TRUTH]\nfn hybrid_self_heal_due(\n    status: Option<&str>,\n    controllable: Option<bool>,\n    pid_alive: Option<bool>,\n) -> bool {\n    status == Some(STATUS_ONLINE) && controllable == Some(false) && pid_alive == Some(false)\n}\n\n/// Converge ONE harness-hosted endpoint whose relay process is provably gone\n/// (REQ-RELAY-DEATH-CONVERGENCE). Returns whether the endpoint was converged.\n///\n/// The verdict comes from [`spt_store::liveness::relay_liveness`] — the single\n/// relay-liveness authority, reading the process-identity oracle. `Held` and\n/// `Unproven` change nothing (a heal may be missed; it is never mis-fired).\n///\n/// On `Gone` the endpoint's whole projection is retired in the order that keeps a\n/// racing re-bind safe: the RECORD first, under\n/// [`converge_dead_relay`](spt_store::info::converge_dead_relay)'s session+pid\n/// guard, and the ready marker + registered address only once that write actually\n/// applied. `endpoint stop` can afford the opposite order (marker first, as a\n/// stop-accepting gate) because the operator's command IS the death observation;\n/// here the observation is a snapshot, so a re-bound endpoint must be able to keep\n/// its marker and address. What lands is exactly the projection a clean\n/// `endpoint stop` leaves: status offline, rest suspended, no ready marker, no\n/// registered address — so `ready`/`alive`/`address` all converge together instead\n/// of one reader disagreeing with another.\n// [impl->REQ-RELAY-DEATH-CONVERGENCE]\nfn converge_dead_relay(\n    id: &str,\n    perch_path: &Path,\n    info: &spt_store::info::InfoJson,\n    owlery: &Path,\n) -> bool { … }\n\npub fn reconcile_hosted_liveness(owlery: &Path, live_sessions: &BTreeSet<String>) -> Vec<String> { … }\n\n/// The livehost session census is request/reply only, never a subscription.\n/// Retain its healthy carrier across 5s polls; discard it on EVERY query error\n/// so a failed/partially consumed reply can never contaminate the next census.\n/// `None` still means skip reconciliation, not an empty authoritative census.\n// [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n#[derive(Default)]\nstruct SessionPoll { … }\n\nimpl SessionPoll { … }\n\n/// The normalized program basename a perch's adapter would spawn its Psyche as —\n/// the first whitespace token of `psyche_init.command`, file-name only, run through\n/// [`normalize_basename`](spt_runtime::resolve::normalize_basename) (lowercase /\n/// `.exe`-strip / dot-stem) so it keys the same as a live exe basename. `None` for an\n/// empty command. Adapter-AGNOSTIC: the expected program is the MANIFEST-declared psyche\n/// command, NEVER a hardcoded adapter name (glue-model boundary).\n// [impl->REQ-PSYCHE-LEGACY-RESIDENT-SWEEP]\nfn psyche_program_norm(command: &str) -> Option<String> { … }\n\n/// FAIL-SAFE-DECLINE, positive-match-only guard for a stranded legacy-resident\n/// `<id>-psyche` process (F-030 W5, REQ-PSYCHE-LEGACY-RESIDENT-SWEEP, pin 2). Kill IFF\n/// ALL THREE hold: (a) `pid` is alive, (b) its exe basename == the adapter's\n/// MANIFEST-declared psyche program (`expected_prog_norm`, via `normalize_basename` —\n/// never a hardcoded adapter name), AND (c) its command line contains the FULL\n/// `<id>-psyche` id (baked into the wrapper argv via `{id}`; a sibling\n/// `agent-other-psyche` never contains `agent-psyche`). ANY None — dead pid, unreadable\n/// basename, basename mismatch (a same-basename SIBLING), or unreadable cmdline — returns\n/// `false`, i.e. DECLINE to reap: a missed orphan is a visible idle process (cheap), a\n/// wrong-kill is catastrophic (infra never-kill inside the sweep). Pure over the process\n/// signals so the reap/spare/decline verdict is unit-testable.\n// [impl->REQ-PSYCHE-LEGACY-RESIDENT-SWEEP]\nfn legacy_psyche_should_reap(psyche_id: &str, expected_prog_norm: &str, pid: u32) -> bool { … }\n\n/// F-030 W5 ONE-SHOT legacy-resident sweep (REQ-PSYCHE-LEGACY-RESIDENT-SWEEP) — bury the\n/// residency era's corpse, NOT resurrect its machinery. A box upgraded across F-030 W4\n/// can carry a pre-W3 v0.13.x psyche WRAPPER still ALIVE: it registered its `<id>-psyche`\n/// READY perch then BLOCKED FOREVER on `spt ready <id>-psyche --once` (W4's nested-ready\n/// fix converted the old accidental READY_FAIL self-reap into a silent hang — KH 2.6\n/// invisible-loop, one level up). The per-event driver never delivers to a psyche perch,\n/// so it lingers idle, no exit, no fault. This sweep is its INTENDED replacement.\n///\n/// Runs ONCE at brain start (pin 1), BEFORE the first reconcile — so no per-event turn is\n/// in flight yet, and any `<id>-psyche` psyche-program process alive at this instant is\n/// unambiguously a stranded pre-W3 resident (fratricide closed by TIMING: the F-030 shim\n/// is daemon-spawned PER EVENT and its only alive-window is an in-flight resume turn,\n/// which cannot exist before the first reconcile). For each self-perch live-agent id it\n/// reads the pid the wrapper's ready registration recorded in the nested `<id>-psyche`\n/// info.json — a pure F-030 psyche writes NO info.json (custody is a dedicated\n/// `psyche-custody.json`, FORK-2), so `read_pid` is `None` for it and the sweep declines,\n/// making the info.json-with-pid the exact discriminator of a legacy wrapper.\n// [impl->REQ-PSYCHE-LEGACY-RESIDENT-SWEEP]\npub fn sweep_legacy_resident_psyches(\n    owlery: &Path,\n    registered: &[(AdapterRecord, Manifest)],\n    adapters_dir: &Path,\n) { … }\n\n/// Sweep ONE endpoint's stranded legacy-resident psyche (see\n/// [`sweep_legacy_resident_psyches`]). Positive-match-gated by\n/// [`legacy_psyche_should_reap`]; on a match, tree-kill the wrapper (so any harness\n/// children it spawned die with it, not orphaned) THEN clear the stale `<id>-psyche` ready\n/// registration (pin 3 residue: the info.json + `ready` marker the wrapper wrote, else a\n/// dead-pid phantom ready perch lingers — the REMOTE-TRUTH presence-lie class). Loud,\n/// non-fatal.\n// [impl->REQ-PSYCHE-LEGACY-RESIDENT-SWEEP]\nfn sweep_legacy_resident_psyche_for(\n    id: &str,\n    registered: &[(AdapterRecord, Manifest)],\n    adapters_dir: &Path,\n) { … }\n\n/// F-030 W5 ONE-SHOT brain-start GC of stranded `.live-bin` own-copy dirs (design\n/// §W5, [`REQ-PSYCHE-LEGACY-RESIDENT-SWEEP`]). The retired resident spawn kept a\n/// per-endpoint `<perch>/.live-bin/<prog>` PRIVATE COPY of the psyche binary; the\n/// ephemeral per-event turn needs none (it runs from the install-dir/PATH snapshot,\n/// image-lock exposure now turn-bounded), so a cross-version upgrade leaves a multi-MB\n/// `.live-bin` per pre-W5 endpoint forever. Runs ONCE at brain start AFTER\n/// [`sweep_legacy_resident_psyches`] (which kills any wrapper running FROM the image →\n/// the dir then unlocks). A still-image-locked dir → F-028 C1 DISPLACE (a locked exe\n/// can be MOVED even when it can't be unlinked): rename the dir aside to a fresh\n/// `.live-bin.old-N` sibling + best-effort remove; a still-locked displaced dir stays\n/// uniquely-named for a later boot's GC. NEVER delete-fail-abort. Also sweeps stale\n/// `.live-bin.old*` litter a prior locked displace left behind.\n// [impl->REQ-PSYCHE-LEGACY-RESIDENT-SWEEP]\npub fn gc_live_bin_dirs(owlery: &Path) {\n    gc_live_bin_dirs_with(owlery, &|d| std::fs::remove_dir_all(d));\n}\n\n/// [`gc_live_bin_dirs`] over an injected `remove` (the crc_swap injectable-seam\n/// pattern) so the image-locked DISPLACE path is unit-testable deterministically\n/// without a real OS lock: a unit feeds a `remove` that fails for `.live-bin` and\n/// asserts the dir is DISPLACED (renamed aside), never errored/aborted.\nfn gc_live_bin_dirs_with(owlery: &Path, remove: &dyn Fn(&Path) -> std::io::Result<()>) { … }\n\n/// A fresh, non-colliding `.live-bin.old[-N]` sibling of a `.live-bin` dir (mirrors\n/// crc_swap's `fresh_old_path` so a still-mapped prior `.old` is never renamed OVER).\nfn fresh_live_bin_old(live_bin: &Path) -> PathBuf { … }\n\n/// Delete every EXPIRED attachment snapshot and its registry entry together, on\n/// the live host's existing tick.\n///\n/// Serve-time masking already refuses an expired entry, so this is not what\n/// stops an expired link from working — it is what stops\n/// `$SPT_HOME/serve/snapshots/` from growing without bound, which masking alone\n/// never does.\n///\n/// **The order is deliberate: bytes first, then the entry.** Neither step is\n/// atomic with the other, so the question is which crash heals. Deleting the\n/// file and then failing to save leaves an expired entry naming a file that is\n/// gone: the next tick finds it expired again and `remove_file` is idempotent,\n/// so the pair completes on its own. Saving first and then failing to delete\n/// orphans the bytes with no entry left to name them, and no later tick can\n/// repair that — it is exactly the unbounded growth this exists to prevent.\n///\n/// **Only a snapshot is ever deleted.** `ServedKind::Attachment` marks an entry\n/// whose `path` this node minted into its own snapshot store; a `File` or `Dir`\n/// entry's `path` is the USER'S OWN file, which a reaper must never touch. The\n/// containment check is belt-and-braces on that same point: a hand-edited or\n/// corrupted registry must not be able to aim a delete at an arbitrary path.\n// [impl->REQ-WEB-ATTACHMENT-PULL]\nfn reap_expired_attachments() { … }\n\n/// Spawn the brain's live host: one thread sweeping [`reconcile_once`] at boot\n/// and every [`LIVE_RECONCILE_INTERVAL_MS`] until `stop`. The registered set +\n/// config are re-read each sweep (adapter / pulse-period changes land between\n/// ticks). Mirror of [`crate::shellwake::spawn_wake_host`]; the stop flag is held\n/// for symmetry — the threads die with the brain process and the next brain\n/// rehydrates from disk.\n// [impl->REQ-DAEMON-1]\npub fn spawn_live_host(stop: Arc<AtomicBool>, reason: StartReason) -> JoinHandle<()> { … }\n\n#[cfg(test)]\nmod tests { … }","startLine":1},"summary":{"lines":447,"elidedSpans":26,"elidedLines":2446},"fileSize":150325,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-daemon\\src\\livehost.rs"}}},"isError":false,"timestamp":1788999503250}}
{"type":"message","id":"c9717f8a","parentId":"5ddc7f24","timestamp":"2026-09-10T00:18:23.256Z","message":{"role":"toolResult","toolCallId":"call_Ztw97bKWnzOuGCeWR03Lkl8W|fc_04ba494197989c86016aa1f74f140087d19a31b610d33cd16d","toolName":"read","content":[{"type":"text","text":"diff --git a/CHANGELOG.md b/CHANGELOG.md\nindex eafe18d6..5b4a7156 100644\n--- a/CHANGELOG.md\n+++ b/CHANGELOG.md\n@@ -10,6 +10,17 @@ breaks something, or changes the observable behavior of existing surfaces\n broadly; **patch** for fixes, and for additive opt-in capability — a new key,\n flag, or page that no existing user can encounter without opting into it.\n \n+## [Unreleased]\n+\n+### Fixed\n+\n+<!-- [doc->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED] -->\n+- Daemon logs summarize healthy IPC connection starts and closes once per\n+  minute instead of logging each one. Failures and subscriber lifecycle\n+  events keep their per-connection attribution, including the original\n+  first-write time. Livehost session polling now reuses its healthy\n+  connection and names the reason when it opens a replacement.\n+\n ## [0.68.0] - 2026-09-08\n \n Web serving. Files, directories, adapter documentation and the changelog are\ndiff --git a/crates/spt-daemon/src/conn.rs b/crates/spt-daemon/src/conn.rs\nindex 9c51ad4e..163fecb6 100644\n--- a/crates/spt-daemon/src/conn.rs\n+++ b/crates/spt-daemon/src/conn.rs\n@@ -61,7 +61,7 @@\n \n use std::io;\n use std::sync::atomic::{AtomicBool, AtomicU64, Ordering};\n-use std::sync::{Arc, Condvar, Mutex, MutexGuard, OnceLock};\n+use std::sync::{Arc, Condvar, LazyLock, Mutex, MutexGuard, OnceLock};\n use std::thread::{self, JoinHandle};\n use std::time::{Duration, Instant, SystemTime, UNIX_EPOCH};\n \n@@ -99,12 +99,107 @@ pub(crate) fn init_log_anchor() {\n /// to field incident times (wall).\n // [impl->REQ-CONN-POISON-ATTRIBUTION]\n pub(crate) fn log_stamp() -> String {\n-    let wall_ms = SystemTime::now()\n-        .duration_since(UNIX_EPOCH)\n-        .map(|d| d.as_millis())\n-        .unwrap_or(0);\n-    let mono_ms = MONO_ANCHOR.get_or_init(Instant::now).elapsed().as_millis();\n-    format!(\"wall_ms={wall_ms} mono_ms={mono_ms}\")\n+    let stamp = LogTime::now();\n+    format!(\"wall_ms={} mono_ms={}\", stamp.wall_ms, stamp.mono_ms)\n+}\n+\n+/// Retain the original write instant even when its healthy line is aggregated.\n+#[derive(Clone, Copy)]\n+struct LogTime {\n+    wall_ms: u128,\n+    mono_ms: u128,\n+}\n+\n+impl LogTime {\n+    fn now() -> Self {\n+        Self {\n+            wall_ms: SystemTime::now()\n+                .duration_since(UNIX_EPOCH)\n+                .map(|d| d.as_millis())\n+                .unwrap_or(0),\n+            mono_ms: MONO_ANCHOR.get_or_init(Instant::now).elapsed().as_millis(),\n+        }\n+    }\n+}\n+\n+const HEALTHY_LOG_INTERVAL: Duration = Duration::from_secs(60);\n+\n+struct HealthyWindow {\n+    since: Instant,\n+    starts: u64,\n+    closes: u64,\n+}\n+\n+/// One process-wide counter window, not one limiter per connection. No thread,\n+/// queue or per-conn healthy string: the next event after 60s emits the totals.\n+/// Exceptional and role-transition records bypass this window entirely.\n+// [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+struct LifecycleLog {\n+    healthy: Mutex<HealthyWindow>,\n+    #[cfg(test)]\n+    capture: Option<Mutex<Vec<String>>>,\n+    #[cfg(test)]\n+    clock: Option<Mutex<Instant>>,\n+}\n+\n+impl LifecycleLog {\n+    fn new() -> Self {\n+        Self {\n+            healthy: Mutex::new(HealthyWindow {\n+                since: Instant::now(),\n+                starts: 0,\n+                closes: 0,\n+            }),\n+            #[cfg(test)]\n+            capture: None,\n+            #[cfg(test)]\n+            clock: None,\n+        }\n+    }\n+\n+    #[cfg(test)]\n+    fn captured(now: Instant) -> Self {\n+        let mut log = Self::new();\n+        log.healthy.get_mut().unwrap().since = now;\n+        log.capture = Some(Mutex::new(Vec::new()));\n+        log.clock = Some(Mutex::new(now));\n+        log\n+    }\n+\n+    fn emit(&self, args: std::fmt::Arguments<'_>) {\n+        #[cfg(test)]\n+        if let Some(lines) = &self.capture {\n+            recover(lines).push(args.to_string());\n+            return;\n+        }\n+        spt_proto::emit_line_err!(\"{args}\");\n+    }\n+\n+    fn healthy(&self, start: bool) {\n+        let mut window = recover(&self.healthy);\n+        // Read time under the counter lock so concurrent callers cannot move\n+        // the window backwards after waiting for a newer caller to flush.\n+        let now = Instant::now();\n+        #[cfg(test)]\n+        let now = self.clock.as_ref().map(|c| *recover(c)).unwrap_or(now);\n+        if start {\n+            window.starts = window.starts.saturating_add(1);\n+        } else {\n+            window.closes = window.closes.saturating_add(1);\n+        }\n+        let elapsed = now.saturating_duration_since(window.since);\n+        if elapsed < HEALTHY_LOG_INTERVAL {\n+            return;\n+        }\n+        let (starts, closes) = (window.starts, window.closes);\n+        *window = HealthyWindow { since: now, starts: 0, closes: 0 };\n+        drop(window);\n+        self.emit(format_args!(\n+            \"CONN_LIFECYCLE_SUMMARY: scope=broker-ipc {} interval_ms={} starts={starts} closes={closes} \\\n+             [REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\",\n+            log_stamp(), elapsed.as_millis()\n+        ));\n+    }\n }\n \n /// Attribution-label cap: facts accrete per role attach (a long-lived brain\n@@ -144,8 +239,9 @@ struct Inner {\n     /// Accreted attribution facts (subscriber role, endpoint/session where\n     /// known — [`BrokerConn::describe`]). Bounded by [`LABEL_CAP`].\n     label: Mutex<String>,\n-    /// Latch: the once-per-conn `write-start` lifecycle record fired.\n-    first_write_logged: AtomicBool,\n+    /// Once-per-conn original first committed write time; retained on failures.\n+    first_write: OnceLock<LogTime>,\n+    lifecycle_log: Arc<LifecycleLog>,\n     gate: Mutex<Gate>,\n     gate_cv: Condvar,\n     dog: Mutex<Dog>,\n@@ -223,7 +319,13 @@ impl Inner {\n     fn attribution(&self) -> String {\n         let label = recover(&self.label);\n         let facts: &str = if label.is_empty() { \"role=unattributed\" } else { &label };\n-        format!(\"conn={} {} {}\", self.id, facts, log_stamp())\n+        match self.first_write.get() {\n+            Some(start) => format!(\n+                \"conn={} {} {} first_write_wall_ms={} first_write_mono_ms={}\",\n+                self.id, facts, log_stamp(), start.wall_ms, start.mono_ms\n+            ),\n+            None => format!(\"conn={} {} {} first_write=none\", self.id, facts, log_stamp()),\n+        }\n     }\n \n     /// Render the once-per-conn retirement record — the F-039 token split\n@@ -261,18 +363,17 @@ impl Inner {\n         }\n     }\n \n-    /// One BOUNDED per-conn lifecycle record (leg d, doyle-confirmed\n-    /// UNCONDITIONAL): `write-start` (first write only), `transport-close`\n-    /// (drop), `writer-exit` / `*-attach` / `*-replaced` (emitted by the\n-    /// broker's sink machinery through [`BrokerConn::lifecycle_event`]). The\n-    /// timeout/cancel leg is the retirement record itself.\n+    /// Per-conn exceptional and role-transition records remain unconditional.\n+    /// releases#286 supersedes only healthy start/close emission, which uses\n+    /// the process-wide interval counter instead. Retirement retains the\n+    /// original first-write time even when the start line was aggregated.\n     // [impl->REQ-CONN-POISON-ATTRIBUTION]\n     fn lifecycle(&self, event: &str, extra: &str) {\n         let sep = if extra.is_empty() { \"\" } else { \" \" };\n-        spt_proto::emit_line_err!(\n+        self.lifecycle_log.emit(format_args!(\n             \"CONN_LIFECYCLE: {} event={event}{sep}{extra} [REQ-CONN-POISON-ATTRIBUTION]\",\n             self.attribution()\n-        );\n+        ));\n     }\n \n     /// Abort the connection's in-flight I/O OUT OF BAND (both directions —\n@@ -287,10 +388,10 @@ impl Inner {\n             // First poison of this conn: loud, once. The token is class-split\n             // (F-039 leg a): deadline → CONN_WRITE_POISONED (the wedge\n             // observable), organic fast-fail → CONN_WRITE_RETIRED.\n-            eprintln!(\n+            self.lifecycle_log.emit(format_args!(\n                 \"{}\",\n                 self.render_retirement(self.timed_out.load(Ordering::Acquire), cause)\n-            );\n+            ));\n         }\n         #[cfg(windows)]\n         // SAFETY: scalar kernel32 call on a handle whose owning half is kept\n@@ -424,11 +525,17 @@ impl BrokerConn {\n     /// deadline (gate-wait + OS write completion) — the broker passes its\n     /// `brain_write_deadline()` (`SPT_BRAIN_WRITE_DEADLINE_MS` knob).\n     pub(crate) fn new(half: SendHalf, bound: Duration) -> Self {\n+        static LOG: LazyLock<Arc<LifecycleLog>> = LazyLock::new(|| Arc::new(LifecycleLog::new()));\n+        Self::with_log(half, bound, Arc::clone(&LOG))\n+    }\n+\n+    fn with_log(half: SendHalf, bound: Duration, lifecycle_log: Arc<LifecycleLog>) -> Self {\n         let raw = raw_of(&half);\n         let inner = Arc::new(Inner {\n             id: CONN_ID_SEQ.fetch_add(1, Ordering::Relaxed) + 1,\n             label: Mutex::new(String::new()),\n-            first_write_logged: AtomicBool::new(false),\n+            first_write: OnceLock::new(),\n+            lifecycle_log,\n             gate: Mutex::new(Gate { half: Some(half) }),\n             gate_cv: Condvar::new(),\n             dog: Mutex::new(Dog {\n@@ -539,12 +646,12 @@ impl BrokerConn {\n             return Ok(ConnWrite::Superseded);\n         }\n \n-        // Once-per-conn `write-start` lifecycle record (leg d): marks the conn's\n-        // first committed write, so a conn that later retires can be correlated\n-        // to when it went active (fresh-carrier churn shows as open→start→retire\n-        // triplets with fresh ids).\n-        if !inner.first_write_logged.swap(true, Ordering::AcqRel) {\n-            inner.lifecycle(\"write-start\", \"\");\n+        // The gate serializes first-write initialization. Save the instant\n+        // BEFORE I/O so a later failure still names when this conn went active.\n+        // [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+        if inner.first_write.get().is_none() {\n+            let _ = inner.first_write.set(LogTime::now());\n+            inner.lifecycle_log.healthy(true);\n         }\n \n         // ── Check out the half + arm the watchdog, then write OUT of locks. ──\n@@ -638,17 +745,20 @@ impl BrokerConn {\n \n impl Drop for BrokerConn {\n     fn drop(&mut self) {\n-        // Once-per-conn `transport-close` lifecycle record (leg d): the handle\n-        // is about to close — the terminal record of this conn id's life, with\n-        // how it ended (organic EOF vs poisoned vs deadline-poisoned).\n-        self.inner.lifecycle(\n-            \"transport-close\",\n-            &format!(\n-                \"poisoned={} timed_out={}\",\n-                self.inner.poisoned.load(Ordering::Acquire),\n-                self.inner.timed_out.load(Ordering::Acquire)\n-            ),\n-        );\n+        // Only a healthy physical close is counted instead of logged per conn.\n+        // Error/timeout closes keep their full attribution and original start.\n+        // [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+        if self.inner.poisoned.load(Ordering::Acquire) {\n+            self.inner.lifecycle(\n+                \"transport-close\",\n+                &format!(\n+                    \"poisoned=true timed_out={}\",\n+                    self.inner.timed_out.load(Ordering::Acquire)\n+                ),\n+            );\n+        } else {\n+            self.inner.lifecycle_log.healthy(false);\n+        }\n         {\n             let mut d = recover(&self.inner.dog);\n             d.shutdown = true;\n@@ -688,6 +798,17 @@ mod tests {\n     /// primitive, the client end returned so a test can hold it alive or drop\n     /// it to force an ORGANIC write failure.\n     fn make_conn() -> (BrokerConn, Stream) {\n+        make_conn_with_log(Arc::new(LifecycleLog::new()))\n+    }\n+\n+    fn make_conn_with_log(log: Arc<LifecycleLog>) -> (BrokerConn, Stream) {\n+        let (conn, client, _recv) = make_duplex_with_log(log);\n+        (conn, client)\n+    }\n+\n+    fn make_duplex_with_log(\n+        log: Arc<LifecycleLog>,\n+    ) -> (BrokerConn, Stream, interprocess::local_socket::RecvHalf) {\n         static SEQ: AtomicU32 = AtomicU32::new(0);\n\n[Showing lines 1-300 of 659. Use :301 to continue]"}],"details":{"resolvedPath":"C:\\Users\\decid\\.omp\\agent\\sessions\\-Documents-projects-spt-core\\2026-09-09T23-45-42-317Z_01a08890-21ad-7000-8f54-208d4f4f2f2c\\62.bash-original.log","contentType":"text/plain","totalLines":659,"displayContent":{"text":"diff --git a/CHANGELOG.md b/CHANGELOG.md\nindex eafe18d6..5b4a7156 100644\n--- a/CHANGELOG.md\n+++ b/CHANGELOG.md\n@@ -10,6 +10,17 @@ breaks something, or changes the observable behavior of existing surfaces\n broadly; **patch** for fixes, and for additive opt-in capability — a new key,\n flag, or page that no existing user can encounter without opting into it.\n \n+## [Unreleased]\n+\n+### Fixed\n+\n+<!-- [doc->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED] -->\n+- Daemon logs summarize healthy IPC connection starts and closes once per\n+  minute instead of logging each one. Failures and subscriber lifecycle\n+  events keep their per-connection attribution, including the original\n+  first-write time. Livehost session polling now reuses its healthy\n+  connection and names the reason when it opens a replacement.\n+\n ## [0.68.0] - 2026-09-08\n \n Web serving. Files, directories, adapter documentation and the changelog are\ndiff --git a/crates/spt-daemon/src/conn.rs b/crates/spt-daemon/src/conn.rs\nindex 9c51ad4e..163fecb6 100644\n--- a/crates/spt-daemon/src/conn.rs\n+++ b/crates/spt-daemon/src/conn.rs\n@@ -61,7 +61,7 @@\n \n use std::io;\n use std::sync::atomic::{AtomicBool, AtomicU64, Ordering};\n-use std::sync::{Arc, Condvar, Mutex, MutexGuard, OnceLock};\n+use std::sync::{Arc, Condvar, LazyLock, Mutex, MutexGuard, OnceLock};\n use std::thread::{self, JoinHandle};\n use std::time::{Duration, Instant, SystemTime, UNIX_EPOCH};\n \n@@ -99,12 +99,107 @@ pub(crate) fn init_log_anchor() {\n /// to field incident times (wall).\n // [impl->REQ-CONN-POISON-ATTRIBUTION]\n pub(crate) fn log_stamp() -> String {\n-    let wall_ms = SystemTime::now()\n-        .duration_since(UNIX_EPOCH)\n-        .map(|d| d.as_millis())\n-        .unwrap_or(0);\n-    let mono_ms = MONO_ANCHOR.get_or_init(Instant::now).elapsed().as_millis();\n-    format!(\"wall_ms={wall_ms} mono_ms={mono_ms}\")\n+    let stamp = LogTime::now();\n+    format!(\"wall_ms={} mono_ms={}\", stamp.wall_ms, stamp.mono_ms)\n+}\n+\n+/// Retain the original write instant even when its healthy line is aggregated.\n+#[derive(Clone, Copy)]\n+struct LogTime {\n+    wall_ms: u128,\n+    mono_ms: u128,\n+}\n+\n+impl LogTime {\n+    fn now() -> Self {\n+        Self {\n+            wall_ms: SystemTime::now()\n+                .duration_since(UNIX_EPOCH)\n+                .map(|d| d.as_millis())\n+                .unwrap_or(0),\n+            mono_ms: MONO_ANCHOR.get_or_init(Instant::now).elapsed().as_millis(),\n+        }\n+    }\n+}\n+\n+const HEALTHY_LOG_INTERVAL: Duration = Duration::from_secs(60);\n+\n+struct HealthyWindow {\n+    since: Instant,\n+    starts: u64,\n+    closes: u64,\n+}\n+\n+/// One process-wide counter window, not one limiter per connection. No thread,\n+/// queue or per-conn healthy string: the next event after 60s emits the totals.\n+/// Exceptional and role-transition records bypass this window entirely.\n+// [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+struct LifecycleLog {\n+    healthy: Mutex<HealthyWindow>,\n+    #[cfg(test)]\n+    capture: Option<Mutex<Vec<String>>>,\n+    #[cfg(test)]\n+    clock: Option<Mutex<Instant>>,\n+}\n+\n+impl LifecycleLog {\n+    fn new() -> Self {\n+        Self {\n+            healthy: Mutex::new(HealthyWindow {\n+                since: Instant::now(),\n+                starts: 0,\n+                closes: 0,\n+            }),\n+            #[cfg(test)]\n+            capture: None,\n+            #[cfg(test)]\n+            clock: None,\n+        }\n+    }\n+\n+    #[cfg(test)]\n+    fn captured(now: Instant) -> Self {\n+        let mut log = Self::new();\n+        log.healthy.get_mut().unwrap().since = now;\n+        log.capture = Some(Mutex::new(Vec::new()));\n+        log.clock = Some(Mutex::new(now));\n+        log\n+    }\n+\n+    fn emit(&self, args: std::fmt::Arguments<'_>) {\n+        #[cfg(test)]\n+        if let Some(lines) = &self.capture {\n+            recover(lines).push(args.to_string());\n+            return;\n+        }\n+        spt_proto::emit_line_err!(\"{args}\");\n+    }\n+\n+    fn healthy(&self, start: bool) {\n+        let mut window = recover(&self.healthy);\n+        // Read time under the counter lock so concurrent callers cannot move\n+        // the window backwards after waiting for a newer caller to flush.\n+        let now = Instant::now();\n+        #[cfg(test)]\n+        let now = self.clock.as_ref().map(|c| *recover(c)).unwrap_or(now);\n+        if start {\n+            window.starts = window.starts.saturating_add(1);\n+        } else {\n+            window.closes = window.closes.saturating_add(1);\n+        }\n+        let elapsed = now.saturating_duration_since(window.since);\n+        if elapsed < HEALTHY_LOG_INTERVAL {\n+            return;\n+        }\n+        let (starts, closes) = (window.starts, window.closes);\n+        *window = HealthyWindow { since: now, starts: 0, closes: 0 };\n+        drop(window);\n+        self.emit(format_args!(\n+            \"CONN_LIFECYCLE_SUMMARY: scope=broker-ipc {} interval_ms={} starts={starts} closes={closes} \\\n+             [REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\",\n+            log_stamp(), elapsed.as_millis()\n+        ));\n+    }\n }\n \n /// Attribution-label cap: facts accrete per role attach (a long-lived brain\n@@ -144,8 +239,9 @@ struct Inner {\n     /// Accreted attribution facts (subscriber role, endpoint/session where\n     /// known — [`BrokerConn::describe`]). Bounded by [`LABEL_CAP`].\n     label: Mutex<String>,\n-    /// Latch: the once-per-conn `write-start` lifecycle record fired.\n-    first_write_logged: AtomicBool,\n+    /// Once-per-conn original first committed write time; retained on failures.\n+    first_write: OnceLock<LogTime>,\n+    lifecycle_log: Arc<LifecycleLog>,\n     gate: Mutex<Gate>,\n     gate_cv: Condvar,\n     dog: Mutex<Dog>,\n@@ -223,7 +319,13 @@ impl Inner {\n     fn attribution(&self) -> String {\n         let label = recover(&self.label);\n         let facts: &str = if label.is_empty() { \"role=unattributed\" } else { &label };\n-        format!(\"conn={} {} {}\", self.id, facts, log_stamp())\n+        match self.first_write.get() {\n+            Some(start) => format!(\n+                \"conn={} {} {} first_write_wall_ms={} first_write_mono_ms={}\",\n+                self.id, facts, log_stamp(), start.wall_ms, start.mono_ms\n+            ),\n+            None => format!(\"conn={} {} {} first_write=none\", self.id, facts, log_stamp()),\n+        }\n     }\n \n     /// Render the once-per-conn retirement record — the F-039 token split\n@@ -261,18 +363,17 @@ impl Inner {\n         }\n     }\n \n-    /// One BOUNDED per-conn lifecycle record (leg d, doyle-confirmed\n-    /// UNCONDITIONAL): `write-start` (first write only), `transport-close`\n-    /// (drop), `writer-exit` / `*-attach` / `*-replaced` (emitted by the\n-    /// broker's sink machinery through [`BrokerConn::lifecycle_event`]). The\n-    /// timeout/cancel leg is the retirement record itself.\n+    /// Per-conn exceptional and role-transition records remain unconditional.\n+    /// releases#286 supersedes only healthy start/close emission, which uses\n+    /// the process-wide interval counter instead. Retirement retains the\n+    /// original first-write time even when the start line was aggregated.\n     // [impl->REQ-CONN-POISON-ATTRIBUTION]\n     fn lifecycle(&self, event: &str, extra: &str) {\n         let sep = if extra.is_empty() { \"\" } else { \" \" };\n-        spt_proto::emit_line_err!(\n+        self.lifecycle_log.emit(format_args!(\n             \"CONN_LIFECYCLE: {} event={event}{sep}{extra} [REQ-CONN-POISON-ATTRIBUTION]\",\n             self.attribution()\n-        );\n+        ));\n     }\n \n     /// Abort the connection's in-flight I/O OUT OF BAND (both directions —\n@@ -287,10 +388,10 @@ impl Inner {\n             // First poison of this conn: loud, once. The token is class-split\n             // (F-039 leg a): deadline → CONN_WRITE_POISONED (the wedge\n             // observable), organic fast-fail → CONN_WRITE_RETIRED.\n-            eprintln!(\n+            self.lifecycle_log.emit(format_args!(\n                 \"{}\",\n                 self.render_retirement(self.timed_out.load(Ordering::Acquire), cause)\n-            );\n+            ));\n         }\n         #[cfg(windows)]\n         // SAFETY: scalar kernel32 call on a handle whose owning half is kept\n@@ -424,11 +525,17 @@ impl BrokerConn {\n     /// deadline (gate-wait + OS write completion) — the broker passes its\n     /// `brain_write_deadline()` (`SPT_BRAIN_WRITE_DEADLINE_MS` knob).\n     pub(crate) fn new(half: SendHalf, bound: Duration) -> Self {\n+        static LOG: LazyLock<Arc<LifecycleLog>> = LazyLock::new(|| Arc::new(LifecycleLog::new()));\n+        Self::with_log(half, bound, Arc::clone(&LOG))\n+    }\n+\n+    fn with_log(half: SendHalf, bound: Duration, lifecycle_log: Arc<LifecycleLog>) -> Self {\n         let raw = raw_of(&half);\n         let inner = Arc::new(Inner {\n             id: CONN_ID_SEQ.fetch_add(1, Ordering::Relaxed) + 1,\n             label: Mutex::new(String::new()),\n-            first_write_logged: AtomicBool::new(false),\n+            first_write: OnceLock::new(),\n+            lifecycle_log,\n             gate: Mutex::new(Gate { half: Some(half) }),\n             gate_cv: Condvar::new(),\n             dog: Mutex::new(Dog {\n@@ -539,12 +646,12 @@ impl BrokerConn {\n             return Ok(ConnWrite::Superseded);\n         }\n \n-        // Once-per-conn `write-start` lifecycle record (leg d): marks the conn's\n-        // first committed write, so a conn that later retires can be correlated\n-        // to when it went active (fresh-carrier churn shows as open→start→retire\n-        // triplets with fresh ids).\n-        if !inner.first_write_logged.swap(true, Ordering::AcqRel) {\n-            inner.lifecycle(\"write-start\", \"\");\n+        // The gate serializes first-write initialization. Save the instant\n+        // BEFORE I/O so a later failure still names when this conn went active.\n+        // [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+        if inner.first_write.get().is_none() {\n+            let _ = inner.first_write.set(LogTime::now());\n+            inner.lifecycle_log.healthy(true);\n         }\n \n         // ── Check out the half + arm the watchdog, then write OUT of locks. ──\n@@ -638,17 +745,20 @@ impl BrokerConn {\n \n impl Drop for BrokerConn {\n     fn drop(&mut self) {\n-        // Once-per-conn `transport-close` lifecycle record (leg d): the handle\n-        // is about to close — the terminal record of this conn id's life, with\n-        // how it ended (organic EOF vs poisoned vs deadline-poisoned).\n-        self.inner.lifecycle(\n-            \"transport-close\",\n-            &format!(\n-                \"poisoned={} timed_out={}\",\n-                self.inner.poisoned.load(Ordering::Acquire),\n-                self.inner.timed_out.load(Ordering::Acquire)\n-            ),\n-        );\n+        // Only a healthy physical close is counted instead of logged per conn.\n+        // Error/timeout closes keep their full attribution and original start.\n+        // [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+        if self.inner.poisoned.load(Ordering::Acquire) {\n+            self.inner.lifecycle(\n+                \"transport-close\",\n+                &format!(\n+                    \"poisoned=true timed_out={}\",\n+                    self.inner.timed_out.load(Ordering::Acquire)\n+                ),\n+            );\n+        } else {\n+            self.inner.lifecycle_log.healthy(false);\n+        }\n         {\n             let mut d = recover(&self.inner.dog);\n             d.shutdown = true;\n@@ -688,6 +798,17 @@ mod tests {\n     /// primitive, the client end returned so a test can hold it alive or drop\n     /// it to force an ORGANIC write failure.\n     fn make_conn() -> (BrokerConn, Stream) {\n+        make_conn_with_log(Arc::new(LifecycleLog::new()))\n+    }\n+\n+    fn make_conn_with_log(log: Arc<LifecycleLog>) -> (BrokerConn, Stream) {\n+        let (conn, client, _recv) = make_duplex_with_log(log);\n+        (conn, client)\n+    }\n+\n+    fn make_duplex_with_log(\n+        log: Arc<LifecycleLog>,\n+    ) -> (BrokerConn, Stream, interprocess::local_socket::RecvHalf) {\n         static SEQ: AtomicU32 = AtomicU32::new(0);","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120,121,122,123,124,125,126,127,128,129,130,131,132,133,134,135,136,137,138,139,140,141,142,143,144,145,146,147,148,149,150,151,152,153,154,155,156,157,158,159,160,161,162,163,164,165,166,167,168,169,170,171,172,173,174,175,176,177,178,179,180,181,182,183,184,185,186,187,188,189,190,191,192,193,194,195,196,197,198,199,200,201,202,203,204,205,206,207,208,209,210,211,212,213,214,215,216,217,218,219,220,221,222,223,224,225,226,227,228,229,230,231,232,233,234,235,236,237,238,239,240,241,242,243,244,245,246,247,248,249,250,251,252,253,254,255,256,257,258,259,260,261,262,263,264,265,266,267,268,269,270,271,272,273,274,275,276,277,278,279,280,281,282,283,284,285,286,287,288,289,290,291,292,293,294,295,296,297,298,299,300]},"truncation":{"content":"diff --git a/CHANGELOG.md b/CHANGELOG.md\nindex eafe18d6..5b4a7156 100644\n--- a/CHANGELOG.md\n+++ b/CHANGELOG.md\n@@ -10,6 +10,17 @@ breaks something, or changes the observable behavior of existing surfaces\n broadly; **patch** for fixes, and for additive opt-in capability — a new key,\n flag, or page that no existing user can encounter without opting into it.\n \n+## [Unreleased]\n+\n+### Fixed\n+\n+<!-- [doc->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED] -->\n+- Daemon logs summarize healthy IPC connection starts and closes once per\n+  minute instead of logging each one. Failures and subscriber lifecycle\n+  events keep their per-connection attribution, including the original\n+  first-write time. Livehost session polling now reuses its healthy\n+  connection and names the reason when it opens a replacement.\n+\n ## [0.68.0] - 2026-09-08\n \n Web serving. Files, directories, adapter documentation and the changelog are\ndiff --git a/crates/spt-daemon/src/conn.rs b/crates/spt-daemon/src/conn.rs\nindex 9c51ad4e..163fecb6 100644\n--- a/crates/spt-daemon/src/conn.rs\n+++ b/crates/spt-daemon/src/conn.rs\n@@ -61,7 +61,7 @@\n \n use std::io;\n use std::sync::atomic::{AtomicBool, AtomicU64, Ordering};\n-use std::sync::{Arc, Condvar, Mutex, MutexGuard, OnceLock};\n+use std::sync::{Arc, Condvar, LazyLock, Mutex, MutexGuard, OnceLock};\n use std::thread::{self, JoinHandle};\n use std::time::{Duration, Instant, SystemTime, UNIX_EPOCH};\n \n@@ -99,12 +99,107 @@ pub(crate) fn init_log_anchor() {\n /// to field incident times (wall).\n // [impl->REQ-CONN-POISON-ATTRIBUTION]\n pub(crate) fn log_stamp() -> String {\n-    let wall_ms = SystemTime::now()\n-        .duration_since(UNIX_EPOCH)\n-        .map(|d| d.as_millis())\n-        .unwrap_or(0);\n-    let mono_ms = MONO_ANCHOR.get_or_init(Instant::now).elapsed().as_millis();\n-    format!(\"wall_ms={wall_ms} mono_ms={mono_ms}\")\n+    let stamp = LogTime::now();\n+    format!(\"wall_ms={} mono_ms={}\", stamp.wall_ms, stamp.mono_ms)\n+}\n+\n+/// Retain the original write instant even when its healthy line is aggregated.\n+#[derive(Clone, Copy)]\n+struct LogTime {\n+    wall_ms: u128,\n+    mono_ms: u128,\n+}\n+\n+impl LogTime {\n+    fn now() -> Self {\n+        Self {\n+            wall_ms: SystemTime::now()\n+                .duration_since(UNIX_EPOCH)\n+                .map(|d| d.as_millis())\n+                .unwrap_or(0),\n+            mono_ms: MONO_ANCHOR.get_or_init(Instant::now).elapsed().as_millis(),\n+        }\n+    }\n+}\n+\n+const HEALTHY_LOG_INTERVAL: Duration = Duration::from_secs(60);\n+\n+struct HealthyWindow {\n+    since: Instant,\n+    starts: u64,\n+    closes: u64,\n+}\n+\n+/// One process-wide counter window, not one limiter per connection. No thread,\n+/// queue or per-conn healthy string: the next event after 60s emits the totals.\n+/// Exceptional and role-transition records bypass this window entirely.\n+// [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+struct LifecycleLog {\n+    healthy: Mutex<HealthyWindow>,\n+    #[cfg(test)]\n+    capture: Option<Mutex<Vec<String>>>,\n+    #[cfg(test)]\n+    clock: Option<Mutex<Instant>>,\n+}\n+\n+impl LifecycleLog {\n+    fn new() -> Self {\n+        Self {\n+            healthy: Mutex::new(HealthyWindow {\n+                since: Instant::now(),\n+                starts: 0,\n+                closes: 0,\n+            }),\n+            #[cfg(test)]\n+            capture: None,\n+            #[cfg(test)]\n+            clock: None,\n+        }\n+    }\n+\n+    #[cfg(test)]\n+    fn captured(now: Instant) -> Self {\n+        let mut log = Self::new();\n+        log.healthy.get_mut().unwrap().since = now;\n+        log.capture = Some(Mutex::new(Vec::new()));\n+        log.clock = Some(Mutex::new(now));\n+        log\n+    }\n+\n+    fn emit(&self, args: std::fmt::Arguments<'_>) {\n+        #[cfg(test)]\n+        if let Some(lines) = &self.capture {\n+            recover(lines).push(args.to_string());\n+            return;\n+        }\n+        spt_proto::emit_line_err!(\"{args}\");\n+    }\n+\n+    fn healthy(&self, start: bool) {\n+        let mut window = recover(&self.healthy);\n+        // Read time under the counter lock so concurrent callers cannot move\n+        // the window backwards after waiting for a newer caller to flush.\n+        let now = Instant::now();\n+        #[cfg(test)]\n+        let now = self.clock.as_ref().map(|c| *recover(c)).unwrap_or(now);\n+        if start {\n+            window.starts = window.starts.saturating_add(1);\n+        } else {\n+            window.closes = window.closes.saturating_add(1);\n+        }\n+        let elapsed = now.saturating_duration_since(window.since);\n+        if elapsed < HEALTHY_LOG_INTERVAL {\n+            return;\n+        }\n+        let (starts, closes) = (window.starts, window.closes);\n+        *window = HealthyWindow { since: now, starts: 0, closes: 0 };\n+        drop(window);\n+        self.emit(format_args!(\n+            \"CONN_LIFECYCLE_SUMMARY: scope=broker-ipc {} interval_ms={} starts={starts} closes={closes} \\\n+             [REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\",\n+            log_stamp(), elapsed.as_millis()\n+        ));\n+    }\n }\n \n /// Attribution-label cap: facts accrete per role attach (a long-lived brain\n@@ -144,8 +239,9 @@ struct Inner {\n     /// Accreted attribution facts (subscriber role, endpoint/session where\n     /// known — [`BrokerConn::describe`]). Bounded by [`LABEL_CAP`].\n     label: Mutex<String>,\n-    /// Latch: the once-per-conn `write-start` lifecycle record fired.\n-    first_write_logged: AtomicBool,\n+    /// Once-per-conn original first committed write time; retained on failures.\n+    first_write: OnceLock<LogTime>,\n+    lifecycle_log: Arc<LifecycleLog>,\n     gate: Mutex<Gate>,\n     gate_cv: Condvar,\n     dog: Mutex<Dog>,\n@@ -223,7 +319,13 @@ impl Inner {\n     fn attribution(&self) -> String {\n         let label = recover(&self.label);\n         let facts: &str = if label.is_empty() { \"role=unattributed\" } else { &label };\n-        format!(\"conn={} {} {}\", self.id, facts, log_stamp())\n+        match self.first_write.get() {\n+            Some(start) => format!(\n+                \"conn={} {} {} first_write_wall_ms={} first_write_mono_ms={}\",\n+                self.id, facts, log_stamp(), start.wall_ms, start.mono_ms\n+            ),\n+            None => format!(\"conn={} {} {} first_write=none\", self.id, facts, log_stamp()),\n+        }\n     }\n \n     /// Render the once-per-conn retirement record — the F-039 token split\n@@ -261,18 +363,17 @@ impl Inner {\n         }\n     }\n \n-    /// One BOUNDED per-conn lifecycle record (leg d, doyle-confirmed\n-    /// UNCONDITIONAL): `write-start` (first write only), `transport-close`\n-    /// (drop), `writer-exit` / `*-attach` / `*-replaced` (emitted by the\n-    /// broker's sink machinery through [`BrokerConn::lifecycle_event`]). The\n-    /// timeout/cancel leg is the retirement record itself.\n+    /// Per-conn exceptional and role-transition records remain unconditional.\n+    /// releases#286 supersedes only healthy start/close emission, which uses\n+    /// the process-wide interval counter instead. Retirement retains the\n+    /// original first-write time even when the start line was aggregated.\n     // [impl->REQ-CONN-POISON-ATTRIBUTION]\n     fn lifecycle(&self, event: &str, extra: &str) {\n         let sep = if extra.is_empty() { \"\" } else { \" \" };\n-        spt_proto::emit_line_err!(\n+        self.lifecycle_log.emit(format_args!(\n             \"CONN_LIFECYCLE: {} event={event}{sep}{extra} [REQ-CONN-POISON-ATTRIBUTION]\",\n             self.attribution()\n-        );\n+        ));\n     }\n \n     /// Abort the connection's in-flight I/O OUT OF BAND (both directions —\n@@ -287,10 +388,10 @@ impl Inner {\n             // First poison of this conn: loud, once. The token is class-split\n             // (F-039 leg a): deadline → CONN_WRITE_POISONED (the wedge\n             // observable), organic fast-fail → CONN_WRITE_RETIRED.\n-            eprintln!(\n+            self.lifecycle_log.emit(format_args!(\n                 \"{}\",\n                 self.render_retirement(self.timed_out.load(Ordering::Acquire), cause)\n-            );\n+            ));\n         }\n         #[cfg(windows)]\n         // SAFETY: scalar kernel32 call on a handle whose owning half is kept\n@@ -424,11 +525,17 @@ impl BrokerConn {\n     /// deadline (gate-wait + OS write completion) — the broker passes its\n     /// `brain_write_deadline()` (`SPT_BRAIN_WRITE_DEADLINE_MS` knob).\n     pub(crate) fn new(half: SendHalf, bound: Duration) -> Self {\n+        static LOG: LazyLock<Arc<LifecycleLog>> = LazyLock::new(|| Arc::new(LifecycleLog::new()));\n+        Self::with_log(half, bound, Arc::clone(&LOG))\n+    }\n+\n+    fn with_log(half: SendHalf, bound: Duration, lifecycle_log: Arc<LifecycleLog>) -> Self {\n         let raw = raw_of(&half);\n         let inner = Arc::new(Inner {\n             id: CONN_ID_SEQ.fetch_add(1, Ordering::Relaxed) + 1,\n             label: Mutex::new(String::new()),\n-            first_write_logged: AtomicBool::new(false),\n+            first_write: OnceLock::new(),\n+            lifecycle_log,\n             gate: Mutex::new(Gate { half: Some(half) }),\n             gate_cv: Condvar::new(),\n             dog: Mutex::new(Dog {\n@@ -539,12 +646,12 @@ impl BrokerConn {\n             return Ok(ConnWrite::Superseded);\n         }\n \n-        // Once-per-conn `write-start` lifecycle record (leg d): marks the conn's\n-        // first committed write, so a conn that later retires can be correlated\n-        // to when it went active (fresh-carrier churn shows as open→start→retire\n-        // triplets with fresh ids).\n-        if !inner.first_write_logged.swap(true, Ordering::AcqRel) {\n-            inner.lifecycle(\"write-start\", \"\");\n+        // The gate serializes first-write initialization. Save the instant\n+        // BEFORE I/O so a later failure still names when this conn went active.\n+        // [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+        if inner.first_write.get().is_none() {\n+            let _ = inner.first_write.set(LogTime::now());\n+            inner.lifecycle_log.healthy(true);\n         }\n \n         // ── Check out the half + arm the watchdog, then write OUT of locks. ──\n@@ -638,17 +745,20 @@ impl BrokerConn {\n \n impl Drop for BrokerConn {\n     fn drop(&mut self) {\n-        // Once-per-conn `transport-close` lifecycle record (leg d): the handle\n-        // is about to close — the terminal record of this conn id's life, with\n-        // how it ended (organic EOF vs poisoned vs deadline-poisoned).\n-        self.inner.lifecycle(\n-            \"transport-close\",\n-            &format!(\n-                \"poisoned={} timed_out={}\",\n-                self.inner.poisoned.load(Ordering::Acquire),\n-                self.inner.timed_out.load(Ordering::Acquire)\n-            ),\n-        );\n+        // Only a healthy physical close is counted instead of logged per conn.\n+        // Error/timeout closes keep their full attribution and original start.\n+        // [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+        if self.inner.poisoned.load(Ordering::Acquire) {\n+            self.inner.lifecycle(\n+                \"transport-close\",\n+                &format!(\n+                    \"poisoned=true timed_out={}\",\n+                    self.inner.timed_out.load(Ordering::Acquire)\n+                ),\n+            );\n+        } else {\n+            self.inner.lifecycle_log.healthy(false);\n+        }\n         {\n             let mut d = recover(&self.inner.dog);\n             d.shutdown = true;\n@@ -688,6 +798,17 @@ mod tests {\n     /// primitive, the client end returned so a test can hold it alive or drop\n     /// it to force an ORGANIC write failure.\n     fn make_conn() -> (BrokerConn, Stream) {\n+        make_conn_with_log(Arc::new(LifecycleLog::new()))\n+    }\n+\n+    fn make_conn_with_log(log: Arc<LifecycleLog>) -> (BrokerConn, Stream) {\n+        let (conn, client, _recv) = make_duplex_with_log(log);\n+        (conn, client)\n+    }\n+\n+    fn make_duplex_with_log(\n+        log: Arc<LifecycleLog>,\n+    ) -> (BrokerConn, Stream, interprocess::local_socket::RecvHalf) {\n         static SEQ: AtomicU32 = AtomicU32::new(0);","truncated":true,"truncatedBy":"lines","totalLines":659,"totalBytes":12067,"outputLines":300,"outputBytes":12067,"lastLinePartial":false,"firstLineExceedsLimit":false},"meta":{"source":{"type":"internal","value":"artifact://62"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":659,"totalBytes":12067,"outputLines":300,"outputBytes":12067,"maxBytes":153600,"shownRange":{"start":1,"end":300},"nextOffset":301}}},"isError":false,"timestamp":1788999503255}}
{"type":"message","id":"f7423ca3","parentId":"c9717f8a","timestamp":"2026-09-10T00:18:23.257Z","message":{"role":"toolResult","toolCallId":"call_jXttUDLEeUbjVDJh0FaN6qQx|fc_04ba494197989c86016aa1f74f13f087d1b7850e25879cf204","toolName":"read","content":[{"type":"text","text":"# spt-core — agent working rules\n\nHarness-independent core for an agent ecosystem (messaging, live-agent lifecycle, terminal hosting, P2P networking, runtime-manifest harness contract). Rust crates + a single `spt` binary. Clean-room rebuild of `claude_skill_owl`.\n\n**Orientation — read before working:** `PRD.md` (requirements), `ROADMAP.md` (build path), `CONTEXT.md` (glossary/model, authoritative for meaning), `docs/adr/` (decisions), `docs/KNOWN-HAZARDS.md` (invariants we must not re-break), `docs/{STORAGE,MANIFEST,CONTEXT-MEMORY,DOCS-STRATEGY,TRACEABILITY}.md`.\n\nFor maintainer/debug update work, also read `docs/DEBUG-ROLLOUT.md` before touching the update-set or rollout path.\n\nFor work needing Linux tests and proof, this machine is authorized to ssh into a Linux box: `reavus@kitsubito`\n\nIMPORTANT: All agents you work alongside run on spt-core itself — the legacy-SPT (`$OWL`/`$LIVE`) era ended with the 2026-07 fleet migration; those commands are gone. Reach other agents with `spt send` (or the `/sptc:send` skill); keep message bodies short, and send long bodies from a file rather than inline. Everything else can be learned from the corresponding `/sptc:` skill doc.\n\n## Requirement traceability (binding)\n\nThis project uses `traceable-reqs` (`traceable-reqs.toml` = the authoritative `REQ-*` registry). The full contract is `docs/TRACEABILITY.md`. The rules you must follow:\n\n1. **Tag evidence in the same change.** When you write a function/test/doc-section that satisfies a requirement stage, add its tag in that same commit:\n   `// [impl->REQ-FOO]` · `// [unit->REQ-FOO]` · `<!-- [doc->REQ-FOO] -->`\n   Stages: `doc` / `impl` / `unit` / `int`. Tag *on or immediately above* the real evidence — never at file tops to satisfy coverage.\n2. **Run `traceable-reqs check` before declaring work done.** Exit-1 means missing/invalid evidence — fix it, don't ship it.\n3. **New requirement → add it to `traceable-reqs.toml` first** (with a `REQ-*` id), then satisfy it. No untracked work; no untagged evidence.\n4. **KNOWN-HAZARDS are `REQ-HAZARD-*` requirements** — each needs a test before it's \"covered.\" Treat the hazard list as a conformance checklist you must satisfy, not advice.\n5. **Activate, don't pre-fail.** Requirements you aren't yet working stay `required_stages = []`. Activate (set real stages) only when starting the milestone that delivers them.\n\n## Issue tracking (unique case)\n\nCodebase issues are maintained on **`BigscreenVR/spt-bs-releases`** — the less-private *releases* repo — NOT on `spt-bs-core` where the code lives. This is a deliberate exception (operator-ruled 2026-07-27): file, triage, and reference issues against `spt-bs-releases`.\n\nRequests and milestones are tracked through the **alchemy shell** (`alchemy-0`, the Requests shell bound to `spt-bs-releases`) on top of those GitHub issues. The label taxonomy is the workflow: `state:` BACKLOG → EVAL → GREENLIT → WIP → ACCEPTANCE → DONE (plus CUT / SHELVED), `type:` BUGFIX / ADDITION / CHANGE, `flag: NEEDS-OPERATOR`, and per-node `kind: SPT-TRACKING` pins (marked DO NOT DELETE — leave them alone).\n\n## Other conventions\n\n- Match surrounding code style; copy-verbatim the sister project's stable wire/schema formats (ADR-0001), clean-room everything else.\n- Honor every `docs/KNOWN-HAZARDS.md` invariant — these are real bugs already paid for once.\n- Docs are dual-audience (human + AI dev-agent) per `docs/DOCS-STRATEGY.md`; doc generation is CI-gated against drift.\n- Commit messages end with the project's co-author trailer, exact spelling `Co-authored by: <your live agent name>` (e.g. `Co-authored by: todlando`) — NOT git's hyphenated `Co-Authored-By:`; CI's runner-informant parses only the space spelling for attribution.\n  - **Audit this trailer with `git log --grep` or a raw body read, NEVER `%(trailers:)`** (nor `interpret-trailers --parse`): git's trailer tokenizer requires a whitespace-free token, so the space spelling is structurally invisible to it and every trailer query returns a confident EMPTY on every commit — a clean zero that reads as total attribution loss. Measured 2026-07-30 (main @`8f3e10b`, Signed-off-by and hyphenated-key controls both parse); re-derived at full cost 2026-08-03 by one of its own original measurers, which is why this sentence lives here at the mandate rather than in a document you must decide to open.\n- **Worktrees (binding, operator-ruled 2026-07-05):** every `git worktree` (gate rigs, agent scratch, anything) goes under `.worktrees/` in THIS project's root (`.gitignore`d) — e.g. `git worktree add --detach .worktrees/gate-<sha> <sha>`. NEVER create worktrees as siblings of the project dir; stray sibling folders accumulate orphan-pinned skeletons and pollute `~/Documents/projects`. Remove yours when done (`git worktree remove` + prune; retry later if a handle pins it).\n<!-- [doc->REQ-POOL-FOREIGN-TREE-REFUSED] -->\n- **Build-cache pools (binding, releases#103):** a cargo `target/` directory belongs to ONE source tree. Two trees building into one pool overwrite each other's artifacts and fabricate reds that belong to neither — the v0.51.0 W1/H1 `E0308` that vanished on re-run. **The rule is about the RELATIONSHIP, not the path:** a junction and a `CARGO_TARGET_DIR` env var reach the same hazard, and the env var leaves no directory entry to notice afterwards, so \"is this dir under `.worktrees/`\" never was the question. A build script enforces it (`crates/spt-store/build.rs`) and refuses with `SPT_POOL_FOREIGN`. **Sharing a pool SEQUENTIALLY is fine and expected** — a finished lane is taken over loudly, not refused; only a LIVE lane blocks you. Claim a pool at lane start …\n- **Target-tree teardown (releases#103, same lane):** never delete a `target/` you have not classified. Check OUTBOUND first (`Get-Item -Force`: a junction is deleted as a link, a real dir as a tree), then sweep INBOUND for reparse points aimed at it before removing a real one, and measure size and free space before and after so the reclaim is a number rather than a claim. Reap the target SUBTREE, never a session or worktree directory around it. Do not point `CARGO_TARGET_DIR` into a scratchpad (v0.50.0 incident) or at another checkout's target (H1 incident).\n- **CI / merge strategy:** golden CI (ADR-0050, ruled 2026-07-27) — milestone-batch golden runs, thin-lane PRs, ff-only main (tested sha == merged sha). Piloting from the next milestone; read the ADR before shaping a PR lane.\n\nIf you finish a significant body of work without need for user intervention, or if your context gets too high, you can clear your own context and keep moving:\n1. Create a JIT plan for the next immediate body of work, if it isn't already planned\n2. If you are a live agent, /sptc:commune with immediate next steps and broad summary of the project's status + end goal\n3. Execute the /self-clear skill\n\n[Some lines truncated to 768 chars]"}],"details":{"totalLines":47,"displayContent":{"text":"# spt-core — agent working rules\r\n\r\nHarness-independent core for an agent ecosystem (messaging, live-agent lifecycle, terminal hosting, P2P networking, runtime-manifest harness contract). Rust crates + a single `spt` binary. Clean-room rebuild of `claude_skill_owl`.\r\n\r\n**Orientation — read before working:** `PRD.md` (requirements), `ROADMAP.md` (build path), `CONTEXT.md` (glossary/model, authoritative for meaning), `docs/adr/` (decisions), `docs/KNOWN-HAZARDS.md` (invariants we must not re-break), `docs/{STORAGE,MANIFEST,CONTEXT-MEMORY,DOCS-STRATEGY,TRACEABILITY}.md`.\r\n\r\nFor maintainer/debug update work, also read `docs/DEBUG-ROLLOUT.md` before touching the update-set or rollout path.\r\n\r\nFor work needing Linux tests and proof, this machine is authorized to ssh into a Linux box: `reavus@kitsubito`\r\n\r\nIMPORTANT: All agents you work alongside run on spt-core itself — the legacy-SPT (`$OWL`/`$LIVE`) era ended with the 2026-07 fleet migration; those commands are gone. Reach other agents with `spt send` (or the `/sptc:send` skill); keep message bodies short, and send long bodies from a file rather than inline. Everything else can be learned from the corresponding `/sptc:` skill doc.\r\n\r\n## Requirement traceability (binding)\r\n\r\nThis project uses `traceable-reqs` (`traceable-reqs.toml` = the authoritative `REQ-*` registry). The full contract is `docs/TRACEABILITY.md`. The rules you must follow:\r\n\r\n1. **Tag evidence in the same change.** When you write a function/test/doc-section that satisfies a requirement stage, add its tag in that same commit:\r\n   `// [impl->REQ-FOO]` · `// [unit->REQ-FOO]` · `<!-- [doc->REQ-FOO] -->`\r\n   Stages: `doc` / `impl` / `unit` / `int`. Tag *on or immediately above* the real evidence — never at file tops to satisfy coverage.\r\n2. **Run `traceable-reqs check` before declaring work done.** Exit-1 means missing/invalid evidence — fix it, don't ship it.\r\n3. **New requirement → add it to `traceable-reqs.toml` first** (with a `REQ-*` id), then satisfy it. No untracked work; no untagged evidence.\r\n4. **KNOWN-HAZARDS are `REQ-HAZARD-*` requirements** — each needs a test before it's \"covered.\" Treat the hazard list as a conformance checklist you must satisfy, not advice.\r\n5. **Activate, don't pre-fail.** Requirements you aren't yet working stay `required_stages = []`. Activate (set real stages) only when starting the milestone that delivers them.\r\n\r\n## Issue tracking (unique case)\r\n\r\nCodebase issues are maintained on **`BigscreenVR/spt-bs-releases`** — the less-private *releases* repo — NOT on `spt-bs-core` where the code lives. This is a deliberate exception (operator-ruled 2026-07-27): file, triage, and reference issues against `spt-bs-releases`.\r\n\r\nRequests and milestones are tracked through the **alchemy shell** (`alchemy-0`, the Requests shell bound to `spt-bs-releases`) on top of those GitHub issues. The label taxonomy is the workflow: `state:` BACKLOG → EVAL → GREENLIT → WIP → ACCEPTANCE → DONE (plus CUT / SHELVED), `type:` BUGFIX / ADDITION / CHANGE, `flag: NEEDS-OPERATOR`, and per-node `kind: SPT-TRACKING` pins (marked DO NOT DELETE — leave them alone).\r\n\r\n## Other conventions\r\n\r\n- Match surrounding code style; copy-verbatim the sister project's stable wire/schema formats (ADR-0001), clean-room everything else.\r\n- Honor every `docs/KNOWN-HAZARDS.md` invariant — these are real bugs already paid for once.\r\n- Docs are dual-audience (human + AI dev-agent) per `docs/DOCS-STRATEGY.md`; doc generation is CI-gated against drift.\r\n- Commit messages end with the project's co-author trailer, exact spelling `Co-authored by: <your live agent name>` (e.g. `Co-authored by: todlando`) — NOT git's hyphenated `Co-Authored-By:`; CI's runner-informant parses only the space spelling for attribution.\r\n  - **Audit this trailer with `git log --grep` or a raw body read, NEVER `%(trailers:)`** (nor `interpret-trailers --parse`): git's trailer tokenizer requires a whitespace-free token, so the space spelling is structurally invisible to it and every trailer query returns a confident EMPTY on every commit — a clean zero that reads as total attribution loss. Measured 2026-07-30 (main @`8f3e10b`, Signed-off-by and hyphenated-key controls both parse); re-derived at full cost 2026-08-03 by one of its own original measurers, which is why this sentence lives here at the mandate rather than in a document you must decide to open.\r\n- **Worktrees (binding, operator-ruled 2026-07-05):** every `git worktree` (gate rigs, agent scratch, anything) goes under `.worktrees/` in THIS project's root (`.gitignore`d) — e.g. `git worktree add --detach .worktrees/gate-<sha> <sha>`. NEVER create worktrees as siblings of the project dir; stray sibling folders accumulate orphan-pinned skeletons and pollute `~/Documents/projects`. Remove yours when done (`git worktree remove` + prune; retry later if a handle pins it).\r\n<!-- [doc->REQ-POOL-FOREIGN-TREE-REFUSED] -->\r\n- **Build-cache pools (binding, releases#103):** a cargo `target/` directory belongs to ONE source tree. Two trees building into one pool overwrite each other's artifacts and fabricate reds that belong to neither — the v0.51.0 W1/H1 `E0308` that vanished on re-run. **The rule is about the RELATIONSHIP, not the path:** a junction and a `CARGO_TARGET_DIR` env var reach the same hazard, and the env var leaves no directory entry to notice afterwards, so \"is this dir under `.worktrees/`\" never was the question. A build script enforces it (`crates/spt-store/build.rs`) and refuses with `SPT_POOL_FOREIGN`. **Sharing a pool SEQUENTIALLY is fine and expected** — a finished lane is taken over loudly, not refused; only a LIVE lane blocks you. Claim a pool at lane start …\n- **Target-tree teardown (releases#103, same lane):** never delete a `target/` you have not classified. Check OUTBOUND first (`Get-Item -Force`: a junction is deleted as a link, a real dir as a tree), then sweep INBOUND for reparse points aimed at it before removing a real one, and measure size and free space before and after so the reclaim is a number rather than a claim. Reap the target SUBTREE, never a session or worktree directory around it. Do not point `CARGO_TARGET_DIR` into a scratchpad (v0.50.0 incident) or at another checkout's target (H1 incident).\r\n- **CI / merge strategy:** golden CI (ADR-0050, ruled 2026-07-27) — milestone-batch golden runs, thin-lane PRs, ff-only main (tested sha == merged sha). Piloting from the next milestone; read the ADR before shaping a PR lane.\r\n\r\nIf you finish a significant body of work without need for user intervention, or if your context gets too high, you can clear your own context and keep moving:\r\n1. Create a JIT plan for the next immediate body of work, if it isn't already planned\r\n2. If you are a live agent, /sptc:commune with immediate next steps and broad summary of the project's status + end goal\r\n3. Execute the /self-clear skill\r","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47]},"fileSize":8436,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\AGENTS.md"},"limits":{"columnTruncated":{"maxColumn":768}}}},"isError":false,"timestamp":1788999503256}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_K2wFCixSKT8qdBaNFENxm8GO|fc_04ba494197989c86016aa1f759aa7887d1b84325e52f8d59b4","toolName":"read","startedAt":"2026-09-10T00:18:33.770Z","args":{"path":"artifact://62:301-659"},"intent":"Reading remaining patch"},"id":"f1e32c96","parentId":"f7423ca3","timestamp":"2026-09-10T00:18:33.770Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_kmSQJbsJUh9pi2Nrmlkh27VR|fc_04ba494197989c86016aa1f759aa9487d1b2ed60d4784dcc55","toolName":"grep","startedAt":"2026-09-10T00:18:33.770Z","args":{"path":".worktrees/286-churn/PRD.md;.worktrees/286-churn/ROADMAP.md;.worktrees/286-churn/CONTEXT.md;.worktrees/286-churn/docs/KNOWN-HAZARDS.md;.worktrees/286-churn/docs/TRACEABILITY.md;.worktrees/286-churn/do…"},"intent":"Locating relevant orientation"},"id":"b017b0b7","parentId":"f1e32c96","timestamp":"2026-09-10T00:18:33.770Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_USkvjwI4jUCmiz4xDHy5Nb67|fc_04ba494197989c86016aa1f759aa9c87d1a5d6a53c1efd5b01","toolName":"read","startedAt":"2026-09-10T00:18:33.770Z","args":{"path":".worktrees/286-churn/crates/spt-daemon/src/conn.rs:230-770"},"intent":"Reading connection implementation"},"id":"73ec763e","parentId":"b017b0b7","timestamp":"2026-09-10T00:18:33.770Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_2V4wFZQ1QRjp0Ge0UwqRe9QB|fc_04ba494197989c86016aa1f759aaa887d18df792da4212f9be","toolName":"read","startedAt":"2026-09-10T00:18:33.770Z","args":{"path":".worktrees/286-churn/crates/spt-daemon/src/livehost.rs:997-1053,1339-1500"},"intent":"Reading livehost poll path"},"id":"2439b377","parentId":"73ec763e","timestamp":"2026-09-10T00:18:33.770Z"}
{"type":"message","id":"a73940d4","parentId":"2439b377","timestamp":"2026-09-10T00:18:33.770Z","message":{"role":"assistant","content":[{"type":"toolCall","id":"call_K2wFCixSKT8qdBaNFENxm8GO|fc_04ba494197989c86016aa1f759aa7887d1b84325e52f8d59b4","name":"read","arguments":{"i":"Reading remaining patch","path":"artifact://62:301-659"},"intent":"Reading remaining patch"},{"type":"toolCall","id":"call_kmSQJbsJUh9pi2Nrmlkh27VR|fc_04ba494197989c86016aa1f759aa9487d1b2ed60d4784dcc55","name":"grep","arguments":{"i":"Locating relevant orientation","pattern":"(?i)brain|lifecycle|poison|conn_write|blackhole|livehost|ipc|stderr","path":".worktrees/286-churn/PRD.md;.worktrees/286-churn/ROADMAP.md;.worktrees/286-churn/CONTEXT.md;.worktrees/286-churn/docs/KNOWN-HAZARDS.md;.worktrees/286-churn/docs/TRACEABILITY.md;.worktrees/286-churn/docs/DOCS-STRATEGY.md;.worktrees/286-churn/docs/STORAGE.md;.worktrees/286-churn/docs/MANIFEST.md;.worktrees/286-churn/docs/CONTEXT-MEMORY.md","case":false,"gitignore":false,"skip":null},"intent":"Locating relevant orientation"},{"type":"toolCall","id":"call_USkvjwI4jUCmiz4xDHy5Nb67|fc_04ba494197989c86016aa1f759aa9c87d1a5d6a53c1efd5b01","name":"read","arguments":{"i":"Reading connection implementation","path":".worktrees/286-churn/crates/spt-daemon/src/conn.rs:230-770"},"intent":"Reading connection implementation"},{"type":"toolCall","id":"call_2V4wFZQ1QRjp0Ge0UwqRe9QB|fc_04ba494197989c86016aa1f759aaa887d18df792da4212f9be","name":"read","arguments":{"i":"Reading livehost poll path","path":".worktrees/286-churn/crates/spt-daemon/src/livehost.rs:997-1053,1339-1500"},"intent":"Reading livehost poll path"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":13920,"output":305,"cacheRead":17664,"cacheWrite":0,"totalTokens":31889,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1788999503303,"responseId":"resp_04ba494197989c86016aa1f74f853887d19dab1553b08b827e","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"fc_04ba494197989c86016aa1f759aa7887d1b84325e52f8d59b4","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading remaining patch\",\"path\":\"artifact://62:301-659\"}","call_id":"call_K2wFCixSKT8qdBaNFENxm8GO","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"read","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f759aa9487d1b2ed60d4784dcc55","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating relevant orientation\",\"pattern\":\"(?i)brain|lifecycle|poison|conn_write|blackhole|livehost|ipc|stderr\",\"path\":\".worktrees/286-churn/PRD.md;.worktrees/286-churn/ROADMAP.md;.worktrees/286-churn/CONTEXT.md;.worktrees/286-churn/docs/KNOWN-HAZARDS.md;.worktrees/286-churn/docs/TRACEABILITY.md;.worktrees/286-churn/docs/DOCS-STRATEGY.md;.worktrees/286-churn/docs/STORAGE.md;.worktrees/286-churn/docs/MANIFEST.md;.worktrees/286-churn/docs/CONTEXT-MEMORY.md\",\"case\":false,\"gitignore\":false,\"skip\":null}","call_id":"call_kmSQJbsJUh9pi2Nrmlkh27VR","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"grep","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f759aa9c87d1a5d6a53c1efd5b01","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading connection implementation\",\"path\":\".worktrees/286-churn/crates/spt-daemon/src/conn.rs:230-770\"}","call_id":"call_USkvjwI4jUCmiz4xDHy5Nb67","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"read","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f759aaa887d18df792da4212f9be","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading livehost poll path\",\"path\":\".worktrees/286-churn/crates/spt-daemon/src/livehost.rs:997-1053,1339-1500\"}","call_id":"call_2V4wFZQ1QRjp0Ge0UwqRe9QB","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"read","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}}]},"duration":10464.87950000004,"ttft":10391.009800000116,"completedAt":1788999513769,"contextSnapshot":{"promptTokens":31584,"nonMessageTokens":13648,"compactionEpoch":0}}}
{"type":"message","id":"57b1a2a7","parentId":"a73940d4","timestamp":"2026-09-10T00:18:33.780Z","message":{"role":"toolResult","toolCallId":"call_USkvjwI4jUCmiz4xDHy5Nb67|fc_04ba494197989c86016aa1f759aa9c87d1a5d6a53c1efd5b01","toolName":"read","content":[{"type":"text","text":"struct Dog {\n…\n    inflight: Option<(u64, Instant)>,\n    /// Raised by [`BrokerConn::drop`]; the watchdog exits.\n    shutdown: bool,\n}\n\n/// The state shared with the watchdog thread (no `Arc<BrokerConn>` cycle —\n/// the watchdog must not keep the conn alive).\nstruct Inner {\n    /// Stable per-physical-conn id ([`CONN_ID_SEQ`]) — on every lifecycle record.\n    id: u64,\n    /// Accreted attribution facts (subscriber role, endpoint/session where\n    /// known — [`BrokerConn::describe`]). Bounded by [`LABEL_CAP`].\n    label: Mutex<String>,\n    /// Once-per-conn original first committed write time; retained on failures.\n    first_write: OnceLock<LogTime>,\n    lifecycle_log: Arc<LifecycleLog>,\n    gate: Mutex<Gate>,\n    gate_cv: Condvar,\n    dog: Mutex<Dog>,\n    dog_cv: Condvar,\n    /// Idempotently latched on the first deadline/failure/cancel. A poisoned\n    /// conn refuses every future write and is never re-gated.\n    poisoned: AtomicBool,\n    /// Latched when the retirement was DEADLINE-driven (watchdog fired, or a\n    /// gate-waiter exhausted its bound) — as opposed to an organic I/O error\n    /// (peer death). Callers see it as `io::ErrorKind::TimedOut` and can\n    /// count the wedge (the W2 stall-evict tally) without conflating it with\n    /// an ordinary disconnect.\n    timed_out: AtomicBool,\n    /// The physical sever ran (exactly once; Windows `DisconnectNamedPipe`).\n    severed: AtomicBool,\n    /// A retired (poisoned) half parks here so its OS handle stays alive\n    /// until [`BrokerConn::drop`] — after the watchdog is joined, so no raw\n    /// `CancelIoEx`/`DisconnectNamedPipe` can race the `CloseHandle`.\n    graveyard: Mutex<Option<SendHalf>>,\n    /// Monotonic per-write op identity (the watchdog cancels a SPECIFIC op).\n    op_seq: AtomicU64,\n    /// The absolute per-write bound (gate-wait + OS write completion).\n    bound: Duration,\n    /// The raw OS handle/fd of the underlying stream, captured at\n    /// construction for the out-of-band abort. The recv and send halves share\n    /// ONE underlying stream object (`interprocess` splits by ref-clone), so\n    /// aborting this handle covers a parked read too. Valid for the whole\n    /// conn lifetime: the half lives in `gate`, a writer's stack, or\n    /// `graveyard` until drop.\n    #[cfg(windows)]\n    raw: isize,\n    #[cfg(unix)]\n    raw: i32,\n}\n\n/// Recover a mutex guard even if a holder panicked — the state these guards\n/// protect is a plain slot/flag pair, valid on every panic boundary.\nfn recover<T>(m: &Mutex<T>) -> MutexGuard<'_, T> {\n    match m.lock() {\n        Ok(g) => g,\n        Err(p) => p.into_inner(),\n    }\n}\n\n#[cfg(windows)]\nfn raw_of(half: &SendHalf) -> isize {\n    use std::os::windows::io::{AsHandle, AsRawHandle};\n    match half {\n        SendHalf::NamedPipe(h) => h.as_handle().as_raw_handle() as isize,\n    }\n}\n\n#[cfg(unix)]\nfn raw_of(half: &SendHalf) -> i32 {\n    use std::os::fd::{AsFd, AsRawFd};\n    match half {\n        SendHalf::UdSocket(h) => h.as_fd().as_raw_fd(),\n    }\n}\n\n#[cfg(windows)]\nmod win {\n    // kernel32 is in the MSVC default link set — raw externs, no new dep\n    // (the deelevate.rs / daemon.rs posture).\n    extern \"system\" {\n        pub fn CancelIoEx(handle: isize, overlapped: *mut core::ffi::c_void) -> i32;\n        pub fn DisconnectNamedPipe(handle: isize) -> i32;\n    }\n}\n\nimpl Inner {\n    /// The attribution snapshot every lifecycle record carries: stable conn id,\n    /// accreted role/endpoint/session facts, wall + monotonic stamps.\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\n    fn attribution(&self) -> String {\n        let label = recover(&self.label);\n        let facts: &str = if label.is_empty() { \"role=unattributed\" } else { &label };\n        match self.first_write.get() {\n            Some(start) => format!(\n                \"conn={} {} {} first_write_wall_ms={} first_write_mono_ms={}\",\n                self.id, facts, log_stamp(), start.wall_ms, start.mono_ms\n            ),\n            None => format!(\"conn={} {} {} first_write=none\", self.id, facts, log_stamp()),\n        }\n    }\n\n    /// Render the once-per-conn retirement record — the F-039 token split\n    /// (leg a): `CONN_WRITE_POISONED` is RESERVED for the deadline class\n    /// (`timed_out` latched — the wedge observable blackhole-controller.ps1\n    /// watches); an ORGANIC write failure (peer death / broken pipe on an\n    /// already-dead counterpart, the offline-peer dial churn) emits the\n    /// distinct `CONN_WRITE_RETIRED` token carrying the error kind and the\n    /// SAME attribution fields — quieter class, never thinner (the retired\n    /// class is what discriminates fresh-carrier churn from interleave\n    /// artifact in the field). Behavior is identical either way: poison,\n    /// abort, retire, never reuse.\n    // [impl->REQ-CONN-POISON-DIAL-SCOPE]\n    fn render_retirement(&self, timed_out: bool, cause: Option<&io::Error>) -> String {\n        if timed_out {\n            format!(\n                \"CONN_WRITE_POISONED: {} — physical broker conn retired: a write \\\n                 exceeded its {}ms bound; in-flight I/O aborted, conn never reused \\\n                 [REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\",\n                self.attribution(),\n                self.bound.as_millis()\n            )\n        } else {\n            let cause = match cause {\n                Some(e) => format!(\"{:?}: {e}\", e.kind()),\n                None => \"write failed\".to_string(),\n            };\n            format!(\n                \"CONN_WRITE_RETIRED: {} — organic write failure ({cause}); physical \\\n                 broker conn poisoned + retired, never reused (same retire behavior \\\n                 as POISONED — only the failure class differs) \\\n                 [REQ-CONN-POISON-DIAL-SCOPE]\",\n                self.attribution()\n            )\n        }\n    }\n\n    /// Per-conn exceptional and role-transition records remain unconditional.\n    /// releases#286 supersedes only healthy start/close emission, which uses\n    /// the process-wide interval counter instead. Retirement retains the\n    /// original first-write time even when the start line was aggregated.\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\n    fn lifecycle(&self, event: &str, extra: &str) {\n        let sep = if extra.is_empty() { \"\" } else { \" \" };\n        self.lifecycle_log.emit(format_args!(\n            \"CONN_LIFECYCLE: {} event={event}{sep}{extra} [REQ-CONN-POISON-ATTRIBUTION]\",\n            self.attribution()\n        ));\n    }\n\n    /// Abort the connection's in-flight I/O OUT OF BAND (both directions —\n    /// the halves share one OS object) and latch the poison. Idempotent and\n    /// safe to repeat: `CancelIoEx` on a handle with nothing pending is a\n    /// no-op, and re-cancel closes the race where a write was issued a beat\n    /// AFTER a cancel (cancellation only reaches ALREADY-issued ops).\n    /// `cause` (the organic caller's write error) rides the retirement record\n    /// when the retirement is NOT deadline-driven.\n    fn poison_and_cancel(&self, cause: Option<&io::Error>) {\n        if !self.poisoned.swap(true, Ordering::AcqRel) {\n            // First poison of this conn: loud, once. The token is class-split\n            // (F-039 leg a): deadline → CONN_WRITE_POISONED (the wedge\n            // observable), organic fast-fail → CONN_WRITE_RETIRED.\n            self.lifecycle_log.emit(format_args!(\n                \"{}\",\n                self.render_retirement(self.timed_out.load(Ordering::Acquire), cause)\n            ));\n        }\n        #[cfg(windows)]\n        // SAFETY: scalar kernel32 call on a handle whose owning half is kept\n        // alive until Drop (gate / writer stack / graveyard — see `raw`).\n        unsafe {\n            let _ = win::CancelIoEx(self.raw, core::ptr::null_mut());\n        }\n        #[cfg(unix)]\n        // SAFETY: scalar shutdown(2) on an fd alive until Drop. SHUT_RDWR\n        // fails the blocked write (EPIPE), EOFs the parked read, and fails\n        // every future op — on Unix this one call IS the whole teardown.\n        unsafe {\n            let _ = libc::shutdown(self.raw, libc::SHUT_RDWR);\n        }\n        // Wake gate-waiters so they observe the poison instead of sleeping\n        // out their full deadline, and the watchdog so it re-evaluates.\n        self.gate_cv.notify_all();\n        self.dog_cv.notify_all();\n    }\n\n    /// Sever the physical connection (exactly once), AFTER the canceled op\n    /// reported completion: on Windows `DisconnectNamedPipe` forces the\n    /// client off (its next access fails), makes a re-issued server-side\n    /// read fail (so `handle_conn` reaches EOF cleanup even if it slipped a\n    /// new read in after the cancel), and defuses the drop-limbo flush; the\n    /// trailing re-cancel covers exactly that slipped-in read. On Unix the\n    /// `shutdown` in [`poison_and_cancel`] already did all of this.\n    // On Unix the cfg(windows) teardown strips out and the idempotence\n    // guard's early `return` becomes the function tail — needless_return\n    // there, load-bearing here. Restructuring (e.g. `if !swapped { … }`)\n    // would leave Unix an empty if-body (`clippy::needless_if`) instead.\n    #[cfg_attr(unix, allow(clippy::needless_return))]\n    fn sever(&self) {\n        if self.severed.swap(true, Ordering::AcqRel) {\n            return;\n        }\n        #[cfg(windows)]\n        // SAFETY: as in poison_and_cancel — handle outlives the watchdog.\n        unsafe {\n            let _ = win::DisconnectNamedPipe(self.raw);\n            let _ = win::CancelIoEx(self.raw, core::ptr::null_mut());\n        }\n    }\n\n    /// The watchdog body: park until a write is in flight, then sleep to its\n    /// absolute deadline. If the op is still in flight at the deadline,\n    /// poison + cancel, re-canceling periodically until the op reports\n    /// completion (the writer clears `inflight` the instant `write_frame`\n    /// returns — that return IS the buffer-release handshake), then sever.\n    fn watchdog(&self) {\n        let mut d = recover(&self.dog);\n        loop {\n            // Park until armed (or shutdown).\n            let (op, deadline) = loop {\n                if d.shutdown {\n                    return;\n                }\n                if let Some(armed) = d.inflight {\n                    break armed;\n                }\n                d = self\n                    .dog_cv\n                    .wait(d)\n                    .unwrap_or_else(|p| p.into_inner());\n            };\n            // Sleep toward the deadline while THIS op stays in flight.\n            let fired = loop {\n                if d.shutdown {\n                    return;\n                }\n                match d.inflight {\n                    Some((cur, _)) if cur == op => {}\n                    // Completed (or a newer op armed) before the deadline.\n                    _ => break false,\n                }\n                let now = Instant::now();\n                if now >= deadline {\n                    break true;\n                }\n                let (g, _) = self\n                    .dog_cv\n                    .wait_timeout(d, deadline - now)\n                    .unwrap_or_else(|p| p.into_inner());\n                d = g;\n            };\n            if !fired {\n                continue;\n            }\n            // Deadline: abort out of band, then wait for the canceled op to\n            // report completion before severing (re-cancel each slice — a\n            // cancel that raced the WriteFileEx issuance reaches nothing, so\n            // one shot is not enough). Latch timed_out FIRST so the aborted\n            // writer observes its failure as a DEADLINE retirement (wedge),\n            // not an organic conn death.\n            drop(d);\n            self.timed_out.store(true, Ordering::Release);\n            self.poison_and_cancel(None);\n            d = recover(&self.dog);\n            while matches!(d.inflight, Some((cur, _)) if cur == op) && !d.shutdown {\n                let (g, _) = self\n                    .dog_cv\n                    .wait_timeout(d, Duration::from_millis(100))\n                    .unwrap_or_else(|p| p.into_inner());\n                d = g;\n                drop(d);\n                self.poison_and_cancel(None);\n                d = recover(&self.dog);\n            }\n            if d.shutdown {\n                return;\n            }\n            drop(d);\n            self.sever();\n            d = recover(&self.dog);\n        }\n    }\n}\n\n/// The shared, serialized, BOUNDED send side of one physical broker\n/// connection — what [`crate::broker::SharedSend`] now aliases. Owns the\n/// [`SendHalf`] (it never leaves this module) and the per-conn watchdog.\npub(crate) struct BrokerConn {\n    inner: Arc<Inner>,\n    /// Joined on drop — after `shutdown` is raised — so the raw handle the\n    /// watchdog aborts through strictly outlives every raw use.\n    dog: Option<JoinHandle<()>>,\n}\n\nimpl BrokerConn {\n    /// Wrap a freshly-split send half. `bound` is the ABSOLUTE per-write\n    /// deadline (gate-wait + OS write completion) — the broker passes its\n    /// `brain_write_deadline()` (`SPT_BRAIN_WRITE_DEADLINE_MS` knob).\n    pub(crate) fn new(half: SendHalf, bound: Duration) -> Self {\n        static LOG: LazyLock<Arc<LifecycleLog>> = LazyLock::new(|| Arc::new(LifecycleLog::new()));\n        Self::with_log(half, bound, Arc::clone(&LOG))\n    }\n\n    fn with_log(half: SendHalf, bound: Duration, lifecycle_log: Arc<LifecycleLog>) -> Self {\n        let raw = raw_of(&half);\n        let inner = Arc::new(Inner {\n            id: CONN_ID_SEQ.fetch_add(1, Ordering::Relaxed) + 1,\n            label: Mutex::new(String::new()),\n            first_write: OnceLock::new(),\n            lifecycle_log,\n            gate: Mutex::new(Gate { half: Some(half) }),\n            gate_cv: Condvar::new(),\n            dog: Mutex::new(Dog {\n                inflight: None,\n                shutdown: false,\n            }),\n            dog_cv: Condvar::new(),\n            poisoned: AtomicBool::new(false),\n            timed_out: AtomicBool::new(false),\n            severed: AtomicBool::new(false),\n            graveyard: Mutex::new(None),\n            op_seq: AtomicU64::new(0),\n            bound,\n            raw,\n        });\n        let dog_inner = Arc::clone(&inner);\n        let dog = thread::Builder::new()\n            .name(\"conn-watchdog\".into())\n            .spawn(move || dog_inner.watchdog())\n            .ok();\n        BrokerConn { inner, dog }\n    }\n\n    /// The stable per-physical-conn id — what lifecycle records key on\n    /// (`Arc::ptr_eq` identity does not survive a log line).\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\n    pub(crate) fn id(&self) -> u64 {\n        self.inner.id\n    }\n\n    /// Accrete an attribution fact onto this conn's label (subscriber role,\n    /// endpoint/session where known — e.g. `role=brain`,\n    /// `controller session=7 by=cli@host`). Facts accumulate space-separated\n    /// and ride EVERY subsequent lifecycle record for this conn; growth is\n    /// bounded by [`LABEL_CAP`] (one `…` marker, then no-op).\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\n    pub(crate) fn describe(&self, fact: &str) {\n        let mut l = recover(&self.inner.label);\n        if l.ends_with('…') {\n            return;\n        }\n        if l.len() + fact.len() + 1 > LABEL_CAP {\n            l.push('…');\n            return;\n        }\n        if !l.is_empty() {\n            l.push(' ');\n        }\n        l.push_str(fact);\n    }\n\n    /// Emit a bounded per-conn lifecycle record through this conn's shared\n    /// attribution format — the broker's sink machinery uses it for\n    /// `writer-exit`, attach/replace, and detach events (leg d).\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\n    pub(crate) fn lifecycle_event(&self, event: &str, extra: &str) {\n        self.inner.lifecycle(event, extra);\n    }\n\n    /// Bounded framed write. See [`Self::write_gated`].\n    pub(crate) fn write(&self, env: &Envelope) -> io::Result<ConnWrite> {\n        self.write_gated(env, || true)\n    }\n\n    /// Bounded framed write with a pre-write gate closure, run AFTER the\n    /// serialized gate is acquired and BEFORE any byte is written — the same\n    /// serialization point the old epoch-check-under-`send.lock()` had, so a\n    /// superseded controller writer still bails atomically with respect to\n    /// its successor's writes (P1c, REQ-HAZARD-CONTROLLER-WRITER-REORDER).\n    /// `pre` returning false → `Ok(ConnWrite::Superseded)`, nothing written.\n    pub(crate) fn write_gated(\n        &self,\n        env: &Envelope,\n        pre: impl FnOnce() -> bool,\n    ) -> io::Result<ConnWrite> {\n        let inner = &*self.inner;\n        let deadline = Instant::now() + inner.bound;\n\n        // ── Acquire the gate, bounded by the SAME absolute deadline. ──\n        let mut g = recover(&inner.gate);\n        loop {\n            if inner.poisoned.load(Ordering::Acquire) {\n                return Err(poisoned_err());\n            }\n            if g.half.is_some() {\n                break;\n            }\n            let now = Instant::now();\n            if now >= deadline {\n                // The write we queued behind is stuck past OUR whole bound:\n                // that op is the wedge. Abort it out of band and fail — the\n                // watchdog's own deadline (armed at that op's entry) is also\n                // in flight; both paths converge on the same idempotent\n                // poison/cancel.\n                drop(g);\n                inner.timed_out.store(true, Ordering::Release);\n                inner.poison_and_cancel(None);\n                return Err(timeout_err());\n            }\n            let (g2, _) = inner\n                .gate_cv\n                .wait_timeout(g, deadline - now)\n                .unwrap_or_else(|p| p.into_inner());\n            g = g2;\n        }\n\n        if !pre() {\n            return Ok(ConnWrite::Superseded);\n        }\n\n        // The gate serializes first-write initialization. Save the instant\n        // BEFORE I/O so a later failure still names when this conn went active.\n        // [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n        if inner.first_write.get().is_none() {\n            let _ = inner.first_write.set(LogTime::now());\n            inner.lifecycle_log.healthy(true);\n        }\n\n        // ── Check out the half + arm the watchdog, then write OUT of locks. ──\n        let mut half = g.half.take().expect(\"gate half present at checkout\");\n        let op = inner.op_seq.fetch_add(1, Ordering::Relaxed) + 1;\n        {\n            let mut d = recover(&inner.dog);\n            d.inflight = Some((op, deadline));\n            inner.dog_cv.notify_all();\n        }\n        drop(g);\n\n        let res = write_frame(&mut half, env);\n\n        // Completion handshake: `write_frame` has returned, so the OS op —\n        // canceled or not — reported completion and the frame buffer is\n        // released. Clear the arm BEFORE deciding health, so the watchdog\n        // never severs while the op is formally in flight.\n        {\n            let mut d = recover(&inner.dog);\n            d.inflight = None;\n            inner.dog_cv.notify_all();\n        }\n\n        let mut g = recover(&inner.gate);\n        if res.is_ok() && !inner.poisoned.load(Ordering::Acquire) {\n            g.half = Some(half);\n            drop(g);\n            inner.gate_cv.notify_all();\n            return Ok(ConnWrite::Done);\n        }\n\n        // ── Retire: poison the whole conn, abort the read side so\n        // `handle_conn` reaches EOF cleanup, sever, and park the half in the\n        // graveyard (its handle must outlive any late raw abort; never\n        // re-gated, never reused — a timed-out length-prefixed frame may be\n        // partially written). ──\n        drop(g);\n        inner.poison_and_cancel(res.as_ref().err());\n        inner.sever();\n        *recover(&inner.graveyard) = Some(half);\n        inner.gate_cv.notify_all();\n        // A DEADLINE retirement (the watchdog aborted OUR op, or the conn was\n        // already deadline-poisoned) surfaces as TimedOut — the caller can\n        // count the wedge (W2 stall-evict tally). An organic I/O failure\n        // (peer death) keeps its own error.\n        if inner.timed_out.load(Ordering::Acquire) {\n            return Err(timeout_err());\n        }\n        Err(match res {\n            Err(e) => e,\n            Ok(()) => poisoned_err(),\n        })\n    }\n}\n\n/// TEST-ONLY determinism lever: holds the serialized write gate (the half is\n/// checked out) so every writer parks AT THE GATE until the pin drops — the\n/// round-4 analog of pinning the old `Mutex<SendHalf>` across a registration\n/// window (the P1c epoch-gate keystone unit relies on it).\n#[cfg(test)]\npub(crate) struct TestGatePin {\n    inner: Arc<Inner>,\n    half: Option<SendHalf>,\n}\n\n#[cfg(test)]\nimpl Drop for TestGatePin {\n    fn drop(&mut self) {\n        let mut g = recover(&self.inner.gate);\n        g.half = self.half.take();\n        drop(g);\n        self.inner.gate_cv.notify_all();\n    }\n}\n\n#[cfg(test)]\nimpl BrokerConn {\n    /// Check the half out of the gate without arming the watchdog — writers\n    /// queue behind the gate (bounded by their own deadlines) until the\n    /// returned pin drops. Test-only; never a production path.\n    pub(crate) fn pin_gate_for_test(&self) -> TestGatePin {\n        let mut g = recover(&self.inner.gate);\n        let half = g.half.take().expect(\"pin: gate half present\");\n        TestGatePin {\n            inner: Arc::clone(&self.inner),\n            half: Some(half),\n        }\n    }\n}\n\nimpl Drop for BrokerConn {\n    fn drop(&mut self) {\n        // Only a healthy physical close is counted instead of logged per conn.\n        // Error/timeout closes keep their full attribution and original start.\n        // [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n        if self.inner.poisoned.load(Ordering::Acquire) {\n            self.inner.lifecycle(\n                \"transport-close\",\n                &format!(\n                    \"poisoned=true timed_out={}\",\n                    self.inner.timed_out.load(Ordering::Acquire)\n                ),\n            );\n        } else {\n            self.inner.lifecycle_log.healthy(false);\n        }\n        {\n            let mut d = recover(&self.inner.dog);\n            d.shutdown = true;\n            self.inner.dog_cv.notify_all();\n        }\n        if let Some(dog) = self.dog.take() {\n            let _ = dog.join();\n        }\n        // Halves (gate or graveyard) drop AFTER the watchdog is gone — no raw\n        // handle use can race the CloseHandle/limbo path.\n    }\n}\n\n[Showing lines 229-773 of 1000. Use :774 to continue]"}],"details":{"truncation":{"content":"    inflight: Option<(u64, Instant)>,\r\n    /// Raised by [`BrokerConn::drop`]; the watchdog exits.\r\n    shutdown: bool,\r\n}\r\n\r\n/// The state shared with the watchdog thread (no `Arc<BrokerConn>` cycle —\r\n/// the watchdog must not keep the conn alive).\r\nstruct Inner {\r\n    /// Stable per-physical-conn id ([`CONN_ID_SEQ`]) — on every lifecycle record.\r\n    id: u64,\r\n    /// Accreted attribution facts (subscriber role, endpoint/session where\r\n    /// known — [`BrokerConn::describe`]). Bounded by [`LABEL_CAP`].\r\n    label: Mutex<String>,\r\n    /// Once-per-conn original first committed write time; retained on failures.\r\n    first_write: OnceLock<LogTime>,\r\n    lifecycle_log: Arc<LifecycleLog>,\r\n    gate: Mutex<Gate>,\r\n    gate_cv: Condvar,\r\n    dog: Mutex<Dog>,\r\n    dog_cv: Condvar,\r\n    /// Idempotently latched on the first deadline/failure/cancel. A poisoned\r\n    /// conn refuses every future write and is never re-gated.\r\n    poisoned: AtomicBool,\r\n    /// Latched when the retirement was DEADLINE-driven (watchdog fired, or a\r\n    /// gate-waiter exhausted its bound) — as opposed to an organic I/O error\r\n    /// (peer death). Callers see it as `io::ErrorKind::TimedOut` and can\r\n    /// count the wedge (the W2 stall-evict tally) without conflating it with\r\n    /// an ordinary disconnect.\r\n    timed_out: AtomicBool,\r\n    /// The physical sever ran (exactly once; Windows `DisconnectNamedPipe`).\r\n    severed: AtomicBool,\r\n    /// A retired (poisoned) half parks here so its OS handle stays alive\r\n    /// until [`BrokerConn::drop`] — after the watchdog is joined, so no raw\r\n    /// `CancelIoEx`/`DisconnectNamedPipe` can race the `CloseHandle`.\r\n    graveyard: Mutex<Option<SendHalf>>,\r\n    /// Monotonic per-write op identity (the watchdog cancels a SPECIFIC op).\r\n    op_seq: AtomicU64,\r\n    /// The absolute per-write bound (gate-wait + OS write completion).\r\n    bound: Duration,\r\n    /// The raw OS handle/fd of the underlying stream, captured at\r\n    /// construction for the out-of-band abort. The recv and send halves share\r\n    /// ONE underlying stream object (`interprocess` splits by ref-clone), so\r\n    /// aborting this handle covers a parked read too. Valid for the whole\r\n    /// conn lifetime: the half lives in `gate`, a writer's stack, or\r\n    /// `graveyard` until drop.\r\n    #[cfg(windows)]\r\n    raw: isize,\r\n    #[cfg(unix)]\r\n    raw: i32,\r\n}\r\n\r\n/// Recover a mutex guard even if a holder panicked — the state these guards\r\n/// protect is a plain slot/flag pair, valid on every panic boundary.\r\nfn recover<T>(m: &Mutex<T>) -> MutexGuard<'_, T> {\r\n    match m.lock() {\r\n        Ok(g) => g,\r\n        Err(p) => p.into_inner(),\r\n    }\r\n}\r\n\r\n#[cfg(windows)]\r\nfn raw_of(half: &SendHalf) -> isize {\r\n    use std::os::windows::io::{AsHandle, AsRawHandle};\r\n    match half {\r\n        SendHalf::NamedPipe(h) => h.as_handle().as_raw_handle() as isize,\r\n    }\r\n}\r\n\r\n#[cfg(unix)]\r\nfn raw_of(half: &SendHalf) -> i32 {\r\n    use std::os::fd::{AsFd, AsRawFd};\r\n    match half {\r\n        SendHalf::UdSocket(h) => h.as_fd().as_raw_fd(),\r\n    }\r\n}\r\n\r\n#[cfg(windows)]\r\nmod win {\r\n    // kernel32 is in the MSVC default link set — raw externs, no new dep\r\n    // (the deelevate.rs / daemon.rs posture).\r\n    extern \"system\" {\r\n        pub fn CancelIoEx(handle: isize, overlapped: *mut core::ffi::c_void) -> i32;\r\n        pub fn DisconnectNamedPipe(handle: isize) -> i32;\r\n    }\r\n}\r\n\r\nimpl Inner {\r\n    /// The attribution snapshot every lifecycle record carries: stable conn id,\r\n    /// accreted role/endpoint/session facts, wall + monotonic stamps.\r\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n    fn attribution(&self) -> String {\r\n        let label = recover(&self.label);\r\n        let facts: &str = if label.is_empty() { \"role=unattributed\" } else { &label };\r\n        match self.first_write.get() {\r\n            Some(start) => format!(\r\n                \"conn={} {} {} first_write_wall_ms={} first_write_mono_ms={}\",\r\n                self.id, facts, log_stamp(), start.wall_ms, start.mono_ms\r\n            ),\r\n            None => format!(\"conn={} {} {} first_write=none\", self.id, facts, log_stamp()),\r\n        }\r\n    }\r\n\r\n    /// Render the once-per-conn retirement record — the F-039 token split\r\n    /// (leg a): `CONN_WRITE_POISONED` is RESERVED for the deadline class\r\n    /// (`timed_out` latched — the wedge observable blackhole-controller.ps1\r\n    /// watches); an ORGANIC write failure (peer death / broken pipe on an\r\n    /// already-dead counterpart, the offline-peer dial churn) emits the\r\n    /// distinct `CONN_WRITE_RETIRED` token carrying the error kind and the\r\n    /// SAME attribution fields — quieter class, never thinner (the retired\r\n    /// class is what discriminates fresh-carrier churn from interleave\r\n    /// artifact in the field). Behavior is identical either way: poison,\r\n    /// abort, retire, never reuse.\r\n    // [impl->REQ-CONN-POISON-DIAL-SCOPE]\r\n    fn render_retirement(&self, timed_out: bool, cause: Option<&io::Error>) -> String {\r\n        if timed_out {\r\n            format!(\r\n                \"CONN_WRITE_POISONED: {} — physical broker conn retired: a write \\\r\n                 exceeded its {}ms bound; in-flight I/O aborted, conn never reused \\\r\n                 [REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\",\r\n                self.attribution(),\r\n                self.bound.as_millis()\r\n            )\r\n        } else {\r\n            let cause = match cause {\r\n                Some(e) => format!(\"{:?}: {e}\", e.kind()),\r\n                None => \"write failed\".to_string(),\r\n            };\r\n            format!(\r\n                \"CONN_WRITE_RETIRED: {} — organic write failure ({cause}); physical \\\r\n                 broker conn poisoned + retired, never reused (same retire behavior \\\r\n                 as POISONED — only the failure class differs) \\\r\n                 [REQ-CONN-POISON-DIAL-SCOPE]\",\r\n                self.attribution()\r\n            )\r\n        }\r\n    }\r\n\r\n    /// Per-conn exceptional and role-transition records remain unconditional.\r\n    /// releases#286 supersedes only healthy start/close emission, which uses\r\n    /// the process-wide interval counter instead. Retirement retains the\r\n    /// original first-write time even when the start line was aggregated.\r\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n    fn lifecycle(&self, event: &str, extra: &str) {\r\n        let sep = if extra.is_empty() { \"\" } else { \" \" };\r\n        self.lifecycle_log.emit(format_args!(\r\n            \"CONN_LIFECYCLE: {} event={event}{sep}{extra} [REQ-CONN-POISON-ATTRIBUTION]\",\r\n            self.attribution()\r\n        ));\r\n    }\r\n\r\n    /// Abort the connection's in-flight I/O OUT OF BAND (both directions —\r\n    /// the halves share one OS object) and latch the poison. Idempotent and\r\n    /// safe to repeat: `CancelIoEx` on a handle with nothing pending is a\r\n    /// no-op, and re-cancel closes the race where a write was issued a beat\r\n    /// AFTER a cancel (cancellation only reaches ALREADY-issued ops).\r\n    /// `cause` (the organic caller's write error) rides the retirement record\r\n    /// when the retirement is NOT deadline-driven.\r\n    fn poison_and_cancel(&self, cause: Option<&io::Error>) {\r\n        if !self.poisoned.swap(true, Ordering::AcqRel) {\r\n            // First poison of this conn: loud, once. The token is class-split\r\n            // (F-039 leg a): deadline → CONN_WRITE_POISONED (the wedge\r\n            // observable), organic fast-fail → CONN_WRITE_RETIRED.\r\n            self.lifecycle_log.emit(format_args!(\r\n                \"{}\",\r\n                self.render_retirement(self.timed_out.load(Ordering::Acquire), cause)\r\n            ));\r\n        }\r\n        #[cfg(windows)]\r\n        // SAFETY: scalar kernel32 call on a handle whose owning half is kept\r\n        // alive until Drop (gate / writer stack / graveyard — see `raw`).\r\n        unsafe {\r\n            let _ = win::CancelIoEx(self.raw, core::ptr::null_mut());\r\n        }\r\n        #[cfg(unix)]\r\n        // SAFETY: scalar shutdown(2) on an fd alive until Drop. SHUT_RDWR\r\n        // fails the blocked write (EPIPE), EOFs the parked read, and fails\r\n        // every future op — on Unix this one call IS the whole teardown.\r\n        unsafe {\r\n            let _ = libc::shutdown(self.raw, libc::SHUT_RDWR);\r\n        }\r\n        // Wake gate-waiters so they observe the poison instead of sleeping\r\n        // out their full deadline, and the watchdog so it re-evaluates.\r\n        self.gate_cv.notify_all();\r\n        self.dog_cv.notify_all();\r\n    }\r\n\r\n    /// Sever the physical connection (exactly once), AFTER the canceled op\r\n    /// reported completion: on Windows `DisconnectNamedPipe` forces the\r\n    /// client off (its next access fails), makes a re-issued server-side\r\n    /// read fail (so `handle_conn` reaches EOF cleanup even if it slipped a\r\n    /// new read in after the cancel), and defuses the drop-limbo flush; the\r\n    /// trailing re-cancel covers exactly that slipped-in read. On Unix the\r\n    /// `shutdown` in [`poison_and_cancel`] already did all of this.\r\n    // On Unix the cfg(windows) teardown strips out and the idempotence\r\n    // guard's early `return` becomes the function tail — needless_return\r\n    // there, load-bearing here. Restructuring (e.g. `if !swapped { … }`)\r\n    // would leave Unix an empty if-body (`clippy::needless_if`) instead.\r\n    #[cfg_attr(unix, allow(clippy::needless_return))]\r\n    fn sever(&self) {\r\n        if self.severed.swap(true, Ordering::AcqRel) {\r\n            return;\r\n        }\r\n        #[cfg(windows)]\r\n        // SAFETY: as in poison_and_cancel — handle outlives the watchdog.\r\n        unsafe {\r\n            let _ = win::DisconnectNamedPipe(self.raw);\r\n            let _ = win::CancelIoEx(self.raw, core::ptr::null_mut());\r\n        }\r\n    }\r\n\r\n    /// The watchdog body: park until a write is in flight, then sleep to its\r\n    /// absolute deadline. If the op is still in flight at the deadline,\r\n    /// poison + cancel, re-canceling periodically until the op reports\r\n    /// completion (the writer clears `inflight` the instant `write_frame`\r\n    /// returns — that return IS the buffer-release handshake), then sever.\r\n    fn watchdog(&self) {\r\n        let mut d = recover(&self.dog);\r\n        loop {\r\n            // Park until armed (or shutdown).\r\n            let (op, deadline) = loop {\r\n                if d.shutdown {\r\n                    return;\r\n                }\r\n                if let Some(armed) = d.inflight {\r\n                    break armed;\r\n                }\r\n                d = self\r\n                    .dog_cv\r\n                    .wait(d)\r\n                    .unwrap_or_else(|p| p.into_inner());\r\n            };\r\n            // Sleep toward the deadline while THIS op stays in flight.\r\n            let fired = loop {\r\n                if d.shutdown {\r\n                    return;\r\n                }\r\n                match d.inflight {\r\n                    Some((cur, _)) if cur == op => {}\r\n                    // Completed (or a newer op armed) before the deadline.\r\n                    _ => break false,\r\n                }\r\n                let now = Instant::now();\r\n                if now >= deadline {\r\n                    break true;\r\n                }\r\n                let (g, _) = self\r\n                    .dog_cv\r\n                    .wait_timeout(d, deadline - now)\r\n                    .unwrap_or_else(|p| p.into_inner());\r\n                d = g;\r\n            };\r\n            if !fired {\r\n                continue;\r\n            }\r\n            // Deadline: abort out of band, then wait for the canceled op to\r\n            // report completion before severing (re-cancel each slice — a\r\n            // cancel that raced the WriteFileEx issuance reaches nothing, so\r\n            // one shot is not enough). Latch timed_out FIRST so the aborted\r\n            // writer observes its failure as a DEADLINE retirement (wedge),\r\n            // not an organic conn death.\r\n            drop(d);\r\n            self.timed_out.store(true, Ordering::Release);\r\n            self.poison_and_cancel(None);\r\n            d = recover(&self.dog);\r\n            while matches!(d.inflight, Some((cur, _)) if cur == op) && !d.shutdown {\r\n                let (g, _) = self\r\n                    .dog_cv\r\n                    .wait_timeout(d, Duration::from_millis(100))\r\n                    .unwrap_or_else(|p| p.into_inner());\r\n                d = g;\r\n                drop(d);\r\n                self.poison_and_cancel(None);\r\n                d = recover(&self.dog);\r\n            }\r\n            if d.shutdown {\r\n                return;\r\n            }\r\n            drop(d);\r\n            self.sever();\r\n            d = recover(&self.dog);\r\n        }\r\n    }\r\n}\r\n\r\n/// The shared, serialized, BOUNDED send side of one physical broker\r\n/// connection — what [`crate::broker::SharedSend`] now aliases. Owns the\r\n/// [`SendHalf`] (it never leaves this module) and the per-conn watchdog.\r\npub(crate) struct BrokerConn {\r\n    inner: Arc<Inner>,\r\n    /// Joined on drop — after `shutdown` is raised — so the raw handle the\r\n    /// watchdog aborts through strictly outlives every raw use.\r\n    dog: Option<JoinHandle<()>>,\r\n}\r\n\r\nimpl BrokerConn {\r\n    /// Wrap a freshly-split send half. `bound` is the ABSOLUTE per-write\r\n    /// deadline (gate-wait + OS write completion) — the broker passes its\r\n    /// `brain_write_deadline()` (`SPT_BRAIN_WRITE_DEADLINE_MS` knob).\r\n    pub(crate) fn new(half: SendHalf, bound: Duration) -> Self {\r\n        static LOG: LazyLock<Arc<LifecycleLog>> = LazyLock::new(|| Arc::new(LifecycleLog::new()));\r\n        Self::with_log(half, bound, Arc::clone(&LOG))\r\n    }\r\n\r\n    fn with_log(half: SendHalf, bound: Duration, lifecycle_log: Arc<LifecycleLog>) -> Self {\r\n        let raw = raw_of(&half);\r\n        let inner = Arc::new(Inner {\r\n            id: CONN_ID_SEQ.fetch_add(1, Ordering::Relaxed) + 1,\r\n            label: Mutex::new(String::new()),\r\n            first_write: OnceLock::new(),\r\n            lifecycle_log,\r\n            gate: Mutex::new(Gate { half: Some(half) }),\r\n            gate_cv: Condvar::new(),\r\n            dog: Mutex::new(Dog {\r\n                inflight: None,\r\n                shutdown: false,\r\n            }),\r\n            dog_cv: Condvar::new(),\r\n            poisoned: AtomicBool::new(false),\r\n            timed_out: AtomicBool::new(false),\r\n            severed: AtomicBool::new(false),\r\n            graveyard: Mutex::new(None),\r\n            op_seq: AtomicU64::new(0),\r\n            bound,\r\n            raw,\r\n        });\r\n        let dog_inner = Arc::clone(&inner);\r\n        let dog = thread::Builder::new()\r\n            .name(\"conn-watchdog\".into())\r\n            .spawn(move || dog_inner.watchdog())\r\n            .ok();\r\n        BrokerConn { inner, dog }\r\n    }\r\n\r\n    /// The stable per-physical-conn id — what lifecycle records key on\r\n    /// (`Arc::ptr_eq` identity does not survive a log line).\r\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n    pub(crate) fn id(&self) -> u64 {\r\n        self.inner.id\r\n    }\r\n\r\n    /// Accrete an attribution fact onto this conn's label (subscriber role,\r\n    /// endpoint/session where known — e.g. `role=brain`,\r\n    /// `controller session=7 by=cli@host`). Facts accumulate space-separated\r\n    /// and ride EVERY subsequent lifecycle record for this conn; growth is\r\n    /// bounded by [`LABEL_CAP`] (one `…` marker, then no-op).\r\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n    pub(crate) fn describe(&self, fact: &str) {\r\n        let mut l = recover(&self.inner.label);\r\n        if l.ends_with('…') {\r\n            return;\r\n        }\r\n        if l.len() + fact.len() + 1 > LABEL_CAP {\r\n            l.push('…');\r\n            return;\r\n        }\r\n        if !l.is_empty() {\r\n            l.push(' ');\r\n        }\r\n        l.push_str(fact);\r\n    }\r\n\r\n    /// Emit a bounded per-conn lifecycle record through this conn's shared\r\n    /// attribution format — the broker's sink machinery uses it for\r\n    /// `writer-exit`, attach/replace, and detach events (leg d).\r\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n    pub(crate) fn lifecycle_event(&self, event: &str, extra: &str) {\r\n        self.inner.lifecycle(event, extra);\r\n    }\r\n\r\n    /// Bounded framed write. See [`Self::write_gated`].\r\n    pub(crate) fn write(&self, env: &Envelope) -> io::Result<ConnWrite> {\r\n        self.write_gated(env, || true)\r\n    }\r\n\r\n    /// Bounded framed write with a pre-write gate closure, run AFTER the\r\n    /// serialized gate is acquired and BEFORE any byte is written — the same\r\n    /// serialization point the old epoch-check-under-`send.lock()` had, so a\r\n    /// superseded controller writer still bails atomically with respect to\r\n    /// its successor's writes (P1c, REQ-HAZARD-CONTROLLER-WRITER-REORDER).\r\n    /// `pre` returning false → `Ok(ConnWrite::Superseded)`, nothing written.\r\n    pub(crate) fn write_gated(\r\n        &self,\r\n        env: &Envelope,\r\n        pre: impl FnOnce() -> bool,\r\n    ) -> io::Result<ConnWrite> {\r\n        let inner = &*self.inner;\r\n        let deadline = Instant::now() + inner.bound;\r\n\r\n        // ── Acquire the gate, bounded by the SAME absolute deadline. ──\r\n        let mut g = recover(&inner.gate);\r\n        loop {\r\n            if inner.poisoned.load(Ordering::Acquire) {\r\n                return Err(poisoned_err());\r\n            }\r\n            if g.half.is_some() {\r\n                break;\r\n            }\r\n            let now = Instant::now();\r\n            if now >= deadline {\r\n                // The write we queued behind is stuck past OUR whole bound:\r\n                // that op is the wedge. Abort it out of band and fail — the\r\n                // watchdog's own deadline (armed at that op's entry) is also\r\n                // in flight; both paths converge on the same idempotent\r\n                // poison/cancel.\r\n                drop(g);\r\n                inner.timed_out.store(true, Ordering::Release);\r\n                inner.poison_and_cancel(None);\r\n                return Err(timeout_err());\r\n            }\r\n            let (g2, _) = inner\r\n                .gate_cv\r\n                .wait_timeout(g, deadline - now)\r\n                .unwrap_or_else(|p| p.into_inner());\r\n            g = g2;\r\n        }\r\n\r\n        if !pre() {\r\n            return Ok(ConnWrite::Superseded);\r\n        }\r\n\r\n        // The gate serializes first-write initialization. Save the instant\r\n        // BEFORE I/O so a later failure still names when this conn went active.\r\n        // [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\r\n        if inner.first_write.get().is_none() {\r\n            let _ = inner.first_write.set(LogTime::now());\r\n            inner.lifecycle_log.healthy(true);\r\n        }\r\n\r\n        // ── Check out the half + arm the watchdog, then write OUT of locks. ──\r\n        let mut half = g.half.take().expect(\"gate half present at checkout\");\r\n        let op = inner.op_seq.fetch_add(1, Ordering::Relaxed) + 1;\r\n        {\r\n            let mut d = recover(&inner.dog);\r\n            d.inflight = Some((op, deadline));\r\n            inner.dog_cv.notify_all();\r\n        }\r\n        drop(g);\r\n\r\n        let res = write_frame(&mut half, env);\r\n\r\n        // Completion handshake: `write_frame` has returned, so the OS op —\r\n        // canceled or not — reported completion and the frame buffer is\r\n        // released. Clear the arm BEFORE deciding health, so the watchdog\r\n        // never severs while the op is formally in flight.\r\n        {\r\n            let mut d = recover(&inner.dog);\r\n            d.inflight = None;\r\n            inner.dog_cv.notify_all();\r\n        }\r\n\r\n        let mut g = recover(&inner.gate);\r\n        if res.is_ok() && !inner.poisoned.load(Ordering::Acquire) {\r\n            g.half = Some(half);\r\n            drop(g);\r\n            inner.gate_cv.notify_all();\r\n            return Ok(ConnWrite::Done);\r\n        }\r\n\r\n        // ── Retire: poison the whole conn, abort the read side so\r\n        // `handle_conn` reaches EOF cleanup, sever, and park the half in the\r\n        // graveyard (its handle must outlive any late raw abort; never\r\n        // re-gated, never reused — a timed-out length-prefixed frame may be\r\n        // partially written). ──\r\n        drop(g);\r\n        inner.poison_and_cancel(res.as_ref().err());\r\n        inner.sever();\r\n        *recover(&inner.graveyard) = Some(half);\r\n        inner.gate_cv.notify_all();\r\n        // A DEADLINE retirement (the watchdog aborted OUR op, or the conn was\r\n        // already deadline-poisoned) surfaces as TimedOut — the caller can\r\n        // count the wedge (W2 stall-evict tally). An organic I/O failure\r\n        // (peer death) keeps its own error.\r\n        if inner.timed_out.load(Ordering::Acquire) {\r\n            return Err(timeout_err());\r\n        }\r\n        Err(match res {\r\n            Err(e) => e,\r\n            Ok(()) => poisoned_err(),\r\n        })\r\n    }\r\n}\r\n\r\n/// TEST-ONLY determinism lever: holds the serialized write gate (the half is\r\n/// checked out) so every writer parks AT THE GATE until the pin drops — the\r\n/// round-4 analog of pinning the old `Mutex<SendHalf>` across a registration\r\n/// window (the P1c epoch-gate keystone unit relies on it).\r\n#[cfg(test)]\r\npub(crate) struct TestGatePin {\r\n    inner: Arc<Inner>,\r\n    half: Option<SendHalf>,\r\n}\r\n\r\n#[cfg(test)]\r\nimpl Drop for TestGatePin {\r\n    fn drop(&mut self) {\r\n        let mut g = recover(&self.inner.gate);\r\n        g.half = self.half.take();\r\n        drop(g);\r\n        self.inner.gate_cv.notify_all();\r\n    }\r\n}\r\n\r\n#[cfg(test)]\r\nimpl BrokerConn {\r\n    /// Check the half out of the gate without arming the watchdog — writers\r\n    /// queue behind the gate (bounded by their own deadlines) until the\r\n    /// returned pin drops. Test-only; never a production path.\r\n    pub(crate) fn pin_gate_for_test(&self) -> TestGatePin {\r\n        let mut g = recover(&self.inner.gate);\r\n        let half = g.half.take().expect(\"pin: gate half present\");\r\n        TestGatePin {\r\n            inner: Arc::clone(&self.inner),\r\n            half: Some(half),\r\n        }\r\n    }\r\n}\r\n\r\nimpl Drop for BrokerConn {\r\n    fn drop(&mut self) {\r\n        // Only a healthy physical close is counted instead of logged per conn.\r\n        // Error/timeout closes keep their full attribution and original start.\r\n        // [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\r\n        if self.inner.poisoned.load(Ordering::Acquire) {\r\n            self.inner.lifecycle(\r\n                \"transport-close\",\r\n                &format!(\r\n                    \"poisoned=true timed_out={}\",\r\n                    self.inner.timed_out.load(Ordering::Acquire)\r\n                ),\r\n            );\r\n        } else {\r\n            self.inner.lifecycle_log.healthy(false);\r\n        }\r\n        {\r\n            let mut d = recover(&self.inner.dog);\r\n            d.shutdown = true;\r\n            self.inner.dog_cv.notify_all();\r\n        }\r\n        if let Some(dog) = self.dog.take() {\r\n            let _ = dog.join();\r\n        }\r\n        // Halves (gate or graveyard) drop AFTER the watchdog is gone — no raw\r\n        // handle use can race the CloseHandle/limbo path.\r\n    }\r\n}\r","truncated":true,"truncatedBy":"lines","totalLines":772,"totalBytes":23120,"outputLines":545,"outputBytes":23120,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":1000,"displayContent":{"text":"struct Dog {\n…\n    inflight: Option<(u64, Instant)>,\r\n    /// Raised by [`BrokerConn::drop`]; the watchdog exits.\r\n    shutdown: bool,\r\n}\r\n\r\n/// The state shared with the watchdog thread (no `Arc<BrokerConn>` cycle —\r\n/// the watchdog must not keep the conn alive).\r\nstruct Inner {\r\n    /// Stable per-physical-conn id ([`CONN_ID_SEQ`]) — on every lifecycle record.\r\n    id: u64,\r\n    /// Accreted attribution facts (subscriber role, endpoint/session where\r\n    /// known — [`BrokerConn::describe`]). Bounded by [`LABEL_CAP`].\r\n    label: Mutex<String>,\r\n    /// Once-per-conn original first committed write time; retained on failures.\r\n    first_write: OnceLock<LogTime>,\r\n    lifecycle_log: Arc<LifecycleLog>,\r\n    gate: Mutex<Gate>,\r\n    gate_cv: Condvar,\r\n    dog: Mutex<Dog>,\r\n    dog_cv: Condvar,\r\n    /// Idempotently latched on the first deadline/failure/cancel. A poisoned\r\n    /// conn refuses every future write and is never re-gated.\r\n    poisoned: AtomicBool,\r\n    /// Latched when the retirement was DEADLINE-driven (watchdog fired, or a\r\n    /// gate-waiter exhausted its bound) — as opposed to an organic I/O error\r\n    /// (peer death). Callers see it as `io::ErrorKind::TimedOut` and can\r\n    /// count the wedge (the W2 stall-evict tally) without conflating it with\r\n    /// an ordinary disconnect.\r\n    timed_out: AtomicBool,\r\n    /// The physical sever ran (exactly once; Windows `DisconnectNamedPipe`).\r\n    severed: AtomicBool,\r\n    /// A retired (poisoned) half parks here so its OS handle stays alive\r\n    /// until [`BrokerConn::drop`] — after the watchdog is joined, so no raw\r\n    /// `CancelIoEx`/`DisconnectNamedPipe` can race the `CloseHandle`.\r\n    graveyard: Mutex<Option<SendHalf>>,\r\n    /// Monotonic per-write op identity (the watchdog cancels a SPECIFIC op).\r\n    op_seq: AtomicU64,\r\n    /// The absolute per-write bound (gate-wait + OS write completion).\r\n    bound: Duration,\r\n    /// The raw OS handle/fd of the underlying stream, captured at\r\n    /// construction for the out-of-band abort. The recv and send halves share\r\n    /// ONE underlying stream object (`interprocess` splits by ref-clone), so\r\n    /// aborting this handle covers a parked read too. Valid for the whole\r\n    /// conn lifetime: the half lives in `gate`, a writer's stack, or\r\n    /// `graveyard` until drop.\r\n    #[cfg(windows)]\r\n    raw: isize,\r\n    #[cfg(unix)]\r\n    raw: i32,\r\n}\r\n\r\n/// Recover a mutex guard even if a holder panicked — the state these guards\r\n/// protect is a plain slot/flag pair, valid on every panic boundary.\r\nfn recover<T>(m: &Mutex<T>) -> MutexGuard<'_, T> {\r\n    match m.lock() {\r\n        Ok(g) => g,\r\n        Err(p) => p.into_inner(),\r\n    }\r\n}\r\n\r\n#[cfg(windows)]\r\nfn raw_of(half: &SendHalf) -> isize {\r\n    use std::os::windows::io::{AsHandle, AsRawHandle};\r\n    match half {\r\n        SendHalf::NamedPipe(h) => h.as_handle().as_raw_handle() as isize,\r\n    }\r\n}\r\n\r\n#[cfg(unix)]\r\nfn raw_of(half: &SendHalf) -> i32 {\r\n    use std::os::fd::{AsFd, AsRawFd};\r\n    match half {\r\n        SendHalf::UdSocket(h) => h.as_fd().as_raw_fd(),\r\n    }\r\n}\r\n\r\n#[cfg(windows)]\r\nmod win {\r\n    // kernel32 is in the MSVC default link set — raw externs, no new dep\r\n    // (the deelevate.rs / daemon.rs posture).\r\n    extern \"system\" {\r\n        pub fn CancelIoEx(handle: isize, overlapped: *mut core::ffi::c_void) -> i32;\r\n        pub fn DisconnectNamedPipe(handle: isize) -> i32;\r\n    }\r\n}\r\n\r\nimpl Inner {\r\n    /// The attribution snapshot every lifecycle record carries: stable conn id,\r\n    /// accreted role/endpoint/session facts, wall + monotonic stamps.\r\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n    fn attribution(&self) -> String {\r\n        let label = recover(&self.label);\r\n        let facts: &str = if label.is_empty() { \"role=unattributed\" } else { &label };\r\n        match self.first_write.get() {\r\n            Some(start) => format!(\r\n                \"conn={} {} {} first_write_wall_ms={} first_write_mono_ms={}\",\r\n                self.id, facts, log_stamp(), start.wall_ms, start.mono_ms\r\n            ),\r\n            None => format!(\"conn={} {} {} first_write=none\", self.id, facts, log_stamp()),\r\n        }\r\n    }\r\n\r\n    /// Render the once-per-conn retirement record — the F-039 token split\r\n    /// (leg a): `CONN_WRITE_POISONED` is RESERVED for the deadline class\r\n    /// (`timed_out` latched — the wedge observable blackhole-controller.ps1\r\n    /// watches); an ORGANIC write failure (peer death / broken pipe on an\r\n    /// already-dead counterpart, the offline-peer dial churn) emits the\r\n    /// distinct `CONN_WRITE_RETIRED` token carrying the error kind and the\r\n    /// SAME attribution fields — quieter class, never thinner (the retired\r\n    /// class is what discriminates fresh-carrier churn from interleave\r\n    /// artifact in the field). Behavior is identical either way: poison,\r\n    /// abort, retire, never reuse.\r\n    // [impl->REQ-CONN-POISON-DIAL-SCOPE]\r\n    fn render_retirement(&self, timed_out: bool, cause: Option<&io::Error>) -> String {\r\n        if timed_out {\r\n            format!(\r\n                \"CONN_WRITE_POISONED: {} — physical broker conn retired: a write \\\r\n                 exceeded its {}ms bound; in-flight I/O aborted, conn never reused \\\r\n                 [REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\",\r\n                self.attribution(),\r\n                self.bound.as_millis()\r\n            )\r\n        } else {\r\n            let cause = match cause {\r\n                Some(e) => format!(\"{:?}: {e}\", e.kind()),\r\n                None => \"write failed\".to_string(),\r\n            };\r\n            format!(\r\n                \"CONN_WRITE_RETIRED: {} — organic write failure ({cause}); physical \\\r\n                 broker conn poisoned + retired, never reused (same retire behavior \\\r\n                 as POISONED — only the failure class differs) \\\r\n                 [REQ-CONN-POISON-DIAL-SCOPE]\",\r\n                self.attribution()\r\n            )\r\n        }\r\n    }\r\n\r\n    /// Per-conn exceptional and role-transition records remain unconditional.\r\n    /// releases#286 supersedes only healthy start/close emission, which uses\r\n    /// the process-wide interval counter instead. Retirement retains the\r\n    /// original first-write time even when the start line was aggregated.\r\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n    fn lifecycle(&self, event: &str, extra: &str) {\r\n        let sep = if extra.is_empty() { \"\" } else { \" \" };\r\n        self.lifecycle_log.emit(format_args!(\r\n            \"CONN_LIFECYCLE: {} event={event}{sep}{extra} [REQ-CONN-POISON-ATTRIBUTION]\",\r\n            self.attribution()\r\n        ));\r\n    }\r\n\r\n    /// Abort the connection's in-flight I/O OUT OF BAND (both directions —\r\n    /// the halves share one OS object) and latch the poison. Idempotent and\r\n    /// safe to repeat: `CancelIoEx` on a handle with nothing pending is a\r\n    /// no-op, and re-cancel closes the race where a write was issued a beat\r\n    /// AFTER a cancel (cancellation only reaches ALREADY-issued ops).\r\n    /// `cause` (the organic caller's write error) rides the retirement record\r\n    /// when the retirement is NOT deadline-driven.\r\n    fn poison_and_cancel(&self, cause: Option<&io::Error>) {\r\n        if !self.poisoned.swap(true, Ordering::AcqRel) {\r\n            // First poison of this conn: loud, once. The token is class-split\r\n            // (F-039 leg a): deadline → CONN_WRITE_POISONED (the wedge\r\n            // observable), organic fast-fail → CONN_WRITE_RETIRED.\r\n            self.lifecycle_log.emit(format_args!(\r\n                \"{}\",\r\n                self.render_retirement(self.timed_out.load(Ordering::Acquire), cause)\r\n            ));\r\n        }\r\n        #[cfg(windows)]\r\n        // SAFETY: scalar kernel32 call on a handle whose owning half is kept\r\n        // alive until Drop (gate / writer stack / graveyard — see `raw`).\r\n        unsafe {\r\n            let _ = win::CancelIoEx(self.raw, core::ptr::null_mut());\r\n        }\r\n        #[cfg(unix)]\r\n        // SAFETY: scalar shutdown(2) on an fd alive until Drop. SHUT_RDWR\r\n        // fails the blocked write (EPIPE), EOFs the parked read, and fails\r\n        // every future op — on Unix this one call IS the whole teardown.\r\n        unsafe {\r\n            let _ = libc::shutdown(self.raw, libc::SHUT_RDWR);\r\n        }\r\n        // Wake gate-waiters so they observe the poison instead of sleeping\r\n        // out their full deadline, and the watchdog so it re-evaluates.\r\n        self.gate_cv.notify_all();\r\n        self.dog_cv.notify_all();\r\n    }\r\n\r\n    /// Sever the physical connection (exactly once), AFTER the canceled op\r\n    /// reported completion: on Windows `DisconnectNamedPipe` forces the\r\n    /// client off (its next access fails), makes a re-issued server-side\r\n    /// read fail (so `handle_conn` reaches EOF cleanup even if it slipped a\r\n    /// new read in after the cancel), and defuses the drop-limbo flush; the\r\n    /// trailing re-cancel covers exactly that slipped-in read. On Unix the\r\n    /// `shutdown` in [`poison_and_cancel`] already did all of this.\r\n    // On Unix the cfg(windows) teardown strips out and the idempotence\r\n    // guard's early `return` becomes the function tail — needless_return\r\n    // there, load-bearing here. Restructuring (e.g. `if !swapped { … }`)\r\n    // would leave Unix an empty if-body (`clippy::needless_if`) instead.\r\n    #[cfg_attr(unix, allow(clippy::needless_return))]\r\n    fn sever(&self) {\r\n        if self.severed.swap(true, Ordering::AcqRel) {\r\n            return;\r\n        }\r\n        #[cfg(windows)]\r\n        // SAFETY: as in poison_and_cancel — handle outlives the watchdog.\r\n        unsafe {\r\n            let _ = win::DisconnectNamedPipe(self.raw);\r\n            let _ = win::CancelIoEx(self.raw, core::ptr::null_mut());\r\n        }\r\n    }\r\n\r\n    /// The watchdog body: park until a write is in flight, then sleep to its\r\n    /// absolute deadline. If the op is still in flight at the deadline,\r\n    /// poison + cancel, re-canceling periodically until the op reports\r\n    /// completion (the writer clears `inflight` the instant `write_frame`\r\n    /// returns — that return IS the buffer-release handshake), then sever.\r\n    fn watchdog(&self) {\r\n        let mut d = recover(&self.dog);\r\n        loop {\r\n            // Park until armed (or shutdown).\r\n            let (op, deadline) = loop {\r\n                if d.shutdown {\r\n                    return;\r\n                }\r\n                if let Some(armed) = d.inflight {\r\n                    break armed;\r\n                }\r\n                d = self\r\n                    .dog_cv\r\n                    .wait(d)\r\n                    .unwrap_or_else(|p| p.into_inner());\r\n            };\r\n            // Sleep toward the deadline while THIS op stays in flight.\r\n            let fired = loop {\r\n                if d.shutdown {\r\n                    return;\r\n                }\r\n                match d.inflight {\r\n                    Some((cur, _)) if cur == op => {}\r\n                    // Completed (or a newer op armed) before the deadline.\r\n                    _ => break false,\r\n                }\r\n                let now = Instant::now();\r\n                if now >= deadline {\r\n                    break true;\r\n                }\r\n                let (g, _) = self\r\n                    .dog_cv\r\n                    .wait_timeout(d, deadline - now)\r\n                    .unwrap_or_else(|p| p.into_inner());\r\n                d = g;\r\n            };\r\n            if !fired {\r\n                continue;\r\n            }\r\n            // Deadline: abort out of band, then wait for the canceled op to\r\n            // report completion before severing (re-cancel each slice — a\r\n            // cancel that raced the WriteFileEx issuance reaches nothing, so\r\n            // one shot is not enough). Latch timed_out FIRST so the aborted\r\n            // writer observes its failure as a DEADLINE retirement (wedge),\r\n            // not an organic conn death.\r\n            drop(d);\r\n            self.timed_out.store(true, Ordering::Release);\r\n            self.poison_and_cancel(None);\r\n            d = recover(&self.dog);\r\n            while matches!(d.inflight, Some((cur, _)) if cur == op) && !d.shutdown {\r\n                let (g, _) = self\r\n                    .dog_cv\r\n                    .wait_timeout(d, Duration::from_millis(100))\r\n                    .unwrap_or_else(|p| p.into_inner());\r\n                d = g;\r\n                drop(d);\r\n                self.poison_and_cancel(None);\r\n                d = recover(&self.dog);\r\n            }\r\n            if d.shutdown {\r\n                return;\r\n            }\r\n            drop(d);\r\n            self.sever();\r\n            d = recover(&self.dog);\r\n        }\r\n    }\r\n}\r\n\r\n/// The shared, serialized, BOUNDED send side of one physical broker\r\n/// connection — what [`crate::broker::SharedSend`] now aliases. Owns the\r\n/// [`SendHalf`] (it never leaves this module) and the per-conn watchdog.\r\npub(crate) struct BrokerConn {\r\n    inner: Arc<Inner>,\r\n    /// Joined on drop — after `shutdown` is raised — so the raw handle the\r\n    /// watchdog aborts through strictly outlives every raw use.\r\n    dog: Option<JoinHandle<()>>,\r\n}\r\n\r\nimpl BrokerConn {\r\n    /// Wrap a freshly-split send half. `bound` is the ABSOLUTE per-write\r\n    /// deadline (gate-wait + OS write completion) — the broker passes its\r\n    /// `brain_write_deadline()` (`SPT_BRAIN_WRITE_DEADLINE_MS` knob).\r\n    pub(crate) fn new(half: SendHalf, bound: Duration) -> Self {\r\n        static LOG: LazyLock<Arc<LifecycleLog>> = LazyLock::new(|| Arc::new(LifecycleLog::new()));\r\n        Self::with_log(half, bound, Arc::clone(&LOG))\r\n    }\r\n\r\n    fn with_log(half: SendHalf, bound: Duration, lifecycle_log: Arc<LifecycleLog>) -> Self {\r\n        let raw = raw_of(&half);\r\n        let inner = Arc::new(Inner {\r\n            id: CONN_ID_SEQ.fetch_add(1, Ordering::Relaxed) + 1,\r\n            label: Mutex::new(String::new()),\r\n            first_write: OnceLock::new(),\r\n            lifecycle_log,\r\n            gate: Mutex::new(Gate { half: Some(half) }),\r\n            gate_cv: Condvar::new(),\r\n            dog: Mutex::new(Dog {\r\n                inflight: None,\r\n                shutdown: false,\r\n            }),\r\n            dog_cv: Condvar::new(),\r\n            poisoned: AtomicBool::new(false),\r\n            timed_out: AtomicBool::new(false),\r\n            severed: AtomicBool::new(false),\r\n            graveyard: Mutex::new(None),\r\n            op_seq: AtomicU64::new(0),\r\n            bound,\r\n            raw,\r\n        });\r\n        let dog_inner = Arc::clone(&inner);\r\n        let dog = thread::Builder::new()\r\n            .name(\"conn-watchdog\".into())\r\n            .spawn(move || dog_inner.watchdog())\r\n            .ok();\r\n        BrokerConn { inner, dog }\r\n    }\r\n\r\n    /// The stable per-physical-conn id — what lifecycle records key on\r\n    /// (`Arc::ptr_eq` identity does not survive a log line).\r\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n    pub(crate) fn id(&self) -> u64 {\r\n        self.inner.id\r\n    }\r\n\r\n    /// Accrete an attribution fact onto this conn's label (subscriber role,\r\n    /// endpoint/session where known — e.g. `role=brain`,\r\n    /// `controller session=7 by=cli@host`). Facts accumulate space-separated\r\n    /// and ride EVERY subsequent lifecycle record for this conn; growth is\r\n    /// bounded by [`LABEL_CAP`] (one `…` marker, then no-op).\r\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n    pub(crate) fn describe(&self, fact: &str) {\r\n        let mut l = recover(&self.inner.label);\r\n        if l.ends_with('…') {\r\n            return;\r\n        }\r\n        if l.len() + fact.len() + 1 > LABEL_CAP {\r\n            l.push('…');\r\n            return;\r\n        }\r\n        if !l.is_empty() {\r\n            l.push(' ');\r\n        }\r\n        l.push_str(fact);\r\n    }\r\n\r\n    /// Emit a bounded per-conn lifecycle record through this conn's shared\r\n    /// attribution format — the broker's sink machinery uses it for\r\n    /// `writer-exit`, attach/replace, and detach events (leg d).\r\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n    pub(crate) fn lifecycle_event(&self, event: &str, extra: &str) {\r\n        self.inner.lifecycle(event, extra);\r\n    }\r\n\r\n    /// Bounded framed write. See [`Self::write_gated`].\r\n    pub(crate) fn write(&self, env: &Envelope) -> io::Result<ConnWrite> {\r\n        self.write_gated(env, || true)\r\n    }\r\n\r\n    /// Bounded framed write with a pre-write gate closure, run AFTER the\r\n    /// serialized gate is acquired and BEFORE any byte is written — the same\r\n    /// serialization point the old epoch-check-under-`send.lock()` had, so a\r\n    /// superseded controller writer still bails atomically with respect to\r\n    /// its successor's writes (P1c, REQ-HAZARD-CONTROLLER-WRITER-REORDER).\r\n    /// `pre` returning false → `Ok(ConnWrite::Superseded)`, nothing written.\r\n    pub(crate) fn write_gated(\r\n        &self,\r\n        env: &Envelope,\r\n        pre: impl FnOnce() -> bool,\r\n    ) -> io::Result<ConnWrite> {\r\n        let inner = &*self.inner;\r\n        let deadline = Instant::now() + inner.bound;\r\n\r\n        // ── Acquire the gate, bounded by the SAME absolute deadline. ──\r\n        let mut g = recover(&inner.gate);\r\n        loop {\r\n            if inner.poisoned.load(Ordering::Acquire) {\r\n                return Err(poisoned_err());\r\n            }\r\n            if g.half.is_some() {\r\n                break;\r\n            }\r\n            let now = Instant::now();\r\n            if now >= deadline {\r\n                // The write we queued behind is stuck past OUR whole bound:\r\n                // that op is the wedge. Abort it out of band and fail — the\r\n                // watchdog's own deadline (armed at that op's entry) is also\r\n                // in flight; both paths converge on the same idempotent\r\n                // poison/cancel.\r\n                drop(g);\r\n                inner.timed_out.store(true, Ordering::Release);\r\n                inner.poison_and_cancel(None);\r\n                return Err(timeout_err());\r\n            }\r\n            let (g2, _) = inner\r\n                .gate_cv\r\n                .wait_timeout(g, deadline - now)\r\n                .unwrap_or_else(|p| p.into_inner());\r\n            g = g2;\r\n        }\r\n\r\n        if !pre() {\r\n            return Ok(ConnWrite::Superseded);\r\n        }\r\n\r\n        // The gate serializes first-write initialization. Save the instant\r\n        // BEFORE I/O so a later failure still names when this conn went active.\r\n        // [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\r\n        if inner.first_write.get().is_none() {\r\n            let _ = inner.first_write.set(LogTime::now());\r\n            inner.lifecycle_log.healthy(true);\r\n        }\r\n\r\n        // ── Check out the half + arm the watchdog, then write OUT of locks. ──\r\n        let mut half = g.half.take().expect(\"gate half present at checkout\");\r\n        let op = inner.op_seq.fetch_add(1, Ordering::Relaxed) + 1;\r\n        {\r\n            let mut d = recover(&inner.dog);\r\n            d.inflight = Some((op, deadline));\r\n            inner.dog_cv.notify_all();\r\n        }\r\n        drop(g);\r\n\r\n        let res = write_frame(&mut half, env);\r\n\r\n        // Completion handshake: `write_frame` has returned, so the OS op —\r\n        // canceled or not — reported completion and the frame buffer is\r\n        // released. Clear the arm BEFORE deciding health, so the watchdog\r\n        // never severs while the op is formally in flight.\r\n        {\r\n            let mut d = recover(&inner.dog);\r\n            d.inflight = None;\r\n            inner.dog_cv.notify_all();\r\n        }\r\n\r\n        let mut g = recover(&inner.gate);\r\n        if res.is_ok() && !inner.poisoned.load(Ordering::Acquire) {\r\n            g.half = Some(half);\r\n            drop(g);\r\n            inner.gate_cv.notify_all();\r\n            return Ok(ConnWrite::Done);\r\n        }\r\n\r\n        // ── Retire: poison the whole conn, abort the read side so\r\n        // `handle_conn` reaches EOF cleanup, sever, and park the half in the\r\n        // graveyard (its handle must outlive any late raw abort; never\r\n        // re-gated, never reused — a timed-out length-prefixed frame may be\r\n        // partially written). ──\r\n        drop(g);\r\n        inner.poison_and_cancel(res.as_ref().err());\r\n        inner.sever();\r\n        *recover(&inner.graveyard) = Some(half);\r\n        inner.gate_cv.notify_all();\r\n        // A DEADLINE retirement (the watchdog aborted OUR op, or the conn was\r\n        // already deadline-poisoned) surfaces as TimedOut — the caller can\r\n        // count the wedge (W2 stall-evict tally). An organic I/O failure\r\n        // (peer death) keeps its own error.\r\n        if inner.timed_out.load(Ordering::Acquire) {\r\n            return Err(timeout_err());\r\n        }\r\n        Err(match res {\r\n            Err(e) => e,\r\n            Ok(()) => poisoned_err(),\r\n        })\r\n    }\r\n}\r\n\r\n/// TEST-ONLY determinism lever: holds the serialized write gate (the half is\r\n/// checked out) so every writer parks AT THE GATE until the pin drops — the\r\n/// round-4 analog of pinning the old `Mutex<SendHalf>` across a registration\r\n/// window (the P1c epoch-gate keystone unit relies on it).\r\n#[cfg(test)]\r\npub(crate) struct TestGatePin {\r\n    inner: Arc<Inner>,\r\n    half: Option<SendHalf>,\r\n}\r\n\r\n#[cfg(test)]\r\nimpl Drop for TestGatePin {\r\n    fn drop(&mut self) {\r\n        let mut g = recover(&self.inner.gate);\r\n        g.half = self.half.take();\r\n        drop(g);\r\n        self.inner.gate_cv.notify_all();\r\n    }\r\n}\r\n\r\n#[cfg(test)]\r\nimpl BrokerConn {\r\n    /// Check the half out of the gate without arming the watchdog — writers\r\n    /// queue behind the gate (bounded by their own deadlines) until the\r\n    /// returned pin drops. Test-only; never a production path.\r\n    pub(crate) fn pin_gate_for_test(&self) -> TestGatePin {\r\n        let mut g = recover(&self.inner.gate);\r\n        let half = g.half.take().expect(\"pin: gate half present\");\r\n        TestGatePin {\r\n            inner: Arc::clone(&self.inner),\r\n            half: Some(half),\r\n        }\r\n    }\r\n}\r\n\r\nimpl Drop for BrokerConn {\r\n    fn drop(&mut self) {\r\n        // Only a healthy physical close is counted instead of logged per conn.\r\n        // Error/timeout closes keep their full attribution and original start.\r\n        // [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\r\n        if self.inner.poisoned.load(Ordering::Acquire) {\r\n            self.inner.lifecycle(\r\n                \"transport-close\",\r\n                &format!(\r\n                    \"poisoned=true timed_out={}\",\r\n                    self.inner.timed_out.load(Ordering::Acquire)\r\n                ),\r\n            );\r\n        } else {\r\n            self.inner.lifecycle_log.healthy(false);\r\n        }\r\n        {\r\n            let mut d = recover(&self.inner.dog);\r\n            d.shutdown = true;\r\n            self.inner.dog_cv.notify_all();\r\n        }\r\n        if let Some(dog) = self.dog.take() {\r\n            let _ = dog.join();\r\n        }\r\n        // Halves (gate or graveyard) drop AFTER the watchdog is gone — no raw\r\n        // handle use can race the CloseHandle/limbo path.\r\n    }\r\n}\r","startLine":227,"lineNumbers":[227,null,229,230,231,232,233,234,235,236,237,238,239,240,241,242,243,244,245,246,247,248,249,250,251,252,253,254,255,256,257,258,259,260,261,262,263,264,265,266,267,268,269,270,271,272,273,274,275,276,277,278,279,280,281,282,283,284,285,286,287,288,289,290,291,292,293,294,295,296,297,298,299,300,301,302,303,304,305,306,307,308,309,310,311,312,313,314,315,316,317,318,319,320,321,322,323,324,325,326,327,328,329,330,331,332,333,334,335,336,337,338,339,340,341,342,343,344,345,346,347,348,349,350,351,352,353,354,355,356,357,358,359,360,361,362,363,364,365,366,367,368,369,370,371,372,373,374,375,376,377,378,379,380,381,382,383,384,385,386,387,388,389,390,391,392,393,394,395,396,397,398,399,400,401,402,403,404,405,406,407,408,409,410,411,412,413,414,415,416,417,418,419,420,421,422,423,424,425,426,427,428,429,430,431,432,433,434,435,436,437,438,439,440,441,442,443,444,445,446,447,448,449,450,451,452,453,454,455,456,457,458,459,460,461,462,463,464,465,466,467,468,469,470,471,472,473,474,475,476,477,478,479,480,481,482,483,484,485,486,487,488,489,490,491,492,493,494,495,496,497,498,499,500,501,502,503,504,505,506,507,508,509,510,511,512,513,514,515,516,517,518,519,520,521,522,523,524,525,526,527,528,529,530,531,532,533,534,535,536,537,538,539,540,541,542,543,544,545,546,547,548,549,550,551,552,553,554,555,556,557,558,559,560,561,562,563,564,565,566,567,568,569,570,571,572,573,574,575,576,577,578,579,580,581,582,583,584,585,586,587,588,589,590,591,592,593,594,595,596,597,598,599,600,601,602,603,604,605,606,607,608,609,610,611,612,613,614,615,616,617,618,619,620,621,622,623,624,625,626,627,628,629,630,631,632,633,634,635,636,637,638,639,640,641,642,643,644,645,646,647,648,649,650,651,652,653,654,655,656,657,658,659,660,661,662,663,664,665,666,667,668,669,670,671,672,673,674,675,676,677,678,679,680,681,682,683,684,685,686,687,688,689,690,691,692,693,694,695,696,697,698,699,700,701,702,703,704,705,706,707,708,709,710,711,712,713,714,715,716,717,718,719,720,721,722,723,724,725,726,727,728,729,730,731,732,733,734,735,736,737,738,739,740,741,742,743,744,745,746,747,748,749,750,751,752,753,754,755,756,757,758,759,760,761,762,763,764,765,766,767,768,769,770,771,772,773]},"fileSize":43321,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-daemon\\src\\conn.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":1000,"totalBytes":23120,"outputLines":545,"outputBytes":23120,"shownRange":{"start":229,"end":773},"nextOffset":774}}},"isError":false,"timestamp":1788999513779}}
{"type":"message","id":"af833770","parentId":"57b1a2a7","timestamp":"2026-09-10T00:18:33.822Z","message":{"role":"toolResult","toolCallId":"call_2V4wFZQ1QRjp0Ge0UwqRe9QB|fc_04ba494197989c86016aa1f759aaa887d18df792da4212f9be","toolName":"read","content":[{"type":"text","text":"\n/// The livehost session census is request/reply only, never a subscription.\n/// Retain its healthy carrier across 5s polls; discard it on EVERY query error\n/// so a failed/partially consumed reply can never contaminate the next census.\n/// `None` still means skip reconciliation, not an empty authoritative census.\n// [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n#[derive(Default)]\nstruct SessionPoll {\n    brain: Option<Brain>,\n    retry: bool,\n}\n\nimpl SessionPoll {\n    fn query(&mut self, broker_name: &str) -> Option<BTreeSet<String>> {\n        if self.brain.is_none() {\n            let reason = if self.retry { \"previous-query-error\" } else { \"initial-query\" };\n            match Brain::cold_start(broker_name, now_ms()) {\n                Ok(brain) => {\n                    spt_proto::emit_line_err!(\n                        \"BRAIN_CONN_OPEN: caller=livehost-session-query reason={reason} \\\n                         policy=reuse-until-query-error {}\",\n                        crate::conn::log_stamp()\n                    );\n                    self.brain = Some(brain);\n                }\n                Err(e) => {\n                    spt_proto::emit_line_err!(\n                        \"BRAIN_CONN_OPEN_FAIL: caller=livehost-session-query reason={reason} \\\n                         kind={:?}: {e} {}\", e.kind(), crate::conn::log_stamp()\n                    );\n                    return None;\n                }\n            }\n        }\n        match self.brain.as_mut().expect(\"connected session census\").sessions() {\n            Ok(reply) => Some(\n                reply.sessions.into_iter().map(|s| s.endpoint)\n                    .filter(|e| !e.is_empty()).collect()\n            ),\n            Err(e) => {\n                // No same-tick retry: the caller skips this census exactly as\n                // before. Next tick starts a fresh physical connection.\n                self.brain = None;\n                self.retry = true;\n                spt_proto::emit_line_err!(\n                    \"BRAIN_CONN_RETIRED: caller=livehost-session-query reason=query-error \\\n                     kind={:?}: {e} {}\", e.kind(), crate::conn::log_stamp()\n                );\n                None\n            }\n        }\n    }\n}\n\n/// The normalized program basename a perch's adapter would spawn its Psyche as —\n/// the first whitespace token of `psyche_init.command`, file-name only, run through\n/// [`normalize_basename`](spt_runtime::resolve::normalize_basename) (lowercase /\n…\n/// Spawn the brain's live host: one thread sweeping [`reconcile_once`] at boot\n/// and every [`LIVE_RECONCILE_INTERVAL_MS`] until `stop`. The registered set +\n/// config are re-read each sweep (adapter / pulse-period changes land between\n/// ticks). Mirror of [`crate::shellwake::spawn_wake_host`]; the stop flag is held\n/// for symmetry — the threads die with the brain process and the next brain\n/// rehydrates from disk.\n// [impl->REQ-DAEMON-1]\npub fn spawn_live_host(stop: Arc<AtomicBool>, reason: StartReason) -> JoinHandle<()> {\n    std::thread::spawn(move || {\n        let set = Arc::new(LiveSet::new());\n        {\n            // F-030 W5 ONE-SHOT legacy-resident sweep (REQ-PSYCHE-LEGACY-RESIDENT-SWEEP):\n            // BEFORE the first reconcile, bury any pre-W3 psyche WRAPPER a cross-version\n            // upgrade left ALIVE-but-hung (it registered its `<id>-psyche` ready perch then\n            // blocked forever on `spt ready --once`; the per-event driver never delivers to\n            // it). Fires exactly once at brain start (never per-reconcile) — the retired\n            // residency machinery's corpse-burial, not its resurrection.\n            // [impl->REQ-PSYCHE-LEGACY-RESIDENT-SWEEP]\n            let owlery = perch::owlery_dir();\n            let adapters_dir = perch::adapters_dir();\n            let registered = spt_runtime::registry::registered(&adapters_dir);\n            sweep_legacy_resident_psyches(&owlery, &registered, &adapters_dir);\n            // F-030 W5 ONE-SHOT `.live-bin` GC — ORDERED AFTER the sweep so any wrapper\n            // running FROM the own-copy image is killed first (its dir then unlocks).\n            gc_live_bin_dirs(&owlery);\n            // LIFECYCLE-TRUTH W3 ONE-SHOT daemon-restart endpoint survival\n            // (REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL): re-run the harness of every\n            // previously-ONLINE spt-hosted endpoint this (re)start orphaned, BEFORE the\n            // reconcile loop's phantom-clear offlines the stale-online records (the \"was\n            // online\" signal). A brain restart with a surviving broker re-runs nothing\n            // (its sessions are live, not orphaned). Belts inside the gate.\n            // [impl->REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL]\n            resume_restart_orphaned_endpoints(&owlery, &registered, &adapters_dir);\n        }\n        let mut session_poll = SessionPoll::default();\n        while !stop.load(Ordering::SeqCst) {\n            // TEST-ONLY gate `SPT_LIVEHOST_RECONCILE_DISABLE` (runtime env, DEFAULT\n            // unset = normal production, untouched): skip the reconcile poll+body so an\n            // int rig can prove REQ-UPDATE-TRIAL-DRAIN-DRIVE's core-loop KIND_SESSIONS\n            // reap-driver in ISOLATION. This loop's own `SessionPoll::query`\n            // → `brain.sessions()` (every LIVE_RECONCILE_INTERVAL_MS) would OTHERWISE\n            // drive the same broker reap and mask the fix (todlando 2026-07-09). Setting\n            // it faithfully REPRODUCES THE FIELD'S livehost-silent update-trial condition\n            // (PIN Q2: the trial-brain livehost's first reconcile tick landed around the\n            // 30s kill, so it never drove the reap) — the exact failure mode the fix\n            // makes the wedge-reap independent of. Runtime env (not cfg(test)): the rig\n            // spawns a real release `spt daemon brain` child cfg(test) can't reach — the\n            // SPT_BRAIN_WRITE_DEADLINE_MS precedent. Still honors `stop` on a tight tick.\n            // [impl->REQ-UPDATE-TRIAL-DRAIN-DRIVE]\n            if std::env::var_os(\"SPT_LIVEHOST_RECONCILE_DISABLE\").is_some() {\n                std::thread::sleep(Duration::from_millis(100));\n                continue;\n            }\n            let owlery = perch::owlery_dir();\n            let adapters_dir = perch::adapters_dir();\n            let registered = spt_runtime::registry::registered(&adapters_dir);\n            let cfg = DaemonConfig::load();\n            // B2 keystone (REQ-HAZARD-HOSTED-LIVENESS-RECONCILE) + B5 boot-gate\n            // (REQ-HAZARD-LIVEHOST-BOOT-LIVENESS-GATE): clear the status=online\n            // latch for any spt-hosted endpoint whose broker session is gone — and\n            // do it from BOOT tick 1, NOT after a grace. Runs BEFORE reconcile_once\n            // so the same-tick stop-side un-hosts the now-offline perch and the\n            // start-side never REVIVES its Psyche. The boot-gate is what stops a\n            // cold start after an unclean stop from spawning N phantom Psyches for N\n            // dead-but-latched perches. Skipped only when the broker is unreachable\n            // (None) — never mass-offline on a hiccup; safe at tick 1 because\n            // run_brain connects to the broker BEFORE this host spawns, so the query\n            // returns the REAL session set: empty at a cold start (every online\n            // controllable perch is a phantom → offlined), the surviving sessions at\n            // a brain restart (live agents stay, only the dead are offlined). The\n            // controllable==Some(true) gate inside keeps relay/legacy perches exempt.\n            // [impl->REQ-HAZARD-HOSTED-LIVENESS-RECONCILE]\n            // [impl->REQ-HAZARD-LIVEHOST-BOOT-LIVENESS-GATE]\n            if let Some(live) = session_poll.query(&crate::endpoint::broker_socket_name()) {\n                reconcile_hosted_liveness(&owlery, &live);\n            }\n            reconcile_once(&owlery, &registered, &adapters_dir, &set, &cfg, reason);\n            // W2: the attachment TTL reaper rides this same 5 s pulse — snapshot\n            // bytes and the entry that names them are retired together, so an\n            // expired attachment stops costing disk as well as stops serving.\n            // [impl->REQ-WEB-ATTACHMENT-PULL]\n            reap_expired_attachments();\n            // Sleep in slices so a stop lands promptly.\n            let mut left = LIVE_RECONCILE_INTERVAL_MS;\n            while left > 0 && !stop.load(Ordering::SeqCst) {\n                let step = left.min(100);\n                std::thread::sleep(Duration::from_millis(step));\n                left -= step;\n            }\n        }\n    })\n}\n\n#[cfg(test)]\nmod tests {\n    use super::*;\n    use crate::test_home::with_home;\n    use std::time::{Duration, Instant};\n\n    // [unit->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n    #[test]\n    fn session_poll_reuses_healthy_carrier_and_reopens_after_query_error() {\n        use crate::codec::{read_frame, write_frame};\n        use crate::frame::{Envelope, Role};\n        use crate::transport::{recv_hello, DaemonTransport, LocalSocketTransport};\n        use std::sync::atomic::AtomicU32;\n\n        static SEQ: AtomicU32 = AtomicU32::new(0);\n        let name = format!(\n            \"spt-livehost-poll-{}-{}\", std::process::id(),\n            SEQ.fetch_add(1, Ordering::Relaxed)\n        );\n        let listener = LocalSocketTransport::bind(&name).unwrap();\n        const N: usize = 16;\n        let server = std::thread::spawn(move || {\n            let mut accepted = 0;\n            let mut conn = listener.accept().unwrap();\n            accepted += 1;\n            recv_hello(&mut conn, Role::Brain).unwrap();\n            for _ in 0..N {\n                assert_eq!(read_frame(&mut conn).unwrap().kind, crate::msg::KIND_SESSIONS);\n                write_frame(&mut conn, &Envelope::new(\n                    crate::msg::KIND_SESSIONS_REPLY, serde_json::json!({\"sessions\": []})\n                )).unwrap();\n            }\n            // A malformed reply is a query error even while the transport is\n            // still alive. It must retire, not leave a desynchronized cache.\n            assert_eq!(read_frame(&mut conn).unwrap().kind, crate::msg::KIND_SESSIONS);\n            write_frame(&mut conn, &Envelope::new(\n                crate::msg::KIND_SESSIONS_REPLY, serde_json::json!({\"sessions\": false})\n            )).unwrap();\n            assert!(read_frame(&mut conn).is_err(), \"failed query closes old carrier\");\n            let mut conn = listener.accept().unwrap();\n            accepted += 1;\n            recv_hello(&mut conn, Role::Brain).unwrap();\n            assert_eq!(read_frame(&mut conn).unwrap().kind, crate::msg::KIND_SESSIONS);\n            write_frame(&mut conn, &Envelope::new(\n                crate::msg::KIND_SESSIONS_REPLY, serde_json::json!({\"sessions\": []})\n            )).unwrap();\n            accepted\n        });\n        let mut poll = SessionPoll::default();\n        for _ in 0..N {\n            assert_eq!(poll.query(&name), Some(BTreeSet::new()));\n        }\n        assert_eq!(poll.query(&name), None, \"failure is not an empty session census\");\n        assert_eq!(poll.query(&name), Some(BTreeSet::new()));\n        drop(poll);\n        assert_eq!(server.join().unwrap(), 2, \"one healthy carrier plus one recovery\");\n    }\n\n    // A no-op Psyche summarizer command: spawns + exits 0 (we never depend on its\n    // output — the daemon online-stamp is what hosts it, not the pid).\n    #[cfg(windows)]\n    const NOOP_CMD: &str = \"cmd /C rem\";\n    #[cfg(unix)]\n    const NOOP_CMD: &str = \"true\";\n\n    /// Register a live-capable adapter `name` (declares `[session.psyche_init]` +\n    /// a watched commune/signoff dir) so `resolve_option_in` finds it from disk.\n    fn register_live_adapter(name: &str, drops: &Path) {\n        let dir = drops.to_string_lossy().replace('\\\\', \"/\");\n…\n    }\n…\n}"}],"details":{"displayContent":{"text":"\n/// The livehost session census is request/reply only, never a subscription.\n/// Retain its healthy carrier across 5s polls; discard it on EVERY query error\n/// so a failed/partially consumed reply can never contaminate the next census.\n/// `None` still means skip reconciliation, not an empty authoritative census.\n// [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n#[derive(Default)]\nstruct SessionPoll {\n    brain: Option<Brain>,\n    retry: bool,\n}\n\nimpl SessionPoll {\n    fn query(&mut self, broker_name: &str) -> Option<BTreeSet<String>> {\n        if self.brain.is_none() {\n            let reason = if self.retry { \"previous-query-error\" } else { \"initial-query\" };\n            match Brain::cold_start(broker_name, now_ms()) {\n                Ok(brain) => {\n                    spt_proto::emit_line_err!(\n                        \"BRAIN_CONN_OPEN: caller=livehost-session-query reason={reason} \\\n                         policy=reuse-until-query-error {}\",\n                        crate::conn::log_stamp()\n                    );\n                    self.brain = Some(brain);\n                }\n                Err(e) => {\n                    spt_proto::emit_line_err!(\n                        \"BRAIN_CONN_OPEN_FAIL: caller=livehost-session-query reason={reason} \\\n                         kind={:?}: {e} {}\", e.kind(), crate::conn::log_stamp()\n                    );\n                    return None;\n                }\n            }\n        }\n        match self.brain.as_mut().expect(\"connected session census\").sessions() {\n            Ok(reply) => Some(\n                reply.sessions.into_iter().map(|s| s.endpoint)\n                    .filter(|e| !e.is_empty()).collect()\n            ),\n            Err(e) => {\n                // No same-tick retry: the caller skips this census exactly as\n                // before. Next tick starts a fresh physical connection.\n                self.brain = None;\n                self.retry = true;\n                spt_proto::emit_line_err!(\n                    \"BRAIN_CONN_RETIRED: caller=livehost-session-query reason=query-error \\\n                     kind={:?}: {e} {}\", e.kind(), crate::conn::log_stamp()\n                );\n                None\n            }\n        }\n    }\n}\n\n/// The normalized program basename a perch's adapter would spawn its Psyche as —\n/// the first whitespace token of `psyche_init.command`, file-name only, run through\n/// [`normalize_basename`](spt_runtime::resolve::normalize_basename) (lowercase /\n…\n/// Spawn the brain's live host: one thread sweeping [`reconcile_once`] at boot\n/// and every [`LIVE_RECONCILE_INTERVAL_MS`] until `stop`. The registered set +\n/// config are re-read each sweep (adapter / pulse-period changes land between\n/// ticks). Mirror of [`crate::shellwake::spawn_wake_host`]; the stop flag is held\n/// for symmetry — the threads die with the brain process and the next brain\n/// rehydrates from disk.\n// [impl->REQ-DAEMON-1]\npub fn spawn_live_host(stop: Arc<AtomicBool>, reason: StartReason) -> JoinHandle<()> {\n    std::thread::spawn(move || {\n        let set = Arc::new(LiveSet::new());\n        {\n            // F-030 W5 ONE-SHOT legacy-resident sweep (REQ-PSYCHE-LEGACY-RESIDENT-SWEEP):\n            // BEFORE the first reconcile, bury any pre-W3 psyche WRAPPER a cross-version\n            // upgrade left ALIVE-but-hung (it registered its `<id>-psyche` ready perch then\n            // blocked forever on `spt ready --once`; the per-event driver never delivers to\n            // it). Fires exactly once at brain start (never per-reconcile) — the retired\n            // residency machinery's corpse-burial, not its resurrection.\n            // [impl->REQ-PSYCHE-LEGACY-RESIDENT-SWEEP]\n            let owlery = perch::owlery_dir();\n            let adapters_dir = perch::adapters_dir();\n            let registered = spt_runtime::registry::registered(&adapters_dir);\n            sweep_legacy_resident_psyches(&owlery, &registered, &adapters_dir);\n            // F-030 W5 ONE-SHOT `.live-bin` GC — ORDERED AFTER the sweep so any wrapper\n            // running FROM the own-copy image is killed first (its dir then unlocks).\n            gc_live_bin_dirs(&owlery);\n            // LIFECYCLE-TRUTH W3 ONE-SHOT daemon-restart endpoint survival\n            // (REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL): re-run the harness of every\n            // previously-ONLINE spt-hosted endpoint this (re)start orphaned, BEFORE the\n            // reconcile loop's phantom-clear offlines the stale-online records (the \"was\n            // online\" signal). A brain restart with a surviving broker re-runs nothing\n            // (its sessions are live, not orphaned). Belts inside the gate.\n            // [impl->REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL]\n            resume_restart_orphaned_endpoints(&owlery, &registered, &adapters_dir);\n        }\n        let mut session_poll = SessionPoll::default();\n        while !stop.load(Ordering::SeqCst) {\n            // TEST-ONLY gate `SPT_LIVEHOST_RECONCILE_DISABLE` (runtime env, DEFAULT\n            // unset = normal production, untouched): skip the reconcile poll+body so an\n            // int rig can prove REQ-UPDATE-TRIAL-DRAIN-DRIVE's core-loop KIND_SESSIONS\n            // reap-driver in ISOLATION. This loop's own `SessionPoll::query`\n            // → `brain.sessions()` (every LIVE_RECONCILE_INTERVAL_MS) would OTHERWISE\n            // drive the same broker reap and mask the fix (todlando 2026-07-09). Setting\n            // it faithfully REPRODUCES THE FIELD'S livehost-silent update-trial condition\n            // (PIN Q2: the trial-brain livehost's first reconcile tick landed around the\n            // 30s kill, so it never drove the reap) — the exact failure mode the fix\n            // makes the wedge-reap independent of. Runtime env (not cfg(test)): the rig\n            // spawns a real release `spt daemon brain` child cfg(test) can't reach — the\n            // SPT_BRAIN_WRITE_DEADLINE_MS precedent. Still honors `stop` on a tight tick.\n            // [impl->REQ-UPDATE-TRIAL-DRAIN-DRIVE]\n            if std::env::var_os(\"SPT_LIVEHOST_RECONCILE_DISABLE\").is_some() {\n                std::thread::sleep(Duration::from_millis(100));\n                continue;\n            }\n            let owlery = perch::owlery_dir();\n            let adapters_dir = perch::adapters_dir();\n            let registered = spt_runtime::registry::registered(&adapters_dir);\n            let cfg = DaemonConfig::load();\n            // B2 keystone (REQ-HAZARD-HOSTED-LIVENESS-RECONCILE) + B5 boot-gate\n            // (REQ-HAZARD-LIVEHOST-BOOT-LIVENESS-GATE): clear the status=online\n            // latch for any spt-hosted endpoint whose broker session is gone — and\n            // do it from BOOT tick 1, NOT after a grace. Runs BEFORE reconcile_once\n            // so the same-tick stop-side un-hosts the now-offline perch and the\n            // start-side never REVIVES its Psyche. The boot-gate is what stops a\n            // cold start after an unclean stop from spawning N phantom Psyches for N\n            // dead-but-latched perches. Skipped only when the broker is unreachable\n            // (None) — never mass-offline on a hiccup; safe at tick 1 because\n            // run_brain connects to the broker BEFORE this host spawns, so the query\n            // returns the REAL session set: empty at a cold start (every online\n            // controllable perch is a phantom → offlined), the surviving sessions at\n            // a brain restart (live agents stay, only the dead are offlined). The\n            // controllable==Some(true) gate inside keeps relay/legacy perches exempt.\n            // [impl->REQ-HAZARD-HOSTED-LIVENESS-RECONCILE]\n            // [impl->REQ-HAZARD-LIVEHOST-BOOT-LIVENESS-GATE]\n            if let Some(live) = session_poll.query(&crate::endpoint::broker_socket_name()) {\n                reconcile_hosted_liveness(&owlery, &live);\n            }\n            reconcile_once(&owlery, &registered, &adapters_dir, &set, &cfg, reason);\n            // W2: the attachment TTL reaper rides this same 5 s pulse — snapshot\n            // bytes and the entry that names them are retired together, so an\n            // expired attachment stops costing disk as well as stops serving.\n            // [impl->REQ-WEB-ATTACHMENT-PULL]\n            reap_expired_attachments();\n            // Sleep in slices so a stop lands promptly.\n            let mut left = LIVE_RECONCILE_INTERVAL_MS;\n            while left > 0 && !stop.load(Ordering::SeqCst) {\n                let step = left.min(100);\n                std::thread::sleep(Duration::from_millis(step));\n                left -= step;\n            }\n        }\n    })\n}\n\n#[cfg(test)]\nmod tests {\n    use super::*;\n    use crate::test_home::with_home;\n    use std::time::{Duration, Instant};\n\n    // [unit->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n    #[test]\n    fn session_poll_reuses_healthy_carrier_and_reopens_after_query_error() {\n        use crate::codec::{read_frame, write_frame};\n        use crate::frame::{Envelope, Role};\n        use crate::transport::{recv_hello, DaemonTransport, LocalSocketTransport};\n        use std::sync::atomic::AtomicU32;\n\n        static SEQ: AtomicU32 = AtomicU32::new(0);\n        let name = format!(\n            \"spt-livehost-poll-{}-{}\", std::process::id(),\n            SEQ.fetch_add(1, Ordering::Relaxed)\n        );\n        let listener = LocalSocketTransport::bind(&name).unwrap();\n        const N: usize = 16;\n        let server = std::thread::spawn(move || {\n            let mut accepted = 0;\n            let mut conn = listener.accept().unwrap();\n            accepted += 1;\n            recv_hello(&mut conn, Role::Brain).unwrap();\n            for _ in 0..N {\n                assert_eq!(read_frame(&mut conn).unwrap().kind, crate::msg::KIND_SESSIONS);\n                write_frame(&mut conn, &Envelope::new(\n                    crate::msg::KIND_SESSIONS_REPLY, serde_json::json!({\"sessions\": []})\n                )).unwrap();\n            }\n            // A malformed reply is a query error even while the transport is\n            // still alive. It must retire, not leave a desynchronized cache.\n            assert_eq!(read_frame(&mut conn).unwrap().kind, crate::msg::KIND_SESSIONS);\n            write_frame(&mut conn, &Envelope::new(\n                crate::msg::KIND_SESSIONS_REPLY, serde_json::json!({\"sessions\": false})\n            )).unwrap();\n            assert!(read_frame(&mut conn).is_err(), \"failed query closes old carrier\");\n            let mut conn = listener.accept().unwrap();\n            accepted += 1;\n            recv_hello(&mut conn, Role::Brain).unwrap();\n            assert_eq!(read_frame(&mut conn).unwrap().kind, crate::msg::KIND_SESSIONS);\n            write_frame(&mut conn, &Envelope::new(\n                crate::msg::KIND_SESSIONS_REPLY, serde_json::json!({\"sessions\": []})\n            )).unwrap();\n            accepted\n        });\n        let mut poll = SessionPoll::default();\n        for _ in 0..N {\n            assert_eq!(poll.query(&name), Some(BTreeSet::new()));\n        }\n        assert_eq!(poll.query(&name), None, \"failure is not an empty session census\");\n        assert_eq!(poll.query(&name), Some(BTreeSet::new()));\n        drop(poll);\n        assert_eq!(server.join().unwrap(), 2, \"one healthy carrier plus one recovery\");\n    }\n\n    // A no-op Psyche summarizer command: spawns + exits 0 (we never depend on its\n    // output — the daemon online-stamp is what hosts it, not the pid).\n    #[cfg(windows)]\n    const NOOP_CMD: &str = \"cmd /C rem\";\n    #[cfg(unix)]\n    const NOOP_CMD: &str = \"true\";\n\n    /// Register a live-capable adapter `name` (declares `[session.psyche_init]` +\n    /// a watched commune/signoff dir) so `resolve_option_in` finds it from disk.\n    fn register_live_adapter(name: &str, drops: &Path) {\n        let dir = drops.to_string_lossy().replace('\\\\', \"/\");\n…\n    }\n…\n}","startLine":997,"lineNumbers":[997,998,999,1000,1001,1002,1003,1004,1005,1006,1007,1008,1009,1010,1011,1012,1013,1014,1015,1016,1017,1018,1019,1020,1021,1022,1023,1024,1025,1026,1027,1028,1029,1030,1031,1032,1033,1034,1035,1036,1037,1038,1039,1040,1041,1042,1043,1044,1045,1046,1047,1048,1049,1050,1051,1052,1053,null,1339,1340,1341,1342,1343,1344,1345,1346,1347,1348,1349,1350,1351,1352,1353,1354,1355,1356,1357,1358,1359,1360,1361,1362,1363,1364,1365,1366,1367,1368,1369,1370,1371,1372,1373,1374,1375,1376,1377,1378,1379,1380,1381,1382,1383,1384,1385,1386,1387,1388,1389,1390,1391,1392,1393,1394,1395,1396,1397,1398,1399,1400,1401,1402,1403,1404,1405,1406,1407,1408,1409,1410,1411,1412,1413,1414,1415,1416,1417,1418,1419,1420,1421,1422,1423,1424,1425,1426,1427,1428,1429,1430,1431,1432,1433,1434,1435,1436,1437,1438,1439,1440,1441,1442,1443,1444,1445,1446,1447,1448,1449,1450,1451,1452,1453,1454,1455,1456,1457,1458,1459,1460,1461,1462,1463,1464,1465,1466,1467,1468,1469,1470,1471,1472,1473,1474,1475,1476,1477,1478,1479,1480,1481,1482,1483,1484,1485,1486,1487,1488,1489,1490,1491,1492,1493,1494,1495,1496,1497,1498,1499,1500,null,1511,null,2917]},"fileSize":150325,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-daemon\\src\\livehost.rs"}}},"isError":false,"timestamp":1788999513821}}
{"type":"message","id":"177013f5","parentId":"af833770","timestamp":"2026-09-10T00:18:33.830Z","message":{"role":"toolResult","toolCallId":"call_K2wFCixSKT8qdBaNFENxm8GO|fc_04ba494197989c86016aa1f759aa7887d1b84325e52f8d59b4","toolName":"read","content":[{"type":"text","text":"         static SEQ: AtomicU32 = AtomicU32::new(0);\n         let name = format!(\n             \"spt-daemon-connattr-{}-{}.sock\",\n@@ -697,13 +818,78 @@ mod tests {\n         let listener = LocalSocketTransport::bind(&name).expect(\"bind\");\n         let client = LocalSocketTransport::connect(&name).expect(\"connect\");\n         let server = listener.accept().expect(\"accept\");\n-        let (_recv, send) = server.split();\n+        let (recv, send) = server.split();\n         (\n-            BrokerConn::new(send, Duration::from_millis(2000)),\n+            BrokerConn::with_log(send, Duration::from_millis(2000), log),\n             client,\n+            recv,\n         )\n     }\n \n+    // [unit->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+    #[test]\n+    fn healthy_round_trips_and_reconnections_emit_one_interval_summary() {\n+        let now = Instant::now();\n+        let log = Arc::new(LifecycleLog::captured(now));\n+        // The real framed-write/drop path, both a persistent carrier and fresh\n+        // per-poll carriers. An isolated clock/sink avoids global test races.\n+        const N: u64 = 32;\n+        for round in 0..N {\n+            let (conn, mut client, mut recv) = make_duplex_with_log(Arc::clone(&log));\n+            conn.describe(\"role=brain\");\n+            for seq in 0..4 {\n+                let request = Envelope::new(\"poll\", serde_json::json!([round, seq]));\n+                write_frame(&mut client, &request).unwrap();\n+                let received = crate::codec::read_frame(&mut recv).unwrap();\n+                assert_eq!(received.payload, request.payload);\n+                assert_eq!(conn.write(&request).unwrap(), ConnWrite::Done);\n+                let reply = crate::codec::read_frame(&mut client).unwrap();\n+                assert_eq!(reply.payload, request.payload);\n+            }\n+            drop(conn);\n+        }\n+        assert!(recover(log.capture.as_ref().unwrap()).is_empty());\n+        *recover(log.clock.as_ref().unwrap()) = now + HEALTHY_LOG_INTERVAL;\n+        let (conn, mut client) = make_conn_with_log(Arc::clone(&log));\n+        conn.describe(\"role=brain\");\n+        conn.write(&Envelope::new(\"poll\", serde_json::json!(N))).unwrap();\n+        assert_eq!(crate::codec::read_frame(&mut client).unwrap().payload, serde_json::json!(N));\n+        drop(conn);\n+        let lines = recover(log.capture.as_ref().unwrap());\n+        assert_eq!(lines.len(), 1, \"healthy volume is per interval, not per conn\");\n+        assert!(lines[0].contains(&format!(\"starts={}\", N + 1)));\n+        assert!(lines[0].contains(&format!(\"closes={N}\")));\n+    }\n+\n+    // [unit->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+    // [unit->REQ-CONN-POISON-ATTRIBUTION]\n+    #[test]\n+    fn failures_and_role_events_bypass_healthy_aggregation_with_original_start() {\n+        let log = Arc::new(LifecycleLog::captured(Instant::now()));\n+        let (conn, mut client) = make_conn_with_log(Arc::clone(&log));\n+        conn.describe(\"role=brain controller session=7\");\n+        conn.write(&Envelope::new(\"poll\", serde_json::json!({}))).unwrap();\n+        crate::codec::read_frame(&mut client).unwrap();\n+        let start = *conn.inner.first_write.get().expect(\"first committed write\");\n+        conn.lifecycle_event(\"writer-exit\", \"reason=write-failed\");\n+        conn.inner.poison_and_cancel(Some(&io::Error::new(io::ErrorKind::BrokenPipe, \"peer gone\")));\n+        // Repeated cancellation must not repeat the retirement record.\n+        conn.inner.poison_and_cancel(None);\n+        let id = conn.id();\n+        drop(conn);\n+        let lines = recover(log.capture.as_ref().unwrap());\n+        assert_eq!(lines.len(), 3, \"role event, retirement and poisoned close survive\");\n+        for line in lines.iter() {\n+            assert!(line.contains(&format!(\"conn={id}\")));\n+            assert!(line.contains(\"role=brain controller session=7\"));\n+            assert!(line.contains(&format!(\"first_write_wall_ms={}\", start.wall_ms)));\n+            assert!(line.contains(&format!(\"first_write_mono_ms={}\", start.mono_ms)));\n+        }\n+        assert!(lines[1].contains(\"CONN_WRITE_RETIRED:\"));\n+        assert!(lines[1].contains(\"BrokenPipe\"));\n+        assert!(lines[2].contains(\"poisoned=true timed_out=false\"));\n+    }\n+\n     // [unit->REQ-CONN-POISON-DIAL-SCOPE] the F-039 leg-(a) token split: the loud\n     // CONN_WRITE_POISONED token is RESERVED for the deadline (timed_out) class;\n     // an organic write failure renders the distinct CONN_WRITE_RETIRED token\ndiff --git a/crates/spt-daemon/src/livehost.rs b/crates/spt-daemon/src/livehost.rs\nindex 63f479d9..151c5e2e 100644\n--- a/crates/spt-daemon/src/livehost.rs\n+++ b/crates/spt-daemon/src/livehost.rs\n@@ -569,7 +569,7 @@ pub fn resume_restart_orphaned_endpoints(\n     registered: &[(AdapterRecord, Manifest)],\n     adapters_dir: &Path,\n ) {\n-    let Some(live) = query_live_session_endpoints() else {\n+    let Some(live) = SessionPoll::default().query(&crate::endpoint::broker_socket_name()) else {\n         return; // broker unreachable — skip (never mass-respawn on a hiccup)\n     };\n     for id in perch::list_self_perch_ids(owlery) {\n@@ -995,20 +995,57 @@ pub fn reconcile_hosted_liveness(owlery: &Path, live_sessions: &BTreeSet<String>\n     offlined\n }\n \n-/// Query the broker for the set of endpoint ids it currently hosts a session for\n-/// (`KIND_SESSIONS`) — the B2 pull signal. `None` when the broker is unreachable\n-/// (the caller then SKIPS the offline pass this tick rather than mass-offlining).\n-fn query_live_session_endpoints() -> Option<BTreeSet<String>> {\n-    let mut brain = Brain::cold_start(&crate::endpoint::broker_socket_name(), now_ms()).ok()?;\n-    let reply = brain.sessions().ok()?;\n-    Some(\n-        reply\n-            .sessions\n-            .into_iter()\n-            .map(|s| s.endpoint)\n-            .filter(|e| !e.is_empty())\n-            .collect(),\n-    )\n+/// The livehost session census is request/reply only, never a subscription.\n+/// Retain its healthy carrier across 5s polls; discard it on EVERY query error\n+/// so a failed/partially consumed reply can never contaminate the next census.\n+/// `None` still means skip reconciliation, not an empty authoritative census.\n+// [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+#[derive(Default)]\n+struct SessionPoll {\n+    brain: Option<Brain>,\n+    retry: bool,\n+}\n+\n+impl SessionPoll {\n+    fn query(&mut self, broker_name: &str) -> Option<BTreeSet<String>> {\n+        if self.brain.is_none() {\n+            let reason = if self.retry { \"previous-query-error\" } else { \"initial-query\" };\n+            match Brain::cold_start(broker_name, now_ms()) {\n+                Ok(brain) => {\n+                    spt_proto::emit_line_err!(\n+                        \"BRAIN_CONN_OPEN: caller=livehost-session-query reason={reason} \\\n+                         policy=reuse-until-query-error {}\",\n+                        crate::conn::log_stamp()\n+                    );\n+                    self.brain = Some(brain);\n+                }\n+                Err(e) => {\n+                    spt_proto::emit_line_err!(\n+                        \"BRAIN_CONN_OPEN_FAIL: caller=livehost-session-query reason={reason} \\\n+                         kind={:?}: {e} {}\", e.kind(), crate::conn::log_stamp()\n+                    );\n+                    return None;\n+                }\n+            }\n+        }\n+        match self.brain.as_mut().expect(\"connected session census\").sessions() {\n+            Ok(reply) => Some(\n+                reply.sessions.into_iter().map(|s| s.endpoint)\n+                    .filter(|e| !e.is_empty()).collect()\n+            ),\n+            Err(e) => {\n+                // No same-tick retry: the caller skips this census exactly as\n+                // before. Next tick starts a fresh physical connection.\n+                self.brain = None;\n+                self.retry = true;\n+                spt_proto::emit_line_err!(\n+                    \"BRAIN_CONN_RETIRED: caller=livehost-session-query reason=query-error \\\n+                     kind={:?}: {e} {}\", e.kind(), crate::conn::log_stamp()\n+                );\n+                None\n+            }\n+        }\n+    }\n }\n \n /// The normalized program basename a perch's adapter would spawn its Psyche as —\n@@ -1333,11 +1370,12 @@ pub fn spawn_live_host(stop: Arc<AtomicBool>, reason: StartReason) -> JoinHandle\n             // [impl->REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL]\n             resume_restart_orphaned_endpoints(&owlery, &registered, &adapters_dir);\n         }\n+        let mut session_poll = SessionPoll::default();\n         while !stop.load(Ordering::SeqCst) {\n             // TEST-ONLY gate `SPT_LIVEHOST_RECONCILE_DISABLE` (runtime env, DEFAULT\n             // unset = normal production, untouched): skip the reconcile poll+body so an\n             // int rig can prove REQ-UPDATE-TRIAL-DRAIN-DRIVE's core-loop KIND_SESSIONS\n-            // reap-driver in ISOLATION. This loop's own `query_live_session_endpoints()`\n+            // reap-driver in ISOLATION. This loop's own `SessionPoll::query`\n             // → `brain.sessions()` (every LIVE_RECONCILE_INTERVAL_MS) would OTHERWISE\n             // drive the same broker reap and mask the fix (todlando 2026-07-09). Setting\n             // it faithfully REPRODUCES THE FIELD'S livehost-silent update-trial condition\n@@ -1371,7 +1409,7 @@ pub fn spawn_live_host(stop: Arc<AtomicBool>, reason: StartReason) -> JoinHandle\n             // controllable==Some(true) gate inside keeps relay/legacy perches exempt.\n             // [impl->REQ-HAZARD-HOSTED-LIVENESS-RECONCILE]\n             // [impl->REQ-HAZARD-LIVEHOST-BOOT-LIVENESS-GATE]\n-            if let Some(live) = query_live_session_endpoints() {\n+            if let Some(live) = session_poll.query(&crate::endpoint::broker_socket_name()) {\n                 reconcile_hosted_liveness(&owlery, &live);\n             }\n             reconcile_once(&owlery, &registered, &adapters_dir, &set, &cfg, reason);\n@@ -1397,6 +1435,58 @@ mod tests {\n     use crate::test_home::with_home;\n     use std::time::{Duration, Instant};\n \n+    // [unit->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+    #[test]\n+    fn session_poll_reuses_healthy_carrier_and_reopens_after_query_error() {\n+        use crate::codec::{read_frame, write_frame};\n+        use crate::frame::{Envelope, Role};\n+        use crate::transport::{recv_hello, DaemonTransport, LocalSocketTransport};\n+        use std::sync::atomic::AtomicU32;\n+\n+        static SEQ: AtomicU32 = AtomicU32::new(0);\n+        let name = format!(\n+            \"spt-livehost-poll-{}-{}\", std::process::id(),\n+            SEQ.fetch_add(1, Ordering::Relaxed)\n+        );\n+        let listener = LocalSocketTransport::bind(&name).unwrap();\n+        const N: usize = 16;\n+        let server = std::thread::spawn(move || {\n+            let mut accepted = 0;\n+            let mut conn = listener.accept().unwrap();\n+            accepted += 1;\n+            recv_hello(&mut conn, Role::Brain).unwrap();\n+            for _ in 0..N {\n+                assert_eq!(read_frame(&mut conn).unwrap().kind, crate::msg::KIND_SESSIONS);\n+                write_frame(&mut conn, &Envelope::new(\n+                    crate::msg::KIND_SESSIONS_REPLY, serde_json::json!({\"sessions\": []})\n+                )).unwrap();\n+            }\n+            // A malformed reply is a query error even while the transport is\n+            // still alive. It must retire, not leave a desynchronized cache.\n+            assert_eq!(read_frame(&mut conn).unwrap().kind, crate::msg::KIND_SESSIONS);\n+            write_frame(&mut conn, &Envelope::new(\n+                crate::msg::KIND_SESSIONS_REPLY, serde_json::json!({\"sessions\": false})\n+            )).unwrap();\n+            assert!(read_frame(&mut conn).is_err(), \"failed query closes old carrier\");\n+            let mut conn = listener.accept().unwrap();\n+            accepted += 1;\n+            recv_hello(&mut conn, Role::Brain).unwrap();\n+            assert_eq!(read_frame(&mut conn).unwrap().kind, crate::msg::KIND_SESSIONS);\n+            write_frame(&mut conn, &Envelope::new(\n+                crate::msg::KIND_SESSIONS_REPLY, serde_json::json!({\"sessions\": []})\n+            )).unwrap();\n+            accepted\n+        });\n+        let mut poll = SessionPoll::default();\n+        for _ in 0..N {\n+            assert_eq!(poll.query(&name), Some(BTreeSet::new()));\n+        }\n+        assert_eq!(poll.query(&name), None, \"failure is not an empty session census\");\n+        assert_eq!(poll.query(&name), Some(BTreeSet::new()));\n+        drop(poll);\n+        assert_eq!(server.join().unwrap(), 2, \"one healthy carrier plus one recovery\");\n+    }\n+\n     // A no-op Psyche summarizer command: spawns + exits 0 (we never depend on its\n     // output — the daemon online-stamp is what hosts it, not the pid).\n     #[cfg(windows)]\ndiff --git a/docs/RCA-FLEET-DAEMON-14444.md b/docs/RCA-FLEET-DAEMON-14444.md\nindex f7e897d1..010613fd 100644\n--- a/docs/RCA-FLEET-DAEMON-14444.md\n+++ b/docs/RCA-FLEET-DAEMON-14444.md\n@@ -69,3 +69,76 @@ per 2 MB of log, under `[REQ-CONN-POISON-ATTRIBUTION]`, is its own finding.\n    (image path or start time), which is the only way reuse becomes a kill. My own standing rule from\n    the orphan reap is exactly this: re-verify the path AT kill time.\n 3. **Bound the log volume** independently of the cause.\n+\n+## releases#286 — stderr follow-up (2026-09-09)\n+\n+<!-- [doc->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED] -->\n+\n+The historical 303 MB incident above remains the reported baseline, not a\n+newly reproduced event. A read-only sample of the currently installed\n+`%LOCALAPPDATA%/spt-core/logs/daemon.stderr.log` ended at byte **4,163,267**:\n+the last **1,999,958 complete-line bytes** spanned wall stamps\n+**1788997126331–1788998151874** (1025.543 s). It contained **5127** first writes\n+(656,158 bytes), **5092** unpoisoned/non-timeout closes (850,622 bytes),\n+**3** organic retirements, **0** deadline poisons, **107** family-gate lines\n+and **108** rendezvous-up lines. Healthy start/close records alone consumed\n+**75.3%** of that sample. A subsequent live append observation over\n+**307.203 s** added **79,741 bytes** (**259.57 B/s**), with **104** starts,\n+**104** healthy closes and **30** family gates paired with **30** rendezvous\n+binds. These are different observation windows, not a claimed constant rate.\n+No daemon was restarted or reconfigured for either measurement.\n+\n+**Cause discrimination.** `role=brain` identifies *every broker IPC client*,\n+not the supervised coordinator. `run_brain` already keeps its heartbeat\n+`Brain` alive. In contrast, the pre-fix\n+`livehost::query_live_session_endpoints` constructed `Brain::cold_start`,\n+called `sessions`, and dropped the carrier on every invocation from the\n+five-second reconcile loop. This is a source-confirmed avoidable reopen\n+cause, consistent with the recurring five-second bare-brain start/close\n+pairs in the live log, **not an attribution of all 5127 connections to that\n+caller**. Other short-lived callers include dispatch idle queries and\n+per-operation dispatch/CLI clients; their lifetimes are not changed here.\n+\n+The livehost census now owns one carrier across polls. Its `BRAIN_CONN_OPEN`\n+breadcrumb names `caller=livehost-session-query` and distinguishes\n+`reason=initial-query` from `reason=previous-query-error`. Every query error\n+discards the carrier, returns an unavailable census (never an empty\n+authoritative census), and permits a fresh connection only on the next\n+scheduled query. The one-shot boot orphan census remains one-shot.\n+No read/write deadline, poison semantics, or retry cadence changed.\n+\n+**Family gate exonerated.** The sole diagnostic site in\n+`NetEndpoint::bind_scoped` is already once per real bind attempt, not per IPC\n+connection. `pairhost::spawn_meet_rotation` binds a new derived identity for\n+each attached subnet at every 30-second TOTP boundary; the observed\n+`BIGNET`, `SPT_DEV`, and `SPT_MANTLE` windows account for the three lines per\n+boundary. These listeners cannot share a stable identity without changing\n+the pairing protocol. Bind and diagnostic policies are intentionally\n+unchanged; a process-global \"log once\" would hide later real binds.\n+\n+**Explicit contract amendment.** Issue #286 supersedes only the\n+`REQ-CONN-POISON-ATTRIBUTION` leg-(d) requirement for unconditional\n+*healthy start/close lines*. Across the broker process these become\n+`CONN_LIFECYCLE_SUMMARY` totals, at most once per 60-second monotonic interval,\n+flushed by the first subsequent lifecycle counter event. Starts count the\n+first committed write attempt; closes count unpoisoned physical drops.\n+There is no timer thread and no shutdown flush, so a quiet tail remains\n+pending until another counter event (and can be lost on process exit).\n+Poison, organic failure, poisoned close, and all role-transition records\n+remain unconditional per connection. They retain conn id, role facts,\n+wall/monotonic stamps **and the original first-write wall/monotonic time**;\n+a connection that never wrote says `first_write=none`.\n+\n+The deterministic regression cells exercise real framed socket exchanges\n+over both retained and fresh carriers with an isolated fixed log clock/sink,\n+assert one interval summary for N healthy connections, preserve exceptional\n+records through the same sink, and drive livehost census reuse plus recovery\n+from a malformed reply. Focused commands (not executed in the preparation\n+lane; integration owner runs them after landing):\n+\n+```text\n+cargo test -p spt-daemon --lib conn::tests::\n+cargo test -p spt-daemon --lib livehost::tests::session_poll_reuses_healthy_carrier_and_reopens_after_query_error\n+cargo test -p spt-daemon --test conn_blackhole_lifecycle\n+traceable-reqs check --json\n+```\ndiff --git a/traceable-reqs.toml b/traceable-reqs.toml\nindex 21917ea9..9224b3aa 100644\n--- a/traceable-reqs.toml\n+++ b/traceable-reqs.toml\n@@ -2569,6 +2569,14 @@ id = \"REQ-CONN-POISON-ATTRIBUTION\"\n title = \"MSG-IDENTITY W6 / F-039 legs b-d (doyle W6 LOCK 2026-07-10, minted per amendment 3): every broker-conn lifecycle record is ATTRIBUTABLE — the W6 RCA's terminal undecidability (per-line 1:1 CONN_WRITE_POISONED churn = fresh-carrier churn OR stderr interleave artifact) exists because records carry no stable conn identity, no role/endpoint/session context, and no timestamps, and the once-per-conn poison latch hides multiplicity. THREE LEGS. (b) IDENTITY: mint a stable per-physical-conn id (monotonic u64 at conn construction — Arc::ptr_eq is the only identity today and it does not survive a log line) plus subscriber role and endpoint/session where known, stamped on CONN_WRITE_POISONED, CONN_WRITE_RETIRED, logical stall-evict, attach/resume/detach, and write-retirement records (RCA attach sites: presence nethost.rs:379, stream nethost.rs:258, controller broker.rs:891, viewer broker.rs:1073). (c) TIME: daemon stderr correlation records carry wall-clock AND monotonic timestamps (stderrlog has neither; broker+brain share one file — interleave is unresolvable without them). (d) LIFECYCLE (doyle-confirmed UNCONDITIONAL, not debug-gated): one BOUNDED set of per-conn lifecycle events — write start/timeout-cancel/transport close/writer exit/replacement-reattach (hertz RCA fix-shape items 1-3). Constraint (doyle LOCK): the split/attribution must not REDUCE total information, only correct its attribution; NO timeout-value changes; NO suppression-as-fix. Gate: unit — lifecycle records carry conn id + role + timestamps; the id is unique per physical conn and stable across that conn's records. Kin REQ-CONN-POISON-DIAL-SCOPE (leg a, the token split these fields ride on), REQ-CONN-BLACKHOLE-LIFECYCLE-HARNESS (leg e, consumes these records), REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK (behavior invariant preserved).\"\n required_stages = [\"impl\", \"unit\"]  # FLIPPED in the W6 build commit carrying the evidence (todlando 2026-07-10). impl = conn id mint + describe/label accretion + log_stamp wall+mono (anchored at both daemon entries) + bounded lifecycle events (write-start, transport-close, writer-exit, attach/replace/detach) + attribution on stall-evict records. unit = conn.rs record-shape assertions (id/role/time present + parity across both retirement tokens, id unique per conn, label bounded).\n \n+# releases#286 narrowly supersedes leg (d)'s unconditional healthy start/close\n+# lines. Failure/poison, role transitions and original first-write time remain\n+# attributable per physical conn; healthy counts are interval totals instead.\n+[[requirements]]\n+id = \"REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED\"\n+title = \"Healthy broker IPC starts and closes aggregate once per monotonic interval across connections, independently of request volume; poison, failure and role-transition evidence remains per-connection with original first-write time. Livehost session polling reuses its healthy carrier and names why it opens a new one; role=brain alone never identifies the supervised coordinator. Network family diagnostics remain once per real endpoint bind attempt, not IPC churn. Narrow healthy-start/close supersession of REQ-CONN-POISON-ATTRIBUTION leg d (releases#286).\"\n+required_stages = [\"doc\", \"impl\", \"unit\"]\n+\n [[requirements]]\n id = \"REQ-CONN-BLACKHOLE-LIFECYCLE-HARNESS\"\n title = \"MSG-IDENTITY W6 / F-039 leg e (doyle W6 LOCK 2026-07-10, minted per amendment 3 — hertz's five invariants VERBATIM from his RCA fix-shape item 5): 'Build a deterministic black-holed-controller harness against current v0.30.6 semantics and assert: unrelated sessions continue; the bad physical connection is canceled/closed within the bound; its writer exits; a fresh viewer can attach; no lock or task remains owned by the retired connection.' The harness is the standing conformance rig for the r4 SHAREDSEND fix-class — hertz's RCA discipline: only after a timestamped incident maps to a FAILING lifecycle invariant does an ownership/cancellation defect get fixed (the likely shape being complete physical-connection cancellation and writer-task join/retirement, never a broader timeout increase). Consumes REQ-CONN-POISON-ATTRIBUTION's records (conn id + lifecycle events are what make the five assertions checkable deterministically). Kin REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK (the invariant class under test — its brain_decouple int stays the Windows-mandatory gate leg), REQ-CONN-POISON-DIAL-SCOPE.\"\n@@ -4860,6 +4868,7 @@ name = \"message-format-identity\"\n requirements = [\n   \"REQ-CONN-BLACKHOLE-LIFECYCLE-HARNESS\",\n   \"REQ-CONN-POISON-ATTRIBUTION\",\n+  \"REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED\",\n   \"REQ-CONN-POISON-DIAL-SCOPE\",\n   \"REQ-DAEMON-SERVICE-INSTALL\",\n   \"REQ-EP-3\","}],"details":{"resolvedPath":"C:\\Users\\decid\\.omp\\agent\\sessions\\-Documents-projects-spt-core\\2026-09-09T23-45-42-317Z_01a08890-21ad-7000-8f54-208d4f4f2f2c\\62.bash-original.log","contentType":"text/plain","totalLines":658,"displayContent":{"text":"         static SEQ: AtomicU32 = AtomicU32::new(0);\n         let name = format!(\n             \"spt-daemon-connattr-{}-{}.sock\",\n@@ -697,13 +818,78 @@ mod tests {\n         let listener = LocalSocketTransport::bind(&name).expect(\"bind\");\n         let client = LocalSocketTransport::connect(&name).expect(\"connect\");\n         let server = listener.accept().expect(\"accept\");\n-        let (_recv, send) = server.split();\n+        let (recv, send) = server.split();\n         (\n-            BrokerConn::new(send, Duration::from_millis(2000)),\n+            BrokerConn::with_log(send, Duration::from_millis(2000), log),\n             client,\n+            recv,\n         )\n     }\n \n+    // [unit->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+    #[test]\n+    fn healthy_round_trips_and_reconnections_emit_one_interval_summary() {\n+        let now = Instant::now();\n+        let log = Arc::new(LifecycleLog::captured(now));\n+        // The real framed-write/drop path, both a persistent carrier and fresh\n+        // per-poll carriers. An isolated clock/sink avoids global test races.\n+        const N: u64 = 32;\n+        for round in 0..N {\n+            let (conn, mut client, mut recv) = make_duplex_with_log(Arc::clone(&log));\n+            conn.describe(\"role=brain\");\n+            for seq in 0..4 {\n+                let request = Envelope::new(\"poll\", serde_json::json!([round, seq]));\n+                write_frame(&mut client, &request).unwrap();\n+                let received = crate::codec::read_frame(&mut recv).unwrap();\n+                assert_eq!(received.payload, request.payload);\n+                assert_eq!(conn.write(&request).unwrap(), ConnWrite::Done);\n+                let reply = crate::codec::read_frame(&mut client).unwrap();\n+                assert_eq!(reply.payload, request.payload);\n+            }\n+            drop(conn);\n+        }\n+        assert!(recover(log.capture.as_ref().unwrap()).is_empty());\n+        *recover(log.clock.as_ref().unwrap()) = now + HEALTHY_LOG_INTERVAL;\n+        let (conn, mut client) = make_conn_with_log(Arc::clone(&log));\n+        conn.describe(\"role=brain\");\n+        conn.write(&Envelope::new(\"poll\", serde_json::json!(N))).unwrap();\n+        assert_eq!(crate::codec::read_frame(&mut client).unwrap().payload, serde_json::json!(N));\n+        drop(conn);\n+        let lines = recover(log.capture.as_ref().unwrap());\n+        assert_eq!(lines.len(), 1, \"healthy volume is per interval, not per conn\");\n+        assert!(lines[0].contains(&format!(\"starts={}\", N + 1)));\n+        assert!(lines[0].contains(&format!(\"closes={N}\")));\n+    }\n+\n+    // [unit->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+    // [unit->REQ-CONN-POISON-ATTRIBUTION]\n+    #[test]\n+    fn failures_and_role_events_bypass_healthy_aggregation_with_original_start() {\n+        let log = Arc::new(LifecycleLog::captured(Instant::now()));\n+        let (conn, mut client) = make_conn_with_log(Arc::clone(&log));\n+        conn.describe(\"role=brain controller session=7\");\n+        conn.write(&Envelope::new(\"poll\", serde_json::json!({}))).unwrap();\n+        crate::codec::read_frame(&mut client).unwrap();\n+        let start = *conn.inner.first_write.get().expect(\"first committed write\");\n+        conn.lifecycle_event(\"writer-exit\", \"reason=write-failed\");\n+        conn.inner.poison_and_cancel(Some(&io::Error::new(io::ErrorKind::BrokenPipe, \"peer gone\")));\n+        // Repeated cancellation must not repeat the retirement record.\n+        conn.inner.poison_and_cancel(None);\n+        let id = conn.id();\n+        drop(conn);\n+        let lines = recover(log.capture.as_ref().unwrap());\n+        assert_eq!(lines.len(), 3, \"role event, retirement and poisoned close survive\");\n+        for line in lines.iter() {\n+            assert!(line.contains(&format!(\"conn={id}\")));\n+            assert!(line.contains(\"role=brain controller session=7\"));\n+            assert!(line.contains(&format!(\"first_write_wall_ms={}\", start.wall_ms)));\n+            assert!(line.contains(&format!(\"first_write_mono_ms={}\", start.mono_ms)));\n+        }\n+        assert!(lines[1].contains(\"CONN_WRITE_RETIRED:\"));\n+        assert!(lines[1].contains(\"BrokenPipe\"));\n+        assert!(lines[2].contains(\"poisoned=true timed_out=false\"));\n+    }\n+\n     // [unit->REQ-CONN-POISON-DIAL-SCOPE] the F-039 leg-(a) token split: the loud\n     // CONN_WRITE_POISONED token is RESERVED for the deadline (timed_out) class;\n     // an organic write failure renders the distinct CONN_WRITE_RETIRED token\ndiff --git a/crates/spt-daemon/src/livehost.rs b/crates/spt-daemon/src/livehost.rs\nindex 63f479d9..151c5e2e 100644\n--- a/crates/spt-daemon/src/livehost.rs\n+++ b/crates/spt-daemon/src/livehost.rs\n@@ -569,7 +569,7 @@ pub fn resume_restart_orphaned_endpoints(\n     registered: &[(AdapterRecord, Manifest)],\n     adapters_dir: &Path,\n ) {\n-    let Some(live) = query_live_session_endpoints() else {\n+    let Some(live) = SessionPoll::default().query(&crate::endpoint::broker_socket_name()) else {\n         return; // broker unreachable — skip (never mass-respawn on a hiccup)\n     };\n     for id in perch::list_self_perch_ids(owlery) {\n@@ -995,20 +995,57 @@ pub fn reconcile_hosted_liveness(owlery: &Path, live_sessions: &BTreeSet<String>\n     offlined\n }\n \n-/// Query the broker for the set of endpoint ids it currently hosts a session for\n-/// (`KIND_SESSIONS`) — the B2 pull signal. `None` when the broker is unreachable\n-/// (the caller then SKIPS the offline pass this tick rather than mass-offlining).\n-fn query_live_session_endpoints() -> Option<BTreeSet<String>> {\n-    let mut brain = Brain::cold_start(&crate::endpoint::broker_socket_name(), now_ms()).ok()?;\n-    let reply = brain.sessions().ok()?;\n-    Some(\n-        reply\n-            .sessions\n-            .into_iter()\n-            .map(|s| s.endpoint)\n-            .filter(|e| !e.is_empty())\n-            .collect(),\n-    )\n+/// The livehost session census is request/reply only, never a subscription.\n+/// Retain its healthy carrier across 5s polls; discard it on EVERY query error\n+/// so a failed/partially consumed reply can never contaminate the next census.\n+/// `None` still means skip reconciliation, not an empty authoritative census.\n+// [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+#[derive(Default)]\n+struct SessionPoll {\n+    brain: Option<Brain>,\n+    retry: bool,\n+}\n+\n+impl SessionPoll {\n+    fn query(&mut self, broker_name: &str) -> Option<BTreeSet<String>> {\n+        if self.brain.is_none() {\n+            let reason = if self.retry { \"previous-query-error\" } else { \"initial-query\" };\n+            match Brain::cold_start(broker_name, now_ms()) {\n+                Ok(brain) => {\n+                    spt_proto::emit_line_err!(\n+                        \"BRAIN_CONN_OPEN: caller=livehost-session-query reason={reason} \\\n+                         policy=reuse-until-query-error {}\",\n+                        crate::conn::log_stamp()\n+                    );\n+                    self.brain = Some(brain);\n+                }\n+                Err(e) => {\n+                    spt_proto::emit_line_err!(\n+                        \"BRAIN_CONN_OPEN_FAIL: caller=livehost-session-query reason={reason} \\\n+                         kind={:?}: {e} {}\", e.kind(), crate::conn::log_stamp()\n+                    );\n+                    return None;\n+                }\n+            }\n+        }\n+        match self.brain.as_mut().expect(\"connected session census\").sessions() {\n+            Ok(reply) => Some(\n+                reply.sessions.into_iter().map(|s| s.endpoint)\n+                    .filter(|e| !e.is_empty()).collect()\n+            ),\n+            Err(e) => {\n+                // No same-tick retry: the caller skips this census exactly as\n+                // before. Next tick starts a fresh physical connection.\n+                self.brain = None;\n+                self.retry = true;\n+                spt_proto::emit_line_err!(\n+                    \"BRAIN_CONN_RETIRED: caller=livehost-session-query reason=query-error \\\n+                     kind={:?}: {e} {}\", e.kind(), crate::conn::log_stamp()\n+                );\n+                None\n+            }\n+        }\n+    }\n }\n \n /// The normalized program basename a perch's adapter would spawn its Psyche as —\n@@ -1333,11 +1370,12 @@ pub fn spawn_live_host(stop: Arc<AtomicBool>, reason: StartReason) -> JoinHandle\n             // [impl->REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL]\n             resume_restart_orphaned_endpoints(&owlery, &registered, &adapters_dir);\n         }\n+        let mut session_poll = SessionPoll::default();\n         while !stop.load(Ordering::SeqCst) {\n             // TEST-ONLY gate `SPT_LIVEHOST_RECONCILE_DISABLE` (runtime env, DEFAULT\n             // unset = normal production, untouched): skip the reconcile poll+body so an\n             // int rig can prove REQ-UPDATE-TRIAL-DRAIN-DRIVE's core-loop KIND_SESSIONS\n-            // reap-driver in ISOLATION. This loop's own `query_live_session_endpoints()`\n+            // reap-driver in ISOLATION. This loop's own `SessionPoll::query`\n             // → `brain.sessions()` (every LIVE_RECONCILE_INTERVAL_MS) would OTHERWISE\n             // drive the same broker reap and mask the fix (todlando 2026-07-09). Setting\n             // it faithfully REPRODUCES THE FIELD'S livehost-silent update-trial condition\n@@ -1371,7 +1409,7 @@ pub fn spawn_live_host(stop: Arc<AtomicBool>, reason: StartReason) -> JoinHandle\n             // controllable==Some(true) gate inside keeps relay/legacy perches exempt.\n             // [impl->REQ-HAZARD-HOSTED-LIVENESS-RECONCILE]\n             // [impl->REQ-HAZARD-LIVEHOST-BOOT-LIVENESS-GATE]\n-            if let Some(live) = query_live_session_endpoints() {\n+            if let Some(live) = session_poll.query(&crate::endpoint::broker_socket_name()) {\n                 reconcile_hosted_liveness(&owlery, &live);\n             }\n             reconcile_once(&owlery, &registered, &adapters_dir, &set, &cfg, reason);\n@@ -1397,6 +1435,58 @@ mod tests {\n     use crate::test_home::with_home;\n     use std::time::{Duration, Instant};\n \n+    // [unit->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n+    #[test]\n+    fn session_poll_reuses_healthy_carrier_and_reopens_after_query_error() {\n+        use crate::codec::{read_frame, write_frame};\n+        use crate::frame::{Envelope, Role};\n+        use crate::transport::{recv_hello, DaemonTransport, LocalSocketTransport};\n+        use std::sync::atomic::AtomicU32;\n+\n+        static SEQ: AtomicU32 = AtomicU32::new(0);\n+        let name = format!(\n+            \"spt-livehost-poll-{}-{}\", std::process::id(),\n+            SEQ.fetch_add(1, Ordering::Relaxed)\n+        );\n+        let listener = LocalSocketTransport::bind(&name).unwrap();\n+        const N: usize = 16;\n+        let server = std::thread::spawn(move || {\n+            let mut accepted = 0;\n+            let mut conn = listener.accept().unwrap();\n+            accepted += 1;\n+            recv_hello(&mut conn, Role::Brain).unwrap();\n+            for _ in 0..N {\n+                assert_eq!(read_frame(&mut conn).unwrap().kind, crate::msg::KIND_SESSIONS);\n+                write_frame(&mut conn, &Envelope::new(\n+                    crate::msg::KIND_SESSIONS_REPLY, serde_json::json!({\"sessions\": []})\n+                )).unwrap();\n+            }\n+            // A malformed reply is a query error even while the transport is\n+            // still alive. It must retire, not leave a desynchronized cache.\n+            assert_eq!(read_frame(&mut conn).unwrap().kind, crate::msg::KIND_SESSIONS);\n+            write_frame(&mut conn, &Envelope::new(\n+                crate::msg::KIND_SESSIONS_REPLY, serde_json::json!({\"sessions\": false})\n+            )).unwrap();\n+            assert!(read_frame(&mut conn).is_err(), \"failed query closes old carrier\");\n+            let mut conn = listener.accept().unwrap();\n+            accepted += 1;\n+            recv_hello(&mut conn, Role::Brain).unwrap();\n+            assert_eq!(read_frame(&mut conn).unwrap().kind, crate::msg::KIND_SESSIONS);\n+            write_frame(&mut conn, &Envelope::new(\n+                crate::msg::KIND_SESSIONS_REPLY, serde_json::json!({\"sessions\": []})\n+            )).unwrap();\n+            accepted\n+        });\n+        let mut poll = SessionPoll::default();\n+        for _ in 0..N {\n+            assert_eq!(poll.query(&name), Some(BTreeSet::new()));\n+        }\n+        assert_eq!(poll.query(&name), None, \"failure is not an empty session census\");\n+        assert_eq!(poll.query(&name), Some(BTreeSet::new()));\n+        drop(poll);\n+        assert_eq!(server.join().unwrap(), 2, \"one healthy carrier plus one recovery\");\n+    }\n+\n     // A no-op Psyche summarizer command: spawns + exits 0 (we never depend on its\n     // output — the daemon online-stamp is what hosts it, not the pid).\n     #[cfg(windows)]\ndiff --git a/docs/RCA-FLEET-DAEMON-14444.md b/docs/RCA-FLEET-DAEMON-14444.md\nindex f7e897d1..010613fd 100644\n--- a/docs/RCA-FLEET-DAEMON-14444.md\n+++ b/docs/RCA-FLEET-DAEMON-14444.md\n@@ -69,3 +69,76 @@ per 2 MB of log, under `[REQ-CONN-POISON-ATTRIBUTION]`, is its own finding.\n    (image path or start time), which is the only way reuse becomes a kill. My own standing rule from\n    the orphan reap is exactly this: re-verify the path AT kill time.\n 3. **Bound the log volume** independently of the cause.\n+\n+## releases#286 — stderr follow-up (2026-09-09)\n+\n+<!-- [doc->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED] -->\n+\n+The historical 303 MB incident above remains the reported baseline, not a\n+newly reproduced event. A read-only sample of the currently installed\n+`%LOCALAPPDATA%/spt-core/logs/daemon.stderr.log` ended at byte **4,163,267**:\n+the last **1,999,958 complete-line bytes** spanned wall stamps\n+**1788997126331–1788998151874** (1025.543 s). It contained **5127** first writes\n+(656,158 bytes), **5092** unpoisoned/non-timeout closes (850,622 bytes),\n+**3** organic retirements, **0** deadline poisons, **107** family-gate lines\n+and **108** rendezvous-up lines. Healthy start/close records alone consumed\n+**75.3%** of that sample. A subsequent live append observation over\n+**307.203 s** added **79,741 bytes** (**259.57 B/s**), with **104** starts,\n+**104** healthy closes and **30** family gates paired with **30** rendezvous\n+binds. These are different observation windows, not a claimed constant rate.\n+No daemon was restarted or reconfigured for either measurement.\n+\n+**Cause discrimination.** `role=brain` identifies *every broker IPC client*,\n+not the supervised coordinator. `run_brain` already keeps its heartbeat\n+`Brain` alive. In contrast, the pre-fix\n+`livehost::query_live_session_endpoints` constructed `Brain::cold_start`,\n+called `sessions`, and dropped the carrier on every invocation from the\n+five-second reconcile loop. This is a source-confirmed avoidable reopen\n+cause, consistent with the recurring five-second bare-brain start/close\n+pairs in the live log, **not an attribution of all 5127 connections to that\n+caller**. Other short-lived callers include dispatch idle queries and\n+per-operation dispatch/CLI clients; their lifetimes are not changed here.\n+\n+The livehost census now owns one carrier across polls. Its `BRAIN_CONN_OPEN`\n+breadcrumb names `caller=livehost-session-query` and distinguishes\n+`reason=initial-query` from `reason=previous-query-error`. Every query error\n+discards the carrier, returns an unavailable census (never an empty\n+authoritative census), and permits a fresh connection only on the next\n+scheduled query. The one-shot boot orphan census remains one-shot.\n+No read/write deadline, poison semantics, or retry cadence changed.\n+\n+**Family gate exonerated.** The sole diagnostic site in\n+`NetEndpoint::bind_scoped` is already once per real bind attempt, not per IPC\n+connection. `pairhost::spawn_meet_rotation` binds a new derived identity for\n+each attached subnet at every 30-second TOTP boundary; the observed\n+`BIGNET`, `SPT_DEV`, and `SPT_MANTLE` windows account for the three lines per\n+boundary. These listeners cannot share a stable identity without changing\n+the pairing protocol. Bind and diagnostic policies are intentionally\n+unchanged; a process-global \"log once\" would hide later real binds.\n+\n+**Explicit contract amendment.** Issue #286 supersedes only the\n+`REQ-CONN-POISON-ATTRIBUTION` leg-(d) requirement for unconditional\n+*healthy start/close lines*. Across the broker process these become\n+`CONN_LIFECYCLE_SUMMARY` totals, at most once per 60-second monotonic interval,\n+flushed by the first subsequent lifecycle counter event. Starts count the\n+first committed write attempt; closes count unpoisoned physical drops.\n+There is no timer thread and no shutdown flush, so a quiet tail remains\n+pending until another counter event (and can be lost on process exit).\n+Poison, organic failure, poisoned close, and all role-transition records\n+remain unconditional per connection. They retain conn id, role facts,\n+wall/monotonic stamps **and the original first-write wall/monotonic time**;\n+a connection that never wrote says `first_write=none`.\n+\n+The deterministic regression cells exercise real framed socket exchanges\n+over both retained and fresh carriers with an isolated fixed log clock/sink,\n+assert one interval summary for N healthy connections, preserve exceptional\n+records through the same sink, and drive livehost census reuse plus recovery\n+from a malformed reply. Focused commands (not executed in the preparation\n+lane; integration owner runs them after landing):\n+\n+```text\n+cargo test -p spt-daemon --lib conn::tests::\n+cargo test -p spt-daemon --lib livehost::tests::session_poll_reuses_healthy_carrier_and_reopens_after_query_error\n+cargo test -p spt-daemon --test conn_blackhole_lifecycle\n+traceable-reqs check --json\n+```\ndiff --git a/traceable-reqs.toml b/traceable-reqs.toml\nindex 21917ea9..9224b3aa 100644\n--- a/traceable-reqs.toml\n+++ b/traceable-reqs.toml\n@@ -2569,6 +2569,14 @@ id = \"REQ-CONN-POISON-ATTRIBUTION\"\n title = \"MSG-IDENTITY W6 / F-039 legs b-d (doyle W6 LOCK 2026-07-10, minted per amendment 3): every broker-conn lifecycle record is ATTRIBUTABLE — the W6 RCA's terminal undecidability (per-line 1:1 CONN_WRITE_POISONED churn = fresh-carrier churn OR stderr interleave artifact) exists because records carry no stable conn identity, no role/endpoint/session context, and no timestamps, and the once-per-conn poison latch hides multiplicity. THREE LEGS. (b) IDENTITY: mint a stable per-physical-conn id (monotonic u64 at conn construction — Arc::ptr_eq is the only identity today and it does not survive a log line) plus subscriber role and endpoint/session where known, stamped on CONN_WRITE_POISONED, CONN_WRITE_RETIRED, logical stall-evict, attach/resume/detach, and write-retirement records (RCA attach sites: presence nethost.rs:379, stream nethost.rs:258, controller broker.rs:891, viewer broker.rs:1073). (c) TIME: daemon stderr correlation records carry wall-clock AND monotonic timestamps (stderrlog has neither; broker+brain share one file — interleave is unresolvable without them). (d) LIFECYCLE (doyle-confirmed UNCONDITIONAL, not debug-gated): one BOUNDED set of per-conn lifecycle events — write start/timeout-cancel/transport close/writer exit/replacement-reattach (hertz RCA fix-shape items 1-3). Constraint (doyle LOCK): the split/attribution must not REDUCE total information, only correct its attribution; NO timeout-value changes; NO suppression-as-fix. Gate: unit — lifecycle records carry conn id + role + timestamps; the id is unique per physical conn and stable across that conn's records. Kin REQ-CONN-POISON-DIAL-SCOPE (leg a, the token split these fields ride on), REQ-CONN-BLACKHOLE-LIFECYCLE-HARNESS (leg e, consumes these records), REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK (behavior invariant preserved).\"\n required_stages = [\"impl\", \"unit\"]  # FLIPPED in the W6 build commit carrying the evidence (todlando 2026-07-10). impl = conn id mint + describe/label accretion + log_stamp wall+mono (anchored at both daemon entries) + bounded lifecycle events (write-start, transport-close, writer-exit, attach/replace/detach) + attribution on stall-evict records. unit = conn.rs record-shape assertions (id/role/time present + parity across both retirement tokens, id unique per conn, label bounded).\n \n+# releases#286 narrowly supersedes leg (d)'s unconditional healthy start/close\n+# lines. Failure/poison, role transitions and original first-write time remain\n+# attributable per physical conn; healthy counts are interval totals instead.\n+[[requirements]]\n+id = \"REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED\"\n+title = \"Healthy broker IPC starts and closes aggregate once per monotonic interval across connections, independently of request volume; poison, failure and role-transition evidence remains per-connection with original first-write time. Livehost session polling reuses its healthy carrier and names why it opens a new one; role=brain alone never identifies the supervised coordinator. Network family diagnostics remain once per real endpoint bind attempt, not IPC churn. Narrow healthy-start/close supersession of REQ-CONN-POISON-ATTRIBUTION leg d (releases#286).\"\n+required_stages = [\"doc\", \"impl\", \"unit\"]\n+\n [[requirements]]\n id = \"REQ-CONN-BLACKHOLE-LIFECYCLE-HARNESS\"\n title = \"MSG-IDENTITY W6 / F-039 leg e (doyle W6 LOCK 2026-07-10, minted per amendment 3 — hertz's five invariants VERBATIM from his RCA fix-shape item 5): 'Build a deterministic black-holed-controller harness against current v0.30.6 semantics and assert: unrelated sessions continue; the bad physical connection is canceled/closed within the bound; its writer exits; a fresh viewer can attach; no lock or task remains owned by the retired connection.' The harness is the standing conformance rig for the r4 SHAREDSEND fix-class — hertz's RCA discipline: only after a timestamped incident maps to a FAILING lifecycle invariant does an ownership/cancellation defect get fixed (the likely shape being complete physical-connection cancellation and writer-task join/retirement, never a broader timeout increase). Consumes REQ-CONN-POISON-ATTRIBUTION's records (conn id + lifecycle events are what make the five assertions checkable deterministically). Kin REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK (the invariant class under test — its brain_decouple int stays the Windows-mandatory gate leg), REQ-CONN-POISON-DIAL-SCOPE.\"\n@@ -4860,6 +4868,7 @@ name = \"message-format-identity\"\n requirements = [\n   \"REQ-CONN-BLACKHOLE-LIFECYCLE-HARNESS\",\n   \"REQ-CONN-POISON-ATTRIBUTION\",\n+  \"REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED\",\n   \"REQ-CONN-POISON-DIAL-SCOPE\",\n   \"REQ-DAEMON-SERVICE-INSTALL\",\n   \"REQ-EP-3\",","startLine":300,"lineNumbers":[300,301,302,303,304,305,306,307,308,309,310,311,312,313,314,315,316,317,318,319,320,321,322,323,324,325,326,327,328,329,330,331,332,333,334,335,336,337,338,339,340,341,342,343,344,345,346,347,348,349,350,351,352,353,354,355,356,357,358,359,360,361,362,363,364,365,366,367,368,369,370,371,372,373,374,375,376,377,378,379,380,381,382,383,384,385,386,387,388,389,390,391,392,393,394,395,396,397,398,399,400,401,402,403,404,405,406,407,408,409,410,411,412,413,414,415,416,417,418,419,420,421,422,423,424,425,426,427,428,429,430,431,432,433,434,435,436,437,438,439,440,441,442,443,444,445,446,447,448,449,450,451,452,453,454,455,456,457,458,459,460,461,462,463,464,465,466,467,468,469,470,471,472,473,474,475,476,477,478,479,480,481,482,483,484,485,486,487,488,489,490,491,492,493,494,495,496,497,498,499,500,501,502,503,504,505,506,507,508,509,510,511,512,513,514,515,516,517,518,519,520,521,522,523,524,525,526,527,528,529,530,531,532,533,534,535,536,537,538,539,540,541,542,543,544,545,546,547,548,549,550,551,552,553,554,555,556,557,558,559,560,561,562,563,564,565,566,567,568,569,570,571,572,573,574,575,576,577,578,579,580,581,582,583,584,585,586,587,588,589,590,591,592,593,594,595,596,597,598,599,600,601,602,603,604,605,606,607,608,609,610,611,612,613,614,615,616,617,618,619,620,621,622,623,624,625,626,627,628,629,630,631,632,633,634,635,636,637,638,639,640,641,642,643,644,645,646,647,648,649,650,651,652,653,654,655,656,657,658]},"meta":{"source":{"type":"internal","value":"artifact://62"}}},"isError":false,"timestamp":1788999513830}}
{"type":"message","id":"9e606103","parentId":"177013f5","timestamp":"2026-09-10T00:18:33.861Z","message":{"role":"toolResult","toolCallId":"call_kmSQJbsJUh9pi2Nrmlkh27VR|fc_04ba494197989c86016aa1f759aa9487d1b2ed60d4784dcc55","toolName":"grep","content":[{"type":"text","text":"# .worktrees/286-churn/\n## PRD.md\n 9|\n*10|spt-core is a **harness-independent core for an agent ecosystem**: inter-agent messaging, live-agent lifecycle, terminal hosting, seamless self-update, and zero-config cross-machine networking — shipped as both a Rust library workspace and a single canonical binary (`spt` / `spt.exe`).\n 11|\n 12|It is the ground-up rebuild of `claude_skill_owl` (\"modern SPT\"), untethering the system from Claude Code so that **any** agent runtime — Claude Code, Codex, Cursor, headless, or novel surfaces — can participate by either shelling out to the binary or linking the crates. The first thing built atop it is a rebuilt `spt` plugin reaching parity with today's system; the networking, terminal-wrapper, and multi-instance capabilities then make an agent reachable and operable across all of a user's machines.\n 13|\n...\n 16|- Today's modern SPT works but is **fragile** — each release introduces 2–5 regressions because it grew rapidly from a narrow starting point and carries the resulting tech debt (ADR-0001).\n*17|- It is **welded to Claude Code**. Other harnesses can't reuse the messaging/lifecycle substrate.\n 18|- Networking, terminal hosting (\"capsule\"), and cross-machine operation were planned piecemeal atop a CC-tied core; they deserve to be first-class in a clean foundation.\n 19|\n 20|## 3. Goals\n...\n 54|- **R-ARCH-3** `spt-proto` carries an explicit **wire-protocol version** independent of crate semver, with a documented **N-1 compat window**. Binary and library consumers interoperate by proto version, never by release version.\n*55|- **R-ARCH-4** Copy-verbatim the commodity layer from the sister project (envelope grammar, spool DDL, registry RPC, `info.json` shape); clean-room only the architecture/lifecycle (ADR-0001).\n 56|\n*57|### 7.2 The daemon (broker/brain)\n 58|- **R-DAEMON-1** Exactly **one logical `spt-daemon` per machine**, always-on, owning PTYs, network identity + WAN endpoint, subnet registry, all spools, all poll-listener logic, and all Psyche/pulse loops. No separate listener or wrapper processes (ADR-0004).\n*59|- **R-DAEMON-2** Internally split into a stable **broker** (holds PTY masters, harness child processes, listening sockets; versioned IPC; rarely updates) and an **updatable brain** (all logic; rehydrates from disk + re-attaches to broker handles).\n 60|- **R-DAEMON-3** Any `spt api` invocation that needs the daemon **auto-starts it** if absent.\n 61|- **R-DAEMON-4** Honor every invariant in [`docs/KNOWN-HAZARDS.md`](./docs/KNOWN-HAZARDS.md) as a conformance checklist.\n 62|\n...\n 155|- **R-UPD-2** **All binaries signature-verified before handoff**, regardless of source (spt-core release key).\n*156|- **R-UPD-3** **No-endpoint-termination invariant, scoped by update class (ADR-0004 §A).** For the **brain-only** class — the routine case — the guarantee is **absolute**: no endpoint process terminates or suspends, satisfied by the broker/brain split (the broker keeps held PTYs/children/sockets alive across brain replacement). **broker-compatible** swaps the broker behind the versioned IPC with endpoints still surviving. **broker-breaking** (held-resource-type / OS-API change) MAY suspend endpoints...\n 157|- **R-UPD-4** **Not auto by default** — gated on user confirmation delivered to the most-recently-active live session, with an opt-in full-auto.\n 158|- **R-UPD-5** spt-core conducts **adapter ripple-updates** after self-update, per each adapter's update declaration and `min_spt_core_version` gate.\n 159|\n...\n 176|- **R-INSTALL-3** First run idempotent + interactive-optional: generate node identity, start daemon; pairing + subnet-naming are separate explicit steps.\n*177|- **R-INSTALL-4** (CONTEXT \"adapter registration\") **Adapter registration lifecycle:** `spt adapter add <path>` (or `--github <user/repo>`, manifest-first then install via the declared `[update]` avenue — *install is the first update*) validates the manifest and records it under `{SPT_HOME}/…/adapters/` (copy for `file_pull`, pointer for `delegated`); one command/dir for `kind=\"harness\"` and `kind=\"shell\"`. Node-local (drive/launch capability, not subnet advertisement). `spt adapter remove` is **soft...\n 178|- **R-MIGRATE-1** Standalone install **auto-detects a `claude_skill_owl` install and offers migration** (identity, agents, tracked Psyche context). First-class supported path.\n 179|\n 180|---\n...\n 195|\n*196|**In v1:** harness-independent core (library + binary); manifest + `api` harness contract; both startup topologies; one-daemon broker/brain; LiveAgent/ReadyAgent/Psyche/Worker/SptNode endpoints + Shell-vs-agent seam + typed/binary payloads; multi-instance data model + subnet registry + resolution policy + dormant/suspended + two-tier context sync + remote-drive of running instances + off-node file transfer; first-class WAN networking + TOTP-SPAKE2 pairing + subnet naming; cross-node Psyche sync (replaces...\n 197|\n 198|**Deferred (seam-compatible):** see [`docs/DEFERRED.md`](./docs/DEFERRED.md) — concrete Shells, PresenceChannel impl, presence gossip, instantiate-anywhere + consent, remote command exec + security gate, sidecar adapters, manifest includes, scrollback spillover, predictive/state-sync terminal layer, OS-level presence signal, macOS.\n 199|\n...\n 201|\n*202|1. A non-CC harness (or a headless spike) drives the full live-agent lifecycle through manifest + `api` alone — proving harness independence.\n 203|2. The rebuilt spt-plugin reaches **feature parity** with frozen `claude_skill_owl`.\n 204|3. Two of a user's machines, freshly paired via one TOTP code, exchange messages and the user drives an agent on machine A from machine B — zero config beyond pairing.\n 205|4. A self-update rolls across the subnet **without terminating any endpoint**.\n## ROADMAP.md\n 6|\n*7|A harness-independent core for an agent ecosystem — messaging, live-agent lifecycle, terminal hosting, **first-class P2P networking**, and a runtime-manifest harness contract — shipped as Rust crates + a single `spt` binary. The proof: the same agent runs natively across all of a user's machines with a synced mind, reachable and operable from anywhere, with seamless self-update and zero-config cross-machine messaging. First thing built atop it: a rebuilt `spt` plugin at parity with today's modern SPT.\n 8|\n 9|## Scope decisions (locked)\n 10|\n...\n 28|- **M1 — Local messaging + binary.** `spt-msg` (delivery TCP+spool, routing, send/ring/ready) + the `spt` binary + CLI. **= the killer quickstart (local): two agents exchange a message.** Task breakdown: [`M1-PLAN.md`](./M1-PLAN.md).\n*29|- **M2 — Harness contract + lifecycle.** **Strictly harness-agnostic — zero Claude Code conventions in this repo.** spt-core defines the adapter contract; the only adapter-shaped artifact in-tree is a generic mock/test adapter (manifest fixture + fake `api` caller) that exercises the contract (PRD R-DOCS-2 dev-agent quickstart). **The rebuilt `spt` plugin (Claude Code) is a SEPARATE downstream project** built atop these crates/binary — spt-core v1's acceptance proof (PRD §\"feature parity\"), not an...\n 30|  - **M2a — Harness contract. ✅ delivered (2026-06-01).** `spt-runtime` (AgentRuntime, ManifestRuntime, manifest schema) + the `api` subcommand surface (seed/listen/bind/state/echo-gate/poll/worker-*/boundary/session-end/presence/history-log/emit/capability) + the **harness-hosted** startup topology (seed→listen→bind, interim seed-file) + local-api-auth + the generic mock adapter + contract E2E. All M2a reqs activated + green. Task breakdown: [`M2a-PLAN.md`](./M2a-PLAN.md).\n*31|  - **M2b — Live-agent lifecycle. ✅ delivered (2026-06-01).** `spt-live` (fifth layer): history subsystem (fetcher/locate-normalize/native) + spawn-psyche (nested perch, spt-core-owned `$psyche_prompt`) + echo-commune (history brief → single-writer commune drop) + drop-file ingest (single-writer 6.4 + direct-write precedence 6.5) + interim pulse + signoff/boundary ordering invariants (echo-before-signoff 3.3, grace-before-signoff 1.1, stale-signoff sweep 3.2) + resume seam (continue-existing / fres...\n 32|  - **Interim, no daemon** (like M1): the `api listen` process holds the perch + relays; the seed bridges via a file. **The spt-hosted topology (PTY-launch) + the consolidated `spt-daemon` move to M3** (they need the `spt-term` broker) — this narrows the earlier \"both topologies in M2\" wording.\n*33|- **M3 — Terminal wrapper + daemon + self-update.** `spt-term` (session-surface, PTY, broker) + the consolidated `spt-daemon` (broker/brain split) + peer-propagated signed self-update. Split M3a/M3b/M3c. Scope + carried-forward Spike #1 gaps: [`M3-PLAN.md`](./M3-PLAN.md).\n 34|  - **M3a — `spt-term` session-surface mechanism. ✅ delivered (2026-06-02).** The sixth crate (`…→spt-msg→spt-term`, sibling of spt-runtime): the `SessionSurface` trait + native PTY backends (ConPTY/forkpty via `portable-pty`, REQ-TERM-1) · ConPTY-DSR auto-answer drain pump (REQ-HAZARD-CONPTY-DSR) · send-keys/send-line injection (REQ-TERM-2) · bounded byte-stream (REQ-TERM-3) · the PTY-digest parser primitive (REQ-TERM-4 impl/unit, ADR-0008). OS-neutral per Spike #4/#5 (no OS's stream contr...\n*35|  - **M3b — `spt-daemon` broker/brain process. ✅ delivered (2026-06-03).** The seventh crate (`…→spt-live→spt-daemon→spt`) hosts many `spt-term` surfaces behind a **versioned local IPC** (named-pipe/UDS behind one `DaemonTransport` trait, forward-compatible frame schema): a **broker** (stable kernel — PTY masters, harness children, sockets; no logic) + a **brain** (restartable logic) so a brain kill/restart leaves the hosted PTY child + its output stream **intact and gapless** (Spike #1 mad...\n*36|  - **M3c — signed self-update. ✅ delivered (2026-06-03).** The gated, signed, ripple-capable self-update engine (no new crate — it lives in `spt-daemon`, signing primitives in `release.rs`): the **update-class taxonomy** (brain-only / broker-compatible / broker-breaking, ADR-0004 §A) with the **brain-only zero-interruption swap** that drives the M3b handoff substrate (snapshot→drop→re-attach) so a live PTY child + its output stream survive the logic swap **untouched** (REQ-UPD-3, proven E2E `...\n*37|  - **M3 COMPLETE** (M3a + M3b + M3c all ✅, 2026-06-03): the real ADR-0004 architecture is live — native PTY hosting, the consolidated broker/brain daemon, and seamless signed self-update. Next milestone: M4.\n 38|- **M4 — Networking + instances. ✅ delivered (2026-06-04).** The eighth crate `spt-net` (Iroh WAN endpoint bound to the node's own Ed25519 identity, mDNS LAN discovery, TOTP-SPAKE2 pairing over a dedicated pre-trust ALPN with transcript binding / rate limiting / seed rotation+transfer, subnet registry + NDJSON replication under the per-node epoch lease) · the multi-instance model (registry rows + bare-id resolution refuse-and-qualify + visibility/sync-membership gates + rename/collision + **resting-...\n*39|- **M5 — Shells, presence, deferred capabilities. ✅ delivered (2026-06-04;** scope locked with user 2026-06-04, see [`M5-PLAN.md`](./M5-PLAN.md) + per-task JIT plans D0–D9). The consent framework seam (grant store + interactive escalation + pre-consent flags; remote-exec / instantiate-anywhere reserved-but-refusing) · adapter registration lifecycle (`spt adapter add/remove`, manifest-first) · **shell hosting machinery + sleep/wake** (link-token channels with per-frame MAC, broker-launched binarie...\n 40|\n 41|## Post-v1-core milestones\n 42|\n...\n 46|\n*47|- **✅ Broker/brain process-isolation restoration — delivered (2026-06-11).** Design: [`docs/BROKER-BRAIN-SPLIT-RESTORATION.md`](./docs/BROKER-BRAIN-SPLIT-RESTORATION.md) (verified-with-amendments by `doyle`); decision: **ADR-0018**; build plan: [`RESTORATION-PLAN.md`](./RESTORATION-PLAN.md) (D1–D7). Corrects an unintended regression — the broker/brain split (ADR-0004) was specced + spiked as two *processes*, but the production daemon runs the broker as an in-process *thread* (`daemon.rs:165-170`)...\n 48|\n 49|- **🔜 NEXT milestone (planned 2026-06-12) — M9: Adapter customization & session surfaces.** Product of the two-day **Gateway grill** (operator + `doyle`; CONTEXT.md mints @ `8dc22ff`/`bd469b9`; ADR-0008 amended; the lecturn device design parked in the sibling `spt-lecturn/` repo for a later build). Shell-free, independently-shippable features: **adapter profiles** (sparse leaf-replace overlays on the parent manifest; shipped + local sources, local surviving adapter updates; composite `<adapter>:<pro...\n 50|- **🔜 M10 — Digest extractor & thread-spanning** (planned 2026-06-13; NEXT). Product of the digest grill (operator + `doyle`); governed by **ADR-0019** (reverses M9's \"no manifest seam\"), CONTEXT.md \"session digest\" rewritten. An `spt-claude-code` enabler, so it takes the M10 slot ahead of the Shell/remote-attach work (which shift to M11/M12). Adapter-declared **`[digest]` extractor seam** (imperative, no DSL; same source files as `[history]` by default + own-source escape; `api digest-entry` push f...\n...\n 57|\n*58|1. **Broker/brain no-terminate handoff** — prove a PTY child + open socket survive a daemon-*logic* restart (FD-passing or stable-broker shim). The hardest invariant in the design (ADR-0004).\n 59|2. **Iroh + mDNS smoke test** — two nodes, separate NATs, one message end-to-end; measure direct-vs-relay + binary size (research brief §5). Validates the whole networking premise.\n 60|3. **ConPTY + `portable-pty`** — host and drive a real session in a broker PTY on Windows + Linux.\n*61|4. **Manifest-driven lifecycle** — a manifest drives one real Claude Code session through spawn → bind → commune → signoff, proving harness-independence works in practice.\n 62|\n 63|## Immediate next steps\n 64|\n...\n 66|- ✅ **Stage A — codex adversarial pass** — done. 14 findings (5 FATAL, 9 SERIOUS) in `docs/reviews/STAGE-A-codex-redteam.md`.\n*67|- ✅ **Spike #1 (broker/brain handoff)** — PASS on Windows ConPTY. `docs/spikes/SPIKE-01-broker-handoff.md`. Found+fixed the ConPTY DSR-stall hazard. 4 open gaps carried to M3.\n 68|- ✅ **Spike #2 (Iroh smoke test)** — PASS (single-host). `docs/spikes/SPIKE-02-iroh-smoke.md`. iroh 0.98 API workable; 13.57 MB release. Cross-NAT/relay proof deferred to M4.\n 69|- ✅ **Findings folded** into ADR-0002/0003/0004/0005 + KNOWN-HAZARDS (+3) + traceable-reqs.toml (+4 REQ-HAZARD).\n 70|\n...\n 72|1. **Stage A — `plan-eng-review`** (interactive architecture pass; second half of Stage A).\n*73|2. **Spike #4 (manifest-driven lifecycle)** — a manifest drives one real session spawn→bind→commune→signoff. (Spike #3 ConPTY hosting is largely covered by Spike #1.)\n 74|3. Resolve the open user decisions logged in the ADR amendments (context-merge model, seed-rotation, R-UPD-3 rewording).\n 75|4. **M0** (workspace + `spt-proto` + `spt-store`) — **started**; task breakdown in [`M0-PLAN.md`](./M0-PLAN.md). Each milestone gets its own just-in-time plan (this one is M0); the roadmap stays the milestone sequence, the JIT plan is the task layer.\n 76|\n## CONTEXT.md\n 6|\n*7|Harness-independent core for the SPT ecosystem. Provides inter-agent messaging, live-agent lifecycle, terminal wrapping, self-update, and networking primitives — as both a Rust library workspace and a canonical reference binary. Designed so any agent runtime (Claude Code, Codex, Cursor, headless, future harnesses) can interface with the SPT ecosystem either by shelling out to the binary or by linking the crates directly.\n 8|\n 9|Successor to `claude_skill_owl` (today's \"modern SPT\"), which is being rebuilt as `spt-core` to untether the system from Claude Code and lift it to a general-purpose agent-ecosystem core.\n 10|\n...\n 22|**spt plugin** (separate downstream project — NOT an spt-core deliverable):\n*23|A rebuilt version of today's Claude Code `spt` plugin. It is the **first consumer** built *atop* spt-core and the **acceptance proof** of spt-core v1 (it reaches feature parity with modern SPT while delegating all core functionality to spt-core, primarily via `spt.exe`, with deeper hooks where useful) — but it **lives and builds in its own repository, outside spt-core**. It is a Claude-Code-specific *adapter*: it holds the Claude Code conventions (hooks, slash-commands, skill/plugin layout, `claude` se...\n 24|\n 25|**Pi** (disambiguation — two meanings, never conflate):\n 26|(1) **Pi, the coding agent/harness** (`badlogic/pi-mono`) — a harness example alongside Claude Code and Codex; this is the meaning in user-facing harness lists. (2) **Pi-class node** — Raspberry-Pi-class low-power hardware hosting a Shell-only or headless SPT node; an incidental hardware descriptor, never an explicit product example. Public-facing docs must disambiguate or avoid the bare word.\n...\n 35|Internally the logical daemon is split into two implementation layers for seamless self-update (see Self-update):\n*36|- **broker** (stable \"kernel\") — holds *only* the un-transferable, must-not-die resources: PTY master fds, the spawned harness child processes, and listening network sockets. Minimal, dumb, versioned local IPC. Almost never updates.\n*37|- **daemon brain** (\"userspace\") — all logic (routing, registry, pulse/psyche loops, manifest parsing, update orchestration). Restarts freely on update; rehydrates from disk state and re-attaches to the broker's held handles.\n 38|\n 39|Logical addressing is unchanged — still one per-machine `spt-daemon`; the broker is an internal layer, not separately addressable. There is exactly **one broker per machine** (per `SPT_HOME`) — *not* one per endpoint: a single broker holds every hosted endpoint's resources, and it is present whenever the daemon runs, even with zero endpoints online (the bare-daemon case). It is therefore the always-present per-machine layer, which is why the single-daemon lock + liveness anchor belong to it.\n 40|\n 41|**in-session relay**:\n*42|A thin, stateless `spt.exe` task that exists only in **harness-hosted** sessions (where the agent harness is the parent process and spt cannot reach into its process tree — today's Monitor model). It streams the daemon brain's events into the session's stdout. All *stateful* listener logic lives in the daemon; the relay is a dumb pipe, freely killable and respawnable. **spt-hosted** sessions need no separate harness-owned relay — the daemon owns the PTY and consumes the same poll feed itself. Idle de...\n 43|\n 44|### Deliverable shape\n 45|\n...\n 109|<!-- [doc->REQ-ADAPTER-GH-TRANSPORT] -->\n*110|A `.spt` may be **multi-platform**: shared `manifest.toml` + `strings/` at the root, role binaries under per-target-triple subdirectories (`x86_64-pc-windows-msvc/`, …); install/update extracts the shared root plus only the current node's triple, flattened into `install_dir`, so flat `<install_dir>/<program>` resolution is unchanged. It stays one signed asset (`adapter.spt`, plain-tar or gzip); a multi-platform archive missing the recipient's triple is a typed `NoArtifactForPlatform`. Large adapters ma...\n 111|<!-- [doc->REQ-ADAPTER-UPDATE-MESSAGE] -->\n 112|An optional **`[update].message`** (avenue-agnostic) is a plain multi-line operator notice surfaced to stdout, markdown-rendered (the helpfmt prose path), **only when an update is actually applied** (the version changed) — never on a no-op. It is read from the newly-installed manifest with no `{key}` substitution; its use is to announce a post-update action (e.g. \"run `/reload-plugins` in any ongoing sessions\").\n 113|\n 114|**composite update — `[update.post]`** (ratified 2026-06-25, v0.16.0 update-arc grill; ADR-0029):\n*115|An optional **avenue-agnostic** post-step `{ command, self_verifies }` spt-core runs **after** the primary avenue resolves — in the same `spt adapter update` **and at `spt adapter add`** (install is the first update, so a fresh install conducts the post-step too; bug-#1 operator ruling, v0.19.0 — the eager-extract acquisition runs it post-registration, a delegated acquisition after the acquisition succeeds, and only the payload-less `file_pull` PENDING add defers it to the payload's arrival) — so o...\n 116|\n 117|**resident adapter binary**: an adapter-owned process spt-core keeps alive for an endpoint's lifetime (today the `[message-idle-translation-binary]`), as opposed to **ephemeral** adapter binaries — the Psyche loop (daemon-hosted, ADR-0004), the `[digest]` extractor, `[session.*]` runners, hooks — which spawn on demand and pick up an update on their next invocation. Only resident binaries are stopped/restarted on a live update; ephemerals self-heal.\n 118|_Avoid_: calling the Psyche loop or an on-demand extractor a \"resident\" binary; \"restart the endpoint\" for what is a per-binary cycle.\n...\n 124|\n*125|Governing principle: **SPT is not a harness.** Model choice, billing shape, harness-internal env, and harness-internal context are entirely the adapter's concern, expressed inside the adapter's own command templates. spt-core owns only the template *mechanism* (substitution keys), the substitution *values* it is responsible for, and the surrounding lifecycle. Env for the *endpoint binary itself* is auto-handled by spt-core/broker; env for the *agent running inside* that binary is the adapter's config (e....\n 126|\n 127|**spawn-session seam** — launch a new agent session on this node. Manifest provides: a command template; `cwd`/project; a `headless` flag (optional, default false — for the GUI's resume-of-compatible-adapters); a `resume` flag (optional); and the `commune` + `signoff` file directories relative to `cwd` (so the daemon knows where to watch). Substitution keys spt-core can supply: `{id}` and, optionally, a spt-core-generated valid session UUID (e.g. injected as `--session-id {uuid}`) so an adapter can s...\n 128|- **id resolution:** `id` is optional. With no id, spt-core reproduces today's no-id `/spt:live` behavior — run the lone live agent if that's all the project has; show a picker with proposed default IDs if the project has none; let the user choose if there are several.\n...\n 175|- **Command templates are opaque.** spt-core never parses out a model/tool/flag — the adapter writes the whole command line; spt-core fills substitution keys and runs it.\n*176|- **A command template's program token resolves against the adapter install dir before PATH (since v0.8.0).** A `.spt` adapter ships its built binaries to the adapter's install dir (`adapters/_github/<safe>/` via `--release`/`--github`, or the record's `source_dir` under copy-mode), so a bare program name (e.g. `claude-spt-digest …`) binds to the shipped binary first and falls back to PATH when absent — a `.spt` that ships its binaries is **self-contained**, needing no PATH placement. <!-- [doc->REQ-...\n 177|- **Hook output capability is declared per harness-event** (`can_inject`). CC's Stop hook cannot inject context — that single fact drives the echo-gate sentinel + relay fallback. The manifest expresses it so spt-core knows when to fall back.\n 178|- **Env injection is asymmetric** (file-bridge-only-when-not-launcher, applied to env): spt-hosted sessions inherit env from the broker that spawns them; harness-hosted sessions need the harness's declared env channel. With `spt` on PATH the env table is small.\n 179|- **Cross-adapter fallback** is a **node-wide setting**, not a manifest field: if a Psyche/echo invocation under one adapter is rate-limited, spt-core falls back to another adapter (e.g. `ccs` — its own adapter, not a binary-swap). <!-- [doc->REQ-MANIFEST-6] --> A fallback **target is addressed as `<adapter>:<profile>`** (not just a bare adapter_name) and resolves through the one composite-addressing resolver (`registry::resolve_option`), so a fallback may select a shipped or local profile (`ccs`, `ccs...\n...\n 249|**Psyche**:\n*250|The Psyche companion's own perch, distinct from its paired LiveAgent's perch. First-class endpoint type so messages addressed to a LiveAgent's Psyche route directly without ambiguity. **A Psyche is a bounded per-event turn, not a resident process (since v0.25.0).** Each psyche-relevant event (a pulse fire, a commune/signoff drop, a session-custody transition) runs **exactly one** bounded turn through the psyche role template, spawned by the daemon, which exits at turn end — there is no long-lived psych...\n 251|\n 252|*I/O & trust boundary (ADR-0012):* the Psyche is a **sandboxed** actor — it may read and write files but **cannot send messages or reach the network itself**. Its inbound context arrives two ways: events/messages the daemon hands it, and **commune/signoff file-drops** (Self → daemon → Psyche; the *Summarizer* authors the commune delta). Its **sole outbound** is **reply/notify intents** the daemon relays as its **outbound proxy** — emitted as `<EVENT type=\"reply\">`/`<EVENT type=\"notify\">` (the sha...\n 253|\n*254|*Psyche-host health — harness-reachable failure signal (v0.8.1, REQ-HAZARD-LIVEHOST-BOOT-RACE):* a LiveAgent's `status=online` is daemon-authoritative liveness and **stays authoritative** — but it does not by itself prove the daemon hosted a Psyche. When the brain's live-host reconcile fails to spawn the Psyche (e.g. the adapter's psyche binary is absent from its install dir, or the net-less boot-race starves the host), that failure was previously **silent** — only an `eprintln!` on the brain's inv...\n 255|\n 256|**Summarizer**:\n 257|The ephemeral, cheap model that builds a **commune delta** from a Self's recent turns and feeds it *into* the **Psyche** as inbound context. A distinct actor from the Psyche — different (cheaper) model, fire-and-forget, **no perch** (not an endpoint type). It authors *commune* deltas only, **never** *reply*/*notify*.\n...\n 293|**ResidentService** (concept ratified 2026-07-26, ADR-0049 — the supervised substrate; first consumer spt-alchemy's Hub Daemon):\n*294|A **daemon-supervised binary an adapter declares, with no perch, no identity, no address.** Core owns the process from birth: the daemon spawns it **job-neutrally** (never a shell's child — a shell's tree-kill and a launching terminal's Job Object cannot reach it) and supervises it with the wake-watcher scaffolding (backoff, give-up latch, one-per-instance lock, orphan-kill, brain-side reconcile), running **independent of any agent's liveness**. Declared by the adapter manifest's **`[service]`** sectio...\n 295|_Avoid_: calling it an endpoint (no identity, no address — that is the [[AlwaysOnEndpoint]] layered on top); calling it a Shell (owner-less, not driven); \"daemon\" unqualified (the node has one daemon; this is a supervised service under it); treating its supervision record as liveness truth (there is none — liveness is the child handle).\n 296|\n 297|**AlwaysOnEndpoint** (always-on endpoint; concept ratified 2026-06-21, re-based on the substrate 2026-07-26 per ADR-0049 — core kind, first instance downstream `spt-discord`):\n...\n 311|\n*312|**resume custody**: The record a wake-resume spawn leaves (`resume.pid`) so lifecycle readers defer to a revival in flight rather than normalizing over it. Custody is an **identity — the (pid, process-creation-time) pair — never a bare PID** (ADR-0047): a recycled pid must read **not ours**, and the reader that discovers a mismatched pair deletes the record and proceeds (self-heal). Bind and reap clear custody atomically with their own outcome. _Avoid_: `is_process_alive(pid)` as a custody test (the ...\n 313|\n 314|**operator-stop inhibit**: The durable machine-scoped marker `daemon stop` records *before* teardown, which every implicit daemon-ensure path (REQ-DAEMON-3's anchor, as amended by ADR-0047) consults and honors by **declining to spawn** with one line naming the remedy. Cleared only by intent verbs (`daemon start`, the update paths that restart by design) — never by TTL. _Avoid_: reading the anchor as unconditional (that was the respawn-convoy bug — KNOWN-HAZARDS 7.52), or a TTL \"safety\" clear (a surpr...\n 315|\n...\n 332|\n*333|**Shells differ structurally from agent endpoints** (full treatment below): a Shell has a node-local perch but **no `tracked/` context** (no mind to sync); its logs are node-local; it is **adapter/platform-bound, not adapter-agnostic**; and its lifecycle is link/teardown, not the dormant/suspended resting model of agent instances.\n 334|\n 335|#### Shell model (detailed)\n 336|\n...\n 347|<!-- [doc->REQ-SHELL-4] shell tunnel: a long-lived reliable-ordered link-bound QUIC stream pair carrying opaque bytes the taxonomy never reinterprets; manifest opt-in, not enveloped/MAC-framed/spooled; link-break closes it; reliable-ordered ⇒ on-LAN posture -->\n*348|Channels carry typed, taxonomy-interpreted payloads. Distinct from them, an owner↔shell link may also hold a **shell tunnel**: a long-lived, **reliable, ordered** byte stream (a dedicated QUIC stream pair bound to the link) for protocol traffic the channel taxonomy must NOT reinterpret — opaque wire protocols spoken end-to-end (first consumer: USB/IP URB traffic to a usbip shell). Not spooled, not enveloped; the link's lifecycle governs it — a link-break closes the tunnel. Reliable-ordered is the p...\n 349|\n 350|<!-- [doc->REQ-SHELL-5] owner-type-agnostic: control-exclusivity keys on the owner endpoint_id, never the owner's endpoint type -->\n 351|**Owner-linked, exclusive.** Spawned by an **owning endpoint**; both linked; **only that owner `endpoint_id` may control it**. Agent endpoints are the common owner, but ownership is NOT agent-exclusive — any non-Shell endpoint type may own shells (ratified 2026-06-11, the Gateway grill: e.g. a **Gateway** owns the driven surfaces it steers; a future **Resource** endpoint could too). Control-exclusivity does not mean interaction-exclusivity: a Shell stays 2-way with its *environment* (user messages on t...\n...\n 354|\n*355|**Lifecycle = online / offline / torn-down.** A **link-break always closes the broker + shell binary** (with an optional manifest-declared **pre-close instruction** to the binary + a **termination timeout** for graceful shutdown). Then: an **ephemeral** shell (a manifest property, not an agent choice) is fully torn down **and removed from the agent's shell history**; a **persistent** shell keeps its perch **offline** (re-linkable — the binary is re-spawned on relink). **The binary never survives a link...\n 356|\n 357|**Broadcast policy** (manifest): the shell advertises its availability to all agents on the subnet, only same-node agents, or not at all (agents learn of it only by user instruction). Shells run only on nodes where the shell adapter is installed/advertised.\n 358|\n...\n 360|\n*361|**`shell spawn` mints a new instance, it is not the online switch.** `spt shell spawn <adapter> [--id]` creates a *new* Shell identity (`<adapter>-<n>` + its perch + the owner link). Bringing an *existing* offline (persistent) instance back online is **`shell relink <id>`** (re-spawns the binary), or happens automatically via `persistent` (auto-online with the owner) / `wake_command` (offline wake-watcher). `shell teardown <id>` destroys it. So **`spawn`:create :: `relink`/`persistent`/`wake`:online** ...\n 362|\n 363|**per-shell instantiation approval (`require_approval`, manifest enum):** gates `shell spawn`, reusing the consent plumbing (grant store + interactive escalation; see Consent & security gates).\n 364|- `none` (**default** — matches the system's everything-opt-in posture) — the agent spawns freely within scope.\n...\n 411|<!-- [doc->REQ-RCVIEW-1] [doc->REQ-KICK-1] [doc->REQ-VIEWER-SKIP-TO-LIVE-ON-EVICT] [doc->REQ-HAZARD-VIEWER-RING-ROLL-SNAP] -->\n*412|**BUILT (M12 W2.5).** The controller/viewer model is implemented end-to-end. Attach intent is **three-valued** (`AttachIntent = Viewer | Control | Take`, wire-default `Control`): `Control` to a FREE endpoint becomes controller; `Control` to a CONTROLLED endpoint is **refused with guidance** (`--view` to watch, `--take` to control) — never auto-viewer, never silent-displace; `Take` (`spt rc --take` / picker \"Kick\") kicks the incumbent with a **loud `Displaced{by}` notice** and full detach (not demote). ...\n 413|\n*414|**redispatch** (REDISPATCH-TRUTH, ADR-0038): the fresh brain's reconstruction of target-side attach workers over broker-held peer streams after a **brain cycle** (`spt daemon refresh` / `spt update` apply — the PTYs and QUIC streams survive; the brain-owned serve workers do not). Redispatch eligibility is **lifecycle-gated**: finished/terminal stream rows are retired from the dispatcher's enumeration (a fresh dispatcher never re-serves a terminal Attach — the KNOWN-HAZARDS 7.41 frozen-PTY steal/clear...\n*415|_Avoid_: conflating redispatch with msg dispatch or session-cursor resume; filtering replay by origin identity (the legitimate same-`by` successor re-take must keep working — lifecycle is the discriminator).\n 416|\n 417|<!-- [doc->REQ-RC-KEY-VT-TRANSLATE] -->\n 418|**rc keyboard input (Windows VT translation, v0.13.0 bug 2).** On **Windows** an interactive `spt rc` console reads crossterm **key events** and translates each to **standard xterm VT** (`translate_key_event`) — arrows / Home / End / PgUp/Dn / Insert / Delete / F-keys + modifiers all reach the harness as the universal terminal contract (**agnostic**, NOT win32-input-mode; the legacy console delivers those keys as events, not bytes, so the old byte-pump left them DEAD). **Unix passes through** (its raw-...\n\n## docs/\n### KNOWN-HAZARDS.md\n 4|\n*5|**Architecture-translation note.** The sister project runs poll listeners and Psyche wrappers as *separate processes*. spt-core consolidates both into the one `spt-daemon` (brain), with a stable broker beneath it (ADR-0004). Many hazards below were inter-process races in the sister project; in spt-core some become intra-daemon concerns (potentially easier) while others move to the daemon↔broker IPC boundary or the network boundary (potentially new failure surface). Each entry notes the mapping where it...\n 6|\n 7|---\n 8|\n...\n 34|\n*35|### 1.5 Worker (working-perch) lifecycle path consistency\n 36|- **Failure:** subagent-start creates the perch at one path layout; later hooks read it at another → not found; stop-hook scan misses nested perches.\n 37|- **Invariant:** all Worker/Psyche child-perch path composition routes through one central resolver; no divergent path construction across hooks.\n 38|- **spt-core mapping:** `Worker` is a day-one endpoint type; the daemon owns the registry, so perch location is a registry lookup, not ad-hoc path math. Single source of truth for instance→location.\n...\n 59|- **Invariant:** every newly-added handoff positional has a default; state-file rehydration happens *after* argv parse; defaults survive intermediate versions.\n*60|- **spt-core mapping:** CRITICAL — self-update (ADR-0004) makes handoff routine. The broker↔brain IPC and any brain-relaunch argv must be versioned and forward/backward tolerant (a newer brain talks to an older broker). This is the single most update-frequency-sensitive invariant.\n*61|- **Sister cite:** `src/live/wrapper/lifecycle.rs:17-106`; `src/cli.rs` defaults; CHANGELOG v1.11.10.\n 62|\n 63|### 2.4 Generation `gen_start` always = now() on cold-start AND handoff\n 64|- **Failure:** stale gen_start from a rehydrated state file fires time-based discriminators on the new process.\n...\n 66|- **spt-core mapping:** carries to the daemon's per-instance generation tracking.\n*67|- **Restoration D3/D4 (ADR-0018):** the generation *counter* custody moved to the broker (D3-2 — it observes every brain spawn, planned or crash, and hands `{generation, start-reason}` at spawn; `gen_start` stays `now()`-fresh, never rehydrated). The brain→brain **`BrainState` *message*** (`{session_id, generation, next_seq, gen_start_ms}`) that previously carried continuity across a handoff is **retired from the production path in D4-2**: a brain the supervisor respawns cold-starts and reconstructs ...\n*68|- **Sister cite:** `src/live/wrapper/lifecycle.rs:70`; `src/common/wrapper_state.rs`.\n 69|\n 70|### 2.5 Daemon-hosted endpoints have no dedicated liveness PID\n 71|- **Failure:** the sister evaluates Psyche/perch liveness via a dedicated process PID — the wrapper's own pid in `info.json`, checked with `is_process_alive`. Under ADR-0004 the Psyche (and any spt-hosted Self) is a **loop inside the daemon**, not a separate process: it holds no dedicated pid, and its `claude`/summarizer subprocess is ephemeral (spawned per pulse/commune, then exits). If a daemon-hosted perch's `info.json` carries the **daemon's** pid, then *every* hosted endpoint shares one pid, and `...\n...\n 95|\n*96|## 3. Lifecycle\n 97|\n 98|### 3.1 Ephemeral perch cleanup on every `ring` exit path\n 99|- **Failure:** `ring` creates an ephemeral perch; early-exit paths (no-perch, empty-msg, timeout) skip cleanup → stale dirs accumulate.\n 100|- **Invariant:** every code path that creates an ephemeral perch cleans it before exit; exception: if the caller already had an active perch, do not treat as ephemeral and do not clean up.\n*101|- **spt-core mapping:** `ring` semantics carry; the daemon owns ephemeral-perch lifecycle, so a single guaranteed-cleanup (drop guard / RAII) is achievable in-process.\n 102|- **Sister cite:** `src/owl/ring.rs:58-294`.\n 103|\n 104|### 3.2 Stale signoff sentinel must not kill a fresh start\n...\n 196|- **Failure:** the registry's only superseding mechanism is the per-`(endpoint_id, node)` epoch lease (4.8) — a row is replaced only by a newer row *from the same node*. When a node identity dies permanently (machine retired, or `node.key` regenerated so the \"node\" never speaks again), its rows are never superseded and never expire: they sit in the in-memory registries and the `identity/registry/<subnet>.json` snapshots forever. A bare-id send then resolves the same endpoint id on both the live and the...\n*197|- **Invariant:** registry rows authored by a **silent** peer node decay: a node not *heard* (admitted inbound feed — the M7 D2 heard-map, REQ-SUBNET-1) within the eviction window (`registry_evict_after_ms`, default 300s ≈ 10 default pump cadences) has its rows **evicted** from every subnet registry, snapshots rewritten. Own rows never decay (the node always hears itself implicitly — it authors them each pump tick). Eviction is safe under the lease: v1 has **no transitive gossip**, so any future upd...\n 198|- **spt-core mapping:** `spt_net::net::registry::SubnetRegistry::evict_nodes` (model) + `spt_daemon::registryhost::RegistryHost::evict_silent_peers` (heard-map TTL) driven from the registry pump tick (`peerloop`). Trust rows are NOT auto-evicted (trust is a user decision; a stale trust row only costs dead dials) — pruning those is a separate verb.\n 199|- **Source:** M7 acceptance run 2026-06-06 (DEFERRED.md \"Ghost registry row eviction\"); the AMBIGUOUS render fix rode along.\n 200|- **Mesh note (ADR-0017, 2026-06-08):** the subnet mesh **preserves** this invariant rather than superseding it. \"No transitive gossip\" sharpens to **no transitive *row* gossip** — the mesh relays only the member *roster* (discovery), while registry **rows stay own-authored and are fetched directly** from each member over a handshake. So \"any future update for a node comes from that node itself, alive\" still holds and the eviction lease is untouched. (The plan's rejected alternative — signed transiti...\n...\n 227|- **Failure:** a hung subprocess (git on slow net) blocks the supervisor indefinitely.\n*228|- **Invariant:** every metadata-producing subprocess has a timeout; timeout yields `None` + rate-limited stderr, never a hang.\n 229|- **spt-core mapping:** generalize to all manifest-declared harness invocations (delegated commands, adapter updates) — timeouts mandatory.\n 230|- **Sister cite:** `src/common/git.rs`.\n 231|\n...\n 406|\n*407|### 6.7 Broker and brain MUST be separate processes (in-process collapse silently breaks no-endpoint-drop update) `[REQ-HAZARD-BROKER-PROCESS-ISOLATION]`\n*408|- **Failure:** the daemon hosts the broker as a background *thread* in the single `spt daemon` process (`daemon.rs:165-170`, `Arc<Broker>` + `thread::spawn(serve)`) instead of a separate process. A brain restart onto a swapped binary then cannot happen without killing the broker thread — closing every PTY, orphaning every harness child, dropping every socket. So `spt update apply` degrades to an in-process `Brain::handoff` no-op: the binary swaps on disk but the running daemon keeps executing the old c...\n*409|- **Invariant:** the broker runs as its own long-lived process that survives every brain restart; the brain restarts onto the new binary and re-attaches via the versioned IPC. A routine (brain-only) update must leave every hosted endpoint untouched at the *process* level — not merely re-subscribe a brain within the same process. The evidence for REQ-UPD-3 / REQ-DAEMON-2 must prove process-level survival (a PTY child + a live QUIC conn survive a brain-process restart onto a swapped binary — SPIKE-01/0...\n*410|- **spt-core mapping:** restoration is ADR-0018 (next milestone). The current `int` tags on REQ-DAEMON-2 / REQ-UPD-3 are regression-masked and re-point at restoration; the broker becomes the always-up per-machine anchor (seed-lock + liveness + brain supervisor). Two-process supervision, generation custody, durable-deadline loop timing, broker-cursor-of-record, and readiness-gated auto-rollback all hang off this.\n*411|- **Origin:** unintended spec/impl drift from ADR-0004 (the broker *process* was specced + spiked but built in-process), discovered during the v0.3.2 fleet update verify. Full audit + decisions: `docs/BROKER-BRAIN-SPLIT-RESTORATION.md` (verified) + ADR-0018.\n 412|<!-- [doc->REQ-HAZARD-BROKER-PROCESS-ISOLATION] -->\n*413|- **D1 (restoration skeleton, ADR-0018 Q2/Q3):** the process boundary is restored — `spt daemon run` is the broker process and spawns a supervised `spt daemon brain` child (`brainproc.rs`); the broker survives the brain dying and respawns it (proven in production topology by `crates/spt/tests/brain_split.rs`). The logic loops still run broker-side (D2 migrates them); the `int` process-level survival E2E + the in-process re-point land at D7.\n*414|- **Closed out (2026-06-11, v0.4.0–v0.4.2):** the two-process model shipped (v0.4.0); the D7 `int` E2E (`brain_survive.rs`) + the N-1 gate prove process-level survival onto swapped bytes and re-pointed REQ-DAEMON-2 / REQ-UPD-3. The v0.4.1 fleet-verify proved this Windows-seamless (hfenduleam: brain pid rolls, broker held, `exe_hash` flips, no manual bounce) but exposed a **Linux** respawn-path gap — the resident broker respawned the brain via per-spawn `current_exe()`, which on Linux follows the `app...\n 415|\n 416|### 6.8 No irreversible durable-state migration before update ready-promotion `[REQ-HAZARD-ROLLBACK-STATE-COMPAT]`\n*417|- **Failure:** the readiness-gated auto-rollback (ADR-0018 Q7) spawns the *previous* binary against durable state the *new* brain already wrote. The first release that migrates a durable-state schema in place would make the old binary unable to read it — silently bricking rollback exactly when it is needed (a logic-bricking update that can no longer fall back).\n*418|- **Invariant:** a brain must not irreversibly migrate durable state before it is ready-promoted; equivalently, every pre-ready write must remain readable by the N-1 brain. Schema migrations are gated behind ready-promotion (or written in an N-1-tolerant additive form).\n 419|- **spt-core mapping:** lands with ADR-0018's auto-rollback. Free to assert now (a 2026-06-09 source audit confirmed zero state-migration code exists); unmintable retroactively once a migration ships.\n 420|- **D5 conformance (2026-06-10):** the new durable timing state `<spt_home>/deadline-<key>.json` (restoration D5-1) is **additive** — a rolled-back pre-D5 binary does not know the file and simply ignores it (re-phasing on its own flat-sleep cadence, the pre-D5 behavior). No existing-file schema migration, no irreversible pre-ready write → the new file is rollback-N-1-safe by construction. The thing for a future D6 guard to gate is a *migration* of this file's shape, not its introduction.\n*421|- **D6 guard (2026-06-10, restoration D6-3):** the invariant is now **asserted**, not just noted. The pre-ready durable writes are **enumerated in one place** — `spt-daemon::PRE_READY_DURABLE_FILES` (`rollback_compat.rs`): `deadline-<key>.json` (D5, `DeadlineAnchor`), `applied-state.json` (D6-1, the two-phase `AppliedRecord`), and the generation-stamped `brain.ready` breadcrumb (D6-1b, `{pid, generation}`). A **tripwire unit test** pins each one's additive / N-1-readable contract (load-bearing field na...\n### TRACEABILITY.md\n 20|\n*21|`doc` (prose/design/rustdoc) · `impl` (production code) · `unit` (unit test) · `int` (integration / e2e / cross-node). Default activation policy: `[\"doc\",\"impl\",\"unit\"]`; networking/lifecycle/cross-node reqs add `int`.\n 22|\n 23|### Tag examples\n 24|\n...\n 70|\n*71|## Lifecycle\n 72|\n 73|- **Now (planning):** manifest seeded, all reqs inactive. `check` is green (nothing required yet).\n 74|- **M0:** install the CLI, validate the seed against it, wire the CI workflow, **activate M0 reqs** (`REQ-ARCH-*`, `REQ-HAZARD-*` for the commodity layer), tag as code lands.\n### DOCS-STRATEGY.md\n 17|7. **Consistent conversational voice** — adopt the Google developer style guide: second person (\"you\"), active voice, knowledgeable-friend tone.\n*18|8. **Explain *why*, not just *what*** — conceptual docs + diagrams for the messaging model, the agent lifecycle states, instances/dormancy, and P2P topology.\n 19|9. **Stable, never-renamed anchors/URLs** — agents cache links.\n 20|10. **Docs-as-product, gated in CI** — generation (rustdoc, `llms.txt`, schema, CLI help) is part of the build so docs can never drift from code. Drift is the #1 most-cited docs failure; this kills it structurally.\n 21|\n...\n 25|\n*26|**Messaging · Live-agent lifecycle · Terminal hosting · Networking (P2P + pairing) · Harness contract (manifest + `api`) · Instances · Shells · Self-update.**\n 27|\n 28|## Killer quickstart targets (two, one per audience)\n 29|\n 30|- **Human dev:** two agents exchange a message end-to-end in **one command + minimal config**, < 10 min — the core primitive, whole-thing-first, then decomposed.\n*31|- **Dev-agent integrator:** a minimal **adapter** that satisfies the manifest + `api` contract and drives one lifecycle operation — the \"build a harness for spt-core\" hello-world.\n 32|\n 33|## Agent-consumable docs (spt-core specifics)\n 34|\n...\n 74|A harness-independent core for an agent ecosystem: messaging, live-agent\n*75|lifecycle, terminal hosting, P2P networking, a runtime-manifest harness contract.\n 76|\n 77|> Each capability links to its own llms.txt. Full concatenated docs: /llms-full.txt.\n 78|> Machine-readable contract: /manifest.schema.json. Append `.md` to any doc URL\n...\n 87|- [Messaging](…/messaging/llms.txt): typed ops + file blobs; send / ring / ready.\n*88|- [Live-agent lifecycle](…/lifecycle/llms.txt): skeleton → live → dormant → suspended.\n 89|- … Terminal · Networking · Harness contract · Instances · Shells · Self-update\n 90|\n 91|## Reference (auto-generated)\n...\n 122|**Vertical ordering (dependency-sequenced):**\n*123|`Messaging → Lifecycle → Terminal → Networking → Harness contract → Instances → Shells → Self-update`\n 124|(messaging is the primitive the quickstart needs; self-update last because its no-termination invariant depends on the terminal/broker model.)\n 125|**Per-vertical internal template (Django labels × Cloudflare ordering):**\n 126|`Overview (why+diagram) · Quickstart/Tutorial · How-to guides · Reference (rustdoc+CLI+schema) · llms.txt`\n...\n 133|//! # spt-runtime — one-sentence value prop.\n*134|//! ## A tour of spt-runtime   - [`Manifest`] · [`ApiCommand`] · [`Lifecycle`]\n 135|//! ## Example                 (runnable, real values, near the top)\n 136|//! ## Feature flags           `net` (default) · `unstable`\n 137|//! ## Wire-protocol compatibility   interop by proto version, never crate version.\n### STORAGE.md\n 43|    project-index.json           endpoint→project attribution (schema-versioned; daemon-written)\n*44|  daemon/                        ── ephemeral ── brain rehydration state, broker handle metadata\n 45|  updates/                       ── ephemeral ── staged binaries pending handoff\n 46|  adapters/                      registered adapter manifests / pointers\n 47|    <adapter>/web/               core-owned adapter served root; output bytes survive deactivation\n...\n 173|| subnet registry (snapshot + distribution) | **BranchStore** | audited; rides the precedence merge driver instead of a separate replication scheme (M4) |\n*174|| daemon/brain coarse checkpoint (rehydrate anchor, ADR-0004) | **BranchStore** | commit = checkpoint, ref tip = resume |\n*175|| endpoint lifecycle / status *audit trail* | **BranchStore** (or a derived view) | the *trajectory* matters, not just current value |\n 176|| spool (queue, drain-by-status, `id→address`) | **SQLite** | indexed queries + hot churn |\n 177|| broker effect journal (per-PTY-effect, B5) | **fsync-append** | hot path, per-keystroke — a commit per write is far too heavy |\n 178|| `info.json` / `status` (current value) | flat + atomic-rename (today) | tiny, read-hot; *may* move under BranchStore later, low priority |\n...\n 185|\n*186|**Reading the effect journal (`<spt_home>/daemon-effects.log`) — operator note.** The broker fsync-appends two line kinds per durable effect: `PENDING <key> <kind>` at attempt, `DONE <key>` on success (durable kinds only — an ephemeral `PtyWrite` writes no lines). A line with **`PENDING` and no matching `DONE`** is **expected**, not a leak: it is a durable effect that did NOT complete — for a `net-send` (a WAN write to an offline/unreachable peer) it is an **unconfirmed send retained as the re-driv...\n 187|\n 188|**Delivery vs propagation — latency envelope (operator note).** Two classes of cross-node traffic have deliberately different latency floors, so a minutes-scale delay is not automatically a defect: (1) **messages** (`spt send`) are **event-driven** — delivered on the events the daemon owns (WAN ingress + the ACTIVE→IDLE edge, F-023), so a healthy path is near-immediate and never waits on an adapter hook-poll cadence; (2) **anti-entropy context-sync** (the mind pull) is **periodic** — it rides the...\n 189|\n...\n 274|  one refresh. Fire-and-forget and daemon-offline-safe (events accumulate\n*275|  and are consumed at the next brain start).\n 276|  <!-- [doc->REQ-PROJECT-INDEX-INVALIDATION] -->\n 277|\n 278|## perches/ — node-local runtime\n...\n 312|\n\n[Showing lines 1-346 of 397 (49.7KB limit). Use :347 to continue. Some lines truncated to 512 chars]"}],"details":{"scopePath":".worktrees/286-churn/PRD.md, .worktrees/286-churn/ROADMAP.md, .worktrees/286-churn/CONTEXT.md, .worktrees/286-churn/docs/KNOWN-HAZARDS.md, .worktrees/286-churn/docs/TRACEABILITY.md, .worktrees/286-churn/docs/DOCS-STRATEGY.md, .worktrees/286-churn/docs/STORAGE.md, .worktrees/286-churn/docs/MANIFEST.md, .worktrees/286-churn/docs/CONTEXT-MEMORY.md","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":85,"fileCount":8,"files":[".worktrees/286-churn/PRD.md",".worktrees/286-churn/ROADMAP.md",".worktrees/286-churn/CONTEXT.md",".worktrees/286-churn/docs/KNOWN-HAZARDS.md",".worktrees/286-churn/docs/TRACEABILITY.md",".worktrees/286-churn/docs/DOCS-STRATEGY.md",".worktrees/286-churn/docs/STORAGE.md",".worktrees/286-churn/docs/MANIFEST.md"],"fileMatches":[{"path":".worktrees/286-churn/PRD.md","count":9},{"path":".worktrees/286-churn/ROADMAP.md","count":13},{"path":".worktrees/286-churn/CONTEXT.md","count":20},{"path":".worktrees/286-churn/docs/KNOWN-HAZARDS.md","count":20},{"path":".worktrees/286-churn/docs/TRACEABILITY.md","count":2},{"path":".worktrees/286-churn/docs/DOCS-STRATEGY.md","count":7},{"path":".worktrees/286-churn/docs/STORAGE.md","count":7},{"path":".worktrees/286-churn/docs/MANIFEST.md","count":7}],"truncated":true,"perFileLimitReached":20,"displayContent":"# .worktrees/286-churn/\n## PRD.md\n   9│\n *10│spt-core is a **harness-independent core for an agent ecosystem**: inter-agent messaging, live-agent lifecycle, terminal hosting, seamless self-update, and zero-config cross-machine networking — shipped as both a Rust library workspace and a single canonical binary (`spt` / `spt.exe`).\n  11│\n  12│It is the ground-up rebuild of `claude_skill_owl` (\"modern SPT\"), untethering the system from Claude Code so that **any** agent runtime — Claude Code, Codex, Cursor, headless, or novel surfaces — can participate by either shelling out to the binary or linking the crates. The first thing built atop it is a rebuilt `spt` plugin reaching parity with today's system; the networking, terminal-wrapper, and multi-instance capabilities then make an agent reachable and operable across all of a user's machines.\n  13│\n    │...\n  16│- Today's modern SPT works but is **fragile** — each release introduces 2–5 regressions because it grew rapidly from a narrow starting point and carries the resulting tech debt (ADR-0001).\n *17│- It is **welded to Claude Code**. Other harnesses can't reuse the messaging/lifecycle substrate.\n  18│- Networking, terminal hosting (\"capsule\"), and cross-machine operation were planned piecemeal atop a CC-tied core; they deserve to be first-class in a clean foundation.\n  19│\n  20│## 3. Goals\n    │...\n  54│- **R-ARCH-3** `spt-proto` carries an explicit **wire-protocol version** independent of crate semver, with a documented **N-1 compat window**. Binary and library consumers interoperate by proto version, never by release version.\n *55│- **R-ARCH-4** Copy-verbatim the commodity layer from the sister project (envelope grammar, spool DDL, registry RPC, `info.json` shape); clean-room only the architecture/lifecycle (ADR-0001).\n  56│\n *57│### 7.2 The daemon (broker/brain)\n  58│- **R-DAEMON-1** Exactly **one logical `spt-daemon` per machine**, always-on, owning PTYs, network identity + WAN endpoint, subnet registry, all spools, all poll-listener logic, and all Psyche/pulse loops. No separate listener or wrapper processes (ADR-0004).\n *59│- **R-DAEMON-2** Internally split into a stable **broker** (holds PTY masters, harness child processes, listening sockets; versioned IPC; rarely updates) and an **updatable brain** (all logic; rehydrates from disk + re-attaches to broker handles).\n  60│- **R-DAEMON-3** Any `spt api` invocation that needs the daemon **auto-starts it** if absent.\n  61│- **R-DAEMON-4** Honor every invariant in [`docs/KNOWN-HAZARDS.md`](./docs/KNOWN-HAZARDS.md) as a conformance checklist.\n  62│\n    │...\n 155│- **R-UPD-2** **All binaries signature-verified before handoff**, regardless of source (spt-core release key).\n*156│- **R-UPD-3** **No-endpoint-termination invariant, scoped by update class (ADR-0004 §A).** For the **brain-only** class — the routine case — the guarantee is **absolute**: no endpoint process terminates or suspends, satisfied by the broker/brain split (the broker keeps held PTYs/children/sockets alive across brain replacement). **broker-compatible** swaps the broker behind the versioned IPC with endpoints still surviving. **broker-breaking** (held-resource-type / OS-API change) MAY suspend endpoints...\n 157│- **R-UPD-4** **Not auto by default** — gated on user confirmation delivered to the most-recently-active live session, with an opt-in full-auto.\n 158│- **R-UPD-5** spt-core conducts **adapter ripple-updates** after self-update, per each adapter's update declaration and `min_spt_core_version` gate.\n 159│\n    │...\n 176│- **R-INSTALL-3** First run idempotent + interactive-optional: generate node identity, start daemon; pairing + subnet-naming are separate explicit steps.\n*177│- **R-INSTALL-4** (CONTEXT \"adapter registration\") **Adapter registration lifecycle:** `spt adapter add <path>` (or `--github <user/repo>`, manifest-first then install via the declared `[update]` avenue — *install is the first update*) validates the manifest and records it under `{SPT_HOME}/…/adapters/` (copy for `file_pull`, pointer for `delegated`); one command/dir for `kind=\"harness\"` and `kind=\"shell\"`. Node-local (drive/launch capability, not subnet advertisement). `spt adapter remove` is **soft...\n 178│- **R-MIGRATE-1** Standalone install **auto-detects a `claude_skill_owl` install and offers migration** (identity, agents, tracked Psyche context). First-class supported path.\n 179│\n 180│---\n    │...\n 195│\n*196│**In v1:** harness-independent core (library + binary); manifest + `api` harness contract; both startup topologies; one-daemon broker/brain; LiveAgent/ReadyAgent/Psyche/Worker/SptNode endpoints + Shell-vs-agent seam + typed/binary payloads; multi-instance data model + subnet registry + resolution policy + dormant/suspended + two-tier context sync + remote-drive of running instances + off-node file transfer; first-class WAN networking + TOTP-SPAKE2 pairing + subnet naming; cross-node Psyche sync (replaces...\n 197│\n 198│**Deferred (seam-compatible):** see [`docs/DEFERRED.md`](./docs/DEFERRED.md) — concrete Shells, PresenceChannel impl, presence gossip, instantiate-anywhere + consent, remote command exec + security gate, sidecar adapters, manifest includes, scrollback spillover, predictive/state-sync terminal layer, OS-level presence signal, macOS.\n 199│\n    │...\n 201│\n*202│1. A non-CC harness (or a headless spike) drives the full live-agent lifecycle through manifest + `api` alone — proving harness independence.\n 203│2. The rebuilt spt-plugin reaches **feature parity** with frozen `claude_skill_owl`.\n 204│3. Two of a user's machines, freshly paired via one TOTP code, exchange messages and the user drives an agent on machine A from machine B — zero config beyond pairing.\n 205│4. A self-update rolls across the subnet **without terminating any endpoint**.\n## ROADMAP.md\n  6│\n *7│A harness-independent core for an agent ecosystem — messaging, live-agent lifecycle, terminal hosting, **first-class P2P networking**, and a runtime-manifest harness contract — shipped as Rust crates + a single `spt` binary. The proof: the same agent runs natively across all of a user's machines with a synced mind, reachable and operable from anywhere, with seamless self-update and zero-config cross-machine messaging. First thing built atop it: a rebuilt `spt` plugin at parity with today's modern SPT.\n  8│\n  9│## Scope decisions (locked)\n 10│\n   │...\n 28│- **M1 — Local messaging + binary.** `spt-msg` (delivery TCP+spool, routing, send/ring/ready) + the `spt` binary + CLI. **= the killer quickstart (local): two agents exchange a message.** Task breakdown: [`M1-PLAN.md`](./M1-PLAN.md).\n*29│- **M2 — Harness contract + lifecycle.** **Strictly harness-agnostic — zero Claude Code conventions in this repo.** spt-core defines the adapter contract; the only adapter-shaped artifact in-tree is a generic mock/test adapter (manifest fixture + fake `api` caller) that exercises the contract (PRD R-DOCS-2 dev-agent quickstart). **The rebuilt `spt` plugin (Claude Code) is a SEPARATE downstream project** built atop these crates/binary — spt-core v1's acceptance proof (PRD §\"feature parity\"), not an...\n 30│  - **M2a — Harness contract. ✅ delivered (2026-06-01).** `spt-runtime` (AgentRuntime, ManifestRuntime, manifest schema) + the `api` subcommand surface (seed/listen/bind/state/echo-gate/poll/worker-*/boundary/session-end/presence/history-log/emit/capability) + the **harness-hosted** startup topology (seed→listen→bind, interim seed-file) + local-api-auth + the generic mock adapter + contract E2E. All M2a reqs activated + green. Task breakdown: [`M2a-PLAN.md`](./M2a-PLAN.md).\n*31│  - **M2b — Live-agent lifecycle. ✅ delivered (2026-06-01).** `spt-live` (fifth layer): history subsystem (fetcher/locate-normalize/native) + spawn-psyche (nested perch, spt-core-owned `$psyche_prompt`) + echo-commune (history brief → single-writer commune drop) + drop-file ingest (single-writer 6.4 + direct-write precedence 6.5) + interim pulse + signoff/boundary ordering invariants (echo-before-signoff 3.3, grace-before-signoff 1.1, stale-signoff sweep 3.2) + resume seam (continue-existing / fres...\n 32│  - **Interim, no daemon** (like M1): the `api listen` process holds the perch + relays; the seed bridges via a file. **The spt-hosted topology (PTY-launch) + the consolidated `spt-daemon` move to M3** (they need the `spt-term` broker) — this narrows the earlier \"both topologies in M2\" wording.\n*33│- **M3 — Terminal wrapper + daemon + self-update.** `spt-term` (session-surface, PTY, broker) + the consolidated `spt-daemon` (broker/brain split) + peer-propagated signed self-update. Split M3a/M3b/M3c. Scope + carried-forward Spike #1 gaps: [`M3-PLAN.md`](./M3-PLAN.md).\n 34│  - **M3a — `spt-term` session-surface mechanism. ✅ delivered (2026-06-02).** The sixth crate (`…→spt-msg→spt-term`, sibling of spt-runtime): the `SessionSurface` trait + native PTY backends (ConPTY/forkpty via `portable-pty`, REQ-TERM-1) · ConPTY-DSR auto-answer drain pump (REQ-HAZARD-CONPTY-DSR) · send-keys/send-line injection (REQ-TERM-2) · bounded byte-stream (REQ-TERM-3) · the PTY-digest parser primitive (REQ-TERM-4 impl/unit, ADR-0008). OS-neutral per Spike #4/#5 (no OS's stream contr...\n*35│  - **M3b — `spt-daemon` broker/brain process. ✅ delivered (2026-06-03).** The seventh crate (`…→spt-live→spt-daemon→spt`) hosts many `spt-term` surfaces behind a **versioned local IPC** (named-pipe/UDS behind one `DaemonTransport` trait, forward-compatible frame schema): a **broker** (stable kernel — PTY masters, harness children, sockets; no logic) + a **brain** (restartable logic) so a brain kill/restart leaves the hosted PTY child + its output stream **intact and gapless** (Spike #1 mad...\n*36│  - **M3c — signed self-update. ✅ delivered (2026-06-03).** The gated, signed, ripple-capable self-update engine (no new crate — it lives in `spt-daemon`, signing primitives in `release.rs`): the **update-class taxonomy** (brain-only / broker-compatible / broker-breaking, ADR-0004 §A) with the **brain-only zero-interruption swap** that drives the M3b handoff substrate (snapshot→drop→re-attach) so a live PTY child + its output stream survive the logic swap **untouched** (REQ-UPD-3, proven E2E `...\n*37│  - **M3 COMPLETE** (M3a + M3b + M3c all ✅, 2026-06-03): the real ADR-0004 architecture is live — native PTY hosting, the consolidated broker/brain daemon, and seamless signed self-update. Next milestone: M4.\n 38│- **M4 — Networking + instances. ✅ delivered (2026-06-04).** The eighth crate `spt-net` (Iroh WAN endpoint bound to the node's own Ed25519 identity, mDNS LAN discovery, TOTP-SPAKE2 pairing over a dedicated pre-trust ALPN with transcript binding / rate limiting / seed rotation+transfer, subnet registry + NDJSON replication under the per-node epoch lease) · the multi-instance model (registry rows + bare-id resolution refuse-and-qualify + visibility/sync-membership gates + rename/collision + **resting-...\n*39│- **M5 — Shells, presence, deferred capabilities. ✅ delivered (2026-06-04;** scope locked with user 2026-06-04, see [`M5-PLAN.md`](./M5-PLAN.md) + per-task JIT plans D0–D9). The consent framework seam (grant store + interactive escalation + pre-consent flags; remote-exec / instantiate-anywhere reserved-but-refusing) · adapter registration lifecycle (`spt adapter add/remove`, manifest-first) · **shell hosting machinery + sleep/wake** (link-token channels with per-frame MAC, broker-launched binarie...\n 40│\n 41│## Post-v1-core milestones\n 42│\n   │...\n 46│\n*47│- **✅ Broker/brain process-isolation restoration — delivered (2026-06-11).** Design: [`docs/BROKER-BRAIN-SPLIT-RESTORATION.md`](./docs/BROKER-BRAIN-SPLIT-RESTORATION.md) (verified-with-amendments by `doyle`); decision: **ADR-0018**; build plan: [`RESTORATION-PLAN.md`](./RESTORATION-PLAN.md) (D1–D7). Corrects an unintended regression — the broker/brain split (ADR-0004) was specced + spiked as two *processes*, but the production daemon runs the broker as an in-process *thread* (`daemon.rs:165-170`)...\n 48│\n 49│- **🔜 NEXT milestone (planned 2026-06-12) — M9: Adapter customization & session surfaces.** Product of the two-day **Gateway grill** (operator + `doyle`; CONTEXT.md mints @ `8dc22ff`/`bd469b9`; ADR-0008 amended; the lecturn device design parked in the sibling `spt-lecturn/` repo for a later build). Shell-free, independently-shippable features: **adapter profiles** (sparse leaf-replace overlays on the parent manifest; shipped + local sources, local surviving adapter updates; composite `<adapter>:<pro...\n 50│- **🔜 M10 — Digest extractor & thread-spanning** (planned 2026-06-13; NEXT). Product of the digest grill (operator + `doyle`); governed by **ADR-0019** (reverses M9's \"no manifest seam\"), CONTEXT.md \"session digest\" rewritten. An `spt-claude-code` enabler, so it takes the M10 slot ahead of the Shell/remote-attach work (which shift to M11/M12). Adapter-declared **`[digest]` extractor seam** (imperative, no DSL; same source files as `[history]` by default + own-source escape; `api digest-entry` push f...\n   │...\n 57│\n*58│1. **Broker/brain no-terminate handoff** — prove a PTY child + open socket survive a daemon-*logic* restart (FD-passing or stable-broker shim). The hardest invariant in the design (ADR-0004).\n 59│2. **Iroh + mDNS smoke test** — two nodes, separate NATs, one message end-to-end; measure direct-vs-relay + binary size (research brief §5). Validates the whole networking premise.\n 60│3. **ConPTY + `portable-pty`** — host and drive a real session in a broker PTY on Windows + Linux.\n*61│4. **Manifest-driven lifecycle** — a manifest drives one real Claude Code session through spawn → bind → commune → signoff, proving harness-independence works in practice.\n 62│\n 63│## Immediate next steps\n 64│\n   │...\n 66│- ✅ **Stage A — codex adversarial pass** — done. 14 findings (5 FATAL, 9 SERIOUS) in `docs/reviews/STAGE-A-codex-redteam.md`.\n*67│- ✅ **Spike #1 (broker/brain handoff)** — PASS on Windows ConPTY. `docs/spikes/SPIKE-01-broker-handoff.md`. Found+fixed the ConPTY DSR-stall hazard. 4 open gaps carried to M3.\n 68│- ✅ **Spike #2 (Iroh smoke test)** — PASS (single-host). `docs/spikes/SPIKE-02-iroh-smoke.md`. iroh 0.98 API workable; 13.57 MB release. Cross-NAT/relay proof deferred to M4.\n 69│- ✅ **Findings folded** into ADR-0002/0003/0004/0005 + KNOWN-HAZARDS (+3) + traceable-reqs.toml (+4 REQ-HAZARD).\n 70│\n   │...\n 72│1. **Stage A — `plan-eng-review`** (interactive architecture pass; second half of Stage A).\n*73│2. **Spike #4 (manifest-driven lifecycle)** — a manifest drives one real session spawn→bind→commune→signoff. (Spike #3 ConPTY hosting is largely covered by Spike #1.)\n 74│3. Resolve the open user decisions logged in the ADR amendments (context-merge model, seed-rotation, R-UPD-3 rewording).\n 75│4. **M0** (workspace + `spt-proto` + `spt-store`) — **started**; task breakdown in [`M0-PLAN.md`](./M0-PLAN.md). Each milestone gets its own just-in-time plan (this one is M0); the roadmap stays the milestone sequence, the JIT plan is the task layer.\n 76│\n## CONTEXT.md\n   6│\n  *7│Harness-independent core for the SPT ecosystem. Provides inter-agent messaging, live-agent lifecycle, terminal wrapping, self-update, and networking primitives — as both a Rust library workspace and a canonical reference binary. Designed so any agent runtime (Claude Code, Codex, Cursor, headless, future harnesses) can interface with the SPT ecosystem either by shelling out to the binary or by linking the crates directly.\n   8│\n   9│Successor to `claude_skill_owl` (today's \"modern SPT\"), which is being rebuilt as `spt-core` to untether the system from Claude Code and lift it to a general-purpose agent-ecosystem core.\n  10│\n    │...\n  22│**spt plugin** (separate downstream project — NOT an spt-core deliverable):\n *23│A rebuilt version of today's Claude Code `spt` plugin. It is the **first consumer** built *atop* spt-core and the **acceptance proof** of spt-core v1 (it reaches feature parity with modern SPT while delegating all core functionality to spt-core, primarily via `spt.exe`, with deeper hooks where useful) — but it **lives and builds in its own repository, outside spt-core**. It is a Claude-Code-specific *adapter*: it holds the Claude Code conventions (hooks, slash-commands, skill/plugin layout, `claude` se...\n  24│\n  25│**Pi** (disambiguation — two meanings, never conflate):\n  26│(1) **Pi, the coding agent/harness** (`badlogic/pi-mono`) — a harness example alongside Claude Code and Codex; this is the meaning in user-facing harness lists. (2) **Pi-class node** — Raspberry-Pi-class low-power hardware hosting a Shell-only or headless SPT node; an incidental hardware descriptor, never an explicit product example. Public-facing docs must disambiguate or avoid the bare word.\n    │...\n  35│Internally the logical daemon is split into two implementation layers for seamless self-update (see Self-update):\n *36│- **broker** (stable \"kernel\") — holds *only* the un-transferable, must-not-die resources: PTY master fds, the spawned harness child processes, and listening network sockets. Minimal, dumb, versioned local IPC. Almost never updates.\n *37│- **daemon brain** (\"userspace\") — all logic (routing, registry, pulse/psyche loops, manifest parsing, update orchestration). Restarts freely on update; rehydrates from disk state and re-attaches to the broker's held handles.\n  38│\n  39│Logical addressing is unchanged — still one per-machine `spt-daemon`; the broker is an internal layer, not separately addressable. There is exactly **one broker per machine** (per `SPT_HOME`) — *not* one per endpoint: a single broker holds every hosted endpoint's resources, and it is present whenever the daemon runs, even with zero endpoints online (the bare-daemon case). It is therefore the always-present per-machine layer, which is why the single-daemon lock + liveness anchor belong to it.\n  40│\n  41│**in-session relay**:\n *42│A thin, stateless `spt.exe` task that exists only in **harness-hosted** sessions (where the agent harness is the parent process and spt cannot reach into its process tree — today's Monitor model). It streams the daemon brain's events into the session's stdout. All *stateful* listener logic lives in the daemon; the relay is a dumb pipe, freely killable and respawnable. **spt-hosted** sessions need no separate harness-owned relay — the daemon owns the PTY and consumes the same poll feed itself. Idle de...\n  43│\n  44│### Deliverable shape\n  45│\n    │...\n 109│<!-- [doc->REQ-ADAPTER-GH-TRANSPORT] -->\n*110│A `.spt` may be **multi-platform**: shared `manifest.toml` + `strings/` at the root, role binaries under per-target-triple subdirectories (`x86_64-pc-windows-msvc/`, …); install/update extracts the shared root plus only the current node's triple, flattened into `install_dir`, so flat `<install_dir>/<program>` resolution is unchanged. It stays one signed asset (`adapter.spt`, plain-tar or gzip); a multi-platform archive missing the recipient's triple is a typed `NoArtifactForPlatform`. Large adapters ma...\n 111│<!-- [doc->REQ-ADAPTER-UPDATE-MESSAGE] -->\n 112│An optional **`[update].message`** (avenue-agnostic) is a plain multi-line operator notice surfaced to stdout, markdown-rendered (the helpfmt prose path), **only when an update is actually applied** (the version changed) — never on a no-op. It is read from the newly-installed manifest with no `{key}` substitution; its use is to announce a post-update action (e.g. \"run `/reload-plugins` in any ongoing sessions\").\n 113│\n 114│**composite update — `[update.post]`** (ratified 2026-06-25, v0.16.0 update-arc grill; ADR-0029):\n*115│An optional **avenue-agnostic** post-step `{ command, self_verifies }` spt-core runs **after** the primary avenue resolves — in the same `spt adapter update` **and at `spt adapter add`** (install is the first update, so a fresh install conducts the post-step too; bug-#1 operator ruling, v0.19.0 — the eager-extract acquisition runs it post-registration, a delegated acquisition after the acquisition succeeds, and only the payload-less `file_pull` PENDING add defers it to the payload's arrival) — so o...\n 116│\n 117│**resident adapter binary**: an adapter-owned process spt-core keeps alive for an endpoint's lifetime (today the `[message-idle-translation-binary]`), as opposed to **ephemeral** adapter binaries — the Psyche loop (daemon-hosted, ADR-0004), the `[digest]` extractor, `[session.*]` runners, hooks — which spawn on demand and pick up an update on their next invocation. Only resident binaries are stopped/restarted on a live update; ephemerals self-heal.\n 118│_Avoid_: calling the Psyche loop or an on-demand extractor a \"resident\" binary; \"restart the endpoint\" for what is a per-binary cycle.\n    │...\n 124│\n*125│Governing principle: **SPT is not a harness.** Model choice, billing shape, harness-internal env, and harness-internal context are entirely the adapter's concern, expressed inside the adapter's own command templates. spt-core owns only the template *mechanism* (substitution keys), the substitution *values* it is responsible for, and the surrounding lifecycle. Env for the *endpoint binary itself* is auto-handled by spt-core/broker; env for the *agent running inside* that binary is the adapter's config (e....\n 126│\n 127│**spawn-session seam** — launch a new agent session on this node. Manifest provides: a command template; `cwd`/project; a `headless` flag (optional, default false — for the GUI's resume-of-compatible-adapters); a `resume` flag (optional); and the `commune` + `signoff` file directories relative to `cwd` (so the daemon knows where to watch). Substitution keys spt-core can supply: `{id}` and, optionally, a spt-core-generated valid session UUID (e.g. injected as `--session-id {uuid}`) so an adapter can s...\n 128│- **id resolution:** `id` is optional. With no id, spt-core reproduces today's no-id `/spt:live` behavior — run the lone live agent if that's all the project has; show a picker with proposed default IDs if the project has none; let the user choose if there are several.\n    │...\n 175│- **Command templates are opaque.** spt-core never parses out a model/tool/flag — the adapter writes the whole command line; spt-core fills substitution keys and runs it.\n*176│- **A command template's program token resolves against the adapter install dir before PATH (since v0.8.0).** A `.spt` adapter ships its built binaries to the adapter's install dir (`adapters/_github/<safe>/` via `--release`/`--github`, or the record's `source_dir` under copy-mode), so a bare program name (e.g. `claude-spt-digest …`) binds to the shipped binary first and falls back to PATH when absent — a `.spt` that ships its binaries is **self-contained**, needing no PATH placement. <!-- [doc->REQ-...\n 177│- **Hook output capability is declared per harness-event** (`can_inject`). CC's Stop hook cannot inject context — that single fact drives the echo-gate sentinel + relay fallback. The manifest expresses it so spt-core knows when to fall back.\n 178│- **Env injection is asymmetric** (file-bridge-only-when-not-launcher, applied to env): spt-hosted sessions inherit env from the broker that spawns them; harness-hosted sessions need the harness's declared env channel. With `spt` on PATH the env table is small.\n 179│- **Cross-adapter fallback** is a **node-wide setting**, not a manifest field: if a Psyche/echo invocation under one adapter is rate-limited, spt-core falls back to another adapter (e.g. `ccs` — its own adapter, not a binary-swap). <!-- [doc->REQ-MANIFEST-6] --> A fallback **target is addressed as `<adapter>:<profile>`** (not just a bare adapter_name) and resolves through the one composite-addressing resolver (`registry::resolve_option`), so a fallback may select a shipped or local profile (`ccs`, `ccs...\n    │...\n 249│**Psyche**:\n*250│The Psyche companion's own perch, distinct from its paired LiveAgent's perch. First-class endpoint type so messages addressed to a LiveAgent's Psyche route directly without ambiguity. **A Psyche is a bounded per-event turn, not a resident process (since v0.25.0).** Each psyche-relevant event (a pulse fire, a commune/signoff drop, a session-custody transition) runs **exactly one** bounded turn through the psyche role template, spawned by the daemon, which exits at turn end — there is no long-lived psych...\n 251│\n 252│*I/O & trust boundary (ADR-0012):* the Psyche is a **sandboxed** actor — it may read and write files but **cannot send messages or reach the network itself**. Its inbound context arrives two ways: events/messages the daemon hands it, and **commune/signoff file-drops** (Self → daemon → Psyche; the *Summarizer* authors the commune delta). Its **sole outbound** is **reply/notify intents** the daemon relays as its **outbound proxy** — emitted as `<EVENT type=\"reply\">`/`<EVENT type=\"notify\">` (the sha...\n 253│\n*254│*Psyche-host health — harness-reachable failure signal (v0.8.1, REQ-HAZARD-LIVEHOST-BOOT-RACE):* a LiveAgent's `status=online` is daemon-authoritative liveness and **stays authoritative** — but it does not by itself prove the daemon hosted a Psyche. When the brain's live-host reconcile fails to spawn the Psyche (e.g. the adapter's psyche binary is absent from its install dir, or the net-less boot-race starves the host), that failure was previously **silent** — only an `eprintln!` on the brain's inv...\n 255│\n 256│**Summarizer**:\n 257│The ephemeral, cheap model that builds a **commune delta** from a Self's recent turns and feeds it *into* the **Psyche** as inbound context. A distinct actor from the Psyche — different (cheaper) model, fire-and-forget, **no perch** (not an endpoint type). It authors *commune* deltas only, **never** *reply*/*notify*.\n    │...\n 293│**ResidentService** (concept ratified 2026-07-26, ADR-0049 — the supervised substrate; first consumer spt-alchemy's Hub Daemon):\n*294│A **daemon-supervised binary an adapter declares, with no perch, no identity, no address.** Core owns the process from birth: the daemon spawns it **job-neutrally** (never a shell's child — a shell's tree-kill and a launching terminal's Job Object cannot reach it) and supervises it with the wake-watcher scaffolding (backoff, give-up latch, one-per-instance lock, orphan-kill, brain-side reconcile), running **independent of any agent's liveness**. Declared by the adapter manifest's **`[service]`** sectio...\n 295│_Avoid_: calling it an endpoint (no identity, no address — that is the [[AlwaysOnEndpoint]] layered on top); calling it a Shell (owner-less, not driven); \"daemon\" unqualified (the node has one daemon; this is a supervised service under it); treating its supervision record as liveness truth (there is none — liveness is the child handle).\n 296│\n 297│**AlwaysOnEndpoint** (always-on endpoint; concept ratified 2026-06-21, re-based on the substrate 2026-07-26 per ADR-0049 — core kind, first instance downstream `spt-discord`):\n    │...\n 311│\n*312│**resume custody**: The record a wake-resume spawn leaves (`resume.pid`) so lifecycle readers defer to a revival in flight rather than normalizing over it. Custody is an **identity — the (pid, process-creation-time) pair — never a bare PID** (ADR-0047): a recycled pid must read **not ours**, and the reader that discovers a mismatched pair deletes the record and proceeds (self-heal). Bind and reap clear custody atomically with their own outcome. _Avoid_: `is_process_alive(pid)` as a custody test (the ...\n 313│\n 314│**operator-stop inhibit**: The durable machine-scoped marker `daemon stop` records *before* teardown, which every implicit daemon-ensure path (REQ-DAEMON-3's anchor, as amended by ADR-0047) consults and honors by **declining to spawn** with one line naming the remedy. Cleared only by intent verbs (`daemon start`, the update paths that restart by design) — never by TTL. _Avoid_: reading the anchor as unconditional (that was the respawn-convoy bug — KNOWN-HAZARDS 7.52), or a TTL \"safety\" clear (a surpr...\n 315│\n    │...\n 332│\n*333│**Shells differ structurally from agent endpoints** (full treatment below): a Shell has a node-local perch but **no `tracked/` context** (no mind to sync); its logs are node-local; it is **adapter/platform-bound, not adapter-agnostic**; and its lifecycle is link/teardown, not the dormant/suspended resting model of agent instances.\n 334│\n 335│#### Shell model (detailed)\n 336│\n    │...\n 347│<!-- [doc->REQ-SHELL-4] shell tunnel: a long-lived reliable-ordered link-bound QUIC stream pair carrying opaque bytes the taxonomy never reinterprets; manifest opt-in, not enveloped/MAC-framed/spooled; link-break closes it; reliable-ordered ⇒ on-LAN posture -->\n*348│Channels carry typed, taxonomy-interpreted payloads. Distinct from them, an owner↔shell link may also hold a **shell tunnel**: a long-lived, **reliable, ordered** byte stream (a dedicated QUIC stream pair bound to the link) for protocol traffic the channel taxonomy must NOT reinterpret — opaque wire protocols spoken end-to-end (first consumer: USB/IP URB traffic to a usbip shell). Not spooled, not enveloped; the link's lifecycle governs it — a link-break closes the tunnel. Reliable-ordered is the p...\n 349│\n 350│<!-- [doc->REQ-SHELL-5] owner-type-agnostic: control-exclusivity keys on the owner endpoint_id, never the owner's endpoint type -->\n 351│**Owner-linked, exclusive.** Spawned by an **owning endpoint**; both linked; **only that owner `endpoint_id` may control it**. Agent endpoints are the common owner, but ownership is NOT agent-exclusive — any non-Shell endpoint type may own shells (ratified 2026-06-11, the Gateway grill: e.g. a **Gateway** owns the driven surfaces it steers; a future **Resource** endpoint could too). Control-exclusivity does not mean interaction-exclusivity: a Shell stays 2-way with its *environment* (user messages on t...\n    │...\n 354│\n*355│**Lifecycle = online / offline / torn-down.** A **link-break always closes the broker + shell binary** (with an optional manifest-declared **pre-close instruction** to the binary + a **termination timeout** for graceful shutdown). Then: an **ephemeral** shell (a manifest property, not an agent choice) is fully torn down **and removed from the agent's shell history**; a **persistent** shell keeps its perch **offline** (re-linkable — the binary is re-spawned on relink). **The binary never survives a link...\n 356│\n 357│**Broadcast policy** (manifest): the shell advertises its availability to all agents on the subnet, only same-node agents, or not at all (agents learn of it only by user instruction). Shells run only on nodes where the shell adapter is installed/advertised.\n 358│\n    │...\n 360│\n*361│**`shell spawn` mints a new instance, it is not the online switch.** `spt shell spawn <adapter> [--id]` creates a *new* Shell identity (`<adapter>-<n>` + its perch + the owner link). Bringing an *existing* offline (persistent) instance back online is **`shell relink <id>`** (re-spawns the binary), or happens automatically via `persistent` (auto-online with the owner) / `wake_command` (offline wake-watcher). `shell teardown <id>` destroys it. So **`spawn`:create :: `relink`/`persistent`/`wake`:online** ...\n 362│\n 363│**per-shell instantiation approval (`require_approval`, manifest enum):** gates `shell spawn`, reusing the consent plumbing (grant store + interactive escalation; see Consent & security gates).\n 364│- `none` (**default** — matches the system's everything-opt-in posture) — the agent spawns freely within scope.\n    │...\n 411│<!-- [doc->REQ-RCVIEW-1] [doc->REQ-KICK-1] [doc->REQ-VIEWER-SKIP-TO-LIVE-ON-EVICT] [doc->REQ-HAZARD-VIEWER-RING-ROLL-SNAP] -->\n*412│**BUILT (M12 W2.5).** The controller/viewer model is implemented end-to-end. Attach intent is **three-valued** (`AttachIntent = Viewer | Control | Take`, wire-default `Control`): `Control` to a FREE endpoint becomes controller; `Control` to a CONTROLLED endpoint is **refused with guidance** (`--view` to watch, `--take` to control) — never auto-viewer, never silent-displace; `Take` (`spt rc --take` / picker \"Kick\") kicks the incumbent with a **loud `Displaced{by}` notice** and full detach (not demote). ...\n 413│\n*414│**redispatch** (REDISPATCH-TRUTH, ADR-0038): the fresh brain's reconstruction of target-side attach workers over broker-held peer streams after a **brain cycle** (`spt daemon refresh` / `spt update` apply — the PTYs and QUIC streams survive; the brain-owned serve workers do not). Redispatch eligibility is **lifecycle-gated**: finished/terminal stream rows are retired from the dispatcher's enumeration (a fresh dispatcher never re-serves a terminal Attach — the KNOWN-HAZARDS 7.41 frozen-PTY steal/clear...\n*415│_Avoid_: conflating redispatch with msg dispatch or session-cursor resume; filtering replay by origin identity (the legitimate same-`by` successor re-take must keep working — lifecycle is the discriminator).\n 416│\n 417│<!-- [doc->REQ-RC-KEY-VT-TRANSLATE] -->\n 418│**rc keyboard input (Windows VT translation, v0.13.0 bug 2).** On **Windows** an interactive `spt rc` console reads crossterm **key events** and translates each to **standard xterm VT** (`translate_key_event`) — arrows / Home / End / PgUp/Dn / Insert / Delete / F-keys + modifiers all reach the harness as the universal terminal contract (**agnostic**, NOT win32-input-mode; the legacy console delivers those keys as events, not bytes, so the old byte-pump left them DEAD). **Unix passes through** (its raw-...\n\n## docs/\n### KNOWN-HAZARDS.md\n   4│\n  *5│**Architecture-translation note.** The sister project runs poll listeners and Psyche wrappers as *separate processes*. spt-core consolidates both into the one `spt-daemon` (brain), with a stable broker beneath it (ADR-0004). Many hazards below were inter-process races in the sister project; in spt-core some become intra-daemon concerns (potentially easier) while others move to the daemon↔broker IPC boundary or the network boundary (potentially new failure surface). Each entry notes the mapping where it...\n   6│\n   7│---\n   8│\n    │...\n  34│\n *35│### 1.5 Worker (working-perch) lifecycle path consistency\n  36│- **Failure:** subagent-start creates the perch at one path layout; later hooks read it at another → not found; stop-hook scan misses nested perches.\n  37│- **Invariant:** all Worker/Psyche child-perch path composition routes through one central resolver; no divergent path construction across hooks.\n  38│- **spt-core mapping:** `Worker` is a day-one endpoint type; the daemon owns the registry, so perch location is a registry lookup, not ad-hoc path math. Single source of truth for instance→location.\n    │...\n  59│- **Invariant:** every newly-added handoff positional has a default; state-file rehydration happens *after* argv parse; defaults survive intermediate versions.\n *60│- **spt-core mapping:** CRITICAL — self-update (ADR-0004) makes handoff routine. The broker↔brain IPC and any brain-relaunch argv must be versioned and forward/backward tolerant (a newer brain talks to an older broker). This is the single most update-frequency-sensitive invariant.\n *61│- **Sister cite:** `src/live/wrapper/lifecycle.rs:17-106`; `src/cli.rs` defaults; CHANGELOG v1.11.10.\n  62│\n  63│### 2.4 Generation `gen_start` always = now() on cold-start AND handoff\n  64│- **Failure:** stale gen_start from a rehydrated state file fires time-based discriminators on the new process.\n    │...\n  66│- **spt-core mapping:** carries to the daemon's per-instance generation tracking.\n *67│- **Restoration D3/D4 (ADR-0018):** the generation *counter* custody moved to the broker (D3-2 — it observes every brain spawn, planned or crash, and hands `{generation, start-reason}` at spawn; `gen_start` stays `now()`-fresh, never rehydrated). The brain→brain **`BrainState` *message*** (`{session_id, generation, next_seq, gen_start_ms}`) that previously carried continuity across a handoff is **retired from the production path in D4-2**: a brain the supervisor respawns cold-starts and reconstructs ...\n *68│- **Sister cite:** `src/live/wrapper/lifecycle.rs:70`; `src/common/wrapper_state.rs`.\n  69│\n  70│### 2.5 Daemon-hosted endpoints have no dedicated liveness PID\n  71│- **Failure:** the sister evaluates Psyche/perch liveness via a dedicated process PID — the wrapper's own pid in `info.json`, checked with `is_process_alive`. Under ADR-0004 the Psyche (and any spt-hosted Self) is a **loop inside the daemon**, not a separate process: it holds no dedicated pid, and its `claude`/summarizer subprocess is ephemeral (spawned per pulse/commune, then exits). If a daemon-hosted perch's `info.json` carries the **daemon's** pid, then *every* hosted endpoint shares one pid, and `...\n    │...\n  95│\n *96│## 3. Lifecycle\n  97│\n  98│### 3.1 Ephemeral perch cleanup on every `ring` exit path\n  99│- **Failure:** `ring` creates an ephemeral perch; early-exit paths (no-perch, empty-msg, timeout) skip cleanup → stale dirs accumulate.\n 100│- **Invariant:** every code path that creates an ephemeral perch cleans it before exit; exception: if the caller already had an active perch, do not treat as ephemeral and do not clean up.\n*101│- **spt-core mapping:** `ring` semantics carry; the daemon owns ephemeral-perch lifecycle, so a single guaranteed-cleanup (drop guard / RAII) is achievable in-process.\n 102│- **Sister cite:** `src/owl/ring.rs:58-294`.\n 103│\n 104│### 3.2 Stale signoff sentinel must not kill a fresh start\n    │...\n 196│- **Failure:** the registry's only superseding mechanism is the per-`(endpoint_id, node)` epoch lease (4.8) — a row is replaced only by a newer row *from the same node*. When a node identity dies permanently (machine retired, or `node.key` regenerated so the \"node\" never speaks again), its rows are never superseded and never expire: they sit in the in-memory registries and the `identity/registry/<subnet>.json` snapshots forever. A bare-id send then resolves the same endpoint id on both the live and the...\n*197│- **Invariant:** registry rows authored by a **silent** peer node decay: a node not *heard* (admitted inbound feed — the M7 D2 heard-map, REQ-SUBNET-1) within the eviction window (`registry_evict_after_ms`, default 300s ≈ 10 default pump cadences) has its rows **evicted** from every subnet registry, snapshots rewritten. Own rows never decay (the node always hears itself implicitly — it authors them each pump tick). Eviction is safe under the lease: v1 has **no transitive gossip**, so any future upd...\n 198│- **spt-core mapping:** `spt_net::net::registry::SubnetRegistry::evict_nodes` (model) + `spt_daemon::registryhost::RegistryHost::evict_silent_peers` (heard-map TTL) driven from the registry pump tick (`peerloop`). Trust rows are NOT auto-evicted (trust is a user decision; a stale trust row only costs dead dials) — pruning those is a separate verb.\n 199│- **Source:** M7 acceptance run 2026-06-06 (DEFERRED.md \"Ghost registry row eviction\"); the AMBIGUOUS render fix rode along.\n 200│- **Mesh note (ADR-0017, 2026-06-08):** the subnet mesh **preserves** this invariant rather than superseding it. \"No transitive gossip\" sharpens to **no transitive *row* gossip** — the mesh relays only the member *roster* (discovery), while registry **rows stay own-authored and are fetched directly** from each member over a handshake. So \"any future update for a node comes from that node itself, alive\" still holds and the eviction lease is untouched. (The plan's rejected alternative — signed transiti...\n    │...\n 227│- **Failure:** a hung subprocess (git on slow net) blocks the supervisor indefinitely.\n*228│- **Invariant:** every metadata-producing subprocess has a timeout; timeout yields `None` + rate-limited stderr, never a hang.\n 229│- **spt-core mapping:** generalize to all manifest-declared harness invocations (delegated commands, adapter updates) — timeouts mandatory.\n 230│- **Sister cite:** `src/common/git.rs`.\n 231│\n    │...\n 406│\n*407│### 6.7 Broker and brain MUST be separate processes (in-process collapse silently breaks no-endpoint-drop update) `[REQ-HAZARD-BROKER-PROCESS-ISOLATION]`\n*408│- **Failure:** the daemon hosts the broker as a background *thread* in the single `spt daemon` process (`daemon.rs:165-170`, `Arc<Broker>` + `thread::spawn(serve)`) instead of a separate process. A brain restart onto a swapped binary then cannot happen without killing the broker thread — closing every PTY, orphaning every harness child, dropping every socket. So `spt update apply` degrades to an in-process `Brain::handoff` no-op: the binary swaps on disk but the running daemon keeps executing the old c...\n*409│- **Invariant:** the broker runs as its own long-lived process that survives every brain restart; the brain restarts onto the new binary and re-attaches via the versioned IPC. A routine (brain-only) update must leave every hosted endpoint untouched at the *process* level — not merely re-subscribe a brain within the same process. The evidence for REQ-UPD-3 / REQ-DAEMON-2 must prove process-level survival (a PTY child + a live QUIC conn survive a brain-process restart onto a swapped binary — SPIKE-01/0...\n*410│- **spt-core mapping:** restoration is ADR-0018 (next milestone). The current `int` tags on REQ-DAEMON-2 / REQ-UPD-3 are regression-masked and re-point at restoration; the broker becomes the always-up per-machine anchor (seed-lock + liveness + brain supervisor). Two-process supervision, generation custody, durable-deadline loop timing, broker-cursor-of-record, and readiness-gated auto-rollback all hang off this.\n*411│- **Origin:** unintended spec/impl drift from ADR-0004 (the broker *process* was specced + spiked but built in-process), discovered during the v0.3.2 fleet update verify. Full audit + decisions: `docs/BROKER-BRAIN-SPLIT-RESTORATION.md` (verified) + ADR-0018.\n 412│<!-- [doc->REQ-HAZARD-BROKER-PROCESS-ISOLATION] -->\n*413│- **D1 (restoration skeleton, ADR-0018 Q2/Q3):** the process boundary is restored — `spt daemon run` is the broker process and spawns a supervised `spt daemon brain` child (`brainproc.rs`); the broker survives the brain dying and respawns it (proven in production topology by `crates/spt/tests/brain_split.rs`). The logic loops still run broker-side (D2 migrates them); the `int` process-level survival E2E + the in-process re-point land at D7.\n*414│- **Closed out (2026-06-11, v0.4.0–v0.4.2):** the two-process model shipped (v0.4.0); the D7 `int` E2E (`brain_survive.rs`) + the N-1 gate prove process-level survival onto swapped bytes and re-pointed REQ-DAEMON-2 / REQ-UPD-3. The v0.4.1 fleet-verify proved this Windows-seamless (hfenduleam: brain pid rolls, broker held, `exe_hash` flips, no manual bounce) but exposed a **Linux** respawn-path gap — the resident broker respawned the brain via per-spawn `current_exe()`, which on Linux follows the `app...\n 415│\n 416│### 6.8 No irreversible durable-state migration before update ready-promotion `[REQ-HAZARD-ROLLBACK-STATE-COMPAT]`\n*417│- **Failure:** the readiness-gated auto-rollback (ADR-0018 Q7) spawns the *previous* binary against durable state the *new* brain already wrote. The first release that migrates a durable-state schema in place would make the old binary unable to read it — silently bricking rollback exactly when it is needed (a logic-bricking update that can no longer fall back).\n*418│- **Invariant:** a brain must not irreversibly migrate durable state before it is ready-promoted; equivalently, every pre-ready write must remain readable by the N-1 brain. Schema migrations are gated behind ready-promotion (or written in an N-1-tolerant additive form).\n 419│- **spt-core mapping:** lands with ADR-0018's auto-rollback. Free to assert now (a 2026-06-09 source audit confirmed zero state-migration code exists); unmintable retroactively once a migration ships.\n 420│- **D5 conformance (2026-06-10):** the new durable timing state `<spt_home>/deadline-<key>.json` (restoration D5-1) is **additive** — a rolled-back pre-D5 binary does not know the file and simply ignores it (re-phasing on its own flat-sleep cadence, the pre-D5 behavior). No existing-file schema migration, no irreversible pre-ready write → the new file is rollback-N-1-safe by construction. The thing for a future D6 guard to gate is a *migration* of this file's shape, not its introduction.\n*421│- **D6 guard (2026-06-10, restoration D6-3):** the invariant is now **asserted**, not just noted. The pre-ready durable writes are **enumerated in one place** — `spt-daemon::PRE_READY_DURABLE_FILES` (`rollback_compat.rs`): `deadline-<key>.json` (D5, `DeadlineAnchor`), `applied-state.json` (D6-1, the two-phase `AppliedRecord`), and the generation-stamped `brain.ready` breadcrumb (D6-1b, `{pid, generation}`). A **tripwire unit test** pins each one's additive / N-1-readable contract (load-bearing field na...\n### TRACEABILITY.md\n 20│\n*21│`doc` (prose/design/rustdoc) · `impl` (production code) · `unit` (unit test) · `int` (integration / e2e / cross-node). Default activation policy: `[\"doc\",\"impl\",\"unit\"]`; networking/lifecycle/cross-node reqs add `int`.\n 22│\n 23│### Tag examples\n 24│\n   │...\n 70│\n*71│## Lifecycle\n 72│\n 73│- **Now (planning):** manifest seeded, all reqs inactive. `check` is green (nothing required yet).\n 74│- **M0:** install the CLI, validate the seed against it, wire the CI workflow, **activate M0 reqs** (`REQ-ARCH-*`, `REQ-HAZARD-*` for the commodity layer), tag as code lands.\n### DOCS-STRATEGY.md\n  17│7. **Consistent conversational voice** — adopt the Google developer style guide: second person (\"you\"), active voice, knowledgeable-friend tone.\n *18│8. **Explain *why*, not just *what*** — conceptual docs + diagrams for the messaging model, the agent lifecycle states, instances/dormancy, and P2P topology.\n  19│9. **Stable, never-renamed anchors/URLs** — agents cache links.\n  20│10. **Docs-as-product, gated in CI** — generation (rustdoc, `llms.txt`, schema, CLI help) is part of the build so docs can never drift from code. Drift is the #1 most-cited docs failure; this kills it structurally.\n  21│\n    │...\n  25│\n *26│**Messaging · Live-agent lifecycle · Terminal hosting · Networking (P2P + pairing) · Harness contract (manifest + `api`) · Instances · Shells · Self-update.**\n  27│\n  28│## Killer quickstart targets (two, one per audience)\n  29│\n  30│- **Human dev:** two agents exchange a message end-to-end in **one command + minimal config**, < 10 min — the core primitive, whole-thing-first, then decomposed.\n *31│- **Dev-agent integrator:** a minimal **adapter** that satisfies the manifest + `api` contract and drives one lifecycle operation — the \"build a harness for spt-core\" hello-world.\n  32│\n  33│## Agent-consumable docs (spt-core specifics)\n  34│\n    │...\n  74│A harness-independent core for an agent ecosystem: messaging, live-agent\n *75│lifecycle, terminal hosting, P2P networking, a runtime-manifest harness contract.\n  76│\n  77│> Each capability links to its own llms.txt. Full concatenated docs: /llms-full.txt.\n  78│> Machine-readable contract: /manifest.schema.json. Append `.md` to any doc URL\n    │...\n  87│- [Messaging](…/messaging/llms.txt): typed ops + file blobs; send / ring / ready.\n *88│- [Live-agent lifecycle](…/lifecycle/llms.txt): skeleton → live → dormant → suspended.\n  89│- … Terminal · Networking · Harness contract · Instances · Shells · Self-update\n  90│\n  91│## Reference (auto-generated)\n    │...\n 122│**Vertical ordering (dependency-sequenced):**\n*123│`Messaging → Lifecycle → Terminal → Networking → Harness contract → Instances → Shells → Self-update`\n 124│(messaging is the primitive the quickstart needs; self-update last because its no-termination invariant depends on the terminal/broker model.)\n 125│**Per-vertical internal template (Django labels × Cloudflare ordering):**\n 126│`Overview (why+diagram) · Quickstart/Tutorial · How-to guides · Reference (rustdoc+CLI+schema) · llms.txt`\n    │...\n 133│//! # spt-runtime — one-sentence value prop.\n*134│//! ## A tour of spt-runtime   - [`Manifest`] · [`ApiCommand`] · [`Lifecycle`]\n 135│//! ## Example                 (runnable, real values, near the top)\n 136│//! ## Feature flags           `net` (default) · `unstable`\n 137│//! ## Wire-protocol compatibility   interop by proto version, never crate version.\n### STORAGE.md\n  43│    project-index.json           endpoint→project attribution (schema-versioned; daemon-written)\n *44│  daemon/                        ── ephemeral ── brain rehydration state, broker handle metadata\n  45│  updates/                       ── ephemeral ── staged binaries pending handoff\n  46│  adapters/                      registered adapter manifests / pointers\n  47│    <adapter>/web/               core-owned adapter served root; output bytes survive deactivation\n    │...\n 173│| subnet registry (snapshot + distribution) | **BranchStore** | audited; rides the precedence merge driver instead of a separate replication scheme (M4) |\n*174│| daemon/brain coarse checkpoint (rehydrate anchor, ADR-0004) | **BranchStore** | commit = checkpoint, ref tip = resume |\n*175│| endpoint lifecycle / status *audit trail* | **BranchStore** (or a derived view) | the *trajectory* matters, not just current value |\n 176│| spool (queue, drain-by-status, `id→address`) | **SQLite** | indexed queries + hot churn |\n 177│| broker effect journal (per-PTY-effect, B5) | **fsync-append** | hot path, per-keystroke — a commit per write is far too heavy |\n 178│| `info.json` / `status` (current value) | flat + atomic-rename (today) | tiny, read-hot; *may* move under BranchStore later, low priority |\n    │...\n 185│\n*186│**Reading the effect journal (`<spt_home>/daemon-effects.log`) — operator note.** The broker fsync-appends two line kinds per durable effect: `PENDING <key> <kind>` at attempt, `DONE <key>` on success (durable kinds only — an ephemeral `PtyWrite` writes no lines). A line with **`PENDING` and no matching `DONE`** is **expected**, not a leak: it is a durable effect that did NOT complete — for a `net-send` (a WAN write to an offline/unreachable peer) it is an **unconfirmed send retained as the re-driv...\n 187│\n 188│**Delivery vs propagation — latency envelope (operator note).** Two classes of cross-node traffic have deliberately different latency floors, so a minutes-scale delay is not automatically a defect: (1) **messages** (`spt send`) are **event-driven** — delivered on the events the daemon owns (WAN ingress + the ACTIVE→IDLE edge, F-023), so a healthy path is near-immediate and never waits on an adapter hook-poll cadence; (2) **anti-entropy context-sync** (the mind pull) is **periodic** — it rides the...\n 189│\n    │...\n 274│  one refresh. Fire-and-forget and daemon-offline-safe (events accumulate\n*275│  and are consumed at the next brain start).\n 276│  <!-- [doc->REQ-PROJECT-INDEX-INVALIDATION] -->\n 277│\n 278│## perches/ — node-local runtime\n    │...\n 312│\n*313│- **Relay-death convergence** (`livehost::reconcile_hosted_liveness`) is reached only by rows with `controllable != Some(true)`, and now acts only on `pid_role == relay`. So: harness relay rows converge on a proven relay death; `binder` rows (every `api bind` product of an agent or non-agent path, including gateway and shell binds) are **never** convergeable no matter how dead their pid; legacy rows are never convergeable; rows with a live broker PTY (`controllable == Some(true)`) never reach the branch ...\n 314│- **`controllable` carry-forward** (`api/startup.rs`) drops an earned `Some(true)` on a listener re-bind **only** when the prior row's `relay` pid is provably `Gone`. A dead `binder` pid, an `Unproven` verdict, and a legacy row all inherit.\n 315│\n 316│Both consumers route their liveness verdict through the single authority (`liveness::relay_liveness` over the `proc::process_identity` oracle); `Unproven` never kills.\n    │...\n 331│- **Adapter/platform-bound, not adapter-agnostic** — capability surface comes from the providing shell adapter; it does not move harnesses.\n*332│- **Lifecycle = online / offline / torn-down.** Link-break always closes the binary (optional pre-close instruction + termination timeout). **Ephemeral** (manifest property) ⇒ perch torn down *and removed from the agent's shell history*; **persistent** ⇒ perch kept offline for relink.\n 333│- **Cross-node links:** the shell perch lives on the shell's node (nested under the owner id as a label there); commands ride Iroh from the owner's node.\n 334│\n 335│## Migration & backup\n### MANIFEST.md\n  88│Unknown keys are ignored, never refused (no `deny_unknown_fields`, by design);\n *89│`spt adapter add` names each on stderr — `manifest: unknown key [<table>].<key>\n  90│(ignored)` — and proceeds, so a misspelled optional key is not visible only by\n  91│its absence. Free-form tables (`[profiles.*]`, `[strings]`) are not swept.\n  92│\n    │...\n 204│spawn provisional on record — spt-core does **not** feed the provisional to your\n*205│template: it prints `RESUME_NO_HARNESS_SESSION:<id>` on stderr and starts a\n 206│**fresh** session through `[session.self]`. So a declared `[session.resume]`\n 207│legitimately does not run in that case, and the notice — not silence — is how\n 208│you tell that apart from a template that failed.\n    │...\n 232│|---|---|---|\n*233│| `{node}` | this node's **advertised label** — the label store value (OS-hostname default, re-checked at daemon startup), OS-hostname fallback when no store label is known | **SINGLE-TOKEN fill only** — tokenize-then-fill cannot produce a space-carrying argv element, so a composite display name like `<id> @ <node>` stays adapter-shim territory (`{node}` complements the shim for tokenizable args, e.g. `--remote-control {id}--{node}`, it does not replace it). **Caveat:** the daemon-side lifecycle reso...\n 234│\n 235│<!-- [doc->REQ-MANIFEST-NODE-KEY] -->\n 236│```toml\n    │...\n 422│Opt-in. The adapter's **idle-delivery translation binary**: a pure stdin→stdout\n*423│JSON-lines filter spt-core lifecycle-manages (spawned when the spt-hosted endpoint\n 424│comes up, terminated when it goes down). spt-core feeds it the inbound `<EVENT>`\n 425│feed and reads back keystroke-commands, which spt-core applies to the broker-held\n 426│PTY **atomically** — controller input is buffered during the emitted sequence and\n    │...\n 439│- **`path` (deprecated)** — the original bare-program form. It still **parses** (forward/back-compat) but **warns at registration**, steering authors to `command`. **Exactly one of `{path, command}`** may be set: both-set is **refused** at registration; neither set = **no translation binary** (the endpoint falls back to its `[inject]` hook path).\n*440│- Whichever form is used, the resolved program is spawned at the initial spt-hosted bringup **and** at a live-update respawn. (A declared binary that fails to spawn is logged `TRANSLATION_SPAWN_FAILED` on the daemon's stderr; before diagnosing idle-delivery behavior, confirm the binary actually spawned.) <!-- [doc->REQ-INSTALL-11] -->\n 441│- **stdin** (spt-core → binary, one JSON object per line): `{\"type\":\"init\",\"endpoint_id\":…,\"node\":…}` first · `{\"type\":\"event\",\"envelope\":\"<EVENT…>\"}` per inbound message (the ADR-0020 envelope) · `{\"type\":\"input\"}` a **content-free** ping each time the operator types (so the binary can track user-idle; the PTY input content is **never** duplicated to the binary).\n 442│- **stdout** (binary → spt-core, one per line): `{\"key\":\"ctrl+s\"}` · `{\"delay_ms\":50}` · `{\"text\":\"<payload>\"}` · `{\"key\":\"enter\"}` · **`{\"commit\":true}`**, … (extensible vocabulary).\n 443│- **`{\"commit\":true}` is the mandatory sequence TERMINATOR, required for EVERY `{\"type\":\"event\"}`.** While a sequence is in flight spt-core buffers the live `spt rc` controller's keystrokes (the *inject floor*) and applies the emitted commands atomically; the trailing `{\"commit\":true}` releases the floor and flushes the buffered controller input *after* the sequence. The submit keystroke is NOT the terminator (a choreography may keep typing after `{\"key\":\"enter\"}`). **An empty response is a protocol viol...\n 444│- **Commit deadline + miss tolerance (F029 C-1, ADR-0022 amendment).** If no `{\"commit\":true}` arrives within the **commit deadline (5 s, `INJECT_COMMIT_DEADLINE`)**, spt-core flushes the buffered operator input (never stranded) and re-spools that message once — but a SINGLE miss does NOT fault a healthy binary (it is preserved; the next event delivers through it). Only **3 consecutive** missed commits, or a real binary death, faults the binary — and a real fault gets a **bounded eager respawn** (a h...\n*445│- **Known keys today:** `command` (preferred) and the deprecated `path`. spt-core does **not** `deny_unknown_fields` here — a newer adapter declaring a future key against an older spt-core parses fine (the unknown key is ignored), so the lifecycle-binary contract degrades gracefully.\n 446│- A minimal binary just emits `{\"text\":payload}{\"key\":\"enter\"}{\"commit\":true}` with no choreography (spt-hosted idle delivery is translation-binary-only since v0.14.3 — no raw-inject fallback).\n 447│\n 448│### `[service]` — a daemon-supervised resident service (ADR-0049)\n    │...\n 535│  | `status-advisory` | **the service** (optional) | one **advisory** line surfaced by `spt adapter service status` |\n*536│  | `startup.capture` | spt-core | the service's own stdout+stderr from its startup window — the evidence a `STARTUP_FAULT` carries |\n 537│\n 538│  `status-advisory` is the one file in this contract an adapter writes and core\n 539│  reads. It is **display-only and never consulted for a decision** — no","truncation":{"content":"# .worktrees/286-churn/\n## PRD.md\n 9|\n*10|spt-core is a **harness-independent core for an agent ecosystem**: inter-agent messaging, live-agent lifecycle, terminal hosting, seamless self-update, and zero-config cross-machine networking — shipped as both a Rust library workspace and a single canonical binary (`spt` / `spt.exe`).\n 11|\n 12|It is the ground-up rebuild of `claude_skill_owl` (\"modern SPT\"), untethering the system from Claude Code so that **any** agent runtime — Claude Code, Codex, Cursor, headless, or novel surfaces — can participate by either shelling out to the binary or linking the crates. The first thing built atop it is a rebuilt `spt` plugin reaching parity with today's system; the networking, terminal-wrapper, and multi-instance capabilities then make an agent reachable and operable across all of a user's machines.\n 13|\n...\n 16|- Today's modern SPT works but is **fragile** — each release introduces 2–5 regressions because it grew rapidly from a narrow starting point and carries the resulting tech debt (ADR-0001).\n*17|- It is **welded to Claude Code**. Other harnesses can't reuse the messaging/lifecycle substrate.\n 18|- Networking, terminal hosting (\"capsule\"), and cross-machine operation were planned piecemeal atop a CC-tied core; they deserve to be first-class in a clean foundation.\n 19|\n 20|## 3. Goals\n...\n 54|- **R-ARCH-3** `spt-proto` carries an explicit **wire-protocol version** independent of crate semver, with a documented **N-1 compat window**. Binary and library consumers interoperate by proto version, never by release version.\n*55|- **R-ARCH-4** Copy-verbatim the commodity layer from the sister project (envelope grammar, spool DDL, registry RPC, `info.json` shape); clean-room only the architecture/lifecycle (ADR-0001).\n 56|\n*57|### 7.2 The daemon (broker/brain)\n 58|- **R-DAEMON-1** Exactly **one logical `spt-daemon` per machine**, always-on, owning PTYs, network identity + WAN endpoint, subnet registry, all spools, all poll-listener logic, and all Psyche/pulse loops. No separate listener or wrapper processes (ADR-0004).\n*59|- **R-DAEMON-2** Internally split into a stable **broker** (holds PTY masters, harness child processes, listening sockets; versioned IPC; rarely updates) and an **updatable brain** (all logic; rehydrates from disk + re-attaches to broker handles).\n 60|- **R-DAEMON-3** Any `spt api` invocation that needs the daemon **auto-starts it** if absent.\n 61|- **R-DAEMON-4** Honor every invariant in [`docs/KNOWN-HAZARDS.md`](./docs/KNOWN-HAZARDS.md) as a conformance checklist.\n 62|\n...\n 155|- **R-UPD-2** **All binaries signature-verified before handoff**, regardless of source (spt-core release key).\n*156|- **R-UPD-3** **No-endpoint-termination invariant, scoped by update class (ADR-0004 §A).** For the **brain-only** class — the routine case — the guarantee is **absolute**: no endpoint process terminates or suspends, satisfied by the broker/brain split (the broker keeps held PTYs/children/sockets alive across brain replacement). **broker-compatible** swaps the broker behind the versioned IPC with endpoints still surviving. **broker-breaking** (held-resource-type / OS-API change) MAY suspend endpoints...\n 157|- **R-UPD-4** **Not auto by default** — gated on user confirmation delivered to the most-recently-active live session, with an opt-in full-auto.\n 158|- **R-UPD-5** spt-core conducts **adapter ripple-updates** after self-update, per each adapter's update declaration and `min_spt_core_version` gate.\n 159|\n...\n 176|- **R-INSTALL-3** First run idempotent + interactive-optional: generate node identity, start daemon; pairing + subnet-naming are separate explicit steps.\n*177|- **R-INSTALL-4** (CONTEXT \"adapter registration\") **Adapter registration lifecycle:** `spt adapter add <path>` (or `--github <user/repo>`, manifest-first then install via the declared `[update]` avenue — *install is the first update*) validates the manifest and records it under `{SPT_HOME}/…/adapters/` (copy for `file_pull`, pointer for `delegated`); one command/dir for `kind=\"harness\"` and `kind=\"shell\"`. Node-local (drive/launch capability, not subnet advertisement). `spt adapter remove` is **soft...\n 178|- **R-MIGRATE-1** Standalone install **auto-detects a `claude_skill_owl` install and offers migration** (identity, agents, tracked Psyche context). First-class supported path.\n 179|\n 180|---\n...\n 195|\n*196|**In v1:** harness-independent core (library + binary); manifest + `api` harness contract; both startup topologies; one-daemon broker/brain; LiveAgent/ReadyAgent/Psyche/Worker/SptNode endpoints + Shell-vs-agent seam + typed/binary payloads; multi-instance data model + subnet registry + resolution policy + dormant/suspended + two-tier context sync + remote-drive of running instances + off-node file transfer; first-class WAN networking + TOTP-SPAKE2 pairing + subnet naming; cross-node Psyche sync (replaces...\n 197|\n 198|**Deferred (seam-compatible):** see [`docs/DEFERRED.md`](./docs/DEFERRED.md) — concrete Shells, PresenceChannel impl, presence gossip, instantiate-anywhere + consent, remote command exec + security gate, sidecar adapters, manifest includes, scrollback spillover, predictive/state-sync terminal layer, OS-level presence signal, macOS.\n 199|\n...\n 201|\n*202|1. A non-CC harness (or a headless spike) drives the full live-agent lifecycle through manifest + `api` alone — proving harness independence.\n 203|2. The rebuilt spt-plugin reaches **feature parity** with frozen `claude_skill_owl`.\n 204|3. Two of a user's machines, freshly paired via one TOTP code, exchange messages and the user drives an agent on machine A from machine B — zero config beyond pairing.\n 205|4. A self-update rolls across the subnet **without terminating any endpoint**.\n## ROADMAP.md\n 6|\n*7|A harness-independent core for an agent ecosystem — messaging, live-agent lifecycle, terminal hosting, **first-class P2P networking**, and a runtime-manifest harness contract — shipped as Rust crates + a single `spt` binary. The proof: the same agent runs natively across all of a user's machines with a synced mind, reachable and operable from anywhere, with seamless self-update and zero-config cross-machine messaging. First thing built atop it: a rebuilt `spt` plugin at parity with today's modern SPT.\n 8|\n 9|## Scope decisions (locked)\n 10|\n...\n 28|- **M1 — Local messaging + binary.** `spt-msg` (delivery TCP+spool, routing, send/ring/ready) + the `spt` binary + CLI. **= the killer quickstart (local): two agents exchange a message.** Task breakdown: [`M1-PLAN.md`](./M1-PLAN.md).\n*29|- **M2 — Harness contract + lifecycle.** **Strictly harness-agnostic — zero Claude Code conventions in this repo.** spt-core defines the adapter contract; the only adapter-shaped artifact in-tree is a generic mock/test adapter (manifest fixture + fake `api` caller) that exercises the contract (PRD R-DOCS-2 dev-agent quickstart). **The rebuilt `spt` plugin (Claude Code) is a SEPARATE downstream project** built atop these crates/binary — spt-core v1's acceptance proof (PRD §\"feature parity\"), not an...\n 30|  - **M2a — Harness contract. ✅ delivered (2026-06-01).** `spt-runtime` (AgentRuntime, ManifestRuntime, manifest schema) + the `api` subcommand surface (seed/listen/bind/state/echo-gate/poll/worker-*/boundary/session-end/presence/history-log/emit/capability) + the **harness-hosted** startup topology (seed→listen→bind, interim seed-file) + local-api-auth + the generic mock adapter + contract E2E. All M2a reqs activated + green. Task breakdown: [`M2a-PLAN.md`](./M2a-PLAN.md).\n*31|  - **M2b — Live-agent lifecycle. ✅ delivered (2026-06-01).** `spt-live` (fifth layer): history subsystem (fetcher/locate-normalize/native) + spawn-psyche (nested perch, spt-core-owned `$psyche_prompt`) + echo-commune (history brief → single-writer commune drop) + drop-file ingest (single-writer 6.4 + direct-write precedence 6.5) + interim pulse + signoff/boundary ordering invariants (echo-before-signoff 3.3, grace-before-signoff 1.1, stale-signoff sweep 3.2) + resume seam (continue-existing / fres...\n 32|  - **Interim, no daemon** (like M1): the `api listen` process holds the perch + relays; the seed bridges via a file. **The spt-hosted topology (PTY-launch) + the consolidated `spt-daemon` move to M3** (they need the `spt-term` broker) — this narrows the earlier \"both topologies in M2\" wording.\n*33|- **M3 — Terminal wrapper + daemon + self-update.** `spt-term` (session-surface, PTY, broker) + the consolidated `spt-daemon` (broker/brain split) + peer-propagated signed self-update. Split M3a/M3b/M3c. Scope + carried-forward Spike #1 gaps: [`M3-PLAN.md`](./M3-PLAN.md).\n 34|  - **M3a — `spt-term` session-surface mechanism. ✅ delivered (2026-06-02).** The sixth crate (`…→spt-msg→spt-term`, sibling of spt-runtime): the `SessionSurface` trait + native PTY backends (ConPTY/forkpty via `portable-pty`, REQ-TERM-1) · ConPTY-DSR auto-answer drain pump (REQ-HAZARD-CONPTY-DSR) · send-keys/send-line injection (REQ-TERM-2) · bounded byte-stream (REQ-TERM-3) · the PTY-digest parser primitive (REQ-TERM-4 impl/unit, ADR-0008). OS-neutral per Spike #4/#5 (no OS's stream contr...\n*35|  - **M3b — `spt-daemon` broker/brain process. ✅ delivered (2026-06-03).** The seventh crate (`…→spt-live→spt-daemon→spt`) hosts many `spt-term` surfaces behind a **versioned local IPC** (named-pipe/UDS behind one `DaemonTransport` trait, forward-compatible frame schema): a **broker** (stable kernel — PTY masters, harness children, sockets; no logic) + a **brain** (restartable logic) so a brain kill/restart leaves the hosted PTY child + its output stream **intact and gapless** (Spike #1 mad...\n*36|  - **M3c — signed self-update. ✅ delivered (2026-06-03).** The gated, signed, ripple-capable self-update engine (no new crate — it lives in `spt-daemon`, signing primitives in `release.rs`): the **update-class taxonomy** (brain-only / broker-compatible / broker-breaking, ADR-0004 §A) with the **brain-only zero-interruption swap** that drives the M3b handoff substrate (snapshot→drop→re-attach) so a live PTY child + its output stream survive the logic swap **untouched** (REQ-UPD-3, proven E2E `...\n*37|  - **M3 COMPLETE** (M3a + M3b + M3c all ✅, 2026-06-03): the real ADR-0004 architecture is live — native PTY hosting, the consolidated broker/brain daemon, and seamless signed self-update. Next milestone: M4.\n 38|- **M4 — Networking + instances. ✅ delivered (2026-06-04).** The eighth crate `spt-net` (Iroh WAN endpoint bound to the node's own Ed25519 identity, mDNS LAN discovery, TOTP-SPAKE2 pairing over a dedicated pre-trust ALPN with transcript binding / rate limiting / seed rotation+transfer, subnet registry + NDJSON replication under the per-node epoch lease) · the multi-instance model (registry rows + bare-id resolution refuse-and-qualify + visibility/sync-membership gates + rename/collision + **resting-...\n*39|- **M5 — Shells, presence, deferred capabilities. ✅ delivered (2026-06-04;** scope locked with user 2026-06-04, see [`M5-PLAN.md`](./M5-PLAN.md) + per-task JIT plans D0–D9). The consent framework seam (grant store + interactive escalation + pre-consent flags; remote-exec / instantiate-anywhere reserved-but-refusing) · adapter registration lifecycle (`spt adapter add/remove`, manifest-first) · **shell hosting machinery + sleep/wake** (link-token channels with per-frame MAC, broker-launched binarie...\n 40|\n 41|## Post-v1-core milestones\n 42|\n...\n 46|\n*47|- **✅ Broker/brain process-isolation restoration — delivered (2026-06-11).** Design: [`docs/BROKER-BRAIN-SPLIT-RESTORATION.md`](./docs/BROKER-BRAIN-SPLIT-RESTORATION.md) (verified-with-amendments by `doyle`); decision: **ADR-0018**; build plan: [`RESTORATION-PLAN.md`](./RESTORATION-PLAN.md) (D1–D7). Corrects an unintended regression — the broker/brain split (ADR-0004) was specced + spiked as two *processes*, but the production daemon runs the broker as an in-process *thread* (`daemon.rs:165-170`)...\n 48|\n 49|- **🔜 NEXT milestone (planned 2026-06-12) — M9: Adapter customization & session surfaces.** Product of the two-day **Gateway grill** (operator + `doyle`; CONTEXT.md mints @ `8dc22ff`/`bd469b9`; ADR-0008 amended; the lecturn device design parked in the sibling `spt-lecturn/` repo for a later build). Shell-free, independently-shippable features: **adapter profiles** (sparse leaf-replace overlays on the parent manifest; shipped + local sources, local surviving adapter updates; composite `<adapter>:<pro...\n 50|- **🔜 M10 — Digest extractor & thread-spanning** (planned 2026-06-13; NEXT). Product of the digest grill (operator + `doyle`); governed by **ADR-0019** (reverses M9's \"no manifest seam\"), CONTEXT.md \"session digest\" rewritten. An `spt-claude-code` enabler, so it takes the M10 slot ahead of the Shell/remote-attach work (which shift to M11/M12). Adapter-declared **`[digest]` extractor seam** (imperative, no DSL; same source files as `[history]` by default + own-source escape; `api digest-entry` push f...\n...\n 57|\n*58|1. **Broker/brain no-terminate handoff** — prove a PTY child + open socket survive a daemon-*logic* restart (FD-passing or stable-broker shim). The hardest invariant in the design (ADR-0004).\n 59|2. **Iroh + mDNS smoke test** — two nodes, separate NATs, one message end-to-end; measure direct-vs-relay + binary size (research brief §5). Validates the whole networking premise.\n 60|3. **ConPTY + `portable-pty`** — host and drive a real session in a broker PTY on Windows + Linux.\n*61|4. **Manifest-driven lifecycle** — a manifest drives one real Claude Code session through spawn → bind → commune → signoff, proving harness-independence works in practice.\n 62|\n 63|## Immediate next steps\n 64|\n...\n 66|- ✅ **Stage A — codex adversarial pass** — done. 14 findings (5 FATAL, 9 SERIOUS) in `docs/reviews/STAGE-A-codex-redteam.md`.\n*67|- ✅ **Spike #1 (broker/brain handoff)** — PASS on Windows ConPTY. `docs/spikes/SPIKE-01-broker-handoff.md`. Found+fixed the ConPTY DSR-stall hazard. 4 open gaps carried to M3.\n 68|- ✅ **Spike #2 (Iroh smoke test)** — PASS (single-host). `docs/spikes/SPIKE-02-iroh-smoke.md`. iroh 0.98 API workable; 13.57 MB release. Cross-NAT/relay proof deferred to M4.\n 69|- ✅ **Findings folded** into ADR-0002/0003/0004/0005 + KNOWN-HAZARDS (+3) + traceable-reqs.toml (+4 REQ-HAZARD).\n 70|\n...\n 72|1. **Stage A — `plan-eng-review`** (interactive architecture pass; second half of Stage A).\n*73|2. **Spike #4 (manifest-driven lifecycle)** — a manifest drives one real session spawn→bind→commune→signoff. (Spike #3 ConPTY hosting is largely covered by Spike #1.)\n 74|3. Resolve the open user decisions logged in the ADR amendments (context-merge model, seed-rotation, R-UPD-3 rewording).\n 75|4. **M0** (workspace + `spt-proto` + `spt-store`) — **started**; task breakdown in [`M0-PLAN.md`](./M0-PLAN.md). Each milestone gets its own just-in-time plan (this one is M0); the roadmap stays the milestone sequence, the JIT plan is the task layer.\n 76|\n## CONTEXT.md\n 6|\n*7|Harness-independent core for the SPT ecosystem. Provides inter-agent messaging, live-agent lifecycle, terminal wrapping, self-update, and networking primitives — as both a Rust library workspace and a canonical reference binary. Designed so any agent runtime (Claude Code, Codex, Cursor, headless, future harnesses) can interface with the SPT ecosystem either by shelling out to the binary or by linking the crates directly.\n 8|\n 9|Successor to `claude_skill_owl` (today's \"modern SPT\"), which is being rebuilt as `spt-core` to untether the system from Claude Code and lift it to a general-purpose agent-ecosystem core.\n 10|\n...\n 22|**spt plugin** (separate downstream project — NOT an spt-core deliverable):\n*23|A rebuilt version of today's Claude Code `spt` plugin. It is the **first consumer** built *atop* spt-core and the **acceptance proof** of spt-core v1 (it reaches feature parity with modern SPT while delegating all core functionality to spt-core, primarily via `spt.exe`, with deeper hooks where useful) — but it **lives and builds in its own repository, outside spt-core**. It is a Claude-Code-specific *adapter*: it holds the Claude Code conventions (hooks, slash-commands, skill/plugin layout, `claude` se...\n 24|\n 25|**Pi** (disambiguation — two meanings, never conflate):\n 26|(1) **Pi, the coding agent/harness** (`badlogic/pi-mono`) — a harness example alongside Claude Code and Codex; this is the meaning in user-facing harness lists. (2) **Pi-class node** — Raspberry-Pi-class low-power hardware hosting a Shell-only or headless SPT node; an incidental hardware descriptor, never an explicit product example. Public-facing docs must disambiguate or avoid the bare word.\n...\n 35|Internally the logical daemon is split into two implementation layers for seamless self-update (see Self-update):\n*36|- **broker** (stable \"kernel\") — holds *only* the un-transferable, must-not-die resources: PTY master fds, the spawned harness child processes, and listening network sockets. Minimal, dumb, versioned local IPC. Almost never updates.\n*37|- **daemon brain** (\"userspace\") — all logic (routing, registry, pulse/psyche loops, manifest parsing, update orchestration). Restarts freely on update; rehydrates from disk state and re-attaches to the broker's held handles.\n 38|\n 39|Logical addressing is unchanged — still one per-machine `spt-daemon`; the broker is an internal layer, not separately addressable. There is exactly **one broker per machine** (per `SPT_HOME`) — *not* one per endpoint: a single broker holds every hosted endpoint's resources, and it is present whenever the daemon runs, even with zero endpoints online (the bare-daemon case). It is therefore the always-present per-machine layer, which is why the single-daemon lock + liveness anchor belong to it.\n 40|\n 41|**in-session relay**:\n*42|A thin, stateless `spt.exe` task that exists only in **harness-hosted** sessions (where the agent harness is the parent process and spt cannot reach into its process tree — today's Monitor model). It streams the daemon brain's events into the session's stdout. All *stateful* listener logic lives in the daemon; the relay is a dumb pipe, freely killable and respawnable. **spt-hosted** sessions need no separate harness-owned relay — the daemon owns the PTY and consumes the same poll feed itself. Idle de...\n 43|\n 44|### Deliverable shape\n 45|\n...\n 109|<!-- [doc->REQ-ADAPTER-GH-TRANSPORT] -->\n*110|A `.spt` may be **multi-platform**: shared `manifest.toml` + `strings/` at the root, role binaries under per-target-triple subdirectories (`x86_64-pc-windows-msvc/`, …); install/update extracts the shared root plus only the current node's triple, flattened into `install_dir`, so flat `<install_dir>/<program>` resolution is unchanged. It stays one signed asset (`adapter.spt`, plain-tar or gzip); a multi-platform archive missing the recipient's triple is a typed `NoArtifactForPlatform`. Large adapters ma...\n 111|<!-- [doc->REQ-ADAPTER-UPDATE-MESSAGE] -->\n 112|An optional **`[update].message`** (avenue-agnostic) is a plain multi-line operator notice surfaced to stdout, markdown-rendered (the helpfmt prose path), **only when an update is actually applied** (the version changed) — never on a no-op. It is read from the newly-installed manifest with no `{key}` substitution; its use is to announce a post-update action (e.g. \"run `/reload-plugins` in any ongoing sessions\").\n 113|\n 114|**composite update — `[update.post]`** (ratified 2026-06-25, v0.16.0 update-arc grill; ADR-0029):\n*115|An optional **avenue-agnostic** post-step `{ command, self_verifies }` spt-core runs **after** the primary avenue resolves — in the same `spt adapter update` **and at `spt adapter add`** (install is the first update, so a fresh install conducts the post-step too; bug-#1 operator ruling, v0.19.0 — the eager-extract acquisition runs it post-registration, a delegated acquisition after the acquisition succeeds, and only the payload-less `file_pull` PENDING add defers it to the payload's arrival) — so o...\n 116|\n 117|**resident adapter binary**: an adapter-owned process spt-core keeps alive for an endpoint's lifetime (today the `[message-idle-translation-binary]`), as opposed to **ephemeral** adapter binaries — the Psyche loop (daemon-hosted, ADR-0004), the `[digest]` extractor, `[session.*]` runners, hooks — which spawn on demand and pick up an update on their next invocation. Only resident binaries are stopped/restarted on a live update; ephemerals self-heal.\n 118|_Avoid_: calling the Psyche loop or an on-demand extractor a \"resident\" binary; \"restart the endpoint\" for what is a per-binary cycle.\n...\n 124|\n*125|Governing principle: **SPT is not a harness.** Model choice, billing shape, harness-internal env, and harness-internal context are entirely the adapter's concern, expressed inside the adapter's own command templates. spt-core owns only the template *mechanism* (substitution keys), the substitution *values* it is responsible for, and the surrounding lifecycle. Env for the *endpoint binary itself* is auto-handled by spt-core/broker; env for the *agent running inside* that binary is the adapter's config (e....\n 126|\n 127|**spawn-session seam** — launch a new agent session on this node. Manifest provides: a command template; `cwd`/project; a `headless` flag (optional, default false — for the GUI's resume-of-compatible-adapters); a `resume` flag (optional); and the `commune` + `signoff` file directories relative to `cwd` (so the daemon knows where to watch). Substitution keys spt-core can supply: `{id}` and, optionally, a spt-core-generated valid session UUID (e.g. injected as `--session-id {uuid}`) so an adapter can s...\n 128|- **id resolution:** `id` is optional. With no id, spt-core reproduces today's no-id `/spt:live` behavior — run the lone live agent if that's all the project has; show a picker with proposed default IDs if the project has none; let the user choose if there are several.\n...\n 175|- **Command templates are opaque.** spt-core never parses out a model/tool/flag — the adapter writes the whole command line; spt-core fills substitution keys and runs it.\n*176|- **A command template's program token resolves against the adapter install dir before PATH (since v0.8.0).** A `.spt` adapter ships its built binaries to the adapter's install dir (`adapters/_github/<safe>/` via `--release`/`--github`, or the record's `source_dir` under copy-mode), so a bare program name (e.g. `claude-spt-digest …`) binds to the shipped binary first and falls back to PATH when absent — a `.spt` that ships its binaries is **self-contained**, needing no PATH placement. <!-- [doc->REQ-...\n 177|- **Hook output capability is declared per harness-event** (`can_inject`). CC's Stop hook cannot inject context — that single fact drives the echo-gate sentinel + relay fallback. The manifest expresses it so spt-core knows when to fall back.\n 178|- **Env injection is asymmetric** (file-bridge-only-when-not-launcher, applied to env): spt-hosted sessions inherit env from the broker that spawns them; harness-hosted sessions need the harness's declared env channel. With `spt` on PATH the env table is small.\n 179|- **Cross-adapter fallback** is a **node-wide setting**, not a manifest field: if a Psyche/echo invocation under one adapter is rate-limited, spt-core falls back to another adapter (e.g. `ccs` — its own adapter, not a binary-swap). <!-- [doc->REQ-MANIFEST-6] --> A fallback **target is addressed as `<adapter>:<profile>`** (not just a bare adapter_name) and resolves through the one composite-addressing resolver (`registry::resolve_option`), so a fallback may select a shipped or local profile (`ccs`, `ccs...\n...\n 249|**Psyche**:\n*250|The Psyche companion's own perch, distinct from its paired LiveAgent's perch. First-class endpoint type so messages addressed to a LiveAgent's Psyche route directly without ambiguity. **A Psyche is a bounded per-event turn, not a resident process (since v0.25.0).** Each psyche-relevant event (a pulse fire, a commune/signoff drop, a session-custody transition) runs **exactly one** bounded turn through the psyche role template, spawned by the daemon, which exits at turn end — there is no long-lived psych...\n 251|\n 252|*I/O & trust boundary (ADR-0012):* the Psyche is a **sandboxed** actor — it may read and write files but **cannot send messages or reach the network itself**. Its inbound context arrives two ways: events/messages the daemon hands it, and **commune/signoff file-drops** (Self → daemon → Psyche; the *Summarizer* authors the commune delta). Its **sole outbound** is **reply/notify intents** the daemon relays as its **outbound proxy** — emitted as `<EVENT type=\"reply\">`/`<EVENT type=\"notify\">` (the sha...\n 253|\n*254|*Psyche-host health — harness-reachable failure signal (v0.8.1, REQ-HAZARD-LIVEHOST-BOOT-RACE):* a LiveAgent's `status=online` is daemon-authoritative liveness and **stays authoritative** — but it does not by itself prove the daemon hosted a Psyche. When the brain's live-host reconcile fails to spawn the Psyche (e.g. the adapter's psyche binary is absent from its install dir, or the net-less boot-race starves the host), that failure was previously **silent** — only an `eprintln!` on the brain's inv...\n 255|\n 256|**Summarizer**:\n 257|The ephemeral, cheap model that builds a **commune delta** from a Self's recent turns and feeds it *into* the **Psyche** as inbound context. A distinct actor from the Psyche — different (cheaper) model, fire-and-forget, **no perch** (not an endpoint type). It authors *commune* deltas only, **never** *reply*/*notify*.\n...\n 293|**ResidentService** (concept ratified 2026-07-26, ADR-0049 — the supervised substrate; first consumer spt-alchemy's Hub Daemon):\n*294|A **daemon-supervised binary an adapter declares, with no perch, no identity, no address.** Core owns the process from birth: the daemon spawns it **job-neutrally** (never a shell's child — a shell's tree-kill and a launching terminal's Job Object cannot reach it) and supervises it with the wake-watcher scaffolding (backoff, give-up latch, one-per-instance lock, orphan-kill, brain-side reconcile), running **independent of any agent's liveness**. Declared by the adapter manifest's **`[service]`** sectio...\n 295|_Avoid_: calling it an endpoint (no identity, no address — that is the [[AlwaysOnEndpoint]] layered on top); calling it a Shell (owner-less, not driven); \"daemon\" unqualified (the node has one daemon; this is a supervised service under it); treating its supervision record as liveness truth (there is none — liveness is the child handle).\n 296|\n 297|**AlwaysOnEndpoint** (always-on endpoint; concept ratified 2026-06-21, re-based on the substrate 2026-07-26 per ADR-0049 — core kind, first instance downstream `spt-discord`):\n...\n 311|\n*312|**resume custody**: The record a wake-resume spawn leaves (`resume.pid`) so lifecycle readers defer to a revival in flight rather than normalizing over it. Custody is an **identity — the (pid, process-creation-time) pair — never a bare PID** (ADR-0047): a recycled pid must read **not ours**, and the reader that discovers a mismatched pair deletes the record and proceeds (self-heal). Bind and reap clear custody atomically with their own outcome. _Avoid_: `is_process_alive(pid)` as a custody test (the ...\n 313|\n 314|**operator-stop inhibit**: The durable machine-scoped marker `daemon stop` records *before* teardown, which every implicit daemon-ensure path (REQ-DAEMON-3's anchor, as amended by ADR-0047) consults and honors by **declining to spawn** with one line naming the remedy. Cleared only by intent verbs (`daemon start`, the update paths that restart by design) — never by TTL. _Avoid_: reading the anchor as unconditional (that was the respawn-convoy bug — KNOWN-HAZARDS 7.52), or a TTL \"safety\" clear (a surpr...\n 315|\n...\n 332|\n*333|**Shells differ structurally from agent endpoints** (full treatment below): a Shell has a node-local perch but **no `tracked/` context** (no mind to sync); its logs are node-local; it is **adapter/platform-bound, not adapter-agnostic**; and its lifecycle is link/teardown, not the dormant/suspended resting model of agent instances.\n 334|\n 335|#### Shell model (detailed)\n 336|\n...\n 347|<!-- [doc->REQ-SHELL-4] shell tunnel: a long-lived reliable-ordered link-bound QUIC stream pair carrying opaque bytes the taxonomy never reinterprets; manifest opt-in, not enveloped/MAC-framed/spooled; link-break closes it; reliable-ordered ⇒ on-LAN posture -->\n*348|Channels carry typed, taxonomy-interpreted payloads. Distinct from them, an owner↔shell link may also hold a **shell tunnel**: a long-lived, **reliable, ordered** byte stream (a dedicated QUIC stream pair bound to the link) for protocol traffic the channel taxonomy must NOT reinterpret — opaque wire protocols spoken end-to-end (first consumer: USB/IP URB traffic to a usbip shell). Not spooled, not enveloped; the link's lifecycle governs it — a link-break closes the tunnel. Reliable-ordered is the p...\n 349|\n 350|<!-- [doc->REQ-SHELL-5] owner-type-agnostic: control-exclusivity keys on the owner endpoint_id, never the owner's endpoint type -->\n 351|**Owner-linked, exclusive.** Spawned by an **owning endpoint**; both linked; **only that owner `endpoint_id` may control it**. Agent endpoints are the common owner, but ownership is NOT agent-exclusive — any non-Shell endpoint type may own shells (ratified 2026-06-11, the Gateway grill: e.g. a **Gateway** owns the driven surfaces it steers; a future **Resource** endpoint could too). Control-exclusivity does not mean interaction-exclusivity: a Shell stays 2-way with its *environment* (user messages on t...\n...\n 354|\n*355|**Lifecycle = online / offline / torn-down.** A **link-break always closes the broker + shell binary** (with an optional manifest-declared **pre-close instruction** to the binary + a **termination timeout** for graceful shutdown). Then: an **ephemeral** shell (a manifest property, not an agent choice) is fully torn down **and removed from the agent's shell history**; a **persistent** shell keeps its perch **offline** (re-linkable — the binary is re-spawned on relink). **The binary never survives a link...\n 356|\n 357|**Broadcast policy** (manifest): the shell advertises its availability to all agents on the subnet, only same-node agents, or not at all (agents learn of it only by user instruction). Shells run only on nodes where the shell adapter is installed/advertised.\n 358|\n...\n 360|\n*361|**`shell spawn` mints a new instance, it is not the online switch.** `spt shell spawn <adapter> [--id]` creates a *new* Shell identity (`<adapter>-<n>` + its perch + the owner link). Bringing an *existing* offline (persistent) instance back online is **`shell relink <id>`** (re-spawns the binary), or happens automatically via `persistent` (auto-online with the owner) / `wake_command` (offline wake-watcher). `shell teardown <id>` destroys it. So **`spawn`:create :: `relink`/`persistent`/`wake`:online** ...\n 362|\n 363|**per-shell instantiation approval (`require_approval`, manifest enum):** gates `shell spawn`, reusing the consent plumbing (grant store + interactive escalation; see Consent & security gates).\n 364|- `none` (**default** — matches the system's everything-opt-in posture) — the agent spawns freely within scope.\n...\n 411|<!-- [doc->REQ-RCVIEW-1] [doc->REQ-KICK-1] [doc->REQ-VIEWER-SKIP-TO-LIVE-ON-EVICT] [doc->REQ-HAZARD-VIEWER-RING-ROLL-SNAP] -->\n*412|**BUILT (M12 W2.5).** The controller/viewer model is implemented end-to-end. Attach intent is **three-valued** (`AttachIntent = Viewer | Control | Take`, wire-default `Control`): `Control` to a FREE endpoint becomes controller; `Control` to a CONTROLLED endpoint is **refused with guidance** (`--view` to watch, `--take` to control) — never auto-viewer, never silent-displace; `Take` (`spt rc --take` / picker \"Kick\") kicks the incumbent with a **loud `Displaced{by}` notice** and full detach (not demote). ...\n 413|\n*414|**redispatch** (REDISPATCH-TRUTH, ADR-0038): the fresh brain's reconstruction of target-side attach workers over broker-held peer streams after a **brain cycle** (`spt daemon refresh` / `spt update` apply — the PTYs and QUIC streams survive; the brain-owned serve workers do not). Redispatch eligibility is **lifecycle-gated**: finished/terminal stream rows are retired from the dispatcher's enumeration (a fresh dispatcher never re-serves a terminal Attach — the KNOWN-HAZARDS 7.41 frozen-PTY steal/clear...\n*415|_Avoid_: conflating redispatch with msg dispatch or session-cursor resume; filtering replay by origin identity (the legitimate same-`by` successor re-take must keep working — lifecycle is the discriminator).\n 416|\n 417|<!-- [doc->REQ-RC-KEY-VT-TRANSLATE] -->\n 418|**rc keyboard input (Windows VT translation, v0.13.0 bug 2).** On **Windows** an interactive `spt rc` console reads crossterm **key events** and translates each to **standard xterm VT** (`translate_key_event`) — arrows / Home / End / PgUp/Dn / Insert / Delete / F-keys + modifiers all reach the harness as the universal terminal contract (**agnostic**, NOT win32-input-mode; the legacy console delivers those keys as events, not bytes, so the old byte-pump left them DEAD). **Unix passes through** (its raw-...\n\n## docs/\n### KNOWN-HAZARDS.md\n 4|\n*5|**Architecture-translation note.** The sister project runs poll listeners and Psyche wrappers as *separate processes*. spt-core consolidates both into the one `spt-daemon` (brain), with a stable broker beneath it (ADR-0004). Many hazards below were inter-process races in the sister project; in spt-core some become intra-daemon concerns (potentially easier) while others move to the daemon↔broker IPC boundary or the network boundary (potentially new failure surface). Each entry notes the mapping where it...\n 6|\n 7|---\n 8|\n...\n 34|\n*35|### 1.5 Worker (working-perch) lifecycle path consistency\n 36|- **Failure:** subagent-start creates the perch at one path layout; later hooks read it at another → not found; stop-hook scan misses nested perches.\n 37|- **Invariant:** all Worker/Psyche child-perch path composition routes through one central resolver; no divergent path construction across hooks.\n 38|- **spt-core mapping:** `Worker` is a day-one endpoint type; the daemon owns the registry, so perch location is a registry lookup, not ad-hoc path math. Single source of truth for instance→location.\n...\n 59|- **Invariant:** every newly-added handoff positional has a default; state-file rehydration happens *after* argv parse; defaults survive intermediate versions.\n*60|- **spt-core mapping:** CRITICAL — self-update (ADR-0004) makes handoff routine. The broker↔brain IPC and any brain-relaunch argv must be versioned and forward/backward tolerant (a newer brain talks to an older broker). This is the single most update-frequency-sensitive invariant.\n*61|- **Sister cite:** `src/live/wrapper/lifecycle.rs:17-106`; `src/cli.rs` defaults; CHANGELOG v1.11.10.\n 62|\n 63|### 2.4 Generation `gen_start` always = now() on cold-start AND handoff\n 64|- **Failure:** stale gen_start from a rehydrated state file fires time-based discriminators on the new process.\n...\n 66|- **spt-core mapping:** carries to the daemon's per-instance generation tracking.\n*67|- **Restoration D3/D4 (ADR-0018):** the generation *counter* custody moved to the broker (D3-2 — it observes every brain spawn, planned or crash, and hands `{generation, start-reason}` at spawn; `gen_start` stays `now()`-fresh, never rehydrated). The brain→brain **`BrainState` *message*** (`{session_id, generation, next_seq, gen_start_ms}`) that previously carried continuity across a handoff is **retired from the production path in D4-2**: a brain the supervisor respawns cold-starts and reconstructs ...\n*68|- **Sister cite:** `src/live/wrapper/lifecycle.rs:70`; `src/common/wrapper_state.rs`.\n 69|\n 70|### 2.5 Daemon-hosted endpoints have no dedicated liveness PID\n 71|- **Failure:** the sister evaluates Psyche/perch liveness via a dedicated process PID — the wrapper's own pid in `info.json`, checked with `is_process_alive`. Under ADR-0004 the Psyche (and any spt-hosted Self) is a **loop inside the daemon**, not a separate process: it holds no dedicated pid, and its `claude`/summarizer subprocess is ephemeral (spawned per pulse/commune, then exits). If a daemon-hosted perch's `info.json` carries the **daemon's** pid, then *every* hosted endpoint shares one pid, and `...\n...\n 95|\n*96|## 3. Lifecycle\n 97|\n 98|### 3.1 Ephemeral perch cleanup on every `ring` exit path\n 99|- **Failure:** `ring` creates an ephemeral perch; early-exit paths (no-perch, empty-msg, timeout) skip cleanup → stale dirs accumulate.\n 100|- **Invariant:** every code path that creates an ephemeral perch cleans it before exit; exception: if the caller already had an active perch, do not treat as ephemeral and do not clean up.\n*101|- **spt-core mapping:** `ring` semantics carry; the daemon owns ephemeral-perch lifecycle, so a single guaranteed-cleanup (drop guard / RAII) is achievable in-process.\n 102|- **Sister cite:** `src/owl/ring.rs:58-294`.\n 103|\n 104|### 3.2 Stale signoff sentinel must not kill a fresh start\n...\n 196|- **Failure:** the registry's only superseding mechanism is the per-`(endpoint_id, node)` epoch lease (4.8) — a row is replaced only by a newer row *from the same node*. When a node identity dies permanently (machine retired, or `node.key` regenerated so the \"node\" never speaks again), its rows are never superseded and never expire: they sit in the in-memory registries and the `identity/registry/<subnet>.json` snapshots forever. A bare-id send then resolves the same endpoint id on both the live and the...\n*197|- **Invariant:** registry rows authored by a **silent** peer node decay: a node not *heard* (admitted inbound feed — the M7 D2 heard-map, REQ-SUBNET-1) within the eviction window (`registry_evict_after_ms`, default 300s ≈ 10 default pump cadences) has its rows **evicted** from every subnet registry, snapshots rewritten. Own rows never decay (the node always hears itself implicitly — it authors them each pump tick). Eviction is safe under the lease: v1 has **no transitive gossip**, so any future upd...\n 198|- **spt-core mapping:** `spt_net::net::registry::SubnetRegistry::evict_nodes` (model) + `spt_daemon::registryhost::RegistryHost::evict_silent_peers` (heard-map TTL) driven from the registry pump tick (`peerloop`). Trust rows are NOT auto-evicted (trust is a user decision; a stale trust row only costs dead dials) — pruning those is a separate verb.\n 199|- **Source:** M7 acceptance run 2026-06-06 (DEFERRED.md \"Ghost registry row eviction\"); the AMBIGUOUS render fix rode along.\n 200|- **Mesh note (ADR-0017, 2026-06-08):** the subnet mesh **preserves** this invariant rather than superseding it. \"No transitive gossip\" sharpens to **no transitive *row* gossip** — the mesh relays only the member *roster* (discovery), while registry **rows stay own-authored and are fetched directly** from each member over a handshake. So \"any future update for a node comes from that node itself, alive\" still holds and the eviction lease is untouched. (The plan's rejected alternative — signed transiti...\n...\n 227|- **Failure:** a hung subprocess (git on slow net) blocks the supervisor indefinitely.\n*228|- **Invariant:** every metadata-producing subprocess has a timeout; timeout yields `None` + rate-limited stderr, never a hang.\n 229|- **spt-core mapping:** generalize to all manifest-declared harness invocations (delegated commands, adapter updates) — timeouts mandatory.\n 230|- **Sister cite:** `src/common/git.rs`.\n 231|\n...\n 406|\n*407|### 6.7 Broker and brain MUST be separate processes (in-process collapse silently breaks no-endpoint-drop update) `[REQ-HAZARD-BROKER-PROCESS-ISOLATION]`\n*408|- **Failure:** the daemon hosts the broker as a background *thread* in the single `spt daemon` process (`daemon.rs:165-170`, `Arc<Broker>` + `thread::spawn(serve)`) instead of a separate process. A brain restart onto a swapped binary then cannot happen without killing the broker thread — closing every PTY, orphaning every harness child, dropping every socket. So `spt update apply` degrades to an in-process `Brain::handoff` no-op: the binary swaps on disk but the running daemon keeps executing the old c...\n*409|- **Invariant:** the broker runs as its own long-lived process that survives every brain restart; the brain restarts onto the new binary and re-attaches via the versioned IPC. A routine (brain-only) update must leave every hosted endpoint untouched at the *process* level — not merely re-subscribe a brain within the same process. The evidence for REQ-UPD-3 / REQ-DAEMON-2 must prove process-level survival (a PTY child + a live QUIC conn survive a brain-process restart onto a swapped binary — SPIKE-01/0...\n*410|- **spt-core mapping:** restoration is ADR-0018 (next milestone). The current `int` tags on REQ-DAEMON-2 / REQ-UPD-3 are regression-masked and re-point at restoration; the broker becomes the always-up per-machine anchor (seed-lock + liveness + brain supervisor). Two-process supervision, generation custody, durable-deadline loop timing, broker-cursor-of-record, and readiness-gated auto-rollback all hang off this.\n*411|- **Origin:** unintended spec/impl drift from ADR-0004 (the broker *process* was specced + spiked but built in-process), discovered during the v0.3.2 fleet update verify. Full audit + decisions: `docs/BROKER-BRAIN-SPLIT-RESTORATION.md` (verified) + ADR-0018.\n 412|<!-- [doc->REQ-HAZARD-BROKER-PROCESS-ISOLATION] -->\n*413|- **D1 (restoration skeleton, ADR-0018 Q2/Q3):** the process boundary is restored — `spt daemon run` is the broker process and spawns a supervised `spt daemon brain` child (`brainproc.rs`); the broker survives the brain dying and respawns it (proven in production topology by `crates/spt/tests/brain_split.rs`). The logic loops still run broker-side (D2 migrates them); the `int` process-level survival E2E + the in-process re-point land at D7.\n*414|- **Closed out (2026-06-11, v0.4.0–v0.4.2):** the two-process model shipped (v0.4.0); the D7 `int` E2E (`brain_survive.rs`) + the N-1 gate prove process-level survival onto swapped bytes and re-pointed REQ-DAEMON-2 / REQ-UPD-3. The v0.4.1 fleet-verify proved this Windows-seamless (hfenduleam: brain pid rolls, broker held, `exe_hash` flips, no manual bounce) but exposed a **Linux** respawn-path gap — the resident broker respawned the brain via per-spawn `current_exe()`, which on Linux follows the `app...\n 415|\n 416|### 6.8 No irreversible durable-state migration before update ready-promotion `[REQ-HAZARD-ROLLBACK-STATE-COMPAT]`\n*417|- **Failure:** the readiness-gated auto-rollback (ADR-0018 Q7) spawns the *previous* binary against durable state the *new* brain already wrote. The first release that migrates a durable-state schema in place would make the old binary unable to read it — silently bricking rollback exactly when it is needed (a logic-bricking update that can no longer fall back).\n*418|- **Invariant:** a brain must not irreversibly migrate durable state before it is ready-promoted; equivalently, every pre-ready write must remain readable by the N-1 brain. Schema migrations are gated behind ready-promotion (or written in an N-1-tolerant additive form).\n 419|- **spt-core mapping:** lands with ADR-0018's auto-rollback. Free to assert now (a 2026-06-09 source audit confirmed zero state-migration code exists); unmintable retroactively once a migration ships.\n 420|- **D5 conformance (2026-06-10):** the new durable timing state `<spt_home>/deadline-<key>.json` (restoration D5-1) is **additive** — a rolled-back pre-D5 binary does not know the file and simply ignores it (re-phasing on its own flat-sleep cadence, the pre-D5 behavior). No existing-file schema migration, no irreversible pre-ready write → the new file is rollback-N-1-safe by construction. The thing for a future D6 guard to gate is a *migration* of this file's shape, not its introduction.\n*421|- **D6 guard (2026-06-10, restoration D6-3):** the invariant is now **asserted**, not just noted. The pre-ready durable writes are **enumerated in one place** — `spt-daemon::PRE_READY_DURABLE_FILES` (`rollback_compat.rs`): `deadline-<key>.json` (D5, `DeadlineAnchor`), `applied-state.json` (D6-1, the two-phase `AppliedRecord`), and the generation-stamped `brain.ready` breadcrumb (D6-1b, `{pid, generation}`). A **tripwire unit test** pins each one's additive / N-1-readable contract (load-bearing field na...\n### TRACEABILITY.md\n 20|\n*21|`doc` (prose/design/rustdoc) · `impl` (production code) · `unit` (unit test) · `int` (integration / e2e / cross-node). Default activation policy: `[\"doc\",\"impl\",\"unit\"]`; networking/lifecycle/cross-node reqs add `int`.\n 22|\n 23|### Tag examples\n 24|\n...\n 70|\n*71|## Lifecycle\n 72|\n 73|- **Now (planning):** manifest seeded, all reqs inactive. `check` is green (nothing required yet).\n 74|- **M0:** install the CLI, validate the seed against it, wire the CI workflow, **activate M0 reqs** (`REQ-ARCH-*`, `REQ-HAZARD-*` for the commodity layer), tag as code lands.\n### DOCS-STRATEGY.md\n 17|7. **Consistent conversational voice** — adopt the Google developer style guide: second person (\"you\"), active voice, knowledgeable-friend tone.\n*18|8. **Explain *why*, not just *what*** — conceptual docs + diagrams for the messaging model, the agent lifecycle states, instances/dormancy, and P2P topology.\n 19|9. **Stable, never-renamed anchors/URLs** — agents cache links.\n 20|10. **Docs-as-product, gated in CI** — generation (rustdoc, `llms.txt`, schema, CLI help) is part of the build so docs can never drift from code. Drift is the #1 most-cited docs failure; this kills it structurally.\n 21|\n...\n 25|\n*26|**Messaging · Live-agent lifecycle · Terminal hosting · Networking (P2P + pairing) · Harness contract (manifest + `api`) · Instances · Shells · Self-update.**\n 27|\n 28|## Killer quickstart targets (two, one per audience)\n 29|\n 30|- **Human dev:** two agents exchange a message end-to-end in **one command + minimal config**, < 10 min — the core primitive, whole-thing-first, then decomposed.\n*31|- **Dev-agent integrator:** a minimal **adapter** that satisfies the manifest + `api` contract and drives one lifecycle operation — the \"build a harness for spt-core\" hello-world.\n 32|\n 33|## Agent-consumable docs (spt-core specifics)\n 34|\n...\n 74|A harness-independent core for an agent ecosystem: messaging, live-agent\n*75|lifecycle, terminal hosting, P2P networking, a runtime-manifest harness contract.\n 76|\n 77|> Each capability links to its own llms.txt. Full concatenated docs: /llms-full.txt.\n 78|> Machine-readable contract: /manifest.schema.json. Append `.md` to any doc URL\n...\n 87|- [Messaging](…/messaging/llms.txt): typed ops + file blobs; send / ring / ready.\n*88|- [Live-agent lifecycle](…/lifecycle/llms.txt): skeleton → live → dormant → suspended.\n 89|- … Terminal · Networking · Harness contract · Instances · Shells · Self-update\n 90|\n 91|## Reference (auto-generated)\n...\n 122|**Vertical ordering (dependency-sequenced):**\n*123|`Messaging → Lifecycle → Terminal → Networking → Harness contract → Instances → Shells → Self-update`\n 124|(messaging is the primitive the quickstart needs; self-update last because its no-termination invariant depends on the terminal/broker model.)\n 125|**Per-vertical internal template (Django labels × Cloudflare ordering):**\n 126|`Overview (why+diagram) · Quickstart/Tutorial · How-to guides · Reference (rustdoc+CLI+schema) · llms.txt`\n...\n 133|//! # spt-runtime — one-sentence value prop.\n*134|//! ## A tour of spt-runtime   - [`Manifest`] · [`ApiCommand`] · [`Lifecycle`]\n 135|//! ## Example                 (runnable, real values, near the top)\n 136|//! ## Feature flags           `net` (default) · `unstable`\n 137|//! ## Wire-protocol compatibility   interop by proto version, never crate version.\n### STORAGE.md\n 43|    project-index.json           endpoint→project attribution (schema-versioned; daemon-written)\n*44|  daemon/                        ── ephemeral ── brain rehydration state, broker handle metadata\n 45|  updates/                       ── ephemeral ── staged binaries pending handoff\n 46|  adapters/                      registered adapter manifests / pointers\n 47|    <adapter>/web/               core-owned adapter served root; output bytes survive deactivation\n...\n 173|| subnet registry (snapshot + distribution) | **BranchStore** | audited; rides the precedence merge driver instead of a separate replication scheme (M4) |\n*174|| daemon/brain coarse checkpoint (rehydrate anchor, ADR-0004) | **BranchStore** | commit = checkpoint, ref tip = resume |\n*175|| endpoint lifecycle / status *audit trail* | **BranchStore** (or a derived view) | the *trajectory* matters, not just current value |\n 176|| spool (queue, drain-by-status, `id→address`) | **SQLite** | indexed queries + hot churn |\n 177|| broker effect journal (per-PTY-effect, B5) | **fsync-append** | hot path, per-keystroke — a commit per write is far too heavy |\n 178|| `info.json` / `status` (current value) | flat + atomic-rename (today) | tiny, read-hot; *may* move under BranchStore later, low priority |\n...\n 185|\n*186|**Reading the effect journal (`<spt_home>/daemon-effects.log`) — operator note.** The broker fsync-appends two line kinds per durable effect: `PENDING <key> <kind>` at attempt, `DONE <key>` on success (durable kinds only — an ephemeral `PtyWrite` writes no lines). A line with **`PENDING` and no matching `DONE`** is **expected**, not a leak: it is a durable effect that did NOT complete — for a `net-send` (a WAN write to an offline/unreachable peer) it is an **unconfirmed send retained as the re-driv...\n 187|\n 188|**Delivery vs propagation — latency envelope (operator note).** Two classes of cross-node traffic have deliberately different latency floors, so a minutes-scale delay is not automatically a defect: (1) **messages** (`spt send`) are **event-driven** — delivered on the events the daemon owns (WAN ingress + the ACTIVE→IDLE edge, F-023), so a healthy path is near-immediate and never waits on an adapter hook-poll cadence; (2) **anti-entropy context-sync** (the mind pull) is **periodic** — it rides the...\n 189|\n...\n 274|  one refresh. Fire-and-forget and daemon-offline-safe (events accumulate\n*275|  and are consumed at the next brain start).\n 276|  <!-- [doc->REQ-PROJECT-INDEX-INVALIDATION] -->\n 277|\n 278|## perches/ — node-local runtime\n...\n 312|","truncated":true,"truncatedBy":"bytes","totalLines":397,"totalBytes":57610,"outputLines":346,"outputBytes":50908,"lastLinePartial":false,"firstLineExceedsLimit":false},"linesTruncated":true,"meta":{"limits":{"columnTruncated":{"maxColumn":512}},"truncation":{"direction":"head","truncatedBy":"bytes","totalLines":397,"totalBytes":57610,"outputLines":346,"outputBytes":50908,"shownRange":{"start":1,"end":346},"nextOffset":347}}},"isError":false,"timestamp":1788999513860}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_l9O7zIiPJselAVqRlROjBVHy|fc_04ba494197989c86016aa1f765fe8087d1bebdabb56e4b7334","toolName":"grep","startedAt":"2026-09-10T00:18:46.071Z","args":{"path":".worktrees/286-churn/crates/spt-daemon/src/brain.rs;.worktrees/286-churn/crates/spt-daemon/src/broker.rs;.worktrees/286-churn/crates/spt-daemon/src/transport.rs"},"intent":"Locating session request dispatch"},"id":"1b1cdb9b","parentId":"9e606103","timestamp":"2026-09-10T00:18:46.071Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_FGPiUYpnCxmu07fdnxGYa1Y8|fc_04ba494197989c86016aa1f765fe9487d18b6a1bad3b65dc2a","toolName":"read","startedAt":"2026-09-10T00:18:46.072Z","args":{"path":".worktrees/286-churn/crates/spt-daemon/tests/conn_blackhole_lifecycle.rs"},"intent":"Reading lifecycle harness"},"id":"23093480","parentId":"1b1cdb9b","timestamp":"2026-09-10T00:18:46.072Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_SU9lYmkPNdJD3f5bY6nbP7PW|fc_04ba494197989c86016aa1f765fea087d1b2e143277e6a452b","toolName":"read","startedAt":"2026-09-10T00:18:46.072Z","args":{"path":".worktrees/286-churn/crates/spt-daemon/src/conn.rs:780-1000"},"intent":"Reading connection regressions"},"id":"9cc45045","parentId":"23093480","timestamp":"2026-09-10T00:18:46.072Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_f4zLEDdKAHP55bJ7LsJfMXDW|fc_04ba494197989c86016aa1f765fea487d191759e69bdaef2b1","toolName":"grep","startedAt":"2026-09-10T00:18:46.072Z","args":{"path":".worktrees/286-churn/crates/spt-daemon/tests;.worktrees/286-churn/scripts;.worktrees/286-churn/crates/spt-daemon/src/broker.rs"},"intent":"Locating lifecycle event consumers"},"id":"2ecc8714","parentId":"9cc45045","timestamp":"2026-09-10T00:18:46.072Z"}
{"type":"message","id":"df795c38","parentId":"2ecc8714","timestamp":"2026-09-10T00:18:46.072Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":"**Inspecting persistent session state**"},{"type":"toolCall","id":"call_l9O7zIiPJselAVqRlROjBVHy|fc_04ba494197989c86016aa1f765fe8087d1bebdabb56e4b7334","name":"grep","arguments":{"i":"Locating session request dispatch","pattern":"pub fn (cold_start|sessions)|fn (request|rpc|recv)|impl Drop|KIND_SESSIONS|Role::Brain|set_read_timeout","path":".worktrees/286-churn/crates/spt-daemon/src/brain.rs;.worktrees/286-churn/crates/spt-daemon/src/broker.rs;.worktrees/286-churn/crates/spt-daemon/src/transport.rs","case":true,"gitignore":false,"skip":null},"intent":"Locating session request dispatch"},{"type":"toolCall","id":"call_FGPiUYpnCxmu07fdnxGYa1Y8|fc_04ba494197989c86016aa1f765fe9487d18b6a1bad3b65dc2a","name":"read","arguments":{"i":"Reading lifecycle harness","path":".worktrees/286-churn/crates/spt-daemon/tests/conn_blackhole_lifecycle.rs"},"intent":"Reading lifecycle harness"},{"type":"toolCall","id":"call_SU9lYmkPNdJD3f5bY6nbP7PW|fc_04ba494197989c86016aa1f765fea087d1b2e143277e6a452b","name":"read","arguments":{"i":"Reading connection regressions","path":".worktrees/286-churn/crates/spt-daemon/src/conn.rs:780-1000"},"intent":"Reading connection regressions"},{"type":"toolCall","id":"call_f4zLEDdKAHP55bJ7LsJfMXDW|fc_04ba494197989c86016aa1f765fea487d191759e69bdaef2b1","name":"grep","arguments":{"i":"Locating lifecycle event consumers","pattern":"write-start|transport-close|CONN_LIFECYCLE|BRAIN_CONN_OPEN|first_write_","path":".worktrees/286-churn/crates/spt-daemon/tests;.worktrees/286-churn/scripts;.worktrees/286-churn/crates/spt-daemon/src/broker.rs","case":true,"gitignore":false,"skip":null},"intent":"Locating lifecycle event consumers"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":27576,"output":357,"cacheRead":31488,"cacheWrite":0,"totalTokens":59421,"reasoningTokens":38,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1788999513905,"responseId":"resp_04ba494197989c86016aa1f75a2de887d1afa4a4c835a7b6f7","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_04ba494197989c86016aa1f75b4da487d1b8deaba27ce9d2c1","type":"reasoning","content":[],"encrypted_content":"gAAAAABqofdlFudjG3xNY87JK-FClZLWiH0Lpzk-kFBUomYkWE8FEE2H3lzMSiG3ADe0XNK7UgC6kLQAhyznWYqys3eDuWqfcTAuFfk6lH7WJSaKxmCMhoOFaBAoB_g-BqwOsGErt8cQB9FkCNK-2X_K7dO20wAra4Im7Su77kNFVHes0Y18pGnlvGFlj4RV07tExax0ULkeI__tJdp1LaAp_ryGJeZHec-euUYVeai2LHgmzT4irEyC18QA2yqmZiZG6DO3ApB_n7UeFa1czHY4iNK6MJK7CDDSLwitbbd7-jiQ1ETXno6oGOOvC7011uzyfkRYEAQTOGx-Wn9_AEUxYmEfNnXS3m5Lf8IUO7anXXkowJXaVwYVSGd_zhp-GTqD7FkXAX_z_kI8ByUFzkgURvMV1Ooy93DNCNr-H-vJbI6K4lE6UUtvaA_dWLO9jeXa0adlog8B9LLSeAHCwre79beWRTpaWcylB5AjehqJjAOdaUKeGm0-4Do0vZf9Klajd4z_J5uYdx-nJFpHVjCwkVz8VaXBHoQa-izLPe1U6GbM2IeAA74LpRmzya08xrrwdFWrQepuKkY53wOi9uWaHSLliqE7zb8idRQGQz3j2Sg7uzp5YjZPXCuCUTnbpT0n4YcEegSKWKg0RPkMR5R6CXW00dLeN914jdowm_DI5KxiUqlO7EF1sJHbxmVtpN58GiFW4qGGE9vxGTKSPO540pKSdqor1EYVAgRrqJmo5qadhPwXI_AoFm3VsdyYNiWmMdeIxbduJTuG6msRno4WQa-D-MH-eL20eHJ1Zv4G0Vosanh4HNOh_pViChvrkuwYVasmKT5tgaXTVU6T9yg9gHm5jMRHc13_s8ryD3yZrJkxaZlUzBr8ZL3CNr2-tyMi9s86X4te6rrEtK2kltnTdN2EDMlv5GjTCzFDJCg-i1cgZtYhAZJkwT1k7Csmlj8HcYuGkUJJrY_j3p5MXv09wR0fvRhXL2eR7TQUAHMivGzR6fgg5opV4vhSp6zcL9_ZcxwfF_LgdQotal7mE6fVp3RYL5eFDj2Ip5_qjkS_s5WqzreHJZyhO2E_JUzEyh-ScwjQFyQAS4BCODu1SY1qVD1ABeDGaPEpxsB0Jx1XvsKW7CV4fmu5kzLrCteK-LaaveiMH8rs1QB-FrhlcbO7LN_v1ate-z7UVlRi42p7nNIyTyfFZl_XOnC3QYTins-2Fb7L0bJAPFe2-h6zsla3nc6rcRg0gB11As-yZSDn4hKmU4kta_IRIzEXcIqy67eVq2bgnpHYybFc7zeK_2Wltl9opeC3s2kq-EnG_zoAMbNprGRfH0FWsr1aEcGsi3yn81z88EE2gdfe_n-wpzhCozs9nInV1oKKN22-kUIZGVtG2M_WZs7NMLg4TtXhq1XMyDtbsYi5wBM6EPV18CBCbervWJnGekGq2YGk2y4ExKcDlNqBPI7wx3zqK3FVAyU81hB-mZAviu4gdLaNbgcK84OFMtJP99AQu1OigSCseZdGv583Zf-f1RkiNE18HjbYVAkFysF25Ss3uvoIjZHBVXs98ITTVP0RPdfvbI_ynT3SWvXz_-I=","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"summary":[{"type":"summary_text","text":"**Inspecting persistent session state**"}],"metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f765fe8087d1bebdabb56e4b7334","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating session request dispatch\",\"pattern\":\"pub fn (cold_start|sessions)|fn (request|rpc|recv)|impl Drop|KIND_SESSIONS|Role::Brain|set_read_timeout\",\"path\":\".worktrees/286-churn/crates/spt-daemon/src/brain.rs;.worktrees/286-churn/crates/spt-daemon/src/broker.rs;.worktrees/286-churn/crates/spt-daemon/src/transport.rs\",\"case\":true,\"gitignore\":false,\"skip\":null}","call_id":"call_l9O7zIiPJselAVqRlROjBVHy","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"grep","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f765fe9487d18b6a1bad3b65dc2a","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading lifecycle harness\",\"path\":\".worktrees/286-churn/crates/spt-daemon/tests/conn_blackhole_lifecycle.rs\"}","call_id":"call_FGPiUYpnCxmu07fdnxGYa1Y8","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"read","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f765fea087d1b2e143277e6a452b","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading connection regressions\",\"path\":\".worktrees/286-churn/crates/spt-daemon/src/conn.rs:780-1000\"}","call_id":"call_SU9lYmkPNdJD3f5bY6nbP7PW","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"read","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f765fea487d191759e69bdaef2b1","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating lifecycle event consumers\",\"pattern\":\"write-start|transport-close|CONN_LIFECYCLE|BRAIN_CONN_OPEN|first_write_\",\"path\":\".worktrees/286-churn/crates/spt-daemon/tests;.worktrees/286-churn/scripts;.worktrees/286-churn/crates/spt-daemon/src/broker.rs\",\"case\":true,\"gitignore\":false,\"skip\":null}","call_id":"call_f4zLEDdKAHP55bJ7LsJfMXDW","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"grep","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}}]},"duration":12165.225899999961,"ttft":1437.8608999999706,"completedAt":1788999526070,"contextSnapshot":{"promptTokens":59064,"nonMessageTokens":13648,"compactionEpoch":0}}}
{"type":"message","id":"470c0610","parentId":"df795c38","timestamp":"2026-09-10T00:18:46.082Z","message":{"role":"toolResult","toolCallId":"call_SU9lYmkPNdJD3f5bY6nbP7PW|fc_04ba494197989c86016aa1f765fea087d1b2e143277e6a452b","toolName":"read","content":[{"type":"text","text":"fn poisoned_err() -> io::Error {\n    io::Error::new(\n…\n    )\n}\n\nfn timeout_err() -> io::Error {\n    io::Error::new(\n        io::ErrorKind::TimedOut,\n        \"broker conn write exceeded its bound (conn poisoned + aborted)\",\n    )\n}\n\n#[cfg(test)]\nmod tests {\n    use super::*;\n    use crate::transport::{DaemonTransport, LocalSocketTransport};\n    use interprocess::local_socket::traits::Stream as _;\n    use interprocess::local_socket::Stream;\n    use std::sync::atomic::AtomicU32;\n\n    /// A real (loopback local-socket) conn pair — the send half wrapped in the\n    /// primitive, the client end returned so a test can hold it alive or drop\n    /// it to force an ORGANIC write failure.\n    fn make_conn() -> (BrokerConn, Stream) {\n        make_conn_with_log(Arc::new(LifecycleLog::new()))\n    }\n\n    fn make_conn_with_log(log: Arc<LifecycleLog>) -> (BrokerConn, Stream) {\n        let (conn, client, _recv) = make_duplex_with_log(log);\n        (conn, client)\n    }\n\n    fn make_duplex_with_log(\n        log: Arc<LifecycleLog>,\n    ) -> (BrokerConn, Stream, interprocess::local_socket::RecvHalf) {\n        static SEQ: AtomicU32 = AtomicU32::new(0);\n        let name = format!(\n            \"spt-daemon-connattr-{}-{}.sock\",\n            std::process::id(),\n            SEQ.fetch_add(1, Ordering::Relaxed)\n        );\n        let listener = LocalSocketTransport::bind(&name).expect(\"bind\");\n        let client = LocalSocketTransport::connect(&name).expect(\"connect\");\n        let server = listener.accept().expect(\"accept\");\n        let (recv, send) = server.split();\n        (\n            BrokerConn::with_log(send, Duration::from_millis(2000), log),\n            client,\n            recv,\n        )\n    }\n\n    // [unit->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n    #[test]\n    fn healthy_round_trips_and_reconnections_emit_one_interval_summary() {\n        let now = Instant::now();\n        let log = Arc::new(LifecycleLog::captured(now));\n        // The real framed-write/drop path, both a persistent carrier and fresh\n        // per-poll carriers. An isolated clock/sink avoids global test races.\n        const N: u64 = 32;\n        for round in 0..N {\n            let (conn, mut client, mut recv) = make_duplex_with_log(Arc::clone(&log));\n            conn.describe(\"role=brain\");\n            for seq in 0..4 {\n                let request = Envelope::new(\"poll\", serde_json::json!([round, seq]));\n                write_frame(&mut client, &request).unwrap();\n                let received = crate::codec::read_frame(&mut recv).unwrap();\n                assert_eq!(received.payload, request.payload);\n                assert_eq!(conn.write(&request).unwrap(), ConnWrite::Done);\n                let reply = crate::codec::read_frame(&mut client).unwrap();\n                assert_eq!(reply.payload, request.payload);\n            }\n            drop(conn);\n        }\n        assert!(recover(log.capture.as_ref().unwrap()).is_empty());\n        *recover(log.clock.as_ref().unwrap()) = now + HEALTHY_LOG_INTERVAL;\n        let (conn, mut client) = make_conn_with_log(Arc::clone(&log));\n        conn.describe(\"role=brain\");\n        conn.write(&Envelope::new(\"poll\", serde_json::json!(N))).unwrap();\n        assert_eq!(crate::codec::read_frame(&mut client).unwrap().payload, serde_json::json!(N));\n        drop(conn);\n        let lines = recover(log.capture.as_ref().unwrap());\n        assert_eq!(lines.len(), 1, \"healthy volume is per interval, not per conn\");\n        assert!(lines[0].contains(&format!(\"starts={}\", N + 1)));\n        assert!(lines[0].contains(&format!(\"closes={N}\")));\n    }\n\n    // [unit->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n    // [unit->REQ-CONN-POISON-ATTRIBUTION]\n    #[test]\n    fn failures_and_role_events_bypass_healthy_aggregation_with_original_start() {\n        let log = Arc::new(LifecycleLog::captured(Instant::now()));\n        let (conn, mut client) = make_conn_with_log(Arc::clone(&log));\n        conn.describe(\"role=brain controller session=7\");\n        conn.write(&Envelope::new(\"poll\", serde_json::json!({}))).unwrap();\n        crate::codec::read_frame(&mut client).unwrap();\n        let start = *conn.inner.first_write.get().expect(\"first committed write\");\n        conn.lifecycle_event(\"writer-exit\", \"reason=write-failed\");\n        conn.inner.poison_and_cancel(Some(&io::Error::new(io::ErrorKind::BrokenPipe, \"peer gone\")));\n        // Repeated cancellation must not repeat the retirement record.\n        conn.inner.poison_and_cancel(None);\n        let id = conn.id();\n        drop(conn);\n        let lines = recover(log.capture.as_ref().unwrap());\n        assert_eq!(lines.len(), 3, \"role event, retirement and poisoned close survive\");\n        for line in lines.iter() {\n            assert!(line.contains(&format!(\"conn={id}\")));\n            assert!(line.contains(\"role=brain controller session=7\"));\n            assert!(line.contains(&format!(\"first_write_wall_ms={}\", start.wall_ms)));\n            assert!(line.contains(&format!(\"first_write_mono_ms={}\", start.mono_ms)));\n        }\n        assert!(lines[1].contains(\"CONN_WRITE_RETIRED:\"));\n        assert!(lines[1].contains(\"BrokenPipe\"));\n        assert!(lines[2].contains(\"poisoned=true timed_out=false\"));\n    }\n\n    // [unit->REQ-CONN-POISON-DIAL-SCOPE] the F-039 leg-(a) token split: the loud\n    // CONN_WRITE_POISONED token is RESERVED for the deadline (timed_out) class;\n    // an organic write failure renders the distinct CONN_WRITE_RETIRED token\n    // carrying the error kind — and (doyle amendment 1) the SAME attribution\n    // fields as POISONED: conn id, accreted role/endpoint/session facts, wall +\n    // monotonic stamps. Quieter class, never thinner.\n    // [unit->REQ-CONN-POISON-ATTRIBUTION]\n    #[test]\n    fn retirement_token_splits_deadline_from_organic_with_attribution_parity() {\n        let (conn, _client) = make_conn();\n        conn.describe(\"role=brain\");\n        conn.describe(\"controller session=7 endpoint=ep by=cli@host\");\n\n        let poisoned = conn.inner.render_retirement(true, None);\n        assert!(\n            poisoned.starts_with(\"CONN_WRITE_POISONED:\"),\n            \"deadline class keeps the loud wedge-observable token: {poisoned}\"\n        );\n        let retired = conn\n            .inner\n            .render_retirement(false, Some(&io::Error::new(io::ErrorKind::BrokenPipe, \"peer gone\")));\n        assert!(\n            retired.starts_with(\"CONN_WRITE_RETIRED:\"),\n            \"organic class emits the DISTINCT retired token: {retired}\"\n        );\n        assert!(\n            retired.contains(\"BrokenPipe\") && retired.contains(\"peer gone\"),\n            \"the organic record carries the error kind + message: {retired}\"\n        );\n        // Attribution parity: every field present on BOTH tokens.\n        for (name, line) in [(\"POISONED\", &poisoned), (\"RETIRED\", &retired)] {\n            assert!(\n                line.contains(&format!(\"conn={}\", conn.id())),\n                \"{name} carries the stable conn id: {line}\"\n            );\n            assert!(\n                line.contains(\"role=brain\") && line.contains(\"controller session=7\"),\n                \"{name} carries the accreted role/session facts: {line}\"\n            );\n            assert!(\n                line.contains(\"wall_ms=\") && line.contains(\"mono_ms=\"),\n                \"{name} carries wall + monotonic stamps: {line}\"\n            );\n        }\n    }\n\n    // [unit->REQ-CONN-POISON-DIAL-SCOPE] behavior is UNCHANGED by the token\n    // split (the REQ-HAZARD-SHAREDSEND invariant): an organic fast-fail still\n    // poisons + retires the conn — it is only the CLASS that differs (the error\n    // surfaces with its own kind, never TimedOut, and the next write refuses).\n    #[test]\n    fn organic_write_failure_still_poisons_and_is_not_the_timeout_class() {\n        let (conn, client) = make_conn();\n        drop(client); // peer dies → writes fail organically (no deadline involved)\n        let env = Envelope::new(\"attr-test\", serde_json::json!({\"n\": 1}));\n        let mut organic = None;\n        // A first write can land in the OS buffer before the peer-death is\n        // observable; loop until the failure surfaces (bounded).\n        for _ in 0..200 {\n            match conn.write(&env) {\n                Err(e) => {\n                    organic = Some(e);\n                    break;\n                }\n                Ok(_) => std::thread::yield_now(),\n            }\n        }\n        let e = organic.expect(\"a write to a dead peer eventually fails\");\n        assert_ne!(\n            e.kind(),\n            io::ErrorKind::TimedOut,\n            \"an organic failure must NOT surface as the deadline class\"\n        );\n        // Poisoned + retired: the next write refuses immediately.\n        let refused = conn.write(&env).expect_err(\"poisoned conn refuses writes\");\n        assert_eq!(refused.kind(), io::ErrorKind::BrokenPipe);\n    }\n\n    // [unit->REQ-CONN-POISON-ATTRIBUTION] the stable conn id is unique across\n    // conns (monotonic mint — Arc::ptr_eq identity does not survive a log\n    // line), and the accreted label is BOUNDED: one truncation marker at\n    // LABEL_CAP, then no further growth.\n    #[test]\n    fn conn_ids_are_unique_and_labels_stay_bounded() {\n        let (a, _ca) = make_conn();\n        let (b, _cb) = make_conn();\n        assert_ne!(a.id(), b.id(), \"each physical conn mints a distinct id\");\n\n        for i in 0..100 {\n            a.describe(&format!(\"fact-{i}-padding-padding-padding\"));\n        }\n        let label = recover(&a.inner.label).clone();\n        assert!(\n            label.len() <= LABEL_CAP + '…'.len_utf8(),\n            \"label growth is bounded at LABEL_CAP (+1 marker), got {} bytes\",\n            label.len()\n        );\n        assert!(\n            label.ends_with('…'),\n            \"the bound is marked once with an ellipsis\"\n        );\n        // The id (and the capped label) stay stable across records.\n        let one = a.inner.render_retirement(true, None);\n        let two = a.inner.render_retirement(true, None);\n        assert!(one.contains(&format!(\"conn={}\", a.id())));\n        assert!(two.contains(&format!(\"conn={}\", a.id())));\n    }\n}"}],"details":{"totalLines":1000,"displayContent":{"text":"fn poisoned_err() -> io::Error {\n    io::Error::new(\n…\n    )\r\n}\r\n\r\nfn timeout_err() -> io::Error {\r\n    io::Error::new(\r\n        io::ErrorKind::TimedOut,\r\n        \"broker conn write exceeded its bound (conn poisoned + aborted)\",\r\n    )\r\n}\r\n\r\n#[cfg(test)]\r\nmod tests {\r\n    use super::*;\r\n    use crate::transport::{DaemonTransport, LocalSocketTransport};\r\n    use interprocess::local_socket::traits::Stream as _;\r\n    use interprocess::local_socket::Stream;\r\n    use std::sync::atomic::AtomicU32;\r\n\r\n    /// A real (loopback local-socket) conn pair — the send half wrapped in the\r\n    /// primitive, the client end returned so a test can hold it alive or drop\r\n    /// it to force an ORGANIC write failure.\r\n    fn make_conn() -> (BrokerConn, Stream) {\r\n        make_conn_with_log(Arc::new(LifecycleLog::new()))\r\n    }\r\n\r\n    fn make_conn_with_log(log: Arc<LifecycleLog>) -> (BrokerConn, Stream) {\r\n        let (conn, client, _recv) = make_duplex_with_log(log);\r\n        (conn, client)\r\n    }\r\n\r\n    fn make_duplex_with_log(\r\n        log: Arc<LifecycleLog>,\r\n    ) -> (BrokerConn, Stream, interprocess::local_socket::RecvHalf) {\r\n        static SEQ: AtomicU32 = AtomicU32::new(0);\r\n        let name = format!(\r\n            \"spt-daemon-connattr-{}-{}.sock\",\r\n            std::process::id(),\r\n            SEQ.fetch_add(1, Ordering::Relaxed)\r\n        );\r\n        let listener = LocalSocketTransport::bind(&name).expect(\"bind\");\r\n        let client = LocalSocketTransport::connect(&name).expect(\"connect\");\r\n        let server = listener.accept().expect(\"accept\");\r\n        let (recv, send) = server.split();\r\n        (\r\n            BrokerConn::with_log(send, Duration::from_millis(2000), log),\r\n            client,\r\n            recv,\r\n        )\r\n    }\r\n\r\n    // [unit->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\r\n    #[test]\r\n    fn healthy_round_trips_and_reconnections_emit_one_interval_summary() {\r\n        let now = Instant::now();\r\n        let log = Arc::new(LifecycleLog::captured(now));\r\n        // The real framed-write/drop path, both a persistent carrier and fresh\r\n        // per-poll carriers. An isolated clock/sink avoids global test races.\r\n        const N: u64 = 32;\r\n        for round in 0..N {\r\n            let (conn, mut client, mut recv) = make_duplex_with_log(Arc::clone(&log));\r\n            conn.describe(\"role=brain\");\r\n            for seq in 0..4 {\r\n                let request = Envelope::new(\"poll\", serde_json::json!([round, seq]));\r\n                write_frame(&mut client, &request).unwrap();\r\n                let received = crate::codec::read_frame(&mut recv).unwrap();\r\n                assert_eq!(received.payload, request.payload);\r\n                assert_eq!(conn.write(&request).unwrap(), ConnWrite::Done);\r\n                let reply = crate::codec::read_frame(&mut client).unwrap();\r\n                assert_eq!(reply.payload, request.payload);\r\n            }\r\n            drop(conn);\r\n        }\r\n        assert!(recover(log.capture.as_ref().unwrap()).is_empty());\r\n        *recover(log.clock.as_ref().unwrap()) = now + HEALTHY_LOG_INTERVAL;\r\n        let (conn, mut client) = make_conn_with_log(Arc::clone(&log));\r\n        conn.describe(\"role=brain\");\r\n        conn.write(&Envelope::new(\"poll\", serde_json::json!(N))).unwrap();\r\n        assert_eq!(crate::codec::read_frame(&mut client).unwrap().payload, serde_json::json!(N));\r\n        drop(conn);\r\n        let lines = recover(log.capture.as_ref().unwrap());\r\n        assert_eq!(lines.len(), 1, \"healthy volume is per interval, not per conn\");\r\n        assert!(lines[0].contains(&format!(\"starts={}\", N + 1)));\r\n        assert!(lines[0].contains(&format!(\"closes={N}\")));\r\n    }\r\n\r\n    // [unit->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\r\n    // [unit->REQ-CONN-POISON-ATTRIBUTION]\r\n    #[test]\r\n    fn failures_and_role_events_bypass_healthy_aggregation_with_original_start() {\r\n        let log = Arc::new(LifecycleLog::captured(Instant::now()));\r\n        let (conn, mut client) = make_conn_with_log(Arc::clone(&log));\r\n        conn.describe(\"role=brain controller session=7\");\r\n        conn.write(&Envelope::new(\"poll\", serde_json::json!({}))).unwrap();\r\n        crate::codec::read_frame(&mut client).unwrap();\r\n        let start = *conn.inner.first_write.get().expect(\"first committed write\");\r\n        conn.lifecycle_event(\"writer-exit\", \"reason=write-failed\");\r\n        conn.inner.poison_and_cancel(Some(&io::Error::new(io::ErrorKind::BrokenPipe, \"peer gone\")));\r\n        // Repeated cancellation must not repeat the retirement record.\r\n        conn.inner.poison_and_cancel(None);\r\n        let id = conn.id();\r\n        drop(conn);\r\n        let lines = recover(log.capture.as_ref().unwrap());\r\n        assert_eq!(lines.len(), 3, \"role event, retirement and poisoned close survive\");\r\n        for line in lines.iter() {\r\n            assert!(line.contains(&format!(\"conn={id}\")));\r\n            assert!(line.contains(\"role=brain controller session=7\"));\r\n            assert!(line.contains(&format!(\"first_write_wall_ms={}\", start.wall_ms)));\r\n            assert!(line.contains(&format!(\"first_write_mono_ms={}\", start.mono_ms)));\r\n        }\r\n        assert!(lines[1].contains(\"CONN_WRITE_RETIRED:\"));\r\n        assert!(lines[1].contains(\"BrokenPipe\"));\r\n        assert!(lines[2].contains(\"poisoned=true timed_out=false\"));\r\n    }\r\n\r\n    // [unit->REQ-CONN-POISON-DIAL-SCOPE] the F-039 leg-(a) token split: the loud\r\n    // CONN_WRITE_POISONED token is RESERVED for the deadline (timed_out) class;\r\n    // an organic write failure renders the distinct CONN_WRITE_RETIRED token\r\n    // carrying the error kind — and (doyle amendment 1) the SAME attribution\r\n    // fields as POISONED: conn id, accreted role/endpoint/session facts, wall +\r\n    // monotonic stamps. Quieter class, never thinner.\r\n    // [unit->REQ-CONN-POISON-ATTRIBUTION]\r\n    #[test]\r\n    fn retirement_token_splits_deadline_from_organic_with_attribution_parity() {\r\n        let (conn, _client) = make_conn();\r\n        conn.describe(\"role=brain\");\r\n        conn.describe(\"controller session=7 endpoint=ep by=cli@host\");\r\n\r\n        let poisoned = conn.inner.render_retirement(true, None);\r\n        assert!(\r\n            poisoned.starts_with(\"CONN_WRITE_POISONED:\"),\r\n            \"deadline class keeps the loud wedge-observable token: {poisoned}\"\r\n        );\r\n        let retired = conn\r\n            .inner\r\n            .render_retirement(false, Some(&io::Error::new(io::ErrorKind::BrokenPipe, \"peer gone\")));\r\n        assert!(\r\n            retired.starts_with(\"CONN_WRITE_RETIRED:\"),\r\n            \"organic class emits the DISTINCT retired token: {retired}\"\r\n        );\r\n        assert!(\r\n            retired.contains(\"BrokenPipe\") && retired.contains(\"peer gone\"),\r\n            \"the organic record carries the error kind + message: {retired}\"\r\n        );\r\n        // Attribution parity: every field present on BOTH tokens.\r\n        for (name, line) in [(\"POISONED\", &poisoned), (\"RETIRED\", &retired)] {\r\n            assert!(\r\n                line.contains(&format!(\"conn={}\", conn.id())),\r\n                \"{name} carries the stable conn id: {line}\"\r\n            );\r\n            assert!(\r\n                line.contains(\"role=brain\") && line.contains(\"controller session=7\"),\r\n                \"{name} carries the accreted role/session facts: {line}\"\r\n            );\r\n            assert!(\r\n                line.contains(\"wall_ms=\") && line.contains(\"mono_ms=\"),\r\n                \"{name} carries wall + monotonic stamps: {line}\"\r\n            );\r\n        }\r\n    }\r\n\r\n    // [unit->REQ-CONN-POISON-DIAL-SCOPE] behavior is UNCHANGED by the token\r\n    // split (the REQ-HAZARD-SHAREDSEND invariant): an organic fast-fail still\r\n    // poisons + retires the conn — it is only the CLASS that differs (the error\r\n    // surfaces with its own kind, never TimedOut, and the next write refuses).\r\n    #[test]\r\n    fn organic_write_failure_still_poisons_and_is_not_the_timeout_class() {\r\n        let (conn, client) = make_conn();\r\n        drop(client); // peer dies → writes fail organically (no deadline involved)\r\n        let env = Envelope::new(\"attr-test\", serde_json::json!({\"n\": 1}));\r\n        let mut organic = None;\r\n        // A first write can land in the OS buffer before the peer-death is\r\n        // observable; loop until the failure surfaces (bounded).\r\n        for _ in 0..200 {\r\n            match conn.write(&env) {\r\n                Err(e) => {\r\n                    organic = Some(e);\r\n                    break;\r\n                }\r\n                Ok(_) => std::thread::yield_now(),\r\n            }\r\n        }\r\n        let e = organic.expect(\"a write to a dead peer eventually fails\");\r\n        assert_ne!(\r\n            e.kind(),\r\n            io::ErrorKind::TimedOut,\r\n            \"an organic failure must NOT surface as the deadline class\"\r\n        );\r\n        // Poisoned + retired: the next write refuses immediately.\r\n        let refused = conn.write(&env).expect_err(\"poisoned conn refuses writes\");\r\n        assert_eq!(refused.kind(), io::ErrorKind::BrokenPipe);\r\n    }\r\n\r\n    // [unit->REQ-CONN-POISON-ATTRIBUTION] the stable conn id is unique across\r\n    // conns (monotonic mint — Arc::ptr_eq identity does not survive a log\r\n    // line), and the accreted label is BOUNDED: one truncation marker at\r\n    // LABEL_CAP, then no further growth.\r\n    #[test]\r\n    fn conn_ids_are_unique_and_labels_stay_bounded() {\r\n        let (a, _ca) = make_conn();\r\n        let (b, _cb) = make_conn();\r\n        assert_ne!(a.id(), b.id(), \"each physical conn mints a distinct id\");\r\n\r\n        for i in 0..100 {\r\n            a.describe(&format!(\"fact-{i}-padding-padding-padding\"));\r\n        }\r\n        let label = recover(&a.inner.label).clone();\r\n        assert!(\r\n            label.len() <= LABEL_CAP + '…'.len_utf8(),\r\n            \"label growth is bounded at LABEL_CAP (+1 marker), got {} bytes\",\r\n            label.len()\r\n        );\r\n        assert!(\r\n            label.ends_with('…'),\r\n            \"the bound is marked once with an ellipsis\"\r\n        );\r\n        // The id (and the capped label) stay stable across records.\r\n        let one = a.inner.render_retirement(true, None);\r\n        let two = a.inner.render_retirement(true, None);\r\n        assert!(one.contains(&format!(\"conn={}\", a.id())));\r\n        assert!(two.contains(&format!(\"conn={}\", a.id())));\r\n    }\r\n}\r","startLine":775,"lineNumbers":[775,776,null,779,780,781,782,783,784,785,786,787,788,789,790,791,792,793,794,795,796,797,798,799,800,801,802,803,804,805,806,807,808,809,810,811,812,813,814,815,816,817,818,819,820,821,822,823,824,825,826,827,828,829,830,831,832,833,834,835,836,837,838,839,840,841,842,843,844,845,846,847,848,849,850,851,852,853,854,855,856,857,858,859,860,861,862,863,864,865,866,867,868,869,870,871,872,873,874,875,876,877,878,879,880,881,882,883,884,885,886,887,888,889,890,891,892,893,894,895,896,897,898,899,900,901,902,903,904,905,906,907,908,909,910,911,912,913,914,915,916,917,918,919,920,921,922,923,924,925,926,927,928,929,930,931,932,933,934,935,936,937,938,939,940,941,942,943,944,945,946,947,948,949,950,951,952,953,954,955,956,957,958,959,960,961,962,963,964,965,966,967,968,969,970,971,972,973,974,975,976,977,978,979,980,981,982,983,984,985,986,987,988,989,990,991,992,993,994,995,996,997,998,999,1000]},"fileSize":43321,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-daemon\\src\\conn.rs"}}},"isError":false,"timestamp":1788999526081}}
{"type":"message","id":"648734f9","parentId":"470c0610","timestamp":"2026-09-10T00:18:46.087Z","message":{"role":"toolResult","toolCallId":"call_FGPiUYpnCxmu07fdnxGYa1Y8|fc_04ba494197989c86016aa1f765fe9487d18b6a1bad3b65dc2a","toolName":"read","content":[{"type":"text","text":"//! MSG-IDENTITY W6 / F-039 leg (e) — the DETERMINISTIC BLACK-HOLED-CONTROLLER\n//! LIFECYCLE HARNESS (REQ-CONN-BLACKHOLE-LIFECYCLE-HARNESS), on a REAL broker\n//! hosting REAL PTY children over the real local IPC.\n//!\n//! ## What this is\n//!\n//! This is hertz's RCA fix-shape item 5 made executable (`.claude/\n//! hertz-pty-freeze-rca.md` §Proposed fix shape, item 5): the standing\n//! conformance rig for the round-4 `SharedSend` fix class. It reproduces the\n//! EXACT Windows named-pipe black-hole (a controller that stops draining while a\n//! finite burst floods its socket, so the broker's writer parks inside\n//! `write_frame` on `WriteFileEx`+`SleepEx(INFINITE, alertable)` — no supported\n//! write timeout) and, against CURRENT v0.30.6 semantics (`crates/spt-daemon/\n//! src/conn.rs` bounded/cancelable/poison-on-failure writes, landed fc048bb),\n//! asserts hertz's FIVE invariants for ONE incident:\n//!\n//!   1. UNRELATED SESSIONS CONTINUE — a second, independent session's viewer keeps\n//!      receiving output throughout the black-hole incident (the fan-out never\n//!      blocks the whole broker on the one wedged conn);\n//!   2. THE BAD PHYSICAL CONNECTION IS CANCELED/CLOSED WITHIN THE BOUND — a write\n//!      on the retired old-client connection FAILS, and the conn's `handle_conn`\n//!      thread finishes within the write bound of logical release (the conn was\n//!      physically retired, not merely logically stall-evicted);\n//!   3. ITS WRITER EXITS — `broker.conn_handler_count()` falls to the\n//!      unrelated-viewer floor (<= 1; the persistent invariant-1 viewer holds one\n//!      handler) WITHOUT the black-holed client ever dropping its end (the parked\n//!      writer was aborted + the reader EOF-cleaned, so the handler thread ran to\n//!      completion), and the viewer flows output AFTER the sample latches — proving\n//!      the survivor is the viewer's handler, not a leaked one;\n//!   4. A FRESH VIEWER CAN ATTACH — a brand-new Viewer subscribe on the retired\n//!      session COMPLETES and RECEIVES replayed output from the frozen cursor\n//!      (the fan-out accepts + serves a new subscriber after retirement);\n//!   5. NO LOCK OR TASK REMAINS OWNED BY THE RETIRED CONNECTION — asserted via\n//!      PROXIES (this invariant has no direct seam): after retirement the broker\n//!      still answers a control-plane `KIND_SESSIONS` request PROMPTLY (no lock\n//!      is held by the retired writer), `conn_handler_count()` returns to the\n//!      unrelated-viewer floor without the client dropping its end (no task\n//!      lingers), AND both the\n//!      unrelated session's viewer and the fresh viewer flow output (the shared\n//!      fan-out lock is free). The conjunction of these three is the strongest\n//!      deterministic proxy available for \"nothing is still owned.\"\n//!\n//! ## Relationship to `brain_decouple.rs`\n//!\n//! This harness GENERALIZES `brain_decouple.rs`'s r4 gate\n//! (`non_draining_controller_stall_evict_releases_writer_and_connection`, the\n//! Windows-mandatory physical-retirement leg for\n//! REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK, which stays the primary\n//! gate for that hazard). It reuses the same SEED/trigger/burst choreography,\n//! the hand-back-the-unread-Stream pattern, the `conn_handler_count` and\n//! `stall_evict_count` seams, the old-client-write-fails probe, and the\n//! teardown-before-assert discipline. It ADDS the two pieces the r4 gate does\n//! not have, which are exactly hertz's invariants (1) and (4): a SECOND\n//! UNRELATED session whose viewer keeps ticking across the incident, and a fresh\n//! VIEWER attach (not just a controller take) proven to receive output.\n//!\n//! ## Why Windows-mandatory (`#![cfg(windows)]` — the whole file)\n//!\n//! Invariants 2, 3, and the task-ownership proxy of 5 are the PHYSICAL-retirement\n//! asserts, and the deterministic black-hole only reproduces on the Windows\n//! named-pipe path (`interprocess 2.4.2` → `WriteFileEx`+`SleepEx(INFINITE)`,\n//! `set_timeout` → `Unsupported`) — the exact field failure. On Unix the\n//! poison/retire path is a `shutdown(SHUT_RDWR)` abort of a buffer-fill block: a\n//! genuinely DIFFERENT mechanism, not this incident. Faking a Unix companion\n//! would test a different thing, so — mirroring the r4 gate's own\n//! `#[cfg(windows)]` leg — the entire harness is Windows-only. On other targets\n//! this test binary compiles to zero tests (no fakery).\n//!\n//! ## RUNNER CONTRACT — run process-per-test (nextest) or single-threaded\n//!\n//! Shares ONE process-global `SPT_HOME` + `SPT_BRAIN_WRITE_DEADLINE_MS`, spins a\n//! REAL broker + TWO PTY children. HEAVY class. Run isolated with\n//! `cargo nextest run -p spt-daemon --test conn_blackhole_lifecycle` (CI's path)\n//! or `cargo test --test conn_blackhole_lifecycle -- --test-threads=1`. NEVER a\n//! bare `cargo test --test conn_blackhole_lifecycle` on Windows. Every read in\n//! this test rides a bounded thread+channel or a polled deadline; the test\n//! captures plain values as it goes and only asserts AFTER teardown, so it can\n//! never hang even on broken code.\n#![cfg(windows)]\n\nuse std::sync::atomic::{AtomicU32, AtomicU64, Ordering};\n…\nuse spt_test_support::TestHome;\n\nstatic SEQ: AtomicU32 = AtomicU32::new(0);\nfn unique_name() -> String { … }\n\nfn init_home() -> TestHome {\n    TestHome::new()\n}\n\nfn kill_pid(pid: u32) { … }\n\nfn send(conn: &mut Stream, kind: &str, payload: serde_json::Value) {\n    write_frame(conn, &Envelope::new(kind, payload)).expect(\"send frame\");\n}\n\nfn connect(name: &str) -> Stream { … }\n\n/// A child that FLOODS stdout forever (never reads stdin) — drives the UNRELATED\n/// session so its viewer has a continuous frame stream to keep ticking on\n/// throughout the black-hole incident (invariant 1).\nfn flood_spawn_req(endpoint: &str) -> SpawnReq { … }\n\n/// A real PTY child (PowerShell) that emits `SEED`, parks on a first `Read-Host`,\n/// then on a stdin trigger emits a finite ~1.4 MiB burst (1400 lines × 1023 `B`s)\n/// and parks again. The seq-0 boot output gives a real `seq>0` controller frame;\n/// the burst floods a non-draining controller's socket so its writer parks INSIDE\n/// `write_frame` — the exact black-hole the RCA pinned. Verbatim sibling of the\n/// r4 gate's `seed_burst_spawn_req` in `brain_decouple.rs`.\nfn seed_burst_spawn_req(endpoint: &str) -> SpawnReq { … }\n\n/// Spawn a PTY child from a fresh conn and return `(spawn_conn, session_id)`.\nfn spawn_session(name: &str, req: SpawnReq) -> (Stream, u64) { … }\n\n/// Query the broker's stall-evict tally. Bounded thread + channel (a wedge\n/// surfaces as a timeout, not a hang). Verbatim sibling of the r4 gate helper.\nfn stall_evict_count(name: &str, deadline: Duration) -> u64 { … }\n\n/// One `KIND_SESSIONS` round-trip, returning the reply. The dispatch runs the\n/// `reap_dead_controller` → time-based `stall_evict_controller`, so polling this\n/// both DRIVES the logical eviction and reads the frozen `resume_seq` cursor.\n/// Bounded thread + channel. Verbatim sibling of the r4 gate helper.\nfn query_sessions(name: &str, deadline: Duration) -> Option<SessionsReply> { … }\n\n// [int->REQ-CONN-BLACKHOLE-LIFECYCLE-HARNESS]\n#[test]\nfn blackholed_controller_lifecycle_five_invariants() { … }\n\n[…571ln elided; re-read needed ranges, e.g. .worktrees/286-churn/crates/spt-daemon/tests/conn_blackhole_lifecycle.rs:82-97,101-104]"}],"details":{"displayContent":{"text":"//! MSG-IDENTITY W6 / F-039 leg (e) — the DETERMINISTIC BLACK-HOLED-CONTROLLER\n//! LIFECYCLE HARNESS (REQ-CONN-BLACKHOLE-LIFECYCLE-HARNESS), on a REAL broker\n//! hosting REAL PTY children over the real local IPC.\n//!\n//! ## What this is\n//!\n//! This is hertz's RCA fix-shape item 5 made executable (`.claude/\n//! hertz-pty-freeze-rca.md` §Proposed fix shape, item 5): the standing\n//! conformance rig for the round-4 `SharedSend` fix class. It reproduces the\n//! EXACT Windows named-pipe black-hole (a controller that stops draining while a\n//! finite burst floods its socket, so the broker's writer parks inside\n//! `write_frame` on `WriteFileEx`+`SleepEx(INFINITE, alertable)` — no supported\n//! write timeout) and, against CURRENT v0.30.6 semantics (`crates/spt-daemon/\n//! src/conn.rs` bounded/cancelable/poison-on-failure writes, landed fc048bb),\n//! asserts hertz's FIVE invariants for ONE incident:\n//!\n//!   1. UNRELATED SESSIONS CONTINUE — a second, independent session's viewer keeps\n//!      receiving output throughout the black-hole incident (the fan-out never\n//!      blocks the whole broker on the one wedged conn);\n//!   2. THE BAD PHYSICAL CONNECTION IS CANCELED/CLOSED WITHIN THE BOUND — a write\n//!      on the retired old-client connection FAILS, and the conn's `handle_conn`\n//!      thread finishes within the write bound of logical release (the conn was\n//!      physically retired, not merely logically stall-evicted);\n//!   3. ITS WRITER EXITS — `broker.conn_handler_count()` falls to the\n//!      unrelated-viewer floor (<= 1; the persistent invariant-1 viewer holds one\n//!      handler) WITHOUT the black-holed client ever dropping its end (the parked\n//!      writer was aborted + the reader EOF-cleaned, so the handler thread ran to\n//!      completion), and the viewer flows output AFTER the sample latches — proving\n//!      the survivor is the viewer's handler, not a leaked one;\n//!   4. A FRESH VIEWER CAN ATTACH — a brand-new Viewer subscribe on the retired\n//!      session COMPLETES and RECEIVES replayed output from the frozen cursor\n//!      (the fan-out accepts + serves a new subscriber after retirement);\n//!   5. NO LOCK OR TASK REMAINS OWNED BY THE RETIRED CONNECTION — asserted via\n//!      PROXIES (this invariant has no direct seam): after retirement the broker\n//!      still answers a control-plane `KIND_SESSIONS` request PROMPTLY (no lock\n//!      is held by the retired writer), `conn_handler_count()` returns to the\n//!      unrelated-viewer floor without the client dropping its end (no task\n//!      lingers), AND both the\n//!      unrelated session's viewer and the fresh viewer flow output (the shared\n//!      fan-out lock is free). The conjunction of these three is the strongest\n//!      deterministic proxy available for \"nothing is still owned.\"\n//!\n//! ## Relationship to `brain_decouple.rs`\n//!\n//! This harness GENERALIZES `brain_decouple.rs`'s r4 gate\n//! (`non_draining_controller_stall_evict_releases_writer_and_connection`, the\n//! Windows-mandatory physical-retirement leg for\n//! REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK, which stays the primary\n//! gate for that hazard). It reuses the same SEED/trigger/burst choreography,\n//! the hand-back-the-unread-Stream pattern, the `conn_handler_count` and\n//! `stall_evict_count` seams, the old-client-write-fails probe, and the\n//! teardown-before-assert discipline. It ADDS the two pieces the r4 gate does\n//! not have, which are exactly hertz's invariants (1) and (4): a SECOND\n//! UNRELATED session whose viewer keeps ticking across the incident, and a fresh\n//! VIEWER attach (not just a controller take) proven to receive output.\n//!\n//! ## Why Windows-mandatory (`#![cfg(windows)]` — the whole file)\n//!\n//! Invariants 2, 3, and the task-ownership proxy of 5 are the PHYSICAL-retirement\n//! asserts, and the deterministic black-hole only reproduces on the Windows\n//! named-pipe path (`interprocess 2.4.2` → `WriteFileEx`+`SleepEx(INFINITE)`,\n//! `set_timeout` → `Unsupported`) — the exact field failure. On Unix the\n//! poison/retire path is a `shutdown(SHUT_RDWR)` abort of a buffer-fill block: a\n//! genuinely DIFFERENT mechanism, not this incident. Faking a Unix companion\n//! would test a different thing, so — mirroring the r4 gate's own\n//! `#[cfg(windows)]` leg — the entire harness is Windows-only. On other targets\n//! this test binary compiles to zero tests (no fakery).\n//!\n//! ## RUNNER CONTRACT — run process-per-test (nextest) or single-threaded\n//!\n//! Shares ONE process-global `SPT_HOME` + `SPT_BRAIN_WRITE_DEADLINE_MS`, spins a\n//! REAL broker + TWO PTY children. HEAVY class. Run isolated with\n//! `cargo nextest run -p spt-daemon --test conn_blackhole_lifecycle` (CI's path)\n//! or `cargo test --test conn_blackhole_lifecycle -- --test-threads=1`. NEVER a\n//! bare `cargo test --test conn_blackhole_lifecycle` on Windows. Every read in\n//! this test rides a bounded thread+channel or a polled deadline; the test\n//! captures plain values as it goes and only asserts AFTER teardown, so it can\n//! never hang even on broken code.\n#![cfg(windows)]\n\nuse std::sync::atomic::{AtomicU32, AtomicU64, Ordering};\n…\nuse spt_test_support::TestHome;\n\nstatic SEQ: AtomicU32 = AtomicU32::new(0);\nfn unique_name() -> String { … }\n\nfn init_home() -> TestHome {\n    TestHome::new()\n}\n\nfn kill_pid(pid: u32) { … }\n\nfn send(conn: &mut Stream, kind: &str, payload: serde_json::Value) {\n    write_frame(conn, &Envelope::new(kind, payload)).expect(\"send frame\");\n}\n\nfn connect(name: &str) -> Stream { … }\n\n/// A child that FLOODS stdout forever (never reads stdin) — drives the UNRELATED\n/// session so its viewer has a continuous frame stream to keep ticking on\n/// throughout the black-hole incident (invariant 1).\nfn flood_spawn_req(endpoint: &str) -> SpawnReq { … }\n\n/// A real PTY child (PowerShell) that emits `SEED`, parks on a first `Read-Host`,\n/// then on a stdin trigger emits a finite ~1.4 MiB burst (1400 lines × 1023 `B`s)\n/// and parks again. The seq-0 boot output gives a real `seq>0` controller frame;\n/// the burst floods a non-draining controller's socket so its writer parks INSIDE\n/// `write_frame` — the exact black-hole the RCA pinned. Verbatim sibling of the\n/// r4 gate's `seed_burst_spawn_req` in `brain_decouple.rs`.\nfn seed_burst_spawn_req(endpoint: &str) -> SpawnReq { … }\n\n/// Spawn a PTY child from a fresh conn and return `(spawn_conn, session_id)`.\nfn spawn_session(name: &str, req: SpawnReq) -> (Stream, u64) { … }\n\n/// Query the broker's stall-evict tally. Bounded thread + channel (a wedge\n/// surfaces as a timeout, not a hang). Verbatim sibling of the r4 gate helper.\nfn stall_evict_count(name: &str, deadline: Duration) -> u64 { … }\n\n/// One `KIND_SESSIONS` round-trip, returning the reply. The dispatch runs the\n/// `reap_dead_controller` → time-based `stall_evict_controller`, so polling this\n/// both DRIVES the logical eviction and reads the frozen `resume_seq` cursor.\n/// Bounded thread + channel. Verbatim sibling of the r4 gate helper.\nfn query_sessions(name: &str, deadline: Duration) -> Option<SessionsReply> { … }\n\n// [int->REQ-CONN-BLACKHOLE-LIFECYCLE-HARNESS]\n#[test]\nfn blackholed_controller_lifecycle_five_invariants() { … }","startLine":1},"summary":{"lines":128,"elidedSpans":10,"elidedLines":571},"fileSize":34165,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-daemon\\tests\\conn_blackhole_lifecycle.rs"}}},"isError":false,"timestamp":1788999526087}}
{"type":"message","id":"bef24cc1","parentId":"648734f9","timestamp":"2026-09-10T00:18:46.098Z","message":{"role":"toolResult","toolCallId":"call_l9O7zIiPJselAVqRlROjBVHy|fc_04ba494197989c86016aa1f765fe8087d1bebdabb56e4b7334","toolName":"grep","content":[{"type":"text","text":"# .worktrees/286-churn/crates/spt-daemon/src/\n## brain.rs\n 57|    KIND_MET_MEMBER, KIND_NET_STREAM_SEND, KIND_NET_STREAM_SUBSCRIBE, KIND_OUTPUT, KIND_PAIR_CODE_SUBMIT, KIND_PAIR_JOIN, KIND_PAIR_JOINED, KIND_PAIR_MEET,\n*58|    KIND_RESIZE, KIND_SESSIONS, KIND_SESSIONS_REPLY, KIND_SIZE, KIND_SPAWN, KIND_SPAWNED,\n 59|    KIND_SPAWN_CONFLICT, KIND_SPAWN_FRESH,\n 60|    KIND_SUBSCRIBE, KIND_SUBSCRIBED, KIND_UNSUBSCRIBE, UnsubscribeReq,\n 61|    BringUpReq, BroughtUpReply, KIND_BRING_UP, KIND_BROUGHT_UP,\n...\n 446|    /// except the peer pump, which uses [`Brain::cold_start_pump`].\n*447|    pub fn cold_start(name: &str, now_ms: u64) -> io::Result<Self> {\n 448|        let mut conn = LocalSocketTransport::connect(name)?;\n*449|        send_hello(&mut conn, Role::Brain)?;\n 450|        Ok(Brain {\n 451|            conn: BrainConn::Whole(conn),\n 452|            generation: 0,\n...\n 479|    // [impl->REQ-RC-DISPLAY-SOLE-WRITER]\n*480|    pub fn cold_start_pump(\n 481|        name: &str,\n 482|        now_ms: u64,\n 483|        io_timeout: Duration,\n...\n 486|        let mut stream = LocalSocketTransport::connect(name)?;\n*487|        send_hello(&mut stream, Role::Brain)?;\n 488|        Ok(Brain {\n 489|            conn: BrainConn::split_with_reader(stream, trace),\n 490|            generation: 0,\n...\n 525|        let mut conn = LocalSocketTransport::connect(name)?;\n*526|        send_hello(&mut conn, Role::Brain)?;\n 527|        // DEFENSE-IN-DEPTH (P1c): seed the dedup cursor at the resume point so\n 528|        // `read_event` routes this session through the dedup-below + snap-above\n 529|        // path (matching `resume_sessions`), not the legacy reject-gap path that\n...\n 1268|                }\n*1269|                KIND_SESSIONS_REPLY => {\n 1270|                    let ev: SessionsReply = serde_json::from_value(env.payload)\n 1271|                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n 1272|                    return Ok(BrokerEvent::Sessions(ev));\n...\n 1632|    /// `delivered_through`-blind, broker.rs). Drives the serve_attach controller\n*1633|    /// gap-resume — no mid-stream `KIND_SESSIONS` round-trip (which would re-fatal\n 1634|    /// on the same gap and discard Output).\n 1635|    // [impl->REQ-HAZARD-CONTROLLER-GAP-RESUME]\n 1636|    pub fn controller_resume_floor(&self) -> u64 {\n...\n 1699|    /// the generation does not match the child the broker currently supervises\n*1700|    /// (any other `Role::Brain` client lands here), and `false` too when the\n 1701|    /// broker is too old to know the kind (`unknown command kind`) — either way\n 1702|    /// `daemon status` reports the coordinator image as not reported, never a\n 1703|    /// masquerading version. Round-trip on purpose: it consumes an older\n...\n 2180|    // [impl->REQ-HAZARD-BROKER-PROCESS-ISOLATION]\n*2181|    pub fn request_brain_restart(&mut self) -> io::Result<bool> {\n 2182|        self.send(KIND_BRAIN_RESTART, serde_json::Value::Null)?;\n 2183|        loop {\n 2184|            match self.read_event()? {\n...\n 2237|    /// dispatcher's session→endpoint resolution (D9-1).\n*2238|    pub fn sessions(&mut self) -> io::Result<SessionsReply> {\n*2239|        self.send(KIND_SESSIONS, serde_json::Value::Null)?;\n 2240|        loop {\n 2241|            match self.read_event()? {\n 2242|                BrokerEvent::Sessions(ev) => return Ok(ev),\n...\n 2627|        let mut server = listener.accept().expect(\"accept the brain\");\n*2628|        crate::transport::recv_hello(&mut server, Role::Brain).expect(\"stub recv_hello\");\n 2629|        let brain = brain_handle.join().expect(\"brain thread\");\n 2630|\n 2631|        // Split: the stub writes fed frames through the send half; the recv half is\n## broker.rs\n 71|    KIND_NET_STREAM_UNSUBSCRIBE, KIND_NET_STREAM_UNSUBSCRIBED,\n*72|    KIND_MET_MEMBER, KIND_PAIR_CODE_SUBMIT, KIND_PAIR_JOIN, KIND_PAIR_JOINED, KIND_PAIR_MEET, KIND_RESIZE, KIND_SESSIONS, KIND_SESSIONS_REPLY, KIND_SPAWN,\n 73|    BringUpReq, BroughtUpReply, BRING_UP_ADMITTED, BRING_UP_ADMIT_PREFIX, BRING_UP_ALREADY_LIVE, BRING_UP_REFUSED, KIND_BRING_UP, KIND_BROUGHT_UP,\n 74|    SealCeremonyReq, SealCeremonyReply, SealCeremonyOpenEvent, SealCeremonyCodeReq, SealCeremonyResultEvent,\n 75|    KIND_SEAL_CEREMONY, KIND_SEAL_CEREMONY_REPLY, KIND_SEAL_CEREMONY_OPEN, KIND_SEAL_CEREMONY_CODE, KIND_SEAL_CEREMONY_RESULT,\n...\n 211|/// write — the brain subscriber suspended / black-holed — before a take/reattach\n*212|/// or the `KIND_SESSIONS` reap stall-evicts it (LIFECYCLE-TRUTH W2,\n 213|/// REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE). Distinct from and 3× longer than\n 214|/// [`CONTROLLER_WRITE_DEADLINE`] on PURPOSE (doyle ruling Q1): that 5 s bound\n 215|/// protects the fan-out against ONE per-client conn falling behind and is\n...\n 286|\n*287|impl Drop for WakeClaimGuard<'_> {\n 288|    fn drop(&mut self) {\n 289|        release_claim(&mut recover(self.map), &self.endpoint, self.generation);\n 290|    }\n...\n 429|/// `spawned_ms_ago` comes from the session's `spawned_at`; callers on the\n*430|/// KIND_SESSIONS reply use the additive `SessionInfo` fields, so the CLI\n 431|/// preflight and the broker gate consult the SAME authority.\n 432|/// PURE resolution of the two liveness probes for [`session_is_zombie`], so the\n 433|/// NO-SNAPSHOT arm is unit-testable WITHOUT a platform that lacks a process\n...\n 548|/// whole reason the surface is trustworthy. The connection role cannot do this\n*549|/// job: the broker admits every client as `Role::Brain`, and the loudest\n 550|/// impostor is the CLI's own `Brain::cold_start`, which is by definition running\n 551|/// the freshly-installed version — accepting on role would report a healthy\n 552|/// version as the coordinator's on exactly the node whose coordinator is stale.\n...\n 1002|    /// **survives `detach_if`** — a dead brain leaves its resume point behind, and\n*1003|    /// a cold-starting brain reads it (via `KIND_SESSIONS` → `resume_seq`) instead\n 1004|    /// of replaying the whole ring from 0. **Viewer writes NEVER move it**\n 1005|    /// (REQ-RCVIEW-1: the brain-resume cursor tracks the controller only).\n 1006|    ///\n...\n 1106|/// no-session truth (the exit-waiter reap) BUMPS, and a poll-derived stamp\n*1107|/// write VALIDATES against — so a `KIND_SESSIONS` snapshot taken PRE-reap can\n 1108|/// never relatch `controlled=true` after the reap cleared it (the session is\n 1109|/// gone; no later poll would ever heal the latch). The `write` mutex\n 1110|/// serializes the validate-then-write pair against the bump-then-clear pair —\n...\n 1138|/// `controlled:false` and no later edge re-stamps, so the endpoint reads uncontrolled\n*1139|/// forever while driven. This DIVERGENCE-GATED re-assert (on the `KIND_SESSIONS` poll,\n 1140|/// bounded by the reconcile cadence) writes ONLY on a real difference — no fsync storm.\n 1141|/// Best-effort: an unbound / corrupt / perch-less endpoint has nowhere to converge yet.\n 1142|///\n...\n 2697|    /// THREAD has EXITED — the record-vs-truth honesty step, run on the authoritative\n*2698|    /// `KIND_SESSIONS` poll BEFORE `controller_by`/`has_controller` are read. The\n 2699|    /// writer exits when a socket write fails (a severed / keepalive-killed controller\n 2700|    /// conn) or its `tx` drops; while it is parked on `rx.recv()` (a LIVE but idle\n 2701|    /// controller) it is NOT finished, so a live controller is never touched (the\n...\n 2733|        // writer is BLOCKED (not exited) on a suspended brain conn past\n*2734|        // BRAIN_WRITE_DEADLINE is also reaped here, so the honest `KIND_SESSIONS`\n 2735|        // truth + the off-lock converge_perch_stamps release its stale `driven_by`\n 2736|        // stamp even when nobody take/reattaches. Deferred stamp: the reap runs in\n 2737|        // the multi-session snapshot closure (no I/O under the shared sessions lock).\n...\n 2785|    /// the tally surfaces on `daemon status`). The take/reattach path and the\n*2786|    /// `KIND_SESSIONS` reap call this so a suspended brain subscriber can never hold\n 2787|    /// control on an idle session (the output-driven [`mark_controller_gone`] never\n 2788|    /// fires without new output). Returns true iff it evicted. `stamp` selects whether\n 2789|    /// the perch re-stamp happens inline (the take/reattach path, which already stamps\n...\n 3328|    ///   * controller Disconnected: the writer is already gone — the next append\n*3329|    ///     or the KIND_SESSIONS liveness reap detaches the row, unchanged.\n 3330|    ///\n 3331|    /// `last_ok` is re-armed here regardless of outcome: the suppressed window\n 3332|    /// attempted no handoff, so its width must never count toward the\n...\n 3608|        // the endpoint's stamp generation UNDER the stamp-write serial, then\n*3609|        // clear. A KIND_SESSIONS poll that snapshotted this session PRE-reap now\n 3610|        // fails its generation check and can never relatch `controlled=true`\n 3611|        // onto the perch of a session that no longer exists (the emphasys C3\n 3612|        // poll-vs-reap window: no later poll would ever heal the latch).\n...\n 3661|    /// The next seq a re-attaching controller still needs (the persisted resume\n*3662|    /// cursor surfaced in `KIND_SESSIONS` → `resume_seq`). W1: read from the\n 3663|    /// atomic the controller writer thread advances.\n 3664|    fn delivered_through(&self) -> u64 {\n 3665|        self.delivered_through.load(Ordering::Acquire)\n...\n 3688|    }\n*3689|    impl Drop for ExitNote {\n 3690|        fn drop(&mut self) {\n 3691|            self.send.lifecycle_event(\n 3692|                \"writer-exit\",\n...\n 5704|    /// drains through the REAL logical paths (a take's `stall_evict_controller`\n*5705|    /// or a `KIND_SESSIONS` reap), exactly like a production stall-evict.\n 5706|    /// Returns false if the session doesn't exist or the seed can't be built.\n 5707|    /// `#[doc(hidden)]` — test support, not a product surface.\n 5708|    // [impl->REQ-UPDATE-PROMOTE-DRAINED]\n...\n 5870|    /// closed old conn is simply absent → `false` at once (the normal fast promote).\n*5871|    /// Follows the established sessions→log lock order (KIND_SESSIONS reap); reads\n 5872|    /// only `write_blocked_since` under each log lock, never any I/O.\n 5873|    // [impl->REQ-UPDATE-PROMOTE-DRAINED]\n 5874|    pub fn any_local_controller_wedged(&self) -> bool {\n...\n 5924|        struct HandlerTally<'a>(&'a AtomicUsize);\n*5925|        impl Drop for HandlerTally<'_> {\n 5926|            fn drop(&mut self) {\n 5927|                self.0.fetch_sub(1, Ordering::AcqRel);\n 5928|            }\n...\n 5944|        // Version handshake: this connection must be a brain (2.3).\n*5945|        recv_hello(&mut recv, Role::Brain)?;\n 5946|\n 5947|        let mut my_subs: Vec<u64> = Vec::new();\n 5948|        let mut my_stream_subs: Vec<u64> = Vec::new();\n...\n 6132|                }\n*6133|                KIND_SESSIONS => {\n 6134|                    // Clone the row facts out from under the sessions lock, then\n 6135|                    // read each log's cursor under the log lock alone — never\n 6136|                    // both locks at once (the dispatch_subscribe discipline).\n## transport.rs\n 80|// [impl->REQ-HAZARD-HANDOFF-ARGV-COMPAT]\n*81|pub fn recv_hello<R: Read>(r: &mut R, expected_peer: Role) -> io::Result<Hello> {\n 82|    let env = read_frame(r)?;\n 83|    if env.kind != \"hello\" {\n 84|        return Err(io::Error::new(\n...\n 119|            let mut conn = LocalSocketTransport::connect(&client_name).unwrap();\n*120|            send_hello(&mut conn, Role::Brain).unwrap();\n 121|            write_frame(&mut conn, &Envelope::new(\"ping\", json!({ \"seq\": 1 }))).unwrap();\n 122|            // Read the broker's reply.\n 123|            read_frame(&mut conn).unwrap()\n...\n 126|        let mut server = broker.accept().unwrap();\n*127|        let hello = recv_hello(&mut server, Role::Brain).unwrap();\n*128|        assert_eq!(hello.role, Role::Brain);\n 129|        let ping = read_frame(&mut server).unwrap();\n 130|        assert_eq!(ping.kind, \"ping\");\n 131|        write_frame(&mut server, &Envelope::new(\"pong\", json!({ \"seq\": 1 }))).unwrap();\n...\n 151|        let mut server = broker.accept().unwrap();\n*152|        let err = recv_hello(&mut server, Role::Brain).unwrap_err();\n 153|        assert_eq!(err.kind(), io::ErrorKind::InvalidData);\n 154|        client.join().unwrap();\n 155|    }"}],"details":{"scopePath":".worktrees/286-churn/crates/spt-daemon/src/brain.rs, .worktrees/286-churn/crates/spt-daemon/src/broker.rs, .worktrees/286-churn/crates/spt-daemon/src/transport.rs","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-daemon\\src","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":38,"fileCount":3,"files":[".worktrees/286-churn/crates/spt-daemon/src/brain.rs",".worktrees/286-churn/crates/spt-daemon/src/broker.rs",".worktrees/286-churn/crates/spt-daemon/src/transport.rs"],"fileMatches":[{"path":".worktrees/286-churn/crates/spt-daemon/src/brain.rs","count":13},{"path":".worktrees/286-churn/crates/spt-daemon/src/broker.rs","count":20},{"path":".worktrees/286-churn/crates/spt-daemon/src/transport.rs","count":5}],"truncated":true,"perFileLimitReached":20,"displayContent":"# .worktrees/286-churn/crates/spt-daemon/src/\n## brain.rs\n   57│    KIND_MET_MEMBER, KIND_NET_STREAM_SEND, KIND_NET_STREAM_SUBSCRIBE, KIND_OUTPUT, KIND_PAIR_CODE_SUBMIT, KIND_PAIR_JOIN, KIND_PAIR_JOINED, KIND_PAIR_MEET,\n  *58│    KIND_RESIZE, KIND_SESSIONS, KIND_SESSIONS_REPLY, KIND_SIZE, KIND_SPAWN, KIND_SPAWNED,\n   59│    KIND_SPAWN_CONFLICT, KIND_SPAWN_FRESH,\n   60│    KIND_SUBSCRIBE, KIND_SUBSCRIBED, KIND_UNSUBSCRIBE, UnsubscribeReq,\n   61│    BringUpReq, BroughtUpReply, KIND_BRING_UP, KIND_BROUGHT_UP,\n     │...\n  446│    /// except the peer pump, which uses [`Brain::cold_start_pump`].\n *447│    pub fn cold_start(name: &str, now_ms: u64) -> io::Result<Self> {\n  448│        let mut conn = LocalSocketTransport::connect(name)?;\n *449│        send_hello(&mut conn, Role::Brain)?;\n  450│        Ok(Brain {\n  451│            conn: BrainConn::Whole(conn),\n  452│            generation: 0,\n     │...\n  479│    // [impl->REQ-RC-DISPLAY-SOLE-WRITER]\n *480│    pub fn cold_start_pump(\n  481│        name: &str,\n  482│        now_ms: u64,\n  483│        io_timeout: Duration,\n     │...\n  486│        let mut stream = LocalSocketTransport::connect(name)?;\n *487│        send_hello(&mut stream, Role::Brain)?;\n  488│        Ok(Brain {\n  489│            conn: BrainConn::split_with_reader(stream, trace),\n  490│            generation: 0,\n     │...\n  525│        let mut conn = LocalSocketTransport::connect(name)?;\n *526│        send_hello(&mut conn, Role::Brain)?;\n  527│        // DEFENSE-IN-DEPTH (P1c): seed the dedup cursor at the resume point so\n  528│        // `read_event` routes this session through the dedup-below + snap-above\n  529│        // path (matching `resume_sessions`), not the legacy reject-gap path that\n     │...\n 1268│                }\n*1269│                KIND_SESSIONS_REPLY => {\n 1270│                    let ev: SessionsReply = serde_json::from_value(env.payload)\n 1271│                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n 1272│                    return Ok(BrokerEvent::Sessions(ev));\n     │...\n 1632│    /// `delivered_through`-blind, broker.rs). Drives the serve_attach controller\n*1633│    /// gap-resume — no mid-stream `KIND_SESSIONS` round-trip (which would re-fatal\n 1634│    /// on the same gap and discard Output).\n 1635│    // [impl->REQ-HAZARD-CONTROLLER-GAP-RESUME]\n 1636│    pub fn controller_resume_floor(&self) -> u64 {\n     │...\n 1699│    /// the generation does not match the child the broker currently supervises\n*1700│    /// (any other `Role::Brain` client lands here), and `false` too when the\n 1701│    /// broker is too old to know the kind (`unknown command kind`) — either way\n 1702│    /// `daemon status` reports the coordinator image as not reported, never a\n 1703│    /// masquerading version. Round-trip on purpose: it consumes an older\n     │...\n 2180│    // [impl->REQ-HAZARD-BROKER-PROCESS-ISOLATION]\n*2181│    pub fn request_brain_restart(&mut self) -> io::Result<bool> {\n 2182│        self.send(KIND_BRAIN_RESTART, serde_json::Value::Null)?;\n 2183│        loop {\n 2184│            match self.read_event()? {\n     │...\n 2237│    /// dispatcher's session→endpoint resolution (D9-1).\n*2238│    pub fn sessions(&mut self) -> io::Result<SessionsReply> {\n*2239│        self.send(KIND_SESSIONS, serde_json::Value::Null)?;\n 2240│        loop {\n 2241│            match self.read_event()? {\n 2242│                BrokerEvent::Sessions(ev) => return Ok(ev),\n     │...\n 2627│        let mut server = listener.accept().expect(\"accept the brain\");\n*2628│        crate::transport::recv_hello(&mut server, Role::Brain).expect(\"stub recv_hello\");\n 2629│        let brain = brain_handle.join().expect(\"brain thread\");\n 2630│\n 2631│        // Split: the stub writes fed frames through the send half; the recv half is\n## broker.rs\n   71│    KIND_NET_STREAM_UNSUBSCRIBE, KIND_NET_STREAM_UNSUBSCRIBED,\n  *72│    KIND_MET_MEMBER, KIND_PAIR_CODE_SUBMIT, KIND_PAIR_JOIN, KIND_PAIR_JOINED, KIND_PAIR_MEET, KIND_RESIZE, KIND_SESSIONS, KIND_SESSIONS_REPLY, KIND_SPAWN,\n   73│    BringUpReq, BroughtUpReply, BRING_UP_ADMITTED, BRING_UP_ADMIT_PREFIX, BRING_UP_ALREADY_LIVE, BRING_UP_REFUSED, KIND_BRING_UP, KIND_BROUGHT_UP,\n   74│    SealCeremonyReq, SealCeremonyReply, SealCeremonyOpenEvent, SealCeremonyCodeReq, SealCeremonyResultEvent,\n   75│    KIND_SEAL_CEREMONY, KIND_SEAL_CEREMONY_REPLY, KIND_SEAL_CEREMONY_OPEN, KIND_SEAL_CEREMONY_CODE, KIND_SEAL_CEREMONY_RESULT,\n     │...\n  211│/// write — the brain subscriber suspended / black-holed — before a take/reattach\n *212│/// or the `KIND_SESSIONS` reap stall-evicts it (LIFECYCLE-TRUTH W2,\n  213│/// REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE). Distinct from and 3× longer than\n  214│/// [`CONTROLLER_WRITE_DEADLINE`] on PURPOSE (doyle ruling Q1): that 5 s bound\n  215│/// protects the fan-out against ONE per-client conn falling behind and is\n     │...\n  286│\n *287│impl Drop for WakeClaimGuard<'_> {\n  288│    fn drop(&mut self) {\n  289│        release_claim(&mut recover(self.map), &self.endpoint, self.generation);\n  290│    }\n     │...\n  429│/// `spawned_ms_ago` comes from the session's `spawned_at`; callers on the\n *430│/// KIND_SESSIONS reply use the additive `SessionInfo` fields, so the CLI\n  431│/// preflight and the broker gate consult the SAME authority.\n  432│/// PURE resolution of the two liveness probes for [`session_is_zombie`], so the\n  433│/// NO-SNAPSHOT arm is unit-testable WITHOUT a platform that lacks a process\n     │...\n  548│/// whole reason the surface is trustworthy. The connection role cannot do this\n *549│/// job: the broker admits every client as `Role::Brain`, and the loudest\n  550│/// impostor is the CLI's own `Brain::cold_start`, which is by definition running\n  551│/// the freshly-installed version — accepting on role would report a healthy\n  552│/// version as the coordinator's on exactly the node whose coordinator is stale.\n     │...\n 1002│    /// **survives `detach_if`** — a dead brain leaves its resume point behind, and\n*1003│    /// a cold-starting brain reads it (via `KIND_SESSIONS` → `resume_seq`) instead\n 1004│    /// of replaying the whole ring from 0. **Viewer writes NEVER move it**\n 1005│    /// (REQ-RCVIEW-1: the brain-resume cursor tracks the controller only).\n 1006│    ///\n     │...\n 1106│/// no-session truth (the exit-waiter reap) BUMPS, and a poll-derived stamp\n*1107│/// write VALIDATES against — so a `KIND_SESSIONS` snapshot taken PRE-reap can\n 1108│/// never relatch `controlled=true` after the reap cleared it (the session is\n 1109│/// gone; no later poll would ever heal the latch). The `write` mutex\n 1110│/// serializes the validate-then-write pair against the bump-then-clear pair —\n     │...\n 1138│/// `controlled:false` and no later edge re-stamps, so the endpoint reads uncontrolled\n*1139│/// forever while driven. This DIVERGENCE-GATED re-assert (on the `KIND_SESSIONS` poll,\n 1140│/// bounded by the reconcile cadence) writes ONLY on a real difference — no fsync storm.\n 1141│/// Best-effort: an unbound / corrupt / perch-less endpoint has nowhere to converge yet.\n 1142│///\n     │...\n 2697│    /// THREAD has EXITED — the record-vs-truth honesty step, run on the authoritative\n*2698│    /// `KIND_SESSIONS` poll BEFORE `controller_by`/`has_controller` are read. The\n 2699│    /// writer exits when a socket write fails (a severed / keepalive-killed controller\n 2700│    /// conn) or its `tx` drops; while it is parked on `rx.recv()` (a LIVE but idle\n 2701│    /// controller) it is NOT finished, so a live controller is never touched (the\n     │...\n 2733│        // writer is BLOCKED (not exited) on a suspended brain conn past\n*2734│        // BRAIN_WRITE_DEADLINE is also reaped here, so the honest `KIND_SESSIONS`\n 2735│        // truth + the off-lock converge_perch_stamps release its stale `driven_by`\n 2736│        // stamp even when nobody take/reattaches. Deferred stamp: the reap runs in\n 2737│        // the multi-session snapshot closure (no I/O under the shared sessions lock).\n     │...\n 2785│    /// the tally surfaces on `daemon status`). The take/reattach path and the\n*2786│    /// `KIND_SESSIONS` reap call this so a suspended brain subscriber can never hold\n 2787│    /// control on an idle session (the output-driven [`mark_controller_gone`] never\n 2788│    /// fires without new output). Returns true iff it evicted. `stamp` selects whether\n 2789│    /// the perch re-stamp happens inline (the take/reattach path, which already stamps\n     │...\n 3328│    ///   * controller Disconnected: the writer is already gone — the next append\n*3329│    ///     or the KIND_SESSIONS liveness reap detaches the row, unchanged.\n 3330│    ///\n 3331│    /// `last_ok` is re-armed here regardless of outcome: the suppressed window\n 3332│    /// attempted no handoff, so its width must never count toward the\n     │...\n 3608│        // the endpoint's stamp generation UNDER the stamp-write serial, then\n*3609│        // clear. A KIND_SESSIONS poll that snapshotted this session PRE-reap now\n 3610│        // fails its generation check and can never relatch `controlled=true`\n 3611│        // onto the perch of a session that no longer exists (the emphasys C3\n 3612│        // poll-vs-reap window: no later poll would ever heal the latch).\n     │...\n 3661│    /// The next seq a re-attaching controller still needs (the persisted resume\n*3662│    /// cursor surfaced in `KIND_SESSIONS` → `resume_seq`). W1: read from the\n 3663│    /// atomic the controller writer thread advances.\n 3664│    fn delivered_through(&self) -> u64 {\n 3665│        self.delivered_through.load(Ordering::Acquire)\n     │...\n 3688│    }\n*3689│    impl Drop for ExitNote {\n 3690│        fn drop(&mut self) {\n 3691│            self.send.lifecycle_event(\n 3692│                \"writer-exit\",\n     │...\n 5704│    /// drains through the REAL logical paths (a take's `stall_evict_controller`\n*5705│    /// or a `KIND_SESSIONS` reap), exactly like a production stall-evict.\n 5706│    /// Returns false if the session doesn't exist or the seed can't be built.\n 5707│    /// `#[doc(hidden)]` — test support, not a product surface.\n 5708│    // [impl->REQ-UPDATE-PROMOTE-DRAINED]\n     │...\n 5870│    /// closed old conn is simply absent → `false` at once (the normal fast promote).\n*5871│    /// Follows the established sessions→log lock order (KIND_SESSIONS reap); reads\n 5872│    /// only `write_blocked_since` under each log lock, never any I/O.\n 5873│    // [impl->REQ-UPDATE-PROMOTE-DRAINED]\n 5874│    pub fn any_local_controller_wedged(&self) -> bool {\n     │...\n 5924│        struct HandlerTally<'a>(&'a AtomicUsize);\n*5925│        impl Drop for HandlerTally<'_> {\n 5926│            fn drop(&mut self) {\n 5927│                self.0.fetch_sub(1, Ordering::AcqRel);\n 5928│            }\n     │...\n 5944│        // Version handshake: this connection must be a brain (2.3).\n*5945│        recv_hello(&mut recv, Role::Brain)?;\n 5946│\n 5947│        let mut my_subs: Vec<u64> = Vec::new();\n 5948│        let mut my_stream_subs: Vec<u64> = Vec::new();\n     │...\n 6132│                }\n*6133│                KIND_SESSIONS => {\n 6134│                    // Clone the row facts out from under the sessions lock, then\n 6135│                    // read each log's cursor under the log lock alone — never\n 6136│                    // both locks at once (the dispatch_subscribe discipline).\n## transport.rs\n  80│// [impl->REQ-HAZARD-HANDOFF-ARGV-COMPAT]\n *81│pub fn recv_hello<R: Read>(r: &mut R, expected_peer: Role) -> io::Result<Hello> {\n  82│    let env = read_frame(r)?;\n  83│    if env.kind != \"hello\" {\n  84│        return Err(io::Error::new(\n    │...\n 119│            let mut conn = LocalSocketTransport::connect(&client_name).unwrap();\n*120│            send_hello(&mut conn, Role::Brain).unwrap();\n 121│            write_frame(&mut conn, &Envelope::new(\"ping\", json!({ \"seq\": 1 }))).unwrap();\n 122│            // Read the broker's reply.\n 123│            read_frame(&mut conn).unwrap()\n    │...\n 126│        let mut server = broker.accept().unwrap();\n*127│        let hello = recv_hello(&mut server, Role::Brain).unwrap();\n*128│        assert_eq!(hello.role, Role::Brain);\n 129│        let ping = read_frame(&mut server).unwrap();\n 130│        assert_eq!(ping.kind, \"ping\");\n 131│        write_frame(&mut server, &Envelope::new(\"pong\", json!({ \"seq\": 1 }))).unwrap();\n    │...\n 151│        let mut server = broker.accept().unwrap();\n*152│        let err = recv_hello(&mut server, Role::Brain).unwrap_err();\n 153│        assert_eq!(err.kind(), io::ErrorKind::InvalidData);\n 154│        client.join().unwrap();\n 155│    }"},"isError":false,"timestamp":1788999526098}}
{"type":"message","id":"76daa677","parentId":"bef24cc1","timestamp":"2026-09-10T00:18:46.103Z","message":{"role":"toolResult","toolCallId":"call_f4zLEDdKAHP55bJ7LsJfMXDW|fc_04ba494197989c86016aa1f765fea487d191759e69bdaef2b1","toolName":"grep","content":[{"type":"text","text":"No matches found\nSkipped missing paths: .worktrees/286-churn/scripts"}],"details":{"scopePath":".worktrees/286-churn/crates/spt-daemon/tests, .worktrees/286-churn/crates/spt-daemon/src/broker.rs","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-daemon","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":0,"fileCount":0,"files":[],"truncated":false,"missingPaths":[".worktrees/286-churn/scripts"]},"isError":false,"useless":true,"timestamp":1788999526102}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_TN9DCQEx0BUKRWMlcwwNG6C4|fc_04ba494197989c86016aa1f773101087d19f294de6106e7363","toolName":"read","startedAt":"2026-09-10T00:18:59.139Z","args":{"path":".worktrees/286-churn/crates/spt-daemon/src/brain.rs:436-460,1150-1280,2228-2257"},"intent":"Reading census client contract"},"id":"821c2504","parentId":"76daa677","timestamp":"2026-09-10T00:18:59.139Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_N4mYWqwsUEKTIo71FTmmrHQu|fc_04ba494197989c86016aa1f773102087d1b07a2ff28833cde0","toolName":"read","startedAt":"2026-09-10T00:18:59.139Z","args":{"path":".worktrees/286-churn/crates/spt-daemon/src/broker.rs:5900-6030,6100-6230"},"intent":"Reading broker census routing"},"id":"4e08f695","parentId":"821c2504","timestamp":"2026-09-10T00:18:59.139Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_2TECCqwuAU1wojbDRz3AyGFO|fc_04ba494197989c86016aa1f773102487d1aefb07c9f14a8a15","toolName":"grep","startedAt":"2026-09-10T00:18:59.139Z","args":{"path":".worktrees/286-churn/crates/spt-proto;.worktrees/286-churn/tools;.worktrees/286-churn/crates/spt-daemon/tests"},"intent":"Locating log output implementation"},"id":"ec796bf5","parentId":"4e08f695","timestamp":"2026-09-10T00:18:59.139Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_uFKff8ET1Eegxa2m8FpNqTKD|fc_04ba494197989c86016aa1f773102c87d1935d5239d68cfcd6","toolName":"grep","startedAt":"2026-09-10T00:18:59.139Z","args":{"path":".worktrees/286-churn/docs/KNOWN-HAZARDS.md:.420-1200"},"intent":"Locating safety contract sections"},"id":"420acc7d","parentId":"ec796bf5","timestamp":"2026-09-10T00:18:59.139Z"}
{"type":"message","id":"93fa797b","parentId":"420acc7d","timestamp":"2026-09-10T00:18:59.140Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":"**Reading Brain files**"},{"type":"toolCall","id":"call_TN9DCQEx0BUKRWMlcwwNG6C4|fc_04ba494197989c86016aa1f773101087d19f294de6106e7363","name":"read","arguments":{"i":"Reading census client contract","path":".worktrees/286-churn/crates/spt-daemon/src/brain.rs:436-460,1150-1280,2228-2257"},"intent":"Reading census client contract"},{"type":"toolCall","id":"call_N4mYWqwsUEKTIo71FTmmrHQu|fc_04ba494197989c86016aa1f773102087d1b07a2ff28833cde0","name":"read","arguments":{"i":"Reading broker census routing","path":".worktrees/286-churn/crates/spt-daemon/src/broker.rs:5900-6030,6100-6230"},"intent":"Reading broker census routing"},{"type":"toolCall","id":"call_2TECCqwuAU1wojbDRz3AyGFO|fc_04ba494197989c86016aa1f773102487d1aefb07c9f14a8a15","name":"grep","arguments":{"i":"Locating log output implementation","pattern":"macro_rules! emit_line_err|fn emit_line_err|CONN_LIFECYCLE|write-start|transport-close","path":".worktrees/286-churn/crates/spt-proto;.worktrees/286-churn/tools;.worktrees/286-churn/crates/spt-daemon/tests","case":true,"gitignore":false,"skip":null},"intent":"Locating log output implementation"},{"type":"toolCall","id":"call_uFKff8ET1Eegxa2m8FpNqTKD|fc_04ba494197989c86016aa1f773102c87d1935d5239d68cfcd6","name":"grep","arguments":{"i":"Locating safety contract sections","pattern":"^###.*(SHAREDSEND|LIVEHOST|CONN|BROKER)|^##|^###","path":".worktrees/286-churn/docs/KNOWN-HAZARDS.md:.420-1200","case":true,"gitignore":false,"skip":null},"intent":"Locating safety contract sections"},{"type":"toolCall","id":"call_3SM5dfY95KDRDo1ca7XanZi9|fc_04ba494197989c86016aa1f773103487d1a3e44e79cdf4e932","name":"read","arguments":{"i":"Reading architecture orientation","path":".worktrees/286-churn/docs/adr/0004-broker-brain-split.md"},"intent":"Reading architecture orientation"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":8436,"output":364,"cacheRead":58880,"cacheWrite":0,"totalTokens":67680,"reasoningTokens":28,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1788999526124,"responseId":"resp_04ba494197989c86016aa1f76660b087d1884bce0b5def3fea","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_04ba494197989c86016aa1f7682acc87d1b91a7ec653f107cc","type":"reasoning","content":[],"encrypted_content":"gAAAAABqofdzenAsVrKVOndzUsnh8vbedZIsIplcQODTS5drc_HsTCIL_9q_Wy7NHrkMcBepLv7aw0OPgeNfogLs5GrXc9lHNoT0Oyu2yKTXfG7koqmdsTu_kl-3fw03T6jjmVjdFgC9gGdyVpHjuv9eeQbAL-xwR_2o2eQFeava6Kl55NwXTpBRyZmQ5qkiMyoLiULBQVRnf8BsWcVPaobKOLtRkPs4aWn7-A4ukrXb3hx1aWh_Ohyuqz5R3zXGHYZJwnJiTWsySeM-hLNhe2Ic0e27gjDxzPDaeA6Y4ZdcxsD9oPJMY-Gyfzi28tw0iVt49iaORKVlhDnzyxcZoITUhW4y9PsViuAugQbx1CHHOsK5WDgIv11tmUftDGP8T8wWhlFhNGLshHT6yKDpp6B9KlqyEZUrF8kUv33ZrCQUgioV9VvGv0XrDo34stesDs2LNG4uvzv2Vp17s1gql9RVYu9XJuunIo8sO8axaAqK_PWGDOm7nUudMEEjS84OSsfDnZ3wFUEH0CJBNEtgMpv4eGYd-iPQRo0jIQMlaKG02DLxL4gfa35f45vH741qI0VUtFt16zy59zgkCwen8UuLRK7D5FveFNWfPixcLyO6YHnxMA4TVJtwLgOX5sU549s0EEDHnVlhFlHhICW0J5wcoEE05fGVFJBK7jXSj63yVRYxekF-ifUF4Pzoo-m1Tkj9WUTnPsT0t97Sy4Xo7MaW8zxKNCtRCl564mp9VnqHhQrGX_5HUtu1lejwcnMwNWFZngQ_SqLwMnwnrQhsfDHJlqzc443CFpucKrp2-3y6wqx8jhOrji8suzF7Leulq4v3xLPLmwLA_iZssfG6JtE647JcithSEKyrWgRDqEwHYbkGfvH4kn5S_y6t2jJcvl5lrdC-YyFGohjdDFP_SU-LEw8KnbuKCTp6tKHGoUW12dwmpCsGKUWRuYacUpWwpcyvD2QnmcbJE67NM5j3r82A2Hy1oqWH1bMTfXq0JBNq1diZZVd1bibE0-YLoJjs2oWWC0KRdxUS6hLY0-VOs1jDXUvaZa4hCjLW1pHtIADgLrIFODr58cyR1iuBbh7SfRFBq_snBDzTXEHfD_oSrgdEfLpo0m677VFOIk1dub_j-gGlhgP8j3q1hnyDcPVgkLdpHro_lnMkeNmiCwyYQJw5mHf2xlEeBiDBxZH_nWtJs8kp53TSl6xzIFaRdZU3UsYvD4kZmxobqptAZZd-kANkoS9TWp9853xjQkncGeDtoXwbZdJkXQwbUiFbdXy78NC6Me4oDVs2NKY8HHva4QuzbBt1lpOMbzS7zwyWxowMGE4n1DAqxgsLbf_Hj8ubQzIrY2ifnNCDbkxtenx3FLE287pT7YCXEJF6kx0zUnxGBvFEITwMV27HYMUPao-k3z6lCtc_d2UHmIjuzK5vrgjZxES9KhcziUuiNJ8NGBK1qCXvppkZFgZ8l2WslrjmWrvcbKWVzyCY","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"summary":[{"type":"summary_text","text":"**Reading Brain files**"}],"metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f773101087d19f294de6106e7363","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading census client contract\",\"path\":\".worktrees/286-churn/crates/spt-daemon/src/brain.rs:436-460,1150-1280,2228-2257\"}","call_id":"call_TN9DCQEx0BUKRWMlcwwNG6C4","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"read","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f773102087d1b07a2ff28833cde0","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading broker census routing\",\"path\":\".worktrees/286-churn/crates/spt-daemon/src/broker.rs:5900-6030,6100-6230\"}","call_id":"call_N4mYWqwsUEKTIo71FTmmrHQu","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"read","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f773102487d1aefb07c9f14a8a15","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating log output implementation\",\"pattern\":\"macro_rules! emit_line_err|fn emit_line_err|CONN_LIFECYCLE|write-start|transport-close\",\"path\":\".worktrees/286-churn/crates/spt-proto;.worktrees/286-churn/tools;.worktrees/286-churn/crates/spt-daemon/tests\",\"case\":true,\"gitignore\":false,\"skip\":null}","call_id":"call_2TECCqwuAU1wojbDRz3AyGFO","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"grep","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f773102c87d1935d5239d68cfcd6","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating safety contract sections\",\"pattern\":\"^###.*(SHAREDSEND|LIVEHOST|CONN|BROKER)|^##|^###\",\"path\":\".worktrees/286-churn/docs/KNOWN-HAZARDS.md:.420-1200\",\"case\":true,\"gitignore\":false,\"skip\":null}","call_id":"call_uFKff8ET1Eegxa2m8FpNqTKD","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"grep","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f773103487d1a3e44e79cdf4e932","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading architecture orientation\",\"path\":\".worktrees/286-churn/docs/adr/0004-broker-brain-split.md\"}","call_id":"call_3SM5dfY95KDRDo1ca7XanZi9","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"read","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}}]},"duration":13012.571899999864,"ttft":2098.1713000000454,"completedAt":1788999539137,"contextSnapshot":{"promptTokens":67316,"nonMessageTokens":13648,"compactionEpoch":0}}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_3SM5dfY95KDRDo1ca7XanZi9|fc_04ba494197989c86016aa1f773103487d1a3e44e79cdf4e932","toolName":"read","startedAt":"2026-09-10T00:18:59.140Z","args":{"path":".worktrees/286-churn/docs/adr/0004-broker-brain-split.md"},"intent":"Reading architecture orientation"},"id":"5925ccaa","parentId":"93fa797b","timestamp":"2026-09-10T00:18:59.140Z"}
{"type":"message","id":"0246e215","parentId":"5925ccaa","timestamp":"2026-09-10T00:18:59.145Z","message":{"role":"toolResult","toolCallId":"call_uFKff8ET1Eegxa2m8FpNqTKD|fc_04ba494197989c86016aa1f773102c87d1935d5239d68cfcd6","toolName":"grep","content":[{"type":"text","text":"Path not found: .worktrees/286-churn/docs/KNOWN-HAZARDS.md:.420-1200"}],"details":{},"isError":true,"timestamp":1788999539145}}
{"type":"message","id":"6d22eb3f","parentId":"0246e215","timestamp":"2026-09-10T00:18:59.160Z","message":{"role":"toolResult","toolCallId":"call_TN9DCQEx0BUKRWMlcwwNG6C4|fc_04ba494197989c86016aa1f773101087d19f294de6106e7363","toolName":"read","content":[{"type":"text","text":"pub struct Brain {\n…\n    /// re-fetches — B2 live exactly-once preserved). Only ever consulted on the\n    /// legacy single-session path; a resume-mode brain (snap-above) ignores it. A\n    /// resume-from-floor re-fetch (`from_seq > 0`) leaves this `false` → raw-ring\n    /// replay + strict reject-gap, entirely unchanged.\n    baseline_next_output: bool,\n}\n\nimpl Brain {\n    /// Cold start: connect, handshake, fresh generation 0, `gen_start = now`.\n    /// The ordinary blocking carrier ([`BrainConn::Whole`]) — every client\n    /// except the peer pump, which uses [`Brain::cold_start_pump`].\n    pub fn cold_start(name: &str, now_ms: u64) -> io::Result<Self> {\n        let mut conn = LocalSocketTransport::connect(name)?;\n        send_hello(&mut conn, Role::Brain)?;\n        Ok(Brain {\n            conn: BrainConn::Whole(conn),\n            generation: 0,\n            gen_start_ms: now_ms,\n            session_id: None,\n            next_seq: 0,\n            net_cursors: std::collections::HashMap::new(),\n            presence_cursor: None,\n            session_cursors: std::collections::HashMap::new(),\n            io_timeout: None,\n            baseline_next_output: false,\n        })\n    }\n…\n                KIND_SEAL_CEREMONY_REPLY => {\n…\n                    return Ok(BrokerEvent::SealCeremonyDone {\n                        outcome: ev.outcome,\n                        detail: ev.detail,\n                        token: ev.token,\n                    });\n                }\n                // [impl->REQ-SEAL-CEREMONY-RC-CLIENT]\n                KIND_SEAL_CEREMONY_OPEN => {\n                    let ev: SealCeremonyOpenEvent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::SealCeremonyOpen {\n                        session_id: ev.session_id,\n                        ceremony_id: ev.ceremony_id,\n                        content_b64: ev.content_b64,\n                        subnet: ev.subnet,\n                        destination: ev.destination,\n                        payload_to_sign_b64: ev.payload_to_sign_b64,\n                        fido2_node: ev.fido2_node,\n                        offer_enroll: ev.offer_enroll,\n                    });\n                }\n                KIND_SEAL_CEREMONY_RESULT => {\n                    let ev: SealCeremonyResultEvent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::SealCeremonyResult {\n                        session_id: ev.session_id,\n                        ceremony_id: ev.ceremony_id,\n                        outcome: ev.outcome,\n                        detail: ev.detail,\n                    });\n                }\n                KIND_SPAWNED => {\n                    let ev: Spawned = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Spawned {\n                        session_id: ev.session_id,\n                        pid: ev.pid,\n                    });\n                }\n                KIND_EXIT => {\n                    let ev: ExitEvent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Exit {\n                        session_id: ev.session_id,\n                        code: ev.code,\n                    });\n                }\n                KIND_SUBSCRIBED => {\n                    let ev: SubscribedReply = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Subscribed {\n                        session_id: ev.session_id,\n                        outcome: ev.outcome,\n                        note: ev.note,\n                    });\n                }\n                KIND_SIZE => {\n                    let ev: SizeEvent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Size {\n                        session_id: ev.session_id,\n                        rows: ev.rows,\n                        cols: ev.cols,\n                    });\n                }\n                KIND_DISPLACED => {\n                    let ev: DisplacedEvent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Displaced {\n                        session_id: ev.session_id,\n                        by: ev.by,\n                    });\n                }\n                // [impl->REQ-VIEWER-SKIP-TO-LIVE-ON-EVICT]\n                KIND_VIEWER_EVICTED => {\n                    let ev: ViewerEvictedEvent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::ViewerEvicted {\n                        session_id: ev.session_id,\n                    });\n                }\n                KIND_APPLIED => {\n                    let ev: AppliedEvent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Applied {\n                        session_id: ev.session_id,\n                        op_id: ev.op_id,\n                        applied_now: ev.applied_now,\n                    });\n                }\n                KIND_NET_STATUS_REPLY => {\n                    let ev: NetStatusReply = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::NetStatus(ev));\n                }\n                KIND_NET_DIALED => {\n                    let ev: NetDialed = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::NetDialed(ev));\n                }\n                // The bare ack for a non-blocking submit-dial (PUMP-TRUTH W2) —\n                // null payload; the dial's outcome rides the presence stream.\n                // [impl->REQ-PUMP-PEER-ISOLATION]\n                KIND_NET_DIAL_SUBMITTED => return Ok(BrokerEvent::NetDialSubmitted),\n                KIND_NET_STREAM_OPENED => {\n                    let ev: NetStreamOpened = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::NetStreamOpened(ev));\n                }\n                KIND_NET_SENT => {\n                    let ev: NetSent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::NetSent(ev));\n                }\n                KIND_NET_STREAMS_REPLY => {\n                    let ev: NetStreamsReply = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::NetStreams(ev));\n                }\n                KIND_SESSIONS_REPLY => {\n                    let ev: SessionsReply = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Sessions(ev));\n                }\n                KIND_PAIR_JOINED => {\n                    let ev: PairJoinReply = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::PairJoined(ev));\n                }\n                KIND_MET_MEMBER => {\n                    let ev: MetMember = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n…\n                }\n…\n    pub fn inject_endpoint(\n…\n    ) -> io::Result<(bool, bool)> {\n…\n        loop {\n            match self.read_event()? {\n                BrokerEvent::Other(env) if env.kind == KIND_ENDPOINT_INJECTED => {\n…\n                    return Ok((reply.delivered, reply.spool_deferred));\n                }\n                BrokerEvent::Error { message } => return Err(io::Error::other(message)),\n                _ => continue,\n            }\n        }\n    }\n\n    /// List the broker's hosted sessions (id + owning endpoint) — the inbound\n    /// dispatcher's session→endpoint resolution (D9-1).\n    pub fn sessions(&mut self) -> io::Result<SessionsReply> {\n        self.send(KIND_SESSIONS, serde_json::Value::Null)?;\n        loop {\n            match self.read_event()? {\n                BrokerEvent::Sessions(ev) => return Ok(ev),\n                BrokerEvent::Error { message } => return Err(io::Error::other(message)),\n                _ => continue,\n            }\n        }\n    }\n\n    /// (Re)subscribe to a stream's buffered reads from `from_seq`, setting the\n    /// local cursor (gapless, exactly-once resume — the PTY `subscribe`\n    /// discipline on a QUIC stream, D4b).\n    // [impl->REQ-DAEMON-2]\n    pub fn net_stream_subscribe(&mut self, stream_id: u64, from_seq: u64) -> io::Result<()> {\n        self.net_cursors.insert(stream_id, from_seq);\n        self.send(\n            KIND_NET_STREAM_SUBSCRIBE,\n            serde_json::to_value(NetStreamSubscribeReq {\n…\n            })\n            .expect(\"NetStreamSubscribeReq serializes\"),\n        )\n    }\n…\n}"}],"details":{"displayContent":{"text":"pub struct Brain {\n…\n    /// re-fetches — B2 live exactly-once preserved). Only ever consulted on the\n    /// legacy single-session path; a resume-mode brain (snap-above) ignores it. A\n    /// resume-from-floor re-fetch (`from_seq > 0`) leaves this `false` → raw-ring\n    /// replay + strict reject-gap, entirely unchanged.\n    baseline_next_output: bool,\n}\n\nimpl Brain {\n    /// Cold start: connect, handshake, fresh generation 0, `gen_start = now`.\n    /// The ordinary blocking carrier ([`BrainConn::Whole`]) — every client\n    /// except the peer pump, which uses [`Brain::cold_start_pump`].\n    pub fn cold_start(name: &str, now_ms: u64) -> io::Result<Self> {\n        let mut conn = LocalSocketTransport::connect(name)?;\n        send_hello(&mut conn, Role::Brain)?;\n        Ok(Brain {\n            conn: BrainConn::Whole(conn),\n            generation: 0,\n            gen_start_ms: now_ms,\n            session_id: None,\n            next_seq: 0,\n            net_cursors: std::collections::HashMap::new(),\n            presence_cursor: None,\n            session_cursors: std::collections::HashMap::new(),\n            io_timeout: None,\n            baseline_next_output: false,\n        })\n    }\n…\n                KIND_SEAL_CEREMONY_REPLY => {\n…\n                    return Ok(BrokerEvent::SealCeremonyDone {\n                        outcome: ev.outcome,\n                        detail: ev.detail,\n                        token: ev.token,\n                    });\n                }\n                // [impl->REQ-SEAL-CEREMONY-RC-CLIENT]\n                KIND_SEAL_CEREMONY_OPEN => {\n                    let ev: SealCeremonyOpenEvent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::SealCeremonyOpen {\n                        session_id: ev.session_id,\n                        ceremony_id: ev.ceremony_id,\n                        content_b64: ev.content_b64,\n                        subnet: ev.subnet,\n                        destination: ev.destination,\n                        payload_to_sign_b64: ev.payload_to_sign_b64,\n                        fido2_node: ev.fido2_node,\n                        offer_enroll: ev.offer_enroll,\n                    });\n                }\n                KIND_SEAL_CEREMONY_RESULT => {\n                    let ev: SealCeremonyResultEvent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::SealCeremonyResult {\n                        session_id: ev.session_id,\n                        ceremony_id: ev.ceremony_id,\n                        outcome: ev.outcome,\n                        detail: ev.detail,\n                    });\n                }\n                KIND_SPAWNED => {\n                    let ev: Spawned = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Spawned {\n                        session_id: ev.session_id,\n                        pid: ev.pid,\n                    });\n                }\n                KIND_EXIT => {\n                    let ev: ExitEvent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Exit {\n                        session_id: ev.session_id,\n                        code: ev.code,\n                    });\n                }\n                KIND_SUBSCRIBED => {\n                    let ev: SubscribedReply = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Subscribed {\n                        session_id: ev.session_id,\n                        outcome: ev.outcome,\n                        note: ev.note,\n                    });\n                }\n                KIND_SIZE => {\n                    let ev: SizeEvent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Size {\n                        session_id: ev.session_id,\n                        rows: ev.rows,\n                        cols: ev.cols,\n                    });\n                }\n                KIND_DISPLACED => {\n                    let ev: DisplacedEvent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Displaced {\n                        session_id: ev.session_id,\n                        by: ev.by,\n                    });\n                }\n                // [impl->REQ-VIEWER-SKIP-TO-LIVE-ON-EVICT]\n                KIND_VIEWER_EVICTED => {\n                    let ev: ViewerEvictedEvent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::ViewerEvicted {\n                        session_id: ev.session_id,\n                    });\n                }\n                KIND_APPLIED => {\n                    let ev: AppliedEvent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Applied {\n                        session_id: ev.session_id,\n                        op_id: ev.op_id,\n                        applied_now: ev.applied_now,\n                    });\n                }\n                KIND_NET_STATUS_REPLY => {\n                    let ev: NetStatusReply = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::NetStatus(ev));\n                }\n                KIND_NET_DIALED => {\n                    let ev: NetDialed = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::NetDialed(ev));\n                }\n                // The bare ack for a non-blocking submit-dial (PUMP-TRUTH W2) —\n                // null payload; the dial's outcome rides the presence stream.\n                // [impl->REQ-PUMP-PEER-ISOLATION]\n                KIND_NET_DIAL_SUBMITTED => return Ok(BrokerEvent::NetDialSubmitted),\n                KIND_NET_STREAM_OPENED => {\n                    let ev: NetStreamOpened = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::NetStreamOpened(ev));\n                }\n                KIND_NET_SENT => {\n                    let ev: NetSent = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::NetSent(ev));\n                }\n                KIND_NET_STREAMS_REPLY => {\n                    let ev: NetStreamsReply = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::NetStreams(ev));\n                }\n                KIND_SESSIONS_REPLY => {\n                    let ev: SessionsReply = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::Sessions(ev));\n                }\n                KIND_PAIR_JOINED => {\n                    let ev: PairJoinReply = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n                    return Ok(BrokerEvent::PairJoined(ev));\n                }\n                KIND_MET_MEMBER => {\n                    let ev: MetMember = serde_json::from_value(env.payload)\n                        .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n…\n                }\n…\n    pub fn inject_endpoint(\n…\n    ) -> io::Result<(bool, bool)> {\n…\n        loop {\n            match self.read_event()? {\n                BrokerEvent::Other(env) if env.kind == KIND_ENDPOINT_INJECTED => {\n…\n                    return Ok((reply.delivered, reply.spool_deferred));\n                }\n                BrokerEvent::Error { message } => return Err(io::Error::other(message)),\n                _ => continue,\n            }\n        }\n    }\n\n    /// List the broker's hosted sessions (id + owning endpoint) — the inbound\n    /// dispatcher's session→endpoint resolution (D9-1).\n    pub fn sessions(&mut self) -> io::Result<SessionsReply> {\n        self.send(KIND_SESSIONS, serde_json::Value::Null)?;\n        loop {\n            match self.read_event()? {\n                BrokerEvent::Sessions(ev) => return Ok(ev),\n                BrokerEvent::Error { message } => return Err(io::Error::other(message)),\n                _ => continue,\n            }\n        }\n    }\n\n    /// (Re)subscribe to a stream's buffered reads from `from_seq`, setting the\n    /// local cursor (gapless, exactly-once resume — the PTY `subscribe`\n    /// discipline on a QUIC stream, D4b).\n    // [impl->REQ-DAEMON-2]\n    pub fn net_stream_subscribe(&mut self, stream_id: u64, from_seq: u64) -> io::Result<()> {\n        self.net_cursors.insert(stream_id, from_seq);\n        self.send(\n            KIND_NET_STREAM_SUBSCRIBE,\n            serde_json::to_value(NetStreamSubscribeReq {\n…\n            })\n            .expect(\"NetStreamSubscribeReq serializes\"),\n        )\n    }\n…\n}","startLine":391,"lineNumbers":[391,null,436,437,438,439,440,441,442,443,444,445,446,447,448,449,450,451,452,453,454,455,456,457,458,459,460,461,462,null,1147,null,1150,1151,1152,1153,1154,1155,1156,1157,1158,1159,1160,1161,1162,1163,1164,1165,1166,1167,1168,1169,1170,1171,1172,1173,1174,1175,1176,1177,1178,1179,1180,1181,1182,1183,1184,1185,1186,1187,1188,1189,1190,1191,1192,1193,1194,1195,1196,1197,1198,1199,1200,1201,1202,1203,1204,1205,1206,1207,1208,1209,1210,1211,1212,1213,1214,1215,1216,1217,1218,1219,1220,1221,1222,1223,1224,1225,1226,1227,1228,1229,1230,1231,1232,1233,1234,1235,1236,1237,1238,1239,1240,1241,1242,1243,1244,1245,1246,1247,1248,1249,1250,1251,1252,1253,1254,1255,1256,1257,1258,1259,1260,1261,1262,1263,1264,1265,1266,1267,1268,1269,1270,1271,1272,1273,1274,1275,1276,1277,1278,1279,1280,1281,null,1283,null,2208,null,2213,null,2223,2224,2225,null,2228,2229,2230,2231,2232,2233,2234,2235,2236,2237,2238,2239,2240,2241,2242,2243,2244,2245,2246,2247,2248,2249,2250,2251,2252,2253,2254,2255,2256,2257,null,2260,2261,2262,2263,null,2498]},"fileSize":160723,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-daemon\\src\\brain.rs"}}},"isError":false,"timestamp":1788999539160}}
{"type":"message","id":"94d16ba7","parentId":"6d22eb3f","timestamp":"2026-09-10T00:18:59.161Z","message":{"role":"toolResult","toolCallId":"call_3SM5dfY95KDRDo1ca7XanZi9|fc_04ba494197989c86016aa1f773103487d1a3e44e79cdf4e932","toolName":"read","content":[{"type":"text","text":"Path '.worktrees/286-churn/docs/adr/0004-broker-brain-split.md' not found"}],"details":{},"isError":true,"timestamp":1788999539161}}
{"type":"message","id":"7003304d","parentId":"94d16ba7","timestamp":"2026-09-10T00:18:59.220Z","message":{"role":"toolResult","toolCallId":"call_N4mYWqwsUEKTIo71FTmmrHQu|fc_04ba494197989c86016aa1f773102087d1b07a2ff28833cde0","toolName":"read","content":[{"type":"text","text":"\n    /// The output-log high-water (`next_seq`) of a hosted session, or `None` if\n    /// absent. It advances every time the child emits output — the\n    /// **functional**-survival seam for the D7-1 process-level E2E (doyle F1):\n    /// after a brain-PROCESS swap, writing input to the broker-held PTY child and\n    /// watching THIS advance proves the path is alive end-to-end (input → child →\n    /// echo → broker ring), not merely that the handle is held. A read-only direct\n    /// peek, like [`session_pid`](Self::session_pid).\n    // [impl->REQ-HAZARD-BROKER-PROCESS-ISOLATION]\n    pub fn session_output_seq(&self, id: u64) -> Option<u64> {\n        self.sessions\n            .lock()\n            .unwrap()\n            .get(&id)\n            .map(|h| recover_log(&h.log).next_seq)\n    }\n\n    /// Drive one brain connection: handshake, then dispatch its command frames\n    /// until it disconnects. On disconnect, detach this connection from any\n    /// sessions it subscribed to — the sessions themselves persist (B2: a brain\n    /// restart must not disturb the hosted children).\n    fn handle_conn(self: &Arc<Self>, conn: Stream) -> io::Result<()> {\n        // RAII handler tally (panic-safe): the round-4 gate's seam for \"the\n        // poisoned conn's handler reached EOF cleanup\" (see conn_handler_count).\n        struct HandlerTally<'a>(&'a AtomicUsize);\n        impl Drop for HandlerTally<'_> {\n            fn drop(&mut self) {\n                self.0.fetch_sub(1, Ordering::AcqRel);\n            }\n        }\n        self.conn_handlers.fetch_add(1, Ordering::AcqRel);\n        let _tally = HandlerTally(&self.conn_handlers);\n\n        let (mut recv, send) = conn.split();\n        // Round-4: the send half is owned by the bounded/cancelable conn\n        // primitive from birth — every write on this physical connection is\n        // deadline-bounded, and a wedge poisons + retires the conn.\n        // [impl->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\n        let send: SharedSend = Arc::new(BrokerConn::new(send, brain_write_deadline()));\n        // Base attribution fact (F-039 leg b): broker clients are brains; role\n        // facts (controller/viewer/presence/stream subscriber) accrete at the\n        // attach sites. [impl->REQ-CONN-POISON-ATTRIBUTION]\n        send.describe(\"role=brain\");\n\n        // Version handshake: this connection must be a brain (2.3).\n        recv_hello(&mut recv, Role::Brain)?;\n\n        let mut my_subs: Vec<u64> = Vec::new();\n        let mut my_stream_subs: Vec<u64> = Vec::new();\n        // Streams THIS conn opened as ConnectionBound (ADR-0040 decision 6):\n        // they die with the conn — the exit sweep FINs each toward its target\n        // and terminal-retires the local row. Durable rows never enter it.\n        // [impl->REQ-STREAM-LEASE-CLASSES]\n        let mut my_cb_streams: Vec<u64> = Vec::new();\n        let mut my_presence_sub = false;\n\n        loop {\n            let env = match read_frame(&mut recv) {\n                Ok(e) => e,\n                // Any read failure (EOF included) ends the connection.\n                Err(_) => break,\n            };\n            match env.kind.as_str() {\n                KIND_SPAWN => match self.dispatch_spawn(env, &send) {\n                    Ok(id) => my_subs.push(id),\n                    Err(msg) => send_error(&send, &msg),\n                },\n                // Fresh/RejectExisting spawn (ADR-0041 decision 5): a typed\n                // conflict reply carries no session — only a real Created\n                // auto-subscribes. TEST-RIG-ONLY valve (the N-1 loud-fail leg):\n                // with SPT_TEST_SPAWN_FRESH_UNSUPPORTED set this broker answers\n                // EXACTLY like an N-1 broker that predates the verb — the\n                // forward-compat unknown-kind reply, byte-identical. Never set\n                // in production.\n                // [impl->REQ-SPAWN-FRESH-TRUTHFUL]\n                KIND_SPAWN_FRESH => {\n                    if std::env::var(\"SPT_TEST_SPAWN_FRESH_UNSUPPORTED\").is_ok() {\n                        send_error(&send, &format!(\"unknown command kind {:?}\", env.kind));\n                    } else {\n                        match self.dispatch_spawn_fresh(env, &send) {\n                            Ok(Some(id)) => my_subs.push(id),\n                            Ok(None) => {} // typed conflict already sent\n                            Err(msg) => send_error(&send, &msg),\n                        }\n                    }\n                }\n                // The engine room's entrance (releases#71). Its own verb rather\n                // than a flag on subscribe, because with nothing hosted there is\n                // nothing to subscribe to — which is precisely why the code the\n                // operator typed used to reach no gate at all.\n                // [impl->REQ-ER-BRINGUP-SPAWNS-SESSION]\n                KIND_BRING_UP => {\n                    if let Err(msg) = self.dispatch_bring_up(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                // The wax-seal mint entrance (WAX-SEAL W2, releases#21). The\n                // reply is DEFERRED — it rides this conn from the ceremony's\n                // end, not from this dispatch.\n                // [impl->REQ-SEAL-CEREMONY-TOTP]\n                KIND_SEAL_CEREMONY => {\n                    if let Err(msg) = self.dispatch_seal_ceremony(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                // The enrollment entrance (SIGNET W1) — same ceremony\n                // machinery, enroll purpose, same deferred reply.\n                // [impl->REQ-SEAL-ENROLL-VERB]\n                KIND_SEAL_ENROLL => {\n                    if let Err(msg) = self.dispatch_seal_enroll(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                // The presented code (or cancel) from the controller serve.\n                KIND_SEAL_CEREMONY_CODE => {\n                    if let Err(msg) = self.dispatch_seal_ceremony_code(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                KIND_SUBSCRIBE => match self.dispatch_subscribe(env, &send) {\n                    Ok(id) => my_subs.push(id),\n                    Err(msg) => send_error(&send, &msg),\n                },\n                KIND_UNSUBSCRIBE => {\n                    if let Err(msg) = self.dispatch_unsubscribe(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                KIND_INPUT => {\n                    if let Err(msg) = self.dispatch_input(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n…\n                KIND_NET_STREAM_OPEN => match self.dispatch_net_stream_open(env, &send) {\n…\n                    Ok(Some(id)) if !my_cb_streams.contains(&id) => my_cb_streams.push(id),\n                    Ok(_) => {}\n                    Err(msg) => send_error(&send, &msg),\n                },\n                KIND_NET_STREAM_SEND => {\n                    if let Err(msg) = self.dispatch_net_stream_send(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                KIND_NET_STREAMS => {\n                    if let Err(msg) = self.dispatch_net_streams(&send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                KIND_NET_STREAM_RETIRE => {\n                    if let Err(msg) = self.dispatch_net_stream_retire(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                KIND_NET_STREAM_OPENER => {\n                    // TEST-RIG-ONLY (T7 mixed-image N-1): with\n                    // SPT_TEST_STREAM_OPENER_UNSUPPORTED set this broker\n                    // answers the verb EXACTLY like an N-1 broker that\n                    // predates it — the forward-compat unknown-kind reply\n                    // below, byte-identical (the stable KH-2.3 wire fact the\n                    // dispatcher's fallback gate keys on). Never set in\n                    // production.\n                    if std::env::var(\"SPT_TEST_STREAM_OPENER_UNSUPPORTED\").is_ok() {\n                        send_error(&send, &format!(\"unknown command kind {:?}\", env.kind));\n                    } else if let Err(msg) = self.dispatch_net_stream_opener(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                KIND_SESSIONS => {\n                    // Clone the row facts out from under the sessions lock, then\n                    // read each log's cursor under the log lock alone — never\n                    // both locks at once (the dispatch_subscribe discipline).\n                    // pid/adapter/age ride along (ADR-0041 decision 6): the ONE\n                    // cycle-verb liveness authority's inputs, additive on the\n                    // reply so the CLI preflight and the shutdown edge consult\n                    // the SAME zombie verdict the spawn gate enforces.\n                    // [impl->REQ-ENDPOINT-CYCLE-HONEST]\n                    type SessRow = (u64, String, Arc<Mutex<OutputLog>>, Option<u32>, String, u64);\n                    let rows: Vec<SessRow> = {\n                        let sessions = recover(&self.sessions);\n                        sessions\n                            .iter()\n                            .map(|(id, h)| {\n                                (\n                                    *id,\n                                    h.endpoint.clone(),\n                                    Arc::clone(&h.log),\n                                    h.session.process_id(),\n                                    h.adapter.clone(),\n                                    h.spawned_at.elapsed().as_millis() as u64,\n                                )\n                            })\n                            .collect()\n                    };\n                    // Snapshot each session's truth under its OWN log lock (cursor +\n                    // controller identity + the control/viewer stamp truth), doing NO\n                    // file I/O under the lock (the lock-across-effect discipline, KH\n                    // 7.12/5.16). The convergence re-assert + the reply then run OFF\n                    // the lock.\n                    struct SessSnap {\n                        id: u64,\n                        endpoint: String,\n                        resume_seq: u64,\n                        controller_by: Option<String>,\n                        has_controller: bool,\n                        viewers: u32,\n                        // The endpoint's stamp generation AT snapshot time\n                        // (ADR-0041 decision 4) — the off-lock converge below\n                        // validates against it so a reap that lands between\n                        // this snapshot and the write refuses the stale stamp.\n                        stamp_gen: u64,\n                        // ADR-0041 decision 6: the zombie-verdict inputs, carried\n                        // to the reply (additive SessionInfo fields).\n                        pid: Option<u32>,\n                        adapter: String,\n                        spawned_ms_ago: u64,\n                    }\n                    let snaps: Vec<SessSnap> = rows\n                        .into_iter()\n                        .map(|(id, endpoint, log, pid, adapter, spawned_ms_ago)| {\n                            let mut log = recover_log(&log);\n                            // B-2 (REQ-CONTROLLER-LIVENESS-REAP): make the controller\n                            // truth HONEST before reading it — reap a dead-writer\n                            // (severed) controller so both the reply's controller_by\n                            // AND the off-lock converge_perch_stamps below see the\n                            // cleared state and the stale info.json stamp clears.\n                            let _ = log.reap_dead_controller();\n                            let stamp_gen =\n                                stamp_slot(&endpoint).gen.load(Ordering::Acquire);\n                            SessSnap {\n                                id,\n                                endpoint,\n                                resume_seq: log.delivered_through(),\n                                controller_by: log.controller_by().map(str::to_string),\n                                has_controller: log.has_controller(),\n                                viewers: log.live_viewer_count(),\n                                stamp_gen,\n                                pid,\n                                adapter,\n                                spawned_ms_ago,\n                            }\n                        })\n                        .collect();\n                    // STAMP CONVERGENCE (REQ-HAZARD-CONTROL-STAMP-CONVERGENCE): re-assert\n                    // each live session's control/viewer stamps to broker truth,\n                    // divergence-gated + OFF the log lock. Closes the stamp-before-bind\n                    // loss for a picker-created endpoint (its spawn stamp was swallowed\n                    // pre-bind); converges within one reconcile-poll window after bind.\n                    // [impl->REQ-HAZARD-CONTROL-STAMP-CONVERGENCE]\n                    // [impl->REQ-CONTROL-STAMP-CONVERGENCE]\n                    for s in &snaps {\n                        let _ = converge_perch_stamps(\n                            &s.endpoint,\n                            s.controller_by.as_deref(),\n                            s.has_controller,\n                            s.viewers,\n                            s.stamp_gen,\n                        );\n                    }\n                    let reply = SessionsReply {\n                        sessions: snaps\n                            .into_iter()\n                            .map(|s| SessionInfo {\n                                session_id: s.id,\n                                endpoint: s.endpoint,\n                                resume_seq: s.resume_seq,\n…\n                            })\n                            .collect(),\n…\n                    };\n…\n                }\n…\n            }\n        }\n…\n    }"}],"details":{"displayContent":{"text":"\n    /// The output-log high-water (`next_seq`) of a hosted session, or `None` if\n    /// absent. It advances every time the child emits output — the\n    /// **functional**-survival seam for the D7-1 process-level E2E (doyle F1):\n    /// after a brain-PROCESS swap, writing input to the broker-held PTY child and\n    /// watching THIS advance proves the path is alive end-to-end (input → child →\n    /// echo → broker ring), not merely that the handle is held. A read-only direct\n    /// peek, like [`session_pid`](Self::session_pid).\n    // [impl->REQ-HAZARD-BROKER-PROCESS-ISOLATION]\n    pub fn session_output_seq(&self, id: u64) -> Option<u64> {\n        self.sessions\n            .lock()\n            .unwrap()\n            .get(&id)\n            .map(|h| recover_log(&h.log).next_seq)\n    }\n\n    /// Drive one brain connection: handshake, then dispatch its command frames\n    /// until it disconnects. On disconnect, detach this connection from any\n    /// sessions it subscribed to — the sessions themselves persist (B2: a brain\n    /// restart must not disturb the hosted children).\n    fn handle_conn(self: &Arc<Self>, conn: Stream) -> io::Result<()> {\n        // RAII handler tally (panic-safe): the round-4 gate's seam for \"the\n        // poisoned conn's handler reached EOF cleanup\" (see conn_handler_count).\n        struct HandlerTally<'a>(&'a AtomicUsize);\n        impl Drop for HandlerTally<'_> {\n            fn drop(&mut self) {\n                self.0.fetch_sub(1, Ordering::AcqRel);\n            }\n        }\n        self.conn_handlers.fetch_add(1, Ordering::AcqRel);\n        let _tally = HandlerTally(&self.conn_handlers);\n\n        let (mut recv, send) = conn.split();\n        // Round-4: the send half is owned by the bounded/cancelable conn\n        // primitive from birth — every write on this physical connection is\n        // deadline-bounded, and a wedge poisons + retires the conn.\n        // [impl->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\n        let send: SharedSend = Arc::new(BrokerConn::new(send, brain_write_deadline()));\n        // Base attribution fact (F-039 leg b): broker clients are brains; role\n        // facts (controller/viewer/presence/stream subscriber) accrete at the\n        // attach sites. [impl->REQ-CONN-POISON-ATTRIBUTION]\n        send.describe(\"role=brain\");\n\n        // Version handshake: this connection must be a brain (2.3).\n        recv_hello(&mut recv, Role::Brain)?;\n\n        let mut my_subs: Vec<u64> = Vec::new();\n        let mut my_stream_subs: Vec<u64> = Vec::new();\n        // Streams THIS conn opened as ConnectionBound (ADR-0040 decision 6):\n        // they die with the conn — the exit sweep FINs each toward its target\n        // and terminal-retires the local row. Durable rows never enter it.\n        // [impl->REQ-STREAM-LEASE-CLASSES]\n        let mut my_cb_streams: Vec<u64> = Vec::new();\n        let mut my_presence_sub = false;\n\n        loop {\n            let env = match read_frame(&mut recv) {\n                Ok(e) => e,\n                // Any read failure (EOF included) ends the connection.\n                Err(_) => break,\n            };\n            match env.kind.as_str() {\n                KIND_SPAWN => match self.dispatch_spawn(env, &send) {\n                    Ok(id) => my_subs.push(id),\n                    Err(msg) => send_error(&send, &msg),\n                },\n                // Fresh/RejectExisting spawn (ADR-0041 decision 5): a typed\n                // conflict reply carries no session — only a real Created\n                // auto-subscribes. TEST-RIG-ONLY valve (the N-1 loud-fail leg):\n                // with SPT_TEST_SPAWN_FRESH_UNSUPPORTED set this broker answers\n                // EXACTLY like an N-1 broker that predates the verb — the\n                // forward-compat unknown-kind reply, byte-identical. Never set\n                // in production.\n                // [impl->REQ-SPAWN-FRESH-TRUTHFUL]\n                KIND_SPAWN_FRESH => {\n                    if std::env::var(\"SPT_TEST_SPAWN_FRESH_UNSUPPORTED\").is_ok() {\n                        send_error(&send, &format!(\"unknown command kind {:?}\", env.kind));\n                    } else {\n                        match self.dispatch_spawn_fresh(env, &send) {\n                            Ok(Some(id)) => my_subs.push(id),\n                            Ok(None) => {} // typed conflict already sent\n                            Err(msg) => send_error(&send, &msg),\n                        }\n                    }\n                }\n                // The engine room's entrance (releases#71). Its own verb rather\n                // than a flag on subscribe, because with nothing hosted there is\n                // nothing to subscribe to — which is precisely why the code the\n                // operator typed used to reach no gate at all.\n                // [impl->REQ-ER-BRINGUP-SPAWNS-SESSION]\n                KIND_BRING_UP => {\n                    if let Err(msg) = self.dispatch_bring_up(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                // The wax-seal mint entrance (WAX-SEAL W2, releases#21). The\n                // reply is DEFERRED — it rides this conn from the ceremony's\n                // end, not from this dispatch.\n                // [impl->REQ-SEAL-CEREMONY-TOTP]\n                KIND_SEAL_CEREMONY => {\n                    if let Err(msg) = self.dispatch_seal_ceremony(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                // The enrollment entrance (SIGNET W1) — same ceremony\n                // machinery, enroll purpose, same deferred reply.\n                // [impl->REQ-SEAL-ENROLL-VERB]\n                KIND_SEAL_ENROLL => {\n                    if let Err(msg) = self.dispatch_seal_enroll(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                // The presented code (or cancel) from the controller serve.\n                KIND_SEAL_CEREMONY_CODE => {\n                    if let Err(msg) = self.dispatch_seal_ceremony_code(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                KIND_SUBSCRIBE => match self.dispatch_subscribe(env, &send) {\n                    Ok(id) => my_subs.push(id),\n                    Err(msg) => send_error(&send, &msg),\n                },\n                KIND_UNSUBSCRIBE => {\n                    if let Err(msg) = self.dispatch_unsubscribe(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                KIND_INPUT => {\n                    if let Err(msg) = self.dispatch_input(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n…\n                KIND_NET_STREAM_OPEN => match self.dispatch_net_stream_open(env, &send) {\n…\n                    Ok(Some(id)) if !my_cb_streams.contains(&id) => my_cb_streams.push(id),\n                    Ok(_) => {}\n                    Err(msg) => send_error(&send, &msg),\n                },\n                KIND_NET_STREAM_SEND => {\n                    if let Err(msg) = self.dispatch_net_stream_send(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                KIND_NET_STREAMS => {\n                    if let Err(msg) = self.dispatch_net_streams(&send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                KIND_NET_STREAM_RETIRE => {\n                    if let Err(msg) = self.dispatch_net_stream_retire(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                KIND_NET_STREAM_OPENER => {\n                    // TEST-RIG-ONLY (T7 mixed-image N-1): with\n                    // SPT_TEST_STREAM_OPENER_UNSUPPORTED set this broker\n                    // answers the verb EXACTLY like an N-1 broker that\n                    // predates it — the forward-compat unknown-kind reply\n                    // below, byte-identical (the stable KH-2.3 wire fact the\n                    // dispatcher's fallback gate keys on). Never set in\n                    // production.\n                    if std::env::var(\"SPT_TEST_STREAM_OPENER_UNSUPPORTED\").is_ok() {\n                        send_error(&send, &format!(\"unknown command kind {:?}\", env.kind));\n                    } else if let Err(msg) = self.dispatch_net_stream_opener(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                KIND_SESSIONS => {\n                    // Clone the row facts out from under the sessions lock, then\n                    // read each log's cursor under the log lock alone — never\n                    // both locks at once (the dispatch_subscribe discipline).\n                    // pid/adapter/age ride along (ADR-0041 decision 6): the ONE\n                    // cycle-verb liveness authority's inputs, additive on the\n                    // reply so the CLI preflight and the shutdown edge consult\n                    // the SAME zombie verdict the spawn gate enforces.\n                    // [impl->REQ-ENDPOINT-CYCLE-HONEST]\n                    type SessRow = (u64, String, Arc<Mutex<OutputLog>>, Option<u32>, String, u64);\n                    let rows: Vec<SessRow> = {\n                        let sessions = recover(&self.sessions);\n                        sessions\n                            .iter()\n                            .map(|(id, h)| {\n                                (\n                                    *id,\n                                    h.endpoint.clone(),\n                                    Arc::clone(&h.log),\n                                    h.session.process_id(),\n                                    h.adapter.clone(),\n                                    h.spawned_at.elapsed().as_millis() as u64,\n                                )\n                            })\n                            .collect()\n                    };\n                    // Snapshot each session's truth under its OWN log lock (cursor +\n                    // controller identity + the control/viewer stamp truth), doing NO\n                    // file I/O under the lock (the lock-across-effect discipline, KH\n                    // 7.12/5.16). The convergence re-assert + the reply then run OFF\n                    // the lock.\n                    struct SessSnap {\n                        id: u64,\n                        endpoint: String,\n                        resume_seq: u64,\n                        controller_by: Option<String>,\n                        has_controller: bool,\n                        viewers: u32,\n                        // The endpoint's stamp generation AT snapshot time\n                        // (ADR-0041 decision 4) — the off-lock converge below\n                        // validates against it so a reap that lands between\n                        // this snapshot and the write refuses the stale stamp.\n                        stamp_gen: u64,\n                        // ADR-0041 decision 6: the zombie-verdict inputs, carried\n                        // to the reply (additive SessionInfo fields).\n                        pid: Option<u32>,\n                        adapter: String,\n                        spawned_ms_ago: u64,\n                    }\n                    let snaps: Vec<SessSnap> = rows\n                        .into_iter()\n                        .map(|(id, endpoint, log, pid, adapter, spawned_ms_ago)| {\n                            let mut log = recover_log(&log);\n                            // B-2 (REQ-CONTROLLER-LIVENESS-REAP): make the controller\n                            // truth HONEST before reading it — reap a dead-writer\n                            // (severed) controller so both the reply's controller_by\n                            // AND the off-lock converge_perch_stamps below see the\n                            // cleared state and the stale info.json stamp clears.\n                            let _ = log.reap_dead_controller();\n                            let stamp_gen =\n                                stamp_slot(&endpoint).gen.load(Ordering::Acquire);\n                            SessSnap {\n                                id,\n                                endpoint,\n                                resume_seq: log.delivered_through(),\n                                controller_by: log.controller_by().map(str::to_string),\n                                has_controller: log.has_controller(),\n                                viewers: log.live_viewer_count(),\n                                stamp_gen,\n                                pid,\n                                adapter,\n                                spawned_ms_ago,\n                            }\n                        })\n                        .collect();\n                    // STAMP CONVERGENCE (REQ-HAZARD-CONTROL-STAMP-CONVERGENCE): re-assert\n                    // each live session's control/viewer stamps to broker truth,\n                    // divergence-gated + OFF the log lock. Closes the stamp-before-bind\n                    // loss for a picker-created endpoint (its spawn stamp was swallowed\n                    // pre-bind); converges within one reconcile-poll window after bind.\n                    // [impl->REQ-HAZARD-CONTROL-STAMP-CONVERGENCE]\n                    // [impl->REQ-CONTROL-STAMP-CONVERGENCE]\n                    for s in &snaps {\n                        let _ = converge_perch_stamps(\n                            &s.endpoint,\n                            s.controller_by.as_deref(),\n                            s.has_controller,\n                            s.viewers,\n                            s.stamp_gen,\n                        );\n                    }\n                    let reply = SessionsReply {\n                        sessions: snaps\n                            .into_iter()\n                            .map(|s| SessionInfo {\n                                session_id: s.id,\n                                endpoint: s.endpoint,\n                                resume_seq: s.resume_seq,\n…\n                            })\n                            .collect(),\n…\n                    };\n…\n                }\n…\n            }\n        }\n…\n    }","startLine":5900,"lineNumbers":[5900,5901,5902,5903,5904,5905,5906,5907,5908,5909,5910,5911,5912,5913,5914,5915,5916,5917,5918,5919,5920,5921,5922,5923,5924,5925,5926,5927,5928,5929,5930,5931,5932,5933,5934,5935,5936,5937,5938,5939,5940,5941,5942,5943,5944,5945,5946,5947,5948,5949,5950,5951,5952,5953,5954,5955,5956,5957,5958,5959,5960,5961,5962,5963,5964,5965,5966,5967,5968,5969,5970,5971,5972,5973,5974,5975,5976,5977,5978,5979,5980,5981,5982,5983,5984,5985,5986,5987,5988,5989,5990,5991,5992,5993,5994,5995,5996,5997,5998,5999,6000,6001,6002,6003,6004,6005,6006,6007,6008,6009,6010,6011,6012,6013,6014,6015,6016,6017,6018,6019,6020,6021,6022,6023,6024,6025,6026,6027,6028,6029,6030,6031,6032,null,6095,null,6100,6101,6102,6103,6104,6105,6106,6107,6108,6109,6110,6111,6112,6113,6114,6115,6116,6117,6118,6119,6120,6121,6122,6123,6124,6125,6126,6127,6128,6129,6130,6131,6132,6133,6134,6135,6136,6137,6138,6139,6140,6141,6142,6143,6144,6145,6146,6147,6148,6149,6150,6151,6152,6153,6154,6155,6156,6157,6158,6159,6160,6161,6162,6163,6164,6165,6166,6167,6168,6169,6170,6171,6172,6173,6174,6175,6176,6177,6178,6179,6180,6181,6182,6183,6184,6185,6186,6187,6188,6189,6190,6191,6192,6193,6194,6195,6196,6197,6198,6199,6200,6201,6202,6203,6204,6205,6206,6207,6208,6209,6210,6211,6212,6213,6214,6215,6216,6217,6218,6219,6220,6221,6222,6223,6224,6225,6226,6227,6228,6229,6230,null,6235,6236,null,6245,null,6251,null,6280,6281,null,6335]},"fileSize":844842,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-daemon\\src\\broker.rs"}}},"isError":false,"timestamp":1788999539220}}
{"type":"message","id":"5b59b0de","parentId":"7003304d","timestamp":"2026-09-10T00:18:59.232Z","message":{"role":"toolResult","toolCallId":"call_2TECCqwuAU1wojbDRz3AyGFO|fc_04ba494197989c86016aa1f773102487d1aefb07c9f14a8a15","toolName":"grep","content":[{"type":"text","text":"# .worktrees/286-churn/crates/spt-proto/src/\n## emit.rs\n 186|#[macro_export]\n*187|macro_rules! emit_line_err {\n 188|    ($($arg:tt)*) => {{\n 189|        let _ = $crate::emit::write_line(\n 190|            &mut ::std::io::stderr().lock(),\n\nSkipped missing paths: .worktrees/286-churn/tools"}],"details":{"scopePath":".worktrees/286-churn/crates/spt-proto, .worktrees/286-churn/crates/spt-daemon/tests","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":1,"fileCount":1,"files":[".worktrees/286-churn/crates/spt-proto/src/emit.rs"],"fileMatches":[{"path":".worktrees/286-churn/crates/spt-proto/src/emit.rs","count":1}],"truncated":false,"displayContent":"# .worktrees/286-churn/crates/spt-proto/src/\n## emit.rs\n 186│#[macro_export]\n*187│macro_rules! emit_line_err {\n 188│    ($($arg:tt)*) => {{\n 189│        let _ = $crate::emit::write_line(\n 190│            &mut ::std::io::stderr().lock(),","missingPaths":[".worktrees/286-churn/tools"]},"isError":false,"timestamp":1788999539232}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_Ep2k17QHxISUknQNs4G8klUT|fc_04ba494197989c86016aa1f78cabf887d1b4c18052cc0f2330","toolName":"grep","startedAt":"2026-09-10T00:19:24.754Z","args":{"path":".worktrees/286-churn/docs/KNOWN-HAZARDS.md"},"intent":"Locating load-bearing hazards"},"id":"83c1f4d0","parentId":"5b59b0de","timestamp":"2026-09-10T00:19:24.754Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_3KMS7iOQs9n4KXqYMINFM2Aq|fc_04ba494197989c86016aa1f78cac0887d1a7309b9e20a751df","toolName":"glob","startedAt":"2026-09-10T00:19:24.755Z","args":{"path":".worktrees/286-churn/docs/adr/0004*;.worktrees/286-churn/docs/adr/0018*"},"intent":"Locating architecture decisions"},"id":"69661e48","parentId":"83c1f4d0","timestamp":"2026-09-10T00:19:24.755Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_xmN2hN8ydJUQTV7bHZeoOBU8|fc_04ba494197989c86016aa1f78cac1087d18d8cc71da3a04b65","toolName":"read","startedAt":"2026-09-10T00:19:24.755Z","args":{"path":".worktrees/286-churn/crates/spt-proto/src/emit.rs:1-77,175-200"},"intent":"Reading stderr write contract"},"id":"bc003753","parentId":"69661e48","timestamp":"2026-09-10T00:19:24.755Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_JdGvmSXZnv2rOXLxueGjOYJQ|fc_04ba494197989c86016aa1f78cac1887d1bddfacbb2b0b6af2","toolName":"read","startedAt":"2026-09-10T00:19:24.755Z","args":{"path":".worktrees/286-churn/crates/spt-daemon/src/broker.rs:6213-6259"},"intent":"Reading census reply forwarding"},"id":"f86c949c","parentId":"bc003753","timestamp":"2026-09-10T00:19:24.755Z"}
{"type":"message","id":"6b9a765d","parentId":"f86c949c","timestamp":"2026-09-10T00:19:24.755Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":"**Checking output consistency**\n\n**Checking connection shutdown**\n\n**Locating relevant headings**"},{"type":"toolCall","id":"call_Ep2k17QHxISUknQNs4G8klUT|fc_04ba494197989c86016aa1f78cabf887d1b4c18052cc0f2330","name":"grep","arguments":{"i":"Locating load-bearing hazards","pattern":"^###.*(SHAREDSEND|LIVEHOST|CONN|BROKER|IPC|controller|session)","path":".worktrees/286-churn/docs/KNOWN-HAZARDS.md","case":false,"gitignore":false,"skip":null},"intent":"Locating load-bearing hazards"},{"type":"toolCall","id":"call_3KMS7iOQs9n4KXqYMINFM2Aq|fc_04ba494197989c86016aa1f78cac0887d1a7309b9e20a751df","name":"glob","arguments":{"i":"Locating architecture decisions","path":".worktrees/286-churn/docs/adr/0004*;.worktrees/286-churn/docs/adr/0018*","hidden":true,"gitignore":false,"limit":10},"intent":"Locating architecture decisions"},{"type":"toolCall","id":"call_xmN2hN8ydJUQTV7bHZeoOBU8|fc_04ba494197989c86016aa1f78cac1087d18d8cc71da3a04b65","name":"read","arguments":{"i":"Reading stderr write contract","path":".worktrees/286-churn/crates/spt-proto/src/emit.rs:1-77,175-200"},"intent":"Reading stderr write contract"},{"type":"toolCall","id":"call_JdGvmSXZnv2rOXLxueGjOYJQ|fc_04ba494197989c86016aa1f78cac1887d1bddfacbb2b0b6af2","name":"read","arguments":{"i":"Reading census reply forwarding","path":".worktrees/286-churn/crates/spt-daemon/src/broker.rs:6213-6259"},"intent":"Reading census reply forwarding"},{"type":"toolCall","id":"call_u4IiOAuqAiFe4hGZF5BnXWtz|fc_04ba494197989c86016aa1f78cac2087d184e51af69e863486","name":"grep","arguments":{"i":"Tracing connection lifetime consumers","pattern":"fn send\\(|fn read_event|impl Drop for Brain|conn_handler_count\\(|my_presence_sub|CONN_LIFECYCLE_SUMMARY|CONN_WRITE_POISONED|first_write","path":".worktrees/286-churn/crates/spt-daemon/src/brain.rs;.worktrees/286-churn/crates/spt-daemon/src/broker.rs;.worktrees/286-churn/crates/spt-daemon/tests/conn_blackhole_lifecycle.rs","case":true,"gitignore":false,"skip":null},"intent":"Tracing connection lifetime consumers"},{"type":"toolCall","id":"call_u8hO9OrBm9bV0D0jwvTMTy6k|fc_04ba494197989c86016aa1f78cac2487d183b6f317cd4f7762","name":"read","arguments":{"i":"Reading traceability orientation","path":".worktrees/286-churn/docs/TRACEABILITY.md:1-85"},"intent":"Reading traceability orientation"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":5552,"output":805,"cacheRead":67200,"cacheWrite":0,"totalTokens":73557,"reasoningTokens":395,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1788999539280,"responseId":"resp_04ba494197989c86016aa1f7737e3087d18ec85b2afdd5d0eb","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_04ba494197989c86016aa1f7748bf487d18fe6b12e17c823a0","type":"reasoning","content":[],"encrypted_content":"gAAAAABqofeMKo0ghgLnyWU8NWL9YTRWP-24RNZn_G_wxAMCS7-4aq6fkGvsLCwxzKNQw6y7ft2throJwwUtzmCF9w9Z0Q5eGBY2A39ZJTaa1V23wus7XG27E7kUXSz1krj-VsdJyzrafuc1twm_NAnuw8sn02wQkY5PVCo5Cjnx3nRy69YUcvL_6uBjHqS1I1Nwkl1qyojWEkQZ3wPxpnSNCxR1v7DSvYndk6W2XGTvvPKr46WHXdxeH7yS7tzByeR2ulgKFlcF9lUanOgjLcuWa5oNERtg9GRMp_gcD35EP9BHt_I1KX0u46xbGq6wV_6rdI2qJmOkpPcKkORjHIsJdxvcxyjT_EPXXc8xQp0CsOsw7gdNWBBc7fDjDWiaeMgUovnabzCtrtzptnzFLkvhdHtmk8kdm1S-TrL3mNbKqq3UsI8soOI14D2_Jew3ff574x5gLCXq95ZJQ_UL1_affoLFIFdQXkJUeYDZIpX4wVqD1m2OzexYo4BqfYM4dEIHrO9mjhD35gNq-CYz9njhIuaJIq5CCyQXAy5lxrcIgFAXKdJbbhqw04II0PJmiQfxaieYFYx7snkzuECJjhxjCaCzIryaQ4Jz1QlLKmXTH6n_bpZUipMCAE8TZcH9kmu4qa4Lskp-BT_T-41TpgPzHRrjpR31BY_Hr4meEFsJz0LJWrFTTx8ySIXzKlLnGhyC6JZpgcWbk_26IcKIz60M0Ozt-QMrZiCEwJneFkVywu79YtWS-JZUKkONZjAVwQINrjDctmH7Hu9d3vUQCAt0YUIEo1L3BcFVacOaSA7sobLHPjn05meojDUXVCXVp00jh33e8tc16SIw6T8GXIWkWLsqNJ7fZr3SSn8uP9YjhtBGCGxaYoqo867322Oz1D6tN2G--vugx-Gmy1-nwSnBvB7-zjW2Arwk-S0LbCffAqNz9zGnxgLaCMkP42oOOlHwpahEizEru01Y4An-vY7xBiF-0yvHeJsKp554-N_lz8-YJD2DH-LI_1QqxRuD1Odh54YTA6onGOmKiYb_1Eo8ZdBjSmL9iWj7807b-DsbC12MmQefi2NGje9dFwSe93FI5_BSiTUMf7MAggv4JSVV1K1_s6jJc5L3F8A9t-HWtRX9GYqNQGmtirhTzybt0pjnUs4Aa7_xPDcmu4EZpl1d2KUqW-AsvGSA8VNVQXYN4q2fSK0sUCxma5VNRRq4jj6Ckf3mBkEK0rMy91y1dKh3MwFXQwb8869gAWvTOoxjtoCjQdgbLPnaSeEfxv9ZwSWUumwEMbJaZC-SyX2IlaWXeBYQ7MccmhUt7BfaSZ9mbWaKLxKr1p8UAqFab01op0twz09mlBUmFnrWvh02dvsOfM1erJJAeOlLNdEaGTazGMYpIsuwhoGTdRvhdFWblh4C-8zwgkIaXu2R4wWM2NZFtFWAHxVgIapcNQ6oPjyCAp44-NSve2kO3FQefsodQ31x6kpsuvuQkQZGadxcSmeIO7dZAtjaSFUpyjpwNXhIsy-hSQCFL6DW6sC3fMvQkBPQ9rQsl6Ofg-DvNbexszp8RKxw3tE_MxKe-EzRSmMi3x0Gg4_qfCr_eL1iZUAfBQDdQ9hZiTQiPEdmcGlqls4Q4iJBgyjw46u4JJwJPTf8MOxTswJ9fnGX29oz69MRfkqdBz6qKq5yp78F0NxgFcjRr3dV_F9KOrv5RsG40mFCAFjJI8fG7gbj_zvwwtLr07OMyZGSEftB3IbrY_guQ7pR1VI5weNxWlpxLG1KdEY7pp5joods17nNVWe9vUSMl36ReGhf1qFYLsO4qU1nT9NCa2D2nzBWxLCIw_9mldxqvFgk9eWKonvrg2FrRJWSOjiQ8WlQwbTXIko7SFtkOIQK7S4pOqnaCpPZFxBw7X2uBbjAkuhJ5ehdagY1I7gYkfbDcR-ux6KjtL0R7sipU7VmEKDDZNS0EM7bZv1Sgwk6KzUkyzIoxMNpPCF1xRPBCkX_DmhKoW0Wp-i352bmR0RrMXyc2QDYLazbaClER4ogF46sX7zrd3rie-at4tqY56NFQnPkbYmRUXso0NRtK-RtaQwG7oDpBYYkj0j57rZWlPt94FksUz5dURyrkPy3CUlFqpJ5ppeCHopFzG5z0959jPwgIkS5vCX6ueciJUIksfpjMTPki289Eb_gYXi_tsEXQ2n2goakSs8WPB72kf4GTrLjzAyN1rl2EsnaUz5w5NxR2XR2tvT89jOy5iJCnpS2rJDbbFYytcvLohXttpmB0wy3gGaDkG6Bl4ii23Q5sat8YeyGePG-v02n5xavKzioUx8D37Sl-XGisb2rOKL7hy_cRrrSwMu9KbMRZLtmtVXZpqxM5j4hddvhgLZ1zums1RC_rYw-w4hSPE-_KFT7hHkHk0_YhSRbp1XR5XRRZCnf9bSqCOIQTZS6pbIw_FWZfRpFtCyE-fUYG25-U5UvuSq0DG2WAb-hZ3P485VLh0hbnGuVRYBfuYeN0zDHfwvDJvTs3vyh1tKC3jEtRLtBZMY2UQz35VE3loVkf4wZiqVLR6OYmWNovvE-4BaY9h3dAdnuVfxGKhIRbFuPBRgGZ7wIHKcqPXQfwtgJlgb6wYUpGCyOwMgXp7y5zRpdDexQVV_CyYyz6N84VjiCBH6exxmnrTwCakk1KJTlb4P5PlzuqxExyjuk4CZHT1mqNQ5eW_a9DrAf6Z-YpsBAIXKcgVVSJu5ik1MToMj6zcnCfOKVA6Y1lTa66EqBbWiJ-JTuCbKvJFVmU2xldPGGdhBsAEzdV_hksBN3GY3IZ4zqrrG-TFw4ZBzdvXF_D-5glpl283FCdkTvXFzsnUALSn-scp2LcNhhPt3JnOrh83poXhTRYQmGNFmmzFlArqCMTlnq8CjqGEnac9GQa6IXi7bPzHGD_tQcafuZWj2yDmxamDpuT4-7p7L4dt_qrpeB06IEnuJP3o4GkaTlxBEcTax27zkLKHcp0q9cKeqPH3h8BI9OaxKLLHMAP7t3Vc6eN2DJaVk3eFFZBkIm1F9xiPLu56iu3qxo6nctq5AGYfauk5pNfFQEwnbCtQWaXGpSz_P8mrUnaGlFeFHVXpw6K-Pgs-YCRfRAyxTZ0zPRf36UW1fJwQsMqbmI49JSPtBJoGvEWiFoQUnoSBFqziXtP1Ibz82cQyV6qhIE3yU7E0GyZJOjErHUsJazQKssyYrQdAbSLtKFnBIJr5WcAoC23c9NA-_1f2s-uKhxRm9rHdI-zkWoZAxB-haMxfPQ4MOz7pv7dy2Vr9NL6qIcGdNPU3q_sARbhdSM_trQwsXrohg48SVtubJmgOPRIeoMVIXiTd9MtWhSP2_feV_eDqDj0_SYxVtvdCV55s1DkS1tH8gmSsXuoy4kAEqKHD5quA8ommmjMoH5VNJo4mGgmrTOSMMjRyNBy-Nl5aodQjrvAUW_sIXJMT0146SwcdMpUkrAY5drSF6RP2_Urd5v7S_AiSopEGxnGLQiMlK2MWM7FDlLmO3GdyRmUdAxzhz1D_NH4JVq0MS90BmJxqFywVFFrYgC4RfO9nIOfFUBVhVz8RfBvHsprhc51Lf8MDU-aUYnTqKAKk4Ptv0TLjVrZBOa60n10JXwyNII71PXIvW4_tE3DhaIKuoBQvAeG_1E5AAvv1lalkp7jizGEa3xo20GhIFW1FH-XW61DDLkwDF_w_T1tbou9BvEpgUMbCzjQJGBQRYnQl4Bub9DMcEJfqqXYvMCJKkn0GSQfF_FakZbNozZ0Fx-Iwn2QvZWHt9n8QzViLAVipww4rTlykSY5wuJCJPN9dcpjEMz81yRRPWIdmCDlGBzJ03c90lAqhq4FwrOTLIgDp644xWxpoluCeGRviXCx1k0ELdWm3IQ6SNgVsVvYfSo9runYz_9MaYdFLDr2vkyhJ7euVER3GbhyBVAx7jIBronIeOZF0ivvZKieBovMQA44UD0zLwrph4WzNsML_GGg2-Pd68AKo5FI0Sr3f13j0qtHMYTPoIvDt8EGaxpCxka6c_JiEikSIcCtJwFxRe-61T-ZUZvQY5bKpvp9Y7sGfkvZ1SU9UaSa6SRqquDdOjk0U6MK2I7CXcFh1uJB3aIkqByI8W_575oXRj7wdVEeC5ojTukgOCS4HpnergVcXmJaVhu4tGaicXLUZ6EwhzBxzSHCUt4Zrn0f6-AcnJQhjy3n4jhRUwHDvO6tY1f0WdA1VNQK-hiaSaTk-r3OrFBRFBK9HB7Vx9uRw0DFwjK7vxuYBxh2uWqPDCHw66u5e-E8mcp85Q929s6bb1tyC5mybQJ4POt","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"summary":[{"type":"summary_text","text":"**Checking output consistency**"},{"type":"summary_text","text":"**Checking connection shutdown**"},{"type":"summary_text","text":"**Locating relevant headings**"}],"metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f78cabf887d1b4c18052cc0f2330","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating load-bearing hazards\",\"pattern\":\"^###.*(SHAREDSEND|LIVEHOST|CONN|BROKER|IPC|controller|session)\",\"path\":\".worktrees/286-churn/docs/KNOWN-HAZARDS.md\",\"case\":false,\"gitignore\":false,\"skip\":null}","call_id":"call_Ep2k17QHxISUknQNs4G8klUT","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"grep","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f78cac0887d1a7309b9e20a751df","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating architecture decisions\",\"path\":\".worktrees/286-churn/docs/adr/0004*;.worktrees/286-churn/docs/adr/0018*\",\"hidden\":true,\"gitignore\":false,\"limit\":10}","call_id":"call_3KMS7iOQs9n4KXqYMINFM2Aq","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"glob","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f78cac1087d18d8cc71da3a04b65","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading stderr write contract\",\"path\":\".worktrees/286-churn/crates/spt-proto/src/emit.rs:1-77,175-200\"}","call_id":"call_xmN2hN8ydJUQTV7bHZeoOBU8","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"read","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f78cac1887d1bddfacbb2b0b6af2","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading census reply forwarding\",\"path\":\".worktrees/286-churn/crates/spt-daemon/src/broker.rs:6213-6259\"}","call_id":"call_JdGvmSXZnv2rOXLxueGjOYJQ","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"read","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f78cac2087d184e51af69e863486","type":"function_call","status":"completed","arguments":"{\"i\":\"Tracing connection lifetime consumers\",\"pattern\":\"fn send\\\\(|fn read_event|impl Drop for Brain|conn_handler_count\\\\(|my_presence_sub|CONN_LIFECYCLE_SUMMARY|CONN_WRITE_POISONED|first_write\",\"path\":\".worktrees/286-churn/crates/spt-daemon/src/brain.rs;.worktrees/286-churn/crates/spt-daemon/src/broker.rs;.worktrees/286-churn/crates/spt-daemon/tests/conn_blackhole_lifecycle.rs\",\"case\":true,\"gitignore\":false,\"skip\":null}","call_id":"call_u4IiOAuqAiFe4hGZF5BnXWtz","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"grep","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f78cac2487d183b6f317cd4f7762","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading traceability orientation\",\"path\":\".worktrees/286-churn/docs/TRACEABILITY.md:1-85\"}","call_id":"call_u8hO9OrBm9bV0D0jwvTMTy6k","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"read","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}}]},"duration":25471.47849999997,"ttft":1307.6842000000179,"completedAt":1788999564752,"contextSnapshot":{"promptTokens":72752,"nonMessageTokens":13648,"compactionEpoch":0}}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_u4IiOAuqAiFe4hGZF5BnXWtz|fc_04ba494197989c86016aa1f78cac2087d184e51af69e863486","toolName":"grep","startedAt":"2026-09-10T00:19:24.756Z","args":{"path":".worktrees/286-churn/crates/spt-daemon/src/brain.rs;.worktrees/286-churn/crates/spt-daemon/src/broker.rs;.worktrees/286-churn/crates/spt-daemon/tests/conn_blackhole_lifecycle.rs"},"intent":"Tracing connection lifetime consumers"},"id":"af7af710","parentId":"6b9a765d","timestamp":"2026-09-10T00:19:24.756Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_u8hO9OrBm9bV0D0jwvTMTy6k|fc_04ba494197989c86016aa1f78cac2487d183b6f317cd4f7762","toolName":"read","startedAt":"2026-09-10T00:19:24.756Z","args":{"path":".worktrees/286-churn/docs/TRACEABILITY.md:1-85"},"intent":"Reading traceability orientation"},"id":"1e55bfd6","parentId":"af7af710","timestamp":"2026-09-10T00:19:24.756Z"}
{"type":"message","id":"2e6e8748","parentId":"1e55bfd6","timestamp":"2026-09-10T00:19:24.793Z","message":{"role":"toolResult","toolCallId":"call_xmN2hN8ydJUQTV7bHZeoOBU8|fc_04ba494197989c86016aa1f78cac1087d18d8cc71da3a04b65","toolName":"read","content":[{"type":"text","text":"//! Single-write emission of machine-parsed diagnostic lines.\n//!\n//! A diagnostic a test parses (`TOKEN:...`) must reach the handle as ONE write\n//! of the whole rendered text, newline included. `eprintln!` does not do that:\n//! `io::Stderr` is unbuffered, and `Write::write_fmt` single-writes only a\n//! statically known str — anything interpolated goes through an adapter that\n//! `write_all`s PER FORMAT FRAGMENT, the trailing newline separately. Threads of\n//! one process are serialized by a reentrant lock, but the daemon's children\n//! inherit the same handle and no lock spans processes, so a reader can observe\n//! a line torn between fragments and completed by another process's newline.\n//!\n//! MEASURED on this toolchain, not inferred (W1 step 0, the lane's gate\n//! evidence): an interpolated `write_fmt` issued **5** writes —\n//! `\"SUBSCRIBE_DECISION:\"`, `\"endpoint-7\"`, `\": \"`, `\"the store did not\n//! persist\"`, `\"\\n\"` — while the static-literal control issued **1**.\n//!\n//! Two named arms, because the population has two shapes and a weak assertion\n//! would let them blur:\n//!\n//! - [`emit_line!`] — the single-line contract. It refuses an interior newline:\n//!   `debug_assert` in debug, and in release it ESCAPES rather than emitting the\n//!   text as-is, so one emission is always one logical line. Emitting anyway\n//!   would put multi-line text back on the single-line surface in the one build\n//!   where no assert watches and no test looks.\n//! - [`emit_block!`] — the explicit opt-in for a composed multi-line block,\n//!   still exactly one write. Two sites in the tree need it and now say so by\n//!   name: `firewall.rs` INBOUND_REACHABILITY (its newline rides in the payload)\n//!   and `servicehost.rs` SERVICE_STARTUP_FAULT (it appends a captured tail).\n//!\n//! The handle is a parameter everywhere here. A hardcoded `stderr()` would make\n//! the write count unobservable, and an unobservable property is not a property.\n\nuse std::io::{self, Write};\n\n/// The escape an interior newline is rewritten to by [`compose_line`], pinned\n/// here and asserted in a unit so it cannot drift: a reader who finds these two\n/// characters inside a parsed line has the offending caller named by the\n/// artifact alone, with no repro.\n// [impl->REQ-EMIT-SINGLE-WRITE]\npub const ESCAPED_NEWLINE: &str = r\"\\n\";\n\n/// The same for a bare carriage return, which splits a line for plenty of\n/// readers and would otherwise survive the newline escape as a stray byte.\n// [impl->REQ-EMIT-SINGLE-WRITE]\npub const ESCAPED_RETURN: &str = r\"\\r\";\n\n/// Compose a SINGLE-LINE emission: exactly one logical line, newline-terminated.\n///\n/// Pure on purpose — the text is unit-pinnable without reaching through any\n/// writer, which is what makes the escape spelling testable at all. One trailing\n/// line break is the caller's own terminator and is absorbed rather than\n/// escaped; everything left inside is rewritten, because a payload carrying its\n/// own newline would otherwise pass as two logical lines inside one write —\n/// indistinguishable downstream from the tear this module exists to close.\n// [impl->REQ-EMIT-SINGLE-WRITE]\npub fn compose_line(rendered: &str) -> String {\n    let body = rendered\n        .strip_suffix('\\n')\n        .map(|b| b.strip_suffix('\\r').unwrap_or(b))\n        .unwrap_or(rendered);\n    let mut out = String::with_capacity(body.len() + 1);\n    for ch in body.chars() {\n        match ch {\n            '\\n' => out.push_str(ESCAPED_NEWLINE),\n            '\\r' => out.push_str(ESCAPED_RETURN),\n            other => out.push(other),\n        }\n    }\n    out.push('\\n');\n    out\n}\n\n/// Compose a MULTI-LINE block emission: interior line breaks are the point, and\n/// the block is still handed over as one write. Terminated exactly once.\n// [impl->REQ-EMIT-SINGLE-WRITE]\npub fn compose_block(rendered: &str) -> String {\n    if rendered.ends_with('\\n') {\n…\n    } else {\n…\n    }\n}\n…\n/// writing to stderr fails, so a closed or full stderr can take down a process\n/// at the exact moment it was trying to explain itself. That is the same wedge\n/// class this project has refused before: a reporting surface must never become\n/// the thing that kills the operation it reports on.\n///\n/// The cost is named, not hidden: this makes emission a place where absence can\n/// be SILENT — a diagnostic that never arrives leaves no trace of having been\n/// attempted. That is an observability gap, accepted here only because the\n/// alternative is a panic, and it belongs in the conversion notes so no later\n/// reader rediscovers it as a defect.\n// [impl->REQ-EMIT-SINGLE-WRITE]\n#[macro_export]\nmacro_rules! emit_line_err {\n    ($($arg:tt)*) => {{\n        let _ = $crate::emit::write_line(\n            &mut ::std::io::stderr().lock(),\n            &::std::format!($($arg)*),\n        );\n    }};\n}\n\n/// [`emit_block!`] aimed at this process's stderr, on the same terms as\n/// [`emit_line_err!`]: separate name, error discarded.\n// [impl->REQ-EMIT-SINGLE-WRITE]\n#[macro_export]\nmacro_rules! emit_block_err {\n…\n}"}],"details":{"displayContent":{"text":"//! Single-write emission of machine-parsed diagnostic lines.\n//!\n//! A diagnostic a test parses (`TOKEN:...`) must reach the handle as ONE write\n//! of the whole rendered text, newline included. `eprintln!` does not do that:\n//! `io::Stderr` is unbuffered, and `Write::write_fmt` single-writes only a\n//! statically known str — anything interpolated goes through an adapter that\n//! `write_all`s PER FORMAT FRAGMENT, the trailing newline separately. Threads of\n//! one process are serialized by a reentrant lock, but the daemon's children\n//! inherit the same handle and no lock spans processes, so a reader can observe\n//! a line torn between fragments and completed by another process's newline.\n//!\n//! MEASURED on this toolchain, not inferred (W1 step 0, the lane's gate\n//! evidence): an interpolated `write_fmt` issued **5** writes —\n//! `\"SUBSCRIBE_DECISION:\"`, `\"endpoint-7\"`, `\": \"`, `\"the store did not\n//! persist\"`, `\"\\n\"` — while the static-literal control issued **1**.\n//!\n//! Two named arms, because the population has two shapes and a weak assertion\n//! would let them blur:\n//!\n//! - [`emit_line!`] — the single-line contract. It refuses an interior newline:\n//!   `debug_assert` in debug, and in release it ESCAPES rather than emitting the\n//!   text as-is, so one emission is always one logical line. Emitting anyway\n//!   would put multi-line text back on the single-line surface in the one build\n//!   where no assert watches and no test looks.\n//! - [`emit_block!`] — the explicit opt-in for a composed multi-line block,\n//!   still exactly one write. Two sites in the tree need it and now say so by\n//!   name: `firewall.rs` INBOUND_REACHABILITY (its newline rides in the payload)\n//!   and `servicehost.rs` SERVICE_STARTUP_FAULT (it appends a captured tail).\n//!\n//! The handle is a parameter everywhere here. A hardcoded `stderr()` would make\n//! the write count unobservable, and an unobservable property is not a property.\n\nuse std::io::{self, Write};\n\n/// The escape an interior newline is rewritten to by [`compose_line`], pinned\n/// here and asserted in a unit so it cannot drift: a reader who finds these two\n/// characters inside a parsed line has the offending caller named by the\n/// artifact alone, with no repro.\n// [impl->REQ-EMIT-SINGLE-WRITE]\npub const ESCAPED_NEWLINE: &str = r\"\\n\";\n\n/// The same for a bare carriage return, which splits a line for plenty of\n/// readers and would otherwise survive the newline escape as a stray byte.\n// [impl->REQ-EMIT-SINGLE-WRITE]\npub const ESCAPED_RETURN: &str = r\"\\r\";\n\n/// Compose a SINGLE-LINE emission: exactly one logical line, newline-terminated.\n///\n/// Pure on purpose — the text is unit-pinnable without reaching through any\n/// writer, which is what makes the escape spelling testable at all. One trailing\n/// line break is the caller's own terminator and is absorbed rather than\n/// escaped; everything left inside is rewritten, because a payload carrying its\n/// own newline would otherwise pass as two logical lines inside one write —\n/// indistinguishable downstream from the tear this module exists to close.\n// [impl->REQ-EMIT-SINGLE-WRITE]\npub fn compose_line(rendered: &str) -> String {\n    let body = rendered\n        .strip_suffix('\\n')\n        .map(|b| b.strip_suffix('\\r').unwrap_or(b))\n        .unwrap_or(rendered);\n    let mut out = String::with_capacity(body.len() + 1);\n    for ch in body.chars() {\n        match ch {\n            '\\n' => out.push_str(ESCAPED_NEWLINE),\n            '\\r' => out.push_str(ESCAPED_RETURN),\n            other => out.push(other),\n        }\n    }\n    out.push('\\n');\n    out\n}\n\n/// Compose a MULTI-LINE block emission: interior line breaks are the point, and\n/// the block is still handed over as one write. Terminated exactly once.\n// [impl->REQ-EMIT-SINGLE-WRITE]\npub fn compose_block(rendered: &str) -> String {\n    if rendered.ends_with('\\n') {\n…\n    } else {\n…\n    }\n}\n…\n/// writing to stderr fails, so a closed or full stderr can take down a process\n/// at the exact moment it was trying to explain itself. That is the same wedge\n/// class this project has refused before: a reporting surface must never become\n/// the thing that kills the operation it reports on.\n///\n/// The cost is named, not hidden: this makes emission a place where absence can\n/// be SILENT — a diagnostic that never arrives leaves no trace of having been\n/// attempted. That is an observability gap, accepted here only because the\n/// alternative is a panic, and it belongs in the conversion notes so no later\n/// reader rediscovers it as a defect.\n// [impl->REQ-EMIT-SINGLE-WRITE]\n#[macro_export]\nmacro_rules! emit_line_err {\n    ($($arg:tt)*) => {{\n        let _ = $crate::emit::write_line(\n            &mut ::std::io::stderr().lock(),\n            &::std::format!($($arg)*),\n        );\n    }};\n}\n\n/// [`emit_block!`] aimed at this process's stderr, on the same terms as\n/// [`emit_line_err!`]: separate name, error discarded.\n// [impl->REQ-EMIT-SINGLE-WRITE]\n#[macro_export]\nmacro_rules! emit_block_err {\n…\n}","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,null,79,null,84,85,null,175,176,177,178,179,180,181,182,183,184,185,186,187,188,189,190,191,192,193,194,195,196,197,198,199,200,null,207]},"fileSize":22720,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-proto\\src\\emit.rs"}}},"isError":false,"timestamp":1788999564793}}
{"type":"message","id":"1b8b42dd","parentId":"2e6e8748","timestamp":"2026-09-10T00:19:24.794Z","message":{"role":"toolResult","toolCallId":"call_u8hO9OrBm9bV0D0jwvTMTy6k|fc_04ba494197989c86016aa1f78cac2487d183b6f317cd4f7762","toolName":"read","content":[{"type":"text","text":"# Traceability development contract\n\n> How spt-core uses [`traceable-reqs`](https://github.com/BigscreenVR/traceable-reqs) to keep every requirement traced from doc → impl → test. The manifest is `traceable-reqs.toml` (seeded from PRD `R-*` + KNOWN-HAZARDS invariants). This contract makes the trace pay off instead of rotting.\n\n## Why\n\nThe whole point of the rebuild is to *not* re-introduce the sister project's bugs and to ship the PRD's scope verifiably. A requirement that isn't traced to real evidence (a doc, an implementation, a test) is a requirement you're *hoping* is met. This contract converts that hope into a build that fails when evidence is missing.\n\n## The contract (rules)\n\n1. **The manifest is the authoritative requirement registry.** Every PRD `R-*` requirement and every KNOWN-HAZARDS invariant exists in `traceable-reqs.toml` as a `REQ-*` id. No work without a REQ; no REQ without intent to satisfy. The PRD holds the prose; the manifest holds the id + `required_stages`.\n2. **Tag in the same change as the evidence.** When you write the function / test / doc-section that satisfies a stage, add its `[<stage>-><REQ-ID>]` tag in *that same commit*. Never \"tag later\" — this single rule is what stops the trace from drifting.\n3. **Evidence-proximate tags, not file-tops.** A tag sits on or immediately above the real function, test, doc section, config entry, or workflow step that proves the stage. One tag = one piece of real evidence. Tags at file tops to satisfy coverage are noise and usually wrong.\n4. **KNOWN-HAZARDS are first-class requirements.** Each invariant is a `REQ-HAZARD-*` requiring `unit` (and `int` where cross-process/cross-node). A hazard cannot be \"covered\" without a test tagging it — the anti-regression promise becomes mechanical (PRD success criterion #6).\n5. **Activation, not premature failure.** Every requirement starts `required_stages = []` (inactive) so `check` stays green pre-code. A **milestone (M0–M5) activates** the requirements it delivers by setting their real `required_stages`. Deferred items (`docs/DEFERRED.md`) stay `[]` until promoted.\n6. **Scan roots stay honest.** `[scan].roots` includes every evidence location (`src`, `tests`, `docs`, and later `.github/workflows`, `scripts`). Audit roots whenever a new evidence dir appears — a missing root makes evidence *silently* vanish from the trace.\n7. **Scope at mint.** Every new `REQ-*` joins at least one `[[groups]]` subject scope in the same commit that mints it. The groups partition is what the confusion audit (Gate 4, 2026-07-30) turned into manifest structure: the `req-confusion-audit` skill scores confusability only *within* groups, so an unscoped requirement's confusable pairs go unexamined — roughly half of them, by the audit's measurement. `span-*` groups are the declared cross-scope cover for measured-hot pairs; add one when a new requirement is known to sit confusably against another scope. Do not derive scopes from id prefixes — subject matter, not spelling.\n\n### Stages\n\n`doc` (prose/design/rustdoc) · `impl` (production code) · `unit` (unit test) · `int` (integration / e2e / cross-node). Default activation policy: `[\"doc\",\"impl\",\"unit\"]`; networking/lifecycle/cross-node reqs add `int`.\n\n### Tag examples\n\nReal tags are `[<stage>->REQ-ID]` with a bare stage word. The examples below\nangle-bracket the stage (`[<impl>->...]`) only so the scanner treats them as\nillustrations, not live evidence — write a bare stage word in real code.\n\n```rust\n// [<impl>->REQ-PAIR-1]\nfn spake2_pair(...) { ... }\n```\n```rust\n// [<unit>->REQ-HAZARD-ENVELOPE-DECODE-ORDER]\n#[test] fn decodes_amp_last() { ... }\n```\n```markdown\n<!-- [<doc>->REQ-NET-2] -->\n## Relay dependency\n```\n\n## Enforcement (four layers — defense in depth)\n\n1. **CI gate (hard backstop).** `traceable-reqs check --json` runs on every PR/push; exit-1 (`missing_stage` / `undeclared_id` / `parse_error` / `manifest_error`) **fails the build**. A PR can't merge if it leaves an *activated* req's required stage uncovered. Runs alongside the DOCS-STRATEGY CI gates.\n2. **Milestone activation gate.** Starting a milestone flips its reqs from `[]` to the real policy; the milestone **is not done until `check` is green for its reqs**. Coverage grows *with* the project instead of failing on day one.\n3. **Agent-facing rule.** The repo `CLAUDE.md` instructs every dev-agent: tag `[<stage>-><REQ-ID>]` in the same change as the evidence, and run `traceable-reqs check` before declaring done. Since spt-core is built *with* agents, the contract must be machine-followable — agent discipline + CI as the net.\n4. **Quality audit (anti-box-ticking).** `check` proves a tag *exists*; `traceable-reqs review` + `lint` audit whether tags sit near *real* evidence and whether titles are meaningful. Run advisory in CI (or periodically). Presence = hard gate; quality = audited.\n\nOptional fast-feedback: a pre-push git hook running `traceable-reqs check` locally.\n\n## CI snippet (wire at M0, once there is a Rust project + CI)\n\n```yaml\n# .github/workflows/traceability.yml\nname: traceability\non: [push, pull_request]\njobs:\n  check:\n    runs-on: ubuntu-latest\n    steps:\n      - uses: actions/checkout@v4\n      - name: Install traceable-reqs\n        run: gh release download --repo BigscreenVR/traceable-reqs --pattern '*linux-x86_64' --output traceable-reqs && chmod +x traceable-reqs\n        env: { GH_TOKEN: ${{ github.token }} }\n      - name: Coverage gate (hard)\n        run: ./traceable-reqs check --json\n      - name: Quality audit (advisory)\n        run: ./traceable-reqs lint || true\n```\n\n## Lifecycle\n\n- **Now (planning):** manifest seeded, all reqs inactive. `check` is green (nothing required yet).\n- **M0:** install the CLI, validate the seed against it, wire the CI workflow, **activate M0 reqs** (`REQ-ARCH-*`, `REQ-HAZARD-*` for the commodity layer), tag as code lands.\n- **M1–M5:** each milestone activates + covers its requirements before it's called done."}],"details":{"totalLines":75,"displayContent":{"text":"# Traceability development contract\r\n\r\n> How spt-core uses [`traceable-reqs`](https://github.com/BigscreenVR/traceable-reqs) to keep every requirement traced from doc → impl → test. The manifest is `traceable-reqs.toml` (seeded from PRD `R-*` + KNOWN-HAZARDS invariants). This contract makes the trace pay off instead of rotting.\r\n\r\n## Why\r\n\r\nThe whole point of the rebuild is to *not* re-introduce the sister project's bugs and to ship the PRD's scope verifiably. A requirement that isn't traced to real evidence (a doc, an implementation, a test) is a requirement you're *hoping* is met. This contract converts that hope into a build that fails when evidence is missing.\r\n\r\n## The contract (rules)\r\n\r\n1. **The manifest is the authoritative requirement registry.** Every PRD `R-*` requirement and every KNOWN-HAZARDS invariant exists in `traceable-reqs.toml` as a `REQ-*` id. No work without a REQ; no REQ without intent to satisfy. The PRD holds the prose; the manifest holds the id + `required_stages`.\r\n2. **Tag in the same change as the evidence.** When you write the function / test / doc-section that satisfies a stage, add its `[<stage>-><REQ-ID>]` tag in *that same commit*. Never \"tag later\" — this single rule is what stops the trace from drifting.\r\n3. **Evidence-proximate tags, not file-tops.** A tag sits on or immediately above the real function, test, doc section, config entry, or workflow step that proves the stage. One tag = one piece of real evidence. Tags at file tops to satisfy coverage are noise and usually wrong.\r\n4. **KNOWN-HAZARDS are first-class requirements.** Each invariant is a `REQ-HAZARD-*` requiring `unit` (and `int` where cross-process/cross-node). A hazard cannot be \"covered\" without a test tagging it — the anti-regression promise becomes mechanical (PRD success criterion #6).\r\n5. **Activation, not premature failure.** Every requirement starts `required_stages = []` (inactive) so `check` stays green pre-code. A **milestone (M0–M5) activates** the requirements it delivers by setting their real `required_stages`. Deferred items (`docs/DEFERRED.md`) stay `[]` until promoted.\r\n6. **Scan roots stay honest.** `[scan].roots` includes every evidence location (`src`, `tests`, `docs`, and later `.github/workflows`, `scripts`). Audit roots whenever a new evidence dir appears — a missing root makes evidence *silently* vanish from the trace.\r\n7. **Scope at mint.** Every new `REQ-*` joins at least one `[[groups]]` subject scope in the same commit that mints it. The groups partition is what the confusion audit (Gate 4, 2026-07-30) turned into manifest structure: the `req-confusion-audit` skill scores confusability only *within* groups, so an unscoped requirement's confusable pairs go unexamined — roughly half of them, by the audit's measurement. `span-*` groups are the declared cross-scope cover for measured-hot pairs; add one when a new requirement is known to sit confusably against another scope. Do not derive scopes from id prefixes — subject matter, not spelling.\r\n\r\n### Stages\r\n\r\n`doc` (prose/design/rustdoc) · `impl` (production code) · `unit` (unit test) · `int` (integration / e2e / cross-node). Default activation policy: `[\"doc\",\"impl\",\"unit\"]`; networking/lifecycle/cross-node reqs add `int`.\r\n\r\n### Tag examples\r\n\r\nReal tags are `[<stage>->REQ-ID]` with a bare stage word. The examples below\r\nangle-bracket the stage (`[<impl>->...]`) only so the scanner treats them as\r\nillustrations, not live evidence — write a bare stage word in real code.\r\n\r\n```rust\r\n// [<impl>->REQ-PAIR-1]\r\nfn spake2_pair(...) { ... }\r\n```\r\n```rust\r\n// [<unit>->REQ-HAZARD-ENVELOPE-DECODE-ORDER]\r\n#[test] fn decodes_amp_last() { ... }\r\n```\r\n```markdown\r\n<!-- [<doc>->REQ-NET-2] -->\r\n## Relay dependency\r\n```\r\n\r\n## Enforcement (four layers — defense in depth)\r\n\r\n1. **CI gate (hard backstop).** `traceable-reqs check --json` runs on every PR/push; exit-1 (`missing_stage` / `undeclared_id` / `parse_error` / `manifest_error`) **fails the build**. A PR can't merge if it leaves an *activated* req's required stage uncovered. Runs alongside the DOCS-STRATEGY CI gates.\r\n2. **Milestone activation gate.** Starting a milestone flips its reqs from `[]` to the real policy; the milestone **is not done until `check` is green for its reqs**. Coverage grows *with* the project instead of failing on day one.\r\n3. **Agent-facing rule.** The repo `CLAUDE.md` instructs every dev-agent: tag `[<stage>-><REQ-ID>]` in the same change as the evidence, and run `traceable-reqs check` before declaring done. Since spt-core is built *with* agents, the contract must be machine-followable — agent discipline + CI as the net.\r\n4. **Quality audit (anti-box-ticking).** `check` proves a tag *exists*; `traceable-reqs review` + `lint` audit whether tags sit near *real* evidence and whether titles are meaningful. Run advisory in CI (or periodically). Presence = hard gate; quality = audited.\r\n\r\nOptional fast-feedback: a pre-push git hook running `traceable-reqs check` locally.\r\n\r\n## CI snippet (wire at M0, once there is a Rust project + CI)\r\n\r\n```yaml\r\n# .github/workflows/traceability.yml\r\nname: traceability\r\non: [push, pull_request]\r\njobs:\r\n  check:\r\n    runs-on: ubuntu-latest\r\n    steps:\r\n      - uses: actions/checkout@v4\r\n      - name: Install traceable-reqs\r\n        run: gh release download --repo BigscreenVR/traceable-reqs --pattern '*linux-x86_64' --output traceable-reqs && chmod +x traceable-reqs\r\n        env: { GH_TOKEN: ${{ github.token }} }\r\n      - name: Coverage gate (hard)\r\n        run: ./traceable-reqs check --json\r\n      - name: Quality audit (advisory)\r\n        run: ./traceable-reqs lint || true\r\n```\r\n\r\n## Lifecycle\r\n\r\n- **Now (planning):** manifest seeded, all reqs inactive. `check` is green (nothing required yet).\r\n- **M0:** install the CLI, validate the seed against it, wire the CI workflow, **activate M0 reqs** (`REQ-ARCH-*`, `REQ-HAZARD-*` for the commodity layer), tag as code lands.\r\n- **M1–M5:** each milestone activates + covers its requirements before it's called done.\r","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75]},"fileSize":6057,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\docs\\TRACEABILITY.md"}}},"isError":false,"timestamp":1788999564794}}
{"type":"message","id":"818340c0","parentId":"1b8b42dd","timestamp":"2026-09-10T00:19:24.799Z","message":{"role":"toolResult","toolCallId":"call_JdGvmSXZnv2rOXLxueGjOYJQ|fc_04ba494197989c86016aa1f78cac1887d1bddfacbb2b0b6af2","toolName":"read","content":[{"type":"text","text":"                KIND_SESSIONS => {\n…\n                    // pre-bind); converges within one reconcile-poll window after bind.\n                    // [impl->REQ-HAZARD-CONTROL-STAMP-CONVERGENCE]\n                    // [impl->REQ-CONTROL-STAMP-CONVERGENCE]\n                    for s in &snaps {\n                        let _ = converge_perch_stamps(\n                            &s.endpoint,\n                            s.controller_by.as_deref(),\n                            s.has_controller,\n                            s.viewers,\n                            s.stamp_gen,\n                        );\n                    }\n                    let reply = SessionsReply {\n                        sessions: snaps\n                            .into_iter()\n                            .map(|s| SessionInfo {\n                                session_id: s.id,\n                                endpoint: s.endpoint,\n                                resume_seq: s.resume_seq,\n                                controller_by: s.controller_by,\n                                pid: s.pid,\n                                adapter: s.adapter,\n                                spawned_ms_ago: s.spawned_ms_ago,\n                            })\n                            .collect(),\n                        // The other half of the answer (releases#199): the\n                        // bring-ups this broker has started and not yet\n                        // registered. Read at reply time rather than snapshotted\n                        // with the sessions above, so a bring-up that registered\n                        // mid-handler is reported by the row it now has instead\n                        // of by an entry it no longer needs.\n                        // [impl->REQ-ER-BRINGUP-INFLIGHT-LEDGER]\n                        bringing_up: self.bringups_in_flight_at(Instant::now()),\n                    };\n                    let frame = Envelope::new(\n                        KIND_SESSIONS_REPLY,\n                        serde_json::to_value(reply).expect(\"SessionsReply serializes\"),\n                    );\n                    send_frame(&send, &frame);\n                }\n                KIND_ENDPOINT_INPUT => {\n                    if let Err(msg) = self.dispatch_endpoint_input(env, &send) {\n                        send_error(&send, &msg);\n                    }\n                }\n                KIND_NET_STREAM_SUBSCRIBE => match self.dispatch_net_stream_subscribe(env, &send) {\n                    Ok(id) => my_stream_subs.push(id),\n                    Err(msg) => send_error(&send, &msg),\n                },\n                // Explicit seat release (ADR-0040 decision 5): the released\n                // stream also leaves this conn's sweep list, so the exit\n\n[Showing lines 6212-6262 of 16233. Use :6263 to continue]"}],"details":{"truncation":{"content":"                    // pre-bind); converges within one reconcile-poll window after bind.\r\n                    // [impl->REQ-HAZARD-CONTROL-STAMP-CONVERGENCE]\r\n                    // [impl->REQ-CONTROL-STAMP-CONVERGENCE]\r\n                    for s in &snaps {\r\n                        let _ = converge_perch_stamps(\r\n                            &s.endpoint,\r\n                            s.controller_by.as_deref(),\r\n                            s.has_controller,\r\n                            s.viewers,\r\n                            s.stamp_gen,\r\n                        );\r\n                    }\r\n                    let reply = SessionsReply {\r\n                        sessions: snaps\r\n                            .into_iter()\r\n                            .map(|s| SessionInfo {\r\n                                session_id: s.id,\r\n                                endpoint: s.endpoint,\r\n                                resume_seq: s.resume_seq,\r\n                                controller_by: s.controller_by,\r\n                                pid: s.pid,\r\n                                adapter: s.adapter,\r\n                                spawned_ms_ago: s.spawned_ms_ago,\r\n                            })\r\n                            .collect(),\r\n                        // The other half of the answer (releases#199): the\r\n                        // bring-ups this broker has started and not yet\r\n                        // registered. Read at reply time rather than snapshotted\r\n                        // with the sessions above, so a bring-up that registered\r\n                        // mid-handler is reported by the row it now has instead\r\n                        // of by an entry it no longer needs.\r\n                        // [impl->REQ-ER-BRINGUP-INFLIGHT-LEDGER]\r\n                        bringing_up: self.bringups_in_flight_at(Instant::now()),\r\n                    };\r\n                    let frame = Envelope::new(\r\n                        KIND_SESSIONS_REPLY,\r\n                        serde_json::to_value(reply).expect(\"SessionsReply serializes\"),\r\n                    );\r\n                    send_frame(&send, &frame);\r\n                }\r\n                KIND_ENDPOINT_INPUT => {\r\n                    if let Err(msg) = self.dispatch_endpoint_input(env, &send) {\r\n                        send_error(&send, &msg);\r\n                    }\r\n                }\r\n                KIND_NET_STREAM_SUBSCRIBE => match self.dispatch_net_stream_subscribe(env, &send) {\r\n                    Ok(id) => my_stream_subs.push(id),\r\n                    Err(msg) => send_error(&send, &msg),\r\n                },\r\n                // Explicit seat release (ADR-0040 decision 5): the released\r\n                // stream also leaves this conn's sweep list, so the exit\r","truncated":true,"truncatedBy":"lines","totalLines":10022,"totalBytes":2759,"outputLines":51,"outputBytes":2759,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":16233,"displayContent":{"text":"                KIND_SESSIONS => {\n…\n                    // pre-bind); converges within one reconcile-poll window after bind.\r\n                    // [impl->REQ-HAZARD-CONTROL-STAMP-CONVERGENCE]\r\n                    // [impl->REQ-CONTROL-STAMP-CONVERGENCE]\r\n                    for s in &snaps {\r\n                        let _ = converge_perch_stamps(\r\n                            &s.endpoint,\r\n                            s.controller_by.as_deref(),\r\n                            s.has_controller,\r\n                            s.viewers,\r\n                            s.stamp_gen,\r\n                        );\r\n                    }\r\n                    let reply = SessionsReply {\r\n                        sessions: snaps\r\n                            .into_iter()\r\n                            .map(|s| SessionInfo {\r\n                                session_id: s.id,\r\n                                endpoint: s.endpoint,\r\n                                resume_seq: s.resume_seq,\r\n                                controller_by: s.controller_by,\r\n                                pid: s.pid,\r\n                                adapter: s.adapter,\r\n                                spawned_ms_ago: s.spawned_ms_ago,\r\n                            })\r\n                            .collect(),\r\n                        // The other half of the answer (releases#199): the\r\n                        // bring-ups this broker has started and not yet\r\n                        // registered. Read at reply time rather than snapshotted\r\n                        // with the sessions above, so a bring-up that registered\r\n                        // mid-handler is reported by the row it now has instead\r\n                        // of by an entry it no longer needs.\r\n                        // [impl->REQ-ER-BRINGUP-INFLIGHT-LEDGER]\r\n                        bringing_up: self.bringups_in_flight_at(Instant::now()),\r\n                    };\r\n                    let frame = Envelope::new(\r\n                        KIND_SESSIONS_REPLY,\r\n                        serde_json::to_value(reply).expect(\"SessionsReply serializes\"),\r\n                    );\r\n                    send_frame(&send, &frame);\r\n                }\r\n                KIND_ENDPOINT_INPUT => {\r\n                    if let Err(msg) = self.dispatch_endpoint_input(env, &send) {\r\n                        send_error(&send, &msg);\r\n                    }\r\n                }\r\n                KIND_NET_STREAM_SUBSCRIBE => match self.dispatch_net_stream_subscribe(env, &send) {\r\n                    Ok(id) => my_stream_subs.push(id),\r\n                    Err(msg) => send_error(&send, &msg),\r\n                },\r\n                // Explicit seat release (ADR-0040 decision 5): the released\r\n                // stream also leaves this conn's sweep list, so the exit\r","startLine":6133,"lineNumbers":[6133,null,6212,6213,6214,6215,6216,6217,6218,6219,6220,6221,6222,6223,6224,6225,6226,6227,6228,6229,6230,6231,6232,6233,6234,6235,6236,6237,6238,6239,6240,6241,6242,6243,6244,6245,6246,6247,6248,6249,6250,6251,6252,6253,6254,6255,6256,6257,6258,6259,6260,6261,6262]},"fileSize":844842,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-daemon\\src\\broker.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":16233,"totalBytes":2759,"outputLines":51,"outputBytes":2759,"shownRange":{"start":6212,"end":6262},"nextOffset":6263}}},"isError":false,"timestamp":1788999564799}}
{"type":"message","id":"d3395329","parentId":"818340c0","timestamp":"2026-09-10T00:19:24.801Z","message":{"role":"toolResult","toolCallId":"call_Ep2k17QHxISUknQNs4G8klUT|fc_04ba494197989c86016aa1f78cabf887d1b4c18052cc0f2330","toolName":"grep","content":[{"type":"text","text":" 50|\n*51|### 2.2 Stdin session_id precedence over env\n 52|- **Failure:** subagent inherits a stale `OWL_SESSION_ID` env across `/clear`; hook gets two session_ids (fresh stdin, stale env) → wrong-agent binding.\n 53|- **Invariant:** stdin-provided session_id wins; env is fallback only.\n 54|- **spt-core mapping:** the harness-contract subcommand surface must define the same precedence for whatever identity fields hooks pass in.\n...\n 406|\n*407|### 6.7 Broker and brain MUST be separate processes (in-process collapse silently breaks no-endpoint-drop update) `[REQ-HAZARD-BROKER-PROCESS-ISOLATION]`\n 408|- **Failure:** the daemon hosts the broker as a background *thread* in the single `spt daemon` process (`daemon.rs:165-170`, `Arc<Broker>` + `thread::spawn(serve)`) instead of a separate process. A brain restart onto a swapped binary then cannot happen without killing the broker thread — closing every PTY, orphaning every harness child, dropping every socket. So `spt update apply` degrades to an in-process `Brain::handoff` no-op: the binary swaps on disk but the running daemon keeps executing the old c...\n 409|- **Invariant:** the broker runs as its own long-lived process that survives every brain restart; the brain restarts onto the new binary and re-attaches via the versioned IPC. A routine (brain-only) update must leave every hosted endpoint untouched at the *process* level — not merely re-subscribe a brain within the same process. The evidence for REQ-UPD-3 / REQ-DAEMON-2 must prove process-level survival (a PTY child + a live QUIC conn survive a brain-process restart onto a swapped binary — SPIKE-01/0...\n 410|- **spt-core mapping:** restoration is ADR-0018 (next milestone). The current `int` tags on REQ-DAEMON-2 / REQ-UPD-3 are regression-masked and re-point at restoration; the broker becomes the always-up per-machine anchor (seed-lock + liveness + brain supervisor). Two-process supervision, generation custody, durable-deadline loop timing, broker-cursor-of-record, and readiness-gated auto-rollback all hang off this.\n...\n 425|\n*426|### 6.9 Resume-mode brain: a blocking spawn/command wait silently discards OTHER sessions' output\n 427|- **Failure:** a resume-mode brain (per-session `session_cursors` populated by `resume_sessions`) drives `Brain` over **blocking** `read_event` calls. Any command that loops `read_event` until its own reply — `spawn_session_pid` waiting for `Spawned`, `net_status`, `sessions`, etc. — calls `read_event` on *every* interleaved frame, so an OUTPUT frame for a **different** session is **cursor-processed** (its `session_cursors` entry snaps forward; the broker already counted it delivered via `delivered_t...\n 428|- **Invariant:** the daemon-hosted multi-session event loop must not consume a session's OUTPUT inside another session's blocking wait. The **live-agent adapter milestone must restructure the brain event pump** so command/`spawn` is non-blocking (a single demux loop owns `read_event` and routes every frame to its session's consumer), OR a blocking wait must re-queue/route the frames it reads for other sessions rather than dropping them.\n 429|- **spt-core mapping:** the **other-session** face is unreachable today — the supervised daemon brain hosts no PTY sessions and spawns none; a single-session seat (legacy, empty map) has no other session to starve. The **same-session** face is reachable in direct `Brain` users: `spawn_session_pid` can consume and discard that session's own first `Output` frame while waiting for `Spawned` (measured in-process on Linux, 2026-08-25), so a caller must not treat the producer's first chunk as guaranteed post...\n...\n 431|\n*432|### 6.10 Phase-significant loop timing must be a durable absolute-deadline grid, not phase-relative sleep `[REQ-HAZARD-BROKER-PROCESS-ISOLATION]`\n 433|- **Failure:** a periodic loop that sleeps a flat `period` each iteration (`pulse_tick` then `sleep(pulse_period)`) is **phase-relative** — every brain restart silently re-phases the grid to the restart instant. Under the seamless-update model (the supervisor respawns the brain onto a swapped binary, ADR-0018 D3-3), a routine update would shift the cadence of every phase-significant loop, and continuity cannot ride a brain→brain frame (the outgoing brain is gone before the new one starts — the same...\n 434|- **Invariant:** phase-significant periodic timing lives as durable absolute-deadline state on disk (`(anchor, interval)`), rehydrated on every brain start, with fires **derived functionally** (`next_fire = anchor + interval·⌈max(0,now−anchor)/interval⌉`) and **no per-fire write**. An **Update** restart re-reads the anchor and keeps deriving (phase preserved, lands mid-grid); a **Crash**/**Cold** restart re-bases the anchor to `now` (phase reset acceptable — the loop is idempotent catch-up). The...\n 435|- **spt-core mapping:** ADR-0018 Q4/V3/V4, restoration D5. Mechanism in `spt-daemon::deadline` (`DeadlineAnchor` periodic + `OneShotDeadline` rule-only pure helper); the pulse loop (`lifecycle::run_pulse_loop`) consumes it. The one-shot **machinery** (a durable in-daemon alarm scheduler) is the deferred alarm port (`docs/DEFERRED.md`) — the daemon has no one-shot consumer today, so building the timer now would ship untested dead code (activate-don't-pre-fail); D5 fixes the *rule* as a tested-unwired he...\n...\n 497|\n*498|### 7.6 Pump brain-IPC reads must be deadline-bounded (a blocked read wedges the whole pump)  `[REQ-HAZARD-PUMP-IPC-DEADLINE]`\n 499|<!-- [doc->REQ-HAZARD-PUMP-IPC-DEADLINE] -->\n 500|- **Failure:** the peer pump is a SINGLE thread driving every leg (registry/notif/sync/update) against every peer over ONE brain-IPC client. Its reply reads (`net_open_stream`, `net_stream_send`, `net_dial`, and the sync/update pull `read_event` loops) were `loop { read_event() }` with no deadline. When a peer's QUIC path black-holes, the broker's stream-open/send awaits the dead peer and never sends the reply, so the brain's `read_frame` blocks FOREVER and the pump freezes mid-round. The heartbeat (loop...\n 501|- **Invariant:** in PUMP mode the brain carrier is **SPLIT at construction** — a dedicated `pump-ipc-reader` thread does blocking `read_frame` on the `RecvHalf` and forwards each framed result down a channel; the main thread writes on the `SendHalf` and reads with `Receiver::recv_timeout`. Every IPC reply read is bounded by a per-call **total-wait** deadline (`PUMP_PEER_IO_TIMEOUT` = 30s, > any legitimate round-trip, < the 60s QUIC idle; re-armed on stream progress for the streaming pull legs so a heal...\n...\n 504|\n*505|### 7.7 A slow/dead/hostile remote VIEWER must never stall the controller, child, or drain  `[REQ-HAZARD-VIEWER-ISOLATION]`\n 506|<!-- [doc->REQ-HAZARD-VIEWER-ISOLATION] -->\n 507|- **Failure:** the W2.5 controller/viewer model lets ANY number of read-only `--view` attachers ride one session's broker `OutputLog`. The single drain thread fans each output chunk to every attacher. If a viewer's socket is fanned out with a **blocking** write under the log lock (the controller's authoritative path), one wedged viewer (a slow terminal, a black-holed WAN peer, a hostile non-reader) stalls the drain — freezing the controller's stream and backing up the PTY child. A single watcher must n...\n 508|- **Invariant:** the drain writes the **controller** on the authoritative blocking bounded path (it alone advances `delivered_through`), but each **viewer** gets an **isolated bounded SPSC queue + a dedicated writer thread**; the drain `try_send`s under the log lock and **evicts** any viewer whose queue is `Full` (fell behind the live stream) or `Disconnected` (its writer died on a dead socket) — the drain thread **never touches a viewer socket**, so no viewer write can backpressure it. A **soft cap** ...\n...\n 511|\n*512|### 7.8 The broker must never make a brain wait UNBOUNDED on a QUIC op (the pump-IPC-deadline B-half)  `[REQ-HAZARD-BROKER-QUIC-DEADLINE]`\n 513|<!-- [doc->REQ-HAZARD-BROKER-QUIC-DEADLINE] -->\n 514|- **Failure:** the broker's brain-facing QUIC handlers (`dispatch_net_dial` / `dispatch_net_stream_open` / `dispatch_net_stream_send`) call into `NetHost::dial` / `open_stream` / `send_stream`, whose iroh awaits (`endpoint.connect` + `prove_membership`; `open_bi`; `write_all`/`finish`) had NO bound of their own. A dead/black-holed roster peer (its process gone, or a mixed-pair that accepts the conn but never answers the seed-proof) makes the broker await its QUIC path FOREVER, so the brain escapes only v...\n 515|- **Invariant:** every brain-waiting QUIC op is wrapped in a broker-side deadline (`NetHost::bounded_block_on` → `tokio::time::timeout`, `BROKER_QUIC_OP_TIMEOUT_MS` = 10s). On elapse the future is DROPPED (cancelling the in-flight connect/stream op, so nothing is half-registered) and a non-`TimedOut` `io::Error` is returned, which the broker REPLIES as an ordinary error frame. The bound (10s) sits comfortably above any legitimate LAN/relay round-trip and 20s below the brain's 30s `PUMP_PEER_IO_TIMEOUT`...\n...\n 518|\n*519|### 7.9 A daemon-state wire change needs a deliberate BROKER restart (the broker is resident across a brain self-update)  `[REQ-HAZARD-BROKER-SEED-WIRE-SKEW]`\n 520|<!-- [doc->REQ-HAZARD-BROKER-SEED-WIRE-SKEW] -->\n 521|- **Failure:** the broker serves the seed-control channel and is RESIDENT across a brain-only self-update (ADR-0004's no-terminate-during-update pillar forbids auto-killing it — 6.7). A self-update that changes a daemon-state WIRE FORMAT — e.g. the v0.9.0 adapter-agnostic `Seed` (the `adapter` field dropped) — therefore lands a NEW-version CLI talking to the STILL-RESIDENT OLD broker. The old broker cannot deserialize the new `Seed` (its formerly-required `adapter` is absent), so it drops the seed-...\n 522|- **Invariant:** (a) spt-core surfaces an ACTIONABLE diagnostic on the seed-ack EOF — naming the stale-broker cause + the fix (`spt daemon stop`; the broker restarts on the next `spt api` call) — never the bare io error (scoped to `UnexpectedEof` on the seed-ack path, so it never mis-fires on an unrelated error like a refused connect). (b) A daemon-state wire change requires a DELIBERATE full broker restart; this is NOT automatic — ADR-0004 forbids auto-killing the resident broker, and a brain-only...\n...\n 532|\n*533|### 7.11 A dead PTY child + a dropped operator pump must NOT wedge the broker for other clients  `[REQ-HAZARD-ATTACH-WEDGE]`\n 534|- **Failure (hypothesized, v0.12.0):** a legitimately dead PTY child (crash/kill) plus an `rc` pump dropped without a clean detach (closed tab) was thought to make the broker's loopback forward `write_all` block forever on a full 64 KB duplex, park a worker in the 2-worker net runtime, saturate both, and stall every new attach / `endpoint run` (with `daemon stop` unable to join).\n 535|- **Invariant — DISPOSITION = PROVE-DON'T-CHANGE** (doyle GATE-PASS @e883f45, 2026-06-18): the post-L0 code ALREADY prevents the wedge; NO fail-fast / worker-count code was added. (1) `serve_attach` forwards fire-and-forget (`net_stream_send` `op_id=None`) and the broker-side `send_stream` is already deadline-bounded (`bounded_block_on`, `BROKER_QUIC_OP_TIMEOUT_MS` = 10s — hazard 7.8), not forever. (2) the loopback duplex is drained broker-INTERNALLY by the operator row's OWN read pump (`nethost.rs` ...\n 536|- **spt-core mapping:** int `crates/spt/tests/attach_wedge_e2e.rs` (REAL detached daemon + dummy-harness fixture): serve the victim (rc sees its tick), abruptly kill rc (dropped pump) + kill the PTY child → a NEW endpoint still comes online + is served (no wedge), the dead endpoint is offlined within a tick (`LIVENESS_RECONCILE_OFFLINE`), `daemon stop` bounded. Leans on 7.8 (the QUIC-op deadline) + 7.7 (viewer isolation).\n...\n 538|\n*539|### 7.12 Controller output must NOT be written inline on the drain thread — a backed-up controller wedges the session  `[REQ-HAZARD-INJECT-CONTROL-COEXIST]`\n 540|- **Failure (v0.12.x, operator dogfooding + doyle /diagnose 2026-06-19):** `OutputLog::append` fanned each live chunk to the CONTROLLER via a SYNCHRONOUS, blocking `write_frame` held INLINE on the session's single drain thread while `Mutex<OutputLog>` was locked (viewers already had a dedicated writer thread + bounded evicting channel). A backed-up controller socket — a slow operator, or the full 64 KB `rc` loopback duplex under heavy TUI redraw — parked the drain thread WITH THE LOG LOCK HELD: outpu...\n 541|- **Invariant:** controller delivery runs OFF the drain thread — a dedicated controller writer thread + bounded channel (the 7.7 viewer_writer pattern) does the blocking socket write; the drain hands each chunk off with a BOUNDED, OFF-LOCK send (deadline → detach + `clear_controller`, never park forever, never silently evict a LIVE operator's authoritative view). The controller is AUTHORITATIVE (unlike a viewer): its writer advances the `delivered_through` resume cursor (atomic, only-on-success, mono...\n 542|- **spt-core mapping:** impl `broker.rs` `OutputLog::append`/`controller_writer`/`become_controller`/`ControllerJob::deliver`/`clear_controller`; int `crates/spt-daemon/tests/inject_control_wedge.rs` (a fully-backed-up controller keeps `KIND_SESSIONS` answering = no wedge); unit the bounded-deliver + monotonic-cursor kernels. The input-side parks (write_input `write_all` on a full buffer; DSR-answer writer-mutex contention) are BENIGN on Windows ConPTY (absorbs a large inject), real only on Unix forkpty ...\n...\n 557|\n*558|### 7.15 An OFFLINE spt-hosted endpoint must NOT render phantom `ONLINE+CONTROLLED` — clear `driven_by` when its session is gone  `[REQ-HAZARD-DRIVEN-BY-SELFHEAL]`\n 559|- **Failure:** `driven_by` (the `info.json` `ONLINE+CONTROLLED` latch) is single-written by the broker via `clear_controller`/`stamp_driven_by`, which only fire on a controller change. When an spt-hosted endpoint's broker session is GONE (harness dead — the B2 case, 7-series sibling `REQ-HAZARD-HOSTED-LIVENESS-RECONCILE`), no controller event ever fires, so a stale `driven_by=Some(node)` persists: the picker renders a phantom \"controlled by X\" on an endpoint that is actually OFFLINE. The B2 reconcile a...\n 560|- **Invariant:** `reconcile_hosted_liveness`, when it offlines a sessionless controllable perch (the B2 keystone — no live broker session ⇒ dead harness), ALSO clears `driven_by` (`set_driven_by(perch, None)`). RACE-FREE and single-writer-safe: with NO live broker session there is no controller to re-stamp `driven_by` concurrently, so the brain may write it here without contending the broker. (The LIVE-session leg — a controller gone while its session survives — is NOT this hazard: a clean discon...\n 561|- **MEASURED 2026-07-22 (DAEMON-LIFECYCLE W2 Leg B, `transport_death_eof.rs`) — the residual SPLITS in two, and only one half is still open.** On REAL QUIC (two brokers, `BindScope::Loopback` + a real `net_dial` — NOT the in-process duplex, which has no idle timeout at all), a remote controller whose QUIC STACK DIES WITHOUT A FIN (peer frozen: endpoint alive, UDP port still bound, keepalives unanswered — the sleeping-laptop shape) DOES self-heal: the transport's own idle timeout tears the conn, the...\n...\n 581|<!-- [doc->REQ-HAZARD-INPUT-ACK-BACKPRESSURE] -->\n*582|### 7.19 An operator input FLOOD must not deadlock the broker via the applied-ack on the same conn  `[REQ-HAZARD-INPUT-ACK-BACKPRESSURE]`\n 583|- **Failure (operator HITL, the ctrl+V re-open):** a flood of operator input on one brain↔broker conn wedged the WHOLE broker PERMANENTLY (no new/existing attach; the controller stayed latched — the per-conn handler couldn't process the detach). `serve_attach` processes a whole `NetStreamData` batch of N `Input` records in its inner loop, calling `send_effect` N times WITHOUT returning to `read_event()`; the broker answers each with `send_frame(applied_envelope)` on the SAME conn. Brain not reading ...\n 584|- **Invariant:** the applied-ack is OPT-IN. `InputReq` carries `ack: bool` (serde `default = true`, N-1-safe). The fire-and-forward operator/rc path (`serve_attach`) sends `ack=false` via `Brain::send_effect_no_ack`; `dispatch_input` writes NO applied frame when `ack=false`, so the per-conn handler never writes back while servicing the flood → it always drains → no deadlock (cures ANY input flood). `shellchan` (one-at-a-time spool delivery, WAITS on `BrokerEvent::Applied`) keeps `send_effect` (`ack=t...\n 585|- **spt-core mapping:** impl `msg.rs` `InputReq.ack` (`default_true`) + `brain.rs` `send_effect_no_ack`/`send_effect_inner` + `attach.rs` `serve_attach` operator path → `send_effect_no_ack` + `broker.rs` `dispatch_input` gates `send_frame(applied)` on `req.ack`; unit the serde default + the ack-emitted-iff-true + the no-ack-path exactly-once dedup; int (keystone, repro-first) a flood of N>pipe-buffer input frames through `serve_attach` on one conn — PRE-FIX deadlocks, POST-FIX drains all N + the sess...\n...\n 595|<!-- [doc->REQ-HAZARD-CONTROLLER-WRITER-REORDER] -->\n*596|### 7.21 Exactly ONE `controller_writer` per brain↔broker connection — a superseded writer must write nothing further  `[REQ-HAZARD-CONTROLLER-WRITER-REORDER]`\n 597|- **Failure (doyle instrumented RACEDIAG repro, kitsubito):** on a brain-restart re-serve the handoff brain registered as controller on the SAME session TWICE over the SAME socket — `Brain::handoff` eagerly `subscribe(prior.next_seq=1)` → `become_controller(from_seq=1)`, spawning writer-A (writes seq 1); then `serve_attach` re-handled the replayed `Request{from_seq:0}` → `attach_as(sid,0)` → `become_controller(from_seq=0)`, spawning writer-B (writes 0,1). `become_controller` dropped the prior `Co...\n 598|- **Invariant:** on a single brain↔broker connection exactly ONE `controller_writer` is ever the LIVE writer; a SUPERSEDED writer writes no further frames after the epoch bump it observes; and **every `controller_writer` emits a strictly ASCENDING seq stream** (sorted initial batch + ascending live frames). The CORRECTNESS guarantee that falls out: a snap-above consumer over any interleaving of ascending writers — where the surviving writer (`serve_attach`'s `attach_as(sid,0)`) offers the COMPLETE ra...\n 599|- **spt-core mapping:** impl `brain.rs` `Brain::handoff` (KEEP the eager `subscribe`; seed `session_cursors`) + `Brain::subscribe_with` (resume-mode dedup-cursor reset to `from_seq`, shared by `attach`/`attach_as` — the operator-stream boundary fix) + `broker.rs` `OutputLog.controller_epoch: Arc<AtomicU64>` / `become_controller` (atomic `fetch_add`, passes the new epoch + `Arc::clone` into the writer) / `controller_writer` (epoch gate read UNDER the lock on both loops) / `mark_controller_gone` + the `C...\n...\n 638|\n*639|### 7.25 A perch PINNED to a DEAD session self-heals (dead-owner re-pin) instead of wedging forever; a LIVE-owner rotation still refuses  `[REQ-HAZARD-SESSION-PIN-WEDGE]`\n 640|- **Failure (F-024C/F-024D, ENLYZEAM field + clean-room repro 2026-07-02):** `authenticate()` (auth.rs) gates `api poll`/`state`/`boundary` on `proof.session_id == info.json.session_id`. If ONE boundary rotation is LOST — the departing session dies (crash / tab-close) or its `/clear`-era `boundary` call never lands — the perch stays PINNED to the dead sid, and every id-scoped hook call thereafter AUTH_REFUSES *including `boundary` itself* (the very call that would re-pin presents the NEW sid). A perm...\n 641|- **Invariant:** `authenticate()` gains a DEAD-OWNER fallback — when the caller's sid MISMATCHES the pin **AND** the perch's recorded owner pid is DEAD (`proc::is_process_alive == false`), ACCEPT the caller's sid and RE-PIN (rotate `info.json.session_id` via the locked `mutate_info` RMW + a LOUD `SESSION_REPIN` stderr line). Same trust model as `establish_perch`'s conflict gate, which already allows a rebind exactly when `owner_alive == false` (an orphaned perch accepts a new LOCAL owner). This is **AD...\n 642|- **spt-core mapping:** impl `spt/src/api/auth.rs` `authenticate` dead-owner-mismatch branch (`read_pid` → `proc::is_process_alive` → `mutate_info` rotate + `SESSION_REPIN` log); unit `auth.rs` `pinned_to_dead_sid_mismatched_poll_repins` (AUTH_REFUSED → drains+re-pins, RED pre-fix) + `live_owner_mismatch_still_refuses` (the `/clear` live-rotation squat guard, unwidened) + `token_auth_path_unchanged` (correct token wins first, no rotation).\n...\n 655|<!-- [doc->REQ-HAZARD-CONTROL-STAMP-LIFETIME] -->\n*656|### 7.27 A control/viewer stamp must NEVER outlive its session — every teardown path clears what attach stamped  `[REQ-HAZARD-CONTROL-STAMP-LIFETIME]`\n 657|- **Failure (F-026 #2, live evidence HFENDULEAM):** an spt-hosted endpoint stayed `ONLINE+CONTROLLED` after the operator's RC `/exit` — hours later hall-a's `info.json` still read `controlled:true` (status offline, dormant) and hfenduleam still GOSSIPED `controller_node=self` cross-node (SPT_DEV.json snapshot). ROOT: `/exit` kills the CHILD, not the controller connection, so the endpoint suspends via child termination and the teardown runs the REAP path (`broker.rs` exit-waiter → `sessions.remove(&id...\n 658|- **Invariant:** on session reap the broker clears the perch's control/viewer stamps — `OutputLog::stamp_reaped()` = `set_driven_by(None)` + `set_controlled(false)` + `set_viewer_count(0)` via the known endpoint id — so no teardown path leaves a stamp behind. The broker stays the SINGLE writer; the clear is unconditional (idempotent when already clear) and RACE-FREE — the session is dead, so there is no live controller to re-stamp concurrently (the same argument as 7.15's no-session leg). Best-effo...\n 659|- **spt-core mapping:** impl `broker.rs` `OutputLog::stamp_reaped` + the exit-waiter reap call (before `sessions.remove`); int (real broker+PTY, no mocks) `crates/spt-daemon/tests/control_stamp_lifetime.rs` `reap_clears_control_and_viewer_stamps` (spawn hosted session → the live controller stamps `controlled=true` → latch `driven_by`+`viewer_count` on disk → KILL the child → the reap clears all three; RED-first verified: fix disabled leaves `driven_by=Some`, `controlled=true`, `viewer_count=Some(...\n...\n 674|<!-- [doc->REQ-HAZARD-CONTROL-STAMP-CONVERGENCE] -->\n*675|### 7.29 Control/viewer stamps CONVERGE to broker session-table truth, not merely edge-trigger  `[REQ-HAZARD-CONTROL-STAMP-CONVERGENCE]`\n 676|- **Failure (F-026 stamp-gap, hall-b + the original ball-b):** a picker-created endpoint (`endpoint run` → new) read plain `ONLINE` in the list + picker while genuinely driven — `info.json` `controlled:false` throughout. ROOT (the UPWARD companion to 7.27's downward edge-clear): the broker spawn path's `become_controller` → `stamp_driven_by` → `set_controlled(true)` fires at SPAWN time, but a FRESH endpoint has NO PERCH yet (the adapter binds it after claude boots), so `mutate_info` returns `NotF...\n 677|- **Invariant:** the broker (SINGLE writer) re-asserts each live session's control/viewer stamps to session-table TRUTH, DIVERGENCE-GATED — read the perch's recorded `(driven_by, controlled, viewer_count)`, compare to the live truth (`controller_by()` / `has_controller()` / live `viewer_count`), and write ONLY on a real difference (no per-poll fsync storm). The re-assert rides the `KIND_SESSIONS` handler (the daemon reconcile + picker poll it on cadence), so a fresh perch CONVERGES within one reconcile...\n 678|- **Coverage boundary (who MAY clear vs who MUST NOT — cross-ref 7.15):** the convergence's `driven_by`-clear leg fires ONLY when `has_controller()==false` — the DEFINITIVE broker signal (it reads its OWN controller slot, not the ambiguous `controller_by`). That is why the broker MAY clear a stale `driven_by` where the BRAIN reconcile (7.15) MUST NOT: the brain only sees `controller_by==None`, which a live LOCAL controller also reads, so a brain clear would false-heal a genuinely-driven session. This...\n...\n 702|\n*703|### 7.33 NO bare `.lock().unwrap()` on a broker-resident lock reachable from serve/dispatch — a poison permanently wedges every attach  `[REQ-HAZARD-BROKER-FLOOR-LOCK-POISON]`\n 704|<!-- [doc->REQ-HAZARD-BROKER-FLOOR-LOCK-POISON] -->\n 705|- **Failure (paid-for class):** a brain-only self-update restarts the BRAIN but keeps the BROKER process — and every one of its `Mutex`es — ALIVE by design (REQ-UPD-3). So a single panic while another thread held a broker-resident lock POISONS it permanently: the next `.lock().unwrap()` panics too, kills its per-conn reply thread, and EVERY subsequent attach silently deadlines (`spt rc` → \"brain IPC read deadline elapsed\") while non-locked ops keep working. The effect journal (bug #16, `[REQ-HAZARD...\n 706|- **Invariant:** NO broker-resident lock reachable from serve/dispatch may use a bare `.lock().unwrap()`. Recover via the shared `recover` (`into_inner` — safe for the short, coherent-on-recovery map ops of `sessions`/`pair_holds`) or, for the `OutputLog` ring, `recover_log` — which adds a COHERENCE CLAMP on the recovery path: a panic mid-`append` can leave the ring torn (over-cap, a last seq not below `next_seq`, non-monotonic), and serving those bytes risks garbage, so `clamp_or_reset` cheap-checks...\n...\n 722|\n*723|### 7.36 The broker control plane and PTY fan-out must NEVER block on a single subscriber connection — a suspended brain conn must not wedge control  `[REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]`\n 724|<!-- [doc->REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE] -->\n 725|- **Failure (paid-for, rig-CONFIRMED 2026-07-06/07 — `NtSuspendProcess` on the brain, no update involved):** a controller's writer thread does a BLOCKING socket write to its brain subscriber conn. When that brain is suspended (or black-holed) the write never returns. The output-driven eviction path (`append` → `mark_controller_gone`, bounded by 7.12's `CONTROLLER_WRITE_DEADLINE`) only fires on NEW output, and `reap_dead_controller` (7.29 companion) only caught a writer that had EXITED — so a contro...\n 726|- **Invariant:** no broker client ever observes a wedge because the BRAIN stalled. (1) The controller's writer publishes its IN-FLIGHT-write window (`write_blocked_since`: `Some(Instant)` set immediately before the blocking `write_frame`, `None` after — the mutex never held ACROSS the write); a reader sees a `Some` older than `BRAIN_WRITE_DEADLINE` (15 s = 3× the 7.12 controller bound, measured full-with-**zero writer progress**, so a slow-but-DRAINING brain never trips it) as a WEDGED writer and stal...\n...\n 729|\n*730|### 7.37 The Layer-1 settle-gate must RE-ARM per delivery on an observable PTY — a mid-session reader reattach (`/clear`) re-creates the head-swallow race  `[REQ-HAZARD-INJECT-SETTLE-REARM]`\n 731|<!-- [doc->REQ-HAZARD-INJECT-SETTLE-REARM] -->\n 732|- **Failure (paid-for, field-CONFIRMED on 0.29.0 — doyle diagnosis + perri screenshot `WindowsTerminal_6iSjya8pMt.png`):** the shipped W5-A settle-gate (7-of-W5, `REQ-INJECT-MULTILINE-INTEGRITY`) gated Layer 1 behind a worker-local ONE-SHOT (`settled_once`) on the premise that the head-swallow race is a STARTUP condition (the input reader not yet attached after spawn). That premise is FALSE: a mid-session `/clear` re-enters the harness's raw-mode input reader, re-attaching it and RE-CREATING the pre-se...\n 733|- **Invariant:** the settle-gate re-arms before EVERY delivery on an OBSERVABLE (echoing/interactive) PTY — the bug-prone class, where re-settling is cheap (a tick or two once the reader answers) and necessary (a mid-session reader reattach must be re-confirmed). The steady-state settle is latch-skipped ONLY where the probe is known UNOBSERVABLE on this PTY (a non-echoing ConPTY whose DSR answer never surfaces): there is no reader-reattach race to guard, and each settle would burn the full `INJECT_SETT...\n...\n 736|\n*737|### 7.38 EVERY write on a physical broker connection is bounded + cancelable + poison-on-failure — no writer holds the send gate across an UNBOUNDED OS write  `[REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]`\n 738|<!-- [doc->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK] -->\n 739|- **Failure (paid-for, field capture + deterministic Windows repro 2026-07-09 — hertz RCA, UPDATE-WEDGE round-4):** `controller_writer` held the `SharedSend = Arc<Mutex<SendHalf>>` guard ACROSS a blocking `write_frame`. On Windows, `interprocess 2.4.2` routes the send to `WriteFileEx` + `SleepEx(INFINITE, alertable)` — this local-socket send path exposes **no supported write timeout** (`set_timeout` → `Unsupported`). An `rc --take` controller consumer that stopped reading blocked the write **indefi...\n 740|- **Invariant:** every write on a physical broker connection rides ONE broker-owned bounded/cancelable framed-write primitive with an **independent out-of-band aborter**. The `SendHalf` never leaves the conn object (a bypass cannot compile). An ABSOLUTE deadline stamped at write entry covers BOTH the serialized gate-wait AND OS write completion; a per-conn watchdog fires at that deadline and aborts the in-flight op — never relying on the write returning or on opportunistic stall-evict sampling. On dead...\n...\n 745|\n*746|### 7.39 Per-session identity env (`SPT_ENDPOINT_ID`/`OWL_SESSION_ID`/`SPT_AGENT_ID`) is NEVER inherited — the daemon scrubs it at startup AND on every role spawn, regardless of any role's declared `env_remove`  `[REQ-HAZARD-DAEMON-IDENTITY-ENV-SANITIZE]`\n 747|<!-- [doc->REQ-HAZARD-DAEMON-IDENTITY-ENV-SANITIZE] -->\n 748|- **Failure (paid-for, perri field RCA 2026-07-09/10 — F-036 psyche seat-theft):** a daemon restarted from inside an agent session (routine during core dev / `spt update apply`) carried the session's `SPT_ENDPOINT_ID=doyle` and passed it verbatim into every `[session.psyche_resume]` spawn — core stripped only each role's DECLARED `env_remove` list (runtime.rs `command_for`), so ONE adapter `env_remove` miss infected the whole node. Every psyche claude turn fired SessionStart, the adapter hook saw the...\n 749|- **Invariant:** the identity set (`spt_runtime::IDENTITY_ENV_VARS`) is per-session state and NEVER correct inherited state for a daemon or any child it spawns. The daemon (broker AND brain entries) scrubs its own process env first thing (`scrub_identity_env`); the shared role-spawn command builder scrubs the set unconditionally and LAST (after `env_remove`, read-env stamps, and the recursion guard), so no adapter declaration — missing or pathological — can leak identity into a role child. A spawn th...\n...\n 763|\n*764|### 7.41 A fresh dispatcher must NEVER re-serve a terminal stream — historical replay must not steal or clear a live controller  `[REQ-HAZARD-REDISPATCH-CONTROL-STEAL]`\n 765|<!-- [doc->REQ-HAZARD-REDISPATCH-CONTROL-STEAL] -->\n 766|- **Failure (paid-for, hertz field RCA 2026-07-16 — 4/5 endpoints frozen, twice in one day):** every brain cycle (`spt daemon refresh`, `spt update` apply — same `applyhost.rs` path) launches a fresh `dispatch::run_dispatch_loop` whose EMPTY per-process `claimed` set enumerates every broker-held peer stream. `NetShared.streams` has NO removal path (`StreamLog::finish` only marks) and the claim condition has no finished filter → the fresh dispatcher replays FINISHED historical attach streams. A repl...\n 767|- **Invariant:** redispatch eligibility is LIFECYCLE-gated: a finished/terminal stream row is retired from the dispatcher's enumeration (a fresh dispatcher never re-serves a terminal Attach); classification identity is restart-durable and independent of the evictable data ring (immutable opener fact pinned per stream until close); claims are retryable on transient worker-setup failure and terminal on terminal outcomes (no hot-loop). The legitimate same-`by` successor re-take after a brain restart still s...\n...\n 789|\n*790|### 7.44 Streams and seats on a long-lived connection must have bounded lifetime — one-way rows terminal at FIN, seats released at serve completion, no per-chunk full-state rewrites in a drain loop  `[REQ-HAZARD-REGISTRY-STALL]`\n 791|<!-- [doc->REQ-HAZARD-REGISTRY-STALL] -->\n 792|- **Failure (paid-for, hertz post-close v0.36 field RCA 2026-07-17 — live prod box):** FOUR compounding consequences of \"physical teardown waits for conn close\" on connections that never close. The registry pump opens one fresh stream per feed (~30s) on the persistent pump conn; nothing retires the rows (the dispatcher — sole `retire_stream` caller — skips `initiated_locally`) → sender history unbounded. `stream_infos` filters `retired` only → every dispatcher poll serializes O(history) rows ov...\n 793|- **Invariant:** a one-way (fire-and-forget) family's row is TERMINAL at successful FIN — the sender retires it; eligibility filtering happens server-side before IPC serialization; feed application is transactional per feed (snapshot writes O(feeds), never O(chunks × record-kinds) — no synchronous full-state rewrite inside a per-chunk drain iteration); a deadline-poisoned one-way replay retires terminal at a per-stream strike budget (safe: the next pump round re-advertises); every completed serve re...\n...\n 803|\n*804|### 7.46 An rc surface answers from live session authority, never a stale persisted projection — and a resuming perch is UNBOUND, not offline  `[REQ-HAZARD-RC-ATTACH-TRUTH]`\n 805|<!-- [doc->REQ-HAZARD-RC-ATTACH-TRUTH] -->\n 806|- **Failure (paid-for, hertz perri contradiction RCA 2026-07-17/18 + operator field recovery):** the broker hosted an honest live session (client tree alive, `SessionProbe::has_live_session_honest` = true) while the persisted perch row said `offline` — the resumed harness sat wedged at a failed native resume, SessionStart never fired, the perch never bound, inbound messages black-holed. `spt rc` trusted the offline row and refused pre-IPC while `endpoint run --resume` reattached to the very same sessio...\n 807|- **Invariant:** every rc-family reader consults the single honest-session authority (ADR-0041 `SessionProbe`) before any persisted-status fast-fail — honest session ⇒ attach; no session ⇒ refuse; dead client tree ⇒ refuse/reap, never attach. UNBOUND = broker session exists + harness not bound, fresh/resume-invariant: resume stamps an existing offline perch UNBOUND (generation-safe rollback on spawn-fail/session death), bind owns UNBOUND→ONLINE. Refusal copy states known facts (harness-hosted ...\n...\n 817|\n*818|### 7.48 At most one input-capable controller lease per PTY session — takeover revokes atomically and loudly, input is fenced to the active lease, node identity is never a lease  `[REQ-HAZARD-CONTROLLER-LEASE]`\n 819|<!-- [doc->REQ-HAZARD-CONTROLLER-LEASE] -->\n 820|- **Failure (paid-for, hertz same-machine `--take` RCA, field repro 2026-07-16):** terminal A controlled an endpoint; terminal B on the SAME machine ran `spt rc --take`. Local loopback attaches carry only NODE identity, so `resolve_subscribe` computed `same_identity=true` and took the silent successor path for a distinct `--take` — intent never consulted; sink replaced with no `Displaced`, no stream close (the only loud branch is different-remote+Take). `dispatch_input` is session-addressed with no con...\n 821|- **Invariant:** each rc invocation/attach mints a unique controller LEASE id carried through subscribe, the controller slot, and Input/Resize; the only silent successor is same-lease + equal-or-newer generation (dispatcher recovery, ADR-0038 fix-6 preserved); explicit Take on a DISTINCT lease always revokes ATOMICALLY and authoritatively (old lease fenced + old attach stream FORCE-closed; the Displaced notice via the old writer is best-effort — a Full-queue notice drop must never leave the incumbent i...\n...\n 981|\n*982|### 7.64 Session-scoped authority does not survive a session-id rotation at rest — a boundary CARRIES it as a move, never leaves a copy  `[REQ-HAZARD-EMPOWER-SESSION-RESURRECT]`\n 983|\n 984|<!-- [doc->REQ-HAZARD-EMPOWER-SESSION-RESURRECT] -->\n 985|- **Failure (LOCKSMITH t1, releases#117; MEASURED by red probe 2026-08-03, not inferred):** an engine-room empowerment is session-state — ADR-0052 says \"nothing to revoke at rest\" — and is filed under the session id the engine room's perch is bound to. `spt api boundary` rotates exactly that id, to an ARBITRARY `--to-session-id` past a guard that refuses only psyche-custody squats, and it is NEITHER of the two revocation edges ADR-0052 §3a names. The record was therefore left filed under the OLD ses..."}],"details":{"scopePath":".worktrees/286-churn/docs/KNOWN-HAZARDS.md","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\docs\\KNOWN-HAZARDS.md","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":26,"fileCount":1,"files":[".worktrees/286-churn/docs/KNOWN-HAZARDS.md"],"fileMatches":[{"path":".worktrees/286-churn/docs/KNOWN-HAZARDS.md","count":26}],"truncated":false,"displayContent":"  50│\n *51│### 2.2 Stdin session_id precedence over env\n  52│- **Failure:** subagent inherits a stale `OWL_SESSION_ID` env across `/clear`; hook gets two session_ids (fresh stdin, stale env) → wrong-agent binding.\n  53│- **Invariant:** stdin-provided session_id wins; env is fallback only.\n  54│- **spt-core mapping:** the harness-contract subcommand surface must define the same precedence for whatever identity fields hooks pass in.\n    │...\n 406│\n*407│### 6.7 Broker and brain MUST be separate processes (in-process collapse silently breaks no-endpoint-drop update) `[REQ-HAZARD-BROKER-PROCESS-ISOLATION]`\n 408│- **Failure:** the daemon hosts the broker as a background *thread* in the single `spt daemon` process (`daemon.rs:165-170`, `Arc<Broker>` + `thread::spawn(serve)`) instead of a separate process. A brain restart onto a swapped binary then cannot happen without killing the broker thread — closing every PTY, orphaning every harness child, dropping every socket. So `spt update apply` degrades to an in-process `Brain::handoff` no-op: the binary swaps on disk but the running daemon keeps executing the old c...\n 409│- **Invariant:** the broker runs as its own long-lived process that survives every brain restart; the brain restarts onto the new binary and re-attaches via the versioned IPC. A routine (brain-only) update must leave every hosted endpoint untouched at the *process* level — not merely re-subscribe a brain within the same process. The evidence for REQ-UPD-3 / REQ-DAEMON-2 must prove process-level survival (a PTY child + a live QUIC conn survive a brain-process restart onto a swapped binary — SPIKE-01/0...\n 410│- **spt-core mapping:** restoration is ADR-0018 (next milestone). The current `int` tags on REQ-DAEMON-2 / REQ-UPD-3 are regression-masked and re-point at restoration; the broker becomes the always-up per-machine anchor (seed-lock + liveness + brain supervisor). Two-process supervision, generation custody, durable-deadline loop timing, broker-cursor-of-record, and readiness-gated auto-rollback all hang off this.\n    │...\n 425│\n*426│### 6.9 Resume-mode brain: a blocking spawn/command wait silently discards OTHER sessions' output\n 427│- **Failure:** a resume-mode brain (per-session `session_cursors` populated by `resume_sessions`) drives `Brain` over **blocking** `read_event` calls. Any command that loops `read_event` until its own reply — `spawn_session_pid` waiting for `Spawned`, `net_status`, `sessions`, etc. — calls `read_event` on *every* interleaved frame, so an OUTPUT frame for a **different** session is **cursor-processed** (its `session_cursors` entry snaps forward; the broker already counted it delivered via `delivered_t...\n 428│- **Invariant:** the daemon-hosted multi-session event loop must not consume a session's OUTPUT inside another session's blocking wait. The **live-agent adapter milestone must restructure the brain event pump** so command/`spawn` is non-blocking (a single demux loop owns `read_event` and routes every frame to its session's consumer), OR a blocking wait must re-queue/route the frames it reads for other sessions rather than dropping them.\n 429│- **spt-core mapping:** the **other-session** face is unreachable today — the supervised daemon brain hosts no PTY sessions and spawns none; a single-session seat (legacy, empty map) has no other session to starve. The **same-session** face is reachable in direct `Brain` users: `spawn_session_pid` can consume and discard that session's own first `Output` frame while waiting for `Spawned` (measured in-process on Linux, 2026-08-25), so a caller must not treat the producer's first chunk as guaranteed post...\n    │...\n 431│\n*432│### 6.10 Phase-significant loop timing must be a durable absolute-deadline grid, not phase-relative sleep `[REQ-HAZARD-BROKER-PROCESS-ISOLATION]`\n 433│- **Failure:** a periodic loop that sleeps a flat `period` each iteration (`pulse_tick` then `sleep(pulse_period)`) is **phase-relative** — every brain restart silently re-phases the grid to the restart instant. Under the seamless-update model (the supervisor respawns the brain onto a swapped binary, ADR-0018 D3-3), a routine update would shift the cadence of every phase-significant loop, and continuity cannot ride a brain→brain frame (the outgoing brain is gone before the new one starts — the same...\n 434│- **Invariant:** phase-significant periodic timing lives as durable absolute-deadline state on disk (`(anchor, interval)`), rehydrated on every brain start, with fires **derived functionally** (`next_fire = anchor + interval·⌈max(0,now−anchor)/interval⌉`) and **no per-fire write**. An **Update** restart re-reads the anchor and keeps deriving (phase preserved, lands mid-grid); a **Crash**/**Cold** restart re-bases the anchor to `now` (phase reset acceptable — the loop is idempotent catch-up). The...\n 435│- **spt-core mapping:** ADR-0018 Q4/V3/V4, restoration D5. Mechanism in `spt-daemon::deadline` (`DeadlineAnchor` periodic + `OneShotDeadline` rule-only pure helper); the pulse loop (`lifecycle::run_pulse_loop`) consumes it. The one-shot **machinery** (a durable in-daemon alarm scheduler) is the deferred alarm port (`docs/DEFERRED.md`) — the daemon has no one-shot consumer today, so building the timer now would ship untested dead code (activate-don't-pre-fail); D5 fixes the *rule* as a tested-unwired he...\n    │...\n 497│\n*498│### 7.6 Pump brain-IPC reads must be deadline-bounded (a blocked read wedges the whole pump)  `[REQ-HAZARD-PUMP-IPC-DEADLINE]`\n 499│<!-- [doc->REQ-HAZARD-PUMP-IPC-DEADLINE] -->\n 500│- **Failure:** the peer pump is a SINGLE thread driving every leg (registry/notif/sync/update) against every peer over ONE brain-IPC client. Its reply reads (`net_open_stream`, `net_stream_send`, `net_dial`, and the sync/update pull `read_event` loops) were `loop { read_event() }` with no deadline. When a peer's QUIC path black-holes, the broker's stream-open/send awaits the dead peer and never sends the reply, so the brain's `read_frame` blocks FOREVER and the pump freezes mid-round. The heartbeat (loop...\n 501│- **Invariant:** in PUMP mode the brain carrier is **SPLIT at construction** — a dedicated `pump-ipc-reader` thread does blocking `read_frame` on the `RecvHalf` and forwards each framed result down a channel; the main thread writes on the `SendHalf` and reads with `Receiver::recv_timeout`. Every IPC reply read is bounded by a per-call **total-wait** deadline (`PUMP_PEER_IO_TIMEOUT` = 30s, > any legitimate round-trip, < the 60s QUIC idle; re-armed on stream progress for the streaming pull legs so a heal...\n    │...\n 504│\n*505│### 7.7 A slow/dead/hostile remote VIEWER must never stall the controller, child, or drain  `[REQ-HAZARD-VIEWER-ISOLATION]`\n 506│<!-- [doc->REQ-HAZARD-VIEWER-ISOLATION] -->\n 507│- **Failure:** the W2.5 controller/viewer model lets ANY number of read-only `--view` attachers ride one session's broker `OutputLog`. The single drain thread fans each output chunk to every attacher. If a viewer's socket is fanned out with a **blocking** write under the log lock (the controller's authoritative path), one wedged viewer (a slow terminal, a black-holed WAN peer, a hostile non-reader) stalls the drain — freezing the controller's stream and backing up the PTY child. A single watcher must n...\n 508│- **Invariant:** the drain writes the **controller** on the authoritative blocking bounded path (it alone advances `delivered_through`), but each **viewer** gets an **isolated bounded SPSC queue + a dedicated writer thread**; the drain `try_send`s under the log lock and **evicts** any viewer whose queue is `Full` (fell behind the live stream) or `Disconnected` (its writer died on a dead socket) — the drain thread **never touches a viewer socket**, so no viewer write can backpressure it. A **soft cap** ...\n    │...\n 511│\n*512│### 7.8 The broker must never make a brain wait UNBOUNDED on a QUIC op (the pump-IPC-deadline B-half)  `[REQ-HAZARD-BROKER-QUIC-DEADLINE]`\n 513│<!-- [doc->REQ-HAZARD-BROKER-QUIC-DEADLINE] -->\n 514│- **Failure:** the broker's brain-facing QUIC handlers (`dispatch_net_dial` / `dispatch_net_stream_open` / `dispatch_net_stream_send`) call into `NetHost::dial` / `open_stream` / `send_stream`, whose iroh awaits (`endpoint.connect` + `prove_membership`; `open_bi`; `write_all`/`finish`) had NO bound of their own. A dead/black-holed roster peer (its process gone, or a mixed-pair that accepts the conn but never answers the seed-proof) makes the broker await its QUIC path FOREVER, so the brain escapes only v...\n 515│- **Invariant:** every brain-waiting QUIC op is wrapped in a broker-side deadline (`NetHost::bounded_block_on` → `tokio::time::timeout`, `BROKER_QUIC_OP_TIMEOUT_MS` = 10s). On elapse the future is DROPPED (cancelling the in-flight connect/stream op, so nothing is half-registered) and a non-`TimedOut` `io::Error` is returned, which the broker REPLIES as an ordinary error frame. The bound (10s) sits comfortably above any legitimate LAN/relay round-trip and 20s below the brain's 30s `PUMP_PEER_IO_TIMEOUT`...\n    │...\n 518│\n*519│### 7.9 A daemon-state wire change needs a deliberate BROKER restart (the broker is resident across a brain self-update)  `[REQ-HAZARD-BROKER-SEED-WIRE-SKEW]`\n 520│<!-- [doc->REQ-HAZARD-BROKER-SEED-WIRE-SKEW] -->\n 521│- **Failure:** the broker serves the seed-control channel and is RESIDENT across a brain-only self-update (ADR-0004's no-terminate-during-update pillar forbids auto-killing it — 6.7). A self-update that changes a daemon-state WIRE FORMAT — e.g. the v0.9.0 adapter-agnostic `Seed` (the `adapter` field dropped) — therefore lands a NEW-version CLI talking to the STILL-RESIDENT OLD broker. The old broker cannot deserialize the new `Seed` (its formerly-required `adapter` is absent), so it drops the seed-...\n 522│- **Invariant:** (a) spt-core surfaces an ACTIONABLE diagnostic on the seed-ack EOF — naming the stale-broker cause + the fix (`spt daemon stop`; the broker restarts on the next `spt api` call) — never the bare io error (scoped to `UnexpectedEof` on the seed-ack path, so it never mis-fires on an unrelated error like a refused connect). (b) A daemon-state wire change requires a DELIBERATE full broker restart; this is NOT automatic — ADR-0004 forbids auto-killing the resident broker, and a brain-only...\n    │...\n 532│\n*533│### 7.11 A dead PTY child + a dropped operator pump must NOT wedge the broker for other clients  `[REQ-HAZARD-ATTACH-WEDGE]`\n 534│- **Failure (hypothesized, v0.12.0):** a legitimately dead PTY child (crash/kill) plus an `rc` pump dropped without a clean detach (closed tab) was thought to make the broker's loopback forward `write_all` block forever on a full 64 KB duplex, park a worker in the 2-worker net runtime, saturate both, and stall every new attach / `endpoint run` (with `daemon stop` unable to join).\n 535│- **Invariant — DISPOSITION = PROVE-DON'T-CHANGE** (doyle GATE-PASS @e883f45, 2026-06-18): the post-L0 code ALREADY prevents the wedge; NO fail-fast / worker-count code was added. (1) `serve_attach` forwards fire-and-forget (`net_stream_send` `op_id=None`) and the broker-side `send_stream` is already deadline-bounded (`bounded_block_on`, `BROKER_QUIC_OP_TIMEOUT_MS` = 10s — hazard 7.8), not forever. (2) the loopback duplex is drained broker-INTERNALLY by the operator row's OWN read pump (`nethost.rs` ...\n 536│- **spt-core mapping:** int `crates/spt/tests/attach_wedge_e2e.rs` (REAL detached daemon + dummy-harness fixture): serve the victim (rc sees its tick), abruptly kill rc (dropped pump) + kill the PTY child → a NEW endpoint still comes online + is served (no wedge), the dead endpoint is offlined within a tick (`LIVENESS_RECONCILE_OFFLINE`), `daemon stop` bounded. Leans on 7.8 (the QUIC-op deadline) + 7.7 (viewer isolation).\n    │...\n 538│\n*539│### 7.12 Controller output must NOT be written inline on the drain thread — a backed-up controller wedges the session  `[REQ-HAZARD-INJECT-CONTROL-COEXIST]`\n 540│- **Failure (v0.12.x, operator dogfooding + doyle /diagnose 2026-06-19):** `OutputLog::append` fanned each live chunk to the CONTROLLER via a SYNCHRONOUS, blocking `write_frame` held INLINE on the session's single drain thread while `Mutex<OutputLog>` was locked (viewers already had a dedicated writer thread + bounded evicting channel). A backed-up controller socket — a slow operator, or the full 64 KB `rc` loopback duplex under heavy TUI redraw — parked the drain thread WITH THE LOG LOCK HELD: outpu...\n 541│- **Invariant:** controller delivery runs OFF the drain thread — a dedicated controller writer thread + bounded channel (the 7.7 viewer_writer pattern) does the blocking socket write; the drain hands each chunk off with a BOUNDED, OFF-LOCK send (deadline → detach + `clear_controller`, never park forever, never silently evict a LIVE operator's authoritative view). The controller is AUTHORITATIVE (unlike a viewer): its writer advances the `delivered_through` resume cursor (atomic, only-on-success, mono...\n 542│- **spt-core mapping:** impl `broker.rs` `OutputLog::append`/`controller_writer`/`become_controller`/`ControllerJob::deliver`/`clear_controller`; int `crates/spt-daemon/tests/inject_control_wedge.rs` (a fully-backed-up controller keeps `KIND_SESSIONS` answering = no wedge); unit the bounded-deliver + monotonic-cursor kernels. The input-side parks (write_input `write_all` on a full buffer; DSR-answer writer-mutex contention) are BENIGN on Windows ConPTY (absorbs a large inject), real only on Unix forkpty ...\n    │...\n 557│\n*558│### 7.15 An OFFLINE spt-hosted endpoint must NOT render phantom `ONLINE+CONTROLLED` — clear `driven_by` when its session is gone  `[REQ-HAZARD-DRIVEN-BY-SELFHEAL]`\n 559│- **Failure:** `driven_by` (the `info.json` `ONLINE+CONTROLLED` latch) is single-written by the broker via `clear_controller`/`stamp_driven_by`, which only fire on a controller change. When an spt-hosted endpoint's broker session is GONE (harness dead — the B2 case, 7-series sibling `REQ-HAZARD-HOSTED-LIVENESS-RECONCILE`), no controller event ever fires, so a stale `driven_by=Some(node)` persists: the picker renders a phantom \"controlled by X\" on an endpoint that is actually OFFLINE. The B2 reconcile a...\n 560│- **Invariant:** `reconcile_hosted_liveness`, when it offlines a sessionless controllable perch (the B2 keystone — no live broker session ⇒ dead harness), ALSO clears `driven_by` (`set_driven_by(perch, None)`). RACE-FREE and single-writer-safe: with NO live broker session there is no controller to re-stamp `driven_by` concurrently, so the brain may write it here without contending the broker. (The LIVE-session leg — a controller gone while its session survives — is NOT this hazard: a clean discon...\n 561│- **MEASURED 2026-07-22 (DAEMON-LIFECYCLE W2 Leg B, `transport_death_eof.rs`) — the residual SPLITS in two, and only one half is still open.** On REAL QUIC (two brokers, `BindScope::Loopback` + a real `net_dial` — NOT the in-process duplex, which has no idle timeout at all), a remote controller whose QUIC STACK DIES WITHOUT A FIN (peer frozen: endpoint alive, UDP port still bound, keepalives unanswered — the sleeping-laptop shape) DOES self-heal: the transport's own idle timeout tears the conn, the...\n    │...\n 581│<!-- [doc->REQ-HAZARD-INPUT-ACK-BACKPRESSURE] -->\n*582│### 7.19 An operator input FLOOD must not deadlock the broker via the applied-ack on the same conn  `[REQ-HAZARD-INPUT-ACK-BACKPRESSURE]`\n 583│- **Failure (operator HITL, the ctrl+V re-open):** a flood of operator input on one brain↔broker conn wedged the WHOLE broker PERMANENTLY (no new/existing attach; the controller stayed latched — the per-conn handler couldn't process the detach). `serve_attach` processes a whole `NetStreamData` batch of N `Input` records in its inner loop, calling `send_effect` N times WITHOUT returning to `read_event()`; the broker answers each with `send_frame(applied_envelope)` on the SAME conn. Brain not reading ...\n 584│- **Invariant:** the applied-ack is OPT-IN. `InputReq` carries `ack: bool` (serde `default = true`, N-1-safe). The fire-and-forward operator/rc path (`serve_attach`) sends `ack=false` via `Brain::send_effect_no_ack`; `dispatch_input` writes NO applied frame when `ack=false`, so the per-conn handler never writes back while servicing the flood → it always drains → no deadlock (cures ANY input flood). `shellchan` (one-at-a-time spool delivery, WAITS on `BrokerEvent::Applied`) keeps `send_effect` (`ack=t...\n 585│- **spt-core mapping:** impl `msg.rs` `InputReq.ack` (`default_true`) + `brain.rs` `send_effect_no_ack`/`send_effect_inner` + `attach.rs` `serve_attach` operator path → `send_effect_no_ack` + `broker.rs` `dispatch_input` gates `send_frame(applied)` on `req.ack`; unit the serde default + the ack-emitted-iff-true + the no-ack-path exactly-once dedup; int (keystone, repro-first) a flood of N>pipe-buffer input frames through `serve_attach` on one conn — PRE-FIX deadlocks, POST-FIX drains all N + the sess...\n    │...\n 595│<!-- [doc->REQ-HAZARD-CONTROLLER-WRITER-REORDER] -->\n*596│### 7.21 Exactly ONE `controller_writer` per brain↔broker connection — a superseded writer must write nothing further  `[REQ-HAZARD-CONTROLLER-WRITER-REORDER]`\n 597│- **Failure (doyle instrumented RACEDIAG repro, kitsubito):** on a brain-restart re-serve the handoff brain registered as controller on the SAME session TWICE over the SAME socket — `Brain::handoff` eagerly `subscribe(prior.next_seq=1)` → `become_controller(from_seq=1)`, spawning writer-A (writes seq 1); then `serve_attach` re-handled the replayed `Request{from_seq:0}` → `attach_as(sid,0)` → `become_controller(from_seq=0)`, spawning writer-B (writes 0,1). `become_controller` dropped the prior `Co...\n 598│- **Invariant:** on a single brain↔broker connection exactly ONE `controller_writer` is ever the LIVE writer; a SUPERSEDED writer writes no further frames after the epoch bump it observes; and **every `controller_writer` emits a strictly ASCENDING seq stream** (sorted initial batch + ascending live frames). The CORRECTNESS guarantee that falls out: a snap-above consumer over any interleaving of ascending writers — where the surviving writer (`serve_attach`'s `attach_as(sid,0)`) offers the COMPLETE ra...\n 599│- **spt-core mapping:** impl `brain.rs` `Brain::handoff` (KEEP the eager `subscribe`; seed `session_cursors`) + `Brain::subscribe_with` (resume-mode dedup-cursor reset to `from_seq`, shared by `attach`/`attach_as` — the operator-stream boundary fix) + `broker.rs` `OutputLog.controller_epoch: Arc<AtomicU64>` / `become_controller` (atomic `fetch_add`, passes the new epoch + `Arc::clone` into the writer) / `controller_writer` (epoch gate read UNDER the lock on both loops) / `mark_controller_gone` + the `C...\n    │...\n 638│\n*639│### 7.25 A perch PINNED to a DEAD session self-heals (dead-owner re-pin) instead of wedging forever; a LIVE-owner rotation still refuses  `[REQ-HAZARD-SESSION-PIN-WEDGE]`\n 640│- **Failure (F-024C/F-024D, ENLYZEAM field + clean-room repro 2026-07-02):** `authenticate()` (auth.rs) gates `api poll`/`state`/`boundary` on `proof.session_id == info.json.session_id`. If ONE boundary rotation is LOST — the departing session dies (crash / tab-close) or its `/clear`-era `boundary` call never lands — the perch stays PINNED to the dead sid, and every id-scoped hook call thereafter AUTH_REFUSES *including `boundary` itself* (the very call that would re-pin presents the NEW sid). A perm...\n 641│- **Invariant:** `authenticate()` gains a DEAD-OWNER fallback — when the caller's sid MISMATCHES the pin **AND** the perch's recorded owner pid is DEAD (`proc::is_process_alive == false`), ACCEPT the caller's sid and RE-PIN (rotate `info.json.session_id` via the locked `mutate_info` RMW + a LOUD `SESSION_REPIN` stderr line). Same trust model as `establish_perch`'s conflict gate, which already allows a rebind exactly when `owner_alive == false` (an orphaned perch accepts a new LOCAL owner). This is **AD...\n 642│- **spt-core mapping:** impl `spt/src/api/auth.rs` `authenticate` dead-owner-mismatch branch (`read_pid` → `proc::is_process_alive` → `mutate_info` rotate + `SESSION_REPIN` log); unit `auth.rs` `pinned_to_dead_sid_mismatched_poll_repins` (AUTH_REFUSED → drains+re-pins, RED pre-fix) + `live_owner_mismatch_still_refuses` (the `/clear` live-rotation squat guard, unwidened) + `token_auth_path_unchanged` (correct token wins first, no rotation).\n    │...\n 655│<!-- [doc->REQ-HAZARD-CONTROL-STAMP-LIFETIME] -->\n*656│### 7.27 A control/viewer stamp must NEVER outlive its session — every teardown path clears what attach stamped  `[REQ-HAZARD-CONTROL-STAMP-LIFETIME]`\n 657│- **Failure (F-026 #2, live evidence HFENDULEAM):** an spt-hosted endpoint stayed `ONLINE+CONTROLLED` after the operator's RC `/exit` — hours later hall-a's `info.json` still read `controlled:true` (status offline, dormant) and hfenduleam still GOSSIPED `controller_node=self` cross-node (SPT_DEV.json snapshot). ROOT: `/exit` kills the CHILD, not the controller connection, so the endpoint suspends via child termination and the teardown runs the REAP path (`broker.rs` exit-waiter → `sessions.remove(&id...\n 658│- **Invariant:** on session reap the broker clears the perch's control/viewer stamps — `OutputLog::stamp_reaped()` = `set_driven_by(None)` + `set_controlled(false)` + `set_viewer_count(0)` via the known endpoint id — so no teardown path leaves a stamp behind. The broker stays the SINGLE writer; the clear is unconditional (idempotent when already clear) and RACE-FREE — the session is dead, so there is no live controller to re-stamp concurrently (the same argument as 7.15's no-session leg). Best-effo...\n 659│- **spt-core mapping:** impl `broker.rs` `OutputLog::stamp_reaped` + the exit-waiter reap call (before `sessions.remove`); int (real broker+PTY, no mocks) `crates/spt-daemon/tests/control_stamp_lifetime.rs` `reap_clears_control_and_viewer_stamps` (spawn hosted session → the live controller stamps `controlled=true` → latch `driven_by`+`viewer_count` on disk → KILL the child → the reap clears all three; RED-first verified: fix disabled leaves `driven_by=Some`, `controlled=true`, `viewer_count=Some(...\n    │...\n 674│<!-- [doc->REQ-HAZARD-CONTROL-STAMP-CONVERGENCE] -->\n*675│### 7.29 Control/viewer stamps CONVERGE to broker session-table truth, not merely edge-trigger  `[REQ-HAZARD-CONTROL-STAMP-CONVERGENCE]`\n 676│- **Failure (F-026 stamp-gap, hall-b + the original ball-b):** a picker-created endpoint (`endpoint run` → new) read plain `ONLINE` in the list + picker while genuinely driven — `info.json` `controlled:false` throughout. ROOT (the UPWARD companion to 7.27's downward edge-clear): the broker spawn path's `become_controller` → `stamp_driven_by` → `set_controlled(true)` fires at SPAWN time, but a FRESH endpoint has NO PERCH yet (the adapter binds it after claude boots), so `mutate_info` returns `NotF...\n 677│- **Invariant:** the broker (SINGLE writer) re-asserts each live session's control/viewer stamps to session-table TRUTH, DIVERGENCE-GATED — read the perch's recorded `(driven_by, controlled, viewer_count)`, compare to the live truth (`controller_by()` / `has_controller()` / live `viewer_count`), and write ONLY on a real difference (no per-poll fsync storm). The re-assert rides the `KIND_SESSIONS` handler (the daemon reconcile + picker poll it on cadence), so a fresh perch CONVERGES within one reconcile...\n 678│- **Coverage boundary (who MAY clear vs who MUST NOT — cross-ref 7.15):** the convergence's `driven_by`-clear leg fires ONLY when `has_controller()==false` — the DEFINITIVE broker signal (it reads its OWN controller slot, not the ambiguous `controller_by`). That is why the broker MAY clear a stale `driven_by` where the BRAIN reconcile (7.15) MUST NOT: the brain only sees `controller_by==None`, which a live LOCAL controller also reads, so a brain clear would false-heal a genuinely-driven session. This...\n    │...\n 702│\n*703│### 7.33 NO bare `.lock().unwrap()` on a broker-resident lock reachable from serve/dispatch — a poison permanently wedges every attach  `[REQ-HAZARD-BROKER-FLOOR-LOCK-POISON]`\n 704│<!-- [doc->REQ-HAZARD-BROKER-FLOOR-LOCK-POISON] -->\n 705│- **Failure (paid-for class):** a brain-only self-update restarts the BRAIN but keeps the BROKER process — and every one of its `Mutex`es — ALIVE by design (REQ-UPD-3). So a single panic while another thread held a broker-resident lock POISONS it permanently: the next `.lock().unwrap()` panics too, kills its per-conn reply thread, and EVERY subsequent attach silently deadlines (`spt rc` → \"brain IPC read deadline elapsed\") while non-locked ops keep working. The effect journal (bug #16, `[REQ-HAZARD...\n 706│- **Invariant:** NO broker-resident lock reachable from serve/dispatch may use a bare `.lock().unwrap()`. Recover via the shared `recover` (`into_inner` — safe for the short, coherent-on-recovery map ops of `sessions`/`pair_holds`) or, for the `OutputLog` ring, `recover_log` — which adds a COHERENCE CLAMP on the recovery path: a panic mid-`append` can leave the ring torn (over-cap, a last seq not below `next_seq`, non-monotonic), and serving those bytes risks garbage, so `clamp_or_reset` cheap-checks...\n    │...\n 722│\n*723│### 7.36 The broker control plane and PTY fan-out must NEVER block on a single subscriber connection — a suspended brain conn must not wedge control  `[REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]`\n 724│<!-- [doc->REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE] -->\n 725│- **Failure (paid-for, rig-CONFIRMED 2026-07-06/07 — `NtSuspendProcess` on the brain, no update involved):** a controller's writer thread does a BLOCKING socket write to its brain subscriber conn. When that brain is suspended (or black-holed) the write never returns. The output-driven eviction path (`append` → `mark_controller_gone`, bounded by 7.12's `CONTROLLER_WRITE_DEADLINE`) only fires on NEW output, and `reap_dead_controller` (7.29 companion) only caught a writer that had EXITED — so a contro...\n 726│- **Invariant:** no broker client ever observes a wedge because the BRAIN stalled. (1) The controller's writer publishes its IN-FLIGHT-write window (`write_blocked_since`: `Some(Instant)` set immediately before the blocking `write_frame`, `None` after — the mutex never held ACROSS the write); a reader sees a `Some` older than `BRAIN_WRITE_DEADLINE` (15 s = 3× the 7.12 controller bound, measured full-with-**zero writer progress**, so a slow-but-DRAINING brain never trips it) as a WEDGED writer and stal...\n    │...\n 729│\n*730│### 7.37 The Layer-1 settle-gate must RE-ARM per delivery on an observable PTY — a mid-session reader reattach (`/clear`) re-creates the head-swallow race  `[REQ-HAZARD-INJECT-SETTLE-REARM]`\n 731│<!-- [doc->REQ-HAZARD-INJECT-SETTLE-REARM] -->\n 732│- **Failure (paid-for, field-CONFIRMED on 0.29.0 — doyle diagnosis + perri screenshot `WindowsTerminal_6iSjya8pMt.png`):** the shipped W5-A settle-gate (7-of-W5, `REQ-INJECT-MULTILINE-INTEGRITY`) gated Layer 1 behind a worker-local ONE-SHOT (`settled_once`) on the premise that the head-swallow race is a STARTUP condition (the input reader not yet attached after spawn). That premise is FALSE: a mid-session `/clear` re-enters the harness's raw-mode input reader, re-attaching it and RE-CREATING the pre-se...\n 733│- **Invariant:** the settle-gate re-arms before EVERY delivery on an OBSERVABLE (echoing/interactive) PTY — the bug-prone class, where re-settling is cheap (a tick or two once the reader answers) and necessary (a mid-session reader reattach must be re-confirmed). The steady-state settle is latch-skipped ONLY where the probe is known UNOBSERVABLE on this PTY (a non-echoing ConPTY whose DSR answer never surfaces): there is no reader-reattach race to guard, and each settle would burn the full `INJECT_SETT...\n    │...\n 736│\n*737│### 7.38 EVERY write on a physical broker connection is bounded + cancelable + poison-on-failure — no writer holds the send gate across an UNBOUNDED OS write  `[REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]`\n 738│<!-- [doc->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK] -->\n 739│- **Failure (paid-for, field capture + deterministic Windows repro 2026-07-09 — hertz RCA, UPDATE-WEDGE round-4):** `controller_writer` held the `SharedSend = Arc<Mutex<SendHalf>>` guard ACROSS a blocking `write_frame`. On Windows, `interprocess 2.4.2` routes the send to `WriteFileEx` + `SleepEx(INFINITE, alertable)` — this local-socket send path exposes **no supported write timeout** (`set_timeout` → `Unsupported`). An `rc --take` controller consumer that stopped reading blocked the write **indefi...\n 740│- **Invariant:** every write on a physical broker connection rides ONE broker-owned bounded/cancelable framed-write primitive with an **independent out-of-band aborter**. The `SendHalf` never leaves the conn object (a bypass cannot compile). An ABSOLUTE deadline stamped at write entry covers BOTH the serialized gate-wait AND OS write completion; a per-conn watchdog fires at that deadline and aborts the in-flight op — never relying on the write returning or on opportunistic stall-evict sampling. On dead...\n    │...\n 745│\n*746│### 7.39 Per-session identity env (`SPT_ENDPOINT_ID`/`OWL_SESSION_ID`/`SPT_AGENT_ID`) is NEVER inherited — the daemon scrubs it at startup AND on every role spawn, regardless of any role's declared `env_remove`  `[REQ-HAZARD-DAEMON-IDENTITY-ENV-SANITIZE]`\n 747│<!-- [doc->REQ-HAZARD-DAEMON-IDENTITY-ENV-SANITIZE] -->\n 748│- **Failure (paid-for, perri field RCA 2026-07-09/10 — F-036 psyche seat-theft):** a daemon restarted from inside an agent session (routine during core dev / `spt update apply`) carried the session's `SPT_ENDPOINT_ID=doyle` and passed it verbatim into every `[session.psyche_resume]` spawn — core stripped only each role's DECLARED `env_remove` list (runtime.rs `command_for`), so ONE adapter `env_remove` miss infected the whole node. Every psyche claude turn fired SessionStart, the adapter hook saw the...\n 749│- **Invariant:** the identity set (`spt_runtime::IDENTITY_ENV_VARS`) is per-session state and NEVER correct inherited state for a daemon or any child it spawns. The daemon (broker AND brain entries) scrubs its own process env first thing (`scrub_identity_env`); the shared role-spawn command builder scrubs the set unconditionally and LAST (after `env_remove`, read-env stamps, and the recursion guard), so no adapter declaration — missing or pathological — can leak identity into a role child. A spawn th...\n    │...\n 763│\n*764│### 7.41 A fresh dispatcher must NEVER re-serve a terminal stream — historical replay must not steal or clear a live controller  `[REQ-HAZARD-REDISPATCH-CONTROL-STEAL]`\n 765│<!-- [doc->REQ-HAZARD-REDISPATCH-CONTROL-STEAL] -->\n 766│- **Failure (paid-for, hertz field RCA 2026-07-16 — 4/5 endpoints frozen, twice in one day):** every brain cycle (`spt daemon refresh`, `spt update` apply — same `applyhost.rs` path) launches a fresh `dispatch::run_dispatch_loop` whose EMPTY per-process `claimed` set enumerates every broker-held peer stream. `NetShared.streams` has NO removal path (`StreamLog::finish` only marks) and the claim condition has no finished filter → the fresh dispatcher replays FINISHED historical attach streams. A repl...\n 767│- **Invariant:** redispatch eligibility is LIFECYCLE-gated: a finished/terminal stream row is retired from the dispatcher's enumeration (a fresh dispatcher never re-serves a terminal Attach); classification identity is restart-durable and independent of the evictable data ring (immutable opener fact pinned per stream until close); claims are retryable on transient worker-setup failure and terminal on terminal outcomes (no hot-loop). The legitimate same-`by` successor re-take after a brain restart still s...\n    │...\n 789│\n*790│### 7.44 Streams and seats on a long-lived connection must have bounded lifetime — one-way rows terminal at FIN, seats released at serve completion, no per-chunk full-state rewrites in a drain loop  `[REQ-HAZARD-REGISTRY-STALL]`\n 791│<!-- [doc->REQ-HAZARD-REGISTRY-STALL] -->\n 792│- **Failure (paid-for, hertz post-close v0.36 field RCA 2026-07-17 — live prod box):** FOUR compounding consequences of \"physical teardown waits for conn close\" on connections that never close. The registry pump opens one fresh stream per feed (~30s) on the persistent pump conn; nothing retires the rows (the dispatcher — sole `retire_stream` caller — skips `initiated_locally`) → sender history unbounded. `stream_infos` filters `retired` only → every dispatcher poll serializes O(history) rows ov...\n 793│- **Invariant:** a one-way (fire-and-forget) family's row is TERMINAL at successful FIN — the sender retires it; eligibility filtering happens server-side before IPC serialization; feed application is transactional per feed (snapshot writes O(feeds), never O(chunks × record-kinds) — no synchronous full-state rewrite inside a per-chunk drain iteration); a deadline-poisoned one-way replay retires terminal at a per-stream strike budget (safe: the next pump round re-advertises); every completed serve re...\n    │...\n 803│\n*804│### 7.46 An rc surface answers from live session authority, never a stale persisted projection — and a resuming perch is UNBOUND, not offline  `[REQ-HAZARD-RC-ATTACH-TRUTH]`\n 805│<!-- [doc->REQ-HAZARD-RC-ATTACH-TRUTH] -->\n 806│- **Failure (paid-for, hertz perri contradiction RCA 2026-07-17/18 + operator field recovery):** the broker hosted an honest live session (client tree alive, `SessionProbe::has_live_session_honest` = true) while the persisted perch row said `offline` — the resumed harness sat wedged at a failed native resume, SessionStart never fired, the perch never bound, inbound messages black-holed. `spt rc` trusted the offline row and refused pre-IPC while `endpoint run --resume` reattached to the very same sessio...\n 807│- **Invariant:** every rc-family reader consults the single honest-session authority (ADR-0041 `SessionProbe`) before any persisted-status fast-fail — honest session ⇒ attach; no session ⇒ refuse; dead client tree ⇒ refuse/reap, never attach. UNBOUND = broker session exists + harness not bound, fresh/resume-invariant: resume stamps an existing offline perch UNBOUND (generation-safe rollback on spawn-fail/session death), bind owns UNBOUND→ONLINE. Refusal copy states known facts (harness-hosted ...\n    │...\n 817│\n*818│### 7.48 At most one input-capable controller lease per PTY session — takeover revokes atomically and loudly, input is fenced to the active lease, node identity is never a lease  `[REQ-HAZARD-CONTROLLER-LEASE]`\n 819│<!-- [doc->REQ-HAZARD-CONTROLLER-LEASE] -->\n 820│- **Failure (paid-for, hertz same-machine `--take` RCA, field repro 2026-07-16):** terminal A controlled an endpoint; terminal B on the SAME machine ran `spt rc --take`. Local loopback attaches carry only NODE identity, so `resolve_subscribe` computed `same_identity=true` and took the silent successor path for a distinct `--take` — intent never consulted; sink replaced with no `Displaced`, no stream close (the only loud branch is different-remote+Take). `dispatch_input` is session-addressed with no con...\n 821│- **Invariant:** each rc invocation/attach mints a unique controller LEASE id carried through subscribe, the controller slot, and Input/Resize; the only silent successor is same-lease + equal-or-newer generation (dispatcher recovery, ADR-0038 fix-6 preserved); explicit Take on a DISTINCT lease always revokes ATOMICALLY and authoritatively (old lease fenced + old attach stream FORCE-closed; the Displaced notice via the old writer is best-effort — a Full-queue notice drop must never leave the incumbent i...\n    │...\n 981│\n*982│### 7.64 Session-scoped authority does not survive a session-id rotation at rest — a boundary CARRIES it as a move, never leaves a copy  `[REQ-HAZARD-EMPOWER-SESSION-RESURRECT]`\n 983│\n 984│<!-- [doc->REQ-HAZARD-EMPOWER-SESSION-RESURRECT] -->\n 985│- **Failure (LOCKSMITH t1, releases#117; MEASURED by red probe 2026-08-03, not inferred):** an engine-room empowerment is session-state — ADR-0052 says \"nothing to revoke at rest\" — and is filed under the session id the engine room's perch is bound to. `spt api boundary` rotates exactly that id, to an ARBITRARY `--to-session-id` past a guard that refuses only psyche-custody squats, and it is NEITHER of the two revocation edges ADR-0052 §3a names. The record was therefore left filed under the OLD ses..."},"isError":false,"timestamp":1788999564801}}
{"type":"message","id":"c74275ba","parentId":"d3395329","timestamp":"2026-09-10T00:19:24.802Z","message":{"role":"toolResult","toolCallId":"call_3KMS7iOQs9n4KXqYMINFM2Aq|fc_04ba494197989c86016aa1f78cac0887d1a7309b9e20a751df","toolName":"glob","content":[{"type":"text","text":"# .worktrees/286-churn/docs/adr/\n0004-single-daemon-broker-brain-split-and-self-update.md\n0018-broker-brain-process-isolation-restoration.md"}],"details":{"scopePath":".worktrees/286-churn/docs/adr/0004*, .worktrees/286-churn/docs/adr/0018*","fileCount":2,"files":[".worktrees/286-churn/docs/adr/0004-single-daemon-broker-brain-split-and-self-update.md",".worktrees/286-churn/docs/adr/0018-broker-brain-process-isolation-restoration.md"],"truncated":false,"cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core"},"isError":false,"timestamp":1788999564802}}
{"type":"message","id":"8eb253d4","parentId":"c74275ba","timestamp":"2026-09-10T00:19:24.850Z","message":{"role":"toolResult","toolCallId":"call_u4IiOAuqAiFe4hGZF5BnXWtz|fc_04ba494197989c86016aa1f78cac2087d184e51af69e863486","toolName":"grep","content":[{"type":"text","text":"# .worktrees/286-churn/crates/spt-daemon/\n\n## src/\n### brain.rs\n 1024|\n*1025|    pub fn read_event(&mut self) -> io::Result<BrokerEvent> {\n 1026|        self.read_event_until(None)\n 1027|    }\n 1028|\n...\n 1035|    // [impl->REQ-HAZARD-PUMP-IPC-DEADLINE]\n*1036|    pub fn read_event_until(&mut self, deadline: Option<Instant>) -> io::Result<BrokerEvent> {\n 1037|        loop {\n 1038|            let env = self.read_frame_until(deadline)?;\n 1039|            match env.kind.as_str() {\n...\n 2422|\n*2423|    fn send(&mut self, kind: &str, payload: serde_json::Value) -> io::Result<()> {\n 2424|        let env = Envelope::new(kind, payload);\n 2425|        match &mut self.conn {\n 2426|            BrainConn::Whole(stream) => write_frame(stream, &env),\n### broker.rs\n 5817|    // [impl->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\n*5818|    pub fn conn_handler_count(&self) -> usize {\n 5819|        self.conn_handlers.load(Ordering::Acquire)\n 5820|    }\n 5821|\n...\n 5953|        let mut my_cb_streams: Vec<u64> = Vec::new();\n*5954|        let mut my_presence_sub = false;\n 5955|\n 5956|        loop {\n 5957|            let env = match read_frame(&mut recv) {\n...\n 6272|                    match self.dispatch_net_presence_subscribe(env, &send) {\n*6273|                        Ok(()) => my_presence_sub = true,\n 6274|                        Err(msg) => send_error(&send, &msg),\n 6275|                    }\n 6276|                }\n...\n 6329|            // And presence: the liveness log + its ring persist (D4c).\n*6330|            if my_presence_sub {\n 6331|                host.detach_presence_sub(&send);\n 6332|            }\n 6333|        }\n\n## tests/\n### conn_blackhole_lifecycle.rs\n 23|//!      physically retired, not merely logically stall-evicted);\n*24|//!   3. ITS WRITER EXITS — `broker.conn_handler_count()` falls to the\n 25|//!      unrelated-viewer floor (<= 1; the persistent invariant-1 viewer holds one\n 26|//!      handler) WITHOUT the black-holed client ever dropping its end (the parked\n 27|//!      writer was aborted + the reader EOF-cleaned, so the handler thread ran to\n...\n 35|//!      still answers a control-plane `KIND_SESSIONS` request PROMPTLY (no lock\n*36|//!      is held by the retired writer), `conn_handler_count()` returns to the\n 37|//!      unrelated-viewer floor without the client dropping its end (no task\n 38|//!      lingers), AND both the\n 39|//!      unrelated session's viewer and the fresh viewer flow output (the shared\n...\n 115|\n*116|fn send(conn: &mut Stream, kind: &str, payload: serde_json::Value) {\n 117|    write_frame(conn, &Envelope::new(kind, payload)).expect(\"send frame\");\n 118|}\n 119|\n...\n 466|    // ── INVARIANT 3 (its writer exits): WITHOUT dropping R's Stream, poll\n*467|    //    conn_handler_count() until it drops to the UNRELATED-VIEWER FLOOR (<= 1),\n 468|    //    bounded at t_release + 2 s. `conn_handlers` is a GLOBAL atomic of live\n 469|    //    handle_conn threads, incremented once per physical conn. Unlike the r4\n 470|    //    gate (whose only persistent conn was the black-holed R, so its floor was\n...\n 490|    while Instant::now() < step7_deadline {\n*491|        if broker.conn_handler_count() <= 1 {\n 492|            handler_released = true;\n 493|            u_at_release = u_frames.load(Ordering::Relaxed);\n 494|            break;"}],"details":{"scopePath":".worktrees/286-churn/crates/spt-daemon/src/brain.rs, .worktrees/286-churn/crates/spt-daemon/src/broker.rs, .worktrees/286-churn/crates/spt-daemon/tests/conn_blackhole_lifecycle.rs","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-daemon","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":12,"fileCount":3,"files":[".worktrees/286-churn/crates/spt-daemon/src/brain.rs",".worktrees/286-churn/crates/spt-daemon/src/broker.rs",".worktrees/286-churn/crates/spt-daemon/tests/conn_blackhole_lifecycle.rs"],"fileMatches":[{"path":".worktrees/286-churn/crates/spt-daemon/src/brain.rs","count":3},{"path":".worktrees/286-churn/crates/spt-daemon/src/broker.rs","count":4},{"path":".worktrees/286-churn/crates/spt-daemon/tests/conn_blackhole_lifecycle.rs","count":5}],"truncated":false,"displayContent":"# .worktrees/286-churn/crates/spt-daemon/\n\n## src/\n### brain.rs\n 1024│\n*1025│    pub fn read_event(&mut self) -> io::Result<BrokerEvent> {\n 1026│        self.read_event_until(None)\n 1027│    }\n 1028│\n     │...\n 1035│    // [impl->REQ-HAZARD-PUMP-IPC-DEADLINE]\n*1036│    pub fn read_event_until(&mut self, deadline: Option<Instant>) -> io::Result<BrokerEvent> {\n 1037│        loop {\n 1038│            let env = self.read_frame_until(deadline)?;\n 1039│            match env.kind.as_str() {\n     │...\n 2422│\n*2423│    fn send(&mut self, kind: &str, payload: serde_json::Value) -> io::Result<()> {\n 2424│        let env = Envelope::new(kind, payload);\n 2425│        match &mut self.conn {\n 2426│            BrainConn::Whole(stream) => write_frame(stream, &env),\n### broker.rs\n 5817│    // [impl->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\n*5818│    pub fn conn_handler_count(&self) -> usize {\n 5819│        self.conn_handlers.load(Ordering::Acquire)\n 5820│    }\n 5821│\n     │...\n 5953│        let mut my_cb_streams: Vec<u64> = Vec::new();\n*5954│        let mut my_presence_sub = false;\n 5955│\n 5956│        loop {\n 5957│            let env = match read_frame(&mut recv) {\n     │...\n 6272│                    match self.dispatch_net_presence_subscribe(env, &send) {\n*6273│                        Ok(()) => my_presence_sub = true,\n 6274│                        Err(msg) => send_error(&send, &msg),\n 6275│                    }\n 6276│                }\n     │...\n 6329│            // And presence: the liveness log + its ring persist (D4c).\n*6330│            if my_presence_sub {\n 6331│                host.detach_presence_sub(&send);\n 6332│            }\n 6333│        }\n\n## tests/\n### conn_blackhole_lifecycle.rs\n  23│//!      physically retired, not merely logically stall-evicted);\n *24│//!   3. ITS WRITER EXITS — `broker.conn_handler_count()` falls to the\n  25│//!      unrelated-viewer floor (<= 1; the persistent invariant-1 viewer holds one\n  26│//!      handler) WITHOUT the black-holed client ever dropping its end (the parked\n  27│//!      writer was aborted + the reader EOF-cleaned, so the handler thread ran to\n    │...\n  35│//!      still answers a control-plane `KIND_SESSIONS` request PROMPTLY (no lock\n *36│//!      is held by the retired writer), `conn_handler_count()` returns to the\n  37│//!      unrelated-viewer floor without the client dropping its end (no task\n  38│//!      lingers), AND both the\n  39│//!      unrelated session's viewer and the fresh viewer flow output (the shared\n    │...\n 115│\n*116│fn send(conn: &mut Stream, kind: &str, payload: serde_json::Value) {\n 117│    write_frame(conn, &Envelope::new(kind, payload)).expect(\"send frame\");\n 118│}\n 119│\n    │...\n 466│    // ── INVARIANT 3 (its writer exits): WITHOUT dropping R's Stream, poll\n*467│    //    conn_handler_count() until it drops to the UNRELATED-VIEWER FLOOR (<= 1),\n 468│    //    bounded at t_release + 2 s. `conn_handlers` is a GLOBAL atomic of live\n 469│    //    handle_conn threads, incremented once per physical conn. Unlike the r4\n 470│    //    gate (whose only persistent conn was the black-holed R, so its floor was\n    │...\n 490│    while Instant::now() < step7_deadline {\n*491│        if broker.conn_handler_count() <= 1 {\n 492│            handler_released = true;\n 493│            u_at_release = u_frames.load(Ordering::Relaxed);\n 494│            break;"},"isError":false,"timestamp":1788999564849}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_Re3ZZoKyQLmXfXuonP61fipo|fc_04ba494197989c86016aa1f79f906c87d18f29a2507bac3fb0","toolName":"read","startedAt":"2026-09-10T00:19:43.647Z","args":{"path":".worktrees/286-churn/docs/adr/0004-single-daemon-broker-brain-split-and-self-update.md:1-90"},"intent":"Reading split architecture principles"},"id":"3b3ebfe1","parentId":"8eb253d4","timestamp":"2026-09-10T00:19:43.647Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_ABSGllMeYjx6QoXD6yL0arzG|fc_04ba494197989c86016aa1f79f908087d1b69e384c27cdb319","toolName":"read","startedAt":"2026-09-10T00:19:43.647Z","args":{"path":".worktrees/286-churn/docs/adr/0018-broker-brain-process-isolation-restoration.md:1-82"},"intent":"Reading process isolation decisions"},"id":"37631dbc","parentId":"3b3ebfe1","timestamp":"2026-09-10T00:19:43.647Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_G5Y4Noa6C1q1JJBlBCCnIEjU|fc_04ba494197989c86016aa1f79f908887d1b7309f531a7ece25","toolName":"read","startedAt":"2026-09-10T00:19:43.647Z","args":{"path":".worktrees/286-churn/docs/KNOWN-HAZARDS.md:723-744"},"intent":"Reading bounded write invariants"},"id":"9539b8a2","parentId":"37631dbc","timestamp":"2026-09-10T00:19:43.647Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_juk9D2uUpTKzSUivlQbRcFm2|fc_04ba494197989c86016aa1f79f908c87d1b6d2402630ddd951","toolName":"read","startedAt":"2026-09-10T00:19:43.647Z","args":{"path":".worktrees/286-churn/crates/spt-daemon/tests/conn_blackhole_lifecycle.rs:190-652"},"intent":"Reading harness teardown assertions"},"id":"ced767fa","parentId":"9539b8a2","timestamp":"2026-09-10T00:19:43.647Z"}
{"type":"message","id":"35a9fdc3","parentId":"ced767fa","timestamp":"2026-09-10T00:19:43.648Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":"**Checking restart log bounds**\n\n**Checking compile compatibility**"},{"type":"toolCall","id":"call_Re3ZZoKyQLmXfXuonP61fipo|fc_04ba494197989c86016aa1f79f906c87d18f29a2507bac3fb0","name":"read","arguments":{"i":"Reading split architecture principles","path":".worktrees/286-churn/docs/adr/0004-single-daemon-broker-brain-split-and-self-update.md:1-90"},"intent":"Reading split architecture principles"},{"type":"toolCall","id":"call_ABSGllMeYjx6QoXD6yL0arzG|fc_04ba494197989c86016aa1f79f908087d1b69e384c27cdb319","name":"read","arguments":{"i":"Reading process isolation decisions","path":".worktrees/286-churn/docs/adr/0018-broker-brain-process-isolation-restoration.md:1-82"},"intent":"Reading process isolation decisions"},{"type":"toolCall","id":"call_G5Y4Noa6C1q1JJBlBCCnIEjU|fc_04ba494197989c86016aa1f79f908887d1b7309f531a7ece25","name":"read","arguments":{"i":"Reading bounded write invariants","path":".worktrees/286-churn/docs/KNOWN-HAZARDS.md:723-744"},"intent":"Reading bounded write invariants"},{"type":"toolCall","id":"call_juk9D2uUpTKzSUivlQbRcFm2|fc_04ba494197989c86016aa1f79f908c87d1b6d2402630ddd951","name":"read","arguments":{"i":"Reading harness teardown assertions","path":".worktrees/286-churn/crates/spt-daemon/tests/conn_blackhole_lifecycle.rs:190-652"},"intent":"Reading harness teardown assertions"},{"type":"toolCall","id":"call_X5zYm59djDnj7Mn5roAGvf4q|fc_04ba494197989c86016aa1f79f909487d19af158c34208bd6e","name":"grep","arguments":{"i":"Locating session response schema","pattern":"struct SessionsReply|fn send_frame|fn write_frame|fn write_line","path":".worktrees/286-churn/crates/spt-daemon/src/msg.rs;.worktrees/286-churn/crates/spt-daemon/src/broker.rs;.worktrees/286-churn/crates/spt-daemon/src/codec.rs;.worktrees/286-churn/crates/spt-proto/src/emit.rs","case":true,"gitignore":false,"skip":null},"intent":"Locating session response schema"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":14787,"output":575,"cacheRead":72576,"cacheWrite":0,"totalTokens":87938,"reasoningTokens":256,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1788999564905,"responseId":"resp_04ba494197989c86016aa1f78d209487d189ef80061c934fad","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_04ba494197989c86016aa1f78e578887d1884e0a244ff90296","type":"reasoning","content":[],"encrypted_content":"gAAAAABqofefbujQYAYcUObpgCCp3FIk7nWHBaofxj9sgGQ-tvkXAWWUpZEWb8xMfMTBwzerCo5JDq5KheGHF8ggkLnM5UxgX4id4hirjcq4RlxHOBvEJf2xv4yZpZBnFxp0Rka77HKIM_DvDRXLNpmXrXdKwboCTp_C5_mjo8IcGDfXuUNka4FcPQrUbG20OReBLjDtYFI4YmkOL3lD0u6MQnoFq1WY24e8PDclPPQqrJe9VGpud1_fihvMnzuHHcxMeLDJX__VJ1oggujhrNgcB2KUC50a9PhHqUmg_t136cCZjIoB3nH13mB9wmZYIKjHvIxssZhhJpjcBOIo97e60uy03LnOWSbkJ43bVzotqZfTUDTidrGzw1j5WEbnLCrLEv2HRe6WZEY--ZccfnoG5G7eJKUJnAhA8id_u4llSyEU7jJ6FRU4FaZWhiHfvLqPewYPoV0ZVg3CHIAJZfdcuze6ovPqSIqJyDBTN6Zl7rir4BZvHiFYeHqkCnCYU70MG6O5ExKEjJXkbUJMw-xQKaXxHHKDfwKE0kdWLxp4GyDdER9u6BrEYxuLa7RScHjw5sdHWXTlGkK1JA7ycypmgUCvCBBcKo6-rPoxrhEOJ0uR135KQi3m8TtYyxB4vzLrum3HMjeW3lA2IclYjXLOCYdy__3_Ee_es63dvDveJeOrVUQ44ZBSqqNCNIXjZbslsDxXiyTWEJHd0lx9bwPnagqNBcuFeQe2g6HopN8ixLe1S7jLd6Bjs-tYgoQqGNUAvqwO-gPeEljrHU_5MDzZAC9PDBQZZTwqb-2mp7pvTXF9pSYX_LXv7fZ1xQAQCo9ZODK5OECn6jqT5uyM17HEPc444iudbRcpd7pQPbANgu0wsddctC7QH-nsijQBqtRZIqVVVyNbvzl3UARh43eggDiA4OyW2xLHSREnM691pRJDMJjRodEIQASPQFH_NK2i9uDtg5MtGMBdcirliB-EDgQm1vtbPtobPaq-bTNYJTqH14vlvyOyy86PC9FdB_PUkwKrXXq_iP6DJ-7YtyEYq60hIJpG4kgSiUDz2E6KfJMs0eAFoVcXuZ50h0ZwxCf_koLEZWFYp-gbtuK3yuMZx9r0n6mAlUWT8SGAkH323-QHZcJG9wyMpjPMBWvNlZaCA3XT5P09OCDTyXKVQ2vCRYFj4GM8fprCkq1Lu0is11edP3hxq6rkWwWBVlQTTDx19d4KCBBzAYDqFjOk7U2Pz8Ygzkd8IYSfx6wFMtnAA8c7FYtTzQHU2PEj2fLh7gGydbw73UNneIafzD_BQBWU36YmKktsau7U-TeN41lh87r3anpWtVK0-H4J3PXKdW9j3C61DNegk-4n2JL_CuRM4VXaWbrhWfI-CbzF7ZBVgMI2edzAonSyUvueVypKlcobUunh4cxWVuY9pZMic0GBYdY6HXES2YxQINDCcZBX78VjT4KJ1yd1g8EuYDlD4NVd-4umtSE1hQ7_AMu2rVhMQtJIt8jgtQiy4Aa-7mAaPW-8G3pNMSg7mvvlmo7Vo4ItWs_xdeebfeAK3SQFC4UEVcKhjZpbgUBNJ-9KOZWVTkP3Ope-PRYwQZRz81L2AEByVwtAUVZ_GuenhUlI2reqCTDHiTbWUmsTLidpRy82Rr9RcCwqO1-cFyHnJEzWJHl7XyK37ywnrW7a9ZrbgHQHP_r7ok9opin9uj7M5FIctIP_f4rKjicG7Tn5-sPpGmDhHzXMGPGGcC5xdl6mHI2Y2vpjJd_SxDH04HNC-LnEBGcQMc4j9NUzEQjYrB-vi1T3s7BIBsBn2vFzotVCTVsZP7b03XrsSS0ZFf2WJsQqqda5fPhQURRfnDOmuXPetAPNfs7EuoofW6emu0WDyr0TeB7PIghzFlQa8Tn3jFKKmBy5l5WiEPsWaQz-mLDrJgQ3ycHbWZONtbk9G3qA0-b01kUBW8nnonvseegsx_2gB9_YzCND25kyWUkC-1x4niFHmu4gKKf5EL3oHxKu9BcF_I_Ho8fs3Cftxm9Qb2ZsCEAwFLzt2kxXTt3gzJipauBTtvqltxXxpaF1a60ZmXhJujsQsqsIWB0R092NSEX1YbV5-NgUx6Sclp39FdP9PdYkZLJlvQ2y8EyAWe9E_cbI-fJZmksCoqgkzRzAtlLQpdbH63aV4z6rCo9zv_ldr9JzHi9j_ESJBgbsHcG7WKWUUBFziB_unYzqpuLWqQz6CLJQx7C3qMYVg0DX4PRcJSQ1kdcxVDS2q-dOBY5xpsrjHO0_vgk9eVjnT97dv-iwB3TwUj_7-Ten5ZnhPNFQUkkRqiGy9HRyq64wfKQxNSbVj5tlFfi0w63kKhzOth19vQUxvBDMdIpzfErP3DPgXPSoyZuzfTNAVwWTQ3P7etzFllspSOn8RwL1RWQYA2N0ZftfPYvSu5jzyjm_8dkqwHpSoAgCbP1XsPZ9ExV1IAD0VhUn957lGtJSJR3U12dwxf62abXgVVYd3OWG7VEc0BSKNLpC_aQFGkTSTqJolIKcdfb4-LmUmsLa_m_piuWYiOTPD5bNJ5TAK82Ffz_wOZgHl8fYZvPxU1zvUNaBz6QwN3n3HVgPqZC04brcfynr_eT4GkUp3kevTFHgbpc1Dv-aK_QcI_7BEazDb2eHHE3lLF8qel3K0EAbifFBrW2CeR-SXxEHaTmkLowpiO32WV5bZbDoqm98z7JN8gpiwx5_Ix0vlTf49boYjmPldxzvOlRCoXytJ8Bz72zSVfeEjcceYx6_mwTLGuT0iFpC-TTGnt6OCh2UAEM9sH6hyF-54HnUlhCnIKfzCJbBIQcWjK8PRhYA_8Ucfm4c7-3b2Ya2BTZcwLFdIWv2LuI-Ij7bpwUlEQ5cIXISAIZfHH43c_RpCuwxpfnPUUDvkHJm18qWktVCGonTKuJguqQc7A4BJN0_fGfP5xkUr4sDLd-kWM9aqVvZcbIB0UHZqIBJtpUQYCigklwSWKVBL7jvb9ubnBziWYLD6Ml5Dl9E4Ze5HjwyrflGpLGTjNwHQRSjr58a3gk_5tpDtA==","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"summary":[{"type":"summary_text","text":"**Checking restart log bounds**"},{"type":"summary_text","text":"**Checking compile compatibility**"}],"metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f79f906c87d18f29a2507bac3fb0","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading split architecture principles\",\"path\":\".worktrees/286-churn/docs/adr/0004-single-daemon-broker-brain-split-and-self-update.md:1-90\"}","call_id":"call_Re3ZZoKyQLmXfXuonP61fipo","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"read","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f79f908087d1b69e384c27cdb319","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading process isolation decisions\",\"path\":\".worktrees/286-churn/docs/adr/0018-broker-brain-process-isolation-restoration.md:1-82\"}","call_id":"call_ABSGllMeYjx6QoXD6yL0arzG","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"read","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f79f908887d1b7309f531a7ece25","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading bounded write invariants\",\"path\":\".worktrees/286-churn/docs/KNOWN-HAZARDS.md:723-744\"}","call_id":"call_G5Y4Noa6C1q1JJBlBCCnIEjU","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"read","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f79f908c87d1b6d2402630ddd951","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading harness teardown assertions\",\"path\":\".worktrees/286-churn/crates/spt-daemon/tests/conn_blackhole_lifecycle.rs:190-652\"}","call_id":"call_juk9D2uUpTKzSUivlQbRcFm2","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"read","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f79f909487d19af158c34208bd6e","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating session response schema\",\"pattern\":\"struct SessionsReply|fn send_frame|fn write_frame|fn write_line\",\"path\":\".worktrees/286-churn/crates/spt-daemon/src/msg.rs;.worktrees/286-churn/crates/spt-daemon/src/broker.rs;.worktrees/286-churn/crates/spt-daemon/src/codec.rs;.worktrees/286-churn/crates/spt-proto/src/emit.rs\",\"case\":true,\"gitignore\":false,\"skip\":null}","call_id":"call_X5zYm59djDnj7Mn5roAGvf4q","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"grep","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}}]},"duration":18739.222499999916,"ttft":1463.3941999999806,"completedAt":1788999583645,"contextSnapshot":{"promptTokens":87363,"nonMessageTokens":13648,"compactionEpoch":0}}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_X5zYm59djDnj7Mn5roAGvf4q|fc_04ba494197989c86016aa1f79f909487d19af158c34208bd6e","toolName":"grep","startedAt":"2026-09-10T00:19:43.648Z","args":{"path":".worktrees/286-churn/crates/spt-daemon/src/msg.rs;.worktrees/286-churn/crates/spt-daemon/src/broker.rs;.worktrees/286-churn/crates/spt-daemon/src/codec.rs;.worktrees/286-churn/crates/spt-proto/src/emi…"},"intent":"Locating session response schema"},"id":"ca1b98a6","parentId":"35a9fdc3","timestamp":"2026-09-10T00:19:43.648Z"}
{"type":"message","id":"96a1043d","parentId":"ca1b98a6","timestamp":"2026-09-10T00:19:43.655Z","message":{"role":"toolResult","toolCallId":"call_Re3ZZoKyQLmXfXuonP61fipo|fc_04ba494197989c86016aa1f79f906c87d18f29a2507bac3fb0","toolName":"read","content":[{"type":"text","text":"# Single consolidated daemon with broker/brain split; peer-propagated gated self-update\n\n## Status\n\naccepted (2026-05-29)\n\n## Context\n\nADR-0002 collapsed PTY-hosting and network-hosting into one per-machine `spt-daemon`. Two further forces refine its shape:\n\n1. **Consolidation.** The sister project runs poll listeners as ephemeral per-session background tasks and Psyche wrappers as detached per-live-agent supervisor processes. But poll listeners already interact directly with the agent session (capsule/idle), and Psyche wrappers already invoke harness binaries directly. Once the daemon owns every PTY, keeping these as separate processes is unjustified.\n\n2. **Seamless self-update with a hard no-terminate constraint.** Self-update is a day-one pillar. The constraint: *no endpoint process may terminate or suspend during an spt-core update* — we cannot assume every endpoint can safely suspend. The naive \"drain + restart the daemon\" approach violates this for spt-hosted sessions (the daemon owns their PTY; killing the daemon SIGHUPs the child).\n\n## Decision\n\n**Consolidate all per-machine logic into the one daemon.** Poll-listener logic and Psyche/pulse loops move into the `spt-daemon` — no separate listener or wrapper processes. The only residue is a thin, stateless **in-session relay** for harness-hosted sessions (topology 1), where spt cannot reach into a process tree it doesn't own; the relay just pipes the daemon's events into the session's stdout and is freely killable.\n\n**Split the logical daemon into two implementation layers** to satisfy the no-terminate invariant:\n- **broker** (stable kernel) — holds *only* un-transferable, must-not-die resources: PTY master fds, spawned harness child processes, listening network sockets. Minimal, versioned local IPC. Almost never updates.\n- **daemon brain** (userspace) — all logic (routing, registry, pulse/psyche loops, manifest parsing, update orchestration). Restarts freely; rehydrates from disk state and re-attaches to broker-held handles.\n\nRoutine self-updates replace only the brain → endpoints never notice. The broker updates rarely (IPC-contract change, held-resource-type change, OS PTY/socket API change, or broker bugfix); the IPC is versioned so a newer brain talks to an older broker. Logical addressing is unchanged — still one per-machine `spt-daemon`; the broker is internal, not separately addressable.\n\nRejected alternatives for the invariant: **drain + restart** (violates no-terminate for spt-hosted sessions); **whole-daemon live FD-passing** (ConPTY handle transfer mid-swap is hard and platform-divergent — pushing that complexity into the rare broker update instead of every routine update is strictly better); **topology-1-only invariant with spt-hosted as fast-follow** (silently breaks the stated hard constraint for the exact new topology spt-core introduces).\n\n**Self-update delivery:** peer-propagated over P2P, layered on self-fetch, out-of-band still supported. All binaries signature-verified before handoff (spt-core's own release key) regardless of source — peer-propagation otherwise lets one compromised node poison the subnet. spt-core conducts updates for the whole stack: self first, then ripple-update each registered adapter via the adapter manifest's update declaration (file-pull or delegated command). The plugin's role shrinks to initial bootstrap only.\n\n**Cadence/consent:** not fully automatic by default; gated on user confirmation delivered to the most-recently-active live session, with an opt-in full-auto choice.\n\n## Consequences\n\n- The daemon is the single brain for a machine; crash-recovery and update logic must cover PTYs, networking, registry, spools, listeners, and psyche loops together.\n- A small internal broker process exists beneath the daemon — a deliberate, bounded walk-back of \"literally one process,\" preserving B1's *intent* (one network identity, one supervisor, one firewall prompt) while guaranteeing endpoint survival across updates.\n- Peer-propagated updates make release signing mandatory, not optional.\n- spt-core becomes the update conductor for adapters too; adapter manifests must declare an update avenue.\n- The \"deliver to most-recently-active session\" mechanism is a v1 building block that the deferred PresenceChannel will later generalize.\n- Whole-daemon live FD-passing (zero-interruption even for broker updates) remains a possible future polish but is explicitly not required for v1.\n\n## Validation & amendments (2026-05-31 — Stage A red-team + Spike #1)\n\nCodex adversarial review (`docs/reviews/STAGE-A-codex-redteam.md`) put 4 FATAL findings on this ADR; Spike #1 (`docs/spikes/SPIKE-01-broker-handoff.md`, `spt-spikes/spike-01-broker-handoff`) tested the hardest path. Resulting binding amendments:\n\n### A. Update-class taxonomy (resolves FATAL #1 — R-UPD-3 self-contradiction)\nR-UPD-3's \"no endpoint terminates during update\" is **only** absolute for the *brain-only* class. Three classes, each with its own invariant:\n- **brain-only** (the routine case) — endpoints MUST survive untouched. **Spike #1: PROVEN on Windows ConPTY.**\n- **broker-compatible** — broker binary swap behind a versioned IPC the running broker can hot-accept; endpoints survive; no spike yet.\n- **broker-breaking** — held-resource-type or OS-API change; requires a *planned endpoint-cycle*. R-UPD-3 is explicitly weakened here: this class MAY suspend endpoints, with consent + scheduling. PRD R-UPD-3 must be reworded to scope the absolute guarantee to brain-only.\n\n### B. Ownership table (resolves FATAL #4 — \"any live stream owned by brain breaks the invariant\")\nThe boundary is normative, not \"broker holds fds, brain holds logic.\" A resource is broker-owned **iff** a live consumer would lose continuity on brain restart.\n\n| Resource | Owner | Spiked? |\n|---|---|---|\n| PTY master read+write handle | **broker** | ✅ #1 |\n| Spawned harness child process | **broker** | ✅ #1 |\n| Accepted local client socket (`api listen`, relay) | **broker** | ✅ #1 (plain TCP) |\n| Iroh endpoint + accepted QUIC streams | **broker** | ✅ #3 (loopback shape) → **implemented M4-D4a** (`spt-daemon::nethost::NetHost`: dedicated tokio runtime in the broker owns the endpoint + conn table; brain drives status/dial over IPC net frames) |\n| mDNS socket / relay session | **broker** | ✅ closed by construction at M4-D4a (2026-06-03): `MdnsAddressLookup` and the relay session are constructed *inside* the broker-owned iroh endpoint — there is no separate socket to own |\n| Listening sockets | **broker** | ✅ #1 |\n| Routing, registry state, pulse/psyche loops, manifest parse, update orchestration | brain (rehydrate from disk) | n/a |\n\nConsequence codex forced into the open: pushing live Iroh/QUIC ownership into the broker means networking stream-state lives in the \"stable kernel,\" enlarging it. Accepted as the cost of the invariant; the broker is \"stable\" in *update cadence*, not in *narrowness*.\n\n### C. ConPTY DSR gotcha (new hazard, discovered by Spike #1)\nConPTY withholds **all** child stdout until the terminal answers its startup cursor-position query (`ESC [ 6 n`). A broker reading a ConPTY master that does not reply `ESC [ <r>;<c> R` sees only the 4-byte query and then nothing — looks like a hung/silent child. Spike got 0 bytes of program output until the reader answered DSR, then full output. → `REQ-HAZARD-CONPTY-DSR`; every ConPTY reader must auto-answer DSR. Added to KNOWN-HAZARDS §5.\n\n### D. Self-update delivery hardening (resolves SERIOUS #5)\nSignature-verify alone is insufficient. v1 adds: monotonic version field (rollback rejection), release-metadata expiry, channel pinning, key rotation/revocation list, and **adapter content signing** (adapter file-pull / delegated-command updates were uncovered by \"binary signature\"). Track as `REQ-HAZARD-UPDATE-ROLLBACK`.\n\n### E. Open spike gaps (must close before M3 builds the daemon)\n1. ✅ **CLOSED (Spike #3, 2026-06-01).** Live Iroh + file-transfer stream survival across brain restart (FATAL #2). Spike #1 only proved PTY + plain TCP; Spike #3 stood up a broker-owned Iroh endpoint + uni QUIC transfer and proved a live peer download survives a brain restart gapless + exactly-once. Loopback shape validated; off-node transport + the QUIC-ownership *implementation* defer to M4. (`docs/spikes/SPIKE-03-quic-survival.md`.)\n2. ✅ **CLOSED (Spike #4, 2026-06-01).** Linux `forkpty` parity. Spike #1's binary, unchanged, passed all four invariants on `gravity-linux` (Ubuntu 22.04) — `portable-pty` selects `forkpty` on Unix; the ConPTY-DSR branch is inert there. Invariant B passed with the *strict* contiguity check that ConPTY-under-resize fails (Spike #5): `forkpty` is a raw pipe, ConPTY a screen buffer — different stream contracts, promoted into M3a's `SessionSurface` design. (`docs/spikes/SPIKE-04-forkpty-parity.md`.)\n3. ✅ **CLOSED (Spike #5, 2026-06-01).** 100× restart + resize-under-load stress (codex #3). The broker survived 100 rapid brain kill/restart cycles + a resize thread racing the reader with no leak (one child, clean reap), no hang (watchdog deadline), no lost byte (full value-set coverage in the client's observation window). Surfaced two binding ConPTY findings for M3a/REQ-TERM-3: resize triggers a repaint that reorders+dups the stream (terminal-stream contract ≠ exactly-once transfer contract); mid-session attach yields the viewport only, not scrollback. (`docs/spikes/SPIKE-05-restart-stress.md`.)\n4. ✅ **CLOSED (Spike #6, 2026-06-01).** Idempotent/exactly-once delivery across brain restart (FATAL-adjacent #14). Durable-ID + WAL + dedup-at-effect proven exactly-once across a crash at any protocol point (before-intent / before-effect / after-effect). Two binding design constraints surfaced for `REQ-HAZARD-RESTART-IDEMPOTENT` / M3b-B5: broker-owned recovery anchor; dedup-at-effect keyed by durable ID. (`docs/spikes/SPIKE-06-idempotent-boundary.md`.)\n\n**Gate status (2026-06-01): ALL FOUR §E GAPS CLOSED PASS** (#3, #4, #5, #6) → the M3-PLAN Phase-0 spike-gate is **complete**. The broker/brain split is validated on both OSes (ConPTY + `forkpty`), across QUIC-stream survival, restart+resize churn, and exactly-once idempotency. M3a and M3b are both unblocked; the QUIC-ownership *implementation* and off-node transport remain explicitly deferred to M4 (only the *shape* was spiked)."}],"details":{"totalLines":77,"displayContent":{"text":"# Single consolidated daemon with broker/brain split; peer-propagated gated self-update\r\n\r\n## Status\r\n\r\naccepted (2026-05-29)\r\n\r\n## Context\r\n\r\nADR-0002 collapsed PTY-hosting and network-hosting into one per-machine `spt-daemon`. Two further forces refine its shape:\r\n\r\n1. **Consolidation.** The sister project runs poll listeners as ephemeral per-session background tasks and Psyche wrappers as detached per-live-agent supervisor processes. But poll listeners already interact directly with the agent session (capsule/idle), and Psyche wrappers already invoke harness binaries directly. Once the daemon owns every PTY, keeping these as separate processes is unjustified.\r\n\r\n2. **Seamless self-update with a hard no-terminate constraint.** Self-update is a day-one pillar. The constraint: *no endpoint process may terminate or suspend during an spt-core update* — we cannot assume every endpoint can safely suspend. The naive \"drain + restart the daemon\" approach violates this for spt-hosted sessions (the daemon owns their PTY; killing the daemon SIGHUPs the child).\r\n\r\n## Decision\r\n\r\n**Consolidate all per-machine logic into the one daemon.** Poll-listener logic and Psyche/pulse loops move into the `spt-daemon` — no separate listener or wrapper processes. The only residue is a thin, stateless **in-session relay** for harness-hosted sessions (topology 1), where spt cannot reach into a process tree it doesn't own; the relay just pipes the daemon's events into the session's stdout and is freely killable.\r\n\r\n**Split the logical daemon into two implementation layers** to satisfy the no-terminate invariant:\r\n- **broker** (stable kernel) — holds *only* un-transferable, must-not-die resources: PTY master fds, spawned harness child processes, listening network sockets. Minimal, versioned local IPC. Almost never updates.\r\n- **daemon brain** (userspace) — all logic (routing, registry, pulse/psyche loops, manifest parsing, update orchestration). Restarts freely; rehydrates from disk state and re-attaches to broker-held handles.\r\n\r\nRoutine self-updates replace only the brain → endpoints never notice. The broker updates rarely (IPC-contract change, held-resource-type change, OS PTY/socket API change, or broker bugfix); the IPC is versioned so a newer brain talks to an older broker. Logical addressing is unchanged — still one per-machine `spt-daemon`; the broker is internal, not separately addressable.\r\n\r\nRejected alternatives for the invariant: **drain + restart** (violates no-terminate for spt-hosted sessions); **whole-daemon live FD-passing** (ConPTY handle transfer mid-swap is hard and platform-divergent — pushing that complexity into the rare broker update instead of every routine update is strictly better); **topology-1-only invariant with spt-hosted as fast-follow** (silently breaks the stated hard constraint for the exact new topology spt-core introduces).\r\n\r\n**Self-update delivery:** peer-propagated over P2P, layered on self-fetch, out-of-band still supported. All binaries signature-verified before handoff (spt-core's own release key) regardless of source — peer-propagation otherwise lets one compromised node poison the subnet. spt-core conducts updates for the whole stack: self first, then ripple-update each registered adapter via the adapter manifest's update declaration (file-pull or delegated command). The plugin's role shrinks to initial bootstrap only.\r\n\r\n**Cadence/consent:** not fully automatic by default; gated on user confirmation delivered to the most-recently-active live session, with an opt-in full-auto choice.\r\n\r\n## Consequences\r\n\r\n- The daemon is the single brain for a machine; crash-recovery and update logic must cover PTYs, networking, registry, spools, listeners, and psyche loops together.\r\n- A small internal broker process exists beneath the daemon — a deliberate, bounded walk-back of \"literally one process,\" preserving B1's *intent* (one network identity, one supervisor, one firewall prompt) while guaranteeing endpoint survival across updates.\r\n- Peer-propagated updates make release signing mandatory, not optional.\r\n- spt-core becomes the update conductor for adapters too; adapter manifests must declare an update avenue.\r\n- The \"deliver to most-recently-active session\" mechanism is a v1 building block that the deferred PresenceChannel will later generalize.\r\n- Whole-daemon live FD-passing (zero-interruption even for broker updates) remains a possible future polish but is explicitly not required for v1.\r\n\r\n## Validation & amendments (2026-05-31 — Stage A red-team + Spike #1)\r\n\r\nCodex adversarial review (`docs/reviews/STAGE-A-codex-redteam.md`) put 4 FATAL findings on this ADR; Spike #1 (`docs/spikes/SPIKE-01-broker-handoff.md`, `spt-spikes/spike-01-broker-handoff`) tested the hardest path. Resulting binding amendments:\r\n\r\n### A. Update-class taxonomy (resolves FATAL #1 — R-UPD-3 self-contradiction)\r\nR-UPD-3's \"no endpoint terminates during update\" is **only** absolute for the *brain-only* class. Three classes, each with its own invariant:\r\n- **brain-only** (the routine case) — endpoints MUST survive untouched. **Spike #1: PROVEN on Windows ConPTY.**\r\n- **broker-compatible** — broker binary swap behind a versioned IPC the running broker can hot-accept; endpoints survive; no spike yet.\r\n- **broker-breaking** — held-resource-type or OS-API change; requires a *planned endpoint-cycle*. R-UPD-3 is explicitly weakened here: this class MAY suspend endpoints, with consent + scheduling. PRD R-UPD-3 must be reworded to scope the absolute guarantee to brain-only.\r\n\r\n### B. Ownership table (resolves FATAL #4 — \"any live stream owned by brain breaks the invariant\")\r\nThe boundary is normative, not \"broker holds fds, brain holds logic.\" A resource is broker-owned **iff** a live consumer would lose continuity on brain restart.\r\n\r\n| Resource | Owner | Spiked? |\r\n|---|---|---|\r\n| PTY master read+write handle | **broker** | ✅ #1 |\r\n| Spawned harness child process | **broker** | ✅ #1 |\r\n| Accepted local client socket (`api listen`, relay) | **broker** | ✅ #1 (plain TCP) |\r\n| Iroh endpoint + accepted QUIC streams | **broker** | ✅ #3 (loopback shape) → **implemented M4-D4a** (`spt-daemon::nethost::NetHost`: dedicated tokio runtime in the broker owns the endpoint + conn table; brain drives status/dial over IPC net frames) |\r\n| mDNS socket / relay session | **broker** | ✅ closed by construction at M4-D4a (2026-06-03): `MdnsAddressLookup` and the relay session are constructed *inside* the broker-owned iroh endpoint — there is no separate socket to own |\r\n| Listening sockets | **broker** | ✅ #1 |\r\n| Routing, registry state, pulse/psyche loops, manifest parse, update orchestration | brain (rehydrate from disk) | n/a |\r\n\r\nConsequence codex forced into the open: pushing live Iroh/QUIC ownership into the broker means networking stream-state lives in the \"stable kernel,\" enlarging it. Accepted as the cost of the invariant; the broker is \"stable\" in *update cadence*, not in *narrowness*.\r\n\r\n### C. ConPTY DSR gotcha (new hazard, discovered by Spike #1)\r\nConPTY withholds **all** child stdout until the terminal answers its startup cursor-position query (`ESC [ 6 n`). A broker reading a ConPTY master that does not reply `ESC [ <r>;<c> R` sees only the 4-byte query and then nothing — looks like a hung/silent child. Spike got 0 bytes of program output until the reader answered DSR, then full output. → `REQ-HAZARD-CONPTY-DSR`; every ConPTY reader must auto-answer DSR. Added to KNOWN-HAZARDS §5.\r\n\r\n### D. Self-update delivery hardening (resolves SERIOUS #5)\r\nSignature-verify alone is insufficient. v1 adds: monotonic version field (rollback rejection), release-metadata expiry, channel pinning, key rotation/revocation list, and **adapter content signing** (adapter file-pull / delegated-command updates were uncovered by \"binary signature\"). Track as `REQ-HAZARD-UPDATE-ROLLBACK`.\r\n\r\n### E. Open spike gaps (must close before M3 builds the daemon)\r\n1. ✅ **CLOSED (Spike #3, 2026-06-01).** Live Iroh + file-transfer stream survival across brain restart (FATAL #2). Spike #1 only proved PTY + plain TCP; Spike #3 stood up a broker-owned Iroh endpoint + uni QUIC transfer and proved a live peer download survives a brain restart gapless + exactly-once. Loopback shape validated; off-node transport + the QUIC-ownership *implementation* defer to M4. (`docs/spikes/SPIKE-03-quic-survival.md`.)\r\n2. ✅ **CLOSED (Spike #4, 2026-06-01).** Linux `forkpty` parity. Spike #1's binary, unchanged, passed all four invariants on `gravity-linux` (Ubuntu 22.04) — `portable-pty` selects `forkpty` on Unix; the ConPTY-DSR branch is inert there. Invariant B passed with the *strict* contiguity check that ConPTY-under-resize fails (Spike #5): `forkpty` is a raw pipe, ConPTY a screen buffer — different stream contracts, promoted into M3a's `SessionSurface` design. (`docs/spikes/SPIKE-04-forkpty-parity.md`.)\r\n3. ✅ **CLOSED (Spike #5, 2026-06-01).** 100× restart + resize-under-load stress (codex #3). The broker survived 100 rapid brain kill/restart cycles + a resize thread racing the reader with no leak (one child, clean reap), no hang (watchdog deadline), no lost byte (full value-set coverage in the client's observation window). Surfaced two binding ConPTY findings for M3a/REQ-TERM-3: resize triggers a repaint that reorders+dups the stream (terminal-stream contract ≠ exactly-once transfer contract); mid-session attach yields the viewport only, not scrollback. (`docs/spikes/SPIKE-05-restart-stress.md`.)\r\n4. ✅ **CLOSED (Spike #6, 2026-06-01).** Idempotent/exactly-once delivery across brain restart (FATAL-adjacent #14). Durable-ID + WAL + dedup-at-effect proven exactly-once across a crash at any protocol point (before-intent / before-effect / after-effect). Two binding design constraints surfaced for `REQ-HAZARD-RESTART-IDEMPOTENT` / M3b-B5: broker-owned recovery anchor; dedup-at-effect keyed by durable ID. (`docs/spikes/SPIKE-06-idempotent-boundary.md`.)\r\n\r\n**Gate status (2026-06-01): ALL FOUR §E GAPS CLOSED PASS** (#3, #4, #5, #6) → the M3-PLAN Phase-0 spike-gate is **complete**. The broker/brain split is validated on both OSes (ConPTY + `forkpty`), across QUIC-stream survival, restart+resize churn, and exactly-once idempotency. M3a and M3b are both unblocked; the QUIC-ownership *implementation* and off-node transport remain explicitly deferred to M4 (only the *shape* was spiked).\r","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77]},"fileSize":10463,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\docs\\adr\\0004-single-daemon-broker-brain-split-and-self-update.md"}}},"isError":false,"timestamp":1788999583654}}
{"type":"message","id":"a7a9e3eb","parentId":"96a1043d","timestamp":"2026-09-10T00:19:43.656Z","message":{"role":"toolResult","toolCallId":"call_ABSGllMeYjx6QoXD6yL0arzG|fc_04ba494197989c86016aa1f79f908087d1b69e384c27cdb319","toolName":"read","content":[{"type":"text","text":"# Restore broker/brain process isolation (correct the in-process-collapse regression)\n\n## Status\n\nAccepted (2026-06-09) — **extends and amends ADR-0004.** The design was independently verified (agent `doyle`, verified-with-amendments); implementation is a dedicated future milestone, sequenced next (before `spt-claude-code`). Full rationale, the `file:line` audit, and the per-decision alternatives live in `docs/BROKER-BRAIN-SPLIT-RESTORATION.md`.\n\n## Context\n\nADR-0004 decided the broker/brain split **as two processes** — a stable broker kernel (PTY masters, harness children, sockets, the Iroh/QUIC endpoint) beneath a freely-restartable brain — specifically to satisfy the no-endpoint-drop self-update invariant (REQ-UPD-3): a routine (brain-only) update restarts the brain while the broker process survives, so no hosted endpoint terminates. The spikes (01/03/04/05/06) proved this with two separate binaries.\n\nOn 2026-06-09, while verifying the v0.3.2 cross-OS update fix on the live fleet, a regression was discovered: **the production daemon runs the broker as a background thread inside the single `spt daemon` process** (`daemon.rs:165-170`: `Broker::bind_in_with_net` → `Arc<Broker>` served on `thread::spawn`; no broker child-spawn exists anywhere in the repo). This is **unintended drift** — no doc, plan, commit, or comment marks the collapse as a deliberate interim (the operator confirmed no such decision was made).\n\nConsequences of the drift:\n- The brain cannot restart onto a new binary without killing the in-process broker thread, which would close every PTY, orphan every harness child, and drop every listening socket. **The no-endpoint-drop self-update pillar is therefore silently unrealized.**\n- `spt update apply` performs an in-process `Brain::handoff` that re-attaches a subscriber within the *same old process* — a no-op for a binary swap. `update.rs:233-234`'s \"the live daemon execs the verified new binary\" is aspirational and was never wired.\n- New code does not run until an unrelated restart/logon. Observed live: `enlyzeam` ran 0.3.0 with the valid 0.3.2 binary on disk for ~a day, continuing to reproduce the `\\r`-corruption bug the update was meant to fix.\n- **REQ-DAEMON-2 and REQ-UPD-3 carry `int` evidence that proves only the in-process handoff shape** (`tests/update.rs`, `brain_swap.rs`, the M3b-B9 daemon E2E) — i.e. the regression is masked in the requirement registry: the tests pass while proving the wrong thing.\n\n## Decision\n\n**Restore the ADR-0004 two-process model.** The brain runs new code by restarting onto the swapped binary while the broker process — holding every continuity-bearing resource — survives. The design decisions (grill Q2–Q8, hardened by six verification amendments `[V1]`–`[V6]`):\n\n1. **Supervision (Q2).** The broker is the always-up per-machine anchor (one per `SPT_HOME`, present even with zero endpoints). It owns the seed-control lock + liveness (`ensure_running`/`is_running` target) and supervises/spawns the brain as its child.\n2. **Update trigger (Q3).** `apply` swaps the binary on disk, then signals the brain to snapshot + self-exit; the broker auto-respawns from the executable path (now the new binary); the new brain re-attaches. Reuses the existing snapshot→drop→re-attach primitive; an update is a *planned* crash on the path the broker already recovers from. **`[AMENDED 2026-06-11, v0.4.2]`** \"the executable path\" silently assumed path-string semantics. The Linux apply renames the running `spt` → `spt.old-N`; a per-spawn `std::env::current_exe()` (`readlink(/proc/self/exe)`, inode-tracking) then **follows the rename to `.old-N`**, so the broker respawned the brain onto the OLD bytes while recording `applied` (caught live on kitsubito, v0.4.1 roll). Fix: capture the canonical ex…\n3. **Loop-timing continuity (Q4).** Durable absolute-deadline state on disk, not a handoff snapshot. Periodic loops persist `(anchor, interval)` once on a fresh/crash start and derive `next_fire` functionally (no per-fire writes); an update restart keeps deriving (phase preserved), a crash restart resets the anchor. One-shot deadlines persist at creation and survive crash + update. **`[V3]`** The one-shot *rule* is fixed here; its *machinery* is built with the alarm port (no untested dead code). **`[V4]`** Only phase-significant loops convert — the idempotent pump cadences need none.\n4. **IPC boundary (Q5).** ADR-0004 §B adjudicates: net bring-up → broker (already broker-owned via `broker.rs:175` `OnceLock` — near-free), digest hub → broker / parse → brain, seed-lock → broker; every brain→broker call becomes a versioned IPC verb (no shared `Arc`). **Exception:** shellwake watcher children stay brain-side, re-reconciled from disk on start (rare updates make the window tolerable). Broker enlargement accepted (ADR-0004:63 precedent). **`[V6]`** Steady state after any update is new-brain × old-broker, so a CI-real old-broker × new-brain compat test across the verb surface is required (else KH-2.3 returns).\n5. **Multi-session handoff (Q6).** The broker becomes cursor-of-record per session; the new brain re-attaches **all** sessions in resume mode. Output is at-least-once (matches the SPIKE-05 terminal-stream contract); input/effects stay exactly-once via the broker-owned `EffectJournal`. The explicit `BrainState` handoff *message* retires (no brain→brain channel exists under the self-exit trigger).\n6. **Failure atomicity (Q7).** Bounded-retry → auto-rollback to the last-known-good binary, gated on a brain `ready` signal (reuse the `peerloop.rs:805` supervise-backoff + healthy-run). Quarantine the failed version; loud notif. The applied record becomes **two-phase** (`applied-pending` → `applied` on ready, or `rolled-back` on failure) — fixing today's optimistic `record_applied` before boot (`applyhost.rs:176`). **`[V1]`** Forward invariant minted: a brain must not irreversibly migrate durable state before ready-promotion (pre-ready writes stay N-1-readable), or auto-rollback silently breaks on the first schema migration.\n7. **Generation custody (Q2/`[V2]`).** Retiring `BrainState` would orphan the KH-2.4 generation counter; the **broker** owns it, increments on every spawn, and hands `{generation, start-reason}` to the brain at spawn via a versioned argv/hello field (KH-2.3 compat). The same channel carries Q4's update-vs-crash discriminator.\n8. **Cross-platform uniformity (Q8).** The broker spawns the brain as a child process (`Command::spawn`) + socket IPC — no `exec`, so no Windows/Unix divergence. ConPTY/forkpty handles never leave the broker.\n\n## Consequences\n\n- **Decomposition:** `spt daemon run` becomes the broker process entry (binds seed-control + broker socket + NetHost + digest hub; holds children + `EffectJournal`; spawns the brain); a new hidden `spt daemon brain` is the brain entry (connects, runs the logic loops, rehydrates from disk, emits `ready`). `ensure_running`/`is_running`/`daemon stop` contracts unchanged. The de-elevation guard applies at the broker entry.\n- **Requirement registry:** REQ-DAEMON-2 and REQ-UPD-3's `int` evidence must be **re-pointed** to a productionized SPIKE-01/03 E2E proving *process-level* endpoint survival (`[V5]`). Two new hazards are minted: `REQ-HAZARD-BROKER-PROCESS-ISOLATION` (a brain restart must never drop a hosted endpoint) and `REQ-HAZARD-ROLLBACK-STATE-COMPAT`. A KNOWN-HAZARDS §6.7 entry records the regression.\n- **Sequencing:** the restoration is the next milestone, before `spt-claude-code` scoping (operator-accepted 2026-06-09). Rationale: it is the *last* release that needs a manual fleet daemon bounce — paid 3× for v0.3.2 — so every adapter-era release rolls seamlessly; and the adapter is better built on the final topology than atop a daemon under later surgery while it hosts the user's daily driver. The split changes daemon internals, not the M8-frozen CLI/api surface.\n- **Out of scope (unchanged):** the broker-touching update classes (broker-compatible / broker-breaking) remain as ADR-0004 left them; whole-daemon FD-passing stays the deferred \"future polish\" (ADR-0004:38); a durable in-daemon alarm scheduler is a separate gap (alarms are legacy-listener-only today).\n\n[Some lines truncated to 768 chars]"}],"details":{"totalLines":37,"displayContent":{"text":"# Restore broker/brain process isolation (correct the in-process-collapse regression)\r\n\r\n## Status\r\n\r\nAccepted (2026-06-09) — **extends and amends ADR-0004.** The design was independently verified (agent `doyle`, verified-with-amendments); implementation is a dedicated future milestone, sequenced next (before `spt-claude-code`). Full rationale, the `file:line` audit, and the per-decision alternatives live in `docs/BROKER-BRAIN-SPLIT-RESTORATION.md`.\r\n\r\n## Context\r\n\r\nADR-0004 decided the broker/brain split **as two processes** — a stable broker kernel (PTY masters, harness children, sockets, the Iroh/QUIC endpoint) beneath a freely-restartable brain — specifically to satisfy the no-endpoint-drop self-update invariant (REQ-UPD-3): a routine (brain-only) update restarts the brain while the broker process survives, so no hosted endpoint terminates. The spikes (01/03/04/05/06) proved this with two separate binaries.\r\n\r\nOn 2026-06-09, while verifying the v0.3.2 cross-OS update fix on the live fleet, a regression was discovered: **the production daemon runs the broker as a background thread inside the single `spt daemon` process** (`daemon.rs:165-170`: `Broker::bind_in_with_net` → `Arc<Broker>` served on `thread::spawn`; no broker child-spawn exists anywhere in the repo). This is **unintended drift** — no doc, plan, commit, or comment marks the collapse as a deliberate interim (the operator confirmed no such decision was made).\r\n\r\nConsequences of the drift:\r\n- The brain cannot restart onto a new binary without killing the in-process broker thread, which would close every PTY, orphan every harness child, and drop every listening socket. **The no-endpoint-drop self-update pillar is therefore silently unrealized.**\r\n- `spt update apply` performs an in-process `Brain::handoff` that re-attaches a subscriber within the *same old process* — a no-op for a binary swap. `update.rs:233-234`'s \"the live daemon execs the verified new binary\" is aspirational and was never wired.\r\n- New code does not run until an unrelated restart/logon. Observed live: `enlyzeam` ran 0.3.0 with the valid 0.3.2 binary on disk for ~a day, continuing to reproduce the `\\r`-corruption bug the update was meant to fix.\r\n- **REQ-DAEMON-2 and REQ-UPD-3 carry `int` evidence that proves only the in-process handoff shape** (`tests/update.rs`, `brain_swap.rs`, the M3b-B9 daemon E2E) — i.e. the regression is masked in the requirement registry: the tests pass while proving the wrong thing.\r\n\r\n## Decision\r\n\r\n**Restore the ADR-0004 two-process model.** The brain runs new code by restarting onto the swapped binary while the broker process — holding every continuity-bearing resource — survives. The design decisions (grill Q2–Q8, hardened by six verification amendments `[V1]`–`[V6]`):\r\n\r\n1. **Supervision (Q2).** The broker is the always-up per-machine anchor (one per `SPT_HOME`, present even with zero endpoints). It owns the seed-control lock + liveness (`ensure_running`/`is_running` target) and supervises/spawns the brain as its child.\r\n2. **Update trigger (Q3).** `apply` swaps the binary on disk, then signals the brain to snapshot + self-exit; the broker auto-respawns from the executable path (now the new binary); the new brain re-attaches. Reuses the existing snapshot→drop→re-attach primitive; an update is a *planned* crash on the path the broker already recovers from. **`[AMENDED 2026-06-11, v0.4.2]`** \"the executable path\" silently assumed path-string semantics. The Linux apply renames the running `spt` → `spt.old-N`; a per-spawn `std::env::current_exe()` (`readlink(/proc/self/exe)`, inode-tracking) then **follows the rename to `.old-N`**, so the broker respawned the brain onto the OLD bytes while recording `applied` (caught live on kitsubito, v0.4.1 roll). Fix: capture the canonical ex…\n3. **Loop-timing continuity (Q4).** Durable absolute-deadline state on disk, not a handoff snapshot. Periodic loops persist `(anchor, interval)` once on a fresh/crash start and derive `next_fire` functionally (no per-fire writes); an update restart keeps deriving (phase preserved), a crash restart resets the anchor. One-shot deadlines persist at creation and survive crash + update. **`[V3]`** The one-shot *rule* is fixed here; its *machinery* is built with the alarm port (no untested dead code). **`[V4]`** Only phase-significant loops convert — the idempotent pump cadences need none.\r\n4. **IPC boundary (Q5).** ADR-0004 §B adjudicates: net bring-up → broker (already broker-owned via `broker.rs:175` `OnceLock` — near-free), digest hub → broker / parse → brain, seed-lock → broker; every brain→broker call becomes a versioned IPC verb (no shared `Arc`). **Exception:** shellwake watcher children stay brain-side, re-reconciled from disk on start (rare updates make the window tolerable). Broker enlargement accepted (ADR-0004:63 precedent). **`[V6]`** Steady state after any update is new-brain × old-broker, so a CI-real old-broker × new-brain compat test across the verb surface is required (else KH-2.3 returns).\r\n5. **Multi-session handoff (Q6).** The broker becomes cursor-of-record per session; the new brain re-attaches **all** sessions in resume mode. Output is at-least-once (matches the SPIKE-05 terminal-stream contract); input/effects stay exactly-once via the broker-owned `EffectJournal`. The explicit `BrainState` handoff *message* retires (no brain→brain channel exists under the self-exit trigger).\r\n6. **Failure atomicity (Q7).** Bounded-retry → auto-rollback to the last-known-good binary, gated on a brain `ready` signal (reuse the `peerloop.rs:805` supervise-backoff + healthy-run). Quarantine the failed version; loud notif. The applied record becomes **two-phase** (`applied-pending` → `applied` on ready, or `rolled-back` on failure) — fixing today's optimistic `record_applied` before boot (`applyhost.rs:176`). **`[V1]`** Forward invariant minted: a brain must not irreversibly migrate durable state before ready-promotion (pre-ready writes stay N-1-readable), or auto-rollback silently breaks on the first schema migration.\r\n7. **Generation custody (Q2/`[V2]`).** Retiring `BrainState` would orphan the KH-2.4 generation counter; the **broker** owns it, increments on every spawn, and hands `{generation, start-reason}` to the brain at spawn via a versioned argv/hello field (KH-2.3 compat). The same channel carries Q4's update-vs-crash discriminator.\r\n8. **Cross-platform uniformity (Q8).** The broker spawns the brain as a child process (`Command::spawn`) + socket IPC — no `exec`, so no Windows/Unix divergence. ConPTY/forkpty handles never leave the broker.\r\n\r\n## Consequences\r\n\r\n- **Decomposition:** `spt daemon run` becomes the broker process entry (binds seed-control + broker socket + NetHost + digest hub; holds children + `EffectJournal`; spawns the brain); a new hidden `spt daemon brain` is the brain entry (connects, runs the logic loops, rehydrates from disk, emits `ready`). `ensure_running`/`is_running`/`daemon stop` contracts unchanged. The de-elevation guard applies at the broker entry.\r\n- **Requirement registry:** REQ-DAEMON-2 and REQ-UPD-3's `int` evidence must be **re-pointed** to a productionized SPIKE-01/03 E2E proving *process-level* endpoint survival (`[V5]`). Two new hazards are minted: `REQ-HAZARD-BROKER-PROCESS-ISOLATION` (a brain restart must never drop a hosted endpoint) and `REQ-HAZARD-ROLLBACK-STATE-COMPAT`. A KNOWN-HAZARDS §6.7 entry records the regression.\r\n- **Sequencing:** the restoration is the next milestone, before `spt-claude-code` scoping (operator-accepted 2026-06-09). Rationale: it is the *last* release that needs a manual fleet daemon bounce — paid 3× for v0.3.2 — so every adapter-era release rolls seamlessly; and the adapter is better built on the final topology than atop a daemon under later surgery while it hosts the user's daily driver. The split changes daemon internals, not the M8-frozen CLI/api surface.\r\n- **Out of scope (unchanged):** the broker-touching update classes (broker-compatible / broker-breaking) remain as ADR-0004 left them; whole-daemon FD-passing stays the deferred \"future polish\" (ADR-0004:38); a durable in-daemon alarm scheduler is a separate gap (alarms are legacy-listener-only today).\r","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37]},"fileSize":9995,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\docs\\adr\\0018-broker-brain-process-isolation-restoration.md"},"limits":{"columnTruncated":{"maxColumn":768}}}},"isError":false,"timestamp":1788999583656}}
{"type":"message","id":"f1bf6e11","parentId":"a7a9e3eb","timestamp":"2026-09-10T00:19:43.660Z","message":{"role":"toolResult","toolCallId":"call_juk9D2uUpTKzSUivlQbRcFm2|fc_04ba494197989c86016aa1f79f908c87d1b6d2402630ddd951","toolName":"read","content":[{"type":"text","text":"fn spawn_session(name: &str, req: SpawnReq) -> (Stream, u64) {\n…\n    let sid = loop {\n…\n        if f.kind == KIND_SPAWNED {\n…\n        }\n    };\n    (c, sid)\n}\n\n/// Query the broker's stall-evict tally. Bounded thread + channel (a wedge\n/// surfaces as a timeout, not a hang). Verbatim sibling of the r4 gate helper.\nfn stall_evict_count(name: &str, deadline: Duration) -> u64 {\n    let name = name.to_string();\n    let (tx, rx) = std::sync::mpsc::channel();\n    thread::spawn(move || {\n        let mut c = connect(&name);\n        send(&mut c, KIND_STALL_EVICTS, serde_json::Value::Null);\n        loop {\n            match read_frame(&mut c) {\n                Ok(f) if f.kind == KIND_STALL_EVICTS_REPLY => {\n                    let r: StallEvictsReply = serde_json::from_value(f.payload).unwrap();\n                    let _ = tx.send(r.count);\n                    return;\n                }\n                Ok(_) => continue,\n                Err(_) => {\n                    let _ = tx.send(0);\n                    return;\n                }\n            }\n        }\n    });\n    rx.recv_timeout(deadline).unwrap_or(0)\n}\n\n/// One `KIND_SESSIONS` round-trip, returning the reply. The dispatch runs the\n/// `reap_dead_controller` → time-based `stall_evict_controller`, so polling this\n/// both DRIVES the logical eviction and reads the frozen `resume_seq` cursor.\n/// Bounded thread + channel. Verbatim sibling of the r4 gate helper.\nfn query_sessions(name: &str, deadline: Duration) -> Option<SessionsReply> {\n    let name = name.to_string();\n    let (tx, rx) = std::sync::mpsc::channel();\n    thread::spawn(move || {\n        let mut c = connect(&name);\n        send(&mut c, KIND_SESSIONS, serde_json::Value::Null);\n        loop {\n            match read_frame(&mut c) {\n                Ok(f) if f.kind == KIND_SESSIONS_REPLY => {\n                    let r: SessionsReply = serde_json::from_value(f.payload).unwrap();\n                    let _ = tx.send(r);\n                    return;\n                }\n                Ok(_) => continue,\n                Err(_) => return,\n            }\n        }\n    });\n    rx.recv_timeout(deadline).ok()\n}\n\n// [int->REQ-CONN-BLACKHOLE-LIFECYCLE-HARNESS]\n#[test]\nfn blackholed_controller_lifecycle_five_invariants() {\n    let _home = init_home();\n    // Shrink the brain-write deadline so the wedged writer is \"stalled\" in ~400 ms\n    // instead of 15 s — the structural outcome, fast + deterministic (the knob\n    // grants wall-clock only; it never relaxes an assertion).\n    std::env::set_var(\"SPT_BRAIN_WRITE_DEADLINE_MS\", \"400\");\n\n    let name = unique_name();\n    let broker = Broker::bind(&name).expect(\"bind broker\");\n    let serve = Arc::clone(&broker);\n    thread::spawn(move || {\n        let _ = serve.serve();\n    });\n\n    // ── UNRELATED session U (invariant 1): a second, independent flood child on\n    //    its OWN endpoint. Drop its spawner conn (so U has no controller to wedge\n    //    — the PTY keeps producing headless), and attach a SEPARATE DRAINING\n    //    viewer whose frame count must keep advancing ACROSS the whole incident. ──\n    let (u_spawn, sid_u) = spawn_session(&name, flood_spawn_req(\"blackhole-unrelated-ep\"));\n    drop(u_spawn);\n    let u_name = name.clone();\n    let u_frames = Arc::new(AtomicU64::new(0));\n    let u_frames_w = Arc::clone(&u_frames);\n    thread::spawn(move || {\n        let mut v = connect(&u_name);\n        send(\n            &mut v,\n            KIND_SUBSCRIBE,\n            serde_json::to_value(SubscribeReq {\n                session_id: sid_u,\n                from_seq: 0,\n                intent: AttachIntent::Viewer,\n                by: Some(\"unrelated-viewer\".to_string()),\n                gen: 0, code: None, seal_ceremony: false, })\n            .unwrap(),\n        );\n        while let Ok(f) = read_frame(&mut v) {\n            if f.kind == KIND_OUTPUT {\n                u_frames_w.fetch_add(1, Ordering::Relaxed);\n            }\n        }\n    });\n\n    // ── BLACK-HOLED session B: the real SEED+burst PTY child from conn A, then\n    //    DROP A so the ONLY persistent brain conn is the black-holed controller R. ──\n    let (a, sid_b) = spawn_session(&name, seed_burst_spawn_req(\"blackhole-victim-ep\"));\n    drop(a);\n\n    // ── Remote controller R: subscribes Control from_seq 0, drains until it has\n    //    seen `SEED` (a real controller frame flowed), signals main so main can\n    //    freeze the PRE-burst high-water, then on main's go fires the burst trigger,\n    //    drains a few burst frames (a real seq>0 frame delivered), and STOPS\n    //    reading — handing the Stream back to main UNREAD. Main holds it, never\n    //    draining, so the broker's R-writer parks in `write_frame` exactly as\n    //    against a suspended brain. R's Stream is NEVER dropped until the very end. ──\n    let r_name = name.clone();\n    let (seed_tx, seed_rx) = std::sync::mpsc::channel::<(bool, u64)>();\n    let (go_tx, go_rx) = std::sync::mpsc::channel::<()>();\n    let (hand_tx, hand_rx) = std::sync::mpsc::channel::<(Stream, u64)>();\n    thread::spawn(move || {\n        let mut r = connect(&r_name);\n        send(\n            &mut r,\n            KIND_SUBSCRIBE,\n            serde_json::to_value(SubscribeReq {\n                session_id: sid_b,\n                from_seq: 0,\n                intent: AttachIntent::Control,\n                by: Some(\"operator-one\".to_string()),\n                gen: 0, code: None, seal_ceremony: false, })\n            .unwrap(),\n        );\n        loop {\n            match read_frame(&mut r) {\n                Ok(f) if f.kind == KIND_SUBSCRIBED => break,\n                Ok(_) => continue,\n                Err(_) => {\n                    let _ = seed_tx.send((false, 0));\n                    return;\n                }\n            }\n        }\n        // Drain live output until the CONCATENATED decoded bytes contain `SEED`\n        // (ConPTY echoes input + emits VT noise, so per-frame matching is\n        // unreliable). Track the max seq seen.\n        let mut acc: Vec<u8> = Vec::new();\n        let mut max_seq: u64 = 0;\n        let mut saw_seed = false;\n        let seed_deadline = Instant::now() + Duration::from_secs(30);\n        while Instant::now() < seed_deadline {\n            match read_frame(&mut r) {\n                Ok(f) if f.kind == KIND_OUTPUT => {\n                    let ev: OutputEvent = serde_json::from_value(f.payload).unwrap();\n                    max_seq = max_seq.max(ev.seq);\n                    if let Ok(bytes) = decode_bytes(&ev.data_b64) {\n                        acc.extend_from_slice(&bytes);\n                    }\n                    if acc.windows(4).any(|w| w == b\"SEED\") {\n                        saw_seed = true;\n                        break;\n                    }\n                }\n                Ok(_) => continue,\n                Err(_) => break,\n            }\n        }\n        let _ = seed_tx.send((saw_seed, max_seq));\n        // Wait for main's go, then fire the burst trigger (`\\r\\n` satisfies the\n        // first Read-Host; ack=false = the rc/operator fire-and-forward path).\n        let _ = go_rx.recv_timeout(Duration::from_secs(15));\n        send(\n            &mut r,\n            KIND_INPUT,\n            serde_json::to_value(InputReq {\n                session_id: sid_b,\n                data_b64: encode_bytes(b\"\\r\\n\"),\n                op_id: None,\n                minter: None,\n                ack: false,\n            })\n            .unwrap(),\n        );\n        // Drain a bounded handful of burst frames so a REAL seq>0 controller frame\n        // is proven delivered, then STOP reading forever (the black-hole). The\n        // remaining ~1300 frames flood the now-unread socket → R-writer wedges.\n        let post_deadline = Instant::now() + Duration::from_secs(5);\n        let mut post = 0u32;\n        while (max_seq == 0 || post < 8) && Instant::now() < post_deadline {\n            match read_frame(&mut r) {\n                Ok(f) if f.kind == KIND_OUTPUT => {\n                    let ev: OutputEvent = serde_json::from_value(f.payload).unwrap();\n                    max_seq = max_seq.max(ev.seq);\n                    post += 1;\n                }\n                Ok(_) => continue,\n                Err(_) => break,\n            }\n        }\n        // Hand the Stream back to main UNREAD (main holds it = the black-holed\n        // client). The broker's R-writer stays wedged on the remaining burst.\n        let _ = hand_tx.send((r, max_seq));\n    });\n\n    // Main: recv SEED-reached, freeze the PRE-burst high-water (trigger NOT fired —\n    // R is parked on go_rx), release R.\n    let (saw_seed, _seed_max) = seed_rx\n        .recv_timeout(Duration::from_secs(35))\n        .unwrap_or((false, 0));\n    let seed_high = broker.session_output_seq(sid_b).unwrap_or(0);\n    let _ = go_tx.send(());\n    // Recv the handed-back black-holed Stream (held UNREAD for the rest of the test).\n    let handed = hand_rx.recv_timeout(Duration::from_secs(25)).ok();\n    let mut r_stream: Option<Stream> = None;\n    let mut max_drained_seq: u64 = 0;\n    if let Some((stream, mds)) = handed {\n        r_stream = Some(stream);\n        max_drained_seq = mds;\n    }\n\n    // Prove the burst is real: the victim's ring high-water advances ≥64 frames\n    // past the trigger (independent of the wedged controller).\n    let mut burst_advanced = false;\n    let step_deadline = Instant::now() + Duration::from_secs(20);\n    while Instant::now() < step_deadline {\n        if let Some(hw) = broker.session_output_seq(sid_b) {\n            if hw >= seed_high + 64 {\n                burst_advanced = true;\n                break;\n            }\n        }\n        thread::sleep(Duration::from_millis(100));\n    }\n\n    // INVARIANT 1 baseline: the unrelated viewer must ALREADY be ticking before\n    // the incident window opens (so the \"kept advancing across the incident\"\n    // assertion below is meaningful, not vacuous).\n    let mut u_baseline = 0u64;\n    let u_base_deadline = Instant::now() + Duration::from_secs(15);\n    while Instant::now() < u_base_deadline {\n        u_baseline = u_frames.load(Ordering::Relaxed);\n        if u_baseline > 0 {\n            break;\n        }\n        thread::sleep(Duration::from_millis(50));\n    }\n\n    // ── Observe exactly ONE logical stall eviction via an INDEPENDENT DRAINING\n    //    control-plane connection (the KIND_SESSIONS reap runs the 400 ms\n    //    stall-evict). Record t_release + freeze the resume cursor for the sid_b\n    //    row. Accept a tally of {1, 2}: POISON-BEATS-EVICT (F-039 rider,\n    //    doyle-ruled) — the physical watchdog can poison first and record the same\n    //    release class, so the tally can step 0→2 when watchdog races the reap. ──\n    let mut stall_evicts_seen: u64 = 0;\n    let mut t_release: Option<Instant> = None;\n    let mut frozen_cursor: Option<u64> = None;\n    let step6_deadline = Instant::now() + Duration::from_secs(10);\n    while Instant::now() < step6_deadline {\n        let sessions = query_sessions(&name, Duration::from_secs(3));\n        let count = stall_evict_count(&name, Duration::from_secs(3));\n        if count >= 1 {\n            stall_evicts_seen = count;\n            t_release = Some(Instant::now());\n            frozen_cursor = sessions\n                .as_ref()\n                .and_then(|r| r.sessions.iter().find(|s| s.session_id == sid_b))\n                .map(|s| s.resume_seq)\n                .or_else(|| {\n                    query_sessions(&name, Duration::from_secs(3)).and_then(|r| {\n                        r.sessions\n                            .iter()\n                            .find(|s| s.session_id == sid_b)\n                            .map(|s| s.resume_seq)\n                    })\n                });\n            break;\n        }\n        thread::sleep(Duration::from_millis(150));\n    }\n\n    // ── INVARIANT 3 (its writer exits): WITHOUT dropping R's Stream, poll\n    //    conn_handler_count() until it drops to the UNRELATED-VIEWER FLOOR (<= 1),\n    //    bounded at t_release + 2 s. `conn_handlers` is a GLOBAL atomic of live\n    //    handle_conn threads, incremented once per physical conn. Unlike the r4\n    //    gate (whose only persistent conn was the black-holed R, so its floor was\n    //    0), THIS harness keeps a persistent UNRELATED viewer conn attached for\n    //    invariant 1 — so while BOTH R's handler and the unrelated viewer's\n    //    handler are alive the count is atomically >= 2, and it CANNOT reach 1\n    //    until R's handler exits. A sample of <= 1 therefore proves R's\n    //    handle_conn thread FINISHED (the parked writer was aborted + the reader\n    //    EOF-cleaned) WITHOUT the client dropping its end — airtight because\n    //    invariant 1 (below) independently proves the unrelated viewer stayed\n    //    alive across this window, so the surviving handler is the viewer's, not\n    //    R's. (The fresh viewer has not yet attached at this point.) ──\n    let step7_deadline = t_release.unwrap_or_else(Instant::now) + Duration::from_secs(2);\n    let mut handler_released = false;\n    // R1 (doyle gate rework): snapshot the unrelated viewer's frame count AT THE\n    // INSTANT the <=1 sample latches. `u_final > u_baseline` alone cannot prove\n    // WHICH handler survived — the viewer could accrue frames past baseline and\n    // then die BEFORE this sample, leaving R's LEAKED handler as the count==1 and\n    // every invariant false-green. Frames observed AFTER this snapshot\n    // (`u_final > u_at_release`, asserted with invariant 3 below) prove the\n    // surviving handler is the VIEWER's — forcing R's to be the one that exited.\n    let mut u_at_release: u64 = 0;\n    while Instant::now() < step7_deadline {\n        if broker.conn_handler_count() <= 1 {\n            handler_released = true;\n            u_at_release = u_frames.load(Ordering::Relaxed);\n            break;\n        }\n        thread::sleep(Duration::from_millis(50));\n    }\n\n    // ── INVARIANT 2 (the bad physical connection is canceled/closed within the\n    //    bound): a write on the OLD client Stream must FAIL — the physical conn\n    //    was retired (poisoned + severed), not merely logically stall-evicted. ──\n    let old_client_write: Option<std::io::Result<()>> = r_stream\n        .as_mut()\n        .map(|s| write_frame(s, &Envelope::new(KIND_SESSIONS, serde_json::Value::Null)));\n    let old_client_write_is_err = old_client_write.as_ref().map(|r| r.is_err());\n\n    // ── INVARIANT 5 proxy (no lock/task remains owned by the retired conn): after\n    //    retirement the broker must still answer a control-plane request PROMPTLY.\n    //    A lock still held by the retired writer would stall this; a well under the\n    //    watchdog reply proves the control plane is free. (Combined with\n    //    handler_released above = no lingering task, and the two viewers flowing\n    //    below = the fan-out lock is free.) ──\n    let cp_start = Instant::now();\n    let cp_reply = query_sessions(&name, Duration::from_secs(5));\n    let control_plane_elapsed = cp_start.elapsed();\n    let control_plane_ok = cp_reply\n        .as_ref()\n        .map(|r| r.sessions.iter().any(|s| s.session_id == sid_b))\n        .unwrap_or(false);\n\n    // ── INVARIANT 4 (a fresh VIEWER can attach): a brand-new Viewer subscribe on\n    //    the retired session from the frozen cursor must COMPLETE and RECEIVE\n    //    replayed output (the burst frames buffered in the ring after the cursor).\n    //    Bounded thread + channel. ──\n    let resume_seq = frozen_cursor.unwrap_or(0);\n    let fv_name = name.clone();\n    let (fv_tx, fv_rx) =\n        std::sync::mpsc::channel::<(Option<SubscribeOutcome>, bool, Option<u64>)>();\n    thread::spawn(move || {\n        let mut c = connect(&fv_name);\n        send(\n            &mut c,\n            KIND_SUBSCRIBE,\n            serde_json::to_value(SubscribeReq {\n                session_id: sid_b,\n                from_seq: resume_seq,\n                intent: AttachIntent::Viewer,\n                by: Some(\"fresh-viewer\".to_string()),\n                gen: 0, code: None, seal_ceremony: false, })\n            .unwrap(),\n        );\n        // The writer thread resolve_subscribe spawns races the inline SUBSCRIBED\n        // reply for the same send lock, so a replay frame can hit the wire BEFORE\n        // SUBSCRIBED. Capture BOTH the outcome and the first replayed seq; break\n        // once both are known (or the deadline elapses).\n        let deadline = Instant::now() + Duration::from_secs(15);\n        let mut outcome: Option<SubscribeOutcome> = None;\n        let mut got_output = false;\n        let mut first_seq: Option<u64> = None;\n        while Instant::now() < deadline && (outcome.is_none() || !got_output) {\n            match read_frame(&mut c) {\n                Ok(f) if f.kind == KIND_SUBSCRIBED => {\n                    let reply: SubscribedReply = serde_json::from_value(f.payload).unwrap();\n                    outcome = Some(reply.outcome);\n                }\n                Ok(f) if f.kind == KIND_OUTPUT => {\n                    if !got_output {\n                        let ev: OutputEvent = serde_json::from_value(f.payload).unwrap();\n                        first_seq = Some(ev.seq);\n                        got_output = true;\n                    }\n                }\n                Ok(_) => continue,\n                Err(_) => break,\n            }\n        }\n        let _ = fv_tx.send((outcome, got_output, first_seq));\n    });\n    let (fv_outcome, fv_got_output, fv_first_seq) = fv_rx\n        .recv_timeout(Duration::from_secs(18))\n        .unwrap_or((None, false, None));\n\n    // ── INVARIANT 1 final sample: the unrelated viewer's frame count must have\n    //    ADVANCED across the whole incident window (baseline → now). It kept\n    //    receiving output while B's controller conn was black-holed + retired. ──\n    let u_final = u_frames.load(Ordering::Relaxed);\n\n    // ── Teardown BEFORE asserting (so a failing assert leaves no orphan child):\n    //    kill both PTY children, stop the broker, drop the black-holed Stream. ──\n    if let Some(pid) = broker.session_pid(sid_b) {\n        kill_pid(pid);\n    }\n    if let Some(pid) = broker.session_pid(sid_u) {\n        kill_pid(pid);\n    }\n    broker.stop();\n    drop(r_stream);\n\n    eprintln!(\n        \"=== BLACKHOLE-LC HARNESS: saw_seed={saw_seed} seed_high={seed_high} \\\n         max_drained_seq={max_drained_seq} burst_advanced={burst_advanced} \\\n         u_baseline={u_baseline} u_at_release={u_at_release} u_final={u_final} \\\n         stall_evicts_seen={stall_evicts_seen} \\\n         handler_released={handler_released} old_client_write_is_err={old_client_write_is_err:?} \\\n         control_plane_ok={control_plane_ok} control_plane_elapsed={control_plane_elapsed:?} \\\n         resume_seq={resume_seq} fv_outcome={fv_outcome:?} fv_got_output={fv_got_output} \\\n         fv_first_seq={fv_first_seq:?} ===\"\n    );\n\n    // ── HARNESS preconditions (test-environment, distinct from the product\n    //    invariants): the black-hole must have actually been established, else the\n    //    five-invariant observation is vacuous. ──\n    assert!(\n        old_client_write_is_err.is_some(),\n        \"HARNESS: the black-holed controller never handed its stream back — the \\\n         PowerShell child likely never produced SEED (saw_seed={saw_seed}). This is a \\\n         test-environment failure, not a product regression.\"\n    );\n    assert!(\n        burst_advanced,\n        \"HARNESS: the victim PTY/ring high-water must advance ≥64 frames past the \\\n         burst trigger (seed_high={seed_high}) — the finite burst must actually flow.\"\n    );\n    assert!(\n        max_drained_seq > 0,\n        \"HARNESS: the controller must have drained a real seq>0 frame before \\\n         black-holing (load-gated repro); got max_drained_seq={max_drained_seq}\"\n    );\n    assert!(\n        (1..=2).contains(&stall_evicts_seen),\n        \"HARNESS: the one wedged controller's release must be tallied once (logical \\\n         reap OR physical poison) or twice (both raced — poison-beats-evict, \\\n         field-accepted v0.30.6); got {stall_evicts_seen}\"\n    );\n\n    // ── INVARIANT 1: UNRELATED SESSIONS CONTINUE. ──\n    assert!(\n        u_baseline > 0,\n        \"INVARIANT 1: the unrelated session's viewer must have been receiving output \\\n         BEFORE the incident (else the 'kept ticking' proof is vacuous); got \\\n         u_baseline={u_baseline}\"\n    );\n    assert!(\n        u_final > u_baseline,\n        \"INVARIANT 1: the UNRELATED session's viewer must keep receiving output \\\n         ACROSS the black-hole incident — the fan-out must never block the whole \\\n         broker on the one wedged conn; got u_baseline={u_baseline} u_final={u_final}\"\n    );\n\n    // ── INVARIANT 2: THE BAD PHYSICAL CONNECTION IS CANCELED/CLOSED WITHIN THE\n    //    BOUND — a write on the OLD client conn FAILS (it was physically retired,\n    //    not merely logically stall-evicted). ──\n    assert_eq!(\n        old_client_write_is_err,\n        Some(true),\n        \"INVARIANT 2: a write on the OLD black-holed client connection must FAIL — \\\n         the physical conn must be canceled/closed within the write bound of logical \\\n         release. A success means the conn is still physically alive (the r4 RED).\"\n    );\n\n    // ── INVARIANT 3: ITS WRITER EXITS — the black-holed conn's handle_conn thread\n    //    finishes (conn_handler_count → the <=1 unrelated-viewer floor) within 2 s\n    //    of logical release, WITHOUT the client dropping its end. ──\n    assert!(\n        handler_released,\n…\n    );\n…\n}\n\n[Showing lines 189-655 of 707. Use :656 to continue]"}],"details":{"truncation":{"content":"        }\r\n    };\r\n    (c, sid)\r\n}\r\n\r\n/// Query the broker's stall-evict tally. Bounded thread + channel (a wedge\r\n/// surfaces as a timeout, not a hang). Verbatim sibling of the r4 gate helper.\r\nfn stall_evict_count(name: &str, deadline: Duration) -> u64 {\r\n    let name = name.to_string();\r\n    let (tx, rx) = std::sync::mpsc::channel();\r\n    thread::spawn(move || {\r\n        let mut c = connect(&name);\r\n        send(&mut c, KIND_STALL_EVICTS, serde_json::Value::Null);\r\n        loop {\r\n            match read_frame(&mut c) {\r\n                Ok(f) if f.kind == KIND_STALL_EVICTS_REPLY => {\r\n                    let r: StallEvictsReply = serde_json::from_value(f.payload).unwrap();\r\n                    let _ = tx.send(r.count);\r\n                    return;\r\n                }\r\n                Ok(_) => continue,\r\n                Err(_) => {\r\n                    let _ = tx.send(0);\r\n                    return;\r\n                }\r\n            }\r\n        }\r\n    });\r\n    rx.recv_timeout(deadline).unwrap_or(0)\r\n}\r\n\r\n/// One `KIND_SESSIONS` round-trip, returning the reply. The dispatch runs the\r\n/// `reap_dead_controller` → time-based `stall_evict_controller`, so polling this\r\n/// both DRIVES the logical eviction and reads the frozen `resume_seq` cursor.\r\n/// Bounded thread + channel. Verbatim sibling of the r4 gate helper.\r\nfn query_sessions(name: &str, deadline: Duration) -> Option<SessionsReply> {\r\n    let name = name.to_string();\r\n    let (tx, rx) = std::sync::mpsc::channel();\r\n    thread::spawn(move || {\r\n        let mut c = connect(&name);\r\n        send(&mut c, KIND_SESSIONS, serde_json::Value::Null);\r\n        loop {\r\n            match read_frame(&mut c) {\r\n                Ok(f) if f.kind == KIND_SESSIONS_REPLY => {\r\n                    let r: SessionsReply = serde_json::from_value(f.payload).unwrap();\r\n                    let _ = tx.send(r);\r\n                    return;\r\n                }\r\n                Ok(_) => continue,\r\n                Err(_) => return,\r\n            }\r\n        }\r\n    });\r\n    rx.recv_timeout(deadline).ok()\r\n}\r\n\r\n// [int->REQ-CONN-BLACKHOLE-LIFECYCLE-HARNESS]\r\n#[test]\r\nfn blackholed_controller_lifecycle_five_invariants() {\r\n    let _home = init_home();\r\n    // Shrink the brain-write deadline so the wedged writer is \"stalled\" in ~400 ms\r\n    // instead of 15 s — the structural outcome, fast + deterministic (the knob\r\n    // grants wall-clock only; it never relaxes an assertion).\r\n    std::env::set_var(\"SPT_BRAIN_WRITE_DEADLINE_MS\", \"400\");\r\n\r\n    let name = unique_name();\r\n    let broker = Broker::bind(&name).expect(\"bind broker\");\r\n    let serve = Arc::clone(&broker);\r\n    thread::spawn(move || {\r\n        let _ = serve.serve();\r\n    });\r\n\r\n    // ── UNRELATED session U (invariant 1): a second, independent flood child on\r\n    //    its OWN endpoint. Drop its spawner conn (so U has no controller to wedge\r\n    //    — the PTY keeps producing headless), and attach a SEPARATE DRAINING\r\n    //    viewer whose frame count must keep advancing ACROSS the whole incident. ──\r\n    let (u_spawn, sid_u) = spawn_session(&name, flood_spawn_req(\"blackhole-unrelated-ep\"));\r\n    drop(u_spawn);\r\n    let u_name = name.clone();\r\n    let u_frames = Arc::new(AtomicU64::new(0));\r\n    let u_frames_w = Arc::clone(&u_frames);\r\n    thread::spawn(move || {\r\n        let mut v = connect(&u_name);\r\n        send(\r\n            &mut v,\r\n            KIND_SUBSCRIBE,\r\n            serde_json::to_value(SubscribeReq {\r\n                session_id: sid_u,\r\n                from_seq: 0,\r\n                intent: AttachIntent::Viewer,\r\n                by: Some(\"unrelated-viewer\".to_string()),\r\n                gen: 0, code: None, seal_ceremony: false, })\r\n            .unwrap(),\r\n        );\r\n        while let Ok(f) = read_frame(&mut v) {\r\n            if f.kind == KIND_OUTPUT {\r\n                u_frames_w.fetch_add(1, Ordering::Relaxed);\r\n            }\r\n        }\r\n    });\r\n\r\n    // ── BLACK-HOLED session B: the real SEED+burst PTY child from conn A, then\r\n    //    DROP A so the ONLY persistent brain conn is the black-holed controller R. ──\r\n    let (a, sid_b) = spawn_session(&name, seed_burst_spawn_req(\"blackhole-victim-ep\"));\r\n    drop(a);\r\n\r\n    // ── Remote controller R: subscribes Control from_seq 0, drains until it has\r\n    //    seen `SEED` (a real controller frame flowed), signals main so main can\r\n    //    freeze the PRE-burst high-water, then on main's go fires the burst trigger,\r\n    //    drains a few burst frames (a real seq>0 frame delivered), and STOPS\r\n    //    reading — handing the Stream back to main UNREAD. Main holds it, never\r\n    //    draining, so the broker's R-writer parks in `write_frame` exactly as\r\n    //    against a suspended brain. R's Stream is NEVER dropped until the very end. ──\r\n    let r_name = name.clone();\r\n    let (seed_tx, seed_rx) = std::sync::mpsc::channel::<(bool, u64)>();\r\n    let (go_tx, go_rx) = std::sync::mpsc::channel::<()>();\r\n    let (hand_tx, hand_rx) = std::sync::mpsc::channel::<(Stream, u64)>();\r\n    thread::spawn(move || {\r\n        let mut r = connect(&r_name);\r\n        send(\r\n            &mut r,\r\n            KIND_SUBSCRIBE,\r\n            serde_json::to_value(SubscribeReq {\r\n                session_id: sid_b,\r\n                from_seq: 0,\r\n                intent: AttachIntent::Control,\r\n                by: Some(\"operator-one\".to_string()),\r\n                gen: 0, code: None, seal_ceremony: false, })\r\n            .unwrap(),\r\n        );\r\n        loop {\r\n            match read_frame(&mut r) {\r\n                Ok(f) if f.kind == KIND_SUBSCRIBED => break,\r\n                Ok(_) => continue,\r\n                Err(_) => {\r\n                    let _ = seed_tx.send((false, 0));\r\n                    return;\r\n                }\r\n            }\r\n        }\r\n        // Drain live output until the CONCATENATED decoded bytes contain `SEED`\r\n        // (ConPTY echoes input + emits VT noise, so per-frame matching is\r\n        // unreliable). Track the max seq seen.\r\n        let mut acc: Vec<u8> = Vec::new();\r\n        let mut max_seq: u64 = 0;\r\n        let mut saw_seed = false;\r\n        let seed_deadline = Instant::now() + Duration::from_secs(30);\r\n        while Instant::now() < seed_deadline {\r\n            match read_frame(&mut r) {\r\n                Ok(f) if f.kind == KIND_OUTPUT => {\r\n                    let ev: OutputEvent = serde_json::from_value(f.payload).unwrap();\r\n                    max_seq = max_seq.max(ev.seq);\r\n                    if let Ok(bytes) = decode_bytes(&ev.data_b64) {\r\n                        acc.extend_from_slice(&bytes);\r\n                    }\r\n                    if acc.windows(4).any(|w| w == b\"SEED\") {\r\n                        saw_seed = true;\r\n                        break;\r\n                    }\r\n                }\r\n                Ok(_) => continue,\r\n                Err(_) => break,\r\n            }\r\n        }\r\n        let _ = seed_tx.send((saw_seed, max_seq));\r\n        // Wait for main's go, then fire the burst trigger (`\\r\\n` satisfies the\r\n        // first Read-Host; ack=false = the rc/operator fire-and-forward path).\r\n        let _ = go_rx.recv_timeout(Duration::from_secs(15));\r\n        send(\r\n            &mut r,\r\n            KIND_INPUT,\r\n            serde_json::to_value(InputReq {\r\n                session_id: sid_b,\r\n                data_b64: encode_bytes(b\"\\r\\n\"),\r\n                op_id: None,\r\n                minter: None,\r\n                ack: false,\r\n            })\r\n            .unwrap(),\r\n        );\r\n        // Drain a bounded handful of burst frames so a REAL seq>0 controller frame\r\n        // is proven delivered, then STOP reading forever (the black-hole). The\r\n        // remaining ~1300 frames flood the now-unread socket → R-writer wedges.\r\n        let post_deadline = Instant::now() + Duration::from_secs(5);\r\n        let mut post = 0u32;\r\n        while (max_seq == 0 || post < 8) && Instant::now() < post_deadline {\r\n            match read_frame(&mut r) {\r\n                Ok(f) if f.kind == KIND_OUTPUT => {\r\n                    let ev: OutputEvent = serde_json::from_value(f.payload).unwrap();\r\n                    max_seq = max_seq.max(ev.seq);\r\n                    post += 1;\r\n                }\r\n                Ok(_) => continue,\r\n                Err(_) => break,\r\n            }\r\n        }\r\n        // Hand the Stream back to main UNREAD (main holds it = the black-holed\r\n        // client). The broker's R-writer stays wedged on the remaining burst.\r\n        let _ = hand_tx.send((r, max_seq));\r\n    });\r\n\r\n    // Main: recv SEED-reached, freeze the PRE-burst high-water (trigger NOT fired —\r\n    // R is parked on go_rx), release R.\r\n    let (saw_seed, _seed_max) = seed_rx\r\n        .recv_timeout(Duration::from_secs(35))\r\n        .unwrap_or((false, 0));\r\n    let seed_high = broker.session_output_seq(sid_b).unwrap_or(0);\r\n    let _ = go_tx.send(());\r\n    // Recv the handed-back black-holed Stream (held UNREAD for the rest of the test).\r\n    let handed = hand_rx.recv_timeout(Duration::from_secs(25)).ok();\r\n    let mut r_stream: Option<Stream> = None;\r\n    let mut max_drained_seq: u64 = 0;\r\n    if let Some((stream, mds)) = handed {\r\n        r_stream = Some(stream);\r\n        max_drained_seq = mds;\r\n    }\r\n\r\n    // Prove the burst is real: the victim's ring high-water advances ≥64 frames\r\n    // past the trigger (independent of the wedged controller).\r\n    let mut burst_advanced = false;\r\n    let step_deadline = Instant::now() + Duration::from_secs(20);\r\n    while Instant::now() < step_deadline {\r\n        if let Some(hw) = broker.session_output_seq(sid_b) {\r\n            if hw >= seed_high + 64 {\r\n                burst_advanced = true;\r\n                break;\r\n            }\r\n        }\r\n        thread::sleep(Duration::from_millis(100));\r\n    }\r\n\r\n    // INVARIANT 1 baseline: the unrelated viewer must ALREADY be ticking before\r\n    // the incident window opens (so the \"kept advancing across the incident\"\r\n    // assertion below is meaningful, not vacuous).\r\n    let mut u_baseline = 0u64;\r\n    let u_base_deadline = Instant::now() + Duration::from_secs(15);\r\n    while Instant::now() < u_base_deadline {\r\n        u_baseline = u_frames.load(Ordering::Relaxed);\r\n        if u_baseline > 0 {\r\n            break;\r\n        }\r\n        thread::sleep(Duration::from_millis(50));\r\n    }\r\n\r\n    // ── Observe exactly ONE logical stall eviction via an INDEPENDENT DRAINING\r\n    //    control-plane connection (the KIND_SESSIONS reap runs the 400 ms\r\n    //    stall-evict). Record t_release + freeze the resume cursor for the sid_b\r\n    //    row. Accept a tally of {1, 2}: POISON-BEATS-EVICT (F-039 rider,\r\n    //    doyle-ruled) — the physical watchdog can poison first and record the same\r\n    //    release class, so the tally can step 0→2 when watchdog races the reap. ──\r\n    let mut stall_evicts_seen: u64 = 0;\r\n    let mut t_release: Option<Instant> = None;\r\n    let mut frozen_cursor: Option<u64> = None;\r\n    let step6_deadline = Instant::now() + Duration::from_secs(10);\r\n    while Instant::now() < step6_deadline {\r\n        let sessions = query_sessions(&name, Duration::from_secs(3));\r\n        let count = stall_evict_count(&name, Duration::from_secs(3));\r\n        if count >= 1 {\r\n            stall_evicts_seen = count;\r\n            t_release = Some(Instant::now());\r\n            frozen_cursor = sessions\r\n                .as_ref()\r\n                .and_then(|r| r.sessions.iter().find(|s| s.session_id == sid_b))\r\n                .map(|s| s.resume_seq)\r\n                .or_else(|| {\r\n                    query_sessions(&name, Duration::from_secs(3)).and_then(|r| {\r\n                        r.sessions\r\n                            .iter()\r\n                            .find(|s| s.session_id == sid_b)\r\n                            .map(|s| s.resume_seq)\r\n                    })\r\n                });\r\n            break;\r\n        }\r\n        thread::sleep(Duration::from_millis(150));\r\n    }\r\n\r\n    // ── INVARIANT 3 (its writer exits): WITHOUT dropping R's Stream, poll\r\n    //    conn_handler_count() until it drops to the UNRELATED-VIEWER FLOOR (<= 1),\r\n    //    bounded at t_release + 2 s. `conn_handlers` is a GLOBAL atomic of live\r\n    //    handle_conn threads, incremented once per physical conn. Unlike the r4\r\n    //    gate (whose only persistent conn was the black-holed R, so its floor was\r\n    //    0), THIS harness keeps a persistent UNRELATED viewer conn attached for\r\n    //    invariant 1 — so while BOTH R's handler and the unrelated viewer's\r\n    //    handler are alive the count is atomically >= 2, and it CANNOT reach 1\r\n    //    until R's handler exits. A sample of <= 1 therefore proves R's\r\n    //    handle_conn thread FINISHED (the parked writer was aborted + the reader\r\n    //    EOF-cleaned) WITHOUT the client dropping its end — airtight because\r\n    //    invariant 1 (below) independently proves the unrelated viewer stayed\r\n    //    alive across this window, so the surviving handler is the viewer's, not\r\n    //    R's. (The fresh viewer has not yet attached at this point.) ──\r\n    let step7_deadline = t_release.unwrap_or_else(Instant::now) + Duration::from_secs(2);\r\n    let mut handler_released = false;\r\n    // R1 (doyle gate rework): snapshot the unrelated viewer's frame count AT THE\r\n    // INSTANT the <=1 sample latches. `u_final > u_baseline` alone cannot prove\r\n    // WHICH handler survived — the viewer could accrue frames past baseline and\r\n    // then die BEFORE this sample, leaving R's LEAKED handler as the count==1 and\r\n    // every invariant false-green. Frames observed AFTER this snapshot\r\n    // (`u_final > u_at_release`, asserted with invariant 3 below) prove the\r\n    // surviving handler is the VIEWER's — forcing R's to be the one that exited.\r\n    let mut u_at_release: u64 = 0;\r\n    while Instant::now() < step7_deadline {\r\n        if broker.conn_handler_count() <= 1 {\r\n            handler_released = true;\r\n            u_at_release = u_frames.load(Ordering::Relaxed);\r\n            break;\r\n        }\r\n        thread::sleep(Duration::from_millis(50));\r\n    }\r\n\r\n    // ── INVARIANT 2 (the bad physical connection is canceled/closed within the\r\n    //    bound): a write on the OLD client Stream must FAIL — the physical conn\r\n    //    was retired (poisoned + severed), not merely logically stall-evicted. ──\r\n    let old_client_write: Option<std::io::Result<()>> = r_stream\r\n        .as_mut()\r\n        .map(|s| write_frame(s, &Envelope::new(KIND_SESSIONS, serde_json::Value::Null)));\r\n    let old_client_write_is_err = old_client_write.as_ref().map(|r| r.is_err());\r\n\r\n    // ── INVARIANT 5 proxy (no lock/task remains owned by the retired conn): after\r\n    //    retirement the broker must still answer a control-plane request PROMPTLY.\r\n    //    A lock still held by the retired writer would stall this; a well under the\r\n    //    watchdog reply proves the control plane is free. (Combined with\r\n    //    handler_released above = no lingering task, and the two viewers flowing\r\n    //    below = the fan-out lock is free.) ──\r\n    let cp_start = Instant::now();\r\n    let cp_reply = query_sessions(&name, Duration::from_secs(5));\r\n    let control_plane_elapsed = cp_start.elapsed();\r\n    let control_plane_ok = cp_reply\r\n        .as_ref()\r\n        .map(|r| r.sessions.iter().any(|s| s.session_id == sid_b))\r\n        .unwrap_or(false);\r\n\r\n    // ── INVARIANT 4 (a fresh VIEWER can attach): a brand-new Viewer subscribe on\r\n    //    the retired session from the frozen cursor must COMPLETE and RECEIVE\r\n    //    replayed output (the burst frames buffered in the ring after the cursor).\r\n    //    Bounded thread + channel. ──\r\n    let resume_seq = frozen_cursor.unwrap_or(0);\r\n    let fv_name = name.clone();\r\n    let (fv_tx, fv_rx) =\r\n        std::sync::mpsc::channel::<(Option<SubscribeOutcome>, bool, Option<u64>)>();\r\n    thread::spawn(move || {\r\n        let mut c = connect(&fv_name);\r\n        send(\r\n            &mut c,\r\n            KIND_SUBSCRIBE,\r\n            serde_json::to_value(SubscribeReq {\r\n                session_id: sid_b,\r\n                from_seq: resume_seq,\r\n                intent: AttachIntent::Viewer,\r\n                by: Some(\"fresh-viewer\".to_string()),\r\n                gen: 0, code: None, seal_ceremony: false, })\r\n            .unwrap(),\r\n        );\r\n        // The writer thread resolve_subscribe spawns races the inline SUBSCRIBED\r\n        // reply for the same send lock, so a replay frame can hit the wire BEFORE\r\n        // SUBSCRIBED. Capture BOTH the outcome and the first replayed seq; break\r\n        // once both are known (or the deadline elapses).\r\n        let deadline = Instant::now() + Duration::from_secs(15);\r\n        let mut outcome: Option<SubscribeOutcome> = None;\r\n        let mut got_output = false;\r\n        let mut first_seq: Option<u64> = None;\r\n        while Instant::now() < deadline && (outcome.is_none() || !got_output) {\r\n            match read_frame(&mut c) {\r\n                Ok(f) if f.kind == KIND_SUBSCRIBED => {\r\n                    let reply: SubscribedReply = serde_json::from_value(f.payload).unwrap();\r\n                    outcome = Some(reply.outcome);\r\n                }\r\n                Ok(f) if f.kind == KIND_OUTPUT => {\r\n                    if !got_output {\r\n                        let ev: OutputEvent = serde_json::from_value(f.payload).unwrap();\r\n                        first_seq = Some(ev.seq);\r\n                        got_output = true;\r\n                    }\r\n                }\r\n                Ok(_) => continue,\r\n                Err(_) => break,\r\n            }\r\n        }\r\n        let _ = fv_tx.send((outcome, got_output, first_seq));\r\n    });\r\n    let (fv_outcome, fv_got_output, fv_first_seq) = fv_rx\r\n        .recv_timeout(Duration::from_secs(18))\r\n        .unwrap_or((None, false, None));\r\n\r\n    // ── INVARIANT 1 final sample: the unrelated viewer's frame count must have\r\n    //    ADVANCED across the whole incident window (baseline → now). It kept\r\n    //    receiving output while B's controller conn was black-holed + retired. ──\r\n    let u_final = u_frames.load(Ordering::Relaxed);\r\n\r\n    // ── Teardown BEFORE asserting (so a failing assert leaves no orphan child):\r\n    //    kill both PTY children, stop the broker, drop the black-holed Stream. ──\r\n    if let Some(pid) = broker.session_pid(sid_b) {\r\n        kill_pid(pid);\r\n    }\r\n    if let Some(pid) = broker.session_pid(sid_u) {\r\n        kill_pid(pid);\r\n    }\r\n    broker.stop();\r\n    drop(r_stream);\r\n\r\n    eprintln!(\r\n        \"=== BLACKHOLE-LC HARNESS: saw_seed={saw_seed} seed_high={seed_high} \\\r\n         max_drained_seq={max_drained_seq} burst_advanced={burst_advanced} \\\r\n         u_baseline={u_baseline} u_at_release={u_at_release} u_final={u_final} \\\r\n         stall_evicts_seen={stall_evicts_seen} \\\r\n         handler_released={handler_released} old_client_write_is_err={old_client_write_is_err:?} \\\r\n         control_plane_ok={control_plane_ok} control_plane_elapsed={control_plane_elapsed:?} \\\r\n         resume_seq={resume_seq} fv_outcome={fv_outcome:?} fv_got_output={fv_got_output} \\\r\n         fv_first_seq={fv_first_seq:?} ===\"\r\n    );\r\n\r\n    // ── HARNESS preconditions (test-environment, distinct from the product\r\n    //    invariants): the black-hole must have actually been established, else the\r\n    //    five-invariant observation is vacuous. ──\r\n    assert!(\r\n        old_client_write_is_err.is_some(),\r\n        \"HARNESS: the black-holed controller never handed its stream back — the \\\r\n         PowerShell child likely never produced SEED (saw_seed={saw_seed}). This is a \\\r\n         test-environment failure, not a product regression.\"\r\n    );\r\n    assert!(\r\n        burst_advanced,\r\n        \"HARNESS: the victim PTY/ring high-water must advance ≥64 frames past the \\\r\n         burst trigger (seed_high={seed_high}) — the finite burst must actually flow.\"\r\n    );\r\n    assert!(\r\n        max_drained_seq > 0,\r\n        \"HARNESS: the controller must have drained a real seq>0 frame before \\\r\n         black-holing (load-gated repro); got max_drained_seq={max_drained_seq}\"\r\n    );\r\n    assert!(\r\n        (1..=2).contains(&stall_evicts_seen),\r\n        \"HARNESS: the one wedged controller's release must be tallied once (logical \\\r\n         reap OR physical poison) or twice (both raced — poison-beats-evict, \\\r\n         field-accepted v0.30.6); got {stall_evicts_seen}\"\r\n    );\r\n\r\n    // ── INVARIANT 1: UNRELATED SESSIONS CONTINUE. ──\r\n    assert!(\r\n        u_baseline > 0,\r\n        \"INVARIANT 1: the unrelated session's viewer must have been receiving output \\\r\n         BEFORE the incident (else the 'kept ticking' proof is vacuous); got \\\r\n         u_baseline={u_baseline}\"\r\n    );\r\n    assert!(\r\n        u_final > u_baseline,\r\n        \"INVARIANT 1: the UNRELATED session's viewer must keep receiving output \\\r\n         ACROSS the black-hole incident — the fan-out must never block the whole \\\r\n         broker on the one wedged conn; got u_baseline={u_baseline} u_final={u_final}\"\r\n    );\r\n\r\n    // ── INVARIANT 2: THE BAD PHYSICAL CONNECTION IS CANCELED/CLOSED WITHIN THE\r\n    //    BOUND — a write on the OLD client conn FAILS (it was physically retired,\r\n    //    not merely logically stall-evicted). ──\r\n    assert_eq!(\r\n        old_client_write_is_err,\r\n        Some(true),\r\n        \"INVARIANT 2: a write on the OLD black-holed client connection must FAIL — \\\r\n         the physical conn must be canceled/closed within the write bound of logical \\\r\n         release. A success means the conn is still physically alive (the r4 RED).\"\r\n    );\r\n\r\n    // ── INVARIANT 3: ITS WRITER EXITS — the black-holed conn's handle_conn thread\r\n    //    finishes (conn_handler_count → the <=1 unrelated-viewer floor) within 2 s\r\n    //    of logical release, WITHOUT the client dropping its end. ──\r\n    assert!(\r\n        handler_released,\r","truncated":true,"truncatedBy":"lines","totalLines":519,"totalBytes":21953,"outputLines":467,"outputBytes":21953,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":707,"displayContent":{"text":"fn spawn_session(name: &str, req: SpawnReq) -> (Stream, u64) {\n…\n    let sid = loop {\n…\n        if f.kind == KIND_SPAWNED {\n…\n        }\r\n    };\r\n    (c, sid)\r\n}\r\n\r\n/// Query the broker's stall-evict tally. Bounded thread + channel (a wedge\r\n/// surfaces as a timeout, not a hang). Verbatim sibling of the r4 gate helper.\r\nfn stall_evict_count(name: &str, deadline: Duration) -> u64 {\r\n    let name = name.to_string();\r\n    let (tx, rx) = std::sync::mpsc::channel();\r\n    thread::spawn(move || {\r\n        let mut c = connect(&name);\r\n        send(&mut c, KIND_STALL_EVICTS, serde_json::Value::Null);\r\n        loop {\r\n            match read_frame(&mut c) {\r\n                Ok(f) if f.kind == KIND_STALL_EVICTS_REPLY => {\r\n                    let r: StallEvictsReply = serde_json::from_value(f.payload).unwrap();\r\n                    let _ = tx.send(r.count);\r\n                    return;\r\n                }\r\n                Ok(_) => continue,\r\n                Err(_) => {\r\n                    let _ = tx.send(0);\r\n                    return;\r\n                }\r\n            }\r\n        }\r\n    });\r\n    rx.recv_timeout(deadline).unwrap_or(0)\r\n}\r\n\r\n/// One `KIND_SESSIONS` round-trip, returning the reply. The dispatch runs the\r\n/// `reap_dead_controller` → time-based `stall_evict_controller`, so polling this\r\n/// both DRIVES the logical eviction and reads the frozen `resume_seq` cursor.\r\n/// Bounded thread + channel. Verbatim sibling of the r4 gate helper.\r\nfn query_sessions(name: &str, deadline: Duration) -> Option<SessionsReply> {\r\n    let name = name.to_string();\r\n    let (tx, rx) = std::sync::mpsc::channel();\r\n    thread::spawn(move || {\r\n        let mut c = connect(&name);\r\n        send(&mut c, KIND_SESSIONS, serde_json::Value::Null);\r\n        loop {\r\n            match read_frame(&mut c) {\r\n                Ok(f) if f.kind == KIND_SESSIONS_REPLY => {\r\n                    let r: SessionsReply = serde_json::from_value(f.payload).unwrap();\r\n                    let _ = tx.send(r);\r\n                    return;\r\n                }\r\n                Ok(_) => continue,\r\n                Err(_) => return,\r\n            }\r\n        }\r\n    });\r\n    rx.recv_timeout(deadline).ok()\r\n}\r\n\r\n// [int->REQ-CONN-BLACKHOLE-LIFECYCLE-HARNESS]\r\n#[test]\r\nfn blackholed_controller_lifecycle_five_invariants() {\r\n    let _home = init_home();\r\n    // Shrink the brain-write deadline so the wedged writer is \"stalled\" in ~400 ms\r\n    // instead of 15 s — the structural outcome, fast + deterministic (the knob\r\n    // grants wall-clock only; it never relaxes an assertion).\r\n    std::env::set_var(\"SPT_BRAIN_WRITE_DEADLINE_MS\", \"400\");\r\n\r\n    let name = unique_name();\r\n    let broker = Broker::bind(&name).expect(\"bind broker\");\r\n    let serve = Arc::clone(&broker);\r\n    thread::spawn(move || {\r\n        let _ = serve.serve();\r\n    });\r\n\r\n    // ── UNRELATED session U (invariant 1): a second, independent flood child on\r\n    //    its OWN endpoint. Drop its spawner conn (so U has no controller to wedge\r\n    //    — the PTY keeps producing headless), and attach a SEPARATE DRAINING\r\n    //    viewer whose frame count must keep advancing ACROSS the whole incident. ──\r\n    let (u_spawn, sid_u) = spawn_session(&name, flood_spawn_req(\"blackhole-unrelated-ep\"));\r\n    drop(u_spawn);\r\n    let u_name = name.clone();\r\n    let u_frames = Arc::new(AtomicU64::new(0));\r\n    let u_frames_w = Arc::clone(&u_frames);\r\n    thread::spawn(move || {\r\n        let mut v = connect(&u_name);\r\n        send(\r\n            &mut v,\r\n            KIND_SUBSCRIBE,\r\n            serde_json::to_value(SubscribeReq {\r\n                session_id: sid_u,\r\n                from_seq: 0,\r\n                intent: AttachIntent::Viewer,\r\n                by: Some(\"unrelated-viewer\".to_string()),\r\n                gen: 0, code: None, seal_ceremony: false, })\r\n            .unwrap(),\r\n        );\r\n        while let Ok(f) = read_frame(&mut v) {\r\n            if f.kind == KIND_OUTPUT {\r\n                u_frames_w.fetch_add(1, Ordering::Relaxed);\r\n            }\r\n        }\r\n    });\r\n\r\n    // ── BLACK-HOLED session B: the real SEED+burst PTY child from conn A, then\r\n    //    DROP A so the ONLY persistent brain conn is the black-holed controller R. ──\r\n    let (a, sid_b) = spawn_session(&name, seed_burst_spawn_req(\"blackhole-victim-ep\"));\r\n    drop(a);\r\n\r\n    // ── Remote controller R: subscribes Control from_seq 0, drains until it has\r\n    //    seen `SEED` (a real controller frame flowed), signals main so main can\r\n    //    freeze the PRE-burst high-water, then on main's go fires the burst trigger,\r\n    //    drains a few burst frames (a real seq>0 frame delivered), and STOPS\r\n    //    reading — handing the Stream back to main UNREAD. Main holds it, never\r\n    //    draining, so the broker's R-writer parks in `write_frame` exactly as\r\n    //    against a suspended brain. R's Stream is NEVER dropped until the very end. ──\r\n    let r_name = name.clone();\r\n    let (seed_tx, seed_rx) = std::sync::mpsc::channel::<(bool, u64)>();\r\n    let (go_tx, go_rx) = std::sync::mpsc::channel::<()>();\r\n    let (hand_tx, hand_rx) = std::sync::mpsc::channel::<(Stream, u64)>();\r\n    thread::spawn(move || {\r\n        let mut r = connect(&r_name);\r\n        send(\r\n            &mut r,\r\n            KIND_SUBSCRIBE,\r\n            serde_json::to_value(SubscribeReq {\r\n                session_id: sid_b,\r\n                from_seq: 0,\r\n                intent: AttachIntent::Control,\r\n                by: Some(\"operator-one\".to_string()),\r\n                gen: 0, code: None, seal_ceremony: false, })\r\n            .unwrap(),\r\n        );\r\n        loop {\r\n            match read_frame(&mut r) {\r\n                Ok(f) if f.kind == KIND_SUBSCRIBED => break,\r\n                Ok(_) => continue,\r\n                Err(_) => {\r\n                    let _ = seed_tx.send((false, 0));\r\n                    return;\r\n                }\r\n            }\r\n        }\r\n        // Drain live output until the CONCATENATED decoded bytes contain `SEED`\r\n        // (ConPTY echoes input + emits VT noise, so per-frame matching is\r\n        // unreliable). Track the max seq seen.\r\n        let mut acc: Vec<u8> = Vec::new();\r\n        let mut max_seq: u64 = 0;\r\n        let mut saw_seed = false;\r\n        let seed_deadline = Instant::now() + Duration::from_secs(30);\r\n        while Instant::now() < seed_deadline {\r\n            match read_frame(&mut r) {\r\n                Ok(f) if f.kind == KIND_OUTPUT => {\r\n                    let ev: OutputEvent = serde_json::from_value(f.payload).unwrap();\r\n                    max_seq = max_seq.max(ev.seq);\r\n                    if let Ok(bytes) = decode_bytes(&ev.data_b64) {\r\n                        acc.extend_from_slice(&bytes);\r\n                    }\r\n                    if acc.windows(4).any(|w| w == b\"SEED\") {\r\n                        saw_seed = true;\r\n                        break;\r\n                    }\r\n                }\r\n                Ok(_) => continue,\r\n                Err(_) => break,\r\n            }\r\n        }\r\n        let _ = seed_tx.send((saw_seed, max_seq));\r\n        // Wait for main's go, then fire the burst trigger (`\\r\\n` satisfies the\r\n        // first Read-Host; ack=false = the rc/operator fire-and-forward path).\r\n        let _ = go_rx.recv_timeout(Duration::from_secs(15));\r\n        send(\r\n            &mut r,\r\n            KIND_INPUT,\r\n            serde_json::to_value(InputReq {\r\n                session_id: sid_b,\r\n                data_b64: encode_bytes(b\"\\r\\n\"),\r\n                op_id: None,\r\n                minter: None,\r\n                ack: false,\r\n            })\r\n            .unwrap(),\r\n        );\r\n        // Drain a bounded handful of burst frames so a REAL seq>0 controller frame\r\n        // is proven delivered, then STOP reading forever (the black-hole). The\r\n        // remaining ~1300 frames flood the now-unread socket → R-writer wedges.\r\n        let post_deadline = Instant::now() + Duration::from_secs(5);\r\n        let mut post = 0u32;\r\n        while (max_seq == 0 || post < 8) && Instant::now() < post_deadline {\r\n            match read_frame(&mut r) {\r\n                Ok(f) if f.kind == KIND_OUTPUT => {\r\n                    let ev: OutputEvent = serde_json::from_value(f.payload).unwrap();\r\n                    max_seq = max_seq.max(ev.seq);\r\n                    post += 1;\r\n                }\r\n                Ok(_) => continue,\r\n                Err(_) => break,\r\n            }\r\n        }\r\n        // Hand the Stream back to main UNREAD (main holds it = the black-holed\r\n        // client). The broker's R-writer stays wedged on the remaining burst.\r\n        let _ = hand_tx.send((r, max_seq));\r\n    });\r\n\r\n    // Main: recv SEED-reached, freeze the PRE-burst high-water (trigger NOT fired —\r\n    // R is parked on go_rx), release R.\r\n    let (saw_seed, _seed_max) = seed_rx\r\n        .recv_timeout(Duration::from_secs(35))\r\n        .unwrap_or((false, 0));\r\n    let seed_high = broker.session_output_seq(sid_b).unwrap_or(0);\r\n    let _ = go_tx.send(());\r\n    // Recv the handed-back black-holed Stream (held UNREAD for the rest of the test).\r\n    let handed = hand_rx.recv_timeout(Duration::from_secs(25)).ok();\r\n    let mut r_stream: Option<Stream> = None;\r\n    let mut max_drained_seq: u64 = 0;\r\n    if let Some((stream, mds)) = handed {\r\n        r_stream = Some(stream);\r\n        max_drained_seq = mds;\r\n    }\r\n\r\n    // Prove the burst is real: the victim's ring high-water advances ≥64 frames\r\n    // past the trigger (independent of the wedged controller).\r\n    let mut burst_advanced = false;\r\n    let step_deadline = Instant::now() + Duration::from_secs(20);\r\n    while Instant::now() < step_deadline {\r\n        if let Some(hw) = broker.session_output_seq(sid_b) {\r\n            if hw >= seed_high + 64 {\r\n                burst_advanced = true;\r\n                break;\r\n            }\r\n        }\r\n        thread::sleep(Duration::from_millis(100));\r\n    }\r\n\r\n    // INVARIANT 1 baseline: the unrelated viewer must ALREADY be ticking before\r\n    // the incident window opens (so the \"kept advancing across the incident\"\r\n    // assertion below is meaningful, not vacuous).\r\n    let mut u_baseline = 0u64;\r\n    let u_base_deadline = Instant::now() + Duration::from_secs(15);\r\n    while Instant::now() < u_base_deadline {\r\n        u_baseline = u_frames.load(Ordering::Relaxed);\r\n        if u_baseline > 0 {\r\n            break;\r\n        }\r\n        thread::sleep(Duration::from_millis(50));\r\n    }\r\n\r\n    // ── Observe exactly ONE logical stall eviction via an INDEPENDENT DRAINING\r\n    //    control-plane connection (the KIND_SESSIONS reap runs the 400 ms\r\n    //    stall-evict). Record t_release + freeze the resume cursor for the sid_b\r\n    //    row. Accept a tally of {1, 2}: POISON-BEATS-EVICT (F-039 rider,\r\n    //    doyle-ruled) — the physical watchdog can poison first and record the same\r\n    //    release class, so the tally can step 0→2 when watchdog races the reap. ──\r\n    let mut stall_evicts_seen: u64 = 0;\r\n    let mut t_release: Option<Instant> = None;\r\n    let mut frozen_cursor: Option<u64> = None;\r\n    let step6_deadline = Instant::now() + Duration::from_secs(10);\r\n    while Instant::now() < step6_deadline {\r\n        let sessions = query_sessions(&name, Duration::from_secs(3));\r\n        let count = stall_evict_count(&name, Duration::from_secs(3));\r\n        if count >= 1 {\r\n            stall_evicts_seen = count;\r\n            t_release = Some(Instant::now());\r\n            frozen_cursor = sessions\r\n                .as_ref()\r\n                .and_then(|r| r.sessions.iter().find(|s| s.session_id == sid_b))\r\n                .map(|s| s.resume_seq)\r\n                .or_else(|| {\r\n                    query_sessions(&name, Duration::from_secs(3)).and_then(|r| {\r\n                        r.sessions\r\n                            .iter()\r\n                            .find(|s| s.session_id == sid_b)\r\n                            .map(|s| s.resume_seq)\r\n                    })\r\n                });\r\n            break;\r\n        }\r\n        thread::sleep(Duration::from_millis(150));\r\n    }\r\n\r\n    // ── INVARIANT 3 (its writer exits): WITHOUT dropping R's Stream, poll\r\n    //    conn_handler_count() until it drops to the UNRELATED-VIEWER FLOOR (<= 1),\r\n    //    bounded at t_release + 2 s. `conn_handlers` is a GLOBAL atomic of live\r\n    //    handle_conn threads, incremented once per physical conn. Unlike the r4\r\n    //    gate (whose only persistent conn was the black-holed R, so its floor was\r\n    //    0), THIS harness keeps a persistent UNRELATED viewer conn attached for\r\n    //    invariant 1 — so while BOTH R's handler and the unrelated viewer's\r\n    //    handler are alive the count is atomically >= 2, and it CANNOT reach 1\r\n    //    until R's handler exits. A sample of <= 1 therefore proves R's\r\n    //    handle_conn thread FINISHED (the parked writer was aborted + the reader\r\n    //    EOF-cleaned) WITHOUT the client dropping its end — airtight because\r\n    //    invariant 1 (below) independently proves the unrelated viewer stayed\r\n    //    alive across this window, so the surviving handler is the viewer's, not\r\n    //    R's. (The fresh viewer has not yet attached at this point.) ──\r\n    let step7_deadline = t_release.unwrap_or_else(Instant::now) + Duration::from_secs(2);\r\n    let mut handler_released = false;\r\n    // R1 (doyle gate rework): snapshot the unrelated viewer's frame count AT THE\r\n    // INSTANT the <=1 sample latches. `u_final > u_baseline` alone cannot prove\r\n    // WHICH handler survived — the viewer could accrue frames past baseline and\r\n    // then die BEFORE this sample, leaving R's LEAKED handler as the count==1 and\r\n    // every invariant false-green. Frames observed AFTER this snapshot\r\n    // (`u_final > u_at_release`, asserted with invariant 3 below) prove the\r\n    // surviving handler is the VIEWER's — forcing R's to be the one that exited.\r\n    let mut u_at_release: u64 = 0;\r\n    while Instant::now() < step7_deadline {\r\n        if broker.conn_handler_count() <= 1 {\r\n            handler_released = true;\r\n            u_at_release = u_frames.load(Ordering::Relaxed);\r\n            break;\r\n        }\r\n        thread::sleep(Duration::from_millis(50));\r\n    }\r\n\r\n    // ── INVARIANT 2 (the bad physical connection is canceled/closed within the\r\n    //    bound): a write on the OLD client Stream must FAIL — the physical conn\r\n    //    was retired (poisoned + severed), not merely logically stall-evicted. ──\r\n    let old_client_write: Option<std::io::Result<()>> = r_stream\r\n        .as_mut()\r\n        .map(|s| write_frame(s, &Envelope::new(KIND_SESSIONS, serde_json::Value::Null)));\r\n    let old_client_write_is_err = old_client_write.as_ref().map(|r| r.is_err());\r\n\r\n    // ── INVARIANT 5 proxy (no lock/task remains owned by the retired conn): after\r\n    //    retirement the broker must still answer a control-plane request PROMPTLY.\r\n    //    A lock still held by the retired writer would stall this; a well under the\r\n    //    watchdog reply proves the control plane is free. (Combined with\r\n    //    handler_released above = no lingering task, and the two viewers flowing\r\n    //    below = the fan-out lock is free.) ──\r\n    let cp_start = Instant::now();\r\n    let cp_reply = query_sessions(&name, Duration::from_secs(5));\r\n    let control_plane_elapsed = cp_start.elapsed();\r\n    let control_plane_ok = cp_reply\r\n        .as_ref()\r\n        .map(|r| r.sessions.iter().any(|s| s.session_id == sid_b))\r\n        .unwrap_or(false);\r\n\r\n    // ── INVARIANT 4 (a fresh VIEWER can attach): a brand-new Viewer subscribe on\r\n    //    the retired session from the frozen cursor must COMPLETE and RECEIVE\r\n    //    replayed output (the burst frames buffered in the ring after the cursor).\r\n    //    Bounded thread + channel. ──\r\n    let resume_seq = frozen_cursor.unwrap_or(0);\r\n    let fv_name = name.clone();\r\n    let (fv_tx, fv_rx) =\r\n        std::sync::mpsc::channel::<(Option<SubscribeOutcome>, bool, Option<u64>)>();\r\n    thread::spawn(move || {\r\n        let mut c = connect(&fv_name);\r\n        send(\r\n            &mut c,\r\n            KIND_SUBSCRIBE,\r\n            serde_json::to_value(SubscribeReq {\r\n                session_id: sid_b,\r\n                from_seq: resume_seq,\r\n                intent: AttachIntent::Viewer,\r\n                by: Some(\"fresh-viewer\".to_string()),\r\n                gen: 0, code: None, seal_ceremony: false, })\r\n            .unwrap(),\r\n        );\r\n        // The writer thread resolve_subscribe spawns races the inline SUBSCRIBED\r\n        // reply for the same send lock, so a replay frame can hit the wire BEFORE\r\n        // SUBSCRIBED. Capture BOTH the outcome and the first replayed seq; break\r\n        // once both are known (or the deadline elapses).\r\n        let deadline = Instant::now() + Duration::from_secs(15);\r\n        let mut outcome: Option<SubscribeOutcome> = None;\r\n        let mut got_output = false;\r\n        let mut first_seq: Option<u64> = None;\r\n        while Instant::now() < deadline && (outcome.is_none() || !got_output) {\r\n            match read_frame(&mut c) {\r\n                Ok(f) if f.kind == KIND_SUBSCRIBED => {\r\n                    let reply: SubscribedReply = serde_json::from_value(f.payload).unwrap();\r\n                    outcome = Some(reply.outcome);\r\n                }\r\n                Ok(f) if f.kind == KIND_OUTPUT => {\r\n                    if !got_output {\r\n                        let ev: OutputEvent = serde_json::from_value(f.payload).unwrap();\r\n                        first_seq = Some(ev.seq);\r\n                        got_output = true;\r\n                    }\r\n                }\r\n                Ok(_) => continue,\r\n                Err(_) => break,\r\n            }\r\n        }\r\n        let _ = fv_tx.send((outcome, got_output, first_seq));\r\n    });\r\n    let (fv_outcome, fv_got_output, fv_first_seq) = fv_rx\r\n        .recv_timeout(Duration::from_secs(18))\r\n        .unwrap_or((None, false, None));\r\n\r\n    // ── INVARIANT 1 final sample: the unrelated viewer's frame count must have\r\n    //    ADVANCED across the whole incident window (baseline → now). It kept\r\n    //    receiving output while B's controller conn was black-holed + retired. ──\r\n    let u_final = u_frames.load(Ordering::Relaxed);\r\n\r\n    // ── Teardown BEFORE asserting (so a failing assert leaves no orphan child):\r\n    //    kill both PTY children, stop the broker, drop the black-holed Stream. ──\r\n    if let Some(pid) = broker.session_pid(sid_b) {\r\n        kill_pid(pid);\r\n    }\r\n    if let Some(pid) = broker.session_pid(sid_u) {\r\n        kill_pid(pid);\r\n    }\r\n    broker.stop();\r\n    drop(r_stream);\r\n\r\n    eprintln!(\r\n        \"=== BLACKHOLE-LC HARNESS: saw_seed={saw_seed} seed_high={seed_high} \\\r\n         max_drained_seq={max_drained_seq} burst_advanced={burst_advanced} \\\r\n         u_baseline={u_baseline} u_at_release={u_at_release} u_final={u_final} \\\r\n         stall_evicts_seen={stall_evicts_seen} \\\r\n         handler_released={handler_released} old_client_write_is_err={old_client_write_is_err:?} \\\r\n         control_plane_ok={control_plane_ok} control_plane_elapsed={control_plane_elapsed:?} \\\r\n         resume_seq={resume_seq} fv_outcome={fv_outcome:?} fv_got_output={fv_got_output} \\\r\n         fv_first_seq={fv_first_seq:?} ===\"\r\n    );\r\n\r\n    // ── HARNESS preconditions (test-environment, distinct from the product\r\n    //    invariants): the black-hole must have actually been established, else the\r\n    //    five-invariant observation is vacuous. ──\r\n    assert!(\r\n        old_client_write_is_err.is_some(),\r\n        \"HARNESS: the black-holed controller never handed its stream back — the \\\r\n         PowerShell child likely never produced SEED (saw_seed={saw_seed}). This is a \\\r\n         test-environment failure, not a product regression.\"\r\n    );\r\n    assert!(\r\n        burst_advanced,\r\n        \"HARNESS: the victim PTY/ring high-water must advance ≥64 frames past the \\\r\n         burst trigger (seed_high={seed_high}) — the finite burst must actually flow.\"\r\n    );\r\n    assert!(\r\n        max_drained_seq > 0,\r\n        \"HARNESS: the controller must have drained a real seq>0 frame before \\\r\n         black-holing (load-gated repro); got max_drained_seq={max_drained_seq}\"\r\n    );\r\n    assert!(\r\n        (1..=2).contains(&stall_evicts_seen),\r\n        \"HARNESS: the one wedged controller's release must be tallied once (logical \\\r\n         reap OR physical poison) or twice (both raced — poison-beats-evict, \\\r\n         field-accepted v0.30.6); got {stall_evicts_seen}\"\r\n    );\r\n\r\n    // ── INVARIANT 1: UNRELATED SESSIONS CONTINUE. ──\r\n    assert!(\r\n        u_baseline > 0,\r\n        \"INVARIANT 1: the unrelated session's viewer must have been receiving output \\\r\n         BEFORE the incident (else the 'kept ticking' proof is vacuous); got \\\r\n         u_baseline={u_baseline}\"\r\n    );\r\n    assert!(\r\n        u_final > u_baseline,\r\n        \"INVARIANT 1: the UNRELATED session's viewer must keep receiving output \\\r\n         ACROSS the black-hole incident — the fan-out must never block the whole \\\r\n         broker on the one wedged conn; got u_baseline={u_baseline} u_final={u_final}\"\r\n    );\r\n\r\n    // ── INVARIANT 2: THE BAD PHYSICAL CONNECTION IS CANCELED/CLOSED WITHIN THE\r\n    //    BOUND — a write on the OLD client conn FAILS (it was physically retired,\r\n    //    not merely logically stall-evicted). ──\r\n    assert_eq!(\r\n        old_client_write_is_err,\r\n        Some(true),\r\n        \"INVARIANT 2: a write on the OLD black-holed client connection must FAIL — \\\r\n         the physical conn must be canceled/closed within the write bound of logical \\\r\n         release. A success means the conn is still physically alive (the r4 RED).\"\r\n    );\r\n\r\n    // ── INVARIANT 3: ITS WRITER EXITS — the black-holed conn's handle_conn thread\r\n    //    finishes (conn_handler_count → the <=1 unrelated-viewer floor) within 2 s\r\n    //    of logical release, WITHOUT the client dropping its end. ──\r\n    assert!(\r\n        handler_released,\r\n…\n    );\n…\n}","startLine":180,"lineNumbers":[180,null,183,null,185,null,189,190,191,192,193,194,195,196,197,198,199,200,201,202,203,204,205,206,207,208,209,210,211,212,213,214,215,216,217,218,219,220,221,222,223,224,225,226,227,228,229,230,231,232,233,234,235,236,237,238,239,240,241,242,243,244,245,246,247,248,249,250,251,252,253,254,255,256,257,258,259,260,261,262,263,264,265,266,267,268,269,270,271,272,273,274,275,276,277,278,279,280,281,282,283,284,285,286,287,288,289,290,291,292,293,294,295,296,297,298,299,300,301,302,303,304,305,306,307,308,309,310,311,312,313,314,315,316,317,318,319,320,321,322,323,324,325,326,327,328,329,330,331,332,333,334,335,336,337,338,339,340,341,342,343,344,345,346,347,348,349,350,351,352,353,354,355,356,357,358,359,360,361,362,363,364,365,366,367,368,369,370,371,372,373,374,375,376,377,378,379,380,381,382,383,384,385,386,387,388,389,390,391,392,393,394,395,396,397,398,399,400,401,402,403,404,405,406,407,408,409,410,411,412,413,414,415,416,417,418,419,420,421,422,423,424,425,426,427,428,429,430,431,432,433,434,435,436,437,438,439,440,441,442,443,444,445,446,447,448,449,450,451,452,453,454,455,456,457,458,459,460,461,462,463,464,465,466,467,468,469,470,471,472,473,474,475,476,477,478,479,480,481,482,483,484,485,486,487,488,489,490,491,492,493,494,495,496,497,498,499,500,501,502,503,504,505,506,507,508,509,510,511,512,513,514,515,516,517,518,519,520,521,522,523,524,525,526,527,528,529,530,531,532,533,534,535,536,537,538,539,540,541,542,543,544,545,546,547,548,549,550,551,552,553,554,555,556,557,558,559,560,561,562,563,564,565,566,567,568,569,570,571,572,573,574,575,576,577,578,579,580,581,582,583,584,585,586,587,588,589,590,591,592,593,594,595,596,597,598,599,600,601,602,603,604,605,606,607,608,609,610,611,612,613,614,615,616,617,618,619,620,621,622,623,624,625,626,627,628,629,630,631,632,633,634,635,636,637,638,639,640,641,642,643,644,645,646,647,648,649,650,651,652,653,654,655,null,662,null,707]},"fileSize":34165,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-daemon\\tests\\conn_blackhole_lifecycle.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":707,"totalBytes":21953,"outputLines":467,"outputBytes":21953,"shownRange":{"start":189,"end":655},"nextOffset":656}}},"isError":false,"timestamp":1788999583660}}
{"type":"message","id":"1e5b6781","parentId":"f1bf6e11","timestamp":"2026-09-10T00:19:43.725Z","message":{"role":"toolResult","toolCallId":"call_G5Y4Noa6C1q1JJBlBCCnIEjU|fc_04ba494197989c86016aa1f79f908887d1b7309f531a7ece25","toolName":"read","content":[{"type":"text","text":"### 7.35 The cached ceremony-clock NTP offset must NOT survive an OS clock STEP — an offset measured against the pre-step clock strands every pairing for the TTL  `[REQ-HAZARD-CEREMONY-CLOCK-STEP]`\n- **Failure (paid-for, field evidence — the enlyzeam BIGNET-join RCA 2026-07-06):** enlyzeam's `w32time` was STOPPED, so its system clock drifted +3m30s; the pairing ceremony offset (REQ-PAIR-8, `ntp.rs`) correctly corrected the TOTP clock — until an operator `w32tm /resync` STEPPED the OS clock back −210s under the LIVE daemon. The cached offset had been measured against the OLD (drifted) clock and its snapshot was monotonic-only (`when: Option<Instant>`), so the step was invisible to the cache: for up to the 15-min TTL the ceremony clock stayed ~7 TOTP steps in the past, every `subnet join` returned NO_SEED_HOLDER while fresh-process probes met in <500ms, and only a daemon bounce (which forced a fresh query) healed it. Timeline-proven: refresh cadence 15:0…\n…\n- **Source:** JOIN-TRUTH W1/D3 (doyle /diagnose enlyzeam RCA; the offset defeated by a live clock step is the D3 of four defects, D1 multihome + D2 loud-fail + D4 verbose-clock the siblings).\n\n### 7.36 The broker control plane and PTY fan-out must NEVER block on a single subscriber connection — a suspended brain conn must not wedge control  `[REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]`\n<!-- [doc->REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE] -->\n- **Failure (paid-for, rig-CONFIRMED 2026-07-06/07 — `NtSuspendProcess` on the brain, no update involved):** a controller's writer thread does a BLOCKING socket write to its brain subscriber conn. When that brain is suspended (or black-holed) the write never returns. The output-driven eviction path (`append` → `mark_controller_gone`, bounded by 7.12's `CONTROLLER_WRITE_DEADLINE`) only fires on NEW output, and `reap_dead_controller` (7.29 companion) only caught a writer that had EXITED — so a controller BLOCKED (not exited) on an IDLE session was evicted by NEITHER. Within seconds: attached output froze, `detach` did not release the control stamp (release routed through the wedged conn), reattach was REFUSED (`controlled-by` latched), `rc --take` hung, and `d…\n- **Invariant:** no broker client ever observes a wedge because the BRAIN stalled. (1) The controller's writer publishes its IN-FLIGHT-write window (`write_blocked_since`: `Some(Instant)` set immediately before the blocking `write_frame`, `None` after — the mutex never held ACROSS the write); a reader sees a `Some` older than `BRAIN_WRITE_DEADLINE` (15 s = 3× the 7.12 controller bound, measured full-with-**zero writer progress**, so a slow-but-DRAINING brain never trips it) as a WEDGED writer and stall-evicts it. Eviction is triggered TIME-based, not output-driven: at the take/reattach path (`resolve_subscribe`, so control self-heals against broker truth even on an idle session) AND at the `KIND_SESSIONS` reap. (2) The stamp release is BROKER-side — `clear_c…\n- **spt-core mapping:** `broker.rs` — `ControllerSink::write_blocked_since` + `controller_writer` (marks the in-flight window around both the initial-batch and live-loop writes) + `controller_write_stalled` (pure predicate over injected `now`/`deadline`) + `stall_evict_controller` (Inline stamp at `resolve_subscribe`, Deferred stamp at `reap_dead_controller` so the reap closure does no I/O under the shared sessions lock) + `record_stall_evict`/`stall_evict_stats` (the tally) + `dispatch_stall_evicts` (`KIND_STALL_EVICTS` IPC); `brain.rs` `Brain::stall_evicts`; `cli.rs` `render_stall_evict_line` (daemon-status surface, public wording). unit = `controller_write_stall_predicate_distinguishes_wedged_from_idle` + `stalled_incumbent_is_evicted_on_reattach_and_cont…\n- **Source:** LIFECYCLE-TRUTH W2 (doyle rig `NtSuspendProcess` root + design ruling `docs/W2-DESIGN-RULING.md` @60ec0ed — minimal-plus scope, `BRAIN_WRITE_DEADLINE` 15 s, passive resubscribe + observability rider; todlando build 2026-07-07). The v0.27→v0.28 update-wedge night's flagship defect.\n\n### 7.37 The Layer-1 settle-gate must RE-ARM per delivery on an observable PTY — a mid-session reader reattach (`/clear`) re-creates the head-swallow race  `[REQ-HAZARD-INJECT-SETTLE-REARM]`\n<!-- [doc->REQ-HAZARD-INJECT-SETTLE-REARM] -->\n- **Failure (paid-for, field-CONFIRMED on 0.29.0 — doyle diagnosis + perri screenshot `WindowsTerminal_6iSjya8pMt.png`):** the shipped W5-A settle-gate (7-of-W5, `REQ-INJECT-MULTILINE-INTEGRITY`) gated Layer 1 behind a worker-local ONE-SHOT (`settled_once`) on the premise that the head-swallow race is a STARTUP condition (the input reader not yet attached after spawn). That premise is FALSE: a mid-session `/clear` re-enters the harness's raw-mode input reader, re-attaching it and RE-CREATING the pre-settle window — but the one-shot had already fired at spawn, so `settle_before_inject` was SKIPPED and the head was eaten again. A checkpoint-wake payload injected right after `/clear` arrived head-truncated (mid-path `spt/Cargo.toml)`); echo-verify (Layer 2) is …\n- **Invariant:** the settle-gate re-arms before EVERY delivery on an OBSERVABLE (echoing/interactive) PTY — the bug-prone class, where re-settling is cheap (a tick or two once the reader answers) and necessary (a mid-session reader reattach must be re-confirmed). The steady-state settle is latch-skipped ONLY where the probe is known UNOBSERVABLE on this PTY (a non-echoing ConPTY whose DSR answer never surfaces): there is no reader-reattach race to guard, and each settle would burn the full `INJECT_SETTLE` deadline. The class is discriminated by the settle's OWN return value — `settle_before_inject` returns `true` iff the session ring advanced (probe observed) — latched on the FIRST attempt only; a RE-DRIVE (`attempt > 1`, reached only on a swallowed head) AL…\n- **spt-core mapping:** `broker.rs` `run_inject_worker` — the `settled_once: bool` one-shot is replaced by a `probe_unobservable: bool` latch (init `false` = assume observable → settle first delivery); the gate is the pure `should_settle(attempt, probe_unobservable) = attempt > 1 || !probe_unobservable`, and the first-attempt settle's `observed` return sets `probe_unobservable = !observed`. `settle_before_inject` gains a `bool` return (observed vs timed-out). unit = `should_settle_rearms_on_observable_pty` (observable PTY re-settles each delivery; unobservable-probe PTY skips the steady-state settle; a re-drive settles in both classes — RED-first: restore the one-shot and the observable-PTY assert flips). No int edit: the `inject_control_wedge.rs` mocks (`fi…\n- **Source:** post-0.29.0 field fix (doyle diagnosis + design ruling `BUILD-SPEC-inject-settle-rearm`; todlando build 2026-07-08). An impl bug in the shipped W5-A Layer-1 fix, not a new invariant class — the settle-gate was correct but under-armed.\n\n### 7.38 EVERY write on a physical broker connection is bounded + cancelable + poison-on-failure — no writer holds the send gate across an UNBOUNDED OS write  `[REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]`\n<!-- [doc->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK] -->\n- **Failure (paid-for, field capture + deterministic Windows repro 2026-07-09 — hertz RCA, UPDATE-WEDGE round-4):** `controller_writer` held the `SharedSend = Arc<Mutex<SendHalf>>` guard ACROSS a blocking `write_frame`. On Windows, `interprocess 2.4.2` routes the send to `WriteFileEx` + `SleepEx(INFINITE, alertable)` — this local-socket send path exposes **no supported write timeout** (`set_timeout` → `Unsupported`). An `rc --take` controller consumer that stopped reading blocked the write **indefinitely** (~127.95 s in the capture, released only when a brain restart tore the connections down). The 7.36 logical `stall_evict_controller` removed the controller ROLE but neither canceled the in-flight pipe write, closed the physical connection, nor called `Cance…\n- **Invariant:** every write on a physical broker connection rides ONE broker-owned bounded/cancelable framed-write primitive with an **independent out-of-band aborter**. The `SendHalf` never leaves the conn object (a bypass cannot compile). An ABSOLUTE deadline stamped at write entry covers BOTH the serialized gate-wait AND OS write completion; a per-conn watchdog fires at that deadline and aborts the in-flight op — never relying on the write returning or on opportunistic stall-evict sampling. On deadline / partial write / cancellation / unknown completion: (1) poison the whole physical conn idempotently, (2) abort read+write so `handle_conn` reaches existing EOF cleanup, (3) wait for the canceled op to report completion before releasing its buffer, (4) NEV…\n- **spt-core mapping:** `spt-daemon/src/conn.rs` `BrokerConn` (Condvar-gated serialized write gate + poisoned state + per-in-flight op id + watchdog + `abort_physical()`: cfg(windows) `CancelIoEx` → completion handshake → `DisconnectNamedPipe` raw kernel32 externs; cfg(unix) `UnixStream::shutdown(Both)`); `broker.rs` `SharedSend = Arc<BrokerConn>` + `controller_writer`/`viewer_writer`/`send_frame`/`send_error` routed; `nethost.rs` stream-log/presence sends routed. int = `brain_decouple.rs` `non_draining_controller_stall_evict_releases_writer_and_connection` (`#[cfg(windows)]`, the hertz 9-step: real PTY burst, black-holed remote controller, exactly one logical stall-evict via an independent draining conn, the black-holed conn's `handle_conn` finishes ≤2 s of…\n- **Source:** UPDATE-WEDGE round-4 (hertz root doc `docs/UPDATE-WEDGE-ROOT-CAUSE.md` @f8596ca; doyle gate ruling `docs/UPDATE-WEDGE-RCA-DISPATCH-todlando.md`; todlando build 2026-07-09). The residual 7.36 left open: logical eviction without physical retirement.\n\n---\n\n### 7.39 Per-session identity env (`SPT_ENDPOINT_ID`/`OWL_SESSION_ID`/`SPT_AGENT_ID`) is NEVER inherited — the daemon scrubs it at startup AND on every role spawn, regardless of any role's declared `env_remove`  `[REQ-HAZARD-DAEMON-IDENTITY-ENV-SANITIZE]`\n<!-- [doc->REQ-HAZARD-DAEMON-IDENTITY-ENV-SANITIZE] -->\n…\n\n\n[Showing lines 722-747 of 1091. Use :748 to continue. Some lines truncated to 768 chars]"}],"details":{"truncation":{"content":"\r\n### 7.36 The broker control plane and PTY fan-out must NEVER block on a single subscriber connection — a suspended brain conn must not wedge control  `[REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]`\r\n<!-- [doc->REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE] -->\r\n- **Failure (paid-for, rig-CONFIRMED 2026-07-06/07 — `NtSuspendProcess` on the brain, no update involved):** a controller's writer thread does a BLOCKING socket write to its brain subscriber conn. When that brain is suspended (or black-holed) the write never returns. The output-driven eviction path (`append` → `mark_controller_gone`, bounded by 7.12's `CONTROLLER_WRITE_DEADLINE`) only fires on NEW output, and `reap_dead_controller` (7.29 companion) only caught a writer that had EXITED — so a controller BLOCKED (not exited) on an IDLE session was evicted by NEITHER. Within seconds: attached output froze, `detach` did not release the control stamp (release routed through the wedged conn), reattach was REFUSED (`controlled-by` latched), `rc --take` hung, and `d…\n- **Invariant:** no broker client ever observes a wedge because the BRAIN stalled. (1) The controller's writer publishes its IN-FLIGHT-write window (`write_blocked_since`: `Some(Instant)` set immediately before the blocking `write_frame`, `None` after — the mutex never held ACROSS the write); a reader sees a `Some` older than `BRAIN_WRITE_DEADLINE` (15 s = 3× the 7.12 controller bound, measured full-with-**zero writer progress**, so a slow-but-DRAINING brain never trips it) as a WEDGED writer and stall-evicts it. Eviction is triggered TIME-based, not output-driven: at the take/reattach path (`resolve_subscribe`, so control self-heals against broker truth even on an idle session) AND at the `KIND_SESSIONS` reap. (2) The stamp release is BROKER-side — `clear_c…\n- **spt-core mapping:** `broker.rs` — `ControllerSink::write_blocked_since` + `controller_writer` (marks the in-flight window around both the initial-batch and live-loop writes) + `controller_write_stalled` (pure predicate over injected `now`/`deadline`) + `stall_evict_controller` (Inline stamp at `resolve_subscribe`, Deferred stamp at `reap_dead_controller` so the reap closure does no I/O under the shared sessions lock) + `record_stall_evict`/`stall_evict_stats` (the tally) + `dispatch_stall_evicts` (`KIND_STALL_EVICTS` IPC); `brain.rs` `Brain::stall_evicts`; `cli.rs` `render_stall_evict_line` (daemon-status surface, public wording). unit = `controller_write_stall_predicate_distinguishes_wedged_from_idle` + `stalled_incumbent_is_evicted_on_reattach_and_cont…\n- **Source:** LIFECYCLE-TRUTH W2 (doyle rig `NtSuspendProcess` root + design ruling `docs/W2-DESIGN-RULING.md` @60ec0ed — minimal-plus scope, `BRAIN_WRITE_DEADLINE` 15 s, passive resubscribe + observability rider; todlando build 2026-07-07). The v0.27→v0.28 update-wedge night's flagship defect.\r\n\r\n### 7.37 The Layer-1 settle-gate must RE-ARM per delivery on an observable PTY — a mid-session reader reattach (`/clear`) re-creates the head-swallow race  `[REQ-HAZARD-INJECT-SETTLE-REARM]`\r\n<!-- [doc->REQ-HAZARD-INJECT-SETTLE-REARM] -->\r\n- **Failure (paid-for, field-CONFIRMED on 0.29.0 — doyle diagnosis + perri screenshot `WindowsTerminal_6iSjya8pMt.png`):** the shipped W5-A settle-gate (7-of-W5, `REQ-INJECT-MULTILINE-INTEGRITY`) gated Layer 1 behind a worker-local ONE-SHOT (`settled_once`) on the premise that the head-swallow race is a STARTUP condition (the input reader not yet attached after spawn). That premise is FALSE: a mid-session `/clear` re-enters the harness's raw-mode input reader, re-attaching it and RE-CREATING the pre-settle window — but the one-shot had already fired at spawn, so `settle_before_inject` was SKIPPED and the head was eaten again. A checkpoint-wake payload injected right after `/clear` arrived head-truncated (mid-path `spt/Cargo.toml)`); echo-verify (Layer 2) is …\n- **Invariant:** the settle-gate re-arms before EVERY delivery on an OBSERVABLE (echoing/interactive) PTY — the bug-prone class, where re-settling is cheap (a tick or two once the reader answers) and necessary (a mid-session reader reattach must be re-confirmed). The steady-state settle is latch-skipped ONLY where the probe is known UNOBSERVABLE on this PTY (a non-echoing ConPTY whose DSR answer never surfaces): there is no reader-reattach race to guard, and each settle would burn the full `INJECT_SETTLE` deadline. The class is discriminated by the settle's OWN return value — `settle_before_inject` returns `true` iff the session ring advanced (probe observed) — latched on the FIRST attempt only; a RE-DRIVE (`attempt > 1`, reached only on a swallowed head) AL…\n- **spt-core mapping:** `broker.rs` `run_inject_worker` — the `settled_once: bool` one-shot is replaced by a `probe_unobservable: bool` latch (init `false` = assume observable → settle first delivery); the gate is the pure `should_settle(attempt, probe_unobservable) = attempt > 1 || !probe_unobservable`, and the first-attempt settle's `observed` return sets `probe_unobservable = !observed`. `settle_before_inject` gains a `bool` return (observed vs timed-out). unit = `should_settle_rearms_on_observable_pty` (observable PTY re-settles each delivery; unobservable-probe PTY skips the steady-state settle; a re-drive settles in both classes — RED-first: restore the one-shot and the observable-PTY assert flips). No int edit: the `inject_control_wedge.rs` mocks (`fi…\n- **Source:** post-0.29.0 field fix (doyle diagnosis + design ruling `BUILD-SPEC-inject-settle-rearm`; todlando build 2026-07-08). An impl bug in the shipped W5-A Layer-1 fix, not a new invariant class — the settle-gate was correct but under-armed.\r\n\r\n### 7.38 EVERY write on a physical broker connection is bounded + cancelable + poison-on-failure — no writer holds the send gate across an UNBOUNDED OS write  `[REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]`\r\n<!-- [doc->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK] -->\r\n- **Failure (paid-for, field capture + deterministic Windows repro 2026-07-09 — hertz RCA, UPDATE-WEDGE round-4):** `controller_writer` held the `SharedSend = Arc<Mutex<SendHalf>>` guard ACROSS a blocking `write_frame`. On Windows, `interprocess 2.4.2` routes the send to `WriteFileEx` + `SleepEx(INFINITE, alertable)` — this local-socket send path exposes **no supported write timeout** (`set_timeout` → `Unsupported`). An `rc --take` controller consumer that stopped reading blocked the write **indefinitely** (~127.95 s in the capture, released only when a brain restart tore the connections down). The 7.36 logical `stall_evict_controller` removed the controller ROLE but neither canceled the in-flight pipe write, closed the physical connection, nor called `Cance…\n- **Invariant:** every write on a physical broker connection rides ONE broker-owned bounded/cancelable framed-write primitive with an **independent out-of-band aborter**. The `SendHalf` never leaves the conn object (a bypass cannot compile). An ABSOLUTE deadline stamped at write entry covers BOTH the serialized gate-wait AND OS write completion; a per-conn watchdog fires at that deadline and aborts the in-flight op — never relying on the write returning or on opportunistic stall-evict sampling. On deadline / partial write / cancellation / unknown completion: (1) poison the whole physical conn idempotently, (2) abort read+write so `handle_conn` reaches existing EOF cleanup, (3) wait for the canceled op to report completion before releasing its buffer, (4) NEV…\n- **spt-core mapping:** `spt-daemon/src/conn.rs` `BrokerConn` (Condvar-gated serialized write gate + poisoned state + per-in-flight op id + watchdog + `abort_physical()`: cfg(windows) `CancelIoEx` → completion handshake → `DisconnectNamedPipe` raw kernel32 externs; cfg(unix) `UnixStream::shutdown(Both)`); `broker.rs` `SharedSend = Arc<BrokerConn>` + `controller_writer`/`viewer_writer`/`send_frame`/`send_error` routed; `nethost.rs` stream-log/presence sends routed. int = `brain_decouple.rs` `non_draining_controller_stall_evict_releases_writer_and_connection` (`#[cfg(windows)]`, the hertz 9-step: real PTY burst, black-holed remote controller, exactly one logical stall-evict via an independent draining conn, the black-holed conn's `handle_conn` finishes ≤2 s of…\n- **Source:** UPDATE-WEDGE round-4 (hertz root doc `docs/UPDATE-WEDGE-ROOT-CAUSE.md` @f8596ca; doyle gate ruling `docs/UPDATE-WEDGE-RCA-DISPATCH-todlando.md`; todlando build 2026-07-09). The residual 7.36 left open: logical eviction without physical retirement.\r\n\r\n---\r\n\r\n### 7.39 Per-session identity env (`SPT_ENDPOINT_ID`/`OWL_SESSION_ID`/`SPT_AGENT_ID`) is NEVER inherited — the daemon scrubs it at startup AND on every role spawn, regardless of any role's declared `env_remove`  `[REQ-HAZARD-DAEMON-IDENTITY-ENV-SANITIZE]`\r\n<!-- [doc->REQ-HAZARD-DAEMON-IDENTITY-ENV-SANITIZE] -->\r","truncated":true,"truncatedBy":"lines","totalLines":370,"totalBytes":14823,"outputLines":26,"outputBytes":14823,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":1091,"displayContent":{"text":"### 7.35 The cached ceremony-clock NTP offset must NOT survive an OS clock STEP — an offset measured against the pre-step clock strands every pairing for the TTL  `[REQ-HAZARD-CEREMONY-CLOCK-STEP]`\n- **Failure (paid-for, field evidence — the enlyzeam BIGNET-join RCA 2026-07-06):** enlyzeam's `w32time` was STOPPED, so its system clock drifted +3m30s; the pairing ceremony offset (REQ-PAIR-8, `ntp.rs`) correctly corrected the TOTP clock — until an operator `w32tm /resync` STEPPED the OS clock back −210s under the LIVE daemon. The cached offset had been measured against the OLD (drifted) clock and its snapshot was monotonic-only (`when: Option<Instant>`), so the step was invisible to the cache: for up to the 15-min TTL the ceremony clock stayed ~7 TOTP steps in the past, every `subnet join` returned NO_SEED_HOLDER while fresh-process probes met in <500ms, and only a daemon bounce (which forced a fresh query) healed it. Timeline-proven: refresh cadence 15:0…\n…\n- **Source:** JOIN-TRUTH W1/D3 (doyle /diagnose enlyzeam RCA; the offset defeated by a live clock step is the D3 of four defects, D1 multihome + D2 loud-fail + D4 verbose-clock the siblings).\n\r\n### 7.36 The broker control plane and PTY fan-out must NEVER block on a single subscriber connection — a suspended brain conn must not wedge control  `[REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]`\r\n<!-- [doc->REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE] -->\r\n- **Failure (paid-for, rig-CONFIRMED 2026-07-06/07 — `NtSuspendProcess` on the brain, no update involved):** a controller's writer thread does a BLOCKING socket write to its brain subscriber conn. When that brain is suspended (or black-holed) the write never returns. The output-driven eviction path (`append` → `mark_controller_gone`, bounded by 7.12's `CONTROLLER_WRITE_DEADLINE`) only fires on NEW output, and `reap_dead_controller` (7.29 companion) only caught a writer that had EXITED — so a controller BLOCKED (not exited) on an IDLE session was evicted by NEITHER. Within seconds: attached output froze, `detach` did not release the control stamp (release routed through the wedged conn), reattach was REFUSED (`controlled-by` latched), `rc --take` hung, and `d…\n- **Invariant:** no broker client ever observes a wedge because the BRAIN stalled. (1) The controller's writer publishes its IN-FLIGHT-write window (`write_blocked_since`: `Some(Instant)` set immediately before the blocking `write_frame`, `None` after — the mutex never held ACROSS the write); a reader sees a `Some` older than `BRAIN_WRITE_DEADLINE` (15 s = 3× the 7.12 controller bound, measured full-with-**zero writer progress**, so a slow-but-DRAINING brain never trips it) as a WEDGED writer and stall-evicts it. Eviction is triggered TIME-based, not output-driven: at the take/reattach path (`resolve_subscribe`, so control self-heals against broker truth even on an idle session) AND at the `KIND_SESSIONS` reap. (2) The stamp release is BROKER-side — `clear_c…\n- **spt-core mapping:** `broker.rs` — `ControllerSink::write_blocked_since` + `controller_writer` (marks the in-flight window around both the initial-batch and live-loop writes) + `controller_write_stalled` (pure predicate over injected `now`/`deadline`) + `stall_evict_controller` (Inline stamp at `resolve_subscribe`, Deferred stamp at `reap_dead_controller` so the reap closure does no I/O under the shared sessions lock) + `record_stall_evict`/`stall_evict_stats` (the tally) + `dispatch_stall_evicts` (`KIND_STALL_EVICTS` IPC); `brain.rs` `Brain::stall_evicts`; `cli.rs` `render_stall_evict_line` (daemon-status surface, public wording). unit = `controller_write_stall_predicate_distinguishes_wedged_from_idle` + `stalled_incumbent_is_evicted_on_reattach_and_cont…\n- **Source:** LIFECYCLE-TRUTH W2 (doyle rig `NtSuspendProcess` root + design ruling `docs/W2-DESIGN-RULING.md` @60ec0ed — minimal-plus scope, `BRAIN_WRITE_DEADLINE` 15 s, passive resubscribe + observability rider; todlando build 2026-07-07). The v0.27→v0.28 update-wedge night's flagship defect.\r\n\r\n### 7.37 The Layer-1 settle-gate must RE-ARM per delivery on an observable PTY — a mid-session reader reattach (`/clear`) re-creates the head-swallow race  `[REQ-HAZARD-INJECT-SETTLE-REARM]`\r\n<!-- [doc->REQ-HAZARD-INJECT-SETTLE-REARM] -->\r\n- **Failure (paid-for, field-CONFIRMED on 0.29.0 — doyle diagnosis + perri screenshot `WindowsTerminal_6iSjya8pMt.png`):** the shipped W5-A settle-gate (7-of-W5, `REQ-INJECT-MULTILINE-INTEGRITY`) gated Layer 1 behind a worker-local ONE-SHOT (`settled_once`) on the premise that the head-swallow race is a STARTUP condition (the input reader not yet attached after spawn). That premise is FALSE: a mid-session `/clear` re-enters the harness's raw-mode input reader, re-attaching it and RE-CREATING the pre-settle window — but the one-shot had already fired at spawn, so `settle_before_inject` was SKIPPED and the head was eaten again. A checkpoint-wake payload injected right after `/clear` arrived head-truncated (mid-path `spt/Cargo.toml)`); echo-verify (Layer 2) is …\n- **Invariant:** the settle-gate re-arms before EVERY delivery on an OBSERVABLE (echoing/interactive) PTY — the bug-prone class, where re-settling is cheap (a tick or two once the reader answers) and necessary (a mid-session reader reattach must be re-confirmed). The steady-state settle is latch-skipped ONLY where the probe is known UNOBSERVABLE on this PTY (a non-echoing ConPTY whose DSR answer never surfaces): there is no reader-reattach race to guard, and each settle would burn the full `INJECT_SETTLE` deadline. The class is discriminated by the settle's OWN return value — `settle_before_inject` returns `true` iff the session ring advanced (probe observed) — latched on the FIRST attempt only; a RE-DRIVE (`attempt > 1`, reached only on a swallowed head) AL…\n- **spt-core mapping:** `broker.rs` `run_inject_worker` — the `settled_once: bool` one-shot is replaced by a `probe_unobservable: bool` latch (init `false` = assume observable → settle first delivery); the gate is the pure `should_settle(attempt, probe_unobservable) = attempt > 1 || !probe_unobservable`, and the first-attempt settle's `observed` return sets `probe_unobservable = !observed`. `settle_before_inject` gains a `bool` return (observed vs timed-out). unit = `should_settle_rearms_on_observable_pty` (observable PTY re-settles each delivery; unobservable-probe PTY skips the steady-state settle; a re-drive settles in both classes — RED-first: restore the one-shot and the observable-PTY assert flips). No int edit: the `inject_control_wedge.rs` mocks (`fi…\n- **Source:** post-0.29.0 field fix (doyle diagnosis + design ruling `BUILD-SPEC-inject-settle-rearm`; todlando build 2026-07-08). An impl bug in the shipped W5-A Layer-1 fix, not a new invariant class — the settle-gate was correct but under-armed.\r\n\r\n### 7.38 EVERY write on a physical broker connection is bounded + cancelable + poison-on-failure — no writer holds the send gate across an UNBOUNDED OS write  `[REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]`\r\n<!-- [doc->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK] -->\r\n- **Failure (paid-for, field capture + deterministic Windows repro 2026-07-09 — hertz RCA, UPDATE-WEDGE round-4):** `controller_writer` held the `SharedSend = Arc<Mutex<SendHalf>>` guard ACROSS a blocking `write_frame`. On Windows, `interprocess 2.4.2` routes the send to `WriteFileEx` + `SleepEx(INFINITE, alertable)` — this local-socket send path exposes **no supported write timeout** (`set_timeout` → `Unsupported`). An `rc --take` controller consumer that stopped reading blocked the write **indefinitely** (~127.95 s in the capture, released only when a brain restart tore the connections down). The 7.36 logical `stall_evict_controller` removed the controller ROLE but neither canceled the in-flight pipe write, closed the physical connection, nor called `Cance…\n- **Invariant:** every write on a physical broker connection rides ONE broker-owned bounded/cancelable framed-write primitive with an **independent out-of-band aborter**. The `SendHalf` never leaves the conn object (a bypass cannot compile). An ABSOLUTE deadline stamped at write entry covers BOTH the serialized gate-wait AND OS write completion; a per-conn watchdog fires at that deadline and aborts the in-flight op — never relying on the write returning or on opportunistic stall-evict sampling. On deadline / partial write / cancellation / unknown completion: (1) poison the whole physical conn idempotently, (2) abort read+write so `handle_conn` reaches existing EOF cleanup, (3) wait for the canceled op to report completion before releasing its buffer, (4) NEV…\n- **spt-core mapping:** `spt-daemon/src/conn.rs` `BrokerConn` (Condvar-gated serialized write gate + poisoned state + per-in-flight op id + watchdog + `abort_physical()`: cfg(windows) `CancelIoEx` → completion handshake → `DisconnectNamedPipe` raw kernel32 externs; cfg(unix) `UnixStream::shutdown(Both)`); `broker.rs` `SharedSend = Arc<BrokerConn>` + `controller_writer`/`viewer_writer`/`send_frame`/`send_error` routed; `nethost.rs` stream-log/presence sends routed. int = `brain_decouple.rs` `non_draining_controller_stall_evict_releases_writer_and_connection` (`#[cfg(windows)]`, the hertz 9-step: real PTY burst, black-holed remote controller, exactly one logical stall-evict via an independent draining conn, the black-holed conn's `handle_conn` finishes ≤2 s of…\n- **Source:** UPDATE-WEDGE round-4 (hertz root doc `docs/UPDATE-WEDGE-ROOT-CAUSE.md` @f8596ca; doyle gate ruling `docs/UPDATE-WEDGE-RCA-DISPATCH-todlando.md`; todlando build 2026-07-09). The residual 7.36 left open: logical eviction without physical retirement.\r\n\r\n---\r\n\r\n### 7.39 Per-session identity env (`SPT_ENDPOINT_ID`/`OWL_SESSION_ID`/`SPT_AGENT_ID`) is NEVER inherited — the daemon scrubs it at startup AND on every role spawn, regardless of any role's declared `env_remove`  `[REQ-HAZARD-DAEMON-IDENTITY-ENV-SANITIZE]`\r\n<!-- [doc->REQ-HAZARD-DAEMON-IDENTITY-ENV-SANITIZE] -->\r\n…\n","startLine":717,"lineNumbers":[717,718,null,721,722,723,724,725,726,727,728,729,730,731,732,733,734,735,736,737,738,739,740,741,742,743,744,745,746,747,null,754]},"fileSize":356704,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\docs\\KNOWN-HAZARDS.md"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":1091,"totalBytes":14823,"outputLines":26,"outputBytes":14823,"shownRange":{"start":722,"end":747},"nextOffset":748},"limits":{"columnTruncated":{"maxColumn":768}}}},"isError":false,"timestamp":1788999583724}}
{"type":"message","id":"d7c71a8b","parentId":"1e5b6781","timestamp":"2026-09-10T00:19:43.733Z","message":{"role":"toolResult","toolCallId":"call_X5zYm59djDnj7Mn5roAGvf4q|fc_04ba494197989c86016aa1f79f909487d19af158c34208bd6e","toolName":"grep","content":[{"type":"text","text":"# .worktrees/286-churn/crates/\n\n## spt-daemon/src/\n### msg.rs\n 1543|#[derive(Debug, Clone, Serialize, Deserialize)]\n*1544|pub struct SessionsReply {\n 1545|    #[serde(default)]\n 1546|    pub sessions: Vec<SessionInfo>,\n 1547|    /// Endpoints whose bring-up the broker has STARTED but whose session row\n### broker.rs\n 9935|// [impl->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\n*9936|fn send_frame(send: &SharedSend, frame: &Envelope) {\n 9937|    let _ = send.write(frame);\n 9938|}\n 9939|\n### codec.rs\n 21|/// Serialize and write one [`Envelope`] as a length-prefixed JSON frame.\n*22|pub fn write_frame<W: Write>(w: &mut W, env: &Envelope) -> io::Result<()> {\n 23|    let body =\n 24|        serde_json::to_vec(env).map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n 25|    if body.len() > MAX_FRAME_BYTES {\n\n## spt-proto/src/\n### emit.rs\n 121|// [impl->REQ-EMIT-SINGLE-WRITE]\n*122|pub fn write_line<W: Write + ?Sized>(w: &mut W, rendered: &str) -> io::Result<()> {\n 123|    debug_assert!(\n 124|        !has_interior_break(rendered),\n 125|        \"emit_line! got an interior line break — use emit_block! if the emission \\\n...\n 433|    #[should_panic(expected = \"interior line break\")]\n*434|    fn write_line_debug_asserts_on_an_interior_newline() {\n 435|        let mut w = CountingWriter::new();\n 436|        let _ = write_line(&mut w, \"SERVICE_STARTUP_FAULT:svc: 3 exits\\ntail line\");\n 437|    }"}],"details":{"scopePath":".worktrees/286-churn/crates/spt-daemon/src/msg.rs, .worktrees/286-churn/crates/spt-daemon/src/broker.rs, .worktrees/286-churn/crates/spt-daemon/src/codec.rs, .worktrees/286-churn/crates/spt-proto/src/emit.rs","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":5,"fileCount":4,"files":[".worktrees/286-churn/crates/spt-daemon/src/msg.rs",".worktrees/286-churn/crates/spt-daemon/src/broker.rs",".worktrees/286-churn/crates/spt-daemon/src/codec.rs",".worktrees/286-churn/crates/spt-proto/src/emit.rs"],"fileMatches":[{"path":".worktrees/286-churn/crates/spt-daemon/src/msg.rs","count":1},{"path":".worktrees/286-churn/crates/spt-daemon/src/broker.rs","count":1},{"path":".worktrees/286-churn/crates/spt-daemon/src/codec.rs","count":1},{"path":".worktrees/286-churn/crates/spt-proto/src/emit.rs","count":2}],"truncated":false,"displayContent":"# .worktrees/286-churn/crates/\n\n## spt-daemon/src/\n### msg.rs\n 1543│#[derive(Debug, Clone, Serialize, Deserialize)]\n*1544│pub struct SessionsReply {\n 1545│    #[serde(default)]\n 1546│    pub sessions: Vec<SessionInfo>,\n 1547│    /// Endpoints whose bring-up the broker has STARTED but whose session row\n### broker.rs\n 9935│// [impl->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\n*9936│fn send_frame(send: &SharedSend, frame: &Envelope) {\n 9937│    let _ = send.write(frame);\n 9938│}\n 9939│\n### codec.rs\n 21│/// Serialize and write one [`Envelope`] as a length-prefixed JSON frame.\n*22│pub fn write_frame<W: Write>(w: &mut W, env: &Envelope) -> io::Result<()> {\n 23│    let body =\n 24│        serde_json::to_vec(env).map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;\n 25│    if body.len() > MAX_FRAME_BYTES {\n\n## spt-proto/src/\n### emit.rs\n 121│// [impl->REQ-EMIT-SINGLE-WRITE]\n*122│pub fn write_line<W: Write + ?Sized>(w: &mut W, rendered: &str) -> io::Result<()> {\n 123│    debug_assert!(\n 124│        !has_interior_break(rendered),\n 125│        \"emit_line! got an interior line break — use emit_block! if the emission \\\n    │...\n 433│    #[should_panic(expected = \"interior line break\")]\n*434│    fn write_line_debug_asserts_on_an_interior_newline() {\n 435│        let mut w = CountingWriter::new();\n 436│        let _ = write_line(&mut w, \"SERVICE_STARTUP_FAULT:svc: 3 exits\\ntail line\");\n 437│    }"},"isError":false,"timestamp":1788999583732}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_4913vnwkKG0bUghZtggB4iYo|fc_04ba494197989c86016aa1f7b1658487d1a20b355e3b14df21","toolName":"read","startedAt":"2026-09-10T00:20:01.492Z","args":{"path":".worktrees/286-churn/crates/spt-daemon/src/livehost.rs:566-631"},"intent":"Reading restart census lifetime"},"id":"3d074194","parentId":"d7c71a8b","timestamp":"2026-09-10T00:20:01.492Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_cSxP5ySWyMIv4WN063wa3Kby|fc_04ba494197989c86016aa1f7b1659887d1b666239ba1d6a042","toolName":"read","startedAt":"2026-09-10T00:20:01.492Z","args":{"path":".worktrees/286-churn/crates/spt-daemon/tests/conn_blackhole_lifecycle.rs:647-707"},"intent":"Reading harness final assertions"},"id":"d1c3aae3","parentId":"3d074194","timestamp":"2026-09-10T00:20:01.492Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_oj0v4ZgW9LyDTrJxvv5Q2Mzm|fc_04ba494197989c86016aa1f7b165a087d19e53342fcda6b0d2","toolName":"read","startedAt":"2026-09-10T00:20:01.492Z","args":{"path":".worktrees/286-churn/crates/spt-daemon/src/msg.rs:1535-1555"},"intent":"Reading session schema defaults"},"id":"30a320f0","parentId":"d1c3aae3","timestamp":"2026-09-10T00:20:01.492Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_06blQdJTfWykY1FZ5t6BYiZF|fc_04ba494197989c86016aa1f7b165a487d18c3a721382deb941","toolName":"read","startedAt":"2026-09-10T00:20:01.492Z","args":{"path":".worktrees/286-churn/crates/spt-daemon/src/brain.rs:2390-2442"},"intent":"Reading client failure propagation"},"id":"58695c0e","parentId":"30a320f0","timestamp":"2026-09-10T00:20:01.492Z"}
{"type":"message","id":"d0ec22a3","parentId":"58695c0e","timestamp":"2026-09-10T00:20:01.492Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":"**Verifying source-file lines**"},{"type":"toolCall","id":"call_4913vnwkKG0bUghZtggB4iYo|fc_04ba494197989c86016aa1f7b1658487d1a20b355e3b14df21","name":"read","arguments":{"i":"Reading restart census lifetime","path":".worktrees/286-churn/crates/spt-daemon/src/livehost.rs:566-631"},"intent":"Reading restart census lifetime"},{"type":"toolCall","id":"call_cSxP5ySWyMIv4WN063wa3Kby|fc_04ba494197989c86016aa1f7b1659887d1b666239ba1d6a042","name":"read","arguments":{"i":"Reading harness final assertions","path":".worktrees/286-churn/crates/spt-daemon/tests/conn_blackhole_lifecycle.rs:647-707"},"intent":"Reading harness final assertions"},{"type":"toolCall","id":"call_oj0v4ZgW9LyDTrJxvv5Q2Mzm|fc_04ba494197989c86016aa1f7b165a087d19e53342fcda6b0d2","name":"read","arguments":{"i":"Reading session schema defaults","path":".worktrees/286-churn/crates/spt-daemon/src/msg.rs:1535-1555"},"intent":"Reading session schema defaults"},{"type":"toolCall","id":"call_06blQdJTfWykY1FZ5t6BYiZF|fc_04ba494197989c86016aa1f7b165a487d18c3a721382deb941","name":"read","arguments":{"i":"Reading client failure propagation","path":".worktrees/286-churn/crates/spt-daemon/src/brain.rs:2390-2442"},"intent":"Reading client failure propagation"},{"type":"toolCall","id":"call_KipxCifOChjOo2YZI1P2Fl3h|fc_04ba494197989c86016aa1f7b165ac87d19160c2877e8a1f71","name":"read","arguments":{"i":"Reading broker disconnect cleanup","path":".worktrees/286-churn/crates/spt-daemon/src/broker.rs:6288-6350"},"intent":"Reading broker disconnect cleanup"},{"type":"toolCall","id":"call_jffJAXKOmCgDVTa7XDUNq6vh|fc_04ba494197989c86016aa1f7b165b087d1af8af05c68443252","name":"read","arguments":{"i":"Reading counter lock boundaries","path":".worktrees/286-churn/crates/spt-daemon/src/conn.rs:98-203"},"intent":"Reading counter lock boundaries"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":13879,"output":484,"cacheRead":87168,"cacheWrite":0,"totalTokens":101531,"reasoningTokens":213,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1788999583762,"responseId":"resp_04ba494197989c86016aa1f7a003e487d1b62fb3f5e248472c","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_04ba494197989c86016aa1f7a2e7f887d19477d3b1c42a1dfd","type":"reasoning","content":[],"encrypted_content":"gAAAAABqofexUErvG1mQKPmT4SL-vGYEl9fxImEWNYWaxbHU6w7kBIjXtm-l0QVgmyEFuY2vm_DK4ZACR212-l3zf41TQnlnd70hmKZRFfrXlNjCGUiPdzuaKcIp02COCuJICiL1atgmHxzQd52BA-WifKGD_QGVsp0JBT3Qt_up5wUNcVjT8Sl8wesq6QDtsIL3Mu0NUpU5eZ5BF5ObpfLYflHwtIE1nuAgEjzzNSn55yJxHwUetX20JW8psTs0VCCfqgkGbHKcXDk7YvhT9x0CW-jsZZMuIq8gGVIsxDdizTbddikD6WjHJeBWrmXropkVlMPAAA0P9AHlvLw4IWMJoeQ25SRBGaioSWz5pj7Je8qnKV6dB_q6pTBC1uXru0vj4OtjmK9wBUZ_amEHs7VAUPoK2HwhCosDANVI4YgHCSGp_hU49Y1Bl3iF8xtzuAeKxNZmRD7soaL-_xwua3AJWu5E7psnQbu1WiV6q_w-9ucmWvayiY6QA_-G0GnCDfgCMg2tQUKaBkwN4rHkgNNPXLA4jDNr0IObQupeNTNha-lS5Roz_TEap25IGciAWYUt07NrCzIjnPY1POjq3ORaQB22HSIsG0A2pkJbhGlnIPe4wx-KNeSjAkE8FVKRrMbvVSH7JArI3AVDKOq4uX3Quc1XW2qRAPqHGIQ5iUGb58co293KShgLen_nkd4DIjVbOm_og9G714IwuCRad1Z9hokGWFAJoovBa2CmOecVbwxyCRbQ3hDOsQEexFYKE1EQeJPg-lti9jWLatwX8v9XZ8_DEUj906TxYzj6mm7ELIIOB2j2M4F6b-Q4cj91kHvWCq97TP0CcwYpZ8AYNoI4ga3_P_Lu-NoYnao5bdojpRo2h10QXzMNV4Jur1npGIdUk57ODMGNsX67Qu5DLEEjH86lSgLeYv3yH783v3vvC3EQWnLejzgCYsznzHEvrmW0Uvx7K0z8UrKidZdtNt23FLmmIk1D6OY3-EIYZo4Jxv2PHhf_bm_OlHIlw1EzKgdJoqraUAD-87q7ZdoFyS6OfbLvGLZnIPP_J8rHeXCGp2Me87mMTbRGLf804pyXviK9JDgGQEdmthxo18Ekxm4ui-w3YfKX9Hobx3VPSocZoLDRIVf-orYCfJZghsPEUQBkU0OK9gfDfyda8Fi0RcGGn9wVj-NjLSSzc6m3n2hT024FCyKjmfvepHe-REBQwFAI6bop7UiyLl32UEPCbGUVYBhawNJ1FuYf5u5gvEc5e6SY3fLaGF8qD9xfa-Z7SsO_TUOa7jOJ5melRds13h_3LYYqnKDZ8bk0f4PMsUNsWn9Jri07VZHLVCuJTU3Kj30dySp6mVTDbICfhCB3NhZitS9_XykboTCRcLtRwKPmiUTFTpF-LeUzJOmXPo7u4ct5VXcUTw3Uuo4zj-8zLcmJK6cXPZURdWcTZwsqIZ87kMR4vKjEKR2jrPC4gyqWl8fiKWblEU900K3IC3LyY2SbwHwbiktBffqYndN1nmdeU7uuXiouEqvQUVBepApiFFcRoXZXDULqDJochOHMzrYVxZqLYQbYyaL8wbGzBBIfY7IWzTZ9CD_p72YpKRoQ1s5Pj9_IqBYnOJcEFPMMAaexQtuP2jFo5K0SNzNRZeItbNb_qkt9opNkx0t7CrgkrzAR8xQ1TLzMPBtCLqZV52UhHvFXVKWN90zZ6cNuOERQMPRvNyn9xxyQAYJxK4bry3g62YoPS-aMVhCF3Kcly1CNNVBaIRjYwaoCyJUHA6jx3hJOiDWGgTXB68SEWsXYezuOBHfq6Jv9v2tA-Xz1Vj0Nc4BrmT0rvJsMBcsExKQrdmyurXK-br_iK4Ob0qu7NLmPSKhsR4YJpiQdZ7uDT3MUPrVOzjeWq1sJosgscNIgSO6WlaY5Nd6yKZABKz0JrQpPWNK88LfutOfYfa7AqcijN3u8wfbk2z4v_s3RASPLCmLT-lPsjao5A-2n3jcjKDpBJAKyUa1z1jk9qBjkeOEDfhFvonviqJk7LoSjuKwt9NnS0Dmco3jFAjIL2W1vxatHR3K2QM6Hweq-2NuZZnamxauqTJrQBuOM_r7A0HIoS_YEviQ2d57HMvNG0rBJKt63JTyKkdEFDgK_gv6a6NJytfsj8isjRaO5XybqoOMnBVaj_9Jm4zbcely5ifCo7la51Ytv6FBdVzetlU-R37zZ5C0LvRYSTuWQLPt8X-8y-kjlnqTad2asDTv3W70Zw30EJaNrbRZLjEjjhhpg90u6u8Xc6PzaW8CUC71iIaDJnrPCw8nfTiI0dk_0cXyIH-8ssiJU-Snn54CJqwWTPKjBGiSvS3YMphcr0KeCB8Hea2xOIzNinv56rfRBMW1C4VRnt2TSgwRnCrbbPx02k6zXT94z0LplhPUlEkrx5wBdc00C0ylLIeZACS5eUC2NZNErpH17wxGFARrQHOn7mjNOtE6QhHQaMqNUaddJk5GQiZKxsf3IJN_7EC9vMJ4O6YALExsefX4z1QGkxGFwaPyfL15rMhF69RmvxF-y-Xp0o3pTTFUu3MJ31gpKslzc_CVDrykodd6tB2N4gkmkUrj80Kgj-_LxCCuUSf7h99SNyGcwdhwhzrxjfQ2yRWwg7LhuIjPT52CQnqfI75NiuJWVnTZURr4xmtz9HlcT79eDL_0nOTSas06n1UEsBhqfnmk-c0R5Lncq61ixnGT3eyNQsNYZerQus4SRqsdubfOqS8fJEZCDwvloyc8lFyIUDT3pZ9YvDjFfzoI4UqoO4sx1NB_cR76KICMVIuC4ulemLzYzvGcPt4xfiumijMbKetXzsDNT0TDU","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"summary":[{"type":"summary_text","text":"**Verifying source-file lines**"}],"metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f7b1658487d1a20b355e3b14df21","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading restart census lifetime\",\"path\":\".worktrees/286-churn/crates/spt-daemon/src/livehost.rs:566-631\"}","call_id":"call_4913vnwkKG0bUghZtggB4iYo","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"read","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f7b1659887d1b666239ba1d6a042","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading harness final assertions\",\"path\":\".worktrees/286-churn/crates/spt-daemon/tests/conn_blackhole_lifecycle.rs:647-707\"}","call_id":"call_cSxP5ySWyMIv4WN063wa3Kby","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"read","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f7b165a087d19e53342fcda6b0d2","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading session schema defaults\",\"path\":\".worktrees/286-churn/crates/spt-daemon/src/msg.rs:1535-1555\"}","call_id":"call_oj0v4ZgW9LyDTrJxvv5Q2Mzm","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"read","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f7b165a487d18c3a721382deb941","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading client failure propagation\",\"path\":\".worktrees/286-churn/crates/spt-daemon/src/brain.rs:2390-2442\"}","call_id":"call_06blQdJTfWykY1FZ5t6BYiZF","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"read","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f7b165ac87d19160c2877e8a1f71","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading broker disconnect cleanup\",\"path\":\".worktrees/286-churn/crates/spt-daemon/src/broker.rs:6288-6350\"}","call_id":"call_KipxCifOChjOo2YZI1P2Fl3h","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"read","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f7b165b087d1af8af05c68443252","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading counter lock boundaries\",\"path\":\".worktrees/286-churn/crates/spt-daemon/src/conn.rs:98-203\"}","call_id":"call_jffJAXKOmCgDVTa7XDUNq6vh","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"read","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}}]},"duration":17728.064499999862,"ttft":3177.745399999898,"completedAt":1788999601490,"contextSnapshot":{"promptTokens":101047,"nonMessageTokens":13648,"compactionEpoch":0}}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_KipxCifOChjOo2YZI1P2Fl3h|fc_04ba494197989c86016aa1f7b165ac87d19160c2877e8a1f71","toolName":"read","startedAt":"2026-09-10T00:20:01.492Z","args":{"path":".worktrees/286-churn/crates/spt-daemon/src/broker.rs:6288-6350"},"intent":"Reading broker disconnect cleanup"},"id":"2ba9d0de","parentId":"d0ec22a3","timestamp":"2026-09-10T00:20:01.492Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_jffJAXKOmCgDVTa7XDUNq6vh|fc_04ba494197989c86016aa1f7b165b087d1af8af05c68443252","toolName":"read","startedAt":"2026-09-10T00:20:01.493Z","args":{"path":".worktrees/286-churn/crates/spt-daemon/src/conn.rs:98-203"},"intent":"Reading counter lock boundaries"},"id":"5533ec2c","parentId":"2ba9d0de","timestamp":"2026-09-10T00:20:01.493Z"}
{"type":"message","id":"83f95c23","parentId":"5533ec2c","timestamp":"2026-09-10T00:20:01.504Z","message":{"role":"toolResult","toolCallId":"call_cSxP5ySWyMIv4WN063wa3Kby|fc_04ba494197989c86016aa1f7b1659887d1b666239ba1d6a042","toolName":"read","content":[{"type":"text","text":"fn blackholed_controller_lifecycle_five_invariants() {\n…\n    assert_eq!(\n…\n        \"INVARIANT 2: a write on the OLD black-holed client connection must FAIL — \\\n         the physical conn must be canceled/closed within the write bound of logical \\\n         release. A success means the conn is still physically alive (the r4 RED).\"\n    );\n\n    // ── INVARIANT 3: ITS WRITER EXITS — the black-holed conn's handle_conn thread\n    //    finishes (conn_handler_count → the <=1 unrelated-viewer floor) within 2 s\n    //    of logical release, WITHOUT the client dropping its end. ──\n    assert!(\n        handler_released,\n        \"INVARIANT 3: the black-holed connection's writer must EXIT — \\\n         conn_handler_count must fall to the unrelated-viewer floor (<= 1) within 2 s \\\n         of logical release WITHOUT the client dropping its end (while R's handler is \\\n         alive the global count is >= 2, so <= 1 proves R's handler finished). A stuck \\\n         count means the detached writer stays wedged in write_frame and the handler \\\n         never exits (the r4 RED).\"\n    );\n    // R1 survivor-identity assert: frames observed AFTER the <=1 sample latched\n    // prove the surviving handler is the unrelated VIEWER's (a dead viewer conn\n    // accrues nothing past the sample), so the handler that exited was R's — not\n    // R leaked while the viewer died pre-sample with its earlier frames making\n    // `u_final > u_baseline` false-green.\n    assert!(\n        u_final > u_at_release,\n        \"INVARIANT 3: the unrelated viewer must receive output AFTER the <=1 handler \\\n         sample latched — otherwise the surviving handler could be R's LEAKED one \\\n         with the viewer already dead; got u_at_release={u_at_release} u_final={u_final}\"\n    );\n\n    // ── INVARIANT 4: A FRESH VIEWER CAN ATTACH — completes AND receives output. ──\n    assert!(\n        matches!(fv_outcome, Some(SubscribeOutcome::Viewer)),\n        \"INVARIANT 4: a fresh Viewer subscribe on the retired session must COMPLETE \\\n         as a Viewer; got {fv_outcome:?}\"\n    );\n    assert!(\n        fv_got_output,\n        \"INVARIANT 4: the fresh viewer must RECEIVE replayed output from the frozen \\\n         cursor (resume_seq={resume_seq}) — the fan-out must serve a new subscriber \\\n         after retirement; got fv_got_output={fv_got_output} fv_first_seq={fv_first_seq:?}\"\n    );\n\n    // ── INVARIANT 5: NO LOCK OR TASK REMAINS OWNED BY THE RETIRED CONNECTION —\n    //    asserted via the conjunction of proxies (no direct seam). The prompt\n    //    control-plane reply (no lock held), the released handler (invariant 3, no\n    //    task lingers), and both viewers flowing output (invariants 1 + 4, the\n    //    fan-out lock is free) together are the strongest deterministic proxy for\n    //    \"nothing is still owned.\" Here we pin the control-plane leg. ──\n    assert!(\n        control_plane_ok,\n        \"INVARIANT 5 (proxy): after retirement the broker must still answer a \\\n         control-plane KIND_SESSIONS request (listing the retired session) — a lock \\\n         still owned by the retired writer would stall the control plane; got \\\n         control_plane_ok={control_plane_ok}\"\n    );\n    assert!(\n        control_plane_elapsed < Duration::from_secs(3),\n        \"INVARIANT 5 (proxy): the post-retirement control-plane reply must be PROMPT \\\n         (well under the watchdog) — a slow reply signals a lock still contended by \\\n         the retired connection; took {control_plane_elapsed:?}\"\n    );\n}"}],"details":{"totalLines":707,"displayContent":{"text":"fn blackholed_controller_lifecycle_five_invariants() {\n…\n    assert_eq!(\n…\n        \"INVARIANT 2: a write on the OLD black-holed client connection must FAIL — \\\r\n         the physical conn must be canceled/closed within the write bound of logical \\\r\n         release. A success means the conn is still physically alive (the r4 RED).\"\r\n    );\r\n\r\n    // ── INVARIANT 3: ITS WRITER EXITS — the black-holed conn's handle_conn thread\r\n    //    finishes (conn_handler_count → the <=1 unrelated-viewer floor) within 2 s\r\n    //    of logical release, WITHOUT the client dropping its end. ──\r\n    assert!(\r\n        handler_released,\r\n        \"INVARIANT 3: the black-holed connection's writer must EXIT — \\\r\n         conn_handler_count must fall to the unrelated-viewer floor (<= 1) within 2 s \\\r\n         of logical release WITHOUT the client dropping its end (while R's handler is \\\r\n         alive the global count is >= 2, so <= 1 proves R's handler finished). A stuck \\\r\n         count means the detached writer stays wedged in write_frame and the handler \\\r\n         never exits (the r4 RED).\"\r\n    );\r\n    // R1 survivor-identity assert: frames observed AFTER the <=1 sample latched\r\n    // prove the surviving handler is the unrelated VIEWER's (a dead viewer conn\r\n    // accrues nothing past the sample), so the handler that exited was R's — not\r\n    // R leaked while the viewer died pre-sample with its earlier frames making\r\n    // `u_final > u_baseline` false-green.\r\n    assert!(\r\n        u_final > u_at_release,\r\n        \"INVARIANT 3: the unrelated viewer must receive output AFTER the <=1 handler \\\r\n         sample latched — otherwise the surviving handler could be R's LEAKED one \\\r\n         with the viewer already dead; got u_at_release={u_at_release} u_final={u_final}\"\r\n    );\r\n\r\n    // ── INVARIANT 4: A FRESH VIEWER CAN ATTACH — completes AND receives output. ──\r\n    assert!(\r\n        matches!(fv_outcome, Some(SubscribeOutcome::Viewer)),\r\n        \"INVARIANT 4: a fresh Viewer subscribe on the retired session must COMPLETE \\\r\n         as a Viewer; got {fv_outcome:?}\"\r\n    );\r\n    assert!(\r\n        fv_got_output,\r\n        \"INVARIANT 4: the fresh viewer must RECEIVE replayed output from the frozen \\\r\n         cursor (resume_seq={resume_seq}) — the fan-out must serve a new subscriber \\\r\n         after retirement; got fv_got_output={fv_got_output} fv_first_seq={fv_first_seq:?}\"\r\n    );\r\n\r\n    // ── INVARIANT 5: NO LOCK OR TASK REMAINS OWNED BY THE RETIRED CONNECTION —\r\n    //    asserted via the conjunction of proxies (no direct seam). The prompt\r\n    //    control-plane reply (no lock held), the released handler (invariant 3, no\r\n    //    task lingers), and both viewers flowing output (invariants 1 + 4, the\r\n    //    fan-out lock is free) together are the strongest deterministic proxy for\r\n    //    \"nothing is still owned.\" Here we pin the control-plane leg. ──\r\n    assert!(\r\n        control_plane_ok,\r\n        \"INVARIANT 5 (proxy): after retirement the broker must still answer a \\\r\n         control-plane KIND_SESSIONS request (listing the retired session) — a lock \\\r\n         still owned by the retired writer would stall the control plane; got \\\r\n         control_plane_ok={control_plane_ok}\"\r\n    );\r\n    assert!(\r\n        control_plane_elapsed < Duration::from_secs(3),\r\n        \"INVARIANT 5 (proxy): the post-retirement control-plane reply must be PROMPT \\\r\n         (well under the watchdog) — a slow reply signals a lock still contended by \\\r\n         the retired connection; took {control_plane_elapsed:?}\"\r\n    );\r\n}\r","startLine":247,"lineNumbers":[247,null,643,null,646,647,648,649,650,651,652,653,654,655,656,657,658,659,660,661,662,663,664,665,666,667,668,669,670,671,672,673,674,675,676,677,678,679,680,681,682,683,684,685,686,687,688,689,690,691,692,693,694,695,696,697,698,699,700,701,702,703,704,705,706,707]},"fileSize":34165,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-daemon\\tests\\conn_blackhole_lifecycle.rs"}}},"isError":false,"timestamp":1788999601503}}
{"type":"message","id":"55db2a2c","parentId":"83f95c23","timestamp":"2026-09-10T00:20:01.525Z","message":{"role":"toolResult","toolCallId":"call_4913vnwkKG0bUghZtggB4iYo|fc_04ba494197989c86016aa1f7b1658487d1a20b355e3b14df21","toolName":"read","content":[{"type":"text","text":"/// hiccup — the safe direction, same posture as [`reconcile_hosted_liveness`]).\n// [impl->REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL]\npub fn resume_restart_orphaned_endpoints(\n    owlery: &Path,\n    registered: &[(AdapterRecord, Manifest)],\n    adapters_dir: &Path,\n) {\n    let Some(live) = SessionPoll::default().query(&crate::endpoint::broker_socket_name()) else {\n        return; // broker unreachable — skip (never mass-respawn on a hiccup)\n    };\n    for id in perch::list_self_perch_ids(owlery) {\n        let perch = perch::resolve_perch_path(&id, ParentHint::Infer);\n        let Some(info) = spt_store::info::read_info(&perch) else {\n            continue;\n        };\n        let relay_present = spt_store::registry::resolve_address(&id, owlery).is_some();\n        // [impl->REQ-RESUME-CUSTODY-IDENTITY] the PAIR test, not a bare-pid probe.\n        let custody_pid_alive = resume_in_flight(&perch);\n        match restart_resume_gate(\n            info.status.as_deref(),\n            info.controllable,\n            live.contains(&id),\n            relay_present,\n            custody_pid_alive,\n        ) {\n            RestartResume::Skip => {}\n            RestartResume::RefuseLivePid => {\n                // BELT 2 breadcrumb (feeds the W4 dup-wake rig, seed #7): never re-run\n                // over a live custody process — the dup-spawn / stomped-info.json class.\n                spt_proto::emit_line_err!(\n                    \"DAEMON_RESTART_RESUME_REFUSED:{id}: custody pid still alive — refusing \\\n                     to re-run over a live process (dup guard)\"\n                );\n            }\n            RestartResume::Resume => {\n                // Resume material from the newest ledger row (session + adapter + cwd).\n                let Some(last) = spt_store::sessions::last_k(&perch, 1).into_iter().next_back()\n                else {\n                    spt_proto::emit_line_err!(\n                        \"DAEMON_RESTART_RESUME_SKIP:{id}: online spt-hosted but no ledger \\\n                         session to resume\"\n                    );\n                    continue;\n                };\n                let Some(adapter) = last.adapter.as_deref().or(info.adapter.as_deref()) else {\n                    spt_proto::emit_line_err!(\"DAEMON_RESTART_RESUME_SKIP:{id}: no adapter recorded to resume under\");\n                    continue;\n                };\n                launch_ledger_resume(\n                    &perch,\n                    &id,\n                    adapter,\n                    &last.session_id,\n                    last.cwd.as_deref(),\n                    registered,\n                    adapters_dir,\n                    \"DAEMON_RESTART\",\n                );\n            }\n        }\n    }\n}\n\n/// Host one live endpoint: sweep a stale signoff (3.2), spawn the Psyche (the\n/// daemon online-stamp makes liveness daemon-authoritative), then start the\n/// config-paced pulse driver on its own supervised thread with a held stop flag.\n/// A Psyche-spawn failure tables nothing — the next sweep retries.\n// [impl->REQ-DAEMON-1]\nfn host_one(\n    set: &Arc<LiveSet>,\n…\n) {\n…\n}\n\n[Showing lines 565-634 of 2917. Use :635 to continue]"}],"details":{"truncation":{"content":"/// hiccup — the safe direction, same posture as [`reconcile_hosted_liveness`]).\r\n// [impl->REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL]\r\npub fn resume_restart_orphaned_endpoints(\r\n    owlery: &Path,\r\n    registered: &[(AdapterRecord, Manifest)],\r\n    adapters_dir: &Path,\r\n) {\r\n    let Some(live) = SessionPoll::default().query(&crate::endpoint::broker_socket_name()) else {\r\n        return; // broker unreachable — skip (never mass-respawn on a hiccup)\r\n    };\r\n    for id in perch::list_self_perch_ids(owlery) {\r\n        let perch = perch::resolve_perch_path(&id, ParentHint::Infer);\r\n        let Some(info) = spt_store::info::read_info(&perch) else {\r\n            continue;\r\n        };\r\n        let relay_present = spt_store::registry::resolve_address(&id, owlery).is_some();\r\n        // [impl->REQ-RESUME-CUSTODY-IDENTITY] the PAIR test, not a bare-pid probe.\r\n        let custody_pid_alive = resume_in_flight(&perch);\r\n        match restart_resume_gate(\r\n            info.status.as_deref(),\r\n            info.controllable,\r\n            live.contains(&id),\r\n            relay_present,\r\n            custody_pid_alive,\r\n        ) {\r\n            RestartResume::Skip => {}\r\n            RestartResume::RefuseLivePid => {\r\n                // BELT 2 breadcrumb (feeds the W4 dup-wake rig, seed #7): never re-run\r\n                // over a live custody process — the dup-spawn / stomped-info.json class.\r\n                spt_proto::emit_line_err!(\r\n                    \"DAEMON_RESTART_RESUME_REFUSED:{id}: custody pid still alive — refusing \\\r\n                     to re-run over a live process (dup guard)\"\r\n                );\r\n            }\r\n            RestartResume::Resume => {\r\n                // Resume material from the newest ledger row (session + adapter + cwd).\r\n                let Some(last) = spt_store::sessions::last_k(&perch, 1).into_iter().next_back()\r\n                else {\r\n                    spt_proto::emit_line_err!(\r\n                        \"DAEMON_RESTART_RESUME_SKIP:{id}: online spt-hosted but no ledger \\\r\n                         session to resume\"\r\n                    );\r\n                    continue;\r\n                };\r\n                let Some(adapter) = last.adapter.as_deref().or(info.adapter.as_deref()) else {\r\n                    spt_proto::emit_line_err!(\"DAEMON_RESTART_RESUME_SKIP:{id}: no adapter recorded to resume under\");\r\n                    continue;\r\n                };\r\n                launch_ledger_resume(\r\n                    &perch,\r\n                    &id,\r\n                    adapter,\r\n                    &last.session_id,\r\n                    last.cwd.as_deref(),\r\n                    registered,\r\n                    adapters_dir,\r\n                    \"DAEMON_RESTART\",\r\n                );\r\n            }\r\n        }\r\n    }\r\n}\r\n\r\n/// Host one live endpoint: sweep a stale signoff (3.2), spawn the Psyche (the\r\n/// daemon online-stamp makes liveness daemon-authoritative), then start the\r\n/// config-paced pulse driver on its own supervised thread with a held stop flag.\r\n/// A Psyche-spawn failure tables nothing — the next sweep retries.\r\n// [impl->REQ-DAEMON-1]\r\nfn host_one(\r\n    set: &Arc<LiveSet>,\r","truncated":true,"truncatedBy":"lines","totalLines":2353,"totalBytes":3165,"outputLines":70,"outputBytes":3165,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":2917,"displayContent":{"text":"/// hiccup — the safe direction, same posture as [`reconcile_hosted_liveness`]).\r\n// [impl->REQ-UPDATE-FINISH-ENDPOINT-SURVIVAL]\r\npub fn resume_restart_orphaned_endpoints(\r\n    owlery: &Path,\r\n    registered: &[(AdapterRecord, Manifest)],\r\n    adapters_dir: &Path,\r\n) {\r\n    let Some(live) = SessionPoll::default().query(&crate::endpoint::broker_socket_name()) else {\r\n        return; // broker unreachable — skip (never mass-respawn on a hiccup)\r\n    };\r\n    for id in perch::list_self_perch_ids(owlery) {\r\n        let perch = perch::resolve_perch_path(&id, ParentHint::Infer);\r\n        let Some(info) = spt_store::info::read_info(&perch) else {\r\n            continue;\r\n        };\r\n        let relay_present = spt_store::registry::resolve_address(&id, owlery).is_some();\r\n        // [impl->REQ-RESUME-CUSTODY-IDENTITY] the PAIR test, not a bare-pid probe.\r\n        let custody_pid_alive = resume_in_flight(&perch);\r\n        match restart_resume_gate(\r\n            info.status.as_deref(),\r\n            info.controllable,\r\n            live.contains(&id),\r\n            relay_present,\r\n            custody_pid_alive,\r\n        ) {\r\n            RestartResume::Skip => {}\r\n            RestartResume::RefuseLivePid => {\r\n                // BELT 2 breadcrumb (feeds the W4 dup-wake rig, seed #7): never re-run\r\n                // over a live custody process — the dup-spawn / stomped-info.json class.\r\n                spt_proto::emit_line_err!(\r\n                    \"DAEMON_RESTART_RESUME_REFUSED:{id}: custody pid still alive — refusing \\\r\n                     to re-run over a live process (dup guard)\"\r\n                );\r\n            }\r\n            RestartResume::Resume => {\r\n                // Resume material from the newest ledger row (session + adapter + cwd).\r\n                let Some(last) = spt_store::sessions::last_k(&perch, 1).into_iter().next_back()\r\n                else {\r\n                    spt_proto::emit_line_err!(\r\n                        \"DAEMON_RESTART_RESUME_SKIP:{id}: online spt-hosted but no ledger \\\r\n                         session to resume\"\r\n                    );\r\n                    continue;\r\n                };\r\n                let Some(adapter) = last.adapter.as_deref().or(info.adapter.as_deref()) else {\r\n                    spt_proto::emit_line_err!(\"DAEMON_RESTART_RESUME_SKIP:{id}: no adapter recorded to resume under\");\r\n                    continue;\r\n                };\r\n                launch_ledger_resume(\r\n                    &perch,\r\n                    &id,\r\n                    adapter,\r\n                    &last.session_id,\r\n                    last.cwd.as_deref(),\r\n                    registered,\r\n                    adapters_dir,\r\n                    \"DAEMON_RESTART\",\r\n                );\r\n            }\r\n        }\r\n    }\r\n}\r\n\r\n/// Host one live endpoint: sweep a stale signoff (3.2), spawn the Psyche (the\r\n/// daemon online-stamp makes liveness daemon-authoritative), then start the\r\n/// config-paced pulse driver on its own supervised thread with a held stop flag.\r\n/// A Psyche-spawn failure tables nothing — the next sweep retries.\r\n// [impl->REQ-DAEMON-1]\r\nfn host_one(\r\n    set: &Arc<LiveSet>,\r\n…\n) {\n…\n}","startLine":565,"lineNumbers":[565,566,567,568,569,570,571,572,573,574,575,576,577,578,579,580,581,582,583,584,585,586,587,588,589,590,591,592,593,594,595,596,597,598,599,600,601,602,603,604,605,606,607,608,609,610,611,612,613,614,615,616,617,618,619,620,621,622,623,624,625,626,627,628,629,630,631,632,633,634,null,639,null,658]},"fileSize":150325,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-daemon\\src\\livehost.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":2917,"totalBytes":3165,"outputLines":70,"outputBytes":3165,"shownRange":{"start":565,"end":634},"nextOffset":635}}},"isError":false,"timestamp":1788999601525}}
{"type":"message","id":"50b6c073","parentId":"55db2a2c","timestamp":"2026-09-10T00:20:01.544Z","message":{"role":"toolResult","toolCallId":"call_06blQdJTfWykY1FZ5t6BYiZF|fc_04ba494197989c86016aa1f7b165a487d18c3a721382deb941","toolName":"read","content":[{"type":"text","text":"    fn subscribe_with(\n…\n    ) -> io::Result<()> {\n        // P1c (operator-stream boundary, kitsubito RACEDIAG): a resume-mode brain\n        // (seeded map) that subscribes from `from_seq` must RESET this session's\n        // dedup cursor to `from_seq` — the broker's replay from `from_seq` is then\n        // re-delivered (and the operator's render cursor dedups the overlap),\n        // never lost. WHY THIS, NOT just `read_event`'s `or_insert`: on a re-serve,\n        // `serve_attach` may CONSUME a handoff-replay frame at the K boundary\n        // BEFORE it is `attached` (the frame is dropped by the `if attached` gate\n        // but the snap-above cursor still advanced past K); `attach_as(sid, 0)`\n        // then re-subscribes from 0, and without this reset the broker's re-send of\n        // seq K arrives below the advanced cursor and is deduped → seq K never\n        // reaches the operator viewport → a forward gap. Resetting to `from_seq`\n        // makes the re-subscribe re-deliver everything from its floor. A cold-start\n        // brain (empty map — e.g. the production dispatch serve brain) keeps the\n        // legacy `next_seq` path; this only arms in resume mode.\n        // [impl->REQ-HAZARD-CONTROLLER-WRITER-REORDER]\n        if !self.session_cursors.is_empty() {\n            self.session_cursors.insert(session_id, from_seq);\n        }\n        self.send(\n            KIND_SUBSCRIBE,\n            serde_json::to_value(SubscribeReq {\n                session_id,\n                from_seq,\n                intent,\n                by,\n                gen,\n                code,\n                seal_ceremony,\n            })\n            .expect(\"SubscribeReq serializes\"),\n        )\n    }\n\n    fn send(&mut self, kind: &str, payload: serde_json::Value) -> io::Result<()> {\n        let env = Envelope::new(kind, payload);\n        match &mut self.conn {\n            BrainConn::Whole(stream) => write_frame(stream, &env),\n            BrainConn::Split { send, .. } => write_frame(send, &env),\n        }\n    }\n\n    /// Read one framed [`Envelope`], bounded by an optional wall-clock\n    /// `deadline`. The carrier dispatch (REQ-HAZARD-PUMP-IPC-DEADLINE):\n    ///\n    /// - [`BrainConn::Whole`] — the ordinary blocking client. `None` is the\n    ///   ordinary blocking [`read_frame`]. `Some(deadline)` is REFUSED BY NAME\n    ///   (releases#190): this carrier reads the frame inline, so there is no\n    ///   thread to hand a deadline to and a timeout would have to abandon a\n    ///   half-read frame — a desynced carrier, which is exactly why the pump\n    ///   carrier was SPLIT instead of given a socket timeout. Pre-fix this arm\n    ///   dropped the parameter on the floor, so a caller that asked for a\n    ///   bounded read silently got an unbounded one; the defect was the\n    ///   silence, not the missing timeout.\n    /// - [`BrainConn::Split`] — the pump carrier: the reader thread already owns\n    ///   the blocking read, so a bounded read is just [`Receiver::recv_timeout`]\n    ///   on the remaining budget. A `Timeout` surfaces [`io::ErrorKind::TimedOut`]\n\n[Showing lines 2389-2445 of 3085. Use :2446 to continue]"}],"details":{"truncation":{"content":"    ) -> io::Result<()> {\r\n        // P1c (operator-stream boundary, kitsubito RACEDIAG): a resume-mode brain\r\n        // (seeded map) that subscribes from `from_seq` must RESET this session's\r\n        // dedup cursor to `from_seq` — the broker's replay from `from_seq` is then\r\n        // re-delivered (and the operator's render cursor dedups the overlap),\r\n        // never lost. WHY THIS, NOT just `read_event`'s `or_insert`: on a re-serve,\r\n        // `serve_attach` may CONSUME a handoff-replay frame at the K boundary\r\n        // BEFORE it is `attached` (the frame is dropped by the `if attached` gate\r\n        // but the snap-above cursor still advanced past K); `attach_as(sid, 0)`\r\n        // then re-subscribes from 0, and without this reset the broker's re-send of\r\n        // seq K arrives below the advanced cursor and is deduped → seq K never\r\n        // reaches the operator viewport → a forward gap. Resetting to `from_seq`\r\n        // makes the re-subscribe re-deliver everything from its floor. A cold-start\r\n        // brain (empty map — e.g. the production dispatch serve brain) keeps the\r\n        // legacy `next_seq` path; this only arms in resume mode.\r\n        // [impl->REQ-HAZARD-CONTROLLER-WRITER-REORDER]\r\n        if !self.session_cursors.is_empty() {\r\n            self.session_cursors.insert(session_id, from_seq);\r\n        }\r\n        self.send(\r\n            KIND_SUBSCRIBE,\r\n            serde_json::to_value(SubscribeReq {\r\n                session_id,\r\n                from_seq,\r\n                intent,\r\n                by,\r\n                gen,\r\n                code,\r\n                seal_ceremony,\r\n            })\r\n            .expect(\"SubscribeReq serializes\"),\r\n        )\r\n    }\r\n\r\n    fn send(&mut self, kind: &str, payload: serde_json::Value) -> io::Result<()> {\r\n        let env = Envelope::new(kind, payload);\r\n        match &mut self.conn {\r\n            BrainConn::Whole(stream) => write_frame(stream, &env),\r\n            BrainConn::Split { send, .. } => write_frame(send, &env),\r\n        }\r\n    }\r\n\r\n    /// Read one framed [`Envelope`], bounded by an optional wall-clock\r\n    /// `deadline`. The carrier dispatch (REQ-HAZARD-PUMP-IPC-DEADLINE):\r\n    ///\r\n    /// - [`BrainConn::Whole`] — the ordinary blocking client. `None` is the\r\n    ///   ordinary blocking [`read_frame`]. `Some(deadline)` is REFUSED BY NAME\r\n    ///   (releases#190): this carrier reads the frame inline, so there is no\r\n    ///   thread to hand a deadline to and a timeout would have to abandon a\r\n    ///   half-read frame — a desynced carrier, which is exactly why the pump\r\n    ///   carrier was SPLIT instead of given a socket timeout. Pre-fix this arm\r\n    ///   dropped the parameter on the floor, so a caller that asked for a\r\n    ///   bounded read silently got an unbounded one; the defect was the\r\n    ///   silence, not the missing timeout.\r\n    /// - [`BrainConn::Split`] — the pump carrier: the reader thread already owns\r\n    ///   the blocking read, so a bounded read is just [`Receiver::recv_timeout`]\r\n    ///   on the remaining budget. A `Timeout` surfaces [`io::ErrorKind::TimedOut`]\r","truncated":true,"truncatedBy":"lines","totalLines":697,"totalBytes":3133,"outputLines":57,"outputBytes":3133,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":3085,"displayContent":{"text":"    fn subscribe_with(\n…\n    ) -> io::Result<()> {\r\n        // P1c (operator-stream boundary, kitsubito RACEDIAG): a resume-mode brain\r\n        // (seeded map) that subscribes from `from_seq` must RESET this session's\r\n        // dedup cursor to `from_seq` — the broker's replay from `from_seq` is then\r\n        // re-delivered (and the operator's render cursor dedups the overlap),\r\n        // never lost. WHY THIS, NOT just `read_event`'s `or_insert`: on a re-serve,\r\n        // `serve_attach` may CONSUME a handoff-replay frame at the K boundary\r\n        // BEFORE it is `attached` (the frame is dropped by the `if attached` gate\r\n        // but the snap-above cursor still advanced past K); `attach_as(sid, 0)`\r\n        // then re-subscribes from 0, and without this reset the broker's re-send of\r\n        // seq K arrives below the advanced cursor and is deduped → seq K never\r\n        // reaches the operator viewport → a forward gap. Resetting to `from_seq`\r\n        // makes the re-subscribe re-deliver everything from its floor. A cold-start\r\n        // brain (empty map — e.g. the production dispatch serve brain) keeps the\r\n        // legacy `next_seq` path; this only arms in resume mode.\r\n        // [impl->REQ-HAZARD-CONTROLLER-WRITER-REORDER]\r\n        if !self.session_cursors.is_empty() {\r\n            self.session_cursors.insert(session_id, from_seq);\r\n        }\r\n        self.send(\r\n            KIND_SUBSCRIBE,\r\n            serde_json::to_value(SubscribeReq {\r\n                session_id,\r\n                from_seq,\r\n                intent,\r\n                by,\r\n                gen,\r\n                code,\r\n                seal_ceremony,\r\n            })\r\n            .expect(\"SubscribeReq serializes\"),\r\n        )\r\n    }\r\n\r\n    fn send(&mut self, kind: &str, payload: serde_json::Value) -> io::Result<()> {\r\n        let env = Envelope::new(kind, payload);\r\n        match &mut self.conn {\r\n            BrainConn::Whole(stream) => write_frame(stream, &env),\r\n            BrainConn::Split { send, .. } => write_frame(send, &env),\r\n        }\r\n    }\r\n\r\n    /// Read one framed [`Envelope`], bounded by an optional wall-clock\r\n    /// `deadline`. The carrier dispatch (REQ-HAZARD-PUMP-IPC-DEADLINE):\r\n    ///\r\n    /// - [`BrainConn::Whole`] — the ordinary blocking client. `None` is the\r\n    ///   ordinary blocking [`read_frame`]. `Some(deadline)` is REFUSED BY NAME\r\n    ///   (releases#190): this carrier reads the frame inline, so there is no\r\n    ///   thread to hand a deadline to and a timeout would have to abandon a\r\n    ///   half-read frame — a desynced carrier, which is exactly why the pump\r\n    ///   carrier was SPLIT instead of given a socket timeout. Pre-fix this arm\r\n    ///   dropped the parameter on the floor, so a caller that asked for a\r\n    ///   bounded read silently got an unbounded one; the defect was the\r\n    ///   silence, not the missing timeout.\r\n    /// - [`BrainConn::Split`] — the pump carrier: the reader thread already owns\r\n    ///   the blocking read, so a bounded read is just [`Receiver::recv_timeout`]\r\n    ///   on the remaining budget. A `Timeout` surfaces [`io::ErrorKind::TimedOut`]\r","startLine":2380,"lineNumbers":[2380,null,2389,2390,2391,2392,2393,2394,2395,2396,2397,2398,2399,2400,2401,2402,2403,2404,2405,2406,2407,2408,2409,2410,2411,2412,2413,2414,2415,2416,2417,2418,2419,2420,2421,2422,2423,2424,2425,2426,2427,2428,2429,2430,2431,2432,2433,2434,2435,2436,2437,2438,2439,2440,2441,2442,2443,2444,2445]},"fileSize":160723,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-daemon\\src\\brain.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":3085,"totalBytes":3133,"outputLines":57,"outputBytes":3133,"shownRange":{"start":2389,"end":2445},"nextOffset":2446}}},"isError":false,"timestamp":1788999601543}}
{"type":"message","id":"9c8c36a2","parentId":"50b6c073","timestamp":"2026-09-10T00:20:01.550Z","message":{"role":"toolResult","toolCallId":"call_jffJAXKOmCgDVTa7XDUNq6vh|fc_04ba494197989c86016aa1f7b165b087d1af8af05c68443252","toolName":"read","content":[{"type":"text","text":"/// stderr has no per-line timestamps and broker+brain share one file — without\n/// BOTH clocks, interleaved records cannot be ordered (monotonic) or matched\n/// to field incident times (wall).\n// [impl->REQ-CONN-POISON-ATTRIBUTION]\npub(crate) fn log_stamp() -> String {\n    let stamp = LogTime::now();\n    format!(\"wall_ms={} mono_ms={}\", stamp.wall_ms, stamp.mono_ms)\n}\n\n/// Retain the original write instant even when its healthy line is aggregated.\n#[derive(Clone, Copy)]\nstruct LogTime {\n    wall_ms: u128,\n    mono_ms: u128,\n}\n\nimpl LogTime {\n    fn now() -> Self {\n        Self {\n            wall_ms: SystemTime::now()\n                .duration_since(UNIX_EPOCH)\n                .map(|d| d.as_millis())\n                .unwrap_or(0),\n            mono_ms: MONO_ANCHOR.get_or_init(Instant::now).elapsed().as_millis(),\n        }\n    }\n}\n\nconst HEALTHY_LOG_INTERVAL: Duration = Duration::from_secs(60);\n\nstruct HealthyWindow {\n    since: Instant,\n    starts: u64,\n    closes: u64,\n}\n\n/// One process-wide counter window, not one limiter per connection. No thread,\n/// queue or per-conn healthy string: the next event after 60s emits the totals.\n/// Exceptional and role-transition records bypass this window entirely.\n// [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\nstruct LifecycleLog {\n    healthy: Mutex<HealthyWindow>,\n    #[cfg(test)]\n    capture: Option<Mutex<Vec<String>>>,\n    #[cfg(test)]\n    clock: Option<Mutex<Instant>>,\n}\n\nimpl LifecycleLog {\n    fn new() -> Self {\n        Self {\n            healthy: Mutex::new(HealthyWindow {\n                since: Instant::now(),\n                starts: 0,\n                closes: 0,\n            }),\n            #[cfg(test)]\n            capture: None,\n            #[cfg(test)]\n            clock: None,\n        }\n    }\n\n    #[cfg(test)]\n    fn captured(now: Instant) -> Self {\n        let mut log = Self::new();\n        log.healthy.get_mut().unwrap().since = now;\n        log.capture = Some(Mutex::new(Vec::new()));\n        log.clock = Some(Mutex::new(now));\n        log\n    }\n\n    fn emit(&self, args: std::fmt::Arguments<'_>) {\n        #[cfg(test)]\n        if let Some(lines) = &self.capture {\n            recover(lines).push(args.to_string());\n            return;\n        }\n        spt_proto::emit_line_err!(\"{args}\");\n    }\n\n    fn healthy(&self, start: bool) {\n        let mut window = recover(&self.healthy);\n        // Read time under the counter lock so concurrent callers cannot move\n        // the window backwards after waiting for a newer caller to flush.\n        let now = Instant::now();\n        #[cfg(test)]\n        let now = self.clock.as_ref().map(|c| *recover(c)).unwrap_or(now);\n        if start {\n            window.starts = window.starts.saturating_add(1);\n        } else {\n            window.closes = window.closes.saturating_add(1);\n        }\n        let elapsed = now.saturating_duration_since(window.since);\n        if elapsed < HEALTHY_LOG_INTERVAL {\n            return;\n        }\n        let (starts, closes) = (window.starts, window.closes);\n        *window = HealthyWindow { since: now, starts: 0, closes: 0 };\n        drop(window);\n        self.emit(format_args!(\n            \"CONN_LIFECYCLE_SUMMARY: scope=broker-ipc {} interval_ms={} starts={starts} closes={closes} \\\n             [REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\",\n            log_stamp(), elapsed.as_millis()\n        ));\n    }\n}\n\n/// Attribution-label cap: facts accrete per role attach (a long-lived brain\n/// conn can be controller + viewer + presence subscriber over its life), and\n\n[Showing lines 97-206 of 1000. Use :207 to continue]"}],"details":{"truncation":{"content":"/// stderr has no per-line timestamps and broker+brain share one file — without\r\n/// BOTH clocks, interleaved records cannot be ordered (monotonic) or matched\r\n/// to field incident times (wall).\r\n// [impl->REQ-CONN-POISON-ATTRIBUTION]\r\npub(crate) fn log_stamp() -> String {\r\n    let stamp = LogTime::now();\r\n    format!(\"wall_ms={} mono_ms={}\", stamp.wall_ms, stamp.mono_ms)\r\n}\r\n\r\n/// Retain the original write instant even when its healthy line is aggregated.\r\n#[derive(Clone, Copy)]\r\nstruct LogTime {\r\n    wall_ms: u128,\r\n    mono_ms: u128,\r\n}\r\n\r\nimpl LogTime {\r\n    fn now() -> Self {\r\n        Self {\r\n            wall_ms: SystemTime::now()\r\n                .duration_since(UNIX_EPOCH)\r\n                .map(|d| d.as_millis())\r\n                .unwrap_or(0),\r\n            mono_ms: MONO_ANCHOR.get_or_init(Instant::now).elapsed().as_millis(),\r\n        }\r\n    }\r\n}\r\n\r\nconst HEALTHY_LOG_INTERVAL: Duration = Duration::from_secs(60);\r\n\r\nstruct HealthyWindow {\r\n    since: Instant,\r\n    starts: u64,\r\n    closes: u64,\r\n}\r\n\r\n/// One process-wide counter window, not one limiter per connection. No thread,\r\n/// queue or per-conn healthy string: the next event after 60s emits the totals.\r\n/// Exceptional and role-transition records bypass this window entirely.\r\n// [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\r\nstruct LifecycleLog {\r\n    healthy: Mutex<HealthyWindow>,\r\n    #[cfg(test)]\r\n    capture: Option<Mutex<Vec<String>>>,\r\n    #[cfg(test)]\r\n    clock: Option<Mutex<Instant>>,\r\n}\r\n\r\nimpl LifecycleLog {\r\n    fn new() -> Self {\r\n        Self {\r\n            healthy: Mutex::new(HealthyWindow {\r\n                since: Instant::now(),\r\n                starts: 0,\r\n                closes: 0,\r\n            }),\r\n            #[cfg(test)]\r\n            capture: None,\r\n            #[cfg(test)]\r\n            clock: None,\r\n        }\r\n    }\r\n\r\n    #[cfg(test)]\r\n    fn captured(now: Instant) -> Self {\r\n        let mut log = Self::new();\r\n        log.healthy.get_mut().unwrap().since = now;\r\n        log.capture = Some(Mutex::new(Vec::new()));\r\n        log.clock = Some(Mutex::new(now));\r\n        log\r\n    }\r\n\r\n    fn emit(&self, args: std::fmt::Arguments<'_>) {\r\n        #[cfg(test)]\r\n        if let Some(lines) = &self.capture {\r\n            recover(lines).push(args.to_string());\r\n            return;\r\n        }\r\n        spt_proto::emit_line_err!(\"{args}\");\r\n    }\r\n\r\n    fn healthy(&self, start: bool) {\r\n        let mut window = recover(&self.healthy);\r\n        // Read time under the counter lock so concurrent callers cannot move\r\n        // the window backwards after waiting for a newer caller to flush.\r\n        let now = Instant::now();\r\n        #[cfg(test)]\r\n        let now = self.clock.as_ref().map(|c| *recover(c)).unwrap_or(now);\r\n        if start {\r\n            window.starts = window.starts.saturating_add(1);\r\n        } else {\r\n            window.closes = window.closes.saturating_add(1);\r\n        }\r\n        let elapsed = now.saturating_duration_since(window.since);\r\n        if elapsed < HEALTHY_LOG_INTERVAL {\r\n            return;\r\n        }\r\n        let (starts, closes) = (window.starts, window.closes);\r\n        *window = HealthyWindow { since: now, starts: 0, closes: 0 };\r\n        drop(window);\r\n        self.emit(format_args!(\r\n            \"CONN_LIFECYCLE_SUMMARY: scope=broker-ipc {} interval_ms={} starts={starts} closes={closes} \\\r\n             [REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\",\r\n            log_stamp(), elapsed.as_millis()\r\n        ));\r\n    }\r\n}\r\n\r\n/// Attribution-label cap: facts accrete per role attach (a long-lived brain\r\n/// conn can be controller + viewer + presence subscriber over its life), and\r","truncated":true,"truncatedBy":"lines","totalLines":904,"totalBytes":3635,"outputLines":110,"outputBytes":3635,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":1000,"displayContent":{"text":"/// stderr has no per-line timestamps and broker+brain share one file — without\r\n/// BOTH clocks, interleaved records cannot be ordered (monotonic) or matched\r\n/// to field incident times (wall).\r\n// [impl->REQ-CONN-POISON-ATTRIBUTION]\r\npub(crate) fn log_stamp() -> String {\r\n    let stamp = LogTime::now();\r\n    format!(\"wall_ms={} mono_ms={}\", stamp.wall_ms, stamp.mono_ms)\r\n}\r\n\r\n/// Retain the original write instant even when its healthy line is aggregated.\r\n#[derive(Clone, Copy)]\r\nstruct LogTime {\r\n    wall_ms: u128,\r\n    mono_ms: u128,\r\n}\r\n\r\nimpl LogTime {\r\n    fn now() -> Self {\r\n        Self {\r\n            wall_ms: SystemTime::now()\r\n                .duration_since(UNIX_EPOCH)\r\n                .map(|d| d.as_millis())\r\n                .unwrap_or(0),\r\n            mono_ms: MONO_ANCHOR.get_or_init(Instant::now).elapsed().as_millis(),\r\n        }\r\n    }\r\n}\r\n\r\nconst HEALTHY_LOG_INTERVAL: Duration = Duration::from_secs(60);\r\n\r\nstruct HealthyWindow {\r\n    since: Instant,\r\n    starts: u64,\r\n    closes: u64,\r\n}\r\n\r\n/// One process-wide counter window, not one limiter per connection. No thread,\r\n/// queue or per-conn healthy string: the next event after 60s emits the totals.\r\n/// Exceptional and role-transition records bypass this window entirely.\r\n// [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\r\nstruct LifecycleLog {\r\n    healthy: Mutex<HealthyWindow>,\r\n    #[cfg(test)]\r\n    capture: Option<Mutex<Vec<String>>>,\r\n    #[cfg(test)]\r\n    clock: Option<Mutex<Instant>>,\r\n}\r\n\r\nimpl LifecycleLog {\r\n    fn new() -> Self {\r\n        Self {\r\n            healthy: Mutex::new(HealthyWindow {\r\n                since: Instant::now(),\r\n                starts: 0,\r\n                closes: 0,\r\n            }),\r\n            #[cfg(test)]\r\n            capture: None,\r\n            #[cfg(test)]\r\n            clock: None,\r\n        }\r\n    }\r\n\r\n    #[cfg(test)]\r\n    fn captured(now: Instant) -> Self {\r\n        let mut log = Self::new();\r\n        log.healthy.get_mut().unwrap().since = now;\r\n        log.capture = Some(Mutex::new(Vec::new()));\r\n        log.clock = Some(Mutex::new(now));\r\n        log\r\n    }\r\n\r\n    fn emit(&self, args: std::fmt::Arguments<'_>) {\r\n        #[cfg(test)]\r\n        if let Some(lines) = &self.capture {\r\n            recover(lines).push(args.to_string());\r\n            return;\r\n        }\r\n        spt_proto::emit_line_err!(\"{args}\");\r\n    }\r\n\r\n    fn healthy(&self, start: bool) {\r\n        let mut window = recover(&self.healthy);\r\n        // Read time under the counter lock so concurrent callers cannot move\r\n        // the window backwards after waiting for a newer caller to flush.\r\n        let now = Instant::now();\r\n        #[cfg(test)]\r\n        let now = self.clock.as_ref().map(|c| *recover(c)).unwrap_or(now);\r\n        if start {\r\n            window.starts = window.starts.saturating_add(1);\r\n        } else {\r\n            window.closes = window.closes.saturating_add(1);\r\n        }\r\n        let elapsed = now.saturating_duration_since(window.since);\r\n        if elapsed < HEALTHY_LOG_INTERVAL {\r\n            return;\r\n        }\r\n        let (starts, closes) = (window.starts, window.closes);\r\n        *window = HealthyWindow { since: now, starts: 0, closes: 0 };\r\n        drop(window);\r\n        self.emit(format_args!(\r\n            \"CONN_LIFECYCLE_SUMMARY: scope=broker-ipc {} interval_ms={} starts={starts} closes={closes} \\\r\n             [REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\",\r\n            log_stamp(), elapsed.as_millis()\r\n        ));\r\n    }\r\n}\r\n\r\n/// Attribution-label cap: facts accrete per role attach (a long-lived brain\r\n/// conn can be controller + viewer + presence subscriber over its life), and\r","startLine":97,"lineNumbers":[97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120,121,122,123,124,125,126,127,128,129,130,131,132,133,134,135,136,137,138,139,140,141,142,143,144,145,146,147,148,149,150,151,152,153,154,155,156,157,158,159,160,161,162,163,164,165,166,167,168,169,170,171,172,173,174,175,176,177,178,179,180,181,182,183,184,185,186,187,188,189,190,191,192,193,194,195,196,197,198,199,200,201,202,203,204,205,206]},"fileSize":43321,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-daemon\\src\\conn.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":1000,"totalBytes":3635,"outputLines":110,"outputBytes":3635,"shownRange":{"start":97,"end":206},"nextOffset":207}}},"isError":false,"timestamp":1788999601550}}
{"type":"message","id":"430b7bbe","parentId":"9c8c36a2","timestamp":"2026-09-10T00:20:01.640Z","message":{"role":"toolResult","toolCallId":"call_KipxCifOChjOo2YZI1P2Fl3h|fc_04ba494197989c86016aa1f7b165ac87d19160c2877e8a1f71","toolName":"read","content":[{"type":"text","text":"    fn handle_conn(self: &Arc<Self>, conn: Stream) -> io::Result<()> {\n…\n        // [impl->REQ-SEAL-CEREMONY-ESC-CANCEL]\n        self.abort_ceremonies_for_conn(send.id());\n        // Brain gone: detach as subscriber but leave the sessions running.\n        // Clone the log Arcs out from under the sessions lock FIRST (ADR-0038\n        // Amendment fix 8, KH 7.12 kin): `detach_if` re-stamps info.json\n        // (`viewer_count`/`driven_by`) and emits a lifecycle record — file +\n        // conn I/O that must never run under the shared sessions table lock,\n        // or one slow perch write stalls every spawn/attach/kill dispatch.\n        // [impl->REQ-DISPATCH-HYGIENE-TELEMETRY]\n        let logs: Vec<Arc<Mutex<OutputLog>>> = {\n            let sessions = recover(&self.sessions);\n            my_subs\n                .iter()\n                .filter_map(|id| sessions.get(id).map(|h| Arc::clone(&h.log)))\n                .collect()\n        };\n        for log in &logs {\n            recover_log(log).detach_if(&send);\n        }\n        // Same for net streams: the streams + their rings persist (D4b).\n        if let Some(host) = self.net.get() {\n            for id in &my_stream_subs {\n                host.detach_stream_sub(*id, &send);\n            }\n            // ConnectionBound teardown (ADR-0040 decision 6): rows THIS conn\n            // opened as ConnectionBound die with it. FIN the stream first — the\n            // remote target's read pump sees EOF, `serve_attach` runs\n            // `detach_session` under its own serve generation (fix-6 validated:\n            // a stale opener's late close can never evict a newer controller),\n            // and the controller/viewer stamps clear — then terminal-retire the\n            // local row (eligibility + physical teardown, the W1 machinery).\n            // This is exactly the emphasys C2 hole: an rc viewport brain dying\n            // WITHOUT FIN left its attach stream restart-durable forever, the\n            // target never saw EOF, and CONTROLLED survived even a full broker\n            // restart. Durable rows (inter-brain) are untouched — brain\n            // disconnect must NEVER globally retire (brain-swap continuity).\n            // [impl->REQ-STREAM-LEASE-CLASSES]\n            for id in &my_cb_streams {\n                let _ = host.send_stream(*id, &[], true);\n                let _ = host.retire_stream_terminal(*id);\n                spt_proto::emit_line_err!(\"STREAM_CONNBOUND_RETIRE:{id}: opener conn exited — FIN + terminal retire\");\n            }\n            // And presence: the liveness log + its ring persist (D4c).\n            if my_presence_sub {\n                host.detach_presence_sub(&send);\n            }\n        }\n        Ok(())\n    }\n\n    /// Spawn a child under a fresh PTY, register it, auto-subscribe the spawning\n    /// connection, start its output drain and exit waiter, and reply `spawned`.\n    /// Every hosted session gets its **own** drain + exit-waiter threads (KH\n    /// 7.4): a hung child blocks only its own threads — shell instances ride\n    /// exactly this isolation (one owner's hung shell binary stalls neither\n    /// another owner's delivery nor the broker's control dispatch).\n    // [impl->REQ-HAZARD-DAEMON-SCHED-NONBLOCKING]\n    fn dispatch_spawn(&self, env: Envelope, send: &SharedSend) -> Result<u64, String> {\n        let req: SpawnReq =\n            serde_json::from_value(env.payload).map_err(|e| format!(\"bad spawn payload: {e}\"))?;\n        self.dispatch_spawn_policy(req, send, false)\n            .map(|id| id.expect(\"the ResumeOrReuse policy always yields a session\"))\n    }\n\n    /// The FRESH/RejectExisting policy (ADR-0041 decision 5,\n    /// REQ-SPAWN-FRESH-TRUTHFUL): same gate + spawn body as [`Self::dispatch_spawn`],\n    /// but an endpoint that already holds a live session answers the TYPED\n\n[Showing lines 6287-6353 of 16233. Use :6354 to continue]"}],"details":{"truncation":{"content":"        // [impl->REQ-SEAL-CEREMONY-ESC-CANCEL]\r\n        self.abort_ceremonies_for_conn(send.id());\r\n        // Brain gone: detach as subscriber but leave the sessions running.\r\n        // Clone the log Arcs out from under the sessions lock FIRST (ADR-0038\r\n        // Amendment fix 8, KH 7.12 kin): `detach_if` re-stamps info.json\r\n        // (`viewer_count`/`driven_by`) and emits a lifecycle record — file +\r\n        // conn I/O that must never run under the shared sessions table lock,\r\n        // or one slow perch write stalls every spawn/attach/kill dispatch.\r\n        // [impl->REQ-DISPATCH-HYGIENE-TELEMETRY]\r\n        let logs: Vec<Arc<Mutex<OutputLog>>> = {\r\n            let sessions = recover(&self.sessions);\r\n            my_subs\r\n                .iter()\r\n                .filter_map(|id| sessions.get(id).map(|h| Arc::clone(&h.log)))\r\n                .collect()\r\n        };\r\n        for log in &logs {\r\n            recover_log(log).detach_if(&send);\r\n        }\r\n        // Same for net streams: the streams + their rings persist (D4b).\r\n        if let Some(host) = self.net.get() {\r\n            for id in &my_stream_subs {\r\n                host.detach_stream_sub(*id, &send);\r\n            }\r\n            // ConnectionBound teardown (ADR-0040 decision 6): rows THIS conn\r\n            // opened as ConnectionBound die with it. FIN the stream first — the\r\n            // remote target's read pump sees EOF, `serve_attach` runs\r\n            // `detach_session` under its own serve generation (fix-6 validated:\r\n            // a stale opener's late close can never evict a newer controller),\r\n            // and the controller/viewer stamps clear — then terminal-retire the\r\n            // local row (eligibility + physical teardown, the W1 machinery).\r\n            // This is exactly the emphasys C2 hole: an rc viewport brain dying\r\n            // WITHOUT FIN left its attach stream restart-durable forever, the\r\n            // target never saw EOF, and CONTROLLED survived even a full broker\r\n            // restart. Durable rows (inter-brain) are untouched — brain\r\n            // disconnect must NEVER globally retire (brain-swap continuity).\r\n            // [impl->REQ-STREAM-LEASE-CLASSES]\r\n            for id in &my_cb_streams {\r\n                let _ = host.send_stream(*id, &[], true);\r\n                let _ = host.retire_stream_terminal(*id);\r\n                spt_proto::emit_line_err!(\"STREAM_CONNBOUND_RETIRE:{id}: opener conn exited — FIN + terminal retire\");\r\n            }\r\n            // And presence: the liveness log + its ring persist (D4c).\r\n            if my_presence_sub {\r\n                host.detach_presence_sub(&send);\r\n            }\r\n        }\r\n        Ok(())\r\n    }\r\n\r\n    /// Spawn a child under a fresh PTY, register it, auto-subscribe the spawning\r\n    /// connection, start its output drain and exit waiter, and reply `spawned`.\r\n    /// Every hosted session gets its **own** drain + exit-waiter threads (KH\r\n    /// 7.4): a hung child blocks only its own threads — shell instances ride\r\n    /// exactly this isolation (one owner's hung shell binary stalls neither\r\n    /// another owner's delivery nor the broker's control dispatch).\r\n    // [impl->REQ-HAZARD-DAEMON-SCHED-NONBLOCKING]\r\n    fn dispatch_spawn(&self, env: Envelope, send: &SharedSend) -> Result<u64, String> {\r\n        let req: SpawnReq =\r\n            serde_json::from_value(env.payload).map_err(|e| format!(\"bad spawn payload: {e}\"))?;\r\n        self.dispatch_spawn_policy(req, send, false)\r\n            .map(|id| id.expect(\"the ResumeOrReuse policy always yields a session\"))\r\n    }\r\n\r\n    /// The FRESH/RejectExisting policy (ADR-0041 decision 5,\r\n    /// REQ-SPAWN-FRESH-TRUTHFUL): same gate + spawn body as [`Self::dispatch_spawn`],\r\n    /// but an endpoint that already holds a live session answers the TYPED\r","truncated":true,"truncatedBy":"lines","totalLines":9947,"totalBytes":3834,"outputLines":67,"outputBytes":3834,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":16233,"displayContent":{"text":"    fn handle_conn(self: &Arc<Self>, conn: Stream) -> io::Result<()> {\n…\n        // [impl->REQ-SEAL-CEREMONY-ESC-CANCEL]\r\n        self.abort_ceremonies_for_conn(send.id());\r\n        // Brain gone: detach as subscriber but leave the sessions running.\r\n        // Clone the log Arcs out from under the sessions lock FIRST (ADR-0038\r\n        // Amendment fix 8, KH 7.12 kin): `detach_if` re-stamps info.json\r\n        // (`viewer_count`/`driven_by`) and emits a lifecycle record — file +\r\n        // conn I/O that must never run under the shared sessions table lock,\r\n        // or one slow perch write stalls every spawn/attach/kill dispatch.\r\n        // [impl->REQ-DISPATCH-HYGIENE-TELEMETRY]\r\n        let logs: Vec<Arc<Mutex<OutputLog>>> = {\r\n            let sessions = recover(&self.sessions);\r\n            my_subs\r\n                .iter()\r\n                .filter_map(|id| sessions.get(id).map(|h| Arc::clone(&h.log)))\r\n                .collect()\r\n        };\r\n        for log in &logs {\r\n            recover_log(log).detach_if(&send);\r\n        }\r\n        // Same for net streams: the streams + their rings persist (D4b).\r\n        if let Some(host) = self.net.get() {\r\n            for id in &my_stream_subs {\r\n                host.detach_stream_sub(*id, &send);\r\n            }\r\n            // ConnectionBound teardown (ADR-0040 decision 6): rows THIS conn\r\n            // opened as ConnectionBound die with it. FIN the stream first — the\r\n            // remote target's read pump sees EOF, `serve_attach` runs\r\n            // `detach_session` under its own serve generation (fix-6 validated:\r\n            // a stale opener's late close can never evict a newer controller),\r\n            // and the controller/viewer stamps clear — then terminal-retire the\r\n            // local row (eligibility + physical teardown, the W1 machinery).\r\n            // This is exactly the emphasys C2 hole: an rc viewport brain dying\r\n            // WITHOUT FIN left its attach stream restart-durable forever, the\r\n            // target never saw EOF, and CONTROLLED survived even a full broker\r\n            // restart. Durable rows (inter-brain) are untouched — brain\r\n            // disconnect must NEVER globally retire (brain-swap continuity).\r\n            // [impl->REQ-STREAM-LEASE-CLASSES]\r\n            for id in &my_cb_streams {\r\n                let _ = host.send_stream(*id, &[], true);\r\n                let _ = host.retire_stream_terminal(*id);\r\n                spt_proto::emit_line_err!(\"STREAM_CONNBOUND_RETIRE:{id}: opener conn exited — FIN + terminal retire\");\r\n            }\r\n            // And presence: the liveness log + its ring persist (D4c).\r\n            if my_presence_sub {\r\n                host.detach_presence_sub(&send);\r\n            }\r\n        }\r\n        Ok(())\r\n    }\r\n\r\n    /// Spawn a child under a fresh PTY, register it, auto-subscribe the spawning\r\n    /// connection, start its output drain and exit waiter, and reply `spawned`.\r\n    /// Every hosted session gets its **own** drain + exit-waiter threads (KH\r\n    /// 7.4): a hung child blocks only its own threads — shell instances ride\r\n    /// exactly this isolation (one owner's hung shell binary stalls neither\r\n    /// another owner's delivery nor the broker's control dispatch).\r\n    // [impl->REQ-HAZARD-DAEMON-SCHED-NONBLOCKING]\r\n    fn dispatch_spawn(&self, env: Envelope, send: &SharedSend) -> Result<u64, String> {\r\n        let req: SpawnReq =\r\n            serde_json::from_value(env.payload).map_err(|e| format!(\"bad spawn payload: {e}\"))?;\r\n        self.dispatch_spawn_policy(req, send, false)\r\n            .map(|id| id.expect(\"the ResumeOrReuse policy always yields a session\"))\r\n    }\r\n\r\n    /// The FRESH/RejectExisting policy (ADR-0041 decision 5,\r\n    /// REQ-SPAWN-FRESH-TRUTHFUL): same gate + spawn body as [`Self::dispatch_spawn`],\r\n    /// but an endpoint that already holds a live session answers the TYPED\r","startLine":5921,"lineNumbers":[5921,null,6287,6288,6289,6290,6291,6292,6293,6294,6295,6296,6297,6298,6299,6300,6301,6302,6303,6304,6305,6306,6307,6308,6309,6310,6311,6312,6313,6314,6315,6316,6317,6318,6319,6320,6321,6322,6323,6324,6325,6326,6327,6328,6329,6330,6331,6332,6333,6334,6335,6336,6337,6338,6339,6340,6341,6342,6343,6344,6345,6346,6347,6348,6349,6350,6351,6352,6353]},"fileSize":844842,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-daemon\\src\\broker.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":16233,"totalBytes":3834,"outputLines":67,"outputBytes":3834,"shownRange":{"start":6287,"end":6353},"nextOffset":6354}}},"isError":false,"timestamp":1788999601639}}
{"type":"message","id":"c3d8ca79","parentId":"430b7bbe","timestamp":"2026-09-10T00:20:01.662Z","message":{"role":"toolResult","toolCallId":"call_oj0v4ZgW9LyDTrJxvv5Q2Mzm|fc_04ba494197989c86016aa1f7b165a087d19e53342fcda6b0d2","toolName":"read","content":[{"type":"text","text":"pub struct SessionInfo {\n…\n    /// verdict's boot-grace input (a just-spawned harness has no client chain\n    /// YET; it must never read as a dead tree). Additive: an older broker\n    /// omits it → `0` → the grace gate keeps the verdict conservative (alive).\n    // [impl->REQ-ENDPOINT-CYCLE-HONEST]\n    #[serde(default)]\n    pub spawned_ms_ago: u64,\n}\n\n/// `sessions-reply` payload — every session the broker currently hosts.\n#[derive(Debug, Clone, Serialize, Deserialize)]\npub struct SessionsReply {\n    #[serde(default)]\n    pub sessions: Vec<SessionInfo>,\n    /// Endpoints whose bring-up the broker has STARTED but whose session row\n    /// does not exist yet (releases#199).\n    ///\n    /// The sessions list above records completed-ness; this records\n    /// started-ness, and the reserved-id bind gate needs the second fact because\n    /// its own premise is \"the completion of a bring-up the broker started\". A\n    /// harness spawned by a bring-up asks that gate before its row lands, so\n    /// without this field the broker's own child is refused its own identity.\n    ///\n    /// A LIST rather than a bool, and endpoint-keyed rather than\n    /// engine-room-shaped, so this field states a fact about the broker instead\n    /// of encoding one caller's question — it is read exactly the way `sessions`\n…\n}\n\n[Showing lines 1534-1558 of 2315. Use :1559 to continue]"}],"details":{"truncation":{"content":"    /// verdict's boot-grace input (a just-spawned harness has no client chain\r\n    /// YET; it must never read as a dead tree). Additive: an older broker\r\n    /// omits it → `0` → the grace gate keeps the verdict conservative (alive).\r\n    // [impl->REQ-ENDPOINT-CYCLE-HONEST]\r\n    #[serde(default)]\r\n    pub spawned_ms_ago: u64,\r\n}\r\n\r\n/// `sessions-reply` payload — every session the broker currently hosts.\r\n#[derive(Debug, Clone, Serialize, Deserialize)]\r\npub struct SessionsReply {\r\n    #[serde(default)]\r\n    pub sessions: Vec<SessionInfo>,\r\n    /// Endpoints whose bring-up the broker has STARTED but whose session row\r\n    /// does not exist yet (releases#199).\r\n    ///\r\n    /// The sessions list above records completed-ness; this records\r\n    /// started-ness, and the reserved-id bind gate needs the second fact because\r\n    /// its own premise is \"the completion of a bring-up the broker started\". A\r\n    /// harness spawned by a bring-up asks that gate before its row lands, so\r\n    /// without this field the broker's own child is refused its own identity.\r\n    ///\r\n    /// A LIST rather than a bool, and endpoint-keyed rather than\r\n    /// engine-room-shaped, so this field states a fact about the broker instead\r\n    /// of encoding one caller's question — it is read exactly the way `sessions`\r","truncated":true,"truncatedBy":"lines","totalLines":782,"totalBytes":1321,"outputLines":25,"outputBytes":1321,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":2315,"displayContent":{"text":"pub struct SessionInfo {\n…\n    /// verdict's boot-grace input (a just-spawned harness has no client chain\r\n    /// YET; it must never read as a dead tree). Additive: an older broker\r\n    /// omits it → `0` → the grace gate keeps the verdict conservative (alive).\r\n    // [impl->REQ-ENDPOINT-CYCLE-HONEST]\r\n    #[serde(default)]\r\n    pub spawned_ms_ago: u64,\r\n}\r\n\r\n/// `sessions-reply` payload — every session the broker currently hosts.\r\n#[derive(Debug, Clone, Serialize, Deserialize)]\r\npub struct SessionsReply {\r\n    #[serde(default)]\r\n    pub sessions: Vec<SessionInfo>,\r\n    /// Endpoints whose bring-up the broker has STARTED but whose session row\r\n    /// does not exist yet (releases#199).\r\n    ///\r\n    /// The sessions list above records completed-ness; this records\r\n    /// started-ness, and the reserved-id bind gate needs the second fact because\r\n    /// its own premise is \"the completion of a bring-up the broker started\". A\r\n    /// harness spawned by a bring-up asks that gate before its row lands, so\r\n    /// without this field the broker's own child is refused its own identity.\r\n    ///\r\n    /// A LIST rather than a bool, and endpoint-keyed rather than\r\n    /// engine-room-shaped, so this field states a fact about the broker instead\r\n    /// of encoding one caller's question — it is read exactly the way `sessions`\r\n…\n}","startLine":1480,"lineNumbers":[1480,null,1534,1535,1536,1537,1538,1539,1540,1541,1542,1543,1544,1545,1546,1547,1548,1549,1550,1551,1552,1553,1554,1555,1556,1557,1558,null,1573]},"fileSize":117548,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\286-churn\\crates\\spt-daemon\\src\\msg.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":2315,"totalBytes":1321,"outputLines":25,"outputBytes":1321,"shownRange":{"start":1534,"end":1558},"nextOffset":1559}}},"isError":false,"timestamp":1788999601660}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_cnr6bw8jk04FI0xS9Q9PlrYL|fc_04ba494197989c86016aa1f7c16de487d1a6f685c4072635da","toolName":"yield","startedAt":"2026-09-10T00:20:23.438Z"},"id":"b42d809a","parentId":"c3d8ca79","timestamp":"2026-09-10T00:20:23.439Z"}
{"type":"message","id":"6a57c38a","parentId":"b42d809a","timestamp":"2026-09-10T00:20:23.439Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":"**Checking retained-carrier tests**\n\n**Reviewing retained counts**"},{"type":"toolCall","id":"call_cnr6bw8jk04FI0xS9Q9PlrYL|fc_04ba494197989c86016aa1f7c16de487d1a6f685c4072635da","name":"yield","arguments":{"type":null,"data":{"overall_correctness":"correct","explanation":"Clean Spec-axis review: no actionable, patch-introduced defects found against issue #286 in de5a44bc…19a90979. The patch explains the five-second livehost reopen cause, retains its request/reply carrier, and discards failed queries without treating failure as an empty census (livehost.rs:1004–1047,1373,1412); the consuming KIND_SESSIONS dispatch remains reusable and subscription-free (broker.rs:6133–6251), while process-wide, mutex-serialized counters bound healthy lifecycle records and preserve exceptional attribution plus original first-write timestamps (conn.rs:178–201,319–395,649–655,746–765). The new cells exercise real framed round-trips, interval aggregation, exceptional records, and malformed-reply recovery (conn.rs:832–893; livehost.rs:1439–1486); the existing blackhole lifecycle harness was inspected, family-gate policy correctly remains unchanged, and no builds, tests, validation commands, or edits were performed.","confidence":0.92,"findings":[]},"error":null}}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":5149,"output":656,"cacheRead":100864,"cacheWrite":0,"totalTokens":106669,"reasoningTokens":400,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1788999601715,"responseId":"resp_04ba494197989c86016aa1f7b210c487d1917662f847a2d68b","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_04ba494197989c86016aa1f7b3bf7887d18a503539420d8c65","type":"reasoning","content":[],"encrypted_content":"gAAAAABqoffBzZbiujqP0U7h4EgjBDn9yEcL7gek0dBcf-amceV1SBTjVtIOQhXV0nFw5P7-tkGW6KVaHaE2cGwZATcRI92Nc-WUqtHDZKKfyMvmS_hLjxlmAGWwxpwBtmbx7Niwjl83YWJkjcyiZZsHOMlK9-cIqW2_M0Kxq-hPZYWFa3JkF96GmGMGYZpi6WRMuczyiVX6MZNMBNiXXR6HRy6ajmNTsaFctfm5PfvfmIWj7zQK7yNNZCt3V0sGYI2CudqGUD_Xk7ifaPIKWMimuDtvtQf3Ct8id7XH75XIYwYo4XdEGRi1L7ReXx3YjcoGOx0RHp299cJ4-d3rlUs4kBpdYPaHNuM7jrHC81R4ApelVQoy7CSCulNHuRsedfxGAKeMhHC_lgOYgMBP7N52R-rMdctW-VdPb-K9JPWj8Jhw3YauHZfOEbnc7OfIlgVw0RzOH1ZqU9WO3jDfWKfOwh9zzU2aXEitSeOKV5FYNa-Aaa9j_-EWiQjaB8CGarpPAlb-y_0AlkDNn1GFsqe_9XTFvdHTiARIZ1aCejsV-wczS6HMUi5-WebXi1t7DcbwezHLUduwgtzB3eBig3PPsaKBh1QRJ-nAU2mvdea5vQx6rBuhLLzvR8m6vczIUdEA4u6PfM15rT-8or7TGFTryIG01ww8X8uHAMq8Gu344KUutE-r5fPbbsuuE4FihixBK9uF5HPSWKEVQIr0c-PB4iJHPjSlYYx2lS8jhWTHQGqlO4EG522lZLioWJAl932_-MqXdR1FG_gJDhXBl_8Kqa8adjJLpFBpFk_CbFwzXve2gd_d28Bwj8yDpbr8CKaYpz_zftpUSRVX0Olj-vV2Js2vrtK8lm6Mj0x68pXPpeuIuM3EmkfIf3tH8PKNT4U5ctrEHChiKoSEpV6xPuIM4flWRo4lniILUGsidVWtk5ZqbLtlNbkLjSwltZ5c6LJq3TZi63ckuMKywJtkRqOmunOO9cgGgFmaGiq91WUKPCdQ-SPWDT1m6V7WApT720O9Vrkcdh1ntbuUBg3n-NuxUxpudZJfUYcUMaf-2BWMrVRBdHNIGyVEc2SZ3E3uupr4mek_Lozv-gSYZN3UhH44x-3_9G5cfjgfYssNWuq3pFOWA0PnFR2LnWcZTVA7cVyHyIOmnIPlFPgK84A8MlTRFCkWb4caEL8T6TuIz98BEZ38zyAf2V31pmvCnGgVaa9nIvALhMJDmkT6gMhuji8apu99K94_sQHpD5xfb7smey4_Fuetrj_eYpEmlY1VleX5DjjMuzsAquiwdmmyXq0ZmbsMSDcwjsfUu5c_ThrWn18ktWjInO8K371SoHU2BAhDVGUfgeW6V62_gPKGdmS50IUVh6BUU-YDmttJQiiyfk-APp-kD72Q-Kpqtg5-Ex2pfLNaWyIoDUyQDWj3yAO7xSb7ueAR2nNCQ_ciBZy7hsI-mFWGqj1jS4oWNOn9d87hMaFTs4g9pzWCzePBr-qWPh74SKM_kZPhZTApEFHZ7nJnvpnpFSHBqQROGiiuk8wY2LvkobOmha8VTPb92GiWKqhgbtucre4acS15RIdBMTzYD9BFprM4nW-vUNyxiW8_WyoGWN9OLNVVeB_do-QU1DIHQWWNuoOiKw3b2nxC6pJ4q3Go_ZM5NehlApMR_XXFZOuFsBXPC1Xq68QTUW_6uMpT62SA-gDnffsNTczDe40oQS58elYbC4RuYygSdLhmZ-aqUabS8isNg8iVJZaIL86SLSEDGEPZkMiPWUkeP-5vxW_1l27AwH-B7RoQIGOkyK8cOCbNic9urutwzWwudikNm58JMjLeNXyAwuKq8GkjGn3XdEHKuRuFvwC0f4yIi-vG5naGJ1kA0PXgJjwwfU4EqseSw7JCkTi9Y_5j6eZIRfHbNbU-zF9oGmxPBhzHvOh143koqtEpTLnpeJBbY_Vh3qa7LdeLK5JH1iNZ_w5DMFseczVP3QQq0afuRyeBAJVA_CRZkIsdUMvlsqqAndETToDk8xAQLCxxEtCmkZftU9wFTuRsjMCySWz9_EAz5fGyYoazQtBlNKNCt894i6xHvEZ9ZFMwAS54_VBw8Cq7caQ9-rmhKxuihYnC-WKgoE1Jxi54_fj3f-szg4AF0zxOHbUI8g-hGib-Z7agRFAobxtZDE43dbuQK3WGkmkSQZ8zGgNm_4W1G8yC0xMAV8NGTjOYl-IKNpwBNOJ8ZMq68hZmgnXyEeqwFFEltWE09MFwKXRXTVqDHf4t6PssS4VfIG3epU7SgoU3Z1qbChjZ6AmPydSe1KCJFMtNxIrJJT57pZiVRGbe12nLUuNNg7erMXvxrOlsBdhZgw0C9LdUk2wbbi_u8gaXMetg0s0ZFvqWx_6cHrJAF-TX8Vd1D2MqABtEqPIlh10lNhxSvfYkQ5eTBJoE1zrFabyLqdL3TjS81jsZ7grhMuV5dtFWqBeGBEtUvnWlOCcGF6F-qFY_KkrkwVSmvZ0_btifGaDmpg8yRhG9OWYdHKGdtgGyDLGkRuU2NwXlFI0Hz4MZOQw1MBIGLMtPInAKKGA_6c6N86i09LTGYLxUFoptbxii8_UJ7HUxQ7dfVwNK46SMQ_5Mx2viIK2xsejawkGRlbGkGJ6sBdhDIMMBxkUy3hZCE9skAmOTVBuITM2DsrPUTG3jt8wcpn84FUljc2UvEGEHaXL06IZF1JNZ5NsBErcZRKHoapd_D_8XDj5es4mVDxD2jUhwvVt6Hbf8oD8tZJoPbH0vzDL_Gp8PqgBmSZkCmVpbdX8Cc3PXj136DnMeQjpu02KgYsAhGNKdJj2Ae5CvlWFxSCOGK6r9O8192DUOLY5o0arKH3K6TmPDE8b4mOzHm1rtpTowTLREiLlNBfbo1jXJ3FvkDVoYAlQYr-K3Ct6b7NsvZbLRx1-8DC626tCkz-YRJqOoeRsCLKwbEQ-A2HdWq0k5XJSSFqG6ocR8lUw1r_zSko7Da8E3n3RE6KOcI0RJUddzFPGonPyR91F4t9ZWvHiXvVyl96LrKvTy4dwCpEKH90xGui7HFnaudNAI4k7OCcumhR9Bi3Ianu2wJE_qTjcBGkwweuRlZhM2ZkMUVaqXJwTYTKSy0rPOaY3BG5-gZhhCmoPdUbdeG1MQhOD4VSKkywiCZKrRzTyreypZ3nrCR-xEcioUm4wK3i2lRvcKxnH9yGU70Q1K9Tng4St1L68FgLk1-op8XCR6OL9PRs7yjDYFD1oCjym1OCnmZLK3XsshNmHWDttuOk4uSYyczwONr3O451SYphhBNfEbPNZpTLcvbh5sHMi9LmhgBguuhbJKnErYLrfCrPlfs2C1ywpFSqYAcqlDBUDo3wlcY9Ft9TN6mBFO_GwUinyoCrcFcc3AdUJVqIq0CdROmRwz-OoLOPcjYeaaWXI6HzBRonQegh3W6vLzh3TvskRhGNdvld3Elu9_ORC5gxDJ6KJV1iobCUdN942NQxuLDgaN-b3iKLyVQYxl35isLX4d2RroTGqL_uxpt25-yeMF7H8z1l7Oe0GMbF8wrx8SK6o4pCHOz0Ok_M6HQC8d--zpOfVnfLQ9p4Zjqn7dEMc3eyvCJtDchhIYN0aR1OVyedgMDLmOnm5s8CqKw0m2ySAcPBDKEIPLMoCPZvphY5_Mp7olm4gDsFzeoGOERYkTbqXXJ5TWKUafZawVdA-7297CIiX_U3EgNHI04yEtZ1K5-Ld5ubjG8trQUxBQVvq-5N2k4VQuDFmgxvU2zOuDpwYmy1YcML7h5rS_pg7XJ-xk-c7FMJJEKgyS23dh8vDkaqzAEs3tJY5zvmm7oAj4iB5mbEOdVAVrp5hioWn7vjNfXk0BWlq_NnJJwU-4mzLJiqqh6WGwjATNgDq912_Cu0CdkR_m60Dfv2iNc4_RNtK2F30tQAhYxRiwDr9zX0hHt_E5k_gI7o740ap3QlhP8Irpjlo-FLDZp2_tAFUEpwwtwKYh8rA-GbGsCeTB3B_gEtvqkU6GnPRStbw_cjYX5bOgSnqA0R5VcwimGCoFQY7naIQYwyzdCu0wBtMYmnkee5E7xCDj0cH3xRutXcvYBjg15KCPjKKF_64tZ8PTeBsKPitp2-6E1lhb0Y_B12hDqYhJUOaXAwTo_1bYxqemyZZwHATUYEMwalikbUFKduPgSSNXlgQbA-bbHySYA0Ou8AF_3LGZsvb3dHLGb1TrVOouzqEdyzBgfwuOketPpaiSkBz7h03gVJ6lWnFUXcaMTsHhKulKxNx7yRC7EZ7oU0wgop59Nj4k9g_6dT4A084Fddb2BSFnn8wxKDnGkJJgAPSZXKPOOU-v3e-shhHjXIy4a6gmwwoWnZIGj-38LOpphlJEZoPu","internal_chat_message_metadata_passthrough":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"summary":[{"type":"summary_text","text":"**Checking retained-carrier tests**"},{"type":"summary_text","text":"**Reviewing retained counts**"}],"metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}},{"id":"fc_04ba494197989c86016aa1f7c16de487d1a6f685c4072635da","type":"function_call","status":"completed","arguments":"{\"type\":null,\"data\":{\"overall_correctness\":\"correct\",\"explanation\":\"Clean Spec-axis review: no actionable, patch-introduced defects found against issue #286 in de5a44bc…19a90979. The patch explains the five-second livehost reopen cause, retains its request/reply carrier, and discards failed queries without treating failure as an empty census (livehost.rs:1004–1047,1373,1412); the consuming KIND_SESSIONS dispatch remains reusable and subscription-free (broker.rs:6133–6251), while process-wide, mutex-serialized counters bound healthy lifecycle records and preserve exceptional attribution plus original first-write timestamps (conn.rs:178–201,319–395,649–655,746–765). The new cells exercise real framed round-trips, interval aggregation, exceptional records, and malformed-reply recovery (conn.rs:832–893; livehost.rs:1439–1486); the existing blackhole lifecycle harness was inspected, family-gate policy correctly remains unchanged, and no builds, tests, validation commands, or edits were performed.\",\"confidence\":0.92,\"findings\":[]},\"error\":null}","call_id":"call_cnr6bw8jk04FI0xS9Q9PlrYL","internal_chat_message_metadata_passthrough":{"create_time":1788999602.330165,"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"},"name":"yield","metadata":{"turn_id":"6edd6eb1-c26a-417c-a436-18a20bb991b3"}}]},"duration":21720.808900000062,"ttft":2059.129799999995,"completedAt":1788999623437,"contextSnapshot":{"promptTokens":106013,"nonMessageTokens":13648,"compactionEpoch":0}}}
{"type":"message","id":"7a7b5542","parentId":"6a57c38a","timestamp":"2026-09-10T00:20:23.442Z","message":{"role":"toolResult","toolCallId":"call_cnr6bw8jk04FI0xS9Q9PlrYL|fc_04ba494197989c86016aa1f7c16de487d1a6f685c4072635da","toolName":"yield","content":[{"type":"text","text":"Result submitted."}],"details":{"data":{"overall_correctness":"correct","explanation":"Clean Spec-axis review: no actionable, patch-introduced defects found against issue #286 in de5a44bc…19a90979. The patch explains the five-second livehost reopen cause, retains its request/reply carrier, and discards failed queries without treating failure as an empty census (livehost.rs:1004–1047,1373,1412); the consuming KIND_SESSIONS dispatch remains reusable and subscription-free (broker.rs:6133–6251), while process-wide, mutex-serialized counters bound healthy lifecycle records and preserve exceptional attribution plus original first-write timestamps (conn.rs:178–201,319–395,649–655,746–765). The new cells exercise real framed round-trips, interval aggregation, exceptional records, and malformed-reply recovery (conn.rs:832–893; livehost.rs:1439–1486); the existing blackhole lifecycle harness was inspected, family-gate policy correctly remains unchanged, and no builds, tests, validation commands, or edits were performed.","confidence":0.92,"findings":[]},"status":"success"},"isError":false,"timestamp":1788999623440}}
{"type":"custom","customType":"session_exit","data":{"reason":"dispose","kind":"normal","recordedAt":"2026-09-10T00:27:23.453Z"},"id":"bfd7dfee","parentId":"7a7b5542","timestamp":"2026-09-10T00:27:23.453Z"}
