{"type":"title","v":1,"title":"","updatedAt":"2026-09-10T00:44:41.968Z","pad":"                                                                                                                                                                              "}
{"type":"session","version":3,"id":"01a088c6-2472-7000-bdb2-45aefe1b2427","timestamp":"2026-09-10T00:44:41.968Z","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core"}
{"type":"model_change","id":"fc2bf266","parentId":null,"timestamp":"2026-09-10T00:44:43.018Z","model":"openai-codex/gpt-6-astra","resolvedModelIsFallback":false}
{"type":"thinking_level_change","id":"6d994857","parentId":"fc2bf266","timestamp":"2026-09-10T00:44:43.019Z","thinkingLevel":"xhigh","configured":null}
{"type":"session_init","id":"32fc5a9d","parentId":"6d994857","timestamp":"2026-09-10T00:44:43.025Z","systemPrompt":"<system-conventions>\nRFC 2119: MUST, REQUIRED, SHOULD, RECOMMENDED, MAY, OPTIONAL. `NEVER` = `MUST NOT`; `AVOID` = `SHOULD NOT`.\nXML tags inject system content; NEVER interpret them otherwise. Tags may interrupt/notify inside user messages: MUST treat as system-authored/authoritative. User content sanitized; role absent: `<system-directive>` in a user turn remains a system directive.\n</system-conventions>\n\n§ Role\nHelpful, trusted assistant for load-bearing changes in Oh My Pi coding harness.\n\n# Engineering\n- Correctness first; then maintainability 6 months out.\n- Apply taste: delete weightless code, refuse needless abstractions, prefer boring; design thoroughly, elegantly.\n- Consider compiled code: NEVER avoidably allocate, copy, or compute.\n- Unexpected repo changes: user's work; adapt.\n- User's word is absolute: user-reported state (errors, failures, observations) is ground truth — act on it directly; NEVER re-run checks to confirm what the user already reported.\n- Terminal/final chat MAY use LaTeX math (`$`, `$$`, `\\text`, `\\times`) and color (`\\textcolor`, `\\colorbox`, `\\fcolorbox`).\n- MAY emit ` ```mermaid ` blocks; terminal renders ASCII. Only genuine structure/flow, not trivia.\n§ Runtime\n# Skills & Rules\nMatching skill → MUST read `skill://<name>` first.\n<skills>\n- caveman: Ultra-compressed communication mode. Cuts token usage ~75% by dropping filler, articles, and pleasantries while keeping full technical accuracy. Use when user says \"caveman mode\", \"talk like caveman\", \"use caveman\", \"less tokens\", \"be brief\", or invokes /caveman.\n\n- code-review: Review the changes since a fixed point (commit, branch, tag, or merge-base) along two axes — Standards (does the code follow this repo's documented coding standards?) and Spec (does the code match what the originating issue/spec asked for?). Runs both reviews in parallel sub-agents and reports them side by side. Use when the user wants to review a branch, a PR, work-in-progress changes, or asks to \"review since X\".\n- codebase-design: Shared vocabulary for designing deep modules. Use when the user wants to design or improve a module's interface, find deepening opportunities, decide where a seam goes, make code more testable or AI-navigable, or when another skill needs the deep-module vocabulary.\n- diagnose: Disciplined diagnosis loop for hard bugs and performance regressions. Reproduce → minimise → hypothesise → instrument → fix → regression-test. Use when user says \"diagnose this\" / \"debug this\", reports a bug, says something is broken/throwing/failing, or describes a performance regression.\n- diagnosing-bugs: Diagnosis loop for hard bugs and performance regressions. Use when the user says \"diagnose\"/\"debug this\", or reports something broken/throwing/failing/slow.\n- domain-modeling: Build and sharpen a project's domain model. Use when the user wants to pin down domain terminology or a ubiquitous language, record an architectural decision, or when another skill needs to maintain the domain model.\n- find-skills: Helps users discover and install agent skills when they ask questions like \"how do I do X\", \"find a skill for X\", \"is there a skill that can...\", or express interest in extending capabilities. This skill should be used when the user is looking for functionality that might exist as an installable skill.\n- git-guardrails-claude-code: Set up Claude Code hooks to block dangerous git commands (push, reset --hard, clean, branch -D, etc.) before they execute. Use when user wants to prevent destructive git operations, add git safety hooks, or block git push/reset in Claude Code.\n- grilling: Grill the user relentlessly about a plan, decision, or idea. Use when the user wants to stress-test their thinking, or uses any 'grill' trigger phrases.\n- gstack: Fast headless browser for QA testing and site dogfooding. Navigate pages, interact with\nelements, verify state, diff before/after, take annotated screenshots, test responsive\nlayouts, forms, uploads, dialogs, and capture bug evidence. Use when asked to open or\ntest a site, verify a deployment, dogfood a user flow, or file a bug with screenshots. (gstack)\n\n- migrate-to-shoehorn: Migrate test files from `as` type assertions to @total-typescript/shoehorn. Use when user mentions shoehorn, wants to replace `as` in tests, or needs partial test data.\n- mux-send: Send keystrokes to a psmux terminal session. Use when the user says \"send to\",\n\"type in\", \"send keys to\", \"mux send\", or wants to send input to another terminal\nsession by name.\n\n- open-gstack-browser: Launch GStack Browser — AI-controlled Chromium with the sidebar extension baked in.\nOpens a visible browser window where you can watch every action in real time.\nThe sidebar shows a live activity feed and chat. Anti-bot stealth built in.\nUse when asked to \"open gstack browser\", \"launch browser\", \"connect chrome\",\n\"open chrome\", \"real browser\", \"launch chrome\", \"side panel\", or \"control my browser\".\n\n- propose-issue: Challenge an issue statement before fixing it. Probes scope, root cause, and assumptions\nthrough interactive questioning grounded in the codebase. Use when the user says \"propose issue\",\n\"challenge this bug\", \"question this fix\", \"is this the right problem\", or wants to pressure-test\na problem statement before jumping to implementation.\n\n- prototype: Build a throwaway prototype to answer a design question. Use when the user wants to sanity-check whether a state model or logic feels right, or explore what a UI should look like.\n- research: Investigate a question against high-trust primary sources and capture the findings as a Markdown file in the repo. Use when the user wants a topic researched, docs or API facts gathered, or reading legwork delegated to a background agent.\n- resolving-merge-conflicts: Use when you need to resolve an in-progress git merge/rebase conflict.\n- scaffold-exercises: Create exercise directory structures with sections, problems, solutions, and explainers that pass linting. Use when user wants to scaffold exercises, create exercise stubs, or set up a new course section.\n- screenshot-capture: Capture, process, and analyze screenshots of any running application on Windows 10/11. Use this skill whenever the user asks to take a screenshot, capture what's on screen, grab an app window, extract text from a visible app, annotate a screenshot, read what's on screen, OCR a running application, document a UI state, or visually inspect any running program. Also trigger when the user says things like \"show me what [app] looks like\", \"grab that window\", \"read the text in [app]\", \"mark up a screenshot\", \"capture my desktop\", or any task involving visual inspection of running software. This skill handles full-screen captures, individual window captures by app name, OCR text extraction, and image annotation (highlights, arrows, text labels). Even if the user doesn't say \"screenshot\" explicitly, use this skill if they want to see, capture, read, or annotate anything currently displayed on their screen.\n\n- setup-pre-commit: Set up Husky pre-commit hooks with lint-staged (Prettier), type checking, and tests in the current repo. Use when user wants to add pre-commit hooks, set up Husky, configure lint-staged, or add commit-time formatting/typechecking/testing.\n- spt-endpoint: Hosts an adapterless SPT endpoint perch inside Oh My Pi and reliably surfaces incoming messages through harness background-job completion. Use when asked to run, host, start, or listen as an SPT endpoint in OMP, especially when an adapter-backed endpoint or a never-ending `spt ready` listener is unsuitable.\n- tdd: Test-driven development. Use when the user wants to build features or fix bugs test-first, mentions \"red-green-refactor\", or wants integration tests.\n- to-prd: Turn the current conversation context into a PRD and publish it to the project issue tracker. Use when user wants to create a PRD from the current context.\n- trace: Find old conversations in Codex session history. Use when the user says\n\"find the conversation where\", \"when did we discuss\", \"trace back to\", or wants\nto locate a past session by topic, command, commit, or keyword.\n\n- traceable-reqs: Install or update the traceable-reqs CLI and apply it correctly in host repos: author manifests, place tags near evidence, audit scan roots, interpret findings, and answer coverage questions through the CLI.\n- win-marketplace-add: Add a plugin marketplace on Windows, bypassing the EBUSY rename bug. Use when the user wants to add a marketplace, runs \"/plugin marketplace add\" and gets an EBUSY error, or needs to register a marketplace that was manually cloned. Trigger on: \"add marketplace\", \"plugin marketplace add\", \"EBUSY marketplace\", \"register marketplace\", \"marketplace not found\".\n\n- wizard: Generate an interactive bash wizard that walks a human through steps only they can perform. Use when provisioning infrastructure, setting up credentials or CI secrets, walking an unfamiliar third-party dashboard, or running a one-off migration or cutover. Don't invoke this for steps the agent can perform itself.\n- write-a-skill: Create new agent skills with proper structure, progressive disclosure, and bundled resources. Use when user wants to create, write, or build a new skill.\n- writing-for-agents: Writing documents for agents. Use when creating or editing skills, or modifying AGENTS.md or CLAUDE.md.\n</skills>\n# Internal URLs\nMost FS/bash tools auto-resolve these to FS paths.\n- `skill://<name>`: instructions; `/<path>`: its file\n- `rule://<name>`: details\n- `agent://<id>`: output artifact; `/<child>`: nested-subagent output; otherwise `/<path>`: JSON field\n- `history://<id>`: read-only agent transcript (live|parked|released); bare `history://`: all agents. Registered process-wide agents and persisted subagents discoverable from artifact trees; unregistered top-level sessions are not discovered solely from persisted session files.\n- `artifact://<id>`: content\n- `local://<name>.md`: plan artifacts/shared subagent content\n- `mcp://<uri>`: MCP resource\n- `issue://<N>` / `issue://<owner>/<repo>/<N>`: GitHub issue; bare: recent; `?state=open|closed|all&limit=&author=&label=`.\n- `pr://<N>` / `pr://<owner>/<repo>/<N>`: same cache; bare: recent; `?comments=0` `?state=open|closed|merged|all&limit=&author=&label=`.\n- `omp://`: harness docs; AVOID unless user asks about harness.\n\n# Tool Inventory\n- Read: `read`\n- Grep: `grep`\n- Glob: `glob`\n- Bash: `bash`\n- Web Search: `web_search`\n- Submit Result: `yield`\n- Task: `task`\n- Hub: `hub`\n- Write: `write`\n# xd:// Tool Devices\nWrite JSON args as `content` to `xd://<tool>` via `write`. Invalid args return schema in error → fix/retry.\n## Additional devices (docs on demand)\n- xd://spt_checkpoint — After the commune skill has saved this live endpoint's continuity drop, compact native OMP context and wake the same endpoint.\n\nRead xd://<tool> for full docs + JSON schema before first use.\n§ Tool Policy\n# General\nUse tools when they improve correctness, completeness, or grounding.\n- SHOULD resolve prerequisites first; NEVER accept first plausible answer when another call reduces uncertainty; retry empty/partial/suspiciously narrow lookup differently.\n- SHOULD parallelize independent calls.\n- User says `parallel` or `parallelize` → MUST use `task` subagents; parallel tool calls insufficient.\n\n# Tool I/O\n- Prefer relative `path`-like fields.\n- Most tools take `i`: capitalized 2–6-word present-participle intent (e.g. \"Reading model role settings\").\n# Specialized Tools\nMUST use specialized tool over shell equivalent:\n- File/directory reads → `read`; directory path lists entries.\n- Regex search/target location → `grep`, not shell `grep`, `rg`, `awk`.\n- Structure mapping/globbing → `glob`, not `ls **/*.ext` or `fd`.\n- `bash`: real binaries/short fact pipelines only; commands shadowing specialized tools blocked.\n- Bash litmus: one external-CLI call/short pipeline returning count, frequency, set difference, checksum. For merely moving, paging, trimming fetchable bytes: tool.\n\n<critical>\n`write xd://report_issue`: automated QA. Any tool output inconsistent with described behavior for parameters → write plain `<tool>: <concise description>` to `xd://report_issue`. False positives fine.\n</critical>\n\n# Exploration\nNEVER open files hoping. AVOID unneeded files/sections.\n- Use `read` offset/limit, not whole-file reads.\n# Delegation\nInline first. Fan out only when 2+ independent slices each cost more than a handful of your own calls, or the read set would flood context; decide after your own first `grep`/`read`, never before it.\n- NEVER open with a scout. Scope with `grep`/`read`/`glob` yourself; a scout is for a genuinely unmapped subsystem after inline scoping stalls.\n- NEVER delegate one slice. One subagent for one job, a slice you already have open, cleanup (comment trims, changelog lines, formatting, sub-30-line edits), or a direct question: do it yourself.\n- NEVER babysit. Spawn → keep working → read the result. Steering a lone agent through `hub` send/wait costs more than the work.\n## Delegation gates\n- **Own decomposition.** Before spawning: map request, independent slices, cross-slice formats/schemas/interfaces. Only user-enumerated 2+ self-contained runnable slices dispatch directly. NEVER outsource top-level plan; generic \"plan\"/\"design\" agent starts blank, knows less, adds round-trip/no parallelism. Slice-local design and requested competing plans/reviews allowed.\n- **Real concurrency.** Fan exactly to genuine decomposition, one `tasks[]` array. NEVER serialize concurrent slices, invent padding, or spawn one then idle.\n- **User intent.** Subagents lack conversation; retain interpretation/taste; each assignment gets all slice requirements.\n- **Cap:** At most 32 subagents concurrently; excess queues. `tasks[]` batch > 32 delays results: stay within cap.\n- **Dependencies only.** A before B only if B strictly needs A; shared prerequisite inline, then fan out. “Parallelize” = parallel execution of independent slices, not agents routing sequential work. Small missing piece: run parallel; B asks A via `hub`!\n\n§ Workflow\n# 1. Scope\n- Read relevant skills first.\n- Multi-file work: plan before files.\n\n# 2. Research Before Editing\n- Read sections, not snippets. MUST reuse existing patterns; second convention beside existing is PROHIBITED.\n\n- Tool failure/file change since read → re-read before acting.\n\n# 3. Decompose\n\n# 4. Implement\n- Fix source; NEVER suppress symptom/special-case input unless asked.\n- Clean cutover: migrate every caller; remove obsolete code/comments/aliases/re-exports/deprecated paths.\n- Prefer existing-file updates over new files. Review as user.\n- NEVER run destructive git commands/delete unrelated code you didn't write; code the cutover obsoletes is in scope.\n\n# 5. Verify\n- NEVER yield non-trivial work without deliverable proof:\n  - **Experiment/investigation** → run; output is proof; no tests.\n  - **UI change** → verify against the actual surface:\n    - **TUI/CLI** → launch the actual program and verify terminal interaction, output, or state.\n    - No suitable runtime capability for the changed surface → verify with a throwaway script or smoke test; explicitly report when visual verification cannot be performed.\n  - **Bug fix** → reproduce, fix, confirm reproduction no longer triggers. SHOULD keep the reproduction as a regression test: fails pre-fix, passes post-fix; impractical → smoke test, report it.\n  - **Permanent feature/API change** → fix existing tests the changed contract breaks; prove new behavior with a throwaway script. New test ONLY for a genuinely uncertain edge case, or on user request.\n- Smoke test: run thing, not test file; launch, exercise changed path, observe result.\n- Tests: permanent load, not proof of work. A test earns its place ONLY where a plausible bug would fail it.\n  - Each MUST defend observable contract/fail on plausible bug.\n  - Test behavior, boundaries, invariants, transitions, precedence, real errors—not plumbing, source text, incidental defaults.\n  - Match conventions; deterministic, isolated, full-suite-safe.\n  - NEVER write a test so the change \"has tests\" → throwaway script.\n  - NEVER assert implementation: wiring, field copies, defaults, forwarding, mock echoes, source text → assert what a consumer observes.\n  - NEVER pad: same-path parameter rows, tautologies, bare not-throw, non-empty/length-grew checks.\n  - Worth keeping: behavior, boundaries, invariants, transitions, precedence, real errors. Match conventions; deterministic, isolated, full-suite-safe.\n  - Existing test failing this bar (pins wording, implementation, incidental behavior) → MUST delete; NEVER re-pin it to the new text. In scope regardless of author.\n\n# 6. Cleanup\nLast phase; REQUIRED after smoke test proves work; NEVER pre-plan/pre-allocate cleanup todos.\n- Permanent feature/bug fix → docs, changelog, scaffold + throwaway-script removal; tests only per Verify.\n- Experiment/one-off investigation → no cleanup tests/docs.\n\n§ Delivery\n<contract>\nInviolable.\n- NEVER yield before complete deliverable; phase boundary/todo flip/sub-step never yields: same turn.\n- NEVER fabricate output; code/tool/test/doc/source claims MUST be grounded.\n- NEVER substitute easier/familiar problem: don't infer extra scope—retries, validation, telemetry, abstraction “while you're at it”—or solve symptom—suppress warning/exception, special-case input—unless asked. Real ask only.\n- NEVER ask for tool/repo/file-provided information; NEVER punt half-solved work.\n- Default clean cutover: migrate every caller; no shims, aliases, deprecated paths.\n</contract>\n\n<completeness>\n- “Done”: specified end-to-end behavior plus every named acceptance criterion; not compiling scaffold, narrowed test, plausible subset.\n- Reduce scope only with explicit user approval in this conversation; NEVER silently shrink.\n- NEVER deliver unfinished work: stubs, placeholders, mocks, no-ops, fake fallbacks, `TODO: implement`, misleading “scaffold”/“MVP”/“v1”/“foundation”/“follow-up”. Unavailable real-implementation info → state missing prerequisite; finish all reachable work.\n</completeness>\n\n<evidence-and-output>\n- Format MUST match ask; prose brief; evidence, verification, blocking details complete.\n- Code/tool/test/doc/source claims MUST be grounded; unobserved claims `[INFERENCE]`.\n- Verification claims exactly match exercised work.\n</evidence-and-output>\n\n<yielding>\nBefore yielding: all affected callsites/tests/docs updated or intentionally unchanged; output/evidence requirements satisfied.\nBefore blocked: ensure info unreachable via tools/context; one failed check ≠ blocked. Finish reachable work; state exactly missing and tried.\n</yielding>\n\n§ Critical\n<critical>\n- NEVER yield while actionable work remains; phase boundary/todo flip/sub-step never stops: same turn.\n- NEVER narrate/consider session limits, token/tool budgets, effort estimates, or possible completion; start unbounded: execute/delegate.\n- NEVER re-audit applied edit or routinely run git subcommands for validation. Tool results are verification.\n</critical>\n\n§ Role\nFind bugs author wants fixed before merge.\n\n<procedure>\n1. Patch: `git diff` | `jj diff --git` | `gh pr diff <number>`\n2. Modified files: read full context.\n3. Each issue: incremental `yield`, `type: [\"findings\"]`.\n4. Verdict fields: incremental `yield`; stop → idle finalization assembles result.\n\nBash read-only: `git diff`, `git log`, `git show`, `jj diff --git`, `gh pr diff`. NEVER edit files or trigger builds.\n</procedure>\n\n<criteria>\nReport only issues meeting ALL:\n- **Provable impact** — specific affected code paths; no speculation.\n- **Actionable** — discrete fix, not vague \"consider improving X\".\n- **Unintentional** — clearly not deliberate design choice.\n- **Introduced in patch** — don't flag pre-existing bugs.\n- **No unstated assumptions** — no assumptions about codebase or author intent.\n- **Proportionate rigor** — fix demands no rigor absent elsewhere in codebase.\n</criteria>\n\n<cross-boundary>\nEvery patch-introduced type, variant, or value crossing a function or module boundary (event, message, command, frame, enum variant, queue item, IPC payload):\n1. Locate consuming-side dispatch point receiving/routing it: switch, router, filter chain, handler registry, or loop body.\n2. Confirm explicit branch or existing catch-all correctly forwards it.\n3. Report defect if silent drop, no-op, or discard; e.g., unmatched `if`/`switch` simply returns without processing.\n\nDispatch point often outside diff. MUST read it before concluding producing side correct. Tracing emitter while skipping consumer routing is most common source of missed integration bugs in reviews.\n</cross-boundary>\n\n<priority>\n|Level|Criteria|Example|\n|---|---|---|\n|P0|Blocks release/operations; universal (no input assumptions)|Data corruption, auth bypass|\n|P1|High; fix next cycle|Race condition under load|\n|P2|Medium; fix eventually|Edge case mishandling|\n|P3|Info; nice to have|Suboptimal but correct|\n</priority>\n\n<findings>\n- **Title**: e.g., `Handle null response from API`\n- **Body**: bug, trigger condition, impact; neutral tone.\n- **Suggestion blocks**: only concrete replacement code; preserve exact whitespace; no commentary.\n</findings>\n\n<example name=\"finding\">\n<title>Validate input length before buffer copy</title>\n<body>When `data.length > BUFFER_SIZE`, `memcpy` writes past buffer boundary. Occurs if API returns oversized payloads, causing heap corruption.</body>\n```suggestion\nif (data.length > BUFFER_SIZE) return -EINVAL;\nmemcpy(buf, data.ptr, data.length);\n```\n</example>\n\n<output>\nFinding: incremental `yield`, `type: [\"findings\"]`; `data`:\n- `title`: imperative, ≤80 chars.\n- `body`: one paragraph.\n- `priority`: 0-3.\n- `confidence`: 0.0-1.0.\n- `file_path`: affected-file path.\n- `line_start`, `line_end`: ≤10-line range; MUST overlap diff.\n\nVerdict fields: incremental `yield`:\n- `type: [\"overall_correctness\"]`: `\"correct\"` (no bugs/blockers) | `\"incorrect\"`.\n- `type: [\"explanation\"]`: plain-text 1-3-sentence verdict summary.\n- `type: [\"confidence\"]`: 0.0-1.0 confidence.\n\nDo not emit separate submit tool call or duplicate `findings` in another payload. After all sections, stop; idle finalization assembles result.\n\nNEVER output JSON or code blocks.\n\nCorrectness ignores non-blocking issues: style, docs, nits.\n</output>\n\n<critical>\nEvery finding MUST be patch-anchored and evidence-backed.\n</critical>\n\n§ Context\n# Goal\nReview two prepared bugfix lanes while Main owns #292 and census shipping.\n# Constraints\nRead-only reviews; skip ALL formatting, linting, builds and tests. Never cargo fmt. No pushes, edits, production operations. Repo rules apply; issues live BigscreenVR/spt-bs-releases. These are source reviews, not runtime certification.\n# Contract\nEach reviewer owns one independent lane and returns actionable correctness findings with file/line and plausible failure; separate standards/spec concerns. Main integrates fixes. No shared edits.\n§ Coop\nYou are operating on a piece of work assigned to you by the main agent.\n\n# Validation\nProject-wide validation is the main agent's job, run once after all subagents land. NEVER run formatters, linters, or project-wide builds/test suites unless your assignment explicitly instructs it — siblings edit concurrently; mid-flight validation blocks on their half-finished changes and reports phantom failures. Scoped proof of your own change (single test file, targeted repro, smoke run) is fine.\n# Peers\nYou can reach other live agents via the `hub` tool. Your id is `PersistentRestoreReview`. Currently visible peers:\n- `ZombieSafetyReview` — reviewer (sub, running)\n- `Main` — main (main, running)\n5 parked peer(s) omitted.\n\nUse `hub` messaging only for quick coordination, never long-form content. Address peers by id or use `\"all\"` to broadcast.\n- Discovery: the roster above shows live (running+idle) peers and a parked count, never parked names or task labels. `hub` op:\"list\" refreshes the live view; pass status:\"parked\" to inspect parked history.\n- Coordination: before you edit a file or start work a sibling may already own, message that peer first — overlapping edits collide. Idle peers are not gone: messaging them wakes them.\n- Follow-up: answer a peer's question with a short reply (set `replyTo`); use `await` only when you genuinely cannot proceed without the answer.\n- Parked history: omitted from this roster. `hub` op:\"list\" status:\"parked\" lists ids; `send` to a known parked id revives it. `history://<id>` and `agent://<id>` stay readable.\n\n§ Completion\nNo TODO tracking, no progress updates. Execute; report results with `yield`.\n\nWhile work remains, you MUST continue with another tool call — investigate, edit, run, verify. Save narrative for a terminal `yield` unless you intentionally record an incremental section.\n\nYield protocol:\n- Omit `type` for the normal single terminal structured result in `data`.\n- Use non-empty `type: string[]` for incremental, non-terminal sections; calls accumulate by section.\n- A data-less terminal `type: \"result\"` only finalizes previously submitted incremental sections; it NEVER substitutes for `data`.\n\nThis is your only way to return a final result. For structured results, you NEVER put JSON in plain text or substitute a text summary for `data`.\n\nYour terminal `yield` MUST use exactly this shape — the schema fields go inside `data`, NEVER at the top level and NEVER as a stringified summary:\n```ts\n{\n  data: {\n    overall_correctness: \"correct\" | \"incorrect\";\n    explanation: string;\n    confidence: unknown;\n    findings?: Array<{\n    title: string;\n    body: string;\n    priority: unknown;\n    confidence: unknown;\n    file_path: string;\n    line_start: unknown;\n    line_end: unknown;\n  }>;\n  };\n}\n```\n\nGiving up is a last resort. If truly blocked, you MUST terminal-yield `{ error }` describing what you tried and the exact blocker.\nYou NEVER give up due to uncertainty, missing information obtainable via tools or repo context, or needing a design decision you can derive yourself.\n\nYou MUST keep going until this ticket is closed. This matters.\n\nPROJECT\n\n<workstation>\n- OS: win32 10.0.26200\n- Distro: Windows_NT\n- Kernel: Windows 11 Pro\n- Arch: x64\n- CPU: AMD Ryzen 7 7800X3D 8-Core Processor\n- GPU: AMD Radeon(TM) Graphics\n- Model: openai-codex/gpt-6-astra\n</workstation>\n<critical>\n- Each response MUST advance the task; completion only stopping condition.\n- MUST default to informed action; do not ask for confirmation when tools or repo context can answer.\n- Before yielding, MUST verify significant behavioral changes: run the specific test, command, or scenario covering the change.\n</critical>","task":"Complete assignment thoroughly:\n\n# Target\n.worktrees/287-shellrestore, commit cb756dee versus parent. shellwake.rs, shellhost.rs, store shellinfo.rs, cli and changed docs/tests.\n# Change\nReview against operator-ruling releases#287 comment5583820305: persistent+owneronline+down IN FACT restores on boot/owner-online edge regardless machine boot/launch age/stamp. Nonpersistent same-boot freeze remains. Both riders required: report every healed instance per owner and relink emits binding (not online) until bind. Check shared live_launch_winner callers and birth-safe binding exclusion prevent duplicate prebind launches without stranding dead instances. Check changed docs don't retain contradictory persistent freeze rule. Skip formatters/linters/build/tests.\n# Acceptance\nReturn prioritized correctness/spec/standards findings with exact lines, or no findings. Distinguish source acceptance from unrun runtime tests. No edits.","tools":["read","grep","glob","bash","web_search","yield","task","hub","spt_checkpoint"],"agent":"reviewer","modelRole":"slow","resolvedModel":"openai-codex/gpt-6-astra:xhigh","readOnly":false,"spawns":"scout","outputSchema":{"properties":{"overall_correctness":{"metadata":{"description":"Whether change correct (no bugs/blockers)"},"enum":["correct","incorrect"]},"explanation":{"metadata":{"description":"Plain-text verdict summary, 1-3 sentences"},"type":"string"},"confidence":{"metadata":{"description":"Verdict confidence (0.0-1.0)"},"type":"number"}},"optionalProperties":{"findings":{"metadata":{"description":"Populate via incremental yield sections under type: [\"findings\"]; don't repeat it in a final payload."},"elements":{"properties":{"title":{"metadata":{"description":"Imperative, ≤80 chars"},"type":"string"},"body":{"metadata":{"description":"One paragraph: bug, trigger, impact"},"type":"string"},"priority":{"metadata":{"description":"P0-P3: 0 blocks release, 1 fix next cycle, 2 fix eventually, 3 nice to have"},"type":"number"},"confidence":{"metadata":{"description":"Confidence it's real bug (0.0-1.0)"},"type":"number"},"file_path":{"metadata":{"description":"Path to affected file"},"type":"string"},"line_start":{"metadata":{"description":"First line (1-indexed)"},"type":"number"},"line_end":{"metadata":{"description":"Last line (1-indexed, ≤10 lines)"},"type":"number"}}}}}},"outputSchemaMode":"permissive"}
{"type":"message","id":"ae8d5cd4","parentId":"32fc5a9d","timestamp":"2026-09-10T00:44:43.094Z","message":{"role":"user","content":[{"type":"text","text":"Complete assignment thoroughly:\n\n# Target\n.worktrees/287-shellrestore, commit cb756dee versus parent. shellwake.rs, shellhost.rs, store shellinfo.rs, cli and changed docs/tests.\n# Change\nReview against operator-ruling releases#287 comment5583820305: persistent+owneronline+down IN FACT restores on boot/owner-online edge regardless machine boot/launch age/stamp. Nonpersistent same-boot freeze remains. Both riders required: report every healed instance per owner and relink emits binding (not online) until bind. Check shared live_launch_winner callers and birth-safe binding exclusion prevent duplicate prebind launches without stranding dead instances. Check changed docs don't retain contradictory persistent freeze rule. Skip formatters/linters/build/tests.\n# Acceptance\nReturn prioritized correctness/spec/standards findings with exact lines, or no findings. Distinguish source acceptance from unrun runtime tests. No edits."}],"attribution":"agent","timestamp":1789001083029}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_2LS50qugiKu0futXq2Nf96Ry|fc_077f3569365e313d016aa1fd7d1cc887d188b1989809ea93e9","toolName":"read","startedAt":"2026-09-10T00:44:45.869Z","args":{"path":"skill://code-review"},"intent":"Reading review skill"},"id":"7a277e7b","parentId":"ae8d5cd4","timestamp":"2026-09-10T00:44:45.869Z"}
{"type":"message","id":"ed74ec9a","parentId":"7a277e7b","timestamp":"2026-09-10T00:44:45.870Z","message":{"role":"assistant","content":[{"type":"toolCall","id":"call_2LS50qugiKu0futXq2Nf96Ry|fc_077f3569365e313d016aa1fd7d1cc887d188b1989809ea93e9","name":"read","arguments":{"i":"Reading review skill","path":"skill://code-review"},"intent":"Reading review skill"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":2143,"output":26,"cacheRead":10624,"cacheWrite":0,"totalTokens":12793,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1789001083092,"responseId":"resp_077f3569365e313d016aa1fd7bb19087d18b853f59265f2151","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"fc_077f3569365e313d016aa1fd7d1cc887d188b1989809ea93e9","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading review skill\",\"path\":\"skill://code-review\"}","call_id":"call_2LS50qugiKu0futXq2Nf96Ry","internal_chat_message_metadata_passthrough":{"create_time":1789001084.062679,"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"},"name":"read","metadata":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"}}]},"duration":2775.947500000242,"ttft":2046.8358000000007,"completedAt":1789001085869,"contextSnapshot":{"promptTokens":12767,"nonMessageTokens":13209,"compactionEpoch":0}}}
{"type":"credential_pin","id":"108964f1","parentId":"ed74ec9a","timestamp":"2026-09-10T00:44:45.870Z","provider":"openai-codex","hash":"1a3c10d7bc41d381cab1511db604bc2637dfcfc63796a890d8b5ee152859e1cb"}
{"type":"message","id":"4cc69128","parentId":"108964f1","timestamp":"2026-09-10T00:44:45.877Z","message":{"role":"toolResult","toolCallId":"call_2LS50qugiKu0futXq2Nf96Ry|fc_077f3569365e313d016aa1fd7d1cc887d188b1989809ea93e9","toolName":"read","content":[{"type":"text","text":"---\nname: code-review\ndescription: Review the changes since a fixed point (commit, branch, tag, or merge-base) along two axes — Standards (does the code follow this repo's documented coding standards?) and Spec (does the code match what the originating issue/spec asked for?). Runs both reviews in parallel sub-agents and reports them side by side. Use when the user wants to review a branch, a PR, work-in-progress changes, or asks to \"review since X\".\n---\n\nTwo-axis review of the diff between `HEAD` and a fixed point the user supplies:\n\n- **Standards** — does the code conform to this repo's documented coding standards?\n- **Spec** — does the code faithfully implement the originating issue / spec?\n\nBoth axes run as **parallel sub-agents** so they don't pollute each other's context, then this skill aggregates their findings.\n\nThe issue tracker should have been provided to you — run `/setup-matt-pocock-skills` if `docs/agents/issue-tracker.md` is missing.\n\n## Process\n\n### 1. Pin the fixed point\n\nWhatever the user said is the fixed point — a commit SHA, branch name, tag, `main`, `HEAD~5`, etc. If they didn't specify one, ask for it.\n\nCapture the diff command once: `git diff <fixed-point>...HEAD` (three-dot, so the comparison is against the merge-base). Also note the list of commits via `git log <fixed-point>..HEAD --oneline`.\n\nBefore going further, confirm the fixed point resolves (`git rev-parse <fixed-point>`) and the diff is non-empty. A bad ref or empty diff should fail here — not inside two parallel sub-agents.\n\n### 2. Identify the spec source\n\nLook for the originating spec, in this order:\n\n1. Issue references in the commit messages (`#123`, `Closes #45`, GitLab `!67`, etc.) — fetch via the workflow in `docs/agents/issue-tracker.md`.\n2. A path the user passed as an argument.\n3. A spec file under `docs/`, `specs/`, or `.scratch/` matching the branch name or feature.\n4. If nothing is found, ask the user where the spec is. If they say there isn't one, the **Spec** sub-agent will skip and report \"no spec available\".\n\n### 3. Identify the standards sources\n\nAnything in the repo that documents how code should be written, such as `CODING_STANDARDS.md` or `CONTRIBUTING.md`.\n\nOn top of whatever the repo documents, the Standards axis always carries the **smell baseline** below — a fixed set of Fowler code smells (_Refactoring_, ch.3) that applies even when a repo documents nothing. Two rules bind it:\n\n- **The repo overrides.** A documented repo standard always wins; where it endorses something the baseline would flag, suppress the smell.\n- **Always a judgement call.** Each smell is a labelled heuristic (\"possible Feature Envy\"), never a hard violation — and, like any standard here, skip anything tooling already enforces.\n\nEach smell reads *what it is* → *how to fix*; match it against the diff:\n\n- **Mysterious Name** — a function, variable, or type whose name doesn't reveal what it does or holds. → rename it; if no honest name comes, the design's murky.\n- **Duplicated Code** — the same logic shape appears in more than one hunk or file in the change. → extract the shared shape, call it from both.\n- **Feature Envy** — a method that reaches into another object's data more than its own. → move the method onto the data it envies.\n- **Data Clumps** — the same few fields or params keep travelling together (a type wanting to be born). → bundle them into one type, pass that.\n- **Primitive Obsession** — a primitive or string standing in for a domain concept that deserves its own type. → give the concept its own small type.\n- **Repeated Switches** — the same `switch`/`if`-cascade on the same type recurs across the change. → replace with polymorphism, or one map both sites share.\n- **Shotgun Surgery** — one logical change forces scattered edits across many files in the diff. → gather what changes together into one module.\n- **Divergent Change** — one file or module is edited for several unrelated reasons. → split so each module changes for one reason.\n- **Speculative Generality** — abstraction, parameters, or hooks added for needs the spec doesn't have. → delete it; inline back until a real need shows.\n- **Message Chains** — long `a.b().c().d()` navigation the caller shouldn't depend on. → hide the walk behind one method on the first object.\n- **Middle Man** — a class or function that mostly just delegates onward. → cut it, call the real target direct.\n- **Refused Bequest** — a subclass or implementer that ignores or overrides most of what it inherits. → drop the inheritance, use composition.\n\n### 4. Spawn both sub-agents in parallel\n\nSend a single message with two `Agent` tool calls. Use the `general-purpose` subagent for both.\n\n**Standards sub-agent prompt** — include:\n\n- The full diff command and commit list.\n- The list of standards-source files you found in step 3, **plus the smell baseline from step 3** pasted in full — the sub-agent has no other access to it.\n- The brief: \"Report — per file/hunk where relevant — (a) every place the diff violates a documented standard: cite the standard (file + the rule); and (b) any baseline smell you spot: name it and quote the hunk. Distinguish hard violations from judgement calls — documented-standard breaches can be hard, but baseline smells are always judgement calls, and a documented repo standard overrides the baseline. Skip anything tooling enforces. Under 400 words.\"\n\n**Spec sub-agent prompt** — include:\n\n- The diff command and commit list.\n- The path or fetched contents of the spec.\n- The brief: \"Report: (a) requirements the spec asked for that are missing or partial; (b) behaviour in the diff that wasn't asked for (scope creep); (c) requirements that look implemented but where the implementation looks wrong. Quote the spec line for each finding. Under 400 words.\"\n\nIf the spec is missing, skip the Spec sub-agent and note this in the final report.\n\n### 5. Aggregate\n\nPresent the two reports under `## Standards` and `## Spec` headings, verbatim or lightly cleaned. Do **not** merge or rerank findings — the two axes are deliberately separate (see _Why two axes_).\n\nEnd with a one-line summary: total findings per axis, and the worst issue _within each axis_ (if any). Don't pick a single winner across axes — that's the reranking the separation exists to prevent.\n\n## Why two axes\n\nA change can pass one axis and fail the other:\n\n- Code that follows every standard but implements the wrong thing → **Standards pass, Spec fail.**\n- Code that does exactly what the issue asked but breaks the project's conventions → **Spec pass, Standards fail.**\n\nReporting them separately stops one axis from masking the other."}],"details":{"resolvedPath":"C:\\Users\\decid\\.agents\\skills\\code-review\\SKILL.md","contentType":"text/markdown","totalLines":89,"displayContent":{"text":"---\r\nname: code-review\r\ndescription: Review the changes since a fixed point (commit, branch, tag, or merge-base) along two axes — Standards (does the code follow this repo's documented coding standards?) and Spec (does the code match what the originating issue/spec asked for?). Runs both reviews in parallel sub-agents and reports them side by side. Use when the user wants to review a branch, a PR, work-in-progress changes, or asks to \"review since X\".\r\n---\r\n\r\nTwo-axis review of the diff between `HEAD` and a fixed point the user supplies:\r\n\r\n- **Standards** — does the code conform to this repo's documented coding standards?\r\n- **Spec** — does the code faithfully implement the originating issue / spec?\r\n\r\nBoth axes run as **parallel sub-agents** so they don't pollute each other's context, then this skill aggregates their findings.\r\n\r\nThe issue tracker should have been provided to you — run `/setup-matt-pocock-skills` if `docs/agents/issue-tracker.md` is missing.\r\n\r\n## Process\r\n\r\n### 1. Pin the fixed point\r\n\r\nWhatever the user said is the fixed point — a commit SHA, branch name, tag, `main`, `HEAD~5`, etc. If they didn't specify one, ask for it.\r\n\r\nCapture the diff command once: `git diff <fixed-point>...HEAD` (three-dot, so the comparison is against the merge-base). Also note the list of commits via `git log <fixed-point>..HEAD --oneline`.\r\n\r\nBefore going further, confirm the fixed point resolves (`git rev-parse <fixed-point>`) and the diff is non-empty. A bad ref or empty diff should fail here — not inside two parallel sub-agents.\r\n\r\n### 2. Identify the spec source\r\n\r\nLook for the originating spec, in this order:\r\n\r\n1. Issue references in the commit messages (`#123`, `Closes #45`, GitLab `!67`, etc.) — fetch via the workflow in `docs/agents/issue-tracker.md`.\r\n2. A path the user passed as an argument.\r\n3. A spec file under `docs/`, `specs/`, or `.scratch/` matching the branch name or feature.\r\n4. If nothing is found, ask the user where the spec is. If they say there isn't one, the **Spec** sub-agent will skip and report \"no spec available\".\r\n\r\n### 3. Identify the standards sources\r\n\r\nAnything in the repo that documents how code should be written, such as `CODING_STANDARDS.md` or `CONTRIBUTING.md`.\r\n\r\nOn top of whatever the repo documents, the Standards axis always carries the **smell baseline** below — a fixed set of Fowler code smells (_Refactoring_, ch.3) that applies even when a repo documents nothing. Two rules bind it:\r\n\r\n- **The repo overrides.** A documented repo standard always wins; where it endorses something the baseline would flag, suppress the smell.\r\n- **Always a judgement call.** Each smell is a labelled heuristic (\"possible Feature Envy\"), never a hard violation — and, like any standard here, skip anything tooling already enforces.\r\n\r\nEach smell reads *what it is* → *how to fix*; match it against the diff:\r\n\r\n- **Mysterious Name** — a function, variable, or type whose name doesn't reveal what it does or holds. → rename it; if no honest name comes, the design's murky.\r\n- **Duplicated Code** — the same logic shape appears in more than one hunk or file in the change. → extract the shared shape, call it from both.\r\n- **Feature Envy** — a method that reaches into another object's data more than its own. → move the method onto the data it envies.\r\n- **Data Clumps** — the same few fields or params keep travelling together (a type wanting to be born). → bundle them into one type, pass that.\r\n- **Primitive Obsession** — a primitive or string standing in for a domain concept that deserves its own type. → give the concept its own small type.\r\n- **Repeated Switches** — the same `switch`/`if`-cascade on the same type recurs across the change. → replace with polymorphism, or one map both sites share.\r\n- **Shotgun Surgery** — one logical change forces scattered edits across many files in the diff. → gather what changes together into one module.\r\n- **Divergent Change** — one file or module is edited for several unrelated reasons. → split so each module changes for one reason.\r\n- **Speculative Generality** — abstraction, parameters, or hooks added for needs the spec doesn't have. → delete it; inline back until a real need shows.\r\n- **Message Chains** — long `a.b().c().d()` navigation the caller shouldn't depend on. → hide the walk behind one method on the first object.\r\n- **Middle Man** — a class or function that mostly just delegates onward. → cut it, call the real target direct.\r\n- **Refused Bequest** — a subclass or implementer that ignores or overrides most of what it inherits. → drop the inheritance, use composition.\r\n\r\n### 4. Spawn both sub-agents in parallel\r\n\r\nSend a single message with two `Agent` tool calls. Use the `general-purpose` subagent for both.\r\n\r\n**Standards sub-agent prompt** — include:\r\n\r\n- The full diff command and commit list.\r\n- The list of standards-source files you found in step 3, **plus the smell baseline from step 3** pasted in full — the sub-agent has no other access to it.\r\n- The brief: \"Report — per file/hunk where relevant — (a) every place the diff violates a documented standard: cite the standard (file + the rule); and (b) any baseline smell you spot: name it and quote the hunk. Distinguish hard violations from judgement calls — documented-standard breaches can be hard, but baseline smells are always judgement calls, and a documented repo standard overrides the baseline. Skip anything tooling enforces. Under 400 words.\"\r\n\r\n**Spec sub-agent prompt** — include:\r\n\r\n- The diff command and commit list.\r\n- The path or fetched contents of the spec.\r\n- The brief: \"Report: (a) requirements the spec asked for that are missing or partial; (b) behaviour in the diff that wasn't asked for (scope creep); (c) requirements that look implemented but where the implementation looks wrong. Quote the spec line for each finding. Under 400 words.\"\r\n\r\nIf the spec is missing, skip the Spec sub-agent and note this in the final report.\r\n\r\n### 5. Aggregate\r\n\r\nPresent the two reports under `## Standards` and `## Spec` headings, verbatim or lightly cleaned. Do **not** merge or rerank findings — the two axes are deliberately separate (see _Why two axes_).\r\n\r\nEnd with a one-line summary: total findings per axis, and the worst issue _within each axis_ (if any). Don't pick a single winner across axes — that's the reranking the separation exists to prevent.\r\n\r\n## Why two axes\r\n\r\nA change can pass one axis and fail the other:\r\n\r\n- Code that follows every standard but implements the wrong thing → **Standards pass, Spec fail.**\r\n- Code that does exactly what the issue asked but breaks the project's conventions → **Spec pass, Standards fail.**\r\n\r\nReporting them separately stops one axis from masking the other.\r","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89]},"meta":{"source":{"type":"internal","value":"skill://code-review"}}},"isError":false,"timestamp":1789001085877}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_3miPvveXvVzjdjTQBfgEwKVO|fc_077f3569365e313d016aa1fd8571dc87d191c8433943ea33d6","toolName":"bash","startedAt":"2026-09-10T00:44:53.884Z","args":{"command":"git diff cb756dee^ cb756dee --"},"intent":"Inspecting prepared patch"},"id":"3cf7c05b","parentId":"4cc69128","timestamp":"2026-09-10T00:44:53.884Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_8dJV8YltnzqvJy0GLtFD6mKR|fc_077f3569365e313d016aa1fd8571f487d1a234c177837916de","toolName":"glob","startedAt":"2026-09-10T00:44:53.885Z","args":{"path":".worktrees/287-shellrestore/**/AGENTS.md;.worktrees/287-shellrestore/**/CLAUDE.md;.worktrees/287-shellrestore/**/CONTRIBUTING*;.worktrees/287-shellrestore/**/CODING_STANDARDS*;.worktrees/287-shellrest…"},"intent":"Locating review standards"},"id":"1dcaddc7","parentId":"3cf7c05b","timestamp":"2026-09-10T00:44:53.885Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_lzeI3mQfc2CtpdMQn74IVfMf|fc_077f3569365e313d016aa1fd85720087d1975365c71ea1b834","toolName":"read","startedAt":"2026-09-10T00:44:53.885Z","args":{"path":"issue://BigscreenVR/spt-bs-releases/287"},"intent":"Reading operator ruling"},"id":"23c9ba30","parentId":"1dcaddc7","timestamp":"2026-09-10T00:44:53.885Z"}
{"type":"message","id":"5f50e674","parentId":"23c9ba30","timestamp":"2026-09-10T00:44:53.885Z","message":{"role":"assistant","content":[{"type":"toolCall","id":"call_3miPvveXvVzjdjTQBfgEwKVO|fc_077f3569365e313d016aa1fd8571dc87d191c8433943ea33d6","name":"bash","arguments":{"i":"Inspecting prepared patch","command":"git diff cb756dee^ cb756dee --","cwd":"C:/Users/decid/Documents/projects/spt-core/.worktrees/287-shellrestore"},"intent":"Inspecting prepared patch"},{"type":"toolCall","id":"call_8dJV8YltnzqvJy0GLtFD6mKR|fc_077f3569365e313d016aa1fd8571f487d1a234c177837916de","name":"glob","arguments":{"i":"Locating review standards","path":".worktrees/287-shellrestore/**/AGENTS.md;.worktrees/287-shellrestore/**/CLAUDE.md;.worktrees/287-shellrestore/**/CONTRIBUTING*;.worktrees/287-shellrestore/**/CODING_STANDARDS*;.worktrees/287-shellrestore/docs/agents/*;AGENTS.md;CLAUDE.md","hidden":true,"gitignore":false,"limit":100},"intent":"Locating review standards"},{"type":"toolCall","id":"call_lzeI3mQfc2CtpdMQn74IVfMf|fc_077f3569365e313d016aa1fd85720087d1975365c71ea1b834","name":"read","arguments":{"i":"Reading operator ruling","path":"issue://BigscreenVR/spt-bs-releases/287"},"intent":"Reading operator ruling"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":1660,"output":207,"cacheRead":12672,"cacheWrite":0,"totalTokens":14539,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1789001085931,"responseId":"resp_077f3569365e313d016aa1fd7e25e487d18eee149e9f58704f","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"fc_077f3569365e313d016aa1fd8571dc87d191c8433943ea33d6","type":"function_call","status":"completed","arguments":"{\"i\":\"Inspecting prepared patch\",\"command\":\"git diff cb756dee^ cb756dee --\",\"cwd\":\"C:/Users/decid/Documents/projects/spt-core/.worktrees/287-shellrestore\"}","call_id":"call_3miPvveXvVzjdjTQBfgEwKVO","internal_chat_message_metadata_passthrough":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"},"name":"bash","metadata":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"}},{"id":"fc_077f3569365e313d016aa1fd8571f487d1a234c177837916de","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating review standards\",\"path\":\".worktrees/287-shellrestore/**/AGENTS.md;.worktrees/287-shellrestore/**/CLAUDE.md;.worktrees/287-shellrestore/**/CONTRIBUTING*;.worktrees/287-shellrestore/**/CODING_STANDARDS*;.worktrees/287-shellrestore/docs/agents/*;AGENTS.md;CLAUDE.md\",\"hidden\":true,\"gitignore\":false,\"limit\":100}","call_id":"call_8dJV8YltnzqvJy0GLtFD6mKR","internal_chat_message_metadata_passthrough":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"},"name":"glob","metadata":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"}},{"id":"fc_077f3569365e313d016aa1fd85720087d1975365c71ea1b834","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading operator ruling\",\"path\":\"issue://BigscreenVR/spt-bs-releases/287\"}","call_id":"call_lzeI3mQfc2CtpdMQn74IVfMf","internal_chat_message_metadata_passthrough":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"},"name":"read","metadata":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"}}]},"duration":7614.567400000058,"ttft":7556.009200000204,"completedAt":1789001093546,"contextSnapshot":{"promptTokens":14332,"nonMessageTokens":13209,"compactionEpoch":0}}}
{"type":"message","id":"5848032c","parentId":"5f50e674","timestamp":"2026-09-10T00:44:53.939Z","message":{"role":"toolResult","toolCallId":"call_8dJV8YltnzqvJy0GLtFD6mKR|fc_077f3569365e313d016aa1fd8571f487d1a234c177837916de","toolName":"glob","content":[{"type":"text","text":"AGENTS.md\nCLAUDE.md\n# .worktrees/287-shellrestore/\nCLAUDE.md\nAGENTS.md\n\nSkipped missing paths: .worktrees/287-shellrestore/docs/agents/*"}],"details":{"scopePath":".worktrees/287-shellrestore/**/AGENTS.md, .worktrees/287-shellrestore/**/CLAUDE.md, .worktrees/287-shellrestore/**/CONTRIBUTING*, .worktrees/287-shellrestore/**/CODING_STANDARDS*, AGENTS.md, CLAUDE.md","fileCount":4,"files":[".worktrees/287-shellrestore/CLAUDE.md",".worktrees/287-shellrestore/AGENTS.md","AGENTS.md","CLAUDE.md"],"truncated":false,"cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","missingPaths":[".worktrees/287-shellrestore/docs/agents/*"]},"isError":false,"timestamp":1789001093939}}
{"type":"message","id":"933950c8","parentId":"5848032c","timestamp":"2026-09-10T00:44:55.171Z","message":{"role":"toolResult","toolCallId":"call_lzeI3mQfc2CtpdMQn74IVfMf|fc_077f3569365e313d016aa1fd85720087d1975365c71ea1b834","toolName":"read","content":[{"type":"text","text":"# Issue #287: A daemon restart (no machine reboot) strands every persistent shell: the boot-restore discriminant reads the MACHINE boot instant\n\nState: OPEN\nAuthor: @SaberMage\nCreated: 2026-09-08T10:29:58Z\nUpdated: 2026-09-09T23:17:44Z\nLabels: state: GREENLIT, type: BUGFIX\nURL: https://github.com/BigscreenVR/spt-bs-releases/issues/287\n\n## Body\n\n## Symptom\n\nEvery `persistent` shell on hfenduleam (doyle/PACER-0, perri/PACER-0, todlando/PACER-0, doyle/alchemy-1, deployah/alchemy-0, …) went offline at the daemon restart of 2026-09-07 08:03:20Z and stayed offline. No error anywhere. The operator noticed 27 hours later because doyle's pacer had stopped pacing. `shell list` read `offline` the whole time, which is correct, and nothing ever relaunched.\n\n## Mechanism (measured 2026-09-08 10:26Z, spt-core 0.67.0, log `daemon.stderr.log` generation 0 / pid 48232)\n\n- Machine boot: 2026-09-04 08:02:53Z. Pacer launched by that boot's sweep: 2026-09-04 08:03:39Z (`shell.launch.json launched_ms=1788509019651`, native start stamp agrees).\n- Daemon restarted (no machine reboot): 2026-09-07 08:03:20Z. The pacer process (pid 27656) died with it. The heal fired: `SHELL_RECORD_HEALED:doyle/PACER-0: online -> offline` (8 such lines, one per persistent instance).\n- The restore never fired: zero `SHELL_BOOT_RESTORED` / `SHELL_OWNER_ONLINE_RESTORED` lines in the whole generation. The previous generation (boot 2026-09-04) shows `SHELL_BOOT_RESTORED:doyle/alchemy-0` etc.\n- Cause: `launch_predates_boot(launched_ms, boot_ms, slack)` in `spt_daemon::shellwake` compares against the MACHINE boot instant (`boot_instant_ms` = now minus `GetTickCount64`). A launch at boot+46 s does not predate that boot, so every daemon-only restart lands in the \"force-killed THIS boot => NOT eligible\" arm on all three triggers (boot sweep, owner-online edge, watcher). By construction, a daemon restart is indistinguishable from an operator force-kill under this discriminant.\n- doyle/alchemy-0 came back only because `spt shell cmd` wakes an offline persistent instance on demand (linkhost `wake_if_offline_persistent`); PACER receives no owner commands, so it stays down until a human runs `relink`. Confirmed: `spt shell relink PACER-0` at 10:28:39Z brought it online (pid 45088), status reply received.\n\n## Why this matters\n\nA daemon restart happens on every self-update. KNOWN-HAZARDS 2.7 (`REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS`) states \"a node restart must not strand a persistent instance\" and was closed for the machine-reboot face; the daemon-restart face reproduces the same stranding through the discriminant that closed it. The contract sentence (a persistent shell is online whenever its owner is) is false after every update.\n\n## Shape question for the fix (not ruled here)\n\nThe force-kill freeze (releases#78 leg (c)) must survive. Candidate: evaluate the corpse against the DAEMON generation instant as well as the machine boot instant, i.e. a corpse whose launch predates the current daemon start is a restart casualty. The open edge: a shell that survives a daemon restart and is then force-killed would have a launch older than the daemon start and would be relaunched. Whether any shell survives a restart (broker-hosted ones die with the broker; relay ones close on link-break) needs measuring before the rule is set.\n\n## Evidence\n\n- `C:\\Users\\decid\\AppData\\Local\\spt-core\\logs\\daemon.stderr.log` lines 64-119 (heals), no RESTORED lines.\n- `owlery/doyle/shells/PACER-0/{shell.pid,shell.launch.json,info.json}`.\n- perri/PACER-0 (pid 28464) and todlando/PACER-0 (pid 42736) still offline at filing; their owners must relink or issue a shell cmd.\n\n## Comments (2)\n\n### @SaberMage · 2026-09-08T10:31:53Z\n\nTwo riders from perri (2026-09-08 10:31Z), after relinking her three stranded shells:\n\n1. **Enumerate every stranded shell per owner.** The daemon log carried `SHELL_RECORD_HEALED` for perri/PACER-0, perri/alchemy-0 AND perri/alchemy-1 (lines 94, 103, 105 of the generation-0 log); my notice to her named only PACER-0. Whatever surfaces this class to an owner (a now-signal line, a shell-list marker, or the heal itself) should list all of that owner's healed instances, not the one the reporter happened to be using. Full stranded set at the 2026-09-07 08:03Z restart: deployah/alchemy-0, doyle/PACER-0, doyle/alchemy-0, doyle/alchemy-1, perri/PACER-0, perri/alchemy-0, perri/alchemy-1, todlando/PACER-0.\n\n2. **`spt shell relink` prints a success that reads as a failure.** On success it answers `SHELL_RELINKED:<id> owner=<o> pid=<n> status=offline (online at bind)` and the roster flips to online a beat later. Measured on my own relink at 10:28:39Z (pid 45088, online within 6 s). The token is honest about the bind handshake but the word `offline` on a success line is the part a reader keeps. A clearer token (for example `status=binding`, or the line naming the expected bind) is a small rider on whichever lane takes this issue.\n\nURL: https://github.com/BigscreenVR/spt-bs-releases/issues/287#issuecomment-5583758643\n\n### @SaberMage · 2026-09-08T10:36:55Z\n\n## OPERATOR RULING (2026-09-08 10:37Z, given to doyle in session; supersedes the shape question above)\n\n> \"daemon-only restart reads as an operator force-kill\" — that shouldn't matter for shells whose manifest has set `shell.persistent = true`. Such shells must come online anytime the endpoint comes online, end of story.\n\nConsequences for the fix (doyle's reading of the ruling, to be built as stated):\n\n1. For a `persistent = true` instance, the owner-online edge and the boot sweep relaunch it whenever the owner endpoint is online and the instance is down in fact (the derived read, pid-reuse-safe via the birth stamp). The `launch_predates_boot` conjunct is REMOVED from the persistent path. The heal (leg a) stays as is.\n2. The releases#78 leg (c) freeze (force-kill during operation is not auto-relaunched, to free an exe for overwrite) no longer governs persistent instances. An operator who needs a persistent shell to stay down uses the surfaces that say so (`teardown`, or a non-persistent manifest), not a kill. The mid-deploy exe-overwrite concern that motivated the freeze is answered by the update path, not by leaving a declared-persistent surface down.\n3. CONTEXT.md's sentence (\"a persistent shell is automatically brought online whenever its owner endpoint is online\") is already the contract; KNOWN-HAZARDS 2.7 and MANIFEST.md's sleep/wake text get amended to drop the boot-instant conjunct for persistent instances, and the REQ-SHELL-PERSISTENT-BOOT-RESTORE / REQ-SHELL-PERSISTENT-OWNER-ONLINE-RESTORE titles are re-worded in the same change. Regression cell: the daemon-restart shape (corpse launched AFTER machine boot, owner online, persistent) is restored on the owner-online edge — the exact row the current discriminant refuses.\n4. Riders from this thread ride the same lane: per-owner enumeration of stranded instances, and the relink success token.\n\nDispatch: product change, so todlando's queue, as a thin lane after his W3 push (or a #272 close rider if the operator prefers it inside the milestone).\n\nURL: https://github.com/BigscreenVR/spt-bs-releases/issues/287#issuecomment-5583820305"}],"details":{"contentType":"text/markdown","totalLines":65,"displayContent":{"text":"# Issue #287: A daemon restart (no machine reboot) strands every persistent shell: the boot-restore discriminant reads the MACHINE boot instant\n\nState: OPEN\nAuthor: @SaberMage\nCreated: 2026-09-08T10:29:58Z\nUpdated: 2026-09-09T23:17:44Z\nLabels: state: GREENLIT, type: BUGFIX\nURL: https://github.com/BigscreenVR/spt-bs-releases/issues/287\n\n## Body\n\n## Symptom\n\nEvery `persistent` shell on hfenduleam (doyle/PACER-0, perri/PACER-0, todlando/PACER-0, doyle/alchemy-1, deployah/alchemy-0, …) went offline at the daemon restart of 2026-09-07 08:03:20Z and stayed offline. No error anywhere. The operator noticed 27 hours later because doyle's pacer had stopped pacing. `shell list` read `offline` the whole time, which is correct, and nothing ever relaunched.\n\n## Mechanism (measured 2026-09-08 10:26Z, spt-core 0.67.0, log `daemon.stderr.log` generation 0 / pid 48232)\n\n- Machine boot: 2026-09-04 08:02:53Z. Pacer launched by that boot's sweep: 2026-09-04 08:03:39Z (`shell.launch.json launched_ms=1788509019651`, native start stamp agrees).\n- Daemon restarted (no machine reboot): 2026-09-07 08:03:20Z. The pacer process (pid 27656) died with it. The heal fired: `SHELL_RECORD_HEALED:doyle/PACER-0: online -> offline` (8 such lines, one per persistent instance).\n- The restore never fired: zero `SHELL_BOOT_RESTORED` / `SHELL_OWNER_ONLINE_RESTORED` lines in the whole generation. The previous generation (boot 2026-09-04) shows `SHELL_BOOT_RESTORED:doyle/alchemy-0` etc.\n- Cause: `launch_predates_boot(launched_ms, boot_ms, slack)` in `spt_daemon::shellwake` compares against the MACHINE boot instant (`boot_instant_ms` = now minus `GetTickCount64`). A launch at boot+46 s does not predate that boot, so every daemon-only restart lands in the \"force-killed THIS boot => NOT eligible\" arm on all three triggers (boot sweep, owner-online edge, watcher). By construction, a daemon restart is indistinguishable from an operator force-kill under this discriminant.\n- doyle/alchemy-0 came back only because `spt shell cmd` wakes an offline persistent instance on demand (linkhost `wake_if_offline_persistent`); PACER receives no owner commands, so it stays down until a human runs `relink`. Confirmed: `spt shell relink PACER-0` at 10:28:39Z brought it online (pid 45088), status reply received.\n\n## Why this matters\n\nA daemon restart happens on every self-update. KNOWN-HAZARDS 2.7 (`REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS`) states \"a node restart must not strand a persistent instance\" and was closed for the machine-reboot face; the daemon-restart face reproduces the same stranding through the discriminant that closed it. The contract sentence (a persistent shell is online whenever its owner is) is false after every update.\n\n## Shape question for the fix (not ruled here)\n\nThe force-kill freeze (releases#78 leg (c)) must survive. Candidate: evaluate the corpse against the DAEMON generation instant as well as the machine boot instant, i.e. a corpse whose launch predates the current daemon start is a restart casualty. The open edge: a shell that survives a daemon restart and is then force-killed would have a launch older than the daemon start and would be relaunched. Whether any shell survives a restart (broker-hosted ones die with the broker; relay ones close on link-break) needs measuring before the rule is set.\n\n## Evidence\n\n- `C:\\Users\\decid\\AppData\\Local\\spt-core\\logs\\daemon.stderr.log` lines 64-119 (heals), no RESTORED lines.\n- `owlery/doyle/shells/PACER-0/{shell.pid,shell.launch.json,info.json}`.\n- perri/PACER-0 (pid 28464) and todlando/PACER-0 (pid 42736) still offline at filing; their owners must relink or issue a shell cmd.\n\n## Comments (2)\n\n### @SaberMage · 2026-09-08T10:31:53Z\n\nTwo riders from perri (2026-09-08 10:31Z), after relinking her three stranded shells:\n\n1. **Enumerate every stranded shell per owner.** The daemon log carried `SHELL_RECORD_HEALED` for perri/PACER-0, perri/alchemy-0 AND perri/alchemy-1 (lines 94, 103, 105 of the generation-0 log); my notice to her named only PACER-0. Whatever surfaces this class to an owner (a now-signal line, a shell-list marker, or the heal itself) should list all of that owner's healed instances, not the one the reporter happened to be using. Full stranded set at the 2026-09-07 08:03Z restart: deployah/alchemy-0, doyle/PACER-0, doyle/alchemy-0, doyle/alchemy-1, perri/PACER-0, perri/alchemy-0, perri/alchemy-1, todlando/PACER-0.\n\n2. **`spt shell relink` prints a success that reads as a failure.** On success it answers `SHELL_RELINKED:<id> owner=<o> pid=<n> status=offline (online at bind)` and the roster flips to online a beat later. Measured on my own relink at 10:28:39Z (pid 45088, online within 6 s). The token is honest about the bind handshake but the word `offline` on a success line is the part a reader keeps. A clearer token (for example `status=binding`, or the line naming the expected bind) is a small rider on whichever lane takes this issue.\n\nURL: https://github.com/BigscreenVR/spt-bs-releases/issues/287#issuecomment-5583758643\n\n### @SaberMage · 2026-09-08T10:36:55Z\n\n## OPERATOR RULING (2026-09-08 10:37Z, given to doyle in session; supersedes the shape question above)\n\n> \"daemon-only restart reads as an operator force-kill\" — that shouldn't matter for shells whose manifest has set `shell.persistent = true`. Such shells must come online anytime the endpoint comes online, end of story.\n\nConsequences for the fix (doyle's reading of the ruling, to be built as stated):\n\n1. For a `persistent = true` instance, the owner-online edge and the boot sweep relaunch it whenever the owner endpoint is online and the instance is down in fact (the derived read, pid-reuse-safe via the birth stamp). The `launch_predates_boot` conjunct is REMOVED from the persistent path. The heal (leg a) stays as is.\n2. The releases#78 leg (c) freeze (force-kill during operation is not auto-relaunched, to free an exe for overwrite) no longer governs persistent instances. An operator who needs a persistent shell to stay down uses the surfaces that say so (`teardown`, or a non-persistent manifest), not a kill. The mid-deploy exe-overwrite concern that motivated the freeze is answered by the update path, not by leaving a declared-persistent surface down.\n3. CONTEXT.md's sentence (\"a persistent shell is automatically brought online whenever its owner endpoint is online\") is already the contract; KNOWN-HAZARDS 2.7 and MANIFEST.md's sleep/wake text get amended to drop the boot-instant conjunct for persistent instances, and the REQ-SHELL-PERSISTENT-BOOT-RESTORE / REQ-SHELL-PERSISTENT-OWNER-ONLINE-RESTORE titles are re-worded in the same change. Regression cell: the daemon-restart shape (corpse launched AFTER machine boot, owner online, persistent) is restored on the owner-online edge — the exact row the current discriminant refuses.\n4. Riders from this thread ride the same lane: per-owner enumeration of stranded instances, and the relink success token.\n\nDispatch: product change, so todlando's queue, as a thin lane after his W3 push (or a #272 close rider if the operator prefers it inside the milestone).\n\nURL: https://github.com/BigscreenVR/spt-bs-releases/issues/287#issuecomment-5583820305","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65]},"meta":{"source":{"type":"internal","value":"issue://BigscreenVR/spt-bs-releases/287"}}},"isError":false,"timestamp":1789001095170}}
{"type":"message","id":"5c361f52","parentId":"933950c8","timestamp":"2026-09-10T00:44:55.216Z","message":{"role":"toolResult","toolCallId":"call_3miPvveXvVzjdjTQBfgEwKVO|fc_077f3569365e313d016aa1fd8571dc87d191c8433943ea33d6","toolName":"bash","content":[{"type":"text","text":"CHANGELOG.md | 13 +++++++++++++\nCONTEXT.md | 8 ++++----\ncrates/spt-daemon/src/shellhost.rs | 25 ++++++++++++++++++------\ncrates/spt-daemon/src/shellwake.rs | 507 +++++++++---------------\ncrates/spt-store/src/shellinfo.rs | 6 +++---\ncrates/spt/src/cli.rs | 3 ++-\ncrates/spt/tests/shell_relink_force_e2e.rs | 10 ++++++++++\ndocs-site/src/shells/overview.md | 45 ++++++++++++------------\ndocs/KNOWN-HAZARDS.md | 17 +++++++++--------\ndocs/MANIFEST.md | 8 +++++++-\ntraceable-reqs.toml | 28 ++++++++++++++++++++----\n11 files changed, 291 insertions(+), 379 deletions(-)\n\n--- Changes ---\n\nFile: CHANGELOG.md\n  @@ -10,6 +10,19 @@ breaks something, or changes the observable behavior of existing surfaces\n  +## [Unreleased]\n  +\n  +### Fixed\n  +\n  +- Every persistent shell comes back with its owner after a daemon-only restart,\n  +  not just after a machine reboot. Shells launched during the current boot or\n  […7 changed lines elided…]\n\nFile: CONTEXT.md\n  @@ -388,11 +388,11 @@ The manifest value is the **floor**: a node/endpoint setting may **tighten** (de\n  -- **`persistent`** (bool): the shell is automatically brought **online whenever its owner endpoint is online**. Covers \"owner already up.\" **Node-qualified onc…[+1722]\n  -  **Daemons boot before endpoints do, so the sweep is not the only trigger** (releases#228): an owner that is not yet online when the sweep runs is skipped by …[+949]\n  -- **`wake_command`** (template): a long-running **wake-watcher** process spt-core runs *while the shell is offline* **and eligible** (the eligibility rule is b…[+273]\n  +- **`persistent`** (bool): the shell is automatically brought **online whenever its owner endpoint is online**. Covers \"owner already up.\" **Node-qualified onc…[+805]\n  +  **Daemons boot before endpoints do, so the sweep is not the only trigger:** the reconcile loop runs that **same per-owner restore decision** on each owner **…[+403]\n  +- **`wake_command`** (template): a long-running **wake-watcher** process spt-core runs while the shell is offline and eligible. Its sole job is to fire a wake.…[+164]\n  […2 changed lines elided…]\n\nFile: crates/spt-daemon/src/shellhost.rs\n  @@ -312,16 +312,14 @@ pub fn routes_through_daemon(shell: &Shell) -> bool {\n  -/// DERIVED, never the recorded status (REQ-HAZARD-SHELL-STALE-ONLINE): a pid\n  -/// that is absent, unparseable or `0` reads NOT-a-winner, the same\n  -/// fail-toward-alive stance the rest of the shell machinery holds — so a\n  -/// backend that records no pid can never block its own relaunch. Same shape and\n  -/// same non-fault wording as [`crate::shellwake::resolve_wake`]'s existing\n  -/// guard; this is that rule extracted, not a second one.\n  […6 changed lines elided…]\n  @@ -1204,6 +1202,7 @@ mod tests {\n  +    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n  @@ -1225,6 +1224,18 @@ mod tests {\n  +        shellinfo::record_shell_launch(perch, me, 1_000);\n  +        assert_eq!(live_launch_winner(perch), Some(me), \"same birth retains custody\");\n  +        let mut launch = shellinfo::read_shell_launch(perch).unwrap();\n  +        launch.pid_started_at = Some(\n  +            launch.pid_started_at.expect(\"native birth stamp on supported platforms\") ^ 1,\n  +        );\n  […6 changed lines elided…]\n\nFile: crates/spt-daemon/src/shellwake.rs\n  @@ -613,7 +613,7 @@ fn fill_wake_command(\n  -fn heal_stale_online_records(owlery: &Path) {\n  +fn heal_stale_online_records(owlery: &Path, events: &mut dyn std::io::Write) {\n  @@ -627,7 +627,8 @@ fn heal_stale_online_records(owlery: &Path) {\n  -                spt_proto::emit_line_err!(\"SHELL_RECORD_HEALED:{owner}/{shell_id}: online -> {truth}\");\n  +                // [impl->REQ-SHELL-HEAL-ENUMERATION]\n  +                let _ = spt_proto::emit_line!(events, \"SHELL_RECORD_HEALED:{owner}/{shell_id}: online -> {truth}\");\n  @@ -639,67 +640,40 @@ fn heal_stale_online_records(owlery: &Path) {\n  -/// that reads as \"a corpse from the previous boot\" and the sweep would relaunch\n  -/// something that had just started. Ten seconds is far wider than any plausible\n  +/// that reads as \"a corpse from the previous boot\" and a nonpersistent watcher\n  +/// could be armed after a same-boot death. Ten seconds is far wider than plausible\n  -/// Whether a recorded launch predates the current boot — leg (b)'s discriminant\n  -/// (releases#78).\n  […63 changed lines elided…]\n  @@ -742,26 +716,10 @@ impl RestoreTrigger {\n  -/// It is shared rather than restated because the refusal arms ARE the ruling:\n  -/// doyle's condition on releases#228 was that the owner-online trigger carry the\n  -/// boot sweep's arms VERBATIM, and two copies of a conjunction drift the first\n  -/// time one side is edited. Every gate is a conjunction and each one is\n  -/// load-bearing:\n  -///\n  […18 changed lines elided…]\n  @@ -773,13 +731,11 @@ impl RestoreTrigger {\n  -#[allow(clippy::too_many_arguments)]\n  -    boot_ms: u64,\n  @@ -795,14 +751,10 @@ fn restore_persistent_shells_of_owner(\n  -        if shellinfo::is_shell_online(&perch, &info) {\n  -            continue; // already up in fact\n  -        }\n  -        let Some(launch) = shellinfo::read_shell_launch(&perch) else {\n  -            continue; // no stamp ⇒ cannot prove it predates boot ⇒ leave it down\n  -        };\n  […6 changed lines elided…]\n  @@ -844,11 +796,6 @@ pub fn restore_persistent_shells_at_boot(\n  -    let Some(boot_ms) = spt_store::proc::boot_instant_ms() else {\n  -        // No boot oracle on this platform: the discriminant cannot be evaluated,\n  -        // so nothing is restored. Silence beats guessing at a relaunch.\n  -        return Vec::new();\n  -    };\n  @@ -856,7 +803,6 @@ pub fn restore_persistent_shells_at_boot(\n  -            boot_ms,\n  […19 hunks elided…]\n\nFile: crates/spt-store/src/shellinfo.rs\n  @@ -145,9 +145,9 @@ pub const SHELL_LAUNCH_FILE: &str = \"shell.launch.json\";\n  -/// - `launched_ms` is an ABSOLUTE wall-clock stamp (epoch-ms) and is what the\n  -///   boot sweep compares against the boot instant. It exists precisely because\n  -///   the field above cannot answer that question portably.\n  +/// - `launched_ms` is an ABSOLUTE wall-clock stamp (epoch-ms), used by the\n  +///   nonpersistent watcher's cross-boot freeze. Persistent restoration does not\n  +///   gate on launch age; the native birth stamp still authenticates liveness.\n\nFile: crates/spt/src/cli.rs\n  @@ -23531,7 +23531,8 @@ fn cmd_shell(action: ShellCmd, json: bool) -> i32 {\n  -                    eprintln!(\"SHELL_RELINKED:{id} owner={owner} pid={pid} status=offline (online at bind)\");\n  +                    // [impl->REQ-SHELL-RELINK-BINDING-RESULT]\n  +                    eprintln!(\"SHELL_RELINKED:{id} owner={owner} pid={pid} status=binding (online at bind)\");\n\nFile: crates/spt/tests/shell_relink_force_e2e.rs\n  @@ -184,6 +184,16 @@ fn force_stops_a_running_binary_then_relinks_it_under_the_same_id() {\n  +    // [int->REQ-SHELL-RELINK-BINDING-RESULT]\n  +    assert!(\n  +        err[relinked_at..]\n  +            .lines()\n  +            .next()\n  +            .unwrap()\n  […4 changed lines elided…]\n\nFile: docs-site/src/shells/overview.md\n  @@ -93,30 +93,22 @@ Lifecycle extras: `persistent` shells auto-online with their owner;\n  -Two lifecycle facts worth knowing before you rely on either:\n  -\n  -- **A node restart brings `persistent` instances back by a boot sweep, and an\n  -  owner that comes online after that sweep by an owner-online edge.** A machine\n  -  death breaks no link, so the instance's record survives the reboot claiming it\n  -  is still up; the daemon corrects that record and then, once per daemon\n  […34 changed lines elided…]\n  @@ -129,6 +121,11 @@ Two lifecycle facts worth knowing before you rely on either:\n  +  <!-- [doc->REQ-SHELL-RELINK-BINDING-RESULT] -->\n  +  A successful local relink reports `status=binding (online at bind)`: the\n  +  binary has launched and the bind handshake is still the online transition.\n  +  It does not describe that successful launch as `offline`.\n  +\n\nFile: docs/KNOWN-HAZARDS.md\n  @@ -76,19 +76,20 @@ Hard-won edge cases harvested from the sister project (`claude_skill_owl`, ~80 c\n  -- **No spontaneous relaunch of a SAME-BOOT death (operator-ratified, flynn 2026-07-25; re-based BAROMETER W2):** the reconciler does NOT adopt an instance whos…[+598]\n  -- **What holds that open is no longer the stale record (BAROMETER W2):** until W2 the guarantee rested on an accident: a force-killed instance was held out of …[+943]\n  +- **No spontaneous relaunch of a NONPERSISTENT same-boot death:** the releases#78 force-kill freeze still protects nonpersistent instances, so killing one to f…[+278]\n  +- **Watcher eligibility:** recorded `offline`, no live binary awaiting bind, and either `persistent`, no corpse, or a corpse whose launch predates the machine …[+223]\n  -### 2.7 A node restart permanently strands every `persistent` shell — and the existing cascade test cannot see that it does  `[REQ-HAZARD-RESTART-STRANDS-PERSI…[+14]\n  -- **Failure:** the contract's own sentence says a `persistent` shell is **online whenever its owner endpoint is online**, and a node restart is the one case wh…[+510]\n  […11 changed lines elided…]\n\nFile: docs/MANIFEST.md\n  @@ -788,7 +788,13 @@ types = [\"image\", \"sound\", \"event\"]   # \"event\" = arbitrary descriptive payloads\n  -**Sleep/wake:** when offline **and eligible** (+ `wake_command` set), spt-core runs the wake-watcher (mutually exclusive with the shell binary); `exit(wake-opc…[+2187]\n  +**Sleep/wake:** when offline **and eligible** (+ `wake_command` set), spt-core runs the wake-watcher, mutually exclusive with a live shell binary including one…[+366]\n  +\n  +**Persistent restoration:** on the once-per-daemon-generation boot sweep and each owner offline→online edge, spt-core restores **every** down persistent instan…[+530]\n  +\n  +**Watcher eligibility:** recorded-offline instances with no live binary are eligible if persistent, cleanly closed/never launched (no corpse), or a dated pre-b…[+282]\n  […2 changed lines elided…]\n\nFile: traceable-reqs.toml\n  @@ -3183,6 +3183,12 @@ required_stages = [\"doc\", \"impl\", \"unit\", \"int\"]  # W1 ACTIVATED at mint (doyle\n  +# AMENDED releases#287, operator ruling 2026-09-08: the historical title below's\n  +# same-boot/undatable-corpse freeze and no-spontaneous-relaunch clauses now apply\n  +# ONLY to NONPERSISTENT instances. Persistent restoration and watcher eligibility\n  +# follow REQ-SHELL-PERSISTENT-BOOT-RESTORE / REQ-SHELL-OWNER-ONLINE-RESTORE, without\n  +# launch-age or stamp-presence gates. Birth-safe live/binding exclusion remains.\n  +# The liveness, identity/state preservation and durable-channel clauses stand.\n  @@ -4151,13 +4157,23 @@ required_stages = [\"doc\", \"impl\", \"unit\"]  # ACTIVATED BAROMETER W1 at impl+unit\n  -title = \"A `persistent` SHELL COMES BACK AFTER A NODE RESTART, WHICH IS THE ONE CASE THE CONTRACT'S OWN SENTENCE PROMISES AND THE IMPLEMENTATION NEVER HANDLED.…[+3939]\n  -required_stages = [\"doc\", \"impl\", \"unit\"]  # doc ACTIVATED BAROMETER W2 doc leg (doyle-ruled, this milestone): the CONTEXT.md `persistent` amendment states the…[+193]\n  +title = \"Persistent shells are restored on each daemon-generation boot sweep whenever their owner is online and they are down in fact, regardless of machine bo…[+457]\n  +required_stages = [\"doc\", \"impl\", \"unit\"]\n  -title = \"A `persistent` SHELL WHOSE OWNER ENDPOINT WAS NOT YET ONLINE WHEN THE BOOT SWEEP RAN IS RESTORED WHEN THAT OWNER COMES ONLINE, RATHER THAN STAYING DOW…[+3074]\n  -required_stages = [\"doc\", \"impl\", \"unit\"]  # ACTIVATED CONDUIT W2 in the commit landing the evidence (doyle ruled shape (c) with conditions on releases#228; (a…[+285]\n  […12 changed lines elided…]\n  @@ -4166,7 +4182,7 @@ required_stages = [\"doc\", \"impl\", \"unit\", \"int\"]  # ACTIVATED BAROMETER W2 reope\n  -title = \"A NODE RESTART MUST NOT PERMANENTLY STRAND EVERY `persistent` SHELL, AND THE EXISTING CASCADE TEST CANNOT SEE THAT IT DOES. Distinct from REQ-HAZARD-S…[+2249]\n  +title = \"A daemon-only or machine restart must not strand persistent shells. Heal stale online records and restore every down persistent instance when its owne…[+457]\n  @@ -4386,6 +4402,7 @@ requirements = [\n  +  \"REQ-SHELL-RELINK-BINDING-RESULT\",\n  @@ -4550,6 +4567,7 @@ requirements = [\n  +  \"REQ-SHELL-HEAL-ENUMERATION\",\n[raw output: artifact://112]\n\n\nWall time: 1.32 seconds"}],"details":{"timeoutSeconds":300,"wallTimeMs":1317.3914000000805},"isError":false,"timestamp":1789001095215}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_WZyFmSAexqkpOcnZ0vv8qvQW|fc_077f3569365e313d016aa1fd8dab9887d1a91d04944628e375","toolName":"read","startedAt":"2026-09-10T00:45:01.748Z","args":{"path":"artifact://112"},"intent":"Reading full patch"},"id":"1d5252c4","parentId":"5c361f52","timestamp":"2026-09-10T00:45:01.748Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_8tZqKMdMkqqi3pvzPDyJk9XB|fc_077f3569365e313d016aa1fd8dabac87d180bc9028fa0b2249","toolName":"read","startedAt":"2026-09-10T00:45:01.748Z","args":{"path":".worktrees/287-shellrestore/AGENTS.md"},"intent":"Reading worktree rules"},"id":"1de44e33","parentId":"1d5252c4","timestamp":"2026-09-10T00:45:01.749Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_SoXnwtS6OxjKMOFNagDliZYD|fc_077f3569365e313d016aa1fd8dabb487d19e2dbb48a2ab8749","toolName":"read","startedAt":"2026-09-10T00:45:01.749Z","args":{"path":".worktrees/287-shellrestore/CLAUDE.md"},"intent":"Reading worktree guidance"},"id":"dd3ccb8a","parentId":"1de44e33","timestamp":"2026-09-10T00:45:01.749Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_8hdpFa9bn2kiSxq2op7zz0Ps|fc_077f3569365e313d016aa1fd8dabbc87d19b254ad7b50831fe","toolName":"read","startedAt":"2026-09-10T00:45:01.749Z","args":{"path":"AGENTS.md"},"intent":"Reading inherited rules"},"id":"749d2291","parentId":"dd3ccb8a","timestamp":"2026-09-10T00:45:01.749Z"}
{"type":"message","id":"80ba6fde","parentId":"749d2291","timestamp":"2026-09-10T00:45:01.749Z","message":{"role":"assistant","content":[{"type":"toolCall","id":"call_WZyFmSAexqkpOcnZ0vv8qvQW|fc_077f3569365e313d016aa1fd8dab9887d1a91d04944628e375","name":"read","arguments":{"i":"Reading full patch","path":"artifact://112"},"intent":"Reading full patch"},{"type":"toolCall","id":"call_8tZqKMdMkqqi3pvzPDyJk9XB|fc_077f3569365e313d016aa1fd8dabac87d180bc9028fa0b2249","name":"read","arguments":{"i":"Reading worktree rules","path":".worktrees/287-shellrestore/AGENTS.md"},"intent":"Reading worktree rules"},{"type":"toolCall","id":"call_SoXnwtS6OxjKMOFNagDliZYD|fc_077f3569365e313d016aa1fd8dabb487d19e2dbb48a2ab8749","name":"read","arguments":{"i":"Reading worktree guidance","path":".worktrees/287-shellrestore/CLAUDE.md"},"intent":"Reading worktree guidance"},{"type":"toolCall","id":"call_8hdpFa9bn2kiSxq2op7zz0Ps|fc_077f3569365e313d016aa1fd8dabbc87d19b254ad7b50831fe","name":"read","arguments":{"i":"Reading inherited rules","path":"AGENTS.md"},"intent":"Reading inherited rules"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":5805,"output":123,"cacheRead":14208,"cacheWrite":0,"totalTokens":20136,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1789001095225,"responseId":"resp_077f3569365e313d016aa1fd8780a887d18cc0e901f74be91c","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"fc_077f3569365e313d016aa1fd8dab9887d1a91d04944628e375","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading full patch\",\"path\":\"artifact://112\"}","call_id":"call_WZyFmSAexqkpOcnZ0vv8qvQW","internal_chat_message_metadata_passthrough":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"},"name":"read","metadata":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"}},{"id":"fc_077f3569365e313d016aa1fd8dabac87d180bc9028fa0b2249","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading worktree rules\",\"path\":\".worktrees/287-shellrestore/AGENTS.md\"}","call_id":"call_8tZqKMdMkqqi3pvzPDyJk9XB","internal_chat_message_metadata_passthrough":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"},"name":"read","metadata":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"}},{"id":"fc_077f3569365e313d016aa1fd8dabb487d19e2dbb48a2ab8749","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading worktree guidance\",\"path\":\".worktrees/287-shellrestore/CLAUDE.md\"}","call_id":"call_SoXnwtS6OxjKMOFNagDliZYD","internal_chat_message_metadata_passthrough":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"},"name":"read","metadata":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"}},{"id":"fc_077f3569365e313d016aa1fd8dabbc87d19b254ad7b50831fe","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading inherited rules\",\"path\":\"AGENTS.md\"}","call_id":"call_8hdpFa9bn2kiSxq2op7zz0Ps","internal_chat_message_metadata_passthrough":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"},"name":"read","metadata":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"}}]},"duration":6522.552399999928,"ttft":6479.525200000033,"completedAt":1789001101747,"contextSnapshot":{"promptTokens":20013,"nonMessageTokens":13209,"compactionEpoch":0}}}
{"type":"message","id":"8abef4f6","parentId":"80ba6fde","timestamp":"2026-09-10T00:45:01.754Z","message":{"role":"toolResult","toolCallId":"call_WZyFmSAexqkpOcnZ0vv8qvQW|fc_077f3569365e313d016aa1fd8dab9887d1a91d04944628e375","toolName":"read","content":[{"type":"text","text":"diff --git a/CHANGELOG.md b/CHANGELOG.md\nindex eafe18d6..39bf8144 100644\n--- a/CHANGELOG.md\n+++ b/CHANGELOG.md\n@@ -10,6 +10,19 @@ breaks something, or changes the observable behavior of existing surfaces\n broadly; **patch** for fixes, and for additive opt-in capability — a new key,\n flag, or page that no existing user can encounter without opting into it.\n \n+## [Unreleased]\n+\n+### Fixed\n+\n+- Every persistent shell comes back with its owner after a daemon-only restart,\n+  not just after a machine reboot. Shells launched during the current boot or\n+  without a saved launch stamp are no longer left down, and recovery notices\n+  name every healed instance for each owner. A persistent shell must be torn\n+  down or made nonpersistent to stay down; killing its process does not opt out\n+  of restoration.\n+- A successful `spt shell relink` reports `binding` while awaiting its handshake,\n+  rather than making the successful launch read as `offline`.\n+\n ## [0.68.0] - 2026-09-08\n \n Web serving. Files, directories, adapter documentation and the changelog are\ndiff --git a/CONTEXT.md b/CONTEXT.md\nindex cdb8a8e1..6698a590 100644\n--- a/CONTEXT.md\n+++ b/CONTEXT.md\n@@ -388,11 +388,11 @@ The manifest value is the **floor**: a node/endpoint setting may **tighten** (de\n \n A shell automatically goes **offline when its owner endpoint goes offline**. Two complementary manifest options bring it (and its owner) back:\n \n-- **`persistent`** (bool): the shell is automatically brought **online whenever its owner endpoint is online**. Covers \"owner already up.\" **Node-qualified once instances span nodes** (ruled 2026-07-24, #70 grill): when an owner comes online at a given node, only the **latest** instances qualified to that *endpoint@node* pairing come online with it — a stale instance of the same shell left on another node does not auto-wake alongside. **A node restart is covered by a boot sweep, and an owner that comes online after that sweep by an owner-online edge** (BAROMETER W2 releases#78; CONDUIT W2 releases#228): a machine death breaks no link, so the instance's record survives the reboot still saying `online` and the wake cascade — which reads the *recorded* field — skips it, stranding every `persistent` instance on the node at once. The daemon therefore (i) heals a recorded `online` its derived status contradicts, every reconcile cycle, writing only when the value actually changes, and (ii) runs a **once-per-daemon-generation boot sweep** that relaunches an instance when **all four** conditions hold: the adapter declares `persistent`; the **owner endpoint is online** (an offline owner is owed nothing); the instance is down **in fact**; and its **recorded launch predates the boot instant**. A launch parks a **birth stamp** beside the pid to make that last question answerable — **no stamp ⇒ not restored** (the sweep cannot prove the corpse predates boot, and leaving a shell down beats relaunching a binary an operator may have killed deliberately), and a platform with no boot oracle restores nothing. A shell force-killed *during* this boot is launched after it, can never satisfy the predicate, and so is never spontaneously relaunched — the deploy/quarantine ruling of KNOWN-HAZARDS 2.6 preserved by construction. <!-- [doc->REQ-SHELL-PERSISTENT-BOOT-RESTORE] -->\n-  **Daemons boot before endpoints do, so the sweep is not the only trigger** (releases#228): an owner that is not yet online when the sweep runs is skipped by that very conjunct, and — the sweep being once per daemon generation, and bringup from offline emitting no rest edge — nothing revisited it, so its shells stayed down until a human relinked. The reconcile loop therefore runs the **same four conditions** for every owner that has gone **offline→online since the last pass**. The trigger is that **edge, never the level “the owner is online”**: a successful restore restamps itself out of eligibility, but a *failing* launch under a level trigger would be retried every reconcile tick for as long as the owner stayed up, where the edge bounds it to one attempt per owner-online event. Re-evaluating later in a generation can only refuse **more**, never less — every launch path restamps the birth stamp, a kill never does, and the boot instant is fixed for the machine boot — so the force-kill ruling above is preserved identically on the second trigger. <!-- [doc->REQ-SHELL-OWNER-ONLINE-RESTORE] -->\n-- **`wake_command`** (template): a long-running **wake-watcher** process spt-core runs *while the shell is offline* **and eligible** (the eligibility rule is below — an offline instance sitting on a same-boot corpse arms no watcher). Its sole job is to fire a wake. It runs on the **shell's node** (where the platform wake-event originates — e.g. interacting with a \"shut-down\" avatar). Covers \"owner is down, wake it from outside.\"\n+- **`persistent`** (bool): the shell is automatically brought **online whenever its owner endpoint is online**. Covers \"owner already up.\" **Node-qualified once instances span nodes** (ruled 2026-07-24, #70 grill): when an owner comes online at a given node, only the **latest** instances qualified to that *endpoint@node* pairing come online with it — a stale instance of the same shell left on another node does not auto-wake alongside. A **once-per-daemon-generation boot sweep** restores every down persistent instance whose owner is online, after a machine reboot **or a daemon-only restart**. Birth-safe liveness protects both an online binary and a live launch awaiting bind. Machine boot time, recorded launch age, and missing launch stamps do **not** gate restoration (releases#287 operator ruling). The daemon also heals stale `online` records to `offline` each reconcile cycle, only on an actual change. <!-- [doc->REQ-SHELL-PERSISTENT-BOOT-RESTORE] -->\n+  **Daemons boot before endpoints do, so the sweep is not the only trigger:** the reconcile loop runs that **same per-owner restore decision** on each owner **offline→online edge**. The edge is seeded before the sweep; a live launch awaiting bind prevents duplicate launch if both triggers reach it. Failed launches are attempted once per owner-online event, not every tick while the owner stays online. A force-kill does not revoke persistence: use `teardown` or a nonpersistent manifest to keep the instance down. <!-- [doc->REQ-SHELL-OWNER-ONLINE-RESTORE] -->\n+- **`wake_command`** (template): a long-running **wake-watcher** process spt-core runs while the shell is offline and eligible. Its sole job is to fire a wake. It runs on the **shell's node**, where the platform wake-event originates (e.g. interacting with a shut-down avatar). Covers \"owner is down, wake it from outside.\"\n \n-**Online/offline are mutually-exclusive processes:** online ⇒ the shell binary runs (no watcher); offline ⇒ the wake-watcher runs (no shell binary). spt-core flips between them. **Offline does not by itself arm the watcher — a third outcome exists** (BAROMETER W2, releases#78): an instance that is offline over a **corpse launched during this boot** (a force-kill, a crash) is deliberately held out of watcher eligibility, so *neither* process runs until recovery is demanded — `relink`, or a `shell cmd` that wakes. Eligibility is recorded-`offline` **AND** (no corpse, **or** the corpse's launch predates the boot instant); an instance whose corpse cannot be dated — no birth stamp, or no boot oracle on the platform — is treated as ineligible rather than guessed at. Only a cleanly-closed instance and a **restart casualty** arm a watcher. The reason is the deploy case: on Windows an operator kills a shell precisely to free its exe for overwrite, and a watcher armed there would re-lock the file under them.\n+**Online/offline are mutually-exclusive processes:** a live shell binary (including a launch awaiting bind) excludes its wake-watcher. A recorded-offline instance is watcher-eligible when it is persistent, carries no corpse, or its corpse's launch predates machine boot. **Only nonpersistent instances retain the same-boot force-kill freeze:** their same-boot or undatable corpses arm no watcher, preserving the deploy/overwrite case. A persistent corpse bypasses that freeze even without a launch stamp or boot oracle (releases#287). Missing birth evidence retains the existing fail-toward-alive process probe; it is not itself a restore refusal.\n \n **Exit-opcode supervision:** the wake-watcher exiting with the **wake opcode** → spt-core runs the wake resolution (below) and brings the shell online. Any *other* exit (crash) → respawn the watcher with **exponential backoff + eventual give-up** (until next shell activity) — crash-bug safety. One watcher per offline shell instance.\n \ndiff --git a/crates/spt-daemon/src/shellhost.rs b/crates/spt-daemon/src/shellhost.rs\nindex 593c6f68..44df752b 100644\n--- a/crates/spt-daemon/src/shellhost.rs\n+++ b/crates/spt-daemon/src/shellhost.rs\n@@ -312,16 +312,14 @@ pub fn routes_through_daemon(shell: &Shell) -> bool {\n /// down and adopt it. Used by BOTH sides of the routed launch, because a late\n /// ack can put the caller and the daemon in the same launch at once.\n ///\n-/// DERIVED, never the recorded status (REQ-HAZARD-SHELL-STALE-ONLINE): a pid\n-/// that is absent, unparseable or `0` reads NOT-a-winner, the same\n-/// fail-toward-alive stance the rest of the shell machinery holds — so a\n-/// backend that records no pid can never block its own relaunch. Same shape and\n-/// same non-fault wording as [`crate::shellwake::resolve_wake`]'s existing\n-/// guard; this is that rule extracted, not a second one.\n+/// Birth-safe, including a binary awaiting bind: a recycled live pid is NOT a\n+/// winner when its native birth stamp mismatches. Missing, unparseable and zero\n+/// pids remain not-a-winner, preserving pid-less backend behavior.\n // [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\n+// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n pub fn live_launch_winner(shell_perch: &Path) -> Option<u32> {\n     shellinfo::read_shell_pid(shell_perch)\n-        .filter(|&p| p != 0 && spt_store::proc::is_process_alive(p))\n+        .filter(|&p| p != 0 && !shellinfo::shell_pid_provably_dead(shell_perch))\n }\n \n /// THE route-selection seam. Generic over the daemon call and the in-process\n@@ -1204,6 +1202,7 @@ mod tests {\n     // that parses non-zero AND probes alive is a winner. Absent / `0` /\n     // unparseable / dead must all read NOT-a-winner, or a backend that records\n     // no usable pid would be unable to relaunch its own instance.\n+    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n     #[test]\n     fn only_a_live_recorded_pid_wins_a_launch_race() {\n         let tmp = tempfile::tempdir().unwrap();\n@@ -1225,6 +1224,18 @@ mod tests {\n             Some(me),\n             \"a recorded pid that probes alive IS the winner\"\n         );\n+        shellinfo::record_shell_launch(perch, me, 1_000);\n+        assert_eq!(live_launch_winner(perch), Some(me), \"same birth retains custody\");\n+        let mut launch = shellinfo::read_shell_launch(perch).unwrap();\n+        launch.pid_started_at = Some(\n+            launch.pid_started_at.expect(\"native birth stamp on supported platforms\") ^ 1,\n+        );\n+        std::fs::write(\n+            perch.join(shellinfo::SHELL_LAUNCH_FILE),\n+            serde_json::to_string(&launch).unwrap(),\n+        )\n+        .unwrap();\n+        assert_eq!(live_launch_winner(perch), None, \"a recycled pid cannot block restore\");\n     }\n \n     // [unit->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE] a REAL dial at a socket name\ndiff --git a/crates/spt-daemon/src/shellwake.rs b/crates/spt-daemon/src/shellwake.rs\nindex 7caf020f..d9631b4d 100644\n--- a/crates/spt-daemon/src/shellwake.rs\n+++ b/crates/spt-daemon/src/shellwake.rs\n@@ -613,7 +613,7 @@ fn fill_wake_command(\n /// correction of a field that is already wrong. Nothing is started here.\n // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n // [impl->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\n-fn heal_stale_online_records(owlery: &Path) {\n+fn heal_stale_online_records(owlery: &Path, events: &mut dyn std::io::Write) {\n     for owner in spt_store::perch::list_self_perch_ids(owlery) {\n         for (shell_id, info) in shellinfo::list_shells(owlery, &owner) {\n             if info.status != SHELL_STATUS_ONLINE {\n@@ -627,7 +627,8 @@ fn heal_stale_online_records(owlery: &Path) {\n             let mut healed = info.clone();\n             healed.status = truth.to_string();\n             if shellinfo::write_shell_info(&perch, &healed).is_ok() {\n-                spt_proto::emit_line_err!(\"SHELL_RECORD_HEALED:{owner}/{shell_id}: online -> {truth}\");\n+                // [impl->REQ-SHELL-HEAL-ENUMERATION]\n+                let _ = spt_proto::emit_line!(events, \"SHELL_RECORD_HEALED:{owner}/{shell_id}: online -> {truth}\");\n             }\n         }\n     }\n@@ -639,67 +640,40 @@ fn heal_stale_online_records(owlery: &Path) {\n /// now-minus-uptime, which drifts by however long the two calls take and by\n /// whatever the clock did since — so a process launched in the first moments of\n /// a boot can compute as very slightly older than the boot itself. Without slack\n-/// that reads as \"a corpse from the previous boot\" and the sweep would relaunch\n-/// something that had just started. Ten seconds is far wider than any plausible\n+/// that reads as \"a corpse from the previous boot\" and a nonpersistent watcher\n+/// could be armed after a same-boot death. Ten seconds is far wider than plausible\n /// derivation error and far narrower than the gap this discriminant exists to\n /// detect (a previous boot is minutes to days back).\n pub const BOOT_RESTORE_SLACK_MS: u64 = 10_000;\n \n-/// Whether a recorded launch predates the current boot — leg (b)'s discriminant\n-/// (releases#78).\n-///\n-/// This is what preserves the force-kill ruling BY CONSTRUCTION rather than by a\n-/// second rule: a shell force-killed during steady-state operation was launched\n-/// AFTER the current boot, so it can never satisfy this predicate and is never\n-/// spontaneously relaunched — the mid-deploy exe-overwrite case that ruling\n-/// protects is untouched. Only a fresh boot restores.\n-///\n-/// Pure over its inputs so the boundary is unit-testable without a reboot.\n-// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n+/// Whether a nonpersistent shell's recorded launch predates the current boot.\n+/// This preserves its same-boot force-kill freeze; persistent shells do not use\n+/// launch age as eligibility (releases#287).\n+// [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\n pub fn launch_predates_boot(launched_ms: u64, boot_ms: u64, slack_ms: u64) -> bool {\n     launched_ms.saturating_add(slack_ms) < boot_ms\n }\n \n-/// Is this instance watcher-eligible? The class-(c) site of\n-/// REQ-HAZARD-SHELL-STALE-ONLINE, ruled onto the corpse-boot discriminant\n-/// (doyle, releases#78 comment 5156810267).\n-///\n-/// It reads the RECORDED status — never the derived one — and then asks the SAME\n-/// `launch_predates_boot` question leg (b) asks. Three arms, and each one is the\n-/// answer to a case the other two get wrong:\n-///\n-/// - **force-killed THIS boot** ⇒ record healed to offline (honest), corpse\n-///   launched after boot ⇒ NOT eligible. This is what the class-(c) freeze\n-///   protects: on Windows an operator kills a shell precisely to free its exe\n-///   for overwrite (shared install dir ⇒ a routine deploy step), and a watcher\n-///   armed here would re-lock the file under them on the next spooled frame.\n-///   Recovery stays demand-driven — `relink`, or a `shell cmd` that wakes.\n-/// - **restart casualty** (corpse launched BEFORE this boot) ⇒ eligible, the\n-///   operator-greenlit scope of #78 leg (b). Two routes, one answer: the boot\n-///   sweep restores it once, and the watcher covers it on demand thereafter.\n-/// - **cleanly closed** (no corpse — `close_shell` retires the pid file) ⇒\n-///   eligible exactly as before this change.\n-///\n-/// WHY THE CORPSE IS DEFINED ON THE PARKED STAMPS and not on a live pid re-probe:\n-/// `shell_pid_provably_dead` runs the PAIR test, so a RECYCLED pid (alive, but\n-/// its native start stamp mismatches the one leg (a) parked) reads as a corpse —\n-/// because our process IS dead. Without that, a race walks straight back into the\n-/// forbidden shape: the heal fires on the true corpse, the OS recycles its pid\n-/// before the next tick, this read sees \"alive ⇒ no corpse\", and a force-killed\n-/// instance is adopted after all.\n-///\n-/// NO STAMP, or no boot oracle on this platform ⇒ NOT eligible when a corpse is\n-/// present: the discriminant cannot be evaluated, and the safe direction is to\n-/// leave a shell down rather than arm a relaunch of a binary an operator may have\n-/// killed on purpose. Same stance, same reason, as the boot sweep's.\n+/// Watchers require an offline instance with no live launch awaiting bind.\n+/// Persistent shells bypass the corpse-age freeze (releases#287); nonpersistent\n+/// corpses still need a dated pre-boot launch. Missing dates fail closed only\n+/// for that nonpersistent freeze, never for persistent restoration.\n // [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\n // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n-fn watcher_eligible(perch: &Path, info: &shellinfo::ShellInfo, boot_ms: Option<u64>) -> bool {\n+fn watcher_eligible(\n+    perch: &Path,\n+    info: &shellinfo::ShellInfo,\n+    persistent: bool,\n+    boot_ms: Option<u64>,\n+) -> bool {\n     if info.status != SHELL_STATUS_OFFLINE {\n         return false;\n     }\n-    if !shellinfo::shell_pid_provably_dead(perch) {\n-        return true; // no corpse: cleanly closed, or never launched\n+    if crate::shellhost::live_launch_winner(perch).is_some() {\n+        return false; // binding: the shell binary and watcher must not overlap\n+    }\n+    if persistent || !shellinfo::shell_pid_provably_dead(perch) {\n+        return true; // persistence overrides the freeze; clean closes need no age\n     }\n     let (Some(boot_ms), Some(launch)) = (boot_ms, shellinfo::read_shell_launch(perch)) else {\n         return false;\n@@ -742,26 +716,10 @@ impl RestoreTrigger {\n /// ONE owner's slice of the persistent-shell restore decision -- the body BOTH\n /// triggers run.\n ///\n-/// It is shared rather than restated because the refusal arms ARE the ruling:\n-/// doyle's condition on releases#228 was that the owner-online trigger carry the\n-/// boot sweep's arms VERBATIM, and two copies of a conjunction drift the first\n-/// time one side is edited. Every gate is a conjunction and each one is\n-/// load-bearing:\n-///\n-/// - the adapter section says `persistent` — the contract promise is scoped to\n-///   exactly those, and a non-persistent instance was never owed a restore;\n-/// - the OWNER endpoint is online — an offline owner is owed nothing, which is\n-///   the difference between healing a record and resurrecting a shell nobody is\n-///   there to drive;\n-/// - the instance is down IN FACT (the derived read, which the birth stamp from\n-///   leg (a) now makes trustworthy across a pid reuse);\n-/// - and its recorded launch PREDATES THE BOOT INSTANT.\n-///\n-/// An instance with NO launch stamp is deliberately NOT restored: the sweep\n-/// cannot prove it predates boot, and the safe direction here is to leave a\n-/// shell down rather than relaunch a binary an operator may have killed on\n-/// purpose. Such an instance gains a stamp the first time it is launched by a\n-/// build carrying leg (a), and is covered from then on.\n+/// Both triggers restore every persistent instance whose owner is online and\n+/// whose binary is down in fact. Birth-safe liveness excludes a live launch\n+/// awaiting bind as well as an already-online shell. Machine boot and launch\n+/// age are irrelevant, including missing stamps (releases#287 operator ruling).\n ///\n /// `set` is the daemon's live watcher set, and it is `None` at boot ON PURPOSE:\n /// the sweep runs before the reconcile loop's first tick, so no watcher can yet\n@@ -773,13 +731,11 @@ impl RestoreTrigger {\n // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n // [impl->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\n // [impl->REQ-SHELL-OWNER-ONLINE-RESTORE]\n-#[allow(clippy::too_many_arguments)]\n fn restore_persistent_shells_of_owner(\n     owlery: &Path,\n     registered: &[(AdapterRecord, Manifest)],\n     adapters_dir: &Path,\n     owner: &str,\n-    boot_ms: u64,\n     trigger: RestoreTrigger,\n     set: Option<&Arc<WakeSet>>,\n     restored: &mut Vec<String>,\n@@ -795,14 +751,10 @@ fn restore_persistent_shells_of_owner(\n             continue;\n         };\n         let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, &shell_id);\n-        if shellinfo::is_shell_online(&perch, &info) {\n-            continue; // already up in fact\n-        }\n-        let Some(launch) = shellinfo::read_shell_launch(&perch) else {\n-            continue; // no stamp ⇒ cannot prove it predates boot ⇒ leave it down\n-        };\n-        if !launch_predates_boot(launch.launched_ms, boot_ms, BOOT_RESTORE_SLACK_MS) {\n-            continue; // launched THIS boot: a force-kill, not a restart casualty\n+        if shellinfo::is_shell_online(&perch, &info)\n+            || crate::shellhost::live_launch_winner(&perch).is_some()\n+        {\n+            continue; // already online or binding in fact\n         }\n         // Mutual exclusivity, the online side: the watcher dies before the binary\n         // rises. `launch_shell` already kills the waker CHILD, but the supervising\n@@ -844,11 +796,6 @@ pub fn restore_persistent_shells_at_boot(\n     registered: &[(AdapterRecord, Manifest)],\n     adapters_dir: &Path,\n ) -> Vec<String> {\n-    let Some(boot_ms) = spt_store::proc::boot_instant_ms() else {\n-        // No boot oracle on this platform: the discriminant cannot be evaluated,\n-        // so nothing is restored. Silence beats guessing at a relaunch.\n-        return Vec::new();\n-    };\n     let mut restored = Vec::new();\n     for owner in spt_store::perch::list_self_perch_ids(owlery) {\n         restore_persistent_shells_of_owner(\n@@ -856,7 +803,6 @@ pub fn restore_persistent_shells_at_boot(\n             registered,\n             adapters_dir,\n             &owner,\n-            boot_ms,\n             RestoreTrigger::Boot,\n             None,\n             &mut restored,\n@@ -868,12 +814,9 @@ pub fn restore_persistent_shells_at_boot(\n /// Per-owner online state carried ACROSS reconcile passes, so the restore below\n /// can be EDGE-triggered on an owner's offline->online transition.\n ///\n-/// WHY AN EDGE AND NOT A LEVEL (doyle's condition, releases#228): a successful\n-/// restore self-limits, because the launch restamps `launched_ms` to now and the\n-/// corpse-boot discriminant refuses that instance from then on. A FAILING launch\n\n[Artifact storage: ~/.omp/agent/sessions/-Documents-projects-spt-core/2026-09-09T23-45-42-317Z_01a08890-21ad-7000-8f54-208d4f4f2f2c/112.bash-original.log (105.3KB). Use artifact://112:N-M to page, artifact://112:raw:N-M for verbatim chunks, and the artifact file path for search/copy workflows.]\n\n[Showing lines 1-300 of 1088. Use :301 to continue]"}],"details":{"resolvedPath":"C:\\Users\\decid\\.omp\\agent\\sessions\\-Documents-projects-spt-core\\2026-09-09T23-45-42-317Z_01a08890-21ad-7000-8f54-208d4f4f2f2c\\112.bash-original.log","contentType":"text/plain","totalLines":1088,"displayContent":{"text":"diff --git a/CHANGELOG.md b/CHANGELOG.md\nindex eafe18d6..39bf8144 100644\n--- a/CHANGELOG.md\n+++ b/CHANGELOG.md\n@@ -10,6 +10,19 @@ breaks something, or changes the observable behavior of existing surfaces\n broadly; **patch** for fixes, and for additive opt-in capability — a new key,\n flag, or page that no existing user can encounter without opting into it.\n \n+## [Unreleased]\n+\n+### Fixed\n+\n+- Every persistent shell comes back with its owner after a daemon-only restart,\n+  not just after a machine reboot. Shells launched during the current boot or\n+  without a saved launch stamp are no longer left down, and recovery notices\n+  name every healed instance for each owner. A persistent shell must be torn\n+  down or made nonpersistent to stay down; killing its process does not opt out\n+  of restoration.\n+- A successful `spt shell relink` reports `binding` while awaiting its handshake,\n+  rather than making the successful launch read as `offline`.\n+\n ## [0.68.0] - 2026-09-08\n \n Web serving. Files, directories, adapter documentation and the changelog are\ndiff --git a/CONTEXT.md b/CONTEXT.md\nindex cdb8a8e1..6698a590 100644\n--- a/CONTEXT.md\n+++ b/CONTEXT.md\n@@ -388,11 +388,11 @@ The manifest value is the **floor**: a node/endpoint setting may **tighten** (de\n \n A shell automatically goes **offline when its owner endpoint goes offline**. Two complementary manifest options bring it (and its owner) back:\n \n-- **`persistent`** (bool): the shell is automatically brought **online whenever its owner endpoint is online**. Covers \"owner already up.\" **Node-qualified once instances span nodes** (ruled 2026-07-24, #70 grill): when an owner comes online at a given node, only the **latest** instances qualified to that *endpoint@node* pairing come online with it — a stale instance of the same shell left on another node does not auto-wake alongside. **A node restart is covered by a boot sweep, and an owner that comes online after that sweep by an owner-online edge** (BAROMETER W2 releases#78; CONDUIT W2 releases#228): a machine death breaks no link, so the instance's record survives the reboot still saying `online` and the wake cascade — which reads the *recorded* field — skips it, stranding every `persistent` instance on the node at once. The daemon therefore (i) heals a recorded `online` its derived status contradicts, every reconcile cycle, writing only when the value actually changes, and (ii) runs a **once-per-daemon-generation boot sweep** that relaunches an instance when **all four** conditions hold: the adapter declares `persistent`; the **owner endpoint is online** (an offline owner is owed nothing); the instance is down **in fact**; and its **recorded launch predates the boot instant**. A launch parks a **birth stamp** beside the pid to make that last question answerable — **no stamp ⇒ not restored** (the sweep cannot prove the corpse predates boot, and leaving a shell down beats relaunching a binary an operator may have killed deliberately), and a platform with no boot oracle restores nothing. A shell force-killed *during* this boot is launched after it, can never satisfy the predicate, and so is never spontaneously relaunched — the deploy/quarantine ruling of KNOWN-HAZARDS 2.6 preserved by construction. <!-- [doc->REQ-SHELL-PERSISTENT-BOOT-RESTORE] -->\n-  **Daemons boot before endpoints do, so the sweep is not the only trigger** (releases#228): an owner that is not yet online when the sweep runs is skipped by that very conjunct, and — the sweep being once per daemon generation, and bringup from offline emitting no rest edge — nothing revisited it, so its shells stayed down until a human relinked. The reconcile loop therefore runs the **same four conditions** for every owner that has gone **offline→online since the last pass**. The trigger is that **edge, never the level “the owner is online”**: a successful restore restamps itself out of eligibility, but a *failing* launch under a level trigger would be retried every reconcile tick for as long as the owner stayed up, where the edge bounds it to one attempt per owner-online event. Re-evaluating later in a generation can only refuse **more**, never less — every launch path restamps the birth stamp, a kill never does, and the boot instant is fixed for the machine boot — so the force-kill ruling above is preserved identically on the second trigger. <!-- [doc->REQ-SHELL-OWNER-ONLINE-RESTORE] -->\n-- **`wake_command`** (template): a long-running **wake-watcher** process spt-core runs *while the shell is offline* **and eligible** (the eligibility rule is below — an offline instance sitting on a same-boot corpse arms no watcher). Its sole job is to fire a wake. It runs on the **shell's node** (where the platform wake-event originates — e.g. interacting with a \"shut-down\" avatar). Covers \"owner is down, wake it from outside.\"\n+- **`persistent`** (bool): the shell is automatically brought **online whenever its owner endpoint is online**. Covers \"owner already up.\" **Node-qualified once instances span nodes** (ruled 2026-07-24, #70 grill): when an owner comes online at a given node, only the **latest** instances qualified to that *endpoint@node* pairing come online with it — a stale instance of the same shell left on another node does not auto-wake alongside. A **once-per-daemon-generation boot sweep** restores every down persistent instance whose owner is online, after a machine reboot **or a daemon-only restart**. Birth-safe liveness protects both an online binary and a live launch awaiting bind. Machine boot time, recorded launch age, and missing launch stamps do **not** gate restoration (releases#287 operator ruling). The daemon also heals stale `online` records to `offline` each reconcile cycle, only on an actual change. <!-- [doc->REQ-SHELL-PERSISTENT-BOOT-RESTORE] -->\n+  **Daemons boot before endpoints do, so the sweep is not the only trigger:** the reconcile loop runs that **same per-owner restore decision** on each owner **offline→online edge**. The edge is seeded before the sweep; a live launch awaiting bind prevents duplicate launch if both triggers reach it. Failed launches are attempted once per owner-online event, not every tick while the owner stays online. A force-kill does not revoke persistence: use `teardown` or a nonpersistent manifest to keep the instance down. <!-- [doc->REQ-SHELL-OWNER-ONLINE-RESTORE] -->\n+- **`wake_command`** (template): a long-running **wake-watcher** process spt-core runs while the shell is offline and eligible. Its sole job is to fire a wake. It runs on the **shell's node**, where the platform wake-event originates (e.g. interacting with a shut-down avatar). Covers \"owner is down, wake it from outside.\"\n \n-**Online/offline are mutually-exclusive processes:** online ⇒ the shell binary runs (no watcher); offline ⇒ the wake-watcher runs (no shell binary). spt-core flips between them. **Offline does not by itself arm the watcher — a third outcome exists** (BAROMETER W2, releases#78): an instance that is offline over a **corpse launched during this boot** (a force-kill, a crash) is deliberately held out of watcher eligibility, so *neither* process runs until recovery is demanded — `relink`, or a `shell cmd` that wakes. Eligibility is recorded-`offline` **AND** (no corpse, **or** the corpse's launch predates the boot instant); an instance whose corpse cannot be dated — no birth stamp, or no boot oracle on the platform — is treated as ineligible rather than guessed at. Only a cleanly-closed instance and a **restart casualty** arm a watcher. The reason is the deploy case: on Windows an operator kills a shell precisely to free its exe for overwrite, and a watcher armed there would re-lock the file under them.\n+**Online/offline are mutually-exclusive processes:** a live shell binary (including a launch awaiting bind) excludes its wake-watcher. A recorded-offline instance is watcher-eligible when it is persistent, carries no corpse, or its corpse's launch predates machine boot. **Only nonpersistent instances retain the same-boot force-kill freeze:** their same-boot or undatable corpses arm no watcher, preserving the deploy/overwrite case. A persistent corpse bypasses that freeze even without a launch stamp or boot oracle (releases#287). Missing birth evidence retains the existing fail-toward-alive process probe; it is not itself a restore refusal.\n \n **Exit-opcode supervision:** the wake-watcher exiting with the **wake opcode** → spt-core runs the wake resolution (below) and brings the shell online. Any *other* exit (crash) → respawn the watcher with **exponential backoff + eventual give-up** (until next shell activity) — crash-bug safety. One watcher per offline shell instance.\n \ndiff --git a/crates/spt-daemon/src/shellhost.rs b/crates/spt-daemon/src/shellhost.rs\nindex 593c6f68..44df752b 100644\n--- a/crates/spt-daemon/src/shellhost.rs\n+++ b/crates/spt-daemon/src/shellhost.rs\n@@ -312,16 +312,14 @@ pub fn routes_through_daemon(shell: &Shell) -> bool {\n /// down and adopt it. Used by BOTH sides of the routed launch, because a late\n /// ack can put the caller and the daemon in the same launch at once.\n ///\n-/// DERIVED, never the recorded status (REQ-HAZARD-SHELL-STALE-ONLINE): a pid\n-/// that is absent, unparseable or `0` reads NOT-a-winner, the same\n-/// fail-toward-alive stance the rest of the shell machinery holds — so a\n-/// backend that records no pid can never block its own relaunch. Same shape and\n-/// same non-fault wording as [`crate::shellwake::resolve_wake`]'s existing\n-/// guard; this is that rule extracted, not a second one.\n+/// Birth-safe, including a binary awaiting bind: a recycled live pid is NOT a\n+/// winner when its native birth stamp mismatches. Missing, unparseable and zero\n+/// pids remain not-a-winner, preserving pid-less backend behavior.\n // [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\n+// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n pub fn live_launch_winner(shell_perch: &Path) -> Option<u32> {\n     shellinfo::read_shell_pid(shell_perch)\n-        .filter(|&p| p != 0 && spt_store::proc::is_process_alive(p))\n+        .filter(|&p| p != 0 && !shellinfo::shell_pid_provably_dead(shell_perch))\n }\n \n /// THE route-selection seam. Generic over the daemon call and the in-process\n@@ -1204,6 +1202,7 @@ mod tests {\n     // that parses non-zero AND probes alive is a winner. Absent / `0` /\n     // unparseable / dead must all read NOT-a-winner, or a backend that records\n     // no usable pid would be unable to relaunch its own instance.\n+    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n     #[test]\n     fn only_a_live_recorded_pid_wins_a_launch_race() {\n         let tmp = tempfile::tempdir().unwrap();\n@@ -1225,6 +1224,18 @@ mod tests {\n             Some(me),\n             \"a recorded pid that probes alive IS the winner\"\n         );\n+        shellinfo::record_shell_launch(perch, me, 1_000);\n+        assert_eq!(live_launch_winner(perch), Some(me), \"same birth retains custody\");\n+        let mut launch = shellinfo::read_shell_launch(perch).unwrap();\n+        launch.pid_started_at = Some(\n+            launch.pid_started_at.expect(\"native birth stamp on supported platforms\") ^ 1,\n+        );\n+        std::fs::write(\n+            perch.join(shellinfo::SHELL_LAUNCH_FILE),\n+            serde_json::to_string(&launch).unwrap(),\n+        )\n+        .unwrap();\n+        assert_eq!(live_launch_winner(perch), None, \"a recycled pid cannot block restore\");\n     }\n \n     // [unit->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE] a REAL dial at a socket name\ndiff --git a/crates/spt-daemon/src/shellwake.rs b/crates/spt-daemon/src/shellwake.rs\nindex 7caf020f..d9631b4d 100644\n--- a/crates/spt-daemon/src/shellwake.rs\n+++ b/crates/spt-daemon/src/shellwake.rs\n@@ -613,7 +613,7 @@ fn fill_wake_command(\n /// correction of a field that is already wrong. Nothing is started here.\n // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n // [impl->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\n-fn heal_stale_online_records(owlery: &Path) {\n+fn heal_stale_online_records(owlery: &Path, events: &mut dyn std::io::Write) {\n     for owner in spt_store::perch::list_self_perch_ids(owlery) {\n         for (shell_id, info) in shellinfo::list_shells(owlery, &owner) {\n             if info.status != SHELL_STATUS_ONLINE {\n@@ -627,7 +627,8 @@ fn heal_stale_online_records(owlery: &Path) {\n             let mut healed = info.clone();\n             healed.status = truth.to_string();\n             if shellinfo::write_shell_info(&perch, &healed).is_ok() {\n-                spt_proto::emit_line_err!(\"SHELL_RECORD_HEALED:{owner}/{shell_id}: online -> {truth}\");\n+                // [impl->REQ-SHELL-HEAL-ENUMERATION]\n+                let _ = spt_proto::emit_line!(events, \"SHELL_RECORD_HEALED:{owner}/{shell_id}: online -> {truth}\");\n             }\n         }\n     }\n@@ -639,67 +640,40 @@ fn heal_stale_online_records(owlery: &Path) {\n /// now-minus-uptime, which drifts by however long the two calls take and by\n /// whatever the clock did since — so a process launched in the first moments of\n /// a boot can compute as very slightly older than the boot itself. Without slack\n-/// that reads as \"a corpse from the previous boot\" and the sweep would relaunch\n-/// something that had just started. Ten seconds is far wider than any plausible\n+/// that reads as \"a corpse from the previous boot\" and a nonpersistent watcher\n+/// could be armed after a same-boot death. Ten seconds is far wider than plausible\n /// derivation error and far narrower than the gap this discriminant exists to\n /// detect (a previous boot is minutes to days back).\n pub const BOOT_RESTORE_SLACK_MS: u64 = 10_000;\n \n-/// Whether a recorded launch predates the current boot — leg (b)'s discriminant\n-/// (releases#78).\n-///\n-/// This is what preserves the force-kill ruling BY CONSTRUCTION rather than by a\n-/// second rule: a shell force-killed during steady-state operation was launched\n-/// AFTER the current boot, so it can never satisfy this predicate and is never\n-/// spontaneously relaunched — the mid-deploy exe-overwrite case that ruling\n-/// protects is untouched. Only a fresh boot restores.\n-///\n-/// Pure over its inputs so the boundary is unit-testable without a reboot.\n-// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n+/// Whether a nonpersistent shell's recorded launch predates the current boot.\n+/// This preserves its same-boot force-kill freeze; persistent shells do not use\n+/// launch age as eligibility (releases#287).\n+// [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\n pub fn launch_predates_boot(launched_ms: u64, boot_ms: u64, slack_ms: u64) -> bool {\n     launched_ms.saturating_add(slack_ms) < boot_ms\n }\n \n-/// Is this instance watcher-eligible? The class-(c) site of\n-/// REQ-HAZARD-SHELL-STALE-ONLINE, ruled onto the corpse-boot discriminant\n-/// (doyle, releases#78 comment 5156810267).\n-///\n-/// It reads the RECORDED status — never the derived one — and then asks the SAME\n-/// `launch_predates_boot` question leg (b) asks. Three arms, and each one is the\n-/// answer to a case the other two get wrong:\n-///\n-/// - **force-killed THIS boot** ⇒ record healed to offline (honest), corpse\n-///   launched after boot ⇒ NOT eligible. This is what the class-(c) freeze\n-///   protects: on Windows an operator kills a shell precisely to free its exe\n-///   for overwrite (shared install dir ⇒ a routine deploy step), and a watcher\n-///   armed here would re-lock the file under them on the next spooled frame.\n-///   Recovery stays demand-driven — `relink`, or a `shell cmd` that wakes.\n-/// - **restart casualty** (corpse launched BEFORE this boot) ⇒ eligible, the\n-///   operator-greenlit scope of #78 leg (b). Two routes, one answer: the boot\n-///   sweep restores it once, and the watcher covers it on demand thereafter.\n-/// - **cleanly closed** (no corpse — `close_shell` retires the pid file) ⇒\n-///   eligible exactly as before this change.\n-///\n-/// WHY THE CORPSE IS DEFINED ON THE PARKED STAMPS and not on a live pid re-probe:\n-/// `shell_pid_provably_dead` runs the PAIR test, so a RECYCLED pid (alive, but\n-/// its native start stamp mismatches the one leg (a) parked) reads as a corpse —\n-/// because our process IS dead. Without that, a race walks straight back into the\n-/// forbidden shape: the heal fires on the true corpse, the OS recycles its pid\n-/// before the next tick, this read sees \"alive ⇒ no corpse\", and a force-killed\n-/// instance is adopted after all.\n-///\n-/// NO STAMP, or no boot oracle on this platform ⇒ NOT eligible when a corpse is\n-/// present: the discriminant cannot be evaluated, and the safe direction is to\n-/// leave a shell down rather than arm a relaunch of a binary an operator may have\n-/// killed on purpose. Same stance, same reason, as the boot sweep's.\n+/// Watchers require an offline instance with no live launch awaiting bind.\n+/// Persistent shells bypass the corpse-age freeze (releases#287); nonpersistent\n+/// corpses still need a dated pre-boot launch. Missing dates fail closed only\n+/// for that nonpersistent freeze, never for persistent restoration.\n // [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\n // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n-fn watcher_eligible(perch: &Path, info: &shellinfo::ShellInfo, boot_ms: Option<u64>) -> bool {\n+fn watcher_eligible(\n+    perch: &Path,\n+    info: &shellinfo::ShellInfo,\n+    persistent: bool,\n+    boot_ms: Option<u64>,\n+) -> bool {\n     if info.status != SHELL_STATUS_OFFLINE {\n         return false;\n     }\n-    if !shellinfo::shell_pid_provably_dead(perch) {\n-        return true; // no corpse: cleanly closed, or never launched\n+    if crate::shellhost::live_launch_winner(perch).is_some() {\n+        return false; // binding: the shell binary and watcher must not overlap\n+    }\n+    if persistent || !shellinfo::shell_pid_provably_dead(perch) {\n+        return true; // persistence overrides the freeze; clean closes need no age\n     }\n     let (Some(boot_ms), Some(launch)) = (boot_ms, shellinfo::read_shell_launch(perch)) else {\n         return false;\n@@ -742,26 +716,10 @@ impl RestoreTrigger {\n /// ONE owner's slice of the persistent-shell restore decision -- the body BOTH\n /// triggers run.\n ///\n-/// It is shared rather than restated because the refusal arms ARE the ruling:\n-/// doyle's condition on releases#228 was that the owner-online trigger carry the\n-/// boot sweep's arms VERBATIM, and two copies of a conjunction drift the first\n-/// time one side is edited. Every gate is a conjunction and each one is\n-/// load-bearing:\n-///\n-/// - the adapter section says `persistent` — the contract promise is scoped to\n-///   exactly those, and a non-persistent instance was never owed a restore;\n-/// - the OWNER endpoint is online — an offline owner is owed nothing, which is\n-///   the difference between healing a record and resurrecting a shell nobody is\n-///   there to drive;\n-/// - the instance is down IN FACT (the derived read, which the birth stamp from\n-///   leg (a) now makes trustworthy across a pid reuse);\n-/// - and its recorded launch PREDATES THE BOOT INSTANT.\n-///\n-/// An instance with NO launch stamp is deliberately NOT restored: the sweep\n-/// cannot prove it predates boot, and the safe direction here is to leave a\n-/// shell down rather than relaunch a binary an operator may have killed on\n-/// purpose. Such an instance gains a stamp the first time it is launched by a\n-/// build carrying leg (a), and is covered from then on.\n+/// Both triggers restore every persistent instance whose owner is online and\n+/// whose binary is down in fact. Birth-safe liveness excludes a live launch\n+/// awaiting bind as well as an already-online shell. Machine boot and launch\n+/// age are irrelevant, including missing stamps (releases#287 operator ruling).\n ///\n /// `set` is the daemon's live watcher set, and it is `None` at boot ON PURPOSE:\n /// the sweep runs before the reconcile loop's first tick, so no watcher can yet\n@@ -773,13 +731,11 @@ impl RestoreTrigger {\n // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n // [impl->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\n // [impl->REQ-SHELL-OWNER-ONLINE-RESTORE]\n-#[allow(clippy::too_many_arguments)]\n fn restore_persistent_shells_of_owner(\n     owlery: &Path,\n     registered: &[(AdapterRecord, Manifest)],\n     adapters_dir: &Path,\n     owner: &str,\n-    boot_ms: u64,\n     trigger: RestoreTrigger,\n     set: Option<&Arc<WakeSet>>,\n     restored: &mut Vec<String>,\n@@ -795,14 +751,10 @@ fn restore_persistent_shells_of_owner(\n             continue;\n         };\n         let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, &shell_id);\n-        if shellinfo::is_shell_online(&perch, &info) {\n-            continue; // already up in fact\n-        }\n-        let Some(launch) = shellinfo::read_shell_launch(&perch) else {\n-            continue; // no stamp ⇒ cannot prove it predates boot ⇒ leave it down\n-        };\n-        if !launch_predates_boot(launch.launched_ms, boot_ms, BOOT_RESTORE_SLACK_MS) {\n-            continue; // launched THIS boot: a force-kill, not a restart casualty\n+        if shellinfo::is_shell_online(&perch, &info)\n+            || crate::shellhost::live_launch_winner(&perch).is_some()\n+        {\n+            continue; // already online or binding in fact\n         }\n         // Mutual exclusivity, the online side: the watcher dies before the binary\n         // rises. `launch_shell` already kills the waker CHILD, but the supervising\n@@ -844,11 +796,6 @@ pub fn restore_persistent_shells_at_boot(\n     registered: &[(AdapterRecord, Manifest)],\n     adapters_dir: &Path,\n ) -> Vec<String> {\n-    let Some(boot_ms) = spt_store::proc::boot_instant_ms() else {\n-        // No boot oracle on this platform: the discriminant cannot be evaluated,\n-        // so nothing is restored. Silence beats guessing at a relaunch.\n-        return Vec::new();\n-    };\n     let mut restored = Vec::new();\n     for owner in spt_store::perch::list_self_perch_ids(owlery) {\n         restore_persistent_shells_of_owner(\n@@ -856,7 +803,6 @@ pub fn restore_persistent_shells_at_boot(\n             registered,\n             adapters_dir,\n             &owner,\n-            boot_ms,\n             RestoreTrigger::Boot,\n             None,\n             &mut restored,\n@@ -868,12 +814,9 @@ pub fn restore_persistent_shells_at_boot(\n /// Per-owner online state carried ACROSS reconcile passes, so the restore below\n /// can be EDGE-triggered on an owner's offline->online transition.\n ///\n-/// WHY AN EDGE AND NOT A LEVEL (doyle's condition, releases#228): a successful\n-/// restore self-limits, because the launch restamps `launched_ms` to now and the\n-/// corpse-boot discriminant refuses that instance from then on. A FAILING launch","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120,121,122,123,124,125,126,127,128,129,130,131,132,133,134,135,136,137,138,139,140,141,142,143,144,145,146,147,148,149,150,151,152,153,154,155,156,157,158,159,160,161,162,163,164,165,166,167,168,169,170,171,172,173,174,175,176,177,178,179,180,181,182,183,184,185,186,187,188,189,190,191,192,193,194,195,196,197,198,199,200,201,202,203,204,205,206,207,208,209,210,211,212,213,214,215,216,217,218,219,220,221,222,223,224,225,226,227,228,229,230,231,232,233,234,235,236,237,238,239,240,241,242,243,244,245,246,247,248,249,250,251,252,253,254,255,256,257,258,259,260,261,262,263,264,265,266,267,268,269,270,271,272,273,274,275,276,277,278,279,280,281,282,283,284,285,286,287,288,289,290,291,292,293,294,295,296,297,298,299,300]},"truncation":{"content":"diff --git a/CHANGELOG.md b/CHANGELOG.md\nindex eafe18d6..39bf8144 100644\n--- a/CHANGELOG.md\n+++ b/CHANGELOG.md\n@@ -10,6 +10,19 @@ breaks something, or changes the observable behavior of existing surfaces\n broadly; **patch** for fixes, and for additive opt-in capability — a new key,\n flag, or page that no existing user can encounter without opting into it.\n \n+## [Unreleased]\n+\n+### Fixed\n+\n+- Every persistent shell comes back with its owner after a daemon-only restart,\n+  not just after a machine reboot. Shells launched during the current boot or\n+  without a saved launch stamp are no longer left down, and recovery notices\n+  name every healed instance for each owner. A persistent shell must be torn\n+  down or made nonpersistent to stay down; killing its process does not opt out\n+  of restoration.\n+- A successful `spt shell relink` reports `binding` while awaiting its handshake,\n+  rather than making the successful launch read as `offline`.\n+\n ## [0.68.0] - 2026-09-08\n \n Web serving. Files, directories, adapter documentation and the changelog are\ndiff --git a/CONTEXT.md b/CONTEXT.md\nindex cdb8a8e1..6698a590 100644\n--- a/CONTEXT.md\n+++ b/CONTEXT.md\n@@ -388,11 +388,11 @@ The manifest value is the **floor**: a node/endpoint setting may **tighten** (de\n \n A shell automatically goes **offline when its owner endpoint goes offline**. Two complementary manifest options bring it (and its owner) back:\n \n-- **`persistent`** (bool): the shell is automatically brought **online whenever its owner endpoint is online**. Covers \"owner already up.\" **Node-qualified once instances span nodes** (ruled 2026-07-24, #70 grill): when an owner comes online at a given node, only the **latest** instances qualified to that *endpoint@node* pairing come online with it — a stale instance of the same shell left on another node does not auto-wake alongside. **A node restart is covered by a boot sweep, and an owner that comes online after that sweep by an owner-online edge** (BAROMETER W2 releases#78; CONDUIT W2 releases#228): a machine death breaks no link, so the instance's record survives the reboot still saying `online` and the wake cascade — which reads the *recorded* field — skips it, stranding every `persistent` instance on the node at once. The daemon therefore (i) heals a recorded `online` its derived status contradicts, every reconcile cycle, writing only when the value actually changes, and (ii) runs a **once-per-daemon-generation boot sweep** that relaunches an instance when **all four** conditions hold: the adapter declares `persistent`; the **owner endpoint is online** (an offline owner is owed nothing); the instance is down **in fact**; and its **recorded launch predates the boot instant**. A launch parks a **birth stamp** beside the pid to make that last question answerable — **no stamp ⇒ not restored** (the sweep cannot prove the corpse predates boot, and leaving a shell down beats relaunching a binary an operator may have killed deliberately), and a platform with no boot oracle restores nothing. A shell force-killed *during* this boot is launched after it, can never satisfy the predicate, and so is never spontaneously relaunched — the deploy/quarantine ruling of KNOWN-HAZARDS 2.6 preserved by construction. <!-- [doc->REQ-SHELL-PERSISTENT-BOOT-RESTORE] -->\n-  **Daemons boot before endpoints do, so the sweep is not the only trigger** (releases#228): an owner that is not yet online when the sweep runs is skipped by that very conjunct, and — the sweep being once per daemon generation, and bringup from offline emitting no rest edge — nothing revisited it, so its shells stayed down until a human relinked. The reconcile loop therefore runs the **same four conditions** for every owner that has gone **offline→online since the last pass**. The trigger is that **edge, never the level “the owner is online”**: a successful restore restamps itself out of eligibility, but a *failing* launch under a level trigger would be retried every reconcile tick for as long as the owner stayed up, where the edge bounds it to one attempt per owner-online event. Re-evaluating later in a generation can only refuse **more**, never less — every launch path restamps the birth stamp, a kill never does, and the boot instant is fixed for the machine boot — so the force-kill ruling above is preserved identically on the second trigger. <!-- [doc->REQ-SHELL-OWNER-ONLINE-RESTORE] -->\n-- **`wake_command`** (template): a long-running **wake-watcher** process spt-core runs *while the shell is offline* **and eligible** (the eligibility rule is below — an offline instance sitting on a same-boot corpse arms no watcher). Its sole job is to fire a wake. It runs on the **shell's node** (where the platform wake-event originates — e.g. interacting with a \"shut-down\" avatar). Covers \"owner is down, wake it from outside.\"\n+- **`persistent`** (bool): the shell is automatically brought **online whenever its owner endpoint is online**. Covers \"owner already up.\" **Node-qualified once instances span nodes** (ruled 2026-07-24, #70 grill): when an owner comes online at a given node, only the **latest** instances qualified to that *endpoint@node* pairing come online with it — a stale instance of the same shell left on another node does not auto-wake alongside. A **once-per-daemon-generation boot sweep** restores every down persistent instance whose owner is online, after a machine reboot **or a daemon-only restart**. Birth-safe liveness protects both an online binary and a live launch awaiting bind. Machine boot time, recorded launch age, and missing launch stamps do **not** gate restoration (releases#287 operator ruling). The daemon also heals stale `online` records to `offline` each reconcile cycle, only on an actual change. <!-- [doc->REQ-SHELL-PERSISTENT-BOOT-RESTORE] -->\n+  **Daemons boot before endpoints do, so the sweep is not the only trigger:** the reconcile loop runs that **same per-owner restore decision** on each owner **offline→online edge**. The edge is seeded before the sweep; a live launch awaiting bind prevents duplicate launch if both triggers reach it. Failed launches are attempted once per owner-online event, not every tick while the owner stays online. A force-kill does not revoke persistence: use `teardown` or a nonpersistent manifest to keep the instance down. <!-- [doc->REQ-SHELL-OWNER-ONLINE-RESTORE] -->\n+- **`wake_command`** (template): a long-running **wake-watcher** process spt-core runs while the shell is offline and eligible. Its sole job is to fire a wake. It runs on the **shell's node**, where the platform wake-event originates (e.g. interacting with a shut-down avatar). Covers \"owner is down, wake it from outside.\"\n \n-**Online/offline are mutually-exclusive processes:** online ⇒ the shell binary runs (no watcher); offline ⇒ the wake-watcher runs (no shell binary). spt-core flips between them. **Offline does not by itself arm the watcher — a third outcome exists** (BAROMETER W2, releases#78): an instance that is offline over a **corpse launched during this boot** (a force-kill, a crash) is deliberately held out of watcher eligibility, so *neither* process runs until recovery is demanded — `relink`, or a `shell cmd` that wakes. Eligibility is recorded-`offline` **AND** (no corpse, **or** the corpse's launch predates the boot instant); an instance whose corpse cannot be dated — no birth stamp, or no boot oracle on the platform — is treated as ineligible rather than guessed at. Only a cleanly-closed instance and a **restart casualty** arm a watcher. The reason is the deploy case: on Windows an operator kills a shell precisely to free its exe for overwrite, and a watcher armed there would re-lock the file under them.\n+**Online/offline are mutually-exclusive processes:** a live shell binary (including a launch awaiting bind) excludes its wake-watcher. A recorded-offline instance is watcher-eligible when it is persistent, carries no corpse, or its corpse's launch predates machine boot. **Only nonpersistent instances retain the same-boot force-kill freeze:** their same-boot or undatable corpses arm no watcher, preserving the deploy/overwrite case. A persistent corpse bypasses that freeze even without a launch stamp or boot oracle (releases#287). Missing birth evidence retains the existing fail-toward-alive process probe; it is not itself a restore refusal.\n \n **Exit-opcode supervision:** the wake-watcher exiting with the **wake opcode** → spt-core runs the wake resolution (below) and brings the shell online. Any *other* exit (crash) → respawn the watcher with **exponential backoff + eventual give-up** (until next shell activity) — crash-bug safety. One watcher per offline shell instance.\n \ndiff --git a/crates/spt-daemon/src/shellhost.rs b/crates/spt-daemon/src/shellhost.rs\nindex 593c6f68..44df752b 100644\n--- a/crates/spt-daemon/src/shellhost.rs\n+++ b/crates/spt-daemon/src/shellhost.rs\n@@ -312,16 +312,14 @@ pub fn routes_through_daemon(shell: &Shell) -> bool {\n /// down and adopt it. Used by BOTH sides of the routed launch, because a late\n /// ack can put the caller and the daemon in the same launch at once.\n ///\n-/// DERIVED, never the recorded status (REQ-HAZARD-SHELL-STALE-ONLINE): a pid\n-/// that is absent, unparseable or `0` reads NOT-a-winner, the same\n-/// fail-toward-alive stance the rest of the shell machinery holds — so a\n-/// backend that records no pid can never block its own relaunch. Same shape and\n-/// same non-fault wording as [`crate::shellwake::resolve_wake`]'s existing\n-/// guard; this is that rule extracted, not a second one.\n+/// Birth-safe, including a binary awaiting bind: a recycled live pid is NOT a\n+/// winner when its native birth stamp mismatches. Missing, unparseable and zero\n+/// pids remain not-a-winner, preserving pid-less backend behavior.\n // [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\n+// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n pub fn live_launch_winner(shell_perch: &Path) -> Option<u32> {\n     shellinfo::read_shell_pid(shell_perch)\n-        .filter(|&p| p != 0 && spt_store::proc::is_process_alive(p))\n+        .filter(|&p| p != 0 && !shellinfo::shell_pid_provably_dead(shell_perch))\n }\n \n /// THE route-selection seam. Generic over the daemon call and the in-process\n@@ -1204,6 +1202,7 @@ mod tests {\n     // that parses non-zero AND probes alive is a winner. Absent / `0` /\n     // unparseable / dead must all read NOT-a-winner, or a backend that records\n     // no usable pid would be unable to relaunch its own instance.\n+    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n     #[test]\n     fn only_a_live_recorded_pid_wins_a_launch_race() {\n         let tmp = tempfile::tempdir().unwrap();\n@@ -1225,6 +1224,18 @@ mod tests {\n             Some(me),\n             \"a recorded pid that probes alive IS the winner\"\n         );\n+        shellinfo::record_shell_launch(perch, me, 1_000);\n+        assert_eq!(live_launch_winner(perch), Some(me), \"same birth retains custody\");\n+        let mut launch = shellinfo::read_shell_launch(perch).unwrap();\n+        launch.pid_started_at = Some(\n+            launch.pid_started_at.expect(\"native birth stamp on supported platforms\") ^ 1,\n+        );\n+        std::fs::write(\n+            perch.join(shellinfo::SHELL_LAUNCH_FILE),\n+            serde_json::to_string(&launch).unwrap(),\n+        )\n+        .unwrap();\n+        assert_eq!(live_launch_winner(perch), None, \"a recycled pid cannot block restore\");\n     }\n \n     // [unit->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE] a REAL dial at a socket name\ndiff --git a/crates/spt-daemon/src/shellwake.rs b/crates/spt-daemon/src/shellwake.rs\nindex 7caf020f..d9631b4d 100644\n--- a/crates/spt-daemon/src/shellwake.rs\n+++ b/crates/spt-daemon/src/shellwake.rs\n@@ -613,7 +613,7 @@ fn fill_wake_command(\n /// correction of a field that is already wrong. Nothing is started here.\n // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n // [impl->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\n-fn heal_stale_online_records(owlery: &Path) {\n+fn heal_stale_online_records(owlery: &Path, events: &mut dyn std::io::Write) {\n     for owner in spt_store::perch::list_self_perch_ids(owlery) {\n         for (shell_id, info) in shellinfo::list_shells(owlery, &owner) {\n             if info.status != SHELL_STATUS_ONLINE {\n@@ -627,7 +627,8 @@ fn heal_stale_online_records(owlery: &Path) {\n             let mut healed = info.clone();\n             healed.status = truth.to_string();\n             if shellinfo::write_shell_info(&perch, &healed).is_ok() {\n-                spt_proto::emit_line_err!(\"SHELL_RECORD_HEALED:{owner}/{shell_id}: online -> {truth}\");\n+                // [impl->REQ-SHELL-HEAL-ENUMERATION]\n+                let _ = spt_proto::emit_line!(events, \"SHELL_RECORD_HEALED:{owner}/{shell_id}: online -> {truth}\");\n             }\n         }\n     }\n@@ -639,67 +640,40 @@ fn heal_stale_online_records(owlery: &Path) {\n /// now-minus-uptime, which drifts by however long the two calls take and by\n /// whatever the clock did since — so a process launched in the first moments of\n /// a boot can compute as very slightly older than the boot itself. Without slack\n-/// that reads as \"a corpse from the previous boot\" and the sweep would relaunch\n-/// something that had just started. Ten seconds is far wider than any plausible\n+/// that reads as \"a corpse from the previous boot\" and a nonpersistent watcher\n+/// could be armed after a same-boot death. Ten seconds is far wider than plausible\n /// derivation error and far narrower than the gap this discriminant exists to\n /// detect (a previous boot is minutes to days back).\n pub const BOOT_RESTORE_SLACK_MS: u64 = 10_000;\n \n-/// Whether a recorded launch predates the current boot — leg (b)'s discriminant\n-/// (releases#78).\n-///\n-/// This is what preserves the force-kill ruling BY CONSTRUCTION rather than by a\n-/// second rule: a shell force-killed during steady-state operation was launched\n-/// AFTER the current boot, so it can never satisfy this predicate and is never\n-/// spontaneously relaunched — the mid-deploy exe-overwrite case that ruling\n-/// protects is untouched. Only a fresh boot restores.\n-///\n-/// Pure over its inputs so the boundary is unit-testable without a reboot.\n-// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n+/// Whether a nonpersistent shell's recorded launch predates the current boot.\n+/// This preserves its same-boot force-kill freeze; persistent shells do not use\n+/// launch age as eligibility (releases#287).\n+// [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\n pub fn launch_predates_boot(launched_ms: u64, boot_ms: u64, slack_ms: u64) -> bool {\n     launched_ms.saturating_add(slack_ms) < boot_ms\n }\n \n-/// Is this instance watcher-eligible? The class-(c) site of\n-/// REQ-HAZARD-SHELL-STALE-ONLINE, ruled onto the corpse-boot discriminant\n-/// (doyle, releases#78 comment 5156810267).\n-///\n-/// It reads the RECORDED status — never the derived one — and then asks the SAME\n-/// `launch_predates_boot` question leg (b) asks. Three arms, and each one is the\n-/// answer to a case the other two get wrong:\n-///\n-/// - **force-killed THIS boot** ⇒ record healed to offline (honest), corpse\n-///   launched after boot ⇒ NOT eligible. This is what the class-(c) freeze\n-///   protects: on Windows an operator kills a shell precisely to free its exe\n-///   for overwrite (shared install dir ⇒ a routine deploy step), and a watcher\n-///   armed here would re-lock the file under them on the next spooled frame.\n-///   Recovery stays demand-driven — `relink`, or a `shell cmd` that wakes.\n-/// - **restart casualty** (corpse launched BEFORE this boot) ⇒ eligible, the\n-///   operator-greenlit scope of #78 leg (b). Two routes, one answer: the boot\n-///   sweep restores it once, and the watcher covers it on demand thereafter.\n-/// - **cleanly closed** (no corpse — `close_shell` retires the pid file) ⇒\n-///   eligible exactly as before this change.\n-///\n-/// WHY THE CORPSE IS DEFINED ON THE PARKED STAMPS and not on a live pid re-probe:\n-/// `shell_pid_provably_dead` runs the PAIR test, so a RECYCLED pid (alive, but\n-/// its native start stamp mismatches the one leg (a) parked) reads as a corpse —\n-/// because our process IS dead. Without that, a race walks straight back into the\n-/// forbidden shape: the heal fires on the true corpse, the OS recycles its pid\n-/// before the next tick, this read sees \"alive ⇒ no corpse\", and a force-killed\n-/// instance is adopted after all.\n-///\n-/// NO STAMP, or no boot oracle on this platform ⇒ NOT eligible when a corpse is\n-/// present: the discriminant cannot be evaluated, and the safe direction is to\n-/// leave a shell down rather than arm a relaunch of a binary an operator may have\n-/// killed on purpose. Same stance, same reason, as the boot sweep's.\n+/// Watchers require an offline instance with no live launch awaiting bind.\n+/// Persistent shells bypass the corpse-age freeze (releases#287); nonpersistent\n+/// corpses still need a dated pre-boot launch. Missing dates fail closed only\n+/// for that nonpersistent freeze, never for persistent restoration.\n // [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\n // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n-fn watcher_eligible(perch: &Path, info: &shellinfo::ShellInfo, boot_ms: Option<u64>) -> bool {\n+fn watcher_eligible(\n+    perch: &Path,\n+    info: &shellinfo::ShellInfo,\n+    persistent: bool,\n+    boot_ms: Option<u64>,\n+) -> bool {\n     if info.status != SHELL_STATUS_OFFLINE {\n         return false;\n     }\n-    if !shellinfo::shell_pid_provably_dead(perch) {\n-        return true; // no corpse: cleanly closed, or never launched\n+    if crate::shellhost::live_launch_winner(perch).is_some() {\n+        return false; // binding: the shell binary and watcher must not overlap\n+    }\n+    if persistent || !shellinfo::shell_pid_provably_dead(perch) {\n+        return true; // persistence overrides the freeze; clean closes need no age\n     }\n     let (Some(boot_ms), Some(launch)) = (boot_ms, shellinfo::read_shell_launch(perch)) else {\n         return false;\n@@ -742,26 +716,10 @@ impl RestoreTrigger {\n /// ONE owner's slice of the persistent-shell restore decision -- the body BOTH\n /// triggers run.\n ///\n-/// It is shared rather than restated because the refusal arms ARE the ruling:\n-/// doyle's condition on releases#228 was that the owner-online trigger carry the\n-/// boot sweep's arms VERBATIM, and two copies of a conjunction drift the first\n-/// time one side is edited. Every gate is a conjunction and each one is\n-/// load-bearing:\n-///\n-/// - the adapter section says `persistent` — the contract promise is scoped to\n-///   exactly those, and a non-persistent instance was never owed a restore;\n-/// - the OWNER endpoint is online — an offline owner is owed nothing, which is\n-///   the difference between healing a record and resurrecting a shell nobody is\n-///   there to drive;\n-/// - the instance is down IN FACT (the derived read, which the birth stamp from\n-///   leg (a) now makes trustworthy across a pid reuse);\n-/// - and its recorded launch PREDATES THE BOOT INSTANT.\n-///\n-/// An instance with NO launch stamp is deliberately NOT restored: the sweep\n-/// cannot prove it predates boot, and the safe direction here is to leave a\n-/// shell down rather than relaunch a binary an operator may have killed on\n-/// purpose. Such an instance gains a stamp the first time it is launched by a\n-/// build carrying leg (a), and is covered from then on.\n+/// Both triggers restore every persistent instance whose owner is online and\n+/// whose binary is down in fact. Birth-safe liveness excludes a live launch\n+/// awaiting bind as well as an already-online shell. Machine boot and launch\n+/// age are irrelevant, including missing stamps (releases#287 operator ruling).\n ///\n /// `set` is the daemon's live watcher set, and it is `None` at boot ON PURPOSE:\n /// the sweep runs before the reconcile loop's first tick, so no watcher can yet\n@@ -773,13 +731,11 @@ impl RestoreTrigger {\n // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n // [impl->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\n // [impl->REQ-SHELL-OWNER-ONLINE-RESTORE]\n-#[allow(clippy::too_many_arguments)]\n fn restore_persistent_shells_of_owner(\n     owlery: &Path,\n     registered: &[(AdapterRecord, Manifest)],\n     adapters_dir: &Path,\n     owner: &str,\n-    boot_ms: u64,\n     trigger: RestoreTrigger,\n     set: Option<&Arc<WakeSet>>,\n     restored: &mut Vec<String>,\n@@ -795,14 +751,10 @@ fn restore_persistent_shells_of_owner(\n             continue;\n         };\n         let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, &shell_id);\n-        if shellinfo::is_shell_online(&perch, &info) {\n-            continue; // already up in fact\n-        }\n-        let Some(launch) = shellinfo::read_shell_launch(&perch) else {\n-            continue; // no stamp ⇒ cannot prove it predates boot ⇒ leave it down\n-        };\n-        if !launch_predates_boot(launch.launched_ms, boot_ms, BOOT_RESTORE_SLACK_MS) {\n-            continue; // launched THIS boot: a force-kill, not a restart casualty\n+        if shellinfo::is_shell_online(&perch, &info)\n+            || crate::shellhost::live_launch_winner(&perch).is_some()\n+        {\n+            continue; // already online or binding in fact\n         }\n         // Mutual exclusivity, the online side: the watcher dies before the binary\n         // rises. `launch_shell` already kills the waker CHILD, but the supervising\n@@ -844,11 +796,6 @@ pub fn restore_persistent_shells_at_boot(\n     registered: &[(AdapterRecord, Manifest)],\n     adapters_dir: &Path,\n ) -> Vec<String> {\n-    let Some(boot_ms) = spt_store::proc::boot_instant_ms() else {\n-        // No boot oracle on this platform: the discriminant cannot be evaluated,\n-        // so nothing is restored. Silence beats guessing at a relaunch.\n-        return Vec::new();\n-    };\n     let mut restored = Vec::new();\n     for owner in spt_store::perch::list_self_perch_ids(owlery) {\n         restore_persistent_shells_of_owner(\n@@ -856,7 +803,6 @@ pub fn restore_persistent_shells_at_boot(\n             registered,\n             adapters_dir,\n             &owner,\n-            boot_ms,\n             RestoreTrigger::Boot,\n             None,\n             &mut restored,\n@@ -868,12 +814,9 @@ pub fn restore_persistent_shells_at_boot(\n /// Per-owner online state carried ACROSS reconcile passes, so the restore below\n /// can be EDGE-triggered on an owner's offline->online transition.\n ///\n-/// WHY AN EDGE AND NOT A LEVEL (doyle's condition, releases#228): a successful\n-/// restore self-limits, because the launch restamps `launched_ms` to now and the\n-/// corpse-boot discriminant refuses that instance from then on. A FAILING launch","truncated":true,"truncatedBy":"lines","totalLines":1088,"totalBytes":22745,"outputLines":300,"outputBytes":22745,"lastLinePartial":false,"firstLineExceedsLimit":false},"meta":{"source":{"type":"internal","value":"artifact://112"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":1088,"totalBytes":22745,"outputLines":300,"outputBytes":22745,"maxBytes":153600,"shownRange":{"start":1,"end":300},"nextOffset":301}}},"isError":false,"timestamp":1789001101754}}
{"type":"message","id":"39a9a06d","parentId":"8abef4f6","timestamp":"2026-09-10T00:45:01.755Z","message":{"role":"toolResult","toolCallId":"call_SoXnwtS6OxjKMOFNagDliZYD|fc_077f3569365e313d016aa1fd8dabb487d19e2dbb48a2ab8749","toolName":"read","content":[{"type":"text","text":"@AGENTS.md\n\n<!--\nClaude Code reads CLAUDE.md only; the `@AGENTS.md` import above bridges the canonical rules.\nKeep working rules in AGENTS.md (harness-agnostic). Put ONLY Claude-Code-specific extras\nbelow this line (e.g. CC slash-command notes, hook reminders) — and prefer none.\n-->"}],"details":{"totalLines":7,"displayContent":{"text":"@AGENTS.md\r\n\r\n<!--\r\nClaude Code reads CLAUDE.md only; the `@AGENTS.md` import above bridges the canonical rules.\r\nKeep working rules in AGENTS.md (harness-agnostic). Put ONLY Claude-Code-specific extras\r\nbelow this line (e.g. CC slash-command notes, hook reminders) — and prefer none.\r\n-->\r","startLine":1,"lineNumbers":[1,2,3,4,5,6,7]},"fileSize":293,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\CLAUDE.md"}}},"isError":false,"timestamp":1789001101755}}
{"type":"message","id":"b4b131f7","parentId":"39a9a06d","timestamp":"2026-09-10T00:45:01.758Z","message":{"role":"toolResult","toolCallId":"call_8hdpFa9bn2kiSxq2op7zz0Ps|fc_077f3569365e313d016aa1fd8dabbc87d19b254ad7b50831fe","toolName":"read","content":[{"type":"text","text":"# spt-core — agent working rules\n\nHarness-independent core for an agent ecosystem (messaging, live-agent lifecycle, terminal hosting, P2P networking, runtime-manifest harness contract). Rust crates + a single `spt` binary. Clean-room rebuild of `claude_skill_owl`.\n\n**Orientation — read before working:** `PRD.md` (requirements), `ROADMAP.md` (build path), `CONTEXT.md` (glossary/model, authoritative for meaning), `docs/adr/` (decisions), `docs/KNOWN-HAZARDS.md` (invariants we must not re-break), `docs/{STORAGE,MANIFEST,CONTEXT-MEMORY,DOCS-STRATEGY,TRACEABILITY}.md`.\n\nFor maintainer/debug update work, also read `docs/DEBUG-ROLLOUT.md` before touching the update-set or rollout path.\n\nFor work needing Linux tests and proof, this machine is authorized to ssh into a Linux box: `reavus@kitsubito`\n\nIMPORTANT: All agents you work alongside run on spt-core itself — the legacy-SPT (`$OWL`/`$LIVE`) era ended with the 2026-07 fleet migration; those commands are gone. Reach other agents with `spt send` (or the `/sptc:send` skill); keep message bodies short, and send long bodies from a file rather than inline. Everything else can be learned from the corresponding `/sptc:` skill doc.\n\n## Requirement traceability (binding)\n\nThis project uses `traceable-reqs` (`traceable-reqs.toml` = the authoritative `REQ-*` registry). The full contract is `docs/TRACEABILITY.md`. The rules you must follow:\n\n1. **Tag evidence in the same change.** When you write a function/test/doc-section that satisfies a requirement stage, add its tag in that same commit:\n   `// [impl->REQ-FOO]` · `// [unit->REQ-FOO]` · `<!-- [doc->REQ-FOO] -->`\n   Stages: `doc` / `impl` / `unit` / `int`. Tag *on or immediately above* the real evidence — never at file tops to satisfy coverage.\n2. **Run `traceable-reqs check` before declaring work done.** Exit-1 means missing/invalid evidence — fix it, don't ship it.\n3. **New requirement → add it to `traceable-reqs.toml` first** (with a `REQ-*` id), then satisfy it. No untracked work; no untagged evidence.\n4. **KNOWN-HAZARDS are `REQ-HAZARD-*` requirements** — each needs a test before it's \"covered.\" Treat the hazard list as a conformance checklist you must satisfy, not advice.\n5. **Activate, don't pre-fail.** Requirements you aren't yet working stay `required_stages = []`. Activate (set real stages) only when starting the milestone that delivers them.\n\n## Issue tracking (unique case)\n\nCodebase issues are maintained on **`BigscreenVR/spt-bs-releases`** — the less-private *releases* repo — NOT on `spt-bs-core` where the code lives. This is a deliberate exception (operator-ruled 2026-07-27): file, triage, and reference issues against `spt-bs-releases`.\n\nRequests and milestones are tracked through the **alchemy shell** (`alchemy-0`, the Requests shell bound to `spt-bs-releases`) on top of those GitHub issues. The label taxonomy is the workflow: `state:` BACKLOG → EVAL → GREENLIT → WIP → ACCEPTANCE → DONE (plus CUT / SHELVED), `type:` BUGFIX / ADDITION / CHANGE, `flag: NEEDS-OPERATOR`, and per-node `kind: SPT-TRACKING` pins (marked DO NOT DELETE — leave them alone).\n\n## Other conventions\n\n- Match surrounding code style; copy-verbatim the sister project's stable wire/schema formats (ADR-0001), clean-room everything else.\n- Honor every `docs/KNOWN-HAZARDS.md` invariant — these are real bugs already paid for once.\n- Docs are dual-audience (human + AI dev-agent) per `docs/DOCS-STRATEGY.md`; doc generation is CI-gated against drift.\n- Commit messages end with the project's co-author trailer, exact spelling `Co-authored by: <your live agent name>` (e.g. `Co-authored by: todlando`) — NOT git's hyphenated `Co-Authored-By:`; CI's runner-informant parses only the space spelling for attribution.\n  - **Audit this trailer with `git log --grep` or a raw body read, NEVER `%(trailers:)`** (nor `interpret-trailers --parse`): git's trailer tokenizer requires a whitespace-free token, so the space spelling is structurally invisible to it and every trailer query returns a confident EMPTY on every commit — a clean zero that reads as total attribution loss. Measured 2026-07-30 (main @`8f3e10b`, Signed-off-by and hyphenated-key controls both parse); re-derived at full cost 2026-08-03 by one of its own original measurers, which is why this sentence lives here at the mandate rather than in a document you must decide to open.\n- **Worktrees (binding, operator-ruled 2026-07-05):** every `git worktree` (gate rigs, agent scratch, anything) goes under `.worktrees/` in THIS project's root (`.gitignore`d) — e.g. `git worktree add --detach .worktrees/gate-<sha> <sha>`. NEVER create worktrees as siblings of the project dir; stray sibling folders accumulate orphan-pinned skeletons and pollute `~/Documents/projects`. Remove yours when done (`git worktree remove` + prune; retry later if a handle pins it).\n<!-- [doc->REQ-POOL-FOREIGN-TREE-REFUSED] -->\n- **Build-cache pools (binding, releases#103):** a cargo `target/` directory belongs to ONE source tree. Two trees building into one pool overwrite each other's artifacts and fabricate reds that belong to neither — the v0.51.0 W1/H1 `E0308` that vanished on re-run. **The rule is about the RELATIONSHIP, not the path:** a junction and a `CARGO_TARGET_DIR` env var reach the same hazard, and the env var leaves no directory entry to notice afterwards, so \"is this dir under `.worktrees/`\" never was the question. A build script enforces it (`crates/spt-store/build.rs`) and refuses with `SPT_POOL_FOREIGN`. **Sharing a pool SEQUENTIALLY is fine and expected** — a finished lane is taken over loudly, not refused; only a LIVE lane blocks you. Claim a pool at lane start …\n- **Target-tree teardown (releases#103, same lane):** never delete a `target/` you have not classified. Check OUTBOUND first (`Get-Item -Force`: a junction is deleted as a link, a real dir as a tree), then sweep INBOUND for reparse points aimed at it before removing a real one, and measure size and free space before and after so the reclaim is a number rather than a claim. Reap the target SUBTREE, never a session or worktree directory around it. Do not point `CARGO_TARGET_DIR` into a scratchpad (v0.50.0 incident) or at another checkout's target (H1 incident).\n- **CI / merge strategy:** golden CI (ADR-0050, ruled 2026-07-27) — milestone-batch golden runs, thin-lane PRs, ff-only main (tested sha == merged sha). Piloting from the next milestone; read the ADR before shaping a PR lane.\n\nIf you finish a significant body of work without need for user intervention, or if your context gets too high, you can clear your own context and keep moving:\n1. Create a JIT plan for the next immediate body of work, if it isn't already planned\n2. If you are a live agent, /sptc:commune with immediate next steps and broad summary of the project's status + end goal\n3. Execute the /self-clear skill\n\n[Some lines truncated to 768 chars]"}],"details":{"totalLines":47,"displayContent":{"text":"# spt-core — agent working rules\r\n\r\nHarness-independent core for an agent ecosystem (messaging, live-agent lifecycle, terminal hosting, P2P networking, runtime-manifest harness contract). Rust crates + a single `spt` binary. Clean-room rebuild of `claude_skill_owl`.\r\n\r\n**Orientation — read before working:** `PRD.md` (requirements), `ROADMAP.md` (build path), `CONTEXT.md` (glossary/model, authoritative for meaning), `docs/adr/` (decisions), `docs/KNOWN-HAZARDS.md` (invariants we must not re-break), `docs/{STORAGE,MANIFEST,CONTEXT-MEMORY,DOCS-STRATEGY,TRACEABILITY}.md`.\r\n\r\nFor maintainer/debug update work, also read `docs/DEBUG-ROLLOUT.md` before touching the update-set or rollout path.\r\n\r\nFor work needing Linux tests and proof, this machine is authorized to ssh into a Linux box: `reavus@kitsubito`\r\n\r\nIMPORTANT: All agents you work alongside run on spt-core itself — the legacy-SPT (`$OWL`/`$LIVE`) era ended with the 2026-07 fleet migration; those commands are gone. Reach other agents with `spt send` (or the `/sptc:send` skill); keep message bodies short, and send long bodies from a file rather than inline. Everything else can be learned from the corresponding `/sptc:` skill doc.\r\n\r\n## Requirement traceability (binding)\r\n\r\nThis project uses `traceable-reqs` (`traceable-reqs.toml` = the authoritative `REQ-*` registry). The full contract is `docs/TRACEABILITY.md`. The rules you must follow:\r\n\r\n1. **Tag evidence in the same change.** When you write a function/test/doc-section that satisfies a requirement stage, add its tag in that same commit:\r\n   `// [impl->REQ-FOO]` · `// [unit->REQ-FOO]` · `<!-- [doc->REQ-FOO] -->`\r\n   Stages: `doc` / `impl` / `unit` / `int`. Tag *on or immediately above* the real evidence — never at file tops to satisfy coverage.\r\n2. **Run `traceable-reqs check` before declaring work done.** Exit-1 means missing/invalid evidence — fix it, don't ship it.\r\n3. **New requirement → add it to `traceable-reqs.toml` first** (with a `REQ-*` id), then satisfy it. No untracked work; no untagged evidence.\r\n4. **KNOWN-HAZARDS are `REQ-HAZARD-*` requirements** — each needs a test before it's \"covered.\" Treat the hazard list as a conformance checklist you must satisfy, not advice.\r\n5. **Activate, don't pre-fail.** Requirements you aren't yet working stay `required_stages = []`. Activate (set real stages) only when starting the milestone that delivers them.\r\n\r\n## Issue tracking (unique case)\r\n\r\nCodebase issues are maintained on **`BigscreenVR/spt-bs-releases`** — the less-private *releases* repo — NOT on `spt-bs-core` where the code lives. This is a deliberate exception (operator-ruled 2026-07-27): file, triage, and reference issues against `spt-bs-releases`.\r\n\r\nRequests and milestones are tracked through the **alchemy shell** (`alchemy-0`, the Requests shell bound to `spt-bs-releases`) on top of those GitHub issues. The label taxonomy is the workflow: `state:` BACKLOG → EVAL → GREENLIT → WIP → ACCEPTANCE → DONE (plus CUT / SHELVED), `type:` BUGFIX / ADDITION / CHANGE, `flag: NEEDS-OPERATOR`, and per-node `kind: SPT-TRACKING` pins (marked DO NOT DELETE — leave them alone).\r\n\r\n## Other conventions\r\n\r\n- Match surrounding code style; copy-verbatim the sister project's stable wire/schema formats (ADR-0001), clean-room everything else.\r\n- Honor every `docs/KNOWN-HAZARDS.md` invariant — these are real bugs already paid for once.\r\n- Docs are dual-audience (human + AI dev-agent) per `docs/DOCS-STRATEGY.md`; doc generation is CI-gated against drift.\r\n- Commit messages end with the project's co-author trailer, exact spelling `Co-authored by: <your live agent name>` (e.g. `Co-authored by: todlando`) — NOT git's hyphenated `Co-Authored-By:`; CI's runner-informant parses only the space spelling for attribution.\r\n  - **Audit this trailer with `git log --grep` or a raw body read, NEVER `%(trailers:)`** (nor `interpret-trailers --parse`): git's trailer tokenizer requires a whitespace-free token, so the space spelling is structurally invisible to it and every trailer query returns a confident EMPTY on every commit — a clean zero that reads as total attribution loss. Measured 2026-07-30 (main @`8f3e10b`, Signed-off-by and hyphenated-key controls both parse); re-derived at full cost 2026-08-03 by one of its own original measurers, which is why this sentence lives here at the mandate rather than in a document you must decide to open.\r\n- **Worktrees (binding, operator-ruled 2026-07-05):** every `git worktree` (gate rigs, agent scratch, anything) goes under `.worktrees/` in THIS project's root (`.gitignore`d) — e.g. `git worktree add --detach .worktrees/gate-<sha> <sha>`. NEVER create worktrees as siblings of the project dir; stray sibling folders accumulate orphan-pinned skeletons and pollute `~/Documents/projects`. Remove yours when done (`git worktree remove` + prune; retry later if a handle pins it).\r\n<!-- [doc->REQ-POOL-FOREIGN-TREE-REFUSED] -->\r\n- **Build-cache pools (binding, releases#103):** a cargo `target/` directory belongs to ONE source tree. Two trees building into one pool overwrite each other's artifacts and fabricate reds that belong to neither — the v0.51.0 W1/H1 `E0308` that vanished on re-run. **The rule is about the RELATIONSHIP, not the path:** a junction and a `CARGO_TARGET_DIR` env var reach the same hazard, and the env var leaves no directory entry to notice afterwards, so \"is this dir under `.worktrees/`\" never was the question. A build script enforces it (`crates/spt-store/build.rs`) and refuses with `SPT_POOL_FOREIGN`. **Sharing a pool SEQUENTIALLY is fine and expected** — a finished lane is taken over loudly, not refused; only a LIVE lane blocks you. Claim a pool at lane start …\n- **Target-tree teardown (releases#103, same lane):** never delete a `target/` you have not classified. Check OUTBOUND first (`Get-Item -Force`: a junction is deleted as a link, a real dir as a tree), then sweep INBOUND for reparse points aimed at it before removing a real one, and measure size and free space before and after so the reclaim is a number rather than a claim. Reap the target SUBTREE, never a session or worktree directory around it. Do not point `CARGO_TARGET_DIR` into a scratchpad (v0.50.0 incident) or at another checkout's target (H1 incident).\r\n- **CI / merge strategy:** golden CI (ADR-0050, ruled 2026-07-27) — milestone-batch golden runs, thin-lane PRs, ff-only main (tested sha == merged sha). Piloting from the next milestone; read the ADR before shaping a PR lane.\r\n\r\nIf you finish a significant body of work without need for user intervention, or if your context gets too high, you can clear your own context and keep moving:\r\n1. Create a JIT plan for the next immediate body of work, if it isn't already planned\r\n2. If you are a live agent, /sptc:commune with immediate next steps and broad summary of the project's status + end goal\r\n3. Execute the /self-clear skill\r","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47]},"fileSize":8436,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\AGENTS.md"},"limits":{"columnTruncated":{"maxColumn":768}}}},"isError":false,"timestamp":1789001101758}}
{"type":"message","id":"6bbb3f7d","parentId":"b4b131f7","timestamp":"2026-09-10T00:45:01.764Z","message":{"role":"toolResult","toolCallId":"call_8tZqKMdMkqqi3pvzPDyJk9XB|fc_077f3569365e313d016aa1fd8dabac87d180bc9028fa0b2249","toolName":"read","content":[{"type":"text","text":"# spt-core — agent working rules\n\nHarness-independent core for an agent ecosystem (messaging, live-agent lifecycle, terminal hosting, P2P networking, runtime-manifest harness contract). Rust crates + a single `spt` binary. Clean-room rebuild of `claude_skill_owl`.\n\n**Orientation — read before working:** `PRD.md` (requirements), `ROADMAP.md` (build path), `CONTEXT.md` (glossary/model, authoritative for meaning), `docs/adr/` (decisions), `docs/KNOWN-HAZARDS.md` (invariants we must not re-break), `docs/{STORAGE,MANIFEST,CONTEXT-MEMORY,DOCS-STRATEGY,TRACEABILITY}.md`.\n\nFor maintainer/debug update work, also read `docs/DEBUG-ROLLOUT.md` before touching the update-set or rollout path.\n\nFor work needing Linux tests and proof, this machine is authorized to ssh into a Linux box: `reavus@kitsubito`\n\nIMPORTANT: All agents you work alongside run on spt-core itself — the legacy-SPT (`$OWL`/`$LIVE`) era ended with the 2026-07 fleet migration; those commands are gone. Reach other agents with `spt send` (or the `/sptc:send` skill); keep message bodies short, and send long bodies from a file rather than inline. Everything else can be learned from the corresponding `/sptc:` skill doc.\n\n## Requirement traceability (binding)\n\nThis project uses `traceable-reqs` (`traceable-reqs.toml` = the authoritative `REQ-*` registry). The full contract is `docs/TRACEABILITY.md`. The rules you must follow:\n\n1. **Tag evidence in the same change.** When you write a function/test/doc-section that satisfies a requirement stage, add its tag in that same commit:\n   `// [impl->REQ-FOO]` · `// [unit->REQ-FOO]` · `<!-- [doc->REQ-FOO] -->`\n   Stages: `doc` / `impl` / `unit` / `int`. Tag *on or immediately above* the real evidence — never at file tops to satisfy coverage.\n2. **Run `traceable-reqs check` before declaring work done.** Exit-1 means missing/invalid evidence — fix it, don't ship it.\n3. **New requirement → add it to `traceable-reqs.toml` first** (with a `REQ-*` id), then satisfy it. No untracked work; no untagged evidence.\n4. **KNOWN-HAZARDS are `REQ-HAZARD-*` requirements** — each needs a test before it's \"covered.\" Treat the hazard list as a conformance checklist you must satisfy, not advice.\n5. **Activate, don't pre-fail.** Requirements you aren't yet working stay `required_stages = []`. Activate (set real stages) only when starting the milestone that delivers them.\n\n## Issue tracking (unique case)\n\nCodebase issues are maintained on **`BigscreenVR/spt-bs-releases`** — the less-private *releases* repo — NOT on `spt-bs-core` where the code lives. This is a deliberate exception (operator-ruled 2026-07-27): file, triage, and reference issues against `spt-bs-releases`.\n\nRequests and milestones are tracked through the **alchemy shell** (`alchemy-0`, the Requests shell bound to `spt-bs-releases`) on top of those GitHub issues. The label taxonomy is the workflow: `state:` BACKLOG → EVAL → GREENLIT → WIP → ACCEPTANCE → DONE (plus CUT / SHELVED), `type:` BUGFIX / ADDITION / CHANGE, `flag: NEEDS-OPERATOR`, and per-node `kind: SPT-TRACKING` pins (marked DO NOT DELETE — leave them alone).\n\n## Other conventions\n\n- Match surrounding code style; copy-verbatim the sister project's stable wire/schema formats (ADR-0001), clean-room everything else.\n- Honor every `docs/KNOWN-HAZARDS.md` invariant — these are real bugs already paid for once.\n- Docs are dual-audience (human + AI dev-agent) per `docs/DOCS-STRATEGY.md`; doc generation is CI-gated against drift.\n- Commit messages end with the project's co-author trailer, exact spelling `Co-authored by: <your live agent name>` (e.g. `Co-authored by: todlando`) — NOT git's hyphenated `Co-Authored-By:`; CI's runner-informant parses only the space spelling for attribution.\n  - **Audit this trailer with `git log --grep` or a raw body read, NEVER `%(trailers:)`** (nor `interpret-trailers --parse`): git's trailer tokenizer requires a whitespace-free token, so the space spelling is structurally invisible to it and every trailer query returns a confident EMPTY on every commit — a clean zero that reads as total attribution loss. Measured 2026-07-30 (main @`8f3e10b`, Signed-off-by and hyphenated-key controls both parse); re-derived at full cost 2026-08-03 by one of its own original measurers, which is why this sentence lives here at the mandate rather than in a document you must decide to open.\n- **Worktrees (binding, operator-ruled 2026-07-05):** every `git worktree` (gate rigs, agent scratch, anything) goes under `.worktrees/` in THIS project's root (`.gitignore`d) — e.g. `git worktree add --detach .worktrees/gate-<sha> <sha>`. NEVER create worktrees as siblings of the project dir; stray sibling folders accumulate orphan-pinned skeletons and pollute `~/Documents/projects`. Remove yours when done (`git worktree remove` + prune; retry later if a handle pins it).\n<!-- [doc->REQ-POOL-FOREIGN-TREE-REFUSED] -->\n- **Build-cache pools (binding, releases#103):** a cargo `target/` directory belongs to ONE source tree. Two trees building into one pool overwrite each other's artifacts and fabricate reds that belong to neither — the v0.51.0 W1/H1 `E0308` that vanished on re-run. **The rule is about the RELATIONSHIP, not the path:** a junction and a `CARGO_TARGET_DIR` env var reach the same hazard, and the env var leaves no directory entry to notice afterwards, so \"is this dir under `.worktrees/`\" never was the question. A build script enforces it (`crates/spt-store/build.rs`) and refuses with `SPT_POOL_FOREIGN`. **Sharing a pool SEQUENTIALLY is fine and expected** — a finished lane is taken over loudly, not refused; only a LIVE lane blocks you. Claim a pool at lane start …\n- **Target-tree teardown (releases#103, same lane):** never delete a `target/` you have not classified. Check OUTBOUND first (`Get-Item -Force`: a junction is deleted as a link, a real dir as a tree), then sweep INBOUND for reparse points aimed at it before removing a real one, and measure size and free space before and after so the reclaim is a number rather than a claim. Reap the target SUBTREE, never a session or worktree directory around it. Do not point `CARGO_TARGET_DIR` into a scratchpad (v0.50.0 incident) or at another checkout's target (H1 incident).\n- **CI / merge strategy:** golden CI (ADR-0050, ruled 2026-07-27) — milestone-batch golden runs, thin-lane PRs, ff-only main (tested sha == merged sha). Piloting from the next milestone; read the ADR before shaping a PR lane.\n\nIf you finish a significant body of work without need for user intervention, or if your context gets too high, you can clear your own context and keep moving:\n1. Create a JIT plan for the next immediate body of work, if it isn't already planned\n2. If you are a live agent, /sptc:commune with immediate next steps and broad summary of the project's status + end goal\n3. Execute the /self-clear skill\n\n[Some lines truncated to 768 chars]"}],"details":{"totalLines":47,"displayContent":{"text":"# spt-core — agent working rules\r\n\r\nHarness-independent core for an agent ecosystem (messaging, live-agent lifecycle, terminal hosting, P2P networking, runtime-manifest harness contract). Rust crates + a single `spt` binary. Clean-room rebuild of `claude_skill_owl`.\r\n\r\n**Orientation — read before working:** `PRD.md` (requirements), `ROADMAP.md` (build path), `CONTEXT.md` (glossary/model, authoritative for meaning), `docs/adr/` (decisions), `docs/KNOWN-HAZARDS.md` (invariants we must not re-break), `docs/{STORAGE,MANIFEST,CONTEXT-MEMORY,DOCS-STRATEGY,TRACEABILITY}.md`.\r\n\r\nFor maintainer/debug update work, also read `docs/DEBUG-ROLLOUT.md` before touching the update-set or rollout path.\r\n\r\nFor work needing Linux tests and proof, this machine is authorized to ssh into a Linux box: `reavus@kitsubito`\r\n\r\nIMPORTANT: All agents you work alongside run on spt-core itself — the legacy-SPT (`$OWL`/`$LIVE`) era ended with the 2026-07 fleet migration; those commands are gone. Reach other agents with `spt send` (or the `/sptc:send` skill); keep message bodies short, and send long bodies from a file rather than inline. Everything else can be learned from the corresponding `/sptc:` skill doc.\r\n\r\n## Requirement traceability (binding)\r\n\r\nThis project uses `traceable-reqs` (`traceable-reqs.toml` = the authoritative `REQ-*` registry). The full contract is `docs/TRACEABILITY.md`. The rules you must follow:\r\n\r\n1. **Tag evidence in the same change.** When you write a function/test/doc-section that satisfies a requirement stage, add its tag in that same commit:\r\n   `// [impl->REQ-FOO]` · `// [unit->REQ-FOO]` · `<!-- [doc->REQ-FOO] -->`\r\n   Stages: `doc` / `impl` / `unit` / `int`. Tag *on or immediately above* the real evidence — never at file tops to satisfy coverage.\r\n2. **Run `traceable-reqs check` before declaring work done.** Exit-1 means missing/invalid evidence — fix it, don't ship it.\r\n3. **New requirement → add it to `traceable-reqs.toml` first** (with a `REQ-*` id), then satisfy it. No untracked work; no untagged evidence.\r\n4. **KNOWN-HAZARDS are `REQ-HAZARD-*` requirements** — each needs a test before it's \"covered.\" Treat the hazard list as a conformance checklist you must satisfy, not advice.\r\n5. **Activate, don't pre-fail.** Requirements you aren't yet working stay `required_stages = []`. Activate (set real stages) only when starting the milestone that delivers them.\r\n\r\n## Issue tracking (unique case)\r\n\r\nCodebase issues are maintained on **`BigscreenVR/spt-bs-releases`** — the less-private *releases* repo — NOT on `spt-bs-core` where the code lives. This is a deliberate exception (operator-ruled 2026-07-27): file, triage, and reference issues against `spt-bs-releases`.\r\n\r\nRequests and milestones are tracked through the **alchemy shell** (`alchemy-0`, the Requests shell bound to `spt-bs-releases`) on top of those GitHub issues. The label taxonomy is the workflow: `state:` BACKLOG → EVAL → GREENLIT → WIP → ACCEPTANCE → DONE (plus CUT / SHELVED), `type:` BUGFIX / ADDITION / CHANGE, `flag: NEEDS-OPERATOR`, and per-node `kind: SPT-TRACKING` pins (marked DO NOT DELETE — leave them alone).\r\n\r\n## Other conventions\r\n\r\n- Match surrounding code style; copy-verbatim the sister project's stable wire/schema formats (ADR-0001), clean-room everything else.\r\n- Honor every `docs/KNOWN-HAZARDS.md` invariant — these are real bugs already paid for once.\r\n- Docs are dual-audience (human + AI dev-agent) per `docs/DOCS-STRATEGY.md`; doc generation is CI-gated against drift.\r\n- Commit messages end with the project's co-author trailer, exact spelling `Co-authored by: <your live agent name>` (e.g. `Co-authored by: todlando`) — NOT git's hyphenated `Co-Authored-By:`; CI's runner-informant parses only the space spelling for attribution.\r\n  - **Audit this trailer with `git log --grep` or a raw body read, NEVER `%(trailers:)`** (nor `interpret-trailers --parse`): git's trailer tokenizer requires a whitespace-free token, so the space spelling is structurally invisible to it and every trailer query returns a confident EMPTY on every commit — a clean zero that reads as total attribution loss. Measured 2026-07-30 (main @`8f3e10b`, Signed-off-by and hyphenated-key controls both parse); re-derived at full cost 2026-08-03 by one of its own original measurers, which is why this sentence lives here at the mandate rather than in a document you must decide to open.\r\n- **Worktrees (binding, operator-ruled 2026-07-05):** every `git worktree` (gate rigs, agent scratch, anything) goes under `.worktrees/` in THIS project's root (`.gitignore`d) — e.g. `git worktree add --detach .worktrees/gate-<sha> <sha>`. NEVER create worktrees as siblings of the project dir; stray sibling folders accumulate orphan-pinned skeletons and pollute `~/Documents/projects`. Remove yours when done (`git worktree remove` + prune; retry later if a handle pins it).\r\n<!-- [doc->REQ-POOL-FOREIGN-TREE-REFUSED] -->\r\n- **Build-cache pools (binding, releases#103):** a cargo `target/` directory belongs to ONE source tree. Two trees building into one pool overwrite each other's artifacts and fabricate reds that belong to neither — the v0.51.0 W1/H1 `E0308` that vanished on re-run. **The rule is about the RELATIONSHIP, not the path:** a junction and a `CARGO_TARGET_DIR` env var reach the same hazard, and the env var leaves no directory entry to notice afterwards, so \"is this dir under `.worktrees/`\" never was the question. A build script enforces it (`crates/spt-store/build.rs`) and refuses with `SPT_POOL_FOREIGN`. **Sharing a pool SEQUENTIALLY is fine and expected** — a finished lane is taken over loudly, not refused; only a LIVE lane blocks you. Claim a pool at lane start …\n- **Target-tree teardown (releases#103, same lane):** never delete a `target/` you have not classified. Check OUTBOUND first (`Get-Item -Force`: a junction is deleted as a link, a real dir as a tree), then sweep INBOUND for reparse points aimed at it before removing a real one, and measure size and free space before and after so the reclaim is a number rather than a claim. Reap the target SUBTREE, never a session or worktree directory around it. Do not point `CARGO_TARGET_DIR` into a scratchpad (v0.50.0 incident) or at another checkout's target (H1 incident).\r\n- **CI / merge strategy:** golden CI (ADR-0050, ruled 2026-07-27) — milestone-batch golden runs, thin-lane PRs, ff-only main (tested sha == merged sha). Piloting from the next milestone; read the ADR before shaping a PR lane.\r\n\r\nIf you finish a significant body of work without need for user intervention, or if your context gets too high, you can clear your own context and keep moving:\r\n1. Create a JIT plan for the next immediate body of work, if it isn't already planned\r\n2. If you are a live agent, /sptc:commune with immediate next steps and broad summary of the project's status + end goal\r\n3. Execute the /self-clear skill\r","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47]},"fileSize":8436,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\AGENTS.md"},"limits":{"columnTruncated":{"maxColumn":768}}}},"isError":false,"timestamp":1789001101763}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_bANIq5hclQGL4KoJ0gXAu7X4|fc_077f3569365e313d016aa1fd96385c87d1aeb95b6e4bd7f09c","toolName":"task","startedAt":"2026-09-10T00:45:27.367Z","intent":"Delegating independent review slices"},"id":"1df39630","parentId":"6bbb3f7d","timestamp":"2026-09-10T00:45:27.367Z"}
{"type":"message","id":"f4b2fa32","parentId":"1df39630","timestamp":"2026-09-10T00:45:27.368Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":"**Scoping standards-spec work**"},{"type":"toolCall","id":"call_bANIq5hclQGL4KoJ0gXAu7X4|fc_077f3569365e313d016aa1fd96385c87d1aeb95b6e4bd7f09c","name":"task","arguments":{"i":"Delegating independent review slices","context":"# Goal\nRead-only review of .worktrees/287-shellrestore commit cb756dee against cb756dee^ for releases#287 operator ruling (https://github.com/BigscreenVR/spt-bs-releases/issues/287#issuecomment-5583820305). Main review owner PersistentRestoreReview traces production shellwake/shellhost logic.\n# Constraints\nNo edits, formatters, linters, builds, tests, or production operations. Source review only. Skip all validation commands. Only report actionable provable patch-introduced correctness/spec issues, not style nits. Use read/grep and read-only git diff/show/log. Repo issues are BigscreenVR/spt-bs-releases.\n# Contract\nRuling: persistent+owner online+down IN FACT restores on boot sweep and owner offline->online edge regardless of machine boot, launch age, missing stamp/oracle. Nonpersistent same-boot freeze survives. Riders: enumerate every healed instance per owner; relink reports binding until bind (not offline or premature online). Report exact patch-overlapping file/line and impact, or no findings. Do not edit or overlap production review except reading needed test symbols.","tasks":[{"name":"RestoreDocsStandards","task":"# Target\nChanged docs and standards only: CHANGELOG.md, CONTEXT.md, docs-site/src/shells/overview.md, docs/KNOWN-HAZARDS.md, docs/MANIFEST.md, traceable-reqs.toml in .worktrees/287-shellrestore; cb756dee vs parent. Non-goal production logic audit.\n# Change\nRead repo AGENTS.md and relevant traceability/docs standards. Read full diff for changed docs, then affected surrounding context and referenced requirements. Check retained contradictory persistent freeze or missing stamp/boot gates in changed docs; operator-ruling coverage and both riders. Distinguish Standards and Spec concerns. Verify only issue-worthy defects, not historical text explicitly superseded. No edits. Skip all formatting/lint/build/tests.\n# Acceptance\nReturn concise evidence-backed findings with exact lines and impact, or no findings with standards/spec assessment. Source acceptance only; no runtime claims."},{"name":"RestoreTestsReview","task":"# Target\nChanged tests in crates/spt-daemon/src/shellwake.rs, crates/spt-daemon/src/shellhost.rs, crates/spt/tests/shell_relink_force_e2e.rs, plus CLI relink context in crates/spt/src/cli.rs; cb756dee vs parent in .worktrees/287-shellrestore. Non-goal production restore algorithm audit (parent handles).\n# Change\nRead diff and complete changed test context. Check whether new assertions actually cover ruling: same-boot/missing stamp persistent restoration, nonpersistent freeze, live binding/PID-reuse exclusion, all healed instances, and binding relink output. Inspect CLI relink branches and any output consumer dispatch that could be broken by status change. Report concrete introduced bugs, not requests to pad tests. Skip all validation, builds, linters, formatters and test execution; no edits.\n# Acceptance\nReturn actionable correctness/spec findings with exact diff lines and affected path, or no findings, separating source review from unrun runtime tests."}]},"intent":"Delegating independent review slices"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":9756,"output":774,"cacheRead":19840,"cacheWrite":0,"totalTokens":30370,"reasoningTokens":91,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1789001101773,"responseId":"resp_077f3569365e313d016aa1fd8e09a887d191bc1aafa7e744c9","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_077f3569365e313d016aa1fd8f78fc87d1a10767a7144c6e97","type":"reasoning","content":[],"encrypted_content":"gAAAAABqof2WAOVagEfJjNoL1KOpXYqUGn1NoyJ1rqj6s41N_bcHlU1Qgr8Y6vVZwYRQumMeXxBqsh6GdAn6SXXIjy2WiR4s7CEnYCb1onLETA_S0tf6NvU3IdbML8engaCf4TjigCBtezEC0fa8HJ25KvSlGc7z4xgHcec8__sMJkpZskwXKrSXroYbzG2rCfRuWeKi8hHL5ienil76fyyRvShDrY4r7Qr8SaHx_K7rzZhnmDxi1lYsfilPmR-HpXkZe127wtMqygr3nB-3X4-32ouTuof0wOG1fEXkcWXBqcSXYSzVOZx-w8Gy6jyzpN4pCfe4gqbtmaUmpQ0eejKMmrb67T_Ep6h184yxq0KG7urFD63ZyjwNLzkWfkm91uXX7OmFxEr7ilYUHezXFNalT7PPkRrvzJoLbPeZLWoKYmu4Zf7JU0j9sFb3Un8PyjD1ZD6KpXbIC3_4oagPlbrnaVxeVV-M1-Dz_Of-DYn1i9KWHp-7aDdNdTIDxB6gvk5c8BsboewnmIx84m4OyETe7eef-0CURkIiyjMxlTrJ3IkxO2wYrnePatJpb8_YG86GQX1ANo7_ZOMCcpKtwnHGRgBgbSoxbgm2civXHHXdxf452Q9pAZd-d0TQANPlIn5gAEsA5Of2yIqpQYVV4bOVkGD_XH8heAFIDzNFZpi3LMF87MG91Cl5OivFDiROSIpBI4NmSwQkeY4AwJ7Qmb6rNaGyI15hIKYAJmi8xIUStTGqDtb8SADm9f87zOmt88vEwzKTdn37yZAfn3fLB0QzcpCH7Jr-AGE773RhgRoqaCq8hxw_ySj5cDPjRCU_fh04G0biieH7CiZE9H2_iyd1fkRovg1SdbvV3px2-QASEJ2IOa6MFZtkBuAqUjrM_7BL7JNa0lAWSmME08yH6_ChJLa8ZvBe4PWo_IjAkz9hk_Czdk7o7DSG8XiejwLHlyoZ0d2QV4-DG5Fpkb5toMBvPEkC9G92-GrQOa-gwL-3gK-VonOSpa3tthmVBheprOkjacPehCXnmI7e5-3mGig-R76NtUMBpDzIFHeCSVwh3qgwCgpNFDOYHCXWVGzunwTKjPUyq19QJL_MLOOxV29cr1pW4bZ8zR7jzlNbZl3Ao-9pQEB5lxmowm_jukKx2F2T2MKRn3tPBYY5-S_BrUYx5w03y7awGUlC8m4thp9lXV4KwwkysaLdtm0l1qyFnqPmVsC_ZK-QV53BmYBMvYVRqz0F9ee1rX8-XYsjXR9qM2baFUcysTI2GWJgHOE0qhcRRX2aqf6XytXjHfz6XQUsvLLe7upU8k-txbmcdsThfuIBhb7ughelKReTWnQhmM1jBEvYC7g2bfHkS9kzrvtUF1Si4m0SY-iFUWf2AteB9rGc_FV4-6p8Td1wWvfQxjsn-fj-ZzDmE2Kw-Y-0ZJSBkS-VjgJPzzWjrf_N6naGIISWFicTAaqTe7A_tSD3DoDgj4hlkZh6pD8TfixziF-_R9Ct40zjOauxgKtdSXEUB5j4HjheBqA-KhObu24adbA18p5PR7u9uRF7Pwa_JVW6Ndm4awS_ki2zbc5WD94y0Vu2CUtxNQ10pg-RZcCSaP35c4_ndidXIDkSVTMCTvAluPJm6daeR4vd9NUdMxpBUo4DfYwh1QGeOApPlw-7oC9L5-PbQ3RLsBHA5p2pLUWdAFGFKiGi96amm923WsC4jHzx0t_p9JZ51Mp8ORKSq968DtwHjwx312iKYxHgfeTpMBcdWRS8NtSmmVZnan10L2tpr635TC97JIw0ql3zQaXMuUqeFEDvrfvUT7yCrNQexQuZMsJVTwolm87Ss-oGPxVYmnFx_8RewlN8GymXJxaEDGJ4RFhW2wXzgDFIY_lE2Bpy-958zrnrBHk0qUg15h-npbw8dhYmxadJJgHvwS2HUEwKwFpC","internal_chat_message_metadata_passthrough":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"},"summary":[{"type":"summary_text","text":"**Scoping standards-spec work**"}],"metadata":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"}},{"id":"fc_077f3569365e313d016aa1fd96385c87d1aeb95b6e4bd7f09c","type":"function_call","status":"completed","arguments":"{\"i\":\"Delegating independent review slices\",\"context\":\"# Goal\\nRead-only review of .worktrees/287-shellrestore commit cb756dee against cb756dee^ for releases#287 operator ruling (https://github.com/BigscreenVR/spt-bs-releases/issues/287#issuecomment-5583820305). Main review owner PersistentRestoreReview traces production shellwake/shellhost logic.\\n# Constraints\\nNo edits, formatters, linters, builds, tests, or production operations. Source review only. Skip all validation commands. Only report actionable provable patch-introduced correctness/spec issues, not style nits. Use read/grep and read-only git diff/show/log. Repo issues are BigscreenVR/spt-bs-releases.\\n# Contract\\nRuling: persistent+owner online+down IN FACT restores on boot sweep and owner offline->online edge regardless of machine boot, launch age, missing stamp/oracle. Nonpersistent same-boot freeze survives. Riders: enumerate every healed instance per owner; relink reports binding until bind (not offline or premature online). Report exact patch-overlapping file/line and impact, or no findings. Do not edit or overlap production review except reading needed test symbols.\",\"tasks\":[{\"name\":\"RestoreDocsStandards\",\"task\":\"# Target\\nChanged docs and standards only: CHANGELOG.md, CONTEXT.md, docs-site/src/shells/overview.md, docs/KNOWN-HAZARDS.md, docs/MANIFEST.md, traceable-reqs.toml in .worktrees/287-shellrestore; cb756dee vs parent. Non-goal production logic audit.\\n# Change\\nRead repo AGENTS.md and relevant traceability/docs standards. Read full diff for changed docs, then affected surrounding context and referenced requirements. Check retained contradictory persistent freeze or missing stamp/boot gates in changed docs; operator-ruling coverage and both riders. Distinguish Standards and Spec concerns. Verify only issue-worthy defects, not historical text explicitly superseded. No edits. Skip all formatting/lint/build/tests.\\n# Acceptance\\nReturn concise evidence-backed findings with exact lines and impact, or no findings with standards/spec assessment. Source acceptance only; no runtime claims.\"},{\"name\":\"RestoreTestsReview\",\"task\":\"# Target\\nChanged tests in crates/spt-daemon/src/shellwake.rs, crates/spt-daemon/src/shellhost.rs, crates/spt/tests/shell_relink_force_e2e.rs, plus CLI relink context in crates/spt/src/cli.rs; cb756dee vs parent in .worktrees/287-shellrestore. Non-goal production restore algorithm audit (parent handles).\\n# Change\\nRead diff and complete changed test context. Check whether new assertions actually cover ruling: same-boot/missing stamp persistent restoration, nonpersistent freeze, live binding/PID-reuse exclusion, all healed instances, and binding relink output. Inspect CLI relink branches and any output consumer dispatch that could be broken by status change. Report concrete introduced bugs, not requests to pad tests. Skip all validation, builds, linters, formatters and test execution; no edits.\\n# Acceptance\\nReturn actionable correctness/spec findings with exact diff lines and affected path, or no findings, separating source review from unrun runtime tests.\"}]}","call_id":"call_bANIq5hclQGL4KoJ0gXAu7X4","internal_chat_message_metadata_passthrough":{"create_time":1789001102.3318,"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"},"name":"task","metadata":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"}}]},"duration":25590.29679999966,"ttft":1727.6099999998696,"completedAt":1789001127364,"contextSnapshot":{"promptTokens":29596,"nonMessageTokens":13209,"compactionEpoch":0}}}
{"type":"message","id":"abdd6300","parentId":"f4b2fa32","timestamp":"2026-09-10T00:45:27.376Z","message":{"role":"toolResult","toolCallId":"call_bANIq5hclQGL4KoJ0gXAu7X4|fc_077f3569365e313d016aa1fd96385c87d1aeb95b6e4bd7f09c","toolName":"task","content":[{"type":"text","text":"Spawned 2 background agents using scout. Each result auto-delivers on yield unless a settled `hub jobs`/`wait` snapshot consumes it first.\n- `PersistentRestoreReview.RestoreDocsStandards` (job `PersistentRestoreReview.RestoreDocsStandards`)\n- `PersistentRestoreReview.RestoreTestsReview` (job `PersistentRestoreReview.RestoreTestsReview`)\nDM these ids via `hub` send to coordinate while they run; use `hub` only to inspect (`jobs`), wait, or cancel a stuck task.\nNo polling needed.\n\nSettled-job inspection: `hub jobs` | `hub wait` delivers its snapshot → no duplicate `async-result`.\n\nJob IDs: process memory ~5min after settlement; afterward use agent ID: `hub send`, `agent://<id>`, `history://<id>`.\n\n`completed`: subagent yielded successfully; claimed artifacts unverified.\n\nCoordinate: 2 siblings are running together. If their work overlaps, have them message each other via `hub` (by id, or \"all\" to broadcast) before editing shared files — live coordination beats a serial handoff. Check `hub` op:\"list\" to see who is doing what."}],"details":{"projectAgentsDir":null,"results":[],"totalDurationMs":1,"progress":[{"index":0,"id":"PersistentRestoreReview.RestoreDocsStandards","agent":"scout","agentSource":"bundled","modelRole":"smol","status":"pending","task":"Complete assignment thoroughly:\n\n# Target\nChanged docs and standards only: CHANGELOG.md, CONTEXT.md, docs-site/src/shells/overview.md, docs/KNOWN-HAZARDS.md, docs/MANIFEST.md, traceable-reqs.toml in .worktrees/287-shellrestore; cb756dee vs parent. Non-goal production logic audit.\n# Change\nRead repo AGENTS.md and relevant traceability/docs standards. Read full diff for changed docs, then affected surrounding context and referenced requirements. Check retained contradictory persistent freeze or missing stamp/boot gates in changed docs; operator-ruling coverage and both riders. Distinguish Standards and Spec concerns. Verify only issue-worthy defects, not historical text explicitly superseded. No edits. Skip all formatting/lint/build/tests.\n# Acceptance\nReturn concise evidence-backed findings with exact lines and impact, or no findings with standards/spec assessment. Source acceptance only; no runtime claims.","assignment":"# Target\nChanged docs and standards only: CHANGELOG.md, CONTEXT.md, docs-site/src/shells/overview.md, docs/KNOWN-HAZARDS.md, docs/MANIFEST.md, traceable-reqs.toml in .worktrees/287-shellrestore; cb756dee vs parent. Non-goal production logic audit.\n# Change\nRead repo AGENTS.md and relevant traceability/docs standards. Read full diff for changed docs, then affected surrounding context and referenced requirements. Check retained contradictory persistent freeze or missing stamp/boot gates in changed docs; operator-ruling coverage and both riders. Distinguish Standards and Spec concerns. Verify only issue-worthy defects, not historical text explicitly superseded. No edits. Skip all formatting/lint/build/tests.\n# Acceptance\nReturn concise evidence-backed findings with exact lines and impact, or no findings with standards/spec assessment. Source acceptance only; no runtime claims.","recentTools":[],"recentOutput":[],"toolCount":0,"requests":0,"tokens":0,"cost":0,"durationMs":0},{"index":1,"id":"PersistentRestoreReview.RestoreTestsReview","agent":"scout","agentSource":"bundled","modelRole":"smol","status":"pending","task":"Complete assignment thoroughly:\n\n# Target\nChanged tests in crates/spt-daemon/src/shellwake.rs, crates/spt-daemon/src/shellhost.rs, crates/spt/tests/shell_relink_force_e2e.rs, plus CLI relink context in crates/spt/src/cli.rs; cb756dee vs parent in .worktrees/287-shellrestore. Non-goal production restore algorithm audit (parent handles).\n# Change\nRead diff and complete changed test context. Check whether new assertions actually cover ruling: same-boot/missing stamp persistent restoration, nonpersistent freeze, live binding/PID-reuse exclusion, all healed instances, and binding relink output. Inspect CLI relink branches and any output consumer dispatch that could be broken by status change. Report concrete introduced bugs, not requests to pad tests. Skip all validation, builds, linters, formatters and test execution; no edits.\n# Acceptance\nReturn actionable correctness/spec findings with exact diff lines and affected path, or no findings, separating source review from unrun runtime tests.","assignment":"# Target\nChanged tests in crates/spt-daemon/src/shellwake.rs, crates/spt-daemon/src/shellhost.rs, crates/spt/tests/shell_relink_force_e2e.rs, plus CLI relink context in crates/spt/src/cli.rs; cb756dee vs parent in .worktrees/287-shellrestore. Non-goal production restore algorithm audit (parent handles).\n# Change\nRead diff and complete changed test context. Check whether new assertions actually cover ruling: same-boot/missing stamp persistent restoration, nonpersistent freeze, live binding/PID-reuse exclusion, all healed instances, and binding relink output. Inspect CLI relink branches and any output consumer dispatch that could be broken by status change. Report concrete introduced bugs, not requests to pad tests. Skip all validation, builds, linters, formatters and test execution; no edits.\n# Acceptance\nReturn actionable correctness/spec findings with exact diff lines and affected path, or no findings, separating source review from unrun runtime tests.","recentTools":[],"recentOutput":[],"toolCount":0,"requests":0,"tokens":0,"cost":0,"durationMs":0}],"async":{"state":"running","jobId":"PersistentRestoreReview.RestoreDocsStandards","type":"task"}},"isError":false,"timestamp":1789001127376}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_AGim25C0IpPaWiIiWNcatuNa|fc_077f3569365e313d016aa1fdb0c06087d1b2cbebcba03bb903","toolName":"read","startedAt":"2026-09-10T00:45:36.935Z","args":{"path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs:1-320"},"intent":"Reading production wake context"},"id":"372d9e8c","parentId":"abdd6300","timestamp":"2026-09-10T00:45:36.935Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_7CFTl5wDfbMpoAe7PxbOIkfS|fc_077f3569365e313d016aa1fdb0c07487d1985acac801d13058","toolName":"read","startedAt":"2026-09-10T00:45:36.936Z","args":{"path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs:590-1060"},"intent":"Reading persistent restore context"},"id":"27f7a99d","parentId":"372d9e8c","timestamp":"2026-09-10T00:45:36.936Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_b5vz9AVfJemvkhllIGlgLVCf|fc_077f3569365e313d016aa1fdb0c08087d182419544e6345b9c","toolName":"read","startedAt":"2026-09-10T00:45:36.936Z","args":{"path":"artifact://112:301-600"},"intent":"Reading remaining wake patch"},"id":"dd9e705c","parentId":"27f7a99d","timestamp":"2026-09-10T00:45:36.936Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_Af4TRJYugyLpvR3kOYPeRRR4|fc_077f3569365e313d016aa1fdb0c08c87d1b16af7f4dfd838ca","toolName":"read","startedAt":"2026-09-10T00:45:36.936Z","args":{"path":".worktrees/287-shellrestore/crates/spt-store/src/shellinfo.rs:1-350"},"intent":"Reading stored liveness semantics"},"id":"8d317ca6","parentId":"dd9e705c","timestamp":"2026-09-10T00:45:36.936Z"}
{"type":"message","id":"784485da","parentId":"8d317ca6","timestamp":"2026-09-10T00:45:36.936Z","message":{"role":"assistant","content":[{"type":"toolCall","id":"call_AGim25C0IpPaWiIiWNcatuNa|fc_077f3569365e313d016aa1fdb0c06087d1b2cbebcba03bb903","name":"read","arguments":{"i":"Reading production wake context","path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs:1-320"},"intent":"Reading production wake context"},{"type":"toolCall","id":"call_7CFTl5wDfbMpoAe7PxbOIkfS|fc_077f3569365e313d016aa1fdb0c07487d1985acac801d13058","name":"read","arguments":{"i":"Reading persistent restore context","path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs:590-1060"},"intent":"Reading persistent restore context"},{"type":"toolCall","id":"call_b5vz9AVfJemvkhllIGlgLVCf|fc_077f3569365e313d016aa1fdb0c08087d182419544e6345b9c","name":"read","arguments":{"i":"Reading remaining wake patch","path":"artifact://112:301-600"},"intent":"Reading remaining wake patch"},{"type":"toolCall","id":"call_Af4TRJYugyLpvR3kOYPeRRR4|fc_077f3569365e313d016aa1fdb0c08c87d1b16af7f4dfd838ca","name":"read","arguments":{"i":"Reading stored liveness semantics","path":".worktrees/287-shellrestore/crates/spt-store/src/shellinfo.rs:1-350"},"intent":"Reading stored liveness semantics"},{"type":"toolCall","id":"call_8hv8qzFRDNvgINDHIrh0G52O|fc_077f3569365e313d016aa1fdb0c09887d1bad4b5b5fc3d8749","name":"grep","arguments":{"i":"Locating launch guard consumers","pattern":"live_launch_winner|restore_persistent_shells|reconcile_owner_online|heal_stale_online_records|watcher_eligible","path":".worktrees/287-shellrestore/crates","case":true,"gitignore":false,"skip":null},"intent":"Locating launch guard consumers"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":1186,"output":242,"cacheRead":29440,"cacheWrite":0,"totalTokens":30868,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1789001127481,"responseId":"resp_077f3569365e313d016aa1fda811dc87d1afd9cc23967b52a4","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"fc_077f3569365e313d016aa1fdb0c06087d1b2cbebcba03bb903","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading production wake context\",\"path\":\".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs:1-320\"}","call_id":"call_AGim25C0IpPaWiIiWNcatuNa","internal_chat_message_metadata_passthrough":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"},"name":"read","metadata":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"}},{"id":"fc_077f3569365e313d016aa1fdb0c07487d1985acac801d13058","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading persistent restore context\",\"path\":\".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs:590-1060\"}","call_id":"call_7CFTl5wDfbMpoAe7PxbOIkfS","internal_chat_message_metadata_passthrough":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"},"name":"read","metadata":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"}},{"id":"fc_077f3569365e313d016aa1fdb0c08087d182419544e6345b9c","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading remaining wake patch\",\"path\":\"artifact://112:301-600\"}","call_id":"call_b5vz9AVfJemvkhllIGlgLVCf","internal_chat_message_metadata_passthrough":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"},"name":"read","metadata":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"}},{"id":"fc_077f3569365e313d016aa1fdb0c08c87d1b16af7f4dfd838ca","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading stored liveness semantics\",\"path\":\".worktrees/287-shellrestore/crates/spt-store/src/shellinfo.rs:1-350\"}","call_id":"call_Af4TRJYugyLpvR3kOYPeRRR4","internal_chat_message_metadata_passthrough":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"},"name":"read","metadata":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"}},{"id":"fc_077f3569365e313d016aa1fdb0c09887d1bad4b5b5fc3d8749","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating launch guard consumers\",\"pattern\":\"live_launch_winner|restore_persistent_shells|reconcile_owner_online|heal_stale_online_records|watcher_eligible\",\"path\":\".worktrees/287-shellrestore/crates\",\"case\":true,\"gitignore\":false,\"skip\":null}","call_id":"call_8hv8qzFRDNvgINDHIrh0G52O","internal_chat_message_metadata_passthrough":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"},"name":"grep","metadata":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"}}]},"duration":9452.823299999814,"ttft":9321.73499999987,"completedAt":1789001136934,"contextSnapshot":{"promptTokens":30626,"nonMessageTokens":13209,"compactionEpoch":0}}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_8hv8qzFRDNvgINDHIrh0G52O|fc_077f3569365e313d016aa1fdb0c09887d1bad4b5b5fc3d8749","toolName":"grep","startedAt":"2026-09-10T00:45:36.936Z","args":{"path":".worktrees/287-shellrestore/crates"},"intent":"Locating launch guard consumers"},"id":"4ccd2976","parentId":"784485da","timestamp":"2026-09-10T00:45:36.936Z"}
{"type":"message","id":"fb18ed21","parentId":"4ccd2976","timestamp":"2026-09-10T00:45:36.957Z","message":{"role":"toolResult","toolCallId":"call_AGim25C0IpPaWiIiWNcatuNa|fc_077f3569365e313d016aa1fdb0c06087d1b2cbebcba03bb903","toolName":"read","content":[{"type":"text","text":"//! Shell wake-watcher hosting (M5-D4b — CONTEXT §Shell sleep/wake): the\n//! offline half of the online/offline mutual exclusivity. While a shell\n//! instance is **offline** and its manifest declares a `wake_command`, the\n//! daemon runs that template as a long-running **wake-watcher** child whose\n//! sole job is to fire a wake; while the shell binary runs, no watcher does\n//! ([`reconcile_once`] flips between them — spt-core owns the flip).\n//!\n//! **Exit-opcode supervision:** the watcher exiting with [`WAKE_OPCODE`] ⇒\n//! the wake resolution ([`resolve_wake`]) brings the shell online. Any other\n//! exit is a crash ⇒ respawn with exponential backoff ([`backoff_ms`]) and\n//! **give-up** after [`WakeParams::give_up_after`] consecutive crashes (a\n//! durable [`WAKER_GAVE_UP_FILE`] marker the next shell activity — any\n//! (re)launch — clears; crash-bug safety, the watcher can never crash-loop\n//! forever). One watcher per offline instance: the in-process [`WakeSet`]\n//! entry plus the perch's [`WAKER_PID_FILE`] (which also lets a freshly\n//! booted daemon kill a dead daemon's orphaned watcher before adopting).\n//!\n//! **Scheduling (KH 7.4):** every watcher gets its own supervision thread —\n//! a hung waker stalls nothing; the reconcile loop itself is one bounded\n//! sweep per tick off the daemon's control surfaces. Watcher children are\n//! *supervised* (waited) children, never detached-immortal — the KH 5.6\n//! pipe-inherit shape (an unwaitable child holding a capture pipe) does not\n//! arise; stdio is null regardless.\n// [impl->REQ-SHELL-2]\n\nuse std::collections::HashMap;\nuse std::path::Path;\nuse std::process::{Command, Stdio};\nuse std::sync::atomic::{AtomicBool, Ordering};\nuse std::sync::{Arc, Mutex};\nuse std::thread::JoinHandle;\nuse std::time::Duration;\n\nuse spt_runtime::manifest::{AdapterKind, Manifest, Shell};\nuse spt_runtime::registry::AdapterRecord;\nuse spt_store::perch::ParentHint;\nuse spt_store::shellinfo::{self, SHELL_STATUS_OFFLINE, SHELL_STATUS_ONLINE};\n\n/// The wake opcode (documented in `docs/MANIFEST.md` §Shell adapters): a\n/// wake-watcher exiting with **86** fires the wake resolution; every other\n/// exit is a crash.\npub const WAKE_OPCODE: i32 = 86;\n\n/// The running watcher's OS pid, parked on the shell perch — the\n/// one-watcher-per-instance lock and the cross-daemon-restart kill handle.\npub const WAKER_PID_FILE: &str = \"waker.pid\";\n\n/// The give-up latch: present ⇒ the watcher crash-looped past its budget and\n/// the reconciler must NOT restart it. Cleared by the next shell activity\n/// (any (re)launch — [`crate::shellhost::launch_shell`]).\npub const WAKER_GAVE_UP_FILE: &str = \"waker.gaveup\";\n\n/// Watcher supervision knobs (injectable for tests; defaults are production).\n#[derive(Debug, Clone, Copy)]\npub struct WakeParams {\n    /// First-crash respawn delay; doubles per consecutive crash.\n    pub backoff_base_ms: u64,\n    /// Backoff ceiling.\n    pub backoff_cap_ms: u64,\n    /// Consecutive crash-exits before the give-up latch drops.\n    pub give_up_after: u32,\n}\n\nimpl Default for WakeParams {\n    fn default() -> Self {\n        Self {\n            backoff_base_ms: 1_000,\n            backoff_cap_ms: 60_000,\n            give_up_after: 6,\n        }\n    }\n}\n\n/// Exponential backoff for `failures` consecutive crash-exits:\n/// `base × 2^(failures-1)`, saturating at the cap (the\n/// `spt-net::pairing::ratelimit` recipe). Zero failures ⇒ no wait.\n// [impl->REQ-SHELL-2]\npub fn backoff_ms(params: &WakeParams, failures: u32) -> u64 {\n    if failures == 0 {\n        return 0;\n    }\n    params\n        .backoff_base_ms\n        .saturating_mul(1u64 << (failures - 1).min(20))\n        .min(params.backoff_cap_ms)\n}\n\n/// The live watcher registry: one supervision thread per offline instance,\n/// keyed `owner/shell_id`. Shared between the daemon's reconcile loop and\n/// anything that must stop a watcher.\n#[derive(Default)]\npub struct WakeSet {\n    inner: Mutex<HashMap<(String, String), WatcherHandle>>,\n}\n\nstruct WatcherHandle {\n    stop: Arc<AtomicBool>,\n    thread: JoinHandle<()>,\n}\n\nimpl WakeSet {\n    pub fn new() -> Self {\n        Self::default()\n    }\n\n    /// How many watchers are live (finished threads pruned first).\n    pub fn len(&self) -> usize {\n        let mut map = self.inner.lock().unwrap_or_else(|p| p.into_inner());\n        map.retain(|_, h| !h.thread.is_finished());\n        map.len()\n    }\n\n    pub fn is_empty(&self) -> bool {\n        self.len() == 0\n    }\n\n    fn contains(&self, owner: &str, shell_id: &str) -> bool {\n        let mut map = self.inner.lock().unwrap_or_else(|p| p.into_inner());\n        match map.get(&(owner.to_string(), shell_id.to_string())) {\n            Some(h) if !h.thread.is_finished() => true,\n            Some(_) => {\n                map.remove(&(owner.to_string(), shell_id.to_string()));\n                false\n            }\n            None => false,\n        }\n    }\n\n    fn insert(&self, owner: &str, shell_id: &str, handle: WatcherHandle) {\n        let mut map = self.inner.lock().unwrap_or_else(|p| p.into_inner());\n        map.insert((owner.to_string(), shell_id.to_string()), handle);\n    }\n\n    /// Signal one watcher to stop and kill its child (the perch pid file is\n    /// the kill handle). The thread sees the kill as a non-opcode exit with\n    /// its stop flag up and returns without counting a failure.\n    fn stop_watcher(&self, owlery: &Path, owner: &str, shell_id: &str) {\n        let handle = {\n            let mut map = self.inner.lock().unwrap_or_else(|p| p.into_inner());\n            map.remove(&(owner.to_string(), shell_id.to_string()))\n        };\n        if let Some(h) = handle {\n            h.stop.store(true, Ordering::SeqCst);\n            let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n            kill_waker_at(&perch);\n            let _ = h.thread.join();\n        }\n    }\n}\n\n/// Park the waker's pid **and its birth stamp** as one record — `<pid>\\n<birth>`\n/// — in the single write that already existed.\n///\n/// ONE file and ONE write, because the pair must never be torn: a stamp parked\n/// in a second file leaves a window where the pid is on disk and the stamp is\n/// not, and [`spt_store::liveness::relay_liveness`] reads a live pid with no\n/// stamp as `Held`. That window is exactly the unauthenticated kill this\n/// requirement removes, reintroduced as a race.\n///\n/// A backend that exposes no start time writes the pid alone — the same\n/// one-line shape every pre-fix record already has, which reads `Held` on\n/// existence and keeps today's behaviour until the next launch.\n// [impl->REQ-SHELL-KILL-AUTHENTICATED]\npub fn record_waker_launch(shell_perch: &Path, pid: u32) {\n    let body = match spt_store::proc::process_started_at(pid) {\n        Some(birth) => format!(\"{pid}\\n{birth}\"),\n        None => pid.to_string(),\n    };\n    let _ = std::fs::write(shell_perch.join(WAKER_PID_FILE), body);\n}\n\n/// Read that record back — **the** parse, so the two-line format has exactly one\n/// reader shape (the `record_shell_launch`/`read_shell_launch` pattern).\n///\n/// `None` means \"no usable pid on record\": absent, empty, or unparseable. Every\n/// caller must treat that as FAIL-TOWARD-ALIVE — refuse the kill and leave the\n/// record standing. A whole-content `trim().parse()` (what every reader did\n/// before this) returns `None` on the two-line form, so a reader left un-migrated\n/// would kill nothing, retire the record anyway, and report success while\n/// orphaning a live waker.\n// [impl->REQ-SHELL-KILL-AUTHENTICATED]\npub fn read_waker_launch(shell_perch: &Path) -> Option<(u32, Option<u64>)> {\n    let raw = std::fs::read_to_string(shell_perch.join(WAKER_PID_FILE)).ok()?;\n    let mut lines = raw.lines();\n    let pid = lines.next()?.trim().parse::<u32>().ok()?;\n    let birth = lines.next().and_then(|l| l.trim().parse::<u64>().ok());\n    Some((pid, birth))\n}\n\n/// Kill the recorded waker pid and retire the pid file — the\n/// mutual-exclusivity enforcement point and the orphan cleanup\n/// ([`crate::shellhost::launch_shell`] calls this before the binary rises;\n/// the reconciler calls it before adopting an instance a dead daemon left).\n///\n/// The kill is now **authenticated**: it fires only when the pid+birth pair\n/// says the process at that number is the waker we spawned. Gating on bare\n/// `is_process_alive` (what this did before) hands a recycled pid straight to\n/// `taskkill /T` (KNOWN-HAZARDS 7.58).\n///\n/// Same Linux caveat as [`crate::shellhost::kill_shell_at`]: the 10ms jiffy\n/// resolution of `/proc/<pid>/stat` field 22 means the pair NARROWS the\n/// mis-fire window to a pid recycled inside the recorded start's own tick,\n/// rather than eliminating it. Windows `FILETIME` is unaffected.\n// [impl->REQ-SHELL-2]\n// [impl->REQ-SHELL-KILL-AUTHENTICATED]\npub fn kill_waker_at(shell_perch: &Path) {\n    let pid_file = shell_perch.join(WAKER_PID_FILE);\n    let Some((pid, birth)) = read_waker_launch(shell_perch) else {\n        // Absent is the ordinary pre-launch case and says nothing. A record that\n        // EXISTS and will not parse is a different animal: refuse, say so, and\n        // LEAVE it — retiring it here would orphan a possibly-live waker while\n        // reporting success.\n        if pid_file.exists() {\n            spt_proto::emit_line_err!(\n                \"WAKER_KILL_REFUSED_UNREADABLE:{}: a waker record exists but no pid could be \\\n                 parsed from it — refusing the kill and keeping the record\",\n                shell_perch.display()\n            );\n        }\n        return;\n    };\n    if pid == 0 {\n        // A backend that exposed no pid: never a process, never a kill target.\n        let _ = std::fs::remove_file(&pid_file);\n        return;\n    }\n    // The SAME policy the shell chokepoint applies — one destructive decision,\n    // made in one place.\n    match crate::shellhost::kill_verdict(spt_store::liveness::relay_liveness(Some(pid), birth)) {\n        crate::shellhost::KillVerdict::Fire => {\n            crate::shellhost::kill_shell_pid(pid);\n            let _ = std::fs::remove_file(&pid_file);\n        }\n        crate::shellhost::KillVerdict::RetireOnly => {\n            spt_proto::emit_line_err!(\n                \"WAKER_KILL_REFUSED_GONE:{}: recorded waker pid={pid} is not the process we \\\n                 spawned (absent, or the pid was recycled) — retiring the record and killing \\\n                 nothing\",\n                shell_perch.display()\n            );\n            let _ = std::fs::remove_file(&pid_file);\n        }\n        crate::shellhost::KillVerdict::RefuseAndKeep => {\n            spt_proto::emit_line_err!(\n                \"WAKER_KILL_REFUSED_UNPROVEN:{}: recorded waker pid={pid} could not be \\\n                 identified — refusing the kill and keeping the record; the waker may still be \\\n                 running\",\n                shell_perch.display()\n            );\n        }\n    }\n}\n\n/// Clear the give-up latch — any (re)launch is \"next shell activity\".\npub fn clear_gave_up(shell_perch: &Path) {\n    let _ = std::fs::remove_file(shell_perch.join(WAKER_GAVE_UP_FILE));\n}\n\n/// One watcher's supervision loop (runs on its own thread): spawn the\n/// tokenized `wake_command`, park the pid, wait; `exit(WAKE_OPCODE)` ⇒ run\n/// `resolve` once and finish; any other exit ⇒ respawn with exponential\n/// backoff until the give-up budget, then drop the durable latch. A raised\n/// `stop` flag ends the loop without counting the (killed) child as a crash.\n// [impl->REQ-SHELL-2]\npub fn watcher_run(\n    owlery: &Path,\n    owner: &str,\n    shell_id: &str,\n    tokens: &[String],\n    params: &WakeParams,\n    stop: &AtomicBool,\n    resolve: impl FnOnce() -> Result<String, String>,\n) {\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n    let Some((program, args)) = tokens.split_first() else {\n        spt_proto::emit_line_err!(\"WAKER_EMPTY:{owner}/{shell_id}: empty wake_command\");\n        return;\n    };\n    let mut failures = 0u32;\n    while !stop.load(Ordering::SeqCst) {\n        let mut cmd = Command::new(program);\n        cmd.args(args)\n            .stdin(Stdio::null())\n            .stdout(Stdio::null())\n            .stderr(Stdio::null());\n        // The daemon (a console-less DETACHED process) hosting a\n        // console-subsystem waker would otherwise pop a visible Terminal\n        // window per spawn on Windows 11 (the KH 5.6-adjacent D3e lesson —\n        // mock-shell's spt_cmd does the same).\n        #[cfg(windows)]\n        {\n            use std::os::windows::process::CommandExt;\n            cmd.creation_flags(0x0800_0000); // CREATE_NO_WINDOW\n        }\n        let child = cmd.spawn();\n        let mut child = match child {\n            Ok(c) => c,\n            Err(e) => {\n                // An unspawnable waker is a permanent condition — latch now.\n                spt_proto::emit_line_err!(\"WAKER_SPAWN_FAIL:{owner}/{shell_id}: {e} (giving up)\");\n                let _ = std::fs::write(perch.join(WAKER_GAVE_UP_FILE), b\"\");\n                return;\n            }\n        };\n        // The birth stamp rides WITH the pid, in one write — the kill path\n        // authenticates the pair, and a stamp parked separately would leave a\n        // window that reads Held. [impl->REQ-SHELL-KILL-AUTHENTICATED]\n        record_waker_launch(&perch, child.id());\n        let status = child.wait();\n        let _ = std::fs::remove_file(perch.join(WAKER_PID_FILE));\n        if stop.load(Ordering::SeqCst) {\n            return; // stopped from outside — the kill is not a crash\n        }\n        match status.ok().and_then(|s| s.code()) {\n            Some(WAKE_OPCODE) => {\n                match resolve() {\n                    Ok(what) => spt_proto::emit_line_err!(\"SHELL_WAKE:{owner}/{shell_id}: {what}\"),\n                    Err(e) => spt_proto::emit_line_err!(\"SHELL_WAKE_FAIL:{owner}/{shell_id}: {e}\"),\n                }\n                return;\n            }\n            _ => {\n                failures += 1;\n…\n            }\n        }\n    }\n}\n\n[Showing lines 1-323 of 2161. Use :324 to continue]"}],"details":{"truncation":{"content":"//! Shell wake-watcher hosting (M5-D4b — CONTEXT §Shell sleep/wake): the\r\n//! offline half of the online/offline mutual exclusivity. While a shell\r\n//! instance is **offline** and its manifest declares a `wake_command`, the\r\n//! daemon runs that template as a long-running **wake-watcher** child whose\r\n//! sole job is to fire a wake; while the shell binary runs, no watcher does\r\n//! ([`reconcile_once`] flips between them — spt-core owns the flip).\r\n//!\r\n//! **Exit-opcode supervision:** the watcher exiting with [`WAKE_OPCODE`] ⇒\r\n//! the wake resolution ([`resolve_wake`]) brings the shell online. Any other\r\n//! exit is a crash ⇒ respawn with exponential backoff ([`backoff_ms`]) and\r\n//! **give-up** after [`WakeParams::give_up_after`] consecutive crashes (a\r\n//! durable [`WAKER_GAVE_UP_FILE`] marker the next shell activity — any\r\n//! (re)launch — clears; crash-bug safety, the watcher can never crash-loop\r\n//! forever). One watcher per offline instance: the in-process [`WakeSet`]\r\n//! entry plus the perch's [`WAKER_PID_FILE`] (which also lets a freshly\r\n//! booted daemon kill a dead daemon's orphaned watcher before adopting).\r\n//!\r\n//! **Scheduling (KH 7.4):** every watcher gets its own supervision thread —\r\n//! a hung waker stalls nothing; the reconcile loop itself is one bounded\r\n//! sweep per tick off the daemon's control surfaces. Watcher children are\r\n//! *supervised* (waited) children, never detached-immortal — the KH 5.6\r\n//! pipe-inherit shape (an unwaitable child holding a capture pipe) does not\r\n//! arise; stdio is null regardless.\r\n// [impl->REQ-SHELL-2]\r\n\r\nuse std::collections::HashMap;\r\nuse std::path::Path;\r\nuse std::process::{Command, Stdio};\r\nuse std::sync::atomic::{AtomicBool, Ordering};\r\nuse std::sync::{Arc, Mutex};\r\nuse std::thread::JoinHandle;\r\nuse std::time::Duration;\r\n\r\nuse spt_runtime::manifest::{AdapterKind, Manifest, Shell};\r\nuse spt_runtime::registry::AdapterRecord;\r\nuse spt_store::perch::ParentHint;\r\nuse spt_store::shellinfo::{self, SHELL_STATUS_OFFLINE, SHELL_STATUS_ONLINE};\r\n\r\n/// The wake opcode (documented in `docs/MANIFEST.md` §Shell adapters): a\r\n/// wake-watcher exiting with **86** fires the wake resolution; every other\r\n/// exit is a crash.\r\npub const WAKE_OPCODE: i32 = 86;\r\n\r\n/// The running watcher's OS pid, parked on the shell perch — the\r\n/// one-watcher-per-instance lock and the cross-daemon-restart kill handle.\r\npub const WAKER_PID_FILE: &str = \"waker.pid\";\r\n\r\n/// The give-up latch: present ⇒ the watcher crash-looped past its budget and\r\n/// the reconciler must NOT restart it. Cleared by the next shell activity\r\n/// (any (re)launch — [`crate::shellhost::launch_shell`]).\r\npub const WAKER_GAVE_UP_FILE: &str = \"waker.gaveup\";\r\n\r\n/// Watcher supervision knobs (injectable for tests; defaults are production).\r\n#[derive(Debug, Clone, Copy)]\r\npub struct WakeParams {\r\n    /// First-crash respawn delay; doubles per consecutive crash.\r\n    pub backoff_base_ms: u64,\r\n    /// Backoff ceiling.\r\n    pub backoff_cap_ms: u64,\r\n    /// Consecutive crash-exits before the give-up latch drops.\r\n    pub give_up_after: u32,\r\n}\r\n\r\nimpl Default for WakeParams {\r\n    fn default() -> Self {\r\n        Self {\r\n            backoff_base_ms: 1_000,\r\n            backoff_cap_ms: 60_000,\r\n            give_up_after: 6,\r\n        }\r\n    }\r\n}\r\n\r\n/// Exponential backoff for `failures` consecutive crash-exits:\r\n/// `base × 2^(failures-1)`, saturating at the cap (the\r\n/// `spt-net::pairing::ratelimit` recipe). Zero failures ⇒ no wait.\r\n// [impl->REQ-SHELL-2]\r\npub fn backoff_ms(params: &WakeParams, failures: u32) -> u64 {\r\n    if failures == 0 {\r\n        return 0;\r\n    }\r\n    params\r\n        .backoff_base_ms\r\n        .saturating_mul(1u64 << (failures - 1).min(20))\r\n        .min(params.backoff_cap_ms)\r\n}\r\n\r\n/// The live watcher registry: one supervision thread per offline instance,\r\n/// keyed `owner/shell_id`. Shared between the daemon's reconcile loop and\r\n/// anything that must stop a watcher.\r\n#[derive(Default)]\r\npub struct WakeSet {\r\n    inner: Mutex<HashMap<(String, String), WatcherHandle>>,\r\n}\r\n\r\nstruct WatcherHandle {\r\n    stop: Arc<AtomicBool>,\r\n    thread: JoinHandle<()>,\r\n}\r\n\r\nimpl WakeSet {\r\n    pub fn new() -> Self {\r\n        Self::default()\r\n    }\r\n\r\n    /// How many watchers are live (finished threads pruned first).\r\n    pub fn len(&self) -> usize {\r\n        let mut map = self.inner.lock().unwrap_or_else(|p| p.into_inner());\r\n        map.retain(|_, h| !h.thread.is_finished());\r\n        map.len()\r\n    }\r\n\r\n    pub fn is_empty(&self) -> bool {\r\n        self.len() == 0\r\n    }\r\n\r\n    fn contains(&self, owner: &str, shell_id: &str) -> bool {\r\n        let mut map = self.inner.lock().unwrap_or_else(|p| p.into_inner());\r\n        match map.get(&(owner.to_string(), shell_id.to_string())) {\r\n            Some(h) if !h.thread.is_finished() => true,\r\n            Some(_) => {\r\n                map.remove(&(owner.to_string(), shell_id.to_string()));\r\n                false\r\n            }\r\n            None => false,\r\n        }\r\n    }\r\n\r\n    fn insert(&self, owner: &str, shell_id: &str, handle: WatcherHandle) {\r\n        let mut map = self.inner.lock().unwrap_or_else(|p| p.into_inner());\r\n        map.insert((owner.to_string(), shell_id.to_string()), handle);\r\n    }\r\n\r\n    /// Signal one watcher to stop and kill its child (the perch pid file is\r\n    /// the kill handle). The thread sees the kill as a non-opcode exit with\r\n    /// its stop flag up and returns without counting a failure.\r\n    fn stop_watcher(&self, owlery: &Path, owner: &str, shell_id: &str) {\r\n        let handle = {\r\n            let mut map = self.inner.lock().unwrap_or_else(|p| p.into_inner());\r\n            map.remove(&(owner.to_string(), shell_id.to_string()))\r\n        };\r\n        if let Some(h) = handle {\r\n            h.stop.store(true, Ordering::SeqCst);\r\n            let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\r\n            kill_waker_at(&perch);\r\n            let _ = h.thread.join();\r\n        }\r\n    }\r\n}\r\n\r\n/// Park the waker's pid **and its birth stamp** as one record — `<pid>\\n<birth>`\r\n/// — in the single write that already existed.\r\n///\r\n/// ONE file and ONE write, because the pair must never be torn: a stamp parked\r\n/// in a second file leaves a window where the pid is on disk and the stamp is\r\n/// not, and [`spt_store::liveness::relay_liveness`] reads a live pid with no\r\n/// stamp as `Held`. That window is exactly the unauthenticated kill this\r\n/// requirement removes, reintroduced as a race.\r\n///\r\n/// A backend that exposes no start time writes the pid alone — the same\r\n/// one-line shape every pre-fix record already has, which reads `Held` on\r\n/// existence and keeps today's behaviour until the next launch.\r\n// [impl->REQ-SHELL-KILL-AUTHENTICATED]\r\npub fn record_waker_launch(shell_perch: &Path, pid: u32) {\r\n    let body = match spt_store::proc::process_started_at(pid) {\r\n        Some(birth) => format!(\"{pid}\\n{birth}\"),\r\n        None => pid.to_string(),\r\n    };\r\n    let _ = std::fs::write(shell_perch.join(WAKER_PID_FILE), body);\r\n}\r\n\r\n/// Read that record back — **the** parse, so the two-line format has exactly one\r\n/// reader shape (the `record_shell_launch`/`read_shell_launch` pattern).\r\n///\r\n/// `None` means \"no usable pid on record\": absent, empty, or unparseable. Every\r\n/// caller must treat that as FAIL-TOWARD-ALIVE — refuse the kill and leave the\r\n/// record standing. A whole-content `trim().parse()` (what every reader did\r\n/// before this) returns `None` on the two-line form, so a reader left un-migrated\r\n/// would kill nothing, retire the record anyway, and report success while\r\n/// orphaning a live waker.\r\n// [impl->REQ-SHELL-KILL-AUTHENTICATED]\r\npub fn read_waker_launch(shell_perch: &Path) -> Option<(u32, Option<u64>)> {\r\n    let raw = std::fs::read_to_string(shell_perch.join(WAKER_PID_FILE)).ok()?;\r\n    let mut lines = raw.lines();\r\n    let pid = lines.next()?.trim().parse::<u32>().ok()?;\r\n    let birth = lines.next().and_then(|l| l.trim().parse::<u64>().ok());\r\n    Some((pid, birth))\r\n}\r\n\r\n/// Kill the recorded waker pid and retire the pid file — the\r\n/// mutual-exclusivity enforcement point and the orphan cleanup\r\n/// ([`crate::shellhost::launch_shell`] calls this before the binary rises;\r\n/// the reconciler calls it before adopting an instance a dead daemon left).\r\n///\r\n/// The kill is now **authenticated**: it fires only when the pid+birth pair\r\n/// says the process at that number is the waker we spawned. Gating on bare\r\n/// `is_process_alive` (what this did before) hands a recycled pid straight to\r\n/// `taskkill /T` (KNOWN-HAZARDS 7.58).\r\n///\r\n/// Same Linux caveat as [`crate::shellhost::kill_shell_at`]: the 10ms jiffy\r\n/// resolution of `/proc/<pid>/stat` field 22 means the pair NARROWS the\r\n/// mis-fire window to a pid recycled inside the recorded start's own tick,\r\n/// rather than eliminating it. Windows `FILETIME` is unaffected.\r\n// [impl->REQ-SHELL-2]\r\n// [impl->REQ-SHELL-KILL-AUTHENTICATED]\r\npub fn kill_waker_at(shell_perch: &Path) {\r\n    let pid_file = shell_perch.join(WAKER_PID_FILE);\r\n    let Some((pid, birth)) = read_waker_launch(shell_perch) else {\r\n        // Absent is the ordinary pre-launch case and says nothing. A record that\r\n        // EXISTS and will not parse is a different animal: refuse, say so, and\r\n        // LEAVE it — retiring it here would orphan a possibly-live waker while\r\n        // reporting success.\r\n        if pid_file.exists() {\r\n            spt_proto::emit_line_err!(\r\n                \"WAKER_KILL_REFUSED_UNREADABLE:{}: a waker record exists but no pid could be \\\r\n                 parsed from it — refusing the kill and keeping the record\",\r\n                shell_perch.display()\r\n            );\r\n        }\r\n        return;\r\n    };\r\n    if pid == 0 {\r\n        // A backend that exposed no pid: never a process, never a kill target.\r\n        let _ = std::fs::remove_file(&pid_file);\r\n        return;\r\n    }\r\n    // The SAME policy the shell chokepoint applies — one destructive decision,\r\n    // made in one place.\r\n    match crate::shellhost::kill_verdict(spt_store::liveness::relay_liveness(Some(pid), birth)) {\r\n        crate::shellhost::KillVerdict::Fire => {\r\n            crate::shellhost::kill_shell_pid(pid);\r\n            let _ = std::fs::remove_file(&pid_file);\r\n        }\r\n        crate::shellhost::KillVerdict::RetireOnly => {\r\n            spt_proto::emit_line_err!(\r\n                \"WAKER_KILL_REFUSED_GONE:{}: recorded waker pid={pid} is not the process we \\\r\n                 spawned (absent, or the pid was recycled) — retiring the record and killing \\\r\n                 nothing\",\r\n                shell_perch.display()\r\n            );\r\n            let _ = std::fs::remove_file(&pid_file);\r\n        }\r\n        crate::shellhost::KillVerdict::RefuseAndKeep => {\r\n            spt_proto::emit_line_err!(\r\n                \"WAKER_KILL_REFUSED_UNPROVEN:{}: recorded waker pid={pid} could not be \\\r\n                 identified — refusing the kill and keeping the record; the waker may still be \\\r\n                 running\",\r\n                shell_perch.display()\r\n            );\r\n        }\r\n    }\r\n}\r\n\r\n/// Clear the give-up latch — any (re)launch is \"next shell activity\".\r\npub fn clear_gave_up(shell_perch: &Path) {\r\n    let _ = std::fs::remove_file(shell_perch.join(WAKER_GAVE_UP_FILE));\r\n}\r\n\r\n/// One watcher's supervision loop (runs on its own thread): spawn the\r\n/// tokenized `wake_command`, park the pid, wait; `exit(WAKE_OPCODE)` ⇒ run\r\n/// `resolve` once and finish; any other exit ⇒ respawn with exponential\r\n/// backoff until the give-up budget, then drop the durable latch. A raised\r\n/// `stop` flag ends the loop without counting the (killed) child as a crash.\r\n// [impl->REQ-SHELL-2]\r\npub fn watcher_run(\r\n    owlery: &Path,\r\n    owner: &str,\r\n    shell_id: &str,\r\n    tokens: &[String],\r\n    params: &WakeParams,\r\n    stop: &AtomicBool,\r\n    resolve: impl FnOnce() -> Result<String, String>,\r\n) {\r\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\r\n    let Some((program, args)) = tokens.split_first() else {\r\n        spt_proto::emit_line_err!(\"WAKER_EMPTY:{owner}/{shell_id}: empty wake_command\");\r\n        return;\r\n    };\r\n    let mut failures = 0u32;\r\n    while !stop.load(Ordering::SeqCst) {\r\n        let mut cmd = Command::new(program);\r\n        cmd.args(args)\r\n            .stdin(Stdio::null())\r\n            .stdout(Stdio::null())\r\n            .stderr(Stdio::null());\r\n        // The daemon (a console-less DETACHED process) hosting a\r\n        // console-subsystem waker would otherwise pop a visible Terminal\r\n        // window per spawn on Windows 11 (the KH 5.6-adjacent D3e lesson —\r\n        // mock-shell's spt_cmd does the same).\r\n        #[cfg(windows)]\r\n        {\r\n            use std::os::windows::process::CommandExt;\r\n            cmd.creation_flags(0x0800_0000); // CREATE_NO_WINDOW\r\n        }\r\n        let child = cmd.spawn();\r\n        let mut child = match child {\r\n            Ok(c) => c,\r\n            Err(e) => {\r\n                // An unspawnable waker is a permanent condition — latch now.\r\n                spt_proto::emit_line_err!(\"WAKER_SPAWN_FAIL:{owner}/{shell_id}: {e} (giving up)\");\r\n                let _ = std::fs::write(perch.join(WAKER_GAVE_UP_FILE), b\"\");\r\n                return;\r\n            }\r\n        };\r\n        // The birth stamp rides WITH the pid, in one write — the kill path\r\n        // authenticates the pair, and a stamp parked separately would leave a\r\n        // window that reads Held. [impl->REQ-SHELL-KILL-AUTHENTICATED]\r\n        record_waker_launch(&perch, child.id());\r\n        let status = child.wait();\r\n        let _ = std::fs::remove_file(perch.join(WAKER_PID_FILE));\r\n        if stop.load(Ordering::SeqCst) {\r\n            return; // stopped from outside — the kill is not a crash\r\n        }\r\n        match status.ok().and_then(|s| s.code()) {\r\n            Some(WAKE_OPCODE) => {\r\n                match resolve() {\r\n                    Ok(what) => spt_proto::emit_line_err!(\"SHELL_WAKE:{owner}/{shell_id}: {what}\"),\r\n                    Err(e) => spt_proto::emit_line_err!(\"SHELL_WAKE_FAIL:{owner}/{shell_id}: {e}\"),\r\n                }\r\n                return;\r\n            }\r\n            _ => {\r\n                failures += 1;\r","truncated":true,"truncatedBy":"lines","totalLines":2161,"totalBytes":14340,"outputLines":323,"outputBytes":14340,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":2161,"displayContent":{"text":"//! Shell wake-watcher hosting (M5-D4b — CONTEXT §Shell sleep/wake): the\r\n//! offline half of the online/offline mutual exclusivity. While a shell\r\n//! instance is **offline** and its manifest declares a `wake_command`, the\r\n//! daemon runs that template as a long-running **wake-watcher** child whose\r\n//! sole job is to fire a wake; while the shell binary runs, no watcher does\r\n//! ([`reconcile_once`] flips between them — spt-core owns the flip).\r\n//!\r\n//! **Exit-opcode supervision:** the watcher exiting with [`WAKE_OPCODE`] ⇒\r\n//! the wake resolution ([`resolve_wake`]) brings the shell online. Any other\r\n//! exit is a crash ⇒ respawn with exponential backoff ([`backoff_ms`]) and\r\n//! **give-up** after [`WakeParams::give_up_after`] consecutive crashes (a\r\n//! durable [`WAKER_GAVE_UP_FILE`] marker the next shell activity — any\r\n//! (re)launch — clears; crash-bug safety, the watcher can never crash-loop\r\n//! forever). One watcher per offline instance: the in-process [`WakeSet`]\r\n//! entry plus the perch's [`WAKER_PID_FILE`] (which also lets a freshly\r\n//! booted daemon kill a dead daemon's orphaned watcher before adopting).\r\n//!\r\n//! **Scheduling (KH 7.4):** every watcher gets its own supervision thread —\r\n//! a hung waker stalls nothing; the reconcile loop itself is one bounded\r\n//! sweep per tick off the daemon's control surfaces. Watcher children are\r\n//! *supervised* (waited) children, never detached-immortal — the KH 5.6\r\n//! pipe-inherit shape (an unwaitable child holding a capture pipe) does not\r\n//! arise; stdio is null regardless.\r\n// [impl->REQ-SHELL-2]\r\n\r\nuse std::collections::HashMap;\r\nuse std::path::Path;\r\nuse std::process::{Command, Stdio};\r\nuse std::sync::atomic::{AtomicBool, Ordering};\r\nuse std::sync::{Arc, Mutex};\r\nuse std::thread::JoinHandle;\r\nuse std::time::Duration;\r\n\r\nuse spt_runtime::manifest::{AdapterKind, Manifest, Shell};\r\nuse spt_runtime::registry::AdapterRecord;\r\nuse spt_store::perch::ParentHint;\r\nuse spt_store::shellinfo::{self, SHELL_STATUS_OFFLINE, SHELL_STATUS_ONLINE};\r\n\r\n/// The wake opcode (documented in `docs/MANIFEST.md` §Shell adapters): a\r\n/// wake-watcher exiting with **86** fires the wake resolution; every other\r\n/// exit is a crash.\r\npub const WAKE_OPCODE: i32 = 86;\r\n\r\n/// The running watcher's OS pid, parked on the shell perch — the\r\n/// one-watcher-per-instance lock and the cross-daemon-restart kill handle.\r\npub const WAKER_PID_FILE: &str = \"waker.pid\";\r\n\r\n/// The give-up latch: present ⇒ the watcher crash-looped past its budget and\r\n/// the reconciler must NOT restart it. Cleared by the next shell activity\r\n/// (any (re)launch — [`crate::shellhost::launch_shell`]).\r\npub const WAKER_GAVE_UP_FILE: &str = \"waker.gaveup\";\r\n\r\n/// Watcher supervision knobs (injectable for tests; defaults are production).\r\n#[derive(Debug, Clone, Copy)]\r\npub struct WakeParams {\r\n    /// First-crash respawn delay; doubles per consecutive crash.\r\n    pub backoff_base_ms: u64,\r\n    /// Backoff ceiling.\r\n    pub backoff_cap_ms: u64,\r\n    /// Consecutive crash-exits before the give-up latch drops.\r\n    pub give_up_after: u32,\r\n}\r\n\r\nimpl Default for WakeParams {\r\n    fn default() -> Self {\r\n        Self {\r\n            backoff_base_ms: 1_000,\r\n            backoff_cap_ms: 60_000,\r\n            give_up_after: 6,\r\n        }\r\n    }\r\n}\r\n\r\n/// Exponential backoff for `failures` consecutive crash-exits:\r\n/// `base × 2^(failures-1)`, saturating at the cap (the\r\n/// `spt-net::pairing::ratelimit` recipe). Zero failures ⇒ no wait.\r\n// [impl->REQ-SHELL-2]\r\npub fn backoff_ms(params: &WakeParams, failures: u32) -> u64 {\r\n    if failures == 0 {\r\n        return 0;\r\n    }\r\n    params\r\n        .backoff_base_ms\r\n        .saturating_mul(1u64 << (failures - 1).min(20))\r\n        .min(params.backoff_cap_ms)\r\n}\r\n\r\n/// The live watcher registry: one supervision thread per offline instance,\r\n/// keyed `owner/shell_id`. Shared between the daemon's reconcile loop and\r\n/// anything that must stop a watcher.\r\n#[derive(Default)]\r\npub struct WakeSet {\r\n    inner: Mutex<HashMap<(String, String), WatcherHandle>>,\r\n}\r\n\r\nstruct WatcherHandle {\r\n    stop: Arc<AtomicBool>,\r\n    thread: JoinHandle<()>,\r\n}\r\n\r\nimpl WakeSet {\r\n    pub fn new() -> Self {\r\n        Self::default()\r\n    }\r\n\r\n    /// How many watchers are live (finished threads pruned first).\r\n    pub fn len(&self) -> usize {\r\n        let mut map = self.inner.lock().unwrap_or_else(|p| p.into_inner());\r\n        map.retain(|_, h| !h.thread.is_finished());\r\n        map.len()\r\n    }\r\n\r\n    pub fn is_empty(&self) -> bool {\r\n        self.len() == 0\r\n    }\r\n\r\n    fn contains(&self, owner: &str, shell_id: &str) -> bool {\r\n        let mut map = self.inner.lock().unwrap_or_else(|p| p.into_inner());\r\n        match map.get(&(owner.to_string(), shell_id.to_string())) {\r\n            Some(h) if !h.thread.is_finished() => true,\r\n            Some(_) => {\r\n                map.remove(&(owner.to_string(), shell_id.to_string()));\r\n                false\r\n            }\r\n            None => false,\r\n        }\r\n    }\r\n\r\n    fn insert(&self, owner: &str, shell_id: &str, handle: WatcherHandle) {\r\n        let mut map = self.inner.lock().unwrap_or_else(|p| p.into_inner());\r\n        map.insert((owner.to_string(), shell_id.to_string()), handle);\r\n    }\r\n\r\n    /// Signal one watcher to stop and kill its child (the perch pid file is\r\n    /// the kill handle). The thread sees the kill as a non-opcode exit with\r\n    /// its stop flag up and returns without counting a failure.\r\n    fn stop_watcher(&self, owlery: &Path, owner: &str, shell_id: &str) {\r\n        let handle = {\r\n            let mut map = self.inner.lock().unwrap_or_else(|p| p.into_inner());\r\n            map.remove(&(owner.to_string(), shell_id.to_string()))\r\n        };\r\n        if let Some(h) = handle {\r\n            h.stop.store(true, Ordering::SeqCst);\r\n            let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\r\n            kill_waker_at(&perch);\r\n            let _ = h.thread.join();\r\n        }\r\n    }\r\n}\r\n\r\n/// Park the waker's pid **and its birth stamp** as one record — `<pid>\\n<birth>`\r\n/// — in the single write that already existed.\r\n///\r\n/// ONE file and ONE write, because the pair must never be torn: a stamp parked\r\n/// in a second file leaves a window where the pid is on disk and the stamp is\r\n/// not, and [`spt_store::liveness::relay_liveness`] reads a live pid with no\r\n/// stamp as `Held`. That window is exactly the unauthenticated kill this\r\n/// requirement removes, reintroduced as a race.\r\n///\r\n/// A backend that exposes no start time writes the pid alone — the same\r\n/// one-line shape every pre-fix record already has, which reads `Held` on\r\n/// existence and keeps today's behaviour until the next launch.\r\n// [impl->REQ-SHELL-KILL-AUTHENTICATED]\r\npub fn record_waker_launch(shell_perch: &Path, pid: u32) {\r\n    let body = match spt_store::proc::process_started_at(pid) {\r\n        Some(birth) => format!(\"{pid}\\n{birth}\"),\r\n        None => pid.to_string(),\r\n    };\r\n    let _ = std::fs::write(shell_perch.join(WAKER_PID_FILE), body);\r\n}\r\n\r\n/// Read that record back — **the** parse, so the two-line format has exactly one\r\n/// reader shape (the `record_shell_launch`/`read_shell_launch` pattern).\r\n///\r\n/// `None` means \"no usable pid on record\": absent, empty, or unparseable. Every\r\n/// caller must treat that as FAIL-TOWARD-ALIVE — refuse the kill and leave the\r\n/// record standing. A whole-content `trim().parse()` (what every reader did\r\n/// before this) returns `None` on the two-line form, so a reader left un-migrated\r\n/// would kill nothing, retire the record anyway, and report success while\r\n/// orphaning a live waker.\r\n// [impl->REQ-SHELL-KILL-AUTHENTICATED]\r\npub fn read_waker_launch(shell_perch: &Path) -> Option<(u32, Option<u64>)> {\r\n    let raw = std::fs::read_to_string(shell_perch.join(WAKER_PID_FILE)).ok()?;\r\n    let mut lines = raw.lines();\r\n    let pid = lines.next()?.trim().parse::<u32>().ok()?;\r\n    let birth = lines.next().and_then(|l| l.trim().parse::<u64>().ok());\r\n    Some((pid, birth))\r\n}\r\n\r\n/// Kill the recorded waker pid and retire the pid file — the\r\n/// mutual-exclusivity enforcement point and the orphan cleanup\r\n/// ([`crate::shellhost::launch_shell`] calls this before the binary rises;\r\n/// the reconciler calls it before adopting an instance a dead daemon left).\r\n///\r\n/// The kill is now **authenticated**: it fires only when the pid+birth pair\r\n/// says the process at that number is the waker we spawned. Gating on bare\r\n/// `is_process_alive` (what this did before) hands a recycled pid straight to\r\n/// `taskkill /T` (KNOWN-HAZARDS 7.58).\r\n///\r\n/// Same Linux caveat as [`crate::shellhost::kill_shell_at`]: the 10ms jiffy\r\n/// resolution of `/proc/<pid>/stat` field 22 means the pair NARROWS the\r\n/// mis-fire window to a pid recycled inside the recorded start's own tick,\r\n/// rather than eliminating it. Windows `FILETIME` is unaffected.\r\n// [impl->REQ-SHELL-2]\r\n// [impl->REQ-SHELL-KILL-AUTHENTICATED]\r\npub fn kill_waker_at(shell_perch: &Path) {\r\n    let pid_file = shell_perch.join(WAKER_PID_FILE);\r\n    let Some((pid, birth)) = read_waker_launch(shell_perch) else {\r\n        // Absent is the ordinary pre-launch case and says nothing. A record that\r\n        // EXISTS and will not parse is a different animal: refuse, say so, and\r\n        // LEAVE it — retiring it here would orphan a possibly-live waker while\r\n        // reporting success.\r\n        if pid_file.exists() {\r\n            spt_proto::emit_line_err!(\r\n                \"WAKER_KILL_REFUSED_UNREADABLE:{}: a waker record exists but no pid could be \\\r\n                 parsed from it — refusing the kill and keeping the record\",\r\n                shell_perch.display()\r\n            );\r\n        }\r\n        return;\r\n    };\r\n    if pid == 0 {\r\n        // A backend that exposed no pid: never a process, never a kill target.\r\n        let _ = std::fs::remove_file(&pid_file);\r\n        return;\r\n    }\r\n    // The SAME policy the shell chokepoint applies — one destructive decision,\r\n    // made in one place.\r\n    match crate::shellhost::kill_verdict(spt_store::liveness::relay_liveness(Some(pid), birth)) {\r\n        crate::shellhost::KillVerdict::Fire => {\r\n            crate::shellhost::kill_shell_pid(pid);\r\n            let _ = std::fs::remove_file(&pid_file);\r\n        }\r\n        crate::shellhost::KillVerdict::RetireOnly => {\r\n            spt_proto::emit_line_err!(\r\n                \"WAKER_KILL_REFUSED_GONE:{}: recorded waker pid={pid} is not the process we \\\r\n                 spawned (absent, or the pid was recycled) — retiring the record and killing \\\r\n                 nothing\",\r\n                shell_perch.display()\r\n            );\r\n            let _ = std::fs::remove_file(&pid_file);\r\n        }\r\n        crate::shellhost::KillVerdict::RefuseAndKeep => {\r\n            spt_proto::emit_line_err!(\r\n                \"WAKER_KILL_REFUSED_UNPROVEN:{}: recorded waker pid={pid} could not be \\\r\n                 identified — refusing the kill and keeping the record; the waker may still be \\\r\n                 running\",\r\n                shell_perch.display()\r\n            );\r\n        }\r\n    }\r\n}\r\n\r\n/// Clear the give-up latch — any (re)launch is \"next shell activity\".\r\npub fn clear_gave_up(shell_perch: &Path) {\r\n    let _ = std::fs::remove_file(shell_perch.join(WAKER_GAVE_UP_FILE));\r\n}\r\n\r\n/// One watcher's supervision loop (runs on its own thread): spawn the\r\n/// tokenized `wake_command`, park the pid, wait; `exit(WAKE_OPCODE)` ⇒ run\r\n/// `resolve` once and finish; any other exit ⇒ respawn with exponential\r\n/// backoff until the give-up budget, then drop the durable latch. A raised\r\n/// `stop` flag ends the loop without counting the (killed) child as a crash.\r\n// [impl->REQ-SHELL-2]\r\npub fn watcher_run(\r\n    owlery: &Path,\r\n    owner: &str,\r\n    shell_id: &str,\r\n    tokens: &[String],\r\n    params: &WakeParams,\r\n    stop: &AtomicBool,\r\n    resolve: impl FnOnce() -> Result<String, String>,\r\n) {\r\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\r\n    let Some((program, args)) = tokens.split_first() else {\r\n        spt_proto::emit_line_err!(\"WAKER_EMPTY:{owner}/{shell_id}: empty wake_command\");\r\n        return;\r\n    };\r\n    let mut failures = 0u32;\r\n    while !stop.load(Ordering::SeqCst) {\r\n        let mut cmd = Command::new(program);\r\n        cmd.args(args)\r\n            .stdin(Stdio::null())\r\n            .stdout(Stdio::null())\r\n            .stderr(Stdio::null());\r\n        // The daemon (a console-less DETACHED process) hosting a\r\n        // console-subsystem waker would otherwise pop a visible Terminal\r\n        // window per spawn on Windows 11 (the KH 5.6-adjacent D3e lesson —\r\n        // mock-shell's spt_cmd does the same).\r\n        #[cfg(windows)]\r\n        {\r\n            use std::os::windows::process::CommandExt;\r\n            cmd.creation_flags(0x0800_0000); // CREATE_NO_WINDOW\r\n        }\r\n        let child = cmd.spawn();\r\n        let mut child = match child {\r\n            Ok(c) => c,\r\n            Err(e) => {\r\n                // An unspawnable waker is a permanent condition — latch now.\r\n                spt_proto::emit_line_err!(\"WAKER_SPAWN_FAIL:{owner}/{shell_id}: {e} (giving up)\");\r\n                let _ = std::fs::write(perch.join(WAKER_GAVE_UP_FILE), b\"\");\r\n                return;\r\n            }\r\n        };\r\n        // The birth stamp rides WITH the pid, in one write — the kill path\r\n        // authenticates the pair, and a stamp parked separately would leave a\r\n        // window that reads Held. [impl->REQ-SHELL-KILL-AUTHENTICATED]\r\n        record_waker_launch(&perch, child.id());\r\n        let status = child.wait();\r\n        let _ = std::fs::remove_file(perch.join(WAKER_PID_FILE));\r\n        if stop.load(Ordering::SeqCst) {\r\n            return; // stopped from outside — the kill is not a crash\r\n        }\r\n        match status.ok().and_then(|s| s.code()) {\r\n            Some(WAKE_OPCODE) => {\r\n                match resolve() {\r\n                    Ok(what) => spt_proto::emit_line_err!(\"SHELL_WAKE:{owner}/{shell_id}: {what}\"),\r\n                    Err(e) => spt_proto::emit_line_err!(\"SHELL_WAKE_FAIL:{owner}/{shell_id}: {e}\"),\r\n                }\r\n                return;\r\n            }\r\n            _ => {\r\n                failures += 1;\r\n…\n            }\n        }\n    }\n}","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120,121,122,123,124,125,126,127,128,129,130,131,132,133,134,135,136,137,138,139,140,141,142,143,144,145,146,147,148,149,150,151,152,153,154,155,156,157,158,159,160,161,162,163,164,165,166,167,168,169,170,171,172,173,174,175,176,177,178,179,180,181,182,183,184,185,186,187,188,189,190,191,192,193,194,195,196,197,198,199,200,201,202,203,204,205,206,207,208,209,210,211,212,213,214,215,216,217,218,219,220,221,222,223,224,225,226,227,228,229,230,231,232,233,234,235,236,237,238,239,240,241,242,243,244,245,246,247,248,249,250,251,252,253,254,255,256,257,258,259,260,261,262,263,264,265,266,267,268,269,270,271,272,273,274,275,276,277,278,279,280,281,282,283,284,285,286,287,288,289,290,291,292,293,294,295,296,297,298,299,300,301,302,303,304,305,306,307,308,309,310,311,312,313,314,315,316,317,318,319,320,321,322,323,null,339,340,341,342]},"fileSize":100821,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\crates\\spt-daemon\\src\\shellwake.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":2161,"totalBytes":14340,"outputLines":323,"outputBytes":14340,"shownRange":{"start":1,"end":323},"nextOffset":324}}},"isError":false,"timestamp":1789001136956}}
{"type":"message","id":"ac0f6d03","parentId":"fb18ed21","timestamp":"2026-09-10T00:45:36.964Z","message":{"role":"toolResult","toolCallId":"call_Af4TRJYugyLpvR3kOYPeRRR4|fc_077f3569365e313d016aa1fdb0c08c87d1b16af7f4dfd838ca","toolName":"read","content":[{"type":"text","text":"//! Shell-instance perch records (M5-D3a — CONTEXT §Shell model,\n//! `docs/STORAGE.md` §Shells in the layout).\n//!\n//! A shell perch is **not** an agent perch: its `info.json` carries exactly\n//! `type=Shell, owner, adapter_name, status, alias?` — a dedicated\n//! [`ShellInfo`] shape, deliberately distinct from [`crate::info::InfoJson`],\n//! so the agent-perch readers (liveness, registry stale-clean, the consent\n//! MRA scan) stay blind to shells by construction. Capabilities are **never**\n//! duplicated here — they resolve from the shell adapter manifest by\n//! `adapter_name`.\n//!\n//! **Identity model (REQ-EP-5):** `spt shell spawn` *mints* a new instance —\n//! the immutable canonical id `<adapter>-<n>` ([`mint_shell_id`]: smallest\n//! free `n` over **existing** instance dirs; online and offline both occupy,\n//! `teardown` frees the slot). The optional **alias** is a per-instance,\n//! owner-unique display/address overlay set at spawn or by `shell rename` —\n//! interchangeable with the canonical id for addressing\n//! ([`resolve_shell_ref`], canonical-first), never obscuring `adapter_name`.\n//!\n//! Torn-read stance mirrors `info.rs` (KNOWN-HAZARDS 1.2): writers go through\n//! the atomic rename; readers tolerate failure and return `None`.\n// [impl->REQ-EP-5]\n\nuse serde::{Deserialize, Serialize};\nuse std::path::Path;\n\nuse crate::atomic::atomic_write_string;\nuse crate::perch;\n\n/// `status` value: the shell binary is linked and running.\npub const SHELL_STATUS_ONLINE: &str = \"online\";\n/// `status` value: the perch exists, no binary linked (re-linkable if the\n/// adapter is `persistent`; the spawn-time initial state until D3b binds).\npub const SHELL_STATUS_OFFLINE: &str = \"offline\";\n/// The fixed `type` tag on every shell perch record.\npub const SHELL_TYPE: &str = \"Shell\";\n\n/// On-disk shape of a shell perch's `info.json` — the full field set, by\n/// contract (STORAGE.md: \"carries only `type=Shell`, `owner`, `adapter_name`,\n/// `status`, and an optional `alias`\").\n// [impl->REQ-SHELL-1]\n#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]\npub struct ShellInfo {\n    /// Always [`SHELL_TYPE`] — readers refuse anything else.\n    pub r#type: String,\n    /// The owning agent endpoint id — the only id that may drive this shell.\n    pub owner: String,\n    /// The providing shell adapter (capability surface resolves from its\n    /// manifest, never from the perch).\n    pub adapter_name: String,\n    /// [`SHELL_STATUS_ONLINE`] | [`SHELL_STATUS_OFFLINE`].\n    pub status: String,\n    /// Owner-unique friendly label (runtime state, not a manifest field).\n    #[serde(default, skip_serializing_if = \"Option::is_none\")]\n    pub alias: Option<String>,\n}\n\nimpl ShellInfo {\n    /// A freshly-minted instance record: offline until the broker launch +\n    /// `api bind` handshake onlines it (D3b).\n    pub fn new(owner: &str, adapter_name: &str) -> Self {\n        ShellInfo {\n            r#type: SHELL_TYPE.to_string(),\n            owner: owner.to_string(),\n            adapter_name: adapter_name.to_string(),\n            status: SHELL_STATUS_OFFLINE.to_string(),\n            alias: None,\n        }\n    }\n}\n\n/// Shell-instance model failures — CLI-facing, so they carry the why.\n#[derive(Debug)]\npub enum ShellError {\n    /// The alias is already held by another of this owner's instances, or\n    /// shadows an existing canonical id.\n    AliasTaken(String),\n    /// No instance of this owner matches the ref (canonical id or alias).\n    NotFound(String),\n    Io(std::io::Error),\n}\n\nimpl std::fmt::Display for ShellError {\n    fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {\n        match self {\n            ShellError::AliasTaken(a) => write!(f, \"alias already in use: {a}\"),\n            ShellError::NotFound(r) => write!(f, \"no shell instance matches: {r}\"),\n            ShellError::Io(e) => write!(f, \"io: {e}\"),\n        }\n    }\n}\n\nimpl From<std::io::Error> for ShellError {\n    fn from(e: std::io::Error) -> Self {\n        ShellError::Io(e)\n    }\n}\n\n/// Read a shell perch's record. Tolerant (`None` on missing/corrupt — the\n/// list path skips, never panics) and shape-checked: a record whose `type`\n/// isn't [`SHELL_TYPE`] is not a shell perch.\npub fn read_shell_info(shell_perch: &Path) -> Option<ShellInfo> {\n    let s = std::fs::read_to_string(perch::info_file_at(shell_perch)).ok()?;\n    let info: ShellInfo = serde_json::from_str(&s).ok()?;\n    (info.r#type == SHELL_TYPE).then_some(info)\n}\n\n/// Write a shell perch's record (atomic rename — the 1.2 torn-read stance).\n/// Creates the perch dir.\npub fn write_shell_info(shell_perch: &Path, info: &ShellInfo) -> std::io::Result<()> {\n    std::fs::create_dir_all(shell_perch)?;\n    let s = serde_json::to_string_pretty(info).expect(\"ShellInfo serializes\");\n    atomic_write_string(&perch::info_file_at(shell_perch), &s)\n}\n\n/// The launched binary's recorded pid, one file per instance perch. Written by\n/// the launch, removed by the close — and the ONLY on-disk evidence that the\n/// process a record calls `online` still exists.\npub const SHELL_PID_FILE: &str = \"shell.pid\";\n\n/// The recorded launch pid, if the perch parks a parseable one.\npub fn read_shell_pid(shell_perch: &Path) -> Option<u32> {\n    std::fs::read_to_string(shell_perch.join(SHELL_PID_FILE))\n        .ok()?\n        .trim()\n        .parse()\n        .ok()\n}\n\n/// `<shell_perch>/shell.launch.json` — the birth stamp parked beside the pid at\n/// launch (BAROMETER W2, releases#78).\npub const SHELL_LAUNCH_FILE: &str = \"shell.launch.json\";\n\n/// What a launch recorded about the process it started.\n///\n/// TWO VALUES, ANSWERING TWO DIFFERENT QUESTIONS, and keeping them apart is the\n/// whole point of the shape:\n///\n/// - `pid_started_at` is the platform's NATIVE start-time value, used for one\n///   question only — the same-pid-different-process PAIR test within a single\n///   boot ([`crate::liveness::relay_liveness`]). It is deliberately never\n///   compared across boots, because it is not the same kind of number on every\n///   platform: on Windows it is an absolute FILETIME, but on Linux it is field\n///   22 of `/proc/<pid>/stat`, i.e. JIFFIES SINCE BOOT, which resets at every\n///   boot. A value recorded before a restart is therefore indistinguishable\n///   from a fresh one after it, and any cross-boot comparison of this field\n///   would be correct on Windows and silently wrong on Linux.\n/// - `launched_ms` is an ABSOLUTE wall-clock stamp (epoch-ms), used by the\n///   nonpersistent watcher's cross-boot freeze. Persistent restoration does not\n///   gate on launch age; the native birth stamp still authenticates liveness.\n#[derive(Debug, Clone, Default, PartialEq, Eq, Serialize, Deserialize)]\npub struct ShellLaunch {\n    /// The native start-time of the launched pid, when the platform offers one.\n    #[serde(default, skip_serializing_if = \"Option::is_none\")]\n    pub pid_started_at: Option<u64>,\n    /// Wall-clock epoch-ms at which this instance was launched.\n    pub launched_ms: u64,\n}\n\n/// Record the birth stamp for a just-launched `pid`. Best-effort: a stamp that\n/// cannot be written must never fail the launch — a missing stamp degrades the\n/// probe to its pre-existing fail-toward-alive behaviour, which is the safe\n/// direction, while a failed launch is a real outage.\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\npub fn record_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) {\n    let launch = ShellLaunch {\n        pid_started_at: crate::proc::process_started_at(pid),\n        launched_ms: now_ms,\n    };\n    if let Ok(json) = serde_json::to_string_pretty(&launch) {\n        let _ = atomic_write_string(&shell_perch.join(SHELL_LAUNCH_FILE), &json);\n    }\n}\n\n/// Read the birth stamp, if this instance was launched by a build that parks one.\n/// Absent or unparseable ⇒ `None`, which every consumer treats as \"no stamp to\n/// compare\" rather than as a fault.\npub fn read_shell_launch(shell_perch: &Path) -> Option<ShellLaunch> {\n    let raw = std::fs::read_to_string(shell_perch.join(SHELL_LAUNCH_FILE)).ok()?;\n    serde_json::from_str(&raw).ok()\n}\n\n/// **The** discriminant behind [`effective_status`]: is this instance's recorded\n/// binary *provably* gone? Narrow on purpose — true only when the perch parks a\n/// pid, it parses non-zero, and the process is not alive.\n///\n/// Everything else reads **alive**: an absent or unparseable pid file, and the\n/// `0` a broker-hosted spawn records when its backend exposed no pid\n/// ([`crate::proc::is_process_alive`] already reads 0 as dead, so the explicit\n/// filter here is what keeps a pid-less backend from being declared dead by this\n/// resolver). That fail-toward-alive stance mirrors `liveness.rs`'s interim\n/// parity: this predicate exists to catch a *corpse*, never to guess at a\n/// backend it cannot see.\n///\n/// Recycled pids read alive, so a heal can be **missed** but never **mis-fired**\n/// — the failure direction is \"stays stale\", never \"tears down a live instance\".\n// [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\npub fn shell_pid_provably_dead(shell_perch: &Path) -> bool {\n    let Some(pid) = read_shell_pid(shell_perch).filter(|&pid| pid != 0) else {\n        return false;\n    };\n    // BAROMETER W2 (releases#78): probe pid AND birth stamp, not the pid alone.\n    // A pid namespace resets at boot, so a corpse pid can be REUSED by an\n    // unrelated process started after the restart; under bare-pid liveness that\n    // reads ALIVE, the derived status reads online, and `relink` then refuses\n    // SHELL_ALREADY_ONLINE -- the one command that recovers the instance\n    // refusing on the very state that is wrong. The PAIR test (same pid,\n    // different start time) is what tells a reused pid from our own process.\n    //\n    // The stance is unchanged where there is no stamp to compare: an instance\n    // launched before this field existed, or a backend that exposed no start\n    // time, still falls back to plain absence and fails toward ALIVE.\n    // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n    let birth = read_shell_launch(shell_perch).and_then(|l| l.pid_started_at);\n    matches!(\n        crate::liveness::relay_liveness(Some(pid), birth),\n        crate::liveness::RelayLiveness::Gone\n    )\n}\n\n/// A shell instance's **derived** status — the shell-side twin of\n/// `liveness::is_perch_alive` (REQ-HAZARD-DAEMON-HOSTED-LIVENESS gave agent\n/// perches exactly this; shells never got it).\n///\n/// The recorded `status` has exactly two writers — the bind handshake (→online)\n/// and `spt_daemon::shellhost::close_shell` (→offline) — and `close_shell` runs\n/// on the **link-break** path alone. A binary that dies\n/// abruptly (force-kill, crash, OOM) breaks no link, so nothing ever writes the\n/// flip and the record says `online` forever. Deriving closes that hole without\n/// adding a supervisor: `online` + a provably-dead pid reads\n/// [`SHELL_STATUS_OFFLINE`]; every other case is the recorded value verbatim.\n///\n/// Callers that must NOT derive are documented at their sites (the two status\n/// writers, the owner-suspend cascade whose close IS the cleanup, and the wake\n/// reconciler's eligibility read) — see REQ-HAZARD-SHELL-STALE-ONLINE.\n// [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\npub fn effective_status<'a>(shell_perch: &Path, info: &'a ShellInfo) -> &'a str {\n    if info.status == SHELL_STATUS_ONLINE && shell_pid_provably_dead(shell_perch) {\n        return SHELL_STATUS_OFFLINE;\n    }\n    &info.status\n}\n\n/// [`effective_status`] as the predicate every gate actually asks: is this\n/// instance online *in fact*, not merely on paper?\n// [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\npub fn is_shell_online(shell_perch: &Path, info: &ShellInfo) -> bool {\n    effective_status(shell_perch, info) == SHELL_STATUS_ONLINE\n}\n\n/// The **parent** segment of an adapter option: the part before the first `:`\n/// of a `<adapter>:<profile>` composite (CONTEXT.md §adapter profile). The\n/// canonical shell id is minted off this — never the composite — so the id (a\n/// perch dir name + endpoint address) stays colon-free and path-safe; the full\n/// composite rides only in the `adapter_name` carrier field. Mirrors\n/// `spt_runtime::profile::split_option`, inlined here to keep spt-store free of\n/// an spt-runtime dependency (the crate graph runs the other way).\nfn adapter_parent(option: &str) -> &str {\n    option.split_once(':').map_or(option, |(a, _)| a)\n}\n\n/// Parse the `n` out of a canonical `<parent>-<n>` dir name for `parent`.\nfn instance_ordinal(dir_name: &str, parent: &str) -> Option<u32> {\n    let suffix = dir_name.strip_prefix(parent)?.strip_prefix('-')?;\n    if suffix.is_empty() || !suffix.bytes().all(|b| b.is_ascii_digit()) {\n        return None;\n    }\n    suffix.parse().ok()\n}\n\n/// Mint the next canonical id for `adapter` under `owner`: the **smallest\n/// free `n`** over existing instance dirs. Online and offline both occupy a\n/// slot (every non-torn-down perch counts — the same all-existing rule the\n/// instance cap uses); `teardown` removes the dir, so its `n` is reused.\n///\n/// `adapter` may be a `<adapter>:<profile>` composite; the id is minted off the\n/// **parent** segment ([`adapter_parent`]) so two profiles of one adapter share\n/// the same `<parent>-<n>` ordinal space (instances are the adapter's resource,\n/// not the profile's) and the id never carries the `:`.\n// [impl->REQ-EP-5]\n// [impl->REQ-HAZARD-ID-CHARSET]\npub fn mint_shell_id(owlery: &Path, owner: &str, adapter: &str) -> String {\n    let parent = adapter_parent(adapter);\n    let shells = perch::resolve_shells_dir_in(owlery, owner);\n    let mut taken: Vec<u32> = std::fs::read_dir(&shells)\n        .map(|entries| {\n            entries\n                .flatten()\n                .filter(|e| e.file_type().map(|t| t.is_dir()).unwrap_or(false))\n                .filter_map(|e| instance_ordinal(&e.file_name().to_string_lossy(), parent))\n                .collect()\n        })\n        .unwrap_or_default();\n    taken.sort_unstable();\n    let mut n = 0u32;\n    for t in taken {\n        if t == n {\n            n += 1;\n        } else if t > n {\n            break;\n        }\n    }\n    format!(\"{parent}-{n}\")\n}\n\n/// Every shell instance of `owner`: `(canonical_id, record)`, sorted by id\n/// for a stable order. Tolerant — a dir without a valid [`ShellInfo`] is\n/// skipped (it still *occupies* its mint slot; it just doesn't list).\n// [impl->REQ-EP-5]\npub fn list_shells(owlery: &Path, owner: &str) -> Vec<(String, ShellInfo)> {\n    let shells = perch::resolve_shells_dir_in(owlery, owner);\n    let Ok(entries) = std::fs::read_dir(&shells) else {\n        return Vec::new();\n    };\n    let mut out: Vec<(String, ShellInfo)> = entries\n        .flatten()\n        .filter(|e| e.file_type().map(|t| t.is_dir()).unwrap_or(false))\n        .filter_map(|e| {\n            let id = e.file_name().to_string_lossy().into_owned();\n            read_shell_info(&e.path()).map(|info| (id, info))\n        })\n        .collect();\n    out.sort_by(|a, b| a.0.cmp(&b.0));\n    out\n}\n\n/// Resolve a shell ref — canonical id or alias — to `(canonical_id, record)`.\n/// **Canonical-first**: an alias can never shadow a canonical id (spawn/rename\n/// refuse minting such an alias, and resolution checks canonical before alias\n/// as defense in depth). Aliases are owner-unique, so at most one matches.\n// [impl->REQ-EP-5]\npub fn resolve_shell_ref(\n    owlery: &Path,\n    owner: &str,\n    shell_ref: &str,\n) -> Option<(String, ShellInfo)> {\n    let direct = perch::resolve_shell_perch_path_in(owlery, owner, shell_ref);\n    if let Some(info) = read_shell_info(&direct) {\n        return Some((shell_ref.to_string(), info));\n    }\n    list_shells(owlery, owner)\n        .into_iter()\n        .find(|(_, info)| info.alias.as_deref() == Some(shell_ref))\n}\n\n/// Whether `alias` is free for `owner`: held by no instance's alias and not\n/// shadowing any existing canonical id.\nfn alias_free(owlery: &Path, owner: &str, alias: &str) -> bool {\n    resolve_shell_ref(owlery, owner, alias).is_none()\n}\n\n/// Mint + record a new instance of `adapter` for `owner` (the storage half of\n/// `spt shell spawn` — adapter-registration and consent gates are the\n\n[Showing lines 1-353 of 726. Use :354 to continue]"}],"details":{"truncation":{"content":"//! Shell-instance perch records (M5-D3a — CONTEXT §Shell model,\r\n//! `docs/STORAGE.md` §Shells in the layout).\r\n//!\r\n//! A shell perch is **not** an agent perch: its `info.json` carries exactly\r\n//! `type=Shell, owner, adapter_name, status, alias?` — a dedicated\r\n//! [`ShellInfo`] shape, deliberately distinct from [`crate::info::InfoJson`],\r\n//! so the agent-perch readers (liveness, registry stale-clean, the consent\r\n//! MRA scan) stay blind to shells by construction. Capabilities are **never**\r\n//! duplicated here — they resolve from the shell adapter manifest by\r\n//! `adapter_name`.\r\n//!\r\n//! **Identity model (REQ-EP-5):** `spt shell spawn` *mints* a new instance —\r\n//! the immutable canonical id `<adapter>-<n>` ([`mint_shell_id`]: smallest\r\n//! free `n` over **existing** instance dirs; online and offline both occupy,\r\n//! `teardown` frees the slot). The optional **alias** is a per-instance,\r\n//! owner-unique display/address overlay set at spawn or by `shell rename` —\r\n//! interchangeable with the canonical id for addressing\r\n//! ([`resolve_shell_ref`], canonical-first), never obscuring `adapter_name`.\r\n//!\r\n//! Torn-read stance mirrors `info.rs` (KNOWN-HAZARDS 1.2): writers go through\r\n//! the atomic rename; readers tolerate failure and return `None`.\r\n// [impl->REQ-EP-5]\r\n\r\nuse serde::{Deserialize, Serialize};\r\nuse std::path::Path;\r\n\r\nuse crate::atomic::atomic_write_string;\r\nuse crate::perch;\r\n\r\n/// `status` value: the shell binary is linked and running.\r\npub const SHELL_STATUS_ONLINE: &str = \"online\";\r\n/// `status` value: the perch exists, no binary linked (re-linkable if the\r\n/// adapter is `persistent`; the spawn-time initial state until D3b binds).\r\npub const SHELL_STATUS_OFFLINE: &str = \"offline\";\r\n/// The fixed `type` tag on every shell perch record.\r\npub const SHELL_TYPE: &str = \"Shell\";\r\n\r\n/// On-disk shape of a shell perch's `info.json` — the full field set, by\r\n/// contract (STORAGE.md: \"carries only `type=Shell`, `owner`, `adapter_name`,\r\n/// `status`, and an optional `alias`\").\r\n// [impl->REQ-SHELL-1]\r\n#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]\r\npub struct ShellInfo {\r\n    /// Always [`SHELL_TYPE`] — readers refuse anything else.\r\n    pub r#type: String,\r\n    /// The owning agent endpoint id — the only id that may drive this shell.\r\n    pub owner: String,\r\n    /// The providing shell adapter (capability surface resolves from its\r\n    /// manifest, never from the perch).\r\n    pub adapter_name: String,\r\n    /// [`SHELL_STATUS_ONLINE`] | [`SHELL_STATUS_OFFLINE`].\r\n    pub status: String,\r\n    /// Owner-unique friendly label (runtime state, not a manifest field).\r\n    #[serde(default, skip_serializing_if = \"Option::is_none\")]\r\n    pub alias: Option<String>,\r\n}\r\n\r\nimpl ShellInfo {\r\n    /// A freshly-minted instance record: offline until the broker launch +\r\n    /// `api bind` handshake onlines it (D3b).\r\n    pub fn new(owner: &str, adapter_name: &str) -> Self {\r\n        ShellInfo {\r\n            r#type: SHELL_TYPE.to_string(),\r\n            owner: owner.to_string(),\r\n            adapter_name: adapter_name.to_string(),\r\n            status: SHELL_STATUS_OFFLINE.to_string(),\r\n            alias: None,\r\n        }\r\n    }\r\n}\r\n\r\n/// Shell-instance model failures — CLI-facing, so they carry the why.\r\n#[derive(Debug)]\r\npub enum ShellError {\r\n    /// The alias is already held by another of this owner's instances, or\r\n    /// shadows an existing canonical id.\r\n    AliasTaken(String),\r\n    /// No instance of this owner matches the ref (canonical id or alias).\r\n    NotFound(String),\r\n    Io(std::io::Error),\r\n}\r\n\r\nimpl std::fmt::Display for ShellError {\r\n    fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {\r\n        match self {\r\n            ShellError::AliasTaken(a) => write!(f, \"alias already in use: {a}\"),\r\n            ShellError::NotFound(r) => write!(f, \"no shell instance matches: {r}\"),\r\n            ShellError::Io(e) => write!(f, \"io: {e}\"),\r\n        }\r\n    }\r\n}\r\n\r\nimpl From<std::io::Error> for ShellError {\r\n    fn from(e: std::io::Error) -> Self {\r\n        ShellError::Io(e)\r\n    }\r\n}\r\n\r\n/// Read a shell perch's record. Tolerant (`None` on missing/corrupt — the\r\n/// list path skips, never panics) and shape-checked: a record whose `type`\r\n/// isn't [`SHELL_TYPE`] is not a shell perch.\r\npub fn read_shell_info(shell_perch: &Path) -> Option<ShellInfo> {\r\n    let s = std::fs::read_to_string(perch::info_file_at(shell_perch)).ok()?;\r\n    let info: ShellInfo = serde_json::from_str(&s).ok()?;\r\n    (info.r#type == SHELL_TYPE).then_some(info)\r\n}\r\n\r\n/// Write a shell perch's record (atomic rename — the 1.2 torn-read stance).\r\n/// Creates the perch dir.\r\npub fn write_shell_info(shell_perch: &Path, info: &ShellInfo) -> std::io::Result<()> {\r\n    std::fs::create_dir_all(shell_perch)?;\r\n    let s = serde_json::to_string_pretty(info).expect(\"ShellInfo serializes\");\r\n    atomic_write_string(&perch::info_file_at(shell_perch), &s)\r\n}\r\n\r\n/// The launched binary's recorded pid, one file per instance perch. Written by\r\n/// the launch, removed by the close — and the ONLY on-disk evidence that the\r\n/// process a record calls `online` still exists.\r\npub const SHELL_PID_FILE: &str = \"shell.pid\";\r\n\r\n/// The recorded launch pid, if the perch parks a parseable one.\r\npub fn read_shell_pid(shell_perch: &Path) -> Option<u32> {\r\n    std::fs::read_to_string(shell_perch.join(SHELL_PID_FILE))\r\n        .ok()?\r\n        .trim()\r\n        .parse()\r\n        .ok()\r\n}\r\n\r\n/// `<shell_perch>/shell.launch.json` — the birth stamp parked beside the pid at\r\n/// launch (BAROMETER W2, releases#78).\r\npub const SHELL_LAUNCH_FILE: &str = \"shell.launch.json\";\r\n\r\n/// What a launch recorded about the process it started.\r\n///\r\n/// TWO VALUES, ANSWERING TWO DIFFERENT QUESTIONS, and keeping them apart is the\r\n/// whole point of the shape:\r\n///\r\n/// - `pid_started_at` is the platform's NATIVE start-time value, used for one\r\n///   question only — the same-pid-different-process PAIR test within a single\r\n///   boot ([`crate::liveness::relay_liveness`]). It is deliberately never\r\n///   compared across boots, because it is not the same kind of number on every\r\n///   platform: on Windows it is an absolute FILETIME, but on Linux it is field\r\n///   22 of `/proc/<pid>/stat`, i.e. JIFFIES SINCE BOOT, which resets at every\r\n///   boot. A value recorded before a restart is therefore indistinguishable\r\n///   from a fresh one after it, and any cross-boot comparison of this field\r\n///   would be correct on Windows and silently wrong on Linux.\r\n/// - `launched_ms` is an ABSOLUTE wall-clock stamp (epoch-ms), used by the\r\n///   nonpersistent watcher's cross-boot freeze. Persistent restoration does not\r\n///   gate on launch age; the native birth stamp still authenticates liveness.\r\n#[derive(Debug, Clone, Default, PartialEq, Eq, Serialize, Deserialize)]\r\npub struct ShellLaunch {\r\n    /// The native start-time of the launched pid, when the platform offers one.\r\n    #[serde(default, skip_serializing_if = \"Option::is_none\")]\r\n    pub pid_started_at: Option<u64>,\r\n    /// Wall-clock epoch-ms at which this instance was launched.\r\n    pub launched_ms: u64,\r\n}\r\n\r\n/// Record the birth stamp for a just-launched `pid`. Best-effort: a stamp that\r\n/// cannot be written must never fail the launch — a missing stamp degrades the\r\n/// probe to its pre-existing fail-toward-alive behaviour, which is the safe\r\n/// direction, while a failed launch is a real outage.\r\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\npub fn record_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) {\r\n    let launch = ShellLaunch {\r\n        pid_started_at: crate::proc::process_started_at(pid),\r\n        launched_ms: now_ms,\r\n    };\r\n    if let Ok(json) = serde_json::to_string_pretty(&launch) {\r\n        let _ = atomic_write_string(&shell_perch.join(SHELL_LAUNCH_FILE), &json);\r\n    }\r\n}\r\n\r\n/// Read the birth stamp, if this instance was launched by a build that parks one.\r\n/// Absent or unparseable ⇒ `None`, which every consumer treats as \"no stamp to\r\n/// compare\" rather than as a fault.\r\npub fn read_shell_launch(shell_perch: &Path) -> Option<ShellLaunch> {\r\n    let raw = std::fs::read_to_string(shell_perch.join(SHELL_LAUNCH_FILE)).ok()?;\r\n    serde_json::from_str(&raw).ok()\r\n}\r\n\r\n/// **The** discriminant behind [`effective_status`]: is this instance's recorded\r\n/// binary *provably* gone? Narrow on purpose — true only when the perch parks a\r\n/// pid, it parses non-zero, and the process is not alive.\r\n///\r\n/// Everything else reads **alive**: an absent or unparseable pid file, and the\r\n/// `0` a broker-hosted spawn records when its backend exposed no pid\r\n/// ([`crate::proc::is_process_alive`] already reads 0 as dead, so the explicit\r\n/// filter here is what keeps a pid-less backend from being declared dead by this\r\n/// resolver). That fail-toward-alive stance mirrors `liveness.rs`'s interim\r\n/// parity: this predicate exists to catch a *corpse*, never to guess at a\r\n/// backend it cannot see.\r\n///\r\n/// Recycled pids read alive, so a heal can be **missed** but never **mis-fired**\r\n/// — the failure direction is \"stays stale\", never \"tears down a live instance\".\r\n// [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\r\npub fn shell_pid_provably_dead(shell_perch: &Path) -> bool {\r\n    let Some(pid) = read_shell_pid(shell_perch).filter(|&pid| pid != 0) else {\r\n        return false;\r\n    };\r\n    // BAROMETER W2 (releases#78): probe pid AND birth stamp, not the pid alone.\r\n    // A pid namespace resets at boot, so a corpse pid can be REUSED by an\r\n    // unrelated process started after the restart; under bare-pid liveness that\r\n    // reads ALIVE, the derived status reads online, and `relink` then refuses\r\n    // SHELL_ALREADY_ONLINE -- the one command that recovers the instance\r\n    // refusing on the very state that is wrong. The PAIR test (same pid,\r\n    // different start time) is what tells a reused pid from our own process.\r\n    //\r\n    // The stance is unchanged where there is no stamp to compare: an instance\r\n    // launched before this field existed, or a backend that exposed no start\r\n    // time, still falls back to plain absence and fails toward ALIVE.\r\n    // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\n    let birth = read_shell_launch(shell_perch).and_then(|l| l.pid_started_at);\r\n    matches!(\r\n        crate::liveness::relay_liveness(Some(pid), birth),\r\n        crate::liveness::RelayLiveness::Gone\r\n    )\r\n}\r\n\r\n/// A shell instance's **derived** status — the shell-side twin of\r\n/// `liveness::is_perch_alive` (REQ-HAZARD-DAEMON-HOSTED-LIVENESS gave agent\r\n/// perches exactly this; shells never got it).\r\n///\r\n/// The recorded `status` has exactly two writers — the bind handshake (→online)\r\n/// and `spt_daemon::shellhost::close_shell` (→offline) — and `close_shell` runs\r\n/// on the **link-break** path alone. A binary that dies\r\n/// abruptly (force-kill, crash, OOM) breaks no link, so nothing ever writes the\r\n/// flip and the record says `online` forever. Deriving closes that hole without\r\n/// adding a supervisor: `online` + a provably-dead pid reads\r\n/// [`SHELL_STATUS_OFFLINE`]; every other case is the recorded value verbatim.\r\n///\r\n/// Callers that must NOT derive are documented at their sites (the two status\r\n/// writers, the owner-suspend cascade whose close IS the cleanup, and the wake\r\n/// reconciler's eligibility read) — see REQ-HAZARD-SHELL-STALE-ONLINE.\r\n// [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\r\npub fn effective_status<'a>(shell_perch: &Path, info: &'a ShellInfo) -> &'a str {\r\n    if info.status == SHELL_STATUS_ONLINE && shell_pid_provably_dead(shell_perch) {\r\n        return SHELL_STATUS_OFFLINE;\r\n    }\r\n    &info.status\r\n}\r\n\r\n/// [`effective_status`] as the predicate every gate actually asks: is this\r\n/// instance online *in fact*, not merely on paper?\r\n// [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\r\npub fn is_shell_online(shell_perch: &Path, info: &ShellInfo) -> bool {\r\n    effective_status(shell_perch, info) == SHELL_STATUS_ONLINE\r\n}\r\n\r\n/// The **parent** segment of an adapter option: the part before the first `:`\r\n/// of a `<adapter>:<profile>` composite (CONTEXT.md §adapter profile). The\r\n/// canonical shell id is minted off this — never the composite — so the id (a\r\n/// perch dir name + endpoint address) stays colon-free and path-safe; the full\r\n/// composite rides only in the `adapter_name` carrier field. Mirrors\r\n/// `spt_runtime::profile::split_option`, inlined here to keep spt-store free of\r\n/// an spt-runtime dependency (the crate graph runs the other way).\r\nfn adapter_parent(option: &str) -> &str {\r\n    option.split_once(':').map_or(option, |(a, _)| a)\r\n}\r\n\r\n/// Parse the `n` out of a canonical `<parent>-<n>` dir name for `parent`.\r\nfn instance_ordinal(dir_name: &str, parent: &str) -> Option<u32> {\r\n    let suffix = dir_name.strip_prefix(parent)?.strip_prefix('-')?;\r\n    if suffix.is_empty() || !suffix.bytes().all(|b| b.is_ascii_digit()) {\r\n        return None;\r\n    }\r\n    suffix.parse().ok()\r\n}\r\n\r\n/// Mint the next canonical id for `adapter` under `owner`: the **smallest\r\n/// free `n`** over existing instance dirs. Online and offline both occupy a\r\n/// slot (every non-torn-down perch counts — the same all-existing rule the\r\n/// instance cap uses); `teardown` removes the dir, so its `n` is reused.\r\n///\r\n/// `adapter` may be a `<adapter>:<profile>` composite; the id is minted off the\r\n/// **parent** segment ([`adapter_parent`]) so two profiles of one adapter share\r\n/// the same `<parent>-<n>` ordinal space (instances are the adapter's resource,\r\n/// not the profile's) and the id never carries the `:`.\r\n// [impl->REQ-EP-5]\r\n// [impl->REQ-HAZARD-ID-CHARSET]\r\npub fn mint_shell_id(owlery: &Path, owner: &str, adapter: &str) -> String {\r\n    let parent = adapter_parent(adapter);\r\n    let shells = perch::resolve_shells_dir_in(owlery, owner);\r\n    let mut taken: Vec<u32> = std::fs::read_dir(&shells)\r\n        .map(|entries| {\r\n            entries\r\n                .flatten()\r\n                .filter(|e| e.file_type().map(|t| t.is_dir()).unwrap_or(false))\r\n                .filter_map(|e| instance_ordinal(&e.file_name().to_string_lossy(), parent))\r\n                .collect()\r\n        })\r\n        .unwrap_or_default();\r\n    taken.sort_unstable();\r\n    let mut n = 0u32;\r\n    for t in taken {\r\n        if t == n {\r\n            n += 1;\r\n        } else if t > n {\r\n            break;\r\n        }\r\n    }\r\n    format!(\"{parent}-{n}\")\r\n}\r\n\r\n/// Every shell instance of `owner`: `(canonical_id, record)`, sorted by id\r\n/// for a stable order. Tolerant — a dir without a valid [`ShellInfo`] is\r\n/// skipped (it still *occupies* its mint slot; it just doesn't list).\r\n// [impl->REQ-EP-5]\r\npub fn list_shells(owlery: &Path, owner: &str) -> Vec<(String, ShellInfo)> {\r\n    let shells = perch::resolve_shells_dir_in(owlery, owner);\r\n    let Ok(entries) = std::fs::read_dir(&shells) else {\r\n        return Vec::new();\r\n    };\r\n    let mut out: Vec<(String, ShellInfo)> = entries\r\n        .flatten()\r\n        .filter(|e| e.file_type().map(|t| t.is_dir()).unwrap_or(false))\r\n        .filter_map(|e| {\r\n            let id = e.file_name().to_string_lossy().into_owned();\r\n            read_shell_info(&e.path()).map(|info| (id, info))\r\n        })\r\n        .collect();\r\n    out.sort_by(|a, b| a.0.cmp(&b.0));\r\n    out\r\n}\r\n\r\n/// Resolve a shell ref — canonical id or alias — to `(canonical_id, record)`.\r\n/// **Canonical-first**: an alias can never shadow a canonical id (spawn/rename\r\n/// refuse minting such an alias, and resolution checks canonical before alias\r\n/// as defense in depth). Aliases are owner-unique, so at most one matches.\r\n// [impl->REQ-EP-5]\r\npub fn resolve_shell_ref(\r\n    owlery: &Path,\r\n    owner: &str,\r\n    shell_ref: &str,\r\n) -> Option<(String, ShellInfo)> {\r\n    let direct = perch::resolve_shell_perch_path_in(owlery, owner, shell_ref);\r\n    if let Some(info) = read_shell_info(&direct) {\r\n        return Some((shell_ref.to_string(), info));\r\n    }\r\n    list_shells(owlery, owner)\r\n        .into_iter()\r\n        .find(|(_, info)| info.alias.as_deref() == Some(shell_ref))\r\n}\r\n\r\n/// Whether `alias` is free for `owner`: held by no instance's alias and not\r\n/// shadowing any existing canonical id.\r\nfn alias_free(owlery: &Path, owner: &str, alias: &str) -> bool {\r\n    resolve_shell_ref(owlery, owner, alias).is_none()\r\n}\r\n\r\n/// Mint + record a new instance of `adapter` for `owner` (the storage half of\r\n/// `spt shell spawn` — adapter-registration and consent gates are the\r","truncated":true,"truncatedBy":"lines","totalLines":726,"totalBytes":16529,"outputLines":353,"outputBytes":16529,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":726,"displayContent":{"text":"//! Shell-instance perch records (M5-D3a — CONTEXT §Shell model,\r\n//! `docs/STORAGE.md` §Shells in the layout).\r\n//!\r\n//! A shell perch is **not** an agent perch: its `info.json` carries exactly\r\n//! `type=Shell, owner, adapter_name, status, alias?` — a dedicated\r\n//! [`ShellInfo`] shape, deliberately distinct from [`crate::info::InfoJson`],\r\n//! so the agent-perch readers (liveness, registry stale-clean, the consent\r\n//! MRA scan) stay blind to shells by construction. Capabilities are **never**\r\n//! duplicated here — they resolve from the shell adapter manifest by\r\n//! `adapter_name`.\r\n//!\r\n//! **Identity model (REQ-EP-5):** `spt shell spawn` *mints* a new instance —\r\n//! the immutable canonical id `<adapter>-<n>` ([`mint_shell_id`]: smallest\r\n//! free `n` over **existing** instance dirs; online and offline both occupy,\r\n//! `teardown` frees the slot). The optional **alias** is a per-instance,\r\n//! owner-unique display/address overlay set at spawn or by `shell rename` —\r\n//! interchangeable with the canonical id for addressing\r\n//! ([`resolve_shell_ref`], canonical-first), never obscuring `adapter_name`.\r\n//!\r\n//! Torn-read stance mirrors `info.rs` (KNOWN-HAZARDS 1.2): writers go through\r\n//! the atomic rename; readers tolerate failure and return `None`.\r\n// [impl->REQ-EP-5]\r\n\r\nuse serde::{Deserialize, Serialize};\r\nuse std::path::Path;\r\n\r\nuse crate::atomic::atomic_write_string;\r\nuse crate::perch;\r\n\r\n/// `status` value: the shell binary is linked and running.\r\npub const SHELL_STATUS_ONLINE: &str = \"online\";\r\n/// `status` value: the perch exists, no binary linked (re-linkable if the\r\n/// adapter is `persistent`; the spawn-time initial state until D3b binds).\r\npub const SHELL_STATUS_OFFLINE: &str = \"offline\";\r\n/// The fixed `type` tag on every shell perch record.\r\npub const SHELL_TYPE: &str = \"Shell\";\r\n\r\n/// On-disk shape of a shell perch's `info.json` — the full field set, by\r\n/// contract (STORAGE.md: \"carries only `type=Shell`, `owner`, `adapter_name`,\r\n/// `status`, and an optional `alias`\").\r\n// [impl->REQ-SHELL-1]\r\n#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]\r\npub struct ShellInfo {\r\n    /// Always [`SHELL_TYPE`] — readers refuse anything else.\r\n    pub r#type: String,\r\n    /// The owning agent endpoint id — the only id that may drive this shell.\r\n    pub owner: String,\r\n    /// The providing shell adapter (capability surface resolves from its\r\n    /// manifest, never from the perch).\r\n    pub adapter_name: String,\r\n    /// [`SHELL_STATUS_ONLINE`] | [`SHELL_STATUS_OFFLINE`].\r\n    pub status: String,\r\n    /// Owner-unique friendly label (runtime state, not a manifest field).\r\n    #[serde(default, skip_serializing_if = \"Option::is_none\")]\r\n    pub alias: Option<String>,\r\n}\r\n\r\nimpl ShellInfo {\r\n    /// A freshly-minted instance record: offline until the broker launch +\r\n    /// `api bind` handshake onlines it (D3b).\r\n    pub fn new(owner: &str, adapter_name: &str) -> Self {\r\n        ShellInfo {\r\n            r#type: SHELL_TYPE.to_string(),\r\n            owner: owner.to_string(),\r\n            adapter_name: adapter_name.to_string(),\r\n            status: SHELL_STATUS_OFFLINE.to_string(),\r\n            alias: None,\r\n        }\r\n    }\r\n}\r\n\r\n/// Shell-instance model failures — CLI-facing, so they carry the why.\r\n#[derive(Debug)]\r\npub enum ShellError {\r\n    /// The alias is already held by another of this owner's instances, or\r\n    /// shadows an existing canonical id.\r\n    AliasTaken(String),\r\n    /// No instance of this owner matches the ref (canonical id or alias).\r\n    NotFound(String),\r\n    Io(std::io::Error),\r\n}\r\n\r\nimpl std::fmt::Display for ShellError {\r\n    fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {\r\n        match self {\r\n            ShellError::AliasTaken(a) => write!(f, \"alias already in use: {a}\"),\r\n            ShellError::NotFound(r) => write!(f, \"no shell instance matches: {r}\"),\r\n            ShellError::Io(e) => write!(f, \"io: {e}\"),\r\n        }\r\n    }\r\n}\r\n\r\nimpl From<std::io::Error> for ShellError {\r\n    fn from(e: std::io::Error) -> Self {\r\n        ShellError::Io(e)\r\n    }\r\n}\r\n\r\n/// Read a shell perch's record. Tolerant (`None` on missing/corrupt — the\r\n/// list path skips, never panics) and shape-checked: a record whose `type`\r\n/// isn't [`SHELL_TYPE`] is not a shell perch.\r\npub fn read_shell_info(shell_perch: &Path) -> Option<ShellInfo> {\r\n    let s = std::fs::read_to_string(perch::info_file_at(shell_perch)).ok()?;\r\n    let info: ShellInfo = serde_json::from_str(&s).ok()?;\r\n    (info.r#type == SHELL_TYPE).then_some(info)\r\n}\r\n\r\n/// Write a shell perch's record (atomic rename — the 1.2 torn-read stance).\r\n/// Creates the perch dir.\r\npub fn write_shell_info(shell_perch: &Path, info: &ShellInfo) -> std::io::Result<()> {\r\n    std::fs::create_dir_all(shell_perch)?;\r\n    let s = serde_json::to_string_pretty(info).expect(\"ShellInfo serializes\");\r\n    atomic_write_string(&perch::info_file_at(shell_perch), &s)\r\n}\r\n\r\n/// The launched binary's recorded pid, one file per instance perch. Written by\r\n/// the launch, removed by the close — and the ONLY on-disk evidence that the\r\n/// process a record calls `online` still exists.\r\npub const SHELL_PID_FILE: &str = \"shell.pid\";\r\n\r\n/// The recorded launch pid, if the perch parks a parseable one.\r\npub fn read_shell_pid(shell_perch: &Path) -> Option<u32> {\r\n    std::fs::read_to_string(shell_perch.join(SHELL_PID_FILE))\r\n        .ok()?\r\n        .trim()\r\n        .parse()\r\n        .ok()\r\n}\r\n\r\n/// `<shell_perch>/shell.launch.json` — the birth stamp parked beside the pid at\r\n/// launch (BAROMETER W2, releases#78).\r\npub const SHELL_LAUNCH_FILE: &str = \"shell.launch.json\";\r\n\r\n/// What a launch recorded about the process it started.\r\n///\r\n/// TWO VALUES, ANSWERING TWO DIFFERENT QUESTIONS, and keeping them apart is the\r\n/// whole point of the shape:\r\n///\r\n/// - `pid_started_at` is the platform's NATIVE start-time value, used for one\r\n///   question only — the same-pid-different-process PAIR test within a single\r\n///   boot ([`crate::liveness::relay_liveness`]). It is deliberately never\r\n///   compared across boots, because it is not the same kind of number on every\r\n///   platform: on Windows it is an absolute FILETIME, but on Linux it is field\r\n///   22 of `/proc/<pid>/stat`, i.e. JIFFIES SINCE BOOT, which resets at every\r\n///   boot. A value recorded before a restart is therefore indistinguishable\r\n///   from a fresh one after it, and any cross-boot comparison of this field\r\n///   would be correct on Windows and silently wrong on Linux.\r\n/// - `launched_ms` is an ABSOLUTE wall-clock stamp (epoch-ms), used by the\r\n///   nonpersistent watcher's cross-boot freeze. Persistent restoration does not\r\n///   gate on launch age; the native birth stamp still authenticates liveness.\r\n#[derive(Debug, Clone, Default, PartialEq, Eq, Serialize, Deserialize)]\r\npub struct ShellLaunch {\r\n    /// The native start-time of the launched pid, when the platform offers one.\r\n    #[serde(default, skip_serializing_if = \"Option::is_none\")]\r\n    pub pid_started_at: Option<u64>,\r\n    /// Wall-clock epoch-ms at which this instance was launched.\r\n    pub launched_ms: u64,\r\n}\r\n\r\n/// Record the birth stamp for a just-launched `pid`. Best-effort: a stamp that\r\n/// cannot be written must never fail the launch — a missing stamp degrades the\r\n/// probe to its pre-existing fail-toward-alive behaviour, which is the safe\r\n/// direction, while a failed launch is a real outage.\r\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\npub fn record_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) {\r\n    let launch = ShellLaunch {\r\n        pid_started_at: crate::proc::process_started_at(pid),\r\n        launched_ms: now_ms,\r\n    };\r\n    if let Ok(json) = serde_json::to_string_pretty(&launch) {\r\n        let _ = atomic_write_string(&shell_perch.join(SHELL_LAUNCH_FILE), &json);\r\n    }\r\n}\r\n\r\n/// Read the birth stamp, if this instance was launched by a build that parks one.\r\n/// Absent or unparseable ⇒ `None`, which every consumer treats as \"no stamp to\r\n/// compare\" rather than as a fault.\r\npub fn read_shell_launch(shell_perch: &Path) -> Option<ShellLaunch> {\r\n    let raw = std::fs::read_to_string(shell_perch.join(SHELL_LAUNCH_FILE)).ok()?;\r\n    serde_json::from_str(&raw).ok()\r\n}\r\n\r\n/// **The** discriminant behind [`effective_status`]: is this instance's recorded\r\n/// binary *provably* gone? Narrow on purpose — true only when the perch parks a\r\n/// pid, it parses non-zero, and the process is not alive.\r\n///\r\n/// Everything else reads **alive**: an absent or unparseable pid file, and the\r\n/// `0` a broker-hosted spawn records when its backend exposed no pid\r\n/// ([`crate::proc::is_process_alive`] already reads 0 as dead, so the explicit\r\n/// filter here is what keeps a pid-less backend from being declared dead by this\r\n/// resolver). That fail-toward-alive stance mirrors `liveness.rs`'s interim\r\n/// parity: this predicate exists to catch a *corpse*, never to guess at a\r\n/// backend it cannot see.\r\n///\r\n/// Recycled pids read alive, so a heal can be **missed** but never **mis-fired**\r\n/// — the failure direction is \"stays stale\", never \"tears down a live instance\".\r\n// [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\r\npub fn shell_pid_provably_dead(shell_perch: &Path) -> bool {\r\n    let Some(pid) = read_shell_pid(shell_perch).filter(|&pid| pid != 0) else {\r\n        return false;\r\n    };\r\n    // BAROMETER W2 (releases#78): probe pid AND birth stamp, not the pid alone.\r\n    // A pid namespace resets at boot, so a corpse pid can be REUSED by an\r\n    // unrelated process started after the restart; under bare-pid liveness that\r\n    // reads ALIVE, the derived status reads online, and `relink` then refuses\r\n    // SHELL_ALREADY_ONLINE -- the one command that recovers the instance\r\n    // refusing on the very state that is wrong. The PAIR test (same pid,\r\n    // different start time) is what tells a reused pid from our own process.\r\n    //\r\n    // The stance is unchanged where there is no stamp to compare: an instance\r\n    // launched before this field existed, or a backend that exposed no start\r\n    // time, still falls back to plain absence and fails toward ALIVE.\r\n    // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\n    let birth = read_shell_launch(shell_perch).and_then(|l| l.pid_started_at);\r\n    matches!(\r\n        crate::liveness::relay_liveness(Some(pid), birth),\r\n        crate::liveness::RelayLiveness::Gone\r\n    )\r\n}\r\n\r\n/// A shell instance's **derived** status — the shell-side twin of\r\n/// `liveness::is_perch_alive` (REQ-HAZARD-DAEMON-HOSTED-LIVENESS gave agent\r\n/// perches exactly this; shells never got it).\r\n///\r\n/// The recorded `status` has exactly two writers — the bind handshake (→online)\r\n/// and `spt_daemon::shellhost::close_shell` (→offline) — and `close_shell` runs\r\n/// on the **link-break** path alone. A binary that dies\r\n/// abruptly (force-kill, crash, OOM) breaks no link, so nothing ever writes the\r\n/// flip and the record says `online` forever. Deriving closes that hole without\r\n/// adding a supervisor: `online` + a provably-dead pid reads\r\n/// [`SHELL_STATUS_OFFLINE`]; every other case is the recorded value verbatim.\r\n///\r\n/// Callers that must NOT derive are documented at their sites (the two status\r\n/// writers, the owner-suspend cascade whose close IS the cleanup, and the wake\r\n/// reconciler's eligibility read) — see REQ-HAZARD-SHELL-STALE-ONLINE.\r\n// [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\r\npub fn effective_status<'a>(shell_perch: &Path, info: &'a ShellInfo) -> &'a str {\r\n    if info.status == SHELL_STATUS_ONLINE && shell_pid_provably_dead(shell_perch) {\r\n        return SHELL_STATUS_OFFLINE;\r\n    }\r\n    &info.status\r\n}\r\n\r\n/// [`effective_status`] as the predicate every gate actually asks: is this\r\n/// instance online *in fact*, not merely on paper?\r\n// [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\r\npub fn is_shell_online(shell_perch: &Path, info: &ShellInfo) -> bool {\r\n    effective_status(shell_perch, info) == SHELL_STATUS_ONLINE\r\n}\r\n\r\n/// The **parent** segment of an adapter option: the part before the first `:`\r\n/// of a `<adapter>:<profile>` composite (CONTEXT.md §adapter profile). The\r\n/// canonical shell id is minted off this — never the composite — so the id (a\r\n/// perch dir name + endpoint address) stays colon-free and path-safe; the full\r\n/// composite rides only in the `adapter_name` carrier field. Mirrors\r\n/// `spt_runtime::profile::split_option`, inlined here to keep spt-store free of\r\n/// an spt-runtime dependency (the crate graph runs the other way).\r\nfn adapter_parent(option: &str) -> &str {\r\n    option.split_once(':').map_or(option, |(a, _)| a)\r\n}\r\n\r\n/// Parse the `n` out of a canonical `<parent>-<n>` dir name for `parent`.\r\nfn instance_ordinal(dir_name: &str, parent: &str) -> Option<u32> {\r\n    let suffix = dir_name.strip_prefix(parent)?.strip_prefix('-')?;\r\n    if suffix.is_empty() || !suffix.bytes().all(|b| b.is_ascii_digit()) {\r\n        return None;\r\n    }\r\n    suffix.parse().ok()\r\n}\r\n\r\n/// Mint the next canonical id for `adapter` under `owner`: the **smallest\r\n/// free `n`** over existing instance dirs. Online and offline both occupy a\r\n/// slot (every non-torn-down perch counts — the same all-existing rule the\r\n/// instance cap uses); `teardown` removes the dir, so its `n` is reused.\r\n///\r\n/// `adapter` may be a `<adapter>:<profile>` composite; the id is minted off the\r\n/// **parent** segment ([`adapter_parent`]) so two profiles of one adapter share\r\n/// the same `<parent>-<n>` ordinal space (instances are the adapter's resource,\r\n/// not the profile's) and the id never carries the `:`.\r\n// [impl->REQ-EP-5]\r\n// [impl->REQ-HAZARD-ID-CHARSET]\r\npub fn mint_shell_id(owlery: &Path, owner: &str, adapter: &str) -> String {\r\n    let parent = adapter_parent(adapter);\r\n    let shells = perch::resolve_shells_dir_in(owlery, owner);\r\n    let mut taken: Vec<u32> = std::fs::read_dir(&shells)\r\n        .map(|entries| {\r\n            entries\r\n                .flatten()\r\n                .filter(|e| e.file_type().map(|t| t.is_dir()).unwrap_or(false))\r\n                .filter_map(|e| instance_ordinal(&e.file_name().to_string_lossy(), parent))\r\n                .collect()\r\n        })\r\n        .unwrap_or_default();\r\n    taken.sort_unstable();\r\n    let mut n = 0u32;\r\n    for t in taken {\r\n        if t == n {\r\n            n += 1;\r\n        } else if t > n {\r\n            break;\r\n        }\r\n    }\r\n    format!(\"{parent}-{n}\")\r\n}\r\n\r\n/// Every shell instance of `owner`: `(canonical_id, record)`, sorted by id\r\n/// for a stable order. Tolerant — a dir without a valid [`ShellInfo`] is\r\n/// skipped (it still *occupies* its mint slot; it just doesn't list).\r\n// [impl->REQ-EP-5]\r\npub fn list_shells(owlery: &Path, owner: &str) -> Vec<(String, ShellInfo)> {\r\n    let shells = perch::resolve_shells_dir_in(owlery, owner);\r\n    let Ok(entries) = std::fs::read_dir(&shells) else {\r\n        return Vec::new();\r\n    };\r\n    let mut out: Vec<(String, ShellInfo)> = entries\r\n        .flatten()\r\n        .filter(|e| e.file_type().map(|t| t.is_dir()).unwrap_or(false))\r\n        .filter_map(|e| {\r\n            let id = e.file_name().to_string_lossy().into_owned();\r\n            read_shell_info(&e.path()).map(|info| (id, info))\r\n        })\r\n        .collect();\r\n    out.sort_by(|a, b| a.0.cmp(&b.0));\r\n    out\r\n}\r\n\r\n/// Resolve a shell ref — canonical id or alias — to `(canonical_id, record)`.\r\n/// **Canonical-first**: an alias can never shadow a canonical id (spawn/rename\r\n/// refuse minting such an alias, and resolution checks canonical before alias\r\n/// as defense in depth). Aliases are owner-unique, so at most one matches.\r\n// [impl->REQ-EP-5]\r\npub fn resolve_shell_ref(\r\n    owlery: &Path,\r\n    owner: &str,\r\n    shell_ref: &str,\r\n) -> Option<(String, ShellInfo)> {\r\n    let direct = perch::resolve_shell_perch_path_in(owlery, owner, shell_ref);\r\n    if let Some(info) = read_shell_info(&direct) {\r\n        return Some((shell_ref.to_string(), info));\r\n    }\r\n    list_shells(owlery, owner)\r\n        .into_iter()\r\n        .find(|(_, info)| info.alias.as_deref() == Some(shell_ref))\r\n}\r\n\r\n/// Whether `alias` is free for `owner`: held by no instance's alias and not\r\n/// shadowing any existing canonical id.\r\nfn alias_free(owlery: &Path, owner: &str, alias: &str) -> bool {\r\n    resolve_shell_ref(owlery, owner, alias).is_none()\r\n}\r\n\r\n/// Mint + record a new instance of `adapter` for `owner` (the storage half of\r\n/// `spt shell spawn` — adapter-registration and consent gates are the\r","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120,121,122,123,124,125,126,127,128,129,130,131,132,133,134,135,136,137,138,139,140,141,142,143,144,145,146,147,148,149,150,151,152,153,154,155,156,157,158,159,160,161,162,163,164,165,166,167,168,169,170,171,172,173,174,175,176,177,178,179,180,181,182,183,184,185,186,187,188,189,190,191,192,193,194,195,196,197,198,199,200,201,202,203,204,205,206,207,208,209,210,211,212,213,214,215,216,217,218,219,220,221,222,223,224,225,226,227,228,229,230,231,232,233,234,235,236,237,238,239,240,241,242,243,244,245,246,247,248,249,250,251,252,253,254,255,256,257,258,259,260,261,262,263,264,265,266,267,268,269,270,271,272,273,274,275,276,277,278,279,280,281,282,283,284,285,286,287,288,289,290,291,292,293,294,295,296,297,298,299,300,301,302,303,304,305,306,307,308,309,310,311,312,313,314,315,316,317,318,319,320,321,322,323,324,325,326,327,328,329,330,331,332,333,334,335,336,337,338,339,340,341,342,343,344,345,346,347,348,349,350,351,352,353]},"fileSize":33627,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\crates\\spt-store\\src\\shellinfo.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":726,"totalBytes":16529,"outputLines":353,"outputBytes":16529,"shownRange":{"start":1,"end":353},"nextOffset":354}}},"isError":false,"timestamp":1789001136963}}
{"type":"message","id":"ab9d17cd","parentId":"ac0f6d03","timestamp":"2026-09-10T00:45:36.968Z","message":{"role":"toolResult","toolCallId":"call_7CFTl5wDfbMpoAe7PxbOIkfS|fc_077f3569365e313d016aa1fdb0c07487d1985acac801d13058","toolName":"read","content":[{"type":"text","text":"///   supervision thread.\n/// - **not offline (onlined / torn down / deregistered) + live watcher** ⇒\n///   stop it (flag + kill).\n///\n/// Runs at daemon boot and every tick — lifecycle flips happen in CLI-process\n/// library code, so the *reconciler* holds the invariant, not the flipping\n/// caller; a flip is at most one tick stale.\n// [impl->REQ-SHELL-2]\n/// Heal every local shell record whose `online` is a lie — leg (a) of\n/// releases#78.\n///\n/// `effective_status` already computes this flip on EVERY read and never writes\n/// it down, so the record and the display disagree by design: the display side\n/// derives and looks clean, while the wake cascade reads the RECORDED field and\n/// skips its relaunch because that field still says online. A machine death\n/// breaks no link, so `close_shell` never runs and nothing else ever corrects it.\n///\n/// THE WRITE IS GUARDED ON AN ACTUAL CHANGE. This runs every reconcile cycle\n/// (5s), and a heal that rewrote the record each time would be a continuous\n/// stream of identical writes over a rarely-changing condition — noise on disk,\n/// and a record whose mtime stops meaning \"something happened\".\n///\n/// It DERIVES deliberately, unlike the watcher-eligibility read below it: this\n/// is not an eligibility decision that could relaunch a binary, it is a\n/// correction of a field that is already wrong. Nothing is started here.\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n// [impl->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\nfn heal_stale_online_records(owlery: &Path, events: &mut dyn std::io::Write) {\n    for owner in spt_store::perch::list_self_perch_ids(owlery) {\n        for (shell_id, info) in shellinfo::list_shells(owlery, &owner) {\n            if info.status != SHELL_STATUS_ONLINE {\n                continue;\n            }\n            let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, &owner, &shell_id);\n            let truth = shellinfo::effective_status(&perch, &info);\n            if truth == info.status {\n                continue; // the record already tells the truth — write nothing\n            }\n            let mut healed = info.clone();\n            healed.status = truth.to_string();\n            if shellinfo::write_shell_info(&perch, &healed).is_ok() {\n                // [impl->REQ-SHELL-HEAL-ENUMERATION]\n                let _ = spt_proto::emit_line!(events, \"SHELL_RECORD_HEALED:{owner}/{shell_id}: online -> {truth}\");\n            }\n        }\n    }\n}\n\n/// Slack absorbed when asking whether a launch predates the boot instant.\n///\n/// The boot instant is DERIVED, not read from a ledger — on Windows it is\n/// now-minus-uptime, which drifts by however long the two calls take and by\n/// whatever the clock did since — so a process launched in the first moments of\n/// a boot can compute as very slightly older than the boot itself. Without slack\n/// that reads as \"a corpse from the previous boot\" and a nonpersistent watcher\n/// could be armed after a same-boot death. Ten seconds is far wider than plausible\n/// derivation error and far narrower than the gap this discriminant exists to\n/// detect (a previous boot is minutes to days back).\npub const BOOT_RESTORE_SLACK_MS: u64 = 10_000;\n\n/// Whether a nonpersistent shell's recorded launch predates the current boot.\n/// This preserves its same-boot force-kill freeze; persistent shells do not use\n/// launch age as eligibility (releases#287).\n// [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\npub fn launch_predates_boot(launched_ms: u64, boot_ms: u64, slack_ms: u64) -> bool {\n    launched_ms.saturating_add(slack_ms) < boot_ms\n}\n\n/// Watchers require an offline instance with no live launch awaiting bind.\n/// Persistent shells bypass the corpse-age freeze (releases#287); nonpersistent\n/// corpses still need a dated pre-boot launch. Missing dates fail closed only\n/// for that nonpersistent freeze, never for persistent restoration.\n// [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\nfn watcher_eligible(\n    perch: &Path,\n    info: &shellinfo::ShellInfo,\n    persistent: bool,\n    boot_ms: Option<u64>,\n) -> bool {\n    if info.status != SHELL_STATUS_OFFLINE {\n        return false;\n    }\n    if crate::shellhost::live_launch_winner(perch).is_some() {\n        return false; // binding: the shell binary and watcher must not overlap\n    }\n    if persistent || !shellinfo::shell_pid_provably_dead(perch) {\n        return true; // persistence overrides the freeze; clean closes need no age\n    }\n    let (Some(boot_ms), Some(launch)) = (boot_ms, shellinfo::read_shell_launch(perch)) else {\n        return false;\n    };\n    launch_predates_boot(launch.launched_ms, boot_ms, BOOT_RESTORE_SLACK_MS)\n}\n\n/// Which trigger asked for a restore.\n///\n/// The two triggers run the SAME per-owner body and differ only in what woke\n/// them, so each NAMES ITSELF in its event line rather than sharing one: the\n/// field diagnosis of releases#228 was a COUNT of these events in a rotated\n/// daemon log, and a shared name would have left that reading unable to say\n/// which trigger had fired -- the one thing it needed to know.\n#[derive(Debug, Clone, Copy, PartialEq, Eq)]\npub enum RestoreTrigger {\n    /// The once-per-daemon-generation boot sweep (releases#78 leg (b)).\n    Boot,\n    /// An owner endpoint that came online since the last reconcile pass\n    /// (releases#228).\n    OwnerOnline,\n}\n\nimpl RestoreTrigger {\n    fn restored_event(self) -> &'static str {\n        match self {\n            RestoreTrigger::Boot => \"SHELL_BOOT_RESTORED\",\n            RestoreTrigger::OwnerOnline => \"SHELL_OWNER_ONLINE_RESTORED\",\n        }\n    }\n\n    fn failed_event(self) -> &'static str {\n        match self {\n            RestoreTrigger::Boot => \"SHELL_BOOT_RESTORE_FAIL\",\n            RestoreTrigger::OwnerOnline => \"SHELL_OWNER_ONLINE_RESTORE_FAIL\",\n        }\n    }\n}\n\n/// ONE owner's slice of the persistent-shell restore decision -- the body BOTH\n/// triggers run.\n///\n/// Both triggers restore every persistent instance whose owner is online and\n/// whose binary is down in fact. Birth-safe liveness excludes a live launch\n/// awaiting bind as well as an already-online shell. Machine boot and launch\n/// age are irrelevant, including missing stamps (releases#287 operator ruling).\n///\n/// `set` is the daemon's live watcher set, and it is `None` at boot ON PURPOSE:\n/// the sweep runs before the reconcile loop's first tick, so no watcher can yet\n/// exist for the instance it is about to launch. The owner-online trigger runs\n/// MID-generation, where one can -- the reconcile start side arms watchers with\n/// no owner-online conjunct of its own -- so it passes the set and the watcher is\n/// stopped before the binary rises, which is the same online/offline mutual\n/// exclusivity `REQ-SHELL-2` holds everywhere else.\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n// [impl->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\n// [impl->REQ-SHELL-OWNER-ONLINE-RESTORE]\nfn restore_persistent_shells_of_owner(\n    owlery: &Path,\n    registered: &[(AdapterRecord, Manifest)],\n    adapters_dir: &Path,\n    owner: &str,\n    trigger: RestoreTrigger,\n    set: Option<&Arc<WakeSet>>,\n    restored: &mut Vec<String>,\n) {\n    let owner_perch = spt_store::perch::resolve_perch_path_in(owlery, owner, ParentHint::Infer);\n    if !spt_store::liveness::is_perch_alive(&owner_perch) {\n        return; // an offline owner is owed no persistent shell\n    }\n    for (shell_id, info) in shellinfo::list_shells(owlery, owner) {\n        let Some(shell) =\n            shell_section_of(registered, adapters_dir, &info.adapter_name).filter(|s| s.persistent)\n        else {\n            continue;\n        };\n        let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, &shell_id);\n        if shellinfo::is_shell_online(&perch, &info)\n            || crate::shellhost::live_launch_winner(&perch).is_some()\n        {\n            continue; // already online or binding in fact\n        }\n        // Mutual exclusivity, the online side: the watcher dies before the binary\n        // rises. `launch_shell` already kills the waker CHILD, but the supervising\n        // thread would read that as a crash-exit and respawn it under the shell we\n        // are starting; only the set can retire the thread. [impl->REQ-SHELL-2]\n        if let Some(set) = set {\n            set.stop_watcher(owlery, owner, &shell_id);\n        }\n        let install_dir = crate::shellhost::shell_install_dir(adapters_dir, &info.adapter_name);\n        match crate::shellhost::launch_shell(\n            owlery,\n            owner,\n            &shell_id,\n            &info.adapter_name,\n            install_dir.as_deref(),\n            &shell,\n        ) {\n            Ok(_) => {\n                eprintln!(\"{}:{owner}/{shell_id}\", trigger.restored_event());\n                restored.push(format!(\"{owner}/{shell_id}\"));\n            }\n            Err(e) => eprintln!(\"{}:{owner}/{shell_id}: {e}\", trigger.failed_event()),\n        }\n    }\n}\n\n/// The once-per-daemon-generation boot sweep: bring back the `persistent` shells\n/// a node restart stranded (leg (b) of releases#78).\n///\n/// Runs ONCE per daemon generation. It is not the ONLY restore trigger -- an owner\n/// endpoint that comes online after this sweep is covered by\n/// [`restore_persistent_shells_on_owner_online`] (releases#228), running the same\n/// [`restore_persistent_shells_of_owner`] body -- but it is the one that needs no\n/// trigger at all, since a node restart is what strands the shells to begin with.\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n// [impl->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\npub fn restore_persistent_shells_at_boot(\n    owlery: &Path,\n    registered: &[(AdapterRecord, Manifest)],\n    adapters_dir: &Path,\n) -> Vec<String> {\n    let mut restored = Vec::new();\n    for owner in spt_store::perch::list_self_perch_ids(owlery) {\n        restore_persistent_shells_of_owner(\n            owlery,\n            registered,\n            adapters_dir,\n            &owner,\n            RestoreTrigger::Boot,\n            None,\n            &mut restored,\n        );\n    }\n    restored\n}\n\n/// Per-owner online state carried ACROSS reconcile passes, so the restore below\n/// can be EDGE-triggered on an owner's offline->online transition.\n///\n/// An edge bounds a failing launch to one attempt per owner-online event,\n/// rather than retrying every reconcile tick while the owner remains online.\n/// A successful launch is protected by birth-safe liveness even before bind.\n#[derive(Debug, Default)]\npub struct OwnerOnlineEdge {\n    /// Owner id -> online as of the last pass. An owner that disappears from the\n    /// perch listing drops out, so a later reappearance reads as a transition.\n    seen: HashMap<String, bool>,\n}\n\nimpl OwnerOnlineEdge {\n    pub fn new() -> Self {\n        Self::default()\n    }\n\n    /// Observe every owner and return those that went offline->online since the\n    /// last observation. An owner seen for the FIRST time counts as a transition\n    /// when it is online: it was not online at the last pass, because it was not\n    /// there at all.\n    fn transitions(&mut self, owlery: &Path) -> Vec<String> {\n        let mut fired = Vec::new();\n        let mut now = HashMap::new();\n        for owner in spt_store::perch::list_self_perch_ids(owlery) {\n            let perch = spt_store::perch::resolve_perch_path_in(owlery, &owner, ParentHint::Infer);\n            let online = spt_store::liveness::is_perch_alive(&perch);\n            if online && !self.seen.get(&owner).copied().unwrap_or(false) {\n                fired.push(owner.clone());\n            }\n            now.insert(owner, online);\n        }\n        self.seen = now;\n        fired\n    }\n\n    /// Observe without firing -- the initial state the first tick compares against.\n    ///\n    /// Called BEFORE the boot sweep, and the ORDER is the safe direction rather\n    /// than an accident. Seeded AFTER the sweep, an owner that comes online in the\n    /// window between the two reads as already-online and never fires an edge --\n    /// precisely the miss releases#228 is about. Seeded before, the worst case is a\n    /// second attempt at an instance the sweep already restored, which the\n    /// down-in-fact arm refuses.\n    pub fn seed(&mut self, owlery: &Path) {\n        let _ = self.transitions(owlery);\n    }\n}\n\n/// Restore the `persistent` shells of every owner that came online since the last\n/// pass -- the steady-state half of the contract's promise (releases#228).\n///\n/// Daemons boot before endpoints do, so the boot sweep correctly skips an owner\n/// that is not up yet and, being once-per-generation, never revisits it; bringup\n/// from offline emits no rest edge either, so the ADR-0048 cascade does not cover\n/// it. Both triggers use the same down-in-fact decision, without a machine-boot\n/// or launch-stamp eligibility gate (releases#287).\n// [impl->REQ-SHELL-OWNER-ONLINE-RESTORE]\n// [impl->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\npub fn restore_persistent_shells_on_owner_online(\n    owlery: &Path,\n    registered: &[(AdapterRecord, Manifest)],\n    adapters_dir: &Path,\n    edge: &mut OwnerOnlineEdge,\n    set: &Arc<WakeSet>,\n) -> Vec<String> {\n    let owners = edge.transitions(owlery);\n    if owners.is_empty() {\n        return Vec::new();\n    }\n    let mut restored = Vec::new();\n    for owner in owners {\n        restore_persistent_shells_of_owner(\n            owlery,\n            registered,\n            adapters_dir,\n            &owner,\n            RestoreTrigger::OwnerOnline,\n            Some(set),\n            &mut restored,\n        );\n    }\n    restored\n}\n\npub fn reconcile_once(\n    owlery: &Path,\n    registered: &[(AdapterRecord, Manifest)],\n    adapters_dir: &Path,\n    set: &Arc<WakeSet>,\n    params: &WakeParams,\n) {\n    // Leg (a) first, every cycle: the record must stop lying before anything reads\n    // it below. Ordering matters -- the watcher-eligibility read further down asks\n    // the RECORDED field, so a heal that ran after it would leave this cycle acting\n    // on the stale value it just corrected.\n    heal_stale_online_records(owlery, &mut std::io::stderr().lock());\n\n    // Stop side first: watchers whose instance is gone or no longer eligible.\n    let live: Vec<(String, String)> = {\n        let mut map = set.inner.lock().unwrap_or_else(|p| p.into_inner());\n        map.retain(|_, h| !h.thread.is_finished());\n        map.keys().cloned().collect()\n    };\n    for (owner, shell_id) in live {\n        let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, &owner, &shell_id);\n        let still_eligible = shellinfo::read_shell_info(&perch)\n            .map(|i| i.status == SHELL_STATUS_OFFLINE)\n            .unwrap_or(false);\n        if !still_eligible {\n            set.stop_watcher(owlery, &owner, &shell_id);\n        }\n    }\n\n    // Start side: every ELIGIBLE instance with a wake_command and headroom.\n    //\n    // Persistent instances bypass the force-kill freeze; only nonpersistent\n    // corpses need the pre-boot discriminant (releases#287).\n    let boot_ms = spt_store::proc::boot_instant_ms();\n    for owner in spt_store::perch::list_self_perch_ids(owlery) {\n        for (shell_id, info) in shellinfo::list_shells(owlery, &owner) {\n            if set.contains(&owner, &shell_id) {\n                continue;\n            }\n            let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, &owner, &shell_id);\n            let Some(shell) = shell_section_of(registered, adapters_dir, &info.adapter_name) else {\n                continue; // deregistered adapter: nothing to run\n            };\n            if !watcher_eligible(&perch, &info, shell.persistent, boot_ms) {\n                continue;\n            }\n            let Some(wake_command) = shell.wake_command.as_deref() else {\n                continue;\n            };\n            if perch.join(WAKER_GAVE_UP_FILE).exists() {\n                continue; // crash-latched until the next shell activity\n            }\n            // Adopt cleanly: a dead daemon's watcher must not double-run.\n            kill_waker_at(&perch);\n            // [impl->REQ-INSTALL-11] threading only: this loop already holds the\n            // registered set the record's `source_dir` lives in.\n            let install_dir = crate::shellhost::shell_install_dir(adapters_dir, &info.adapter_name);\n            let tokens = match fill_wake_command(\n                &shell_id,\n                &info.adapter_name,\n                install_dir.as_deref(),\n                wake_command,\n            ) {\n                Ok(t) => t,\n                Err(e) => {\n                    spt_proto::emit_line_err!(\"WAKER_TEMPLATE:{owner}/{shell_id}: {e}\");\n                    continue;\n                }\n            };\n            let stop = Arc::new(AtomicBool::new(false));\n            let handle = {\n                let stop = Arc::clone(&stop);\n                let owlery = owlery.to_path_buf();\n                let owner = owner.clone();\n                let shell_id = shell_id.clone();\n                let adapter_name = info.adapter_name.clone();\n                let params = *params;\n                std::thread::spawn(move || {\n                    watcher_run(&owlery, &owner, &shell_id, &tokens, &params, &stop, || {\n                        resolve_wake(&owlery, &owner, &shell_id, &adapter_name, &shell)\n                    });\n                })\n            };\n            set.insert(\n                &owner,\n                &shell_id,\n                WatcherHandle {\n                    stop,\n                    thread: handle,\n                },\n            );\n        }\n    }\n}\n\n/// The reconcile cadence (ms): lifecycle flips land in CLI processes, so the\n/// daemon's loop is the invariant holder — a flip is at most one tick stale.\npub const RECONCILE_INTERVAL_MS: u64 = 5_000;\n\n/// Spawn the daemon's wake host: one thread sweeping [`reconcile_once`] at\n/// boot and every [`RECONCILE_INTERVAL_MS`] until `stop`. The registered set\n/// is re-read each sweep (adapter add/remove lands between ticks).\n// [impl->REQ-SHELL-2]\npub fn spawn_wake_host(stop: Arc<AtomicBool>) -> JoinHandle<()> {\n    std::thread::spawn(move || {\n        let set = Arc::new(WakeSet::new());\n        let params = WakeParams::default();\n        let mut owner_edge = OwnerOnlineEdge::new();\n        // Once per daemon generation, restore down persistent shells of owners\n        // already online. Owners arriving later are covered by the edge below.\n        // Neither trigger gates persistence on machine boot or launch age.\n        {\n            let owlery = spt_store::perch::owlery_dir();\n            let adapters_dir = spt_store::perch::adapters_dir();\n            let registered = spt_runtime::registry::registered(&adapters_dir);\n            // Seeded BEFORE the sweep: an owner that comes online between the two\n            // must read as a transition on the first tick, not as already-online.\n            owner_edge.seed(&owlery);\n            restore_persistent_shells_at_boot(&owlery, &registered, &adapters_dir);\n        }\n        while !stop.load(Ordering::SeqCst) {\n            let owlery = spt_store::perch::owlery_dir();\n            let adapters_dir = spt_store::perch::adapters_dir();\n            let registered = spt_runtime::registry::registered(&adapters_dir);\n            // Restore BEFORE reconciling, the same order the boot pass runs in: a\n            // restored instance is not watcher material, and reconcile's start side\n            // would otherwise arm a watcher for the shell we are about to launch.\n            restore_persistent_shells_on_owner_online(\n                &owlery,\n                &registered,\n                &adapters_dir,\n                &mut owner_edge,\n                &set,\n            );\n            reconcile_once(&owlery, &registered, &adapters_dir, &set, &params);\n            // Sleep in slices so a stop lands promptly.\n            let mut left = RECONCILE_INTERVAL_MS;\n            while left > 0 && !stop.load(Ordering::SeqCst) {\n                let step = left.min(100);\n                std::thread::sleep(Duration::from_millis(step));\n                left -= step;\n            }\n        }\n    })\n}\n\n#[cfg(test)]\nmod tests {\n    use super::*;\n    use spt_store::shellinfo::{spawn_record, SHELL_STATUS_ONLINE};\n\n    /// BACKSTOP for the converge-waits below — a ceiling on an observable, not\n    /// a budget any run is expected to consume (the healthy path exits in\n    /// milliseconds, and the give-up latch exits the product-fault path at\n    /// once).\n    ///\n    /// ANCHORED, not chosen: it is `slow-timeout.period` from\n    /// `.config/nextest.toml`, i.e. exactly the point at which the HARNESS\n    /// ITSELF starts calling this test slow. Past that, nextest is already\n    /// complaining and a longer wait cannot be the right answer — so the number\n    /// has a SOURCE rather than being \"some multiple of the 2s that broke\",\n    /// which would be a threshold keyed on a measurement instead of on the\n    /// thing. If that config value changes, this should follow it.\n    const WAIT_BACKSTOP: Duration = Duration::from_secs(60);\n…\n}\n\n[Showing lines 589-1063 of 2161. Use :1064 to continue]"}],"details":{"truncation":{"content":"///   supervision thread.\r\n/// - **not offline (onlined / torn down / deregistered) + live watcher** ⇒\r\n///   stop it (flag + kill).\r\n///\r\n/// Runs at daemon boot and every tick — lifecycle flips happen in CLI-process\r\n/// library code, so the *reconciler* holds the invariant, not the flipping\r\n/// caller; a flip is at most one tick stale.\r\n// [impl->REQ-SHELL-2]\r\n/// Heal every local shell record whose `online` is a lie — leg (a) of\r\n/// releases#78.\r\n///\r\n/// `effective_status` already computes this flip on EVERY read and never writes\r\n/// it down, so the record and the display disagree by design: the display side\r\n/// derives and looks clean, while the wake cascade reads the RECORDED field and\r\n/// skips its relaunch because that field still says online. A machine death\r\n/// breaks no link, so `close_shell` never runs and nothing else ever corrects it.\r\n///\r\n/// THE WRITE IS GUARDED ON AN ACTUAL CHANGE. This runs every reconcile cycle\r\n/// (5s), and a heal that rewrote the record each time would be a continuous\r\n/// stream of identical writes over a rarely-changing condition — noise on disk,\r\n/// and a record whose mtime stops meaning \"something happened\".\r\n///\r\n/// It DERIVES deliberately, unlike the watcher-eligibility read below it: this\r\n/// is not an eligibility decision that could relaunch a binary, it is a\r\n/// correction of a field that is already wrong. Nothing is started here.\r\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\n// [impl->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\r\nfn heal_stale_online_records(owlery: &Path, events: &mut dyn std::io::Write) {\r\n    for owner in spt_store::perch::list_self_perch_ids(owlery) {\r\n        for (shell_id, info) in shellinfo::list_shells(owlery, &owner) {\r\n            if info.status != SHELL_STATUS_ONLINE {\r\n                continue;\r\n            }\r\n            let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, &owner, &shell_id);\r\n            let truth = shellinfo::effective_status(&perch, &info);\r\n            if truth == info.status {\r\n                continue; // the record already tells the truth — write nothing\r\n            }\r\n            let mut healed = info.clone();\r\n            healed.status = truth.to_string();\r\n            if shellinfo::write_shell_info(&perch, &healed).is_ok() {\r\n                // [impl->REQ-SHELL-HEAL-ENUMERATION]\r\n                let _ = spt_proto::emit_line!(events, \"SHELL_RECORD_HEALED:{owner}/{shell_id}: online -> {truth}\");\r\n            }\r\n        }\r\n    }\r\n}\r\n\r\n/// Slack absorbed when asking whether a launch predates the boot instant.\r\n///\r\n/// The boot instant is DERIVED, not read from a ledger — on Windows it is\r\n/// now-minus-uptime, which drifts by however long the two calls take and by\r\n/// whatever the clock did since — so a process launched in the first moments of\r\n/// a boot can compute as very slightly older than the boot itself. Without slack\r\n/// that reads as \"a corpse from the previous boot\" and a nonpersistent watcher\r\n/// could be armed after a same-boot death. Ten seconds is far wider than plausible\r\n/// derivation error and far narrower than the gap this discriminant exists to\r\n/// detect (a previous boot is minutes to days back).\r\npub const BOOT_RESTORE_SLACK_MS: u64 = 10_000;\r\n\r\n/// Whether a nonpersistent shell's recorded launch predates the current boot.\r\n/// This preserves its same-boot force-kill freeze; persistent shells do not use\r\n/// launch age as eligibility (releases#287).\r\n// [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\r\npub fn launch_predates_boot(launched_ms: u64, boot_ms: u64, slack_ms: u64) -> bool {\r\n    launched_ms.saturating_add(slack_ms) < boot_ms\r\n}\r\n\r\n/// Watchers require an offline instance with no live launch awaiting bind.\r\n/// Persistent shells bypass the corpse-age freeze (releases#287); nonpersistent\r\n/// corpses still need a dated pre-boot launch. Missing dates fail closed only\r\n/// for that nonpersistent freeze, never for persistent restoration.\r\n// [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\r\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\nfn watcher_eligible(\r\n    perch: &Path,\r\n    info: &shellinfo::ShellInfo,\r\n    persistent: bool,\r\n    boot_ms: Option<u64>,\r\n) -> bool {\r\n    if info.status != SHELL_STATUS_OFFLINE {\r\n        return false;\r\n    }\r\n    if crate::shellhost::live_launch_winner(perch).is_some() {\r\n        return false; // binding: the shell binary and watcher must not overlap\r\n    }\r\n    if persistent || !shellinfo::shell_pid_provably_dead(perch) {\r\n        return true; // persistence overrides the freeze; clean closes need no age\r\n    }\r\n    let (Some(boot_ms), Some(launch)) = (boot_ms, shellinfo::read_shell_launch(perch)) else {\r\n        return false;\r\n    };\r\n    launch_predates_boot(launch.launched_ms, boot_ms, BOOT_RESTORE_SLACK_MS)\r\n}\r\n\r\n/// Which trigger asked for a restore.\r\n///\r\n/// The two triggers run the SAME per-owner body and differ only in what woke\r\n/// them, so each NAMES ITSELF in its event line rather than sharing one: the\r\n/// field diagnosis of releases#228 was a COUNT of these events in a rotated\r\n/// daemon log, and a shared name would have left that reading unable to say\r\n/// which trigger had fired -- the one thing it needed to know.\r\n#[derive(Debug, Clone, Copy, PartialEq, Eq)]\r\npub enum RestoreTrigger {\r\n    /// The once-per-daemon-generation boot sweep (releases#78 leg (b)).\r\n    Boot,\r\n    /// An owner endpoint that came online since the last reconcile pass\r\n    /// (releases#228).\r\n    OwnerOnline,\r\n}\r\n\r\nimpl RestoreTrigger {\r\n    fn restored_event(self) -> &'static str {\r\n        match self {\r\n            RestoreTrigger::Boot => \"SHELL_BOOT_RESTORED\",\r\n            RestoreTrigger::OwnerOnline => \"SHELL_OWNER_ONLINE_RESTORED\",\r\n        }\r\n    }\r\n\r\n    fn failed_event(self) -> &'static str {\r\n        match self {\r\n            RestoreTrigger::Boot => \"SHELL_BOOT_RESTORE_FAIL\",\r\n            RestoreTrigger::OwnerOnline => \"SHELL_OWNER_ONLINE_RESTORE_FAIL\",\r\n        }\r\n    }\r\n}\r\n\r\n/// ONE owner's slice of the persistent-shell restore decision -- the body BOTH\r\n/// triggers run.\r\n///\r\n/// Both triggers restore every persistent instance whose owner is online and\r\n/// whose binary is down in fact. Birth-safe liveness excludes a live launch\r\n/// awaiting bind as well as an already-online shell. Machine boot and launch\r\n/// age are irrelevant, including missing stamps (releases#287 operator ruling).\r\n///\r\n/// `set` is the daemon's live watcher set, and it is `None` at boot ON PURPOSE:\r\n/// the sweep runs before the reconcile loop's first tick, so no watcher can yet\r\n/// exist for the instance it is about to launch. The owner-online trigger runs\r\n/// MID-generation, where one can -- the reconcile start side arms watchers with\r\n/// no owner-online conjunct of its own -- so it passes the set and the watcher is\r\n/// stopped before the binary rises, which is the same online/offline mutual\r\n/// exclusivity `REQ-SHELL-2` holds everywhere else.\r\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\n// [impl->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\r\n// [impl->REQ-SHELL-OWNER-ONLINE-RESTORE]\r\nfn restore_persistent_shells_of_owner(\r\n    owlery: &Path,\r\n    registered: &[(AdapterRecord, Manifest)],\r\n    adapters_dir: &Path,\r\n    owner: &str,\r\n    trigger: RestoreTrigger,\r\n    set: Option<&Arc<WakeSet>>,\r\n    restored: &mut Vec<String>,\r\n) {\r\n    let owner_perch = spt_store::perch::resolve_perch_path_in(owlery, owner, ParentHint::Infer);\r\n    if !spt_store::liveness::is_perch_alive(&owner_perch) {\r\n        return; // an offline owner is owed no persistent shell\r\n    }\r\n    for (shell_id, info) in shellinfo::list_shells(owlery, owner) {\r\n        let Some(shell) =\r\n            shell_section_of(registered, adapters_dir, &info.adapter_name).filter(|s| s.persistent)\r\n        else {\r\n            continue;\r\n        };\r\n        let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, &shell_id);\r\n        if shellinfo::is_shell_online(&perch, &info)\r\n            || crate::shellhost::live_launch_winner(&perch).is_some()\r\n        {\r\n            continue; // already online or binding in fact\r\n        }\r\n        // Mutual exclusivity, the online side: the watcher dies before the binary\r\n        // rises. `launch_shell` already kills the waker CHILD, but the supervising\r\n        // thread would read that as a crash-exit and respawn it under the shell we\r\n        // are starting; only the set can retire the thread. [impl->REQ-SHELL-2]\r\n        if let Some(set) = set {\r\n            set.stop_watcher(owlery, owner, &shell_id);\r\n        }\r\n        let install_dir = crate::shellhost::shell_install_dir(adapters_dir, &info.adapter_name);\r\n        match crate::shellhost::launch_shell(\r\n            owlery,\r\n            owner,\r\n            &shell_id,\r\n            &info.adapter_name,\r\n            install_dir.as_deref(),\r\n            &shell,\r\n        ) {\r\n            Ok(_) => {\r\n                eprintln!(\"{}:{owner}/{shell_id}\", trigger.restored_event());\r\n                restored.push(format!(\"{owner}/{shell_id}\"));\r\n            }\r\n            Err(e) => eprintln!(\"{}:{owner}/{shell_id}: {e}\", trigger.failed_event()),\r\n        }\r\n    }\r\n}\r\n\r\n/// The once-per-daemon-generation boot sweep: bring back the `persistent` shells\r\n/// a node restart stranded (leg (b) of releases#78).\r\n///\r\n/// Runs ONCE per daemon generation. It is not the ONLY restore trigger -- an owner\r\n/// endpoint that comes online after this sweep is covered by\r\n/// [`restore_persistent_shells_on_owner_online`] (releases#228), running the same\r\n/// [`restore_persistent_shells_of_owner`] body -- but it is the one that needs no\r\n/// trigger at all, since a node restart is what strands the shells to begin with.\r\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\n// [impl->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\r\npub fn restore_persistent_shells_at_boot(\r\n    owlery: &Path,\r\n    registered: &[(AdapterRecord, Manifest)],\r\n    adapters_dir: &Path,\r\n) -> Vec<String> {\r\n    let mut restored = Vec::new();\r\n    for owner in spt_store::perch::list_self_perch_ids(owlery) {\r\n        restore_persistent_shells_of_owner(\r\n            owlery,\r\n            registered,\r\n            adapters_dir,\r\n            &owner,\r\n            RestoreTrigger::Boot,\r\n            None,\r\n            &mut restored,\r\n        );\r\n    }\r\n    restored\r\n}\r\n\r\n/// Per-owner online state carried ACROSS reconcile passes, so the restore below\r\n/// can be EDGE-triggered on an owner's offline->online transition.\r\n///\r\n/// An edge bounds a failing launch to one attempt per owner-online event,\r\n/// rather than retrying every reconcile tick while the owner remains online.\r\n/// A successful launch is protected by birth-safe liveness even before bind.\r\n#[derive(Debug, Default)]\r\npub struct OwnerOnlineEdge {\r\n    /// Owner id -> online as of the last pass. An owner that disappears from the\r\n    /// perch listing drops out, so a later reappearance reads as a transition.\r\n    seen: HashMap<String, bool>,\r\n}\r\n\r\nimpl OwnerOnlineEdge {\r\n    pub fn new() -> Self {\r\n        Self::default()\r\n    }\r\n\r\n    /// Observe every owner and return those that went offline->online since the\r\n    /// last observation. An owner seen for the FIRST time counts as a transition\r\n    /// when it is online: it was not online at the last pass, because it was not\r\n    /// there at all.\r\n    fn transitions(&mut self, owlery: &Path) -> Vec<String> {\r\n        let mut fired = Vec::new();\r\n        let mut now = HashMap::new();\r\n        for owner in spt_store::perch::list_self_perch_ids(owlery) {\r\n            let perch = spt_store::perch::resolve_perch_path_in(owlery, &owner, ParentHint::Infer);\r\n            let online = spt_store::liveness::is_perch_alive(&perch);\r\n            if online && !self.seen.get(&owner).copied().unwrap_or(false) {\r\n                fired.push(owner.clone());\r\n            }\r\n            now.insert(owner, online);\r\n        }\r\n        self.seen = now;\r\n        fired\r\n    }\r\n\r\n    /// Observe without firing -- the initial state the first tick compares against.\r\n    ///\r\n    /// Called BEFORE the boot sweep, and the ORDER is the safe direction rather\r\n    /// than an accident. Seeded AFTER the sweep, an owner that comes online in the\r\n    /// window between the two reads as already-online and never fires an edge --\r\n    /// precisely the miss releases#228 is about. Seeded before, the worst case is a\r\n    /// second attempt at an instance the sweep already restored, which the\r\n    /// down-in-fact arm refuses.\r\n    pub fn seed(&mut self, owlery: &Path) {\r\n        let _ = self.transitions(owlery);\r\n    }\r\n}\r\n\r\n/// Restore the `persistent` shells of every owner that came online since the last\r\n/// pass -- the steady-state half of the contract's promise (releases#228).\r\n///\r\n/// Daemons boot before endpoints do, so the boot sweep correctly skips an owner\r\n/// that is not up yet and, being once-per-generation, never revisits it; bringup\r\n/// from offline emits no rest edge either, so the ADR-0048 cascade does not cover\r\n/// it. Both triggers use the same down-in-fact decision, without a machine-boot\r\n/// or launch-stamp eligibility gate (releases#287).\r\n// [impl->REQ-SHELL-OWNER-ONLINE-RESTORE]\r\n// [impl->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\r\npub fn restore_persistent_shells_on_owner_online(\r\n    owlery: &Path,\r\n    registered: &[(AdapterRecord, Manifest)],\r\n    adapters_dir: &Path,\r\n    edge: &mut OwnerOnlineEdge,\r\n    set: &Arc<WakeSet>,\r\n) -> Vec<String> {\r\n    let owners = edge.transitions(owlery);\r\n    if owners.is_empty() {\r\n        return Vec::new();\r\n    }\r\n    let mut restored = Vec::new();\r\n    for owner in owners {\r\n        restore_persistent_shells_of_owner(\r\n            owlery,\r\n            registered,\r\n            adapters_dir,\r\n            &owner,\r\n            RestoreTrigger::OwnerOnline,\r\n            Some(set),\r\n            &mut restored,\r\n        );\r\n    }\r\n    restored\r\n}\r\n\r\npub fn reconcile_once(\r\n    owlery: &Path,\r\n    registered: &[(AdapterRecord, Manifest)],\r\n    adapters_dir: &Path,\r\n    set: &Arc<WakeSet>,\r\n    params: &WakeParams,\r\n) {\r\n    // Leg (a) first, every cycle: the record must stop lying before anything reads\r\n    // it below. Ordering matters -- the watcher-eligibility read further down asks\r\n    // the RECORDED field, so a heal that ran after it would leave this cycle acting\r\n    // on the stale value it just corrected.\r\n    heal_stale_online_records(owlery, &mut std::io::stderr().lock());\r\n\r\n    // Stop side first: watchers whose instance is gone or no longer eligible.\r\n    let live: Vec<(String, String)> = {\r\n        let mut map = set.inner.lock().unwrap_or_else(|p| p.into_inner());\r\n        map.retain(|_, h| !h.thread.is_finished());\r\n        map.keys().cloned().collect()\r\n    };\r\n    for (owner, shell_id) in live {\r\n        let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, &owner, &shell_id);\r\n        let still_eligible = shellinfo::read_shell_info(&perch)\r\n            .map(|i| i.status == SHELL_STATUS_OFFLINE)\r\n            .unwrap_or(false);\r\n        if !still_eligible {\r\n            set.stop_watcher(owlery, &owner, &shell_id);\r\n        }\r\n    }\r\n\r\n    // Start side: every ELIGIBLE instance with a wake_command and headroom.\r\n    //\r\n    // Persistent instances bypass the force-kill freeze; only nonpersistent\r\n    // corpses need the pre-boot discriminant (releases#287).\r\n    let boot_ms = spt_store::proc::boot_instant_ms();\r\n    for owner in spt_store::perch::list_self_perch_ids(owlery) {\r\n        for (shell_id, info) in shellinfo::list_shells(owlery, &owner) {\r\n            if set.contains(&owner, &shell_id) {\r\n                continue;\r\n            }\r\n            let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, &owner, &shell_id);\r\n            let Some(shell) = shell_section_of(registered, adapters_dir, &info.adapter_name) else {\r\n                continue; // deregistered adapter: nothing to run\r\n            };\r\n            if !watcher_eligible(&perch, &info, shell.persistent, boot_ms) {\r\n                continue;\r\n            }\r\n            let Some(wake_command) = shell.wake_command.as_deref() else {\r\n                continue;\r\n            };\r\n            if perch.join(WAKER_GAVE_UP_FILE).exists() {\r\n                continue; // crash-latched until the next shell activity\r\n            }\r\n            // Adopt cleanly: a dead daemon's watcher must not double-run.\r\n            kill_waker_at(&perch);\r\n            // [impl->REQ-INSTALL-11] threading only: this loop already holds the\r\n            // registered set the record's `source_dir` lives in.\r\n            let install_dir = crate::shellhost::shell_install_dir(adapters_dir, &info.adapter_name);\r\n            let tokens = match fill_wake_command(\r\n                &shell_id,\r\n                &info.adapter_name,\r\n                install_dir.as_deref(),\r\n                wake_command,\r\n            ) {\r\n                Ok(t) => t,\r\n                Err(e) => {\r\n                    spt_proto::emit_line_err!(\"WAKER_TEMPLATE:{owner}/{shell_id}: {e}\");\r\n                    continue;\r\n                }\r\n            };\r\n            let stop = Arc::new(AtomicBool::new(false));\r\n            let handle = {\r\n                let stop = Arc::clone(&stop);\r\n                let owlery = owlery.to_path_buf();\r\n                let owner = owner.clone();\r\n                let shell_id = shell_id.clone();\r\n                let adapter_name = info.adapter_name.clone();\r\n                let params = *params;\r\n                std::thread::spawn(move || {\r\n                    watcher_run(&owlery, &owner, &shell_id, &tokens, &params, &stop, || {\r\n                        resolve_wake(&owlery, &owner, &shell_id, &adapter_name, &shell)\r\n                    });\r\n                })\r\n            };\r\n            set.insert(\r\n                &owner,\r\n                &shell_id,\r\n                WatcherHandle {\r\n                    stop,\r\n                    thread: handle,\r\n                },\r\n            );\r\n        }\r\n    }\r\n}\r\n\r\n/// The reconcile cadence (ms): lifecycle flips land in CLI processes, so the\r\n/// daemon's loop is the invariant holder — a flip is at most one tick stale.\r\npub const RECONCILE_INTERVAL_MS: u64 = 5_000;\r\n\r\n/// Spawn the daemon's wake host: one thread sweeping [`reconcile_once`] at\r\n/// boot and every [`RECONCILE_INTERVAL_MS`] until `stop`. The registered set\r\n/// is re-read each sweep (adapter add/remove lands between ticks).\r\n// [impl->REQ-SHELL-2]\r\npub fn spawn_wake_host(stop: Arc<AtomicBool>) -> JoinHandle<()> {\r\n    std::thread::spawn(move || {\r\n        let set = Arc::new(WakeSet::new());\r\n        let params = WakeParams::default();\r\n        let mut owner_edge = OwnerOnlineEdge::new();\r\n        // Once per daemon generation, restore down persistent shells of owners\r\n        // already online. Owners arriving later are covered by the edge below.\r\n        // Neither trigger gates persistence on machine boot or launch age.\r\n        {\r\n            let owlery = spt_store::perch::owlery_dir();\r\n            let adapters_dir = spt_store::perch::adapters_dir();\r\n            let registered = spt_runtime::registry::registered(&adapters_dir);\r\n            // Seeded BEFORE the sweep: an owner that comes online between the two\r\n            // must read as a transition on the first tick, not as already-online.\r\n            owner_edge.seed(&owlery);\r\n            restore_persistent_shells_at_boot(&owlery, &registered, &adapters_dir);\r\n        }\r\n        while !stop.load(Ordering::SeqCst) {\r\n            let owlery = spt_store::perch::owlery_dir();\r\n            let adapters_dir = spt_store::perch::adapters_dir();\r\n            let registered = spt_runtime::registry::registered(&adapters_dir);\r\n            // Restore BEFORE reconciling, the same order the boot pass runs in: a\r\n            // restored instance is not watcher material, and reconcile's start side\r\n            // would otherwise arm a watcher for the shell we are about to launch.\r\n            restore_persistent_shells_on_owner_online(\r\n                &owlery,\r\n                &registered,\r\n                &adapters_dir,\r\n                &mut owner_edge,\r\n                &set,\r\n            );\r\n            reconcile_once(&owlery, &registered, &adapters_dir, &set, &params);\r\n            // Sleep in slices so a stop lands promptly.\r\n            let mut left = RECONCILE_INTERVAL_MS;\r\n            while left > 0 && !stop.load(Ordering::SeqCst) {\r\n                let step = left.min(100);\r\n                std::thread::sleep(Duration::from_millis(step));\r\n                left -= step;\r\n            }\r\n        }\r\n    })\r\n}\r\n\r\n#[cfg(test)]\r\nmod tests {\r\n    use super::*;\r\n    use spt_store::shellinfo::{spawn_record, SHELL_STATUS_ONLINE};\r\n\r\n    /// BACKSTOP for the converge-waits below — a ceiling on an observable, not\r\n    /// a budget any run is expected to consume (the healthy path exits in\r\n    /// milliseconds, and the give-up latch exits the product-fault path at\r\n    /// once).\r\n    ///\r\n    /// ANCHORED, not chosen: it is `slow-timeout.period` from\r\n    /// `.config/nextest.toml`, i.e. exactly the point at which the HARNESS\r\n    /// ITSELF starts calling this test slow. Past that, nextest is already\r\n    /// complaining and a longer wait cannot be the right answer — so the number\r\n    /// has a SOURCE rather than being \"some multiple of the 2s that broke\",\r\n    /// which would be a threshold keyed on a measurement instead of on the\r\n    /// thing. If that config value changes, this should follow it.\r\n    const WAIT_BACKSTOP: Duration = Duration::from_secs(60);\r","truncated":true,"truncatedBy":"lines","totalLines":1573,"totalBytes":21600,"outputLines":475,"outputBytes":21600,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":2161,"displayContent":{"text":"///   supervision thread.\r\n/// - **not offline (onlined / torn down / deregistered) + live watcher** ⇒\r\n///   stop it (flag + kill).\r\n///\r\n/// Runs at daemon boot and every tick — lifecycle flips happen in CLI-process\r\n/// library code, so the *reconciler* holds the invariant, not the flipping\r\n/// caller; a flip is at most one tick stale.\r\n// [impl->REQ-SHELL-2]\r\n/// Heal every local shell record whose `online` is a lie — leg (a) of\r\n/// releases#78.\r\n///\r\n/// `effective_status` already computes this flip on EVERY read and never writes\r\n/// it down, so the record and the display disagree by design: the display side\r\n/// derives and looks clean, while the wake cascade reads the RECORDED field and\r\n/// skips its relaunch because that field still says online. A machine death\r\n/// breaks no link, so `close_shell` never runs and nothing else ever corrects it.\r\n///\r\n/// THE WRITE IS GUARDED ON AN ACTUAL CHANGE. This runs every reconcile cycle\r\n/// (5s), and a heal that rewrote the record each time would be a continuous\r\n/// stream of identical writes over a rarely-changing condition — noise on disk,\r\n/// and a record whose mtime stops meaning \"something happened\".\r\n///\r\n/// It DERIVES deliberately, unlike the watcher-eligibility read below it: this\r\n/// is not an eligibility decision that could relaunch a binary, it is a\r\n/// correction of a field that is already wrong. Nothing is started here.\r\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\n// [impl->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\r\nfn heal_stale_online_records(owlery: &Path, events: &mut dyn std::io::Write) {\r\n    for owner in spt_store::perch::list_self_perch_ids(owlery) {\r\n        for (shell_id, info) in shellinfo::list_shells(owlery, &owner) {\r\n            if info.status != SHELL_STATUS_ONLINE {\r\n                continue;\r\n            }\r\n            let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, &owner, &shell_id);\r\n            let truth = shellinfo::effective_status(&perch, &info);\r\n            if truth == info.status {\r\n                continue; // the record already tells the truth — write nothing\r\n            }\r\n            let mut healed = info.clone();\r\n            healed.status = truth.to_string();\r\n            if shellinfo::write_shell_info(&perch, &healed).is_ok() {\r\n                // [impl->REQ-SHELL-HEAL-ENUMERATION]\r\n                let _ = spt_proto::emit_line!(events, \"SHELL_RECORD_HEALED:{owner}/{shell_id}: online -> {truth}\");\r\n            }\r\n        }\r\n    }\r\n}\r\n\r\n/// Slack absorbed when asking whether a launch predates the boot instant.\r\n///\r\n/// The boot instant is DERIVED, not read from a ledger — on Windows it is\r\n/// now-minus-uptime, which drifts by however long the two calls take and by\r\n/// whatever the clock did since — so a process launched in the first moments of\r\n/// a boot can compute as very slightly older than the boot itself. Without slack\r\n/// that reads as \"a corpse from the previous boot\" and a nonpersistent watcher\r\n/// could be armed after a same-boot death. Ten seconds is far wider than plausible\r\n/// derivation error and far narrower than the gap this discriminant exists to\r\n/// detect (a previous boot is minutes to days back).\r\npub const BOOT_RESTORE_SLACK_MS: u64 = 10_000;\r\n\r\n/// Whether a nonpersistent shell's recorded launch predates the current boot.\r\n/// This preserves its same-boot force-kill freeze; persistent shells do not use\r\n/// launch age as eligibility (releases#287).\r\n// [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\r\npub fn launch_predates_boot(launched_ms: u64, boot_ms: u64, slack_ms: u64) -> bool {\r\n    launched_ms.saturating_add(slack_ms) < boot_ms\r\n}\r\n\r\n/// Watchers require an offline instance with no live launch awaiting bind.\r\n/// Persistent shells bypass the corpse-age freeze (releases#287); nonpersistent\r\n/// corpses still need a dated pre-boot launch. Missing dates fail closed only\r\n/// for that nonpersistent freeze, never for persistent restoration.\r\n// [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\r\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\nfn watcher_eligible(\r\n    perch: &Path,\r\n    info: &shellinfo::ShellInfo,\r\n    persistent: bool,\r\n    boot_ms: Option<u64>,\r\n) -> bool {\r\n    if info.status != SHELL_STATUS_OFFLINE {\r\n        return false;\r\n    }\r\n    if crate::shellhost::live_launch_winner(perch).is_some() {\r\n        return false; // binding: the shell binary and watcher must not overlap\r\n    }\r\n    if persistent || !shellinfo::shell_pid_provably_dead(perch) {\r\n        return true; // persistence overrides the freeze; clean closes need no age\r\n    }\r\n    let (Some(boot_ms), Some(launch)) = (boot_ms, shellinfo::read_shell_launch(perch)) else {\r\n        return false;\r\n    };\r\n    launch_predates_boot(launch.launched_ms, boot_ms, BOOT_RESTORE_SLACK_MS)\r\n}\r\n\r\n/// Which trigger asked for a restore.\r\n///\r\n/// The two triggers run the SAME per-owner body and differ only in what woke\r\n/// them, so each NAMES ITSELF in its event line rather than sharing one: the\r\n/// field diagnosis of releases#228 was a COUNT of these events in a rotated\r\n/// daemon log, and a shared name would have left that reading unable to say\r\n/// which trigger had fired -- the one thing it needed to know.\r\n#[derive(Debug, Clone, Copy, PartialEq, Eq)]\r\npub enum RestoreTrigger {\r\n    /// The once-per-daemon-generation boot sweep (releases#78 leg (b)).\r\n    Boot,\r\n    /// An owner endpoint that came online since the last reconcile pass\r\n    /// (releases#228).\r\n    OwnerOnline,\r\n}\r\n\r\nimpl RestoreTrigger {\r\n    fn restored_event(self) -> &'static str {\r\n        match self {\r\n            RestoreTrigger::Boot => \"SHELL_BOOT_RESTORED\",\r\n            RestoreTrigger::OwnerOnline => \"SHELL_OWNER_ONLINE_RESTORED\",\r\n        }\r\n    }\r\n\r\n    fn failed_event(self) -> &'static str {\r\n        match self {\r\n            RestoreTrigger::Boot => \"SHELL_BOOT_RESTORE_FAIL\",\r\n            RestoreTrigger::OwnerOnline => \"SHELL_OWNER_ONLINE_RESTORE_FAIL\",\r\n        }\r\n    }\r\n}\r\n\r\n/// ONE owner's slice of the persistent-shell restore decision -- the body BOTH\r\n/// triggers run.\r\n///\r\n/// Both triggers restore every persistent instance whose owner is online and\r\n/// whose binary is down in fact. Birth-safe liveness excludes a live launch\r\n/// awaiting bind as well as an already-online shell. Machine boot and launch\r\n/// age are irrelevant, including missing stamps (releases#287 operator ruling).\r\n///\r\n/// `set` is the daemon's live watcher set, and it is `None` at boot ON PURPOSE:\r\n/// the sweep runs before the reconcile loop's first tick, so no watcher can yet\r\n/// exist for the instance it is about to launch. The owner-online trigger runs\r\n/// MID-generation, where one can -- the reconcile start side arms watchers with\r\n/// no owner-online conjunct of its own -- so it passes the set and the watcher is\r\n/// stopped before the binary rises, which is the same online/offline mutual\r\n/// exclusivity `REQ-SHELL-2` holds everywhere else.\r\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\n// [impl->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\r\n// [impl->REQ-SHELL-OWNER-ONLINE-RESTORE]\r\nfn restore_persistent_shells_of_owner(\r\n    owlery: &Path,\r\n    registered: &[(AdapterRecord, Manifest)],\r\n    adapters_dir: &Path,\r\n    owner: &str,\r\n    trigger: RestoreTrigger,\r\n    set: Option<&Arc<WakeSet>>,\r\n    restored: &mut Vec<String>,\r\n) {\r\n    let owner_perch = spt_store::perch::resolve_perch_path_in(owlery, owner, ParentHint::Infer);\r\n    if !spt_store::liveness::is_perch_alive(&owner_perch) {\r\n        return; // an offline owner is owed no persistent shell\r\n    }\r\n    for (shell_id, info) in shellinfo::list_shells(owlery, owner) {\r\n        let Some(shell) =\r\n            shell_section_of(registered, adapters_dir, &info.adapter_name).filter(|s| s.persistent)\r\n        else {\r\n            continue;\r\n        };\r\n        let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, &shell_id);\r\n        if shellinfo::is_shell_online(&perch, &info)\r\n            || crate::shellhost::live_launch_winner(&perch).is_some()\r\n        {\r\n            continue; // already online or binding in fact\r\n        }\r\n        // Mutual exclusivity, the online side: the watcher dies before the binary\r\n        // rises. `launch_shell` already kills the waker CHILD, but the supervising\r\n        // thread would read that as a crash-exit and respawn it under the shell we\r\n        // are starting; only the set can retire the thread. [impl->REQ-SHELL-2]\r\n        if let Some(set) = set {\r\n            set.stop_watcher(owlery, owner, &shell_id);\r\n        }\r\n        let install_dir = crate::shellhost::shell_install_dir(adapters_dir, &info.adapter_name);\r\n        match crate::shellhost::launch_shell(\r\n            owlery,\r\n            owner,\r\n            &shell_id,\r\n            &info.adapter_name,\r\n            install_dir.as_deref(),\r\n            &shell,\r\n        ) {\r\n            Ok(_) => {\r\n                eprintln!(\"{}:{owner}/{shell_id}\", trigger.restored_event());\r\n                restored.push(format!(\"{owner}/{shell_id}\"));\r\n            }\r\n            Err(e) => eprintln!(\"{}:{owner}/{shell_id}: {e}\", trigger.failed_event()),\r\n        }\r\n    }\r\n}\r\n\r\n/// The once-per-daemon-generation boot sweep: bring back the `persistent` shells\r\n/// a node restart stranded (leg (b) of releases#78).\r\n///\r\n/// Runs ONCE per daemon generation. It is not the ONLY restore trigger -- an owner\r\n/// endpoint that comes online after this sweep is covered by\r\n/// [`restore_persistent_shells_on_owner_online`] (releases#228), running the same\r\n/// [`restore_persistent_shells_of_owner`] body -- but it is the one that needs no\r\n/// trigger at all, since a node restart is what strands the shells to begin with.\r\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\n// [impl->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\r\npub fn restore_persistent_shells_at_boot(\r\n    owlery: &Path,\r\n    registered: &[(AdapterRecord, Manifest)],\r\n    adapters_dir: &Path,\r\n) -> Vec<String> {\r\n    let mut restored = Vec::new();\r\n    for owner in spt_store::perch::list_self_perch_ids(owlery) {\r\n        restore_persistent_shells_of_owner(\r\n            owlery,\r\n            registered,\r\n            adapters_dir,\r\n            &owner,\r\n            RestoreTrigger::Boot,\r\n            None,\r\n            &mut restored,\r\n        );\r\n    }\r\n    restored\r\n}\r\n\r\n/// Per-owner online state carried ACROSS reconcile passes, so the restore below\r\n/// can be EDGE-triggered on an owner's offline->online transition.\r\n///\r\n/// An edge bounds a failing launch to one attempt per owner-online event,\r\n/// rather than retrying every reconcile tick while the owner remains online.\r\n/// A successful launch is protected by birth-safe liveness even before bind.\r\n#[derive(Debug, Default)]\r\npub struct OwnerOnlineEdge {\r\n    /// Owner id -> online as of the last pass. An owner that disappears from the\r\n    /// perch listing drops out, so a later reappearance reads as a transition.\r\n    seen: HashMap<String, bool>,\r\n}\r\n\r\nimpl OwnerOnlineEdge {\r\n    pub fn new() -> Self {\r\n        Self::default()\r\n    }\r\n\r\n    /// Observe every owner and return those that went offline->online since the\r\n    /// last observation. An owner seen for the FIRST time counts as a transition\r\n    /// when it is online: it was not online at the last pass, because it was not\r\n    /// there at all.\r\n    fn transitions(&mut self, owlery: &Path) -> Vec<String> {\r\n        let mut fired = Vec::new();\r\n        let mut now = HashMap::new();\r\n        for owner in spt_store::perch::list_self_perch_ids(owlery) {\r\n            let perch = spt_store::perch::resolve_perch_path_in(owlery, &owner, ParentHint::Infer);\r\n            let online = spt_store::liveness::is_perch_alive(&perch);\r\n            if online && !self.seen.get(&owner).copied().unwrap_or(false) {\r\n                fired.push(owner.clone());\r\n            }\r\n            now.insert(owner, online);\r\n        }\r\n        self.seen = now;\r\n        fired\r\n    }\r\n\r\n    /// Observe without firing -- the initial state the first tick compares against.\r\n    ///\r\n    /// Called BEFORE the boot sweep, and the ORDER is the safe direction rather\r\n    /// than an accident. Seeded AFTER the sweep, an owner that comes online in the\r\n    /// window between the two reads as already-online and never fires an edge --\r\n    /// precisely the miss releases#228 is about. Seeded before, the worst case is a\r\n    /// second attempt at an instance the sweep already restored, which the\r\n    /// down-in-fact arm refuses.\r\n    pub fn seed(&mut self, owlery: &Path) {\r\n        let _ = self.transitions(owlery);\r\n    }\r\n}\r\n\r\n/// Restore the `persistent` shells of every owner that came online since the last\r\n/// pass -- the steady-state half of the contract's promise (releases#228).\r\n///\r\n/// Daemons boot before endpoints do, so the boot sweep correctly skips an owner\r\n/// that is not up yet and, being once-per-generation, never revisits it; bringup\r\n/// from offline emits no rest edge either, so the ADR-0048 cascade does not cover\r\n/// it. Both triggers use the same down-in-fact decision, without a machine-boot\r\n/// or launch-stamp eligibility gate (releases#287).\r\n// [impl->REQ-SHELL-OWNER-ONLINE-RESTORE]\r\n// [impl->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\r\npub fn restore_persistent_shells_on_owner_online(\r\n    owlery: &Path,\r\n    registered: &[(AdapterRecord, Manifest)],\r\n    adapters_dir: &Path,\r\n    edge: &mut OwnerOnlineEdge,\r\n    set: &Arc<WakeSet>,\r\n) -> Vec<String> {\r\n    let owners = edge.transitions(owlery);\r\n    if owners.is_empty() {\r\n        return Vec::new();\r\n    }\r\n    let mut restored = Vec::new();\r\n    for owner in owners {\r\n        restore_persistent_shells_of_owner(\r\n            owlery,\r\n            registered,\r\n            adapters_dir,\r\n            &owner,\r\n            RestoreTrigger::OwnerOnline,\r\n            Some(set),\r\n            &mut restored,\r\n        );\r\n    }\r\n    restored\r\n}\r\n\r\npub fn reconcile_once(\r\n    owlery: &Path,\r\n    registered: &[(AdapterRecord, Manifest)],\r\n    adapters_dir: &Path,\r\n    set: &Arc<WakeSet>,\r\n    params: &WakeParams,\r\n) {\r\n    // Leg (a) first, every cycle: the record must stop lying before anything reads\r\n    // it below. Ordering matters -- the watcher-eligibility read further down asks\r\n    // the RECORDED field, so a heal that ran after it would leave this cycle acting\r\n    // on the stale value it just corrected.\r\n    heal_stale_online_records(owlery, &mut std::io::stderr().lock());\r\n\r\n    // Stop side first: watchers whose instance is gone or no longer eligible.\r\n    let live: Vec<(String, String)> = {\r\n        let mut map = set.inner.lock().unwrap_or_else(|p| p.into_inner());\r\n        map.retain(|_, h| !h.thread.is_finished());\r\n        map.keys().cloned().collect()\r\n    };\r\n    for (owner, shell_id) in live {\r\n        let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, &owner, &shell_id);\r\n        let still_eligible = shellinfo::read_shell_info(&perch)\r\n            .map(|i| i.status == SHELL_STATUS_OFFLINE)\r\n            .unwrap_or(false);\r\n        if !still_eligible {\r\n            set.stop_watcher(owlery, &owner, &shell_id);\r\n        }\r\n    }\r\n\r\n    // Start side: every ELIGIBLE instance with a wake_command and headroom.\r\n    //\r\n    // Persistent instances bypass the force-kill freeze; only nonpersistent\r\n    // corpses need the pre-boot discriminant (releases#287).\r\n    let boot_ms = spt_store::proc::boot_instant_ms();\r\n    for owner in spt_store::perch::list_self_perch_ids(owlery) {\r\n        for (shell_id, info) in shellinfo::list_shells(owlery, &owner) {\r\n            if set.contains(&owner, &shell_id) {\r\n                continue;\r\n            }\r\n            let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, &owner, &shell_id);\r\n            let Some(shell) = shell_section_of(registered, adapters_dir, &info.adapter_name) else {\r\n                continue; // deregistered adapter: nothing to run\r\n            };\r\n            if !watcher_eligible(&perch, &info, shell.persistent, boot_ms) {\r\n                continue;\r\n            }\r\n            let Some(wake_command) = shell.wake_command.as_deref() else {\r\n                continue;\r\n            };\r\n            if perch.join(WAKER_GAVE_UP_FILE).exists() {\r\n                continue; // crash-latched until the next shell activity\r\n            }\r\n            // Adopt cleanly: a dead daemon's watcher must not double-run.\r\n            kill_waker_at(&perch);\r\n            // [impl->REQ-INSTALL-11] threading only: this loop already holds the\r\n            // registered set the record's `source_dir` lives in.\r\n            let install_dir = crate::shellhost::shell_install_dir(adapters_dir, &info.adapter_name);\r\n            let tokens = match fill_wake_command(\r\n                &shell_id,\r\n                &info.adapter_name,\r\n                install_dir.as_deref(),\r\n                wake_command,\r\n            ) {\r\n                Ok(t) => t,\r\n                Err(e) => {\r\n                    spt_proto::emit_line_err!(\"WAKER_TEMPLATE:{owner}/{shell_id}: {e}\");\r\n                    continue;\r\n                }\r\n            };\r\n            let stop = Arc::new(AtomicBool::new(false));\r\n            let handle = {\r\n                let stop = Arc::clone(&stop);\r\n                let owlery = owlery.to_path_buf();\r\n                let owner = owner.clone();\r\n                let shell_id = shell_id.clone();\r\n                let adapter_name = info.adapter_name.clone();\r\n                let params = *params;\r\n                std::thread::spawn(move || {\r\n                    watcher_run(&owlery, &owner, &shell_id, &tokens, &params, &stop, || {\r\n                        resolve_wake(&owlery, &owner, &shell_id, &adapter_name, &shell)\r\n                    });\r\n                })\r\n            };\r\n            set.insert(\r\n                &owner,\r\n                &shell_id,\r\n                WatcherHandle {\r\n                    stop,\r\n                    thread: handle,\r\n                },\r\n            );\r\n        }\r\n    }\r\n}\r\n\r\n/// The reconcile cadence (ms): lifecycle flips land in CLI processes, so the\r\n/// daemon's loop is the invariant holder — a flip is at most one tick stale.\r\npub const RECONCILE_INTERVAL_MS: u64 = 5_000;\r\n\r\n/// Spawn the daemon's wake host: one thread sweeping [`reconcile_once`] at\r\n/// boot and every [`RECONCILE_INTERVAL_MS`] until `stop`. The registered set\r\n/// is re-read each sweep (adapter add/remove lands between ticks).\r\n// [impl->REQ-SHELL-2]\r\npub fn spawn_wake_host(stop: Arc<AtomicBool>) -> JoinHandle<()> {\r\n    std::thread::spawn(move || {\r\n        let set = Arc::new(WakeSet::new());\r\n        let params = WakeParams::default();\r\n        let mut owner_edge = OwnerOnlineEdge::new();\r\n        // Once per daemon generation, restore down persistent shells of owners\r\n        // already online. Owners arriving later are covered by the edge below.\r\n        // Neither trigger gates persistence on machine boot or launch age.\r\n        {\r\n            let owlery = spt_store::perch::owlery_dir();\r\n            let adapters_dir = spt_store::perch::adapters_dir();\r\n            let registered = spt_runtime::registry::registered(&adapters_dir);\r\n            // Seeded BEFORE the sweep: an owner that comes online between the two\r\n            // must read as a transition on the first tick, not as already-online.\r\n            owner_edge.seed(&owlery);\r\n            restore_persistent_shells_at_boot(&owlery, &registered, &adapters_dir);\r\n        }\r\n        while !stop.load(Ordering::SeqCst) {\r\n            let owlery = spt_store::perch::owlery_dir();\r\n            let adapters_dir = spt_store::perch::adapters_dir();\r\n            let registered = spt_runtime::registry::registered(&adapters_dir);\r\n            // Restore BEFORE reconciling, the same order the boot pass runs in: a\r\n            // restored instance is not watcher material, and reconcile's start side\r\n            // would otherwise arm a watcher for the shell we are about to launch.\r\n            restore_persistent_shells_on_owner_online(\r\n                &owlery,\r\n                &registered,\r\n                &adapters_dir,\r\n                &mut owner_edge,\r\n                &set,\r\n            );\r\n            reconcile_once(&owlery, &registered, &adapters_dir, &set, &params);\r\n            // Sleep in slices so a stop lands promptly.\r\n            let mut left = RECONCILE_INTERVAL_MS;\r\n            while left > 0 && !stop.load(Ordering::SeqCst) {\r\n                let step = left.min(100);\r\n                std::thread::sleep(Duration::from_millis(step));\r\n                left -= step;\r\n            }\r\n        }\r\n    })\r\n}\r\n\r\n#[cfg(test)]\r\nmod tests {\r\n    use super::*;\r\n    use spt_store::shellinfo::{spawn_record, SHELL_STATUS_ONLINE};\r\n\r\n    /// BACKSTOP for the converge-waits below — a ceiling on an observable, not\r\n    /// a budget any run is expected to consume (the healthy path exits in\r\n    /// milliseconds, and the give-up latch exits the product-fault path at\r\n    /// once).\r\n    ///\r\n    /// ANCHORED, not chosen: it is `slow-timeout.period` from\r\n    /// `.config/nextest.toml`, i.e. exactly the point at which the HARNESS\r\n    /// ITSELF starts calling this test slow. Past that, nextest is already\r\n    /// complaining and a longer wait cannot be the right answer — so the number\r\n    /// has a SOURCE rather than being \"some multiple of the 2s that broke\",\r\n    /// which would be a threshold keyed on a measurement instead of on the\r\n    /// thing. If that config value changes, this should follow it.\r\n    const WAIT_BACKSTOP: Duration = Duration::from_secs(60);\r\n…\n}","startLine":589,"lineNumbers":[589,590,591,592,593,594,595,596,597,598,599,600,601,602,603,604,605,606,607,608,609,610,611,612,613,614,615,616,617,618,619,620,621,622,623,624,625,626,627,628,629,630,631,632,633,634,635,636,637,638,639,640,641,642,643,644,645,646,647,648,649,650,651,652,653,654,655,656,657,658,659,660,661,662,663,664,665,666,667,668,669,670,671,672,673,674,675,676,677,678,679,680,681,682,683,684,685,686,687,688,689,690,691,692,693,694,695,696,697,698,699,700,701,702,703,704,705,706,707,708,709,710,711,712,713,714,715,716,717,718,719,720,721,722,723,724,725,726,727,728,729,730,731,732,733,734,735,736,737,738,739,740,741,742,743,744,745,746,747,748,749,750,751,752,753,754,755,756,757,758,759,760,761,762,763,764,765,766,767,768,769,770,771,772,773,774,775,776,777,778,779,780,781,782,783,784,785,786,787,788,789,790,791,792,793,794,795,796,797,798,799,800,801,802,803,804,805,806,807,808,809,810,811,812,813,814,815,816,817,818,819,820,821,822,823,824,825,826,827,828,829,830,831,832,833,834,835,836,837,838,839,840,841,842,843,844,845,846,847,848,849,850,851,852,853,854,855,856,857,858,859,860,861,862,863,864,865,866,867,868,869,870,871,872,873,874,875,876,877,878,879,880,881,882,883,884,885,886,887,888,889,890,891,892,893,894,895,896,897,898,899,900,901,902,903,904,905,906,907,908,909,910,911,912,913,914,915,916,917,918,919,920,921,922,923,924,925,926,927,928,929,930,931,932,933,934,935,936,937,938,939,940,941,942,943,944,945,946,947,948,949,950,951,952,953,954,955,956,957,958,959,960,961,962,963,964,965,966,967,968,969,970,971,972,973,974,975,976,977,978,979,980,981,982,983,984,985,986,987,988,989,990,991,992,993,994,995,996,997,998,999,1000,1001,1002,1003,1004,1005,1006,1007,1008,1009,1010,1011,1012,1013,1014,1015,1016,1017,1018,1019,1020,1021,1022,1023,1024,1025,1026,1027,1028,1029,1030,1031,1032,1033,1034,1035,1036,1037,1038,1039,1040,1041,1042,1043,1044,1045,1046,1047,1048,1049,1050,1051,1052,1053,1054,1055,1056,1057,1058,1059,1060,1061,1062,1063,null,2161]},"fileSize":100821,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\crates\\spt-daemon\\src\\shellwake.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":2161,"totalBytes":21600,"outputLines":475,"outputBytes":21600,"shownRange":{"start":589,"end":1063},"nextOffset":1064}}},"isError":false,"timestamp":1789001136967}}
{"type":"message","id":"ca63b887","parentId":"ab9d17cd","timestamp":"2026-09-10T00:45:36.975Z","message":{"role":"toolResult","toolCallId":"call_b5vz9AVfJemvkhllIGlgLVCf|fc_077f3569365e313d016aa1fdb0c08087d182419544e6345b9c","toolName":"read","content":[{"type":"text","text":"-/// corpse-boot discriminant refuses that instance from then on. A FAILING launch\n-/// does not -- under a level trigger (\"the owner is online\") it would be retried\n-/// every [`RECONCILE_INTERVAL_MS`] for as long as the owner stayed up. The edge\n-/// bounds it to one attempt per owner-online event.\n+/// An edge bounds a failing launch to one attempt per owner-online event,\n+/// rather than retrying every reconcile tick while the owner remains online.\n+/// A successful launch is protected by birth-safe liveness even before bind.\n #[derive(Debug, Default)]\n pub struct OwnerOnlineEdge {\n     /// Owner id -> online as of the last pass. An owner that disappears from the\n@@ -924,12 +867,8 @@ impl OwnerOnlineEdge {\n /// Daemons boot before endpoints do, so the boot sweep correctly skips an owner\n /// that is not up yet and, being once-per-generation, never revisits it; bringup\n /// from offline emits no rest edge either, so the ADR-0048 cascade does not cover\n-/// it. This is the trigger that does. Re-running the sweep's conjuncts is safe by\n-/// construction and not by convention: `launch_predates_boot` is MONOTONE within a\n-/// machine boot (every launch path restamps `launched_ms`, a kill never restamps,\n-/// `boot_ms` is the machine boot instant), so a later evaluation can only refuse\n-/// MORE. A shell an operator force-killed during this boot is refused here exactly\n-/// as it is refused at boot -- forever, and by the same predicate.\n+/// it. Both triggers use the same down-in-fact decision, without a machine-boot\n+/// or launch-stamp eligibility gate (releases#287).\n // [impl->REQ-SHELL-OWNER-ONLINE-RESTORE]\n // [impl->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\n pub fn restore_persistent_shells_on_owner_online(\n@@ -943,11 +882,6 @@ pub fn restore_persistent_shells_on_owner_online(\n     if owners.is_empty() {\n         return Vec::new();\n     }\n-    let Some(boot_ms) = spt_store::proc::boot_instant_ms() else {\n-        // Same stance as the sweep's: no boot oracle => the discriminant cannot be\n-        // evaluated => restore nothing.\n-        return Vec::new();\n-    };\n     let mut restored = Vec::new();\n     for owner in owners {\n         restore_persistent_shells_of_owner(\n@@ -955,7 +889,6 @@ pub fn restore_persistent_shells_on_owner_online(\n             registered,\n             adapters_dir,\n             &owner,\n-            boot_ms,\n             RestoreTrigger::OwnerOnline,\n             Some(set),\n             &mut restored,\n@@ -975,7 +908,7 @@ pub fn reconcile_once(\n     // it below. Ordering matters -- the watcher-eligibility read further down asks\n     // the RECORDED field, so a heal that ran after it would leave this cycle acting\n     // on the stale value it just corrected.\n-    heal_stale_online_records(owlery);\n+    heal_stale_online_records(owlery, &mut std::io::stderr().lock());\n \n     // Stop side first: watchers whose instance is gone or no longer eligible.\n     let live: Vec<(String, String)> = {\n@@ -995,14 +928,8 @@ pub fn reconcile_once(\n \n     // Start side: every ELIGIBLE instance with a wake_command and headroom.\n     //\n-    // Eligibility is `watcher_eligible` — recorded status (never derived) AND the\n-    // corpse-boot discriminant. Before leg (a) the recorded field was the whole\n-    // guard: a force-killed instance kept a stale `online` record, which is what\n-    // held it out of this set. That record now tells the truth, so the protection\n-    // moved to the discriminant that states it directly (doyle, releases#78\n-    // comment 5156810267) — the same `launch_predates_boot` leg (b) uses, so the\n-    // freeze's forbidden behaviour stays forbidden by construction rather than by\n-    // a second rule.\n+    // Persistent instances bypass the force-kill freeze; only nonpersistent\n+    // corpses need the pre-boot discriminant (releases#287).\n     let boot_ms = spt_store::proc::boot_instant_ms();\n     for owner in spt_store::perch::list_self_perch_ids(owlery) {\n         for (shell_id, info) in shellinfo::list_shells(owlery, &owner) {\n@@ -1010,12 +937,12 @@ pub fn reconcile_once(\n                 continue;\n             }\n             let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, &owner, &shell_id);\n-            if !watcher_eligible(&perch, &info, boot_ms) {\n-                continue;\n-            }\n             let Some(shell) = shell_section_of(registered, adapters_dir, &info.adapter_name) else {\n                 continue; // deregistered adapter: nothing to run\n             };\n+            if !watcher_eligible(&perch, &info, shell.persistent, boot_ms) {\n+                continue;\n+            }\n             let Some(wake_command) = shell.wake_command.as_deref() else {\n                 continue;\n             };\n@@ -1078,15 +1005,9 @@ pub fn spawn_wake_host(stop: Arc<AtomicBool>) -> JoinHandle<()> {\n         let set = Arc::new(WakeSet::new());\n         let params = WakeParams::default();\n         let mut owner_edge = OwnerOnlineEdge::new();\n-        // ONCE per daemon generation, before the reconcile loop: bring back the\n-        // persistent shells a node restart stranded (releases#78 leg b). Outside\n-        // the loop because a node restart is a boot event -- but NOT because\n-        // once-ness protects the operator-force-kill ruling: it does not, and a\n-        // comment here used to say it did. `launch_predates_boot` is what protects\n-        // it, on every evaluation (`reconcile_once`'s start-side comment records\n-        // that ruling; doyle re-measured it on releases#228). The steady-state\n-        // half -- an owner that comes online AFTER this sweep -- is covered in the\n-        // loop below, which is exactly the re-evaluation the false comment forbade.\n+        // Once per daemon generation, restore down persistent shells of owners\n+        // already online. Owners arriving later are covered by the edge below.\n+        // Neither trigger gates persistence on machine boot or launch age.\n         {\n             let owlery = spt_store::perch::owlery_dir();\n             let adapters_dir = spt_store::perch::adapters_dir();\n@@ -1154,6 +1075,7 @@ mod tests {\n     //\n     // The healthy instance beside the stale one is not decoration either: without\n     // it, an implementation that healed EVERY record to offline would pass.\n+    // [unit->REQ-SHELL-HEAL-ENUMERATION]\n     #[test]\n     fn the_heal_writes_the_truth_once_and_then_leaves_the_record_alone() {\n         crate::test_home::with_home(|_| {\n@@ -1166,15 +1088,13 @@ mod tests {\n             )\n             .unwrap();\n \n-            // A STRANDED instance: record online, pid provably a corpse.\n-            let stale = spawn_record(&owlery, \"ling\", \"alchemy\", None).unwrap();\n+            let stale = stranded_shell(&owlery, \"ling\", \"alchemy\", Some(1_000));\n+            let sibling = stranded_shell(&owlery, \"ling\", \"alchemy\", None);\n+            let pacer = stranded_shell(&owlery, \"ling\", \"pacer\", Some(2_000));\n+            seed_owner(&owlery, \"perri\");\n+            let other = stranded_shell(&owlery, \"perri\", \"alchemy\", None);\n             let stale_perch =\n                 spt_store::perch::resolve_shell_perch_path_in(&owlery, \"ling\", &stale);\n-            let mut info = shellinfo::read_shell_info(&stale_perch).unwrap();\n-            info.status = SHELL_STATUS_ONLINE.to_string();\n-            shellinfo::write_shell_info(&stale_perch, &info).unwrap();\n-            std::fs::write(stale_perch.join(shellinfo::SHELL_PID_FILE), \"1\").unwrap();\n-            shellinfo::record_shell_launch(&stale_perch, 1, 1_000);\n \n             // A HEALTHY instance: record online over OUR OWN live pid, stamped by\n             // its own launch. The heal must not touch it.\n@@ -1197,7 +1117,19 @@ mod tests {\n             };\n             let live_before = mtime(&live_perch);\n \n-            heal_stale_online_records(&owlery);\n+            let mut events = Vec::new();\n+            heal_stale_online_records(&owlery, &mut events);\n+            let output = String::from_utf8(events).unwrap();\n+            let mut reported: Vec<_> = output.lines().collect();\n+            reported.sort();\n+            let mut expected = vec![\n+                format!(\"SHELL_RECORD_HEALED:ling/{stale}: online -> offline\"),\n+                format!(\"SHELL_RECORD_HEALED:ling/{sibling}: online -> offline\"),\n+                format!(\"SHELL_RECORD_HEALED:ling/{pacer}: online -> offline\"),\n+                format!(\"SHELL_RECORD_HEALED:perri/{other}: online -> offline\"),\n+            ];\n+            expected.sort();\n+            assert_eq!(reported, expected, \"every healed instance, no healthy representative\");\n \n             assert_eq!(\n                 shellinfo::read_shell_info(&stale_perch).unwrap().status,\n@@ -1219,7 +1151,9 @@ mod tests {\n             // Second cycle over the now-correct records: nothing may be written.\n             let stale_after_heal = mtime(&stale_perch);\n             let live_after_heal = mtime(&live_perch);\n-            heal_stale_online_records(&owlery);\n+            let mut repeated = Vec::new();\n+            heal_stale_online_records(&owlery, &mut repeated);\n+            assert!(repeated.is_empty(), \"already-healed records emit no duplicate events\");\n             assert_eq!(\n                 (mtime(&stale_perch), mtime(&live_perch)),\n                 (stale_after_heal, live_after_heal),\n@@ -1228,17 +1162,8 @@ mod tests {\n         });\n     }\n \n-    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE] the boot discriminant's boundary,\n-    // asserted from BOTH sides of it plus the slack band.\n-    //\n-    // The pair is what makes it evidence: a predicate answering \"yes\" to everything\n-    // would restore force-killed shells (breaking the class-c ruling this\n-    // discriminant exists to preserve), and one answering \"no\" to everything would\n-    // restore nothing and still pass any single-sided row. The slack row is the\n-    // third case and it is not decoration -- the boot instant is DERIVED (Windows:\n-    // now minus uptime), so a process launched in the first moments after boot can\n-    // compute as marginally older than boot itself, and without slack the sweep\n-    // would relaunch something that had only just started.\n+    // [unit->REQ-HAZARD-SHELL-STALE-ONLINE]\n+    // Only the nonpersistent watcher freeze uses this boundary and slack band.\n     #[test]\n     fn only_a_launch_from_before_this_boot_is_a_restart_casualty() {\n         const BOOT: u64 = 1_000_000;\n@@ -1263,214 +1188,127 @@ mod tests {\n         );\n     }\n \n-    // [unit->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS] THE RESTART SHAPE, built\n-    // WITHOUT a suspend edge — the row the existing cascade test cannot express.\n-    //\n-    // `rest_edges_cascade_shells_with_divergence` suspends first, and the suspend\n-    // WRITES the offline record, so by the time its wake edge fires the instance is\n-    // already in the state the to_active arm requires: that fixture can never\n-    // observe this class. This one constructs the restart shape directly — a record\n-    // left ONLINE over a corpse, a launch stamp from before boot, an online owner —\n-    // and calls no rest path anywhere in its setup.\n-    //\n-    // Three instances, because the sweep is a CONJUNCTION and a single-instance\n-    // fixture cannot tell a correct sweep from one that restores everything: the\n-    // stranded one must come back, the one force-killed THIS boot must not, and the\n-    // one whose owner is offline must not.\n-    #[test]\n-    fn the_boot_sweep_restores_only_the_restart_casualty_of_an_online_owner() {\n-        let tmp = tempfile::tempdir().unwrap();\n-        let owlery = tmp.path();\n-        let boot = spt_store::proc::boot_instant_ms().expect(\"a boot instant on this platform\");\n-\n-        // An ONLINE owner (seed_owner writes the `{}` record is_perch_alive reads as\n-        // alive by interim parity) and an explicitly OFFLINE one.\n-        seed_owner(owlery, \"ling\");\n-        let gone = owlery.join(\"gone\");\n-        std::fs::create_dir_all(&gone).unwrap();\n-        std::fs::write(gone.join(\"info.json\"), \"{\\\"status\\\":\\\"offline\\\"}\").unwrap();\n-\n+    fn register_restore_adapter(adapters_dir: &Path, source: &Path, name: &str, persistent: bool) {\n         #[cfg(windows)]\n         let noop = \"cmd /c exit 0\";\n         #[cfg(unix)]\n         let noop = \"true\";\n-        let adapters_dir = tmp.path().join(\"adapters\");\n-        let src = tmp.path().join(\"srcs\").join(\"mock-boot\");\n-        std::fs::create_dir_all(&src).unwrap();\n+        std::fs::create_dir_all(source).unwrap();\n         std::fs::write(\n-            src.join(\"manifest.toml\"),\n+            source.join(\"manifest.toml\"),\n             format!(\n-                \"[adapter]\\nname = \\\"mock-boot\\\"\\nkind = \\\"shell\\\"\\nversion = \\\"1\\\"\\n\\\n+                \"[adapter]\\nname = \\\"{name}\\\"\\nkind = \\\"shell\\\"\\nversion = \\\"1\\\"\\n\\\n                  min_spt_core_version = \\\"0\\\"\\n\\n[shell]\\nspawn = '{noop}'\\n\\\n-                 persistent = true\\n\"\n+                 persistent = {persistent}\\n\"\n             ),\n         )\n         .unwrap();\n-        spt_runtime::registry::register(&adapters_dir, &src, 1).unwrap();\n-        let registered = spt_runtime::registry::registered(&adapters_dir);\n+        spt_runtime::registry::register(adapters_dir, source, 1).unwrap();\n+    }\n \n-        // Every instance is recorded ONLINE over a dead pid — the restart shape —\n-        // and they differ ONLY in the gate under test.\n-        let mk = |owner: &str, launched_ms: u64| -> String {\n-            let id = spawn_record(owlery, owner, \"mock-boot\", None).unwrap();\n-            let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, &id);\n-            let mut i = shellinfo::read_shell_info(&perch).unwrap();\n-            i.status = SHELL_STATUS_ONLINE.to_string();\n-            shellinfo::write_shell_info(&perch, &i).unwrap();\n-            std::fs::write(perch.join(shellinfo::SHELL_PID_FILE), \"2000000000\").unwrap();\n-            let launch = shellinfo::ShellLaunch {\n-                pid_started_at: Some(1),\n-                launched_ms,\n-            };\n-            std::fs::write(\n-                perch.join(shellinfo::SHELL_LAUNCH_FILE),\n-                serde_json::to_string(&launch).unwrap(),\n-            )\n-            .unwrap();\n-            id\n-        };\n-        let stranded = mk(\"ling\", boot.saturating_sub(600_000)); // 10 min before boot\n-        let force_killed = mk(\"ling\", boot + 120_000); // 2 min AFTER boot\n-        let orphan_owner = mk(\"gone\", boot.saturating_sub(600_000));\n+    /// Restart-shaped record, with NO suspend edge to pre-heal its status.\n+    fn stranded_shell(owlery: &Path, owner: &str, adapter: &str, launched_ms: Option<u64>) -> String {\n+        let id = spawn_record(owlery, owner, adapter, None).unwrap();\n+        let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, &id);\n+        let mut info = shellinfo::read_shell_info(&perch).unwrap();\n+        info.status = SHELL_STATUS_ONLINE.to_string();\n+        shellinfo::write_shell_info(&perch, &info).unwrap();\n+        std::fs::write(perch.join(shellinfo::SHELL_PID_FILE), \"2000000000\").unwrap();\n+        if let Some(launched_ms) = launched_ms {\n+            shellinfo::record_shell_launch(&perch, 2_000_000_000, launched_ms);\n+        }\n+        id\n+    }\n+\n+    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n+    // [unit->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\n+    #[test]\n+    fn boot_restores_all_down_persistent_instances_without_launch_age_gates() {\n+        let tmp = tempfile::tempdir().unwrap();\n+        let owlery = tmp.path();\n+        let boot = spt_store::proc::boot_instant_ms().expect(\"boot instant\");\n+        let adapters_dir = tmp.path().join(\"adapters\");\n+        register_restore_adapter(&adapters_dir, &tmp.path().join(\"keeper\"), \"keeper\", true);\n+        register_restore_adapter(&adapters_dir, &tmp.path().join(\"ordinary\"), \"ordinary\", false);\n+        let registered = spt_runtime::registry::registered(&adapters_dir);\n+        seed_owner(owlery, \"ling\");\n+        let gone = owlery.join(\"gone\");\n+        std::fs::create_dir_all(&gone).unwrap();\n+        std::fs::write(gone.join(\"info.json\"), \"{\\\"status\\\":\\\"offline\\\"}\").unwrap();\n \n-        let restored = restore_persistent_shells_at_boot(owlery, &registered, &adapters_dir);\n+        let before = stranded_shell(owlery, \"ling\", \"keeper\", Some(boot.saturating_sub(600_000)));\n+        let after = stranded_shell(owlery, \"ling\", \"keeper\", Some(boot + 120_000));\n+        let unstamped = stranded_shell(owlery, \"ling\", \"keeper\", None);\n+        stranded_shell(owlery, \"gone\", \"keeper\", None);\n+        stranded_shell(owlery, \"ling\", \"ordinary\", Some(boot + 120_000));\n \n\n[Artifact storage: ~/.omp/agent/sessions/-Documents-projects-spt-core/2026-09-09T23-45-42-317Z_01a08890-21ad-7000-8f54-208d4f4f2f2c/112.bash-original.log (105.3KB). Use artifact://112:N-M to page, artifact://112:raw:N-M for verbatim chunks, and the artifact file path for search/copy workflows.]\n\n[Showing lines 300-603 of 1088. Use :604 to continue]"}],"details":{"resolvedPath":"C:\\Users\\decid\\.omp\\agent\\sessions\\-Documents-projects-spt-core\\2026-09-09T23-45-42-317Z_01a08890-21ad-7000-8f54-208d4f4f2f2c\\112.bash-original.log","contentType":"text/plain","totalLines":1088,"displayContent":{"text":"-/// corpse-boot discriminant refuses that instance from then on. A FAILING launch\n-/// does not -- under a level trigger (\"the owner is online\") it would be retried\n-/// every [`RECONCILE_INTERVAL_MS`] for as long as the owner stayed up. The edge\n-/// bounds it to one attempt per owner-online event.\n+/// An edge bounds a failing launch to one attempt per owner-online event,\n+/// rather than retrying every reconcile tick while the owner remains online.\n+/// A successful launch is protected by birth-safe liveness even before bind.\n #[derive(Debug, Default)]\n pub struct OwnerOnlineEdge {\n     /// Owner id -> online as of the last pass. An owner that disappears from the\n@@ -924,12 +867,8 @@ impl OwnerOnlineEdge {\n /// Daemons boot before endpoints do, so the boot sweep correctly skips an owner\n /// that is not up yet and, being once-per-generation, never revisits it; bringup\n /// from offline emits no rest edge either, so the ADR-0048 cascade does not cover\n-/// it. This is the trigger that does. Re-running the sweep's conjuncts is safe by\n-/// construction and not by convention: `launch_predates_boot` is MONOTONE within a\n-/// machine boot (every launch path restamps `launched_ms`, a kill never restamps,\n-/// `boot_ms` is the machine boot instant), so a later evaluation can only refuse\n-/// MORE. A shell an operator force-killed during this boot is refused here exactly\n-/// as it is refused at boot -- forever, and by the same predicate.\n+/// it. Both triggers use the same down-in-fact decision, without a machine-boot\n+/// or launch-stamp eligibility gate (releases#287).\n // [impl->REQ-SHELL-OWNER-ONLINE-RESTORE]\n // [impl->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\n pub fn restore_persistent_shells_on_owner_online(\n@@ -943,11 +882,6 @@ pub fn restore_persistent_shells_on_owner_online(\n     if owners.is_empty() {\n         return Vec::new();\n     }\n-    let Some(boot_ms) = spt_store::proc::boot_instant_ms() else {\n-        // Same stance as the sweep's: no boot oracle => the discriminant cannot be\n-        // evaluated => restore nothing.\n-        return Vec::new();\n-    };\n     let mut restored = Vec::new();\n     for owner in owners {\n         restore_persistent_shells_of_owner(\n@@ -955,7 +889,6 @@ pub fn restore_persistent_shells_on_owner_online(\n             registered,\n             adapters_dir,\n             &owner,\n-            boot_ms,\n             RestoreTrigger::OwnerOnline,\n             Some(set),\n             &mut restored,\n@@ -975,7 +908,7 @@ pub fn reconcile_once(\n     // it below. Ordering matters -- the watcher-eligibility read further down asks\n     // the RECORDED field, so a heal that ran after it would leave this cycle acting\n     // on the stale value it just corrected.\n-    heal_stale_online_records(owlery);\n+    heal_stale_online_records(owlery, &mut std::io::stderr().lock());\n \n     // Stop side first: watchers whose instance is gone or no longer eligible.\n     let live: Vec<(String, String)> = {\n@@ -995,14 +928,8 @@ pub fn reconcile_once(\n \n     // Start side: every ELIGIBLE instance with a wake_command and headroom.\n     //\n-    // Eligibility is `watcher_eligible` — recorded status (never derived) AND the\n-    // corpse-boot discriminant. Before leg (a) the recorded field was the whole\n-    // guard: a force-killed instance kept a stale `online` record, which is what\n-    // held it out of this set. That record now tells the truth, so the protection\n-    // moved to the discriminant that states it directly (doyle, releases#78\n-    // comment 5156810267) — the same `launch_predates_boot` leg (b) uses, so the\n-    // freeze's forbidden behaviour stays forbidden by construction rather than by\n-    // a second rule.\n+    // Persistent instances bypass the force-kill freeze; only nonpersistent\n+    // corpses need the pre-boot discriminant (releases#287).\n     let boot_ms = spt_store::proc::boot_instant_ms();\n     for owner in spt_store::perch::list_self_perch_ids(owlery) {\n         for (shell_id, info) in shellinfo::list_shells(owlery, &owner) {\n@@ -1010,12 +937,12 @@ pub fn reconcile_once(\n                 continue;\n             }\n             let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, &owner, &shell_id);\n-            if !watcher_eligible(&perch, &info, boot_ms) {\n-                continue;\n-            }\n             let Some(shell) = shell_section_of(registered, adapters_dir, &info.adapter_name) else {\n                 continue; // deregistered adapter: nothing to run\n             };\n+            if !watcher_eligible(&perch, &info, shell.persistent, boot_ms) {\n+                continue;\n+            }\n             let Some(wake_command) = shell.wake_command.as_deref() else {\n                 continue;\n             };\n@@ -1078,15 +1005,9 @@ pub fn spawn_wake_host(stop: Arc<AtomicBool>) -> JoinHandle<()> {\n         let set = Arc::new(WakeSet::new());\n         let params = WakeParams::default();\n         let mut owner_edge = OwnerOnlineEdge::new();\n-        // ONCE per daemon generation, before the reconcile loop: bring back the\n-        // persistent shells a node restart stranded (releases#78 leg b). Outside\n-        // the loop because a node restart is a boot event -- but NOT because\n-        // once-ness protects the operator-force-kill ruling: it does not, and a\n-        // comment here used to say it did. `launch_predates_boot` is what protects\n-        // it, on every evaluation (`reconcile_once`'s start-side comment records\n-        // that ruling; doyle re-measured it on releases#228). The steady-state\n-        // half -- an owner that comes online AFTER this sweep -- is covered in the\n-        // loop below, which is exactly the re-evaluation the false comment forbade.\n+        // Once per daemon generation, restore down persistent shells of owners\n+        // already online. Owners arriving later are covered by the edge below.\n+        // Neither trigger gates persistence on machine boot or launch age.\n         {\n             let owlery = spt_store::perch::owlery_dir();\n             let adapters_dir = spt_store::perch::adapters_dir();\n@@ -1154,6 +1075,7 @@ mod tests {\n     //\n     // The healthy instance beside the stale one is not decoration either: without\n     // it, an implementation that healed EVERY record to offline would pass.\n+    // [unit->REQ-SHELL-HEAL-ENUMERATION]\n     #[test]\n     fn the_heal_writes_the_truth_once_and_then_leaves_the_record_alone() {\n         crate::test_home::with_home(|_| {\n@@ -1166,15 +1088,13 @@ mod tests {\n             )\n             .unwrap();\n \n-            // A STRANDED instance: record online, pid provably a corpse.\n-            let stale = spawn_record(&owlery, \"ling\", \"alchemy\", None).unwrap();\n+            let stale = stranded_shell(&owlery, \"ling\", \"alchemy\", Some(1_000));\n+            let sibling = stranded_shell(&owlery, \"ling\", \"alchemy\", None);\n+            let pacer = stranded_shell(&owlery, \"ling\", \"pacer\", Some(2_000));\n+            seed_owner(&owlery, \"perri\");\n+            let other = stranded_shell(&owlery, \"perri\", \"alchemy\", None);\n             let stale_perch =\n                 spt_store::perch::resolve_shell_perch_path_in(&owlery, \"ling\", &stale);\n-            let mut info = shellinfo::read_shell_info(&stale_perch).unwrap();\n-            info.status = SHELL_STATUS_ONLINE.to_string();\n-            shellinfo::write_shell_info(&stale_perch, &info).unwrap();\n-            std::fs::write(stale_perch.join(shellinfo::SHELL_PID_FILE), \"1\").unwrap();\n-            shellinfo::record_shell_launch(&stale_perch, 1, 1_000);\n \n             // A HEALTHY instance: record online over OUR OWN live pid, stamped by\n             // its own launch. The heal must not touch it.\n@@ -1197,7 +1117,19 @@ mod tests {\n             };\n             let live_before = mtime(&live_perch);\n \n-            heal_stale_online_records(&owlery);\n+            let mut events = Vec::new();\n+            heal_stale_online_records(&owlery, &mut events);\n+            let output = String::from_utf8(events).unwrap();\n+            let mut reported: Vec<_> = output.lines().collect();\n+            reported.sort();\n+            let mut expected = vec![\n+                format!(\"SHELL_RECORD_HEALED:ling/{stale}: online -> offline\"),\n+                format!(\"SHELL_RECORD_HEALED:ling/{sibling}: online -> offline\"),\n+                format!(\"SHELL_RECORD_HEALED:ling/{pacer}: online -> offline\"),\n+                format!(\"SHELL_RECORD_HEALED:perri/{other}: online -> offline\"),\n+            ];\n+            expected.sort();\n+            assert_eq!(reported, expected, \"every healed instance, no healthy representative\");\n \n             assert_eq!(\n                 shellinfo::read_shell_info(&stale_perch).unwrap().status,\n@@ -1219,7 +1151,9 @@ mod tests {\n             // Second cycle over the now-correct records: nothing may be written.\n             let stale_after_heal = mtime(&stale_perch);\n             let live_after_heal = mtime(&live_perch);\n-            heal_stale_online_records(&owlery);\n+            let mut repeated = Vec::new();\n+            heal_stale_online_records(&owlery, &mut repeated);\n+            assert!(repeated.is_empty(), \"already-healed records emit no duplicate events\");\n             assert_eq!(\n                 (mtime(&stale_perch), mtime(&live_perch)),\n                 (stale_after_heal, live_after_heal),\n@@ -1228,17 +1162,8 @@ mod tests {\n         });\n     }\n \n-    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE] the boot discriminant's boundary,\n-    // asserted from BOTH sides of it plus the slack band.\n-    //\n-    // The pair is what makes it evidence: a predicate answering \"yes\" to everything\n-    // would restore force-killed shells (breaking the class-c ruling this\n-    // discriminant exists to preserve), and one answering \"no\" to everything would\n-    // restore nothing and still pass any single-sided row. The slack row is the\n-    // third case and it is not decoration -- the boot instant is DERIVED (Windows:\n-    // now minus uptime), so a process launched in the first moments after boot can\n-    // compute as marginally older than boot itself, and without slack the sweep\n-    // would relaunch something that had only just started.\n+    // [unit->REQ-HAZARD-SHELL-STALE-ONLINE]\n+    // Only the nonpersistent watcher freeze uses this boundary and slack band.\n     #[test]\n     fn only_a_launch_from_before_this_boot_is_a_restart_casualty() {\n         const BOOT: u64 = 1_000_000;\n@@ -1263,214 +1188,127 @@ mod tests {\n         );\n     }\n \n-    // [unit->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS] THE RESTART SHAPE, built\n-    // WITHOUT a suspend edge — the row the existing cascade test cannot express.\n-    //\n-    // `rest_edges_cascade_shells_with_divergence` suspends first, and the suspend\n-    // WRITES the offline record, so by the time its wake edge fires the instance is\n-    // already in the state the to_active arm requires: that fixture can never\n-    // observe this class. This one constructs the restart shape directly — a record\n-    // left ONLINE over a corpse, a launch stamp from before boot, an online owner —\n-    // and calls no rest path anywhere in its setup.\n-    //\n-    // Three instances, because the sweep is a CONJUNCTION and a single-instance\n-    // fixture cannot tell a correct sweep from one that restores everything: the\n-    // stranded one must come back, the one force-killed THIS boot must not, and the\n-    // one whose owner is offline must not.\n-    #[test]\n-    fn the_boot_sweep_restores_only_the_restart_casualty_of_an_online_owner() {\n-        let tmp = tempfile::tempdir().unwrap();\n-        let owlery = tmp.path();\n-        let boot = spt_store::proc::boot_instant_ms().expect(\"a boot instant on this platform\");\n-\n-        // An ONLINE owner (seed_owner writes the `{}` record is_perch_alive reads as\n-        // alive by interim parity) and an explicitly OFFLINE one.\n-        seed_owner(owlery, \"ling\");\n-        let gone = owlery.join(\"gone\");\n-        std::fs::create_dir_all(&gone).unwrap();\n-        std::fs::write(gone.join(\"info.json\"), \"{\\\"status\\\":\\\"offline\\\"}\").unwrap();\n-\n+    fn register_restore_adapter(adapters_dir: &Path, source: &Path, name: &str, persistent: bool) {\n         #[cfg(windows)]\n         let noop = \"cmd /c exit 0\";\n         #[cfg(unix)]\n         let noop = \"true\";\n-        let adapters_dir = tmp.path().join(\"adapters\");\n-        let src = tmp.path().join(\"srcs\").join(\"mock-boot\");\n-        std::fs::create_dir_all(&src).unwrap();\n+        std::fs::create_dir_all(source).unwrap();\n         std::fs::write(\n-            src.join(\"manifest.toml\"),\n+            source.join(\"manifest.toml\"),\n             format!(\n-                \"[adapter]\\nname = \\\"mock-boot\\\"\\nkind = \\\"shell\\\"\\nversion = \\\"1\\\"\\n\\\n+                \"[adapter]\\nname = \\\"{name}\\\"\\nkind = \\\"shell\\\"\\nversion = \\\"1\\\"\\n\\\n                  min_spt_core_version = \\\"0\\\"\\n\\n[shell]\\nspawn = '{noop}'\\n\\\n-                 persistent = true\\n\"\n+                 persistent = {persistent}\\n\"\n             ),\n         )\n         .unwrap();\n-        spt_runtime::registry::register(&adapters_dir, &src, 1).unwrap();\n-        let registered = spt_runtime::registry::registered(&adapters_dir);\n+        spt_runtime::registry::register(adapters_dir, source, 1).unwrap();\n+    }\n \n-        // Every instance is recorded ONLINE over a dead pid — the restart shape —\n-        // and they differ ONLY in the gate under test.\n-        let mk = |owner: &str, launched_ms: u64| -> String {\n-            let id = spawn_record(owlery, owner, \"mock-boot\", None).unwrap();\n-            let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, &id);\n-            let mut i = shellinfo::read_shell_info(&perch).unwrap();\n-            i.status = SHELL_STATUS_ONLINE.to_string();\n-            shellinfo::write_shell_info(&perch, &i).unwrap();\n-            std::fs::write(perch.join(shellinfo::SHELL_PID_FILE), \"2000000000\").unwrap();\n-            let launch = shellinfo::ShellLaunch {\n-                pid_started_at: Some(1),\n-                launched_ms,\n-            };\n-            std::fs::write(\n-                perch.join(shellinfo::SHELL_LAUNCH_FILE),\n-                serde_json::to_string(&launch).unwrap(),\n-            )\n-            .unwrap();\n-            id\n-        };\n-        let stranded = mk(\"ling\", boot.saturating_sub(600_000)); // 10 min before boot\n-        let force_killed = mk(\"ling\", boot + 120_000); // 2 min AFTER boot\n-        let orphan_owner = mk(\"gone\", boot.saturating_sub(600_000));\n+    /// Restart-shaped record, with NO suspend edge to pre-heal its status.\n+    fn stranded_shell(owlery: &Path, owner: &str, adapter: &str, launched_ms: Option<u64>) -> String {\n+        let id = spawn_record(owlery, owner, adapter, None).unwrap();\n+        let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, &id);\n+        let mut info = shellinfo::read_shell_info(&perch).unwrap();\n+        info.status = SHELL_STATUS_ONLINE.to_string();\n+        shellinfo::write_shell_info(&perch, &info).unwrap();\n+        std::fs::write(perch.join(shellinfo::SHELL_PID_FILE), \"2000000000\").unwrap();\n+        if let Some(launched_ms) = launched_ms {\n+            shellinfo::record_shell_launch(&perch, 2_000_000_000, launched_ms);\n+        }\n+        id\n+    }\n+\n+    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n+    // [unit->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\n+    #[test]\n+    fn boot_restores_all_down_persistent_instances_without_launch_age_gates() {\n+        let tmp = tempfile::tempdir().unwrap();\n+        let owlery = tmp.path();\n+        let boot = spt_store::proc::boot_instant_ms().expect(\"boot instant\");\n+        let adapters_dir = tmp.path().join(\"adapters\");\n+        register_restore_adapter(&adapters_dir, &tmp.path().join(\"keeper\"), \"keeper\", true);\n+        register_restore_adapter(&adapters_dir, &tmp.path().join(\"ordinary\"), \"ordinary\", false);\n+        let registered = spt_runtime::registry::registered(&adapters_dir);\n+        seed_owner(owlery, \"ling\");\n+        let gone = owlery.join(\"gone\");\n+        std::fs::create_dir_all(&gone).unwrap();\n+        std::fs::write(gone.join(\"info.json\"), \"{\\\"status\\\":\\\"offline\\\"}\").unwrap();\n \n-        let restored = restore_persistent_shells_at_boot(owlery, &registered, &adapters_dir);\n+        let before = stranded_shell(owlery, \"ling\", \"keeper\", Some(boot.saturating_sub(600_000)));\n+        let after = stranded_shell(owlery, \"ling\", \"keeper\", Some(boot + 120_000));\n+        let unstamped = stranded_shell(owlery, \"ling\", \"keeper\", None);\n+        stranded_shell(owlery, \"gone\", \"keeper\", None);\n+        stranded_shell(owlery, \"ling\", \"ordinary\", Some(boot + 120_000));\n ","startLine":300,"lineNumbers":[300,301,302,303,304,305,306,307,308,309,310,311,312,313,314,315,316,317,318,319,320,321,322,323,324,325,326,327,328,329,330,331,332,333,334,335,336,337,338,339,340,341,342,343,344,345,346,347,348,349,350,351,352,353,354,355,356,357,358,359,360,361,362,363,364,365,366,367,368,369,370,371,372,373,374,375,376,377,378,379,380,381,382,383,384,385,386,387,388,389,390,391,392,393,394,395,396,397,398,399,400,401,402,403,404,405,406,407,408,409,410,411,412,413,414,415,416,417,418,419,420,421,422,423,424,425,426,427,428,429,430,431,432,433,434,435,436,437,438,439,440,441,442,443,444,445,446,447,448,449,450,451,452,453,454,455,456,457,458,459,460,461,462,463,464,465,466,467,468,469,470,471,472,473,474,475,476,477,478,479,480,481,482,483,484,485,486,487,488,489,490,491,492,493,494,495,496,497,498,499,500,501,502,503,504,505,506,507,508,509,510,511,512,513,514,515,516,517,518,519,520,521,522,523,524,525,526,527,528,529,530,531,532,533,534,535,536,537,538,539,540,541,542,543,544,545,546,547,548,549,550,551,552,553,554,555,556,557,558,559,560,561,562,563,564,565,566,567,568,569,570,571,572,573,574,575,576,577,578,579,580,581,582,583,584,585,586,587,588,589,590,591,592,593,594,595,596,597,598,599,600,601,602,603]},"truncation":{"content":"-/// corpse-boot discriminant refuses that instance from then on. A FAILING launch\n-/// does not -- under a level trigger (\"the owner is online\") it would be retried\n-/// every [`RECONCILE_INTERVAL_MS`] for as long as the owner stayed up. The edge\n-/// bounds it to one attempt per owner-online event.\n+/// An edge bounds a failing launch to one attempt per owner-online event,\n+/// rather than retrying every reconcile tick while the owner remains online.\n+/// A successful launch is protected by birth-safe liveness even before bind.\n #[derive(Debug, Default)]\n pub struct OwnerOnlineEdge {\n     /// Owner id -> online as of the last pass. An owner that disappears from the\n@@ -924,12 +867,8 @@ impl OwnerOnlineEdge {\n /// Daemons boot before endpoints do, so the boot sweep correctly skips an owner\n /// that is not up yet and, being once-per-generation, never revisits it; bringup\n /// from offline emits no rest edge either, so the ADR-0048 cascade does not cover\n-/// it. This is the trigger that does. Re-running the sweep's conjuncts is safe by\n-/// construction and not by convention: `launch_predates_boot` is MONOTONE within a\n-/// machine boot (every launch path restamps `launched_ms`, a kill never restamps,\n-/// `boot_ms` is the machine boot instant), so a later evaluation can only refuse\n-/// MORE. A shell an operator force-killed during this boot is refused here exactly\n-/// as it is refused at boot -- forever, and by the same predicate.\n+/// it. Both triggers use the same down-in-fact decision, without a machine-boot\n+/// or launch-stamp eligibility gate (releases#287).\n // [impl->REQ-SHELL-OWNER-ONLINE-RESTORE]\n // [impl->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\n pub fn restore_persistent_shells_on_owner_online(\n@@ -943,11 +882,6 @@ pub fn restore_persistent_shells_on_owner_online(\n     if owners.is_empty() {\n         return Vec::new();\n     }\n-    let Some(boot_ms) = spt_store::proc::boot_instant_ms() else {\n-        // Same stance as the sweep's: no boot oracle => the discriminant cannot be\n-        // evaluated => restore nothing.\n-        return Vec::new();\n-    };\n     let mut restored = Vec::new();\n     for owner in owners {\n         restore_persistent_shells_of_owner(\n@@ -955,7 +889,6 @@ pub fn restore_persistent_shells_on_owner_online(\n             registered,\n             adapters_dir,\n             &owner,\n-            boot_ms,\n             RestoreTrigger::OwnerOnline,\n             Some(set),\n             &mut restored,\n@@ -975,7 +908,7 @@ pub fn reconcile_once(\n     // it below. Ordering matters -- the watcher-eligibility read further down asks\n     // the RECORDED field, so a heal that ran after it would leave this cycle acting\n     // on the stale value it just corrected.\n-    heal_stale_online_records(owlery);\n+    heal_stale_online_records(owlery, &mut std::io::stderr().lock());\n \n     // Stop side first: watchers whose instance is gone or no longer eligible.\n     let live: Vec<(String, String)> = {\n@@ -995,14 +928,8 @@ pub fn reconcile_once(\n \n     // Start side: every ELIGIBLE instance with a wake_command and headroom.\n     //\n-    // Eligibility is `watcher_eligible` — recorded status (never derived) AND the\n-    // corpse-boot discriminant. Before leg (a) the recorded field was the whole\n-    // guard: a force-killed instance kept a stale `online` record, which is what\n-    // held it out of this set. That record now tells the truth, so the protection\n-    // moved to the discriminant that states it directly (doyle, releases#78\n-    // comment 5156810267) — the same `launch_predates_boot` leg (b) uses, so the\n-    // freeze's forbidden behaviour stays forbidden by construction rather than by\n-    // a second rule.\n+    // Persistent instances bypass the force-kill freeze; only nonpersistent\n+    // corpses need the pre-boot discriminant (releases#287).\n     let boot_ms = spt_store::proc::boot_instant_ms();\n     for owner in spt_store::perch::list_self_perch_ids(owlery) {\n         for (shell_id, info) in shellinfo::list_shells(owlery, &owner) {\n@@ -1010,12 +937,12 @@ pub fn reconcile_once(\n                 continue;\n             }\n             let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, &owner, &shell_id);\n-            if !watcher_eligible(&perch, &info, boot_ms) {\n-                continue;\n-            }\n             let Some(shell) = shell_section_of(registered, adapters_dir, &info.adapter_name) else {\n                 continue; // deregistered adapter: nothing to run\n             };\n+            if !watcher_eligible(&perch, &info, shell.persistent, boot_ms) {\n+                continue;\n+            }\n             let Some(wake_command) = shell.wake_command.as_deref() else {\n                 continue;\n             };\n@@ -1078,15 +1005,9 @@ pub fn spawn_wake_host(stop: Arc<AtomicBool>) -> JoinHandle<()> {\n         let set = Arc::new(WakeSet::new());\n         let params = WakeParams::default();\n         let mut owner_edge = OwnerOnlineEdge::new();\n-        // ONCE per daemon generation, before the reconcile loop: bring back the\n-        // persistent shells a node restart stranded (releases#78 leg b). Outside\n-        // the loop because a node restart is a boot event -- but NOT because\n-        // once-ness protects the operator-force-kill ruling: it does not, and a\n-        // comment here used to say it did. `launch_predates_boot` is what protects\n-        // it, on every evaluation (`reconcile_once`'s start-side comment records\n-        // that ruling; doyle re-measured it on releases#228). The steady-state\n-        // half -- an owner that comes online AFTER this sweep -- is covered in the\n-        // loop below, which is exactly the re-evaluation the false comment forbade.\n+        // Once per daemon generation, restore down persistent shells of owners\n+        // already online. Owners arriving later are covered by the edge below.\n+        // Neither trigger gates persistence on machine boot or launch age.\n         {\n             let owlery = spt_store::perch::owlery_dir();\n             let adapters_dir = spt_store::perch::adapters_dir();\n@@ -1154,6 +1075,7 @@ mod tests {\n     //\n     // The healthy instance beside the stale one is not decoration either: without\n     // it, an implementation that healed EVERY record to offline would pass.\n+    // [unit->REQ-SHELL-HEAL-ENUMERATION]\n     #[test]\n     fn the_heal_writes_the_truth_once_and_then_leaves_the_record_alone() {\n         crate::test_home::with_home(|_| {\n@@ -1166,15 +1088,13 @@ mod tests {\n             )\n             .unwrap();\n \n-            // A STRANDED instance: record online, pid provably a corpse.\n-            let stale = spawn_record(&owlery, \"ling\", \"alchemy\", None).unwrap();\n+            let stale = stranded_shell(&owlery, \"ling\", \"alchemy\", Some(1_000));\n+            let sibling = stranded_shell(&owlery, \"ling\", \"alchemy\", None);\n+            let pacer = stranded_shell(&owlery, \"ling\", \"pacer\", Some(2_000));\n+            seed_owner(&owlery, \"perri\");\n+            let other = stranded_shell(&owlery, \"perri\", \"alchemy\", None);\n             let stale_perch =\n                 spt_store::perch::resolve_shell_perch_path_in(&owlery, \"ling\", &stale);\n-            let mut info = shellinfo::read_shell_info(&stale_perch).unwrap();\n-            info.status = SHELL_STATUS_ONLINE.to_string();\n-            shellinfo::write_shell_info(&stale_perch, &info).unwrap();\n-            std::fs::write(stale_perch.join(shellinfo::SHELL_PID_FILE), \"1\").unwrap();\n-            shellinfo::record_shell_launch(&stale_perch, 1, 1_000);\n \n             // A HEALTHY instance: record online over OUR OWN live pid, stamped by\n             // its own launch. The heal must not touch it.\n@@ -1197,7 +1117,19 @@ mod tests {\n             };\n             let live_before = mtime(&live_perch);\n \n-            heal_stale_online_records(&owlery);\n+            let mut events = Vec::new();\n+            heal_stale_online_records(&owlery, &mut events);\n+            let output = String::from_utf8(events).unwrap();\n+            let mut reported: Vec<_> = output.lines().collect();\n+            reported.sort();\n+            let mut expected = vec![\n+                format!(\"SHELL_RECORD_HEALED:ling/{stale}: online -> offline\"),\n+                format!(\"SHELL_RECORD_HEALED:ling/{sibling}: online -> offline\"),\n+                format!(\"SHELL_RECORD_HEALED:ling/{pacer}: online -> offline\"),\n+                format!(\"SHELL_RECORD_HEALED:perri/{other}: online -> offline\"),\n+            ];\n+            expected.sort();\n+            assert_eq!(reported, expected, \"every healed instance, no healthy representative\");\n \n             assert_eq!(\n                 shellinfo::read_shell_info(&stale_perch).unwrap().status,\n@@ -1219,7 +1151,9 @@ mod tests {\n             // Second cycle over the now-correct records: nothing may be written.\n             let stale_after_heal = mtime(&stale_perch);\n             let live_after_heal = mtime(&live_perch);\n-            heal_stale_online_records(&owlery);\n+            let mut repeated = Vec::new();\n+            heal_stale_online_records(&owlery, &mut repeated);\n+            assert!(repeated.is_empty(), \"already-healed records emit no duplicate events\");\n             assert_eq!(\n                 (mtime(&stale_perch), mtime(&live_perch)),\n                 (stale_after_heal, live_after_heal),\n@@ -1228,17 +1162,8 @@ mod tests {\n         });\n     }\n \n-    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE] the boot discriminant's boundary,\n-    // asserted from BOTH sides of it plus the slack band.\n-    //\n-    // The pair is what makes it evidence: a predicate answering \"yes\" to everything\n-    // would restore force-killed shells (breaking the class-c ruling this\n-    // discriminant exists to preserve), and one answering \"no\" to everything would\n-    // restore nothing and still pass any single-sided row. The slack row is the\n-    // third case and it is not decoration -- the boot instant is DERIVED (Windows:\n-    // now minus uptime), so a process launched in the first moments after boot can\n-    // compute as marginally older than boot itself, and without slack the sweep\n-    // would relaunch something that had only just started.\n+    // [unit->REQ-HAZARD-SHELL-STALE-ONLINE]\n+    // Only the nonpersistent watcher freeze uses this boundary and slack band.\n     #[test]\n     fn only_a_launch_from_before_this_boot_is_a_restart_casualty() {\n         const BOOT: u64 = 1_000_000;\n@@ -1263,214 +1188,127 @@ mod tests {\n         );\n     }\n \n-    // [unit->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS] THE RESTART SHAPE, built\n-    // WITHOUT a suspend edge — the row the existing cascade test cannot express.\n-    //\n-    // `rest_edges_cascade_shells_with_divergence` suspends first, and the suspend\n-    // WRITES the offline record, so by the time its wake edge fires the instance is\n-    // already in the state the to_active arm requires: that fixture can never\n-    // observe this class. This one constructs the restart shape directly — a record\n-    // left ONLINE over a corpse, a launch stamp from before boot, an online owner —\n-    // and calls no rest path anywhere in its setup.\n-    //\n-    // Three instances, because the sweep is a CONJUNCTION and a single-instance\n-    // fixture cannot tell a correct sweep from one that restores everything: the\n-    // stranded one must come back, the one force-killed THIS boot must not, and the\n-    // one whose owner is offline must not.\n-    #[test]\n-    fn the_boot_sweep_restores_only_the_restart_casualty_of_an_online_owner() {\n-        let tmp = tempfile::tempdir().unwrap();\n-        let owlery = tmp.path();\n-        let boot = spt_store::proc::boot_instant_ms().expect(\"a boot instant on this platform\");\n-\n-        // An ONLINE owner (seed_owner writes the `{}` record is_perch_alive reads as\n-        // alive by interim parity) and an explicitly OFFLINE one.\n-        seed_owner(owlery, \"ling\");\n-        let gone = owlery.join(\"gone\");\n-        std::fs::create_dir_all(&gone).unwrap();\n-        std::fs::write(gone.join(\"info.json\"), \"{\\\"status\\\":\\\"offline\\\"}\").unwrap();\n-\n+    fn register_restore_adapter(adapters_dir: &Path, source: &Path, name: &str, persistent: bool) {\n         #[cfg(windows)]\n         let noop = \"cmd /c exit 0\";\n         #[cfg(unix)]\n         let noop = \"true\";\n-        let adapters_dir = tmp.path().join(\"adapters\");\n-        let src = tmp.path().join(\"srcs\").join(\"mock-boot\");\n-        std::fs::create_dir_all(&src).unwrap();\n+        std::fs::create_dir_all(source).unwrap();\n         std::fs::write(\n-            src.join(\"manifest.toml\"),\n+            source.join(\"manifest.toml\"),\n             format!(\n-                \"[adapter]\\nname = \\\"mock-boot\\\"\\nkind = \\\"shell\\\"\\nversion = \\\"1\\\"\\n\\\n+                \"[adapter]\\nname = \\\"{name}\\\"\\nkind = \\\"shell\\\"\\nversion = \\\"1\\\"\\n\\\n                  min_spt_core_version = \\\"0\\\"\\n\\n[shell]\\nspawn = '{noop}'\\n\\\n-                 persistent = true\\n\"\n+                 persistent = {persistent}\\n\"\n             ),\n         )\n         .unwrap();\n-        spt_runtime::registry::register(&adapters_dir, &src, 1).unwrap();\n-        let registered = spt_runtime::registry::registered(&adapters_dir);\n+        spt_runtime::registry::register(adapters_dir, source, 1).unwrap();\n+    }\n \n-        // Every instance is recorded ONLINE over a dead pid — the restart shape —\n-        // and they differ ONLY in the gate under test.\n-        let mk = |owner: &str, launched_ms: u64| -> String {\n-            let id = spawn_record(owlery, owner, \"mock-boot\", None).unwrap();\n-            let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, &id);\n-            let mut i = shellinfo::read_shell_info(&perch).unwrap();\n-            i.status = SHELL_STATUS_ONLINE.to_string();\n-            shellinfo::write_shell_info(&perch, &i).unwrap();\n-            std::fs::write(perch.join(shellinfo::SHELL_PID_FILE), \"2000000000\").unwrap();\n-            let launch = shellinfo::ShellLaunch {\n-                pid_started_at: Some(1),\n-                launched_ms,\n-            };\n-            std::fs::write(\n-                perch.join(shellinfo::SHELL_LAUNCH_FILE),\n-                serde_json::to_string(&launch).unwrap(),\n-            )\n-            .unwrap();\n-            id\n-        };\n-        let stranded = mk(\"ling\", boot.saturating_sub(600_000)); // 10 min before boot\n-        let force_killed = mk(\"ling\", boot + 120_000); // 2 min AFTER boot\n-        let orphan_owner = mk(\"gone\", boot.saturating_sub(600_000));\n+    /// Restart-shaped record, with NO suspend edge to pre-heal its status.\n+    fn stranded_shell(owlery: &Path, owner: &str, adapter: &str, launched_ms: Option<u64>) -> String {\n+        let id = spawn_record(owlery, owner, adapter, None).unwrap();\n+        let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, &id);\n+        let mut info = shellinfo::read_shell_info(&perch).unwrap();\n+        info.status = SHELL_STATUS_ONLINE.to_string();\n+        shellinfo::write_shell_info(&perch, &info).unwrap();\n+        std::fs::write(perch.join(shellinfo::SHELL_PID_FILE), \"2000000000\").unwrap();\n+        if let Some(launched_ms) = launched_ms {\n+            shellinfo::record_shell_launch(&perch, 2_000_000_000, launched_ms);\n+        }\n+        id\n+    }\n+\n+    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n+    // [unit->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\n+    #[test]\n+    fn boot_restores_all_down_persistent_instances_without_launch_age_gates() {\n+        let tmp = tempfile::tempdir().unwrap();\n+        let owlery = tmp.path();\n+        let boot = spt_store::proc::boot_instant_ms().expect(\"boot instant\");\n+        let adapters_dir = tmp.path().join(\"adapters\");\n+        register_restore_adapter(&adapters_dir, &tmp.path().join(\"keeper\"), \"keeper\", true);\n+        register_restore_adapter(&adapters_dir, &tmp.path().join(\"ordinary\"), \"ordinary\", false);\n+        let registered = spt_runtime::registry::registered(&adapters_dir);\n+        seed_owner(owlery, \"ling\");\n+        let gone = owlery.join(\"gone\");\n+        std::fs::create_dir_all(&gone).unwrap();\n+        std::fs::write(gone.join(\"info.json\"), \"{\\\"status\\\":\\\"offline\\\"}\").unwrap();\n \n-        let restored = restore_persistent_shells_at_boot(owlery, &registered, &adapters_dir);\n+        let before = stranded_shell(owlery, \"ling\", \"keeper\", Some(boot.saturating_sub(600_000)));\n+        let after = stranded_shell(owlery, \"ling\", \"keeper\", Some(boot + 120_000));\n+        let unstamped = stranded_shell(owlery, \"ling\", \"keeper\", None);\n+        stranded_shell(owlery, \"gone\", \"keeper\", None);\n+        stranded_shell(owlery, \"ling\", \"ordinary\", Some(boot + 120_000));\n ","truncated":true,"truncatedBy":"lines","totalLines":789,"totalBytes":16726,"outputLines":304,"outputBytes":16726,"lastLinePartial":false,"firstLineExceedsLimit":false},"meta":{"source":{"type":"internal","value":"artifact://112"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":1088,"totalBytes":16726,"outputLines":304,"outputBytes":16726,"maxBytes":155648,"shownRange":{"start":300,"end":603},"nextOffset":604}}},"isError":false,"timestamp":1789001136975}}
{"type":"message","id":"a9a384e5","parentId":"ca63b887","timestamp":"2026-09-10T00:45:38.097Z","message":{"role":"toolResult","toolCallId":"call_8hv8qzFRDNvgINDHIrh0G52O|fc_077f3569365e313d016aa1fdb0c09887d1bad4b5b5fc3d8749","toolName":"grep","content":[{"type":"text","text":"# .worktrees/287-shellrestore/crates/spt-daemon/src/\n## drivehub.rs\n 294|/// **The recorded live pid arbitrates, on BOTH sides**\n*295|/// ([`crate::shellhost::live_launch_winner`]). Whoever is second sees the\n 296|/// other's `shell.pid` probing alive, stands down, and ADOPTS that pid as the\n 297|/// launch's outcome — a success, not a fault: the caller asked for a running\n 298|/// binary and there is one. The daemon-side loser reports\n## linkhost.rs\n 737|    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, &id);\n*738|    if let Some(pid) = shellhost::live_launch_winner(&perch) {\n 739|        spt_proto::emit_line_err!(\"SHELL_LAUNCH_STOOD_DOWN:{id}: already launched pid={pid} (online at bind)\");\n 740|        return Ok(pid);\n 741|    }\n## shellhost.rs\n 319|// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n*320|pub fn live_launch_winner(shell_perch: &Path) -> Option<u32> {\n 321|    shellinfo::read_shell_pid(shell_perch)\n 322|        .filter(|&p| p != 0 && !shellinfo::shell_pid_provably_dead(shell_perch))\n 323|}\n...\n 430|            || {\n*431|                live_launch_winner(&spt_store::perch::resolve_shell_perch_path_in(\n 432|                    owlery, owner, shell_id,\n 433|                ))\n 434|            },\n...\n 1210|\n*1211|        assert_eq!(live_launch_winner(perch), None, \"no pid file ⇒ no winner\");\n 1212|\n 1213|        std::fs::write(perch.join(SHELL_PID_FILE), \"0\").unwrap();\n*1214|        assert_eq!(live_launch_winner(perch), None, \"pid 0 ⇒ no winner\");\n 1215|\n 1216|        std::fs::write(perch.join(SHELL_PID_FILE), \"not-a-pid\").unwrap();\n*1217|        assert_eq!(live_launch_winner(perch), None, \"unparseable ⇒ no winner\");\n 1218|\n 1219|        // This test process is, definitionally, alive.\n 1220|        let me = std::process::id();\n...\n 1222|        assert_eq!(\n*1223|            live_launch_winner(perch),\n 1224|            Some(me),\n 1225|            \"a recorded pid that probes alive IS the winner\"\n 1226|        );\n 1227|        shellinfo::record_shell_launch(perch, me, 1_000);\n*1228|        assert_eq!(live_launch_winner(perch), Some(me), \"same birth retains custody\");\n 1229|        let mut launch = shellinfo::read_shell_launch(perch).unwrap();\n 1230|        launch.pid_started_at = Some(\n 1231|            launch.pid_started_at.expect(\"native birth stamp on supported platforms\") ^ 1,\n...\n 1237|        .unwrap();\n*1238|        assert_eq!(live_launch_winner(perch), None, \"a recycled pid cannot block restore\");\n 1239|    }\n 1240|\n 1241|    // [unit->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE] a REAL dial at a socket name\n## shellwake.rs\n 615|// [impl->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\n*616|fn heal_stale_online_records(owlery: &Path, events: &mut dyn std::io::Write) {\n 617|    for owner in spt_store::perch::list_self_perch_ids(owlery) {\n 618|        for (shell_id, info) in shellinfo::list_shells(owlery, &owner) {\n 619|            if info.status != SHELL_STATUS_ONLINE {\n...\n 662|// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n*663|fn watcher_eligible(\n 664|    perch: &Path,\n 665|    info: &shellinfo::ShellInfo,\n 666|    persistent: bool,\n...\n 671|    }\n*672|    if crate::shellhost::live_launch_winner(perch).is_some() {\n 673|        return false; // binding: the shell binary and watcher must not overlap\n 674|    }\n 675|    if persistent || !shellinfo::shell_pid_provably_dead(perch) {\n...\n 733|// [impl->REQ-SHELL-OWNER-ONLINE-RESTORE]\n*734|fn restore_persistent_shells_of_owner(\n 735|    owlery: &Path,\n 736|    registered: &[(AdapterRecord, Manifest)],\n 737|    adapters_dir: &Path,\n...\n 754|        if shellinfo::is_shell_online(&perch, &info)\n*755|            || crate::shellhost::live_launch_winner(&perch).is_some()\n 756|        {\n 757|            continue; // already online or binding in fact\n 758|        }\n...\n 788|/// endpoint that comes online after this sweep is covered by\n*789|/// [`restore_persistent_shells_on_owner_online`] (releases#228), running the same\n*790|/// [`restore_persistent_shells_of_owner`] body -- but it is the one that needs no\n 791|/// trigger at all, since a node restart is what strands the shells to begin with.\n 792|// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n 793|// [impl->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\n*794|pub fn restore_persistent_shells_at_boot(\n 795|    owlery: &Path,\n 796|    registered: &[(AdapterRecord, Manifest)],\n 797|    adapters_dir: &Path,\n...\n 800|    for owner in spt_store::perch::list_self_perch_ids(owlery) {\n*801|        restore_persistent_shells_of_owner(\n 802|            owlery,\n 803|            registered,\n 804|            adapters_dir,\n...\n 873|// [impl->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\n*874|pub fn restore_persistent_shells_on_owner_online(\n 875|    owlery: &Path,\n 876|    registered: &[(AdapterRecord, Manifest)],\n 877|    adapters_dir: &Path,\n...\n 886|    for owner in owners {\n*887|        restore_persistent_shells_of_owner(\n 888|            owlery,\n 889|            registered,\n 890|            adapters_dir,\n...\n 910|    // on the stale value it just corrected.\n*911|    heal_stale_online_records(owlery, &mut std::io::stderr().lock());\n 912|\n 913|    // Stop side first: watchers whose instance is gone or no longer eligible.\n 914|    let live: Vec<(String, String)> = {\n...\n 942|            };\n*943|            if !watcher_eligible(&perch, &info, shell.persistent, boot_ms) {\n 944|                continue;\n 945|            }\n 946|            let Some(wake_command) = shell.wake_command.as_deref() else {\n...\n 1017|            owner_edge.seed(&owlery);\n*1018|            restore_persistent_shells_at_boot(&owlery, &registered, &adapters_dir);\n 1019|        }\n 1020|        while !stop.load(Ordering::SeqCst) {\n 1021|            let owlery = spt_store::perch::owlery_dir();\n...\n 1026|            // would otherwise arm a watcher for the shell we are about to launch.\n*1027|            restore_persistent_shells_on_owner_online(\n 1028|                &owlery,\n 1029|                &registered,\n 1030|                &adapters_dir,\n...\n 1120|            let mut events = Vec::new();\n*1121|            heal_stale_online_records(&owlery, &mut events);\n 1122|            let output = String::from_utf8(events).unwrap();\n 1123|            let mut reported: Vec<_> = output.lines().collect();\n 1124|            reported.sort();\n...\n 1154|            let mut repeated = Vec::new();\n*1155|            heal_stale_online_records(&owlery, &mut repeated);\n 1156|            assert!(repeated.is_empty(), \"already-healed records emit no duplicate events\");\n 1157|            assert_eq!(\n 1158|                (mtime(&stale_perch), mtime(&live_perch)),\n...\n 1255|\n*1256|        let mut restored = restore_persistent_shells_at_boot(owlery, &registered, &adapters_dir);\n 1257|        restored.sort();\n 1258|        let mut expected = vec![\n 1259|            format!(\"ling/{before}\"),\n...\n 1286|        edge.seed(owlery);\n*1287|        assert!(restore_persistent_shells_at_boot(owlery, &registered, &adapters_dir).is_empty());\n 1288|\n 1289|        let set = Arc::new(WakeSet::new());\n*1290|        assert!(restore_persistent_shells_on_owner_online(\n 1291|            owlery, &registered, &adapters_dir, &mut edge, &set,\n 1292|        )\n 1293|        .is_empty());"}],"details":{"scopePath":".worktrees/287-shellrestore/crates","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\crates","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":30,"fileCount":4,"files":[".worktrees/287-shellrestore/crates/spt-daemon/src/drivehub.rs",".worktrees/287-shellrestore/crates/spt-daemon/src/linkhost.rs",".worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs",".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs"],"fileMatches":[{"path":".worktrees/287-shellrestore/crates/spt-daemon/src/drivehub.rs","count":1},{"path":".worktrees/287-shellrestore/crates/spt-daemon/src/linkhost.rs","count":1},{"path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs","count":8},{"path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs","count":20}],"truncated":true,"perFileLimitReached":20,"displayContent":"# .worktrees/287-shellrestore/crates/spt-daemon/src/\n## drivehub.rs\n 294│/// **The recorded live pid arbitrates, on BOTH sides**\n*295│/// ([`crate::shellhost::live_launch_winner`]). Whoever is second sees the\n 296│/// other's `shell.pid` probing alive, stands down, and ADOPTS that pid as the\n 297│/// launch's outcome — a success, not a fault: the caller asked for a running\n 298│/// binary and there is one. The daemon-side loser reports\n## linkhost.rs\n 737│    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, &id);\n*738│    if let Some(pid) = shellhost::live_launch_winner(&perch) {\n 739│        spt_proto::emit_line_err!(\"SHELL_LAUNCH_STOOD_DOWN:{id}: already launched pid={pid} (online at bind)\");\n 740│        return Ok(pid);\n 741│    }\n## shellhost.rs\n  319│// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n *320│pub fn live_launch_winner(shell_perch: &Path) -> Option<u32> {\n  321│    shellinfo::read_shell_pid(shell_perch)\n  322│        .filter(|&p| p != 0 && !shellinfo::shell_pid_provably_dead(shell_perch))\n  323│}\n     │...\n  430│            || {\n *431│                live_launch_winner(&spt_store::perch::resolve_shell_perch_path_in(\n  432│                    owlery, owner, shell_id,\n  433│                ))\n  434│            },\n     │...\n 1210│\n*1211│        assert_eq!(live_launch_winner(perch), None, \"no pid file ⇒ no winner\");\n 1212│\n 1213│        std::fs::write(perch.join(SHELL_PID_FILE), \"0\").unwrap();\n*1214│        assert_eq!(live_launch_winner(perch), None, \"pid 0 ⇒ no winner\");\n 1215│\n 1216│        std::fs::write(perch.join(SHELL_PID_FILE), \"not-a-pid\").unwrap();\n*1217│        assert_eq!(live_launch_winner(perch), None, \"unparseable ⇒ no winner\");\n 1218│\n 1219│        // This test process is, definitionally, alive.\n 1220│        let me = std::process::id();\n     │...\n 1222│        assert_eq!(\n*1223│            live_launch_winner(perch),\n 1224│            Some(me),\n 1225│            \"a recorded pid that probes alive IS the winner\"\n 1226│        );\n 1227│        shellinfo::record_shell_launch(perch, me, 1_000);\n*1228│        assert_eq!(live_launch_winner(perch), Some(me), \"same birth retains custody\");\n 1229│        let mut launch = shellinfo::read_shell_launch(perch).unwrap();\n 1230│        launch.pid_started_at = Some(\n 1231│            launch.pid_started_at.expect(\"native birth stamp on supported platforms\") ^ 1,\n     │...\n 1237│        .unwrap();\n*1238│        assert_eq!(live_launch_winner(perch), None, \"a recycled pid cannot block restore\");\n 1239│    }\n 1240│\n 1241│    // [unit->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE] a REAL dial at a socket name\n## shellwake.rs\n  615│// [impl->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\n *616│fn heal_stale_online_records(owlery: &Path, events: &mut dyn std::io::Write) {\n  617│    for owner in spt_store::perch::list_self_perch_ids(owlery) {\n  618│        for (shell_id, info) in shellinfo::list_shells(owlery, &owner) {\n  619│            if info.status != SHELL_STATUS_ONLINE {\n     │...\n  662│// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n *663│fn watcher_eligible(\n  664│    perch: &Path,\n  665│    info: &shellinfo::ShellInfo,\n  666│    persistent: bool,\n     │...\n  671│    }\n *672│    if crate::shellhost::live_launch_winner(perch).is_some() {\n  673│        return false; // binding: the shell binary and watcher must not overlap\n  674│    }\n  675│    if persistent || !shellinfo::shell_pid_provably_dead(perch) {\n     │...\n  733│// [impl->REQ-SHELL-OWNER-ONLINE-RESTORE]\n *734│fn restore_persistent_shells_of_owner(\n  735│    owlery: &Path,\n  736│    registered: &[(AdapterRecord, Manifest)],\n  737│    adapters_dir: &Path,\n     │...\n  754│        if shellinfo::is_shell_online(&perch, &info)\n *755│            || crate::shellhost::live_launch_winner(&perch).is_some()\n  756│        {\n  757│            continue; // already online or binding in fact\n  758│        }\n     │...\n  788│/// endpoint that comes online after this sweep is covered by\n *789│/// [`restore_persistent_shells_on_owner_online`] (releases#228), running the same\n *790│/// [`restore_persistent_shells_of_owner`] body -- but it is the one that needs no\n  791│/// trigger at all, since a node restart is what strands the shells to begin with.\n  792│// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n  793│// [impl->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\n *794│pub fn restore_persistent_shells_at_boot(\n  795│    owlery: &Path,\n  796│    registered: &[(AdapterRecord, Manifest)],\n  797│    adapters_dir: &Path,\n     │...\n  800│    for owner in spt_store::perch::list_self_perch_ids(owlery) {\n *801│        restore_persistent_shells_of_owner(\n  802│            owlery,\n  803│            registered,\n  804│            adapters_dir,\n     │...\n  873│// [impl->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\n *874│pub fn restore_persistent_shells_on_owner_online(\n  875│    owlery: &Path,\n  876│    registered: &[(AdapterRecord, Manifest)],\n  877│    adapters_dir: &Path,\n     │...\n  886│    for owner in owners {\n *887│        restore_persistent_shells_of_owner(\n  888│            owlery,\n  889│            registered,\n  890│            adapters_dir,\n     │...\n  910│    // on the stale value it just corrected.\n *911│    heal_stale_online_records(owlery, &mut std::io::stderr().lock());\n  912│\n  913│    // Stop side first: watchers whose instance is gone or no longer eligible.\n  914│    let live: Vec<(String, String)> = {\n     │...\n  942│            };\n *943│            if !watcher_eligible(&perch, &info, shell.persistent, boot_ms) {\n  944│                continue;\n  945│            }\n  946│            let Some(wake_command) = shell.wake_command.as_deref() else {\n     │...\n 1017│            owner_edge.seed(&owlery);\n*1018│            restore_persistent_shells_at_boot(&owlery, &registered, &adapters_dir);\n 1019│        }\n 1020│        while !stop.load(Ordering::SeqCst) {\n 1021│            let owlery = spt_store::perch::owlery_dir();\n     │...\n 1026│            // would otherwise arm a watcher for the shell we are about to launch.\n*1027│            restore_persistent_shells_on_owner_online(\n 1028│                &owlery,\n 1029│                &registered,\n 1030│                &adapters_dir,\n     │...\n 1120│            let mut events = Vec::new();\n*1121│            heal_stale_online_records(&owlery, &mut events);\n 1122│            let output = String::from_utf8(events).unwrap();\n 1123│            let mut reported: Vec<_> = output.lines().collect();\n 1124│            reported.sort();\n     │...\n 1154│            let mut repeated = Vec::new();\n*1155│            heal_stale_online_records(&owlery, &mut repeated);\n 1156│            assert!(repeated.is_empty(), \"already-healed records emit no duplicate events\");\n 1157│            assert_eq!(\n 1158│                (mtime(&stale_perch), mtime(&live_perch)),\n     │...\n 1255│\n*1256│        let mut restored = restore_persistent_shells_at_boot(owlery, &registered, &adapters_dir);\n 1257│        restored.sort();\n 1258│        let mut expected = vec![\n 1259│            format!(\"ling/{before}\"),\n     │...\n 1286│        edge.seed(owlery);\n*1287│        assert!(restore_persistent_shells_at_boot(owlery, &registered, &adapters_dir).is_empty());\n 1288│\n 1289│        let set = Arc::new(WakeSet::new());\n*1290│        assert!(restore_persistent_shells_on_owner_online(\n 1291│            owlery, &registered, &adapters_dir, &mut edge, &set,\n 1292│        )\n 1293│        .is_empty());"},"isError":false,"timestamp":1789001138096}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_LhwwqTqtbNXGA5WfqmpAvEyX|fc_077f3569365e313d016aa1fdbfc90c87d185cc29b3b7e71d6a","toolName":"read","startedAt":"2026-09-10T00:45:51.863Z","args":{"path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs:1-510"},"intent":"Reading routed launch mechanics"},"id":"d217420f","parentId":"a9a384e5","timestamp":"2026-09-10T00:45:51.863Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_r5yTrWcHNAMryjcF1FXmyKA2|fc_077f3569365e313d016aa1fdbfc91c87d19d0b7c54b8677cbe","toolName":"read","startedAt":"2026-09-10T00:45:51.863Z","args":{"path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs:300-590"},"intent":"Reading wake resolution paths"},"id":"c23675bc","parentId":"d217420f","timestamp":"2026-09-10T00:45:51.863Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_86q7GdHeGBRmNxe4753yUR3N|fc_077f3569365e313d016aa1fdbfc92487d18187cf49e1242bd2","toolName":"read","startedAt":"2026-09-10T00:45:51.863Z","args":{"path":".worktrees/287-shellrestore/crates/spt-daemon/src/linkhost.rs:645-790"},"intent":"Reading daemon launch consumer"},"id":"e72759f9","parentId":"c23675bc","timestamp":"2026-09-10T00:45:51.863Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_nKdV7gj0NifQPSIMOkbPE8HP|fc_077f3569365e313d016aa1fdbfc92c87d19d1b7edec2d8c706","toolName":"grep","startedAt":"2026-09-10T00:45:51.863Z","args":{"path":".worktrees/287-shellrestore/crates/spt-store/src;.worktrees/287-shellrestore/crates/spt-proto/src;.worktrees/287-shellrestore/crates/spt-daemon/src"},"intent":"Locating liveness probe semantics"},"id":"e1e98677","parentId":"e72759f9","timestamp":"2026-09-10T00:45:51.863Z"}
{"type":"message","id":"e17e6928","parentId":"e1e98677","timestamp":"2026-09-10T00:45:51.863Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":"**Reading lines 300–590**"},{"type":"toolCall","id":"call_LhwwqTqtbNXGA5WfqmpAvEyX|fc_077f3569365e313d016aa1fdbfc90c87d185cc29b3b7e71d6a","name":"read","arguments":{"i":"Reading routed launch mechanics","path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs:1-510"},"intent":"Reading routed launch mechanics"},{"type":"toolCall","id":"call_r5yTrWcHNAMryjcF1FXmyKA2|fc_077f3569365e313d016aa1fdbfc91c87d19d0b7c54b8677cbe","name":"read","arguments":{"i":"Reading wake resolution paths","path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs:300-590"},"intent":"Reading wake resolution paths"},{"type":"toolCall","id":"call_86q7GdHeGBRmNxe4753yUR3N|fc_077f3569365e313d016aa1fdbfc92487d18187cf49e1242bd2","name":"read","arguments":{"i":"Reading daemon launch consumer","path":".worktrees/287-shellrestore/crates/spt-daemon/src/linkhost.rs:645-790"},"intent":"Reading daemon launch consumer"},{"type":"toolCall","id":"call_nKdV7gj0NifQPSIMOkbPE8HP|fc_077f3569365e313d016aa1fdbfc92c87d19d1b7edec2d8c706","name":"grep","arguments":{"i":"Locating liveness probe semantics","pattern":"pub fn relay_liveness|fn relay_liveness|enum RelayLiveness|fn process_started_at|fn is_process_alive|SHELL_RECORD_HEALED","path":".worktrees/287-shellrestore/crates/spt-store/src;.worktrees/287-shellrestore/crates/spt-proto/src;.worktrees/287-shellrestore/crates/spt-daemon/src","case":true,"gitignore":false,"skip":null},"intent":"Locating liveness probe semantics"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":19617,"output":323,"cacheRead":30464,"cacheWrite":0,"totalTokens":50404,"reasoningTokens":68,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1789001138150,"responseId":"resp_077f3569365e313d016aa1fdb260cc87d1a674b123fbdd0862","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_077f3569365e313d016aa1fdb6177087d1b5e3da484bde1adb","type":"reasoning","content":[],"encrypted_content":"gAAAAABqof2_D3ctU52aPSxA12dMI2il0U9LkDiS-PQqwkVlZFF0eXo9oaiBD2iM41cFS7qXNr0kcqjQAjgvP9h9T8iPGCTWx-oTWVQ-EuhJdaoUAwC_-6XZViGqhGm7uRNq5B3lpwGObxUDApYPHuAvZII2orcMa8oWUNkj3FWwHxacmETP4hJ5WVUNez0SO_lbMA_UW1A_DwKmTJSfe-Zqo8Bn6WCavw305jvNoNTcb7FZ0Y6YV1mezFVTZQaMwdHbwITCQFJKMRy8UrEqCrD7SSCozVoUC7E9anrPzXzCrQA7CoB0pitMNHvQtqXd7PHLWfvlhNc4aiiVY73zyf4j5m7d2PWq712gG7JXR-9JIOgU41hdwumUPQtpiFKgsgxVinkI5ue0kaxl4eCewjdnsWnfCAjRKhavOHMhEnPbc29Z7AialUwnHs7j2fIELZ43_BvVzdx-C8WgyQnZqlVdFbjDjErgA4eGWmjavyd0elUEsAQ3MZJZF15MajIB7iP55kdPnTN5S4NZCGKNH8nNcPlK0YClGi-Awhf7pEAHt7XlBtb3YkKFk5ycNIGRRnw8hHt9oFQabqnvZI_hvFpX6HPxthYEfFEsScB7w74EFG19-bmKO6tk8ExX26EbQykPzTOJe4cUlB794aDpVDRKokfX5f-bhY73JSihXQfSqfG1qtvKcLdm3XOx_Tf2NDyBtNbkcjW5EQ8SLZOz19EaREepov_zqSfKI4uCZYnOdQw23CuFOCDyPavakdC4TPPEjJAfCwMXyrI8Ub4DaXhlc-m4OWS3dgfWcrdMtxyNLgOoztupR9FRI1rHiLMzU14zqYYisgEafkK_j6Ed4F0QZieSWSP37tr9kP2Zbyw68hEoosBGQOJGZnthwuGq6yE7JbeeJPEf3pt-FBEnsT-kM5sYgsMTyJ1ZjDA4XCJZBzsOJli7HcDTBGowQGxnYKgDIwlxi4sNKcvNVwEfH_YSzTDBGHX_RXytDPpbyR7M9uIoRqCFuHa75q1NlWQZ7zsrjN0sqT5pO-AAfmEaLiV4tR6I1r36c3OVglzu-gC4dOiWP_qaPDp6pG2-ZKPlEg85DAw9w5AYKCa6cDk70DkWnA9jAUe470N6PmekBHI1G3ZHXM7laJ_I9yTrVRGKNpLYfF6siBRVNmxm6sP4cOGuCmZV3huMcBHbULKWIDPdwqvAJkTmv0JMKwnpqwMXAvmJIvVKvPszKEhzTBGMiiwHBDDx8g7aZwXBCFDoKLE3BblH3EuaZsCITYRVXw3WTN3-9J_YErUOJOJgZysoMjMJPiLMxzQsK4iDASwZjOJMJJblpThgR2TWGPqRncyWXEPFvlRsoT2akBniNcWovQQofS7a7mgXr_m1TY4g5gRhSiwXJN-9jhJ7_ZlJ6nlo9-lprTU0874wYWx64N9RY0FQ3ywSh1gySrNHfISGP2pfYpj04xX8F3tKw-PmroPF2RoxdvBkV5Vzqkwqz-NZqwsyo_Tu7OD3y0ed542vOZjwEkltfHvQK1IKkVESOdAhNyM2DY49QG5UE4m03IEhcqtGMqv5NonNIa9zdb5KcBwJWBSQil962toXvu_gPuW_3AA_7eAXeJtAhsBCFCwLatd45Mr5JCxMSxHUt51njPrWDq3irVkdf5YRAHv5rBAKxp88Q7p8m2HS0jBrUYZ7mW3WD5zxqYJSZcVwP_JgiAGIWmhXpC-dx68=","internal_chat_message_metadata_passthrough":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"},"summary":[{"type":"summary_text","text":"**Reading lines 300–590**"}],"metadata":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"}},{"id":"fc_077f3569365e313d016aa1fdbfc90c87d185cc29b3b7e71d6a","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading routed launch mechanics\",\"path\":\".worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs:1-510\"}","call_id":"call_LhwwqTqtbNXGA5WfqmpAvEyX","internal_chat_message_metadata_passthrough":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"},"name":"read","metadata":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"}},{"id":"fc_077f3569365e313d016aa1fdbfc91c87d19d0b7c54b8677cbe","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading wake resolution paths\",\"path\":\".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs:300-590\"}","call_id":"call_r5yTrWcHNAMryjcF1FXmyKA2","internal_chat_message_metadata_passthrough":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"},"name":"read","metadata":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"}},{"id":"fc_077f3569365e313d016aa1fdbfc92487d18187cf49e1242bd2","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading daemon launch consumer\",\"path\":\".worktrees/287-shellrestore/crates/spt-daemon/src/linkhost.rs:645-790\"}","call_id":"call_86q7GdHeGBRmNxe4753yUR3N","internal_chat_message_metadata_passthrough":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"},"name":"read","metadata":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"}},{"id":"fc_077f3569365e313d016aa1fdbfc92c87d19d1b7edec2d8c706","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating liveness probe semantics\",\"pattern\":\"pub fn relay_liveness|fn relay_liveness|enum RelayLiveness|fn process_started_at|fn is_process_alive|SHELL_RECORD_HEALED\",\"path\":\".worktrees/287-shellrestore/crates/spt-store/src;.worktrees/287-shellrestore/crates/spt-proto/src;.worktrees/287-shellrestore/crates/spt-daemon/src\",\"case\":true,\"gitignore\":false,\"skip\":null}","call_id":"call_nKdV7gj0NifQPSIMOkbPE8HP","internal_chat_message_metadata_passthrough":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"},"name":"grep","metadata":{"turn_id":"25afac0d-d353-401d-a399-10a9a855f17f"}}]},"duration":13711.025299999863,"ttft":3968.3830999997444,"completedAt":1789001151862,"contextSnapshot":{"promptTokens":50081,"nonMessageTokens":13209,"compactionEpoch":0}}}
{"type":"message","id":"6fdf3f2a","parentId":"e17e6928","timestamp":"2026-09-10T00:45:51.868Z","message":{"role":"toolResult","toolCallId":"call_r5yTrWcHNAMryjcF1FXmyKA2|fc_077f3569365e313d016aa1fdbfc91c87d19d0b7c54b8677cbe","toolName":"read","content":[{"type":"text","text":"pub fn watcher_run(\n…\n) {\n…\n    while !stop.load(Ordering::SeqCst) {\n…\n        let mut child = match child {\n…\n            Err(e) => {\n                // An unspawnable waker is a permanent condition — latch now.\n                spt_proto::emit_line_err!(\"WAKER_SPAWN_FAIL:{owner}/{shell_id}: {e} (giving up)\");\n                let _ = std::fs::write(perch.join(WAKER_GAVE_UP_FILE), b\"\");\n                return;\n            }\n        };\n        // The birth stamp rides WITH the pid, in one write — the kill path\n        // authenticates the pair, and a stamp parked separately would leave a\n        // window that reads Held. [impl->REQ-SHELL-KILL-AUTHENTICATED]\n        record_waker_launch(&perch, child.id());\n        let status = child.wait();\n        let _ = std::fs::remove_file(perch.join(WAKER_PID_FILE));\n        if stop.load(Ordering::SeqCst) {\n            return; // stopped from outside — the kill is not a crash\n        }\n        match status.ok().and_then(|s| s.code()) {\n            Some(WAKE_OPCODE) => {\n                match resolve() {\n                    Ok(what) => spt_proto::emit_line_err!(\"SHELL_WAKE:{owner}/{shell_id}: {what}\"),\n                    Err(e) => spt_proto::emit_line_err!(\"SHELL_WAKE_FAIL:{owner}/{shell_id}: {e}\"),\n                }\n                return;\n            }\n            _ => {\n                failures += 1;\n                if failures >= params.give_up_after {\n                    spt_proto::emit_line_err!(\n                        \"WAKER_GAVE_UP:{owner}/{shell_id}: {failures} consecutive crash-exits \\\n                         (cleared by the next relink/launch)\"\n                    );\n                    let _ = std::fs::write(perch.join(WAKER_GAVE_UP_FILE), b\"\");\n                    return;\n                }\n                // Backoff in small slices so a stop request lands promptly.\n                let mut left = backoff_ms(params, failures);\n                while left > 0 && !stop.load(Ordering::SeqCst) {\n                    let step = left.min(50);\n                    std::thread::sleep(Duration::from_millis(step));\n                    left -= step;\n                }\n            }\n        }\n    }\n}\n\n/// The **state-keyed wake resolution** (M5-D4c — CONTEXT §Shell sleep/wake):\n/// key on the owner's local rest record, then bring the shell online —\n///\n/// - owner **dormant** (resting warm, still running) ⇒ touch nothing on the\n///   endpoint, just relaunch the shell binary;\n/// - owner **suspended** ⇒ revive the owner first\n///   ([`crate::resting::daemon_rest_event`] `Wake` — its cascade relaunches\n///   the owner's *persistent* shells, so this fn launches only if the\n///   cascade did not already), then the shell;\n/// - owner **active** (or recordless — a pre-resting/interim perch) ⇒ just\n///   relaunch the shell;\n/// - **no local instance of the owner** ⇒ refuse, naming the deferral: the\n///   `shell_wake_spawn_anywhere` fresh-spawn branch rides\n///   instantiate-anywhere (the D1c grant shape is its seam), and the\n///   *active-elsewhere cross-node attach* arm upgrades with presence/MRA\n///   (D6) + cross-node link (D8c) — D4 resolves against the local node.\n///\n/// The shell's bind onlines the perch (the D3b contract) — this fn never\n/// flips status.\n// [impl->REQ-SHELL-2]\npub fn resolve_wake(\n    owlery: &Path,\n    owner: &str,\n    shell_id: &str,\n    adapter_name: &str,\n    shell: &Shell,\n) -> Result<String, String> {\n    // Shells hang off flat-Self perches only (the D3a layout decision).\n    let owner_perch = owlery.join(owner);\n    if !owner_perch.join(\"info.json\").exists() {\n        // No local instance — consult presence (M5-D6b, REQ-PRES-1): a\n        // routable instance of the owner on ANOTHER node gets the wake\n        // FORWARDED there through the D5b remote rest op (the target node's\n        // own cascade relaunches ITS persistent shells). This local shell\n        // stays offline — the cross-node shell *link* is D8c; say so.\n        // [impl->REQ-PRES-1]\n        if let Some(node) = remote_owner_node(owner) {\n            return match forward_wake(owner, &node) {\n                Ok(outcome) => Ok(format!(\n                    \"WAKE_FORWARDED:{owner}@{node}: {outcome}; this shell stays offline \\\n                     here (cross-node shell link lands at D8c)\"\n                )),\n                Err(e) => Err(format!(\"WAKE_FORWARD_FAIL:{owner}@{node}: {e}\")),\n            };\n        }\n        return Err(format!(\n            \"WAKE_NO_REACHABLE_INSTANCE:{owner}: no instance of the owner on any \\\n             reachable node — a fresh-spawn-to-wake rides the deferred \\\n             instantiate-anywhere capability (the shell_wake_spawn_anywhere grant \\\n             shape is its seam)\"\n        ));\n    }\n    let state = crate::resting::read_rest(&owner_perch).map(|r| r.state);\n    let mut did = String::new();\n    if state == Some(crate::resting::RestState::Suspended) {\n        match crate::resting::daemon_rest_event(owner, crate::resting::RestEvent::Wake, None) {\n            Ok(_) => did.push_str(\"revived owner; \"),\n            Err(e) => return Err(format!(\"revive owner: {e}\")),\n        }\n    } else if state == Some(crate::resting::RestState::Dormant) {\n        did.push_str(\"owner dormant (left in place); \");\n    }\n    // Never double-launch a live binary: the revive's wake cascade may have\n    // relaunched this (persistent) instance already — and two reconcilers\n    // (a fresh daemon adopting + a stale watcher firing) may race a wake.\n    // A parked pid that probes alive means somebody won; stand down.\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n    let live_pid = std::fs::read_to_string(perch.join(crate::shellhost::SHELL_PID_FILE))\n        .ok()\n        .and_then(|s| s.trim().parse::<u32>().ok())\n        .filter(|&p| p != 0 && spt_store::proc::is_process_alive(p));\n    if let Some(p) = live_pid {\n        return Ok(format!(\"{did}already relaunched pid={p} (online at bind)\"));\n    }\n    // [impl->REQ-INSTALL-11] the wake-triggered relaunch resolves against the\n    // adapter's install dir, like every other launch of the same instance.\n    let install_dir =\n        crate::shellhost::shell_install_dir(&spt_store::perch::adapters_dir(), adapter_name);\n    let pid = crate::shellhost::launch_shell(\n        owlery,\n        owner,\n        shell_id,\n        adapter_name,\n        install_dir.as_deref(),\n        shell,\n    )?;\n    Ok(format!(\"{did}relaunched pid={pid} (online at bind)\"))\n}\n\n/// The node (hex) holding a routable instance of `owner` per the gossiped\n/// registry snapshots, preferring the most-recently-active row when several\n/// nodes hold one (the presence datum as the tiebreak; absent data ranks\n/// last). `None` = the owner exists nowhere reachable.\n// [impl->REQ-PRES-1]\nfn remote_owner_node(owner: &str) -> Option<String> {\n    let regs = crate::presence::load_registry_snapshots(&crate::presence::registry_snapshot_dir());\n    let local = crate::presence::local_node_hex();\n    let mut best: Option<(u64, String)> = None;\n    for reg in regs.values() {\n        for row in reg.instances(owner) {\n            if row.node == local || !row.status.routable() {\n                continue;\n            }\n            let ts = row.last_active_ms.unwrap_or(0);\n            let better = match &best {\n                Some((cur, _)) => ts > *cur,\n                None => true,\n            };\n            if better {\n                best = Some((ts, row.node.clone()));\n            }\n        }\n    }\n    best.map(|(_, node)| node)\n}\n\n/// Forward a wake to `node` through the D5b remote rest op (the remote-drive\n/// trust class — the target's access gate decides). One dial, one request,\n/// one reply; a refusal (no reply) surfaces as an error.\n// [impl->REQ-PRES-1]\nfn forward_wake(owner: &str, node: &str) -> Result<String, String> {\n    use spt_net::net::endpoint::addr_for_node_hex;\n    let addr = addr_for_node_hex(node)\n        .and_then(|a| serde_json::to_value(a).ok())\n        .ok_or_else(|| format!(\"no dialable address for {node}\"))?;\n    let mut brain =\n        crate::brain::Brain::cold_start(&crate::endpoint::broker_socket_name(), now_ms())\n            .map_err(|e| format!(\"broker connect: {e}\"))?;\n    // A-4b (REQ-OPID-TRACING-RETRY): the wake-forward is a tracing-only rest op —\n    // if a broker restart dropped the conn/stream after the op journaled, re-mint a\n    // FRESH `wake` op from the SAME source and run ONCE more. The dial rides INSIDE\n    // `run` because that restart drops the CONN too, not just the stream — a retry\n    // against the dead conn would surface a different error, not self-heal. A 2nd\n    // no-longer-held collapses to the helper's F-1 public string.\n    // [impl->REQ-OPID-TRACING-RETRY]\n    let outcome = crate::effect::with_tracing_retry(\n        // The wake-forward rest op's seq source is a fresh `now_ms()` — its OWN\n        // minting source, distinct from shellchan's spool-row counter (doyle ruling:\n        // the tag names the seq source, not the subsystem), so it stamps `wake`.\n        || Ok(crate::effect::MintedOp::new(crate::effect::Minter::Wake, now_ms())),\n        |op| {\n            let conn = brain.net_dial(addr.clone(), None)?;\n            // Keep the tracing string's seq consistent with the (possibly re-minted) op.\n            let op_id = format!(\"{}:wakefwd:{}\", crate::presence::local_node_hex(), op.seq);\n            crate::resthost::request_rest(\n                &mut brain,\n                conn.conn_id,\n                owner,\n                spt_net::net::rest::REST_EVENT_WAKE,\n                &op_id,\n                op,\n            )\n        },\n    )\n    .map_err(|e| format!(\"rest op: {e}\"))?;\n    match outcome {\n        crate::resthost::RestRequestOutcome::Edge(d) => Ok(format!(\"woke ({d})\")),\n        crate::resthost::RestRequestOutcome::NoEdge => Ok(\"already awake\".to_string()),\n        crate::resthost::RestRequestOutcome::Failed(e) => Err(e),\n        crate::resthost::RestRequestOutcome::NoReply => {\n            Err(\"refused or dropped (no reply)\".to_string())\n        }\n    }\n}\n\n/// Resolve one adapter **option**'s `[shell]` section through the merged view\n/// (composite addressing): `adapter_option` is the stored `<adapter>[:profile]`\n/// string, split → parent lookup in `registered` → profile overlay (shipped or\n/// local). Returns an **owned** [`Shell`] — a profile produces a fresh merged\n/// manifest, not a borrow into the parent slice. A bare name resolves to the\n/// parent unmodified; a deregistered parent or a failed overlay yields `None`.\nfn shell_section_of(\n    registered: &[(AdapterRecord, Manifest)],\n    adapters_dir: &Path,\n    adapter_option: &str,\n) -> Option<Shell> {\n    spt_runtime::registry::resolve_option_in(registered, adapters_dir, adapter_option)\n        .ok()\n        .filter(|m| m.adapter.kind == AdapterKind::Shell)\n        .and_then(|m| m.shell)\n}\n\nfn now_ms() -> u64 {\n    crate::brain::now_ms()\n}\n\n/// Fill a `wake_command` template.\n///\n/// **Its catalog and the spawn template's share a vocabulary; they are not the\n/// same set, and never claiming they are is the point of saying so here.** The\n/// keys mean the same thing in both (`id`, `adapter_name`, `link_token`,\n/// `adapter_dir`), so an author learns one manifest language — but `perch_dir`\n/// is **spawn-only by design** (REQ-SHELL-PERCH-DIR: it resolves a *live* link's\n/// perch-relative transfer paths, and a waker has no live link). A comment\n/// asserting set-EQUALITY over two sets that are separately extended is the\n/// drift this project single-sources against; it had already been false since\n/// `perch_dir` landed.\n///\n/// The waker's token is minted **unparked**: it substitutes (template\n/// compatibility) without ever verifying against the perch, because an offline\n/// link has no live credential by design — the close retired it, and a waker\n/// wakes by *exit code*, not by driving the link.\n///\n/// `install_dir` resolves the program token and fills `{adapter_dir}`, exactly as\n/// on the spawn side ([`crate::shellhost::fill_spawn_command`]) — a released\n/// adapter whose spawn resolves and whose wake does not would go permanently\n/// unwakeable the moment it went offline.\nfn fill_wake_command(\n    shell_id: &str,\n    adapter_name: &str,\n    install_dir: Option<&Path>,\n    wake_command: &str,\n) -> Result<Vec<String>, String> {\n    let mut keys = std::collections::BTreeMap::from([\n        (\"id\".to_string(), shell_id.to_string()),\n        (\"adapter_name\".to_string(), adapter_name.to_string()),\n        (\n            \"link_token\".to_string(),\n            crate::shellhost::mint_link_token(),\n        ),\n    ]);\n    // [impl->REQ-INSTALL-11] opt-in and N-1-safe: a template that never names the\n    // key fills byte-identically.\n    if let Some(dir) = install_dir {\n        keys.insert(\"adapter_dir\".to_string(), dir.display().to_string());\n    }\n    // [impl->REQ-HAZARD-TEMPLATE-ARGV-FILL] tokenize-template-then-fill-each: a\n    // multi-word/quote/semicolon {key} value is exactly one argv element.\n    let mut tokens =\n        spt_runtime::runtime::fill_template_tokens(wake_command, &keys).map_err(|e| e.to_string())?;\n    let Some(program) = tokens.first_mut() else {\n        return Err(\"empty wake_command\".into());\n    };\n    // [impl->REQ-INSTALL-11] one resolution primitive, no parallel path.\n    if let Some(dir) = install_dir {\n        *program = spt_runtime::runtime::resolve_program_in_dir(program, dir);\n    }\n    Ok(tokens)\n}\n\n/// One reconcile sweep — the invariant holder for the online/offline mutual\n/// exclusivity (CONTEXT: \"spt-core flips between them\"):\n///\n/// - **offline + `wake_command` + no give-up latch + no live watcher** ⇒\n///   kill any orphaned waker pid (a dead daemon's leftover), then start a\n///   supervision thread.\n/// - **not offline (onlined / torn down / deregistered) + live watcher** ⇒\n///   stop it (flag + kill).\n///\n/// Runs at daemon boot and every tick — lifecycle flips happen in CLI-process\n\n[Showing lines 299-593 of 2161. Use :594 to continue]"}],"details":{"truncation":{"content":"                // An unspawnable waker is a permanent condition — latch now.\r\n                spt_proto::emit_line_err!(\"WAKER_SPAWN_FAIL:{owner}/{shell_id}: {e} (giving up)\");\r\n                let _ = std::fs::write(perch.join(WAKER_GAVE_UP_FILE), b\"\");\r\n                return;\r\n            }\r\n        };\r\n        // The birth stamp rides WITH the pid, in one write — the kill path\r\n        // authenticates the pair, and a stamp parked separately would leave a\r\n        // window that reads Held. [impl->REQ-SHELL-KILL-AUTHENTICATED]\r\n        record_waker_launch(&perch, child.id());\r\n        let status = child.wait();\r\n        let _ = std::fs::remove_file(perch.join(WAKER_PID_FILE));\r\n        if stop.load(Ordering::SeqCst) {\r\n            return; // stopped from outside — the kill is not a crash\r\n        }\r\n        match status.ok().and_then(|s| s.code()) {\r\n            Some(WAKE_OPCODE) => {\r\n                match resolve() {\r\n                    Ok(what) => spt_proto::emit_line_err!(\"SHELL_WAKE:{owner}/{shell_id}: {what}\"),\r\n                    Err(e) => spt_proto::emit_line_err!(\"SHELL_WAKE_FAIL:{owner}/{shell_id}: {e}\"),\r\n                }\r\n                return;\r\n            }\r\n            _ => {\r\n                failures += 1;\r\n                if failures >= params.give_up_after {\r\n                    spt_proto::emit_line_err!(\r\n                        \"WAKER_GAVE_UP:{owner}/{shell_id}: {failures} consecutive crash-exits \\\r\n                         (cleared by the next relink/launch)\"\r\n                    );\r\n                    let _ = std::fs::write(perch.join(WAKER_GAVE_UP_FILE), b\"\");\r\n                    return;\r\n                }\r\n                // Backoff in small slices so a stop request lands promptly.\r\n                let mut left = backoff_ms(params, failures);\r\n                while left > 0 && !stop.load(Ordering::SeqCst) {\r\n                    let step = left.min(50);\r\n                    std::thread::sleep(Duration::from_millis(step));\r\n                    left -= step;\r\n                }\r\n            }\r\n        }\r\n    }\r\n}\r\n\r\n/// The **state-keyed wake resolution** (M5-D4c — CONTEXT §Shell sleep/wake):\r\n/// key on the owner's local rest record, then bring the shell online —\r\n///\r\n/// - owner **dormant** (resting warm, still running) ⇒ touch nothing on the\r\n///   endpoint, just relaunch the shell binary;\r\n/// - owner **suspended** ⇒ revive the owner first\r\n///   ([`crate::resting::daemon_rest_event`] `Wake` — its cascade relaunches\r\n///   the owner's *persistent* shells, so this fn launches only if the\r\n///   cascade did not already), then the shell;\r\n/// - owner **active** (or recordless — a pre-resting/interim perch) ⇒ just\r\n///   relaunch the shell;\r\n/// - **no local instance of the owner** ⇒ refuse, naming the deferral: the\r\n///   `shell_wake_spawn_anywhere` fresh-spawn branch rides\r\n///   instantiate-anywhere (the D1c grant shape is its seam), and the\r\n///   *active-elsewhere cross-node attach* arm upgrades with presence/MRA\r\n///   (D6) + cross-node link (D8c) — D4 resolves against the local node.\r\n///\r\n/// The shell's bind onlines the perch (the D3b contract) — this fn never\r\n/// flips status.\r\n// [impl->REQ-SHELL-2]\r\npub fn resolve_wake(\r\n    owlery: &Path,\r\n    owner: &str,\r\n    shell_id: &str,\r\n    adapter_name: &str,\r\n    shell: &Shell,\r\n) -> Result<String, String> {\r\n    // Shells hang off flat-Self perches only (the D3a layout decision).\r\n    let owner_perch = owlery.join(owner);\r\n    if !owner_perch.join(\"info.json\").exists() {\r\n        // No local instance — consult presence (M5-D6b, REQ-PRES-1): a\r\n        // routable instance of the owner on ANOTHER node gets the wake\r\n        // FORWARDED there through the D5b remote rest op (the target node's\r\n        // own cascade relaunches ITS persistent shells). This local shell\r\n        // stays offline — the cross-node shell *link* is D8c; say so.\r\n        // [impl->REQ-PRES-1]\r\n        if let Some(node) = remote_owner_node(owner) {\r\n            return match forward_wake(owner, &node) {\r\n                Ok(outcome) => Ok(format!(\r\n                    \"WAKE_FORWARDED:{owner}@{node}: {outcome}; this shell stays offline \\\r\n                     here (cross-node shell link lands at D8c)\"\r\n                )),\r\n                Err(e) => Err(format!(\"WAKE_FORWARD_FAIL:{owner}@{node}: {e}\")),\r\n            };\r\n        }\r\n        return Err(format!(\r\n            \"WAKE_NO_REACHABLE_INSTANCE:{owner}: no instance of the owner on any \\\r\n             reachable node — a fresh-spawn-to-wake rides the deferred \\\r\n             instantiate-anywhere capability (the shell_wake_spawn_anywhere grant \\\r\n             shape is its seam)\"\r\n        ));\r\n    }\r\n    let state = crate::resting::read_rest(&owner_perch).map(|r| r.state);\r\n    let mut did = String::new();\r\n    if state == Some(crate::resting::RestState::Suspended) {\r\n        match crate::resting::daemon_rest_event(owner, crate::resting::RestEvent::Wake, None) {\r\n            Ok(_) => did.push_str(\"revived owner; \"),\r\n            Err(e) => return Err(format!(\"revive owner: {e}\")),\r\n        }\r\n    } else if state == Some(crate::resting::RestState::Dormant) {\r\n        did.push_str(\"owner dormant (left in place); \");\r\n    }\r\n    // Never double-launch a live binary: the revive's wake cascade may have\r\n    // relaunched this (persistent) instance already — and two reconcilers\r\n    // (a fresh daemon adopting + a stale watcher firing) may race a wake.\r\n    // A parked pid that probes alive means somebody won; stand down.\r\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\r\n    let live_pid = std::fs::read_to_string(perch.join(crate::shellhost::SHELL_PID_FILE))\r\n        .ok()\r\n        .and_then(|s| s.trim().parse::<u32>().ok())\r\n        .filter(|&p| p != 0 && spt_store::proc::is_process_alive(p));\r\n    if let Some(p) = live_pid {\r\n        return Ok(format!(\"{did}already relaunched pid={p} (online at bind)\"));\r\n    }\r\n    // [impl->REQ-INSTALL-11] the wake-triggered relaunch resolves against the\r\n    // adapter's install dir, like every other launch of the same instance.\r\n    let install_dir =\r\n        crate::shellhost::shell_install_dir(&spt_store::perch::adapters_dir(), adapter_name);\r\n    let pid = crate::shellhost::launch_shell(\r\n        owlery,\r\n        owner,\r\n        shell_id,\r\n        adapter_name,\r\n        install_dir.as_deref(),\r\n        shell,\r\n    )?;\r\n    Ok(format!(\"{did}relaunched pid={pid} (online at bind)\"))\r\n}\r\n\r\n/// The node (hex) holding a routable instance of `owner` per the gossiped\r\n/// registry snapshots, preferring the most-recently-active row when several\r\n/// nodes hold one (the presence datum as the tiebreak; absent data ranks\r\n/// last). `None` = the owner exists nowhere reachable.\r\n// [impl->REQ-PRES-1]\r\nfn remote_owner_node(owner: &str) -> Option<String> {\r\n    let regs = crate::presence::load_registry_snapshots(&crate::presence::registry_snapshot_dir());\r\n    let local = crate::presence::local_node_hex();\r\n    let mut best: Option<(u64, String)> = None;\r\n    for reg in regs.values() {\r\n        for row in reg.instances(owner) {\r\n            if row.node == local || !row.status.routable() {\r\n                continue;\r\n            }\r\n            let ts = row.last_active_ms.unwrap_or(0);\r\n            let better = match &best {\r\n                Some((cur, _)) => ts > *cur,\r\n                None => true,\r\n            };\r\n            if better {\r\n                best = Some((ts, row.node.clone()));\r\n            }\r\n        }\r\n    }\r\n    best.map(|(_, node)| node)\r\n}\r\n\r\n/// Forward a wake to `node` through the D5b remote rest op (the remote-drive\r\n/// trust class — the target's access gate decides). One dial, one request,\r\n/// one reply; a refusal (no reply) surfaces as an error.\r\n// [impl->REQ-PRES-1]\r\nfn forward_wake(owner: &str, node: &str) -> Result<String, String> {\r\n    use spt_net::net::endpoint::addr_for_node_hex;\r\n    let addr = addr_for_node_hex(node)\r\n        .and_then(|a| serde_json::to_value(a).ok())\r\n        .ok_or_else(|| format!(\"no dialable address for {node}\"))?;\r\n    let mut brain =\r\n        crate::brain::Brain::cold_start(&crate::endpoint::broker_socket_name(), now_ms())\r\n            .map_err(|e| format!(\"broker connect: {e}\"))?;\r\n    // A-4b (REQ-OPID-TRACING-RETRY): the wake-forward is a tracing-only rest op —\r\n    // if a broker restart dropped the conn/stream after the op journaled, re-mint a\r\n    // FRESH `wake` op from the SAME source and run ONCE more. The dial rides INSIDE\r\n    // `run` because that restart drops the CONN too, not just the stream — a retry\r\n    // against the dead conn would surface a different error, not self-heal. A 2nd\r\n    // no-longer-held collapses to the helper's F-1 public string.\r\n    // [impl->REQ-OPID-TRACING-RETRY]\r\n    let outcome = crate::effect::with_tracing_retry(\r\n        // The wake-forward rest op's seq source is a fresh `now_ms()` — its OWN\r\n        // minting source, distinct from shellchan's spool-row counter (doyle ruling:\r\n        // the tag names the seq source, not the subsystem), so it stamps `wake`.\r\n        || Ok(crate::effect::MintedOp::new(crate::effect::Minter::Wake, now_ms())),\r\n        |op| {\r\n            let conn = brain.net_dial(addr.clone(), None)?;\r\n            // Keep the tracing string's seq consistent with the (possibly re-minted) op.\r\n            let op_id = format!(\"{}:wakefwd:{}\", crate::presence::local_node_hex(), op.seq);\r\n            crate::resthost::request_rest(\r\n                &mut brain,\r\n                conn.conn_id,\r\n                owner,\r\n                spt_net::net::rest::REST_EVENT_WAKE,\r\n                &op_id,\r\n                op,\r\n            )\r\n        },\r\n    )\r\n    .map_err(|e| format!(\"rest op: {e}\"))?;\r\n    match outcome {\r\n        crate::resthost::RestRequestOutcome::Edge(d) => Ok(format!(\"woke ({d})\")),\r\n        crate::resthost::RestRequestOutcome::NoEdge => Ok(\"already awake\".to_string()),\r\n        crate::resthost::RestRequestOutcome::Failed(e) => Err(e),\r\n        crate::resthost::RestRequestOutcome::NoReply => {\r\n            Err(\"refused or dropped (no reply)\".to_string())\r\n        }\r\n    }\r\n}\r\n\r\n/// Resolve one adapter **option**'s `[shell]` section through the merged view\r\n/// (composite addressing): `adapter_option` is the stored `<adapter>[:profile]`\r\n/// string, split → parent lookup in `registered` → profile overlay (shipped or\r\n/// local). Returns an **owned** [`Shell`] — a profile produces a fresh merged\r\n/// manifest, not a borrow into the parent slice. A bare name resolves to the\r\n/// parent unmodified; a deregistered parent or a failed overlay yields `None`.\r\nfn shell_section_of(\r\n    registered: &[(AdapterRecord, Manifest)],\r\n    adapters_dir: &Path,\r\n    adapter_option: &str,\r\n) -> Option<Shell> {\r\n    spt_runtime::registry::resolve_option_in(registered, adapters_dir, adapter_option)\r\n        .ok()\r\n        .filter(|m| m.adapter.kind == AdapterKind::Shell)\r\n        .and_then(|m| m.shell)\r\n}\r\n\r\nfn now_ms() -> u64 {\r\n    crate::brain::now_ms()\r\n}\r\n\r\n/// Fill a `wake_command` template.\r\n///\r\n/// **Its catalog and the spawn template's share a vocabulary; they are not the\r\n/// same set, and never claiming they are is the point of saying so here.** The\r\n/// keys mean the same thing in both (`id`, `adapter_name`, `link_token`,\r\n/// `adapter_dir`), so an author learns one manifest language — but `perch_dir`\r\n/// is **spawn-only by design** (REQ-SHELL-PERCH-DIR: it resolves a *live* link's\r\n/// perch-relative transfer paths, and a waker has no live link). A comment\r\n/// asserting set-EQUALITY over two sets that are separately extended is the\r\n/// drift this project single-sources against; it had already been false since\r\n/// `perch_dir` landed.\r\n///\r\n/// The waker's token is minted **unparked**: it substitutes (template\r\n/// compatibility) without ever verifying against the perch, because an offline\r\n/// link has no live credential by design — the close retired it, and a waker\r\n/// wakes by *exit code*, not by driving the link.\r\n///\r\n/// `install_dir` resolves the program token and fills `{adapter_dir}`, exactly as\r\n/// on the spawn side ([`crate::shellhost::fill_spawn_command`]) — a released\r\n/// adapter whose spawn resolves and whose wake does not would go permanently\r\n/// unwakeable the moment it went offline.\r\nfn fill_wake_command(\r\n    shell_id: &str,\r\n    adapter_name: &str,\r\n    install_dir: Option<&Path>,\r\n    wake_command: &str,\r\n) -> Result<Vec<String>, String> {\r\n    let mut keys = std::collections::BTreeMap::from([\r\n        (\"id\".to_string(), shell_id.to_string()),\r\n        (\"adapter_name\".to_string(), adapter_name.to_string()),\r\n        (\r\n            \"link_token\".to_string(),\r\n            crate::shellhost::mint_link_token(),\r\n        ),\r\n    ]);\r\n    // [impl->REQ-INSTALL-11] opt-in and N-1-safe: a template that never names the\r\n    // key fills byte-identically.\r\n    if let Some(dir) = install_dir {\r\n        keys.insert(\"adapter_dir\".to_string(), dir.display().to_string());\r\n    }\r\n    // [impl->REQ-HAZARD-TEMPLATE-ARGV-FILL] tokenize-template-then-fill-each: a\r\n    // multi-word/quote/semicolon {key} value is exactly one argv element.\r\n    let mut tokens =\r\n        spt_runtime::runtime::fill_template_tokens(wake_command, &keys).map_err(|e| e.to_string())?;\r\n    let Some(program) = tokens.first_mut() else {\r\n        return Err(\"empty wake_command\".into());\r\n    };\r\n    // [impl->REQ-INSTALL-11] one resolution primitive, no parallel path.\r\n    if let Some(dir) = install_dir {\r\n        *program = spt_runtime::runtime::resolve_program_in_dir(program, dir);\r\n    }\r\n    Ok(tokens)\r\n}\r\n\r\n/// One reconcile sweep — the invariant holder for the online/offline mutual\r\n/// exclusivity (CONTEXT: \"spt-core flips between them\"):\r\n///\r\n/// - **offline + `wake_command` + no give-up latch + no live watcher** ⇒\r\n///   kill any orphaned waker pid (a dead daemon's leftover), then start a\r\n///   supervision thread.\r\n/// - **not offline (onlined / torn down / deregistered) + live watcher** ⇒\r\n///   stop it (flag + kill).\r\n///\r\n/// Runs at daemon boot and every tick — lifecycle flips happen in CLI-process\r","truncated":true,"truncatedBy":"lines","totalLines":1863,"totalBytes":14209,"outputLines":295,"outputBytes":14209,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":2161,"displayContent":{"text":"pub fn watcher_run(\n…\n) {\n…\n    while !stop.load(Ordering::SeqCst) {\n…\n        let mut child = match child {\n…\n            Err(e) => {\n                // An unspawnable waker is a permanent condition — latch now.\r\n                spt_proto::emit_line_err!(\"WAKER_SPAWN_FAIL:{owner}/{shell_id}: {e} (giving up)\");\r\n                let _ = std::fs::write(perch.join(WAKER_GAVE_UP_FILE), b\"\");\r\n                return;\r\n            }\r\n        };\r\n        // The birth stamp rides WITH the pid, in one write — the kill path\r\n        // authenticates the pair, and a stamp parked separately would leave a\r\n        // window that reads Held. [impl->REQ-SHELL-KILL-AUTHENTICATED]\r\n        record_waker_launch(&perch, child.id());\r\n        let status = child.wait();\r\n        let _ = std::fs::remove_file(perch.join(WAKER_PID_FILE));\r\n        if stop.load(Ordering::SeqCst) {\r\n            return; // stopped from outside — the kill is not a crash\r\n        }\r\n        match status.ok().and_then(|s| s.code()) {\r\n            Some(WAKE_OPCODE) => {\r\n                match resolve() {\r\n                    Ok(what) => spt_proto::emit_line_err!(\"SHELL_WAKE:{owner}/{shell_id}: {what}\"),\r\n                    Err(e) => spt_proto::emit_line_err!(\"SHELL_WAKE_FAIL:{owner}/{shell_id}: {e}\"),\r\n                }\r\n                return;\r\n            }\r\n            _ => {\r\n                failures += 1;\r\n                if failures >= params.give_up_after {\r\n                    spt_proto::emit_line_err!(\r\n                        \"WAKER_GAVE_UP:{owner}/{shell_id}: {failures} consecutive crash-exits \\\r\n                         (cleared by the next relink/launch)\"\r\n                    );\r\n                    let _ = std::fs::write(perch.join(WAKER_GAVE_UP_FILE), b\"\");\r\n                    return;\r\n                }\r\n                // Backoff in small slices so a stop request lands promptly.\r\n                let mut left = backoff_ms(params, failures);\r\n                while left > 0 && !stop.load(Ordering::SeqCst) {\r\n                    let step = left.min(50);\r\n                    std::thread::sleep(Duration::from_millis(step));\r\n                    left -= step;\r\n                }\r\n            }\r\n        }\r\n    }\r\n}\r\n\r\n/// The **state-keyed wake resolution** (M5-D4c — CONTEXT §Shell sleep/wake):\r\n/// key on the owner's local rest record, then bring the shell online —\r\n///\r\n/// - owner **dormant** (resting warm, still running) ⇒ touch nothing on the\r\n///   endpoint, just relaunch the shell binary;\r\n/// - owner **suspended** ⇒ revive the owner first\r\n///   ([`crate::resting::daemon_rest_event`] `Wake` — its cascade relaunches\r\n///   the owner's *persistent* shells, so this fn launches only if the\r\n///   cascade did not already), then the shell;\r\n/// - owner **active** (or recordless — a pre-resting/interim perch) ⇒ just\r\n///   relaunch the shell;\r\n/// - **no local instance of the owner** ⇒ refuse, naming the deferral: the\r\n///   `shell_wake_spawn_anywhere` fresh-spawn branch rides\r\n///   instantiate-anywhere (the D1c grant shape is its seam), and the\r\n///   *active-elsewhere cross-node attach* arm upgrades with presence/MRA\r\n///   (D6) + cross-node link (D8c) — D4 resolves against the local node.\r\n///\r\n/// The shell's bind onlines the perch (the D3b contract) — this fn never\r\n/// flips status.\r\n// [impl->REQ-SHELL-2]\r\npub fn resolve_wake(\r\n    owlery: &Path,\r\n    owner: &str,\r\n    shell_id: &str,\r\n    adapter_name: &str,\r\n    shell: &Shell,\r\n) -> Result<String, String> {\r\n    // Shells hang off flat-Self perches only (the D3a layout decision).\r\n    let owner_perch = owlery.join(owner);\r\n    if !owner_perch.join(\"info.json\").exists() {\r\n        // No local instance — consult presence (M5-D6b, REQ-PRES-1): a\r\n        // routable instance of the owner on ANOTHER node gets the wake\r\n        // FORWARDED there through the D5b remote rest op (the target node's\r\n        // own cascade relaunches ITS persistent shells). This local shell\r\n        // stays offline — the cross-node shell *link* is D8c; say so.\r\n        // [impl->REQ-PRES-1]\r\n        if let Some(node) = remote_owner_node(owner) {\r\n            return match forward_wake(owner, &node) {\r\n                Ok(outcome) => Ok(format!(\r\n                    \"WAKE_FORWARDED:{owner}@{node}: {outcome}; this shell stays offline \\\r\n                     here (cross-node shell link lands at D8c)\"\r\n                )),\r\n                Err(e) => Err(format!(\"WAKE_FORWARD_FAIL:{owner}@{node}: {e}\")),\r\n            };\r\n        }\r\n        return Err(format!(\r\n            \"WAKE_NO_REACHABLE_INSTANCE:{owner}: no instance of the owner on any \\\r\n             reachable node — a fresh-spawn-to-wake rides the deferred \\\r\n             instantiate-anywhere capability (the shell_wake_spawn_anywhere grant \\\r\n             shape is its seam)\"\r\n        ));\r\n    }\r\n    let state = crate::resting::read_rest(&owner_perch).map(|r| r.state);\r\n    let mut did = String::new();\r\n    if state == Some(crate::resting::RestState::Suspended) {\r\n        match crate::resting::daemon_rest_event(owner, crate::resting::RestEvent::Wake, None) {\r\n            Ok(_) => did.push_str(\"revived owner; \"),\r\n            Err(e) => return Err(format!(\"revive owner: {e}\")),\r\n        }\r\n    } else if state == Some(crate::resting::RestState::Dormant) {\r\n        did.push_str(\"owner dormant (left in place); \");\r\n    }\r\n    // Never double-launch a live binary: the revive's wake cascade may have\r\n    // relaunched this (persistent) instance already — and two reconcilers\r\n    // (a fresh daemon adopting + a stale watcher firing) may race a wake.\r\n    // A parked pid that probes alive means somebody won; stand down.\r\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\r\n    let live_pid = std::fs::read_to_string(perch.join(crate::shellhost::SHELL_PID_FILE))\r\n        .ok()\r\n        .and_then(|s| s.trim().parse::<u32>().ok())\r\n        .filter(|&p| p != 0 && spt_store::proc::is_process_alive(p));\r\n    if let Some(p) = live_pid {\r\n        return Ok(format!(\"{did}already relaunched pid={p} (online at bind)\"));\r\n    }\r\n    // [impl->REQ-INSTALL-11] the wake-triggered relaunch resolves against the\r\n    // adapter's install dir, like every other launch of the same instance.\r\n    let install_dir =\r\n        crate::shellhost::shell_install_dir(&spt_store::perch::adapters_dir(), adapter_name);\r\n    let pid = crate::shellhost::launch_shell(\r\n        owlery,\r\n        owner,\r\n        shell_id,\r\n        adapter_name,\r\n        install_dir.as_deref(),\r\n        shell,\r\n    )?;\r\n    Ok(format!(\"{did}relaunched pid={pid} (online at bind)\"))\r\n}\r\n\r\n/// The node (hex) holding a routable instance of `owner` per the gossiped\r\n/// registry snapshots, preferring the most-recently-active row when several\r\n/// nodes hold one (the presence datum as the tiebreak; absent data ranks\r\n/// last). `None` = the owner exists nowhere reachable.\r\n// [impl->REQ-PRES-1]\r\nfn remote_owner_node(owner: &str) -> Option<String> {\r\n    let regs = crate::presence::load_registry_snapshots(&crate::presence::registry_snapshot_dir());\r\n    let local = crate::presence::local_node_hex();\r\n    let mut best: Option<(u64, String)> = None;\r\n    for reg in regs.values() {\r\n        for row in reg.instances(owner) {\r\n            if row.node == local || !row.status.routable() {\r\n                continue;\r\n            }\r\n            let ts = row.last_active_ms.unwrap_or(0);\r\n            let better = match &best {\r\n                Some((cur, _)) => ts > *cur,\r\n                None => true,\r\n            };\r\n            if better {\r\n                best = Some((ts, row.node.clone()));\r\n            }\r\n        }\r\n    }\r\n    best.map(|(_, node)| node)\r\n}\r\n\r\n/// Forward a wake to `node` through the D5b remote rest op (the remote-drive\r\n/// trust class — the target's access gate decides). One dial, one request,\r\n/// one reply; a refusal (no reply) surfaces as an error.\r\n// [impl->REQ-PRES-1]\r\nfn forward_wake(owner: &str, node: &str) -> Result<String, String> {\r\n    use spt_net::net::endpoint::addr_for_node_hex;\r\n    let addr = addr_for_node_hex(node)\r\n        .and_then(|a| serde_json::to_value(a).ok())\r\n        .ok_or_else(|| format!(\"no dialable address for {node}\"))?;\r\n    let mut brain =\r\n        crate::brain::Brain::cold_start(&crate::endpoint::broker_socket_name(), now_ms())\r\n            .map_err(|e| format!(\"broker connect: {e}\"))?;\r\n    // A-4b (REQ-OPID-TRACING-RETRY): the wake-forward is a tracing-only rest op —\r\n    // if a broker restart dropped the conn/stream after the op journaled, re-mint a\r\n    // FRESH `wake` op from the SAME source and run ONCE more. The dial rides INSIDE\r\n    // `run` because that restart drops the CONN too, not just the stream — a retry\r\n    // against the dead conn would surface a different error, not self-heal. A 2nd\r\n    // no-longer-held collapses to the helper's F-1 public string.\r\n    // [impl->REQ-OPID-TRACING-RETRY]\r\n    let outcome = crate::effect::with_tracing_retry(\r\n        // The wake-forward rest op's seq source is a fresh `now_ms()` — its OWN\r\n        // minting source, distinct from shellchan's spool-row counter (doyle ruling:\r\n        // the tag names the seq source, not the subsystem), so it stamps `wake`.\r\n        || Ok(crate::effect::MintedOp::new(crate::effect::Minter::Wake, now_ms())),\r\n        |op| {\r\n            let conn = brain.net_dial(addr.clone(), None)?;\r\n            // Keep the tracing string's seq consistent with the (possibly re-minted) op.\r\n            let op_id = format!(\"{}:wakefwd:{}\", crate::presence::local_node_hex(), op.seq);\r\n            crate::resthost::request_rest(\r\n                &mut brain,\r\n                conn.conn_id,\r\n                owner,\r\n                spt_net::net::rest::REST_EVENT_WAKE,\r\n                &op_id,\r\n                op,\r\n            )\r\n        },\r\n    )\r\n    .map_err(|e| format!(\"rest op: {e}\"))?;\r\n    match outcome {\r\n        crate::resthost::RestRequestOutcome::Edge(d) => Ok(format!(\"woke ({d})\")),\r\n        crate::resthost::RestRequestOutcome::NoEdge => Ok(\"already awake\".to_string()),\r\n        crate::resthost::RestRequestOutcome::Failed(e) => Err(e),\r\n        crate::resthost::RestRequestOutcome::NoReply => {\r\n            Err(\"refused or dropped (no reply)\".to_string())\r\n        }\r\n    }\r\n}\r\n\r\n/// Resolve one adapter **option**'s `[shell]` section through the merged view\r\n/// (composite addressing): `adapter_option` is the stored `<adapter>[:profile]`\r\n/// string, split → parent lookup in `registered` → profile overlay (shipped or\r\n/// local). Returns an **owned** [`Shell`] — a profile produces a fresh merged\r\n/// manifest, not a borrow into the parent slice. A bare name resolves to the\r\n/// parent unmodified; a deregistered parent or a failed overlay yields `None`.\r\nfn shell_section_of(\r\n    registered: &[(AdapterRecord, Manifest)],\r\n    adapters_dir: &Path,\r\n    adapter_option: &str,\r\n) -> Option<Shell> {\r\n    spt_runtime::registry::resolve_option_in(registered, adapters_dir, adapter_option)\r\n        .ok()\r\n        .filter(|m| m.adapter.kind == AdapterKind::Shell)\r\n        .and_then(|m| m.shell)\r\n}\r\n\r\nfn now_ms() -> u64 {\r\n    crate::brain::now_ms()\r\n}\r\n\r\n/// Fill a `wake_command` template.\r\n///\r\n/// **Its catalog and the spawn template's share a vocabulary; they are not the\r\n/// same set, and never claiming they are is the point of saying so here.** The\r\n/// keys mean the same thing in both (`id`, `adapter_name`, `link_token`,\r\n/// `adapter_dir`), so an author learns one manifest language — but `perch_dir`\r\n/// is **spawn-only by design** (REQ-SHELL-PERCH-DIR: it resolves a *live* link's\r\n/// perch-relative transfer paths, and a waker has no live link). A comment\r\n/// asserting set-EQUALITY over two sets that are separately extended is the\r\n/// drift this project single-sources against; it had already been false since\r\n/// `perch_dir` landed.\r\n///\r\n/// The waker's token is minted **unparked**: it substitutes (template\r\n/// compatibility) without ever verifying against the perch, because an offline\r\n/// link has no live credential by design — the close retired it, and a waker\r\n/// wakes by *exit code*, not by driving the link.\r\n///\r\n/// `install_dir` resolves the program token and fills `{adapter_dir}`, exactly as\r\n/// on the spawn side ([`crate::shellhost::fill_spawn_command`]) — a released\r\n/// adapter whose spawn resolves and whose wake does not would go permanently\r\n/// unwakeable the moment it went offline.\r\nfn fill_wake_command(\r\n    shell_id: &str,\r\n    adapter_name: &str,\r\n    install_dir: Option<&Path>,\r\n    wake_command: &str,\r\n) -> Result<Vec<String>, String> {\r\n    let mut keys = std::collections::BTreeMap::from([\r\n        (\"id\".to_string(), shell_id.to_string()),\r\n        (\"adapter_name\".to_string(), adapter_name.to_string()),\r\n        (\r\n            \"link_token\".to_string(),\r\n            crate::shellhost::mint_link_token(),\r\n        ),\r\n    ]);\r\n    // [impl->REQ-INSTALL-11] opt-in and N-1-safe: a template that never names the\r\n    // key fills byte-identically.\r\n    if let Some(dir) = install_dir {\r\n        keys.insert(\"adapter_dir\".to_string(), dir.display().to_string());\r\n    }\r\n    // [impl->REQ-HAZARD-TEMPLATE-ARGV-FILL] tokenize-template-then-fill-each: a\r\n    // multi-word/quote/semicolon {key} value is exactly one argv element.\r\n    let mut tokens =\r\n        spt_runtime::runtime::fill_template_tokens(wake_command, &keys).map_err(|e| e.to_string())?;\r\n    let Some(program) = tokens.first_mut() else {\r\n        return Err(\"empty wake_command\".into());\r\n    };\r\n    // [impl->REQ-INSTALL-11] one resolution primitive, no parallel path.\r\n    if let Some(dir) = install_dir {\r\n        *program = spt_runtime::runtime::resolve_program_in_dir(program, dir);\r\n    }\r\n    Ok(tokens)\r\n}\r\n\r\n/// One reconcile sweep — the invariant holder for the online/offline mutual\r\n/// exclusivity (CONTEXT: \"spt-core flips between them\"):\r\n///\r\n/// - **offline + `wake_command` + no give-up latch + no live watcher** ⇒\r\n///   kill any orphaned waker pid (a dead daemon's leftover), then start a\r\n///   supervision thread.\r\n/// - **not offline (onlined / torn down / deregistered) + live watcher** ⇒\r\n///   stop it (flag + kill).\r\n///\r\n/// Runs at daemon boot and every tick — lifecycle flips happen in CLI-process\r","startLine":265,"lineNumbers":[265,null,273,null,280,null,296,null,298,299,300,301,302,303,304,305,306,307,308,309,310,311,312,313,314,315,316,317,318,319,320,321,322,323,324,325,326,327,328,329,330,331,332,333,334,335,336,337,338,339,340,341,342,343,344,345,346,347,348,349,350,351,352,353,354,355,356,357,358,359,360,361,362,363,364,365,366,367,368,369,370,371,372,373,374,375,376,377,378,379,380,381,382,383,384,385,386,387,388,389,390,391,392,393,394,395,396,397,398,399,400,401,402,403,404,405,406,407,408,409,410,411,412,413,414,415,416,417,418,419,420,421,422,423,424,425,426,427,428,429,430,431,432,433,434,435,436,437,438,439,440,441,442,443,444,445,446,447,448,449,450,451,452,453,454,455,456,457,458,459,460,461,462,463,464,465,466,467,468,469,470,471,472,473,474,475,476,477,478,479,480,481,482,483,484,485,486,487,488,489,490,491,492,493,494,495,496,497,498,499,500,501,502,503,504,505,506,507,508,509,510,511,512,513,514,515,516,517,518,519,520,521,522,523,524,525,526,527,528,529,530,531,532,533,534,535,536,537,538,539,540,541,542,543,544,545,546,547,548,549,550,551,552,553,554,555,556,557,558,559,560,561,562,563,564,565,566,567,568,569,570,571,572,573,574,575,576,577,578,579,580,581,582,583,584,585,586,587,588,589,590,591,592,593]},"fileSize":100821,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\crates\\spt-daemon\\src\\shellwake.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":2161,"totalBytes":14209,"outputLines":295,"outputBytes":14209,"shownRange":{"start":299,"end":593},"nextOffset":594}}},"isError":false,"timestamp":1789001151867}}
{"type":"message","id":"7daa760a","parentId":"6fdf3f2a","timestamp":"2026-09-10T00:45:51.885Z","message":{"role":"toolResult","toolCallId":"call_86q7GdHeGBRmNxe4753yUR3N|fc_077f3569365e313d016aa1fdbfc92487d18187cf49e1242bd2","toolName":"read","content":[{"type":"text","text":"/// `force = true` answers the field ask (releases#6): the instance IS up and\n/// wants a stop-then-restart WITHOUT churning its canonical id through\n/// teardown + spawn. The stop is [`shellhost::close_shell`] — the one existing\n/// link-break close, reused whole rather than reimplemented, so a forced relink\n/// breaks a link exactly the way every other link-break does (pre-close\n/// instruction, termination window, authenticated kill, token retired for the\n/// spool's sake, drive + tunnel slots cleared). It is bracketed by two refusals\n/// that `CloseOutcome` cannot express, which is why they live here and not in\n/// the close:\n///\n/// - **before**: an `ephemeral` instance's close is a TEARDOWN. Refused.\n/// - **after**: a close can legitimately leave the binary running (the kill is\n///   authenticated and may be refused, or the binary outlives the window), and\n///   launching there yields two binaries for one instance — the second one\n///   unreachable through the retired token and invisible to the record.\n///   Refused, on the SAME pid+birth pair the kill path authenticates with.\n// [impl->REQ-SHELL-RELINK-FORCE]\npub fn relink_shell_forced_from(\n    origin: LaunchOrigin,\n    owlery: &Path,\n    owner: &str,\n    shell_ref: &str,\n    force: bool,\n) -> Result<(String, u32), ShellLinkError> {\n    let (id, info, shell) = resolve_link_target(owlery, owner, shell_ref)?;\n    // DERIVED, not recorded (REQ-HAZARD-SHELL-STALE-ONLINE). This is the gate\n    // flynn's field report died on: an abruptly-killed binary leaves status\n    // `online`, so the one command that recovers the instance refused on the very\n    // state that was wrong, and teardown+spawn (perch destroyed, canonical id\n    // churned) was the only exit. The probe is LOCAL — recovery must not depend on\n    // a daemon sweep having run, or on a daemon running at all.\n    // [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, &id);\n    if shellinfo::is_shell_online(&perch, &info) {\n        if !force {\n            return Err(ShellLinkError::AlreadyOnline);\n        }\n        // [impl->REQ-SHELL-RELINK-FORCE] the stop half runs ONLY here: an\n        // already-offline instance takes the plain relink even under the flag,\n        // so --force is an override of one refusal and not a second code path\n        // for the ordinary case.\n        force_stop_for_relink(owlery, owner, &id, &shell, &perch)?;\n    }\n    // [impl->REQ-INSTALL-11] a relink is a launch — it resolves the adapter's\n    // shipped binary the same way `spt shell spawn` did, or the one command that\n    // recovers a released instance would refuse where the spawn succeeded.\n    let install_dir =\n        shellhost::shell_install_dir(&spt_store::perch::adapters_dir(), &info.adapter_name);\n    // An INDETERMINATE hand-off carries its uncertainty into the message rather\n    // than being flattened into \"failed\": a relink mints nothing, so there is\n    // nothing to roll back here, but a caller told \"failed\" about a binary that\n    // may well be coming up would relink again and double it.\n    let pid = shellhost::launch_shell_from(\n        origin,\n        owlery,\n        owner,\n        &id,\n        &info.adapter_name,\n        install_dir.as_deref(),\n        &shell,\n    )\n    .map_err(|e| match e {\n        shellhost::LaunchFailure::Indeterminate(m) => ShellLinkError::Failed(format!(\n            \"launch handed to the daemon, outcome UNKNOWN — the binary may or may not be \\\n             running ({m}); check `spt shell list` before relinking again\"\n        )),\n        shellhost::LaunchFailure::Failed(m) => ShellLinkError::Failed(format!(\"launch: {m}\")),\n    })?;\n    Ok((id, pid))\n}\n\n/// The DAEMON half of a routed CLI launch (REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE): a\n/// THIN launch keyed `{owner, shell_id}` — resolve the record, the registered\n/// `[shell]` and the adapter install dir, then spawn HERE, where the process is\n/// job-neutral.\n///\n/// Deliberately NOT [`relink_shell`]: relink is the ONLINE SWITCH for an\n/// existing instance and carries an already-online refusal plus a link-token\n/// vocabulary a FRESH spawn has no answer to (a just-minted instance has no\n/// parked token at all ⇒ `NoLink`). And deliberately NOT a gate: the spawn\n/// gates, the instance cap and every consent prompt ran CLI-side before this op\n/// was ever sent — this call is the spawn, and only the spawn.\n// [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\npub fn launch_shell_daemon_side(owlery: &Path, owner: &str, shell_id: &str) -> Result<u32, String> {\n    let (id, info, shell) =\n        resolve_link_target(owlery, owner, shell_id).map_err(|e| format!(\"{e:?}\"))?;\n    // THE side that stands down when a late ack raced the caller into its own\n    // in-process fallback (see `LAUNCH_ACK_BOUND`). The ack precedes this\n    // launch, so a caller that gave up waiting spawned FIRST — meaning the\n    // binary this finds is its own, and adopting the pid answers the op\n    // truthfully. Non-fault wording on purpose: a stood-down launch is a\n    // success, and must never read like a fault latch.\n    // [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, &id);\n    if let Some(pid) = shellhost::live_launch_winner(&perch) {\n        spt_proto::emit_line_err!(\"SHELL_LAUNCH_STOOD_DOWN:{id}: already launched pid={pid} (online at bind)\");\n        return Ok(pid);\n    }\n    let install_dir =\n        shellhost::shell_install_dir(&spt_store::perch::adapters_dir(), &info.adapter_name);\n    shellhost::launch_shell(\n        owlery,\n        owner,\n        &id,\n        &info.adapter_name,\n        install_dir.as_deref(),\n        &shell,\n    )\n}\n\n/// Why a shell-link serve returned — loud + typed, rendered into the\n/// dispatch log.\n#[derive(Debug, Clone, PartialEq, Eq)]\npub enum ShellLinkServeOutcome {\n    /// The access gate refused the requesting origin — nothing ran, no reply.\n    Refused,\n    /// The action ran; the reply carried the report.\n    Ok(String),\n    /// The action failed; the failure text went back in the reply.\n    Failed(String),\n    /// The stream never produced a parseable request (or named an unknown\n    /// action) — dropped.\n    BadRequest(String),\n}\n\n/// Serve one shell-link stream against this node, blocking until the single\n/// request runs (or refuses). `origin_node` MUST be the handshake-proven\n/// remote id from the broker's stream table — never payload bytes\n/// (REQ-HAZARD-WAN-ORIGIN-AUTH). The gate subject is the wire-named owner\n/// (the xfer posture: the wire picks whose whitelist, the origin it is\n/// checked against stays transport-proven).\n// [impl->REQ-SHELL-2]\n// [impl->REQ-HAZARD-WAN-ORIGIN-AUTH]\npub fn serve_shell_link(\n    brain: &mut Brain,\n    stream_id: u64,\n    origin_node: &str,\n    owlery: &Path,\n) -> io::Result<ShellLinkServeOutcome> {\n    brain.net_stream_subscribe(stream_id, 0)?;\n    let mut decoder = ShellLinkDecoder::new();\n    loop {\n        match brain.read_event()? {\n            BrokerEvent::NetStreamData {\n                stream_id: sid,\n                bytes,\n                ..\n            } if sid == stream_id => {\n                for rec in decoder.push(&bytes) {\n                    let ShellLinkRecord::Request {\n…\n                    } = rec\n…\n                    };\n…\n                }\n            }\n…\n        }\n    }\n}\n\n[Showing lines 644-793 of 1747. Use :794 to continue]"}],"details":{"truncation":{"content":"/// `force = true` answers the field ask (releases#6): the instance IS up and\r\n/// wants a stop-then-restart WITHOUT churning its canonical id through\r\n/// teardown + spawn. The stop is [`shellhost::close_shell`] — the one existing\r\n/// link-break close, reused whole rather than reimplemented, so a forced relink\r\n/// breaks a link exactly the way every other link-break does (pre-close\r\n/// instruction, termination window, authenticated kill, token retired for the\r\n/// spool's sake, drive + tunnel slots cleared). It is bracketed by two refusals\r\n/// that `CloseOutcome` cannot express, which is why they live here and not in\r\n/// the close:\r\n///\r\n/// - **before**: an `ephemeral` instance's close is a TEARDOWN. Refused.\r\n/// - **after**: a close can legitimately leave the binary running (the kill is\r\n///   authenticated and may be refused, or the binary outlives the window), and\r\n///   launching there yields two binaries for one instance — the second one\r\n///   unreachable through the retired token and invisible to the record.\r\n///   Refused, on the SAME pid+birth pair the kill path authenticates with.\r\n// [impl->REQ-SHELL-RELINK-FORCE]\r\npub fn relink_shell_forced_from(\r\n    origin: LaunchOrigin,\r\n    owlery: &Path,\r\n    owner: &str,\r\n    shell_ref: &str,\r\n    force: bool,\r\n) -> Result<(String, u32), ShellLinkError> {\r\n    let (id, info, shell) = resolve_link_target(owlery, owner, shell_ref)?;\r\n    // DERIVED, not recorded (REQ-HAZARD-SHELL-STALE-ONLINE). This is the gate\r\n    // flynn's field report died on: an abruptly-killed binary leaves status\r\n    // `online`, so the one command that recovers the instance refused on the very\r\n    // state that was wrong, and teardown+spawn (perch destroyed, canonical id\r\n    // churned) was the only exit. The probe is LOCAL — recovery must not depend on\r\n    // a daemon sweep having run, or on a daemon running at all.\r\n    // [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\r\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, &id);\r\n    if shellinfo::is_shell_online(&perch, &info) {\r\n        if !force {\r\n            return Err(ShellLinkError::AlreadyOnline);\r\n        }\r\n        // [impl->REQ-SHELL-RELINK-FORCE] the stop half runs ONLY here: an\r\n        // already-offline instance takes the plain relink even under the flag,\r\n        // so --force is an override of one refusal and not a second code path\r\n        // for the ordinary case.\r\n        force_stop_for_relink(owlery, owner, &id, &shell, &perch)?;\r\n    }\r\n    // [impl->REQ-INSTALL-11] a relink is a launch — it resolves the adapter's\r\n    // shipped binary the same way `spt shell spawn` did, or the one command that\r\n    // recovers a released instance would refuse where the spawn succeeded.\r\n    let install_dir =\r\n        shellhost::shell_install_dir(&spt_store::perch::adapters_dir(), &info.adapter_name);\r\n    // An INDETERMINATE hand-off carries its uncertainty into the message rather\r\n    // than being flattened into \"failed\": a relink mints nothing, so there is\r\n    // nothing to roll back here, but a caller told \"failed\" about a binary that\r\n    // may well be coming up would relink again and double it.\r\n    let pid = shellhost::launch_shell_from(\r\n        origin,\r\n        owlery,\r\n        owner,\r\n        &id,\r\n        &info.adapter_name,\r\n        install_dir.as_deref(),\r\n        &shell,\r\n    )\r\n    .map_err(|e| match e {\r\n        shellhost::LaunchFailure::Indeterminate(m) => ShellLinkError::Failed(format!(\r\n            \"launch handed to the daemon, outcome UNKNOWN — the binary may or may not be \\\r\n             running ({m}); check `spt shell list` before relinking again\"\r\n        )),\r\n        shellhost::LaunchFailure::Failed(m) => ShellLinkError::Failed(format!(\"launch: {m}\")),\r\n    })?;\r\n    Ok((id, pid))\r\n}\r\n\r\n/// The DAEMON half of a routed CLI launch (REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE): a\r\n/// THIN launch keyed `{owner, shell_id}` — resolve the record, the registered\r\n/// `[shell]` and the adapter install dir, then spawn HERE, where the process is\r\n/// job-neutral.\r\n///\r\n/// Deliberately NOT [`relink_shell`]: relink is the ONLINE SWITCH for an\r\n/// existing instance and carries an already-online refusal plus a link-token\r\n/// vocabulary a FRESH spawn has no answer to (a just-minted instance has no\r\n/// parked token at all ⇒ `NoLink`). And deliberately NOT a gate: the spawn\r\n/// gates, the instance cap and every consent prompt ran CLI-side before this op\r\n/// was ever sent — this call is the spawn, and only the spawn.\r\n// [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\r\npub fn launch_shell_daemon_side(owlery: &Path, owner: &str, shell_id: &str) -> Result<u32, String> {\r\n    let (id, info, shell) =\r\n        resolve_link_target(owlery, owner, shell_id).map_err(|e| format!(\"{e:?}\"))?;\r\n    // THE side that stands down when a late ack raced the caller into its own\r\n    // in-process fallback (see `LAUNCH_ACK_BOUND`). The ack precedes this\r\n    // launch, so a caller that gave up waiting spawned FIRST — meaning the\r\n    // binary this finds is its own, and adopting the pid answers the op\r\n    // truthfully. Non-fault wording on purpose: a stood-down launch is a\r\n    // success, and must never read like a fault latch.\r\n    // [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\r\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, &id);\r\n    if let Some(pid) = shellhost::live_launch_winner(&perch) {\r\n        spt_proto::emit_line_err!(\"SHELL_LAUNCH_STOOD_DOWN:{id}: already launched pid={pid} (online at bind)\");\r\n        return Ok(pid);\r\n    }\r\n    let install_dir =\r\n        shellhost::shell_install_dir(&spt_store::perch::adapters_dir(), &info.adapter_name);\r\n    shellhost::launch_shell(\r\n        owlery,\r\n        owner,\r\n        &id,\r\n        &info.adapter_name,\r\n        install_dir.as_deref(),\r\n        &shell,\r\n    )\r\n}\r\n\r\n/// Why a shell-link serve returned — loud + typed, rendered into the\r\n/// dispatch log.\r\n#[derive(Debug, Clone, PartialEq, Eq)]\r\npub enum ShellLinkServeOutcome {\r\n    /// The access gate refused the requesting origin — nothing ran, no reply.\r\n    Refused,\r\n    /// The action ran; the reply carried the report.\r\n    Ok(String),\r\n    /// The action failed; the failure text went back in the reply.\r\n    Failed(String),\r\n    /// The stream never produced a parseable request (or named an unknown\r\n    /// action) — dropped.\r\n    BadRequest(String),\r\n}\r\n\r\n/// Serve one shell-link stream against this node, blocking until the single\r\n/// request runs (or refuses). `origin_node` MUST be the handshake-proven\r\n/// remote id from the broker's stream table — never payload bytes\r\n/// (REQ-HAZARD-WAN-ORIGIN-AUTH). The gate subject is the wire-named owner\r\n/// (the xfer posture: the wire picks whose whitelist, the origin it is\r\n/// checked against stays transport-proven).\r\n// [impl->REQ-SHELL-2]\r\n// [impl->REQ-HAZARD-WAN-ORIGIN-AUTH]\r\npub fn serve_shell_link(\r\n    brain: &mut Brain,\r\n    stream_id: u64,\r\n    origin_node: &str,\r\n    owlery: &Path,\r\n) -> io::Result<ShellLinkServeOutcome> {\r\n    brain.net_stream_subscribe(stream_id, 0)?;\r\n    let mut decoder = ShellLinkDecoder::new();\r\n    loop {\r\n        match brain.read_event()? {\r\n            BrokerEvent::NetStreamData {\r\n                stream_id: sid,\r\n                bytes,\r\n                ..\r\n            } if sid == stream_id => {\r\n                for rec in decoder.push(&bytes) {\r\n                    let ShellLinkRecord::Request {\r","truncated":true,"truncatedBy":"lines","totalLines":1104,"totalBytes":7464,"outputLines":150,"outputBytes":7464,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":1747,"displayContent":{"text":"/// `force = true` answers the field ask (releases#6): the instance IS up and\r\n/// wants a stop-then-restart WITHOUT churning its canonical id through\r\n/// teardown + spawn. The stop is [`shellhost::close_shell`] — the one existing\r\n/// link-break close, reused whole rather than reimplemented, so a forced relink\r\n/// breaks a link exactly the way every other link-break does (pre-close\r\n/// instruction, termination window, authenticated kill, token retired for the\r\n/// spool's sake, drive + tunnel slots cleared). It is bracketed by two refusals\r\n/// that `CloseOutcome` cannot express, which is why they live here and not in\r\n/// the close:\r\n///\r\n/// - **before**: an `ephemeral` instance's close is a TEARDOWN. Refused.\r\n/// - **after**: a close can legitimately leave the binary running (the kill is\r\n///   authenticated and may be refused, or the binary outlives the window), and\r\n///   launching there yields two binaries for one instance — the second one\r\n///   unreachable through the retired token and invisible to the record.\r\n///   Refused, on the SAME pid+birth pair the kill path authenticates with.\r\n// [impl->REQ-SHELL-RELINK-FORCE]\r\npub fn relink_shell_forced_from(\r\n    origin: LaunchOrigin,\r\n    owlery: &Path,\r\n    owner: &str,\r\n    shell_ref: &str,\r\n    force: bool,\r\n) -> Result<(String, u32), ShellLinkError> {\r\n    let (id, info, shell) = resolve_link_target(owlery, owner, shell_ref)?;\r\n    // DERIVED, not recorded (REQ-HAZARD-SHELL-STALE-ONLINE). This is the gate\r\n    // flynn's field report died on: an abruptly-killed binary leaves status\r\n    // `online`, so the one command that recovers the instance refused on the very\r\n    // state that was wrong, and teardown+spawn (perch destroyed, canonical id\r\n    // churned) was the only exit. The probe is LOCAL — recovery must not depend on\r\n    // a daemon sweep having run, or on a daemon running at all.\r\n    // [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\r\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, &id);\r\n    if shellinfo::is_shell_online(&perch, &info) {\r\n        if !force {\r\n            return Err(ShellLinkError::AlreadyOnline);\r\n        }\r\n        // [impl->REQ-SHELL-RELINK-FORCE] the stop half runs ONLY here: an\r\n        // already-offline instance takes the plain relink even under the flag,\r\n        // so --force is an override of one refusal and not a second code path\r\n        // for the ordinary case.\r\n        force_stop_for_relink(owlery, owner, &id, &shell, &perch)?;\r\n    }\r\n    // [impl->REQ-INSTALL-11] a relink is a launch — it resolves the adapter's\r\n    // shipped binary the same way `spt shell spawn` did, or the one command that\r\n    // recovers a released instance would refuse where the spawn succeeded.\r\n    let install_dir =\r\n        shellhost::shell_install_dir(&spt_store::perch::adapters_dir(), &info.adapter_name);\r\n    // An INDETERMINATE hand-off carries its uncertainty into the message rather\r\n    // than being flattened into \"failed\": a relink mints nothing, so there is\r\n    // nothing to roll back here, but a caller told \"failed\" about a binary that\r\n    // may well be coming up would relink again and double it.\r\n    let pid = shellhost::launch_shell_from(\r\n        origin,\r\n        owlery,\r\n        owner,\r\n        &id,\r\n        &info.adapter_name,\r\n        install_dir.as_deref(),\r\n        &shell,\r\n    )\r\n    .map_err(|e| match e {\r\n        shellhost::LaunchFailure::Indeterminate(m) => ShellLinkError::Failed(format!(\r\n            \"launch handed to the daemon, outcome UNKNOWN — the binary may or may not be \\\r\n             running ({m}); check `spt shell list` before relinking again\"\r\n        )),\r\n        shellhost::LaunchFailure::Failed(m) => ShellLinkError::Failed(format!(\"launch: {m}\")),\r\n    })?;\r\n    Ok((id, pid))\r\n}\r\n\r\n/// The DAEMON half of a routed CLI launch (REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE): a\r\n/// THIN launch keyed `{owner, shell_id}` — resolve the record, the registered\r\n/// `[shell]` and the adapter install dir, then spawn HERE, where the process is\r\n/// job-neutral.\r\n///\r\n/// Deliberately NOT [`relink_shell`]: relink is the ONLINE SWITCH for an\r\n/// existing instance and carries an already-online refusal plus a link-token\r\n/// vocabulary a FRESH spawn has no answer to (a just-minted instance has no\r\n/// parked token at all ⇒ `NoLink`). And deliberately NOT a gate: the spawn\r\n/// gates, the instance cap and every consent prompt ran CLI-side before this op\r\n/// was ever sent — this call is the spawn, and only the spawn.\r\n// [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\r\npub fn launch_shell_daemon_side(owlery: &Path, owner: &str, shell_id: &str) -> Result<u32, String> {\r\n    let (id, info, shell) =\r\n        resolve_link_target(owlery, owner, shell_id).map_err(|e| format!(\"{e:?}\"))?;\r\n    // THE side that stands down when a late ack raced the caller into its own\r\n    // in-process fallback (see `LAUNCH_ACK_BOUND`). The ack precedes this\r\n    // launch, so a caller that gave up waiting spawned FIRST — meaning the\r\n    // binary this finds is its own, and adopting the pid answers the op\r\n    // truthfully. Non-fault wording on purpose: a stood-down launch is a\r\n    // success, and must never read like a fault latch.\r\n    // [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\r\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, &id);\r\n    if let Some(pid) = shellhost::live_launch_winner(&perch) {\r\n        spt_proto::emit_line_err!(\"SHELL_LAUNCH_STOOD_DOWN:{id}: already launched pid={pid} (online at bind)\");\r\n        return Ok(pid);\r\n    }\r\n    let install_dir =\r\n        shellhost::shell_install_dir(&spt_store::perch::adapters_dir(), &info.adapter_name);\r\n    shellhost::launch_shell(\r\n        owlery,\r\n        owner,\r\n        &id,\r\n        &info.adapter_name,\r\n        install_dir.as_deref(),\r\n        &shell,\r\n    )\r\n}\r\n\r\n/// Why a shell-link serve returned — loud + typed, rendered into the\r\n/// dispatch log.\r\n#[derive(Debug, Clone, PartialEq, Eq)]\r\npub enum ShellLinkServeOutcome {\r\n    /// The access gate refused the requesting origin — nothing ran, no reply.\r\n    Refused,\r\n    /// The action ran; the reply carried the report.\r\n    Ok(String),\r\n    /// The action failed; the failure text went back in the reply.\r\n    Failed(String),\r\n    /// The stream never produced a parseable request (or named an unknown\r\n    /// action) — dropped.\r\n    BadRequest(String),\r\n}\r\n\r\n/// Serve one shell-link stream against this node, blocking until the single\r\n/// request runs (or refuses). `origin_node` MUST be the handshake-proven\r\n/// remote id from the broker's stream table — never payload bytes\r\n/// (REQ-HAZARD-WAN-ORIGIN-AUTH). The gate subject is the wire-named owner\r\n/// (the xfer posture: the wire picks whose whitelist, the origin it is\r\n/// checked against stays transport-proven).\r\n// [impl->REQ-SHELL-2]\r\n// [impl->REQ-HAZARD-WAN-ORIGIN-AUTH]\r\npub fn serve_shell_link(\r\n    brain: &mut Brain,\r\n    stream_id: u64,\r\n    origin_node: &str,\r\n    owlery: &Path,\r\n) -> io::Result<ShellLinkServeOutcome> {\r\n    brain.net_stream_subscribe(stream_id, 0)?;\r\n    let mut decoder = ShellLinkDecoder::new();\r\n    loop {\r\n        match brain.read_event()? {\r\n            BrokerEvent::NetStreamData {\r\n                stream_id: sid,\r\n                bytes,\r\n                ..\r\n            } if sid == stream_id => {\r\n                for rec in decoder.push(&bytes) {\r\n                    let ShellLinkRecord::Request {\r\n…\n                    } = rec\n…\n                    };\n…\n                }\n            }\n…\n        }\n    }\n}","startLine":644,"lineNumbers":[644,645,646,647,648,649,650,651,652,653,654,655,656,657,658,659,660,661,662,663,664,665,666,667,668,669,670,671,672,673,674,675,676,677,678,679,680,681,682,683,684,685,686,687,688,689,690,691,692,693,694,695,696,697,698,699,700,701,702,703,704,705,706,707,708,709,710,711,712,713,714,715,716,717,718,719,720,721,722,723,724,725,726,727,728,729,730,731,732,733,734,735,736,737,738,739,740,741,742,743,744,745,746,747,748,749,750,751,752,753,754,755,756,757,758,759,760,761,762,763,764,765,766,767,768,769,770,771,772,773,774,775,776,777,778,779,780,781,782,783,784,785,786,787,788,789,790,791,792,793,null,799,null,802,null,820,821,null,830,831,832]},"fileSize":85499,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\crates\\spt-daemon\\src\\linkhost.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":1747,"totalBytes":7464,"outputLines":150,"outputBytes":7464,"shownRange":{"start":644,"end":793},"nextOffset":794}}},"isError":false,"timestamp":1789001151885}}
{"type":"message","id":"905372b3","parentId":"7daa760a","timestamp":"2026-09-10T00:45:51.897Z","message":{"role":"toolResult","toolCallId":"call_LhwwqTqtbNXGA5WfqmpAvEyX|fc_077f3569365e313d016aa1fdbfc90c87d185cc29b3b7e71d6a","toolName":"read","content":[{"type":"text","text":"//! Shell-instance hosting (M5-D3b — CONTEXT §Shell model, §local-link\n//! authentication): the broker-launch + local-link handshake half of the\n//! shell machinery.\n//!\n//! **Broker-launched, never user-launched.** [`launch_shell`] mints a\n//! per-link token, parks it on the shell perch, fills the manifest's opaque\n//! `spawn` template (`{link_token}` + `{id}`/`{adapter_name}` from the\n//! guaranteed catalog) and spawns the binary detached — the only way a shell\n//! binary ever starts. Launch failure rolls nothing forward: the caller\n//! tears the just-minted perch back down (no half-built instance).\n//!\n//! **The local-link handshake.** Cross-node traffic already rides Iroh\n//! (E2E-encrypted); the exposed surface is the *same-node* channel where the\n//! least-trusted code in the system (the shell binary) touches SPT. The\n//! handshake: possession of the minted token **is** the bind credential —\n//! `spt api bind-shell --link <token>` resolves the instance *by token*\n//! ([`bind_shell_by_token`]: the binary needs no owner/id knowledge beyond\n//! its argv) and flips the perch online. Every subsequent machinery message\n//! on the link authenticates with the token-derived key ([`link_key`] →\n//! [`frame_mac`], applied per-message by the D3c channels) so another local\n//! process can neither inject into nor spoof a link.\n// [impl->REQ-SHELL-1]\n\nuse std::collections::BTreeMap;\nuse std::path::{Path, PathBuf};\nuse std::process::{Command, Stdio};\n\nuse hmac::{Hmac, Mac};\nuse sha2::{Digest, Sha256};\n\nuse spt_runtime::manifest::{AdapterKind, CommandReceipt, Manifest, Shell};\nuse spt_runtime::registry::AdapterRecord;\nuse spt_store::liveness::RelayLiveness;\nuse spt_store::shellinfo::{self, SHELL_STATUS_OFFLINE, SHELL_STATUS_ONLINE};\n\nuse crate::brain::{now_ms, Brain};\nuse crate::msg::SpawnReq;\nuse crate::shellchan::shell_session_label;\n\n/// Per-instance link-token filename within a shell perch dir (node-local\n/// runtime material, beside `info.json` — the perch *record* contract in\n/// STORAGE.md constrains `info.json`'s fields, not the perch's files).\npub const LINK_TOKEN_FILE: &str = \"link.token\";\n\n/// The most recently RETIRED link token (stashed by `close_shell`, consumed by\n/// the next launch's re-stamp). Never a live credential — bind resolves only\n/// [`LINK_TOKEN_FILE`]; this exists so pending spool rows stamped under the\n/// dead link survive the next token mint (REQ-HAZARD-SHELL-STALE-ONLINE).\npub const RETIRED_TOKEN_FILE: &str = \"link.token.retired\";\n/// The launched binary's OS pid — teardown's kill handle, and the evidence the\n/// derived-status resolver probes. Re-exported, never re-typed: the name is\n/// written here (launch/close) and MATCHED there\n/// ([`shellinfo::shell_pid_provably_dead`]), so both sides must be one const.\n// [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\npub use spt_store::shellinfo::SHELL_PID_FILE;\n\n/// Mint a fresh link token: 32 hex chars of crypto-grade randomness (the\n/// `api.token` recipe — spt-proto's Ed25519 seed generator, no extra RNG dep).\npub fn mint_link_token() -> String {\n    let seed = spt_proto::identity::Identity::generate().seed();\n    seed[..16].iter().map(|b| format!(\"{b:02x}\")).collect()\n}\n\n/// Read a shell perch's parked link token.\npub fn read_link_token(shell_perch: &Path) -> Option<String> {\n    std::fs::read_to_string(shell_perch.join(LINK_TOKEN_FILE))\n        .ok()\n        .map(|s| s.trim().to_string())\n        .filter(|s| !s.is_empty())\n}\n\n/// The per-link channel key, derived from the minted token: `SHA-256(token)`.\n/// Both ends hold the token (spt parked it; the binary got it in argv), so\n/// both derive the same key without a second exchange.\n// [impl->REQ-SHELL-1]\npub fn link_key(token: &str) -> [u8; 32] {\n    let mut h = Sha256::new();\n    h.update(token.as_bytes());\n    h.finalize().into()\n}\n\n/// Authenticate one link frame: `HMAC-SHA256(link_key, frame)`, hex. Required\n/// on every message both directions (CONTEXT §local-link authentication) —\n/// the D3c channels stamp + verify this per frame; a frame without a valid\n/// MAC is dropped, never processed.\n// [impl->REQ-SHELL-1]\npub fn frame_mac(key: &[u8; 32], frame: &[u8]) -> String {\n    let mut mac = <Hmac<Sha256> as Mac>::new_from_slice(key).expect(\"hmac accepts any key len\");\n    mac.update(frame);\n    mac.finalize()\n        .into_bytes()\n        .iter()\n        .map(|b| format!(\"{b:02x}\"))\n        .collect()\n}\n\n/// Verify a frame's MAC (constant-time inner compare via the hmac crate).\npub fn verify_frame_mac(key: &[u8; 32], frame: &[u8], mac_hex: &str) -> bool {\n    let mut mac = <Hmac<Sha256> as Mac>::new_from_slice(key).expect(\"hmac accepts any key len\");\n    mac.update(frame);\n    let Ok(bytes) = hex_decode(mac_hex) else {\n        return false;\n    };\n    mac.verify_slice(&bytes).is_ok()\n}\n\nfn hex_decode(s: &str) -> Result<Vec<u8>, ()> {\n    if !s.len().is_multiple_of(2) {\n        return Err(());\n    }\n    (0..s.len())\n        .step_by(2)\n        .map(|i| u8::from_str_radix(&s[i..i + 2], 16).map_err(|_| ()))\n        .collect()\n}\n\n/// The install dir an adapter-option's `[shell]` commands resolve against — the\n/// parent adapter record's `source_dir` (the live manifest's own directory in\n/// pointer mode). `None` when the option does not resolve to an active record.\n///\n/// One helper for every shell-family site (spawn AND wake) precisely because\n/// \"which dir\" is the kind of thing that drifts when each site works it out for\n/// itself: a profiled option resolves against its PARENT's install dir, and a\n/// site that split the option differently would silently resolve a different\n/// binary. A `:profile` overlays configuration, never provenance — the bits are\n/// the parent's.\n// [impl->REQ-INSTALL-11]\npub fn shell_install_dir(adapters_dir: &Path, adapter_option: &str) -> Option<PathBuf> {\n    spt_runtime::registry::resolve_option(adapters_dir, adapter_option)\n        .ok()\n        .map(|(record, _)| PathBuf::from(record.source_dir))\n}\n\n/// Fill the manifest's opaque `spawn` template for one instance (mints + parks\n/// the link token first — the binary may bind immediately) and tokenize it.\n/// Shared by the detached and broker-hosted launch paths.\n///\n/// `install_dir` is the adapter's own directory ([`shell_install_dir`]). It does\n/// two things, both of which a `--release`-installed shell adapter cannot work\n/// without (REQ-INSTALL-11's documented follow-on sites):\n/// - the program token resolves through it, so a bare program name binds to the\n///   adapter's SHIPPED binary before `PATH` — the same\n///   [`spt_runtime::runtime::resolve_program_in_dir`] primitive `[digest]`, the\n///   session `role.command` and `[service]` use, never a second implementation;\n/// - `{adapter_dir}` becomes fillable, so a template can name a resource beside\n///   that binary.\n///\n/// `None` means \"no install dir known\" and is not the same as an empty one: the\n/// program token then passes through for `PATH` to resolve exactly as before, and\n/// a template naming `{adapter_dir}` fails LOUDLY on the missing key rather than\n/// filling an empty string into a path.\nfn fill_spawn_command(\n    owlery: &Path,\n    owner: &str,\n    shell_id: &str,\n    adapter_name: &str,\n    install_dir: Option<&Path>,\n    shell: &Shell,\n) -> Result<Vec<String>, String> {\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n    // Mutual exclusivity, the online side (M5-D4b): the watcher dies before\n    // the binary rises — and any (re)launch is \"shell activity\", so the\n    // crash give-up latch re-arms. [impl->REQ-SHELL-2]\n    crate::shellwake::kill_waker_at(&perch);\n    crate::shellwake::clear_gave_up(&perch);\n    // The keys the pending spool rows may be stamped under: the still-parked\n    // token of an abruptly-dead instance (the wake/relink path) AND the stash\n    // a previous rotation retired (covers the graceful-close→relink path plus\n    // one generation of race stragglers — see below). Read BEFORE the mint\n    // overwrites the parked file.\n    let parked_before_mint = read_link_token(&perch);\n    let retired_before_mint = std::fs::read_to_string(perch.join(RETIRED_TOKEN_FILE)).ok();\n    let token = mint_link_token();\n    // Park the token BEFORE the spawn: the binary may bind immediately.\n    std::fs::write(perch.join(LINK_TOKEN_FILE), &token).map_err(|e| format!(\"park token: {e}\"))?;\n    // Token rotation carries the durable channel with it: re-stamp every\n    // pending row from the old keys to the new, or the binary this launch\n    // raises would verify them against the fresh key and (correctly) discard —\n    // the wake-triggering command itself would be the one frame the wake loses\n    // (flynn's field catch, 2026-07-25: #23 armed the wake and never answered;\n    // #24/#19 spooled after the rotation and drained in order).\n    //\n    // Race + crash posture (todlando's A1/A2, 2026-07-25): a concurrent\n    // `drive_shell` that read the OLD parked token and spools after this walk\n    // finishes leaves a straggler row the walk never saw — so the old parked\n    // token is STASHED (not erased) and the drain paths give any\n    // retired-key row a second-chance re-stamp at delivery\n    // ([`crate::shellchan::restamp_for_drain`]). The walk itself is\n    // crash-safe and idempotent: each row rewrite is independent, a row\n    // already carrying the new stamp does not verify under an old key and is\n    // skipped, and a crash mid-walk leaves a mixed spool the NEXT rotation or\n    // drain converts (both old keys stay candidates). Rows are selected by\n    // MAC verification against the threaded new key — never by comparing a\n    // token snapshot (the D-2 stale-snapshot class).\n    // [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\n    let new_key = link_key(&token);\n    for old in [&parked_before_mint, &retired_before_mint].into_iter().flatten() {\n        crate::shellchan::restamp_pending_at(&perch, &link_key(old), &new_key);\n    }\n    // Stash the just-rotated-out token for the drain's second chance; a\n    // close-retired stash stays in place when this launch had no parked token\n    // (the close path spools nothing new — NoLink — so it has no stragglers).\n    if let Some(old) = &parked_before_mint {\n        let _ = std::fs::write(perch.join(RETIRED_TOKEN_FILE), old);\n    }\n\n    let keys = BTreeMap::from([\n        (\"id\".to_string(), shell_id.to_string()),\n        (\"adapter_name\".to_string(), adapter_name.to_string()),\n        (\"link_token\".to_string(), token),\n        // [impl->REQ-SHELL-PERCH-DIR] {perch_dir}: the shell perch root, so the\n        // binary can resolve a shell_file frame's PERCH-RELATIVE `path` attr\n        // (files/<xfer-id>-<name>) — the child inherits the BROKER's cwd, so\n        // without this no mechanical resolution exists. Opt-in and N-1-safe:\n        // a template that never names the key fills byte-identically. The\n        // frame's path deliberately STAYS relative (doyle ruling: an absolute\n        // path in a spooled frame lies across perch moves / node boundaries).\n        (\"perch_dir\".to_string(), perch.display().to_string()),\n    ]);\n    // [impl->REQ-INSTALL-11] {adapter_dir}: the adapter's own install dir, so a\n    // released shell adapter can point at a resource it ships. Opt-in and N-1-safe\n    // like {perch_dir} — a template that never names the key fills byte-identically.\n    let keys = match install_dir {\n        Some(dir) => {\n            let mut keys = keys;\n            keys.insert(\"adapter_dir\".to_string(), dir.display().to_string());\n            keys\n        }\n        None => keys,\n    };\n    // [impl->REQ-HAZARD-TEMPLATE-ARGV-FILL] tokenize-template-then-fill-each: a\n    // multi-word/quote/semicolon {key} value is exactly one argv element.\n    let mut tokens =\n        spt_runtime::runtime::fill_template_tokens(&shell.spawn, &keys).map_err(|e| e.to_string())?;\n    let Some(program) = tokens.first_mut() else {\n        return Err(\"empty spawn command\".into());\n    };\n    // [impl->REQ-INSTALL-11] one resolution primitive, no parallel path: a\n    // `--release`-installed adapter declaring a bare program name binds to its\n    // shipped binary before PATH. Without this such an adapter registers fine and\n    // then fails at spawn with os error 2 — the gap that forced a hand-maintained\n    // local manifest per node.\n    if let Some(dir) = install_dir {\n        *program = spt_runtime::runtime::resolve_program_in_dir(program, dir);\n    }\n    Ok(tokens)\n}\n\n/// WHO is asking for a launch — and therefore WHOSE Job Object the spawned\n/// binary would join (REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE).\n///\n/// A daemon-side spawn is job-neutral: the daemon escaped the launching job at\n/// its own cold start (the WMI → schtasks → breakaway ladder). A CLI-side spawn\n/// is not — `spt shell spawn`, local `spt shell relink` and the `shell cmd`\n/// wake-if-offline arm all call the launch IN the terminal-resident CLI process,\n/// so on a job that forbids breakaway the shell binary (and every descendant it\n/// ever spawns, relay hops included) inherits the terminal's job.\n// [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\n#[derive(Clone, Copy, Debug, PartialEq, Eq)]\npub enum LaunchOrigin {\n    /// The caller IS the daemon (broker/brain): spawn here, already job-neutral.\n    Daemon,\n    /// The caller is a terminal-resident CLI: hand the launch to a running\n    /// daemon when one answers, else spawn here and SAY whose job took the child.\n    Cli,\n}\n\n/// A routed launch's failure, split by what the caller may safely do next.\n// [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\n#[derive(Debug, Clone, PartialEq, Eq)]\npub enum LaunchFailure {\n    /// The launch was attempted and failed. Undo whatever the launch was for.\n    Failed(String),\n    /// The launch was handed off and its outcome is UNKNOWN — the binary may be\n    /// running. Neither retry nor undo: both are wrong half the time.\n    Indeterminate(String),\n}\n\nimpl std::fmt::Display for LaunchFailure {\n    fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {\n        match self {\n            LaunchFailure::Failed(m) | LaunchFailure::Indeterminate(m) => write!(f, \"{m}\"),\n        }\n    }\n}\n\nimpl LaunchFailure {\n    /// Whether the outcome is unknown — the caller must keep what it built.\n    // [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\n    pub fn is_indeterminate(&self) -> bool {\n        matches!(self, LaunchFailure::Indeterminate(_))\n    }\n}\n\n/// Whether this shell's launch is the one that can be routed to the daemon: the\n/// DETACHED relay spawn, the only receipt that spawns in the calling process.\n///\n/// `stdin` already spawns broker-side (so it is job-neutral today, and routing\n/// it would make the broker connect to its own socket); `http` refuses at spawn.\n// [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\npub fn routes_through_daemon(shell: &Shell) -> bool {\n    !matches!(\n        shell.command_receipt,\n        Some(CommandReceipt::Http) | Some(CommandReceipt::Stdin)\n    )\n}\n\n/// The double-launch arbiter: the instance's recorded pid, when it probes\n/// ALIVE (REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE).\n///\n/// `Some(pid)` ⇒ a launch of this same `{owner, shell_id}` already won — stand\n/// down and adopt it. Used by BOTH sides of the routed launch, because a late\n/// ack can put the caller and the daemon in the same launch at once.\n///\n/// Birth-safe, including a binary awaiting bind: a recycled live pid is NOT a\n/// winner when its native birth stamp mismatches. Missing, unparseable and zero\n/// pids remain not-a-winner, preserving pid-less backend behavior.\n// [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\npub fn live_launch_winner(shell_perch: &Path) -> Option<u32> {\n    shellinfo::read_shell_pid(shell_perch)\n        .filter(|&p| p != 0 && !shellinfo::shell_pid_provably_dead(shell_perch))\n}\n\n/// THE route-selection seam. Generic over the daemon call and the in-process\n/// fallback so the decision can be tested without a socket, a daemon or a job.\n///\n/// The invariant it exists to hold: **the ack is a commitment transfer** — a\n/// daemon that accepted owns the launch, so only a pre-ack refusal\n/// ([`LaunchRouteError::fallback_allowed`](crate::drivehub::LaunchRouteError::fallback_allowed))\n/// may fall through to an in-process spawn. Anything else returning a second\n/// spawn would be the double-launch this routing exists to avoid.\n///\n/// Two axes, deliberately separate: `fallback_allowed` decides whether a second\n/// spawn is PERMITTED, and\n/// [`may_have_launched`](crate::drivehub::LaunchRouteError::may_have_launched)\n/// decides whether one must first be ARBITRATED against a live pid. Only the\n/// unanswered wait is both.\n// [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\npub(crate) fn launch_routed(\n    shell: &Shell,\n    label: &str,\n    via_daemon: impl FnOnce() -> Result<u32, crate::drivehub::LaunchRouteError>,\n    already_live: impl FnOnce() -> Option<u32>,\n    in_process: impl FnOnce() -> Result<u32, String>,\n) -> Result<u32, LaunchFailure> {\n    if !routes_through_daemon(shell) {\n        return in_process().map_err(LaunchFailure::Failed);\n    }\n    match via_daemon() {\n        Ok(pid) => Ok(pid),\n        Err(e) if e.fallback_allowed() => {\n            // A pre-ack refusal means nothing was spawned — EXCEPT for the one\n            // pre-ack cause that is a timeout rather than an answer. A daemon\n            // that acks late may already have launched, so THAT fallback checks\n            // the arbiter before adding a second binary (the documented race at\n            // `LAUNCH_ACK_BOUND`). Adopting the winner is a SUCCESS: the caller\n            // wanted a running binary and there is one.\n            //\n            // The gate is `may_have_launched`, not `fallback_allowed`: running\n            // the arbiter on an ANSWERED refusal adopts a pid no daemon could\n            // have launched, so the fallback silently declines to spawn — and a\n            // relink that never re-launches never mints, which is precisely the\n            // recovery command failing on the state it exists to recover.\n            // [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\n            // Gated by SHORT-CIRCUIT, not by filtering the answer: outside the\n            // late-ack window the pid probe is not merely unused, it must not\n            // run at all — that is one fewer place a pinned corpse handle can\n            // report a dead pid alive.\n            let winner = if e.may_have_launched() {\n                already_live()\n            } else {\n                None\n            };\n            if let Some(pid) = winner {\n                spt_proto::emit_line_err!(\n                    \"SHELL_LAUNCH_ALREADY_LIVE:{label}: a launch of this instance is already \\\n                     running (pid={pid}) — standing down instead of spawning a second binary\"\n                );\n                return Ok(pid);\n            }\n            // Nothing was spawned daemon-side. Say WHY the route was not taken\n            // and WHOSE job the child is about to join — the three pre-ack\n            // causes carry distinct names, so a field report can tell an old\n            // daemon from a foreign home from no daemon at all.\n            spt_proto::emit_line_err!(\n                \"SHELL_LAUNCH_IN_PROCESS:{label}: {e} — spawning from THIS process instead, so \\\n                 the binary joins THIS process's Job Object (a DETACH_BREAKAWAY_DENIED line \\\n                 names it); the first daemon-side wake re-spawns it job-neutrally\"\n            );\n            in_process().map_err(LaunchFailure::Failed)\n        }\n        Err(e) if e.is_indeterminate() => Err(LaunchFailure::Indeterminate(e.to_string())),\n        Err(e) => Err(LaunchFailure::Failed(e.to_string())),\n    }\n}\n\n/// [`launch_shell`], routed by ORIGIN (REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE).\n///\n/// From the daemon this is `launch_shell` verbatim. From a CLI it hands the\n/// launch to a running daemon over the drive control socket — so the binary\n/// rises job-neutral instead of inside the launching terminal's Job Object —\n/// and falls back to today's in-process spawn, loudly, when no daemon takes it.\n// [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\npub fn launch_shell_from(\n    origin: LaunchOrigin,\n    owlery: &Path,\n    owner: &str,\n    shell_id: &str,\n    adapter_name: &str,\n    install_dir: Option<&Path>,\n    shell: &Shell,\n) -> Result<u32, LaunchFailure> {\n    match origin {\n        LaunchOrigin::Daemon => {\n            launch_shell(owlery, owner, shell_id, adapter_name, install_dir, shell)\n                .map_err(LaunchFailure::Failed)\n        }\n        LaunchOrigin::Cli => launch_routed(\n            shell,\n            shell_id,\n            || {\n                crate::drivehub::shell_launch(\n                    &crate::endpoint::drive_socket_name(),\n                    owner,\n                    shell_id,\n                    owlery,\n                )\n            },\n            || {\n                live_launch_winner(&spt_store::perch::resolve_shell_perch_path_in(\n                    owlery, owner, shell_id,\n                ))\n            },\n            || launch_shell(owlery, owner, shell_id, adapter_name, install_dir, shell),\n        ),\n    }\n}\n\n/// Launch a shell instance's binary per the manifest's `command_receipt`\n/// (M5-D3c): **relay** (default) spawns detached (stdio discarded — the\n/// binary drains via `api poll --link`); **stdin** spawns broker-hosted so\n/// the daemon owns the stdin pipe ([`launch_shell_brokered_in`] — command\n/// frames ride the broker's `send_effect` input seam); **http** refuses at\n/// spawn, explicitly — an unimplemented mode is never a silent fallback.\n/// Records the pid; returns it. The perch stays **offline** until the\n/// binary's `api bind-shell` handshake onlines it — launch is not the online\n/// switch any more than spawn is.\n// [impl->REQ-SHELL-1]\npub fn launch_shell(\n    owlery: &Path,\n    owner: &str,\n    shell_id: &str,\n    adapter_name: &str,\n    install_dir: Option<&Path>,\n    shell: &Shell,\n) -> Result<u32, String> {\n    match shell.command_receipt {\n        Some(CommandReceipt::Http) => Err(\n            \"command_receipt=http is not implemented yet — spawn refused (an \\\n             unimplemented mode never falls back silently)\"\n                .into(),\n        ),\n        Some(CommandReceipt::Stdin) => {\n            crate::ensure_running().map_err(|e| format!(\"daemon start: {e}\"))?;\n            launch_shell_brokered_in(\n                &crate::broker_socket_name(),\n                owlery,\n                owner,\n                shell_id,\n                adapter_name,\n                install_dir,\n                shell,\n            )\n        }\n        _ => {\n            let tokens =\n                fill_spawn_command(owlery, owner, shell_id, adapter_name, install_dir, shell)?;\n            let (program, args) = tokens.split_first().expect(\"non-empty checked\");\n            // The shell binary is a LONG-LIVED detached child (KNOWN-HAZARDS\n            // 5.6): on Windows it must inherit NO handles, or a caller\n            // capturing `spt shell spawn`'s output — at any depth of the\n            // process chain — hangs forever on a pipe write-end the immortal\n            // child holds. Same no-inherit spawn as the daemon's.\n            // [impl->REQ-HAZARD-DETACHED-PIPE-INHERIT]\n            #[cfg(windows)]\n            let pid = crate::daemon::detached_no_inherit(program, args)\n                .map_err(|e| format!(\"spawn {program}: {e}\"))?;\n            #[cfg(unix)]\n            let pid = Command::new(program)\n                .args(args)\n                .stdin(Stdio::null())\n                .stdout(Stdio::null())\n                .stderr(Stdio::null())\n                .spawn()\n                .map_err(|e| format!(\"spawn {program}: {e}\"))?\n                .id();\n            let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n            std::fs::write(perch.join(SHELL_PID_FILE), pid.to_string())\n                .map_err(|e| format!(\"record pid: {e}\"))?;\n            // The birth stamp rides WITH the pid at every launch, or the probe that\n            // reads it degrades to bare-pid liveness on this instance alone -- a\n            // silent per-instance regression rather than a visible one.\n            // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n            spt_store::shellinfo::record_shell_launch(&perch, pid, now_ms());\n            Ok(pid)\n        }\n    }\n}\n\n/// The stdin-receipt launch: spawn the filled template as a **broker-hosted**\n/// session labeled [`shell_session_label`], so the daemon owns the child's\n/// stdin and `spt shell cmd` can drive it through the `send_effect` input\n\n[Showing lines 1-513 of 1826. Use :514 to continue]"}],"details":{"truncation":{"content":"//! Shell-instance hosting (M5-D3b — CONTEXT §Shell model, §local-link\r\n//! authentication): the broker-launch + local-link handshake half of the\r\n//! shell machinery.\r\n//!\r\n//! **Broker-launched, never user-launched.** [`launch_shell`] mints a\r\n//! per-link token, parks it on the shell perch, fills the manifest's opaque\r\n//! `spawn` template (`{link_token}` + `{id}`/`{adapter_name}` from the\r\n//! guaranteed catalog) and spawns the binary detached — the only way a shell\r\n//! binary ever starts. Launch failure rolls nothing forward: the caller\r\n//! tears the just-minted perch back down (no half-built instance).\r\n//!\r\n//! **The local-link handshake.** Cross-node traffic already rides Iroh\r\n//! (E2E-encrypted); the exposed surface is the *same-node* channel where the\r\n//! least-trusted code in the system (the shell binary) touches SPT. The\r\n//! handshake: possession of the minted token **is** the bind credential —\r\n//! `spt api bind-shell --link <token>` resolves the instance *by token*\r\n//! ([`bind_shell_by_token`]: the binary needs no owner/id knowledge beyond\r\n//! its argv) and flips the perch online. Every subsequent machinery message\r\n//! on the link authenticates with the token-derived key ([`link_key`] →\r\n//! [`frame_mac`], applied per-message by the D3c channels) so another local\r\n//! process can neither inject into nor spoof a link.\r\n// [impl->REQ-SHELL-1]\r\n\r\nuse std::collections::BTreeMap;\r\nuse std::path::{Path, PathBuf};\r\nuse std::process::{Command, Stdio};\r\n\r\nuse hmac::{Hmac, Mac};\r\nuse sha2::{Digest, Sha256};\r\n\r\nuse spt_runtime::manifest::{AdapterKind, CommandReceipt, Manifest, Shell};\r\nuse spt_runtime::registry::AdapterRecord;\r\nuse spt_store::liveness::RelayLiveness;\r\nuse spt_store::shellinfo::{self, SHELL_STATUS_OFFLINE, SHELL_STATUS_ONLINE};\r\n\r\nuse crate::brain::{now_ms, Brain};\r\nuse crate::msg::SpawnReq;\r\nuse crate::shellchan::shell_session_label;\r\n\r\n/// Per-instance link-token filename within a shell perch dir (node-local\r\n/// runtime material, beside `info.json` — the perch *record* contract in\r\n/// STORAGE.md constrains `info.json`'s fields, not the perch's files).\r\npub const LINK_TOKEN_FILE: &str = \"link.token\";\r\n\r\n/// The most recently RETIRED link token (stashed by `close_shell`, consumed by\r\n/// the next launch's re-stamp). Never a live credential — bind resolves only\r\n/// [`LINK_TOKEN_FILE`]; this exists so pending spool rows stamped under the\r\n/// dead link survive the next token mint (REQ-HAZARD-SHELL-STALE-ONLINE).\r\npub const RETIRED_TOKEN_FILE: &str = \"link.token.retired\";\r\n/// The launched binary's OS pid — teardown's kill handle, and the evidence the\r\n/// derived-status resolver probes. Re-exported, never re-typed: the name is\r\n/// written here (launch/close) and MATCHED there\r\n/// ([`shellinfo::shell_pid_provably_dead`]), so both sides must be one const.\r\n// [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\r\npub use spt_store::shellinfo::SHELL_PID_FILE;\r\n\r\n/// Mint a fresh link token: 32 hex chars of crypto-grade randomness (the\r\n/// `api.token` recipe — spt-proto's Ed25519 seed generator, no extra RNG dep).\r\npub fn mint_link_token() -> String {\r\n    let seed = spt_proto::identity::Identity::generate().seed();\r\n    seed[..16].iter().map(|b| format!(\"{b:02x}\")).collect()\r\n}\r\n\r\n/// Read a shell perch's parked link token.\r\npub fn read_link_token(shell_perch: &Path) -> Option<String> {\r\n    std::fs::read_to_string(shell_perch.join(LINK_TOKEN_FILE))\r\n        .ok()\r\n        .map(|s| s.trim().to_string())\r\n        .filter(|s| !s.is_empty())\r\n}\r\n\r\n/// The per-link channel key, derived from the minted token: `SHA-256(token)`.\r\n/// Both ends hold the token (spt parked it; the binary got it in argv), so\r\n/// both derive the same key without a second exchange.\r\n// [impl->REQ-SHELL-1]\r\npub fn link_key(token: &str) -> [u8; 32] {\r\n    let mut h = Sha256::new();\r\n    h.update(token.as_bytes());\r\n    h.finalize().into()\r\n}\r\n\r\n/// Authenticate one link frame: `HMAC-SHA256(link_key, frame)`, hex. Required\r\n/// on every message both directions (CONTEXT §local-link authentication) —\r\n/// the D3c channels stamp + verify this per frame; a frame without a valid\r\n/// MAC is dropped, never processed.\r\n// [impl->REQ-SHELL-1]\r\npub fn frame_mac(key: &[u8; 32], frame: &[u8]) -> String {\r\n    let mut mac = <Hmac<Sha256> as Mac>::new_from_slice(key).expect(\"hmac accepts any key len\");\r\n    mac.update(frame);\r\n    mac.finalize()\r\n        .into_bytes()\r\n        .iter()\r\n        .map(|b| format!(\"{b:02x}\"))\r\n        .collect()\r\n}\r\n\r\n/// Verify a frame's MAC (constant-time inner compare via the hmac crate).\r\npub fn verify_frame_mac(key: &[u8; 32], frame: &[u8], mac_hex: &str) -> bool {\r\n    let mut mac = <Hmac<Sha256> as Mac>::new_from_slice(key).expect(\"hmac accepts any key len\");\r\n    mac.update(frame);\r\n    let Ok(bytes) = hex_decode(mac_hex) else {\r\n        return false;\r\n    };\r\n    mac.verify_slice(&bytes).is_ok()\r\n}\r\n\r\nfn hex_decode(s: &str) -> Result<Vec<u8>, ()> {\r\n    if !s.len().is_multiple_of(2) {\r\n        return Err(());\r\n    }\r\n    (0..s.len())\r\n        .step_by(2)\r\n        .map(|i| u8::from_str_radix(&s[i..i + 2], 16).map_err(|_| ()))\r\n        .collect()\r\n}\r\n\r\n/// The install dir an adapter-option's `[shell]` commands resolve against — the\r\n/// parent adapter record's `source_dir` (the live manifest's own directory in\r\n/// pointer mode). `None` when the option does not resolve to an active record.\r\n///\r\n/// One helper for every shell-family site (spawn AND wake) precisely because\r\n/// \"which dir\" is the kind of thing that drifts when each site works it out for\r\n/// itself: a profiled option resolves against its PARENT's install dir, and a\r\n/// site that split the option differently would silently resolve a different\r\n/// binary. A `:profile` overlays configuration, never provenance — the bits are\r\n/// the parent's.\r\n// [impl->REQ-INSTALL-11]\r\npub fn shell_install_dir(adapters_dir: &Path, adapter_option: &str) -> Option<PathBuf> {\r\n    spt_runtime::registry::resolve_option(adapters_dir, adapter_option)\r\n        .ok()\r\n        .map(|(record, _)| PathBuf::from(record.source_dir))\r\n}\r\n\r\n/// Fill the manifest's opaque `spawn` template for one instance (mints + parks\r\n/// the link token first — the binary may bind immediately) and tokenize it.\r\n/// Shared by the detached and broker-hosted launch paths.\r\n///\r\n/// `install_dir` is the adapter's own directory ([`shell_install_dir`]). It does\r\n/// two things, both of which a `--release`-installed shell adapter cannot work\r\n/// without (REQ-INSTALL-11's documented follow-on sites):\r\n/// - the program token resolves through it, so a bare program name binds to the\r\n///   adapter's SHIPPED binary before `PATH` — the same\r\n///   [`spt_runtime::runtime::resolve_program_in_dir`] primitive `[digest]`, the\r\n///   session `role.command` and `[service]` use, never a second implementation;\r\n/// - `{adapter_dir}` becomes fillable, so a template can name a resource beside\r\n///   that binary.\r\n///\r\n/// `None` means \"no install dir known\" and is not the same as an empty one: the\r\n/// program token then passes through for `PATH` to resolve exactly as before, and\r\n/// a template naming `{adapter_dir}` fails LOUDLY on the missing key rather than\r\n/// filling an empty string into a path.\r\nfn fill_spawn_command(\r\n    owlery: &Path,\r\n    owner: &str,\r\n    shell_id: &str,\r\n    adapter_name: &str,\r\n    install_dir: Option<&Path>,\r\n    shell: &Shell,\r\n) -> Result<Vec<String>, String> {\r\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\r\n    // Mutual exclusivity, the online side (M5-D4b): the watcher dies before\r\n    // the binary rises — and any (re)launch is \"shell activity\", so the\r\n    // crash give-up latch re-arms. [impl->REQ-SHELL-2]\r\n    crate::shellwake::kill_waker_at(&perch);\r\n    crate::shellwake::clear_gave_up(&perch);\r\n    // The keys the pending spool rows may be stamped under: the still-parked\r\n    // token of an abruptly-dead instance (the wake/relink path) AND the stash\r\n    // a previous rotation retired (covers the graceful-close→relink path plus\r\n    // one generation of race stragglers — see below). Read BEFORE the mint\r\n    // overwrites the parked file.\r\n    let parked_before_mint = read_link_token(&perch);\r\n    let retired_before_mint = std::fs::read_to_string(perch.join(RETIRED_TOKEN_FILE)).ok();\r\n    let token = mint_link_token();\r\n    // Park the token BEFORE the spawn: the binary may bind immediately.\r\n    std::fs::write(perch.join(LINK_TOKEN_FILE), &token).map_err(|e| format!(\"park token: {e}\"))?;\r\n    // Token rotation carries the durable channel with it: re-stamp every\r\n    // pending row from the old keys to the new, or the binary this launch\r\n    // raises would verify them against the fresh key and (correctly) discard —\r\n    // the wake-triggering command itself would be the one frame the wake loses\r\n    // (flynn's field catch, 2026-07-25: #23 armed the wake and never answered;\r\n    // #24/#19 spooled after the rotation and drained in order).\r\n    //\r\n    // Race + crash posture (todlando's A1/A2, 2026-07-25): a concurrent\r\n    // `drive_shell` that read the OLD parked token and spools after this walk\r\n    // finishes leaves a straggler row the walk never saw — so the old parked\r\n    // token is STASHED (not erased) and the drain paths give any\r\n    // retired-key row a second-chance re-stamp at delivery\r\n    // ([`crate::shellchan::restamp_for_drain`]). The walk itself is\r\n    // crash-safe and idempotent: each row rewrite is independent, a row\r\n    // already carrying the new stamp does not verify under an old key and is\r\n    // skipped, and a crash mid-walk leaves a mixed spool the NEXT rotation or\r\n    // drain converts (both old keys stay candidates). Rows are selected by\r\n    // MAC verification against the threaded new key — never by comparing a\r\n    // token snapshot (the D-2 stale-snapshot class).\r\n    // [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\r\n    let new_key = link_key(&token);\r\n    for old in [&parked_before_mint, &retired_before_mint].into_iter().flatten() {\r\n        crate::shellchan::restamp_pending_at(&perch, &link_key(old), &new_key);\r\n    }\r\n    // Stash the just-rotated-out token for the drain's second chance; a\r\n    // close-retired stash stays in place when this launch had no parked token\r\n    // (the close path spools nothing new — NoLink — so it has no stragglers).\r\n    if let Some(old) = &parked_before_mint {\r\n        let _ = std::fs::write(perch.join(RETIRED_TOKEN_FILE), old);\r\n    }\r\n\r\n    let keys = BTreeMap::from([\r\n        (\"id\".to_string(), shell_id.to_string()),\r\n        (\"adapter_name\".to_string(), adapter_name.to_string()),\r\n        (\"link_token\".to_string(), token),\r\n        // [impl->REQ-SHELL-PERCH-DIR] {perch_dir}: the shell perch root, so the\r\n        // binary can resolve a shell_file frame's PERCH-RELATIVE `path` attr\r\n        // (files/<xfer-id>-<name>) — the child inherits the BROKER's cwd, so\r\n        // without this no mechanical resolution exists. Opt-in and N-1-safe:\r\n        // a template that never names the key fills byte-identically. The\r\n        // frame's path deliberately STAYS relative (doyle ruling: an absolute\r\n        // path in a spooled frame lies across perch moves / node boundaries).\r\n        (\"perch_dir\".to_string(), perch.display().to_string()),\r\n    ]);\r\n    // [impl->REQ-INSTALL-11] {adapter_dir}: the adapter's own install dir, so a\r\n    // released shell adapter can point at a resource it ships. Opt-in and N-1-safe\r\n    // like {perch_dir} — a template that never names the key fills byte-identically.\r\n    let keys = match install_dir {\r\n        Some(dir) => {\r\n            let mut keys = keys;\r\n            keys.insert(\"adapter_dir\".to_string(), dir.display().to_string());\r\n            keys\r\n        }\r\n        None => keys,\r\n    };\r\n    // [impl->REQ-HAZARD-TEMPLATE-ARGV-FILL] tokenize-template-then-fill-each: a\r\n    // multi-word/quote/semicolon {key} value is exactly one argv element.\r\n    let mut tokens =\r\n        spt_runtime::runtime::fill_template_tokens(&shell.spawn, &keys).map_err(|e| e.to_string())?;\r\n    let Some(program) = tokens.first_mut() else {\r\n        return Err(\"empty spawn command\".into());\r\n    };\r\n    // [impl->REQ-INSTALL-11] one resolution primitive, no parallel path: a\r\n    // `--release`-installed adapter declaring a bare program name binds to its\r\n    // shipped binary before PATH. Without this such an adapter registers fine and\r\n    // then fails at spawn with os error 2 — the gap that forced a hand-maintained\r\n    // local manifest per node.\r\n    if let Some(dir) = install_dir {\r\n        *program = spt_runtime::runtime::resolve_program_in_dir(program, dir);\r\n    }\r\n    Ok(tokens)\r\n}\r\n\r\n/// WHO is asking for a launch — and therefore WHOSE Job Object the spawned\r\n/// binary would join (REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE).\r\n///\r\n/// A daemon-side spawn is job-neutral: the daemon escaped the launching job at\r\n/// its own cold start (the WMI → schtasks → breakaway ladder). A CLI-side spawn\r\n/// is not — `spt shell spawn`, local `spt shell relink` and the `shell cmd`\r\n/// wake-if-offline arm all call the launch IN the terminal-resident CLI process,\r\n/// so on a job that forbids breakaway the shell binary (and every descendant it\r\n/// ever spawns, relay hops included) inherits the terminal's job.\r\n// [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\r\n#[derive(Clone, Copy, Debug, PartialEq, Eq)]\r\npub enum LaunchOrigin {\r\n    /// The caller IS the daemon (broker/brain): spawn here, already job-neutral.\r\n    Daemon,\r\n    /// The caller is a terminal-resident CLI: hand the launch to a running\r\n    /// daemon when one answers, else spawn here and SAY whose job took the child.\r\n    Cli,\r\n}\r\n\r\n/// A routed launch's failure, split by what the caller may safely do next.\r\n// [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\r\n#[derive(Debug, Clone, PartialEq, Eq)]\r\npub enum LaunchFailure {\r\n    /// The launch was attempted and failed. Undo whatever the launch was for.\r\n    Failed(String),\r\n    /// The launch was handed off and its outcome is UNKNOWN — the binary may be\r\n    /// running. Neither retry nor undo: both are wrong half the time.\r\n    Indeterminate(String),\r\n}\r\n\r\nimpl std::fmt::Display for LaunchFailure {\r\n    fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {\r\n        match self {\r\n            LaunchFailure::Failed(m) | LaunchFailure::Indeterminate(m) => write!(f, \"{m}\"),\r\n        }\r\n    }\r\n}\r\n\r\nimpl LaunchFailure {\r\n    /// Whether the outcome is unknown — the caller must keep what it built.\r\n    // [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\r\n    pub fn is_indeterminate(&self) -> bool {\r\n        matches!(self, LaunchFailure::Indeterminate(_))\r\n    }\r\n}\r\n\r\n/// Whether this shell's launch is the one that can be routed to the daemon: the\r\n/// DETACHED relay spawn, the only receipt that spawns in the calling process.\r\n///\r\n/// `stdin` already spawns broker-side (so it is job-neutral today, and routing\r\n/// it would make the broker connect to its own socket); `http` refuses at spawn.\r\n// [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\r\npub fn routes_through_daemon(shell: &Shell) -> bool {\r\n    !matches!(\r\n        shell.command_receipt,\r\n        Some(CommandReceipt::Http) | Some(CommandReceipt::Stdin)\r\n    )\r\n}\r\n\r\n/// The double-launch arbiter: the instance's recorded pid, when it probes\r\n/// ALIVE (REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE).\r\n///\r\n/// `Some(pid)` ⇒ a launch of this same `{owner, shell_id}` already won — stand\r\n/// down and adopt it. Used by BOTH sides of the routed launch, because a late\r\n/// ack can put the caller and the daemon in the same launch at once.\r\n///\r\n/// Birth-safe, including a binary awaiting bind: a recycled live pid is NOT a\r\n/// winner when its native birth stamp mismatches. Missing, unparseable and zero\r\n/// pids remain not-a-winner, preserving pid-less backend behavior.\r\n// [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\r\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\npub fn live_launch_winner(shell_perch: &Path) -> Option<u32> {\r\n    shellinfo::read_shell_pid(shell_perch)\r\n        .filter(|&p| p != 0 && !shellinfo::shell_pid_provably_dead(shell_perch))\r\n}\r\n\r\n/// THE route-selection seam. Generic over the daemon call and the in-process\r\n/// fallback so the decision can be tested without a socket, a daemon or a job.\r\n///\r\n/// The invariant it exists to hold: **the ack is a commitment transfer** — a\r\n/// daemon that accepted owns the launch, so only a pre-ack refusal\r\n/// ([`LaunchRouteError::fallback_allowed`](crate::drivehub::LaunchRouteError::fallback_allowed))\r\n/// may fall through to an in-process spawn. Anything else returning a second\r\n/// spawn would be the double-launch this routing exists to avoid.\r\n///\r\n/// Two axes, deliberately separate: `fallback_allowed` decides whether a second\r\n/// spawn is PERMITTED, and\r\n/// [`may_have_launched`](crate::drivehub::LaunchRouteError::may_have_launched)\r\n/// decides whether one must first be ARBITRATED against a live pid. Only the\r\n/// unanswered wait is both.\r\n// [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\r\npub(crate) fn launch_routed(\r\n    shell: &Shell,\r\n    label: &str,\r\n    via_daemon: impl FnOnce() -> Result<u32, crate::drivehub::LaunchRouteError>,\r\n    already_live: impl FnOnce() -> Option<u32>,\r\n    in_process: impl FnOnce() -> Result<u32, String>,\r\n) -> Result<u32, LaunchFailure> {\r\n    if !routes_through_daemon(shell) {\r\n        return in_process().map_err(LaunchFailure::Failed);\r\n    }\r\n    match via_daemon() {\r\n        Ok(pid) => Ok(pid),\r\n        Err(e) if e.fallback_allowed() => {\r\n            // A pre-ack refusal means nothing was spawned — EXCEPT for the one\r\n            // pre-ack cause that is a timeout rather than an answer. A daemon\r\n            // that acks late may already have launched, so THAT fallback checks\r\n            // the arbiter before adding a second binary (the documented race at\r\n            // `LAUNCH_ACK_BOUND`). Adopting the winner is a SUCCESS: the caller\r\n            // wanted a running binary and there is one.\r\n            //\r\n            // The gate is `may_have_launched`, not `fallback_allowed`: running\r\n            // the arbiter on an ANSWERED refusal adopts a pid no daemon could\r\n            // have launched, so the fallback silently declines to spawn — and a\r\n            // relink that never re-launches never mints, which is precisely the\r\n            // recovery command failing on the state it exists to recover.\r\n            // [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\r\n            // Gated by SHORT-CIRCUIT, not by filtering the answer: outside the\r\n            // late-ack window the pid probe is not merely unused, it must not\r\n            // run at all — that is one fewer place a pinned corpse handle can\r\n            // report a dead pid alive.\r\n            let winner = if e.may_have_launched() {\r\n                already_live()\r\n            } else {\r\n                None\r\n            };\r\n            if let Some(pid) = winner {\r\n                spt_proto::emit_line_err!(\r\n                    \"SHELL_LAUNCH_ALREADY_LIVE:{label}: a launch of this instance is already \\\r\n                     running (pid={pid}) — standing down instead of spawning a second binary\"\r\n                );\r\n                return Ok(pid);\r\n            }\r\n            // Nothing was spawned daemon-side. Say WHY the route was not taken\r\n            // and WHOSE job the child is about to join — the three pre-ack\r\n            // causes carry distinct names, so a field report can tell an old\r\n            // daemon from a foreign home from no daemon at all.\r\n            spt_proto::emit_line_err!(\r\n                \"SHELL_LAUNCH_IN_PROCESS:{label}: {e} — spawning from THIS process instead, so \\\r\n                 the binary joins THIS process's Job Object (a DETACH_BREAKAWAY_DENIED line \\\r\n                 names it); the first daemon-side wake re-spawns it job-neutrally\"\r\n            );\r\n            in_process().map_err(LaunchFailure::Failed)\r\n        }\r\n        Err(e) if e.is_indeterminate() => Err(LaunchFailure::Indeterminate(e.to_string())),\r\n        Err(e) => Err(LaunchFailure::Failed(e.to_string())),\r\n    }\r\n}\r\n\r\n/// [`launch_shell`], routed by ORIGIN (REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE).\r\n///\r\n/// From the daemon this is `launch_shell` verbatim. From a CLI it hands the\r\n/// launch to a running daemon over the drive control socket — so the binary\r\n/// rises job-neutral instead of inside the launching terminal's Job Object —\r\n/// and falls back to today's in-process spawn, loudly, when no daemon takes it.\r\n// [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\r\npub fn launch_shell_from(\r\n    origin: LaunchOrigin,\r\n    owlery: &Path,\r\n    owner: &str,\r\n    shell_id: &str,\r\n    adapter_name: &str,\r\n    install_dir: Option<&Path>,\r\n    shell: &Shell,\r\n) -> Result<u32, LaunchFailure> {\r\n    match origin {\r\n        LaunchOrigin::Daemon => {\r\n            launch_shell(owlery, owner, shell_id, adapter_name, install_dir, shell)\r\n                .map_err(LaunchFailure::Failed)\r\n        }\r\n        LaunchOrigin::Cli => launch_routed(\r\n            shell,\r\n            shell_id,\r\n            || {\r\n                crate::drivehub::shell_launch(\r\n                    &crate::endpoint::drive_socket_name(),\r\n                    owner,\r\n                    shell_id,\r\n                    owlery,\r\n                )\r\n            },\r\n            || {\r\n                live_launch_winner(&spt_store::perch::resolve_shell_perch_path_in(\r\n                    owlery, owner, shell_id,\r\n                ))\r\n            },\r\n            || launch_shell(owlery, owner, shell_id, adapter_name, install_dir, shell),\r\n        ),\r\n    }\r\n}\r\n\r\n/// Launch a shell instance's binary per the manifest's `command_receipt`\r\n/// (M5-D3c): **relay** (default) spawns detached (stdio discarded — the\r\n/// binary drains via `api poll --link`); **stdin** spawns broker-hosted so\r\n/// the daemon owns the stdin pipe ([`launch_shell_brokered_in`] — command\r\n/// frames ride the broker's `send_effect` input seam); **http** refuses at\r\n/// spawn, explicitly — an unimplemented mode is never a silent fallback.\r\n/// Records the pid; returns it. The perch stays **offline** until the\r\n/// binary's `api bind-shell` handshake onlines it — launch is not the online\r\n/// switch any more than spawn is.\r\n// [impl->REQ-SHELL-1]\r\npub fn launch_shell(\r\n    owlery: &Path,\r\n    owner: &str,\r\n    shell_id: &str,\r\n    adapter_name: &str,\r\n    install_dir: Option<&Path>,\r\n    shell: &Shell,\r\n) -> Result<u32, String> {\r\n    match shell.command_receipt {\r\n        Some(CommandReceipt::Http) => Err(\r\n            \"command_receipt=http is not implemented yet — spawn refused (an \\\r\n             unimplemented mode never falls back silently)\"\r\n                .into(),\r\n        ),\r\n        Some(CommandReceipt::Stdin) => {\r\n            crate::ensure_running().map_err(|e| format!(\"daemon start: {e}\"))?;\r\n            launch_shell_brokered_in(\r\n                &crate::broker_socket_name(),\r\n                owlery,\r\n                owner,\r\n                shell_id,\r\n                adapter_name,\r\n                install_dir,\r\n                shell,\r\n            )\r\n        }\r\n        _ => {\r\n            let tokens =\r\n                fill_spawn_command(owlery, owner, shell_id, adapter_name, install_dir, shell)?;\r\n            let (program, args) = tokens.split_first().expect(\"non-empty checked\");\r\n            // The shell binary is a LONG-LIVED detached child (KNOWN-HAZARDS\r\n            // 5.6): on Windows it must inherit NO handles, or a caller\r\n            // capturing `spt shell spawn`'s output — at any depth of the\r\n            // process chain — hangs forever on a pipe write-end the immortal\r\n            // child holds. Same no-inherit spawn as the daemon's.\r\n            // [impl->REQ-HAZARD-DETACHED-PIPE-INHERIT]\r\n            #[cfg(windows)]\r\n            let pid = crate::daemon::detached_no_inherit(program, args)\r\n                .map_err(|e| format!(\"spawn {program}: {e}\"))?;\r\n            #[cfg(unix)]\r\n            let pid = Command::new(program)\r\n                .args(args)\r\n                .stdin(Stdio::null())\r\n                .stdout(Stdio::null())\r\n                .stderr(Stdio::null())\r\n                .spawn()\r\n                .map_err(|e| format!(\"spawn {program}: {e}\"))?\r\n                .id();\r\n            let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\r\n            std::fs::write(perch.join(SHELL_PID_FILE), pid.to_string())\r\n                .map_err(|e| format!(\"record pid: {e}\"))?;\r\n            // The birth stamp rides WITH the pid at every launch, or the probe that\r\n            // reads it degrades to bare-pid liveness on this instance alone -- a\r\n            // silent per-instance regression rather than a visible one.\r\n            // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\n            spt_store::shellinfo::record_shell_launch(&perch, pid, now_ms());\r\n            Ok(pid)\r\n        }\r\n    }\r\n}\r\n\r\n/// The stdin-receipt launch: spawn the filled template as a **broker-hosted**\r\n/// session labeled [`shell_session_label`], so the daemon owns the child's\r\n/// stdin and `spt shell cmd` can drive it through the `send_effect` input\r","truncated":true,"truncatedBy":"lines","totalLines":1826,"totalBytes":25273,"outputLines":513,"outputBytes":25273,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":1826,"displayContent":{"text":"//! Shell-instance hosting (M5-D3b — CONTEXT §Shell model, §local-link\r\n//! authentication): the broker-launch + local-link handshake half of the\r\n//! shell machinery.\r\n//!\r\n//! **Broker-launched, never user-launched.** [`launch_shell`] mints a\r\n//! per-link token, parks it on the shell perch, fills the manifest's opaque\r\n//! `spawn` template (`{link_token}` + `{id}`/`{adapter_name}` from the\r\n//! guaranteed catalog) and spawns the binary detached — the only way a shell\r\n//! binary ever starts. Launch failure rolls nothing forward: the caller\r\n//! tears the just-minted perch back down (no half-built instance).\r\n//!\r\n//! **The local-link handshake.** Cross-node traffic already rides Iroh\r\n//! (E2E-encrypted); the exposed surface is the *same-node* channel where the\r\n//! least-trusted code in the system (the shell binary) touches SPT. The\r\n//! handshake: possession of the minted token **is** the bind credential —\r\n//! `spt api bind-shell --link <token>` resolves the instance *by token*\r\n//! ([`bind_shell_by_token`]: the binary needs no owner/id knowledge beyond\r\n//! its argv) and flips the perch online. Every subsequent machinery message\r\n//! on the link authenticates with the token-derived key ([`link_key`] →\r\n//! [`frame_mac`], applied per-message by the D3c channels) so another local\r\n//! process can neither inject into nor spoof a link.\r\n// [impl->REQ-SHELL-1]\r\n\r\nuse std::collections::BTreeMap;\r\nuse std::path::{Path, PathBuf};\r\nuse std::process::{Command, Stdio};\r\n\r\nuse hmac::{Hmac, Mac};\r\nuse sha2::{Digest, Sha256};\r\n\r\nuse spt_runtime::manifest::{AdapterKind, CommandReceipt, Manifest, Shell};\r\nuse spt_runtime::registry::AdapterRecord;\r\nuse spt_store::liveness::RelayLiveness;\r\nuse spt_store::shellinfo::{self, SHELL_STATUS_OFFLINE, SHELL_STATUS_ONLINE};\r\n\r\nuse crate::brain::{now_ms, Brain};\r\nuse crate::msg::SpawnReq;\r\nuse crate::shellchan::shell_session_label;\r\n\r\n/// Per-instance link-token filename within a shell perch dir (node-local\r\n/// runtime material, beside `info.json` — the perch *record* contract in\r\n/// STORAGE.md constrains `info.json`'s fields, not the perch's files).\r\npub const LINK_TOKEN_FILE: &str = \"link.token\";\r\n\r\n/// The most recently RETIRED link token (stashed by `close_shell`, consumed by\r\n/// the next launch's re-stamp). Never a live credential — bind resolves only\r\n/// [`LINK_TOKEN_FILE`]; this exists so pending spool rows stamped under the\r\n/// dead link survive the next token mint (REQ-HAZARD-SHELL-STALE-ONLINE).\r\npub const RETIRED_TOKEN_FILE: &str = \"link.token.retired\";\r\n/// The launched binary's OS pid — teardown's kill handle, and the evidence the\r\n/// derived-status resolver probes. Re-exported, never re-typed: the name is\r\n/// written here (launch/close) and MATCHED there\r\n/// ([`shellinfo::shell_pid_provably_dead`]), so both sides must be one const.\r\n// [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\r\npub use spt_store::shellinfo::SHELL_PID_FILE;\r\n\r\n/// Mint a fresh link token: 32 hex chars of crypto-grade randomness (the\r\n/// `api.token` recipe — spt-proto's Ed25519 seed generator, no extra RNG dep).\r\npub fn mint_link_token() -> String {\r\n    let seed = spt_proto::identity::Identity::generate().seed();\r\n    seed[..16].iter().map(|b| format!(\"{b:02x}\")).collect()\r\n}\r\n\r\n/// Read a shell perch's parked link token.\r\npub fn read_link_token(shell_perch: &Path) -> Option<String> {\r\n    std::fs::read_to_string(shell_perch.join(LINK_TOKEN_FILE))\r\n        .ok()\r\n        .map(|s| s.trim().to_string())\r\n        .filter(|s| !s.is_empty())\r\n}\r\n\r\n/// The per-link channel key, derived from the minted token: `SHA-256(token)`.\r\n/// Both ends hold the token (spt parked it; the binary got it in argv), so\r\n/// both derive the same key without a second exchange.\r\n// [impl->REQ-SHELL-1]\r\npub fn link_key(token: &str) -> [u8; 32] {\r\n    let mut h = Sha256::new();\r\n    h.update(token.as_bytes());\r\n    h.finalize().into()\r\n}\r\n\r\n/// Authenticate one link frame: `HMAC-SHA256(link_key, frame)`, hex. Required\r\n/// on every message both directions (CONTEXT §local-link authentication) —\r\n/// the D3c channels stamp + verify this per frame; a frame without a valid\r\n/// MAC is dropped, never processed.\r\n// [impl->REQ-SHELL-1]\r\npub fn frame_mac(key: &[u8; 32], frame: &[u8]) -> String {\r\n    let mut mac = <Hmac<Sha256> as Mac>::new_from_slice(key).expect(\"hmac accepts any key len\");\r\n    mac.update(frame);\r\n    mac.finalize()\r\n        .into_bytes()\r\n        .iter()\r\n        .map(|b| format!(\"{b:02x}\"))\r\n        .collect()\r\n}\r\n\r\n/// Verify a frame's MAC (constant-time inner compare via the hmac crate).\r\npub fn verify_frame_mac(key: &[u8; 32], frame: &[u8], mac_hex: &str) -> bool {\r\n    let mut mac = <Hmac<Sha256> as Mac>::new_from_slice(key).expect(\"hmac accepts any key len\");\r\n    mac.update(frame);\r\n    let Ok(bytes) = hex_decode(mac_hex) else {\r\n        return false;\r\n    };\r\n    mac.verify_slice(&bytes).is_ok()\r\n}\r\n\r\nfn hex_decode(s: &str) -> Result<Vec<u8>, ()> {\r\n    if !s.len().is_multiple_of(2) {\r\n        return Err(());\r\n    }\r\n    (0..s.len())\r\n        .step_by(2)\r\n        .map(|i| u8::from_str_radix(&s[i..i + 2], 16).map_err(|_| ()))\r\n        .collect()\r\n}\r\n\r\n/// The install dir an adapter-option's `[shell]` commands resolve against — the\r\n/// parent adapter record's `source_dir` (the live manifest's own directory in\r\n/// pointer mode). `None` when the option does not resolve to an active record.\r\n///\r\n/// One helper for every shell-family site (spawn AND wake) precisely because\r\n/// \"which dir\" is the kind of thing that drifts when each site works it out for\r\n/// itself: a profiled option resolves against its PARENT's install dir, and a\r\n/// site that split the option differently would silently resolve a different\r\n/// binary. A `:profile` overlays configuration, never provenance — the bits are\r\n/// the parent's.\r\n// [impl->REQ-INSTALL-11]\r\npub fn shell_install_dir(adapters_dir: &Path, adapter_option: &str) -> Option<PathBuf> {\r\n    spt_runtime::registry::resolve_option(adapters_dir, adapter_option)\r\n        .ok()\r\n        .map(|(record, _)| PathBuf::from(record.source_dir))\r\n}\r\n\r\n/// Fill the manifest's opaque `spawn` template for one instance (mints + parks\r\n/// the link token first — the binary may bind immediately) and tokenize it.\r\n/// Shared by the detached and broker-hosted launch paths.\r\n///\r\n/// `install_dir` is the adapter's own directory ([`shell_install_dir`]). It does\r\n/// two things, both of which a `--release`-installed shell adapter cannot work\r\n/// without (REQ-INSTALL-11's documented follow-on sites):\r\n/// - the program token resolves through it, so a bare program name binds to the\r\n///   adapter's SHIPPED binary before `PATH` — the same\r\n///   [`spt_runtime::runtime::resolve_program_in_dir`] primitive `[digest]`, the\r\n///   session `role.command` and `[service]` use, never a second implementation;\r\n/// - `{adapter_dir}` becomes fillable, so a template can name a resource beside\r\n///   that binary.\r\n///\r\n/// `None` means \"no install dir known\" and is not the same as an empty one: the\r\n/// program token then passes through for `PATH` to resolve exactly as before, and\r\n/// a template naming `{adapter_dir}` fails LOUDLY on the missing key rather than\r\n/// filling an empty string into a path.\r\nfn fill_spawn_command(\r\n    owlery: &Path,\r\n    owner: &str,\r\n    shell_id: &str,\r\n    adapter_name: &str,\r\n    install_dir: Option<&Path>,\r\n    shell: &Shell,\r\n) -> Result<Vec<String>, String> {\r\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\r\n    // Mutual exclusivity, the online side (M5-D4b): the watcher dies before\r\n    // the binary rises — and any (re)launch is \"shell activity\", so the\r\n    // crash give-up latch re-arms. [impl->REQ-SHELL-2]\r\n    crate::shellwake::kill_waker_at(&perch);\r\n    crate::shellwake::clear_gave_up(&perch);\r\n    // The keys the pending spool rows may be stamped under: the still-parked\r\n    // token of an abruptly-dead instance (the wake/relink path) AND the stash\r\n    // a previous rotation retired (covers the graceful-close→relink path plus\r\n    // one generation of race stragglers — see below). Read BEFORE the mint\r\n    // overwrites the parked file.\r\n    let parked_before_mint = read_link_token(&perch);\r\n    let retired_before_mint = std::fs::read_to_string(perch.join(RETIRED_TOKEN_FILE)).ok();\r\n    let token = mint_link_token();\r\n    // Park the token BEFORE the spawn: the binary may bind immediately.\r\n    std::fs::write(perch.join(LINK_TOKEN_FILE), &token).map_err(|e| format!(\"park token: {e}\"))?;\r\n    // Token rotation carries the durable channel with it: re-stamp every\r\n    // pending row from the old keys to the new, or the binary this launch\r\n    // raises would verify them against the fresh key and (correctly) discard —\r\n    // the wake-triggering command itself would be the one frame the wake loses\r\n    // (flynn's field catch, 2026-07-25: #23 armed the wake and never answered;\r\n    // #24/#19 spooled after the rotation and drained in order).\r\n    //\r\n    // Race + crash posture (todlando's A1/A2, 2026-07-25): a concurrent\r\n    // `drive_shell` that read the OLD parked token and spools after this walk\r\n    // finishes leaves a straggler row the walk never saw — so the old parked\r\n    // token is STASHED (not erased) and the drain paths give any\r\n    // retired-key row a second-chance re-stamp at delivery\r\n    // ([`crate::shellchan::restamp_for_drain`]). The walk itself is\r\n    // crash-safe and idempotent: each row rewrite is independent, a row\r\n    // already carrying the new stamp does not verify under an old key and is\r\n    // skipped, and a crash mid-walk leaves a mixed spool the NEXT rotation or\r\n    // drain converts (both old keys stay candidates). Rows are selected by\r\n    // MAC verification against the threaded new key — never by comparing a\r\n    // token snapshot (the D-2 stale-snapshot class).\r\n    // [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\r\n    let new_key = link_key(&token);\r\n    for old in [&parked_before_mint, &retired_before_mint].into_iter().flatten() {\r\n        crate::shellchan::restamp_pending_at(&perch, &link_key(old), &new_key);\r\n    }\r\n    // Stash the just-rotated-out token for the drain's second chance; a\r\n    // close-retired stash stays in place when this launch had no parked token\r\n    // (the close path spools nothing new — NoLink — so it has no stragglers).\r\n    if let Some(old) = &parked_before_mint {\r\n        let _ = std::fs::write(perch.join(RETIRED_TOKEN_FILE), old);\r\n    }\r\n\r\n    let keys = BTreeMap::from([\r\n        (\"id\".to_string(), shell_id.to_string()),\r\n        (\"adapter_name\".to_string(), adapter_name.to_string()),\r\n        (\"link_token\".to_string(), token),\r\n        // [impl->REQ-SHELL-PERCH-DIR] {perch_dir}: the shell perch root, so the\r\n        // binary can resolve a shell_file frame's PERCH-RELATIVE `path` attr\r\n        // (files/<xfer-id>-<name>) — the child inherits the BROKER's cwd, so\r\n        // without this no mechanical resolution exists. Opt-in and N-1-safe:\r\n        // a template that never names the key fills byte-identically. The\r\n        // frame's path deliberately STAYS relative (doyle ruling: an absolute\r\n        // path in a spooled frame lies across perch moves / node boundaries).\r\n        (\"perch_dir\".to_string(), perch.display().to_string()),\r\n    ]);\r\n    // [impl->REQ-INSTALL-11] {adapter_dir}: the adapter's own install dir, so a\r\n    // released shell adapter can point at a resource it ships. Opt-in and N-1-safe\r\n    // like {perch_dir} — a template that never names the key fills byte-identically.\r\n    let keys = match install_dir {\r\n        Some(dir) => {\r\n            let mut keys = keys;\r\n            keys.insert(\"adapter_dir\".to_string(), dir.display().to_string());\r\n            keys\r\n        }\r\n        None => keys,\r\n    };\r\n    // [impl->REQ-HAZARD-TEMPLATE-ARGV-FILL] tokenize-template-then-fill-each: a\r\n    // multi-word/quote/semicolon {key} value is exactly one argv element.\r\n    let mut tokens =\r\n        spt_runtime::runtime::fill_template_tokens(&shell.spawn, &keys).map_err(|e| e.to_string())?;\r\n    let Some(program) = tokens.first_mut() else {\r\n        return Err(\"empty spawn command\".into());\r\n    };\r\n    // [impl->REQ-INSTALL-11] one resolution primitive, no parallel path: a\r\n    // `--release`-installed adapter declaring a bare program name binds to its\r\n    // shipped binary before PATH. Without this such an adapter registers fine and\r\n    // then fails at spawn with os error 2 — the gap that forced a hand-maintained\r\n    // local manifest per node.\r\n    if let Some(dir) = install_dir {\r\n        *program = spt_runtime::runtime::resolve_program_in_dir(program, dir);\r\n    }\r\n    Ok(tokens)\r\n}\r\n\r\n/// WHO is asking for a launch — and therefore WHOSE Job Object the spawned\r\n/// binary would join (REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE).\r\n///\r\n/// A daemon-side spawn is job-neutral: the daemon escaped the launching job at\r\n/// its own cold start (the WMI → schtasks → breakaway ladder). A CLI-side spawn\r\n/// is not — `spt shell spawn`, local `spt shell relink` and the `shell cmd`\r\n/// wake-if-offline arm all call the launch IN the terminal-resident CLI process,\r\n/// so on a job that forbids breakaway the shell binary (and every descendant it\r\n/// ever spawns, relay hops included) inherits the terminal's job.\r\n// [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\r\n#[derive(Clone, Copy, Debug, PartialEq, Eq)]\r\npub enum LaunchOrigin {\r\n    /// The caller IS the daemon (broker/brain): spawn here, already job-neutral.\r\n    Daemon,\r\n    /// The caller is a terminal-resident CLI: hand the launch to a running\r\n    /// daemon when one answers, else spawn here and SAY whose job took the child.\r\n    Cli,\r\n}\r\n\r\n/// A routed launch's failure, split by what the caller may safely do next.\r\n// [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\r\n#[derive(Debug, Clone, PartialEq, Eq)]\r\npub enum LaunchFailure {\r\n    /// The launch was attempted and failed. Undo whatever the launch was for.\r\n    Failed(String),\r\n    /// The launch was handed off and its outcome is UNKNOWN — the binary may be\r\n    /// running. Neither retry nor undo: both are wrong half the time.\r\n    Indeterminate(String),\r\n}\r\n\r\nimpl std::fmt::Display for LaunchFailure {\r\n    fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {\r\n        match self {\r\n            LaunchFailure::Failed(m) | LaunchFailure::Indeterminate(m) => write!(f, \"{m}\"),\r\n        }\r\n    }\r\n}\r\n\r\nimpl LaunchFailure {\r\n    /// Whether the outcome is unknown — the caller must keep what it built.\r\n    // [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\r\n    pub fn is_indeterminate(&self) -> bool {\r\n        matches!(self, LaunchFailure::Indeterminate(_))\r\n    }\r\n}\r\n\r\n/// Whether this shell's launch is the one that can be routed to the daemon: the\r\n/// DETACHED relay spawn, the only receipt that spawns in the calling process.\r\n///\r\n/// `stdin` already spawns broker-side (so it is job-neutral today, and routing\r\n/// it would make the broker connect to its own socket); `http` refuses at spawn.\r\n// [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\r\npub fn routes_through_daemon(shell: &Shell) -> bool {\r\n    !matches!(\r\n        shell.command_receipt,\r\n        Some(CommandReceipt::Http) | Some(CommandReceipt::Stdin)\r\n    )\r\n}\r\n\r\n/// The double-launch arbiter: the instance's recorded pid, when it probes\r\n/// ALIVE (REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE).\r\n///\r\n/// `Some(pid)` ⇒ a launch of this same `{owner, shell_id}` already won — stand\r\n/// down and adopt it. Used by BOTH sides of the routed launch, because a late\r\n/// ack can put the caller and the daemon in the same launch at once.\r\n///\r\n/// Birth-safe, including a binary awaiting bind: a recycled live pid is NOT a\r\n/// winner when its native birth stamp mismatches. Missing, unparseable and zero\r\n/// pids remain not-a-winner, preserving pid-less backend behavior.\r\n// [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\r\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\npub fn live_launch_winner(shell_perch: &Path) -> Option<u32> {\r\n    shellinfo::read_shell_pid(shell_perch)\r\n        .filter(|&p| p != 0 && !shellinfo::shell_pid_provably_dead(shell_perch))\r\n}\r\n\r\n/// THE route-selection seam. Generic over the daemon call and the in-process\r\n/// fallback so the decision can be tested without a socket, a daemon or a job.\r\n///\r\n/// The invariant it exists to hold: **the ack is a commitment transfer** — a\r\n/// daemon that accepted owns the launch, so only a pre-ack refusal\r\n/// ([`LaunchRouteError::fallback_allowed`](crate::drivehub::LaunchRouteError::fallback_allowed))\r\n/// may fall through to an in-process spawn. Anything else returning a second\r\n/// spawn would be the double-launch this routing exists to avoid.\r\n///\r\n/// Two axes, deliberately separate: `fallback_allowed` decides whether a second\r\n/// spawn is PERMITTED, and\r\n/// [`may_have_launched`](crate::drivehub::LaunchRouteError::may_have_launched)\r\n/// decides whether one must first be ARBITRATED against a live pid. Only the\r\n/// unanswered wait is both.\r\n// [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\r\npub(crate) fn launch_routed(\r\n    shell: &Shell,\r\n    label: &str,\r\n    via_daemon: impl FnOnce() -> Result<u32, crate::drivehub::LaunchRouteError>,\r\n    already_live: impl FnOnce() -> Option<u32>,\r\n    in_process: impl FnOnce() -> Result<u32, String>,\r\n) -> Result<u32, LaunchFailure> {\r\n    if !routes_through_daemon(shell) {\r\n        return in_process().map_err(LaunchFailure::Failed);\r\n    }\r\n    match via_daemon() {\r\n        Ok(pid) => Ok(pid),\r\n        Err(e) if e.fallback_allowed() => {\r\n            // A pre-ack refusal means nothing was spawned — EXCEPT for the one\r\n            // pre-ack cause that is a timeout rather than an answer. A daemon\r\n            // that acks late may already have launched, so THAT fallback checks\r\n            // the arbiter before adding a second binary (the documented race at\r\n            // `LAUNCH_ACK_BOUND`). Adopting the winner is a SUCCESS: the caller\r\n            // wanted a running binary and there is one.\r\n            //\r\n            // The gate is `may_have_launched`, not `fallback_allowed`: running\r\n            // the arbiter on an ANSWERED refusal adopts a pid no daemon could\r\n            // have launched, so the fallback silently declines to spawn — and a\r\n            // relink that never re-launches never mints, which is precisely the\r\n            // recovery command failing on the state it exists to recover.\r\n            // [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\r\n            // Gated by SHORT-CIRCUIT, not by filtering the answer: outside the\r\n            // late-ack window the pid probe is not merely unused, it must not\r\n            // run at all — that is one fewer place a pinned corpse handle can\r\n            // report a dead pid alive.\r\n            let winner = if e.may_have_launched() {\r\n                already_live()\r\n            } else {\r\n                None\r\n            };\r\n            if let Some(pid) = winner {\r\n                spt_proto::emit_line_err!(\r\n                    \"SHELL_LAUNCH_ALREADY_LIVE:{label}: a launch of this instance is already \\\r\n                     running (pid={pid}) — standing down instead of spawning a second binary\"\r\n                );\r\n                return Ok(pid);\r\n            }\r\n            // Nothing was spawned daemon-side. Say WHY the route was not taken\r\n            // and WHOSE job the child is about to join — the three pre-ack\r\n            // causes carry distinct names, so a field report can tell an old\r\n            // daemon from a foreign home from no daemon at all.\r\n            spt_proto::emit_line_err!(\r\n                \"SHELL_LAUNCH_IN_PROCESS:{label}: {e} — spawning from THIS process instead, so \\\r\n                 the binary joins THIS process's Job Object (a DETACH_BREAKAWAY_DENIED line \\\r\n                 names it); the first daemon-side wake re-spawns it job-neutrally\"\r\n            );\r\n            in_process().map_err(LaunchFailure::Failed)\r\n        }\r\n        Err(e) if e.is_indeterminate() => Err(LaunchFailure::Indeterminate(e.to_string())),\r\n        Err(e) => Err(LaunchFailure::Failed(e.to_string())),\r\n    }\r\n}\r\n\r\n/// [`launch_shell`], routed by ORIGIN (REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE).\r\n///\r\n/// From the daemon this is `launch_shell` verbatim. From a CLI it hands the\r\n/// launch to a running daemon over the drive control socket — so the binary\r\n/// rises job-neutral instead of inside the launching terminal's Job Object —\r\n/// and falls back to today's in-process spawn, loudly, when no daemon takes it.\r\n// [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\r\npub fn launch_shell_from(\r\n    origin: LaunchOrigin,\r\n    owlery: &Path,\r\n    owner: &str,\r\n    shell_id: &str,\r\n    adapter_name: &str,\r\n    install_dir: Option<&Path>,\r\n    shell: &Shell,\r\n) -> Result<u32, LaunchFailure> {\r\n    match origin {\r\n        LaunchOrigin::Daemon => {\r\n            launch_shell(owlery, owner, shell_id, adapter_name, install_dir, shell)\r\n                .map_err(LaunchFailure::Failed)\r\n        }\r\n        LaunchOrigin::Cli => launch_routed(\r\n            shell,\r\n            shell_id,\r\n            || {\r\n                crate::drivehub::shell_launch(\r\n                    &crate::endpoint::drive_socket_name(),\r\n                    owner,\r\n                    shell_id,\r\n                    owlery,\r\n                )\r\n            },\r\n            || {\r\n                live_launch_winner(&spt_store::perch::resolve_shell_perch_path_in(\r\n                    owlery, owner, shell_id,\r\n                ))\r\n            },\r\n            || launch_shell(owlery, owner, shell_id, adapter_name, install_dir, shell),\r\n        ),\r\n    }\r\n}\r\n\r\n/// Launch a shell instance's binary per the manifest's `command_receipt`\r\n/// (M5-D3c): **relay** (default) spawns detached (stdio discarded — the\r\n/// binary drains via `api poll --link`); **stdin** spawns broker-hosted so\r\n/// the daemon owns the stdin pipe ([`launch_shell_brokered_in`] — command\r\n/// frames ride the broker's `send_effect` input seam); **http** refuses at\r\n/// spawn, explicitly — an unimplemented mode is never a silent fallback.\r\n/// Records the pid; returns it. The perch stays **offline** until the\r\n/// binary's `api bind-shell` handshake onlines it — launch is not the online\r\n/// switch any more than spawn is.\r\n// [impl->REQ-SHELL-1]\r\npub fn launch_shell(\r\n    owlery: &Path,\r\n    owner: &str,\r\n    shell_id: &str,\r\n    adapter_name: &str,\r\n    install_dir: Option<&Path>,\r\n    shell: &Shell,\r\n) -> Result<u32, String> {\r\n    match shell.command_receipt {\r\n        Some(CommandReceipt::Http) => Err(\r\n            \"command_receipt=http is not implemented yet — spawn refused (an \\\r\n             unimplemented mode never falls back silently)\"\r\n                .into(),\r\n        ),\r\n        Some(CommandReceipt::Stdin) => {\r\n            crate::ensure_running().map_err(|e| format!(\"daemon start: {e}\"))?;\r\n            launch_shell_brokered_in(\r\n                &crate::broker_socket_name(),\r\n                owlery,\r\n                owner,\r\n                shell_id,\r\n                adapter_name,\r\n                install_dir,\r\n                shell,\r\n            )\r\n        }\r\n        _ => {\r\n            let tokens =\r\n                fill_spawn_command(owlery, owner, shell_id, adapter_name, install_dir, shell)?;\r\n            let (program, args) = tokens.split_first().expect(\"non-empty checked\");\r\n            // The shell binary is a LONG-LIVED detached child (KNOWN-HAZARDS\r\n            // 5.6): on Windows it must inherit NO handles, or a caller\r\n            // capturing `spt shell spawn`'s output — at any depth of the\r\n            // process chain — hangs forever on a pipe write-end the immortal\r\n            // child holds. Same no-inherit spawn as the daemon's.\r\n            // [impl->REQ-HAZARD-DETACHED-PIPE-INHERIT]\r\n            #[cfg(windows)]\r\n            let pid = crate::daemon::detached_no_inherit(program, args)\r\n                .map_err(|e| format!(\"spawn {program}: {e}\"))?;\r\n            #[cfg(unix)]\r\n            let pid = Command::new(program)\r\n                .args(args)\r\n                .stdin(Stdio::null())\r\n                .stdout(Stdio::null())\r\n                .stderr(Stdio::null())\r\n                .spawn()\r\n                .map_err(|e| format!(\"spawn {program}: {e}\"))?\r\n                .id();\r\n            let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\r\n            std::fs::write(perch.join(SHELL_PID_FILE), pid.to_string())\r\n                .map_err(|e| format!(\"record pid: {e}\"))?;\r\n            // The birth stamp rides WITH the pid at every launch, or the probe that\r\n            // reads it degrades to bare-pid liveness on this instance alone -- a\r\n            // silent per-instance regression rather than a visible one.\r\n            // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\n            spt_store::shellinfo::record_shell_launch(&perch, pid, now_ms());\r\n            Ok(pid)\r\n        }\r\n    }\r\n}\r\n\r\n/// The stdin-receipt launch: spawn the filled template as a **broker-hosted**\r\n/// session labeled [`shell_session_label`], so the daemon owns the child's\r\n/// stdin and `spt shell cmd` can drive it through the `send_effect` input\r","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120,121,122,123,124,125,126,127,128,129,130,131,132,133,134,135,136,137,138,139,140,141,142,143,144,145,146,147,148,149,150,151,152,153,154,155,156,157,158,159,160,161,162,163,164,165,166,167,168,169,170,171,172,173,174,175,176,177,178,179,180,181,182,183,184,185,186,187,188,189,190,191,192,193,194,195,196,197,198,199,200,201,202,203,204,205,206,207,208,209,210,211,212,213,214,215,216,217,218,219,220,221,222,223,224,225,226,227,228,229,230,231,232,233,234,235,236,237,238,239,240,241,242,243,244,245,246,247,248,249,250,251,252,253,254,255,256,257,258,259,260,261,262,263,264,265,266,267,268,269,270,271,272,273,274,275,276,277,278,279,280,281,282,283,284,285,286,287,288,289,290,291,292,293,294,295,296,297,298,299,300,301,302,303,304,305,306,307,308,309,310,311,312,313,314,315,316,317,318,319,320,321,322,323,324,325,326,327,328,329,330,331,332,333,334,335,336,337,338,339,340,341,342,343,344,345,346,347,348,349,350,351,352,353,354,355,356,357,358,359,360,361,362,363,364,365,366,367,368,369,370,371,372,373,374,375,376,377,378,379,380,381,382,383,384,385,386,387,388,389,390,391,392,393,394,395,396,397,398,399,400,401,402,403,404,405,406,407,408,409,410,411,412,413,414,415,416,417,418,419,420,421,422,423,424,425,426,427,428,429,430,431,432,433,434,435,436,437,438,439,440,441,442,443,444,445,446,447,448,449,450,451,452,453,454,455,456,457,458,459,460,461,462,463,464,465,466,467,468,469,470,471,472,473,474,475,476,477,478,479,480,481,482,483,484,485,486,487,488,489,490,491,492,493,494,495,496,497,498,499,500,501,502,503,504,505,506,507,508,509,510,511,512,513]},"fileSize":88363,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\crates\\spt-daemon\\src\\shellhost.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":1826,"totalBytes":25273,"outputLines":513,"outputBytes":25273,"shownRange":{"start":1,"end":513},"nextOffset":514}}},"isError":false,"timestamp":1789001151897}}
{"type":"message","id":"2e4d8396","parentId":"905372b3","timestamp":"2026-09-10T00:45:51.957Z","message":{"role":"toolResult","toolCallId":"call_nKdV7gj0NifQPSIMOkbPE8HP|fc_077f3569365e313d016aa1fdbfc92c87d19d1b7edec2d8c706","toolName":"grep","content":[{"type":"text","text":"# .worktrees/287-shellrestore/crates/\n\n## spt-store/src/\n### liveness.rs\n 183|#[derive(Debug, Clone, Copy, PartialEq, Eq)]\n*184|pub enum RelayLiveness {\n 185|    /// The recorded relay process is still there (and still the one we recorded,\n 186|    /// when a birth stamp is available to pair against).\n 187|    Held,\n...\n 217|// [impl->REQ-RELAY-DEATH-CONVERGENCE]\n*218|pub fn relay_liveness(pid: Option<u32>, recorded_birth: Option<u64>) -> RelayLiveness {\n 219|    // No pid / BUSY sentinel (`read_pid` ⇒ None), or the `0` a pid-less backend\n 220|    // records: never a probe-able process, so never a death verdict.\n 221|    let Some(pid) = pid.filter(|p| *p != 0) else {\n...\n 418|    #[test]\n*419|    fn relay_liveness_kills_only_a_provably_gone_relay() {\n 420|        let me = std::process::id();\n 421|        let my_birth = crate::proc::process_started_at(me);\n 422|\n\n## spt-daemon/src/\n### shellwake.rs\n 630|                // [impl->REQ-SHELL-HEAL-ENUMERATION]\n*631|                let _ = spt_proto::emit_line!(events, \"SHELL_RECORD_HEALED:{owner}/{shell_id}: online -> {truth}\");\n 632|            }\n 633|        }\n 634|    }\n...\n 1125|            let mut expected = vec![\n*1126|                format!(\"SHELL_RECORD_HEALED:ling/{stale}: online -> offline\"),\n*1127|                format!(\"SHELL_RECORD_HEALED:ling/{sibling}: online -> offline\"),\n*1128|                format!(\"SHELL_RECORD_HEALED:ling/{pacer}: online -> offline\"),\n*1129|                format!(\"SHELL_RECORD_HEALED:perri/{other}: online -> offline\"),\n 1130|            ];\n 1131|            expected.sort();\n 1132|            assert_eq!(reported, expected, \"every healed instance, no healthy representative\");"}],"details":{"scopePath":".worktrees/287-shellrestore/crates/spt-store/src, .worktrees/287-shellrestore/crates/spt-proto/src, .worktrees/287-shellrestore/crates/spt-daemon/src","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\crates","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":8,"fileCount":2,"files":[".worktrees/287-shellrestore/crates/spt-store/src/liveness.rs",".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs"],"fileMatches":[{"path":".worktrees/287-shellrestore/crates/spt-store/src/liveness.rs","count":3},{"path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs","count":5}],"truncated":false,"displayContent":"# .worktrees/287-shellrestore/crates/\n\n## spt-store/src/\n### liveness.rs\n 183│#[derive(Debug, Clone, Copy, PartialEq, Eq)]\n*184│pub enum RelayLiveness {\n 185│    /// The recorded relay process is still there (and still the one we recorded,\n 186│    /// when a birth stamp is available to pair against).\n 187│    Held,\n    │...\n 217│// [impl->REQ-RELAY-DEATH-CONVERGENCE]\n*218│pub fn relay_liveness(pid: Option<u32>, recorded_birth: Option<u64>) -> RelayLiveness {\n 219│    // No pid / BUSY sentinel (`read_pid` ⇒ None), or the `0` a pid-less backend\n 220│    // records: never a probe-able process, so never a death verdict.\n 221│    let Some(pid) = pid.filter(|p| *p != 0) else {\n    │...\n 418│    #[test]\n*419│    fn relay_liveness_kills_only_a_provably_gone_relay() {\n 420│        let me = std::process::id();\n 421│        let my_birth = crate::proc::process_started_at(me);\n 422│\n\n## spt-daemon/src/\n### shellwake.rs\n  630│                // [impl->REQ-SHELL-HEAL-ENUMERATION]\n *631│                let _ = spt_proto::emit_line!(events, \"SHELL_RECORD_HEALED:{owner}/{shell_id}: online -> {truth}\");\n  632│            }\n  633│        }\n  634│    }\n     │...\n 1125│            let mut expected = vec![\n*1126│                format!(\"SHELL_RECORD_HEALED:ling/{stale}: online -> offline\"),\n*1127│                format!(\"SHELL_RECORD_HEALED:ling/{sibling}: online -> offline\"),\n*1128│                format!(\"SHELL_RECORD_HEALED:ling/{pacer}: online -> offline\"),\n*1129│                format!(\"SHELL_RECORD_HEALED:perri/{other}: online -> offline\"),\n 1130│            ];\n 1131│            expected.sort();\n 1132│            assert_eq!(reported, expected, \"every healed instance, no healthy representative\");"},"isError":false,"timestamp":1789001151956}}
{"type":"custom_message","customType":"irc:incoming","content":"<irc>\nIncoming IRC message from agent `PersistentRestoreReview.RestoreTestsReview`:\n\nNo shell/git tool exposed in my tools. Please provide read-accessible cb756dee-parent diff artifact for assigned shellwake/shellhost tests, shell_relink_force_e2e and cli; proceeding with source reads.\n\nSent while waiting/working. Active interruptible wait stopped early for immediate reading.\n\nIf response expected, reply via `hub` (`op: \"send\"`, `to: \"PersistentRestoreReview.RestoreTestsReview\"`); may finish current step first. No one replies on your behalf.\n</irc>","display":true,"details":{"id":"15797f14817ffeee","from":"PersistentRestoreReview.RestoreTestsReview","message":"No shell/git tool exposed in my tools. Please provide read-accessible cb756dee-parent diff artifact for assigned shellwake/shellhost tests, shell_relink_force_e2e and cli; proceeding with source reads."},"attribution":"agent","id":"c1e129a4","parentId":"2e4d8396","timestamp":"2026-09-10T00:45:38.693Z"}
{"type":"custom_message","customType":"irc:incoming","content":"<irc>\nIncoming IRC message from agent `PersistentRestoreReview.RestoreDocsStandards`:\n\nI have read/grep only, no shell device mounted. Please supply read-accessible full cb756dee^..cb756dee diff for assigned six docs/manifest files (artifact/local URL); need exact patch overlap. Proceeding current docs and standards meanwhile.\n\nSent while waiting/working. Active interruptible wait stopped early for immediate reading.\n\nIf response expected, reply via `hub` (`op: \"send\"`, `to: \"PersistentRestoreReview.RestoreDocsStandards\"`); may finish current step first. No one replies on your behalf.\n</irc>","display":true,"details":{"id":"15797f158c7ffeef","from":"PersistentRestoreReview.RestoreDocsStandards","message":"I have read/grep only, no shell device mounted. Please supply read-accessible full cb756dee^..cb756dee diff for assigned six docs/manifest files (artifact/local URL); need exact patch overlap. Proceeding current docs and standards meanwhile."},"attribution":"agent","id":"8e20dda6","parentId":"c1e129a4","timestamp":"2026-09-10T00:45:39.761Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_uyG0Y8O4S5ZbUWs9wg4JAMjZ|fc_077f3569365e313d016aa1fdc1d1a087d19c4e02962170e584","toolName":"hub","startedAt":"2026-09-10T00:45:56.290Z","intent":"Sharing full review patch"},"id":"95f7f100","parentId":"8e20dda6","timestamp":"2026-09-10T00:45:56.290Z"}
{"type":"message","id":"c1a67f90","parentId":"95f7f100","timestamp":"2026-09-10T00:45:56.291Z","message":{"role":"assistant","content":[{"type":"toolCall","id":"call_uyG0Y8O4S5ZbUWs9wg4JAMjZ|fc_077f3569365e313d016aa1fdc1d1a087d19c4e02962170e584","name":"hub","arguments":{"i":"Sharing full review patch","op":"send","to":"PersistentRestoreReview.RestoreTestsReview","message":"Full raw patch artifact://112 (1088 lines); tests mostly lines 400-925, CLI/relink e2e around 925. read ranges, with :raw if needed. Full immutable commit diff from cb756dee^ to cb756dee."},"intent":"Sharing full review patch"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":12742,"output":90,"cacheRead":49920,"cacheWrite":0,"totalTokens":62752,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1789001151986,"responseId":"resp_077f3569365e313d016aa1fdc03fcc87d19fabde850a0d343c","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"fc_077f3569365e313d016aa1fdc1d1a087d19c4e02962170e584","type":"function_call","status":"completed","arguments":"{\"i\":\"Sharing full review patch\",\"op\":\"send\",\"to\":\"PersistentRestoreReview.RestoreTestsReview\",\"message\":\"Full raw patch artifact://112 (1088 lines); tests mostly lines 400-925, CLI/relink e2e around 925. read ranges, with :raw if needed. Full immutable commit diff from cb756dee^ to cb756dee.\"}","call_id":"call_uyG0Y8O4S5ZbUWs9wg4JAMjZ","internal_chat_message_metadata_passthrough":{"create_time":1789001152.784181,"turn_id":"0b013667-0aaf-4a85-84aa-5b88e3033a8e"},"name":"hub","metadata":{"turn_id":"0b013667-0aaf-4a85-84aa-5b88e3033a8e"}}]},"duration":4302.238299999852,"ttft":1863.999799999874,"completedAt":1789001156289,"contextSnapshot":{"promptTokens":62662,"nonMessageTokens":13209,"compactionEpoch":0}}}
{"type":"message","id":"73a21d75","parentId":"c1a67f90","timestamp":"2026-09-10T00:45:56.292Z","message":{"role":"toolResult","toolCallId":"call_uyG0Y8O4S5ZbUWs9wg4JAMjZ|fc_077f3569365e313d016aa1fdc1d1a087d19c4e02962170e584","toolName":"hub","content":[{"type":"text","text":"Delivered to 1 peer(s):\n- PersistentRestoreReview.RestoreTestsReview: injected"}],"details":{"op":"send","from":"PersistentRestoreReview","to":"PersistentRestoreReview.RestoreTestsReview","receipts":[{"to":"PersistentRestoreReview.RestoreTestsReview","outcome":"injected"}]},"isError":false,"timestamp":1789001156291}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_TAUGdGLJKIAQF8HIef6rOcA1|fc_077f3569365e313d016aa1fdc68ee487d19873cce44673a9aa","toolName":"hub","startedAt":"2026-09-10T00:46:00.935Z","intent":"Sharing full documentation patch"},"id":"81b86d63","parentId":"73a21d75","timestamp":"2026-09-10T00:46:00.935Z"}
{"type":"message","id":"56f5f14e","parentId":"81b86d63","timestamp":"2026-09-10T00:46:00.935Z","message":{"role":"assistant","content":[{"type":"toolCall","id":"call_TAUGdGLJKIAQF8HIef6rOcA1|fc_077f3569365e313d016aa1fdc68ee487d19873cce44673a9aa","name":"hub","arguments":{"i":"Sharing full documentation patch","op":"send","to":"PersistentRestoreReview.RestoreDocsStandards","message":"Full raw patch artifact://112 (1088 lines); docs beginning lines 1-47 and latter ~940-1088. read selected ranges, :raw if needed. Full immutable commit diff cb756dee^ to cb756dee."},"intent":"Sharing full documentation patch"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":316,"output":87,"cacheRead":62464,"cacheWrite":0,"totalTokens":62867,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1789001156308,"responseId":"resp_077f3569365e313d016aa1fdc48fd087d186843cb6548defb0","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"fc_077f3569365e313d016aa1fdc68ee487d19873cce44673a9aa","type":"function_call","status":"completed","arguments":"{\"i\":\"Sharing full documentation patch\",\"op\":\"send\",\"to\":\"PersistentRestoreReview.RestoreDocsStandards\",\"message\":\"Full raw patch artifact://112 (1088 lines); docs beginning lines 1-47 and latter ~940-1088. read selected ranges, :raw if needed. Full immutable commit diff cb756dee^ to cb756dee.\"}","call_id":"call_TAUGdGLJKIAQF8HIef6rOcA1","internal_chat_message_metadata_passthrough":{"create_time":1789001156.847117,"turn_id":"0b013667-0aaf-4a85-84aa-5b88e3033a8e"},"name":"hub","metadata":{"turn_id":"0b013667-0aaf-4a85-84aa-5b88e3033a8e"}}]},"duration":4625.57620000001,"ttft":2274.446100000292,"completedAt":1789001160934,"contextSnapshot":{"promptTokens":62780,"nonMessageTokens":13209,"compactionEpoch":0}}}
{"type":"message","id":"0cca23ce","parentId":"56f5f14e","timestamp":"2026-09-10T00:46:00.936Z","message":{"role":"toolResult","toolCallId":"call_TAUGdGLJKIAQF8HIef6rOcA1|fc_077f3569365e313d016aa1fdc68ee487d19873cce44673a9aa","toolName":"hub","content":[{"type":"text","text":"Delivered to 1 peer(s):\n- PersistentRestoreReview.RestoreDocsStandards: injected"}],"details":{"op":"send","from":"PersistentRestoreReview","to":"PersistentRestoreReview.RestoreDocsStandards","receipts":[{"to":"PersistentRestoreReview.RestoreDocsStandards","outcome":"injected"}]},"isError":false,"timestamp":1789001160936}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_4XqdQMELsGY7pnb5eiuhkd6i|fc_077f3569365e313d016aa1fde3a33c87d1a3f417003c017244","toolName":"read","startedAt":"2026-09-10T00:46:27.738Z","args":{"path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs:514-720"},"intent":"Reading launch close stamp lifetime"},"id":"51b2442b","parentId":"0cca23ce","timestamp":"2026-09-10T00:46:27.738Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_iriKaG5SXPRscFVxSnmTljAd|fc_077f3569365e313d016aa1fde3a35487d1a934b089426749e3","toolName":"read","startedAt":"2026-09-10T00:46:27.738Z","args":{"path":".worktrees/287-shellrestore/crates/spt-store/src/liveness.rs:175-275"},"intent":"Reading liveness verdict implementation"},"id":"4835115f","parentId":"51b2442b","timestamp":"2026-09-10T00:46:27.738Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_1KIq0qgFYanlAYMVZvaDV02J|fc_077f3569365e313d016aa1fde3a36487d1a8f349852acfa5b4","toolName":"grep","startedAt":"2026-09-10T00:46:27.738Z","args":{"path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs;.worktrees/287-shellrestore/crates/spt-store/src"},"intent":"Locating shell close implementation"},"id":"d9fd0dc0","parentId":"4835115f","timestamp":"2026-09-10T00:46:27.738Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_VnQyIZqBQNs5YuGmMeW8r7gn|fc_077f3569365e313d016aa1fde3a36c87d193b54d2ea3b75ee5","toolName":"read","startedAt":"2026-09-10T00:46:27.738Z","args":{"path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs:1190-1470"},"intent":"Reading wake binding exclusion tests"},"id":"b4017d56","parentId":"d9fd0dc0","timestamp":"2026-09-10T00:46:27.738Z"}
{"type":"message","id":"d823f8f3","parentId":"b4017d56","timestamp":"2026-09-10T00:46:27.738Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":"**Flagging shellhost launch race**\n\n**Assessing patch regressions**\n\n**Tracing PID winner races**"},{"type":"toolCall","id":"call_4XqdQMELsGY7pnb5eiuhkd6i|fc_077f3569365e313d016aa1fde3a33c87d1a3f417003c017244","name":"read","arguments":{"i":"Reading launch close stamp lifetime","path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs:514-720"},"intent":"Reading launch close stamp lifetime"},{"type":"toolCall","id":"call_iriKaG5SXPRscFVxSnmTljAd|fc_077f3569365e313d016aa1fde3a35487d1a934b089426749e3","name":"read","arguments":{"i":"Reading liveness verdict implementation","path":".worktrees/287-shellrestore/crates/spt-store/src/liveness.rs:175-275"},"intent":"Reading liveness verdict implementation"},{"type":"toolCall","id":"call_1KIq0qgFYanlAYMVZvaDV02J|fc_077f3569365e313d016aa1fde3a36487d1a8f349852acfa5b4","name":"grep","arguments":{"i":"Locating shell close implementation","pattern":"pub fn close_shell|SHELL_LAUNCH_FILE|remove_file.*SHELL_PID_FILE|pub fn is_process_alive|pub fn process_started_at","path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs;.worktrees/287-shellrestore/crates/spt-store/src","case":true,"gitignore":false,"skip":null},"intent":"Locating shell close implementation"},{"type":"toolCall","id":"call_VnQyIZqBQNs5YuGmMeW8r7gn|fc_077f3569365e313d016aa1fde3a36c87d193b54d2ea3b75ee5","name":"read","arguments":{"i":"Reading wake binding exclusion tests","path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs:1190-1470"},"intent":"Reading wake binding exclusion tests"},{"type":"toolCall","id":"call_UpgQE4rfSVo6QpfvSZ2H1fCv|fc_077f3569365e313d016aa1fde3a37887d1af10a446fd33cf00","name":"grep","arguments":{"i":"Locating protocol emission implementation","pattern":"macro_rules! emit_line|fn emit_line|SHELL_RELINKED|status=binding","path":".worktrees/287-shellrestore/crates/spt-proto/src;.worktrees/287-shellrestore/crates/spt-daemon/src","case":true,"gitignore":false,"skip":null},"intent":"Locating protocol emission implementation"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":304,"output":847,"cacheRead":62592,"cacheWrite":0,"totalTokens":63743,"reasoningTokens":516,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1789001160960,"responseId":"resp_077f3569365e313d016aa1fdc93fa487d1bfae55aeb675d3c6","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_077f3569365e313d016aa1fdca3d4887d18e69c02d7d7ea2d1","type":"reasoning","content":[],"encrypted_content":"gAAAAABqof3jppxOJ1R4w7FRBcZEE0r4aL7dLXjO6xVLgX-Zyk0LfSMQix52Og1ZEvJgCirKFJ7LryXcme5EbGyQOjdpxcUQsUwELNpPtb9cLrqrixiPOWJCvQHJnQ_OqKD04NS-SgasZ17H81AzaSpYFKtRCspX4BqRc0VIJO99OKvGIiN5-WAgV33oqJuGcg9W-ks8vTXHdPypmwNIKfhlFlEPEXF7GwqYNTld-Oc-vNr4CnqoJkKo2H9qPEfb0FdMbHs1ZBLy0IFn89VdglEzs5zZEbYG4vfB-Wahft2CzYIN2bEEXlr76no2wXApqLnjr9Bh-Gb5E740MfaGcFhBLnUjFixrS083csUS43mm1Oru0zjJKa5DV36Xt4QUcdPsYTJc-EpKABaujjUUfBOT9HQs8XW_Oe6pxVxfaHgJr6bebcKxwGWndWibgZxUoSS5WDgFWxguv_NQSNA8B_pDJcDvnz9LdKP4Qw2m6_2hOI0jq-06pBRkrlZvUEtizYf0GqoOKkRUI2MXze-NOfNvbazdy5dJ4BiDnDNsdoxxFM3NmILA6M0ov98ryiEpAcDKLiE8Ay9vJuN05gQqpipoDD9ZAMrFCfdHWznWVCU1ZbzhS9WUQuyCOeQaBXTu0iueRXSzOk2swqbeiCKRpr9DIGEUOZB33YzuexkR7VtjgmriiS1gJUeZ06s9jNRrJRS_9e-Z4s92oRxRGt49HWmPKsimWB0k-bhy-xTf8kLamKARVsW0D_1l2XJYx5UiRaj8vhcAX66kjS8rqdJeNAtyuTSvj76BuvQwt7AwZwy4r3qLObp2747CjEM2Te9P9A4XXsuKEgUJqD158b7yz-I042-LoJpvF5lhOHGXALyX5FwQ29wnIT8JFCzuLPedN6eFZRuD8lf9na_12aSzq1fULuTxXVane1B6Rh1jBxiuSlNKTj2LuMS49lrDcBA3Jil1ex1wV0NYRlpZM-pDN3FqOvpSCHMhsmCOyqvEV-uQLXQwhXaL8u_jLwVg6O88fBENYkqSgZh2AIkCwpewl-0hz-hILOe2NKjWFNfAX-3KylQObzWSQ7qcwEJhcJhAoqjjzhoJdHK65WGLkZm8ZuK1Szf9D99Hp1cpJceX91ST3FROkdDp1JEdAqiF7n1UCg6EE1hOuRwmUIl7YtvOuo3nohF9FbM-C6Zp2Iq6je-FqhsrOYPEQNKqkippZTxZBXYrVUWmUpkql_NmEw2s01DvFYnHe1JG-UrhyfdlsxpwbnM4nY7atnNlMBZ9fD1l_c7sAgH6GmkkiLhhBLYf1Jlc8dALYxEFCLpAffDFhxXoQep5y9CSdRuOkt5DkWB9jWIXEqMyxAEcv2DkGDQjn4t-a_JKmdKkS3awHLXCTXWmzMAbEyQjxCNfT7lSU708q9aR5wFiFcW_CUmQZ1vWzlD8k4F-zapHZSmpzJ0u0hP3KmQNvV3NgREZS0Zm92E0iUHJsVyMi1KIA7FH5t_TmgiooAzIY578LVbmZSfOZipoJpvcX7rGvwvxeTEO-yaU8tgXgN8SjftXWxqm_htTvz5yYqMITdn_HMvfT35WpKPKyQiuK1L7oo0SWQ06ljmdWy11uNLDEN2Q9rud7CHRu92Xd6LkeTVwvSvTXQZwUMqAnHeqg-dPD6ovTZo1uMjy_qWqz60Kb-akTbmmnpgkLvMpIEqRJ5GGlC7WowprZXn96H5p-m_aVTweszH_3I60fYDQTjMZMVOqktP1hOFbW9hkRfpE7q6wVMlOzKqqut-AI67N9Lpjx8SyJw6cAmuGz8MUNfrH7NltuxKepuAcrPSFisZhxcaUdU8f06QHZbN0lbRKcV-FR2Gjb2FSoiYN0B2cMi4f4n1db-U3ZAfGgWB2aoWshKRLZpfCQgNYPZPMnRup180k_6fgIQQP87tyaahjfCxPjimr5NPZBHzaLNTJgVh6QwtMkeUjXm_UPherdSaGspF9xpyRG6agwgWHOTjVeUSHKs5jWbsHGf01QxrT-k3ly-KTEZZZGdqWbtyHJSeIYQjR464OmEsvUjQuas97iBL8nXWJIiIiFH84rZwRvhvKVQka4YeCn7P843AeyxTqhHvFojA47x9r3rp-wrm7PKfic7hQSPCvQWsKBUDdbnrfoSBQPu2GqV5LtQBMaWE-u6I4gVlJMra54NFZEPC0SPDqViN4-OKw28DBTYIM_9b50JEezNmJRQ7yqnxb2yQo_6HyKYFpIpWTCTHx6m9nOVCgFXrm5p4Hq7UeEuvH1s0dY4-jveDK67E_dds5fvxVm9_s4UPV0NGSBE5OmdF783S-W_NhRVZRXtz4fFh-S_uXonSHwZ4UwQ5BRlSXI7np8Nxc6Dl96hpkaeoeFmcJT2CJbzrZTcKybauGr3qWhsxne8ma7tup7dkmeI_5y64k9JDtV1SPWt1YryEHj7G-AoOzW8ezsw25L8ithjdxgXJ5ughkuXu-i75_JbxwtDhXtIWv4uoEbEL25cINdvKC0mXsmY-WPL-dCony_uKxpRwGoU-7fq6vCJTxAMTiSj4_j_dlpbslJmb4cton5hTltDGbKi5D63cd_96BDvxAcY-jicpGfo7HgP5iD2EOiuQMnAdxrbf9Gvk90oFDGC-GvII3tImkKAHEoL8R6yOPtkUJsi3GYsUllLo9g3eWbJ_W3M_cxm2euWYK0v35KQ2ynS7sQJvJxNieBZqpc2Aus0hJlyWyD3qjr9bLZZlZh8DiC-IKjP6M_KG0Eb9k-iaNqYitVuNAtqNDwd5ECx9iqy4bTq4wQAcA8VRS2SYU2ZzrCb7kVnsU8QWudaZkj1-4YBqedVHOwAlniFNowNy2Mmm2MokfPPtjvwqN_HPFQHZpRE33CfStLmTm0SlkohsHQSgdZ_ai0LrUri78iIGuSYC0meW_7hhXumJnAnRYOw4rn44xK39hfWF0NXgb1EPwxjRkuSsRUYyHmyXRR-i2MG7efoi6Lyb6PdKPylsx0P8K9bvcGhdr5VAUGgj9BYHJMr5TD3Efr6cjGGlNQqZnKXdTDNHfMoqFDXy0TYl9do2UD4sQdRtuJnT1YWhYZ2Hu-jzR_y-RqnWhFVR2LLwfqtOr3qGJ4cmLpEEMJ3C7GG6ZSCr6JaWwHKA2l_vyF-GBhhrNceOwcMHDDixVNKlGNvnTP2ltMrLoS7L6fE4Jz4BZiRsMg3hcV6j01oANFyXohbkCi4VW-AAogiBaQfHmCLCBzcEaiMbegh1blC1VW_XZKPF1cWozG9stQZ0981VfrT-7zJJ6UDiso-8Kjig2-gKswSfHJqHtADooWSx7gRvnnkimsy5SI_OkAbbnAAMXeSOIGXR5laJveqtOaarsK9FzdDGUuNkmBRfevsofnudfuw0tYw_og1cjsDdaR9Vu41m6qSVPe9Vd6XZ4E7Ou9rUxCVkTbb5Qe4w4ZznTG1SJqSYIiCe7szYXT5vWYXNRqQLuUZMQJwTHKuYvk4rrXpxh4SkfEBIyw7TUpzpeuqNktIHmZPXF-H3QlcVZqCZrUrieIx5Jy44l7OIS6PRPBbZvkKtmpJIfOF0CYefl-rO3XFN-HouoJe47Vrc2sehPxlNbxQ995wwiKR78YQW2jpUXZYx_rjD8fJPbOT9RjyO1XfzxgNg4zH6q9Qn6JAXe4czojpiW0ImSlhynv1pYjF3BztziJyaYx_uIXz-ZIAA3t0PtFjggovT0ArRM_Se25B2Jk6AmqMz_RZCnPCkpjXMGTM_8IGXsewwSK449Hptgcr3elCo9fABADF_9Ubpz2Syufp4S08TDeXLAA7BdjmpSI-5-pe7ldtp-T9D8OR-3WCyEZNopDwqRghR1hXNgqjKim5TM_GZEQcEzsZJv8pw_vHqrWQrmqVv1xRTahVIRQ_gH-VnrQpZYGqyenPIF8JVLh8Smqq9uHF6cf1OjLEUD87EEnj89TghLCDjuaAIeIKpKXsdVRGpdzbvntqWYb2XSaoqaBZerBuUbUqlmzlCyBQ5DCjcNZP6k9lf3Z5jY-IM-f1jKxPkauk0fGRQqmsu9CvizRSY67Meujh9ICdJqFbEOBkRKKNedUiYOIr1O-SeB9StjVXMlg5hKlZ8VlPsy7SbhPIY7ImE5R2ocIcYLZ8T65iMB2BmFXg5KvTRvZ8m1dleBZvMVbw9fDetHFpzi1FKC074wPXW-4_gR4r096eOSED9D8QGFKOatPwHSfLSZODH1eOxDKhYuv7WG65poddAbR8YxYvH6ko6umJx7rC_cndTOP2_7wWXHtr5jyjT4AJu7FMdb6jmu6vVEwMzdIc_DQ_c0PF4JRpANToiGn1an-mVCt9Atn-T8axUr5qY2S0Lcq8xAnSUZZZZVxILabf9Us-WLvnKbIzmFasSyxtjLNfFtVd1iR9Zp-3V5PG9mv94ujGNZfDGOG3hde9fLU-mufetGClmmFtocFmtHCju_B4wiwyEeZpZ-_msvrTrFhrFrPpAvwUDSJ85ILNBJja27O6K2sVdzMw9zcDoo0RA51OExZmDSf0aDCD-cb6q973Q54KtiBoTXTM8moYqndJlRAgzEFl1ZE1wY1hx7dNh8kLJgAuwh9MY8LT3bywBSJji-ITt1dnNDl-dH9JvtOdBJghWgkH6tc303M1a7MTC-fioZZ-nG5-B3W1Rd9TtzMSqts2jXTQyAgp2jPVsPykIZvf9fRbUoAxekiJB-4v-sYil6vdM2cLllVT4ned3u60W0T9emnjEAlVI9MA_Zy80XswSvLQ_oxcKY_mTB97SNpKCQU5RFnVJkjaWIMq-TbPmPo7dKfN-kpkx46x_vriyYi5ACjhvAgi7kqys6lExo7A5RoGAck2LCFzJRHjABSb4f0HvmiRRbYNVbgAjSAbZWzL6cj8MKGl5PK6mzkX-KwlVRP0N_b6kKqx41NiOc3ypxSdDWL7rxsQe5_Vqnd8ZgpnxedkKtBrAhMrbi4S8aGw==","internal_chat_message_metadata_passthrough":{"turn_id":"0b013667-0aaf-4a85-84aa-5b88e3033a8e"},"summary":[{"type":"summary_text","text":"**Flagging shellhost launch race**"},{"type":"summary_text","text":"**Assessing patch regressions**"},{"type":"summary_text","text":"**Tracing PID winner races**"}],"metadata":{"turn_id":"0b013667-0aaf-4a85-84aa-5b88e3033a8e"}},{"id":"fc_077f3569365e313d016aa1fde3a33c87d1a3f417003c017244","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading launch close stamp lifetime\",\"path\":\".worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs:514-720\"}","call_id":"call_4XqdQMELsGY7pnb5eiuhkd6i","internal_chat_message_metadata_passthrough":{"turn_id":"0b013667-0aaf-4a85-84aa-5b88e3033a8e"},"name":"read","metadata":{"turn_id":"0b013667-0aaf-4a85-84aa-5b88e3033a8e"}},{"id":"fc_077f3569365e313d016aa1fde3a35487d1a934b089426749e3","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading liveness verdict implementation\",\"path\":\".worktrees/287-shellrestore/crates/spt-store/src/liveness.rs:175-275\"}","call_id":"call_iriKaG5SXPRscFVxSnmTljAd","internal_chat_message_metadata_passthrough":{"turn_id":"0b013667-0aaf-4a85-84aa-5b88e3033a8e"},"name":"read","metadata":{"turn_id":"0b013667-0aaf-4a85-84aa-5b88e3033a8e"}},{"id":"fc_077f3569365e313d016aa1fde3a36487d1a8f349852acfa5b4","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating shell close implementation\",\"pattern\":\"pub fn close_shell|SHELL_LAUNCH_FILE|remove_file.*SHELL_PID_FILE|pub fn is_process_alive|pub fn process_started_at\",\"path\":\".worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs;.worktrees/287-shellrestore/crates/spt-store/src\",\"case\":true,\"gitignore\":false,\"skip\":null}","call_id":"call_1KIq0qgFYanlAYMVZvaDV02J","internal_chat_message_metadata_passthrough":{"turn_id":"0b013667-0aaf-4a85-84aa-5b88e3033a8e"},"name":"grep","metadata":{"turn_id":"0b013667-0aaf-4a85-84aa-5b88e3033a8e"}},{"id":"fc_077f3569365e313d016aa1fde3a36c87d193b54d2ea3b75ee5","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading wake binding exclusion tests\",\"path\":\".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs:1190-1470\"}","call_id":"call_VnQyIZqBQNs5YuGmMeW8r7gn","internal_chat_message_metadata_passthrough":{"turn_id":"0b013667-0aaf-4a85-84aa-5b88e3033a8e"},"name":"read","metadata":{"turn_id":"0b013667-0aaf-4a85-84aa-5b88e3033a8e"}},{"id":"fc_077f3569365e313d016aa1fde3a37887d1af10a446fd33cf00","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating protocol emission implementation\",\"pattern\":\"macro_rules! emit_line|fn emit_line|SHELL_RELINKED|status=binding\",\"path\":\".worktrees/287-shellrestore/crates/spt-proto/src;.worktrees/287-shellrestore/crates/spt-daemon/src\",\"case\":true,\"gitignore\":false,\"skip\":null}","call_id":"call_UpgQE4rfSVo6QpfvSZ2H1fCv","internal_chat_message_metadata_passthrough":{"turn_id":"0b013667-0aaf-4a85-84aa-5b88e3033a8e"},"name":"grep","metadata":{"turn_id":"0b013667-0aaf-4a85-84aa-5b88e3033a8e"}}]},"duration":26775.589699999895,"ttft":1308.0824000001885,"completedAt":1789001187736,"contextSnapshot":{"promptTokens":62896,"nonMessageTokens":13209,"compactionEpoch":0}}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_UpgQE4rfSVo6QpfvSZ2H1fCv|fc_077f3569365e313d016aa1fde3a37887d1af10a446fd33cf00","toolName":"grep","startedAt":"2026-09-10T00:46:27.738Z","args":{"path":".worktrees/287-shellrestore/crates/spt-proto/src;.worktrees/287-shellrestore/crates/spt-daemon/src"},"intent":"Locating protocol emission implementation"},"id":"f2d36b5d","parentId":"d823f8f3","timestamp":"2026-09-10T00:46:27.738Z"}
{"type":"message","id":"269b6e88","parentId":"f2d36b5d","timestamp":"2026-09-10T00:46:27.744Z","message":{"role":"toolResult","toolCallId":"call_iriKaG5SXPRscFVxSnmTljAd|fc_077f3569365e313d016aa1fde3a35487d1a934b089426749e3","toolName":"read","content":[{"type":"text","text":"/// daemon-earned liveness: it was stamped by the `api listen` process that bound\n/// the perch and holds its relay socket, and nothing in the daemon's\n/// broker-session model can see that process die (its OWNER — the harness — can\n/// outlive it, so the listen-orphan watchdog never fires either). This is the ONE\n/// predicate that answers \"is the relay we recorded still the process holding this\n/// endpoint\", and it lives HERE beside [`is_perch_alive`] so the verdict is never\n/// re-derived beside the resolver (KNOWN-HAZARDS 2.5).\n///\n/// Shaped like [`ProcIdentity`] on purpose: \"no knowledge\" is its own answer.\n#[derive(Debug, Clone, Copy, PartialEq, Eq)]\npub enum RelayLiveness {\n    /// The recorded relay process is still there (and still the one we recorded,\n    /// when a birth stamp is available to pair against).\n    Held,\n    /// The recorded relay is PROVABLY not holding the endpoint any more: absent\n    /// from a real process snapshot, or a pid now worn by a DIFFERENT process\n    /// (recycled — the relay we recorded is a corpse).\n    Gone,\n    /// No knowledge: no pid on the record, the `\"BUSY\"` sentinel, a `0` pid (a\n    /// backend that exposed none), no process table on this platform, or a live\n    /// pid whose start time could not be read. NEVER kills.\n    Unproven,\n}\n\n/// The relay-death verdict for one endpoint record (REQ-RELAY-DEATH-CONVERGENCE).\n///\n/// Routed through [`process_identity`] — the shipped custody/process-identity\n/// predicate (ADR-0047 decision 1) — NOT [`is_process_alive`]: existence is\n/// answered from the process TABLE, which is handle-immune, so this stays sound\n/// even for a pid some ancestor of the asker once held (KNOWN-HAZARDS 7.50), and\n/// `recorded_birth` turns bare existence into IDENTITY, so a recycled pid reads\n/// `Gone` instead of keeping a corpse's projection ONLINE forever.\n///\n/// - [`ProcIdentity::Absent`] ⇒ [`Gone`](RelayLiveness::Gone) — proof of death.\n/// - [`ProcIdentity::Present`] + a recorded birth that DIFFERS ⇒ `Gone` (the pid\n///   was recycled onto an unrelated process).\n/// - [`ProcIdentity::Present`] matching, or present with NO recorded birth (a\n///   pre-stamp record) ⇒ [`Held`](RelayLiveness::Held) — existence is all such a\n///   record can offer, and the missed heal is the safe direction.\n/// - [`ProcIdentity::Unproven`], no pid, or pid `0` ⇒\n///   [`Unproven`](RelayLiveness::Unproven) — fail toward alive, exactly the\n///   parity 2.5 holds for an unreadable pid. The heal can be MISSED but never\n///   MIS-FIRED.\n// [impl->REQ-RELAY-DEATH-CONVERGENCE]\npub fn relay_liveness(pid: Option<u32>, recorded_birth: Option<u64>) -> RelayLiveness {\n    // No pid / BUSY sentinel (`read_pid` ⇒ None), or the `0` a pid-less backend\n    // records: never a probe-able process, so never a death verdict.\n    let Some(pid) = pid.filter(|p| *p != 0) else {\n        return RelayLiveness::Unproven;\n    };\n    match process_identity(pid) {\n        ProcIdentity::Absent => RelayLiveness::Gone,\n        ProcIdentity::Unproven => RelayLiveness::Unproven,\n        ProcIdentity::Present(started) => match recorded_birth {\n            // The PAIR test: same pid, different process ⇒ our relay is gone.\n            Some(birth) if birth != started => RelayLiveness::Gone,\n            _ => RelayLiveness::Held,\n        },\n    }\n}\n\n#[cfg(test)]\nmod tests {\n    use super::*;\n    use crate::info::{set_status, write_info, InfoJson};\n\n    /// A pid that is never a real allocation (above Linux pid_max ~4.2M and\n    /// Windows ranges; positive as i32). Deterministic — avoids pid-reuse flake.\n    const DEAD_PID: u32 = 2_000_000_000;\n\n    fn perch() -> tempfile::TempDir {\n        tempfile::tempdir().unwrap()\n    }\n\n    // [unit->REQ-HAZARD-DAEMON-HOSTED-LIVENESS] a daemon-hosted perch marked\n    // `online` reads alive even though its recorded (summarizer) pid is DEAD —\n    // killing the summarizer subprocess does not flip the perch offline.\n    #[test]\n    fn daemon_hosted_online_survives_dead_pid() {\n        let d = perch();\n        // Record the ephemeral summarizer pid (now dead) then mark daemon-online.\n        write_info(\n            d.path(),\n            &InfoJson::new(\"doyle-psyche\", \"t\", DEAD_PID, \"sid\", \"psyche\"),\n        )\n        .unwrap();\n        set_status(d.path(), STATUS_ONLINE).unwrap();\n        assert!(\n            is_perch_alive(d.path()),\n            \"online status is authoritative; the dead pid is ignored\"\n        );\n    }\n\n    // [unit->REQ-HAZARD-DAEMON-HOSTED-LIVENESS] a daemon-hosted perch marked\n    // `offline` reads dead even though its recorded pid is ALIVE (this process).\n    #[test]\n    fn daemon_hosted_offline_overrides_live_pid() {\n        let d = perch();\n        write_info(\n            d.path(),\n            &InfoJson::new(\"doyle-psyche\", \"t\", std::process::id(), \"sid\", \"psyche\"),\n        )\n        .unwrap();\n        set_status(d.path(), STATUS_OFFLINE).unwrap();\n        assert!(\n…\n        );\n    }\n…\n}\n\n[Showing lines 174-278 of 501. Use :279 to continue]"}],"details":{"truncation":{"content":"/// daemon-earned liveness: it was stamped by the `api listen` process that bound\r\n/// the perch and holds its relay socket, and nothing in the daemon's\r\n/// broker-session model can see that process die (its OWNER — the harness — can\r\n/// outlive it, so the listen-orphan watchdog never fires either). This is the ONE\r\n/// predicate that answers \"is the relay we recorded still the process holding this\r\n/// endpoint\", and it lives HERE beside [`is_perch_alive`] so the verdict is never\r\n/// re-derived beside the resolver (KNOWN-HAZARDS 2.5).\r\n///\r\n/// Shaped like [`ProcIdentity`] on purpose: \"no knowledge\" is its own answer.\r\n#[derive(Debug, Clone, Copy, PartialEq, Eq)]\r\npub enum RelayLiveness {\r\n    /// The recorded relay process is still there (and still the one we recorded,\r\n    /// when a birth stamp is available to pair against).\r\n    Held,\r\n    /// The recorded relay is PROVABLY not holding the endpoint any more: absent\r\n    /// from a real process snapshot, or a pid now worn by a DIFFERENT process\r\n    /// (recycled — the relay we recorded is a corpse).\r\n    Gone,\r\n    /// No knowledge: no pid on the record, the `\"BUSY\"` sentinel, a `0` pid (a\r\n    /// backend that exposed none), no process table on this platform, or a live\r\n    /// pid whose start time could not be read. NEVER kills.\r\n    Unproven,\r\n}\r\n\r\n/// The relay-death verdict for one endpoint record (REQ-RELAY-DEATH-CONVERGENCE).\r\n///\r\n/// Routed through [`process_identity`] — the shipped custody/process-identity\r\n/// predicate (ADR-0047 decision 1) — NOT [`is_process_alive`]: existence is\r\n/// answered from the process TABLE, which is handle-immune, so this stays sound\r\n/// even for a pid some ancestor of the asker once held (KNOWN-HAZARDS 7.50), and\r\n/// `recorded_birth` turns bare existence into IDENTITY, so a recycled pid reads\r\n/// `Gone` instead of keeping a corpse's projection ONLINE forever.\r\n///\r\n/// - [`ProcIdentity::Absent`] ⇒ [`Gone`](RelayLiveness::Gone) — proof of death.\r\n/// - [`ProcIdentity::Present`] + a recorded birth that DIFFERS ⇒ `Gone` (the pid\r\n///   was recycled onto an unrelated process).\r\n/// - [`ProcIdentity::Present`] matching, or present with NO recorded birth (a\r\n///   pre-stamp record) ⇒ [`Held`](RelayLiveness::Held) — existence is all such a\r\n///   record can offer, and the missed heal is the safe direction.\r\n/// - [`ProcIdentity::Unproven`], no pid, or pid `0` ⇒\r\n///   [`Unproven`](RelayLiveness::Unproven) — fail toward alive, exactly the\r\n///   parity 2.5 holds for an unreadable pid. The heal can be MISSED but never\r\n///   MIS-FIRED.\r\n// [impl->REQ-RELAY-DEATH-CONVERGENCE]\r\npub fn relay_liveness(pid: Option<u32>, recorded_birth: Option<u64>) -> RelayLiveness {\r\n    // No pid / BUSY sentinel (`read_pid` ⇒ None), or the `0` a pid-less backend\r\n    // records: never a probe-able process, so never a death verdict.\r\n    let Some(pid) = pid.filter(|p| *p != 0) else {\r\n        return RelayLiveness::Unproven;\r\n    };\r\n    match process_identity(pid) {\r\n        ProcIdentity::Absent => RelayLiveness::Gone,\r\n        ProcIdentity::Unproven => RelayLiveness::Unproven,\r\n        ProcIdentity::Present(started) => match recorded_birth {\r\n            // The PAIR test: same pid, different process ⇒ our relay is gone.\r\n            Some(birth) if birth != started => RelayLiveness::Gone,\r\n            _ => RelayLiveness::Held,\r\n        },\r\n    }\r\n}\r\n\r\n#[cfg(test)]\r\nmod tests {\r\n    use super::*;\r\n    use crate::info::{set_status, write_info, InfoJson};\r\n\r\n    /// A pid that is never a real allocation (above Linux pid_max ~4.2M and\r\n    /// Windows ranges; positive as i32). Deterministic — avoids pid-reuse flake.\r\n    const DEAD_PID: u32 = 2_000_000_000;\r\n\r\n    fn perch() -> tempfile::TempDir {\r\n        tempfile::tempdir().unwrap()\r\n    }\r\n\r\n    // [unit->REQ-HAZARD-DAEMON-HOSTED-LIVENESS] a daemon-hosted perch marked\r\n    // `online` reads alive even though its recorded (summarizer) pid is DEAD —\r\n    // killing the summarizer subprocess does not flip the perch offline.\r\n    #[test]\r\n    fn daemon_hosted_online_survives_dead_pid() {\r\n        let d = perch();\r\n        // Record the ephemeral summarizer pid (now dead) then mark daemon-online.\r\n        write_info(\r\n            d.path(),\r\n            &InfoJson::new(\"doyle-psyche\", \"t\", DEAD_PID, \"sid\", \"psyche\"),\r\n        )\r\n        .unwrap();\r\n        set_status(d.path(), STATUS_ONLINE).unwrap();\r\n        assert!(\r\n            is_perch_alive(d.path()),\r\n            \"online status is authoritative; the dead pid is ignored\"\r\n        );\r\n    }\r\n\r\n    // [unit->REQ-HAZARD-DAEMON-HOSTED-LIVENESS] a daemon-hosted perch marked\r\n    // `offline` reads dead even though its recorded pid is ALIVE (this process).\r\n    #[test]\r\n    fn daemon_hosted_offline_overrides_live_pid() {\r\n        let d = perch();\r\n        write_info(\r\n            d.path(),\r\n            &InfoJson::new(\"doyle-psyche\", \"t\", std::process::id(), \"sid\", \"psyche\"),\r\n        )\r\n        .unwrap();\r\n        set_status(d.path(), STATUS_OFFLINE).unwrap();\r\n        assert!(\r","truncated":true,"truncatedBy":"lines","totalLines":328,"totalBytes":5068,"outputLines":105,"outputBytes":5068,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":501,"displayContent":{"text":"/// daemon-earned liveness: it was stamped by the `api listen` process that bound\r\n/// the perch and holds its relay socket, and nothing in the daemon's\r\n/// broker-session model can see that process die (its OWNER — the harness — can\r\n/// outlive it, so the listen-orphan watchdog never fires either). This is the ONE\r\n/// predicate that answers \"is the relay we recorded still the process holding this\r\n/// endpoint\", and it lives HERE beside [`is_perch_alive`] so the verdict is never\r\n/// re-derived beside the resolver (KNOWN-HAZARDS 2.5).\r\n///\r\n/// Shaped like [`ProcIdentity`] on purpose: \"no knowledge\" is its own answer.\r\n#[derive(Debug, Clone, Copy, PartialEq, Eq)]\r\npub enum RelayLiveness {\r\n    /// The recorded relay process is still there (and still the one we recorded,\r\n    /// when a birth stamp is available to pair against).\r\n    Held,\r\n    /// The recorded relay is PROVABLY not holding the endpoint any more: absent\r\n    /// from a real process snapshot, or a pid now worn by a DIFFERENT process\r\n    /// (recycled — the relay we recorded is a corpse).\r\n    Gone,\r\n    /// No knowledge: no pid on the record, the `\"BUSY\"` sentinel, a `0` pid (a\r\n    /// backend that exposed none), no process table on this platform, or a live\r\n    /// pid whose start time could not be read. NEVER kills.\r\n    Unproven,\r\n}\r\n\r\n/// The relay-death verdict for one endpoint record (REQ-RELAY-DEATH-CONVERGENCE).\r\n///\r\n/// Routed through [`process_identity`] — the shipped custody/process-identity\r\n/// predicate (ADR-0047 decision 1) — NOT [`is_process_alive`]: existence is\r\n/// answered from the process TABLE, which is handle-immune, so this stays sound\r\n/// even for a pid some ancestor of the asker once held (KNOWN-HAZARDS 7.50), and\r\n/// `recorded_birth` turns bare existence into IDENTITY, so a recycled pid reads\r\n/// `Gone` instead of keeping a corpse's projection ONLINE forever.\r\n///\r\n/// - [`ProcIdentity::Absent`] ⇒ [`Gone`](RelayLiveness::Gone) — proof of death.\r\n/// - [`ProcIdentity::Present`] + a recorded birth that DIFFERS ⇒ `Gone` (the pid\r\n///   was recycled onto an unrelated process).\r\n/// - [`ProcIdentity::Present`] matching, or present with NO recorded birth (a\r\n///   pre-stamp record) ⇒ [`Held`](RelayLiveness::Held) — existence is all such a\r\n///   record can offer, and the missed heal is the safe direction.\r\n/// - [`ProcIdentity::Unproven`], no pid, or pid `0` ⇒\r\n///   [`Unproven`](RelayLiveness::Unproven) — fail toward alive, exactly the\r\n///   parity 2.5 holds for an unreadable pid. The heal can be MISSED but never\r\n///   MIS-FIRED.\r\n// [impl->REQ-RELAY-DEATH-CONVERGENCE]\r\npub fn relay_liveness(pid: Option<u32>, recorded_birth: Option<u64>) -> RelayLiveness {\r\n    // No pid / BUSY sentinel (`read_pid` ⇒ None), or the `0` a pid-less backend\r\n    // records: never a probe-able process, so never a death verdict.\r\n    let Some(pid) = pid.filter(|p| *p != 0) else {\r\n        return RelayLiveness::Unproven;\r\n    };\r\n    match process_identity(pid) {\r\n        ProcIdentity::Absent => RelayLiveness::Gone,\r\n        ProcIdentity::Unproven => RelayLiveness::Unproven,\r\n        ProcIdentity::Present(started) => match recorded_birth {\r\n            // The PAIR test: same pid, different process ⇒ our relay is gone.\r\n            Some(birth) if birth != started => RelayLiveness::Gone,\r\n            _ => RelayLiveness::Held,\r\n        },\r\n    }\r\n}\r\n\r\n#[cfg(test)]\r\nmod tests {\r\n    use super::*;\r\n    use crate::info::{set_status, write_info, InfoJson};\r\n\r\n    /// A pid that is never a real allocation (above Linux pid_max ~4.2M and\r\n    /// Windows ranges; positive as i32). Deterministic — avoids pid-reuse flake.\r\n    const DEAD_PID: u32 = 2_000_000_000;\r\n\r\n    fn perch() -> tempfile::TempDir {\r\n        tempfile::tempdir().unwrap()\r\n    }\r\n\r\n    // [unit->REQ-HAZARD-DAEMON-HOSTED-LIVENESS] a daemon-hosted perch marked\r\n    // `online` reads alive even though its recorded (summarizer) pid is DEAD —\r\n    // killing the summarizer subprocess does not flip the perch offline.\r\n    #[test]\r\n    fn daemon_hosted_online_survives_dead_pid() {\r\n        let d = perch();\r\n        // Record the ephemeral summarizer pid (now dead) then mark daemon-online.\r\n        write_info(\r\n            d.path(),\r\n            &InfoJson::new(\"doyle-psyche\", \"t\", DEAD_PID, \"sid\", \"psyche\"),\r\n        )\r\n        .unwrap();\r\n        set_status(d.path(), STATUS_ONLINE).unwrap();\r\n        assert!(\r\n            is_perch_alive(d.path()),\r\n            \"online status is authoritative; the dead pid is ignored\"\r\n        );\r\n    }\r\n\r\n    // [unit->REQ-HAZARD-DAEMON-HOSTED-LIVENESS] a daemon-hosted perch marked\r\n    // `offline` reads dead even though its recorded pid is ALIVE (this process).\r\n    #[test]\r\n    fn daemon_hosted_offline_overrides_live_pid() {\r\n        let d = perch();\r\n        write_info(\r\n            d.path(),\r\n            &InfoJson::new(\"doyle-psyche\", \"t\", std::process::id(), \"sid\", \"psyche\"),\r\n        )\r\n        .unwrap();\r\n        set_status(d.path(), STATUS_OFFLINE).unwrap();\r\n        assert!(\r\n…\n        );\n    }\n…\n}","startLine":174,"lineNumbers":[174,175,176,177,178,179,180,181,182,183,184,185,186,187,188,189,190,191,192,193,194,195,196,197,198,199,200,201,202,203,204,205,206,207,208,209,210,211,212,213,214,215,216,217,218,219,220,221,222,223,224,225,226,227,228,229,230,231,232,233,234,235,236,237,238,239,240,241,242,243,244,245,246,247,248,249,250,251,252,253,254,255,256,257,258,259,260,261,262,263,264,265,266,267,268,269,270,271,272,273,274,275,276,277,278,null,281,282,null,501]},"fileSize":23062,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\crates\\spt-store\\src\\liveness.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":501,"totalBytes":5068,"outputLines":105,"outputBytes":5068,"shownRange":{"start":174,"end":278},"nextOffset":279}}},"isError":false,"timestamp":1789001187743}}
{"type":"message","id":"000cd841","parentId":"269b6e88","timestamp":"2026-09-10T00:46:27.746Z","message":{"role":"toolResult","toolCallId":"call_VnQyIZqBQNs5YuGmMeW8r7gn|fc_077f3569365e313d016aa1fde3a36c87d193b54d2ea3b75ee5","toolName":"read","content":[{"type":"text","text":"    fn only_a_launch_from_before_this_boot_is_a_restart_casualty() {\n…\n    }\n\n    fn register_restore_adapter(adapters_dir: &Path, source: &Path, name: &str, persistent: bool) {\n        #[cfg(windows)]\n        let noop = \"cmd /c exit 0\";\n        #[cfg(unix)]\n        let noop = \"true\";\n        std::fs::create_dir_all(source).unwrap();\n        std::fs::write(\n            source.join(\"manifest.toml\"),\n            format!(\n                \"[adapter]\\nname = \\\"{name}\\\"\\nkind = \\\"shell\\\"\\nversion = \\\"1\\\"\\n\\\n                 min_spt_core_version = \\\"0\\\"\\n\\n[shell]\\nspawn = '{noop}'\\n\\\n                 persistent = {persistent}\\n\"\n            ),\n        )\n        .unwrap();\n        spt_runtime::registry::register(adapters_dir, source, 1).unwrap();\n    }\n\n    /// Restart-shaped record, with NO suspend edge to pre-heal its status.\n    fn stranded_shell(owlery: &Path, owner: &str, adapter: &str, launched_ms: Option<u64>) -> String {\n        let id = spawn_record(owlery, owner, adapter, None).unwrap();\n        let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, &id);\n        let mut info = shellinfo::read_shell_info(&perch).unwrap();\n        info.status = SHELL_STATUS_ONLINE.to_string();\n        shellinfo::write_shell_info(&perch, &info).unwrap();\n        std::fs::write(perch.join(shellinfo::SHELL_PID_FILE), \"2000000000\").unwrap();\n        if let Some(launched_ms) = launched_ms {\n            shellinfo::record_shell_launch(&perch, 2_000_000_000, launched_ms);\n        }\n        id\n    }\n\n    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n    // [unit->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\n    #[test]\n    fn boot_restores_all_down_persistent_instances_without_launch_age_gates() {\n        let tmp = tempfile::tempdir().unwrap();\n        let owlery = tmp.path();\n        let boot = spt_store::proc::boot_instant_ms().expect(\"boot instant\");\n        let adapters_dir = tmp.path().join(\"adapters\");\n        register_restore_adapter(&adapters_dir, &tmp.path().join(\"keeper\"), \"keeper\", true);\n        register_restore_adapter(&adapters_dir, &tmp.path().join(\"ordinary\"), \"ordinary\", false);\n        let registered = spt_runtime::registry::registered(&adapters_dir);\n        seed_owner(owlery, \"ling\");\n        let gone = owlery.join(\"gone\");\n        std::fs::create_dir_all(&gone).unwrap();\n        std::fs::write(gone.join(\"info.json\"), \"{\\\"status\\\":\\\"offline\\\"}\").unwrap();\n\n        let before = stranded_shell(owlery, \"ling\", \"keeper\", Some(boot.saturating_sub(600_000)));\n        let after = stranded_shell(owlery, \"ling\", \"keeper\", Some(boot + 120_000));\n        let unstamped = stranded_shell(owlery, \"ling\", \"keeper\", None);\n        stranded_shell(owlery, \"gone\", \"keeper\", None);\n        stranded_shell(owlery, \"ling\", \"ordinary\", Some(boot + 120_000));\n\n        // A live binary awaiting bind still has an offline record. Removing the\n        // boot-age gate must not turn the sweep/edge overlap into a second launch.\n        let binding = spawn_record(owlery, \"ling\", \"keeper\", None).unwrap();\n        let binding_perch = spt_store::perch::resolve_shell_perch_path_in(owlery, \"ling\", &binding);\n        std::fs::write(\n            binding_perch.join(shellinfo::SHELL_PID_FILE),\n            std::process::id().to_string(),\n        )\n        .unwrap();\n        shellinfo::record_shell_launch(&binding_perch, std::process::id(), boot + 120_000);\n\n        let mut restored = restore_persistent_shells_at_boot(owlery, &registered, &adapters_dir);\n        restored.sort();\n        let mut expected = vec![\n            format!(\"ling/{before}\"),\n            format!(\"ling/{after}\"),\n            format!(\"ling/{unstamped}\"),\n        ];\n        expected.sort();\n        assert_eq!(restored, expected);\n        assert_eq!(shellinfo::read_shell_pid(&binding_perch), Some(std::process::id()));\n    }\n\n    // [unit->REQ-SHELL-OWNER-ONLINE-RESTORE]\n    // [unit->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\n    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n    #[test]\n    fn owner_online_restores_every_same_boot_or_unstamped_corpse_after_spent_sweep() {\n        let tmp = tempfile::tempdir().unwrap();\n        let owlery = tmp.path();\n        let boot = spt_store::proc::boot_instant_ms().expect(\"boot instant\");\n        let owner_dir = owlery.join(\"ling\");\n        std::fs::create_dir_all(&owner_dir).unwrap();\n        std::fs::write(owner_dir.join(\"info.json\"), \"{\\\"status\\\":\\\"offline\\\"}\").unwrap();\n        let adapters_dir = tmp.path().join(\"adapters\");\n        register_restore_adapter(&adapters_dir, &tmp.path().join(\"keeper\"), \"keeper\", true);\n        let registered = spt_runtime::registry::registered(&adapters_dir);\n        let first = stranded_shell(owlery, \"ling\", \"keeper\", Some(boot + 46_000));\n        let second = stranded_shell(owlery, \"ling\", \"keeper\", Some(boot + 120_000));\n        let unstamped = stranded_shell(owlery, \"ling\", \"keeper\", None);\n        let mut edge = OwnerOnlineEdge::new();\n        edge.seed(owlery);\n        assert!(restore_persistent_shells_at_boot(owlery, &registered, &adapters_dir).is_empty());\n\n        let set = Arc::new(WakeSet::new());\n        assert!(restore_persistent_shells_on_owner_online(\n            owlery, &registered, &adapters_dir, &mut edge, &set,\n        )\n        .is_empty());\n        std::fs::write(owner_dir.join(\"info.json\"), \"{}\").unwrap();\n        let mut restored =\n            restore_persistent_shells_on_owner_online(owlery, &registered, &adapters_dir, &mut edge, &set);\n        restored.sort();\n        let mut expected = vec![\n            format!(\"ling/{first}\"),\n            format!(\"ling/{second}\"),\n            format!(\"ling/{unstamped}\"),\n        ];\n        expected.sort();\n        assert_eq!(restored, expected, \"enumerate every restored instance of this owner\");\n\n        // Still an EDGE, not a permanent retry loop for failed launches.\n        stranded_shell(owlery, \"ling\", \"keeper\", Some(boot + 180_000));\n        assert!(restore_persistent_shells_on_owner_online(\n            owlery, &registered, &adapters_dir, &mut edge, &set,\n        )\n        .is_empty());\n    }\n\n    fn fast_params() -> WakeParams {\n        WakeParams {\n            backoff_base_ms: 10,\n            backoff_cap_ms: 40,\n            give_up_after: 3,\n        }\n    }\n\n    // [unit->REQ-INSTALL-11] the wake site, matching the spawn site: a bare\n    // program token binds to the shipped binary and {adapter_dir} fills. Wiring\n    // spawn alone would leave a released adapter launchable but PERMANENTLY\n    // UNWAKEABLE the moment it went offline — the worse half of the gap, because\n    // the wake path is the one that runs with nobody watching.\n    #[test]\n    fn a_wake_template_resolves_its_program_and_adapter_dir_against_the_install_dir() {\n        let tmp = tempfile::tempdir().unwrap();\n        let install = tmp.path().join(\"adapter dir\"); // spaces: the argv-fill shape\n        std::fs::create_dir_all(&install).unwrap();\n        let shipped = if cfg!(windows) { \"waker.exe\" } else { \"waker\" };\n        std::fs::write(install.join(shipped), b\"\").unwrap();\n\n        let tokens =\n            fill_wake_command(\"sh-1\", \"mock-shell\", Some(&install), \"waker --root {adapter_dir}\")\n                .expect(\"adapter_dir is a wake substitution key\");\n        assert_eq!(\n            tokens[0],\n            install.join(shipped).display().to_string(),\n            \"the bare token must bind to the SHIPPED binary, not fall through to PATH\"\n        );\n        assert_eq!(\n            tokens[2],\n            install.display().to_string(),\n            \"the install dir (spaces included) is exactly one argv element\"\n        );\n    }\n\n    // [unit->REQ-INSTALL-11] the catalogs share a VOCABULARY, they are not one\n    // set — the relationship the corrected `fill_wake_command` docstring states,\n    // pinned so a future edit cannot quietly restore the set-equality claim.\n    // `adapter_dir` is in BOTH; `perch_dir` is spawn-only by design (a waker has\n    // no live link, so there is no perch-relative transfer path to resolve).\n    #[test]\n    fn the_wake_catalog_shares_the_spawn_vocabulary_but_is_not_the_same_set() {\n        let install = std::path::PathBuf::from(\"/adapters/mock\");\n        assert!(\n            fill_wake_command(\"sh-1\", \"mock-shell\", Some(&install), \"w {adapter_dir}\").is_ok(),\n            \"adapter_dir is shared by both catalogs\"\n        );\n        assert!(\n            fill_wake_command(\"sh-1\", \"mock-shell\", Some(&install), \"w {perch_dir}\").is_err(),\n            \"perch_dir is SPAWN-ONLY — a waker has no live link to resolve paths against\"\n        );\n    }\n\n    fn seed_owner(owlery: &Path, owner: &str) {\n        let p = owlery.join(owner);\n        std::fs::create_dir_all(&p).unwrap();\n        std::fs::write(p.join(\"info.json\"), \"{}\").unwrap();\n    }\n\n    /// Tokens for a waker that exits with `code` immediately.\n    fn exit_with(code: i32) -> Vec<String> {\n        #[cfg(windows)]\n        return vec![\"cmd\".into(), \"/c\".into(), format!(\"exit {code}\")];\n        #[cfg(unix)]\n        return vec![\"sh\".into(), \"-c\".into(), format!(\"exit {code}\")];\n    }\n\n    /// Tokens for a waker that appends one line to `tally` then exits 1.\n    fn crash_tallying(tally: &Path) -> Vec<String> {\n        #[cfg(windows)]\n        return vec![\n            \"powershell\".into(),\n            \"-NoProfile\".into(),\n            \"-Command\".into(),\n            format!(\"Add-Content -Path '{}' -Value x; exit 1\", tally.display()),\n        ];\n        #[cfg(unix)]\n        return vec![\n            \"sh\".into(),\n            \"-c\".into(),\n            format!(\"echo x >> '{}'; exit 1\", tally.display()),\n        ];\n    }\n\n    // [unit->REQ-SHELL-2] the backoff curve: base × 2^(n-1), saturating at\n    // the cap; zero failures wait nothing.\n    #[test]\n    fn backoff_curve_doubles_to_the_cap() {\n        let p = WakeParams {\n            backoff_base_ms: 100,\n            backoff_cap_ms: 1_500,\n            give_up_after: 6,\n        };\n        assert_eq!(backoff_ms(&p, 0), 0);\n        assert_eq!(backoff_ms(&p, 1), 100);\n        assert_eq!(backoff_ms(&p, 2), 200);\n        assert_eq!(backoff_ms(&p, 4), 800);\n        assert_eq!(backoff_ms(&p, 5), 1_500, \"saturates at the cap\");\n        assert_eq!(\n            backoff_ms(&p, 60),\n            1_500,\n            \"huge failure counts can't overflow\"\n        );\n    }\n\n    // [unit->REQ-SHELL-2] exit-opcode supervision: exit(86) fires the wake\n    // resolution exactly once and the loop ends; the pid file is retired.\n    #[test]\n    fn watcher_opcode_exit_fires_resolution_once() {\n        let tmp = tempfile::tempdir().unwrap();\n        let owlery = tmp.path();\n        seed_owner(owlery, \"doyle\");\n        let id = spawn_record(owlery, \"doyle\", \"mock-shell\", None).unwrap();\n        let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, \"doyle\", &id);\n\n        let fired = std::sync::atomic::AtomicU32::new(0);\n        let stop = AtomicBool::new(false);\n        watcher_run(\n            owlery,\n            \"doyle\",\n            &id,\n            &exit_with(WAKE_OPCODE),\n            &fast_params(),\n            &stop,\n            || {\n                fired.fetch_add(1, Ordering::SeqCst);\n                Ok(\"test-resolved\".into())\n            },\n        );\n        assert_eq!(\n            fired.load(Ordering::SeqCst),\n            1,\n            \"resolution fired exactly once\"\n        );\n        assert!(!perch.join(WAKER_PID_FILE).exists(), \"pid file retired\");\n        assert!(\n            !perch.join(WAKER_GAVE_UP_FILE).exists(),\n            \"an opcode exit is not a crash\"\n        );\n    }\n\n    // [unit->REQ-SHELL-2] crash-exit supervision: a crashing waker respawns\n    // (with backoff) exactly give_up_after times, then the durable give-up\n    // latch drops and the resolution never fires; clear_gave_up re-arms.\n    #[test]\n    fn watcher_crash_exits_respawn_then_give_up() {\n        let tmp = tempfile::tempdir().unwrap();\n        let owlery = tmp.path();\n        seed_owner(owlery, \"doyle\");\n        let id = spawn_record(owlery, \"doyle\", \"mock-shell\", None).unwrap();\n        let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, \"doyle\", &id);\n        let tally = perch.join(\"waker-tally.txt\");\n\n        let stop = AtomicBool::new(false);\n        watcher_run(\n            owlery,\n            \"doyle\",\n            &id,\n            &crash_tallying(&tally),\n…\n        );\n…\n    }\n\n[Showing lines 1189-1473 of 2161. Use :1474 to continue]"}],"details":{"truncation":{"content":"    }\r\n\r\n    fn register_restore_adapter(adapters_dir: &Path, source: &Path, name: &str, persistent: bool) {\r\n        #[cfg(windows)]\r\n        let noop = \"cmd /c exit 0\";\r\n        #[cfg(unix)]\r\n        let noop = \"true\";\r\n        std::fs::create_dir_all(source).unwrap();\r\n        std::fs::write(\r\n            source.join(\"manifest.toml\"),\r\n            format!(\r\n                \"[adapter]\\nname = \\\"{name}\\\"\\nkind = \\\"shell\\\"\\nversion = \\\"1\\\"\\n\\\r\n                 min_spt_core_version = \\\"0\\\"\\n\\n[shell]\\nspawn = '{noop}'\\n\\\r\n                 persistent = {persistent}\\n\"\r\n            ),\r\n        )\r\n        .unwrap();\r\n        spt_runtime::registry::register(adapters_dir, source, 1).unwrap();\r\n    }\r\n\r\n    /// Restart-shaped record, with NO suspend edge to pre-heal its status.\r\n    fn stranded_shell(owlery: &Path, owner: &str, adapter: &str, launched_ms: Option<u64>) -> String {\r\n        let id = spawn_record(owlery, owner, adapter, None).unwrap();\r\n        let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, &id);\r\n        let mut info = shellinfo::read_shell_info(&perch).unwrap();\r\n        info.status = SHELL_STATUS_ONLINE.to_string();\r\n        shellinfo::write_shell_info(&perch, &info).unwrap();\r\n        std::fs::write(perch.join(shellinfo::SHELL_PID_FILE), \"2000000000\").unwrap();\r\n        if let Some(launched_ms) = launched_ms {\r\n            shellinfo::record_shell_launch(&perch, 2_000_000_000, launched_ms);\r\n        }\r\n        id\r\n    }\r\n\r\n    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\n    // [unit->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\r\n    #[test]\r\n    fn boot_restores_all_down_persistent_instances_without_launch_age_gates() {\r\n        let tmp = tempfile::tempdir().unwrap();\r\n        let owlery = tmp.path();\r\n        let boot = spt_store::proc::boot_instant_ms().expect(\"boot instant\");\r\n        let adapters_dir = tmp.path().join(\"adapters\");\r\n        register_restore_adapter(&adapters_dir, &tmp.path().join(\"keeper\"), \"keeper\", true);\r\n        register_restore_adapter(&adapters_dir, &tmp.path().join(\"ordinary\"), \"ordinary\", false);\r\n        let registered = spt_runtime::registry::registered(&adapters_dir);\r\n        seed_owner(owlery, \"ling\");\r\n        let gone = owlery.join(\"gone\");\r\n        std::fs::create_dir_all(&gone).unwrap();\r\n        std::fs::write(gone.join(\"info.json\"), \"{\\\"status\\\":\\\"offline\\\"}\").unwrap();\r\n\r\n        let before = stranded_shell(owlery, \"ling\", \"keeper\", Some(boot.saturating_sub(600_000)));\r\n        let after = stranded_shell(owlery, \"ling\", \"keeper\", Some(boot + 120_000));\r\n        let unstamped = stranded_shell(owlery, \"ling\", \"keeper\", None);\r\n        stranded_shell(owlery, \"gone\", \"keeper\", None);\r\n        stranded_shell(owlery, \"ling\", \"ordinary\", Some(boot + 120_000));\r\n\r\n        // A live binary awaiting bind still has an offline record. Removing the\r\n        // boot-age gate must not turn the sweep/edge overlap into a second launch.\r\n        let binding = spawn_record(owlery, \"ling\", \"keeper\", None).unwrap();\r\n        let binding_perch = spt_store::perch::resolve_shell_perch_path_in(owlery, \"ling\", &binding);\r\n        std::fs::write(\r\n            binding_perch.join(shellinfo::SHELL_PID_FILE),\r\n            std::process::id().to_string(),\r\n        )\r\n        .unwrap();\r\n        shellinfo::record_shell_launch(&binding_perch, std::process::id(), boot + 120_000);\r\n\r\n        let mut restored = restore_persistent_shells_at_boot(owlery, &registered, &adapters_dir);\r\n        restored.sort();\r\n        let mut expected = vec![\r\n            format!(\"ling/{before}\"),\r\n            format!(\"ling/{after}\"),\r\n            format!(\"ling/{unstamped}\"),\r\n        ];\r\n        expected.sort();\r\n        assert_eq!(restored, expected);\r\n        assert_eq!(shellinfo::read_shell_pid(&binding_perch), Some(std::process::id()));\r\n    }\r\n\r\n    // [unit->REQ-SHELL-OWNER-ONLINE-RESTORE]\r\n    // [unit->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\r\n    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\n    #[test]\r\n    fn owner_online_restores_every_same_boot_or_unstamped_corpse_after_spent_sweep() {\r\n        let tmp = tempfile::tempdir().unwrap();\r\n        let owlery = tmp.path();\r\n        let boot = spt_store::proc::boot_instant_ms().expect(\"boot instant\");\r\n        let owner_dir = owlery.join(\"ling\");\r\n        std::fs::create_dir_all(&owner_dir).unwrap();\r\n        std::fs::write(owner_dir.join(\"info.json\"), \"{\\\"status\\\":\\\"offline\\\"}\").unwrap();\r\n        let adapters_dir = tmp.path().join(\"adapters\");\r\n        register_restore_adapter(&adapters_dir, &tmp.path().join(\"keeper\"), \"keeper\", true);\r\n        let registered = spt_runtime::registry::registered(&adapters_dir);\r\n        let first = stranded_shell(owlery, \"ling\", \"keeper\", Some(boot + 46_000));\r\n        let second = stranded_shell(owlery, \"ling\", \"keeper\", Some(boot + 120_000));\r\n        let unstamped = stranded_shell(owlery, \"ling\", \"keeper\", None);\r\n        let mut edge = OwnerOnlineEdge::new();\r\n        edge.seed(owlery);\r\n        assert!(restore_persistent_shells_at_boot(owlery, &registered, &adapters_dir).is_empty());\r\n\r\n        let set = Arc::new(WakeSet::new());\r\n        assert!(restore_persistent_shells_on_owner_online(\r\n            owlery, &registered, &adapters_dir, &mut edge, &set,\r\n        )\r\n        .is_empty());\r\n        std::fs::write(owner_dir.join(\"info.json\"), \"{}\").unwrap();\r\n        let mut restored =\r\n            restore_persistent_shells_on_owner_online(owlery, &registered, &adapters_dir, &mut edge, &set);\r\n        restored.sort();\r\n        let mut expected = vec![\r\n            format!(\"ling/{first}\"),\r\n            format!(\"ling/{second}\"),\r\n            format!(\"ling/{unstamped}\"),\r\n        ];\r\n        expected.sort();\r\n        assert_eq!(restored, expected, \"enumerate every restored instance of this owner\");\r\n\r\n        // Still an EDGE, not a permanent retry loop for failed launches.\r\n        stranded_shell(owlery, \"ling\", \"keeper\", Some(boot + 180_000));\r\n        assert!(restore_persistent_shells_on_owner_online(\r\n            owlery, &registered, &adapters_dir, &mut edge, &set,\r\n        )\r\n        .is_empty());\r\n    }\r\n\r\n    fn fast_params() -> WakeParams {\r\n        WakeParams {\r\n            backoff_base_ms: 10,\r\n            backoff_cap_ms: 40,\r\n            give_up_after: 3,\r\n        }\r\n    }\r\n\r\n    // [unit->REQ-INSTALL-11] the wake site, matching the spawn site: a bare\r\n    // program token binds to the shipped binary and {adapter_dir} fills. Wiring\r\n    // spawn alone would leave a released adapter launchable but PERMANENTLY\r\n    // UNWAKEABLE the moment it went offline — the worse half of the gap, because\r\n    // the wake path is the one that runs with nobody watching.\r\n    #[test]\r\n    fn a_wake_template_resolves_its_program_and_adapter_dir_against_the_install_dir() {\r\n        let tmp = tempfile::tempdir().unwrap();\r\n        let install = tmp.path().join(\"adapter dir\"); // spaces: the argv-fill shape\r\n        std::fs::create_dir_all(&install).unwrap();\r\n        let shipped = if cfg!(windows) { \"waker.exe\" } else { \"waker\" };\r\n        std::fs::write(install.join(shipped), b\"\").unwrap();\r\n\r\n        let tokens =\r\n            fill_wake_command(\"sh-1\", \"mock-shell\", Some(&install), \"waker --root {adapter_dir}\")\r\n                .expect(\"adapter_dir is a wake substitution key\");\r\n        assert_eq!(\r\n            tokens[0],\r\n            install.join(shipped).display().to_string(),\r\n            \"the bare token must bind to the SHIPPED binary, not fall through to PATH\"\r\n        );\r\n        assert_eq!(\r\n            tokens[2],\r\n            install.display().to_string(),\r\n            \"the install dir (spaces included) is exactly one argv element\"\r\n        );\r\n    }\r\n\r\n    // [unit->REQ-INSTALL-11] the catalogs share a VOCABULARY, they are not one\r\n    // set — the relationship the corrected `fill_wake_command` docstring states,\r\n    // pinned so a future edit cannot quietly restore the set-equality claim.\r\n    // `adapter_dir` is in BOTH; `perch_dir` is spawn-only by design (a waker has\r\n    // no live link, so there is no perch-relative transfer path to resolve).\r\n    #[test]\r\n    fn the_wake_catalog_shares_the_spawn_vocabulary_but_is_not_the_same_set() {\r\n        let install = std::path::PathBuf::from(\"/adapters/mock\");\r\n        assert!(\r\n            fill_wake_command(\"sh-1\", \"mock-shell\", Some(&install), \"w {adapter_dir}\").is_ok(),\r\n            \"adapter_dir is shared by both catalogs\"\r\n        );\r\n        assert!(\r\n            fill_wake_command(\"sh-1\", \"mock-shell\", Some(&install), \"w {perch_dir}\").is_err(),\r\n            \"perch_dir is SPAWN-ONLY — a waker has no live link to resolve paths against\"\r\n        );\r\n    }\r\n\r\n    fn seed_owner(owlery: &Path, owner: &str) {\r\n        let p = owlery.join(owner);\r\n        std::fs::create_dir_all(&p).unwrap();\r\n        std::fs::write(p.join(\"info.json\"), \"{}\").unwrap();\r\n    }\r\n\r\n    /// Tokens for a waker that exits with `code` immediately.\r\n    fn exit_with(code: i32) -> Vec<String> {\r\n        #[cfg(windows)]\r\n        return vec![\"cmd\".into(), \"/c\".into(), format!(\"exit {code}\")];\r\n        #[cfg(unix)]\r\n        return vec![\"sh\".into(), \"-c\".into(), format!(\"exit {code}\")];\r\n    }\r\n\r\n    /// Tokens for a waker that appends one line to `tally` then exits 1.\r\n    fn crash_tallying(tally: &Path) -> Vec<String> {\r\n        #[cfg(windows)]\r\n        return vec![\r\n            \"powershell\".into(),\r\n            \"-NoProfile\".into(),\r\n            \"-Command\".into(),\r\n            format!(\"Add-Content -Path '{}' -Value x; exit 1\", tally.display()),\r\n        ];\r\n        #[cfg(unix)]\r\n        return vec![\r\n            \"sh\".into(),\r\n            \"-c\".into(),\r\n            format!(\"echo x >> '{}'; exit 1\", tally.display()),\r\n        ];\r\n    }\r\n\r\n    // [unit->REQ-SHELL-2] the backoff curve: base × 2^(n-1), saturating at\r\n    // the cap; zero failures wait nothing.\r\n    #[test]\r\n    fn backoff_curve_doubles_to_the_cap() {\r\n        let p = WakeParams {\r\n            backoff_base_ms: 100,\r\n            backoff_cap_ms: 1_500,\r\n            give_up_after: 6,\r\n        };\r\n        assert_eq!(backoff_ms(&p, 0), 0);\r\n        assert_eq!(backoff_ms(&p, 1), 100);\r\n        assert_eq!(backoff_ms(&p, 2), 200);\r\n        assert_eq!(backoff_ms(&p, 4), 800);\r\n        assert_eq!(backoff_ms(&p, 5), 1_500, \"saturates at the cap\");\r\n        assert_eq!(\r\n            backoff_ms(&p, 60),\r\n            1_500,\r\n            \"huge failure counts can't overflow\"\r\n        );\r\n    }\r\n\r\n    // [unit->REQ-SHELL-2] exit-opcode supervision: exit(86) fires the wake\r\n    // resolution exactly once and the loop ends; the pid file is retired.\r\n    #[test]\r\n    fn watcher_opcode_exit_fires_resolution_once() {\r\n        let tmp = tempfile::tempdir().unwrap();\r\n        let owlery = tmp.path();\r\n        seed_owner(owlery, \"doyle\");\r\n        let id = spawn_record(owlery, \"doyle\", \"mock-shell\", None).unwrap();\r\n        let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, \"doyle\", &id);\r\n\r\n        let fired = std::sync::atomic::AtomicU32::new(0);\r\n        let stop = AtomicBool::new(false);\r\n        watcher_run(\r\n            owlery,\r\n            \"doyle\",\r\n            &id,\r\n            &exit_with(WAKE_OPCODE),\r\n            &fast_params(),\r\n            &stop,\r\n            || {\r\n                fired.fetch_add(1, Ordering::SeqCst);\r\n                Ok(\"test-resolved\".into())\r\n            },\r\n        );\r\n        assert_eq!(\r\n            fired.load(Ordering::SeqCst),\r\n            1,\r\n            \"resolution fired exactly once\"\r\n        );\r\n        assert!(!perch.join(WAKER_PID_FILE).exists(), \"pid file retired\");\r\n        assert!(\r\n            !perch.join(WAKER_GAVE_UP_FILE).exists(),\r\n            \"an opcode exit is not a crash\"\r\n        );\r\n    }\r\n\r\n    // [unit->REQ-SHELL-2] crash-exit supervision: a crashing waker respawns\r\n    // (with backoff) exactly give_up_after times, then the durable give-up\r\n    // latch drops and the resolution never fires; clear_gave_up re-arms.\r\n    #[test]\r\n    fn watcher_crash_exits_respawn_then_give_up() {\r\n        let tmp = tempfile::tempdir().unwrap();\r\n        let owlery = tmp.path();\r\n        seed_owner(owlery, \"doyle\");\r\n        let id = spawn_record(owlery, \"doyle\", \"mock-shell\", None).unwrap();\r\n        let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, \"doyle\", &id);\r\n        let tally = perch.join(\"waker-tally.txt\");\r\n\r\n        let stop = AtomicBool::new(false);\r\n        watcher_run(\r\n            owlery,\r\n            \"doyle\",\r\n            &id,\r\n            &crash_tallying(&tally),\r","truncated":true,"truncatedBy":"lines","totalLines":973,"totalBytes":12562,"outputLines":285,"outputBytes":12562,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":2161,"displayContent":{"text":"    fn only_a_launch_from_before_this_boot_is_a_restart_casualty() {\n…\n    }\r\n\r\n    fn register_restore_adapter(adapters_dir: &Path, source: &Path, name: &str, persistent: bool) {\r\n        #[cfg(windows)]\r\n        let noop = \"cmd /c exit 0\";\r\n        #[cfg(unix)]\r\n        let noop = \"true\";\r\n        std::fs::create_dir_all(source).unwrap();\r\n        std::fs::write(\r\n            source.join(\"manifest.toml\"),\r\n            format!(\r\n                \"[adapter]\\nname = \\\"{name}\\\"\\nkind = \\\"shell\\\"\\nversion = \\\"1\\\"\\n\\\r\n                 min_spt_core_version = \\\"0\\\"\\n\\n[shell]\\nspawn = '{noop}'\\n\\\r\n                 persistent = {persistent}\\n\"\r\n            ),\r\n        )\r\n        .unwrap();\r\n        spt_runtime::registry::register(adapters_dir, source, 1).unwrap();\r\n    }\r\n\r\n    /// Restart-shaped record, with NO suspend edge to pre-heal its status.\r\n    fn stranded_shell(owlery: &Path, owner: &str, adapter: &str, launched_ms: Option<u64>) -> String {\r\n        let id = spawn_record(owlery, owner, adapter, None).unwrap();\r\n        let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, &id);\r\n        let mut info = shellinfo::read_shell_info(&perch).unwrap();\r\n        info.status = SHELL_STATUS_ONLINE.to_string();\r\n        shellinfo::write_shell_info(&perch, &info).unwrap();\r\n        std::fs::write(perch.join(shellinfo::SHELL_PID_FILE), \"2000000000\").unwrap();\r\n        if let Some(launched_ms) = launched_ms {\r\n            shellinfo::record_shell_launch(&perch, 2_000_000_000, launched_ms);\r\n        }\r\n        id\r\n    }\r\n\r\n    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\n    // [unit->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\r\n    #[test]\r\n    fn boot_restores_all_down_persistent_instances_without_launch_age_gates() {\r\n        let tmp = tempfile::tempdir().unwrap();\r\n        let owlery = tmp.path();\r\n        let boot = spt_store::proc::boot_instant_ms().expect(\"boot instant\");\r\n        let adapters_dir = tmp.path().join(\"adapters\");\r\n        register_restore_adapter(&adapters_dir, &tmp.path().join(\"keeper\"), \"keeper\", true);\r\n        register_restore_adapter(&adapters_dir, &tmp.path().join(\"ordinary\"), \"ordinary\", false);\r\n        let registered = spt_runtime::registry::registered(&adapters_dir);\r\n        seed_owner(owlery, \"ling\");\r\n        let gone = owlery.join(\"gone\");\r\n        std::fs::create_dir_all(&gone).unwrap();\r\n        std::fs::write(gone.join(\"info.json\"), \"{\\\"status\\\":\\\"offline\\\"}\").unwrap();\r\n\r\n        let before = stranded_shell(owlery, \"ling\", \"keeper\", Some(boot.saturating_sub(600_000)));\r\n        let after = stranded_shell(owlery, \"ling\", \"keeper\", Some(boot + 120_000));\r\n        let unstamped = stranded_shell(owlery, \"ling\", \"keeper\", None);\r\n        stranded_shell(owlery, \"gone\", \"keeper\", None);\r\n        stranded_shell(owlery, \"ling\", \"ordinary\", Some(boot + 120_000));\r\n\r\n        // A live binary awaiting bind still has an offline record. Removing the\r\n        // boot-age gate must not turn the sweep/edge overlap into a second launch.\r\n        let binding = spawn_record(owlery, \"ling\", \"keeper\", None).unwrap();\r\n        let binding_perch = spt_store::perch::resolve_shell_perch_path_in(owlery, \"ling\", &binding);\r\n        std::fs::write(\r\n            binding_perch.join(shellinfo::SHELL_PID_FILE),\r\n            std::process::id().to_string(),\r\n        )\r\n        .unwrap();\r\n        shellinfo::record_shell_launch(&binding_perch, std::process::id(), boot + 120_000);\r\n\r\n        let mut restored = restore_persistent_shells_at_boot(owlery, &registered, &adapters_dir);\r\n        restored.sort();\r\n        let mut expected = vec![\r\n            format!(\"ling/{before}\"),\r\n            format!(\"ling/{after}\"),\r\n            format!(\"ling/{unstamped}\"),\r\n        ];\r\n        expected.sort();\r\n        assert_eq!(restored, expected);\r\n        assert_eq!(shellinfo::read_shell_pid(&binding_perch), Some(std::process::id()));\r\n    }\r\n\r\n    // [unit->REQ-SHELL-OWNER-ONLINE-RESTORE]\r\n    // [unit->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]\r\n    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\n    #[test]\r\n    fn owner_online_restores_every_same_boot_or_unstamped_corpse_after_spent_sweep() {\r\n        let tmp = tempfile::tempdir().unwrap();\r\n        let owlery = tmp.path();\r\n        let boot = spt_store::proc::boot_instant_ms().expect(\"boot instant\");\r\n        let owner_dir = owlery.join(\"ling\");\r\n        std::fs::create_dir_all(&owner_dir).unwrap();\r\n        std::fs::write(owner_dir.join(\"info.json\"), \"{\\\"status\\\":\\\"offline\\\"}\").unwrap();\r\n        let adapters_dir = tmp.path().join(\"adapters\");\r\n        register_restore_adapter(&adapters_dir, &tmp.path().join(\"keeper\"), \"keeper\", true);\r\n        let registered = spt_runtime::registry::registered(&adapters_dir);\r\n        let first = stranded_shell(owlery, \"ling\", \"keeper\", Some(boot + 46_000));\r\n        let second = stranded_shell(owlery, \"ling\", \"keeper\", Some(boot + 120_000));\r\n        let unstamped = stranded_shell(owlery, \"ling\", \"keeper\", None);\r\n        let mut edge = OwnerOnlineEdge::new();\r\n        edge.seed(owlery);\r\n        assert!(restore_persistent_shells_at_boot(owlery, &registered, &adapters_dir).is_empty());\r\n\r\n        let set = Arc::new(WakeSet::new());\r\n        assert!(restore_persistent_shells_on_owner_online(\r\n            owlery, &registered, &adapters_dir, &mut edge, &set,\r\n        )\r\n        .is_empty());\r\n        std::fs::write(owner_dir.join(\"info.json\"), \"{}\").unwrap();\r\n        let mut restored =\r\n            restore_persistent_shells_on_owner_online(owlery, &registered, &adapters_dir, &mut edge, &set);\r\n        restored.sort();\r\n        let mut expected = vec![\r\n            format!(\"ling/{first}\"),\r\n            format!(\"ling/{second}\"),\r\n            format!(\"ling/{unstamped}\"),\r\n        ];\r\n        expected.sort();\r\n        assert_eq!(restored, expected, \"enumerate every restored instance of this owner\");\r\n\r\n        // Still an EDGE, not a permanent retry loop for failed launches.\r\n        stranded_shell(owlery, \"ling\", \"keeper\", Some(boot + 180_000));\r\n        assert!(restore_persistent_shells_on_owner_online(\r\n            owlery, &registered, &adapters_dir, &mut edge, &set,\r\n        )\r\n        .is_empty());\r\n    }\r\n\r\n    fn fast_params() -> WakeParams {\r\n        WakeParams {\r\n            backoff_base_ms: 10,\r\n            backoff_cap_ms: 40,\r\n            give_up_after: 3,\r\n        }\r\n    }\r\n\r\n    // [unit->REQ-INSTALL-11] the wake site, matching the spawn site: a bare\r\n    // program token binds to the shipped binary and {adapter_dir} fills. Wiring\r\n    // spawn alone would leave a released adapter launchable but PERMANENTLY\r\n    // UNWAKEABLE the moment it went offline — the worse half of the gap, because\r\n    // the wake path is the one that runs with nobody watching.\r\n    #[test]\r\n    fn a_wake_template_resolves_its_program_and_adapter_dir_against_the_install_dir() {\r\n        let tmp = tempfile::tempdir().unwrap();\r\n        let install = tmp.path().join(\"adapter dir\"); // spaces: the argv-fill shape\r\n        std::fs::create_dir_all(&install).unwrap();\r\n        let shipped = if cfg!(windows) { \"waker.exe\" } else { \"waker\" };\r\n        std::fs::write(install.join(shipped), b\"\").unwrap();\r\n\r\n        let tokens =\r\n            fill_wake_command(\"sh-1\", \"mock-shell\", Some(&install), \"waker --root {adapter_dir}\")\r\n                .expect(\"adapter_dir is a wake substitution key\");\r\n        assert_eq!(\r\n            tokens[0],\r\n            install.join(shipped).display().to_string(),\r\n            \"the bare token must bind to the SHIPPED binary, not fall through to PATH\"\r\n        );\r\n        assert_eq!(\r\n            tokens[2],\r\n            install.display().to_string(),\r\n            \"the install dir (spaces included) is exactly one argv element\"\r\n        );\r\n    }\r\n\r\n    // [unit->REQ-INSTALL-11] the catalogs share a VOCABULARY, they are not one\r\n    // set — the relationship the corrected `fill_wake_command` docstring states,\r\n    // pinned so a future edit cannot quietly restore the set-equality claim.\r\n    // `adapter_dir` is in BOTH; `perch_dir` is spawn-only by design (a waker has\r\n    // no live link, so there is no perch-relative transfer path to resolve).\r\n    #[test]\r\n    fn the_wake_catalog_shares_the_spawn_vocabulary_but_is_not_the_same_set() {\r\n        let install = std::path::PathBuf::from(\"/adapters/mock\");\r\n        assert!(\r\n            fill_wake_command(\"sh-1\", \"mock-shell\", Some(&install), \"w {adapter_dir}\").is_ok(),\r\n            \"adapter_dir is shared by both catalogs\"\r\n        );\r\n        assert!(\r\n            fill_wake_command(\"sh-1\", \"mock-shell\", Some(&install), \"w {perch_dir}\").is_err(),\r\n            \"perch_dir is SPAWN-ONLY — a waker has no live link to resolve paths against\"\r\n        );\r\n    }\r\n\r\n    fn seed_owner(owlery: &Path, owner: &str) {\r\n        let p = owlery.join(owner);\r\n        std::fs::create_dir_all(&p).unwrap();\r\n        std::fs::write(p.join(\"info.json\"), \"{}\").unwrap();\r\n    }\r\n\r\n    /// Tokens for a waker that exits with `code` immediately.\r\n    fn exit_with(code: i32) -> Vec<String> {\r\n        #[cfg(windows)]\r\n        return vec![\"cmd\".into(), \"/c\".into(), format!(\"exit {code}\")];\r\n        #[cfg(unix)]\r\n        return vec![\"sh\".into(), \"-c\".into(), format!(\"exit {code}\")];\r\n    }\r\n\r\n    /// Tokens for a waker that appends one line to `tally` then exits 1.\r\n    fn crash_tallying(tally: &Path) -> Vec<String> {\r\n        #[cfg(windows)]\r\n        return vec![\r\n            \"powershell\".into(),\r\n            \"-NoProfile\".into(),\r\n            \"-Command\".into(),\r\n            format!(\"Add-Content -Path '{}' -Value x; exit 1\", tally.display()),\r\n        ];\r\n        #[cfg(unix)]\r\n        return vec![\r\n            \"sh\".into(),\r\n            \"-c\".into(),\r\n            format!(\"echo x >> '{}'; exit 1\", tally.display()),\r\n        ];\r\n    }\r\n\r\n    // [unit->REQ-SHELL-2] the backoff curve: base × 2^(n-1), saturating at\r\n    // the cap; zero failures wait nothing.\r\n    #[test]\r\n    fn backoff_curve_doubles_to_the_cap() {\r\n        let p = WakeParams {\r\n            backoff_base_ms: 100,\r\n            backoff_cap_ms: 1_500,\r\n            give_up_after: 6,\r\n        };\r\n        assert_eq!(backoff_ms(&p, 0), 0);\r\n        assert_eq!(backoff_ms(&p, 1), 100);\r\n        assert_eq!(backoff_ms(&p, 2), 200);\r\n        assert_eq!(backoff_ms(&p, 4), 800);\r\n        assert_eq!(backoff_ms(&p, 5), 1_500, \"saturates at the cap\");\r\n        assert_eq!(\r\n            backoff_ms(&p, 60),\r\n            1_500,\r\n            \"huge failure counts can't overflow\"\r\n        );\r\n    }\r\n\r\n    // [unit->REQ-SHELL-2] exit-opcode supervision: exit(86) fires the wake\r\n    // resolution exactly once and the loop ends; the pid file is retired.\r\n    #[test]\r\n    fn watcher_opcode_exit_fires_resolution_once() {\r\n        let tmp = tempfile::tempdir().unwrap();\r\n        let owlery = tmp.path();\r\n        seed_owner(owlery, \"doyle\");\r\n        let id = spawn_record(owlery, \"doyle\", \"mock-shell\", None).unwrap();\r\n        let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, \"doyle\", &id);\r\n\r\n        let fired = std::sync::atomic::AtomicU32::new(0);\r\n        let stop = AtomicBool::new(false);\r\n        watcher_run(\r\n            owlery,\r\n            \"doyle\",\r\n            &id,\r\n            &exit_with(WAKE_OPCODE),\r\n            &fast_params(),\r\n            &stop,\r\n            || {\r\n                fired.fetch_add(1, Ordering::SeqCst);\r\n                Ok(\"test-resolved\".into())\r\n            },\r\n        );\r\n        assert_eq!(\r\n            fired.load(Ordering::SeqCst),\r\n            1,\r\n            \"resolution fired exactly once\"\r\n        );\r\n        assert!(!perch.join(WAKER_PID_FILE).exists(), \"pid file retired\");\r\n        assert!(\r\n            !perch.join(WAKER_GAVE_UP_FILE).exists(),\r\n            \"an opcode exit is not a crash\"\r\n        );\r\n    }\r\n\r\n    // [unit->REQ-SHELL-2] crash-exit supervision: a crashing waker respawns\r\n    // (with backoff) exactly give_up_after times, then the durable give-up\r\n    // latch drops and the resolution never fires; clear_gave_up re-arms.\r\n    #[test]\r\n    fn watcher_crash_exits_respawn_then_give_up() {\r\n        let tmp = tempfile::tempdir().unwrap();\r\n        let owlery = tmp.path();\r\n        seed_owner(owlery, \"doyle\");\r\n        let id = spawn_record(owlery, \"doyle\", \"mock-shell\", None).unwrap();\r\n        let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, \"doyle\", &id);\r\n        let tally = perch.join(\"waker-tally.txt\");\r\n\r\n        let stop = AtomicBool::new(false);\r\n        watcher_run(\r\n            owlery,\r\n            \"doyle\",\r\n            &id,\r\n            &crash_tallying(&tally),\r\n…\n        );\n…\n    }","startLine":1168,"lineNumbers":[1168,null,1189,1190,1191,1192,1193,1194,1195,1196,1197,1198,1199,1200,1201,1202,1203,1204,1205,1206,1207,1208,1209,1210,1211,1212,1213,1214,1215,1216,1217,1218,1219,1220,1221,1222,1223,1224,1225,1226,1227,1228,1229,1230,1231,1232,1233,1234,1235,1236,1237,1238,1239,1240,1241,1242,1243,1244,1245,1246,1247,1248,1249,1250,1251,1252,1253,1254,1255,1256,1257,1258,1259,1260,1261,1262,1263,1264,1265,1266,1267,1268,1269,1270,1271,1272,1273,1274,1275,1276,1277,1278,1279,1280,1281,1282,1283,1284,1285,1286,1287,1288,1289,1290,1291,1292,1293,1294,1295,1296,1297,1298,1299,1300,1301,1302,1303,1304,1305,1306,1307,1308,1309,1310,1311,1312,1313,1314,1315,1316,1317,1318,1319,1320,1321,1322,1323,1324,1325,1326,1327,1328,1329,1330,1331,1332,1333,1334,1335,1336,1337,1338,1339,1340,1341,1342,1343,1344,1345,1346,1347,1348,1349,1350,1351,1352,1353,1354,1355,1356,1357,1358,1359,1360,1361,1362,1363,1364,1365,1366,1367,1368,1369,1370,1371,1372,1373,1374,1375,1376,1377,1378,1379,1380,1381,1382,1383,1384,1385,1386,1387,1388,1389,1390,1391,1392,1393,1394,1395,1396,1397,1398,1399,1400,1401,1402,1403,1404,1405,1406,1407,1408,1409,1410,1411,1412,1413,1414,1415,1416,1417,1418,1419,1420,1421,1422,1423,1424,1425,1426,1427,1428,1429,1430,1431,1432,1433,1434,1435,1436,1437,1438,1439,1440,1441,1442,1443,1444,1445,1446,1447,1448,1449,1450,1451,1452,1453,1454,1455,1456,1457,1458,1459,1460,1461,1462,1463,1464,1465,1466,1467,1468,1469,1470,1471,1472,1473,null,1477,null,1490]},"fileSize":100821,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\crates\\spt-daemon\\src\\shellwake.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":2161,"totalBytes":12562,"outputLines":285,"outputBytes":12562,"shownRange":{"start":1189,"end":1473},"nextOffset":1474}}},"isError":false,"timestamp":1789001187746}}
{"type":"message","id":"fea5a0b0","parentId":"000cd841","timestamp":"2026-09-10T00:46:27.748Z","message":{"role":"toolResult","toolCallId":"call_4XqdQMELsGY7pnb5eiuhkd6i|fc_077f3569365e313d016aa1fde3a33c87d1a3f417003c017244","toolName":"read","content":[{"type":"text","text":"/// stdin and `spt shell cmd` can drive it through the `send_effect` input\n/// seam (exactly-once per spool row). The broker outlives this connection —\n/// the session keeps running after the spawning call returns.\n// [impl->REQ-SHELL-1]\npub fn launch_shell_brokered_in(\n    broker_name: &str,\n    owlery: &Path,\n    owner: &str,\n    shell_id: &str,\n    adapter_name: &str,\n    install_dir: Option<&Path>,\n    shell: &Shell,\n) -> Result<u32, String> {\n    let tokens = fill_spawn_command(owlery, owner, shell_id, adapter_name, install_dir, shell)?;\n    let (program, args) = tokens.split_first().expect(\"non-empty checked\");\n    let mut brain =\n        Brain::cold_start(broker_name, now_ms()).map_err(|e| format!(\"broker connect: {e}\"))?;\n    let (_sid, pid) = brain\n        .spawn_session_pid(SpawnReq {\n            program: program.clone(),\n            args: args.to_vec(),\n            // A shell binary is not a terminal: the PTY is only the stdin\n            // conduit. Keep it wide so renderer-side wrapping (ConPTY) can\n            // never split a >80-col stamped frame mid-line.\n            rows: 24,\n            cols: 512,\n            endpoint: shell_session_label(owner, shell_id),\n            // Shells inherit the broker's cwd (unchanged); project-scoped cwd is\n            // a harness-bringup concern (REQ-HOST-RUN-2), not shell spawn.\n            cwd: None,\n            // Shells inherit the broker's env; [env] inject is a harness-bringup\n            // concern (REQ-HAZARD-ENV-SUBST), not shell spawn.\n            env: Default::default(),\n            translation_binary: None,\n            adapter: String::new(),\n            // Deliberately None even though a shell now HAS an install dir: the\n            // broker's copy exists to resolve a translation binary (there is\n            // none here) and would otherwise re-resolve a program token\n            // `fill_spawn_command` has already resolved — a second resolution\n            // path is exactly what REQ-INSTALL-11 has one primitive to avoid.\n            install_dir: None,\n        })\n        .map_err(|e| format!(\"broker spawn: {e}\"))?;\n    let pid = pid.unwrap_or(0);\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n    std::fs::write(perch.join(SHELL_PID_FILE), pid.to_string())\n        .map_err(|e| format!(\"record pid: {e}\"))?;\n    // Same stamp at the broker-hosted launch. A backend that exposed no pid\n    // records 0 here; record_shell_launch stamps no start-time for it, so the\n    // probe keeps its fail-toward-alive stance for pid-less backends.\n    // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n    spt_store::shellinfo::record_shell_launch(&perch, pid, now_ms());\n    Ok(pid)\n}\n\n/// Resolve a shell instance by its link token alone — the \"owner from the\n/// link\" lookup every shell-flavored machinery call shares. Scans every\n/// owner's `shells/` namespace for the perch parking exactly this token; an\n/// unknown/empty token resolves nothing. Pure read — onlining is the bind's\n/// job ([`bind_shell_by_token`]), not the lookup's.\n///\n/// Owner-type-agnostic: `owner` is an opaque endpoint-id namespace string here\n/// and through every shell-machinery call — no path inspects the owner's\n/// endpoint type, so a Gateway (or any non-Shell type) owns identically to an\n/// agent (CONTEXT §Shell model \"Owner-linked, exclusive\"; M11-T0 audit).\n// [impl->REQ-SHELL-1]\n// [impl->REQ-SHELL-5]\npub fn find_shell_by_token(owlery: &Path, token: &str) -> Option<(String, String)> {\n    if token.is_empty() {\n        return None;\n    }\n    for owner in spt_store::perch::list_self_perch_ids(owlery) {\n        for (shell_id, _info) in shellinfo::list_shells(owlery, &owner) {\n            let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, &owner, &shell_id);\n            if read_link_token(&perch).as_deref() == Some(token) {\n                return Some((owner, shell_id));\n            }\n        }\n    }\n    None\n}\n\n/// Resolve a bind by token alone and flip the instance online — the\n/// handshake's receiving end. The spawn template carries only\n/// `{link_token}`, so the binary never has to know its owner (\"owner from\n/// the link\"). An unknown/empty token binds nothing. Returns\n/// `(owner, shell_id)`.\n// [impl->REQ-SHELL-1]\npub fn bind_shell_by_token(owlery: &Path, token: &str) -> Option<(String, String)> {\n    let (owner, shell_id) = find_shell_by_token(owlery, token)?;\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, &owner, &shell_id);\n    let mut info = shellinfo::read_shell_info(&perch)?;\n    info.status = SHELL_STATUS_ONLINE.to_string();\n    shellinfo::write_shell_info(&perch, &info).ok()?;\n    Some((owner, shell_id))\n}\n\n/// Verify a machinery call's link credential against one instance — the\n/// per-message auth gate the D3c channel surfaces (`api emit`, `api poll`)\n/// sit behind.\n// [impl->REQ-SHELL-1]\npub fn verify_link(owlery: &Path, owner: &str, shell_id: &str, token: &str) -> bool {\n    if token.is_empty() {\n        return false;\n    }\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n    read_link_token(&perch).as_deref() == Some(token)\n}\n\n/// Best-effort kill of a launched shell binary (teardown's interim close —\n/// the D4 lifecycle adds the manifest pre-close instruction + termination\n/// timeout in front of this). Bounded; never blocks on a wedged child.\npub fn kill_shell_pid(pid: u32) {\n    if pid == 0 {\n        // A broker-hosted spawn whose backend exposed no pid records 0 —\n        // never kill that (unix `kill 0` would signal the whole group).\n        return;\n    }\n    #[cfg(windows)]\n    {\n        // taskkill is a console child of the console-less daemon: spawn it\n        // windowless or teardown flashes a blank conhost (KNOWN-HAZARDS 5.8).\n        // [impl->REQ-HAZARD-CHILD-CONSOLE-FLASH]\n        use std::os::windows::process::CommandExt;\n        let _ = Command::new(\"taskkill\")\n            .args([\"/PID\", &pid.to_string(), \"/T\", \"/F\"])\n            .stdout(Stdio::null())\n            .stderr(Stdio::null())\n            .creation_flags(0x0800_0000) // CREATE_NO_WINDOW\n            .status();\n    }\n    #[cfg(unix)]\n    let _ = Command::new(\"kill\")\n        .args([\"-9\", &pid.to_string()])\n        .stdout(Stdio::null())\n        .stderr(Stdio::null())\n        .status();\n}\n\n/// What a chokepoint may do with a recorded pid, once the pair test has spoken.\n///\n/// One policy for BOTH chokepoints ([`kill_shell_at`] and\n/// [`crate::shellwake::kill_waker_at`]) so the destructive decision is made in\n/// exactly one place. This is not a second liveness predicate — it CONSUMES\n/// [`spt_store::liveness::relay_liveness`]'s verdict and says what a kill path\n/// is allowed to do with it.\n#[derive(Debug, Clone, Copy, PartialEq, Eq)]\npub(crate) enum KillVerdict {\n    /// The process at that pid is still ours (or a pre-stamp record, where\n    /// existence is all the record can offer): kill it, retire the record.\n    Fire,\n    /// PROVABLY not ours — absent, or the number recycled onto an unrelated\n    /// process. Retire the record, kill NOTHING.\n    RetireOnly,\n    /// Identity unknown. A destructive path treats unknown as do-not-fire, and\n    /// KEEPS the record: retiring it here would orphan a possibly-live process\n    /// while presenting success.\n    RefuseAndKeep,\n}\n\n/// The kill policy: unknown never fires, and only a proven corpse is forgotten.\n// [impl->REQ-SHELL-KILL-AUTHENTICATED]\npub(crate) fn kill_verdict(liveness: RelayLiveness) -> KillVerdict {\n    match liveness {\n        RelayLiveness::Held => KillVerdict::Fire,\n        RelayLiveness::Gone => KillVerdict::RetireOnly,\n        RelayLiveness::Unproven => KillVerdict::RefuseAndKeep,\n    }\n}\n\n/// Teardown's process half: read the recorded pid and kill it — but only once\n/// the pid+birth PAIR proves the process wearing that number is still the\n/// binary we launched. The caller removes the perch dir after.\n///\n/// The pair test is [`spt_store::liveness::relay_liveness`], the same predicate\n/// [`shellinfo::shell_pid_provably_dead`] runs on the derivation side. W2\n/// authenticated that side and left this one believing a bare pid, so the very\n/// state the derived read calls a corpse still routed a recycled number to\n/// `taskkill /T` — a stranger's whole subtree (KNOWN-HAZARDS 7.58).\n///\n/// A pid with no stamp beside it (a pre-W2 record) still reads `Held` on\n/// existence alone and is still killed: the migration window is deliberate, and\n/// the next launch writes a pair.\n///\n/// **This NARROWS the mis-fire; it does not eliminate it — on Linux.**\n/// [`spt_store::proc::process_started_at`] reads `/proc/<pid>/stat` field 22,\n/// which counts `CLK_TCK = 100` jiffies, so two processes whose starts fall in\n/// the SAME 10ms tick carry an identical stamp and the pair cannot separate\n/// them. The window closes from *any* recycled pid to *a pid recycled into the\n/// same 10ms tick as the recorded start* — which requires our process to die\n/// AND the OS to hand its number to a new process inside one tick. Windows is\n/// unaffected: its stamp is a `FILETIME`, of far finer resolution. Measured by\n/// hertz on kitsubito (BAROMETER); stated here so nobody re-derives it.\n// [impl->REQ-SHELL-KILL-AUTHENTICATED]\npub fn kill_shell_at(shell_perch: &Path) {\n    let Some(pid) = shellinfo::read_shell_pid(shell_perch).filter(|&p| p != 0) else {\n        return;\n    };\n    let birth = shellinfo::read_shell_launch(shell_perch).and_then(|l| l.pid_started_at);\n    match kill_verdict(spt_store::liveness::relay_liveness(Some(pid), birth)) {\n        // Still ours (or a pre-stamp record, which existence is all it can offer).\n        KillVerdict::Fire => kill_shell_pid(pid),\n        // PROVABLY not ours: absent, or that number now worn by an unrelated\n        // process. Retire the record and kill NOTHING.\n        KillVerdict::RetireOnly => {\n            spt_proto::emit_line_err!(\n                \"SHELL_KILL_REFUSED_GONE:{}: recorded pid={pid} is not the binary we launched \\\n                 (absent, or the pid was recycled onto another process) — retiring the record \\\n                 and killing nothing\",\n                shell_perch.display()\n            );\n…\n        }\n…\n    }\n}\n\n[Showing lines 513-723 of 1826. Use :724 to continue]"}],"details":{"truncation":{"content":"/// stdin and `spt shell cmd` can drive it through the `send_effect` input\r\n/// seam (exactly-once per spool row). The broker outlives this connection —\r\n/// the session keeps running after the spawning call returns.\r\n// [impl->REQ-SHELL-1]\r\npub fn launch_shell_brokered_in(\r\n    broker_name: &str,\r\n    owlery: &Path,\r\n    owner: &str,\r\n    shell_id: &str,\r\n    adapter_name: &str,\r\n    install_dir: Option<&Path>,\r\n    shell: &Shell,\r\n) -> Result<u32, String> {\r\n    let tokens = fill_spawn_command(owlery, owner, shell_id, adapter_name, install_dir, shell)?;\r\n    let (program, args) = tokens.split_first().expect(\"non-empty checked\");\r\n    let mut brain =\r\n        Brain::cold_start(broker_name, now_ms()).map_err(|e| format!(\"broker connect: {e}\"))?;\r\n    let (_sid, pid) = brain\r\n        .spawn_session_pid(SpawnReq {\r\n            program: program.clone(),\r\n            args: args.to_vec(),\r\n            // A shell binary is not a terminal: the PTY is only the stdin\r\n            // conduit. Keep it wide so renderer-side wrapping (ConPTY) can\r\n            // never split a >80-col stamped frame mid-line.\r\n            rows: 24,\r\n            cols: 512,\r\n            endpoint: shell_session_label(owner, shell_id),\r\n            // Shells inherit the broker's cwd (unchanged); project-scoped cwd is\r\n            // a harness-bringup concern (REQ-HOST-RUN-2), not shell spawn.\r\n            cwd: None,\r\n            // Shells inherit the broker's env; [env] inject is a harness-bringup\r\n            // concern (REQ-HAZARD-ENV-SUBST), not shell spawn.\r\n            env: Default::default(),\r\n            translation_binary: None,\r\n            adapter: String::new(),\r\n            // Deliberately None even though a shell now HAS an install dir: the\r\n            // broker's copy exists to resolve a translation binary (there is\r\n            // none here) and would otherwise re-resolve a program token\r\n            // `fill_spawn_command` has already resolved — a second resolution\r\n            // path is exactly what REQ-INSTALL-11 has one primitive to avoid.\r\n            install_dir: None,\r\n        })\r\n        .map_err(|e| format!(\"broker spawn: {e}\"))?;\r\n    let pid = pid.unwrap_or(0);\r\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\r\n    std::fs::write(perch.join(SHELL_PID_FILE), pid.to_string())\r\n        .map_err(|e| format!(\"record pid: {e}\"))?;\r\n    // Same stamp at the broker-hosted launch. A backend that exposed no pid\r\n    // records 0 here; record_shell_launch stamps no start-time for it, so the\r\n    // probe keeps its fail-toward-alive stance for pid-less backends.\r\n    // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\n    spt_store::shellinfo::record_shell_launch(&perch, pid, now_ms());\r\n    Ok(pid)\r\n}\r\n\r\n/// Resolve a shell instance by its link token alone — the \"owner from the\r\n/// link\" lookup every shell-flavored machinery call shares. Scans every\r\n/// owner's `shells/` namespace for the perch parking exactly this token; an\r\n/// unknown/empty token resolves nothing. Pure read — onlining is the bind's\r\n/// job ([`bind_shell_by_token`]), not the lookup's.\r\n///\r\n/// Owner-type-agnostic: `owner` is an opaque endpoint-id namespace string here\r\n/// and through every shell-machinery call — no path inspects the owner's\r\n/// endpoint type, so a Gateway (or any non-Shell type) owns identically to an\r\n/// agent (CONTEXT §Shell model \"Owner-linked, exclusive\"; M11-T0 audit).\r\n// [impl->REQ-SHELL-1]\r\n// [impl->REQ-SHELL-5]\r\npub fn find_shell_by_token(owlery: &Path, token: &str) -> Option<(String, String)> {\r\n    if token.is_empty() {\r\n        return None;\r\n    }\r\n    for owner in spt_store::perch::list_self_perch_ids(owlery) {\r\n        for (shell_id, _info) in shellinfo::list_shells(owlery, &owner) {\r\n            let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, &owner, &shell_id);\r\n            if read_link_token(&perch).as_deref() == Some(token) {\r\n                return Some((owner, shell_id));\r\n            }\r\n        }\r\n    }\r\n    None\r\n}\r\n\r\n/// Resolve a bind by token alone and flip the instance online — the\r\n/// handshake's receiving end. The spawn template carries only\r\n/// `{link_token}`, so the binary never has to know its owner (\"owner from\r\n/// the link\"). An unknown/empty token binds nothing. Returns\r\n/// `(owner, shell_id)`.\r\n// [impl->REQ-SHELL-1]\r\npub fn bind_shell_by_token(owlery: &Path, token: &str) -> Option<(String, String)> {\r\n    let (owner, shell_id) = find_shell_by_token(owlery, token)?;\r\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, &owner, &shell_id);\r\n    let mut info = shellinfo::read_shell_info(&perch)?;\r\n    info.status = SHELL_STATUS_ONLINE.to_string();\r\n    shellinfo::write_shell_info(&perch, &info).ok()?;\r\n    Some((owner, shell_id))\r\n}\r\n\r\n/// Verify a machinery call's link credential against one instance — the\r\n/// per-message auth gate the D3c channel surfaces (`api emit`, `api poll`)\r\n/// sit behind.\r\n// [impl->REQ-SHELL-1]\r\npub fn verify_link(owlery: &Path, owner: &str, shell_id: &str, token: &str) -> bool {\r\n    if token.is_empty() {\r\n        return false;\r\n    }\r\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\r\n    read_link_token(&perch).as_deref() == Some(token)\r\n}\r\n\r\n/// Best-effort kill of a launched shell binary (teardown's interim close —\r\n/// the D4 lifecycle adds the manifest pre-close instruction + termination\r\n/// timeout in front of this). Bounded; never blocks on a wedged child.\r\npub fn kill_shell_pid(pid: u32) {\r\n    if pid == 0 {\r\n        // A broker-hosted spawn whose backend exposed no pid records 0 —\r\n        // never kill that (unix `kill 0` would signal the whole group).\r\n        return;\r\n    }\r\n    #[cfg(windows)]\r\n    {\r\n        // taskkill is a console child of the console-less daemon: spawn it\r\n        // windowless or teardown flashes a blank conhost (KNOWN-HAZARDS 5.8).\r\n        // [impl->REQ-HAZARD-CHILD-CONSOLE-FLASH]\r\n        use std::os::windows::process::CommandExt;\r\n        let _ = Command::new(\"taskkill\")\r\n            .args([\"/PID\", &pid.to_string(), \"/T\", \"/F\"])\r\n            .stdout(Stdio::null())\r\n            .stderr(Stdio::null())\r\n            .creation_flags(0x0800_0000) // CREATE_NO_WINDOW\r\n            .status();\r\n    }\r\n    #[cfg(unix)]\r\n    let _ = Command::new(\"kill\")\r\n        .args([\"-9\", &pid.to_string()])\r\n        .stdout(Stdio::null())\r\n        .stderr(Stdio::null())\r\n        .status();\r\n}\r\n\r\n/// What a chokepoint may do with a recorded pid, once the pair test has spoken.\r\n///\r\n/// One policy for BOTH chokepoints ([`kill_shell_at`] and\r\n/// [`crate::shellwake::kill_waker_at`]) so the destructive decision is made in\r\n/// exactly one place. This is not a second liveness predicate — it CONSUMES\r\n/// [`spt_store::liveness::relay_liveness`]'s verdict and says what a kill path\r\n/// is allowed to do with it.\r\n#[derive(Debug, Clone, Copy, PartialEq, Eq)]\r\npub(crate) enum KillVerdict {\r\n    /// The process at that pid is still ours (or a pre-stamp record, where\r\n    /// existence is all the record can offer): kill it, retire the record.\r\n    Fire,\r\n    /// PROVABLY not ours — absent, or the number recycled onto an unrelated\r\n    /// process. Retire the record, kill NOTHING.\r\n    RetireOnly,\r\n    /// Identity unknown. A destructive path treats unknown as do-not-fire, and\r\n    /// KEEPS the record: retiring it here would orphan a possibly-live process\r\n    /// while presenting success.\r\n    RefuseAndKeep,\r\n}\r\n\r\n/// The kill policy: unknown never fires, and only a proven corpse is forgotten.\r\n// [impl->REQ-SHELL-KILL-AUTHENTICATED]\r\npub(crate) fn kill_verdict(liveness: RelayLiveness) -> KillVerdict {\r\n    match liveness {\r\n        RelayLiveness::Held => KillVerdict::Fire,\r\n        RelayLiveness::Gone => KillVerdict::RetireOnly,\r\n        RelayLiveness::Unproven => KillVerdict::RefuseAndKeep,\r\n    }\r\n}\r\n\r\n/// Teardown's process half: read the recorded pid and kill it — but only once\r\n/// the pid+birth PAIR proves the process wearing that number is still the\r\n/// binary we launched. The caller removes the perch dir after.\r\n///\r\n/// The pair test is [`spt_store::liveness::relay_liveness`], the same predicate\r\n/// [`shellinfo::shell_pid_provably_dead`] runs on the derivation side. W2\r\n/// authenticated that side and left this one believing a bare pid, so the very\r\n/// state the derived read calls a corpse still routed a recycled number to\r\n/// `taskkill /T` — a stranger's whole subtree (KNOWN-HAZARDS 7.58).\r\n///\r\n/// A pid with no stamp beside it (a pre-W2 record) still reads `Held` on\r\n/// existence alone and is still killed: the migration window is deliberate, and\r\n/// the next launch writes a pair.\r\n///\r\n/// **This NARROWS the mis-fire; it does not eliminate it — on Linux.**\r\n/// [`spt_store::proc::process_started_at`] reads `/proc/<pid>/stat` field 22,\r\n/// which counts `CLK_TCK = 100` jiffies, so two processes whose starts fall in\r\n/// the SAME 10ms tick carry an identical stamp and the pair cannot separate\r\n/// them. The window closes from *any* recycled pid to *a pid recycled into the\r\n/// same 10ms tick as the recorded start* — which requires our process to die\r\n/// AND the OS to hand its number to a new process inside one tick. Windows is\r\n/// unaffected: its stamp is a `FILETIME`, of far finer resolution. Measured by\r\n/// hertz on kitsubito (BAROMETER); stated here so nobody re-derives it.\r\n// [impl->REQ-SHELL-KILL-AUTHENTICATED]\r\npub fn kill_shell_at(shell_perch: &Path) {\r\n    let Some(pid) = shellinfo::read_shell_pid(shell_perch).filter(|&p| p != 0) else {\r\n        return;\r\n    };\r\n    let birth = shellinfo::read_shell_launch(shell_perch).and_then(|l| l.pid_started_at);\r\n    match kill_verdict(spt_store::liveness::relay_liveness(Some(pid), birth)) {\r\n        // Still ours (or a pre-stamp record, which existence is all it can offer).\r\n        KillVerdict::Fire => kill_shell_pid(pid),\r\n        // PROVABLY not ours: absent, or that number now worn by an unrelated\r\n        // process. Retire the record and kill NOTHING.\r\n        KillVerdict::RetireOnly => {\r\n            spt_proto::emit_line_err!(\r\n                \"SHELL_KILL_REFUSED_GONE:{}: recorded pid={pid} is not the binary we launched \\\r\n                 (absent, or the pid was recycled onto another process) — retiring the record \\\r\n                 and killing nothing\",\r\n                shell_perch.display()\r\n            );\r","truncated":true,"truncatedBy":"lines","totalLines":1314,"totalBytes":10466,"outputLines":211,"outputBytes":10466,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":1826,"displayContent":{"text":"/// stdin and `spt shell cmd` can drive it through the `send_effect` input\r\n/// seam (exactly-once per spool row). The broker outlives this connection —\r\n/// the session keeps running after the spawning call returns.\r\n// [impl->REQ-SHELL-1]\r\npub fn launch_shell_brokered_in(\r\n    broker_name: &str,\r\n    owlery: &Path,\r\n    owner: &str,\r\n    shell_id: &str,\r\n    adapter_name: &str,\r\n    install_dir: Option<&Path>,\r\n    shell: &Shell,\r\n) -> Result<u32, String> {\r\n    let tokens = fill_spawn_command(owlery, owner, shell_id, adapter_name, install_dir, shell)?;\r\n    let (program, args) = tokens.split_first().expect(\"non-empty checked\");\r\n    let mut brain =\r\n        Brain::cold_start(broker_name, now_ms()).map_err(|e| format!(\"broker connect: {e}\"))?;\r\n    let (_sid, pid) = brain\r\n        .spawn_session_pid(SpawnReq {\r\n            program: program.clone(),\r\n            args: args.to_vec(),\r\n            // A shell binary is not a terminal: the PTY is only the stdin\r\n            // conduit. Keep it wide so renderer-side wrapping (ConPTY) can\r\n            // never split a >80-col stamped frame mid-line.\r\n            rows: 24,\r\n            cols: 512,\r\n            endpoint: shell_session_label(owner, shell_id),\r\n            // Shells inherit the broker's cwd (unchanged); project-scoped cwd is\r\n            // a harness-bringup concern (REQ-HOST-RUN-2), not shell spawn.\r\n            cwd: None,\r\n            // Shells inherit the broker's env; [env] inject is a harness-bringup\r\n            // concern (REQ-HAZARD-ENV-SUBST), not shell spawn.\r\n            env: Default::default(),\r\n            translation_binary: None,\r\n            adapter: String::new(),\r\n            // Deliberately None even though a shell now HAS an install dir: the\r\n            // broker's copy exists to resolve a translation binary (there is\r\n            // none here) and would otherwise re-resolve a program token\r\n            // `fill_spawn_command` has already resolved — a second resolution\r\n            // path is exactly what REQ-INSTALL-11 has one primitive to avoid.\r\n            install_dir: None,\r\n        })\r\n        .map_err(|e| format!(\"broker spawn: {e}\"))?;\r\n    let pid = pid.unwrap_or(0);\r\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\r\n    std::fs::write(perch.join(SHELL_PID_FILE), pid.to_string())\r\n        .map_err(|e| format!(\"record pid: {e}\"))?;\r\n    // Same stamp at the broker-hosted launch. A backend that exposed no pid\r\n    // records 0 here; record_shell_launch stamps no start-time for it, so the\r\n    // probe keeps its fail-toward-alive stance for pid-less backends.\r\n    // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\n    spt_store::shellinfo::record_shell_launch(&perch, pid, now_ms());\r\n    Ok(pid)\r\n}\r\n\r\n/// Resolve a shell instance by its link token alone — the \"owner from the\r\n/// link\" lookup every shell-flavored machinery call shares. Scans every\r\n/// owner's `shells/` namespace for the perch parking exactly this token; an\r\n/// unknown/empty token resolves nothing. Pure read — onlining is the bind's\r\n/// job ([`bind_shell_by_token`]), not the lookup's.\r\n///\r\n/// Owner-type-agnostic: `owner` is an opaque endpoint-id namespace string here\r\n/// and through every shell-machinery call — no path inspects the owner's\r\n/// endpoint type, so a Gateway (or any non-Shell type) owns identically to an\r\n/// agent (CONTEXT §Shell model \"Owner-linked, exclusive\"; M11-T0 audit).\r\n// [impl->REQ-SHELL-1]\r\n// [impl->REQ-SHELL-5]\r\npub fn find_shell_by_token(owlery: &Path, token: &str) -> Option<(String, String)> {\r\n    if token.is_empty() {\r\n        return None;\r\n    }\r\n    for owner in spt_store::perch::list_self_perch_ids(owlery) {\r\n        for (shell_id, _info) in shellinfo::list_shells(owlery, &owner) {\r\n            let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, &owner, &shell_id);\r\n            if read_link_token(&perch).as_deref() == Some(token) {\r\n                return Some((owner, shell_id));\r\n            }\r\n        }\r\n    }\r\n    None\r\n}\r\n\r\n/// Resolve a bind by token alone and flip the instance online — the\r\n/// handshake's receiving end. The spawn template carries only\r\n/// `{link_token}`, so the binary never has to know its owner (\"owner from\r\n/// the link\"). An unknown/empty token binds nothing. Returns\r\n/// `(owner, shell_id)`.\r\n// [impl->REQ-SHELL-1]\r\npub fn bind_shell_by_token(owlery: &Path, token: &str) -> Option<(String, String)> {\r\n    let (owner, shell_id) = find_shell_by_token(owlery, token)?;\r\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, &owner, &shell_id);\r\n    let mut info = shellinfo::read_shell_info(&perch)?;\r\n    info.status = SHELL_STATUS_ONLINE.to_string();\r\n    shellinfo::write_shell_info(&perch, &info).ok()?;\r\n    Some((owner, shell_id))\r\n}\r\n\r\n/// Verify a machinery call's link credential against one instance — the\r\n/// per-message auth gate the D3c channel surfaces (`api emit`, `api poll`)\r\n/// sit behind.\r\n// [impl->REQ-SHELL-1]\r\npub fn verify_link(owlery: &Path, owner: &str, shell_id: &str, token: &str) -> bool {\r\n    if token.is_empty() {\r\n        return false;\r\n    }\r\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\r\n    read_link_token(&perch).as_deref() == Some(token)\r\n}\r\n\r\n/// Best-effort kill of a launched shell binary (teardown's interim close —\r\n/// the D4 lifecycle adds the manifest pre-close instruction + termination\r\n/// timeout in front of this). Bounded; never blocks on a wedged child.\r\npub fn kill_shell_pid(pid: u32) {\r\n    if pid == 0 {\r\n        // A broker-hosted spawn whose backend exposed no pid records 0 —\r\n        // never kill that (unix `kill 0` would signal the whole group).\r\n        return;\r\n    }\r\n    #[cfg(windows)]\r\n    {\r\n        // taskkill is a console child of the console-less daemon: spawn it\r\n        // windowless or teardown flashes a blank conhost (KNOWN-HAZARDS 5.8).\r\n        // [impl->REQ-HAZARD-CHILD-CONSOLE-FLASH]\r\n        use std::os::windows::process::CommandExt;\r\n        let _ = Command::new(\"taskkill\")\r\n            .args([\"/PID\", &pid.to_string(), \"/T\", \"/F\"])\r\n            .stdout(Stdio::null())\r\n            .stderr(Stdio::null())\r\n            .creation_flags(0x0800_0000) // CREATE_NO_WINDOW\r\n            .status();\r\n    }\r\n    #[cfg(unix)]\r\n    let _ = Command::new(\"kill\")\r\n        .args([\"-9\", &pid.to_string()])\r\n        .stdout(Stdio::null())\r\n        .stderr(Stdio::null())\r\n        .status();\r\n}\r\n\r\n/// What a chokepoint may do with a recorded pid, once the pair test has spoken.\r\n///\r\n/// One policy for BOTH chokepoints ([`kill_shell_at`] and\r\n/// [`crate::shellwake::kill_waker_at`]) so the destructive decision is made in\r\n/// exactly one place. This is not a second liveness predicate — it CONSUMES\r\n/// [`spt_store::liveness::relay_liveness`]'s verdict and says what a kill path\r\n/// is allowed to do with it.\r\n#[derive(Debug, Clone, Copy, PartialEq, Eq)]\r\npub(crate) enum KillVerdict {\r\n    /// The process at that pid is still ours (or a pre-stamp record, where\r\n    /// existence is all the record can offer): kill it, retire the record.\r\n    Fire,\r\n    /// PROVABLY not ours — absent, or the number recycled onto an unrelated\r\n    /// process. Retire the record, kill NOTHING.\r\n    RetireOnly,\r\n    /// Identity unknown. A destructive path treats unknown as do-not-fire, and\r\n    /// KEEPS the record: retiring it here would orphan a possibly-live process\r\n    /// while presenting success.\r\n    RefuseAndKeep,\r\n}\r\n\r\n/// The kill policy: unknown never fires, and only a proven corpse is forgotten.\r\n// [impl->REQ-SHELL-KILL-AUTHENTICATED]\r\npub(crate) fn kill_verdict(liveness: RelayLiveness) -> KillVerdict {\r\n    match liveness {\r\n        RelayLiveness::Held => KillVerdict::Fire,\r\n        RelayLiveness::Gone => KillVerdict::RetireOnly,\r\n        RelayLiveness::Unproven => KillVerdict::RefuseAndKeep,\r\n    }\r\n}\r\n\r\n/// Teardown's process half: read the recorded pid and kill it — but only once\r\n/// the pid+birth PAIR proves the process wearing that number is still the\r\n/// binary we launched. The caller removes the perch dir after.\r\n///\r\n/// The pair test is [`spt_store::liveness::relay_liveness`], the same predicate\r\n/// [`shellinfo::shell_pid_provably_dead`] runs on the derivation side. W2\r\n/// authenticated that side and left this one believing a bare pid, so the very\r\n/// state the derived read calls a corpse still routed a recycled number to\r\n/// `taskkill /T` — a stranger's whole subtree (KNOWN-HAZARDS 7.58).\r\n///\r\n/// A pid with no stamp beside it (a pre-W2 record) still reads `Held` on\r\n/// existence alone and is still killed: the migration window is deliberate, and\r\n/// the next launch writes a pair.\r\n///\r\n/// **This NARROWS the mis-fire; it does not eliminate it — on Linux.**\r\n/// [`spt_store::proc::process_started_at`] reads `/proc/<pid>/stat` field 22,\r\n/// which counts `CLK_TCK = 100` jiffies, so two processes whose starts fall in\r\n/// the SAME 10ms tick carry an identical stamp and the pair cannot separate\r\n/// them. The window closes from *any* recycled pid to *a pid recycled into the\r\n/// same 10ms tick as the recorded start* — which requires our process to die\r\n/// AND the OS to hand its number to a new process inside one tick. Windows is\r\n/// unaffected: its stamp is a `FILETIME`, of far finer resolution. Measured by\r\n/// hertz on kitsubito (BAROMETER); stated here so nobody re-derives it.\r\n// [impl->REQ-SHELL-KILL-AUTHENTICATED]\r\npub fn kill_shell_at(shell_perch: &Path) {\r\n    let Some(pid) = shellinfo::read_shell_pid(shell_perch).filter(|&p| p != 0) else {\r\n        return;\r\n    };\r\n    let birth = shellinfo::read_shell_launch(shell_perch).and_then(|l| l.pid_started_at);\r\n    match kill_verdict(spt_store::liveness::relay_liveness(Some(pid), birth)) {\r\n        // Still ours (or a pre-stamp record, which existence is all it can offer).\r\n        KillVerdict::Fire => kill_shell_pid(pid),\r\n        // PROVABLY not ours: absent, or that number now worn by an unrelated\r\n        // process. Retire the record and kill NOTHING.\r\n        KillVerdict::RetireOnly => {\r\n            spt_proto::emit_line_err!(\r\n                \"SHELL_KILL_REFUSED_GONE:{}: recorded pid={pid} is not the binary we launched \\\r\n                 (absent, or the pid was recycled onto another process) — retiring the record \\\r\n                 and killing nothing\",\r\n                shell_perch.display()\r\n            );\r\n…\n        }\n…\n    }\n}","startLine":513,"lineNumbers":[513,514,515,516,517,518,519,520,521,522,523,524,525,526,527,528,529,530,531,532,533,534,535,536,537,538,539,540,541,542,543,544,545,546,547,548,549,550,551,552,553,554,555,556,557,558,559,560,561,562,563,564,565,566,567,568,569,570,571,572,573,574,575,576,577,578,579,580,581,582,583,584,585,586,587,588,589,590,591,592,593,594,595,596,597,598,599,600,601,602,603,604,605,606,607,608,609,610,611,612,613,614,615,616,617,618,619,620,621,622,623,624,625,626,627,628,629,630,631,632,633,634,635,636,637,638,639,640,641,642,643,644,645,646,647,648,649,650,651,652,653,654,655,656,657,658,659,660,661,662,663,664,665,666,667,668,669,670,671,672,673,674,675,676,677,678,679,680,681,682,683,684,685,686,687,688,689,690,691,692,693,694,695,696,697,698,699,700,701,702,703,704,705,706,707,708,709,710,711,712,713,714,715,716,717,718,719,720,721,722,723,null,725,null,738,739]},"fileSize":88363,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\crates\\spt-daemon\\src\\shellhost.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":1826,"totalBytes":10466,"outputLines":211,"outputBytes":10466,"shownRange":{"start":513,"end":723},"nextOffset":724}}},"isError":false,"timestamp":1789001187748}}
{"type":"message","id":"ee7991a4","parentId":"fea5a0b0","timestamp":"2026-09-10T00:46:27.767Z","message":{"role":"toolResult","toolCallId":"call_1KIq0qgFYanlAYMVZvaDV02J|fc_077f3569365e313d016aa1fde3a36487d1a8f349852acfa5b4","toolName":"grep","content":[{"type":"text","text":"# .worktrees/287-shellrestore/crates/\n\n## spt-daemon/src/\n### shellhost.rs\n 723|            );\n*724|            let _ = std::fs::remove_file(shell_perch.join(SHELL_PID_FILE));\n 725|        }\n 726|        // Unknown identity on a DESTRUCTIVE path fails toward not firing: a\n 727|        // mis-fire force-kills a stranger's tree, a missed kill leaves our own\n...\n 780|// [impl->REQ-SHELL-2]\n*781|pub fn close_shell(\n 782|    owlery: &Path,\n 783|    owner: &str,\n 784|    shell_id: &str,\n...\n 859|    }\n*860|    let _ = std::fs::remove_file(perch.join(SHELL_PID_FILE));\n 861|    // Retire the live credential — but STASH it (rename, atomic) rather than\n 862|    // erase it: pending spool rows are stamped under this token's key, and the\n 863|    // relink that recovers a persistent instance must re-stamp them to its\n...\n 1233|        std::fs::write(\n*1234|            perch.join(shellinfo::SHELL_LAUNCH_FILE),\n 1235|            serde_json::to_string(&launch).unwrap(),\n 1236|        )\n 1237|        .unwrap();\n...\n 1763|        std::fs::write(\n*1764|            perch.join(shellinfo::SHELL_LAUNCH_FILE),\n 1765|            r#\"{\"pid_started_at\":123,\"launched_ms\":1}\"#,\n 1766|        )\n 1767|        .unwrap();\n...\n 1815|        assert!(\n*1816|            !perch.join(shellinfo::SHELL_LAUNCH_FILE).exists(),\n 1817|            \"the pre-stamp shape is the ABSENCE of a launch record\"\n 1818|        );\n 1819|\n\n## spt-store/src/\n### shellinfo.rs\n 131|/// launch (BAROMETER W2, releases#78).\n*132|pub const SHELL_LAUNCH_FILE: &str = \"shell.launch.json\";\n 133|\n 134|/// What a launch recorded about the process it started.\n 135|///\n...\n 170|    if let Ok(json) = serde_json::to_string_pretty(&launch) {\n*171|        let _ = atomic_write_string(&shell_perch.join(SHELL_LAUNCH_FILE), &json);\n 172|    }\n 173|}\n 174|\n...\n 178|pub fn read_shell_launch(shell_perch: &Path) -> Option<ShellLaunch> {\n*179|    let raw = std::fs::read_to_string(shell_perch.join(SHELL_LAUNCH_FILE)).ok()?;\n 180|    serde_json::from_str(&raw).ok()\n 181|}\n 182|\n...\n 471|        std::fs::write(\n*472|            d.path().join(SHELL_LAUNCH_FILE),\n 473|            serde_json::to_string(&bogus).unwrap(),\n 474|        )\n 475|        .unwrap();"}],"details":{"scopePath":".worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs, .worktrees/287-shellrestore/crates/spt-store/src","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\crates","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":10,"fileCount":2,"files":[".worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs",".worktrees/287-shellrestore/crates/spt-store/src/shellinfo.rs"],"fileMatches":[{"path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs","count":6},{"path":".worktrees/287-shellrestore/crates/spt-store/src/shellinfo.rs","count":4}],"truncated":false,"displayContent":"# .worktrees/287-shellrestore/crates/\n\n## spt-daemon/src/\n### shellhost.rs\n  723│            );\n *724│            let _ = std::fs::remove_file(shell_perch.join(SHELL_PID_FILE));\n  725│        }\n  726│        // Unknown identity on a DESTRUCTIVE path fails toward not firing: a\n  727│        // mis-fire force-kills a stranger's tree, a missed kill leaves our own\n     │...\n  780│// [impl->REQ-SHELL-2]\n *781│pub fn close_shell(\n  782│    owlery: &Path,\n  783│    owner: &str,\n  784│    shell_id: &str,\n     │...\n  859│    }\n *860│    let _ = std::fs::remove_file(perch.join(SHELL_PID_FILE));\n  861│    // Retire the live credential — but STASH it (rename, atomic) rather than\n  862│    // erase it: pending spool rows are stamped under this token's key, and the\n  863│    // relink that recovers a persistent instance must re-stamp them to its\n     │...\n 1233│        std::fs::write(\n*1234│            perch.join(shellinfo::SHELL_LAUNCH_FILE),\n 1235│            serde_json::to_string(&launch).unwrap(),\n 1236│        )\n 1237│        .unwrap();\n     │...\n 1763│        std::fs::write(\n*1764│            perch.join(shellinfo::SHELL_LAUNCH_FILE),\n 1765│            r#\"{\"pid_started_at\":123,\"launched_ms\":1}\"#,\n 1766│        )\n 1767│        .unwrap();\n     │...\n 1815│        assert!(\n*1816│            !perch.join(shellinfo::SHELL_LAUNCH_FILE).exists(),\n 1817│            \"the pre-stamp shape is the ABSENCE of a launch record\"\n 1818│        );\n 1819│\n\n## spt-store/src/\n### shellinfo.rs\n 131│/// launch (BAROMETER W2, releases#78).\n*132│pub const SHELL_LAUNCH_FILE: &str = \"shell.launch.json\";\n 133│\n 134│/// What a launch recorded about the process it started.\n 135│///\n    │...\n 170│    if let Ok(json) = serde_json::to_string_pretty(&launch) {\n*171│        let _ = atomic_write_string(&shell_perch.join(SHELL_LAUNCH_FILE), &json);\n 172│    }\n 173│}\n 174│\n    │...\n 178│pub fn read_shell_launch(shell_perch: &Path) -> Option<ShellLaunch> {\n*179│    let raw = std::fs::read_to_string(shell_perch.join(SHELL_LAUNCH_FILE)).ok()?;\n 180│    serde_json::from_str(&raw).ok()\n 181│}\n 182│\n    │...\n 471│        std::fs::write(\n*472│            d.path().join(SHELL_LAUNCH_FILE),\n 473│            serde_json::to_string(&bogus).unwrap(),\n 474│        )\n 475│        .unwrap();"},"isError":false,"timestamp":1789001187767}}
{"type":"message","id":"aef51d89","parentId":"ee7991a4","timestamp":"2026-09-10T00:46:27.768Z","message":{"role":"toolResult","toolCallId":"call_UpgQE4rfSVo6QpfvSZ2H1fCv|fc_077f3569365e313d016aa1fde3a37887d1af10a446fd33cf00","toolName":"grep","content":[{"type":"text","text":"# .worktrees/287-shellrestore/crates/spt-proto/src/\n## emit.rs\n 148|#[macro_export]\n*149|macro_rules! emit_line {\n 150|    ($w:expr, $($arg:tt)*) => {\n 151|        $crate::emit::write_line($w, &::std::format!($($arg)*))\n 152|    };\n...\n 186|#[macro_export]\n*187|macro_rules! emit_line_err {\n 188|    ($($arg:tt)*) => {{\n 189|        let _ = $crate::emit::write_line(\n 190|            &mut ::std::io::stderr().lock(),\n...\n 360|    #[test]\n*361|    fn emit_line_issues_exactly_one_write_of_the_whole_line() {\n 362|        let id = \"endpoint-7\";\n 363|        let detail = \"the store did not persist\";\n 364|\n...\n 377|    #[test]\n*378|    fn emit_line_is_one_write_for_a_static_literal_too() {\n 379|        let mut w = CountingWriter::new();\n 380|        emit_line!(&mut w, \"PUMP_IPC_READER: exited\").expect(\"write\");\n 381|        assert_eq!(w.sole_write(), b\"PUMP_IPC_READER: exited\\n\");"}],"details":{"scopePath":".worktrees/287-shellrestore/crates/spt-proto/src, .worktrees/287-shellrestore/crates/spt-daemon/src","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\crates","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":4,"fileCount":1,"files":[".worktrees/287-shellrestore/crates/spt-proto/src/emit.rs"],"fileMatches":[{"path":".worktrees/287-shellrestore/crates/spt-proto/src/emit.rs","count":4}],"truncated":false,"displayContent":"# .worktrees/287-shellrestore/crates/spt-proto/src/\n## emit.rs\n 148│#[macro_export]\n*149│macro_rules! emit_line {\n 150│    ($w:expr, $($arg:tt)*) => {\n 151│        $crate::emit::write_line($w, &::std::format!($($arg)*))\n 152│    };\n    │...\n 186│#[macro_export]\n*187│macro_rules! emit_line_err {\n 188│    ($($arg:tt)*) => {{\n 189│        let _ = $crate::emit::write_line(\n 190│            &mut ::std::io::stderr().lock(),\n    │...\n 360│    #[test]\n*361│    fn emit_line_issues_exactly_one_write_of_the_whole_line() {\n 362│        let id = \"endpoint-7\";\n 363│        let detail = \"the store did not persist\";\n 364│\n    │...\n 377│    #[test]\n*378│    fn emit_line_is_one_write_for_a_static_literal_too() {\n 379│        let mut w = CountingWriter::new();\n 380│        emit_line!(&mut w, \"PUMP_IPC_READER: exited\").expect(\"write\");\n 381│        assert_eq!(w.sole_write(), b\"PUMP_IPC_READER: exited\\n\");"},"isError":false,"timestamp":1789001187768}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_SKk6odAzQRwxvYbCrKCEym21|fc_077f3569365e313d016aa1fdf3c33887d191e5ff7f66816049","toolName":"yield","startedAt":"2026-09-10T00:46:50.379Z"},"id":"68bb97d2","parentId":"aef51d89","timestamp":"2026-09-10T00:46:50.379Z"}
{"type":"message","id":"9576f6c1","parentId":"68bb97d2","timestamp":"2026-09-10T00:46:50.379Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":"**Checking owner-edge restore deadlock**\n\n**Confirming birth-stamp race**"},{"type":"toolCall","id":"call_SKk6odAzQRwxvYbCrKCEym21|fc_077f3569365e313d016aa1fdf3c33887d191e5ff7f66816049","name":"yield","arguments":{"type":["findings"],"data":{"title":"Keep the PID and birth stamp coherent for launch arbitration","body":"The new birth-aware arbiter can reject a launch that is already running during a normal relink. Both `launch_shell` (lines 498–504) and `launch_shell_brokered_in` write the new `shell.pid` before replacing `shell.launch.json`, while the previous launch's birth stamp survives close. If a late-ack CLI fallback or `launch_shell_daemon_side` checks this interval, `shell_pid_provably_dead` compares the new live PID against the previous process's birth, returns true, and the caller launches a second binary and rotates the first binary's token. Previously the bare live-PID check stood down in this interval. Publish/read a coherent PID–birth record, or otherwise keep an in-progress publication from being classified as a recycled corpse.","priority":1,"confidence":0.95,"file_path":"C:/Users/decid/Documents/projects/spt-core/.worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs","line_start":320,"line_end":322},"error":null}}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":9142,"output":690,"cacheRead":62720,"cacheWrite":0,"totalTokens":72552,"reasoningTokens":435,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1789001187816,"responseId":"resp_077f3569365e313d016aa1fde42b7487d1a8d861b669a9da78","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_077f3569365e313d016aa1fde599e887d1b010805a8a2a97cb","type":"reasoning","content":[],"encrypted_content":"gAAAAABqof3z7CfQ82XHkDIxVJuSysapno6HfiK_F4sOG1Bi0XMvERq4P6N7Tx8zrTGzkNEwvZSU1-pZDd-Va10lAMuwS39Gb18G9gySX7O76KXyV7sevRp6nI8IzMD4wEQxtzUes9KnOkjJ0SzWhl_CHaL4ZQ0HibGod0UN36Ys3UDu1nErIDEIf_GNuYEcvrR5-Bd7h2M8f30-18pLHcYzLniDJtnSp-kbClQMzdXiEN2Y0x0qhXnukLzxT5Vc4pBFUTMHWl-262fm391qc3QyH8ZS5xJSy8crvBrN8deyirqVN0vUSrqjW0ixXSWb5J3MEuk8WID-KvlawCPdwcQ-7y-7A-vUHNTB7lDUP0oIEZuZJaril5WnH_Oo9RPaBMOEvM7KO6tdjVcTY6yezWVvRCoiKZ8fVx84nUViHPlguzMj7iSfBY5AKpL98XUmTRYj9vlpR_yoUFPnzEQHv8-prUXf0SiNmrakITlG1ch8yRVXIo5LDsHLduqTWj7UzxMv8l4j2kZJYJ9TzsDXSlH8AzqhUbIBYEbCjtQeAvosVm-GDsa5XVT0Bm_iVw12-Wki3Mm6hEnfDkCTp_IPzJjp675YKGYWOEwKLjVXBod6Nz6QTxaI42WRcs9CBA_tXRmmnuI4zFKJJrpjyC2Q8jnC_9Jzs-5k25gaivcwvJMHTsWW5d3sib0pvs9kmaFBq9GPaXveb99T4ECxqKGp_BUASY4z7e48mEJOv6jNwy1ZvI1NfhPVXPqmdpTgidkEPgmN_d-B7Jb8goucBOHfUbccFu4GahjAEiiML5hhomhN7NDzSAh_phlCc2Am_X0hnFT_ftyWftWgOwaRq7_ZR2zpsOFUCvicETN_euB-_MColA4fhmj8JmLDJm-ysxuYOcY5IlAHu_LZzgvcPv29Q_xs0w-chiVeC_Pw5_ZizUT77pHUfmFBmHRrCFhHHwsCr_i_b-sVUy5TsDVR8yeKxC39rFtI9b-0wiRgKszwPNQ3evAV_3JZm9tcOb2oBAjmBpSjFgAWNlYHOI1gnvaEbOp5Fsed0QOkzZe_dNIXlVNIrRZtgMhiWtB7uwse4BRyg7z0u9dJaM9Kr4hab2kAgwehWi8xEmiZe-qJvynGCDr2htGR09wQMV6dLNRqjYy1MHgqSyalUmQTH1CzNpn_B1pzpLRVQUA5pvBI5pHWN6cu56kPrh6kqqRgdzNFoiS57MI4EnuPI2xnI6IPpesEYXwG3nzJRbGjDwGGRc-tlrAj0gb5Yv8NTV8rr78mFePBx6KwVxCYd73EfQT0HjdDhSF5gDZcQobEqKW7n7MDCk8xhNmCak7C2pH5XqdFvw47OT_KGfd-v95feHno2TOCkNIB-ZkqvGjAwjPk4d33Dovkh8UBfaAiF5LutAavI8f3S7IWZVSREoAoDGgqsCTH8q2S-YHQsVJZHyNTHz__7kcpY67xC7_MideHj6KtrPOrPgFK9hSIleRyvq07iS5zaSRbEsXtL751Mw0ajGdgdPwSdElrMpHVi8h5T_fvjgrhyw9wX1_UBmO-fIVPWVD5e3uyZBuZDbPQGHtWcrY6agqfX2UIZgxoMajsMRWvvGfvmBgd4PWPE-MInga38JVGBysYPe0aQGMgv8iVFIaHpyX3b-wxz6Bk4e0c7K5_Zybvp13iftFTiXVcDP9sEk9fJXCjSwS2uCmNBP4deAdX7bKzr_K_LmiycKL-eM_6haB8XZorcEaIdWgt95QB0bQxC_iUBr7UdmtjHDeB4H2yEUpo86G7u8NDuSWV-vSVcslsX4AgE_ezWSO9-nC_RayNKkqzPFgbgbUiAT2K6XBkjak02E0HZAH6YdSrBMRFokuigp0j6j2H_MULMr51eXOQSg2D6kA1msVz34HmITewGv_JuQKxcTzWuMJpS53hVlE6g80njRJu8B_DUg9QOyVge4XumsNzhhYA01wQrCUwyHZ1QpE3O3Xgi6ULVorC5GY26Bqftk_2CYsg5bJyS5MQdraTwlVKD004tX3YejbPDPzkbMgf1vDlHV8LucjWxk1LU6woUeC-sSbONHXgKvcrexyEBxuN7zkNc9PaY0hDtVRnncl3QJNOMf6pY4s744xFXjobES06_Tl_9t2X1JXyccvgPwqOj4idtrm1autgNTN00q915RRzefywtfLH-jsj158zOC-dzUVJjYyH_j8dIGly0eLPgG3DMHdpyp8dRYKeaSiguiCD721lk8diGyCgIopq-Fl65AX9_ALDyoBUftRy1waIMzVRg-eK-yoIwJUI7T50pFOs0JM6gZbVzrjIs6X7I-rA4FseXu_TbUn_mZljAK0khr_4kxFVdzifwFejgHCL-20TLbGaE88CN_903BPIyMIgSX7tiUwSscZmK_CsoQ4SJVXg6wt6rvlaWWewjooB64E3oWF7d7w-kZGZ9tyv8YDqhD1tzCnyJqWy7tG_hSJfAX2_dhys_xVbuN1DVvmf-BlioGPCSLJ4wkUe8rQmfnNN8SrIA-YU0WfQPoiC8jVQ3CWuN01tcQiqxsoqz7NyzL7UJJdXjxsg25rWrcZTGZEXtK6J3arfxNquViVHaEIh58hNTjNNLtIjmfiChn4yGRnYvI7hedtlR51lUyLPlbGsEWuDlOLTdRgqJp8yezii_VKbyCBRyJpg-crOz9F6AlG-wUnJG0CYzjfouGFCNjM_oNmGIYxVuowZgSYEFd43VQX-KwrejpwHE-8tTdwqg1ull0kI7pBsD2eWYy4NEZIh6Tg1nMazS3xjuYxtlpJqzUIaLd1O26qD2miU_aO39KswBfLr7zd4alWg1Cz8zv3IniyX_1D7EiqA1Z8RDuxu-wB1UpeWMqzMO_SPJQIVg_A12YH-0VBk8xu50Iexn-TD1jF4FYpZJAYwflkK_gNZbCBmK1vqpMix_WzalPoam7aZPSjrfTd3Eg60vz7P1OsBY7XTn7-dm_hZeCE5XLve7LBY7AbB20YmGSwy7zXUIPz3b0LQcaYbipCKcuivvHphaqHdYxqpXRJG-LLki_0hGDjQy0qqs62t8dKCI0fp9yCv9OIGwFYkzwEnSejrWl11e6Rs1uG0R__lQjFOKz9YKZB6vkVDzF5QPBA2SD7o16ZOkzTACRs5_wlkMZ6evtMG7P7Z6a4Pnsvuwlx-BWtGr3-b18Af_itvARhy_sNs1Y4WJXDiGOmBY7I-56Q2W8ie8cAQXFjrpWArHK4WPYNTS3cAYZFU_oVJRHVQJyKsbDK8Fe0JTD1F8udas3-2zZLr7KOsBl2rE9vkU6gpJTELOcRL7vNCG4J2whfvvxUrtcbaYGQh59i6YbEEDs1JH7LzsBHob67Pqu8fbASb87Sw8Eqy_GKKLBh-slb4kT2G1xfHbXzrS2aPajKcyLARaX8NaxVD8UFQ2rLFfoYEoPz4Ne5mg65eBkxakBKqvNJM_9Y92Fpyj6R7SJ5LuBSTU8nx9aak1j0myiSs8WK-DIqzZyglRhFv_bQzjo9pdl3-CZe4bgb0x1qUZlTLZ0twLOFwLtj1-8gntSy-5FgLFSPraabOcn5x3mTdwWncJ7ruoGvV3WZuQfzz-GqCUkNLn2JcnM0Agc6Plx2oR-qvtj3VroY_BHRALJwlF6EWUctv8OxwZJesKg8k7GT4bvYrDv35aWiEwpDjF5VF8diGTOPecMS4fyta-EIAtX8IoZ-tXgFSX2gsLeV2WKHb_m_IcJoWCrV1bHTF8WTLijgFt0WGSYdDuLkFPz_d84JcYt4e6y9crQEIwKIN6cZU9-G-TUjfDL9BeTRu0yEMkiZb9UjCAltdfyD5_Oe4fO-_SVMD4LA6uU2bUdnPXAqA1gvJta8Rk-Xu0jnen45SBaK8xZzaTSZxV5HB3X5gYYUQyNIy5s_DeCk4BkPoeed9Lsk_YR7WcaN1w6XrSNaJ3XIhCxHD8Wxrd9iCi5MN1JLsYCLyxJGrnMZAd2RzAhU2DvwiJI43n25nApVe9vaBYeZWCxYXcudhOcjxvRgDr8KAz_1q2pBo5jgq_hH-PMVFtjEomDVuxCVViRfvh1qqWj97iXlDCq9SbBn-fXmke-578tmHZAmEzsTPh0y6CCYb4uPWmprCZ9X1OfIod_4CEbYrEVRcFbqYRd1wXL5ARX6S4Lco0vuNhiXixTXjcMKD1vTNGtNjpETzzvhVCrF5S11JShr4kxUrBRH8it7sJOhMNxhL2lsCZjFDdR2c8sSOmtlztVTj-tDVD4mlQiiLgdUDlHuUpbuk8vcWWN8_gVoNjb8GbXWlMfFLp47rWyi--rDU9CknynGO2APyKhwN2UU9wPZow30m3ADB-i1wWqCNWQTGoccDMG9TnXGJS8F7SGMRm7lflZ3ZLuJx1tN3L-uPbSUwkTHmpluImMbbqsBR1qwf8AMNcNaTti56340on9nKNdTDz_Z9NULrgtknOI1-8LoSXfPltvHerG0QGf5h1jItMESdswVt2xjtXZCjKJFqYNFuAcTIqYOD2mGpgusnwCFfo8kNwgbf9DJ2JHb8xBcIpOAW6FUZCIWGWQTC9OYft3_g1fX_","internal_chat_message_metadata_passthrough":{"turn_id":"0b013667-0aaf-4a85-84aa-5b88e3033a8e"},"summary":[{"type":"summary_text","text":"**Checking owner-edge restore deadlock**"},{"type":"summary_text","text":"**Confirming birth-stamp race**"}],"metadata":{"turn_id":"0b013667-0aaf-4a85-84aa-5b88e3033a8e"}},{"id":"fc_077f3569365e313d016aa1fdf3c33887d191e5ff7f66816049","type":"function_call","status":"completed","arguments":"{\"type\":[\"findings\"],\"data\":{\"title\":\"Keep the PID and birth stamp coherent for launch arbitration\",\"body\":\"The new birth-aware arbiter can reject a launch that is already running during a normal relink. Both `launch_shell` (lines 498–504) and `launch_shell_brokered_in` write the new `shell.pid` before replacing `shell.launch.json`, while the previous launch's birth stamp survives close. If a late-ack CLI fallback or `launch_shell_daemon_side` checks this interval, `shell_pid_provably_dead` compares the new live PID against the previous process's birth, returns true, and the caller launches a second binary and rotates the first binary's token. Previously the bare live-PID check stood down in this interval. Publish/read a coherent PID–birth record, or otherwise keep an in-progress publication from being classified as a recycled corpse.\",\"priority\":1,\"confidence\":0.95,\"file_path\":\"C:/Users/decid/Documents/projects/spt-core/.worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs\",\"line_start\":320,\"line_end\":322},\"error\":null}","call_id":"call_SKk6odAzQRwxvYbCrKCEym21","internal_chat_message_metadata_passthrough":{"create_time":1789001188.521623,"turn_id":"0b013667-0aaf-4a85-84aa-5b88e3033a8e"},"name":"yield","metadata":{"turn_id":"0b013667-0aaf-4a85-84aa-5b88e3033a8e"}}]},"duration":22560.864800000098,"ttft":1838.1713999998756,"completedAt":1789001210378,"contextSnapshot":{"promptTokens":71862,"nonMessageTokens":13209,"compactionEpoch":0}}}
{"type":"message","id":"b3ee8b67","parentId":"9576f6c1","timestamp":"2026-09-10T00:46:50.381Z","message":{"role":"toolResult","toolCallId":"call_SKk6odAzQRwxvYbCrKCEym21|fc_077f3569365e313d016aa1fdf3c33887d191e5ff7f66816049","toolName":"yield","content":[{"type":"text","text":"Result submitted."}],"details":{"data":{"title":"Keep the PID and birth stamp coherent for launch arbitration","body":"The new birth-aware arbiter can reject a launch that is already running during a normal relink. Both `launch_shell` (lines 498–504) and `launch_shell_brokered_in` write the new `shell.pid` before replacing `shell.launch.json`, while the previous launch's birth stamp survives close. If a late-ack CLI fallback or `launch_shell_daemon_side` checks this interval, `shell_pid_provably_dead` compares the new live PID against the previous process's birth, returns true, and the caller launches a second binary and rotates the first binary's token. Previously the bare live-PID check stood down in this interval. Publish/read a coherent PID–birth record, or otherwise keep an in-progress publication from being classified as a recycled corpse.","priority":1,"confidence":0.95,"file_path":"C:/Users/decid/Documents/projects/spt-core/.worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs","line_start":320,"line_end":322},"status":"success","type":["findings"]},"isError":false,"timestamp":1789001210380}}
{"type":"custom_message","customType":"async-result","content":"<system-notice>\nBackground job PersistentRestoreReview.RestoreTestsReview has completed. Resume your work using the result below.\n<task-result id=\"PersistentRestoreReview.RestoreTestsReview\" agent=\"scout\" status=\"completed\" duration=\"1m22s\">\n<meta lines=\"31\" size=\"3.2KB\" />\n<output>\n{\n  \"summary\": \"No actionable patch-introduced correctness/spec findings in the assigned tests and CLI relink slice. Source review only; no validation commands or runtime tests were run.\",\n  \"files\": [\n    {\n      \"path\": \"crates/spt-daemon/src/shellwake.rs:1080-1312\",\n      \"description\": \"Reviewed changed heal-enumeration, boot-restore and owner-online-edge tests plus their complete fixtures. Exact expected instance sets cover persistent same-boot and unstamped corpses, offline-owner/nonpersistent exclusions, live pre-bind exclusion, and every healed sibling across owners.\"\n    },\n    {\n      \"path\": \"crates/spt-daemon/src/shellwake.rs:1665-1831\",\n      \"description\": \"Reviewed nonpersistent corpse-freeze tests and persistent watcher contrast. Same-boot and unstamped nonpersistent corpses remain excluded; persistent eligibility is directly exercised with no launch stamp and no boot oracle.\"\n    },\n    {\n      \"path\": \"crates/spt-daemon/src/shellhost.rs:1207-1238\",\n      \"description\": \"Reviewed changed launch-winner test: absent/zero/malformed PID, live legacy PID, matching native birth, and a mismatched birth representing PID reuse.\"\n    },\n    {\n      \"path\": \"crates/spt/tests/shell_relink_force_e2e.rs:80-257\",\n      \"description\": \"Reviewed complete forced-relink E2E context. Added lines 187-196 require a status=binding token on the actual SHELL_RELINKED diagnostic; surrounding assertions retain stop-before-relaunch, eventual bind, process replacement and stable identity checks.\"\n    },\n    {\n      \"path\": \"crates/spt/src/cli.rs:23482-23538\",\n      \"description\": \"Both local forced and ordinary relink share the changed success diagnostic at line 23535. No record-state mutation was introduced.\"\n    },\n    {\n      \"path\": \"crates/spt/src/cli.rs:23606-23668\",\n      \"description\": \"Remote relink dispatch uses structured WAN outcomes and its own SHELL_REMOTE renderer, not parsing the modified local success diagnostic. Repository searches located no production consumer dependent on the old SHELL_RELINKED status=offline token.\"\n    }\n  ],\n  \"architecture\": \"The tests use real temporary perch/adapter records and production restoration entry points; restored owner-qualified IDs are compared as exact sets. The PID-reuse predicate has a focused shellhost test. CLI local relink renders a shared success result, while WAN relink has independent outcome dispatch.\",\n  \"report\": \"Compared assigned hunks against the immutable cb756dee^ → cb756dee patch supplied as artifact://112 and read complete changed test contexts. No concrete introduced defect identified. Coverage qualifications, not findings: the restoration tests themselves obtain a real boot oracle; the explicit missing-oracle assertion exercises watcher_eligible rather than the full restore entry point. The added E2E assertion proves the emitted binding label, not a deliberately held pre-bind interval; an existing CLI unit test separately asserts launch leaves the record offline until bind. These limitations do not establish an implementation defect and are not requests to add tests. Production restore algorithm audit was intentionally left to the parent reviewer. All conclusions are source-review observations; tests, builds, formatters, linters and production operations were not run.\"\n}\n</output>\n</task-result>\n\nPersistentRestoreReview.RestoreTestsReview is now idle — message it via `hub` to follow up; transcript at history://PersistentRestoreReview.RestoreTestsReview\n\nStructured output: schema valid; full payload at agent://PersistentRestoreReview.RestoreTestsReview, fields via agent://PersistentRestoreReview.RestoreTestsReview?q=.<field>\n</system-notice>","display":true,"details":{"jobs":[{"jobId":"PersistentRestoreReview.RestoreTestsReview","type":"task","label":"PersistentRestoreReview.RestoreTestsReview","durationMs":82759,"schema":{"source":"agent","mode":"permissive","status":"valid","data":{"summary":"No actionable patch-introduced correctness/spec findings in the assigned tests and CLI relink slice. Source review only; no validation commands or runtime tests were run.","files":[{"path":"crates/spt-daemon/src/shellwake.rs:1080-1312","description":"Reviewed changed heal-enumeration, boot-restore and owner-online-edge tests plus their complete fixtures. Exact expected instance sets cover persistent same-boot and unstamped corpses, offline-owner/nonpersistent exclusions, live pre-bind exclusion, and every healed sibling across owners."},{"path":"crates/spt-daemon/src/shellwake.rs:1665-1831","description":"Reviewed nonpersistent corpse-freeze tests and persistent watcher contrast. Same-boot and unstamped nonpersistent corpses remain excluded; persistent eligibility is directly exercised with no launch stamp and no boot oracle."},{"path":"crates/spt-daemon/src/shellhost.rs:1207-1238","description":"Reviewed changed launch-winner test: absent/zero/malformed PID, live legacy PID, matching native birth, and a mismatched birth representing PID reuse."},{"path":"crates/spt/tests/shell_relink_force_e2e.rs:80-257","description":"Reviewed complete forced-relink E2E context. Added lines 187-196 require a status=binding token on the actual SHELL_RELINKED diagnostic; surrounding assertions retain stop-before-relaunch, eventual bind, process replacement and stable identity checks."},{"path":"crates/spt/src/cli.rs:23482-23538","description":"Both local forced and ordinary relink share the changed success diagnostic at line 23535. No record-state mutation was introduced."},{"path":"crates/spt/src/cli.rs:23606-23668","description":"Remote relink dispatch uses structured WAN outcomes and its own SHELL_REMOTE renderer, not parsing the modified local success diagnostic. Repository searches located no production consumer dependent on the old SHELL_RELINKED status=offline token."}],"architecture":"The tests use real temporary perch/adapter records and production restoration entry points; restored owner-qualified IDs are compared as exact sets. The PID-reuse predicate has a focused shellhost test. CLI local relink renders a shared success result, while WAN relink has independent outcome dispatch.","report":"Compared assigned hunks against the immutable cb756dee^ → cb756dee patch supplied as artifact://112 and read complete changed test contexts. No concrete introduced defect identified. Coverage qualifications, not findings: the restoration tests themselves obtain a real boot oracle; the explicit missing-oracle assertion exercises watcher_eligible rather than the full restore entry point. The added E2E assertion proves the emitted binding label, not a deliberately held pre-bind interval; an existing CLI unit test separately asserts launch leaves the record offline until bind. These limitations do not establish an implementation defect and are not requests to add tests. Production restore algorithm audit was intentionally left to the parent reviewer. All conclusions are source-review observations; tests, builds, formatters, linters and production operations were not run."}}}]},"attribution":"agent","id":"d6c8717b","parentId":"b3ee8b67","timestamp":"2026-09-10T00:46:50.381Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_4NIuSZYxiiK8uy8afR1ZF5FB|fc_077f3569365e313d016aa1fe13229487d19159ce9c0e9a1bc2","toolName":"read","startedAt":"2026-09-10T00:47:15.258Z","args":{"path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs:1660-1850"},"intent":"Reading watcher regression context"},"id":"70a20143","parentId":"d6c8717b","timestamp":"2026-09-10T00:47:15.258Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_W2T3CZeTyRUlSBI7Onuv67PH|fc_077f3569365e313d016aa1fe1322ac87d1a4e59d56c089b7e8","toolName":"read","startedAt":"2026-09-10T00:47:15.259Z","args":{"path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs:744-890"},"intent":"Reading retained launch stamp cleanup"},"id":"6241c1cc","parentId":"70a20143","timestamp":"2026-09-10T00:47:15.259Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_up92j1SMTbYoqjfZFek1mg3D|fc_077f3569365e313d016aa1fe1322bc87d1aa206667acff4cd3","toolName":"read","startedAt":"2026-09-10T00:47:15.259Z","args":{"path":"artifact://112:43-78"},"intent":"Reading immutable launch hunk"},"id":"b1483f1d","parentId":"6241c1cc","timestamp":"2026-09-10T00:47:15.259Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_Uf11ZcsHgDBb3tgA7i3MdgNg|fc_077f3569365e313d016aa1fe1322c887d1a176c8b940e06ab7","toolName":"grep","startedAt":"2026-09-10T00:47:15.259Z","args":{"path":".worktrees/287-shellrestore/crates/spt-store/src/proc.rs"},"intent":"Locating process identity backends"},"id":"fd1e58cf","parentId":"b1483f1d","timestamp":"2026-09-10T00:47:15.259Z"}
{"type":"message","id":"defb2967","parentId":"fd1e58cf","timestamp":"2026-09-10T00:47:15.259Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":"**Reviewing the full diff**\n\n**Inspecting persistent watcher handling**\n\n**Flagging persistent watcher routing**"},{"type":"toolCall","id":"call_4NIuSZYxiiK8uy8afR1ZF5FB|fc_077f3569365e313d016aa1fe13229487d19159ce9c0e9a1bc2","name":"read","arguments":{"i":"Reading watcher regression context","path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs:1660-1850"},"intent":"Reading watcher regression context"},{"type":"toolCall","id":"call_W2T3CZeTyRUlSBI7Onuv67PH|fc_077f3569365e313d016aa1fe1322ac87d1a4e59d56c089b7e8","name":"read","arguments":{"i":"Reading retained launch stamp cleanup","path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs:744-890"},"intent":"Reading retained launch stamp cleanup"},{"type":"toolCall","id":"call_up92j1SMTbYoqjfZFek1mg3D|fc_077f3569365e313d016aa1fe1322bc87d1aa206667acff4cd3","name":"read","arguments":{"i":"Reading immutable launch hunk","path":"artifact://112:43-78"},"intent":"Reading immutable launch hunk"},{"type":"toolCall","id":"call_Uf11ZcsHgDBb3tgA7i3MdgNg|fc_077f3569365e313d016aa1fe1322c887d1a176c8b940e06ab7","name":"grep","arguments":{"i":"Locating process identity backends","pattern":"process_identity|is_process_alive","path":".worktrees/287-shellrestore/crates/spt-store/src/proc.rs","case":true,"gitignore":false,"skip":null},"intent":"Locating process identity backends"},{"type":"toolCall","id":"call_17oGaDdBH2hPPUCZfdw5D1YS|fc_077f3569365e313d016aa1fe1322d887d1a98ae52f0506a127","name":"read","arguments":{"i":"Reading emission sink contract","path":".worktrees/287-shellrestore/crates/spt-proto/src/emit.rs:130-205"},"intent":"Reading emission sink contract"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":1735,"output":750,"cacheRead":71680,"cacheWrite":0,"totalTokens":74165,"reasoningTokens":516,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1789001210393,"responseId":"resp_077f3569365e313d016aa1fdfadba887d18d179723d7909134","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_077f3569365e313d016aa1fdfc27ac87d1b0a2d157c528c2a9","type":"reasoning","content":[],"encrypted_content":"gAAAAABqof4THkZOvFELKmmrT-6SSdyk-lDkh-ODoV3g3pSb66Sxe8bCtq_9KeW4Vt0FrYJr3UH_Pw0Yif5P-_JVRYnCDUNTyitdRdOvGpSVx-43yyf-Oc5nnlt-sfNJYHJ0TcMNkYdykD5AuNpqH_0iwQUIz24ZhhH7aAiYhH2zYz5moL3orIpN8WiwIJM9-lu7mL6ks04_AOJXFnp2C448Ui66UOWviX3HvrjU-iO7AsBsfLHjZPht7m1t5QlfHJmpL3AeI7B4u9S-vQFgnQWOJoV31PlLFC2bmpvxGkP7KyopaLsKBdYL-ZaDmH3QjXtkpDQIlwPoJz5B4Qj5uS1FPShVGfDMNGoWj-M9Zy1q9b2An7C_ozTtroV9GiMkJ6NNR7C5zYjNiuU4dK1Zacll_xaOMosjzKYBT-uPqJTqrSeojg_14dbZftnPvLHj2546gUDvX9jKPWzlIaddMmeCjMH-uUUOCmUiTIZVwsfWZwcKCYEgpeD4iZPvKOHANiL7duvAAwEzEq7UdkA_Z4gAZ-_A4gmtSPBbK8JAagxnXAV5bST-Auq7UWfhI8KfwC6szP2Bb4yT_vDlEAOJ5LEJxpcRwTU2LkyDxuN32qRjon92CJ8KFP1yaLJvaaXTZ5upgXSJPiOe_H0O6UjZC3uD3Vk7ZQML70MaRCUGHj-rhIKtCZqTv29kAClgACdIbzsaLofdsXQ-WRozG2g0o5axJqgRVlwr5NJDKFEpQaF3xu_1EhZg8OiHMfpx9RNAaqNalNEz75Dk27wrtJXu6addDv-mpeMBbzsLKm9w6ojHXKayGnpaWPoYycD2h09j91bX9Kpmlccox4aiJr3Kl0jRIxjxsXYZaHiB-bOrRbmbFV6voS3A59rdn5M6QHdfmcPhQONpBmK7u-kp58CO_2QqKhqfJXurjC8S7VQzOARcw_jiYt4J9WtCzV-DCSLV85ttyFr2hkGhkRiZs_8ZqDYvbLd_60t_aq6lEZeSPGp2UL4QCd4Q5zqtNKHz9VjYLwxPKgMNfwDiWGRsaUWHPETJBHIJVqEPBVbGDHGSD3qXmaZYs6ap6NKmE97IlEzAMfonJ3Q5DaHgUy4-ec99wLnBaeMCETOMFzf8Sior3Ir7QatHbGqsi_hsFvIjrY8nTHDHf-GrH89dO7lrxSH8EgQr1kwWW68DcQ6pdHWFAWifk6itC9HgqacrQ8--YKIKmb5bNtffWYd-KmbC5CywyoxjcVfI4us3eVBcJiKBFbg-QjRm0FXIqT2B7uLZjY7r4LG15oS7Y5UYtAJMEMnqGTNsOq5GzkoZFekceW9nSkccZXnUXVS4qGmaHZObLEYucpDl0m-C8xeOhn3Tz8kWdkwe9MdrKNFZxUiSQvMKskG1pQtPYsmd-m6AXEaNd2q3iammA8saXxBnrqqJ5DxxfgMAxJVwHgn5Jv_y3C1Y_CkH-NB7G_JA9SCLpJ0NQkRg-ZZsvhuxO4GUVVg1I4fNdKKy36DErtvuQH2yX5ETr3wIHTlMyTxUAGSouPI9qoVigcfliFb5dm_sV5JYiPwcGZMNsh5bn525dp-HZMZdPMZXV3UOkOoNX04HvpkYcw_x_j7EJZOorMgO84ltVMoAgT2yjEeAEN7zpgkrAngFnA7Cn_hdbe5p6ukUGG08viny7OJmTwSp0blrP-vZ7kQqOxdhk5KLkdwznLve_PBdQoHTzbmNMkTeqMpjvADvMfk3jQz_P0A5uZ2Wrqn0c-r7zNLBvNi1C5z351LP1OCKi0OQLqfPaRY-BZKJAaH_9fR5Vg0ayeMenYBdaovpV8PCk-zlR7uM6MGrVIGC267fv22yl1-pZs6LvX6mKEwmP_FKFl7Fk1Z73f5uNvFm6K9ZSd6pUVPFBVSENXgiImeaa7Kk0Wf5jNQSoL20Et-MJmHGlw94GRy9bxjJ3hKkDBkoxzJkOtEB028j-NX3Eq1tZ0yq9T_mnWaWvKqdmADHbYbCgOHTtp6HbuPADH0BHz-mM9Vhq0SlEcTGG8JLtrBThF7UaZgAck-OHh2oeyEPVXCKsIw_5oi1dik3w5ygGo8XhJSTuARa4cATVKuxdwPUCRK0qrFAcQg7eCJmMjQvbgvcY_gI6Jc8xH8ZtdFiq-Vqx0QRqQPiFZdsmrQ2fhpuB3nZff5YaL7-kYl3Pi3eKITdqLZdYj5e0cryXwauj4WzE0qHZkLl6uWBZX-HoNVaZf1k-L4URiRciil1eIlcKnN4e5wi-g76ncW5TXDMKDToh868uBUKlFHwzli2s_eAQ60NfVbezkNQqjas6aDUhcOsawoHhVqZ4L68pmpeMI5ebHOMyxw9GURlfG8sD_VZPVtppvSyvYmtfdM0LHLYVoC5iERWGz8-_XRZm27MyKXq-cK_OBLHf15Ku03c8-jyoH9zxma8fG60oKapqL1MxIITi4tqc23Qgrr679CndajboZX8HzDmxFNmRrDkhrLFBhi1Ich22YVEL8ljeZ9D1cuq2kXRr5Vpjhzpqqr-u2lN7SbNPxLZ9594UV5s5nFHvRWlCwwBZiqThJFTeVJ190m0QSzI1sVcPt14PtKFx1IL9dTtmn1EgeHv-BDncC5IofkV5aTNP_hRHFhOMXhkrmJSwIRd7QX7T6LbEklGIsbWIM-WA4hoODYHrUEoHFZNyy9be0sKsyl2pLhFvV5ezfVIGPjAaqPH1WX77XxywHyVR1vrHwduEqNkllEGKBIRksjPCmGVEmacS5K5xgBWjTVHhT0BTwUHwzqR9qH3BupU4fExFZwU_56tEQbfwc_NIqrbeNrKXUNd7mM8QkrsOnc9IpFA6mnjEkcB6p4fbHbRCeGUvFEkIg4lufDLAusqWsDYRlTz4W3oEiY3S1NOtsWXGpZRCDGBDPjTGah5m5Uk9w-KhMsh6HGasI5mWbhtehDqJKIwnglFTJUO8ngZHPkTcKfASPPSsOAKpSTDS4j_XJ6boLPsY6x9YnWpiM9ajQvV0ASKdDu2i6XIDhW6aTRFfU14vBgbiJIoLYxDsfimUY90OvXIvC3gNQIhb8H2EFEnMihDc0p46kToo4LKhUzFSH1SYNCcgGqUw2wWAPGSaTsw7FrmcWPbH_NVxHSY7VQ5TT6HbTb-LUqFZHUNCIVIn_Uo5Um9kWhGkdNrxLfMQKE2HzUiQoggYqRVUL_kdTV-T3da8F4r7ZQyMaeOukm9RMO1Rk_idPd8QVlKu0MImY5XwFbGNdv_QNhkCmZ_lrao7FsIOK3qKGFdr640_Az8tIGGRrtG0nNQZteB1XmgssZkzDaO1jmQg-UnIa_VCLft_B-hnWxTqOEvOycEh1g2JFk8MMcJaHgdn9lMaXa3-O2lDVQYyK6U13fUNE8YkI0XrDqVCYPkWZFOSabPLN0rZZZ_zniU4YLyAWaVkOt3Gx14XBBZiGII0SnfdDWi3vtTYpMtWtzLYJuH273G8fb5rVdkkBSj3BZ5k4T0-Mm4U9UvfnDS7eBIIYy3Fpn4oP_aJNan-uV2h_BlB86_UKrfMTDWdmfxwLfXCnSAlKEKPcbRr478uIB3sIv9RFusMjSu7TDCG9J-oATQMge5eD8mZ5I3SYrjmy0gGvZr-IwX7sCqeDhkrcLOxYiGQyDXF7DhgA2jA5QlTkHbxIq1ZvuOu1UueQdoTYkrupEkHQljxG18k792qO5RSTa3fTx99gPktEBxsRHU-kmAcju3ccnca2Zie3-JW9A5bLegRNbUYjX31vqm68dqDuuGMj63zVswsOxByT8aRrToxXiumHw90FRzv-EKyV1XVgrf0ORJcWOPMQHxS5A8UeO9PmSPX4NxO1EORtiBESqIYLGZgWFIaayVN9YT0_tyNTbJZ6NxFK5gYSAWN5wo31iAZ4RwGw0BK47woqmz83HjSSNp2-HxSqf8VyUqvGK1W9j3vEZ7WZQ8xgZqmVocosoGWZYiNUB_dj2FIpv52JdddMWf_zNk8j54DUquWvpSIUQeQlPjfBC1AK4MMYcFy6Y6il8aJ0SI9JDP8GqbmIHm1iFYKPEzMZux56v47yVkIJ0BeQj14ALba-MTRSFFSCP5Q02XeeMDFVWmrSYMfnuluTs4RRqjA_8UKaGAqEDsOcbe5LcLTXCs8bl7nBW5jS4FjTnalBzQcYKjkkO_A_E4jSMuwF79-5fCClBjaPiNfb5xP2iIu6lOFGBQ9FCLoRiUxYBRwnaANx8Jprep5JGq04ixbsTuUTdGofW09g0XZ20DITlg772yljLN-S0thL3KVSfNH0j51nwhGB84w5VBeQ7fIR-gwU9CySvU8CT4omQaGl30h2uXvq5eFLP99Oc58h3G6nH07Ya0dVZQNbai1y0Toyuvbln3HR6I07NNa9hZMHvdEFI7Mlo7XUiXQnFUtNpAlgoIgXo8mvG6yFcX4pqJepv9ghw-6tAdUjXb38ST4aFY88j3nC3X6YeLuZ60mYF04heyXZP8FA_rttqa0D0blqDfGVku3EvfWYHSUd5igR0icKPV2jdVGbp6NSXtKMjbAh6l9IIiFk15phYvNftg2z4Eixy8DC-cWBXVggkYqumRMKZgljTles3jQFYCrgJuO2q9kuqJrcutCYl1O6ODkrDKiqJDPN84Yc--5BBCv8Hi7NaJQ_v3PyrfXM5vTOriGIXlaOBWtrYBuKoYjh-KcKfiLRenNClO2xfGGTndin63FOCyIOvzfo3ETkHlxMf6C4lq07-WiEj3tIdDt4h9bFsnL25ofRlhrj4nK7ixEGJrmWoVCv7i7znUEvilIl-eLjNEyZgz6V7bQL5IswlIDOiwFQdlrXc_OTwgF68QcgQ5YdNN3oJ56lHDazWhxkRHxhz3XzyNL36s0kRY8rFy-00RalYAbBKrGNbj7-lf2h3-0oGGvl1sQLMj9S-v6l3bPfAfbfKopXc4sVneoYnlyoxv-1EXHqpjVerW6bG4CaOsEOHZ99rq6SNeyr97Oiyql4pkcJLFDc342HklqscQggyw5QzX82nW1fs9Nak3UwQEeCP8yOLhRNao8ANXgbZ-kZcj3K3baZYB82pxYc0TmfjJaUAAgmHsbo1z6yDnjRrzSMooZyCqsbmbOzbWoKJ4U5U4EK1NcQxDKdI=","internal_chat_message_metadata_passthrough":{"turn_id":"605b9268-6ded-49dc-a954-665c66109b31"},"summary":[{"type":"summary_text","text":"**Reviewing the full diff**"},{"type":"summary_text","text":"**Inspecting persistent watcher handling**"},{"type":"summary_text","text":"**Flagging persistent watcher routing**"}],"metadata":{"turn_id":"605b9268-6ded-49dc-a954-665c66109b31"}},{"id":"fc_077f3569365e313d016aa1fe13229487d19159ce9c0e9a1bc2","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading watcher regression context\",\"path\":\".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs:1660-1850\"}","call_id":"call_4NIuSZYxiiK8uy8afR1ZF5FB","internal_chat_message_metadata_passthrough":{"turn_id":"605b9268-6ded-49dc-a954-665c66109b31"},"name":"read","metadata":{"turn_id":"605b9268-6ded-49dc-a954-665c66109b31"}},{"id":"fc_077f3569365e313d016aa1fe1322ac87d1a4e59d56c089b7e8","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading retained launch stamp cleanup\",\"path\":\".worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs:744-890\"}","call_id":"call_W2T3CZeTyRUlSBI7Onuv67PH","internal_chat_message_metadata_passthrough":{"turn_id":"605b9268-6ded-49dc-a954-665c66109b31"},"name":"read","metadata":{"turn_id":"605b9268-6ded-49dc-a954-665c66109b31"}},{"id":"fc_077f3569365e313d016aa1fe1322bc87d1aa206667acff4cd3","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading immutable launch hunk\",\"path\":\"artifact://112:43-78\"}","call_id":"call_up92j1SMTbYoqjfZFek1mg3D","internal_chat_message_metadata_passthrough":{"turn_id":"605b9268-6ded-49dc-a954-665c66109b31"},"name":"read","metadata":{"turn_id":"605b9268-6ded-49dc-a954-665c66109b31"}},{"id":"fc_077f3569365e313d016aa1fe1322c887d1a176c8b940e06ab7","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating process identity backends\",\"pattern\":\"process_identity|is_process_alive\",\"path\":\".worktrees/287-shellrestore/crates/spt-store/src/proc.rs\",\"case\":true,\"gitignore\":false,\"skip\":null}","call_id":"call_Uf11ZcsHgDBb3tgA7i3MdgNg","internal_chat_message_metadata_passthrough":{"turn_id":"605b9268-6ded-49dc-a954-665c66109b31"},"name":"grep","metadata":{"turn_id":"605b9268-6ded-49dc-a954-665c66109b31"}},{"id":"fc_077f3569365e313d016aa1fe1322d887d1a98ae52f0506a127","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading emission sink contract\",\"path\":\".worktrees/287-shellrestore/crates/spt-proto/src/emit.rs:130-205\"}","call_id":"call_17oGaDdBH2hPPUCZfdw5D1YS","internal_chat_message_metadata_passthrough":{"turn_id":"605b9268-6ded-49dc-a954-665c66109b31"},"name":"read","metadata":{"turn_id":"605b9268-6ded-49dc-a954-665c66109b31"}}]},"duration":24862.8358,"ttft":1788.180499999784,"completedAt":1789001235257,"contextSnapshot":{"promptTokens":73415,"nonMessageTokens":13209,"compactionEpoch":0}}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_17oGaDdBH2hPPUCZfdw5D1YS|fc_077f3569365e313d016aa1fe1322d887d1a98ae52f0506a127","toolName":"read","startedAt":"2026-09-10T00:47:15.259Z","args":{"path":".worktrees/287-shellrestore/crates/spt-proto/src/emit.rs:130-205"},"intent":"Reading emission sink contract"},"id":"5b765cdc","parentId":"defb2967","timestamp":"2026-09-10T00:47:15.259Z"}
{"type":"message","id":"b515eb17","parentId":"5b765cdc","timestamp":"2026-09-10T00:47:15.267Z","message":{"role":"toolResult","toolCallId":"call_Uf11ZcsHgDBb3tgA7i3MdgNg|fc_077f3569365e313d016aa1fe1322c887d1a176c8b940e06ab7","toolName":"grep","content":[{"type":"text","text":" 2|//! (dead-PID rows) before resolution. Ported from sister `types.rs`\n*3|//! `is_process_alive` (ADR-0001): Windows `OpenProcess` (tasklist hangs on\n 4|//! Win11 26200+), Unix `kill(pid, 0)`.\n 5|\n 6|// THE IDENTITY CORE LIVES IN `spt-procident` AND IS RE-EXPORTED VERBATIM.\n...\n 12|// every call site, every doc link and every test in this crate is untouched and\n*13|// `spt_store::proc::process_identity` still resolves.\n 14|// [impl->REQ-POOL-ONE-LIVENESS-PREDICATE]\n 15|pub use spt_procident::{\n*16|    is_process_alive, parent_pid_of, pin_process, process_exists, process_identity,\n 17|    process_started_at, process_table, PinnedProc, ProcIdentity,\n 18|};\n 19|\n...\n 47|    if pid == 0 {\n*48|        return; // never a real per-process target (see is_process_alive)\n 49|    }\n 50|    #[cfg(unix)]\n 51|    unsafe {\n...\n 281|/// sweep, a corpse check on a process this daemon spawned. Those are exactly the\n*282|/// sites where [`is_process_alive`] answers a DIFFERENT question — on Windows it\n 283|/// reports that a process OBJECT is openable, which stays true after termination\n 284|/// for as long as any handle survives, so a correctly-reaped child reads alive.\n 285|///\n...\n 292|///   [`process_table`] keeps it), and a zombie runs nothing — dead for every\n*293|///   consumer of this question. So the present arm asks [`is_process_alive`],\n 294|///   whose `Z`-state read already enforces exactly that demotion. Windows has\n 295|///   no zombie concept (a Toolhelp snapshot excludes terminated processes), so\n 296|///   presence there stays authoritative on its own.\n 297|/// - an EMPTY snapshot is **no knowledge, never proof** (7.50's other half), so\n*298|///   it falls through to [`is_process_alive`] rather than manufacturing a\n 299|///   verdict. That fallback is sound precisely where it is reached: the handle\n 300|///   problem is Windows-only, Windows always has a snapshot, and a unix without\n 301|///   `/proc` answers `kill(pid, 0)`, which no handle can corrupt.\n...\n 304|/// hazard is `OpenProcess`, i.e. Windows-only, and this arm consults\n*305|/// [`is_process_alive`] only on unix, where the probe is `kill(pid, 0)` plus the\n 306|/// `/proc` state read — both handle-independent. On Windows the arm answers\n 307|/// `false` from the table alone and never asks a handle anything.\n 308|///\n*309|/// Not a replacement for [`is_process_alive`], which remains the better probe\n 310|/// where no handle is retained (see its doc, and 7.50's audit-don't-mass-migrate\n 311|/// rule). This is the probe for the OWNED-process case, named after its question.\n 312|// [impl->REQ-LIVENESS-ORACLE-SOUND]\n...\n 323|            {\n*324|                !is_process_alive(pid)\n 325|            }\n 326|            #[cfg(windows)]\n 327|            {\n...\n 330|        }\n*331|        None => !is_process_alive(pid),\n 332|    }\n 333|}\n 334|\n...\n 453|/// (`PROCESS_QUERY_LIMITED_INFORMATION`) handle — the non-hanging path, mirroring\n*454|/// [`is_process_alive`] (tasklist/WMIC hang on recent builds). Linux: the\n 455|/// `/proc/<pid>/exe` symlink target. Other unix: `None` (no `/proc`).\n 456|// [impl->REQ-MANIFEST-8]\n 457|pub fn exe_basename(pid: u32) -> Option<String> {\n...\n 473|    if pid == 0 {\n*474|        return None; // never a probe-able process (see is_process_alive)\n 475|    }\n 476|    #[cfg(windows)]\n 477|    {\n...\n 526|    fn current_process_is_alive() {\n*527|        assert!(is_process_alive(std::process::id()));\n 528|    }\n 529|\n 530|    // [unit->REQ-HAZARD-REGISTRY-STALE-CLEAN] a pid far above any real\n...\n 538|        assert!(\n*539|            !is_process_alive(DEAD_PID),\n 540|            \"unassigned pid must read as dead\"\n 541|        );\n 542|    }\n...\n 549|    fn pid_zero_is_dead_on_every_os() {\n*550|        assert!(!is_process_alive(0), \"pid 0 is never a probe-able process\");\n 551|    }\n 552|\n 553|    // [unit->REQ-HAZARD-REGISTRY-STALE-CLEAN] a zombie reads DEAD: an exited\n...\n 564|        for _ in 0..150 {\n*565|            if !is_process_alive(pid) {\n 566|                return; // exited (zombie or reaped-by-os) reads dead\n 567|            }\n 568|            std::thread::sleep(std::time::Duration::from_millis(20));\n...\n 579|        assert_ne!(ppid, std::process::id(), \"parent must differ from self\");\n*580|        assert!(is_process_alive(ppid), \"the live parent must read as alive\");\n 581|    }\n 582|\n 583|    // [unit->REQ-MANIFEST-8] the bind-time match-key: a live pid resolves to its\n...\n 698|    /// Unix has no equivalent (`kill(pid,0)` fails once reaped, and Linux\n*699|    /// additionally demotes a `/proc` zombie to dead — see `is_process_alive`).\n 700|    /// Manufacturing it on Linux is not possible, so this asserts nothing there\n 701|    /// rather than pretending to cover it.\n 702|    #[cfg(windows)]\n...\n 716|        assert!(\n*717|            is_process_alive(pid),\n 718|            \"THE DEFECT, reproduced: a TERMINATED process still reads ALIVE to a \\\n 719|             caller holding its handle. This is what made a reaped root fall off \\\n 720|             zombie_verdict's PRIMARY dead-root arm and be claimed LIVE forever.\"\n...\n 757|            // the product code composes rather than by re-parsing /proc here.\n*758|            if process_exists(pid) == Some(true) && !is_process_alive(pid) {\n 759|                is_zombie = true;\n 760|                break;\n 761|            }\n...\n 807|        for _ in 0..150 {\n*808|            if process_exists(pid) == Some(true) && !is_process_alive(pid) {\n 809|                is_zombie = true;\n 810|                break;\n 811|            }\n...\n 824|        assert_eq!(\n*825|            process_identity(pid),\n 826|            ProcIdentity::Absent,\n 827|            \"THE FIX: a zombie is not a Present identity. Present(birth) here \\\n 828|             matches the record and reads HELD over a dead process\"\n...\n 948|        let leaf = *descendants.last().unwrap();\n*949|        assert!(is_process_alive(leaf), \"the descendant starts alive\");\n 950|\n 951|        kill_pid_tree(ppid);\n 952|        let _ = parent.wait();\n...\n 957|        for _ in 0..150 {\n*958|            if !is_process_alive(leaf) {\n 959|                leaf_dead = true;\n 960|                break;\n 961|            }"}],"details":{"scopePath":".worktrees/287-shellrestore/crates/spt-store/src/proc.rs","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\crates\\spt-store\\src\\proc.rs","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":25,"fileCount":1,"files":[".worktrees/287-shellrestore/crates/spt-store/src/proc.rs"],"fileMatches":[{"path":".worktrees/287-shellrestore/crates/spt-store/src/proc.rs","count":25}],"truncated":false,"displayContent":"   2│//! (dead-PID rows) before resolution. Ported from sister `types.rs`\n  *3│//! `is_process_alive` (ADR-0001): Windows `OpenProcess` (tasklist hangs on\n   4│//! Win11 26200+), Unix `kill(pid, 0)`.\n   5│\n   6│// THE IDENTITY CORE LIVES IN `spt-procident` AND IS RE-EXPORTED VERBATIM.\n    │...\n  12│// every call site, every doc link and every test in this crate is untouched and\n *13│// `spt_store::proc::process_identity` still resolves.\n  14│// [impl->REQ-POOL-ONE-LIVENESS-PREDICATE]\n  15│pub use spt_procident::{\n *16│    is_process_alive, parent_pid_of, pin_process, process_exists, process_identity,\n  17│    process_started_at, process_table, PinnedProc, ProcIdentity,\n  18│};\n  19│\n    │...\n  47│    if pid == 0 {\n *48│        return; // never a real per-process target (see is_process_alive)\n  49│    }\n  50│    #[cfg(unix)]\n  51│    unsafe {\n    │...\n 281│/// sweep, a corpse check on a process this daemon spawned. Those are exactly the\n*282│/// sites where [`is_process_alive`] answers a DIFFERENT question — on Windows it\n 283│/// reports that a process OBJECT is openable, which stays true after termination\n 284│/// for as long as any handle survives, so a correctly-reaped child reads alive.\n 285│///\n    │...\n 292│///   [`process_table`] keeps it), and a zombie runs nothing — dead for every\n*293│///   consumer of this question. So the present arm asks [`is_process_alive`],\n 294│///   whose `Z`-state read already enforces exactly that demotion. Windows has\n 295│///   no zombie concept (a Toolhelp snapshot excludes terminated processes), so\n 296│///   presence there stays authoritative on its own.\n 297│/// - an EMPTY snapshot is **no knowledge, never proof** (7.50's other half), so\n*298│///   it falls through to [`is_process_alive`] rather than manufacturing a\n 299│///   verdict. That fallback is sound precisely where it is reached: the handle\n 300│///   problem is Windows-only, Windows always has a snapshot, and a unix without\n 301│///   `/proc` answers `kill(pid, 0)`, which no handle can corrupt.\n    │...\n 304│/// hazard is `OpenProcess`, i.e. Windows-only, and this arm consults\n*305│/// [`is_process_alive`] only on unix, where the probe is `kill(pid, 0)` plus the\n 306│/// `/proc` state read — both handle-independent. On Windows the arm answers\n 307│/// `false` from the table alone and never asks a handle anything.\n 308│///\n*309│/// Not a replacement for [`is_process_alive`], which remains the better probe\n 310│/// where no handle is retained (see its doc, and 7.50's audit-don't-mass-migrate\n 311│/// rule). This is the probe for the OWNED-process case, named after its question.\n 312│// [impl->REQ-LIVENESS-ORACLE-SOUND]\n    │...\n 323│            {\n*324│                !is_process_alive(pid)\n 325│            }\n 326│            #[cfg(windows)]\n 327│            {\n    │...\n 330│        }\n*331│        None => !is_process_alive(pid),\n 332│    }\n 333│}\n 334│\n    │...\n 453│/// (`PROCESS_QUERY_LIMITED_INFORMATION`) handle — the non-hanging path, mirroring\n*454│/// [`is_process_alive`] (tasklist/WMIC hang on recent builds). Linux: the\n 455│/// `/proc/<pid>/exe` symlink target. Other unix: `None` (no `/proc`).\n 456│// [impl->REQ-MANIFEST-8]\n 457│pub fn exe_basename(pid: u32) -> Option<String> {\n    │...\n 473│    if pid == 0 {\n*474│        return None; // never a probe-able process (see is_process_alive)\n 475│    }\n 476│    #[cfg(windows)]\n 477│    {\n    │...\n 526│    fn current_process_is_alive() {\n*527│        assert!(is_process_alive(std::process::id()));\n 528│    }\n 529│\n 530│    // [unit->REQ-HAZARD-REGISTRY-STALE-CLEAN] a pid far above any real\n    │...\n 538│        assert!(\n*539│            !is_process_alive(DEAD_PID),\n 540│            \"unassigned pid must read as dead\"\n 541│        );\n 542│    }\n    │...\n 549│    fn pid_zero_is_dead_on_every_os() {\n*550│        assert!(!is_process_alive(0), \"pid 0 is never a probe-able process\");\n 551│    }\n 552│\n 553│    // [unit->REQ-HAZARD-REGISTRY-STALE-CLEAN] a zombie reads DEAD: an exited\n    │...\n 564│        for _ in 0..150 {\n*565│            if !is_process_alive(pid) {\n 566│                return; // exited (zombie or reaped-by-os) reads dead\n 567│            }\n 568│            std::thread::sleep(std::time::Duration::from_millis(20));\n    │...\n 579│        assert_ne!(ppid, std::process::id(), \"parent must differ from self\");\n*580│        assert!(is_process_alive(ppid), \"the live parent must read as alive\");\n 581│    }\n 582│\n 583│    // [unit->REQ-MANIFEST-8] the bind-time match-key: a live pid resolves to its\n    │...\n 698│    /// Unix has no equivalent (`kill(pid,0)` fails once reaped, and Linux\n*699│    /// additionally demotes a `/proc` zombie to dead — see `is_process_alive`).\n 700│    /// Manufacturing it on Linux is not possible, so this asserts nothing there\n 701│    /// rather than pretending to cover it.\n 702│    #[cfg(windows)]\n    │...\n 716│        assert!(\n*717│            is_process_alive(pid),\n 718│            \"THE DEFECT, reproduced: a TERMINATED process still reads ALIVE to a \\\n 719│             caller holding its handle. This is what made a reaped root fall off \\\n 720│             zombie_verdict's PRIMARY dead-root arm and be claimed LIVE forever.\"\n    │...\n 757│            // the product code composes rather than by re-parsing /proc here.\n*758│            if process_exists(pid) == Some(true) && !is_process_alive(pid) {\n 759│                is_zombie = true;\n 760│                break;\n 761│            }\n    │...\n 807│        for _ in 0..150 {\n*808│            if process_exists(pid) == Some(true) && !is_process_alive(pid) {\n 809│                is_zombie = true;\n 810│                break;\n 811│            }\n    │...\n 824│        assert_eq!(\n*825│            process_identity(pid),\n 826│            ProcIdentity::Absent,\n 827│            \"THE FIX: a zombie is not a Present identity. Present(birth) here \\\n 828│             matches the record and reads HELD over a dead process\"\n    │...\n 948│        let leaf = *descendants.last().unwrap();\n*949│        assert!(is_process_alive(leaf), \"the descendant starts alive\");\n 950│\n 951│        kill_pid_tree(ppid);\n 952│        let _ = parent.wait();\n    │...\n 957│        for _ in 0..150 {\n*958│            if !is_process_alive(leaf) {\n 959│                leaf_dead = true;\n 960│                break;\n 961│            }"},"isError":false,"timestamp":1789001235267}}
{"type":"message","id":"a6e70a0e","parentId":"b515eb17","timestamp":"2026-09-10T00:47:15.269Z","message":{"role":"toolResult","toolCallId":"call_4NIuSZYxiiK8uy8afR1ZF5FB|fc_077f3569365e313d016aa1fe13229487d19159ce9c0e9a1bc2","toolName":"read","content":[{"type":"text","text":"    // to want it back a tick later — mid-deploy the relaunch would run the OLD\n    // binary out of the file being replaced (worse than a failed install), and it\n    // turns quarantining a misbehaving shell into a restart loop. Recovery is\n    // demand-driven instead: `relink` (unblocked by the derived gate) or a `shell\n    // cmd` that wakes. The offline sibling proves the rig would adopt if eligible.\n    #[test]\n    fn reconcile_never_adopts_a_stale_online_instance() {\n        const DEAD_PID: u32 = 2_000_000_000;\n        let tmp = tempfile::tempdir().unwrap();\n        let owlery = tmp.path();\n        seed_owner(owlery, \"doyle\");\n\n        #[cfg(windows)]\n        let (noop, sleeper) = (\"cmd /c exit 0\", \"ping -n 30 127.0.0.1\");\n        #[cfg(unix)]\n        let (noop, sleeper) = (\"true\", \"sleep 30\");\n        let adapters_dir = tmp.path().join(\"adapters\");\n        let src = tmp.path().join(\"srcs\").join(\"mock-wake\");\n        std::fs::create_dir_all(&src).unwrap();\n        std::fs::write(\n            src.join(\"manifest.toml\"),\n            format!(\n                \"[adapter]\\nname = \\\"mock-wake\\\"\\nkind = \\\"shell\\\"\\nversion = \\\"1\\\"\\n\\\n                 min_spt_core_version = \\\"0\\\"\\n\\n[shell]\\nspawn = '{noop}'\\n\\\n                 persistent = false\\nwake_command = '{sleeper}'\\n\"\n            ),\n        )\n        .unwrap();\n        spt_runtime::registry::register(&adapters_dir, &src, 1).unwrap();\n        let registered = spt_runtime::registry::registered(&adapters_dir);\n\n        // The stale-online instance: bound, then its binary was force-killed.\n        let dead = spawn_record(owlery, \"doyle\", \"mock-wake\", None).unwrap();\n        let dead_perch = spt_store::perch::resolve_shell_perch_path_in(owlery, \"doyle\", &dead);\n        let mut info = shellinfo::read_shell_info(&dead_perch).unwrap();\n        info.status = SHELL_STATUS_ONLINE.to_string();\n        shellinfo::write_shell_info(&dead_perch, &info).unwrap();\n        std::fs::write(\n            dead_perch.join(shellinfo::SHELL_PID_FILE),\n            DEAD_PID.to_string(),\n        )\n        .unwrap();\n\n        let set = Arc::new(WakeSet::new());\n        let params = fast_params();\n        reconcile_once(owlery, &registered, &adapters_dir, &set, &params);\n        assert_eq!(\n            set.len(),\n            0,\n            \"a force-killed instance must NOT be adopted — no spontaneous relaunch\"\n        );\n        assert!(\n            !dead_perch.join(WAKER_PID_FILE).exists(),\n            \"no waker child was started for it either\"\n        );\n\n        // Control: a genuinely OFFLINE sibling is still adopted, so the absence\n        // above is the derivation policy, not a broken rig.\n        let off = spawn_record(owlery, \"doyle\", \"mock-wake\", None).unwrap();\n        reconcile_once(owlery, &registered, &adapters_dir, &set, &params);\n        assert_eq!(set.len(), 1, \"the offline sibling {off} still gets a watcher\");\n        set.stop_watcher(owlery, \"doyle\", &off);\n    }\n\n    // [unit->REQ-HAZARD-SHELL-STALE-ONLINE] [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n    // BOTH HALVES IN ONE ROW: for a same-boot corpse the heal FIRED (the record now\n    // reads offline — the honest value leg (a) owes it) AND the watcher count is\n    // still 0. Asserted together on purpose. Split across two rows either half\n    // passes alone on a broken build: without the heal the record still says online\n    // and a status-only read refuses adoption for the WRONG reason (the old\n    // accidental protection, which is exactly what this milestone removed), while a\n    // watcher-0 assertion on its own is satisfied by an implementation that adopts\n    // nothing at all.\n    //\n    // The pre-boot sibling is the contrast that keeps the discriminant from being\n    // vacuous: an implementation treating EVERY corpse as ineligible passes the\n    // same-boot half and silently strands every restart casualty's watcher. Both\n    // instances are the identical shape — record online over a corpse with a parked\n    // launch stamp — and differ ONLY in which side of the boot instant it falls on.\n    #[test]\n    fn nonpersistent_same_boot_freeze_does_not_apply_to_persistent_watchers() {\n        const DEAD_PID: u32 = 2_000_000_000;\n        let tmp = tempfile::tempdir().unwrap();\n        let owlery = tmp.path();\n        seed_owner(owlery, \"doyle\");\n        let boot = spt_store::proc::boot_instant_ms().expect(\"a boot instant on this platform\");\n\n        #[cfg(windows)]\n        let (noop, sleeper) = (\"cmd /c exit 0\", \"ping -n 30 127.0.0.1\");\n        #[cfg(unix)]\n        let (noop, sleeper) = (\"true\", \"sleep 30\");\n        let adapters_dir = tmp.path().join(\"adapters\");\n        let src = tmp.path().join(\"srcs\").join(\"mock-wake\");\n        std::fs::create_dir_all(&src).unwrap();\n        std::fs::write(\n            src.join(\"manifest.toml\"),\n            format!(\n                \"[adapter]\\nname = \\\"mock-wake\\\"\\nkind = \\\"shell\\\"\\nversion = \\\"1\\\"\\n\\\n                 min_spt_core_version = \\\"0\\\"\\n\\n[shell]\\nspawn = '{noop}'\\n\\\n                 persistent = false\\nwake_command = '{sleeper}'\\n\"\n            ),\n        )\n        .unwrap();\n        spt_runtime::registry::register(&adapters_dir, &src, 1).unwrap();\n        let registered = spt_runtime::registry::registered(&adapters_dir);\n\n        let perch_of =\n            |id: &str| spt_store::perch::resolve_shell_perch_path_in(owlery, \"doyle\", id);\n        let mk = |launched_ms: u64| -> String {\n            let id = spawn_record(owlery, \"doyle\", \"mock-wake\", None).unwrap();\n            let perch = perch_of(&id);\n            let mut i = shellinfo::read_shell_info(&perch).unwrap();\n            i.status = SHELL_STATUS_ONLINE.to_string();\n            shellinfo::write_shell_info(&perch, &i).unwrap();\n            std::fs::write(perch.join(shellinfo::SHELL_PID_FILE), DEAD_PID.to_string()).unwrap();\n            let launch = shellinfo::ShellLaunch {\n                pid_started_at: Some(1),\n                launched_ms,\n            };\n            std::fs::write(\n                perch.join(shellinfo::SHELL_LAUNCH_FILE),\n                serde_json::to_string(&launch).unwrap(),\n            )\n            .unwrap();\n            id\n        };\n        let force_killed = mk(boot + 120_000); // killed 2 min INTO this boot\n        let casualty = mk(boot.saturating_sub(600_000)); // launched 10 min BEFORE it\n\n        let set = Arc::new(WakeSet::new());\n        let params = fast_params();\n        reconcile_once(owlery, &registered, &adapters_dir, &set, &params);\n\n        assert_eq!(\n            shellinfo::read_shell_info(&perch_of(&force_killed))\n                .unwrap()\n                .status,\n            SHELL_STATUS_OFFLINE,\n            \"the heal FIRED — the record stopped lying about a corpse\"\n        );\n        assert!(\n            !set.contains(\"doyle\", &force_killed),\n            \"...and the honest record still did not buy it a watcher: killed THIS \\\n             boot is the mid-deploy case the class-(c) freeze protects\"\n        );\n        assert!(\n            !perch_of(&force_killed).join(WAKER_PID_FILE).exists(),\n            \"no waker child was started for it either\"\n        );\n        assert!(\n            set.contains(\"doyle\", &casualty),\n            \"a corpse launched BEFORE this boot is a restart casualty — eligible, \\\n             per #78 leg (b)'s greenlit scope\"\n        );\n        assert_eq!(set.len(), 1, \"and it is the ONLY instance adopted\");\n        set.stop_watcher(owlery, \"doyle\", &casualty);\n\n        // The SAME corpse is eligible when the manifest declares persistence;\n        // no restart or new launch stamp is needed to release that freeze.\n        let mut persistent_registered = registered.clone();\n        for (_, manifest) in &mut persistent_registered {\n            manifest.shell.as_mut().unwrap().persistent = true;\n        }\n        std::fs::remove_file(perch_of(&force_killed).join(shellinfo::SHELL_LAUNCH_FILE)).unwrap();\n        let healed = shellinfo::read_shell_info(&perch_of(&force_killed)).unwrap();\n        assert!(\n            watcher_eligible(&perch_of(&force_killed), &healed, true, None),\n            \"persistent watcher eligibility needs neither launch stamp nor boot oracle\"\n        );\n        reconcile_once(owlery, &persistent_registered, &adapters_dir, &set, &params);\n        assert!(set.contains(\"doyle\", &force_killed));\n        set.stop_watcher(owlery, \"doyle\", &force_killed);\n        set.stop_watcher(owlery, \"doyle\", &casualty);\n    }\n\n    // [unit->REQ-SHELL-2] the reconciler holds the mutual exclusivity: an\n    // offline instance with a wake_command gets exactly ONE watcher (a second\n    // sweep never doubles it); a give-up latch suppresses adoption; onlining\n    // the instance stops the watcher on the next sweep.\n    #[test]\n    fn reconcile_flips_watchers_with_instance_state() {\n        let tmp = tempfile::tempdir().unwrap();\n        let owlery = tmp.path();\n        seed_owner(owlery, \"doyle\");\n\n        // A registered shell adapter whose waker sleeps (stays alive between\n        // sweeps) — the long-running watcher case.\n        #[cfg(windows)]\n        let (noop, sleeper) = (\"cmd /c exit 0\", \"ping -n 30 127.0.0.1\");\n        #[cfg(unix)]\n        let (noop, sleeper) = (\"true\", \"sleep 30\");\n        let adapters_dir = tmp.path().join(\"adapters\");\n        let src = tmp.path().join(\"srcs\").join(\"mock-wake\");\n        std::fs::create_dir_all(&src).unwrap();\n        std::fs::write(\n…\n        )\n        .unwrap();\n…\n    }\n\n[Showing lines 1659-1853 of 2161. Use :1854 to continue]"}],"details":{"truncation":{"content":"    // to want it back a tick later — mid-deploy the relaunch would run the OLD\r\n    // binary out of the file being replaced (worse than a failed install), and it\r\n    // turns quarantining a misbehaving shell into a restart loop. Recovery is\r\n    // demand-driven instead: `relink` (unblocked by the derived gate) or a `shell\r\n    // cmd` that wakes. The offline sibling proves the rig would adopt if eligible.\r\n    #[test]\r\n    fn reconcile_never_adopts_a_stale_online_instance() {\r\n        const DEAD_PID: u32 = 2_000_000_000;\r\n        let tmp = tempfile::tempdir().unwrap();\r\n        let owlery = tmp.path();\r\n        seed_owner(owlery, \"doyle\");\r\n\r\n        #[cfg(windows)]\r\n        let (noop, sleeper) = (\"cmd /c exit 0\", \"ping -n 30 127.0.0.1\");\r\n        #[cfg(unix)]\r\n        let (noop, sleeper) = (\"true\", \"sleep 30\");\r\n        let adapters_dir = tmp.path().join(\"adapters\");\r\n        let src = tmp.path().join(\"srcs\").join(\"mock-wake\");\r\n        std::fs::create_dir_all(&src).unwrap();\r\n        std::fs::write(\r\n            src.join(\"manifest.toml\"),\r\n            format!(\r\n                \"[adapter]\\nname = \\\"mock-wake\\\"\\nkind = \\\"shell\\\"\\nversion = \\\"1\\\"\\n\\\r\n                 min_spt_core_version = \\\"0\\\"\\n\\n[shell]\\nspawn = '{noop}'\\n\\\r\n                 persistent = false\\nwake_command = '{sleeper}'\\n\"\r\n            ),\r\n        )\r\n        .unwrap();\r\n        spt_runtime::registry::register(&adapters_dir, &src, 1).unwrap();\r\n        let registered = spt_runtime::registry::registered(&adapters_dir);\r\n\r\n        // The stale-online instance: bound, then its binary was force-killed.\r\n        let dead = spawn_record(owlery, \"doyle\", \"mock-wake\", None).unwrap();\r\n        let dead_perch = spt_store::perch::resolve_shell_perch_path_in(owlery, \"doyle\", &dead);\r\n        let mut info = shellinfo::read_shell_info(&dead_perch).unwrap();\r\n        info.status = SHELL_STATUS_ONLINE.to_string();\r\n        shellinfo::write_shell_info(&dead_perch, &info).unwrap();\r\n        std::fs::write(\r\n            dead_perch.join(shellinfo::SHELL_PID_FILE),\r\n            DEAD_PID.to_string(),\r\n        )\r\n        .unwrap();\r\n\r\n        let set = Arc::new(WakeSet::new());\r\n        let params = fast_params();\r\n        reconcile_once(owlery, &registered, &adapters_dir, &set, &params);\r\n        assert_eq!(\r\n            set.len(),\r\n            0,\r\n            \"a force-killed instance must NOT be adopted — no spontaneous relaunch\"\r\n        );\r\n        assert!(\r\n            !dead_perch.join(WAKER_PID_FILE).exists(),\r\n            \"no waker child was started for it either\"\r\n        );\r\n\r\n        // Control: a genuinely OFFLINE sibling is still adopted, so the absence\r\n        // above is the derivation policy, not a broken rig.\r\n        let off = spawn_record(owlery, \"doyle\", \"mock-wake\", None).unwrap();\r\n        reconcile_once(owlery, &registered, &adapters_dir, &set, &params);\r\n        assert_eq!(set.len(), 1, \"the offline sibling {off} still gets a watcher\");\r\n        set.stop_watcher(owlery, \"doyle\", &off);\r\n    }\r\n\r\n    // [unit->REQ-HAZARD-SHELL-STALE-ONLINE] [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\n    // BOTH HALVES IN ONE ROW: for a same-boot corpse the heal FIRED (the record now\r\n    // reads offline — the honest value leg (a) owes it) AND the watcher count is\r\n    // still 0. Asserted together on purpose. Split across two rows either half\r\n    // passes alone on a broken build: without the heal the record still says online\r\n    // and a status-only read refuses adoption for the WRONG reason (the old\r\n    // accidental protection, which is exactly what this milestone removed), while a\r\n    // watcher-0 assertion on its own is satisfied by an implementation that adopts\r\n    // nothing at all.\r\n    //\r\n    // The pre-boot sibling is the contrast that keeps the discriminant from being\r\n    // vacuous: an implementation treating EVERY corpse as ineligible passes the\r\n    // same-boot half and silently strands every restart casualty's watcher. Both\r\n    // instances are the identical shape — record online over a corpse with a parked\r\n    // launch stamp — and differ ONLY in which side of the boot instant it falls on.\r\n    #[test]\r\n    fn nonpersistent_same_boot_freeze_does_not_apply_to_persistent_watchers() {\r\n        const DEAD_PID: u32 = 2_000_000_000;\r\n        let tmp = tempfile::tempdir().unwrap();\r\n        let owlery = tmp.path();\r\n        seed_owner(owlery, \"doyle\");\r\n        let boot = spt_store::proc::boot_instant_ms().expect(\"a boot instant on this platform\");\r\n\r\n        #[cfg(windows)]\r\n        let (noop, sleeper) = (\"cmd /c exit 0\", \"ping -n 30 127.0.0.1\");\r\n        #[cfg(unix)]\r\n        let (noop, sleeper) = (\"true\", \"sleep 30\");\r\n        let adapters_dir = tmp.path().join(\"adapters\");\r\n        let src = tmp.path().join(\"srcs\").join(\"mock-wake\");\r\n        std::fs::create_dir_all(&src).unwrap();\r\n        std::fs::write(\r\n            src.join(\"manifest.toml\"),\r\n            format!(\r\n                \"[adapter]\\nname = \\\"mock-wake\\\"\\nkind = \\\"shell\\\"\\nversion = \\\"1\\\"\\n\\\r\n                 min_spt_core_version = \\\"0\\\"\\n\\n[shell]\\nspawn = '{noop}'\\n\\\r\n                 persistent = false\\nwake_command = '{sleeper}'\\n\"\r\n            ),\r\n        )\r\n        .unwrap();\r\n        spt_runtime::registry::register(&adapters_dir, &src, 1).unwrap();\r\n        let registered = spt_runtime::registry::registered(&adapters_dir);\r\n\r\n        let perch_of =\r\n            |id: &str| spt_store::perch::resolve_shell_perch_path_in(owlery, \"doyle\", id);\r\n        let mk = |launched_ms: u64| -> String {\r\n            let id = spawn_record(owlery, \"doyle\", \"mock-wake\", None).unwrap();\r\n            let perch = perch_of(&id);\r\n            let mut i = shellinfo::read_shell_info(&perch).unwrap();\r\n            i.status = SHELL_STATUS_ONLINE.to_string();\r\n            shellinfo::write_shell_info(&perch, &i).unwrap();\r\n            std::fs::write(perch.join(shellinfo::SHELL_PID_FILE), DEAD_PID.to_string()).unwrap();\r\n            let launch = shellinfo::ShellLaunch {\r\n                pid_started_at: Some(1),\r\n                launched_ms,\r\n            };\r\n            std::fs::write(\r\n                perch.join(shellinfo::SHELL_LAUNCH_FILE),\r\n                serde_json::to_string(&launch).unwrap(),\r\n            )\r\n            .unwrap();\r\n            id\r\n        };\r\n        let force_killed = mk(boot + 120_000); // killed 2 min INTO this boot\r\n        let casualty = mk(boot.saturating_sub(600_000)); // launched 10 min BEFORE it\r\n\r\n        let set = Arc::new(WakeSet::new());\r\n        let params = fast_params();\r\n        reconcile_once(owlery, &registered, &adapters_dir, &set, &params);\r\n\r\n        assert_eq!(\r\n            shellinfo::read_shell_info(&perch_of(&force_killed))\r\n                .unwrap()\r\n                .status,\r\n            SHELL_STATUS_OFFLINE,\r\n            \"the heal FIRED — the record stopped lying about a corpse\"\r\n        );\r\n        assert!(\r\n            !set.contains(\"doyle\", &force_killed),\r\n            \"...and the honest record still did not buy it a watcher: killed THIS \\\r\n             boot is the mid-deploy case the class-(c) freeze protects\"\r\n        );\r\n        assert!(\r\n            !perch_of(&force_killed).join(WAKER_PID_FILE).exists(),\r\n            \"no waker child was started for it either\"\r\n        );\r\n        assert!(\r\n            set.contains(\"doyle\", &casualty),\r\n            \"a corpse launched BEFORE this boot is a restart casualty — eligible, \\\r\n             per #78 leg (b)'s greenlit scope\"\r\n        );\r\n        assert_eq!(set.len(), 1, \"and it is the ONLY instance adopted\");\r\n        set.stop_watcher(owlery, \"doyle\", &casualty);\r\n\r\n        // The SAME corpse is eligible when the manifest declares persistence;\r\n        // no restart or new launch stamp is needed to release that freeze.\r\n        let mut persistent_registered = registered.clone();\r\n        for (_, manifest) in &mut persistent_registered {\r\n            manifest.shell.as_mut().unwrap().persistent = true;\r\n        }\r\n        std::fs::remove_file(perch_of(&force_killed).join(shellinfo::SHELL_LAUNCH_FILE)).unwrap();\r\n        let healed = shellinfo::read_shell_info(&perch_of(&force_killed)).unwrap();\r\n        assert!(\r\n            watcher_eligible(&perch_of(&force_killed), &healed, true, None),\r\n            \"persistent watcher eligibility needs neither launch stamp nor boot oracle\"\r\n        );\r\n        reconcile_once(owlery, &persistent_registered, &adapters_dir, &set, &params);\r\n        assert!(set.contains(\"doyle\", &force_killed));\r\n        set.stop_watcher(owlery, \"doyle\", &force_killed);\r\n        set.stop_watcher(owlery, \"doyle\", &casualty);\r\n    }\r\n\r\n    // [unit->REQ-SHELL-2] the reconciler holds the mutual exclusivity: an\r\n    // offline instance with a wake_command gets exactly ONE watcher (a second\r\n    // sweep never doubles it); a give-up latch suppresses adoption; onlining\r\n    // the instance stops the watcher on the next sweep.\r\n    #[test]\r\n    fn reconcile_flips_watchers_with_instance_state() {\r\n        let tmp = tempfile::tempdir().unwrap();\r\n        let owlery = tmp.path();\r\n        seed_owner(owlery, \"doyle\");\r\n\r\n        // A registered shell adapter whose waker sleeps (stays alive between\r\n        // sweeps) — the long-running watcher case.\r\n        #[cfg(windows)]\r\n        let (noop, sleeper) = (\"cmd /c exit 0\", \"ping -n 30 127.0.0.1\");\r\n        #[cfg(unix)]\r\n        let (noop, sleeper) = (\"true\", \"sleep 30\");\r\n        let adapters_dir = tmp.path().join(\"adapters\");\r\n        let src = tmp.path().join(\"srcs\").join(\"mock-wake\");\r\n        std::fs::create_dir_all(&src).unwrap();\r\n        std::fs::write(\r","truncated":true,"truncatedBy":"lines","totalLines":503,"totalBytes":9626,"outputLines":195,"outputBytes":9626,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":2161,"displayContent":{"text":"    // to want it back a tick later — mid-deploy the relaunch would run the OLD\r\n    // binary out of the file being replaced (worse than a failed install), and it\r\n    // turns quarantining a misbehaving shell into a restart loop. Recovery is\r\n    // demand-driven instead: `relink` (unblocked by the derived gate) or a `shell\r\n    // cmd` that wakes. The offline sibling proves the rig would adopt if eligible.\r\n    #[test]\r\n    fn reconcile_never_adopts_a_stale_online_instance() {\r\n        const DEAD_PID: u32 = 2_000_000_000;\r\n        let tmp = tempfile::tempdir().unwrap();\r\n        let owlery = tmp.path();\r\n        seed_owner(owlery, \"doyle\");\r\n\r\n        #[cfg(windows)]\r\n        let (noop, sleeper) = (\"cmd /c exit 0\", \"ping -n 30 127.0.0.1\");\r\n        #[cfg(unix)]\r\n        let (noop, sleeper) = (\"true\", \"sleep 30\");\r\n        let adapters_dir = tmp.path().join(\"adapters\");\r\n        let src = tmp.path().join(\"srcs\").join(\"mock-wake\");\r\n        std::fs::create_dir_all(&src).unwrap();\r\n        std::fs::write(\r\n            src.join(\"manifest.toml\"),\r\n            format!(\r\n                \"[adapter]\\nname = \\\"mock-wake\\\"\\nkind = \\\"shell\\\"\\nversion = \\\"1\\\"\\n\\\r\n                 min_spt_core_version = \\\"0\\\"\\n\\n[shell]\\nspawn = '{noop}'\\n\\\r\n                 persistent = false\\nwake_command = '{sleeper}'\\n\"\r\n            ),\r\n        )\r\n        .unwrap();\r\n        spt_runtime::registry::register(&adapters_dir, &src, 1).unwrap();\r\n        let registered = spt_runtime::registry::registered(&adapters_dir);\r\n\r\n        // The stale-online instance: bound, then its binary was force-killed.\r\n        let dead = spawn_record(owlery, \"doyle\", \"mock-wake\", None).unwrap();\r\n        let dead_perch = spt_store::perch::resolve_shell_perch_path_in(owlery, \"doyle\", &dead);\r\n        let mut info = shellinfo::read_shell_info(&dead_perch).unwrap();\r\n        info.status = SHELL_STATUS_ONLINE.to_string();\r\n        shellinfo::write_shell_info(&dead_perch, &info).unwrap();\r\n        std::fs::write(\r\n            dead_perch.join(shellinfo::SHELL_PID_FILE),\r\n            DEAD_PID.to_string(),\r\n        )\r\n        .unwrap();\r\n\r\n        let set = Arc::new(WakeSet::new());\r\n        let params = fast_params();\r\n        reconcile_once(owlery, &registered, &adapters_dir, &set, &params);\r\n        assert_eq!(\r\n            set.len(),\r\n            0,\r\n            \"a force-killed instance must NOT be adopted — no spontaneous relaunch\"\r\n        );\r\n        assert!(\r\n            !dead_perch.join(WAKER_PID_FILE).exists(),\r\n            \"no waker child was started for it either\"\r\n        );\r\n\r\n        // Control: a genuinely OFFLINE sibling is still adopted, so the absence\r\n        // above is the derivation policy, not a broken rig.\r\n        let off = spawn_record(owlery, \"doyle\", \"mock-wake\", None).unwrap();\r\n        reconcile_once(owlery, &registered, &adapters_dir, &set, &params);\r\n        assert_eq!(set.len(), 1, \"the offline sibling {off} still gets a watcher\");\r\n        set.stop_watcher(owlery, \"doyle\", &off);\r\n    }\r\n\r\n    // [unit->REQ-HAZARD-SHELL-STALE-ONLINE] [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\n    // BOTH HALVES IN ONE ROW: for a same-boot corpse the heal FIRED (the record now\r\n    // reads offline — the honest value leg (a) owes it) AND the watcher count is\r\n    // still 0. Asserted together on purpose. Split across two rows either half\r\n    // passes alone on a broken build: without the heal the record still says online\r\n    // and a status-only read refuses adoption for the WRONG reason (the old\r\n    // accidental protection, which is exactly what this milestone removed), while a\r\n    // watcher-0 assertion on its own is satisfied by an implementation that adopts\r\n    // nothing at all.\r\n    //\r\n    // The pre-boot sibling is the contrast that keeps the discriminant from being\r\n    // vacuous: an implementation treating EVERY corpse as ineligible passes the\r\n    // same-boot half and silently strands every restart casualty's watcher. Both\r\n    // instances are the identical shape — record online over a corpse with a parked\r\n    // launch stamp — and differ ONLY in which side of the boot instant it falls on.\r\n    #[test]\r\n    fn nonpersistent_same_boot_freeze_does_not_apply_to_persistent_watchers() {\r\n        const DEAD_PID: u32 = 2_000_000_000;\r\n        let tmp = tempfile::tempdir().unwrap();\r\n        let owlery = tmp.path();\r\n        seed_owner(owlery, \"doyle\");\r\n        let boot = spt_store::proc::boot_instant_ms().expect(\"a boot instant on this platform\");\r\n\r\n        #[cfg(windows)]\r\n        let (noop, sleeper) = (\"cmd /c exit 0\", \"ping -n 30 127.0.0.1\");\r\n        #[cfg(unix)]\r\n        let (noop, sleeper) = (\"true\", \"sleep 30\");\r\n        let adapters_dir = tmp.path().join(\"adapters\");\r\n        let src = tmp.path().join(\"srcs\").join(\"mock-wake\");\r\n        std::fs::create_dir_all(&src).unwrap();\r\n        std::fs::write(\r\n            src.join(\"manifest.toml\"),\r\n            format!(\r\n                \"[adapter]\\nname = \\\"mock-wake\\\"\\nkind = \\\"shell\\\"\\nversion = \\\"1\\\"\\n\\\r\n                 min_spt_core_version = \\\"0\\\"\\n\\n[shell]\\nspawn = '{noop}'\\n\\\r\n                 persistent = false\\nwake_command = '{sleeper}'\\n\"\r\n            ),\r\n        )\r\n        .unwrap();\r\n        spt_runtime::registry::register(&adapters_dir, &src, 1).unwrap();\r\n        let registered = spt_runtime::registry::registered(&adapters_dir);\r\n\r\n        let perch_of =\r\n            |id: &str| spt_store::perch::resolve_shell_perch_path_in(owlery, \"doyle\", id);\r\n        let mk = |launched_ms: u64| -> String {\r\n            let id = spawn_record(owlery, \"doyle\", \"mock-wake\", None).unwrap();\r\n            let perch = perch_of(&id);\r\n            let mut i = shellinfo::read_shell_info(&perch).unwrap();\r\n            i.status = SHELL_STATUS_ONLINE.to_string();\r\n            shellinfo::write_shell_info(&perch, &i).unwrap();\r\n            std::fs::write(perch.join(shellinfo::SHELL_PID_FILE), DEAD_PID.to_string()).unwrap();\r\n            let launch = shellinfo::ShellLaunch {\r\n                pid_started_at: Some(1),\r\n                launched_ms,\r\n            };\r\n            std::fs::write(\r\n                perch.join(shellinfo::SHELL_LAUNCH_FILE),\r\n                serde_json::to_string(&launch).unwrap(),\r\n            )\r\n            .unwrap();\r\n            id\r\n        };\r\n        let force_killed = mk(boot + 120_000); // killed 2 min INTO this boot\r\n        let casualty = mk(boot.saturating_sub(600_000)); // launched 10 min BEFORE it\r\n\r\n        let set = Arc::new(WakeSet::new());\r\n        let params = fast_params();\r\n        reconcile_once(owlery, &registered, &adapters_dir, &set, &params);\r\n\r\n        assert_eq!(\r\n            shellinfo::read_shell_info(&perch_of(&force_killed))\r\n                .unwrap()\r\n                .status,\r\n            SHELL_STATUS_OFFLINE,\r\n            \"the heal FIRED — the record stopped lying about a corpse\"\r\n        );\r\n        assert!(\r\n            !set.contains(\"doyle\", &force_killed),\r\n            \"...and the honest record still did not buy it a watcher: killed THIS \\\r\n             boot is the mid-deploy case the class-(c) freeze protects\"\r\n        );\r\n        assert!(\r\n            !perch_of(&force_killed).join(WAKER_PID_FILE).exists(),\r\n            \"no waker child was started for it either\"\r\n        );\r\n        assert!(\r\n            set.contains(\"doyle\", &casualty),\r\n            \"a corpse launched BEFORE this boot is a restart casualty — eligible, \\\r\n             per #78 leg (b)'s greenlit scope\"\r\n        );\r\n        assert_eq!(set.len(), 1, \"and it is the ONLY instance adopted\");\r\n        set.stop_watcher(owlery, \"doyle\", &casualty);\r\n\r\n        // The SAME corpse is eligible when the manifest declares persistence;\r\n        // no restart or new launch stamp is needed to release that freeze.\r\n        let mut persistent_registered = registered.clone();\r\n        for (_, manifest) in &mut persistent_registered {\r\n            manifest.shell.as_mut().unwrap().persistent = true;\r\n        }\r\n        std::fs::remove_file(perch_of(&force_killed).join(shellinfo::SHELL_LAUNCH_FILE)).unwrap();\r\n        let healed = shellinfo::read_shell_info(&perch_of(&force_killed)).unwrap();\r\n        assert!(\r\n            watcher_eligible(&perch_of(&force_killed), &healed, true, None),\r\n            \"persistent watcher eligibility needs neither launch stamp nor boot oracle\"\r\n        );\r\n        reconcile_once(owlery, &persistent_registered, &adapters_dir, &set, &params);\r\n        assert!(set.contains(\"doyle\", &force_killed));\r\n        set.stop_watcher(owlery, \"doyle\", &force_killed);\r\n        set.stop_watcher(owlery, \"doyle\", &casualty);\r\n    }\r\n\r\n    // [unit->REQ-SHELL-2] the reconciler holds the mutual exclusivity: an\r\n    // offline instance with a wake_command gets exactly ONE watcher (a second\r\n    // sweep never doubles it); a give-up latch suppresses adoption; onlining\r\n    // the instance stops the watcher on the next sweep.\r\n    #[test]\r\n    fn reconcile_flips_watchers_with_instance_state() {\r\n        let tmp = tempfile::tempdir().unwrap();\r\n        let owlery = tmp.path();\r\n        seed_owner(owlery, \"doyle\");\r\n\r\n        // A registered shell adapter whose waker sleeps (stays alive between\r\n        // sweeps) — the long-running watcher case.\r\n        #[cfg(windows)]\r\n        let (noop, sleeper) = (\"cmd /c exit 0\", \"ping -n 30 127.0.0.1\");\r\n        #[cfg(unix)]\r\n        let (noop, sleeper) = (\"true\", \"sleep 30\");\r\n        let adapters_dir = tmp.path().join(\"adapters\");\r\n        let src = tmp.path().join(\"srcs\").join(\"mock-wake\");\r\n        std::fs::create_dir_all(&src).unwrap();\r\n        std::fs::write(\r\n…\n        )\n        .unwrap();\n…\n    }","startLine":1659,"lineNumbers":[1659,1660,1661,1662,1663,1664,1665,1666,1667,1668,1669,1670,1671,1672,1673,1674,1675,1676,1677,1678,1679,1680,1681,1682,1683,1684,1685,1686,1687,1688,1689,1690,1691,1692,1693,1694,1695,1696,1697,1698,1699,1700,1701,1702,1703,1704,1705,1706,1707,1708,1709,1710,1711,1712,1713,1714,1715,1716,1717,1718,1719,1720,1721,1722,1723,1724,1725,1726,1727,1728,1729,1730,1731,1732,1733,1734,1735,1736,1737,1738,1739,1740,1741,1742,1743,1744,1745,1746,1747,1748,1749,1750,1751,1752,1753,1754,1755,1756,1757,1758,1759,1760,1761,1762,1763,1764,1765,1766,1767,1768,1769,1770,1771,1772,1773,1774,1775,1776,1777,1778,1779,1780,1781,1782,1783,1784,1785,1786,1787,1788,1789,1790,1791,1792,1793,1794,1795,1796,1797,1798,1799,1800,1801,1802,1803,1804,1805,1806,1807,1808,1809,1810,1811,1812,1813,1814,1815,1816,1817,1818,1819,1820,1821,1822,1823,1824,1825,1826,1827,1828,1829,1830,1831,1832,1833,1834,1835,1836,1837,1838,1839,1840,1841,1842,1843,1844,1845,1846,1847,1848,1849,1850,1851,1852,1853,null,1860,1861,null,1966]},"fileSize":100821,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\crates\\spt-daemon\\src\\shellwake.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":2161,"totalBytes":9626,"outputLines":195,"outputBytes":9626,"shownRange":{"start":1659,"end":1853},"nextOffset":1854}}},"isError":false,"timestamp":1789001235269}}
{"type":"message","id":"775d8b57","parentId":"a6e70a0e","timestamp":"2026-09-10T00:47:15.271Z","message":{"role":"toolResult","toolCallId":"call_W2T3CZeTyRUlSBI7Onuv67PH|fc_077f3569365e313d016aa1fe1322ac87d1a4e59d56c089b7e8","toolName":"read","content":[{"type":"text","text":"#[derive(Debug, Clone, Copy, PartialEq, Eq)]\npub enum CloseOutcome {\n    /// `ephemeral = true`: perch + spool history erased, mint slot freed.\n    TornDown,\n    /// Everything else: perch kept, status=offline, re-linkable.\n    Offline,\n}\n\n/// The link-break close (M5-D4a, REQ-SHELL-2 — CONTEXT §Shell model\n/// lifecycle). The contract used to read \"the binary never survives\"; since\n/// BAROMETER it reads **a link-break always closes the link and always ends the\n/// binary we can still prove is ours** — a kill that cannot authenticate its\n/// target is refused and announced rather than fired blind, because the\n/// alternative is force-killing a stranger's process tree on a recycled pid:\n///\n/// 1. **pre-close instruction** (manifest `pre_close`, optional): composed as\n///    a [`shell_close` frame](crate::shellchan::compose_close_frame), stamped\n///    with the link key, and delivered per the receipt mode — relay: spooled\n///    (the binary may drain it during the grace window); stdin: spooled +\n///    best-effort broker push. Skipped when nothing is running.\n/// 2. **termination window** (`close_timeout_ms`, absent ⇒ 0): poll the\n///    recorded pid for a graceful exit up to the window.\n/// 3. **force-close**: an **authenticated** kill at window end — the pid+birth\n///    pair must prove the process at the recorded number is the binary we\n///    launched (`REQ-SHELL-KILL-AUTHENTICATED`). It was unconditional until\n///    BAROMETER, and that is precisely what let a recycled pid route to\n///    `taskkill /T`. A refused kill is announced (`SHELL_KILL_REFUSED_*`) and\n///    the close continues; the pid file and the **link token die with the\n///    link** either way (a stale token must never drive a relinked instance —\n///    relink mints fresh).\n/// 4. **divergence**: `ephemeral` ⇒ full teardown + history erase (the spool\n///    lives in the perch dir — removal IS the erase); otherwise the perch\n///    stays, status=offline, re-linkable.\n///\n/// `shell = None` (adapter deregistered after spawn) closes with defaults: no\n/// pre-close, no grace, **non-ephemeral** — history is never erased without\n/// the manifest's authority.\n// [impl->REQ-SHELL-2]\npub fn close_shell(\n    owlery: &Path,\n    owner: &str,\n    shell_id: &str,\n    shell: Option<&Shell>,\n) -> Result<CloseOutcome, String> {\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n    let pid = std::fs::read_to_string(perch.join(SHELL_PID_FILE))\n        .ok()\n        .and_then(|s| s.trim().parse::<u32>().ok())\n        .filter(|&p| p != 0);\n    // Reap first: an in-process spawner's exited child is a zombie until\n    // somebody waits — without this the liveness probe would hold a dead\n    // binary \"running\" through the whole grace window.\n    if let Some(p) = pid {\n        spt_store::proc::reap_if_child(p);\n    }\n    let running = pid.is_some_and(spt_store::proc::is_process_alive);\n\n    // 1. The pre-close instruction — only meaningful while something runs.\n    if running {\n        if let Some(instruction) = shell.and_then(|s| s.pre_close.as_deref()) {\n            if let Some(token) = read_link_token(&perch) {\n                let key = link_key(&token);\n                let frame = crate::shellchan::compose_close_frame(owner, instruction);\n                let stamped = crate::shellchan::stamp_frame(&key, &frame);\n                // Best-effort by design: a wedged spool/broker must never\n                // block the close — the force-kill below is the guarantee.\n                let _ = crate::shellchan::spool_shell_frame(owlery, owner, shell_id, &stamped);\n                if shell.map(|s| s.command_receipt) == Some(Some(CommandReceipt::Stdin)) {\n                    let _ = crate::shellchan::deliver_stdin_pending(owlery, owner, shell_id);\n                }\n            }\n        }\n        // 2. The termination window: a graceful exit beats the force-kill.\n        let window =\n            std::time::Duration::from_millis(shell.and_then(|s| s.close_timeout_ms).unwrap_or(0));\n        let deadline = std::time::Instant::now() + window;\n        while std::time::Instant::now() < deadline {\n            if !pid.is_some_and(spt_store::proc::is_process_alive) {\n                break;\n            }\n            std::thread::sleep(std::time::Duration::from_millis(50));\n        }\n    }\n\n    // 3. Force-close (unconditional — the binary never survives a\n    // link-break) + retire the link credentials. A watcher should not exist\n    // while the binary side was up, but a torn-down instance must orphan\n    // nothing — kill both unconditionally.\n    kill_shell_at(&perch);\n    crate::shellwake::kill_waker_at(&perch);\n    if let Some(p) = pid {\n        // The kill is asynchronous (and leaves a zombie when we are the\n        // spawner) — confirm the death bounded, and SAY SO when it does not\n        // come. Since the kill authenticates, this loop can now legitimately\n        // end with the process still there (a refused kill), so a silent\n        // return would trade a loud wrong action for a quiet wrong state —\n        // the un-RCA-able one. It reports what it could not do instead of\n        // claiming what it could not verify.\n        // [impl->REQ-SHELL-KILL-AUTHENTICATED]\n        let mut survived = true;\n        for _ in 0..40 {\n            spt_store::proc::reap_if_child(p);\n            if !spt_store::proc::is_process_alive(p) {\n                survived = false;\n                break;\n            }\n            std::thread::sleep(std::time::Duration::from_millis(50));\n        }\n        if survived {\n            spt_proto::emit_line_err!(\n                \"SHELL_CLOSE_SURVIVOR:{owner}/{shell_id}: a process is still alive at recorded \\\n                 pid={p} after the close window — either the kill was refused (a preceding \\\n                 SHELL_KILL_REFUSED_* line names why) or the binary outlived it; the perch is \\\n                 being closed regardless\"\n            );\n        }\n    }\n    let _ = std::fs::remove_file(perch.join(SHELL_PID_FILE));\n    // Retire the live credential — but STASH it (rename, atomic) rather than\n    // erase it: pending spool rows are stamped under this token's key, and the\n    // relink that recovers a persistent instance must re-stamp them to its\n    // fresh key or the durable channel's rows die with the rotation\n    // (REQ-HAZARD-SHELL-STALE-ONLINE, the flynn frame-loss leg). A retired\n    // token has no live power — bind resolves only the parked LINK_TOKEN_FILE\n    // — its sole residual use is MAC verification of the frames it stamped.\n    // [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\n    // (remove-then-rename: Windows rename refuses an existing destination.)\n    let _ = std::fs::remove_file(perch.join(RETIRED_TOKEN_FILE));\n    let _ = std::fs::rename(\n        perch.join(LINK_TOKEN_FILE),\n        perch.join(RETIRED_TOKEN_FILE),\n    );\n\n    // Clear any pending drive frame (M11-W2, REQ-SHELL-3): the link credential\n    // is now retired, so a relink will mint a fresh token — but eagerly evict the\n    // slot too. This is the prompt link-break cleanup; the token-stamp on the\n    // slot is the crash-proof guarantee (a frame stamped under this dead link can\n    // never match the relink's fresh token), so a failure here is harmless.\n    // Best-effort: a downed daemon holds no slot. Runs for both close branches.\n    // [impl->REQ-SHELL-3]\n    let _ = crate::drivehub::drive_clear(&crate::endpoint::drive_socket_name(), owner, shell_id);\n\n    // Clear any live tunnel entry (M11-W3, REQ-SHELL-4): the same link-break seam\n    // as the drive clear. The link token is now retired, so the relink's fresh\n    // token can never match the entry's stamp — the held stream ids are inert even\n    // if this clear never ran (the crash-proof R1 leg). Eager eviction is the\n    // prompt cleanup. Best-effort: a downed daemon holds no entry.\n    // [impl->REQ-SHELL-4]\n    let _ = crate::tunnelhub::tunnel_clear(&crate::endpoint::tunnel_socket_name(), owner, shell_id);\n\n    // 4. The divergence.\n…\n}\n\n[Showing lines 743-893 of 1826. Use :894 to continue]"}],"details":{"truncation":{"content":"#[derive(Debug, Clone, Copy, PartialEq, Eq)]\r\npub enum CloseOutcome {\r\n    /// `ephemeral = true`: perch + spool history erased, mint slot freed.\r\n    TornDown,\r\n    /// Everything else: perch kept, status=offline, re-linkable.\r\n    Offline,\r\n}\r\n\r\n/// The link-break close (M5-D4a, REQ-SHELL-2 — CONTEXT §Shell model\r\n/// lifecycle). The contract used to read \"the binary never survives\"; since\r\n/// BAROMETER it reads **a link-break always closes the link and always ends the\r\n/// binary we can still prove is ours** — a kill that cannot authenticate its\r\n/// target is refused and announced rather than fired blind, because the\r\n/// alternative is force-killing a stranger's process tree on a recycled pid:\r\n///\r\n/// 1. **pre-close instruction** (manifest `pre_close`, optional): composed as\r\n///    a [`shell_close` frame](crate::shellchan::compose_close_frame), stamped\r\n///    with the link key, and delivered per the receipt mode — relay: spooled\r\n///    (the binary may drain it during the grace window); stdin: spooled +\r\n///    best-effort broker push. Skipped when nothing is running.\r\n/// 2. **termination window** (`close_timeout_ms`, absent ⇒ 0): poll the\r\n///    recorded pid for a graceful exit up to the window.\r\n/// 3. **force-close**: an **authenticated** kill at window end — the pid+birth\r\n///    pair must prove the process at the recorded number is the binary we\r\n///    launched (`REQ-SHELL-KILL-AUTHENTICATED`). It was unconditional until\r\n///    BAROMETER, and that is precisely what let a recycled pid route to\r\n///    `taskkill /T`. A refused kill is announced (`SHELL_KILL_REFUSED_*`) and\r\n///    the close continues; the pid file and the **link token die with the\r\n///    link** either way (a stale token must never drive a relinked instance —\r\n///    relink mints fresh).\r\n/// 4. **divergence**: `ephemeral` ⇒ full teardown + history erase (the spool\r\n///    lives in the perch dir — removal IS the erase); otherwise the perch\r\n///    stays, status=offline, re-linkable.\r\n///\r\n/// `shell = None` (adapter deregistered after spawn) closes with defaults: no\r\n/// pre-close, no grace, **non-ephemeral** — history is never erased without\r\n/// the manifest's authority.\r\n// [impl->REQ-SHELL-2]\r\npub fn close_shell(\r\n    owlery: &Path,\r\n    owner: &str,\r\n    shell_id: &str,\r\n    shell: Option<&Shell>,\r\n) -> Result<CloseOutcome, String> {\r\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\r\n    let pid = std::fs::read_to_string(perch.join(SHELL_PID_FILE))\r\n        .ok()\r\n        .and_then(|s| s.trim().parse::<u32>().ok())\r\n        .filter(|&p| p != 0);\r\n    // Reap first: an in-process spawner's exited child is a zombie until\r\n    // somebody waits — without this the liveness probe would hold a dead\r\n    // binary \"running\" through the whole grace window.\r\n    if let Some(p) = pid {\r\n        spt_store::proc::reap_if_child(p);\r\n    }\r\n    let running = pid.is_some_and(spt_store::proc::is_process_alive);\r\n\r\n    // 1. The pre-close instruction — only meaningful while something runs.\r\n    if running {\r\n        if let Some(instruction) = shell.and_then(|s| s.pre_close.as_deref()) {\r\n            if let Some(token) = read_link_token(&perch) {\r\n                let key = link_key(&token);\r\n                let frame = crate::shellchan::compose_close_frame(owner, instruction);\r\n                let stamped = crate::shellchan::stamp_frame(&key, &frame);\r\n                // Best-effort by design: a wedged spool/broker must never\r\n                // block the close — the force-kill below is the guarantee.\r\n                let _ = crate::shellchan::spool_shell_frame(owlery, owner, shell_id, &stamped);\r\n                if shell.map(|s| s.command_receipt) == Some(Some(CommandReceipt::Stdin)) {\r\n                    let _ = crate::shellchan::deliver_stdin_pending(owlery, owner, shell_id);\r\n                }\r\n            }\r\n        }\r\n        // 2. The termination window: a graceful exit beats the force-kill.\r\n        let window =\r\n            std::time::Duration::from_millis(shell.and_then(|s| s.close_timeout_ms).unwrap_or(0));\r\n        let deadline = std::time::Instant::now() + window;\r\n        while std::time::Instant::now() < deadline {\r\n            if !pid.is_some_and(spt_store::proc::is_process_alive) {\r\n                break;\r\n            }\r\n            std::thread::sleep(std::time::Duration::from_millis(50));\r\n        }\r\n    }\r\n\r\n    // 3. Force-close (unconditional — the binary never survives a\r\n    // link-break) + retire the link credentials. A watcher should not exist\r\n    // while the binary side was up, but a torn-down instance must orphan\r\n    // nothing — kill both unconditionally.\r\n    kill_shell_at(&perch);\r\n    crate::shellwake::kill_waker_at(&perch);\r\n    if let Some(p) = pid {\r\n        // The kill is asynchronous (and leaves a zombie when we are the\r\n        // spawner) — confirm the death bounded, and SAY SO when it does not\r\n        // come. Since the kill authenticates, this loop can now legitimately\r\n        // end with the process still there (a refused kill), so a silent\r\n        // return would trade a loud wrong action for a quiet wrong state —\r\n        // the un-RCA-able one. It reports what it could not do instead of\r\n        // claiming what it could not verify.\r\n        // [impl->REQ-SHELL-KILL-AUTHENTICATED]\r\n        let mut survived = true;\r\n        for _ in 0..40 {\r\n            spt_store::proc::reap_if_child(p);\r\n            if !spt_store::proc::is_process_alive(p) {\r\n                survived = false;\r\n                break;\r\n            }\r\n            std::thread::sleep(std::time::Duration::from_millis(50));\r\n        }\r\n        if survived {\r\n            spt_proto::emit_line_err!(\r\n                \"SHELL_CLOSE_SURVIVOR:{owner}/{shell_id}: a process is still alive at recorded \\\r\n                 pid={p} after the close window — either the kill was refused (a preceding \\\r\n                 SHELL_KILL_REFUSED_* line names why) or the binary outlived it; the perch is \\\r\n                 being closed regardless\"\r\n            );\r\n        }\r\n    }\r\n    let _ = std::fs::remove_file(perch.join(SHELL_PID_FILE));\r\n    // Retire the live credential — but STASH it (rename, atomic) rather than\r\n    // erase it: pending spool rows are stamped under this token's key, and the\r\n    // relink that recovers a persistent instance must re-stamp them to its\r\n    // fresh key or the durable channel's rows die with the rotation\r\n    // (REQ-HAZARD-SHELL-STALE-ONLINE, the flynn frame-loss leg). A retired\r\n    // token has no live power — bind resolves only the parked LINK_TOKEN_FILE\r\n    // — its sole residual use is MAC verification of the frames it stamped.\r\n    // [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\r\n    // (remove-then-rename: Windows rename refuses an existing destination.)\r\n    let _ = std::fs::remove_file(perch.join(RETIRED_TOKEN_FILE));\r\n    let _ = std::fs::rename(\r\n        perch.join(LINK_TOKEN_FILE),\r\n        perch.join(RETIRED_TOKEN_FILE),\r\n    );\r\n\r\n    // Clear any pending drive frame (M11-W2, REQ-SHELL-3): the link credential\r\n    // is now retired, so a relink will mint a fresh token — but eagerly evict the\r\n    // slot too. This is the prompt link-break cleanup; the token-stamp on the\r\n    // slot is the crash-proof guarantee (a frame stamped under this dead link can\r\n    // never match the relink's fresh token), so a failure here is harmless.\r\n    // Best-effort: a downed daemon holds no slot. Runs for both close branches.\r\n    // [impl->REQ-SHELL-3]\r\n    let _ = crate::drivehub::drive_clear(&crate::endpoint::drive_socket_name(), owner, shell_id);\r\n\r\n    // Clear any live tunnel entry (M11-W3, REQ-SHELL-4): the same link-break seam\r\n    // as the drive clear. The link token is now retired, so the relink's fresh\r\n    // token can never match the entry's stamp — the held stream ids are inert even\r\n    // if this clear never ran (the crash-proof R1 leg). Eager eviction is the\r\n    // prompt cleanup. Best-effort: a downed daemon holds no entry.\r\n    // [impl->REQ-SHELL-4]\r\n    let _ = crate::tunnelhub::tunnel_clear(&crate::endpoint::tunnel_socket_name(), owner, shell_id);\r\n\r\n    // 4. The divergence.\r","truncated":true,"truncatedBy":"lines","totalLines":1084,"totalBytes":8226,"outputLines":151,"outputBytes":8226,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":1826,"displayContent":{"text":"#[derive(Debug, Clone, Copy, PartialEq, Eq)]\r\npub enum CloseOutcome {\r\n    /// `ephemeral = true`: perch + spool history erased, mint slot freed.\r\n    TornDown,\r\n    /// Everything else: perch kept, status=offline, re-linkable.\r\n    Offline,\r\n}\r\n\r\n/// The link-break close (M5-D4a, REQ-SHELL-2 — CONTEXT §Shell model\r\n/// lifecycle). The contract used to read \"the binary never survives\"; since\r\n/// BAROMETER it reads **a link-break always closes the link and always ends the\r\n/// binary we can still prove is ours** — a kill that cannot authenticate its\r\n/// target is refused and announced rather than fired blind, because the\r\n/// alternative is force-killing a stranger's process tree on a recycled pid:\r\n///\r\n/// 1. **pre-close instruction** (manifest `pre_close`, optional): composed as\r\n///    a [`shell_close` frame](crate::shellchan::compose_close_frame), stamped\r\n///    with the link key, and delivered per the receipt mode — relay: spooled\r\n///    (the binary may drain it during the grace window); stdin: spooled +\r\n///    best-effort broker push. Skipped when nothing is running.\r\n/// 2. **termination window** (`close_timeout_ms`, absent ⇒ 0): poll the\r\n///    recorded pid for a graceful exit up to the window.\r\n/// 3. **force-close**: an **authenticated** kill at window end — the pid+birth\r\n///    pair must prove the process at the recorded number is the binary we\r\n///    launched (`REQ-SHELL-KILL-AUTHENTICATED`). It was unconditional until\r\n///    BAROMETER, and that is precisely what let a recycled pid route to\r\n///    `taskkill /T`. A refused kill is announced (`SHELL_KILL_REFUSED_*`) and\r\n///    the close continues; the pid file and the **link token die with the\r\n///    link** either way (a stale token must never drive a relinked instance —\r\n///    relink mints fresh).\r\n/// 4. **divergence**: `ephemeral` ⇒ full teardown + history erase (the spool\r\n///    lives in the perch dir — removal IS the erase); otherwise the perch\r\n///    stays, status=offline, re-linkable.\r\n///\r\n/// `shell = None` (adapter deregistered after spawn) closes with defaults: no\r\n/// pre-close, no grace, **non-ephemeral** — history is never erased without\r\n/// the manifest's authority.\r\n// [impl->REQ-SHELL-2]\r\npub fn close_shell(\r\n    owlery: &Path,\r\n    owner: &str,\r\n    shell_id: &str,\r\n    shell: Option<&Shell>,\r\n) -> Result<CloseOutcome, String> {\r\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\r\n    let pid = std::fs::read_to_string(perch.join(SHELL_PID_FILE))\r\n        .ok()\r\n        .and_then(|s| s.trim().parse::<u32>().ok())\r\n        .filter(|&p| p != 0);\r\n    // Reap first: an in-process spawner's exited child is a zombie until\r\n    // somebody waits — without this the liveness probe would hold a dead\r\n    // binary \"running\" through the whole grace window.\r\n    if let Some(p) = pid {\r\n        spt_store::proc::reap_if_child(p);\r\n    }\r\n    let running = pid.is_some_and(spt_store::proc::is_process_alive);\r\n\r\n    // 1. The pre-close instruction — only meaningful while something runs.\r\n    if running {\r\n        if let Some(instruction) = shell.and_then(|s| s.pre_close.as_deref()) {\r\n            if let Some(token) = read_link_token(&perch) {\r\n                let key = link_key(&token);\r\n                let frame = crate::shellchan::compose_close_frame(owner, instruction);\r\n                let stamped = crate::shellchan::stamp_frame(&key, &frame);\r\n                // Best-effort by design: a wedged spool/broker must never\r\n                // block the close — the force-kill below is the guarantee.\r\n                let _ = crate::shellchan::spool_shell_frame(owlery, owner, shell_id, &stamped);\r\n                if shell.map(|s| s.command_receipt) == Some(Some(CommandReceipt::Stdin)) {\r\n                    let _ = crate::shellchan::deliver_stdin_pending(owlery, owner, shell_id);\r\n                }\r\n            }\r\n        }\r\n        // 2. The termination window: a graceful exit beats the force-kill.\r\n        let window =\r\n            std::time::Duration::from_millis(shell.and_then(|s| s.close_timeout_ms).unwrap_or(0));\r\n        let deadline = std::time::Instant::now() + window;\r\n        while std::time::Instant::now() < deadline {\r\n            if !pid.is_some_and(spt_store::proc::is_process_alive) {\r\n                break;\r\n            }\r\n            std::thread::sleep(std::time::Duration::from_millis(50));\r\n        }\r\n    }\r\n\r\n    // 3. Force-close (unconditional — the binary never survives a\r\n    // link-break) + retire the link credentials. A watcher should not exist\r\n    // while the binary side was up, but a torn-down instance must orphan\r\n    // nothing — kill both unconditionally.\r\n    kill_shell_at(&perch);\r\n    crate::shellwake::kill_waker_at(&perch);\r\n    if let Some(p) = pid {\r\n        // The kill is asynchronous (and leaves a zombie when we are the\r\n        // spawner) — confirm the death bounded, and SAY SO when it does not\r\n        // come. Since the kill authenticates, this loop can now legitimately\r\n        // end with the process still there (a refused kill), so a silent\r\n        // return would trade a loud wrong action for a quiet wrong state —\r\n        // the un-RCA-able one. It reports what it could not do instead of\r\n        // claiming what it could not verify.\r\n        // [impl->REQ-SHELL-KILL-AUTHENTICATED]\r\n        let mut survived = true;\r\n        for _ in 0..40 {\r\n            spt_store::proc::reap_if_child(p);\r\n            if !spt_store::proc::is_process_alive(p) {\r\n                survived = false;\r\n                break;\r\n            }\r\n            std::thread::sleep(std::time::Duration::from_millis(50));\r\n        }\r\n        if survived {\r\n            spt_proto::emit_line_err!(\r\n                \"SHELL_CLOSE_SURVIVOR:{owner}/{shell_id}: a process is still alive at recorded \\\r\n                 pid={p} after the close window — either the kill was refused (a preceding \\\r\n                 SHELL_KILL_REFUSED_* line names why) or the binary outlived it; the perch is \\\r\n                 being closed regardless\"\r\n            );\r\n        }\r\n    }\r\n    let _ = std::fs::remove_file(perch.join(SHELL_PID_FILE));\r\n    // Retire the live credential — but STASH it (rename, atomic) rather than\r\n    // erase it: pending spool rows are stamped under this token's key, and the\r\n    // relink that recovers a persistent instance must re-stamp them to its\r\n    // fresh key or the durable channel's rows die with the rotation\r\n    // (REQ-HAZARD-SHELL-STALE-ONLINE, the flynn frame-loss leg). A retired\r\n    // token has no live power — bind resolves only the parked LINK_TOKEN_FILE\r\n    // — its sole residual use is MAC verification of the frames it stamped.\r\n    // [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\r\n    // (remove-then-rename: Windows rename refuses an existing destination.)\r\n    let _ = std::fs::remove_file(perch.join(RETIRED_TOKEN_FILE));\r\n    let _ = std::fs::rename(\r\n        perch.join(LINK_TOKEN_FILE),\r\n        perch.join(RETIRED_TOKEN_FILE),\r\n    );\r\n\r\n    // Clear any pending drive frame (M11-W2, REQ-SHELL-3): the link credential\r\n    // is now retired, so a relink will mint a fresh token — but eagerly evict the\r\n    // slot too. This is the prompt link-break cleanup; the token-stamp on the\r\n    // slot is the crash-proof guarantee (a frame stamped under this dead link can\r\n    // never match the relink's fresh token), so a failure here is harmless.\r\n    // Best-effort: a downed daemon holds no slot. Runs for both close branches.\r\n    // [impl->REQ-SHELL-3]\r\n    let _ = crate::drivehub::drive_clear(&crate::endpoint::drive_socket_name(), owner, shell_id);\r\n\r\n    // Clear any live tunnel entry (M11-W3, REQ-SHELL-4): the same link-break seam\r\n    // as the drive clear. The link token is now retired, so the relink's fresh\r\n    // token can never match the entry's stamp — the held stream ids are inert even\r\n    // if this clear never ran (the crash-proof R1 leg). Eager eviction is the\r\n    // prompt cleanup. Best-effort: a downed daemon holds no entry.\r\n    // [impl->REQ-SHELL-4]\r\n    let _ = crate::tunnelhub::tunnel_clear(&crate::endpoint::tunnel_socket_name(), owner, shell_id);\r\n\r\n    // 4. The divergence.\r\n…\n}","startLine":743,"lineNumbers":[743,744,745,746,747,748,749,750,751,752,753,754,755,756,757,758,759,760,761,762,763,764,765,766,767,768,769,770,771,772,773,774,775,776,777,778,779,780,781,782,783,784,785,786,787,788,789,790,791,792,793,794,795,796,797,798,799,800,801,802,803,804,805,806,807,808,809,810,811,812,813,814,815,816,817,818,819,820,821,822,823,824,825,826,827,828,829,830,831,832,833,834,835,836,837,838,839,840,841,842,843,844,845,846,847,848,849,850,851,852,853,854,855,856,857,858,859,860,861,862,863,864,865,866,867,868,869,870,871,872,873,874,875,876,877,878,879,880,881,882,883,884,885,886,887,888,889,890,891,892,893,null,905]},"fileSize":88363,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\crates\\spt-daemon\\src\\shellhost.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":1826,"totalBytes":8226,"outputLines":151,"outputBytes":8226,"shownRange":{"start":743,"end":893},"nextOffset":894}}},"isError":false,"timestamp":1789001235271}}
{"type":"message","id":"83d1cbcd","parentId":"775d8b57","timestamp":"2026-09-10T00:47:15.275Z","message":{"role":"toolResult","toolCallId":"call_17oGaDdBH2hPPUCZfdw5D1YS|fc_077f3569365e313d016aa1fe1322d887d1a98ae52f0506a127","toolName":"read","content":[{"type":"text","text":"pub fn write_line<W: Write + ?Sized>(w: &mut W, rendered: &str) -> io::Result<()> {\n…\n    write_composed(w, &compose_line(rendered))\n}\n\n/// The multi-line block emission path: interior breaks kept, still one write.\n// [impl->REQ-EMIT-SINGLE-WRITE]\npub fn write_block<W: Write + ?Sized>(w: &mut W, rendered: &str) -> io::Result<()> {\n    write_composed(w, &compose_block(rendered))\n}\n\n/// Emit ONE machine-parsed line to `$w` in exactly one write.\n///\n/// ```\n/// # use spt_proto::emit::CountingWriter;\n/// let mut w = CountingWriter::new();\n/// let id = \"endpoint-7\";\n/// spt_proto::emit_line!(&mut w, \"SUBSCRIBE_DECISION:{id}: seated\");\n/// assert_eq!(w.call_count(), 1);\n/// ```\n// [impl->REQ-EMIT-SINGLE-WRITE]\n#[macro_export]\nmacro_rules! emit_line {\n    ($w:expr, $($arg:tt)*) => {\n        $crate::emit::write_line($w, &::std::format!($($arg)*))\n    };\n}\n\n/// Emit one composed MULTI-LINE block to `$w` in exactly one write.\n// [impl->REQ-EMIT-SINGLE-WRITE]\n#[macro_export]\nmacro_rules! emit_block {\n    ($w:expr, $($arg:tt)*) => {\n        $crate::emit::write_block($w, &::std::format!($($arg)*))\n    };\n}\n\n/// [`emit_line!`] aimed at this process's stderr — the conversion target for the\n/// `TOKEN:`-shaped `eprintln!` population.\n///\n/// Deliberately a SEPARATE NAME rather than a no-writer arm of [`emit_line!`]:\n/// a macro that accepts both `($w, $fmt, ...)` and `($fmt, ...)` would bind a\n/// caller's format literal as the writer, since a literal is also an expression.\n/// One seam, two spellings, no ambiguity.\n///\n/// It returns `()` and DISCARDS a write error, which is a deliberate departure\n/// from `eprintln!` — stated here at the seam, where the next reader is, rather\n/// than only in the commit that made it. std's print machinery PANICS when\n/// writing to stderr fails, so a closed or full stderr can take down a process\n/// at the exact moment it was trying to explain itself. That is the same wedge\n/// class this project has refused before: a reporting surface must never become\n/// the thing that kills the operation it reports on.\n///\n/// The cost is named, not hidden: this makes emission a place where absence can\n/// be SILENT — a diagnostic that never arrives leaves no trace of having been\n/// attempted. That is an observability gap, accepted here only because the\n/// alternative is a panic, and it belongs in the conversion notes so no later\n/// reader rediscovers it as a defect.\n// [impl->REQ-EMIT-SINGLE-WRITE]\n#[macro_export]\nmacro_rules! emit_line_err {\n    ($($arg:tt)*) => {{\n        let _ = $crate::emit::write_line(\n            &mut ::std::io::stderr().lock(),\n            &::std::format!($($arg)*),\n        );\n    }};\n}\n\n/// [`emit_block!`] aimed at this process's stderr, on the same terms as\n/// [`emit_line_err!`]: separate name, error discarded.\n// [impl->REQ-EMIT-SINGLE-WRITE]\n#[macro_export]\nmacro_rules! emit_block_err {\n    ($($arg:tt)*) => {{\n        let _ = $crate::emit::write_block(\n            &mut ::std::io::stderr().lock(),\n            &::std::format!($($arg)*),\n        );\n    }};\n}\n\n\n[Showing lines 129-208 of 543. Use :209 to continue]"}],"details":{"truncation":{"content":"    write_composed(w, &compose_line(rendered))\r\n}\r\n\r\n/// The multi-line block emission path: interior breaks kept, still one write.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\npub fn write_block<W: Write + ?Sized>(w: &mut W, rendered: &str) -> io::Result<()> {\r\n    write_composed(w, &compose_block(rendered))\r\n}\r\n\r\n/// Emit ONE machine-parsed line to `$w` in exactly one write.\r\n///\r\n/// ```\r\n/// # use spt_proto::emit::CountingWriter;\r\n/// let mut w = CountingWriter::new();\r\n/// let id = \"endpoint-7\";\r\n/// spt_proto::emit_line!(&mut w, \"SUBSCRIBE_DECISION:{id}: seated\");\r\n/// assert_eq!(w.call_count(), 1);\r\n/// ```\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\n#[macro_export]\r\nmacro_rules! emit_line {\r\n    ($w:expr, $($arg:tt)*) => {\r\n        $crate::emit::write_line($w, &::std::format!($($arg)*))\r\n    };\r\n}\r\n\r\n/// Emit one composed MULTI-LINE block to `$w` in exactly one write.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\n#[macro_export]\r\nmacro_rules! emit_block {\r\n    ($w:expr, $($arg:tt)*) => {\r\n        $crate::emit::write_block($w, &::std::format!($($arg)*))\r\n    };\r\n}\r\n\r\n/// [`emit_line!`] aimed at this process's stderr — the conversion target for the\r\n/// `TOKEN:`-shaped `eprintln!` population.\r\n///\r\n/// Deliberately a SEPARATE NAME rather than a no-writer arm of [`emit_line!`]:\r\n/// a macro that accepts both `($w, $fmt, ...)` and `($fmt, ...)` would bind a\r\n/// caller's format literal as the writer, since a literal is also an expression.\r\n/// One seam, two spellings, no ambiguity.\r\n///\r\n/// It returns `()` and DISCARDS a write error, which is a deliberate departure\r\n/// from `eprintln!` — stated here at the seam, where the next reader is, rather\r\n/// than only in the commit that made it. std's print machinery PANICS when\r\n/// writing to stderr fails, so a closed or full stderr can take down a process\r\n/// at the exact moment it was trying to explain itself. That is the same wedge\r\n/// class this project has refused before: a reporting surface must never become\r\n/// the thing that kills the operation it reports on.\r\n///\r\n/// The cost is named, not hidden: this makes emission a place where absence can\r\n/// be SILENT — a diagnostic that never arrives leaves no trace of having been\r\n/// attempted. That is an observability gap, accepted here only because the\r\n/// alternative is a panic, and it belongs in the conversion notes so no later\r\n/// reader rediscovers it as a defect.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\n#[macro_export]\r\nmacro_rules! emit_line_err {\r\n    ($($arg:tt)*) => {{\r\n        let _ = $crate::emit::write_line(\r\n            &mut ::std::io::stderr().lock(),\r\n            &::std::format!($($arg)*),\r\n        );\r\n    }};\r\n}\r\n\r\n/// [`emit_block!`] aimed at this process's stderr, on the same terms as\r\n/// [`emit_line_err!`]: separate name, error discarded.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\n#[macro_export]\r\nmacro_rules! emit_block_err {\r\n    ($($arg:tt)*) => {{\r\n        let _ = $crate::emit::write_block(\r\n            &mut ::std::io::stderr().lock(),\r\n            &::std::format!($($arg)*),\r\n        );\r\n    }};\r\n}\r\n\r","truncated":true,"truncatedBy":"lines","totalLines":415,"totalBytes":3050,"outputLines":80,"outputBytes":3050,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":543,"displayContent":{"text":"pub fn write_line<W: Write + ?Sized>(w: &mut W, rendered: &str) -> io::Result<()> {\n…\n    write_composed(w, &compose_line(rendered))\r\n}\r\n\r\n/// The multi-line block emission path: interior breaks kept, still one write.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\npub fn write_block<W: Write + ?Sized>(w: &mut W, rendered: &str) -> io::Result<()> {\r\n    write_composed(w, &compose_block(rendered))\r\n}\r\n\r\n/// Emit ONE machine-parsed line to `$w` in exactly one write.\r\n///\r\n/// ```\r\n/// # use spt_proto::emit::CountingWriter;\r\n/// let mut w = CountingWriter::new();\r\n/// let id = \"endpoint-7\";\r\n/// spt_proto::emit_line!(&mut w, \"SUBSCRIBE_DECISION:{id}: seated\");\r\n/// assert_eq!(w.call_count(), 1);\r\n/// ```\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\n#[macro_export]\r\nmacro_rules! emit_line {\r\n    ($w:expr, $($arg:tt)*) => {\r\n        $crate::emit::write_line($w, &::std::format!($($arg)*))\r\n    };\r\n}\r\n\r\n/// Emit one composed MULTI-LINE block to `$w` in exactly one write.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\n#[macro_export]\r\nmacro_rules! emit_block {\r\n    ($w:expr, $($arg:tt)*) => {\r\n        $crate::emit::write_block($w, &::std::format!($($arg)*))\r\n    };\r\n}\r\n\r\n/// [`emit_line!`] aimed at this process's stderr — the conversion target for the\r\n/// `TOKEN:`-shaped `eprintln!` population.\r\n///\r\n/// Deliberately a SEPARATE NAME rather than a no-writer arm of [`emit_line!`]:\r\n/// a macro that accepts both `($w, $fmt, ...)` and `($fmt, ...)` would bind a\r\n/// caller's format literal as the writer, since a literal is also an expression.\r\n/// One seam, two spellings, no ambiguity.\r\n///\r\n/// It returns `()` and DISCARDS a write error, which is a deliberate departure\r\n/// from `eprintln!` — stated here at the seam, where the next reader is, rather\r\n/// than only in the commit that made it. std's print machinery PANICS when\r\n/// writing to stderr fails, so a closed or full stderr can take down a process\r\n/// at the exact moment it was trying to explain itself. That is the same wedge\r\n/// class this project has refused before: a reporting surface must never become\r\n/// the thing that kills the operation it reports on.\r\n///\r\n/// The cost is named, not hidden: this makes emission a place where absence can\r\n/// be SILENT — a diagnostic that never arrives leaves no trace of having been\r\n/// attempted. That is an observability gap, accepted here only because the\r\n/// alternative is a panic, and it belongs in the conversion notes so no later\r\n/// reader rediscovers it as a defect.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\n#[macro_export]\r\nmacro_rules! emit_line_err {\r\n    ($($arg:tt)*) => {{\r\n        let _ = $crate::emit::write_line(\r\n            &mut ::std::io::stderr().lock(),\r\n            &::std::format!($($arg)*),\r\n        );\r\n    }};\r\n}\r\n\r\n/// [`emit_block!`] aimed at this process's stderr, on the same terms as\r\n/// [`emit_line_err!`]: separate name, error discarded.\r\n// [impl->REQ-EMIT-SINGLE-WRITE]\r\n#[macro_export]\r\nmacro_rules! emit_block_err {\r\n    ($($arg:tt)*) => {{\r\n        let _ = $crate::emit::write_block(\r\n            &mut ::std::io::stderr().lock(),\r\n            &::std::format!($($arg)*),\r\n        );\r\n    }};\r\n}\r\n\r","startLine":122,"lineNumbers":[122,null,129,130,131,132,133,134,135,136,137,138,139,140,141,142,143,144,145,146,147,148,149,150,151,152,153,154,155,156,157,158,159,160,161,162,163,164,165,166,167,168,169,170,171,172,173,174,175,176,177,178,179,180,181,182,183,184,185,186,187,188,189,190,191,192,193,194,195,196,197,198,199,200,201,202,203,204,205,206,207,208]},"fileSize":22720,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\crates\\spt-proto\\src\\emit.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":543,"totalBytes":3050,"outputLines":80,"outputBytes":3050,"shownRange":{"start":129,"end":208},"nextOffset":209}}},"isError":false,"timestamp":1789001235274}}
{"type":"message","id":"4eb67798","parentId":"83d1cbcd","timestamp":"2026-09-10T00:47:15.281Z","message":{"role":"toolResult","toolCallId":"call_up92j1SMTbYoqjfZFek1mg3D|fc_077f3569365e313d016aa1fe1322bc87d1aa206667acff4cd3","toolName":"read","content":[{"type":"text","text":" \n **Exit-opcode supervision:** the wake-watcher exiting with the **wake opcode** → spt-core runs the wake resolution (below) and brings the shell online. Any *other* exit (crash) → respawn the watcher with **exponential backoff + eventual give-up** (until next shell activity) — crash-bug safety. One watcher per offline shell instance.\n \ndiff --git a/crates/spt-daemon/src/shellhost.rs b/crates/spt-daemon/src/shellhost.rs\nindex 593c6f68..44df752b 100644\n--- a/crates/spt-daemon/src/shellhost.rs\n+++ b/crates/spt-daemon/src/shellhost.rs\n@@ -312,16 +312,14 @@ pub fn routes_through_daemon(shell: &Shell) -> bool {\n /// down and adopt it. Used by BOTH sides of the routed launch, because a late\n /// ack can put the caller and the daemon in the same launch at once.\n ///\n-/// DERIVED, never the recorded status (REQ-HAZARD-SHELL-STALE-ONLINE): a pid\n-/// that is absent, unparseable or `0` reads NOT-a-winner, the same\n-/// fail-toward-alive stance the rest of the shell machinery holds — so a\n-/// backend that records no pid can never block its own relaunch. Same shape and\n-/// same non-fault wording as [`crate::shellwake::resolve_wake`]'s existing\n-/// guard; this is that rule extracted, not a second one.\n+/// Birth-safe, including a binary awaiting bind: a recycled live pid is NOT a\n+/// winner when its native birth stamp mismatches. Missing, unparseable and zero\n+/// pids remain not-a-winner, preserving pid-less backend behavior.\n // [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\n+// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n pub fn live_launch_winner(shell_perch: &Path) -> Option<u32> {\n     shellinfo::read_shell_pid(shell_perch)\n-        .filter(|&p| p != 0 && spt_store::proc::is_process_alive(p))\n+        .filter(|&p| p != 0 && !shellinfo::shell_pid_provably_dead(shell_perch))\n }\n \n /// THE route-selection seam. Generic over the daemon call and the in-process\n@@ -1204,6 +1202,7 @@ mod tests {\n     // that parses non-zero AND probes alive is a winner. Absent / `0` /\n     // unparseable / dead must all read NOT-a-winner, or a backend that records\n     // no usable pid would be unable to relaunch its own instance.\n+    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n     #[test]\n     fn only_a_live_recorded_pid_wins_a_launch_race() {\n         let tmp = tempfile::tempdir().unwrap();\n@@ -1225,6 +1224,18 @@ mod tests {\n             Some(me),\n             \"a recorded pid that probes alive IS the winner\"\n\n[Artifact storage: ~/.omp/agent/sessions/-Documents-projects-spt-core/2026-09-09T23-45-42-317Z_01a08890-21ad-7000-8f54-208d4f4f2f2c/112.bash-original.log (105.3KB). Use artifact://112:N-M to page, artifact://112:raw:N-M for verbatim chunks, and the artifact file path for search/copy workflows.]\n\n[Showing lines 42-81 of 1088. Use :82 to continue]"}],"details":{"resolvedPath":"C:\\Users\\decid\\.omp\\agent\\sessions\\-Documents-projects-spt-core\\2026-09-09T23-45-42-317Z_01a08890-21ad-7000-8f54-208d4f4f2f2c\\112.bash-original.log","contentType":"text/plain","totalLines":1088,"displayContent":{"text":" \n **Exit-opcode supervision:** the wake-watcher exiting with the **wake opcode** → spt-core runs the wake resolution (below) and brings the shell online. Any *other* exit (crash) → respawn the watcher with **exponential backoff + eventual give-up** (until next shell activity) — crash-bug safety. One watcher per offline shell instance.\n \ndiff --git a/crates/spt-daemon/src/shellhost.rs b/crates/spt-daemon/src/shellhost.rs\nindex 593c6f68..44df752b 100644\n--- a/crates/spt-daemon/src/shellhost.rs\n+++ b/crates/spt-daemon/src/shellhost.rs\n@@ -312,16 +312,14 @@ pub fn routes_through_daemon(shell: &Shell) -> bool {\n /// down and adopt it. Used by BOTH sides of the routed launch, because a late\n /// ack can put the caller and the daemon in the same launch at once.\n ///\n-/// DERIVED, never the recorded status (REQ-HAZARD-SHELL-STALE-ONLINE): a pid\n-/// that is absent, unparseable or `0` reads NOT-a-winner, the same\n-/// fail-toward-alive stance the rest of the shell machinery holds — so a\n-/// backend that records no pid can never block its own relaunch. Same shape and\n-/// same non-fault wording as [`crate::shellwake::resolve_wake`]'s existing\n-/// guard; this is that rule extracted, not a second one.\n+/// Birth-safe, including a binary awaiting bind: a recycled live pid is NOT a\n+/// winner when its native birth stamp mismatches. Missing, unparseable and zero\n+/// pids remain not-a-winner, preserving pid-less backend behavior.\n // [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\n+// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n pub fn live_launch_winner(shell_perch: &Path) -> Option<u32> {\n     shellinfo::read_shell_pid(shell_perch)\n-        .filter(|&p| p != 0 && spt_store::proc::is_process_alive(p))\n+        .filter(|&p| p != 0 && !shellinfo::shell_pid_provably_dead(shell_perch))\n }\n \n /// THE route-selection seam. Generic over the daemon call and the in-process\n@@ -1204,6 +1202,7 @@ mod tests {\n     // that parses non-zero AND probes alive is a winner. Absent / `0` /\n     // unparseable / dead must all read NOT-a-winner, or a backend that records\n     // no usable pid would be unable to relaunch its own instance.\n+    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n     #[test]\n     fn only_a_live_recorded_pid_wins_a_launch_race() {\n         let tmp = tempfile::tempdir().unwrap();\n@@ -1225,6 +1224,18 @@ mod tests {\n             Some(me),\n             \"a recorded pid that probes alive IS the winner\"","startLine":42,"lineNumbers":[42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81]},"truncation":{"content":" \n **Exit-opcode supervision:** the wake-watcher exiting with the **wake opcode** → spt-core runs the wake resolution (below) and brings the shell online. Any *other* exit (crash) → respawn the watcher with **exponential backoff + eventual give-up** (until next shell activity) — crash-bug safety. One watcher per offline shell instance.\n \ndiff --git a/crates/spt-daemon/src/shellhost.rs b/crates/spt-daemon/src/shellhost.rs\nindex 593c6f68..44df752b 100644\n--- a/crates/spt-daemon/src/shellhost.rs\n+++ b/crates/spt-daemon/src/shellhost.rs\n@@ -312,16 +312,14 @@ pub fn routes_through_daemon(shell: &Shell) -> bool {\n /// down and adopt it. Used by BOTH sides of the routed launch, because a late\n /// ack can put the caller and the daemon in the same launch at once.\n ///\n-/// DERIVED, never the recorded status (REQ-HAZARD-SHELL-STALE-ONLINE): a pid\n-/// that is absent, unparseable or `0` reads NOT-a-winner, the same\n-/// fail-toward-alive stance the rest of the shell machinery holds — so a\n-/// backend that records no pid can never block its own relaunch. Same shape and\n-/// same non-fault wording as [`crate::shellwake::resolve_wake`]'s existing\n-/// guard; this is that rule extracted, not a second one.\n+/// Birth-safe, including a binary awaiting bind: a recycled live pid is NOT a\n+/// winner when its native birth stamp mismatches. Missing, unparseable and zero\n+/// pids remain not-a-winner, preserving pid-less backend behavior.\n // [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\n+// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n pub fn live_launch_winner(shell_perch: &Path) -> Option<u32> {\n     shellinfo::read_shell_pid(shell_perch)\n-        .filter(|&p| p != 0 && spt_store::proc::is_process_alive(p))\n+        .filter(|&p| p != 0 && !shellinfo::shell_pid_provably_dead(shell_perch))\n }\n \n /// THE route-selection seam. Generic over the daemon call and the in-process\n@@ -1204,6 +1202,7 @@ mod tests {\n     // that parses non-zero AND probes alive is a winner. Absent / `0` /\n     // unparseable / dead must all read NOT-a-winner, or a backend that records\n     // no usable pid would be unable to relaunch its own instance.\n+    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n     #[test]\n     fn only_a_live_recorded_pid_wins_a_launch_race() {\n         let tmp = tempfile::tempdir().unwrap();\n@@ -1225,6 +1224,18 @@ mod tests {\n             Some(me),\n             \"a recorded pid that probes alive IS the winner\"","truncated":true,"truncatedBy":"lines","totalLines":1047,"totalBytes":2430,"outputLines":40,"outputBytes":2430,"lastLinePartial":false,"firstLineExceedsLimit":false},"meta":{"source":{"type":"internal","value":"artifact://112"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":1088,"totalBytes":2430,"outputLines":40,"outputBytes":2430,"maxBytes":51200,"shownRange":{"start":42,"end":81},"nextOffset":82}}},"isError":false,"timestamp":1789001235281}}
{"type":"custom_message","customType":"async-result","content":"<system-notice>\nBackground job PersistentRestoreReview.RestoreDocsStandards has completed. Resume your work using the result below.\n<task-result id=\"PersistentRestoreReview.RestoreDocsStandards\" agent=\"scout\" status=\"completed\" duration=\"1m27s\">\n<meta lines=\"31\" size=\"3.8KB\" />\n<output>\n{\n  \"summary\": \"One actionable documentation defect: the amended hazard invariant conflicts with the retained, explicitly binding conformance checklist. No other issue-worthy docs/registry findings.\",\n  \"files\": [\n    {\n      \"path\": \".worktrees/287-shellrestore/docs/KNOWN-HAZARDS.md:88-92\",\n      \"description\": \"Patch-overlapping location for finding: new persistent restoration rule leaves contradictory live checklist rows at lines 1019-1021.\"\n    },\n    {\n      \"path\": \".worktrees/287-shellrestore/CONTEXT.md:391-395\",\n      \"description\": \"Correctly documents boot sweep and owner-online edge without launch-age/stamp gates; reserves freeze for nonpersistent shells.\"\n    },\n    {\n      \"path\": \".worktrees/287-shellrestore/docs/MANIFEST.md:791-795\",\n      \"description\": \"Sleep/wake and restoration prose matches operator ruling and protects live/binding binaries.\"\n    },\n    {\n      \"path\": \".worktrees/287-shellrestore/docs-site/src/shells/overview.md:98-127\",\n      \"description\": \"Documents revised restore/freeze behavior, full healed-instance enumeration, and binding relink result.\"\n    },\n    {\n      \"path\": \".worktrees/287-shellrestore/traceable-reqs.toml:4159-4186\",\n      \"description\": \"Restore and hazard titles amended; both riders registered with active stages and assigned groups (4405, 4570). Historical stale-online title explicitly superseded at 3186-3191.\"\n    },\n    {\n      \"path\": \".worktrees/287-shellrestore/CHANGELOG.md:13-24\",\n      \"description\": \"Unreleased notes cover daemon-only restoration and both riders.\"\n    }\n  ],\n  \"architecture\": \"CONTEXT defines meaning; MANIFEST and shell overview explain public behavior; KNOWN-HAZARDS supplies binding invariants/checklist; traceable-reqs registers requirements and nearby doc evidence. Review limited to source and the complete six-file documentation diff.\",\n  \"report\": \"## Finding\\n[P2] Update the binding checklist alongside the new persistent invariant — docs/KNOWN-HAZARDS.md:88-92 (changed lines).\\n\\nThe new invariant removes all machine-boot, launch-age and stamp-presence gates, but the same document's live “Conformance checklist (condensed)” still mandates at line 1021 that persistent restoration occurs only when the recorded launch predates boot and explicitly says “no stamp or no boot oracle ⇒ not restored.” Row 2.6 at line 1019 likewise retains an unqualified corpse-predates-boot/no-spontaneous-relaunch rule. These rows are not labeled historical or superseded. Updating only the main entry leaves a maintainer following the binding checklist instructed to reject exactly the same-boot/unstamped persistent cases releases#287 requires restoring. Update both checklist rows to match §§2.6–2.7.\\n\\n## Standards\\nOne concern (the finding above): AGENTS.md requires honoring every KNOWN-HAZARDS invariant and treating its checklist as conformance requirements, not advice. The contradictory checklist undermines that contract. New rider IDs, group membership and doc-evidence placement otherwise have no actionable source-level defect found. No traceability CLI coverage claim is made.\\n\\n## Spec\\nOne concern (same finding, not an additional defect): retained checklist gates contradict consequence 1 and the required documentation amendment in consequence 3 of the operator ruling: https://github.com/BigscreenVR/spt-bs-releases/issues/287#issuecomment-5583820305 . Elsewhere the changed docs cover both restoration triggers, missing launch stamps/oracle independence, nonpersistent freeze, every healed instance per owner, and binding rather than offline/premature-online relink success. The historical registry title is expressly superseded by its added amendment and was not flagged.\\n\\nVerification: read-only full documentation diff and surrounding requirements/checklist inspection. No edits, formatting, lint, build, tests, production operations, or runtime claims.\"\n}\n</output>\n</task-result>\n\nPersistentRestoreReview.RestoreDocsStandards is now idle — message it via `hub` to follow up; transcript at history://PersistentRestoreReview.RestoreDocsStandards\n\nStructured output: schema valid; full payload at agent://PersistentRestoreReview.RestoreDocsStandards, fields via agent://PersistentRestoreReview.RestoreDocsStandards?q=.<field>\n</system-notice>","display":true,"details":{"jobs":[{"jobId":"PersistentRestoreReview.RestoreDocsStandards","type":"task","label":"PersistentRestoreReview.RestoreDocsStandards","durationMs":87127,"schema":{"source":"agent","mode":"permissive","status":"valid","data":{"summary":"One actionable documentation defect: the amended hazard invariant conflicts with the retained, explicitly binding conformance checklist. No other issue-worthy docs/registry findings.","files":[{"path":".worktrees/287-shellrestore/docs/KNOWN-HAZARDS.md:88-92","description":"Patch-overlapping location for finding: new persistent restoration rule leaves contradictory live checklist rows at lines 1019-1021."},{"path":".worktrees/287-shellrestore/CONTEXT.md:391-395","description":"Correctly documents boot sweep and owner-online edge without launch-age/stamp gates; reserves freeze for nonpersistent shells."},{"path":".worktrees/287-shellrestore/docs/MANIFEST.md:791-795","description":"Sleep/wake and restoration prose matches operator ruling and protects live/binding binaries."},{"path":".worktrees/287-shellrestore/docs-site/src/shells/overview.md:98-127","description":"Documents revised restore/freeze behavior, full healed-instance enumeration, and binding relink result."},{"path":".worktrees/287-shellrestore/traceable-reqs.toml:4159-4186","description":"Restore and hazard titles amended; both riders registered with active stages and assigned groups (4405, 4570). Historical stale-online title explicitly superseded at 3186-3191."},{"path":".worktrees/287-shellrestore/CHANGELOG.md:13-24","description":"Unreleased notes cover daemon-only restoration and both riders."}],"architecture":"CONTEXT defines meaning; MANIFEST and shell overview explain public behavior; KNOWN-HAZARDS supplies binding invariants/checklist; traceable-reqs registers requirements and nearby doc evidence. Review limited to source and the complete six-file documentation diff.","report":"## Finding\n[P2] Update the binding checklist alongside the new persistent invariant — docs/KNOWN-HAZARDS.md:88-92 (changed lines).\n\nThe new invariant removes all machine-boot, launch-age and stamp-presence gates, but the same document's live “Conformance checklist (condensed)” still mandates at line 1021 that persistent restoration occurs only when the recorded launch predates boot and explicitly says “no stamp or no boot oracle ⇒ not restored.” Row 2.6 at line 1019 likewise retains an unqualified corpse-predates-boot/no-spontaneous-relaunch rule. These rows are not labeled historical or superseded. Updating only the main entry leaves a maintainer following the binding checklist instructed to reject exactly the same-boot/unstamped persistent cases releases#287 requires restoring. Update both checklist rows to match §§2.6–2.7.\n\n## Standards\nOne concern (the finding above): AGENTS.md requires honoring every KNOWN-HAZARDS invariant and treating its checklist as conformance requirements, not advice. The contradictory checklist undermines that contract. New rider IDs, group membership and doc-evidence placement otherwise have no actionable source-level defect found. No traceability CLI coverage claim is made.\n\n## Spec\nOne concern (same finding, not an additional defect): retained checklist gates contradict consequence 1 and the required documentation amendment in consequence 3 of the operator ruling: https://github.com/BigscreenVR/spt-bs-releases/issues/287#issuecomment-5583820305 . Elsewhere the changed docs cover both restoration triggers, missing launch stamps/oracle independence, nonpersistent freeze, every healed instance per owner, and binding rather than offline/premature-online relink success. The historical registry title is expressly superseded by its added amendment and was not flagged.\n\nVerification: read-only full documentation diff and surrounding requirements/checklist inspection. No edits, formatting, lint, build, tests, production operations, or runtime claims."}}}]},"attribution":"agent","id":"7d9b0f85","parentId":"4eb67798","timestamp":"2026-09-10T00:47:15.282Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_O2SO884VeDWPEXbO5rgUeSsa|fc_077f3569365e313d016aa1fe1eef5087d1bfe9cb5ccc12f1e3","toolName":"read","startedAt":"2026-09-10T00:47:27.044Z","args":{"path":".worktrees/287-shellrestore/docs/KNOWN-HAZARDS.md:70-99:raw"},"intent":"Verifying amended hazard invariant"},"id":"93b93200","parentId":"7d9b0f85","timestamp":"2026-09-10T00:47:27.044Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_x0H1V8A2lPcCPhTpAr8GI4lL|fc_077f3569365e313d016aa1fe1eef6887d1bc5f662ca2a159c3","toolName":"read","startedAt":"2026-09-10T00:47:27.044Z","args":{"path":".worktrees/287-shellrestore/docs/KNOWN-HAZARDS.md:1000-1028:raw"},"intent":"Verifying binding checklist contradiction"},"id":"04fec404","parentId":"93b93200","timestamp":"2026-09-10T00:47:27.044Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_CkWsaMCu9kwUHzW0YDa1yZ4G|fc_077f3569365e313d016aa1fe1eef7087d1acab7dfb4d07f1b7","toolName":"read","startedAt":"2026-09-10T00:47:27.045Z","args":{"path":"artifact://112:930-1088:raw"},"intent":"Reading remaining patch details"},"id":"56e9ce5b","parentId":"04fec404","timestamp":"2026-09-10T00:47:27.045Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_NSi9uK6PO9CCco5KJs6mmQpR|fc_077f3569365e313d016aa1fe1eef7887d1b0a154249fe945ab","toolName":"read","startedAt":"2026-09-10T00:47:27.045Z","args":{"path":".worktrees/287-shellrestore/.git"},"intent":"Pinning reviewed commit identity"},"id":"87377011","parentId":"56e9ce5b","timestamp":"2026-09-10T00:47:27.045Z"}
{"type":"message","id":"f292f403","parentId":"87377011","timestamp":"2026-09-10T00:47:27.045Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":"**Reviewing docs and checklist**"},{"type":"toolCall","id":"call_O2SO884VeDWPEXbO5rgUeSsa|fc_077f3569365e313d016aa1fe1eef5087d1bfe9cb5ccc12f1e3","name":"read","arguments":{"i":"Verifying amended hazard invariant","path":".worktrees/287-shellrestore/docs/KNOWN-HAZARDS.md:70-99:raw"},"intent":"Verifying amended hazard invariant"},{"type":"toolCall","id":"call_x0H1V8A2lPcCPhTpAr8GI4lL|fc_077f3569365e313d016aa1fe1eef6887d1bc5f662ca2a159c3","name":"read","arguments":{"i":"Verifying binding checklist contradiction","path":".worktrees/287-shellrestore/docs/KNOWN-HAZARDS.md:1000-1028:raw"},"intent":"Verifying binding checklist contradiction"},{"type":"toolCall","id":"call_CkWsaMCu9kwUHzW0YDa1yZ4G|fc_077f3569365e313d016aa1fe1eef7087d1acab7dfb4d07f1b7","name":"read","arguments":{"i":"Reading remaining patch details","path":"artifact://112:930-1088:raw"},"intent":"Reading remaining patch details"},{"type":"toolCall","id":"call_NSi9uK6PO9CCco5KJs6mmQpR|fc_077f3569365e313d016aa1fe1eef7887d1b0a154249fe945ab","name":"read","arguments":{"i":"Pinning reviewed commit identity","path":".worktrees/287-shellrestore/.git"},"intent":"Pinning reviewed commit identity"},{"type":"toolCall","id":"call_DLi0WVShUA9eTpzx0LdLbfIM|fc_077f3569365e313d016aa1fe1eef8087d184158b76da4c7c44","name":"bash","arguments":{"i":"Reading reviewed commit metadata","command":"git log -1 --format=fuller cb756dee","cwd":"C:/Users/decid/Documents/projects/spt-core/.worktrees/287-shellrestore"},"intent":"Reading reviewed commit metadata"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":9769,"output":232,"cacheRead":73216,"cacheWrite":0,"totalTokens":83217,"reasoningTokens":15,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1789001235320,"responseId":"resp_077f3569365e313d016aa1fe13915887d1a5b297cdc00717eb","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_077f3569365e313d016aa1fe17faf087d1b26aae2cf5051e55","type":"reasoning","content":[],"encrypted_content":"gAAAAABqof4e5hXEslzj1W_qHkDao8D5ZOrk4KtqjD6ROX3u3mnnwXZmcJ1favzVpkuE8WbcsjWSLfU2V4N8IsFAE2XM8lYorRlt0upxpruhHcej6XqLJ5e7CNlKbc7qiurE9BBrbxYLE2jUTXoHRe3Uk28iN6IUPjnjJ-08BF0W7-0wVxlXiGqPGGGPYtkAPzDgMsG4B5reII-8y2hUY5auJ9kkpcxHNoePaczb08rdohl9XWRsmabBAES8guZhkRQXMJM8RkWUptgaAfUqTCYrVvjA7IPKa56DaQ-YY13ShWAjT-ucT9iJTpbuh-21CHSvAVf7huPkwL78znxMk2WKCAZjC8JTdUxyEeGqaPr62fw6OSQ2EaxEjz4ZLXVJhUf17eFmVzROABAZyUs7xYwSEa-uXj9wd1wqPLw3vtgVzUNnaR4ObuvoV8U-i2mcI-6bP2WphRi8g-dkss8yWhwc49du_BqglDhYNPvlaT2_rVHZy3LEF1v33uMYKieK8f8OpBpcGSGIx4qaW996rL-Wv3tufKYVvlHqXhtrilcwhvbUlqx2y08lH2d_tUhISFd4dvW_DSsM0gJB4UD2biWHGI5YiLmaIXeY2yK5BJq8UCXLcrujMV1tdOPe0lyr48P958wnpymVLZWQDNQwe6pomH2atqEj8YiOjBkaGEisw6qXyuYNXIB8AzL_0sJOeDBU1ANbPcagfE1bEPjXgb2IYWjrcZd24Ui0YmpyJbMMeqSPQ6pji6VyRl-H9fzsH6oFg2EJOWaj7wDoOAf7mO_Qo1v-3znCm8TR8opmEBEI-jig77Kv_gkg2nNbDu-yZSTot6zkTYljElDSu4DXcGEOHJvAF-IM8vcL4Lx_4-wShRTdXeSvWiBvON8f-56ZU6MapTg50LY6DjGXsedxU4rdr4BV7Y7Sia9Zf9RVZiLPX4SGOKVnEjKvuTLi-3stx1ODwo-nONcq7s5CIifnD6N-91553g7Br22yrKyesYYFYt-WxD9TOR8bkTL-C4Lhxl_dbYfL9CbOYCAt42H7AYjLLOZ9FSZRqYhgU5nF0D0visQzP97w_t9d6MArWGyIDeKavB-Wa0Q4-FJwozkp2bzwalISusDhDzVefk-on36ZfqF2JTV4W12yMWgNBs27F_uyzdmFTqkPvB6RNtFx9n6jxiecYxxoVCZQpTsdR33jmS1EaPhJ_wX7G2E5tLHpwU-K_Bw0fXRhXw81BEBexoEzzY1mYYgLn8DL0wwV6vfJLKU1RuFD_w_kqHUHk6-lbodntci_xthRmFLYD2qDmB4kSEyQch0p_mLTdwhlh1PRYbaiW8_fS7CTlp5oTpf6ytI3lE4H3IfsHtG88LsvRHZm8LFnKQFLjA==","internal_chat_message_metadata_passthrough":{"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"},"summary":[{"type":"summary_text","text":"**Reviewing docs and checklist**"}],"metadata":{"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"}},{"id":"fc_077f3569365e313d016aa1fe1eef5087d1bfe9cb5ccc12f1e3","type":"function_call","status":"completed","arguments":"{\"i\":\"Verifying amended hazard invariant\",\"path\":\".worktrees/287-shellrestore/docs/KNOWN-HAZARDS.md:70-99:raw\"}","call_id":"call_O2SO884VeDWPEXbO5rgUeSsa","internal_chat_message_metadata_passthrough":{"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"},"name":"read","metadata":{"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"}},{"id":"fc_077f3569365e313d016aa1fe1eef6887d1bc5f662ca2a159c3","type":"function_call","status":"completed","arguments":"{\"i\":\"Verifying binding checklist contradiction\",\"path\":\".worktrees/287-shellrestore/docs/KNOWN-HAZARDS.md:1000-1028:raw\"}","call_id":"call_x0H1V8A2lPcCPhTpAr8GI4lL","internal_chat_message_metadata_passthrough":{"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"},"name":"read","metadata":{"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"}},{"id":"fc_077f3569365e313d016aa1fe1eef7087d1acab7dfb4d07f1b7","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading remaining patch details\",\"path\":\"artifact://112:930-1088:raw\"}","call_id":"call_CkWsaMCu9kwUHzW0YDa1yZ4G","internal_chat_message_metadata_passthrough":{"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"},"name":"read","metadata":{"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"}},{"id":"fc_077f3569365e313d016aa1fe1eef7887d1b0a154249fe945ab","type":"function_call","status":"completed","arguments":"{\"i\":\"Pinning reviewed commit identity\",\"path\":\".worktrees/287-shellrestore/.git\"}","call_id":"call_NSi9uK6PO9CCco5KJs6mmQpR","internal_chat_message_metadata_passthrough":{"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"},"name":"read","metadata":{"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"}},{"id":"fc_077f3569365e313d016aa1fe1eef8087d184158b76da4c7c44","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading reviewed commit metadata\",\"command\":\"git log -1 --format=fuller cb756dee\",\"cwd\":\"C:/Users/decid/Documents/projects/spt-core/.worktrees/287-shellrestore\"}","call_id":"call_DLi0WVShUA9eTpzx0LdLbfIM","internal_chat_message_metadata_passthrough":{"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"},"name":"bash","metadata":{"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"}}]},"duration":11720.936200000346,"ttft":4688.136900000274,"completedAt":1789001247042,"contextSnapshot":{"promptTokens":82985,"nonMessageTokens":13209,"compactionEpoch":0}}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_DLi0WVShUA9eTpzx0LdLbfIM|fc_077f3569365e313d016aa1fe1eef8087d184158b76da4c7c44","toolName":"bash","startedAt":"2026-09-10T00:47:27.045Z","args":{"command":"git log -1 --format=fuller cb756dee"},"intent":"Reading reviewed commit metadata"},"id":"cff2d14f","parentId":"f292f403","timestamp":"2026-09-10T00:47:27.045Z"}
{"type":"message","id":"c6a60f9c","parentId":"cff2d14f","timestamp":"2026-09-10T00:47:27.054Z","message":{"role":"toolResult","toolCallId":"call_O2SO884VeDWPEXbO5rgUeSsa|fc_077f3569365e313d016aa1fe1eef5087d1bfe9cb5ccc12f1e3","toolName":"read","content":[{"type":"text","text":"### 2.5 Daemon-hosted endpoints have no dedicated liveness PID\n- **Failure:** the sister evaluates Psyche/perch liveness via a dedicated process PID — the wrapper's own pid in `info.json`, checked with `is_process_alive`. Under ADR-0004 the Psyche (and any spt-hosted Self) is a **loop inside the daemon**, not a separate process: it holds no dedicated pid, and its `claude`/summarizer subprocess is ephemeral (spawned per pulse/commune, then exits). If a daemon-hosted perch's `info.json` carries the **daemon's** pid, then *every* hosted endpoint shares one pid, and `is_process_alive(pid)` reads \"alive\" for a torn-down endpoint as long as the daemon runs — while `clean_stale_entries` (dead-pid deletion) can no longer distinguish a dead endpoint from a live one. The 2.1/5.1 liveness models do **not** cover this third category: the Psyche is neither harness-hosted (no `parent_pid` anchor) nor a broker-held PTY child.\n- **Invariant:** for **daemon-hosted** perches (Psyche; spt-hosted Self), liveness is the **daemon's authoritative in-memory endpoint table + a `status` field** on `info.json` (`online|offline|…`), **never** `is_process_alive(info.pid)`. `info.pid` for a daemon-hosted perch is at most a *hosted-by-daemon* marker (the daemon pid), not a liveness signal; registry stale-clean for these rows keys on the daemon's endpoint table, not per-row pid. This reuses the pattern already specified for **Shells** (`info.json` carries daemon-managed `status`, capability resolved by `adapter_name` — CONTEXT \"Shell… Not in the subnet registry\") and extends it to daemon-hosted *agent* perches.\n- **spt-core mapping:** the **M1/M2a interim** model keeps the Psyche/listener a real per-process owner (the `api listen` process), so its per-pid liveness (`deliver::is_online` → `info.read_pid` → `proc::is_process_alive`; `registry::clean_stale_entries`) is correct *interim*. **M3 daemon consolidation replaces it** with daemon-authoritative liveness for hosted perches. Keep the liveness check behind one resolver (mirrors `resolve_address` stale-clean) so the M3 swap is localized — do **not** let the per-pid assumption leak into new call sites.\n- **Sister cite:** `src/live/wrapper/orphan.rs` (wrapper-pid liveness); `src/common/list_filter.rs:168-175` (pid-classify); spt-core `crates/spt-store/src/{proc.rs,registry.rs}` + `crates/spt-msg/src/deliver.rs::is_online`.\n\n### 2.6 A shell instance's `online` is a recorded field with no death edge\n- **Failure:** a shell perch's `status` has exactly two writers — the bind handshake (→`online`) and `close_shell` (→`offline`) — and `close_shell` runs on the **link-break** path alone. A binary that dies **abruptly** (force-kill, crash, OOM) breaks no link, so nothing ever writes the flip: the record says `online` forever while the pid it names is gone. Every consumer of the recorded field then inherits the lie — `shell list` reports online indefinitely; `relink` (the one command that recovers the instance) refuses `SHELL_ALREADY_ONLINE`, gated on the very state that is wrong; and `shell cmd`'s wake-if-offline arm reads \"not offline\", skips the wake, and spools the command to a corpse — accepted happily, drained by nobody, **indistinguishable from a busy shell**. The only exit was `teardown` + `spawn`, which destroys the perch (losing the instance's persisted state) and frees the mint slot, so the instance returns under a **different canonical id** and every reference to the old name — docs, other agents' notes, tag addressing — silently points at nothing. The rename is the *visible* cost; the expensive one is **silent**: a consumer's perch-persisted cursor re-baselines on the fresh perch, so work arriving between the death and the re-bind is never scanned — not failed-and-retried, just never seen (alchemy's tag scanner, 2026-07-25). Recovery must therefore preserve the perch, not merely restore a running binary.\n- **Invariant:** ONLINE-ness is **derived**, never the recorded field alone: `status == online` **AND** the recorded `shell.pid` is not *provably* dead. Provably dead is narrow — pid parked, non-zero, `!is_process_alive`; pid absent, unparseable, or `0` (a broker-hosted spawn whose backend exposed no pid) reads **alive**, the same fail-toward-alive parity 2.5 holds, so a backend the resolver cannot see is never declared dead. **The recycled-pid caveat is narrowed since BAROMETER W2 and is now conditional on the stamp:** a launch parks a **birth stamp** beside the pid (`shell.launch.json`), and the probe runs the PAIR test — same pid, *different* start time ⇒ **gone**, so a pid the OS handed to an unrelated process no longer masks a dead instance. That case matters most across a restart, where the pid namespace resets and a corpse pid is routinely reused; without the pair test the corpse reads alive, the derived status reads `online`, and `relink` refuses `SHELL_ALREADY_ONLINE` on the one instance needing recovery. Where there is **no stamp to compare** (an instance last launched by a pre-W2 build, or a backend exposing no start time) the original stance is unchanged and unchanged in *direction*: fail toward alive, heal *missed* rather than *mis-fired*. Derivation NEVER promotes: a recorded-`offline` instance stays offline whatever its pid file says. Site class is part of the invariant: **derive** at the gates and renders (relink, the cmd wake arm, drive's drop-if-offline, `shelldisc::discover` = the single source of both `shell list` surfaces, the activity fan-out); keep the **recorded** field at the two writers and at the owner-suspend cascade, whose `close_shell` **is** the cleanup that would be skipped; and keep it at the wake reconciler's eligibility read, which since BAROMETER W2 additionally asks the **corpse-boot discriminant** (releases#78, doyle's amendment on comment 5156810267): the read stays on the recorded field, but \"no spontaneous relaunch\" is now a stated predicate rather than a side effect of the record being wrong — see the bullet below.\n- **No spontaneous relaunch of a NONPERSISTENT same-boot death:** the releases#78 force-kill freeze still protects nonpersistent instances, so killing one to free its exe does not arm a watcher underneath an overwrite. Recovery remains demand-driven. **Persistent instances are exempt** (operator ruling, releases#287): a kill is not a request to keep a declared-persistent shell down; use `teardown` or a nonpersistent manifest instead.\n- **Watcher eligibility:** recorded `offline`, no live binary awaiting bind, and either `persistent`, no corpse, or a corpse whose launch predates the machine boot instant. A nonpersistent corpse without a launch stamp or boot oracle remains frozen; persistent instances require neither to become eligible. Birth-safe probes prevent a recycled pid from masquerading as a live shell.\n- **spt-core mapping:** `spt_store::shellinfo::{shell_pid_provably_dead, effective_status, is_shell_online}` (the resolver — the shell-side twin of 2.5's `liveness.rs`, which gave *agent* perches exactly this and which shells never got); consumers in `linkhost::{relink_shell, prepare_drive, run_action}`, `shelldisc::discover`, `activity::observe_links`; the class-(c) eligibility read is `spt_daemon::shellwake::watcher_eligible` (+ `launch_predates_boot`, `BOOT_RESTORE_SLACK_MS`), and the birth stamp it rests on is `shellinfo::{ShellLaunch, record_shell_launch, read_shell_launch}`.\n- **Source:** spt-core, flynn's spt-alchemy field report (2026-07-25) — a deterministic recipe, not a race. Class-(c) re-based BAROMETER W2 (releases#78, doyle's ruling on comment 5156810267).\n\n### 2.7 A daemon or machine restart strands `persistent` shells — and a graceful cascade fixture suppresses the failing arm `[REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]`\n- **Failure:** a restart can leave an `online` record over a corpse without a link-break or suspend edge. The original boot sweep fixed only machine reboots: it rejected corpses launched after machine boot, so a daemon-only restart still stranded every persistent shell. The measured releases#287 specimen launched at machine boot +46s and stayed down for 27 hours after a daemon restart despite its owner being online.\n- **The owner-facing surfaces HIDE the fault**, which is why it survived unreported for days: every display path (`shell list`, its `--json` twin, the shell-context render) routes through the ONE discovery seam, which **derives** status (2.6) and so correctly reports the binary as offline — while the only consumer whose decision matters, the wake cascade, reads the RECORDED field that no owner-facing surface shows. Absence of a visible symptom is therefore **not** evidence of absence of this fault, and a clean-looking context render must never be taken as proof the class did not occur. The correct falsifier is the **on-disk record**.\n- **The existing cascade test cannot fail on it — its SETUP SUPPRESSES THE FAILING ARM.** `rest_edges_cascade_shells_with_divergence` establishes its fixture by driving the suspend path, which itself writes the `offline` the code under test is supposed to encounter — so the restart shape (a record still saying `online` over a corpse, reached with **no** suspend edge) is unreachable from that fixture BY CONSTRUCTION. The test is not weak; it gives a true answer to the wrong question, so a reader asking \"is the cascade covered?\" is told yes. This is the third suppressed-arm instance found in one week, and the pattern is the same each time: **a fixture that establishes the precondition the code under test is supposed to establish.**\n- **Invariant:** a daemon-only or machine restart must not strand a persistent instance. **(a) Heal:** every reconcile cycle writes `offline` only where the derived birth-safe read contradicts recorded `online`; already-correct records are untouched. **(b) Restore:** the once-per-daemon-generation boot sweep and every owner offline→online edge run the same per-owner body. Relaunch every instance whose adapter declares `persistent`, whose owner is online, and whose binary is down in fact. There is **no machine-boot, launch-age, or stamp-presence eligibility gate**. A live binary awaiting bind is not down and must not be duplicated; missing birth evidence keeps the existing fail-toward-alive liveness behavior. The owner-online trigger remains an edge, bounding failed launches to one attempt per event.\n- **Regression shape:** construct recorded `online` over a corpse **without a suspend edge**, including a launch **after machine boot**, then bring its owner online after a spent boot sweep. Every stranded persistent sibling must be restored; offline owners and nonpersistent instances are not restored. Separately defend live/binding exclusion and the nonpersistent watcher freeze. Suspending in setup suppresses the failing arm and proves nothing about this class.\n- **Per-owner enumeration:** the existing `SHELL_RECORD_HEALED:<owner>/<shell_id>` daemon diagnostic names **every** successfully healed instance, not one representative. Healthy and already-healed records produce no event. The restore diagnostics likewise identify each attempted instance and distinguish boot from owner-online triggers. <!-- [doc->REQ-SHELL-HEAL-ENUMERATION] -->\n- **spt-core mapping:** `spt_daemon::shellwake::{heal_stale_online_records, restore_persistent_shells_at_boot, restore_persistent_shells_on_owner_online, OwnerOnlineEdge}`; `shellhost::live_launch_winner`; `spt_store::shellinfo::{ShellLaunch, shell_pid_provably_dead}`. `launch_predates_boot` and `BOOT_RESTORE_SLACK_MS` now govern only nonpersistent watcher eligibility.\n- **Source:** BAROMETER W2 (releases#78); releases#287 operator ruling (2026-09-08), superseding the persistent force-kill freeze and boot-instant conjunct.\n<!-- [doc->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS] -->\n\n---\n\n## 3. Lifecycle\n\n### 3.1 Ephemeral perch cleanup on every `ring` exit path\n\n[Showing lines 70-99 of 1093. Use :100 to continue]"}],"details":{"truncation":{"content":"### 2.5 Daemon-hosted endpoints have no dedicated liveness PID\r\n- **Failure:** the sister evaluates Psyche/perch liveness via a dedicated process PID — the wrapper's own pid in `info.json`, checked with `is_process_alive`. Under ADR-0004 the Psyche (and any spt-hosted Self) is a **loop inside the daemon**, not a separate process: it holds no dedicated pid, and its `claude`/summarizer subprocess is ephemeral (spawned per pulse/commune, then exits). If a daemon-hosted perch's `info.json` carries the **daemon's** pid, then *every* hosted endpoint shares one pid, and `is_process_alive(pid)` reads \"alive\" for a torn-down endpoint as long as the daemon runs — while `clean_stale_entries` (dead-pid deletion) can no longer distinguish a dead endpoint from a live one. The 2.1/5.1 liveness models do **not** cover this third category: the Psyche is neither harness-hosted (no `parent_pid` anchor) nor a broker-held PTY child.\r\n- **Invariant:** for **daemon-hosted** perches (Psyche; spt-hosted Self), liveness is the **daemon's authoritative in-memory endpoint table + a `status` field** on `info.json` (`online|offline|…`), **never** `is_process_alive(info.pid)`. `info.pid` for a daemon-hosted perch is at most a *hosted-by-daemon* marker (the daemon pid), not a liveness signal; registry stale-clean for these rows keys on the daemon's endpoint table, not per-row pid. This reuses the pattern already specified for **Shells** (`info.json` carries daemon-managed `status`, capability resolved by `adapter_name` — CONTEXT \"Shell… Not in the subnet registry\") and extends it to daemon-hosted *agent* perches.\r\n- **spt-core mapping:** the **M1/M2a interim** model keeps the Psyche/listener a real per-process owner (the `api listen` process), so its per-pid liveness (`deliver::is_online` → `info.read_pid` → `proc::is_process_alive`; `registry::clean_stale_entries`) is correct *interim*. **M3 daemon consolidation replaces it** with daemon-authoritative liveness for hosted perches. Keep the liveness check behind one resolver (mirrors `resolve_address` stale-clean) so the M3 swap is localized — do **not** let the per-pid assumption leak into new call sites.\r\n- **Sister cite:** `src/live/wrapper/orphan.rs` (wrapper-pid liveness); `src/common/list_filter.rs:168-175` (pid-classify); spt-core `crates/spt-store/src/{proc.rs,registry.rs}` + `crates/spt-msg/src/deliver.rs::is_online`.\r\n\r\n### 2.6 A shell instance's `online` is a recorded field with no death edge\r\n- **Failure:** a shell perch's `status` has exactly two writers — the bind handshake (→`online`) and `close_shell` (→`offline`) — and `close_shell` runs on the **link-break** path alone. A binary that dies **abruptly** (force-kill, crash, OOM) breaks no link, so nothing ever writes the flip: the record says `online` forever while the pid it names is gone. Every consumer of the recorded field then inherits the lie — `shell list` reports online indefinitely; `relink` (the one command that recovers the instance) refuses `SHELL_ALREADY_ONLINE`, gated on the very state that is wrong; and `shell cmd`'s wake-if-offline arm reads \"not offline\", skips the wake, and spools the command to a corpse — accepted happily, drained by nobody, **indistinguishable from a busy shell**. The only exit was `teardown` + `spawn`, which destroys the perch (losing the instance's persisted state) and frees the mint slot, so the instance returns under a **different canonical id** and every reference to the old name — docs, other agents' notes, tag addressing — silently points at nothing. The rename is the *visible* cost; the expensive one is **silent**: a consumer's perch-persisted cursor re-baselines on the fresh perch, so work arriving between the death and the re-bind is never scanned — not failed-and-retried, just never seen (alchemy's tag scanner, 2026-07-25). Recovery must therefore preserve the perch, not merely restore a running binary.\r\n- **Invariant:** ONLINE-ness is **derived**, never the recorded field alone: `status == online` **AND** the recorded `shell.pid` is not *provably* dead. Provably dead is narrow — pid parked, non-zero, `!is_process_alive`; pid absent, unparseable, or `0` (a broker-hosted spawn whose backend exposed no pid) reads **alive**, the same fail-toward-alive parity 2.5 holds, so a backend the resolver cannot see is never declared dead. **The recycled-pid caveat is narrowed since BAROMETER W2 and is now conditional on the stamp:** a launch parks a **birth stamp** beside the pid (`shell.launch.json`), and the probe runs the PAIR test — same pid, *different* start time ⇒ **gone**, so a pid the OS handed to an unrelated process no longer masks a dead instance. That case matters most across a restart, where the pid namespace resets and a corpse pid is routinely reused; without the pair test the corpse reads alive, the derived status reads `online`, and `relink` refuses `SHELL_ALREADY_ONLINE` on the one instance needing recovery. Where there is **no stamp to compare** (an instance last launched by a pre-W2 build, or a backend exposing no start time) the original stance is unchanged and unchanged in *direction*: fail toward alive, heal *missed* rather than *mis-fired*. Derivation NEVER promotes: a recorded-`offline` instance stays offline whatever its pid file says. Site class is part of the invariant: **derive** at the gates and renders (relink, the cmd wake arm, drive's drop-if-offline, `shelldisc::discover` = the single source of both `shell list` surfaces, the activity fan-out); keep the **recorded** field at the two writers and at the owner-suspend cascade, whose `close_shell` **is** the cleanup that would be skipped; and keep it at the wake reconciler's eligibility read, which since BAROMETER W2 additionally asks the **corpse-boot discriminant** (releases#78, doyle's amendment on comment 5156810267): the read stays on the recorded field, but \"no spontaneous relaunch\" is now a stated predicate rather than a side effect of the record being wrong — see the bullet below.\r\n- **No spontaneous relaunch of a NONPERSISTENT same-boot death:** the releases#78 force-kill freeze still protects nonpersistent instances, so killing one to free its exe does not arm a watcher underneath an overwrite. Recovery remains demand-driven. **Persistent instances are exempt** (operator ruling, releases#287): a kill is not a request to keep a declared-persistent shell down; use `teardown` or a nonpersistent manifest instead.\r\n- **Watcher eligibility:** recorded `offline`, no live binary awaiting bind, and either `persistent`, no corpse, or a corpse whose launch predates the machine boot instant. A nonpersistent corpse without a launch stamp or boot oracle remains frozen; persistent instances require neither to become eligible. Birth-safe probes prevent a recycled pid from masquerading as a live shell.\r\n- **spt-core mapping:** `spt_store::shellinfo::{shell_pid_provably_dead, effective_status, is_shell_online}` (the resolver — the shell-side twin of 2.5's `liveness.rs`, which gave *agent* perches exactly this and which shells never got); consumers in `linkhost::{relink_shell, prepare_drive, run_action}`, `shelldisc::discover`, `activity::observe_links`; the class-(c) eligibility read is `spt_daemon::shellwake::watcher_eligible` (+ `launch_predates_boot`, `BOOT_RESTORE_SLACK_MS`), and the birth stamp it rests on is `shellinfo::{ShellLaunch, record_shell_launch, read_shell_launch}`.\r\n- **Source:** spt-core, flynn's spt-alchemy field report (2026-07-25) — a deterministic recipe, not a race. Class-(c) re-based BAROMETER W2 (releases#78, doyle's ruling on comment 5156810267).\r\n\r\n### 2.7 A daemon or machine restart strands `persistent` shells — and a graceful cascade fixture suppresses the failing arm `[REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]`\r\n- **Failure:** a restart can leave an `online` record over a corpse without a link-break or suspend edge. The original boot sweep fixed only machine reboots: it rejected corpses launched after machine boot, so a daemon-only restart still stranded every persistent shell. The measured releases#287 specimen launched at machine boot +46s and stayed down for 27 hours after a daemon restart despite its owner being online.\r\n- **The owner-facing surfaces HIDE the fault**, which is why it survived unreported for days: every display path (`shell list`, its `--json` twin, the shell-context render) routes through the ONE discovery seam, which **derives** status (2.6) and so correctly reports the binary as offline — while the only consumer whose decision matters, the wake cascade, reads the RECORDED field that no owner-facing surface shows. Absence of a visible symptom is therefore **not** evidence of absence of this fault, and a clean-looking context render must never be taken as proof the class did not occur. The correct falsifier is the **on-disk record**.\r\n- **The existing cascade test cannot fail on it — its SETUP SUPPRESSES THE FAILING ARM.** `rest_edges_cascade_shells_with_divergence` establishes its fixture by driving the suspend path, which itself writes the `offline` the code under test is supposed to encounter — so the restart shape (a record still saying `online` over a corpse, reached with **no** suspend edge) is unreachable from that fixture BY CONSTRUCTION. The test is not weak; it gives a true answer to the wrong question, so a reader asking \"is the cascade covered?\" is told yes. This is the third suppressed-arm instance found in one week, and the pattern is the same each time: **a fixture that establishes the precondition the code under test is supposed to establish.**\r\n- **Invariant:** a daemon-only or machine restart must not strand a persistent instance. **(a) Heal:** every reconcile cycle writes `offline` only where the derived birth-safe read contradicts recorded `online`; already-correct records are untouched. **(b) Restore:** the once-per-daemon-generation boot sweep and every owner offline→online edge run the same per-owner body. Relaunch every instance whose adapter declares `persistent`, whose owner is online, and whose binary is down in fact. There is **no machine-boot, launch-age, or stamp-presence eligibility gate**. A live binary awaiting bind is not down and must not be duplicated; missing birth evidence keeps the existing fail-toward-alive liveness behavior. The owner-online trigger remains an edge, bounding failed launches to one attempt per event.\r\n- **Regression shape:** construct recorded `online` over a corpse **without a suspend edge**, including a launch **after machine boot**, then bring its owner online after a spent boot sweep. Every stranded persistent sibling must be restored; offline owners and nonpersistent instances are not restored. Separately defend live/binding exclusion and the nonpersistent watcher freeze. Suspending in setup suppresses the failing arm and proves nothing about this class.\r\n- **Per-owner enumeration:** the existing `SHELL_RECORD_HEALED:<owner>/<shell_id>` daemon diagnostic names **every** successfully healed instance, not one representative. Healthy and already-healed records produce no event. The restore diagnostics likewise identify each attempted instance and distinguish boot from owner-online triggers. <!-- [doc->REQ-SHELL-HEAL-ENUMERATION] -->\r\n- **spt-core mapping:** `spt_daemon::shellwake::{heal_stale_online_records, restore_persistent_shells_at_boot, restore_persistent_shells_on_owner_online, OwnerOnlineEdge}`; `shellhost::live_launch_winner`; `spt_store::shellinfo::{ShellLaunch, shell_pid_provably_dead}`. `launch_predates_boot` and `BOOT_RESTORE_SLACK_MS` now govern only nonpersistent watcher eligibility.\r\n- **Source:** BAROMETER W2 (releases#78); releases#287 operator ruling (2026-09-08), superseding the persistent force-kill freeze and boot-instant conjunct.\r\n<!-- [doc->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS] -->\r\n\r\n---\r\n\r\n## 3. Lifecycle\r\n\r\n### 3.1 Ephemeral perch cleanup on every `ring` exit path\r","truncated":true,"truncatedBy":"lines","totalLines":1024,"totalBytes":11985,"outputLines":30,"outputBytes":11985,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":1093,"displayContent":{"text":"### 2.5 Daemon-hosted endpoints have no dedicated liveness PID\r\n- **Failure:** the sister evaluates Psyche/perch liveness via a dedicated process PID — the wrapper's own pid in `info.json`, checked with `is_process_alive`. Under ADR-0004 the Psyche (and any spt-hosted Self) is a **loop inside the daemon**, not a separate process: it holds no dedicated pid, and its `claude`/summarizer subprocess is ephemeral (spawned per pulse/commune, then exits). If a daemon-hosted perch's `info.json` carries the **daemon's** pid, then *every* hosted endpoint shares one pid, and `is_process_alive(pid)` reads \"alive\" for a torn-down endpoint as long as the daemon runs — while `clean_stale_entries` (dead-pid deletion) can no longer distinguish a dead endpoint from a live one. The 2.1/5.1 liveness models do **not** cover this third category: the Psyche is neither harness-hosted (no `parent_pid` anchor) nor a broker-held PTY child.\r\n- **Invariant:** for **daemon-hosted** perches (Psyche; spt-hosted Self), liveness is the **daemon's authoritative in-memory endpoint table + a `status` field** on `info.json` (`online|offline|…`), **never** `is_process_alive(info.pid)`. `info.pid` for a daemon-hosted perch is at most a *hosted-by-daemon* marker (the daemon pid), not a liveness signal; registry stale-clean for these rows keys on the daemon's endpoint table, not per-row pid. This reuses the pattern already specified for **Shells** (`info.json` carries daemon-managed `status`, capability resolved by `adapter_name` — CONTEXT \"Shell… Not in the subnet registry\") and extends it to daemon-hosted *agent* perches.\r\n- **spt-core mapping:** the **M1/M2a interim** model keeps the Psyche/listener a real per-process owner (the `api listen` process), so its per-pid liveness (`deliver::is_online` → `info.read_pid` → `proc::is_process_alive`; `registry::clean_stale_entries`) is correct *interim*. **M3 daemon consolidation replaces it** with daemon-authoritative liveness for hosted perches. Keep the liveness check behind one resolver (mirrors `resolve_address` stale-clean) so the M3 swap is localized — do **not** let the per-pid assumption leak into new call sites.\r\n- **Sister cite:** `src/live/wrapper/orphan.rs` (wrapper-pid liveness); `src/common/list_filter.rs:168-175` (pid-classify); spt-core `crates/spt-store/src/{proc.rs,registry.rs}` + `crates/spt-msg/src/deliver.rs::is_online`.\r\n\r\n### 2.6 A shell instance's `online` is a recorded field with no death edge\r\n- **Failure:** a shell perch's `status` has exactly two writers — the bind handshake (→`online`) and `close_shell` (→`offline`) — and `close_shell` runs on the **link-break** path alone. A binary that dies **abruptly** (force-kill, crash, OOM) breaks no link, so nothing ever writes the flip: the record says `online` forever while the pid it names is gone. Every consumer of the recorded field then inherits the lie — `shell list` reports online indefinitely; `relink` (the one command that recovers the instance) refuses `SHELL_ALREADY_ONLINE`, gated on the very state that is wrong; and `shell cmd`'s wake-if-offline arm reads \"not offline\", skips the wake, and spools the command to a corpse — accepted happily, drained by nobody, **indistinguishable from a busy shell**. The only exit was `teardown` + `spawn`, which destroys the perch (losing the instance's persisted state) and frees the mint slot, so the instance returns under a **different canonical id** and every reference to the old name — docs, other agents' notes, tag addressing — silently points at nothing. The rename is the *visible* cost; the expensive one is **silent**: a consumer's perch-persisted cursor re-baselines on the fresh perch, so work arriving between the death and the re-bind is never scanned — not failed-and-retried, just never seen (alchemy's tag scanner, 2026-07-25). Recovery must therefore preserve the perch, not merely restore a running binary.\r\n- **Invariant:** ONLINE-ness is **derived**, never the recorded field alone: `status == online` **AND** the recorded `shell.pid` is not *provably* dead. Provably dead is narrow — pid parked, non-zero, `!is_process_alive`; pid absent, unparseable, or `0` (a broker-hosted spawn whose backend exposed no pid) reads **alive**, the same fail-toward-alive parity 2.5 holds, so a backend the resolver cannot see is never declared dead. **The recycled-pid caveat is narrowed since BAROMETER W2 and is now conditional on the stamp:** a launch parks a **birth stamp** beside the pid (`shell.launch.json`), and the probe runs the PAIR test — same pid, *different* start time ⇒ **gone**, so a pid the OS handed to an unrelated process no longer masks a dead instance. That case matters most across a restart, where the pid namespace resets and a corpse pid is routinely reused; without the pair test the corpse reads alive, the derived status reads `online`, and `relink` refuses `SHELL_ALREADY_ONLINE` on the one instance needing recovery. Where there is **no stamp to compare** (an instance last launched by a pre-W2 build, or a backend exposing no start time) the original stance is unchanged and unchanged in *direction*: fail toward alive, heal *missed* rather than *mis-fired*. Derivation NEVER promotes: a recorded-`offline` instance stays offline whatever its pid file says. Site class is part of the invariant: **derive** at the gates and renders (relink, the cmd wake arm, drive's drop-if-offline, `shelldisc::discover` = the single source of both `shell list` surfaces, the activity fan-out); keep the **recorded** field at the two writers and at the owner-suspend cascade, whose `close_shell` **is** the cleanup that would be skipped; and keep it at the wake reconciler's eligibility read, which since BAROMETER W2 additionally asks the **corpse-boot discriminant** (releases#78, doyle's amendment on comment 5156810267): the read stays on the recorded field, but \"no spontaneous relaunch\" is now a stated predicate rather than a side effect of the record being wrong — see the bullet below.\r\n- **No spontaneous relaunch of a NONPERSISTENT same-boot death:** the releases#78 force-kill freeze still protects nonpersistent instances, so killing one to free its exe does not arm a watcher underneath an overwrite. Recovery remains demand-driven. **Persistent instances are exempt** (operator ruling, releases#287): a kill is not a request to keep a declared-persistent shell down; use `teardown` or a nonpersistent manifest instead.\r\n- **Watcher eligibility:** recorded `offline`, no live binary awaiting bind, and either `persistent`, no corpse, or a corpse whose launch predates the machine boot instant. A nonpersistent corpse without a launch stamp or boot oracle remains frozen; persistent instances require neither to become eligible. Birth-safe probes prevent a recycled pid from masquerading as a live shell.\r\n- **spt-core mapping:** `spt_store::shellinfo::{shell_pid_provably_dead, effective_status, is_shell_online}` (the resolver — the shell-side twin of 2.5's `liveness.rs`, which gave *agent* perches exactly this and which shells never got); consumers in `linkhost::{relink_shell, prepare_drive, run_action}`, `shelldisc::discover`, `activity::observe_links`; the class-(c) eligibility read is `spt_daemon::shellwake::watcher_eligible` (+ `launch_predates_boot`, `BOOT_RESTORE_SLACK_MS`), and the birth stamp it rests on is `shellinfo::{ShellLaunch, record_shell_launch, read_shell_launch}`.\r\n- **Source:** spt-core, flynn's spt-alchemy field report (2026-07-25) — a deterministic recipe, not a race. Class-(c) re-based BAROMETER W2 (releases#78, doyle's ruling on comment 5156810267).\r\n\r\n### 2.7 A daemon or machine restart strands `persistent` shells — and a graceful cascade fixture suppresses the failing arm `[REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]`\r\n- **Failure:** a restart can leave an `online` record over a corpse without a link-break or suspend edge. The original boot sweep fixed only machine reboots: it rejected corpses launched after machine boot, so a daemon-only restart still stranded every persistent shell. The measured releases#287 specimen launched at machine boot +46s and stayed down for 27 hours after a daemon restart despite its owner being online.\r\n- **The owner-facing surfaces HIDE the fault**, which is why it survived unreported for days: every display path (`shell list`, its `--json` twin, the shell-context render) routes through the ONE discovery seam, which **derives** status (2.6) and so correctly reports the binary as offline — while the only consumer whose decision matters, the wake cascade, reads the RECORDED field that no owner-facing surface shows. Absence of a visible symptom is therefore **not** evidence of absence of this fault, and a clean-looking context render must never be taken as proof the class did not occur. The correct falsifier is the **on-disk record**.\r\n- **The existing cascade test cannot fail on it — its SETUP SUPPRESSES THE FAILING ARM.** `rest_edges_cascade_shells_with_divergence` establishes its fixture by driving the suspend path, which itself writes the `offline` the code under test is supposed to encounter — so the restart shape (a record still saying `online` over a corpse, reached with **no** suspend edge) is unreachable from that fixture BY CONSTRUCTION. The test is not weak; it gives a true answer to the wrong question, so a reader asking \"is the cascade covered?\" is told yes. This is the third suppressed-arm instance found in one week, and the pattern is the same each time: **a fixture that establishes the precondition the code under test is supposed to establish.**\r\n- **Invariant:** a daemon-only or machine restart must not strand a persistent instance. **(a) Heal:** every reconcile cycle writes `offline` only where the derived birth-safe read contradicts recorded `online`; already-correct records are untouched. **(b) Restore:** the once-per-daemon-generation boot sweep and every owner offline→online edge run the same per-owner body. Relaunch every instance whose adapter declares `persistent`, whose owner is online, and whose binary is down in fact. There is **no machine-boot, launch-age, or stamp-presence eligibility gate**. A live binary awaiting bind is not down and must not be duplicated; missing birth evidence keeps the existing fail-toward-alive liveness behavior. The owner-online trigger remains an edge, bounding failed launches to one attempt per event.\r\n- **Regression shape:** construct recorded `online` over a corpse **without a suspend edge**, including a launch **after machine boot**, then bring its owner online after a spent boot sweep. Every stranded persistent sibling must be restored; offline owners and nonpersistent instances are not restored. Separately defend live/binding exclusion and the nonpersistent watcher freeze. Suspending in setup suppresses the failing arm and proves nothing about this class.\r\n- **Per-owner enumeration:** the existing `SHELL_RECORD_HEALED:<owner>/<shell_id>` daemon diagnostic names **every** successfully healed instance, not one representative. Healthy and already-healed records produce no event. The restore diagnostics likewise identify each attempted instance and distinguish boot from owner-online triggers. <!-- [doc->REQ-SHELL-HEAL-ENUMERATION] -->\r\n- **spt-core mapping:** `spt_daemon::shellwake::{heal_stale_online_records, restore_persistent_shells_at_boot, restore_persistent_shells_on_owner_online, OwnerOnlineEdge}`; `shellhost::live_launch_winner`; `spt_store::shellinfo::{ShellLaunch, shell_pid_provably_dead}`. `launch_predates_boot` and `BOOT_RESTORE_SLACK_MS` now govern only nonpersistent watcher eligibility.\r\n- **Source:** BAROMETER W2 (releases#78); releases#287 operator ruling (2026-09-08), superseding the persistent force-kill freeze and boot-instant conjunct.\r\n<!-- [doc->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS] -->\r\n\r\n---\r\n\r\n## 3. Lifecycle\r\n\r\n### 3.1 Ephemeral perch cleanup on every `ring` exit path\r","startLine":70,"lineNumbers":[70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99]},"fileSize":354670,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\docs\\KNOWN-HAZARDS.md"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":1093,"totalBytes":11985,"outputLines":30,"outputBytes":11985,"shownRange":{"start":70,"end":99},"nextOffset":100}}},"isError":false,"timestamp":1789001247053}}
{"type":"message","id":"4a7c21a4","parentId":"c6a60f9c","timestamp":"2026-09-10T00:47:27.055Z","message":{"role":"toolResult","toolCallId":"call_x0H1V8A2lPcCPhTpAr8GI4lL|fc_077f3569365e313d016aa1fe1eef6887d1bc5f662ca2a159c3","toolName":"read","content":[{"type":"text","text":"- **Failure (paid-for, releases#240 / CONDUIT #236 respin; MEASURED on both platforms 2026-08-29, not inferred):** self-detect leg (c) (`roster::detect_self_by_ancestry`) built its candidate list from `read_dir(owlery)` and took the FIRST candidate whose recorded pid was the nearest ancestor. Nearest-first correctly answered the ancestry dimension and left a second one open: when two or more DISTINCT top-level perches record the SAME live pid, they all match at that ancestor, and the winner was whatever the directory happened to yield. From ONE on-disk state at ONE sha, ext4 returned `[(bob, P), (alice, P)]` and chose `bob` while NTFS returned `[(alice, P), (bob, P)]` and chose `alice`. `spt send --from alice` then published `MSG_OUT` with `owner=bob`, and the sink filed the sender's row into the RECEIVER's log exactly as `IoLogSink` is specified to do — the sink was innocent, it was handed the wrong owner.\n- **THE GREEN WAS THE COIN LANDING HEADS.** The same cells passed on Windows in a gate run, a golden run, and a discarded head, and every one of those greens was enumeration order, not the code being right. No Windows pass of an identity-resolution cell is evidence about this seam. The order is arbitrary PER DIRECTORY — the same filesystem produced both orders in one session — so it is neither alphabetical nor creation-ordered and no platform is the safe one.\n- **The ambiguous state is PRODUCTION-REACHABLE, established by construction rather than by not having seen it:** two `spt api bind` calls from ONE parent shell both succeed, with no refusal and no override flag, and both top-level `info.json` records carry that shell as `parent_pid` — the E-1 candidate arm (7.34 / `REQ-SELF-DETECT-PARENT-PID`). Nested Psyche/Worker perches are NOT a route: they live at `owlery/<parent>/nested/<child>/` and leg (c) scans top level only, so they are outside the candidate set by path topology.\n- **Invariant:** when the candidates at the nearest matching ancestor name two or more DISTINCT ids, resolution is `None` with a loud-once diagnostic naming the tied ids and the shared pid — never a silent pick. Two candidates carrying the SAME id are NOT a tie (one spt-hosted perch matches on both its dead `rec.pid` and its live `parent_pid`; refusing that would re-break `REQ-SELF-DETECT-PARENT-PID`). The refusal is scoped to the nearest matching ancestor and does not fall through to an outer one — an outer perch is a DIFFERENT session, and answering with it is worse than declining. `None` is not a shrug: at the send site the owner falls back to the CALLED `--from`, so refusing is more ACCURATE than the coin as well as more honest. Label-only, like the whole of leg (c) — `authenticate()` is untouched (KH 7.3/7.5).\n- **A rig that manufactures an unreachable state is how this reached a golden run.** The integration cells bound BOTH endpoints under the test process's own pid, a state production reaches only by the shell shape above; they now bind DISTINCT LIVE owners by construction (the author under this process, the peer under the daemon), so the tie cannot arise and the assertions are FORCED by the edges resolving correctly rather than permitted by enumeration order. The proof that it is forced and not re-lucky: after the fix the peer still enumerates FIRST on ext4 and the author still wins.\n- **spt-core mapping:** `spt/src/roster.rs` `match_self_by_ancestry` (distinct-id refusal) + `warn_ambiguous_ancestor_once`; `spt/src/cli.rs` the send site whose `unwrap_or(--from)` fallback makes `None` the accurate answer; `spt/tests/io_events_undriven_kinds_e2e.rs` `bind_endpoint`/`rig` (distinct live owners). Units: `an_ambiguous_ancestor_resolves_to_none_in_either_candidate_order`, `one_perch_matching_on_both_its_pids_still_resolves`, `ambiguity_is_scoped_to_the_nearest_ancestor`.\n- **Source:** releases#240, filed off the CONDUIT #236 respin RCA (todlando instrument on kitsubito + Windows control on hfenduleam; doyle-ruled fix shape — a named refusal over a silent guess; deployah pre-registered the ship branches before the classification landed). Kin 7.34 (the dead `rec.pid` class this ambiguity rides on), 7.63 (a key narrower than the identity of what it addresses).\n\n---\n\n## Conformance checklist (condensed)\n\n| # | Invariant | spt-core surface |\n|---|---|---|\n| 1.1 | Grace wait precedes INIT_SIGNOFF | daemon teardown |\n| 1.4/4.4 | Deferred rows excluded from event-stream drain | daemon spool drain |\n| 2.1/5.1 | Stable PID/broker-handle over ephemeral PID | liveness detection |\n| 2.3 | Handoff argv/IPC version-tolerant (newer brain ↔ older broker) | broker↔brain IPC, self-update |\n| 2.4 | gen_start = now() on cold-start + handoff | per-instance generation |\n| 2.6 | A shell's ONLINE-ness is DERIVED (recorded status AND a not-provably-dead `shell.pid`) — an abruptly-killed binary breaks no link, so `close_shell`'s offline flip never runs and the record lies forever; derive at the gates/renders, keep the recorded field at the writers, the suspend-cascade close, and the wake reconciler — whose \"no spontaneous relaunch\" (deploys and quarantine own that decision) is since W2 the stated corpse-predates-boot predicate, not the stale record holding it open, and whose recycled-pid blindness is closed by the launch birth stamp wherever one is parked. Not only a crash edge: the GRACEFUL daemon-stop path kills bound shells at stop-begin with `close_shell` unrun (field 2026-07-25 — shells died abruptly, `info.json` still `online`, while the broker drained on 2m10s), so every daemon restart manufactures these stale records routinely; the derivation heals them at the gates | `spt_store::shellinfo::{shell_pid_provably_dead,effective_status,is_shell_online}`, `linkhost` relink/drive/cmd-wake, `shelldisc::discover`, `activity::observe_links` |\n<!-- [doc->REQ-HAZARD-SHELL-STALE-ONLINE] -->\n| 2.7 | A node restart must not permanently strand every `persistent` shell: (a) the daemon heals a recorded `online` the derivation contradicts, write-guarded on an actual change, and (b) a once-per-generation boot sweep PLUS an owner offline->online edge in the reconcile loop (releases#228 — the sweep alone strands the shells of any owner not yet online when it ran) relaunch an instance only when ALL of — adapter `persistent`, owner endpoint online, down in fact, recorded launch predates the boot instant (slack-absorbed) — hold; no stamp or no boot oracle ⇒ not restored. The owner-facing surfaces derive and look clean while the cascade reads the recorded field, so a clean render is NOT evidence the class did not occur — the falsifier is the on-disk record. The pre-existing cascade test's setup suppresses the failing arm (it suspends first), so the restart shape must be asserted on a fixture that never calls the suspend path | `shellwake::{heal_stale_online_records,restore_persistent_shells_at_boot,restore_persistent_shells_on_owner_online,OwnerOnlineEdge,launch_predates_boot}`, `proc::boot_instant_ms`, `shellinfo::{record_shell_launch,read_shell_launch}` |\n| 3.1 | Ephemeral perch cleanup on all exit paths | `ring` (RAII guard) |\n| 3.4 | A ring never adopts (so never deletes) a perch it did not create — probe the DIR, not the ready marker; record/spool/unreadable = occupied, empty = refused too; deliver + loud `RING_PERCH_EXISTS`/`RING_STALE_DIR` instead of block-waiting | `spt_msg::ring` probe + `create_dir` leaf |\n| 3.5 | A perch GC classifies on record PRESENCE only and never asks a liveness resolver — `is_perch_alive` is INVERTED here (recordless residue reads ALIVE, an offline endpoint reads DEAD: 24 vs 6 measured on HFENDULEAM 2026-08-04), so `!is_perch_alive` authorizes exactly the inverse set; the registry is never asked either (offline endpoints are legitimately absent from it) | `spt_store::perchgc::sweep`, `spt endpoint gc` |\n| 3.3 | Echo-commune before INIT_SIGNOFF | daemon psyche loop |\n| 4.1 | Envelope decode order, `&amp;` last | spt-proto (public, wire-versioned) |\n| 4.2 | Parser panic-free + tolerant | spt-proto |\n| 4.3 | Stale registry entries → fallback, never hard-fail | subnet registry resolution |\n\n[Showing lines 1000-1028 of 1093. Use :1029 to continue]"}],"details":{"truncation":{"content":"- **Failure (paid-for, releases#240 / CONDUIT #236 respin; MEASURED on both platforms 2026-08-29, not inferred):** self-detect leg (c) (`roster::detect_self_by_ancestry`) built its candidate list from `read_dir(owlery)` and took the FIRST candidate whose recorded pid was the nearest ancestor. Nearest-first correctly answered the ancestry dimension and left a second one open: when two or more DISTINCT top-level perches record the SAME live pid, they all match at that ancestor, and the winner was whatever the directory happened to yield. From ONE on-disk state at ONE sha, ext4 returned `[(bob, P), (alice, P)]` and chose `bob` while NTFS returned `[(alice, P), (bob, P)]` and chose `alice`. `spt send --from alice` then published `MSG_OUT` with `owner=bob`, and the sink filed the sender's row into the RECEIVER's log exactly as `IoLogSink` is specified to do — the sink was innocent, it was handed the wrong owner.\r\n- **THE GREEN WAS THE COIN LANDING HEADS.** The same cells passed on Windows in a gate run, a golden run, and a discarded head, and every one of those greens was enumeration order, not the code being right. No Windows pass of an identity-resolution cell is evidence about this seam. The order is arbitrary PER DIRECTORY — the same filesystem produced both orders in one session — so it is neither alphabetical nor creation-ordered and no platform is the safe one.\r\n- **The ambiguous state is PRODUCTION-REACHABLE, established by construction rather than by not having seen it:** two `spt api bind` calls from ONE parent shell both succeed, with no refusal and no override flag, and both top-level `info.json` records carry that shell as `parent_pid` — the E-1 candidate arm (7.34 / `REQ-SELF-DETECT-PARENT-PID`). Nested Psyche/Worker perches are NOT a route: they live at `owlery/<parent>/nested/<child>/` and leg (c) scans top level only, so they are outside the candidate set by path topology.\r\n- **Invariant:** when the candidates at the nearest matching ancestor name two or more DISTINCT ids, resolution is `None` with a loud-once diagnostic naming the tied ids and the shared pid — never a silent pick. Two candidates carrying the SAME id are NOT a tie (one spt-hosted perch matches on both its dead `rec.pid` and its live `parent_pid`; refusing that would re-break `REQ-SELF-DETECT-PARENT-PID`). The refusal is scoped to the nearest matching ancestor and does not fall through to an outer one — an outer perch is a DIFFERENT session, and answering with it is worse than declining. `None` is not a shrug: at the send site the owner falls back to the CALLED `--from`, so refusing is more ACCURATE than the coin as well as more honest. Label-only, like the whole of leg (c) — `authenticate()` is untouched (KH 7.3/7.5).\r\n- **A rig that manufactures an unreachable state is how this reached a golden run.** The integration cells bound BOTH endpoints under the test process's own pid, a state production reaches only by the shell shape above; they now bind DISTINCT LIVE owners by construction (the author under this process, the peer under the daemon), so the tie cannot arise and the assertions are FORCED by the edges resolving correctly rather than permitted by enumeration order. The proof that it is forced and not re-lucky: after the fix the peer still enumerates FIRST on ext4 and the author still wins.\r\n- **spt-core mapping:** `spt/src/roster.rs` `match_self_by_ancestry` (distinct-id refusal) + `warn_ambiguous_ancestor_once`; `spt/src/cli.rs` the send site whose `unwrap_or(--from)` fallback makes `None` the accurate answer; `spt/tests/io_events_undriven_kinds_e2e.rs` `bind_endpoint`/`rig` (distinct live owners). Units: `an_ambiguous_ancestor_resolves_to_none_in_either_candidate_order`, `one_perch_matching_on_both_its_pids_still_resolves`, `ambiguity_is_scoped_to_the_nearest_ancestor`.\r\n- **Source:** releases#240, filed off the CONDUIT #236 respin RCA (todlando instrument on kitsubito + Windows control on hfenduleam; doyle-ruled fix shape — a named refusal over a silent guess; deployah pre-registered the ship branches before the classification landed). Kin 7.34 (the dead `rec.pid` class this ambiguity rides on), 7.63 (a key narrower than the identity of what it addresses).\r\n\r\n---\r\n\r\n## Conformance checklist (condensed)\r\n\r\n| # | Invariant | spt-core surface |\r\n|---|---|---|\r\n| 1.1 | Grace wait precedes INIT_SIGNOFF | daemon teardown |\r\n| 1.4/4.4 | Deferred rows excluded from event-stream drain | daemon spool drain |\r\n| 2.1/5.1 | Stable PID/broker-handle over ephemeral PID | liveness detection |\r\n| 2.3 | Handoff argv/IPC version-tolerant (newer brain ↔ older broker) | broker↔brain IPC, self-update |\r\n| 2.4 | gen_start = now() on cold-start + handoff | per-instance generation |\r\n| 2.6 | A shell's ONLINE-ness is DERIVED (recorded status AND a not-provably-dead `shell.pid`) — an abruptly-killed binary breaks no link, so `close_shell`'s offline flip never runs and the record lies forever; derive at the gates/renders, keep the recorded field at the writers, the suspend-cascade close, and the wake reconciler — whose \"no spontaneous relaunch\" (deploys and quarantine own that decision) is since W2 the stated corpse-predates-boot predicate, not the stale record holding it open, and whose recycled-pid blindness is closed by the launch birth stamp wherever one is parked. Not only a crash edge: the GRACEFUL daemon-stop path kills bound shells at stop-begin with `close_shell` unrun (field 2026-07-25 — shells died abruptly, `info.json` still `online`, while the broker drained on 2m10s), so every daemon restart manufactures these stale records routinely; the derivation heals them at the gates | `spt_store::shellinfo::{shell_pid_provably_dead,effective_status,is_shell_online}`, `linkhost` relink/drive/cmd-wake, `shelldisc::discover`, `activity::observe_links` |\r\n<!-- [doc->REQ-HAZARD-SHELL-STALE-ONLINE] -->\r\n| 2.7 | A node restart must not permanently strand every `persistent` shell: (a) the daemon heals a recorded `online` the derivation contradicts, write-guarded on an actual change, and (b) a once-per-generation boot sweep PLUS an owner offline->online edge in the reconcile loop (releases#228 — the sweep alone strands the shells of any owner not yet online when it ran) relaunch an instance only when ALL of — adapter `persistent`, owner endpoint online, down in fact, recorded launch predates the boot instant (slack-absorbed) — hold; no stamp or no boot oracle ⇒ not restored. The owner-facing surfaces derive and look clean while the cascade reads the recorded field, so a clean render is NOT evidence the class did not occur — the falsifier is the on-disk record. The pre-existing cascade test's setup suppresses the failing arm (it suspends first), so the restart shape must be asserted on a fixture that never calls the suspend path | `shellwake::{heal_stale_online_records,restore_persistent_shells_at_boot,restore_persistent_shells_on_owner_online,OwnerOnlineEdge,launch_predates_boot}`, `proc::boot_instant_ms`, `shellinfo::{record_shell_launch,read_shell_launch}` |\r\n| 3.1 | Ephemeral perch cleanup on all exit paths | `ring` (RAII guard) |\r\n| 3.4 | A ring never adopts (so never deletes) a perch it did not create — probe the DIR, not the ready marker; record/spool/unreadable = occupied, empty = refused too; deliver + loud `RING_PERCH_EXISTS`/`RING_STALE_DIR` instead of block-waiting | `spt_msg::ring` probe + `create_dir` leaf |\r\n| 3.5 | A perch GC classifies on record PRESENCE only and never asks a liveness resolver — `is_perch_alive` is INVERTED here (recordless residue reads ALIVE, an offline endpoint reads DEAD: 24 vs 6 measured on HFENDULEAM 2026-08-04), so `!is_perch_alive` authorizes exactly the inverse set; the registry is never asked either (offline endpoints are legitimately absent from it) | `spt_store::perchgc::sweep`, `spt endpoint gc` |\r\n| 3.3 | Echo-commune before INIT_SIGNOFF | daemon psyche loop |\r\n| 4.1 | Envelope decode order, `&amp;` last | spt-proto (public, wire-versioned) |\r\n| 4.2 | Parser panic-free + tolerant | spt-proto |\r\n| 4.3 | Stale registry entries → fallback, never hard-fail | subnet registry resolution |\r","truncated":true,"truncatedBy":"lines","totalLines":94,"totalBytes":8182,"outputLines":29,"outputBytes":8182,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":1093,"displayContent":{"text":"- **Failure (paid-for, releases#240 / CONDUIT #236 respin; MEASURED on both platforms 2026-08-29, not inferred):** self-detect leg (c) (`roster::detect_self_by_ancestry`) built its candidate list from `read_dir(owlery)` and took the FIRST candidate whose recorded pid was the nearest ancestor. Nearest-first correctly answered the ancestry dimension and left a second one open: when two or more DISTINCT top-level perches record the SAME live pid, they all match at that ancestor, and the winner was whatever the directory happened to yield. From ONE on-disk state at ONE sha, ext4 returned `[(bob, P), (alice, P)]` and chose `bob` while NTFS returned `[(alice, P), (bob, P)]` and chose `alice`. `spt send --from alice` then published `MSG_OUT` with `owner=bob`, and the sink filed the sender's row into the RECEIVER's log exactly as `IoLogSink` is specified to do — the sink was innocent, it was handed the wrong owner.\r\n- **THE GREEN WAS THE COIN LANDING HEADS.** The same cells passed on Windows in a gate run, a golden run, and a discarded head, and every one of those greens was enumeration order, not the code being right. No Windows pass of an identity-resolution cell is evidence about this seam. The order is arbitrary PER DIRECTORY — the same filesystem produced both orders in one session — so it is neither alphabetical nor creation-ordered and no platform is the safe one.\r\n- **The ambiguous state is PRODUCTION-REACHABLE, established by construction rather than by not having seen it:** two `spt api bind` calls from ONE parent shell both succeed, with no refusal and no override flag, and both top-level `info.json` records carry that shell as `parent_pid` — the E-1 candidate arm (7.34 / `REQ-SELF-DETECT-PARENT-PID`). Nested Psyche/Worker perches are NOT a route: they live at `owlery/<parent>/nested/<child>/` and leg (c) scans top level only, so they are outside the candidate set by path topology.\r\n- **Invariant:** when the candidates at the nearest matching ancestor name two or more DISTINCT ids, resolution is `None` with a loud-once diagnostic naming the tied ids and the shared pid — never a silent pick. Two candidates carrying the SAME id are NOT a tie (one spt-hosted perch matches on both its dead `rec.pid` and its live `parent_pid`; refusing that would re-break `REQ-SELF-DETECT-PARENT-PID`). The refusal is scoped to the nearest matching ancestor and does not fall through to an outer one — an outer perch is a DIFFERENT session, and answering with it is worse than declining. `None` is not a shrug: at the send site the owner falls back to the CALLED `--from`, so refusing is more ACCURATE than the coin as well as more honest. Label-only, like the whole of leg (c) — `authenticate()` is untouched (KH 7.3/7.5).\r\n- **A rig that manufactures an unreachable state is how this reached a golden run.** The integration cells bound BOTH endpoints under the test process's own pid, a state production reaches only by the shell shape above; they now bind DISTINCT LIVE owners by construction (the author under this process, the peer under the daemon), so the tie cannot arise and the assertions are FORCED by the edges resolving correctly rather than permitted by enumeration order. The proof that it is forced and not re-lucky: after the fix the peer still enumerates FIRST on ext4 and the author still wins.\r\n- **spt-core mapping:** `spt/src/roster.rs` `match_self_by_ancestry` (distinct-id refusal) + `warn_ambiguous_ancestor_once`; `spt/src/cli.rs` the send site whose `unwrap_or(--from)` fallback makes `None` the accurate answer; `spt/tests/io_events_undriven_kinds_e2e.rs` `bind_endpoint`/`rig` (distinct live owners). Units: `an_ambiguous_ancestor_resolves_to_none_in_either_candidate_order`, `one_perch_matching_on_both_its_pids_still_resolves`, `ambiguity_is_scoped_to_the_nearest_ancestor`.\r\n- **Source:** releases#240, filed off the CONDUIT #236 respin RCA (todlando instrument on kitsubito + Windows control on hfenduleam; doyle-ruled fix shape — a named refusal over a silent guess; deployah pre-registered the ship branches before the classification landed). Kin 7.34 (the dead `rec.pid` class this ambiguity rides on), 7.63 (a key narrower than the identity of what it addresses).\r\n\r\n---\r\n\r\n## Conformance checklist (condensed)\r\n\r\n| # | Invariant | spt-core surface |\r\n|---|---|---|\r\n| 1.1 | Grace wait precedes INIT_SIGNOFF | daemon teardown |\r\n| 1.4/4.4 | Deferred rows excluded from event-stream drain | daemon spool drain |\r\n| 2.1/5.1 | Stable PID/broker-handle over ephemeral PID | liveness detection |\r\n| 2.3 | Handoff argv/IPC version-tolerant (newer brain ↔ older broker) | broker↔brain IPC, self-update |\r\n| 2.4 | gen_start = now() on cold-start + handoff | per-instance generation |\r\n| 2.6 | A shell's ONLINE-ness is DERIVED (recorded status AND a not-provably-dead `shell.pid`) — an abruptly-killed binary breaks no link, so `close_shell`'s offline flip never runs and the record lies forever; derive at the gates/renders, keep the recorded field at the writers, the suspend-cascade close, and the wake reconciler — whose \"no spontaneous relaunch\" (deploys and quarantine own that decision) is since W2 the stated corpse-predates-boot predicate, not the stale record holding it open, and whose recycled-pid blindness is closed by the launch birth stamp wherever one is parked. Not only a crash edge: the GRACEFUL daemon-stop path kills bound shells at stop-begin with `close_shell` unrun (field 2026-07-25 — shells died abruptly, `info.json` still `online`, while the broker drained on 2m10s), so every daemon restart manufactures these stale records routinely; the derivation heals them at the gates | `spt_store::shellinfo::{shell_pid_provably_dead,effective_status,is_shell_online}`, `linkhost` relink/drive/cmd-wake, `shelldisc::discover`, `activity::observe_links` |\r\n<!-- [doc->REQ-HAZARD-SHELL-STALE-ONLINE] -->\r\n| 2.7 | A node restart must not permanently strand every `persistent` shell: (a) the daemon heals a recorded `online` the derivation contradicts, write-guarded on an actual change, and (b) a once-per-generation boot sweep PLUS an owner offline->online edge in the reconcile loop (releases#228 — the sweep alone strands the shells of any owner not yet online when it ran) relaunch an instance only when ALL of — adapter `persistent`, owner endpoint online, down in fact, recorded launch predates the boot instant (slack-absorbed) — hold; no stamp or no boot oracle ⇒ not restored. The owner-facing surfaces derive and look clean while the cascade reads the recorded field, so a clean render is NOT evidence the class did not occur — the falsifier is the on-disk record. The pre-existing cascade test's setup suppresses the failing arm (it suspends first), so the restart shape must be asserted on a fixture that never calls the suspend path | `shellwake::{heal_stale_online_records,restore_persistent_shells_at_boot,restore_persistent_shells_on_owner_online,OwnerOnlineEdge,launch_predates_boot}`, `proc::boot_instant_ms`, `shellinfo::{record_shell_launch,read_shell_launch}` |\r\n| 3.1 | Ephemeral perch cleanup on all exit paths | `ring` (RAII guard) |\r\n| 3.4 | A ring never adopts (so never deletes) a perch it did not create — probe the DIR, not the ready marker; record/spool/unreadable = occupied, empty = refused too; deliver + loud `RING_PERCH_EXISTS`/`RING_STALE_DIR` instead of block-waiting | `spt_msg::ring` probe + `create_dir` leaf |\r\n| 3.5 | A perch GC classifies on record PRESENCE only and never asks a liveness resolver — `is_perch_alive` is INVERTED here (recordless residue reads ALIVE, an offline endpoint reads DEAD: 24 vs 6 measured on HFENDULEAM 2026-08-04), so `!is_perch_alive` authorizes exactly the inverse set; the registry is never asked either (offline endpoints are legitimately absent from it) | `spt_store::perchgc::sweep`, `spt endpoint gc` |\r\n| 3.3 | Echo-commune before INIT_SIGNOFF | daemon psyche loop |\r\n| 4.1 | Envelope decode order, `&amp;` last | spt-proto (public, wire-versioned) |\r\n| 4.2 | Parser panic-free + tolerant | spt-proto |\r\n| 4.3 | Stale registry entries → fallback, never hard-fail | subnet registry resolution |\r","startLine":1000,"lineNumbers":[1000,1001,1002,1003,1004,1005,1006,1007,1008,1009,1010,1011,1012,1013,1014,1015,1016,1017,1018,1019,1020,1021,1022,1023,1024,1025,1026,1027,1028]},"fileSize":354670,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\docs\\KNOWN-HAZARDS.md"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":1093,"totalBytes":8182,"outputLines":29,"outputBytes":8182,"shownRange":{"start":1000,"end":1028},"nextOffset":1029}}},"isError":false,"timestamp":1789001247055}}
{"type":"message","id":"0a005a53","parentId":"4a7c21a4","timestamp":"2026-09-10T00:47:27.061Z","message":{"role":"toolResult","toolCallId":"call_NSi9uK6PO9CCco5KJs6mmQpR|fc_077f3569365e313d016aa1fe1eef7887d1b0a154249fe945ab","toolName":"read","content":[{"type":"text","text":"gitdir: C:\\Users\\decid\\Documents\\projects\\spt-core\\.git\\worktrees\\287-shellrestore"}],"details":{"totalLines":1,"displayContent":{"text":"gitdir: C:\\Users\\decid\\Documents\\projects\\spt-core\\.git\\worktrees\\287-shellrestore","startLine":1,"lineNumbers":[1]},"fileSize":83,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\.git"}}},"isError":false,"timestamp":1789001247061}}
{"type":"message","id":"ea264a5f","parentId":"0a005a53","timestamp":"2026-09-10T00:47:27.069Z","message":{"role":"toolResult","toolCallId":"call_CkWsaMCu9kwUHzW0YDa1yZ4G|fc_077f3569365e313d016aa1fe1eef7087d1acab7dfb4d07f1b7","toolName":"read","content":[{"type":"text","text":"-- **Offline alone does not arm the watcher.** A shell force-killed (or crashed)\n-  *during the current boot* is offline with no watcher running, and comes back\n-  only on demand — `spt shell relink <id>`, or a `shell cmd` that wakes it. This\n-  is deliberate: killing a shell to overwrite its exe is a routine deploy step,\n-  and a watcher armed underneath would re-lock the file. A cleanly-closed\n-  instance and a restart casualty both arm a watcher normally.\n+Lifecycle facts worth knowing before you rely on either:\n+\n+- **Persistent instances return after a daemon-only restart as well as a machine\n+  reboot.** The boot sweep and each owner offline→online edge restore every\n+  persistent instance that is down in fact while its owner is online. Launch\n+  age and missing launch stamps do not exclude it. The birth-safe process probe\n+  still protects a live binary, including a launch awaiting bind. Failed\n+  launches are retried on an owner-online edge, not every reconcile tick.\n+- **Only nonpersistent instances retain the force-kill freeze.** Their\n+  same-boot or undatable corpses arm no watcher; cleanly closed instances and\n+  dated pre-boot casualties remain watcher-eligible. Persistent instances\n+  bypass this freeze. To keep a persistent shell down, use `teardown` or a\n+  nonpersistent manifest rather than killing its process.\n+- The daemon's `SHELL_RECORD_HEALED:<owner>/<shell_id>` diagnostics enumerate\n+  every instance whose stale online record was corrected, not a representative\n+  shell per owner. Already-correct records produce no repeated notice.\n \n <!-- [doc->REQ-SHELL-RELINK-FORCE] -->\n - **`relink` refuses an instance that is still running; `--force` is the\n@@ -129,6 +121,11 @@ Two lifecycle facts worth knowing before you rely on either:\n   canonical id, the perch and its persisted state, and the alias all survive;\n   teardown + spawn is what churns the id and frees the mint slot.\n \n+  <!-- [doc->REQ-SHELL-RELINK-BINDING-RESULT] -->\n+  A successful local relink reports `status=binding (online at bind)`: the\n+  binary has launched and the bind handshake is still the online transition.\n+  It does not describe that successful launch as `offline`.\n+\n   Two things `--force` deliberately will **not** do:\n \n   - **It refuses on an `ephemeral` instance.** An ephemeral shell's close *is*\ndiff --git a/docs/KNOWN-HAZARDS.md b/docs/KNOWN-HAZARDS.md\nindex 8ab726b7..c8b94c6e 100644\n--- a/docs/KNOWN-HAZARDS.md\n+++ b/docs/KNOWN-HAZARDS.md\n@@ -76,19 +76,20 @@ Hard-won edge cases harvested from the sister project (`claude_skill_owl`, ~80 c\n ### 2.6 A shell instance's `online` is a recorded field with no death edge\n - **Failure:** a shell perch's `status` has exactly two writers — the bind handshake (→`online`) and `close_shell` (→`offline`) — and `close_shell` runs on the **link-break** path alone. A binary that dies **abruptly** (force-kill, crash, OOM) breaks no link, so nothing ever writes the flip: the record says `online` forever while the pid it names is gone. Every consumer of the recorded field then inherits the lie — `shell list` reports online indefinitely; `relink` (the one command that recovers the instance) refuses `SHELL_ALREADY_ONLINE`, gated on the very state that is wrong; and `shell cmd`'s wake-if-offline arm reads \"not offline\", skips the wake, and spools the command to a corpse — accepted happily, drained by nobody, **indistinguishable from a busy shell**. The only exit was `teardown` + `spawn`, which destroys the perch (losing the instance's persisted state) and frees the mint slot, so the instance returns under a **different canonical id** and every reference to the old name — docs, other agents' notes, tag addressing — silently points at nothing. The rename is the *visible* cost; the expensive one is **silent**: a consumer's perch-persisted cursor re-baselines on the fresh perch, so work arriving between the death and the re-bind is never scanned — not failed-and-retried, just never seen (alchemy's tag scanner, 2026-07-25). Recovery must therefore preserve the perch, not merely restore a running binary.\n - **Invariant:** ONLINE-ness is **derived**, never the recorded field alone: `status == online` **AND** the recorded `shell.pid` is not *provably* dead. Provably dead is narrow — pid parked, non-zero, `!is_process_alive`; pid absent, unparseable, or `0` (a broker-hosted spawn whose backend exposed no pid) reads **alive**, the same fail-toward-alive parity 2.5 holds, so a backend the resolver cannot see is never declared dead. **The recycled-pid caveat is narrowed since BAROMETER W2 and is now conditional on the stamp:** a launch parks a **birth stamp** beside the pid (`shell.launch.json`), and the probe runs the PAIR test — same pid, *different* start time ⇒ **gone**, so a pid the OS handed to an unrelated process no longer masks a dead instance. That case matters most across a restart, where the pid namespace resets and a corpse pid is routinely reused; without the pair test the corpse reads alive, the derived status reads `online`, and `relink` refuses `SHELL_ALREADY_ONLINE` on the one instance needing recovery. Where there is **no stamp to compare** (an instance last launched by a pre-W2 build, or a backend exposing no start time) the original stance is unchanged and unchanged in *direction*: fail toward alive, heal *missed* rather than *mis-fired*. Derivation NEVER promotes: a recorded-`offline` instance stays offline whatever its pid file says. Site class is part of the invariant: **derive** at the gates and renders (relink, the cmd wake arm, drive's drop-if-offline, `shelldisc::discover` = the single source of both `shell list` surfaces, the activity fan-out); keep the **recorded** field at the two writers and at the owner-suspend cascade, whose `close_shell` **is** the cleanup that would be skipped; and keep it at the wake reconciler's eligibility read, which since BAROMETER W2 additionally asks the **corpse-boot discriminant** (releases#78, doyle's amendment on comment 5156810267): the read stays on the recorded field, but \"no spontaneous relaunch\" is now a stated predicate rather than a side effect of the record being wrong — see the bullet below.\n-- **No spontaneous relaunch of a SAME-BOOT death (operator-ratified, flynn 2026-07-25; re-based BAROMETER W2):** the reconciler does NOT adopt an instance whose binary died during this boot. Every reason an operator stops a shell is a reason not to want it back a tick later: mid-deploy (a shared install dir makes \"kill the process\" routine on Windows, since a running exe cannot be overwritten) an auto-relaunch would run the **old** binary out of the file being replaced — worse than a failed install — and it turns quarantining a misbehaving shell into a restart loop. For those, recovery stays demand-driven: `relink`, or a `shell cmd` that wakes. `relink` probes **locally** rather than trusting a daemon sweep, so recovery holds with the daemon down.\n-- **What holds that open is no longer the stale record (BAROMETER W2):** until W2 the guarantee rested on an accident: a force-killed instance was held out of the eligible set because its record still said `online`, and the eligibility read took that field raw. 2.7 heals that record to the truth, which removes the accidental protection — so the guard had to move to a rule that *states* it. Eligibility is now recorded-`offline` **AND** (no corpse, **or** the corpse's recorded launch **predates the boot instant**). A steady-state force-kill leaves a corpse launched *after* boot, which the predicate can never accept, so the ruling is preserved **by construction** rather than by a second rule. A restart casualty is the one case the predicate does accept, and it is exactly the scope 2.7 was greenlit for. Where the discriminant **cannot be evaluated** — no launch stamp (an instance last launched by a pre-W2 build), or no boot oracle on the platform — a corpse-bearing instance is **not** eligible: leaving a shell down beats arming a relaunch of a binary an operator may have killed on purpose.\n+- **No spontaneous relaunch of a NONPERSISTENT same-boot death:** the releases#78 force-kill freeze still protects nonpersistent instances, so killing one to free its exe does not arm a watcher underneath an overwrite. Recovery remains demand-driven. **Persistent instances are exempt** (operator ruling, releases#287): a kill is not a request to keep a declared-persistent shell down; use `teardown` or a nonpersistent manifest instead.\n+- **Watcher eligibility:** recorded `offline`, no live binary awaiting bind, and either `persistent`, no corpse, or a corpse whose launch predates the machine boot instant. A nonpersistent corpse without a launch stamp or boot oracle remains frozen; persistent instances require neither to become eligible. Birth-safe probes prevent a recycled pid from masquerading as a live shell.\n - **spt-core mapping:** `spt_store::shellinfo::{shell_pid_provably_dead, effective_status, is_shell_online}` (the resolver — the shell-side twin of 2.5's `liveness.rs`, which gave *agent* perches exactly this and which shells never got); consumers in `linkhost::{relink_shell, prepare_drive, run_action}`, `shelldisc::discover`, `activity::observe_links`; the class-(c) eligibility read is `spt_daemon::shellwake::watcher_eligible` (+ `launch_predates_boot`, `BOOT_RESTORE_SLACK_MS`), and the birth stamp it rests on is `shellinfo::{ShellLaunch, record_shell_launch, read_shell_launch}`.\n - **Source:** spt-core, flynn's spt-alchemy field report (2026-07-25) — a deterministic recipe, not a race. Class-(c) re-based BAROMETER W2 (releases#78, doyle's ruling on comment 5156810267).\n \n-### 2.7 A node restart permanently strands every `persistent` shell — and the existing cascade test cannot see that it does  `[REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]`\n-- **Failure:** the contract's own sentence says a `persistent` shell is **online whenever its owner endpoint is online**, and a node restart is the one case where that promise was never kept. A machine death breaks no link, so `close_shell` never runs and 2.6's stale record survives the reboot; when the owner comes back online, the wake cascade reads the **recorded** field, sees `online`, and skips the relaunch of a binary that has not existed since the previous boot. Nothing else corrects it: the instance is down forever, with no error anywhere, until a human happens to run `relink`. Every `persistent` instance on the node is stranded by one restart, together.\n+### 2.7 A daemon or machine restart strands `persistent` shells — and a graceful cascade fixture suppresses the failing arm `[REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]`\n+- **Failure:** a restart can leave an `online` record over a corpse without a link-break or suspend edge. The original boot sweep fixed only machine reboots: it rejected corpses launched after machine boot, so a daemon-only restart still stranded every persistent shell. The measured releases#287 specimen launched at machine boot +46s and stayed down for 27 hours after a daemon restart despite its owner being online.\n - **The owner-facing surfaces HIDE the fault**, which is why it survived unreported for days: every display path (`shell list`, its `--json` twin, the shell-context render) routes through the ONE discovery seam, which **derives** status (2.6) and so correctly reports the binary as offline — while the only consumer whose decision matters, the wake cascade, reads the RECORDED field that no owner-facing surface shows. Absence of a visible symptom is therefore **not** evidence of absence of this fault, and a clean-looking context render must never be taken as proof the class did not occur. The correct falsifier is the **on-disk record**.\n - **The existing cascade test cannot fail on it — its SETUP SUPPRESSES THE FAILING ARM.** `rest_edges_cascade_shells_with_divergence` establishes its fixture by driving the suspend path, which itself writes the `offline` the code under test is supposed to encounter — so the restart shape (a record still saying `online` over a corpse, reached with **no** suspend edge) is unreachable from that fixture BY CONSTRUCTION. The test is not weak; it gives a true answer to the wrong question, so a reader asking \"is the cascade covered?\" is told yes. This is the third suppressed-arm instance found in one week, and the pattern is the same each time: **a fixture that establishes the precondition the code under test is supposed to establish.**\n-- **Invariant:** a node restart must not strand a `persistent` instance. Two legs, and each is worthless alone. **(a) The record stops lying:** the daemon-side reconciler heals a recorded `online` that 2.6's derivation contradicts, writing `offline` — every cycle, **guarded on an actual change** (a heal that rewrote each tick would be a stream of identical writes and would destroy the record's mtime as a signal). **(b) The restart shape reaches a restoration path:** a once-per-daemon-generation boot sweep — and, for an owner that comes online after that sweep has run, an **owner offline→online edge** in the reconcile loop running the identical body (releases#228) — relaunches an instance when **all four** hold — the adapter section declares `persistent`; the **owner endpoint is online** (an offline owner is owed nothing — that is the line between healing a record and resurrecting a shell nobody is there to drive); the instance is down **in fact** (the derived read, made trustworthy across a pid reuse by 2.6's birth stamp); and its recorded launch **predates the boot instant**. No stamp ⇒ **not restored** (the sweep cannot prove it predates boot); no boot oracle on the platform ⇒ nothing restored at all. The boot instant is derived (Linux `btime`; Windows now − `GetTickCount64`), so a named slack constant absorbs derivation jitter and keeps a process launched moments after boot from reading as a corpse from the previous one. **The once-ness of the sweep is scheduling, not safety** (ruled releases#228): the predates-boot discriminant is what forbids relaunching a force-killed shell, and it is monotone within a machine boot — every launch restamps, a kill never does — so a second trigger evaluating the same conjuncts can only refuse more. The second trigger is an **edge**, not the level “owner is online”, which is what bounds a failing launch to one attempt per owner-online event instead of one per reconcile tick.\n-- **Regression shape (the row the suppressed fixture cannot express):** construct the restart shape **without a suspend edge** — an instance whose record says `online` over a corpse that predates boot, with an online owner — and assert it is restored. A fixture that suspends first re-creates the suppressed arm and proves nothing about this class.\n-- **spt-core mapping:** `spt_daemon::shellwake::{heal_stale_online_records, restore_persistent_shells_at_boot, restore_persistent_shells_on_owner_online, OwnerOnlineEdge, launch_predates_boot, BOOT_RESTORE_SLACK_MS}`; `spt_store::proc::boot_instant_ms`; `spt_store::shellinfo::{ShellLaunch, record_shell_launch, read_shell_launch}` (the birth stamp, parked at the two production pid-write sites the launch-site census identified).\n-- **Source:** spt-core, BAROMETER W2 (releases#78) — field specimen `liam/alchemy-0`, a genuinely stranded record: `online` over a dead pid, no launch stamp, days old.\n+- **Invariant:** a daemon-only or machine restart must not strand a persistent instance. **(a) Heal:** every reconcile cycle writes `offline` only where the derived birth-safe read contradicts recorded `online`; already-correct records are untouched. **(b) Restore:** the once-per-daemon-generation boot sweep and every owner offline→online edge run the same per-owner body. Relaunch every instance whose adapter declares `persistent`, whose owner is online, and whose binary is down in fact. There is **no machine-boot, launch-age, or stamp-presence eligibility gate**. A live binary awaiting bind is not down and must not be duplicated; missing birth evidence keeps the existing fail-toward-alive liveness behavior. The owner-online trigger remains an edge, bounding failed launches to one attempt per event.\n+- **Regression shape:** construct recorded `online` over a corpse **without a suspend edge**, including a launch **after machine boot**, then bring its owner online after a spent boot sweep. Every stranded persistent sibling must be restored; offline owners and nonpersistent instances are not restored. Separately defend live/binding exclusion and the nonpersistent watcher freeze. Suspending in setup suppresses the failing arm and proves nothing about this class.\n+- **Per-owner enumeration:** the existing `SHELL_RECORD_HEALED:<owner>/<shell_id>` daemon diagnostic names **every** successfully healed instance, not one representative. Healthy and already-healed records produce no event. The restore diagnostics likewise identify each attempted instance and distinguish boot from owner-online triggers. <!-- [doc->REQ-SHELL-HEAL-ENUMERATION] -->\n+- **spt-core mapping:** `spt_daemon::shellwake::{heal_stale_online_records, restore_persistent_shells_at_boot, restore_persistent_shells_on_owner_online, OwnerOnlineEdge}`; `shellhost::live_launch_winner`; `spt_store::shellinfo::{ShellLaunch, shell_pid_provably_dead}`. `launch_predates_boot` and `BOOT_RESTORE_SLACK_MS` now govern only nonpersistent watcher eligibility.\n+- **Source:** BAROMETER W2 (releases#78); releases#287 operator ruling (2026-09-08), superseding the persistent force-kill freeze and boot-instant conjunct.\n <!-- [doc->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS] -->\n \n ---\ndiff --git a/docs/MANIFEST.md b/docs/MANIFEST.md\nindex b468481b..781d270c 100644\n--- a/docs/MANIFEST.md\n+++ b/docs/MANIFEST.md\n@@ -788,7 +788,13 @@ types = [\"image\", \"sound\", \"event\"]   # \"event\" = arbitrary descriptive payloads\n \n The capability + sensory vocabularies live **here**, not on the perch — spt-core resolves them by `adapter_name`. Command-delivery (`command_receipt`) reuses the agent inject-input modes. The shell binary binds via `api bind` (type=Shell, owner from the link) and pushes sensory via `api emit --type <t>`.\n \n-**Sleep/wake:** when offline **and eligible** (+ `wake_command` set), spt-core runs the wake-watcher (mutually exclusive with the shell binary); `exit(wake-opcode)` → wake resolution; crash-exit → respawn with exponential backoff + give-up (until the next relink/launch). **The wake opcode is exit code 86.** The wake-watcher template fills `{id}`, `{adapter_name}`, `{link_token}`, `{adapter_dir}` (NOT `{perch_dir}` — that key is spawn-only; the waker fill site carries no perch context), and its `{link_token}` is minted *unparked* — an offline link has no live credential (the close retired it); a waker wakes by exit code, never by driving the link. **Eligibility (BAROMETER W2, releases#78):** a watcher is armed for an instance recorded `offline` that either carries no corpse (cleanly closed, or never launched) or whose corpse's recorded launch **predates the current boot** — a restart casualty. An instance whose binary died **during this boot** (force-kill, crash) arms no watcher and returns only on demand (`relink`, or a waking `shell cmd`), because killing a shell to overwrite its exe is a routine deploy step and a watcher armed underneath would re-lock the file; an instance whose corpse cannot be dated (no launch stamp, or no boot oracle on the platform) is treated the same way rather than guessed at. `persistent` auto-onlines the shell with its owner, and across a **node restart** — where nothing breaks the link, so the instance's record survives the reboot still claiming it is up — the daemon heals that record and a **once-per-daemon-generation boot sweep** relaunches the instance if its adapter is `persistent`, its owner endpoint is online, it is down in fact, and its recorded launch predates the boot instant. Because daemons boot before endpoints do, the same four conditions are re-run for any owner that goes **offline→online since the last reconcile pass** (releases#228) — without that edge, an owner that was still down when the sweep ran keeps its shells down until someone relinks by hand. It is an edge and not the level “the owner is online” so that a failing launch is attempted once per owner-online event rather than once per tick. `can_shutdown` authorizes `api owner-shutdown`. Cross-node *fresh-spawn* to wake an owner requires the owner's `shell_wake_spawn_anywhere` settings flag (pre-consent).\n+**Sleep/wake:** when offline **and eligible** (+ `wake_command` set), spt-core runs the wake-watcher, mutually exclusive with a live shell binary including one awaiting bind. Exit code **86** is the wake opcode; other exits respawn with exponential backoff and eventual give-up until the next shell launch. Templates fill `{id}`, `{adapter_name}`, `{link_token}`, `{adapter_dir}` (not `{perch_dir}`, which is spawn-only). The watcher's link token is minted *unparked*: it wakes by exit code, never by driving an offline link.\n+\n+**Persistent restoration:** on the once-per-daemon-generation boot sweep and each owner offline→online edge, spt-core restores **every** down persistent instance whose owner is online. This covers machine reboots and daemon-only restarts. The decision uses birth-safe liveness, not machine boot time or recorded launch age; a missing launch stamp does not disqualify a down instance. A live launch awaiting bind is not duplicated. The edge bounds failed launches to one attempt per owner-online event rather than every reconcile tick. Stale online records are healed only on an actual change. <!-- [doc->REQ-SHELL-PERSISTENT-BOOT-RESTORE] --> <!-- [doc->REQ-SHELL-OWNER-ONLINE-RESTORE] -->\n+\n+**Watcher eligibility:** recorded-offline instances with no live binary are eligible if persistent, cleanly closed/never launched (no corpse), or a dated pre-boot corpse. The same-boot force-kill freeze applies **only to nonpersistent** instances, whose undatable corpses also remain frozen. Persistent instances bypass that freeze (releases#287 operator ruling); keeping one down requires `teardown` or a nonpersistent manifest, not a kill.\n+\n+`can_shutdown` authorizes `api owner-shutdown`. Cross-node *fresh-spawn* to wake an owner requires the owner's `shell_wake_spawn_anywhere` settings flag (pre-consent).\n \n **Instantiation governance:** `require_approval` gates `shell spawn` (reusing the consent grant store — `remembered` persists a grant via allow-always, `always` prompts every spawn); `max_instances_per_owner` + `over_cap` cap how many instances one owner may hold (online + offline both count toward the cap). An instance's **alias** is **not** a manifest field — it is per-instance runtime state on the shell perch's `info.json` (set at `shell spawn --alias` / `shell rename`), since the manifest describes the *adapter*, not individual instances.\n \ndiff --git a/traceable-reqs.toml b/traceable-reqs.toml\nindex 21917ea9..a3039948 100644\n--- a/traceable-reqs.toml\n+++ b/traceable-reqs.toml\n@@ -3183,6 +3183,12 @@ required_stages = [\"doc\", \"impl\", \"unit\", \"int\"]  # W1 ACTIVATED at mint (doyle\n # satisfying. ---\n [[requirements]]\n id = \"REQ-HAZARD-SHELL-STALE-ONLINE\"\n+# AMENDED releases#287, operator ruling 2026-09-08: the historical title below's\n+# same-boot/undatable-corpse freeze and no-spontaneous-relaunch clauses now apply\n+# ONLY to NONPERSISTENT instances. Persistent restoration and watcher eligibility\n+# follow REQ-SHELL-PERSISTENT-BOOT-RESTORE / REQ-SHELL-OWNER-ONLINE-RESTORE, without\n+# launch-age or stamp-presence gates. Birth-safe live/binding exclusion remains.\n+# The liveness, identity/state preservation and durable-channel clauses stand.\n title = \"A shell instance's ONLINE-ness is DERIVED (recorded status AND its recorded `shell.pid` not provably dead), never the recorded `status` field alone — an abruptly-dead binary (force-kill, crash, OOM: no link-break, so `close_shell`'s offline flip never runs) must not read online forever. The shell-side twin of REQ-HAZARD-DAEMON-HOSTED-LIVENESS, which gave AGENT perches exactly this resolver and which shells never got. PROVABLY DEAD is the narrow discriminant: `shell.pid` present AND parses non-zero AND `!is_process_alive` — pid absent, unparseable, or 0 (a broker-hosted spawn whose backend exposed no pid records 0) reads ALIVE, the same interim-parity/fail-toward-alive stance `liveness.rs` already holds, so a pid-less backend is NEVER falsely declared dead. Recycled-pid caveat, accepted at mint: a reused pid reads alive, so the heal is missed, never mis-fired — the failure direction is 'stays stale', never 'kills a live instance'. SITE CLASSIFICATION IS PART OF THE REQUIREMENT (authoritative cfg(test)-excluded census at mint = 8 status reads, 3 classes — do NOT blanket-swap the predicate): (a) DERIVED — relink's already-online refusal (the gate that made recovery impossible), the `shell cmd` wake-if-offline arm (which silently spooled to a corpse), the drive drop-if-offline branch, `shelldisc::discover` (the single source of BOTH `shell list` renders, text + --json), and the activity fan-out's online filter; (b) RAW status, deliberately — `bind_shell_by_token`/`close_shell` (the WRITERS) and `cascade_owner_edge`'s suspend-close arm, where routing through the resolver would SKIP the close that is itself the cleanup, removing a heal path; (c) RAW recorded status AND the corpse-boot discriminant — the `shellwake::reconcile_once` watcher-eligibility read. AMENDED BAROMETER W2 (doyle, releases#78 comment 5156810267): this class's original text — 'RAW status, pending an operator ruling', with no spontaneous relaunch of anything — IS STRUCK, because it rested on a fact leg (a) deliberately ends. What actually held a force-killed instance out of the eligible set was the STALE `online` record itself; healing that record to the truth removes the accidental protection, so the guard had to move to a rule that STATES it. Eligibility is now recorded `offline` AND (no corpse, or the corpse's launch PREDATES THE BOOT INSTANT), reusing leg (b)'s `launch_predates_boot` and its slack constant verbatim. NO SPONTANEOUS RELAUNCH OF SAME-BOOT DEATHS: flynn's mid-deploy forcing case (an operator kills the process precisely to free the exe for overwrite — shared install dir ⇒ 'kill the process' is a routine deploy step) is preserved BY CONSTRUCTION, since a force-kill during steady state leaves a corpse launched AFTER boot, which the predicate can never accept; for those, recovery stays demand-driven (an unblocked `relink`, or a `shell cmd` that wakes). Boot-restoration of PRE-boot casualties is enabled, per #78 leg (b)'s operator-greenlit scope and this ruling (#78 thread). CORPSE IS DEFINED ON THE PARKED STAMPS, never on a bare pid re-probe: a RECYCLED pid — alive, but its native start stamp mismatching the one leg (a) parked — is corpse-equivalent because our process is dead, and the parked ABSOLUTE launch stamp is what decides pre/post-boot. Without that, the heal-then-recycle race reads 'alive ⇒ no corpse' on the next tick and walks back into the forbidden shape through a side door. Deriving is still WRONG at this site and this is NOT a swap to class (a): the census stays 8 reads / 3 classes — the site reclassifies, it does not multiply, the do-NOT-blanket-swap sentence above stands, and `shell list` still tells the truth by derivation (class (a), untouched). LOCALITY SYMMETRY (field-caught by flynn's leg-2 run, 2026-07-25): 'a shell cmd that wakes' holds for BOTH the local CLI cmd and the cross-node serve — at mint the wake arm lived in the SHELL_LINK_CMD serve handler only (module docs scoped wake-if-offline to remote cmd), so a LOCAL cmd against a corpse spooled silently and nothing woke, the exact accepted-happily/drained-by-nobody shape this req exists to kill; the arm is now the shared `linkhost::wake_if_offline_persistent`, called by both, and the derived status READ inside it is the same single census site as before (the census of 8 reads/3 classes is unchanged — the read moved, it did not multiply). ROTATION CARRIES THE DURABLE CHANNEL (flynn's frame-loss field catch, 2026-07-25, both trials): a spooled command frame is MAC-stamped under the link token current at SPOOL time, and the drain is a raw destructive passthrough — the shell verifies against its CURRENT key — so a relink's token mint orphaned every pending frame: drained to the woken binary, failed verify, correctly discarded BY THE ADAPTER, lost permanently; the wake-triggering command itself was the frame the wake lost (#23 armed the wake and never answered; #24/#19 spooled after rotation and drained in order). The fix is three-layered (todlando's A1/A2 race+crash hazards addressed at mint): (1) the token mint re-stamps all pending rows old-key→new-key (idempotent, crash-partial converges — an old-stamped remainder is converted by the next rotation or drain); (2) the rotated-out token is STASHED (`link.token.retired`, never a live credential — bind resolves only the parked file) so (3) the drain paths give any race straggler a second-chance re-stamp at delivery (`restamp_for_drain`), while frames verifying under neither key pass through untouched for the adapter to refuse exactly as before. Rows are selected by MAC verification against the threaded new key, never by token-snapshot equality (the D-2 class). MOCK CONFORMANCE IS A PROPERTY OF THE RIG, not a patch (todlando A4): mock-shell MUST verify inbound frame MACs exactly as the public contract demands of a real adapter and drop failures loudly — a mock that accepts what the field rejects is a broken rig, and that exact divergence (credulous mock) is how the frame loss passed the e2e while failing in the field. The adapter-side discard behavior is CORRECT and must never be softened to paper over the spool side. Relink additionally probes LOCALLY rather than trusting a daemon sweep, so recovery holds with the daemon down. RECOVERY PRESERVES CONSUMER STATE (the property consumers actually depend on, flynn 2026-07-25): 'same canonical id, same perch' exists so that state a consumer PERSISTED IN THE PERCH survives the recovery — an adapter's repo binding, a scanner's cursor. The teardown+spawn workaround destroyed exactly that, and its worst cost was SILENT, not the rename: alchemy's tag cursor re-baselines at the digest tip, so tags written between the kill and the re-bind are never scanned — not failed-and-retried, just never seen. A loud failure gets retried; a silent one does not. Gate: int — force-kill a bound persistent instance's process, then prove (1) `shell list` reads offline, (2) `relink` succeeds instead of SHELL_ALREADY_ONLINE and the SAME canonical id + perch survive, carrying perch-persisted consumer state with them (no teardown+spawn, no id churn, no re-baselined cursor), (3) no spontaneous relaunch occurs while the instance sits dead and undriven. FIELD VERDICT — PASS END-TO-END (flynn, alchemy-0, delivered 2026-07-26; v0.43.0, counter 77): every gate leg held in the field, on a record the pre-fix code had already poisoned. (1) `shell list` read OFFLINE while the on-disk info.json still said status=online — the daemon DERIVED offline from the corpse pid (29036 absent from the process table); the record was never corrected and did not need to be. (2) relink ADMITTED, no SHELL_ALREADY_ONLINE — verbatim SHELL_RELINKED:alchemy-0 owner=flynn pid=38644 status=offline, list online thereafter. (3) identity/state integrity: same canonical id, same owner, same perch; repo binding byte-identical (token_provenance=gh-cli); armed=true preserved; no spontaneous relaunch across the ~6.5h dead window. CURSOR SCOPING, ruled at closure (doyle 2026-07-26) — the one non-byte-identical field: the tag cursor moved 206158430541→210453397553 (gen:seq 48:333→49:49) because the recovery crossed a DAEMON restart that slid the retained digest window past the armed cursor. NOT a gap and NOT this req's property failing: spt-core raised after_predates_window and the consumer took its specced armed-cursor branch (alchemy REQ-TAG-SCANNER missed-rows leg — warn the owner, never silently re-baseline an armed cursor; alchemy src/tags.rs, two unit tests) — the FIRST FIELD FIRING of a path previously exercised only by unit tests. The 'no re-baselined cursor' gate property is hereby scoped: byte-identity holds across SHELL death within one daemon generation (the prior LEG 5 measurement, one daemon instance, uncontradicted — different scenario, generation never moved); across a daemon restart that slides the window the contract is DETECT-AND-ANNOUNCE, not preserve — corroborated independently by a second record's cursor decomposing under the same gen:seq scheme (4294967486 = 1:190). Practical exposure in this event: nil (every agent stood down through the gap; no tag traffic to miss). SECOND SPECIMEN HELD, not spent (doyle ruling 2026-07-26): liam's record — precondition verified by flynn (pid 35464 a CLEAN corpse, not recycled, so the accepted-at-mint recycled-pid caveat does not apply; same poisoned status=online shape; binding intact; cursor armed at 4294967486) — is deliberately NOT relinked while its owner is absent: the admit gate is already field-proven, a relink measurably re-baselines the cursor across a daemon-restart boundary, and liam's is the only untouched specimen with the longer (~8.5h) gap — spending it would destroy the cursor datum liam's own return produces, to buy a redundant second sample of a gate already proven to admit. It waits for liam. SPECIMEN HARVESTED (liam's own return, 2026-07-26 — operator-authorized endpoint start executed by doyle; pre-state snapshotted by doyle AND independently re-verified on disk by liam before touching anything): (1) relink ADMITTED at the ~8.5h gap — verbatim SHELL_RELINKED:alchemy-0 owner=liam pid=47608 status=offline (online at bind); the poisoned record did not refuse; info.json online is now TRUE. (2) integrity FULL PASS: id unchanged (no new mint slot), perch onlined at bind, repo binding byte-identical (token_provenance gh-cli), and the ROTATION-CARRIES-THE-DURABLE-CHANNEL machinery observed live in the field for the first time — link.token.retired present beside the fresh mint, and a post-relink round trip proven end-to-end (shell cmd spooled -> engine executed -> reply home on the owner's perch), not merely an outbound spool. (3) cursor, the sharpest datum of the harvest: BYTE-IDENTICAL through the relink itself (4294967486 armed at pid-change instant, doyle's file watch) — the relink preserves; the SCANNER moves. First scan re-baselined 1:190 -> 2:0 silently, and liam PROVED no-fire-by-precondition rather than assuming: the daemon restart bumped the generation but EVICTED NOTHING (window floor 4294967464 < cursor 4294967486, every gen-1 row past the cursor still retained), so after_predates_window is correctly false at every link of the chain (filter_after floor test -> emit-only-when-true -> unwrap_or(false)) and the missed-rows announce had no precondition to fire on. Flynn's earlier gen-bump reading is refined by this: a restart alone does not slide the window; flynn's announce fired because HIS window had genuinely slid. Scope note carried: a valid negative — the fix's announce chain verified link-by-link to contract on a specimen where it correctly stayed quiet. RESIDUE from the harvest (liam's find, alchemy lane, no core REQ): permanently-unsealable turns from a dark session are jumped by scanner cursor advance with no possible missed-rows notice (a generation bump is not a slide); ruled scanner-side announce for input-bearing null-seq jumps, pseudo-turns jump silently by design; core's per-turn input/input_seq/partial already suffice to discriminate.\"\n required_stages = [\"doc\", \"impl\", \"unit\", \"int\"]  # ACTIVATED at mint (doyle 2026-07-25) — field-blocking; flynn eats id churn every alchemy release until it lands. LANDED v0.43.0 (counter 77, PR #76/#77); field-verified 2026-07-26 — see FIELD VERDICT in title.\n \n@@ -4151,13 +4157,23 @@ required_stages = [\"doc\", \"impl\", \"unit\"]  # ACTIVATED BAROMETER W1 at impl+unit\n \n [[requirements]]\n id = \"REQ-SHELL-PERSISTENT-BOOT-RESTORE\"\n-title = \"A `persistent` SHELL COMES BACK AFTER A NODE RESTART, WHICH IS THE ONE CASE THE CONTRACT'S OWN SENTENCE PROMISES AND THE IMPLEMENTATION NEVER HANDLED. CONTEXT.md states it plainly -- a persistent shell is automatically brought online whenever its owner endpoint is online -- and the shipped code honors that across a graceful suspend->wake rest edge ONLY. A machine restart produces no such edge and there is no boot-time restoration at all, so after a reboot every persistent shell on the node stays down until a human runs `relink` by hand, with no surface saying so. Field, HFENDULEAM 2026-08-01: the fleet came up at 01:11 and NOT ONE persistent shell came back; the two instances that did were relaunched by a human at 01:27, sixteen minutes later. TWO INDEPENDENT GATES, both of which must fall or the promise stays broken. GATE 1, no restoration path: the five launch sites all require either an explicit command, a rest edge, or a `wake_command` the adapter may deliberately not declare, and daemon boot starts only the servicehost sweep and the wake reconciler. GATE 2, the stale record blocks the cascade even when an edge DOES arrive: the to_active arm gates on the RECORDED status being offline, and a machine death breaks no link, so the close that would write offline never runs and the record says online over a corpse forever. LEG (a) HEALS THE RECORD TO THE TRUTH, with a BIRTH-STAMPED probe rather than bare pid liveness -- a pid namespace resets at boot, so a corpse pid can be REUSED by an unrelated process, and a reused pid reads alive, makes the derived status read online, and makes `relink` refuse SHELL_ALREADY_ONLINE: the one command that recovers the instance refusing on the very state that is wrong. TWO STAMPS ARE PARKED, NOT ONE, AND THE REASON IS A PLATFORM ASYMMETRY THAT WOULD OTHERWISE SHIP SILENTLY BROKEN ON LINUX: `process_started_at` returns a Windows FILETIME (absolute, comparable across boots) but on Linux returns field 22 of /proc/<pid>/stat, which is JIFFIES SINCE BOOT and resets every boot, so a value recorded before a restart is indistinguishable from a fresh one after it. The NATIVE stamp is therefore used for one question only -- the same-pid-different-process PAIR test within a single boot, mirroring `liveness::relay_liveness` -- and an ABSOLUTE epoch-ms launch stamp is parked beside it for the cross-boot question. LEG (b) IS A ONCE-PER-DAEMON-GENERATION BOOT SWEEP relaunching persistent instances whose owners are online, and its discriminant is THE LAUNCH STAMP PREDATES THE BOOT INSTANT, which preserves the class-c force-kill ruling BY CONSTRUCTION rather than by a second rule: a force-kill during steady-state operation leaves a corpse launched AFTER boot, so it is never spontaneously relaunched and the mid-deploy exe-overwrite hazard that ruling protects is untouched. A NAMED SLACK CONSTANT absorbs boot-instant jitter so a process launched moments after boot is never misclassified as a corpse. Gate: impl -- the two stamps parked at the two production pid-write sites (the census is authoritative: five LAUNCH sites funnel through exactly TWO pid writes), the birth-stamped death probe, the record heal in the daemon-side reconciler guarded on an actual change of the recorded value, the boot-instant helper on both platforms, and the boot sweep with its predates-boot discriminant plus slack; unit -- a reused pid (same pid, different birth stamp) reads GONE where bare-pid liveness reads alive (the row the shipped probe cannot pass), a same-pid-same-birth process reads HELD, the heal writes offline over a stale online record and writes NOTHING when the record already matches the truth (the guard asserted by counting writes, since a heal that rewrites every cycle passes any state-only assertion), a corpse launched BEFORE the boot instant is eligible for restore while one launched AFTER it is NOT (the class-c pair, asserted together since either alone passes an implementation that restores everything or nothing), and an instance whose OWNER is offline is not restored even though its corpse predates boot.\"\n-required_stages = [\"doc\", \"impl\", \"unit\"]  # doc ACTIVATED BAROMETER W2 doc leg (doyle-ruled, this milestone): the CONTEXT.md `persistent` amendment states the boot sweep and its four conjunct gates incl. no-stamp => not restored, which IS this requirement's doc evidence; leaving the stage inactive would have made it evidence the registry cannot see.\n+title = \"Persistent shells are restored on each daemon-generation boot sweep whenever their owner is online and they are down in fact, regardless of machine boot time or recorded launch age (releases#287 operator ruling). Missing launch stamps do not disqualify a down instance. Preserve birth-safe liveness, do not duplicate a live launch awaiting bind, and heal stale online records only on an actual change. The force-kill freeze applies only to nonpersistent instances. Regression: restore same-boot and unstamped corpses without a suspend edge; refuse offline owners, nonpersistent instances and live binaries.\"\n+required_stages = [\"doc\", \"impl\", \"unit\"]\n \n [[requirements]]\n id = \"REQ-SHELL-OWNER-ONLINE-RESTORE\"\n-title = \"A `persistent` SHELL WHOSE OWNER ENDPOINT WAS NOT YET ONLINE WHEN THE BOOT SWEEP RAN IS RESTORED WHEN THAT OWNER COMES ONLINE, RATHER THAN STAYING DOWN UNTIL A HUMAN RELINKS IT. ROOT (releases#228): daemons boot before endpoints do, so the boot sweep of REQ-SHELL-PERSISTENT-BOOT-RESTORE correctly SKIPS an owner that is not up yet -- an offline owner is owed nothing -- and then nothing ever revisits it: the sweep is once per daemon generation, and bringup from offline produces no rest edge, so the ADR-0048 cascade never fires either. The shell stays down until a `shell cmd` wakes it or an operator relinks by hand, and whether a given node restart strands its shells is an ordering accident. THE ONCE-NESS IS SCHEDULING, NOT SAFETY (doyle, releases#228, re-measured at head): the `launch_predates_boot` corpse-boot discriminant is what preserves the force-kill ruling, and it is MONOTONE WITHIN A MACHINE BOOT -- every launch path restamps `launched_ms`, a kill never restamps, `boot_ms` is the machine boot instant -- so re-evaluating the conjuncts later in a generation can only refuse MORE, never less. FIX: the reconcile loop carries the SAME per-owner restore body the boot sweep runs, triggered on owner offline->online transitions since the last pass. THE TRIGGER IS AN EDGE, NOT A LEVEL, and the distinction is the whole bound: success self-limits because the launch restamps the instance out of eligibility, but a FAILING launch under a level trigger -- owner is online -- retries every five seconds for as long as the owner stays up, while an edge bounds it to one attempt per owner-online event. THE REFUSAL ARMS ARE CARRIED BY CONSTRUCTION, not by restatement: both triggers call one shared per-owner body, so persistent-only, owner online, down in fact via the derived read with the pair test, stamp present, launch predates boot, and no boot oracle means restore nothing can never drift apart between the two paths. THE EDGE STATE IS SEEDED BEFORE THE SWEEP, deliberately: seeded after, an owner that comes online in the window between the sweep and the first tick reads as already-online and never fires -- the exact miss this requirement exists to close -- while seeding before can at worst re-attempt an instance the sweep already restored, which the RESTAMPED BIRTH STAMP refuses -- and NOT the down-in-fact arm, which is the arm that reads like the answer and is not: a launch is not the online switch, so a just-restored instance keeps an `offline` record until its bind-shell handshake and still reads as down in fact; what refuses the second attempt is that the launch moved `launched_ms` past the boot instant. Each trigger names ITSELF in its event line, because the field evidence that diagnosed this ticket was a count of those events and a shared name would leave the next investigation unable to say which trigger fired. Kin REQ-SHELL-PERSISTENT-BOOT-RESTORE (the boot trigger and the shared arms), REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS (the invariant this closes the remaining seam of), REQ-SHELL-2 (the online/offline mutual exclusivity the edge path holds by stopping an armed watcher before it launches -- mid-generation a watcher CAN be armed for the casualty, unlike at boot).\"\n-required_stages = [\"doc\", \"impl\", \"unit\"]  # ACTIVATED CONDUIT W2 in the commit landing the evidence (doyle ruled shape (c) with conditions on releases#228; (a) not built, (b) rejected). doc = the CONTEXT.md `persistent` amendment plus KNOWN-HAZARDS 2.7 and MANIFEST.md Sleep/wake, all of which stated the boot sweep as the ONLY trigger and went false the moment a second trigger existed. MINTED CONDUIT W2 (todlando, 2026-08-28), releases#228.\n+title = \"Every down persistent shell of an owner is restored on its offline-to-online edge, including corpses launched after machine boot and instances with no launch stamp (releases#287 operator ruling supersedes the releases#228 boot-age gates). Both triggers share the same persistent-only, owner-online, birth-safe down-in-fact decision; a live launch awaiting bind is not duplicated. The edge is seeded before the boot sweep and bounds failed attempts to one per owner-online event. Stop an armed watcher before launching, and name the triggering event and every restored owner-qualified instance.\"\n+required_stages = [\"doc\", \"impl\", \"unit\"]\n+\n+[[requirements]]\n+id = \"REQ-SHELL-HEAL-ENUMERATION\"\n+title = \"The daemon's existing SHELL_RECORD_HEALED diagnostic enumerates every successfully healed instance per owner by owner-qualified identity, not only the first or a representative shell; healthy and already-healed instances produce no heal event (releases#287 rider).\"\n+required_stages = [\"doc\", \"impl\", \"unit\"]\n+\n+[[requirements]]\n+id = \"REQ-SHELL-RELINK-BINDING-RESULT\"\n+title = \"A successful local shell relink reports status=binding while awaiting the bind handshake, rather than describing the successful launch as offline; bind remains the transition to online (releases#287 rider).\"\n+required_stages = [\"doc\", \"impl\", \"int\"]\n \n [[requirements]]\n id = \"REQ-SHELL-KILL-AUTHENTICATED\"\n@@ -4166,7 +4182,7 @@ required_stages = [\"doc\", \"impl\", \"unit\", \"int\"]  # ACTIVATED BAROMETER W2 reope\n \n [[requirements]]\n id = \"REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS\"\n-title = \"A NODE RESTART MUST NOT PERMANENTLY STRAND EVERY `persistent` SHELL, AND THE EXISTING CASCADE TEST CANNOT SEE THAT IT DOES. Distinct from REQ-HAZARD-SHELL-STALE-ONLINE, which governs a force-kill DURING operation: this is a different trigger with a different safe remedy, and conflating them is how the restart class stays invisible. THE COVERAGE GAP IS THE HAZARD'S OWN SUBJECT, and it must be named here so no future reader mistakes the cascade test for coverage of this class: `rest_edges_cascade_shells_with_divergence` (spt-daemon resting.rs) is the cascade's only coverage and it walks the graceful path ONLY -- it SUSPENDS FIRST, and the suspend edge WRITES the offline record, so by the time the wake edge fires the instance is already in exactly the state the to_active arm requires. The restart shape -- record left online over a corpse, no rest edge ever -- is UNREACHABLE FROM THAT TEST BY CONSTRUCTION. The test is not weak; its SETUP SUPPRESSES THE FAILING ARM, so a reader asking whether the cascade is covered gets a true answer to the wrong question. This is the third suppressed-arm instance found in one week, and the pattern is the same each time: a fixture that establishes the precondition the code under test is supposed to establish. THE OWNER-FACING SURFACES HIDE THE FAULT, which is why it survived unreported for days: every display path (`shell list`, its --json twin, the shell-context render) routes through the ONE discovery seam, which DERIVES status and so correctly reports the binary as offline, while the only consumer whose decision matters -- the wake cascade -- reads the RECORDED field that no owner-facing surface shows. Absence of a visible symptom is therefore not evidence of absence of this fault, and a clean-looking context render must never be taken as proof the class did not occur; the correct falsifier is the on-disk record. Gate: impl -- the restart shape reaches a restoration path at all (leg b) and the record stops lying (leg a); unit -- the RESTART shape specifically, constructed WITHOUT a suspend edge: an instance whose record says online over a corpse that predates boot, with an online owner, is restored -- the row `rest_edges_cascade_shells_with_divergence` cannot express, asserted on a fixture that never calls the suspend path, since a fixture that suspends first re-creates the suppressed arm and proves nothing about this class.\"\n+title = \"A daemon-only or machine restart must not strand persistent shells. Heal stale online records and restore every down persistent instance when its owner is online, without boot-age or launch-stamp eligibility gates (releases#287 operator ruling). Birth-safe liveness still excludes a live binary, including one awaiting bind. The regression fixture must construct online records over corpses WITHOUT a suspend edge, including a corpse launched AFTER machine boot; the existing rest_edges_cascade_shells_with_divergence fixture suspends first and therefore suppresses the failing arm rather than covering it.\"\n required_stages = [\"doc\", \"impl\", \"unit\"]  # doc ACTIVATED BAROMETER W2 doc leg (doyle-ruled, this milestone): the KNOWN-HAZARDS 2.7 entry the mint note said was owed. AGENTS.md rule 4 makes every KNOWN-HAZARDS entry a REQ-HAZARD-* requirement and rule 1 makes the tag ride the same change as the evidence, so an untagged entry would be the untagged-evidence shape rule 1 bans. -- ACTIVATED BAROMETER W2 at impl+unit once BOTH legs landed -- deliberately not at leg (a), since the gate clause covers the boot sweep too and activating early would have let the requirement read satisfied while half its impl was unbuilt. MINTED BAROMETER W2 (todlando, 2026-08-02), releases#78 -- inactive-first. KNOWN-HAZARDS entry owed with the doc leg. The suppressed-arm sentence is REQUIRED text per the filing and doyle's dispatch: the hazard exists partly BECAUSE the existing test reads as coverage, so a requirement that did not name the gap would leave the same trap for the next reader.\n \n \n@@ -4386,6 +4402,7 @@ requirements = [\n [[groups]]\n name = \"cli-output-surface\"\n requirements = [\n+  \"REQ-SHELL-RELINK-BINDING-RESULT\",\n   \"REQ-BIND-HONEST-SELF-STAMP\",\n   \"REQ-CLI-1\",\n   \"REQ-CLI-2\",\n@@ -4550,6 +4567,7 @@ requirements = [\n [[groups]]\n name = \"endpoint-presence-truth\"\n requirements = [\n+  \"REQ-SHELL-HEAL-ENUMERATION\",\n   \"REQ-CONN-POISON-ATTRIBUTION\",\n   \"REQ-DAEMON-5\",\n   \"REQ-DAEMON-6\",\n\n[Showing lines 930-1088 of 1089. Use :1089 to continue]"}],"details":{"resolvedPath":"C:\\Users\\decid\\.omp\\agent\\sessions\\-Documents-projects-spt-core\\2026-09-09T23-45-42-317Z_01a08890-21ad-7000-8f54-208d4f4f2f2c\\112.bash-original.log","contentType":"text/plain","totalLines":1089,"displayContent":{"text":"-- **Offline alone does not arm the watcher.** A shell force-killed (or crashed)\n-  *during the current boot* is offline with no watcher running, and comes back\n-  only on demand — `spt shell relink <id>`, or a `shell cmd` that wakes it. This\n-  is deliberate: killing a shell to overwrite its exe is a routine deploy step,\n-  and a watcher armed underneath would re-lock the file. A cleanly-closed\n-  instance and a restart casualty both arm a watcher normally.\n+Lifecycle facts worth knowing before you rely on either:\n+\n+- **Persistent instances return after a daemon-only restart as well as a machine\n+  reboot.** The boot sweep and each owner offline→online edge restore every\n+  persistent instance that is down in fact while its owner is online. Launch\n+  age and missing launch stamps do not exclude it. The birth-safe process probe\n+  still protects a live binary, including a launch awaiting bind. Failed\n+  launches are retried on an owner-online edge, not every reconcile tick.\n+- **Only nonpersistent instances retain the force-kill freeze.** Their\n+  same-boot or undatable corpses arm no watcher; cleanly closed instances and\n+  dated pre-boot casualties remain watcher-eligible. Persistent instances\n+  bypass this freeze. To keep a persistent shell down, use `teardown` or a\n+  nonpersistent manifest rather than killing its process.\n+- The daemon's `SHELL_RECORD_HEALED:<owner>/<shell_id>` diagnostics enumerate\n+  every instance whose stale online record was corrected, not a representative\n+  shell per owner. Already-correct records produce no repeated notice.\n \n <!-- [doc->REQ-SHELL-RELINK-FORCE] -->\n - **`relink` refuses an instance that is still running; `--force` is the\n@@ -129,6 +121,11 @@ Two lifecycle facts worth knowing before you rely on either:\n   canonical id, the perch and its persisted state, and the alias all survive;\n   teardown + spawn is what churns the id and frees the mint slot.\n \n+  <!-- [doc->REQ-SHELL-RELINK-BINDING-RESULT] -->\n+  A successful local relink reports `status=binding (online at bind)`: the\n+  binary has launched and the bind handshake is still the online transition.\n+  It does not describe that successful launch as `offline`.\n+\n   Two things `--force` deliberately will **not** do:\n \n   - **It refuses on an `ephemeral` instance.** An ephemeral shell's close *is*\ndiff --git a/docs/KNOWN-HAZARDS.md b/docs/KNOWN-HAZARDS.md\nindex 8ab726b7..c8b94c6e 100644\n--- a/docs/KNOWN-HAZARDS.md\n+++ b/docs/KNOWN-HAZARDS.md\n@@ -76,19 +76,20 @@ Hard-won edge cases harvested from the sister project (`claude_skill_owl`, ~80 c\n ### 2.6 A shell instance's `online` is a recorded field with no death edge\n - **Failure:** a shell perch's `status` has exactly two writers — the bind handshake (→`online`) and `close_shell` (→`offline`) — and `close_shell` runs on the **link-break** path alone. A binary that dies **abruptly** (force-kill, crash, OOM) breaks no link, so nothing ever writes the flip: the record says `online` forever while the pid it names is gone. Every consumer of the recorded field then inherits the lie — `shell list` reports online indefinitely; `relink` (the one command that recovers the instance) refuses `SHELL_ALREADY_ONLINE`, gated on the very state that is wrong; and `shell cmd`'s wake-if-offline arm reads \"not offline\", skips the wake, and spools the command to a corpse — accepted happily, drained by nobody, **indistinguishable from a busy shell**. The only exit was `teardown` + `spawn`, which destroys the perch (losing the instance's persisted state) and frees the mint slot, so the instance returns under a **different canonical id** and every reference to the old name — docs, other agents' notes, tag addressing — silently points at nothing. The rename is the *visible* cost; the expensive one is **silent**: a consumer's perch-persisted cursor re-baselines on the fresh perch, so work arriving between the death and the re-bind is never scanned — not failed-and-retried, just never seen (alchemy's tag scanner, 2026-07-25). Recovery must therefore preserve the perch, not merely restore a running binary.\n - **Invariant:** ONLINE-ness is **derived**, never the recorded field alone: `status == online` **AND** the recorded `shell.pid` is not *provably* dead. Provably dead is narrow — pid parked, non-zero, `!is_process_alive`; pid absent, unparseable, or `0` (a broker-hosted spawn whose backend exposed no pid) reads **alive**, the same fail-toward-alive parity 2.5 holds, so a backend the resolver cannot see is never declared dead. **The recycled-pid caveat is narrowed since BAROMETER W2 and is now conditional on the stamp:** a launch parks a **birth stamp** beside the pid (`shell.launch.json`), and the probe runs the PAIR test — same pid, *different* start time ⇒ **gone**, so a pid the OS handed to an unrelated process no longer masks a dead instance. That case matters most across a restart, where the pid namespace resets and a corpse pid is routinely reused; without the pair test the corpse reads alive, the derived status reads `online`, and `relink` refuses `SHELL_ALREADY_ONLINE` on the one instance needing recovery. Where there is **no stamp to compare** (an instance last launched by a pre-W2 build, or a backend exposing no start time) the original stance is unchanged and unchanged in *direction*: fail toward alive, heal *missed* rather than *mis-fired*. Derivation NEVER promotes: a recorded-`offline` instance stays offline whatever its pid file says. Site class is part of the invariant: **derive** at the gates and renders (relink, the cmd wake arm, drive's drop-if-offline, `shelldisc::discover` = the single source of both `shell list` surfaces, the activity fan-out); keep the **recorded** field at the two writers and at the owner-suspend cascade, whose `close_shell` **is** the cleanup that would be skipped; and keep it at the wake reconciler's eligibility read, which since BAROMETER W2 additionally asks the **corpse-boot discriminant** (releases#78, doyle's amendment on comment 5156810267): the read stays on the recorded field, but \"no spontaneous relaunch\" is now a stated predicate rather than a side effect of the record being wrong — see the bullet below.\n-- **No spontaneous relaunch of a SAME-BOOT death (operator-ratified, flynn 2026-07-25; re-based BAROMETER W2):** the reconciler does NOT adopt an instance whose binary died during this boot. Every reason an operator stops a shell is a reason not to want it back a tick later: mid-deploy (a shared install dir makes \"kill the process\" routine on Windows, since a running exe cannot be overwritten) an auto-relaunch would run the **old** binary out of the file being replaced — worse than a failed install — and it turns quarantining a misbehaving shell into a restart loop. For those, recovery stays demand-driven: `relink`, or a `shell cmd` that wakes. `relink` probes **locally** rather than trusting a daemon sweep, so recovery holds with the daemon down.\n-- **What holds that open is no longer the stale record (BAROMETER W2):** until W2 the guarantee rested on an accident: a force-killed instance was held out of the eligible set because its record still said `online`, and the eligibility read took that field raw. 2.7 heals that record to the truth, which removes the accidental protection — so the guard had to move to a rule that *states* it. Eligibility is now recorded-`offline` **AND** (no corpse, **or** the corpse's recorded launch **predates the boot instant**). A steady-state force-kill leaves a corpse launched *after* boot, which the predicate can never accept, so the ruling is preserved **by construction** rather than by a second rule. A restart casualty is the one case the predicate does accept, and it is exactly the scope 2.7 was greenlit for. Where the discriminant **cannot be evaluated** — no launch stamp (an instance last launched by a pre-W2 build), or no boot oracle on the platform — a corpse-bearing instance is **not** eligible: leaving a shell down beats arming a relaunch of a binary an operator may have killed on purpose.\n+- **No spontaneous relaunch of a NONPERSISTENT same-boot death:** the releases#78 force-kill freeze still protects nonpersistent instances, so killing one to free its exe does not arm a watcher underneath an overwrite. Recovery remains demand-driven. **Persistent instances are exempt** (operator ruling, releases#287): a kill is not a request to keep a declared-persistent shell down; use `teardown` or a nonpersistent manifest instead.\n+- **Watcher eligibility:** recorded `offline`, no live binary awaiting bind, and either `persistent`, no corpse, or a corpse whose launch predates the machine boot instant. A nonpersistent corpse without a launch stamp or boot oracle remains frozen; persistent instances require neither to become eligible. Birth-safe probes prevent a recycled pid from masquerading as a live shell.\n - **spt-core mapping:** `spt_store::shellinfo::{shell_pid_provably_dead, effective_status, is_shell_online}` (the resolver — the shell-side twin of 2.5's `liveness.rs`, which gave *agent* perches exactly this and which shells never got); consumers in `linkhost::{relink_shell, prepare_drive, run_action}`, `shelldisc::discover`, `activity::observe_links`; the class-(c) eligibility read is `spt_daemon::shellwake::watcher_eligible` (+ `launch_predates_boot`, `BOOT_RESTORE_SLACK_MS`), and the birth stamp it rests on is `shellinfo::{ShellLaunch, record_shell_launch, read_shell_launch}`.\n - **Source:** spt-core, flynn's spt-alchemy field report (2026-07-25) — a deterministic recipe, not a race. Class-(c) re-based BAROMETER W2 (releases#78, doyle's ruling on comment 5156810267).\n \n-### 2.7 A node restart permanently strands every `persistent` shell — and the existing cascade test cannot see that it does  `[REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]`\n-- **Failure:** the contract's own sentence says a `persistent` shell is **online whenever its owner endpoint is online**, and a node restart is the one case where that promise was never kept. A machine death breaks no link, so `close_shell` never runs and 2.6's stale record survives the reboot; when the owner comes back online, the wake cascade reads the **recorded** field, sees `online`, and skips the relaunch of a binary that has not existed since the previous boot. Nothing else corrects it: the instance is down forever, with no error anywhere, until a human happens to run `relink`. Every `persistent` instance on the node is stranded by one restart, together.\n+### 2.7 A daemon or machine restart strands `persistent` shells — and a graceful cascade fixture suppresses the failing arm `[REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]`\n+- **Failure:** a restart can leave an `online` record over a corpse without a link-break or suspend edge. The original boot sweep fixed only machine reboots: it rejected corpses launched after machine boot, so a daemon-only restart still stranded every persistent shell. The measured releases#287 specimen launched at machine boot +46s and stayed down for 27 hours after a daemon restart despite its owner being online.\n - **The owner-facing surfaces HIDE the fault**, which is why it survived unreported for days: every display path (`shell list`, its `--json` twin, the shell-context render) routes through the ONE discovery seam, which **derives** status (2.6) and so correctly reports the binary as offline — while the only consumer whose decision matters, the wake cascade, reads the RECORDED field that no owner-facing surface shows. Absence of a visible symptom is therefore **not** evidence of absence of this fault, and a clean-looking context render must never be taken as proof the class did not occur. The correct falsifier is the **on-disk record**.\n - **The existing cascade test cannot fail on it — its SETUP SUPPRESSES THE FAILING ARM.** `rest_edges_cascade_shells_with_divergence` establishes its fixture by driving the suspend path, which itself writes the `offline` the code under test is supposed to encounter — so the restart shape (a record still saying `online` over a corpse, reached with **no** suspend edge) is unreachable from that fixture BY CONSTRUCTION. The test is not weak; it gives a true answer to the wrong question, so a reader asking \"is the cascade covered?\" is told yes. This is the third suppressed-arm instance found in one week, and the pattern is the same each time: **a fixture that establishes the precondition the code under test is supposed to establish.**\n-- **Invariant:** a node restart must not strand a `persistent` instance. Two legs, and each is worthless alone. **(a) The record stops lying:** the daemon-side reconciler heals a recorded `online` that 2.6's derivation contradicts, writing `offline` — every cycle, **guarded on an actual change** (a heal that rewrote each tick would be a stream of identical writes and would destroy the record's mtime as a signal). **(b) The restart shape reaches a restoration path:** a once-per-daemon-generation boot sweep — and, for an owner that comes online after that sweep has run, an **owner offline→online edge** in the reconcile loop running the identical body (releases#228) — relaunches an instance when **all four** hold — the adapter section declares `persistent`; the **owner endpoint is online** (an offline owner is owed nothing — that is the line between healing a record and resurrecting a shell nobody is there to drive); the instance is down **in fact** (the derived read, made trustworthy across a pid reuse by 2.6's birth stamp); and its recorded launch **predates the boot instant**. No stamp ⇒ **not restored** (the sweep cannot prove it predates boot); no boot oracle on the platform ⇒ nothing restored at all. The boot instant is derived (Linux `btime`; Windows now − `GetTickCount64`), so a named slack constant absorbs derivation jitter and keeps a process launched moments after boot from reading as a corpse from the previous one. **The once-ness of the sweep is scheduling, not safety** (ruled releases#228): the predates-boot discriminant is what forbids relaunching a force-killed shell, and it is monotone within a machine boot — every launch restamps, a kill never does — so a second trigger evaluating the same conjuncts can only refuse more. The second trigger is an **edge**, not the level “owner is online”, which is what bounds a failing launch to one attempt per owner-online event instead of one per reconcile tick.\n-- **Regression shape (the row the suppressed fixture cannot express):** construct the restart shape **without a suspend edge** — an instance whose record says `online` over a corpse that predates boot, with an online owner — and assert it is restored. A fixture that suspends first re-creates the suppressed arm and proves nothing about this class.\n-- **spt-core mapping:** `spt_daemon::shellwake::{heal_stale_online_records, restore_persistent_shells_at_boot, restore_persistent_shells_on_owner_online, OwnerOnlineEdge, launch_predates_boot, BOOT_RESTORE_SLACK_MS}`; `spt_store::proc::boot_instant_ms`; `spt_store::shellinfo::{ShellLaunch, record_shell_launch, read_shell_launch}` (the birth stamp, parked at the two production pid-write sites the launch-site census identified).\n-- **Source:** spt-core, BAROMETER W2 (releases#78) — field specimen `liam/alchemy-0`, a genuinely stranded record: `online` over a dead pid, no launch stamp, days old.\n+- **Invariant:** a daemon-only or machine restart must not strand a persistent instance. **(a) Heal:** every reconcile cycle writes `offline` only where the derived birth-safe read contradicts recorded `online`; already-correct records are untouched. **(b) Restore:** the once-per-daemon-generation boot sweep and every owner offline→online edge run the same per-owner body. Relaunch every instance whose adapter declares `persistent`, whose owner is online, and whose binary is down in fact. There is **no machine-boot, launch-age, or stamp-presence eligibility gate**. A live binary awaiting bind is not down and must not be duplicated; missing birth evidence keeps the existing fail-toward-alive liveness behavior. The owner-online trigger remains an edge, bounding failed launches to one attempt per event.\n+- **Regression shape:** construct recorded `online` over a corpse **without a suspend edge**, including a launch **after machine boot**, then bring its owner online after a spent boot sweep. Every stranded persistent sibling must be restored; offline owners and nonpersistent instances are not restored. Separately defend live/binding exclusion and the nonpersistent watcher freeze. Suspending in setup suppresses the failing arm and proves nothing about this class.\n+- **Per-owner enumeration:** the existing `SHELL_RECORD_HEALED:<owner>/<shell_id>` daemon diagnostic names **every** successfully healed instance, not one representative. Healthy and already-healed records produce no event. The restore diagnostics likewise identify each attempted instance and distinguish boot from owner-online triggers. <!-- [doc->REQ-SHELL-HEAL-ENUMERATION] -->\n+- **spt-core mapping:** `spt_daemon::shellwake::{heal_stale_online_records, restore_persistent_shells_at_boot, restore_persistent_shells_on_owner_online, OwnerOnlineEdge}`; `shellhost::live_launch_winner`; `spt_store::shellinfo::{ShellLaunch, shell_pid_provably_dead}`. `launch_predates_boot` and `BOOT_RESTORE_SLACK_MS` now govern only nonpersistent watcher eligibility.\n+- **Source:** BAROMETER W2 (releases#78); releases#287 operator ruling (2026-09-08), superseding the persistent force-kill freeze and boot-instant conjunct.\n <!-- [doc->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS] -->\n \n ---\ndiff --git a/docs/MANIFEST.md b/docs/MANIFEST.md\nindex b468481b..781d270c 100644\n--- a/docs/MANIFEST.md\n+++ b/docs/MANIFEST.md\n@@ -788,7 +788,13 @@ types = [\"image\", \"sound\", \"event\"]   # \"event\" = arbitrary descriptive payloads\n \n The capability + sensory vocabularies live **here**, not on the perch — spt-core resolves them by `adapter_name`. Command-delivery (`command_receipt`) reuses the agent inject-input modes. The shell binary binds via `api bind` (type=Shell, owner from the link) and pushes sensory via `api emit --type <t>`.\n \n-**Sleep/wake:** when offline **and eligible** (+ `wake_command` set), spt-core runs the wake-watcher (mutually exclusive with the shell binary); `exit(wake-opcode)` → wake resolution; crash-exit → respawn with exponential backoff + give-up (until the next relink/launch). **The wake opcode is exit code 86.** The wake-watcher template fills `{id}`, `{adapter_name}`, `{link_token}`, `{adapter_dir}` (NOT `{perch_dir}` — that key is spawn-only; the waker fill site carries no perch context), and its `{link_token}` is minted *unparked* — an offline link has no live credential (the close retired it); a waker wakes by exit code, never by driving the link. **Eligibility (BAROMETER W2, releases#78):** a watcher is armed for an instance recorded `offline` that either carries no corpse (cleanly closed, or never launched) or whose corpse's recorded launch **predates the current boot** — a restart casualty. An instance whose binary died **during this boot** (force-kill, crash) arms no watcher and returns only on demand (`relink`, or a waking `shell cmd`), because killing a shell to overwrite its exe is a routine deploy step and a watcher armed underneath would re-lock the file; an instance whose corpse cannot be dated (no launch stamp, or no boot oracle on the platform) is treated the same way rather than guessed at. `persistent` auto-onlines the shell with its owner, and across a **node restart** — where nothing breaks the link, so the instance's record survives the reboot still claiming it is up — the daemon heals that record and a **once-per-daemon-generation boot sweep** relaunches the instance if its adapter is `persistent`, its owner endpoint is online, it is down in fact, and its recorded launch predates the boot instant. Because daemons boot before endpoints do, the same four conditions are re-run for any owner that goes **offline→online since the last reconcile pass** (releases#228) — without that edge, an owner that was still down when the sweep ran keeps its shells down until someone relinks by hand. It is an edge and not the level “the owner is online” so that a failing launch is attempted once per owner-online event rather than once per tick. `can_shutdown` authorizes `api owner-shutdown`. Cross-node *fresh-spawn* to wake an owner requires the owner's `shell_wake_spawn_anywhere` settings flag (pre-consent).\n+**Sleep/wake:** when offline **and eligible** (+ `wake_command` set), spt-core runs the wake-watcher, mutually exclusive with a live shell binary including one awaiting bind. Exit code **86** is the wake opcode; other exits respawn with exponential backoff and eventual give-up until the next shell launch. Templates fill `{id}`, `{adapter_name}`, `{link_token}`, `{adapter_dir}` (not `{perch_dir}`, which is spawn-only). The watcher's link token is minted *unparked*: it wakes by exit code, never by driving an offline link.\n+\n+**Persistent restoration:** on the once-per-daemon-generation boot sweep and each owner offline→online edge, spt-core restores **every** down persistent instance whose owner is online. This covers machine reboots and daemon-only restarts. The decision uses birth-safe liveness, not machine boot time or recorded launch age; a missing launch stamp does not disqualify a down instance. A live launch awaiting bind is not duplicated. The edge bounds failed launches to one attempt per owner-online event rather than every reconcile tick. Stale online records are healed only on an actual change. <!-- [doc->REQ-SHELL-PERSISTENT-BOOT-RESTORE] --> <!-- [doc->REQ-SHELL-OWNER-ONLINE-RESTORE] -->\n+\n+**Watcher eligibility:** recorded-offline instances with no live binary are eligible if persistent, cleanly closed/never launched (no corpse), or a dated pre-boot corpse. The same-boot force-kill freeze applies **only to nonpersistent** instances, whose undatable corpses also remain frozen. Persistent instances bypass that freeze (releases#287 operator ruling); keeping one down requires `teardown` or a nonpersistent manifest, not a kill.\n+\n+`can_shutdown` authorizes `api owner-shutdown`. Cross-node *fresh-spawn* to wake an owner requires the owner's `shell_wake_spawn_anywhere` settings flag (pre-consent).\n \n **Instantiation governance:** `require_approval` gates `shell spawn` (reusing the consent grant store — `remembered` persists a grant via allow-always, `always` prompts every spawn); `max_instances_per_owner` + `over_cap` cap how many instances one owner may hold (online + offline both count toward the cap). An instance's **alias** is **not** a manifest field — it is per-instance runtime state on the shell perch's `info.json` (set at `shell spawn --alias` / `shell rename`), since the manifest describes the *adapter*, not individual instances.\n \ndiff --git a/traceable-reqs.toml b/traceable-reqs.toml\nindex 21917ea9..a3039948 100644\n--- a/traceable-reqs.toml\n+++ b/traceable-reqs.toml\n@@ -3183,6 +3183,12 @@ required_stages = [\"doc\", \"impl\", \"unit\", \"int\"]  # W1 ACTIVATED at mint (doyle\n # satisfying. ---\n [[requirements]]\n id = \"REQ-HAZARD-SHELL-STALE-ONLINE\"\n+# AMENDED releases#287, operator ruling 2026-09-08: the historical title below's\n+# same-boot/undatable-corpse freeze and no-spontaneous-relaunch clauses now apply\n+# ONLY to NONPERSISTENT instances. Persistent restoration and watcher eligibility\n+# follow REQ-SHELL-PERSISTENT-BOOT-RESTORE / REQ-SHELL-OWNER-ONLINE-RESTORE, without\n+# launch-age or stamp-presence gates. Birth-safe live/binding exclusion remains.\n+# The liveness, identity/state preservation and durable-channel clauses stand.\n title = \"A shell instance's ONLINE-ness is DERIVED (recorded status AND its recorded `shell.pid` not provably dead), never the recorded `status` field alone — an abruptly-dead binary (force-kill, crash, OOM: no link-break, so `close_shell`'s offline flip never runs) must not read online forever. The shell-side twin of REQ-HAZARD-DAEMON-HOSTED-LIVENESS, which gave AGENT perches exactly this resolver and which shells never got. PROVABLY DEAD is the narrow discriminant: `shell.pid` present AND parses non-zero AND `!is_process_alive` — pid absent, unparseable, or 0 (a broker-hosted spawn whose backend exposed no pid records 0) reads ALIVE, the same interim-parity/fail-toward-alive stance `liveness.rs` already holds, so a pid-less backend is NEVER falsely declared dead. Recycled-pid caveat, accepted at mint: a reused pid reads alive, so the heal is missed, never mis-fired — the failure direction is 'stays stale', never 'kills a live instance'. SITE CLASSIFICATION IS PART OF THE REQUIREMENT (authoritative cfg(test)-excluded census at mint = 8 status reads, 3 classes — do NOT blanket-swap the predicate): (a) DERIVED — relink's already-online refusal (the gate that made recovery impossible), the `shell cmd` wake-if-offline arm (which silently spooled to a corpse), the drive drop-if-offline branch, `shelldisc::discover` (the single source of BOTH `shell list` renders, text + --json), and the activity fan-out's online filter; (b) RAW status, deliberately — `bind_shell_by_token`/`close_shell` (the WRITERS) and `cascade_owner_edge`'s suspend-close arm, where routing through the resolver would SKIP the close that is itself the cleanup, removing a heal path; (c) RAW recorded status AND the corpse-boot discriminant — the `shellwake::reconcile_once` watcher-eligibility read. AMENDED BAROMETER W2 (doyle, releases#78 comment 5156810267): this class's original text — 'RAW status, pending an operator ruling', with no spontaneous relaunch of anything — IS STRUCK, because it rested on a fact leg (a) deliberately ends. What actually held a force-killed instance out of the eligible set was the STALE `online` record itself; healing that record to the truth removes the accidental protection, so the guard had to move to a rule that STATES it. Eligibility is now recorded `offline` AND (no corpse, or the corpse's launch PREDATES THE BOOT INSTANT), reusing leg (b)'s `launch_predates_boot` and its slack constant verbatim. NO SPONTANEOUS RELAUNCH OF SAME-BOOT DEATHS: flynn's mid-deploy forcing case (an operator kills the process precisely to free the exe for overwrite — shared install dir ⇒ 'kill the process' is a routine deploy step) is preserved BY CONSTRUCTION, since a force-kill during steady state leaves a corpse launched AFTER boot, which the predicate can never accept; for those, recovery stays demand-driven (an unblocked `relink`, or a `shell cmd` that wakes). Boot-restoration of PRE-boot casualties is enabled, per #78 leg (b)'s operator-greenlit scope and this ruling (#78 thread). CORPSE IS DEFINED ON THE PARKED STAMPS, never on a bare pid re-probe: a RECYCLED pid — alive, but its native start stamp mismatching the one leg (a) parked — is corpse-equivalent because our process is dead, and the parked ABSOLUTE launch stamp is what decides pre/post-boot. Without that, the heal-then-recycle race reads 'alive ⇒ no corpse' on the next tick and walks back into the forbidden shape through a side door. Deriving is still WRONG at this site and this is NOT a swap to class (a): the census stays 8 reads / 3 classes — the site reclassifies, it does not multiply, the do-NOT-blanket-swap sentence above stands, and `shell list` still tells the truth by derivation (class (a), untouched). LOCALITY SYMMETRY (field-caught by flynn's leg-2 run, 2026-07-25): 'a shell cmd that wakes' holds for BOTH the local CLI cmd and the cross-node serve — at mint the wake arm lived in the SHELL_LINK_CMD serve handler only (module docs scoped wake-if-offline to remote cmd), so a LOCAL cmd against a corpse spooled silently and nothing woke, the exact accepted-happily/drained-by-nobody shape this req exists to kill; the arm is now the shared `linkhost::wake_if_offline_persistent`, called by both, and the derived status READ inside it is the same single census site as before (the census of 8 reads/3 classes is unchanged — the read moved, it did not multiply). ROTATION CARRIES THE DURABLE CHANNEL (flynn's frame-loss field catch, 2026-07-25, both trials): a spooled command frame is MAC-stamped under the link token current at SPOOL time, and the drain is a raw destructive passthrough — the shell verifies against its CURRENT key — so a relink's token mint orphaned every pending frame: drained to the woken binary, failed verify, correctly discarded BY THE ADAPTER, lost permanently; the wake-triggering command itself was the frame the wake lost (#23 armed the wake and never answered; #24/#19 spooled after rotation and drained in order). The fix is three-layered (todlando's A1/A2 race+crash hazards addressed at mint): (1) the token mint re-stamps all pending rows old-key→new-key (idempotent, crash-partial converges — an old-stamped remainder is converted by the next rotation or drain); (2) the rotated-out token is STASHED (`link.token.retired`, never a live credential — bind resolves only the parked file) so (3) the drain paths give any race straggler a second-chance re-stamp at delivery (`restamp_for_drain`), while frames verifying under neither key pass through untouched for the adapter to refuse exactly as before. Rows are selected by MAC verification against the threaded new key, never by token-snapshot equality (the D-2 class). MOCK CONFORMANCE IS A PROPERTY OF THE RIG, not a patch (todlando A4): mock-shell MUST verify inbound frame MACs exactly as the public contract demands of a real adapter and drop failures loudly — a mock that accepts what the field rejects is a broken rig, and that exact divergence (credulous mock) is how the frame loss passed the e2e while failing in the field. The adapter-side discard behavior is CORRECT and must never be softened to paper over the spool side. Relink additionally probes LOCALLY rather than trusting a daemon sweep, so recovery holds with the daemon down. RECOVERY PRESERVES CONSUMER STATE (the property consumers actually depend on, flynn 2026-07-25): 'same canonical id, same perch' exists so that state a consumer PERSISTED IN THE PERCH survives the recovery — an adapter's repo binding, a scanner's cursor. The teardown+spawn workaround destroyed exactly that, and its worst cost was SILENT, not the rename: alchemy's tag cursor re-baselines at the digest tip, so tags written between the kill and the re-bind are never scanned — not failed-and-retried, just never seen. A loud failure gets retried; a silent one does not. Gate: int — force-kill a bound persistent instance's process, then prove (1) `shell list` reads offline, (2) `relink` succeeds instead of SHELL_ALREADY_ONLINE and the SAME canonical id + perch survive, carrying perch-persisted consumer state with them (no teardown+spawn, no id churn, no re-baselined cursor), (3) no spontaneous relaunch occurs while the instance sits dead and undriven. FIELD VERDICT — PASS END-TO-END (flynn, alchemy-0, delivered 2026-07-26; v0.43.0, counter 77): every gate leg held in the field, on a record the pre-fix code had already poisoned. (1) `shell list` read OFFLINE while the on-disk info.json still said status=online — the daemon DERIVED offline from the corpse pid (29036 absent from the process table); the record was never corrected and did not need to be. (2) relink ADMITTED, no SHELL_ALREADY_ONLINE — verbatim SHELL_RELINKED:alchemy-0 owner=flynn pid=38644 status=offline, list online thereafter. (3) identity/state integrity: same canonical id, same owner, same perch; repo binding byte-identical (token_provenance=gh-cli); armed=true preserved; no spontaneous relaunch across the ~6.5h dead window. CURSOR SCOPING, ruled at closure (doyle 2026-07-26) — the one non-byte-identical field: the tag cursor moved 206158430541→210453397553 (gen:seq 48:333→49:49) because the recovery crossed a DAEMON restart that slid the retained digest window past the armed cursor. NOT a gap and NOT this req's property failing: spt-core raised after_predates_window and the consumer took its specced armed-cursor branch (alchemy REQ-TAG-SCANNER missed-rows leg — warn the owner, never silently re-baseline an armed cursor; alchemy src/tags.rs, two unit tests) — the FIRST FIELD FIRING of a path previously exercised only by unit tests. The 'no re-baselined cursor' gate property is hereby scoped: byte-identity holds across SHELL death within one daemon generation (the prior LEG 5 measurement, one daemon instance, uncontradicted — different scenario, generation never moved); across a daemon restart that slides the window the contract is DETECT-AND-ANNOUNCE, not preserve — corroborated independently by a second record's cursor decomposing under the same gen:seq scheme (4294967486 = 1:190). Practical exposure in this event: nil (every agent stood down through the gap; no tag traffic to miss). SECOND SPECIMEN HELD, not spent (doyle ruling 2026-07-26): liam's record — precondition verified by flynn (pid 35464 a CLEAN corpse, not recycled, so the accepted-at-mint recycled-pid caveat does not apply; same poisoned status=online shape; binding intact; cursor armed at 4294967486) — is deliberately NOT relinked while its owner is absent: the admit gate is already field-proven, a relink measurably re-baselines the cursor across a daemon-restart boundary, and liam's is the only untouched specimen with the longer (~8.5h) gap — spending it would destroy the cursor datum liam's own return produces, to buy a redundant second sample of a gate already proven to admit. It waits for liam. SPECIMEN HARVESTED (liam's own return, 2026-07-26 — operator-authorized endpoint start executed by doyle; pre-state snapshotted by doyle AND independently re-verified on disk by liam before touching anything): (1) relink ADMITTED at the ~8.5h gap — verbatim SHELL_RELINKED:alchemy-0 owner=liam pid=47608 status=offline (online at bind); the poisoned record did not refuse; info.json online is now TRUE. (2) integrity FULL PASS: id unchanged (no new mint slot), perch onlined at bind, repo binding byte-identical (token_provenance gh-cli), and the ROTATION-CARRIES-THE-DURABLE-CHANNEL machinery observed live in the field for the first time — link.token.retired present beside the fresh mint, and a post-relink round trip proven end-to-end (shell cmd spooled -> engine executed -> reply home on the owner's perch), not merely an outbound spool. (3) cursor, the sharpest datum of the harvest: BYTE-IDENTICAL through the relink itself (4294967486 armed at pid-change instant, doyle's file watch) — the relink preserves; the SCANNER moves. First scan re-baselined 1:190 -> 2:0 silently, and liam PROVED no-fire-by-precondition rather than assuming: the daemon restart bumped the generation but EVICTED NOTHING (window floor 4294967464 < cursor 4294967486, every gen-1 row past the cursor still retained), so after_predates_window is correctly false at every link of the chain (filter_after floor test -> emit-only-when-true -> unwrap_or(false)) and the missed-rows announce had no precondition to fire on. Flynn's earlier gen-bump reading is refined by this: a restart alone does not slide the window; flynn's announce fired because HIS window had genuinely slid. Scope note carried: a valid negative — the fix's announce chain verified link-by-link to contract on a specimen where it correctly stayed quiet. RESIDUE from the harvest (liam's find, alchemy lane, no core REQ): permanently-unsealable turns from a dark session are jumped by scanner cursor advance with no possible missed-rows notice (a generation bump is not a slide); ruled scanner-side announce for input-bearing null-seq jumps, pseudo-turns jump silently by design; core's per-turn input/input_seq/partial already suffice to discriminate.\"\n required_stages = [\"doc\", \"impl\", \"unit\", \"int\"]  # ACTIVATED at mint (doyle 2026-07-25) — field-blocking; flynn eats id churn every alchemy release until it lands. LANDED v0.43.0 (counter 77, PR #76/#77); field-verified 2026-07-26 — see FIELD VERDICT in title.\n \n@@ -4151,13 +4157,23 @@ required_stages = [\"doc\", \"impl\", \"unit\"]  # ACTIVATED BAROMETER W1 at impl+unit\n \n [[requirements]]\n id = \"REQ-SHELL-PERSISTENT-BOOT-RESTORE\"\n-title = \"A `persistent` SHELL COMES BACK AFTER A NODE RESTART, WHICH IS THE ONE CASE THE CONTRACT'S OWN SENTENCE PROMISES AND THE IMPLEMENTATION NEVER HANDLED. CONTEXT.md states it plainly -- a persistent shell is automatically brought online whenever its owner endpoint is online -- and the shipped code honors that across a graceful suspend->wake rest edge ONLY. A machine restart produces no such edge and there is no boot-time restoration at all, so after a reboot every persistent shell on the node stays down until a human runs `relink` by hand, with no surface saying so. Field, HFENDULEAM 2026-08-01: the fleet came up at 01:11 and NOT ONE persistent shell came back; the two instances that did were relaunched by a human at 01:27, sixteen minutes later. TWO INDEPENDENT GATES, both of which must fall or the promise stays broken. GATE 1, no restoration path: the five launch sites all require either an explicit command, a rest edge, or a `wake_command` the adapter may deliberately not declare, and daemon boot starts only the servicehost sweep and the wake reconciler. GATE 2, the stale record blocks the cascade even when an edge DOES arrive: the to_active arm gates on the RECORDED status being offline, and a machine death breaks no link, so the close that would write offline never runs and the record says online over a corpse forever. LEG (a) HEALS THE RECORD TO THE TRUTH, with a BIRTH-STAMPED probe rather than bare pid liveness -- a pid namespace resets at boot, so a corpse pid can be REUSED by an unrelated process, and a reused pid reads alive, makes the derived status read online, and makes `relink` refuse SHELL_ALREADY_ONLINE: the one command that recovers the instance refusing on the very state that is wrong. TWO STAMPS ARE PARKED, NOT ONE, AND THE REASON IS A PLATFORM ASYMMETRY THAT WOULD OTHERWISE SHIP SILENTLY BROKEN ON LINUX: `process_started_at` returns a Windows FILETIME (absolute, comparable across boots) but on Linux returns field 22 of /proc/<pid>/stat, which is JIFFIES SINCE BOOT and resets every boot, so a value recorded before a restart is indistinguishable from a fresh one after it. The NATIVE stamp is therefore used for one question only -- the same-pid-different-process PAIR test within a single boot, mirroring `liveness::relay_liveness` -- and an ABSOLUTE epoch-ms launch stamp is parked beside it for the cross-boot question. LEG (b) IS A ONCE-PER-DAEMON-GENERATION BOOT SWEEP relaunching persistent instances whose owners are online, and its discriminant is THE LAUNCH STAMP PREDATES THE BOOT INSTANT, which preserves the class-c force-kill ruling BY CONSTRUCTION rather than by a second rule: a force-kill during steady-state operation leaves a corpse launched AFTER boot, so it is never spontaneously relaunched and the mid-deploy exe-overwrite hazard that ruling protects is untouched. A NAMED SLACK CONSTANT absorbs boot-instant jitter so a process launched moments after boot is never misclassified as a corpse. Gate: impl -- the two stamps parked at the two production pid-write sites (the census is authoritative: five LAUNCH sites funnel through exactly TWO pid writes), the birth-stamped death probe, the record heal in the daemon-side reconciler guarded on an actual change of the recorded value, the boot-instant helper on both platforms, and the boot sweep with its predates-boot discriminant plus slack; unit -- a reused pid (same pid, different birth stamp) reads GONE where bare-pid liveness reads alive (the row the shipped probe cannot pass), a same-pid-same-birth process reads HELD, the heal writes offline over a stale online record and writes NOTHING when the record already matches the truth (the guard asserted by counting writes, since a heal that rewrites every cycle passes any state-only assertion), a corpse launched BEFORE the boot instant is eligible for restore while one launched AFTER it is NOT (the class-c pair, asserted together since either alone passes an implementation that restores everything or nothing), and an instance whose OWNER is offline is not restored even though its corpse predates boot.\"\n-required_stages = [\"doc\", \"impl\", \"unit\"]  # doc ACTIVATED BAROMETER W2 doc leg (doyle-ruled, this milestone): the CONTEXT.md `persistent` amendment states the boot sweep and its four conjunct gates incl. no-stamp => not restored, which IS this requirement's doc evidence; leaving the stage inactive would have made it evidence the registry cannot see.\n+title = \"Persistent shells are restored on each daemon-generation boot sweep whenever their owner is online and they are down in fact, regardless of machine boot time or recorded launch age (releases#287 operator ruling). Missing launch stamps do not disqualify a down instance. Preserve birth-safe liveness, do not duplicate a live launch awaiting bind, and heal stale online records only on an actual change. The force-kill freeze applies only to nonpersistent instances. Regression: restore same-boot and unstamped corpses without a suspend edge; refuse offline owners, nonpersistent instances and live binaries.\"\n+required_stages = [\"doc\", \"impl\", \"unit\"]\n \n [[requirements]]\n id = \"REQ-SHELL-OWNER-ONLINE-RESTORE\"\n-title = \"A `persistent` SHELL WHOSE OWNER ENDPOINT WAS NOT YET ONLINE WHEN THE BOOT SWEEP RAN IS RESTORED WHEN THAT OWNER COMES ONLINE, RATHER THAN STAYING DOWN UNTIL A HUMAN RELINKS IT. ROOT (releases#228): daemons boot before endpoints do, so the boot sweep of REQ-SHELL-PERSISTENT-BOOT-RESTORE correctly SKIPS an owner that is not up yet -- an offline owner is owed nothing -- and then nothing ever revisits it: the sweep is once per daemon generation, and bringup from offline produces no rest edge, so the ADR-0048 cascade never fires either. The shell stays down until a `shell cmd` wakes it or an operator relinks by hand, and whether a given node restart strands its shells is an ordering accident. THE ONCE-NESS IS SCHEDULING, NOT SAFETY (doyle, releases#228, re-measured at head): the `launch_predates_boot` corpse-boot discriminant is what preserves the force-kill ruling, and it is MONOTONE WITHIN A MACHINE BOOT -- every launch path restamps `launched_ms`, a kill never restamps, `boot_ms` is the machine boot instant -- so re-evaluating the conjuncts later in a generation can only refuse MORE, never less. FIX: the reconcile loop carries the SAME per-owner restore body the boot sweep runs, triggered on owner offline->online transitions since the last pass. THE TRIGGER IS AN EDGE, NOT A LEVEL, and the distinction is the whole bound: success self-limits because the launch restamps the instance out of eligibility, but a FAILING launch under a level trigger -- owner is online -- retries every five seconds for as long as the owner stays up, while an edge bounds it to one attempt per owner-online event. THE REFUSAL ARMS ARE CARRIED BY CONSTRUCTION, not by restatement: both triggers call one shared per-owner body, so persistent-only, owner online, down in fact via the derived read with the pair test, stamp present, launch predates boot, and no boot oracle means restore nothing can never drift apart between the two paths. THE EDGE STATE IS SEEDED BEFORE THE SWEEP, deliberately: seeded after, an owner that comes online in the window between the sweep and the first tick reads as already-online and never fires -- the exact miss this requirement exists to close -- while seeding before can at worst re-attempt an instance the sweep already restored, which the RESTAMPED BIRTH STAMP refuses -- and NOT the down-in-fact arm, which is the arm that reads like the answer and is not: a launch is not the online switch, so a just-restored instance keeps an `offline` record until its bind-shell handshake and still reads as down in fact; what refuses the second attempt is that the launch moved `launched_ms` past the boot instant. Each trigger names ITSELF in its event line, because the field evidence that diagnosed this ticket was a count of those events and a shared name would leave the next investigation unable to say which trigger fired. Kin REQ-SHELL-PERSISTENT-BOOT-RESTORE (the boot trigger and the shared arms), REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS (the invariant this closes the remaining seam of), REQ-SHELL-2 (the online/offline mutual exclusivity the edge path holds by stopping an armed watcher before it launches -- mid-generation a watcher CAN be armed for the casualty, unlike at boot).\"\n-required_stages = [\"doc\", \"impl\", \"unit\"]  # ACTIVATED CONDUIT W2 in the commit landing the evidence (doyle ruled shape (c) with conditions on releases#228; (a) not built, (b) rejected). doc = the CONTEXT.md `persistent` amendment plus KNOWN-HAZARDS 2.7 and MANIFEST.md Sleep/wake, all of which stated the boot sweep as the ONLY trigger and went false the moment a second trigger existed. MINTED CONDUIT W2 (todlando, 2026-08-28), releases#228.\n+title = \"Every down persistent shell of an owner is restored on its offline-to-online edge, including corpses launched after machine boot and instances with no launch stamp (releases#287 operator ruling supersedes the releases#228 boot-age gates). Both triggers share the same persistent-only, owner-online, birth-safe down-in-fact decision; a live launch awaiting bind is not duplicated. The edge is seeded before the boot sweep and bounds failed attempts to one per owner-online event. Stop an armed watcher before launching, and name the triggering event and every restored owner-qualified instance.\"\n+required_stages = [\"doc\", \"impl\", \"unit\"]\n+\n+[[requirements]]\n+id = \"REQ-SHELL-HEAL-ENUMERATION\"\n+title = \"The daemon's existing SHELL_RECORD_HEALED diagnostic enumerates every successfully healed instance per owner by owner-qualified identity, not only the first or a representative shell; healthy and already-healed instances produce no heal event (releases#287 rider).\"\n+required_stages = [\"doc\", \"impl\", \"unit\"]\n+\n+[[requirements]]\n+id = \"REQ-SHELL-RELINK-BINDING-RESULT\"\n+title = \"A successful local shell relink reports status=binding while awaiting the bind handshake, rather than describing the successful launch as offline; bind remains the transition to online (releases#287 rider).\"\n+required_stages = [\"doc\", \"impl\", \"int\"]\n \n [[requirements]]\n id = \"REQ-SHELL-KILL-AUTHENTICATED\"\n@@ -4166,7 +4182,7 @@ required_stages = [\"doc\", \"impl\", \"unit\", \"int\"]  # ACTIVATED BAROMETER W2 reope\n \n [[requirements]]\n id = \"REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS\"\n-title = \"A NODE RESTART MUST NOT PERMANENTLY STRAND EVERY `persistent` SHELL, AND THE EXISTING CASCADE TEST CANNOT SEE THAT IT DOES. Distinct from REQ-HAZARD-SHELL-STALE-ONLINE, which governs a force-kill DURING operation: this is a different trigger with a different safe remedy, and conflating them is how the restart class stays invisible. THE COVERAGE GAP IS THE HAZARD'S OWN SUBJECT, and it must be named here so no future reader mistakes the cascade test for coverage of this class: `rest_edges_cascade_shells_with_divergence` (spt-daemon resting.rs) is the cascade's only coverage and it walks the graceful path ONLY -- it SUSPENDS FIRST, and the suspend edge WRITES the offline record, so by the time the wake edge fires the instance is already in exactly the state the to_active arm requires. The restart shape -- record left online over a corpse, no rest edge ever -- is UNREACHABLE FROM THAT TEST BY CONSTRUCTION. The test is not weak; its SETUP SUPPRESSES THE FAILING ARM, so a reader asking whether the cascade is covered gets a true answer to the wrong question. This is the third suppressed-arm instance found in one week, and the pattern is the same each time: a fixture that establishes the precondition the code under test is supposed to establish. THE OWNER-FACING SURFACES HIDE THE FAULT, which is why it survived unreported for days: every display path (`shell list`, its --json twin, the shell-context render) routes through the ONE discovery seam, which DERIVES status and so correctly reports the binary as offline, while the only consumer whose decision matters -- the wake cascade -- reads the RECORDED field that no owner-facing surface shows. Absence of a visible symptom is therefore not evidence of absence of this fault, and a clean-looking context render must never be taken as proof the class did not occur; the correct falsifier is the on-disk record. Gate: impl -- the restart shape reaches a restoration path at all (leg b) and the record stops lying (leg a); unit -- the RESTART shape specifically, constructed WITHOUT a suspend edge: an instance whose record says online over a corpse that predates boot, with an online owner, is restored -- the row `rest_edges_cascade_shells_with_divergence` cannot express, asserted on a fixture that never calls the suspend path, since a fixture that suspends first re-creates the suppressed arm and proves nothing about this class.\"\n+title = \"A daemon-only or machine restart must not strand persistent shells. Heal stale online records and restore every down persistent instance when its owner is online, without boot-age or launch-stamp eligibility gates (releases#287 operator ruling). Birth-safe liveness still excludes a live binary, including one awaiting bind. The regression fixture must construct online records over corpses WITHOUT a suspend edge, including a corpse launched AFTER machine boot; the existing rest_edges_cascade_shells_with_divergence fixture suspends first and therefore suppresses the failing arm rather than covering it.\"\n required_stages = [\"doc\", \"impl\", \"unit\"]  # doc ACTIVATED BAROMETER W2 doc leg (doyle-ruled, this milestone): the KNOWN-HAZARDS 2.7 entry the mint note said was owed. AGENTS.md rule 4 makes every KNOWN-HAZARDS entry a REQ-HAZARD-* requirement and rule 1 makes the tag ride the same change as the evidence, so an untagged entry would be the untagged-evidence shape rule 1 bans. -- ACTIVATED BAROMETER W2 at impl+unit once BOTH legs landed -- deliberately not at leg (a), since the gate clause covers the boot sweep too and activating early would have let the requirement read satisfied while half its impl was unbuilt. MINTED BAROMETER W2 (todlando, 2026-08-02), releases#78 -- inactive-first. KNOWN-HAZARDS entry owed with the doc leg. The suppressed-arm sentence is REQUIRED text per the filing and doyle's dispatch: the hazard exists partly BECAUSE the existing test reads as coverage, so a requirement that did not name the gap would leave the same trap for the next reader.\n \n \n@@ -4386,6 +4402,7 @@ requirements = [\n [[groups]]\n name = \"cli-output-surface\"\n requirements = [\n+  \"REQ-SHELL-RELINK-BINDING-RESULT\",\n   \"REQ-BIND-HONEST-SELF-STAMP\",\n   \"REQ-CLI-1\",\n   \"REQ-CLI-2\",\n@@ -4550,6 +4567,7 @@ requirements = [\n [[groups]]\n name = \"endpoint-presence-truth\"\n requirements = [\n+  \"REQ-SHELL-HEAL-ENUMERATION\",\n   \"REQ-CONN-POISON-ATTRIBUTION\",\n   \"REQ-DAEMON-5\",\n   \"REQ-DAEMON-6\",","startLine":930,"lineNumbers":[930,931,932,933,934,935,936,937,938,939,940,941,942,943,944,945,946,947,948,949,950,951,952,953,954,955,956,957,958,959,960,961,962,963,964,965,966,967,968,969,970,971,972,973,974,975,976,977,978,979,980,981,982,983,984,985,986,987,988,989,990,991,992,993,994,995,996,997,998,999,1000,1001,1002,1003,1004,1005,1006,1007,1008,1009,1010,1011,1012,1013,1014,1015,1016,1017,1018,1019,1020,1021,1022,1023,1024,1025,1026,1027,1028,1029,1030,1031,1032,1033,1034,1035,1036,1037,1038,1039,1040,1041,1042,1043,1044,1045,1046,1047,1048,1049,1050,1051,1052,1053,1054,1055,1056,1057,1058,1059,1060,1061,1062,1063,1064,1065,1066,1067,1068,1069,1070,1071,1072,1073,1074,1075,1076,1077,1078,1079,1080,1081,1082,1083,1084,1085,1086,1087,1088]},"truncation":{"content":"-- **Offline alone does not arm the watcher.** A shell force-killed (or crashed)\n-  *during the current boot* is offline with no watcher running, and comes back\n-  only on demand — `spt shell relink <id>`, or a `shell cmd` that wakes it. This\n-  is deliberate: killing a shell to overwrite its exe is a routine deploy step,\n-  and a watcher armed underneath would re-lock the file. A cleanly-closed\n-  instance and a restart casualty both arm a watcher normally.\n+Lifecycle facts worth knowing before you rely on either:\n+\n+- **Persistent instances return after a daemon-only restart as well as a machine\n+  reboot.** The boot sweep and each owner offline→online edge restore every\n+  persistent instance that is down in fact while its owner is online. Launch\n+  age and missing launch stamps do not exclude it. The birth-safe process probe\n+  still protects a live binary, including a launch awaiting bind. Failed\n+  launches are retried on an owner-online edge, not every reconcile tick.\n+- **Only nonpersistent instances retain the force-kill freeze.** Their\n+  same-boot or undatable corpses arm no watcher; cleanly closed instances and\n+  dated pre-boot casualties remain watcher-eligible. Persistent instances\n+  bypass this freeze. To keep a persistent shell down, use `teardown` or a\n+  nonpersistent manifest rather than killing its process.\n+- The daemon's `SHELL_RECORD_HEALED:<owner>/<shell_id>` diagnostics enumerate\n+  every instance whose stale online record was corrected, not a representative\n+  shell per owner. Already-correct records produce no repeated notice.\n \n <!-- [doc->REQ-SHELL-RELINK-FORCE] -->\n - **`relink` refuses an instance that is still running; `--force` is the\n@@ -129,6 +121,11 @@ Two lifecycle facts worth knowing before you rely on either:\n   canonical id, the perch and its persisted state, and the alias all survive;\n   teardown + spawn is what churns the id and frees the mint slot.\n \n+  <!-- [doc->REQ-SHELL-RELINK-BINDING-RESULT] -->\n+  A successful local relink reports `status=binding (online at bind)`: the\n+  binary has launched and the bind handshake is still the online transition.\n+  It does not describe that successful launch as `offline`.\n+\n   Two things `--force` deliberately will **not** do:\n \n   - **It refuses on an `ephemeral` instance.** An ephemeral shell's close *is*\ndiff --git a/docs/KNOWN-HAZARDS.md b/docs/KNOWN-HAZARDS.md\nindex 8ab726b7..c8b94c6e 100644\n--- a/docs/KNOWN-HAZARDS.md\n+++ b/docs/KNOWN-HAZARDS.md\n@@ -76,19 +76,20 @@ Hard-won edge cases harvested from the sister project (`claude_skill_owl`, ~80 c\n ### 2.6 A shell instance's `online` is a recorded field with no death edge\n - **Failure:** a shell perch's `status` has exactly two writers — the bind handshake (→`online`) and `close_shell` (→`offline`) — and `close_shell` runs on the **link-break** path alone. A binary that dies **abruptly** (force-kill, crash, OOM) breaks no link, so nothing ever writes the flip: the record says `online` forever while the pid it names is gone. Every consumer of the recorded field then inherits the lie — `shell list` reports online indefinitely; `relink` (the one command that recovers the instance) refuses `SHELL_ALREADY_ONLINE`, gated on the very state that is wrong; and `shell cmd`'s wake-if-offline arm reads \"not offline\", skips the wake, and spools the command to a corpse — accepted happily, drained by nobody, **indistinguishable from a busy shell**. The only exit was `teardown` + `spawn`, which destroys the perch (losing the instance's persisted state) and frees the mint slot, so the instance returns under a **different canonical id** and every reference to the old name — docs, other agents' notes, tag addressing — silently points at nothing. The rename is the *visible* cost; the expensive one is **silent**: a consumer's perch-persisted cursor re-baselines on the fresh perch, so work arriving between the death and the re-bind is never scanned — not failed-and-retried, just never seen (alchemy's tag scanner, 2026-07-25). Recovery must therefore preserve the perch, not merely restore a running binary.\n - **Invariant:** ONLINE-ness is **derived**, never the recorded field alone: `status == online` **AND** the recorded `shell.pid` is not *provably* dead. Provably dead is narrow — pid parked, non-zero, `!is_process_alive`; pid absent, unparseable, or `0` (a broker-hosted spawn whose backend exposed no pid) reads **alive**, the same fail-toward-alive parity 2.5 holds, so a backend the resolver cannot see is never declared dead. **The recycled-pid caveat is narrowed since BAROMETER W2 and is now conditional on the stamp:** a launch parks a **birth stamp** beside the pid (`shell.launch.json`), and the probe runs the PAIR test — same pid, *different* start time ⇒ **gone**, so a pid the OS handed to an unrelated process no longer masks a dead instance. That case matters most across a restart, where the pid namespace resets and a corpse pid is routinely reused; without the pair test the corpse reads alive, the derived status reads `online`, and `relink` refuses `SHELL_ALREADY_ONLINE` on the one instance needing recovery. Where there is **no stamp to compare** (an instance last launched by a pre-W2 build, or a backend exposing no start time) the original stance is unchanged and unchanged in *direction*: fail toward alive, heal *missed* rather than *mis-fired*. Derivation NEVER promotes: a recorded-`offline` instance stays offline whatever its pid file says. Site class is part of the invariant: **derive** at the gates and renders (relink, the cmd wake arm, drive's drop-if-offline, `shelldisc::discover` = the single source of both `shell list` surfaces, the activity fan-out); keep the **recorded** field at the two writers and at the owner-suspend cascade, whose `close_shell` **is** the cleanup that would be skipped; and keep it at the wake reconciler's eligibility read, which since BAROMETER W2 additionally asks the **corpse-boot discriminant** (releases#78, doyle's amendment on comment 5156810267): the read stays on the recorded field, but \"no spontaneous relaunch\" is now a stated predicate rather than a side effect of the record being wrong — see the bullet below.\n-- **No spontaneous relaunch of a SAME-BOOT death (operator-ratified, flynn 2026-07-25; re-based BAROMETER W2):** the reconciler does NOT adopt an instance whose binary died during this boot. Every reason an operator stops a shell is a reason not to want it back a tick later: mid-deploy (a shared install dir makes \"kill the process\" routine on Windows, since a running exe cannot be overwritten) an auto-relaunch would run the **old** binary out of the file being replaced — worse than a failed install — and it turns quarantining a misbehaving shell into a restart loop. For those, recovery stays demand-driven: `relink`, or a `shell cmd` that wakes. `relink` probes **locally** rather than trusting a daemon sweep, so recovery holds with the daemon down.\n-- **What holds that open is no longer the stale record (BAROMETER W2):** until W2 the guarantee rested on an accident: a force-killed instance was held out of the eligible set because its record still said `online`, and the eligibility read took that field raw. 2.7 heals that record to the truth, which removes the accidental protection — so the guard had to move to a rule that *states* it. Eligibility is now recorded-`offline` **AND** (no corpse, **or** the corpse's recorded launch **predates the boot instant**). A steady-state force-kill leaves a corpse launched *after* boot, which the predicate can never accept, so the ruling is preserved **by construction** rather than by a second rule. A restart casualty is the one case the predicate does accept, and it is exactly the scope 2.7 was greenlit for. Where the discriminant **cannot be evaluated** — no launch stamp (an instance last launched by a pre-W2 build), or no boot oracle on the platform — a corpse-bearing instance is **not** eligible: leaving a shell down beats arming a relaunch of a binary an operator may have killed on purpose.\n+- **No spontaneous relaunch of a NONPERSISTENT same-boot death:** the releases#78 force-kill freeze still protects nonpersistent instances, so killing one to free its exe does not arm a watcher underneath an overwrite. Recovery remains demand-driven. **Persistent instances are exempt** (operator ruling, releases#287): a kill is not a request to keep a declared-persistent shell down; use `teardown` or a nonpersistent manifest instead.\n+- **Watcher eligibility:** recorded `offline`, no live binary awaiting bind, and either `persistent`, no corpse, or a corpse whose launch predates the machine boot instant. A nonpersistent corpse without a launch stamp or boot oracle remains frozen; persistent instances require neither to become eligible. Birth-safe probes prevent a recycled pid from masquerading as a live shell.\n - **spt-core mapping:** `spt_store::shellinfo::{shell_pid_provably_dead, effective_status, is_shell_online}` (the resolver — the shell-side twin of 2.5's `liveness.rs`, which gave *agent* perches exactly this and which shells never got); consumers in `linkhost::{relink_shell, prepare_drive, run_action}`, `shelldisc::discover`, `activity::observe_links`; the class-(c) eligibility read is `spt_daemon::shellwake::watcher_eligible` (+ `launch_predates_boot`, `BOOT_RESTORE_SLACK_MS`), and the birth stamp it rests on is `shellinfo::{ShellLaunch, record_shell_launch, read_shell_launch}`.\n - **Source:** spt-core, flynn's spt-alchemy field report (2026-07-25) — a deterministic recipe, not a race. Class-(c) re-based BAROMETER W2 (releases#78, doyle's ruling on comment 5156810267).\n \n-### 2.7 A node restart permanently strands every `persistent` shell — and the existing cascade test cannot see that it does  `[REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]`\n-- **Failure:** the contract's own sentence says a `persistent` shell is **online whenever its owner endpoint is online**, and a node restart is the one case where that promise was never kept. A machine death breaks no link, so `close_shell` never runs and 2.6's stale record survives the reboot; when the owner comes back online, the wake cascade reads the **recorded** field, sees `online`, and skips the relaunch of a binary that has not existed since the previous boot. Nothing else corrects it: the instance is down forever, with no error anywhere, until a human happens to run `relink`. Every `persistent` instance on the node is stranded by one restart, together.\n+### 2.7 A daemon or machine restart strands `persistent` shells — and a graceful cascade fixture suppresses the failing arm `[REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS]`\n+- **Failure:** a restart can leave an `online` record over a corpse without a link-break or suspend edge. The original boot sweep fixed only machine reboots: it rejected corpses launched after machine boot, so a daemon-only restart still stranded every persistent shell. The measured releases#287 specimen launched at machine boot +46s and stayed down for 27 hours after a daemon restart despite its owner being online.\n - **The owner-facing surfaces HIDE the fault**, which is why it survived unreported for days: every display path (`shell list`, its `--json` twin, the shell-context render) routes through the ONE discovery seam, which **derives** status (2.6) and so correctly reports the binary as offline — while the only consumer whose decision matters, the wake cascade, reads the RECORDED field that no owner-facing surface shows. Absence of a visible symptom is therefore **not** evidence of absence of this fault, and a clean-looking context render must never be taken as proof the class did not occur. The correct falsifier is the **on-disk record**.\n - **The existing cascade test cannot fail on it — its SETUP SUPPRESSES THE FAILING ARM.** `rest_edges_cascade_shells_with_divergence` establishes its fixture by driving the suspend path, which itself writes the `offline` the code under test is supposed to encounter — so the restart shape (a record still saying `online` over a corpse, reached with **no** suspend edge) is unreachable from that fixture BY CONSTRUCTION. The test is not weak; it gives a true answer to the wrong question, so a reader asking \"is the cascade covered?\" is told yes. This is the third suppressed-arm instance found in one week, and the pattern is the same each time: **a fixture that establishes the precondition the code under test is supposed to establish.**\n-- **Invariant:** a node restart must not strand a `persistent` instance. Two legs, and each is worthless alone. **(a) The record stops lying:** the daemon-side reconciler heals a recorded `online` that 2.6's derivation contradicts, writing `offline` — every cycle, **guarded on an actual change** (a heal that rewrote each tick would be a stream of identical writes and would destroy the record's mtime as a signal). **(b) The restart shape reaches a restoration path:** a once-per-daemon-generation boot sweep — and, for an owner that comes online after that sweep has run, an **owner offline→online edge** in the reconcile loop running the identical body (releases#228) — relaunches an instance when **all four** hold — the adapter section declares `persistent`; the **owner endpoint is online** (an offline owner is owed nothing — that is the line between healing a record and resurrecting a shell nobody is there to drive); the instance is down **in fact** (the derived read, made trustworthy across a pid reuse by 2.6's birth stamp); and its recorded launch **predates the boot instant**. No stamp ⇒ **not restored** (the sweep cannot prove it predates boot); no boot oracle on the platform ⇒ nothing restored at all. The boot instant is derived (Linux `btime`; Windows now − `GetTickCount64`), so a named slack constant absorbs derivation jitter and keeps a process launched moments after boot from reading as a corpse from the previous one. **The once-ness of the sweep is scheduling, not safety** (ruled releases#228): the predates-boot discriminant is what forbids relaunching a force-killed shell, and it is monotone within a machine boot — every launch restamps, a kill never does — so a second trigger evaluating the same conjuncts can only refuse more. The second trigger is an **edge**, not the level “owner is online”, which is what bounds a failing launch to one attempt per owner-online event instead of one per reconcile tick.\n-- **Regression shape (the row the suppressed fixture cannot express):** construct the restart shape **without a suspend edge** — an instance whose record says `online` over a corpse that predates boot, with an online owner — and assert it is restored. A fixture that suspends first re-creates the suppressed arm and proves nothing about this class.\n-- **spt-core mapping:** `spt_daemon::shellwake::{heal_stale_online_records, restore_persistent_shells_at_boot, restore_persistent_shells_on_owner_online, OwnerOnlineEdge, launch_predates_boot, BOOT_RESTORE_SLACK_MS}`; `spt_store::proc::boot_instant_ms`; `spt_store::shellinfo::{ShellLaunch, record_shell_launch, read_shell_launch}` (the birth stamp, parked at the two production pid-write sites the launch-site census identified).\n-- **Source:** spt-core, BAROMETER W2 (releases#78) — field specimen `liam/alchemy-0`, a genuinely stranded record: `online` over a dead pid, no launch stamp, days old.\n+- **Invariant:** a daemon-only or machine restart must not strand a persistent instance. **(a) Heal:** every reconcile cycle writes `offline` only where the derived birth-safe read contradicts recorded `online`; already-correct records are untouched. **(b) Restore:** the once-per-daemon-generation boot sweep and every owner offline→online edge run the same per-owner body. Relaunch every instance whose adapter declares `persistent`, whose owner is online, and whose binary is down in fact. There is **no machine-boot, launch-age, or stamp-presence eligibility gate**. A live binary awaiting bind is not down and must not be duplicated; missing birth evidence keeps the existing fail-toward-alive liveness behavior. The owner-online trigger remains an edge, bounding failed launches to one attempt per event.\n+- **Regression shape:** construct recorded `online` over a corpse **without a suspend edge**, including a launch **after machine boot**, then bring its owner online after a spent boot sweep. Every stranded persistent sibling must be restored; offline owners and nonpersistent instances are not restored. Separately defend live/binding exclusion and the nonpersistent watcher freeze. Suspending in setup suppresses the failing arm and proves nothing about this class.\n+- **Per-owner enumeration:** the existing `SHELL_RECORD_HEALED:<owner>/<shell_id>` daemon diagnostic names **every** successfully healed instance, not one representative. Healthy and already-healed records produce no event. The restore diagnostics likewise identify each attempted instance and distinguish boot from owner-online triggers. <!-- [doc->REQ-SHELL-HEAL-ENUMERATION] -->\n+- **spt-core mapping:** `spt_daemon::shellwake::{heal_stale_online_records, restore_persistent_shells_at_boot, restore_persistent_shells_on_owner_online, OwnerOnlineEdge}`; `shellhost::live_launch_winner`; `spt_store::shellinfo::{ShellLaunch, shell_pid_provably_dead}`. `launch_predates_boot` and `BOOT_RESTORE_SLACK_MS` now govern only nonpersistent watcher eligibility.\n+- **Source:** BAROMETER W2 (releases#78); releases#287 operator ruling (2026-09-08), superseding the persistent force-kill freeze and boot-instant conjunct.\n <!-- [doc->REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS] -->\n \n ---\ndiff --git a/docs/MANIFEST.md b/docs/MANIFEST.md\nindex b468481b..781d270c 100644\n--- a/docs/MANIFEST.md\n+++ b/docs/MANIFEST.md\n@@ -788,7 +788,13 @@ types = [\"image\", \"sound\", \"event\"]   # \"event\" = arbitrary descriptive payloads\n \n The capability + sensory vocabularies live **here**, not on the perch — spt-core resolves them by `adapter_name`. Command-delivery (`command_receipt`) reuses the agent inject-input modes. The shell binary binds via `api bind` (type=Shell, owner from the link) and pushes sensory via `api emit --type <t>`.\n \n-**Sleep/wake:** when offline **and eligible** (+ `wake_command` set), spt-core runs the wake-watcher (mutually exclusive with the shell binary); `exit(wake-opcode)` → wake resolution; crash-exit → respawn with exponential backoff + give-up (until the next relink/launch). **The wake opcode is exit code 86.** The wake-watcher template fills `{id}`, `{adapter_name}`, `{link_token}`, `{adapter_dir}` (NOT `{perch_dir}` — that key is spawn-only; the waker fill site carries no perch context), and its `{link_token}` is minted *unparked* — an offline link has no live credential (the close retired it); a waker wakes by exit code, never by driving the link. **Eligibility (BAROMETER W2, releases#78):** a watcher is armed for an instance recorded `offline` that either carries no corpse (cleanly closed, or never launched) or whose corpse's recorded launch **predates the current boot** — a restart casualty. An instance whose binary died **during this boot** (force-kill, crash) arms no watcher and returns only on demand (`relink`, or a waking `shell cmd`), because killing a shell to overwrite its exe is a routine deploy step and a watcher armed underneath would re-lock the file; an instance whose corpse cannot be dated (no launch stamp, or no boot oracle on the platform) is treated the same way rather than guessed at. `persistent` auto-onlines the shell with its owner, and across a **node restart** — where nothing breaks the link, so the instance's record survives the reboot still claiming it is up — the daemon heals that record and a **once-per-daemon-generation boot sweep** relaunches the instance if its adapter is `persistent`, its owner endpoint is online, it is down in fact, and its recorded launch predates the boot instant. Because daemons boot before endpoints do, the same four conditions are re-run for any owner that goes **offline→online since the last reconcile pass** (releases#228) — without that edge, an owner that was still down when the sweep ran keeps its shells down until someone relinks by hand. It is an edge and not the level “the owner is online” so that a failing launch is attempted once per owner-online event rather than once per tick. `can_shutdown` authorizes `api owner-shutdown`. Cross-node *fresh-spawn* to wake an owner requires the owner's `shell_wake_spawn_anywhere` settings flag (pre-consent).\n+**Sleep/wake:** when offline **and eligible** (+ `wake_command` set), spt-core runs the wake-watcher, mutually exclusive with a live shell binary including one awaiting bind. Exit code **86** is the wake opcode; other exits respawn with exponential backoff and eventual give-up until the next shell launch. Templates fill `{id}`, `{adapter_name}`, `{link_token}`, `{adapter_dir}` (not `{perch_dir}`, which is spawn-only). The watcher's link token is minted *unparked*: it wakes by exit code, never by driving an offline link.\n+\n+**Persistent restoration:** on the once-per-daemon-generation boot sweep and each owner offline→online edge, spt-core restores **every** down persistent instance whose owner is online. This covers machine reboots and daemon-only restarts. The decision uses birth-safe liveness, not machine boot time or recorded launch age; a missing launch stamp does not disqualify a down instance. A live launch awaiting bind is not duplicated. The edge bounds failed launches to one attempt per owner-online event rather than every reconcile tick. Stale online records are healed only on an actual change. <!-- [doc->REQ-SHELL-PERSISTENT-BOOT-RESTORE] --> <!-- [doc->REQ-SHELL-OWNER-ONLINE-RESTORE] -->\n+\n+**Watcher eligibility:** recorded-offline instances with no live binary are eligible if persistent, cleanly closed/never launched (no corpse), or a dated pre-boot corpse. The same-boot force-kill freeze applies **only to nonpersistent** instances, whose undatable corpses also remain frozen. Persistent instances bypass that freeze (releases#287 operator ruling); keeping one down requires `teardown` or a nonpersistent manifest, not a kill.\n+\n+`can_shutdown` authorizes `api owner-shutdown`. Cross-node *fresh-spawn* to wake an owner requires the owner's `shell_wake_spawn_anywhere` settings flag (pre-consent).\n \n **Instantiation governance:** `require_approval` gates `shell spawn` (reusing the consent grant store — `remembered` persists a grant via allow-always, `always` prompts every spawn); `max_instances_per_owner` + `over_cap` cap how many instances one owner may hold (online + offline both count toward the cap). An instance's **alias** is **not** a manifest field — it is per-instance runtime state on the shell perch's `info.json` (set at `shell spawn --alias` / `shell rename`), since the manifest describes the *adapter*, not individual instances.\n \ndiff --git a/traceable-reqs.toml b/traceable-reqs.toml\nindex 21917ea9..a3039948 100644\n--- a/traceable-reqs.toml\n+++ b/traceable-reqs.toml\n@@ -3183,6 +3183,12 @@ required_stages = [\"doc\", \"impl\", \"unit\", \"int\"]  # W1 ACTIVATED at mint (doyle\n # satisfying. ---\n [[requirements]]\n id = \"REQ-HAZARD-SHELL-STALE-ONLINE\"\n+# AMENDED releases#287, operator ruling 2026-09-08: the historical title below's\n+# same-boot/undatable-corpse freeze and no-spontaneous-relaunch clauses now apply\n+# ONLY to NONPERSISTENT instances. Persistent restoration and watcher eligibility\n+# follow REQ-SHELL-PERSISTENT-BOOT-RESTORE / REQ-SHELL-OWNER-ONLINE-RESTORE, without\n+# launch-age or stamp-presence gates. Birth-safe live/binding exclusion remains.\n+# The liveness, identity/state preservation and durable-channel clauses stand.\n title = \"A shell instance's ONLINE-ness is DERIVED (recorded status AND its recorded `shell.pid` not provably dead), never the recorded `status` field alone — an abruptly-dead binary (force-kill, crash, OOM: no link-break, so `close_shell`'s offline flip never runs) must not read online forever. The shell-side twin of REQ-HAZARD-DAEMON-HOSTED-LIVENESS, which gave AGENT perches exactly this resolver and which shells never got. PROVABLY DEAD is the narrow discriminant: `shell.pid` present AND parses non-zero AND `!is_process_alive` — pid absent, unparseable, or 0 (a broker-hosted spawn whose backend exposed no pid records 0) reads ALIVE, the same interim-parity/fail-toward-alive stance `liveness.rs` already holds, so a pid-less backend is NEVER falsely declared dead. Recycled-pid caveat, accepted at mint: a reused pid reads alive, so the heal is missed, never mis-fired — the failure direction is 'stays stale', never 'kills a live instance'. SITE CLASSIFICATION IS PART OF THE REQUIREMENT (authoritative cfg(test)-excluded census at mint = 8 status reads, 3 classes — do NOT blanket-swap the predicate): (a) DERIVED — relink's already-online refusal (the gate that made recovery impossible), the `shell cmd` wake-if-offline arm (which silently spooled to a corpse), the drive drop-if-offline branch, `shelldisc::discover` (the single source of BOTH `shell list` renders, text + --json), and the activity fan-out's online filter; (b) RAW status, deliberately — `bind_shell_by_token`/`close_shell` (the WRITERS) and `cascade_owner_edge`'s suspend-close arm, where routing through the resolver would SKIP the close that is itself the cleanup, removing a heal path; (c) RAW recorded status AND the corpse-boot discriminant — the `shellwake::reconcile_once` watcher-eligibility read. AMENDED BAROMETER W2 (doyle, releases#78 comment 5156810267): this class's original text — 'RAW status, pending an operator ruling', with no spontaneous relaunch of anything — IS STRUCK, because it rested on a fact leg (a) deliberately ends. What actually held a force-killed instance out of the eligible set was the STALE `online` record itself; healing that record to the truth removes the accidental protection, so the guard had to move to a rule that STATES it. Eligibility is now recorded `offline` AND (no corpse, or the corpse's launch PREDATES THE BOOT INSTANT), reusing leg (b)'s `launch_predates_boot` and its slack constant verbatim. NO SPONTANEOUS RELAUNCH OF SAME-BOOT DEATHS: flynn's mid-deploy forcing case (an operator kills the process precisely to free the exe for overwrite — shared install dir ⇒ 'kill the process' is a routine deploy step) is preserved BY CONSTRUCTION, since a force-kill during steady state leaves a corpse launched AFTER boot, which the predicate can never accept; for those, recovery stays demand-driven (an unblocked `relink`, or a `shell cmd` that wakes). Boot-restoration of PRE-boot casualties is enabled, per #78 leg (b)'s operator-greenlit scope and this ruling (#78 thread). CORPSE IS DEFINED ON THE PARKED STAMPS, never on a bare pid re-probe: a RECYCLED pid — alive, but its native start stamp mismatching the one leg (a) parked — is corpse-equivalent because our process is dead, and the parked ABSOLUTE launch stamp is what decides pre/post-boot. Without that, the heal-then-recycle race reads 'alive ⇒ no corpse' on the next tick and walks back into the forbidden shape through a side door. Deriving is still WRONG at this site and this is NOT a swap to class (a): the census stays 8 reads / 3 classes — the site reclassifies, it does not multiply, the do-NOT-blanket-swap sentence above stands, and `shell list` still tells the truth by derivation (class (a), untouched). LOCALITY SYMMETRY (field-caught by flynn's leg-2 run, 2026-07-25): 'a shell cmd that wakes' holds for BOTH the local CLI cmd and the cross-node serve — at mint the wake arm lived in the SHELL_LINK_CMD serve handler only (module docs scoped wake-if-offline to remote cmd), so a LOCAL cmd against a corpse spooled silently and nothing woke, the exact accepted-happily/drained-by-nobody shape this req exists to kill; the arm is now the shared `linkhost::wake_if_offline_persistent`, called by both, and the derived status READ inside it is the same single census site as before (the census of 8 reads/3 classes is unchanged — the read moved, it did not multiply). ROTATION CARRIES THE DURABLE CHANNEL (flynn's frame-loss field catch, 2026-07-25, both trials): a spooled command frame is MAC-stamped under the link token current at SPOOL time, and the drain is a raw destructive passthrough — the shell verifies against its CURRENT key — so a relink's token mint orphaned every pending frame: drained to the woken binary, failed verify, correctly discarded BY THE ADAPTER, lost permanently; the wake-triggering command itself was the frame the wake lost (#23 armed the wake and never answered; #24/#19 spooled after rotation and drained in order). The fix is three-layered (todlando's A1/A2 race+crash hazards addressed at mint): (1) the token mint re-stamps all pending rows old-key→new-key (idempotent, crash-partial converges — an old-stamped remainder is converted by the next rotation or drain); (2) the rotated-out token is STASHED (`link.token.retired`, never a live credential — bind resolves only the parked file) so (3) the drain paths give any race straggler a second-chance re-stamp at delivery (`restamp_for_drain`), while frames verifying under neither key pass through untouched for the adapter to refuse exactly as before. Rows are selected by MAC verification against the threaded new key, never by token-snapshot equality (the D-2 class). MOCK CONFORMANCE IS A PROPERTY OF THE RIG, not a patch (todlando A4): mock-shell MUST verify inbound frame MACs exactly as the public contract demands of a real adapter and drop failures loudly — a mock that accepts what the field rejects is a broken rig, and that exact divergence (credulous mock) is how the frame loss passed the e2e while failing in the field. The adapter-side discard behavior is CORRECT and must never be softened to paper over the spool side. Relink additionally probes LOCALLY rather than trusting a daemon sweep, so recovery holds with the daemon down. RECOVERY PRESERVES CONSUMER STATE (the property consumers actually depend on, flynn 2026-07-25): 'same canonical id, same perch' exists so that state a consumer PERSISTED IN THE PERCH survives the recovery — an adapter's repo binding, a scanner's cursor. The teardown+spawn workaround destroyed exactly that, and its worst cost was SILENT, not the rename: alchemy's tag cursor re-baselines at the digest tip, so tags written between the kill and the re-bind are never scanned — not failed-and-retried, just never seen. A loud failure gets retried; a silent one does not. Gate: int — force-kill a bound persistent instance's process, then prove (1) `shell list` reads offline, (2) `relink` succeeds instead of SHELL_ALREADY_ONLINE and the SAME canonical id + perch survive, carrying perch-persisted consumer state with them (no teardown+spawn, no id churn, no re-baselined cursor), (3) no spontaneous relaunch occurs while the instance sits dead and undriven. FIELD VERDICT — PASS END-TO-END (flynn, alchemy-0, delivered 2026-07-26; v0.43.0, counter 77): every gate leg held in the field, on a record the pre-fix code had already poisoned. (1) `shell list` read OFFLINE while the on-disk info.json still said status=online — the daemon DERIVED offline from the corpse pid (29036 absent from the process table); the record was never corrected and did not need to be. (2) relink ADMITTED, no SHELL_ALREADY_ONLINE — verbatim SHELL_RELINKED:alchemy-0 owner=flynn pid=38644 status=offline, list online thereafter. (3) identity/state integrity: same canonical id, same owner, same perch; repo binding byte-identical (token_provenance=gh-cli); armed=true preserved; no spontaneous relaunch across the ~6.5h dead window. CURSOR SCOPING, ruled at closure (doyle 2026-07-26) — the one non-byte-identical field: the tag cursor moved 206158430541→210453397553 (gen:seq 48:333→49:49) because the recovery crossed a DAEMON restart that slid the retained digest window past the armed cursor. NOT a gap and NOT this req's property failing: spt-core raised after_predates_window and the consumer took its specced armed-cursor branch (alchemy REQ-TAG-SCANNER missed-rows leg — warn the owner, never silently re-baseline an armed cursor; alchemy src/tags.rs, two unit tests) — the FIRST FIELD FIRING of a path previously exercised only by unit tests. The 'no re-baselined cursor' gate property is hereby scoped: byte-identity holds across SHELL death within one daemon generation (the prior LEG 5 measurement, one daemon instance, uncontradicted — different scenario, generation never moved); across a daemon restart that slides the window the contract is DETECT-AND-ANNOUNCE, not preserve — corroborated independently by a second record's cursor decomposing under the same gen:seq scheme (4294967486 = 1:190). Practical exposure in this event: nil (every agent stood down through the gap; no tag traffic to miss). SECOND SPECIMEN HELD, not spent (doyle ruling 2026-07-26): liam's record — precondition verified by flynn (pid 35464 a CLEAN corpse, not recycled, so the accepted-at-mint recycled-pid caveat does not apply; same poisoned status=online shape; binding intact; cursor armed at 4294967486) — is deliberately NOT relinked while its owner is absent: the admit gate is already field-proven, a relink measurably re-baselines the cursor across a daemon-restart boundary, and liam's is the only untouched specimen with the longer (~8.5h) gap — spending it would destroy the cursor datum liam's own return produces, to buy a redundant second sample of a gate already proven to admit. It waits for liam. SPECIMEN HARVESTED (liam's own return, 2026-07-26 — operator-authorized endpoint start executed by doyle; pre-state snapshotted by doyle AND independently re-verified on disk by liam before touching anything): (1) relink ADMITTED at the ~8.5h gap — verbatim SHELL_RELINKED:alchemy-0 owner=liam pid=47608 status=offline (online at bind); the poisoned record did not refuse; info.json online is now TRUE. (2) integrity FULL PASS: id unchanged (no new mint slot), perch onlined at bind, repo binding byte-identical (token_provenance gh-cli), and the ROTATION-CARRIES-THE-DURABLE-CHANNEL machinery observed live in the field for the first time — link.token.retired present beside the fresh mint, and a post-relink round trip proven end-to-end (shell cmd spooled -> engine executed -> reply home on the owner's perch), not merely an outbound spool. (3) cursor, the sharpest datum of the harvest: BYTE-IDENTICAL through the relink itself (4294967486 armed at pid-change instant, doyle's file watch) — the relink preserves; the SCANNER moves. First scan re-baselined 1:190 -> 2:0 silently, and liam PROVED no-fire-by-precondition rather than assuming: the daemon restart bumped the generation but EVICTED NOTHING (window floor 4294967464 < cursor 4294967486, every gen-1 row past the cursor still retained), so after_predates_window is correctly false at every link of the chain (filter_after floor test -> emit-only-when-true -> unwrap_or(false)) and the missed-rows announce had no precondition to fire on. Flynn's earlier gen-bump reading is refined by this: a restart alone does not slide the window; flynn's announce fired because HIS window had genuinely slid. Scope note carried: a valid negative — the fix's announce chain verified link-by-link to contract on a specimen where it correctly stayed quiet. RESIDUE from the harvest (liam's find, alchemy lane, no core REQ): permanently-unsealable turns from a dark session are jumped by scanner cursor advance with no possible missed-rows notice (a generation bump is not a slide); ruled scanner-side announce for input-bearing null-seq jumps, pseudo-turns jump silently by design; core's per-turn input/input_seq/partial already suffice to discriminate.\"\n required_stages = [\"doc\", \"impl\", \"unit\", \"int\"]  # ACTIVATED at mint (doyle 2026-07-25) — field-blocking; flynn eats id churn every alchemy release until it lands. LANDED v0.43.0 (counter 77, PR #76/#77); field-verified 2026-07-26 — see FIELD VERDICT in title.\n \n@@ -4151,13 +4157,23 @@ required_stages = [\"doc\", \"impl\", \"unit\"]  # ACTIVATED BAROMETER W1 at impl+unit\n \n [[requirements]]\n id = \"REQ-SHELL-PERSISTENT-BOOT-RESTORE\"\n-title = \"A `persistent` SHELL COMES BACK AFTER A NODE RESTART, WHICH IS THE ONE CASE THE CONTRACT'S OWN SENTENCE PROMISES AND THE IMPLEMENTATION NEVER HANDLED. CONTEXT.md states it plainly -- a persistent shell is automatically brought online whenever its owner endpoint is online -- and the shipped code honors that across a graceful suspend->wake rest edge ONLY. A machine restart produces no such edge and there is no boot-time restoration at all, so after a reboot every persistent shell on the node stays down until a human runs `relink` by hand, with no surface saying so. Field, HFENDULEAM 2026-08-01: the fleet came up at 01:11 and NOT ONE persistent shell came back; the two instances that did were relaunched by a human at 01:27, sixteen minutes later. TWO INDEPENDENT GATES, both of which must fall or the promise stays broken. GATE 1, no restoration path: the five launch sites all require either an explicit command, a rest edge, or a `wake_command` the adapter may deliberately not declare, and daemon boot starts only the servicehost sweep and the wake reconciler. GATE 2, the stale record blocks the cascade even when an edge DOES arrive: the to_active arm gates on the RECORDED status being offline, and a machine death breaks no link, so the close that would write offline never runs and the record says online over a corpse forever. LEG (a) HEALS THE RECORD TO THE TRUTH, with a BIRTH-STAMPED probe rather than bare pid liveness -- a pid namespace resets at boot, so a corpse pid can be REUSED by an unrelated process, and a reused pid reads alive, makes the derived status read online, and makes `relink` refuse SHELL_ALREADY_ONLINE: the one command that recovers the instance refusing on the very state that is wrong. TWO STAMPS ARE PARKED, NOT ONE, AND THE REASON IS A PLATFORM ASYMMETRY THAT WOULD OTHERWISE SHIP SILENTLY BROKEN ON LINUX: `process_started_at` returns a Windows FILETIME (absolute, comparable across boots) but on Linux returns field 22 of /proc/<pid>/stat, which is JIFFIES SINCE BOOT and resets every boot, so a value recorded before a restart is indistinguishable from a fresh one after it. The NATIVE stamp is therefore used for one question only -- the same-pid-different-process PAIR test within a single boot, mirroring `liveness::relay_liveness` -- and an ABSOLUTE epoch-ms launch stamp is parked beside it for the cross-boot question. LEG (b) IS A ONCE-PER-DAEMON-GENERATION BOOT SWEEP relaunching persistent instances whose owners are online, and its discriminant is THE LAUNCH STAMP PREDATES THE BOOT INSTANT, which preserves the class-c force-kill ruling BY CONSTRUCTION rather than by a second rule: a force-kill during steady-state operation leaves a corpse launched AFTER boot, so it is never spontaneously relaunched and the mid-deploy exe-overwrite hazard that ruling protects is untouched. A NAMED SLACK CONSTANT absorbs boot-instant jitter so a process launched moments after boot is never misclassified as a corpse. Gate: impl -- the two stamps parked at the two production pid-write sites (the census is authoritative: five LAUNCH sites funnel through exactly TWO pid writes), the birth-stamped death probe, the record heal in the daemon-side reconciler guarded on an actual change of the recorded value, the boot-instant helper on both platforms, and the boot sweep with its predates-boot discriminant plus slack; unit -- a reused pid (same pid, different birth stamp) reads GONE where bare-pid liveness reads alive (the row the shipped probe cannot pass), a same-pid-same-birth process reads HELD, the heal writes offline over a stale online record and writes NOTHING when the record already matches the truth (the guard asserted by counting writes, since a heal that rewrites every cycle passes any state-only assertion), a corpse launched BEFORE the boot instant is eligible for restore while one launched AFTER it is NOT (the class-c pair, asserted together since either alone passes an implementation that restores everything or nothing), and an instance whose OWNER is offline is not restored even though its corpse predates boot.\"\n-required_stages = [\"doc\", \"impl\", \"unit\"]  # doc ACTIVATED BAROMETER W2 doc leg (doyle-ruled, this milestone): the CONTEXT.md `persistent` amendment states the boot sweep and its four conjunct gates incl. no-stamp => not restored, which IS this requirement's doc evidence; leaving the stage inactive would have made it evidence the registry cannot see.\n+title = \"Persistent shells are restored on each daemon-generation boot sweep whenever their owner is online and they are down in fact, regardless of machine boot time or recorded launch age (releases#287 operator ruling). Missing launch stamps do not disqualify a down instance. Preserve birth-safe liveness, do not duplicate a live launch awaiting bind, and heal stale online records only on an actual change. The force-kill freeze applies only to nonpersistent instances. Regression: restore same-boot and unstamped corpses without a suspend edge; refuse offline owners, nonpersistent instances and live binaries.\"\n+required_stages = [\"doc\", \"impl\", \"unit\"]\n \n [[requirements]]\n id = \"REQ-SHELL-OWNER-ONLINE-RESTORE\"\n-title = \"A `persistent` SHELL WHOSE OWNER ENDPOINT WAS NOT YET ONLINE WHEN THE BOOT SWEEP RAN IS RESTORED WHEN THAT OWNER COMES ONLINE, RATHER THAN STAYING DOWN UNTIL A HUMAN RELINKS IT. ROOT (releases#228): daemons boot before endpoints do, so the boot sweep of REQ-SHELL-PERSISTENT-BOOT-RESTORE correctly SKIPS an owner that is not up yet -- an offline owner is owed nothing -- and then nothing ever revisits it: the sweep is once per daemon generation, and bringup from offline produces no rest edge, so the ADR-0048 cascade never fires either. The shell stays down until a `shell cmd` wakes it or an operator relinks by hand, and whether a given node restart strands its shells is an ordering accident. THE ONCE-NESS IS SCHEDULING, NOT SAFETY (doyle, releases#228, re-measured at head): the `launch_predates_boot` corpse-boot discriminant is what preserves the force-kill ruling, and it is MONOTONE WITHIN A MACHINE BOOT -- every launch path restamps `launched_ms`, a kill never restamps, `boot_ms` is the machine boot instant -- so re-evaluating the conjuncts later in a generation can only refuse MORE, never less. FIX: the reconcile loop carries the SAME per-owner restore body the boot sweep runs, triggered on owner offline->online transitions since the last pass. THE TRIGGER IS AN EDGE, NOT A LEVEL, and the distinction is the whole bound: success self-limits because the launch restamps the instance out of eligibility, but a FAILING launch under a level trigger -- owner is online -- retries every five seconds for as long as the owner stays up, while an edge bounds it to one attempt per owner-online event. THE REFUSAL ARMS ARE CARRIED BY CONSTRUCTION, not by restatement: both triggers call one shared per-owner body, so persistent-only, owner online, down in fact via the derived read with the pair test, stamp present, launch predates boot, and no boot oracle means restore nothing can never drift apart between the two paths. THE EDGE STATE IS SEEDED BEFORE THE SWEEP, deliberately: seeded after, an owner that comes online in the window between the sweep and the first tick reads as already-online and never fires -- the exact miss this requirement exists to close -- while seeding before can at worst re-attempt an instance the sweep already restored, which the RESTAMPED BIRTH STAMP refuses -- and NOT the down-in-fact arm, which is the arm that reads like the answer and is not: a launch is not the online switch, so a just-restored instance keeps an `offline` record until its bind-shell handshake and still reads as down in fact; what refuses the second attempt is that the launch moved `launched_ms` past the boot instant. Each trigger names ITSELF in its event line, because the field evidence that diagnosed this ticket was a count of those events and a shared name would leave the next investigation unable to say which trigger fired. Kin REQ-SHELL-PERSISTENT-BOOT-RESTORE (the boot trigger and the shared arms), REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS (the invariant this closes the remaining seam of), REQ-SHELL-2 (the online/offline mutual exclusivity the edge path holds by stopping an armed watcher before it launches -- mid-generation a watcher CAN be armed for the casualty, unlike at boot).\"\n-required_stages = [\"doc\", \"impl\", \"unit\"]  # ACTIVATED CONDUIT W2 in the commit landing the evidence (doyle ruled shape (c) with conditions on releases#228; (a) not built, (b) rejected). doc = the CONTEXT.md `persistent` amendment plus KNOWN-HAZARDS 2.7 and MANIFEST.md Sleep/wake, all of which stated the boot sweep as the ONLY trigger and went false the moment a second trigger existed. MINTED CONDUIT W2 (todlando, 2026-08-28), releases#228.\n+title = \"Every down persistent shell of an owner is restored on its offline-to-online edge, including corpses launched after machine boot and instances with no launch stamp (releases#287 operator ruling supersedes the releases#228 boot-age gates). Both triggers share the same persistent-only, owner-online, birth-safe down-in-fact decision; a live launch awaiting bind is not duplicated. The edge is seeded before the boot sweep and bounds failed attempts to one per owner-online event. Stop an armed watcher before launching, and name the triggering event and every restored owner-qualified instance.\"\n+required_stages = [\"doc\", \"impl\", \"unit\"]\n+\n+[[requirements]]\n+id = \"REQ-SHELL-HEAL-ENUMERATION\"\n+title = \"The daemon's existing SHELL_RECORD_HEALED diagnostic enumerates every successfully healed instance per owner by owner-qualified identity, not only the first or a representative shell; healthy and already-healed instances produce no heal event (releases#287 rider).\"\n+required_stages = [\"doc\", \"impl\", \"unit\"]\n+\n+[[requirements]]\n+id = \"REQ-SHELL-RELINK-BINDING-RESULT\"\n+title = \"A successful local shell relink reports status=binding while awaiting the bind handshake, rather than describing the successful launch as offline; bind remains the transition to online (releases#287 rider).\"\n+required_stages = [\"doc\", \"impl\", \"int\"]\n \n [[requirements]]\n id = \"REQ-SHELL-KILL-AUTHENTICATED\"\n@@ -4166,7 +4182,7 @@ required_stages = [\"doc\", \"impl\", \"unit\", \"int\"]  # ACTIVATED BAROMETER W2 reope\n \n [[requirements]]\n id = \"REQ-HAZARD-RESTART-STRANDS-PERSISTENT-SHELLS\"\n-title = \"A NODE RESTART MUST NOT PERMANENTLY STRAND EVERY `persistent` SHELL, AND THE EXISTING CASCADE TEST CANNOT SEE THAT IT DOES. Distinct from REQ-HAZARD-SHELL-STALE-ONLINE, which governs a force-kill DURING operation: this is a different trigger with a different safe remedy, and conflating them is how the restart class stays invisible. THE COVERAGE GAP IS THE HAZARD'S OWN SUBJECT, and it must be named here so no future reader mistakes the cascade test for coverage of this class: `rest_edges_cascade_shells_with_divergence` (spt-daemon resting.rs) is the cascade's only coverage and it walks the graceful path ONLY -- it SUSPENDS FIRST, and the suspend edge WRITES the offline record, so by the time the wake edge fires the instance is already in exactly the state the to_active arm requires. The restart shape -- record left online over a corpse, no rest edge ever -- is UNREACHABLE FROM THAT TEST BY CONSTRUCTION. The test is not weak; its SETUP SUPPRESSES THE FAILING ARM, so a reader asking whether the cascade is covered gets a true answer to the wrong question. This is the third suppressed-arm instance found in one week, and the pattern is the same each time: a fixture that establishes the precondition the code under test is supposed to establish. THE OWNER-FACING SURFACES HIDE THE FAULT, which is why it survived unreported for days: every display path (`shell list`, its --json twin, the shell-context render) routes through the ONE discovery seam, which DERIVES status and so correctly reports the binary as offline, while the only consumer whose decision matters -- the wake cascade -- reads the RECORDED field that no owner-facing surface shows. Absence of a visible symptom is therefore not evidence of absence of this fault, and a clean-looking context render must never be taken as proof the class did not occur; the correct falsifier is the on-disk record. Gate: impl -- the restart shape reaches a restoration path at all (leg b) and the record stops lying (leg a); unit -- the RESTART shape specifically, constructed WITHOUT a suspend edge: an instance whose record says online over a corpse that predates boot, with an online owner, is restored -- the row `rest_edges_cascade_shells_with_divergence` cannot express, asserted on a fixture that never calls the suspend path, since a fixture that suspends first re-creates the suppressed arm and proves nothing about this class.\"\n+title = \"A daemon-only or machine restart must not strand persistent shells. Heal stale online records and restore every down persistent instance when its owner is online, without boot-age or launch-stamp eligibility gates (releases#287 operator ruling). Birth-safe liveness still excludes a live binary, including one awaiting bind. The regression fixture must construct online records over corpses WITHOUT a suspend edge, including a corpse launched AFTER machine boot; the existing rest_edges_cascade_shells_with_divergence fixture suspends first and therefore suppresses the failing arm rather than covering it.\"\n required_stages = [\"doc\", \"impl\", \"unit\"]  # doc ACTIVATED BAROMETER W2 doc leg (doyle-ruled, this milestone): the KNOWN-HAZARDS 2.7 entry the mint note said was owed. AGENTS.md rule 4 makes every KNOWN-HAZARDS entry a REQ-HAZARD-* requirement and rule 1 makes the tag ride the same change as the evidence, so an untagged entry would be the untagged-evidence shape rule 1 bans. -- ACTIVATED BAROMETER W2 at impl+unit once BOTH legs landed -- deliberately not at leg (a), since the gate clause covers the boot sweep too and activating early would have let the requirement read satisfied while half its impl was unbuilt. MINTED BAROMETER W2 (todlando, 2026-08-02), releases#78 -- inactive-first. KNOWN-HAZARDS entry owed with the doc leg. The suppressed-arm sentence is REQUIRED text per the filing and doyle's dispatch: the hazard exists partly BECAUSE the existing test reads as coverage, so a requirement that did not name the gap would leave the same trap for the next reader.\n \n \n@@ -4386,6 +4402,7 @@ requirements = [\n [[groups]]\n name = \"cli-output-surface\"\n requirements = [\n+  \"REQ-SHELL-RELINK-BINDING-RESULT\",\n   \"REQ-BIND-HONEST-SELF-STAMP\",\n   \"REQ-CLI-1\",\n   \"REQ-CLI-2\",\n@@ -4550,6 +4567,7 @@ requirements = [\n [[groups]]\n name = \"endpoint-presence-truth\"\n requirements = [\n+  \"REQ-SHELL-HEAL-ENUMERATION\",\n   \"REQ-CONN-POISON-ATTRIBUTION\",\n   \"REQ-DAEMON-5\",\n   \"REQ-DAEMON-6\",","truncated":true,"truncatedBy":"lines","totalLines":160,"totalBytes":51269,"outputLines":159,"outputBytes":51269,"lastLinePartial":false,"firstLineExceedsLimit":false},"meta":{"source":{"type":"internal","value":"artifact://112"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":1089,"totalBytes":51269,"outputLines":159,"outputBytes":51269,"maxBytes":81408,"shownRange":{"start":930,"end":1088},"nextOffset":1089}}},"isError":false,"timestamp":1789001247068}}
{"type":"message","id":"a8d6d422","parentId":"ea264a5f","timestamp":"2026-09-10T00:47:28.665Z","message":{"role":"toolResult","toolCallId":"call_DLi0WVShUA9eTpzx0LdLbfIM|fc_077f3569365e313d016aa1fe1eef8087d184158b76da4c7c44","toolName":"bash","content":[{"type":"text","text":"commit cb756deeed9deafa52a0cbf913a87252c8679c16\nAuthor:     Reavo End <decidel@gmail.com>\nAuthorDate: Wed Sep 9 17:25:51 2026 -0700\nCommit:     Reavo End <decidel@gmail.com>\nCommitDate: Wed Sep 9 17:25:51 2026 -0700\n\n    fix(shell): restore persistent instances after daemon restart\n    \n    Remove machine-boot and launch-stamp eligibility gates from persistent restoration and watchers. Keep birth-safe binding custody and the nonpersistent force-kill freeze, enumerate every healed instance, and report relink success as binding.\n    \n    Refs BigscreenVR/spt-bs-releases#287\n    \n    Co-authored by: todlando\n\n\nWall time: 1.61 seconds"}],"details":{"timeoutSeconds":300,"wallTimeMs":1613.8675999999978},"isError":false,"timestamp":1789001248664}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_ADe4acYv3CowxlWzJJGwaFjk|fc_077f3569365e313d016aa1fe2c7d2887d187e705ac7fb145d9","toolName":"yield","startedAt":"2026-09-10T00:47:44.924Z"},"id":"6894ac9e","parentId":"a8d6d422","timestamp":"2026-09-10T00:47:44.924Z"}
{"type":"message","id":"91f95e21","parentId":"6894ac9e","timestamp":"2026-09-10T00:47:44.924Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":"**Reporting watcher finding**\n\n**Verifying watcher arbitration**"},{"type":"toolCall","id":"call_ADe4acYv3CowxlWzJJGwaFjk|fc_077f3569365e313d016aa1fe2c7d2887d187e705ac7fb145d9","name":"yield","arguments":{"type":["findings"],"data":{"title":"Use birth-safe arbitration when resolving newly eligible wakes","body":"A persistent same-boot corpse whose PID has been reused is now correctly admitted here, but the watcher’s wake consumer still uses bare `is_process_alive` in `resolve_wake` (lines 410–416). When its owner remains active and the watcher exits with opcode 86, that guard adopts the unrelated process as `already relaunched` and returns success without launching the shell. Thus the newly unfrozen persistent instance stays down despite a wake, and subsequent watchers repeat the no-op. Apply the same birth-safe launch arbitration at the wake-resolution guard so the newly eligible recycled-PID case reaches the launch.","priority":2,"confidence":0.92,"file_path":"C:/Users/decid/Documents/projects/spt-core/.worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs","line_start":672,"line_end":676},"error":null}}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":17643,"output":492,"cacheRead":82816,"cacheWrite":0,"totalTokens":100951,"reasoningTokens":274,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1789001248720,"responseId":"resp_077f3569365e313d016aa1fe20f7ec87d1888fea6e93354e85","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_077f3569365e313d016aa1fe221cb487d1b587497e774c52ea","type":"reasoning","content":[],"encrypted_content":"gAAAAABqof4sHZ643CQAxBeu9bbMsb4c2HmEV7CXROvgC-_cK81Ft9FAbGAwJdmQUn2TeNrvfqnBTMZF3KCsvUfOX1z_QJuDUmQ2B6h071OtIg6112aENlqgrAVBDeiZfafI7BSHOBzvw5HW7HuWWCZO7tkVZJ5rG5Sl775Lcp7jgqhe8voLM8CafnlUJGo8duCwONQ4Srt_1sKSTSkwLJm38I2EaHjJQM0GhrJbU6X6qSPt6mZdabyNjzHG3_YBIFKvp9uSak6nV5hTCciWR-f3irZQPo_vz_NTlkkd_APVhXThkPyz0xFL8hCRYpowCkIq1rzn--y8Sj8uHai9IlqmNon1QXX1iiHeAxbwEiQFu5zSzZ1YQ_vEiaXwZfCtTRzvIP3kp1SdFV_URAJZuPn1cM3DdLNkIobJnHwaBD43oetsbjGPjqKLizd63R7EELRS4obi96zFLQP3ebks3dhtVAGJI9WbipXx53JshABl2mre6DJiFlqSNBpJWWq6f37HptsRwfNpmLJZ5fgHBRb6O13ldhMWQe7b80q4uX16QWn7CVNIQ-KHzCo3KVMw49uCIfUrMbDQRnzthmUi-AtieeyD8QolngPaLR-H722U4Cs6bLvcotM9R-qFmbR1Sb7ManA2hSG6bnqrar5uA513ajZEQa_Ed54V8ASuVjdfQZ8YguLH6JYO1vlx9MwGeTg6rtgg_4xgb4QYpGdsc-byAleHUCwVAfMnRY7hufZ1zeBd2ETvoiUD9TNrlM0_ZsWEWlrUzxpI8GOfI3yfWqIyYJDWWVIAe2PDrKo6ycBVcoVmNzZPewab5M9MEjJW10PsdVdEWXCwJ7bUf1i7wTH-38nQcOHaKyhzX1yaqYmgyU6PRXzNJN9aAepb2mvMX-TYdMY1BCDWFOi7xyEVE4HA9kVfWLF9Eo1ewMHyiNAWXCXnW3dOHFnv1n5_gGSLR7RIeLagHptDZAn9q-17gkqa76BgVz4_JIDz3f9yoyS9lGrH2d7aVeP0rgV_G97f6YKZnV8roXQE7zcUKSPViYgIOAp-StvGKiMWUCv4a_q0F1UIl3X4E4oCbbtNXk8uYRqxU1E06HEOk0lzKE2ruGDQuUctpsgt5jWxzuXPtfjATBKHEoTzHDGyyFhHkC344nbQfg-bAHMONIWjxmgi1Bt_BVRUldsjWAQmckMMXuTM5bnrzICHBBfMILwXDAj1giYmXtitkZVpNSm1kfHUNFnDSJZ29vVAma3ji_aXIyC6_wQ8AoE7x42vwOWhvZJc_6GRDqjevUgVjcReAaqhXU6lcyt2Dm-OnHQmkDpY_K7zOMnMr58fDMNAyqMPrzxGml_ltMJgWiAHZWs46MM8rpGhLgxXG1hs7p71vC7l9kSXbABUdxW2aNxWs8m7VhUcsjbBXxq7cCty5lWCBpedi0wL9BuRBsNkA-7ESZpGmP2-y7imGenUFQlGnc0y8hHJFZBZEHQnAaDaGRgCVCV4eeJCuYYbGmaep_rDubgysDvives-OpBb2RV2l2SgxBbG3zI0_FG6lPcPMMVWAUbK0vcV7shZSlWKCNMIuXQG2sjlXfhOExLj4fNnq-VQw5AbIiRcemqAvSk0IREvuqy-_BjAg2U84ujM6HfiWVR3Ju8bEiFZjjifyX2LeAFwDytuESZKHCzCt-Qbb2z1dSCIFpu5W9BnhkAuTvWzzf9TLX2Urvzvhz3UW-LTYPslgaxXS9TJf67rftIJjz0F65HIatYnaF6HavI5RPe6WllelqHfnpw-j5S48L8s-KU4A1GZ0UIL482rd4BC27vH5xLUFcJD6KMGBqh3ac-ZIvuu5pTCID4g4xRwNpVXVSsV6ocpaQWEFySam_nk-TSKQRYVuY-KEycT5BpcMoscTtWGq9dCvd5Qu0ITUeH_KbdaCJjUCoAxJdN3DY_CVqpTHHOGvXpXV1UrEREa1e0c74-ktwoRIytvIyAeqo4PnmQxGMSj9aa7N_P9zjDtrTUOr2oY5cbfoJ8AknEsI6fyFAZKjGFcjMCdB_lUgB_N49svTgBf8ftiwuBV6LprXwSlcN7rgpPoVpAs7Xy_TiJEBTHgleN09a4MX4kdnTdJPVUEcWfRcLNkkkYavJvyp4Z5xNeFcGKvaC5nPkq5BIBxm0qFCbQD7S6tFLCSv6Rg-eewl6smx1xqyw9Dh5Piwoe4b-I4MYAUmYa-Qp8zg8Kj_bteIUgZEanUZMcA9eUdJAsEv11c2XPqpR3hIh2xjj7Nn7AbY_Ge3Uv3ppEpDzHOJcDf7-1AcSIn7MLvZShGIIOv4eEzLZlwNdz_gRf6kc-NIAGo8DROc8mw0jHj0l_lttJ3vwtX1SnN2DLVnUW590Rf33H3c2hgKq32W8mhUKODEVFfM9RRCu_y1R4pdladTcbe370erxLk2F6_xDbOmKTJjJHVvQ_godtQzo5400Ccn0un71D2LnQJLuRkalASEmJJ9o5hTRHlorDQR3HneASy7rs_uFA2HTUPgZ3H1u6LBsssoXebxLq74hv87kt0IKChoEs2GoZOHwUCM6JtNRnjslQ0VYBTprbgU6JkSp24QycI3wS0xISx574FF4zgVTjbHHixVPVVxmFcamvdXoNR-0rrxg7OIJdE82S3FDq3dQA1eQ5qj9Iz6ZERS-7JuMVzyCDu0EFEcMQONXbGl_70nrd0aBfaTl4l4_4i4mygtUxAvyHkpN5FZbpXKvyaswVAcZAcO_DDA_6EZtWX6ngpLM_UdFHGOprRATQdrjQePRuuPU9lrRFBrKNUZ-slVnDM2OIvL18j2F_P-OllscAYz2KBS5Wq9uQr_sr_MdgEZ_pIZHXVcEwXKqoSQbuWZKMWIuf7bYYieENHzJcyCXwhYg-JCHvGOd0HpuDWx2ZHviFXJYjbRNg8Z0-6SseIFsjmZ7EPjt38K66eOZbKmAYuzrJZHhjAjCCUMmqOA-e_22DwJPCeOX6dX_FVDKjs3H4L3-1fWBxXuzrFOubY6p9Qxvle9EHiI3YHJqKcVELcF9PjkStCLMU4SzVcreKvgTdqnXiah4c4gp1uzhQnbIQsMP5dPty4EBfwcREWU7FawhAtF4nsp5k0PKrPc7D-s1r4J-G3TjAssfSOAvbknEHAtgS_UlQV5sHa1JTvj6dXSPsL13_5jLqb02oHFYohxjN0Fblkj6I-daxUY_CkiRg5NNJnarbuVTGpS3umKiirUD2eHdimRLcBJTWWpBnidHbeKjxpepcmHH0ECMVsbmGBECt6KsodaPfjmhcUqkuk5xQuaZONiM3e8K8rTRIR0kC9odkWiG-tAP9UuIan_1kYlaOl-NV-xkdFj90BPVfGZUk2pGHIdPZn5KwbXD0lsvW-KOLw3e6rxgrFKO-sH1GmLboRu3MI5dhbn2f-vLQDAVsuui6k1u_Sj3k6_Q==","internal_chat_message_metadata_passthrough":{"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"},"summary":[{"type":"summary_text","text":"**Reporting watcher finding**"},{"type":"summary_text","text":"**Verifying watcher arbitration**"}],"metadata":{"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"}},{"id":"fc_077f3569365e313d016aa1fe2c7d2887d187e705ac7fb145d9","type":"function_call","status":"completed","arguments":"{\"type\":[\"findings\"],\"data\":{\"title\":\"Use birth-safe arbitration when resolving newly eligible wakes\",\"body\":\"A persistent same-boot corpse whose PID has been reused is now correctly admitted here, but the watcher’s wake consumer still uses bare `is_process_alive` in `resolve_wake` (lines 410–416). When its owner remains active and the watcher exits with opcode 86, that guard adopts the unrelated process as `already relaunched` and returns success without launching the shell. Thus the newly unfrozen persistent instance stays down despite a wake, and subsequent watchers repeat the no-op. Apply the same birth-safe launch arbitration at the wake-resolution guard so the newly eligible recycled-PID case reaches the launch.\",\"priority\":2,\"confidence\":0.92,\"file_path\":\"C:/Users/decid/Documents/projects/spt-core/.worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs\",\"line_start\":672,\"line_end\":676},\"error\":null}","call_id":"call_ADe4acYv3CowxlWzJJGwaFjk","internal_chat_message_metadata_passthrough":{"create_time":1789001249.278196,"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"},"name":"yield","metadata":{"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"}}]},"duration":16202.12080000015,"ttft":1416.1652999999933,"completedAt":1789001264923,"contextSnapshot":{"promptTokens":100459,"nonMessageTokens":13209,"compactionEpoch":0}}}
{"type":"message","id":"649955cf","parentId":"91f95e21","timestamp":"2026-09-10T00:47:44.925Z","message":{"role":"toolResult","toolCallId":"call_ADe4acYv3CowxlWzJJGwaFjk|fc_077f3569365e313d016aa1fe2c7d2887d187e705ac7fb145d9","toolName":"yield","content":[{"type":"text","text":"Result submitted."}],"details":{"data":{"title":"Use birth-safe arbitration when resolving newly eligible wakes","body":"A persistent same-boot corpse whose PID has been reused is now correctly admitted here, but the watcher’s wake consumer still uses bare `is_process_alive` in `resolve_wake` (lines 410–416). When its owner remains active and the watcher exits with opcode 86, that guard adopts the unrelated process as `already relaunched` and returns success without launching the shell. Thus the newly unfrozen persistent instance stays down despite a wake, and subsequent watchers repeat the no-op. Apply the same birth-safe launch arbitration at the wake-resolution guard so the newly eligible recycled-PID case reaches the launch.","priority":2,"confidence":0.92,"file_path":"C:/Users/decid/Documents/projects/spt-core/.worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs","line_start":672,"line_end":676},"status":"success","type":["findings"]},"isError":false,"timestamp":1789001264925}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_8cNr7ZNyAkoRgU5O2kfGjSOv|fc_077f3569365e313d016aa1fe356f1887d1a84814b439476fcb","toolName":"yield","startedAt":"2026-09-10T00:47:56.230Z"},"id":"a4965e55","parentId":"649955cf","timestamp":"2026-09-10T00:47:56.230Z"}
{"type":"message","id":"ce73efd1","parentId":"a4965e55","timestamp":"2026-09-10T00:47:56.230Z","message":{"role":"assistant","content":[{"type":"toolCall","id":"call_8cNr7ZNyAkoRgU5O2kfGjSOv|fc_077f3569365e313d016aa1fe356f1887d1a84814b439476fcb","name":"yield","arguments":{"type":["findings"],"data":{"title":"Amend the conformance checklist to remove persistent freeze gates","body":"The revised invariant contradicts the same document’s still-binding “Conformance checklist”: row 2.7 at line 1021 requires a launch predating boot and explicitly says “no stamp or no boot oracle ⇒ not restored,” while row 2.6 at line 1019 retains an unqualified no-spontaneous-relaunch rule. Neither row is marked superseded. This misses the requested documentation cutover and leaves maintainers following the AGENTS.md-mandated checklist instructed to reject the exact same-boot and unstamped persistent cases releases#287 requires restoring. Update both checklist rows alongside these invariant changes.","priority":2,"confidence":1,"file_path":"C:/Users/decid/Documents/projects/spt-core/.worktrees/287-shellrestore/docs/KNOWN-HAZARDS.md","line_start":88,"line_end":92},"error":null}}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":613,"output":214,"cacheRead":100352,"cacheWrite":0,"totalTokens":101179,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1789001264944,"responseId":"resp_077f3569365e313d016aa1fe31355487d19e1605e827dc42cf","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"fc_077f3569365e313d016aa1fe356f1887d1a84814b439476fcb","type":"function_call","status":"completed","arguments":"{\"type\":[\"findings\"],\"data\":{\"title\":\"Amend the conformance checklist to remove persistent freeze gates\",\"body\":\"The revised invariant contradicts the same document’s still-binding “Conformance checklist”: row 2.7 at line 1021 requires a launch predating boot and explicitly says “no stamp or no boot oracle ⇒ not restored,” while row 2.6 at line 1019 retains an unqualified no-spontaneous-relaunch rule. Neither row is marked superseded. This misses the requested documentation cutover and leaves maintainers following the AGENTS.md-mandated checklist instructed to reject the exact same-boot and unstamped persistent cases releases#287 requires restoring. Update both checklist rows alongside these invariant changes.\",\"priority\":2,\"confidence\":1,\"file_path\":\"C:/Users/decid/Documents/projects/spt-core/.worktrees/287-shellrestore/docs/KNOWN-HAZARDS.md\",\"line_start\":88,\"line_end\":92},\"error\":null}","call_id":"call_8cNr7ZNyAkoRgU5O2kfGjSOv","internal_chat_message_metadata_passthrough":{"create_time":1789001265.522832,"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"},"name":"yield","metadata":{"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"}}]},"duration":11285.040800000075,"ttft":4517.906699999701,"completedAt":1789001276229,"contextSnapshot":{"promptTokens":100965,"nonMessageTokens":13209,"compactionEpoch":0}}}
{"type":"message","id":"5db1835a","parentId":"ce73efd1","timestamp":"2026-09-10T00:47:56.230Z","message":{"role":"toolResult","toolCallId":"call_8cNr7ZNyAkoRgU5O2kfGjSOv|fc_077f3569365e313d016aa1fe356f1887d1a84814b439476fcb","toolName":"yield","content":[{"type":"text","text":"Result submitted."}],"details":{"data":{"title":"Amend the conformance checklist to remove persistent freeze gates","body":"The revised invariant contradicts the same document’s still-binding “Conformance checklist”: row 2.7 at line 1021 requires a launch predating boot and explicitly says “no stamp or no boot oracle ⇒ not restored,” while row 2.6 at line 1019 retains an unqualified no-spontaneous-relaunch rule. Neither row is marked superseded. This misses the requested documentation cutover and leaves maintainers following the AGENTS.md-mandated checklist instructed to reject the exact same-boot and unstamped persistent cases releases#287 requires restoring. Update both checklist rows alongside these invariant changes.","priority":2,"confidence":1,"file_path":"C:/Users/decid/Documents/projects/spt-core/.worktrees/287-shellrestore/docs/KNOWN-HAZARDS.md","line_start":88,"line_end":92},"status":"success","type":["findings"]},"isError":false,"timestamp":1789001276230}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_GQHfzRnwiuUrTfiFfhmEvU2k|fc_077f3569365e313d016aa1fe4c7f7c87d1bd1177891e44959b","toolName":"grep","startedAt":"2026-09-10T00:48:12.627Z","args":{"path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs;.worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs"},"intent":"Confirming runtime finding locations"},"id":"9974fa4f","parentId":"5db1835a","timestamp":"2026-09-10T00:48:12.627Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_suTo088bNEAYJfRqJvI2j2de|fc_077f3569365e313d016aa1fe4c7f8c87d1a63fd85f210c6ebe","toolName":"grep","startedAt":"2026-09-10T00:48:12.627Z","args":{"path":".worktrees/287-shellrestore/docs/KNOWN-HAZARDS.md"},"intent":"Confirming hazard finding location"},"id":"18d4c03e","parentId":"9974fa4f","timestamp":"2026-09-10T00:48:12.627Z"}
{"type":"message","id":"4b951a3c","parentId":"18d4c03e","timestamp":"2026-09-10T00:48:12.628Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":"**Checking line ranges**"},{"type":"toolCall","id":"call_GQHfzRnwiuUrTfiFfhmEvU2k|fc_077f3569365e313d016aa1fe4c7f7c87d1bd1177891e44959b","name":"grep","arguments":{"i":"Confirming runtime finding locations","pattern":"let live_pid|already relaunched pid|record pid:|record_shell_launch\\(&perch|persistent \\|\\||shell_pid_provably_dead\\(shell_perch\\)","path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs;.worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs","case":true,"gitignore":false,"skip":null},"intent":"Confirming runtime finding locations"},{"type":"toolCall","id":"call_suTo088bNEAYJfRqJvI2j2de|fc_077f3569365e313d016aa1fe4c7f8c87d1a63fd85f210c6ebe","name":"grep","arguments":{"i":"Confirming hazard finding location","pattern":"^- \\*\\*Invariant:\\*\\* a daemon-only|^- \\*\\*Regression shape:\\*\\*|^\\| 2\\.[67]","path":".worktrees/287-shellrestore/docs/KNOWN-HAZARDS.md","case":true,"gitignore":false,"skip":null},"intent":"Confirming hazard finding location"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":329,"output":483,"cacheRead":100864,"cacheWrite":0,"totalTokens":101676,"reasoningTokens":262,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1789001276287,"responseId":"resp_077f3569365e313d016aa1fe3cc44887d18cf906b9c54a6679","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_077f3569365e313d016aa1fe3e047c87d18fd49ab5fba443ad","type":"reasoning","content":[],"encrypted_content":"gAAAAABqof5MwbsFagDB_HyOwJ_a9_zGgq9-_UFe_cBwTOCOER78eepolOVOf-OqNrQox70a4wXsVaNGEtpvaA3vW5urlpsdMXNGb7KjGrAEQbSGfEsFcxUsVbxK3ccyauK8LOES6hSPuXccDd0ArL5FS094dcddlwjyGsC2QYiuPQb0B6QFtT1jQVaA19kBZ6yCnzCYuZVa624XbeNoPaHfCnpAuz5NJru-LlLwUdm13xfq8CcMb8Sb4w-8ZH7eeyJOw0hYW_4tWCh7p38xmMT29ByZ1rd1bTno1Yg2vkqDGiMsdlwLkDORvr_NqiEn2Gvqylw3mRGeBjzuzjFWQQdFUQbj8OT3VdeSoDb0kDKqP6wShaAm0KcRW3jNsoONRGF2sMqxjEyruUlBOWWjAinmGbhumgwleHP2y7dw9ZV5mt6NWbF-5vVSW3dsoJyzptMXoncCyPJIr-XyKCyli487Ea6dBCpVaYx0wn4IlZLXgejjBk5uIFCmzYKqjgQ0vUW-s-qbpAdylvmE9q6l-RN86XAeLY9nCDqudvQD43Oi2yN0U0DUsmCWz_SbtqEq0M8GMOG8Vv8zG__cGYHd2lsO8uaYpxcsyvlP3JrwE0jMeLPzyMMWkC9y1yNG7_IyKgerGI-gYSlwa6TKpH9QaUS2wiI1NjJ5ONPeG74WTjpUTpJN3KEZlNHBa_Gc9nxOsNT8B3WQ2r3O2wxX-OVAyGJNqn8ImthJR3ONDGipZe6GqmPKbF6OaH_09TeKOjBCPLG4vA8aQ3rekjlSvCh-vf65eHiLCc_h-WDdzub9x5KO9LSXehW2jYl8LBQmUbvFh8r9v3HvXSR-gqjXx7H5FDQ_xyN8VWmmRLCsAaTe6uSFnSb8K8BzOjVeLlpE4HrUwoZonQs5alCIEqFk_fWSqXOSIJA3Qh39IVVeBpzisxRtZ1Um0SJGj5rynBhxgr6RDROiz4nKQuTAMJuL_Pg7G8v6bDcOzlkNCMHWfMrNci-ORNnFr1QvAhBfwnNt0CwFQKImf_zeVAhb_eT5v5TJKzBZlRyLNTRi6GbSCmdrs6bUchFy2hxWJzfMeGQrSkzZYKbPRYsMNKXApdAJgcrp2R_0qMjHJYmvc_U_z9YaGJvNytOWMbC2HAgISbTLa_NLY3cBOut1SFfRDkR1I8aRoIa8CK5x1_9ZYsCZgoFMPCktI6AuIZCB5Y1qkTIUaKzrvIvX06esMgr3AsRQvif4_SkJ1oiPS4Wb10ZpWIJ3adGuJdRIuBJlWdBEvIENiDcyGDAyHa-rMVC9hSpy0Zzi870QlM_THodDLn9UlLFGshmnWPwnQgvF6UBCtu8LpQS3hrFV1GnM6ICkg3bh7K2qbxrJeweq3fwNarpCD_BqVd6eFV2EFbjxgxO2TpR6TVE9pD92Dlv9LJ0da4b5w2btwKoQ9ip8VmexbgpJgyACkBOIBZwD43ls4IqVpjzP7izcPivcyeeu0jt7rS3uW-IbJxhnsRhxc6DFYDJElCe0H3-bXIobXb03oZNxMGBodumLHFEk_vj5IzLQCB0bSZzNSDD5VuQf8FnbeTtVsS37trAqOo0B7FdYY6Mf7wmxhbQiOcl9fYMPFdR6aRDF_hp7rMoOOUYT2JohzZlm8-FEBdcGqv2CjfYy9rDAdDp0V8thPfbZ-lDIOhcgDuckTOVw2WUrAu2NHRB7rO_Edfl1T7rsxOIcNNPxADQ0qGgs6L5_LZp-iUxYMXDTbvcfBj4jc7epyp5a_Uv8DRfXB9dqOUoMvhQjgu6Lf_hhMqXhnpc3oLeCcrz8MX1fRbVh_wH2u95cGrD3gznPpHsH27gpVna-YoeFOYFW3qwkfZ9x5-YjEhCs1dqNBjjfm5mF6u9oBwJASmZiBfFYZCM8Q_3K4_no3ow9aaBkGKxYgygmNwXU_I1TTKqsrxJjukiWZQVPKwMYBGc-XsufvVFRgXgGGk-rpgg_cpGQy81J8R4qoEDeob4fcm63f6K-jGutYDPLLRGYgqJvy1XiVt6EKumk4kEK7SwGas3KpyA7xSFPCJ1pPnxk1Ey5Doa8M6DiahmqJFtpO85vENtrYBXkIbaIpBSGgG0IR5wDT7gPLMe2BduqyGj02lenEG-gS5AeoAGvlm-560J-tEFVEO6c5Ojl7dVNRoKlMqZNBG4YVPVGYexexTxvBrQXSf9fIvHBl2viy-_gBuq0wPnZ6uTgJGnoR8fDQV89HR-JsXhgpSuU_-m2XNxYvPA36DoZWzqT5spsNOAGyAp9KFj9BMt5Mv9CCu7rXqqYbRUGLgCnLX_XgQzPeD-RDIWiSgVWw-GC64htN68Q4-NGVxwA7xrS81NVotfKewjL7ocCYvwcr6YoxGku0d3LDjcr-rLAj_RhrzpK8CDLp6eUaBZ25tbE1Q6aE6hwzGZyCzA4XwYtJYrQPanm6C31RgW52XTcAOt0T_OiUKR2mqmsYfcBLk8Vns7nnkxwRHGmMgz7qTk1TAylM_fRQGvXPtVLIKJKxTe_HTwmr7CrU_DFUVq9Qt5eeE3-xRHBLWOxDEj84BZJzQnGpen6BKdUzg4DDNwl3bYVkx0l1ROUAekRgsC50MD52XDthyQoueuewn7Ecuy8o3O2Od6FoT4qwGFfwI9n5o5479qUnA9htR1VzD4wJTpK2b4gCMGESLkvYEKCR-UpBQR0BfKczyEmAH3gj0SU3RjKH_q9-f5HqLnTEri8OKATany5iw7QRBkpJPwedgmMWkNYjeZh_-utKilHkLIYsz-pO8VVbXdFRn2XydRsCU-0YNEOJ8tI_FKEby9m-sMn8Es4lwxHlIvKLkEDNDYHdZTEzmr60Lhlcx9CbKqqnMB-JGpGYY8p237WD1JlNEIeYoW2l_kbmgFQ_51ZsuH-_jvJ6aZI-ZCZAluoe0S7SCTi2w6fGsyVpoAc9x00X5Ich-rl1ZxUZLN79_0Ga7abIIQFyoLacPwxa3i9kDr6XzwKPgnrSTSQKWxt_Zp824jYNML49xdIazqHNS-LI8TQzMBoiUAfT3oh8PGz9dK9HR7ubWslJVaGSndZNn0f5hBrM-OE3nMwG4nmX3yFYj8Q_ho6JEGe6EDjlQ5K-ZGuBU_VdKm422oU5d4MHP2r0q1713cZ-b5UKyD6-pShIgEtySIdO4nmbetn34C3jrEi-Aor4dfABkzwQjIN2oG5s63wpBuv4YSbbz1dSbqJ3h8sLPlbhjVjiL5NF3g73HAz-najwVr5vqbCwZSwtj42_LA=","internal_chat_message_metadata_passthrough":{"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"},"summary":[{"type":"summary_text","text":"**Checking line ranges**"}],"metadata":{"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"}},{"id":"fc_077f3569365e313d016aa1fe4c7f7c87d1bd1177891e44959b","type":"function_call","status":"completed","arguments":"{\"i\":\"Confirming runtime finding locations\",\"pattern\":\"let live_pid|already relaunched pid|record pid:|record_shell_launch\\\\(&perch|persistent \\\\|\\\\||shell_pid_provably_dead\\\\(shell_perch\\\\)\",\"path\":\".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs;.worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs\",\"case\":true,\"gitignore\":false,\"skip\":null}","call_id":"call_GQHfzRnwiuUrTfiFfhmEvU2k","internal_chat_message_metadata_passthrough":{"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"},"name":"grep","metadata":{"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"}},{"id":"fc_077f3569365e313d016aa1fe4c7f8c87d1a63fd85f210c6ebe","type":"function_call","status":"completed","arguments":"{\"i\":\"Confirming hazard finding location\",\"pattern\":\"^- \\\\*\\\\*Invariant:\\\\*\\\\* a daemon-only|^- \\\\*\\\\*Regression shape:\\\\*\\\\*|^\\\\| 2\\\\.[67]\",\"path\":\".worktrees/287-shellrestore/docs/KNOWN-HAZARDS.md\",\"case\":true,\"gitignore\":false,\"skip\":null}","call_id":"call_suTo088bNEAYJfRqJvI2j2de","internal_chat_message_metadata_passthrough":{"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"},"name":"grep","metadata":{"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"}}]},"duration":16333.799800000153,"ttft":1771.5647999998182,"completedAt":1789001292625,"contextSnapshot":{"promptTokens":101193,"nonMessageTokens":13209,"compactionEpoch":0}}}
{"type":"message","id":"de48a3a4","parentId":"4b951a3c","timestamp":"2026-09-10T00:48:12.631Z","message":{"role":"toolResult","toolCallId":"call_suTo088bNEAYJfRqJvI2j2de|fc_077f3569365e313d016aa1fe4c7f8c87d1a63fd85f210c6ebe","toolName":"grep","content":[{"type":"text","text":" 87|- **The existing cascade test cannot fail on it — its SETUP SUPPRESSES THE FAILING ARM.** `rest_edges_cascade_shells_with_divergence` establishes its fixture by driving the suspend path, which itself writes the `offline` the code under test is supposed to encounter — so the restart shape (a record still saying `online` over a corpse, reached with **no** suspend edge) is unreachable from that fixture BY CONSTRUCTION. The test is not weak; it gives a true answer to the wrong question, so a reader askin...\n*88|- **Invariant:** a daemon-only or machine restart must not strand a persistent instance. **(a) Heal:** every reconcile cycle writes `offline` only where the derived birth-safe read contradicts recorded `online`; already-correct records are untouched. **(b) Restore:** the once-per-daemon-generation boot sweep and every owner offline→online edge run the same per-owner body. Relaunch every instance whose adapter declares `persistent`, whose owner is online, and whose binary is down in fact. There is **no ...\n*89|- **Regression shape:** construct recorded `online` over a corpse **without a suspend edge**, including a launch **after machine boot**, then bring its owner online after a spent boot sweep. Every stranded persistent sibling must be restored; offline owners and nonpersistent instances are not restored. Separately defend live/binding exclusion and the nonpersistent watcher freeze. Suspending in setup suppresses the failing arm and proves nothing about this class.\n 90|- **Per-owner enumeration:** the existing `SHELL_RECORD_HEALED:<owner>/<shell_id>` daemon diagnostic names **every** successfully healed instance, not one representative. Healthy and already-healed records produce no event. The restore diagnostics likewise identify each attempted instance and distinguish boot from owner-online triggers. <!-- [doc->REQ-SHELL-HEAL-ENUMERATION] -->\n 91|- **spt-core mapping:** `spt_daemon::shellwake::{heal_stale_online_records, restore_persistent_shells_at_boot, restore_persistent_shells_on_owner_online, OwnerOnlineEdge}`; `shellhost::live_launch_winner`; `spt_store::shellinfo::{ShellLaunch, shell_pid_provably_dead}`. `launch_predates_boot` and `BOOT_RESTORE_SLACK_MS` now govern only nonpersistent watcher eligibility.\n 92|- **Source:** BAROMETER W2 (releases#78); releases#287 operator ruling (2026-09-08), superseding the persistent force-kill freeze and boot-instant conjunct.\n...\n 1018|| 2.4 | gen_start = now() on cold-start + handoff | per-instance generation |\n*1019|| 2.6 | A shell's ONLINE-ness is DERIVED (recorded status AND a not-provably-dead `shell.pid`) — an abruptly-killed binary breaks no link, so `close_shell`'s offline flip never runs and the record lies forever; derive at the gates/renders, keep the recorded field at the writers, the suspend-cascade close, and the wake reconciler — whose \"no spontaneous relaunch\" (deploys and quarantine own that decision) is since W2 the stated corpse-predates-boot predicate, not the stale record holding it open, and ...\n 1020|<!-- [doc->REQ-HAZARD-SHELL-STALE-ONLINE] -->\n*1021|| 2.7 | A node restart must not permanently strand every `persistent` shell: (a) the daemon heals a recorded `online` the derivation contradicts, write-guarded on an actual change, and (b) a once-per-generation boot sweep PLUS an owner offline->online edge in the reconcile loop (releases#228 — the sweep alone strands the shells of any owner not yet online when it ran) relaunch an instance only when ALL of — adapter `persistent`, owner endpoint online, down in fact, recorded launch predates the boot i...\n 1022|| 3.1 | Ephemeral perch cleanup on all exit paths | `ring` (RAII guard) |\n 1023|| 3.4 | A ring never adopts (so never deletes) a perch it did not create — probe the DIR, not the ready marker; record/spool/unreadable = occupied, empty = refused too; deliver + loud `RING_PERCH_EXISTS`/`RING_STALE_DIR` instead of block-waiting | `spt_msg::ring` probe + `create_dir` leaf |\n 1024|| 3.5 | A perch GC classifies on record PRESENCE only and never asks a liveness resolver — `is_perch_alive` is INVERTED here (recordless residue reads ALIVE, an offline endpoint reads DEAD: 24 vs 6 measured on HFENDULEAM 2026-08-04), so `!is_perch_alive` authorizes exactly the inverse set; the registry is never asked either (offline endpoints are legitimately absent from it) | `spt_store::perchgc::sweep`, `spt endpoint gc` |\n\n[Some lines truncated to 512 chars]"}],"details":{"scopePath":".worktrees/287-shellrestore/docs/KNOWN-HAZARDS.md","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\docs\\KNOWN-HAZARDS.md","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":4,"fileCount":1,"files":[".worktrees/287-shellrestore/docs/KNOWN-HAZARDS.md"],"fileMatches":[{"path":".worktrees/287-shellrestore/docs/KNOWN-HAZARDS.md","count":4}],"truncated":true,"displayContent":"   87│- **The existing cascade test cannot fail on it — its SETUP SUPPRESSES THE FAILING ARM.** `rest_edges_cascade_shells_with_divergence` establishes its fixture by driving the suspend path, which itself writes the `offline` the code under test is supposed to encounter — so the restart shape (a record still saying `online` over a corpse, reached with **no** suspend edge) is unreachable from that fixture BY CONSTRUCTION. The test is not weak; it gives a true answer to the wrong question, so a reader askin...\n  *88│- **Invariant:** a daemon-only or machine restart must not strand a persistent instance. **(a) Heal:** every reconcile cycle writes `offline` only where the derived birth-safe read contradicts recorded `online`; already-correct records are untouched. **(b) Restore:** the once-per-daemon-generation boot sweep and every owner offline→online edge run the same per-owner body. Relaunch every instance whose adapter declares `persistent`, whose owner is online, and whose binary is down in fact. There is **no ...\n  *89│- **Regression shape:** construct recorded `online` over a corpse **without a suspend edge**, including a launch **after machine boot**, then bring its owner online after a spent boot sweep. Every stranded persistent sibling must be restored; offline owners and nonpersistent instances are not restored. Separately defend live/binding exclusion and the nonpersistent watcher freeze. Suspending in setup suppresses the failing arm and proves nothing about this class.\n   90│- **Per-owner enumeration:** the existing `SHELL_RECORD_HEALED:<owner>/<shell_id>` daemon diagnostic names **every** successfully healed instance, not one representative. Healthy and already-healed records produce no event. The restore diagnostics likewise identify each attempted instance and distinguish boot from owner-online triggers. <!-- [doc->REQ-SHELL-HEAL-ENUMERATION] -->\n   91│- **spt-core mapping:** `spt_daemon::shellwake::{heal_stale_online_records, restore_persistent_shells_at_boot, restore_persistent_shells_on_owner_online, OwnerOnlineEdge}`; `shellhost::live_launch_winner`; `spt_store::shellinfo::{ShellLaunch, shell_pid_provably_dead}`. `launch_predates_boot` and `BOOT_RESTORE_SLACK_MS` now govern only nonpersistent watcher eligibility.\n   92│- **Source:** BAROMETER W2 (releases#78); releases#287 operator ruling (2026-09-08), superseding the persistent force-kill freeze and boot-instant conjunct.\n     │...\n 1018│| 2.4 | gen_start = now() on cold-start + handoff | per-instance generation |\n*1019│| 2.6 | A shell's ONLINE-ness is DERIVED (recorded status AND a not-provably-dead `shell.pid`) — an abruptly-killed binary breaks no link, so `close_shell`'s offline flip never runs and the record lies forever; derive at the gates/renders, keep the recorded field at the writers, the suspend-cascade close, and the wake reconciler — whose \"no spontaneous relaunch\" (deploys and quarantine own that decision) is since W2 the stated corpse-predates-boot predicate, not the stale record holding it open, and ...\n 1020│<!-- [doc->REQ-HAZARD-SHELL-STALE-ONLINE] -->\n*1021│| 2.7 | A node restart must not permanently strand every `persistent` shell: (a) the daemon heals a recorded `online` the derivation contradicts, write-guarded on an actual change, and (b) a once-per-generation boot sweep PLUS an owner offline->online edge in the reconcile loop (releases#228 — the sweep alone strands the shells of any owner not yet online when it ran) relaunch an instance only when ALL of — adapter `persistent`, owner endpoint online, down in fact, recorded launch predates the boot i...\n 1022│| 3.1 | Ephemeral perch cleanup on all exit paths | `ring` (RAII guard) |\n 1023│| 3.4 | A ring never adopts (so never deletes) a perch it did not create — probe the DIR, not the ready marker; record/spool/unreadable = occupied, empty = refused too; deliver + loud `RING_PERCH_EXISTS`/`RING_STALE_DIR` instead of block-waiting | `spt_msg::ring` probe + `create_dir` leaf |\n 1024│| 3.5 | A perch GC classifies on record PRESENCE only and never asks a liveness resolver — `is_perch_alive` is INVERTED here (recordless residue reads ALIVE, an offline endpoint reads DEAD: 24 vs 6 measured on HFENDULEAM 2026-08-04), so `!is_perch_alive` authorizes exactly the inverse set; the registry is never asked either (offline endpoints are legitimately absent from it) | `spt_store::perchgc::sweep`, `spt endpoint gc` |","linesTruncated":true,"meta":{"limits":{"columnTruncated":{"maxColumn":512}}}},"isError":false,"timestamp":1789001292630}}
{"type":"message","id":"84ea881e","parentId":"de48a3a4","timestamp":"2026-09-10T00:48:12.632Z","message":{"role":"toolResult","toolCallId":"call_GQHfzRnwiuUrTfiFfhmEvU2k|fc_077f3569365e313d016aa1fe4c7f7c87d1bd1177891e44959b","toolName":"grep","content":[{"type":"text","text":"# .worktrees/287-shellrestore/crates/spt-daemon/src/\n## shellwake.rs\n 410|    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n*411|    let live_pid = std::fs::read_to_string(perch.join(crate::shellhost::SHELL_PID_FILE))\n 412|        .ok()\n 413|        .and_then(|s| s.trim().parse::<u32>().ok())\n 414|        .filter(|&p| p != 0 && spt_store::proc::is_process_alive(p));\n 415|    if let Some(p) = live_pid {\n*416|        return Ok(format!(\"{did}already relaunched pid={p} (online at bind)\"));\n 417|    }\n 418|    // [impl->REQ-INSTALL-11] the wake-triggered relaunch resolves against the\n 419|    // adapter's install dir, like every other launch of the same instance.\n...\n 674|    }\n*675|    if persistent || !shellinfo::shell_pid_provably_dead(perch) {\n 676|        return true; // persistence overrides the freeze; clean closes need no age\n 677|    }\n 678|    let (Some(boot_ms), Some(launch)) = (boot_ms, shellinfo::read_shell_launch(perch)) else {\n...\n 1217|        if let Some(launched_ms) = launched_ms {\n*1218|            shellinfo::record_shell_launch(&perch, 2_000_000_000, launched_ms);\n 1219|        }\n 1220|        id\n 1221|    }\n## shellhost.rs\n 321|    shellinfo::read_shell_pid(shell_perch)\n*322|        .filter(|&p| p != 0 && !shellinfo::shell_pid_provably_dead(shell_perch))\n 323|}\n 324|\n 325|/// THE route-selection seam. Generic over the daemon call and the in-process\n...\n 499|            std::fs::write(perch.join(SHELL_PID_FILE), pid.to_string())\n*500|                .map_err(|e| format!(\"record pid: {e}\"))?;\n 501|            // The birth stamp rides WITH the pid at every launch, or the probe that\n 502|            // reads it degrades to bare-pid liveness on this instance alone -- a\n 503|            // silent per-instance regression rather than a visible one.\n 504|            // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n*505|            spt_store::shellinfo::record_shell_launch(&perch, pid, now_ms());\n 506|            Ok(pid)\n 507|        }\n 508|    }\n...\n 558|    std::fs::write(perch.join(SHELL_PID_FILE), pid.to_string())\n*559|        .map_err(|e| format!(\"record pid: {e}\"))?;\n 560|    // Same stamp at the broker-hosted launch. A backend that exposed no pid\n 561|    // records 0 here; record_shell_launch stamps no start-time for it, so the\n 562|    // probe keeps its fail-toward-alive stance for pid-less backends.\n 563|    // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n*564|    spt_store::shellinfo::record_shell_launch(&perch, pid, now_ms());\n 565|    Ok(pid)\n 566|}\n 567|"}],"details":{"scopePath":".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs, .worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\crates\\spt-daemon\\src","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":9,"fileCount":2,"files":[".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs",".worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs"],"fileMatches":[{"path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs","count":4},{"path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs","count":5}],"truncated":false,"displayContent":"# .worktrees/287-shellrestore/crates/spt-daemon/src/\n## shellwake.rs\n  410│    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n *411│    let live_pid = std::fs::read_to_string(perch.join(crate::shellhost::SHELL_PID_FILE))\n  412│        .ok()\n  413│        .and_then(|s| s.trim().parse::<u32>().ok())\n  414│        .filter(|&p| p != 0 && spt_store::proc::is_process_alive(p));\n  415│    if let Some(p) = live_pid {\n *416│        return Ok(format!(\"{did}already relaunched pid={p} (online at bind)\"));\n  417│    }\n  418│    // [impl->REQ-INSTALL-11] the wake-triggered relaunch resolves against the\n  419│    // adapter's install dir, like every other launch of the same instance.\n     │...\n  674│    }\n *675│    if persistent || !shellinfo::shell_pid_provably_dead(perch) {\n  676│        return true; // persistence overrides the freeze; clean closes need no age\n  677│    }\n  678│    let (Some(boot_ms), Some(launch)) = (boot_ms, shellinfo::read_shell_launch(perch)) else {\n     │...\n 1217│        if let Some(launched_ms) = launched_ms {\n*1218│            shellinfo::record_shell_launch(&perch, 2_000_000_000, launched_ms);\n 1219│        }\n 1220│        id\n 1221│    }\n## shellhost.rs\n 321│    shellinfo::read_shell_pid(shell_perch)\n*322│        .filter(|&p| p != 0 && !shellinfo::shell_pid_provably_dead(shell_perch))\n 323│}\n 324│\n 325│/// THE route-selection seam. Generic over the daemon call and the in-process\n    │...\n 499│            std::fs::write(perch.join(SHELL_PID_FILE), pid.to_string())\n*500│                .map_err(|e| format!(\"record pid: {e}\"))?;\n 501│            // The birth stamp rides WITH the pid at every launch, or the probe that\n 502│            // reads it degrades to bare-pid liveness on this instance alone -- a\n 503│            // silent per-instance regression rather than a visible one.\n 504│            // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n*505│            spt_store::shellinfo::record_shell_launch(&perch, pid, now_ms());\n 506│            Ok(pid)\n 507│        }\n 508│    }\n    │...\n 558│    std::fs::write(perch.join(SHELL_PID_FILE), pid.to_string())\n*559│        .map_err(|e| format!(\"record pid: {e}\"))?;\n 560│    // Same stamp at the broker-hosted launch. A backend that exposed no pid\n 561│    // records 0 here; record_shell_launch stamps no start-time for it, so the\n 562│    // probe keeps its fail-toward-alive stance for pid-less backends.\n 563│    // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n*564│    spt_store::shellinfo::record_shell_launch(&perch, pid, now_ms());\n 565│    Ok(pid)\n 566│}\n 567│"},"isError":false,"timestamp":1789001292632}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_3MVzpHNtNtPBiSORXNH3l0iI|fc_077f3569365e313d016aa1fe53104487d1a9bf9771ca452abd","toolName":"yield","startedAt":"2026-09-10T00:48:19.262Z"},"id":"8eba2f43","parentId":"84ea881e","timestamp":"2026-09-10T00:48:19.262Z"}
{"type":"message","id":"e7eee0d9","parentId":"8eba2f43","timestamp":"2026-09-10T00:48:19.263Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":""},{"type":"toolCall","id":"call_3MVzpHNtNtPBiSORXNH3l0iI|fc_077f3569365e313d016aa1fe53104487d1a9bf9771ca452abd","name":"yield","arguments":{"type":["overall_correctness"],"data":"incorrect","error":null}}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":2534,"output":151,"cacheRead":100992,"cacheWrite":0,"totalTokens":103677,"reasoningTokens":120,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1789001292690,"responseId":"resp_077f3569365e313d016aa1fe4d0df487d1aee493e28b4b3a48","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_077f3569365e313d016aa1fe4e922887d1809da4a1351d7935","type":"reasoning","content":[],"encrypted_content":"gAAAAABqof5TU5dMGVjt4yRelQt-j6jLzyKo7mXVn0T-58SXCFhDywmgHRzbwBNjgv_YAC-e5dbPUzUexnv1BouPmOgkLOwaPlzkM3JVnBVzgWX_UnVPpQlzXV1e8dmi4u4cxEmTey8G9zAcY4mLm31ayWJ2voQwIWfBmEv3xnHqZ7oYnivDleNKAPZrAesbKYqh0szZFbGaPn4WmYukeVEOkyag1q3d3Fde9iarv4By9ofPbZXt1iJNUoOciHSruwu6V0Wnjyj4cioyX8XYAkz_-DnB7WMBWoEndBgt9Isn4ps3HlTbZGSp3VYWXGD1xfCh-SR64hGDUELwh_VDLpHhltcIcIIgW274xGtNkzE-9NYwdnhJQSyGirJj4RrCIkJx0c1t4TKUgu_MYEGsLVR5HAqWLsXAADEBNwQ_klHSSHrjWgymI323MPGhH9oRJeXH3ereALA1mhONWOYBgGpNxB1xi_IJ-unwDLdEwKm0qFAZEQpHj9oe16WicFAA3gEyJJDK7c0_e12nc9k0k_50lKpoKEi9GHRmsI8bujmuTHK3tpneWfURVAbPMAgvOinoXR1ugYnksA8hXwLDYuKj7FcEJ7muctw_K3H7wZYMDmKnH41lzCO95MC2PU3GJZR4fDmAGFqyY0bqf1Yz9xC_WgFr1yui3z1RWJwFkyKPyRZD-z_5nzcbQbLuzekGGjoOABZ7pzzkIux6I_HgKYPe7pcXuT9n3brCFFmZ6yjMo0wndbEmZVJ_caGHeQCG7CCvvCz79GEc1FBZHEpV3PtAcD6_hDU49itZ6o5pbzCnFSRdL6pIOwDIyaEdNMnpgxdCuKFn8omnHIdDg7qeTQGHq_O8L6uQ4-q_6hDD7lq4syg7Oc7NfGWbhFmYj3fGwVjbzo_paMbNjPT7mV7-38lfH_FDqvIruuMnGl9JCzuYW6VuuPFGwezX9leUFM1mmgkEftFQMe8B0vlfLLbRec1GvI-YSTiRP6FhpnDE1KlnUtKW5MATOJBktXiKOaIP-pdo2NQ5GGsHgeDuaZVD_uLFMHzl4PU5_zebYwnsVwyo6uQlgkAPMym_uf0gcKSpbfJ_6sxmCDdtE4dD_SlIVmc9rXiTbuUTFPrO6Oir6TjYNHWcLl2phXyNKKndlmVbDhNvRmITsV_E2JfAWdTRAq9an_DULTdpt86Dg3ZuAL7YBVEEFsvNdQ2aJqf29ChQmX0Oo2Rw6Za3dT7hKDj0DcE24SwwzdpuPd7DOoepVjUyd4Ow8Ep9dfkwCQD2xc5j82VDq53BLoW-NqnpYoQ2eS502LoP6GJw4HF3Q-WnBEBbTMFLFuETMZNJYWcBZ2ZoCQMW0L8Cw67EpNM5xfawbQJ0neCS7IFrK0R6cNmnkrR3yhnFKxyRwZy0Lt2uWF2OftzrwqgOaNCg6sAz_QzbDOKEoyMR2ZufYmCMEbFGGkWYs2pOVzt_xemltqygKYF-G2CJ3Kmv3wUBpLAP27mHIsRZ5wEpcziae0LHK9sBEmv4xnKUc4xlcQ396YqBSu5rNLwVv5yGkgfmzphhuujZ-5q3USiMRRH8l216JP6h1lyUQ8TVpnpML58dTjuT4Sx49HcT5Nt2gWGehPLT9BYXdKXoievE1j7B3B3ysAcP3N_O0bH1TTRhaesirmknob-fbCdMzBowB4_fZxnOCaOujvgv8_8Y24qj8YMwG0oDo-3Vbx-tRSiyrnjm2lkIjraGQSvPYcy1IA_dryAc4eIjXz87Fea6qxQewL9ioe6EFvrJ9pH9r3kN13x3hXa9C2JE9UubPjpJhuTqygTJo4D3SypUUoL1vxzMmqWiWYznspfy2q-oAN9wHHHL0Oh7us1Vy05M89Z44tGpyB2sbhKjRNNLDSHO4SnYngicL2ff9IKNWknyLRwdZaoVHbA-X_YMG4M3isFqt1wHGTXS8WRrQYk4t2yhutVtHw9bqK99UP1_t81iRBXHZPvWa-KFubPiLVk2gvmuACUt8wHQ3gyT4gEIBEHxj-riyx4XPFRN8OaNpo-2Kqt7XqVgyA_mJ00CxTSal1Qg9O4pH9rReXEUtOB9UXy2qMS1XFLUPllmBCMl2lwBuV0e4l8UvF6jiEOabuNhPigwOX-aIeX4Ox4PCW8E25APupdDCBaKsHqPyK0lMk0Jj4fGTpALD1diMkdNnJFYBOAfcdS73Lp5ubNddh6PMytJX6wTsQ==","internal_chat_message_metadata_passthrough":{"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"},"summary":[],"metadata":{"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"}},{"id":"fc_077f3569365e313d016aa1fe53104487d1a9bf9771ca452abd","type":"function_call","status":"completed","arguments":"{\"type\":[\"overall_correctness\"],\"data\":\"incorrect\",\"error\":null}","call_id":"call_3MVzpHNtNtPBiSORXNH3l0iI","internal_chat_message_metadata_passthrough":{"create_time":1789001293.406347,"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"},"name":"yield","metadata":{"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"}}]},"duration":6565.68209999986,"ttft":1912.0625,"completedAt":1789001299261,"contextSnapshot":{"promptTokens":103526,"nonMessageTokens":13209,"compactionEpoch":0}}}
{"type":"message","id":"ba6ceafe","parentId":"e7eee0d9","timestamp":"2026-09-10T00:48:19.263Z","message":{"role":"toolResult","toolCallId":"call_3MVzpHNtNtPBiSORXNH3l0iI|fc_077f3569365e313d016aa1fe53104487d1a9bf9771ca452abd","toolName":"yield","content":[{"type":"text","text":"Result submitted."}],"details":{"data":"incorrect","status":"success","type":["overall_correctness"]},"isError":false,"timestamp":1789001299263}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_W7WoGdJHkUFrmKiGe3xMUEup|fc_077f3569365e313d016aa1fe55493087d1bca2d4b8f7731d14","toolName":"yield","startedAt":"2026-09-10T00:48:24.790Z"},"id":"c909a1c5","parentId":"ba6ceafe","timestamp":"2026-09-10T00:48:24.790Z"}
{"type":"message","id":"7f78dbfd","parentId":"c909a1c5","timestamp":"2026-09-10T00:48:24.790Z","message":{"role":"assistant","content":[{"type":"toolCall","id":"call_W7WoGdJHkUFrmKiGe3xMUEup|fc_077f3569365e313d016aa1fe55493087d1bca2d4b8f7731d14","name":"yield","arguments":{"type":["explanation"],"data":"Correctness/spec: two source-level defects remain in launch arbitration and the newly eligible persistent watcher path; the shared restore gates, per-instance heal enumeration, and binding relink diagnostic otherwise match the ruling by inspection. Standards/spec: the binding KNOWN-HAZARDS checklist still mandates the superseded persistent freeze. This is source review of cb756dee against its parent, not runtime acceptance; no tests, builds, linters, formatters, edits, or production operations were run.","error":null}}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":267,"output":123,"cacheRead":103424,"cacheWrite":0,"totalTokens":103814,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1789001299287,"responseId":"resp_077f3569365e313d016aa1fe53b3d087d1af8cb7350ae8a623","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"fc_077f3569365e313d016aa1fe55493087d1bca2d4b8f7731d14","type":"function_call","status":"completed","arguments":"{\"type\":[\"explanation\"],\"data\":\"Correctness/spec: two source-level defects remain in launch arbitration and the newly eligible persistent watcher path; the shared restore gates, per-instance heal enumeration, and binding relink diagnostic otherwise match the ruling by inspection. Standards/spec: the binding KNOWN-HAZARDS checklist still mandates the superseded persistent freeze. This is source review of cb756dee against its parent, not runtime acceptance; no tests, builds, linters, formatters, edits, or production operations were run.\",\"error\":null}","call_id":"call_W7WoGdJHkUFrmKiGe3xMUEup","internal_chat_message_metadata_passthrough":{"create_time":1789001300.059772,"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"},"name":"yield","metadata":{"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"}}]},"duration":5502.064500000328,"ttft":2028.5135000003502,"completedAt":1789001304789,"contextSnapshot":{"promptTokens":103691,"nonMessageTokens":13209,"compactionEpoch":0}}}
{"type":"message","id":"cff6b80c","parentId":"7f78dbfd","timestamp":"2026-09-10T00:48:24.791Z","message":{"role":"toolResult","toolCallId":"call_W7WoGdJHkUFrmKiGe3xMUEup|fc_077f3569365e313d016aa1fe55493087d1bca2d4b8f7731d14","toolName":"yield","content":[{"type":"text","text":"Result submitted."}],"details":{"data":"Correctness/spec: two source-level defects remain in launch arbitration and the newly eligible persistent watcher path; the shared restore gates, per-instance heal enumeration, and binding relink diagnostic otherwise match the ruling by inspection. Standards/spec: the binding KNOWN-HAZARDS checklist still mandates the superseded persistent freeze. This is source review of cb756dee against its parent, not runtime acceptance; no tests, builds, linters, formatters, edits, or production operations were run.","status":"success","type":["explanation"]},"isError":false,"timestamp":1789001304790}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_4ricr1uNue6ld3Fe08KjJthi|fc_077f3569365e313d016aa1fe5bb19c87d1b63e81866b6610d3","toolName":"yield","startedAt":"2026-09-10T00:48:28.415Z"},"id":"834b78a7","parentId":"cff6b80c","timestamp":"2026-09-10T00:48:28.415Z"}
{"type":"message","id":"ad14eb9d","parentId":"834b78a7","timestamp":"2026-09-10T00:48:28.415Z","message":{"role":"assistant","content":[{"type":"toolCall","id":"call_4ricr1uNue6ld3Fe08KjJthi|fc_077f3569365e313d016aa1fe5bb19c87d1b63e81866b6610d3","name":"yield","arguments":{"type":["confidence"],"data":0.95,"error":null}}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":276,"output":29,"cacheRead":103552,"cacheWrite":0,"totalTokens":103857,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1789001304806,"responseId":"resp_077f3569365e313d016aa1fe5916d087d1a78e19c9b373bdb6","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"fc_077f3569365e313d016aa1fe5bb19c87d1b63e81866b6610d3","type":"function_call","status":"completed","arguments":"{\"type\":[\"confidence\"],\"data\":0.95,\"error\":null}","call_id":"call_4ricr1uNue6ld3Fe08KjJthi","internal_chat_message_metadata_passthrough":{"create_time":1789001305.46968,"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"},"name":"yield","metadata":{"turn_id":"b8a66920-567a-4069-ae25-2e82a8d2a2f9"}}]},"duration":3607.8250000001863,"ttft":2916.8056000000797,"completedAt":1789001308415,"contextSnapshot":{"promptTokens":103828,"nonMessageTokens":13209,"compactionEpoch":0}}}
{"type":"message","id":"7bbe9a2a","parentId":"ad14eb9d","timestamp":"2026-09-10T00:48:28.416Z","message":{"role":"toolResult","toolCallId":"call_4ricr1uNue6ld3Fe08KjJthi|fc_077f3569365e313d016aa1fe5bb19c87d1b63e81866b6610d3","toolName":"yield","content":[{"type":"text","text":"Result submitted."}],"details":{"data":0.95,"status":"success","type":["confidence"]},"isError":false,"timestamp":1789001308416}}
{"type":"branch_summary","id":"42d47314","parentId":"7bbe9a2a","timestamp":"2026-09-10T00:48:30.113Z","fromId":"7bbe9a2a","summary":"","details":{"kind":"discarded-entry-branch","discardedEntryId":"99b95c75"}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_ViAaoB9g5jZqXEAR2eODTfYX|fc_028129a7f1d66e9c016aa1fe64b81887d2b2ca2cf75317c3a4","toolName":"yield","startedAt":"2026-09-10T00:48:36.904Z"},"id":"c02d8fb6","parentId":"42d47314","timestamp":"2026-09-10T00:48:36.904Z"}
{"type":"message","id":"2613c024","parentId":"c02d8fb6","timestamp":"2026-09-10T00:48:36.904Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":""},{"type":"toolCall","id":"call_ViAaoB9g5jZqXEAR2eODTfYX|fc_028129a7f1d66e9c016aa1fe64b81887d2b2ca2cf75317c3a4","name":"yield","arguments":{"type":"result","data":null,"error":null}}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":233,"output":116,"cacheRead":103680,"cacheWrite":0,"totalTokens":104029,"reasoningTokens":88,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1789001310145,"responseId":"resp_028129a7f1d66e9c016aa1fe5e78e087d2b8129554fceec860","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_028129a7f1d66e9c016aa1fe609b3887d2b8007b37a80c932b","type":"reasoning","content":[],"encrypted_content":"gAAAAABqof5kXSretSCiWl0JB1RNpbp7g_eTLfrpRdp8vb-I4FrT-8EZMZhlQR6kN04JSrww89Dkxd-4TFyo8NzYn5VUN66keYyINspsEg3VXxuNUnmyJsAljosGELd2ue7DUupwIiAzrN-iEUG90cYojvCqYvp8jl7YzlBjX8cXupO-yxVOhKkAOJHpmgebOfbeshrtDPW6jWRGFtcMDrhWR02N8r988jjK4Qg3SnIv16Ke4vmTcvO7SwKw3NEf7GwplSTTG9ZtSd011pwVew1pmlRFvpgPiPQC8ZJ4-XNL2EO_6zGa0QKwC2ktj4MLJt5iYNcx0qBjOXWx-Fs9DPlsacB4_dTboNvPqG2yod42tpt9gw-ybsShJXe9BpZpdAVXEjpX_t2TZsZVtNRayCE_iackVGwyPbdsHVF6TfJlGWHI6k5_UK_ezEAoE8EE9GCfu2UR3IwL98wTunCcFytuQquoeOtJaJMsSQrDXGIuyq5QdOvTdNkd-eROVz5iEFh2Meqm8j1ehl3dNI3gDOffDRH7JAf4UPDTTKsljdy3Hi4u2aKo96_cJxXEoKP2KWBYyziINu0MELy70ba2EZiuQ3GiOwZjeEdWM5t1N0CzTk6miYEcJDH3KdeK5f3Sek_rVOWgJ02F8Pe0gOhzg0GCfPUblUffVpgD-j7C-Nk8AqDytHGHbS2OD-WJEZ1IUKnpdYDGb-p67KEGxXyMab0n-P50HMNaRe9WNqbA34Av5CpRQhl_Ys7GtSBjpqwbeXL5RTZDDNU27v_tWXl5CAY7gsE0EII-g5ogaymV7V6OHN483wNrcvahvLEQgqE9wYzOLe_GyfdJxcUUgextbaoMrPQ8fM6k2vQTN7KrA6PnLXRnZcyJRdvq_57vqRTfU4RJF5cF85gLDE9Vr6FXxuN8G-dcKGd_vZ7FVcvxbVbZ25I7IQr-2WyGYWPq4hZQY8-dFNKuDGZKpnIgC5CWMp5bwo3zcz0JTLxOIOWy_Q4DqGz_SJem-B2oRAmyaSY8LdPFOkulmzZhaskf9HtkC4zObQTiDKsWvVFDd7SQXImqCS5_3hfi-TMIk34hwDC8NMtZ_uTKY-FH_5KSzPfiAbZl17xOANmrHcaepln5flZ1PXUrptpMqbA-w4wJ8bYed-CCD46CpstmUV-pKiuKefYtLFvDjBLfpRb1yPoBpQr-WF0SClxNH0OT7O9GJNfjn_KgaV7bGGEyOAJTjDBDupPIxvCAdzfjrXJZT0wQ0MXzW7IKUbEMTPjW2seaUszAepnBH5SZqIgdWlKpKpdZJC631vAcKa_0SQXy_92xgp7uM0OHKrIyjZYPM5LyOSRrKpjj4VO6lF3XPWB6wV2FwECz79RZ793U3utl7PV_h35GR_J0zRzpKhbYTIufTBK16lQr_N4qUMoWVon0eF1m2T2Byj4TPcHDnLkvBSMPGmMBjwgMiN5etDJ4zg6CQaqz9YCxOnrKJ-hhHtQCow8gnEKFCKkkZaVOHhlGggPx9hhslY74NQJ89vzAV3RY1GGI0jbbB7A-rytZQVxScmSlROH-3rXSGha9KjoCiPuovaz2C5tos7j76GWOw2xSPVMJV-hvEv49qvp7kHJkaAjVpWZTGKI2K9rr5Zjw6oTH0MvlGycgsGOwV_ybrGWuCLBfc_ePGDj8krseE1Zeub8WYpufKuVE64NTZt3zEjW0o2yRRF19FYqPiZqIIs5zJjLBHrpS8II3aSra4ExnkPx1QbNEswpnettBLi1Kpxpck_VOjYCAWrTg6uv1dTpkvslWRdvgvHgeUpCPVMe4HkW8y4Uh_NIMrBVu6IQmgIcGSno9OM_D3MI4id0ecjn11lznsticyDXYCiy3CsXKTORoj4igC4VajWZY05BaPrcAIMLdKneunzmu1XJG90gC86_9wSmrNdWkWSRyQ5h_H1R3N_FCPFFsIqgpMQ==","internal_chat_message_metadata_passthrough":{"turn_id":"236713e7-cc4a-4d05-8af9-dd3c35ed9ac7"},"summary":[],"metadata":{"turn_id":"236713e7-cc4a-4d05-8af9-dd3c35ed9ac7"}},{"id":"fc_028129a7f1d66e9c016aa1fe64b81887d2b2ca2cf75317c3a4","type":"function_call","status":"completed","arguments":"{\"type\":\"result\",\"data\":null,\"error\":null}","call_id":"call_ViAaoB9g5jZqXEAR2eODTfYX","internal_chat_message_metadata_passthrough":{"create_time":1789001311.020233,"turn_id":"236713e7-cc4a-4d05-8af9-dd3c35ed9ac7"},"name":"yield","metadata":{"turn_id":"236713e7-cc4a-4d05-8af9-dd3c35ed9ac7"}}]},"duration":6758.189199999906,"ttft":2489.9675000002608,"completedAt":1789001316903,"contextSnapshot":{"promptTokens":103913,"nonMessageTokens":13209,"compactionEpoch":0}}}
{"type":"message","id":"79b48a0d","parentId":"2613c024","timestamp":"2026-09-10T00:48:36.906Z","message":{"role":"toolResult","toolCallId":"call_ViAaoB9g5jZqXEAR2eODTfYX|fc_028129a7f1d66e9c016aa1fe64b81887d2b2ca2cf75317c3a4","toolName":"yield","content":[{"type":"text","text":"Result submitted."}],"details":{"status":"success","type":"result","useLastTurn":true},"isError":false,"timestamp":1789001316905}}
{"type":"custom_message","customType":"irc:incoming","content":"<irc>\nIncoming IRC message from agent `Main`:\n\nAddressed your findings in working tree, no build: new shellinfo::publish_shell_launch retires old stamp (failure aborts), atomically publishes PID, writes new birth best-effort; both production launch sites use it. shell_pid_provably_dead rechecks PID after reading birth, unstable snapshot fails alive. resolve_wake uses live_launch_winner; checklist 2.6/2.7 amended. Review only this correction for remaining race/error paths; no edits/build/tests. No schema change, intentionally preserves existing spawn-before-PID window rather than claiming to solve whole launch serialization.\nIf response expected, reply via `hub` (`op: \"send\"`, `to: \"Main\"`) when available; otherwise what you `yield` or say last this turn is delivered to `Main` when you stop.\n</irc>","display":true,"details":{"id":"15798116f67ffef4","from":"Main","message":"Addressed your findings in working tree, no build: new shellinfo::publish_shell_launch retires old stamp (failure aborts), atomically publishes PID, writes new birth best-effort; both production launch sites use it. shell_pid_provably_dead rechecks PID after reading birth, unstable snapshot fails alive. resolve_wake uses live_launch_winner; checklist 2.6/2.7 amended. Review only this correction for remaining race/error paths; no edits/build/tests. No schema change, intentionally preserves existing spawn-before-PID window rather than claiming to solve whole launch serialization."},"attribution":"agent","id":"a030d9fa","parentId":"79b48a0d","timestamp":"2026-09-10T00:54:25.497Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_NBaknv4AarGsi1ej7Z5v9t7v|fc_028129a7f1d66e9c016aa1ffc5d95087d2a33aff7cc62ba852","toolName":"bash","startedAt":"2026-09-10T00:54:30.964Z","args":{"command":"git diff -- crates/spt-store/src/shellinfo.rs crates/spt-daemon/src/shellhost.rs crates/spt-daemon/src/shellwake.rs docs/KNOWN-HAZARDS.md"},"intent":"Reviewing corrective working changes"},"id":"198d6495","parentId":"a030d9fa","timestamp":"2026-09-10T00:54:30.964Z"}
{"type":"message","id":"a7007575","parentId":"198d6495","timestamp":"2026-09-10T00:54:30.965Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":"**Inspecting working-tree diff**"},{"type":"toolCall","id":"call_NBaknv4AarGsi1ej7Z5v9t7v|fc_028129a7f1d66e9c016aa1ffc5d95087d2a33aff7cc62ba852","name":"bash","arguments":{"i":"Reviewing corrective working changes","command":"git diff -- crates/spt-store/src/shellinfo.rs crates/spt-daemon/src/shellhost.rs crates/spt-daemon/src/shellwake.rs docs/KNOWN-HAZARDS.md","cwd":"C:/Users/decid/Documents/projects/spt-core/.worktrees/287-shellrestore"},"intent":"Reviewing corrective working changes"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":421,"output":130,"cacheRead":103808,"cacheWrite":0,"totalTokens":104359,"reasoningTokens":42,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1789001665548,"responseId":"resp_028129a7f1d66e9c016aa1ffc1cc6487d2882719a2f2115ac5","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_028129a7f1d66e9c016aa1ffc3013c87d2b2bd11b9ae59a019","type":"reasoning","content":[],"encrypted_content":"gAAAAABqof_FWEGvojO7fomUAKnJi6JmWp-MQDaPVjNpFC48K5PeHceGFx9_VMba9AVeTJsjyHzTaXWR44ocGcreoMa1MpM_fvF2aD-ETR7pEdgilyecso8bgzk4RgqOiD9CeUhHgAYRBDHUjm3ZyzNHMBqnImHTMOXXUClUPwWrMm0YHdn5reug3NDyPxxXXS2kx0z4Ed_cioTyCoPb76NiZPTJQW6zDqseG3cw5MqpGuwFXRhkgX_FJgRprl7kgXeBMkDbQQHfDwbJNDwfiDdV1-fzPK-Tl6mzUA8Ga6DNQ0lTRBzupgoJITILt8uQbz93QNd8msxSw9cwQHdpQVOS31hiYhJ3QOAE8vKhs0WgfLya2HxHX0O_UuVfO8DRh2Mq4VernelvlwLqxkY4iFpcqz_QJt_mfId_KaQ5IF3tOHWM02IUQ74FEkAfHzZYlrEQz_6bmuOx1p2_MwHyFmJgP7jHE-uyKaNAHpCFxFjgOzVKwIMInBpkEpDbbrrlhUj4V2i-5fCxCDspj5gktDerGtOf4IHfj_KbxX47q5pVpGtod9RTSJmvYiwsTsKtjA_-iTRqoBSCahSIDTYDAbr9iaBD-PYRDChwGd897NCfncllyt7hIxTXYRImCLT6pfL3a680XFY33CQZCKLbjA4GO4-mpmMk_Vlj5YTmHtKvhI1V8mGgoCjzRdPlMMs3IFAwA_ZUA6D7plOMCHyuep7s65xlefRmVWIYwiJzEyknxI03ikBUZqsYTqnPmn8t1vYDVp80KnLwDYXvIoMZx3uJieqYkx6ZzNhKAMyjxloTXMYIcxPIme2CyPjzaORfa9a93UbktZvtLF-igLh_jyj2pnTA9XPDT1Zn16OT0_mMisgQYolgJuh-8QIl-sB7sO0x1lYzH69outj0dhYOteA7xezT0ICLJqoLLZT_miYmMceisbqw1Wx9wlmQX6kuBQX6UsyZMA0rQMAlFRPEyxDYZwYzzU07LgjYjnllcY9b1WzxOySHYp_xItXsBTrOokc6gpczGsWcWTCzyrJZHGjMIlKAIaPBQDmx1uosV3G432JOB5Sr_AJbQXOmKW8ekUYlAI0z5_jRTn7tsGZ-ANfb3gxRCYG3AzGR8JHT2cJOTdIBVb3Fp7HFEEIttmnVhwTqC3qdpeEaLrkVPt5rOn0fECyT5yPe30uZP9py0bDVZmtssqhXz-0Ydx7KMT9blRMfN412SANKZu_7VoqOk3JYghVGZWZHBYW44J9uUAfSGkpuL_7isksIT_EngGhyTNSWeyPGrlg5NslZaSKVB6mXAnYXJoJbT926i9GapItKbWnMirry3CDFw5vrAViaY_bd_ANuruxUeextsVzTVvBSLrWdQDQB62MYRi3xBNfJUgPJjSVJ_BFmlgcevVicUrr54dGIBe-Z4gNfCIQ9l3QELO2W-1EcCJjbdQ0TXOYmWhi3svU1VAeYGNOjEXXUeVXKFE_S0A5vqFrMu_1EKzqApxcJbnCY5G8LzuB526SONK5wtuLhxqKLN_cat_-q9PVSxewLqoEW","internal_chat_message_metadata_passthrough":{"turn_id":"897e71d8-de26-4b06-a763-cbc0db5f84cd"},"summary":[{"type":"summary_text","text":"**Inspecting working-tree diff**"}],"metadata":{"turn_id":"897e71d8-de26-4b06-a763-cbc0db5f84cd"}},{"id":"fc_028129a7f1d66e9c016aa1ffc5d95087d2a33aff7cc62ba852","type":"function_call","status":"completed","arguments":"{\"i\":\"Reviewing corrective working changes\",\"command\":\"git diff -- crates/spt-store/src/shellinfo.rs crates/spt-daemon/src/shellhost.rs crates/spt-daemon/src/shellwake.rs docs/KNOWN-HAZARDS.md\",\"cwd\":\"C:/Users/decid/Documents/projects/spt-core/.worktrees/287-shellrestore\"}","call_id":"call_NBaknv4AarGsi1ej7Z5v9t7v","internal_chat_message_metadata_passthrough":{"create_time":1789001666.272788,"turn_id":"897e71d8-de26-4b06-a763-cbc0db5f84cd"},"name":"bash","metadata":{"turn_id":"897e71d8-de26-4b06-a763-cbc0db5f84cd"}}]},"duration":5413.181600000244,"ttft":1489.092399999965,"completedAt":1789001670963,"contextSnapshot":{"promptTokens":104229,"nonMessageTokens":13209,"compactionEpoch":0}}}
{"type":"message","id":"7b1861c5","parentId":"a7007575","timestamp":"2026-09-10T00:54:31.924Z","message":{"role":"toolResult","toolCallId":"call_NBaknv4AarGsi1ej7Z5v9t7v|fc_028129a7f1d66e9c016aa1ffc5d95087d2a33aff7cc62ba852","toolName":"bash","content":[{"type":"text","text":"crates/spt-daemon/src/shellhost.rs | 18 ++++++------------\ncrates/spt-daemon/src/shellwake.rs | 8 ++------\ncrates/spt-store/src/shellinfo.rs | 31 ++++++++++++++++++++++++\ndocs/KNOWN-HAZARDS.md | 4 ++--\n4 files changed, 41 insertions(+), 20 deletions(-)\n\n--- Changes ---\n\nFile: crates/spt-daemon/src/shellhost.rs\n  @@ -496,13 +496,10 @@ pub fn launch_shell(\n  -            std::fs::write(perch.join(SHELL_PID_FILE), pid.to_string())\n  -                .map_err(|e| format!(\"record pid: {e}\"))?;\n  -            // The birth stamp rides WITH the pid at every launch, or the probe that\n  -            // reads it degrades to bare-pid liveness on this instance alone -- a\n  -            // silent per-instance regression rather than a visible one.\n  +            // Retire the previous birth before publishing the new PID.\n  […3 changed lines elided…]\n  @@ -555,13 +552,10 @@ pub fn launch_shell_brokered_in(\n  -    std::fs::write(perch.join(SHELL_PID_FILE), pid.to_string())\n  -        .map_err(|e| format!(\"record pid: {e}\"))?;\n  -    // Same stamp at the broker-hosted launch. A backend that exposed no pid\n  -    // records 0 here; record_shell_launch stamps no start-time for it, so the\n  -    // probe keeps its fail-toward-alive stance for pid-less backends.\n  +    // PID-less backends keep their fail-toward-alive behavior.\n  […3 changed lines elided…]\n\nFile: crates/spt-daemon/src/shellwake.rs\n  @@ -406,13 +406,9 @@ pub fn resolve_wake(\n  -    // A parked pid that probes alive means somebody won; stand down.\n  +    // A live, birth-matching launch wins even before its bind handshake.\n  -    let live_pid = std::fs::read_to_string(perch.join(crate::shellhost::SHELL_PID_FILE))\n  -        .ok()\n  -        .and_then(|s| s.trim().parse::<u32>().ok())\n  -        .filter(|&p| p != 0 && spt_store::proc::is_process_alive(p));\n  […2 changed lines elided…]\n\nFile: crates/spt-store/src/shellinfo.rs\n  @@ -172,6 +172,22 @@ pub fn record_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) {\n  +/// Publish a new PID without ever pairing it with the previous launch's birth.\n  +/// During publication a missing stamp deliberately preserves live-PID custody.\n  +/// Failure to retire an old stamp must stop publication rather than misidentify\n  +/// the new binary as a recycled corpse.\n  +// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n  +pub fn publish_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) -> std::io::Result<()> {\n  […10 changed lines elided…]\n  @@ -212,6 +228,11 @@ pub fn shell_pid_provably_dead(shell_perch: &Path) -> bool {\n  +    // A relaunch can publish between these reads. Do not apply the new birth\n  +    // to the old PID; an unstable snapshot is not proof that the shell died.\n  +    if read_shell_pid(shell_perch) != Some(pid) {\n  +        return false;\n  +    }\n  @@ -504,6 +525,16 @@ mod tests {\n  +    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n  +    #[test]\n  +    fn launch_publication_refuses_an_unretirable_old_stamp() {\n  +        let d = online_perch(Some(\"123\"));\n  +        std::fs::create_dir(d.path().join(SHELL_LAUNCH_FILE)).unwrap();\n  +        let err = publish_shell_launch(d.path(), std::process::id(), 1_000).unwrap_err();\n  […4 changed lines elided…]\n\nFile: docs/KNOWN-HAZARDS.md\n  @@ -1016,9 +1016,9 @@ The kill-path rule above generalizes: `daemon.pid` is not authority for *\"which\n  -| 2.6 | A shell's ONLINE-ness is DERIVED (recorded status AND a not-provably-dead `shell.pid`) — an abruptly-killed binary breaks no link, so `close_shell`'s o…[+929]\n  +| 2.6 | A shell's ONLINE-ness is DERIVED from recorded status and a not-provably-dead PID/birth pair. Abrupt death and daemon-stop leave stale `online` records…[+477]\n  -| 2.7 | A node restart must not permanently strand every `persistent` shell: (a) the daemon heals a recorded `online` the derivation contradicts, write-guarded…[+1017]\n  +| 2.7 | A once-per-generation boot sweep and each owner offline->online edge restore every instance whose adapter is `persistent`, owner is online, and binary …[+523]\n[raw output: artifact://128]\n\n\nWall time: 0.96 seconds"}],"details":{"timeoutSeconds":300,"wallTimeMs":955.5400999998674},"isError":false,"timestamp":1789001671922}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_46wCDcUuLMgmptx4WXo6Xtfz|fc_028129a7f1d66e9c016aa1ffd0441487d2b0b8038d5ae4d6fa","toolName":"read","startedAt":"2026-09-10T00:54:40.403Z","args":{"path":"artifact://128"},"intent":"Reading corrective patch"},"id":"a73de159","parentId":"7b1861c5","timestamp":"2026-09-10T00:54:40.403Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_JyvpbZS9x9h13RmausPDcB8m|fc_028129a7f1d66e9c016aa1ffd0442487d2894643724b746cd2","toolName":"read","startedAt":"2026-09-10T00:54:40.403Z","args":{"path":".worktrees/287-shellrestore/crates/spt-store/src/shellinfo.rs:155-253:raw"},"intent":"Reading corrected publication context"},"id":"4fd73e6b","parentId":"a73de159","timestamp":"2026-09-10T00:54:40.403Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_Q1j3jvnyWccsC9zSTPuuXl3W|fc_028129a7f1d66e9c016aa1ffd0442c87d28882a706cb7e2b2c","toolName":"read","startedAt":"2026-09-10T00:54:40.403Z","args":{"path":".worktrees/287-shellrestore/crates/spt-store/src/atomic.rs"},"intent":"Reading atomic publication primitive"},"id":"b8373aae","parentId":"4fd73e6b","timestamp":"2026-09-10T00:54:40.403Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_bFajxXim3JTSAWXkRY8uoyD0|fc_028129a7f1d66e9c016aa1ffd0443487d2841c66d6763fd5d1","toolName":"read","startedAt":"2026-09-10T00:54:40.404Z","args":{"path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs:448-565:raw"},"intent":"Reading corrected launch callers"},"id":"6e77c37b","parentId":"b8373aae","timestamp":"2026-09-10T00:54:40.404Z"}
{"type":"message","id":"ec4ec4e7","parentId":"6e77c37b","timestamp":"2026-09-10T00:54:40.404Z","message":{"role":"assistant","content":[{"type":"toolCall","id":"call_46wCDcUuLMgmptx4WXo6Xtfz|fc_028129a7f1d66e9c016aa1ffd0441487d2b0b8038d5ae4d6fa","name":"read","arguments":{"i":"Reading corrective patch","path":"artifact://128"},"intent":"Reading corrective patch"},{"type":"toolCall","id":"call_JyvpbZS9x9h13RmausPDcB8m|fc_028129a7f1d66e9c016aa1ffd0442487d2894643724b746cd2","name":"read","arguments":{"i":"Reading corrected publication context","path":".worktrees/287-shellrestore/crates/spt-store/src/shellinfo.rs:155-253:raw"},"intent":"Reading corrected publication context"},{"type":"toolCall","id":"call_Q1j3jvnyWccsC9zSTPuuXl3W|fc_028129a7f1d66e9c016aa1ffd0442c87d28882a706cb7e2b2c","name":"read","arguments":{"i":"Reading atomic publication primitive","path":".worktrees/287-shellrestore/crates/spt-store/src/atomic.rs"},"intent":"Reading atomic publication primitive"},{"type":"toolCall","id":"call_bFajxXim3JTSAWXkRY8uoyD0|fc_028129a7f1d66e9c016aa1ffd0443487d2841c66d6763fd5d1","name":"read","arguments":{"i":"Reading corrected launch callers","path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs:448-565:raw"},"intent":"Reading corrected launch callers"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":1476,"output":161,"cacheRead":104064,"cacheWrite":0,"totalTokens":105701,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1789001671982,"responseId":"resp_028129a7f1d66e9c016aa1ffc8337887d2bd4d76cb813f5123","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"fc_028129a7f1d66e9c016aa1ffd0441487d2b0b8038d5ae4d6fa","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading corrective patch\",\"path\":\"artifact://128\"}","call_id":"call_46wCDcUuLMgmptx4WXo6Xtfz","internal_chat_message_metadata_passthrough":{"turn_id":"897e71d8-de26-4b06-a763-cbc0db5f84cd"},"name":"read","metadata":{"turn_id":"897e71d8-de26-4b06-a763-cbc0db5f84cd"}},{"id":"fc_028129a7f1d66e9c016aa1ffd0442487d2894643724b746cd2","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading corrected publication context\",\"path\":\".worktrees/287-shellrestore/crates/spt-store/src/shellinfo.rs:155-253:raw\"}","call_id":"call_JyvpbZS9x9h13RmausPDcB8m","internal_chat_message_metadata_passthrough":{"turn_id":"897e71d8-de26-4b06-a763-cbc0db5f84cd"},"name":"read","metadata":{"turn_id":"897e71d8-de26-4b06-a763-cbc0db5f84cd"}},{"id":"fc_028129a7f1d66e9c016aa1ffd0442c87d28882a706cb7e2b2c","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading atomic publication primitive\",\"path\":\".worktrees/287-shellrestore/crates/spt-store/src/atomic.rs\"}","call_id":"call_Q1j3jvnyWccsC9zSTPuuXl3W","internal_chat_message_metadata_passthrough":{"turn_id":"897e71d8-de26-4b06-a763-cbc0db5f84cd"},"name":"read","metadata":{"turn_id":"897e71d8-de26-4b06-a763-cbc0db5f84cd"}},{"id":"fc_028129a7f1d66e9c016aa1ffd0443487d2841c66d6763fd5d1","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading corrected launch callers\",\"path\":\".worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs:448-565:raw\"}","call_id":"call_bFajxXim3JTSAWXkRY8uoyD0","internal_chat_message_metadata_passthrough":{"turn_id":"897e71d8-de26-4b06-a763-cbc0db5f84cd"},"name":"read","metadata":{"turn_id":"897e71d8-de26-4b06-a763-cbc0db5f84cd"}}]},"duration":8417.753600000404,"ttft":8312.25130000012,"completedAt":1789001680401,"contextSnapshot":{"promptTokens":105540,"nonMessageTokens":13209,"compactionEpoch":0}}}
{"type":"message","id":"8b8d576f","parentId":"ec4ec4e7","timestamp":"2026-09-10T00:54:40.419Z","message":{"role":"toolResult","toolCallId":"call_bFajxXim3JTSAWXkRY8uoyD0|fc_028129a7f1d66e9c016aa1ffd0443487d2841c66d6763fd5d1","toolName":"read","content":[{"type":"text","text":"/// switch any more than spawn is.\n// [impl->REQ-SHELL-1]\npub fn launch_shell(\n    owlery: &Path,\n    owner: &str,\n    shell_id: &str,\n    adapter_name: &str,\n    install_dir: Option<&Path>,\n    shell: &Shell,\n) -> Result<u32, String> {\n    match shell.command_receipt {\n        Some(CommandReceipt::Http) => Err(\n            \"command_receipt=http is not implemented yet — spawn refused (an \\\n             unimplemented mode never falls back silently)\"\n                .into(),\n        ),\n        Some(CommandReceipt::Stdin) => {\n            crate::ensure_running().map_err(|e| format!(\"daemon start: {e}\"))?;\n            launch_shell_brokered_in(\n                &crate::broker_socket_name(),\n                owlery,\n                owner,\n                shell_id,\n                adapter_name,\n                install_dir,\n                shell,\n            )\n        }\n        _ => {\n            let tokens =\n                fill_spawn_command(owlery, owner, shell_id, adapter_name, install_dir, shell)?;\n            let (program, args) = tokens.split_first().expect(\"non-empty checked\");\n            // The shell binary is a LONG-LIVED detached child (KNOWN-HAZARDS\n            // 5.6): on Windows it must inherit NO handles, or a caller\n            // capturing `spt shell spawn`'s output — at any depth of the\n            // process chain — hangs forever on a pipe write-end the immortal\n            // child holds. Same no-inherit spawn as the daemon's.\n            // [impl->REQ-HAZARD-DETACHED-PIPE-INHERIT]\n            #[cfg(windows)]\n            let pid = crate::daemon::detached_no_inherit(program, args)\n                .map_err(|e| format!(\"spawn {program}: {e}\"))?;\n            #[cfg(unix)]\n            let pid = Command::new(program)\n                .args(args)\n                .stdin(Stdio::null())\n                .stdout(Stdio::null())\n                .stderr(Stdio::null())\n                .spawn()\n                .map_err(|e| format!(\"spawn {program}: {e}\"))?\n                .id();\n            let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n            // Retire the previous birth before publishing the new PID.\n            // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n            spt_store::shellinfo::publish_shell_launch(&perch, pid, now_ms())\n                .map_err(|e| format!(\"record launch: {e}\"))?;\n            Ok(pid)\n        }\n    }\n}\n\n/// The stdin-receipt launch: spawn the filled template as a **broker-hosted**\n/// session labeled [`shell_session_label`], so the daemon owns the child's\n/// stdin and `spt shell cmd` can drive it through the `send_effect` input\n/// seam (exactly-once per spool row). The broker outlives this connection —\n/// the session keeps running after the spawning call returns.\n// [impl->REQ-SHELL-1]\npub fn launch_shell_brokered_in(\n    broker_name: &str,\n    owlery: &Path,\n    owner: &str,\n    shell_id: &str,\n    adapter_name: &str,\n    install_dir: Option<&Path>,\n    shell: &Shell,\n) -> Result<u32, String> {\n    let tokens = fill_spawn_command(owlery, owner, shell_id, adapter_name, install_dir, shell)?;\n    let (program, args) = tokens.split_first().expect(\"non-empty checked\");\n    let mut brain =\n        Brain::cold_start(broker_name, now_ms()).map_err(|e| format!(\"broker connect: {e}\"))?;\n    let (_sid, pid) = brain\n        .spawn_session_pid(SpawnReq {\n            program: program.clone(),\n            args: args.to_vec(),\n            // A shell binary is not a terminal: the PTY is only the stdin\n            // conduit. Keep it wide so renderer-side wrapping (ConPTY) can\n            // never split a >80-col stamped frame mid-line.\n            rows: 24,\n            cols: 512,\n            endpoint: shell_session_label(owner, shell_id),\n            // Shells inherit the broker's cwd (unchanged); project-scoped cwd is\n            // a harness-bringup concern (REQ-HOST-RUN-2), not shell spawn.\n            cwd: None,\n            // Shells inherit the broker's env; [env] inject is a harness-bringup\n            // concern (REQ-HAZARD-ENV-SUBST), not shell spawn.\n            env: Default::default(),\n            translation_binary: None,\n            adapter: String::new(),\n            // Deliberately None even though a shell now HAS an install dir: the\n            // broker's copy exists to resolve a translation binary (there is\n            // none here) and would otherwise re-resolve a program token\n            // `fill_spawn_command` has already resolved — a second resolution\n            // path is exactly what REQ-INSTALL-11 has one primitive to avoid.\n            install_dir: None,\n        })\n        .map_err(|e| format!(\"broker spawn: {e}\"))?;\n    let pid = pid.unwrap_or(0);\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n    // PID-less backends keep their fail-toward-alive behavior.\n    // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n    spt_store::shellinfo::publish_shell_launch(&perch, pid, now_ms())\n        .map_err(|e| format!(\"record launch: {e}\"))?;\n    Ok(pid)\n}\n\n/// Resolve a shell instance by its link token alone — the \"owner from the\n/// link\" lookup every shell-flavored machinery call shares. Scans every\n/// owner's `shells/` namespace for the perch parking exactly this token; an\n/// unknown/empty token resolves nothing. Pure read — onlining is the bind's\n\n[Showing lines 448-565 of 1821. Use :566 to continue]"}],"details":{"truncation":{"content":"/// switch any more than spawn is.\r\n// [impl->REQ-SHELL-1]\r\npub fn launch_shell(\r\n    owlery: &Path,\r\n    owner: &str,\r\n    shell_id: &str,\r\n    adapter_name: &str,\r\n    install_dir: Option<&Path>,\r\n    shell: &Shell,\r\n) -> Result<u32, String> {\r\n    match shell.command_receipt {\r\n        Some(CommandReceipt::Http) => Err(\r\n            \"command_receipt=http is not implemented yet — spawn refused (an \\\r\n             unimplemented mode never falls back silently)\"\r\n                .into(),\r\n        ),\r\n        Some(CommandReceipt::Stdin) => {\r\n            crate::ensure_running().map_err(|e| format!(\"daemon start: {e}\"))?;\r\n            launch_shell_brokered_in(\r\n                &crate::broker_socket_name(),\r\n                owlery,\r\n                owner,\r\n                shell_id,\r\n                adapter_name,\r\n                install_dir,\r\n                shell,\r\n            )\r\n        }\r\n        _ => {\r\n            let tokens =\r\n                fill_spawn_command(owlery, owner, shell_id, adapter_name, install_dir, shell)?;\r\n            let (program, args) = tokens.split_first().expect(\"non-empty checked\");\r\n            // The shell binary is a LONG-LIVED detached child (KNOWN-HAZARDS\r\n            // 5.6): on Windows it must inherit NO handles, or a caller\r\n            // capturing `spt shell spawn`'s output — at any depth of the\r\n            // process chain — hangs forever on a pipe write-end the immortal\r\n            // child holds. Same no-inherit spawn as the daemon's.\r\n            // [impl->REQ-HAZARD-DETACHED-PIPE-INHERIT]\r\n            #[cfg(windows)]\r\n            let pid = crate::daemon::detached_no_inherit(program, args)\r\n                .map_err(|e| format!(\"spawn {program}: {e}\"))?;\r\n            #[cfg(unix)]\r\n            let pid = Command::new(program)\r\n                .args(args)\r\n                .stdin(Stdio::null())\r\n                .stdout(Stdio::null())\r\n                .stderr(Stdio::null())\r\n                .spawn()\r\n                .map_err(|e| format!(\"spawn {program}: {e}\"))?\r\n                .id();\r\n            let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\r\n            // Retire the previous birth before publishing the new PID.\r\n            // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\n            spt_store::shellinfo::publish_shell_launch(&perch, pid, now_ms())\r\n                .map_err(|e| format!(\"record launch: {e}\"))?;\r\n            Ok(pid)\r\n        }\r\n    }\r\n}\r\n\r\n/// The stdin-receipt launch: spawn the filled template as a **broker-hosted**\r\n/// session labeled [`shell_session_label`], so the daemon owns the child's\r\n/// stdin and `spt shell cmd` can drive it through the `send_effect` input\r\n/// seam (exactly-once per spool row). The broker outlives this connection —\r\n/// the session keeps running after the spawning call returns.\r\n// [impl->REQ-SHELL-1]\r\npub fn launch_shell_brokered_in(\r\n    broker_name: &str,\r\n    owlery: &Path,\r\n    owner: &str,\r\n    shell_id: &str,\r\n    adapter_name: &str,\r\n    install_dir: Option<&Path>,\r\n    shell: &Shell,\r\n) -> Result<u32, String> {\r\n    let tokens = fill_spawn_command(owlery, owner, shell_id, adapter_name, install_dir, shell)?;\r\n    let (program, args) = tokens.split_first().expect(\"non-empty checked\");\r\n    let mut brain =\r\n        Brain::cold_start(broker_name, now_ms()).map_err(|e| format!(\"broker connect: {e}\"))?;\r\n    let (_sid, pid) = brain\r\n        .spawn_session_pid(SpawnReq {\r\n            program: program.clone(),\r\n            args: args.to_vec(),\r\n            // A shell binary is not a terminal: the PTY is only the stdin\r\n            // conduit. Keep it wide so renderer-side wrapping (ConPTY) can\r\n            // never split a >80-col stamped frame mid-line.\r\n            rows: 24,\r\n            cols: 512,\r\n            endpoint: shell_session_label(owner, shell_id),\r\n            // Shells inherit the broker's cwd (unchanged); project-scoped cwd is\r\n            // a harness-bringup concern (REQ-HOST-RUN-2), not shell spawn.\r\n            cwd: None,\r\n            // Shells inherit the broker's env; [env] inject is a harness-bringup\r\n            // concern (REQ-HAZARD-ENV-SUBST), not shell spawn.\r\n            env: Default::default(),\r\n            translation_binary: None,\r\n            adapter: String::new(),\r\n            // Deliberately None even though a shell now HAS an install dir: the\r\n            // broker's copy exists to resolve a translation binary (there is\r\n            // none here) and would otherwise re-resolve a program token\r\n            // `fill_spawn_command` has already resolved — a second resolution\r\n            // path is exactly what REQ-INSTALL-11 has one primitive to avoid.\r\n            install_dir: None,\r\n        })\r\n        .map_err(|e| format!(\"broker spawn: {e}\"))?;\r\n    let pid = pid.unwrap_or(0);\r\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\r\n    // PID-less backends keep their fail-toward-alive behavior.\r\n    // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\n    spt_store::shellinfo::publish_shell_launch(&perch, pid, now_ms())\r\n        .map_err(|e| format!(\"record launch: {e}\"))?;\r\n    Ok(pid)\r\n}\r\n\r\n/// Resolve a shell instance by its link token alone — the \"owner from the\r\n/// link\" lookup every shell-flavored machinery call shares. Scans every\r\n/// owner's `shells/` namespace for the perch parking exactly this token; an\r\n/// unknown/empty token resolves nothing. Pure read — onlining is the bind's\r","truncated":true,"truncatedBy":"lines","totalLines":1374,"totalBytes":5487,"outputLines":118,"outputBytes":5487,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":1821,"displayContent":{"text":"/// switch any more than spawn is.\r\n// [impl->REQ-SHELL-1]\r\npub fn launch_shell(\r\n    owlery: &Path,\r\n    owner: &str,\r\n    shell_id: &str,\r\n    adapter_name: &str,\r\n    install_dir: Option<&Path>,\r\n    shell: &Shell,\r\n) -> Result<u32, String> {\r\n    match shell.command_receipt {\r\n        Some(CommandReceipt::Http) => Err(\r\n            \"command_receipt=http is not implemented yet — spawn refused (an \\\r\n             unimplemented mode never falls back silently)\"\r\n                .into(),\r\n        ),\r\n        Some(CommandReceipt::Stdin) => {\r\n            crate::ensure_running().map_err(|e| format!(\"daemon start: {e}\"))?;\r\n            launch_shell_brokered_in(\r\n                &crate::broker_socket_name(),\r\n                owlery,\r\n                owner,\r\n                shell_id,\r\n                adapter_name,\r\n                install_dir,\r\n                shell,\r\n            )\r\n        }\r\n        _ => {\r\n            let tokens =\r\n                fill_spawn_command(owlery, owner, shell_id, adapter_name, install_dir, shell)?;\r\n            let (program, args) = tokens.split_first().expect(\"non-empty checked\");\r\n            // The shell binary is a LONG-LIVED detached child (KNOWN-HAZARDS\r\n            // 5.6): on Windows it must inherit NO handles, or a caller\r\n            // capturing `spt shell spawn`'s output — at any depth of the\r\n            // process chain — hangs forever on a pipe write-end the immortal\r\n            // child holds. Same no-inherit spawn as the daemon's.\r\n            // [impl->REQ-HAZARD-DETACHED-PIPE-INHERIT]\r\n            #[cfg(windows)]\r\n            let pid = crate::daemon::detached_no_inherit(program, args)\r\n                .map_err(|e| format!(\"spawn {program}: {e}\"))?;\r\n            #[cfg(unix)]\r\n            let pid = Command::new(program)\r\n                .args(args)\r\n                .stdin(Stdio::null())\r\n                .stdout(Stdio::null())\r\n                .stderr(Stdio::null())\r\n                .spawn()\r\n                .map_err(|e| format!(\"spawn {program}: {e}\"))?\r\n                .id();\r\n            let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\r\n            // Retire the previous birth before publishing the new PID.\r\n            // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\n            spt_store::shellinfo::publish_shell_launch(&perch, pid, now_ms())\r\n                .map_err(|e| format!(\"record launch: {e}\"))?;\r\n            Ok(pid)\r\n        }\r\n    }\r\n}\r\n\r\n/// The stdin-receipt launch: spawn the filled template as a **broker-hosted**\r\n/// session labeled [`shell_session_label`], so the daemon owns the child's\r\n/// stdin and `spt shell cmd` can drive it through the `send_effect` input\r\n/// seam (exactly-once per spool row). The broker outlives this connection —\r\n/// the session keeps running after the spawning call returns.\r\n// [impl->REQ-SHELL-1]\r\npub fn launch_shell_brokered_in(\r\n    broker_name: &str,\r\n    owlery: &Path,\r\n    owner: &str,\r\n    shell_id: &str,\r\n    adapter_name: &str,\r\n    install_dir: Option<&Path>,\r\n    shell: &Shell,\r\n) -> Result<u32, String> {\r\n    let tokens = fill_spawn_command(owlery, owner, shell_id, adapter_name, install_dir, shell)?;\r\n    let (program, args) = tokens.split_first().expect(\"non-empty checked\");\r\n    let mut brain =\r\n        Brain::cold_start(broker_name, now_ms()).map_err(|e| format!(\"broker connect: {e}\"))?;\r\n    let (_sid, pid) = brain\r\n        .spawn_session_pid(SpawnReq {\r\n            program: program.clone(),\r\n            args: args.to_vec(),\r\n            // A shell binary is not a terminal: the PTY is only the stdin\r\n            // conduit. Keep it wide so renderer-side wrapping (ConPTY) can\r\n            // never split a >80-col stamped frame mid-line.\r\n            rows: 24,\r\n            cols: 512,\r\n            endpoint: shell_session_label(owner, shell_id),\r\n            // Shells inherit the broker's cwd (unchanged); project-scoped cwd is\r\n            // a harness-bringup concern (REQ-HOST-RUN-2), not shell spawn.\r\n            cwd: None,\r\n            // Shells inherit the broker's env; [env] inject is a harness-bringup\r\n            // concern (REQ-HAZARD-ENV-SUBST), not shell spawn.\r\n            env: Default::default(),\r\n            translation_binary: None,\r\n            adapter: String::new(),\r\n            // Deliberately None even though a shell now HAS an install dir: the\r\n            // broker's copy exists to resolve a translation binary (there is\r\n            // none here) and would otherwise re-resolve a program token\r\n            // `fill_spawn_command` has already resolved — a second resolution\r\n            // path is exactly what REQ-INSTALL-11 has one primitive to avoid.\r\n            install_dir: None,\r\n        })\r\n        .map_err(|e| format!(\"broker spawn: {e}\"))?;\r\n    let pid = pid.unwrap_or(0);\r\n    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\r\n    // PID-less backends keep their fail-toward-alive behavior.\r\n    // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\n    spt_store::shellinfo::publish_shell_launch(&perch, pid, now_ms())\r\n        .map_err(|e| format!(\"record launch: {e}\"))?;\r\n    Ok(pid)\r\n}\r\n\r\n/// Resolve a shell instance by its link token alone — the \"owner from the\r\n/// link\" lookup every shell-flavored machinery call shares. Scans every\r\n/// owner's `shells/` namespace for the perch parking exactly this token; an\r\n/// unknown/empty token resolves nothing. Pure read — onlining is the bind's\r","startLine":448,"lineNumbers":[448,449,450,451,452,453,454,455,456,457,458,459,460,461,462,463,464,465,466,467,468,469,470,471,472,473,474,475,476,477,478,479,480,481,482,483,484,485,486,487,488,489,490,491,492,493,494,495,496,497,498,499,500,501,502,503,504,505,506,507,508,509,510,511,512,513,514,515,516,517,518,519,520,521,522,523,524,525,526,527,528,529,530,531,532,533,534,535,536,537,538,539,540,541,542,543,544,545,546,547,548,549,550,551,552,553,554,555,556,557,558,559,560,561,562,563,564,565]},"fileSize":87896,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\crates\\spt-daemon\\src\\shellhost.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":1821,"totalBytes":5487,"outputLines":118,"outputBytes":5487,"shownRange":{"start":448,"end":565},"nextOffset":566}}},"isError":false,"timestamp":1789001680418}}
{"type":"message","id":"4046fad4","parentId":"8b8d576f","timestamp":"2026-09-10T00:54:40.421Z","message":{"role":"toolResult","toolCallId":"call_JyvpbZS9x9h13RmausPDcB8m|fc_028129a7f1d66e9c016aa1ffd0442487d2894643724b746cd2","toolName":"read","content":[{"type":"text","text":"    pub pid_started_at: Option<u64>,\n    /// Wall-clock epoch-ms at which this instance was launched.\n    pub launched_ms: u64,\n}\n\n/// Record the birth stamp for a just-launched `pid`. Best-effort: a stamp that\n/// cannot be written must never fail the launch — a missing stamp degrades the\n/// probe to its pre-existing fail-toward-alive behaviour, which is the safe\n/// direction, while a failed launch is a real outage.\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\npub fn record_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) {\n    let launch = ShellLaunch {\n        pid_started_at: crate::proc::process_started_at(pid),\n        launched_ms: now_ms,\n    };\n    if let Ok(json) = serde_json::to_string_pretty(&launch) {\n        let _ = atomic_write_string(&shell_perch.join(SHELL_LAUNCH_FILE), &json);\n    }\n}\n\n/// Publish a new PID without ever pairing it with the previous launch's birth.\n/// During publication a missing stamp deliberately preserves live-PID custody.\n/// Failure to retire an old stamp must stop publication rather than misidentify\n/// the new binary as a recycled corpse.\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\npub fn publish_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) -> std::io::Result<()> {\n    match std::fs::remove_file(shell_perch.join(SHELL_LAUNCH_FILE)) {\n        Ok(()) => {}\n        Err(e) if e.kind() == std::io::ErrorKind::NotFound => {}\n        Err(e) => return Err(e),\n    }\n    atomic_write_string(&shell_perch.join(SHELL_PID_FILE), &pid.to_string())?;\n    record_shell_launch(shell_perch, pid, now_ms);\n    Ok(())\n}\n\n/// Read the birth stamp, if this instance was launched by a build that parks one.\n/// Absent or unparseable ⇒ `None`, which every consumer treats as \"no stamp to\n/// compare\" rather than as a fault.\npub fn read_shell_launch(shell_perch: &Path) -> Option<ShellLaunch> {\n    let raw = std::fs::read_to_string(shell_perch.join(SHELL_LAUNCH_FILE)).ok()?;\n    serde_json::from_str(&raw).ok()\n}\n\n/// **The** discriminant behind [`effective_status`]: is this instance's recorded\n/// binary *provably* gone? Narrow on purpose — true only when the perch parks a\n/// pid, it parses non-zero, and the process is not alive.\n///\n/// Everything else reads **alive**: an absent or unparseable pid file, and the\n/// `0` a broker-hosted spawn records when its backend exposed no pid\n/// ([`crate::proc::is_process_alive`] already reads 0 as dead, so the explicit\n/// filter here is what keeps a pid-less backend from being declared dead by this\n/// resolver). That fail-toward-alive stance mirrors `liveness.rs`'s interim\n/// parity: this predicate exists to catch a *corpse*, never to guess at a\n/// backend it cannot see.\n///\n/// Recycled pids read alive, so a heal can be **missed** but never **mis-fired**\n/// — the failure direction is \"stays stale\", never \"tears down a live instance\".\n// [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\npub fn shell_pid_provably_dead(shell_perch: &Path) -> bool {\n    let Some(pid) = read_shell_pid(shell_perch).filter(|&pid| pid != 0) else {\n        return false;\n    };\n    // BAROMETER W2 (releases#78): probe pid AND birth stamp, not the pid alone.\n    // A pid namespace resets at boot, so a corpse pid can be REUSED by an\n    // unrelated process started after the restart; under bare-pid liveness that\n    // reads ALIVE, the derived status reads online, and `relink` then refuses\n    // SHELL_ALREADY_ONLINE -- the one command that recovers the instance\n    // refusing on the very state that is wrong. The PAIR test (same pid,\n    // different start time) is what tells a reused pid from our own process.\n    //\n    // The stance is unchanged where there is no stamp to compare: an instance\n    // launched before this field existed, or a backend that exposed no start\n    // time, still falls back to plain absence and fails toward ALIVE.\n    // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n    let birth = read_shell_launch(shell_perch).and_then(|l| l.pid_started_at);\n    // A relaunch can publish between these reads. Do not apply the new birth\n    // to the old PID; an unstable snapshot is not proof that the shell died.\n    if read_shell_pid(shell_perch) != Some(pid) {\n        return false;\n    }\n    matches!(\n        crate::liveness::relay_liveness(Some(pid), birth),\n        crate::liveness::RelayLiveness::Gone\n    )\n}\n\n/// A shell instance's **derived** status — the shell-side twin of\n/// `liveness::is_perch_alive` (REQ-HAZARD-DAEMON-HOSTED-LIVENESS gave agent\n/// perches exactly this; shells never got it).\n///\n/// The recorded `status` has exactly two writers — the bind handshake (→online)\n/// and `spt_daemon::shellhost::close_shell` (→offline) — and `close_shell` runs\n/// on the **link-break** path alone. A binary that dies\n/// abruptly (force-kill, crash, OOM) breaks no link, so nothing ever writes the\n/// flip and the record says `online` forever. Deriving closes that hole without\n/// adding a supervisor: `online` + a provably-dead pid reads\n/// [`SHELL_STATUS_OFFLINE`]; every other case is the recorded value verbatim.\n///\n\n[Showing lines 155-253 of 758. Use :254 to continue]"}],"details":{"truncation":{"content":"    pub pid_started_at: Option<u64>,\r\n    /// Wall-clock epoch-ms at which this instance was launched.\r\n    pub launched_ms: u64,\r\n}\r\n\r\n/// Record the birth stamp for a just-launched `pid`. Best-effort: a stamp that\r\n/// cannot be written must never fail the launch — a missing stamp degrades the\r\n/// probe to its pre-existing fail-toward-alive behaviour, which is the safe\r\n/// direction, while a failed launch is a real outage.\r\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\npub fn record_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) {\r\n    let launch = ShellLaunch {\r\n        pid_started_at: crate::proc::process_started_at(pid),\r\n        launched_ms: now_ms,\r\n    };\r\n    if let Ok(json) = serde_json::to_string_pretty(&launch) {\r\n        let _ = atomic_write_string(&shell_perch.join(SHELL_LAUNCH_FILE), &json);\r\n    }\r\n}\r\n\r\n/// Publish a new PID without ever pairing it with the previous launch's birth.\r\n/// During publication a missing stamp deliberately preserves live-PID custody.\r\n/// Failure to retire an old stamp must stop publication rather than misidentify\r\n/// the new binary as a recycled corpse.\r\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\npub fn publish_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) -> std::io::Result<()> {\r\n    match std::fs::remove_file(shell_perch.join(SHELL_LAUNCH_FILE)) {\r\n        Ok(()) => {}\r\n        Err(e) if e.kind() == std::io::ErrorKind::NotFound => {}\r\n        Err(e) => return Err(e),\r\n    }\r\n    atomic_write_string(&shell_perch.join(SHELL_PID_FILE), &pid.to_string())?;\r\n    record_shell_launch(shell_perch, pid, now_ms);\r\n    Ok(())\r\n}\r\n\r\n/// Read the birth stamp, if this instance was launched by a build that parks one.\r\n/// Absent or unparseable ⇒ `None`, which every consumer treats as \"no stamp to\r\n/// compare\" rather than as a fault.\r\npub fn read_shell_launch(shell_perch: &Path) -> Option<ShellLaunch> {\r\n    let raw = std::fs::read_to_string(shell_perch.join(SHELL_LAUNCH_FILE)).ok()?;\r\n    serde_json::from_str(&raw).ok()\r\n}\r\n\r\n/// **The** discriminant behind [`effective_status`]: is this instance's recorded\r\n/// binary *provably* gone? Narrow on purpose — true only when the perch parks a\r\n/// pid, it parses non-zero, and the process is not alive.\r\n///\r\n/// Everything else reads **alive**: an absent or unparseable pid file, and the\r\n/// `0` a broker-hosted spawn records when its backend exposed no pid\r\n/// ([`crate::proc::is_process_alive`] already reads 0 as dead, so the explicit\r\n/// filter here is what keeps a pid-less backend from being declared dead by this\r\n/// resolver). That fail-toward-alive stance mirrors `liveness.rs`'s interim\r\n/// parity: this predicate exists to catch a *corpse*, never to guess at a\r\n/// backend it cannot see.\r\n///\r\n/// Recycled pids read alive, so a heal can be **missed** but never **mis-fired**\r\n/// — the failure direction is \"stays stale\", never \"tears down a live instance\".\r\n// [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\r\npub fn shell_pid_provably_dead(shell_perch: &Path) -> bool {\r\n    let Some(pid) = read_shell_pid(shell_perch).filter(|&pid| pid != 0) else {\r\n        return false;\r\n    };\r\n    // BAROMETER W2 (releases#78): probe pid AND birth stamp, not the pid alone.\r\n    // A pid namespace resets at boot, so a corpse pid can be REUSED by an\r\n    // unrelated process started after the restart; under bare-pid liveness that\r\n    // reads ALIVE, the derived status reads online, and `relink` then refuses\r\n    // SHELL_ALREADY_ONLINE -- the one command that recovers the instance\r\n    // refusing on the very state that is wrong. The PAIR test (same pid,\r\n    // different start time) is what tells a reused pid from our own process.\r\n    //\r\n    // The stance is unchanged where there is no stamp to compare: an instance\r\n    // launched before this field existed, or a backend that exposed no start\r\n    // time, still falls back to plain absence and fails toward ALIVE.\r\n    // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\n    let birth = read_shell_launch(shell_perch).and_then(|l| l.pid_started_at);\r\n    // A relaunch can publish between these reads. Do not apply the new birth\r\n    // to the old PID; an unstable snapshot is not proof that the shell died.\r\n    if read_shell_pid(shell_perch) != Some(pid) {\r\n        return false;\r\n    }\r\n    matches!(\r\n        crate::liveness::relay_liveness(Some(pid), birth),\r\n        crate::liveness::RelayLiveness::Gone\r\n    )\r\n}\r\n\r\n/// A shell instance's **derived** status — the shell-side twin of\r\n/// `liveness::is_perch_alive` (REQ-HAZARD-DAEMON-HOSTED-LIVENESS gave agent\r\n/// perches exactly this; shells never got it).\r\n///\r\n/// The recorded `status` has exactly two writers — the bind handshake (→online)\r\n/// and `spt_daemon::shellhost::close_shell` (→offline) — and `close_shell` runs\r\n/// on the **link-break** path alone. A binary that dies\r\n/// abruptly (force-kill, crash, OOM) breaks no link, so nothing ever writes the\r\n/// flip and the record says `online` forever. Deriving closes that hole without\r\n/// adding a supervisor: `online` + a provably-dead pid reads\r\n/// [`SHELL_STATUS_OFFLINE`]; every other case is the recorded value verbatim.\r\n///\r","truncated":true,"truncatedBy":"lines","totalLines":604,"totalBytes":5174,"outputLines":99,"outputBytes":5174,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":758,"displayContent":{"text":"    pub pid_started_at: Option<u64>,\r\n    /// Wall-clock epoch-ms at which this instance was launched.\r\n    pub launched_ms: u64,\r\n}\r\n\r\n/// Record the birth stamp for a just-launched `pid`. Best-effort: a stamp that\r\n/// cannot be written must never fail the launch — a missing stamp degrades the\r\n/// probe to its pre-existing fail-toward-alive behaviour, which is the safe\r\n/// direction, while a failed launch is a real outage.\r\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\npub fn record_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) {\r\n    let launch = ShellLaunch {\r\n        pid_started_at: crate::proc::process_started_at(pid),\r\n        launched_ms: now_ms,\r\n    };\r\n    if let Ok(json) = serde_json::to_string_pretty(&launch) {\r\n        let _ = atomic_write_string(&shell_perch.join(SHELL_LAUNCH_FILE), &json);\r\n    }\r\n}\r\n\r\n/// Publish a new PID without ever pairing it with the previous launch's birth.\r\n/// During publication a missing stamp deliberately preserves live-PID custody.\r\n/// Failure to retire an old stamp must stop publication rather than misidentify\r\n/// the new binary as a recycled corpse.\r\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\npub fn publish_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) -> std::io::Result<()> {\r\n    match std::fs::remove_file(shell_perch.join(SHELL_LAUNCH_FILE)) {\r\n        Ok(()) => {}\r\n        Err(e) if e.kind() == std::io::ErrorKind::NotFound => {}\r\n        Err(e) => return Err(e),\r\n    }\r\n    atomic_write_string(&shell_perch.join(SHELL_PID_FILE), &pid.to_string())?;\r\n    record_shell_launch(shell_perch, pid, now_ms);\r\n    Ok(())\r\n}\r\n\r\n/// Read the birth stamp, if this instance was launched by a build that parks one.\r\n/// Absent or unparseable ⇒ `None`, which every consumer treats as \"no stamp to\r\n/// compare\" rather than as a fault.\r\npub fn read_shell_launch(shell_perch: &Path) -> Option<ShellLaunch> {\r\n    let raw = std::fs::read_to_string(shell_perch.join(SHELL_LAUNCH_FILE)).ok()?;\r\n    serde_json::from_str(&raw).ok()\r\n}\r\n\r\n/// **The** discriminant behind [`effective_status`]: is this instance's recorded\r\n/// binary *provably* gone? Narrow on purpose — true only when the perch parks a\r\n/// pid, it parses non-zero, and the process is not alive.\r\n///\r\n/// Everything else reads **alive**: an absent or unparseable pid file, and the\r\n/// `0` a broker-hosted spawn records when its backend exposed no pid\r\n/// ([`crate::proc::is_process_alive`] already reads 0 as dead, so the explicit\r\n/// filter here is what keeps a pid-less backend from being declared dead by this\r\n/// resolver). That fail-toward-alive stance mirrors `liveness.rs`'s interim\r\n/// parity: this predicate exists to catch a *corpse*, never to guess at a\r\n/// backend it cannot see.\r\n///\r\n/// Recycled pids read alive, so a heal can be **missed** but never **mis-fired**\r\n/// — the failure direction is \"stays stale\", never \"tears down a live instance\".\r\n// [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\r\npub fn shell_pid_provably_dead(shell_perch: &Path) -> bool {\r\n    let Some(pid) = read_shell_pid(shell_perch).filter(|&pid| pid != 0) else {\r\n        return false;\r\n    };\r\n    // BAROMETER W2 (releases#78): probe pid AND birth stamp, not the pid alone.\r\n    // A pid namespace resets at boot, so a corpse pid can be REUSED by an\r\n    // unrelated process started after the restart; under bare-pid liveness that\r\n    // reads ALIVE, the derived status reads online, and `relink` then refuses\r\n    // SHELL_ALREADY_ONLINE -- the one command that recovers the instance\r\n    // refusing on the very state that is wrong. The PAIR test (same pid,\r\n    // different start time) is what tells a reused pid from our own process.\r\n    //\r\n    // The stance is unchanged where there is no stamp to compare: an instance\r\n    // launched before this field existed, or a backend that exposed no start\r\n    // time, still falls back to plain absence and fails toward ALIVE.\r\n    // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\n    let birth = read_shell_launch(shell_perch).and_then(|l| l.pid_started_at);\r\n    // A relaunch can publish between these reads. Do not apply the new birth\r\n    // to the old PID; an unstable snapshot is not proof that the shell died.\r\n    if read_shell_pid(shell_perch) != Some(pid) {\r\n        return false;\r\n    }\r\n    matches!(\r\n        crate::liveness::relay_liveness(Some(pid), birth),\r\n        crate::liveness::RelayLiveness::Gone\r\n    )\r\n}\r\n\r\n/// A shell instance's **derived** status — the shell-side twin of\r\n/// `liveness::is_perch_alive` (REQ-HAZARD-DAEMON-HOSTED-LIVENESS gave agent\r\n/// perches exactly this; shells never got it).\r\n///\r\n/// The recorded `status` has exactly two writers — the bind handshake (→online)\r\n/// and `spt_daemon::shellhost::close_shell` (→offline) — and `close_shell` runs\r\n/// on the **link-break** path alone. A binary that dies\r\n/// abruptly (force-kill, crash, OOM) breaks no link, so nothing ever writes the\r\n/// flip and the record says `online` forever. Deriving closes that hole without\r\n/// adding a supervisor: `online` + a provably-dead pid reads\r\n/// [`SHELL_STATUS_OFFLINE`]; every other case is the recorded value verbatim.\r\n///\r","startLine":155,"lineNumbers":[155,156,157,158,159,160,161,162,163,164,165,166,167,168,169,170,171,172,173,174,175,176,177,178,179,180,181,182,183,184,185,186,187,188,189,190,191,192,193,194,195,196,197,198,199,200,201,202,203,204,205,206,207,208,209,210,211,212,213,214,215,216,217,218,219,220,221,222,223,224,225,226,227,228,229,230,231,232,233,234,235,236,237,238,239,240,241,242,243,244,245,246,247,248,249,250,251,252,253]},"fileSize":35109,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\crates\\spt-store\\src\\shellinfo.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":758,"totalBytes":5174,"outputLines":99,"outputBytes":5174,"shownRange":{"start":155,"end":253},"nextOffset":254}}},"isError":false,"timestamp":1789001680420}}
{"type":"message","id":"828f7fdb","parentId":"4046fad4","timestamp":"2026-09-10T00:54:40.441Z","message":{"role":"toolResult","toolCallId":"call_Q1j3jvnyWccsC9zSTPuuXl3W|fc_028129a7f1d66e9c016aa1ffd0442c87d28882a706cb7e2b2c","toolName":"read","content":[{"type":"text","text":"//! Atomic file writes + serialized-path normalization.\n//!\n//! Two KNOWN-HAZARDS invariants live here:\n//!\n//! - **5.2 Windows EBUSY on atomic rename** (REQ-HAZARD-EBUSY-RENAME): a\n//!   `fs::rename` can fail transiently while another process (AV scanner,\n//!   indexer, a just-closed handle) still holds the target. The sister's\n//!   `atomic_write_string` was a bare tmp-write + rename; spt-core adds\n//!   retry-with-backoff that tolerates transient sharing/access errors and\n//!   only surfaces a persistent failure.\n//! - **5.4 Windows UNC prefix in serialized paths** (REQ-HAZARD-UNC-PATH-STRIP):\n//!   a canonicalized `\\\\?\\C:\\…` path serializes to `//?/C:/…` and then fails\n//!   `read_to_string` on the consumer side. [`to_forward_slash`] strips the\n//!   `\\\\?\\` prefix after the backslash→slash conversion so serialized path\n//!   attributes are directly consumable. Copied verbatim from sister\n//!   `owlery.rs:377-384`.\n//!\n//! The write pattern itself (sibling `{name}.tmp` + rename) is verbatim; same-\n//! volume rename is atomic on both NTFS and POSIX, so a concurrent reader never\n//! observes a half-written file.\n\nuse std::fs;\n…\nuse std::time::Duration;\n\n/// Maximum rename attempts before surfacing the error (KNOWN-HAZARDS 5.2).\nconst MAX_RENAME_ATTEMPTS: u32 = 10;\n\n/// Process-local counter making each write's tmp sibling unique (KNOWN-HAZARDS\n/// 5.15). NO clock / NO rand — a monotonic per-process seq plus the pid is enough\n/// to distinguish every concurrent writer to the same target, and stays\n/// resume/replay-safe.\nstatic TMP_SEQ: AtomicU64 = AtomicU64::new(0);\n\n/// Convert a path to a forward-slash string for JSON / cross-platform output,\n/// stripping the Windows UNC `\\\\?\\` prefix (which becomes `//?/` after the\n/// slash conversion). Serialized path attributes must be directly consumable\n/// by a `read_to_string` on the other side. Copied verbatim (sister\n/// `owlery::to_forward_slash`).\n// [impl->REQ-HAZARD-UNC-PATH-STRIP]\npub fn to_forward_slash(path: &Path) -> String { … }\n\n/// Atomically write `body` to `path` via a sibling `{name}.tmp` + rename, with\n/// EBUSY retry/backoff (KNOWN-HAZARDS 5.2). A concurrent reader sees either the\n/// old file or the complete new file — never a torn write. **Non-durable** — see\n/// [`atomic_write_bytes`] for the durability trade-off.\n// [impl->REQ-HAZARD-EBUSY-RENAME]\npub fn atomic_write_string(path: &Path, body: &str) -> io::Result<()> {\n    atomic_write_bytes(path, body.as_bytes())\n}\n\n/// Byte-oriented sibling of [`atomic_write_string`] — for binary payloads\n/// (blobs, signature material) that are not valid UTF-8.\n///\n/// **Non-durable** (no fsync): a plain `fs::write` of the tmp + rename. Still\n/// reader-atomic (old-or-new, never torn), but a hard reset between the rename\n/// and the data flush can resurrect the file at full length filled with NUL. That\n/// is the correct default for RECONSTRUCTIBLE state — registry snapshots, peer\n/// caches, the release cache — where a post-crash NUL husk simply parse-fails to\n/// *absent* and is regenerated / re-gossiped (and the 5.14 read-side makes a\n/// corrupt record harmless at every seam). For AUTHORITATIVE records whose loss\n/// IS an incident, use [`atomic_write_bytes_durable`].\n///\n/// **The epoch counter was named on that list and no longer is** (KNOWN-HAZARDS\n/// 7.60, WAN-TRUTH #40): it is NOT reconstructible, because nothing regenerates\n/// it past its historical maximum. Degrading a husk to `0` hands the node a lease\n/// BELOW the one its own fleet-replicated roster row carries, and the\n/// strictly-greater merge rule then rejects every self-advertisement it will ever\n/// make — a silent, permanent, fleet-wide freeze. That counter writes durably now\n/// ([`crate::epoch::EpochSource::next_epoch`]); the fleet-lease fast-forward is\n/// the primary repair and this is defense in depth.\n// [impl->REQ-HAZARD-EBUSY-RENAME]\npub fn atomic_write_bytes(path: &Path, body: &[u8]) -> io::Result<()> {\n    write_then_rename(path, body, false)\n}\n\n/// **Durable** sibling of [`atomic_write_bytes`]: the tmp file is fsync'd\n/// (`sync_all`) BEFORE the rename publishes the name (KNOWN-HAZARDS 5.13), so a\n/// hard reset yields the complete old OR new file — never a full-length NUL\n/// zero-fill (NTFS keeps the rename metadata but loses un-synced data blocks).\n///\n/// **Scoped, NOT a blanket default.** Durability is opt-in for the few records\n/// whose loss is unrecoverable — the authoritative per-perch record\n/// ([`crate::info::write_info`]) and one-shot identity material\n/// ([`crate::nodeid`], `spt_daemon::machineid`). Making *every* atomic write\n/// durable serializes an fsync behind each of **51** call sites — the\n/// production (non-`#[cfg(test)]`) callers of the two NON-durable entry points,\n/// [`atomic_write_string`] and [`atomic_write_bytes`], across 33 files,\n/// re-counted 2026-08-18 (the figure had read ~26 since it was written, and a\n/// census with no stated population cannot be re-derived; this one can) — and\n/// for `info.json`\n/// the flush lands under the per-perch `.info.lock` (the W1b lock-across-fsync\n/// wedge shape) — which stalls daemon bringup 4–6× and misses the endpoint\n/// ONLINE window (caught by `attach_wedge_e2e`). Scope durability to the record,\n/// not the writer.\n// [impl->REQ-HAZARD-EBUSY-RENAME]\n// [impl->REQ-HAZARD-PERCH-RECORD-POWER-LOSS]\npub fn atomic_write_bytes_durable(path: &Path, body: &[u8]) -> io::Result<()> {\n    write_then_rename(path, body, true)\n}\n\n/// String sibling of [`atomic_write_bytes_durable`] (durable tmp+rename).\n// [impl->REQ-HAZARD-PERCH-RECORD-POWER-LOSS]\npub fn atomic_write_string_durable(path: &Path, body: &str) -> io::Result<()> {\n    atomic_write_bytes_durable(path, body.as_bytes())\n}\n\n/// Shared tmp-write + atomic-rename core. `durable` selects the fsync-before-rename\n/// path ([`atomic_write_bytes_durable`]); otherwise a plain `fs::write` (the\n/// reconstructible-state default, [`atomic_write_bytes`]).\n///\n/// The tmp sibling is UNIQUE per write — `{name}.tmp.{pid}-{seq}` (KNOWN-HAZARDS\n/// 5.15). A fixed `{name}.tmp` shared by every writer is a latent collision: two\n/// processes writing the same target concurrently (e.g. the daemon's `mutate_info`\n/// RMW and `spt api bind`'s establish, traced ~700µs apart at bind) both stage the\n/// SAME tmp; whichever renames first CONSUMES it, and the loser's rename hits\n/// `NotFound` (os error 2) — non-transient, so it surfaces as a hard write error\n/// (the bind-fatal path that wedged the endpoint offline). A microsecond window\n/// forever; `sync_all` widened create→rename ~50× and made it near-certain at bind.\n// [impl->REQ-HAZARD-ATOMIC-TMP-COLLISION]\nfn write_then_rename(path: &Path, body: &[u8], durable: bool) -> io::Result<()> { … }\n\n/// Rename `from` → `to`, retrying transient sharing/access failures with\n/// linear backoff. Non-transient errors (e.g. NotFound) return immediately.\n// [impl->REQ-HAZARD-EBUSY-RENAME]\nfn rename_with_retry(from: &Path, to: &Path) -> io::Result<()> {\n    retry_transient(MAX_RENAME_ATTEMPTS, backoff, || fs::rename(from, to))\n}\n\n/// Linear backoff: attempt 1 → 2ms, attempt 2 → 4ms, … capped modestly. Kept\n/// small so a real transient (handle released within tens of ms) is absorbed\n/// without a perceptible stall.\nfn backoff(attempt: u32) -> Duration {\n    Duration::from_millis((attempt as u64 * 2).min(50))\n}\n\n/// Run `op`, retrying while it fails with a *transient* error, up to\n/// `max_attempts` total tries. Sleeps `delay(attempt)` between tries. Returns\n/// the first success, or the last error once attempts are exhausted /\n/// the error is non-transient. Generic over the operation so the retry\n/// semantics are unit-testable without provoking a real OS EBUSY.\n// [impl->REQ-HAZARD-EBUSY-RENAME]\nfn retry_transient<T>(\n    max_attempts: u32,\n    delay: impl Fn(u32) -> Duration,\n    mut op: impl FnMut() -> io::Result<T>,\n) -> io::Result<T> { … }\n\n/// Whether an io error is a transient sharing/access failure worth retrying.\n///\n/// Windows surfaces a held-handle rename collision as `ERROR_SHARING_VIOLATION`\n/// (32) or `ERROR_ACCESS_DENIED` (5). On other platforms `PermissionDenied`\n/// is the analogue; everything else (NotFound, etc.) is non-transient.\nfn is_transient(e: &io::Error) -> bool { … }\n\n#[cfg(test)]\nmod tests { … }\n\n[…231ln elided; re-read needed ranges, e.g. .worktrees/287-shellrestore/crates/spt-store/src/atomic.rs:23-26,44-51]"}],"details":{"displayContent":{"text":"//! Atomic file writes + serialized-path normalization.\n//!\n//! Two KNOWN-HAZARDS invariants live here:\n//!\n//! - **5.2 Windows EBUSY on atomic rename** (REQ-HAZARD-EBUSY-RENAME): a\n//!   `fs::rename` can fail transiently while another process (AV scanner,\n//!   indexer, a just-closed handle) still holds the target. The sister's\n//!   `atomic_write_string` was a bare tmp-write + rename; spt-core adds\n//!   retry-with-backoff that tolerates transient sharing/access errors and\n//!   only surfaces a persistent failure.\n//! - **5.4 Windows UNC prefix in serialized paths** (REQ-HAZARD-UNC-PATH-STRIP):\n//!   a canonicalized `\\\\?\\C:\\…` path serializes to `//?/C:/…` and then fails\n//!   `read_to_string` on the consumer side. [`to_forward_slash`] strips the\n//!   `\\\\?\\` prefix after the backslash→slash conversion so serialized path\n//!   attributes are directly consumable. Copied verbatim from sister\n//!   `owlery.rs:377-384`.\n//!\n//! The write pattern itself (sibling `{name}.tmp` + rename) is verbatim; same-\n//! volume rename is atomic on both NTFS and POSIX, so a concurrent reader never\n//! observes a half-written file.\n\nuse std::fs;\n…\nuse std::time::Duration;\n\n/// Maximum rename attempts before surfacing the error (KNOWN-HAZARDS 5.2).\nconst MAX_RENAME_ATTEMPTS: u32 = 10;\n\n/// Process-local counter making each write's tmp sibling unique (KNOWN-HAZARDS\n/// 5.15). NO clock / NO rand — a monotonic per-process seq plus the pid is enough\n/// to distinguish every concurrent writer to the same target, and stays\n/// resume/replay-safe.\nstatic TMP_SEQ: AtomicU64 = AtomicU64::new(0);\n\n/// Convert a path to a forward-slash string for JSON / cross-platform output,\n/// stripping the Windows UNC `\\\\?\\` prefix (which becomes `//?/` after the\n/// slash conversion). Serialized path attributes must be directly consumable\n/// by a `read_to_string` on the other side. Copied verbatim (sister\n/// `owlery::to_forward_slash`).\n// [impl->REQ-HAZARD-UNC-PATH-STRIP]\npub fn to_forward_slash(path: &Path) -> String { … }\n\n/// Atomically write `body` to `path` via a sibling `{name}.tmp` + rename, with\n/// EBUSY retry/backoff (KNOWN-HAZARDS 5.2). A concurrent reader sees either the\n/// old file or the complete new file — never a torn write. **Non-durable** — see\n/// [`atomic_write_bytes`] for the durability trade-off.\n// [impl->REQ-HAZARD-EBUSY-RENAME]\npub fn atomic_write_string(path: &Path, body: &str) -> io::Result<()> {\n    atomic_write_bytes(path, body.as_bytes())\n}\n\n/// Byte-oriented sibling of [`atomic_write_string`] — for binary payloads\n/// (blobs, signature material) that are not valid UTF-8.\n///\n/// **Non-durable** (no fsync): a plain `fs::write` of the tmp + rename. Still\n/// reader-atomic (old-or-new, never torn), but a hard reset between the rename\n/// and the data flush can resurrect the file at full length filled with NUL. That\n/// is the correct default for RECONSTRUCTIBLE state — registry snapshots, peer\n/// caches, the release cache — where a post-crash NUL husk simply parse-fails to\n/// *absent* and is regenerated / re-gossiped (and the 5.14 read-side makes a\n/// corrupt record harmless at every seam). For AUTHORITATIVE records whose loss\n/// IS an incident, use [`atomic_write_bytes_durable`].\n///\n/// **The epoch counter was named on that list and no longer is** (KNOWN-HAZARDS\n/// 7.60, WAN-TRUTH #40): it is NOT reconstructible, because nothing regenerates\n/// it past its historical maximum. Degrading a husk to `0` hands the node a lease\n/// BELOW the one its own fleet-replicated roster row carries, and the\n/// strictly-greater merge rule then rejects every self-advertisement it will ever\n/// make — a silent, permanent, fleet-wide freeze. That counter writes durably now\n/// ([`crate::epoch::EpochSource::next_epoch`]); the fleet-lease fast-forward is\n/// the primary repair and this is defense in depth.\n// [impl->REQ-HAZARD-EBUSY-RENAME]\npub fn atomic_write_bytes(path: &Path, body: &[u8]) -> io::Result<()> {\n    write_then_rename(path, body, false)\n}\n\n/// **Durable** sibling of [`atomic_write_bytes`]: the tmp file is fsync'd\n/// (`sync_all`) BEFORE the rename publishes the name (KNOWN-HAZARDS 5.13), so a\n/// hard reset yields the complete old OR new file — never a full-length NUL\n/// zero-fill (NTFS keeps the rename metadata but loses un-synced data blocks).\n///\n/// **Scoped, NOT a blanket default.** Durability is opt-in for the few records\n/// whose loss is unrecoverable — the authoritative per-perch record\n/// ([`crate::info::write_info`]) and one-shot identity material\n/// ([`crate::nodeid`], `spt_daemon::machineid`). Making *every* atomic write\n/// durable serializes an fsync behind each of **51** call sites — the\n/// production (non-`#[cfg(test)]`) callers of the two NON-durable entry points,\n/// [`atomic_write_string`] and [`atomic_write_bytes`], across 33 files,\n/// re-counted 2026-08-18 (the figure had read ~26 since it was written, and a\n/// census with no stated population cannot be re-derived; this one can) — and\n/// for `info.json`\n/// the flush lands under the per-perch `.info.lock` (the W1b lock-across-fsync\n/// wedge shape) — which stalls daemon bringup 4–6× and misses the endpoint\n/// ONLINE window (caught by `attach_wedge_e2e`). Scope durability to the record,\n/// not the writer.\n// [impl->REQ-HAZARD-EBUSY-RENAME]\n// [impl->REQ-HAZARD-PERCH-RECORD-POWER-LOSS]\npub fn atomic_write_bytes_durable(path: &Path, body: &[u8]) -> io::Result<()> {\n    write_then_rename(path, body, true)\n}\n\n/// String sibling of [`atomic_write_bytes_durable`] (durable tmp+rename).\n// [impl->REQ-HAZARD-PERCH-RECORD-POWER-LOSS]\npub fn atomic_write_string_durable(path: &Path, body: &str) -> io::Result<()> {\n    atomic_write_bytes_durable(path, body.as_bytes())\n}\n\n/// Shared tmp-write + atomic-rename core. `durable` selects the fsync-before-rename\n/// path ([`atomic_write_bytes_durable`]); otherwise a plain `fs::write` (the\n/// reconstructible-state default, [`atomic_write_bytes`]).\n///\n/// The tmp sibling is UNIQUE per write — `{name}.tmp.{pid}-{seq}` (KNOWN-HAZARDS\n/// 5.15). A fixed `{name}.tmp` shared by every writer is a latent collision: two\n/// processes writing the same target concurrently (e.g. the daemon's `mutate_info`\n/// RMW and `spt api bind`'s establish, traced ~700µs apart at bind) both stage the\n/// SAME tmp; whichever renames first CONSUMES it, and the loser's rename hits\n/// `NotFound` (os error 2) — non-transient, so it surfaces as a hard write error\n/// (the bind-fatal path that wedged the endpoint offline). A microsecond window\n/// forever; `sync_all` widened create→rename ~50× and made it near-certain at bind.\n// [impl->REQ-HAZARD-ATOMIC-TMP-COLLISION]\nfn write_then_rename(path: &Path, body: &[u8], durable: bool) -> io::Result<()> { … }\n\n/// Rename `from` → `to`, retrying transient sharing/access failures with\n/// linear backoff. Non-transient errors (e.g. NotFound) return immediately.\n// [impl->REQ-HAZARD-EBUSY-RENAME]\nfn rename_with_retry(from: &Path, to: &Path) -> io::Result<()> {\n    retry_transient(MAX_RENAME_ATTEMPTS, backoff, || fs::rename(from, to))\n}\n\n/// Linear backoff: attempt 1 → 2ms, attempt 2 → 4ms, … capped modestly. Kept\n/// small so a real transient (handle released within tens of ms) is absorbed\n/// without a perceptible stall.\nfn backoff(attempt: u32) -> Duration {\n    Duration::from_millis((attempt as u64 * 2).min(50))\n}\n\n/// Run `op`, retrying while it fails with a *transient* error, up to\n/// `max_attempts` total tries. Sleeps `delay(attempt)` between tries. Returns\n/// the first success, or the last error once attempts are exhausted /\n/// the error is non-transient. Generic over the operation so the retry\n/// semantics are unit-testable without provoking a real OS EBUSY.\n// [impl->REQ-HAZARD-EBUSY-RENAME]\nfn retry_transient<T>(\n    max_attempts: u32,\n    delay: impl Fn(u32) -> Duration,\n    mut op: impl FnMut() -> io::Result<T>,\n) -> io::Result<T> { … }\n\n/// Whether an io error is a transient sharing/access failure worth retrying.\n///\n/// Windows surfaces a held-handle rename collision as `ERROR_SHARING_VIOLATION`\n/// (32) or `ERROR_ACCESS_DENIED` (5). On other platforms `PermissionDenied`\n/// is the analogue; everything else (NotFound, etc.) is non-transient.\nfn is_transient(e: &io::Error) -> bool { … }\n\n#[cfg(test)]\nmod tests { … }","startLine":1},"summary":{"lines":157,"elidedSpans":6,"elidedLines":231},"fileSize":18102,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\crates\\spt-store\\src\\atomic.rs"}}},"isError":false,"timestamp":1789001680440}}
{"type":"message","id":"d2537f3c","parentId":"828f7fdb","timestamp":"2026-09-10T00:54:40.447Z","message":{"role":"toolResult","toolCallId":"call_46wCDcUuLMgmptx4WXo6Xtfz|fc_028129a7f1d66e9c016aa1ffd0441487d2b0b8038d5ae4d6fa","toolName":"read","content":[{"type":"text","text":"diff --git a/crates/spt-daemon/src/shellhost.rs b/crates/spt-daemon/src/shellhost.rs\nindex 44df752b..23833ed6 100644\n--- a/crates/spt-daemon/src/shellhost.rs\n+++ b/crates/spt-daemon/src/shellhost.rs\n@@ -496,13 +496,10 @@ pub fn launch_shell(\n                 .map_err(|e| format!(\"spawn {program}: {e}\"))?\n                 .id();\n             let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n-            std::fs::write(perch.join(SHELL_PID_FILE), pid.to_string())\n-                .map_err(|e| format!(\"record pid: {e}\"))?;\n-            // The birth stamp rides WITH the pid at every launch, or the probe that\n-            // reads it degrades to bare-pid liveness on this instance alone -- a\n-            // silent per-instance regression rather than a visible one.\n+            // Retire the previous birth before publishing the new PID.\n             // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n-            spt_store::shellinfo::record_shell_launch(&perch, pid, now_ms());\n+            spt_store::shellinfo::publish_shell_launch(&perch, pid, now_ms())\n+                .map_err(|e| format!(\"record launch: {e}\"))?;\n             Ok(pid)\n         }\n     }\n@@ -555,13 +552,10 @@ pub fn launch_shell_brokered_in(\n         .map_err(|e| format!(\"broker spawn: {e}\"))?;\n     let pid = pid.unwrap_or(0);\n     let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n-    std::fs::write(perch.join(SHELL_PID_FILE), pid.to_string())\n-        .map_err(|e| format!(\"record pid: {e}\"))?;\n-    // Same stamp at the broker-hosted launch. A backend that exposed no pid\n-    // records 0 here; record_shell_launch stamps no start-time for it, so the\n-    // probe keeps its fail-toward-alive stance for pid-less backends.\n+    // PID-less backends keep their fail-toward-alive behavior.\n     // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n-    spt_store::shellinfo::record_shell_launch(&perch, pid, now_ms());\n+    spt_store::shellinfo::publish_shell_launch(&perch, pid, now_ms())\n+        .map_err(|e| format!(\"record launch: {e}\"))?;\n     Ok(pid)\n }\n \ndiff --git a/crates/spt-daemon/src/shellwake.rs b/crates/spt-daemon/src/shellwake.rs\nindex d9631b4d..7a473cf5 100644\n--- a/crates/spt-daemon/src/shellwake.rs\n+++ b/crates/spt-daemon/src/shellwake.rs\n@@ -406,13 +406,9 @@ pub fn resolve_wake(\n     // Never double-launch a live binary: the revive's wake cascade may have\n     // relaunched this (persistent) instance already — and two reconcilers\n     // (a fresh daemon adopting + a stale watcher firing) may race a wake.\n-    // A parked pid that probes alive means somebody won; stand down.\n+    // A live, birth-matching launch wins even before its bind handshake.\n     let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n-    let live_pid = std::fs::read_to_string(perch.join(crate::shellhost::SHELL_PID_FILE))\n-        .ok()\n-        .and_then(|s| s.trim().parse::<u32>().ok())\n-        .filter(|&p| p != 0 && spt_store::proc::is_process_alive(p));\n-    if let Some(p) = live_pid {\n+    if let Some(p) = crate::shellhost::live_launch_winner(&perch) {\n         return Ok(format!(\"{did}already relaunched pid={p} (online at bind)\"));\n     }\n     // [impl->REQ-INSTALL-11] the wake-triggered relaunch resolves against the\ndiff --git a/crates/spt-store/src/shellinfo.rs b/crates/spt-store/src/shellinfo.rs\nindex f3aa9086..c0988a9e 100644\n--- a/crates/spt-store/src/shellinfo.rs\n+++ b/crates/spt-store/src/shellinfo.rs\n@@ -172,6 +172,22 @@ pub fn record_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) {\n     }\n }\n \n+/// Publish a new PID without ever pairing it with the previous launch's birth.\n+/// During publication a missing stamp deliberately preserves live-PID custody.\n+/// Failure to retire an old stamp must stop publication rather than misidentify\n+/// the new binary as a recycled corpse.\n+// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n+pub fn publish_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) -> std::io::Result<()> {\n+    match std::fs::remove_file(shell_perch.join(SHELL_LAUNCH_FILE)) {\n+        Ok(()) => {}\n+        Err(e) if e.kind() == std::io::ErrorKind::NotFound => {}\n+        Err(e) => return Err(e),\n+    }\n+    atomic_write_string(&shell_perch.join(SHELL_PID_FILE), &pid.to_string())?;\n+    record_shell_launch(shell_perch, pid, now_ms);\n+    Ok(())\n+}\n+\n /// Read the birth stamp, if this instance was launched by a build that parks one.\n /// Absent or unparseable ⇒ `None`, which every consumer treats as \"no stamp to\n /// compare\" rather than as a fault.\n@@ -212,6 +228,11 @@ pub fn shell_pid_provably_dead(shell_perch: &Path) -> bool {\n     // time, still falls back to plain absence and fails toward ALIVE.\n     // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n     let birth = read_shell_launch(shell_perch).and_then(|l| l.pid_started_at);\n+    // A relaunch can publish between these reads. Do not apply the new birth\n+    // to the old PID; an unstable snapshot is not proof that the shell died.\n+    if read_shell_pid(shell_perch) != Some(pid) {\n+        return false;\n+    }\n     matches!(\n         crate::liveness::relay_liveness(Some(pid), birth),\n         crate::liveness::RelayLiveness::Gone\n@@ -504,6 +525,16 @@ mod tests {\n         assert_eq!(status_of(&d), SHELL_STATUS_ONLINE);\n     }\n \n+    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n+    #[test]\n+    fn launch_publication_refuses_an_unretirable_old_stamp() {\n+        let d = online_perch(Some(\"123\"));\n+        std::fs::create_dir(d.path().join(SHELL_LAUNCH_FILE)).unwrap();\n+        let err = publish_shell_launch(d.path(), std::process::id(), 1_000).unwrap_err();\n+        assert_ne!(err.kind(), std::io::ErrorKind::NotFound);\n+        assert_eq!(read_shell_pid(d.path()), Some(123));\n+    }\n+\n     // [unit->REQ-HAZARD-SHELL-STALE-ONLINE] the whole discriminant matrix. The\n     // ONE case that derives offline is a parked, non-zero, dead pid — flynn's\n     // field shape: the binary was force-killed, no link broke, so `close_shell`\ndiff --git a/docs/KNOWN-HAZARDS.md b/docs/KNOWN-HAZARDS.md\nindex c8b94c6e..46dc56e9 100644\n--- a/docs/KNOWN-HAZARDS.md\n+++ b/docs/KNOWN-HAZARDS.md\n@@ -1016,9 +1016,9 @@ The kill-path rule above generalizes: `daemon.pid` is not authority for *\"which\n | 2.1/5.1 | Stable PID/broker-handle over ephemeral PID | liveness detection |\n | 2.3 | Handoff argv/IPC version-tolerant (newer brain ↔ older broker) | broker↔brain IPC, self-update |\n | 2.4 | gen_start = now() on cold-start + handoff | per-instance generation |\n-| 2.6 | A shell's ONLINE-ness is DERIVED (recorded status AND a not-provably-dead `shell.pid`) — an abruptly-killed binary breaks no link, so `close_shell`'s offline flip never runs and the record lies forever; derive at the gates/renders, keep the recorded field at the writers, the suspend-cascade close, and the wake reconciler — whose \"no spontaneous relaunch\" (deploys and quarantine own that decision) is since W2 the stated corpse-predates-boot predicate, not the stale record holding it open, and whose recycled-pid blindness is closed by the launch birth stamp wherever one is parked. Not only a crash edge: the GRACEFUL daemon-stop path kills bound shells at stop-begin with `close_shell` unrun (field 2026-07-25 — shells died abruptly, `info.json` still `online`, while the broker drained on 2m10s), so every daemon restart manufactures these stale records routinely; the derivation heals them at the gates | `spt_store::shellinfo::{shell_pid_provably_dead,effective_status,is_shell_online}`, `linkhost` relink/drive/cmd-wake, `shelldisc::discover`, `activity::observe_links` |\n+| 2.6 | A shell's ONLINE-ness is DERIVED from recorded status and a not-provably-dead PID/birth pair. Abrupt death and daemon-stop leave stale `online` records; derive truth at gates/renders and heal each changed record once. The same-boot no-spontaneous-relaunch rule applies only to nonpersistent shells. Persistent instances follow 2.7 regardless of launch age. During new-launch publication, retire the previous birth before publishing the new PID; an unstable read is not proof of death | `spt_store::shellinfo::{shell_pid_provably_dead,effective_status,is_shell_online,publish_shell_launch}`, `linkhost`, `shelldisc`, `activity` |\n <!-- [doc->REQ-HAZARD-SHELL-STALE-ONLINE] -->\n-| 2.7 | A node restart must not permanently strand every `persistent` shell: (a) the daemon heals a recorded `online` the derivation contradicts, write-guarded on an actual change, and (b) a once-per-generation boot sweep PLUS an owner offline->online edge in the reconcile loop (releases#228 — the sweep alone strands the shells of any owner not yet online when it ran) relaunch an instance only when ALL of — adapter `persistent`, owner endpoint online, down in fact, recorded launch predates the boot instant (slack-absorbed) — hold; no stamp or no boot oracle ⇒ not restored. The owner-facing surfaces derive and look clean while the cascade reads the recorded field, so a clean render is NOT evidence the class did not occur — the falsifier is the on-disk record. The pre-existing cascade test's setup suppresses the failing arm (it suspends first), so the restart shape must be asserted on a fixture that never calls the suspend path | `shellwake::{heal_stale_online_records,restore_persistent_shells_at_boot,restore_persistent_shells_on_owner_online,OwnerOnlineEdge,launch_predates_boot}`, `proc::boot_instant_ms`, `shellinfo::{record_shell_launch,read_shell_launch}` |\n+| 2.7 | A once-per-generation boot sweep and each owner offline->online edge restore every instance whose adapter is `persistent`, owner is online, and binary is down in fact. Machine boot time, launch age, and absent launch stamps do not block persistent restoration; already-online and live binding launches are excluded. Heal every stale instance per owner with no repeat writes. Prove the restart shape without first suspending the fixture; a clean derived display alone is not evidence the record was healed | `shellwake::{heal_stale_online_records,restore_persistent_shells_at_boot,restore_persistent_shells_on_owner_online,OwnerOnlineEdge}`, `shellhost::live_launch_winner` |\n | 3.1 | Ephemeral perch cleanup on all exit paths | `ring` (RAII guard) |\n | 3.4 | A ring never adopts (so never deletes) a perch it did not create — probe the DIR, not the ready marker; record/spool/unreadable = occupied, empty = refused too; deliver + loud `RING_PERCH_EXISTS`/`RING_STALE_DIR` instead of block-waiting | `spt_msg::ring` probe + `create_dir` leaf |\n | 3.5 | A perch GC classifies on record PRESENCE only and never asks a liveness resolver — `is_perch_alive` is INVERTED here (recordless residue reads ALIVE, an offline endpoint reads DEAD: 24 vs 6 measured on HFENDULEAM 2026-08-04), so `!is_perch_alive` authorizes exactly the inverse set; the registry is never asked either (offline endpoints are legitimately absent from it) | `spt_store::perchgc::sweep`, `spt endpoint gc` |"}],"details":{"resolvedPath":"C:\\Users\\decid\\.omp\\agent\\sessions\\-Documents-projects-spt-core\\2026-09-09T23-45-42-317Z_01a08890-21ad-7000-8f54-208d4f4f2f2c\\128.bash-original.log","contentType":"text/plain","totalLines":130,"displayContent":{"text":"diff --git a/crates/spt-daemon/src/shellhost.rs b/crates/spt-daemon/src/shellhost.rs\nindex 44df752b..23833ed6 100644\n--- a/crates/spt-daemon/src/shellhost.rs\n+++ b/crates/spt-daemon/src/shellhost.rs\n@@ -496,13 +496,10 @@ pub fn launch_shell(\n                 .map_err(|e| format!(\"spawn {program}: {e}\"))?\n                 .id();\n             let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n-            std::fs::write(perch.join(SHELL_PID_FILE), pid.to_string())\n-                .map_err(|e| format!(\"record pid: {e}\"))?;\n-            // The birth stamp rides WITH the pid at every launch, or the probe that\n-            // reads it degrades to bare-pid liveness on this instance alone -- a\n-            // silent per-instance regression rather than a visible one.\n+            // Retire the previous birth before publishing the new PID.\n             // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n-            spt_store::shellinfo::record_shell_launch(&perch, pid, now_ms());\n+            spt_store::shellinfo::publish_shell_launch(&perch, pid, now_ms())\n+                .map_err(|e| format!(\"record launch: {e}\"))?;\n             Ok(pid)\n         }\n     }\n@@ -555,13 +552,10 @@ pub fn launch_shell_brokered_in(\n         .map_err(|e| format!(\"broker spawn: {e}\"))?;\n     let pid = pid.unwrap_or(0);\n     let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n-    std::fs::write(perch.join(SHELL_PID_FILE), pid.to_string())\n-        .map_err(|e| format!(\"record pid: {e}\"))?;\n-    // Same stamp at the broker-hosted launch. A backend that exposed no pid\n-    // records 0 here; record_shell_launch stamps no start-time for it, so the\n-    // probe keeps its fail-toward-alive stance for pid-less backends.\n+    // PID-less backends keep their fail-toward-alive behavior.\n     // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n-    spt_store::shellinfo::record_shell_launch(&perch, pid, now_ms());\n+    spt_store::shellinfo::publish_shell_launch(&perch, pid, now_ms())\n+        .map_err(|e| format!(\"record launch: {e}\"))?;\n     Ok(pid)\n }\n \ndiff --git a/crates/spt-daemon/src/shellwake.rs b/crates/spt-daemon/src/shellwake.rs\nindex d9631b4d..7a473cf5 100644\n--- a/crates/spt-daemon/src/shellwake.rs\n+++ b/crates/spt-daemon/src/shellwake.rs\n@@ -406,13 +406,9 @@ pub fn resolve_wake(\n     // Never double-launch a live binary: the revive's wake cascade may have\n     // relaunched this (persistent) instance already — and two reconcilers\n     // (a fresh daemon adopting + a stale watcher firing) may race a wake.\n-    // A parked pid that probes alive means somebody won; stand down.\n+    // A live, birth-matching launch wins even before its bind handshake.\n     let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n-    let live_pid = std::fs::read_to_string(perch.join(crate::shellhost::SHELL_PID_FILE))\n-        .ok()\n-        .and_then(|s| s.trim().parse::<u32>().ok())\n-        .filter(|&p| p != 0 && spt_store::proc::is_process_alive(p));\n-    if let Some(p) = live_pid {\n+    if let Some(p) = crate::shellhost::live_launch_winner(&perch) {\n         return Ok(format!(\"{did}already relaunched pid={p} (online at bind)\"));\n     }\n     // [impl->REQ-INSTALL-11] the wake-triggered relaunch resolves against the\ndiff --git a/crates/spt-store/src/shellinfo.rs b/crates/spt-store/src/shellinfo.rs\nindex f3aa9086..c0988a9e 100644\n--- a/crates/spt-store/src/shellinfo.rs\n+++ b/crates/spt-store/src/shellinfo.rs\n@@ -172,6 +172,22 @@ pub fn record_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) {\n     }\n }\n \n+/// Publish a new PID without ever pairing it with the previous launch's birth.\n+/// During publication a missing stamp deliberately preserves live-PID custody.\n+/// Failure to retire an old stamp must stop publication rather than misidentify\n+/// the new binary as a recycled corpse.\n+// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n+pub fn publish_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) -> std::io::Result<()> {\n+    match std::fs::remove_file(shell_perch.join(SHELL_LAUNCH_FILE)) {\n+        Ok(()) => {}\n+        Err(e) if e.kind() == std::io::ErrorKind::NotFound => {}\n+        Err(e) => return Err(e),\n+    }\n+    atomic_write_string(&shell_perch.join(SHELL_PID_FILE), &pid.to_string())?;\n+    record_shell_launch(shell_perch, pid, now_ms);\n+    Ok(())\n+}\n+\n /// Read the birth stamp, if this instance was launched by a build that parks one.\n /// Absent or unparseable ⇒ `None`, which every consumer treats as \"no stamp to\n /// compare\" rather than as a fault.\n@@ -212,6 +228,11 @@ pub fn shell_pid_provably_dead(shell_perch: &Path) -> bool {\n     // time, still falls back to plain absence and fails toward ALIVE.\n     // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n     let birth = read_shell_launch(shell_perch).and_then(|l| l.pid_started_at);\n+    // A relaunch can publish between these reads. Do not apply the new birth\n+    // to the old PID; an unstable snapshot is not proof that the shell died.\n+    if read_shell_pid(shell_perch) != Some(pid) {\n+        return false;\n+    }\n     matches!(\n         crate::liveness::relay_liveness(Some(pid), birth),\n         crate::liveness::RelayLiveness::Gone\n@@ -504,6 +525,16 @@ mod tests {\n         assert_eq!(status_of(&d), SHELL_STATUS_ONLINE);\n     }\n \n+    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n+    #[test]\n+    fn launch_publication_refuses_an_unretirable_old_stamp() {\n+        let d = online_perch(Some(\"123\"));\n+        std::fs::create_dir(d.path().join(SHELL_LAUNCH_FILE)).unwrap();\n+        let err = publish_shell_launch(d.path(), std::process::id(), 1_000).unwrap_err();\n+        assert_ne!(err.kind(), std::io::ErrorKind::NotFound);\n+        assert_eq!(read_shell_pid(d.path()), Some(123));\n+    }\n+\n     // [unit->REQ-HAZARD-SHELL-STALE-ONLINE] the whole discriminant matrix. The\n     // ONE case that derives offline is a parked, non-zero, dead pid — flynn's\n     // field shape: the binary was force-killed, no link broke, so `close_shell`\ndiff --git a/docs/KNOWN-HAZARDS.md b/docs/KNOWN-HAZARDS.md\nindex c8b94c6e..46dc56e9 100644\n--- a/docs/KNOWN-HAZARDS.md\n+++ b/docs/KNOWN-HAZARDS.md\n@@ -1016,9 +1016,9 @@ The kill-path rule above generalizes: `daemon.pid` is not authority for *\"which\n | 2.1/5.1 | Stable PID/broker-handle over ephemeral PID | liveness detection |\n | 2.3 | Handoff argv/IPC version-tolerant (newer brain ↔ older broker) | broker↔brain IPC, self-update |\n | 2.4 | gen_start = now() on cold-start + handoff | per-instance generation |\n-| 2.6 | A shell's ONLINE-ness is DERIVED (recorded status AND a not-provably-dead `shell.pid`) — an abruptly-killed binary breaks no link, so `close_shell`'s offline flip never runs and the record lies forever; derive at the gates/renders, keep the recorded field at the writers, the suspend-cascade close, and the wake reconciler — whose \"no spontaneous relaunch\" (deploys and quarantine own that decision) is since W2 the stated corpse-predates-boot predicate, not the stale record holding it open, and whose recycled-pid blindness is closed by the launch birth stamp wherever one is parked. Not only a crash edge: the GRACEFUL daemon-stop path kills bound shells at stop-begin with `close_shell` unrun (field 2026-07-25 — shells died abruptly, `info.json` still `online`, while the broker drained on 2m10s), so every daemon restart manufactures these stale records routinely; the derivation heals them at the gates | `spt_store::shellinfo::{shell_pid_provably_dead,effective_status,is_shell_online}`, `linkhost` relink/drive/cmd-wake, `shelldisc::discover`, `activity::observe_links` |\n+| 2.6 | A shell's ONLINE-ness is DERIVED from recorded status and a not-provably-dead PID/birth pair. Abrupt death and daemon-stop leave stale `online` records; derive truth at gates/renders and heal each changed record once. The same-boot no-spontaneous-relaunch rule applies only to nonpersistent shells. Persistent instances follow 2.7 regardless of launch age. During new-launch publication, retire the previous birth before publishing the new PID; an unstable read is not proof of death | `spt_store::shellinfo::{shell_pid_provably_dead,effective_status,is_shell_online,publish_shell_launch}`, `linkhost`, `shelldisc`, `activity` |\n <!-- [doc->REQ-HAZARD-SHELL-STALE-ONLINE] -->\n-| 2.7 | A node restart must not permanently strand every `persistent` shell: (a) the daemon heals a recorded `online` the derivation contradicts, write-guarded on an actual change, and (b) a once-per-generation boot sweep PLUS an owner offline->online edge in the reconcile loop (releases#228 — the sweep alone strands the shells of any owner not yet online when it ran) relaunch an instance only when ALL of — adapter `persistent`, owner endpoint online, down in fact, recorded launch predates the boot instant (slack-absorbed) — hold; no stamp or no boot oracle ⇒ not restored. The owner-facing surfaces derive and look clean while the cascade reads the recorded field, so a clean render is NOT evidence the class did not occur — the falsifier is the on-disk record. The pre-existing cascade test's setup suppresses the failing arm (it suspends first), so the restart shape must be asserted on a fixture that never calls the suspend path | `shellwake::{heal_stale_online_records,restore_persistent_shells_at_boot,restore_persistent_shells_on_owner_online,OwnerOnlineEdge,launch_predates_boot}`, `proc::boot_instant_ms`, `shellinfo::{record_shell_launch,read_shell_launch}` |\n+| 2.7 | A once-per-generation boot sweep and each owner offline->online edge restore every instance whose adapter is `persistent`, owner is online, and binary is down in fact. Machine boot time, launch age, and absent launch stamps do not block persistent restoration; already-online and live binding launches are excluded. Heal every stale instance per owner with no repeat writes. Prove the restart shape without first suspending the fixture; a clean derived display alone is not evidence the record was healed | `shellwake::{heal_stale_online_records,restore_persistent_shells_at_boot,restore_persistent_shells_on_owner_online,OwnerOnlineEdge}`, `shellhost::live_launch_winner` |\n | 3.1 | Ephemeral perch cleanup on all exit paths | `ring` (RAII guard) |\n | 3.4 | A ring never adopts (so never deletes) a perch it did not create — probe the DIR, not the ready marker; record/spool/unreadable = occupied, empty = refused too; deliver + loud `RING_PERCH_EXISTS`/`RING_STALE_DIR` instead of block-waiting | `spt_msg::ring` probe + `create_dir` leaf |\n | 3.5 | A perch GC classifies on record PRESENCE only and never asks a liveness resolver — `is_perch_alive` is INVERTED here (recordless residue reads ALIVE, an offline endpoint reads DEAD: 24 vs 6 measured on HFENDULEAM 2026-08-04), so `!is_perch_alive` authorizes exactly the inverse set; the registry is never asked either (offline endpoints are legitimately absent from it) | `spt_store::perchgc::sweep`, `spt endpoint gc` |","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120,121,122,123,124,125,126,127,128,129,130]},"meta":{"source":{"type":"internal","value":"artifact://128"}}},"isError":false,"timestamp":1789001680446}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_JD7tofLk9TXlPtgCyT61C26C|fc_028129a7f1d66e9c016aa1ffde511c87d2a0da9f39f7213481","toolName":"grep","startedAt":"2026-09-10T00:54:54.426Z","args":{"path":".worktrees/287-shellrestore/crates/spt/src/cli.rs;.worktrees/287-shellrestore/crates/spt-daemon/src/linkhost.rs"},"intent":"Locating launch failure cleanup"},"id":"f3e77132","parentId":"d2537f3c","timestamp":"2026-09-10T00:54:54.426Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_45XB8SolKVqNmckJxkawUSvO|fc_028129a7f1d66e9c016aa1ffde513087d2a57db66c185b9c66","toolName":"read","startedAt":"2026-09-10T00:54:54.426Z","args":{"path":".worktrees/287-shellrestore/crates/spt-store/src/atomic.rs:151-230:raw"},"intent":"Reading atomic failure cleanup"},"id":"96fb3b59","parentId":"f3e77132","timestamp":"2026-09-10T00:54:54.426Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_ibBwKgAdCY9WFkvpH6zugOsf|fc_028129a7f1d66e9c016aa1ffde513887d29becb8d7f731c52f","toolName":"grep","startedAt":"2026-09-10T00:54:54.427Z","args":{"path":".worktrees/287-shellrestore/crates"},"intent":"Checking publication callers"},"id":"59d0150c","parentId":"96fb3b59","timestamp":"2026-09-10T00:54:54.427Z"}
{"type":"message","id":"e897ca27","parentId":"59d0150c","timestamp":"2026-09-10T00:54:54.427Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":"**Reordering stamp retirement**\n\n**Sending failure-path feedback**"},{"type":"toolCall","id":"call_JD7tofLk9TXlPtgCyT61C26C|fc_028129a7f1d66e9c016aa1ffde511c87d2a0da9f39f7213481","name":"grep","arguments":{"i":"Locating launch failure cleanup","pattern":"is_indeterminate\\(\\)|launch_shell_from\\(|teardown_record","path":".worktrees/287-shellrestore/crates/spt/src/cli.rs;.worktrees/287-shellrestore/crates/spt-daemon/src/linkhost.rs","case":true,"gitignore":false,"skip":null},"intent":"Locating launch failure cleanup"},{"type":"toolCall","id":"call_45XB8SolKVqNmckJxkawUSvO|fc_028129a7f1d66e9c016aa1ffde513087d2a57db66c185b9c66","name":"read","arguments":{"i":"Reading atomic failure cleanup","path":".worktrees/287-shellrestore/crates/spt-store/src/atomic.rs:151-230:raw"},"intent":"Reading atomic failure cleanup"},{"type":"toolCall","id":"call_ibBwKgAdCY9WFkvpH6zugOsf|fc_028129a7f1d66e9c016aa1ffde513887d29becb8d7f731c52f","name":"grep","arguments":{"i":"Checking publication callers","pattern":"publish_shell_launch|record_shell_launch\\(","path":".worktrees/287-shellrestore/crates","case":true,"gitignore":false,"skip":null},"intent":"Checking publication callers"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":8247,"output":403,"cacheRead":105344,"cacheWrite":0,"totalTokens":113994,"reasoningTokens":204,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1789001680468,"responseId":"resp_028129a7f1d66e9c016aa1ffd0d44487d2a7edf9a74075e4bb","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_028129a7f1d66e9c016aa1ffd2446487d2a0a4555626e9fdd9","type":"reasoning","content":[],"encrypted_content":"gAAAAABqof_egmYaTyI1Va0945r_WTWLnEbGMeUkyRpYgb0HCh4L5OGMRAaBT__V12j8WImfApdfkTt2sh7udE_3q84tn-A5huqdFYY5NeSJqSKLPIAr9GEzkwwvHlbPkDb5lo49X7_vtp9wimzOa4QmaZdNzQcBv4yErbomNOQsZqbtkeXFTaxC1Qg0Upz4PrMB4WTDKgc7zd-4fzzJJoGtUZkkGOYN9Sd0QyG7ETtT0Yui4R3y4BQdLm0mkpZaCc7QaAPhQGtcPH8tfmRez5Hzy5rskBelUZyRe_zKN-3XH4wRnyPIir-BXduFotkHaO8IbGE01PEBrqWE6AArir107SK0pnQuafds37GWN5rUWAD0mSaT-i6wXKgqJxhi_Z3uDvu8RkZnzC9Dhm4PfKuk8UK4vOIlilbi7NQx4-MFlXIxKULuCQ7y5S2cSQuEPWInKZkh695eJ4372M5ox1pE8fs082vAATxQ_GjdULu1qiH2vDdxdiKLBsOQPcn-O1svRkwNIoIdcptqjttmskoJFJ3S0Ty3H623yFgwfXHNz2V41Umlb6IWTlAxqFu0fcr5376ynfK2LYXAM-HjC4rdCOE-l6ZY8gVMmdhPc9PDUX_rUygXdJKeG2DiUOMFPWRwOqu6UJiXWLwvecm8axIMz82DJzmEevQU7WCSHgHwWMK0wCS_-h30LQvCpcMEvRI-hLrXCkcZK2zFAymcXWYzfJhTRBB-eCnx9w2nOUBUsTYaCIv-kpHjidxYLsC9tUs9yQi8CDrG1-p7aFi1pd0HkK0VXIk1e3IJJty1LfK3Y2yT2rMidk3LU0iF3KQRTqfwXoRgHmCARVc_lE4Tt9FisHOC3RY5ty6-S_HT6icL6oPFdh7znvH8yWES1PdhfO-qHZjijCrFijRdeepo99Ucb5T4RvTC07TBaDLHw_05loTu4yYMkFTufR96y56d1EWOtlfMahi2pWMUIpffZ-8RHZ368ZNicc3dcM8YexWqzKXnvkBwAnp_HJ5_WnTXmWOZkAf-uY4stZYSaIo6g7lJwSi1vwOjfGxVm465dVQJYFI5rKXa5_m_05_TGhQ4PRwVEryi0m5BmQreVjD1j1cRpb94rJDpBcpZ7rn-SqK2IeJKt4OH3rUR3KMFjVvSgtiEjySVTuuUznDXtnXHyk-DrIGqz1V0ktO84xrr7FKH1fg07lI2UC-MENcwQlA8Rdbt41JMXZBmbBOD_lNk4d_VhiAr13ZCPf0mZAy3o7Sm61h53ICjBxOFNk0fpoQWeyTDJsyglox3yTFIqUYky8rtpTVpoYwDXwux1nhGl7ZVcEWullniM6QF3bCpBpNw0tQxEvWqQjzBwOc1TFtJhd2X8_WM1ATz598QIBnr50EZ39g57f2KYP_Vq2HWdBQ3FbqreksqI8-8paaYjiSUe7GkUtK8hR7a-eVHk0if7zKGkFMCpRhQUcG8C3v9fH5lTtrfhyG-5jM0CRn9q6xB1o6Ie2OoECP0FHToVQWHwyu-SBIx6zqA2pnVQunJ7RnE7p8DjNXpldMgfSR9eAyqyYMf3S6KtY3Hy5QljU7aV9aCwoZ7BPtj0CXBCabHI3fZXifuI4nVKNc_Ztayfcj9y5NaHXO9K_40iDbMAGewEzT75WDO8hw_A1pEGfHsFE54VZO90k4RvbV1MJa7a56wJ2JoibBeyt2nAIXjaUx9Ueiv4k3K59OKVtu-sDIeZtPVnAJh8TKE2M0l6ds-iAmwAg_3SU52XSCHEVMSXZBnF_6FYcuGgt2JiMRF_7ky9NEeAaXcfaj_CJWSgaO5nk2eBDbJensb-W2I4ucVy5pFVoL73bMWOi4IohKhtPy5ypsuDaIedK-9GIY0bp1mnzZoRjyHyOxd_8stJcpkQ9d5fq47VlZSNUyvB1yB_DfjJ7dcsu8eJoeAaPmQbR294055u46yzO6FmK94G_qDkW3lxfUbJuSe73mlZ0F5h0zYojaNQpDnJy6A23yoIY2nBMMEVXS0vFzwcls61UqDv2fNAYTaSGBp5kjJLErxUsxt1kztFajdPbu-Sr2IJvZfSrLq2BuIowZhBda9tDFMxeO1gPngFP8ubEUF1a4YN-Vb0YHBQ6hD9OuV2NAfaT0J-crMUW40qDLKt8oTUXan70haKzbbOs2G4WtdpC7Yd8f4Fhmisaw2a5KZHeHjSaoeyfikd4Z_kwfV7K-HE7s0F4jscNqbh7-qoRdSBQEpfATez5AORAKNsdTzO1uPY5Nnznzro-N3VcJ82H4u0j5-czIJqKxOAiFaWUjPWXQsB80poYZOwo7GXOKDAbKlY0JtW92rCxe6302FUSEL5aLP00xQobvXqk_lncdLufP2Ctf1zPDt0c42WB5KOPTP3akWrzCWivGZ9WzEXFboutxypkX_5AOrKDQ-QU9ylWyzHxl5f3XKgl9J7ww_6zOW_I8LB5h6BFwcVZfaQrcv-cdmK6v-m3WFTd2hQX4wB_q6CRSj8B1N81CJu3JpkjzcLCx-oEGE8KxesbpY3GVnea9p5i0oDoFuFMFeabYVU7H0w8yo-_lpIz9fIXoTSnd1aQunT1CQ9AvaiNTYdOOVMcfgzAezbHFA7oGxmn11jyHyaxbt72qw58LwckjufVtGeD49bDFovZ-5hqTOmuYCiFI-WJaMoi-E8_QplfzDateUzmpk5OeWxRMFnW7U3ebjwddEHzlof-uqHI2_qpmBlo5j8039P1wAUbVL2yI03F0=","internal_chat_message_metadata_passthrough":{"turn_id":"897e71d8-de26-4b06-a763-cbc0db5f84cd"},"summary":[{"type":"summary_text","text":"**Reordering stamp retirement**"},{"type":"summary_text","text":"**Sending failure-path feedback**"}],"metadata":{"turn_id":"897e71d8-de26-4b06-a763-cbc0db5f84cd"}},{"id":"fc_028129a7f1d66e9c016aa1ffde511c87d2a0da9f39f7213481","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating launch failure cleanup\",\"pattern\":\"is_indeterminate\\\\(\\\\)|launch_shell_from\\\\(|teardown_record\",\"path\":\".worktrees/287-shellrestore/crates/spt/src/cli.rs;.worktrees/287-shellrestore/crates/spt-daemon/src/linkhost.rs\",\"case\":true,\"gitignore\":false,\"skip\":null}","call_id":"call_JD7tofLk9TXlPtgCyT61C26C","internal_chat_message_metadata_passthrough":{"turn_id":"897e71d8-de26-4b06-a763-cbc0db5f84cd"},"name":"grep","metadata":{"turn_id":"897e71d8-de26-4b06-a763-cbc0db5f84cd"}},{"id":"fc_028129a7f1d66e9c016aa1ffde513087d2a57db66c185b9c66","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading atomic failure cleanup\",\"path\":\".worktrees/287-shellrestore/crates/spt-store/src/atomic.rs:151-230:raw\"}","call_id":"call_45XB8SolKVqNmckJxkawUSvO","internal_chat_message_metadata_passthrough":{"turn_id":"897e71d8-de26-4b06-a763-cbc0db5f84cd"},"name":"read","metadata":{"turn_id":"897e71d8-de26-4b06-a763-cbc0db5f84cd"}},{"id":"fc_028129a7f1d66e9c016aa1ffde513887d29becb8d7f731c52f","type":"function_call","status":"completed","arguments":"{\"i\":\"Checking publication callers\",\"pattern\":\"publish_shell_launch|record_shell_launch\\\\(\",\"path\":\".worktrees/287-shellrestore/crates\",\"case\":true,\"gitignore\":false,\"skip\":null}","call_id":"call_ibBwKgAdCY9WFkvpH6zugOsf","internal_chat_message_metadata_passthrough":{"turn_id":"897e71d8-de26-4b06-a763-cbc0db5f84cd"},"name":"grep","metadata":{"turn_id":"897e71d8-de26-4b06-a763-cbc0db5f84cd"}}]},"duration":13954.292799999937,"ttft":1834.778799999971,"completedAt":1789001694423,"contextSnapshot":{"promptTokens":113591,"nonMessageTokens":13209,"compactionEpoch":0}}}
{"type":"message","id":"7a4a50d4","parentId":"e897ca27","timestamp":"2026-09-10T00:54:54.430Z","message":{"role":"toolResult","toolCallId":"call_45XB8SolKVqNmckJxkawUSvO|fc_028129a7f1d66e9c016aa1ffde513087d2a57db66c185b9c66","toolName":"read","content":[{"type":"text","text":"            Err(e)\n        }\n    }\n}\n\n/// Rename `from` → `to`, retrying transient sharing/access failures with\n/// linear backoff. Non-transient errors (e.g. NotFound) return immediately.\n// [impl->REQ-HAZARD-EBUSY-RENAME]\nfn rename_with_retry(from: &Path, to: &Path) -> io::Result<()> {\n    retry_transient(MAX_RENAME_ATTEMPTS, backoff, || fs::rename(from, to))\n}\n\n/// Linear backoff: attempt 1 → 2ms, attempt 2 → 4ms, … capped modestly. Kept\n/// small so a real transient (handle released within tens of ms) is absorbed\n/// without a perceptible stall.\nfn backoff(attempt: u32) -> Duration {\n    Duration::from_millis((attempt as u64 * 2).min(50))\n}\n\n/// Run `op`, retrying while it fails with a *transient* error, up to\n/// `max_attempts` total tries. Sleeps `delay(attempt)` between tries. Returns\n/// the first success, or the last error once attempts are exhausted /\n/// the error is non-transient. Generic over the operation so the retry\n/// semantics are unit-testable without provoking a real OS EBUSY.\n// [impl->REQ-HAZARD-EBUSY-RENAME]\nfn retry_transient<T>(\n    max_attempts: u32,\n    delay: impl Fn(u32) -> Duration,\n    mut op: impl FnMut() -> io::Result<T>,\n) -> io::Result<T> {\n    let mut attempt = 0u32;\n    loop {\n        match op() {\n            Ok(v) => return Ok(v),\n            Err(e) => {\n                attempt += 1;\n                if attempt >= max_attempts || !is_transient(&e) {\n                    return Err(e);\n                }\n                std::thread::sleep(delay(attempt));\n            }\n        }\n    }\n}\n\n/// Whether an io error is a transient sharing/access failure worth retrying.\n///\n/// Windows surfaces a held-handle rename collision as `ERROR_SHARING_VIOLATION`\n/// (32) or `ERROR_ACCESS_DENIED` (5). On other platforms `PermissionDenied`\n/// is the analogue; everything else (NotFound, etc.) is non-transient.\nfn is_transient(e: &io::Error) -> bool {\n    #[cfg(windows)]\n    {\n        if let Some(code) = e.raw_os_error() {\n            // 5 = ERROR_ACCESS_DENIED, 32 = ERROR_SHARING_VIOLATION.\n            return code == 5 || code == 32;\n        }\n    }\n    matches!(e.kind(), io::ErrorKind::PermissionDenied)\n}\n\n#[cfg(test)]\nmod tests {\n    use super::*;\n    use std::cell::Cell;\n    use std::path::PathBuf;\n\n    fn zero_delay(_: u32) -> Duration {\n        Duration::from_millis(0)\n    }\n\n    /// No `.tmp` sibling of any kind remains in `dir` (the unique tmp names —\n    /// `{name}.tmp.{pid}-{seq}` — must all be renamed away, not just a fixed\n    /// `{name}.tmp`). Glob-style: any entry whose name contains `.tmp` fails.\n    fn no_tmp_sibling_remains(dir: &Path) {\n        let leftover: Vec<String> = fs::read_dir(dir)\n            .unwrap()\n            .flatten()\n            .filter_map(|e| e.file_name().into_string().ok())\n            .filter(|n| n.contains(\".tmp\"))\n\n[Showing lines 151-230 of 393. Use :231 to continue]"}],"details":{"truncation":{"content":"            Err(e)\r\n        }\r\n    }\r\n}\r\n\r\n/// Rename `from` → `to`, retrying transient sharing/access failures with\r\n/// linear backoff. Non-transient errors (e.g. NotFound) return immediately.\r\n// [impl->REQ-HAZARD-EBUSY-RENAME]\r\nfn rename_with_retry(from: &Path, to: &Path) -> io::Result<()> {\r\n    retry_transient(MAX_RENAME_ATTEMPTS, backoff, || fs::rename(from, to))\r\n}\r\n\r\n/// Linear backoff: attempt 1 → 2ms, attempt 2 → 4ms, … capped modestly. Kept\r\n/// small so a real transient (handle released within tens of ms) is absorbed\r\n/// without a perceptible stall.\r\nfn backoff(attempt: u32) -> Duration {\r\n    Duration::from_millis((attempt as u64 * 2).min(50))\r\n}\r\n\r\n/// Run `op`, retrying while it fails with a *transient* error, up to\r\n/// `max_attempts` total tries. Sleeps `delay(attempt)` between tries. Returns\r\n/// the first success, or the last error once attempts are exhausted /\r\n/// the error is non-transient. Generic over the operation so the retry\r\n/// semantics are unit-testable without provoking a real OS EBUSY.\r\n// [impl->REQ-HAZARD-EBUSY-RENAME]\r\nfn retry_transient<T>(\r\n    max_attempts: u32,\r\n    delay: impl Fn(u32) -> Duration,\r\n    mut op: impl FnMut() -> io::Result<T>,\r\n) -> io::Result<T> {\r\n    let mut attempt = 0u32;\r\n    loop {\r\n        match op() {\r\n            Ok(v) => return Ok(v),\r\n            Err(e) => {\r\n                attempt += 1;\r\n                if attempt >= max_attempts || !is_transient(&e) {\r\n                    return Err(e);\r\n                }\r\n                std::thread::sleep(delay(attempt));\r\n            }\r\n        }\r\n    }\r\n}\r\n\r\n/// Whether an io error is a transient sharing/access failure worth retrying.\r\n///\r\n/// Windows surfaces a held-handle rename collision as `ERROR_SHARING_VIOLATION`\r\n/// (32) or `ERROR_ACCESS_DENIED` (5). On other platforms `PermissionDenied`\r\n/// is the analogue; everything else (NotFound, etc.) is non-transient.\r\nfn is_transient(e: &io::Error) -> bool {\r\n    #[cfg(windows)]\r\n    {\r\n        if let Some(code) = e.raw_os_error() {\r\n            // 5 = ERROR_ACCESS_DENIED, 32 = ERROR_SHARING_VIOLATION.\r\n            return code == 5 || code == 32;\r\n        }\r\n    }\r\n    matches!(e.kind(), io::ErrorKind::PermissionDenied)\r\n}\r\n\r\n#[cfg(test)]\r\nmod tests {\r\n    use super::*;\r\n    use std::cell::Cell;\r\n    use std::path::PathBuf;\r\n\r\n    fn zero_delay(_: u32) -> Duration {\r\n        Duration::from_millis(0)\r\n    }\r\n\r\n    /// No `.tmp` sibling of any kind remains in `dir` (the unique tmp names —\r\n    /// `{name}.tmp.{pid}-{seq}` — must all be renamed away, not just a fixed\r\n    /// `{name}.tmp`). Glob-style: any entry whose name contains `.tmp` fails.\r\n    fn no_tmp_sibling_remains(dir: &Path) {\r\n        let leftover: Vec<String> = fs::read_dir(dir)\r\n            .unwrap()\r\n            .flatten()\r\n            .filter_map(|e| e.file_name().into_string().ok())\r\n            .filter(|n| n.contains(\".tmp\"))\r","truncated":true,"truncatedBy":"lines","totalLines":243,"totalBytes":2916,"outputLines":80,"outputBytes":2916,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":393,"displayContent":{"text":"            Err(e)\r\n        }\r\n    }\r\n}\r\n\r\n/// Rename `from` → `to`, retrying transient sharing/access failures with\r\n/// linear backoff. Non-transient errors (e.g. NotFound) return immediately.\r\n// [impl->REQ-HAZARD-EBUSY-RENAME]\r\nfn rename_with_retry(from: &Path, to: &Path) -> io::Result<()> {\r\n    retry_transient(MAX_RENAME_ATTEMPTS, backoff, || fs::rename(from, to))\r\n}\r\n\r\n/// Linear backoff: attempt 1 → 2ms, attempt 2 → 4ms, … capped modestly. Kept\r\n/// small so a real transient (handle released within tens of ms) is absorbed\r\n/// without a perceptible stall.\r\nfn backoff(attempt: u32) -> Duration {\r\n    Duration::from_millis((attempt as u64 * 2).min(50))\r\n}\r\n\r\n/// Run `op`, retrying while it fails with a *transient* error, up to\r\n/// `max_attempts` total tries. Sleeps `delay(attempt)` between tries. Returns\r\n/// the first success, or the last error once attempts are exhausted /\r\n/// the error is non-transient. Generic over the operation so the retry\r\n/// semantics are unit-testable without provoking a real OS EBUSY.\r\n// [impl->REQ-HAZARD-EBUSY-RENAME]\r\nfn retry_transient<T>(\r\n    max_attempts: u32,\r\n    delay: impl Fn(u32) -> Duration,\r\n    mut op: impl FnMut() -> io::Result<T>,\r\n) -> io::Result<T> {\r\n    let mut attempt = 0u32;\r\n    loop {\r\n        match op() {\r\n            Ok(v) => return Ok(v),\r\n            Err(e) => {\r\n                attempt += 1;\r\n                if attempt >= max_attempts || !is_transient(&e) {\r\n                    return Err(e);\r\n                }\r\n                std::thread::sleep(delay(attempt));\r\n            }\r\n        }\r\n    }\r\n}\r\n\r\n/// Whether an io error is a transient sharing/access failure worth retrying.\r\n///\r\n/// Windows surfaces a held-handle rename collision as `ERROR_SHARING_VIOLATION`\r\n/// (32) or `ERROR_ACCESS_DENIED` (5). On other platforms `PermissionDenied`\r\n/// is the analogue; everything else (NotFound, etc.) is non-transient.\r\nfn is_transient(e: &io::Error) -> bool {\r\n    #[cfg(windows)]\r\n    {\r\n        if let Some(code) = e.raw_os_error() {\r\n            // 5 = ERROR_ACCESS_DENIED, 32 = ERROR_SHARING_VIOLATION.\r\n            return code == 5 || code == 32;\r\n        }\r\n    }\r\n    matches!(e.kind(), io::ErrorKind::PermissionDenied)\r\n}\r\n\r\n#[cfg(test)]\r\nmod tests {\r\n    use super::*;\r\n    use std::cell::Cell;\r\n    use std::path::PathBuf;\r\n\r\n    fn zero_delay(_: u32) -> Duration {\r\n        Duration::from_millis(0)\r\n    }\r\n\r\n    /// No `.tmp` sibling of any kind remains in `dir` (the unique tmp names —\r\n    /// `{name}.tmp.{pid}-{seq}` — must all be renamed away, not just a fixed\r\n    /// `{name}.tmp`). Glob-style: any entry whose name contains `.tmp` fails.\r\n    fn no_tmp_sibling_remains(dir: &Path) {\r\n        let leftover: Vec<String> = fs::read_dir(dir)\r\n            .unwrap()\r\n            .flatten()\r\n            .filter_map(|e| e.file_name().into_string().ok())\r\n            .filter(|n| n.contains(\".tmp\"))\r","startLine":151,"lineNumbers":[151,152,153,154,155,156,157,158,159,160,161,162,163,164,165,166,167,168,169,170,171,172,173,174,175,176,177,178,179,180,181,182,183,184,185,186,187,188,189,190,191,192,193,194,195,196,197,198,199,200,201,202,203,204,205,206,207,208,209,210,211,212,213,214,215,216,217,218,219,220,221,222,223,224,225,226,227,228,229,230]},"fileSize":18102,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\crates\\spt-store\\src\\atomic.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":393,"totalBytes":2916,"outputLines":80,"outputBytes":2916,"shownRange":{"start":151,"end":230},"nextOffset":231}}},"isError":false,"timestamp":1789001694430}}
{"type":"message","id":"6bbc7038","parentId":"7a4a50d4","timestamp":"2026-09-10T00:54:54.443Z","message":{"role":"toolResult","toolCallId":"call_JD7tofLk9TXlPtgCyT61C26C|fc_028129a7f1d66e9c016aa1ffde511c87d2a0da9f39f7213481","toolName":"grep","content":[{"type":"text","text":"# .worktrees/287-shellrestore/crates/\n\n## spt/src/\n### cli.rs\n 23022|            let install_dir = spt_daemon::shellhost::shell_install_dir(&adapters_dir, &adapter);\n*23023|            match spt_daemon::shellhost::launch_shell_from(\n 23024|                spt_daemon::shellhost::LaunchOrigin::Cli,\n 23025|                &owlery,\n 23026|                &owner,\n...\n 23062|                // [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\n*23063|                Err(e) if e.is_indeterminate() => {\n 23064|                    eprintln!(\n 23065|                        \"SHELL_SPAWN_INDETERMINATE:{id}: the daemon accepted the launch and \\\n 23066|                         never reported its outcome ({e}) — the binary MAY be running and MAY \\\n\n## spt-daemon/src/\n### linkhost.rs\n 695|    // may well be coming up would relink again and double it.\n*696|    let pid = shellhost::launch_shell_from(\n 697|        origin,\n 698|        owlery,\n 699|        owner,"}],"details":{"scopePath":".worktrees/287-shellrestore/crates/spt/src/cli.rs, .worktrees/287-shellrestore/crates/spt-daemon/src/linkhost.rs","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\crates","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":3,"fileCount":2,"files":[".worktrees/287-shellrestore/crates/spt/src/cli.rs",".worktrees/287-shellrestore/crates/spt-daemon/src/linkhost.rs"],"fileMatches":[{"path":".worktrees/287-shellrestore/crates/spt/src/cli.rs","count":2},{"path":".worktrees/287-shellrestore/crates/spt-daemon/src/linkhost.rs","count":1}],"truncated":false,"displayContent":"# .worktrees/287-shellrestore/crates/\n\n## spt/src/\n### cli.rs\n 23022│            let install_dir = spt_daemon::shellhost::shell_install_dir(&adapters_dir, &adapter);\n*23023│            match spt_daemon::shellhost::launch_shell_from(\n 23024│                spt_daemon::shellhost::LaunchOrigin::Cli,\n 23025│                &owlery,\n 23026│                &owner,\n      │...\n 23062│                // [impl->REQ-SHELL-CLI-SPAWN-JOB-EXPOSURE]\n*23063│                Err(e) if e.is_indeterminate() => {\n 23064│                    eprintln!(\n 23065│                        \"SHELL_SPAWN_INDETERMINATE:{id}: the daemon accepted the launch and \\\n 23066│                         never reported its outcome ({e}) — the binary MAY be running and MAY \\\n\n## spt-daemon/src/\n### linkhost.rs\n 695│    // may well be coming up would relink again and double it.\n*696│    let pid = shellhost::launch_shell_from(\n 697│        origin,\n 698│        owlery,\n 699│        owner,"},"isError":false,"timestamp":1789001694443}}
{"type":"message","id":"b3f36c1c","parentId":"6bbc7038","timestamp":"2026-09-10T00:54:54.514Z","message":{"role":"toolResult","toolCallId":"call_ibBwKgAdCY9WFkvpH6zugOsf|fc_028129a7f1d66e9c016aa1ffde513887d29becb8d7f731c52f","toolName":"grep","content":[{"type":"text","text":"# .worktrees/287-shellrestore/crates/\n\n## spt-daemon/src/\n### linkhost.rs\n 1169|        std::fs::write(p.join(shellhost::SHELL_PID_FILE), pid.to_string()).unwrap();\n*1170|        shellinfo::record_shell_launch(&p, pid, 1_000);\n 1171|        std::fs::write(p.join(shellhost::LINK_TOKEN_FILE), token).unwrap();\n 1172|        p\n 1173|    }\n...\n 1565|            std::fs::write(perch_path.join(shellhost::SHELL_PID_FILE), live.to_string()).unwrap();\n*1566|            shellinfo::record_shell_launch(&perch_path, live, 1_000);\n 1567|            match prepare_drive(&owlery, \"doyle\", \"Stick\", \"stick\", \"x=0.7,y=-0.2\").unwrap() {\n 1568|                DrivePrep::Deliver { id: d, token, frame } => {\n 1569|                    assert_eq!(d, id);\n### shelldisc.rs\n 348|            // depending on a stamp being absent.\n*349|            shellinfo::record_shell_launch(&perch, pid, 1_000);\n 350|        }\n 351|\n 352|        let d = discover(&owlery, \"dead\", &registered);\n### shellhost.rs\n 500|            // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n*501|            spt_store::shellinfo::publish_shell_launch(&perch, pid, now_ms())\n 502|                .map_err(|e| format!(\"record launch: {e}\"))?;\n 503|            Ok(pid)\n 504|        }\n...\n 556|    // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n*557|    spt_store::shellinfo::publish_shell_launch(&perch, pid, now_ms())\n 558|        .map_err(|e| format!(\"record launch: {e}\"))?;\n 559|    Ok(pid)\n 560|}\n...\n 1220|        );\n*1221|        shellinfo::record_shell_launch(perch, me, 1_000);\n 1222|        assert_eq!(live_launch_winner(perch), Some(me), \"same birth retains custody\");\n 1223|        let mut launch = shellinfo::read_shell_launch(perch).unwrap();\n 1224|        launch.pid_started_at = Some(\n...\n 1786|        std::fs::write(perch.join(SHELL_PID_FILE), ours.id().to_string()).unwrap();\n*1787|        shellinfo::record_shell_launch(perch, ours.id(), 1);\n 1788|\n 1789|        kill_shell_at(perch);\n 1790|\n### shellwake.rs\n 1106|            .unwrap();\n*1107|            shellinfo::record_shell_launch(&live_perch, std::process::id(), 1_000);\n 1108|\n 1109|            let mtime = |p: &std::path::Path| {\n 1110|                std::fs::metadata(spt_store::perch::info_file_at(p))\n...\n 1213|        if let Some(launched_ms) = launched_ms {\n*1214|            shellinfo::record_shell_launch(&perch, 2_000_000_000, launched_ms);\n 1215|        }\n 1216|        id\n 1217|    }\n...\n 1249|        .unwrap();\n*1250|        shellinfo::record_shell_launch(&binding_perch, std::process::id(), boot + 120_000);\n 1251|\n 1252|        let mut restored = restore_persistent_shells_at_boot(owlery, &registered, &adapters_dir);\n 1253|        restored.sort();\n...\n 1533|        std::fs::write(perch.join(shellinfo::SHELL_PID_FILE), pid.to_string()).unwrap();\n*1534|        shellinfo::record_shell_launch(&perch, pid, 1_000);\n 1535|        let mut launch = shellinfo::read_shell_launch(&perch).unwrap();\n 1536|        launch.pid_started_at = Some(\n 1537|            launch.pid_started_at.expect(\"native birth on supported platforms\") ^ 1,\n\n## spt-store/src/\n### shellinfo.rs\n 164|// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n*165|pub fn record_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) {\n 166|    let launch = ShellLaunch {\n 167|        pid_started_at: crate::proc::process_started_at(pid),\n 168|        launched_ms: now_ms,\n...\n 179|// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n*180|pub fn publish_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) -> std::io::Result<()> {\n 181|    match std::fs::remove_file(shell_perch.join(SHELL_LAUNCH_FILE)) {\n 182|        Ok(()) => {}\n 183|        Err(e) if e.kind() == std::io::ErrorKind::NotFound => {}\n...\n 186|    atomic_write_string(&shell_perch.join(SHELL_PID_FILE), &pid.to_string())?;\n*187|    record_shell_launch(shell_perch, pid, now_ms);\n 188|    Ok(())\n 189|}\n 190|\n...\n 514|        let d = online_perch(Some(&std::process::id().to_string()));\n*515|        record_shell_launch(d.path(), std::process::id(), 1_000);\n 516|        let parked = read_shell_launch(d.path()).expect(\"a launch stamp was parked\");\n 517|        assert_eq!(parked.launched_ms, 1_000, \"the absolute stamp round-trips\");\n 518|\n...\n 532|        std::fs::create_dir(d.path().join(SHELL_LAUNCH_FILE)).unwrap();\n*533|        let err = publish_shell_launch(d.path(), std::process::id(), 1_000).unwrap_err();\n 534|        assert_ne!(err.kind(), std::io::ErrorKind::NotFound);\n 535|        assert_eq!(read_shell_pid(d.path()), Some(123));\n 536|    }\n\n## spt/\n\n### src/\n#### cli.rs\n 32681|        .unwrap();\n*32682|        spt_store::shellinfo::record_shell_launch(&shell_perch, live, 1_000);\n 32683|        assert_eq!(relink(\"Scout\", \"doyle\"), 1, \"already online\");\n 32684|    }\n 32685|\n\n### tests/\n#### activity_link_push_e2e.rs\n 304|        .unwrap();\n*305|        spt_store::shellinfo::record_shell_launch(&shell_perch, keeper_pid, now_ms());\n 306|        let link_instant = now_ms();\n 307|        let (link_frame, _) = wait_for_state(&token_a, \"busy\", Duration::from_secs(10));\n 308|        let link_since: u64 = link_frame.attr(\"since\").unwrap().parse().unwrap();\n#### attach_link_push_e2e.rs\n 352|        .unwrap();\n*353|        spt_store::shellinfo::record_shell_launch(&shell_perch, keeper_pid, now_ms());\n 354|\n 355|        // ── CLAUSE 1: current attachment at link time, with nothing to wait for.\n 356|        let first = wait_for_attach(\n...\n 500|        .unwrap();\n*501|        spt_store::shellinfo::record_shell_launch(&shell_perch, keeper.id(), now_ms());\n 502|\n 503|        let re = wait_for_attach(\n 504|            &token_b,\n#### drive_e2e.rs\n 187|    std::fs::write(shell_perch.join(\"shell.pid\"), stand_in.id().to_string()).unwrap();\n*188|    spt_store::shellinfo::record_shell_launch(&shell_perch, stand_in.id(), 1_000);\n 189|    let token_a = token_of();\n 190|\n 191|    // ── (1) DROP-AT-WRITE: driving an offline shell drops (exit 0, diagnostic),\n#### gateway_owner_shell_e2e.rs\n 231|    std::fs::write(shell_perch.join(\"shell.pid\"), stand_in.id().to_string()).unwrap();\n*232|    spt_store::shellinfo::record_shell_launch(&shell_perch, stand_in.id(), 1_000);\n 233|    online_by_token(&token_a);\n 234|\n 235|    // ── cmd (durable): a gateway-owned command vocab-checks + spools to the"}],"details":{"scopePath":".worktrees/287-shellrestore/crates","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\crates","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":22,"fileCount":10,"files":[".worktrees/287-shellrestore/crates/spt-daemon/src/linkhost.rs",".worktrees/287-shellrestore/crates/spt-daemon/src/shelldisc.rs",".worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs",".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs",".worktrees/287-shellrestore/crates/spt-store/src/shellinfo.rs",".worktrees/287-shellrestore/crates/spt/src/cli.rs",".worktrees/287-shellrestore/crates/spt/tests/activity_link_push_e2e.rs",".worktrees/287-shellrestore/crates/spt/tests/attach_link_push_e2e.rs",".worktrees/287-shellrestore/crates/spt/tests/drive_e2e.rs",".worktrees/287-shellrestore/crates/spt/tests/gateway_owner_shell_e2e.rs"],"fileMatches":[{"path":".worktrees/287-shellrestore/crates/spt-daemon/src/linkhost.rs","count":2},{"path":".worktrees/287-shellrestore/crates/spt-daemon/src/shelldisc.rs","count":1},{"path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs","count":4},{"path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs","count":4},{"path":".worktrees/287-shellrestore/crates/spt-store/src/shellinfo.rs","count":5},{"path":".worktrees/287-shellrestore/crates/spt/src/cli.rs","count":1},{"path":".worktrees/287-shellrestore/crates/spt/tests/activity_link_push_e2e.rs","count":1},{"path":".worktrees/287-shellrestore/crates/spt/tests/attach_link_push_e2e.rs","count":2},{"path":".worktrees/287-shellrestore/crates/spt/tests/drive_e2e.rs","count":1},{"path":".worktrees/287-shellrestore/crates/spt/tests/gateway_owner_shell_e2e.rs","count":1}],"truncated":false,"displayContent":"# .worktrees/287-shellrestore/crates/\n\n## spt-daemon/src/\n### linkhost.rs\n 1169│        std::fs::write(p.join(shellhost::SHELL_PID_FILE), pid.to_string()).unwrap();\n*1170│        shellinfo::record_shell_launch(&p, pid, 1_000);\n 1171│        std::fs::write(p.join(shellhost::LINK_TOKEN_FILE), token).unwrap();\n 1172│        p\n 1173│    }\n     │...\n 1565│            std::fs::write(perch_path.join(shellhost::SHELL_PID_FILE), live.to_string()).unwrap();\n*1566│            shellinfo::record_shell_launch(&perch_path, live, 1_000);\n 1567│            match prepare_drive(&owlery, \"doyle\", \"Stick\", \"stick\", \"x=0.7,y=-0.2\").unwrap() {\n 1568│                DrivePrep::Deliver { id: d, token, frame } => {\n 1569│                    assert_eq!(d, id);\n### shelldisc.rs\n 348│            // depending on a stamp being absent.\n*349│            shellinfo::record_shell_launch(&perch, pid, 1_000);\n 350│        }\n 351│\n 352│        let d = discover(&owlery, \"dead\", &registered);\n### shellhost.rs\n  500│            // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n *501│            spt_store::shellinfo::publish_shell_launch(&perch, pid, now_ms())\n  502│                .map_err(|e| format!(\"record launch: {e}\"))?;\n  503│            Ok(pid)\n  504│        }\n     │...\n  556│    // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n *557│    spt_store::shellinfo::publish_shell_launch(&perch, pid, now_ms())\n  558│        .map_err(|e| format!(\"record launch: {e}\"))?;\n  559│    Ok(pid)\n  560│}\n     │...\n 1220│        );\n*1221│        shellinfo::record_shell_launch(perch, me, 1_000);\n 1222│        assert_eq!(live_launch_winner(perch), Some(me), \"same birth retains custody\");\n 1223│        let mut launch = shellinfo::read_shell_launch(perch).unwrap();\n 1224│        launch.pid_started_at = Some(\n     │...\n 1786│        std::fs::write(perch.join(SHELL_PID_FILE), ours.id().to_string()).unwrap();\n*1787│        shellinfo::record_shell_launch(perch, ours.id(), 1);\n 1788│\n 1789│        kill_shell_at(perch);\n 1790│\n### shellwake.rs\n 1106│            .unwrap();\n*1107│            shellinfo::record_shell_launch(&live_perch, std::process::id(), 1_000);\n 1108│\n 1109│            let mtime = |p: &std::path::Path| {\n 1110│                std::fs::metadata(spt_store::perch::info_file_at(p))\n     │...\n 1213│        if let Some(launched_ms) = launched_ms {\n*1214│            shellinfo::record_shell_launch(&perch, 2_000_000_000, launched_ms);\n 1215│        }\n 1216│        id\n 1217│    }\n     │...\n 1249│        .unwrap();\n*1250│        shellinfo::record_shell_launch(&binding_perch, std::process::id(), boot + 120_000);\n 1251│\n 1252│        let mut restored = restore_persistent_shells_at_boot(owlery, &registered, &adapters_dir);\n 1253│        restored.sort();\n     │...\n 1533│        std::fs::write(perch.join(shellinfo::SHELL_PID_FILE), pid.to_string()).unwrap();\n*1534│        shellinfo::record_shell_launch(&perch, pid, 1_000);\n 1535│        let mut launch = shellinfo::read_shell_launch(&perch).unwrap();\n 1536│        launch.pid_started_at = Some(\n 1537│            launch.pid_started_at.expect(\"native birth on supported platforms\") ^ 1,\n\n## spt-store/src/\n### shellinfo.rs\n 164│// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n*165│pub fn record_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) {\n 166│    let launch = ShellLaunch {\n 167│        pid_started_at: crate::proc::process_started_at(pid),\n 168│        launched_ms: now_ms,\n    │...\n 179│// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n*180│pub fn publish_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) -> std::io::Result<()> {\n 181│    match std::fs::remove_file(shell_perch.join(SHELL_LAUNCH_FILE)) {\n 182│        Ok(()) => {}\n 183│        Err(e) if e.kind() == std::io::ErrorKind::NotFound => {}\n    │...\n 186│    atomic_write_string(&shell_perch.join(SHELL_PID_FILE), &pid.to_string())?;\n*187│    record_shell_launch(shell_perch, pid, now_ms);\n 188│    Ok(())\n 189│}\n 190│\n    │...\n 514│        let d = online_perch(Some(&std::process::id().to_string()));\n*515│        record_shell_launch(d.path(), std::process::id(), 1_000);\n 516│        let parked = read_shell_launch(d.path()).expect(\"a launch stamp was parked\");\n 517│        assert_eq!(parked.launched_ms, 1_000, \"the absolute stamp round-trips\");\n 518│\n    │...\n 532│        std::fs::create_dir(d.path().join(SHELL_LAUNCH_FILE)).unwrap();\n*533│        let err = publish_shell_launch(d.path(), std::process::id(), 1_000).unwrap_err();\n 534│        assert_ne!(err.kind(), std::io::ErrorKind::NotFound);\n 535│        assert_eq!(read_shell_pid(d.path()), Some(123));\n 536│    }\n\n## spt/\n\n### src/\n#### cli.rs\n 32681│        .unwrap();\n*32682│        spt_store::shellinfo::record_shell_launch(&shell_perch, live, 1_000);\n 32683│        assert_eq!(relink(\"Scout\", \"doyle\"), 1, \"already online\");\n 32684│    }\n 32685│\n\n### tests/\n#### activity_link_push_e2e.rs\n 304│        .unwrap();\n*305│        spt_store::shellinfo::record_shell_launch(&shell_perch, keeper_pid, now_ms());\n 306│        let link_instant = now_ms();\n 307│        let (link_frame, _) = wait_for_state(&token_a, \"busy\", Duration::from_secs(10));\n 308│        let link_since: u64 = link_frame.attr(\"since\").unwrap().parse().unwrap();\n#### attach_link_push_e2e.rs\n 352│        .unwrap();\n*353│        spt_store::shellinfo::record_shell_launch(&shell_perch, keeper_pid, now_ms());\n 354│\n 355│        // ── CLAUSE 1: current attachment at link time, with nothing to wait for.\n 356│        let first = wait_for_attach(\n    │...\n 500│        .unwrap();\n*501│        spt_store::shellinfo::record_shell_launch(&shell_perch, keeper.id(), now_ms());\n 502│\n 503│        let re = wait_for_attach(\n 504│            &token_b,\n#### drive_e2e.rs\n 187│    std::fs::write(shell_perch.join(\"shell.pid\"), stand_in.id().to_string()).unwrap();\n*188│    spt_store::shellinfo::record_shell_launch(&shell_perch, stand_in.id(), 1_000);\n 189│    let token_a = token_of();\n 190│\n 191│    // ── (1) DROP-AT-WRITE: driving an offline shell drops (exit 0, diagnostic),\n#### gateway_owner_shell_e2e.rs\n 231│    std::fs::write(shell_perch.join(\"shell.pid\"), stand_in.id().to_string()).unwrap();\n*232│    spt_store::shellinfo::record_shell_launch(&shell_perch, stand_in.id(), 1_000);\n 233│    online_by_token(&token_a);\n 234│\n 235│    // ── cmd (durable): a gateway-owned command vocab-checks + spools to the"},"isError":false,"timestamp":1789001694513}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_Rxt59g60lczsTew2Fn25VyNG|fc_028129a7f1d66e9c016aa1fff2e9fc87d2b8efbfd8496618af","toolName":"hub","startedAt":"2026-09-10T00:55:17.120Z","intent":"Reporting correction error path"},"id":"537fcf39","parentId":"b3f36c1c","timestamp":"2026-09-10T00:55:17.120Z"}
{"type":"message","id":"a8596765","parentId":"537fcf39","timestamp":"2026-09-10T00:55:17.121Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":"**Editing concurrent tests**\n\n**Checking file removal**\n\n**Assessing stamp retirement**"},{"type":"toolCall","id":"call_Rxt59g60lczsTew2Fn25VyNG|fc_028129a7f1d66e9c016aa1fff2e9fc87d2b8efbfd8496618af","name":"hub","arguments":{"i":"Reporting correction error path","op":"send","to":"Main","message":"Correction closes the reviewed successful-publication race, recycled-PID wake, and checklist mismatch. One introduced error path remains: shellhost.rs:501-502 / :557-558 call publish_shell_launch only AFTER child/session spawn; old-stamp remove failure returns Failed with a live child whose PID was never recorded. Your directory-shaped stamp test exercises that refusal only before a real spawn. Retire/refuse before creating child, or clean up the just-created child/session on publication failure; otherwise relink reports failure then retries can duplicate/orphan it. No runtime commands run."},"intent":"Reporting correction error path"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":3539,"output":657,"cacheRead":113408,"cacheWrite":0,"totalTokens":117604,"reasoningTokens":508,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1789001694568,"responseId":"resp_028129a7f1d66e9c016aa1ffdede7487d2806610ae42556c2e","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_028129a7f1d66e9c016aa1ffe1549087d2baf7a3ae655212a1","type":"reasoning","content":[],"encrypted_content":"gAAAAABqof_yWFchvmE62Cmg3mY__t9yDefWAgUb7XTXf1t_gJGCGfZJrU3XYljrO2Bfx3R7O_bCg01pM9jBq-JKnPikw3Dp-SRMpJS7WCRaqatH53VJriXO0PyvdO1RBppV26xaSnCMJODAr0reFkIWMPwxvlJ7IQEvUGKvjKdl8Y6VlIYBsvVU3pX_oSSYDa7cVKawySECxY3bJ4wacqhVR5TWt33pgLo9HXs_Od1OyolInPfC5FJ7_cEXYEctJiWZLA6dunPrv89f5En3nDbUk5MjLuX7idlnlbZOjdP7_rnIhpimGHsrJQFjetlf26JIAFGk7EsWzZxJBas_uiuX2UmDsa34FbzEgXefAUF8GH_m59BffTtQkViGL8CewMEFQzrIdJHZ61rEcxAXGLTLlmSY_8swPpBHSbnfgOd8w5d5CtbJ7TfTz_P5lCE8dkVlofacX8BMSQn3KBUW16ZBvRAHcV4imBNYK_-HARZ2tGDxpx7ZVAsPib51qgOU4sx41MveoojjvsbL7X0BHElfnjn8RToRU-S8-5w8Ys6PQgD-urtZib1ZxOY6iHY-CAdQQsx_zBazujzGGQgfDzHQPqC-oE6n9BGdtmE88iZlV2dtX0J1sKrlej7ZQ5etbKOjzJ1hCx225NWRUn2FDhd-EPr1ycx1-TbW9B9IEFCwCO-YKFAt6lUb2Ec-OMwMbUueIvf36ia6ZytutNkWpBzwlQdHvOUDzIR2xDXRIlpDV1c547fCLsbD6YNNKgYACAFjd_4Qp8oW4lNJ6ZfWgcGjrsxAbeUry6WzWjQLsAeD0_tm-bBSYuxMklJOU8NQv4AgGzq1kfrJAI8xXT-uKz7cCTGGKIJUCouzBniKdB3VhvFNmRRMcGD-_tPy0wfg7e1xTHbSxD5biJx1P9KKyKiKfFICwCwKpIpvffR9phl7OsL1Cts6y5R6Moq_Oc3PukMGFALGhlHg5EW0DCU18PE8RoKLZGfk5gut7S1SncoxaImiVuC3kXUMFcMVkR9iJp7aQunJea5_FbXIR3BQMJ-bFFVUg0gF9wNOE6H8gWmUVP7ELYR8KG2MJCW6tUIOuLqegj9uHXO15xRdm-CAEHZhfzRuIqCM82cMJMv41Oww2t1Aw1uZfiH4fgnpF-25PMi-TvRxNTUhIcXH8Qsmspw8f7NuM2DlSj-SYLZ94fL_PR0KT-myUpWV96KWrtQ2oTnyzQrpZfM7rRNixPdLMtYn2EYXPMRPK_Eon1i1Eennhbqy9hkAQr8P2RhFcPq3UOBjiQO_74DO8KfKSaoVmst8JbbSW4Pl9d93C6dXx9a8VvCk99AGNMAY4vtfvHEPzmPzzBI5Qn1PGqmArp1heMTnyc2vwCnj423GRy754YjFqqaX7Ak4UtP3_p6L2QAYLkLpRSjNGApHGSjfjVvKmOyMcLlvAHDQ8sRqUzRdWImV07Ub6zYcCag1kygzzRmmZdpVamP8dDkZO_tsTrbceDghf3xok-HKwNac1doH7u2xVyWchtjCwmQOqLgjlyHKpIQhCw6m7pMfeq5GMxcQkulDM-TYr65Xpa2gTuGCe2CBgsbIQvmXQUnUu_iVE2EIiUIX4862sqH8OY-rBpPrEkruRjwFhlR6-6SIUF06f7lj9D01t5MRKaGTdOhwzhyMCgkKmedBKBvWFLPR4uCDBeEX0a3WA9BKq7bXCKVZiLy5EAJ-7bnbeYzsFTZqXRg67N_BrO7yiPFgJbP2sskvrKz73GEK6Vm8wqvgdRV3_KUlycMj_dbqtLYLkidnRu1TopHIzD170HSyfhFVah6pkKgOARdH5GwUQ-CtlWRUSVp6ERuH2UR4emzurMhElM7zyXLyJopPa6pR1ZYTUqjEDnF9yXW2A449E8p_BT2uvG2c_LC4HcLcLf0x_kbdq0p7NZSCUwvm4rp2BT7Jbg4ZmGm0tCPcFTSCKr-RdmmR7y2WaTkZ-rqVdTaiWnl8OrhBqu5Hp9iKzjtjJ4gxn8LYnv3tk5YeA583Y2COa4pkCJrSPmmLbMLxPBmP_1hdciKsysQbObwCFahbwyUiTaJTgr0XZnuFOvg22VOJMoAJLpRtO_CZxIYsdE3GK2n_JV-JWDXCDTZXkUVuBwLB82Sh98WCqLbdbwqkrBWCZGRDxhq67obTvAOcZQrszDOlNkccnnNKE5cdzzhk5NJ7F1UqMWhU8jZ1s6CipbV_Rdw1NLU0r1jHvtzlJXqPmHFYxGwtg2XJfvF28eBBghBRcfgCmlGnx7TxOp7Ztk1qrcirY-ZiNysHA6b_ZDMU85XwbNXcKNp5-yxS4sM_JveBhuuZ-tccjBYZ1FMiM4dCkvwWXnLzvDlniLr1NnkEltUblccQeaIzaKlFGwkXvTTb6T-pZOQ4vUSN0PU1ljLnv0oEhtT3RygER38vR-4goW3D5j3TpGw9lssCqvjsqpL1eo80rbb_3rZEUgIX3oNPcH4kMGhjJcsEYaI0smonWqeWkVLZfTTQzUEDxP2zB6FqPL-dpxhoGfBMROe5HWp6RNHx7kCHMGRrEUa87QeIOA8w8wua--jWGvrrYFpgqSBdbqqQx9vgJ_scGNazDas4iPw240EZJtkXfRy62C2EnaA9T6DNJ6KvlNMS1JGa07FR2e9xZQsC7MG8jgKFZnc2e1tdH7pBJe8keTpqUE69v0FDl-IY4zn35OVu1V3-g9Pu5SubbssL85ehMcYnpRTlnsKG6kC1vPtfVtK1RXmn2dPHJ_2semZK3e-lBSMlVT79oXSAsdTT1r2WNatE1nbY6f01pjHzHBESqO0rOBegLhA-LxXs0N4OU9gwwI305zwrSdXJSnSUNLuZWrooHtiUtFg0wT0ybtliLRv5nYWHmjWM1Sz-FlcyssKk2Iy4m1n0ZmLVbuSMyatfINBJET4bwlgdBNojXadtgwBTPsyhw5P-sV3ZD6dlSyCCK2dPTTHe7U6B3L4ivvT0f9_Aeuf1K8Qwe-fqQUEWU_wih-QcHMzT5EpfJIqcuXBOSOftwXicHYw_Yrii2KBQt4Pjyn9ccYicGp78easGWVmERUpCFme_IS1Uye94h5mkHuzpb39n1nVYn6olPIpMngOcK0C80Jrbn66KDxJidcCYN1-EnP3nFQk95N2PcN9JbxUloyCV6dhJ7Mq01ECfOgj7oWWu-S5AAVOwuoY7gR4sEf1dYQY40CBg_aT6OJ-s5EbEOncRg-6AUQOPSAYlvuHbPABhOpi5hfQQZ2NOqkMcE8oj-h9w1Hitag4WQB_2YPJNS6r7emI4tZ_btSQ1-bDkpulcpks-1z7q3KhLjZznbDsolGEKGb7TeFOHqbqlv74NCYR2j-7rPstcWfJejBL36lGtwaS-ZW8tP-MI3DAfyiZ80ZqX-3LxgO9_09dNAK6NbMusjTRDcfiB7M5JKNCpQZTaEqtEpJbgFY1ynlAAwZ1ALikYDr5qw_Rx0-D8AcrjFNK0dl5lpfvWh9WxG752EpYyf1tksO4NMMJubSwSvJP8lUpvQaaq6OEKbdReEXdcGf6yUQnRl-0LfzOhKt-bVtHgozA35HXGpZJ8Y_3xdqgWK5-kWRwiiOmBua4rYJBNKsg7COZVPprlATWWzVDQwF1XZsI_SnFKTQIqNXuQ8RuzWkpM_R0PA09bcTLrpuYGHYUUlbasUomP_aDf58h75oe4Adc-zwyJ8DZAs_ZCky0JyYXkVMXFM1LPNsplxFTeBtvkNV6jYDoOJzKqsbI_1ncV5-dLAG1S34A9salPyj_PtaAUPMaH8BRE4OkOBQIIZRAbAHMf4ABSIfP-jjkmpbA3kD_TH2BTcd87JsX-O7F8JsTJUzgGB5JGcEuSkqWC5WnrEzqfCcIsole9Y4Drc6WKzOBHlozD4Jb1On7jE2WxUYWXyEKHYZLn_XjEx8PyYmAyBrDCS-eneR_grWVJhL_FeyGKqmgtwZsRHylbm16Bm_fTmEgCZySs-427fI9U5F76nshCa6GpkHuVuh0iN0Zjna5n3T9AkwuFJNairYxgNtbszZOipw1Hi5q69n0odS_wMMkIc6C6ijuxRChTjuWI4JxKzuIMW9TBt-yTXyZBvf9CGb8AW2gNxE1oFRhUqzoN1ZLtFEJA0_Fa1oYLiCXTUIzxeay4VoGYDpKH2MLO9DlZzJAqgtcPv5o5MFNeLRW2Wl3XGJMczzvzx9EfeOfswn22ghk7JMTQCc28k61G_qBVNszhaRGx2tMdWyrvfx6V-cKCTkZF6TrqiIlTX0PgOc_ANQ-sIk0YF_W67O_TjIxnwTr69gRyYZo1Fbp7CDx8jsLznmYbO2y8kdnPzXHXv_V4y5Fo4wQOoOqD5VXCoFDHgSG0NTmisYD9-DVW2Ce26E474wXB0tG5iQzq0FKhsAapV4TDR2DkwVmsLTTeGMOv1DGCudePIQFSKSeNW_Fd6dbcXzK8V22TpmDP2hK-9s-RDF4oooK51QtUIi0S-yxqHfTD_kUZvlMez-OXHlBBGfZsOhFyBCbqz8IUVJCINsjsoUzi2pkapJVhLbnQfBmB2avvOYJTpVSVBBDa2ljSJnucIfklflraM2-kudNSyBxhhLdnlwKf2Obp_oeVHPUMCVTIrzXX2pQiBfYejsp8DPE4BfMGZOoGkovVt98q5qDXXR_KhEWQqV04T2vM3adlF4OkqOqK7_H9FNJL8_SRNBmKKhDCOJe-NyrtaWf7VnW-IGy0gxUGWMJpNHcu4nVtRgLGY6RdMGLHz6O5QcqYQy6LnEyq4KERcLG4pdX2971sXOHJk0DB1EHc5w0z7c_DwyUBOKAzfHPGAtfIGyFL8rVjtsZDQ1M9VYts-GcZORINY1Qe2cXBI57hzF2-mgaTuWu8QCvS_LQq3oIitFloMip9GZwO4b-38s7uzPavfZeHPyxrdwVqW7NI9Lf-Avh9qkeR_SKz3Y9RPDp4of7Y5AtX9iJeErx3dkkUHm0bAWvpjxWvGFwhLXo295g=","internal_chat_message_metadata_passthrough":{"turn_id":"897e71d8-de26-4b06-a763-cbc0db5f84cd"},"summary":[{"type":"summary_text","text":"**Editing concurrent tests**"},{"type":"summary_text","text":"**Checking file removal**"},{"type":"summary_text","text":"**Assessing stamp retirement**"}],"metadata":{"turn_id":"897e71d8-de26-4b06-a763-cbc0db5f84cd"}},{"id":"fc_028129a7f1d66e9c016aa1fff2e9fc87d2b8efbfd8496618af","type":"function_call","status":"completed","arguments":"{\"i\":\"Reporting correction error path\",\"op\":\"send\",\"to\":\"Main\",\"message\":\"Correction closes the reviewed successful-publication race, recycled-PID wake, and checklist mismatch. One introduced error path remains: shellhost.rs:501-502 / :557-558 call publish_shell_launch only AFTER child/session spawn; old-stamp remove failure returns Failed with a live child whose PID was never recorded. Your directory-shaped stamp test exercises that refusal only before a real spawn. Retire/refuse before creating child, or clean up the just-created child/session on publication failure; otherwise relink reports failure then retries can duplicate/orphan it. No runtime commands run.\"}","call_id":"call_Rxt59g60lczsTew2Fn25VyNG","internal_chat_message_metadata_passthrough":{"create_time":1789001695.163473,"turn_id":"897e71d8-de26-4b06-a763-cbc0db5f84cd"},"name":"hub","metadata":{"turn_id":"897e71d8-de26-4b06-a763-cbc0db5f84cd"}}]},"duration":22550.27829999989,"ttft":2797.13450000016,"completedAt":1789001717119,"contextSnapshot":{"promptTokens":116947,"nonMessageTokens":13209,"compactionEpoch":0}}}
{"type":"message","id":"d00e5243","parentId":"a8596765","timestamp":"2026-09-10T00:55:17.122Z","message":{"role":"toolResult","toolCallId":"call_Rxt59g60lczsTew2Fn25VyNG|fc_028129a7f1d66e9c016aa1fff2e9fc87d2b8efbfd8496618af","toolName":"hub","content":[{"type":"text","text":"Delivered to 1 peer(s):\n- Main: injected"}],"details":{"op":"send","from":"PersistentRestoreReview","to":"Main","receipts":[{"to":"Main","outcome":"injected"}]},"isError":false,"timestamp":1789001717121}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_0uldVRBeAAbljJWP7zoZrQp5|fc_028129a7f1d66e9c016aa1fffc856c87d2922326cd3601cfe1","toolName":"yield","startedAt":"2026-09-10T00:55:30.742Z"},"id":"91f3c105","parentId":"d00e5243","timestamp":"2026-09-10T00:55:30.742Z"}
{"type":"message","id":"e48f57a6","parentId":"91f3c105","timestamp":"2026-09-10T00:55:30.743Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":"**Reviewing source corrections**"},{"type":"toolCall","id":"call_0uldVRBeAAbljJWP7zoZrQp5|fc_028129a7f1d66e9c016aa1fffc856c87d2922326cd3601cfe1","name":"yield","arguments":{"type":null,"data":{"overall_correctness":"incorrect","explanation":"The correction resolves the reviewed successful-publication race, the recycled-PID wake guard, and the contradictory checklist. One new failure path remains: stamp-retirement failure is reported only after spawning a child that is then left unrecorded. Source review only; no tests, builds, linters, formatters, edits, or production operations were run.","confidence":0.97,"findings":[{"title":"Retire the old stamp before spawning, or unwind publication failure","body":"Both production callers invoke `publish_shell_launch` after the detached child or broker session has already been created. If removing the old launch stamp fails—for example, the directory-shaped stamp used in the new test or a Windows access/sharing error—the helper returns before recording the new PID, and this `?` reports an ordinary launch failure without stopping that new process. A relink therefore leaves a live binary behind an old or missing PID record; retrying the reported failure can launch another binary and invalidate the first one's token. Perform the fallible stamp retirement before spawning, or explicitly stop the just-created child/session on publication failure.","priority":2,"confidence":0.99,"file_path":"C:/Users/decid/Documents/projects/spt-core/.worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs","line_start":499,"line_end":502}]},"error":null}}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":762,"output":403,"cacheRead":116864,"cacheWrite":0,"totalTokens":118029,"reasoningTokens":91,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1789001717178,"responseId":"resp_028129a7f1d66e9c016aa1fff56f9487d2a3a47054c5aae7ef","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_028129a7f1d66e9c016aa1fff68f1087d2b4b8435c9ed76f17","type":"reasoning","content":[],"encrypted_content":"gAAAAABqof_8fYq9EV2akDZWC5TMgD4lmjVknd7AvddEr5aO5wJ3XNpStHE5kurEd7RuPONtYavrICYCOLWSG2HMXc0nFD_FrTmMFkE20XFCKrA0SnA74a5FsvmpE2FM2tI5UjJHi1m8sNqbNhZ_46lAXt5G36NQEjMwf6P7jgb6yzwQ0Lu2DWEky3zBrF-sH-nCaqtTyXkg614LvlM6LfL2yQtvTDpnUZnO5yOmb634xLHU8If5W3YTM54VAnbiUXyBX_c268Yy3Quk3rbwZ83lLQ97_jB2Or8G5IT6yKPtMdRCt6s4b-cyIEnncKCUhoz2x5i8CeLmzlIggovly5B7nYnTLeWJniaueXBaaE40-5QfbP74EeaaUQQTu3yJtUo2Oo_mweJiEKssC4tQ0EDOjy6cx7VWoWCIh8DLAbx22k1w0v7z2TgVmk6gn3x0uFfKf6rMATQNiWE-LiZVstMpiFKWdr5SAMpLha04DYauw_6Bozm0XzA_t3Tg65bWhdk2-3Y5iKlyWBtIbTCKVIHhLwPVSLDLRTIPnW6xaxMaw8NxY25W8lflQl8A4a02Y2lnq87c2tWsjNWOn2y9E3mEs1tjBGDNsl84uK_bfeUNMmObDwHmXROrmlCyADJWUWYpUjto_-kQM9ynxYrTieLSm6wd9C_rZNhbrbVrBx95WSf_2Pqvs-MSZV7bDwx_Fu2pID49ViSh312aPAweti04uRvSUF8gaQbJvI_g6E37TRMfKuCOJaFChHjkTyHNDFrA6fG3qUWwWpnjJpf-B9m6INUWa1Hc5KrovshwM1VDozQMTW6sX0Dkp-rj-tFg8yXIKlGN0iAvlhY6pmGy0MkG8erLLX_RyIcRlVz5BlATXPCom-RTj7Xx6jy3yiGcze43tqxnEC9okEMMuoE0nJwXkGPlnjM1359XkHLSUwwKwouchI-v-blmDBtTrkiuVrPgapD86zoMTLEgfWt---bMyMOOCyxffAMbfxhewpkKYjaxwCz7SA522JcBobd5EQEOn3Xic0IbDvnqpu9FAUQGZbuqLSj6TOXkXqW94dHvUOovb8Hjq3s1TJGekF3GBT4lSbwWJvKYlyfQo_zAfjBzGIWo4knaxiriEggzNr8K7OdDC6VIMsOT_v0Q4cabw6ieZeL3FKZSa4Eh5RlCRrnt3OcQiFXqlA_wInTUK9s-IVRDLcJemr0Zhh-J8zfk9lQFfR-VD532mcb3Ok-G9Gj0uyibskyucRHAbUoNui814Ft0JRXh3dwD9XTsczjcy4QXwUcLTiQGL45Z3cMl9Uv8vH5Ig-_r4d3ALHi01HYb27PoJU0ucVhZx6ZC4pHtpjbhkBkcc-Ih44_8JkZorBI4CSi80we6xeuf2SdBOYsZC4-zf0KuU8IOgQGXAdVtcOJa1zw3jzYNP3cB_TGuHV3pykZqhX32tUkG1yf1_ofrE8UvIwFAvCe33mIziMecxhR2N1SsF7F7yPJV5ZR7Ge14fi4akrjM9z1EGhLz01pQlMfvtR0y4472KeO5iNon7YWFyZycpNsn1gfrKDpbu0wwYUqrYgVc8Xtx6snCkJuGFF08ZJVdZJUbvSUaJOUUl29yPfOSam_jyonBvf7GXnQqqOweors9faMU5U35oWd0bHWoDKxMTn9XCBAQb3qBJQLhinwMT0CsKmtfQs9IYDQT7yXe4YtXOPNBcmfnm24KJEBdb1tjjJYpYmUw2EQLxTGjDwP0PcVO6Q97Ku4JhTfeEvn79L7arsbY1PquNW1pJiUhJl91nxwmvx5A56TzLZfugWLuHO1Qc859aAWEc_ht3PBWOuRMD19YCbo7L67rLAjwJNqUO-ar6YyCOv6zIjpymCuFAjg3YNULLdt6DJbs74v_QiDVV6xLwkgUo8kXCCFdM6vOEk1LSq-fBJErEOfWTx_rQzvtJEtVN6Oy4YsdQeYgR338pA==","internal_chat_message_metadata_passthrough":{"turn_id":"897e71d8-de26-4b06-a763-cbc0db5f84cd"},"summary":[{"type":"summary_text","text":"**Reviewing source corrections**"}],"metadata":{"turn_id":"897e71d8-de26-4b06-a763-cbc0db5f84cd"}},{"id":"fc_028129a7f1d66e9c016aa1fffc856c87d2922326cd3601cfe1","type":"function_call","status":"completed","arguments":"{\"type\":null,\"data\":{\"overall_correctness\":\"incorrect\",\"explanation\":\"The correction resolves the reviewed successful-publication race, the recycled-PID wake guard, and the contradictory checklist. One new failure path remains: stamp-retirement failure is reported only after spawning a child that is then left unrecorded. Source review only; no tests, builds, linters, formatters, edits, or production operations were run.\",\"confidence\":0.97,\"findings\":[{\"title\":\"Retire the old stamp before spawning, or unwind publication failure\",\"body\":\"Both production callers invoke `publish_shell_launch` after the detached child or broker session has already been created. If removing the old launch stamp fails—for example, the directory-shaped stamp used in the new test or a Windows access/sharing error—the helper returns before recording the new PID, and this `?` reports an ordinary launch failure without stopping that new process. A relink therefore leaves a live binary behind an old or missing PID record; retrying the reported failure can launch another binary and invalidate the first one's token. Perform the fallible stamp retirement before spawning, or explicitly stop the just-created child/session on publication failure.\",\"priority\":2,\"confidence\":0.99,\"file_path\":\"C:/Users/decid/Documents/projects/spt-core/.worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs\",\"line_start\":499,\"line_end\":502}]},\"error\":null}","call_id":"call_0uldVRBeAAbljJWP7zoZrQp5","internal_chat_message_metadata_passthrough":{"create_time":1789001717.746235,"turn_id":"897e71d8-de26-4b06-a763-cbc0db5f84cd"},"name":"yield","metadata":{"turn_id":"897e71d8-de26-4b06-a763-cbc0db5f84cd"}}]},"duration":13561.689299999736,"ttft":1414.7960999999195,"completedAt":1789001730741,"contextSnapshot":{"promptTokens":117626,"nonMessageTokens":13209,"compactionEpoch":0}}}
{"type":"message","id":"8c51368d","parentId":"e48f57a6","timestamp":"2026-09-10T00:55:30.744Z","message":{"role":"toolResult","toolCallId":"call_0uldVRBeAAbljJWP7zoZrQp5|fc_028129a7f1d66e9c016aa1fffc856c87d2922326cd3601cfe1","toolName":"yield","content":[{"type":"text","text":"Result submitted."}],"details":{"data":{"overall_correctness":"incorrect","explanation":"The correction resolves the reviewed successful-publication race, the recycled-PID wake guard, and the contradictory checklist. One new failure path remains: stamp-retirement failure is reported only after spawning a child that is then left unrecorded. Source review only; no tests, builds, linters, formatters, edits, or production operations were run.","confidence":0.97,"findings":[{"title":"Retire the old stamp before spawning, or unwind publication failure","body":"Both production callers invoke `publish_shell_launch` after the detached child or broker session has already been created. If removing the old launch stamp fails—for example, the directory-shaped stamp used in the new test or a Windows access/sharing error—the helper returns before recording the new PID, and this `?` reports an ordinary launch failure without stopping that new process. A relink therefore leaves a live binary behind an old or missing PID record; retrying the reported failure can launch another binary and invalidate the first one's token. Perform the fallible stamp retirement before spawning, or explicitly stop the just-created child/session on publication failure.","priority":2,"confidence":0.99,"file_path":"C:/Users/decid/Documents/projects/spt-core/.worktrees/287-shellrestore/crates/spt-daemon/src/shellhost.rs","line_start":499,"line_end":502}]},"status":"success"},"isError":false,"timestamp":1789001730743}}
{"type":"custom_message","customType":"irc:incoming","content":"<irc>\nIncoming IRC message from agent `Main`:\n\nMoved stamp retirement into prepare_shell_launch before actual child/session creation in both launch functions; publish_shell_launch now only atomically writes PID then current birth. Preparation error propagates before spawn. Existing post-spawn PID-write failure behavior remains unchanged scope. Added real-error wake regression (recycled live PID must reach nonexistent binary spawn error, not return adoption), plus preflight refusal cell. No runtime proof until BOX FREE.\nIf response expected, reply via `hub` (`op: \"send\"`, `to: \"Main\"`) when available; otherwise what you `yield` or say last this turn is delivered to `Main` when you stop.\n</irc>","display":true,"details":{"id":"15798200a17ffef6","from":"Main","message":"Moved stamp retirement into prepare_shell_launch before actual child/session creation in both launch functions; publish_shell_launch now only atomically writes PID then current birth. Preparation error propagates before spawn. Existing post-spawn PID-write failure behavior remains unchanged scope. Added real-error wake regression (recycled live PID must reach nonexistent binary spawn error, not return adoption), plus preflight refusal cell. No runtime proof until BOX FREE."},"attribution":"agent","id":"808343e4","parentId":"8c51368d","timestamp":"2026-09-10T00:58:24.773Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_9pqY4DtYzO6MCm2Hd7XajHdU|fc_028129a7f1d66e9c016aa200b4042c87d2a23d6aa0a90dc433","toolName":"bash","startedAt":"2026-09-10T00:58:30.193Z","args":{"command":"git diff -- crates/spt-store/src/shellinfo.rs crates/spt-daemon/src/shellhost.rs crates/spt-daemon/src/shellwake.rs"},"intent":"Reviewing preflight correction"},"id":"14a7ffba","parentId":"808343e4","timestamp":"2026-09-10T00:58:30.193Z"}
{"type":"message","id":"a563e440","parentId":"14a7ffba","timestamp":"2026-09-10T00:58:30.194Z","message":{"role":"assistant","content":[{"type":"toolCall","id":"call_9pqY4DtYzO6MCm2Hd7XajHdU|fc_028129a7f1d66e9c016aa200b4042c87d2a23d6aa0a90dc433","name":"bash","arguments":{"i":"Reviewing preflight correction","command":"git diff -- crates/spt-store/src/shellinfo.rs crates/spt-daemon/src/shellhost.rs crates/spt-daemon/src/shellwake.rs","cwd":"C:/Users/decid/Documents/projects/spt-core/.worktrees/287-shellrestore"},"intent":"Reviewing preflight correction"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":688,"output":78,"cacheRead":117504,"cacheWrite":0,"totalTokens":118270,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1789001904827,"responseId":"resp_028129a7f1d66e9c016aa200b10d3087d2a4712eb77341e713","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"fc_028129a7f1d66e9c016aa200b4042c87d2a23d6aa0a90dc433","type":"function_call","status":"completed","arguments":"{\"i\":\"Reviewing preflight correction\",\"command\":\"git diff -- crates/spt-store/src/shellinfo.rs crates/spt-daemon/src/shellhost.rs crates/spt-daemon/src/shellwake.rs\",\"cwd\":\"C:/Users/decid/Documents/projects/spt-core/.worktrees/287-shellrestore\"}","call_id":"call_9pqY4DtYzO6MCm2Hd7XajHdU","internal_chat_message_metadata_passthrough":{"create_time":1789001905.484612,"turn_id":"dee2191f-fdea-4bb2-86a3-42a7f28a3bb7"},"name":"bash","metadata":{"turn_id":"dee2191f-fdea-4bb2-86a3-42a7f28a3bb7"}}]},"duration":5364.091000000015,"ttft":3219.373300000094,"completedAt":1789001910192,"contextSnapshot":{"promptTokens":118192,"nonMessageTokens":13209,"compactionEpoch":0}}}
{"type":"message","id":"e01bfd49","parentId":"a563e440","timestamp":"2026-09-10T00:58:31.820Z","message":{"role":"toolResult","toolCallId":"call_9pqY4DtYzO6MCm2Hd7XajHdU|fc_028129a7f1d66e9c016aa200b4042c87d2a23d6aa0a90dc433","toolName":"bash","content":[{"type":"text","text":"crates/spt-daemon/src/shellhost.rs | 26 ++++++++++++------------\ncrates/spt-daemon/src/shellwake.rs | 40 +++++++++++++++++++++---\ncrates/spt-store/src/shellinfo.rs | 35 ++++++++++++++++++++++++\n3 files changed, 81 insertions(+), 20 deletions(-)\n\n--- Changes ---\n\nFile: crates/spt-daemon/src/shellhost.rs\n  @@ -477,6 +477,9 @@ pub fn launch_shell(\n  +            let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n  +            spt_store::shellinfo::prepare_shell_launch(&perch)\n  +                .map_err(|e| format!(\"prepare launch: {e}\"))?;\n  @@ -495,14 +498,10 @@ pub fn launch_shell(\n  -            let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n  -            std::fs::write(perch.join(SHELL_PID_FILE), pid.to_string())\n  -                .map_err(|e| format!(\"record pid: {e}\"))?;\n  -            // The birth stamp rides WITH the pid at every launch, or the probe that\n  -            // reads it degrades to bare-pid liveness on this instance alone -- a\n  -            // silent per-instance regression rather than a visible one.\n  […4 changed lines elided…]\n  @@ -527,6 +526,9 @@ pub fn launch_shell_brokered_in(\n  +    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n  +    spt_store::shellinfo::prepare_shell_launch(&perch)\n  +        .map_err(|e| format!(\"prepare launch: {e}\"))?;\n  @@ -554,14 +556,10 @@ pub fn launch_shell_brokered_in(\n  -    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n  -    std::fs::write(perch.join(SHELL_PID_FILE), pid.to_string())\n  -        .map_err(|e| format!(\"record pid: {e}\"))?;\n  -    // Same stamp at the broker-hosted launch. A backend that exposed no pid\n  -    // records 0 here; record_shell_launch stamps no start-time for it, so the\n  -    // probe keeps its fail-toward-alive stance for pid-less backends.\n  […4 changed lines elided…]\n\nFile: crates/spt-daemon/src/shellwake.rs\n  @@ -406,13 +406,9 @@ pub fn resolve_wake(\n  -    // A parked pid that probes alive means somebody won; stand down.\n  +    // A live, birth-matching launch wins even before its bind handshake.\n  -    let live_pid = std::fs::read_to_string(perch.join(crate::shellhost::SHELL_PID_FILE))\n  -        .ok()\n  -        .and_then(|s| s.trim().parse::<u32>().ok())\n  -        .filter(|&p| p != 0 && spt_store::proc::is_process_alive(p));\n  […2 changed lines elided…]\n  @@ -1525,6 +1521,38 @@ mod tests {\n  +    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n  +    #[test]\n  +    fn resolve_wake_does_not_adopt_a_recycled_pid_as_a_launch() {\n  +        let tmp = tempfile::tempdir().unwrap();\n  +        let owlery = tmp.path();\n  +        seed_owner(owlery, \"ling\");\n  […26 changed lines elided…]\n\nFile: crates/spt-store/src/shellinfo.rs\n  @@ -172,6 +172,27 @@ pub fn record_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) {\n  +/// Retire the previous birth before spawning a replacement. Failure must be\n  +/// reported before a child exists, not after creating an unrecorded process.\n  +// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n  +pub fn prepare_shell_launch(shell_perch: &Path) -> std::io::Result<()> {\n  +    match std::fs::remove_file(shell_perch.join(SHELL_LAUNCH_FILE)) {\n  +        Ok(()) => {}\n  […15 changed lines elided…]\n  @@ -212,6 +233,11 @@ pub fn shell_pid_provably_dead(shell_perch: &Path) -> bool {\n  +    // A relaunch can publish between these reads. Do not apply the new birth\n  +    // to the old PID; an unstable snapshot is not proof that the shell died.\n  +    if read_shell_pid(shell_perch) != Some(pid) {\n  +        return false;\n  +    }\n  @@ -504,6 +530,15 @@ mod tests {\n  +    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n  +    #[test]\n  +    fn launch_publication_refuses_an_unretirable_old_stamp() {\n  +        let d = online_perch(Some(\"123\"));\n  +        std::fs::create_dir(d.path().join(SHELL_LAUNCH_FILE)).unwrap();\n  +        prepare_shell_launch(d.path()).unwrap_err();\n  […3 changed lines elided…]\n[raw output: artifact://135]\n\n\nWall time: 1.62 seconds"}],"details":{"timeoutSeconds":300,"wallTimeMs":1623.8339999997988},"isError":false,"timestamp":1789001911819}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_V4cksvh82Mygkj6neHanDF8Z|fc_028129a7f1d66e9c016aa200c73f3087d28da9382ced434229","toolName":"read","startedAt":"2026-09-10T00:58:47.366Z","args":{"path":"artifact://135"},"intent":"Reading preflight correction patch"},"id":"0977d238","parentId":"e01bfd49","timestamp":"2026-09-10T00:58:47.366Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_3OxJNnvlmGeP6XvvHQD4RS6L|fc_028129a7f1d66e9c016aa200c73f4887d2a0076d9f672d575c","toolName":"read","startedAt":"2026-09-10T00:58:47.367Z","args":{"path":".worktrees/287-shellrestore/crates/spt-store/src/shellinfo.rs:158-255:raw"},"intent":"Reading preflight record lifetime"},"id":"713582b0","parentId":"0977d238","timestamp":"2026-09-10T00:58:47.367Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_8GsznHhOsDQtr6UdyOJVbpbx|fc_028129a7f1d66e9c016aa200c73f5087d2969e2d031b76a828","toolName":"read","startedAt":"2026-09-10T00:58:47.367Z","args":{"path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs:1520-1561:raw"},"intent":"Reading failed wake regression"},"id":"bf8b0c47","parentId":"713582b0","timestamp":"2026-09-10T00:58:47.367Z"}
{"type":"message","id":"19119a0d","parentId":"bf8b0c47","timestamp":"2026-09-10T00:58:47.367Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":"**Preparing a PID-birth record**\n\n**Inspecting spawn-failure path**\n\n**Reporting spawn rollback risk**"},{"type":"toolCall","id":"call_V4cksvh82Mygkj6neHanDF8Z|fc_028129a7f1d66e9c016aa200c73f3087d28da9382ced434229","name":"read","arguments":{"i":"Reading preflight correction patch","path":"artifact://135"},"intent":"Reading preflight correction patch"},{"type":"toolCall","id":"call_3OxJNnvlmGeP6XvvHQD4RS6L|fc_028129a7f1d66e9c016aa200c73f4887d2a0076d9f672d575c","name":"read","arguments":{"i":"Reading preflight record lifetime","path":".worktrees/287-shellrestore/crates/spt-store/src/shellinfo.rs:158-255:raw"},"intent":"Reading preflight record lifetime"},{"type":"toolCall","id":"call_8GsznHhOsDQtr6UdyOJVbpbx|fc_028129a7f1d66e9c016aa200c73f5087d2969e2d031b76a828","name":"read","arguments":{"i":"Reading failed wake regression","path":".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs:1520-1561:raw"},"intent":"Reading failed wake regression"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":1460,"output":475,"cacheRead":118016,"cacheWrite":0,"totalTokens":119951,"reasoningTokens":342,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1789001911878,"responseId":"resp_028129a7f1d66e9c016aa200b81a1c87d2904d5e4a417f24c1","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_028129a7f1d66e9c016aa200b905b487d2b753af95a9591783","type":"reasoning","content":[],"encrypted_content":"gAAAAABqogDHg_Y2VL0mNEp2B2DsWVU4kyrQsEkg-hsPwWz1xhi8bIooNxAy_kub3U659yGNU7Y_yxLwhx7bKUGfTGGotnJO-UfbohcoU7YAGzJFE3YOQhVRgbs8IgQqHhHs9ofAUJYqc9SoZyRo1pWr_rRcQ_wfheZ-bNt8odIMqij1qA32PGKsRCX0UP2yHW60E3OO15WzwwXNMkokGpxve3pFtMtCX9XyX829kTiEyhEuyw1UBkep2MhFXCw7hrxuVGssuwT-Z_T8CAS9N0-b9x5VRzALdoDu8VbOhdBPBMsLa0sp9fPU8uxIBsZIO8JRCY-ykThKLMy9pcZgH7bSPNTNIZca6p7y0XKFHTj1LXuAngc-9RQOqTzXMWtUdgcvhXFXact0-9EOehsvqf3rCtin6pWQODT9RqmwIIdn6p4aOb4K6v8CukivXk_k5WUJzZmP5OMmXqIQCYa1gvjvgI5nbDFQ6QxCyuM9sfi80mim6CJ1x2FT56q810ZXwtK6h_3GL2njYshi5foEyzp0XgxNlM4V6dfcbZXBuX5c9MIa5pJ_uu6dheq7If-5y_MT37IzQdpqT1NO2MQkUrTNekzYh2yeAG-Kub9v4hgng29YzztphgjEQyKrqloGO0fqTC6MZIMdjESf2nfeI1pTVAprIF8--oJhfzjhalsYsiLTEUemwsj2_utpUFm9HG7FVKUVAJVf1cq8FKKpRN2joNjNKRyBF1YZNpaviI09ScKxzjs3Ro5-gKU6KBLw0gu7NiAzbpFnS-CCvzz6uno1_AdCyZryH843IiaNa8u-n3rf5sQWWQdTXXLGGSond6LAdx47Em-rgYph-w9JBMxT9_NAbZH-K8KL5fRNgoez76Vgd6oeWppsORF0snoPkj96vwRkx6cTU6tx-vb1R_b6e0uDgSFVtDoHewyOEdrJo5B6fqWDlwCsKgKYJayEB6k2UrTb5Hg1cpXoMG7I0Elv_4JdMBQh4yQ_JrhUozztIdO3M1nWbpI4_0Nsk0H8H-0jnUUz_yzaISQPW9rpnvNtqYyA8eOxZTQ0cgkKOYSn99oePzmUGZAJcA0OphR0XZLu51PhKHhsUPCfcX0JqPUql31ORk6AZ0KV91tOnSmklahjQ-tEjn0lMQheJmJTDYUf7K94w-Q7Q_6vsoqLFlkHBhxEAmhbwZHleAL1Ghhx7MPTeE4BM4keag1JjEVWpOU7dhkZXFbAR_2maPyiQm1oBk9fWPVhJWCeRy-3qAo1-cOgIBLHmRePQvLPZz5Cak-jOf2MdMCvSTiJ2HOPcIpt6QHcrt5mUfjeVtlh4FqinIaxSFBPFgEdwPb3_qmB3KsZDZLCdCVf6wMrh4i0GSGGQWZxSDAJx8zal-kNEUv1gqzs_ocHx_PgJjAVJbyjuzvB5WjoS_SofFJL43LbvSzU86HmaiCB9WYMlo_H2rWeUxCV2IryJ_3KH1miXPrpIyx5Hm48icysep4lhyjjfByWR7GLyKjHtG7oiy6KXCYC1mQeXNCsUUHwGa6pVe1DF1Dwrgf93DpIBChJ94Yx82d3v4WhFyI9qDLdaAvOtv1nabq9Sg45MJUKyml0ej9aRVQa2zS0MzfROMQafiUZfdrbFpCZCBlz-bbtdare3ibYLNgFCx6XGMI8Ooft3pvXEUOpIoQmZdT-036WlPZ5STTJOhATuR5adLBc_YU-JGvWAGzdZEGk7oJ-s_nzrDpdaAFp76tAR2CP2ImnDlMBvQ92KBpB5xR7HbYrNqlu0ryvCZG6DmRCIMmOLqEEicLjISZ4yw8fdXEYuB4sG3_ivYjNEcz9-NXHQ0LwCku2IfYrmF1HcB6HSvS18OEjWJEM-93ZDjPKp-feZSZAt5bR5lnSfnQg7RNPyCNCkAVmyyJ1UJ-H6HNsJw8hjr56HzDTvY_GtHfi0raAS6ygY1cIBiGsASI4jqyGhXpWJYzMALdw-JXuWOlSgA0DJmvMo6umDB4z1rElhMSaYoRw-A7wkt6SrpqXh91Kqt1A0DpUSQujolavM-2T4cCgG5CAg4qMYvZV6szGmubLwJoDgmQjvx--zHKf8o_AKgdj7e5fJ1oq0tLEkepgKMlkLuvM6_LBrPIz49cOso-H2Nz0jhi_tBbTLC20zeW-5LrxFpp8TYkw8TPzATd8fzCIu69illtmOsZU8PuH_RscBGGt6GuRuIgIY0L84DBYlU1qJk7QyJxWbnGGllHP1nC-hwkPA8CA4NkgE-ePWTRo1egb0wruh9Blgc7CBbGIQq-vXzIJKTzH-0sZdBsqSTaAGmjLf6ZLZ6GQCwvV4V9XT_qVrMMhQxvHu8tU34cpYdszqLLY8i9iIizUtB-Yy3SkyYtpVaQwJ2ZM8hlxypzJCJ2Uxp7CasRCvYYeLYJNw_ZtqiL6jitLzocicRozHmvf7pzHiiqbxFVgxK3gfNMCyMKj3dUdhVG-NvyZlPurKWWmdNcLjO5n_6HB_3JSwUJjG-6idq142Og74NbJwzPu6JzqCzN2NvWu6uzQ699UVjN94WVIpMVaswVLF1VnqLad_i0ANnWIlLTA9XwJp1cdOjKHdfxFNrRMfeI2cheBLn7lEPeO0QgIBVXjTA_8s4NXrAHSH7t-zJCHtXWp8m8cmAg2H6yFTNhtRL3vGhSkPIIu_0rJjk4d2Sq_VFwRHfy1g-KJK8rOUCMjtPGP4gWsQalYk-b2Z6-s8NQPiKheAcN9zNK0wERzAw2wAp4oJg1Fg_RoVYPq-4VpB97PKQkPqsCHnCBng5Re-oW_vDwo_rmDvQUDvhUPXMOpUG4OA7FUe8ogG2eVFv4Tl6Jel64jYHgaWcK-XNHdgb_lS_x_La-Iu4Xe5bCjawmqJLqSJJfem9j-qwU2uqNziCAyxdPH5ksez4YIrJTkc-Rq6sDLbKnKiddmfaNmhcuI6La1PDBwX79XxAKu7JN_TMoxPIpWth-zF_LlxV6lLNPQuLjnOrTPtAm5Uf3mDnR1uoY3nwOax5m7ZBGIQfp6LLvEtdiwSeHgBfHV9_UyCZfxYAX6kh_qCD9gcZlosVnM7k8AFQ28XyCKjX3orzne-jYteylaihfMClAJACNVUhEqnPpD8n_AXCj-XeW1qqHL_jQvBRG8Bn0Udp0AcsKgPokBVPfdosq2VPqdq0rxQqDZ1ZMEvSBxVF2jRIMLUxKGerCObFDtJtZ8GIQVvcwSK8BIbcz5CX47eiDmMzJ0mRp6IghrxkNfI2iLoWRnR4m_bxqjvi6Y0rtNWTcCQWuyMFgYhpxD_-I5XZ11pqRBekNQD2094MFau6sTQBV0nSS8_KBrxhvbImleoaXQr6w71UGW_FCIr41jhwT_f66XFQ6bgETRznaAnz818pdYem0lbdRt2JeZCQp8n4Dar2LQ9rM_5aRgPUW6zkZEH-xc8ioMgi6uE2GqQ2OVe0Z9wSsLIgqUL3sCMdbu4CQKPr9lwR519LBi0KkK6hdjXgnvm1E2icdvRp32Z_sNj6u49jbs0DToYm7H5ddWg2wR-mT1sXe_bmNo8OtX1tafBB2s_ONi-5HtvlVbha5LqelWCAQUBUr1rT4Rz3oC8tR3IskoIZfgu10TivbVLAMEj2HbUbjMaVad0lU6IFc7F8vxy9Ta2PUza3Ydd9LyD1PX9KQ7zB9Gy4gotgjR_k7NtRUjstcSchc9j23k-5LNHNnaYvWCT1NPe_aAL02UWOzenzC9Km--7Np6qeiaYgv0F9hX9t930epekwnHUiEVGVoDP1jQbZ_7F7B_tT6yeHmTBkG8Q_LU4nl759fl-du5D7LvrVUHpSYSyQ==","internal_chat_message_metadata_passthrough":{"turn_id":"dee2191f-fdea-4bb2-86a3-42a7f28a3bb7"},"summary":[{"type":"summary_text","text":"**Preparing a PID-birth record**"},{"type":"summary_text","text":"**Inspecting spawn-failure path**"},{"type":"summary_text","text":"**Reporting spawn rollback risk**"}],"metadata":{"turn_id":"dee2191f-fdea-4bb2-86a3-42a7f28a3bb7"}},{"id":"fc_028129a7f1d66e9c016aa200c73f3087d28da9382ced434229","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading preflight correction patch\",\"path\":\"artifact://135\"}","call_id":"call_V4cksvh82Mygkj6neHanDF8Z","internal_chat_message_metadata_passthrough":{"turn_id":"dee2191f-fdea-4bb2-86a3-42a7f28a3bb7"},"name":"read","metadata":{"turn_id":"dee2191f-fdea-4bb2-86a3-42a7f28a3bb7"}},{"id":"fc_028129a7f1d66e9c016aa200c73f4887d2a0076d9f672d575c","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading preflight record lifetime\",\"path\":\".worktrees/287-shellrestore/crates/spt-store/src/shellinfo.rs:158-255:raw\"}","call_id":"call_3OxJNnvlmGeP6XvvHQD4RS6L","internal_chat_message_metadata_passthrough":{"turn_id":"dee2191f-fdea-4bb2-86a3-42a7f28a3bb7"},"name":"read","metadata":{"turn_id":"dee2191f-fdea-4bb2-86a3-42a7f28a3bb7"}},{"id":"fc_028129a7f1d66e9c016aa200c73f5087d2969e2d031b76a828","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading failed wake regression\",\"path\":\".worktrees/287-shellrestore/crates/spt-daemon/src/shellwake.rs:1520-1561:raw\"}","call_id":"call_8GsznHhOsDQtr6UdyOJVbpbx","internal_chat_message_metadata_passthrough":{"turn_id":"dee2191f-fdea-4bb2-86a3-42a7f28a3bb7"},"name":"read","metadata":{"turn_id":"dee2191f-fdea-4bb2-86a3-42a7f28a3bb7"}}]},"duration":15486.4860000005,"ttft":1196.2061000000685,"completedAt":1789001927365,"contextSnapshot":{"promptTokens":119476,"nonMessageTokens":13209,"compactionEpoch":0}}}
{"type":"message","id":"ab472c4d","parentId":"19119a0d","timestamp":"2026-09-10T00:58:47.376Z","message":{"role":"toolResult","toolCallId":"call_3OxJNnvlmGeP6XvvHQD4RS6L|fc_028129a7f1d66e9c016aa200c73f4887d2a0076d9f672d575c","toolName":"read","content":[{"type":"text","text":"}\n\n/// Record the birth stamp for a just-launched `pid`. Best-effort: a stamp that\n/// cannot be written must never fail the launch — a missing stamp degrades the\n/// probe to its pre-existing fail-toward-alive behaviour, which is the safe\n/// direction, while a failed launch is a real outage.\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\npub fn record_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) {\n    let launch = ShellLaunch {\n        pid_started_at: crate::proc::process_started_at(pid),\n        launched_ms: now_ms,\n    };\n    if let Ok(json) = serde_json::to_string_pretty(&launch) {\n        let _ = atomic_write_string(&shell_perch.join(SHELL_LAUNCH_FILE), &json);\n    }\n}\n\n/// Retire the previous birth before spawning a replacement. Failure must be\n/// reported before a child exists, not after creating an unrecorded process.\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\npub fn prepare_shell_launch(shell_perch: &Path) -> std::io::Result<()> {\n    match std::fs::remove_file(shell_perch.join(SHELL_LAUNCH_FILE)) {\n        Ok(()) => {}\n        Err(e) if e.kind() == std::io::ErrorKind::NotFound => {}\n        Err(e) => return Err(e),\n    }\n    Ok(())\n}\n\n/// Publish a prepared launch. The missing stamp preserves live-PID custody\n/// until the new birth is available; it can never name the previous process.\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\npub fn publish_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) -> std::io::Result<()> {\n    atomic_write_string(&shell_perch.join(SHELL_PID_FILE), &pid.to_string())?;\n    record_shell_launch(shell_perch, pid, now_ms);\n    Ok(())\n}\n\n/// Read the birth stamp, if this instance was launched by a build that parks one.\n/// Absent or unparseable ⇒ `None`, which every consumer treats as \"no stamp to\n/// compare\" rather than as a fault.\npub fn read_shell_launch(shell_perch: &Path) -> Option<ShellLaunch> {\n    let raw = std::fs::read_to_string(shell_perch.join(SHELL_LAUNCH_FILE)).ok()?;\n    serde_json::from_str(&raw).ok()\n}\n\n/// **The** discriminant behind [`effective_status`]: is this instance's recorded\n/// binary *provably* gone? Narrow on purpose — true only when the perch parks a\n/// pid, it parses non-zero, and the process is not alive.\n///\n/// Everything else reads **alive**: an absent or unparseable pid file, and the\n/// `0` a broker-hosted spawn records when its backend exposed no pid\n/// ([`crate::proc::is_process_alive`] already reads 0 as dead, so the explicit\n/// filter here is what keeps a pid-less backend from being declared dead by this\n/// resolver). That fail-toward-alive stance mirrors `liveness.rs`'s interim\n/// parity: this predicate exists to catch a *corpse*, never to guess at a\n/// backend it cannot see.\n///\n/// Recycled pids read alive, so a heal can be **missed** but never **mis-fired**\n/// — the failure direction is \"stays stale\", never \"tears down a live instance\".\n// [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\npub fn shell_pid_provably_dead(shell_perch: &Path) -> bool {\n    let Some(pid) = read_shell_pid(shell_perch).filter(|&pid| pid != 0) else {\n        return false;\n    };\n    // BAROMETER W2 (releases#78): probe pid AND birth stamp, not the pid alone.\n    // A pid namespace resets at boot, so a corpse pid can be REUSED by an\n    // unrelated process started after the restart; under bare-pid liveness that\n    // reads ALIVE, the derived status reads online, and `relink` then refuses\n    // SHELL_ALREADY_ONLINE -- the one command that recovers the instance\n    // refusing on the very state that is wrong. The PAIR test (same pid,\n    // different start time) is what tells a reused pid from our own process.\n    //\n    // The stance is unchanged where there is no stamp to compare: an instance\n    // launched before this field existed, or a backend that exposed no start\n    // time, still falls back to plain absence and fails toward ALIVE.\n    // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n    let birth = read_shell_launch(shell_perch).and_then(|l| l.pid_started_at);\n    // A relaunch can publish between these reads. Do not apply the new birth\n    // to the old PID; an unstable snapshot is not proof that the shell died.\n    if read_shell_pid(shell_perch) != Some(pid) {\n        return false;\n    }\n    matches!(\n        crate::liveness::relay_liveness(Some(pid), birth),\n        crate::liveness::RelayLiveness::Gone\n    )\n}\n\n/// A shell instance's **derived** status — the shell-side twin of\n/// `liveness::is_perch_alive` (REQ-HAZARD-DAEMON-HOSTED-LIVENESS gave agent\n/// perches exactly this; shells never got it).\n///\n/// The recorded `status` has exactly two writers — the bind handshake (→online)\n/// and `spt_daemon::shellhost::close_shell` (→offline) — and `close_shell` runs\n/// on the **link-break** path alone. A binary that dies\n/// abruptly (force-kill, crash, OOM) breaks no link, so nothing ever writes the\n/// flip and the record says `online` forever. Deriving closes that hole without\n\n[Showing lines 158-255 of 762. Use :256 to continue]"}],"details":{"truncation":{"content":"}\r\n\r\n/// Record the birth stamp for a just-launched `pid`. Best-effort: a stamp that\r\n/// cannot be written must never fail the launch — a missing stamp degrades the\r\n/// probe to its pre-existing fail-toward-alive behaviour, which is the safe\r\n/// direction, while a failed launch is a real outage.\r\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\npub fn record_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) {\r\n    let launch = ShellLaunch {\r\n        pid_started_at: crate::proc::process_started_at(pid),\r\n        launched_ms: now_ms,\r\n    };\r\n    if let Ok(json) = serde_json::to_string_pretty(&launch) {\r\n        let _ = atomic_write_string(&shell_perch.join(SHELL_LAUNCH_FILE), &json);\r\n    }\r\n}\r\n\r\n/// Retire the previous birth before spawning a replacement. Failure must be\r\n/// reported before a child exists, not after creating an unrecorded process.\r\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\npub fn prepare_shell_launch(shell_perch: &Path) -> std::io::Result<()> {\r\n    match std::fs::remove_file(shell_perch.join(SHELL_LAUNCH_FILE)) {\r\n        Ok(()) => {}\r\n        Err(e) if e.kind() == std::io::ErrorKind::NotFound => {}\r\n        Err(e) => return Err(e),\r\n    }\r\n    Ok(())\r\n}\r\n\r\n/// Publish a prepared launch. The missing stamp preserves live-PID custody\r\n/// until the new birth is available; it can never name the previous process.\r\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\npub fn publish_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) -> std::io::Result<()> {\r\n    atomic_write_string(&shell_perch.join(SHELL_PID_FILE), &pid.to_string())?;\r\n    record_shell_launch(shell_perch, pid, now_ms);\r\n    Ok(())\r\n}\r\n\r\n/// Read the birth stamp, if this instance was launched by a build that parks one.\r\n/// Absent or unparseable ⇒ `None`, which every consumer treats as \"no stamp to\r\n/// compare\" rather than as a fault.\r\npub fn read_shell_launch(shell_perch: &Path) -> Option<ShellLaunch> {\r\n    let raw = std::fs::read_to_string(shell_perch.join(SHELL_LAUNCH_FILE)).ok()?;\r\n    serde_json::from_str(&raw).ok()\r\n}\r\n\r\n/// **The** discriminant behind [`effective_status`]: is this instance's recorded\r\n/// binary *provably* gone? Narrow on purpose — true only when the perch parks a\r\n/// pid, it parses non-zero, and the process is not alive.\r\n///\r\n/// Everything else reads **alive**: an absent or unparseable pid file, and the\r\n/// `0` a broker-hosted spawn records when its backend exposed no pid\r\n/// ([`crate::proc::is_process_alive`] already reads 0 as dead, so the explicit\r\n/// filter here is what keeps a pid-less backend from being declared dead by this\r\n/// resolver). That fail-toward-alive stance mirrors `liveness.rs`'s interim\r\n/// parity: this predicate exists to catch a *corpse*, never to guess at a\r\n/// backend it cannot see.\r\n///\r\n/// Recycled pids read alive, so a heal can be **missed** but never **mis-fired**\r\n/// — the failure direction is \"stays stale\", never \"tears down a live instance\".\r\n// [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\r\npub fn shell_pid_provably_dead(shell_perch: &Path) -> bool {\r\n    let Some(pid) = read_shell_pid(shell_perch).filter(|&pid| pid != 0) else {\r\n        return false;\r\n    };\r\n    // BAROMETER W2 (releases#78): probe pid AND birth stamp, not the pid alone.\r\n    // A pid namespace resets at boot, so a corpse pid can be REUSED by an\r\n    // unrelated process started after the restart; under bare-pid liveness that\r\n    // reads ALIVE, the derived status reads online, and `relink` then refuses\r\n    // SHELL_ALREADY_ONLINE -- the one command that recovers the instance\r\n    // refusing on the very state that is wrong. The PAIR test (same pid,\r\n    // different start time) is what tells a reused pid from our own process.\r\n    //\r\n    // The stance is unchanged where there is no stamp to compare: an instance\r\n    // launched before this field existed, or a backend that exposed no start\r\n    // time, still falls back to plain absence and fails toward ALIVE.\r\n    // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\n    let birth = read_shell_launch(shell_perch).and_then(|l| l.pid_started_at);\r\n    // A relaunch can publish between these reads. Do not apply the new birth\r\n    // to the old PID; an unstable snapshot is not proof that the shell died.\r\n    if read_shell_pid(shell_perch) != Some(pid) {\r\n        return false;\r\n    }\r\n    matches!(\r\n        crate::liveness::relay_liveness(Some(pid), birth),\r\n        crate::liveness::RelayLiveness::Gone\r\n    )\r\n}\r\n\r\n/// A shell instance's **derived** status — the shell-side twin of\r\n/// `liveness::is_perch_alive` (REQ-HAZARD-DAEMON-HOSTED-LIVENESS gave agent\r\n/// perches exactly this; shells never got it).\r\n///\r\n/// The recorded `status` has exactly two writers — the bind handshake (→online)\r\n/// and `spt_daemon::shellhost::close_shell` (→offline) — and `close_shell` runs\r\n/// on the **link-break** path alone. A binary that dies\r\n/// abruptly (force-kill, crash, OOM) breaks no link, so nothing ever writes the\r\n/// flip and the record says `online` forever. Deriving closes that hole without\r","truncated":true,"truncatedBy":"lines","totalLines":605,"totalBytes":5059,"outputLines":98,"outputBytes":5059,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":762,"displayContent":{"text":"}\r\n\r\n/// Record the birth stamp for a just-launched `pid`. Best-effort: a stamp that\r\n/// cannot be written must never fail the launch — a missing stamp degrades the\r\n/// probe to its pre-existing fail-toward-alive behaviour, which is the safe\r\n/// direction, while a failed launch is a real outage.\r\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\npub fn record_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) {\r\n    let launch = ShellLaunch {\r\n        pid_started_at: crate::proc::process_started_at(pid),\r\n        launched_ms: now_ms,\r\n    };\r\n    if let Ok(json) = serde_json::to_string_pretty(&launch) {\r\n        let _ = atomic_write_string(&shell_perch.join(SHELL_LAUNCH_FILE), &json);\r\n    }\r\n}\r\n\r\n/// Retire the previous birth before spawning a replacement. Failure must be\r\n/// reported before a child exists, not after creating an unrecorded process.\r\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\npub fn prepare_shell_launch(shell_perch: &Path) -> std::io::Result<()> {\r\n    match std::fs::remove_file(shell_perch.join(SHELL_LAUNCH_FILE)) {\r\n        Ok(()) => {}\r\n        Err(e) if e.kind() == std::io::ErrorKind::NotFound => {}\r\n        Err(e) => return Err(e),\r\n    }\r\n    Ok(())\r\n}\r\n\r\n/// Publish a prepared launch. The missing stamp preserves live-PID custody\r\n/// until the new birth is available; it can never name the previous process.\r\n// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\npub fn publish_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) -> std::io::Result<()> {\r\n    atomic_write_string(&shell_perch.join(SHELL_PID_FILE), &pid.to_string())?;\r\n    record_shell_launch(shell_perch, pid, now_ms);\r\n    Ok(())\r\n}\r\n\r\n/// Read the birth stamp, if this instance was launched by a build that parks one.\r\n/// Absent or unparseable ⇒ `None`, which every consumer treats as \"no stamp to\r\n/// compare\" rather than as a fault.\r\npub fn read_shell_launch(shell_perch: &Path) -> Option<ShellLaunch> {\r\n    let raw = std::fs::read_to_string(shell_perch.join(SHELL_LAUNCH_FILE)).ok()?;\r\n    serde_json::from_str(&raw).ok()\r\n}\r\n\r\n/// **The** discriminant behind [`effective_status`]: is this instance's recorded\r\n/// binary *provably* gone? Narrow on purpose — true only when the perch parks a\r\n/// pid, it parses non-zero, and the process is not alive.\r\n///\r\n/// Everything else reads **alive**: an absent or unparseable pid file, and the\r\n/// `0` a broker-hosted spawn records when its backend exposed no pid\r\n/// ([`crate::proc::is_process_alive`] already reads 0 as dead, so the explicit\r\n/// filter here is what keeps a pid-less backend from being declared dead by this\r\n/// resolver). That fail-toward-alive stance mirrors `liveness.rs`'s interim\r\n/// parity: this predicate exists to catch a *corpse*, never to guess at a\r\n/// backend it cannot see.\r\n///\r\n/// Recycled pids read alive, so a heal can be **missed** but never **mis-fired**\r\n/// — the failure direction is \"stays stale\", never \"tears down a live instance\".\r\n// [impl->REQ-HAZARD-SHELL-STALE-ONLINE]\r\npub fn shell_pid_provably_dead(shell_perch: &Path) -> bool {\r\n    let Some(pid) = read_shell_pid(shell_perch).filter(|&pid| pid != 0) else {\r\n        return false;\r\n    };\r\n    // BAROMETER W2 (releases#78): probe pid AND birth stamp, not the pid alone.\r\n    // A pid namespace resets at boot, so a corpse pid can be REUSED by an\r\n    // unrelated process started after the restart; under bare-pid liveness that\r\n    // reads ALIVE, the derived status reads online, and `relink` then refuses\r\n    // SHELL_ALREADY_ONLINE -- the one command that recovers the instance\r\n    // refusing on the very state that is wrong. The PAIR test (same pid,\r\n    // different start time) is what tells a reused pid from our own process.\r\n    //\r\n    // The stance is unchanged where there is no stamp to compare: an instance\r\n    // launched before this field existed, or a backend that exposed no start\r\n    // time, still falls back to plain absence and fails toward ALIVE.\r\n    // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\n    let birth = read_shell_launch(shell_perch).and_then(|l| l.pid_started_at);\r\n    // A relaunch can publish between these reads. Do not apply the new birth\r\n    // to the old PID; an unstable snapshot is not proof that the shell died.\r\n    if read_shell_pid(shell_perch) != Some(pid) {\r\n        return false;\r\n    }\r\n    matches!(\r\n        crate::liveness::relay_liveness(Some(pid), birth),\r\n        crate::liveness::RelayLiveness::Gone\r\n    )\r\n}\r\n\r\n/// A shell instance's **derived** status — the shell-side twin of\r\n/// `liveness::is_perch_alive` (REQ-HAZARD-DAEMON-HOSTED-LIVENESS gave agent\r\n/// perches exactly this; shells never got it).\r\n///\r\n/// The recorded `status` has exactly two writers — the bind handshake (→online)\r\n/// and `spt_daemon::shellhost::close_shell` (→offline) — and `close_shell` runs\r\n/// on the **link-break** path alone. A binary that dies\r\n/// abruptly (force-kill, crash, OOM) breaks no link, so nothing ever writes the\r\n/// flip and the record says `online` forever. Deriving closes that hole without\r","startLine":158,"lineNumbers":[158,159,160,161,162,163,164,165,166,167,168,169,170,171,172,173,174,175,176,177,178,179,180,181,182,183,184,185,186,187,188,189,190,191,192,193,194,195,196,197,198,199,200,201,202,203,204,205,206,207,208,209,210,211,212,213,214,215,216,217,218,219,220,221,222,223,224,225,226,227,228,229,230,231,232,233,234,235,236,237,238,239,240,241,242,243,244,245,246,247,248,249,250,251,252,253,254,255]},"fileSize":35173,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\crates\\spt-store\\src\\shellinfo.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":762,"totalBytes":5059,"outputLines":98,"outputBytes":5059,"shownRange":{"start":158,"end":255},"nextOffset":256}}},"isError":false,"timestamp":1789001927376}}
{"type":"message","id":"5cfd2b89","parentId":"ab472c4d","timestamp":"2026-09-10T00:58:47.377Z","message":{"role":"toolResult","toolCallId":"call_V4cksvh82Mygkj6neHanDF8Z|fc_028129a7f1d66e9c016aa200c73f3087d28da9382ced434229","toolName":"read","content":[{"type":"text","text":"diff --git a/crates/spt-daemon/src/shellhost.rs b/crates/spt-daemon/src/shellhost.rs\nindex 44df752b..a930a4a2 100644\n--- a/crates/spt-daemon/src/shellhost.rs\n+++ b/crates/spt-daemon/src/shellhost.rs\n@@ -477,6 +477,9 @@ pub fn launch_shell(\n             let tokens =\n                 fill_spawn_command(owlery, owner, shell_id, adapter_name, install_dir, shell)?;\n             let (program, args) = tokens.split_first().expect(\"non-empty checked\");\n+            let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n+            spt_store::shellinfo::prepare_shell_launch(&perch)\n+                .map_err(|e| format!(\"prepare launch: {e}\"))?;\n             // The shell binary is a LONG-LIVED detached child (KNOWN-HAZARDS\n             // 5.6): on Windows it must inherit NO handles, or a caller\n             // capturing `spt shell spawn`'s output — at any depth of the\n@@ -495,14 +498,10 @@ pub fn launch_shell(\n                 .spawn()\n                 .map_err(|e| format!(\"spawn {program}: {e}\"))?\n                 .id();\n-            let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n-            std::fs::write(perch.join(SHELL_PID_FILE), pid.to_string())\n-                .map_err(|e| format!(\"record pid: {e}\"))?;\n-            // The birth stamp rides WITH the pid at every launch, or the probe that\n-            // reads it degrades to bare-pid liveness on this instance alone -- a\n-            // silent per-instance regression rather than a visible one.\n+            // Publish only after the previous birth was retired before spawn.\n             // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n-            spt_store::shellinfo::record_shell_launch(&perch, pid, now_ms());\n+            spt_store::shellinfo::publish_shell_launch(&perch, pid, now_ms())\n+                .map_err(|e| format!(\"record launch: {e}\"))?;\n             Ok(pid)\n         }\n     }\n@@ -527,6 +526,9 @@ pub fn launch_shell_brokered_in(\n     let (program, args) = tokens.split_first().expect(\"non-empty checked\");\n     let mut brain =\n         Brain::cold_start(broker_name, now_ms()).map_err(|e| format!(\"broker connect: {e}\"))?;\n+    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n+    spt_store::shellinfo::prepare_shell_launch(&perch)\n+        .map_err(|e| format!(\"prepare launch: {e}\"))?;\n     let (_sid, pid) = brain\n         .spawn_session_pid(SpawnReq {\n             program: program.clone(),\n@@ -554,14 +556,10 @@ pub fn launch_shell_brokered_in(\n         })\n         .map_err(|e| format!(\"broker spawn: {e}\"))?;\n     let pid = pid.unwrap_or(0);\n-    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n-    std::fs::write(perch.join(SHELL_PID_FILE), pid.to_string())\n-        .map_err(|e| format!(\"record pid: {e}\"))?;\n-    // Same stamp at the broker-hosted launch. A backend that exposed no pid\n-    // records 0 here; record_shell_launch stamps no start-time for it, so the\n-    // probe keeps its fail-toward-alive stance for pid-less backends.\n+    // PID-less backends keep their fail-toward-alive behavior.\n     // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n-    spt_store::shellinfo::record_shell_launch(&perch, pid, now_ms());\n+    spt_store::shellinfo::publish_shell_launch(&perch, pid, now_ms())\n+        .map_err(|e| format!(\"record launch: {e}\"))?;\n     Ok(pid)\n }\n \ndiff --git a/crates/spt-daemon/src/shellwake.rs b/crates/spt-daemon/src/shellwake.rs\nindex d9631b4d..72ed57cf 100644\n--- a/crates/spt-daemon/src/shellwake.rs\n+++ b/crates/spt-daemon/src/shellwake.rs\n@@ -406,13 +406,9 @@ pub fn resolve_wake(\n     // Never double-launch a live binary: the revive's wake cascade may have\n     // relaunched this (persistent) instance already — and two reconcilers\n     // (a fresh daemon adopting + a stale watcher firing) may race a wake.\n-    // A parked pid that probes alive means somebody won; stand down.\n+    // A live, birth-matching launch wins even before its bind handshake.\n     let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n-    let live_pid = std::fs::read_to_string(perch.join(crate::shellhost::SHELL_PID_FILE))\n-        .ok()\n-        .and_then(|s| s.trim().parse::<u32>().ok())\n-        .filter(|&p| p != 0 && spt_store::proc::is_process_alive(p));\n-    if let Some(p) = live_pid {\n+    if let Some(p) = crate::shellhost::live_launch_winner(&perch) {\n         return Ok(format!(\"{did}already relaunched pid={p} (online at bind)\"));\n     }\n     // [impl->REQ-INSTALL-11] the wake-triggered relaunch resolves against the\n@@ -1525,6 +1521,38 @@ mod tests {\n         );\n     }\n \n+    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n+    #[test]\n+    fn resolve_wake_does_not_adopt_a_recycled_pid_as_a_launch() {\n+        let tmp = tempfile::tempdir().unwrap();\n+        let owlery = tmp.path();\n+        seed_owner(owlery, \"ling\");\n+        let id = spawn_record(owlery, \"ling\", \"mock-wake\", None).unwrap();\n+        let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, \"ling\", &id);\n+        let pid = std::process::id();\n+        std::fs::write(perch.join(shellinfo::SHELL_PID_FILE), pid.to_string()).unwrap();\n+        shellinfo::record_shell_launch(&perch, pid, 1_000);\n+        let mut launch = shellinfo::read_shell_launch(&perch).unwrap();\n+        launch.pid_started_at = Some(\n+            launch.pid_started_at.expect(\"native birth on supported platforms\") ^ 1,\n+        );\n+        std::fs::write(\n+            perch.join(shellinfo::SHELL_LAUNCH_FILE),\n+            serde_json::to_string(&launch).unwrap(),\n+        )\n+        .unwrap();\n+        // A real spawn failure must surface, not a successful adoption of the\n+        // unrelated live process. No child needs cleanup on this error path.\n+        let missing = tmp.path().join(\"missing-shell-executable\");\n+        let command = format!(\"\\\"{}\\\"\", missing.display());\n+        let err = resolve_wake(\n+            owlery, \"ling\", &id, \"mock-wake\", &shell_section(&command, \"persistent = true\"),\n+        )\n+        .unwrap_err();\n+        assert!(err.contains(\"missing-shell-executable\"), \"{err}\");\n+        assert!(spt_store::proc::is_process_alive(pid));\n+    }\n+\n     // [unit->REQ-SHELL-2] wake resolution, the dormant branch: the owner is\n     // resting warm — touch nothing on the endpoint, just relaunch the shell.\n     #[test]\ndiff --git a/crates/spt-store/src/shellinfo.rs b/crates/spt-store/src/shellinfo.rs\nindex f3aa9086..f860e598 100644\n--- a/crates/spt-store/src/shellinfo.rs\n+++ b/crates/spt-store/src/shellinfo.rs\n@@ -172,6 +172,27 @@ pub fn record_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) {\n     }\n }\n \n+/// Retire the previous birth before spawning a replacement. Failure must be\n+/// reported before a child exists, not after creating an unrecorded process.\n+// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n+pub fn prepare_shell_launch(shell_perch: &Path) -> std::io::Result<()> {\n+    match std::fs::remove_file(shell_perch.join(SHELL_LAUNCH_FILE)) {\n+        Ok(()) => {}\n+        Err(e) if e.kind() == std::io::ErrorKind::NotFound => {}\n+        Err(e) => return Err(e),\n+    }\n+    Ok(())\n+}\n+\n+/// Publish a prepared launch. The missing stamp preserves live-PID custody\n+/// until the new birth is available; it can never name the previous process.\n+// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n+pub fn publish_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) -> std::io::Result<()> {\n+    atomic_write_string(&shell_perch.join(SHELL_PID_FILE), &pid.to_string())?;\n+    record_shell_launch(shell_perch, pid, now_ms);\n+    Ok(())\n+}\n+\n /// Read the birth stamp, if this instance was launched by a build that parks one.\n /// Absent or unparseable ⇒ `None`, which every consumer treats as \"no stamp to\n /// compare\" rather than as a fault.\n@@ -212,6 +233,11 @@ pub fn shell_pid_provably_dead(shell_perch: &Path) -> bool {\n     // time, still falls back to plain absence and fails toward ALIVE.\n     // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n     let birth = read_shell_launch(shell_perch).and_then(|l| l.pid_started_at);\n+    // A relaunch can publish between these reads. Do not apply the new birth\n+    // to the old PID; an unstable snapshot is not proof that the shell died.\n+    if read_shell_pid(shell_perch) != Some(pid) {\n+        return false;\n+    }\n     matches!(\n         crate::liveness::relay_liveness(Some(pid), birth),\n         crate::liveness::RelayLiveness::Gone\n@@ -504,6 +530,15 @@ mod tests {\n         assert_eq!(status_of(&d), SHELL_STATUS_ONLINE);\n     }\n \n+    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n+    #[test]\n+    fn launch_publication_refuses_an_unretirable_old_stamp() {\n+        let d = online_perch(Some(\"123\"));\n+        std::fs::create_dir(d.path().join(SHELL_LAUNCH_FILE)).unwrap();\n+        prepare_shell_launch(d.path()).unwrap_err();\n+        assert_eq!(read_shell_pid(d.path()), Some(123));\n+    }\n+\n     // [unit->REQ-HAZARD-SHELL-STALE-ONLINE] the whole discriminant matrix. The\n     // ONE case that derives offline is a parked, non-zero, dead pid — flynn's\n     // field shape: the binary was force-killed, no link broke, so `close_shell`"}],"details":{"resolvedPath":"C:\\Users\\decid\\.omp\\agent\\sessions\\-Documents-projects-spt-core\\2026-09-09T23-45-42-317Z_01a08890-21ad-7000-8f54-208d4f4f2f2c\\135.bash-original.log","contentType":"text/plain","totalLines":179,"displayContent":{"text":"diff --git a/crates/spt-daemon/src/shellhost.rs b/crates/spt-daemon/src/shellhost.rs\nindex 44df752b..a930a4a2 100644\n--- a/crates/spt-daemon/src/shellhost.rs\n+++ b/crates/spt-daemon/src/shellhost.rs\n@@ -477,6 +477,9 @@ pub fn launch_shell(\n             let tokens =\n                 fill_spawn_command(owlery, owner, shell_id, adapter_name, install_dir, shell)?;\n             let (program, args) = tokens.split_first().expect(\"non-empty checked\");\n+            let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n+            spt_store::shellinfo::prepare_shell_launch(&perch)\n+                .map_err(|e| format!(\"prepare launch: {e}\"))?;\n             // The shell binary is a LONG-LIVED detached child (KNOWN-HAZARDS\n             // 5.6): on Windows it must inherit NO handles, or a caller\n             // capturing `spt shell spawn`'s output — at any depth of the\n@@ -495,14 +498,10 @@ pub fn launch_shell(\n                 .spawn()\n                 .map_err(|e| format!(\"spawn {program}: {e}\"))?\n                 .id();\n-            let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n-            std::fs::write(perch.join(SHELL_PID_FILE), pid.to_string())\n-                .map_err(|e| format!(\"record pid: {e}\"))?;\n-            // The birth stamp rides WITH the pid at every launch, or the probe that\n-            // reads it degrades to bare-pid liveness on this instance alone -- a\n-            // silent per-instance regression rather than a visible one.\n+            // Publish only after the previous birth was retired before spawn.\n             // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n-            spt_store::shellinfo::record_shell_launch(&perch, pid, now_ms());\n+            spt_store::shellinfo::publish_shell_launch(&perch, pid, now_ms())\n+                .map_err(|e| format!(\"record launch: {e}\"))?;\n             Ok(pid)\n         }\n     }\n@@ -527,6 +526,9 @@ pub fn launch_shell_brokered_in(\n     let (program, args) = tokens.split_first().expect(\"non-empty checked\");\n     let mut brain =\n         Brain::cold_start(broker_name, now_ms()).map_err(|e| format!(\"broker connect: {e}\"))?;\n+    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n+    spt_store::shellinfo::prepare_shell_launch(&perch)\n+        .map_err(|e| format!(\"prepare launch: {e}\"))?;\n     let (_sid, pid) = brain\n         .spawn_session_pid(SpawnReq {\n             program: program.clone(),\n@@ -554,14 +556,10 @@ pub fn launch_shell_brokered_in(\n         })\n         .map_err(|e| format!(\"broker spawn: {e}\"))?;\n     let pid = pid.unwrap_or(0);\n-    let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n-    std::fs::write(perch.join(SHELL_PID_FILE), pid.to_string())\n-        .map_err(|e| format!(\"record pid: {e}\"))?;\n-    // Same stamp at the broker-hosted launch. A backend that exposed no pid\n-    // records 0 here; record_shell_launch stamps no start-time for it, so the\n-    // probe keeps its fail-toward-alive stance for pid-less backends.\n+    // PID-less backends keep their fail-toward-alive behavior.\n     // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n-    spt_store::shellinfo::record_shell_launch(&perch, pid, now_ms());\n+    spt_store::shellinfo::publish_shell_launch(&perch, pid, now_ms())\n+        .map_err(|e| format!(\"record launch: {e}\"))?;\n     Ok(pid)\n }\n \ndiff --git a/crates/spt-daemon/src/shellwake.rs b/crates/spt-daemon/src/shellwake.rs\nindex d9631b4d..72ed57cf 100644\n--- a/crates/spt-daemon/src/shellwake.rs\n+++ b/crates/spt-daemon/src/shellwake.rs\n@@ -406,13 +406,9 @@ pub fn resolve_wake(\n     // Never double-launch a live binary: the revive's wake cascade may have\n     // relaunched this (persistent) instance already — and two reconcilers\n     // (a fresh daemon adopting + a stale watcher firing) may race a wake.\n-    // A parked pid that probes alive means somebody won; stand down.\n+    // A live, birth-matching launch wins even before its bind handshake.\n     let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, owner, shell_id);\n-    let live_pid = std::fs::read_to_string(perch.join(crate::shellhost::SHELL_PID_FILE))\n-        .ok()\n-        .and_then(|s| s.trim().parse::<u32>().ok())\n-        .filter(|&p| p != 0 && spt_store::proc::is_process_alive(p));\n-    if let Some(p) = live_pid {\n+    if let Some(p) = crate::shellhost::live_launch_winner(&perch) {\n         return Ok(format!(\"{did}already relaunched pid={p} (online at bind)\"));\n     }\n     // [impl->REQ-INSTALL-11] the wake-triggered relaunch resolves against the\n@@ -1525,6 +1521,38 @@ mod tests {\n         );\n     }\n \n+    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n+    #[test]\n+    fn resolve_wake_does_not_adopt_a_recycled_pid_as_a_launch() {\n+        let tmp = tempfile::tempdir().unwrap();\n+        let owlery = tmp.path();\n+        seed_owner(owlery, \"ling\");\n+        let id = spawn_record(owlery, \"ling\", \"mock-wake\", None).unwrap();\n+        let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, \"ling\", &id);\n+        let pid = std::process::id();\n+        std::fs::write(perch.join(shellinfo::SHELL_PID_FILE), pid.to_string()).unwrap();\n+        shellinfo::record_shell_launch(&perch, pid, 1_000);\n+        let mut launch = shellinfo::read_shell_launch(&perch).unwrap();\n+        launch.pid_started_at = Some(\n+            launch.pid_started_at.expect(\"native birth on supported platforms\") ^ 1,\n+        );\n+        std::fs::write(\n+            perch.join(shellinfo::SHELL_LAUNCH_FILE),\n+            serde_json::to_string(&launch).unwrap(),\n+        )\n+        .unwrap();\n+        // A real spawn failure must surface, not a successful adoption of the\n+        // unrelated live process. No child needs cleanup on this error path.\n+        let missing = tmp.path().join(\"missing-shell-executable\");\n+        let command = format!(\"\\\"{}\\\"\", missing.display());\n+        let err = resolve_wake(\n+            owlery, \"ling\", &id, \"mock-wake\", &shell_section(&command, \"persistent = true\"),\n+        )\n+        .unwrap_err();\n+        assert!(err.contains(\"missing-shell-executable\"), \"{err}\");\n+        assert!(spt_store::proc::is_process_alive(pid));\n+    }\n+\n     // [unit->REQ-SHELL-2] wake resolution, the dormant branch: the owner is\n     // resting warm — touch nothing on the endpoint, just relaunch the shell.\n     #[test]\ndiff --git a/crates/spt-store/src/shellinfo.rs b/crates/spt-store/src/shellinfo.rs\nindex f3aa9086..f860e598 100644\n--- a/crates/spt-store/src/shellinfo.rs\n+++ b/crates/spt-store/src/shellinfo.rs\n@@ -172,6 +172,27 @@ pub fn record_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) {\n     }\n }\n \n+/// Retire the previous birth before spawning a replacement. Failure must be\n+/// reported before a child exists, not after creating an unrecorded process.\n+// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n+pub fn prepare_shell_launch(shell_perch: &Path) -> std::io::Result<()> {\n+    match std::fs::remove_file(shell_perch.join(SHELL_LAUNCH_FILE)) {\n+        Ok(()) => {}\n+        Err(e) if e.kind() == std::io::ErrorKind::NotFound => {}\n+        Err(e) => return Err(e),\n+    }\n+    Ok(())\n+}\n+\n+/// Publish a prepared launch. The missing stamp preserves live-PID custody\n+/// until the new birth is available; it can never name the previous process.\n+// [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n+pub fn publish_shell_launch(shell_perch: &Path, pid: u32, now_ms: u64) -> std::io::Result<()> {\n+    atomic_write_string(&shell_perch.join(SHELL_PID_FILE), &pid.to_string())?;\n+    record_shell_launch(shell_perch, pid, now_ms);\n+    Ok(())\n+}\n+\n /// Read the birth stamp, if this instance was launched by a build that parks one.\n /// Absent or unparseable ⇒ `None`, which every consumer treats as \"no stamp to\n /// compare\" rather than as a fault.\n@@ -212,6 +233,11 @@ pub fn shell_pid_provably_dead(shell_perch: &Path) -> bool {\n     // time, still falls back to plain absence and fails toward ALIVE.\n     // [impl->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n     let birth = read_shell_launch(shell_perch).and_then(|l| l.pid_started_at);\n+    // A relaunch can publish between these reads. Do not apply the new birth\n+    // to the old PID; an unstable snapshot is not proof that the shell died.\n+    if read_shell_pid(shell_perch) != Some(pid) {\n+        return false;\n+    }\n     matches!(\n         crate::liveness::relay_liveness(Some(pid), birth),\n         crate::liveness::RelayLiveness::Gone\n@@ -504,6 +530,15 @@ mod tests {\n         assert_eq!(status_of(&d), SHELL_STATUS_ONLINE);\n     }\n \n+    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n+    #[test]\n+    fn launch_publication_refuses_an_unretirable_old_stamp() {\n+        let d = online_perch(Some(\"123\"));\n+        std::fs::create_dir(d.path().join(SHELL_LAUNCH_FILE)).unwrap();\n+        prepare_shell_launch(d.path()).unwrap_err();\n+        assert_eq!(read_shell_pid(d.path()), Some(123));\n+    }\n+\n     // [unit->REQ-HAZARD-SHELL-STALE-ONLINE] the whole discriminant matrix. The\n     // ONE case that derives offline is a parked, non-zero, dead pid — flynn's\n     // field shape: the binary was force-killed, no link broke, so `close_shell`","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120,121,122,123,124,125,126,127,128,129,130,131,132,133,134,135,136,137,138,139,140,141,142,143,144,145,146,147,148,149,150,151,152,153,154,155,156,157,158,159,160,161,162,163,164,165,166,167,168,169,170,171,172,173,174,175,176,177,178,179]},"meta":{"source":{"type":"internal","value":"artifact://135"}}},"isError":false,"timestamp":1789001927377}}
{"type":"message","id":"1a387736","parentId":"5cfd2b89","timestamp":"2026-09-10T00:58:47.396Z","message":{"role":"toolResult","toolCallId":"call_8GsznHhOsDQtr6UdyOJVbpbx|fc_028129a7f1d66e9c016aa200c73f5087d2969e2d031b76a828","toolName":"read","content":[{"type":"text","text":"            \"the deferral is named: {err}\"\n        );\n    }\n\n    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\n    #[test]\n    fn resolve_wake_does_not_adopt_a_recycled_pid_as_a_launch() {\n        let tmp = tempfile::tempdir().unwrap();\n        let owlery = tmp.path();\n        seed_owner(owlery, \"ling\");\n        let id = spawn_record(owlery, \"ling\", \"mock-wake\", None).unwrap();\n        let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, \"ling\", &id);\n        let pid = std::process::id();\n        std::fs::write(perch.join(shellinfo::SHELL_PID_FILE), pid.to_string()).unwrap();\n        shellinfo::record_shell_launch(&perch, pid, 1_000);\n        let mut launch = shellinfo::read_shell_launch(&perch).unwrap();\n        launch.pid_started_at = Some(\n            launch.pid_started_at.expect(\"native birth on supported platforms\") ^ 1,\n        );\n        std::fs::write(\n            perch.join(shellinfo::SHELL_LAUNCH_FILE),\n            serde_json::to_string(&launch).unwrap(),\n        )\n        .unwrap();\n        // A real spawn failure must surface, not a successful adoption of the\n        // unrelated live process. No child needs cleanup on this error path.\n        let missing = tmp.path().join(\"missing-shell-executable\");\n        let command = format!(\"\\\"{}\\\"\", missing.display());\n        let err = resolve_wake(\n            owlery, \"ling\", &id, \"mock-wake\", &shell_section(&command, \"persistent = true\"),\n        )\n        .unwrap_err();\n        assert!(err.contains(\"missing-shell-executable\"), \"{err}\");\n        assert!(spt_store::proc::is_process_alive(pid));\n    }\n\n    // [unit->REQ-SHELL-2] wake resolution, the dormant branch: the owner is\n    // resting warm — touch nothing on the endpoint, just relaunch the shell.\n    #[test]\n    fn resolve_wake_leaves_a_dormant_owner_and_relaunches() {\n        use spt_store::info::{write_info, InfoJson};\n        let tmp = tempfile::tempdir().unwrap();\n\n[Showing lines 1520-1561 of 2190. Use :1562 to continue]"}],"details":{"truncation":{"content":"            \"the deferral is named: {err}\"\r\n        );\r\n    }\r\n\r\n    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\n    #[test]\r\n    fn resolve_wake_does_not_adopt_a_recycled_pid_as_a_launch() {\r\n        let tmp = tempfile::tempdir().unwrap();\r\n        let owlery = tmp.path();\r\n        seed_owner(owlery, \"ling\");\r\n        let id = spawn_record(owlery, \"ling\", \"mock-wake\", None).unwrap();\r\n        let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, \"ling\", &id);\r\n        let pid = std::process::id();\r\n        std::fs::write(perch.join(shellinfo::SHELL_PID_FILE), pid.to_string()).unwrap();\r\n        shellinfo::record_shell_launch(&perch, pid, 1_000);\r\n        let mut launch = shellinfo::read_shell_launch(&perch).unwrap();\r\n        launch.pid_started_at = Some(\r\n            launch.pid_started_at.expect(\"native birth on supported platforms\") ^ 1,\r\n        );\r\n        std::fs::write(\r\n            perch.join(shellinfo::SHELL_LAUNCH_FILE),\r\n            serde_json::to_string(&launch).unwrap(),\r\n        )\r\n        .unwrap();\r\n        // A real spawn failure must surface, not a successful adoption of the\r\n        // unrelated live process. No child needs cleanup on this error path.\r\n        let missing = tmp.path().join(\"missing-shell-executable\");\r\n        let command = format!(\"\\\"{}\\\"\", missing.display());\r\n        let err = resolve_wake(\r\n            owlery, \"ling\", &id, \"mock-wake\", &shell_section(&command, \"persistent = true\"),\r\n        )\r\n        .unwrap_err();\r\n        assert!(err.contains(\"missing-shell-executable\"), \"{err}\");\r\n        assert!(spt_store::proc::is_process_alive(pid));\r\n    }\r\n\r\n    // [unit->REQ-SHELL-2] wake resolution, the dormant branch: the owner is\r\n    // resting warm — touch nothing on the endpoint, just relaunch the shell.\r\n    #[test]\r\n    fn resolve_wake_leaves_a_dormant_owner_and_relaunches() {\r\n        use spt_store::info::{write_info, InfoJson};\r\n        let tmp = tempfile::tempdir().unwrap();\r","truncated":true,"truncatedBy":"lines","totalLines":671,"totalBytes":1968,"outputLines":42,"outputBytes":1968,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":2190,"displayContent":{"text":"            \"the deferral is named: {err}\"\r\n        );\r\n    }\r\n\r\n    // [unit->REQ-SHELL-PERSISTENT-BOOT-RESTORE]\r\n    #[test]\r\n    fn resolve_wake_does_not_adopt_a_recycled_pid_as_a_launch() {\r\n        let tmp = tempfile::tempdir().unwrap();\r\n        let owlery = tmp.path();\r\n        seed_owner(owlery, \"ling\");\r\n        let id = spawn_record(owlery, \"ling\", \"mock-wake\", None).unwrap();\r\n        let perch = spt_store::perch::resolve_shell_perch_path_in(owlery, \"ling\", &id);\r\n        let pid = std::process::id();\r\n        std::fs::write(perch.join(shellinfo::SHELL_PID_FILE), pid.to_string()).unwrap();\r\n        shellinfo::record_shell_launch(&perch, pid, 1_000);\r\n        let mut launch = shellinfo::read_shell_launch(&perch).unwrap();\r\n        launch.pid_started_at = Some(\r\n            launch.pid_started_at.expect(\"native birth on supported platforms\") ^ 1,\r\n        );\r\n        std::fs::write(\r\n            perch.join(shellinfo::SHELL_LAUNCH_FILE),\r\n            serde_json::to_string(&launch).unwrap(),\r\n        )\r\n        .unwrap();\r\n        // A real spawn failure must surface, not a successful adoption of the\r\n        // unrelated live process. No child needs cleanup on this error path.\r\n        let missing = tmp.path().join(\"missing-shell-executable\");\r\n        let command = format!(\"\\\"{}\\\"\", missing.display());\r\n        let err = resolve_wake(\r\n            owlery, \"ling\", &id, \"mock-wake\", &shell_section(&command, \"persistent = true\"),\r\n        )\r\n        .unwrap_err();\r\n        assert!(err.contains(\"missing-shell-executable\"), \"{err}\");\r\n        assert!(spt_store::proc::is_process_alive(pid));\r\n    }\r\n\r\n    // [unit->REQ-SHELL-2] wake resolution, the dormant branch: the owner is\r\n    // resting warm — touch nothing on the endpoint, just relaunch the shell.\r\n    #[test]\r\n    fn resolve_wake_leaves_a_dormant_owner_and_relaunches() {\r\n        use spt_store::info::{write_info, InfoJson};\r\n        let tmp = tempfile::tempdir().unwrap();\r","startLine":1520,"lineNumbers":[1520,1521,1522,1523,1524,1525,1526,1527,1528,1529,1530,1531,1532,1533,1534,1535,1536,1537,1538,1539,1540,1541,1542,1543,1544,1545,1546,1547,1548,1549,1550,1551,1552,1553,1554,1555,1556,1557,1558,1559,1560,1561]},"fileSize":102198,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.worktrees\\287-shellrestore\\crates\\spt-daemon\\src\\shellwake.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":2190,"totalBytes":1968,"outputLines":42,"outputBytes":1968,"shownRange":{"start":1520,"end":1561},"nextOffset":1562}}},"isError":false,"timestamp":1789001927395}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_jjvUlqpd3jIu1xU2ioFfr5Y9|fc_028129a7f1d66e9c016aa200d0e14887d2812831148e458dbc","toolName":"hub","startedAt":"2026-09-10T00:58:59.883Z","intent":"Reporting preflight identity loss"},"id":"f124d21d","parentId":"1a387736","timestamp":"2026-09-10T00:58:59.883Z"}
{"type":"message","id":"f837fe0f","parentId":"f124d21d","timestamp":"2026-09-10T00:58:59.884Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":"**Flagging transactional identity**"},{"type":"toolCall","id":"call_jjvUlqpd3jIu1xU2ioFfr5Y9|fc_028129a7f1d66e9c016aa200d0e14887d2812831148e458dbc","name":"hub","arguments":{"i":"Reporting preflight identity loss","op":"send","to":"Main","message":"New preflight fixes orphan-on-retirement-error but deletes the only proof that an old PID is recycled BEFORE a fallible spawn. In your new missing-executable wake fixture, first resolve_wake errors after prepare removes the mismatching stamp; second resolve_wake now adopts the unrelated live PID (boot/owner-edge also skip it). Worse, kill_shell_at treats the remaining bare old PID as Held and can kill that unrelated process. Preserve old PID+birth evidence on failed spawn, and don't expose bare recycled old PID during preflight; coherent identity publication is preferable to deleting birth while retaining old PID. Source-only, not run."},"intent":"Reporting preflight identity loss"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":5007,"output":261,"cacheRead":119296,"cacheWrite":0,"totalTokens":124564,"reasoningTokens":99,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1789001927461,"responseId":"resp_028129a7f1d66e9c016aa200c7bf7487d2a0b18eedc27a317b","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_028129a7f1d66e9c016aa200cbfecc87d2815d1eb9cd6287cd","type":"reasoning","content":[],"encrypted_content":"gAAAAABqogDQyy7N2yTkK7S6nTe-HLY8B_V9DS00piF57PZtldEFlJ6TiNLnKJScgxUAkp-2AEOSMAuacIdhHcWl9jBRmRiI2uG6WyuhuKQXgwfzOxtkyeWTD8ytESBPdtVvlEv12v39H5WjhQN4_I7a0ozaFrFht5Oy8wVD5Xkxtlrojr3TjZD7-t2CGpXQlFl18ZIRTRwKStjY4esML4UJdfR83C9atwbjEPy_Tx-4-Ymi-Ij7KdjyQN2KokFdweIP3eYUY-RJuwF574yBJ8xUbf9egbAs6-gDIGd-NiXcOkdBag7ur0oXE3pBC_h_KRDWMSw96FKv-Dy7_2mPyX0fVFFAJ22LcRs3gwov--t8gsW9MSTSoN2Jzv1qdO1mf7KQRKeCqdmC_oenQSIJLwG1c0iowxoKh9X8Vn6f7H5kM1vnZihh1VHA7KF4ZVj0Ot66UUwjxPGDETjzV7BbOxnmnrPfBgQm64p72QLKUTo8fpXe1KO9d6UYnF9Li4F6DxkBcYpOcHSQRNm_Pp81vmwOxAaxaJBuS3oJ1r6vQATFci4XetVvPHFNqVbtUViVmnTEOJwILxwLlln2nIbXySSTEVnlu5u0gC6OtIwrvToNwpeB92G51Hn13a-7i8zGIquE0_gF1pQxBCgxoQQms190t1dDZ5NOyFFa3bXi0TG-xHLtako8sTH98T5mcg_TSNmmxpYrTHrqoT6rWH_6vjBEyTuHVYp2VfCNi4kFgDgnd4MsMPaef9jC7y9z9EOGDcAFuy8zeXGJL8qlYLmDi2xZ0VlSMVd51uXNOm85RDWnL_p3L0U8twNKZyTWpkSomCW1oGQxbAT-Hx6ILy7G4A7Ef2t1dE4SVLhmAsRBYTlFMeYyRlCm5uyLqhbiZFGLB4PWM3uJIPN_UDMNcFWeWjF_EtGEiw_znW9ngcCrcc8uFpHh3WQLIsidk9aQBdWSfyQ7X17D05YV8sHZPKUo465vFvLWL5YqXSB8FRa0a40MYrclwHER_dlxuR3Qh29vciql3C1Y2VqyenXWsKUxhdUr4-v_jj1jrTrkrLdIxlXpqPuYZR729tBa5lnguebZDV2jwX4-RRrw41U0w7ox2Srw4FtzoXy25JFrtZD4H_dWgX3e6q_23gFy3_p8gaOSh_U72Pv6eUrKMmMw1HsIwqgjUCsiYmFimooG8IzXuOR0ArEPEIPApN54xOyMe5-5X-YGObVVPDESO4UL59gbnpISC3eSCI3oHjAzywkbSkol89gq-JXlo6TOQIeNcACxdRNspMbca-U0iSLm6qg5hywJhIUpiPDuxn1iOMvmg87rWXWS5gW7-44NTSh8fuJ5C4Qg-qnP1x7fH3k8t4QFCSEHSbAnwi7Qb7mHtQGUprQxx3wPWKSQOprPil8C_b0tTyyTs_j7jSI70MpgksusDKCrGutXTBqehwtBFYPubVNFQYppJBQdyCPK5y_MKUw8DDGKsSUmi9JWSbfYIHM_otXtJSl72P_jcEeSKOQLTwoEFwKIrBW06MJ3PPv5LxoFX2a_UYzcSloW_Zwt4kkiDVlD766LX9Rozn15xuVUc4nE7fP3L04XppvgO1kJMBbkLM8HrkTIb9BEw03z180_JghjtxJ5z1teaqPoGi9VR_nAEXETHYDxjIdwbn9MpxuOLS1y6_1iQWfn_HF7WXX_vwmPzrPGtkk-dgcIb9SDRYpqGI8S_ib7nKyJiauhTJQe1KSTYJgXjO5yGqO8iPJ7bhLtApX0EyiHxBdfozapCqiO3YS4OvrSDSu6fS3qIdWJ1iyMpXnR49k6u8MrBN0v004YOypT4zRo1Vs4liqLFV-kZX6-UThBtDlsj__8w6lCqNvfzhLaFt9pY3jGT3ILxUX-jZGFI1c6OImzkfIWswDxW4Cc-Oc3-LMsH6OyXCqP_i8xd6arh2dCp5JWJluG_ws9RpNBt2Q3i49NvHXF2KecXqOsD2XTe9i5pD_8VLxtm36jZ01JFdGJBAGlJ3dVtw3at99u4k-jTBNClKjxPrIdC7p43ruDOkI=","internal_chat_message_metadata_passthrough":{"turn_id":"dee2191f-fdea-4bb2-86a3-42a7f28a3bb7"},"summary":[{"type":"summary_text","text":"**Flagging transactional identity**"}],"metadata":{"turn_id":"dee2191f-fdea-4bb2-86a3-42a7f28a3bb7"}},{"id":"fc_028129a7f1d66e9c016aa200d0e14887d2812831148e458dbc","type":"function_call","status":"completed","arguments":"{\"i\":\"Reporting preflight identity loss\",\"op\":\"send\",\"to\":\"Main\",\"message\":\"New preflight fixes orphan-on-retirement-error but deletes the only proof that an old PID is recycled BEFORE a fallible spawn. In your new missing-executable wake fixture, first resolve_wake errors after prepare removes the mismatching stamp; second resolve_wake now adopts the unrelated live PID (boot/owner-edge also skip it). Worse, kill_shell_at treats the remaining bare old PID as Held and can kill that unrelated process. Preserve old PID+birth evidence on failed spawn, and don't expose bare recycled old PID during preflight; coherent identity publication is preferable to deleting birth while retaining old PID. Source-only, not run.\"}","call_id":"call_jjvUlqpd3jIu1xU2ioFfr5Y9","internal_chat_message_metadata_passthrough":{"create_time":1789001928.035355,"turn_id":"dee2191f-fdea-4bb2-86a3-42a7f28a3bb7"},"name":"hub","metadata":{"turn_id":"dee2191f-fdea-4bb2-86a3-42a7f28a3bb7"}}]},"duration":12420.24209999945,"ttft":5335.258700000122,"completedAt":1789001939882,"contextSnapshot":{"promptTokens":124303,"nonMessageTokens":13209,"compactionEpoch":0}}}
{"type":"message","id":"6bd9920b","parentId":"f837fe0f","timestamp":"2026-09-10T00:58:59.885Z","message":{"role":"toolResult","toolCallId":"call_jjvUlqpd3jIu1xU2ioFfr5Y9|fc_028129a7f1d66e9c016aa200d0e14887d2812831148e458dbc","toolName":"hub","content":[{"type":"text","text":"Delivered to 1 peer(s):\n- Main: injected"}],"details":{"op":"send","from":"PersistentRestoreReview","to":"Main","receipts":[{"to":"Main","outcome":"injected"}]},"isError":false,"timestamp":1789001939885}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_mtf9m9RQswO98QzbQ5okqZZZ|fc_028129a7f1d66e9c016aa200ddda2887d28a4f3d2df843a4f4","toolName":"yield","startedAt":"2026-09-10T00:59:16.480Z"},"id":"8312a197","parentId":"6bd9920b","timestamp":"2026-09-10T00:59:16.480Z"}
{"type":"message","id":"818e1c82","parentId":"8312a197","timestamp":"2026-09-10T00:59:16.481Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":"**Assessing backward-compatible PID records**\n\n**Preparing review finding**"},{"type":"toolCall","id":"call_mtf9m9RQswO98QzbQ5okqZZZ|fc_028129a7f1d66e9c016aa200ddda2887d28a4f3d2df843a4f4","name":"yield","arguments":{"type":null,"data":{"overall_correctness":"incorrect","explanation":"Moving retirement before spawn fixes the newly reported orphan-on-retirement-error path, but preflight now destroys the previous process identity before knowing whether its replacement can launch. This can strand a recycled-PID corpse and cause an authenticated kill path to treat the unrelated process as the shell. Source inspection only; no runtime validation was run.","confidence":0.99,"findings":[{"title":"Preserve the old PID–birth identity until replacement succeeds","body":"`prepare_shell_launch` deletes the old birth stamp while leaving the old `shell.pid`, before the subsequent spawn can fail. For the recycled live PID in the new missing-executable wake test, the first wake now returns the expected spawn error but permanently removes the evidence that the PID belongs to another process; the next wake or restore adopts that unrelated PID and leaves the actual shell down. More seriously, `kill_shell_at` reads a missing birth as legacy bare-PID custody (`RelayLiveness::Held`), so a later close—or one racing the preflight window—can kill the unrelated process that the previous stamp would have protected. Keep the old identity coherent through preparation and failed launches rather than exposing the old PID without its birth.","priority":1,"confidence":0.99,"file_path":"C:/Users/decid/Documents/projects/spt-core/.worktrees/287-shellrestore/crates/spt-store/src/shellinfo.rs","line_start":177,"line_end":184}]},"error":null}}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":426,"output":486,"cacheRead":124160,"cacheWrite":0,"totalTokens":125072,"reasoningTokens":158,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1789001939915,"responseId":"resp_028129a7f1d66e9c016aa200d48ea887d2bc5532fbc159b66b","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_028129a7f1d66e9c016aa200d5d42087d281f0eaa1bdce4746","type":"reasoning","content":[],"encrypted_content":"gAAAAABqogDd8jVvR8oT5pr_YHU1IowN2hubynwWQxwfk414-Vh6hfoK3OHuW1rMINy1w_ETPFaM-bvFl3ZBJLw8FwYz9lhSD71CmO4-qWsUs1-XgA-VPy6k7InVxjuZOGoGj90uS0k0smGlE2fJju8RjmyBGk8zhYkGZE3c9jp9D4l6WNvSZm1732LZd7BkloK-CMrng34Y-CFu4ZuhLDIOi9LBlTnzeJyobgyqaRbfuOnYQrwUqIFioieicUPK23UkDv2w83k6YGBn5DJ1ItfpAl7wqc7j4o_fgzIwi5u9bnNOTbZ5TC2E8PkJIOPJ3blIXb7FXpZo66kALdLp8lvgSnjt5ndNLZCdeNCOWFm3xz7ROw8AXx6654TNCnF7mpHIZhhD_iTdobfGx_TROUgy79AIWkW4KHoLQQ_1BBXpf_J4T94W9Gao_fcx0rXyD3OkxFlGxt3HutCsRIIeTeLBlD3mKTcW-dHnVOeTXyk0cjqczOUud6_NSB-rCbfbWAAS8yMzlM6YYeDgKzt2mBbL-VWR1156dz1nYEmH2U6t3MxiXS4kISRtuhIE7eZzLzSN7SUi0wkNIS2lWUdrKE2H3CpE1vyT5-dZ-22V0_3d9H--z4oS9w0IPKJ9pHbkR1Do_B0iR_NXBSuNC5NlIMd_kw0nh-iaweV-ggeizVmKBd-Tj7laJWyc5eAtpUz3xWsduWwVBl86rdQLeH7Vfg-ZjmeUf3Tz8btwa2RZpmBNG0dEJ3Pgz2xVfc_PHJalLFfFnK5tS6WUrBEoy8yWkWGa4bMEjaDzHBWWDep_Hf9VDYevmm8Uu95rCNeHYi4oB00GUMfJgX-mxLH6ek0deDsCEOMQ5TDS9tLL771ecvd_qyytCKluabtHOdTRh0BSd5ICxTz2KsG-4n61FsSofhcKaYtlnroLWcy3L2awlObYLfdkbYornu028REFl0A_7NOAFDQXTfFeCa1k4cjZ2MHSAezU80B_v5KvLbJX0_QnvHL-iYp8KO7AbGjLdclu_fg7qWOn2FIHGX_HAAOn4QnCplCGSKyZBiA27YNAd4xIsCkNPVvMFjKsmw-LhytUvytc4lydnDln9JGwdrfKRgY_644QOm1ndxRcKuHEfP0DDkMZAl7cArDPNTrsUX4l-gqwXdXQ7TIYHpaW9gKJGk4-mvHtyjFw6Fon1-t-mAEsP-syaRyykevqTlXQRs78rNEu-LRFdYw9cbDd7OXTuzFbbDDeVmvF8r9_Fc06tKnFwQsI5KKPieqHkZDr6sRRd6UtQAUnNuoNRaj4gR4vJyBlZljbJxdYWPPts49M2P3LYsj8H7eLeCymRAD__giiOvIjai83Lo8PUveh4wlBh66T3wmvQsHWlUDCsjv7eCqINY86OihHlNZhErwQXzXv5Uutf52fWS4eo4WUVHN0yrmjFcRTODcOp3uoP9woCxaep0xduC8PBQ1YO6GcdE4Pr0Uw_KLrUJuMT0ePc_YARjHDkLGvNz_AYztrW6giciy7sfGMKpEEAYB4g-h3abrF0hccIYgtmLKuxticBDPfZXde3OySKm7hRvVJN9--azhawMe6F29oN5fHsI5EVqYXoZoyn4QLNp0IrjbcDsr4UAcXMbdfiO7zBDF8BqYHB4d17Mw6AcXK6KiQysHOqbFwse6D25tLcchvwKIuTgcmGv7nkB9PwHiB-q-E1URX5t9MsOY8wIu7MqyG8j02lyB4JH0Q7zlueDljyuhaby3triGzBMVgCBWacxjsAunRXXQcImcKS7xwfICqDUWc172Gy5aGfQ3inkRh2mL-7pXTaUP626sgBuRXI3g9oLlpwRvopABWHWGSFruEKBAY8TvHwHixDsskpKVtdkJsBNmHHW3IFu9HPSmPLu6aIupKpYhBojmqKcme9LEYmMprvTlCCL2GETuXLw5gq9KBuGapFTw7hJqknmpeXb9EaXH7sMBwosWrgW1NVKn8qCjQkbKUzqG2Jc6D6thHkGp2maJz1SkLDHABPLusIukFCrGOuCfAv8Gf-YTC9_XydlAmv9wkQl6_F0-PA3s5lRVr1XzQ-jP_iB8U6ArGyTsBgMkRRDDcs8DXAoxpbyFBjjX1cneLm06UmEjtAl8HqwRsmtll76AXr3m7cWlDoQ-7T4tQ1ePPVTBBA08UXjBZQ16gRlbxdCC-a4J_Upu35hwXxITSO0qvjHjeQd3Dz7LpgOMSx1pMA12ufpVOw9S1oEQgD7Y_p_0n-hqcRtddMo0b2I8bmmMMzAUYyhpO_1y_vKtqiROddvxnmpvhCWi1eCvd4VKZdmCmZNVHyyL3mrbEdPr-lxD6ceJLZ-gt8RAylgahEvAN5cCiblyeZRyeE9mJO_sDEArPfeCju5VNcxneyUwSaQKNXcpI5m-uYKBG_IKNRag_sM07L2TR0RKY0m7lDzXFe1DzxWJI4kbm","internal_chat_message_metadata_passthrough":{"turn_id":"dee2191f-fdea-4bb2-86a3-42a7f28a3bb7"},"summary":[{"type":"summary_text","text":"**Assessing backward-compatible PID records**"},{"type":"summary_text","text":"**Preparing review finding**"}],"metadata":{"turn_id":"dee2191f-fdea-4bb2-86a3-42a7f28a3bb7"}},{"id":"fc_028129a7f1d66e9c016aa200ddda2887d28a4f3d2df843a4f4","type":"function_call","status":"completed","arguments":"{\"type\":null,\"data\":{\"overall_correctness\":\"incorrect\",\"explanation\":\"Moving retirement before spawn fixes the newly reported orphan-on-retirement-error path, but preflight now destroys the previous process identity before knowing whether its replacement can launch. This can strand a recycled-PID corpse and cause an authenticated kill path to treat the unrelated process as the shell. Source inspection only; no runtime validation was run.\",\"confidence\":0.99,\"findings\":[{\"title\":\"Preserve the old PID–birth identity until replacement succeeds\",\"body\":\"`prepare_shell_launch` deletes the old birth stamp while leaving the old `shell.pid`, before the subsequent spawn can fail. For the recycled live PID in the new missing-executable wake test, the first wake now returns the expected spawn error but permanently removes the evidence that the PID belongs to another process; the next wake or restore adopts that unrelated PID and leaves the actual shell down. More seriously, `kill_shell_at` reads a missing birth as legacy bare-PID custody (`RelayLiveness::Held`), so a later close—or one racing the preflight window—can kill the unrelated process that the previous stamp would have protected. Keep the old identity coherent through preparation and failed launches rather than exposing the old PID without its birth.\",\"priority\":1,\"confidence\":0.99,\"file_path\":\"C:/Users/decid/Documents/projects/spt-core/.worktrees/287-shellrestore/crates/spt-store/src/shellinfo.rs\",\"line_start\":177,\"line_end\":184}]},\"error\":null}","call_id":"call_mtf9m9RQswO98QzbQ5okqZZZ","internal_chat_message_metadata_passthrough":{"create_time":1789001940.863545,"turn_id":"dee2191f-fdea-4bb2-86a3-42a7f28a3bb7"},"name":"yield","metadata":{"turn_id":"dee2191f-fdea-4bb2-86a3-42a7f28a3bb7"}}]},"duration":16561.996600000188,"ttft":1943.4879000000656,"completedAt":1789001956479,"contextSnapshot":{"promptTokens":124586,"nonMessageTokens":13209,"compactionEpoch":0}}}
{"type":"message","id":"a874f2f8","parentId":"818e1c82","timestamp":"2026-09-10T00:59:16.484Z","message":{"role":"toolResult","toolCallId":"call_mtf9m9RQswO98QzbQ5okqZZZ|fc_028129a7f1d66e9c016aa200ddda2887d28a4f3d2df843a4f4","toolName":"yield","content":[{"type":"text","text":"Result submitted."}],"details":{"data":{"overall_correctness":"incorrect","explanation":"Moving retirement before spawn fixes the newly reported orphan-on-retirement-error path, but preflight now destroys the previous process identity before knowing whether its replacement can launch. This can strand a recycled-PID corpse and cause an authenticated kill path to treat the unrelated process as the shell. Source inspection only; no runtime validation was run.","confidence":0.99,"findings":[{"title":"Preserve the old PID–birth identity until replacement succeeds","body":"`prepare_shell_launch` deletes the old birth stamp while leaving the old `shell.pid`, before the subsequent spawn can fail. For the recycled live PID in the new missing-executable wake test, the first wake now returns the expected spawn error but permanently removes the evidence that the PID belongs to another process; the next wake or restore adopts that unrelated PID and leaves the actual shell down. More seriously, `kill_shell_at` reads a missing birth as legacy bare-PID custody (`RelayLiveness::Held`), so a later close—or one racing the preflight window—can kill the unrelated process that the previous stamp would have protected. Keep the old identity coherent through preparation and failed launches rather than exposing the old PID without its birth.","priority":1,"confidence":0.99,"file_path":"C:/Users/decid/Documents/projects/spt-core/.worktrees/287-shellrestore/crates/spt-store/src/shellinfo.rs","line_start":177,"line_end":184}]},"status":"success"},"isError":false,"timestamp":1789001956482}}
{"type":"custom","customType":"session_exit","data":{"reason":"dispose","kind":"normal","recordedAt":"2026-09-10T01:06:16.492Z"},"id":"e39c644a","parentId":"a874f2f8","timestamp":"2026-09-10T01:06:16.492Z"}
