# Provenance probe plan — #300, two windows

> JIT plan (AGENTS.md §Plans). **Nothing here is commissioned.** doyle holds approval; this document
> exists to be reviewed, not executed. No `REQ-*` is minted or activated by it, and no adapter
> behaviour changes on its account.
>
> Mirror-excluded by the root JIT-plan predicate (`:(glob)*-PLAN.md`) on its name alone.

## What this covers

Two of the three windows I offered. **External automation is DEFERRED and is NOT covered here** —
per doyle's ruling it is not equivalent to resume, and a resume answer must not be read as covering
it. It gets its own plan if it is ever wanted.

| Probe | Question |
|---|---|
| **A** | Does a submission landing in the **Stop→next-turn** boundary fire `UserPromptSubmit` at all? |
| **B** | Can a **resumed session** re-submit text, reaching the acceptance surface with no human send? |

Probe A tests a claim in my own source (`hook.rs` ~:3149) that such a submission is folded into the
closing turn with **no** `UserPromptSubmit`. My n=3 enqueue result landed mid-tool-call — a different
window — and does **not** refute it.

## Shared-host hazards that apply to BOTH probes

HFENDULEAM carries live agents and a CI runner. Three specific harms, with the mitigation each needs:

1. **`hook-trace.log` is one rolling node-wide file** (512KB, one previous generation, ~4.2h of
   retention at current rates). Any probe appends to it, and a high-N probe can **roll it and destroy
   another agent's in-flight trace evidence**. Mitigation: snapshot `hook-trace.log` and
   `hook-trace.log.1` to the scratchpad before starting, and cap total injected submissions per the
   budgets below. This is the single most likely way either probe hurts somebody else.
2. **Never kill by image name.** Probe B terminates a Claude Code process. It must target one
   captured PID. A cleanup that matched on image name once killed every `claude-spt.exe` on this
   shared runner (`[[ci-kill-scoping]]`); that must not recur.
3. **Never open-truncate `.claude.json`.** Probe B touches trust state for a spawned root. Writes go
   tmp + `os.replace`, never open-truncate (`[[claudejson-write-hazard]]`).

## Probe A — Stop→next-turn window

**Trigger (and this is the weak part, said out loud rather than discovered mid-probe).**
A watcher process tails `hook-trace.log`; the instant it sees this endpoint's turn-end line it issues
`api state idle` + a self-send, so core PTY-types a submission plus Enter into the input box. The
window is the few hundred ms between the turn closing and the next turn opening. My latency floor is
a file-poll interval plus an `spt` spawn — the same spawn class that measures 28–45ms per stage in
the live trace, so **realistically 100–300ms of trigger latency against a window that may be
shorter**. I cannot state a reliable hit rate in advance. **The plan is therefore many cheap attempts
plus a hit-detector, not one careful shot** — and if the hit rate is zero the honest output is "this
window is not reachable with the trigger available to me", which is a result, not a failure.

**A0 is already answered, from the existing log, read-only — and it kills the trigger as first drafted.**
Across both trace generations the only `Stop` lines are the SUPPRESSED branch ("Stop idle mark
SUPPRESSED for <id> — across-clear quiet window", 19 of them). That is an exceptional path, not a
per-turn-end record: **the turn-end hook runs every turn but writes no trace line in the normal
case.** A watcher tailing `hook-trace.log` therefore cannot see a turn end, and the trigger as I first
wrote it does not exist.

What survives: the turn-end hook's normal path **marks the endpoint IDLE**, and endpoint state is
externally readable. So the trigger becomes **poll endpoint state for the BUSY→IDLE transition and
fire on the edge**, not tail the trace. Same latency class — a spawn per poll — so the 100–300ms
floor and the unknown hit rate stand unchanged. This also means the probe's own forced `state idle`
must be distinguishable from the turn-end idle mark, or the rig will trigger on itself; that is a
design detail to settle before any run, not during one.

**Observable.** Two independent reads, per trial, keyed by a **unique nonce in each payload** so
trials can never be confused with one another:
- `hook-trace.log` — is there a `BEGIN UserPromptSubmit id=perri` between this trial's send and the
  next turn's start?
- the session transcript JSONL — does the nonce appear, and is it positioned **inside** the closing
  turn (folded) or at the head of a **new** turn (missed window)?

| hook fired | transcript position | reading |
|---|---|---|
| yes | folded | hook fires even in this window — **my source comment is wrong** |
| no | folded | **comment confirmed** — silent fold, no acceptance surface |
| yes | new turn | **missed the window** — discard, not a trial |
| no | nonce absent | **send never landed** — instrument failure, discard |

**Controls — what makes a null a measurement.**
- **A trial counts only if its nonce appears in the transcript.** That is what separates "the hook was
  silent" from "I failed to observe a hook", and it is the whole reason a no-fire row can be believed.
- **Denominator guard:** the harvest refuses to print a fold-rate when the counted-trial denominator
  is zero, rather than printing a reassuring `0/0 = 0.00%`. Paid for twice already
  (`[[v0410-hook-deadline-visible]]`).
- **Positive control that must count:** the same rig, fired mid-tool-call, must reproduce the n=3
  enqueue fire. A rig that cannot see the fire it already measured cannot be trusted to report its
  absence.

**Isolation — and this is the reason for the recommendation below.**
Run against MY live session, the probe **necessarily perturbs a live session**: it injects real
submissions into my own TUI, forces my endpoint state, burns a conversation turn per trial and bloats
my context. It does not touch other endpoints (`state`/`send` are scoped to `perri`) and reads
`hook-trace.log` read-only.
**Recommended instead: a disposable endpoint in a scratch project dir**, headless, with its own
config root — no live session perturbed, no context burn, and its trace lines carry a different
endpoint id so they are trivially separable from live traffic. Cost: spawn setup, and the
headless-spawn trust/external-import gates must be pre-seeded on the spawn's **resolved** root
(`[[headless-spawn-external-imports-gate]]`, `[[f027-trust-seam-groundtruth]]`). Spawning CC from a
live agent's shell is a standing trap (`[[v0200-stub-delivery-shipped]]`) — this is the care that
buys it off, and if that care is not wanted, Probe A should not run at all.

**Budget / duration.** A0 is done (above, read-only, no cost). Remaining: trigger rebuild on the
state-edge design ≈ 45 min, then **≤ 40 injected submissions**, ≈ 45–60 min. Total ≈ 2h. Hit rate
unknown and possibly zero.

## Probe B — resumed-session re-submit

**Trigger.** Far easier and far safer than A, and it needs no race. In a disposable session with
hooks active: submit one nonce-bearing prompt, terminate the process **by captured PID** while the
turn is still in flight, then resume the session and observe whether that same text reaches the
acceptance surface again with nobody typing it.

**Observable.** Count `BEGIN UserPromptSubmit` lines carrying that nonce's endpoint id, across the
original session and the resumed one, against **one** human submission. Read from `hook-trace.log`
plus both transcripts. Two fires for one typed submission ⇒ a machine re-submit reaches the surface,
and a token captured at the second fire would attribute a replay to whoever holds the seat then —
which is the thing #300 must not do.

**Controls.**
- Nonce per trial; a trial counts only if the nonce appears in the resumed session's transcript.
- Denominator guard as above.
- Positive control: the FIRST (human) submission must be seen firing, in the same run, before any
  claim is made about the second.
- Run the same sequence **without** a resume as a negative control, to show the second fire is the
  resume's doing and not an artifact of the kill.

**Isolation.** Fully isolated: disposable endpoint, scratch project dir, own config root, nothing
touching live sessions or the CI runner. The two hazards that bite here are the PID-scoped kill and
the trust-store write, both named above.

**Duration.** ≈ 30 min including setup. N is small — 3–5 trials plus the negative control.

## Recommendation

**Probe B is worth running and cheap; Probe A is expensive, may be untriggerable, and A0 decides it.**
If only one is commissioned, take B: it answers a live question in §2b (machine injection reaching
the surface) with a bounded, fully isolated rig, whereas A may spend two hours to conclude the window
is not reachable with the trigger I have.

## Gate

Unchanged and non-negotiable, for any code this ever produces: `sh ci/run-gates.sh` PASS **and**
`traceable-reqs check` exit 0 before anything lands. Probe rigs live under `ci/measure/` beside
`trace-harvest.py` if they are kept at all; a throwaway rig stays in the scratchpad and is not
committed.
