---
phase: 04-server-rebuild-mvp
plan: 09
subsystem: apps/server-shutdown-persistence
tags: [sigterm, graceful-shutdown, persistence, sqlite, wal, kill-9, deferred, phase-5-staging]
dependency-graph:
  requires:
    - 04-05 (RebnoRoom + onLeave timeout hook)
    - 04-07 (Better-Auth — account_id stable per session)
  provides:
    - apps/server/src/sigterm.ts — runGraceShutdown(deps, opts) + installSigtermHandler 6-step grace
    - apps/server/src/persistence.ts — flushAllCharactersToDb (1 Drizzle txn) + persistCharacter single-row
    - RebnoRoom.snapshotPlayers() — synchronous PlayerSnapshot[] view used by both onLeave-timeout and SIGTERM
    - apps/server/test/sigterm.integ.test.ts — 2 in-process integration tests (flush + ordering)
    - apps/server/test/persistence.test.ts — 5 unit tests
  affects:
    - REQ-SRV-08 SIGTERM half closed (impl + unit + int); kill -9 mid-tick half deferred to Phase 5 staging
    - Plan 04-12 RoomRegistry can call persistCharacter on room-transition events
tech-stack:
  added: []
  patterns:
    - "Drizzle better-sqlite3 transaction is SYNCHRONOUS — txn callback uses .all()/.run() chain methods, NOT awaited promises"
    - "Pre-collect PlayerSnapshot[] BEFORE colyseus.gracefullyShutdown — once gracefullyShutdown(false) returns, rooms are destroyed and matchMaker.getLocalRoomById is null"
    - "Idempotent SIGTERM/SIGINT handler — `draining` flag swallows repeat signals so dev Ctrl-C-spam doesn't double-flush"
    - "runGraceShutdown(deps, opts) extracted from installSigtermHandler so tests invoke the 6-step sequence in-process without process.exit tearing down vitest (skipExit + skipLitestreamGrace bypasses)"
    - "Broadcast SERVER_DRAINING BEFORE gracefullyShutdown(false), NOT after — D-16 step ordering re-derived (clients must receive reason string before WS close)"
key-files:
  created:
    - apps/server/src/sigterm.ts (197 lines)
    - apps/server/src/persistence.ts (126 lines)
    - apps/server/test/persistence.test.ts (148 lines)
  modified:
    - apps/server/src/RebnoRoom.ts (snapshotPlayers + db option threading + onLeave persistCharacter; +55 / -15)
    - apps/server/src/index.ts (installSigtermHandler call after listen + db on define options; +32)
    - apps/server/test/sigterm.integ.test.ts (replaced Wave-0 stub; 2 in-process tests; +258 / -X)
decisions:
  - "Task 3 (manual kill -9 mid-tick recoverability) DEFERRED to Phase 5 first Fly.io staging deploy. Rationale: user has no Linux dev environment; Windows cannot reliably send SIGKILL to a process child running a SQLite WAL write (Win32 TerminateProcess differs from POSIX SIGKILL semantics enough that a Windows-host run would not validate the Fly.io target case anyway). Phase 5 is the first time a real Linux host will run this server; the manual procedure is captured verbatim below in §Manual Verification (Phase 5 Debt) so it can run without re-research."
  - "SRV-08 closed in two halves: SIGTERM grace (this plan, automated integ test) ✅ + kill -9 recoverability (Phase 5 staging, manual procedure) ⏳. REQUIREMENTS.md row reflects partial complete with explicit Phase-5 verification debt note."
  - "In-process integ test instead of forked tsx subprocess (Rule 3 deviation per dcd2d0b commit body): tsx esbuild only honors the entry's tsconfig and emits modern TC39 decorators where @colyseus/schema needs the legacy emit. The OS-signal half (process.kill SIGTERM raising the handler) is verified by the runGraceShutdown wrapper being a thin pass-through — what the handler does on SIGTERM is exactly what runGraceShutdown does, just with skipExit=false."
  - "runGraceShutdown returns { snapshotsFlushed, roomsBroadcastTo } so tests can assert side-effect counts without inspecting log lines or DB rows alone."
  - "Drizzle better-sqlite3 transaction callback is synchronous (Rule 1 - Bug; recorded in a35a499 commit body). The plan body sketched `await db.transaction(async (tx) => ...)` with awaited inserts; better-sqlite3 transactions complete synchronously and the await is meaningless — switched to `db.transaction((tx) => { ... })` with `.run()` calls inside."
metrics:
  duration: ~30 min (Tasks 1+2 in single execution session; Task 3 deferred not counted)
  completed: 2026-05-07
  tasks-completed: 2
  tasks-deferred: 1
  files-created: 3
  files-modified: 3
  commits: 2 (a35a499 feat + dcd2d0b test); plus this metadata commit
---

# Phase 04 Plan 09: SIGTERM Grace + Crash Recoverability (REQ-SRV-08) Summary

Wire the SRV-08 acceptance: a 6-step SIGTERM grace handler that broadcasts
SERVER_DRAINING before tearing down rooms, flushes every in-memory
`PlayerSnapshot` to the `characters` table in a single Drizzle txn,
fsyncs the WAL, and exits cleanly. The complementary half — `kill -9`
mid-tick SQLite recoverability — is OS-signal-semantics-dependent and is
deferred to the Phase 5 first Fly.io staging deploy (no local Linux dev
environment available; Windows cannot reliably exercise SIGKILL against
an in-flight SQLite write).

## What Landed

### `apps/server/src/sigterm.ts` (197 lines)

`runGraceShutdown(deps, opts)` extracted from `installSigtermHandler` so
tests invoke the 6-step sequence in-process. Production handler is a
thin wrapper. Step letters from CONTEXT D-16, with the broadcast
re-ordered to (a*) ahead of (b):

| Step  | Action |
|-------|--------|
| (a*)  | Broadcast `s2c.error{code:'SERVER_DRAINING',reconnect_after_ms:30000}` to every active room — must land BEFORE rooms are torn down |
| (b)   | `colyseus.gracefullyShutdown(false)` — stop accepting new connections, close existing rooms |
| (c)   | `flushAllCharactersToDb(db, snapshots)` — single Drizzle txn |
| (d)   | `sqlite.close()` — synchronous WAL fsync |
| (e)   | 2 s grace for Litestream sidecar (Phase 5; this plan documents only) |
| (f)   | `process.exit(0)` |

Idempotent on repeat signals; SIGINT shares the path for dev Ctrl-C.
Default room enumeration walks `matchMaker.query()` →
`matchMaker.getLocalRoomById(roomId)` so only local rooms (the MVP
single-machine deploy is exactly that) are persisted by THIS process.

### `apps/server/src/persistence.ts` (126 lines)

- `flushAllCharactersToDb(db, players)` — bulk upsert all snapshots in
  one synchronous `db.transaction((tx) => { ... })` callback (the Drizzle
  better-sqlite3 driver does NOT honor an async txn callback; switching
  to sync was a Rule-1 bug fix from the plan body's sketch).
- `persistCharacter(db, p)` — single-row variant called by
  `RebnoRoom.onLeave` timeout path (CONTEXT D-14 graceful-disconnect
  trigger). Warn-and-continue on persistence error so a broken DB
  doesn't leak in-memory rows.

### `apps/server/src/RebnoRoom.ts`

- New public method: `snapshotPlayers(): PlayerSnapshot[]` — synchronous
  view of `state.players` used by BOTH `onLeave` timeout and the SIGTERM
  enumeration walk.
- `onCreate` accepts `{ auth, db }` from the matchMaker; `db` is stored
  on a private field for `onLeave` `persistCharacter` calls.
- `onLeave` timeout branch calls `persistCharacter(this.db, snap)`
  before evicting `state.players` + `authBySession`.

### `apps/server/src/index.ts`

- `colyseus.define('rebno', RebnoRoom, { auth, db })` — `db` now passes
  through to room construction.
- After `httpServer.listen`: `installSigtermHandler({ colyseus, sqlite, db })`.

### `apps/server/test/sigterm.integ.test.ts` (258 lines, 2 cases)

In-process via `spawnServer` helper (NOT a tsx-forked subprocess —
deviation Rule 3 documented in commit `dcd2d0b` body):

1. **End-to-end characters flush**: client joins, server receives an
   intent that sets `x=100,y=100`, runGraceShutdown invoked with
   `skipExit:true`; assertion re-opens the DB and confirms the
   `characters` row exists with the expected coordinates.
2. **Step ordering** (D-16): client subscribes to `s2c`, server
   triggers grace; assertion confirms a `SERVER_DRAINING` event was
   received BEFORE the WS close (`onLeave` fires AFTER the broadcast
   handler).

### `apps/server/test/persistence.test.ts` (148 lines, 5 unit tests)

Uses an in-memory SQLite via `better-sqlite3(':memory:')` + the same
Drizzle migrate runner as production. Tests:

1. `flushAllCharactersToDb` empty array — no-op, no error
2. `flushAllCharactersToDb` single snapshot — row inserted with
   correct `account_id, room_id, x, y, sprite_id, last_saved_at`
3. `flushAllCharactersToDb` multi-snapshot in one txn — N rows
   visible after one commit; partial-state failure leaves zero rows
4. `flushAllCharactersToDb` upsert — second flush of same `account_id`
   updates row in place (no duplicate)
5. `persistCharacter` single-row delegation — same row shape as
   `flushAllCharactersToDb([p])`

## Deviations from Plan

### Rule 1 — Bug: Drizzle better-sqlite3 txn callback is synchronous

The plan body sketched `await db.transaction(async (tx) => { await tx.insert(...) })`.
better-sqlite3's transaction model is synchronous; the await on the txn
callback is meaningless and the awaited inserts inside execute as
fire-and-forget. Switched to `db.transaction((tx) => { ... })` with
`.run()` chain methods inside. Rule 1 — Bug. Documented in `a35a499`
commit body.

### Rule 3 — Tooling friction: tsx subprocess forking can't transpile workspace TS with experimentalDecorators

The plan's Task 2 Step A spawned a tsx subprocess to run `apps/server/src/index.ts`
under SIGTERM. This failed because tsx's esbuild only honors the entry
file's tsconfig — workspace packages (`@rebno/protocol` schema files
which use `@type` decorators) compile through a tsconfig that emits
modern TC39 decorators, but `@colyseus/schema` requires the legacy emit.
Refactored sigterm.ts to expose `runGraceShutdown(deps, opts)` and
invoked it in-process from the test with `skipExit:true,
skipLitestreamGrace:true`. The OS-signal pathway (`process.on('SIGTERM',
handler)`) is a one-line wrapper around `runGraceShutdown(deps)` so
the test's coverage of runGraceShutdown is byte-equivalent to what a
real SIGTERM would do, minus the OS-level signal delivery itself.
Documented in `dcd2d0b` commit body.

### Task 3 (manual kill -9 verification) — DEFERRED to Phase 5 staging

[checkpoint:human-verify] task surfaced a checkpoint return. User
decision: defer to first Phase 5 Fly.io staging deploy because (a) user
has no Linux dev environment locally, (b) Windows SIGKILL semantics
differ from POSIX enough that a Windows run wouldn't validate the
Fly.io target case anyway, (c) Fly.io itself is the first real Linux
host the server will land on. The manual procedure below is verbatim
so Phase 5 can run it without re-research. Tracked as Phase-5
verification debt in `04-HUMAN-UAT.md` and on the SRV-08 row in
REQUIREMENTS.md.

## Manual Verification (Phase 5 Debt)

**Verbatim manual procedure for Phase 5 first staging deploy. Run on
the Fly.io machine OR a Linux host (WSL2 / dev VM acceptable as a
proxy if Fly is not yet provisioned).**

1. **Build + start the server:**
   ```bash
   pnpm --filter @rebno/server build
   DATABASE_URL=/tmp/rebno-killtest.db NODE_ENV=production \
     BETTER_AUTH_SECRET=$(openssl rand -hex 32) \
     pnpm --filter @rebno/server start &
   SERVER_PID=$!
   sleep 2
   ```

2. **Connect a client and join.** Use any colyseus.js script or
   `wscat -c ws://localhost:2567/?token=dev-bypass`; confirm join
   succeeds.

3. **While the server is mid-tick, SIGKILL it:**
   ```bash
   kill -9 $SERVER_PID
   ```

4. **Restart the server with the same DATABASE_URL:**
   ```bash
   DATABASE_URL=/tmp/rebno-killtest.db pnpm --filter @rebno/server start
   ```
   Expected: server boots cleanly; no SQLite corruption error in logs.

5. **Open the DB read-only and inspect:**
   ```bash
   sqlite3 /tmp/rebno-killtest.db "PRAGMA integrity_check; SELECT count(*) FROM characters;"
   ```
   Expected: `integrity_check` returns `ok`; the `characters` table
   is queryable (may be empty if the player hadn't been graceful-
   flushed yet — that's the 30 s checkpoint window per CONTEXT D-14,
   acceptable).

6. **Confirm WAL files are clean:**
   ```bash
   ls -la /tmp/rebno-killtest.db*
   ```
   Expected: `.db`, optional `.db-wal`, optional `.db-shm`. No
   corruption-marker files.

**Acceptance.** SQLite remains in a recoverable state after `kill -9`
mid-tick. The `journal_mode=WAL` + `synchronous=NORMAL` pragmas (CONTEXT
D-15) survive by design; this checkpoint verifies that empirically
on the actual Linux target host.

**Recording.** Append the `integrity_check` output and `count(*)`
result to `04-HUMAN-UAT.md` (Test 1 result). If corruption is
observed, escalate to Phase 5 — the WAL pragma config or the Litestream
restore runbook would need to be revisited earlier than planned (DEP-03
+ DEP-07 carry).

## Pre-Existing Carry-Overs (NOT introduced by this plan)

The following deferrals are pre-existing and continue across this
plan boundary unchanged:

1. **Phase-3 protocol-doc Windows CRLF drift** — `pnpm verify:phase-4`
   fails on Windows local at the `protocol-doc:verify` step due to
   one-sided LF normalization. Linux CI (ubuntu-latest) is the
   canonical "green" reference. Tracked in Plan 04-01 SUMMARY Deferred
   Issue #1 and Phase-3 03-HUMAN-UAT.md Test 1.
2. **`trace:check` apps/ scanner limitation** — the traceable-reqs
   scanner does not yet walk `apps/` workspace TS source for `[<impl>->]`
   tags; impl tags in `apps/server/src/*.ts` count as `doc` stage only.
   Tracked for Plan 04-13 final composite gate audit.
3. **Stale `.js` artifacts** in `apps/server/dist/` from prior
   `tsc` runs — gitignored but periodically re-emerge. No active
   impact.

## Self-Check: PASSED

- File `apps/server/src/sigterm.ts` exists at HEAD (committed in
  `a35a499`).
- File `apps/server/src/persistence.ts` exists at HEAD (committed in
  `a35a499`).
- File `apps/server/test/sigterm.integ.test.ts` exists at HEAD
  (committed in `dcd2d0b`).
- File `apps/server/test/persistence.test.ts` exists at HEAD
  (committed in `dcd2d0b`).
- Commit `a35a499` present in `git log` (verified).
- Commit `dcd2d0b` present in `git log` (verified).
- Task 3 deferral reflected in `04-HUMAN-UAT.md` (this commit) +
  SRV-08 row in REQUIREMENTS.md (this commit).
