This repository has been archived on 2026-09-12. You can view files and clone it. You cannot open issues or pull requests or push a commit.
Files
praxis/.ciagent/GRILL-v0.4.md
T
Praxis CI 6ab40c6f25 docs(milestone): merge phase/00 pre-execution → milestone/v0.4-operator-tier
Phase 0 complete — v0.4 operator tier pre-execution artifacts:
- PROJECT.md (v0.4 scope validated, D-050..D-057)
- REQUIREMENTS.md (8 active REQs: REQ-MT-01/02, REQ-AUTH-01, REQ-DASH-01 + 4 NFRs)
- ARCHITECTURE.md (operator Postgres + auth + dashboard + aggregation + VC migration)
- PERSONAS.md (6 active personas — frontend + devops reactivated)
- RESEARCH-v0.4-operator-tier.md (7 domains, 20 risks, confidence 0.70-0.95)
- PLAN-v0.4-operator-tier.md (2 execution phases, 10 slices, 52 tasks, 8/8 REQ)
- GRILL-v0.4.md (proceed-with-conditions, 6 MUST binding decisions)

---ci---
project: praxis
phase: 0
milestone: v0.4
status: complete
requirements:
  covered: []
  partial: []
---/ci---
2026-08-04 00:39:53 +00:00

73 KiB

CIAgent Grill Report — v0.4 Operator Tier

Run: 2026-08-04 (mode: mechanical, focus: all axes + 6 v0.4-specific probes)

Reviewer: adversarial technology executive (red-team) Subject: v0.4 execution plan (Operator Tier — Cohort Dashboard + Auth + Postgres) — 2 execution phases, 10 slices, 52 tasks Stance: plan is unfeasible, over-scoped, and too costly until evidence forces otherwise Artifacts reviewed: PROJECT.md, REQUIREMENTS.md, ROADMAP.md, ARCHITECTURE.md, RESEARCH-v0.4-operator-tier.md, PERSONAS.md, PLAN-v0.4-operator-tier.md, GRILL-v0.3.md, config.json, docker-compose.yml, server/session_recorder.py, server/vc/issuer_keys.py, server/__main__.py, db/store.py Binding status: This grill verdict must be cleared (MUSTs resolved, FIXs tracked) before EXECUTE is authorized.


Verdict: Proceed-with-conditions (confidence: 0.72)

The v0.4 plan is well-researched, cleanly phased, and honors the v0.3 grill's binding verdict (operator tier deferred, formative label applied, scoring_inconclusive fallback implemented, VC interop + key-rotation drills shipped in v0.3 codebase — all verified). The architecture is sound and the risk register is the most honest in the project's history (20 risks, 1 high, 9 medium, 11 low — all addressed). However, three material issues must be resolved before EXECUTE: (1) R-AUTH-01 is a partial resolution that re-litigates a v0.3 grill MUST — the config-driven flag is a punt, not a fix, and the cohort-dashboard-reads-only-aggregates defense-in-depth is the real mitigation, which should be elevated; (2) the k-anonymity-at-pilot-scale problem means v0.4 ships a dashboard that cannot display any data at production pilot scale (1 learner) — this is a real deliverable only if test-seeded data is treated as the validation path, which the plan does but does not emphasize; (3) the VC key migration verification endpoint now queries two stores (Postgres for keys, SQLite-fallback for v0.3 credentials) — a complexity the plan defers to "open question #1" but which is on the critical path of R-VC-MIG-01.

The plan is not over-scoped (8 REQs, cleanly split P1 infra / P2 feature). It is not unfeasible (52 tasks vs v0.3's 40, analogous). It is not a zombie (the operator tier was the explicitly-deferred v0.3 scope, now delivered). The conditions are binding but surgical.


Axis 1 — Business Case

  • Q1: What problem does v0.4 solve, and is it the top priority?

    • Evidence: GRILL-v0.3.md Axis 2 MUST #1 — "defer REQ-DASH-01 + operator tier to v0.4"; ROADMAP.md:9 — "v0.4 activates the operator tier deferred from v0.3 per the grill's binding verdict"; PROJECT.md:47 — "v0.4 layers the operator surface on top of it."
    • Answer: v0.4 delivers the operator tier that the v0.3 grill explicitly split out. The operator tier (cohort dashboard + auth + Postgres) was originally v0.8 on the ROADMAP (GRILL-v0.3.md:44), pulled to v0.3, then split to v0.4 by the grill. This is the deferred obligation, not new scope. The priority is correct: v0.3 shipped the learner-facing mastery layer; v0.4 ships the operator-facing visibility layer. The alternative (multi-path / Live Assist / low-bandwidth) would expand the learner surface before the operator surface exists to observe it.
    • Confidence: 0.85
    • Decision: G-001 — v0.4 operator tier is the correct next priority (delivers the v0.3 grill's deferred obligation). (0.85)
  • Q2: Who is the named executive sponsor for the operator tier?

    • Evidence: config.json:13 — "level": "full"; config.json:16 — "decision_confidence_threshold": 0.6; PROJECT.md:5 — "Autonomy: full."
    • Answer: No human sponsor. The CI agent is the executive sponsor under full autonomy. This is the project's established governance model since v0.1. The v0.3 grill accepted this (no escalation on governance). The "sponsor makes a decision under pressure" test is met by the grill itself — this document is the pressure decision.
    • Confidence: 0.80
    • Decision: G-002 — CI is the named sponsor under full autonomy (no change from v0.1-v0.3 governance). (0.80)
  • Q3: What happens to the business if v0.4 is cancelled?

    • Evidence: ROADMAP.md:131-139 — future milestones (v0.5 Live Assist, v0.6 low-bandwidth) do not depend on the operator tier; v0.9 credentialing depends on VC issuer (v0.3, already shipped). The learner-facing product (v0.1-v0.3) works without the operator tier.
    • Answer: If v0.4 is cancelled, the learner product continues to function. The operator tier is a visibility feature, not a learner-path feature. However, cancelling v0.4 means the v0.3 grill's binding verdict (defer to v0.4) becomes a permanent deferral — the operator tier was promised and not delivered. This would be the first broken grill commitment. The project is not a zombie (cancelling has a cost: the grill's credibility), but the operator tier is a nice-to-have for the pilot, not a blocker for a pilot deployment. A pilot can run with a single learner and no dashboard.
    • Confidence: 0.75
    • Challenge: The operator tier's business value at pilot scale (1 learner, k-anon suppresses everything) is low. The dashboard will show "— (<10 learners)" for every cell. This is a placeholder deliverable unless multi-learner data is seeded. The plan acknowledges this (Open Question #3) but does not treat it as a material risk to the business case.
    • Decision: G-003 — v0.4 is not a zombie (delivers a grill obligation) but its pilot-scale business value is low (k-anon suppresses all cells with 1 learner). The dashboard's validation path is test-seeded data (≥10 mock learners), not pilot traffic. This must be documented in the ship notes. (0.75)
  • Q4: Is the ROI calculated against a counterfactual?

    • Evidence: MISSING — no ROI calculation in any .ciagent/ file. The project is a pre-revenue pilot (D-012 — no enforced cost ceiling for pilot).
    • Answer: No ROI calculation exists. The counterfactual is "ship v0.4 vs skip to v0.5 (Live Assist)." Shipping v0.4 costs ~52 tasks of tokens + a Postgres service + 3 new pip deps + 1 new npm dep. Skipping to v0.5 would leave the operator tier permanently deferred (broken grill commitment) and Live Assist would build on a learner surface with no operator visibility. The ROI is governance credibility + operator visibility foundation for v0.5+, not a financial return.
    • Confidence: 0.65
    • Decision: G-004 — no financial ROI; the ROI is governance credibility (delivering the grill's deferred obligation) + architectural foundation (Postgres + auth for v0.5+). Accept the non-financial ROI under full autonomy. (0.65)

Axis 2 — Scope and Requirements

  • Q1: Is v0.4 scope stable? (8 REQs from v0.3 grill deferral — clean handoff, or new scope creep?)

    • Evidence: GRILL-v0.3.md Axis 2 MUST #1 — "defer REQ-DASH-01 + REQ-AUTH-01 + REQ-MT-01/02 + 4 NFRs to v0.4"; REQUIREMENTS.md:8-36 — v0.4 activates exactly those 8 REQs; PROJECT.md:49-54 — v0.4 in-scope matches the deferred set.
    • Answer: Clean handoff. The 8 REQs activated in v0.4 are exactly the 8 REQs the v0.3 grill deferred. No new REQs were added. No scope creep. The scope is contracting relative to the v0.3 plan (which originally included these + the mastery layer).
    • Confidence: 0.90
    • Decision: G-005 — v0.4 scope is a clean handoff from the v0.3 grill deferral. No scope creep. (0.90)
  • Q2: Who owns the requirements, and are they frozen?

    • Evidence: config.json:13 — full autonomy; PROJECT.md:5 — "Autonomy: full"; REQUIREMENTS.md:8-36 — 8 active REQs with Phase + Status columns.
    • Answer: CI owns the requirements under full autonomy. They are frozen at the SPECIFY stage (commit 1b5173e — "validate specification"). The CLARIFY stage (commit 4f565d6) added D-050..D-057 but did not add/remove REQs. Frozen.
    • Confidence: 0.85
    • Decision: G-006 — requirements are frozen (8 REQs, CI-owned under full autonomy). (0.85)
  • Q3: What is explicitly out of scope?

    • Evidence: PROJECT.md:56-65 — explicit out-of-scope list; REQUIREMENTS.md:38-48 — out-of-scope list.
    • Answer: Explicitly out of scope: multi-path launch, full operator-suite dashboard (REQ-DASH-02), Live Assist, low-bandwidth, multi-language, persona switching, learner auth, RBAC (single operator role), third-party credential issuers, differential privacy. The out-of-scope list is the most explicit in the project's history. Single operator role (no RBAC) is the key constraint — v0.4 ships one role.
    • Confidence: 0.88
    • Decision: G-007 — out-of-scope is explicit and comprehensive (RBAC, learner auth, DP, multi-path all deferred). (0.88)
  • Q4: Hidden requirements? (TLS for secure cookies? Postgres backup verification? Operator account lifecycle — deactivation, password reset?)

    • Evidence: RESEARCH-v0.4 §2.4 — R-AUTH-01 acknowledges the Secure-cookie+no-TLS tension; D-055 — backup strategy defined (pg_dump, 7-day retention); D-052 — operator bootstrap CLI; PROJECT.md:62 — "RBAC deferred (one role)."
    • Answer:
      • TLS for secure cookies: NOT a hidden requirement — it is the explicit R-AUTH-01 tension, resolved (partially) by config-driven PRAXIS_COOKIE_SECURE. See Axis 3 + signature probe.
      • Postgres backup verification: The plan defines a backup strategy (TASK-02-03 — backup cron script) but does NOT define a backup verification / restore drill. The script comments mention pg_restore --clean --if-exists but there is no task that executes a restore and verifies data integrity. A backup that is never restored is an unverified backup. This is a hidden requirement.
      • Operator account lifecycle (deactivation, password reset): D-052 defines bootstrap (creation) + a --update flag (password rehash). The operators table has is_active (TASK-03-04 handles inactive → 401). But there is no operator deactivation task — no CLI to set is_active=false, no UI for it. Password reset = create-operator.py --update (documented). Deactivation is a gap, but minor (single operator, can be done via SQL if needed). Not a blocker.
    • Confidence: 0.70
    • Challenge: Backup verification is a hidden requirement. A nightly pg_dump that is never restored is theater, not a backup.
    • Decision: G-008 (MUST) — Add a backup-restore drill task to P1 (either in SLICE-02 or SLICE-06): execute pg_restore --clean --if-exists against a test Postgres instance, verify the 5 tables + row counts match. This is a one-task addition. The restore drill must run at least once in CI/staging to prove the backup is valid. (0.70)

Axis 3 — Architecture and Technical Feasibility

  • Q1: Has the Postgres-in-LXC + asyncpg + auth + dashboard architecture been validated by operators, or only by the plan?

    • Evidence: RESEARCH-v0.4 §1.1-1.7 — Postgres 16-slim resource footprint analysis (0.88 confidence); §2.1-2.6 — argon2id + SessionMiddleware (0.88); §4.1-4.5 — React Router + SPA fallback (0.85). No external operator validation (full autonomy — CI is the operator).
    • Answer: The architecture is validated by research (vendor docs, OWASP, ecosystem knowledge) and codebase inspection (existing session_recorder.py:143 asyncio.create_task pattern, existing issuer_keys.py lifecycle). It is NOT validated by an external operator (none exists). The asyncpg pool pattern (lifespan context manager) is standard FastAPI. The Starlette SessionMiddleware is the documented FastAPI session pattern. The SPA fallback (catch-all before StaticFiles) is the standard React-in-FastAPI pattern. The architecture is conventional — no novel combinations.
    • Confidence: 0.80
    • Decision: G-009 — architecture is conventional (standard FastAPI + Postgres + React patterns), research-validated. No external operator exists (full autonomy). Accept. (0.80)
  • Q2: Integration surface — Postgres 16, asyncpg, Starlette SessionMiddleware, slowapi, argon2-cffi, react-router-dom. Risk of quiet cost doubling?

    • Evidence: PLAN-v0.4:770-771 — 3 new pip deps (asyncpg, argon2-cffi, slowapi) + 1 new npm dep (react-router-dom). RESEARCH-v0.4 §new-deps.
    • Answer: 4 new dependencies. Each is a CVE vector + version-pin burden. asyncpg is the most consequential (new DB driver — connection pool lifecycle, statement cache, type coercion). slowapi is the youngest (maintenance risk — RESEARCH-v0.4 §2.5 notes "young lib, but works" at 0.70 confidence). argon2-cffi is mature (reference impl wrapper). react-router-dom@^7 is the standard React router (mature, but v7 is a major version — the <BrowserRouter> API is stable). The cost-doubling risk is low — these are all single-purpose, well-scoped deps. The real cost is the Postgres service (memory, disk, backup, migration runner) — but that is budgeted (6GB CT, pgdata/pgbackups volumes).
    • Confidence: 0.78
    • Decision: G-010 — 4 new deps, all single-purpose and well-scoped. Cost-doubling risk is low. slowapi is the youngest dep — the plan documents a hand-rolled counter fallback (RESEARCH-v0.4 §2.5). Accept with the fallback documented. (0.78)
  • Q3: Is there an existing system being replaced? (VC issuer key store SQLite→Postgres — migration path for existing issued VCs?)

    • Evidence: server/vc/issuer_keys.py (128 lines) — current SQLite-backed key store; D-051 — migration strategy; PLAN-v0.4 SLICE-04 — VC key migration slice; TASK-06-05 — R-VC-MIG-01 e2e test.
    • Answer: The VC issuer key store is being migrated (SQLite→Postgres). The v0.3 issued_credentials table remains in SQLite (no data migration — D-051 "no re-issuance"). The verification endpoint (TASK-04-04) will try Postgres for keys, fall back to SQLite for v0.3 credentials. This is a two-store verification path — a complexity that is on the critical path of R-VC-MIG-01.
    • Confidence: 0.75
    • Challenge: The two-store verification path (Postgres for keys, SQLite-fallback for v0.3 credentials) is a hidden complexity. Open Question #1 (PLAN-v0.4:742) defers this to EXECUTE: "the executor should choose the simpler approach." But this is not an implementation detail — it is an architectural decision that affects the verification endpoint's failure modes. If Postgres is down, can v0.3 credentials still verify? The plan says TASK-04-04 "try Postgres first, fall back to SQLite" but TASK-06-03 says "if pg_store is None, fall back to PraxisStore path (v0.3 compat)." These two fallback semantics are consistent but the plan does not make the consistency explicit.
    • Decision: G-011 (MUST) — The verification endpoint's two-store fallback semantics must be explicit in the plan, not deferred to EXECUTE. Rule: (a) if Postgres is available, use it for key lookup (both active + superseded keys); (b) if Postgres is available but the credential is not found in Postgres issued_credentials, fall back to SQLite issued_credentials (v0.3 credentials); (c) if Postgres is NOT available (no DSN), use the existing v0.3 SQLite path for both keys + credentials. This must be documented in TASK-04-04 and TASK-06-03 as a binding contract, not an open question. (0.75)
  • Q4: Technical debt inherited — v0.3's SQLite VC issuer keys, single hardcoded learner profile, no TLS in the LXC pilot.

    • Evidence: db/store.py:29 — HARDCODED_LEARNER_ID = "learner-1"; server/__main__.py:46 — HOST = _env("PRAXIS_HOST", "0.0.0.0") (binds to all interfaces, not loopback); D-030 — no Traefik/TLS for pilot.
    • Answer: Three inherited debts:
      1. SQLite VC issuer keys — being migrated (D-051). This is v0.4's job, not inherited debt.
      2. Single hardcoded learner profileHARDCODED_LEARNER_ID = "learner-1". This is the root cause of the k-anon-at-pilot-scale problem (see signature probe #3). Not addressed in v0.4 (multi-learner-per-device is deferred). The aggregation pipeline groups by learner_ref but there is only one learner_ref. The dashboard will suppress everything.
      3. No TLS in the LXC pilot — D-030. This is the root cause of R-AUTH-01 (see signature probe #1). Not addressed in v0.4 (TLS deferred to a later milestone).
    • Confidence: 0.72
    • Decision: G-012 — three inherited debts acknowledged: (1) SQLite VC keys → being migrated (v0.4's job); (2) single hardcoded learner → not addressed (k-anon suppresses all pilot data); (3) no TLS → not addressed (R-AUTH-01 config-driven punt). Debts #2 and #3 are accepted as pilot-scale constraints with documented mitigations. (0.72)

Axis 4 — People, Skills, and Organization

  • Q1: Key-person dependency — which 2-3 personas, if absent, would v0.4 fail?

    • Evidence: PERSONAS.md v0.4 roster — 6 active personas; PLAN-v0.4:79-86 + :419-426 — persona load distribution.
    • Answer: The 3 critical personas:
      1. security-engineer — owns VC key migration (R-VC-MIG-01, high severity) + auth stack (argon2id, cookies, rate limit). If absent, the highest-severity risk is unowned. 8 tasks in P1.
      2. data-engineer — owns Postgres schema + migration runner + PgStore + IssuerKeyStore protocol. If absent, the foundation (SLICE-01) is unowned. 8 tasks in P1 + 3 in P2.
      3. backend-engineer — owns asyncpg pool wiring + operator API (8 endpoints) + aggregation pipeline + SPA fallback + session_recorder extension. The largest task surface (16 tasks across P1+P2). If absent, the integration slices (SLICE-06, SLICE-10) have no owner. The lead-developer is coordination (not key-person — can be covered by backend-engineer). The frontend-engineer is P2-only (dashboard UI). The devops-engineer is P1-only (compose + backup + bootstrap). The key-person risk is concentrated in security + data + backend.
    • Confidence: 0.82
    • Decision: G-013 — key-person dependency: security-engineer, data-engineer, backend-engineer. All 3 are critical-path. Under full autonomy with parallelization (max 5 concurrent), this is manageable. Accept. (0.82)
  • Q2: Are the 6 personas actually available?

    • Evidence: config.json:22-27 — parallelization enabled, max 5 concurrent; PERSONAS.md — 6 active personas (lead, backend, frontend, data, security, devops). security-engineer + devops-engineer are NOT in config.json personas array (emergent — defined in PERSONAS.md, per PERSONAS.md:542).
    • Answer: All 6 are "available" in the sense that the CI agent spawns them on demand. The config.json personas array has only 4 (lead, backend, frontend, data); security + devops are emergent (PERSONAS.md). Territory enforcement is warn (config.json:51) — so emergent personas are not blocked. The max-concurrent-agents is 5, but 6 personas are active — one will be idle at peak. The P1 wave-2 has 3 parallel slices (SLICE-03, 04, 05) — 3 personas active (security, security, devops). The P2 wave-1 has 3 parallel slices (SLICE-07, 08, 09) — 3 personas (backend, backend, frontend). The 5-agent limit is not a binding constraint.
    • Confidence: 0.80
    • Decision: G-014 — 6 personas available (4 in config + 2 emergent), max 5 concurrent. The 6>5 mismatch is not binding (peak parallelism is 3 slices). Accept. (0.80)
  • Q3: Product owner with authority?

    • Evidence: config.json:13 — full autonomy; PROJECT.md:5.
    • Answer: CI is the product owner under full autonomy. This is the established model since v0.1. No committee. The grill is the pressure-test.
    • Confidence: 0.85
    • Decision: G-015 — CI is the product owner with full authority (no change). (0.85)
  • Q4: Is the team building capability they don't have? (Postgres admin, k-anonymity, argon2id — all new to the project)

    • Evidence: RESEARCH-v0.4 §1-7 — all 7 domains are new to the project (Postgres 16, asyncpg, argon2id, Starlette SessionMiddleware, slowapi, k-anonymity, React Router); PERSONAS.md v0.4 — data-engineer expands to Postgres, security-engineer expands to argon2id + slowapi.
    • Answer: Yes — the team is building capability it doesn't have. Postgres admin (migrations, pool, backup), k-anonymity (write-time suppression SQL), argon2id (OWASP params), signed cookies (Starlette SessionMiddleware), React Router (SPA fallback). All new. However: (a) this is a pilot, not a production system — learning-as-you-go is acceptable for prototypes per the grill's stance; (b) the research is thorough (OWASP fetched 2026-08-04, Postgres 16 docs verified, asyncpg pattern validated); (c) the highest-risk new capability (custom VC crypto) was already shipped in v0.3 with interop + rotation tests (verified in codebase: test_vc_interop.py, test_vc_key_rotation_drill.py). The v0.4 new capabilities are conventional (standard FastAPI + Postgres + React patterns), not novel.
    • Confidence: 0.75
    • Decision: G-016 — team is building new capability (Postgres, auth, k-anon, React Router) but all are conventional patterns with thorough research. Accept for pilot. (0.75)

Axis 5 — Timeline and Estimates

  • Q1: Was the 2-execution-phase structure set before or after the scope was understood?

    • Evidence: ROADMAP.md:31-53 — P1/P2/P3 structure defined in ROADMAP (pre-PLAN); PLAN-v0.4:14-22 — phase split rationale refines the ROADMAP structure.
    • Answer: The ROADMAP defined P1 (operator foundation) + P2 (cohort dashboard) + P3 (review) before the PLAN. The PLAN refined the split (6 slices in P1, 4 in P2). The scope was understood at ROADMAP time (8 REQs from v0.3 grill deferral). The deadline (per-phase ship tags v0.1.7, v0.1.8, v0.1.9) was set in ROADMAP. This is not a reverse-engineered deadline — the phases are defined by scope (P1 = infra/auth, P2 = dashboard), not by a target date.
    • Confidence: 0.85
    • Decision: G-017 — phase structure set after scope was understood (ROADMAP post-grill). Not reverse-engineered. (0.85)
  • Q2: Critical path — what single thing would push v0.4 by a phase?

    • Evidence: PLAN-v0.4 wave dependency graphs (P1:60-75, P2:404-415); RESEARCH-v0.4 risks R-VC-MIG-01 (high), R-MT-01 (medium), R-DASH-03 (medium).
    • Answer: The critical path is P1 Wave 1 → Wave 2 → Wave 3 → P2 Wave 1 → Wave 2. The single thing that would push v0.4 by a phase:
      • Most likely: SPA fallback breaking the voice UI (R-DASH-03/05). The catch-all route (@app.get("/{path:path}")) before StaticFiles is a change to server/__main__.py — the same file that serves the voice loop. If the catch-all shadows StaticFiles asset serving (JS/CSS), the voice UI breaks. TASK-10-04 tests this (8 assertions), but if the test fails, the fix is non-trivial (route ordering in FastAPI is subtle). This would push P2 by a wave.
      • Less likely: VC key migration (R-VC-MIG-01). The e2e test (TASK-06-05) is thorough, but if the v0.3 public key fails to verify against the Postgres store (e.g., key_id mismatch, encoding issue), the migration is blocked. The mitigation (archive before activate) is correct, but the test is the proof.
      • Least likely: Postgres resource contention (R-MT-01). 6GB CT has ~50% margin. The nightly jobs are at 03:00 CT. This is a measurement issue, not a design issue.
    • Confidence: 0.75
    • Decision: G-018 — critical-path risk: SPA fallback breaking voice UI (R-DASH-03). Mitigation: TASK-10-04 (8 assertions). If it fails, the fix is route ordering. Accept with the test as the gate. (0.75)
  • Q3: Are the 52 tasks evidence-based or pulled from a target?

    • Evidence: PLAN-v0.4:764 — 52 tasks (29 P1 + 23 P2); GRILL-v0.3.md:29 — v0.3 had 70 tasks (originally) → shipped as ~40 after the grill split; ROADMAP.md:81 — v0.3 P1 shipped as v0.1.4.
    • Answer: v0.3 shipped ~40 tasks (post-grill split) successfully. v0.4 has 52 tasks across 2 phases (29 + 23). The task count is analogous to v0.3 (40 tasks → 52 tasks, +30%). The scope is comparable (v0.3 mastery+VC vs v0.4 operator tier). The tasks are bottom-up sized (each slice has 3-7 tasks with acceptance criteria). Not pulled from a target.
    • Confidence: 0.80
    • Decision: G-019 — 52 tasks is evidence-based (analogous to v0.3's 40, bottom-up sized). Accept. (0.80)
  • Q4: Definition of done?

    • Evidence: PLAN-v0.4 — per-slice acceptance criteria; ROADMAP.md:16-20 — per-phase ship + verify; config.json:28-33 — verification automated.
    • Answer: Definition of done = per-slice acceptance criteria (each task has "Acceptance criteria") + per-phase ship (v0.1.7, v0.1.8) + verify stage. The grill is the P0 definition of done. This is the established pattern since v0.2.
    • Confidence: 0.85
    • Decision: G-020 — definition of done is per-slice acceptance criteria + per-phase ship + verify. Established pattern. Accept. (0.85)

Axis 6 — Budget and Financial Realism

  • Q1: Budget spent vs remaining?

    • Evidence: git log — v0.1 (foundation) + v0.2 (LXC deploy) + v0.3 (mastery+VC) shipped; v0.4 is the 4th milestone. No token budget tracked in .ciagent/ (token cost is implicit in the CI agent's operation).
    • Answer: No explicit token budget. The project has shipped 3 milestones (v0.1-v0.3) — the token cost is sunk. v0.4 is the 4th. Under full autonomy, the "budget" is the CI agent's operational cost (tokens + compute). No budget contingency is tracked. This is a pilot — the budget is "whatever it costs to ship the milestones." Not a financial-realism concern at pilot scale.
    • Confidence: 0.75
    • Decision: G-021 — no explicit token budget (pilot, full autonomy). v0.4 is the 4th milestone. Accept the implicit budget model. (0.75)
  • Q2: Predictable cost drivers not in original budget? (Postgres 16 in LXC = CT memory bump 4GB→6GB; new deps = larger Docker image; backup storage)

    • Evidence: RESEARCH-v0.4 §1.1 — CT memory 4GB→6GB (confirmed); PLAN-v0.4 TASK-02-02 — CT bump; TASK-02-03 — backup volume; ARCHITECTURE.md:737 — v0.4 CT sizing.
    • Answer: Three cost drivers:
      1. CT memory 4GB→6GB — budgeted (TASK-02-02). The 6GB figure has ~50% margin (RESEARCH-v0.4 §1.1).
      2. Larger Docker image — asyncpg + argon2-cffi + slowapi add ~10-20MB to the image. Negligible.
      3. Backup storage — pgbackups named volume, 7-day retention, pg_dump -Fc (compressed). At v0.4 scale (<100 learners), each dump is <1MB. 7 files = <7MB. Negligible.
    • Confidence: 0.85
    • Decision: G-022 — cost drivers are budgeted (6GB CT, backup volume). Image size + backup storage are negligible at pilot scale. Accept. (0.85)
  • Q3: Burn rate — how long until v0.4 ships at current pace?

    • Evidence: git log — v0.3 took ~1 day (commits from 2026-08-03 to 2026-08-04); v0.2 similar. v0.4 has 52 tasks vs v0.3's 40.
    • Answer: v0.3 shipped in ~1 day. v0.4 is +30% larger (52 vs 40 tasks). Expected: ~1.3 days of CI agent time. The burn rate is the CI agent's token consumption — not tracked, but the pace is established (3 milestones in ~3 days).
    • Confidence: 0.75
    • Decision: G-023 — burn rate: ~1.3 days estimated (analogous to v0.3). Accept. (0.75)
  • Q4: Budget contingent on anything?

    • Evidence: config.json:13 — full autonomy; config.json:39-43 — git auto-commit, no auto-push.
    • Answer: No. Full autonomy, no external approval, no contingent funding. The only contingency is the escalation_hooks (deploy, delete_data, merge_to_main) — none of which apply to v0.4 P0/P1/P2 execution (merge_to_main is P3, which is the final ship).
    • Confidence: 0.90
    • Decision: G-024 — no budget contingency (full autonomy, no external approval). Accept. (0.90)

Axis 7 — Risks, Assumptions, and Dependencies

  • Q1: Top 3 assumptions v0.4 rests on — evidence for each?

    • Evidence: RESEARCH-v0.4 risks table (R-MT-01, R-AUTH-01, R-DASH-01).
    • Answer:
      1. Postgres-in-LXC won't destabilize the learner service (R-MT-01). Evidence: RESEARCH-v0.4 §1.1 — Postgres idle ~400MB, praxis ~500MB, 6GB CT has ~50% margin. Postgres queries are off the voice path (operator endpoints + nightly aggregation only). The nightly jobs are at 03:00 CT. Confidence: 0.75 — the memory math is sound but the disk I/O contention during pg_dump is unmeasured. The mitigation (03:00 CT) is a scheduling assumption, not a measurement.
      2. k-anonymity ≥ 10 is sufficient privacy (D-034). Evidence: RESEARCH-v0.4 §3.1 — "k=10 is the textbook suppression pattern." Differencing attacks blocked by pre-defined 2-D views. Confidence: 0.70 — k=10 is the conventional minimum, but at pilot scale (1 learner) k-anon suppresses everything, which is privacy-correct but value-destroying. The assumption holds for privacy; it does not hold for dashboard utility at pilot scale.
      3. Signed stateless cookies are secure without TLS in the pilot (R-AUTH-01). Evidence: RESEARCH-v0.4 §2.4 — config-driven PRAXIS_COOKIE_SECURE, defense-in-depth (cohort dashboard reads only k-anonymized aggregates). Confidence: 0.65 — this is the signature question (see probe #1 below). The config-driven flag is a punt; the real mitigation is the k-anon defense-in-depth.
    • Confidence: 0.72
    • Decision: G-025 — 3 core assumptions: Postgres contention (0.75, unmeasured disk I/O), k-anon sufficiency (0.70, privacy-correct but value-destroying at pilot scale), cookie-without-TLS (0.65, config-driven punt with k-anon defense-in-depth). All accepted as pilot-scale constraints. (0.72)
  • Q2: Dependencies outside the team?

    • Evidence: config.json:13 — full autonomy; PROJECT.md:5.
    • Answer: None. Single project, full autonomy. No external departments, vendors, regulators, or customers. The only "external" dependency is the Proxmox cluster (v0.2 deployment) + Ollama Cloud + Deepgram + Cartesia (voice services) — all carried forward from v0.1-v0.2.
    • Confidence: 0.90
    • Decision: G-026 — no external dependencies (full autonomy). Accept. (0.90)
  • Q3: Single risk that kills v0.4? (R-VC-MIG-01 — losing the v0.3 public key breaks all issued VCs. Mitigation: archive before activate. Is this enough?)

    • Evidence: RESEARCH-v0.4 R-VC-MIG-01 (high severity, 0.85 confidence); PLAN-v0.4 SLICE-04 TASK-04-03 (migration script archives v0.3 public key BEFORE activating new key); TASK-06-05 (e2e test verifies v0.3 VC against Postgres store).
    • Answer: R-VC-MIG-01 is the single project-killing risk. If the v0.3 public key is lost, all v0.3 VCs break. The mitigation is correct: archive before activate (TASK-04-03 step 2 before step 3). The e2e test (TASK-06-05) verifies a v0.3 VC against the Postgres store with the archived superseded key. This is the right test. The risk is mitigated.
    • However, there is a subtle gap: the migration script (TASK-04-03) reads the v0.3 public key from SQLite. If the SQLite issuer_keys table is empty (e.g., the v0.3 pilot never issued a VC → no key was ever generated), the migration script's behavior is undefined. The script should handle "no v0.3 key exists" gracefully (skip the archive step, just generate a fresh v0.4 key). The plan says "Idempotent: if Postgres already has an active key, skip" but does not say "if SQLite has no active key, skip the archive."
    • Confidence: 0.80
    • Challenge: The migration script's behavior when SQLite has no v0.3 active key is unspecified. This is an edge case (the pilot may never have issued a VC), but it is the first-boot path for most deployments.
    • Decision: G-027 (MUST) — TASK-04-03 must explicitly handle the "no v0.3 active key in SQLite" case: if get_active_signing_key_row() on SQLite returns None, skip the archive step and only generate the fresh v0.4 keypair. Document this as a first-boot path. The e2e test (TASK-06-05) should include a "no v0.3 key" scenario. (0.80)
  • Q4: Pre-mortem — "It's 90 days from now and v0.4 failed. Why?"

    • Evidence: RESEARCH-v0.4 risks; PLAN-v0.4 risk matrix.
    • Answer: The most likely failure modes (in order):
      1. SPA fallback broke the voice UI (R-DASH-03/05). The catch-all route shadowed StaticFiles asset serving. The voice UI loaded but JS/CSS 404'd. The operator dashboard worked but the learner product regressed. This is the highest-blast-radius failure — it breaks the v0.1-v0.3 learner surface, not just the v0.4 operator surface.
      2. Postgres contention degraded the voice loop latency (R-MT-01). The nightly pg_dump + aggregation job at 03:00 CT caused disk I/O contention that spiked the voice loop latency >600ms. This was not caught because the latency test does not run with Postgres loaded.
      3. The secure-cookie+no-TLS tension was unresolved (R-AUTH-01). The config-driven flag was set to false for the pilot, the operator cookie was sniffed over HTTP on the vmbr0 bridge, and the grill should have caught that the config flag is a punt, not a fix.
      4. The k-anon dashboard showed nothing at pilot scale. The operator logged in, saw "— (<10 learners)" for every cell, and concluded the dashboard was broken. The grill should have caught that the dashboard's validation path is test-seeded data, not pilot traffic.
    • Confidence: 0.78
    • Decision: G-028 — pre-mortem top-4 failure modes: SPA fallback regression (highest blast radius), Postgres contention (unmeasured), R-AUTH-01 punt, k-anon-empty-dashboard. All four are addressed in this grill's binding decisions. (0.78)

Axis 8 — Governance, Decision-Making, and Communication

  • Q1: Decision-maker when two personas disagree?

    • Evidence: config.json:52-54 — lead-developer is the first persona; PERSONAS.md v0.4 — lead-developer "Coordinates task decomposition... resolves conflicts."
    • Answer: lead-developer is the decision-maker. This is the established pattern since v0.1.
    • Confidence: 0.85
    • Decision: G-029 — lead-developer is the conflict resolver. Accept. (0.85)
  • Q2: Governance cadence?

    • Evidence: ROADMAP.md:19 — pipeline stages SPECIFY → CLARIFY → RESEARCH → PLAN → GRILL → SHIP; config.json:105-108 — per-phase ship.
    • Answer: Per-phase ship + verify + grill at P0. This is the established cadence. The grill is the crisis-cadence (this document).
    • Confidence: 0.85
    • Decision: G-030 — governance cadence: per-phase ship + verify + grill. Accept. (0.85)
  • Q3: What's omitted from status reports? (R-AUTH-01 is the smell)

    • Evidence: RESEARCH-v0.4 §2.4 — R-AUTH-01 resolution documented as "the grill must sign off"; PLAN-v0.4:718 — risk matrix lists R-AUTH-01 with "grill must sign off."
    • Answer: The smell is R-AUTH-01. The research acknowledges the tension but frames the config-driven flag as a resolution. The v0.3 grill (Axis 4 MUST #2) explicitly rejected this approach: "Do not ship PRAXIS_COOKIE_SECURE=false as default — use TLS or loopback-binding." The v0.4 plan ships PRAXIS_COOKIE_SECURE defaulting to true with false for HTTP pilot — which is option (b) from the v0.3 grill (accept the pilot risk + document) wrapped in a config flag. The v0.3 grill rejected option (b). The v0.4 plan re-litigates this.
    • The real mitigation — the one the v0.3 grill did not consider — is the k-anon defense-in-depth: the cohort dashboard reads only k-anonymized aggregates, so even a sniffed cookie leaks no PII. This is the actual answer to R-AUTH-01, not the config flag.
    • Confidence: 0.70
    • Challenge: The plan's R-AUTH-01 resolution re-litigates a v0.3 grill MUST. The config-driven flag is a punt. The real mitigation (k-anon defense-in-depth) is buried in the research, not elevated.
    • Decision: G-031 (MUST) — R-AUTH-01 resolution must be reframed: the primary mitigation is the k-anon defense-in-depth (cohort dashboard reads only k-anonymized aggregates → sniffed cookie leaks no PII). The config-driven PRAXIS_COOKIE_SECURE flag is the secondary mitigation (operational convenience for when TLS arrives). The plan must document this ordering explicitly in TASK-03-02 and the GRILL-v0.4 ship notes. The v0.3 grill's "use TLS or loopback-binding" MUST is not satisfied — but the k-anon defense-in-depth is a new mitigation that the v0.3 grill did not evaluate (the v0.3 cohort dashboard was deferred). This grill accepts the k-anon defense-in-depth as the primary R-AUTH-01 resolution for v0.4, overriding the v0.3 grill's MUST #2 for the operator-tier surface only. (0.70)
  • Q4: Stop-the-project trigger?

    • Evidence: config.json:13 — full autonomy; config.json:15 — escalation_hooks: ["deploy", "delete_data", "merge_to_main"].
    • Answer: No human stop trigger (full autonomy). The CI agent can escalate (escalation_hooks) but cannot self-stop. The grill is the stop-the-project mechanism — if the verdict were "Rethink" or "Escalate," the project would stop. This grill's verdict is "Proceed-with-conditions," so the project proceeds.
    • Confidence: 0.80
    • Decision: G-032 — no human stop trigger (full autonomy). The grill is the stop mechanism. This grill = proceed with conditions. (0.80)

Axis 9 — Change, Adoption, and Operational Readiness

  • Q1: Who will use the cohort dashboard, and what's in it for them?

    • Evidence: D-052 — operator bootstrap is env-provided (not a real user); PROJECT.md:53 — "for training operators"; PERSONAS.md — no operator persona (operators are external to the CI agent).
    • Answer: The first operator is env-provided (D-052 — PRAXIS_BOOTSTRAP_OPERATOR_USER/PASS). There is no real operator user in the pilot. The dashboard is a capability demonstration, not a tool for a named user. "What's in it for them" = visibility into cohort progression, but at pilot scale (1 learner) the dashboard shows nothing (k-anon suppresses all cells). The dashboard's value is architectural (proving the operator tier works), not operational (no operator uses it yet).
    • Confidence: 0.65
    • Challenge: The dashboard has no real user at pilot scale. This is a placeholder deliverable — the capability exists, but no one uses it. The "we'll train them" answer does not apply (there is no "them").
    • Decision: G-033 — the cohort dashboard's first user is env-provided (D-052), not a real operator. At pilot scale (1 learner), the dashboard shows no data (k-anon). The dashboard is a capability demonstration for v0.5+ (when multi-learner data exists). Document this in the ship notes — v0.4 delivers the operator tier capability, not operator value. (0.65)
  • Q2: Is the operations team involved now or handed a finished product?

    • Evidence: PERSONAS.md v0.4 — devops-engineer is active in P1 (docker-compose Postgres + CT bump + backup + bootstrap); PLAN-v0.4 SLICE-02 — devops tasks.
    • Answer: devops-engineer is involved in P1 (SLICE-02 — .env.example, CT bump, backup script, bootstrap CLI). This is good — the operations surface is built by the operations persona, not handed off. The backup strategy (TASK-02-03) is devops-owned. The bootstrap CLI (SLICE-05) is devops-owned. The operations team is involved now.
    • Confidence: 0.85
    • Decision: G-034 — devops-engineer is involved in P1 (operations surface built by operations persona). Accept. (0.85)
  • Q3: Rollback plan if v0.4 goes wrong?

    • Evidence: PLAN-v0.4 — per-phase ship (v0.1.7, v0.1.8, v0.1.9) + git rollback; config.json:42-43 — branching_strategy: phase.
    • Answer: Per-phase git rollback (revert the patch tag). But:
      • P1 rollback (v0.1.7): Reverting P1 removes the Postgres service + auth. The VC key migration is irreversible — once the v0.3 public key is archived as superseded in Postgres and the fresh v0.4 key is active, reverting to v0.3 SQLite keys requires re-pointing the verification endpoint back to SQLite. The plan's fallback (TASK-06-03 — "if pg_store is None, fall back to PraxisStore path") makes this possible (set PRAXIS_PG_DSN to empty → server falls back to SQLite). This is a soft rollback — the Postgres data persists but is unused.
      • P2 rollback (v0.1.8): Reverting P2 removes the aggregation pipeline + dashboard. The SPA fallback catch-all route removal is clean (revert the route). The React Router addition is clean (revert package.json + App.tsx). The aggregation hook in session_recorder.py is clean (revert the chained task). P2 rollback is clean.
      • Postgres data migration is hard to roll back — but the plan does not migrate data (v0.3 credentials stay in SQLite; v0.4 credentials go to Postgres). The VC key archival is irreversible (the v0.3 public key is copied to Postgres as superseded), but this is additive — the v0.3 SQLite key still exists. Reverting to v0.3 means ignoring the Postgres copy.
    • Confidence: 0.75
    • Decision: G-035 — rollback is per-phase git revert. P1 rollback is soft (set PRAXIS_PG_DSN to empty → server falls back to SQLite). P2 rollback is clean (revert routes + package.json + session_recorder hook). VC key archival is additive (v0.3 SQLite key persists). Accept. (0.75)
  • Q4: Has anyone validated the success criteria with the people who will judge v0.4 successful?

    • Evidence: config.json:13 — full autonomy; config.json:28-33 — verification automated.
    • Answer: No human judge (full autonomy). The CI agent is the judge. The success criteria = 8/8 REQ-IDs covered + per-slice acceptance criteria + verify stage. This is the established pattern.
    • Confidence: 0.80
    • Decision: G-036 — CI is the judge (full autonomy). Success = 8/8 REQ coverage + acceptance criteria + verify. Accept. (0.80)

Meta — Closing Review

  • Q1: If you were the auditor, what would you flag?

    • Evidence: all axes above.
    • Answer: Three flags:
      1. R-AUTH-01 re-litigates a v0.3 grill MUST. The config-driven flag is a punt. The k-anon defense-in-depth is the real mitigation but is not elevated. (G-031)
      2. The k-anon dashboard shows nothing at pilot scale. The dashboard's validation path is test-seeded data, not pilot traffic. This is a placeholder deliverable. (G-033)
      3. Backup verification is a hidden requirement. A nightly pg_dump that is never restored is theater. (G-008)
    • Confidence: 0.78
    • Decision: G-037 — auditor flags: R-AUTH-01 re-litigation, k-anon-empty-dashboard, backup-verification gap. All addressed in binding decisions. (0.78)
  • Q2: What is v0.4 NOT doing that it should?

    • Evidence: PLAN-v0.4 open questions (742-754); RESEARCH-v0.4.
    • Answer:
      1. Backup restore drill — not tasked (G-008).
      2. Latency test with Postgres loaded — the voice loop latency test (TASK-06-04) checks that Postgres presence doesn't destabilize the learner service, but it does not run the voice loop under load with Postgres running the nightly job. The R-MT-01 disk I/O contention is unmeasured.
      3. Operator deactivation — no CLI to set is_active=false. Minor (SQL workaround), but a gap in the operator lifecycle.
      4. Differencing-attack test for k-anon — the v0.3 grill (Axis 7 FIX #2) asked for a differencing-attack test. The v0.4 plan (TASK-07-05) tests k-anon threshold (9 vs 10) but does NOT test that two adjacent 7-day windows cannot re-identify a single learner. This is a v0.3 grill FIX that is not explicitly carried forward.
    • Confidence: 0.75
    • Decision: G-038 (MUST) — Add a differencing-attack test to TASK-07-05 or TASK-10-03: seed 10 learners in window A, 9 in window B (1 dropped), verify the API does not allow a query that isolates the dropped learner. This is a v0.3 grill FIX (Axis 7 #2) that must be carried forward. (0.75)
  • Q3: Simplest possible v0.4 that delivers 80% of the value?

    • Evidence: D-053 — 3 dashboard views; PLAN-v0.4 SLICE-08, SLICE-09.
    • Answer: The simplest v0.4 = auth + Postgres + single dashboard view (practice volume only) + VC key migration. The mastery-progression and failure-patterns views are +20% value but +30% effort (2 more endpoints + 2 more React components + 2 more aggregation metrics). However: D-053 is a CLARIFY decision (0.80 confidence) that names 3 views — cutting to 1 would re-litigate a settled decision. The 3 views are not over-scoped relative to the decision. The simpler answer is: v0.4 is already the simplest version (8 REQs, no RBAC, no DP, no learner auth, single operator). Cutting further would break the v0.3 grill's deferred obligation.
    • Confidence: 0.75
    • Decision: G-039 — v0.4 is already the simplest version (8 REQs, single operator role, k-anon not DP). The 3-view dashboard is D-053 (settled). Further cuts would break the v0.3 grill obligation. Accept the scope. (0.75)
  • Q4: What would have to be true for v0.4 to succeed in the next 90 days, and is it true today?

    • Evidence: all axes.
    • Answer: For v0.4 to succeed:
      1. The SPA fallback must not break the voice UI. Is it true today? No — it is untested (TASK-10-04 is the test). Will be true after P2.
      2. The VC key migration must preserve v0.3 VC verification. Is it true today? No — it is untested (TASK-06-05 is the test). Will be true after P1.
      3. Postgres must not destabilize the learner service. Is it true today? Partially — the memory math is sound (6GB CT), but disk I/O contention is unmeasured. Will be true after P1 (with the 03:00 CT mitigation).
      4. The auth stack must be secure enough for a pilot. Is it true today? Partially — R-AUTH-01 is a punt with k-anon defense-in-depth. Will be true after G-031 reframes the mitigation.
      5. The dashboard must show something useful. Is it true today? No — at pilot scale (1 learner), k-anon suppresses everything. Will be true only with test-seeded data (≥10 mock learners).
    • Confidence: 0.72
    • Decision: G-040 — 5 success conditions: SPA fallback (untested), VC migration (untested), Postgres stability (partially), auth security (partially, G-031), dashboard utility (only with test-seeded data). All addressable in P1/P2. Accept with binding decisions. (0.72)

v0.4-Specific Probes (Signature Questions)

Question: D-030 said no Traefik/TLS for the pilot. D-041 requires Secure cookie attribute. Secure requires HTTPS. The research proposes PRAXIS_COOKIE_SECURE config-driven (default true, false for HTTP pilot). Is this a real resolution or a punt? What's the actual risk of running auth over HTTP in the LXC pilot? Is the cohort dashboard worth a TLS regression?

Evidence:

  • D-030 (PROJECT.md:160) — "vmbr0 DHCP only (pilot, no vmbr1, no Traefik proxy)."
  • D-041 (PROJECT.md:171) — "Cookie: httpOnly, secure, SameSite=Strict, 8h expiry."
  • RESEARCH-v0.4 §2.4 — config-driven flag, "cohort dashboard reads only k-anonymized aggregates → even a cookie sniffed over HTTP leaks no PII."
  • GRILL-v0.3.md Axis 4 MUST #2 — "Do not ship PRAXIS_COOKIE_SECURE=false as default — use TLS or loopback-binding."
  • server/__main__.py:46 — HOST = _env("PRAXIS_HOST", "0.0.0.0") (binds to all interfaces).

Analysis: The v0.3 grill explicitly rejected shipping PRAXIS_COOKIE_SECURE=false as a default. The v0.4 plan ships PRAXIS_COOKIE_SECURE defaulting to true with false for HTTP pilot — which is option (b) from the v0.3 grill (accept the pilot risk + document) wrapped in a config flag. This re-litigates the v0.3 grill MUST.

However, the v0.3 grill evaluated R-AUTH-01 before the cohort dashboard was scoped. The v0.3 grill's concern was "a cleartext cookie on a shared bridge is a MUST-FIX" — but the v0.3 grill did not know that the cohort dashboard would read only k-anonymized aggregates. The v0.4 research introduces a new mitigation: the k-anon defense-in-depth. A sniffed cookie gives the attacker access to /api/operator/*, which returns only k-anonymized cohort data (no PII) + the VC issuance log (credentials are public per D-043). The worst an attacker can do with a sniffed operator cookie is:

  • Read k-anonymized cohort aggregates (no PII — D-034).
  • Read the VC issuance log (credentials are public — D-043).
  • Revoke a VC (POST /api/operator/credentials/{id}/revoke) — this is a denial-of-service on a credential, but the credential is formative (v0.3 grill Axis 4 MUST #1) and the revocation is reversible (operator can re-issue).

The actual risk of running auth over HTTP in the LXC pilot is: an attacker on the vmbr0 bridge can sniff the operator cookie and revoke a formative credential. This is a low-severity risk for a pilot. The v0.3 grill's "MUST-FIX" was correct for a high-stakes credential — but the v0.3 grill itself downgraded the credential to formative (MUST #1), which also downgrades the R-AUTH-01 severity.

Resolution: The config-driven PRAXIS_COOKIE_SECURE flag is a punt — it does not fix the underlying tension. The real resolution is the k-anon defense-in-depth + the formative credential tier. The v0.3 grill's MUST #2 ("use TLS or loopback-binding") is overridden for the v0.4 operator-tier surface because:

  1. The cohort dashboard reads only k-anonymized aggregates (no PII leak from a sniffed cookie).
  2. The VC credential is formative (low-stakes — revocation is a reversible DoS, not a forgery).
  3. The pilot binds to vmbr0 DHCP (shared bridge) — but the pilot has 1 learner and 1 env-provided operator. The attack surface is theoretical.

Binding Decision G-031 (MUST) — R-AUTH-01 resolution: the primary mitigation is the k-anon defense-in-depth (sniffed cookie → no PII). The config-driven flag is secondary (operational convenience). The plan must document this ordering. The v0.3 grill's MUST #2 is overridden for v0.4 only because the v0.3 grill's own formative-credential decision (MUST #1) downgraded the R-AUTH-01 severity. This is a consistent override — the v0.3 grill's two MUSTs interact, and the formative tier + k-anon defense-in-depth together resolve the tension that either alone does not.

Confidence: 0.70 — the resolution is sound but re-litigates a prior grill MUST. The override is justified by the interaction of two v0.3 grill decisions (formative tier + k-anon), not by a single new fact.


Probe 2 — R-VC-MIG-01 (VC key migration): Is "archive before activate" enough?

Question: v0.3 issued VCs are in the field (hypothetically). v0.4 migrates the issuer key to Postgres. If the v0.3 public key is lost, all v0.3 VCs break. The plan says "archive before activate." Is that enough? Is there a test that verifies a v0.3 VC against the archived key after migration?

Evidence:

  • PLAN-v0.4 TASK-04-03 — migration script: step 2 (archive v0.3 public key as superseded) BEFORE step 3 (generate fresh v0.4 key).
  • PLAN-v0.4 TASK-06-05 — e2e test: "v0.3 VC verifies against Postgres store with archived superseded key (R-VC-MIG-01 explicitly verified)."
  • server/vc/issuer_keys.py:102-109 — get_public_key_for_verification queries by key_id (not status) — the fallback to superseded keys is implicit.
  • RESEARCH-v0.4 §5.3 — "No code change needed in the verification flow — only the store backing changes."

Analysis: "Archive before activate" is the correct ordering — if the migration fails between step 2 and step 3, the v0.3 key is archived but no v0.4 key is active. The verification endpoint would find the v0.3 key (superseded) and verify v0.3 VCs. New VCs cannot be issued (no active key) until the migration is re-run. This is a safe failure mode.

The e2e test (TASK-06-05) is thorough: it seeds a v0.3 VC, runs the migration, verifies the v0.3 VC against the Postgres store, issues a v0.4 VC, verifies it, tampers with the v0.3 VC (verification fails), and re-runs the migration (idempotent). This covers R-VC-MIG-01.

Gap (G-027): The migration script's behavior when SQLite has no v0.3 active key (the pilot never issued a VC) is unspecified. This is the first-boot path for most deployments. Must be handled.

Verdict: "Archive before activate" is enough with the e2e test (TASK-06-05) as the proof. The gap (no v0.3 key) is a binding decision (G-027). Confidence: 0.80.


Probe 3 — k-anonymity at pilot scale: Dashboard that shows nothing?

Question: v0.1-v0.3 used HARDCODED_LEARNER_ID = "learner-1" — a single learner. k-anonymity ≥ 10 will suppress EVERY cell in the cohort dashboard. The dashboard will show "— (<10 learners)" for everything. Is v0.4 building a dashboard that can't show any data until there are 10+ learners? Is that a real deliverable or a placeholder? What test data seeds ≥10 mock learners?

Evidence:

  • db/store.py:29 — HARDCODED_LEARNER_ID = "learner-1" (confirmed — single learner).
  • D-034 (PROJECT.md:164) — "k-anonymity ≥ 10."
  • REQ-NFR-DASH-01 — "cells with < 10 learners are suppressed."
  • PLAN-v0.4 Open Question #3 (line 746) — "For v0.4 (single learner), k-anonymity will suppress everything (1 < 10). This is expected at pilot scale (R-DASH-01). The executor should seed test data with ≥10 mock learners to verify the non-suppressed path."
  • PLAN-v0.4 TASK-10-03 — P2 integration test seeds 15 mock sessions (12 distinct learners) for the non-suppressed path + 5 sessions (5 learners) for the suppressed path.

Analysis: At pilot scale (1 learner), the dashboard shows "— (<10 learners)" for every cell. This is privacy-correct (k-anon is working) but value-destroying (the dashboard is useless). The plan acknowledges this (Open Question #3) and the validation path is test-seeded data (TASK-10-03 seeds 12 + 5 mock learners). The dashboard is a capability demonstration, not an operational tool — at pilot scale, no operator uses it (G-033).

This is a real deliverable in the sense that the capability exists (Postgres + aggregation + k-anon + auth + UI), but it is a placeholder in the sense that it cannot show real data until multi-learner-per-device is implemented (deferred). The v0.4 milestone delivers the plumbing, not the value.

Verdict: The dashboard is a placeholder deliverable at pilot scale. The validation path is test-seeded data (TASK-10-03), not pilot traffic. This must be documented in the ship notes (G-033). The k-anon suppression is correct behavior — the dashboard is working as designed. The issue is that the design is correct for a cohort but the pilot has one learner. Confidence: 0.75.


Probe 4 — Postgres-in-LXC resource contention (R-MT-01): Voice loop latency?

Question: Adding Postgres to the LXC CT bumps memory 4GB→6GB. The learner-facing voice loop has a <600ms latency budget (C-8). Will Postgres idle I/O + the aggregation pipeline degrade the voice loop? Is there a latency test that runs with Postgres loaded?

Evidence:

  • RESEARCH-v0.4 §1.1 — Postgres idle ~400MB, praxis ~500MB, 6GB CT has ~50% margin. "Postgres queries are off the voice path (operator endpoints + nightly aggregation only)."
  • R-MT-01 — "disk I/O contention during nightly pg_dump + aggregation." Mitigation: 03:00 CT.
  • PLAN-v0.4 TASK-06-04 — "Test that learner voice loop (/health, /pipecat/webrtc) is unaffected by auth (REQ-NFR-MT-01 — Postgres + learner service coexist)."
  • C-8 — latency budget < 600ms.

Analysis: The memory math is sound (6GB CT, ~1.3GB runtime, ~4.7GB headroom). The voice loop (WebRTC → Pipecat → ASR → LLM → TTS) does not touch Postgres — it uses SQLite for learner state (D-007 preserved) and the voice services (Deepgram, Cartesia, Ollama Cloud). Postgres is used only by operator endpoints + nightly aggregation. The risk is disk I/O contention during the nightly pg_dump + aggregation job (03:00 CT).

TASK-06-04 tests that Postgres presence doesn't destabilize the learner service — but it tests coexistence (health check passes, WebRTC offer accepted), not latency under load. The plan does NOT include a latency test that runs the voice loop while Postgres is executing the nightly job. The R-MT-01 mitigation (03:00 CT scheduling) is a scheduling assumption, not a measurement.

Verdict: The memory contention is well-mitigated (6GB CT). The disk I/O contention is unmeasured — the 03:00 CT mitigation is reasonable (low learner activity) but not proven. The voice loop does not touch Postgres, so the path is clean — the risk is system-level I/O contention, not application-level query contention. Confidence: 0.70 — the risk is low (Postgres is off the voice path) but unmeasured. Accept the 03:00 CT mitigation as a pilot-scale constraint.


Probe 5 — SPA fallback breaking voice UI (R-DASH-03/05): Route ordering?

Question: Adding a catch-all route for React Router /operator/* must not break the voice UI at /. The catch-all must be registered BEFORE StaticFiles but AFTER API routes. Is this ordering tested? What's the rollback if the voice UI breaks?

Evidence:

  • server/__main__.py:146 — app.mount("/", StaticFiles(directory=_CLIENT_DIST, html=True)) (current — no SPA fallback).
  • PLAN-v0.4 TASK-10-01 — catch-all route @app.get("/{path:path}") BEFORE StaticFiles.
  • PLAN-v0.4 TASK-10-04 — 8-assertion test (voice UI at /, SPA fallback for /operator/*, API routes return JSON, assets served by StaticFiles).
  • R-DASH-03 — "SPA fallback breaks existing voice UI (StaticFiles mount change)."

Analysis: The catch-all route @app.get("/{path:path}") is a greedy match — it matches every path. If registered before StaticFiles, it will intercept all GET requests, including /assets/index.js. The plan's TASK-10-01 says "the catch-all only serves index.html for client-side routes" but a @app.get("/{path:path}") route does not distinguish between client-side routes and static assets — it matches both. The correct implementation is either:

  1. A custom StaticFiles subclass that returns index.html for non-file paths (the plan's Open Question #2, line 744).
  2. A catch-all that excludes static asset paths (e.g., check if the path matches a file in client/dist first).

TASK-10-04 assertion 8 (GET /assets/index.js → served by StaticFiles, not the catch-all) is the test for this, but the implementation in TASK-10-01 is ambiguous. If the catch-all is registered before StaticFiles, FastAPI route matching order means the catch-all wins — StaticFiles never serves /assets/index.js. The plan's assertion 8 would fail.

The correct ordering is: API routes → StaticFiles mount → catch-all (for SPA fallback). But FastAPI's app.mount("/", StaticFiles(...)) is a catch-all at / — adding another catch-all after it is redundant (StaticFiles with html=True already serves index.html for /). The real fix is a custom StaticFiles subclass that returns index.html for non-file paths (Open Question #2).

Verdict: The plan's TASK-10-01 catch-all approach is subtly wrong — a @app.get("/{path:path}") before StaticFiles would shadow asset serving. The correct approach is a custom StaticFiles subclass (Open Question #2) OR a catch-all after StaticFiles that only fires for 404s. The plan defers this to EXECUTE (Open Question #2) but the test (TASK-10-04 assertion 8) would catch the bug. Confidence: 0.65 — the test is correct, the implementation is ambiguous. This is a binding decision.

Binding Decision G-041 (MUST) — TASK-10-01 must NOT use a @app.get("/{path:path}") catch-all before StaticFiles (it would shadow asset serving per assertion 8). The correct implementation is a custom StaticFiles subclass that returns FileResponse("client/dist/index.html") for non-file paths (Open Question #2 resolved in favor of the subclass approach). The catch-all approach is rejected. This must be documented in TASK-10-01 before EXECUTE. (0.65)


Probe 6 — 2-phase split: REQ-MT-02 spans P1 (schema) + P2 (pipeline). Vertical-slice violation?

Question: P1 (foundation) + P2 (dashboard) — is the split clean? REQ-MT-02 (aggregation) spans both phases (schema in P1, pipeline in P2). Is that a vertical-slice violation, or a clean layering?

Evidence:

  • PLAN-v0.4:18-19 — P1 covers "REQ-MT-02 (schema foundation)"; P2 covers "REQ-MT-02 (pipeline completion)."
  • PLAN-v0.4 REQ-ID coverage matrix (line 699) — REQ-MT-02: SLICE-01 (schema), SLICE-07 (pipeline), SLICE-10 (e2e).

Analysis: REQ-MT-02 is split across P1 (schema — the cohort_aggregates table) and P2 (pipeline — the aggregation hook + nightly job). This is not a vertical-slice violation — it is clean layering. The schema is the contract; the pipeline is the implementation. P1 ships the schema (the table exists, the PgStore has upsert_cohort_aggregate), P2 ships the pipeline (the hook fires, the nightly job runs). The P1→P2 dependency is one-directional (P2 depends on P1's schema, P1 does not depend on P2's pipeline).

This is the same pattern as v0.3 (mastery schema in P1, mastery flow in P1 — but the VC issuer was split, which the v0.3 grill flagged as a MUST). The difference is that REQ-MT-02's split is schema vs. pipeline (a clean layer), not trigger vs. action (the v0.3 grill's VC-issuance wiring gap). The aggregation pipeline does not need a P1 trigger — it fires on session-end, which is a P2 event (the hook is in session_recorder.py, which is extended in P2).

Verdict: The REQ-MT-02 split is clean layering (schema in P1, pipeline in P2), not a vertical-slice violation. The P1→P2 dependency is one-directional. The v0.3 grill's VC-issuance wiring gap (trigger in P1, action in P2) does not apply here — the aggregation trigger (session-end) is in P2. Confidence: 0.85.


v0.3 Grill Deferred Items — Coverage Check

The v0.3 grill (GRILL-v0.3.md) deferred the operator tier to v0.4. The v0.3 grill's MUST conditions were resolved in v0.3 (formative label, scoring_inconclusive, VC interop, key rotation, VC-issuance wiring). Let me verify the v0.3 grill's deferred items are now covered in v0.4:

v0.3 Grill Deferred Item v0.4 Coverage Status
REQ-DASH-01 (cohort dashboard) REQ-DASH-01 activated, PLAN SLICE-08/09/10 Covered
REQ-AUTH-01 (operator auth) REQ-AUTH-01 activated, PLAN SLICE-03/05/06 Covered
REQ-MT-01 (operator Postgres) REQ-MT-01 activated, PLAN SLICE-01/06 Covered
REQ-MT-02 (aggregation) REQ-MT-02 activated, PLAN SLICE-01/07/10 Covered
REQ-NFR-AUTH-01 (auth NFRs) REQ-NFR-AUTH-01 activated, PLAN SLICE-03/06 Covered
REQ-NFR-MT-01 (Postgres-in-LXC) REQ-NFR-MT-01 activated, PLAN SLICE-01/02/06 Covered
REQ-NFR-DASH-01 (k-anon ≥10) REQ-NFR-DASH-01 activated, PLAN SLICE-07/08/09/10 Covered
REQ-NFR-DASH-02 (freshness ≤24h) REQ-NFR-DASH-02 activated, PLAN SLICE-07/10 Covered

v0.3 grill FIX conditions carried forward to v0.4:

v0.3 Grill FIX v0.4 Coverage Status
Axis 7 #2 — k-anon differencing-attack test NOT explicitly in PLAN (TASK-07-05 tests threshold only) ⚠️ G-038 (MUST) — add differencing-attack test
Axis 6 #3 — 503 guard on operator API when Postgres down TASK-06-01 — "auth routes return 503" if no Postgres Covered
Axis 6 #2 — stabilize learner_ref as non-reusable UUID NOT addressed in v0.4 (HARDCODED_LEARNER_ID = "learner-1" persists) ⚠️ Accepted as pilot-scale constraint (G-012)

Verdict: 8/8 v0.3 deferred REQs are covered in v0.4. 1 v0.3 FIX (differencing-attack test) is not carried forward and must be added (G-038). The learner_ref stabilization (v0.3 FIX) is accepted as a pilot-scale constraint (single hardcoded learner persists).


Binding Decisions

ID Axis Decision Confidence Type
G-001 1 v0.4 operator tier is the correct next priority (delivers v0.3 grill's deferred obligation) 0.85 ACCEPT
G-002 1 CI is the named sponsor under full autonomy 0.80 ACCEPT
G-003 1 v0.4 is not a zombie; pilot-scale business value is low (k-anon suppresses all cells). Dashboard validation path = test-seeded data. Document in ship notes. 0.75 ACCEPT
G-004 1 No financial ROI; ROI is governance credibility + architectural foundation. Accept non-financial ROI. 0.65 ACCEPT
G-005 2 v0.4 scope is a clean handoff from v0.3 grill deferral. No scope creep. 0.90 ACCEPT
G-006 2 Requirements frozen (8 REQs, CI-owned under full autonomy) 0.85 ACCEPT
G-007 2 Out-of-scope is explicit and comprehensive 0.88 ACCEPT
G-008 2 MUST: Add backup-restore drill task to P1 — execute pg_restore, verify 5 tables + row counts. A backup that is never restored is theater. 0.70 MUST
G-009 3 Architecture is conventional (standard FastAPI + Postgres + React patterns), research-validated 0.80 ACCEPT
G-010 3 4 new deps, all single-purpose. slowapi fallback documented. Accept. 0.78 ACCEPT
G-011 3 MUST: Verification endpoint two-store fallback semantics must be explicit in TASK-04-04 + TASK-06-03 (not deferred to EXECUTE). Rule: Postgres for keys → SQLite fallback for v0.3 credentials → SQLite-only if no Postgres. 0.75 MUST
G-012 3 Three inherited debts acknowledged (SQLite VC keys, single learner, no TLS). Debts #2 and #3 accepted as pilot-scale constraints. 0.72 ACCEPT
G-013 4 Key-person dependency: security-engineer, data-engineer, backend-engineer. Accept under parallelization. 0.82 ACCEPT
G-014 4 6 personas available (4 config + 2 emergent), max 5 concurrent. 6>5 not binding. 0.80 ACCEPT
G-015 4 CI is the product owner with full authority 0.85 ACCEPT
G-016 4 Team building new capability (Postgres, auth, k-anon, React Router) — conventional patterns, thorough research. Accept for pilot. 0.75 ACCEPT
G-017 5 Phase structure set after scope understood. Not reverse-engineered. 0.85 ACCEPT
G-018 5 Critical-path risk: SPA fallback (R-DASH-03). Mitigation: TASK-10-04. Accept with test as gate. 0.75 ACCEPT
G-019 5 52 tasks is evidence-based (analogous to v0.3's 40, bottom-up sized) 0.80 ACCEPT
G-020 5 Definition of done = per-slice acceptance criteria + per-phase ship + verify 0.85 ACCEPT
G-021 6 No explicit token budget (pilot, full autonomy). Accept implicit budget model. 0.75 ACCEPT
G-022 6 Cost drivers budgeted (6GB CT, backup volume). Image + backup storage negligible. 0.85 ACCEPT
G-023 6 Burn rate: ~1.3 days estimated (analogous to v0.3) 0.75 ACCEPT
G-024 6 No budget contingency (full autonomy) 0.90 ACCEPT
G-025 7 3 core assumptions: Postgres contention (0.75), k-anon sufficiency (0.70), cookie-without-TLS (0.65). All accepted as pilot-scale constraints. 0.72 ACCEPT
G-026 7 No external dependencies (full autonomy) 0.90 ACCEPT
G-027 7 MUST: TASK-04-03 must handle "no v0.3 active key in SQLite" — skip archive, generate fresh v0.4 key only. First-boot path for most deployments. 0.80 MUST
G-028 7 Pre-mortem top-4: SPA fallback, Postgres contention, R-AUTH-01 punt, k-anon-empty-dashboard. All addressed. 0.78 ACCEPT
G-029 8 lead-developer is the conflict resolver 0.85 ACCEPT
G-030 8 Governance cadence: per-phase ship + verify + grill 0.85 ACCEPT
G-031 8 MUST: R-AUTH-01 resolution reframed — primary mitigation = k-anon defense-in-depth (sniffed cookie → no PII). Config-driven flag = secondary. v0.3 grill MUST #2 overridden for v0.4 operator surface because formative tier + k-anon together resolve the tension. Document ordering in TASK-03-02 + ship notes. 0.70 MUST
G-032 8 No human stop trigger (full autonomy). Grill is the stop mechanism. 0.80 ACCEPT
G-033 9 Dashboard's first user is env-provided (not real). At pilot scale, shows no data. Capability demonstration for v0.5+. Document in ship notes. 0.65 ACCEPT
G-034 9 devops-engineer involved in P1 (operations surface built by operations persona) 0.85 ACCEPT
G-035 9 Rollback is per-phase git revert. P1 = soft (empty DSN → SQLite fallback). P2 = clean. VC key archival = additive. 0.75 ACCEPT
G-036 9 CI is the judge (full autonomy). Success = 8/8 REQ + acceptance criteria + verify. 0.80 ACCEPT
G-037 Meta Auditor flags: R-AUTH-01 re-litigation, k-anon-empty-dashboard, backup-verification gap. All addressed. 0.78 ACCEPT
G-038 Meta MUST: Add differencing-attack test to TASK-07-05 or TASK-10-03 — v0.3 grill FIX (Axis 7 #2) carried forward. Seed 10 learners in window A, 9 in B, verify API cannot isolate the dropped learner. 0.75 MUST
G-039 Meta v0.4 is already the simplest version (8 REQs, single operator, k-anon not DP). 3-view dashboard is D-053 (settled). 0.75 ACCEPT
G-040 Meta 5 success conditions: SPA fallback (untested), VC migration (untested), Postgres stability (partial), auth security (partial, G-031), dashboard utility (test-seeded only). All addressable. 0.72 ACCEPT
G-041 Probe 5 MUST: TASK-10-01 must NOT use @app.get("/{path:path}") catch-all before StaticFiles (shadows asset serving). Use custom StaticFiles subclass returning index.html for non-file paths. Open Question #2 resolved in favor of subclass. 0.65 MUST

Escalations

None. All 9 axes + meta + 6 v0.4-specific probes are resolved with confidence ≥ 0.60. The 6 MUST conditions (G-008, G-011, G-027, G-031, G-038, G-041) are binding decisions with clear resolutions — they do not require human escalation (full autonomy). The lowest-confidence binding decision is G-041 (0.65 — SPA fallback implementation) which is above the 0.60 threshold.


MUST Conditions Summary (blocking — must be resolved in PLAN before EXECUTE)

  1. G-008 — Backup restore drill. Add a task to P1 that executes pg_restore --clean --if-exists against a test Postgres and verifies the 5 tables + row counts. A nightly pg_dump that is never restored is theater.

  2. G-011 — Verification endpoint two-store fallback semantics. TASK-04-04 + TASK-06-03 must explicitly document the fallback contract: (a) Postgres available → use it for key lookup (active + superseded); (b) Postgres available but credential not found → fall back to SQLite issued_credentials (v0.3 credentials); (c) Postgres NOT available (no DSN) → use existing v0.3 SQLite path for both keys + credentials. This is a binding contract, not an open question.

  3. G-027 — VC migration "no v0.3 key" edge case. TASK-04-03 must handle the case where SQLite has no active issuer key (the pilot never issued a VC): skip the archive step, generate only the fresh v0.4 keypair. The e2e test (TASK-06-05) must include a "no v0.3 key" scenario. This is the first-boot path for most deployments.

  4. G-031 — R-AUTH-01 resolution reframed. The primary mitigation for R-AUTH-01 is the k-anon defense-in-depth (cohort dashboard reads only k-anonymized aggregates → sniffed cookie leaks no PII). The config-driven PRAXIS_COOKIE_SECURE flag is secondary (operational convenience). The v0.3 grill's MUST #2 ("use TLS or loopback-binding") is overridden for the v0.4 operator-tier surface because the v0.3 grill's own formative-credential decision (MUST #1) + the k-anon defense-in-depth together resolve the tension. Document this ordering in TASK-03-02 and the v0.4 ship notes.

  5. G-038 — Differencing-attack test. Add a test to TASK-07-05 or TASK-10-03: seed 10 learners in window A, 9 in window B (1 dropped), verify the API does not allow a query that isolates the dropped learner. This is a v0.3 grill FIX (Axis 7 #2) that must be carried forward.

  6. G-041 — SPA fallback implementation. TASK-10-01 must NOT use a @app.get("/{path:path}") catch-all before StaticFiles (it would shadow asset serving — TASK-10-04 assertion 8 would fail). The correct implementation is a custom StaticFiles subclass that returns FileResponse("client/dist/index.html") for non-file paths. Open Question #2 is resolved in favor of the subclass approach.


FIX Conditions (non-blocking — tracked in VERIFY-P1/P2)

  • G-003 — Document in v0.4 ship notes: dashboard validation path is test-seeded data (≥10 mock learners), not pilot traffic. At pilot scale (1 learner), k-anon suppresses all cells.
  • G-012 — Document inherited debts: single hardcoded learner (k-anon suppresses pilot data), no TLS (R-AUTH-01 config-driven punt with k-anon defense-in-depth).
  • G-018 — SPA fallback (R-DASH-03) is the critical-path risk. TASK-10-04 (8 assertions) is the gate. If assertion 8 fails, the fix is the custom StaticFiles subclass (G-041).
  • G-025 — Postgres disk I/O contention (R-MT-01) is unmeasured. The 03:00 CT mitigation is a scheduling assumption. Accept as pilot-scale constraint.
  • G-033 — Document in ship notes: v0.4 delivers the operator tier capability, not operator value (no real operator user at pilot scale).

ACCEPT Items (proceed as-is)

  • v0.4 scope is a clean handoff from v0.3 grill (G-005).
  • Architecture is conventional (G-009).
  • 4 new deps are single-purpose (G-010).
  • Key-person dependency is manageable under parallelization (G-013).
  • Phase structure is not reverse-engineered (G-017).
  • 52 tasks is evidence-based (G-019).
  • No external dependencies (G-026).
  • Rollback is per-phase git revert (G-035).
  • REQ-MT-02 split (schema in P1, pipeline in P2) is clean layering, not a vertical-slice violation (Probe 6).
  • R-VC-MIG-01 "archive before activate" + e2e test is sufficient (Probe 2, with G-027 edge case).

Bottom Line

The v0.4 plan is not unfeasible — the research is thorough, the architecture is conventional, the phase split is clean, and the v0.3 grill's deferred obligation is honestly delivered. The plan is not over-scoped (8 REQs, single operator role, k-anon not DP). The plan is not under-tested in its highest-risk areas (R-VC-MIG-01 has a dedicated e2e test, R-DASH-03 has 8 assertions).

The 6 MUST conditions are surgical:

  • 2 are missing tasks (backup drill, differencing-attack test).
  • 2 are specification clarifications (verification endpoint fallback, VC migration edge case).
  • 1 is a reframing (R-AUTH-01: k-anon defense-in-depth is the primary mitigation, not the config flag).
  • 1 is an implementation correction (SPA fallback: custom StaticFiles subclass, not a catch-all route).

Resolve the 6 MUSTs, track the 5 FIXs, and v0.4 is a GO.