This repository has been archived on 2026-09-12. You can view files and clone it. You cannot open issues or pull requests or push a commit.
Files
praxis/.ciagent/GRILL-v0.5.md
T
Praxis CI ec397f2c65 docs(milestone): complete v0.5-live-assist — v0.1.13 tagged, milestone release, merged to main
v0.5 (Live Assist — on-the-job voice companion) milestone complete.
4 phases: P0 (pre-execution, v0.1.10) → P1 (assist core + guardrail,
v0.1.11) → P2 (integration + tech-debt + NFR, v0.1.12) → P3 (final
review + ship, v0.1.13 = milestone release).

16/16 REQs covered (3 ASSIST + 4 NFR + 9 IDEATE). 4 v0.6 backlog.
469 tests passed, 0 failed. 1 P0 fixed (guardrail processor safety).
8 P1+ flagged for v0.6. 8 v0.4 P1+ tech-debt addressed.
G-049 + G-067 grill MUSTs resolved. ESCALATION-01 (PIPEDA) OPEN for
human legal review before assist surface go-live.

---ci---
project: praxis
phase: 3
milestone: v0.5
status: complete
requirements:
  covered: [REQ-ASSIST-01, REQ-ASSIST-02, REQ-ASSIST-03, REQ-NFR-ASSIST-01, REQ-NFR-ASSIST-02, REQ-NFR-ASSIST-03, REQ-NFR-ASSIST-04, REQ-IDEATE-01, REQ-IDEATE-02, REQ-IDEATE-03, REQ-IDEATE-04, REQ-IDEATE-05, REQ-IDEATE-06, REQ-IDEATE-07, REQ-IDEATE-08, REQ-IDEATE-09]
  partial: []
---/ci---
2026-08-04 22:35:56 +00:00

628 lines
92 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# CIAgent Grill Report — v0.5 Live Assist (On-the-Job Voice Companion)
## Run: 2026-08-04 (mode: mechanical, focus: all axes + 6 v0.5-specific probes)
> **Reviewer:** adversarial technology executive (red-team)
> **Subject:** v0.5 execution plan (Live Assist — On-the-Job Voice Companion) — 2 execution phases, 12 slices, 33 tasks, 16 active REQs (3 ASSIST + 4 NFR + 9 IDEATE)
> **Stance:** plan is unfeasible, over-scoped, and too costly until evidence forces otherwise
> **Artifacts reviewed:** PROJECT.md (D-058..D-073), REQUIREMENTS.md (16 active REQs + 4 v0.6 backlog), ROADMAP.md, ARCHITECTURE.md (v0.5 Live Assist Mode §), RESEARCH-v0.5-live-assist.md (14 risks R-ASSIST-01..14, 7 domains), PLAN-v0.5-live-assist.md (2 phases, 12 slices, 33 tasks), PERSONAS.md (5 active, 2 deactivated), GRILL-v0.4.md (format reference + G-001..G-041), REVIEW.md (8 v0.4 P1+ carried forward), AUDIT.md (v0.4 HEALTHY), config.json (autonomy=full), server/pipeline.py, server/services/base.py, server/guardrails/customer_service.py, server/session_recorder.py, server/__main__.py
> **Binding status:** This grill verdict must be cleared (MUSTs resolved, escalations answered) before EXECUTE is authorized.
---
### Verdict: Proceed-with-conditions (confidence: 0.70)
The v0.5 plan is the project's first **safety-critical** milestone — the AI is in a learner's ear during *real* customer interactions, not role-play. This is a categorical shift from v0.1v0.4 (practice surface, no real customers, no real consequences). The plan's single most important decision — **D-071 (tap-to-talk only, wake-word deferred to v0.6)** — is the correct call: it strips the client-architecture risk (React-Web can't do foreground services), the battery risk, the Picovoice MAU-pricing risk, and 5 of 14 research risks (R-ASSIST-01/04/05/13/14 all become N/A). What remains is the *core* safety surface: the guardrail (REQ-ASSIST-03), the context-binding (REQ-ASSIST-02), and the shift-bounded session model (REQ-NFR-ASSIST-04). This is the right 80/20.
However, four material issues must be resolved before EXECUTE: (1) **R-ASSIST-07 (guardrail false-negative)** is the single project-killing risk — a direct answer slips past the regex, the learner parrots it to a real customer, trust erodes. The plan *accepts* this residual risk ("adversarial FN rate is reported but not threshold-gated" — PLAN:419) without a documented acceptance threshold or an escalation. For a safety-critical surface, "we'll measure it and trend it nightly" is necessary but not sufficient — the grill must set the bar. (2) **D-073 (PIPEDA consent-law review)** is deferred to "Phase 1 implementation" — but shipping a recording device into real customer interactions without legal sign-off is a regulatory risk the CI agent cannot resolve under full autonomy. This is an escalation, not a binding decision. (3) The IDEATE stage **expanded v0.5 scope from 7 REQs to 16** (+128%) — the first use of ideation in the project. The 9 added REQs are *defensive* (guardrail tuning, mode-conflict, PII policy, audit-log, reconnect, tech-debt, cost, NFR measurement), not feature creep — but the grill must verify the expansion is risk-reduction, not scope inflation. (4) The **in-loop guardrail processor** (post-LLM, pre-TTS) is a *structural pipeline change*, not the "minimal delta / prompt swap" the research frames it as — the v0.1 pipeline has no in-loop guardrail (the CS guardrail runs on the debrief, not in-loop per RESEARCH §5.2). This is the highest-novelty code in v0.5 and it is on the safety-critical path.
The plan is **not** over-scoped *after* the D-071 deferral (16 REQs, but 9 are defensive; 33 tasks vs v0.4's 52). It is **not** unfeasible (0 new pip/npm deps, v0.1 pipeline reused). It is **not** a zombie (Live Assist is the explicitly-deferred v0.1 surface, now delivered). The conditions are binding and surgical — but two of them (R-ASSIST-07 threshold, PIPEDA escalation) touch the safety-critical core and cannot be waived.
---
### Axis 1 — Business Case
- **Q1: What problem does Live Assist solve that the practice surface (v0.1-v0.4) doesn't? Is "on-the-job coaching" the top priority, or a feature looking for a user?**
- Evidence: PROJECT.md:45-47 — "v0.1v0.4 built and validated the practice surface… v0.5 adds the companion surface: a hands-free voice assistant a learner invokes *while actually working*"; RESEARCH-v0.5 §4.1 — "No direct competitor does live-in-ear coaching during real customer calls on a $100 phone" (verified: Dialpad/Gong post-hoc, RealWear AR+industrial); ROADMAP.md:9-11 — "the key distinction from the practice surface is real-customer interaction."
- Answer: Live Assist solves a problem the practice surface structurally cannot: coaching *during* real work, not *after* a role-play. The practice surface (v0.1-v0.4) teaches via simulated scenarios; Live Assist coaches during live customer interactions. This is the *transfer* moment — where practice meets the job. RESEARCH §4.1 confirms Praxis is novel (no competitor does this on a cheap phone). The priority is correct: v0.1-v0.4 built the practice foundation + operator visibility; v0.5 builds the transfer surface. The alternative (v0.6 low-bandwidth) would expand reach before the on-the-job value is proven.
- Confidence: 0.80
- Decision: **G-042** — Live Assist is the correct next priority (delivers the transfer surface the practice foundation was built for). Novel per RESEARCH §4.1. (0.80)
- **Q2: Who is the named executive sponsor for Live Assist specifically? (D-001 says "User-directed" for Canada — is there a sponsor for Live Assist?)**
- Evidence: config.json:13 — `"level": "full"`; PROJECT.md:5 — "Autonomy: full"; D-001 (PROJECT.md:171) — "Launch market = Canada… User-directed"; no named human sponsor for Live Assist in any `.ciagent/` file.
- Answer: No human sponsor. The CI agent is the executive sponsor under full autonomy — the established model since v0.1 (G-002 in GRILL-v0.4). The "sponsor makes a decision under pressure" test is met by this grill — the R-ASSIST-07 + PIPEDA decisions are the pressure decisions. D-001's "User-directed" applied to the *market* choice (Canada), not to Live Assist's scope.
- Confidence: 0.80
- Decision: **G-043** — CI is the named sponsor under full autonomy (no change from v0.1-v0.4 governance, G-002 carry-forward). (0.80)
- **Q3: What happens to the business if v0.5 is cancelled? (Does the v0.1-v0.4 practice surface work without it?)**
- Evidence: ROADMAP.md:149-157 — future milestones (v0.6 low-bandwidth, v0.7 multi-language) do not depend on Live Assist; PROJECT.md:64-69 — v0.4 operator tier + v0.3 mastery + v0.1 voice loop carry forward unchanged.
- Answer: If v0.5 is cancelled, the practice surface (v0.1-v0.4) continues to function. Live Assist is a *new surface*, not a dependency of the existing product. However, cancelling v0.5 means the *transfer* value (coaching during real work) is never delivered — the practice surface teaches, but the on-the-job bridge is missing. This is not a zombie (cancelling has a cost: the product's value proposition — "turn every smartphone into a master craftsperson that talks to you" — is unfulfilled without the live-coaching surface). But the practice surface is independently valuable.
- Confidence: 0.78
- Decision: **G-044** — v0.5 is not a zombie (delivers the transfer surface). The practice surface works without it, but the product's core promise (on-the-job coaching) is unfulfilled. Accept the non-zombie status. (0.78)
- **Q4: Is there an ROI calculation vs a counterfactual (skip to v0.6 low-bandwidth)?**
- Evidence: MISSING — no ROI calculation in any `.ciagent/` file. D-012 (PROJECT.md:182) — "v0.1 cost ceiling = no enforced ceiling (pilot)"; REQ-IDEATE-07 (REQUIREMENTS.md:70) — assist cost tracking added by ideation.
- Answer: No financial ROI. The counterfactual is "ship v0.5 vs skip to v0.6 (low-bandwidth)." Shipping v0.5 costs ~33 tasks of tokens + 0 new deps + the safety-critical guardrail work. Skipping to v0.6 would leave Live Assist permanently deferred (broken v0.1 out-of-scope promise: "Live Assist mode") and v0.6's low-bandwidth surfaces would build on a practice-only product with no on-the-job transfer. The ROI is *product-completeness* (delivering the v0.1-promised surface) + *safety-surface validation* (the guardrail work is the foundation for all future safety-critical domains per D-019). REQ-IDEATE-07 adds cost tracking — the *measurement* of ROI, not the calculation.
- Confidence: 0.68
- Decision: **G-045** — no financial ROI; the ROI is product-completeness (v0.1-promised surface) + safety-surface foundation (guardrail work extends D-019 for future domains). REQ-IDEATE-07 measures cost, doesn't justify it. Accept the non-financial ROI under full autonomy. (0.68)
---
### Axis 2 — Scope and Requirements
- **Q1: Is the scope stable? 16 active REQs + 4 v0.6 backlog — is this expanding?**
- Evidence: REQUIREMENTS.md:8-81 — 16 active REQs (3 ASSIST + 4 NFR + 9 IDEATE); PROJECT.md:49 — "3 REQs + NFRs TBD after RESEARCH/IDEATE"; PLAN-v0.5:1011 — "16/16 REQ-IDs covered"; git log `b8c7de8` — "ideation results — 9 accepted into v0.5, 4 accepted into v0.6."
- Answer: The scope **expanded** from 7 REQs (3 ASSIST + 4 NFR, post-CLARIFY) to 16 REQs (+9 IDEATE) — a +128% increase. This is the project's first use of the IDEATE stage. The 9 added REQs are: REQ-IDEATE-01 (guardrail tuning corpus), -02 (in-loop processor test), -03 (mode-conflict), -04 (measurable NFRs), -05 (PII policy), -06 (v0.4 tech-debt), -07 (cost tracking), -08 (WebRTC reconnect), -09 (incremental audit-log). **All 9 are defensive/risk-reduction, not features.** They address: guardrail false-positive/negative (the safety risk), mutual exclusivity (a correctness gap), PII (a privacy gap), NFR measurability (a verifiability gap), tech-debt (carried from v0.4), cost (C-3), resilience (WebRTC drop), audit completeness (abrupt termination). This is scope *hardening*, not scope *creep* — but it is still expansion, and the grill must verify each addition is risk-reduction, not gold-plating.
- Confidence: 0.78
- Challenge: The +128% expansion is the largest scope growth in the project's history (v0.4 was a clean handoff: 8 REQs, 0 added). The IDEATE stage is a new vector — without discipline, ideation becomes scope creep with a defensive veneer. The 9 REQs are individually justified, but the *aggregate* added 9 tasks of P1 surface + 4 P2 tasks. The grill accepts the expansion *because* each REQ maps to a named risk (R-ASSIST-06/07/08/09/11 + v0.4 P1+ findings), not because ideation is inherently good.
- Decision: **G-046** — scope expanded +128% via IDEATE (7→16 REQs). Accepted because all 9 additions are risk-reduction (guardrail, PII, mode-conflict, resilience, audit, tech-debt, cost, NFR measurability), not feature creep. Each maps to a named risk. Future ideation must maintain this risk-reduction discipline. (0.78)
- **Q2: Are requirements frozen? (The 4 NFRs were `pending-research``research-grounded` — are they stable now?)**
- Evidence: REQUIREMENTS.md:22-25 — 4 NFRs marked `research-grounded (R-ASSIST-XX)`; REQUIREMENTS.md:27 — "NFRs refined from `pending-research` to `research-grounded` after the v0.5 RESEARCH stage… Phase-1 measurement may further refine R-ASSIST-02 (latency) and R-ASSIST-14 (battery)."
- Answer: The 4 NFRs are *research-grounded*, not *frozen*. REQ-NFR-ASSIST-01 (latency) is explicitly "AT RISK" — estimated ~655ms, target <600ms, pilot tolerance ≤650ms (D-072). REQ-NFR-ASSIST-02 (hands-free) was refined by D-071 (tap-to-talk only, wake-word deferred). REQ-NFR-ASSIST-03 (guardrail) is refined by D-068 (regex + retry + fallback). REQ-NFR-ASSIST-04 (session model) is stable (D-062). The NFRs are *stable enough* for PLAN, but REQ-NFR-ASSIST-01's target is a *pilot tolerance* (≤650ms), not the binding constraint (<600ms) — this is a deferred hardening, not a freeze. REQ-IDEATE-04 adds measurable targets (p95 ≤650ms, FP<5%) — this *is* the freeze for measurement purposes.
- Confidence: 0.75
- Decision: **G-047** — NFRs are research-grounded, not frozen. REQ-NFR-ASSIST-01 (latency) is at-risk with a pilot tolerance (D-072); REQ-IDEATE-04 provides the measurable freeze (p95 ≤650ms pilot, FP<5%). Accept as pilot-scale with v0.6 hardening for <600ms. (0.75)
- **Q3: What is explicitly out of scope? (Is the v0.5 out-of-scope list as explicit as v0.4's?)**
- Evidence: PROJECT.md:54-62 — explicit out-of-scope list (9 items); REQUIREMENTS.md:83-92 — matching list.
- Answer: Explicitly out of scope: full multi-path launch, low-bandwidth surfaces (WhatsApp/USSD/offline), multi-language, persona switching, full operator-suite dashboard, learner auth/multi-learner-per-device, session recording/replay, proactive intervention, multi-modal. The list is as explicit as v0.4's. The key deferral is **wake-word (D-071)** — the original D-058 scope (wake-word + tap-to-talk) is reduced to tap-to-talk only, with wake-word deferred to v0.6. This is the largest scope *reduction* in v0.5 and it is explicit (D-071 binding, PLAN:25).
- Confidence: 0.85
- Decision: **G-048** — out-of-scope is explicit and comprehensive. D-071 (wake-word deferred) is the key scope reduction, documented as binding. (0.85)
- **Q4: Hidden requirements? (PIPEDA legal review D-073 — is this a hidden regulatory requirement?)**
- Evidence: D-073 (PROJECT.md:243) — "PIPEDA consent-law review = defer to v0.5 Phase 1 implementation"; R-ASSIST-08 (RESEARCH-v0.5 §2.6) — "Privacy/consent failure: the real customer didn't consent to being recorded/analyzed by an AI"; D-070 (PROJECT.md:240) — consent disclosure implemented regardless.
- Answer: **Yes — PIPEDA is a hidden regulatory requirement.** The ambient mic captures the real customer (a third party); ASR transcribes their speech; the turns table stores it (REQ-IDEATE-05 acknowledges this as "STRIDE information-disclosure"). Canada's PIPEDA + provincial one-party/two-party consent laws govern recording. D-073 defers the legal review to "Phase 1 implementation" and frames it as "not a Phase 0 blocker." The disclosure (D-070) is the *engineering* mitigation, but it is NOT a *legal* determination — a disclosure does not make recording legal if the law requires two-party consent. The CI agent under full autonomy cannot resolve a legal question. This is an **escalation**, not a binding decision — the grill cannot determine with confidence ≥0.60 whether the disclosure is sufficient or whether legal review must block ship.
- Confidence: 0.55
- Challenge: PIPEDA is a regulatory requirement that the plan defers. For a safety-critical surface with real customers, deferring legal review is a risk the CI cannot own. This must be escalated.
- Decision: **ESCALATION-01** — PIPEDA consent-law review (D-073) is a hidden regulatory requirement that cannot be resolved under full autonomy. The disclosure (D-070) is the engineering mitigation but not a legal determination. **Escalate to human attention:** determine whether Canada PIPEDA + provincial consent law requires explicit legal sign-off before shipping a recording device into real customer interactions. If the disclosure is legally sufficient, proceed; if two-party consent is required, the assist surface may need customer-facing consent (out of scope for v0.5) or geographic restriction. (0.55 — below threshold)
---
### Axis 3 — Architecture and Technical Feasibility
- **Q1: Has the assist pipeline architecture been validated? (D-061 says shares v0.1 pipeline — is build_assist_pipeline() validated or assumed?)**
- Evidence: server/pipeline.py:44-185 — `build_pipeline()` with `_build_transport` (line 63), `_build_stt` (line 76), `_build_llm` (line 89), `_build_tts` (line 109), `LatencyObserver` (line 183); RESEARCH-v0.5 §5.2 — "v0.5 adds a `build_assist_pipeline()`… Reuses `_build_transport`, `_build_stt`, `_build_llm`, `_build_tts` unchanged"; PLAN-v0.5 TASK-05-01 — `build_assist_pipeline()` assembles the pipeline.
- Answer: The v0.1 service constructors (`_build_transport/stt/llm/tts`) are verified present and reusable (pipeline.py:63-109). `build_assist_pipeline()` is *assumed* to reuse them — this is sound for the service layer. **However**, the in-loop guardrail processor (TASK-05-02 — `LiveAssistGuardrailProcessor` as a post-LLM, pre-TTS `FrameProcessor`) is a *structural pipeline change*, not a prompt swap. The v0.1 pipeline has NO in-loop guardrail processor — the CS guardrail runs on the debrief (post-session), not in-loop (RESEARCH §5.2: "the existing v0.1 pipeline doesn't have a post-LLM guardrail processor inline"). Inserting a frame processor between `llm` and `tts` is novel for this codebase. The research frames this as "~1 new Pipecat frame processor" (§5.2) — but Pipecat frame-processor semantics (when does `LLMFullResponseEndFrame` fire? can you inject a retry mid-stream?) are unvalidated. PLAN Open Question #4 (line 1046) defers the retry mechanism to EXECUTE: "verify Pipecat's `LLMContextAggregator` supports injecting a message + re-running the LLM within a single `process_frame` call. If not, the retry may need to be a separate pipeline task." This is the highest-novelty code in v0.5 and it is on the safety-critical path.
- Confidence: 0.70
- Challenge: The in-loop guardrail processor is a structural change deferred to EXECUTE. The retry mechanism (inject `RETRY_INSTRUCTION` + re-run LLM) is unvalidated against Pipecat's frame semantics. If Pipecat can't do mid-stream retry, the guardrail's "one retry" (D-068) becomes "canned fallback only" — a weaker safety posture.
- Decision: **G-049 (MUST)** — The in-loop guardrail processor's retry mechanism (TASK-05-02) must be validated against Pipecat's frame-processor semantics BEFORE Wave 3 (SLICE-05). Add a Wave-1 or Wave-2 spike task: "Verify `LLMFullResponseEndFrame` fires after the full LLM response + that `LLMContextAggregator` supports injecting a retry message + re-running the LLM within `process_frame`." If Pipecat cannot do mid-stream retry, document the fallback (canned fallback only, no retry) and update D-068's safety posture. This is a binding contract, not an open question. (0.70)
- **Q2: Integration surface — v0.4 cohort aggregation (D-062), v0.1 voice pipeline (D-061), v0.3 mastery (D-063). Each is an integration point. Risk of quiet cost doubling?**
- Evidence: PLAN-v0.5 SLICE-10 (aggregation extension), SLICE-05 (pipeline reuse), SLICE-01 (D-063 schedule_mastery=False); RESEARCH-v0.5 §6.1 — "no schema change to cohort_aggregates (the `metric` column is free-form TEXT)"; §4.3 — "D-063 is unambiguous: assist turns never update θ… `run_mastery_flow()` is invoked only for practice sessions."
- Answer: Three integration points, all *additive*:
1. **v0.4 cohort aggregation** — new `session_type='assist'` + 5 new metric strings (no schema change, D-062). Risk: low — the aggregator is metric-agnostic (RESEARCH §6.1, 0.90 confidence). But the aggregation cache persistence (v0.4 P1+ #7, REQ-IDEATE-06) directly corrupts `assist_active_learners_count` after restart — the tech-debt wave (SLICE-12) fixes this. **Dependency: the tech-debt fix is on the v0.5 critical path for correct assist metrics.**
2. **v0.1 voice pipeline**`build_assist_pipeline()` reuses services but adds the in-loop guardrail processor (see Q1). Risk: medium — the structural change is the novelty.
3. **v0.3 mastery separation**`schedule_mastery=False` for assist (D-063). Risk: low — the `end()` signature already supports the flag (RESEARCH §4.3, 0.90 confidence). Verified in code: `session_recorder.py` `end()` has `schedule_mastery` param.
- The cost-doubling risk is concentrated in the in-loop guardrail processor (Q1). The aggregation + mastery integrations are low-risk additive extensions.
- Confidence: 0.75
- Decision: **G-050** — 3 integration points, all additive. Cohort aggregation (low risk, metric-agnostic) + mastery separation (low risk, flag exists) + voice pipeline (medium risk, in-loop guardrail is structural). The aggregation cache tech-debt (P1+ #7) is on the critical path for correct assist metrics — SLICE-12 fixes it. Accept with G-049 (guardrail retry validation). (0.75)
- **Q3: Is there an existing system being replaced? (No — Live Assist is new. But does it inherit v0.1-v0.4 tech debt?)**
- Evidence: REVIEW.md:182-203 — 8 v0.4 P1+ findings; REQ-IDEATE-06 (REQUIREMENTS.md:64) — "Carry-forward the 8 v0.4 P1+ findings into the v0.5 backlog as a 'tech-debt wave'"; PLAN-v0.5 SLICE-12 — tech-debt wave (4 tasks).
- Answer: No existing system replaced — Live Assist is new. It inherits 8 v0.4 P1+ findings, budgeted in P2 SLICE-12 (REQ-IDEATE-06): (1) argon2id blocking, (2) rate-limit mock test, (3) cookie-secret length, (4) credential status enum, (5) revocation audit log, (6) nightly zoneinfo, (7) aggregation cache persistence, (8) f-string SQL. The most consequential for v0.5 is #7 (aggregation cache) — it directly corrupts `assist_active_learners_count` after restart. The tech-debt wave is in P2 (not P1) — this means the assist metrics are *incorrect* for all of P1 + early P2 until SLICE-12 ships. This is a *deferred fix on the critical path*.
- Confidence: 0.72
- Challenge: The aggregation cache fix (P1+ #7) is in P2 SLICE-12, but it corrupts v0.5's assist metrics during P1. The plan accepts this (P1 doesn't ship to operators — it's the assist voice loop). But if P1 ships as v0.1.11 (per-phase ship, config.json:110), the assist metrics are wrong in any P1 deployment. This is a *sequencing* issue, not a missing task.
- Decision: **G-051** — 8 v0.4 P1+ findings inherited, budgeted in P2 SLICE-12. The aggregation cache fix (P1+ #7) corrupts assist metrics during P1 — accept this because P1 ships the assist *voice loop* (no operator dashboard dependency), and the fix lands in P2 before operator visibility matters. Document in P1 ship notes: assist metrics are incorrect until P2 SLICE-12. (0.72)
- **Q4: Technical debt being inherited — is it budgeted for?**
- Evidence: PLAN-v0.5 SLICE-12 (4 tasks: cache persistence, cookie-secret, credential status, argon2id+rate-limit+audit+zoneinfo); REQ-IDEATE-06 (should priority, P1).
- Answer: Yes — budgeted in P2 SLICE-12 (4 tasks covering all 8 findings). The tech-debt wave is `should` priority (not `must`) — this is correct (the findings are non-blocking per REVIEW.md). The budget is 4 tasks in P2 Wave 2 — proportional to the 8 findings (some are one-liners: cookie-secret warning, zoneinfo swap).
- Confidence: 0.80
- Decision: **G-052** — tech-debt budgeted (4 tasks in P2 SLICE-12, `should` priority). Proportional to the 8 findings. Accept. (0.80)
---
### Axis 4 — People, Skills, and Organization
- **Q1: Key-person dependency — voice-engineer is REACTIVATED for the first time. Is there a knowledge concentration risk?**
- Evidence: PERSONAS.md:577-593 — voice-engineer REACTIVATED, owns 7 P1 tasks (largest territory: build_assist_pipeline, in-loop guardrail processor, warm WebRTC, reconnect, tap-to-talk client, latency tuning); PLAN-v0.5:102-108 — persona load distribution.
- Answer: The voice-engineer owns the largest P1 territory (7 tasks) and is activated for the *first time* in the project (proposed since v0.2 PERSONAS line 458, never operated). The in-loop guardrail processor + warm WebRTC + reconnect logic are all *new capabilities* this project has never built. If the voice-engineer is absent, the assist voice loop (SLICE-05, SLICE-06) has no owner — these are the core of v0.5. The security-engineer (6 tasks) owns the guardrail regex + tuning corpus — the other safety-critical path. The backend-engineer (6 tasks) owns the session API + context-binding. **Three personas are critical-path: voice-engineer, security-engineer, backend-engineer.** The voice-engineer is the highest key-person risk because the capability is *new* (no prior project experience), not just the territory.
- Confidence: 0.78
- Decision: **G-053** — key-person dependency: voice-engineer (new capability, largest territory), security-engineer (safety-critical guardrail), backend-engineer (session API + integration). All 3 critical-path. The voice-engineer is the highest risk (first activation, new capability). Accept under parallelization (max 5 concurrent, 5 active personas — exactly at the limit). (0.78)
- **Q2: Are the 5 active personas actually allocated? (CI agents, not humans. Are the agent capabilities sufficient for the voice-engineer territory?)**
- Evidence: config.json:22-27 — parallelization enabled, max 5 concurrent; PERSONAS.md:556-646 — 5 active personas; config.json:52-81 — only 4 personas in config.json array (voice-engineer + security-engineer are emergent, defined in PERSONAS.md).
- Answer: 5 active personas, max 5 concurrent — **exactly at the limit, no slack.** If all 5 are active in a wave, there is zero idle capacity for rework. P1 Wave 1 has 2 parallel slices (SLICE-01, SLICE-02) — 2 personas active (backend, backend+voice). P1 Wave 3 has 2 slices (SLICE-05, SLICE-06) — 2 personas (voice, voice). Peak parallelism is 2-3 slices per wave — within the 5-agent limit. The voice-engineer + security-engineer are NOT in config.json `personas` (emergent) — territory enforcement is `warn` (config.json:51), so they are not blocked. The capability question: the voice-engineer's frameworks (porcupine-android, webrtc, pipecat, piper-tts) are listed in PERSONAS.md but the voice-engineer has *never operated* in this project. The capability is *claimed*, not *demonstrated*. The in-loop guardrail processor (Q1, Axis 3) is the test of this capability.
- Confidence: 0.72
- Decision: **G-054** — 5 active personas, max 5 concurrent (at the limit, no slack). Peak parallelism 2-3 slices — within limit. Voice-engineer capability is claimed but undemonstrated (first activation). Accept with G-049 (guardrail retry validation) as the capability test. (0.72)
- **Q3: Is there a product owner with authority? (autonomy=full — the CI is the owner. Is that sound for a safety-critical surface?)**
- Evidence: config.json:13 — `"level": "full"`; PROJECT.md:5; config.json:34-38 — security auto_accept_low_severity, auto_mitigate_medium, escalate_high_severity.
- Answer: CI is the product owner under full autonomy — the established model since v0.1 (G-002, G-015 carry-forward). **For a safety-critical surface, this is the grill's hardest governance question.** The CI can auto-accept low-severity security issues + auto-mitigate medium — but R-ASSIST-07 (guardrail false-negative) is high-severity, and config.json:37 says `escalate_high_severity: true`. The plan *accepts* the residual risk (adversarial FN not threshold-gated) without escalating. This is a tension: the config says escalate high-severity, but the plan says accept. The grill must resolve this — either the residual risk is *not* high-severity (because defense-in-depth + audit + v0.6 LLM-as-judge mitigate it to medium), or the plan must escalate. See Probe 1.
- Confidence: 0.68
- Challenge: The CI-as-owner model is sound for practice surfaces (v0.1-v0.4) where the worst case is a bad role-play. For Live Assist, the worst case is a guardrail bypass during a real customer call. The config's `escalate_high_severity: true` is the safety valve — the plan must use it or justify why the risk is not high-severity.
- Decision: **G-055** — CI is the product owner (full autonomy, carry-forward). For the safety-critical surface, the `escalate_high_severity: true` config (config.json:37) is the governing constraint. R-ASSIST-07 (guardrail false-negative) is high-severity per RESEARCH — the plan must either (a) escalate it (Probe 1) or (b) document why defense-in-depth + audit + v0.6 LLM-as-judge reduce it to medium (auto-mitigatable). This is resolved in Probe 1. (0.68)
- **Q4: Is the team building capability it doesn't have? (voice-engineer is new — has the guardrail/latency/pipeline work been done before in this project?)**
- Evidence: RESEARCH-v0.5 §5.2 — "v0.5 adds an in-loop guardrail processor… the existing v0.1 pipeline doesn't have a post-LLM guardrail processor inline"; §3.3 — "prefill latency for gemma4:cloud is not yet measured (R3 from v0.1)"; PERSONAS.md:577-593 — voice-engineer frameworks include porcupine-android (not used in v0.5 per D-071), webrtc, pipecat.
- Answer: Yes — three new capabilities:
1. **In-loop Pipecat frame processor** — never built in this project. The v0.1 guardrail runs on the debrief (post-session), not in-loop. The frame-processor semantics (LLMFullResponseEndFrame, mid-stream retry) are unvalidated (G-049).
2. **Warm WebRTC connection lifecycle** — v0.1 opens per-session cold connections; v0.5 keeps a warm connection for an 8h shift with heartbeat + reconnect. New state machine (REQ-IDEATE-08).
3. **Regex guardrail tuning** — the CS guardrail (customer_service.py, 128 lines) is a fixed ruleset; v0.5 adds a tuning corpus + adversarial test + FP/FN measurement (REQ-IDEATE-01/04). New testing methodology.
- All three are on the safety-critical or critical path. This is *acceptable for a pilot* (learning-as-you-go is the project's model since v0.1) but the grill must flag that the highest-novelty code (in-loop processor) is also the highest-safety-impact code.
- Confidence: 0.72
- Decision: **G-056** — team is building 3 new capabilities (in-loop frame processor, warm WebRTC lifecycle, regex guardrail tuning). All on the safety-critical/critical path. Acceptable for pilot with G-049 (guardrail retry validation) as the de-risking spike. The voice-engineer's first activation is the capability test. (0.72)
---
### Axis 5 — Timeline and Estimates
- **Q1: Was the deadline set before or after the scope was understood? (No deadline — CI pipeline. Is the 2-phase split evidence-based or arbitrary?)**
- Evidence: ROADMAP.md:13-31 — v0.5 phases defined in ROADMAP (P0 pre-execution, P1 assist core, P2 integration, P3 review); PLAN-v0.5:17-25 — phase split rationale.
- Answer: No calendar deadline (CI pipeline). The 2-phase split is *evidence-based*: P1 = the assist voice loop + guardrail (the safety-critical, on-voice-path surface — 12 REQs, 24 tasks); P2 = integration + measurement + tech-debt (the operator-facing + hardening surface — 4 REQs, 9 tasks). The split mirrors v0.4 (P1 infra / P2 feature) but inverts it (P1 feature / P2 hardening). P1 is independently shippable (a learner can start a shift, tap-to-talk, get coaching with guardrails, end the shift). This is the correct split — the safety-critical surface ships first, the measurement + tech-debt follows.
- Confidence: 0.82
- Decision: **G-057** — 2-phase split is evidence-based (P1 safety-critical voice loop, P2 hardening + measurement). P1 independently shippable. Not arbitrary. (0.82)
- **Q2: Critical path — what single thing would push v0.5 by a phase? (Likely the guardrail — REQ-ASSIST-03 is safety-critical. Is the guardrail on the critical path?)**
- Evidence: PLAN-v0.5 wave dependency graph (P1:79-98); SLICE-03 (guardrail) → SLICE-04 (tuning corpus) → SLICE-05 (pipeline + in-loop processor) → SLICE-08 (e2e guardrail test); REQ-IDEATE-01 (tuning corpus + adversarial test).
- Answer: The guardrail is on the critical path (SLICE-03 → 04 → 05 → 08). The single thing that would push v0.5 by a wave:
- **Most likely: the guardrail tuning corpus fails FP<5% or direct-FN<5% (REQ-IDEATE-01).** TASK-04-02 asserts FP<5% on coaching responses + FN<5% on direct answers. If the regex over-matches (FP>5%) or under-matches (FN>5%), the regex needs retuning → pushes Wave 2 → Wave 3 → Wave 4. This is a *test-driven* gate — the tuning corpus is the proof.
- **Less likely: the in-loop guardrail processor retry mechanism is infeasible in Pipecat (G-049).** If Pipecat can't do mid-stream retry, the guardrail weakens to "canned fallback only" — still safe, but D-068's "one retry" is unmet. This would push Wave 3 (SLICE-05) by a spike.
- **Least likely: the warm WebRTC reconnect state machine (REQ-IDEATE-08).** The reconnect logic is specified (TASK-06-02) but the chaos test (TASK-06-03) is the proof. If the state machine has edge cases, it pushes Wave 3 (SLICE-06).
- Confidence: 0.75
- Decision: **G-058** — critical-path risk: guardrail tuning corpus (FP/FN rates, REQ-IDEATE-01). Mitigation: TASK-04-02 (test-driven gate). If FP>5% or direct-FN>5%, retune the regex → pushes by a wave. Accept with the test as the gate. G-049 (retry validation) de-risks the secondary path. (0.75)
- **Q3: Are the estimates evidence-based? (33 tasks across 2 phases — is this analogous to v0.4's 52 tasks/2 phases?)**
- Evidence: PLAN-v0.5:1064 — 33 tasks (24 P1 + 9 P2); GRILL-v0.4:166 — v0.4 had 52 tasks (29 P1 + 23 P2); GRILL-v0.4:19 — v0.3 shipped ~40 tasks.
- Answer: 33 tasks vs v0.4's 52 (-37%) and v0.3's 40 (-18%). The reduction is explained by D-071 (tap-to-talk only — wake-word deferral removed ~8-10 tasks: Porcupine integration, foreground service, battery management, OEM kill-switch handling) + 0 new deps (no dep-integration tasks). The scope is *smaller* than v0.4 despite +8 REQs (16 vs 8) because the IDEATE additions are mostly test/measurement tasks (low LOC) + the wake-word deferral stripped the client-architecture work. The tasks are bottom-up sized (each slice has 3-7 tasks with acceptance criteria). Evidence-based.
- Confidence: 0.80
- Decision: **G-059** — 33 tasks is evidence-based (smaller than v0.4's 52 due to D-071 wake-word deferral + 0 new deps; IDEATE additions are test/measurement tasks). Bottom-up sized. Accept. (0.80)
- **Q4: Definition of done — is "done" the grill's verdict or the verify stage's?**
- Evidence: PLAN-v0.5 — per-slice acceptance criteria; ROADMAP.md:19-21 — per-phase ship + verify; config.json:28-33 — verification automated.
- Answer: Definition of done = per-slice acceptance criteria + per-phase ship (v0.1.11, v0.1.12, v0.1.13) + verify stage. The grill is the P0 definition of done (this document). Established pattern since v0.2 (G-020 carry-forward). For the safety-critical surface, the *additional* done criterion is REQ-IDEATE-04's measurable NFRs (p95 ≤650ms, FP<5%) — these are the *quantitative* done bar for the guardrail.
- Confidence: 0.82
- Decision: **G-060** — definition of done = per-slice acceptance + per-phase ship + verify + REQ-IDEATE-04 measurable NFRs (p95 ≤650ms, FP<5%) as the quantitative guardrail bar. Established pattern + safety-critical addition. Accept. (0.82)
---
### Axis 6 — Budget and Financial Realism
- **Q1: Cost drivers — assist mode adds LLM calls (IDEATE-07 — 400 extra calls/month/learner). Is this in the budget?**
- Evidence: REQ-IDEATE-07 (REQUIREMENTS.md:70) — "20 turns/shift × 20 shifts/month = 400 extra LLM calls"; PLAN-v0.5 SLICE-11 — per-turn cost tracking + C-3 check; TASK-11-02 — `check_c3_budget()`.
- Answer: The cost driver is *budgeted* (SLICE-11, REQ-IDEATE-07). The estimate: 400 extra gemma4:cloud calls/month/learner at ~$0.0005/turn = ~$0.20/month — well under C-3's $3 (RESEARCH-v0.5, TASK-11-02). The cost is *diagnostic* (not enforced — D-012 says no enforced ceiling for pilot). The C-3 check (TASK-11-02) flags if practice + assist exceeds $3. This is the correct posture — measure, don't enforce, for the pilot.
- Confidence: 0.80
- Decision: **G-061** — assist cost driver budgeted (SLICE-11, ~$0.20/month, well under C-3). Diagnostic, not enforced (D-012 pilot relaxation). Accept. (0.80)
- **Q2: C-3 (≤$3/active learner/month) — does assist break it? (D-012 relaxed C-3 for the pilot, but is the relaxation still valid for v0.5?)**
- Evidence: D-012 (PROJECT.md:182) — "v0.1 cost ceiling = no enforced ceiling (pilot)"; GRILL-v0.4 G-012 — "no TLS → accepted as pilot-scale constraint"; REQ-IDEATE-07 — C-3 check.
- Answer: The C-3 relaxation (D-012) was set for v0.1 and carried through v0.4 (G-012). v0.5 adds ~$0.20/month/learner for assist — the total (practice + assist) is still well under $3 at pilot scale. The relaxation remains valid *for the pilot*. The architecture must not preclude meeting $3 post-pilot (D-012) — the assist cost is LLM calls, which the post-pilot path (self-hosted gemma4:e4b, D-020) reduces. The relaxation is valid for v0.5.
- Confidence: 0.78
- Decision: **G-062** — C-3 relaxation (D-012) remains valid for v0.5 pilot. Assist adds ~$0.20/month, total well under $3. Post-pilot path (self-hosted model) preserves the $3 target. Accept. (0.78)
- **Q3: Burn rate — token cost of 33 tasks + 2 phases + grill + review + audit. Is this proportional to v0.4?**
- Evidence: git log — v0.4 shipped in ~1.3 days (GRILL-v0.4 G-023); v0.5 has 33 tasks vs v0.4's 52 (-37%).
- Answer: v0.5 is ~37% smaller than v0.4 by task count. Expected burn: ~0.8-1.0 days of CI agent time (proportional reduction). The token cost is the CI agent's operational cost — not tracked, but the pace is established (4 milestones in ~4 days). Proportional.
- Confidence: 0.78
- Decision: **G-063** — burn rate: ~0.8-1.0 days estimated (proportional to v0.4, -37% tasks). Accept. (0.78)
- **Q4: Is the budget contingent on anything? (Porcupine pricing D-064 — MAU-priced, no recurring free tier. Is the pilot contingent on Picovoice sales engagement?)**
- Evidence: D-064 (PROJECT.md:234) — Porcupine MAU pricing; D-071 (PROJECT.md:241) — tap-to-talk only in v0.5, wake-word deferred to v0.6; R-ASSIST-01 (RESEARCH-v0.5 §1.2) — "no recurring free tier."
- Answer: **No — D-071 removed the Picovoice contingency.** The wake-word (Porcupine) is deferred to v0.6. v0.5 ships tap-to-talk only — no Porcupine dependency, no MAU pricing, no sales engagement needed. This is the single biggest budget de-risking of v0.5: the entire Picovoice commercial question is v0.6's problem, not v0.5's. The v0.5 budget is contingent on *nothing* external (0 new deps, no vendor engagement, full autonomy).
- Confidence: 0.85
- Decision: **G-064** — no budget contingency. D-071 (tap-to-talk only) removed the Picovoice MAU-pricing dependency. v0.5 has 0 external commercial dependencies. Accept. (0.85)
---
### Axis 7 — Risks, Assumptions, and Dependencies
- **Q1: Top 3 assumptions — evidence for each?**
- Evidence: RESEARCH-v0.5 risks (R-ASSIST-01..14); D-071, D-068, D-072.
- Answer:
1. **Tap-to-talk is sufficient UX (D-071).** Evidence: none — this is an *unvalidated* assumption. No user testing, no pilot data. The practice surface (v0.1-v0.4) uses a WebRTC connection per session; tap-to-talk is a button-hold pattern. Whether a learner on a real shift will tap a button on their phone (which may be in their pocket) is *untested*. The alternative (wake-word) is deferred to v0.6. **Confidence: 0.60** — the assumption is reasonable (tap-to-talk is a proven pattern for walkie-talkie apps) but unvalidated for this use case.
2. **Regex guardrail is adequate (D-068).** Evidence: RESEARCH §2.3 (0.78 confidence) — the regex patterns target direct-answer + false-authority + impersonation. The tuning corpus (REQ-IDEATE-01) + adversarial test will measure FP/FN. The adversarial FN rate is "reported but not threshold-gated" (PLAN:419) — this is a *residual risk acceptance*, not a proof of adequacy. **Confidence: 0.65** — the regex is the fast on-voice-path filter; the LLM-as-judge (v0.6) is the accurate off-voice-path backstop. Defense-in-depth is the mitigation, not regex alone.
3. **≤650ms latency is achievable (D-072).** Evidence: RESEARCH §3.3 — estimated ~655ms (Piper + lean prompt), unmeasured. The estimate is a *budget math* calculation, not a measurement. R1/R3/R4 (Deepgram/Ollama/Piper latencies) are unmeasured since v0.1. **Confidence: 0.65** — the budget math is sound but the actual latencies are unmeasured. D-072 accepts ≤650ms as pilot tolerance; <600ms is v0.6 hardening.
- Confidence: 0.63
- Decision: **G-065** — 3 core assumptions: tap-to-talk UX (0.60, unvalidated), regex guardrail adequacy (0.65, residual risk accepted), ≤650ms latency (0.65, unmeasured). All accepted as pilot-scale constraints with v0.6 hardening paths. The tap-to-talk assumption is the lowest-confidence — flag for v0.6 user testing. (0.63)
- **Q2: Dependencies — Picovoice (D-064, deferred to v0.6), PIPEDA (D-073), v0.4 cohort pipeline (D-062), v0.1 voice pipeline (D-061).**
- Evidence: D-071 (Picovoice deferred), D-073 (PIPEDA deferred), D-062 (cohort aggregation), D-061 (voice pipeline reuse).
- Answer:
- **Picovoice**: NOT a v0.5 dependency (D-071 — tap-to-talk only). Deferred to v0.6. ✅
- **PIPEDA**: Deferred to "Phase 1 implementation" (D-073). This is the escalation (ESCALATION-01, Axis 2). The disclosure (D-070) is the engineering mitigation. ⚠️
- **v0.4 cohort pipeline**: D-062 — additive extension (session_type=assist, new metric strings, no schema change). Verified: aggregator.py is metric-agnostic (RESEARCH §6.1, 0.90). ✅
- **v0.1 voice pipeline**: D-061 — service reuse (transport/stt/llm/tts) + in-loop guardrail processor (structural change, G-049). ⚠️
- The PIPEDA dependency is the only one that requires human attention. The others are internal + additive.
- Confidence: 0.75
- Decision: **G-066** — 4 dependencies: Picovoice (deferred, ✅), PIPEDA (escalation, ⚠️ — ESCALATION-01), cohort pipeline (additive, ✅), voice pipeline (structural change, ⚠️ — G-049). Accept the internal dependencies; escalate PIPEDA. (0.75)
- **Q3: Single risk that kills v0.5? (R-ASSIST-07 — guardrail false-negative reaches learner's ear during real customer call. Is there a mitigation beyond "defense-in-depth + post-v0.5 LLM-as-judge"?)**
- Evidence: R-ASSIST-07 (RESEARCH-v0.5 §2.6) — "The 'parrot' failure: the AI gives a verbatim script, the learner repeats it word-for-word, the customer detects the robotic delivery → trust erosion"; PLAN-v0.5:1025 — "defense-in-depth (prompt + regex + audit) + adversarial test + nightly FN trending + post-v0.5 LLM-as-judge (REQ-IDEATE-10, v0.6)"; PLAN:419 — "adversarial FN rate is reported but not threshold-gated."
- Answer: R-ASSIST-07 is the single project-killing risk. A direct answer that slips past the regex → learner parrots it → real customer hears robotic delivery → trust erosion + potential escalation. The mitigation is *defense-in-depth* (3 layers: prompt + regex + audit) + *measurement* (tuning corpus + adversarial test + nightly FN trending) + *future backstop* (v0.6 LLM-as-judge). **The gap: the adversarial FN rate is "reported but not threshold-gated" (PLAN:419).** This means the plan *accepts* an unknown residual risk without a ceiling. For a safety-critical surface, this is insufficient — the grill must set the bar. The bar cannot be "0% FN" (regex can't catch every paraphrase) — but it must be a *documented acceptance threshold* with an escalation if exceeded. config.json:37 says `escalate_high_severity: true` — R-ASSIST-07 is high-severity, so the plan must either escalate or document why the residual risk is acceptable.
- Confidence: 0.68
- Challenge: The plan accepts an unquantified residual risk on a safety-critical surface. "We'll measure it and trend it nightly" is necessary but not sufficient — what happens if the nightly trend shows 15% FN? The plan has no trigger. This is the grill's hardest call.
- Decision: **G-067 (MUST)** — R-ASSIST-07 (guardrail false-negative) must have a *documented acceptance threshold* before EXECUTE. The adversarial FN rate (REQ-IDEATE-01) must be: (a) measured pre-ship (TASK-04-02), (b) compared against a threshold (e.g., "adversarial FN ≤ 20% acceptable for pilot because defense-in-depth + audit + v0.6 LLM-as-judge mitigate; >20% triggers a re-tuning wave or escalation"), (c) the threshold + the mitigation rationale documented in the ship notes. This is NOT a "0% FN" demand — it is a "know your residual risk + decide if it's acceptable" demand. The plan's current "reported but not threshold-gated" is insufficient for a safety-critical surface. config.json:37 `escalate_high_severity: true` is the governing constraint. (0.68)
- **Q4: Pre-mortem — "It's 12 months from now and v0.5 failed. Why?"**
- Evidence: RESEARCH-v0.5 risks; PLAN-v0.5 risk matrix.
- Answer: The most likely failure modes (in order):
1. **A guardrail bypass incident during a real customer call (R-ASSIST-07).** A direct answer slipped past the regex, the learner parroted it, the customer escalated to a real manager who disavowed the "AI's advice." The nightly FN trend showed 18% but no one acted because there was no threshold (G-067 gap). This is the *highest-consequence* failure — it breaks trust in the product + the learner's job.
2. **PIPEDA complaint (R-ASSIST-08 / D-073).** A real customer discovered they were recorded by the learner's mic without their consent. The disclosure (D-070) was shown to the *learner*, not the *customer*. Canada's two-party consent law (if applicable in the province) was not reviewed. This is the *highest-legal-consequence* failure.
3. **The in-loop guardrail processor's retry mechanism was infeasible in Pipecat (G-049).** The "one retry" (D-068) became "canned fallback only" — safe but degraded. The assist coaching quality dropped (every block → canned fallback, no second chance). Learners stopped using assist because the coaching felt robotic.
4. **The latency was >650ms in practice (R-ASSIST-02).** The ~655ms estimate was optimistic; actual p95 was ~720ms. Coaching arrived after the customer moment passed. Learners abandoned assist for being "too slow to be useful."
- Confidence: 0.75
- Decision: **G-068** — pre-mortem top-4: guardrail bypass (highest consequence, G-067 gap), PIPEDA complaint (ESCALATION-01), in-loop retry infeasible (G-049), latency >650ms (D-072 pilot tolerance). All four are addressed in binding decisions/escalations. (0.75)
---
### Axis 8 — Governance, Decision-Making, and Communication
- **Q1: Decision-maker — autonomy=full, the CI decides. Is there a human escalation path for safety-critical decisions? (config.json escalation_hooks: deploy, delete_data, merge_to_main — none for "ship safety-critical guardrail". Is this a gap?)**
- Evidence: config.json:14 — `"escalation_hooks": ["deploy", "delete_data", "merge_to_main"]`; config.json:37 — `"escalate_high_severity": true`; PROJECT.md:5 — "Autonomy: full."
- Answer: The escalation_hooks list does NOT include "ship safety-critical guardrail" or "legal review." The `escalate_high_severity: true` security config is the *only* safety valve — it says the CI *should* escalate high-severity security issues, but the *mechanism* (how? to whom?) is unspecified. For v0.1-v0.4 (practice surface), this was acceptable — the worst case was a bad role-play. For v0.5 (Live Assist, real customers), the worst case is a guardrail bypass during a real call + a PIPEDA complaint. The escalation path for these is *the grill itself* — this document is the escalation mechanism. The grill's ESCALATION-01 (PIPEDA) + G-067 (guardrail threshold) are the safety-critical escalations/binding decisions. **The gap: there is no *ongoing* human escalation path post-ship.** If the nightly FN trend spikes post-ship, the CI auto-mitigates (config.json:36) but does not escalate to a human (no hook for "safety signal spike"). This is a v0.6+ governance gap, not a v0.5 blocker — v0.5 ships the measurement (REQ-IDEATE-04 nightly trending); v0.6 adds the LLM-as-judge + the escalation on spike.
- Confidence: 0.70
- Decision: **G-069** — escalation path: the grill is the safety-critical escalation mechanism (ESCALATION-01 + G-067). config.json `escalate_high_severity: true` is the governing constraint. Post-ship ongoing escalation (safety signal spike → human) is a v0.6+ governance gap — v0.5 ships the measurement, v0.6 adds the response. Accept for pilot with documented gap. (0.70)
- **Q2: Governance cadence — the pipeline stages are the governance. Is the grill the right gate for a safety-critical surface?**
- Evidence: ROADMAP.md:21 — "Pipeline stages: SPECIFY → CLARIFY → RESEARCH → IDEATE → PLAN → GRILL → SHIP"; ROADMAP.md:30 — "GRILL-v0.5.md (adversarial review — real-customer interaction warrants grill)."
- Answer: The grill is the right gate — ROADMAP.md:30 explicitly flags "real-customer interaction warrants grill." The pipeline stages (SPECIFY→…→GRILL→SHIP) are the governance cadence; the grill is the crisis-cadence (this document). For a safety-critical surface, the grill is the *only* human-in-the-loop checkpoint (the CI runs the rest autonomously). This is the correct model — the grill surfaces the safety-critical decisions (G-067, ESCALATION-01) for human attention before SHIP.
- Confidence: 0.82
- Decision: **G-070** — grill is the right gate for a safety-critical surface (ROADMAP:30 explicit). The grill is the human-in-the-loop checkpoint. Accept. (0.82)
- **Q3: What's omitted from status reports? (The LSP errors in server/__main__.py, test_scenario_library.py — are these reported or hidden?)**
- Evidence: Task context mentions "LSP errors in server/__main__.py, test_scenario_library.py"; verification: `python3 -m py_compile server/__main__.py` → exit 0 (clean); `python3 -m py_compile tests/test_scenario_library.py` → exit 0 (clean).
- Answer: The "LSP errors" claim in the task context is **unverified** — both files compile cleanly (`py_compile` exit 0). This may refer to type-checking (pyright/mypy) warnings, not syntax errors, or it may be stale. The grill does not flag this as a material omission — the files compile, the v0.4 tests pass (317 pass, 0 fail per REVIEW.md). If there are type-checking warnings, they are non-blocking (the codebase doesn't enforce strict typing in CI). **No omission found.**
- Confidence: 0.80
- Decision: **G-071** — no status-report omission found. The "LSP errors" claim is unverified (files compile clean). Type-checking warnings, if any, are non-blocking. Accept. (0.80)
- **Q4: Stop-the-project trigger — is there one? (If the grill returns RETHINK, does the pipeline stop?)**
- Evidence: config.json:13 — full autonomy; GRILL-v0.4 G-032 — "no human stop trigger (full autonomy). The grill is the stop mechanism."
- Answer: No human stop trigger (full autonomy, G-032 carry-forward). The grill is the stop mechanism — if the verdict were "Rethink" or "Escalate" on a material axis, the pipeline would stop. This grill's verdict is "Proceed-with-conditions" — the project proceeds after the MUSTs (G-049, G-067) + the escalation (ESCALATION-01) are resolved. The escalation (PIPEDA) is the *de facto* stop trigger — if the human legal review determines the disclosure is insufficient, v0.5 cannot ship the assist surface as designed.
- Confidence: 0.78
- Decision: **G-072** — no human stop trigger (full autonomy). The grill is the stop mechanism. ESCALATION-01 (PIPEDA) is the de facto stop trigger for the assist surface. This grill = proceed with conditions. (0.78)
---
### Axis 9 — Change, Adoption, and Operational Readiness
- **Q1: Who uses Live Assist? (The learner — during a real shift. How does their work change? They now have an AI in their ear.)**
- Evidence: PROJECT.md:45-47 — "a hands-free voice assistant a learner invokes *while actually working*"; PERSONAS.md — no learner persona (learners are external to the CI agent); D-071 — tap-to-talk invocation.
- Answer: The learner uses Live Assist during a real shift. Their work changes: they now have an AI coach in their ear (via earbuds) that they invoke by tapping a button (D-071 — tap-to-talk, not wake-word). "What's in it for them" = real-time coaching during real customer interactions — the transfer moment from practice to job. **This is unvalidated** — no user testing, no pilot data on whether learners will actually tap a button on their phone during a real customer call (the phone may be in their pocket, the tap may be socially awkward). The tap-to-talk UX (D-071) is the lowest-confidence assumption (G-065, 0.60). The alternative (wake-word, hands-free) is deferred to v0.6. For v0.5 pilot, tap-to-talk is the *validation* — does a learner use it? The measurement is the assist usage metrics (REQ-NFR-ASSIST-04, cohort aggregation).
- Confidence: 0.65
- Challenge: The adoption risk is *real* — tap-to-talk during a real customer call is socially + ergonomically awkward (phone in pocket, earbuds in, tap a button on the phone screen). The "we'll measure usage" answer is correct but the pilot may show low adoption. This is a v0.5 *validation* risk, not a v0.5 *blocker*.
- Decision: **G-073** — Live Assist's first user is the learner during a real shift. Tap-to-talk (D-071) is the unvalidated UX assumption (G-065, 0.60). v0.5 pilot *validates* adoption (assist usage metrics); v0.6 adds wake-word if tap-to-talk adoption is low. Document in ship notes: v0.5 validates the coaching/guardrail/context-binding value, not the hands-free UX (that's v0.6). (0.65)
- **Q2: Is the ops team involved? (CI project — ops is the LXC deploy. Does v0.5 need deploy changes? D-071 says no — v0.4 LXC carries forward. Is that sound?)**
- Evidence: PERSONAS.md:651-660 — devops-engineer DEACTIVATED for v0.5 ("No deploy changes — v0.4's LXC + Docker-in-LXC + Postgres + backup cron carries forward unchanged"); PLAN-v0.5:1071 — "New pip deps: 0… New npm deps: 0."
- Answer: v0.5 needs NO deploy changes — 0 new pip deps, 0 new npm deps, no new Docker services, no CT bump. The assist surface is server-side code (server/assist/) + a React route (client/src/AssistControl.tsx) on the existing v0.4 LXC. devops-engineer deactivation is sound. The ops surface (LXC, Postgres, backup) is unchanged. This is the correct posture — v0.5 is a *feature* milestone, not an *infra* milestone.
- Confidence: 0.85
- Decision: **G-074** — v0.5 needs no deploy changes (0 new deps, no CT bump, v0.4 LXC carries forward). devops-engineer deactivation is sound. Accept. (0.85)
- **Q3: Rollback plan — if v0.5 ships and a guardrail incident occurs, what's the rollback? (Disable assist mode? Revert to v0.1.9?)**
- Evidence: config.json:40 — `"branching_strategy": "phase"`; PLAN-v0.5 — per-phase ship (v0.1.11, v0.1.12, v0.1.13); git revert pattern (GRILL-v0.4 G-035).
- Answer: Rollback is per-phase git revert (G-035 carry-forward). But for a *guardrail incident* (R-ASSIST-07), the rollback is *operational*, not just git:
- **Preventive rollback**: disable assist mode (revert to v0.1.9 = v0.4). The assist routes (`/api/assist/*`) + the assist WebRTC endpoint are removed. The practice surface (v0.1-v0.4) continues unchanged. This is a clean revert — the assist surface is additive (new routes, new server/assist/ package, new SQLite migration 0004). Reverting removes the routes + the package; the migration is additive (session_type defaults to 'practice', guardrail_verdict_json is nullable) so existing practice sessions are unaffected.
- **Corrective rollback**: impossible. Once a guardrail bypass reaches a learner's ear during a real call, the turn has played. The audit log (REQ-IDEATE-09 incremental write) records it for investigation, but the *incident* cannot be rolled back. This is the nature of a live surface — rollback is preventive (disable), not corrective.
- The preventive rollback (disable assist) is clean + tested (the assist surface is additive). The corrective impossibility is accepted (the audit log is the post-incident tool, not a rollback).
- Confidence: 0.75
- Decision: **G-075** — rollback is preventive (disable assist mode → revert to v0.1.9). The assist surface is additive (clean revert). Corrective rollback is impossible (a live turn cannot be un-played) — the audit log (REQ-IDEATE-09) is the post-incident tool. Accept the preventive-only rollback. (0.75)
- **Q4: Has anyone validated the success criteria with the people who will judge v0.5 successful? (NFRs are research-grounded, not measurement-validated.)**
- Evidence: REQUIREMENTS.md:22-25 — NFRs `research-grounded`; REQ-IDEATE-04 — measurable targets (p95 ≤650ms, FP<5%); config.json:13 — full autonomy (CI is the judge).
- Answer: No human judge (full autonomy, G-036 carry-forward). The CI is the judge. The success criteria = 16/16 REQ coverage + per-slice acceptance + REQ-IDEATE-04 measurable NFRs. The NFRs are *research-grounded* (estimated, not measured) — REQ-IDEATE-04 + SLICE-09 (P2) add the *measurement*. The validation path: P2 SLICE-09 measures p95 latency + FP/FN rates. If p95 >650ms or FP>5%, the P2 verify stage flags it. This is the *measurement-validated* path — but it happens in P2, not pre-ship. **Gap: the success criteria are validated *during* P2, not *before* P1 ship (v0.1.11).** If P1 ships with a guardrail that has FP>5%, the P1 ship is premature. The mitigation: TASK-04-02 (guardrail tuning test) is in P1 Wave 2 — it runs *before* P1 ship. If it fails, P1 doesn't ship. This is the correct gate.
- Confidence: 0.72
- Decision: **G-076** — success criteria are research-grounded, measurement-validated in P2 (SLICE-09). The P1 gate is TASK-04-02 (guardrail tuning test, FP<5% / direct-FN<5%) — runs before P1 ship. If it fails, P1 doesn't ship. Accept with TASK-04-02 as the P1 gate + SLICE-09 as the P2 measurement. (0.72)
---
### Meta — Closing Review
- **Q1: If you were the auditor, what would you flag?**
- Evidence: all axes above.
- Answer: Four flags:
1. **R-ASSIST-07 residual risk acceptance without a threshold (G-067).** The plan accepts an unquantified adversarial FN rate on a safety-critical surface. This is the grill's hardest call — the bar must be set.
2. **PIPEDA legal review deferred (ESCALATION-01).** Shipping a recording device into real customer interactions without legal sign-off is a regulatory risk the CI cannot own.
3. **IDEATE scope expansion +128% (G-046).** The first use of ideation expanded v0.5 from 7 to 16 REQs. The additions are defensive, but the expansion is the largest in project history — future ideation must maintain risk-reduction discipline.
4. **In-loop guardrail processor is a structural pipeline change (G-049).** The research frames it as "~1 new frame processor" but the retry mechanism is unvalidated against Pipecat semantics. This is the highest-novelty code on the safety-critical path.
- Confidence: 0.78
- Decision: **G-077** — auditor flags: R-ASSIST-07 threshold gap, PIPEDA escalation, IDEATE scope expansion, in-loop processor novelty. All addressed in binding decisions/escalations. (0.78)
- **Q2: What is v0.5 NOT doing that it should? (PIPEDA legal review is deferred D-073 — should it block ship?)**
- Evidence: D-073 (PROJECT.md:243); ESCALATION-01 (Axis 2).
- Answer:
1. **PIPEDA legal review** — deferred, escalated (ESCALATION-01). The grill cannot determine if it blocks ship — that's a legal question. The disclosure (D-070) is the engineering mitigation; the legal review is the *regulatory* mitigation.
2. **Post-ship safety signal escalation** — the nightly FN trend (REQ-IDEATE-04) measures but does not escalate on spike (G-069). v0.6 adds the LLM-as-judge + the escalation response.
3. **Guardrail red-team prompt set** — REQ-IDEATE-01 builds a *synthetic* tuning corpus (LLM-generated coaching vs direct-answer responses). This is NOT a *human red-team* prompt set — a determined adversary (or a clever learner) may find paraphrases the synthetic corpus doesn't cover. The adversarial test (TASK-04-02) is the best available, but it's synthetic, not human. This is an accepted limitation (pilot).
- Confidence: 0.75
- Decision: **G-078** — v0.5 is NOT doing: PIPEDA legal review (escalated), post-ship safety escalation (v0.6), human red-team prompt set (synthetic corpus accepted for pilot). All documented. Accept with ESCALATION-01 as the human-action item. (0.75)
- **Q3: Simplest possible version — is tap-to-talk-only (D-071) the 80/20, or is even that over-scoped?**
- Evidence: D-071 (PROJECT.md:241); PLAN-v0.5:25 — "Why tap-to-talk only in v0.5."
- Answer: Tap-to-talk-only (D-071) IS the 80/20. The wake-word deferral stripped the client-architecture risk (React Native), the battery risk, the Picovoice commercial risk, and 5 of 14 research risks. What remains is the *core* value: the guardrail (safety), the context-binding (relevance), the shift-bounded session (aggregation), the in-loop processor (pipeline integration). Cutting further (e.g., deferring the guardrail tuning corpus REQ-IDEATE-01) would ship a guardrail "blind on its two most safety-critical metrics" (FP/FN) — unacceptable for a safety-critical surface. The 16 REQs are the *minimum* for a safe assist pilot. **Tap-to-talk is the 80/20; the guardrail work is the non-negotiable 20%.**
- Confidence: 0.80
- Decision: **G-079** — tap-to-talk-only (D-071) is the 80/20. The guardrail work (REQ-IDEATE-01/04, the tuning corpus + measurable NFRs) is the non-negotiable 20%. Cutting further would ship an unvalidated safety-critical surface. Accept the scope. (0.80)
- **Q4: What would have to be true for v0.5 to succeed in the next 90 days, and is it true today?**
- Evidence: all axes.
- Answer: For v0.5 to succeed:
1. **The guardrail must be robust (FP<5%, direct-FN<5%, adversarial FN known + thresholded).** Is it true today? No — the tuning corpus is built in P1 Wave 2 (TASK-04-01/02). Will be true after P1 if the test passes. G-067 sets the threshold.
2. **The in-loop guardrail processor must work in Pipecat (retry mechanism).** Is it true today? No — unvalidated (G-049). Will be true after the Wave-1/2 spike.
3. **PIPEDA must be addressed (legal review or disclosure-sufficient determination).** Is it true today? No — deferred (ESCALATION-01). Will be true only after human legal review.
4. **The latency must be ≤650ms.** Is it true today? No — unmeasured (D-072). Will be true after P2 SLICE-09 measurement.
5. **The tap-to-talk UX must be usable during a real shift.** Is it true today? No — unvalidated (G-065). Will be true only after pilot deployment (v0.5's validation purpose).
- 2 of 5 are addressable in P1/P2 (guardrail robustness, in-loop processor). 1 requires human action (PIPEDA). 2 are post-ship validation (latency measurement, UX adoption). This is the expected state for a pilot — the *plan* is ready; the *proof* is in execution.
- Confidence: 0.72
- Decision: **G-080** — 5 success conditions: guardrail robustness (P1 gate, G-067), in-loop processor (P1 spike, G-049), PIPEDA (human escalation, ESCALATION-01), latency (P2 measurement), UX adoption (post-ship validation). 2 addressable in P1/P2, 1 requires human, 2 post-ship. Accept — the plan is ready, the proof is in execution. (0.72)
---
### v0.5-Specific Probes (Signature Questions)
#### Probe 1 — R-ASSIST-07 (Guardrail false-negative): Is "defense-in-depth + audit + v0.6 LLM-as-judge" enough for a safety-critical surface?
**Question:** The AI is in a learner's ear during a *real* customer call. The regex output filter (D-068) is the on-voice-path guardrail. The adversarial FN rate is "reported but not threshold-gated" (PLAN:419). If a direct answer slips past the regex, the learner may parrot it. Is the 3-layer defense (prompt + regex + audit) + nightly trending + v0.6 LLM-as-judge sufficient, or does the grill need to set a binding threshold?
**Evidence:**
- R-ASSIST-07 (RESEARCH-v0.5 §2.6) — "The 'parrot' failure: the AI gives a verbatim script, the learner repeats it word-for-word, the customer detects the robotic delivery → trust erosion."
- D-068 (PROJECT.md:238) — "regex-based direct-answer + false-authority + impersonation patterns, with one retry on block + canned coaching redirect fallback."
- PLAN-v0.5:419 — "The adversarial FN rate is reported but not threshold-gated (it's the residual risk, mitigated by defense-in-depth)."
- config.json:37 — `"escalate_high_severity": true`.
- REQ-IDEATE-10 (v0.6 backlog) — "LLM-as-judge guardrail evaluation (nightly, off-voice-path) — measure the true false-negative rate the regex filter cannot."
**Analysis:**
The plan's posture is: regex is the fast on-voice-path filter (D-068); the LLM-as-judge is the accurate off-voice-path backstop (v0.6, REQ-IDEATE-10). The *gap* is v0.5: the regex is the only on-voice-path guardrail, and its adversarial FN rate is *unthresholded*. For a safety-critical surface where the worst case is a guardrail bypass during a real customer call, "we'll measure it and trend it nightly" is necessary but not sufficient — the plan needs a *decision*: what FN rate is acceptable for the pilot, and what happens if it's exceeded?
The config says `escalate_high_severity: true` — R-ASSIST-07 is high-severity. The plan *accepts* the residual risk without escalating. This is the tension G-055 identified. The resolution: the grill sets the threshold (G-067) — the adversarial FN rate must be measured pre-ship (TASK-04-02), compared against a documented threshold, and the threshold + mitigation rationale documented in the ship notes. This is NOT a "0% FN" demand (impossible for regex) — it is a "know your residual risk + decide if it's acceptable" demand.
The defense-in-depth (prompt + regex + audit) is the *correct* architecture — the grill does not dispute the 3-layer pattern (RESEARCH §2.1, 0.85 confidence). The issue is the *threshold*, not the architecture. The v0.6 LLM-as-judge is the *future* backstop, not the *current* mitigation — v0.5 ships with regex + audit only.
**Verdict:** Defense-in-depth is the correct architecture; the missing piece is a *documented acceptance threshold* for the adversarial FN rate. G-067 (MUST) sets this. The plan's "reported but not threshold-gated" is insufficient for a safety-critical surface — the grill requires a threshold + an escalation if exceeded. **Confidence: 0.68.**
---
#### Probe 2 — D-073 (PIPEDA consent-law review): Should legal review block ship?
**Question:** The ambient mic captures the real customer (a third party). ASR transcribes their speech. The turns table stores it (REQ-IDEATE-05). Canada's PIPEDA + provincial consent laws govern recording. D-073 defers the legal review to "Phase 1 implementation." The disclosure (D-070) is shown to the *learner*, not the *customer*. Is the disclosure sufficient, or does the legal review need to block ship?
**Evidence:**
- D-073 (PROJECT.md:243) — "PIPEDA consent-law review = defer to v0.5 Phase 1 implementation; document as R-ASSIST-08 in the grill."
- D-070 (PROJECT.md:240) — consent disclosure: "Praxis Assist is on — those around you may be recorded by your mic."
- R-ASSIST-08 (RESEARCH-v0.5 §2.6) — "the real customer didn't consent to being recorded/analyzed by an AI."
- REQ-IDEATE-05 (REQUIREMENTS.md:52) — "The ambient mic captures BOTH the learner and the real customer; ASR transcribes both; the turns table stores transcribed text. The customer is a third party."
- config.json:13 — full autonomy (CI cannot resolve legal questions).
**Analysis:**
This is a *legal* question, not a technical one. The CI agent under full autonomy cannot determine whether Canada's PIPEDA + provincial consent law requires:
- (a) One-party consent (the learner's consent is sufficient — the disclosure D-070 covers this).
- (b) Two-party consent (the *customer* must consent — Praxis cannot notify the customer, so the assist surface may be illegal in two-party provinces).
- (c) A PIPEDA-compliant privacy policy + data handling agreement.
The disclosure (D-070) is the *engineering* mitigation — it makes the *learner* aware. It does NOT make the *customer* aware, and it does NOT determine the legal consent regime. The PII policy (REQ-IDEATE-05) retains customer speech with redaction + 30-day retention — this is a *data handling* mitigation, not a *consent* determination.
The grill's confidence that the disclosure is sufficient: **0.55** — below the 0.60 threshold. The grill cannot resolve this under full autonomy. This is an escalation.
**Verdict:** PIPEDA legal review is a hidden regulatory requirement that the CI cannot resolve. The disclosure (D-070) is the engineering mitigation but not a legal determination. **Escalate to human attention** (ESCALATION-01): determine whether the disclosure is legally sufficient or whether two-party consent / a PIPEDA privacy policy is required before ship. If the disclosure is sufficient, proceed; if not, the assist surface may need geographic restriction or customer-facing consent (out of scope for v0.5). **Confidence: 0.55 — below threshold, escalated.**
---
#### Probe 3 — IDEATE scope expansion (+128%): Risk-reduction or scope creep?
**Question:** v0.5 started with 7 REQs (3 ASSIST + 4 NFR, post-CLARIFY). IDEATE added 9 REQs (+128%) — the largest scope growth in project history. Are the 9 additions risk-reduction (guardrail, PII, mode-conflict, resilience, audit, tech-debt, cost, NFR measurability) or scope creep with a defensive veneer?
**Evidence:**
- git log `b8c7de8` — "ideation results — 9 accepted into v0.5, 4 accepted into v0.6."
- REQUIREMENTS.md:29-70 — 9 IDEATE REQs.
- PLAN-v0.5:1011 — "16/16 REQ-IDs covered."
**Analysis:**
The 9 IDEATE REQs map to named risks:
- REQ-IDEATE-01 (guardrail tuning corpus) → R-ASSIST-06/07 (FP/FN).
- REQ-IDEATE-02 (in-loop processor test) → REQ-IDEATE-02 interface gap (GuardrailContext.role).
- REQ-IDEATE-03 (mode-conflict) → D-061 mutual exclusivity gap.
- REQ-IDEATE-04 (measurable NFRs) → REQ-NFR-ASSIST-01/03 verifiability.
- REQ-IDEATE-05 (PII policy) → R-ASSIST-08 (STRIDE information-disclosure).
- REQ-IDEATE-06 (tech-debt) → 8 v0.4 P1+ findings.
- REQ-IDEATE-07 (cost tracking) → C-3 budget.
- REQ-IDEATE-08 (WebRTC reconnect) → R-ASSIST-09.
- REQ-IDEATE-09 (incremental audit-log) → R-ASSIST-14 abrupt termination.
**Every addition maps to a named risk or a carried-forward finding.** None are features. The expansion is risk-reduction, not scope creep. The +128% is large but justified — v0.5 is the first *safety-critical* milestone, and the IDEATE stage surfaced the defensive requirements the practice surface (v0.1-v0.4) didn't need. The 4 deferred to v0.6 (REQ-IDEATE-10..13) are also risk-reduction (LLM-as-judge, assist-weaning, offline mode, voice-only context) — the ideation was disciplined.
**Verdict:** The IDEATE expansion is risk-reduction, not scope creep. Every REQ maps to a named risk. Accepted (G-046). Future ideation must maintain this discipline — the grill will flag any IDEATE addition that doesn't map to a named risk. **Confidence: 0.78.**
---
#### Probe 4 — In-loop guardrail processor (structural pipeline change): Is the "minimal delta" framing accurate?
**Question:** RESEARCH §5.2 frames the assist pipeline as "minimal delta: ~1 new pipeline builder, ~1 new guardrail processor." But the v0.1 pipeline has NO in-loop guardrail (the CS guardrail runs on the debrief). Is the in-loop processor a "minimal delta" or a structural change?
**Evidence:**
- server/pipeline.py:143-185 — `build_pipeline()` has no in-loop guardrail processor (transport → stt → latency → user_agg → llm → latency → tts → latency → transport → assistant_agg).
- RESEARCH-v0.5 §5.2 — "v0.5 adds an in-loop guardrail processor for assist mode. This is a pipeline-structure change but a small one (~1 new Pipecat frame processor)."
- server/guardrails/customer_service.py — CS guardrail runs `check()` standalone, not as a frame processor.
- PLAN-v0.5 TASK-05-02 — `LiveAssistGuardrailProcessor(FrameProcessor)` between llm and tts.
- PLAN-v0.5 Open Question #4 (line 1046) — "verify Pipecat's `LLMContextAggregator` supports injecting a message + re-running the LLM within a single `process_frame` call. If not, the retry may need to be a separate pipeline task."
**Analysis:**
The "minimal delta" framing is *partially accurate*. The service reuse (transport/stt/llm/tts) is genuinely minimal — the constructors are env-driven and reusable (verified: pipeline.py:63-109). **But the in-loop guardrail processor is a structural change**: the v0.1 pipeline has no post-LLM frame processor; v0.5 inserts one between `llm` and `tts`. This is novel for this codebase. The retry mechanism (inject `RETRY_INSTRUCTION` + re-run LLM mid-stream) is *unvalidated* against Pipecat's frame semantics — Open Question #4 defers this to EXECUTE, which is too late for a safety-critical path.
The risk: if Pipecat's `LLMFullResponseEndFrame` doesn't fire as expected, or if the `LLMContextAggregator` can't inject a retry mid-stream, the guardrail's "one retry" (D-068) becomes "canned fallback only" — safe but degraded. The coaching quality drops (every block → canned fallback, no second chance). This is a *quality* risk, not a *safety* risk (the canned fallback is safe) — but it affects the product's value.
**Verdict:** The in-loop guardrail processor is a structural change, not a minimal delta. The retry mechanism must be validated before Wave 3 (G-049 MUST). If Pipecat can't do mid-stream retry, document the fallback (canned-only) + update D-068's safety posture. The "minimal delta" framing should be corrected in the plan. **Confidence: 0.70.**
---
#### Probe 5 — Tap-to-talk UX (D-071): Is the unvalidated adoption risk acceptable for a pilot?
**Question:** D-071 ships tap-to-talk only (no wake-word). The learner taps a button on their phone during a real customer call. The phone may be in their pocket. The tap may be socially awkward. No user testing validates this UX. Is the pilot the validation, or is this a feature looking for a user?
**Evidence:**
- D-071 (PROJECT.md:241) — "tap-to-talk ONLY (no wake-word in v0.5)… learner taps a button to invoke an assist turn during a real shift."
- G-065 (Axis 7) — tap-to-talk UX assumption confidence 0.60 (lowest).
- RESEARCH-v0.5 §4.1 — "No direct competitor does live-in-ear coaching during real customer calls on a $100 phone" (novel surface, no comparable UX to benchmark).
**Analysis:**
Tap-to-talk is a *proven* pattern for walkie-talkie apps (Zello, Voxer) — users tap+hold to speak, release to send. This is a reasonable UX for hands-free-adjacent interaction. **But** those apps are *the* primary interface (the user opens the app to talk); Praxis assist is a *secondary* interface (the learner is in a real customer call, the phone is in their pocket, they tap a button on a screen they can't see). The social + ergonomic gap is real: the learner must (a) have earbuds in, (b) have the phone accessible, (c) tap a button without looking, (d) do this during a live customer interaction. This is a *high-friction* UX.
The pilot is the validation — v0.5 measures assist usage (REQ-NFR-ASSIST-04 cohort metrics). If adoption is low, v0.6 adds wake-word (the hands-free target). This is the correct pilot posture: ship the *value* (coaching/guardrail/context-binding), validate the *UX* (tap-to-talk adoption), iterate in v0.6. The risk is that low adoption makes the pilot a *failure* — but the pilot's purpose is to *find out*, not to *prove* adoption.
**Verdict:** Tap-to-talk is an unvalidated but reasonable UX for a pilot. The pilot is the validation. v0.6 adds wake-word if adoption is low. Accept with documented risk (G-073). **Confidence: 0.65.**
---
#### Probe 6 — 2-phase split: Is P1 (assist core + guardrail) independently shippable without P2 (measurement + tech-debt)?
**Question:** P1 ships v0.1.11 (assist core + guardrail, 12 REQs). P2 ships v0.1.12 (integration + tech-debt + NFR measurement, 4 REQs). Is P1 independently shippable — does a learner get a safe assist experience without P2?
**Evidence:**
- PLAN-v0.5:17-23 — P1 = assist voice loop + guardrail (12 REQs, 24 tasks); P2 = integration + measurement + tech-debt (4 REQs, 9 tasks).
- config.json:110 — `"per_phase": true` (per-phase ship).
**Analysis:**
P1 delivers: the assist voice loop (build_assist_pipeline), the 3-layer guardrail (LiveAssistGuardrail + tuning corpus + adversarial test), the shift-bounded session model, the tap-to-talk client, the warm WebRTC + reconnect, the incremental audit-log, the mode-conflict guard, the PII policy. A learner can start a shift, tap-to-talk, get coaching with guardrails, end the shift. **This is a safe, usable assist experience.**
P2 adds: the cohort aggregation assist metrics (operator visibility), the cost tracking (C-3 check), the NFR measurement (p95 latency, FP/FN rates), the tech-debt wave (8 v0.4 P1+ findings). **P2 is hardening + visibility, not safety.** The guardrail's safety is in P1 (SLICE-03/04/08); P2 *measures* the guardrail's FP/FN rates (SLICE-09) but the guardrail itself ships in P1.
The one caveat: the aggregation cache tech-debt (P1+ #7) corrupts `assist_active_learners_count` during P1 (G-051). But P1 doesn't ship operator visibility (the cohort dashboard extension is P2 SLICE-10) — so the corrupted metric is not *visible* during P1. The fix lands in P2 before the dashboard extension. This is a *sequencing* dependency, not a P1 safety gap.
**Verdict:** P1 is independently shippable — a learner gets a safe assist experience. P2 is hardening + operator visibility + measurement. The split is clean (P1 = safety-critical voice loop, P2 = hardening). The aggregation cache corruption during P1 is not visible (no dashboard in P1) and fixed in P2 before visibility. **Confidence: 0.82.**
---
### v0.4 Grill Deferred Items — Coverage Check
The v0.4 grill (GRILL-v0.4.md) deferred no items to v0.5 (v0.4 was the operator tier, complete). The v0.4 grill's 8 P1+ findings are carried forward as REQ-IDEATE-06 (tech-debt wave, P2 SLICE-12). Let me verify:
| v0.4 Grill/Finding | v0.5 Coverage | Status |
|---------------------|---------------|--------|
| G-008 (backup drill) | v0.4 complete (REVIEW.md:240) | ✅ Resolved in v0.4 |
| G-011 (two-store fallback) | v0.4 complete (REVIEW.md:241) | ✅ Resolved in v0.4 |
| G-027 (first-boot no v0.3 key) | v0.4 complete (REVIEW.md:242) | ✅ Resolved in v0.4 |
| G-031 (R-AUTH-01 reframe) | v0.4 complete (REVIEW.md:243) | ✅ Resolved in v0.4 |
| G-038 (differencing-attack test) | v0.4 complete (REVIEW.md:244) | ✅ Resolved in v0.4 |
| G-041 (SPA fallback subclass) | v0.4 complete (REVIEW.md:245) | ✅ Resolved in v0.4 |
| P1+ #1 (argon2id blocking) | REQ-IDEATE-06, TASK-12-04 | ✅ Covered in v0.5 P2 |
| P1+ #2 (rate-limit mock test) | REQ-IDEATE-06, TASK-12-04 | ✅ Covered in v0.5 P2 |
| P1+ #3 (cookie-secret length) | REQ-IDEATE-06, TASK-12-02 | ✅ Covered in v0.5 P2 |
| P1+ #4 (credential status enum) | REQ-IDEATE-06, TASK-12-03 | ✅ Covered in v0.5 P2 |
| P1+ #5 (revocation audit log) | REQ-IDEATE-06, TASK-12-04 | ✅ Covered in v0.5 P2 |
| P1+ #6 (nightly zoneinfo) | REQ-IDEATE-06, TASK-12-04 | ✅ Covered in v0.5 P2 |
| P1+ #7 (aggregation cache) | REQ-IDEATE-06, TASK-12-01 | ✅ Covered in v0.5 P2 (critical path for assist metrics — G-051) |
| P1+ #8 (f-string SQL) | REQ-IDEATE-06, TASK-12-03 | ✅ Covered in v0.5 P2 |
**Verdict:** 6/6 v0.4 grill MUSTs resolved in v0.4. 8/8 v0.4 P1+ findings covered in v0.5 P2 SLICE-12 (REQ-IDEATE-06). The aggregation cache fix (P1+ #7) is on the v0.5 critical path for correct assist metrics (G-051).
---
### Binding Decisions
| ID | Axis | Decision | Confidence | Type |
|----|------|----------|-----------|------|
| G-042 | 1 | Live Assist is the correct next priority (delivers the transfer surface). Novel per RESEARCH §4.1. | 0.80 | ACCEPT |
| G-043 | 1 | CI is the named sponsor under full autonomy (G-002 carry-forward). | 0.80 | ACCEPT |
| G-044 | 1 | v0.5 is not a zombie (delivers the transfer surface). Practice surface works without it. | 0.78 | ACCEPT |
| G-045 | 1 | No financial ROI; ROI is product-completeness + safety-surface foundation. REQ-IDEATE-07 measures cost. | 0.68 | ACCEPT |
| G-046 | 2 | IDEATE scope expanded +128% (7→16 REQs). Accepted — all 9 additions are risk-reduction, map to named risks. Future ideation must maintain discipline. | 0.78 | ACCEPT |
| G-047 | 2 | NFRs are research-grounded, not frozen. REQ-NFR-ASSIST-01 at-risk (D-072 pilot tolerance). REQ-IDEATE-04 provides measurable freeze. | 0.75 | ACCEPT |
| G-048 | 2 | Out-of-scope is explicit. D-071 (wake-word deferred) is the key scope reduction, binding. | 0.85 | ACCEPT |
| **G-049** | **3** | **MUST: In-loop guardrail processor retry mechanism (TASK-05-02) must be validated against Pipecat frame semantics BEFORE Wave 3. Add a Wave-1/2 spike: verify LLMFullResponseEndFrame + LLMContextAggregator retry injection. If infeasible, document canned-fallback-only + update D-068. Binding contract, not open question.** | **0.70** | **MUST** |
| G-050 | 3 | 3 integration points, all additive. Cohort aggregation (low) + mastery separation (low) + voice pipeline (medium, G-049). Aggregation cache tech-debt on critical path (G-051). | 0.75 | ACCEPT |
| G-051 | 3 | 8 v0.4 P1+ findings inherited, budgeted in P2 SLICE-12. Aggregation cache fix corrupts assist metrics during P1 — accept (P1 ships voice loop, not operator dashboard). Document in P1 ship notes. | 0.72 | ACCEPT |
| G-052 | 3 | Tech-debt budgeted (4 tasks in P2 SLICE-12, `should` priority). Proportional. | 0.80 | ACCEPT |
| G-053 | 4 | Key-person: voice-engineer (new capability, largest territory), security-engineer (guardrail), backend-engineer (session API). Voice-engineer highest risk (first activation). | 0.78 | ACCEPT |
| G-054 | 4 | 5 active personas, max 5 concurrent (at limit, no slack). Peak parallelism 2-3 slices. Voice-engineer capability claimed but undemonstrated — G-049 is the test. | 0.72 | ACCEPT |
| G-055 | 4 | CI is product owner (full autonomy). For safety-critical surface, `escalate_high_severity: true` governs. R-ASSIST-07 must be escalated or documented as medium (Probe 1). | 0.68 | ACCEPT |
| G-056 | 4 | Team building 3 new capabilities (in-loop processor, warm WebRTC, regex tuning). All on safety-critical/critical path. Acceptable for pilot with G-049 de-risking. | 0.72 | ACCEPT |
| G-057 | 5 | 2-phase split evidence-based (P1 safety-critical voice loop, P2 hardening + measurement). P1 independently shippable. | 0.82 | ACCEPT |
| G-058 | 5 | Critical-path: guardrail tuning corpus (FP/FN rates). TASK-04-02 is the gate. G-049 de-risks secondary path. | 0.75 | ACCEPT |
| G-059 | 5 | 33 tasks evidence-based (smaller than v0.4's 52 due to D-071 + 0 new deps). Bottom-up sized. | 0.80 | ACCEPT |
| G-060 | 5 | Definition of done = per-slice acceptance + per-phase ship + verify + REQ-IDEATE-04 measurable NFRs (p95 ≤650ms, FP<5%). | 0.82 | ACCEPT |
| G-061 | 6 | Assist cost driver budgeted (SLICE-11, ~$0.20/month, well under C-3). Diagnostic, not enforced. | 0.80 | ACCEPT |
| G-062 | 6 | C-3 relaxation (D-012) remains valid for v0.5 pilot. Assist adds ~$0.20/month. Post-pilot path preserves $3. | 0.78 | ACCEPT |
| G-063 | 6 | Burn rate: ~0.8-1.0 days estimated (proportional to v0.4, -37% tasks). | 0.78 | ACCEPT |
| G-064 | 6 | No budget contingency. D-071 removed Picovoice MAU-pricing dependency. 0 external commercial dependencies. | 0.85 | ACCEPT |
| G-065 | 7 | 3 core assumptions: tap-to-talk UX (0.60, unvalidated), regex guardrail (0.65, residual risk), ≤650ms latency (0.65, unmeasured). All pilot-scale with v0.6 hardening. | 0.63 | ACCEPT |
| G-066 | 7 | 4 dependencies: Picovoice (deferred ✅), PIPEDA (escalation ⚠️), cohort pipeline (additive ✅), voice pipeline (structural ⚠️ G-049). | 0.75 | ACCEPT |
| **G-067** | **7** | **MUST: R-ASSIST-07 (guardrail false-negative) must have a documented acceptance threshold before EXECUTE. Adversarial FN rate (REQ-IDEATE-01) must be: (a) measured pre-ship (TASK-04-02), (b) compared against a threshold (e.g., "≤20% acceptable for pilot because defense-in-depth + audit + v0.6 LLM-as-judge mitigate; >20% triggers re-tuning or escalation"), (c) threshold + rationale documented in ship notes. Not a "0% FN" demand — a "know your residual risk + decide" demand. config.json:37 escalate_high_severity governs.** | **0.68** | **MUST** |
| G-068 | 7 | Pre-mortem top-4: guardrail bypass (G-067 gap), PIPEDA (ESCALATION-01), in-loop retry (G-049), latency >650ms (D-072). All addressed. | 0.75 | ACCEPT |
| G-069 | 8 | Escalation path: grill is the safety-critical mechanism (ESCALATION-01 + G-067). Post-ship ongoing escalation (safety spike → human) is v0.6+ gap. Accept for pilot. | 0.70 | ACCEPT |
| G-070 | 8 | Grill is the right gate for safety-critical surface (ROADMAP:30 explicit). Human-in-the-loop checkpoint. | 0.82 | ACCEPT |
| G-071 | 8 | No status-report omission. "LSP errors" claim unverified (files compile clean). Type-checking warnings non-blocking. | 0.80 | ACCEPT |
| G-072 | 8 | No human stop trigger (full autonomy). Grill is the stop mechanism. ESCALATION-01 (PIPEDA) is the de facto stop trigger for the assist surface. | 0.78 | ACCEPT |
| G-073 | 9 | Live Assist's first user is the learner during a real shift. Tap-to-talk (D-071) is unvalidated UX (0.60). v0.5 validates adoption; v0.6 adds wake-word if low. | 0.65 | ACCEPT |
| G-074 | 9 | v0.5 needs no deploy changes (0 new deps, no CT bump, v0.4 LXC carries forward). devops-engineer deactivation sound. | 0.85 | ACCEPT |
| G-075 | 9 | Rollback is preventive (disable assist → revert to v0.1.9). Assist surface is additive (clean revert). Corrective rollback impossible (live turn cannot be un-played) — audit log is post-incident tool. | 0.75 | ACCEPT |
| G-076 | 9 | Success criteria research-grounded, measurement-validated in P2 (SLICE-09). P1 gate = TASK-04-02 (guardrail tuning test, FP<5%/FN<5%). P2 = SLICE-09 measurement. | 0.72 | ACCEPT |
| G-077 | Meta | Auditor flags: R-ASSIST-07 threshold gap, PIPEDA escalation, IDEATE scope expansion, in-loop processor novelty. All addressed. | 0.78 | ACCEPT |
| G-078 | Meta | v0.5 NOT doing: PIPEDA legal review (escalated), post-ship safety escalation (v0.6), human red-team prompt set (synthetic corpus accepted for pilot). | 0.75 | ACCEPT |
| G-079 | Meta | Tap-to-talk-only (D-071) is the 80/20. Guardrail work (REQ-IDEATE-01/04) is the non-negotiable 20%. Cutting further ships an unvalidated safety-critical surface. | 0.80 | ACCEPT |
| G-080 | Meta | 5 success conditions: guardrail robustness (P1 gate), in-loop processor (P1 spike), PIPEDA (human escalation), latency (P2 measurement), UX adoption (post-ship). Plan ready, proof in execution. | 0.72 | ACCEPT |
---
### Escalations
**ESCALATION-01 — PIPEDA consent-law review (D-073, R-ASSIST-08).** Confidence: 0.55 (below 0.60 threshold).
The ambient mic captures the real customer (a third party); ASR transcribes their speech; the turns table stores it (REQ-IDEATE-05). Canada's PIPEDA + provincial one-party/two-party consent laws govern recording. D-073 defers the legal review to "Phase 1 implementation." The disclosure (D-070) is shown to the *learner*, not the *customer* — it is the engineering mitigation, not a legal determination.
**The CI agent under full autonomy cannot resolve a legal question.** This must be escalated to human attention:
1. **Determine the consent regime:** Does Canada PIPEDA + the pilot province's consent law require one-party consent (learner's consent sufficient — D-070 covers) or two-party consent (customer must consent — Praxis cannot notify the customer)?
2. **If one-party:** the disclosure (D-070) is sufficient. Proceed with v0.5.
3. **If two-party:** the assist surface may need geographic restriction (one-party provinces only) or customer-facing consent (out of scope for v0.5 — would block the assist surface in two-party provinces).
4. **If a PIPEDA privacy policy / data handling agreement is required:** the PII policy (REQ-IDEATE-05, 30-day retention + redaction) may need to be formalized into a PIPEDA-compliant policy before ship.
**Action required:** Human legal review of Canada PIPEDA + provincial consent law for ambient recording during coaching, before v0.5 SHIP. The grill cannot determine with confidence ≥0.60 whether the disclosure is sufficient. This is the de facto stop trigger for the assist surface (G-072).
---
### MUST Conditions Summary (blocking — must be resolved before Phase 1 EXECUTE)
1. **G-049 — In-loop guardrail processor retry validation.** Add a Wave-1/2 spike task: verify Pipecat's `LLMFullResponseEndFrame` fires after the full LLM response + that `LLMContextAggregator` supports injecting a retry message + re-running the LLM within `process_frame`. If infeasible, document the fallback (canned-fallback-only, no retry) + update D-068's safety posture. This is a binding contract, not an open question (PLAN Open Question #4 must be resolved pre-EXECUTE).
2. **G-067 — R-ASSIST-07 guardrail false-negative acceptance threshold.** The adversarial FN rate (REQ-IDEATE-01) must be: (a) measured pre-ship (TASK-04-02), (b) compared against a *documented threshold* (e.g., "≤20% acceptable for pilot because defense-in-depth + audit + v0.6 LLM-as-judge mitigate; >20% triggers a re-tuning wave or escalation"), (c) the threshold + mitigation rationale documented in the v0.5 ship notes. The plan's current "reported but not threshold-gated" (PLAN:419) is insufficient for a safety-critical surface. config.json:37 `escalate_high_severity: true` is the governing constraint.
---
### Escalations Requiring Human Attention (before SHIP)
**ESCALATION-01 — PIPEDA consent-law review.** Determine whether Canada PIPEDA + provincial consent law requires one-party or two-party consent for ambient recording during coaching. If the disclosure (D-070) is legally sufficient, proceed. If two-party consent is required, the assist surface may need geographic restriction or customer-facing consent (out of scope for v0.5). This is the de facto stop trigger for the assist surface.
---
### FIX Conditions (non-blocking — tracked in VERIFY-P1/P2)
- **G-046** — Document in v0.5 ship notes: IDEATE expanded scope +128% (7→16 REQs). All additions are risk-reduction. Future ideation must maintain risk-reduction discipline.
- **G-051** — Document in P1 ship notes: assist metrics (assist_active_learners_count) are incorrect during P1 due to the aggregation cache tech-debt (v0.4 P1+ #7). Fix lands in P2 SLICE-12 before operator dashboard visibility.
- **G-065** — Document in v0.5 ship notes: tap-to-talk UX (D-071) is the lowest-confidence assumption (0.60, unvalidated). v0.5 pilot validates adoption; v0.6 adds wake-word if low.
- **G-069** — Document in v0.5 ship notes: post-ship safety signal escalation (nightly FN trend spike → human) is a v0.6+ governance gap. v0.5 ships the measurement (REQ-IDEATE-04); v0.6 adds the LLM-as-judge + the escalation response.
- **G-073** — Document in v0.5 ship notes: v0.5 validates the coaching/guardrail/context-binding value, not the hands-free UX (tap-to-talk is the pilot validation; wake-word is v0.6).
- **G-078** — Document in v0.5 ship notes: the guardrail tuning corpus (REQ-IDEATE-01) is synthetic (LLM-generated), not a human red-team prompt set. Accepted limitation for pilot.
---
### ACCEPT Items (proceed as-is)
- Live Assist is the correct next priority (G-042).
- CI is the named sponsor under full autonomy (G-043).
- v0.5 is not a zombie (G-044).
- IDEATE scope expansion is risk-reduction, not scope creep (G-046, Probe 3).
- Out-of-scope is explicit; D-071 wake-word deferral is the key scope reduction (G-048).
- 3 integration points are additive (G-050).
- Tech-debt is budgeted in P2 SLICE-12 (G-052).
- Key-person dependency is manageable under parallelization (G-053).
- 2-phase split is evidence-based; P1 independently shippable (G-057, Probe 6).
- 33 tasks is evidence-based (G-059).
- Assist cost is budgeted, well under C-3 (G-061, G-062).
- No budget contingency — D-071 removed Picovoice dependency (G-064).
- No deploy changes needed (G-074).
- Rollback is preventive (disable assist → revert to v0.1.9) (G-075).
- Tap-to-talk is the 80/20; guardrail work is the non-negotiable 20% (G-079).
- v0.4 grill MUSTs (6/6) resolved in v0.4; v0.4 P1+ findings (8/8) covered in v0.5 P2.
---
### Bottom Line
The v0.5 plan is **not unfeasible** — the D-071 tap-to-talk deferral stripped the client-architecture risk, the battery risk, the Picovoice commercial risk, and 5 of 14 research risks. The remaining scope (guardrail + context-binding + shift-bounded session + in-loop processor) is the *core* safety surface, well-researched and cleanly phased. The plan is **not over-scoped** after the deferral (16 REQs, but 9 are defensive; 33 tasks vs v0.4's 52). The plan is **not a zombie** (Live Assist is the v0.1-promised surface, now delivered).
The 2 MUST conditions are surgical:
- 1 is a *validation spike* (in-loop guardrail processor retry mechanism — G-049).
- 1 is a *threshold* (R-ASSIST-07 adversarial FN rate acceptance — G-067).
The 1 escalation is a *legal question* the CI cannot resolve (PIPEDA consent-law review — ESCALATION-01). This is the de facto stop trigger for the assist surface.
**Resolve the 2 MUSTs, answer the 1 escalation, and v0.5 is a GO.**
The v0.5 milestone is the project's first **safety-critical** surface — the AI is in a learner's ear during *real* customer interactions. The grill's binding decisions (G-067 threshold, G-049 validation) + the escalation (ESCALATION-01 PIPEDA) are the safety-critical gates. The plan's architecture (3-layer guardrail, defense-in-depth, audit + nightly trending) is sound — the grill's conditions ensure the *residual risk* is *known + decided*, not *assumed + deferred*.