# CIAgent Grill Report — v0.5 Live Assist (On-the-Job Voice Companion) ## Run: 2026-08-04 (mode: mechanical, focus: all axes + 6 v0.5-specific probes) > **Reviewer:** adversarial technology executive (red-team) > **Subject:** v0.5 execution plan (Live Assist — On-the-Job Voice Companion) — 2 execution phases, 12 slices, 33 tasks, 16 active REQs (3 ASSIST + 4 NFR + 9 IDEATE) > **Stance:** plan is unfeasible, over-scoped, and too costly until evidence forces otherwise > **Artifacts reviewed:** PROJECT.md (D-058..D-073), REQUIREMENTS.md (16 active REQs + 4 v0.6 backlog), ROADMAP.md, ARCHITECTURE.md (v0.5 Live Assist Mode §), RESEARCH-v0.5-live-assist.md (14 risks R-ASSIST-01..14, 7 domains), PLAN-v0.5-live-assist.md (2 phases, 12 slices, 33 tasks), PERSONAS.md (5 active, 2 deactivated), GRILL-v0.4.md (format reference + G-001..G-041), REVIEW.md (8 v0.4 P1+ carried forward), AUDIT.md (v0.4 HEALTHY), config.json (autonomy=full), server/pipeline.py, server/services/base.py, server/guardrails/customer_service.py, server/session_recorder.py, server/__main__.py > **Binding status:** This grill verdict must be cleared (MUSTs resolved, escalations answered) before EXECUTE is authorized. --- ### Verdict: Proceed-with-conditions (confidence: 0.70) The v0.5 plan is the project's first **safety-critical** milestone — the AI is in a learner's ear during *real* customer interactions, not role-play. This is a categorical shift from v0.1–v0.4 (practice surface, no real customers, no real consequences). The plan's single most important decision — **D-071 (tap-to-talk only, wake-word deferred to v0.6)** — is the correct call: it strips the client-architecture risk (React-Web can't do foreground services), the battery risk, the Picovoice MAU-pricing risk, and 5 of 14 research risks (R-ASSIST-01/04/05/13/14 all become N/A). What remains is the *core* safety surface: the guardrail (REQ-ASSIST-03), the context-binding (REQ-ASSIST-02), and the shift-bounded session model (REQ-NFR-ASSIST-04). This is the right 80/20. However, four material issues must be resolved before EXECUTE: (1) **R-ASSIST-07 (guardrail false-negative)** is the single project-killing risk — a direct answer slips past the regex, the learner parrots it to a real customer, trust erodes. The plan *accepts* this residual risk ("adversarial FN rate is reported but not threshold-gated" — PLAN:419) without a documented acceptance threshold or an escalation. For a safety-critical surface, "we'll measure it and trend it nightly" is necessary but not sufficient — the grill must set the bar. (2) **D-073 (PIPEDA consent-law review)** is deferred to "Phase 1 implementation" — but shipping a recording device into real customer interactions without legal sign-off is a regulatory risk the CI agent cannot resolve under full autonomy. This is an escalation, not a binding decision. (3) The IDEATE stage **expanded v0.5 scope from 7 REQs to 16** (+128%) — the first use of ideation in the project. The 9 added REQs are *defensive* (guardrail tuning, mode-conflict, PII policy, audit-log, reconnect, tech-debt, cost, NFR measurement), not feature creep — but the grill must verify the expansion is risk-reduction, not scope inflation. (4) The **in-loop guardrail processor** (post-LLM, pre-TTS) is a *structural pipeline change*, not the "minimal delta / prompt swap" the research frames it as — the v0.1 pipeline has no in-loop guardrail (the CS guardrail runs on the debrief, not in-loop per RESEARCH §5.2). This is the highest-novelty code in v0.5 and it is on the safety-critical path. The plan is **not** over-scoped *after* the D-071 deferral (16 REQs, but 9 are defensive; 33 tasks vs v0.4's 52). It is **not** unfeasible (0 new pip/npm deps, v0.1 pipeline reused). It is **not** a zombie (Live Assist is the explicitly-deferred v0.1 surface, now delivered). The conditions are binding and surgical — but two of them (R-ASSIST-07 threshold, PIPEDA escalation) touch the safety-critical core and cannot be waived. --- ### Axis 1 — Business Case - **Q1: What problem does Live Assist solve that the practice surface (v0.1-v0.4) doesn't? Is "on-the-job coaching" the top priority, or a feature looking for a user?** - Evidence: PROJECT.md:45-47 — "v0.1–v0.4 built and validated the practice surface… v0.5 adds the companion surface: a hands-free voice assistant a learner invokes *while actually working*"; RESEARCH-v0.5 §4.1 — "No direct competitor does live-in-ear coaching during real customer calls on a $100 phone" (verified: Dialpad/Gong post-hoc, RealWear AR+industrial); ROADMAP.md:9-11 — "the key distinction from the practice surface is real-customer interaction." - Answer: Live Assist solves a problem the practice surface structurally cannot: coaching *during* real work, not *after* a role-play. The practice surface (v0.1-v0.4) teaches via simulated scenarios; Live Assist coaches during live customer interactions. This is the *transfer* moment — where practice meets the job. RESEARCH §4.1 confirms Praxis is novel (no competitor does this on a cheap phone). The priority is correct: v0.1-v0.4 built the practice foundation + operator visibility; v0.5 builds the transfer surface. The alternative (v0.6 low-bandwidth) would expand reach before the on-the-job value is proven. - Confidence: 0.80 - Decision: **G-042** — Live Assist is the correct next priority (delivers the transfer surface the practice foundation was built for). Novel per RESEARCH §4.1. (0.80) - **Q2: Who is the named executive sponsor for Live Assist specifically? (D-001 says "User-directed" for Canada — is there a sponsor for Live Assist?)** - Evidence: config.json:13 — `"level": "full"`; PROJECT.md:5 — "Autonomy: full"; D-001 (PROJECT.md:171) — "Launch market = Canada… User-directed"; no named human sponsor for Live Assist in any `.ciagent/` file. - Answer: No human sponsor. The CI agent is the executive sponsor under full autonomy — the established model since v0.1 (G-002 in GRILL-v0.4). The "sponsor makes a decision under pressure" test is met by this grill — the R-ASSIST-07 + PIPEDA decisions are the pressure decisions. D-001's "User-directed" applied to the *market* choice (Canada), not to Live Assist's scope. - Confidence: 0.80 - Decision: **G-043** — CI is the named sponsor under full autonomy (no change from v0.1-v0.4 governance, G-002 carry-forward). (0.80) - **Q3: What happens to the business if v0.5 is cancelled? (Does the v0.1-v0.4 practice surface work without it?)** - Evidence: ROADMAP.md:149-157 — future milestones (v0.6 low-bandwidth, v0.7 multi-language) do not depend on Live Assist; PROJECT.md:64-69 — v0.4 operator tier + v0.3 mastery + v0.1 voice loop carry forward unchanged. - Answer: If v0.5 is cancelled, the practice surface (v0.1-v0.4) continues to function. Live Assist is a *new surface*, not a dependency of the existing product. However, cancelling v0.5 means the *transfer* value (coaching during real work) is never delivered — the practice surface teaches, but the on-the-job bridge is missing. This is not a zombie (cancelling has a cost: the product's value proposition — "turn every smartphone into a master craftsperson that talks to you" — is unfulfilled without the live-coaching surface). But the practice surface is independently valuable. - Confidence: 0.78 - Decision: **G-044** — v0.5 is not a zombie (delivers the transfer surface). The practice surface works without it, but the product's core promise (on-the-job coaching) is unfulfilled. Accept the non-zombie status. (0.78) - **Q4: Is there an ROI calculation vs a counterfactual (skip to v0.6 low-bandwidth)?** - Evidence: MISSING — no ROI calculation in any `.ciagent/` file. D-012 (PROJECT.md:182) — "v0.1 cost ceiling = no enforced ceiling (pilot)"; REQ-IDEATE-07 (REQUIREMENTS.md:70) — assist cost tracking added by ideation. - Answer: No financial ROI. The counterfactual is "ship v0.5 vs skip to v0.6 (low-bandwidth)." Shipping v0.5 costs ~33 tasks of tokens + 0 new deps + the safety-critical guardrail work. Skipping to v0.6 would leave Live Assist permanently deferred (broken v0.1 out-of-scope promise: "Live Assist mode") and v0.6's low-bandwidth surfaces would build on a practice-only product with no on-the-job transfer. The ROI is *product-completeness* (delivering the v0.1-promised surface) + *safety-surface validation* (the guardrail work is the foundation for all future safety-critical domains per D-019). REQ-IDEATE-07 adds cost tracking — the *measurement* of ROI, not the calculation. - Confidence: 0.68 - Decision: **G-045** — no financial ROI; the ROI is product-completeness (v0.1-promised surface) + safety-surface foundation (guardrail work extends D-019 for future domains). REQ-IDEATE-07 measures cost, doesn't justify it. Accept the non-financial ROI under full autonomy. (0.68) --- ### Axis 2 — Scope and Requirements - **Q1: Is the scope stable? 16 active REQs + 4 v0.6 backlog — is this expanding?** - Evidence: REQUIREMENTS.md:8-81 — 16 active REQs (3 ASSIST + 4 NFR + 9 IDEATE); PROJECT.md:49 — "3 REQs + NFRs TBD after RESEARCH/IDEATE"; PLAN-v0.5:1011 — "16/16 REQ-IDs covered"; git log `b8c7de8` — "ideation results — 9 accepted into v0.5, 4 accepted into v0.6." - Answer: The scope **expanded** from 7 REQs (3 ASSIST + 4 NFR, post-CLARIFY) to 16 REQs (+9 IDEATE) — a +128% increase. This is the project's first use of the IDEATE stage. The 9 added REQs are: REQ-IDEATE-01 (guardrail tuning corpus), -02 (in-loop processor test), -03 (mode-conflict), -04 (measurable NFRs), -05 (PII policy), -06 (v0.4 tech-debt), -07 (cost tracking), -08 (WebRTC reconnect), -09 (incremental audit-log). **All 9 are defensive/risk-reduction, not features.** They address: guardrail false-positive/negative (the safety risk), mutual exclusivity (a correctness gap), PII (a privacy gap), NFR measurability (a verifiability gap), tech-debt (carried from v0.4), cost (C-3), resilience (WebRTC drop), audit completeness (abrupt termination). This is scope *hardening*, not scope *creep* — but it is still expansion, and the grill must verify each addition is risk-reduction, not gold-plating. - Confidence: 0.78 - Challenge: The +128% expansion is the largest scope growth in the project's history (v0.4 was a clean handoff: 8 REQs, 0 added). The IDEATE stage is a new vector — without discipline, ideation becomes scope creep with a defensive veneer. The 9 REQs are individually justified, but the *aggregate* added 9 tasks of P1 surface + 4 P2 tasks. The grill accepts the expansion *because* each REQ maps to a named risk (R-ASSIST-06/07/08/09/11 + v0.4 P1+ findings), not because ideation is inherently good. - Decision: **G-046** — scope expanded +128% via IDEATE (7→16 REQs). Accepted because all 9 additions are risk-reduction (guardrail, PII, mode-conflict, resilience, audit, tech-debt, cost, NFR measurability), not feature creep. Each maps to a named risk. Future ideation must maintain this risk-reduction discipline. (0.78) - **Q2: Are requirements frozen? (The 4 NFRs were `pending-research` → `research-grounded` — are they stable now?)** - Evidence: REQUIREMENTS.md:22-25 — 4 NFRs marked `research-grounded (R-ASSIST-XX)`; REQUIREMENTS.md:27 — "NFRs refined from `pending-research` to `research-grounded` after the v0.5 RESEARCH stage… Phase-1 measurement may further refine R-ASSIST-02 (latency) and R-ASSIST-14 (battery)." - Answer: The 4 NFRs are *research-grounded*, not *frozen*. REQ-NFR-ASSIST-01 (latency) is explicitly "AT RISK" — estimated ~655ms, target <600ms, pilot tolerance ≤650ms (D-072). REQ-NFR-ASSIST-02 (hands-free) was refined by D-071 (tap-to-talk only, wake-word deferred). REQ-NFR-ASSIST-03 (guardrail) is refined by D-068 (regex + retry + fallback). REQ-NFR-ASSIST-04 (session model) is stable (D-062). The NFRs are *stable enough* for PLAN, but REQ-NFR-ASSIST-01's target is a *pilot tolerance* (≤650ms), not the binding constraint (<600ms) — this is a deferred hardening, not a freeze. REQ-IDEATE-04 adds measurable targets (p95 ≤650ms, FP<5%) — this *is* the freeze for measurement purposes. - Confidence: 0.75 - Decision: **G-047** — NFRs are research-grounded, not frozen. REQ-NFR-ASSIST-01 (latency) is at-risk with a pilot tolerance (D-072); REQ-IDEATE-04 provides the measurable freeze (p95 ≤650ms pilot, FP<5%). Accept as pilot-scale with v0.6 hardening for <600ms. (0.75) - **Q3: What is explicitly out of scope? (Is the v0.5 out-of-scope list as explicit as v0.4's?)** - Evidence: PROJECT.md:54-62 — explicit out-of-scope list (9 items); REQUIREMENTS.md:83-92 — matching list. - Answer: Explicitly out of scope: full multi-path launch, low-bandwidth surfaces (WhatsApp/USSD/offline), multi-language, persona switching, full operator-suite dashboard, learner auth/multi-learner-per-device, session recording/replay, proactive intervention, multi-modal. The list is as explicit as v0.4's. The key deferral is **wake-word (D-071)** — the original D-058 scope (wake-word + tap-to-talk) is reduced to tap-to-talk only, with wake-word deferred to v0.6. This is the largest scope *reduction* in v0.5 and it is explicit (D-071 binding, PLAN:25). - Confidence: 0.85 - Decision: **G-048** — out-of-scope is explicit and comprehensive. D-071 (wake-word deferred) is the key scope reduction, documented as binding. (0.85) - **Q4: Hidden requirements? (PIPEDA legal review D-073 — is this a hidden regulatory requirement?)** - Evidence: D-073 (PROJECT.md:243) — "PIPEDA consent-law review = defer to v0.5 Phase 1 implementation"; R-ASSIST-08 (RESEARCH-v0.5 §2.6) — "Privacy/consent failure: the real customer didn't consent to being recorded/analyzed by an AI"; D-070 (PROJECT.md:240) — consent disclosure implemented regardless. - Answer: **Yes — PIPEDA is a hidden regulatory requirement.** The ambient mic captures the real customer (a third party); ASR transcribes their speech; the turns table stores it (REQ-IDEATE-05 acknowledges this as "STRIDE information-disclosure"). Canada's PIPEDA + provincial one-party/two-party consent laws govern recording. D-073 defers the legal review to "Phase 1 implementation" and frames it as "not a Phase 0 blocker." The disclosure (D-070) is the *engineering* mitigation, but it is NOT a *legal* determination — a disclosure does not make recording legal if the law requires two-party consent. The CI agent under full autonomy cannot resolve a legal question. This is an **escalation**, not a binding decision — the grill cannot determine with confidence ≥0.60 whether the disclosure is sufficient or whether legal review must block ship. - Confidence: 0.55 - Challenge: PIPEDA is a regulatory requirement that the plan defers. For a safety-critical surface with real customers, deferring legal review is a risk the CI cannot own. This must be escalated. - Decision: **ESCALATION-01** — PIPEDA consent-law review (D-073) is a hidden regulatory requirement that cannot be resolved under full autonomy. The disclosure (D-070) is the engineering mitigation but not a legal determination. **Escalate to human attention:** determine whether Canada PIPEDA + provincial consent law requires explicit legal sign-off before shipping a recording device into real customer interactions. If the disclosure is legally sufficient, proceed; if two-party consent is required, the assist surface may need customer-facing consent (out of scope for v0.5) or geographic restriction. (0.55 — below threshold) --- ### Axis 3 — Architecture and Technical Feasibility - **Q1: Has the assist pipeline architecture been validated? (D-061 says shares v0.1 pipeline — is build_assist_pipeline() validated or assumed?)** - Evidence: server/pipeline.py:44-185 — `build_pipeline()` with `_build_transport` (line 63), `_build_stt` (line 76), `_build_llm` (line 89), `_build_tts` (line 109), `LatencyObserver` (line 183); RESEARCH-v0.5 §5.2 — "v0.5 adds a `build_assist_pipeline()`… Reuses `_build_transport`, `_build_stt`, `_build_llm`, `_build_tts` unchanged"; PLAN-v0.5 TASK-05-01 — `build_assist_pipeline()` assembles the pipeline. - Answer: The v0.1 service constructors (`_build_transport/stt/llm/tts`) are verified present and reusable (pipeline.py:63-109). `build_assist_pipeline()` is *assumed* to reuse them — this is sound for the service layer. **However**, the in-loop guardrail processor (TASK-05-02 — `LiveAssistGuardrailProcessor` as a post-LLM, pre-TTS `FrameProcessor`) is a *structural pipeline change*, not a prompt swap. The v0.1 pipeline has NO in-loop guardrail processor — the CS guardrail runs on the debrief (post-session), not in-loop (RESEARCH §5.2: "the existing v0.1 pipeline doesn't have a post-LLM guardrail processor inline"). Inserting a frame processor between `llm` and `tts` is novel for this codebase. The research frames this as "~1 new Pipecat frame processor" (§5.2) — but Pipecat frame-processor semantics (when does `LLMFullResponseEndFrame` fire? can you inject a retry mid-stream?) are unvalidated. PLAN Open Question #4 (line 1046) defers the retry mechanism to EXECUTE: "verify Pipecat's `LLMContextAggregator` supports injecting a message + re-running the LLM within a single `process_frame` call. If not, the retry may need to be a separate pipeline task." This is the highest-novelty code in v0.5 and it is on the safety-critical path. - Confidence: 0.70 - Challenge: The in-loop guardrail processor is a structural change deferred to EXECUTE. The retry mechanism (inject `RETRY_INSTRUCTION` + re-run LLM) is unvalidated against Pipecat's frame semantics. If Pipecat can't do mid-stream retry, the guardrail's "one retry" (D-068) becomes "canned fallback only" — a weaker safety posture. - Decision: **G-049 (MUST)** — The in-loop guardrail processor's retry mechanism (TASK-05-02) must be validated against Pipecat's frame-processor semantics BEFORE Wave 3 (SLICE-05). Add a Wave-1 or Wave-2 spike task: "Verify `LLMFullResponseEndFrame` fires after the full LLM response + that `LLMContextAggregator` supports injecting a retry message + re-running the LLM within `process_frame`." If Pipecat cannot do mid-stream retry, document the fallback (canned fallback only, no retry) and update D-068's safety posture. This is a binding contract, not an open question. (0.70) - **Q2: Integration surface — v0.4 cohort aggregation (D-062), v0.1 voice pipeline (D-061), v0.3 mastery (D-063). Each is an integration point. Risk of quiet cost doubling?** - Evidence: PLAN-v0.5 SLICE-10 (aggregation extension), SLICE-05 (pipeline reuse), SLICE-01 (D-063 schedule_mastery=False); RESEARCH-v0.5 §6.1 — "no schema change to cohort_aggregates (the `metric` column is free-form TEXT)"; §4.3 — "D-063 is unambiguous: assist turns never update θ… `run_mastery_flow()` is invoked only for practice sessions." - Answer: Three integration points, all *additive*: 1. **v0.4 cohort aggregation** — new `session_type='assist'` + 5 new metric strings (no schema change, D-062). Risk: low — the aggregator is metric-agnostic (RESEARCH §6.1, 0.90 confidence). But the aggregation cache persistence (v0.4 P1+ #7, REQ-IDEATE-06) directly corrupts `assist_active_learners_count` after restart — the tech-debt wave (SLICE-12) fixes this. **Dependency: the tech-debt fix is on the v0.5 critical path for correct assist metrics.** 2. **v0.1 voice pipeline** — `build_assist_pipeline()` reuses services but adds the in-loop guardrail processor (see Q1). Risk: medium — the structural change is the novelty. 3. **v0.3 mastery separation** — `schedule_mastery=False` for assist (D-063). Risk: low — the `end()` signature already supports the flag (RESEARCH §4.3, 0.90 confidence). Verified in code: `session_recorder.py` `end()` has `schedule_mastery` param. - The cost-doubling risk is concentrated in the in-loop guardrail processor (Q1). The aggregation + mastery integrations are low-risk additive extensions. - Confidence: 0.75 - Decision: **G-050** — 3 integration points, all additive. Cohort aggregation (low risk, metric-agnostic) + mastery separation (low risk, flag exists) + voice pipeline (medium risk, in-loop guardrail is structural). The aggregation cache tech-debt (P1+ #7) is on the critical path for correct assist metrics — SLICE-12 fixes it. Accept with G-049 (guardrail retry validation). (0.75) - **Q3: Is there an existing system being replaced? (No — Live Assist is new. But does it inherit v0.1-v0.4 tech debt?)** - Evidence: REVIEW.md:182-203 — 8 v0.4 P1+ findings; REQ-IDEATE-06 (REQUIREMENTS.md:64) — "Carry-forward the 8 v0.4 P1+ findings into the v0.5 backlog as a 'tech-debt wave'"; PLAN-v0.5 SLICE-12 — tech-debt wave (4 tasks). - Answer: No existing system replaced — Live Assist is new. It inherits 8 v0.4 P1+ findings, budgeted in P2 SLICE-12 (REQ-IDEATE-06): (1) argon2id blocking, (2) rate-limit mock test, (3) cookie-secret length, (4) credential status enum, (5) revocation audit log, (6) nightly zoneinfo, (7) aggregation cache persistence, (8) f-string SQL. The most consequential for v0.5 is #7 (aggregation cache) — it directly corrupts `assist_active_learners_count` after restart. The tech-debt wave is in P2 (not P1) — this means the assist metrics are *incorrect* for all of P1 + early P2 until SLICE-12 ships. This is a *deferred fix on the critical path*. - Confidence: 0.72 - Challenge: The aggregation cache fix (P1+ #7) is in P2 SLICE-12, but it corrupts v0.5's assist metrics during P1. The plan accepts this (P1 doesn't ship to operators — it's the assist voice loop). But if P1 ships as v0.1.11 (per-phase ship, config.json:110), the assist metrics are wrong in any P1 deployment. This is a *sequencing* issue, not a missing task. - Decision: **G-051** — 8 v0.4 P1+ findings inherited, budgeted in P2 SLICE-12. The aggregation cache fix (P1+ #7) corrupts assist metrics during P1 — accept this because P1 ships the assist *voice loop* (no operator dashboard dependency), and the fix lands in P2 before operator visibility matters. Document in P1 ship notes: assist metrics are incorrect until P2 SLICE-12. (0.72) - **Q4: Technical debt being inherited — is it budgeted for?** - Evidence: PLAN-v0.5 SLICE-12 (4 tasks: cache persistence, cookie-secret, credential status, argon2id+rate-limit+audit+zoneinfo); REQ-IDEATE-06 (should priority, P1). - Answer: Yes — budgeted in P2 SLICE-12 (4 tasks covering all 8 findings). The tech-debt wave is `should` priority (not `must`) — this is correct (the findings are non-blocking per REVIEW.md). The budget is 4 tasks in P2 Wave 2 — proportional to the 8 findings (some are one-liners: cookie-secret warning, zoneinfo swap). - Confidence: 0.80 - Decision: **G-052** — tech-debt budgeted (4 tasks in P2 SLICE-12, `should` priority). Proportional to the 8 findings. Accept. (0.80) --- ### Axis 4 — People, Skills, and Organization - **Q1: Key-person dependency — voice-engineer is REACTIVATED for the first time. Is there a knowledge concentration risk?** - Evidence: PERSONAS.md:577-593 — voice-engineer REACTIVATED, owns 7 P1 tasks (largest territory: build_assist_pipeline, in-loop guardrail processor, warm WebRTC, reconnect, tap-to-talk client, latency tuning); PLAN-v0.5:102-108 — persona load distribution. - Answer: The voice-engineer owns the largest P1 territory (7 tasks) and is activated for the *first time* in the project (proposed since v0.2 PERSONAS line 458, never operated). The in-loop guardrail processor + warm WebRTC + reconnect logic are all *new capabilities* this project has never built. If the voice-engineer is absent, the assist voice loop (SLICE-05, SLICE-06) has no owner — these are the core of v0.5. The security-engineer (6 tasks) owns the guardrail regex + tuning corpus — the other safety-critical path. The backend-engineer (6 tasks) owns the session API + context-binding. **Three personas are critical-path: voice-engineer, security-engineer, backend-engineer.** The voice-engineer is the highest key-person risk because the capability is *new* (no prior project experience), not just the territory. - Confidence: 0.78 - Decision: **G-053** — key-person dependency: voice-engineer (new capability, largest territory), security-engineer (safety-critical guardrail), backend-engineer (session API + integration). All 3 critical-path. The voice-engineer is the highest risk (first activation, new capability). Accept under parallelization (max 5 concurrent, 5 active personas — exactly at the limit). (0.78) - **Q2: Are the 5 active personas actually allocated? (CI agents, not humans. Are the agent capabilities sufficient for the voice-engineer territory?)** - Evidence: config.json:22-27 — parallelization enabled, max 5 concurrent; PERSONAS.md:556-646 — 5 active personas; config.json:52-81 — only 4 personas in config.json array (voice-engineer + security-engineer are emergent, defined in PERSONAS.md). - Answer: 5 active personas, max 5 concurrent — **exactly at the limit, no slack.** If all 5 are active in a wave, there is zero idle capacity for rework. P1 Wave 1 has 2 parallel slices (SLICE-01, SLICE-02) — 2 personas active (backend, backend+voice). P1 Wave 3 has 2 slices (SLICE-05, SLICE-06) — 2 personas (voice, voice). Peak parallelism is 2-3 slices per wave — within the 5-agent limit. The voice-engineer + security-engineer are NOT in config.json `personas` (emergent) — territory enforcement is `warn` (config.json:51), so they are not blocked. The capability question: the voice-engineer's frameworks (porcupine-android, webrtc, pipecat, piper-tts) are listed in PERSONAS.md but the voice-engineer has *never operated* in this project. The capability is *claimed*, not *demonstrated*. The in-loop guardrail processor (Q1, Axis 3) is the test of this capability. - Confidence: 0.72 - Decision: **G-054** — 5 active personas, max 5 concurrent (at the limit, no slack). Peak parallelism 2-3 slices — within limit. Voice-engineer capability is claimed but undemonstrated (first activation). Accept with G-049 (guardrail retry validation) as the capability test. (0.72) - **Q3: Is there a product owner with authority? (autonomy=full — the CI is the owner. Is that sound for a safety-critical surface?)** - Evidence: config.json:13 — `"level": "full"`; PROJECT.md:5; config.json:34-38 — security auto_accept_low_severity, auto_mitigate_medium, escalate_high_severity. - Answer: CI is the product owner under full autonomy — the established model since v0.1 (G-002, G-015 carry-forward). **For a safety-critical surface, this is the grill's hardest governance question.** The CI can auto-accept low-severity security issues + auto-mitigate medium — but R-ASSIST-07 (guardrail false-negative) is high-severity, and config.json:37 says `escalate_high_severity: true`. The plan *accepts* the residual risk (adversarial FN not threshold-gated) without escalating. This is a tension: the config says escalate high-severity, but the plan says accept. The grill must resolve this — either the residual risk is *not* high-severity (because defense-in-depth + audit + v0.6 LLM-as-judge mitigate it to medium), or the plan must escalate. See Probe 1. - Confidence: 0.68 - Challenge: The CI-as-owner model is sound for practice surfaces (v0.1-v0.4) where the worst case is a bad role-play. For Live Assist, the worst case is a guardrail bypass during a real customer call. The config's `escalate_high_severity: true` is the safety valve — the plan must use it or justify why the risk is not high-severity. - Decision: **G-055** — CI is the product owner (full autonomy, carry-forward). For the safety-critical surface, the `escalate_high_severity: true` config (config.json:37) is the governing constraint. R-ASSIST-07 (guardrail false-negative) is high-severity per RESEARCH — the plan must either (a) escalate it (Probe 1) or (b) document why defense-in-depth + audit + v0.6 LLM-as-judge reduce it to medium (auto-mitigatable). This is resolved in Probe 1. (0.68) - **Q4: Is the team building capability it doesn't have? (voice-engineer is new — has the guardrail/latency/pipeline work been done before in this project?)** - Evidence: RESEARCH-v0.5 §5.2 — "v0.5 adds an in-loop guardrail processor… the existing v0.1 pipeline doesn't have a post-LLM guardrail processor inline"; §3.3 — "prefill latency for gemma4:cloud is not yet measured (R3 from v0.1)"; PERSONAS.md:577-593 — voice-engineer frameworks include porcupine-android (not used in v0.5 per D-071), webrtc, pipecat. - Answer: Yes — three new capabilities: 1. **In-loop Pipecat frame processor** — never built in this project. The v0.1 guardrail runs on the debrief (post-session), not in-loop. The frame-processor semantics (LLMFullResponseEndFrame, mid-stream retry) are unvalidated (G-049). 2. **Warm WebRTC connection lifecycle** — v0.1 opens per-session cold connections; v0.5 keeps a warm connection for an 8h shift with heartbeat + reconnect. New state machine (REQ-IDEATE-08). 3. **Regex guardrail tuning** — the CS guardrail (customer_service.py, 128 lines) is a fixed ruleset; v0.5 adds a tuning corpus + adversarial test + FP/FN measurement (REQ-IDEATE-01/04). New testing methodology. - All three are on the safety-critical or critical path. This is *acceptable for a pilot* (learning-as-you-go is the project's model since v0.1) but the grill must flag that the highest-novelty code (in-loop processor) is also the highest-safety-impact code. - Confidence: 0.72 - Decision: **G-056** — team is building 3 new capabilities (in-loop frame processor, warm WebRTC lifecycle, regex guardrail tuning). All on the safety-critical/critical path. Acceptable for pilot with G-049 (guardrail retry validation) as the de-risking spike. The voice-engineer's first activation is the capability test. (0.72) --- ### Axis 5 — Timeline and Estimates - **Q1: Was the deadline set before or after the scope was understood? (No deadline — CI pipeline. Is the 2-phase split evidence-based or arbitrary?)** - Evidence: ROADMAP.md:13-31 — v0.5 phases defined in ROADMAP (P0 pre-execution, P1 assist core, P2 integration, P3 review); PLAN-v0.5:17-25 — phase split rationale. - Answer: No calendar deadline (CI pipeline). The 2-phase split is *evidence-based*: P1 = the assist voice loop + guardrail (the safety-critical, on-voice-path surface — 12 REQs, 24 tasks); P2 = integration + measurement + tech-debt (the operator-facing + hardening surface — 4 REQs, 9 tasks). The split mirrors v0.4 (P1 infra / P2 feature) but inverts it (P1 feature / P2 hardening). P1 is independently shippable (a learner can start a shift, tap-to-talk, get coaching with guardrails, end the shift). This is the correct split — the safety-critical surface ships first, the measurement + tech-debt follows. - Confidence: 0.82 - Decision: **G-057** — 2-phase split is evidence-based (P1 safety-critical voice loop, P2 hardening + measurement). P1 independently shippable. Not arbitrary. (0.82) - **Q2: Critical path — what single thing would push v0.5 by a phase? (Likely the guardrail — REQ-ASSIST-03 is safety-critical. Is the guardrail on the critical path?)** - Evidence: PLAN-v0.5 wave dependency graph (P1:79-98); SLICE-03 (guardrail) → SLICE-04 (tuning corpus) → SLICE-05 (pipeline + in-loop processor) → SLICE-08 (e2e guardrail test); REQ-IDEATE-01 (tuning corpus + adversarial test). - Answer: The guardrail is on the critical path (SLICE-03 → 04 → 05 → 08). The single thing that would push v0.5 by a wave: - **Most likely: the guardrail tuning corpus fails FP<5% or direct-FN<5% (REQ-IDEATE-01).** TASK-04-02 asserts FP<5% on coaching responses + FN<5% on direct answers. If the regex over-matches (FP>5%) or under-matches (FN>5%), the regex needs retuning → pushes Wave 2 → Wave 3 → Wave 4. This is a *test-driven* gate — the tuning corpus is the proof. - **Less likely: the in-loop guardrail processor retry mechanism is infeasible in Pipecat (G-049).** If Pipecat can't do mid-stream retry, the guardrail weakens to "canned fallback only" — still safe, but D-068's "one retry" is unmet. This would push Wave 3 (SLICE-05) by a spike. - **Least likely: the warm WebRTC reconnect state machine (REQ-IDEATE-08).** The reconnect logic is specified (TASK-06-02) but the chaos test (TASK-06-03) is the proof. If the state machine has edge cases, it pushes Wave 3 (SLICE-06). - Confidence: 0.75 - Decision: **G-058** — critical-path risk: guardrail tuning corpus (FP/FN rates, REQ-IDEATE-01). Mitigation: TASK-04-02 (test-driven gate). If FP>5% or direct-FN>5%, retune the regex → pushes by a wave. Accept with the test as the gate. G-049 (retry validation) de-risks the secondary path. (0.75) - **Q3: Are the estimates evidence-based? (33 tasks across 2 phases — is this analogous to v0.4's 52 tasks/2 phases?)** - Evidence: PLAN-v0.5:1064 — 33 tasks (24 P1 + 9 P2); GRILL-v0.4:166 — v0.4 had 52 tasks (29 P1 + 23 P2); GRILL-v0.4:19 — v0.3 shipped ~40 tasks. - Answer: 33 tasks vs v0.4's 52 (-37%) and v0.3's 40 (-18%). The reduction is explained by D-071 (tap-to-talk only — wake-word deferral removed ~8-10 tasks: Porcupine integration, foreground service, battery management, OEM kill-switch handling) + 0 new deps (no dep-integration tasks). The scope is *smaller* than v0.4 despite +8 REQs (16 vs 8) because the IDEATE additions are mostly test/measurement tasks (low LOC) + the wake-word deferral stripped the client-architecture work. The tasks are bottom-up sized (each slice has 3-7 tasks with acceptance criteria). Evidence-based. - Confidence: 0.80 - Decision: **G-059** — 33 tasks is evidence-based (smaller than v0.4's 52 due to D-071 wake-word deferral + 0 new deps; IDEATE additions are test/measurement tasks). Bottom-up sized. Accept. (0.80) - **Q4: Definition of done — is "done" the grill's verdict or the verify stage's?** - Evidence: PLAN-v0.5 — per-slice acceptance criteria; ROADMAP.md:19-21 — per-phase ship + verify; config.json:28-33 — verification automated. - Answer: Definition of done = per-slice acceptance criteria + per-phase ship (v0.1.11, v0.1.12, v0.1.13) + verify stage. The grill is the P0 definition of done (this document). Established pattern since v0.2 (G-020 carry-forward). For the safety-critical surface, the *additional* done criterion is REQ-IDEATE-04's measurable NFRs (p95 ≤650ms, FP<5%) — these are the *quantitative* done bar for the guardrail. - Confidence: 0.82 - Decision: **G-060** — definition of done = per-slice acceptance + per-phase ship + verify + REQ-IDEATE-04 measurable NFRs (p95 ≤650ms, FP<5%) as the quantitative guardrail bar. Established pattern + safety-critical addition. Accept. (0.82) --- ### Axis 6 — Budget and Financial Realism - **Q1: Cost drivers — assist mode adds LLM calls (IDEATE-07 — 400 extra calls/month/learner). Is this in the budget?** - Evidence: REQ-IDEATE-07 (REQUIREMENTS.md:70) — "20 turns/shift × 20 shifts/month = 400 extra LLM calls"; PLAN-v0.5 SLICE-11 — per-turn cost tracking + C-3 check; TASK-11-02 — `check_c3_budget()`. - Answer: The cost driver is *budgeted* (SLICE-11, REQ-IDEATE-07). The estimate: 400 extra gemma4:cloud calls/month/learner at ~$0.0005/turn = ~$0.20/month — well under C-3's $3 (RESEARCH-v0.5, TASK-11-02). The cost is *diagnostic* (not enforced — D-012 says no enforced ceiling for pilot). The C-3 check (TASK-11-02) flags if practice + assist exceeds $3. This is the correct posture — measure, don't enforce, for the pilot. - Confidence: 0.80 - Decision: **G-061** — assist cost driver budgeted (SLICE-11, ~$0.20/month, well under C-3). Diagnostic, not enforced (D-012 pilot relaxation). Accept. (0.80) - **Q2: C-3 (≤$3/active learner/month) — does assist break it? (D-012 relaxed C-3 for the pilot, but is the relaxation still valid for v0.5?)** - Evidence: D-012 (PROJECT.md:182) — "v0.1 cost ceiling = no enforced ceiling (pilot)"; GRILL-v0.4 G-012 — "no TLS → accepted as pilot-scale constraint"; REQ-IDEATE-07 — C-3 check. - Answer: The C-3 relaxation (D-012) was set for v0.1 and carried through v0.4 (G-012). v0.5 adds ~$0.20/month/learner for assist — the total (practice + assist) is still well under $3 at pilot scale. The relaxation remains valid *for the pilot*. The architecture must not preclude meeting $3 post-pilot (D-012) — the assist cost is LLM calls, which the post-pilot path (self-hosted gemma4:e4b, D-020) reduces. The relaxation is valid for v0.5. - Confidence: 0.78 - Decision: **G-062** — C-3 relaxation (D-012) remains valid for v0.5 pilot. Assist adds ~$0.20/month, total well under $3. Post-pilot path (self-hosted model) preserves the $3 target. Accept. (0.78) - **Q3: Burn rate — token cost of 33 tasks + 2 phases + grill + review + audit. Is this proportional to v0.4?** - Evidence: git log — v0.4 shipped in ~1.3 days (GRILL-v0.4 G-023); v0.5 has 33 tasks vs v0.4's 52 (-37%). - Answer: v0.5 is ~37% smaller than v0.4 by task count. Expected burn: ~0.8-1.0 days of CI agent time (proportional reduction). The token cost is the CI agent's operational cost — not tracked, but the pace is established (4 milestones in ~4 days). Proportional. - Confidence: 0.78 - Decision: **G-063** — burn rate: ~0.8-1.0 days estimated (proportional to v0.4, -37% tasks). Accept. (0.78) - **Q4: Is the budget contingent on anything? (Porcupine pricing D-064 — MAU-priced, no recurring free tier. Is the pilot contingent on Picovoice sales engagement?)** - Evidence: D-064 (PROJECT.md:234) — Porcupine MAU pricing; D-071 (PROJECT.md:241) — tap-to-talk only in v0.5, wake-word deferred to v0.6; R-ASSIST-01 (RESEARCH-v0.5 §1.2) — "no recurring free tier." - Answer: **No — D-071 removed the Picovoice contingency.** The wake-word (Porcupine) is deferred to v0.6. v0.5 ships tap-to-talk only — no Porcupine dependency, no MAU pricing, no sales engagement needed. This is the single biggest budget de-risking of v0.5: the entire Picovoice commercial question is v0.6's problem, not v0.5's. The v0.5 budget is contingent on *nothing* external (0 new deps, no vendor engagement, full autonomy). - Confidence: 0.85 - Decision: **G-064** — no budget contingency. D-071 (tap-to-talk only) removed the Picovoice MAU-pricing dependency. v0.5 has 0 external commercial dependencies. Accept. (0.85) --- ### Axis 7 — Risks, Assumptions, and Dependencies - **Q1: Top 3 assumptions — evidence for each?** - Evidence: RESEARCH-v0.5 risks (R-ASSIST-01..14); D-071, D-068, D-072. - Answer: 1. **Tap-to-talk is sufficient UX (D-071).** Evidence: none — this is an *unvalidated* assumption. No user testing, no pilot data. The practice surface (v0.1-v0.4) uses a WebRTC connection per session; tap-to-talk is a button-hold pattern. Whether a learner on a real shift will tap a button on their phone (which may be in their pocket) is *untested*. The alternative (wake-word) is deferred to v0.6. **Confidence: 0.60** — the assumption is reasonable (tap-to-talk is a proven pattern for walkie-talkie apps) but unvalidated for this use case. 2. **Regex guardrail is adequate (D-068).** Evidence: RESEARCH §2.3 (0.78 confidence) — the regex patterns target direct-answer + false-authority + impersonation. The tuning corpus (REQ-IDEATE-01) + adversarial test will measure FP/FN. The adversarial FN rate is "reported but not threshold-gated" (PLAN:419) — this is a *residual risk acceptance*, not a proof of adequacy. **Confidence: 0.65** — the regex is the fast on-voice-path filter; the LLM-as-judge (v0.6) is the accurate off-voice-path backstop. Defense-in-depth is the mitigation, not regex alone. 3. **≤650ms latency is achievable (D-072).** Evidence: RESEARCH §3.3 — estimated ~655ms (Piper + lean prompt), unmeasured. The estimate is a *budget math* calculation, not a measurement. R1/R3/R4 (Deepgram/Ollama/Piper latencies) are unmeasured since v0.1. **Confidence: 0.65** — the budget math is sound but the actual latencies are unmeasured. D-072 accepts ≤650ms as pilot tolerance; <600ms is v0.6 hardening. - Confidence: 0.63 - Decision: **G-065** — 3 core assumptions: tap-to-talk UX (0.60, unvalidated), regex guardrail adequacy (0.65, residual risk accepted), ≤650ms latency (0.65, unmeasured). All accepted as pilot-scale constraints with v0.6 hardening paths. The tap-to-talk assumption is the lowest-confidence — flag for v0.6 user testing. (0.63) - **Q2: Dependencies — Picovoice (D-064, deferred to v0.6), PIPEDA (D-073), v0.4 cohort pipeline (D-062), v0.1 voice pipeline (D-061).** - Evidence: D-071 (Picovoice deferred), D-073 (PIPEDA deferred), D-062 (cohort aggregation), D-061 (voice pipeline reuse). - Answer: - **Picovoice**: NOT a v0.5 dependency (D-071 — tap-to-talk only). Deferred to v0.6. ✅ - **PIPEDA**: Deferred to "Phase 1 implementation" (D-073). This is the escalation (ESCALATION-01, Axis 2). The disclosure (D-070) is the engineering mitigation. ⚠️ - **v0.4 cohort pipeline**: D-062 — additive extension (session_type=assist, new metric strings, no schema change). Verified: aggregator.py is metric-agnostic (RESEARCH §6.1, 0.90). ✅ - **v0.1 voice pipeline**: D-061 — service reuse (transport/stt/llm/tts) + in-loop guardrail processor (structural change, G-049). ⚠️ - The PIPEDA dependency is the only one that requires human attention. The others are internal + additive. - Confidence: 0.75 - Decision: **G-066** — 4 dependencies: Picovoice (deferred, ✅), PIPEDA (escalation, ⚠️ — ESCALATION-01), cohort pipeline (additive, ✅), voice pipeline (structural change, ⚠️ — G-049). Accept the internal dependencies; escalate PIPEDA. (0.75) - **Q3: Single risk that kills v0.5? (R-ASSIST-07 — guardrail false-negative reaches learner's ear during real customer call. Is there a mitigation beyond "defense-in-depth + post-v0.5 LLM-as-judge"?)** - Evidence: R-ASSIST-07 (RESEARCH-v0.5 §2.6) — "The 'parrot' failure: the AI gives a verbatim script, the learner repeats it word-for-word, the customer detects the robotic delivery → trust erosion"; PLAN-v0.5:1025 — "defense-in-depth (prompt + regex + audit) + adversarial test + nightly FN trending + post-v0.5 LLM-as-judge (REQ-IDEATE-10, v0.6)"; PLAN:419 — "adversarial FN rate is reported but not threshold-gated." - Answer: R-ASSIST-07 is the single project-killing risk. A direct answer that slips past the regex → learner parrots it → real customer hears robotic delivery → trust erosion + potential escalation. The mitigation is *defense-in-depth* (3 layers: prompt + regex + audit) + *measurement* (tuning corpus + adversarial test + nightly FN trending) + *future backstop* (v0.6 LLM-as-judge). **The gap: the adversarial FN rate is "reported but not threshold-gated" (PLAN:419).** This means the plan *accepts* an unknown residual risk without a ceiling. For a safety-critical surface, this is insufficient — the grill must set the bar. The bar cannot be "0% FN" (regex can't catch every paraphrase) — but it must be a *documented acceptance threshold* with an escalation if exceeded. config.json:37 says `escalate_high_severity: true` — R-ASSIST-07 is high-severity, so the plan must either escalate or document why the residual risk is acceptable. - Confidence: 0.68 - Challenge: The plan accepts an unquantified residual risk on a safety-critical surface. "We'll measure it and trend it nightly" is necessary but not sufficient — what happens if the nightly trend shows 15% FN? The plan has no trigger. This is the grill's hardest call. - Decision: **G-067 (MUST)** — R-ASSIST-07 (guardrail false-negative) must have a *documented acceptance threshold* before EXECUTE. The adversarial FN rate (REQ-IDEATE-01) must be: (a) measured pre-ship (TASK-04-02), (b) compared against a threshold (e.g., "adversarial FN ≤ 20% acceptable for pilot because defense-in-depth + audit + v0.6 LLM-as-judge mitigate; >20% triggers a re-tuning wave or escalation"), (c) the threshold + the mitigation rationale documented in the ship notes. This is NOT a "0% FN" demand — it is a "know your residual risk + decide if it's acceptable" demand. The plan's current "reported but not threshold-gated" is insufficient for a safety-critical surface. config.json:37 `escalate_high_severity: true` is the governing constraint. (0.68) - **Q4: Pre-mortem — "It's 12 months from now and v0.5 failed. Why?"** - Evidence: RESEARCH-v0.5 risks; PLAN-v0.5 risk matrix. - Answer: The most likely failure modes (in order): 1. **A guardrail bypass incident during a real customer call (R-ASSIST-07).** A direct answer slipped past the regex, the learner parroted it, the customer escalated to a real manager who disavowed the "AI's advice." The nightly FN trend showed 18% but no one acted because there was no threshold (G-067 gap). This is the *highest-consequence* failure — it breaks trust in the product + the learner's job. 2. **PIPEDA complaint (R-ASSIST-08 / D-073).** A real customer discovered they were recorded by the learner's mic without their consent. The disclosure (D-070) was shown to the *learner*, not the *customer*. Canada's two-party consent law (if applicable in the province) was not reviewed. This is the *highest-legal-consequence* failure. 3. **The in-loop guardrail processor's retry mechanism was infeasible in Pipecat (G-049).** The "one retry" (D-068) became "canned fallback only" — safe but degraded. The assist coaching quality dropped (every block → canned fallback, no second chance). Learners stopped using assist because the coaching felt robotic. 4. **The latency was >650ms in practice (R-ASSIST-02).** The ~655ms estimate was optimistic; actual p95 was ~720ms. Coaching arrived after the customer moment passed. Learners abandoned assist for being "too slow to be useful." - Confidence: 0.75 - Decision: **G-068** — pre-mortem top-4: guardrail bypass (highest consequence, G-067 gap), PIPEDA complaint (ESCALATION-01), in-loop retry infeasible (G-049), latency >650ms (D-072 pilot tolerance). All four are addressed in binding decisions/escalations. (0.75) --- ### Axis 8 — Governance, Decision-Making, and Communication - **Q1: Decision-maker — autonomy=full, the CI decides. Is there a human escalation path for safety-critical decisions? (config.json escalation_hooks: deploy, delete_data, merge_to_main — none for "ship safety-critical guardrail". Is this a gap?)** - Evidence: config.json:14 — `"escalation_hooks": ["deploy", "delete_data", "merge_to_main"]`; config.json:37 — `"escalate_high_severity": true`; PROJECT.md:5 — "Autonomy: full." - Answer: The escalation_hooks list does NOT include "ship safety-critical guardrail" or "legal review." The `escalate_high_severity: true` security config is the *only* safety valve — it says the CI *should* escalate high-severity security issues, but the *mechanism* (how? to whom?) is unspecified. For v0.1-v0.4 (practice surface), this was acceptable — the worst case was a bad role-play. For v0.5 (Live Assist, real customers), the worst case is a guardrail bypass during a real call + a PIPEDA complaint. The escalation path for these is *the grill itself* — this document is the escalation mechanism. The grill's ESCALATION-01 (PIPEDA) + G-067 (guardrail threshold) are the safety-critical escalations/binding decisions. **The gap: there is no *ongoing* human escalation path post-ship.** If the nightly FN trend spikes post-ship, the CI auto-mitigates (config.json:36) but does not escalate to a human (no hook for "safety signal spike"). This is a v0.6+ governance gap, not a v0.5 blocker — v0.5 ships the measurement (REQ-IDEATE-04 nightly trending); v0.6 adds the LLM-as-judge + the escalation on spike. - Confidence: 0.70 - Decision: **G-069** — escalation path: the grill is the safety-critical escalation mechanism (ESCALATION-01 + G-067). config.json `escalate_high_severity: true` is the governing constraint. Post-ship ongoing escalation (safety signal spike → human) is a v0.6+ governance gap — v0.5 ships the measurement, v0.6 adds the response. Accept for pilot with documented gap. (0.70) - **Q2: Governance cadence — the pipeline stages are the governance. Is the grill the right gate for a safety-critical surface?** - Evidence: ROADMAP.md:21 — "Pipeline stages: SPECIFY → CLARIFY → RESEARCH → IDEATE → PLAN → GRILL → SHIP"; ROADMAP.md:30 — "GRILL-v0.5.md (adversarial review — real-customer interaction warrants grill)." - Answer: The grill is the right gate — ROADMAP.md:30 explicitly flags "real-customer interaction warrants grill." The pipeline stages (SPECIFY→…→GRILL→SHIP) are the governance cadence; the grill is the crisis-cadence (this document). For a safety-critical surface, the grill is the *only* human-in-the-loop checkpoint (the CI runs the rest autonomously). This is the correct model — the grill surfaces the safety-critical decisions (G-067, ESCALATION-01) for human attention before SHIP. - Confidence: 0.82 - Decision: **G-070** — grill is the right gate for a safety-critical surface (ROADMAP:30 explicit). The grill is the human-in-the-loop checkpoint. Accept. (0.82) - **Q3: What's omitted from status reports? (The LSP errors in server/__main__.py, test_scenario_library.py — are these reported or hidden?)** - Evidence: Task context mentions "LSP errors in server/__main__.py, test_scenario_library.py"; verification: `python3 -m py_compile server/__main__.py` → exit 0 (clean); `python3 -m py_compile tests/test_scenario_library.py` → exit 0 (clean). - Answer: The "LSP errors" claim in the task context is **unverified** — both files compile cleanly (`py_compile` exit 0). This may refer to type-checking (pyright/mypy) warnings, not syntax errors, or it may be stale. The grill does not flag this as a material omission — the files compile, the v0.4 tests pass (317 pass, 0 fail per REVIEW.md). If there are type-checking warnings, they are non-blocking (the codebase doesn't enforce strict typing in CI). **No omission found.** - Confidence: 0.80 - Decision: **G-071** — no status-report omission found. The "LSP errors" claim is unverified (files compile clean). Type-checking warnings, if any, are non-blocking. Accept. (0.80) - **Q4: Stop-the-project trigger — is there one? (If the grill returns RETHINK, does the pipeline stop?)** - Evidence: config.json:13 — full autonomy; GRILL-v0.4 G-032 — "no human stop trigger (full autonomy). The grill is the stop mechanism." - Answer: No human stop trigger (full autonomy, G-032 carry-forward). The grill is the stop mechanism — if the verdict were "Rethink" or "Escalate" on a material axis, the pipeline would stop. This grill's verdict is "Proceed-with-conditions" — the project proceeds after the MUSTs (G-049, G-067) + the escalation (ESCALATION-01) are resolved. The escalation (PIPEDA) is the *de facto* stop trigger — if the human legal review determines the disclosure is insufficient, v0.5 cannot ship the assist surface as designed. - Confidence: 0.78 - Decision: **G-072** — no human stop trigger (full autonomy). The grill is the stop mechanism. ESCALATION-01 (PIPEDA) is the de facto stop trigger for the assist surface. This grill = proceed with conditions. (0.78) --- ### Axis 9 — Change, Adoption, and Operational Readiness - **Q1: Who uses Live Assist? (The learner — during a real shift. How does their work change? They now have an AI in their ear.)** - Evidence: PROJECT.md:45-47 — "a hands-free voice assistant a learner invokes *while actually working*"; PERSONAS.md — no learner persona (learners are external to the CI agent); D-071 — tap-to-talk invocation. - Answer: The learner uses Live Assist during a real shift. Their work changes: they now have an AI coach in their ear (via earbuds) that they invoke by tapping a button (D-071 — tap-to-talk, not wake-word). "What's in it for them" = real-time coaching during real customer interactions — the transfer moment from practice to job. **This is unvalidated** — no user testing, no pilot data on whether learners will actually tap a button on their phone during a real customer call (the phone may be in their pocket, the tap may be socially awkward). The tap-to-talk UX (D-071) is the lowest-confidence assumption (G-065, 0.60). The alternative (wake-word, hands-free) is deferred to v0.6. For v0.5 pilot, tap-to-talk is the *validation* — does a learner use it? The measurement is the assist usage metrics (REQ-NFR-ASSIST-04, cohort aggregation). - Confidence: 0.65 - Challenge: The adoption risk is *real* — tap-to-talk during a real customer call is socially + ergonomically awkward (phone in pocket, earbuds in, tap a button on the phone screen). The "we'll measure usage" answer is correct but the pilot may show low adoption. This is a v0.5 *validation* risk, not a v0.5 *blocker*. - Decision: **G-073** — Live Assist's first user is the learner during a real shift. Tap-to-talk (D-071) is the unvalidated UX assumption (G-065, 0.60). v0.5 pilot *validates* adoption (assist usage metrics); v0.6 adds wake-word if tap-to-talk adoption is low. Document in ship notes: v0.5 validates the coaching/guardrail/context-binding value, not the hands-free UX (that's v0.6). (0.65) - **Q2: Is the ops team involved? (CI project — ops is the LXC deploy. Does v0.5 need deploy changes? D-071 says no — v0.4 LXC carries forward. Is that sound?)** - Evidence: PERSONAS.md:651-660 — devops-engineer DEACTIVATED for v0.5 ("No deploy changes — v0.4's LXC + Docker-in-LXC + Postgres + backup cron carries forward unchanged"); PLAN-v0.5:1071 — "New pip deps: 0… New npm deps: 0." - Answer: v0.5 needs NO deploy changes — 0 new pip deps, 0 new npm deps, no new Docker services, no CT bump. The assist surface is server-side code (server/assist/) + a React route (client/src/AssistControl.tsx) on the existing v0.4 LXC. devops-engineer deactivation is sound. The ops surface (LXC, Postgres, backup) is unchanged. This is the correct posture — v0.5 is a *feature* milestone, not an *infra* milestone. - Confidence: 0.85 - Decision: **G-074** — v0.5 needs no deploy changes (0 new deps, no CT bump, v0.4 LXC carries forward). devops-engineer deactivation is sound. Accept. (0.85) - **Q3: Rollback plan — if v0.5 ships and a guardrail incident occurs, what's the rollback? (Disable assist mode? Revert to v0.1.9?)** - Evidence: config.json:40 — `"branching_strategy": "phase"`; PLAN-v0.5 — per-phase ship (v0.1.11, v0.1.12, v0.1.13); git revert pattern (GRILL-v0.4 G-035). - Answer: Rollback is per-phase git revert (G-035 carry-forward). But for a *guardrail incident* (R-ASSIST-07), the rollback is *operational*, not just git: - **Preventive rollback**: disable assist mode (revert to v0.1.9 = v0.4). The assist routes (`/api/assist/*`) + the assist WebRTC endpoint are removed. The practice surface (v0.1-v0.4) continues unchanged. This is a clean revert — the assist surface is additive (new routes, new server/assist/ package, new SQLite migration 0004). Reverting removes the routes + the package; the migration is additive (session_type defaults to 'practice', guardrail_verdict_json is nullable) so existing practice sessions are unaffected. - **Corrective rollback**: impossible. Once a guardrail bypass reaches a learner's ear during a real call, the turn has played. The audit log (REQ-IDEATE-09 incremental write) records it for investigation, but the *incident* cannot be rolled back. This is the nature of a live surface — rollback is preventive (disable), not corrective. - The preventive rollback (disable assist) is clean + tested (the assist surface is additive). The corrective impossibility is accepted (the audit log is the post-incident tool, not a rollback). - Confidence: 0.75 - Decision: **G-075** — rollback is preventive (disable assist mode → revert to v0.1.9). The assist surface is additive (clean revert). Corrective rollback is impossible (a live turn cannot be un-played) — the audit log (REQ-IDEATE-09) is the post-incident tool. Accept the preventive-only rollback. (0.75) - **Q4: Has anyone validated the success criteria with the people who will judge v0.5 successful? (NFRs are research-grounded, not measurement-validated.)** - Evidence: REQUIREMENTS.md:22-25 — NFRs `research-grounded`; REQ-IDEATE-04 — measurable targets (p95 ≤650ms, FP<5%); config.json:13 — full autonomy (CI is the judge). - Answer: No human judge (full autonomy, G-036 carry-forward). The CI is the judge. The success criteria = 16/16 REQ coverage + per-slice acceptance + REQ-IDEATE-04 measurable NFRs. The NFRs are *research-grounded* (estimated, not measured) — REQ-IDEATE-04 + SLICE-09 (P2) add the *measurement*. The validation path: P2 SLICE-09 measures p95 latency + FP/FN rates. If p95 >650ms or FP>5%, the P2 verify stage flags it. This is the *measurement-validated* path — but it happens in P2, not pre-ship. **Gap: the success criteria are validated *during* P2, not *before* P1 ship (v0.1.11).** If P1 ships with a guardrail that has FP>5%, the P1 ship is premature. The mitigation: TASK-04-02 (guardrail tuning test) is in P1 Wave 2 — it runs *before* P1 ship. If it fails, P1 doesn't ship. This is the correct gate. - Confidence: 0.72 - Decision: **G-076** — success criteria are research-grounded, measurement-validated in P2 (SLICE-09). The P1 gate is TASK-04-02 (guardrail tuning test, FP<5% / direct-FN<5%) — runs before P1 ship. If it fails, P1 doesn't ship. Accept with TASK-04-02 as the P1 gate + SLICE-09 as the P2 measurement. (0.72) --- ### Meta — Closing Review - **Q1: If you were the auditor, what would you flag?** - Evidence: all axes above. - Answer: Four flags: 1. **R-ASSIST-07 residual risk acceptance without a threshold (G-067).** The plan accepts an unquantified adversarial FN rate on a safety-critical surface. This is the grill's hardest call — the bar must be set. 2. **PIPEDA legal review deferred (ESCALATION-01).** Shipping a recording device into real customer interactions without legal sign-off is a regulatory risk the CI cannot own. 3. **IDEATE scope expansion +128% (G-046).** The first use of ideation expanded v0.5 from 7 to 16 REQs. The additions are defensive, but the expansion is the largest in project history — future ideation must maintain risk-reduction discipline. 4. **In-loop guardrail processor is a structural pipeline change (G-049).** The research frames it as "~1 new frame processor" but the retry mechanism is unvalidated against Pipecat semantics. This is the highest-novelty code on the safety-critical path. - Confidence: 0.78 - Decision: **G-077** — auditor flags: R-ASSIST-07 threshold gap, PIPEDA escalation, IDEATE scope expansion, in-loop processor novelty. All addressed in binding decisions/escalations. (0.78) - **Q2: What is v0.5 NOT doing that it should? (PIPEDA legal review is deferred D-073 — should it block ship?)** - Evidence: D-073 (PROJECT.md:243); ESCALATION-01 (Axis 2). - Answer: 1. **PIPEDA legal review** — deferred, escalated (ESCALATION-01). The grill cannot determine if it blocks ship — that's a legal question. The disclosure (D-070) is the engineering mitigation; the legal review is the *regulatory* mitigation. 2. **Post-ship safety signal escalation** — the nightly FN trend (REQ-IDEATE-04) measures but does not escalate on spike (G-069). v0.6 adds the LLM-as-judge + the escalation response. 3. **Guardrail red-team prompt set** — REQ-IDEATE-01 builds a *synthetic* tuning corpus (LLM-generated coaching vs direct-answer responses). This is NOT a *human red-team* prompt set — a determined adversary (or a clever learner) may find paraphrases the synthetic corpus doesn't cover. The adversarial test (TASK-04-02) is the best available, but it's synthetic, not human. This is an accepted limitation (pilot). - Confidence: 0.75 - Decision: **G-078** — v0.5 is NOT doing: PIPEDA legal review (escalated), post-ship safety escalation (v0.6), human red-team prompt set (synthetic corpus accepted for pilot). All documented. Accept with ESCALATION-01 as the human-action item. (0.75) - **Q3: Simplest possible version — is tap-to-talk-only (D-071) the 80/20, or is even that over-scoped?** - Evidence: D-071 (PROJECT.md:241); PLAN-v0.5:25 — "Why tap-to-talk only in v0.5." - Answer: Tap-to-talk-only (D-071) IS the 80/20. The wake-word deferral stripped the client-architecture risk (React Native), the battery risk, the Picovoice commercial risk, and 5 of 14 research risks. What remains is the *core* value: the guardrail (safety), the context-binding (relevance), the shift-bounded session (aggregation), the in-loop processor (pipeline integration). Cutting further (e.g., deferring the guardrail tuning corpus REQ-IDEATE-01) would ship a guardrail "blind on its two most safety-critical metrics" (FP/FN) — unacceptable for a safety-critical surface. The 16 REQs are the *minimum* for a safe assist pilot. **Tap-to-talk is the 80/20; the guardrail work is the non-negotiable 20%.** - Confidence: 0.80 - Decision: **G-079** — tap-to-talk-only (D-071) is the 80/20. The guardrail work (REQ-IDEATE-01/04, the tuning corpus + measurable NFRs) is the non-negotiable 20%. Cutting further would ship an unvalidated safety-critical surface. Accept the scope. (0.80) - **Q4: What would have to be true for v0.5 to succeed in the next 90 days, and is it true today?** - Evidence: all axes. - Answer: For v0.5 to succeed: 1. **The guardrail must be robust (FP<5%, direct-FN<5%, adversarial FN known + thresholded).** Is it true today? No — the tuning corpus is built in P1 Wave 2 (TASK-04-01/02). Will be true after P1 if the test passes. G-067 sets the threshold. 2. **The in-loop guardrail processor must work in Pipecat (retry mechanism).** Is it true today? No — unvalidated (G-049). Will be true after the Wave-1/2 spike. 3. **PIPEDA must be addressed (legal review or disclosure-sufficient determination).** Is it true today? No — deferred (ESCALATION-01). Will be true only after human legal review. 4. **The latency must be ≤650ms.** Is it true today? No — unmeasured (D-072). Will be true after P2 SLICE-09 measurement. 5. **The tap-to-talk UX must be usable during a real shift.** Is it true today? No — unvalidated (G-065). Will be true only after pilot deployment (v0.5's validation purpose). - 2 of 5 are addressable in P1/P2 (guardrail robustness, in-loop processor). 1 requires human action (PIPEDA). 2 are post-ship validation (latency measurement, UX adoption). This is the expected state for a pilot — the *plan* is ready; the *proof* is in execution. - Confidence: 0.72 - Decision: **G-080** — 5 success conditions: guardrail robustness (P1 gate, G-067), in-loop processor (P1 spike, G-049), PIPEDA (human escalation, ESCALATION-01), latency (P2 measurement), UX adoption (post-ship validation). 2 addressable in P1/P2, 1 requires human, 2 post-ship. Accept — the plan is ready, the proof is in execution. (0.72) --- ### v0.5-Specific Probes (Signature Questions) #### Probe 1 — R-ASSIST-07 (Guardrail false-negative): Is "defense-in-depth + audit + v0.6 LLM-as-judge" enough for a safety-critical surface? **Question:** The AI is in a learner's ear during a *real* customer call. The regex output filter (D-068) is the on-voice-path guardrail. The adversarial FN rate is "reported but not threshold-gated" (PLAN:419). If a direct answer slips past the regex, the learner may parrot it. Is the 3-layer defense (prompt + regex + audit) + nightly trending + v0.6 LLM-as-judge sufficient, or does the grill need to set a binding threshold? **Evidence:** - R-ASSIST-07 (RESEARCH-v0.5 §2.6) — "The 'parrot' failure: the AI gives a verbatim script, the learner repeats it word-for-word, the customer detects the robotic delivery → trust erosion." - D-068 (PROJECT.md:238) — "regex-based direct-answer + false-authority + impersonation patterns, with one retry on block + canned coaching redirect fallback." - PLAN-v0.5:419 — "The adversarial FN rate is reported but not threshold-gated (it's the residual risk, mitigated by defense-in-depth)." - config.json:37 — `"escalate_high_severity": true`. - REQ-IDEATE-10 (v0.6 backlog) — "LLM-as-judge guardrail evaluation (nightly, off-voice-path) — measure the true false-negative rate the regex filter cannot." **Analysis:** The plan's posture is: regex is the fast on-voice-path filter (D-068); the LLM-as-judge is the accurate off-voice-path backstop (v0.6, REQ-IDEATE-10). The *gap* is v0.5: the regex is the only on-voice-path guardrail, and its adversarial FN rate is *unthresholded*. For a safety-critical surface where the worst case is a guardrail bypass during a real customer call, "we'll measure it and trend it nightly" is necessary but not sufficient — the plan needs a *decision*: what FN rate is acceptable for the pilot, and what happens if it's exceeded? The config says `escalate_high_severity: true` — R-ASSIST-07 is high-severity. The plan *accepts* the residual risk without escalating. This is the tension G-055 identified. The resolution: the grill sets the threshold (G-067) — the adversarial FN rate must be measured pre-ship (TASK-04-02), compared against a documented threshold, and the threshold + mitigation rationale documented in the ship notes. This is NOT a "0% FN" demand (impossible for regex) — it is a "know your residual risk + decide if it's acceptable" demand. The defense-in-depth (prompt + regex + audit) is the *correct* architecture — the grill does not dispute the 3-layer pattern (RESEARCH §2.1, 0.85 confidence). The issue is the *threshold*, not the architecture. The v0.6 LLM-as-judge is the *future* backstop, not the *current* mitigation — v0.5 ships with regex + audit only. **Verdict:** Defense-in-depth is the correct architecture; the missing piece is a *documented acceptance threshold* for the adversarial FN rate. G-067 (MUST) sets this. The plan's "reported but not threshold-gated" is insufficient for a safety-critical surface — the grill requires a threshold + an escalation if exceeded. **Confidence: 0.68.** --- #### Probe 2 — D-073 (PIPEDA consent-law review): Should legal review block ship? **Question:** The ambient mic captures the real customer (a third party). ASR transcribes their speech. The turns table stores it (REQ-IDEATE-05). Canada's PIPEDA + provincial consent laws govern recording. D-073 defers the legal review to "Phase 1 implementation." The disclosure (D-070) is shown to the *learner*, not the *customer*. Is the disclosure sufficient, or does the legal review need to block ship? **Evidence:** - D-073 (PROJECT.md:243) — "PIPEDA consent-law review = defer to v0.5 Phase 1 implementation; document as R-ASSIST-08 in the grill." - D-070 (PROJECT.md:240) — consent disclosure: "Praxis Assist is on — those around you may be recorded by your mic." - R-ASSIST-08 (RESEARCH-v0.5 §2.6) — "the real customer didn't consent to being recorded/analyzed by an AI." - REQ-IDEATE-05 (REQUIREMENTS.md:52) — "The ambient mic captures BOTH the learner and the real customer; ASR transcribes both; the turns table stores transcribed text. The customer is a third party." - config.json:13 — full autonomy (CI cannot resolve legal questions). **Analysis:** This is a *legal* question, not a technical one. The CI agent under full autonomy cannot determine whether Canada's PIPEDA + provincial consent law requires: - (a) One-party consent (the learner's consent is sufficient — the disclosure D-070 covers this). - (b) Two-party consent (the *customer* must consent — Praxis cannot notify the customer, so the assist surface may be illegal in two-party provinces). - (c) A PIPEDA-compliant privacy policy + data handling agreement. The disclosure (D-070) is the *engineering* mitigation — it makes the *learner* aware. It does NOT make the *customer* aware, and it does NOT determine the legal consent regime. The PII policy (REQ-IDEATE-05) retains customer speech with redaction + 30-day retention — this is a *data handling* mitigation, not a *consent* determination. The grill's confidence that the disclosure is sufficient: **0.55** — below the 0.60 threshold. The grill cannot resolve this under full autonomy. This is an escalation. **Verdict:** PIPEDA legal review is a hidden regulatory requirement that the CI cannot resolve. The disclosure (D-070) is the engineering mitigation but not a legal determination. **Escalate to human attention** (ESCALATION-01): determine whether the disclosure is legally sufficient or whether two-party consent / a PIPEDA privacy policy is required before ship. If the disclosure is sufficient, proceed; if not, the assist surface may need geographic restriction or customer-facing consent (out of scope for v0.5). **Confidence: 0.55 — below threshold, escalated.** --- #### Probe 3 — IDEATE scope expansion (+128%): Risk-reduction or scope creep? **Question:** v0.5 started with 7 REQs (3 ASSIST + 4 NFR, post-CLARIFY). IDEATE added 9 REQs (+128%) — the largest scope growth in project history. Are the 9 additions risk-reduction (guardrail, PII, mode-conflict, resilience, audit, tech-debt, cost, NFR measurability) or scope creep with a defensive veneer? **Evidence:** - git log `b8c7de8` — "ideation results — 9 accepted into v0.5, 4 accepted into v0.6." - REQUIREMENTS.md:29-70 — 9 IDEATE REQs. - PLAN-v0.5:1011 — "16/16 REQ-IDs covered." **Analysis:** The 9 IDEATE REQs map to named risks: - REQ-IDEATE-01 (guardrail tuning corpus) → R-ASSIST-06/07 (FP/FN). - REQ-IDEATE-02 (in-loop processor test) → REQ-IDEATE-02 interface gap (GuardrailContext.role). - REQ-IDEATE-03 (mode-conflict) → D-061 mutual exclusivity gap. - REQ-IDEATE-04 (measurable NFRs) → REQ-NFR-ASSIST-01/03 verifiability. - REQ-IDEATE-05 (PII policy) → R-ASSIST-08 (STRIDE information-disclosure). - REQ-IDEATE-06 (tech-debt) → 8 v0.4 P1+ findings. - REQ-IDEATE-07 (cost tracking) → C-3 budget. - REQ-IDEATE-08 (WebRTC reconnect) → R-ASSIST-09. - REQ-IDEATE-09 (incremental audit-log) → R-ASSIST-14 abrupt termination. **Every addition maps to a named risk or a carried-forward finding.** None are features. The expansion is risk-reduction, not scope creep. The +128% is large but justified — v0.5 is the first *safety-critical* milestone, and the IDEATE stage surfaced the defensive requirements the practice surface (v0.1-v0.4) didn't need. The 4 deferred to v0.6 (REQ-IDEATE-10..13) are also risk-reduction (LLM-as-judge, assist-weaning, offline mode, voice-only context) — the ideation was disciplined. **Verdict:** The IDEATE expansion is risk-reduction, not scope creep. Every REQ maps to a named risk. Accepted (G-046). Future ideation must maintain this discipline — the grill will flag any IDEATE addition that doesn't map to a named risk. **Confidence: 0.78.** --- #### Probe 4 — In-loop guardrail processor (structural pipeline change): Is the "minimal delta" framing accurate? **Question:** RESEARCH §5.2 frames the assist pipeline as "minimal delta: ~1 new pipeline builder, ~1 new guardrail processor." But the v0.1 pipeline has NO in-loop guardrail (the CS guardrail runs on the debrief). Is the in-loop processor a "minimal delta" or a structural change? **Evidence:** - server/pipeline.py:143-185 — `build_pipeline()` has no in-loop guardrail processor (transport → stt → latency → user_agg → llm → latency → tts → latency → transport → assistant_agg). - RESEARCH-v0.5 §5.2 — "v0.5 adds an in-loop guardrail processor for assist mode. This is a pipeline-structure change but a small one (~1 new Pipecat frame processor)." - server/guardrails/customer_service.py — CS guardrail runs `check()` standalone, not as a frame processor. - PLAN-v0.5 TASK-05-02 — `LiveAssistGuardrailProcessor(FrameProcessor)` between llm and tts. - PLAN-v0.5 Open Question #4 (line 1046) — "verify Pipecat's `LLMContextAggregator` supports injecting a message + re-running the LLM within a single `process_frame` call. If not, the retry may need to be a separate pipeline task." **Analysis:** The "minimal delta" framing is *partially accurate*. The service reuse (transport/stt/llm/tts) is genuinely minimal — the constructors are env-driven and reusable (verified: pipeline.py:63-109). **But the in-loop guardrail processor is a structural change**: the v0.1 pipeline has no post-LLM frame processor; v0.5 inserts one between `llm` and `tts`. This is novel for this codebase. The retry mechanism (inject `RETRY_INSTRUCTION` + re-run LLM mid-stream) is *unvalidated* against Pipecat's frame semantics — Open Question #4 defers this to EXECUTE, which is too late for a safety-critical path. The risk: if Pipecat's `LLMFullResponseEndFrame` doesn't fire as expected, or if the `LLMContextAggregator` can't inject a retry mid-stream, the guardrail's "one retry" (D-068) becomes "canned fallback only" — safe but degraded. The coaching quality drops (every block → canned fallback, no second chance). This is a *quality* risk, not a *safety* risk (the canned fallback is safe) — but it affects the product's value. **Verdict:** The in-loop guardrail processor is a structural change, not a minimal delta. The retry mechanism must be validated before Wave 3 (G-049 MUST). If Pipecat can't do mid-stream retry, document the fallback (canned-only) + update D-068's safety posture. The "minimal delta" framing should be corrected in the plan. **Confidence: 0.70.** --- #### Probe 5 — Tap-to-talk UX (D-071): Is the unvalidated adoption risk acceptable for a pilot? **Question:** D-071 ships tap-to-talk only (no wake-word). The learner taps a button on their phone during a real customer call. The phone may be in their pocket. The tap may be socially awkward. No user testing validates this UX. Is the pilot the validation, or is this a feature looking for a user? **Evidence:** - D-071 (PROJECT.md:241) — "tap-to-talk ONLY (no wake-word in v0.5)… learner taps a button to invoke an assist turn during a real shift." - G-065 (Axis 7) — tap-to-talk UX assumption confidence 0.60 (lowest). - RESEARCH-v0.5 §4.1 — "No direct competitor does live-in-ear coaching during real customer calls on a $100 phone" (novel surface, no comparable UX to benchmark). **Analysis:** Tap-to-talk is a *proven* pattern for walkie-talkie apps (Zello, Voxer) — users tap+hold to speak, release to send. This is a reasonable UX for hands-free-adjacent interaction. **But** those apps are *the* primary interface (the user opens the app to talk); Praxis assist is a *secondary* interface (the learner is in a real customer call, the phone is in their pocket, they tap a button on a screen they can't see). The social + ergonomic gap is real: the learner must (a) have earbuds in, (b) have the phone accessible, (c) tap a button without looking, (d) do this during a live customer interaction. This is a *high-friction* UX. The pilot is the validation — v0.5 measures assist usage (REQ-NFR-ASSIST-04 cohort metrics). If adoption is low, v0.6 adds wake-word (the hands-free target). This is the correct pilot posture: ship the *value* (coaching/guardrail/context-binding), validate the *UX* (tap-to-talk adoption), iterate in v0.6. The risk is that low adoption makes the pilot a *failure* — but the pilot's purpose is to *find out*, not to *prove* adoption. **Verdict:** Tap-to-talk is an unvalidated but reasonable UX for a pilot. The pilot is the validation. v0.6 adds wake-word if adoption is low. Accept with documented risk (G-073). **Confidence: 0.65.** --- #### Probe 6 — 2-phase split: Is P1 (assist core + guardrail) independently shippable without P2 (measurement + tech-debt)? **Question:** P1 ships v0.1.11 (assist core + guardrail, 12 REQs). P2 ships v0.1.12 (integration + tech-debt + NFR measurement, 4 REQs). Is P1 independently shippable — does a learner get a safe assist experience without P2? **Evidence:** - PLAN-v0.5:17-23 — P1 = assist voice loop + guardrail (12 REQs, 24 tasks); P2 = integration + measurement + tech-debt (4 REQs, 9 tasks). - config.json:110 — `"per_phase": true` (per-phase ship). **Analysis:** P1 delivers: the assist voice loop (build_assist_pipeline), the 3-layer guardrail (LiveAssistGuardrail + tuning corpus + adversarial test), the shift-bounded session model, the tap-to-talk client, the warm WebRTC + reconnect, the incremental audit-log, the mode-conflict guard, the PII policy. A learner can start a shift, tap-to-talk, get coaching with guardrails, end the shift. **This is a safe, usable assist experience.** P2 adds: the cohort aggregation assist metrics (operator visibility), the cost tracking (C-3 check), the NFR measurement (p95 latency, FP/FN rates), the tech-debt wave (8 v0.4 P1+ findings). **P2 is hardening + visibility, not safety.** The guardrail's safety is in P1 (SLICE-03/04/08); P2 *measures* the guardrail's FP/FN rates (SLICE-09) but the guardrail itself ships in P1. The one caveat: the aggregation cache tech-debt (P1+ #7) corrupts `assist_active_learners_count` during P1 (G-051). But P1 doesn't ship operator visibility (the cohort dashboard extension is P2 SLICE-10) — so the corrupted metric is not *visible* during P1. The fix lands in P2 before the dashboard extension. This is a *sequencing* dependency, not a P1 safety gap. **Verdict:** P1 is independently shippable — a learner gets a safe assist experience. P2 is hardening + operator visibility + measurement. The split is clean (P1 = safety-critical voice loop, P2 = hardening). The aggregation cache corruption during P1 is not visible (no dashboard in P1) and fixed in P2 before visibility. **Confidence: 0.82.** --- ### v0.4 Grill Deferred Items — Coverage Check The v0.4 grill (GRILL-v0.4.md) deferred no items to v0.5 (v0.4 was the operator tier, complete). The v0.4 grill's 8 P1+ findings are carried forward as REQ-IDEATE-06 (tech-debt wave, P2 SLICE-12). Let me verify: | v0.4 Grill/Finding | v0.5 Coverage | Status | |---------------------|---------------|--------| | G-008 (backup drill) | v0.4 complete (REVIEW.md:240) | ✅ Resolved in v0.4 | | G-011 (two-store fallback) | v0.4 complete (REVIEW.md:241) | ✅ Resolved in v0.4 | | G-027 (first-boot no v0.3 key) | v0.4 complete (REVIEW.md:242) | ✅ Resolved in v0.4 | | G-031 (R-AUTH-01 reframe) | v0.4 complete (REVIEW.md:243) | ✅ Resolved in v0.4 | | G-038 (differencing-attack test) | v0.4 complete (REVIEW.md:244) | ✅ Resolved in v0.4 | | G-041 (SPA fallback subclass) | v0.4 complete (REVIEW.md:245) | ✅ Resolved in v0.4 | | P1+ #1 (argon2id blocking) | REQ-IDEATE-06, TASK-12-04 | ✅ Covered in v0.5 P2 | | P1+ #2 (rate-limit mock test) | REQ-IDEATE-06, TASK-12-04 | ✅ Covered in v0.5 P2 | | P1+ #3 (cookie-secret length) | REQ-IDEATE-06, TASK-12-02 | ✅ Covered in v0.5 P2 | | P1+ #4 (credential status enum) | REQ-IDEATE-06, TASK-12-03 | ✅ Covered in v0.5 P2 | | P1+ #5 (revocation audit log) | REQ-IDEATE-06, TASK-12-04 | ✅ Covered in v0.5 P2 | | P1+ #6 (nightly zoneinfo) | REQ-IDEATE-06, TASK-12-04 | ✅ Covered in v0.5 P2 | | P1+ #7 (aggregation cache) | REQ-IDEATE-06, TASK-12-01 | ✅ Covered in v0.5 P2 (critical path for assist metrics — G-051) | | P1+ #8 (f-string SQL) | REQ-IDEATE-06, TASK-12-03 | ✅ Covered in v0.5 P2 | **Verdict:** 6/6 v0.4 grill MUSTs resolved in v0.4. 8/8 v0.4 P1+ findings covered in v0.5 P2 SLICE-12 (REQ-IDEATE-06). The aggregation cache fix (P1+ #7) is on the v0.5 critical path for correct assist metrics (G-051). --- ### Binding Decisions | ID | Axis | Decision | Confidence | Type | |----|------|----------|-----------|------| | G-042 | 1 | Live Assist is the correct next priority (delivers the transfer surface). Novel per RESEARCH §4.1. | 0.80 | ACCEPT | | G-043 | 1 | CI is the named sponsor under full autonomy (G-002 carry-forward). | 0.80 | ACCEPT | | G-044 | 1 | v0.5 is not a zombie (delivers the transfer surface). Practice surface works without it. | 0.78 | ACCEPT | | G-045 | 1 | No financial ROI; ROI is product-completeness + safety-surface foundation. REQ-IDEATE-07 measures cost. | 0.68 | ACCEPT | | G-046 | 2 | IDEATE scope expanded +128% (7→16 REQs). Accepted — all 9 additions are risk-reduction, map to named risks. Future ideation must maintain discipline. | 0.78 | ACCEPT | | G-047 | 2 | NFRs are research-grounded, not frozen. REQ-NFR-ASSIST-01 at-risk (D-072 pilot tolerance). REQ-IDEATE-04 provides measurable freeze. | 0.75 | ACCEPT | | G-048 | 2 | Out-of-scope is explicit. D-071 (wake-word deferred) is the key scope reduction, binding. | 0.85 | ACCEPT | | **G-049** | **3** | **MUST: In-loop guardrail processor retry mechanism (TASK-05-02) must be validated against Pipecat frame semantics BEFORE Wave 3. Add a Wave-1/2 spike: verify LLMFullResponseEndFrame + LLMContextAggregator retry injection. If infeasible, document canned-fallback-only + update D-068. Binding contract, not open question.** | **0.70** | **MUST** | | G-050 | 3 | 3 integration points, all additive. Cohort aggregation (low) + mastery separation (low) + voice pipeline (medium, G-049). Aggregation cache tech-debt on critical path (G-051). | 0.75 | ACCEPT | | G-051 | 3 | 8 v0.4 P1+ findings inherited, budgeted in P2 SLICE-12. Aggregation cache fix corrupts assist metrics during P1 — accept (P1 ships voice loop, not operator dashboard). Document in P1 ship notes. | 0.72 | ACCEPT | | G-052 | 3 | Tech-debt budgeted (4 tasks in P2 SLICE-12, `should` priority). Proportional. | 0.80 | ACCEPT | | G-053 | 4 | Key-person: voice-engineer (new capability, largest territory), security-engineer (guardrail), backend-engineer (session API). Voice-engineer highest risk (first activation). | 0.78 | ACCEPT | | G-054 | 4 | 5 active personas, max 5 concurrent (at limit, no slack). Peak parallelism 2-3 slices. Voice-engineer capability claimed but undemonstrated — G-049 is the test. | 0.72 | ACCEPT | | G-055 | 4 | CI is product owner (full autonomy). For safety-critical surface, `escalate_high_severity: true` governs. R-ASSIST-07 must be escalated or documented as medium (Probe 1). | 0.68 | ACCEPT | | G-056 | 4 | Team building 3 new capabilities (in-loop processor, warm WebRTC, regex tuning). All on safety-critical/critical path. Acceptable for pilot with G-049 de-risking. | 0.72 | ACCEPT | | G-057 | 5 | 2-phase split evidence-based (P1 safety-critical voice loop, P2 hardening + measurement). P1 independently shippable. | 0.82 | ACCEPT | | G-058 | 5 | Critical-path: guardrail tuning corpus (FP/FN rates). TASK-04-02 is the gate. G-049 de-risks secondary path. | 0.75 | ACCEPT | | G-059 | 5 | 33 tasks evidence-based (smaller than v0.4's 52 due to D-071 + 0 new deps). Bottom-up sized. | 0.80 | ACCEPT | | G-060 | 5 | Definition of done = per-slice acceptance + per-phase ship + verify + REQ-IDEATE-04 measurable NFRs (p95 ≤650ms, FP<5%). | 0.82 | ACCEPT | | G-061 | 6 | Assist cost driver budgeted (SLICE-11, ~$0.20/month, well under C-3). Diagnostic, not enforced. | 0.80 | ACCEPT | | G-062 | 6 | C-3 relaxation (D-012) remains valid for v0.5 pilot. Assist adds ~$0.20/month. Post-pilot path preserves $3. | 0.78 | ACCEPT | | G-063 | 6 | Burn rate: ~0.8-1.0 days estimated (proportional to v0.4, -37% tasks). | 0.78 | ACCEPT | | G-064 | 6 | No budget contingency. D-071 removed Picovoice MAU-pricing dependency. 0 external commercial dependencies. | 0.85 | ACCEPT | | G-065 | 7 | 3 core assumptions: tap-to-talk UX (0.60, unvalidated), regex guardrail (0.65, residual risk), ≤650ms latency (0.65, unmeasured). All pilot-scale with v0.6 hardening. | 0.63 | ACCEPT | | G-066 | 7 | 4 dependencies: Picovoice (deferred ✅), PIPEDA (escalation ⚠️), cohort pipeline (additive ✅), voice pipeline (structural ⚠️ G-049). | 0.75 | ACCEPT | | **G-067** | **7** | **MUST: R-ASSIST-07 (guardrail false-negative) must have a documented acceptance threshold before EXECUTE. Adversarial FN rate (REQ-IDEATE-01) must be: (a) measured pre-ship (TASK-04-02), (b) compared against a threshold (e.g., "≤20% acceptable for pilot because defense-in-depth + audit + v0.6 LLM-as-judge mitigate; >20% triggers re-tuning or escalation"), (c) threshold + rationale documented in ship notes. Not a "0% FN" demand — a "know your residual risk + decide" demand. config.json:37 escalate_high_severity governs.** | **0.68** | **MUST** | | G-068 | 7 | Pre-mortem top-4: guardrail bypass (G-067 gap), PIPEDA (ESCALATION-01), in-loop retry (G-049), latency >650ms (D-072). All addressed. | 0.75 | ACCEPT | | G-069 | 8 | Escalation path: grill is the safety-critical mechanism (ESCALATION-01 + G-067). Post-ship ongoing escalation (safety spike → human) is v0.6+ gap. Accept for pilot. | 0.70 | ACCEPT | | G-070 | 8 | Grill is the right gate for safety-critical surface (ROADMAP:30 explicit). Human-in-the-loop checkpoint. | 0.82 | ACCEPT | | G-071 | 8 | No status-report omission. "LSP errors" claim unverified (files compile clean). Type-checking warnings non-blocking. | 0.80 | ACCEPT | | G-072 | 8 | No human stop trigger (full autonomy). Grill is the stop mechanism. ESCALATION-01 (PIPEDA) is the de facto stop trigger for the assist surface. | 0.78 | ACCEPT | | G-073 | 9 | Live Assist's first user is the learner during a real shift. Tap-to-talk (D-071) is unvalidated UX (0.60). v0.5 validates adoption; v0.6 adds wake-word if low. | 0.65 | ACCEPT | | G-074 | 9 | v0.5 needs no deploy changes (0 new deps, no CT bump, v0.4 LXC carries forward). devops-engineer deactivation sound. | 0.85 | ACCEPT | | G-075 | 9 | Rollback is preventive (disable assist → revert to v0.1.9). Assist surface is additive (clean revert). Corrective rollback impossible (live turn cannot be un-played) — audit log is post-incident tool. | 0.75 | ACCEPT | | G-076 | 9 | Success criteria research-grounded, measurement-validated in P2 (SLICE-09). P1 gate = TASK-04-02 (guardrail tuning test, FP<5%/FN<5%). P2 = SLICE-09 measurement. | 0.72 | ACCEPT | | G-077 | Meta | Auditor flags: R-ASSIST-07 threshold gap, PIPEDA escalation, IDEATE scope expansion, in-loop processor novelty. All addressed. | 0.78 | ACCEPT | | G-078 | Meta | v0.5 NOT doing: PIPEDA legal review (escalated), post-ship safety escalation (v0.6), human red-team prompt set (synthetic corpus accepted for pilot). | 0.75 | ACCEPT | | G-079 | Meta | Tap-to-talk-only (D-071) is the 80/20. Guardrail work (REQ-IDEATE-01/04) is the non-negotiable 20%. Cutting further ships an unvalidated safety-critical surface. | 0.80 | ACCEPT | | G-080 | Meta | 5 success conditions: guardrail robustness (P1 gate), in-loop processor (P1 spike), PIPEDA (human escalation), latency (P2 measurement), UX adoption (post-ship). Plan ready, proof in execution. | 0.72 | ACCEPT | --- ### Escalations **ESCALATION-01 — PIPEDA consent-law review (D-073, R-ASSIST-08).** Confidence: 0.55 (below 0.60 threshold). The ambient mic captures the real customer (a third party); ASR transcribes their speech; the turns table stores it (REQ-IDEATE-05). Canada's PIPEDA + provincial one-party/two-party consent laws govern recording. D-073 defers the legal review to "Phase 1 implementation." The disclosure (D-070) is shown to the *learner*, not the *customer* — it is the engineering mitigation, not a legal determination. **The CI agent under full autonomy cannot resolve a legal question.** This must be escalated to human attention: 1. **Determine the consent regime:** Does Canada PIPEDA + the pilot province's consent law require one-party consent (learner's consent sufficient — D-070 covers) or two-party consent (customer must consent — Praxis cannot notify the customer)? 2. **If one-party:** the disclosure (D-070) is sufficient. Proceed with v0.5. 3. **If two-party:** the assist surface may need geographic restriction (one-party provinces only) or customer-facing consent (out of scope for v0.5 — would block the assist surface in two-party provinces). 4. **If a PIPEDA privacy policy / data handling agreement is required:** the PII policy (REQ-IDEATE-05, 30-day retention + redaction) may need to be formalized into a PIPEDA-compliant policy before ship. **Action required:** Human legal review of Canada PIPEDA + provincial consent law for ambient recording during coaching, before v0.5 SHIP. The grill cannot determine with confidence ≥0.60 whether the disclosure is sufficient. This is the de facto stop trigger for the assist surface (G-072). --- ### MUST Conditions Summary (blocking — must be resolved before Phase 1 EXECUTE) 1. **G-049 — In-loop guardrail processor retry validation.** Add a Wave-1/2 spike task: verify Pipecat's `LLMFullResponseEndFrame` fires after the full LLM response + that `LLMContextAggregator` supports injecting a retry message + re-running the LLM within `process_frame`. If infeasible, document the fallback (canned-fallback-only, no retry) + update D-068's safety posture. This is a binding contract, not an open question (PLAN Open Question #4 must be resolved pre-EXECUTE). 2. **G-067 — R-ASSIST-07 guardrail false-negative acceptance threshold.** The adversarial FN rate (REQ-IDEATE-01) must be: (a) measured pre-ship (TASK-04-02), (b) compared against a *documented threshold* (e.g., "≤20% acceptable for pilot because defense-in-depth + audit + v0.6 LLM-as-judge mitigate; >20% triggers a re-tuning wave or escalation"), (c) the threshold + mitigation rationale documented in the v0.5 ship notes. The plan's current "reported but not threshold-gated" (PLAN:419) is insufficient for a safety-critical surface. config.json:37 `escalate_high_severity: true` is the governing constraint. --- ### Escalations Requiring Human Attention (before SHIP) **ESCALATION-01 — PIPEDA consent-law review.** Determine whether Canada PIPEDA + provincial consent law requires one-party or two-party consent for ambient recording during coaching. If the disclosure (D-070) is legally sufficient, proceed. If two-party consent is required, the assist surface may need geographic restriction or customer-facing consent (out of scope for v0.5). This is the de facto stop trigger for the assist surface. --- ### FIX Conditions (non-blocking — tracked in VERIFY-P1/P2) - **G-046** — Document in v0.5 ship notes: IDEATE expanded scope +128% (7→16 REQs). All additions are risk-reduction. Future ideation must maintain risk-reduction discipline. - **G-051** — Document in P1 ship notes: assist metrics (assist_active_learners_count) are incorrect during P1 due to the aggregation cache tech-debt (v0.4 P1+ #7). Fix lands in P2 SLICE-12 before operator dashboard visibility. - **G-065** — Document in v0.5 ship notes: tap-to-talk UX (D-071) is the lowest-confidence assumption (0.60, unvalidated). v0.5 pilot validates adoption; v0.6 adds wake-word if low. - **G-069** — Document in v0.5 ship notes: post-ship safety signal escalation (nightly FN trend spike → human) is a v0.6+ governance gap. v0.5 ships the measurement (REQ-IDEATE-04); v0.6 adds the LLM-as-judge + the escalation response. - **G-073** — Document in v0.5 ship notes: v0.5 validates the coaching/guardrail/context-binding value, not the hands-free UX (tap-to-talk is the pilot validation; wake-word is v0.6). - **G-078** — Document in v0.5 ship notes: the guardrail tuning corpus (REQ-IDEATE-01) is synthetic (LLM-generated), not a human red-team prompt set. Accepted limitation for pilot. --- ### ACCEPT Items (proceed as-is) - Live Assist is the correct next priority (G-042). - CI is the named sponsor under full autonomy (G-043). - v0.5 is not a zombie (G-044). - IDEATE scope expansion is risk-reduction, not scope creep (G-046, Probe 3). - Out-of-scope is explicit; D-071 wake-word deferral is the key scope reduction (G-048). - 3 integration points are additive (G-050). - Tech-debt is budgeted in P2 SLICE-12 (G-052). - Key-person dependency is manageable under parallelization (G-053). - 2-phase split is evidence-based; P1 independently shippable (G-057, Probe 6). - 33 tasks is evidence-based (G-059). - Assist cost is budgeted, well under C-3 (G-061, G-062). - No budget contingency — D-071 removed Picovoice dependency (G-064). - No deploy changes needed (G-074). - Rollback is preventive (disable assist → revert to v0.1.9) (G-075). - Tap-to-talk is the 80/20; guardrail work is the non-negotiable 20% (G-079). - v0.4 grill MUSTs (6/6) resolved in v0.4; v0.4 P1+ findings (8/8) covered in v0.5 P2. --- ### Bottom Line The v0.5 plan is **not unfeasible** — the D-071 tap-to-talk deferral stripped the client-architecture risk, the battery risk, the Picovoice commercial risk, and 5 of 14 research risks. The remaining scope (guardrail + context-binding + shift-bounded session + in-loop processor) is the *core* safety surface, well-researched and cleanly phased. The plan is **not over-scoped** after the deferral (16 REQs, but 9 are defensive; 33 tasks vs v0.4's 52). The plan is **not a zombie** (Live Assist is the v0.1-promised surface, now delivered). The 2 MUST conditions are surgical: - 1 is a *validation spike* (in-loop guardrail processor retry mechanism — G-049). - 1 is a *threshold* (R-ASSIST-07 adversarial FN rate acceptance — G-067). The 1 escalation is a *legal question* the CI cannot resolve (PIPEDA consent-law review — ESCALATION-01). This is the de facto stop trigger for the assist surface. **Resolve the 2 MUSTs, answer the 1 escalation, and v0.5 is a GO.** The v0.5 milestone is the project's first **safety-critical** surface — the AI is in a learner's ear during *real* customer interactions. The grill's binding decisions (G-067 threshold, G-049 validation) + the escalation (ESCALATION-01 PIPEDA) are the safety-critical gates. The plan's architecture (3-layer guardrail, defense-in-depth, audit + nightly trending) is sound — the grill's conditions ensure the *residual risk* is *known + decided*, not *assumed + deferred*.