This repository has been archived on 2026-09-12. You can view files and clone it. You cannot open issues or pull requests or push a commit.
Files
Praxis CI ec397f2c65 docs(milestone): complete v0.5-live-assist — v0.1.13 tagged, milestone release, merged to main
v0.5 (Live Assist — on-the-job voice companion) milestone complete.
4 phases: P0 (pre-execution, v0.1.10) → P1 (assist core + guardrail,
v0.1.11) → P2 (integration + tech-debt + NFR, v0.1.12) → P3 (final
review + ship, v0.1.13 = milestone release).

16/16 REQs covered (3 ASSIST + 4 NFR + 9 IDEATE). 4 v0.6 backlog.
469 tests passed, 0 failed. 1 P0 fixed (guardrail processor safety).
8 P1+ flagged for v0.6. 8 v0.4 P1+ tech-debt addressed.
G-049 + G-067 grill MUSTs resolved. ESCALATION-01 (PIPEDA) OPEN for
human legal review before assist surface go-live.

---ci---
project: praxis
phase: 3
milestone: v0.5
status: complete
requirements:
  covered: [REQ-ASSIST-01, REQ-ASSIST-02, REQ-ASSIST-03, REQ-NFR-ASSIST-01, REQ-NFR-ASSIST-02, REQ-NFR-ASSIST-03, REQ-NFR-ASSIST-04, REQ-IDEATE-01, REQ-IDEATE-02, REQ-IDEATE-03, REQ-IDEATE-04, REQ-IDEATE-05, REQ-IDEATE-06, REQ-IDEATE-07, REQ-IDEATE-08, REQ-IDEATE-09]
  partial: []
---/ci---
2026-08-04 22:35:56 +00:00

92 KiB
Raw Permalink Blame History

CIAgent Grill Report — v0.5 Live Assist (On-the-Job Voice Companion)

Run: 2026-08-04 (mode: mechanical, focus: all axes + 6 v0.5-specific probes)

Reviewer: adversarial technology executive (red-team) Subject: v0.5 execution plan (Live Assist — On-the-Job Voice Companion) — 2 execution phases, 12 slices, 33 tasks, 16 active REQs (3 ASSIST + 4 NFR + 9 IDEATE) Stance: plan is unfeasible, over-scoped, and too costly until evidence forces otherwise Artifacts reviewed: PROJECT.md (D-058..D-073), REQUIREMENTS.md (16 active REQs + 4 v0.6 backlog), ROADMAP.md, ARCHITECTURE.md (v0.5 Live Assist Mode §), RESEARCH-v0.5-live-assist.md (14 risks R-ASSIST-01..14, 7 domains), PLAN-v0.5-live-assist.md (2 phases, 12 slices, 33 tasks), PERSONAS.md (5 active, 2 deactivated), GRILL-v0.4.md (format reference + G-001..G-041), REVIEW.md (8 v0.4 P1+ carried forward), AUDIT.md (v0.4 HEALTHY), config.json (autonomy=full), server/pipeline.py, server/services/base.py, server/guardrails/customer_service.py, server/session_recorder.py, server/__main__.py Binding status: This grill verdict must be cleared (MUSTs resolved, escalations answered) before EXECUTE is authorized.


Verdict: Proceed-with-conditions (confidence: 0.70)

The v0.5 plan is the project's first safety-critical milestone — the AI is in a learner's ear during real customer interactions, not role-play. This is a categorical shift from v0.1v0.4 (practice surface, no real customers, no real consequences). The plan's single most important decision — D-071 (tap-to-talk only, wake-word deferred to v0.6) — is the correct call: it strips the client-architecture risk (React-Web can't do foreground services), the battery risk, the Picovoice MAU-pricing risk, and 5 of 14 research risks (R-ASSIST-01/04/05/13/14 all become N/A). What remains is the core safety surface: the guardrail (REQ-ASSIST-03), the context-binding (REQ-ASSIST-02), and the shift-bounded session model (REQ-NFR-ASSIST-04). This is the right 80/20.

However, four material issues must be resolved before EXECUTE: (1) R-ASSIST-07 (guardrail false-negative) is the single project-killing risk — a direct answer slips past the regex, the learner parrots it to a real customer, trust erodes. The plan accepts this residual risk ("adversarial FN rate is reported but not threshold-gated" — PLAN:419) without a documented acceptance threshold or an escalation. For a safety-critical surface, "we'll measure it and trend it nightly" is necessary but not sufficient — the grill must set the bar. (2) D-073 (PIPEDA consent-law review) is deferred to "Phase 1 implementation" — but shipping a recording device into real customer interactions without legal sign-off is a regulatory risk the CI agent cannot resolve under full autonomy. This is an escalation, not a binding decision. (3) The IDEATE stage expanded v0.5 scope from 7 REQs to 16 (+128%) — the first use of ideation in the project. The 9 added REQs are defensive (guardrail tuning, mode-conflict, PII policy, audit-log, reconnect, tech-debt, cost, NFR measurement), not feature creep — but the grill must verify the expansion is risk-reduction, not scope inflation. (4) The in-loop guardrail processor (post-LLM, pre-TTS) is a structural pipeline change, not the "minimal delta / prompt swap" the research frames it as — the v0.1 pipeline has no in-loop guardrail (the CS guardrail runs on the debrief, not in-loop per RESEARCH §5.2). This is the highest-novelty code in v0.5 and it is on the safety-critical path.

The plan is not over-scoped after the D-071 deferral (16 REQs, but 9 are defensive; 33 tasks vs v0.4's 52). It is not unfeasible (0 new pip/npm deps, v0.1 pipeline reused). It is not a zombie (Live Assist is the explicitly-deferred v0.1 surface, now delivered). The conditions are binding and surgical — but two of them (R-ASSIST-07 threshold, PIPEDA escalation) touch the safety-critical core and cannot be waived.


Axis 1 — Business Case

  • Q1: What problem does Live Assist solve that the practice surface (v0.1-v0.4) doesn't? Is "on-the-job coaching" the top priority, or a feature looking for a user?

    • Evidence: PROJECT.md:45-47 — "v0.1v0.4 built and validated the practice surface… v0.5 adds the companion surface: a hands-free voice assistant a learner invokes while actually working"; RESEARCH-v0.5 §4.1 — "No direct competitor does live-in-ear coaching during real customer calls on a $100 phone" (verified: Dialpad/Gong post-hoc, RealWear AR+industrial); ROADMAP.md:9-11 — "the key distinction from the practice surface is real-customer interaction."
    • Answer: Live Assist solves a problem the practice surface structurally cannot: coaching during real work, not after a role-play. The practice surface (v0.1-v0.4) teaches via simulated scenarios; Live Assist coaches during live customer interactions. This is the transfer moment — where practice meets the job. RESEARCH §4.1 confirms Praxis is novel (no competitor does this on a cheap phone). The priority is correct: v0.1-v0.4 built the practice foundation + operator visibility; v0.5 builds the transfer surface. The alternative (v0.6 low-bandwidth) would expand reach before the on-the-job value is proven.
    • Confidence: 0.80
    • Decision: G-042 — Live Assist is the correct next priority (delivers the transfer surface the practice foundation was built for). Novel per RESEARCH §4.1. (0.80)
  • Q2: Who is the named executive sponsor for Live Assist specifically? (D-001 says "User-directed" for Canada — is there a sponsor for Live Assist?)

    • Evidence: config.json:13 — "level": "full"; PROJECT.md:5 — "Autonomy: full"; D-001 (PROJECT.md:171) — "Launch market = Canada… User-directed"; no named human sponsor for Live Assist in any .ciagent/ file.
    • Answer: No human sponsor. The CI agent is the executive sponsor under full autonomy — the established model since v0.1 (G-002 in GRILL-v0.4). The "sponsor makes a decision under pressure" test is met by this grill — the R-ASSIST-07 + PIPEDA decisions are the pressure decisions. D-001's "User-directed" applied to the market choice (Canada), not to Live Assist's scope.
    • Confidence: 0.80
    • Decision: G-043 — CI is the named sponsor under full autonomy (no change from v0.1-v0.4 governance, G-002 carry-forward). (0.80)
  • Q3: What happens to the business if v0.5 is cancelled? (Does the v0.1-v0.4 practice surface work without it?)

    • Evidence: ROADMAP.md:149-157 — future milestones (v0.6 low-bandwidth, v0.7 multi-language) do not depend on Live Assist; PROJECT.md:64-69 — v0.4 operator tier + v0.3 mastery + v0.1 voice loop carry forward unchanged.
    • Answer: If v0.5 is cancelled, the practice surface (v0.1-v0.4) continues to function. Live Assist is a new surface, not a dependency of the existing product. However, cancelling v0.5 means the transfer value (coaching during real work) is never delivered — the practice surface teaches, but the on-the-job bridge is missing. This is not a zombie (cancelling has a cost: the product's value proposition — "turn every smartphone into a master craftsperson that talks to you" — is unfulfilled without the live-coaching surface). But the practice surface is independently valuable.
    • Confidence: 0.78
    • Decision: G-044 — v0.5 is not a zombie (delivers the transfer surface). The practice surface works without it, but the product's core promise (on-the-job coaching) is unfulfilled. Accept the non-zombie status. (0.78)
  • Q4: Is there an ROI calculation vs a counterfactual (skip to v0.6 low-bandwidth)?

    • Evidence: MISSING — no ROI calculation in any .ciagent/ file. D-012 (PROJECT.md:182) — "v0.1 cost ceiling = no enforced ceiling (pilot)"; REQ-IDEATE-07 (REQUIREMENTS.md:70) — assist cost tracking added by ideation.
    • Answer: No financial ROI. The counterfactual is "ship v0.5 vs skip to v0.6 (low-bandwidth)." Shipping v0.5 costs ~33 tasks of tokens + 0 new deps + the safety-critical guardrail work. Skipping to v0.6 would leave Live Assist permanently deferred (broken v0.1 out-of-scope promise: "Live Assist mode") and v0.6's low-bandwidth surfaces would build on a practice-only product with no on-the-job transfer. The ROI is product-completeness (delivering the v0.1-promised surface) + safety-surface validation (the guardrail work is the foundation for all future safety-critical domains per D-019). REQ-IDEATE-07 adds cost tracking — the measurement of ROI, not the calculation.
    • Confidence: 0.68
    • Decision: G-045 — no financial ROI; the ROI is product-completeness (v0.1-promised surface) + safety-surface foundation (guardrail work extends D-019 for future domains). REQ-IDEATE-07 measures cost, doesn't justify it. Accept the non-financial ROI under full autonomy. (0.68)

Axis 2 — Scope and Requirements

  • Q1: Is the scope stable? 16 active REQs + 4 v0.6 backlog — is this expanding?

    • Evidence: REQUIREMENTS.md:8-81 — 16 active REQs (3 ASSIST + 4 NFR + 9 IDEATE); PROJECT.md:49 — "3 REQs + NFRs TBD after RESEARCH/IDEATE"; PLAN-v0.5:1011 — "16/16 REQ-IDs covered"; git log b8c7de8 — "ideation results — 9 accepted into v0.5, 4 accepted into v0.6."
    • Answer: The scope expanded from 7 REQs (3 ASSIST + 4 NFR, post-CLARIFY) to 16 REQs (+9 IDEATE) — a +128% increase. This is the project's first use of the IDEATE stage. The 9 added REQs are: REQ-IDEATE-01 (guardrail tuning corpus), -02 (in-loop processor test), -03 (mode-conflict), -04 (measurable NFRs), -05 (PII policy), -06 (v0.4 tech-debt), -07 (cost tracking), -08 (WebRTC reconnect), -09 (incremental audit-log). All 9 are defensive/risk-reduction, not features. They address: guardrail false-positive/negative (the safety risk), mutual exclusivity (a correctness gap), PII (a privacy gap), NFR measurability (a verifiability gap), tech-debt (carried from v0.4), cost (C-3), resilience (WebRTC drop), audit completeness (abrupt termination). This is scope hardening, not scope creep — but it is still expansion, and the grill must verify each addition is risk-reduction, not gold-plating.
    • Confidence: 0.78
    • Challenge: The +128% expansion is the largest scope growth in the project's history (v0.4 was a clean handoff: 8 REQs, 0 added). The IDEATE stage is a new vector — without discipline, ideation becomes scope creep with a defensive veneer. The 9 REQs are individually justified, but the aggregate added 9 tasks of P1 surface + 4 P2 tasks. The grill accepts the expansion because each REQ maps to a named risk (R-ASSIST-06/07/08/09/11 + v0.4 P1+ findings), not because ideation is inherently good.
    • Decision: G-046 — scope expanded +128% via IDEATE (7→16 REQs). Accepted because all 9 additions are risk-reduction (guardrail, PII, mode-conflict, resilience, audit, tech-debt, cost, NFR measurability), not feature creep. Each maps to a named risk. Future ideation must maintain this risk-reduction discipline. (0.78)
  • Q2: Are requirements frozen? (The 4 NFRs were pending-researchresearch-grounded — are they stable now?)

    • Evidence: REQUIREMENTS.md:22-25 — 4 NFRs marked research-grounded (R-ASSIST-XX); REQUIREMENTS.md:27 — "NFRs refined from pending-research to research-grounded after the v0.5 RESEARCH stage… Phase-1 measurement may further refine R-ASSIST-02 (latency) and R-ASSIST-14 (battery)."
    • Answer: The 4 NFRs are research-grounded, not frozen. REQ-NFR-ASSIST-01 (latency) is explicitly "AT RISK" — estimated ~655ms, target <600ms, pilot tolerance ≤650ms (D-072). REQ-NFR-ASSIST-02 (hands-free) was refined by D-071 (tap-to-talk only, wake-word deferred). REQ-NFR-ASSIST-03 (guardrail) is refined by D-068 (regex + retry + fallback). REQ-NFR-ASSIST-04 (session model) is stable (D-062). The NFRs are stable enough for PLAN, but REQ-NFR-ASSIST-01's target is a pilot tolerance (≤650ms), not the binding constraint (<600ms) — this is a deferred hardening, not a freeze. REQ-IDEATE-04 adds measurable targets (p95 ≤650ms, FP<5%) — this is the freeze for measurement purposes.
    • Confidence: 0.75
    • Decision: G-047 — NFRs are research-grounded, not frozen. REQ-NFR-ASSIST-01 (latency) is at-risk with a pilot tolerance (D-072); REQ-IDEATE-04 provides the measurable freeze (p95 ≤650ms pilot, FP<5%). Accept as pilot-scale with v0.6 hardening for <600ms. (0.75)
  • Q3: What is explicitly out of scope? (Is the v0.5 out-of-scope list as explicit as v0.4's?)

    • Evidence: PROJECT.md:54-62 — explicit out-of-scope list (9 items); REQUIREMENTS.md:83-92 — matching list.
    • Answer: Explicitly out of scope: full multi-path launch, low-bandwidth surfaces (WhatsApp/USSD/offline), multi-language, persona switching, full operator-suite dashboard, learner auth/multi-learner-per-device, session recording/replay, proactive intervention, multi-modal. The list is as explicit as v0.4's. The key deferral is wake-word (D-071) — the original D-058 scope (wake-word + tap-to-talk) is reduced to tap-to-talk only, with wake-word deferred to v0.6. This is the largest scope reduction in v0.5 and it is explicit (D-071 binding, PLAN:25).
    • Confidence: 0.85
    • Decision: G-048 — out-of-scope is explicit and comprehensive. D-071 (wake-word deferred) is the key scope reduction, documented as binding. (0.85)
  • Q4: Hidden requirements? (PIPEDA legal review D-073 — is this a hidden regulatory requirement?)

    • Evidence: D-073 (PROJECT.md:243) — "PIPEDA consent-law review = defer to v0.5 Phase 1 implementation"; R-ASSIST-08 (RESEARCH-v0.5 §2.6) — "Privacy/consent failure: the real customer didn't consent to being recorded/analyzed by an AI"; D-070 (PROJECT.md:240) — consent disclosure implemented regardless.
    • Answer: Yes — PIPEDA is a hidden regulatory requirement. The ambient mic captures the real customer (a third party); ASR transcribes their speech; the turns table stores it (REQ-IDEATE-05 acknowledges this as "STRIDE information-disclosure"). Canada's PIPEDA + provincial one-party/two-party consent laws govern recording. D-073 defers the legal review to "Phase 1 implementation" and frames it as "not a Phase 0 blocker." The disclosure (D-070) is the engineering mitigation, but it is NOT a legal determination — a disclosure does not make recording legal if the law requires two-party consent. The CI agent under full autonomy cannot resolve a legal question. This is an escalation, not a binding decision — the grill cannot determine with confidence ≥0.60 whether the disclosure is sufficient or whether legal review must block ship.
    • Confidence: 0.55
    • Challenge: PIPEDA is a regulatory requirement that the plan defers. For a safety-critical surface with real customers, deferring legal review is a risk the CI cannot own. This must be escalated.
    • Decision: ESCALATION-01 — PIPEDA consent-law review (D-073) is a hidden regulatory requirement that cannot be resolved under full autonomy. The disclosure (D-070) is the engineering mitigation but not a legal determination. Escalate to human attention: determine whether Canada PIPEDA + provincial consent law requires explicit legal sign-off before shipping a recording device into real customer interactions. If the disclosure is legally sufficient, proceed; if two-party consent is required, the assist surface may need customer-facing consent (out of scope for v0.5) or geographic restriction. (0.55 — below threshold)

Axis 3 — Architecture and Technical Feasibility

  • Q1: Has the assist pipeline architecture been validated? (D-061 says shares v0.1 pipeline — is build_assist_pipeline() validated or assumed?)

    • Evidence: server/pipeline.py:44-185 — build_pipeline() with _build_transport (line 63), _build_stt (line 76), _build_llm (line 89), _build_tts (line 109), LatencyObserver (line 183); RESEARCH-v0.5 §5.2 — "v0.5 adds a build_assist_pipeline()… Reuses _build_transport, _build_stt, _build_llm, _build_tts unchanged"; PLAN-v0.5 TASK-05-01 — build_assist_pipeline() assembles the pipeline.
    • Answer: The v0.1 service constructors (_build_transport/stt/llm/tts) are verified present and reusable (pipeline.py:63-109). build_assist_pipeline() is assumed to reuse them — this is sound for the service layer. However, the in-loop guardrail processor (TASK-05-02 — LiveAssistGuardrailProcessor as a post-LLM, pre-TTS FrameProcessor) is a structural pipeline change, not a prompt swap. The v0.1 pipeline has NO in-loop guardrail processor — the CS guardrail runs on the debrief (post-session), not in-loop (RESEARCH §5.2: "the existing v0.1 pipeline doesn't have a post-LLM guardrail processor inline"). Inserting a frame processor between llm and tts is novel for this codebase. The research frames this as "~1 new Pipecat frame processor" (§5.2) — but Pipecat frame-processor semantics (when does LLMFullResponseEndFrame fire? can you inject a retry mid-stream?) are unvalidated. PLAN Open Question #4 (line 1046) defers the retry mechanism to EXECUTE: "verify Pipecat's LLMContextAggregator supports injecting a message + re-running the LLM within a single process_frame call. If not, the retry may need to be a separate pipeline task." This is the highest-novelty code in v0.5 and it is on the safety-critical path.
    • Confidence: 0.70
    • Challenge: The in-loop guardrail processor is a structural change deferred to EXECUTE. The retry mechanism (inject RETRY_INSTRUCTION + re-run LLM) is unvalidated against Pipecat's frame semantics. If Pipecat can't do mid-stream retry, the guardrail's "one retry" (D-068) becomes "canned fallback only" — a weaker safety posture.
    • Decision: G-049 (MUST) — The in-loop guardrail processor's retry mechanism (TASK-05-02) must be validated against Pipecat's frame-processor semantics BEFORE Wave 3 (SLICE-05). Add a Wave-1 or Wave-2 spike task: "Verify LLMFullResponseEndFrame fires after the full LLM response + that LLMContextAggregator supports injecting a retry message + re-running the LLM within process_frame." If Pipecat cannot do mid-stream retry, document the fallback (canned fallback only, no retry) and update D-068's safety posture. This is a binding contract, not an open question. (0.70)
  • Q2: Integration surface — v0.4 cohort aggregation (D-062), v0.1 voice pipeline (D-061), v0.3 mastery (D-063). Each is an integration point. Risk of quiet cost doubling?

    • Evidence: PLAN-v0.5 SLICE-10 (aggregation extension), SLICE-05 (pipeline reuse), SLICE-01 (D-063 schedule_mastery=False); RESEARCH-v0.5 §6.1 — "no schema change to cohort_aggregates (the metric column is free-form TEXT)"; §4.3 — "D-063 is unambiguous: assist turns never update θ… run_mastery_flow() is invoked only for practice sessions."
    • Answer: Three integration points, all additive:
      1. v0.4 cohort aggregation — new session_type='assist' + 5 new metric strings (no schema change, D-062). Risk: low — the aggregator is metric-agnostic (RESEARCH §6.1, 0.90 confidence). But the aggregation cache persistence (v0.4 P1+ #7, REQ-IDEATE-06) directly corrupts assist_active_learners_count after restart — the tech-debt wave (SLICE-12) fixes this. Dependency: the tech-debt fix is on the v0.5 critical path for correct assist metrics.
      2. v0.1 voice pipelinebuild_assist_pipeline() reuses services but adds the in-loop guardrail processor (see Q1). Risk: medium — the structural change is the novelty.
      3. v0.3 mastery separationschedule_mastery=False for assist (D-063). Risk: low — the end() signature already supports the flag (RESEARCH §4.3, 0.90 confidence). Verified in code: session_recorder.py end() has schedule_mastery param.
    • The cost-doubling risk is concentrated in the in-loop guardrail processor (Q1). The aggregation + mastery integrations are low-risk additive extensions.
    • Confidence: 0.75
    • Decision: G-050 — 3 integration points, all additive. Cohort aggregation (low risk, metric-agnostic) + mastery separation (low risk, flag exists) + voice pipeline (medium risk, in-loop guardrail is structural). The aggregation cache tech-debt (P1+ #7) is on the critical path for correct assist metrics — SLICE-12 fixes it. Accept with G-049 (guardrail retry validation). (0.75)
  • Q3: Is there an existing system being replaced? (No — Live Assist is new. But does it inherit v0.1-v0.4 tech debt?)

    • Evidence: REVIEW.md:182-203 — 8 v0.4 P1+ findings; REQ-IDEATE-06 (REQUIREMENTS.md:64) — "Carry-forward the 8 v0.4 P1+ findings into the v0.5 backlog as a 'tech-debt wave'"; PLAN-v0.5 SLICE-12 — tech-debt wave (4 tasks).
    • Answer: No existing system replaced — Live Assist is new. It inherits 8 v0.4 P1+ findings, budgeted in P2 SLICE-12 (REQ-IDEATE-06): (1) argon2id blocking, (2) rate-limit mock test, (3) cookie-secret length, (4) credential status enum, (5) revocation audit log, (6) nightly zoneinfo, (7) aggregation cache persistence, (8) f-string SQL. The most consequential for v0.5 is #7 (aggregation cache) — it directly corrupts assist_active_learners_count after restart. The tech-debt wave is in P2 (not P1) — this means the assist metrics are incorrect for all of P1 + early P2 until SLICE-12 ships. This is a deferred fix on the critical path.
    • Confidence: 0.72
    • Challenge: The aggregation cache fix (P1+ #7) is in P2 SLICE-12, but it corrupts v0.5's assist metrics during P1. The plan accepts this (P1 doesn't ship to operators — it's the assist voice loop). But if P1 ships as v0.1.11 (per-phase ship, config.json:110), the assist metrics are wrong in any P1 deployment. This is a sequencing issue, not a missing task.
    • Decision: G-051 — 8 v0.4 P1+ findings inherited, budgeted in P2 SLICE-12. The aggregation cache fix (P1+ #7) corrupts assist metrics during P1 — accept this because P1 ships the assist voice loop (no operator dashboard dependency), and the fix lands in P2 before operator visibility matters. Document in P1 ship notes: assist metrics are incorrect until P2 SLICE-12. (0.72)
  • Q4: Technical debt being inherited — is it budgeted for?

    • Evidence: PLAN-v0.5 SLICE-12 (4 tasks: cache persistence, cookie-secret, credential status, argon2id+rate-limit+audit+zoneinfo); REQ-IDEATE-06 (should priority, P1).
    • Answer: Yes — budgeted in P2 SLICE-12 (4 tasks covering all 8 findings). The tech-debt wave is should priority (not must) — this is correct (the findings are non-blocking per REVIEW.md). The budget is 4 tasks in P2 Wave 2 — proportional to the 8 findings (some are one-liners: cookie-secret warning, zoneinfo swap).
    • Confidence: 0.80
    • Decision: G-052 — tech-debt budgeted (4 tasks in P2 SLICE-12, should priority). Proportional to the 8 findings. Accept. (0.80)

Axis 4 — People, Skills, and Organization

  • Q1: Key-person dependency — voice-engineer is REACTIVATED for the first time. Is there a knowledge concentration risk?

    • Evidence: PERSONAS.md:577-593 — voice-engineer REACTIVATED, owns 7 P1 tasks (largest territory: build_assist_pipeline, in-loop guardrail processor, warm WebRTC, reconnect, tap-to-talk client, latency tuning); PLAN-v0.5:102-108 — persona load distribution.
    • Answer: The voice-engineer owns the largest P1 territory (7 tasks) and is activated for the first time in the project (proposed since v0.2 PERSONAS line 458, never operated). The in-loop guardrail processor + warm WebRTC + reconnect logic are all new capabilities this project has never built. If the voice-engineer is absent, the assist voice loop (SLICE-05, SLICE-06) has no owner — these are the core of v0.5. The security-engineer (6 tasks) owns the guardrail regex + tuning corpus — the other safety-critical path. The backend-engineer (6 tasks) owns the session API + context-binding. Three personas are critical-path: voice-engineer, security-engineer, backend-engineer. The voice-engineer is the highest key-person risk because the capability is new (no prior project experience), not just the territory.
    • Confidence: 0.78
    • Decision: G-053 — key-person dependency: voice-engineer (new capability, largest territory), security-engineer (safety-critical guardrail), backend-engineer (session API + integration). All 3 critical-path. The voice-engineer is the highest risk (first activation, new capability). Accept under parallelization (max 5 concurrent, 5 active personas — exactly at the limit). (0.78)
  • Q2: Are the 5 active personas actually allocated? (CI agents, not humans. Are the agent capabilities sufficient for the voice-engineer territory?)

    • Evidence: config.json:22-27 — parallelization enabled, max 5 concurrent; PERSONAS.md:556-646 — 5 active personas; config.json:52-81 — only 4 personas in config.json array (voice-engineer + security-engineer are emergent, defined in PERSONAS.md).
    • Answer: 5 active personas, max 5 concurrent — exactly at the limit, no slack. If all 5 are active in a wave, there is zero idle capacity for rework. P1 Wave 1 has 2 parallel slices (SLICE-01, SLICE-02) — 2 personas active (backend, backend+voice). P1 Wave 3 has 2 slices (SLICE-05, SLICE-06) — 2 personas (voice, voice). Peak parallelism is 2-3 slices per wave — within the 5-agent limit. The voice-engineer + security-engineer are NOT in config.json personas (emergent) — territory enforcement is warn (config.json:51), so they are not blocked. The capability question: the voice-engineer's frameworks (porcupine-android, webrtc, pipecat, piper-tts) are listed in PERSONAS.md but the voice-engineer has never operated in this project. The capability is claimed, not demonstrated. The in-loop guardrail processor (Q1, Axis 3) is the test of this capability.
    • Confidence: 0.72
    • Decision: G-054 — 5 active personas, max 5 concurrent (at the limit, no slack). Peak parallelism 2-3 slices — within limit. Voice-engineer capability is claimed but undemonstrated (first activation). Accept with G-049 (guardrail retry validation) as the capability test. (0.72)
  • Q3: Is there a product owner with authority? (autonomy=full — the CI is the owner. Is that sound for a safety-critical surface?)

    • Evidence: config.json:13 — "level": "full"; PROJECT.md:5; config.json:34-38 — security auto_accept_low_severity, auto_mitigate_medium, escalate_high_severity.
    • Answer: CI is the product owner under full autonomy — the established model since v0.1 (G-002, G-015 carry-forward). For a safety-critical surface, this is the grill's hardest governance question. The CI can auto-accept low-severity security issues + auto-mitigate medium — but R-ASSIST-07 (guardrail false-negative) is high-severity, and config.json:37 says escalate_high_severity: true. The plan accepts the residual risk (adversarial FN not threshold-gated) without escalating. This is a tension: the config says escalate high-severity, but the plan says accept. The grill must resolve this — either the residual risk is not high-severity (because defense-in-depth + audit + v0.6 LLM-as-judge mitigate it to medium), or the plan must escalate. See Probe 1.
    • Confidence: 0.68
    • Challenge: The CI-as-owner model is sound for practice surfaces (v0.1-v0.4) where the worst case is a bad role-play. For Live Assist, the worst case is a guardrail bypass during a real customer call. The config's escalate_high_severity: true is the safety valve — the plan must use it or justify why the risk is not high-severity.
    • Decision: G-055 — CI is the product owner (full autonomy, carry-forward). For the safety-critical surface, the escalate_high_severity: true config (config.json:37) is the governing constraint. R-ASSIST-07 (guardrail false-negative) is high-severity per RESEARCH — the plan must either (a) escalate it (Probe 1) or (b) document why defense-in-depth + audit + v0.6 LLM-as-judge reduce it to medium (auto-mitigatable). This is resolved in Probe 1. (0.68)
  • Q4: Is the team building capability it doesn't have? (voice-engineer is new — has the guardrail/latency/pipeline work been done before in this project?)

    • Evidence: RESEARCH-v0.5 §5.2 — "v0.5 adds an in-loop guardrail processor… the existing v0.1 pipeline doesn't have a post-LLM guardrail processor inline"; §3.3 — "prefill latency for gemma4:cloud is not yet measured (R3 from v0.1)"; PERSONAS.md:577-593 — voice-engineer frameworks include porcupine-android (not used in v0.5 per D-071), webrtc, pipecat.
    • Answer: Yes — three new capabilities:
      1. In-loop Pipecat frame processor — never built in this project. The v0.1 guardrail runs on the debrief (post-session), not in-loop. The frame-processor semantics (LLMFullResponseEndFrame, mid-stream retry) are unvalidated (G-049).
      2. Warm WebRTC connection lifecycle — v0.1 opens per-session cold connections; v0.5 keeps a warm connection for an 8h shift with heartbeat + reconnect. New state machine (REQ-IDEATE-08).
      3. Regex guardrail tuning — the CS guardrail (customer_service.py, 128 lines) is a fixed ruleset; v0.5 adds a tuning corpus + adversarial test + FP/FN measurement (REQ-IDEATE-01/04). New testing methodology.
    • All three are on the safety-critical or critical path. This is acceptable for a pilot (learning-as-you-go is the project's model since v0.1) but the grill must flag that the highest-novelty code (in-loop processor) is also the highest-safety-impact code.
    • Confidence: 0.72
    • Decision: G-056 — team is building 3 new capabilities (in-loop frame processor, warm WebRTC lifecycle, regex guardrail tuning). All on the safety-critical/critical path. Acceptable for pilot with G-049 (guardrail retry validation) as the de-risking spike. The voice-engineer's first activation is the capability test. (0.72)

Axis 5 — Timeline and Estimates

  • Q1: Was the deadline set before or after the scope was understood? (No deadline — CI pipeline. Is the 2-phase split evidence-based or arbitrary?)

    • Evidence: ROADMAP.md:13-31 — v0.5 phases defined in ROADMAP (P0 pre-execution, P1 assist core, P2 integration, P3 review); PLAN-v0.5:17-25 — phase split rationale.
    • Answer: No calendar deadline (CI pipeline). The 2-phase split is evidence-based: P1 = the assist voice loop + guardrail (the safety-critical, on-voice-path surface — 12 REQs, 24 tasks); P2 = integration + measurement + tech-debt (the operator-facing + hardening surface — 4 REQs, 9 tasks). The split mirrors v0.4 (P1 infra / P2 feature) but inverts it (P1 feature / P2 hardening). P1 is independently shippable (a learner can start a shift, tap-to-talk, get coaching with guardrails, end the shift). This is the correct split — the safety-critical surface ships first, the measurement + tech-debt follows.
    • Confidence: 0.82
    • Decision: G-057 — 2-phase split is evidence-based (P1 safety-critical voice loop, P2 hardening + measurement). P1 independently shippable. Not arbitrary. (0.82)
  • Q2: Critical path — what single thing would push v0.5 by a phase? (Likely the guardrail — REQ-ASSIST-03 is safety-critical. Is the guardrail on the critical path?)

    • Evidence: PLAN-v0.5 wave dependency graph (P1:79-98); SLICE-03 (guardrail) → SLICE-04 (tuning corpus) → SLICE-05 (pipeline + in-loop processor) → SLICE-08 (e2e guardrail test); REQ-IDEATE-01 (tuning corpus + adversarial test).
    • Answer: The guardrail is on the critical path (SLICE-03 → 04 → 05 → 08). The single thing that would push v0.5 by a wave:
      • Most likely: the guardrail tuning corpus fails FP<5% or direct-FN<5% (REQ-IDEATE-01). TASK-04-02 asserts FP<5% on coaching responses + FN<5% on direct answers. If the regex over-matches (FP>5%) or under-matches (FN>5%), the regex needs retuning → pushes Wave 2 → Wave 3 → Wave 4. This is a test-driven gate — the tuning corpus is the proof.
      • Less likely: the in-loop guardrail processor retry mechanism is infeasible in Pipecat (G-049). If Pipecat can't do mid-stream retry, the guardrail weakens to "canned fallback only" — still safe, but D-068's "one retry" is unmet. This would push Wave 3 (SLICE-05) by a spike.
      • Least likely: the warm WebRTC reconnect state machine (REQ-IDEATE-08). The reconnect logic is specified (TASK-06-02) but the chaos test (TASK-06-03) is the proof. If the state machine has edge cases, it pushes Wave 3 (SLICE-06).
    • Confidence: 0.75
    • Decision: G-058 — critical-path risk: guardrail tuning corpus (FP/FN rates, REQ-IDEATE-01). Mitigation: TASK-04-02 (test-driven gate). If FP>5% or direct-FN>5%, retune the regex → pushes by a wave. Accept with the test as the gate. G-049 (retry validation) de-risks the secondary path. (0.75)
  • Q3: Are the estimates evidence-based? (33 tasks across 2 phases — is this analogous to v0.4's 52 tasks/2 phases?)

    • Evidence: PLAN-v0.5:1064 — 33 tasks (24 P1 + 9 P2); GRILL-v0.4:166 — v0.4 had 52 tasks (29 P1 + 23 P2); GRILL-v0.4:19 — v0.3 shipped ~40 tasks.
    • Answer: 33 tasks vs v0.4's 52 (-37%) and v0.3's 40 (-18%). The reduction is explained by D-071 (tap-to-talk only — wake-word deferral removed ~8-10 tasks: Porcupine integration, foreground service, battery management, OEM kill-switch handling) + 0 new deps (no dep-integration tasks). The scope is smaller than v0.4 despite +8 REQs (16 vs 8) because the IDEATE additions are mostly test/measurement tasks (low LOC) + the wake-word deferral stripped the client-architecture work. The tasks are bottom-up sized (each slice has 3-7 tasks with acceptance criteria). Evidence-based.
    • Confidence: 0.80
    • Decision: G-059 — 33 tasks is evidence-based (smaller than v0.4's 52 due to D-071 wake-word deferral + 0 new deps; IDEATE additions are test/measurement tasks). Bottom-up sized. Accept. (0.80)
  • Q4: Definition of done — is "done" the grill's verdict or the verify stage's?

    • Evidence: PLAN-v0.5 — per-slice acceptance criteria; ROADMAP.md:19-21 — per-phase ship + verify; config.json:28-33 — verification automated.
    • Answer: Definition of done = per-slice acceptance criteria + per-phase ship (v0.1.11, v0.1.12, v0.1.13) + verify stage. The grill is the P0 definition of done (this document). Established pattern since v0.2 (G-020 carry-forward). For the safety-critical surface, the additional done criterion is REQ-IDEATE-04's measurable NFRs (p95 ≤650ms, FP<5%) — these are the quantitative done bar for the guardrail.
    • Confidence: 0.82
    • Decision: G-060 — definition of done = per-slice acceptance + per-phase ship + verify + REQ-IDEATE-04 measurable NFRs (p95 ≤650ms, FP<5%) as the quantitative guardrail bar. Established pattern + safety-critical addition. Accept. (0.82)

Axis 6 — Budget and Financial Realism

  • Q1: Cost drivers — assist mode adds LLM calls (IDEATE-07 — 400 extra calls/month/learner). Is this in the budget?

    • Evidence: REQ-IDEATE-07 (REQUIREMENTS.md:70) — "20 turns/shift × 20 shifts/month = 400 extra LLM calls"; PLAN-v0.5 SLICE-11 — per-turn cost tracking + C-3 check; TASK-11-02 — check_c3_budget().
    • Answer: The cost driver is budgeted (SLICE-11, REQ-IDEATE-07). The estimate: 400 extra gemma4:cloud calls/month/learner at ~$0.0005/turn = ~$0.20/month — well under C-3's $3 (RESEARCH-v0.5, TASK-11-02). The cost is diagnostic (not enforced — D-012 says no enforced ceiling for pilot). The C-3 check (TASK-11-02) flags if practice + assist exceeds $3. This is the correct posture — measure, don't enforce, for the pilot.
    • Confidence: 0.80
    • Decision: G-061 — assist cost driver budgeted (SLICE-11, ~$0.20/month, well under C-3). Diagnostic, not enforced (D-012 pilot relaxation). Accept. (0.80)
  • Q2: C-3 (≤$3/active learner/month) — does assist break it? (D-012 relaxed C-3 for the pilot, but is the relaxation still valid for v0.5?)

    • Evidence: D-012 (PROJECT.md:182) — "v0.1 cost ceiling = no enforced ceiling (pilot)"; GRILL-v0.4 G-012 — "no TLS → accepted as pilot-scale constraint"; REQ-IDEATE-07 — C-3 check.
    • Answer: The C-3 relaxation (D-012) was set for v0.1 and carried through v0.4 (G-012). v0.5 adds ~$0.20/month/learner for assist — the total (practice + assist) is still well under $3 at pilot scale. The relaxation remains valid for the pilot. The architecture must not preclude meeting $3 post-pilot (D-012) — the assist cost is LLM calls, which the post-pilot path (self-hosted gemma4:e4b, D-020) reduces. The relaxation is valid for v0.5.
    • Confidence: 0.78
    • Decision: G-062 — C-3 relaxation (D-012) remains valid for v0.5 pilot. Assist adds ~$0.20/month, total well under $3. Post-pilot path (self-hosted model) preserves the $3 target. Accept. (0.78)
  • Q3: Burn rate — token cost of 33 tasks + 2 phases + grill + review + audit. Is this proportional to v0.4?

    • Evidence: git log — v0.4 shipped in ~1.3 days (GRILL-v0.4 G-023); v0.5 has 33 tasks vs v0.4's 52 (-37%).
    • Answer: v0.5 is ~37% smaller than v0.4 by task count. Expected burn: ~0.8-1.0 days of CI agent time (proportional reduction). The token cost is the CI agent's operational cost — not tracked, but the pace is established (4 milestones in ~4 days). Proportional.
    • Confidence: 0.78
    • Decision: G-063 — burn rate: ~0.8-1.0 days estimated (proportional to v0.4, -37% tasks). Accept. (0.78)
  • Q4: Is the budget contingent on anything? (Porcupine pricing D-064 — MAU-priced, no recurring free tier. Is the pilot contingent on Picovoice sales engagement?)

    • Evidence: D-064 (PROJECT.md:234) — Porcupine MAU pricing; D-071 (PROJECT.md:241) — tap-to-talk only in v0.5, wake-word deferred to v0.6; R-ASSIST-01 (RESEARCH-v0.5 §1.2) — "no recurring free tier."
    • Answer: No — D-071 removed the Picovoice contingency. The wake-word (Porcupine) is deferred to v0.6. v0.5 ships tap-to-talk only — no Porcupine dependency, no MAU pricing, no sales engagement needed. This is the single biggest budget de-risking of v0.5: the entire Picovoice commercial question is v0.6's problem, not v0.5's. The v0.5 budget is contingent on nothing external (0 new deps, no vendor engagement, full autonomy).
    • Confidence: 0.85
    • Decision: G-064 — no budget contingency. D-071 (tap-to-talk only) removed the Picovoice MAU-pricing dependency. v0.5 has 0 external commercial dependencies. Accept. (0.85)

Axis 7 — Risks, Assumptions, and Dependencies

  • Q1: Top 3 assumptions — evidence for each?

    • Evidence: RESEARCH-v0.5 risks (R-ASSIST-01..14); D-071, D-068, D-072.
    • Answer:
      1. Tap-to-talk is sufficient UX (D-071). Evidence: none — this is an unvalidated assumption. No user testing, no pilot data. The practice surface (v0.1-v0.4) uses a WebRTC connection per session; tap-to-talk is a button-hold pattern. Whether a learner on a real shift will tap a button on their phone (which may be in their pocket) is untested. The alternative (wake-word) is deferred to v0.6. Confidence: 0.60 — the assumption is reasonable (tap-to-talk is a proven pattern for walkie-talkie apps) but unvalidated for this use case.
      2. Regex guardrail is adequate (D-068). Evidence: RESEARCH §2.3 (0.78 confidence) — the regex patterns target direct-answer + false-authority + impersonation. The tuning corpus (REQ-IDEATE-01) + adversarial test will measure FP/FN. The adversarial FN rate is "reported but not threshold-gated" (PLAN:419) — this is a residual risk acceptance, not a proof of adequacy. Confidence: 0.65 — the regex is the fast on-voice-path filter; the LLM-as-judge (v0.6) is the accurate off-voice-path backstop. Defense-in-depth is the mitigation, not regex alone.
      3. ≤650ms latency is achievable (D-072). Evidence: RESEARCH §3.3 — estimated ~655ms (Piper + lean prompt), unmeasured. The estimate is a budget math calculation, not a measurement. R1/R3/R4 (Deepgram/Ollama/Piper latencies) are unmeasured since v0.1. Confidence: 0.65 — the budget math is sound but the actual latencies are unmeasured. D-072 accepts ≤650ms as pilot tolerance; <600ms is v0.6 hardening.
    • Confidence: 0.63
    • Decision: G-065 — 3 core assumptions: tap-to-talk UX (0.60, unvalidated), regex guardrail adequacy (0.65, residual risk accepted), ≤650ms latency (0.65, unmeasured). All accepted as pilot-scale constraints with v0.6 hardening paths. The tap-to-talk assumption is the lowest-confidence — flag for v0.6 user testing. (0.63)
  • Q2: Dependencies — Picovoice (D-064, deferred to v0.6), PIPEDA (D-073), v0.4 cohort pipeline (D-062), v0.1 voice pipeline (D-061).

    • Evidence: D-071 (Picovoice deferred), D-073 (PIPEDA deferred), D-062 (cohort aggregation), D-061 (voice pipeline reuse).
    • Answer:
      • Picovoice: NOT a v0.5 dependency (D-071 — tap-to-talk only). Deferred to v0.6.
      • PIPEDA: Deferred to "Phase 1 implementation" (D-073). This is the escalation (ESCALATION-01, Axis 2). The disclosure (D-070) is the engineering mitigation. ⚠️
      • v0.4 cohort pipeline: D-062 — additive extension (session_type=assist, new metric strings, no schema change). Verified: aggregator.py is metric-agnostic (RESEARCH §6.1, 0.90).
      • v0.1 voice pipeline: D-061 — service reuse (transport/stt/llm/tts) + in-loop guardrail processor (structural change, G-049). ⚠️
    • The PIPEDA dependency is the only one that requires human attention. The others are internal + additive.
    • Confidence: 0.75
    • Decision: G-066 — 4 dependencies: Picovoice (deferred, ), PIPEDA (escalation, ⚠️ — ESCALATION-01), cohort pipeline (additive, ), voice pipeline (structural change, ⚠️ — G-049). Accept the internal dependencies; escalate PIPEDA. (0.75)
  • Q3: Single risk that kills v0.5? (R-ASSIST-07 — guardrail false-negative reaches learner's ear during real customer call. Is there a mitigation beyond "defense-in-depth + post-v0.5 LLM-as-judge"?)

    • Evidence: R-ASSIST-07 (RESEARCH-v0.5 §2.6) — "The 'parrot' failure: the AI gives a verbatim script, the learner repeats it word-for-word, the customer detects the robotic delivery → trust erosion"; PLAN-v0.5:1025 — "defense-in-depth (prompt + regex + audit) + adversarial test + nightly FN trending + post-v0.5 LLM-as-judge (REQ-IDEATE-10, v0.6)"; PLAN:419 — "adversarial FN rate is reported but not threshold-gated."
    • Answer: R-ASSIST-07 is the single project-killing risk. A direct answer that slips past the regex → learner parrots it → real customer hears robotic delivery → trust erosion + potential escalation. The mitigation is defense-in-depth (3 layers: prompt + regex + audit) + measurement (tuning corpus + adversarial test + nightly FN trending) + future backstop (v0.6 LLM-as-judge). The gap: the adversarial FN rate is "reported but not threshold-gated" (PLAN:419). This means the plan accepts an unknown residual risk without a ceiling. For a safety-critical surface, this is insufficient — the grill must set the bar. The bar cannot be "0% FN" (regex can't catch every paraphrase) — but it must be a documented acceptance threshold with an escalation if exceeded. config.json:37 says escalate_high_severity: true — R-ASSIST-07 is high-severity, so the plan must either escalate or document why the residual risk is acceptable.
    • Confidence: 0.68
    • Challenge: The plan accepts an unquantified residual risk on a safety-critical surface. "We'll measure it and trend it nightly" is necessary but not sufficient — what happens if the nightly trend shows 15% FN? The plan has no trigger. This is the grill's hardest call.
    • Decision: G-067 (MUST) — R-ASSIST-07 (guardrail false-negative) must have a documented acceptance threshold before EXECUTE. The adversarial FN rate (REQ-IDEATE-01) must be: (a) measured pre-ship (TASK-04-02), (b) compared against a threshold (e.g., "adversarial FN ≤ 20% acceptable for pilot because defense-in-depth + audit + v0.6 LLM-as-judge mitigate; >20% triggers a re-tuning wave or escalation"), (c) the threshold + the mitigation rationale documented in the ship notes. This is NOT a "0% FN" demand — it is a "know your residual risk + decide if it's acceptable" demand. The plan's current "reported but not threshold-gated" is insufficient for a safety-critical surface. config.json:37 escalate_high_severity: true is the governing constraint. (0.68)
  • Q4: Pre-mortem — "It's 12 months from now and v0.5 failed. Why?"

    • Evidence: RESEARCH-v0.5 risks; PLAN-v0.5 risk matrix.
    • Answer: The most likely failure modes (in order):
      1. A guardrail bypass incident during a real customer call (R-ASSIST-07). A direct answer slipped past the regex, the learner parroted it, the customer escalated to a real manager who disavowed the "AI's advice." The nightly FN trend showed 18% but no one acted because there was no threshold (G-067 gap). This is the highest-consequence failure — it breaks trust in the product + the learner's job.
      2. PIPEDA complaint (R-ASSIST-08 / D-073). A real customer discovered they were recorded by the learner's mic without their consent. The disclosure (D-070) was shown to the learner, not the customer. Canada's two-party consent law (if applicable in the province) was not reviewed. This is the highest-legal-consequence failure.
      3. The in-loop guardrail processor's retry mechanism was infeasible in Pipecat (G-049). The "one retry" (D-068) became "canned fallback only" — safe but degraded. The assist coaching quality dropped (every block → canned fallback, no second chance). Learners stopped using assist because the coaching felt robotic.
      4. The latency was >650ms in practice (R-ASSIST-02). The ~655ms estimate was optimistic; actual p95 was ~720ms. Coaching arrived after the customer moment passed. Learners abandoned assist for being "too slow to be useful."
    • Confidence: 0.75
    • Decision: G-068 — pre-mortem top-4: guardrail bypass (highest consequence, G-067 gap), PIPEDA complaint (ESCALATION-01), in-loop retry infeasible (G-049), latency >650ms (D-072 pilot tolerance). All four are addressed in binding decisions/escalations. (0.75)

Axis 8 — Governance, Decision-Making, and Communication

  • Q1: Decision-maker — autonomy=full, the CI decides. Is there a human escalation path for safety-critical decisions? (config.json escalation_hooks: deploy, delete_data, merge_to_main — none for "ship safety-critical guardrail". Is this a gap?)

    • Evidence: config.json:14 — "escalation_hooks": ["deploy", "delete_data", "merge_to_main"]; config.json:37 — "escalate_high_severity": true; PROJECT.md:5 — "Autonomy: full."
    • Answer: The escalation_hooks list does NOT include "ship safety-critical guardrail" or "legal review." The escalate_high_severity: true security config is the only safety valve — it says the CI should escalate high-severity security issues, but the mechanism (how? to whom?) is unspecified. For v0.1-v0.4 (practice surface), this was acceptable — the worst case was a bad role-play. For v0.5 (Live Assist, real customers), the worst case is a guardrail bypass during a real call + a PIPEDA complaint. The escalation path for these is the grill itself — this document is the escalation mechanism. The grill's ESCALATION-01 (PIPEDA) + G-067 (guardrail threshold) are the safety-critical escalations/binding decisions. The gap: there is no ongoing human escalation path post-ship. If the nightly FN trend spikes post-ship, the CI auto-mitigates (config.json:36) but does not escalate to a human (no hook for "safety signal spike"). This is a v0.6+ governance gap, not a v0.5 blocker — v0.5 ships the measurement (REQ-IDEATE-04 nightly trending); v0.6 adds the LLM-as-judge + the escalation on spike.
    • Confidence: 0.70
    • Decision: G-069 — escalation path: the grill is the safety-critical escalation mechanism (ESCALATION-01 + G-067). config.json escalate_high_severity: true is the governing constraint. Post-ship ongoing escalation (safety signal spike → human) is a v0.6+ governance gap — v0.5 ships the measurement, v0.6 adds the response. Accept for pilot with documented gap. (0.70)
  • Q2: Governance cadence — the pipeline stages are the governance. Is the grill the right gate for a safety-critical surface?

    • Evidence: ROADMAP.md:21 — "Pipeline stages: SPECIFY → CLARIFY → RESEARCH → IDEATE → PLAN → GRILL → SHIP"; ROADMAP.md:30 — "GRILL-v0.5.md (adversarial review — real-customer interaction warrants grill)."
    • Answer: The grill is the right gate — ROADMAP.md:30 explicitly flags "real-customer interaction warrants grill." The pipeline stages (SPECIFY→…→GRILL→SHIP) are the governance cadence; the grill is the crisis-cadence (this document). For a safety-critical surface, the grill is the only human-in-the-loop checkpoint (the CI runs the rest autonomously). This is the correct model — the grill surfaces the safety-critical decisions (G-067, ESCALATION-01) for human attention before SHIP.
    • Confidence: 0.82
    • Decision: G-070 — grill is the right gate for a safety-critical surface (ROADMAP:30 explicit). The grill is the human-in-the-loop checkpoint. Accept. (0.82)
  • Q3: What's omitted from status reports? (The LSP errors in server/__main__.py, test_scenario_library.py — are these reported or hidden?)

    • Evidence: Task context mentions "LSP errors in server/__main__.py, test_scenario_library.py"; verification: python3 -m py_compile server/__main__.py → exit 0 (clean); python3 -m py_compile tests/test_scenario_library.py → exit 0 (clean).
    • Answer: The "LSP errors" claim in the task context is unverified — both files compile cleanly (py_compile exit 0). This may refer to type-checking (pyright/mypy) warnings, not syntax errors, or it may be stale. The grill does not flag this as a material omission — the files compile, the v0.4 tests pass (317 pass, 0 fail per REVIEW.md). If there are type-checking warnings, they are non-blocking (the codebase doesn't enforce strict typing in CI). No omission found.
    • Confidence: 0.80
    • Decision: G-071 — no status-report omission found. The "LSP errors" claim is unverified (files compile clean). Type-checking warnings, if any, are non-blocking. Accept. (0.80)
  • Q4: Stop-the-project trigger — is there one? (If the grill returns RETHINK, does the pipeline stop?)

    • Evidence: config.json:13 — full autonomy; GRILL-v0.4 G-032 — "no human stop trigger (full autonomy). The grill is the stop mechanism."
    • Answer: No human stop trigger (full autonomy, G-032 carry-forward). The grill is the stop mechanism — if the verdict were "Rethink" or "Escalate" on a material axis, the pipeline would stop. This grill's verdict is "Proceed-with-conditions" — the project proceeds after the MUSTs (G-049, G-067) + the escalation (ESCALATION-01) are resolved. The escalation (PIPEDA) is the de facto stop trigger — if the human legal review determines the disclosure is insufficient, v0.5 cannot ship the assist surface as designed.
    • Confidence: 0.78
    • Decision: G-072 — no human stop trigger (full autonomy). The grill is the stop mechanism. ESCALATION-01 (PIPEDA) is the de facto stop trigger for the assist surface. This grill = proceed with conditions. (0.78)

Axis 9 — Change, Adoption, and Operational Readiness

  • Q1: Who uses Live Assist? (The learner — during a real shift. How does their work change? They now have an AI in their ear.)

    • Evidence: PROJECT.md:45-47 — "a hands-free voice assistant a learner invokes while actually working"; PERSONAS.md — no learner persona (learners are external to the CI agent); D-071 — tap-to-talk invocation.
    • Answer: The learner uses Live Assist during a real shift. Their work changes: they now have an AI coach in their ear (via earbuds) that they invoke by tapping a button (D-071 — tap-to-talk, not wake-word). "What's in it for them" = real-time coaching during real customer interactions — the transfer moment from practice to job. This is unvalidated — no user testing, no pilot data on whether learners will actually tap a button on their phone during a real customer call (the phone may be in their pocket, the tap may be socially awkward). The tap-to-talk UX (D-071) is the lowest-confidence assumption (G-065, 0.60). The alternative (wake-word, hands-free) is deferred to v0.6. For v0.5 pilot, tap-to-talk is the validation — does a learner use it? The measurement is the assist usage metrics (REQ-NFR-ASSIST-04, cohort aggregation).
    • Confidence: 0.65
    • Challenge: The adoption risk is real — tap-to-talk during a real customer call is socially + ergonomically awkward (phone in pocket, earbuds in, tap a button on the phone screen). The "we'll measure usage" answer is correct but the pilot may show low adoption. This is a v0.5 validation risk, not a v0.5 blocker.
    • Decision: G-073 — Live Assist's first user is the learner during a real shift. Tap-to-talk (D-071) is the unvalidated UX assumption (G-065, 0.60). v0.5 pilot validates adoption (assist usage metrics); v0.6 adds wake-word if tap-to-talk adoption is low. Document in ship notes: v0.5 validates the coaching/guardrail/context-binding value, not the hands-free UX (that's v0.6). (0.65)
  • Q2: Is the ops team involved? (CI project — ops is the LXC deploy. Does v0.5 need deploy changes? D-071 says no — v0.4 LXC carries forward. Is that sound?)

    • Evidence: PERSONAS.md:651-660 — devops-engineer DEACTIVATED for v0.5 ("No deploy changes — v0.4's LXC + Docker-in-LXC + Postgres + backup cron carries forward unchanged"); PLAN-v0.5:1071 — "New pip deps: 0… New npm deps: 0."
    • Answer: v0.5 needs NO deploy changes — 0 new pip deps, 0 new npm deps, no new Docker services, no CT bump. The assist surface is server-side code (server/assist/) + a React route (client/src/AssistControl.tsx) on the existing v0.4 LXC. devops-engineer deactivation is sound. The ops surface (LXC, Postgres, backup) is unchanged. This is the correct posture — v0.5 is a feature milestone, not an infra milestone.
    • Confidence: 0.85
    • Decision: G-074 — v0.5 needs no deploy changes (0 new deps, no CT bump, v0.4 LXC carries forward). devops-engineer deactivation is sound. Accept. (0.85)
  • Q3: Rollback plan — if v0.5 ships and a guardrail incident occurs, what's the rollback? (Disable assist mode? Revert to v0.1.9?)

    • Evidence: config.json:40 — "branching_strategy": "phase"; PLAN-v0.5 — per-phase ship (v0.1.11, v0.1.12, v0.1.13); git revert pattern (GRILL-v0.4 G-035).
    • Answer: Rollback is per-phase git revert (G-035 carry-forward). But for a guardrail incident (R-ASSIST-07), the rollback is operational, not just git:
      • Preventive rollback: disable assist mode (revert to v0.1.9 = v0.4). The assist routes (/api/assist/*) + the assist WebRTC endpoint are removed. The practice surface (v0.1-v0.4) continues unchanged. This is a clean revert — the assist surface is additive (new routes, new server/assist/ package, new SQLite migration 0004). Reverting removes the routes + the package; the migration is additive (session_type defaults to 'practice', guardrail_verdict_json is nullable) so existing practice sessions are unaffected.
      • Corrective rollback: impossible. Once a guardrail bypass reaches a learner's ear during a real call, the turn has played. The audit log (REQ-IDEATE-09 incremental write) records it for investigation, but the incident cannot be rolled back. This is the nature of a live surface — rollback is preventive (disable), not corrective.
    • The preventive rollback (disable assist) is clean + tested (the assist surface is additive). The corrective impossibility is accepted (the audit log is the post-incident tool, not a rollback).
    • Confidence: 0.75
    • Decision: G-075 — rollback is preventive (disable assist mode → revert to v0.1.9). The assist surface is additive (clean revert). Corrective rollback is impossible (a live turn cannot be un-played) — the audit log (REQ-IDEATE-09) is the post-incident tool. Accept the preventive-only rollback. (0.75)
  • Q4: Has anyone validated the success criteria with the people who will judge v0.5 successful? (NFRs are research-grounded, not measurement-validated.)

    • Evidence: REQUIREMENTS.md:22-25 — NFRs research-grounded; REQ-IDEATE-04 — measurable targets (p95 ≤650ms, FP<5%); config.json:13 — full autonomy (CI is the judge).
    • Answer: No human judge (full autonomy, G-036 carry-forward). The CI is the judge. The success criteria = 16/16 REQ coverage + per-slice acceptance + REQ-IDEATE-04 measurable NFRs. The NFRs are research-grounded (estimated, not measured) — REQ-IDEATE-04 + SLICE-09 (P2) add the measurement. The validation path: P2 SLICE-09 measures p95 latency + FP/FN rates. If p95 >650ms or FP>5%, the P2 verify stage flags it. This is the measurement-validated path — but it happens in P2, not pre-ship. Gap: the success criteria are validated during P2, not before P1 ship (v0.1.11). If P1 ships with a guardrail that has FP>5%, the P1 ship is premature. The mitigation: TASK-04-02 (guardrail tuning test) is in P1 Wave 2 — it runs before P1 ship. If it fails, P1 doesn't ship. This is the correct gate.
    • Confidence: 0.72
    • Decision: G-076 — success criteria are research-grounded, measurement-validated in P2 (SLICE-09). The P1 gate is TASK-04-02 (guardrail tuning test, FP<5% / direct-FN<5%) — runs before P1 ship. If it fails, P1 doesn't ship. Accept with TASK-04-02 as the P1 gate + SLICE-09 as the P2 measurement. (0.72)

Meta — Closing Review

  • Q1: If you were the auditor, what would you flag?

    • Evidence: all axes above.
    • Answer: Four flags:
      1. R-ASSIST-07 residual risk acceptance without a threshold (G-067). The plan accepts an unquantified adversarial FN rate on a safety-critical surface. This is the grill's hardest call — the bar must be set.
      2. PIPEDA legal review deferred (ESCALATION-01). Shipping a recording device into real customer interactions without legal sign-off is a regulatory risk the CI cannot own.
      3. IDEATE scope expansion +128% (G-046). The first use of ideation expanded v0.5 from 7 to 16 REQs. The additions are defensive, but the expansion is the largest in project history — future ideation must maintain risk-reduction discipline.
      4. In-loop guardrail processor is a structural pipeline change (G-049). The research frames it as "~1 new frame processor" but the retry mechanism is unvalidated against Pipecat semantics. This is the highest-novelty code on the safety-critical path.
    • Confidence: 0.78
    • Decision: G-077 — auditor flags: R-ASSIST-07 threshold gap, PIPEDA escalation, IDEATE scope expansion, in-loop processor novelty. All addressed in binding decisions/escalations. (0.78)
  • Q2: What is v0.5 NOT doing that it should? (PIPEDA legal review is deferred D-073 — should it block ship?)

    • Evidence: D-073 (PROJECT.md:243); ESCALATION-01 (Axis 2).
    • Answer:
      1. PIPEDA legal review — deferred, escalated (ESCALATION-01). The grill cannot determine if it blocks ship — that's a legal question. The disclosure (D-070) is the engineering mitigation; the legal review is the regulatory mitigation.
      2. Post-ship safety signal escalation — the nightly FN trend (REQ-IDEATE-04) measures but does not escalate on spike (G-069). v0.6 adds the LLM-as-judge + the escalation response.
      3. Guardrail red-team prompt set — REQ-IDEATE-01 builds a synthetic tuning corpus (LLM-generated coaching vs direct-answer responses). This is NOT a human red-team prompt set — a determined adversary (or a clever learner) may find paraphrases the synthetic corpus doesn't cover. The adversarial test (TASK-04-02) is the best available, but it's synthetic, not human. This is an accepted limitation (pilot).
    • Confidence: 0.75
    • Decision: G-078 — v0.5 is NOT doing: PIPEDA legal review (escalated), post-ship safety escalation (v0.6), human red-team prompt set (synthetic corpus accepted for pilot). All documented. Accept with ESCALATION-01 as the human-action item. (0.75)
  • Q3: Simplest possible version — is tap-to-talk-only (D-071) the 80/20, or is even that over-scoped?

    • Evidence: D-071 (PROJECT.md:241); PLAN-v0.5:25 — "Why tap-to-talk only in v0.5."
    • Answer: Tap-to-talk-only (D-071) IS the 80/20. The wake-word deferral stripped the client-architecture risk (React Native), the battery risk, the Picovoice commercial risk, and 5 of 14 research risks. What remains is the core value: the guardrail (safety), the context-binding (relevance), the shift-bounded session (aggregation), the in-loop processor (pipeline integration). Cutting further (e.g., deferring the guardrail tuning corpus REQ-IDEATE-01) would ship a guardrail "blind on its two most safety-critical metrics" (FP/FN) — unacceptable for a safety-critical surface. The 16 REQs are the minimum for a safe assist pilot. Tap-to-talk is the 80/20; the guardrail work is the non-negotiable 20%.
    • Confidence: 0.80
    • Decision: G-079 — tap-to-talk-only (D-071) is the 80/20. The guardrail work (REQ-IDEATE-01/04, the tuning corpus + measurable NFRs) is the non-negotiable 20%. Cutting further would ship an unvalidated safety-critical surface. Accept the scope. (0.80)
  • Q4: What would have to be true for v0.5 to succeed in the next 90 days, and is it true today?

    • Evidence: all axes.
    • Answer: For v0.5 to succeed:
      1. The guardrail must be robust (FP<5%, direct-FN<5%, adversarial FN known + thresholded). Is it true today? No — the tuning corpus is built in P1 Wave 2 (TASK-04-01/02). Will be true after P1 if the test passes. G-067 sets the threshold.
      2. The in-loop guardrail processor must work in Pipecat (retry mechanism). Is it true today? No — unvalidated (G-049). Will be true after the Wave-1/2 spike.
      3. PIPEDA must be addressed (legal review or disclosure-sufficient determination). Is it true today? No — deferred (ESCALATION-01). Will be true only after human legal review.
      4. The latency must be ≤650ms. Is it true today? No — unmeasured (D-072). Will be true after P2 SLICE-09 measurement.
      5. The tap-to-talk UX must be usable during a real shift. Is it true today? No — unvalidated (G-065). Will be true only after pilot deployment (v0.5's validation purpose).
    • 2 of 5 are addressable in P1/P2 (guardrail robustness, in-loop processor). 1 requires human action (PIPEDA). 2 are post-ship validation (latency measurement, UX adoption). This is the expected state for a pilot — the plan is ready; the proof is in execution.
    • Confidence: 0.72
    • Decision: G-080 — 5 success conditions: guardrail robustness (P1 gate, G-067), in-loop processor (P1 spike, G-049), PIPEDA (human escalation, ESCALATION-01), latency (P2 measurement), UX adoption (post-ship validation). 2 addressable in P1/P2, 1 requires human, 2 post-ship. Accept — the plan is ready, the proof is in execution. (0.72)

v0.5-Specific Probes (Signature Questions)

Probe 1 — R-ASSIST-07 (Guardrail false-negative): Is "defense-in-depth + audit + v0.6 LLM-as-judge" enough for a safety-critical surface?

Question: The AI is in a learner's ear during a real customer call. The regex output filter (D-068) is the on-voice-path guardrail. The adversarial FN rate is "reported but not threshold-gated" (PLAN:419). If a direct answer slips past the regex, the learner may parrot it. Is the 3-layer defense (prompt + regex + audit) + nightly trending + v0.6 LLM-as-judge sufficient, or does the grill need to set a binding threshold?

Evidence:

  • R-ASSIST-07 (RESEARCH-v0.5 §2.6) — "The 'parrot' failure: the AI gives a verbatim script, the learner repeats it word-for-word, the customer detects the robotic delivery → trust erosion."
  • D-068 (PROJECT.md:238) — "regex-based direct-answer + false-authority + impersonation patterns, with one retry on block + canned coaching redirect fallback."
  • PLAN-v0.5:419 — "The adversarial FN rate is reported but not threshold-gated (it's the residual risk, mitigated by defense-in-depth)."
  • config.json:37 — "escalate_high_severity": true.
  • REQ-IDEATE-10 (v0.6 backlog) — "LLM-as-judge guardrail evaluation (nightly, off-voice-path) — measure the true false-negative rate the regex filter cannot."

Analysis: The plan's posture is: regex is the fast on-voice-path filter (D-068); the LLM-as-judge is the accurate off-voice-path backstop (v0.6, REQ-IDEATE-10). The gap is v0.5: the regex is the only on-voice-path guardrail, and its adversarial FN rate is unthresholded. For a safety-critical surface where the worst case is a guardrail bypass during a real customer call, "we'll measure it and trend it nightly" is necessary but not sufficient — the plan needs a decision: what FN rate is acceptable for the pilot, and what happens if it's exceeded?

The config says escalate_high_severity: true — R-ASSIST-07 is high-severity. The plan accepts the residual risk without escalating. This is the tension G-055 identified. The resolution: the grill sets the threshold (G-067) — the adversarial FN rate must be measured pre-ship (TASK-04-02), compared against a documented threshold, and the threshold + mitigation rationale documented in the ship notes. This is NOT a "0% FN" demand (impossible for regex) — it is a "know your residual risk + decide if it's acceptable" demand.

The defense-in-depth (prompt + regex + audit) is the correct architecture — the grill does not dispute the 3-layer pattern (RESEARCH §2.1, 0.85 confidence). The issue is the threshold, not the architecture. The v0.6 LLM-as-judge is the future backstop, not the current mitigation — v0.5 ships with regex + audit only.

Verdict: Defense-in-depth is the correct architecture; the missing piece is a documented acceptance threshold for the adversarial FN rate. G-067 (MUST) sets this. The plan's "reported but not threshold-gated" is insufficient for a safety-critical surface — the grill requires a threshold + an escalation if exceeded. Confidence: 0.68.


Question: The ambient mic captures the real customer (a third party). ASR transcribes their speech. The turns table stores it (REQ-IDEATE-05). Canada's PIPEDA + provincial consent laws govern recording. D-073 defers the legal review to "Phase 1 implementation." The disclosure (D-070) is shown to the learner, not the customer. Is the disclosure sufficient, or does the legal review need to block ship?

Evidence:

  • D-073 (PROJECT.md:243) — "PIPEDA consent-law review = defer to v0.5 Phase 1 implementation; document as R-ASSIST-08 in the grill."
  • D-070 (PROJECT.md:240) — consent disclosure: "Praxis Assist is on — those around you may be recorded by your mic."
  • R-ASSIST-08 (RESEARCH-v0.5 §2.6) — "the real customer didn't consent to being recorded/analyzed by an AI."
  • REQ-IDEATE-05 (REQUIREMENTS.md:52) — "The ambient mic captures BOTH the learner and the real customer; ASR transcribes both; the turns table stores transcribed text. The customer is a third party."
  • config.json:13 — full autonomy (CI cannot resolve legal questions).

Analysis: This is a legal question, not a technical one. The CI agent under full autonomy cannot determine whether Canada's PIPEDA + provincial consent law requires:

  • (a) One-party consent (the learner's consent is sufficient — the disclosure D-070 covers this).
  • (b) Two-party consent (the customer must consent — Praxis cannot notify the customer, so the assist surface may be illegal in two-party provinces).
  • (c) A PIPEDA-compliant privacy policy + data handling agreement.

The disclosure (D-070) is the engineering mitigation — it makes the learner aware. It does NOT make the customer aware, and it does NOT determine the legal consent regime. The PII policy (REQ-IDEATE-05) retains customer speech with redaction + 30-day retention — this is a data handling mitigation, not a consent determination.

The grill's confidence that the disclosure is sufficient: 0.55 — below the 0.60 threshold. The grill cannot resolve this under full autonomy. This is an escalation.

Verdict: PIPEDA legal review is a hidden regulatory requirement that the CI cannot resolve. The disclosure (D-070) is the engineering mitigation but not a legal determination. Escalate to human attention (ESCALATION-01): determine whether the disclosure is legally sufficient or whether two-party consent / a PIPEDA privacy policy is required before ship. If the disclosure is sufficient, proceed; if not, the assist surface may need geographic restriction or customer-facing consent (out of scope for v0.5). Confidence: 0.55 — below threshold, escalated.


Probe 3 — IDEATE scope expansion (+128%): Risk-reduction or scope creep?

Question: v0.5 started with 7 REQs (3 ASSIST + 4 NFR, post-CLARIFY). IDEATE added 9 REQs (+128%) — the largest scope growth in project history. Are the 9 additions risk-reduction (guardrail, PII, mode-conflict, resilience, audit, tech-debt, cost, NFR measurability) or scope creep with a defensive veneer?

Evidence:

  • git log b8c7de8 — "ideation results — 9 accepted into v0.5, 4 accepted into v0.6."
  • REQUIREMENTS.md:29-70 — 9 IDEATE REQs.
  • PLAN-v0.5:1011 — "16/16 REQ-IDs covered."

Analysis: The 9 IDEATE REQs map to named risks:

  • REQ-IDEATE-01 (guardrail tuning corpus) → R-ASSIST-06/07 (FP/FN).
  • REQ-IDEATE-02 (in-loop processor test) → REQ-IDEATE-02 interface gap (GuardrailContext.role).
  • REQ-IDEATE-03 (mode-conflict) → D-061 mutual exclusivity gap.
  • REQ-IDEATE-04 (measurable NFRs) → REQ-NFR-ASSIST-01/03 verifiability.
  • REQ-IDEATE-05 (PII policy) → R-ASSIST-08 (STRIDE information-disclosure).
  • REQ-IDEATE-06 (tech-debt) → 8 v0.4 P1+ findings.
  • REQ-IDEATE-07 (cost tracking) → C-3 budget.
  • REQ-IDEATE-08 (WebRTC reconnect) → R-ASSIST-09.
  • REQ-IDEATE-09 (incremental audit-log) → R-ASSIST-14 abrupt termination.

Every addition maps to a named risk or a carried-forward finding. None are features. The expansion is risk-reduction, not scope creep. The +128% is large but justified — v0.5 is the first safety-critical milestone, and the IDEATE stage surfaced the defensive requirements the practice surface (v0.1-v0.4) didn't need. The 4 deferred to v0.6 (REQ-IDEATE-10..13) are also risk-reduction (LLM-as-judge, assist-weaning, offline mode, voice-only context) — the ideation was disciplined.

Verdict: The IDEATE expansion is risk-reduction, not scope creep. Every REQ maps to a named risk. Accepted (G-046). Future ideation must maintain this discipline — the grill will flag any IDEATE addition that doesn't map to a named risk. Confidence: 0.78.


Probe 4 — In-loop guardrail processor (structural pipeline change): Is the "minimal delta" framing accurate?

Question: RESEARCH §5.2 frames the assist pipeline as "minimal delta: ~1 new pipeline builder, ~1 new guardrail processor." But the v0.1 pipeline has NO in-loop guardrail (the CS guardrail runs on the debrief). Is the in-loop processor a "minimal delta" or a structural change?

Evidence:

  • server/pipeline.py:143-185 — build_pipeline() has no in-loop guardrail processor (transport → stt → latency → user_agg → llm → latency → tts → latency → transport → assistant_agg).
  • RESEARCH-v0.5 §5.2 — "v0.5 adds an in-loop guardrail processor for assist mode. This is a pipeline-structure change but a small one (~1 new Pipecat frame processor)."
  • server/guardrails/customer_service.py — CS guardrail runs check() standalone, not as a frame processor.
  • PLAN-v0.5 TASK-05-02 — LiveAssistGuardrailProcessor(FrameProcessor) between llm and tts.
  • PLAN-v0.5 Open Question #4 (line 1046) — "verify Pipecat's LLMContextAggregator supports injecting a message + re-running the LLM within a single process_frame call. If not, the retry may need to be a separate pipeline task."

Analysis: The "minimal delta" framing is partially accurate. The service reuse (transport/stt/llm/tts) is genuinely minimal — the constructors are env-driven and reusable (verified: pipeline.py:63-109). But the in-loop guardrail processor is a structural change: the v0.1 pipeline has no post-LLM frame processor; v0.5 inserts one between llm and tts. This is novel for this codebase. The retry mechanism (inject RETRY_INSTRUCTION + re-run LLM mid-stream) is unvalidated against Pipecat's frame semantics — Open Question #4 defers this to EXECUTE, which is too late for a safety-critical path.

The risk: if Pipecat's LLMFullResponseEndFrame doesn't fire as expected, or if the LLMContextAggregator can't inject a retry mid-stream, the guardrail's "one retry" (D-068) becomes "canned fallback only" — safe but degraded. The coaching quality drops (every block → canned fallback, no second chance). This is a quality risk, not a safety risk (the canned fallback is safe) — but it affects the product's value.

Verdict: The in-loop guardrail processor is a structural change, not a minimal delta. The retry mechanism must be validated before Wave 3 (G-049 MUST). If Pipecat can't do mid-stream retry, document the fallback (canned-only) + update D-068's safety posture. The "minimal delta" framing should be corrected in the plan. Confidence: 0.70.


Probe 5 — Tap-to-talk UX (D-071): Is the unvalidated adoption risk acceptable for a pilot?

Question: D-071 ships tap-to-talk only (no wake-word). The learner taps a button on their phone during a real customer call. The phone may be in their pocket. The tap may be socially awkward. No user testing validates this UX. Is the pilot the validation, or is this a feature looking for a user?

Evidence:

  • D-071 (PROJECT.md:241) — "tap-to-talk ONLY (no wake-word in v0.5)… learner taps a button to invoke an assist turn during a real shift."
  • G-065 (Axis 7) — tap-to-talk UX assumption confidence 0.60 (lowest).
  • RESEARCH-v0.5 §4.1 — "No direct competitor does live-in-ear coaching during real customer calls on a $100 phone" (novel surface, no comparable UX to benchmark).

Analysis: Tap-to-talk is a proven pattern for walkie-talkie apps (Zello, Voxer) — users tap+hold to speak, release to send. This is a reasonable UX for hands-free-adjacent interaction. But those apps are the primary interface (the user opens the app to talk); Praxis assist is a secondary interface (the learner is in a real customer call, the phone is in their pocket, they tap a button on a screen they can't see). The social + ergonomic gap is real: the learner must (a) have earbuds in, (b) have the phone accessible, (c) tap a button without looking, (d) do this during a live customer interaction. This is a high-friction UX.

The pilot is the validation — v0.5 measures assist usage (REQ-NFR-ASSIST-04 cohort metrics). If adoption is low, v0.6 adds wake-word (the hands-free target). This is the correct pilot posture: ship the value (coaching/guardrail/context-binding), validate the UX (tap-to-talk adoption), iterate in v0.6. The risk is that low adoption makes the pilot a failure — but the pilot's purpose is to find out, not to prove adoption.

Verdict: Tap-to-talk is an unvalidated but reasonable UX for a pilot. The pilot is the validation. v0.6 adds wake-word if adoption is low. Accept with documented risk (G-073). Confidence: 0.65.


Probe 6 — 2-phase split: Is P1 (assist core + guardrail) independently shippable without P2 (measurement + tech-debt)?

Question: P1 ships v0.1.11 (assist core + guardrail, 12 REQs). P2 ships v0.1.12 (integration + tech-debt + NFR measurement, 4 REQs). Is P1 independently shippable — does a learner get a safe assist experience without P2?

Evidence:

  • PLAN-v0.5:17-23 — P1 = assist voice loop + guardrail (12 REQs, 24 tasks); P2 = integration + measurement + tech-debt (4 REQs, 9 tasks).
  • config.json:110 — "per_phase": true (per-phase ship).

Analysis: P1 delivers: the assist voice loop (build_assist_pipeline), the 3-layer guardrail (LiveAssistGuardrail + tuning corpus + adversarial test), the shift-bounded session model, the tap-to-talk client, the warm WebRTC + reconnect, the incremental audit-log, the mode-conflict guard, the PII policy. A learner can start a shift, tap-to-talk, get coaching with guardrails, end the shift. This is a safe, usable assist experience.

P2 adds: the cohort aggregation assist metrics (operator visibility), the cost tracking (C-3 check), the NFR measurement (p95 latency, FP/FN rates), the tech-debt wave (8 v0.4 P1+ findings). P2 is hardening + visibility, not safety. The guardrail's safety is in P1 (SLICE-03/04/08); P2 measures the guardrail's FP/FN rates (SLICE-09) but the guardrail itself ships in P1.

The one caveat: the aggregation cache tech-debt (P1+ #7) corrupts assist_active_learners_count during P1 (G-051). But P1 doesn't ship operator visibility (the cohort dashboard extension is P2 SLICE-10) — so the corrupted metric is not visible during P1. The fix lands in P2 before the dashboard extension. This is a sequencing dependency, not a P1 safety gap.

Verdict: P1 is independently shippable — a learner gets a safe assist experience. P2 is hardening + operator visibility + measurement. The split is clean (P1 = safety-critical voice loop, P2 = hardening). The aggregation cache corruption during P1 is not visible (no dashboard in P1) and fixed in P2 before visibility. Confidence: 0.82.


v0.4 Grill Deferred Items — Coverage Check

The v0.4 grill (GRILL-v0.4.md) deferred no items to v0.5 (v0.4 was the operator tier, complete). The v0.4 grill's 8 P1+ findings are carried forward as REQ-IDEATE-06 (tech-debt wave, P2 SLICE-12). Let me verify:

v0.4 Grill/Finding v0.5 Coverage Status
G-008 (backup drill) v0.4 complete (REVIEW.md:240) Resolved in v0.4
G-011 (two-store fallback) v0.4 complete (REVIEW.md:241) Resolved in v0.4
G-027 (first-boot no v0.3 key) v0.4 complete (REVIEW.md:242) Resolved in v0.4
G-031 (R-AUTH-01 reframe) v0.4 complete (REVIEW.md:243) Resolved in v0.4
G-038 (differencing-attack test) v0.4 complete (REVIEW.md:244) Resolved in v0.4
G-041 (SPA fallback subclass) v0.4 complete (REVIEW.md:245) Resolved in v0.4
P1+ #1 (argon2id blocking) REQ-IDEATE-06, TASK-12-04 Covered in v0.5 P2
P1+ #2 (rate-limit mock test) REQ-IDEATE-06, TASK-12-04 Covered in v0.5 P2
P1+ #3 (cookie-secret length) REQ-IDEATE-06, TASK-12-02 Covered in v0.5 P2
P1+ #4 (credential status enum) REQ-IDEATE-06, TASK-12-03 Covered in v0.5 P2
P1+ #5 (revocation audit log) REQ-IDEATE-06, TASK-12-04 Covered in v0.5 P2
P1+ #6 (nightly zoneinfo) REQ-IDEATE-06, TASK-12-04 Covered in v0.5 P2
P1+ #7 (aggregation cache) REQ-IDEATE-06, TASK-12-01 Covered in v0.5 P2 (critical path for assist metrics — G-051)
P1+ #8 (f-string SQL) REQ-IDEATE-06, TASK-12-03 Covered in v0.5 P2

Verdict: 6/6 v0.4 grill MUSTs resolved in v0.4. 8/8 v0.4 P1+ findings covered in v0.5 P2 SLICE-12 (REQ-IDEATE-06). The aggregation cache fix (P1+ #7) is on the v0.5 critical path for correct assist metrics (G-051).


Binding Decisions

ID Axis Decision Confidence Type
G-042 1 Live Assist is the correct next priority (delivers the transfer surface). Novel per RESEARCH §4.1. 0.80 ACCEPT
G-043 1 CI is the named sponsor under full autonomy (G-002 carry-forward). 0.80 ACCEPT
G-044 1 v0.5 is not a zombie (delivers the transfer surface). Practice surface works without it. 0.78 ACCEPT
G-045 1 No financial ROI; ROI is product-completeness + safety-surface foundation. REQ-IDEATE-07 measures cost. 0.68 ACCEPT
G-046 2 IDEATE scope expanded +128% (7→16 REQs). Accepted — all 9 additions are risk-reduction, map to named risks. Future ideation must maintain discipline. 0.78 ACCEPT
G-047 2 NFRs are research-grounded, not frozen. REQ-NFR-ASSIST-01 at-risk (D-072 pilot tolerance). REQ-IDEATE-04 provides measurable freeze. 0.75 ACCEPT
G-048 2 Out-of-scope is explicit. D-071 (wake-word deferred) is the key scope reduction, binding. 0.85 ACCEPT
G-049 3 MUST: In-loop guardrail processor retry mechanism (TASK-05-02) must be validated against Pipecat frame semantics BEFORE Wave 3. Add a Wave-1/2 spike: verify LLMFullResponseEndFrame + LLMContextAggregator retry injection. If infeasible, document canned-fallback-only + update D-068. Binding contract, not open question. 0.70 MUST
G-050 3 3 integration points, all additive. Cohort aggregation (low) + mastery separation (low) + voice pipeline (medium, G-049). Aggregation cache tech-debt on critical path (G-051). 0.75 ACCEPT
G-051 3 8 v0.4 P1+ findings inherited, budgeted in P2 SLICE-12. Aggregation cache fix corrupts assist metrics during P1 — accept (P1 ships voice loop, not operator dashboard). Document in P1 ship notes. 0.72 ACCEPT
G-052 3 Tech-debt budgeted (4 tasks in P2 SLICE-12, should priority). Proportional. 0.80 ACCEPT
G-053 4 Key-person: voice-engineer (new capability, largest territory), security-engineer (guardrail), backend-engineer (session API). Voice-engineer highest risk (first activation). 0.78 ACCEPT
G-054 4 5 active personas, max 5 concurrent (at limit, no slack). Peak parallelism 2-3 slices. Voice-engineer capability claimed but undemonstrated — G-049 is the test. 0.72 ACCEPT
G-055 4 CI is product owner (full autonomy). For safety-critical surface, escalate_high_severity: true governs. R-ASSIST-07 must be escalated or documented as medium (Probe 1). 0.68 ACCEPT
G-056 4 Team building 3 new capabilities (in-loop processor, warm WebRTC, regex tuning). All on safety-critical/critical path. Acceptable for pilot with G-049 de-risking. 0.72 ACCEPT
G-057 5 2-phase split evidence-based (P1 safety-critical voice loop, P2 hardening + measurement). P1 independently shippable. 0.82 ACCEPT
G-058 5 Critical-path: guardrail tuning corpus (FP/FN rates). TASK-04-02 is the gate. G-049 de-risks secondary path. 0.75 ACCEPT
G-059 5 33 tasks evidence-based (smaller than v0.4's 52 due to D-071 + 0 new deps). Bottom-up sized. 0.80 ACCEPT
G-060 5 Definition of done = per-slice acceptance + per-phase ship + verify + REQ-IDEATE-04 measurable NFRs (p95 ≤650ms, FP<5%). 0.82 ACCEPT
G-061 6 Assist cost driver budgeted (SLICE-11, ~$0.20/month, well under C-3). Diagnostic, not enforced. 0.80 ACCEPT
G-062 6 C-3 relaxation (D-012) remains valid for v0.5 pilot. Assist adds ~$0.20/month. Post-pilot path preserves $3. 0.78 ACCEPT
G-063 6 Burn rate: ~0.8-1.0 days estimated (proportional to v0.4, -37% tasks). 0.78 ACCEPT
G-064 6 No budget contingency. D-071 removed Picovoice MAU-pricing dependency. 0 external commercial dependencies. 0.85 ACCEPT
G-065 7 3 core assumptions: tap-to-talk UX (0.60, unvalidated), regex guardrail (0.65, residual risk), ≤650ms latency (0.65, unmeasured). All pilot-scale with v0.6 hardening. 0.63 ACCEPT
G-066 7 4 dependencies: Picovoice (deferred ), PIPEDA (escalation ⚠️), cohort pipeline (additive ), voice pipeline (structural ⚠️ G-049). 0.75 ACCEPT
G-067 7 MUST: R-ASSIST-07 (guardrail false-negative) must have a documented acceptance threshold before EXECUTE. Adversarial FN rate (REQ-IDEATE-01) must be: (a) measured pre-ship (TASK-04-02), (b) compared against a threshold (e.g., "≤20% acceptable for pilot because defense-in-depth + audit + v0.6 LLM-as-judge mitigate; >20% triggers re-tuning or escalation"), (c) threshold + rationale documented in ship notes. Not a "0% FN" demand — a "know your residual risk + decide" demand. config.json:37 escalate_high_severity governs. 0.68 MUST
G-068 7 Pre-mortem top-4: guardrail bypass (G-067 gap), PIPEDA (ESCALATION-01), in-loop retry (G-049), latency >650ms (D-072). All addressed. 0.75 ACCEPT
G-069 8 Escalation path: grill is the safety-critical mechanism (ESCALATION-01 + G-067). Post-ship ongoing escalation (safety spike → human) is v0.6+ gap. Accept for pilot. 0.70 ACCEPT
G-070 8 Grill is the right gate for safety-critical surface (ROADMAP:30 explicit). Human-in-the-loop checkpoint. 0.82 ACCEPT
G-071 8 No status-report omission. "LSP errors" claim unverified (files compile clean). Type-checking warnings non-blocking. 0.80 ACCEPT
G-072 8 No human stop trigger (full autonomy). Grill is the stop mechanism. ESCALATION-01 (PIPEDA) is the de facto stop trigger for the assist surface. 0.78 ACCEPT
G-073 9 Live Assist's first user is the learner during a real shift. Tap-to-talk (D-071) is unvalidated UX (0.60). v0.5 validates adoption; v0.6 adds wake-word if low. 0.65 ACCEPT
G-074 9 v0.5 needs no deploy changes (0 new deps, no CT bump, v0.4 LXC carries forward). devops-engineer deactivation sound. 0.85 ACCEPT
G-075 9 Rollback is preventive (disable assist → revert to v0.1.9). Assist surface is additive (clean revert). Corrective rollback impossible (live turn cannot be un-played) — audit log is post-incident tool. 0.75 ACCEPT
G-076 9 Success criteria research-grounded, measurement-validated in P2 (SLICE-09). P1 gate = TASK-04-02 (guardrail tuning test, FP<5%/FN<5%). P2 = SLICE-09 measurement. 0.72 ACCEPT
G-077 Meta Auditor flags: R-ASSIST-07 threshold gap, PIPEDA escalation, IDEATE scope expansion, in-loop processor novelty. All addressed. 0.78 ACCEPT
G-078 Meta v0.5 NOT doing: PIPEDA legal review (escalated), post-ship safety escalation (v0.6), human red-team prompt set (synthetic corpus accepted for pilot). 0.75 ACCEPT
G-079 Meta Tap-to-talk-only (D-071) is the 80/20. Guardrail work (REQ-IDEATE-01/04) is the non-negotiable 20%. Cutting further ships an unvalidated safety-critical surface. 0.80 ACCEPT
G-080 Meta 5 success conditions: guardrail robustness (P1 gate), in-loop processor (P1 spike), PIPEDA (human escalation), latency (P2 measurement), UX adoption (post-ship). Plan ready, proof in execution. 0.72 ACCEPT

Escalations

ESCALATION-01 — PIPEDA consent-law review (D-073, R-ASSIST-08). Confidence: 0.55 (below 0.60 threshold).

The ambient mic captures the real customer (a third party); ASR transcribes their speech; the turns table stores it (REQ-IDEATE-05). Canada's PIPEDA + provincial one-party/two-party consent laws govern recording. D-073 defers the legal review to "Phase 1 implementation." The disclosure (D-070) is shown to the learner, not the customer — it is the engineering mitigation, not a legal determination.

The CI agent under full autonomy cannot resolve a legal question. This must be escalated to human attention:

  1. Determine the consent regime: Does Canada PIPEDA + the pilot province's consent law require one-party consent (learner's consent sufficient — D-070 covers) or two-party consent (customer must consent — Praxis cannot notify the customer)?
  2. If one-party: the disclosure (D-070) is sufficient. Proceed with v0.5.
  3. If two-party: the assist surface may need geographic restriction (one-party provinces only) or customer-facing consent (out of scope for v0.5 — would block the assist surface in two-party provinces).
  4. If a PIPEDA privacy policy / data handling agreement is required: the PII policy (REQ-IDEATE-05, 30-day retention + redaction) may need to be formalized into a PIPEDA-compliant policy before ship.

Action required: Human legal review of Canada PIPEDA + provincial consent law for ambient recording during coaching, before v0.5 SHIP. The grill cannot determine with confidence ≥0.60 whether the disclosure is sufficient. This is the de facto stop trigger for the assist surface (G-072).


MUST Conditions Summary (blocking — must be resolved before Phase 1 EXECUTE)

  1. G-049 — In-loop guardrail processor retry validation. Add a Wave-1/2 spike task: verify Pipecat's LLMFullResponseEndFrame fires after the full LLM response + that LLMContextAggregator supports injecting a retry message + re-running the LLM within process_frame. If infeasible, document the fallback (canned-fallback-only, no retry) + update D-068's safety posture. This is a binding contract, not an open question (PLAN Open Question #4 must be resolved pre-EXECUTE).

  2. G-067 — R-ASSIST-07 guardrail false-negative acceptance threshold. The adversarial FN rate (REQ-IDEATE-01) must be: (a) measured pre-ship (TASK-04-02), (b) compared against a documented threshold (e.g., "≤20% acceptable for pilot because defense-in-depth + audit + v0.6 LLM-as-judge mitigate; >20% triggers a re-tuning wave or escalation"), (c) the threshold + mitigation rationale documented in the v0.5 ship notes. The plan's current "reported but not threshold-gated" (PLAN:419) is insufficient for a safety-critical surface. config.json:37 escalate_high_severity: true is the governing constraint.


Escalations Requiring Human Attention (before SHIP)

ESCALATION-01 — PIPEDA consent-law review. Determine whether Canada PIPEDA + provincial consent law requires one-party or two-party consent for ambient recording during coaching. If the disclosure (D-070) is legally sufficient, proceed. If two-party consent is required, the assist surface may need geographic restriction or customer-facing consent (out of scope for v0.5). This is the de facto stop trigger for the assist surface.


FIX Conditions (non-blocking — tracked in VERIFY-P1/P2)

  • G-046 — Document in v0.5 ship notes: IDEATE expanded scope +128% (7→16 REQs). All additions are risk-reduction. Future ideation must maintain risk-reduction discipline.
  • G-051 — Document in P1 ship notes: assist metrics (assist_active_learners_count) are incorrect during P1 due to the aggregation cache tech-debt (v0.4 P1+ #7). Fix lands in P2 SLICE-12 before operator dashboard visibility.
  • G-065 — Document in v0.5 ship notes: tap-to-talk UX (D-071) is the lowest-confidence assumption (0.60, unvalidated). v0.5 pilot validates adoption; v0.6 adds wake-word if low.
  • G-069 — Document in v0.5 ship notes: post-ship safety signal escalation (nightly FN trend spike → human) is a v0.6+ governance gap. v0.5 ships the measurement (REQ-IDEATE-04); v0.6 adds the LLM-as-judge + the escalation response.
  • G-073 — Document in v0.5 ship notes: v0.5 validates the coaching/guardrail/context-binding value, not the hands-free UX (tap-to-talk is the pilot validation; wake-word is v0.6).
  • G-078 — Document in v0.5 ship notes: the guardrail tuning corpus (REQ-IDEATE-01) is synthetic (LLM-generated), not a human red-team prompt set. Accepted limitation for pilot.

ACCEPT Items (proceed as-is)

  • Live Assist is the correct next priority (G-042).
  • CI is the named sponsor under full autonomy (G-043).
  • v0.5 is not a zombie (G-044).
  • IDEATE scope expansion is risk-reduction, not scope creep (G-046, Probe 3).
  • Out-of-scope is explicit; D-071 wake-word deferral is the key scope reduction (G-048).
  • 3 integration points are additive (G-050).
  • Tech-debt is budgeted in P2 SLICE-12 (G-052).
  • Key-person dependency is manageable under parallelization (G-053).
  • 2-phase split is evidence-based; P1 independently shippable (G-057, Probe 6).
  • 33 tasks is evidence-based (G-059).
  • Assist cost is budgeted, well under C-3 (G-061, G-062).
  • No budget contingency — D-071 removed Picovoice dependency (G-064).
  • No deploy changes needed (G-074).
  • Rollback is preventive (disable assist → revert to v0.1.9) (G-075).
  • Tap-to-talk is the 80/20; guardrail work is the non-negotiable 20% (G-079).
  • v0.4 grill MUSTs (6/6) resolved in v0.4; v0.4 P1+ findings (8/8) covered in v0.5 P2.

Bottom Line

The v0.5 plan is not unfeasible — the D-071 tap-to-talk deferral stripped the client-architecture risk, the battery risk, the Picovoice commercial risk, and 5 of 14 research risks. The remaining scope (guardrail + context-binding + shift-bounded session + in-loop processor) is the core safety surface, well-researched and cleanly phased. The plan is not over-scoped after the deferral (16 REQs, but 9 are defensive; 33 tasks vs v0.4's 52). The plan is not a zombie (Live Assist is the v0.1-promised surface, now delivered).

The 2 MUST conditions are surgical:

  • 1 is a validation spike (in-loop guardrail processor retry mechanism — G-049).
  • 1 is a threshold (R-ASSIST-07 adversarial FN rate acceptance — G-067).

The 1 escalation is a legal question the CI cannot resolve (PIPEDA consent-law review — ESCALATION-01). This is the de facto stop trigger for the assist surface.

Resolve the 2 MUSTs, answer the 1 escalation, and v0.5 is a GO.

The v0.5 milestone is the project's first safety-critical surface — the AI is in a learner's ear during real customer interactions. The grill's binding decisions (G-067 threshold, G-049 validation) + the escalation (ESCALATION-01 PIPEDA) are the safety-critical gates. The plan's architecture (3-layer guardrail, defense-in-depth, audit + nightly trending) is sound — the grill's conditions ensure the residual risk is known + decided, not assumed + deferred.