813bd586d6
v0.3 milestone merged to main. Mastery scoring + competency rubrics + verifiable credentials (formative-tier) shipped. 13/13 REQ-IDs covered. Next milestone: v0.4 (operator tier — cohort dashboard + auth + Postgres). ---ci--- project: praxis phase: 2 milestone: v0.3 status: complete milestone_complete: true milestone_merged_to_main: true ---/ci---
291 lines
33 KiB
Markdown
291 lines
33 KiB
Markdown
# Praxis v0.3 CIAgent Plan — GRILL Verdict (Red-Team Review)
|
||
|
||
> **Reviewer:** adversarial technology executive (red-team)
|
||
> **Subject:** v0.3 execution plan (Mastery Scoring + Competency Rubrics) — 2 phases, 15 slices, 70 tasks
|
||
> **Stance:** plan is unfeasible, over-scoped, and too costly until evidence forces otherwise
|
||
> **Date:** 2026-08-03
|
||
> **Binding status:** This GRILL verdict must be cleared (MUSTs resolved, FIXs tracked) before EXECUTE is authorized.
|
||
> **Artifacts reviewed:** PLAN.md, PROJECT.md, REQUIREMENTS.md, RESEARCH.md (v0.3 section), ARCHITECTURE.md (v0.3 section), ROADMAP.md
|
||
|
||
---
|
||
|
||
## Verdict Legend
|
||
|
||
- **MUST** — blocks execution until fixed. The plan cannot enter EXECUTE with this issue open.
|
||
- **FIX** — fix during execution, non-blocking. Tracked as a P1 condition in VERIFY.
|
||
- **ACCEPT** — proceed as-is. The evidence clears the challenge.
|
||
|
||
---
|
||
|
||
## Axis 1 — Feasibility
|
||
|
||
**Forcing question:** Can this actually be built in 2 execution phases (70 tasks)? Is the scope realistic for one milestone, or is it 2 milestones pretending to be one?
|
||
|
||
**Challenge:** The v0.3 scope spans *seven* independent subsystems (rubric/mastery engine, IRT, scenario library + ≥6 authored scenarios, 6-week path engine, W3C VC 2.0 issuer with Ed25519 + Status List, operator auth + argon2id + slowapi, Postgres-in-LXC + asyncpg, cohort dashboard + k-anonymity aggregation + React UI). This is not a milestone — it is a *program*. The PLAN.md phase-split rationale (lines 14-23) openly admits the scope "is too large for one execution phase" and splits into P1/P2, but both phases ship under the *same* v0.3 milestone tag (v0.1.6). The 70-task count is artificially compressed: SLICE-12 (VC issuer) is 6 tasks for a W3C VC 2.0 + Ed25519 + JCS + Bitstring Status List + public verification endpoint + key rotation — that is *at minimum* a 10-12 task slice on its own, and SLICE-13 (cohort aggregation with k-anonymity + nightly reconciliation + on-session-end hook) is similarly under-tasked at 4 tasks.
|
||
|
||
**Evidence:**
|
||
- PLAN.md:14-23 — "The v0.3 scope … is too large for one execution phase."
|
||
- PLAN.md:32, 334 — P1 = 38 tasks, P2 = 32 tasks, total 70 (excludes P3 review).
|
||
- REQUIREMENTS.md:20-66 — 11 functional REQ-IDs + 9 NFRs = 20 active requirements, the largest single-milestone REQ surface in the project's history (v0.1 was ~14, v0.2 was 16 deploy + 4 NFR).
|
||
- RESEARCH.md:769 (R-VC-01) — "No batteries-included Python VC lib → ~200 LOC custom code" — 200 LOC of custom crypto code is not a 6-task slice; it is a liability that demands more tests than the plan allocates (only 2 test tasks: TASK-12-05, TASK-12-06).
|
||
- SLICE-13 (PLAN.md:516-546) — 4 tasks for: on-session-end hook, k-anonymity suppression SQL, nightly reconciliation cron, and tests. The nightly reconciliation job alone (recompute all 7-day windows from raw events, correct drift, idempotent upsert) is a 2-3 task effort.
|
||
|
||
**Binding verdict: FIX** — The plan *is* feasible as a 2-phase *program*, but only if it is honestly re-labeled. The milestone should ship as v0.3 (P1 mastery core, v0.1.4) and v0.3.1 (P2 operator tier, v0.1.5), with the v0.3 milestone release (v0.1.6) being the *merge* of two separately-shipped, separately-verified patches. Do not pretend P1+P2 is one milestone release. Additionally, re-task SLICE-12 and SLICE-13: add 2 tasks each (one for VC Status List edge cases + key rotation drill, one for reconciliation idempotency + race-condition test). This is non-blocking — the wave structure survives — but the task counts must be honest before EXECUTE.
|
||
|
||
---
|
||
|
||
## Axis 2 — Scope
|
||
|
||
**Forcing question:** Is REQ-DASH-01 (cohort dashboard + multi-tenant + auth) really v0.3, or was it correctly deferred in v0.1/v0.2 for a reason? Does D-031 (override D-007) open a Pandora's box?
|
||
|
||
**Challenge:** REQ-DASH-01 was explicitly deferred in v0.1 (REQUIREMENTS.md:152, "later/deferred") and v0.2. ROADMAP.md:94 places the "Employer / program dashboard" at **v0.8**. The v0.3 plan pulls it forward *three milestones* with the justification that mastery scoring "needs" the operator view. But mastery scoring (REQ-MAST-01/02) and VC issuance (REQ-MAST-03) work *without* a cohort dashboard — the dashboard is an *operator* feature, not a *learner* feature. D-031 overrides D-007 (single-learner/no-auth) and introduces a hybrid SQLite+Postgres topology, operator auth, argon2id, slowapi, asyncpg, a second Docker service, k-anonymity aggregation, and a React operator UI — *none* of which is required for the learner-facing mastery gate to function. This is scope creep dressed as a dependency.
|
||
|
||
**Evidence:**
|
||
- ROADMAP.md:94 — "v0.8 | Employer / program dashboard" (original placement).
|
||
- PROJECT.md:129 (D-031) — "overrides D-007 for the cohort-dashboard surface" — confidence 0.75, the *lowest*-confidence decision that expands scope.
|
||
- REQUIREMENTS.md:44 (REQ-DASH-01) — "Forces multi-tenant + operator auth (D-031)" — the word "forces" is doing a lot of work. The mastery gate (REQ-MAST-02) does not depend on the dashboard.
|
||
- PLAN.md:18 — P1 "works standalone (learner can practice, score, progress) without the operator tier." — *This is an admission that the operator tier is separable.*
|
||
- PROJECT.md:147 (D-049) — failure-injection stays off, further confirming the learner-facing mastery layer is the *real* v0.3 deliverable.
|
||
|
||
**Binding verdict: MUST** — Split the milestone. Ship **v0.3 = P1 only** (mastery core + IRT + scenarios + paths + VC issuance, since VC issuance *is* triggered by the mastery gate and is learner-facing per D-048). Defer **REQ-DASH-01 + REQ-AUTH-01 + REQ-MT-01/02 + REQ-NFR-DASH-01/02 + REQ-NFR-AUTH-01 + REQ-NFR-MT-01** to **v0.4** (operator tier), restoring the original ROADMAP intent. D-031 does open a Pandora's box: every hybrid-DB system eventually faces the "which store is the source of truth?" question, and shipping it under a learner-milestone tag hides that risk. If the team insists on keeping the dashboard in v0.3, rebrand the milestone as "v0.3: Mastery + Operator Tier" and accept that this is a 2-milestone program — but the cleaner answer is to defer the dashboard.
|
||
|
||
---
|
||
|
||
## Axis 3 — Cost
|
||
|
||
**Forcing question:** What is the maintenance cost of Postgres-in-LXC, asyncpg, argon2, pynacl, slowapi, and ~200 LOC custom VC code? Is R-VC-01 (custom VC code) a liability vs using a library?
|
||
|
||
**Challenge:** The v0.3 dependency surface grows by *at least* 5 new pip packages (asyncpg, argon2-cffi, slowapi, pynacl, canonicaljson, base58 — actually 6) plus a Postgres service. Each is a CVE vector, a version-pin maintenance burden, and a CI complexity adder. The ~200 LOC custom VC code (R-VC-01) is the most concerning: cryptographic code written by an AI agent is a *liability* regardless of test coverage. The W3C VC 2.0 + eddsa-jcs-2022 cryptosuite has subtle canonicalization edge cases (e.g., JSON number representation, key ordering, URI normalization) that unit-test round-trips do *not* catch — only interop tests against an independent verifier do, and the plan has *zero* interop tests.
|
||
|
||
**Evidence:**
|
||
- RESEARCH.md:769 (R-VC-01) — "~200 LOC custom code" — confidence 0.75. The mitigation is "unit-test signature/verify round-trip," which only proves the code is self-consistent, not that it is W3C-compliant.
|
||
- PLAN.md:504-513 (TASK-12-05, TASK-12-06) — VC tests are sign/verify round-trip, tamper detection, JCS determinism, status list, revocation, key rotation. *No interop test against an external verifier* (e.g., Verifiable Credential JS verifier, Digital Credentials Verifier).
|
||
- PROJECT.md:140 (D-042) — issuer key encrypted at rest with a root key from secrets. Key management is hand-rolled (init_issuer_key, encrypt, store, rotate). This is a security-engineer task, not a backend task, and the plan assigns it to security-engineer (good), but the *rotation drill* (D-042 "new key + old marked superseded") is not tested end-to-end except in TASK-12-06 which only checks "old VC still verifies against archived public key" — it does *not* test the operational rotation procedure (generate new key, archive old, re-sign new VCs, update verificationMethod URL).
|
||
- RESEARCH.md:772 (R-MT-01) — Postgres-in-LXC resource contention, confidence 0.65 — the *lowest*-confidence technical risk. Memory bump to 6GB is a guess, not a measurement.
|
||
|
||
**Binding verdict: MUST** — Two conditions before EXECUTE:
|
||
1. **Add a VC interop test** (TASK-12-07): verify a Praxis-issued VC against at least one *external* W3C VC verifier (e.g., the `digitalbazaar/vc-verifier` or a JS `@digitalcredentials/vc` verifier). Round-trip self-verification is insufficient for cryptographic claims. Without this, R-VC-01 is an unmitigated liability.
|
||
2. **Add a key-rotation operational test** (TASK-12-08): end-to-end drill — issue N VCs with key A, rotate to key B, issue M VCs with key B, verify all N+M VCs still verify (N against archived key A, M against active key B), revoke one of each, verify revocation. This is the *one* crypto procedure that, if broken, silently invalidates every credential ever issued.
|
||
|
||
The Postgres/argon2/slowapi maintenance cost is **ACCEPT** — these are well-maintained, widely-used libraries. The liability is concentrated in the custom VC code.
|
||
|
||
---
|
||
|
||
## Axis 4 — Technical Risk
|
||
|
||
**Forcing question:** R-MAST-01 (N=3 thin for credential), R-AUTH-01 (Secure cookie + no TLS), R-MAST-02 (LLM hallucinated quotes), R-IRT-01 (cold start) — which are MUST-FIX before execution vs ACCEPT?
|
||
|
||
**Challenge:** The plan treats all four as "Open Questions Deferred to EXECUTE" (PLAN.md:687-693). That is insufficient. R-MAST-01 is a *credibility* risk: if the VC is labeled as a mastery credential and employers treat it as high-stakes, N=3 with G≈0.5-0.6 is defensible only if the credential is explicitly labeled *formative*. R-AUTH-01 is a *security* risk: relaxing the Secure cookie flag for a no-TLS pilot means session cookies travel in cleartext — if the operator bridge IP is on a shared network (vmbr0 DHCP), any host on the bridge can sniff the operator session. R-MAST-02 is the *highest*-confidence mitigation (fuzzy-match quotes), but the plan's fallback ("empty evidence + log warning") means a session could silently score as a *zero* with no learner-visible signal. R-IRT-01 is benign (cold-start fallback to fixed difficulty).
|
||
|
||
**Evidence:**
|
||
- RESEARCH.md:766 (R-MAST-01) — confidence 0.62, *below* the 0.70 decision threshold. Mitigation: "Label v0.3 VC as formative." This label is *not* in the PLAN.md VC payload (TASK-12-02) or the REQ-MAST-03 requirement text.
|
||
- RESEARCH.md:771 (R-AUTH-01) — "Secure cookie flag fails without TLS." PLAN.md:450 resolves this with `PRAXIS_COOKIE_SECURE=false` env default. This ships a known-insecure default.
|
||
- PLAN.md:154 (TASK-03-01) — "on final failure, fall back to empty evidence + log warning." Empty evidence → rule scorer has no signals → every criterion scores level 1 (fail) → scenario fails → learner sees a failed session with *no explanation*. This is a UX and fairness bug.
|
||
- RESEARCH.md:774 (R-IRT-01) — mitigation confidence 0.75, "fall back to scenario.difficulty until ≥5 observations." ACCEPT.
|
||
|
||
**Binding verdict: MUST** — Three conditions:
|
||
1. **R-MAST-01**: Add `credentialTier: "formative"` (or equivalent) to the VC payload (TASK-12-02) and to the verification endpoint response (TASK-12-04). Update REQ-MAST-03 to require this label. Without it, the credential is misleading.
|
||
2. **R-AUTH-01**: Do *not* ship `PRAXIS_COOKIE_SECURE=false` as a default. Either (a) require TLS for the operator surface (add a Traefik sidecar or Caddy in front of `/api/operator/*`), or (b) bind the operator surface to `127.0.0.1` only (loopback) so cookies never traverse the bridge. A cleartext cookie on a shared bridge is a MUST-FIX.
|
||
3. **R-MAST-02**: Change the fallback in TASK-03-01 from "empty evidence + log warning" to "empty evidence → mark scenario as `scoring_inconclusive` → do not count toward gate, do not penalize learner, surface 'technical issue, please retry' in the debrief." A silent fail-to-zero is unacceptable.
|
||
|
||
R-IRT-01: **ACCEPT** — cold-start fallback is sound.
|
||
|
||
---
|
||
|
||
## Axis 5 — Requirements Coverage
|
||
|
||
**Forcing question:** Does the plan actually cover all 20 REQ-IDs, or are some hand-waved? Check the coverage matrix in PLAN.md against REQUIREMENTS.md.
|
||
|
||
**Challenge:** The PLAN.md coverage matrix (lines 657-683) claims "20 REQ-IDs covered, 0 partial, 0 deferred." Let me audit the suspicious ones.
|
||
|
||
**Evidence (audit):**
|
||
|
||
| REQ-ID | Claimed coverage | Actual coverage | Verdict |
|
||
|--------|-----------------|-----------------|---------|
|
||
| REQ-MAST-03 | P2 SLICE-12 "VC issuer" | SLICE-12 implements issuance + verification + revocation. But REQ-MAST-03 says "Issued when a mastery gate opens" — the *trigger* is in P1 SLICE-07 (TASK-07-01, "path_engine.check_gate + advance_week") and the *issuance* is in P2 SLICE-12. The P1→P2 handoff for VC issuance is not in any task — who calls `issuer.issue_credential()` when the week-final gate opens? TASK-07-01 says "(6) record mastery_gate_event" but does NOT call the VC issuer (VC issuer is P2). D-048 says "issue VC if week-final gate" but the plan splits the gate-open (P1) from the issuance (P2). **Gap: no task wires the P1 gate-open event to the P2 VC issuer.** | **FIX** — add a task (either in SLICE-07 or SLICE-12) that defines the P1→P2 VC-issuance contract: a `mastery_gate_events` row with `gate_opened_at` is the trigger; P2's VC issuer polls/receives this event and issues. |
|
||
| REQ-NFR-MAST-02 | P1+P2 SLICE-07, 09 "gate auditability (SQLite + Postgres)" | SLICE-07 records the event in SQLite; SLICE-09 defines the Postgres `mastery_gate_events` table; but *no task mirrors* the SQLite event to Postgres. The "mirror" is implied but not tasked. | **FIX** — TASK-13-01 (cohort aggregation hook) should explicitly mirror `mastery_gate_events` from SQLite to Postgres, or add a dedicated mirroring task. |
|
||
| REQ-SCEN-04 | P1 SLICE-02, 06 "expert-authored format + AI variation hooks" | SLICE-02 adds `generated_from` and `intent_hash` fields (the hook). SLICE-06 authors expert scenarios. But *no task implements the AI-variation review pipeline* (`_pending/` dir → expert review → library promotion). RESEARCH.md:758 describes it; PLAN.md does not task it. | **ACCEPT** — REQ-SCEN-04 says "AI-generated variations" with "expert review" — the *hook* is the schema field; the *pipeline* can be deferred. The plan is honest that AI variations are "in P2 or later" (SLICE-06 goal line 246). |
|
||
| REQ-NFR-DASH-02 | P2 SLICE-13 "freshness ≤24h" | SLICE-13 has a nightly reconciliation job (TASK-13-03) at 02:00. If the on-session-end hook (TASK-13-01) fails or lags, freshness depends on the nightly job. ≤24h is satisfied *if* the nightly job runs. But there is no task for *monitoring* or *alerting* on job failure. | **FIX** — add a health check for the nightly job (log last-run timestamp, surface in operator dashboard or `/health`). Non-blocking. |
|
||
| REQ-NFR-VC-02 | P2 SLICE-12 "revocation latency — within 1 sync of status list" | "1 sync" is undefined. Is it 1 sync of the status list blob? Is the status list in-memory or fetched on every verify? TASK-12-04 (verification endpoint) does not specify caching of the status list. | **FIX** — clarify in TASK-12-04: status list is fetched from Postgres on every verification (no cache), so revocation latency = next verify call. Non-blocking. |
|
||
|
||
**Binding verdict: FIX** — The coverage matrix is *mostly* honest (18/20 fully covered), but the P1→P2 VC-issuance wiring gap (REQ-MAST-03) is a real hole — without a task that defines the trigger contract, the VC issuer will be built but never called. Add the wiring task. The other three FIXs are minor clarifications.
|
||
|
||
---
|
||
|
||
## Axis 6 — Architecture
|
||
|
||
**Forcing question:** Is hybrid SQLite + Postgres (D-031) a maintainable pattern or a future migration nightmare? Is "no cross-DB joins" realistic for the cohort dashboard queries?
|
||
|
||
**Challenge:** Hybrid polyglot persistence is a *known* anti-pattern when the two stores hold related data and there is no canonical source of truth. Here, `mastery_gate_events` exists in *both* SQLite (P1, the learner's local record) and Postgres (P2, the operator audit log). Which is canonical? If they diverge (e.g., SQLite write succeeds, Postgres mirror fails due to pool exhaustion), the cohort dashboard shows *stale* data while the learner sees *correct* data — and there is no reconciliation except the nightly job (which recomputes from Postgres `mastery_gate_events`, not from SQLite). This means the nightly job recomputes from a *possibly-incomplete* Postgres copy. The "no cross-DB joins" rule is realistic *only* if the cohort dashboard never needs to join learner-local data (e.g., θ distribution by path) with operator data — but the dashboard's "progression" and "failure patterns" views implicitly need *both* the learner's session outcomes (SQLite) and the operator's aggregate view (Postgres). The plan resolves this by aggregating at session-end (writing the aggregate to Postgres), so the dashboard reads *only* Postgres — but this means the aggregate is a *derived* copy, and the "no cross-DB joins" rule is maintained by *duplicating data*, not by query-time joins. This is workable but fragile.
|
||
|
||
**Evidence:**
|
||
- ARCHITECTURE.md:356-379 — "The two stores never share a session and never join via cross-DB FKs (`learner_ref` is an opaque string in Postgres)." — the design is clean *if* the mirror is reliable.
|
||
- RESEARCH.md:746 — "No cross-DB joins via `learner_ref` — `learner_ref` is an opaque string, not a FK." — correct, but `learner_ref` is still a *logical* join key. If the SQLite learner is deleted and re-created, the Postgres `learner_ref` dangles.
|
||
- PLAN.md:528-529 (TASK-13-01) — "after P1's mastery hooks fire, call `aggregator.upsert_aggregate(...)`" — this is a *synchronous* call after the SQLite write, in the session-end path. If Postgres is down, does the session-end fail? The plan does not specify failure semantics.
|
||
- RESEARCH.md:748 — migration strategy: "SQLite volume untouched → learner path never regresses." Good, but the *operator* path regresses if Postgres is down.
|
||
|
||
**Binding verdict: FIX** — Three conditions:
|
||
1. **Define failure semantics for the Postgres mirror** (in TASK-13-01): if `upsert_aggregate` fails (Postgres down, pool exhausted), the learner session-end must *still succeed* (SQLite write is canonical for the learner). The aggregate failure is logged and reconciled by the nightly job. This makes SQLite the *learner-canonical* store and Postgres the *operator-derived* store — state this explicitly in ARCHITECTURE.md.
|
||
2. **Make `learner_ref` a stable, opaque, non-reusable identifier** (e.g., a UUID generated once and stored in SQLite, never reused). Add this to TASK-02-01 or a new task. Without it, the "no FK" rule is a leaky abstraction.
|
||
3. **Add a Postgres-readiness guard to the operator API**: if Postgres is down, `/api/operator/cohort/*` returns 503 (not 500 with a stack trace). Add to TASK-14-02.
|
||
|
||
The hybrid pattern is **ACCEPT** *with* these conditions — it is the correct pilot choice (don't migrate learner state to Postgres prematurely), but the failure semantics must be explicit.
|
||
|
||
---
|
||
|
||
## Axis 7 — Testing
|
||
|
||
**Forcing question:** 70 tasks, but how many have tests? Is the test strategy (mocked LLM for evidence extraction, testcontainers for Postgres) viable, or are there untestable critical paths?
|
||
|
||
**Challenge:** Let me count test tasks across the plan.
|
||
|
||
**Evidence (test task audit):**
|
||
|
||
| Slice | Tasks | Test tasks | Test ratio |
|
||
|-------|-------|-----------|------------|
|
||
| SLICE-01 | 4 | 1 (TASK-01-04) | 25% |
|
||
| SLICE-02 | 4 | 1 (TASK-02-04) | 25% |
|
||
| SLICE-03 | 5 | 2 (TASK-03-04, 03-05) | 40% |
|
||
| SLICE-04 | 4 | 2 (TASK-04-03, 04-04) | 50% |
|
||
| SLICE-05 | 4 | 1 (TASK-05-04) | 25% |
|
||
| SLICE-06 | 3 | 1 (TASK-06-03) | 33% |
|
||
| SLICE-07 | 4 | 2 (TASK-07-03, 07-04) | 50% |
|
||
| SLICE-08 | 3 | 3 (all test/verification) | 100% |
|
||
| SLICE-09 | 4 | 1 (TASK-09-04) | 25% |
|
||
| SLICE-10 | 3 | 0 | 0% — infra, acceptable |
|
||
| SLICE-11 | 5 | 2 (TASK-11-04, 11-05) | 40% |
|
||
| SLICE-12 | 6 | 2 (TASK-12-05, 12-06) | 33% — *too low for crypto code* |
|
||
| SLICE-13 | 4 | 1 (TASK-13-04) | 25% — *too low for k-anonymity* |
|
||
| SLICE-14 | 4 | 1 (TASK-14-04) | 25% |
|
||
| SLICE-15 | 4 | 1 (TASK-15-04) | 25% |
|
||
| SLICE-16 | 3 | 3 (all integration/verification) | 100% |
|
||
| **Total** | **70** | **24** | **34%** |
|
||
|
||
**Critical untestable paths:**
|
||
1. **LLM evidence extraction (TASK-03-01)** — the plan mocks the LLM (good for unit tests), but there is *no* test that runs against the *real* LLM with a real transcript. A mocked LLM proves the scoring logic, not that the extraction prompt works. This is a *fundamentally untestable in CI* path — the only test is manual/staging.
|
||
2. **k-anonymity suppression (TASK-13-02)** — the test (TASK-13-04) checks "cell with 9 learners → suppressed, 10 → shown." But it does *not* test the differencing attack (comparing two adjacent windows to re-identify a learner who appears in one but not the other). RESEARCH.md:754 says "limit to pre-defined 2-D views to block differencing attacks" — but there is no test that the API *enforces* only pre-defined views (i.e., that an operator cannot request an arbitrary `path × week × outcome` 3-D view).
|
||
3. **Nightly reconciliation (TASK-13-03)** — no test for "reconciliation corrects drift." TASK-13-04 tests "reconciliation correctness" but not *drift correction* (insert a bad aggregate, run reconcile, verify it's fixed).
|
||
4. **VC verification endpoint (TASK-12-04)** — tested via TASK-12-06, but only with Praxis-issued VCs. No interop test (see Axis 3).
|
||
|
||
**Binding verdict: FIX** — Four conditions:
|
||
1. **Add a real-LLM smoke test** (in SLICE-08 or SLICE-16): run one session transcript through the *actual* deepseek-v4-flash:cloud evidence extractor and verify the output is valid JSON with fuzzy-matching quotes. This runs only in staging (requires OLLAMA_API_KEY), gated by an env flag. The mocked-LLM tests stay in CI.
|
||
2. **Add a k-anonymity differencing-attack test** (TASK-13-04 extension): verify that the cohort API rejects arbitrary 3-D view requests, and that two adjacent 7-day windows cannot re-identify a single learner appearing in only one.
|
||
3. **Add a reconciliation drift-correction test** (TASK-13-04 extension): insert a deliberately-wrong aggregate, run `reconcile_cohort()`, verify it's corrected.
|
||
4. **Add the VC interop test** (per Axis 3, MUST condition).
|
||
|
||
The mocked-LLM + testcontainers strategy is **ACCEPT** *for CI*. The gaps are in *integration* and *security* testing, not unit testing.
|
||
|
||
---
|
||
|
||
## Axis 8 — Phase Split
|
||
|
||
**Forcing question:** Is P1/P2 the right split? Should VC issuance (P2 SLICE-12) be in P1 with mastery gates (P1 SLICE-07) since they trigger on the same event? Is the P1→P2 dependency clean?
|
||
|
||
**Challenge:** The plan splits VC issuance (P2) from mastery-gate-open (P1) even though D-048 says "issue VC if week-final gate." This means P1 ships (v0.1.4) with mastery gates that open but *no credential is issued* — the learner reaches mastery and gets... nothing portable. The VC issuer arrives in P2 (v0.1.5). This is a *user-visible gap*: a learner who completes the path in v0.1.4 has no credential. The plan's phase-split rationale (lines 14-23) says P1 "works standalone (learner can practice, score, progress) without the operator tier" — but VC issuance is *not* the operator tier; it is a learner-facing consequence of mastery (D-048). VC issuance should be in P1.
|
||
|
||
Conversely, the operator auth + Postgres + cohort dashboard is correctly P2 — those are operator-tier.
|
||
|
||
**Evidence:**
|
||
- PROJECT.md:146 (D-048) — "issue VC if week-final gate" — VC issuance is a *mastery-gate consequence*, not an operator feature.
|
||
- PLAN.md:18 — "P1 works standalone" — but "standalone" here silently drops the VC, which is a REQ-MAST-03 requirement.
|
||
- PLAN.md:663 (coverage matrix) — REQ-MAST-03 is listed as P2 SLICE-12. But REQ-MAST-03 is a *mastery* requirement, not an *operator* requirement.
|
||
- PLAN.md:282 (TASK-07-01) — P1 session-end hook does steps 1-6 but step 6 is "record mastery_gate_event" — no VC issuance call. The VC issuance is orphaned in P2 with no trigger from P1.
|
||
|
||
**Binding verdict: MUST** — Move SLICE-12 (VC issuer) to **P1**, *after* SLICE-07 (mastery gates), as a new Wave-4 slice in P1 (parallel with SLICE-08). This requires:
|
||
1. VC issuer needs `issuer_keys` storage — use *SQLite* for P1 (the issuer_keys table moves to SQLite for v0.3; Postgres takes over in v0.4 when the operator tier arrives). Or, if Postgres is required for VC, then Postgres must also move to P1 — which inflates P1 further and reinforces the Axis 2 verdict (split the milestone).
|
||
2. The cleaner resolution: **defer VC issuance to v0.3.1 (P2)** *and* accept that v0.1.4 (P1) ships mastery gates without credentials — but *label this explicitly* in the P1 ship notes ("VC issuance in v0.1.5"). Do not claim REQ-MAST-03 is covered in P1.
|
||
|
||
Either resolution is acceptable. The *current* plan — which implies VC issuance is triggered by P1's gate-open but tasks it in P2 with no wiring — is **not acceptable**. Pick one: (a) VC in P1 with SQLite-backed issuer keys, or (b) VC explicitly deferred to P2 with P1 shipping "mastery gates, no credential yet."
|
||
|
||
The P1→P2 dependency is otherwise clean (P2 reads P1's `mastery_gate_events` and session outcomes). **ACCEPT** on the dependency structure.
|
||
|
||
---
|
||
|
||
## Axis 9 — Decisions
|
||
|
||
**Forcing question:** Are D-031..D-049 (12 clarify + 7 specify decisions) well-grounded, or are any below the 0.60 confidence threshold? Is D-049 (failure-injection stays off) a mistake given mastery scoring scores recovery from failure branches?
|
||
|
||
**Challenge:** Let me audit confidences against the 0.70 threshold (the project's apparent decision-acceptance floor).
|
||
|
||
**Evidence (confidence audit of D-031..D-049):**
|
||
|
||
| ID | Confidence | Below 0.70? | Verdict |
|
||
|----|------------|-------------|---------|
|
||
| D-031 | 0.75 | No | ACCEPT — but see Axis 2 (scope creep). |
|
||
| D-032 | 0.70 | At threshold | ACCEPT — N=3 is formative-only per R-MAST-01. |
|
||
| D-033 | 0.70 | At threshold | ACCEPT — W3C VC 2.0 is a stable standard. |
|
||
| D-034 | 0.70 | At threshold | ACCEPT — k=10 is the conventional minimum. |
|
||
| D-035 | 0.70 | At threshold | ACCEPT — 1PL/Rasch is the simplest IRT. |
|
||
| D-036 | 0.80 | No | ACCEPT. |
|
||
| D-037 | 0.75 | No | ACCEPT. |
|
||
| D-038 | 0.80 | No | ACCEPT — deterministic scoring is the right call. |
|
||
| D-039 | 0.80 | No | ACCEPT. |
|
||
| D-040 | 0.80 | No | ACCEPT. |
|
||
| D-041 | 0.75 | No | ACCEPT — but R-AUTH-01 (Secure cookie) is a MUST-FIX (Axis 4). |
|
||
| D-042 | 0.70 | At threshold | ACCEPT — but key rotation drill is a MUST (Axis 3). |
|
||
| D-043 | 0.80 | No | ACCEPT. |
|
||
| D-044 | 0.75 | No | ACCEPT. |
|
||
| D-045 | 0.70 | At threshold | ACCEPT — but failure semantics are a FIX (Axis 6). |
|
||
| D-046 | 0.80 | No | ACCEPT — θ in SQLite is correct. |
|
||
| D-047 | 0.70 | At threshold | ACCEPT — 6 scenarios is tight but defensible for formative. |
|
||
| D-048 | 0.75 | No | ACCEPT — but the P1/P2 split breaks the trigger wiring (Axis 8). |
|
||
| D-049 | 0.80 | No | See below. |
|
||
|
||
**D-049 (failure-injection stays off):** The challenge is whether this is a mistake. The rubric (SLICE-01) has a "de-escalation" criterion (weight 0.20), and RESEARCH.md:718 says "de-escalation up-weights to ~0.40 if the escalate branch triggers." The `escalate` branch is a *naturally-occurring* failure branch in `cs_refund_ca_v01` (D-010), not an AI-provoked failure. So mastery scoring *does* score recovery from a failure branch — the *naturally-occurring* one. D-049 keeps AI-provoked failure injection off, which is correct: the rubric's de-escalation criterion is exercised by the existing branch, and adding AI-provoked failures would couple mastery scoring to a new feature (scope creep). D-049 is well-grounded.
|
||
|
||
**However**, there is a subtle gap: the rubric weights are *static* in the YAML (empathy 0.35, resolution 0.30, de-escalation 0.20, professionalism 0.15 per TASK-01-01). RESEARCH says de-escalation "up-weights to ~0.40 if the escalate branch triggers" — but TASK-01-01 does not mention dynamic re-weighting based on branch outcome. Either the weights are static (and the "up-weight" is a future feature) or they are dynamic (and the plan is missing a task). This is a **FIX** — clarify in TASK-01-01 whether weights are static or branch-dependent. If static, update RESEARCH.md to note the up-weight is deferred.
|
||
|
||
**Binding verdict: ACCEPT** on all D-031..D-049 confidences (none below 0.60; the floor is 0.70, which is the project's threshold). **FIX** on the de-escalation weight ambiguity (static vs dynamic) in TASK-01-01. D-049 is **ACCEPT** — failure-injection stays off is the correct call; the naturally-occurring `escalate` branch exercises the de-escalation criterion.
|
||
|
||
---
|
||
|
||
## Summary Table
|
||
|
||
| # | Axis | Forcing question (short) | Verdict |
|
||
|---|------|---------------------------|---------|
|
||
| 1 | Feasibility | 2 phases / 70 tasks realistic? | **FIX** — re-label as 2-milestone program; re-task SLICE-12/13 (+2 tasks each) |
|
||
| 2 | Scope | REQ-DASH-01 really v0.3? D-031 Pandora's box? | **MUST** — split milestone; defer dashboard to v0.4 (or rebrand honestly) |
|
||
| 3 | Cost | Custom VC code liability? Maintenance burden? | **MUST** — add VC interop test + key-rotation operational test before EXECUTE |
|
||
| 4 | Technical risk | R-MAST-01/R-AUTH-01/R-MAST-02/R-IRT-01 | **MUST** — label VC formative; fix Secure cookie; fix silent-fail-to-zero fallback |
|
||
| 5 | Requirements coverage | 20 REQ-IDs fully covered? | **FIX** — wire P1→P2 VC-issuance trigger; mirror SQLite→Postgres gate events; minor NFR clarifications |
|
||
| 6 | Architecture | Hybrid SQLite+Postgres maintainable? | **FIX** — define Postgres-failure semantics; stabilize learner_ref; add 503 guard |
|
||
| 7 | Testing | Test strategy viable? Untestable paths? | **FIX** — add real-LLM smoke test, differencing-attack test, drift-correction test, VC interop test |
|
||
| 8 | Phase split | VC issuance in P2 but triggers on P1 event? | **MUST** — move VC to P1 (SQLite-backed) OR explicitly defer to P2 with honest labeling |
|
||
| 9 | Decisions | D-031..D-049 below 0.60? D-049 a mistake? | **ACCEPT** — all confidences ≥0.70; D-049 correct; FIX de-escalation weight ambiguity |
|
||
|
||
---
|
||
|
||
## Final Recommendation: **GO-WITH-CONDITIONS**
|
||
|
||
The v0.3 plan is **not approved for EXECUTE as-is**. It is a well-researched, well-structured plan that suffers from two structural flaws: (1) it is two milestones pretending to be one, and (2) it splits a learner-facing consequence (VC issuance) from its trigger (mastery gate) across a phase boundary without wiring.
|
||
|
||
### MUST conditions (blocking — must be resolved in PLAN before EXECUTE):
|
||
|
||
1. **Axis 2 — Split the milestone.** Either (a) defer REQ-DASH-01 + operator tier to v0.4, shipping v0.3 = P1 + VC issuance only; or (b) rebrand v0.3 as a 2-milestone program (v0.3 + v0.3.1) with separate ship/verify cycles. Do not ship P1+P2 under one milestone tag.
|
||
|
||
2. **Axis 3 — Add VC interop test + key-rotation operational test.** Custom crypto code without interop verification is an unmitigated liability. Add TASK-12-07 (interop) and TASK-12-08 (rotation drill).
|
||
|
||
3. **Axis 4 — Fix three technical risks.** (a) Label VC as `formative` in payload + verification response + REQ-MAST-03 text. (b) Do not ship `PRAXIS_COOKIE_SECURE=false` as default — use TLS or loopback-binding for the operator surface. (c) Change evidence-extraction fallback from silent-fail-to-zero to `scoring_inconclusive` with learner-visible retry signal.
|
||
|
||
4. **Axis 8 — Resolve the VC-issuance phase split.** Either move SLICE-12 to P1 (with SQLite-backed issuer keys) or explicitly defer REQ-MAST-03 to P2 and label P1 as "mastery gates, no credential yet." The current plan's implicit wiring is a gap.
|
||
|
||
### FIX conditions (non-blocking — tracked in VERIFY-P1/P2):
|
||
|
||
5. **Axis 1 — Re-task SLICE-12 and SLICE-13.** Add 2 tasks each to honestly reflect the effort (VC edge cases + rotation drill; reconciliation idempotency + race test).
|
||
6. **Axis 5 — Wire the P1→P2 VC-issuance trigger** (if VC stays in P2) and **mirror SQLite→Postgres gate events** explicitly in TASK-13-01.
|
||
7. **Axis 6 — Define Postgres-failure semantics** (SQLite is learner-canonical, Postgres is operator-derived); stabilize `learner_ref` as a non-reusable UUID; add 503 guard on operator API.
|
||
8. **Axis 7 — Add four tests**: real-LLM smoke (staging-gated), k-anonymity differencing-attack, reconciliation drift-correction, VC interop (already a MUST).
|
||
9. **Axis 9 — Clarify de-escalation weight** (static vs branch-dependent) in TASK-01-01.
|
||
|
||
### ACCEPT items (proceed as-is):
|
||
|
||
- IRT 1PL/Rasch cold-start fallback (R-IRT-01).
|
||
- All decision confidences (D-031..D-049 ≥ 0.70, none below 0.60).
|
||
- D-049 (failure-injection stays off) — correct call.
|
||
- Mocked-LLM + testcontainers CI strategy.
|
||
- Hybrid SQLite+Postgres topology (with failure-semantics FIX).
|
||
- P1→P2 dependency structure (clean except for VC-issuance wiring).
|
||
|
||
### Bottom line:
|
||
|
||
The plan is **not unfeasible** — the research is thorough, the architecture is sound, and the slice decomposition is reasonable. But it is **over-scoped** (two milestones in one tag) and **under-tested** in its highest-risk areas (custom crypto, k-anonymity, real-LLM extraction). Resolve the 4 MUST conditions, track the 5 FIX conditions, and this becomes a **GO**. |