---ci--- phase: 7 milestone: v0.2 status: complete requirements: covered: [REQ-2-001, REQ-2-002, REQ-2-003, REQ-2-004, REQ-2-005, REQ-2-006, REQ-2-007, REQ-2-008, REQ-2-009, REQ-2-010, REQ-2-011, REQ-2-012] partial: [] ---/ci--- Milestone v0.2 (ai-tutor-architecture) merged to main. Escalation record (audit remediation, durable): P1 executor delegation failed twice (empty subagent results, zero files created); auto-resolved at full autonomy to inline execution with identical plan fidelity (commit 3271373, reflog-only after phase branch squash-delete).
35 KiB
Nextcraft v0.2 — PLAN.md
Overview
This plan covers execution phases 1-6 of milestone v0.2 (AI Tutor Architecture): the six AI tutor agents (Coach, Tutor, Lab, Assessor, Proctor, Mentor) as real LLM-backed services in a new apps/ai-service Python FastAPI application, wired into the existing v0.1 learner surface with streaming responses. Phases are strictly sequential (P1→P6); within each phase, Wave 1 tasks are parallelizable (no cross-file dependencies) and later waves depend on earlier ones.
Environment facts (apply throughout): Python via python3 -m venv (no system pip, no uv); pnpm 12.3.4 via corepack; ai-service port 8420; default model gemma4:31b (config via AI_TUTOR_MODEL); ollama-cloud base https://ollama.com/v1 (OpenAI-compatible, Bearer auth); API keys live only in gitignored .ciagent/.env.secrets, exported by scripts/dev.sh — never in code, commits, or logs; all automated tests use the deterministic mock provider and never call the cloud.
| Phase | Name | Requirements | Waves | Personas |
|---|---|---|---|---|
| 1 | AI service scaffolding | REQ-2-001, 002, 003 | 3 | backend-engineer, ai-engineer |
| 2 | Agent framework | REQ-2-004 | 2 | ai-engineer, backend-engineer |
| 3 | Coach + Tutor agents | REQ-2-005, 006 | 3 | ai-engineer, backend-engineer |
| 4 | Lab + Assessor agents | REQ-2-007, 008 | 4 | ai-engineer, backend-engineer, data-engineer |
| 5 | Proctor + Mentor agents | REQ-2-009, 010 | 3 | ai-engineer, backend-engineer |
| 6 | Learner surface integration | REQ-2-011, 012 | 3 | frontend-engineer, design-system-engineer |
Phase 1: AI Service Scaffolding
Requirements: REQ-2-001, REQ-2-002, REQ-2-003 Goal: apps/ai-service runs under uvicorn, /health responds, provider-agnostic LLM layer with ollama-cloud/local/mock providers, SSE chat streaming verified, pytest suite green with mock provider, turbo integration wired
Wave 1: Service shell + LLM core (parallel — no shared files)
Task 1-1-01: FastAPI app scaffolding
- Persona: backend-engineer — REQ: REQ-2-001
- Files:
apps/ai-service/pyproject.toml,apps/ai-service/ai_service/main.py,apps/ai-service/ai_service/config.py,apps/ai-service/ai_service/__init__.py,apps/ai-service/tests/conftest.py,apps/ai-service/tests/test_health.py,apps/ai-service/.env.example,apps/ai-service/README.md - Action: pydantic-settings
Settings(env_prefixAI_, env_file,SecretStrkey, port 8420, provider select,AI_TUTOR_MODELdefaultgemma4:31b). FastAPI app factory inmain.pywith lifespan stub (httpx client pool comes in Wave 2), CORS localhost-only (A-008),GET /health. pyproject with pinned deps (fastapi, uvicorn, pydantic, pydantic-settings, httpx, sse-starlette, pytest, pytest-asyncio) and dev extra. conftest: settings override + TestClient fixture..env.exampledocuments allAI_*vars; README documents venv setup and dev workflow. - Verify:
scripts/bootstrap.sh && scripts/test.sh— test_health passes;curl localhost:8420/healthreturns 200
Task 1-1-02: LLM types, protocol, mock provider
- Persona: ai-engineer — REQ: REQ-2-002
- Files:
apps/ai-service/ai_service/llm/types.py,apps/ai-service/ai_service/llm/base.py,apps/ai-service/ai_service/llm/mock.py,apps/ai-service/ai_service/llm/__init__.py - Action:
types.py: pydanticMessage(role/content),ChatDelta(OpenAI-compatible chunk shape).base.py:LLMProviderprotocol — asyncstream_chat(messages, model, response_format=None) -> AsyncIterator[ChatDelta]; the provider is a dumb pipe, no envelope logic (D-016 keeps envelope in API layer).mock.py: deterministic scripted provider (hash-seeded token streams, scripted failure modes: connect error, mid-stream error, malformed JSON) for tests and CI. - Verify: mock provider importable and deterministic; two identical calls yield identical streams
Task 1-1-03: Monorepo integration (shim + turbo + scripts)
- Persona: backend-engineer — REQ: REQ-2-001
- Files:
apps/ai-service/package.json,apps/ai-service/scripts/bootstrap.sh,apps/ai-service/scripts/dev.sh,apps/ai-service/scripts/test.sh,turbo.json(update),package.json(root, update) - Action: zero-dependency shim
package.jsonin apps/ai-service withdev/test/bootstrapscript entries. Turbo passthrough tasksai#dev,ai#test,ai#bootstrap(cache: false,outputs: []). Root scriptsai:dev,ai:test,ai:bootstrap.bootstrap.sh: idempotentpython3 -m venv .venv+ pip install.dev.sh: exports keys from.ciagent/.env.secrets→ uvicorn on 8420.test.sh: pytest via venv. - Verify:
corepack pnpm install && pnpm ai:bootstrap && pnpm ai:testruns pytest through turbo; re-running bootstrap is a no-op
Task 1-1-04: Python lint (ruff, check-only)
- Persona: backend-engineer — REQ: REQ-2-001
- Files:
apps/ai-service/pyproject.toml(update:[tool.ruff]config +ruffin dev extra),apps/ai-service/scripts/lint.sh,package.json(root, update),turbo.json(update) - Action: Add
ruff(check-only, no formatter) to the dev extra;[tool.ruff]with line-length 100, target py311.scripts/lint.sh:.venv/bin/ruff check .with repo-root path resolution via script-relative dirname (not CWD). Root scriptai:lint, turbo passthroughai#lint(cache: false,outputs: []). Run over the entire ai-service tree; fix all findings before phase ship (G-3). - Verify:
pnpm ai:lintexits 0 on the Phase 1 codebase
Wave 2: Real providers + SSE endpoint (depends on Wave 1)
Task 1-2-01: OpenAI-compatible provider + factory
- Persona: ai-engineer — REQ: REQ-2-002
- Files:
apps/ai-service/ai_service/llm/openai_compat.py,apps/ai-service/ai_service/llm/factory.py - Action:
openai_compat.py: singleOpenAICompatProviderfor ollama-cloud (https://ollama.com/v1, Bearer) and local endpoints (base URL from settings); raw httpx against/v1/chat/completionswithstream: true, byte-identical delta passthrough, tolerant of ollama-cloud quirks. Uses the lifespan-managedhttpx.AsyncClient(10s connect / 300s read, D-017) — no openai SDK.factory.py: select provider from settings (ollama-cloud|local|mock). - Verify: provider constructs from settings for all 3 names; manual probe against ollama-cloud streams tokens (documented in README, not a test)
Task 1-2-02: Lifespan wiring + SSE chat endpoint
- Persona: backend-engineer — REQ: REQ-2-003
- Files:
apps/ai-service/ai_service/main.py(update),apps/ai-service/ai_service/api/deps.py,apps/ai-service/ai_service/api/chat.py,apps/ai-service/ai_service/api/__init__.py - Action: Lifespan creates the shared
httpx.AsyncClientand provider factory; deps.py provides provider via DI.POST /v1/chat/streamin chat.py implements the D-016 envelope:metaevent (agent/session/model) flushed before first token → raw OpenAI chunks passed through asdata: {json}→doneevent →errorevent before[DONE]on mid-stream failure; pre-first-byte failures return proper HTTP status codes. HeadersCache-Control: no-cache,X-Accel-Buffering: no; sse-starlette ping keep-alive. - Verify:
curl -N -X POST localhost:8420/v1/chat/streamwith mock provider shows meta event, token deltas, done,[DONE]
Wave 3: Provider + endpoint test suites (depends on Wave 2)
Task 1-3-01: LLM provider tests (byte-exact, no cloud)
- Persona: ai-engineer — REQ: REQ-2-002
- Files:
apps/ai-service/tests/llm/test_openai_compat.py,apps/ai-service/tests/llm/test_mock.py,apps/ai-service/tests/llm/__init__.py - Action: httpx
MockTransporttests parsing byte-exact fixture streams (happy path, empty delta,[DONE], malformed line, mid-stream disconnect). Mock provider tests: determinism, scripted failure modes, response_format echo. - Verify:
pnpm ai:test— llm suite green; zero network calls in tests
Task 1-3-02: SSE stream endpoint tests
- Persona: backend-engineer — REQ: REQ-2-003
- Files:
apps/ai-service/tests/api/test_chat_stream.py,apps/ai-service/tests/api/__init__.py - Action: TestClient
client.stream()tests: meta-first ordering, delta passthrough, done +[DONE]sentinel, mid-stream error event, pre-first-byte failure → HTTP status, required headers. pytest-asyncio auto mode (D-023). - Verify:
pnpm ai:test— api suite green
Must-Haves (Phase 1)
scripts/bootstrap.shis idempotent; creates venv + installs deps without system pippnpm ai:devstarts uvicorn;curl localhost:8420/healthreturns 200pnpm ai:lintexits 0 (ruff check over the ai-service tree) (G-3)pnpm ai:testruns the full pytest suite via turbo and passes (mock provider only — no network)- SSE stream delivers tokens: meta event, incremental deltas, done,
[DONE]observed viacurl -N - Mid-stream failure emits
errorevent before[DONE]; pre-first-byte failure returns HTTP error status - Provider factory resolves ollama-cloud / local / mock from settings; manual ollama-cloud probe documented in README
llm/imports nothing fromagents/orapi/(boundary rule holds)
Phase 2: Agent Framework
Requirements: REQ-2-004 Goal: Shared framework all six agents use: BaseAgent contract, session store, prompt library, registry, structured outputs — all tested against the mock provider
Wave 1: Framework primitives (parallel — no shared files)
Task 2-1-01: BaseAgent ABC
- Persona: ai-engineer — REQ: REQ-2-004
- Files:
apps/ai-service/ai_service/agents/base.py,apps/ai-service/tests/agents/test_base.py,apps/ai-service/ai_service/agents/__init__.py,apps/ai-service/tests/agents/__init__.py - Action:
BaseAgentABC (D-018):name,system_prompt,build_messages(history, learner_context),stream_reply(...) -> AsyncIterator[ChatDelta](delegates to provider),structured_reply(...)(delegates to structured module, landed Wave 2). Subclass contract tested with a stub agent + mock provider. - Verify:
pnpm ai:test— test_base green
Task 2-1-02: SessionStore protocol + in-memory implementation
- Persona: ai-engineer — REQ: REQ-2-004
- Files:
apps/ai-service/ai_service/agents/session.py,apps/ai-service/tests/test_session.py - Action:
SessionStoreprotocol +InMemorySessionStore(D-019): asyncio.Lock-guarded dict, agent-scoped session keys, 20-message rolling window, 500-cap LRU eviction. Protocol shape is DB-migration-ready (A-003). - Verify: test_session covers create/append/window-trim/LRU-eviction/agent scoping
Task 2-1-03: Prompt library scaffolding
- Persona: ai-engineer — REQ: REQ-2-004
- Files:
apps/ai-service/ai_service/prompts/coach.py,apps/ai-service/ai_service/prompts/tutor.py,apps/ai-service/ai_service/prompts/lab.py,apps/ai-service/ai_service/prompts/assessor.py,apps/ai-service/ai_service/prompts/proctor.py,apps/ai-service/ai_service/prompts/mentor.py,apps/ai-service/ai_service/prompts/__init__.py - Action: Per-agent module with a versioned
SYSTEM_PROMPTconstant +render_context(learner_context) -> dictusingstr.format_mapfor learner-context injection (D-018: prompts are code, versioned in git). Initial drafts for all six; final personas land in Phases 3-5. - Verify: all six prompt modules import; render_context fills placeholders without KeyError
Task 2-1-04: Learner context corpus
- Persona: ai-engineer — REQ: REQ-2-004
- Files:
apps/ai-service/ai_service/corpus/learner_context.py,apps/ai-service/ai_service/corpus/__init__.py - Action: Pydantic-typed learner context (active stack, competencies, progress, recent artifacts) mirroring TS
packages/mock-dataIDs per D-021 convention (cross-referencing header comment, identicalcomp-*/stack-*ID strings). - Verify: context renders into prompt placeholders; IDs match packages/mock-data strings
Wave 2: Composition layers (depends on Wave 1)
Task 2-2-01: Structured output defense
- Persona: ai-engineer — REQ: REQ-2-004
- Files:
apps/ai-service/ai_service/agents/structured.py,apps/ai-service/tests/test_structured.py - Action: 4-layer defense (D-020): (1)
response_formatrequest with auto-degrade on provider 400; (2) prompt-embedded JSON schema; (3) parse: strip code fences → first balanced JSON object; (4) single bounded retry with validation-error feedback. Returns pydantic-validated model or raisesStructuredOutputError. - Verify: test_structured covers fenced/unfenced/invalid JSON, retry path, degrade path — all against mock provider
Task 2-2-02: Agent registry
- Persona: ai-engineer — REQ: REQ-2-004
- Files:
apps/ai-service/ai_service/agents/registry.py,apps/ai-service/tests/test_registry.py - Action: Explicit registry: name → agent factory map with
register(name, factory)/get(name); raises on unknown agent. Agents are registered in their own phases (P3-P5). - Verify: test_registry: register/get round-trip, unknown-agent error, duplicate registration error
Task 2-2-03: Session + agent DI wiring into API layer
- Persona: backend-engineer — REQ: REQ-2-004
- Files:
apps/ai-service/ai_service/api/deps.py(update),apps/ai-service/ai_service/api/chat.py(update) - Action: deps.py exposes SessionStore and provider singletons via DI. chat.py persists turn history through the session store (agent-scoped) and includes session ID in the meta event. API composes agents via DI — agents never import api/.
- Verify: chat request appends to and replays windowed history;
pnpm ai:testgreen
Must-Haves (Phase 2)
- BaseAgent unit tests pass (stub agent streams via mock provider)
- Session store tested: create/append, 20-message window trim, 500-cap LRU eviction, agent-scoped keys
- Structured output parsing tested against mock provider: fence-strip, first-balanced-object, invalid JSON, one bounded retry, response_format auto-degrade
- Registry tested: register/get/unknown/duplicate
- Six prompt modules render learner context without errors
- Module boundaries hold:
agents/never importsapi/;llm/never importsagents/orapi/
Phase 3: Coach + Tutor Agents
Requirements: REQ-2-005, REQ-2-006 Goal: Both learner-facing conversational agents fully implemented with distinct personas, registered, routed through the chat streaming endpoint
Wave 1: Agent implementations (parallel — no shared files)
Task 3-1-01: Coach agent
- Persona: ai-engineer — REQ: REQ-2-005
- Files:
apps/ai-service/ai_service/prompts/coach.py(finalize),apps/ai-service/ai_service/agents/coach.py,apps/ai-service/tests/test_coach.py - Action: Final Coach persona: pacing guidance, motivation, retrieval practice prompts; system prompt injects learner context (active stack, progress).
CoachAgent(BaseAgent)streams replies. Mock provider scripts a distinct coach-voice response for tests. - Verify: test_coach: build_messages includes system prompt + windowed history; stream_reply yields deltas; on-persona content asserted against mock script
Task 3-1-02: Tutor agent
- Persona: ai-engineer — REQ: REQ-2-006
- Files:
apps/ai-service/ai_service/prompts/tutor.py(finalize),apps/ai-service/ai_service/agents/tutor.py,apps/ai-service/tests/test_tutor.py - Action: Final Tutor persona: concept delivery, Socratic questioning, worked examples.
TutorAgent(BaseAgent)streams replies; mock scripts a distinct tutor-voice response. - Verify: test_tutor mirrors test_coach; Coach and Tutor mock outputs are observably distinct
Wave 2: Registration + persona verification (depends on Wave 1)
Task 3-2-01: Register Coach + Tutor; document cloud probe
- Persona: ai-engineer — REQ: REQ-2-005, REQ-2-006
- Files:
apps/ai-service/ai_service/agents/registry.py(update),apps/ai-service/tests/test_registry.py(update),apps/ai-service/README.md(update) - Action: Register both agents in the explicit registry. Extend test_registry to assert both resolve. Document the manual ollama-cloud persona probe in README (curl commands with
AI_PROVIDER=ollama-cloud): Coach and Tutor produce distinct on-persona responses; tests remain cloud-free. - Verify:
pnpm ai:testgreen; manual probe against ollama-cloud shows distinct personas (documented, not automated)
Wave 3: Chat endpoint agent routing (depends on Wave 2)
Task 3-3-01: Agent routing on /v1/chat/stream
- Persona: backend-engineer — REQ: REQ-2-005, REQ-2-006
- Files:
apps/ai-service/ai_service/api/chat.py(update),apps/ai-service/ai_service/api/deps.py(update),apps/ai-service/tests/api/test_chat_stream.py(update) - Action: Chat request gains
agentfield (validated against the registry; unknown agent → 422). Endpoint resolves the agent via DI, persists to the agent-scoped session, meta event carries the agent name. No autonomous routing in v0.2 (A-007). - Verify: TestClient tests:
agent=coachandagent=tutorroute correctly, session scoped per agent, unknown agent rejected
Must-Haves (Phase 3)
- Both agents produce distinct, on-persona responses (mock-asserted; manual ollama-cloud probe documented in README)
- Agent routing tested: coach/tutor resolve via registry; unknown agent returns 422
- Both agents exposed end-to-end via
POST /v1/chat/streamwith agent-scoped session history pnpm ai:testgreen; no cloud calls in tests
Phase 4: Lab + Assessor Agents
Requirements: REQ-2-007, REQ-2-008 Goal: Lab consumes simulated sandbox telemetry and streams in-flow feedback; Assessor applies rubrics to pre-baked artifacts and returns structured scores — both over mock engine inputs
Wave 1: Mock engine inputs (parallel — no shared files)
Task 4-1-01: Simulated sandbox telemetry corpus
- Persona: ai-engineer — REQ: REQ-2-007
- Files:
apps/ai-service/ai_service/corpus/telemetry.py - Action: Pydantic-typed Lab telemetry scenarios: scripted build-session event streams (keystrokes, commits, test runs, errors, idle gaps) keyed by scenario ID, aligned with packages/mock-data IDs (D-021).
- Verify: scenarios import, validate, and are addressable by ID
Task 4-1-02: Pre-baked artifacts + rubrics corpus
- Persona: ai-engineer — REQ: REQ-2-008
- Files:
apps/ai-service/ai_service/corpus/artifacts.py - Action: Pre-baked artifacts (code, design, simulation), assessment rubrics (criteria, levels, weights), and defense transcripts keyed by ID — the Assessor's mock inputs (real engines are v0.3+).
- Verify: rubric/artifact/transcript fixtures validate; IDs align with TS mock data
Task 4-1-03: TS mock-data alignment for engine inputs
- Persona: data-engineer — REQ: REQ-2-012
- Files:
packages/mock-data/ai-scenarios.ts(new),packages/mock-data/index.ts(update) - Action: Export scenario/artifact ID constants + display metadata used by the Phase 6 learner panels, mirroring
ai_service/corpus/IDs exactly (D-021). Data-engineer owns the TS side; headers cross-reference the Python corpus. - Verify:
pnpm typecheckpasses; IDs string-equal to corpus IDs
Wave 2: Agent implementations (depends on Wave 1)
Task 4-2-01: Lab agent
- Persona: ai-engineer — REQ: REQ-2-007
- Files:
apps/ai-service/ai_service/prompts/lab.py(finalize),apps/ai-service/ai_service/agents/lab.py,apps/ai-service/tests/test_lab.py - Action:
LabAgent(BaseAgent)consumes a telemetry scenario, builds messages summarizing the event stream, streams concrete in-flow feedback (what happened, what to adjust, next step). No session chat — scenario-driven. - Verify: test_lab: given a mock scenario, feedback references scenario events (mock-scripted assertions)
Task 4-2-02: Assessor agent
- Persona: ai-engineer — REQ: REQ-2-008
- Files:
apps/ai-service/ai_service/prompts/assessor.py(finalize),apps/ai-service/ai_service/agents/assessor.py,apps/ai-service/tests/test_assessor.py - Action:
AssessorAgent(BaseAgent)applies a rubric to an artifact + defense transcript viastructured_reply, returning a pydantic-validated rubric score model (per-criterion scores, strengths, gaps, verdict). - Verify: test_assessor: structured output validates against the rubric model; failure modes exercise the 4-layer defense
Wave 3: Registration (depends on Wave 2)
Task 4-3-01: Register Lab + Assessor
- Persona: ai-engineer — REQ: REQ-2-007, REQ-2-008
- Files:
apps/ai-service/ai_service/agents/registry.py(update),apps/ai-service/tests/test_registry.py(update) - Action: Register both agents; extend registry tests.
- Verify: registry resolves coach/tutor/lab/assessor;
pnpm ai:testgreen
Wave 4: Endpoints (depends on Wave 3)
Task 4-4-01: Lab feedback endpoint
- Persona: backend-engineer — REQ: REQ-2-007
- Files:
apps/ai-service/ai_service/api/lab.py,apps/ai-service/tests/api/test_lab.py - Action:
POST /v1/lab/feedbackwith scenario ID → resolves corpus scenario + Lab agent → SSE stream using the D-016 envelope (meta names agent=lab). Unknown scenario → 404. - Verify: TestClient streams meta + deltas + done +
[DONE]; unknown scenario 404
Task 4-4-02: Assessment evaluate endpoint
- Persona: backend-engineer — REQ: REQ-2-008
- Files:
apps/ai-service/ai_service/api/assessment.py,apps/ai-service/tests/api/test_assessment.py - Action:
POST /v1/assessment/evaluatewith artifact ID → resolves corpus artifact/rubric/transcript + Assessor agent → JSON response (validated rubric score model). Unknown artifact → 404. - Verify: TestClient returns validated rubric JSON; unknown artifact 404;
pnpm ai:testgreen
Must-Haves (Phase 4)
- Lab produces scenario-relevant in-flow feedback for mock telemetry scenarios (mock provider, tested)
- Assessor returns structured rubric scores (pydantic-validated JSON) for pre-baked artifacts/transcripts
POST /v1/lab/feedbackstreams (meta → deltas → done →[DONE]);POST /v1/assessment/evaluatereturns validated JSON- Unknown scenario/artifact IDs return 404
- Corpus IDs align with packages/mock-data (D-021);
pnpm ai:testandpnpm typecheckgreen
Phase 5: Proctor + Mentor Agents
Requirements: REQ-2-009, REQ-2-010 Goal: Proctor classifies integrity signals with coaching interventions from mock telemetry; Mentor generates long-horizon career narrative; both exposed via endpoints
Wave 1: Proctor scenarios + Mentor agent (parallel — no shared files)
Task 5-1-01: Proctor telemetry scenarios
- Persona: ai-engineer — REQ: REQ-2-009
- Files:
apps/ai-service/ai_service/corpus/telemetry.py(update) - Action: Add proctor scenarios: tab switches, idle time, paste events, focus loss — scripted integrity-relevant event sets keyed by scenario ID.
- Verify: proctor scenarios validate; distinguishable from lab scenarios by type
Task 5-1-02: Mentor agent (+ registration)
- Persona: ai-engineer — REQ: REQ-2-010
- Files:
apps/ai-service/ai_service/prompts/mentor.py(finalize),apps/ai-service/ai_service/agents/mentor.py,apps/ai-service/tests/test_mentor.py,apps/ai-service/ai_service/agents/registry.py(update) - Action:
MentorAgent(BaseAgent): long-horizon career narrative — trajectory story, competency-stack progression guidance, market positioning — streaming, session-backed. Registered centrally inregistry.py(single registration pattern, G-4). - Verify: test_mentor: narrative references learner context (mock-scripted); registry resolves mentor
Wave 2: Proctor agent + Mentor endpoint (depends on Wave 1)
Task 5-2-01: Proctor agent (+ registration)
- Persona: ai-engineer — REQ: REQ-2-009
- Files:
apps/ai-service/ai_service/prompts/proctor.py(finalize),apps/ai-service/ai_service/agents/proctor.py,apps/ai-service/tests/test_proctor.py,apps/ai-service/ai_service/agents/registry.py(update) - Action:
ProctorAgent(BaseAgent): consumes proctor scenario →structured_replyreturns pydantic-validated signal classification (severity, signal type) + recommended coaching intervention (supportive, not punitive). Registered centrally inregistry.py(single registration pattern, G-4). - Verify: test_proctor: classified signals + interventions validate for each mock scenario; registry resolves all six agents
Task 5-2-02: Mentor narrative endpoint
- Persona: backend-engineer — REQ: REQ-2-010
- Files:
apps/ai-service/ai_service/api/mentor.py,apps/ai-service/tests/api/test_mentor.py - Action:
POST /v1/mentor/narrative→ Mentor agent → SSE stream with D-016 envelope, session-backed. - Verify: TestClient streams meta (agent=mentor) → deltas → done →
[DONE]
Wave 3: Proctor endpoint (depends on Wave 2)
Task 5-3-01: Proctor signals endpoint
- Persona: backend-engineer — REQ: REQ-2-009
- Files:
apps/ai-service/ai_service/api/proctor.py,apps/ai-service/tests/api/test_proctor.py - Action:
POST /v1/proctor/signalswith scenario ID → resolves corpus scenario + Proctor agent → JSON response (validated signals + interventions). Unknown scenario → 404. - Verify: TestClient returns classified signals JSON;
pnpm ai:testgreen — full suite (all six agents registered)
Must-Haves (Phase 5)
- Proctor produces classified signals with recommended coaching interventions for each mock scenario (structured JSON, validated)
- Mentor produces coherent long-horizon career narrative (streaming, session-backed)
POST /v1/proctor/signalsreturns validated JSON;POST /v1/mentor/narrativestreams- Registry resolves all six agents; full ai-service test suite green, cloud-free
Phase 6: Learner Surface Integration
Requirements: REQ-2-011, REQ-2-012
Goal: v0.1 learner surfaces wired to the real ai-service: streaming chat with agent switcher, Lab/Assessor/Proctor/Mentor outputs surfaced, error/loading states, build + typecheck green
Note (G-2): End-to-end verification may run with AI_PROVIDER=mock as a fallback — the requirement is the real ai-service over HTTP (not canned client-side responses); provider choice is service-internal. This prevents an ollama-cloud outage from blocking P6 verification. Cloud persona probes remain separate (Task 3-2-01).
Wave 1: Client plumbing + primitives (parallel — no shared files)
Task 6-1-01: useChatStream hook
- Persona: frontend-engineer — REQ: REQ-2-011
- Files:
apps/web/hooks/use-chat-stream.ts,apps/web/.env.example(update) - Action:
useChatStream(agent)hook:fetchPOST to${NEXT_PUBLIC_AI_SERVICE_URL}/v1/chat/stream(defaulthttp://localhost:8420, A-002 — no API-route proxy); consumesReadableStreamwith byte buffering, frame split on\n\n, joineddata:lines; ignores frames containing nodata:lines (sse-starlette: pingkeep-alive comment frames) — TestClient streams are too short to surface pings, but real cloud delta gaps emit them (G-1); handles meta / delta / done / error events and[DONE]sentinel; idempotentAbortController.abort()in effect cleanup; exposes{messages, isStreaming, error, send, retry, abort}..env.examplegainsNEXT_PUBLIC_AI_SERVICE_URL. - Verify: hook unit-tested or exercised via the chat UI; unmount mid-stream aborts cleanly (no state updates after unmount)
Task 6-1-02: Agent switcher + streaming primitives
- Persona: design-system-engineer — REQ: REQ-2-011
- Files:
packages/ui/src/primitives/agent-switcher.tsx,packages/ui/src/primitives/stream-status.tsx,packages/ui/src/primitives/toast.tsx,packages/ui/src/primitives/index.ts(update),packages/ui/src/index.ts(update) - Action: Token-driven primitives: AgentSwitcher (segmented coach/tutor control with active state), StreamStatus (idle/streaming/error indicator), Toast with error variant. Dark mode + WCAG AA contrast; exported from
@nextcraft/ui. - Verify: primitives import from
@nextcraft/ui; storybook stories render (dark + light)
Wave 2: Chat rewrite + Byte viewer panel (depends on Wave 1)
Task 6-2-01: Real streaming learner chat
- Persona: frontend-engineer — REQ: REQ-2-011
- Files:
apps/web/components/learner/ai-tutor-chat.tsx(rewrite),apps/web/components/learner/agent-switcher.tsx(composition wrapper) - Action: Replace the canned
aiTutorResponsesbehavior withuseChatStream: agent switcher (Coach/Tutor per A-007), token-by-token rendering, streaming cursor + loading state, error state with retry button when ai-service is down (A-010), suggested-action chips from the meta event. Seed welcome message stays static. - Verify: with ai-service running, messages stream visibly token-by-token; with ai-service stopped, error state + retry appears (no crash, no console errors)
Task 6-2-02: Byte viewer Tutor panel
- Persona: frontend-engineer — REQ: REQ-2-011 (agent routing: byte viewer always uses Tutor, A-007)
- Files:
apps/web/app/(learner)/learn/[competencyId]/page.tsx(update),apps/web/components/learner/byte-tutor-panel.tsx(new) - Action: Byte viewer gains a Tutor explanation panel (byte viewer always uses Tutor, A-007): "Explain this byte" streams a Socratic concept walkthrough for the current competency via useChatStream (agent fixed to tutor).
- Verify: on a byte page, the panel streams a Tutor explanation; error state when service down
Wave 3: Lab / Assessment / Mentor panels (depends on Wave 1 hook; parallel — no shared files)
Task 6-3-01: Sandbox Lab feedback panel
- Persona: frontend-engineer — REQ: REQ-2-012
- Files:
apps/web/app/(learner)/build/[competencyId]/page.tsx(update),apps/web/components/learner/lab-feedback-panel.tsx(new) - Action: Sandbox telemetry sidebar gains a Lab feedback panel: posts the scenario ID (from
packages/mock-data/ai-scenarios) to/v1/lab/feedback, streams in-flow feedback into the panel; loading + error states. - Verify: sandbox page streams Lab feedback for the mock scenario; error state when service down
Task 6-3-02: Assessment Assessor + Proctor surfaces
- Persona: frontend-engineer — REQ: REQ-2-012
- Files:
apps/web/app/(learner)/defend/[competencyId]/page.tsx(update),apps/web/components/learner/assessor-results-panel.tsx(new),apps/web/components/learner/proctor-banner.tsx(new) - Action: Assessment mockup: AI reviewer panel calls
/v1/assessment/evaluatewith the artifact ID and renders the structured rubric scores (per-criterion bars, strengths, gaps, verdict); a Proctor integrity banner surfaces/v1/proctor/signalsclassifications with coaching tone. Loading skeletons + error states. - Verify: defend page renders real Assessor rubric output + Proctor banner; error states when service down
Task 6-3-03: Dashboard Mentor panel
- Persona: frontend-engineer — REQ: REQ-2-012
- Files:
apps/web/app/(learner)/dashboard/page.tsx(update),apps/web/components/learner/mentor-panel.tsx(new) - Action: Learner dashboard gains a Mentor panel: streams career narrative from
/v1/mentor/narrative(learner progress context), with regenerate button, loading + error states. Sits alongside the existing AI tutor chat. - Verify: dashboard shows streaming Mentor narrative; error state when service down
Must-Haves (Phase 6)
- With ai-service running: learner chat at http://localhost:3000/dashboard streams real responses token-by-token (mock provider fallback allowed per G-2 — service over HTTP is the requirement)
- Agent switcher flips Coach ↔ Tutor and the response persona changes accordingly
- Hook tolerates keep-alive comment frames (
: ping, no data lines) during live streams (G-1) - With ai-service stopped: all chat/panels show error states with retry — no crashes, no unhandled promise rejections, no console errors
- Byte viewer, sandbox, and assessment mockups surface Tutor/Lab/Assessor/Proctor outputs; dashboard shows Mentor narrative
- Unmounting/navigating mid-stream aborts cleanly (no post-unmount state updates)
pnpm buildandpnpm typecheckpass;pnpm ai:teststill green
Phase 7: Final Review + Ship (no planned tasks)
Orchestrated by the SHIP stage, not this plan: multi-persona code review (correctness, testing, secrets hygiene — key absent from code/logs/commits/errors, localhost-only CORS, no PII in prompts), project health audit (reconstruction test, .ciagent/ discipline, branch/commit hygiene), then merge milestone → main, tag v0.2.0, create Gitea release, mark all 12 v0.2 requirements complete.
Release-note honesty (G-5): the v0.2.0 release note must explicitly state that Lab/Assessor/Proctor operate on mock engine inputs (real engines are v0.3+), per D-015.
Dead-code disposal (G-5): review must dispose of aiTutorResponses (packages/mock-data/ai-tutor-responses.ts) — its only consumer is rewritten in Task 6-2-01; remove the export or mark it deprecated.
User-Facing Surface
The primary user-facing surface is the learner dashboard and learning flow at http://localhost:3000, backed by the real ai-service at http://localhost:8420:
/dashboard— AI tutor chat (Coach/Tutor switcher, streaming) + Mentor career-narrative panel/learn/[competencyId]— byte viewer with streaming Tutor explanations/build/[competencyId]— sandbox with Lab in-flow feedback panel (mock telemetry)/defend/[competencyId]— assessment with live Assessor rubric scores + Proctor integrity banner
The marketplace, employer, and admin surfaces are unchanged from v0.1.
Happy Path
- Learner opens
/dashboard→ chat shows welcome message; meta event confirms coach/model in the stream - Learner types "I'm stuck on multi-agent communication" → reply streams token-by-token with pacing guidance + a retrieval-practice prompt
- Learner switches to Tutor → asks the same question → gets a Socratic concept walkthrough instead
- Learner opens a byte tutorial → Tutor panel streams an explanation of the current competency
- Learner opens the build sandbox → Lab panel streams feedback on the simulated telemetry scenario
- Learner opens the defense mockup → Assessor panel shows structured rubric scores; Proctor banner shows integrity signals in coaching tone
- Back on
/dashboard, the Mentor panel streams a career narrative tied to the learner's progress - Learner kills ai-service (or it crashes) → next message shows an inline error state with Retry; restarting the service and retrying resumes streaming
UX Acceptance Criteria
- Streaming is visibly incremental — tokens appear as they arrive, not as one blob
- Agent switcher shows the active agent (Coach/Tutor) and the response persona visibly changes
- Loading state during connection (streaming cursor / skeleton) before first token
- When ai-service is unreachable: inline error state + retry action on every chat/panel — no crashes, no console errors, no blank UI
[DONE]reliably ends the stream (input re-enables, no stuck "typing" state)- Navigating away mid-stream aborts cleanly — no leaked requests or post-unmount updates
- All new UI uses design tokens, supports dark mode, meets WCAG AA contrast
- Responsive at 375px, 768px, 1280px
- No hardcoded model names or URLs in UI code — all via
NEXT_PUBLIC_AI_SERVICE_URLand server meta events pnpm buildandpnpm typecheckpass with zero errors