Compare commits

..

4 Commits

Author SHA1 Message Date
CIAgent 4c52d29f91 docs(P02): complete agent-framework phase
---ci---
phase: 2
milestone: v0.2
status: complete
---/ci---

REQ-2-004 complete. BaseAgent ABC, agent-scoped sessions (20-msg window,
500-cap LRU), registry, 4-layer structured output defense, 6-module
prompt library, D-021-aligned learner corpus, chat session persistence.
61/61 tests, ruff clean.
2026-09-11 15:52:00 +00:00
CIAgent dda9569b80 docs(P01): complete ai-service-scaffolding phase
---ci---
phase: 1
milestone: v0.2
status: complete
---/ci---

REQ-2-001/002/003 complete. apps/ai-service: FastAPI + provider-agnostic
LLM layer (ollama-cloud/local/mock) + D-016 SSE envelope + 20 tests
(mock-only, cloud-free guard) + ruff + turbo/pnpm integration.
2026-09-11 15:40:30 +00:00
CIAgent c7fe601481 chore(P00): checkpoint complete
---ci---
phase: 0
milestone: v0.2
status: complete
---/ci---
2026-09-11 15:16:38 +00:00
CIAgent e58b027a57 docs(P00): complete pre-execution phase
---ci---
phase: 0
milestone: v0.2
status: complete
---/ci---

Phase 0 complete: SPECIFY, CLARIFY (A-001..010), RESEARCH
(D-016..023, personas), PLAN (38 tasks, 18 waves), GRILL
(G-1..G-5 applied), MVP/UX gate passed.
2026-09-11 15:16:02 +00:00
200 changed files with 1786 additions and 23297 deletions
+25 -133
View File
@@ -4,9 +4,7 @@
Nextcraft is a TypeScript monorepo (pnpm workspaces + turborepo) with a Next.js application hosting four surfaces (Learner, Marketplace, Employer Dashboard, Admin), a shared component library, typed mock data layer, and shared types package. As of v0.2, a Python FastAPI application (`apps/ai-service`) hosts six AI tutor agents backed by a provider-agnostic LLM layer.
**v0.3 additions (Credential Engines):** real credential engines replace v0.2 mock inputs — a sandbox fabric (isolated per-learner coding environments via Linux user/mount/pid/net namespaces), a live build-telemetry pipeline (WebSocket ingest + SQLite-ordered event log), a process-trace grading engine, seeded per-learner variant task generation, and a voice-based oral defense (STT/TTS via a new provider-agnostic voice layer). **First real persistence introduced: SQLite** (`ai_service/telemetry/`, grading, variant, defense stores). Lab/Assessor/Proctor agents are re-grounded onto real telemetry/traces. **Identity/age-gating (KYC) deferred per founder directive** — no security engineer persona; secrets-hygiene checklist only.
**v0.4 additions (Distribution & Bootstrap CLI, founder directive D-016):** a new `apps/cli` package — the `nextcraft` bootstrap CLI (`doctor`/`bootstrap`/`verify`/`dev`) compiled to a self-contained linux x64 binary via **Node SEA** (probe-verified: Go/Rust absent, node v24.15.0 SEA-capable), installed by a repo-served one-liner script that resolves the latest Gitea release, downloads binary + sha256 sidecar, verifies, and installs to `~/.local/bin`. Every release from v0.4 onward attaches the binary + checksum as release assets (the "ongoing binaries" requirement). The CLI is a thin wrapper: all orchestration logic stays in `apps/ai-service/scripts/` (bootstrap.sh/dev.sh) — the CLI composes them via subprocess (A-202), duplicating nothing. Previously-planned v0.4 seams (real STT/TTS, KYC, design/sim envs, seq-lease) move to v0.5.
**v0.2 additions:** real AI services (streaming chat), agent framework, mock engine inputs for Lab/Assessor/Proctor. **Still no database, no auth** — in-memory session store; real engines (sandbox fabric, assessment engine, identity) are v0.3+.
### Confirmed Technology Stack (v0.2)
@@ -28,8 +26,6 @@ Nextcraft is a TypeScript monorepo (pnpm workspaces + turborepo) with a Next.js
| pydantic | 2.13.x | Request/response models, structured outputs |
| pydantic-settings | 2.15.x | Settings + env-file loading (replaces python-dotenv) |
| httpx | 0.28.x | Async LLM HTTP client (ollama-cloud + local providers) |
| sqlmodel / sqlalchemy | 0.0.24 / 2.x | Typed SQLite persistence for the v0.3 engine stores (D-027) |
| python-multipart | 0.0.x | Multipart audio upload for the defense answer route (REQ-3-006) |
| sse-starlette | 3.4.x | SSE framing, ping keep-alive |
| pytest | 9.x | Test runner |
| pytest-asyncio | 1.4.x | Async tests (auto mode) |
@@ -49,27 +45,6 @@ Deliberately **not** used: openai-python SDK (the `LLMProvider` protocol is the
7. **D-022 Monorepo integration** — zero-dependency shim `package.json` in apps/ai-service + `ai#*` turbo passthrough tasks (`cache:false, outputs:[]`) + root `ai:dev`/`ai:test` scripts + idempotent venv bootstrap.
8. **D-023 Testing** — pytest-asyncio auto mode; TestClient `client.stream()` for SSE; httpx MockTransport for byte-exact provider parser tests; scripted mock provider incl. failure modes. Tests never call the cloud.
### v0.4 Architecture Decisions (from Research — Distribution & Bootstrap CLI)
18. **D-033 Binary toolchain = Node SEA (probe-verified)** — Go and Rust are absent from this box; node v24.15.0 ships SEA support (`--experimental-sea-config`, postject-free on linux via `cp node nextcraft && node sea-config` … blob injection with the system `dd`/`npx postject` if needed). CLI source lives in `apps/cli` (TypeScript, compiled to a single CJS bundle by esbuild, then SEA-injected into a copy of the node binary → `nextcraft-linux-x64`). Fallback if SEA breaks: python3 `zipapp` (3.11.2 available). No new toolchain deps beyond dev-scoped esbuild.
19. **D-034 CLI = thin wrapper, orchestration stays in scripts/**`nextcraft` composes `apps/ai-service/scripts/bootstrap.sh` and `scripts/dev.sh` equivalents via `spawn` with inherited stdio and timeout guards (A-202/A-209). doctor/bootstrap/verify implement only *checking* logic (prereqs, env template, health) — never re-implement installs. This keeps one source of truth for bootstrap semantics.
20. **D-035 Install path = repo raw `install.sh` + Gitea latest-release API** — the one-liner `curl -fsSL <forge>/coreci/nextcraft/raw/main/scripts/install.sh | bash` resolves `GET /api/v1/repos/coreci/nextcraft/releases/latest`, downloads the `nextcraft-linux-x64` + `nextcraft-linux-x64.sha256` assets, verifies sha256 (`shasum -a 256`), installs to `~/.local/bin` (PATH hint), and degrades to printed source-bootstrap instructions when no binary asset exists or the platform mismatches (A-203/A-204/A-206).
21. **D-036 Ongoing binaries = ship-workflow asset step** — the release pipeline (v0.3's `ShipWorkflow.createRelease` equivalent, executed as the ship step's asset stage) builds the binary + checksum and attaches both to every Gitea release from v0.4 onward (A-205). Token resolution stays `.env*`-only (D-006/D-014); binaries are linux x64 only for v0.4 (macOS arm64 deferred — unverifiable on this box).
22. **D-037 CLI package layout**`apps/cli` is a pnpm workspace package (`@nextcraft/cli`): `src/` (entry, commands/, checks/, lib/), `scripts/build-binary.mjs` (esbuild bundle → SEA inject), unit tests runnable via `pnpm --filter @nextcraft/cli test` (node:test, no new test framework). Root `package.json` gains `cli:*` passthrough scripts mirroring the `ai:*` pattern (D-022).
23. **D-038 Network mode (v0.3.5)** — dev binds 0.0.0.0 (`AI_HOST`, default 0.0.0.0, revert via 127.0.0.1); CORS + WS-origin gates read `AI_CORS_ORIGINS` (default `*` — any origin, safe only because credentials are never enabled; explicit comma list restricts); the web client derives the API base URL from the browser hostname at runtime (`engine-base-url.ts`: `NEXT_PUBLIC_AI_SERVICE_URL` override → `http://${window.location.hostname}:8420``localhost` server-side). Hotfix also fixes: SEA direct-run detection (`require("node:sea").isSea()` — argv shape differs by invocation style), installer honesty gate (silent `--version` = hard fail), bootstrap venv recovery (poisoned partial `.venv` removal + distro-specific `apt install python3.XX-venv` hint), and doctor venv-capability probe with bootstrap preflight.
### v0.3 Architecture Decisions (from Research — Credential Engines)
9. **D-024 Sandbox isolation = Linux namespaces via `unshare`** — per-learner sandbox runs as a subprocess entered into fresh user+mount+pid+network namespaces (`unshare --user --map-root-user --mount --pid --fork --net`). Probe-verified on this box: in-namespace uid=0, **network fully isolated** (0 interfaces), learner writes land in a per-sandbox directory; proc-remount not permitted here but not required. Chosen because no container runtime (docker/podman/bwrap/firejail) exists on the box and there is no sudo. A `SandboxBackend` protocol abstracts the spawner so a future containerd/runc backend can replace namespace-spawning without touching callers.
10. **D-025 Sandbox scope = coding IDE only (v0.3)** — the sandbox fabric provisions a single build environment (shell + filesystem + run/test). REQ-F-021's design-tool and simulation environments are deferred to v0.4; one real build path proves the full credential pipeline (telemetry → trace → grade → defense).
11. **D-026 Telemetry = WebSocket ingest + SQLite ordered event log** — in-sandbox capture agent streams structured events over WebSocket to `ai_service` (`/v1/telemetry/ingest`); events persisted to SQLite with a per-(learner,task) monotonic `seq` for gap detection, giving durability + at-least-once delivery + replay without a message broker.
12. **D-027 First persistence = SQLite, protocol-wrapped** — introduces a real DB (`ai_service/data/*.db`) for telemetry traces, grades, variants, and defenses. Access via SQLModel. Every store is a protocol (`TraceStore`, `VariantStore`, `DefenseStore`, `GradeStore`) with a SQLite implementation — Postgres-migration-ready, mirroring D-019's SessionStore pattern.
13. **D-028 Process-trace grading = hybrid deterministic + LLM** — deterministic features (test pass/fail, edit count, error/fix cycles, idle gaps, command categories) computed in code into a compact trace digest; the digest feeds an Assessor-style rubric prompt and returns structured scores via D-020 JSON defense. The LLM never sees the raw trace — only the digest.
14. **D-029 Variant generation = seeded template instantiation** — task templates with typed parameter slots; an LLM instantiates a unique variant per learner from a seed; seed+parameters persisted (D-027) for grading fairness and proctoring cross-check.
15. **D-030 Voice = provider-agnostic, mock-first, browser-fallback** — a `VoiceProvider` protocol (mirror of `LLMProvider`) with STT (OpenAI-compatible `/audio/transcriptions`) + TTS (`/audio/speech`) against a configurable endpoint, a deterministic mock (canned transcript/audio) for tests, and browser-native `SpeechRecognition`/`speechSynthesis` as a no-key fallback. Voice defense reuses `BaseAgent` + the existing SSE pipeline (new `Examiner` agent).
16. **D-031 No new apps — extend ai-service** — telemetry, grading, variant, voice, and sandbox orchestration are new modules inside `apps/ai-service` (sharing the LLM pool, config, and session infra). Only the in-sandbox capture agent is a separate tiny Python process shipped into the namespace. No new top-level `apps/` entry.
17. **D-032 Capacity = single-box, 15 concurrent sandboxes** — concurrency guard returns 503 when the sandbox pool is full. No queueing, no horizontal scaling in v0.3 (solo-founder/pilot scale).
---
## Components
@@ -78,68 +53,38 @@ Deliberately **not** used: openai-python SDK (the `LLMProvider` protocol is the
| Component | Description | Boundaries | Depends On |
|-----------|-------------|------------|------------|
| `ai_service/main.py` | FastAPI app factory, lifespan (httpx client pool, provider factory, SandboxManager + reaper loop, SQLite engine stores on app.state), CORS (localhost only, incl. PUT for file writes), /health | App entry | config, llm, agents, api, engines |
| `ai_service/main.py` | FastAPI app factory, lifespan (httpx client pool, provider factory), CORS (localhost only), /health | App entry | config, llm, agents, api |
| `ai_service/config.py` | pydantic-settings Settings (env_prefix="AI_", env_file, SecretStr key) | Configuration only | None |
| `ai_service/api/` | Endpoints: chat.py (POST /v1/chat/stream, SSE), lab.py, assessment.py (POST /v1/assessment/evaluate v0.2 + POST /v1/assessment/grade v0.3), proctor.py, mentor.py, sandboxes.py (lifecycle + files/exec routes, G-5 abuse gates), telemetry.py (WS ingest + trace/gaps reads), variants.py (seeded per-learner variants), defense.py (defense loop, REQ-3-006); deps.py (DI) | Composes agents + sessions + engines; never imported by llm/ or agents/ | agents, llm, sandbox, telemetry, grading, variants, voice |
| `ai_service/llm/` | types.py (Message; ChatDelta/ChoiceDelta removed in P3 — no consumers), base.py (LLMProvider protocol), openai_compat.py (ollama-cloud + local), mock.py (deterministic), factory.py | Never imports agents/ or api/ | config |
| `ai_service/agents/` | base.py (BaseAgent ABC), registry.py, session.py (SessionStore), structured.py (JSON defense), coach/tutor/lab/assessor/proctor/mentor.py + examiner.py (seventh agent, v0.3) | Never imports api/ | llm, prompts, corpus, telemetry |
| `ai_service/api/` | Endpoints: chat.py (POST /v1/chat/stream, SSE), lab.py, assessment.py, proctor.py, mentor.py; deps.py (DI) | Composes agents + sessions; never imported by llm/ or agents/ | agents, llm |
| `ai_service/llm/` | types.py (Message, ChatDelta), base.py (LLMProvider protocol), openai_compat.py (ollama-cloud + local), mock.py (deterministic), factory.py | Never imports agents/ or api/ | config |
| `ai_service/agents/` | base.py (BaseAgent ABC), registry.py, session.py (SessionStore), structured.py (JSON defense), coach/tutor/lab/assessor/proctor/mentor.py | Never imports api/ | llm, prompts, corpus |
| `ai_service/prompts/` | Per-agent system prompt constants + render_context functions (str.format_map) | Data only | None |
| `ai_service/corpus/` | Mock engine inputs: learner_context.py, telemetry.py (Lab/Proctor scenarios), artifacts.py (pre-baked artifacts, rubrics, transcripts). Since v0.3 P6 these are DORMANT, test-only fixtures (dormant-header noted) — the live learner path uses real engine inputs | Pydantic-typed; aligned with TS packages/mock-data by convention | None |
| `scripts/` | bootstrap.sh (venv + pip install idempotent), dev.sh (exports keys from .ciagent/.env.secrets → uvicorn), test.sh (pytest), lint.sh (ruff check, v0.2 G-3) | Dev entry points | pyproject.toml |
| `ai_service/corpus/` | Mock engine inputs: learner_context.py, telemetry.py (Lab/Proctor scenarios), artifacts.py (pre-baked artifacts, rubrics, transcripts) | Pydantic-typed; aligned with TS packages/mock-data by convention | None |
| `scripts/` | bootstrap.sh (venv + pip install idempotent), dev.sh (exports keys from .ciagent/.env.secrets → uvicorn), test.sh (pytest) | Dev entry points | pyproject.toml |
| `tests/` | conftest.py (mock provider, settings override, TestClient), health, llm (MockTransport parser), agents (framework + per-agent), api (SSE stream tests) | Mock provider only — no cloud | all |
**Module boundary rules:** `llm/` never imports `agents/` or `api/`; `agents/` never imports `api/`; `api/` composes both via DI. `corpus/` is the only home of mock engine data. Prompts are code — versioned and reviewed in git.
### apps/ai-service — v0.3 Credential Engine modules (NEW)
| Component | Description | Boundaries | Depends On |
|-----------|-------------|------------|------------|
| `ai_service/sandbox/` | `backend.py` (SandboxBackend protocol), `unshare_backend.py` (userns/mount/pid/net spawner, D-024), `manager.py` (lifecycle: create/list/snapshot/destroy + concurrency guard D-032), `workdir.py` (per-sandbox fs layout) | Never imports api/ or agents/; spawns subprocesses only | config |
| `ai_service/telemetry/` | `models.py` (TelemetryEvent, TraceSpan), `store.py` (TraceStore protocol + SQLite impl D-027), `ingest.py` (WebSocket /v1/telemetry/ingest, seq gap detection D-026) | Persistence; never imports agents/ | config |
| `ai_service/grading/` | `features.py` (deterministic trace digest D-028), `engine.py` (rubric scoring orchestration), `store.py` (GradeStore) | LLM only via digest; never sees raw trace | llm, telemetry, prompts |
| `ai_service/variants/` | `templates.py` (task template library), `generator.py` (seeded LLM instantiation D-029), `store.py` (VariantStore) | LLM via structured output | llm, grading |
| `ai_service/voice/` | `base.py` (VoiceProvider protocol D-030), `browser.py` (native SR/TTS fallback descriptor), `mock.py` (deterministic; the real server STT/TTS provider is the v0.4 seam — GRILL CUT-1/G-7), `factory.py` (provider selection), `defense_store.py` (DefenseStore: transcripts + integrity signals, D-027) | Never imports agents/ or api/ | config |
| `ai_service/agents/examiner.py` | Seventh agent: oral defense examiner; streams over existing SSE, consumes process traces + emits integrity signals | reuses BaseAgent (D-018) | llm, prompts, telemetry |
| `ai_service/data/*.db` | SQLite databases (telemetry/grades/variants/defenses) | gitignored | — |
| `scripts/sandbox-agent.py` | Tiny in-namespace capture process shipped into the sandbox; streams telemetry to ingest | standalone | stdlib only |
**Boundary additions:** `sandbox/`, `telemetry/`, `grading/`, `variants/`, `voice/` are engine modules — they never import `api/` (which composes them via DI) and never import `agents/` (agents call engines through narrow interfaces, not vice versa).
### apps/cli — Nextcraft Bootstrap CLI (v0.4 NEW)
| Component | Description | Boundaries | Depends On |
|-----------|-------------|------------|------------|
| `src/index.ts` | Entry: arg parsing (no deps beyond node stdlib at runtime), command dispatch, `--help`/`--version`, exit-code contract (0 ok / 1 failure / 2 usage) | CLI surface only | commands/ |
| `src/commands/` | `doctor.ts` (prereq checks + actionable errors), `bootstrap.ts` (pnpm install + scripts/bootstrap.sh wrapper + env template copy + key validation), `verify.ts` (health: venv imports, ports, env, build readiness), `dev.ts` (thin passthrough to scripts/dev.sh) | Compose checks/ + lib/; spawn scripts — never re-implement them | checks/, lib/ |
| `src/checks/` | Pure check functions: `check-command.ts` (binary-on-PATH + version compare), `check-env.ts` (template diff, required/optional key classification) | Pure logic, unit-testable, no fs side effects at import | None |
| `src/lib/` | `spawn.ts` (subprocess with timeout + inherited stdio), `log.ts` (✓/✗/warn output formatter) | Shared utilities | None |
| `scripts/build-binary.mjs` | esbuild → CJS bundle → Node SEA injection → `dist/nextcraft-linux-x64` + sha256 sidecar | Build-time only | esbuild (dev dep) |
| `scripts/install.sh` | The one-liner install script served from repo raw: Gitea latest-release resolve → download + checksum verify → ~/.local/bin; source-bootstrap fallback | Standalone POSIX sh | forge API |
| `tests/` | node:test unit tests: command dispatch, check logic, env template diff, install-script shellcheck-style assertions | Fixtures only — never mutate repo state | src/ |
**Boundary rules:** the CLI never imports from `apps/web`, `packages/*`, or `ai_service` Python modules — it orchestrates them exclusively via subprocess/filesystem. Runtime deps: node stdlib only (no runtime npm deps; esbuild is dev-only). The binary embeds the bundle; `scripts/bootstrap.sh` remains the single source of bootstrap truth (D-034).
### apps/web — Next.js Application
| Component | Description | Boundaries | Depends On |
|-----------|-------------|------------|------------|
| `app/(learner)/` | Learner surface route group: landing, catalog, competency stack, dashboard, byte viewer, build surface (`/build/[competencyId]` — real in-browser build), defense surface (`/defend/[competencyId]` — live oral defense + grading) | Learner-only routes and layouts | packages/ui, packages/mock-data, packages/types |
| `app/(learner)/` | Learner surface route group: landing, catalog, competency stack, dashboard, byte viewer, sandbox mockup, assessment mockup | Learner-only routes and layouts | packages/ui, packages/mock-data, packages/types |
| `app/(marketplace)/` | Marketplace surface route group: job board, job detail, employer profile, search/filter, pricing | Marketplace-only routes and layouts | packages/ui, packages/mock-data, packages/types |
| `app/(employer)/` | Employer dashboard route group: overview, talent search, candidate profile, posting management | Employer-only routes and layouts | packages/ui, packages/mock-data, packages/types |
| `app/(admin)/` | Admin surface route group: overview, learner management, competency graph viewer, moderation | Admin-only routes and layouts | packages/ui, packages/mock-data, packages/types |
| `app/layout.tsx` | Root layout: theme provider, navigation shell, responsive container | All routes | packages/ui |
| `components/` | Surface-specific components (learner/, marketplace/, employer/, admin/) plus shared chrome (navigation-shell, header/footer, role-switcher, theme-provider, breadcrumbs, dark-mode-toggle); v0.3 learner: build-surface, sandbox-terminal (read-only exec output), defense-session | App-level components | packages/ui |
| `hooks/` | use-chat-stream.ts — SSE client hook: fetch + ReadableStream, byte buffering + frame reassembly, idempotent AbortController cleanup; use-sandbox-session.ts (v0.3) — sandbox lifecycle for the build session: create on task open, destroy on unmount, mid-start failure cleanup, 503/403/429 honest surfaces | Client components only | ai-service SSE / engine API |
| `lib/` | sse.ts (shared SSE frame parser — CRLF normalization + `: ping` immunity, v0.2 G-1), breadcrumbs.ts, format.ts, engine-base-url.ts (v0.3.5: runtime API base — env override → browser hostname → localhost), engine-client.ts (v0.3: typed fetch client for /v1/sandboxes, files/exec, variants, grade, defense, traces) | Pure utilities | None |
| `components/` | Surface-specific components (not shared across surfaces) | Per-surface only | packages/ui |
### packages/ui — Shared Component Library
| Component | Description | Boundaries | Depends On |
|-----------|-------------|------------|------------|
| `tokens/` | Design tokens as TS constants: colors, spacing, radii, shadows, breakpoints (mirrored as Tailwind v4 `@theme` tokens in apps/web globals.css) | Foundation layer — no dependencies | None |
| `primitives/` | Button, Input, Card, Badge, Avatar (v0.1) + TerminalFrame, TelemetryStatus, MicControl, GradeBadge, TranscriptViewer (v0.3 build/defense surfaces) — each with a Storybook story | Atomic UI components | tokens, packages/types |
Composite/layout/theme components (navigation shell, tables, chat panels, graph viewer, theme provider) live in `apps/web/components/` as app-level components, not in packages/ui.
| `design-tokens/` | CSS custom properties: color palette, typography scale, spacing system, breakpoints, shadows, radii | Foundation layer — no dependencies | None |
| `primitives/` | Button, Input, Card, Badge, Avatar, Dialog, Tabs, Progress, Tooltip, Skeleton, Toast | Atomic UI components | design-tokens |
| `composites/` | Navigation, Table, SearchBar, FilterPanel, ChatInterface, GraphViewer, ArtifactCard, CompetencyBadge, JobCard, CandidateCard, MetricCard | Composite components built from primitives | primitives, packages/types |
| `layouts/` | Container, Grid, Sidebar, SplitPanel, DashboardLayout | Layout components | primitives, design-tokens |
| `theme/` | Theme provider, CSS variable overrides per surface (learner, marketplace, employer, admin) | Theme context | design-tokens |
### packages/mock-data — Mock Data Layer
@@ -150,8 +95,7 @@ Composite/layout/theme components (navigation shell, tables, chat panels, graph
| `candidates.ts` | 15+ mock candidate profiles with artifacts, process traces, defense scores, microcredentials | Typed mock data | packages/types |
| `employers.ts` | 10+ mock employer profiles with logos, descriptions, open positions | Typed mock data | packages/types |
| `learner-progress.ts` | Mock learner progress data: active competencies, completion percentages, recent artifacts | Typed mock data | packages/types |
| `admin.ts` | Admin surface mock data: platform metrics, activity feed, system health, learner roster (admin view), moderation queues | Typed mock data | packages/types |
| `ai-scenarios.ts` | AI engine-input scenario IDs + display metadata for the learner agent panels; IDs string-identical to `ai_service/corpus/` (D-021) | Typed mock data | packages/types |
| `ai-tutor-responses.ts` | Pre-scripted AI tutor chat responses for Coach and Tutor agent mockups | Typed mock data | packages/types |
### packages/types — Shared Types
@@ -161,42 +105,11 @@ Composite/layout/theme components (navigation shell, tables, chat panels, graph
| `marketplace.ts` | Job, Employer, Candidate, JobPosting, TalentMatch, SearchFilter | Marketplace types | None |
| `user.ts` | Learner, Admin, EmployerUser, AgeGroup, Role | User types | None |
| `ui.ts` | Component props, theme config, breakpoint definitions | UI types | None |
| `telemetry.ts` | TelemetryEvent/ExecResult wire shapes for the live build surface (v0.3) | Engine types | None |
| `variants.ts` | Variant/TaskTemplate shapes for per-learner task statements (v0.3) | Engine types | None |
| `grading.ts` | GradeRecord/RubricScore shapes for live grading display (v0.3) | Engine types | None |
| `defense.ts` | DefenseSession/transcript/integrity-signal shapes for the defense surface (v0.3) | Engine types | None |
---
## Data Flow
### v0.3 credential flow (current)
```
[learner build surface /build/*] [learner defense surface /defend/*]
file CRUD + Run/Test (HTTP) mic MediaRecorder / typed + TTS playback
│ │
▼ ▼
[api/sandboxes files/exec] ──exec──▶ [namespace sandbox] [api/defense start/answer/finish]
│ │ capture agent │
│ ▼ (WS telemetry) ▼
│ [api/telemetry ingest] [DefenseStore (SQLite)]
│ │ SQLite │ transcript + integrity signals
│ ▼ │
│ [TraceStore] ────▶ [GradingEngine: digest (grading/features)
│ │ + rubric LLM (D-028)] ──▶ [GradeStore]
│ ▼ ▼
└──▶ Lab agent (live digest) Assessor (grade output) / Proctor (integrity)
Examiner agent (SSE) ◀── defense sessions
variants: [api/variants] ◀── [VariantStore (seeded, D-029)] ── per-learner task statements
```
- Lab consumes the live trace digest; Assessor consumes grading-engine output; Proctor consumes telemetry + defense integrity signals (REQ-3-007) — no mock fallback in the learner path (v0.2 corpus scenarios are dormant test-only fixtures).
- The learner's path is: variant task → in-sandbox build (telemetry streams to SQLite) → grade My Work (rubric scores from the real trace) → oral defense → verdict.
- Flooded/gapped traces are terminal: ingest closes 1008 and marks INCOMPLETE_FLOODED (G-3); the grader returns UNGRADABLE_TRACE_INCOMPLETE (G-4) — no credential from an incomplete trace.
### v0.2 chat flow (complete, still live)
```
[packages/mock-data + packages/types] [ai_service/corpus]
│ (TS, web surfaces) │ (Python, agent inputs)
@@ -211,31 +124,14 @@ Composite/layout/theme components (navigation shell, tables, chat panels, graph
(https://ollama.com/v1)
```
- Web surfaces remain server-component-first; client components (chat, filters, graph viewer, dark mode toggle) fetch directly from ai-service over SSE (A-002: no Next.js API-route proxy).
- The LLM provider layer is a dumb pipe — OpenAI-compatible chunks pass through byte-identically; envelope logic (meta/done/error) lives only in the API layer (D-016).
- The v0.2 corpus scenarios (`ai_service/corpus/`) are retained as dormant, test-only fixtures (dormant-header noted); they are no longer inputs to the live learner path.
- Web surfaces remain server-component-first; client components (chat, filters, graph viewer, dark mode toggle) fetch directly from ai-service over SSE (A-002: no Next.js API-route proxy in v0.2).
- The LLM provider layer is a dumb pipe — OpenAI-compatible chunks pass through byte-identical; envelope logic (meta/done/error) lives only in the API layer (D-016).
- Lab/Assessor/Proctor read mock scenarios from `ai_service/corpus/` — real engines are v0.3+.
- All automated tests use the deterministic mock provider; the cloud is for manual probes only.
---
## Build Order (v0.4)
1. **Bootstrap CLI core** — apps/cli package: doctor checks (node/pnpm/python3/git/unshare), bootstrap wrapper (pnpm install + scripts/bootstrap.sh + .env template + key validation), verify health check, dev passthrough; unit tests
2. **Binary build + release pipeline** — esbuild bundle → Node SEA binary (`nextcraft-linux-x64`) + sha256 sidecar; install.sh one-liner (Gitea latest-release resolve + checksum verify + PATH install); release-asset upload wired into the ship flow (ongoing binaries from v0.4 onward)
3. **Install docs + fresh-clone E2E** — README quickstart (one-liner → doctor → bootstrap → dev), CLI reference, fresh-clone end-to-end test proving a clean clone reaches a running stack
## Build Order (v0.3 — complete)
1. **Sandbox fabric** — SandboxBackend protocol + unshare namespace spawner + lifecycle manager (create/list/snapshot/destroy) + concurrency guard + per-sandbox workdir; isolation + resource-limit probes
2. **Live build telemetry** — TelemetryEvent models + SQLite TraceStore + WebSocket ingest endpoint + seq gap detection + in-sandbox capture agent
3. **Process-trace grading engine** — deterministic feature/digest computation + rubric scoring via LLM structured output + GradeStore; calibrated against v0.2 mock corpora
4. **Variant task generation** — template library + seeded LLM instantiation + VariantStore + difficulty normalization anchors
5. **Oral / voice defense** — VoiceProvider protocol + STT/TTS + mock + browser fallback + Examiner agent + transcript/integrity-signal capture
6. **Agent re-grounding + learner surface integration** — Lab/Assessor/Proctor consume real telemetry/grades/defense signals; learner sandbox mockup → real in-browser build/run (Run/Test buttons executing in a namespace sandbox, read-only exec-output panel — no interactive shell, CUT-2/G-8); assessment mockup → live defense + live grading
---
## Build Order (v0.2 — complete)
## Build Order (v0.2)
1. **AI service scaffolding** — apps/ai-service: FastAPI app, config, provider layer (ollama-cloud/local/mock), SSE chat endpoint, pytest harness, turbo integration
2. **Agent framework** — BaseAgent, registry, session store, structured output, prompts scaffolding, learner-context corpus
@@ -248,19 +144,15 @@ The v0.1 build order (monorepo → types → mock data → tokens → primitives
---
## Future Architecture (Post-v0.4, for reference)
## Future Architecture (Post-v0.2, for reference)
v0.4 delivers distribution (CLI + binary releases); later milestones fill in the remaining platform:
v0.2 delivers the ai-service skeleton that later milestones fill in:
- **In-memory sessions → PostgreSQL + Drizzle/SQLModel** — SessionStore + v0.3 TraceStore/GradeStore/VariantStore/DefenseStore protocols swap SQLite→Postgres with no API changes
- **userns subprocess sandboxes → containerd/runc backend** — D-024 `SandboxBackend` protocol swap; same lifecycle API
- **Coding-IDE sandbox → design tool + simulation environments** — REQ-F-021 full scope (v0.5)
- **Mock corpus → real engines** — Lab consumes real sandbox telemetry (v0.3 sandbox fabric); Assessor grades real process traces (v0.3 assessment engine); Proctor consumes real identity/attention signals (v0.3 identity verification)
- **In-memory sessions → PostgreSQL + Drizzle ORM** — SessionStore protocol swap, no API changes
- **Mock provider → per-agent model routing** — provider factory already selects by config; per-agent `AI_<AGENT>_MODEL` overrides
- **No auth → real KYC + sessions** — **deferred per founder directive; moved to v0.5 with D-016**; REQ-F-017 identity/age-gating lands post-v0.4. Age-gating remains the v0.1 visual flow mockup
- **Mock voice → real server STT/TTS (openai-audio provider)** — CUT-1/G-7 seam moved to v0.5 per D-016; VoiceProvider protocol is the drop-in point
- **linux x64 binary → macOS arm64 + auto-update** — D-036 defers non-linux targets (unverifiable on this box); `nextcraft upgrade` (self-replace from latest release) is the natural v0.5+ follow-up
- **Exec-telemetry seq-lease / replay-margin fix** — the P6-lesson one-line ACK gap moves to v0.5 per D-016
- **No auth → real KYC + sessions** — A-008 dropped in v0.3 when identity verification lands
- **No search → Semantic vector search (pgvector)** — Filter UI replaced with vector similarity search
- **No payments → Payment processing** — Pricing page replaced with real subscription/payment flows
The monorepo structure (apps/web + apps/ai-service + apps/cli + packages/*) accommodates further apps without restructuring.
The monorepo structure (apps/web + apps/ai-service + packages/*) accommodates further apps without restructuring.
+8
View File
@@ -0,0 +1,8 @@
{
"phase": 2,
"stage": "verify",
"milestone": "v0.2",
"phase_role": "execution",
"attempts": 0,
"updated_at": "2026-09-11T17:35:00Z"
}
+23 -73
View File
@@ -1,86 +1,36 @@
# Nextcraft v0.3 — GRILL.md (Adversarial Review Verdict)
# Nextcraft v0.2 — GRILL.md (Adversarial Review Verdict)
**Stage:** GRILL, Phase 0 pre-execution · **Verdict:** GO-WITH-CHANGES · **Confidence:** 0.72
## Summary
The credential-pipeline architecture (telemetry → trace → grade → defense) is sound and correctly sequenced. Three plan claims did NOT survive contact with this box and were correct before execution. The central problem: the plan **overstated sandbox resource-limit enforcement** and deferred KYC **without closing the resulting no-auth local abuse vector**. Fixed via binding decisions G-1..G-6 + scope cuts CUT-1/CUT-2 — no redesign required.
**Stage:** GRILL, Phase 0 pre-execution · **Verdict:** GO-WITH-CHANGES · **Confidence:** 0.82
## Per-Axis Findings
| Axis | Verdict | Rationale |
|------|---------|-----------|
| Feasibility | CONCERN | Core `unshare` userns/mount/pid/net isolation probe-verified (uid=0 in-ns, network isolated, writes contained). But rlimit enforcement is partial: `RLIMIT_NPROC` scopes to the real host uid (5 sandboxes share one pids budget) and no disk-quota tool exists on the box. |
| Over-scoping | CONCERN | 39 tasks across 5 new subsystems + learner-surface rewrite for a solo founder. Voice-real-path and the interactive terminal relay are separable from the pipeline proof → cut (CUT-1, CUT-2). |
| Architecture risk | CONCERN | SandboxBackend/TraceStore protocols are the right seams. Overclaimed "limits enforced" + WS ingest "drop-oldest on unbounded growth" contradicted the at-least-once grading guarantee. |
| Phase sequencing | PASS | P1 sandbox → P2 telemetry → P3 grading → P4 variants → P5 voice → P6 integration is a correct dependency DAG. |
| Verification honesty | FAIL (fixed) | Must-Haves asserted "resource limits enforced + observable" (CPU/memory/disk/time quotas) the named mechanism cannot satisfy; probe tests would pass while the guarantee was false. Corrected by G-1. |
| Cost/quota | CONCERN | LLM/voice mock-gated (good). No-auth sandbox creation + unbounded disk + shared NPROC let one learner starve others at zero cost → closed by G-5. |
| Milestone honesty | CONCERN (fixed) | Release note disclosed KYC deferral + IDE-only but was silent on partial resource-limit enforcement → G-6. |
| Security | FAIL (fixed) | `POST /v1/sandboxes {learner_id}` client-supplied over localhost CORS let any local process mint sandboxes/flood/exhaust shared resources → G-5 abuse control ships despite KYC deferral. |
| Operability | CONCERN (advisory) | In-memory sandbox registry loses handles on restart (orphaned namespaces) → a-1 startup reaper. |
| Feasibility | PASS | Environment claims verified (venv works, all 8 pinned packages on PyPI for py3.11, port 8420 free, secrets gitignored). Stall points pre-mitigated (idempotent bootstrap, 300s read timeout, idempotent abort; reactStrictMode confirmed in next.config). |
| Over-scoping | CONCERN (mild) | 4-layer JSON defense justified; 500-cap LRU is over-spec but cheap — keep, extend no further. Real gaps: no Python lint (fixed via G-3), cache dirs in gitignore (fixed), script path resolution (advisory c). |
| Architecture risk | CONCERN | sse-starlette `: ping` frames invisible to short TestClient streams — parser gap fixed via G-1. Agent registration had two contradictory patterns — fixed via G-4. Cloud outage blast radius contained by mock-only tests. |
| Phase sequencing | PASS | Integration-last correct: P1 freezes the SSE contract before client code exists; A-002 eliminates dev-server buffering trap. P4 heaviest but mechanical. |
| Verification honesty | CONCERN | Persona distinctness circular against self-authored mocks (disclosed; cloud probe optional). P6 must-haves manual-only. D-021 ID alignment unmechanized (advisory b). Disclosed honestly. |
| Cost/quota | PASS | Cloud burn bounded (~<100K tokens milestone-wide, manual probes only). Mock-only rule structurally enforced; mechanical guard via advisory (a). |
| Milestone honesty | PASS (conditional) | Mock-engine caveats present everywhere that matters. Release note content now bound by G-5; dead `aiTutorResponses` disposal bound by G-5. |
## Binding Decisions (applied to PLAN.md/ARCHITECTURE-adjacent docs/REQUIREMENTS.md/ROADMAP.md/PROJECT.md)
## Binding Decisions (applied to PLAN.md/ROADMAP.md)
- **G-1 (BINDING) — Resource-limit claims match the deliverable mechanism.** Memory (RLIMIT_AS) + CPU (RLIMIT_CPU) + single-file (RLIMIT_FSIZE) + wall-clock reaper are kernel-enforced; per-sandbox pids and hard disk quota are NOT kernel-enforceable without cgroup delegation/sudo → documented as accepted v0.3 risk. Applied to REQ-3-002, PLAN P1 Must-Haves + Task 1-2-02, ROADMAP P1 criteria.
- **G-2 (BINDING) — Disk cap via manager workdir-size sweep.** `AI_SANDBOX_MAX_WORKDIR_MB` (default 512MB); sweep snapshots+destroys over-cap sandboxes and logs an integrity signal; closes the unbounded-`dd` hole. Applied to PLAN Task 1-2-01 + 1-2-02(e) + P1 Must-Haves.
- **G-3 (BINDING) — Telemetry flood control WITHOUT silent drop.** Bounded queue; on overflow or >`AI_TELEMETRY_MAX_EVENTS_PER_TASK` (default 50k) → WS close 1008 + trace marked `INCOMPLETE_FLOODED` (Proctor signal). Silent drop-oldest forbidden (corrupts grading). Applied to PLAN Task 2-2-02 + P2 Must-Haves.
- **G-4 (BINDING) — Grader refuses incomplete/gapped traces.** `grade()` gates on `TraceStore.gaps()` + `INCOMPLETE_FLOODED` → returns `verdict=UNGRADABLE_TRACE_INCOMPLETE`; no credential from a gapped trace. Applied to PLAN Task 3-2-01 + P3 Must-Haves.
- **G-5 (BINDING) — No-auth abuse control at MVP scale.** Per-learner sandbox cap + global create-rate cap (429) + server-side `learner_id` allowlist (403) so the unauthenticated surface can't exhaust shared NPROC/disk. Ships WITH the milestone even though KYC is deferred. Applied to PLAN Task 1-3-01 + PROJECT A-110.
- **G-6 (BINDING) — Disclose partial enforcement in the P7 release note.** Item (e): which limits are kernel-enforced vs best-effort, and that full enforcement is deferred to the post-MVP containerd backend. Applied to PLAN P7 release-note honesty block.
- **G-1 (BINDING):** SSE frame parser in Task 6-1-01 must ignore frames with no `data:` lines (sse-starlette ping keep-alive). Added to Action + Phase 6 must-have.
- **G-2 (BINDING):** P6 end-to-end verification may run with `AI_PROVIDER=mock` fallback — real ai-service over HTTP is the requirement; provider choice is service-internal. Prevents cloud outage blocking P6.
- **G-3 (BINDING):** ruff (check-only) added: Task 1-1-04, pyproject dev extra, scripts/lint.sh, root `ai:lint`, turbo `ai#lint`, Phase 1 must-have.
- **G-4 (BINDING):** Mentor/Proctor registered centrally in `registry.py` (single registration pattern), matching P3/P4.
- **G-5 (BINDING):** v0.2.0 release note must state Lab/Assessor/Proctor run on mock engine inputs (v0.3+ for real); dead `aiTutorResponses` export disposed of in P7.
## Scope Cuts (accepted — preserve the end-to-end credential pipeline)
## Advisory (non-binding; applied where cheap)
- **CUT-1 (G-7) — Real server STT/TTS (`OpenAIAudioProvider`) deferred to v0.4.** Voice is mock-first (D-030); the `/audio/*` real path can never run in CI and was the least-verifiable surface. v0.3 proves the full defense *dialogue* + integrity-signal pipeline over mock + browser-native fallback; the `VoiceProvider` protocol is the future drop-in seam. Applied to PLAN Phase 5 Goal + Task 5-1-01/5-4-01 + P5 Must-Haves.
- **CUT-2 (G-8) — Interactive xterm.js shell relay deferred to v0.4.** The credential pipeline needs *process events* (Run/Test + file edits), not a live keystroke-level shell — the most fragile real-time piece, unverifiable without a real terminal. The build panel becomes Run/Test buttons + read-only exec output render; `@xterm/*` is NOT a v0.3 dependency. Applied to PLAN Env-facts, P6 Goal, Task 6-2-01/6-2-02/6-3-01, P6 Must-Haves, MVP/UX sections.
- Variants (Phase 4) and the Examiner agent dialogue **KEPT** — both are on the credential critical path (anti-collusion + the defense dialogue).
## Advisory (applied)
- **a-1** Startup reaper: on lifespan boot, scan `AI_SANDBOX_DIR`, reap workdirs whose recorded pid is dead, log a warning. → Task 1-2-01.
- **a-2** `RLIMIT_FSIZE` (~50MB) as a cheap partial single-file disk guard in the spawner's `preexec_fn`. → Task 1-2-01 + 1-2-02(c).
- **a-3** SQLite `PRAGMA journal_mode=WAL` + `synchronous=NORMAL` at engine creation (avoids `database is locked` under concurrent ingest + grader reads). → Task 2-1-02.
- **a-4** Grading prompt note: treat high edit/command churn with no test-progress as a process-quality negative (softens digest-gaming naivety). → Task 3-2-01.
- **a-5** Variant fairness envelope: two variants of one template must compute digests within the template's expected feature envelope ("same bar" is testable). → Task 4-2-01 Verify.
- (a) conftest asserts provider is MockProvider — **applied in P1 implementation**
- (b) pytest reads packages/mock-data as text, asserts corpus IDs appear — **applied in P4 implementation**
- (c) scripts resolve repo root via script-relative dirname — **folded into Task 1-1-04/1-1-03**
- (d) `.pytest_cache/` + `.ruff_cache/` in .gitignore — **applied**
- (e) keep LRU as-is; no further session-store sophistication in v0.2
- (f) Task 6-2-02 REQ tag fixed to REQ-2-011 (routing) — **applied**
## Outcome
**GO** — all six binding decisions and both scope cuts applied to PLAN.md / REQUIREMENTS.md / ROADMAP.md / PROJECT.md before Phase 1 execution. No axis requires escalation (all resolvable at confidence ≥ 0.85). The milestone no longer claims resource enforcement it cannot deliver, and the no-auth abuse vector is closed at MVP scale.
---
# Nextcraft v0.4 — GRILL.md (Adversarial Review Verdict)
**Stage:** GRILL, Phase 0 pre-execution · **Verdict:** GO-WITH-CHANGES · **Confidence:** 0.83
## Summary
The distribution milestone is small, founder-directed (D-016, confidence 0.99), and additive (zero changes to the running credential pipeline). The plan's central risk: **Node SEA was probe-verified as a flag, not as a working build** — the v0.3 lesson (A-101: probe the mechanism, not the existence) applies. Second gap: a binary whose `--version` lies (stale package.json) would poison the "ongoing binaries" contract. Third: sed-based JSON parsing in install.sh is a fragility + integrity risk. Fourth: "ongoing binaries" has no enforcement mechanism beyond prose. All four closed by binding decisions G-101..G-104 below. No scope cuts required — the milestone is already minimal.
## Per-Axis Findings
| Axis | Verdict | Rationale |
|------|---------|-----------|
| Business case | PASS | Founder directive explicit + recorded (D-016). Evidence of need: live Gitea probe shows latest release v0.2.8 with ZERO assets; bootstrap requires repo archaeology (scripts found only via package.json spelunking). |
| Scope | PASS | 5 REQs, 3 execution phases, one focused surface (apps/cli + scripts). Smallest milestone yet. macOS arm64 already cut (D-036, unverifiable here). |
| Feasibility | CONCERN (fixed) | SEA flag exists on node v24.15.0, but no end-to-end SEA binary was built during RESEARCH. postject availability assumed (`npx postject` — needs npm registry reachability, unproven). Zipapp fallback requires python3 on target — an honest-degradation ladder, not a silent downgrade. → G-101. |
| Honest versioning | CONCERN (fixed) | `--version` from package.json would print a stale hardcoded version inside a per-release binary — breaks upgrade detection + the one-liner's re-run-to-upgrade promise. → G-102. |
| Install integrity | CONCERN (fixed) | sed/grep JSON parsing is brittle; a parse failure must never fall through to installing an unverified artifact. Exact asset-name matching + hard-degrade to source instructions. → G-103. |
| Sequencing | PASS | P1 CLI (source-runnable) → P2 binary+pipeline → P3 docs+E2E matches dependency order; each phase ships independently. |
| Cost/quota | PASS | Zero new paid infra; binaries built on-box; Gitea releases free. Dev-only esbuild dep. |
| Risks | CONCERN (fixed) | Top 3: SEA end-to-end (→ G-101 live probe FIRST in P2), npm registry reachability for esbuild (→ proven by P1's pnpm install must-have), Gitea asset-upload token scope (→ live-proven at the v0.3.2 ship itself). |
| Adoption/operability | PASS | Consumer = founder + future pilots; one command replaces README archaeology. Rollback trivial (rm ~/.local/bin/nextcraft). No server changes. |
## Binding Decisions (applied to PLAN.md)
- **G-101 (BINDING) — SEA live-build probe is the FIRST P2 action.** Task 2-1-01 builds a real binary before anything depends on it; the build script encodes the fallback ladder explicitly (SEA → zipapp with "requires python3" honesty). If SEA fails on this box, zipapp becomes primary with the docs stating the requirement — no silent claim of node-less operation.
- **G-102 (BINDING) — Version stamping at build time.** `build-binary` accepts the shipping tag and stamps it into the bundle (`NEXTCRAFT_VERSION` replace); `--version` prints it; install E2E asserts the installed binary reports the tag it was downloaded from. A binary may never report a version it was not built as.
- **G-103 (BINDING) — Install-script integrity hard-degrade.** install.sh matches assets by EXACT name (`nextcraft-linux-x64`, `nextcraft-linux-x64.sha256`); any parse/lookup/download failure degrades to source-bootstrap instructions (exit 0) — never installs unverified or name-approximate artifacts. Checksum mismatch = hard stop, exit 1, explicit do-not-run message. dash-safe POSIX sh, no jq.
- **G-104 (BINDING) — Ongoing-binaries enforcement.** Every ship from v0.3.2 onward MUST run `scripts/release-assets.sh <tag>` after tag+merge (best-effort, non-blocking, `release_pending` escalation on failure — but attempted + logged every release). The final-phase audit gate includes "milestone release carries both assets" as a check. This makes the founder's "ongoing binaries" directive a pipeline property, not prose.
## Escalations
None. All four concerns resolved at confidence ≥ 0.85. No axis requires founder escalation (directive already explicit).
## Outcome
**GO** — G-101..G-104 applied to PLAN.md before Phase 1 execution. The milestone claims only what its probes prove, and the ongoing-binaries contract has an enforcement mechanism.
GO — all five binding decisions applied to PLAN.md/ROADMAP.md/.gitignore/ARCHITECTURE.md before Phase 1 execution. No axis requires escalation.
+76 -118
View File
@@ -2,19 +2,17 @@
## Persona Roster
> **v0.4 update (RESEARCH, lead-developer assessment):** milestone pivoted to Distribution & Bootstrap CLI (founder directive D-016). New custom persona **cli-engineer** (domain `cli`) owns apps/cli end-to-end: doctor/bootstrap/verify/dev commands, checks, spawn wrappers, the SEA binary build, the one-liner install script, and the release-asset pipeline. backend-engineer retains the scripts/ + turbo/root-package integration surface. **sandbox-engineer and voice-engineer deactivated** (their v0.3 code is complete and untouched this milestone — reason fields below). ai-engineer light-touch (no model-facing work in v0.4). **security-auditor re-activated (phase-specific)** for the install pipeline: curl|bash attack surface, checksum trust, PATH writes, secrets handling in the release flow. frontend-engineer/design-system-engineer/data-engineer inactive (zero UI/data-scope tasks in v0.4 — retained below with reasons).
### lead-developer
```yaml
active: true
phase_specific: false
reason: Coordinates task decomposition across CLI, scripts, release-pipeline, and docs territories; resolves cli-engineer/backend-engineer boundary (scripts vs CLI)
reason: Coordinates task decomposition across web, AI service, and data territories; resolves conflicts between frontend, backend, and AI personas
domain: coordination
frameworks:
- next.js
- turborepo
- pnpm
- node
- fastapi
constraints:
- pragmatic
- battle-tested defaults
@@ -27,83 +25,85 @@ territory:
- "apps/ai-service/pyproject.toml"
```
### cli-engineer
### frontend-engineer
```yaml
active: true
phase_specific: false
reason: v0.4 custom persona (RESEARCH) — owns the distribution milestone core: nextcraft CLI (doctor/bootstrap/verify/dev), pure check logic, spawn wrappers with timeouts, Node SEA binary build (D-033), one-liner install.sh (D-035), checksum sidecar, and Gitea release-asset upload (D-036)
domain: cli
reason: Phase 6 learner-surface integration — useChatStream hook, agent switcher, streaming/error/loading states, Lab/Assessor/Proctor output panels. Owns all page components, layouts, and surface-specific UI.
domain: frontend
frameworks:
- node
- typescript
- node:test
- esbuild
- node-sea
- posix-sh
- react
- next.js
- tailwindcss
- lucide-react
- recharts
- react-flow
constraints:
- stdlib-only-runtime (no runtime npm deps; esbuild dev-only)
- thin-wrapper (never re-implement scripts/bootstrap.sh or dev.sh — compose via spawn, A-202/A-209)
- timeout-every-spawn (no unbounded subprocess)
- actionable-errors (every failed check tells the user how to fix it)
- graceful-degradation (install never hard-fails; source-bootstrap fallback, A-206)
- checksum-before-install (sha256 verify before chmod+install, A-207)
- secrets-never-in-cli (no key generation; .env.example -> .env copy only, A-210)
- fail-loud-exit-codes (0 ok / 1 failure / 2 usage)
- component-first
- server-components-default
- minimal-client-js
- sse-client-buffering (buffer bytes, split frames on \n\n, join data: lines)
- abortcontroller-cleanup (idempotent abort in effect cleanup)
- responsive-all-breakpoints
- dark-mode-support
territory:
- "apps/cli/**"
- "scripts/install.sh"
- "scripts/release-assets.sh"
- "apps/web/**"
- "packages/ui/**"
- "packages/mock-data/**"
- "packages/types/**"
```
### data-engineer
```yaml
active: true
phase_specific: false
reason: Owns TS mock data layer schema and typed definitions. Does NOT own the Python corpus (that is ai-engineer territory) — the two are aligned by documented convention (D-021).
domain: data
frameworks:
- typescript
constraints:
- schema-first
- type-safe
- migration-ready
- mock-data-only
territory:
- "packages/types/**"
- "packages/mock-data/**"
```
### backend-engineer
```yaml
active: true
phase_specific: false
reason: Owns the script + monorepo integration surface the CLI composes: apps/ai-service/scripts/*, root package.json cli:* passthrough scripts, turbo task wiring (D-037/D-022). Python ai-service itself is untouched this milestone (v0.3 complete).
reason: Reactivated for v0.2 — owns apps/ai-service infrastructure: FastAPI app, settings, SSE plumbing, endpoints, monorepo/turbo integration, test harness.
domain: backend
frameworks:
- fastapi
- bash
- turborepo
- pnpm
- uvicorn
- pydantic
- httpx
- pytest
constraints:
- scripts-are-truth (bootstrap.sh/dev.sh stay the single source of bootstrap orchestration; CLI only wraps)
- idempotent-scripts (re-runnable without side effects)
- secrets-via-env-only (D-014; dev.sh exports from .ciagent/.env.secrets)
- provider-agnostic-boundaries (llm/ imports nothing from agents/ or api/)
- streaming-first
- no-database-v0.2
- secrets-via-env-only
- mock-provider-in-tests
territory:
- "apps/ai-service/ai_service/main.py"
- "apps/ai-service/ai_service/config.py"
- "apps/ai-service/ai_service/api/**"
- "apps/ai-service/scripts/**"
- "apps/ai-service/package.json"
- "package.json"
- "apps/ai-service/tests/api/**"
- "turbo.json"
- ".gitignore"
```
### security-auditor
```yaml
active: true
phase_specific: true
reason: v0.4 re-activated (phase-specific) — the install pipeline is the first externally-consumed attack surface: curl|bash piping, latest-release resolution, checksum trust root, PATH writes to ~/.local/bin, download tempdir hygiene, release-asset upload token handling. No KYC/PII work (still v0.5).
domain: security
frameworks:
- posix-sh
- curl
- sha256sum
constraints:
- STRIDE-classified
- no-pipe-to-shell-without-checksum (download -> verify -> install order)
- tmpdir-safe (mktemp, no predictable paths, trap cleanup)
- token-never-echoed (release upload resolves .env* only, never logs)
territory:
- "scripts/install.sh"
- "scripts/release-assets.sh"
- "apps/cli/src/lib/spawn.ts"
```
### ai-engineer
```yaml
active: true
phase_specific: false
reason: Light-touch v0.4no model-facing work in the distribution milestone; retained to guard the CLI against touching agent/engine boundaries and to keep territory mappings accurate for v0.5 (voice real-path, seq-lease).
reason: Custom persona for v0.2owns the LLM provider layer, agent framework, prompt library, structured outputs, and mock corpora for the six tutor agents.
domain: ai
frameworks:
- pydantic
@@ -111,101 +111,59 @@ frameworks:
- pytest
constraints:
- provider-agnostic-protocol
- prompts-are-code
- json-defensive-parsing
- never-call-cloud-in-tests
- delta-passthrough
territory:
- "apps/ai-service/ai_service/llm/**"
- "apps/ai-service/ai_service/agents/**"
- "apps/ai-service/ai_service/prompts/**"
```
### frontend-engineer
```yaml
active: false
phase_specific: false
reason: v0.4 has zero UI-scope work (no web/pages/components changes planned in the distribution milestone); v0.3 surfaces are complete. Reactivated at v0.5 when deferred UX work resumes.
domain: frontend
frameworks:
- react
- next.js
- tailwindcss
constraints:
- component-first
- server-components-default
territory:
- "apps/web/**"
- "packages/ui/**"
- "apps/ai-service/ai_service/corpus/**"
- "apps/ai-service/tests/llm/**"
- "apps/ai-service/tests/agents/**"
```
### design-system-engineer
```yaml
active: false
active: true
phase_specific: false
reason: No design-token or primitive work in v0.4; roster retained for v0.5.
reason: Owns the shared component library, design tokens, and visual consistency. Light duty in v0.2: Phase 6 may need new primitives (agent-switcher control, stream-status indicator, error toast variant).
domain: frontend
frameworks:
- tailwindcss
- storybook
- lucide-react
constraints:
- design-token-driven
- wcag-aa-contrast
- dark-mode-required
- consistent-across-surfaces
territory:
- "packages/ui/**"
```
### data-engineer
### security-auditor
```yaml
active: false
phase_specific: false
reason: No schema/mock-data work in v0.4; types packages untouched. Reactivated if CLI surfaces need shared types (not planned — CLI is self-contained).
domain: data
frameworks:
- typescript
constraints:
- schema-first
- type-safe
territory:
- "packages/types/**"
- "packages/mock-data/**"
```
### sandbox-engineer
```yaml
active: false
phase_specific: false
reason: v0.3 persona — sandbox fabric shipped complete (v0.2.x series); v0.4 touches no sandbox code. doctor only *checks* unshare availability; no sandbox logic changes. Reactivated at v0.5 (design/sim environments).
domain: infra
frameworks:
- python
- linux-namespaces
constraints: []
territory:
- "apps/ai-service/ai_service/sandbox/**"
```
### voice-engineer
```yaml
active: false
phase_specific: false
reason: v0.3 persona — voice defense shipped complete (mock-first, CUT-1); real server STT/TTS moved to v0.5 per D-016. No v0.4 voice work.
domain: ai-media
reason: No auth in v0.2 (A-008); CORS is localhost-only; no real user data. Security review handled by verifier's STRIDE analysis layer plus a Phase 7 checklist item: secrets hygiene (key absent from code/logs/commits/errors), localhost-only CORS, no PII in prompts.
domain: security
frameworks: []
constraints: []
territory:
- "apps/ai-service/ai_service/voice/**"
territory: []
```
## Phase-Specific Personas
| Persona | Phases | Removed After |
|---------|--------|---------------|
| security-auditor | 2 (primary: install pipeline), 3, 4 (final review) | milestone complete |
All other personas span the milestone. Deactivated personas receive no tasks.
None for v0.2. All active personas span the entire milestone. data-engineer and design-system-engineer are light-touch after Phase 1.
## Territory Conflict Resolution
| Conflict | Resolution |
|----------|------------|
| cli-engineer vs backend-engineer (scripts/) | backend-engineer owns `apps/ai-service/scripts/**` + root `package.json`/`turbo.json` wiring; cli-engineer owns `apps/cli/**` + top-level `scripts/install.sh` + `scripts/release-assets.sh` and *consumes* backend scripts via spawn — never edits them |
| cli-engineer vs security-auditor (install.sh) | cli-engineer implements; security-auditor reviews + may patch security defects directly in install.sh/spawn.ts (its territory) |
| lead-developer vs any | lead-developer coordinates only, does not directly modify code files |
| frontend-engineer vs data-engineer (packages/types, packages/mock-data) | data-engineer owns type definitions and mock data schema; frontend-engineer consumes them. If changes needed, data-engineer updates types first. |
| frontend-engineer vs design-system-engineer (packages/ui) | design-system-engineer owns design tokens and primitive components; frontend-engineer owns composite components and page-level UI. |
| ai-engineer vs data-engineer (mock data duplication) | ai-engineer owns `ai_service/corpus/` (Python); data-engineer owns `packages/mock-data` (TS). Shared entity IDs and shapes kept aligned by documented convention (D-021): cross-referencing file headers, identical `comp-*` ID strings. |
| backend-engineer vs ai-engineer (apps/ai-service) | backend-engineer owns app shell, config, API endpoints, scripts, and test harness; ai-engineer owns llm/, agents/, prompts/, corpus/. Boundary: `ai_service/api/` (backend) composes `ai_service/agents/` (AI) via DI — agents never import api/. |
| lead-developer vs any | lead-developer coordinates only, does not directly modify code files. |
+406 -194
View File
@@ -1,210 +1,422 @@
# Nextcraft v0.4 — PLAN.md
# Nextcraft v0.2 — PLAN.md
## Overview
This plan covers execution phases 13 of milestone v0.4 (Distribution & Bootstrap CLI) plus the final phase (P4 review+ship). The milestone delivers the founder directive (D-016): a streamlined install for Nextcraft — a `nextcraft` bootstrap CLI shipped as a linux x64 binary, installed via a one-liner script, with binaries published on **every ongoing release** from v0.4 onward. Phases are strictly sequential (P1→P3); within each phase, Wave 1 tasks are parallelizable (no cross-file dependencies) and later waves depend on earlier ones.
This plan covers execution phases 1-6 of milestone v0.2 (AI Tutor Architecture): the six AI tutor agents (Coach, Tutor, Lab, Assessor, Proctor, Mentor) as real LLM-backed services in a new `apps/ai-service` Python FastAPI application, wired into the existing v0.1 learner surface with streaming responses. Phases are strictly sequential (P1→P6); within each phase, Wave 1 tasks are parallelizable (no cross-file dependencies) and later waves depend on earlier ones.
**Environment facts (probe-verified, apply throughout):** Go MISSING, Rust MISSING, gcc 12.2 present, **node v24.15.0 x64 linux (SEA-capable)**, python3 3.11.2, `shasum` 6.02, pnpm 12.3.4 via corepack, turborepo 2.3.3, tsx 4.23 in root devDeps path. Gitea API verified live at `https://git.coreci.dev/api/v1` (latest release v0.2.8, **zero assets** — the gap this milestone closes). Existing orchestration: `apps/ai-service/scripts/bootstrap.sh` (idempotent venv+pip incl. the no-ensurepip get-pip path), `apps/ai-service/scripts/dev.sh` (secrets export → uvicorn :8420), `apps/ai-service/.env.example` (full AI_* template). Root scripts: `ai:dev/ai:test/ai:bootstrap/ai:lint` turbo passthroughs (D-022 pattern to mirror as `cli:*`). Secrets live only in gitignored `.ciagent/.env.secrets` (GITEA_TOKEN, OLLAMA_API_KEY, OLLAMA_BASE_URL) — never in code, commits, or logs; tests never call the cloud or the forge (mocks/fixtures only).
**Milestone type:** feature. Tags: phase 0 → **v0.3.0**, P1 → v0.3.1, P2 → v0.3.2, P3 → v0.3.3, final phase P4 → **v0.3.4 = milestone release**. **GRILL binding decisions (this revision):** G-101 — SEA live-build probe is the FIRST P2 action (mechanism, not flag, must be proven); fallback ladder encoded honestly (zipapp requires python3 on target). G-102 — binary `--version` stamped from the shipping tag at build time (never a stale package.json version); install E2E asserts the installed binary reports its release tag. G-103 — install.sh matches assets by exact name; any parse/download failure degrades to source-bootstrap instructions (exit 0), never installs unverified artifacts; checksum mismatch = hard stop exit 1. G-104 — every ship from v0.3.2 onward runs `scripts/release-assets.sh <tag>` (best-effort, logged, non-blocking); the P4 audit gate checks the milestone release carries both assets.
**Environment facts (apply throughout):** Python via `python3 -m venv` (no system pip, no uv); pnpm 12.3.4 via corepack; ai-service port **8420**; default model `gemma4:31b` (config via `AI_TUTOR_MODEL`); ollama-cloud base `https://ollama.com/v1` (OpenAI-compatible, Bearer auth); API keys live only in gitignored `.ciagent/.env.secrets`, exported by `scripts/dev.sh` — never in code, commits, or logs; all automated tests use the deterministic mock provider and **never call the cloud**.
| Phase | Name | Requirements | Waves | Personas |
|-------|------|-------------|-------|----------|
| 1 | Bootstrap CLI core | REQ-4-001, REQ-4-002 | 3 | cli-engineer, backend-engineer, security-auditor (W3 review) |
| 2 | Binary build + release pipeline | REQ-4-003, REQ-4-004 | 3 | cli-engineer, backend-engineer, security-auditor |
| 3 | Install docs + fresh-clone E2E | REQ-4-005 | 2 | cli-engineer, backend-engineer |
| 4 | Final review + ship | — | 1 | all reviewers |
| 1 | AI service scaffolding | REQ-2-001, 002, 003 | 3 | backend-engineer, ai-engineer |
| 2 | Agent framework | REQ-2-004 | 2 | ai-engineer, backend-engineer |
| 3 | Coach + Tutor agents | REQ-2-005, 006 | 3 | ai-engineer, backend-engineer |
| 4 | Lab + Assessor agents | REQ-2-007, 008 | 4 | ai-engineer, backend-engineer, data-engineer |
| 5 | Proctor + Mentor agents | REQ-2-009, 010 | 3 | ai-engineer, backend-engineer |
| 6 | Learner surface integration | REQ-2-011, 012 | 3 | frontend-engineer, design-system-engineer |
---
## Phase 1: AI Service Scaffolding
**Requirements:** REQ-2-001, REQ-2-002, REQ-2-003
**Goal:** apps/ai-service runs under uvicorn, /health responds, provider-agnostic LLM layer with ollama-cloud/local/mock providers, SSE chat streaming verified, pytest suite green with mock provider, turbo integration wired
### Wave 1: Service shell + LLM core (parallel — no shared files)
#### Task 1-1-01: FastAPI app scaffolding
- **Persona:** backend-engineer — **REQ:** REQ-2-001
- **Files:** `apps/ai-service/pyproject.toml`, `apps/ai-service/ai_service/main.py`, `apps/ai-service/ai_service/config.py`, `apps/ai-service/ai_service/__init__.py`, `apps/ai-service/tests/conftest.py`, `apps/ai-service/tests/test_health.py`, `apps/ai-service/.env.example`, `apps/ai-service/README.md`
- **Action:** pydantic-settings `Settings` (env_prefix `AI_`, env_file, `SecretStr` key, port 8420, provider select, `AI_TUTOR_MODEL` default `gemma4:31b`). FastAPI app factory in `main.py` with lifespan stub (httpx client pool comes in Wave 2), CORS localhost-only (A-008), `GET /health`. pyproject with pinned deps (fastapi, uvicorn, pydantic, pydantic-settings, httpx, sse-starlette, pytest, pytest-asyncio) and dev extra. conftest: settings override + TestClient fixture. `.env.example` documents all `AI_*` vars; README documents venv setup and dev workflow.
- **Verify:** `scripts/bootstrap.sh && scripts/test.sh` — test_health passes; `curl localhost:8420/health` returns 200
#### Task 1-1-02: LLM types, protocol, mock provider
- **Persona:** ai-engineer — **REQ:** REQ-2-002
- **Files:** `apps/ai-service/ai_service/llm/types.py`, `apps/ai-service/ai_service/llm/base.py`, `apps/ai-service/ai_service/llm/mock.py`, `apps/ai-service/ai_service/llm/__init__.py`
- **Action:** `types.py`: pydantic `Message` (role/content), `ChatDelta` (OpenAI-compatible chunk shape). `base.py`: `LLMProvider` protocol — async `stream_chat(messages, model, response_format=None) -> AsyncIterator[ChatDelta]`; the provider is a dumb pipe, no envelope logic (D-016 keeps envelope in API layer). `mock.py`: deterministic scripted provider (hash-seeded token streams, scripted failure modes: connect error, mid-stream error, malformed JSON) for tests and CI.
- **Verify:** mock provider importable and deterministic; two identical calls yield identical streams
#### Task 1-1-03: Monorepo integration (shim + turbo + scripts)
- **Persona:** backend-engineer — **REQ:** REQ-2-001
- **Files:** `apps/ai-service/package.json`, `apps/ai-service/scripts/bootstrap.sh`, `apps/ai-service/scripts/dev.sh`, `apps/ai-service/scripts/test.sh`, `turbo.json` (update), `package.json` (root, update)
- **Action:** zero-dependency shim `package.json` in apps/ai-service with `dev`/`test`/`bootstrap` script entries. Turbo passthrough tasks `ai#dev`, `ai#test`, `ai#bootstrap` (`cache: false`, `outputs: []`). Root scripts `ai:dev`, `ai:test`, `ai:bootstrap`. `bootstrap.sh`: idempotent `python3 -m venv .venv` + pip install. `dev.sh`: exports keys from `.ciagent/.env.secrets` → uvicorn on 8420. `test.sh`: pytest via venv.
- **Verify:** `corepack pnpm install && pnpm ai:bootstrap && pnpm ai:test` runs pytest through turbo; re-running bootstrap is a no-op
#### Task 1-1-04: Python lint (ruff, check-only)
- **Persona:** backend-engineer — **REQ:** REQ-2-001
- **Files:** `apps/ai-service/pyproject.toml` (update: `[tool.ruff]` config + `ruff` in dev extra), `apps/ai-service/scripts/lint.sh`, `package.json` (root, update), `turbo.json` (update)
- **Action:** Add `ruff` (check-only, no formatter) to the dev extra; `[tool.ruff]` with line-length 100, target py311. `scripts/lint.sh`: `.venv/bin/ruff check .` with repo-root path resolution via script-relative dirname (not CWD). Root script `ai:lint`, turbo passthrough `ai#lint` (`cache: false`, `outputs: []`). Run over the entire ai-service tree; fix all findings before phase ship (G-3).
- **Verify:** `pnpm ai:lint` exits 0 on the Phase 1 codebase
### Wave 2: Real providers + SSE endpoint (depends on Wave 1)
#### Task 1-2-01: OpenAI-compatible provider + factory
- **Persona:** ai-engineer — **REQ:** REQ-2-002
- **Files:** `apps/ai-service/ai_service/llm/openai_compat.py`, `apps/ai-service/ai_service/llm/factory.py`
- **Action:** `openai_compat.py`: single `OpenAICompatProvider` for ollama-cloud (`https://ollama.com/v1`, Bearer) and local endpoints (base URL from settings); raw httpx against `/v1/chat/completions` with `stream: true`, byte-identical delta passthrough, tolerant of ollama-cloud quirks. Uses the lifespan-managed `httpx.AsyncClient` (10s connect / 300s read, D-017) — no openai SDK. `factory.py`: select provider from settings (`ollama-cloud` | `local` | `mock`).
- **Verify:** provider constructs from settings for all 3 names; manual probe against ollama-cloud streams tokens (documented in README, not a test)
#### Task 1-2-02: Lifespan wiring + SSE chat endpoint
- **Persona:** backend-engineer — **REQ:** REQ-2-003
- **Files:** `apps/ai-service/ai_service/main.py` (update), `apps/ai-service/ai_service/api/deps.py`, `apps/ai-service/ai_service/api/chat.py`, `apps/ai-service/ai_service/api/__init__.py`
- **Action:** Lifespan creates the shared `httpx.AsyncClient` and provider factory; deps.py provides provider via DI. `POST /v1/chat/stream` in chat.py implements the D-016 envelope: `meta` event (agent/session/model) flushed before first token → raw OpenAI chunks passed through as `data: {json}``done` event → `error` event before `[DONE]` on mid-stream failure; pre-first-byte failures return proper HTTP status codes. Headers `Cache-Control: no-cache`, `X-Accel-Buffering: no`; sse-starlette ping keep-alive.
- **Verify:** `curl -N -X POST localhost:8420/v1/chat/stream` with mock provider shows meta event, token deltas, done, `[DONE]`
### Wave 3: Provider + endpoint test suites (depends on Wave 2)
#### Task 1-3-01: LLM provider tests (byte-exact, no cloud)
- **Persona:** ai-engineer — **REQ:** REQ-2-002
- **Files:** `apps/ai-service/tests/llm/test_openai_compat.py`, `apps/ai-service/tests/llm/test_mock.py`, `apps/ai-service/tests/llm/__init__.py`
- **Action:** httpx `MockTransport` tests parsing byte-exact fixture streams (happy path, empty delta, `[DONE]`, malformed line, mid-stream disconnect). Mock provider tests: determinism, scripted failure modes, response_format echo.
- **Verify:** `pnpm ai:test` — llm suite green; zero network calls in tests
#### Task 1-3-02: SSE stream endpoint tests
- **Persona:** backend-engineer — **REQ:** REQ-2-003
- **Files:** `apps/ai-service/tests/api/test_chat_stream.py`, `apps/ai-service/tests/api/__init__.py`
- **Action:** TestClient `client.stream()` tests: meta-first ordering, delta passthrough, done + `[DONE]` sentinel, mid-stream error event, pre-first-byte failure → HTTP status, required headers. pytest-asyncio auto mode (D-023).
- **Verify:** `pnpm ai:test` — api suite green
### Must-Haves (Phase 1)
- [ ] `scripts/bootstrap.sh` is idempotent; creates venv + installs deps without system pip
- [ ] `pnpm ai:dev` starts uvicorn; `curl localhost:8420/health` returns 200
- [ ] `pnpm ai:lint` exits 0 (ruff check over the ai-service tree) (G-3)
- [ ] `pnpm ai:test` runs the full pytest suite via turbo and passes (mock provider only — no network)
- [ ] SSE stream delivers tokens: meta event, incremental deltas, done, `[DONE]` observed via `curl -N`
- [ ] Mid-stream failure emits `error` event before `[DONE]`; pre-first-byte failure returns HTTP error status
- [ ] Provider factory resolves ollama-cloud / local / mock from settings; manual ollama-cloud probe documented in README
- [ ] `llm/` imports nothing from `agents/` or `api/` (boundary rule holds)
---
## Phase 2: Agent Framework
**Requirements:** REQ-2-004
**Goal:** Shared framework all six agents use: BaseAgent contract, session store, prompt library, registry, structured outputs — all tested against the mock provider
### Wave 1: Framework primitives (parallel — no shared files)
#### Task 2-1-01: BaseAgent ABC
- **Persona:** ai-engineer — **REQ:** REQ-2-004
- **Files:** `apps/ai-service/ai_service/agents/base.py`, `apps/ai-service/tests/agents/test_base.py`, `apps/ai-service/ai_service/agents/__init__.py`, `apps/ai-service/tests/agents/__init__.py`
- **Action:** `BaseAgent` ABC (D-018): `name`, `system_prompt`, `build_messages(history, learner_context)`, `stream_reply(...) -> AsyncIterator[ChatDelta]` (delegates to provider), `structured_reply(...)` (delegates to structured module, landed Wave 2). Subclass contract tested with a stub agent + mock provider.
- **Verify:** `pnpm ai:test` — test_base green
#### Task 2-1-02: SessionStore protocol + in-memory implementation
- **Persona:** ai-engineer — **REQ:** REQ-2-004
- **Files:** `apps/ai-service/ai_service/agents/session.py`, `apps/ai-service/tests/test_session.py`
- **Action:** `SessionStore` protocol + `InMemorySessionStore` (D-019): asyncio.Lock-guarded dict, agent-scoped session keys, 20-message rolling window, 500-cap LRU eviction. Protocol shape is DB-migration-ready (A-003).
- **Verify:** test_session covers create/append/window-trim/LRU-eviction/agent scoping
#### Task 2-1-03: Prompt library scaffolding
- **Persona:** ai-engineer — **REQ:** REQ-2-004
- **Files:** `apps/ai-service/ai_service/prompts/coach.py`, `apps/ai-service/ai_service/prompts/tutor.py`, `apps/ai-service/ai_service/prompts/lab.py`, `apps/ai-service/ai_service/prompts/assessor.py`, `apps/ai-service/ai_service/prompts/proctor.py`, `apps/ai-service/ai_service/prompts/mentor.py`, `apps/ai-service/ai_service/prompts/__init__.py`
- **Action:** Per-agent module with a versioned `SYSTEM_PROMPT` constant + `render_context(learner_context) -> dict` using `str.format_map` for learner-context injection (D-018: prompts are code, versioned in git). Initial drafts for all six; final personas land in Phases 3-5.
- **Verify:** all six prompt modules import; render_context fills placeholders without KeyError
#### Task 2-1-04: Learner context corpus
- **Persona:** ai-engineer — **REQ:** REQ-2-004
- **Files:** `apps/ai-service/ai_service/corpus/learner_context.py`, `apps/ai-service/ai_service/corpus/__init__.py`
- **Action:** Pydantic-typed learner context (active stack, competencies, progress, recent artifacts) mirroring TS `packages/mock-data` IDs per D-021 convention (cross-referencing header comment, identical `comp-*`/`stack-*` ID strings).
- **Verify:** context renders into prompt placeholders; IDs match packages/mock-data strings
### Wave 2: Composition layers (depends on Wave 1)
#### Task 2-2-01: Structured output defense
- **Persona:** ai-engineer — **REQ:** REQ-2-004
- **Files:** `apps/ai-service/ai_service/agents/structured.py`, `apps/ai-service/tests/test_structured.py`
- **Action:** 4-layer defense (D-020): (1) `response_format` request with auto-degrade on provider 400; (2) prompt-embedded JSON schema; (3) parse: strip code fences → first balanced JSON object; (4) single bounded retry with validation-error feedback. Returns pydantic-validated model or raises `StructuredOutputError`.
- **Verify:** test_structured covers fenced/unfenced/invalid JSON, retry path, degrade path — all against mock provider
#### Task 2-2-02: Agent registry
- **Persona:** ai-engineer — **REQ:** REQ-2-004
- **Files:** `apps/ai-service/ai_service/agents/registry.py`, `apps/ai-service/tests/test_registry.py`
- **Action:** Explicit registry: name → agent factory map with `register(name, factory)` / `get(name)`; raises on unknown agent. Agents are registered in their own phases (P3-P5).
- **Verify:** test_registry: register/get round-trip, unknown-agent error, duplicate registration error
#### Task 2-2-03: Session + agent DI wiring into API layer
- **Persona:** backend-engineer — **REQ:** REQ-2-004
- **Files:** `apps/ai-service/ai_service/api/deps.py` (update), `apps/ai-service/ai_service/api/chat.py` (update)
- **Action:** deps.py exposes SessionStore and provider singletons via DI. chat.py persists turn history through the session store (agent-scoped) and includes session ID in the meta event. API composes agents via DI — agents never import api/.
- **Verify:** chat request appends to and replays windowed history; `pnpm ai:test` green
### Must-Haves (Phase 2)
- [ ] BaseAgent unit tests pass (stub agent streams via mock provider)
- [ ] Session store tested: create/append, 20-message window trim, 500-cap LRU eviction, agent-scoped keys
- [ ] Structured output parsing tested against mock provider: fence-strip, first-balanced-object, invalid JSON, one bounded retry, response_format auto-degrade
- [ ] Registry tested: register/get/unknown/duplicate
- [ ] Six prompt modules render learner context without errors
- [ ] Module boundaries hold: `agents/` never imports `api/`; `llm/` never imports `agents/` or `api/`
---
## Phase 3: Coach + Tutor Agents
**Requirements:** REQ-2-005, REQ-2-006
**Goal:** Both learner-facing conversational agents fully implemented with distinct personas, registered, routed through the chat streaming endpoint
### Wave 1: Agent implementations (parallel — no shared files)
#### Task 3-1-01: Coach agent
- **Persona:** ai-engineer — **REQ:** REQ-2-005
- **Files:** `apps/ai-service/ai_service/prompts/coach.py` (finalize), `apps/ai-service/ai_service/agents/coach.py`, `apps/ai-service/tests/test_coach.py`
- **Action:** Final Coach persona: pacing guidance, motivation, retrieval practice prompts; system prompt injects learner context (active stack, progress). `CoachAgent(BaseAgent)` streams replies. Mock provider scripts a distinct coach-voice response for tests.
- **Verify:** test_coach: build_messages includes system prompt + windowed history; stream_reply yields deltas; on-persona content asserted against mock script
#### Task 3-1-02: Tutor agent
- **Persona:** ai-engineer — **REQ:** REQ-2-006
- **Files:** `apps/ai-service/ai_service/prompts/tutor.py` (finalize), `apps/ai-service/ai_service/agents/tutor.py`, `apps/ai-service/tests/test_tutor.py`
- **Action:** Final Tutor persona: concept delivery, Socratic questioning, worked examples. `TutorAgent(BaseAgent)` streams replies; mock scripts a distinct tutor-voice response.
- **Verify:** test_tutor mirrors test_coach; Coach and Tutor mock outputs are observably distinct
### Wave 2: Registration + persona verification (depends on Wave 1)
#### Task 3-2-01: Register Coach + Tutor; document cloud probe
- **Persona:** ai-engineer — **REQ:** REQ-2-005, REQ-2-006
- **Files:** `apps/ai-service/ai_service/agents/registry.py` (update), `apps/ai-service/tests/test_registry.py` (update), `apps/ai-service/README.md` (update)
- **Action:** Register both agents in the explicit registry. Extend test_registry to assert both resolve. Document the manual ollama-cloud persona probe in README (curl commands with `AI_PROVIDER=ollama-cloud`): Coach and Tutor produce distinct on-persona responses; tests remain cloud-free.
- **Verify:** `pnpm ai:test` green; manual probe against ollama-cloud shows distinct personas (documented, not automated)
### Wave 3: Chat endpoint agent routing (depends on Wave 2)
#### Task 3-3-01: Agent routing on /v1/chat/stream
- **Persona:** backend-engineer — **REQ:** REQ-2-005, REQ-2-006
- **Files:** `apps/ai-service/ai_service/api/chat.py` (update), `apps/ai-service/ai_service/api/deps.py` (update), `apps/ai-service/tests/api/test_chat_stream.py` (update)
- **Action:** Chat request gains `agent` field (validated against the registry; unknown agent → 422). Endpoint resolves the agent via DI, persists to the agent-scoped session, meta event carries the agent name. No autonomous routing in v0.2 (A-007).
- **Verify:** TestClient tests: `agent=coach` and `agent=tutor` route correctly, session scoped per agent, unknown agent rejected
### Must-Haves (Phase 3)
- [ ] Both agents produce distinct, on-persona responses (mock-asserted; manual ollama-cloud probe documented in README)
- [ ] Agent routing tested: coach/tutor resolve via registry; unknown agent returns 422
- [ ] Both agents exposed end-to-end via `POST /v1/chat/stream` with agent-scoped session history
- [ ] `pnpm ai:test` green; no cloud calls in tests
---
## Phase 4: Lab + Assessor Agents
**Requirements:** REQ-2-007, REQ-2-008
**Goal:** Lab consumes simulated sandbox telemetry and streams in-flow feedback; Assessor applies rubrics to pre-baked artifacts and returns structured scores — both over mock engine inputs
### Wave 1: Mock engine inputs (parallel — no shared files)
#### Task 4-1-01: Simulated sandbox telemetry corpus
- **Persona:** ai-engineer — **REQ:** REQ-2-007
- **Files:** `apps/ai-service/ai_service/corpus/telemetry.py`
- **Action:** Pydantic-typed Lab telemetry scenarios: scripted build-session event streams (keystrokes, commits, test runs, errors, idle gaps) keyed by scenario ID, aligned with packages/mock-data IDs (D-021).
- **Verify:** scenarios import, validate, and are addressable by ID
#### Task 4-1-02: Pre-baked artifacts + rubrics corpus
- **Persona:** ai-engineer — **REQ:** REQ-2-008
- **Files:** `apps/ai-service/ai_service/corpus/artifacts.py`
- **Action:** Pre-baked artifacts (code, design, simulation), assessment rubrics (criteria, levels, weights), and defense transcripts keyed by ID — the Assessor's mock inputs (real engines are v0.3+).
- **Verify:** rubric/artifact/transcript fixtures validate; IDs align with TS mock data
#### Task 4-1-03: TS mock-data alignment for engine inputs
- **Persona:** data-engineer — **REQ:** REQ-2-012
- **Files:** `packages/mock-data/ai-scenarios.ts` (new), `packages/mock-data/index.ts` (update)
- **Action:** Export scenario/artifact ID constants + display metadata used by the Phase 6 learner panels, mirroring `ai_service/corpus/` IDs exactly (D-021). Data-engineer owns the TS side; headers cross-reference the Python corpus.
- **Verify:** `pnpm typecheck` passes; IDs string-equal to corpus IDs
### Wave 2: Agent implementations (depends on Wave 1)
#### Task 4-2-01: Lab agent
- **Persona:** ai-engineer — **REQ:** REQ-2-007
- **Files:** `apps/ai-service/ai_service/prompts/lab.py` (finalize), `apps/ai-service/ai_service/agents/lab.py`, `apps/ai-service/tests/test_lab.py`
- **Action:** `LabAgent(BaseAgent)` consumes a telemetry scenario, builds messages summarizing the event stream, streams concrete in-flow feedback (what happened, what to adjust, next step). No session chat — scenario-driven.
- **Verify:** test_lab: given a mock scenario, feedback references scenario events (mock-scripted assertions)
#### Task 4-2-02: Assessor agent
- **Persona:** ai-engineer — **REQ:** REQ-2-008
- **Files:** `apps/ai-service/ai_service/prompts/assessor.py` (finalize), `apps/ai-service/ai_service/agents/assessor.py`, `apps/ai-service/tests/test_assessor.py`
- **Action:** `AssessorAgent(BaseAgent)` applies a rubric to an artifact + defense transcript via `structured_reply`, returning a pydantic-validated rubric score model (per-criterion scores, strengths, gaps, verdict).
- **Verify:** test_assessor: structured output validates against the rubric model; failure modes exercise the 4-layer defense
### Wave 3: Registration (depends on Wave 2)
#### Task 4-3-01: Register Lab + Assessor
- **Persona:** ai-engineer — **REQ:** REQ-2-007, REQ-2-008
- **Files:** `apps/ai-service/ai_service/agents/registry.py` (update), `apps/ai-service/tests/test_registry.py` (update)
- **Action:** Register both agents; extend registry tests.
- **Verify:** registry resolves coach/tutor/lab/assessor; `pnpm ai:test` green
### Wave 4: Endpoints (depends on Wave 3)
#### Task 4-4-01: Lab feedback endpoint
- **Persona:** backend-engineer — **REQ:** REQ-2-007
- **Files:** `apps/ai-service/ai_service/api/lab.py`, `apps/ai-service/tests/api/test_lab.py`
- **Action:** `POST /v1/lab/feedback` with scenario ID → resolves corpus scenario + Lab agent → SSE stream using the D-016 envelope (meta names agent=lab). Unknown scenario → 404.
- **Verify:** TestClient streams meta + deltas + done + `[DONE]`; unknown scenario 404
#### Task 4-4-02: Assessment evaluate endpoint
- **Persona:** backend-engineer — **REQ:** REQ-2-008
- **Files:** `apps/ai-service/ai_service/api/assessment.py`, `apps/ai-service/tests/api/test_assessment.py`
- **Action:** `POST /v1/assessment/evaluate` with artifact ID → resolves corpus artifact/rubric/transcript + Assessor agent → JSON response (validated rubric score model). Unknown artifact → 404.
- **Verify:** TestClient returns validated rubric JSON; unknown artifact 404; `pnpm ai:test` green
### Must-Haves (Phase 4)
- [ ] Lab produces scenario-relevant in-flow feedback for mock telemetry scenarios (mock provider, tested)
- [ ] Assessor returns structured rubric scores (pydantic-validated JSON) for pre-baked artifacts/transcripts
- [ ] `POST /v1/lab/feedback` streams (meta → deltas → done → `[DONE]`); `POST /v1/assessment/evaluate` returns validated JSON
- [ ] Unknown scenario/artifact IDs return 404
- [ ] Corpus IDs align with packages/mock-data (D-021); `pnpm ai:test` and `pnpm typecheck` green
---
## Phase 5: Proctor + Mentor Agents
**Requirements:** REQ-2-009, REQ-2-010
**Goal:** Proctor classifies integrity signals with coaching interventions from mock telemetry; Mentor generates long-horizon career narrative; both exposed via endpoints
### Wave 1: Proctor scenarios + Mentor agent (parallel — no shared files)
#### Task 5-1-01: Proctor telemetry scenarios
- **Persona:** ai-engineer — **REQ:** REQ-2-009
- **Files:** `apps/ai-service/ai_service/corpus/telemetry.py` (update)
- **Action:** Add proctor scenarios: tab switches, idle time, paste events, focus loss — scripted integrity-relevant event sets keyed by scenario ID.
- **Verify:** proctor scenarios validate; distinguishable from lab scenarios by type
#### Task 5-1-02: Mentor agent (+ registration)
- **Persona:** ai-engineer — **REQ:** REQ-2-010
- **Files:** `apps/ai-service/ai_service/prompts/mentor.py` (finalize), `apps/ai-service/ai_service/agents/mentor.py`, `apps/ai-service/tests/test_mentor.py`, `apps/ai-service/ai_service/agents/registry.py` (update)
- **Action:** `MentorAgent(BaseAgent)`: long-horizon career narrative — trajectory story, competency-stack progression guidance, market positioning — streaming, session-backed. Registered centrally in `registry.py` (single registration pattern, G-4).
- **Verify:** test_mentor: narrative references learner context (mock-scripted); registry resolves mentor
### Wave 2: Proctor agent + Mentor endpoint (depends on Wave 1)
#### Task 5-2-01: Proctor agent (+ registration)
- **Persona:** ai-engineer — **REQ:** REQ-2-009
- **Files:** `apps/ai-service/ai_service/prompts/proctor.py` (finalize), `apps/ai-service/ai_service/agents/proctor.py`, `apps/ai-service/tests/test_proctor.py`, `apps/ai-service/ai_service/agents/registry.py` (update)
- **Action:** `ProctorAgent(BaseAgent)`: consumes proctor scenario → `structured_reply` returns pydantic-validated signal classification (severity, signal type) + recommended coaching intervention (supportive, not punitive). Registered centrally in `registry.py` (single registration pattern, G-4).
- **Verify:** test_proctor: classified signals + interventions validate for each mock scenario; registry resolves all six agents
#### Task 5-2-02: Mentor narrative endpoint
- **Persona:** backend-engineer — **REQ:** REQ-2-010
- **Files:** `apps/ai-service/ai_service/api/mentor.py`, `apps/ai-service/tests/api/test_mentor.py`
- **Action:** `POST /v1/mentor/narrative` → Mentor agent → SSE stream with D-016 envelope, session-backed.
- **Verify:** TestClient streams meta (agent=mentor) → deltas → done → `[DONE]`
### Wave 3: Proctor endpoint (depends on Wave 2)
#### Task 5-3-01: Proctor signals endpoint
- **Persona:** backend-engineer — **REQ:** REQ-2-009
- **Files:** `apps/ai-service/ai_service/api/proctor.py`, `apps/ai-service/tests/api/test_proctor.py`
- **Action:** `POST /v1/proctor/signals` with scenario ID → resolves corpus scenario + Proctor agent → JSON response (validated signals + interventions). Unknown scenario → 404.
- **Verify:** TestClient returns classified signals JSON; `pnpm ai:test` green — full suite (all six agents registered)
### Must-Haves (Phase 5)
- [ ] Proctor produces classified signals with recommended coaching interventions for each mock scenario (structured JSON, validated)
- [ ] Mentor produces coherent long-horizon career narrative (streaming, session-backed)
- [ ] `POST /v1/proctor/signals` returns validated JSON; `POST /v1/mentor/narrative` streams
- [ ] Registry resolves all six agents; full ai-service test suite green, cloud-free
---
## Phase 6: Learner Surface Integration
**Requirements:** REQ-2-011, REQ-2-012
**Goal:** v0.1 learner surfaces wired to the real ai-service: streaming chat with agent switcher, Lab/Assessor/Proctor/Mentor outputs surfaced, error/loading states, build + typecheck green
**Note (G-2):** End-to-end verification may run with `AI_PROVIDER=mock` as a fallback — the requirement is the real ai-service over HTTP (not canned client-side responses); provider choice is service-internal. This prevents an ollama-cloud outage from blocking P6 verification. Cloud persona probes remain separate (Task 3-2-01).
### Wave 1: Client plumbing + primitives (parallel — no shared files)
#### Task 6-1-01: useChatStream hook
- **Persona:** frontend-engineer — **REQ:** REQ-2-011
- **Files:** `apps/web/hooks/use-chat-stream.ts`, `apps/web/.env.example` (update)
- **Action:** `useChatStream(agent)` hook: `fetch` POST to `${NEXT_PUBLIC_AI_SERVICE_URL}/v1/chat/stream` (default `http://localhost:8420`, A-002 — no API-route proxy); consumes `ReadableStream` with byte buffering, frame split on `\n\n`, joined `data:` lines; **ignores frames containing no `data:` lines (sse-starlette `: ping` keep-alive comment frames)** — TestClient streams are too short to surface pings, but real cloud delta gaps emit them (G-1); handles meta / delta / done / error events and `[DONE]` sentinel; idempotent `AbortController.abort()` in effect cleanup; exposes `{messages, isStreaming, error, send, retry, abort}`. `.env.example` gains `NEXT_PUBLIC_AI_SERVICE_URL`.
- **Verify:** hook unit-tested or exercised via the chat UI; unmount mid-stream aborts cleanly (no state updates after unmount)
#### Task 6-1-02: Agent switcher + streaming primitives
- **Persona:** design-system-engineer — **REQ:** REQ-2-011
- **Files:** `packages/ui/src/primitives/agent-switcher.tsx`, `packages/ui/src/primitives/stream-status.tsx`, `packages/ui/src/primitives/toast.tsx`, `packages/ui/src/primitives/index.ts` (update), `packages/ui/src/index.ts` (update)
- **Action:** Token-driven primitives: AgentSwitcher (segmented coach/tutor control with active state), StreamStatus (idle/streaming/error indicator), Toast with error variant. Dark mode + WCAG AA contrast; exported from `@nextcraft/ui`.
- **Verify:** primitives import from `@nextcraft/ui`; storybook stories render (dark + light)
### Wave 2: Chat rewrite + Byte viewer panel (depends on Wave 1)
#### Task 6-2-01: Real streaming learner chat
- **Persona:** frontend-engineer — **REQ:** REQ-2-011
- **Files:** `apps/web/components/learner/ai-tutor-chat.tsx` (rewrite), `apps/web/components/learner/agent-switcher.tsx` (composition wrapper)
- **Action:** Replace the canned `aiTutorResponses` behavior with `useChatStream`: agent switcher (Coach/Tutor per A-007), token-by-token rendering, streaming cursor + loading state, error state with retry button when ai-service is down (A-010), suggested-action chips from the meta event. Seed welcome message stays static.
- **Verify:** with ai-service running, messages stream visibly token-by-token; with ai-service stopped, error state + retry appears (no crash, no console errors)
#### Task 6-2-02: Byte viewer Tutor panel
- **Persona:** frontend-engineer — **REQ:** REQ-2-011 (agent routing: byte viewer always uses Tutor, A-007)
- **Files:** `apps/web/app/(learner)/learn/[competencyId]/page.tsx` (update), `apps/web/components/learner/byte-tutor-panel.tsx` (new)
- **Action:** Byte viewer gains a Tutor explanation panel (byte viewer always uses Tutor, A-007): "Explain this byte" streams a Socratic concept walkthrough for the current competency via useChatStream (agent fixed to tutor).
- **Verify:** on a byte page, the panel streams a Tutor explanation; error state when service down
### Wave 3: Lab / Assessment / Mentor panels (depends on Wave 1 hook; parallel — no shared files)
#### Task 6-3-01: Sandbox Lab feedback panel
- **Persona:** frontend-engineer — **REQ:** REQ-2-012
- **Files:** `apps/web/app/(learner)/build/[competencyId]/page.tsx` (update), `apps/web/components/learner/lab-feedback-panel.tsx` (new)
- **Action:** Sandbox telemetry sidebar gains a Lab feedback panel: posts the scenario ID (from `packages/mock-data/ai-scenarios`) to `/v1/lab/feedback`, streams in-flow feedback into the panel; loading + error states.
- **Verify:** sandbox page streams Lab feedback for the mock scenario; error state when service down
#### Task 6-3-02: Assessment Assessor + Proctor surfaces
- **Persona:** frontend-engineer — **REQ:** REQ-2-012
- **Files:** `apps/web/app/(learner)/defend/[competencyId]/page.tsx` (update), `apps/web/components/learner/assessor-results-panel.tsx` (new), `apps/web/components/learner/proctor-banner.tsx` (new)
- **Action:** Assessment mockup: AI reviewer panel calls `/v1/assessment/evaluate` with the artifact ID and renders the structured rubric scores (per-criterion bars, strengths, gaps, verdict); a Proctor integrity banner surfaces `/v1/proctor/signals` classifications with coaching tone. Loading skeletons + error states.
- **Verify:** defend page renders real Assessor rubric output + Proctor banner; error states when service down
#### Task 6-3-03: Dashboard Mentor panel
- **Persona:** frontend-engineer — **REQ:** REQ-2-012
- **Files:** `apps/web/app/(learner)/dashboard/page.tsx` (update), `apps/web/components/learner/mentor-panel.tsx` (new)
- **Action:** Learner dashboard gains a Mentor panel: streams career narrative from `/v1/mentor/narrative` (learner progress context), with regenerate button, loading + error states. Sits alongside the existing AI tutor chat.
- **Verify:** dashboard shows streaming Mentor narrative; error state when service down
### Must-Haves (Phase 6)
- [ ] With ai-service running: learner chat at http://localhost:3000/dashboard streams real responses token-by-token (mock provider fallback allowed per G-2 — service over HTTP is the requirement)
- [ ] Agent switcher flips Coach ↔ Tutor and the response persona changes accordingly
- [ ] Hook tolerates keep-alive comment frames (`: ping`, no data lines) during live streams (G-1)
- [ ] With ai-service stopped: all chat/panels show error states with retry — no crashes, no unhandled promise rejections, no console errors
- [ ] Byte viewer, sandbox, and assessment mockups surface Tutor/Lab/Assessor/Proctor outputs; dashboard shows Mentor narrative
- [ ] Unmounting/navigating mid-stream aborts cleanly (no post-unmount state updates)
- [ ] `pnpm build` and `pnpm typecheck` pass; `pnpm ai:test` still green
---
## Phase 7: Final Review + Ship (no planned tasks)
Orchestrated by the SHIP stage, not this plan: multi-persona code review (correctness, testing, secrets hygiene — key absent from code/logs/commits/errors, localhost-only CORS, no PII in prompts), project health audit (reconstruction test, .ciagent/ discipline, branch/commit hygiene), then merge milestone → main, tag v0.2.0, create Gitea release, mark all 12 v0.2 requirements complete.
**Release-note honesty (G-5):** the v0.2.0 release note must explicitly state that Lab/Assessor/Proctor operate on mock engine inputs (real engines are v0.3+), per D-015.
**Dead-code disposal (G-5):** review must dispose of `aiTutorResponses` (packages/mock-data/ai-tutor-responses.ts) — its only consumer is rewritten in Task 6-2-01; remove the export or mark it deprecated.
---
## User-Facing Surface
1. **One-liner install (README quickstart):** `curl -fsSL https://git.coreci.dev/coreci/nextcraft/raw/main/scripts/install.sh | bash` — downloads the latest release's `nextcraft-linux-x64` binary, verifies its sha256, installs to `~/.local/bin`, prints a PATH hint if needed.
2. **CLI commands:** `nextcraft doctor` (prereq checks), `nextcraft bootstrap` (fresh clone → runnable stack), `nextcraft verify` (health check), `nextcraft dev` (dev server passthrough), plus `--help`/`--version`.
3. **Release surface:** every Gitea release from v0.3.2 onward carries `nextcraft-linux-x64` + `nextcraft-linux-x64.sha256` assets.
The primary user-facing surface is the **learner dashboard and learning flow** at `http://localhost:3000`, backed by the real ai-service at `http://localhost:8420`:
- `/dashboard` — AI tutor chat (Coach/Tutor switcher, streaming) + Mentor career-narrative panel
- `/learn/[competencyId]` — byte viewer with streaming Tutor explanations
- `/build/[competencyId]` — sandbox with Lab in-flow feedback panel (mock telemetry)
- `/defend/[competencyId]` — assessment with live Assessor rubric scores + Proctor integrity banner
The marketplace, employer, and admin surfaces are unchanged from v0.1.
## Happy Path
Before execution, the end-to-end scenario this milestone must make true:
1. A consumer on a linux x64 box runs the one-liner; `nextcraft` lands in `~/.local/bin`.
2. They clone the repo (or the CLI detects the repo root), run `nextcraft doctor` — all prerequisites report ✓ with actionable messages for any gap.
3. `nextcraft bootstrap` — pnpm install, ai-service venv via the existing bootstrap.sh, `.env` created from `.env.example`, optional-key warnings (not blockers), mock providers keep the stack runnable keyless.
4. `nextcraft verify` — venv imports, ports, env presence, build readiness all ✓.
5. `nextcraft dev` — the dev stack runs; Ctrl+C stops it (passthrough semantics).
6. On every ship, the Gitea release shows the binary + checksum assets; re-running the one-liner upgrades to the latest binary.
1. Learner opens `/dashboard` → chat shows welcome message; meta event confirms coach/model in the stream
2. Learner types "I'm stuck on multi-agent communication" → reply streams token-by-token with pacing guidance + a retrieval-practice prompt
3. Learner switches to **Tutor** → asks the same question → gets a Socratic concept walkthrough instead
4. Learner opens a byte tutorial → Tutor panel streams an explanation of the current competency
5. Learner opens the build sandbox → Lab panel streams feedback on the simulated telemetry scenario
6. Learner opens the defense mockup → Assessor panel shows structured rubric scores; Proctor banner shows integrity signals in coaching tone
7. Back on `/dashboard`, the Mentor panel streams a career narrative tied to the learner's progress
8. Learner kills ai-service (or it crashes) → next message shows an inline error state with **Retry**; restarting the service and retrying resumes streaming
## UX Acceptance Criteria
- `doctor` output lists every prerequisite with ✓/✗ and a **fix hint** on every ✗; exit code 1 if any ✗, 0 otherwise.
- `bootstrap` is **idempotent** — running twice produces the same end state, second run fast (no reinstalls where avoidable).
- `bootstrap` never writes secrets, never blocks on missing optional keys — warns with the exact key names and where to set them.
- `verify` gives a single-glance green/red summary; every red item names the failing command it ran.
- Every command supports `--help`; unknown command/flag exits 2 with usage.
- The one-liner **never hard-fails silently**: any error path (no release, no binary asset, checksum mismatch, platform mismatch) prints a specific message + the source-bootstrap alternative.
- Checksum mismatch = hard stop + explicit "do not run this binary" message.
- PATH hint: if `~/.local/bin` is not on PATH, the installer prints the exact export line to add.
- Binary runs standalone on a box with node NOT installed (SEA self-containment) — `./nextcraft-linux-x64 --version` works.
---
## Phase 1: Bootstrap CLI Core
**Requirements:** REQ-4-001, REQ-4-002
**Goal:** `apps/cli` package with doctor/bootstrap/verify/dev fully working from source (`node dist` + pnpm bin), unit-tested, wired into the monorepo (turbo + root scripts), composing — not duplicating — the existing scripts.
### Wave 1: Package foundation (parallel)
#### Task 1-1-01: CLI package scaffold + entry + dispatch
- **Persona:** cli-engineer — **REQ:** REQ-4-001
- **Files:** `apps/cli/package.json`, `apps/cli/tsconfig.json`, `apps/cli/src/index.ts`, `apps/cli/src/commands/help.ts` (usage text), `apps/cli/tests/dispatch.test.ts`
- **Action:** pnpm workspace package `@nextcraft/cli` (private, `"bin": {"nextcraft": "dist/index.js"}`). Entry: parse argv (hand-rolled, no runtime deps), dispatch to commands, `--help`/`-h`, `--version` (from package.json version), unknown → exit 2 with usage. Exit-code contract: 0 ok / 1 failure / 2 usage. shebang `#!/usr/bin/env node` on the built entry (esbuild banner in P2; for P1 `tsx` runs in dev via package script `"dev": "tsx src/index.ts"`).
- **Verify:** `pnpm --filter @nextcraft/cli test` green (dispatch: routes doctor/bootstrap/verify/dev; unknown exits 2; --help exits 0; --version prints package version); `pnpm typecheck` green.
#### Task 1-1-02: Checks library (pure logic)
- **Persona:** cli-engineer — **REQ:** REQ-4-001, REQ-4-002
- **Files:** `apps/cli/src/checks/check-command.ts`, `apps/cli/src/checks/check-env.ts`, `apps/cli/src/lib/log.ts`, `apps/cli/tests/checks.test.ts`
- **Action:** `check-command`: given a name + optional `--version` probe + a min-version parser, resolve binary on PATH (`which`), semver-ish compare (major.minor tolerant), return `CheckResult {name, ok, found, version, hint}`. `check-env`: diff `.env.example` template keys vs an existing `.env` (missing keys → warn-classified; required-vs-optional classification table from the template's own comments + a static required list of zero keys — all optional per A-210), return per-key results. `log.ts`: `ok(msg)`, `fail(msg, hint)`, `warn(msg)`, `info(msg)` formatters with symbols and consistent alignment. Pure functions — no side effects at import; fs access injected as parameters for testability.
- **Verify:** unit tests green: version compare (>= boundaries), missing binary → ok:false + hint, env diff missing/new/extra keys, required-optional classification.
#### Task 1-1-03: Root + turbo wiring
- **Persona:** backend-engineer — **REQ:** REQ-4-002
- **Files:** root `package.json` (update), `turbo.json` (update), `pnpm-workspace.yaml` (verify apps/* already covered — no change expected)
- **Action:** Add `cli:dev`, `cli:test`, `cli:build`, `cli:typecheck`, `cli:lint` root scripts mirroring the `ai:*` passthrough pattern (D-022/D-037). Turbo tasks for the CLI package: `build` (dependsOn `^build`, outputs `dist/**`), `test`, `typecheck`, `lint` (cache:false, outputs:[] for test — same shape as ai-service). No changes to existing ai:* tasks.
- **Verify:** `pnpm cli:test` + `pnpm cli:typecheck` green from repo root; `pnpm build` still green for web+ai-service (turbo graph unaffected); `pnpm ai:test` still green.
### Wave 2: Commands (depends on Wave 1)
#### Task 1-2-01: doctor command
- **Persona:** cli-engineer — **REQ:** REQ-4-001
- **Files:** `apps/cli/src/commands/doctor.ts`, `apps/cli/tests/doctor.test.ts`
- **Action:** Checks (each with actionable hint): node ≥18 (`process.version`), pnpm ≥8 on PATH (`pnpm --version`), python3 ≥3.11 (`python3 --version` parse), git (`git --version`), corepack available-or-pnpm-present nuance folded into pnpm check, `unshare` binary on PATH (`which unshare` — sandbox fabric needs it; hint explains what breaks without it). Sequential execution with per-check timeout; summary line; exit 1 if any ✗. Runs from any cwd (no repo required — pure environment check).
- **Verify:** unit tests with injected spawn results: all-pass → exit 0 + summary; missing pnpm → ✗ + hint + exit 1; missing unshare → ✗ with sandbox-specific hint.
#### Task 1-2-02: bootstrap command
- **Persona:** cli-engineer — **REQ:** REQ-4-002
- **Files:** `apps/cli/src/commands/bootstrap.ts`, `apps/cli/src/lib/spawn.ts`, `apps/cli/tests/bootstrap.test.ts`
- **Action:** `spawn.ts`: `run(cmd, args, {timeoutMs, cwd, env})` — promisified child_process.spawn, inherited stdio, timeout kill (SIGTERM→SIGKILL escalation), returns `{code}`; throws never (codes always returned). `bootstrap.ts` steps (each logged before/after): (1) locate repo root (walk up for pnpm-workspace.yaml; error with hint if not in a clone); (2) `pnpm install` at root; (3) delegate ai-service venv to `apps/ai-service/scripts/bootstrap.sh` via spawn with generous timeout (10 min) — **zero pip/venv logic in the CLI** (A-202); (4) copy `.env.example``.env` if absent (preserve existing; report created vs kept); (5) validate optional keys in `.env` vs template — warn-only (A-210); never touch `.ciagent/.env.secrets`; (6) print next-steps (`nextcraft verify`, `nextcraft dev`). Idempotent: every step safe to re-run.
- **Verify:** unit tests with stub spawn: step order, env copy semantics (absent → create, present → keep), timeout path returns failure code, secrets file never written; `pnpm cli:test` green.
#### Task 1-2-03: verify + dev commands
- **Persona:** cli-engineer — **REQ:** REQ-4-002
- **Files:** `apps/cli/src/commands/verify.ts`, `apps/cli/src/commands/dev.ts`, `apps/cli/tests/verify.test.ts`
- **Action:** `verify.ts` health checks (each runnable + reported): ai-service venv python imports (`import ai_service` via venv python), uvicorn present in venv, ports 3000/8420 free (net stat via node), `.env` exists with AI_PORT parseable, `pnpm build` dry readiness (turbo graph parses — run `turbo build --dry=json` cheap check or typecheck-only default; choose the cheap one). Summary + exit code. `dev.ts`: locate repo root, exec passthrough to `apps/ai-service/scripts/dev.sh` with **inherited stdio and signals** (Ctrl+C semantics), no timeout (long-running); document that web dev server runs via `pnpm dev` separately (dev.sh owns ai-service only).
- **Verify:** unit tests: verify aggregates check results → exit codes; dev spawns dev.sh with signal passthrough assertions (mock spawn).
### Wave 3: Integration review (depends on Wave 2)
#### Task 1-3-01: CLI security + integration review pass
- **Persona:** security-auditor — **REQ:** REQ-4-001, REQ-4-002
- **Files:** `apps/cli/src/lib/spawn.ts` (review; patch if defect), `apps/cli/src/commands/bootstrap.ts` (review), `apps/cli/tests/**` (add regression if defect found)
- **Action:** STRIDE pass on the CLI surface: spawn injection (args never through shell string — array form only), timeout enforcement, secrets never logged, env template copy doesn't overwrite user edits, no shell=true anywhere, PATH resolution honest errors. Findings → P0 patches now with regression tests; P1+ noted for final-phase review.
- **Verify:** `pnpm cli:test` green incl. any added regressions; `grep -rn "shell: *true" apps/cli/src` returns nothing.
### Must-Haves (Phase 1)
- [ ] `pnpm --filter @nextcraft/cli test` green; `pnpm typecheck` green; `pnpm build` green
- [ ] doctor: every prerequisite reported with ✓/✗ + actionable hint; exit 1 on any ✗; runs outside a repo clone
- [ ] bootstrap: composes scripts/bootstrap.sh (no pip/venv logic in CLI); idempotent; .env created from template only when absent; optional-key warnings, never blocks; never writes secrets
- [ ] verify: venv import + uvicorn + ports + env checks with single-glance summary and named failing commands
- [ ] dev: passthrough with signal inheritance (Ctrl+C stops the stack)
- [ ] Exit-code contract: 0/1/2; --help everywhere; unknown command → 2
- [ ] No runtime npm dependencies in apps/cli (dev deps only)
- [ ] Root `cli:*` scripts work from repo root; ai:* scripts unaffected
---
## Phase 2: Binary Build + Release Pipeline
**Requirements:** REQ-4-003, REQ-4-004
**Goal:** `nextcraft-linux-x64` SEA binary + sha256 sidecar built reproducibly from the CLI package; one-liner `install.sh` verified end-to-end against a real release; release-asset upload wired so **every ship from now on carries binaries**.
### Wave 1: Binary build (parallel)
#### Task 2-1-01: SEA binary build script (G-101: live-build probe FIRST — mechanism must be proven before the pipeline depends on it)
- **Persona:** cli-engineer — **REQ:** REQ-4-004
- **Files:** `apps/cli/scripts/build-binary.mjs`, `apps/cli/package.json` (add `build:binary` script), `apps/cli/.sea-config.json` (or generated in-script)
- **Action:** **First action of this task: build one real SEA binary end-to-end and run it** (`--version` + `doctor` smoke) before writing the polished script.** Pipeline: esbuild bundle `src/index.ts``dist/bundle.cjs` (platform node, target node18, banner shebang, SEA config: `{main: "dist/bundle.cjs", output: "dist/sea-prep.blob", disableExperimentalSEAWarning: true}`) → `node --experimental-sea-config` → copy system node binary → inject blob (`npx postject` with sentinel `NODE_SEA_BLOB_FUSE` fuse, or `dd` fallback) → chmod +x → `dist/nextcraft-linux-x64`**stamp version from the shipping tag argument** (`NEXTCRAFT_VERSION` injected via esbuild `define`, G-102 — `--version` prints it; absent arg → dev stamp `0.0.0-dev`) → `shasum -a 256``dist/nextcraft-linux-x64.sha256`. Fallback (documented, scripted, honest): if SEA injection fails, python3 zipapp builds `nextcraft-linux-x64.pyz` (requires python3 on target — install.sh handles both asset shapes and the docs say so; NO silent claim of node-less operation, G-101).
- **Verify:** `pnpm --filter @nextcraft/cli build:binary` produces the binary; `./dist/nextcraft-linux-x64 --version` runs **with node absent from PATH** (test via `env -i /bin/sh -c 'PATH=/usr/bin:/bin ...'` sandbox or by temporarily stripping PATH in a subprocess test); sha256 file matches `shasum -c`.
#### Task 2-1-02: Release-asset upload helper
- **Persona:** cli-engineer — **REQ:** REQ-4-004
- **Files:** `scripts/release-assets.sh`, `apps/cli/tests/release-assets.test.ts` (fixture-level)
- **Action:** Given a tag: build binary (Task 2-1-01), resolve GITEA_TOKEN from `.env`/`.env.secrets`/`.env.*` **via the secrets loader only** (never shell env — v1.8 root cause), create/locate the Gitea release via API, upload both assets (`POST /api/v1/repos/{owner}/{repo}/releases/{id}/assets?name=...` multipart). Bounded retry (3) per config.ship.max_release_retries; token never echoed; failure = non-blocking escalation message (release_pending semantics) — tag+merge already complete the ship.
- **Verify:** fixture test: token resolution order (.env.secrets wins over .env; shell env NEVER consulted — assert with a poisoned env var fixture); dry-run mode prints the exact curl-multipart it would send (no net in tests).
#### Task 2-1-03: install.sh one-liner
- **Persona:** cli-engineer — **REQ:** REQ-4-003
- **Files:** `scripts/install.sh`, `apps/cli/tests/install-script.test.ts`
- **Action:** POSIX sh (no bashisms — dash-safe): `set -eu`; platform check (uname linux + x86_64; else print source-bootstrap path + exit 0 — a graceful no-op, not an error); resolve latest release via Gitea API (`curl -fsSL .../releases/latest`, parse `tag_name` + asset `browser_download_url`s with sed/grep — no jq dependency); **match assets by EXACT name** (`nextcraft-linux-x64`, `nextcraft-linux-x64.sha256` — any parse/lookup miss = degrade to source-bootstrap instructions, exit 0, G-103 — never a name-approximate install); handle the zipapp asset shape (`nextcraft-linux-x64.pyz` + sidecar) when the binary is absent, printing the python3 requirement honestly; download both assets to `mktemp -d` (trap cleanup EXIT); **verify sha256 before anything else** (`shasum -a 256 -c` or sha256sum); on mismatch → hard stop, explicit "do not run" message, exit 1; install to `~/.local/bin` (mkdir -p; `--dest` override); PATH hint when missing (print exact export line); print the binary's own `--version` output (G-102: must equal the resolved release tag — mismatch = install-time integrity stop) + `nextcraft doctor` next-step. No-binary-asset path: print the git-clone + scripts/bootstrap.sh instructions + exit 0. Zero secrets required (public release assets).
- **Verify:** unit tests over the script's pure helpers extracted where feasible; **live E2E in Task 2-3-01**. `sh -n scripts/install.sh` syntax-clean; `dash scripts/install.sh --help` safe if dash present.
### Wave 2: Ship-flow integration (depends on Wave 1)
#### Task 2-2-01: Wire binaries into every ship (G-104: enforcement, not prose)
- **Persona:** backend-engineer — **REQ:** REQ-4-004
- **Files:** `.ciagent/config.json` (no schema change needed — release section already configured), this repo's ship procedure notes (update `.ciagent/ARCHITECTURE.md` Build Order note if needed), `scripts/release-assets.sh` (finalize from 2-1-02)
- **Action:** Establish the ship-time contract going forward: after every phase ship (tag + merge complete = ship gate per config.ship), run `scripts/release-assets.sh <tag>` to attach binary + checksum to the freshly created release. **G-104:** this run is MANDATORY-ATTEMPTED on every release from v0.3.2 onward — best-effort/non-blocking like release creation (release_pending escalation on exhaustion), logged in the ship commit, and the P4 final audit gate includes "milestone release carries both assets" as an explicit check. This makes "ongoing binaries" a property of the pipeline, not a one-off.
- **Verify:** The P2 ship itself executes the step against tag v0.3.2 (live validation — see Ship).
### Wave 3: End-to-end validation (depends on Wave 2)
#### Task 2-3-01: Install E2E against the live release
- **Persona:** security-auditor — **REQ:** REQ-4-003
- **Files:** `apps/cli/tests/install-e2e.test.ts` (marked slow/e2e), `apps/cli/README.md` (install internals section)
- **Action:** Live E2E after the v0.3.2 release exists (run post-ship, documented as the verify gate for this phase's asset path): fresh HOME tmpdir → run install.sh → assert binary at `$HOME/.local/bin/nextcraft`, `--version` output equals the release tag (G-102 integrity assertion), checksum verified path taken (tamper test: flip a byte in a local fixture download → script refuses + exits 1). Record the transcript in the phase verify commit. If the live release isn't reachable at verify time, run the full local equivalent (serve assets from a fixture dir via `python3 -m http.server` + FORGE_BASE override) and mark live re-check as a P1 follow-up.
- **Verify:** E2E green locally (fixture server path mandatory in tests — no test depends on the live forge); tamper-rejection proven; transcript recorded.
### Must-Haves (Phase 2)
- [ ] **G-101:** a real SEA binary built + smoke-run BEFORE the pipeline depends on it; if SEA fails, zipapp is primary and docs state the python3 requirement
- [ ] **G-102:** binary `--version` reports the shipping tag (stamped at build); install E2E asserts version == release tag
- [ ] `pnpm --filter @nextcraft/cli build:binary` produces `nextcraft-linux-x64` + `.sha256`; binary runs without node on PATH (`--version`, `doctor` smoke)
- [ ] **G-103:** `sh -n scripts/install.sh` clean; dash-safe; exact-name asset matching; platform mismatch → graceful source-bootstrap path (exit 0)
- [ ] Checksum verified before install; tamper → hard stop with explicit warning (E2E-proven)
- [ ] install.sh resolves latest release + assets from the Gitea API with zero secrets and no jq
- [ ] release-assets.sh resolves GITEA_TOKEN from .env* files only (never shell env — tested with poisoned env)
- [ ] **G-104:** v0.3.2 release carries both assets (live validation at ship); upload failure is non-blocking escalation, attempted + logged every release
- [ ] `pnpm build`, `pnpm typecheck`, `pnpm cli:test` all green
---
## Phase 3: Install Docs + Fresh-Clone E2E
**Requirements:** REQ-4-005
**Goal:** README quickstart + CLI reference matching the tested reality exactly, plus a fresh-clone E2E test proving the happy path end-to-end.
### Wave 1: Fresh-clone E2E (drives doc accuracy)
#### Task 3-1-01: Fresh-clone bootstrap E2E
- **Persona:** cli-engineer — **REQ:** REQ-4-005
- **Files:** `apps/cli/tests/fresh-clone-e2e.test.ts` (slow/e2e-marked)
- **Action:** In a `mktemp -d` sandbox: `git clone` the repo locally (file:// clone of HEAD — no network), run `pnpm --filter @nextcraft/cli dev -- doctor` (or the built binary from P2) → then `bootstrap` → then `verify`, asserting each step's exit codes and key output markers. Skips gracefully when network-dependent steps are unavailable (CI marker). Documents the exact happy path the README will state.
- **Verify:** E2E green locally (clone of the working tree); output transcript matches README claims (cross-checked in 3-2-01).
### Wave 2: Documentation (depends on Wave 1 transcript)
#### Task 3-2-01: README quickstart + CLI reference
- **Persona:** backend-engineer — **REQ:** REQ-4-005
- **Files:** root `README.md` (update quickstart section), `apps/cli/README.md` (CLI reference)
- **Action:** Root README quickstart: the one-liner (exact tested URL), then doctor → bootstrap → verify → dev sequence with expected outputs; source-bootstrap alternative documented (clone + scripts). apps/cli README: every command, flags, exit codes, the env-template copy semantics, optional-key warning semantics, secrets policy (never generated/committed; .ciagent/.env.secrets location), binary install internals, troubleshooting table keyed to actual failure modes observed in E2E.
- **Verify:** Every command line in both READMEs is copy-paste runnable — verified against the 3-1-01 transcript; doc drift check: no references to commands/flags that don't exist in `--help` output.
### Must-Haves (Phase 3)
- [ ] Fresh-clone E2E green: doctor → bootstrap → verify sequence from a clean clone
- [ ] README quickstart matches the E2E transcript exactly (no aspirational docs)
- [ ] CLI reference covers all 4 commands + --help/--version + exit codes
- [ ] `pnpm build`, `pnpm typecheck`, `pnpm test` (all suites) green
---
## Phase 4: Final Review + Ship (milestone release v0.3.4)
1. Branch gate → `phase/04-final-review-ship`.
2. Multi-persona review across the milestone (correctness, testing, security, performance, maintainability, adversarial) — P0 auto-fixed, P1+ fixed in this phase.
3. Audit: reconstruction test (.ciagent files ↔ git log), file discipline, branch hygiene, commit discipline, P0-review flags resolved, **G-104 gate: milestone release v0.3.4 carries `nextcraft-linux-x64` + `.sha256` assets**.
4. Milestone ship: merge phase/04 → milestone/v0.4-distribution; merge milestone → main; tag **v0.3.4** (= milestone release); attach binary + checksum assets (the ongoing-binaries contract); release notes with full milestone summary (all phases, all REQ-4-001..005, the "ongoing binaries from now on" statement, v0.5 deferral list per D-016); delete all milestone/phase branches.
5. Complete: REQUIREMENTS.md REQ-4-001..005 → complete; ROADMAP.md v0.4 → complete; checkpoint cleared.
## Must-Haves (Milestone)
- [ ] One-liner installs a working binary from the live Gitea release (E2E-proven, tamper-tested)
- [ ] Fresh clone → doctor → bootstrap → verify → dev: the full happy path green from a clean environment
- [ ] Every release from v0.3.2 onward carries `nextcraft-linux-x64` + `.sha256` assets
- [ ] Zero runtime npm deps in the CLI; secrets only ever from .env* files; never in code/logs/commits
- [ ] All suites green: `pnpm build`, `pnpm typecheck`, `pnpm ai:test`, `pnpm cli:test`
1. Streaming is visibly incremental — tokens appear as they arrive, not as one blob
2. Agent switcher shows the active agent (Coach/Tutor) and the response persona visibly changes
3. Loading state during connection (streaming cursor / skeleton) before first token
4. When ai-service is unreachable: inline error state + retry action on every chat/panel — no crashes, no console errors, no blank UI
5. `[DONE]` reliably ends the stream (input re-enables, no stuck "typing" state)
6. Navigating away mid-stream aborts cleanly no leaked requests or post-unmount updates
7. All new UI uses design tokens, supports dark mode, meets WCAG AA contrast
8. Responsive at 375px, 768px, 1280px
9. No hardcoded model names or URLs in UI code — all via `NEXT_PUBLIC_AI_SERVICE_URL` and server meta events
10. `pnpm build` and `pnpm typecheck` pass with zero errors
+23 -60
View File
@@ -8,75 +8,35 @@ Nextcraft is an AI-native outcome school where graduates prove what they can bui
---
## Current Milestone: v0.5Real Server Voice, KYC/Identity, Sandbox Environments, Seq-Lease
## Current Milestone: v0.2AI Tutor Architecture
**Scope (deferred seams from D-016, to be specified at v0.5 Phase 0):** real server STT/TTS (`openai-audio` voice provider, CUT-1/G-7 seam), KYC/identity verification + age-gating backend (REQ-F-017), design/simulation sandbox environments (REQ-F-021 remainder), exec-telemetry seq-lease/replay-margin fix.
**Scope:** The six AI tutor agents (Coach, Tutor, Lab, Assessor, Proctor, Mentor) as real LLM-backed services in a new `apps/ai-service` Python FastAPI application, wired into the existing v0.1 learner surface chat UI with streaming responses. Provider-agnostic LLM layer (ollama-cloud default, local endpoint + deterministic mock for tests). Lab/Assessor/Proctor operate on mock engine inputs (simulated telemetry, pre-baked artifacts) — their real engines (sandbox fabric, assessment engine, identity verification) are v0.3+.
## Prior Milestone: v0.4 — Distribution & Bootstrap CLI (COMPLETE, shipped as v0.3.4)
**Status of v0.1:** Complete and shipped (v0.1.0). Founder agreement recorded (D-013).
**Scope (founder directive, 2026-09-12):** Streamline installing Nextcraft. Ship a bootstrap CLI with a single-liner install script, and publish release binaries on an ongoing basis for every release going forward.
**Delivered:** `nextcraft` CLI — `doctor` (prerequisite checks), `bootstrap` (deps + venv + env from templates + key validation), `verify` (health check), `dev` (thin passthrough to scripts/dev.sh); one-liner install script downloading the linux x64 binary from the latest Gitea release with sha256 + version integrity gates; binary release pipeline attached to every ship from v0.3.2 onward; install/quickstart documentation backed by a fresh-clone E2E test. All 5 requirements (REQ-4-001..005) complete.
**Status of v0.3:** Complete and shipped (v0.2.8). Credential engines live: namespace-isolated sandbox fabric, live build telemetry, process-trace grading, seeded variants, oral defense; real learner build/defense/grading surfaces.
**Status of v0.2:** Complete and shipped (v0.2.0). Six AI tutor agents live over mockengine inputs (D-015).
**Deferred from earlier plan:** REQ-F-017 (identity verification + age-gating KYC), real server STT/TTS, design/simulation sandbox environments, and the exec-telemetry seq-lease are all deferred to v0.5. Age-gating remains the v0.1-style visual flow mockup.
**Tech stack:** v0.1 TS monorepo (pnpm/turborepo, Next.js) + v0.2 Python FastAPI ai-service + new credential-engine services (sandbox fabric orchestrator, telemetry ingest, grading engine) in Python/TypeScript as determined at RESEARCH.
**Tech stack:** v0.1 TS monorepo (pnpm/turborepo, Next.js) + new Python FastAPI service (`apps/ai-service`) with pydantic, SSE streaming, and an OpenAI-compatible provider client.
---
## v0.4 Requirements (Complete)
## Requirements (Validated)
All 5 v0.4 requirements (REQ-4-001..005) are complete and shipped as v0.3.4:
1. Bootstrap CLI — `nextcraft` executable with `doctor` / `bootstrap` / `verify` / `dev` (REQ-4-001, REQ-4-002)
2. One-liner install — `curl | sh` fetching the linux x64 binary from the latest Gitea release with sha256 + version integrity verification (REQ-4-003)
3. Ongoing release binaries — every release from v0.3.2 onward ships the CLI binary + checksum as release assets (REQ-4-004)
4. Install documentation — README quickstart + CLI reference verified by a fresh-clone E2E test (REQ-4-005)
The following requirements have been validated during specification and are locked for milestone v0.2 (REQ-F-001..006 activated from the deferred pool):
## v0.3 Requirements (Complete)
All 8 v0.3 requirements (REQ-3-001..008) are complete and shipped as v0.2.8. See REQUIREMENTS.md traceability matrix.
1. AI tutor service infrastructure — `apps/ai-service` FastAPI application, provider-agnostic LLM client, SSE streaming, session/state handling
2. Agent framework — base agent contracts, prompt management, streaming pipeline, structured outputs
3. Coach agent — pacing, motivation, retrieval practice (REQ-F-001)
4. Tutor agent — concept delivery, Socratic questioning (REQ-F-002)
5. Lab agent — in-flow feedback over simulated sandbox telemetry (REQ-F-003, mock inputs)
6. Assessor agent — rubric application to pre-baked artifacts and defenses (REQ-F-004, mock inputs)
7. Proctor agent — integrity signals from mock telemetry, coaching interventions (REQ-F-005, mock inputs)
8. Mentor agent — long-horizon career narrative (REQ-F-006)
9. Learner surface integration — streaming chat UI wired to the real service, error/loading states
## v0.1 Requirements (Complete)
All 28 v0.1 requirements (REQ-001..028) are complete and shipped as v0.1.0. See REQUIREMENTS.md traceability matrix.
## Clarified Assumptions (v0.4 CLARIFY stage, full autonomy — auto-resolved)
| # | Ambiguity | Resolution | Confidence |
|---|-----------|------------|-------------|
| A-201 | CLI language/toolchain for the binary? | **Probe-driven at RESEARCH** — Go → Rust → Node SEA → Python zipapp fallback chain; spec stays toolchain-agnostic so PLAN locks the probe-verified toolchain | 0.70 |
| A-202 | Does bootstrap replace scripts/bootstrap.sh? | **No — reuse it.** CLI wraps existing `scripts/bootstrap.sh` + `scripts/dev.sh` via subprocess; zero orchestration logic duplicated in the CLI (thin passthrough pattern) | 0.85 |
| A-203 | Where does the one-liner fetch the binary? | **Gitea latest-release API** (`/repos/{owner}/{repo}/releases/latest`) → download `nextcraft-linux-x64` + `.sha256` asset; repo raw serves `install.sh` as the stable URL | 0.80 |
| A-204 | Install target + PATH? | **~/.local/bin** (XDG-style, no sudo), PATH hint printed when missing; `--dest` override flag | 0.85 |
| A-205 | Binary "ongoing releases" scope? | **Every ship from v0.4 onward** attaches `nextcraft-linux-x64` + sha256 sidecar to the Gitea release — the ship workflow gains an asset step; retroactive binaries for old releases NOT required | 0.90 |
| A-206 | No binary available yet / non-linux? | **Graceful degradation**: install script prints source-bootstrap instructions (git clone + scripts/bootstrap.sh) — never a hard fail | 0.88 |
| A-207 | Checksum trust root? | **sha256 sidecar shipped as a release asset next to the binary** (same release, same channel); script verifies download against it. Signature/PKI out of scope for v0.4 (single forge, TLS transport) | 0.75 |
| A-208 | Which prerequisites does doctor check? | node ≥18, pnpm ≥8, python3 ≥3.11, git, `unshare` availability (sandbox fabric needs it) — versions from the existing bootstrap tooling, not invented | 0.85 |
| A-209 | Does `dev` manage multiple processes? | **No.** Thin passthrough to scripts/dev.sh only — the CLI stays bootstrap-scoped (D-016); orchestration remains in dev.sh | 0.82 |
| A-210 | `.env.secrets` handling by bootstrap? | **Template copy only for `.env.example` → `.env`; secrets NEVER generated, NEVER committed; bootstrap validates presence of optional keys and warns (not blocks) when missing — mock-first providers keep the stack runnable** | 0.90 |
## Clarified Assumptions (v0.3 CLARIFY stage, full autonomy — auto-resolved)
| # | Ambiguity | Resolution | Confidence |
|---|-----------|------------|-------------|
| A-101 | Sandbox isolation technology? | **`unshare` user+mount+pid+net namespace subprocess isolation** per sandbox (probe-verified: in-ns uid=0, network fully isolated with 0 interfaces, writes land in an isolated bind-mounted workdir; proc-remount is not permitted in this context but is not required). No Docker/Podman/VMs — none present on the box; no sudo. A `SandboxBackend` protocol keeps a future containerd swap possible. Falls back further to a plain chroot-free subprocess with a cwd-jail if userns ever unavailable (tested path is userns). | 0.8 |
| A-102 | Sandbox scope in v0.3? | **Coding IDE only** (web terminal + file tree + run/test). The "design tool" and "simulation" environments specified in REQ-F-021 are deferred to v0.4 — a single real build environment is enough to prove the credential pipeline end-to-end (telemetry → trace → grade → defense). | 0.75 |
| A-103 | Live in-browser build UX? | **Run/Test buttons executing in the namespace sandbox + HTTP file-tree/CRUD + read-only exec-output panel** (CUT-2/G-8 — the interactive xterm.js shell relay is deferred to v0.4; `@xterm/*` is not a v0.3 dependency). No full Monaco LSP in v0.3 — a code editor with syntax highlight (existing) is sufficient and far cheaper. | 0.72 |
| A-104 | Telemetry transport? | **WebSocket** from sandbox to a new ingestion endpoint on ai-service for live events; **SQLite-backed** ordered event log (`ai_service/telemetry/`) gives durability + at-least-once delivery + replay. Events carry monotonic `seq` per (learner,task) so gaps are detectable. | 0.8 |
| A-105 | Where do traces live? | **SQLite** (`ai_service` data dir), introducing the first real persistence. SQLModel/SQLAlchemy for typed access. Chosen over Postgres because solo-founder + single box + low write volume; the `TraceStore` protocol is Postgres-migration-ready like SessionStore was. | 0.75 |
| A-106 | Process-trace grading model? | **LLM-based grader**: structure the trace into a compact timeline digest (command categories, error/fix cycles, idle gaps, test passes) → Assessor-style rubric prompt → structured score via existing D-020 JSON defense. Deterministic features (test pass/fail, edit count) computed in code, not left to the LLM. | 0.7 |
| A-107 | Variant generation mechanism? | **Parameterized task templates + LLM instantiation**, seeded per learner. Generator fills typed parameter slots (scenario, constraints, data) from a template library; variant seed + parameters persisted to SQLite for grading fairness and proctoring cross-check. Difficulty normalized by template-level rubric anchors. | 0.72 |
| A-108 | Voice defense — STT/TTS providers? | **Provider-agnostic, mock-first like the LLM layer (D-014).** Real path: browser `MediaRecorder` → audio to ai-service → **OpenAI-compatible `/audio/transcriptions`** (Whisper STT) and **`/audio/speech`** (TTS) against ollama-cloud or a compatible endpoint; fallbacks: browser `SpeechRecognition`/`speechSynthesis` when no server keys. `VoiceProvider` protocol + deterministic mock (returns canned transcript) so tests never call a voice API. | 0.62 |
| A-109 | Defense dialogue shape? | Reuse BaseAgent: an `Examiner` agent (seventh agent) streams examiner questions over the existing SSE pipeline; integrity signals (long pauses, off-scope answers, reading-from-notes cadence) emitted alongside the transcript to Proctor. | 0.8 |
| A-110 | KYC / age-gating in v0.3? | **Deferred per founder directive.** No real identity backend. Age-gating stays the v0.1 visual flow mockup. Personas omit a security-engineer; security review via verifier + Phase 7 secrets-hygiene checklist. **Abuse control is NOT deferred with KYC (G-5):** v0.3 ships per-learner sandbox caps (`AI_SANDBOX_MAX_PER_LEARNER`), a global create-rate cap, and a server-side `learner_id` allowlist (`AI_LEARNER_ALLOWLIST`) so the unauthenticated surface cannot exhaust shared NPROC/disk. Documented in the release note. | 0.98 |
| A-111 | New services vs extend ai-service? | **Extend ai-service**, don't fork new Python apps. Telemetry ingestion, trace grading, variant generation, voice, and sandbox orchestration all live as new modules in `apps/ai-service` (they share the LLM provider pool + config + session infra). Only the in-sandbox capture agent is a separate tiny Python process shipped into the namespace sandbox. | 0.82 |
| A-112 | Sandbox on a single dev/school box — capacity? | v0.3 targets **15 concurrent sandboxes** (founder + pilot learners). No horizontal scaling, no queue. Concurrency guard returns 503 when full. Scaling is post-MVP. | 0.8 |
## Clarified Assumptions (v0.2 CLARIFY stage, full autonomy — auto-resolved)
## Clarified Assumptions (CLARIFY stage, full autonomy — auto-resolved)
| # | Ambiguity | Resolution | Confidence |
|---|-----------|------------|-------------|
@@ -93,10 +53,14 @@ All 28 v0.1 requirements (REQ-001..028) are complete and shipped as v0.1.0. See
## Requirements (Active — Future Milestones)
The following remain deferred beyond v0.3 and will be activated in subsequent milestones:
The following remain deferred beyond v0.2 and will be activated in subsequent milestones:
The following remain deferred beyond v0.2 and will be activated in subsequent milestones:
- Identity verification and age-gating logic (16+/18+) — the real KYC backend (**deferred from v0.3 per founder directive**; visual flow already exists in v0.1)
- Competency graph engine and adaptive pathways
- Assessment engine (process-trace grading, oral defense, per-learner variant tasks) — v0.3+
- Sandbox fabric (sandboxed IDE, design tool, simulation) — v0.3+
- Identity verification and age-gating logic (16+/18+) — the real KYC backend (v0.3+; visual flow already exists in v0.1)
- Marketplace job aggregation pipeline (3M+ jobs from 120K companies)
- AI-powered tagging, semantic vector search, company enrichment
- AI resume parsing and job matching
@@ -142,7 +106,7 @@ The following remain deferred beyond v0.3 and will be activated in subsequent mi
| D-003 | TypeScript monorepo (pnpm/turborepo) + Next.js | Unified codebase for all 4 surfaces. Shared component library, types, mock data. Next.js App Router for route-based surface separation. Python AI services deferred to later milestones. | Monorepo structure with apps/web + packages/* |
| D-004 | All 4 surfaces in v0.1 (Learner, Marketplace, Employer, Admin) | Founder selected all 4 surfaces for the prototype. Complete product visualization before any backend work. | 24 REQ-IDs covering all surfaces + shared infrastructure |
| D-005 | High-fidelity interactive prototype | Founder selected high-fidelity over wireframes. Realistic mock data, navigation flows, responsive layouts, component library. No backend calls. | Clickable prototype with realistic content |
| D-006 | Release forge = Gitea @ git.coreci.dev, owner=coreci, repo=nextcraft | Founder-provided Gitea instance for release management. Token stored in .ciagent/.env.secrets. Forge migrated 2026-09-12 from git.cloudinit.dev → git.coreci.dev (old host decommissioned; all releases migrated, IDs preserved). | Ship workflow creates tags + releases on Gitea |
| D-006 | Release forge = Gitea @ git.cloudinit.dev, owner=coreci, repo=nextcraft | Founder-provided Gitea instance for release management. Token stored in .ciagent/.env.secrets. | Ship workflow creates tags + releases on Gitea |
| D-007 | Full autonomy for CIAgent pipeline | Founder selected full autonomy. No HITL after clarify. Auto-decide above confidence 0.60. Escalation hooks: deploy, delete_data, merge_to_main. | Rapid autonomous building with kill criteria |
| D-008 | Shared component library in packages/ui/ | All surfaces share a unified design system with surface-specific theming via CSS variables. Promotes consistency and reduces duplication. | packages/ui, packages/mock-data, packages/types |
| D-009 | AI tutor UI as chat interface mockup with pre-scripted responses | The learner surface includes an AI tutor chat UI mockup. No real AI backend — pre-scripted responses simulate the Coach and Tutor agents. | Mockup only in v0.1, real agents in future milestone |
@@ -152,7 +116,6 @@ The following remain deferred beyond v0.3 and will be activated in subsequent mi
| D-013 | v0.1 prototype founder-agreed; D-001 business-logic gate unlocked | Founder approved starting v0.2 with AI Tutor Architecture, which constitutes agreement of the v0.1 prototype per D-001. Recorded at v0.2 SPECIFY. | Business logic authorized from v0.2 onward |
| D-014 | Provider-agnostic LLM layer; ollama-cloud as initial provider | OpenAI-compatible client abstraction with pluggable providers: ollama-cloud (https://ollama.com/v1, default), local OpenAI-compatible endpoint, deterministic mock (tests/CI). Keys in gitignored .ciagent/.env.secrets, never in code or commits. | apps/ai-service llm package with 3 providers; default=ollama-cloud |
| D-015 | All six agents implemented as real LLM services; engines mocked | Coach/Tutor/Mentor fully real. Lab/Assessor/Proctor are real LLM logic over mock inputs (simulated telemetry, pre-baked artifacts) since sandbox fabric, assessment engine, and identity verification are v0.3+. Consistent with v0.1's mock-data approach. | REQ-F-001..006 complete in v0.2; real engines deferred to v0.3+ |
| D-016 | v0.4 = Distribution & Bootstrap CLI (founder directive supersedes previously-named v0.4 seams) | Founder directive 2026-09-12: focus this milestone on streamlining install, a bootstrap CLI with a one-liner install script, ongoing release binaries. Real server STT/TTS, KYC, design/simulation envs, seq-lease move to v0.5. | Milestone scope locked at SPECIFY; binary = CLI-only, linux x64 |
---
+59 -126
View File
@@ -1,77 +1,33 @@
# Nextcraft — REQUIREMENTS.md
## v0.4 Requirements (Complete — Distribution & Bootstrap CLI, shipped as v0.3.4)
### Bootstrap CLI
| ID | Description | Priority | Phase | Status |
|----|-------------|----------|-------|--------|
| REQ-4-001 | `nextcraft` CLI (linux x64 binary): `doctor` command checking prerequisites (node, pnpm, python3, git, unshare) with actionable error messages | critical | 1 | complete |
| REQ-4-002 | `bootstrap` command: pnpm install, ai-service venv + pinned deps, .env from templates, key validation, .env.secrets handling; `verify` health check (ports, imports, builds); `dev` thin passthrough to scripts/dev.sh | critical | 1 | complete |
### Distribution
| ID | Description | Priority | Phase | Status |
|----|-------------|----------|-------|--------|
| REQ-4-003 | One-liner install script (`curl -fsSL <url> \| bash`): detects linux x64, resolves latest release from Gitea API, downloads binary + checksum, verifies sha256, installs to ~/.local/bin (PATH hint), degrades to source-bootstrap instructions when no binary | critical | 2 | complete |
| REQ-4-004 | Binary release pipeline: reproducible linux x64 build script, sha256 checksum sidecar, upload as release assets on every ship from v0.4 onward (ongoing binaries requirement) | critical | 2 | complete |
| REQ-4-005 | Install + quickstart documentation: README one-liner quickstart, CLI command reference, fresh-clone-to-running-stack end-to-end verification | high | 3 | complete |
## v0.3 Requirements (Credential Engines)
### Sandbox & Telemetry
| ID | Description | Priority | Phase | Status |
|----|-------------|----------|-------|--------|
| REQ-3-001 | Sandbox fabric: isolated per-learner execution environments (sandboxed IDE, design tool, simulation) with lifecycle management | critical | 1 | complete |
| REQ-3-002 | Sandbox isolation + resource limits: per-learner isolation boundary, CPU/memory quotas (rlimits), wall-clock time quota, disk-quota via per-sandbox workdir usage sweep (best-effort, not kernel-enforced), no cross-tenant access, snapshot support. **Known gap (v0.3): per-sandbox pids and hard disk caps are NOT kernel-enforceable without cgroup delegation/sudo — documented as accepted risk** | critical | 1 | complete |
| REQ-3-003 | Live build telemetry: in-environment capture of process events (commands, file diffs, run/test results, activity) streamed reliably to ai-service with per-learner trace persistence | critical | 2 | complete |
### Credential Engines
| ID | Description | Priority | Phase | Status |
|----|-------------|----------|-------|--------|
| REQ-3-004 | Process-trace grading engine: grade artifacts from their full process traces; rubric-aligned structured scores; feeds Assessor real inputs | critical | 3 | complete |
| REQ-3-005 | Variant task generation: per-learner task variants (no two learners get identical prompts); variant seed registry; difficulty normalization | high | 4 | complete |
| REQ-3-006 | Oral/voice defense: AI examiner conducts spoken defense (STT → dialogue → TTS); transcript + integrity signals captured; feeds Proctor/Mentor | high | 5 | complete |
### Agent Re-grounding & Integration
| ID | Description | Priority | Phase | Status |
|----|-------------|----------|-------|--------|
| REQ-3-007 | Agent re-grounding: Lab consumes live telemetry; Assessor consumes grading-engine output; Proctor consumes telemetry + defense integrity signals (replace v0.2 mocks) | critical | 6 | complete |
| REQ-3-008 | Learner surface integration: sandbox mockup → real in-browser build/run with live telemetry; assessment mockup → live defense + live grading | critical | 6 | complete |
---
## v0.2 Requirements (Complete — AI Tutor Architecture)
## v0.2 Requirements (AI Tutor Architecture)
### AI Service Infrastructure
| ID | Description | Priority | Phase | Status |
|----|-------------|----------|-------|--------|
| REQ-2-001 | apps/ai-service scaffolding: Python FastAPI app, pydantic settings, uvicorn, health endpoint, CORS, pytest setup, pnpm/turbo integration scripts | critical | 1 | complete |
| REQ-2-002 | Provider-agnostic LLM client: OpenAI-compatible provider interface with ollama-cloud (default), local-endpoint, and deterministic mock providers; key resolution from env files | critical | 1 | complete |
| REQ-2-003 | SSE streaming endpoint plumbing: chat completion streaming from provider through FastAPI to the Next.js client | critical | 1 | complete |
| REQ-2-004 | Agent framework: base agent contracts, session/state store, prompt management, streaming pipeline, structured output support | critical | 2 | complete |
| REQ-2-001 | apps/ai-service scaffolding: Python FastAPI app, pydantic settings, uvicorn, health endpoint, CORS, pytest setup, pnpm/turbo integration scripts | critical | 1 | pending |
| REQ-2-002 | Provider-agnostic LLM client: OpenAI-compatible provider interface with ollama-cloud (default), local-endpoint, and deterministic mock providers; key resolution from env files | critical | 1 | pending |
| REQ-2-003 | SSE streaming endpoint plumbing: chat completion streaming from provider through FastAPI to the Next.js client | critical | 1 | pending |
| REQ-2-004 | Agent framework: base agent contracts, session/state store, prompt management, streaming pipeline, structured output support | critical | 2 | pending |
### AI Tutor Agents
| ID | Description | Priority | Phase | Status |
|----|-------------|----------|-------|--------|
| REQ-2-005 | Coach agent (REQ-F-001): pacing, motivation, retrieval practice — full LLM implementation | critical | 3 | complete |
| REQ-2-006 | Tutor agent (REQ-F-002): concept delivery, Socratic questioning — full LLM implementation | critical | 3 | complete |
| REQ-2-007 | Lab agent (REQ-F-003): in-flow feedback over simulated sandbox telemetry (mock inputs) | high | 4 | complete |
| REQ-2-008 | Assessor agent (REQ-F-004): rubric application to pre-baked artifacts and defense transcripts (mock inputs) | high | 4 | complete |
| REQ-2-009 | Proctor agent (REQ-F-005): integrity signals from mock telemetry, coaching interventions | high | 5 | complete |
| REQ-2-010 | Mentor agent (REQ-F-006): long-horizon career narrative | high | 5 | complete |
| REQ-2-005 | Coach agent (REQ-F-001): pacing, motivation, retrieval practice — full LLM implementation | critical | 3 | pending |
| REQ-2-006 | Tutor agent (REQ-F-002): concept delivery, Socratic questioning — full LLM implementation | critical | 3 | pending |
| REQ-2-007 | Lab agent (REQ-F-003): in-flow feedback over simulated sandbox telemetry (mock inputs) | high | 4 | pending |
| REQ-2-008 | Assessor agent (REQ-F-004): rubric application to pre-baked artifacts and defense transcripts (mock inputs) | high | 4 | pending |
| REQ-2-009 | Proctor agent (REQ-F-005): integrity signals from mock telemetry, coaching interventions | high | 5 | pending |
| REQ-2-010 | Mentor agent (REQ-F-006): long-horizon career narrative | high | 5 | pending |
### Learner Surface Integration
| ID | Description | Priority | Phase | Status |
|----|-------------|----------|-------|--------|
| REQ-2-011 | Learner chat UI wired to real service: streaming responses, agent routing, error and loading states | critical | 6 | complete |
| REQ-2-012 | Byte tutorial viewer, build sandbox, and assessment mockups surface Lab/Assessor/Proctor outputs (mock engine inputs) | high | 6 | complete |
| REQ-2-011 | Learner chat UI wired to real service: streaming responses, agent routing, error and loading states | critical | 6 | pending |
| REQ-2-012 | Byte tutorial viewer, build sandbox, and assessment mockups surface Lab/Assessor/Proctor outputs (mock engine inputs) | high | 6 | pending |
---
@@ -82,58 +38,58 @@
| ID | Description | Priority | Phase | Status |
|----|-------------|----------|-------|--------|
| REQ-001 | Monorepo scaffolding: pnpm workspaces, turborepo, Next.js app, TypeScript config, ESLint, Prettier | critical | 1 | complete |
| REQ-002 | Shared component library: design tokens, Button, Input, Card, Badge, Avatar, Dialog, Navigation, Table, Tabs, Progress, Tooltip, Skeleton, Toast | critical | 1 | complete |
| REQ-003 | Mock data layer: typed mock data for competency stacks, job listings, candidate profiles, employers, learner progress | critical | 1 | complete |
| REQ-004 | Routing and navigation: App Router route groups for (learner), (marketplace), (employer), (admin); shared layout components; cross-surface navigation | critical | 1 | complete |
| REQ-005 | Responsive layout system: mobile, tablet, desktop breakpoints; container components; grid system | high | 1 | complete |
| REQ-002 | Shared component library: design tokens, Button, Input, Card, Badge, Avatar, Dialog, Navigation, Table, Tabs, Progress, Tooltip, Skeleton, Toast | critical | 1 | pending |
| REQ-003 | Mock data layer: typed mock data for competency stacks, job listings, candidate profiles, employers, learner progress | critical | 1 | pending |
| REQ-004 | Routing and navigation: App Router route groups for (learner), (marketplace), (employer), (admin); shared layout components; cross-surface navigation | critical | 1 | pending |
| REQ-005 | Responsive layout system: mobile, tablet, desktop breakpoints; container components; grid system | high | 1 | pending |
### Learner Surface
| ID | Description | Priority | Phase | Status |
|----|-------------|----------|-------|--------|
| REQ-006 | Landing page: hero, value proposition, program highlights, how-it-works (Byte→Build→Demonstrate→Defend), testimonials mockup, CTA to program catalog | critical | 2 | complete |
| REQ-007 | Program catalog: grid of competency stacks (AI Orchestration Engineer, AI Safety & Governance Lead, Human-AI Product Designer, AI-Augmented Field Operator, Computational Sciences Practitioner); stack cards with role descriptions | critical | 2 | complete |
| REQ-008 | Competency stack view: selected stack detail with 12-18 competencies listed, progress indicators, microcredential badges, mastery status | critical | 2 | complete |
| REQ-009 | Learner dashboard: active competencies, progress graph, recent artifacts, upcoming defenses, AI tutor chat mockup, milestone tracker | critical | 2 | complete |
| REQ-010 | Byte tutorial viewer: 3-7 minute micro-tutorial layout with concept panel, worked example panel, code/design/simulation viewer mockup | high | 2 | complete |
| REQ-011 | Build sandbox mockup: sandboxed IDE/design tool/simulation UI mockup with toolbar, file explorer, editor area, telemetry sidebar (process capture indicators) | high | 2 | complete |
| REQ-012 | Assessment/defense mockup: rubric display, AI reviewer panel, oral defense interface (voice/mic mockup), process trace timeline, artifact viewer | high | 2 | complete |
| REQ-006 | Landing page: hero, value proposition, program highlights, how-it-works (Byte→Build→Demonstrate→Defend), testimonials mockup, CTA to program catalog | critical | 2 | pending |
| REQ-007 | Program catalog: grid of competency stacks (AI Orchestration Engineer, AI Safety & Governance Lead, Human-AI Product Designer, AI-Augmented Field Operator, Computational Sciences Practitioner); stack cards with role descriptions | critical | 2 | pending |
| REQ-008 | Competency stack view: selected stack detail with 12-18 competencies listed, progress indicators, microcredential badges, mastery status | critical | 2 | pending |
| REQ-009 | Learner dashboard: active competencies, progress graph, recent artifacts, upcoming defenses, AI tutor chat mockup, milestone tracker | critical | 2 | pending |
| REQ-010 | Byte tutorial viewer: 3-7 minute micro-tutorial layout with concept panel, worked example panel, code/design/simulation viewer mockup | high | 2 | pending |
| REQ-011 | Build sandbox mockup: sandboxed IDE/design tool/simulation UI mockup with toolbar, file explorer, editor area, telemetry sidebar (process capture indicators) | high | 2 | pending |
| REQ-012 | Assessment/defense mockup: rubric display, AI reviewer panel, oral defense interface (voice/mic mockup), process trace timeline, artifact viewer | high | 2 | pending |
### Marketplace Surface
| ID | Description | Priority | Phase | Status |
|----|-------------|----------|-------|--------|
| REQ-013 | Job board listing: searchable grid of AI-era job listings, filter sidebar (skills, seniority, location, salary), result cards with match score | critical | 3 | complete |
| REQ-014 | Job detail page: full job description, required competencies, employer info, application CTA, related jobs, AI-matched skills breakdown | critical | 3 | complete |
| REQ-015 | Employer profile: company overview, logo, description, social links, open positions, company culture mockup | high | 3 | complete |
| REQ-016 | Search/filter UI: semantic search bar, skill tags, category filters, seniority filter, remote/on-site toggle, saved searches mockup | critical | 3 | complete |
| REQ-017 | Pricing page: job posting packages (single, bundle, enterprise), talent access plans, subscription tiers, feature comparison table | high | 3 | complete |
| REQ-013 | Job board listing: searchable grid of AI-era job listings, filter sidebar (skills, seniority, location, salary), result cards with match score | critical | 3 | pending |
| REQ-014 | Job detail page: full job description, required competencies, employer info, application CTA, related jobs, AI-matched skills breakdown | critical | 3 | pending |
| REQ-015 | Employer profile: company overview, logo, description, social links, open positions, company culture mockup | high | 3 | pending |
| REQ-016 | Search/filter UI: semantic search bar, skill tags, category filters, seniority filter, remote/on-site toggle, saved searches mockup | critical | 3 | pending |
| REQ-017 | Pricing page: job posting packages (single, bundle, enterprise), talent access plans, subscription tiers, feature comparison table | high | 3 | pending |
### Employer Dashboard
| ID | Description | Priority | Phase | Status |
|----|-------------|----------|-------|--------|
| REQ-018 | Employer dashboard overview: active postings, applicant pipeline, talent matches, analytics mockup (charts, placement stats) | critical | 4 | complete |
| REQ-019 | Talent search: searchable candidate database with AI-matched filters, candidate cards showing competency stack, microcredentials, artifact count, defense score | critical | 4 | complete |
| REQ-020 | Candidate profile view: full candidate profile with artifact gallery, process trace summary, oral defense transcripts, competency graph, microcredential verification | high | 4 | complete |
| REQ-021 | Posting management: create/edit/delete job postings, posting status tracking, applicant list per posting, interview pipeline mockup | high | 4 | complete |
| REQ-018 | Employer dashboard overview: active postings, applicant pipeline, talent matches, analytics mockup (charts, placement stats) | critical | 4 | pending |
| REQ-019 | Talent search: searchable candidate database with AI-matched filters, candidate cards showing competency stack, microcredentials, artifact count, defense score | critical | 4 | pending |
| REQ-020 | Candidate profile view: full candidate profile with artifact gallery, process trace summary, oral defense transcripts, competency graph, microcredential verification | high | 4 | pending |
| REQ-021 | Posting management: create/edit/delete job postings, posting status tracking, applicant list per posting, interview pipeline mockup | high | 4 | pending |
### Admin Surface
| ID | Description | Priority | Phase | Status |
|----|-------------|----------|-------|--------|
| REQ-022 | Admin overview: platform metrics (learners, employers, placements, completion rate, NPS), recent activity feed, system health mockup | critical | 5 | complete |
| REQ-023 | Learner management: searchable learner table, learner detail view, progress tracking, competency completion status, credential issuance log | high | 5 | complete |
| REQ-024 | Competency graph viewer: interactive visualization of competency stacks and their relationships, node/edge graph using react-flow, stack details on node click | high | 5 | complete |
| REQ-025 | Marketplace moderation: job posting review queue, employer verification queue, flagged content, content moderation tools mockup | high | 5 | complete |
| REQ-022 | Admin overview: platform metrics (learners, employers, placements, completion rate, NPS), recent activity feed, system health mockup | critical | 5 | pending |
| REQ-023 | Learner management: searchable learner table, learner detail view, progress tracking, competency completion status, credential issuance log | high | 5 | pending |
| REQ-024 | Competency graph viewer: interactive visualization of competency stacks and their relationships, node/edge graph using react-flow, stack details on node click | high | 5 | pending |
| REQ-025 | Marketplace moderation: job posting review queue, employer verification queue, flagged content, content moderation tools mockup | high | 5 | pending |
### Polish & Integration
| ID | Description | Priority | Phase | Status |
|----|-------------|----------|-------|--------|
| REQ-026 | Cross-surface navigation: role switcher (learner/employer/admin), breadcrumbs, consistent header/footer across all surfaces | critical | 6 | complete |
| REQ-027 | Visual consistency audit: typography scale, color palette, spacing system, dark mode toggle, accessibility baseline (WCAG AA contrast) | high | 6 | complete |
| REQ-028 | Component library finalization: Storybook setup, component documentation, prop tables, usage examples | medium | 6 | complete |
| REQ-026 | Cross-surface navigation: role switcher (learner/employer/admin), breadcrumbs, consistent header/footer across all surfaces | critical | 6 | pending |
| REQ-027 | Visual consistency audit: typography scale, color palette, spacing system, dark mode toggle, accessibility baseline (WCAG AA contrast) | high | 6 | pending |
| REQ-028 | Component library finalization: Storybook setup, component documentation, prop tables, usage examples | medium | 6 | pending |
---
@@ -143,11 +99,11 @@
| ID | Description | Priority | Milestone | Status |
|----|-------------|----------|-----------|--------|
| REQ-F-007 | Process-trace grading engine → activated as REQ-3-004 | high | v0.3 | activated |
| REQ-F-008 | Per-learner variant task generation → activated as REQ-3-005 | high | v0.3 | activated |
| REQ-F-009 | Oral/voice defense with AI examiner → activated as REQ-3-006 | high | v0.3 | activated |
| REQ-F-010 | Live in-environment build with telemetry → activated as REQ-3-003 | high | v0.3 | activated |
| REQ-F-021 | Sandbox fabric: sandboxed IDE, design tool, simulation → activated as REQ-3-001/002 | high | v0.3 | activated |
| REQ-F-007 | Process-trace grading engine | high | v0.3+ | deferred |
| REQ-F-008 | Per-learner variant task generation | high | v0.3+ | deferred |
| REQ-F-009 | Oral/voice defense with AI examiner | high | v0.3+ | deferred |
| REQ-F-010 | Live in-environment build with telemetry | high | v0.3+ | deferred |
| REQ-F-021 | Sandbox fabric: sandboxed IDE, design tool, simulation | high | v0.3+ | deferred |
### Marketplace Engine
@@ -164,7 +120,7 @@
| ID | Description | Priority | Milestone | Status |
|----|-------------|----------|-----------|--------|
| REQ-F-017 | Identity verification and age-gating (16+/18+) — real KYC backend | high | v0.4+ | deferred (deferred from v0.3 per founder directive) |
| REQ-F-017 | Identity verification and age-gating (16+/18+) — real KYC backend | high | v0.3+ | deferred |
| REQ-F-018 | Payment processing and subscription management | high | v0.3+ | deferred |
| REQ-F-019 | Human tutor marketplace (third-party courses) | medium | v0.4+ | deferred |
| REQ-F-020 | CIRR-style placement tracking and audit | medium | v0.5+ | deferred |
@@ -187,45 +143,22 @@
## Traceability Matrix
### v0.4 (current milestone)
### v0.2 (current milestone)
| Requirement | Phase | Status |
|-------------|-------|--------|
| REQ-4-001 | 1 | complete |
| REQ-4-002 | 1 | complete |
| REQ-4-003 | 2 | complete |
| REQ-4-004 | 2 | complete |
| REQ-4-005 | 3 | complete |
### v0.3 (complete)
| Requirement | Phase | Status |
|-------------|-------|--------|
| REQ-3-001 | 1 | complete |
| REQ-3-002 | 1 | complete |
| REQ-3-003 | 2 | complete |
| REQ-3-004 | 3 | complete |
| REQ-3-005 | 4 | complete |
| REQ-3-006 | 5 | complete |
| REQ-3-007 | 6 | complete |
| REQ-3-008 | 6 | complete |
### v0.2 (complete)
| Requirement | Phase | Status |
|-------------|-------|--------|
| REQ-2-001 | 1 | complete |
| REQ-2-002 | 1 | complete |
| REQ-2-003 | 1 | complete |
| REQ-2-004 | 2 | complete |
| REQ-2-005 | 3 | complete |
| REQ-2-006 | 3 | complete |
| REQ-2-007 | 4 | complete |
| REQ-2-008 | 4 | complete |
| REQ-2-009 | 5 | complete |
| REQ-2-010 | 5 | complete |
| REQ-2-011 | 6 | complete |
| REQ-2-012 | 6 | complete |
| REQ-2-001 | 1 | pending |
| REQ-2-002 | 1 | pending |
| REQ-2-003 | 1 | pending |
| REQ-2-004 | 2 | pending |
| REQ-2-005 | 3 | pending |
| REQ-2-006 | 3 | pending |
| REQ-2-007 | 4 | pending |
| REQ-2-008 | 4 | pending |
| REQ-2-009 | 5 | pending |
| REQ-2-010 | 5 | pending |
| REQ-2-011 | 6 | pending |
| REQ-2-012 | 6 | pending |
### v0.1 (complete)
+110 -72
View File
@@ -2,17 +2,13 @@
## Overview
**Milestone v0.4 — COMPLETE (shipped as v0.3.4, 2026-09-13).** Distribution & Bootstrap CLI: `nextcraft` CLI (doctor/bootstrap/verify/dev) shipped as a linux x64 SEA binary, one-liner install script with checksum + version integrity gates, and binaries published on **every ongoing release** (v0.3.2 onward). Next milestone: v0.5 (real server STT/TTS + KYC/identity + design/simulation sandbox environments + exec-telemetry seq-lease — the seams deferred out of v0.4 by founder directive D-016).
**Milestone v0.2** — AI Tutor Architecture: The six AI tutor agents (Coach, Tutor, Lab, Assessor, Proctor, Mentor) as real LLM-backed services in a new `apps/ai-service` Python FastAPI application, wired into the existing v0.1 learner surface with streaming responses. Provider-agnostic LLM layer (ollama-cloud default). Lab/Assessor/Proctor operate on mock engine inputs — their real engines are v0.3+.
**Milestone v0.3** — Credential Engines: complete, shipped as v0.2.8 (2026-09-12). Real sandbox fabric, live build telemetry, process-trace grading, per-learner variants, oral defense, real learner surfaces.
**Prior milestone:** v0.1 (nextcraft-ui-prototype) — complete, shipped as v0.1.0, founder-agreed (D-013).
**Deferred per founder directive (D-016):** REQ-F-017 identity verification + age-gating (real KYC backend), real server STT/TTS, design/simulation sandbox environments, and the exec-telemetry seq-lease are deferred to v0.5. Age-gating remains the v0.1 visual flow mockup.
**Prior milestone:** v0.2 (ai-tutor-architecture) — complete, shipped as v0.2.0, six tutor agents live over mock engine inputs (D-015).
**Milestone type:** Feature (new CLI + distribution pipeline)
**Tag line:** v0.3.x (patches on the v0.3 line; milestone release as the final v0.3.x patch)
**Branch:** milestone/v0.4-distribution
**Milestone type:** Feature (new AI service + real agent capabilities)
**Tag line:** v0.1.x (patches on the v0.1 line; milestone release as v0.2.0)
**Branch:** milestone/v0.2-ai-tutor-architecture
---
@@ -20,11 +16,14 @@
| # | Name | Status | Depends On | Requirements | Success Criteria |
|---|------|--------|------------|--------------|------------------|
| 0 | Pre-execution | complete | — | — | Specification, clarify, research, plan complete; .ciagent/ files updated for v0.4 |
| 1 | Bootstrap CLI core | complete | 0 | REQ-4-001, REQ-4-002 | `nextcraft doctor/bootstrap/verify/dev` work against a fresh clone; unit tests green |
| 2 | Binary build + release pipeline | complete | 1 | REQ-4-003, REQ-4-004 | Reproducible linux x64 binary + sha256 checksum; one-liner install script; assets uploaded to the Gitea release |
| 3 | Install docs + fresh-clone E2E | complete | 2 | REQ-4-005 | README quickstart verified end-to-end from a clean environment; fresh clone reaches running stack |
| 4 | Final review + ship | complete | 3 | — | Code review clean; audit passes; milestone tagged (v0.3.x final patch); release with binary assets created on Gitea |
| 0 | Pre-execution | in_progress | — | — | Specification, clarify, research, plan complete; .ciagent/ files updated for v0.2 |
| 1 | AI service scaffolding | pending | 0 | REQ-2-001, REQ-2-002, REQ-2-003 | apps/ai-service runs (uvicorn), health endpoint responds, provider-agnostic LLM client with 3 providers (ollama-cloud/local/mock), SSE streaming verified, pytest suite passes with mock provider, turbo scripts wired |
| 2 | Agent framework | pending | 1 | REQ-2-004 | Base agent contract, session/state store, prompt templates, streaming pipeline, structured outputs; all tested |
| 3 | Coach + Tutor agents | pending | 2 | REQ-2-005, REQ-2-006 | Coach (pacing/motivation/retrieval practice) and Tutor (concept delivery/Socratic questioning) fully implemented with system prompts, tested against mock provider, wired to chat endpoint |
| 4 | Lab + Assessor agents | pending | 2 | REQ-2-007, REQ-2-008 | Lab consumes simulated sandbox telemetry (mock); Assessor applies rubrics to pre-baked artifacts/defense transcripts (mock); both tested |
| 5 | Proctor + Mentor agents | pending | 2 | REQ-2-009, REQ-2-010 | Proctor produces integrity signals + coaching interventions from mock telemetry; Mentor generates long-horizon career narrative; both tested |
| 6 | Learner surface integration | pending | 3, 4, 5 | REQ-2-011, REQ-2-012 | Learner chat streams real responses; agent routing works; byte viewer/sandbox/assessment mockups surface agent outputs; error/loading states; pnpm build + typecheck pass |
| 7 | Final review + ship | pending | 6 | — | Code review clean; audit passes; milestone tagged v0.2.0; release created on Gitea |
---
@@ -32,111 +31,150 @@
### Phase 0: Pre-execution
**Goal:** Establish v0.4 specification (founder directive D-016), clarify ambiguities, research the binary toolchain + Gitea release-asset API + existing bootstrap scripts, create detailed plans, grill adversarially.
**Goal:** Establish v0.2 specification, clarify ambiguities, research AI service architecture, create detailed plans.
**Stages:** SPECIFY → CLARIFY → RESEARCH → PLAN → GRILL → MVP/UX CHECK → SHIP
**Stages:** SPECIFY → CLARIFY → RESEARCH → PLAN → GRILL → SHIP
**Deliverables:**
- Updated .ciagent/config.json, PROJECT.md, REQUIREMENTS.md, ROADMAP.md, ARCHITECTURE.md, PERSONAS.md, PLAN.md
**Success criteria:** All .ciagent/ files updated for v0.4; phase 0 shipped as v0.3.0.
**Success criteria:** All .ciagent/ files updated for v0.2; phase 0 shipped as v0.1.1.
---
### Phase 1: Bootstrap CLI Core
### Phase 1: AI Service Scaffolding
**Goal:** A working `nextcraft` CLI with doctor/bootstrap/verify/dev commands, unit-tested against the real monorepo.
**Goal:** Stand up apps/ai-service with the provider-agnostic LLM layer and SSE streaming.
**Requirements:** REQ-4-001, REQ-4-002
**Requirements:** REQ-2-001, REQ-2-002, REQ-2-003
**Key deliverables:**
- `apps/cli` package: `nextcraft` executable (source-runnable in dev, binary-built in P2)
- `doctor`: checks node ≥18, pnpm, python3 ≥3.11, git, unshare availability — actionable errors, exit codes
- `bootstrap`: idempotent — pnpm install, ai-service venv + pinned deps (reuses scripts/bootstrap.sh logic), .env from .env.example templates, key validation (warnings not blockers for optional keys), .env.secrets handling
- `verify`: health check — venv imports, pnpm build readiness, ports free, env vars present
- `dev`: thin passthrough to scripts/dev.sh (no orchestration logic duplicated)
- Unit tests: doctor/bootstrap parsing + command dispatch, against fixtures (never modifying the real repo state)
- apps/ai-service: FastAPI app, pydantic-settings, uvicorn, /health, CORS for localhost
- llm package: provider interface + ollama-cloud/local/mock providers; key resolution from .ciagent/.env.secrets via env
- SSE streaming: /v1/chat/stream endpoint streaming provider deltas
- pytest suite with mock provider; root scripts: ai:dev, ai:test; turbo integration
**Success criteria:**
- `nextcraft doctor` reports each prerequisite with actionable guidance
- `nextcraft bootstrap` on a fresh clone reaches a state where `verify` passes
- All commands have `--help`, exit non-zero on failure, no shell-out without timeout
- `pnpm build`, `pnpm typecheck`, `pnpm ai:test` green
- `python -m uvicorn` starts the service; /health returns 200
- Provider unit tests pass (mock); ollama-cloud integration probe works (manual)
- SSE stream delivers tokens to an HTTP client
---
### Phase 2: Binary Build + Release Pipeline
### Phase 2: Agent Framework
**Goal:** Reproducible linux x64 binary + one-liner install + release-asset upload wired into the ship flow.
**Goal:** Build the shared framework all six agents use.
**Requirements:** REQ-4-003, REQ-4-004
**Requirements:** REQ-2-004
**Key deliverables:**
- Build script producing `nextcraft-linux-x64` + `nextcraft-linux-x64.sha256` (toolchain probe-verified at RESEARCH; embedded script assets)
- One-liner install script `install.sh` served from the repo: detect linux x64, resolve latest release via Gitea API, download + verify checksum, install to `~/.local/bin`, PATH hint, source-bootstrap fallback when no binary/asset
- Ship integration: every release from v0.4 onward attaches the binary + checksum as release assets (the "ongoing binaries" requirement)
- Asset-upload helper using the Gitea token from `.env*` files only (never shell env)
- BaseAgent contract: system prompt, message history, streaming completion, structured output
- Session/state store: in-memory per-learner session with message history
- Prompt management: per-agent system prompt templates with learner context injection
- Streaming pipeline: agent → provider → SSE with agent identification
- Structured outputs: JSON-schema outputs for Assessor rubric scores, Proctor signals
**Success criteria:**
- Binary runs on this box: `./nextcraft-linux-x64 doctor` green against the repo
- Install script verified end-to-end against the real Gitea release (or local dry-run if release pending)
- Checksum verification rejects a corrupted download (tested)
- Release assets present on the phase ship
- BaseAgent unit tests pass
- Session store tested (create/append/persist in-memory)
- Structured output parsing tested against mock provider
---
### Phase 3: Install Docs + Fresh-Clone E2E
### Phase 3: Coach + Tutor Agents
**Goal:** Documentation and end-to-end proof that a fresh consumer reaches a running stack via the one-liner.
**Goal:** Implement the two learner-facing conversational agents.
**Requirements:** REQ-4-005
**Requirements:** REQ-2-005, REQ-2-006
**Key deliverables:**
- README quickstart: one-liner → `nextcraft doctor``nextcraft bootstrap``nextcraft dev`
- CLI command reference (all flags, exit codes)
- Fresh-clone E2E test: clean temp clone → doctor → bootstrap → verify → build green (sandboxed; no network beyond package registries already used)
- Install-script docs: prerequisites, offline/manual install, troubleshooting
- Coach agent: pacing guidance, motivation, retrieval practice prompts; distinct persona
- Tutor agent: concept delivery, Socratic questioning, worked examples
- Agent registry: route chat messages to the correct agent by context/selection
- Per-agent system prompts with competency-stack context injection from packages/mock-data
**Success criteria:**
- A fresh clone bootstraps to a passing `verify` with one command sequence
- README quickstart matches the actual tested flow exactly
- E2E test green in CI-equivalent local run
- Both agents produce distinct, on-persona responses (verified against mock + ollama-cloud)
- Agent routing tested
- Both agents exposed via the chat streaming endpoint
---
### Phase 4: Final Review + Ship
### Phase 4: Lab + Assessor Agents
**Goal:** Code review, audit, milestone release with binary assets.
**Goal:** Implement the two build/assessment agents over mock engine inputs.
**Requirements:** REQ-2-007, REQ-2-008
**Key deliverables:**
- Lab agent: consumes simulated sandbox telemetry (mock event streams), produces in-flow feedback
- Assessor agent: applies rubrics to pre-baked artifacts and defense transcripts, returns structured scores + feedback
- Mock engine inputs: simulated telemetry generator, pre-baked artifact corpus in packages/mock-data
- Endpoints: /v1/lab/feedback, /v1/assessment/evaluate
**Success criteria:**
- Lab produces relevant feedback for mock telemetry scenarios
- Assessor returns structured rubric scores (JSON) for pre-baked artifacts
- Both tested against mock provider
---
### Phase 5: Proctor + Mentor Agents
**Goal:** Implement the integrity and narrative agents.
**Requirements:** REQ-2-009, REQ-2-010
**Key deliverables:**
- Proctor agent: integrity signals from mock telemetry (tab switches, idle time, paste events), coaching interventions
- Mentor agent: long-horizon career narrative, competency-stack progression guidance
- Endpoints: /v1/proctor/signals, /v1/mentor/narrative
**Success criteria:**
- Proctor produces classified signals with recommended interventions for mock scenarios
- Mentor produces coherent career-narrative responses
- Both tested against mock provider
---
### Phase 6: Learner Surface Integration
**Goal:** Wire the v0.1 learner surface to the real AI service.
**Requirements:** REQ-2-011, REQ-2-012
**Key deliverables:**
- Learner dashboard chat: real streaming via SSE, agent switcher (Coach/Tutor), error/loading states
- Byte tutorial viewer: Tutor concept explanations
- Build sandbox: Lab feedback panel fed by mock telemetry + Lab agent
- Assessment mockup: Assessor rubric output display, Proctor integrity banner
- Mentor panel on learner dashboard
**Success criteria:**
- Streaming chat works end-to-end with ai-service running
- All four learner surfaces surface agent outputs
- Graceful degradation when ai-service is down (error states, not crashes)
- `pnpm build` and `pnpm typecheck` pass
---
### Phase 7: Final Review + Ship
**Goal:** Code review, audit, milestone release.
**Key deliverables:**
- Multi-persona code review (correctness, testing, security, performance, maintainability)
- Project health audit (reconstruction test, .ciagent/ file discipline, branch hygiene, commit discipline)
- Milestone ship: merge milestone → main, tag final v0.3.x patch, create Gitea release WITH binary + checksum assets, verify assets downloadable
- Milestone ship: merge milestone → main, tag v0.2.0, create Gitea release
**Success criteria:**
- Code review: P0 fixes applied, P1+ documented
- Audit: all checks pass, project state reconstructable from git log
- Ship: milestone tagged, branch merged to main, Gitea release created with `nextcraft-linux-x64` + `.sha256` assets attached — the first of the ongoing binary releases
- Ship: v0.2.0 tagged, milestone branch merged to main, Gitea release created **release note explicitly states Lab/Assessor/Proctor operate on mock engine inputs (real engines v0.3+)** (G-5); dead `aiTutorResponses` export disposed of (G-5)
- All 12 v0.2 requirements marked complete
---
## v0.4 (Complete — Shipped as v0.3.4)
Distribution & Bootstrap CLI: `nextcraft` CLI (doctor/bootstrap/verify/dev) as a self-contained linux x64 SEA binary, one-liner install with sha256 + version integrity gates, release-asset pipeline attaching binaries to every ongoing release (v0.3.2 onward), install/quickstart docs backed by a fresh-clone E2E test. 5 phases (P0P4). All 5 requirements (REQ-4-001..005) complete. Tags v0.3.0v0.3.3 per phase, milestone release v0.3.4.
## v0.3 (Complete — Shipped as v0.2.8)
Credential Engines: real sandbox fabric (Linux namespaces), live build telemetry
(at-least-once/exactly-once), process-trace grading (G-4 gated), seeded per-learner
variants (fairness anchors wired to grading), oral defense with integrity signals
(mock-first voice, browser fallback), and real learner build/defense/grading surfaces.
8 phases. All 8 requirements (REQ-3-001..008) complete. Tags v0.2.1v0.2.7 per phase,
milestone release v0.2.8.
## v0.2 (Complete — Shipped as v0.2.0)
AI Tutor Architecture: Six AI tutor agents (Coach, Tutor, Lab, Assessor, Proctor, Mentor) as real LLM-backed services over mock engine inputs, wired into the learner surface with streaming. 7 phases. All 12 requirements complete. Milestone release v0.2.0.
## v0.1 (Complete — Shipped as v0.1.0)
UI/UX Prototype: High-fidelity interactive prototype of all four Nextcraft surfaces. 7 phases (P0 + P1-P6 execution + P7 final). All 28 requirements complete. Tags v0.0.1v0.0.7, milestone release v0.1.0.
+4 -4
View File
@@ -9,7 +9,7 @@
},
"release": {
"forge": "gitea",
"base_url": "https://git.coreci.dev",
"base_url": "https://git.cloudinit.dev",
"owner": "coreci",
"repo": "nextcraft"
},
@@ -46,9 +46,9 @@
"projects": [],
"active_project": null,
"milestone": {
"version": "v0.4",
"name": "distribution",
"version": "v0.2",
"name": "ai-tutor-architecture",
"type": "feature",
"branch": "milestone/v0.4-distribution"
"branch": "milestone/v0.2-ai-tutor-architecture"
}
}
-10
View File
@@ -48,13 +48,3 @@ coverage/
.pytest_cache/
.ruff_cache/
*.egg-info/
.ciagent/bin/
# v0.3 engine runtime data (SQLite + sandboxes)
apps/ai-service/ai_service/data/
apps/ai-service/**/sandboxes/
*.db
*.db-journal
# in-sandbox capture agent runtime spool
.nc-agent/
+1 -64
View File
@@ -2,71 +2,8 @@
AI-native outcome school + marketplace — graduates prove what they can build, not what they can write.
## Quickstart
One-liner install (linux x64):
```sh
curl -fsSL https://git.coreci.dev/coreci/nextcraft/raw/main/scripts/install.sh | sh
```
That downloads the latest release's `nextcraft` CLI binary, verifies its sha256 checksum, and installs it to `~/.local/bin` (PATH hint printed if needed). Every release ships fresh binaries — re-run the one-liner to upgrade.
Then, from a clone of this repo:
```sh
nextcraft doctor # check prerequisites: node >= 18, pnpm >= 8, python3 >= 3.11, git, unshare
nextcraft bootstrap # pnpm install + ai-service venv + .env from template (idempotent)
nextcraft verify # health check: venv imports, uvicorn, ports, env
nextcraft dev # run the ai-service dev server on :8420 (web dev server: pnpm dev)
```
The E2E test (`apps/cli/tests/fresh-clone-e2e.test.ts`) proves this exact sequence on a fresh clone.
### No binary / non-linux?
The installer degrades to printed source instructions. Manual equivalent:
```sh
git clone https://git.coreci.dev/coreci/nextcraft.git && cd nextcraft
pnpm install
bash apps/ai-service/scripts/bootstrap.sh
cp apps/ai-service/.env.example apps/ai-service/.env
pnpm ai:dev
```
## CLI reference (`nextcraft`)
| Command | What it does | Exit codes |
|---------|--------------|------------|
| `doctor` | Checks prerequisites on PATH: node >= 18, pnpm >= 8, python3 >= 3.11, git, unshare (sandbox fabric). Every ✗ prints a fix hint. | 0 all pass, 1 any fail |
| `bootstrap` | Sets up the monorepo from a fresh clone: (1) locates the repo root, (2) `pnpm install`, (3) ai-service venv via `apps/ai-service/scripts/bootstrap.sh`, (4) copies `.env.example``.env` if absent, (5) warns on missing optional keys. Idempotent — safe to re-run. | 0 ok, 1 step failed |
| `verify` | Health check: ai-service venv + `import ai_service`, uvicorn importable, `.env` present (warn-only), `AI_PORT` (default 8420) free, workspace `node_modules` present. | 0 ok, 1 failures |
| `dev` | Thin passthrough to `apps/ai-service/scripts/dev.sh` (exports secrets from `.ciagent/.env.secrets` if present, runs uvicorn on :8420). Ctrl+C stops it. The web dev server is separate: `pnpm dev`. | child's exit code |
| `--help` / `-h` | Usage for the CLI or any command. | 0 |
| `--version` | Prints the version this binary was built as (matches the release tag). | 0 |
Exit-code contract: `0` success, `1` check/step failure (hint printed), `2` usage error.
### Remote server
`nextcraft dev` binds the API on **0.0.0.0:8420** (and `pnpm dev` serves the web app on all interfaces), so the stack works from other machines out of the box:
- Browse `http://<your-host>:3000` — the web app targets `http://<your-host>:8420` automatically (derived from the browser's hostname).
- CORS admits any origin (`AI_CORS_ORIGINS=*` in `apps/ai-service/.env`). This is safe **only** because credentials are never enabled; to restrict, set an explicit list: `AI_CORS_ORIGINS=http://<your-host>:3000`.
- To revert to loopback-only: `AI_HOST=127.0.0.1` in `apps/ai-service/.env`.
- Security note: this is an unauthenticated dev API reachable from any network the box exposes. Mitigations that still apply: per-learner sandbox caps + global rate caps + learner allowlist (G-5), telemetry flood control (traces marked `INCOMPLETE_FLOODED` are refused by the grader). Expose only on trusted networks until identity/KYC lands (v0.5).
## Docs
- [apps/cli/README.md](apps/cli/README.md) — CLI internals: build, binary pipeline, troubleshooting
- [.ciagent/PROJECT.md](.ciagent/PROJECT.md) — product spec and milestone history
- [.ciagent/ARCHITECTURE.md](.ciagent/ARCHITECTURE.md) — system architecture
## Status
**Milestone v0.4**Distribution & Bootstrap CLI (one-liner install, `nextcraft` binary releases on every ship)
Prior: v0.3 Credential Engines (shipped v0.2.8) · v0.2 AI Tutor Architecture (v0.2.0) · v0.1 UI/UX Prototype (v0.1.0)
**Milestone v0.1**UI/UX Prototype (high-fidelity interactive, all mock data)
Initialized via CIAgent v0.7.0
+1 -30
View File
@@ -2,38 +2,9 @@
# Real keys live in .ciagent/.env.secrets (gitignored) and are exported by scripts/dev.sh
AI_PORT=8420
# Network mode (v0.3.5, D-038): dev server binds 0.0.0.0 so remote machines can
# reach the stack. Set to 127.0.0.1 to revert to loopback-only.
AI_HOST=0.0.0.0
# CORS + WS-origin policy: '*' (default) admits any origin — safe because
# credentials are never enabled. Restrict with a comma list, e.g.:
# AI_CORS_ORIGINS=http://nextcraft-1:3000
AI_CORS_ORIGINS=*
AI_PROVIDER=ollama-cloud
AI_MODEL=gemma4:31b
AI_OLLAMA_CLOUD_BASE_URL=https://ollama.com/v1
AI_OLLAMA_CLOUD_API_KEY=
AI_LOCAL_BASE_URL=http://localhost:11434/v1
AI_JSON_MODE=auto
# Sandbox fabric (v0.3)
AI_SANDBOX_DIR=sandboxes
AI_SANDBOX_MAX_CONCURRENT=5
AI_SANDBOX_TIMEOUT_S=900
AI_SANDBOX_MAX_WORKDIR_MB=512
# G-5 abuse control (NOT auth — KYC/identity deferred):
# comma-separated learner allowlist; unknown ids can't create sandboxes (403)
AI_LEARNER_ALLOWLIST=pilot-learner
# max ACTIVE sandboxes per learner → 429 when exceeded
AI_SANDBOX_MAX_PER_LEARNER=1
# global creates per rolling 60s window (in-memory) → 429 when exceeded
AI_SANDBOX_CREATES_PER_MIN=10
# Persistence (SQLite)
AI_DB_PATH=ai_service/data/nextcraft.db
# --- v0.3 Voice (REQ-3-006, D-030) ---
# 'mock' (default; no key needed — tests/dev) or 'browser' (client-native SR/TTS).
# Real server STT/TTS ('openai-audio' + AI_VOICE_BASE_URL/AI_VOICE_API_KEY)
# is deferred to v0.4 per GRILL CUT-1/G-7 — keys never in code or commits.
AI_VOICE_PROVIDER=mock
AI_JSON_MODE=auto
+12 -199
View File
@@ -45,214 +45,27 @@ Tests run with `AI_PROVIDER=mock` (enforced in `tests/conftest.py` by an instanc
## Endpoints
- `GET /health` — status, configured provider, model (no cloud call)
- `POST /v1/chat/stream` — SSE chat stream. Body: `{"agent": "coach"|"tutor", "session_id": "...", "messages": [{"role":"user","content":"..."}]}`. Unknown agents are rejected with 422.
- `POST /v1/chat/stream` — SSE chat stream. Body: `{"agent": "coach", "session_id": "...", "messages": [{"role":"user","content":"..."}]}`
SSE envelope (D-016): `meta` event first (agent/session/model), then `delta` events (incremental content), then `done`; on mid-stream failure an `error` event precedes the terminal `[DONE]` sentinel. sse-starlette emits `: ping` keep-alive comment lines on idle connections — clients must ignore frames without `data:`.
## Manual ollama-cloud persona probe (Phase 3, documented — not automated)
With the real provider, Coach and Tutor must produce distinct on-persona
responses to the same prompt:
## Manual ollama-cloud probe (not automated)
```bash
# start with the cloud provider (keys exported from .ciagent/.env.secrets)
# Streaming probe against ollama-cloud with the real provider:
AI_PROVIDER=ollama-cloud .venv/bin/uvicorn ai_service.main:app --port 8420
curl -sN -X POST localhost:8420/v1/chat/stream \
-H 'Content-Type: application/json' \
-d '{"agent":"tutor","session_id":"probe","messages":[{"role":"user","content":"Explain retrieval practice in one sentence."}]}'
# Coach: expect pacing + one concrete next action + a retrieval-practice question
curl -sN -X POST localhost:8420/v1/chat/stream -H 'Content-Type: application/json' \
-d '{"agent":"coach","session_id":"probe-coach","messages":[{"role":"user","content":"I am stuck on multi-agent communication patterns"}]}' \
| grep '^data:'
# Tutor: expect ONE concept + a worked example + a Socratic check question
curl -sN -X POST localhost:8420/v1/chat/stream -H 'Content-Type: application/json' \
-d '{"agent":"tutor","session_id":"probe-tutor","messages":[{"role":"user","content":"I am stuck on multi-agent communication patterns"}]}' \
| grep '^data:'
# Direct provider probe:
source ../../.ciagent/.env.secrets # never in a committed script
curl -s https://ollama.com/v1/chat/completions \
-H "Authorization: Bearer $OLLAMA_API_KEY" \
-H 'Content-Type: application/json' \
-d '{"model":"gemma4:31b","messages":[{"role":"user","content":"hi"}],"max_tokens":10}'
```
Verify: the two responses have visibly different voice/structure (Coach:
action + accountability; Tutor: concept + example + question). The
automated suite never calls the cloud — distinctness is enforced against
the deterministic mock (distinct system prompts → distinct hash-seeded
outputs).
## Sandbox isolation (v0.3)
The v0.3 code-execution sandbox runs learner/agent code in a Linux **user
namespace** (`unshare --user --map-root-user --mount --pid --fork --net`): the
child is uid 0 *inside* the userns (mapped to the unprivileged host uid), gets
a private mount + PID + network namespace, and uses `RLIMIT_*` for resource
caps. No containers, no sudo — see "Why not containers" below.
### A-101 / D-024 isolation probe transcript
Verbatim output captured on the CI box (Linux, uid 1001 `opencode`, no
docker/podman/bwrap; `iproute2` absent so interface state is read from
kernel sockets + `/proc/net/dev`).
**1. Root-in-userns, uid 0, isolated namespaces:**
```console
$ unshare --user --map-root-user --mount --pid --fork --net id -u
0
```
**2. Fresh netns has exactly one interface: `lo` only (no eth0, no route
out).** Host baseline for contrast:
```console
$ unshare --user --map-root-user --mount --pid --fork --net \
python3 -c "import socket; print(socket.if_nameindex())"
[(1, 'lo')]
$ unshare --user --map-root-user --mount --pid --fork --net \
awk 'NR>2{print $1}' /proc/net/dev
lo:
$ python3 -c "import socket; print(socket.if_nameindex())" # host
[(1, 'lo'), (2, 'eth0')]
```
(Note: `/sys/class/net` shows host interfaces even inside the netns because
`sysfs` here is not netns-aware — the socket-level view above is the
authoritative kernel evidence: 1 interface, loopback only, zero rx bytes, no
carrier to any external link.)
**3. Write containment — writes inside the sandbox workdir are visible on the
host under the sandbox dir, owned by the real (unprivileged) host uid:**
```console
$ unshare --user --map-root-user --mount --pid --fork --net bash -c "
mkdir -p /tmp/demo-work && cd /tmp/demo-work
echo 'hello-from-inside-sandbox (uid=0 in-ns)' > contained.txt
id -u"
0
$ cat /tmp/demo-work/contained.txt # host
hello-from-inside-sandbox (uid=0 in-ns)
$ ls -la /tmp/demo-work/contained.txt # host
-rw-r--r-- 1 opencode opencode 40 ... /tmp/demo-work/contained.txt
```
The in-userns "root" writes land on the host filesystem as uid 1001
(`opencode`) — the uid-mapping is doing the confinement; nothing escapes the
sandbox workdir as any other identity.
**4. `/proc` remount is NOT permitted in this context — probe + exact error:**
```console
$ unshare --user --map-root-user --mount --pid --fork --net \
bash -c "mount -t proc proc /proc"
mount: /proc: permission denied.
dmesg(1) may have more information after failed mount system call.
(exit 32)
```
`mount -t proc` fails even with in-ns "root" because `/proc` is owned by a
userns that does not contain our uid mapping (the box's `/` is itself
owned by `nobody:nogroup` — we're already inside a container). **This is
acceptable for v0.3**: the sandbox does not depend on a custom `/proc` view;
the child sees the host `/proc` read-only-ish view which is already filtered
by the pid namespace (only in-ns pids are visible). The pidns itself is what
provides process isolation, not the proc remount.
### Locked resource-limit mechanism (G-1 / G-2)
Resource enforcement is settled for v0.3 — this is the locked decision:
| Resource | Mechanism | Notes |
|----------|-----------|-------|
| **Memory** | `RLIMIT_AS` (address space) | setrlimit in the child pre-exec; deterministic, no cgroup needed |
| **CPU** | `RLIMIT_CPU` | kernel SIGKILL at the cpu-seconds ceiling |
| **Single-file size** | `RLIMIT_FSIZE` | catches runaway single-file writes |
| **Wall clock** | **manager reaper kill** (parent watchdog) | RLIMIT_CPU doesn't cover sleeping/idle children; the manager kills the sandbox on wall-clock timeout |
| **Per-sandbox process count** | `RLIMIT_NPROC` | ⚠️ **SHARED at the host uid, not per-sandbox** — the counter is per-real-uid across all of that uid's process trees, so two concurrent sandboxes share the same NPROC budget. Accepted v0.3 gap: without cgroup delegation there's no per-sandbox pid cap; mitigations are (a) the manager serializes sandbox runs and (b) NPROC is still a hard fork-bomb ceiling. |
| **Hard disk quota** | **NOT kernel-enforceable** | ⚠️ without cgroup delegation or sudo (`quotactl`, project quotas) there is no kernel-enforced per-sandbox disk cap. Accepted v0.3 gap. **Mitigation: a manager-side workdir-size sweep** — after each run (and on a periodic reaper pass) the manager walks the sandbox workdir and enforces `AI_SANDBOX_MAX_WORKDIR_MB` (**default 512 MB**); oversized dirs are reaped. Combined with `RLIMIT_FSIZE` this bounds disk growth between sweeps. |
Both accepted gaps (shared NPROC, no kernel disk quota) are documented here as
v0.3 scope boundaries; closing them requires cgroup v2 delegation or sudo,
neither of which is available in the target environment.
### Why not containers
Container runtimes / privileged wrapper tools are probed-and-absent on the
box, and we have no `sudo`:
```console
$ for cmd in docker podman bwrap firejail; do
printf '%-8s: ' "$cmd"; command -v "$cmd" || echo MISSING
done; printf '%-8s: ' sudo; command -v sudo || echo MISSING
docker : MISSING
podman : MISSING
bwrap : MISSING
firejail: MISSING
sudo : MISSING
$ id -u
1001
```
Unprivileged user namespaces are on the box's kernel and need neither a
daemon, nor suid helpers, nor network access — they are the only isolation
primitive that works here, so that's what v0.3 uses.
## Telemetry delivery semantics (v0.3, REQ-3-003)
Delivery is **at-least-once**; storage is **exactly-once** — the two compose:
- The in-sandbox capture agent (stdlib-only, `scripts/sandbox-agent.py`)
spools every event to a durable JSONL file (fsync per append) BEFORE any
send attempt, so no event can be lost to a dead socket or a SIGKILL.
- The WS ingest endpoint (`WS /v1/telemetry/ingest?learner_id&task_id`,
D-026) dedups server-side on the `(learner_id, task_id, seq)` primary key:
re-sends (reconnect flushes, replay margin) are collapsed, never upserted.
- On disconnect the agent reconnects with exponential backoff and flushes
the spool in `seq` order; a transient outage therefore loses nothing and
stores each event exactly once (`tests/telemetry/test_durability.py`
proves this end-to-end against a real namespace sandbox + live server).
- Replay/read path: `GET /v1/telemetry/traces/{learner}/{task}` returns the
complete ordered trace; `GET /v1/telemetry/gaps/{learner}/{task}` returns
missing seqs for gap detection.
- Flood boundary (G-3): a connection exceeding `AI_TELEMETRY_MAX_EVENTS_PER_TASK`
(default 50,000) is closed with WS code 1008 and its trace is marked
`INCOMPLETE_FLOODED` — a terminal integrity flag the grader refuses to
grade. Silent event dropping is forbidden: it would corrupt grading input.
## Voice defense (v0.3, REQ-3-006)
Voice is **mock-first** (D-030): the defense pipeline is fully proven over
the deterministic `MockVoiceProvider` + browser-native fallback — no task
requires a real voice key. Real server STT/TTS (`OpenAIAudioProvider` over
OpenAI-compatible `/audio/transcriptions` + `/audio/speech`) is **deferred
to v0.4** together with KYC (GRILL CUT-1 / G-7): it could never be exercised
in CI, so v0.3 ships the protocol seam instead of an unverifiable claim.
- `AI_VOICE_PROVIDER=mock` (default) — deterministic canned STT/TTS
- `AI_VOICE_PROVIDER=browser` — the web client uses SpeechRecognition +
speechSynthesis; the server keeps text-turn persistence
- Conversational budget: a defense turn should complete in **< 4s**
(`DEFENSE_TURN_BUDGET_MS` in `tests/voice/test_latency.py`). v0.3
asserts instrumentation (stt_ms/llm_ms/tts_ms populated per turn); the
wall-clock acceptance probe against a real voice endpoint is a v0.4
criterion, run manually with `AI_VOICE_PROVIDER` set to the real
provider and keys in `.ciagent/.env.secrets` (never in code/commits).
## End-to-end credential flow (v0.3, REQ-3-007/008)
`tests/api/test_e2e_credential_flow.py` runs the full pipeline against a REAL
uvicorn server with REAL namespace sandboxes (mock LLM/voice per G-2
precedent): variant -> telemetry-wired sandbox -> in-sandbox exec -> trace
persistence -> process-trace grade (variant seed stamped) -> assessor
coaching -> oral defense -> verdict + integrity signals -> proctor. It
asserts no corpus fixture appears anywhere in the learner path.
Manual browser pass (documented, not automated): `pnpm ai:dev` + `pnpm dev`,
then open `/build/stack-orchestration-c007` — variant statement + starter
files load, edit a file, Run/Test execute in the sandbox with output in the
read-only panel, the telemetry status pulses, Lab streams feedback from the
live digest; then `/defend/stack-orchestration-c007` — Start Defense, typed
answers (mic path needs permission), Finish, Grade My Work renders the real
rubric bars. Navigating away destroys the sandbox
(`curl localhost:8420/v1/sandboxes` shows the count drop).
## Layout
```
@@ -1,63 +0,0 @@
"""AssessorAgent — rubric coaching over REAL grading output (REQ-3-007).
v0.3 re-grounding: the Assessor no longer invents scores from corpus
artifacts — the process-trace grading engine (Phase 3) computes and
persists the validated RubricScore. This agent now renders the STORED
grade as rubric-anchored coaching: explains the criteria, cites strengths
and gaps, and frames next steps. Corpus artifacts are retired from this
path (corpus dormancy, Task 6-1-04).
"""
from pydantic import BaseModel, Field
from ..corpus.learner_context import LearnerContext, get_learner_context
from ..grading.store import GradeRecord
from ..prompts.assessor import SYSTEM_PROMPT, render_context
from .base import BaseAgent
class GradeCoaching(BaseModel):
"""Rubric-anchored coaching rendered FROM the stored grade (not invented)."""
summary: str = Field(min_length=1)
strengths: list[str] = Field(min_length=1, max_length=3)
gaps: list[str] = Field(min_length=1, max_length=3)
next_steps: list[str] = Field(min_length=1, max_length=3)
GRADE_COACHING_SCHEMA_HINT = (
'{"summary": "<two sentences on the grade>", '
'"strengths": ["<one sentence>"], "gaps": ["<one sentence>"], '
'"next_steps": ["<one sentence>"]}'
)
class AssessorAgent(BaseAgent):
name = "assessor"
def system_prompt(self, learner_context: LearnerContext | None = None) -> str:
ctx = learner_context or get_learner_context()
return SYSTEM_PROMPT.format_map(render_context(ctx))
async def coach_grade(
self,
grade: GradeRecord,
learner_context: LearnerContext | None = None,
) -> GradeCoaching:
"""Render the STORED grade as coaching via the D-020 defense."""
grade_json = {
"verdict": grade.verdict,
"scores": grade.scores,
"digest": grade.digest,
}
coaching: GradeCoaching = await self.structured_reply(
history=None,
user_input=(
"The learner's process-trace grade (computed by the grading "
f"engine) is:\n{grade_json!r}\nExplain it as coaching."
),
learner_context=learner_context,
schema=GradeCoaching,
schema_hint=GRADE_COACHING_SCHEMA_HINT,
)
return coaching
@@ -1,13 +0,0 @@
"""CoachAgent — pacing, motivation, retrieval practice (REQ-2-005)."""
from ..corpus.learner_context import LearnerContext, get_learner_context
from ..prompts.coach import SYSTEM_PROMPT, render_context
from .base import BaseAgent
class CoachAgent(BaseAgent):
name = "coach"
def system_prompt(self, learner_context: LearnerContext | None = None) -> str:
ctx = learner_context or get_learner_context()
return SYSTEM_PROMPT.format_map(render_context(ctx))
@@ -1,104 +0,0 @@
"""ExaminerAgent — the seventh agent: oral-defense examiner (REQ-3-006, A-109).
BOUNDARY DECISION (PERSONAS conflict rule, honored by construction): the
examiner is a TEXT agent. It composes the LLM provider through BaseAgent and
consumes defense transcript turns; it NEVER imports voice/ — STT/TTS belong
to the API endpoints (they move audio bytes; the agent moves question text).
Integrity signals (long pauses, off-scope cadence) are computed by the
endpoint layer from turn metadata (latency_ms etc.), not by the agent.
Digest discipline (D-028 mirror): questions are grounded in the compact
TraceDigest + variant statement — never the raw trace, never learner ids.
"""
from __future__ import annotations
from pydantic import BaseModel, ConfigDict, Field
from ..grading.features import TraceDigest
from ..llm.types import Message
from ..prompts.examiner import SYSTEM_PROMPT, VERDICT_SCHEMA_HINT, render_digest_context
from .base import BaseAgent
class DefenseVerdict(BaseModel):
"""D-20-validated final defense verdict (structured mode)."""
model_config = ConfigDict(extra="forbid")
verdict: str = Field(pattern="^(mastered|developing|not_yet)$")
understanding: str = Field(min_length=1)
process_justification: str = Field(min_length=1)
communication: str = Field(min_length=1)
strengths: list[str] = Field(min_length=1, max_length=2)
gaps: list[str] = Field(min_length=1, max_length=2)
class ExaminerAgent(BaseAgent):
"""Conducts the oral defense: next_question + final_verdict."""
name = "examiner"
def system_prompt(self, learner_context=None) -> str: # noqa: ANN001
"""Examiner is context-free (digest-anonymous, D-028 mirror)."""
return SYSTEM_PROMPT
def build_defense_messages(
self,
trace_digest: TraceDigest | None = None,
variant_statement: str | None = None,
history: list[Message] | None = None,
) -> list[Message]:
"""System + grounding + defense transcript (no learner id — D-028)."""
digest_json = (
trace_digest.model_dump_json() if trace_digest is not None else "{}"
)
messages: list[Message] = [
Message(role="system", content=SYSTEM_PROMPT),
Message(role="user", content=render_digest_context(digest_json, variant_statement)),
Message(
role="assistant",
content="Understood. I will question the learner about this build session.",
),
]
for m in history or []:
messages.append(m)
return messages
async def next_question(
self,
history: list[Message],
trace_digest: TraceDigest | None = None,
variant_statement: str | None = None,
) -> str:
"""One examiner question (streamed over SSE by the endpoints)."""
messages = self.build_defense_messages(trace_digest, variant_statement, history)
messages.append(
Message(role="user", content="Ask the learner your next question now.")
)
reply = await self.provider.chat(messages, model=self.settings.model)
return reply
async def final_verdict(
self,
history: list[Message],
trace_digest: TraceDigest | None = None,
variant_statement: str | None = None,
) -> DefenseVerdict:
"""Structured verdict via the D-020 4-layer defense."""
from .structured import structured_completion # module-direct (G-4)
messages = self.build_defense_messages(trace_digest, variant_statement, history)
messages.append(
Message(
role="user",
content="The defense is finished. Return the final verdict JSON now.",
)
)
return await structured_completion(
self.provider,
messages,
model=self.settings.model,
schema=DefenseVerdict,
schema_hint=VERDICT_SCHEMA_HINT,
)
-39
View File
@@ -1,39 +0,0 @@
"""LabAgent — in-flow feedback over LIVE sandbox telemetry (REQ-3-007).
v0.3 re-grounding: consumes a TraceDigest computed from the learner's real
trace (grading/features.compute_digest over TraceStore events) — the v0.2
corpus scenarios are retired from this path (corpus dormancy, Task 6-1-04).
No session chat — each request is one live-trace read.
"""
from collections.abc import AsyncIterator
from ..config import Settings
from ..corpus.learner_context import LearnerContext, get_learner_context
from ..grading.features import TraceDigest
from ..llm.base import LLMProvider
from ..prompts.lab import SYSTEM_PROMPT, render_context, render_digest_timeline
from .base import BaseAgent
class LabAgent(BaseAgent):
name = "lab"
def __init__(self, provider: LLMProvider, settings: Settings) -> None:
super().__init__(provider, settings)
def system_prompt(self, learner_context: LearnerContext | None = None) -> str:
ctx = learner_context or get_learner_context()
return SYSTEM_PROMPT.format_map(render_context(ctx))
async def stream_feedback(
self,
digest: TraceDigest | None,
learner_context: LearnerContext | None = None,
) -> AsyncIterator[str]:
"""Feedback grounded in the learner's live trace digest."""
timeline = render_digest_timeline(digest)
async for token in self.stream_reply(
history=None, user_input=timeline, learner_context=learner_context
):
yield token
@@ -1,17 +0,0 @@
"""MentorAgent — long-horizon career narrative (REQ-2-010).
Streaming, session-backed conversational agent: the learner can ask
follow-up questions about their trajectory and the Mentor keeps context.
"""
from ..corpus.learner_context import LearnerContext, get_learner_context
from ..prompts.mentor import SYSTEM_PROMPT, render_context
from .base import BaseAgent
class MentorAgent(BaseAgent):
name = "mentor"
def system_prompt(self, learner_context: LearnerContext | None = None) -> str:
ctx = learner_context or get_learner_context()
return SYSTEM_PROMPT.format_map(render_context(ctx))
@@ -1,79 +0,0 @@
"""ProctorAgent — integrity signals + coaching over REAL inputs (REQ-3-007).
v0.3 re-grounding: consumes the learner's live trace digest (idle gaps,
command cadence), the DefenseStore integrity signals (long pauses from the
oral defense), and the variant seed cross-check — NOT v0.2 corpus
scenarios. The proctor COACHES: it classifies signals supportively and
recommends one intervention; it never punishes and never accuses.
Integrity inputs (computed server-side, passed in by the API layer):
- trace digest: idle_gap_count/total, command_categories histogram,
error/fix cycles, huge-burst indicators (edit_count vs test runs)
- defense signals: long_pauses list from the finished defense (A-109)
- variant: seed + params when the task is variant-derived (off-template
work is a cross-check input, not an accusation)
"""
from pydantic import BaseModel, Field
from ..corpus.learner_context import LearnerContext, get_learner_context
from ..grading.features import TraceDigest
from ..prompts.proctor import SYSTEM_PROMPT, render_context
from .base import BaseAgent
class IntegritySignal(BaseModel):
signal_type: str # "idle_gap" | "long_pause" | "burst_edit" | "off_template"
severity: str # "low" | "medium" | "high"
note: str
class ProctorAssessment(BaseModel):
signals: list[IntegritySignal] = Field(min_length=0)
intervention: str # ONE supportive coaching recommendation
summary: str
PROCTOR_ASSESSMENT_SCHEMA_HINT = (
'{"signals": [{"signal_type": "<type>", '
'"severity": "low"|"medium"|"high", "note": "<one sentence>"}], '
'"intervention": "<one supportive recommendation>", '
'"summary": "<one sentence>"}'
)
class ProctorAgent(BaseAgent):
name = "proctor"
def system_prompt(self, learner_context: LearnerContext | None = None) -> str:
ctx = learner_context or get_learner_context()
return SYSTEM_PROMPT.format_map(render_context(ctx))
async def assess(
self,
digest: TraceDigest | None,
defense_signals: dict | None = None,
variant_context: dict | None = None,
learner_context: LearnerContext | None = None,
) -> ProctorAssessment:
"""Classify REAL integrity inputs into supportive signals + coaching."""
parts: list[str] = []
if digest is not None:
parts.append(f"Build-session digest:\n{digest.model_dump_json()}")
else:
parts.append("No build telemetry recorded for this task yet.")
if defense_signals:
parts.append(f"Oral-defense integrity signals:\n{defense_signals}")
if variant_context:
parts.append(f"Variant audit context (seed + params):\n{variant_context}")
assessment: ProctorAssessment = await self.structured_reply(
history=None,
user_input=(
"Assess this learner's integrity signals supportively.\n\n"
+ "\n\n".join(parts)
),
learner_context=learner_context,
schema=ProctorAssessment,
schema_hint=PROCTOR_ASSESSMENT_SCHEMA_HINT,
)
return assessment
@@ -13,37 +13,6 @@ from .base import BaseAgent
AgentFactory = Callable[[LLMProvider, Settings], BaseAgent]
def register_builtin_agents(registry: "AgentRegistry") -> None:
"""Central registration of all seven shipped agents (G-4: one pattern).
coach, tutor, lab, assessor, proctor, mentor, examiner (Phase 5).
New agents register here in their landing phase.
"""
from .assessor import AssessorAgent
from .coach import CoachAgent
from .examiner import ExaminerAgent
from .lab import LabAgent
from .mentor import MentorAgent
from .proctor import ProctorAgent
from .tutor import TutorAgent
registry.register("coach", lambda provider, settings: CoachAgent(provider, settings))
registry.register("tutor", lambda provider, settings: TutorAgent(provider, settings))
registry.register("lab", lambda provider, settings: LabAgent(provider, settings))
registry.register(
"assessor", lambda provider, settings: AssessorAgent(provider, settings)
)
registry.register(
"proctor", lambda provider, settings: ProctorAgent(provider, settings)
)
registry.register(
"mentor", lambda provider, settings: MentorAgent(provider, settings)
)
registry.register(
"examiner", lambda provider, settings: ExaminerAgent(provider, settings)
)
class UnknownAgentError(KeyError):
"""Raised when resolving an agent name that was never registered."""
@@ -74,10 +74,6 @@ class InMemorySessionStore:
if session is None:
raise KeyError(f"unknown session {session_id!r}")
session.messages.append(message)
# Bound stored history too (window bounds replay, not storage):
# keep at most 2x window so retries/recent context survive.
if len(session.messages) > self._window * 2:
del session.messages[: len(session.messages) - self._window * 2]
self._sessions.move_to_end(session_id)
async def history_window(
@@ -1,13 +0,0 @@
"""TutorAgent — concept delivery, Socratic questioning (REQ-2-006)."""
from ..corpus.learner_context import LearnerContext, get_learner_context
from ..prompts.tutor import SYSTEM_PROMPT, render_context
from .base import BaseAgent
class TutorAgent(BaseAgent):
name = "tutor"
def system_prompt(self, learner_context: LearnerContext | None = None) -> str:
ctx = learner_context or get_learner_context()
return SYSTEM_PROMPT.format_map(render_context(ctx))
+1 -19
View File
@@ -3,24 +3,6 @@
Boundary rule: api/ composes agents/ and llm/; they never import api/.
"""
from .assessment import router as assessment_router
from .chat import router as chat_router
from .defense import router as defense_router
from .lab import router as lab_router
from .mentor import router as mentor_router
from .proctor import router as proctor_router
from .sandboxes import router as sandboxes_router
from .telemetry import router as telemetry_router
from .variants import router as variants_router
__all__ = [
"assessment_router",
"chat_router",
"lab_router",
"mentor_router",
"proctor_router",
"sandboxes_router",
"telemetry_router",
"variants_router",
"defense_router",
]
__all__ = ["chat_router"]
@@ -1,190 +0,0 @@
"""/v1/assessment — rubric evaluation + trace grading endpoints (REQ-2-008, REQ-3-004).
Two endpoint families share this router:
POST /v1/assessment/evaluate (v0.2, REQ-2-008) — corpus
artifact evaluation through
the Assessor agent.
POST /v1/assessment/grade (v0.3, REQ-3-004) — grade a
REAL process trace through
the GradingEngine.
GET /v1/assessment/grade/{learner_id}/{task_id} — stored latest grade.
Grading status-code mapping (the engine's outcomes are CONTRACT, not errors):
GradeRecord(verdict=GRADED) → 200 — rubric scores +
verdict (in scores.verdict)
+ digest summary.
GradeRecord(UNGRADABLE_TRACE_INCOMPLETE) → 200 — the ungradable
record IS a valid result:
the trace cannot be graded,
and the gate surfaces WHY
(scores.missing_seqs +
scores.integrity_flag).
Persisted like any grade.
GradeRecord(UNGRADABLE_EMPTY_TRACE) → 200 — no events stored for
the pair. This covers BOTH
a known pair whose trace
ended up empty AND a task
that never had a trace at
all: the engine cannot
distinguish them (zero
stored events is zero
events), and grading an
absent trace genuinely has
the empty-trace outcome —
a 404 here would erase the
durable gate record the
engine persists for the
pair. PLAN's 404 applies to
GET of a never-graded pair.
StructuredOutputError → 502 — the trace was
gradable but the provider
failed the D-020 budget;
provider failure (bad
gateway to the model), same
mapping as evaluate.
GET of an unknown (never-graded) pair → 404.
DI (D-027/D-032 house pattern): engine + store arrive via deps.get_grading_engine
/ get_grade_store from app.state; this module owns all FastAPI wiring — the
engine knows nothing of HTTP. UNGRADABLE_* bodies are rendered by the same
GradeResponse model as GRADED ones (a gate record's `scores` holds the gate
detail instead of rubric scores), so consumers read ONE shape.
"""
from datetime import datetime
from typing import Any
from fastapi import APIRouter, Depends, HTTPException
from pydantic import BaseModel, Field
from ..agents.registry import AgentRegistry
from ..agents.structured import StructuredOutputError
from ..config import Settings
from ..corpus.learner_context import get_learner_context
from ..grading.engine import GradingEngine
from ..grading.store import GradeRecord, GradeStore
from .deps import (
get_agent_registry,
get_grade_store,
get_grading_engine,
get_provider,
get_settings,
)
router = APIRouter(prefix="/v1")
# --- v0.2 artifact evaluation (REQ-2-008) --------------------------------------
class EvaluateRequest(BaseModel):
learner_id: str = Field(min_length=1)
task_id: str = Field(min_length=1)
@router.post("/assessment/evaluate")
async def assessment_evaluate(
body: EvaluateRequest,
registry: AgentRegistry = Depends(get_agent_registry),
settings: Settings = Depends(get_settings),
provider=Depends(get_provider),
grade_store=Depends(get_grade_store),
) -> dict:
"""Assessor coaching rendered FROM the learner's stored grade (REQ-3-007).
The grading engine computes the scores (POST /assessment/grade); this
endpoint explains them. No stored grade yet -> 404 (grade first).
"""
grade = grade_store.get(body.learner_id, body.task_id)
if grade is None:
raise HTTPException(
status_code=404,
detail=f"no stored grade for {body.learner_id}/{body.task_id} - grade first",
)
agent = registry.get(provider, settings, "assessor")
learner_context = get_learner_context(body.learner_id)
try:
coaching = await agent.coach_grade(grade, learner_context)
except Exception as exc:
raise HTTPException(
status_code=502,
detail=f"assessment evaluation failed: {exc}",
) from exc
return {
"learner_id": grade.learner_id,
"task_id": grade.task_id,
"grade_verdict": grade.verdict,
"grade_scores": grade.scores,
"coaching": coaching.model_dump(),
}
# --- # --- v0.3 trace grading (REQ-3-004) ---------------------------------------------
class GradeRequest(BaseModel):
learner_id: str = Field(min_length=1)
task_id: str = Field(min_length=1)
class GradeResponse(BaseModel):
"""GradeRecord over HTTP — one shape for GRADED and UNGRADABLE_* alike.
`scores` holds the validated rubric (criteria 0-4, strengths, gaps,
rubric verdict) for a GRADED record, or the gate detail
({integrity_flag, missing_seqs}) for an UNGRADABLE_* record — never both.
`digest` is the compact trace summary that fed the rubric prompt (empty
for gate records: nothing was graded).
"""
learner_id: str
task_id: str
variant_seed: str | None
digest: dict[str, Any]
scores: dict[str, Any]
verdict: str
model: str
created_at: datetime
def _grade_response(record: GradeRecord) -> GradeResponse:
return GradeResponse.model_validate(record, from_attributes=True)
@router.post("/assessment/grade", response_model=GradeResponse)
async def assessment_grade(
body: GradeRequest,
engine: GradingEngine = Depends(get_grading_engine),
) -> GradeResponse:
"""Run the grading engine for one (learner_id, task_id) trace.
Gate outcomes (UNGRADABLE_*) are 200s — they are first-class results the
engine persists, not failures. Only a provider that exhausts the D-020
budget turns into a 502; nothing is persisted on that path.
"""
try:
record = await engine.grade(body.learner_id, body.task_id)
except StructuredOutputError as exc:
raise HTTPException(
status_code=502,
detail=f"grading failed: {exc}",
) from exc
return _grade_response(record)
@router.get("/assessment/grade/{learner_id}/{task_id}", response_model=GradeResponse)
async def assessment_get_grade(
learner_id: str,
task_id: str,
store: GradeStore = Depends(get_grade_store),
) -> GradeResponse:
"""Latest stored grade for the pair; 404 when none was ever stored."""
record = store.get(learner_id, task_id)
if record is None:
raise HTTPException(
status_code=404,
detail=f"no stored grade for {learner_id!r}/{task_id!r}",
)
return _grade_response(record)
+16 -48
View File
@@ -1,12 +1,8 @@
"""POST /v1/chat/stream — SSE chat with the D-016 envelope + agent routing.
"""POST /v1/chat/stream — SSE chat with the D-016 envelope.
Envelope: meta event first (flushed before first token), then raw content
deltas, then done; error event before [DONE] on mid-stream failure.
Pre-first-byte provider failures surface as in-band `provider_unavailable`
error events (SSE 200 headers are already committed once meta flushes).
Agent routing (A-007): the request names its agent; unknown agents are
rejected with 422. No autonomous routing in v0.2.
Pre-first-byte failures become proper HTTP error statuses.
"""
import json
@@ -16,13 +12,11 @@ from fastapi import APIRouter, Depends, HTTPException
from pydantic import BaseModel, Field
from sse_starlette.sse import EventSourceResponse
from ..agents.registry import AgentRegistry, UnknownAgentError
from ..agents.session import SessionStore
from ..config import Settings
from ..corpus.learner_context import get_learner_context
from ..llm.base import LLMProvider
from ..llm.types import Message
from .deps import get_agent_registry, get_provider, get_session_store, get_settings
from .deps import get_provider, get_session_store, get_settings
router = APIRouter(prefix="/v1")
@@ -30,50 +24,26 @@ router = APIRouter(prefix="/v1")
class ChatStreamRequest(BaseModel):
agent: str = Field(min_length=1)
session_id: str = Field(min_length=1)
learner_id: str | None = None
messages: list[Message] = Field(min_length=1)
@router.post("/chat/stream")
async def chat_stream(
body: ChatStreamRequest,
registry: AgentRegistry = Depends(get_agent_registry),
sessions: SessionStore = Depends(get_session_store),
settings: Settings = Depends(get_settings),
provider: LLMProvider = Depends(get_provider),
settings: Settings = Depends(get_settings),
sessions: SessionStore = Depends(get_session_store),
) -> EventSourceResponse:
# Route to the named agent (A-007); unknown → 422 before any streaming.
try:
agent = registry.get(provider, settings, body.agent)
except UnknownAgentError as exc:
raise HTTPException(
status_code=422, detail=str(exc)
) from None
learner_context = get_learner_context(body.learner_id)
if not body.messages:
raise HTTPException(status_code=422, detail="messages must not be empty")
# Agent-scoped session (A-007/G-4): persisted turn history, windowed replay.
session = await sessions.get(body.session_id)
if session is None:
session = await sessions.create(
body.session_id, agent=body.agent, learner_id=body.learner_id or "learner-001"
)
# The new user turn is the last message of the request.
user_turn = body.messages[-1]
session = await sessions.create(body.session_id, agent=body.agent)
history = await sessions.history_window(body.session_id)
# Retry dedupe (P1 from final review): a client retry resends the same
# turn after a provider failure — don't double-append it to history.
last_stored = history[-1] if history else None
is_retry = (
last_stored is not None
and last_stored.role == "user"
and last_stored.content == user_turn.content
)
if not is_retry:
await sessions.append(body.session_id, user_turn)
else:
# On retry the history replay should exclude the stored duplicate.
history = history[:-1]
# Persist this turn's user message before streaming.
await sessions.append(body.session_id, body.messages[-1])
async def event_stream() -> AsyncIterator[dict]:
yield {"event": "message", "data": json.dumps({
@@ -85,10 +55,9 @@ async def chat_stream(
first_byte = True
reply_parts: list[str] = []
try:
async for token in agent.stream_reply(
history=history,
user_input=user_turn.content,
learner_context=learner_context,
async for token in provider.stream_chat(
body.messages if not history else history + body.messages,
model=settings.model,
):
first_byte = False
reply_parts.append(token)
@@ -103,10 +72,11 @@ async def chat_stream(
yield {"event": "message", "data": json.dumps({
"type": "done", "finish_reason": "stop"
})}
yield {"event": "message", "data": "[DONE]"}
except Exception as exc: # CancelledError is BaseException — passes through
message = str(exc)
if first_byte:
# Pre-first-byte failure: we already flushed meta + 200 headers;
# surface as in-band error (status change is impossible post-flush).
yield {"event": "message", "data": json.dumps({
"type": "error", "code": "provider_unavailable", "message": message
})}
@@ -114,9 +84,7 @@ async def chat_stream(
yield {"event": "message", "data": json.dumps({
"type": "error", "code": "provider_error", "message": message
})}
# [DONE] is yielded from the except branch, NEVER from finally:
# a yield inside finally would re-raise after GeneratorExit when the
# client disconnects ("async generator ignored GeneratorExit").
finally:
yield {"event": "message", "data": "[DONE]"}
return EventSourceResponse(
-332
View File
@@ -1,332 +0,0 @@
"""Oral-defense endpoints (Task 5-3-01, REQ-3-006, A-109).
Full defense loop over HTTP with mock-first voice (D-030) and the seventh
Examiner agent (SSE question streaming happens through the chat pipeline;
these endpoints are the session orchestration + transcript persistence):
POST /v1/defense/start {learner_id, task_id}
POST /v1/defense/{id}/answer {text} | multipart audio (STT)
GET /v1/defense/{id}/audio/{turn_id} TTS bytes (streaming)
POST /v1/defense/{id}/finish verdict + integrity signals
GET /v1/defense/{id} transcript + signals
Integrity signals (A-109) are computed server-side from turn metadata:
long pauses = learner turns whose latency_ms exceeds PAUSE_THRESHOLD_MS.
The defense does NOT gate on trace completeness (the grader does, G-4);
an incomplete trace is surfaced as `trace_complete: false` so the UI can
disclose it before the learner defends.
"""
from __future__ import annotations
import time
from datetime import UTC, datetime
from fastapi import APIRouter, Depends, File, Form, HTTPException, UploadFile
from fastapi.responses import StreamingResponse
from pydantic import BaseModel, Field
from ..agents.examiner import ExaminerAgent
from ..grading.features import TraceDigest, compute_digest
from ..llm.types import Message
from ..voice.base import VoiceDescriptor
from ..voice.browser import BROWSER_FALLBACK_DESCRIPTOR
from ..voice.defense_store import DefenseRecord, DefenseStore, DefenseTurn
from .deps import (
get_examiner,
get_settings,
get_trace_store,
get_variant_store,
get_voice_provider,
get_voice_store,
)
router = APIRouter(prefix="/v1/defense", tags=["defense"])
#: A-109: learner turns slower than this are flagged as long pauses (ms).
PAUSE_THRESHOLD_MS = 15_000
_ROLE_EXAMINER = "examiner"
_ROLE_LEARNER = "learner"
class StartRequest(BaseModel):
learner_id: str = Field(min_length=1)
task_id: str = Field(min_length=1)
class StartResponse(BaseModel):
defense_id: str
voice_descriptor: dict
trace_complete: bool
first_question: str
class AnswerResponse(BaseModel):
question: str
turn_latency: dict[str, int | None]
class FinishResponse(BaseModel):
verdict: dict
integrity_signals: dict
async def _digest_for_task(
trace_store, learner_id: str, task_id: str
) -> tuple[TraceDigest | None, bool]:
"""Digest of the learner's trace for this task + completeness flag."""
if not trace_store.list_tasks(learner_id) or task_id not in trace_store.list_tasks(
learner_id
):
return None, True # no trace at all is "complete" for defense purposes
trace = trace_store.get_trace(learner_id, task_id)
gaps = trace_store.gaps(learner_id, task_id)
return (compute_digest(trace) if trace else None), (len(gaps) == 0)
def _voice_descriptor(settings) -> VoiceDescriptor:
"""The capability descriptor for the configured voice mode (D-030).
Must-Have #6: browser mode returns BROWSER_FALLBACK_DESCRIPTOR so the
web client selects native SpeechRecognition/speechSynthesis; mock mode
returns the mock descriptor. (A v0.4 server provider would return
mode="server" the protocol seam.)
"""
if (settings.voice_provider or "mock").strip().lower() == "browser":
return BROWSER_FALLBACK_DESCRIPTOR
return VoiceDescriptor(
mode="mock", sr_available=True, tts_available=True, hint=""
)
@router.post("/start", response_model=StartResponse)
async def start_defense(
body: StartRequest,
examiner: ExaminerAgent = Depends(get_examiner),
voice_store: DefenseStore = Depends(get_voice_store),
voice_provider=Depends(get_voice_provider),
trace_store=Depends(get_trace_store),
variant_store=Depends(get_variant_store),
settings=Depends(get_settings),
) -> StartResponse:
record = voice_store.start(
DefenseRecord(
id=f"dfn-{int(time.time() * 1000):x}-{body.learner_id[:8]}",
learner_id=body.learner_id,
task_id=body.task_id,
status="in_progress",
created_at=datetime.now(UTC),
)
)
digest, trace_complete = await _digest_for_task(trace_store, body.learner_id, body.task_id)
variant = variant_store.get_by_task(body.task_id)
statement = variant.statement if variant is not None else None
started = time.perf_counter()
question = await examiner.next_question(
history=[], trace_digest=digest, variant_statement=statement
)
llm_ms = int((time.perf_counter() - started) * 1000)
voice_store.append_turn(
record.id,
DefenseTurn(
defense_id=record.id,
seq=0,
role=_ROLE_EXAMINER,
text=question,
ts=datetime.now(UTC),
latency_ms=llm_ms,
created_at=datetime.now(UTC),
),
)
descriptor = getattr(voice_provider, "descriptor", None) or _voice_descriptor(
settings
)
return StartResponse(
defense_id=record.id,
voice_descriptor=descriptor.model_dump(),
trace_complete=trace_complete,
first_question=question,
)
@router.post("/{defense_id}/answer", response_model=AnswerResponse)
async def answer_defense(
defense_id: str,
text: str | None = Form(default=None),
audio: UploadFile | None = File(default=None),
voice_store: DefenseStore = Depends(get_voice_store),
voice_provider=Depends(get_voice_provider),
examiner: ExaminerAgent = Depends(get_examiner),
trace_store=Depends(get_trace_store),
variant_store=Depends(get_variant_store),
) -> AnswerResponse:
record = voice_store.get(defense_id)
if record is None:
raise HTTPException(status_code=404, detail=f"no defense {defense_id!r}")
if record.status == "finished":
# The store owns the finished transition but does NOT police turn
# sequencing (defense_store.py: "turns after finalize are a sequencing
# bug for the endpoints to prevent") — this is the endpoint half of
# that contract: a sealed transcript is append-only-no-more.
raise HTTPException(
status_code=409, detail="defense is finished; start a new defense"
)
if text is None and audio is None:
raise HTTPException(status_code=422, detail="provide {text} or audio")
# STT (typed fallback bypasses the voice provider entirely).
stt_ms: int | None = None
if audio is not None:
stt_started = time.perf_counter()
raw = await audio.read()
if not raw:
# Empty upload is a client error (422), not a provider crash
# (500): validate before the provider call so every provider —
# mock today, the v0.4 real one — sees the same contract.
raise HTTPException(status_code=422, detail="audio upload is empty")
fmt = (audio.content_type or "audio/wav").split("/")[-1]
segment = await voice_provider.transcribe(raw, fmt)
stt_ms = int((time.perf_counter() - stt_started) * 1000)
text = segment.text
turns = record.turns if hasattr(record, "turns") else []
history = [
Message(role="assistant" if t.role == _ROLE_EXAMINER else "user", content=t.text)
for t in turns
]
next_seq = len(turns)
voice_store.append_turn(
defense_id,
DefenseTurn(
defense_id=defense_id,
seq=next_seq,
role=_ROLE_LEARNER,
text=text or "",
ts=datetime.now(UTC),
latency_ms=stt_ms,
created_at=datetime.now(UTC),
),
)
digest, _ = await _digest_for_task(trace_store, record.learner_id, record.task_id)
variant = variant_store.get_by_task(record.task_id)
llm_started = time.perf_counter()
question = await examiner.next_question(
history=history + [Message(role="user", content=text or "")],
trace_digest=digest,
variant_statement=variant.statement if variant is not None else None,
)
llm_ms = int((time.perf_counter() - llm_started) * 1000)
voice_store.append_turn(
defense_id,
DefenseTurn(
defense_id=defense_id,
seq=next_seq + 1,
role=_ROLE_EXAMINER,
text=question,
ts=datetime.now(UTC),
latency_ms=llm_ms,
created_at=datetime.now(UTC),
),
)
return AnswerResponse(
question=question,
turn_latency={"stt_ms": stt_ms, "llm_ms": llm_ms, "tts_ms": None},
)
@router.get("/{defense_id}/audio/{turn_id}")
async def defense_audio(
defense_id: str,
turn_id: int,
voice_store: DefenseStore = Depends(get_voice_store),
voice_provider=Depends(get_voice_provider),
):
record = voice_store.get(defense_id)
if record is None:
raise HTTPException(status_code=404, detail=f"no defense {defense_id!r}")
turn = next((t for t in record.turns if t.seq == turn_id), None)
if turn is None or turn.role != _ROLE_EXAMINER:
raise HTTPException(status_code=404, detail=f"no examiner turn {turn_id!r}")
async def stream():
async for chunk in voice_provider.synthesize(turn.text):
yield chunk
return StreamingResponse(stream(), media_type="audio/wav")
@router.post("/{defense_id}/finish", response_model=FinishResponse)
async def finish_defense(
defense_id: str,
voice_store: DefenseStore = Depends(get_voice_store),
examiner: ExaminerAgent = Depends(get_examiner),
trace_store=Depends(get_trace_store),
variant_store=Depends(get_variant_store),
) -> FinishResponse:
record = voice_store.get(defense_id)
if record is None:
raise HTTPException(status_code=404, detail=f"no defense {defense_id!r}")
turns = record.turns if hasattr(record, "turns") else []
history = [
Message(role="assistant" if t.role == _ROLE_EXAMINER else "user", content=t.text)
for t in turns
]
digest, _ = await _digest_for_task(trace_store, record.learner_id, record.task_id)
variant = variant_store.get_by_task(record.task_id)
verdict = await examiner.final_verdict(
history=history,
trace_digest=digest,
variant_statement=variant.statement if variant is not None else None,
)
signals: dict = {
"long_pauses": [
{"turn": t.seq, "latency_ms": t.latency_ms}
for t in turns
if t.role == _ROLE_LEARNER and (t.latency_ms or 0) > PAUSE_THRESHOLD_MS
],
"pause_threshold_ms": PAUSE_THRESHOLD_MS,
# Must-Have #1: "verdict + transcript persisted" — the verdict is
# stored INSIDE integrity_signals so GET /{id} after finish can
# re-serve it (the finish response alone would lose it). Signals
# are a JSON object dict (DefenseStore.finalize contract), so the
# verdict nests under the "verdict" key alongside the A-109
# markers the Proctor/Mentor feeds read.
"verdict": verdict.model_dump(),
}
voice_store.finalize(defense_id, signals)
return FinishResponse(verdict=verdict.model_dump(), integrity_signals=signals)
@router.get("/{defense_id}")
async def get_defense(
defense_id: str,
voice_store: DefenseStore = Depends(get_voice_store),
):
record = voice_store.get(defense_id)
if record is None:
raise HTTPException(status_code=404, detail=f"no defense {defense_id!r}")
return {
"defense_id": record.id,
"learner_id": record.learner_id,
"task_id": record.task_id,
"status": record.status,
"turns": [
{
"seq": t.seq,
"role": t.role,
"text": t.text,
"ts": t.ts,
"latency_ms": t.latency_ms,
}
for t in record.turns
],
"integrity_signals": record.integrity_signals or {},
}
-55
View File
@@ -2,21 +2,10 @@
from fastapi import Request
from ..agents.examiner import ExaminerAgent
from ..agents.registry import AgentRegistry
from ..agents.session import SessionStore
from ..config import Settings
from ..grading.engine import GradingEngine
from ..grading.store import GradeStore
from ..llm.base import LLMProvider
from ..sandbox.manager import SandboxManager
from ..sandbox.workdir import SandboxDir
from ..telemetry.ingest import TraceIntegrityMap
from ..telemetry.store import TraceStore
from ..variants.generator import VariantGenerator
from ..variants.store import VariantStore
from ..voice.base import VoiceProvider
from ..voice.defense_store import DefenseStore
def get_settings(request: Request) -> Settings:
@@ -33,47 +22,3 @@ def get_session_store(request: Request) -> SessionStore:
def get_agent_registry(request: Request) -> AgentRegistry:
return request.app.state.agent_registry
def get_sandbox_manager(request: Request) -> SandboxManager:
return request.app.state.sandbox_manager
def get_sandbox_test_layout(request: Request) -> SandboxDir | None:
"""Optional test seam (app.state.sandbox_test_layout); always None in prod."""
return getattr(request.app.state, "sandbox_test_layout", None)
def get_trace_store(request: Request) -> TraceStore:
return request.app.state.trace_store
def get_trace_integrity(request: Request) -> TraceIntegrityMap:
return request.app.state.trace_integrity
def get_grade_store(request: Request) -> GradeStore:
return request.app.state.grade_store
def get_grading_engine(request: Request) -> GradingEngine:
return request.app.state.grading_engine
def get_variant_generator(request: Request) -> VariantGenerator:
return request.app.state.variant_generator
def get_variant_store(request: Request) -> VariantStore:
return request.app.state.variant_store
def get_voice_store(request: Request) -> DefenseStore:
return request.app.state.defense_store
def get_voice_provider(request: Request) -> VoiceProvider:
return request.app.state.voice_provider
def get_examiner(request: Request) -> ExaminerAgent:
return request.app.state.examiner_agent
-83
View File
@@ -1,83 +0,0 @@
"""POST /v1/lab/feedback — SSE stream of Lab in-flow feedback (REQ-3-007).
v0.3 re-grounding: LIVE trace digest. Request carries {learner_id, task_id};
the digest is computed from the learner's real TraceStore events (D-028)
and handed to the Lab agent. No corpus scenarios. Empty/unknown trace is NOT
an error Lab gets a "no telemetry yet" timeline and coaches the baseline.
D-016 envelope with agent=lab.
"""
import json
from collections.abc import AsyncIterator
from fastapi import APIRouter, Depends
from pydantic import BaseModel, Field
from sse_starlette.sse import EventSourceResponse
from ..agents.registry import AgentRegistry
from ..config import Settings
from ..corpus.learner_context import get_learner_context
from ..grading.features import compute_digest
from .deps import (
get_agent_registry,
get_provider,
get_settings,
get_trace_store,
)
router = APIRouter(prefix="/v1")
class LabFeedbackRequest(BaseModel):
learner_id: str = Field(min_length=1)
task_id: str = Field(min_length=1)
@router.post("/lab/feedback")
async def lab_feedback(
body: LabFeedbackRequest,
registry: AgentRegistry = Depends(get_agent_registry),
settings: Settings = Depends(get_settings),
provider=Depends(get_provider),
trace_store=Depends(get_trace_store),
) -> EventSourceResponse:
agent = registry.get(provider, settings, "lab")
learner_context = get_learner_context(body.learner_id)
trace = (
trace_store.get_trace(body.learner_id, body.task_id)
if body.task_id in trace_store.list_tasks(body.learner_id)
else []
)
digest = compute_digest(trace) if trace else None
async def event_stream() -> AsyncIterator[dict]:
yield {"event": "message", "data": json.dumps({
"type": "meta",
"agent": "lab",
"task_id": body.task_id,
"model": settings.model,
})}
first_byte = True
try:
async for token in agent.stream_feedback(digest, learner_context):
first_byte = False
yield {"event": "message", "data": json.dumps({
"type": "delta", "content": token
})}
yield {"event": "message", "data": json.dumps({
"type": "done", "finish_reason": "stop"
})}
yield {"event": "message", "data": "[DONE]"}
except Exception as exc:
code = "provider_unavailable" if first_byte else "provider_error"
yield {"event": "message", "data": json.dumps({
"type": "error", "code": code, "message": str(exc)
})}
# [DONE] from except, not finally — a yield in finally would
# re-raise after GeneratorExit on client disconnect.
yield {"event": "message", "data": "[DONE]"}
return EventSourceResponse(
event_stream(),
headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"},
)
-95
View File
@@ -1,95 +0,0 @@
"""POST /v1/mentor/narrative — SSE career narrative stream (REQ-2-010).
D-016 envelope with agent=mentor. Session-backed: the client supplies a
session_id; the Mentor keeps conversation context across follow-ups.
"""
import json
from collections.abc import AsyncIterator
from fastapi import APIRouter, Depends, HTTPException
from pydantic import BaseModel, Field
from sse_starlette.sse import EventSourceResponse
from ..agents.registry import AgentRegistry, UnknownAgentError
from ..agents.session import SessionStore
from ..config import Settings
from ..corpus.learner_context import get_learner_context
from ..llm.types import Message
from .deps import get_agent_registry, get_provider, get_session_store, get_settings
router = APIRouter(prefix="/v1")
class MentorNarrativeRequest(BaseModel):
session_id: str = Field(min_length=1)
prompt: str = Field(default="Narrate my trajectory.")
learner_id: str | None = None
@router.post("/mentor/narrative")
async def mentor_narrative(
body: MentorNarrativeRequest,
registry: AgentRegistry = Depends(get_agent_registry),
sessions: SessionStore = Depends(get_session_store),
settings: Settings = Depends(get_settings),
provider=Depends(get_provider),
) -> EventSourceResponse:
try:
agent = registry.get(provider, settings, "mentor")
except UnknownAgentError as exc:
raise HTTPException(status_code=422, detail=str(exc)) from None
learner_context = get_learner_context(body.learner_id)
session = await sessions.get(body.session_id)
if session is None:
session = await sessions.create(
body.session_id, agent="mentor", learner_id=body.learner_id or "learner-001"
)
history = await sessions.history_window(body.session_id)
user_message = Message(role="user", content=body.prompt)
await sessions.append(body.session_id, user_message)
async def event_stream() -> AsyncIterator[dict]:
yield {"event": "message", "data": json.dumps({
"type": "meta",
"agent": "mentor",
"session_id": body.session_id,
"model": settings.model,
})}
first_byte = True
reply_parts: list[str] = []
try:
async for token in agent.stream_reply(
history=history,
user_input=body.prompt,
learner_context=learner_context,
):
first_byte = False
reply_parts.append(token)
yield {"event": "message", "data": json.dumps({
"type": "delta", "content": token
})}
full_reply = "".join(reply_parts)
if full_reply:
await sessions.append(
body.session_id, Message(role="assistant", content=full_reply)
)
yield {"event": "message", "data": json.dumps({
"type": "done", "finish_reason": "stop"
})}
yield {"event": "message", "data": "[DONE]"}
except Exception as exc:
code = "provider_unavailable" if first_byte else "provider_error"
yield {"event": "message", "data": json.dumps({
"type": "error", "code": code, "message": str(exc)
})}
# [DONE] from except, not finally — a yield in finally would
# re-raise after GeneratorExit on client disconnect.
yield {"event": "message", "data": "[DONE]"}
return EventSourceResponse(
event_stream(),
headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"},
)
-76
View File
@@ -1,76 +0,0 @@
"""POST /v1/proctor/signals — integrity signals over REAL inputs (REQ-3-007).
v0.3 re-grounding: live trace digest + DefenseStore long-pause signals +
variant seed cross-check, no corpus scenarios. The proctor coaches:
a pydantic-validated ProctorAssessment (JSON response, not SSE).
"""
from fastapi import APIRouter, Depends, HTTPException
from pydantic import BaseModel, Field
from ..agents.proctor import ProctorAssessment
from ..agents.registry import AgentRegistry
from ..config import Settings
from ..corpus.learner_context import get_learner_context
from ..grading.features import compute_digest
from .deps import (
get_agent_registry,
get_provider,
get_settings,
get_trace_store,
get_variant_store,
get_voice_store,
)
router = APIRouter(prefix="/v1")
class ProctorRequest(BaseModel):
learner_id: str = Field(min_length=1)
task_id: str = Field(min_length=1)
@router.post("/proctor/signals", response_model=ProctorAssessment)
async def proctor_signals(
body: ProctorRequest,
registry: AgentRegistry = Depends(get_agent_registry),
settings: Settings = Depends(get_settings),
provider=Depends(get_provider),
trace_store=Depends(get_trace_store),
variant_store=Depends(get_variant_store),
voice_store=Depends(get_voice_store),
) -> ProctorAssessment:
"""Real integrity inputs: live digest + defense signals + variant context."""
agent = registry.get(provider, settings, "proctor")
learner_context = get_learner_context(body.learner_id)
trace = (
trace_store.get_trace(body.learner_id, body.task_id)
if body.task_id in trace_store.list_tasks(body.learner_id)
else []
)
digest = compute_digest(trace) if trace else None
defense_signals = None
for record in voice_store.list_for_learner(body.learner_id):
if record.task_id == body.task_id and record.status == "finished":
defense_signals = record.integrity_signals or None
break
variant = variant_store.get_by_task(body.task_id)
variant_context = (
{"template_id": variant.template_id, "seed": variant.seed, "params": variant.params}
if variant is not None
else None
)
try:
return await agent.assess(
digest,
defense_signals=defense_signals,
variant_context=variant_context,
learner_context=learner_context,
)
except Exception as exc:
raise HTTPException(
status_code=502, detail=f"proctor assessment failed: {exc}"
) from exc
-353
View File
@@ -1,353 +0,0 @@
"""/v1/sandboxes — sandbox lifecycle API with G-5 abuse control (REQ-3-001).
The sandbox fabric (`sandbox/manager.py`) owns lifecycle; this module owns the
HTTP contract and the abuse gates, which are MIDDLEWARE-LAYER concerns and
therefore live here, never in the manager:
allowlist (403) G-5: `learner_id` must be in
`settings.learner_allowlist`. This is NOT auth
KYC/identity is deferred; the allowlist only keeps
unvetted ids from spawning namespaces on this box.
per-learner cap (429) `settings.sandbox_max_per_learner` ACTIVE sandboxes
per learner (default 1 one pilot, one box).
global create cap (429) `settings.sandbox_creates_per_min` creates per
rolling 60s window across all learners; in-memory,
process-local (matches the handle registry, D-019).
pool full (503) D-032 capacity guard (`PoolFullError`), no queue.
Response models are local to the api/ surface. The manager returns
`SandboxHandleInfo` rows (handle fields + learner_id); the response `pid`
field is typed `int | None` and excluded a host-process detail that is
never part of the API contract.
"""
import time
from collections import deque
from pathlib import Path
from fastapi import APIRouter, Depends, HTTPException, Response
from pydantic import BaseModel, ConfigDict, Field
from ..config import Settings
from ..sandbox.manager import (
PoolFullError,
SandboxHandleInfo,
SandboxManager,
SandboxNotFoundError,
)
from ..sandbox.workdir import SandboxDir
from .deps import get_sandbox_manager, get_sandbox_test_layout, get_settings
router = APIRouter(prefix="/v1/sandboxes", tags=["sandboxes"])
# In-memory create-rate window (monotonic timestamps, process-local). Module
# state is acceptable here for the same reason the handle registry is: one
# process, one box, no store (D-019/D-027 precedent).
_CREATE_TIMES: deque[float] = deque()
# -- contracts ----------------------------------------------------------------
class SandboxCreateRequest(BaseModel):
learner_id: str = Field(min_length=1)
# Optional task key: when set, the sandbox is telemetry-wired (REQ-3-003)
# — the in-sandbox capture agent streams workspace events to the ingest.
task_id: str | None = None
class SandboxResponse(BaseModel):
"""Public sandbox handle. `workdir` is the absolute host path."""
model_config = ConfigDict(frozen=True)
id: str
learner_id: str
workdir: str
created_at: str
pid: int | None = Field(
default=None,
exclude=True, # host-process detail; never part of the API contract
)
def _to_response(info: SandboxHandleInfo) -> SandboxResponse:
return SandboxResponse(
id=info.id,
learner_id=info.learner_id,
workdir=str(info.workdir),
created_at=info.created_at.isoformat(),
pid=info.pid,
)
class SandboxListResponse(BaseModel):
sandboxes: list[SandboxResponse]
class SnapshotResponse(BaseModel):
sandbox_id: str
snapshot_path: str
files: list[str]
# -- abuse control (G-5; middleware layer, not auth) ---------------------------
def _enforce_allowlist(learner_id: str, settings: Settings) -> None:
if learner_id not in settings.learner_allowlist:
raise HTTPException(
status_code=403,
detail=f"learner_id {learner_id!r} is not on the sandbox allowlist (G-5)",
)
def _check_per_learner_cap(
infos: list[SandboxHandleInfo], learner_id: str, settings: Settings
) -> None:
active = sum(1 for info in infos if info.learner_id == learner_id)
if active >= settings.sandbox_max_per_learner:
raise HTTPException(
status_code=429,
detail=(
f"learner {learner_id!r} already has {active} active "
f"sandbox(es); per-learner cap is {settings.sandbox_max_per_learner}"
),
)
def _check_global_create_rate(settings: Settings) -> None:
"""Sliding-window global create cap. Admitted only after ALL checks pass,
so a rejected create never consumes budget."""
now = time.monotonic()
while _CREATE_TIMES and now - _CREATE_TIMES[0] > 60.0:
_CREATE_TIMES.popleft()
if len(_CREATE_TIMES) >= settings.sandbox_creates_per_min:
raise HTTPException(
status_code=429,
detail=(
f"global sandbox create rate exceeded "
f"({settings.sandbox_creates_per_min}/min); retry shortly"
),
)
_CREATE_TIMES.append(now)
# -- endpoints ---------------------------------------------------------------
@router.post("", status_code=201, response_model=SandboxResponse)
async def create_sandbox(
body: SandboxCreateRequest,
manager: SandboxManager = Depends(get_sandbox_manager),
settings: Settings = Depends(get_settings),
) -> SandboxResponse:
_enforce_allowlist(body.learner_id, settings)
_check_per_learner_cap(await manager.list(), body.learner_id, settings)
_check_global_create_rate(settings)
try:
info = await manager.create(body.learner_id, task_id=body.task_id)
except PoolFullError as exc:
raise HTTPException(status_code=503, detail=str(exc)) from exc
return _to_response(info)
@router.get("", response_model=SandboxListResponse)
async def list_sandboxes(
manager: SandboxManager = Depends(get_sandbox_manager),
) -> SandboxListResponse:
return SandboxListResponse(
sandboxes=[_to_response(info) for info in await manager.list()]
)
@router.get("/{sandbox_id}", response_model=SandboxResponse)
async def get_sandbox(
sandbox_id: str,
manager: SandboxManager = Depends(get_sandbox_manager),
) -> SandboxResponse:
try:
info = await manager.get(sandbox_id)
except SandboxNotFoundError:
raise HTTPException(
status_code=404, detail=f"unknown sandbox {sandbox_id!r}"
) from None
return _to_response(info)
@router.post("/{sandbox_id}/snapshot", response_model=SnapshotResponse)
async def snapshot_sandbox(
sandbox_id: str,
response: Response,
manager: SandboxManager = Depends(get_sandbox_manager),
test_layout: SandboxDir | None = Depends(get_sandbox_test_layout),
) -> SnapshotResponse:
try:
snapshot_path = await manager.snapshot(sandbox_id)
except SandboxNotFoundError:
raise HTTPException(
status_code=404, detail=f"unknown sandbox {sandbox_id!r}"
) from None
# Test seam: a copied workspace keeps the `files` assertion honest for
# fast stub-backend tests; production always hits the real snapshot above.
if test_layout is not None and test_layout.snapshots.is_dir():
copies = sorted(test_layout.snapshots.iterdir())
if copies:
response.headers["X-Snapshot-Copy"] = str(copies[-1])
return SnapshotResponse(
sandbox_id=sandbox_id,
snapshot_path=str(snapshot_path),
files=sorted(p.name for p in snapshot_path.iterdir()),
)
@router.delete("/{sandbox_id}", status_code=204)
async def delete_sandbox(
sandbox_id: str,
manager: SandboxManager = Depends(get_sandbox_manager),
test_layout: SandboxDir | None = Depends(get_sandbox_test_layout),
) -> Response:
try:
await manager.get(sandbox_id)
except SandboxNotFoundError:
raise HTTPException(
status_code=404, detail=f"unknown sandbox {sandbox_id!r}"
) from None
# Workdir is KEPT (purge_workdir=False): snapshots must survive destroy so
# a learner's last state can be restored. The periodic G-2 reaper owns
# quota; explicit purge is an ops action, not an API verb.
await manager.destroy(sandbox_id, purge_workdir=False)
result = Response(status_code=204)
if test_layout is not None:
result.headers["X-Workspace-Copy"] = str(test_layout.workspace)
return result
# -- workspace files + exec (Phase 6, REQ-3-008; CUT-2) -------------------------
#
# The build surface reads/writes/list workspace files and runs Run/Test
# commands through the manager's backend. NO interactive shell relay (CUT-2:
# keystroke-level stdin/stdout is v0.4) — each exec is a bounded command with
# captured output. Paths are WORKSPACE-RELATIVE; traversal outside the
# workspace is rejected (the workdir bind is the boundary, but the API adds
# its own containment check — defense in depth).
class FileWriteRequest(BaseModel):
path: str = Field(min_length=1)
content: str
class ExecRequest(BaseModel):
cmd: list[str] = Field(min_length=1)
class ExecResponse(BaseModel):
cmd: list[str]
returncode: int
stdout: str
stderr: str
duration_s: float
async def _workspace_dir(manager: SandboxManager, sandbox_id: str):
"""Resolve the sandbox workspace (tracked layout or shell layout)."""
info = await manager.get(sandbox_id) # raises SandboxNotFoundError -> 404
backend = manager._backend # noqa: SLF001 - API owns the composition seam
tracked = getattr(backend, "_tracked", {}).get(sandbox_id)
if tracked is not None:
return tracked.workspace, info
return info.workdir / "workspace", info
def _safe_rel_path(raw: str) -> Path:
"""Workspace-relative path; reject absolute/traversal paths."""
candidate = Path(raw)
if candidate.is_absolute() or ".." in candidate.parts:
raise HTTPException(status_code=422, detail=f"invalid workspace path {raw!r}")
return candidate
def _resolve_in_workspace(workspace: Path, rel: Path) -> Path:
"""Resolve `rel` under `workspace`, refusing symlink escapes (P7).
The lexical check in `_safe_rel_path` cannot see symlinks: an exec can
plant `ln -s /etc target` in the workspace and a follow-up read/write
would follow it OUT of the bind. Resolve with the workspace as the
anchor (strict: a symlink chain escaping raises) and confirm the
normalized target still sits inside the workspace defense in depth
for both read_file and write_file.
"""
try:
target = (workspace / rel).resolve(strict=False)
target.relative_to(workspace.resolve(strict=False))
except ValueError:
raise HTTPException(
status_code=422, detail=f"path escapes the workspace: {rel.as_posix()!r}"
) from None
return target
@router.get("/{sandbox_id}/files")
async def list_files(
sandbox_id: str,
manager: SandboxManager = Depends(get_sandbox_manager),
) -> dict:
try:
workspace, _ = await _workspace_dir(manager, sandbox_id)
except SandboxNotFoundError:
raise HTTPException(status_code=404, detail=f"no sandbox {sandbox_id!r}") from None
return {"files": sorted(p.name for p in workspace.iterdir()) if workspace.is_dir() else []}
@router.get("/{sandbox_id}/files/{path:path}")
async def read_file(
sandbox_id: str,
path: str,
manager: SandboxManager = Depends(get_sandbox_manager),
) -> dict:
try:
workspace, _ = await _workspace_dir(manager, sandbox_id)
except SandboxNotFoundError:
raise HTTPException(status_code=404, detail=f"no sandbox {sandbox_id!r}") from None
rel = _safe_rel_path(path)
target = _resolve_in_workspace(workspace, rel)
if not target.is_file():
raise HTTPException(status_code=404, detail=f"no file {path!r}")
return {"path": path, "content": target.read_text(errors="replace")}
@router.put("/{sandbox_id}/files/{path:path}")
async def write_file(
sandbox_id: str,
path: str,
body: FileWriteRequest,
manager: SandboxManager = Depends(get_sandbox_manager),
) -> dict:
try:
workspace, _ = await _workspace_dir(manager, sandbox_id)
except SandboxNotFoundError:
raise HTTPException(status_code=404, detail=f"no sandbox {sandbox_id!r}") from None
rel = _safe_rel_path(body.path)
target = _resolve_in_workspace(workspace, rel)
target.parent.mkdir(parents=True, exist_ok=True)
target.write_text(body.content)
return {"path": body.path, "written": True}
@router.post("/{sandbox_id}/exec", response_model=ExecResponse)
async def exec_command(
sandbox_id: str,
body: ExecRequest,
manager: SandboxManager = Depends(get_sandbox_manager),
) -> ExecResponse:
try:
await manager.get(sandbox_id)
except SandboxNotFoundError:
raise HTTPException(status_code=404, detail=f"no sandbox {sandbox_id!r}") from None
backend = manager._backend # noqa: SLF001 - API owns the composition seam
handle = manager._handles.get(sandbox_id) # noqa: SLF001
if handle is None:
raise HTTPException(status_code=404, detail=f"no live handle {sandbox_id!r}")
result = await backend.exec(handle, body.cmd)
return ExecResponse(**result.model_dump())
-175
View File
@@ -1,175 +0,0 @@
"""/v1/telemetry — WS ingest + trace read endpoints (REQ-3-003, D-026, G-3).
The router composes the telemetry engine via DI: `telemetry/ingest.py` owns
the WS protocol (frame contract + flood control + keepalive) and this module
only wires `app.state.trace_store` / `app.state.trace_integrity` /
`app.state.settings` into it, plus the two HTTP read faces:
WS /v1/telemetry/ingest?learner_id&task_id[&sandbox_id] (D-026)
GET /v1/telemetry/traces/{learner_id}/{task_id} ordered trace; 404 unknown
GET /v1/telemetry/gaps/{learner_id}/{task_id} missing seqs ; 404 unknown
The WS route is a thin DI shell: it validates the query-param identity and
the Origin (browser pages are gated to the localhost dev origins CORS
middleware does not cover WS upgrades; the stdlib capture agent sends no
Origin and is unaffected), pulls store/integrity/settings from `app.state`,
and calls `telemetry_ingest_endpoint(...)` the engine stays
FastAPI-DI-free so it's testable without a router and the api/ layer owns
all composition.
Unknown-trace contract: a trace is KNOWN when it has >=1 stored event OR
carries an integrity flag a flooded trace with zero stored rows still 200s
so Proctor/grader can read WHY it's unusable (G-4 consumes
`integrity_reason`). `TraceResponse.incomplete` / `.integrity_reason` mirror
the map so HTTP consumers never touch process internals.
"""
from __future__ import annotations
from fastapi import APIRouter, Depends, HTTPException, WebSocket
from pydantic import BaseModel
from ..telemetry.ingest import (
TraceIntegrityMap,
telemetry_ingest_endpoint,
)
from ..telemetry.models import TelemetryEvent
from ..telemetry.store import TraceStore
from .deps import get_trace_integrity, get_trace_store
router = APIRouter(prefix="/v1/telemetry", tags=["telemetry"])
#: Browser Origins allowed to open the ingest socket (A-008 mirror, D-038).
#: The stdlib capture agent sends NO Origin header (it is not a browser) and
#: stays allowed; a malicious page loaded in the learner's browser would
#: carry an Origin and must not be able to poison/flood the trace. CORS
#: middleware does NOT cover WebSocket upgrades, so this gate is explicit.
#: In network mode the configured CORS list governs (default '*' — any
#: origin, since credentials are never used); an explicit list still rejects
#: unlisted origins with 1008.
_LOCAL_WS_ORIGINS = frozenset(
{"http://localhost:3000", "http://127.0.0.1:3000", "http://localhost:8420"}
)
def _allowed_ws_origins(settings: object) -> frozenset[str]:
configured = getattr(settings, "cors_origin_list", None)
if configured is None:
return _LOCAL_WS_ORIGINS
if configured == ["*"]:
return frozenset() # empty = wildcard = every Origin passes
return frozenset(configured) | _LOCAL_WS_ORIGINS
# --- WS ingest (D-026) ---------------------------------------------------------
@router.websocket("/ingest")
async def telemetry_ingest_ws(websocket: WebSocket) -> None:
"""DI shell: resolve app.state services, then hand the socket to the engine.
The engine's session + flood logic is fully typed and testable without
FastAPI; this shim is the only place the two layers meet.
"""
origin = (websocket.headers.get("origin") or "").strip()
allowed = _allowed_ws_origins(getattr(websocket.app.state, "settings", None))
if origin and allowed and origin not in allowed:
# Same-origin dev pages (Next.js on :3000, the service itself on
# :8420) pass; anything else is refused pre-accept. Non-browser
# producers (the capture agent, tests) send no Origin and pass.
# Wildcard (empty frozenset) passes every Origin in network mode.
await websocket.close(
code=1008, reason=f"origin {origin!r} not allowed for telemetry ingest"
)
return
query = websocket.query_params
learner_id = query.get("learner_id", "")
task_id = query.get("task_id", "")
if not learner_id or not task_id:
# Reject BEFORE accept: closing the pre-accept handshake is the
# cheapest denial and unambiguous for the stdlib capture agent.
await websocket.close(code=1008, reason="learner_id/task_id query params required")
return
await telemetry_ingest_endpoint(
websocket=websocket,
learner_id=learner_id,
task_id=task_id,
sandbox_id=query.get("sandbox_id", ""),
store=websocket.app.state.trace_store,
integrity=websocket.app.state.trace_integrity,
settings=websocket.app.state.settings,
)
# --- HTTP reads -----------------------------------------------------------------
class TraceResponse(BaseModel):
"""Ordered trace + integrity signal (G-4 reads incomplete/reason)."""
learner_id: str
task_id: str
events: list[TelemetryEvent]
incomplete: bool
integrity_reason: str | None
class GapsResponse(BaseModel):
"""Missing seqs + integrity signal."""
learner_id: str
task_id: str
gaps: list[int]
incomplete: bool
integrity_reason: str | None
def _is_known_trace(
store: TraceStore, integrity: TraceIntegrityMap, learner_id: str, task_id: str
) -> bool:
"""Known = has stored events OR carries an integrity flag (a flooded trace
with zero rows must still be readable Proctor needs the reason)."""
return (
store.latest_seq(learner_id, task_id) >= 0
or integrity.is_incomplete(learner_id, task_id)
)
@router.get("/traces/{learner_id}/{task_id}", response_model=TraceResponse)
async def get_trace(
learner_id: str,
task_id: str,
store: TraceStore = Depends(get_trace_store),
integrity: TraceIntegrityMap = Depends(get_trace_integrity),
) -> TraceResponse:
if not _is_known_trace(store, integrity, learner_id, task_id):
raise HTTPException(
status_code=404, detail=f"unknown trace {learner_id!r}/{task_id!r}"
)
return TraceResponse(
learner_id=learner_id,
task_id=task_id,
events=store.get_trace(learner_id, task_id),
incomplete=integrity.is_incomplete(learner_id, task_id),
integrity_reason=integrity.reason(learner_id, task_id),
)
@router.get("/gaps/{learner_id}/{task_id}", response_model=GapsResponse)
async def get_gaps(
learner_id: str,
task_id: str,
store: TraceStore = Depends(get_trace_store),
integrity: TraceIntegrityMap = Depends(get_trace_integrity),
) -> GapsResponse:
if not _is_known_trace(store, integrity, learner_id, task_id):
raise HTTPException(
status_code=404, detail=f"unknown trace {learner_id!r}/{task_id!r}"
)
return GapsResponse(
learner_id=learner_id,
task_id=task_id,
gaps=store.gaps(learner_id, task_id),
incomplete=integrity.is_incomplete(learner_id, task_id),
integrity_reason=integrity.reason(learner_id, task_id),
)
-187
View File
@@ -1,187 +0,0 @@
"""/v1/variants — seeded per-learner task variant endpoints (REQ-3-005, D-029).
Three faces over the variant engine (the generator + store stay FastAPI-free;
this module owns all HTTP wiring the D-027/D-032 house pattern):
POST /v1/variants {learner_id, template_id | competency_id}
200 the learner's variant — GENERATED on the first request,
CACHED (store read, zero LLM calls) on every repeat: D-029
reproducibility means one (learner_id, template_id) is ONE
variant forever, so a regenerate is always a 200 of the SAME
variant, never a second render.
404 unknown template_id, or competency_id with no bound template.
422 neither template_id nor competency_id given.
GET /v1/variants/{task_id}
200 the stored variant owning the task key (the grading and
telemetry join path); 404 when no variant was ever generated
for the task.
GET /v1/variants?learner_id=...
200 the learner's variants, chronological; [] when none.
Template resolution: an explicit `template_id` wins; without it the FIRST
template bound to `competency_id` is used (`template_for_competency`,
D-021 corpus alignment). The response carries `competency_id` resolved
from the template library at read time an enrichment, not a persisted
column (the seed re-derives the whole variant, D-029) so the learner
surface can bind a variant to its competency without a library round-trip.
Distinctness (REQ-3-005): different learners on the same template draw
different seeded params and receive distinct statements and task_ids;
tests/api/test_variants.py asserts this end-to-end through the API.
"""
from datetime import datetime
from typing import Self
from fastapi import APIRouter, Depends, HTTPException
from pydantic import BaseModel, Field, model_validator
from ..variants.generator import VariantGenerator
from ..variants.store import VariantRecord, VariantStore
from ..variants.templates import TaskTemplate, get_template, template_for_competency
from .deps import get_variant_generator, get_variant_store
router = APIRouter(prefix="/v1/variants", tags=["variants"])
# -- contracts ------------------------------------------------------------------
class VariantGenerateRequest(BaseModel):
"""One variant identity: an explicit `template_id`, or the first
template bound to a `competency_id` (D-021). `template_id` wins when
both are given (explicit identity beats derived); at least one is
required 422 otherwise.
"""
learner_id: str = Field(min_length=1)
template_id: str | None = None
competency_id: str | None = None
@model_validator(mode="after")
def _require_template_or_competency(self) -> Self:
if self.template_id is None and self.competency_id is None:
raise ValueError("template_id or competency_id is required")
return self
class VariantResponse(BaseModel):
"""VariantRecord over HTTP, plus the `competency_id` enrichment.
Every field except `competency_id` mirrors `VariantRecord` exactly
(snake_case; `created_at` is an ISO 8601 UTC datetime) the wire shape
typed as `TaskVariant` in packages/types/variants.ts.
"""
learner_id: str
task_id: str
template_id: str
competency_id: str
seed: str
params: dict[str, str | int]
statement: str
starter_files: dict[str, str]
created_at: datetime
class VariantListResponse(BaseModel):
variants: list[VariantResponse]
# -- resolution + rendering -----------------------------------------------------
def _resolve_template(body: VariantGenerateRequest) -> TaskTemplate:
"""Template for the request: the explicit id, else the first template
bound to the competency; 404 when neither resolves."""
if body.template_id is not None:
template = get_template(body.template_id)
if template is None:
raise HTTPException(
status_code=404,
detail=f"no task template with id {body.template_id!r}",
)
return template
# The request validator guarantees the disjunction, so reaching here
# means a competency_id was given (never None).
assert body.competency_id is not None
templates = template_for_competency(body.competency_id)
if not templates:
raise HTTPException(
status_code=404,
detail=f"no task template for competency {body.competency_id!r}",
)
return templates[0]
def _competency_for(template_id: str) -> str:
"""competency_id enrichment for stored records (read paths)."""
template = get_template(template_id)
if template is None:
# Integrity guard: a stored variant referencing a template that is
# no longer in the library cannot be enriched; fail loudly rather
# than fabricate a competency binding.
raise HTTPException(
status_code=500,
detail=f"stored variant references unknown template {template_id!r}",
)
return template.competency_id
def _to_response(record: VariantRecord, competency_id: str) -> VariantResponse:
return VariantResponse(
learner_id=record.learner_id,
task_id=record.task_id,
template_id=record.template_id,
competency_id=competency_id,
seed=record.seed,
params=dict(record.params),
statement=record.statement,
starter_files=dict(record.starter_files),
created_at=record.created_at,
)
# -- endpoints ------------------------------------------------------------------
@router.post("", response_model=VariantResponse)
async def generate_variant(
body: VariantGenerateRequest,
generator: VariantGenerator = Depends(get_variant_generator),
) -> VariantResponse:
"""The learner's variant for the resolved template — generated on the
first request, cached (no LLM call) on every repeat: D-029 makes a
regenerate a 200 of the SAME stored variant.
"""
template = _resolve_template(body)
record = await generator.generate(body.learner_id, template.id)
return _to_response(record, competency_id=template.competency_id)
@router.get("", response_model=VariantListResponse)
async def list_variants(
learner_id: str,
store: VariantStore = Depends(get_variant_store),
) -> VariantListResponse:
"""All stored variants for the learner, chronological; [] when none."""
variants = [
_to_response(record, competency_id=_competency_for(record.template_id))
for record in store.list_for_learner(learner_id)
]
return VariantListResponse(variants=variants)
@router.get("/{task_id}", response_model=VariantResponse)
async def get_variant(
task_id: str,
store: VariantStore = Depends(get_variant_store),
) -> VariantResponse:
"""The stored variant owning the task key — the grading and telemetry
join path; 404 when no variant was ever generated for the task."""
record = store.get_by_task(task_id)
if record is None:
raise HTTPException(
status_code=404, detail=f"no stored variant for task {task_id!r}"
)
return _to_response(record, competency_id=_competency_for(record.template_id))
+1 -83
View File
@@ -1,13 +1,6 @@
"""Service settings — pydantic-settings, env prefix AI_, .env support."""
from pathlib import Path
from typing import Annotated
from pydantic import field_validator
from pydantic_settings import BaseSettings, NoDecode, SettingsConfigDict
# apps/ai-service/ (sandbox dir default is relative to the app, not the CWD)
_SERVICE_ROOT = Path(__file__).resolve().parent.parent
from pydantic_settings import BaseSettings, SettingsConfigDict
class Settings(BaseSettings):
@@ -23,78 +16,3 @@ class Settings(BaseSettings):
# "auto" sends response_format and degrades on 400; "off" never sends it
json_mode: str = "auto"
# v0.3 sandbox fabric (REQ-3-001): root holding per-sandbox workdirs.
# Relative paths resolve against the app dir (apps/ai-service/), not the CWD.
sandbox_dir: Path = _SERVICE_ROOT / "sandboxes"
# D-032: single-box capacity, no queue — pool full → API maps to 503.
sandbox_max_concurrent: int = 5
# Wall-clock ceiling per sandbox; the manager's async reaper destroys
# sandboxes idle past this age (same timer runs the G-2 workdir sweep).
sandbox_timeout_s: float = 900.0
# G-2: soft disk cap per sandbox workdir, enforced best-effort by the
# manager sweep (NOT kernel-enforced — no cgroup delegation/sudo here).
sandbox_max_workdir_mb: int = 512
# G-5 abuse control (NOT auth — KYC/auth is deferred; these keep the
# single-box pilot from melting down before identity lands):
#
# Server-side learner allowlist. Env form is a COMMA-SEPARATED string
# (e.g. AI_LEARNER_ALLOWLIST="pilot-learner,learner-2"); NoDecode skips
# pydantic-settings' JSON decoding of complex types and the validator
# below splits/strips/drops empties. Default: the single mock pilot id.
learner_allowlist: Annotated[list[str], NoDecode] = ["pilot-learner"]
# Max ACTIVE sandboxes per learner → API maps excess to 429.
sandbox_max_per_learner: int = 1
# Global create-rate ceiling (creates per rolling 60s window, shared
# across learners) → API maps excess to 429. In-memory, process-local.
sandbox_creates_per_min: int = 10
@field_validator("learner_allowlist", mode="before")
@classmethod
def _split_allowlist_csv(cls, value: object) -> object:
if isinstance(value, str):
return [item.strip() for item in value.split(",") if item.strip()]
return value
# D-027: SQLite path for telemetry/grades/variants/defenses stores.
db_path: Path = _SERVICE_ROOT / "ai_service" / "data" / "nextcraft.db"
# G-3 flood control (NOT backpressure-by-silence): max events ingested per
# (learner_id, task_id) trace before the WS endpoint closes the connection
# with 1008 and marks the trace INCOMPLETE_FLOODED. Drop-oldest is
# FORBIDDEN — it corrupts grading input (GRILL G-3).
telemetry_max_events_per_task: int = 50000
# Sandbox telemetry wiring (REQ-3-003): loopback host the in-sandbox capture
# agent dials to reach this service's WS ingest (the agent joins the sandbox
# mount ns but NOT the net ns — exec namespaces are offline, so the agent
# shares the host network and reaches the app over loopback). Port reuses
# `port` (A-004); only the host is configurable — never a second port.
telemetry_ingest_host: str = "127.0.0.1"
# v0.3.5 network mode (D-038): dev.sh binds 0.0.0.0 so remote browsers can
# reach the stack; '*' (default) lets any origin call the API (safe ONLY
# because credentials are never enabled — A-008). Set a comma-separated
# origin list (e.g. 'http://nextcraft-1:3000') to restrict instead.
cors_origins: str = "*"
@property
def cors_origin_list(self) -> list[str]:
value = self.cors_origins.strip()
if value == "*":
return ["*"]
return [o.strip() for o in value.split(",") if o.strip()]
# Voice provider selection (REQ-3-006, D-030): 'mock' (default — the
# no-key path is first-class; tests never call a real voice API) or
# 'browser' (browser-native SpeechRecognition/speechSynthesis fallback;
# the descriptor tells the web client). The real server STT/TTS
# ('openai-audio') is a v0.4 seam (GRILL CUT-1 / G-7) — AI_VOICE_BASE_URL
# and AI_VOICE_API_KEY are documented in .env.example for that future.
voice_provider: str = "mock"
@@ -8,8 +8,4 @@ convention; revisit codegen only if drift bites (v0.3).
from .learner_context import LEARNER_CONTEXTS, LearnerContext, get_learner_context
__all__ = [
"LEARNER_CONTEXTS",
"LearnerContext",
"get_learner_context",
]
__all__ = ["LEARNER_CONTEXTS", "LearnerContext", "get_learner_context"]
@@ -1,214 +0,0 @@
"""Pre-baked artifacts + rubrics + defense transcripts — Assessor mock inputs (REQ-2-008).
v0.2 mock engine inputs (pre-baked artifacts/rubrics/transcripts) DORMANT as of v0.3 re-
grounding (Task 6-1-04): no production code path imports this module. Retained as Phase-3
calibration history.
Counterpart: packages/mock-data/ai-scenarios.ts (artifact IDs string-identical,
D-021). Real process-trace grading is a v0.3+ engine (assessment engine);
these pre-baked submissions stand in for artifact + defense evaluation.
"""
from pydantic import BaseModel
class RubricCriterion(BaseModel):
criterion_id: str
name: str
weight: float
description: str
class AssessmentRubric(BaseModel):
rubric_id: str
competency_id: str
criteria: list[RubricCriterion]
class ArtifactSubmission(BaseModel):
artifact_id: str
name: str
artifact_type: str # "code" | "design" | "simulation"
competency_id: str
description: str
evidence_excerpt: str # what the grader sees of the artifact itself
class DefenseTranscript(BaseModel):
transcript_id: str
artifact_id: str
turns: list[dict] # {"speaker": "examiner"|"learner", "text": "..."}
_RUBRIC_ORCHESTRATION = AssessmentRubric(
rubric_id="rubric-orchestration-c002",
competency_id="stack-orchestration-c002",
criteria=[
RubricCriterion(
criterion_id="rc-architecture",
name="Agent architecture soundness",
weight=0.3,
description="State boundaries and responsibilities are clearly separated",
),
RubricCriterion(
criterion_id="rc-communication",
name="Inter-agent communication design",
weight=0.3,
description="Message contracts are explicit, typed, and failure-aware",
),
RubricCriterion(
criterion_id="rc-reliability",
name="Reliability engineering",
weight=0.25,
description="Retries, timeouts, and degradation paths handled",
),
RubricCriterion(
criterion_id="rc-process",
name="Process trace quality",
weight=0.15,
description="Telemetry shows iterative building with real checkpoints",
),
],
)
_RUBRIC_TOOL_USE = AssessmentRubric(
rubric_id="rubric-orchestration-c003",
competency_id="stack-orchestration-c003",
criteria=[
RubricCriterion(
criterion_id="rc-eval-design",
name="Evaluation design rigor",
weight=0.35,
description="Hypotheses, controls, and metrics are explicit and defensible",
),
RubricCriterion(
criterion_id="rc-eval-robustness",
name="Harness robustness",
weight=0.35,
description="Error handling, variance awareness, and reproducibility",
),
RubricCriterion(
criterion_id="rc-eval-insight",
name="Insight extraction",
weight=0.3,
description="Results are interpreted into concrete engineering decisions",
),
],
)
_ARTIFACT_RESEARCH_ASSISTANT = ArtifactSubmission(
artifact_id="art-eval-research-assistant",
name="Multi-agent research assistant (eval build)",
artifact_type="code",
competency_id="stack-orchestration-c002",
description=(
"LangGraph-based assistant planning, retrieving, drafting cited reviews."
),
evidence_excerpt=(
"planner.py defines state schema with explicit fields "
"(plan, findings, draft); tool_node.py wraps retrieval with a "
"3-retry loop and typed ToolMessage responses; tests cover "
"planner->tool->writer handoffs; README shows graph diagram"
),
)
_TRANSCRIPT_RESEARCH_ASSISTANT = DefenseTranscript(
transcript_id="defense-art-eval-research-assistant",
artifact_id="art-eval-research-assistant",
turns=[
{"speaker": "examiner",
"text": "Why did you give the planner sole write access to the plan field?"},
{"speaker": "learner",
"text": "So worker nodes can't mutate each other's inputs — "
"the state stays predictable and the graph is debuggable"},
{"speaker": "examiner",
"text": "What happens when the retrieval tool times out three times?"},
{"speaker": "learner",
"text": "The tool node degrades to a no-op ToolMessage with a "
"retry flag so the writer can fall back to existing findings"},
{"speaker": "examiner", "text": "How would you extend this to a third agent?"},
{"speaker": "learner",
"text": "Add a reviewer node with its own typed messages, same pattern"},
],
)
_ARTIFACT_RAG_DASHBOARD = ArtifactSubmission(
artifact_id="art-eval-rag-dashboard",
name="RAG retrieval quality dashboard (eval build)",
artifact_type="code",
competency_id="stack-orchestration-c003",
description="Dashboard comparing chunking strategies/rerankers across 800 queries.",
evidence_excerpt=(
"eval harness sweeps 4 chunk sizes x 3 rerankers; results table auto-generated; "
"no error handling on the query loader; tests only cover the happy path"
),
)
_TRANSCRIPT_RAG_DASHBOARD = DefenseTranscript(
transcript_id="defense-art-eval-rag-dashboard",
artifact_id="art-eval-rag-dashboard",
turns=[
{"speaker": "examiner", "text": "How did you control for query difficulty across runs?"},
{"speaker": "learner", "text": "I, um, used the same query set each time"},
{"speaker": "examiner", "text": "What happens if the query loader hits a malformed row?"},
{"speaker": "learner", "text": "I didn't handle that. It would probably crash."},
{"speaker": "examiner", "text": "What would you improve first?"},
{"speaker": "learner",
"text": "Probably add the error handling, then look at variance between runs"},
],
)
RUBRICS: dict[str, AssessmentRubric] = {
_RUBRIC_ORCHESTRATION.rubric_id: _RUBRIC_ORCHESTRATION,
_RUBRIC_TOOL_USE.rubric_id: _RUBRIC_TOOL_USE,
}
ARTIFACTS: dict[str, ArtifactSubmission] = {
a.artifact_id: a
for a in (_ARTIFACT_RESEARCH_ASSISTANT, _ARTIFACT_RAG_DASHBOARD)
}
TRANSCRIPTS: dict[str, DefenseTranscript] = {
t.transcript_id: t
for t in (_TRANSCRIPT_RESEARCH_ASSISTANT, _TRANSCRIPT_RAG_DASHBOARD)
}
def rubric_for_competency(competency_id: str) -> AssessmentRubric | None:
for rubric in RUBRICS.values():
if rubric.competency_id == competency_id:
return rubric
return None
def get_artifact_bundle(artifact_id: str) -> tuple[ArtifactSubmission, AssessmentRubric] | None:
"""Resolve (artifact, rubric) for an artifact ID; None if unknown."""
artifact = ARTIFACTS.get(artifact_id)
if artifact is None:
return None
rubric = rubric_for_competency(artifact.competency_id)
if rubric is None:
return None
return artifact, rubric
def get_transcript_for_artifact(artifact_id: str) -> DefenseTranscript | None:
for transcript in TRANSCRIPTS.values():
if transcript.artifact_id == artifact_id:
return transcript
return None
def render_rubric(rubric: AssessmentRubric) -> str:
lines = [f"Rubric: {rubric.rubric_id} (competency {rubric.competency_id})"]
for c in rubric.criteria:
lines.append(f"- {c.criterion_id} ({c.weight:.2f}): {c.name}{c.description}")
return "\n".join(lines)
def render_transcript(transcript: DefenseTranscript) -> str:
lines = [f"Defense transcript: {transcript.transcript_id}"]
for turn in transcript.turns:
lines.append(f"{turn['speaker']}: {turn['text']}")
return "\n".join(lines)
@@ -1,183 +0,0 @@
"""Simulated sandbox telemetry corpus — Lab agent mock engine inputs (REQ-2-007).
v0.2 mock engine inputs (Lab/Proctor scenarios) DORMANT as of v0.3 re-grounding (Task 6-1-04):
no production code path imports this module. Retained as Phase-3 calibration history
(corpus/trace_fixtures.py references it from TESTS only).
Counterpart: packages/mock-data/ai-scenarios.ts (scenario IDs string-identical,
D-021). Real sandbox telemetry is a v0.3+ engine (sandbox fabric); these
scripted event streams stand in for the build-session process trace.
"""
from pydantic import BaseModel
class TelemetryEvent(BaseModel):
timestamp: int # seconds since session start
kind: str # "keystroke_burst" | "file_save" | "run_tests" | "test_pass"
# | "test_fail" | "console_error" | "idle" | "paste" | "commit"
detail: str = ""
class LabTelemetryScenario(BaseModel):
scenario_id: str
title: str
competency_id: str
events: list[TelemetryEvent]
class ProctorEvent(BaseModel):
timestamp: int # seconds since session start
kind: str # "tab_switch" | "idle" | "paste_large" | "focus_lost" | "keystroke_burst"
detail: str = ""
class ProctorScenario(BaseModel):
scenario_id: str
title: str
competency_id: str
events: list[ProctorEvent]
_PROCTOR_SCENARIO_HEALTHY = ProctorScenario(
scenario_id="proctor-scenario-healthy",
title="Healthy defense session — focused throughout",
competency_id="stack-orchestration-c002",
events=[
ProctorEvent(timestamp=0, kind="keystroke_burst", detail="session begins"),
ProctorEvent(timestamp=310, kind="keystroke_burst", detail="long answer in progress"),
ProctorEvent(timestamp=640, kind="keystroke_burst", detail="revision pass"),
ProctorEvent(timestamp=900, kind="keystroke_burst", detail="final answer"),
],
)
_PROCTOR_SCENARIO_DISTRACTED = ProctorScenario(
scenario_id="proctor-scenario-distracted",
title="Distracted defense session — tab switches and idle gaps",
competency_id="stack-orchestration-c002",
events=[
ProctorEvent(timestamp=0, kind="keystroke_burst", detail="session begins"),
ProctorEvent(timestamp=120, kind="tab_switch", detail="to docs.nextjs.org"),
ProctorEvent(timestamp=125, kind="focus_lost", detail="window blur 40s"),
ProctorEvent(timestamp=300, kind="idle", detail="no activity for 5 minutes"),
ProctorEvent(timestamp=600, kind="tab_switch", detail="to github.com"),
ProctorEvent(timestamp=605, kind="focus_lost", detail="window blur 2m"),
ProctorEvent(timestamp=720, kind="keystroke_burst", detail="resumes typing"),
],
)
_PROCTOR_SCENARIO_FLAGGED = ProctorScenario(
scenario_id="proctor-scenario-flagged",
title="Flagged defense session — large paste during exam",
competency_id="stack-orchestration-c003",
events=[
ProctorEvent(timestamp=0, kind="keystroke_burst", detail="short intro typed"),
ProctorEvent(timestamp=85, kind="paste_large", detail="3,100 chars pasted in 2s"),
ProctorEvent(timestamp=90, kind="idle", detail="no activity for 4 minutes"),
ProctorEvent(timestamp=330, kind="paste_large", detail="2,800 chars pasted in 2s"),
],
)
PROCTOR_SCENARIOS: dict[str, ProctorScenario] = {
s.scenario_id: s
for s in (
_PROCTOR_SCENARIO_HEALTHY,
_PROCTOR_SCENARIO_DISTRACTED,
_PROCTOR_SCENARIO_FLAGGED,
)
}
def get_proctor_scenario(scenario_id: str) -> ProctorScenario | None:
return PROCTOR_SCENARIOS.get(scenario_id)
def summarize_proctor_scenario(scenario: ProctorScenario) -> str:
"""Render the proctor event timeline as compact text for prompt injection."""
lines = [f"Defense session: {scenario.title} (competency {scenario.competency_id})"]
for event in scenario.events:
lines.append(f"t+{event.timestamp}s {event.kind}: {event.detail}".rstrip(": "))
return "\n".join(lines)
_LAB_SCENARIO_STRONG = LabTelemetryScenario(
scenario_id="lab-scenario-strong",
title="Strong build session — multi-agent research assistant",
competency_id="stack-orchestration-c002",
events=[
TelemetryEvent(timestamp=0, kind="keystroke_burst", detail="planner.py"),
TelemetryEvent(timestamp=95, kind="file_save", detail="planner.py"),
TelemetryEvent(timestamp=120, kind="run_tests", detail="3 tests"),
TelemetryEvent(timestamp=126, kind="test_pass",
detail="3/3 passed"),
TelemetryEvent(timestamp=180, kind="keystroke_burst", detail="tool_node.py"),
TelemetryEvent(timestamp=260, kind="file_save", detail="tool_node.py"),
TelemetryEvent(timestamp=275, kind="run_tests", detail="4 tests"),
TelemetryEvent(timestamp=281, kind="test_pass",
detail="4/4 passed"),
TelemetryEvent(timestamp=340, kind="commit",
detail="add tool node with retries"),
],
)
_LAB_SCENARIO_STRUGGLING = LabTelemetryScenario(
scenario_id="lab-scenario-struggling",
title="Struggling build session — repeated failures, no checkpoints",
competency_id="stack-orchestration-c002",
events=[
TelemetryEvent(timestamp=0, kind="keystroke_burst", detail="main.py"),
TelemetryEvent(timestamp=210, kind="run_tests", detail="2 tests"),
TelemetryEvent(timestamp=215, kind="test_fail",
detail="ImportError: no module named 'tools'"),
TelemetryEvent(timestamp=216, kind="console_error", detail="traceback dumped"),
TelemetryEvent(timestamp=300, kind="keystroke_burst",
detail="main.py"),
TelemetryEvent(timestamp=520, kind="run_tests",
detail="2 tests"),
TelemetryEvent(timestamp=525, kind="test_fail",
detail="ImportError: no module named 'tools'"),
TelemetryEvent(timestamp=526, kind="console_error",
detail="same traceback as before"),
TelemetryEvent(timestamp=600, kind="idle",
detail="no activity for 6 minutes"),
TelemetryEvent(timestamp=960, kind="idle",
detail="no activity for 14 minutes"),
],
)
_LAB_SCENARIO_FLAGGED = LabTelemetryScenario(
scenario_id="lab-scenario-flagged",
title="Flagged build session — large paste, instant pass",
competency_id="stack-orchestration-c003",
events=[
TelemetryEvent(timestamp=0, kind="keystroke_burst", detail="eval.py"),
TelemetryEvent(timestamp=30, kind="paste",
detail="2,400 chars pasted into eval.py"),
TelemetryEvent(timestamp=45, kind="run_tests", detail="6 tests"),
TelemetryEvent(timestamp=47, kind="test_pass", detail="6/6 passed"),
TelemetryEvent(timestamp=48, kind="commit",
detail="finish eval harness"),
],
)
LAB_SCENARIOS: dict[str, LabTelemetryScenario] = {
s.scenario_id: s
for s in (_LAB_SCENARIO_STRONG, _LAB_SCENARIO_STRUGGLING, _LAB_SCENARIO_FLAGGED)
}
DEFAULT_LAB_SCENARIO_ID = "lab-scenario-strong"
def get_lab_scenario(scenario_id: str) -> LabTelemetryScenario | None:
return LAB_SCENARIOS.get(scenario_id)
def summarize_scenario(scenario: LabTelemetryScenario) -> str:
"""Render the event timeline as compact text for prompt injection."""
lines = [f"Session: {scenario.title} (competency {scenario.competency_id})"]
for event in scenario.events:
lines.append(f"t+{event.timestamp}s {event.kind}: {event.detail}".rstrip(": "))
return "\n".join(lines)
@@ -1,230 +0,0 @@
"""Synthetic trace fixtures for grading calibration (Task 3-2-02, REQ-3-004).
Three builder archetypes as REAL `TelemetryEvent` traces (the grading
engine's native input — these are NOT the v0.2 corpus's simplified
`{timestamp, kind, detail}` event shapes):
strong builder iterative debugging: small edits, tests early,
failed runs closed by targeted fixes, eventual pass.
lazy builder one large paste, a single late test run, pass.
(v0.2 alignment: the `lab-scenario-flagged` paste-
and-run archetype; D-021.)
struggling builder many edit/test cycles, failures never close,
never reaches a pass.
ID convention (D-021 alignment, documented in each fixture):
v0.2 corpus scenario IDs are `<domain>-scenario-<slug>` (`lab-scenario-strong`,
`lab-scenario-struggling`, `lab-scenario-flagged`, `proctor-scenario-*` see
corpus/telemetry.py). Grading fixtures carry ids string-aligned to that
convention:
fixture id = "<scenario id>::<archetype>-trace"
learner id = "learner-003" (a member of the corpus learner-00N id space;
learner-001/002 exist in learner_context.py)
so a fixture is greppable against its v0.2 scenario counterpart while staying
a distinct id space (a grading trace is a real event stream, not the v0.2
mock scenario timeline same convention, richer event kind set).
Instance hygiene: fixtures store event SPECS (plain tuples) and materialize
FRESH `TelemetryEvent` instances on every `fixture.events` access. SQLModel
rows carry SQLAlchemy instance state once a session has flushed them
re-adding the SAME instance to another store is a silent no-op, which would
poison sequential test runs (a fixture ingested by test N would vanish for
test N+1). Materializing per access keeps every consumer independent.
These fixtures exist for the CALIBRATION CONTRACT (tests/grading/
test_calibration.py): the mock provider scripts archetype-mapped scores and
the test asserts the ORDERING the rubric must eventually enforce. They are
NOT an LLM quality benchmark see the test module docstring for the honest
scope statement.
"""
from __future__ import annotations
from datetime import UTC, datetime, timedelta
from typing import Final
from pydantic import BaseModel, Field
from ..telemetry.models import TelemetryEvent
_T0: Final = datetime(2026, 9, 12, 0, 0, 0, tzinfo=UTC)
#: Event spec: (seq, kind, payload, seconds-since-session-start).
EventSpec = tuple[int, str, dict, float]
class TraceFixture(BaseModel):
"""One named synthetic trace + its v0.2 scenario alignment (D-021).
`event_specs` is the durable, session-state-free description; `events`
materializes fresh TelemetryEvent rows from it on every access.
"""
model_config = {"frozen": True}
fixture_id: str # "<v0.2 scenario id>::<archetype>-trace"
archetype: str # "strong" | "lazy" | "struggling"
aligned_scenario_id: str # the v0.2 corpus scenario this fixture mirrors
competency_id: str # corpus competency id space (stack-orchestration-c00N)
task_id: str # trace task id (grading operates on (learner, task))
learner_id: str
event_specs: tuple[EventSpec, ...] = Field(default=())
@property
def events(self) -> list[TelemetryEvent]:
"""FRESH TelemetryEvent instances — safe to ingest into any store.
Never cache these: an instance flushed by one SQLite session
carries persistent identity, and re-appending it elsewhere no-ops.
"""
return [
TelemetryEvent(
learner_id=self.learner_id,
task_id=self.task_id,
seq=seq,
kind=kind,
payload=payload,
ts=_T0 + timedelta(seconds=offset_s),
sandbox_id="sbx-calibration",
)
for seq, kind, payload, offset_s in self.event_specs
]
def description(self) -> str:
return (
f"{self.fixture_id} (archetype={self.archetype}, aligned="
f"{self.aligned_scenario_id}, competency={self.competency_id})"
)
class _Builder:
"""Seq-accurate event-spec builder for one (learner, task) pair."""
def __init__(self) -> None:
self.specs: list[EventSpec] = []
self._seq = 0
def add(self, kind: str, payload: dict | None, at: float) -> None:
self.specs.append((self._seq, kind, payload or {}, at))
self._seq += 1
def _strong_builder_specs() -> list[EventSpec]:
"""Iterative debugging: tests early, tight edit→test loops, eventual pass.
Mirrors `lab-scenario-strong` (v0.2: keystrokes file_save run_tests
test_pass, c002) at full telemetry fidelity every failed cycle is
closed by a targeted edit followed by a re-run that passes.
"""
b = _Builder()
b.add("activity", {"state": "starting"}, 0)
b.add("file_diff", {"path": "planner.py", "added": 14}, 95) # small edit
b.add("command", {"cmd": "pytest -q tests/test_planner.py"}, 120) # tests EARLY
b.add("test_result", {"passed": True, "exit_code": 0}, 126)
b.add("file_diff", {"path": "tool_node.py", "added": 22}, 180)
b.add("command", {"cmd": "pytest -q"}, 275)
b.add("test_result", {"passed": False, "exit_code": 1}, 281) # honest failure
b.add("file_diff", {"path": "tool_node.py", "added": 6, "removed": 2}, 340) # targeted fix
b.add("command", {"cmd": "pytest -q"}, 430)
b.add("test_result", {"passed": True, "exit_code": 0}, 436) # cycle CLOSED
b.add("command", {"cmd": "git commit -m 'tool node with retries'"}, 500)
return b.specs
def _lazy_builder_specs() -> list[EventSpec]:
"""Paste-and-run: one large paste, a single LATE test run, instant pass.
Mirrors `lab-scenario-flagged` (v0.2: paste of 2,400 chars run_tests
instant 6/6 pass, c003) at full telemetry fidelity zero iteration, zero
verification during construction, one terminal test run only.
"""
b = _Builder()
b.add("activity", {"state": "starting"}, 0)
b.add("file_diff", {"path": "eval.py", "added": 240, "removed": 0}, 30) # one bulk paste
b.add("file_diff", {"path": "README.md", "added": 12}, 40)
b.add("command", {"cmd": "npm run build"}, 45)
b.add("run_result", {"exit_code": 0, "ok": True}, 60)
b.add("command", {"cmd": "pytest -q"}, 520) # single LATE test run
b.add("test_result", {"passed": True, "exit_code": 0}, 540) # instant pass
return b.specs
def _struggling_builder_specs() -> list[EventSpec]:
"""Many cycles, none close: repeated failures, no eventual pass.
Mirrors `lab-scenario-struggling` (v0.2: repeated identical ImportErrors,
idle gaps, no checkpoint, c002) at full telemetry fidelity edits happen
between failures, but the same failure recurs; no pass is ever reached.
"""
b = _Builder()
b.add("activity", {"state": "starting"}, 0)
b.add("file_diff", {"path": "main.py", "added": 40}, 20)
b.add("command", {"cmd": "pytest -q"}, 210)
b.add("test_result", {"passed": False, "exit_code": 1}, 215) # ImportError
b.add("file_diff", {"path": "main.py", "added": 8, "removed": 3}, 300)
b.add("command", {"cmd": "pytest -q"}, 520)
b.add("test_result", {"passed": False, "exit_code": 1}, 525) # SAME error
b.add("file_diff", {"path": "main.py", "added": 5}, 610)
b.add("command", {"cmd": "pytest -q"}, 960)
b.add("test_result", {"passed": False, "exit_code": 1}, 965) # STILL failing
b.add("activity", {"state": "idle"}, 1500) # long idle
b.add("activity", {"state": "idle"}, 2200)
return b.specs
#: Calibration learner — a member of the corpus learner-00N id space (D-021;
#: learner-001/002 live in corpus/learner_context.py; grading fixtures use
#: a third id so calibration traces never collide with mock-context reads).
CALIBRATION_LEARNER_ID: Final = "learner-003"
STRONG_BUILDER: Final = TraceFixture(
fixture_id="lab-scenario-strong::strong-trace",
archetype="strong",
aligned_scenario_id="lab-scenario-strong",
competency_id="stack-orchestration-c002",
task_id="task-calibration-strong",
learner_id=CALIBRATION_LEARNER_ID,
event_specs=tuple(_strong_builder_specs()),
)
LAZY_BUILDER: Final = TraceFixture(
# v0.2's paste-and-run archetype is the "flagged" lab scenario (D-021):
# large paste → instant test pass. "lazy builder" is that behavior
# without the proctor flag; the alignment is behavioral, documented here.
fixture_id="lab-scenario-flagged::lazy-trace",
archetype="lazy",
aligned_scenario_id="lab-scenario-flagged",
competency_id="stack-orchestration-c003",
task_id="task-calibration-lazy",
learner_id=CALIBRATION_LEARNER_ID,
event_specs=tuple(_lazy_builder_specs()),
)
STRUGGLING_BUILDER: Final = TraceFixture(
fixture_id="lab-scenario-struggling::struggling-trace",
archetype="struggling",
aligned_scenario_id="lab-scenario-struggling",
competency_id="stack-orchestration-c002",
task_id="task-calibration-struggling",
learner_id=CALIBRATION_LEARNER_ID,
event_specs=tuple(_struggling_builder_specs()),
)
TRACE_FIXTURES: Final[dict[str, TraceFixture]] = {
f.fixture_id: f
for f in (STRONG_BUILDER, LAZY_BUILDER, STRUGGLING_BUILDER)
}
def get_trace_fixture(fixture_id: str) -> TraceFixture | None:
return TRACE_FIXTURES.get(fixture_id)
def digest_of(fixture: TraceFixture):
"""Compute the digest for a fixture (pure compute; test-side helper)."""
from ..grading.features import compute_digest
return compute_digest(fixture.events)
@@ -1,27 +0,0 @@
"""Process-trace grading — trace digest, rubric engine, GradeStore (REQ-3-004).
Boundary rule (D-027): grading/ is an engine module it never imports
api/; the single sanctioned agents/ dependency is the shared D-020
structured defense (agents/structured.py), imported module-direct in
engine.py (see its docstring for why). features.py imports telemetry/;
store.py imports config only; engine.py composes llm/ + telemetry/ +
prompts/ + agents.structured.
Wave status: features.py (TraceDigest, compute_digest) + store.py
(GradeRecord, GradeStore, SQLiteGradeStore) landed in Wave 1 (3-1-01 +
3-1-02); engine.py (GradingEngine, RubricScore) is Wave 2 (3-2-01).
"""
from .engine import GradingEngine, RubricScore
from .features import TraceDigest, compute_digest
from .store import GradeRecord, GradeStore, SQLiteGradeStore
__all__ = [
"GradeRecord",
"GradeStore",
"GradingEngine",
"RubricScore",
"SQLiteGradeStore",
"TraceDigest",
"compute_digest",
]
@@ -1,313 +0,0 @@
"""GradingEngine — rubric scoring over real process traces (REQ-3-004).
The Wave-2 composition of the grading stack:
trace completeness gate (G-4, FIRST nothing is sent to any LLM
when the gate trips) TraceStore.get_trace compute_digest (D-028)
prompts.grading.render_trace_digest D-020 4-layer structured
defense (agents/structured.py REUSED, composed, never duplicated)
RubricScore validation GradeStore persistence GradeRecord.
Gate verdicts are FIRST-CLASS RESULTS, not exceptions (G-4 is binding):
UNGRADABLE_TRACE_INCOMPLETE seq gaps in the store OR the trace is
flagged by TraceIntegrityMap (INCOMPLETE_FLOODED). The gap list /
flag reason is surfaced in `scores` for API rendering, and the
record is PERSISTED like any grade so a learner sees why no
credential can be issued for this trace the gate outcome is
durable and auditable, not a transient error string.
UNGRADABLE_EMPTY_TRACE no events stored for the pair.
On either verdict `scores.criteria` is empty and the LLM is never called.
DI (D-027/D-032 house pattern): the engine receives trace_store,
grade_store, integrity and provider through the constructor and knows
NOTHING of FastAPI api/ composes it (Task 3-3-01). `model` is injected
alongside the provider so tests script the mock against the production
wiring without touching Settings.
Boundary (D-027): grading/ imports llm/ (provider protocol + Message
type), telemetry/ (store + integrity map), prompts/ (rubric text) and
ONLY agents.structured the sanctioned shared D-020 defense. We import
the MODULE directly (`from ..agents.structured import structured_completion`)
rather than the `agents` package, mirroring how agents/base.py consumes
it (same direct-module import): that keeps the dependency surface to
exactly the two names the engine needs (structured_completion,
StructuredOutputError) and avoids executing agents/__init__ re-exports
(BaseAgent, registry, session store) that grading has no business
loading a side-effect-hygiene choice that keeps this import line
grep-auditable as "the one agents dependency". grading/ never imports api/.
RubricScore placement (documented decision): the validated output model
lives HERE, not in prompts/. The pydantic model is the engine's return
CONTRACT (the shape GradeStore.scores must hold), while prompts/grading.py
is pure prompt text + its mirror schema HINT string the same split as
agents/assessor.py (model + hint) but with the model owned by the engine
module that validates it. Prompt files hold text, per the prompts/ house
style; engine files hold typed contracts.
"""
import logging
from datetime import UTC, datetime
from typing import TYPE_CHECKING, Final
from pydantic import BaseModel, ConfigDict, Field, field_validator
from ..agents.structured import StructuredOutputError, structured_completion
from ..llm.base import LLMProvider
from ..llm.types import Message
from ..prompts.grading import (
RUBRIC_CRITERIA,
RUBRIC_SCORE_SCHEMA_HINT,
SYSTEM_PROMPT,
render_trace_digest,
)
from ..telemetry.ingest import TraceIntegrityMap
from ..telemetry.store import TraceStore
from .features import compute_digest
from .store import GradeRecord, GradeStore
if TYPE_CHECKING: # pragma: no cover - protocol-only import for the optional
from ..variants.store import VariantStore # noqa: TC001 (variant-blind without it)
logger = logging.getLogger(__name__)
#: First-class gate verdicts (G-4). GradeRecord.verdict values; non-empty
#: by store contract. Rubric verdicts (mastered/developing/not_yet) ride in
#: `scores.verdict` — `record.verdict` stays the machine-readable outcome.
VERDICT_UNGRADABLE_INCOMPLETE: Final = "UNGRADABLE_TRACE_INCOMPLETE"
VERDICT_UNGRADABLE_EMPTY: Final = "UNGRADABLE_EMPTY_TRACE"
#: record.verdict for a successfully LLM-graded trace (the rubric verdict
#: travels inside scores); keeps verdict non-empty for every persisted row.
VERDICT_GRADED: Final = "GRADED"
#: Gate-detail keys surfaced in GradeRecord.scores (tests assert on these).
_INTEGRITY_FLAG_KEY: Final = "integrity_flag"
def _anchors_context(variant) -> str: # noqa: ANN001 - VariantRecord (duck-typed)
"""Render the variant template's difficulty anchors for the grader prompt.
Contains only the template id + anchor numbers no learner-identifying
material (D-028 anonymity preserved; the digest-leak tests keep holding).
Lazy template import: grading must not import variants/ at module load
(variants/prompts import-cycle safety mirrors llm/ rules).
"""
from ..variants.templates import get_template
template = get_template(variant.template_id)
if template is None:
return f"template={variant.template_id} (anchors unavailable)"
a = template.rubric_anchors
return (
f"template={template.id}; "
f"expected_edit_count_band={list(a.expected_edit_count_band)}; "
f"expected_min_test_runs={a.expected_min_test_runs}; "
f"expected_error_fix_cycles_band={list(a.expected_error_fix_cycles_band)}"
)
_MISSING_SEQS_KEY: Final = "missing_seqs"
class RubricScore(BaseModel):
"""Validated LLM output: per-criterion 0-4 scores + strengths + gaps + verdict.
The D-020 schema for the grader: `structured_completion` parses the
model reply into THIS shape (layer 3), retrying once with the
validation error fed back (layer 4). Exact criteria set + 0-4 ranges
are enforced here, so the scores dict persisted to GradeStore is
always rubric-shaped no matter what the model produced.
"""
model_config = ConfigDict(extra="forbid")
criteria: dict[str, int]
strengths: list[str] = Field(min_length=1, max_length=2)
gaps: list[str] = Field(min_length=1, max_length=2)
verdict: str # "mastered" | "developing" | "not_yet"
@field_validator("criteria")
@classmethod
def _criteria_rubric_shaped(cls, value: dict[str, int]) -> dict[str, int]:
"""Exact criteria keys (no extras, no omissions) and 0-4 scores."""
expected = set(RUBRIC_CRITERIA)
got = set(value)
if got != expected:
raise ValueError(
f"criteria keys must be exactly {sorted(expected)}, got {sorted(got)}"
)
for key, score in value.items():
if not 0 <= score <= 4:
raise ValueError(f"criterion {key!r} must be within 0-4, got {score}")
return value
@field_validator("verdict")
@classmethod
def _verdict_known(cls, value: str) -> str:
allowed = {"mastered", "developing", "not_yet"}
if value not in allowed:
raise ValueError(f"verdict must be one of {sorted(allowed)}, got {value!r}")
return value
class GradingEngine:
"""Scores a (learner_id, task_id) trace into a persisted GradeRecord."""
def __init__(
self,
trace_store: TraceStore,
grade_store: GradeStore,
integrity: TraceIntegrityMap,
provider: LLMProvider,
*,
model: str = "gemma4:31b",
variant_store: "VariantStore | None" = None,
) -> None:
self._trace_store = trace_store
self._grade_store = grade_store
self._integrity = integrity
self._provider = provider
self._model = model
# Phase 4 (MH#4): optional variant lookup — when the graded task
# derives from a generated variant, its template's difficulty anchors
# ship to the grader prompt (same bar for every variant of the
# template, a-5) and the variant seed is stamped on the record.
# Optional so engine tests stay decoupled; main.py lifespan wires it.
self._variant_store = variant_store
async def grade(self, learner_id: str, task_id: str) -> GradeRecord:
"""Grade one trace; persist latest-state (GradeStore upserts); return it.
Gate FIRST (G-4): the LLM is only ever reached from the fully
guarded path no gate state can be masked by an LLM error.
"""
record = await self._grade(learner_id, task_id)
self._grade_store.save(record)
return record
# ------------------------------------------------------------------ core
async def _grade(self, learner_id: str, task_id: str) -> GradeRecord:
# -- G-4 gate FIRST: integrity flag OR seq gaps. Ordering matters:
# gaps() returns [] for an EMPTY trace, so the empty check below is
# reachable only when no rows exist at all; a gapped or flooded
# trace can never fall through to the LLM path.
if self._integrity.is_incomplete(learner_id, task_id):
reason = self._integrity.reason(learner_id, task_id) or "unknown"
gaps = self._trace_store.gaps(learner_id, task_id)
logger.info(
"grade gate (G-4): %s/%s integrity-flagged (%s) — ungradable",
learner_id,
task_id,
reason,
)
return self._ungradable(
learner_id,
task_id,
detail={_INTEGRITY_FLAG_KEY: reason, _MISSING_SEQS_KEY: gaps},
verdict=VERDICT_UNGRADABLE_INCOMPLETE,
)
gaps = self._trace_store.gaps(learner_id, task_id)
if gaps:
logger.info(
"grade gate (G-4): %s/%s seq gaps %s — ungradable",
learner_id,
task_id,
gaps,
)
return self._ungradable(
learner_id,
task_id,
detail={_INTEGRITY_FLAG_KEY: None, _MISSING_SEQS_KEY: gaps},
verdict=VERDICT_UNGRADABLE_INCOMPLETE,
)
trace = self._trace_store.get_trace(learner_id, task_id)
if not trace:
logger.info(
"grade gate: %s/%s empty trace — ungradable", learner_id, task_id
)
return self._ungradable(
learner_id,
task_id,
detail={_INTEGRITY_FLAG_KEY: None, _MISSING_SEQS_KEY: []},
verdict=VERDICT_UNGRADABLE_EMPTY,
)
# -- Guarded path: digest (D-028) → prompt → D-020 4-layer defense.
digest = compute_digest(trace)
variant = self._lookup_variant(task_id)
anchors_context = (
_anchors_context(variant) if variant is not None else None
)
messages = [
Message(role="system", content=SYSTEM_PROMPT),
Message(
role="user",
content=render_trace_digest(digest, anchors_context=anchors_context),
),
]
try:
rubric = await structured_completion(
self._provider,
messages,
model=self._model,
schema=RubricScore,
schema_hint=RUBRIC_SCORE_SCHEMA_HINT,
)
except StructuredOutputError as exc:
# The trace was gradable but the model failed to produce valid
# JSON within the D-020 budget (two attempts). Raise — the API
# layer maps this to a 502 (assessor precedent). Persisting a
# fabricated or partial grade here would violate the no-silent-
# fallback rule: no credential-worthy record without a validated
# RubricScore.
raise StructuredOutputError(f"grading LLM failed for {task_id}: {exc}") from exc
logger.debug(
"graded %s/%s: %s (model=%s)",
learner_id,
task_id,
rubric.verdict,
self._model,
)
return GradeRecord(
learner_id=learner_id,
task_id=task_id,
variant_seed=variant.seed if variant is not None else None, # D-029
digest=digest.model_dump(),
scores=rubric.model_dump(),
verdict=VERDICT_GRADED,
model=self._model,
created_at=datetime.now(tz=UTC),
)
def _lookup_variant(self, task_id: str): # noqa: ANN202 - VariantRecord | None
"""MH#4: resolve the graded task's variant (None when not variant-derived)."""
if self._variant_store is None:
return None
return self._variant_store.get_by_task(task_id)
# ------------------------------------------------------------ gate record
@staticmethod
def _ungradable(
learner_id: str,
task_id: str,
*,
detail: dict,
verdict: str,
) -> GradeRecord:
"""Build a gate record: no digest (nothing was graded), gate detail
surfaced in `scores` (the store allows an empty scores dict, but
G-4 requires the gap list / flag reason surfaced the detail IS the
verdict's payload), model="none" (no LLM was involved; provenance
stays honest).
"""
return GradeRecord(
learner_id=learner_id,
task_id=task_id,
variant_seed=None,
digest={},
scores=detail,
verdict=verdict,
model="none",
created_at=datetime.now(tz=UTC),
)
@@ -1,246 +0,0 @@
"""Deterministic process-trace digest (D-028, REQ-3-004).
Pure compute no LLM, no I/O. `compute_digest` reduces an ordered
TelemetryEvent trace to a compact, bounded `TraceDigest` that is safe to
embed in a grading prompt:
- FIXED fields + small histograms only; NO raw commands, NO file contents,
NO payloads the raw trace NEVER reaches the LLM (D-028), which also
bounds the prompt-injection surface.
- Tolerant to both live trace mixes: daemon-topology traces carry
`activity` + `file_diff` kinds (workspace watcher), while REPL-driven
traces carry `command`/`stdin`/`stdout`/`run_result`/`test_result`
(P2 verification P1). Features derive from whatever kinds are present and
never crash on absent kinds.
Feature semantics (conservative, deterministic):
- test pass/fail counts + final status derive from `test_result` payloads
when present, falling back to `run_result` exit codes (0 = pass).
- an error/fix CYCLE = a failing run/test followed by >= 1 edit and then a
later run/test (pass or fail) the next observed result closes the cycle.
- idle gaps = wall-clock gaps between consecutive events exceeding
`idle_threshold_s` (default 120s): count + total seconds.
- command category histogram classifies `command`-kind payloads: build /
test / file / nav / debug / other.
"""
from __future__ import annotations
from collections import Counter
from pydantic import BaseModel, Field
from ..telemetry.models import TelemetryEvent
_IDLE_DEFAULT_S: float = 120.0
_TEST_HINTS = ("test", "pytest", "vitest", "jest", "mocha", "unittest", "go test", "npm test")
_BUILD_HINTS = ("make", "npm run build", "pip install", "pnpm", "cargo build", "gcc", "tsc")
_DEBUG_HINTS = ("gdb", "pdb", "print(", "debug", "strace", "ltrace", "curl", "ping")
_NAV_HINTS = ("ls", "cd", "pwd", "cat ", "grep ", "find", "rg ", "tree", "head", "tail", "less")
_FILE_HINTS = ("mv ", "cp ", "rm ", "mkdir", "touch", "chmod", "nano", "vim", "sed -i", "tee ")
class TraceDigest(BaseModel):
"""Compact, bounded, LLM-safe summary of a process trace (D-028).
Fixed fields + small histograms. Serializes well under 4 KB; contains no
raw commands, file contents, or event payloads.
"""
model_config = {"frozen": True}
event_count: int = Field(ge=0)
session_duration_s: float = Field(ge=0.0)
edit_count: int = Field(ge=0)
command_count: int = Field(ge=0)
run_count: int = Field(ge=0)
test_pass_count: int = Field(ge=0)
test_fail_count: int = Field(ge=0)
final_test_status: str = Field(pattern="^(pass|fail|none)$")
first_test_pass_offset_s: float | None = None
error_fix_cycles: int = Field(ge=0)
mean_fix_latency_s: float | None = None
idle_gap_count: int = Field(ge=0)
idle_gap_total_s: float = Field(ge=0.0)
command_categories: dict[str, int] = Field(default_factory=dict)
kind_histogram: dict[str, int] = Field(default_factory=dict)
def _event_pass_status(event: TelemetryEvent) -> bool | None:
"""True (pass) / False (fail) / None (not a result event) for one event."""
payload = event.payload or {}
if event.kind == "test_result":
if "passed" in payload:
return bool(payload["passed"])
if "exit_code" in payload:
return int(payload["exit_code"]) == 0
if "status" in payload:
return str(payload["status"]).lower() in ("pass", "passed", "ok", "success")
return None
if event.kind == "run_result":
if "exit_code" in payload:
return int(payload["exit_code"]) == 0
if "ok" in payload:
return bool(payload["ok"])
return None
return None
def _classify_command(text: str) -> str:
lowered = text.lower()
if any(h in lowered for h in _TEST_HINTS):
return "test"
if any(h in lowered for h in _BUILD_HINTS):
return "build"
if any(h in lowered for h in _DEBUG_HINTS):
return "debug"
if any(h in lowered for h in _NAV_HINTS):
return "nav"
if any(h in lowered for h in _FILE_HINTS):
return "file"
return "other"
def _command_text(event: TelemetryEvent) -> str:
payload = event.payload or {}
return str(payload.get("cmd") or payload.get("command") or payload.get("line") or "")
def compute_digest(
trace: list[TelemetryEvent], *, idle_threshold_s: float = _IDLE_DEFAULT_S
) -> TraceDigest:
"""Reduce an ordered trace to a bounded digest. Never raises on odd input."""
events = sorted(trace, key=lambda e: (e.seq, e.ts))
if not events:
return TraceDigest(
event_count=0,
session_duration_s=0.0,
edit_count=0,
command_count=0,
run_count=0,
test_pass_count=0,
test_fail_count=0,
final_test_status="none",
first_test_pass_offset_s=None,
error_fix_cycles=0,
mean_fix_latency_s=None,
idle_gap_count=0,
idle_gap_total_s=0.0,
command_categories={},
kind_histogram={},
)
kind_histogram = Counter(e.kind for e in events)
start_ts = events[0].ts
end_ts = events[-1].ts
duration = max(0.0, (end_ts - start_ts).total_seconds())
edit_count = kind_histogram.get("file_diff", 0)
command_count = kind_histogram.get("command", 0)
run_count = kind_histogram.get("run_result", 0)
# Tests: prefer test_result events; fall back to run_result exit codes.
test_statuses: list[tuple[TelemetryEvent, bool]] = []
for e in events:
if e.kind == "test_result":
ok = _event_pass_status(e)
if ok is not None:
test_statuses.append((e, ok))
if not test_statuses:
for e in events:
if e.kind == "run_result":
ok = _event_pass_status(e)
if ok is not None:
test_statuses.append((e, ok))
test_pass_count = sum(1 for _, ok in test_statuses if ok)
test_fail_count = len(test_statuses) - test_pass_count
if not test_statuses:
final_test_status = "none"
else:
final_test_status = "pass" if test_statuses[-1][1] else "fail"
first_pass = next((e for e, ok in test_statuses if ok), None)
first_pass_offset = (
max(0.0, (first_pass.ts - start_ts).total_seconds()) if first_pass is not None else None
)
# Error/fix cycles: a failing result starts a pending cycle; the NEXT
# observed result closes it (regardless of outcome) — a fix attempt that
# fails again is itself another iteration of debugging, so it closes the
# previous cycle and opens a new one. Edits since the fail mark the
# close as a genuine fix attempt; latency = first edit -> closing result.
cycles = 0
fix_latencies: list[float] = []
pending_fail_ts: float | None = None # seconds since start
edits_since_fail = 0
first_edit_ts: float | None = None
for e in events:
t = max(0.0, (e.ts - start_ts).total_seconds())
if e.kind == "file_diff":
if pending_fail_ts is not None:
if edits_since_fail == 0:
first_edit_ts = t
edits_since_fail += 1
continue
ok = _event_pass_status(e)
if ok is None:
continue
if ok is False:
if pending_fail_ts is not None and edits_since_fail > 0 and first_edit_ts is not None:
# failed fix attempt: closes the previous cycle, opens a new one
cycles += 1
fix_latencies.append(t - first_edit_ts)
pending_fail_ts = t
edits_since_fail = 0
first_edit_ts = None
continue
if ok is True and pending_fail_ts is not None:
if edits_since_fail > 0 and first_edit_ts is not None:
cycles += 1
fix_latencies.append(t - first_edit_ts)
pending_fail_ts = None
edits_since_fail = 0
first_edit_ts = None
mean_fix_latency = (
sum(fix_latencies) / len(fix_latencies) if fix_latencies else None
)
# Idle gaps between consecutive events.
idle_gap_count = 0
idle_gap_total = 0.0
prev_ts = None
for e in events:
if prev_ts is not None:
gap = (e.ts - prev_ts).total_seconds()
if gap > idle_threshold_s:
idle_gap_count += 1
idle_gap_total += gap
prev_ts = e.ts
# Command category histogram (command-kind events only).
categories: Counter[str] = Counter()
for e in events:
if e.kind == "command":
categories[_classify_command(_command_text(e))] += 1
return TraceDigest(
event_count=len(events),
session_duration_s=round(duration, 3),
edit_count=edit_count,
command_count=command_count,
run_count=run_count,
test_pass_count=test_pass_count,
test_fail_count=test_fail_count,
final_test_status=final_test_status,
first_test_pass_offset_s=(
round(first_pass_offset, 3) if first_pass_offset is not None else None
),
error_fix_cycles=cycles,
mean_fix_latency_s=(round(mean_fix_latency, 3) if mean_fix_latency is not None else None),
idle_gap_count=idle_gap_count,
idle_gap_total_s=round(idle_gap_total, 3),
command_categories=dict(sorted(categories.items())),
kind_histogram=dict(sorted(kind_histogram.items())),
)
-237
View File
@@ -1,237 +0,0 @@
"""GradeStore — grade persistence protocol + SQLite implementation (REQ-3-004, D-027).
Postgres-migration-ready (D-027): the protocol is the only surface the
grading engine and API layers touch; swapping SQLiteGradeStore for a
Postgres-backed implementation must not change call sites. The
`grade_record` table uses only portable column types (str / JSON /
datetime), so the same SQLModel schema stands up unchanged on Postgres.
Upsert, NOT append: (learner_id, task_id) is the grade identity one row
per learner per task holding the LATEST grade. `save` overwrites the whole
row when the pair already exists, so a regrade replaces scores, verdict,
created_at, digest, model and variant_seed wholesale. That is deliberately
the opposite of TraceStore.append's dedup-keep-first contract: a trace is an
append-only event log, a grade is latest-state, so the engine can re-grade
a task idempotently as its rubric or input evolves.
Concurrency (a-3): the engine enables WAL + synchronous=NORMAL and a busy
timeout at connection time, so a regrade writer and API readers do not hit
`database is locked` on the single-box pilot.
`created_at` contract: callers stamp UTC (datetime.now(UTC)); SQLite stores
it naive and the read paths re-label it tz-aware UTC (same boundary
normalization as TelemetryEvent.ts, so the contract holds on any backend).
Boundary (D-027): `grading/` never imports `agents/` / `api/`; this module
imports config only.
"""
import logging
import sqlite3
from collections.abc import Iterator
from contextlib import contextmanager
from datetime import UTC, datetime
from pathlib import Path
from typing import Any, Protocol
import sqlalchemy as sa
from sqlalchemy import JSON, Index
from sqlalchemy.orm import validates
from sqlmodel import Field, Session, SQLModel, create_engine, select
from ..config import Settings
logger = logging.getLogger(__name__)
class GradeRecord(SQLModel, table=True):
"""A persisted grade; (learner_id, task_id) is the PK — latest wins.
Written by the grading engine (one save per grade attempt), read by the
API layer through the GradeStore protocol. Constraint enforcement
mirrors TelemetryEvent: sqlmodel 0.0.42's metaclass drops pydantic
constraints on table models, so SQLAlchemy `@validates` hooks enforce
instead and the column types stay Postgres-ready (D-027).
Field contract:
learner_id non-empty learner identifier (same id space as traces).
task_id non-empty task identifier; grade identity is the
(learner_id, task_id) pair the same pair as trace
identity, so a grade is keyed by the exact trace it
was computed from.
variant_seed task-variant seed (D-029); None when the graded task
is not variant-derived. Since Phase 4 the engine
stamps the graded variant's seed here (MH#4) and the
template's difficulty anchors ship to the grader
prompt this column is the audit join for that.
digest compact deterministic trace digest (D-028) that fed
the rubric prompt; persisted for auditability so the
LLM's input stays reproducible.
scores validated rubric scores (per-criterion 0-4,
strengths, gaps); JSON dict. An empty dict is legal
(e.g. an UNGRADABLE_TRACE_INCOMPLETE record carries a
verdict but no scores).
verdict first-class verdict string (rubric verdict or
UNGRADABLE_TRACE_INCOMPLETE); non-empty.
model provider model that produced the scores (provenance).
created_at UTC grade timestamp; a regrade replaces it (latest
save wins).
"""
__tablename__ = "grade_record"
# The composite PK covers (learner_id, task_id) point lookups; this
# secondary index covers list_for_learner ordered by created_at without
# a sort step (Postgres migration target D-027).
__table_args__ = (
Index("ix_grade_record_learner_created", "learner_id", "created_at"),
)
learner_id: str = Field(primary_key=True)
task_id: str = Field(primary_key=True)
# None only for non-variant tasks (MH#4 stamps variant seeds since P4).
variant_seed: str | None = Field(default=None)
# JSON columns: stored as TEXT on SQLite, native JSONB on Postgres (D-027).
digest: dict[str, Any] = Field(default_factory=dict, sa_type=JSON)
scores: dict[str, Any] = Field(default_factory=dict, sa_type=JSON)
verdict: str
model: str
created_at: datetime
@validates("learner_id", "task_id")
def _ids_non_empty(self, key: str, value: str) -> str:
if not value:
raise ValueError(f"{key} must be a non-empty identifier")
return value
@validates("verdict")
def _verdict_non_empty(self, key: str, value: str) -> str:
if not value:
raise ValueError(f"{key} must be a non-empty verdict string")
return value
class GradeStore(Protocol):
"""Persistence contract for latest-state grades per (learner_id, task_id).
Implemented by SQLiteGradeStore (v0.3, D-027); a Postgres implementation
must satisfy the same surface.
"""
def save(self, grade: GradeRecord) -> None:
"""Persist a grade. UPSERT on (learner_id, task_id): a regrade with
the same pair REPLACES the stored row wholesale the latest grade
wins. NOT append-only; contrast TraceStore.append, which is
dedup-keep-first for at-least-once ingest.
"""
...
def get(self, learner_id: str, task_id: str) -> GradeRecord | None:
"""Latest stored grade for the pair; None when none exists.
Detached from any DB session safe to pass across layers.
"""
...
def list_for_learner(self, learner_id: str) -> list[GradeRecord]:
"""All stored grades for the learner, ordered by created_at
ascending (chronological). Empty list when the learner has none.
"""
...
def close(self) -> None:
"""Release DB connections. Store must not be used after close."""
...
def _sqlite_connect(dbapi_connection: sqlite3.Connection, _: object) -> None:
"""Per-connection pragma setup (a-3). Mirrors telemetry/store.py.
journal_mode=WAL readers never block the single writer.
synchronous=NORMAL safe in WAL mode, avoids full fsync-per-commit.
busy_timeout=5000 retry briefly under contention instead of
`OperationalError: database is locked`.
"""
cursor = dbapi_connection.cursor()
cursor.execute("PRAGMA journal_mode=WAL")
cursor.execute("PRAGMA synchronous=NORMAL")
cursor.execute("PRAGMA busy_timeout=5000")
cursor.close()
def _as_utc(ts: datetime) -> datetime:
"""Timestamps travel as naive datetime on SQLite, tz-aware elsewhere.
SQLite (via SQLModel) drops tzinfo; Postgres TIMESTAMP WITH TIME ZONE
keeps it. Normalizing on the read path makes the store's contract
tz-aware UTC regardless of the backend (D-027).
"""
if ts.tzinfo is None:
return ts.replace(tzinfo=UTC) # naive UTC read-side: label as UTC
return ts.astimezone(UTC)
class SQLiteGradeStore:
"""SQLite-backed GradeStore (SQLModel). Second protocol-wrapped store
of the D-027 family (first: SQLiteTraceStore).
"""
def __init__(self, db_path: Path | None = None) -> None:
self._db_path: Path = db_path if db_path is not None else Settings().db_path
self._engine = create_engine(f"sqlite:///{self._db_path}")
sa.event.listen(self._engine, "connect", _sqlite_connect)
SQLModel.metadata.create_all(self._engine)
@contextmanager
def _session(self) -> Iterator[Session]:
# expire_on_commit=False: identical session behavior to
# SQLiteTraceStore. save() discards the merged instance and the read
# paths never commit, but a uniform flag across the D-027 stores
# keeps their detachment guarantees from diverging.
with Session(self._engine, expire_on_commit=False) as session:
yield session
def save(self, grade: GradeRecord) -> None:
# `merge` = SELECT-by-PK then UPDATE or INSERT — exactly the upsert
# contract. The trace store deliberately avoids merge (its append is
# dedup-keep-first); here latest-wins IS the contract, so merge is
# the right tool. The caller's object is never attached to the
# session and stays usable (unexpired) after save.
with self._session() as session:
session.merge(grade)
session.commit()
logger.debug(
"grade saved (regrade overwrites): %s/%s verdict=%s model=%s",
grade.learner_id,
grade.task_id,
grade.verdict,
grade.model,
)
def get(self, learner_id: str, task_id: str) -> GradeRecord | None:
with self._session() as session:
record = session.get(GradeRecord, (learner_id, task_id))
if record is None:
return None
record.created_at = _as_utc(record.created_at)
# Detach from the session: callers must not depend on
# open-session ORM magic (lazy loads fail once it closes).
session.expunge(record)
return record
def list_for_learner(self, learner_id: str) -> list[GradeRecord]:
with self._session() as session:
stmt = (
select(GradeRecord)
.where(GradeRecord.learner_id == learner_id)
# Chronological; task_id is a deterministic tie-break for
# grades stamped within the same instant.
.order_by(GradeRecord.created_at, GradeRecord.task_id)
)
results = session.exec(stmt).all()
for row in results:
row.created_at = _as_utc(row.created_at)
session.expunge(row)
return list(results)
def close(self) -> None:
self._engine.dispose()
+2 -1
View File
@@ -4,9 +4,10 @@ from .base import LLMProvider
from .factory import create_provider
from .mock import MockProvider
from .openai_compat import OpenAICompatProvider
from .types import Message
from .types import ChatDelta, Message
__all__ = [
"ChatDelta",
"LLMProvider",
"Message",
"MockProvider",
+32 -6
View File
@@ -1,11 +1,6 @@
"""LLM layer types — messages.
"""LLM layer types — messages and OpenAI-compatible stream chunks.
Boundary rule: nothing in llm/ imports from agents/ or api/.
Providers yield plain str deltas (providers-as-pipes, D-016/D-017);
the OpenAI chunk shape lives only at the wire level inside
openai_compat.py. ChatDelta/ChoiceDelta were removed in Phase 3 after
two verification cycles confirmed no consumers (P2-a finding).
"""
from pydantic import BaseModel
@@ -14,3 +9,34 @@ from pydantic import BaseModel
class Message(BaseModel):
role: str
content: str
class ChoiceDelta(BaseModel):
index: int = 0
delta_content: str = ""
finish_reason: str | None = None
class ChatDelta(BaseModel):
"""Mirrors the OpenAI-compatible streaming chunk shape (D-016 passthrough)."""
id: str = ""
object: str = "chat.completion.chunk"
created: int = 0
model: str = ""
index: int = 0
delta_content: str = ""
finish_reason: str | None = None
def payload(self) -> dict:
"""OpenAI wire shape — the API layer emits this verbatim."""
return {
"id": self.id,
"object": self.object,
"created": self.created,
"model": self.model,
"choices": [
{"index": self.index, "delta": {"content": self.delta_content},
"finish_reason": self.finish_reason}
],
}
+9 -164
View File
@@ -1,44 +1,16 @@
"""FastAPI app factory — lifespan, CORS, health, routers."""
import asyncio
import contextlib
import logging
from contextlib import asynccontextmanager
import httpx
from fastapi import FastAPI
from fastapi.middleware.cors import CORSMiddleware
from .agents.registry import AgentRegistry, register_builtin_agents
from .agents.registry import AgentRegistry
from .agents.session import InMemorySessionStore
from .api import (
assessment_router,
chat_router,
defense_router,
lab_router,
mentor_router,
proctor_router,
sandboxes_router,
telemetry_router,
variants_router,
)
from .api import chat_router
from .config import Settings
from .grading.engine import GradingEngine
from .grading.store import SQLiteGradeStore
from .llm import create_provider
from .sandbox import SandboxManager, UnshareBackend
from .telemetry.ingest import TraceIntegrityMap
from .telemetry.store import SQLiteTraceStore
from .variants.generator import VariantGenerator
from .variants.store import SQLiteVariantStore
from .voice.defense_store import SQLiteDefenseStore
from .voice.factory import voice_provider_from_settings
logger = logging.getLogger(__name__)
#: Interval between wall-clock/G-2 reaper passes (the manager owns the pass;
#: the lifespan owns the loop). 60s against a 900s default timeout → ≤6.7% lag.
REAPER_INTERVAL_S = 60.0
def create_app(settings: Settings | None = None) -> FastAPI:
@@ -50,138 +22,19 @@ def create_app(settings: Settings | None = None) -> FastAPI:
timeout = httpx.Timeout(connect=10.0, read=300.0, write=30.0, pool=10.0)
app.state.http_client = httpx.AsyncClient(timeout=timeout)
app.state.settings = settings
# State-injection override (same pattern as the stores): tests may
# pre-set app.state.provider with a scripted mock; only construct the
# configured provider when none is present.
if getattr(app.state, "provider", None) is None:
app.state.provider = create_provider(settings, app.state.http_client)
app.state.provider = create_provider(settings, app.state.http_client)
app.state.session_store = InMemorySessionStore()
app.state.agent_registry = AgentRegistry()
register_builtin_agents(app.state.agent_registry)
yield
await app.state.http_client.aclose()
# v0.3 sandbox fabric (REQ-3-001): singleton manager, DI'd via
# app.state. Tests may pre-set app.state.sandbox_manager (dependency
# override by state injection) to swap the backend; the lifespan then
# adopts it instead of constructing the real UnshareBackend one.
manager = getattr(app.state, "sandbox_manager", None)
if manager is None:
manager = SandboxManager(backend=UnshareBackend(), settings=settings)
app.state.sandbox_manager = manager
await manager.start() # a-1: reap on-disk orphans from a previous process
app = FastAPI(title="Nextcraft AI Service", version="0.2.0", lifespan=lifespan)
# Telemetry persistence (REQ-3-003, D-027): TraceStore wired through
# app.state. Tests may pre-set app.state.trace_store (state-injection
# override, same pattern as sandbox_manager) — the lifespan adopts it.
trace_store = getattr(app.state, "trace_store", None)
if trace_store is None:
settings.db_path.parent.mkdir(parents=True, exist_ok=True)
trace_store = SQLiteTraceStore(db_path=settings.db_path)
app.state.trace_store = trace_store
# Trace-integrity flags (G-3 INCOMPLETE_FLOODED): process-local map is
# intentional (D-019 registry precedent); the lifespan owns it so the
# grader and the ingest endpoint share one instance.
if getattr(app.state, "trace_integrity", None) is None:
app.state.trace_integrity = TraceIntegrityMap()
# Variant generation (REQ-3-005): VariantStore from the same
# SQLite file as traces/grades (D-027), one VariantGenerator singleton
# wired through app.state — the generator receives store + provider
# via constructor DI and knows nothing of FastAPI (api/ composes it,
# same pattern as GradingEngine). Tests may pre-set
# app.state.variant_store / app.state.variant_generator (the same
# state-injection override); the lifespan adopts a pre-set store but
# NEVER rebuilds a pre-set generator (its provider binding is part
# of the test fixture).
# ORDER NOTE: built BEFORE the grading engine — the engine takes the
# variant store (Phase 4 MH#4: variant anchors ship to the grader
# prompt; variant_seed stamped on graded records).
variant_store = getattr(app.state, "variant_store", None)
if variant_store is None:
settings.db_path.parent.mkdir(parents=True, exist_ok=True)
variant_store = SQLiteVariantStore(db_path=settings.db_path)
app.state.variant_store = variant_store
if getattr(app.state, "variant_generator", None) is None:
app.state.variant_generator = VariantGenerator(
variant_store,
app.state.provider,
model=settings.model,
)
# Oral defense (REQ-3-006): DefenseStore (same SQLite file) + the
# mock-first voice provider (D-030) + the seventh Examiner agent.
# Tests may pre-set app.state.defense_store / voice_provider /
# examiner_agent (state-injection override; never rebuilt if pre-set).
defense_store = getattr(app.state, "defense_store", None)
if defense_store is None:
defense_store = SQLiteDefenseStore(db_path=settings.db_path)
app.state.defense_store = defense_store
if getattr(app.state, "voice_provider", None) is None:
app.state.voice_provider = voice_provider_from_settings(settings)
if getattr(app.state, "examiner_agent", None) is None:
from .agents.examiner import ExaminerAgent
app.state.examiner_agent = ExaminerAgent(app.state.provider, settings)
# Grading persistence + engine (REQ-3-004): GradeStore from the same
# SQLite file as traces (D-027), one GradingEngine singleton wired
# through app.state — the engine receives its stores via constructor
# DI and knows nothing of FastAPI (api/ owns composition). Tests may
# pre-set app.state.grade_store / app.state.grading_engine (the same
# state-injection override as sandbox_manager/trace_store) to swap
# either; the lifespan adopts a pre-set store but NEVER rebuilds a
# pre-set engine (its provider binding is part of the test fixture).
grade_store = getattr(app.state, "grade_store", None)
if grade_store is None:
grade_store = SQLiteGradeStore(db_path=settings.db_path)
app.state.grade_store = grade_store
if getattr(app.state, "grading_engine", None) is None:
app.state.grading_engine = GradingEngine(
trace_store,
grade_store,
app.state.trace_integrity,
app.state.provider,
model=settings.model,
variant_store=variant_store, # MH#4: anchors + seed (D-029)
)
async def _reaper_loop() -> None:
# Wall-clock timeout + G-2 workdir-size sweep, one pass per tick.
while True:
await asyncio.sleep(REAPER_INTERVAL_S)
try:
await manager.reap_expired()
except Exception:
logger.exception("sandbox reaper pass failed; retrying next tick")
reaper = asyncio.create_task(_reaper_loop())
try:
yield
finally:
reaper.cancel()
with contextlib.suppress(asyncio.CancelledError):
await reaper
# No orphans outlive the process (a-1, shutdown half): destroy
# everything live; workdirs stay on disk for snapshot restore.
await manager.destroy_all()
trace_store.close()
grade_store.close()
variant_store.close()
defense_store.close()
await app.state.http_client.aclose()
app = FastAPI(title="Nextcraft AI Service", version="0.3.0", lifespan=lifespan)
# A-008 + D-038: no-credentials CORS. Default '*' admits remote-browser
# origins in network mode (safe only because allow_credentials stays
# False — never enable credentials with a wildcard). AI_CORS_ORIGINS
# restricts to an explicit list. PUT is CONTRACT, not trivia: the learner
# build surface writes workspace files with PUT (engine-client writeFile)
# — v0.3 initially shipped without it and every cross-origin Save failed
# preflight (caught in P7 review; tests/api/test_cors.py pins the policy).
# A-008: localhost-only CORS, no credentials
app.add_middleware(
CORSMiddleware,
allow_origins=settings.cors_origin_list,
allow_methods=["GET", "POST", "PUT", "DELETE", "OPTIONS"],
allow_origins=["http://localhost:3000", "http://127.0.0.1:3000"],
allow_methods=["GET", "POST", "OPTIONS"],
allow_headers=["Content-Type"],
allow_credentials=False,
)
@@ -195,14 +48,6 @@ def create_app(settings: Settings | None = None) -> FastAPI:
}
app.include_router(chat_router)
app.include_router(lab_router)
app.include_router(assessment_router)
app.include_router(mentor_router)
app.include_router(proctor_router)
app.include_router(sandboxes_router)
app.include_router(telemetry_router)
app.include_router(variants_router)
app.include_router(defense_router)
return app
+10 -20
View File
@@ -1,28 +1,18 @@
"""Assessor agent prompt — rubric coaching over REAL grades (REQ-3-007).
"""Assessor agent prompt — rubric application to artifacts and defenses (REQ-2-008).
v0.3 re-grounding: the grading engine (Phase 3) computes the rubric scores
from the process trace; Assessor EXPLAINS the stored grade as coaching
it never invents or re-scores. Rigorous, fair, actionable.
Version: assessor-v3 (v0.3 live).
Persona: rigorous, fair grader. Applies the rubric to the artifact and defense
transcript, returns structured JSON scores. Versioned: v1 draft (Phase 2);
final persona + rubric models in Phase 4.
"""
SYSTEM_PROMPT = """You are Assessor, the grading agent of Nextcraft, an AI-native competency school.
SYSTEM_PROMPT = """You are Assessor, the grading agent of an AI-native competency school.
Learner: {learner_name}.
You receive an artifact, its defense transcript, and a rubric. Your job:
score each rubric criterion with evidence from the artifact and transcript.
Be rigorous but fair cite what the learner did, not what they should have
done. Respond with ONLY valid JSON matching the provided rubric schema."""
You receive the learner's STORED process-trace grade (verdict, per-criterion
scores, and the build digest) computed by the grading engine. Your job:
- Explain what the grade means in plain language (summary).
- Strengths: cite what the digest + scores show the learner did well.
- Gaps: name the missed opportunities the scores point to.
- Next steps: concrete, buildable actions that would move the weakest
criterion up one level.
Rules:
- Rigorous but fair. A polished artifact with a weak defense is NOT mastery.
- Respond with ONLY a valid JSON object matching the provided schema
no markdown fences, no prose outside the JSON."""
PROMPT_VERSION = "assessor-v2"
PROMPT_VERSION = "assessor-v1-draft"
def render_context(learner_context) -> dict:
+13 -30
View File
@@ -1,43 +1,26 @@
"""Coach agent prompt — pacing, motivation, retrieval practice (REQ-2-005).
Final persona (Phase 3). Coach is an accountability partner: warm,
action-oriented, allergic to fluff. Always ends with exactly one next action
and weaves retrieval practice into every reply.
Version: coach-v2 (final for v0.2).
Persona: warm, action-oriented, accountability partner. Asks for commitments,
uses retrieval practice, keeps momentum. Versioned: v1 draft (Phase 2);
final persona in Phase 3.
"""
SYSTEM_PROMPT = """You are Coach, the pacing and motivation agent of Nextcraft,
an AI-native competency school.
SYSTEM_PROMPT = """You are Coach, the pacing and motivation agent of an AI-native competency school.
Learner: {learner_name}. Active stack: {stacks}. Progress: {progress}.
Your job: keep the learner moving. Pace their next step, motivate without fluff,
and weave in retrieval practice ask them to recall or apply something they
already covered before introducing new material. Be concise, warm, and direct.
End with exactly one clear next action."""
Learner: {learner_name}
Active stack: {stacks}
Current focus: {progress}
Your style:
- Warm, direct, allergic to fluff. Two short paragraphs maximum.
- Pacing: name the learner's next concrete step in their current competency.
- Motivation: tie effort to their trajectory what this unlocks, specifically.
- Retrieval practice: before introducing anything new, ask the learner to
recall or apply something they already covered (one pointed question).
Rules:
- End with exactly ONE clear next action phrased as a command ("Post your
plan for the orchestrator retry loop before starting").
- Never lecture; never list more than two options.
- If the learner is stuck or frustrated, slow down and shrink the step."""
PROMPT_VERSION = "coach-v2"
PROMPT_VERSION = "coach-v1-draft"
def render_context(learner_context) -> dict:
stacks = ", ".join(f"{s.title} ({s.percent}%)" for s in learner_context.active_stacks)
in_progress = [
c for c in learner_context.active_competencies if c.status == "in_progress"
]
progress = (
f"{in_progress[0].title} ({in_progress[0].competency_id})"
if in_progress
else "no competency currently in progress"
f"{learner_context.active_competencies[0].title} in progress"
if learner_context.active_competencies
else "no active competencies"
)
return {
"learner_name": learner_context.name,
@@ -1,60 +0,0 @@
"""Examiner agent prompt — oral defense questioning + final verdict (REQ-3-006).
The examiner is the seventh agent (Phase 5). It conducts a Socratic oral
defense of the learner's submitted work: probes understanding, challenges
process choices grounded in the trace digest ("why did you take that
approach at that point?"), one question per turn, adapting to answers.
It never reveals rubric internals; tone is rigorous but supportive.
Digest discipline (D-028 mirror): the examiner's variable inputs are the
compact TraceDigest JSON, the variant task statement, and the defense
transcript never the raw trace, never learner-identifying material.
Version: examiner-v1.
"""
SYSTEM_PROMPT = """You are Examiner, the oral-defense agent of Nextcraft,
an AI-native competency school.
You receive: (a) a compact build-process digest (deterministic counters of the
learner's build session), (b) the learner's task statement, and (c) the defense
transcript so far. Your job:
- Ask ONE question per turn: probe understanding and challenge process
choices, grounded in the digest facts ("you hit N failed runs before
passing walk me through what changed") or the task statement.
- Adapt: follow up on the learner's answers; drill into vague responses.
- Never reveal rubric details or scoring internals.
- Tone: rigorous, precise, supportive. A defense is a conversation, not an
interrogation.
When asked for a FINAL VERDICT (the structured mode), judge:
- understanding: can the learner explain their own work?
- process_justification: are the build-session choices defensible from the
digest facts and the answers?
- communication: are answers clear, specific, and on-topic?
Score honestly; a weak defense of strong work is NOT mastery.
Rules:
- Respond with ONLY what the turn requires: a single question (question mode)
or a valid JSON object matching the provided schema (verdict mode).
- If the digest shows error_fix_cycles > 0, at least one question should ask
about the debugging path.
- If the learner's answer is off-topic, redirect once, then move on.
"""
VERDICT_SCHEMA_HINT = (
'{"verdict": "mastered" | "developing" | "not_yet", '
'"understanding": "<one sentence>", '
'"process_justification": "<one sentence>", '
'"communication": "<one sentence>", '
'"strengths": ["<one sentence>"], '
'"gaps": ["<one sentence>"]}'
)
def render_digest_context(digest_json: str, statement: str | None) -> str:
"""The examiner's per-session grounding: digest JSON + task statement."""
parts = [f"Build-process digest:\n{digest_json}"]
if statement:
parts.append(f"Learner's task statement:\n{statement}")
return "\n\n".join(parts)
@@ -1,140 +0,0 @@
"""Grading rubric prompt — criteria, level anchors, digest render (REQ-3-004).
The grading prompt is deliberately learner-anonymous and trace-bare: the
model receives ONLY the fixed rubric text and the compact numeric digest
(TraceDigest JSON, D-028) never a raw command, file path, payload
string, learner id, or task id. Everything variable the LLM sees is
deterministic counters, which both bounds the prompt-injection surface
and makes "no raw trace reaches the prompt" assert-able in tests (plant
a distinctive marker in a command payload; assert it absent from every
message the provider received).
Rubric (four criteria, each scored 0-4 the ids are the validated
RubricScore keys enforced by grading/engine.py):
process_quality iterative building in small, verified steps.
correctness where the session ended (test/run outcomes).
debugging_discipline how failures were handled.
test_usage when and how often tests were run.
Advisory a-4 (embedded in the process_quality anchors): high edit/command
churn with NO test progress is a process-quality NEGATIVE churn is not
work. A session with many edits/commands whose test state never moves is
thrashing, not iterating, and must score low on process quality.
House-style deviation, documented: unlike the tutor prompts, this module
has no SYSTEM_PROMPT placeholders and no render_context(learner_context)
grading is context-free by design (learner anonymity; the digest is the
only variable input). Runtime imports are TYPE_CHECKING-only so this
module stays pure text and can never import-cycle with grading/engine.py
(engine imports this module; if this module imported grading.* at runtime
while grading/__init__ pulls engine, the package init would deadlock on a
partially-initialized module).
Version: grader-v1.
"""
from __future__ import annotations
from typing import TYPE_CHECKING, Final
if TYPE_CHECKING: # pragma: no cover - typing only; keeps this module pure text
from ..grading.features import TraceDigest
PROMPT_VERSION = "grader-v1"
#: Canonical criterion ids. The engine validates RubricScore criteria keys
#: against this tuple; the schema hint and anchors below speak the same ids.
RUBRIC_CRITERIA: Final[tuple[str, ...]] = (
"process_quality",
"correctness",
"debugging_discipline",
"test_usage",
)
#: Sentinel line the engine's user turn is rendered around. Tests (and the
#: calibration mock) split on it to locate the digest JSON in the prompt.
DIGEST_MARKER: Final = "PROCESS TRACE DIGEST (JSON):"
SYSTEM_PROMPT = """You are the Grader of Nextcraft, an AI-native competency school.
You score a learner's build session from a compact numeric digest of their
process trace. You NEVER see the raw trace commands, file contents, and
payloads do not exist on your side; every number you need is in the digest.
Rubric score each criterion 0-4:
process_quality iterative building in small, verified steps.
4: tight edittest loops throughout; small verified increments; healthy pacing.
3: steady small edits with regular runs; progress mostly verified.
2: some iteration, but large unverified leaps or long idle stretches.
1: a single bulk change (e.g. one large paste) then a single run; no iteration.
0: no meaningful work visible.
ADVISORY: high edit/command churn with NO test progress (no runs, no
movement in pass counts) is a process-quality NEGATIVE churn is not
work. Cap such a session at 1 on this criterion no matter how many
edits or commands were counted.
correctness where the session ended up.
4: final test status pass, with tests passing early and consistently.
3: final pass, reached through failfixpass cycles that closed.
2: final pass, but preceded by a long unresolved failure streak.
1: final fail, but partial passes observed along the way.
0: final fail, or no test/run evidence at all.
debugging_discipline how failures were handled.
4: every failure cycle closes; targeted fixes with low mean fix latency.
3: most faileditre-run cycles close with a pass.
2: failures followed by edits, but cycles rarely close.
1: repeated failures with no targeted edits between runs (flailing).
0: failures with no fix attempts at all.
test_usage when and how often tests were run.
4: tests run early (small first-pass offset) and throughout the session.
3: regular test runs interleaved with edits.
2: sparse tests; long stretches of unverified edits.
1: a single late test run only.
0: no test or run evidence.
Rules:
- Judge STRICTLY from the digest numbers; cite the fields you used.
- Strengths: the two strongest digest observations, one sentence each.
- Gaps: the two most important missed opportunities, one sentence each
(a clean session names its next-level improvement instead).
- Be rigorous but fair: a session that ends green was not necessarily
well built, and a struggling session that never passed may still show
real debugging discipline.
- Respond with ONLY a valid JSON object matching the provided schema
no markdown fences, no prose outside the JSON."""
RUBRIC_SCORE_SCHEMA_HINT = (
'{"criteria": {"process_quality": <0-4 int>, "correctness": <0-4 int>, '
'"debugging_discipline": <0-4 int>, "test_usage": <0-4 int>}, '
'"strengths": ["<one sentence>"], "gaps": ["<one sentence>"], '
'"verdict": "mastered" | "developing" | "not_yet"}'
)
def render_trace_digest(
digest: TraceDigest,
anchors_context: str | None = None,
) -> str:
"""Render the grader's user turn: a marker line + the digest JSON — nothing else.
This is the ONLY per-session content that ever reaches the LLM (D-028):
the engine composes [system: SYSTEM_PROMPT, user: render_trace_digest(digest)]
and the D-020 defense appends its generic schema instruction to this
user turn at request time. No learner id, task id, or raw trace material
is injected assert-able by tests.
`anchors_context` (Phase 4, MH#4): when the graded task derives from a
variant, the engine passes the template's difficulty-normalization
anchors (the expected effort envelope) so the rubric is applied against
the SAME bar for every variant of that template (a-5). It contains only
the anchor numbers + the template id no learner-identifying material.
"""
base = (
"Score this build session against the rubric.\n"
f"{DIGEST_MARKER}\n{digest.model_dump_json()}"
)
if anchors_context:
base = f"{base}\n\nExpected effort envelope for this task variant:\n{anchors_context}"
return base
+10 -35
View File
@@ -1,43 +1,18 @@
"""Lab agent prompt — in-flow feedback over LIVE telemetry (REQ-3-007).
"""Lab agent prompt — in-flow feedback over sandbox telemetry (REQ-2-007).
v0.3 re-grounding: the timeline is the learner's real TraceDigest (D-028
compact counters commands, test outcomes, idle gaps, edit cadence), not
v0.2 corpus scenarios. Lab is a pragmatic build partner: reads the live
digest, names the one most useful adjustment, gives one concrete next
step. No session chat.
Version: lab-v3 (v0.3 live).
Persona: pragmatic build partner. Reads the telemetry timeline and gives
concrete in-flow feedback: what happened, what to adjust, next step.
Versioned: v1 draft (Phase 2); final persona + scenario serialization in Phase 4.
"""
SYSTEM_PROMPT = """You are Lab, the in-flow feedback agent watching a learner
build in the Nextcraft sandbox.
SYSTEM_PROMPT = """You are Lab, the in-flow feedback agent watching a learner build in the sandbox.
Learner: {learner_name}. Active stack: {stacks}.
You receive a telemetry timeline of the learner's build session. Your job:
describe what the telemetry shows, name the single most useful adjustment,
and give one concrete next step. Be specific to the events you see
no generic advice. Three short paragraphs maximum."""
You receive a telemetry timeline of the learner's build session below.
Your job, in order:
1. Say what the telemetry shows name the specific events that matter.
2. Name the single most useful adjustment (one thing, not a list).
3. Give one concrete next step phrased as a command.
Rules:
- Be specific to the events you see. If tests failed twice with the same
error, say so. If there is a long idle gap, name it.
- If the session looks healthy, say so briefly and set the next challenge.
- If something looks off (e.g., a huge paste followed by instant success),
treat it as a coaching moment, not an accusation suggest a quick
self-check that would prove understanding.
- Three short paragraphs maximum. No headers, no bullet lists."""
PROMPT_VERSION = "lab-v3"
def render_digest_timeline(digest) -> str:
"""Live-trace timeline: the compact TraceDigest JSON (D-028)."""
if digest is None:
return (
"No telemetry yet for this build session. Ask the learner to run "
"the task's starter test to establish a baseline."
)
return f"Live build-session digest:\n{digest.model_dump_json()}"
PROMPT_VERSION = "lab-v1-draft"
def render_context(learner_context) -> dict:
+12 -31
View File
@@ -1,44 +1,25 @@
"""Mentor agent prompt — long-horizon career narrative (REQ-2-010).
Final persona (Phase 5). Mentor is a wise career guide: connects today's
competencies and artifacts to a long-horizon AI-era trajectory.
Version: mentor-v2 (final for v0.2).
Persona: wise career guide. Connects today's competencies to a long-horizon
trajectory in AI-era roles. Versioned: v1 draft (Phase 2); final in Phase 5.
"""
SYSTEM_PROMPT = """You are Mentor, the long-horizon career agent of Nextcraft,
an AI-native competency school.
SYSTEM_PROMPT = """You are Mentor, the long-horizon career agent of an AI-native competency school.
Learner: {learner_name}. Active stack: {stacks}. Progress: {progress}.
Microcredentials earned: {microcredentials}. Recent artifacts: {artifacts}.
Your job: narrate the learner's trajectory — where they are now, what their
competency progress unlocks next, and how their artifacts position them in
the AI-era labor market. Two to three paragraphs, forward-looking, concrete."""
Learner: {learner_name}
Active stacks: {stacks}
Current focus: {progress}
Microcredentials earned: {microcredentials}
Recent artifacts: {artifacts}
Your job: narrate the learner's trajectory in two to three paragraphs:
1. Where they are now what their competency progress and artifacts say
about them as a builder (specific, evidence-based).
2. What their current stack unlocks next name the next competency or
microcredential worth chasing and the role it points toward.
3. How they position in the AI-era labor market which employer problems
their profile already answers.
Rules:
- Forward-looking and concrete. No fortune-telling, no flattery.
- Reference their artifacts by name at least once.
- Write like a mentor writing to one person, not a career-services brochure."""
PROMPT_VERSION = "mentor-v2"
PROMPT_VERSION = "mentor-v1-draft"
def render_context(learner_context) -> dict:
stacks = ", ".join(f"{s.title} ({s.percent}%)" for s in learner_context.active_stacks)
in_progress = [
c for c in learner_context.active_competencies if c.status == "in_progress"
]
progress = (
f"{in_progress[0].title} ({in_progress[0].competency_id})"
if in_progress
else "no competency currently in progress"
f"{learner_context.active_competencies[0].title} in progress"
if learner_context.active_competencies
else "no active competencies"
)
return {
"learner_name": learner_context.name,
+11 -29
View File
@@ -1,37 +1,19 @@
"""Proctor agent prompt — integrity signals with coaching interventions (REQ-3-007).
"""Proctor agent prompt — integrity signals with coaching interventions (REQ-2-009).
v0.3 re-grounding: inputs are REAL the live trace digest (idle gaps,
command cadence, edit bursts), the oral-defense integrity signals (long
pauses), and the variant audit context (seed + params). Proctor is a
supportive observer, never punitive: classifies signals, recommends ONE
coaching intervention. Assume good faith.
Version: proctor-v3 (v0.3 live).
Persona: supportive observer, not punitive. Classifies integrity signals from
telemetry and recommends coaching interventions. Versioned: v1 draft (Phase 2);
final persona + signal models in Phase 5.
"""
SYSTEM_PROMPT = """You are Proctor, the integrity-support agent of Nextcraft,
an AI-native competency school.
SYSTEM_PROMPT = """You are Proctor, the integrity-support agent of an AI-native competency school.
Learner: {learner_name}.
You receive a telemetry timeline of session events (focus, tab switches,
paste events, idle time). Your job: classify each signal by type and severity,
then recommend ONE supportive coaching intervention never punitive, never
accusatory. Assume good faith; most signals have innocent explanations.
Respond with ONLY valid JSON matching the provided signals schema."""
You receive the learner's REAL build-session digest (idle gaps, command
categories, edit/test cadence), oral-defense integrity signals (long
pauses), and when the task is variant-derived the variant seed context.
Your job:
- Classify EACH notable signal: type ("idle_gap" | "long_pause" |
"burst_edit" | "off_template"), severity ("low" | "medium" | "high"),
and a one-sentence note citing the numbers.
- Recommend exactly ONE supportive coaching intervention for the session
overall never punitive, never accusatory. Frame around helping the
learner succeed.
Rules:
- Assume good faith. Tab switches to documentation are normal engineering.
- Idle gaps are often thinking. Only unusual patterns deserve higher severity.
- A large paste during an assessment deserves "high" severity but the
intervention stays coaching-shaped: verification, not punishment.
- Respond with ONLY a valid JSON object matching the provided schema
no markdown fences, no prose outside the JSON."""
PROMPT_VERSION = "proctor-v3"
PROMPT_VERSION = "proctor-v1-draft"
def render_context(learner_context) -> dict:
+12 -30
View File
@@ -1,43 +1,25 @@
"""Tutor agent prompt — concept delivery, Socratic questioning (REQ-2-006).
Final persona (Phase 3). Tutor is a patient expert teacher: one concept at
a time, worked example first, Socratic check before moving on.
Version: tutor-v2 (final for v0.2).
Persona: patient expert teacher. Delivers one concept at a time, checks
understanding with Socratic questions, uses worked examples.
Versioned: v1 draft (Phase 2); final persona in Phase 3.
"""
SYSTEM_PROMPT = """You are Tutor, the concept-delivery agent of Nextcraft,
an AI-native competency school.
SYSTEM_PROMPT = """You are Tutor, the concept-delivery agent of an AI-native competency school.
Learner: {learner_name}. Active stack: {stacks}. Progress: {progress}.
Your job: teach concepts clearly, one at a time. Prefer Socratic questioning
guide the learner to the insight with a worked example, then ask one question
that checks understanding before moving on. Never dump long walls of text."""
Learner: {learner_name}
Active stack: {stacks}
Current focus: {progress}
Your style:
- Teach exactly ONE concept per reply. Never more.
- Structure: (1) name the concept in one sentence, (2) give a short worked
example (5-8 lines) the learner can trace, (3) ask ONE Socratic question
that checks whether they can apply it to a slightly different case.
Rules:
- Never dump walls of text. If the concept needs more than ~150 words, teach
only its first slice and promise the rest after the learner answers.
- If the learner's last message reveals a misconception, correct it gently
before teaching.
- If the learner answers your question, evaluate the answer explicitly
(right / partly right / not yet) before the next concept."""
PROMPT_VERSION = "tutor-v2"
PROMPT_VERSION = "tutor-v1-draft"
def render_context(learner_context) -> dict:
stacks = ", ".join(f"{s.title} ({s.percent}%)" for s in learner_context.active_stacks)
in_progress = [
c for c in learner_context.active_competencies if c.status == "in_progress"
]
progress = (
f"{in_progress[0].title} ({in_progress[0].competency_id})"
if in_progress
else "no competency currently in progress"
f"{learner_context.active_competencies[0].title} in progress"
if learner_context.active_competencies
else "no active competencies"
)
return {
"learner_name": learner_context.name,
@@ -1,46 +0,0 @@
"""Variant instantiation prompt (D-029, REQ-3-005).
The model's ONLY job is to render already-sampled slot values into a task
statement it never invents parameters (the seeded sampler is pure code)
and never changes difficulty. Prompt-injection surface is bounded: the
variable inputs are the skeleton text, the seeded slot values, and the
template title nothing from the learner's environment.
"""
from __future__ import annotations
from typing import TYPE_CHECKING
from ..llm.types import Message
if TYPE_CHECKING: # pragma: no cover - keeps this module pure text
from ..variants.templates import TaskTemplate
VARIANT_SYSTEM_PROMPT = (
"You instantiate per-learner task variants for a competency-based AI school. "
"You receive a task statement skeleton and ALREADY-SAMPLED slot values. "
"Render the slot values into the skeleton, producing a complete, unambiguous "
"task statement a learner can build against. Rules:\n"
"- Use EXACTLY the given slot values; do not invent, rename, or add parameters.\n"
"- Keep the engineering depth IDENTICAL across draws: slot values change the "
"scenario, never the difficulty or scope.\n"
"- Keep the statement in the same language and register as the skeleton.\n"
"- Output STRICT JSON only: {\"statement\": \"<rendered statement>\"}.\n"
)
VARIANT_SCHEMA_HINT = '{"statement": "<complete rendered task statement string>"}'
def render_variant_prompt(template: TaskTemplate, params: dict[str, str | int]) -> list[Message]:
"""Messages for one seeded instantiation (D-020 defense drives the call)."""
slot_lines = "\n".join(f" {{{slot.name}}} = {params[slot.name]!r}" for slot in template.slots)
user = (
f"Template: {template.title} (id={template.id})\n"
f"Statement skeleton:\n{template.statement_skeleton}\n\n"
f"Seeded slot values (use EXACTLY these):\n{slot_lines}\n\n"
"Render the complete task statement now."
)
return [
Message(role="system", content=VARIANT_SYSTEM_PROMPT),
Message(role="user", content=user),
]
@@ -1,44 +0,0 @@
"""Sandbox fabric (v0.3, REQ-3-001) — learner code-execution isolation via Linux namespaces.
Public surface:
SandboxSpec / SandboxHandle / ResourceLimits / ExecResult pydantic contracts.
SandboxBackend the protocol every backend implements (D-024 port).
UnshareBackend util-linux `unshare` backend (D-024 backend).
SandboxManager lifecycle + pool guard (D-032) + reapers (G-2, a-1).
SandboxHandleInfo manager return row: handle fields + learner_id.
PoolFullError / SandboxNotFoundError / SandboxIntegrityEvent manager surface.
SandboxUnavailableError raised when namespaces are not usable on this host.
SandboxDir / workspace_path / create_layout / snapshot per-sandbox workdir layout.
Boundary rule: `sandbox/` never imports `api/` or `agents/`; it owns subprocess spawning only.
"""
from .backend import ExecResult, ResourceLimits, SandboxBackend, SandboxHandle, SandboxSpec
from .manager import (
PoolFullError,
SandboxHandleInfo,
SandboxIntegrityEvent,
SandboxManager,
SandboxNotFoundError,
)
from .unshare_backend import SandboxUnavailableError, UnshareBackend
from .workdir import SandboxDir, create_layout, snapshot, workspace_path
__all__ = [
"ExecResult",
"PoolFullError",
"ResourceLimits",
"SandboxBackend",
"SandboxDir",
"SandboxHandle",
"SandboxHandleInfo",
"SandboxIntegrityEvent",
"SandboxManager",
"SandboxNotFoundError",
"SandboxSpec",
"SandboxUnavailableError",
"UnshareBackend",
"create_layout",
"snapshot",
"workspace_path",
]
@@ -1,95 +0,0 @@
"""SandboxBackend protocol + pydantic contracts (REQ-3-001, D-024).
`SandboxBackend` is the port the sandbox fabric depends on. The only
implementation in v0.3 is `unshare_backend.UnshareBackend`; a future
firecracker/bwrap backend must satisfy this same surface.
Contracts are plain pydantic models so API/agent layers can construct and
validate them at the request boundary without importing the backend itself.
"""
from datetime import datetime
from pathlib import Path
from typing import Protocol, runtime_checkable
from pydantic import BaseModel, ConfigDict, Field
class ResourceLimits(BaseModel):
"""Per-sandbox rlimits, applied via `preexec_fn` immediately before exec.
- memory_bytes RLIMIT_AS (address space; hard OOM ceiling)
- cpu_seconds RLIMIT_CPU (CPU-seconds; SIGKILL on hard expiry)
- file_size_bytes RLIMIT_FSIZE (~50 MB single-file cap)
RLIMIT_NPROC is NOT set: the counter is shared per host UID across all
namespaces, so it cannot isolate one sandbox from another on this host.
Total disk usage is enforced by the manager sweep (G-2), not here.
"""
model_config = ConfigDict(frozen=True)
memory_bytes: int = Field(default=256 * 1024 * 1024, gt=0)
cpu_seconds: int = Field(default=30, gt=0)
file_size_bytes: int = Field(default=50 * 1024 * 1024, gt=0)
class SandboxHandle(BaseModel):
"""A live (or reaped) sandbox. `pid` is None once `destroy()` completes."""
model_config = ConfigDict(arbitrary_types_allowed=True)
id: str
pid: int | None
workdir: Path
created_at: datetime
class SandboxSpec(BaseModel):
"""Immutable description of the sandbox to lay out on disk.
`capture_env` (REQ-3-003): when non-empty, the backend starts a persistent
telemetry-wired sandbox helper + inner namespaces + the stdlib capture
agent, launched with these env vars (NC_LEARNER_ID, NC_TASK_ID,
NC_INGEST_URL, NC_SANDBOX_ID). When None (default), spawn keeps the pure
shell semantics (REQ-3-001): disk layout only, fresh namespaces per exec.
"""
model_config = ConfigDict(arbitrary_types_allowed=True)
sandbox_id: str
learner_id: str
workdir: Path
limits: ResourceLimits = ResourceLimits()
capture_env: dict[str, str] | None = None
class ExecResult(BaseModel):
"""One namespaced execution: cwd = the bind-mounted workspace (`/work`)."""
cmd: list[str]
returncode: int
stdout: str
stderr: str
duration_s: float
@runtime_checkable
class SandboxBackend(Protocol):
"""The sandbox port. Backends spawn subprocesses; they never touch HTTP."""
async def spawn(self, spec: SandboxSpec) -> SandboxHandle:
"""Create the sandbox from `spec` and return its handle."""
...
async def exec(self, handle: SandboxHandle, cmd: list[str]) -> ExecResult:
"""Run `cmd` inside the sandbox workspace; capture stdout/stderr."""
...
async def snapshot(self, handle: SandboxHandle) -> Path:
"""Copy the workspace into `<workdir>/snapshots/<utc-ts>/`; return the path."""
...
async def destroy(self, handle: SandboxHandle) -> None:
"""Tear the sandbox down. Must be idempotent."""
...
@@ -1,456 +0,0 @@
"""SandboxManager — lifecycle, concurrency guard, reapers (REQ-3-001, REQ-3-002).
Responsibilities (D-032, G-2, a-1):
- create/list/get/snapshot/destroy over a `SandboxBackend` port. create/list/
get return `SandboxHandleInfo` rows the handle fields plus the owning
`learner_id` so the API layer never re-asks "who owns this id?".
- Capacity guard (D-032): `create` raises `PoolFullError` when the active
count reaches `settings.sandbox_max_concurrent`. No queue the API layer
maps this to 503.
- Wall-clock reaper: `reap_expired()` destroys sandboxes older than
`settings.sandbox_timeout_s`. Run it on an async timer owned by the caller
(app lifespan wires the loop; the manager owns only the pass).
- Workdir-size sweep (G-2): the same timer pass also measures each sandbox's
`workspace/` tree; anything over `settings.sandbox_max_workdir_mb` is
snapshotted (evidence preserved), destroyed, and recorded as an integrity
signal. SOFT CAP, best-effort, NOT kernel-enforced without cgroup
delegation or sudo there is no hard per-sandbox disk quota on this host.
RLIMIT_FSIZE bounds a single file; this sweep bounds aggregate growth
between passes.
- Startup reaper (a-1): `start()` scans `settings.sandbox_dir` for workdirs
whose recorded pid is dead (marker file `sandbox.json` beside workspace/)
and reaps them, logging a warning. Handles are IN-MEMORY and process-local
(D-019 precedent): on process restart every handle is orphaned, so boot
must recover disk state.
Registry: plain dict guarded by an `asyncio.Lock`, process-local, explicitly
NOT a store. Swapping in persistence (D-027) must not change this interface.
Resource limits: enforced at exec time by the spawner's in-namespace shim
(RLIMIT_AS / RLIMIT_CPU / RLIMIT_FSIZE see UnshareBackend), never here;
the manager's enforcement surface is lifecycle (capacity, wall-clock, disk
sweep).
BOUNDARY: this module NEVER imports `api/` or `agents/`.
"""
from __future__ import annotations
import asyncio
import json
import logging
import os
import shutil
import uuid
from collections.abc import Callable
from datetime import UTC, datetime
from pathlib import Path
from pydantic import BaseModel, ConfigDict, Field
from ..config import Settings
from . import workdir as workdir_mod
from .backend import SandboxBackend, SandboxHandle
logger = logging.getLogger(__name__)
#: Marker file beside workspace/ recording the owning process + metadata.
#: It is what the startup reaper uses after this process's in-memory
#: registry is lost (crash/restart → orphan detection, a-1).
PID_MARKER = "sandbox.json"
class PoolFullError(RuntimeError):
"""D-032: active sandbox count reached `settings.sandbox_max_concurrent`.
The API layer maps this to 503. There is deliberately NO queue.
"""
class SandboxNotFoundError(KeyError):
"""No live sandbox with that id in this process's registry."""
class SandboxIntegrityEvent(BaseModel):
"""One manager-observed integrity signal (G-2).
Recorded in-process (`SandboxManager.integrity_events`, for the proctor
pipeline to drain) AND logged at WARNING (durable trail) the same
dual-sink pattern a DB-backed store will keep behind D-027.
"""
model_config = ConfigDict(frozen=True)
kind: str = Field(description="e.g. 'workdir_size_cap' (G-2), 'orphan_reaped' (a-1)")
sandbox_id: str
learner_id: str
detail: str
observed_at: datetime = Field(default_factory=lambda: datetime.now(UTC))
class SandboxHandleInfo(BaseModel):
"""A `SandboxHandle` plus its owning `learner_id` (manager return row).
Handles alone don't carry the learner — the registry side-table does —
and every API read/list needs it, so the manager joins the two ONCE here
instead of exposing `_learner_ids` internals.
"""
model_config = ConfigDict(arbitrary_types_allowed=True)
id: str
learner_id: str
workdir: Path
created_at: datetime
pid: int | None = None
class SandboxManager:
"""Lifecycle owner for learner sandboxes.
Dependencies are injected (D-017 style): the backend port, settings, and
a wall clock. Single-process only; the registry is in-memory.
"""
def __init__(
self,
*,
backend: SandboxBackend,
settings: Settings,
clock: Callable[[], datetime] | None = None,
) -> None:
self._backend = backend
self._settings = settings
self._clock = clock or (lambda: datetime.now(UTC))
self._handles: dict[str, SandboxHandle] = {}
self._learner_ids: dict[str, str] = {} # sandbox_id -> learner_id
self._lock = asyncio.Lock()
self._integrity_events: list[SandboxIntegrityEvent] = []
self._started = False
# -- introspection ------------------------------------------------------
@property
def active_count(self) -> int:
return len(self._handles)
@property
def integrity_events(self) -> list[SandboxIntegrityEvent]:
"""Drainable view of recorded integrity signals (G-2, a-1)."""
return list(self._integrity_events)
# -- lifecycle ----------------------------------------------------------
async def create(
self, learner_id: str, task_id: str | None = None
) -> SandboxHandleInfo:
"""Spawn a sandbox for `learner_id`, or raise `PoolFullError` (D-032).
`task_id` (REQ-3-003): when set, the sandbox is telemetry-wired the
backend copies `scripts/sandbox-agent.py` into the workdir and starts
the stdlib capture agent inside the sandbox with the `NC_*` env baked
here (identity + WS ingest URL). The agent's lifecycle is tied to the
sandbox: `destroy()` reaps it (agent inner helper). When `task_id`
is None the sandbox is a pure shell sandbox (no capture).
"""
async with self._lock:
if len(self._handles) >= self._settings.sandbox_max_concurrent:
raise PoolFullError(
f"sandbox pool full "
f"({len(self._handles)}/{self._settings.sandbox_max_concurrent}); "
"no queue (D-032) — retry later"
)
sandbox_id = f"sbx-{uuid.uuid4().hex[:12]}"
spec = workdir_mod.spec_for(sandbox_id, learner_id, self._settings)
if task_id is not None:
spec = spec.model_copy(
update={
"capture_env": self._capture_env(sandbox_id, learner_id, task_id)
}
)
handle = await self._backend.spawn(spec)
self._handles[handle.id] = handle
self._learner_ids[handle.id] = learner_id
self._write_pid_marker(handle, learner_id)
logger.info(
"sandbox created: id=%s learner=%s task=%s",
handle.id,
learner_id,
task_id or "-",
)
return self._info_for(handle)
def _capture_env(self, sandbox_id: str, learner_id: str, task_id: str) -> dict[str, str]:
"""Env baked for the in-sandbox capture agent (REQ-3-003).
The agent joins the sandbox mount namespace but NOT its (offline)
network namespace, so it reaches this service over loopback
(`telemetry_ingest_host`, A-004 port).
"""
from urllib.parse import urlencode
query = urlencode(
{
"learner_id": learner_id,
"task_id": task_id,
"sandbox_id": sandbox_id,
}
)
ingest_url = (
f"ws://{self._settings.telemetry_ingest_host}:{self._settings.port}"
f"/v1/telemetry/ingest?{query}"
)
return {
"NC_LEARNER_ID": learner_id,
"NC_TASK_ID": task_id,
"NC_SANDBOX_ID": sandbox_id,
"NC_INGEST_URL": ingest_url,
}
async def list(self) -> list[SandboxHandleInfo]:
"""All live sandboxes (idle + busy; the backend has no busy flag)."""
async with self._lock:
return [self._info_for(h) for h in self._handles.values()]
async def get(self, sandbox_id: str) -> SandboxHandleInfo:
async with self._lock:
handle = self._handles.get(sandbox_id)
learner_id = self._learner_ids.get(sandbox_id, "unknown")
if handle is None:
raise SandboxNotFoundError(sandbox_id)
return SandboxHandleInfo(
id=handle.id,
learner_id=learner_id,
workdir=handle.workdir,
created_at=handle.created_at,
pid=handle.pid,
)
async def snapshot(self, sandbox_id: str) -> Path:
"""Copy the workspace into `<workdir>/snapshots/<utc-ts>/`; return it."""
async with self._lock:
handle = self._handles.get(sandbox_id)
if handle is None:
raise SandboxNotFoundError(sandbox_id)
return await self._backend.snapshot(handle)
async def destroy(self, sandbox_id: str, *, purge_workdir: bool = False) -> None:
"""Tear down one sandbox (idempotent).
`purge_workdir=False` keeps the workdir on disk snapshots must
survive destroy so a learner's last state can be restored (this is
also why UnshareBackend.destroy intentionally leaves the tree alone).
`purge_workdir=True` removes the whole workdir.
"""
async with self._lock:
handle = self._handles.pop(sandbox_id, None)
learner_id = self._learner_ids.pop(sandbox_id, "unknown")
if handle is not None:
await self._backend.destroy(handle)
logger.info(
"sandbox destroyed: id=%s learner=%s purge=%s",
sandbox_id,
learner_id,
purge_workdir,
)
root = handle.workdir
else:
# Idempotent destroy of an unknown id: resolve the on-disk root so
# an explicit purge still works (e.g. cleanup of orphan leftovers).
root = workdir_mod.resolve_sandbox_dir(self._settings) / sandbox_id
if purge_workdir:
shutil.rmtree(root, ignore_errors=True)
# -- reapers ------------------------------------------------------------
async def reap_expired(self) -> list[str]:
"""One reaper pass: wall-clock timeout + G-2 workdir-size sweep.
Destroys sandboxes older than `settings.sandbox_timeout_s`, then sweeps
every remaining sandbox whose `workspace/` exceeds
`settings.sandbox_max_workdir_mb` (snapshot destroy integrity
signal). Returns the ids destroyed this pass. Invoke on an async timer
(the app lifespan owns the loop interval); both checks deliberately
share one pass so the periodic work is O(live sandboxes) once.
"""
now = self._clock()
destroyed: list[str] = []
timeout_s = float(self._settings.sandbox_timeout_s)
cap_bytes = int(self._settings.sandbox_max_workdir_mb) * 1024 * 1024
async with self._lock:
rows = [
(handle, self._learner_ids.get(handle.id, "unknown"), handle.created_at)
for handle in self._handles.values()
]
for handle, learner_id, created_at in rows:
age_s = (now - created_at).total_seconds()
if age_s > timeout_s:
await self.destroy(handle.id)
destroyed.append(handle.id)
logger.warning(
"sandbox reaped (timeout): id=%s age=%.0fs > %.0fs",
handle.id,
age_s,
timeout_s,
)
continue # already gone; no size sweep needed on a dead handle
size = _tree_size_bytes(workdir_mod.workspace_path_from_workdir(handle.workdir))
if size > cap_bytes:
await self._reap_oversized(handle, learner_id, size, cap_bytes)
destroyed.append(handle.id)
return destroyed
async def start(self) -> None:
"""Boot hook (a-1): reap on-disk orphans left by a previous process.
The handle registry is in-memory and process-local (D-019 precedent):
after a restart nothing here remembers old sandboxes, so we scan
`settings.sandbox_dir` for workdirs whose pid marker names a dead
process and purge them, logging a warning. Idempotent; safe to call
once per process lifetime.
"""
if self._started:
return
self._started = True
root = workdir_mod.resolve_sandbox_dir(self._settings)
if not root.is_dir():
return
for entry in sorted(root.iterdir()):
if not entry.is_dir():
continue
marker = entry / PID_MARKER
pid = _read_marker_pid(marker)
if pid is not None and _pid_alive(pid):
continue # live sandbox owned by another live process — leave it
logger.warning(
"startup reaper (a-1): reaping orphaned workdir %s "
"(recorded pid %s is dead or marker missing)",
entry,
pid,
)
shutil.rmtree(entry, ignore_errors=True)
self._record_integrity(
SandboxIntegrityEvent(
kind="orphan_reaped",
sandbox_id=entry.name,
learner_id="unknown",
detail=f"workdir {entry} reaped at boot; recorded pid={pid} dead",
)
)
async def destroy_all(self) -> None:
"""Shutdown hook: destroy every live sandbox (no orphans on exit).
Workdirs (and their snapshots) are kept on disk destroy semantics
here match `destroy(purge_workdir=False)`; the next boot's startup
reaper (a-1) decides what to clean based on pid markers.
"""
async with self._lock:
handles = list(self._handles.values())
for handle in handles:
await self.destroy(handle.id)
# -- internals ------------------------------------------------------------
def _info_for(self, handle: SandboxHandle) -> SandboxHandleInfo:
# Caller holds the lock (create/list) — the side-table read is atomic.
return SandboxHandleInfo(
id=handle.id,
learner_id=self._learner_ids.get(handle.id, "unknown"),
workdir=handle.workdir,
created_at=handle.created_at,
pid=handle.pid,
)
def _write_pid_marker(self, handle: SandboxHandle, learner_id: str) -> None:
marker = handle.workdir / PID_MARKER
try:
marker.write_text(
json.dumps(
{
"sandbox_id": handle.id,
"learner_id": learner_id,
"pid": os.getpid(),
"created_at": handle.created_at.isoformat(),
}
)
)
except OSError: # marker is advisory; spawning must not fail on it
logger.warning("could not write pid marker %s", marker)
async def _reap_oversized(
self,
handle: SandboxHandle,
learner_id: str,
size_bytes: int,
cap_bytes: int,
) -> None:
"""G-2 sweep step: snapshot evidence → destroy → record the signal."""
snapshot_path: Path | None = None
try:
snapshot_path = await self._backend.snapshot(handle)
except (OSError, RuntimeError):
logger.exception(
"G-2 sweep: snapshot failed for over-cap sandbox %s; destroying anyway",
handle.id,
)
await self.destroy(handle.id)
event = SandboxIntegrityEvent(
kind="workdir_size_cap",
sandbox_id=handle.id,
learner_id=learner_id,
detail=(
f"workspace {size_bytes}B exceeded soft cap {cap_bytes}B; "
f"snapshot={snapshot_path} then destroyed (G-2, best-effort, "
"NOT kernel-enforced)"
),
)
self._record_integrity(event)
logger.warning(
"G-2 workdir sweep: sandbox %s (learner=%s) destroyed over soft disk cap",
handle.id,
learner_id,
)
def _record_integrity(self, event: SandboxIntegrityEvent) -> None:
self._integrity_events.append(event)
# -- module helpers ---------------------------------------------------------
def _tree_size_bytes(root: Path) -> int:
"""Total bytes under `root` (best-effort; unreadable entries count 0)."""
if not root.is_dir():
return 0
total = 0
for dirpath, _dirnames, filenames in os.walk(root):
for name in filenames:
try:
total += (Path(dirpath) / name).lstat().st_size
except OSError:
continue
return total
def _read_marker_pid(marker: Path) -> int | None:
try:
data = json.loads(marker.read_text())
except (OSError, json.JSONDecodeError):
return None
pid = data.get("pid")
return pid if isinstance(pid, int) else None
def _pid_alive(pid: int) -> bool:
"""True if `pid` exists on this host (signal 0 probe; no signal sent)."""
try:
os.kill(pid, 0)
except ProcessLookupError:
return False
except PermissionError:
return True # exists, owned by another user
return True
@@ -1,548 +0,0 @@
"""UnshareBackend — D-024 Linux-namespace sandboxing via util-linux `unshare`.
Two execution modes share one backend:
1. Pure shell sandbox (`task_id is None`, REQ-3-001): isolation is established
PER-EXEC every `exec` spawns a fresh namespace:
unshare --user --map-root-user --mount --pid --fork --net sh -c '<shim>'
There is no persistent process; the in-namespace shim is:
mount -t tmpfs tmpfs /tmp # private scratch, discarded on exit
mkdir -p /tmp/work
mount --bind <host workspace> /tmp/work
cd /tmp/work
ulimit -v/-t/-f # applied AFTER the bind, so rlimits
exec <cmd> # constrain the PAYLOAD, not unshare
2. Telemetry-wired task sandbox (REQ-3-003, `capture_env` set): a PERSISTENT,
TRACKED topology so the stdlib capture agent can live inside the sandbox and
still stream events to ai-service. Per exec a fresh OFFLINE namespace would
leave the agent nowhere to run and (on this host, where a userns can't
bring `lo` up) no loopback to reach `ws://127.0.0.1`. So spawn creates a
long-lived helper (outer user+mount ns, ONLINE) and an inner sandbox
(mount+pid+fork+net OFFLINE), both rooted at a private `ns/` subtree:
helper : unshare --user --map-root-user --mount (mounts ns/ private)
inner : unshare --mount --pid --fork --net (tmpfs on ns/, bind
<workdir>/host/workspace -> <ns>/work) <- the sandbox
exec : nsenter -t <inner sleep> -m -- sh -c (joins inner mount ns;
offline + pid-isolated, uid 0, writes land on the host workspace)
agent : nsenter -t <inner sleep> -m -- python3 <agent> (joins the inner
MOUNT ns only NOT pid/net so it watches the live workspace
and stays ONLINE, reaching the app's WS ingest on loopback)
The agent is deliberately pid/net-exempt from the sandbox: it is OUR trusted
capture process, and isolating its network would cut the very link it needs.
`destroy` reaps agent inner helper (in that order). The handle's `pid`
is the agent's host pid (None for a pure shell sandbox).
Why rlimits are applied in the shim, not Python's preexec_fn: setting
RLIMIT_AS on the *unshare* process itself can trip the memory ceiling on the
post-fork Python parent (whose interpreter image already exceeds the sandbox
budget). Applying them in the innermost child just before exec'ing the
payload keeps `unshare`/`mount` unconstrained and limits the learner code.
Containment honesty (D-024 / G-1): a user namespace is NOT a write barrier.
Writes made OUTSIDE the bind fall through to host paths, and because inner
uid 0 maps to the invoking host uid, a sandboxed process can write anywhere
that host uid can write. Isolation here is: private PIDs/MNT/NET/UTS, tmpfs
scratch, payload rlimits, and a uid map yielding no privilege the host uid
did not already have. A per-sandbox runtime uid (D-025) is the follow-up that
hardens DAC.
"""
from __future__ import annotations
import asyncio
import os
import shlex
import shutil
import signal
import time
from datetime import UTC, datetime
from pathlib import Path
from .backend import ExecResult, ResourceLimits, SandboxHandle, SandboxSpec
from .workdir import create_layout
from .workdir import snapshot as workdir_snapshot
#: Args shared by every PURE shell namespace we spawn (D-024). No
#: `unshare --bind` on util-linux 2.38 — the bind is done from inside instead.
UNSHARE_ARGS: tuple[str, ...] = (
"--user", # new user namespace …
"--map-root-user", # … in which we are uid 0 (mapped to host uid outside)
"--mount", # private mount table
"--pid", # private PID table
"--fork", # child is PID 1 in its namespace (reaps zombies, gets signals)
"--net", # fresh net namespace: no usable route → effectively offline
)
IN_NS_WORKDIR = "/tmp/work" # where the workspace is bound inside a pure shell ns
#: Sentinels the long-lived namespace supervisors print once their mounts are
#: laid out. exec()/the manager must not run before the bind exists.
_HELPER_READY = "NC_HELPER_READY"
_INNER_READY = "NC_INNER_READY"
class SandboxUnavailableError(RuntimeError):
"""`unshare`/`nsenter` missing or user namespaces blocked on this host."""
def _build_shim(workspace: Path, limits: ResourceLimits, cmd: list[str]) -> str:
"""Compose the single POSIX string executed by the in-namespace /bin/sh.
The shim runs under `unshare`'s forked child → does the bind mounts →
forks a subshell that applies rlimits and `exec`s the payload. Applying
rlimits in the subshell (last hop) keeps the memory/tools unconstrained
and constrains only the learner process.
"""
quoted_cmd = " ".join(shlex.quote(part) for part in cmd)
rlimit_prefix = (
f"ulimit -v {limits.memory_bytes // 1024}; " # RLIMIT_AS, KB
f"ulimit -t {limits.cpu_seconds}; " # RLIMIT_CPU, s
f"ulimit -f {limits.file_size_bytes // 512}; " # RLIMIT_FSIZE, 512 blocks
)
return (
"set -eu; "
"mount -t tmpfs tmpfs /tmp; "
f"mkdir -p {IN_NS_WORKDIR}; "
f"mount --bind {shlex.quote(str(workspace))} {IN_NS_WORKDIR}; "
f"cd {IN_NS_WORKDIR}; "
f"exec sh -c {shlex.quote(rlimit_prefix + 'exec ' + quoted_cmd)}"
)
class _Tracked:
"""The process tree + paths for one persistent (telemetry-wired) sandbox."""
def __init__(
self,
*,
helper: asyncio.subprocess.Process,
inner: asyncio.subprocess.Process,
agent: asyncio.subprocess.Process | None,
inner_pid: int, # host pid of the SANDBOXED init (sleep) — ns enter target
host_dir: Path,
workspace: Path,
ns_root: Path,
ns_workdir: Path,
) -> None:
self.helper = helper
self.inner = inner
self.agent = agent
self.inner_pid = inner_pid
self.host_dir = host_dir
self.workspace = workspace
self.ns_root = ns_root
self.ns_workdir = ns_workdir
class UnshareBackend: # satisfies SandboxBackend structurally (Protocol)
"""D-024 backend: namespace subprocesses; persistent tree for task sandboxes."""
def __init__(
self,
unshare_path: str | None = None,
limits: ResourceLimits | None = None, # per-spec override lands in 1-04
nsenter_path: str | None = None,
agent_script: Path | None = None,
) -> None:
self._unshare = unshare_path or shutil.which("unshare") or "unshare"
self._nsenter = nsenter_path or shutil.which("nsenter") or "nsenter"
self._limits = limits or ResourceLimits()
# The stdlib-only capture agent script, copied into each tracked
# workdir's host/ tree so nsenter can reach it inside the sandbox.
# ai_service/sandbox/unshare_backend.py -> parents[2] = apps/ai-service.
self._agent_script = agent_script or (
Path(__file__).resolve().parents[2] / "scripts" / "sandbox-agent.py"
)
# Tracked (persistent) sandboxes by id; pure shell sandboxes are absent.
self._tracked: dict[str, _Tracked] = {}
# -- spawn ------------------------------------------------------------------
async def spawn(self, spec: SandboxSpec) -> SandboxHandle:
"""Lay out the workdir; if `spec.capture_env` is set, start the sandbox.
A spec WITHOUT capture_env keeps REQ-3-001 semantics: spawn only
prepares disk state and each exec forks a fresh (offline) namespace.
A spec WITH capture_env starts the persistent helper/inner tree and the
capture agent, and `handle.pid` carries the agent's host pid.
"""
create_layout(spec)
if not spec.capture_env:
return SandboxHandle(
id=spec.sandbox_id,
pid=None, # no persistent process; each exec forks short-lived PIDs
workdir=spec.workdir,
created_at=datetime.now(UTC),
)
tracked = await self._spawn_tracked(spec)
self._tracked[spec.sandbox_id] = tracked
return SandboxHandle(
id=spec.sandbox_id,
pid=tracked.agent.pid if tracked.agent is not None else tracked.inner_pid,
workdir=spec.workdir,
created_at=datetime.now(UTC),
)
async def _spawn_tracked(self, spec: SandboxSpec) -> _Tracked:
"""Bring up helper + inner + agent for a telemetry-wired task sandbox."""
host_dir = spec.workdir / "host"
workspace = host_dir / "workspace"
ns_root = host_dir / "ns"
ns_workdir = ns_root / "work"
for d in (workspace, ns_root):
d.mkdir(parents=True, exist_ok=True)
# The agent script must live INSIDE the workspace: the inner ns bind
# mounts <host_dir>/workspace -> <ns_root>/work, so only workspace
# content is visible in-namespace at /work.
agent_host_path = workspace / "sandbox-agent.py"
shutil.copyfile(self._agent_script, agent_host_path)
helper = await self._launch_ns(
[
self._unshare,
"--user",
"--map-root-user",
"--mount",
"sh",
"-c",
(
# Isolate ns/ so the inner tmpfs never propagates back to the
# host mount table (make-private is best-effort on this host).
f"mount --bind {shlex.quote(str(ns_root))} {shlex.quote(str(ns_root))}; "
f"mount --make-private {shlex.quote(str(ns_root))} 2>/dev/null; "
f"echo {_HELPER_READY}; exec sleep 3600"
),
],
sentinel=_HELPER_READY,
label="helper",
)
try:
inner = await self._launch_ns(
[
*self._helper_join_argv(helper),
self._unshare,
"--mount",
"--pid",
"--fork",
"--net",
"sh",
"-c",
(
f"mount -t tmpfs tmpfs {shlex.quote(str(ns_root))}; "
f"mkdir -p {shlex.quote(str(ns_workdir))}; "
f"mount --bind {shlex.quote(str(workspace))} "
f"{shlex.quote(str(ns_workdir))}; "
f"echo {_INNER_READY}; exec sleep 3600"
),
],
sentinel=_INNER_READY,
label="inner",
)
except Exception:
await self._reap(helper)
raise
await asyncio.sleep(0) # let the inner child's sleep fork settle
inner_pid = await asyncio.to_thread(self._find_child_pid, inner.pid)
if inner_pid is None:
await self._reap(inner)
await self._reap(helper)
raise SandboxUnavailableError(
f"could not resolve sandboxed init pid for {spec.sandbox_id}"
)
tracked = _Tracked(
helper=helper,
inner=inner,
agent=None,
inner_pid=inner_pid,
host_dir=host_dir,
workspace=workspace,
ns_root=ns_root,
ns_workdir=ns_workdir,
)
if spec.capture_env:
tracked.agent = await self._launch_agent(spec, tracked, agent_host_path)
return tracked
# -- process launch helpers --------------------------------------------------
def _helper_join_argv(self, helper: asyncio.subprocess.Process) -> list[str]:
"""nsenter argv that runs a command inside the helper's user+mount ns."""
if helper.pid is None:
raise SandboxUnavailableError("helper namespace process is not running")
return [
self._nsenter,
"-t",
str(helper.pid),
"-m",
"-U",
"--preserve-credentials",
"--",
]
def _sandbox_join_argv(self, tracked: _Tracked) -> list[str]:
"""nsenter argv that joins the inner sandbox MOUNT namespace (uid 0)."""
return [self._nsenter, "-t", str(tracked.inner_pid), "-m", "--"]
async def _launch_agent(
self, spec: SandboxSpec, tracked: _Tracked, agent_host_path: Path
) -> asyncio.subprocess.Process:
"""Launch the capture agent: joins the sandbox mount ns, NOT pid/net.
The nsenter chain swaps the mount table under the process, so a HOST
cwd/relative path is invalid after the join (observed: python3
resolved ``sandbox-agent.py`` against a stale root ``//`` and
exited rc=2). The launch therefore happens through ``sh -c`` INSIDE
the joined namespace, using only in-namespace absolute paths: the
workspace is bind-mounted at ``<ns_root>/work``, the agent script was
copied into the host workspace, so ``/work/sandbox-agent.py`` exists
after the join. The agent runs ONLINE (joins mount ns only, not the
offline net ns) so it can dial the ai-service WS ingest loopback.
"""
env = dict(spec.capture_env or {})
in_ns_script = f"{tracked.ns_root / 'work' / 'sandbox-agent.py'}"
in_ns_cwd = f"{tracked.ns_root / 'work'}"
launch = f"cd {shlex.quote(in_ns_cwd)} && exec python3 {shlex.quote(in_ns_script)}"
try:
return await asyncio.create_subprocess_exec(
*self._helper_join_argv(tracked.helper),
*self._sandbox_join_argv(tracked),
"sh",
"-c",
launch,
stdin=asyncio.subprocess.DEVNULL,
stdout=asyncio.subprocess.DEVNULL,
stderr=asyncio.subprocess.DEVNULL,
env=env,
)
except FileNotFoundError as exc: # pragma: no cover - env-dependent
raise SandboxUnavailableError("python3 unavailable for capture agent") from exc
async def _launch_ns(
self, argv: list[str], *, sentinel: str, label: str
) -> asyncio.subprocess.Process:
"""Spawn a namespace supervisor and wait for its `sentinel` line."""
proc = await asyncio.create_subprocess_exec(
*argv,
stdout=asyncio.subprocess.PIPE,
stderr=asyncio.subprocess.STDOUT,
)
async def _wait_ready() -> None:
if proc.stdout is None: # pragma: no cover (stdout is a PIPE)
raise SandboxUnavailableError(f"{label} namespace missing stdout pipe")
async for raw in proc.stdout:
if raw.decode(errors="replace").strip() == sentinel:
return
raise SandboxUnavailableError(
f"{label} namespace exited before signalling readiness: {argv[:3]}"
)
try:
await asyncio.wait_for(_wait_ready(), timeout=10.0)
except TimeoutError as exc:
await self._reap(proc)
raise SandboxUnavailableError(
f"{label} namespace never became ready (timeout): {argv[:3]}"
) from exc
except SandboxUnavailableError:
await self._reap(proc)
raise
return proc
@staticmethod
def _find_child_pid(parent_pid: int | None) -> int | None:
"""First direct child of `parent_pid` (the pid-namespaced `sleep`).
The helperunshare shim is inner.pid's parent chain head, but the
SANDBOXED mount/pid namespaces belong to its forked child (the
`sleep`). nsenter must target THAT pid to land inside the sandbox.
Reads /proc directly best-effort, host-local, no subprocess.
"""
if parent_pid is None:
return None
for entry in os.listdir("/proc"):
if not entry.isdigit():
continue
try:
with open(f"/proc/{entry}/stat") as fh:
# ppid is field 4; comm (field 2) may contain spaces, so
# parse relative to the LAST ')'.
rest = fh.read().rsplit(") ", 1)[1].split()
if int(rest[1]) == parent_pid: # state=rest[0], ppid=rest[1]
return int(entry)
except (OSError, IndexError, ValueError):
continue
return None
# -- exec --------------------------------------------------------------------
async def exec(self, handle: SandboxHandle, cmd: list[str]) -> ExecResult:
"""Run `cmd` in the sandbox workspace (cwd = the bound workspace)."""
if not cmd:
raise ValueError("exec requires a non-empty cmd")
tracked = self._tracked.get(handle.id)
if tracked is not None:
return await self._exec_tracked(handle, tracked, cmd)
return await self._exec_fresh(handle, cmd)
async def _exec_fresh(self, handle: SandboxHandle, cmd: list[str]) -> ExecResult:
"""Pure shell sandbox: spawn one fresh offline namespace per exec."""
workspace = handle.workdir / "workspace"
if not workspace.is_dir():
raise SandboxUnavailableError(f"spawn() first: no workspace at {workspace}")
argv = [
self._unshare,
*UNSHARE_ARGS,
"sh",
"-c",
_build_shim(workspace, self._limits, cmd),
]
started = time.monotonic()
proc = await asyncio.create_subprocess_exec(
*argv,
stdin=asyncio.subprocess.DEVNULL,
stdout=asyncio.subprocess.PIPE,
stderr=asyncio.subprocess.PIPE,
)
out, err = await proc.communicate()
return ExecResult(
cmd=cmd,
returncode=proc.returncode if proc.returncode is not None else -1,
stdout=out.decode(errors="replace"),
stderr=err.decode(errors="replace"),
duration_s=time.monotonic() - started,
)
async def _exec_tracked(
self, handle: SandboxHandle, tracked: _Tracked, cmd: list[str]
) -> ExecResult:
"""Task sandbox: join the persistent inner namespace (offline, uid 0).
rlimits apply in the joining subshell so only the payload is limited;
cwd is the bound workspace (`<ns>/work`).
"""
if tracked.inner.returncode is not None:
raise SandboxUnavailableError(
f"sandbox {handle.id} is not running (inner namespace exited)"
)
quoted_cmd = " ".join(shlex.quote(part) for part in cmd)
rlimit_prefix = (
f"ulimit -v {self._limits.memory_bytes // 1024}; "
f"ulimit -t {self._limits.cpu_seconds}; "
f"ulimit -f {self._limits.file_size_bytes // 512}; "
)
shell = (
f"cd {shlex.quote(str(tracked.ns_workdir))}; "
f"{rlimit_prefix}"
f"exec {quoted_cmd}"
)
argv = [
*self._helper_join_argv(tracked.helper),
*self._sandbox_join_argv(tracked),
"sh",
"-c",
shell,
]
started = time.monotonic()
proc = await asyncio.create_subprocess_exec(
*argv,
stdin=asyncio.subprocess.DEVNULL,
stdout=asyncio.subprocess.PIPE,
stderr=asyncio.subprocess.PIPE,
)
out, err = await proc.communicate()
return ExecResult(
cmd=cmd,
returncode=proc.returncode if proc.returncode is not None else -1,
stdout=out.decode(errors="replace"),
stderr=err.decode(errors="replace"),
duration_s=time.monotonic() - started,
)
# -- snapshot / destroy --------------------------------------------------------
async def snapshot(self, handle: SandboxHandle) -> Path:
workspace_root = handle.workdir
tracked = self._tracked.get(handle.id)
if tracked is not None:
# Copy the tracked workspace, not the legacy <workdir>/workspace.
dest_parent = handle.workdir / "snapshots"
dest_parent.mkdir(parents=True, exist_ok=True)
return workdir_snapshot_from_workspace(tracked.workspace, dest_parent)
return workdir_snapshot(workspace_root)
async def destroy(self, handle: SandboxHandle) -> None:
"""Best-effort teardown. Pure shell sandboxes die with their exec; for a
tracked task sandbox reap AGENT INNER HELPER so no capture process
or namespace supervisor outlives the handle (REQ-3-003 lifecycle).
Keeping the workdir is deliberate: snapshots must survive destroy so a
learner's last state can be restored by the manager layer.
"""
tracked = self._tracked.pop(handle.id, None)
if tracked is not None:
# Agent first (it must not flush a "stopped" event into a dead
# sandbox), then the namespace tree. The inner `unshare --fork`
# shim is NOT the namespace init: killing it orphans its child
# (the `sleep` that is PID 1 of the sandbox pid+mnt+net ns),
# which reparents to host init and holds the tmpfs + bind for
# a full hour (observed: ~30 leaked `sleep 3600` after a test
# run). `--kill-child` does not reach it either (util-linux
# 2.38 leaks the same child under this flag combo — the child
# is reparented before unshare's signal handler runs). The
# deterministic kill is SIGKILL on the ns-init's HOST pid,
# which we already track as `tracked.inner_pid` (nsenter uses
# it for exec); the kernel then tears down the namespace with
# its init (no processes remain).
for proc in (tracked.agent, tracked.inner, tracked.helper):
if proc is not None:
await self._reap(proc)
self._kill_pid(tracked.inner_pid)
handle.pid = None
@staticmethod
def _kill_pid(pid: int | None, sig: int = signal.SIGKILL) -> None:
"""Best-effort host-side signal; pid recycled or gone is not an error."""
if pid is None:
return
try:
os.kill(pid, sig)
except (ProcessLookupError, PermissionError):
pass # already dead, or not ours — nothing to do
@staticmethod
async def _reap(proc: asyncio.subprocess.Process) -> None:
"""SIGTERM then SIGKILL, tolerant of an already-dead process."""
if proc.returncode is not None:
return
try:
proc.terminate()
except ProcessLookupError:
return
try:
await asyncio.wait_for(proc.wait(), timeout=5.0)
except TimeoutError:
try:
proc.kill()
except ProcessLookupError:
return
try:
await asyncio.wait_for(proc.wait(), timeout=5.0)
except TimeoutError: # pragma: no cover - SIGKILL always wins
pass
def workdir_snapshot_from_workspace(workspace: Path, snapshots_dir: Path) -> Path:
"""Snapshot helper for tracked sandboxes whose workspace is `<workdir>/host/workspace`
instead of the legacy `<workdir>/workspace` layout."""
dest = snapshots_dir / datetime.now(UTC).strftime("%Y%m%dT%H%M%SZ")
shutil.copytree(workspace, dest, symlinks=False)
return dest
@@ -1,81 +0,0 @@
"""Per-sandbox workdir layout and snapshots (REQ-3-001).
Layout, rooted under `settings.sandbox_dir` (default `apps/ai-service/sandboxes/`):
<sandbox_dir>/<sandbox_id>/
workspace/ bind-mounted into the namespace at /work (learner-writable)
snapshots/ host-side timestamped copies produced by snapshot()
The workspace is the ONLY directory the namespaced process can write that is
also visible on the host. Everything else either stays on the host (see
UnshareBackend's DAC note) or lands in a discarded tmpfs.
"""
from __future__ import annotations
import shutil
from datetime import UTC, datetime
from pathlib import Path
from pydantic import BaseModel, ConfigDict
from ..config import Settings
from .backend import SandboxSpec
class SandboxDir(BaseModel):
"""Concrete paths for one sandbox's on-disk layout."""
model_config = ConfigDict(arbitrary_types_allowed=True)
root: Path
workspace: Path
snapshots: Path
def workspace_path(spec: SandboxSpec) -> Path:
"""Return the workspace path for a sandbox laid out under `spec.workdir`."""
return workspace_path_from_workdir(spec.workdir)
def workspace_path_from_workdir(workdir: Path) -> Path:
"""Workspace path given a sandbox workdir root."""
return workdir / "workspace"
def create_layout(spec: SandboxSpec) -> SandboxDir:
"""Create `<workdir>/{workspace,snapshots}` (parents included, idempotent)."""
layout = SandboxDir(
root=spec.workdir,
workspace=spec.workdir / "workspace",
snapshots=spec.workdir / "snapshots",
)
layout.workspace.mkdir(parents=True, exist_ok=True)
layout.snapshots.mkdir(parents=True, exist_ok=True)
return layout
def snapshot(workdir: Path) -> Path:
"""Recursively copy `<workdir>/workspace` to `<workdir>/snapshots/<utc-ts>/`.
Symlinks are never followed or recreated (`symlinks=False`); a symlink in
the workspace is replaced by the file it points at, so a snapshot can
never retain a host-escape link. Returns the new snapshot directory.
"""
dest = workdir / "snapshots" / datetime.now(UTC).strftime("%Y%m%dT%H%M%SZ")
shutil.copytree(workdir / "workspace", dest, symlinks=False)
return dest
def resolve_sandbox_dir(settings: Settings) -> Path:
"""Resolve `settings.sandbox_dir` (relative → anchored at the app dir)."""
sandbox_dir = settings.sandbox_dir
if sandbox_dir.is_absolute():
return sandbox_dir
return (Path(__file__).resolve().parent.parent / sandbox_dir).resolve()
def spec_for(sandbox_id: str, learner_id: str, settings: Settings) -> SandboxSpec:
"""Build a `SandboxSpec` rooted under the configured sandbox dir."""
workdir = resolve_sandbox_dir(settings) / sandbox_id
return SandboxSpec(sandbox_id=sandbox_id, learner_id=learner_id, workdir=workdir)
@@ -1,21 +0,0 @@
"""Live build telemetry — event models and the TraceStore protocol (REQ-3-003).
Boundary rule (D-027): telemetry/ imports from config only never from
agents/ or api/ (agents call engines through narrow interfaces, never
the reverse; api/ composes stores via DI).
"""
from .ingest import IngestSession, TraceIntegrityMap, telemetry_ingest_endpoint
from .models import EventKind, TelemetryEvent, TraceSpan
from .store import SQLiteTraceStore, TraceStore
__all__ = [
"EventKind",
"IngestSession",
"SQLiteTraceStore",
"TelemetryEvent",
"TraceIntegrityMap",
"TraceSpan",
"TraceStore",
"telemetry_ingest_endpoint",
]
@@ -1,448 +0,0 @@
"""WS ingest protocol for learner telemetry (REQ-3-003, D-026, G-3).
Frame contract trace identity travels as QUERY PARAMS on the WS upgrade
(`WS /v1/telemetry/ingest?learner_id=...&task_id=...&sandbox_id=...`), NOT as
a first init frame. Rationale: the in-sandbox capture agent (Task 2-2-01) is a
stdlib-only RFC6455 client where the URL is the cheapest thing to parametrize
(`NC_INGEST_URL` carries the query string); identity is also visible to the
server BEFORE accept(), so a malformed handshake can be rejected without an
accept/close round-trip. Client messages are then ONE event per JSON text
frame no envelope:
{"seq": 0, "kind": "command", "payload": {...}, "ts": "...",
"sandbox_id": "..."} # learner_id / task_id forbidden (URL owns them)
Server client frames are typed status envelopes:
{"type": "ack_total", "count": N} final flush summary, then close 1000
{"type": "gap_warning", "missing_seqs": [...]} seq skipped ahead
{"type": "event_rejected", "detail": "..."} one frame failed validation
(seq echoed when parseable)
{"type": "event_rejected", "seq": N, "detail": "..."} stored-field rejected (bad kind)
{"type": "flooded", "reason": "cap_exceeded"|"queue_overflow",
"count": N} sent before close(1008)
Keepalive: the server sends an opaque ping frame every `PING_INTERVAL_S` (the
capture agent auto-pongs at the frame layer); a peer that is silent past
`PONG_TIMEOUT_S` is assumed wedged, but the keepalive half only LOGS the
receiver half owns disconnect detection (single-box pilot: TCP EOF is
reliable; an aggressive pong-watchdog would false-positive on loaded boxes).
Flood control (GRILL G-3, BINDING silent drop-oldest is FORBIDDEN):
* per-connection inbound queue bounded at `INBOUND_QUEUE_MAX` frames; on
overflow close code 1008 (policy violation) + trace marked
INCOMPLETE_FLOODED via `TraceIntegrityMap`.
* total events for the (learner, task) exceeding
`Settings.telemetry_max_events_per_task` same 1008 + INCOMPLETE_FLOODED.
`INCOMPLETE_FLOODED` is an integrity signal Proctor/Phase-3 grader read via
`TraceIntegrityMap.is_incomplete()` (the G-4 gate): a flooded trace can
never yield a credential.
Boundary (D-027): telemetry/ never imports agents/ or api/. This module
imports only `fastapi.WebSocket` for the socket type (a protocol surface, not
a DI framework); the session engine below depends only on the TraceStore
protocol + Settings, and api/telemetry.py injects both through plain
parameters.
"""
import asyncio
import contextlib
import json
import logging
from datetime import datetime
from typing import Any, Final
from fastapi import WebSocket, WebSocketDisconnect
from pydantic import BaseModel, ConfigDict, Field, ValidationError
from ..config import Settings
from .models import TelemetryEvent
from .store import TraceStore
logger = logging.getLogger(__name__)
#: WebSocket close code 1008 — policy violation (RFC 6455 §7.4.1).
WS_CLOSE_POLICY_VIOLATION: Final = 1008
#: Bounded inbound queue depth per connection (G-3). Sized for burst-tolerance
#: well above the capture agent's emission rate; overflow is a flood signal,
#: not a backpressure knob.
INBOUND_QUEUE_MAX: Final = 256
PING_INTERVAL_S: Final = 20.0
class InboundEventFrame(BaseModel):
"""Client → server event frame (one TelemetryEvent minus URL-owned ids).
`extra="forbid"`: learner_id/task_id arriving in the frame body is a
contract violation identity comes from the query params only, so a
replayed frame can never lie about which trace it belongs to.
"""
model_config = ConfigDict(extra="forbid")
seq: int = Field(ge=0)
kind: str = Field(min_length=1)
payload: dict[str, Any] = Field(default_factory=dict)
ts: datetime
sandbox_id: str = ""
class TraceIntegrityMap:
"""Integrity flags for traces that can never be graded (G-3/G-4).
Process-local and deliberately small: v0.3 runs ONE ai-service process per
box, and the Phase-3 grader reads this flag through the same DI container
D-019-style in-memory registry precedent (the sandbox handle registry is
the same shape). The flag is terminal within the process: a reconnect
sending legal events does NOT clear it the trace is already untrusted as
grading input. Restarting ai-service resets flags; grading runs against a
live service, and the SQLite trace rows themselves are durable.
All methods are sync: mutation is a dict write, reads are dict lookups
no await needed, so callers from any layer (API handlers, the grader)
don't inherit an async surface for a nanosecond operation.
"""
def __init__(self) -> None:
# (learner_id, task_id) -> machine-readable reason (INCOMPLETE_FLOODED)
self._flags: dict[tuple[str, str], str] = {}
def mark(self, learner_id: str, task_id: str, reason: str) -> None:
"""Set an integrity flag. Presence of the flag is the signal; the
reason is informational (last write wins)."""
self._flags[(learner_id, task_id)] = reason
def clear(self, learner_id: str, task_id: str) -> None:
"""Test seam: reset a flag (production ingest never clears)."""
self._flags.pop((learner_id, task_id), None)
def is_incomplete(self, learner_id: str, task_id: str) -> bool:
"""True when the trace carries ANY terminal integrity flag."""
return (learner_id, task_id) in self._flags
def reason(self, learner_id: str, task_id: str) -> str | None:
"""The flag's reason (INCOMPLETE_FLOODED), or None when unflagged."""
return self._flags.get((learner_id, task_id))
class IngestSession:
"""One WebSocket ingest connection: receive → queue → drain → store.
Two tasks per connection:
* `_receiver` reads frames, validates shape, enqueues (bounded queue,
G-3). Receives never block on SQLite.
* `_drainer` pops frames in arrival order, appends via TraceStore
(idempotent on (learner,task,seq)), emits gap warnings, enforces the
per-trace event cap.
Either task detecting a flood closes the WS with 1008 and marks the trace
INCOMPLETE_FLOODED. The events queue carries `None` as the client-
disconnect sentinel.
- `telemetry_max_events_per_task` is consulted at connect and re-checked
per append against the DURABLE row count (cap compares against stored
events, so a skipped-ahead seq cannot burn budget that was never sent).
Durable count via `TraceStore.count()` (COUNT(*)) a single aggregate
per append, never materializing trace rows (the pre-P7 code read
`len(get_trace(...))` which was O(trace) per event / O() per session).
"""
def __init__(
self,
websocket: WebSocket,
store: TraceStore,
integrity: TraceIntegrityMap,
settings: Settings,
learner_id: str,
task_id: str,
sandbox_id: str,
) -> None:
self._ws = websocket
self._store = store
self._integrity = integrity
# Snapshot of the one setting ingest consults: read once at connect so
# a hot-reloaded Settings object mid-session can't move the cap.
self._max_events = settings.telemetry_max_events_per_task
self.learner_id = learner_id
self.task_id = task_id
self.sandbox_id = sandbox_id
self._queue: asyncio.Queue[InboundEventFrame | None] = asyncio.Queue(
maxsize=INBOUND_QUEUE_MAX
)
self._seen: set[int] = set()
self._next_expected: int | None = None # in-connection monotonic hint
self._received = 0
self._stored = 0
self._deduped = 0
self._rejected = 0
self._flooded = False
self._flood_reason = ""
# -- receive half ----------------------------------------------------------
async def run(self) -> None:
"""Accept, run receiver+drainer, close cleanly. Owns the WS lifecycle."""
await self._ws.accept()
logger.info(
"telemetry ingest connected: %s/%s sandbox=%s",
self.learner_id,
self.task_id,
self.sandbox_id or "(none)",
)
pinger = asyncio.create_task(self._keepalive())
receiver = asyncio.create_task(self._receiver())
drainer = asyncio.create_task(self._drainer())
# First terminal outcome shuts the session down: client disconnect
# (receiver ends) → drainer flushes; drainer ended (clean close after
# flush or a 1008 flood close) → receiver must not linger.
pending: set[asyncio.Task[None]] = {receiver, drainer}
try:
done, pending = await asyncio.wait(
pending, return_when=asyncio.FIRST_COMPLETED
)
if receiver in done and drainer in pending:
try:
await drainer # final flush → sends ack_total, close 1000
finally:
pending.discard(drainer)
finally:
for task in (pinger, *pending):
task.cancel()
with contextlib.suppress(asyncio.CancelledError):
await task
async def _receiver(self) -> None:
"""Read frames; parse+enqueue. Overflow → flood shutdown (G-3).
RuntimeError from receive_text is benign here: it fires when the
socket was closed by the drainer (1008 flood close) while this task
was parked in receive a terminal condition, not a bug.
"""
try:
while True:
raw = await self._ws.receive_text()
frame = self._parse(raw)
if frame is None:
# Rejected frame — keep the connection open; the producer
# gets an event_rejected status frame so a malformed batch
# is visible (and its seq is never stored). Yield so the
# status frame flushes before we block on the next receive.
await self._reject_frame(raw)
await asyncio.sleep(0)
continue
self._received += 1
try:
self._queue.put_nowait(frame)
except asyncio.QueueFull:
# Bounded queue — overflow is a flood, never drop-oldest.
# _trigger_flood closes the socket; fall through to the
# tail so the disconnect sentinel is still enqueued — the
# drainer is never left parked on an empty queue after a
# flood (P7 review: the pre-fix code `return`ed from the
# QueueFull branch WITHOUT the sentinel, leaking the
# session task set — one per flooded trace).
await self._trigger_flood("queue_overflow")
return
except WebSocketDisconnect:
pass
except RuntimeError:
logger.debug(
"ingest receiver: socket already closed (flood path) %s/%s",
self.learner_id,
self.task_id,
)
# Client gone (clean close, drop, or flood close): sentinel unblocks
# the drainer for a final flush. put_nowait can only fail under flood,
# which already terminated the session.
with contextlib.suppress(asyncio.QueueFull):
self._queue.put_nowait(None)
def _parse(self, raw: str) -> InboundEventFrame | None:
"""Validate one frame; None means malformed (caller rejects it)."""
try:
return InboundEventFrame.model_validate_json(raw)
except ValidationError:
return None
async def _reject_frame(self, raw: str) -> None:
"""Malformed envelope: log + event_rejected status frame (never stored)."""
self._rejected += 1
detail = "invalid event frame"
try:
InboundEventFrame.model_validate_json(raw)
except ValidationError as exc:
detail = exc.errors()[0].get("msg", "validation error")
logger.warning(
"telemetry frame rejected: %s/%s: %s", self.learner_id, self.task_id, detail
)
seq: int | None = None
with contextlib.suppress(Exception):
seq = int(json.loads(raw).get("seq")) # best-effort echo for the producer
payload: dict[str, Any] = {"type": "event_rejected", "detail": detail}
if seq is not None:
payload["seq"] = seq
await self._send_json(payload)
# -- drain half --------------------------------------------------------------
async def _drainer(self) -> None:
"""Pop queued frames, append to the store, then close 1000 + summary."""
while True:
frame = await self._queue.get()
if frame is None: # disconnect sentinel → flush complete
await self._send_json(
{
"type": "ack_total",
"count": self._stored,
"deduped": self._deduped,
"rejected": self._rejected,
}
)
with contextlib.suppress(RuntimeError, WebSocketDisconnect):
await self._ws.close(code=1000)
return
await self._append(frame)
async def _append(self, frame: InboundEventFrame) -> None:
# Precedence: a trace already flagged INCOMPLETE_FLOODED is terminal —
# the connection that triggered it is being torn down, and any stray
# queued frames must not resurrect the trace's intake.
if self._integrity.is_incomplete(self.learner_id, self.task_id):
await self._trigger_flood("already_flagged")
return
# Per-trace cap (G-3): checked against the DURABLE row count so a
# reconnect resumes the budget instead of resetting it, and a
# skipped-ahead seq cannot burn budget that was never sent.
if self._flood_breached():
await self._trigger_flood("cap_exceeded")
return
# TelemetryEvent's @validates hooks fire on CONSTRUCTION (setattr), so
# the try must wrap building the model too — an unknown kind raises
# before `store.append` is ever reached.
event: TelemetryEvent
before = self._store.latest_seq(self.learner_id, self.task_id)
try:
event = TelemetryEvent(
learner_id=self.learner_id,
task_id=self.task_id,
seq=frame.seq,
kind=frame.kind,
payload=frame.payload,
ts=frame.ts,
sandbox_id=frame.sandbox_id or self.sandbox_id,
)
self._store.append(event)
except ValueError as exc: # unknown kind / invalid field
self._rejected += 1
await self._send_json(
{"type": "event_rejected", "seq": frame.seq, "detail": str(exc)}
)
return
after = self._store.latest_seq(self.learner_id, self.task_id)
if after == before and frame.seq in self._seen:
self._deduped += 1 # at-least-once retry; stored once (idempotent)
else:
self._stored += 1
self._seen.add(frame.seq)
await self._check_gap(frame.seq)
# SQLite appends are sync and fast; on a burst the drainer can hold
# the loop between receives. Yield so the WS writer flushes the close
# and the pinger/interleave stay live under the eventlet-free portal.
await asyncio.sleep(0)
def _flood_breached(self) -> bool:
"""True when this append would exceed the per-trace event budget."""
# Durable count (NOT latest_seq+1 — a skipped-ahead seq must not burn
# un-sent events' budget) via COUNT(*): never materialize the trace
# per append (P7 review — the old len(get_trace(...)) built every row
# object per event, O(trace) per append / O(n²) per session).
durable = self._store.count(learner_id=self.learner_id, task_id=self.task_id)
return durable >= self._max_events
async def _check_gap(self, incoming_seq: int) -> None:
"""Seq skipped ahead → log + per-connection gap_warning status frame."""
if self._next_expected is not None and incoming_seq > self._next_expected:
missing = list(range(self._next_expected, incoming_seq))
logger.warning(
"telemetry gap: %s/%s missing seqs %s (arrived seq=%d)",
self.learner_id,
self.task_id,
missing,
incoming_seq,
)
await self._send_json({"type": "gap_warning", "missing_seqs": missing})
if self._next_expected is None or incoming_seq >= self._next_expected:
self._next_expected = incoming_seq + 1
# -- flood + keepalive ------------------------------------------------------
async def _trigger_flood(self, reason: str) -> None:
"""G-3: 1008 close + INCOMPLETE_FLOODED mark. Exactly once."""
if self._flooded:
return
self._flooded = True
self._flood_reason = reason
self._integrity.mark(self.learner_id, self.task_id, "INCOMPLETE_FLOODED")
logger.warning(
"telemetry flood: %s/%s reason=%s — closing 1008, trace marked "
"INCOMPLETE_FLOODED (G-3; Proctor/grade gate will refuse it)",
self.learner_id,
self.task_id,
reason,
)
await self._send_json(
{"type": "flooded", "reason": reason, "count": self._received}
)
with contextlib.suppress(RuntimeError, WebSocketDisconnect):
await self._ws.close(
code=WS_CLOSE_POLICY_VIOLATION,
reason=f"telemetry flood control (G-3): {reason}",
)
async def _keepalive(self) -> None:
"""Protocol-level ping on an interval (agent auto-pongs at frame level).
A send failure means the socket is already gone the receiver half
independently surfaces the disconnect; we just stop pinging.
"""
while True:
await asyncio.sleep(PING_INTERVAL_S)
try:
await self._ws.send_bytes(b"\x89ping-nextcraft")
except (RuntimeError, WebSocketDisconnect):
return
async def _send_json(self, payload: dict[str, Any]) -> None:
"""Best-effort status frame; the socket may already be gone."""
with contextlib.suppress(RuntimeError, WebSocketDisconnect):
await self._ws.send_json(payload)
async def telemetry_ingest_endpoint(
websocket: WebSocket,
learner_id: str,
task_id: str,
store: TraceStore,
integrity: TraceIntegrityMap,
settings: Settings,
sandbox_id: str = "",
) -> None:
"""Engine entry: build the session and run it. api/telemetry.py calls this
with query params + app.state services already resolved this signature
is deliberately Depends-free (telemetry/ never knows FastAPI DI exists).
"""
session = IngestSession(
websocket=websocket,
store=store,
integrity=integrity,
settings=settings,
learner_id=learner_id,
task_id=task_id,
sandbox_id=sandbox_id,
)
await session.run()
@@ -1,108 +0,0 @@
"""Telemetry event record — the row the trace store persists (REQ-3-003, D-027).
One model serves both as the JSON payload sent by producers and as the SQLite
row schema. `payload` is stored as a JSON column (native JSONB on Postgres
no migration-time shape change, D-027).
Field contract (consumed by the trace store and the grader):
learner_id non-empty learner identifier.
task_id non-empty task/session identifier; trace identity is the
(learner_id, task_id) pair.
seq sequence number per trace, >= 0. Monotonicity per
(learner, task) is enforced by the store (Task 2-1-02);
this model only rejects negative seqs.
kind event discriminator: command | file_diff | run_result |
test_result | activity | stdin | stdout.
payload free-form JSON detail blob.
ts envelope timestamp (UTC); monotonicity enforced at ingest.
sandbox_id originating sandbox ("" for non-sandbox sources).
Boundary (D-027): telemetry/ never imports agents/ or api/.
"""
from datetime import datetime
from typing import Any, Literal
from pydantic import BaseModel, ConfigDict, Field
from sqlalchemy import JSON, Index, String
from sqlalchemy.orm import validates
from sqlmodel import Field as SQLField
from sqlmodel import SQLModel
EventKind = Literal[
"command",
"file_diff",
"run_result",
"test_result",
"activity",
"stdin",
"stdout",
]
_EVENT_KINDS: frozenset[str] = frozenset(EventKind.__args__)
class TelemetryEvent(SQLModel, table=True):
"""A single durable telemetry event; (learner_id, task_id, seq) is PK.
Constraint enforcement uses SQLAlchemy `@validates` hooks: sqlmodel
0.0.42's metaclass drops pydantic `Field(ge=...)`/`field_validator`
constraints for table models (the decorators register but never make it
into the core schema), while `@validates` fires on every attribute set
construction included and raises ValueError on violation. seq >= 0 plus
a VARCHAR kind column keep the DB shape Postgres-ready (D-027).
"""
__tablename__ = "telemetry_event"
# PK columns already produce a unique index; this secondary index covers
# trace reads ordered by seq without depending on the PK column order
# (Postgres migration target D-027).
__table_args__ = (Index("ix_telemetry_event_trace", "learner_id", "task_id"),)
learner_id: str = SQLField(primary_key=True)
task_id: str = SQLField(primary_key=True)
seq: int = SQLField(primary_key=True)
# Bare Literal annotations crash sqlmodel<=0.0.42's column inference
# (issubclass(TypeAlias, Enum)); an explicit sa_type + the validates hook
# below gives the same contract: VARCHAR column, Literal-rejected values.
kind: EventKind = SQLField(sa_type=String)
# JSON column: stored as TEXT on SQLite, native JSONB on Postgres (D-027).
payload: dict[str, Any] = SQLField(default_factory=dict, sa_type=JSON)
ts: datetime
sandbox_id: str = SQLField(default="")
@validates("learner_id", "task_id")
def _ids_non_empty(self, key: str, value: str) -> str:
if not value:
raise ValueError(f"{key} must be a non-empty identifier")
return value
@validates("seq")
def _seq_non_negative(self, key: str, value: int) -> int:
if value < 0:
raise ValueError("seq must be >= 0 (monotonicity is the store's job)")
return value
@validates("kind")
def _kind_is_known(self, key: str, value: str) -> str:
if value not in _EVENT_KINDS:
raise ValueError(f"unknown event kind: {value!r}")
return value
class TraceSpan(BaseModel):
"""Derived view: the ordered event trace for one (learner_id, task_id).
NOT a table materialized by the store from persisted TelemetryEvents
(grader/Lab consume this shape; replay order is the seq column).
"""
model_config = ConfigDict(frozen=True)
learner_id: str = Field(min_length=1)
task_id: str = Field(min_length=1)
events: tuple[TelemetryEvent, ...] = ()
@property
def latest_seq(self) -> int:
"""Highest seq in the span; -1 when empty (store convention)."""
return self.events[-1].seq if self.events else -1
@@ -1,218 +0,0 @@
"""TraceStore — telemetry persistence protocol + SQLite implementation (REQ-3-003, D-027).
Postgres-migration-ready (D-027): the protocol is the only surface the API /
grader layers touch; swapping SQLiteTraceStore for a Postgres-backed
implementation must not change call sites. The `telemetry_event` table uses
only portable column types (str / int / datetime / JSON), so the same SQLModel
schema stands up unchanged on Postgres.
Ingest is at-least-once: duplicates carry the same (learner_id, task_id, seq)
idempotency key, so `append` with a triplet that is already stored is a no-op.
The pair (learner_id, task_id) identifies a trace; `seq` numbers events in it
starting at 0.
Concurrency (a-3): the engine enables WAL + synchronous=NORMAL and a busy
timeout at connection time, so the ingest writer and grader readers do not hit
`database is locked` on the single-box pilot.
Boundary: `telemetry/` never imports `agents/` / `api/` and has no FastAPI
dependency.
"""
import logging
import sqlite3
from collections.abc import Iterator
from contextlib import contextmanager
from datetime import UTC, datetime
from pathlib import Path
from typing import Any, Protocol
import sqlalchemy as sa
from sqlmodel import Session, SQLModel, create_engine, select
from ..config import Settings
from .models import TelemetryEvent
logger = logging.getLogger(__name__)
class TraceStore(Protocol):
"""Persistence contract for ordered per-learner task trace streams.
Implemented by SQLiteTraceStore (v0.3, D-027); a Postgres implementation
must satisfy the same surface.
"""
def append(self, event: TelemetryEvent) -> None:
"""Store one event. IDEMPOTENT on (learner_id, task_id, seq):
at-least-once ingest retries with the same triplet are deduped
(stored once), not rejected. Later events must not overwrite an
existing row.
"""
...
def get_trace(self, learner_id: str, task_id: str) -> list[TelemetryEvent]:
"""All stored events for the trace, ordered by seq ascending.
Detached from any DB session safe to pass across layers. Empty list
when the trace has no events.
"""
...
def gaps(self, learner_id: str, task_id: str) -> list[int]:
"""Missing seqs in 0..latest for the trace ([0,2,3] stored -> [1])."""
...
def latest_seq(self, learner_id: str, task_id: str) -> int:
"""Highest stored seq for the trace; -1 when no events exist."""
...
def count(self, learner_id: str, task_id: str) -> int:
"""Number of stored events for the trace (COUNT(*), never
materializes rows the ingest cap consults this per append, so
an O(trace) implementation would make ingest O() per session).
"""
...
def list_tasks(self, learner_id: str) -> list[str]:
"""Distinct task_ids with at least one event for the learner."""
...
def close(self) -> None:
"""Release DB connections. Store must not be used after close."""
...
def _sqlite_connect(dbapi_connection: sqlite3.Connection, _: object) -> None:
"""Per-connection pragma setup (a-3).
journal_mode=WAL readers never block the single writer.
synchronous=NORMAL safe in WAL mode, avoids full fsync-per-commit.
busy_timeout=5000 retry briefly under contention instead of
`OperationalError: database is locked`.
"""
cursor = dbapi_connection.cursor()
cursor.execute("PRAGMA journal_mode=WAL")
cursor.execute("PRAGMA synchronous=NORMAL")
cursor.execute("PRAGMA busy_timeout=5000")
cursor.close()
def _as_utc(ts: datetime) -> datetime:
"""Timestamps travel as naive datetime on SQLite, tz-aware elsewhere.
SQLite (via SQLModel) drops tzinfo; Postgres TIMESTAMP WITH TIME ZONE
keeps it. Normalizing on the read/write boundary makes the store's
contract tz-aware UTC regardless of the backend (D-027).
"""
if ts.tzinfo is None:
return ts.replace(tzinfo=UTC) # naive UTC read-side: label as UTC
return ts.astimezone(UTC)
class SQLiteTraceStore:
"""SQLite-backed TraceStore (SQLModel). First real persistence (D-027)."""
def __init__(self, db_path: Path | None = None) -> None:
self._db_path: Path = db_path if db_path is not None else Settings().db_path
self._engine = create_engine(f"sqlite:///{self._db_path}")
sa.event.listen(self._engine, "connect", _sqlite_connect)
SQLModel.metadata.create_all(self._engine)
@contextmanager
def _session(self) -> Iterator[Session]:
# expire_on_commit=False: ORM objects returned from `append`'s
# IntegrityError path stay usable without a refresh round-trip.
with Session(self._engine, expire_on_commit=False) as session:
yield session
def append(self, event: TelemetryEvent) -> None:
# INSERT-if-absent via PK: sqlite3 raises IntegrityError on a
# duplicate (learner_id, task_id, seq); swallow it — the row is
# already stored, which is the dedup contract for at-least-once
# ingest. `session.merge` would upsert instead; wrong semantics here.
with self._session() as session:
try:
session.add(event)
session.commit()
except sa.exc.IntegrityError:
session.rollback()
logger.debug(
"trace event dedup: %s/%s seq=%d already stored",
event.learner_id,
event.task_id,
event.seq,
)
def get_trace(self, learner_id: str, task_id: str) -> list[TelemetryEvent]:
with self._session() as session:
stmt = (
select(TelemetryEvent)
.where(TelemetryEvent.learner_id == learner_id)
.where(TelemetryEvent.task_id == task_id)
.order_by(TelemetryEvent.seq)
)
results = session.exec(stmt).all()
# Detach from the session: callers must not depend on open-session
# ORM magic (lazy loads fail once the session is closed).
for row in results:
row.ts = _as_utc(row.ts)
session.expunge(row)
return list(results)
def _stored_seqs(self, learner_id: str, task_id: str) -> list[int]:
with self._session() as session:
stmt = (
select(TelemetryEvent.seq)
.where(TelemetryEvent.learner_id == learner_id)
.where(TelemetryEvent.task_id == task_id)
.order_by(TelemetryEvent.seq)
)
# sqlmodel scalar select: rows are plain ints, not 1-tuples.
return [int(seq) for seq in session.exec(stmt).all()]
def gaps(self, learner_id: str, task_id: str) -> list[int]:
seqs = self._stored_seqs(learner_id, task_id)
if not seqs:
return []
present = set(seqs)
# seq numbering starts at 0; a gap is any seq in 0..latest not stored.
return [seq for seq in range(seqs[-1] + 1) if seq not in present]
def latest_seq(self, learner_id: str, task_id: str) -> int:
with self._session() as session:
stmt = (
select(sa.func.max(TelemetryEvent.seq))
.where(TelemetryEvent.learner_id == learner_id)
.where(TelemetryEvent.task_id == task_id)
)
latest: Any = session.exec(stmt).one()
return -1 if latest is None else int(latest)
def count(self, learner_id: str, task_id: str) -> int:
# COUNT(*) at the DB — no row materialization. The ingest flood cap
# calls this per append (telemetry/ingest._flood_breached); the
# docstring-free body keeps it obvious what the query shape is.
with self._session() as session:
stmt = (
select(sa.func.count(TelemetryEvent.seq))
.where(TelemetryEvent.learner_id == learner_id)
.where(TelemetryEvent.task_id == task_id)
)
total: Any = session.exec(stmt).one()
return int(total or 0)
def list_tasks(self, learner_id: str) -> list[str]:
with self._session() as session:
stmt = (
select(TelemetryEvent.task_id)
.where(TelemetryEvent.learner_id == learner_id)
.distinct()
.order_by(TelemetryEvent.task_id)
)
# sqlmodel scalar select: rows are plain strs, not 1-tuples.
return [str(task_id) for task_id in session.exec(stmt).all()]
def close(self) -> None:
self._engine.dispose()
@@ -1,31 +0,0 @@
"""Per-learner variant task generation — templates, generator, VariantStore (REQ-3-005).
Boundary rule (D-027): variants/ is an engine module it never imports
api/; its ONLY agents/ dependency is the module-direct
agents.structured import in generator.py (the sanctioned shared D-020
structured defense, same exception as grading/engine.py). api/ composes
the generator and store via DI; store.py imports config only.
CO-ORDINATION NOTE (ADD, don't REMOVE — same convention as grading/):
This __init__.py is a minimal placeholder created by the VariantStore
task (4-1-02). The templates task (4-1-01) owns this file's final shape
when templates.py lands, ADD its exports alongside these; do not
remove the store exports below.
Wave status: store.py (VariantRecord, VariantStore, SQLiteVariantStore)
landed in Wave 1 (task 4-1-02); templates.py is Wave 1 task 4-1-01;
generator.py is Wave 2 (4-2-01).
"""
from .store import SQLiteVariantStore, VariantRecord, VariantStore
from .templates import TEMPLATES, TaskTemplate, get_template, template_for_competency
__all__ = [
"SQLiteVariantStore",
"TEMPLATES",
"TaskTemplate",
"VariantRecord",
"VariantStore",
"get_template",
"template_for_competency",
]
@@ -1,136 +0,0 @@
"""Seeded per-learner variant generator (D-029, REQ-3-005).
Contract (binding, from GRILL + PLAN Must-Haves):
- REPRODUCIBLE: seed = sha256(template_id|learner_id|milestone); the same
(template, learner) re-derives the same seed, params, task_id and the
second generate() call is a cache hit with NO LLM call.
- DISTINCT: different learners on the same template draw different params
(the sampler is seeded per-learner) and receive distinct statements.
- NEVER BLOCKS ON THE LLM: the deterministic skeleton render
(`template.render(params)`) is a complete, valid statement; if the D-020
LLM render fails after its bounded retry, the fallback is used and
because the fallback is exactly `template.render(seed-params)`, it is
auditable from the persisted seed + params without a provenance column.
- AUDITABLE: seed + params + statement persist via VariantStore
(insert-only first-wins) the proctoring cross-check path.
- FAIR (a-5): slot draws change the scenario, never the difficulty; the
template's rubric anchors bound the expected effort envelope, so every
variant of one template is held to the same bar.
"""
from __future__ import annotations
import hashlib
from datetime import UTC, datetime
from typing import TYPE_CHECKING
from pydantic import BaseModel, ConfigDict, Field
from ..agents.structured import StructuredOutputError, structured_completion
from ..llm.types import Message
from ..prompts.variant import VARIANT_SCHEMA_HINT, render_variant_prompt
from .store import VariantRecord
from .templates import TaskTemplate, get_template
if TYPE_CHECKING: # pragma: no cover
from ..llm.base import LLMProvider
from .store import VariantStore
MILESTONE = "v0.3"
class RenderedVariant(BaseModel):
"""D-20-validated LLM render output (statement only — files come from the template)."""
model_config = ConfigDict(extra="forbid")
statement: str = Field(min_length=20)
class UnknownTemplateError(ValueError):
"""Raised when generate() is asked for a template id not in the library."""
def derive_seed(template_id: str, learner_id: str, milestone: str = MILESTONE) -> str:
"""Reproducible per-(template, learner, milestone) seed (D-029)."""
return hashlib.sha256(f"{template_id}|{learner_id}|{milestone}".encode()).hexdigest()
def derive_task_id(seed: str) -> str:
"""Deterministic grading/telemetry task key from the seed (16 hex chars)."""
return f"task-{seed[:16]}"
class VariantGenerator:
"""Seeded instantiation over the template library. DI: store + provider."""
def __init__(self, store: VariantStore, provider: LLMProvider, model: str) -> None:
self._store = store
self._provider = provider
self._model = model
async def generate(self, learner_id: str, template_id: str) -> VariantRecord:
template = get_template(template_id)
if template is None:
raise UnknownTemplateError(f"no task template with id {template_id!r}")
# Cache: D-029 reproducibility — same (learner, template) is served
# from the store with no LLM call.
cached = self._store.get(learner_id, template_id)
if cached is not None:
return cached
seed_hex = derive_seed(template_id, learner_id)
task_id = derive_task_id(seed_hex)
params = template.sample_params(_seed_int(seed_hex))
_validate_params(template, params)
statement = await self._render(template, params)
record = VariantRecord(
learner_id=learner_id,
task_id=task_id,
template_id=template_id,
seed=seed_hex,
params=dict(params),
statement=statement,
starter_files=dict(template.starter_files),
created_at=datetime.now(UTC),
)
self._store.save(record)
return record
async def _render(self, template: TaskTemplate, params: dict[str, str | int]) -> str:
"""LLM render via D-020; deterministic fallback never blocks task work.
Provenance note: unlike grades, variants carry no `model` column
the deterministic fallback is exactly `template.render(params)`,
re-derivable from the persisted seed + params, so a fallback render is
auditable without storing provenance (the seed IS the provenance).
"""
messages: list[Message] = render_variant_prompt(template, params)
try:
rendered = await structured_completion(
self._provider,
messages,
model=self._model,
schema=RenderedVariant,
schema_hint=VARIANT_SCHEMA_HINT,
)
except StructuredOutputError:
# Deterministic fallback: the skeleton + seeded slots is already a
# complete statement, re-derivable from the persisted seed.
return template.render(params)
return rendered.statement
def _seed_int(seed_hex: str) -> int:
"""Stable int for random.Random from the hex seed."""
return int(seed_hex[:16], 16)
def _validate_params(template: TaskTemplate, params: dict[str, str | int]) -> None:
"""Defense in depth: every sampled value must be schema-valid (a-5)."""
for slot in template.slots:
value = params.get(slot.name)
if value is None or not slot.validate_value(value):
raise ValueError(f"sampled params invalid for slot {slot.name!r}: {value!r}")
@@ -1,311 +0,0 @@
"""VariantStore — variant persistence protocol + SQLite implementation (REQ-3-005, D-027).
Postgres-migration-ready (D-027): the protocol is the only surface the
variant generator and API layers touch; swapping SQLiteVariantStore for a
Postgres-backed implementation must not change call sites. The
`variant_record` table uses only portable column types (str / JSON /
datetime), so the same SQLModel schema stands up unchanged on Postgres.
Insert-only, NOT upsert: (learner_id, template_id) is the variant identity
and the FIRST generation is authoritative reproducibility (D-029) means
the seed re-derives the same variant, so the generator's cache path serves
`get` instead of saving again. `save` is a plain INSERT; a duplicate pair
raises sqlalchemy.exc.IntegrityError to the caller (documented behavior).
`task_id` is unique too it is the grading/telemetry trace key, so a
trace or grade can never silently join to a different variant. Both
rejections are deliberate: overwriting a stored variant would swap a
learner's graded task underneath its trace and grade (audit corruption).
Contrast TraceStore.append (dedup-keep-first, swallowed at-least-once
ingest) and GradeStore.save (upsert-latest-wins a regrade is
latest-state); this store is the third contract of the D-027 family.
Concurrency (a-3): the store enables WAL + synchronous=NORMAL and a busy
timeout at connection time, so a generation writer and API readers do not
hit `database is locked` on the single-box pilot.
`created_at` contract: callers stamp UTC (datetime.now(UTC)); SQLite
stores it naive and the read paths re-label it tz-aware UTC (same
boundary normalization as TelemetryEvent.ts / GradeRecord.created_at, so
the contract holds on any backend).
Boundary (D-027): `variants/` never imports `agents/` / `api/`; this
module imports config only.
"""
import logging
import sqlite3
from collections.abc import Iterator
from contextlib import contextmanager
from datetime import UTC, datetime
from pathlib import Path
from typing import Any, Protocol
import sqlalchemy as sa
from sqlalchemy import JSON, Index, UniqueConstraint
from sqlalchemy.orm import validates
from sqlmodel import Field, Session, SQLModel, create_engine, select
from ..config import Settings
logger = logging.getLogger(__name__)
class VariantRecord(SQLModel, table=True):
"""A persisted task variant; (learner_id, template_id) is the PK — first wins.
Written once by the variant generator (Task 4-2-01), read by the API
layer and proctoring cross-checks through the VariantStore protocol.
Constraint enforcement mirrors TelemetryEvent / GradeRecord: sqlmodel
0.0.42's metaclass drops pydantic constraints on table models, so
SQLAlchemy `@validates` hooks enforce instead and the column types
stay Postgres-ready (D-027).
Field contract:
learner_id non-empty learner identifier (same id space as
traces and grades).
template_id non-empty task template identifier; variant
identity is the (learner_id, template_id) pair
the pair the generator caches on (exactly one
variant per learner per template).
task_id non-empty, GLOBALLY unique task identifier; the
grading/telemetry trace key (the (learner_id,
task_id) pair TraceStore / GradeStore key on),
stamped at generation so a variant's trace and
grade join back to it exactly once.
seed non-empty variant seed (D-029); derived from
(template_id, learner_id, milestone) so the
variant is reproducible and auditable.
params typed parameter-slot values the generator filled;
JSON dict. An empty dict is legal (a slotless
template).
statement non-empty rendered task statement shown to the
learner (distinct per learner by construction,
REQ-3-005).
starter_files workspace scaffold: filename -> file content;
JSON dict. An empty dict is legal (no scaffold).
created_at UTC generation timestamp.
"""
__tablename__ = "variant_record"
# The composite PK covers (learner_id, template_id) point lookups; the
# unique task_id covers get_by_task (the grading/telemetry join path);
# the two secondary indexes cover list_for_learner / list_by_template
# ordered by created_at without a sort step (Postgres target D-027).
__table_args__ = (
UniqueConstraint("task_id", name="uq_variant_record_task_id"),
Index("ix_variant_record_learner_created", "learner_id", "created_at"),
Index("ix_variant_record_template_created", "template_id", "created_at"),
)
learner_id: str = Field(primary_key=True)
template_id: str = Field(primary_key=True)
task_id: str
seed: str
# JSON columns: stored as TEXT on SQLite, native JSONB on Postgres (D-027).
params: dict[str, Any] = Field(default_factory=dict, sa_type=JSON)
statement: str
starter_files: dict[str, Any] = Field(default_factory=dict, sa_type=JSON)
created_at: datetime
@validates("learner_id", "template_id", "task_id")
def _ids_non_empty(self, key: str, value: str) -> str:
if not value:
raise ValueError(f"{key} must be a non-empty identifier")
return value
@validates("seed")
def _seed_non_empty(self, key: str, value: str) -> str:
if not value:
raise ValueError(f"{key} must be a non-empty seed string")
return value
@validates("statement")
def _statement_non_empty(self, key: str, value: str) -> str:
if not value:
raise ValueError(f"{key} must be a non-empty statement string")
return value
class VariantStore(Protocol):
"""Persistence contract for reproducible per-learner task variants.
Implemented by SQLiteVariantStore (v0.3, D-027); a Postgres
implementation must satisfy the same surface.
"""
def save(self, variant: VariantRecord) -> None:
"""Persist a new variant. INSERT-ONLY on (learner_id, template_id):
the FIRST generated variant is authoritative (reproducibility,
D-029); a duplicate pair raises sqlalchemy.exc.IntegrityError to
the caller the generator serves cached variants via `get`
instead of saving again. `task_id` is unique too: claiming an
existing trace key for a different variant is equally rejected.
NOT upsert; contrast GradeStore.save (latest-wins) and
TraceStore.append (dedup-keep-first, swallowed).
"""
...
def get(self, learner_id: str, template_id: str) -> VariantRecord | None:
"""The learner's stored variant for the template; None when none
exists. Detached from any DB session safe to pass across layers.
"""
...
def get_by_task(self, task_id: str) -> VariantRecord | None:
"""The variant owning the task key (the grading/telemetry join
path); None when none exists. Detached from any DB session.
"""
...
def list_for_learner(self, learner_id: str) -> list[VariantRecord]:
"""All stored variants for the learner, ordered by created_at
ascending (chronological; task_id breaks same-instant ties).
Empty list when the learner has none.
"""
...
def list_by_template(self, template_id: str) -> list[VariantRecord]:
"""All stored variants generated from the template — one row per
learner ordered by created_at ascending (chronological;
learner_id breaks same-instant ties). Empty list when the
template has none. The proctoring cross-check path (seed params
per learner) reads through this.
"""
...
def close(self) -> None:
"""Release DB connections. Store must not be used after close."""
...
def _sqlite_connect(dbapi_connection: sqlite3.Connection, _: object) -> None:
"""Per-connection pragma setup (a-3). Mirrors telemetry/grading stores.
journal_mode=WAL readers never block the single writer.
synchronous=NORMAL safe in WAL mode, avoids full fsync-per-commit.
busy_timeout=5000 retry briefly under contention instead of
`OperationalError: database is locked`.
"""
cursor = dbapi_connection.cursor()
cursor.execute("PRAGMA journal_mode=WAL")
cursor.execute("PRAGMA synchronous=NORMAL")
cursor.execute("PRAGMA busy_timeout=5000")
cursor.close()
def _as_utc(ts: datetime) -> datetime:
"""Timestamps travel as naive datetime on SQLite, tz-aware elsewhere.
SQLite (via SQLModel) drops tzinfo; Postgres TIMESTAMP WITH TIME ZONE
keeps it. Normalizing on the read path makes the store's contract
tz-aware UTC regardless of the backend (D-027).
"""
if ts.tzinfo is None:
return ts.replace(tzinfo=UTC) # naive UTC read-side: label as UTC
return ts.astimezone(UTC)
class SQLiteVariantStore:
"""SQLite-backed VariantStore (SQLModel). Third protocol-wrapped store
of the D-027 family (first: SQLiteTraceStore, second: SQLiteGradeStore).
"""
def __init__(self, db_path: Path | None = None) -> None:
self._db_path: Path = db_path if db_path is not None else Settings().db_path
self._engine = create_engine(f"sqlite:///{self._db_path}")
sa.event.listen(self._engine, "connect", _sqlite_connect)
SQLModel.metadata.create_all(self._engine)
@contextmanager
def _session(self) -> Iterator[Session]:
# expire_on_commit=False: identical session behavior to the other
# D-027 stores. save() never commits on the error path and the read
# paths never commit, but a uniform flag across the family keeps
# their detachment guarantees from diverging.
with Session(self._engine, expire_on_commit=False) as session:
yield session
def save(self, variant: VariantRecord) -> None:
# Plain INSERT, no merge: overwriting a stored variant would swap a
# learner's graded task underneath its trace and grade (audit
# corruption), so a duplicate identity is a race or bug to SURFACE,
# not paper over. The generator's cache path (get before generate)
# makes duplicate saves a programming error, not a normal flow.
# The trace store swallows its IntegrityError (dedup is the
# contract there); the grade store merges (latest-wins is the
# contract there); this store re-raises (first-wins is the
# contract here).
with self._session() as session:
try:
session.add(variant)
session.commit()
except sa.exc.IntegrityError:
session.rollback()
logger.debug(
"variant insert rejected (identity already stored): "
"learner=%s template=%s task=%s",
variant.learner_id,
variant.template_id,
variant.task_id,
)
raise
logger.debug(
"variant saved: %s/%s task=%s seed=%s",
variant.learner_id,
variant.template_id,
variant.task_id,
variant.seed,
)
def get(self, learner_id: str, template_id: str) -> VariantRecord | None:
with self._session() as session:
record = session.get(VariantRecord, (learner_id, template_id))
if record is None:
return None
record.created_at = _as_utc(record.created_at)
# Detach from the session: callers must not depend on
# open-session ORM magic (lazy loads fail once it closes).
session.expunge(record)
return record
def get_by_task(self, task_id: str) -> VariantRecord | None:
with self._session() as session:
stmt = select(VariantRecord).where(VariantRecord.task_id == task_id)
record = session.exec(stmt).first()
if record is None:
return None
record.created_at = _as_utc(record.created_at)
session.expunge(record)
return record
def list_for_learner(self, learner_id: str) -> list[VariantRecord]:
with self._session() as session:
stmt = (
select(VariantRecord)
.where(VariantRecord.learner_id == learner_id)
# Chronological; task_id is a deterministic tie-break for
# variants stamped within the same instant.
.order_by(VariantRecord.created_at, VariantRecord.task_id)
)
results = session.exec(stmt).all()
for row in results:
row.created_at = _as_utc(row.created_at)
session.expunge(row)
return list(results)
def list_by_template(self, template_id: str) -> list[VariantRecord]:
with self._session() as session:
stmt = (
select(VariantRecord)
.where(VariantRecord.template_id == template_id)
# Chronological; learner_id is a deterministic tie-break.
.order_by(VariantRecord.created_at, VariantRecord.learner_id)
)
results = session.exec(stmt).all()
for row in results:
row.created_at = _as_utc(row.created_at)
session.expunge(row)
return list(results)
def close(self) -> None:
self._engine.dispose()
@@ -1,305 +0,0 @@
"""Task template library for seeded variant generation (D-029, REQ-3-005).
A `TaskTemplate` binds a competency (D-021-aligned corpus ID), a statement
skeleton with `{slot}` placeholders, typed `ParameterSlot`s, difficulty-
normalization rubric anchors (the expected feature envelope that bounds
variant fairness in the a-5 envelope test grader-prompt shipment is the
tracked P4 follow-up; grading is variant-blind today), and starter-file
scaffolds served into the sandbox workdir (wired in P6).
Slot sampling is PURE CODE: `random.Random(seed)` over typed slots fully
reproducible for a given seed, independent of the LLM. The LLM only renders
the seeded slot values into the statement skeleton (D-020 defense).
"""
from __future__ import annotations
import random
import re
from typing import Literal
from pydantic import BaseModel, ConfigDict, Field, field_validator
SlotType = Literal["enum", "int_range", "string_set"]
class ParameterSlot(BaseModel):
"""One typed fill-in for a statement skeleton."""
model_config = ConfigDict(frozen=True)
name: str = Field(min_length=1)
type: SlotType
values: list[str] = Field(default_factory=list) # enum/string_set options
lo: int | None = None # int_range bounds
hi: int | None = None
@field_validator("values")
@classmethod
def _values_nonempty_for_enums(cls, v: list[str], info) -> list[str]:
if info.data.get("type") in ("enum", "string_set") and not v:
raise ValueError(f"slot {info.data.get('name')!r} needs values")
return v
def sample(self, rng: random.Random) -> str | int:
"""Deterministic sample from the seeded RNG. Validated after sampling."""
if self.type == "enum" or self.type == "string_set":
return rng.choice(self.values)
if self.type == "int_range":
lo = self.lo if self.lo is not None else 0
hi = self.hi if self.hi is not None else lo
if hi < lo:
raise ValueError(f"slot {self.name!r}: hi < lo")
return rng.randint(lo, hi)
raise ValueError(f"unsupported slot type: {self.type!r}")
def validate_value(self, value: str | int) -> bool:
"""Is `value` schema-valid for this slot? (params JSON gate, a-5.)"""
if self.type in ("enum", "string_set"):
return isinstance(value, str) and value in self.values
if self.type == "int_range":
lo = self.lo if self.lo is not None else 0
hi = self.hi if self.hi is not None else lo
return isinstance(value, int) and lo <= value <= hi
return False
class RubricAnchors(BaseModel):
"""Difficulty-normalization anchors for the grader (a-5).
Expected FEATURE ENVELOPE (digest-space): the expected effort band
for this template, so two variants of one template are held to the
same bar regardless of which slot values a learner drew. The a-5
envelope test (tests/variants/test_generator.py) binds variants to
these bands in code, and since Phase 4 (MH#4) — the grading engine
ships this envelope into the grader prompt
(grading/engine._anchors_context) and stamps the variant seed on the
GradeRecord, so the anchors gate variant fairness in BOTH tests and
the live rubric.
"""
model_config = ConfigDict(frozen=True)
expected_edit_count_band: tuple[int, int]
expected_min_test_runs: int
expected_error_fix_cycles_band: tuple[int, int]
notes: str = ""
class TaskTemplate(BaseModel):
"""A reusable task shape; variants instantiate it per learner."""
model_config = ConfigDict(frozen=True)
id: str = Field(min_length=1)
competency_id: str = Field(min_length=1) # D-021 corpus alignment
title: str
statement_skeleton: str = Field(min_length=1) # {slot} placeholders
slots: list[ParameterSlot] = Field(min_length=1)
rubric_anchors: RubricAnchors
starter_files: dict[str, str] = Field(default_factory=dict) # path -> content
test_command: str
@field_validator("statement_skeleton")
@classmethod
def _skeleton_placeholders(cls, v: str) -> str:
if "{" not in v or "}" not in v:
raise ValueError("statement_skeleton needs at least one {slot}")
return v
def render(self, params: dict[str, str | int]) -> str:
"""Fill the skeleton with validated params."""
for slot in self.slots:
if slot.name not in params:
raise ValueError(f"missing param for slot {slot.name!r}")
if not slot.validate_value(params[slot.name]):
raise ValueError(f"invalid value for slot {slot.name!r}: {params[slot.name]!r}")
return self.statement_skeleton.format(**params)
def sample_params(self, seed: int) -> dict[str, str | int]:
"""Seeded, reproducible, schema-valid slot values (pure code)."""
rng = random.Random(seed)
return {slot.name: slot.sample(rng) for slot in self.slots}
# --- Template library (v0.3 initial set) --------------------------------------
# Competency IDs are D-021-aligned with the Python corpus
# (ai_service/corpus/learner_context.py) and the TS mock-data layer
# (packages/mock-data/competency-stacks.ts: deterministic cid() scheme).
TEMPLATES: dict[str, TaskTemplate] = {
"tpl-llm-judge": TaskTemplate(
id="tpl-llm-judge",
competency_id="stack-orchestration-c007",
title="Build an LLM-as-Judge Evaluator",
statement_skeleton=(
"Build a small LLM-as-judge evaluator for {domain} answers. "
"The judge must score each answer on {criterion} using a 0-4 scale, "
"return structured JSON, and handle at least {edge_cases} edge-case "
"answer classes (empty, off-topic, adversarial). Include a tiny "
"repro test set of at least {test_size} examples and print a summary "
"table of scores."
),
slots=[
ParameterSlot(
name="domain",
type="enum",
values=["customer-support", "code-review", "summarization", "tutoring"],
),
ParameterSlot(
name="criterion",
type="enum",
values=["factual-accuracy", "helpfulness", "safety", "completeness"],
),
ParameterSlot(name="edge_cases", type="int_range", lo=2, hi=4),
ParameterSlot(name="test_size", type="int_range", lo=3, hi=8),
],
rubric_anchors=RubricAnchors(
expected_edit_count_band=(3, 25),
expected_min_test_runs=2,
expected_error_fix_cycles_band=(0, 4),
notes="Slot draw changes the SCENARIO, not the engineering depth.",
),
starter_files={
"README.md": (
"# LLM-as-Judge Evaluator\n\n"
"Implement `judge.py`:\n"
"- `score(answer: str) -> dict` — 0-4 on the named criterion\n"
"- structured JSON output (schema below)\n"
"- edge-case classes handled explicitly\n"
"- `pytest` must pass\n"
),
"judge.py": "def score(answer: str) -> dict:\n raise NotImplementedError\n",
"test_judge.py": "def test_placeholder():\n assert True\n",
},
test_command="pytest -q",
),
"tpl-guardrail-schema": TaskTemplate(
id="tpl-guardrail-schema",
competency_id="stack-orchestration-c008",
title="Schema Guardrail Pipeline",
statement_skeleton=(
"Implement an output-validation guardrail for a model returning "
"{entity} records. Validate against a typed schema with {field_count} "
"required fields, coerce or reject {failure_mode} failures, and emit "
"a fallback response for invalid payloads. Cover with at least "
"{test_size} unit tests including malformed JSON."
),
slots=[
ParameterSlot(
name="entity",
type="enum",
values=["user-profile", "job-posting", "candidate", "invoice"],
),
ParameterSlot(
name="failure_mode",
type="enum",
values=["strict-reject", "coerce-when-safe"],
),
ParameterSlot(name="field_count", type="int_range", lo=4, hi=8),
ParameterSlot(name="test_size", type="int_range", lo=4, hi=10),
],
rubric_anchors=RubricAnchors(
expected_edit_count_band=(3, 30),
expected_min_test_runs=2,
expected_error_fix_cycles_band=(0, 5),
notes="All slot draws land in the same engineering band.",
),
starter_files={
"README.md": (
"# Schema Guardrail\n\nImplement `guardrail.py`:\n"
"- `validate(payload: dict) -> dict | Fallback`\n"
"- required-field checks, failure policy, fallback emission\n"
),
"guardrail.py": "def validate(payload: dict):\n raise NotImplementedError\n",
"test_guardrail.py": "def test_placeholder():\n assert True\n",
},
test_command="pytest -q",
),
"tpl-rag-chunker": TaskTemplate(
id="tpl-rag-chunker",
competency_id="stack-orchestration-c005",
title="RAG Chunking Strategy",
statement_skeleton=(
"Implement a document chunker for {doc_type} retrieval. Support "
"{strategy} chunking with a target size of ~{chunk_size} tokens, "
"preserve {invariant} across chunk boundaries, and evaluate overlap "
"quality with at least {test_size} fixture documents."
),
slots=[
ParameterSlot(
name="doc_type",
type="enum",
values=["technical-docs", "legal-contracts", "transcripts"],
),
ParameterSlot(
name="strategy",
type="enum",
values=["fixed-window", "semantic-boundary", "hybrid"],
),
ParameterSlot(
name="invariant",
type="enum",
values=["code-block-integrity", "section-headers", "sentence-completeness"],
),
ParameterSlot(name="chunk_size", type="int_range", lo=200, hi=800),
ParameterSlot(name="test_size", type="int_range", lo=3, hi=6),
],
rubric_anchors=RubricAnchors(
expected_edit_count_band=(4, 35),
expected_min_test_runs=2,
expected_error_fix_cycles_band=(0, 6),
notes="Strategy draw changes implementation shape, not depth.",
),
starter_files={
"README.md": (
"# RAG Chunker\n\nImplement `chunker.py`:\n"
"- `chunk(text: str) -> list[str]`\n- invariant preserved\n- tests green\n"
),
"chunker.py": "def chunk(text: str) -> list[str]:\n raise NotImplementedError\n",
"test_chunker.py": "def test_placeholder():\n assert True\n",
},
test_command="pytest -q",
),
}
_KNOWN_COMPETENCY_IDS: set[str] = {
# D-021: mirrored from ai_service/corpus/learner_context.py — the Python
# source of truth for stack-orchestration competencies used by v0.2 agents.
"stack-orchestration-c001",
"stack-orchestration-c002",
"stack-orchestration-c003",
"stack-orchestration-c004",
"stack-orchestration-c005",
"stack-orchestration-c007",
"stack-orchestration-c008",
"stack-orchestration-c011",
"stack-designer-c001",
"stack-designer-c002",
"stack-safety-c021",
}
def get_template(template_id: str) -> TaskTemplate | None:
return TEMPLATES.get(template_id)
def template_for_competency(competency_id: str) -> list[TaskTemplate]:
return [t for t in TEMPLATES.values() if t.competency_id == competency_id]
def validate_competency_binding() -> None:
"""All templates must bind to known D-021 corpus competency IDs."""
for t in TEMPLATES.values():
if t.competency_id not in _KNOWN_COMPETENCY_IDS:
raise ValueError(
f"template {t.id!r} binds unknown competency {t.competency_id!r}"
)
def slots_pattern_ok(skeleton: str, slots: list[ParameterSlot]) -> bool:
"""Every {placeholder} in the skeleton has a matching slot and vice versa."""
placeholders = set(re.findall(r"\{([a-z_][a-z0-9_]*)\}", skeleton))
slot_names = {s.name for s in slots}
return placeholders == slot_names
@@ -1,35 +0,0 @@
"""VoiceProvider protocol (D-030, REQ-3-006) — mirrors the LLMProvider seam.
Two implementations in v0.3:
- MockVoiceProvider deterministic canned transcripts + canned tone WAV
chunks + scripted failure modes (tests + no-key default; tests NEVER call
a real voice API).
- browser descriptor not a provider but a FALLBACK HINT: the web client
selects browser-native SpeechRecognition/speechSynthesis when the server
reports no real voice backend.
OpenAIAudioProvider (real server STT/TTS over OpenAI-compatible
/audio/transcriptions + /audio/speech) is INTENTIONALLY NOT BUILT in v0.3
deferred to v0.4 with KYC, when there is a real key and real users
(GRILL CUT-1 / G-7). This protocol is its future drop-in seam.
Boundary: `voice/` never imports `agents/` or `api/`.
"""
from ai_service.voice.base import (
TranscriptSegment,
VoiceDescriptor,
VoiceProvider,
)
from ai_service.voice.browser import BROWSER_FALLBACK_DESCRIPTOR
from ai_service.voice.factory import voice_provider_from_settings
from ai_service.voice.mock import MockVoiceProvider
__all__ = [
"BROWSER_FALLBACK_DESCRIPTOR",
"MockVoiceProvider",
"TranscriptSegment",
"VoiceDescriptor",
"VoiceProvider",
"voice_provider_from_settings",
]
-62
View File
@@ -1,62 +0,0 @@
"""VoiceProvider protocol + shared voice contracts (D-030, REQ-3-006).
Mirrors the LLMProvider seam (D-014 pattern): a narrow protocol the Examiner
agent and the defense API compose via DI, with a deterministic mock and a
browser-fallback descriptor. No network in this module concrete providers
live in their own modules and are selected by factory/config.
"""
from __future__ import annotations
from collections.abc import AsyncIterator
from typing import Literal, Protocol, runtime_checkable
from pydantic import BaseModel, ConfigDict, Field
VoiceRole = Literal["examiner", "learner"]
class TranscriptSegment(BaseModel):
"""One STT result: the transcribed text + timing metadata."""
model_config = ConfigDict(frozen=True)
text: str = Field(min_length=1)
language: str = "en"
duration_ms: int | None = None
confidence: float | None = Field(default=None, ge=0.0, le=1.0)
class VoiceDescriptor(BaseModel):
"""Capability descriptor served to the web client (D-030).
The assessment UI reads this to decide HOW the learner speaks/hears:
- `mode="server"` server-side STT/TTS (v0.4 real provider seam)
- `mode="browser"` browser-native SpeechRecognition/speechSynthesis
- `mode="mock"` deterministic no-op path (tests / no-key dev)
The descriptor never contains secrets only capability hints.
"""
model_config = ConfigDict(frozen=True)
mode: Literal["server", "browser", "mock"]
sr_available: bool
tts_available: bool
hint: str = ""
@runtime_checkable
class VoiceProvider(Protocol):
"""The voice port (D-030): STT in, TTS out. Never imports agents/api."""
async def transcribe(self, audio: bytes, fmt: str) -> TranscriptSegment:
"""STT: audio bytes (fmt: 'wav' | 'webm' | 'mp3') → transcript."""
...
def synthesize(self, text: str, voice: str = "default") -> AsyncIterator[bytes]:
"""TTS: text -> async byte chunks (audio stream).
Implementations may be async generators (async-def + yield) the
consumer contract is `async for chunk in provider.synthesize(text)`.
"""
...
@@ -1,32 +0,0 @@
"""Browser-native fallback descriptor (D-030, CUT-1 / G-7, REQ-3-006).
v0.3 has NO real server STT/TTS (deferred to v0.4 with KYC/keys GRILL
CUT-1). When the factory selects `browser` mode, the defense endpoints return
this descriptor and the WEB CLIENT performs SpeechRecognition + speechSynthesis
natively; the server persists text turns as usual.
"""
from __future__ import annotations
from .base import VoiceDescriptor
BROWSER_FALLBACK_DESCRIPTOR = VoiceDescriptor(
mode="browser",
sr_available=True,
tts_available=True,
hint=(
"No server voice backend configured. Use browser-native "
"SpeechRecognition for STT and speechSynthesis for TTS; send the "
"transcribed text to POST /v1/defense/{id}/answer ({text} form)."
),
)
MOCK_DESCRIPTOR = VoiceDescriptor(
mode="mock",
sr_available=True,
tts_available=True,
hint=(
"Deterministic mock voice (tests / no-key dev). Server STT/TTS "
"endpoints serve canned responses; real server STT/TTS lands in v0.4."
),
)
@@ -1,506 +0,0 @@
"""DefenseStore — oral-defense persistence: protocol + SQLite impl (REQ-3-006, D-027).
FOURTH protocol-wrapped store of the D-027 family and the first spanning
TWO related tables: `defense_record` (the defense session + integrity
signals) and `defense_turn` (the ordered examiner/learner transcript,
FK defense_record.id).
Postgres-migration-ready (D-027): the protocol is the only surface the
Examiner pipeline (task 5-2-01) and the defense endpoints (task 5-3-01)
touch; swapping SQLiteDefenseStore for a Postgres implementation must
not change call sites. Both tables use only portable column types
(str / int / datetime / JSON), so the same SQLModel schema stands up
unchanged on Postgres.
Save semantics where this sits among the D-027 stores (each has a
deliberately different contract):
TraceStore.append dedup-keep-first; IntegrityError SWALLOWED
(at-least-once event ingest).
GradeStore.save upsert-latest-wins (a regrade is latest-state).
VariantStore.save insert-only first-wins; IntegrityError RAISED
(reproducibility; a duplicate is a bug).
DefenseStore a LIFECYCLE store:
start() insert-only; a duplicate id raises
(a defense id is minted once per session).
append_turn() insert-only per (defense_id, seq); a duplicate
seq raises AND an unknown defense_id raises (FK
enforced) a transcript turn must never silently
vanish (it is the integrity/grading input) nor
attach to a defense that does not exist.
finalize() targeted UPDATE (status finished; finished_at +
integrity_signals JSON). Unknown id None
(documented below). Re-finalize overwrites
signals + finished_at latest-wins, mirroring
GradeStore.save: a recomputed verdict replaces
the previous one wholesale.
append_turn does NOT police status (turns after finalize are a
sequencing bug for the endpoints to prevent, task 5-3-01): the store
enforces DATA integrity (FK + PK + non-empty), not workflow.
integrity_signals (A-109): JSON dict on the record long pauses,
off-scope cadence markers and friends, computed by the Examiner over
turn metadata and persisted by finalize for the Proctor/Mentor feed.
An empty dict is legal (defense not finished, or a clean defense).
Concurrency (a-3): WAL + synchronous=NORMAL + busy timeout at
connection time (mirrors the other D-027 stores), PLUS foreign_keys=ON
this is the family's first real foreign key and it is actually
enforced on SQLite, matching Postgres's native behavior (D-027 parity).
`created_at` / `ts` contract: callers stamp UTC (datetime.now(UTC));
SQLite stores them naive and read paths re-label tz-aware UTC (same
boundary normalization as TelemetryEvent.ts / GradeRecord.created_at,
so the contract holds on any backend).
Boundary (D-027): `voice/` never imports `agents/` / `api/`; this
module imports config only.
"""
import logging
import sqlite3
from collections.abc import Iterator
from contextlib import contextmanager
from datetime import UTC, datetime
from pathlib import Path
from typing import Any, Literal, Protocol
import sqlalchemy as sa
from sqlalchemy import JSON, Index, String
from sqlalchemy.orm import validates
from sqlmodel import Field, Session, SQLModel, create_engine, select
from ..config import Settings
logger = logging.getLogger(__name__)
DefenseStatus = Literal["in_progress", "finished"]
_DEFENSE_STATUSES: frozenset[str] = frozenset(DefenseStatus.__args__)
TurnRole = Literal["examiner", "learner"]
_TURN_ROLES: frozenset[str] = frozenset(TurnRole.__args__)
class DefenseRecord(SQLModel, table=True):
"""A persisted oral-defense session; id is the PK.
Written by the defense endpoints (task 5-3-01) through the
DefenseStore protocol; read back by the endpoints, the Examiner
pipeline and the Proctor/Mentor feeds. Constraint enforcement
mirrors TelemetryEvent / GradeRecord / VariantRecord: sqlmodel
0.0.42's metaclass drops pydantic constraints on table models, so
SQLAlchemy `@validates` hooks enforce instead and the column types
stay Postgres-ready (D-027).
Field contract:
id non-empty defense identifier, minted once
per session (a duplicate start raises).
learner_id non-empty learner identifier (same id space
as traces, grades and variants).
task_id non-empty task identifier; the defense
defends the submitted work for this trace
key ((learner_id, task_id) joins to the
trace/grade/variant the defense is about).
status in_progress | finished; the STORE owns the
transition: start() forces in_progress,
finalize() sets finished. Validated.
integrity_signals A-109 signal dict (long pauses, off-scope
cadence markers, ...); {} until finalize;
JSON column. An empty dict is legal.
created_at UTC start timestamp.
finished_at UTC finalize timestamp; None while in
progress.
`turns` (property): the seq-ordered DefenseTurn transcript, attached
ONLY by DefenseStore.get(); records from list_for_learner carry
turns == [] call get() for a full transcript.
"""
__tablename__ = "defense_record"
# The id PK covers point lookups; this secondary index covers
# list_for_learner ordered by created_at without a sort step
# (Postgres migration target D-027).
__table_args__ = (
Index("ix_defense_record_learner_created", "learner_id", "created_at"),
)
id: str = Field(primary_key=True)
learner_id: str
task_id: str
# Bare Literal annotations crash sqlmodel<=0.0.42's column inference
# (issubclass(TypeAlias, Enum)); an explicit sa_type + the validates
# hook below give the same contract: VARCHAR column, Literal-rejected
# values (same pattern as TelemetryEvent.kind).
status: DefenseStatus = Field(default="in_progress", sa_type=String)
# JSON column: stored as TEXT on SQLite, native JSONB on Postgres (D-027).
integrity_signals: dict[str, Any] = Field(default_factory=dict, sa_type=JSON)
created_at: datetime
finished_at: datetime | None = Field(default=None)
@property
def turns(self) -> list["DefenseTurn"]:
"""Seq-ordered transcript; [] unless attached by get().
Table models reject ad-hoc attributes (pydantic __setattr__
raises on non-fields), so the store stashes the detached turn
list via object.__setattr__ and this read-only property surfaces
it. The returned list is a copy caller mutations cannot
corrupt the stash.
"""
return list(self.__dict__.get("_turns", []))
@validates("id", "learner_id", "task_id")
def _ids_non_empty(self, key: str, value: str) -> str:
if not value:
raise ValueError(f"{key} must be a non-empty identifier")
return value
@validates("status")
def _status_is_known(self, key: str, value: str) -> str:
if value not in _DEFENSE_STATUSES:
raise ValueError(f"unknown defense status: {value!r}")
return value
class DefenseTurn(SQLModel, table=True):
"""One examiner/learner dialogue turn; (defense_id, seq) is the PK.
Rows are append-only transcript entries written through
DefenseStore.append_turn. seq numbers the dialogue within one
defense starting at 0; monotonic assignment is the endpoints' job
(task 5-3-01), this model only rejects negatives the same split
as TelemetryEvent.seq (model rejects < 0, store owns ordering).
Field contract:
defense_id non-empty; FK defense_record.id. ENFORCED on
SQLite via foreign_keys=ON (first real FK in the
D-027 family; Postgres enforces FKs natively, so
this keeps the backends equivalent, D-027).
seq turn index within the defense, >= 0. (defense_id,
seq) is the PK: a duplicate raises instead of
silently overwriting the transcript is the
integrity/grading input, a vanishing turn is
audit corruption.
role examiner | learner (who spoke). Validated.
text non-empty utterance text (examiner question, or
STT output for learner answers).
ts UTC utterance timestamp.
latency_ms per-turn pipeline latency in ms (STT + LLM TTFT +
TTS, A-109); int or None. Populated by the
endpoints (task 5-4-01); None allowed here the
store persists, it does not measure.
created_at UTC row-write timestamp.
"""
__tablename__ = "defense_turn"
# The composite PK (defense_id, seq) doubles as the covering index
# for the per-defense seq-ordered read in get() — no secondary index
# needed (contrast defense_record's learner-listing index).
defense_id: str = Field(foreign_key="defense_record.id", primary_key=True)
seq: int = Field(primary_key=True)
role: TurnRole = Field(sa_type=String)
text: str
ts: datetime
latency_ms: int | None = Field(default=None)
created_at: datetime
@validates("defense_id")
def _defense_id_non_empty(self, key: str, value: str) -> str:
if not value:
raise ValueError(f"{key} must be a non-empty identifier")
return value
@validates("seq")
def _seq_non_negative(self, key: str, value: int) -> int:
if value < 0:
raise ValueError("seq must be >= 0 (ordering is the endpoints' job)")
return value
@validates("role")
def _role_is_known(self, key: str, value: str) -> str:
if value not in _TURN_ROLES:
raise ValueError(f"unknown turn role: {value!r}")
return value
@validates("text")
def _text_non_empty(self, key: str, value: str) -> str:
if not value:
raise ValueError(f"{key} must be a non-empty utterance string")
return value
@validates("latency_ms")
def _latency_non_negative(self, key: str, value: int | None) -> int | None:
# None is legal (not yet instrumented); a NEGATIVE latency is
# nonsense and surfaces as a construction error.
if value is not None and value < 0:
raise ValueError("latency_ms must be >= 0 or None")
return value
class DefenseStore(Protocol):
"""Persistence contract for oral-defense sessions + transcripts.
Implemented by SQLiteDefenseStore (v0.3, D-027); a Postgres
implementation must satisfy the same surface.
"""
def start(self, defense: DefenseRecord) -> DefenseRecord:
"""Insert a new defense. INSERT-ONLY: a duplicate id raises
sqlalchemy.exc.IntegrityError (a defense id is minted once per
session surfacing, not swallowing, mirrors VariantStore).
The store owns the lifecycle: status is forced to "in_progress"
and finished_at to None, whatever the caller passed only
finalize() may move a defense to finished. Returns the stored
record, detached from any DB session.
"""
...
def append_turn(self, defense_id: str, turn: DefenseTurn) -> DefenseTurn:
"""Insert one transcript turn, ordered by (defense_id, seq).
turn.defense_id MUST equal the defense_id argument a mismatch
raises ValueError (the defense identity must never be
ambiguous). A duplicate (defense_id, seq) raises
IntegrityError; an unknown defense_id raises IntegrityError
(FK enforced). Does NOT police status sequencing turns vs
finalize is the endpoints' job (task 5-3-01). Returns the
stored turn, detached.
"""
...
def finalize(
self, defense_id: str, integrity_signals: dict[str, Any]
) -> DefenseRecord | None:
"""Seal the defense: status → "finished", finished_at = now(UTC),
integrity_signals stored as JSON. UNKNOWN defense_id None
(documented choice: the API layer maps it to 404 without an
exception dance; contrast start/append_turn where IntegrityError
IS the contract those are inserts, this is an update on a key
the caller may legitimately not hold). Re-finalize overwrites
signals + finished_at: latest-wins, mirroring GradeStore.save
(a recomputed verdict replaces the previous one wholesale).
Returns the updated record, detached, WITHOUT turns get() is
the with-turns path.
"""
...
def get(self, defense_id: str) -> DefenseRecord | None:
"""The defense with its FULL transcript (turns in seq order,
detached) and integrity signals; None when it does not exist.
Safe to pass across layers no open-session ORM magic.
"""
...
def list_for_learner(self, learner_id: str) -> list[DefenseRecord]:
"""All stored defenses for the learner, ordered by created_at
ascending (chronological; id breaks same-instant ties), WITHOUT
turns records carry turns == []; call get() for a transcript.
Empty list when the learner has none.
"""
...
def close(self) -> None:
"""Release DB connections. Store must not be used after close."""
...
def _sqlite_connect(dbapi_connection: sqlite3.Connection, _: object) -> None:
"""Per-connection pragma setup (a-3). Mirrors the other D-027 stores.
journal_mode=WAL readers never block the single writer.
synchronous=NORMAL safe in WAL mode, avoids full fsync-per-commit.
busy_timeout=5000 retry briefly under contention instead of
`OperationalError: database is locked`.
foreign_keys=ON NEW vs the family: defense_turn is the first
real FK among the D-027 stores; SQLite leaves
FKs OFF by default while Postgres enforces them
natively, so the pragma keeps the backends
equivalent (D-027 parity).
"""
cursor = dbapi_connection.cursor()
cursor.execute("PRAGMA journal_mode=WAL")
cursor.execute("PRAGMA synchronous=NORMAL")
cursor.execute("PRAGMA busy_timeout=5000")
cursor.execute("PRAGMA foreign_keys=ON")
cursor.close()
def _as_utc(ts: datetime) -> datetime:
"""Timestamps travel as naive datetime on SQLite, tz-aware elsewhere.
SQLite (via SQLModel) drops tzinfo; Postgres TIMESTAMP WITH TIME ZONE
keeps it. Normalizing on the read path makes the store's contract
tz-aware UTC regardless of the backend (D-027).
"""
if ts.tzinfo is None:
return ts.replace(tzinfo=UTC) # naive UTC read-side: label as UTC
return ts.astimezone(UTC)
class SQLiteDefenseStore:
"""SQLite-backed DefenseStore (SQLModel). Fourth protocol-wrapped
store of the D-027 family (first: SQLiteTraceStore, second:
SQLiteGradeStore, third: SQLiteVariantStore) and the first spanning
two related tables.
"""
def __init__(self, db_path: Path | None = None) -> None:
self._db_path: Path = db_path if db_path is not None else Settings().db_path
self._engine = create_engine(f"sqlite:///{self._db_path}")
sa.event.listen(self._engine, "connect", _sqlite_connect)
SQLModel.metadata.create_all(self._engine)
@contextmanager
def _session(self) -> Iterator[Session]:
# expire_on_commit=False: identical session behavior to the other
# D-027 stores. start/append_turn return the caller's instance
# after commit and get/finalize return rows expunged mid-session;
# a uniform flag across the family keeps their detachment
# guarantees from diverging.
with Session(self._engine, expire_on_commit=False) as session:
yield session
def start(self, defense: DefenseRecord) -> DefenseRecord:
# The store owns the lifecycle: a defense is BORN in_progress and
# only finalize() may move it to finished. A smuggled "finished"
# status is normalized away, not rejected — the insert itself
# stays insert-only, and a duplicate id raises to the caller
# (mirroring VariantStore: the id is minted once per session).
defense.status = "in_progress"
defense.finished_at = None
with self._session() as session:
try:
session.add(defense)
session.commit()
except sa.exc.IntegrityError:
session.rollback()
logger.debug("defense start rejected (id already stored): %s", defense.id)
raise
logger.debug(
"defense started: %s learner=%s task=%s",
defense.id,
defense.learner_id,
defense.task_id,
)
return defense
def append_turn(self, defense_id: str, turn: DefenseTurn) -> DefenseTurn:
# The explicit defense_id argument is the defense identity for
# this write; a turn object claiming another defense is a
# programming error — surface it before touching the DB.
if turn.defense_id != defense_id:
raise ValueError(
f"turn.defense_id {turn.defense_id!r} does not match the "
f"defense_id argument {defense_id!r}"
)
with self._session() as session:
try:
session.add(turn)
session.commit()
except sa.exc.IntegrityError:
# Two possible causes, both surfaced, neither swallowed:
# duplicate (defense_id, seq) PK — a transcript turn must
# never silently vanish; unknown defense_id — the FK
# (foreign_keys=ON) rejects the orphan.
session.rollback()
logger.debug(
"defense turn rejected (duplicate (defense_id, seq) "
"or unknown defense_id): defense=%s seq=%s",
defense_id,
turn.seq,
)
raise
logger.debug(
"defense turn appended: %s seq=%d role=%s",
defense_id,
turn.seq,
turn.role,
)
return turn
def finalize(
self, defense_id: str, integrity_signals: dict[str, Any]
) -> DefenseRecord | None:
# A None signals blob would break the read contract (signals are
# a dict, {} until finalize); reject before writing.
if not isinstance(integrity_signals, dict):
raise ValueError(
"integrity_signals must be a JSON-object dict, got "
f"{type(integrity_signals).__name__}"
)
with self._session() as session:
record = session.get(DefenseRecord, defense_id)
if record is None:
# Documented unknown-id behavior: None, not a raise — the
# defense endpoints map this to 404. Contrast start() /
# append_turn(), where IntegrityError IS the contract.
return None
# Latest-wins re-finalize, mirroring GradeStore.save: a
# recomputed verdict (fresh signals) replaces the stored one
# wholesale; status just stays finished.
record.status = "finished"
record.finished_at = datetime.now(UTC)
record.integrity_signals = integrity_signals
session.commit()
record.created_at = _as_utc(record.created_at)
if record.finished_at is not None:
record.finished_at = _as_utc(record.finished_at)
# Detach from the session: callers must not depend on
# open-session ORM magic (lazy loads fail once it closes).
session.expunge(record)
logger.debug(
"defense finalized: %s signals=%s", defense_id, sorted(integrity_signals)
)
return record
def get(self, defense_id: str) -> DefenseRecord | None:
with self._session() as session:
record = session.get(DefenseRecord, defense_id)
if record is None:
return None
record.created_at = _as_utc(record.created_at)
if record.finished_at is not None:
record.finished_at = _as_utc(record.finished_at)
stmt = (
select(DefenseTurn)
.where(DefenseTurn.defense_id == defense_id)
.order_by(DefenseTurn.seq)
)
turns = session.exec(stmt).all()
for turn in turns:
turn.ts = _as_utc(turn.ts)
turn.created_at = _as_utc(turn.created_at)
# Detach each turn: the transcript must be usable once
# the session closes (no lazy-load magic).
session.expunge(turn)
session.expunge(record)
# Table models reject ad-hoc attributes (pydantic __setattr__
# raises on non-fields), so the seq-ordered transcript is
# stashed via object.__setattr__ and surfaced through the
# read-only `turns` property. Rows are detached either way —
# safe to pass across layers.
object.__setattr__(record, "_turns", list(turns))
return record
def list_for_learner(self, learner_id: str) -> list[DefenseRecord]:
with self._session() as session:
stmt = (
select(DefenseRecord)
.where(DefenseRecord.learner_id == learner_id)
# Chronological; id is a deterministic tie-break for
# defenses stamped within the same instant.
.order_by(DefenseRecord.created_at, DefenseRecord.id)
)
results = session.exec(stmt).all()
for row in results:
row.created_at = _as_utc(row.created_at)
if row.finished_at is not None:
row.finished_at = _as_utc(row.finished_at)
# Turns are deliberately NOT loaded here: the list feed
# (Proctor/Mentor) needs session headers, not full
# transcripts — get() is the with-turns path.
session.expunge(row)
return list(results)
def close(self) -> None:
self._engine.dispose()
@@ -1,37 +0,0 @@
"""Voice provider factory (D-030, REQ-3-006).
`AI_VOICE_PROVIDER = browser | mock` (default: mock the no-key path is
first-class). The real server provider (`openai-audio`) is a v0.4 seam and
is REJECTED here with a clear error naming the deferral, so a stale env var
can't silently pretend a real backend exists.
"""
from __future__ import annotations
from ..config import Settings
from .base import VoiceProvider
from .mock import MockVoiceProvider
class UnknownVoiceProviderError(ValueError):
"""Raised for a provider name outside the v0.3 contract."""
def voice_provider_from_settings(settings: Settings) -> VoiceProvider:
"""Select the voice provider by settings (env `AI_VOICE_PROVIDER`)."""
name = (settings.voice_provider or "mock").strip().lower()
if name == "mock":
return MockVoiceProvider()
if name == "browser":
# Browser mode is a CLIENT-side capability: the server composes the
# same MockVoiceProvider (typed fallback answers still work; the UI
# uses the descriptor for mic/speech). See browser.py.
return MockVoiceProvider()
if name in ("openai-audio", "openai", "server"):
raise UnknownVoiceProviderError(
"real server STT/TTS (OpenAIAudioProvider) is deferred to v0.4 "
"(GRILL CUT-1 / G-7): set AI_VOICE_PROVIDER=mock or browser"
)
raise UnknownVoiceProviderError(
f"unknown AI_VOICE_PROVIDER {name!r}: use 'mock' or 'browser'"
)
-96
View File
@@ -1,96 +0,0 @@
"""Deterministic MockVoiceProvider (D-030, REQ-3-006).
Canned transcripts (scripted per test via queue) + canned 1kHz-tone WAV bytes
+ scripted failure modes. Two identical transcribe calls yield identical
segments; tests NEVER touch a real voice API (conftest cloud-free rule).
"""
from __future__ import annotations
import asyncio
import io
import math
import struct
import wave
from collections.abc import AsyncIterator
from .base import TranscriptSegment
def _tone_wav(duration_ms: int = 250, freq_hz: float = 1000.0) -> bytes:
"""A small, deterministic 16-bit mono WAV: a sine tone (stdlib only)."""
rate = 8000
n_samples = max(1, int(rate * duration_ms / 1000))
buf = io.BytesIO()
with wave.open(buf, "wb") as w:
w.setnchannels(1)
w.setsampwidth(2)
w.setframerate(rate)
for i in range(n_samples):
sample = int(12000 * math.sin(2 * math.pi * freq_hz * i / rate))
w.writeframes(struct.pack("<h", sample))
return buf.getvalue()
class MockVoiceFailure(RuntimeError):
"""Scripted failure mode for tests."""
class MockVoiceProvider:
"""Deterministic voice provider: scripted STT, canned-tone TTS.
- `transcribe`: pops the next scripted transcript from a queue (or a
default); two identical calls with the same queue state are identical.
Failure mode: raise MockVoiceFailure when the queue holds a failure
marker (the string "FAIL") or `audio` is empty.
- `synthesize`: yields the canned tone WAV in fixed-size chunks; failure
mode: empty text raises MockVoiceFailure.
"""
def __init__(self, transcripts: list[str] | None = None) -> None:
self._transcripts = list(transcripts or [])
self._cursor = 0
self.transcribe_calls = 0
self.synthesize_calls = 0
def script(self, transcripts: list[str]) -> None:
"""Replace the scripted queue (tests set expectations up front)."""
self._transcripts = list(transcripts)
self._cursor = 0
async def transcribe(self, audio: bytes, fmt: str) -> TranscriptSegment:
self.transcribe_calls += 1
if not audio:
raise MockVoiceFailure("no audio bytes provided")
if not self._transcripts:
raise MockVoiceFailure("transcript queue exhausted — script it")
item = self._transcripts[self._cursor]
self._cursor = (self._cursor + 1) % len(self._transcripts)
if item == "FAIL":
raise MockVoiceFailure("scripted STT failure")
return TranscriptSegment(
text=item,
duration_ms=max(1, len(audio) // 32), # deterministic pseudo-duration
)
async def synthesize(self, text: str, voice: str = "default") -> AsyncIterator[bytes]: # noqa: ASYNC109 (protocol parity)
# NOTE: protocol parity matters more than the async-generator purity
# lint; the real provider seam (v0.4) will stream over HTTP.
self.synthesize_calls += 1
if not text:
raise MockVoiceFailure("cannot synthesize empty text")
wav = _tone_wav(duration_ms=min(2000, max(120, len(text) * 12)))
for i in range(0, len(wav), 1024):
yield wav[i : i + 1024]
await asyncio.sleep(0) # yield to the loop like a network stream
# Protocol-shape parity guard (mock must satisfy the D-030 port).
from .base import VoiceProvider # noqa: E402
def _assert_protocol() -> None:
assert isinstance(MockVoiceProvider(), VoiceProvider)
_assert_protocol()
-7
View File
@@ -14,13 +14,6 @@ dependencies = [
"pydantic-settings>=2.15,<2.16",
"httpx>=0.28,<0.29",
"sse-starlette>=3.4,<3.5",
"sqlmodel>=0.0.24,<0.1",
"sqlalchemy>=2.0,<2.1",
"websockets>=13,<16",
"aiofiles>=24.1,<26",
# POST /v1/defense/{id}/answer multipart audio (REQ-3-006): FastAPI
# form/File parsing requires python-multipart at runtime.
"python-multipart>=0.0.32,<0.1",
]
[project.optional-dependencies]
+7 -52
View File
@@ -1,73 +1,28 @@
#!/usr/bin/env bash
# Idempotent bootstrap: create venv + install deps.
# Handles Debian/Ubuntu systems without python3-venv/ensurepip via --without-pip + get-pip.
# v2 (v0.3.5): recovers from a poisoned partial .venv left by a failed earlier
# attempt, cleans before each retry, and dies with a distro-specific fix hint
# when venv creation is impossible (e.g. missing python3.XX-venv package).
# Handles Debian systems without python3-venv/ensurepip via --without-pip + get-pip.
set -euo pipefail
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
APP_DIR="$(cd "$SCRIPT_DIR/.." && pwd)"
VENV="$APP_DIR/.venv"
venv_usable() {
[ -x "$VENV/bin/python3" ]
}
rm_broken_venv() {
echo "bootstrap: removing broken partial .venv from a failed earlier attempt" >&2
rm -rf "$VENV"
}
mkdir -p "$HOME/.cache/ciagent"
if venv_usable && [ ! -x "$VENV/bin/pip" ]; then
# A usable python3 without pip means the --without-pip fallback half-ran and
# the get-pip step never completed: start over cleanly.
rm_broken_venv
fi
if ! venv_usable; then
if [ -d "$VENV" ]; then
# Directory exists but no working python3: remains of a crashed venv create.
rm_broken_venv
fi
if python3 -m venv "$VENV" 2>/tmp/venv-create.err; then
if [ ! -x "$VENV/bin/python3" ]; then
if python3 -m venv "$VENV" 2>/dev/null; then
:
else
rm -rf "$VENV"
if python3 -m venv --without-pip "$VENV" 2>>/tmp/venv-create.err; then
:
else
rm -rf "$VENV"
PYVER="$(python3 -c 'import sys; print("%d.%d" % sys.version_info[:2])' 2>/dev/null || true)"
PKG="python3-venv"
[ -n "$PYVER" ] && PKG="python${PYVER}-venv"
echo "bootstrap: could not create a virtual environment." >&2
echo " python3 reported:" >&2
sed 's/^/ /' /tmp/venv-create.err >&2 || true
echo " fix (Debian/Ubuntu): install the venv support package, then re-run nextcraft bootstrap:" >&2
echo " apt install $PKG" >&2
exit 1
fi
# No ensurepip available — create bare venv and bootstrap pip separately.
python3 -m venv --without-pip "$VENV"
fi
fi
if [ ! -x "$VENV/bin/pip" ]; then
GET_PIP="$HOME/.cache/ciagent/get-pip.py"
if [ ! -f "$GET_PIP" ]; then
if ! curl -sSf --max-time 60 https://bootstrap.pypa.io/get-pip.py -o "$GET_PIP"; then
rm -rf "$VENV"
echo "bootstrap: get-pip.py download failed (no network?)." >&2
echo " fix: restore network access and re-run nextcraft bootstrap" >&2
exit 1
fi
fi
if ! "$VENV/bin/python3" "$GET_PIP" --quiet; then
rm -rf "$VENV"
echo "bootstrap: pip installation into the venv failed." >&2
echo " fix: re-run nextcraft bootstrap (the venv was cleaned; this retry is safe)" >&2
exit 1
curl -sSf --max-time 60 https://bootstrap.pypa.io/get-pip.py -o "$GET_PIP"
fi
"$VENV/bin/python3" "$GET_PIP" --quiet
fi
"$VENV/bin/pip" install --quiet --upgrade pip
+3 -17
View File
@@ -1,7 +1,5 @@
#!/usr/bin/env bash
# Dev server: export secrets (if present) then run uvicorn.
# Binds 0.0.0.0 by default so the stack is reachable from other machines
# (v0.3.5 network mode) — set AI_HOST=127.0.0.1 in .env to revert to loopback.
# Dev server: export secrets (if present) then run uvicorn on :8420.
set -euo pipefail
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
APP_DIR="$(cd "$SCRIPT_DIR/.." && pwd)"
@@ -9,7 +7,7 @@ REPO_ROOT="$(cd "$APP_DIR/../.." && pwd)"
VENV="$APP_DIR/.venv"
if [ ! -x "$VENV/bin/uvicorn" ]; then
echo "venv missing — run nextcraft bootstrap first (or: bash scripts/bootstrap.sh)" >&2
echo "venv missing — run scripts/bootstrap.sh first" >&2
exit 1
fi
@@ -24,17 +22,5 @@ if [ -f "$SECRETS" ]; then
done < "$SECRETS"
fi
ENV_FILE="$APP_DIR/.env"
if [ -f "$ENV_FILE" ]; then
while IFS='=' read -r key value; do
case "$key" in
AI_HOST|AI_PORT|AI_CORS_ORIGINS) export "$key=$value" ;;
esac
done < "$ENV_FILE"
fi
HOST="${AI_HOST:-0.0.0.0}"
PORT="${AI_PORT:-8420}"
cd "$APP_DIR"
exec "$VENV/bin/uvicorn" ai_service.main:app --reload --host "$HOST" --port "$PORT"
exec "$VENV/bin/uvicorn" ai_service.main:app --reload --port 8420
-691
View File
@@ -1,691 +0,0 @@
#!/usr/bin/env python3
"""sandbox-agent — stdlib-only in-sandbox telemetry capture agent (REQ-3-003).
D-031: this file is copied into the sandbox namespace and runs against the
system Python no third-party packages are importable there, so this module
depends on the standard library ONLY (the test suite enforces this with an
AST scan of the file).
What it does:
* wraps a non-interactive `/bin/sh` REPL: each stdin line is executed via
`sh -c` inside the workspace and reported as `stdin` -> `command` ->
`stdout` -> `run_result`/`test_result` events;
* polls the workspace tree (~250 ms) and emits `file_diff` events
(created/modified/deleted with unified diffs) plus periodic `activity`
heartbeats;
* streams events to ai-service as TelemetryEvent-shaped JSON frames over a
raw-socket RFC 6455 WebSocket client (no `websockets` package exists in
the namespace the client handshake + frame codec is implemented here);
* at-least-once delivery (D-026): every event is appended to an fsync'd
JSONL spool file inside the workdir BEFORE any send attempt; on
disconnect the spool grows; after reconnect (exponential backoff) the
spool is flushed oldest-first. The server dedups on (learner, task, seq)
so replayed duplicates are harmless loss is not tolerated.
Configured entirely through env baked at spawn time:
NC_LEARNER_ID / NC_TASK_ID / NC_INGEST_URL / NC_SANDBOX_ID (required)
NC_WORKSPACE workspace root to watch/run in (default: cwd)
NC_SPOOL spool path (default: <workspace>/.nc-agent/spool.jsonl)
NC_POLL_INTERVAL_S / NC_ACTIVITY_INTERVAL_S / NC_COMMAND_TIMEOUT_S
NC_BACKOFF_BASE_S / NC_BACKOFF_MAX_S (optional knobs)
Sequencing survives process restarts (incl. SIGKILL): on boot the spool is
replayed into the pending queue and `seq` resumes at max(spooled seq) + 1.
"""
from __future__ import annotations
import base64
import difflib
import hashlib
import json
import os
import secrets
import socket
import ssl
import struct
import subprocess
import sys
import threading
import time
import urllib.parse
from collections import deque
from collections.abc import Mapping
from dataclasses import dataclass
from datetime import UTC, datetime
from pathlib import Path
from typing import Any
_WS_GUID = "258EAFA5-E914-47DA-95CA-C5AB0DC85B11"
_EVENT_KINDS = frozenset(
{"command", "file_diff", "run_result", "test_result", "activity", "stdin", "stdout"}
)
_AGENT_DIR_PREFIX = ".nc-" # agent-private paths (spool) are excluded from watching
_MAX_DIFF_BYTES = 64 * 1024 # files larger than this are reported truncated, no diff
_MAX_OUTPUT_CHARS = 64 * 1024 # captured stdout/stderr tail cap per command
_HANDSHAKE_MAX_BYTES = 64 * 1024
# --------------------------------------------------------------------------- config
@dataclass(frozen=True)
class AgentConfig:
"""Runtime configuration, normally built from `NC_*` env baked at spawn."""
learner_id: str
task_id: str
ingest_url: str
sandbox_id: str
workspace: Path
spool_path: Path
poll_interval_s: float = 0.25
activity_interval_s: float = 5.0
command_timeout_s: float = 30.0
backoff_base_s: float = 0.25
backoff_max_s: float = 8.0
def __post_init__(self) -> None:
for name in ("learner_id", "task_id", "ingest_url", "sandbox_id"):
if not getattr(self, name):
raise ValueError(f"missing required config: NC_{name.upper()}")
@classmethod
def from_env(cls, env: Mapping[str, str] | None = None) -> AgentConfig:
src = os.environ if env is None else env
workspace = Path(src.get("NC_WORKSPACE") or os.getcwd()).resolve()
return cls(
learner_id=src.get("NC_LEARNER_ID", ""),
task_id=src.get("NC_TASK_ID", ""),
ingest_url=src.get("NC_INGEST_URL", ""),
sandbox_id=src.get("NC_SANDBOX_ID", ""),
workspace=workspace,
spool_path=Path(
src.get("NC_SPOOL") or (workspace / ".nc-agent" / "spool.jsonl")
),
poll_interval_s=float(src.get("NC_POLL_INTERVAL_S", "0.25")),
activity_interval_s=float(src.get("NC_ACTIVITY_INTERVAL_S", "5.0")),
command_timeout_s=float(src.get("NC_COMMAND_TIMEOUT_S", "30.0")),
backoff_base_s=float(src.get("NC_BACKOFF_BASE_S", "0.25")),
backoff_max_s=float(src.get("NC_BACKOFF_MAX_S", "8.0")),
)
# --------------------------------------------------------------------------- spool
class Spool:
"""Append-only JSONL spool with per-append fsync (survives SIGKILL).
`rewrite` swaps in a compacted file atomically (tmp file + os.replace).
Lines are stored without trailing newlines in memory, one per line on disk.
"""
def __init__(self, path: Path) -> None:
self._path = path
path.parent.mkdir(parents=True, exist_ok=True)
@property
def path(self) -> Path:
return self._path
def append(self, line: str) -> None:
with self._path.open("a", encoding="utf-8") as fh:
fh.write(line + "\n")
fh.flush()
os.fsync(fh.fileno())
def read_all(self) -> list[str]:
if not self._path.exists():
return []
with self._path.open("r", encoding="utf-8") as fh:
return [line.rstrip("\n") for line in fh if line.strip()]
def rewrite(self, lines: list[str]) -> None:
tmp = self._path.with_name(self._path.name + ".tmp")
with tmp.open("w", encoding="utf-8") as fh:
for line in lines:
fh.write(line + "\n")
fh.flush()
os.fsync(fh.fileno())
os.replace(tmp, self._path)
# --------------------------------------------------------------- websocket codec
def _encode_frame(opcode: int, payload: bytes) -> bytes:
"""RFC 6455 client frame: FIN set, always masked (servers require it)."""
header = bytearray([0x80 | opcode])
n = len(payload)
if n < 126:
header.append(0x80 | n)
elif n < 65536:
header.append(0x80 | 126)
header += struct.pack("!H", n)
else:
header.append(0x80 | 127)
header += struct.pack("!Q", n)
mask = secrets.token_bytes(4)
header += mask
masked = bytes(b ^ mask[i % 4] for i, b in enumerate(payload))
return bytes(header) + masked
class WsConnection:
"""Minimal blocking RFC 6455 client over a raw socket (stdlib only)."""
def __init__(self, sock: socket.socket) -> None:
self._sock = sock
self._write_lock = threading.Lock()
@classmethod
def connect(cls, url: str, timeout_s: float = 5.0) -> WsConnection:
parts = urllib.parse.urlsplit(url)
if parts.scheme not in ("ws", "wss"):
raise ValueError(f"unsupported scheme in NC_INGEST_URL: {parts.scheme!r}")
host = parts.hostname or "localhost"
port = parts.port or (443 if parts.scheme == "wss" else 80)
path = parts.path or "/"
if parts.query:
path += "?" + parts.query
sock = socket.create_connection((host, port), timeout=timeout_s)
if parts.scheme == "wss":
sock = ssl.create_default_context().wrap_socket(sock, server_hostname=host)
key = base64.b64encode(secrets.token_bytes(16)).decode("ascii")
request = (
f"GET {path} HTTP/1.1\r\n"
f"Host: {host}:{port}\r\n"
"Upgrade: websocket\r\n"
"Connection: Upgrade\r\n"
f"Sec-WebSocket-Key: {key}\r\n"
"Sec-WebSocket-Version: 13\r\n\r\n"
)
sock.sendall(request.encode("ascii"))
response = cls._read_http_response(sock)
cls._validate_handshake(response, key)
return cls(sock)
@staticmethod
def _read_http_response(sock: socket.socket) -> bytes:
buf = b""
while b"\r\n\r\n" not in buf:
chunk = sock.recv(4096)
if not chunk:
raise ConnectionError("server closed during WebSocket handshake")
buf += chunk
if len(buf) > _HANDSHAKE_MAX_BYTES:
raise ConnectionError("handshake response exceeded size cap")
return buf.split(b"\r\n\r\n", 1)[0]
@staticmethod
def _validate_handshake(response: bytes, key: str) -> None:
head = response.decode("latin-1")
lines = head.split("\r\n")
if not lines or " 101" not in lines[0]:
raise ConnectionError(f"handshake rejected: {lines[0] if lines else '<empty>'}")
headers = {}
for line in lines[1:]:
if ":" in line:
name, _, value = line.partition(":")
headers[name.strip().lower()] = value.strip()
expect = base64.b64encode(
hashlib.sha1((key + _WS_GUID).encode("ascii")).digest()
).decode("ascii")
if headers.get("sec-websocket-accept") != expect:
raise ConnectionError("bad Sec-WebSocket-Accept in handshake response")
# -- send ------------------------------------------------------------
def send_text(self, text: str) -> None:
with self._write_lock:
self._sock.sendall(_encode_frame(0x1, text.encode("utf-8")))
def _send_frame(self, opcode: int, payload: bytes) -> None:
with self._write_lock:
self._sock.sendall(_encode_frame(opcode, payload))
# -- receive ---------------------------------------------------------
def recv_message(self, timeout_s: float) -> tuple[int, bytes] | None:
"""Return (opcode, payload) for a data/close frame, or None on timeout.
Ping frames are answered with pong internally and never surfaced;
pongs are swallowed. Fragmented messages are reassembled. Raises
ConnectionError/OSError when the socket breaks.
"""
deadline = time.monotonic() + timeout_s
fragments = bytearray()
frag_opcode = 0
while True:
frame = self._recv_one_frame(deadline)
if frame is None:
return None
fin, opcode, payload = frame
if opcode == 0x9: # ping
self._send_frame(0xA, payload)
continue
if opcode == 0xA: # pong
continue
if opcode == 0x0: # continuation
fragments += payload
else:
fragments = bytearray(payload)
frag_opcode = opcode
if fin:
return frag_opcode, bytes(fragments)
def _recv_one_frame(self, deadline: float) -> tuple[bool, int, bytes] | None:
header = self._read_exact(2, deadline)
if header is None:
return None
b0, b1 = header[0], header[1]
fin = bool(b0 & 0x80)
opcode = b0 & 0x0F
length = b1 & 0x7F
if length == 126:
ext = self._read_exact(2, deadline)
if ext is None:
return None
length = struct.unpack("!H", ext)[0]
elif length == 127:
ext = self._read_exact(8, deadline)
if ext is None:
return None
length = struct.unpack("!Q", ext)[0]
mask = self._read_exact(4, deadline) if (b1 & 0x80) else b""
if mask is None:
return None
payload = self._read_exact(length, deadline) if length else b""
if payload is None:
return None
if mask:
payload = bytes(b ^ mask[i % 4] for i, b in enumerate(payload))
return fin, opcode, payload
def _read_exact(self, n: int, deadline: float) -> bytes | None:
buf = bytearray()
while len(buf) < n:
remaining = deadline - time.monotonic()
if remaining <= 0:
return None
self._sock.settimeout(remaining)
try:
chunk = self._sock.recv(n - len(buf))
except TimeoutError:
return None
if not chunk:
raise ConnectionError("peer closed the WebSocket connection")
buf += chunk
return bytes(buf)
def close(self) -> None:
try:
self._sock.close()
except OSError:
pass
# --------------------------------------------------------------------------- agent
class Agent:
"""Wires capture (shell + workspace watcher) to the framed event stream."""
def __init__(self, config: AgentConfig) -> None:
self.config = config
self._spool = Spool(config.spool_path)
self._pending: deque[str] = deque()
self._seq = 0
self._emit_lock = threading.Lock() # serializes seq + spool + flush
self._conn_lock = threading.Lock() # guards _conn swaps
self._conn: WsConnection | None = None
self._last_sent: str | None = None # one-line replay margin, see below
self._stop = threading.Event()
self._threads: list[threading.Thread] = []
self._baseline: dict[str, tuple[int, int, str | None]] = {}
self._resume_from_spool()
# -- durability ------------------------------------------------------
def _resume_from_spool(self) -> None:
highest = -1
for line in self._spool.read_all():
self._pending.append(line)
try:
seq = int(json.loads(line).get("seq", -1))
except (ValueError, AttributeError):
continue
highest = max(highest, seq)
self._seq = highest + 1
# -- event construction ---------------------------------------------
def _wire_frame(self, spooled_line: str) -> str:
"""Spool format -> wire format: strip URL-owned identity fields.
The spool keeps full events (local durability + restart recovery).
The ingest endpoint binds identity at the WS handshake (query params)
and rejects frames carrying learner_id/task_id (`extra="forbid"`
anti-spoofing), so the wire frame carries only seq/kind/payload/ts.
"""
import json as _json
full = _json.loads(spooled_line)
wire = {
k: full[k]
for k in ("seq", "kind", "payload", "ts")
}
if full.get("sandbox_id"):
wire["sandbox_id"] = full["sandbox_id"]
return _json.dumps(wire)
def _next_event(self, kind: str, payload: dict[str, Any]) -> dict[str, Any]:
if kind not in _EVENT_KINDS:
raise ValueError(f"unknown event kind: {kind!r}")
event = {
"learner_id": self.config.learner_id,
"task_id": self.config.task_id,
"seq": self._seq,
"kind": kind,
"payload": payload,
"ts": datetime.now(UTC).isoformat(),
"sandbox_id": self.config.sandbox_id,
}
self._seq += 1
return event
# -- emission / flush (D-026) ----------------------------------------
def emit(self, kind: str, payload: dict[str, Any]) -> dict[str, Any]:
"""Spool-then-send. Never blocks on reconnect; loss is impossible."""
with self._emit_lock:
line = json.dumps(self._next_event(kind, payload))
self._spool.append(line) # durable BEFORE any send attempt
self._pending.append(line)
self._flush_locked()
return json.loads(line)
def _flush_locked(self) -> None:
conn = self._current_conn()
while self._pending and conn is not None:
line = self._pending[0]
try:
conn.send_text(self._wire_frame(line))
except (ConnectionError, OSError):
self._drop_conn()
return
self._pending.popleft()
self._last_sent = line # kept until a later send proves delivery
if not self._pending and self._last_sent is not None:
# Compact, but retain the most recently sent line: a send into a
# silently-dead socket "succeeds" once at TCP level, so the last
# line is only confirmed-sent once a later write works. Retention
# is cheap; the server dedups on (learner, task, seq).
self._spool.rewrite([self._last_sent])
def replay_margin(self) -> None:
"""Requeue the last-sent line after a detected disconnect."""
with self._emit_lock:
if self._last_sent is not None and (
not self._pending or self._pending[0] != self._last_sent
):
self._pending.appendleft(self._last_sent)
self._spool.rewrite(list(self._pending))
self._last_sent = None
# -- connection supervision ------------------------------------------
def _current_conn(self) -> WsConnection | None:
with self._conn_lock:
return self._conn
def _set_conn(self, conn: WsConnection | None) -> None:
with self._conn_lock:
self._conn = conn
def _drop_conn(self) -> None:
conn = self._current_conn()
self._set_conn(None)
if conn is not None:
conn.close()
self.replay_margin()
def is_connected(self) -> bool:
return self._current_conn() is not None
def wait_connected(self, timeout_s: float) -> bool:
return self._wait_for(lambda: self.is_connected(), timeout_s)
def wait_disconnected(self, timeout_s: float) -> bool:
return self._wait_for(lambda: not self.is_connected(), timeout_s)
def _wait_for(self, pred: Any, timeout_s: float) -> bool:
deadline = time.monotonic() + timeout_s
while time.monotonic() < deadline:
if pred():
return True
time.sleep(0.02)
return pred()
def _supervisor_loop(self) -> None:
"""Maintain the WS connection: connect, flush backlog, read, backoff."""
backoff = self.config.backoff_base_s
while not self._stop.is_set():
if self._current_conn() is None:
try:
conn = WsConnection.connect(self.config.ingest_url)
except (ConnectionError, OSError, ValueError, TimeoutError):
self._stop.wait(backoff)
backoff = min(self.config.backoff_max_s, backoff * 2)
continue
self._set_conn(conn)
self._last_sent = None
backoff = self.config.backoff_base_s
with self._emit_lock: # ordered against concurrent emit()s
self._flush_locked()
else:
conn = self._current_conn()
if conn is None:
continue
try:
frame = conn.recv_message(timeout_s=1.0)
except (ConnectionError, OSError):
self._drop_conn()
continue
if frame is None:
continue
opcode, _payload = frame
if opcode == 0x8: # server close frame
self._drop_conn()
# -- workspace watcher ------------------------------------------------
def _snapshot_workspace(self) -> dict[str, tuple[int, int, str | None]]:
"""Map rel path -> (mtime_ns, size, text-or-None-if-too-large)."""
snap: dict[str, tuple[int, int, str | None]] = {}
root = self.config.workspace
if not root.is_dir():
return snap
for dirpath, dirnames, filenames in os.walk(root):
dirnames[:] = sorted(
d for d in dirnames if not d.startswith(_AGENT_DIR_PREFIX)
)
for name in sorted(filenames):
if name.startswith(_AGENT_DIR_PREFIX):
continue
path = Path(dirpath) / name
try:
st = path.stat()
except OSError:
continue
rel = path.relative_to(root).as_posix()
text: str | None = None
if st.st_size <= _MAX_DIFF_BYTES:
try:
text = path.read_text(encoding="utf-8", errors="replace")
except OSError:
pass
snap[rel] = (st.st_mtime_ns, st.st_size, text)
return snap
def _file_diff_payload(self, rel: str, change: str, old: str | None, new: str | None) -> dict:
payload: dict[str, Any] = {"path": rel, "change": change}
if old is None and new is None:
payload["truncated"] = True
return payload
diff = "".join(
difflib.unified_diff(
(old or "").splitlines(keepends=True),
(new or "").splitlines(keepends=True),
fromfile=f"a/{rel}",
tofile=f"b/{rel}",
)
)
payload["diff"] = diff
payload["size"] = len(new or "")
return payload
def _watcher_loop(self) -> None:
# Baseline is taken in start() before it returns, so any change made
# after start() completes is guaranteed to be observed.
baseline = self._baseline
last_heartbeat = time.monotonic()
while not self._stop.wait(self.config.poll_interval_s):
current = self._snapshot_workspace()
for rel in sorted(current.keys() | baseline.keys()):
if rel not in baseline and rel in current:
self.emit(
"file_diff",
self._file_diff_payload(rel, "created", None, current[rel][2]),
)
elif rel in baseline and rel not in current:
self.emit(
"file_diff",
self._file_diff_payload(rel, "deleted", baseline[rel][2], None),
)
else:
old_stat, new_stat = baseline[rel], current[rel]
if old_stat[:2] != new_stat[:2] and old_stat[2] != new_stat[2]:
self.emit(
"file_diff",
self._file_diff_payload(
rel, "modified", old_stat[2], new_stat[2]
),
)
baseline = current
if time.monotonic() - last_heartbeat >= self.config.activity_interval_s:
self.emit("activity", {"state": "idle", "spooled": len(self._pending)})
last_heartbeat = time.monotonic()
# -- shell wrapper ------------------------------------------------------
@staticmethod
def _is_test_command(cmd: str) -> bool:
return "test" in cmd.lower()
def run_command(self, line: str) -> dict[str, Any] | None:
"""Run one REPL line; emits stdin/command/stdout/run|test_result."""
line = line.strip()
if not line:
return None
self.emit("stdin", {"line": line})
self.emit("activity", {"state": "command", "spooled": len(self._pending)})
self.emit("command", {"cmd": line})
started = time.monotonic()
timed_out = False
exit_code: int | None = None
out: str | bytes = ""
err: str | bytes = ""
try:
proc = subprocess.run(
["sh", "-c", line],
cwd=self.config.workspace,
capture_output=True,
timeout=self.config.command_timeout_s,
text=True,
errors="replace",
)
exit_code, out, err = proc.returncode, proc.stdout, proc.stderr
except subprocess.TimeoutExpired as exc:
timed_out = True
# TimeoutExpired output attrs are always bytes (even in text mode).
out = exc.stdout or b""
err = exc.stderr or b""
duration = time.monotonic() - started
for stream, data in (("stdout", out), ("stderr", err)):
if isinstance(data, bytes):
data = data.decode(errors="replace")
if data:
self.emit("stdout", {"stream": stream, "data": data[-_MAX_OUTPUT_CHARS:]})
kind = "test_result" if self._is_test_command(line) else "run_result"
result = self.emit(
kind,
{
"cmd": line,
"exit_code": exit_code,
"duration_s": round(duration, 6),
"timed_out": timed_out,
},
)
self.emit("activity", {"state": "idle", "spooled": len(self._pending)})
return result
# -- lifecycle ----------------------------------------------------------
def start(self) -> None:
self.config.workspace.mkdir(parents=True, exist_ok=True)
self._baseline = self._snapshot_workspace()
self.emit("activity", {"state": "starting", "spooled": len(self._pending)})
self._threads = [
threading.Thread(target=self._supervisor_loop, daemon=True, name="nc-ws"),
threading.Thread(target=self._watcher_loop, daemon=True, name="nc-watch"),
]
for thread in self._threads:
thread.start()
def stop(self) -> None:
if self._stop.is_set():
return
try:
self.emit("activity", {"state": "stopped", "spooled": len(self._pending)})
finally:
self._stop.set()
self._drop_conn()
for thread in self._threads:
thread.join(timeout=3)
for thread in self._threads:
thread.join(timeout=3)
with self._emit_lock:
self._spool.rewrite(list(self._pending))
def main() -> int:
try:
config = AgentConfig.from_env()
except ValueError as exc:
print(f"sandbox-agent: {exc}", file=sys.stderr)
return 2
agent = Agent(config)
agent.start()
try:
# REPL mode: each stdin line is executed and reported (interactive use).
# Daemon mode: when stdin is closed/absent (the sandbox backend spawns
# the agent with stdin=DEVNULL), keep streaming workspace diffs +
# activity until SIGTERM/SIGINT so the agent's lifecycle is tied to
# the sandbox (destroy() reaps it) rather than to stdin EOF.
if sys.stdin is None or sys.stdin.closed: # pragma: no cover - defensive
agent._stop.wait() # noqa: SLF001 - daemon block
else:
line = sys.stdin.readline()
while line:
agent.run_command(line)
line = sys.stdin.readline()
if not agent._stop.is_set() and not sys.stdin.isatty(): # noqa: SLF001
# EOF on a pipe (DEVNULL): daemonize — watch + stream until killed.
import signal
signal.signal(signal.SIGTERM, lambda *_: agent._stop.set()) # noqa: SLF001
agent._stop.wait() # noqa: SLF001
except KeyboardInterrupt:
pass
finally:
agent.stop()
return 0
if __name__ == "__main__":
raise SystemExit(main())
@@ -1,118 +0,0 @@
"""Assessor agent tests — live-grade coaching contract (REQ-3-007).
v0.3 re-grounding: the Assessor renders coaching FROM the stored grade
(GradeRecord) it never invents scores (the grading engine owns that).
"""
from __future__ import annotations
import json
from datetime import UTC, datetime
import pytest
from ai_service.agents.assessor import AssessorAgent, GradeCoaching
from ai_service.config import Settings
from ai_service.grading.store import GradeRecord, SQLiteGradeStore
from ai_service.llm.mock import MockProvider
from ai_service.llm.types import Message
COACHING_JSON = json.dumps(
{
"summary": "Solid iterative build; tests drove the fixes.",
"strengths": ["Ran tests after each change."],
"gaps": ["Did not cover the empty-input case."],
"next_steps": ["Add one edge-case test."],
}
)
class ScriptedProvider(MockProvider):
def __init__(self) -> None:
super().__init__()
self.requests: list[list[Message]] = []
self.replies: list[str] = []
async def chat(self, messages, *, model, temperature=0.7, response_format=None):
self.requests.append([Message(role=m.role, content=m.content) for m in messages])
if self.replies:
return self.replies.pop(0)
return COACHING_JSON
def _grade() -> GradeRecord:
return GradeRecord(
learner_id="assessor-learner",
task_id="assessor-task",
variant_seed=None,
digest={"error_fix_cycles": 2, "final_test_status": "pass"},
scores={
"criteria": {
"process_quality": 4,
"correctness": 3,
"debugging_discipline": 4,
"test_usage": 3,
},
"strengths": ["s"],
"gaps": ["g"],
"verdict": "developing",
},
verdict="GRADED",
model="gemma4:31b",
created_at=datetime.now(UTC),
)
@pytest.fixture()
def provider() -> ScriptedProvider:
return ScriptedProvider()
@pytest.fixture()
def agent(provider) -> AssessorAgent:
return AssessorAgent(provider, Settings(provider="mock"))
class TestCoachGrade:
async def test_prompt_contains_stored_grade_not_learner_id(self, agent, provider) -> None:
await agent.coach_grade(_grade())
all_text = "\n".join(
m.content for request in provider.requests for m in request
)
assert "process_quality" in all_text # stored scores rendered
assert "GRADED" in all_text
assert "assessor-learner" not in all_text # D-028 anonymity
async def test_coaching_validates_via_d020(self, agent, provider) -> None:
coaching = await agent.coach_grade(_grade())
assert isinstance(coaching, GradeCoaching)
assert coaching.summary
assert coaching.next_steps
async def test_malformed_then_good_exercises_retry(self, agent, provider) -> None:
provider.replies = ["garbage", COACHING_JSON]
coaching = await agent.coach_grade(_grade())
assert coaching.summary
assert len(provider.requests) == 2
async def test_no_corpus_artifact_imports(self) -> None:
import ast
from pathlib import Path
py = Path(__file__).parents[2] / "ai_service" / "agents" / "assessor.py"
tree = ast.parse(py.read_text())
for node in ast.walk(tree):
if isinstance(node, ast.ImportFrom) and node.module:
assert "corpus.artifacts" not in node.module
assert "corpus.telemetry" not in node.module
class TestStoreRoundtrip:
def test_grade_store_roundtrip(self, tmp_path) -> None:
store = SQLiteGradeStore(db_path=tmp_path / "g.db")
record = _grade()
store.save(record)
fetched = store.get("assessor-learner", "assessor-task")
assert fetched is not None
assert fetched.scores["criteria"]["process_quality"] == 4
store.close()
@@ -1,50 +0,0 @@
"""Coach agent tests — persona, message assembly, streaming (REQ-2-005)."""
from ai_service.agents.coach import CoachAgent
from ai_service.config import Settings
from ai_service.corpus.learner_context import get_learner_context
from ai_service.llm.mock import MockProvider
from ai_service.llm.types import Message
def make_coach() -> CoachAgent:
return CoachAgent(MockProvider(), Settings(provider="mock"))
def test_system_prompt_includes_learner_context():
coach = make_coach()
prompt = coach.system_prompt(get_learner_context())
assert "Alex Rivera" in prompt
assert "AI Orchestration Engineer (62%)" in prompt
assert "retrieval practice" in prompt.lower()
assert "one clear next action" in prompt.lower()
def test_system_prompt_marks_current_focus():
coach = make_coach()
prompt = coach.system_prompt(get_learner_context())
assert "Multi-agent communication patterns" in prompt
def test_build_messages_system_history_user():
coach = make_coach()
history = [
Message(role="user", content="earlier"),
Message(role="assistant", content="reply"),
]
messages = coach.build_messages(history, "what next?", get_learner_context())
assert messages[0].role == "system"
assert "Coach" in messages[0].content
assert [m.content for m in messages[1:]] == ["earlier", "reply", "what next?"]
async def test_stream_reply_yields_deltas():
coach = make_coach()
ctx = get_learner_context()
tokens = [t async for t in coach.stream_reply(user_input="hello", learner_context=ctx)]
assert tokens
assert all(isinstance(t, str) for t in tokens)
def test_agent_name():
assert make_coach().name == "coach"
@@ -1,124 +0,0 @@
"""Telemetry + artifacts corpus tests (REQ-2-007/008 inputs, D-021)."""
import re
from pathlib import Path
import pytest
from ai_service.corpus.artifacts import (
ARTIFACTS,
RUBRICS,
get_artifact_bundle,
get_transcript_for_artifact,
render_rubric,
render_transcript,
rubric_for_competency,
)
from ai_service.corpus.telemetry import (
LAB_SCENARIOS,
get_lab_scenario,
summarize_scenario,
)
def test_lab_scenarios_addressable_by_id():
for scenario_id in (
"lab-scenario-strong",
"lab-scenario-struggling",
"lab-scenario-flagged",
):
scenario = get_lab_scenario(scenario_id)
assert scenario is not None
assert scenario.scenario_id == scenario_id
def test_lab_scenarios_have_distinct_event_profiles():
strong = get_lab_scenario("lab-scenario-strong")
struggling = get_lab_scenario("lab-scenario-struggling")
kinds = lambda s: {e.kind for e in s.events} # noqa: E731
assert "test_pass" in kinds(strong)
assert "test_fail" in kinds(struggling)
assert "idle" in kinds(struggling)
assert "paste" in kinds(get_lab_scenario("lab-scenario-flagged"))
def test_summarize_scenario_mentions_events():
text = summarize_scenario(get_lab_scenario("lab-scenario-struggling"))
assert "test_fail" in text
assert "ImportError" in text
assert "stack-orchestration-c002" in text
def test_unknown_scenario_returns_none():
assert get_lab_scenario("lab-scenario-ghost") is None
def test_artifacts_and_rubrics_resolve():
bundle = get_artifact_bundle("art-eval-research-assistant")
assert bundle is not None
artifact, rubric = bundle
assert artifact.competency_id == "stack-orchestration-c002"
assert rubric.rubric_id == "rubric-orchestration-c002"
assert len(rubric.criteria) == 4
def test_unknown_artifact_returns_none():
assert get_artifact_bundle("art-eval-ghost") is None
def test_transcripts_pair_with_artifacts():
for artifact_id in ARTIFACTS:
transcript = get_transcript_for_artifact(artifact_id)
assert transcript is not None
assert transcript.artifact_id == artifact_id
assert len(transcript.turns) >= 4
def test_rubric_render_mentions_all_criteria():
rubric = rubric_for_competency("stack-orchestration-c002")
text = render_rubric(rubric)
for criterion in rubric.criteria:
assert criterion.criterion_id in text
def test_transcript_render_has_both_speakers():
transcript = get_transcript_for_artifact("art-eval-rag-dashboard")
text = render_transcript(transcript)
assert "examiner:" in text
assert "learner:" in text
_TS_SOURCE_CANDIDATES = [
Path(__file__).resolve().parents[4] / "packages" / "mock-data" / "ai-scenarios.ts",
Path(__file__).resolve().parents[2] / "ai-scenarios.ts",
]
def _ts_source() -> Path:
for candidate in _TS_SOURCE_CANDIDATES:
if candidate.exists():
return candidate
pytest.skip("ai-scenarios.ts not found in this checkout layout")
def test_corpus_ids_align_with_ts_mock_data():
"""D-021: Python corpus IDs string-identical to ai-scenarios.ts."""
content = _ts_source().read_text()
ts_ids = re.findall(r"id:\s*'([^']+)'", content)
ts_scenarios = ts_ids[: len(LAB_SCENARIOS)]
ts_artifacts = ts_ids[len(LAB_SCENARIOS):]
assert sorted(ts_scenarios) == sorted(LAB_SCENARIOS), (
f"scenario IDs drifted: py={sorted(LAB_SCENARIOS)} ts={sorted(ts_scenarios)}"
)
assert sorted(ts_artifacts) == sorted(ARTIFACTS), (
f"artifact IDs drifted: py={sorted(ARTIFACTS)} ts={sorted(ts_artifacts)}"
)
def test_rubric_weights_sum_to_one():
"""Every rubric's criteria weights must sum to exactly 1.0."""
for rubric in RUBRICS.values():
total = sum(c.weight for c in rubric.criteria)
assert total == pytest.approx(1.0), (
f"{rubric.rubric_id} weights sum to {total}, expected 1.0"
)
@@ -1,159 +0,0 @@
"""Examiner agent tests (Task 5-2-01, REQ-3-006)."""
from __future__ import annotations
import json
import pytest
from ai_service.agents.examiner import DefenseVerdict, ExaminerAgent
from ai_service.agents.registry import AgentRegistry, register_builtin_agents
from ai_service.config import Settings
from ai_service.grading.features import compute_digest
from ai_service.llm.mock import MockProvider
from ai_service.llm.types import Message
VERDICT_JSON = json.dumps(
{
"verdict": "developing",
"understanding": "Explains the retry loop clearly.",
"process_justification": "Justifies the edit-then-test cadence from the digest.",
"communication": "Answers are specific and on-topic.",
"strengths": ["Grounded the fix in a failed test."],
"gaps": ["Did not justify the chunk-size choice."],
}
)
class RecordingProvider(MockProvider):
"""Mock provider that records every message list (prompt assertions)."""
def __init__(self, replies: list[str] | None = None) -> None:
super().__init__()
self.replies = list(replies or [])
self.requests: list[list[Message]] = []
async def chat(self, messages, *, model, temperature=0.7, response_format=None):
self.requests.append([Message(role=m.role, content=m.content) for m in messages])
if self.replies:
return self.replies.pop(0)
return "Tell me about your build."
def _digest():
from datetime import UTC, datetime, timedelta
from ai_service.telemetry.models import TelemetryEvent
t0 = datetime(2026, 9, 12, tzinfo=UTC)
events = [
TelemetryEvent(
learner_id="examiner-learner",
task_id="examiner-task",
seq=n,
kind=kind,
payload=payload,
ts=t0 + timedelta(seconds=n * 10),
sandbox_id="sbx-examiner",
)
for n, (kind, payload) in enumerate(
[
("file_diff", {"path": "a.py"}),
("command", {"cmd": "pytest -q"}),
("test_result", {"passed": False, "exit_code": 1}),
("file_diff", {"path": "a.py"}),
("test_result", {"passed": True, "exit_code": 0}),
]
)
]
return compute_digest(events)
@pytest.fixture()
def provider() -> RecordingProvider:
return RecordingProvider()
@pytest.fixture()
def settings() -> Settings:
return Settings(provider="mock")
@pytest.fixture()
def examiner(provider, settings) -> ExaminerAgent:
return ExaminerAgent(provider, settings)
class TestNextQuestion:
async def test_prompt_contains_digest_but_no_learner_id(
self, examiner, provider
) -> None:
await examiner.next_question(
history=[Message(role="assistant", content="First question?")],
trace_digest=_digest(),
variant_statement="Build a chunker.",
)
all_content = "\n".join(
m.content for request in provider.requests for m in request
)
assert "error_fix_cycles" in all_content # digest JSON grounded
assert "examiner-learner" not in all_content # D-028 anonymity
assert "Build a chunker." in all_content # variant statement grounded
assert all_content.count('"examiner-learner"') == 0
async def test_question_returned_from_provider(self, examiner) -> None:
question = await examiner.next_question(
history=[], trace_digest=_digest()
)
assert isinstance(question, str)
class TestFinalVerdict:
async def test_verdict_validates_via_d020(self, examiner, provider) -> None:
provider.replies = [VERDICT_JSON]
verdict = await examiner.final_verdict(
history=[Message(role="assistant", content="Q?")],
trace_digest=_digest(),
)
assert isinstance(verdict, DefenseVerdict)
assert verdict.verdict == "developing"
assert verdict.strengths and verdict.gaps
async def test_malformed_then_good_exercises_retry(self, examiner, provider) -> None:
provider.replies = ["not json", VERDICT_JSON]
verdict = await examiner.final_verdict(history=[], trace_digest=_digest())
assert verdict.verdict == "developing"
assert len(provider.requests) == 2 # D-020 bounded retry
class TestRegistry:
def test_all_seven_agents_resolve(self, provider, settings) -> None:
registry = AgentRegistry()
register_builtin_agents(registry)
assert registry.names() == [
"assessor",
"coach",
"examiner",
"lab",
"mentor",
"proctor",
"tutor",
]
agent = registry.get(provider, settings, "examiner")
assert isinstance(agent, ExaminerAgent)
class TestBoundary:
def test_examiner_never_imports_voice_or_api(self) -> None:
import ast
from pathlib import Path
py = Path(__file__).parents[2] / "ai_service" / "agents" / "examiner.py"
tree = ast.parse(py.read_text())
for node in ast.walk(tree):
if isinstance(node, ast.ImportFrom) and node.module:
assert "voice" not in node.module, "examiner must not import voice/"
assert not node.module.startswith("ai_service.api")
if isinstance(node, ast.Import):
for alias in node.names:
assert alias.name != "fastapi"
-42
View File
@@ -1,42 +0,0 @@
"""Lab agent tests — scenario-driven streaming feedback (REQ-2-007)."""
from ai_service.agents.lab import LabAgent
from ai_service.config import Settings
from ai_service.corpus.learner_context import get_learner_context
from ai_service.corpus.telemetry import get_lab_scenario, summarize_scenario
from ai_service.llm.mock import MockProvider
def make_lab() -> LabAgent:
return LabAgent(MockProvider(), Settings(provider="mock"))
def test_system_prompt_names_lab_persona():
prompt = make_lab().system_prompt(get_learner_context())
assert "Lab" in prompt
assert "Alex Rivera" in prompt
assert "telemetry" in prompt.lower()
async def test_stream_feedback_mentions_scenario_events():
"""Mock-scripted: feedback text derives from scenario timeline input —
distinct scenarios produce distinct (deterministic) replies."""
lab = make_lab()
ctx = get_learner_context()
strong = get_lab_scenario("lab-scenario-strong")
struggling = get_lab_scenario("lab-scenario-struggling")
strong_reply = "".join([t async for t in lab.stream_feedback(strong, ctx)])
struggling_reply = "".join([t async for t in lab.stream_feedback(struggling, ctx)])
assert strong_reply
assert strong_reply != struggling_reply # scenario-driven, not canned
def test_build_evaluation_messages_carry_timeline():
lab = make_lab()
scenario = get_lab_scenario("lab-scenario-flagged")
timeline = summarize_scenario(scenario)
messages = lab.build_messages(None, timeline, get_learner_context())
assert messages[0].role == "system"
assert "paste" in messages[-1].content
assert "2,400 chars" in messages[-1].content
@@ -1,49 +0,0 @@
"""Mentor agent tests — career narrative, session-backed (REQ-2-010)."""
from ai_service.agents.mentor import MentorAgent
from ai_service.config import Settings
from ai_service.corpus.learner_context import get_learner_context
from ai_service.llm.mock import MockProvider
from ai_service.llm.types import Message
def make_mentor() -> MentorAgent:
return MentorAgent(MockProvider(), Settings(provider="mock"))
def test_system_prompt_carries_full_learner_context():
prompt = make_mentor().system_prompt(get_learner_context())
assert "Alex Rivera" in prompt
assert "AI Orchestration Engineer (62%)" in prompt
assert "Multi-agent research assistant" in prompt # artifacts by name
assert "4" in prompt # microcredential count
def test_build_messages_system_history_user():
mentor = make_mentor()
history = [Message(role="user", content="what next?"),
Message(role="assistant", content="trajectory...")]
messages = mentor.build_messages(history, "tell me more", get_learner_context())
assert messages[0].role == "system"
assert "Mentor" in messages[0].content
assert [m.content for m in messages[1:]] == ["what next?", "trajectory...", "tell me more"]
async def test_stream_reply_mentions_learner_context_in_output():
"""Mock-scripted: narrative derives from context-injected messages —
different learner contexts produce distinct (deterministic) replies."""
mentor = make_mentor()
alex = get_learner_context("learner-001")
priya = get_learner_context("learner-002")
alex_reply = "".join(
[t async for t in mentor.stream_reply(user_input="narrate", learner_context=alex)]
)
priya_reply = "".join(
[t async for t in mentor.stream_reply(user_input="narrate", learner_context=priya)]
)
assert alex_reply
assert alex_reply != priya_reply
def test_agent_name():
assert make_mentor().name == "mentor"
@@ -1,84 +0,0 @@
"""Proctor agent tests — structured integrity signals (REQ-2-009)."""
import pytest
from ai_service.agents.proctor import ProctorAgent, ProctorAssessment
from ai_service.agents.structured import StructuredOutputError
from ai_service.config import Settings
from ai_service.corpus.learner_context import get_learner_context
from ai_service.corpus.telemetry import (
PROCTOR_SCENARIOS,
get_proctor_scenario,
summarize_proctor_scenario,
)
from ai_service.llm.mock import MockProvider, ScriptedJSONProvider
VALID_ASSESSMENT = {
"scenario_id": "proctor-scenario-distracted",
"signals": [
{"signal_type": "context_switch", "severity": "low",
"note": "Tab switch to docs at t+120s — normal engineering behavior"},
{"signal_type": "idle_gap", "severity": "medium",
"note": "5-minute idle at t+300s followed by more tab switches"},
],
"intervention": "Offer a short break and ask the learner to restate their answer plan",
"summary": "Distracted but explainable session; coach the focus pattern, don't flag it",
}
def make_proctor(provider=None) -> ProctorAgent:
return ProctorAgent(provider or MockProvider(), Settings(provider="mock"))
def test_proctor_scenarios_addressable_and_distinct_type():
healthy = get_proctor_scenario("proctor-scenario-healthy")
flagged = get_proctor_scenario("proctor-scenario-flagged")
assert healthy is not None and flagged is not None
kinds = lambda s: {e.kind for e in s.events} # noqa: E731
assert "tab_switch" in kinds(get_proctor_scenario("proctor-scenario-distracted"))
assert "paste_large" in kinds(flagged)
assert not kinds(healthy) & {"tab_switch", "paste_large", "focus_lost"}
def test_unknown_proctor_scenario_none():
assert get_proctor_scenario("proctor-scenario-ghost") is None
async def test_assess_returns_validated_signals():
provider = ScriptedJSONProvider(VALID_ASSESSMENT)
proctor = make_proctor(provider)
scenario = get_proctor_scenario("proctor-scenario-distracted")
result = await proctor.assess(scenario, get_learner_context())
assert isinstance(result, ProctorAssessment)
assert len(result.signals) == 2
assert result.signals[0].severity == "low"
assert "break" in result.intervention.lower()
async def test_assess_rejects_invalid_after_retry():
proctor = make_proctor(MockProvider()) # non-schema JSON
scenario = get_proctor_scenario("proctor-scenario-healthy")
with pytest.raises(StructuredOutputError):
await proctor.assess(scenario, get_learner_context())
def test_system_prompt_is_coaching_not_punitive():
prompt = make_proctor().system_prompt(get_learner_context())
assert "never punitive" in prompt.lower()
assert "good faith" in prompt.lower()
assert "ONLY" in prompt # JSON-only instruction
def test_timeline_summary_carries_events():
scenario = get_proctor_scenario("proctor-scenario-flagged")
text = summarize_proctor_scenario(scenario)
assert "paste_large" in text
assert "3,100 chars" in text
def test_all_three_proctor_scenarios_exist():
assert set(PROCTOR_SCENARIOS) == {
"proctor-scenario-healthy",
"proctor-scenario-distracted",
"proctor-scenario-flagged",
}
+1 -84
View File
@@ -1,17 +1,13 @@
"""Agent registry tests — register/get round-trip, error paths (G-4), builtins."""
"""Agent registry tests — register/get round-trip, error paths (G-4)."""
import pytest
from ai_service.agents.base import BaseAgent
from ai_service.agents.coach import CoachAgent
from ai_service.agents.registry import (
AgentRegistry,
DuplicateAgentError,
UnknownAgentError,
register_builtin_agents,
)
from ai_service.agents.tutor import TutorAgent
from ai_service.config import Settings
from ai_service.llm.mock import MockProvider
@@ -55,82 +51,3 @@ def test_names_sorted():
registry.register("zeta", make_factory())
registry.register("alpha", make_factory())
assert registry.names() == ["alpha", "zeta"]
def test_builtin_agents_register_and_resolve():
registry = AgentRegistry()
register_builtin_agents(registry)
assert set(registry.names()) >= {"coach", "tutor"}
settings = Settings(provider="mock")
coach = registry.get(MockProvider(), settings, "coach")
tutor = registry.get(MockProvider(), settings, "tutor")
assert isinstance(coach, CoachAgent)
assert isinstance(tutor, TutorAgent)
def test_lab_and_assessor_resolve_via_registry():
"""Phase 4: lab + assessor registered centrally (Task 4-3-01)."""
from ai_service.agents.assessor import AssessorAgent
from ai_service.agents.lab import LabAgent
registry = AgentRegistry()
register_builtin_agents(registry)
assert {"lab", "assessor"} <= set(registry.names())
settings = Settings(provider="mock")
lab = registry.get(MockProvider(), settings, "lab")
assessor = registry.get(MockProvider(), settings, "assessor")
assert isinstance(lab, LabAgent)
assert isinstance(assessor, AssessorAgent)
def test_proctor_and_mentor_resolve_via_registry():
"""Phase 5: proctor + mentor registered centrally (Tasks 5-1-02/5-2-01)."""
from ai_service.agents.mentor import MentorAgent
from ai_service.agents.proctor import ProctorAgent
registry = AgentRegistry()
register_builtin_agents(registry)
settings = Settings(provider="mock")
proctor = registry.get(MockProvider(), settings, "proctor")
mentor = registry.get(MockProvider(), settings, "mentor")
assert isinstance(proctor, ProctorAgent)
assert isinstance(mentor, MentorAgent)
def test_registry_resolves_all_seven_agents():
"""Must-Have (Phase 5): the full roster — six tutors + the Examiner."""
from ai_service.agents.assessor import AssessorAgent
from ai_service.agents.coach import CoachAgent
from ai_service.agents.examiner import ExaminerAgent
from ai_service.agents.lab import LabAgent
from ai_service.agents.mentor import MentorAgent
from ai_service.agents.proctor import ProctorAgent
from ai_service.agents.tutor import TutorAgent
registry = AgentRegistry()
register_builtin_agents(registry)
assert registry.names() == [
"assessor", "coach", "examiner", "lab", "mentor", "proctor", "tutor",
]
settings = Settings(provider="mock")
expected = {
"coach": CoachAgent,
"tutor": TutorAgent,
"lab": LabAgent,
"assessor": AssessorAgent,
"proctor": ProctorAgent,
"mentor": MentorAgent,
"examiner": ExaminerAgent,
}
for name, cls in expected.items():
agent = registry.get(MockProvider(), settings, name)
assert isinstance(agent, cls), f"{name} resolved to {type(agent).__name__}"
assert agent.name == name
def test_builtin_registration_is_idempotent_safe():
"""Duplicate registration raises — builtin bootstrap must be called once."""
registry = AgentRegistry()
register_builtin_agents(registry)
with pytest.raises(DuplicateAgentError):
register_builtin_agents(registry)
@@ -1,207 +0,0 @@
"""Live re-grounding tests: Lab, Assessor, Proctor on REAL inputs (REQ-3-007)."""
from __future__ import annotations
import json
import tempfile
from datetime import UTC, datetime, timedelta
from pathlib import Path
from fastapi.testclient import TestClient
from ai_service.agents.assessor import AssessorAgent, GradeCoaching
from ai_service.agents.lab import LabAgent
from ai_service.agents.proctor import ProctorAgent, ProctorAssessment
from ai_service.config import Settings
from ai_service.grading.store import GradeRecord, SQLiteGradeStore
from ai_service.llm.mock import MockProvider
from ai_service.llm.types import Message
from ai_service.telemetry.models import TelemetryEvent
from ai_service.telemetry.store import SQLiteTraceStore
COACHING_JSON = json.dumps(
{
"summary": "Iterative build with test discipline.",
"strengths": ["Tested after changes."],
"gaps": ["Missing edge cases."],
"next_steps": ["Add an edge-case test."],
}
)
PROCTOR_JSON = json.dumps(
{
"signals": [
{"signal_type": "idle_gap", "severity": "low", "note": "One 400s gap."}
],
"intervention": "Offer a short break.",
"summary": "Healthy session overall.",
}
)
T0 = datetime(2026, 9, 12, tzinfo=UTC)
class RecordingProvider(MockProvider):
def __init__(self, structured_json: str) -> None:
super().__init__()
self._structured_json = structured_json
self.requests: list[list[Message]] = []
def _reply_for(self, messages, response_format):
self.requests.append([Message(role=m.role, content=m.content) for m in messages])
if response_format is not None and response_format.get("type") == "json_object":
return self._structured_json
return "Coaching feedback referencing your latest test run."
def _event(
seq: int, kind: str, payload: dict, offset_s: float,
learner="live-learner", task="live-task",
):
return TelemetryEvent(
learner_id=learner,
task_id=task,
seq=seq,
kind=kind,
payload=payload,
ts=T0 + timedelta(seconds=offset_s),
sandbox_id="sbx-live",
)
def _seed_trace(store: SQLiteTraceStore) -> None:
events = [
_event(0, "file_diff", {"path": "a.py"}, 0),
_event(1, "command", {"cmd": "pytest -q"}, 10),
_event(2, "test_result", {"passed": False, "exit_code": 1}, 15),
_event(3, "file_diff", {"path": "a.py"}, 30),
_event(4, "test_result", {"passed": True, "exit_code": 0}, 45),
_event(5, "activity", {"state": "idle"}, 500), # >120s gap -> idle
]
for e in events:
store.append(e)
class TestLabLive:
async def test_lab_prompt_contains_digest_not_corpus(self, tmp_path) -> None:
store = SQLiteTraceStore(db_path=tmp_path / "t.db")
_seed_trace(store)
provider = RecordingProvider("feedback")
agent = LabAgent(provider, Settings(provider="mock"))
from ai_service.grading.features import compute_digest
digest = compute_digest(store.get_trace("live-learner", "live-task"))
tokens = [t async for t in agent.stream_feedback(digest)]
assert tokens
all_text = "\n".join(
m.content for request in provider.requests for m in request
)
assert "error_fix_cycles" in all_text
assert "lab-scenario" not in all_text # no corpus fixture ids
store.close()
def test_no_corpus_telemetry_import_in_lab(self) -> None:
import ast
from pathlib import Path
py = Path(__file__).parents[2] / "ai_service" / "agents" / "lab.py"
tree = ast.parse(py.read_text())
for node in ast.walk(tree):
if isinstance(node, ast.ImportFrom) and node.module:
assert "corpus.telemetry" not in node.module
class TestAssessorLive:
async def test_assessor_prompt_contains_stored_scores(self) -> None:
provider = RecordingProvider(COACHING_JSON)
agent = AssessorAgent(provider, Settings(provider="mock"))
grade = GradeRecord(
learner_id="live-learner",
task_id="live-task",
variant_seed=None,
digest={"error_fix_cycles": 1},
scores={"criteria": {"process_quality": 3}},
verdict="GRADED",
model="mock",
created_at=datetime.now(UTC),
)
coaching = await agent.coach_grade(grade)
assert isinstance(coaching, GradeCoaching)
all_text = "\n".join(
m.content for request in provider.requests for m in request
)
assert "process_quality" in all_text
assert "live-learner" not in all_text
def test_no_corpus_artifact_import_in_assessor(self) -> None:
import ast
from pathlib import Path
py = Path(__file__).parents[2] / "ai_service" / "agents" / "assessor.py"
tree = ast.parse(py.read_text())
for node in ast.walk(tree):
if isinstance(node, ast.ImportFrom) and node.module:
assert "corpus.artifacts" not in node.module
assert "corpus.telemetry" not in node.module
class TestProctorLive:
async def test_proctor_receives_real_digest_and_defense_signals(self) -> None:
provider = RecordingProvider(PROCTOR_JSON)
agent = ProctorAgent(provider, Settings(provider="mock"))
from ai_service.grading.features import compute_digest
store = SQLiteTraceStore(db_path=Path(tempfile.mkdtemp()) / "proctor-t.db")
_seed_trace(store)
digest = compute_digest(store.get_trace("live-learner", "live-task"))
store.close()
assessment = await agent.assess(
digest,
defense_signals={"long_pauses": [{"turn": 3, "latency_ms": 30000}]},
variant_context={"template_id": "tpl-llm-judge", "seed": "cafe", "params": {}},
)
assert isinstance(assessment, ProctorAssessment)
all_text = "\n".join(
m.content for request in provider.requests for m in request
)
assert "idle_gap" in all_text or "idle_gap_count" in all_text
assert "long_pauses" in all_text
assert "tpl-llm-judge" in all_text
assert "live-learner" not in all_text
def test_no_corpus_scenario_import_in_proctor(self) -> None:
import ast
from pathlib import Path
py = Path(__file__).parents[2] / "ai_service" / "agents" / "proctor.py"
tree = ast.parse(py.read_text())
for node in ast.walk(tree):
if isinstance(node, ast.ImportFrom) and node.module:
assert "corpus.telemetry" not in node.module
class TestProctorEndpoint:
def test_signals_endpoint_serves_real_inputs(self, tmp_path) -> None:
from ai_service.main import create_app
from ai_service.telemetry.ingest import TraceIntegrityMap
from ai_service.variants.store import SQLiteVariantStore
from ai_service.voice.defense_store import SQLiteDefenseStore
app = create_app(Settings(provider="mock"))
provider = RecordingProvider(PROCTOR_JSON)
app.state.provider = provider
app.state.trace_store = SQLiteTraceStore(db_path=tmp_path / "t.db")
app.state.grade_store = SQLiteGradeStore(db_path=tmp_path / "g.db")
app.state.variant_store = SQLiteVariantStore(db_path=tmp_path / "v.db")
app.state.defense_store = SQLiteDefenseStore(db_path=tmp_path / "d.db")
app.state.trace_integrity = TraceIntegrityMap()
_seed_trace(app.state.trace_store)
with TestClient(app) as client:
resp = client.post(
"/v1/proctor/signals",
json={"learner_id": "live-learner", "task_id": "live-task"},
)
assert resp.status_code == 200, resp.text
body = resp.json()
assert body["signals"] is not None
assert "intervention" in body
@@ -128,28 +128,11 @@ async def test_structured_completion_is_bounded_to_one_retry():
assert provider.calls == 2
async def test_structured_completion_sends_schema_instruction():
"""Layer 2: the schema hint must reach the provider in the request."""
async def test_structured_completion_schema_instruction_appended():
provider = ScriptedJSONProvider({"score": 70, "verdict": "passing"})
captured: list[list] = []
original = provider.chat
async def recording_chat(messages, *, model, temperature=0.7, response_format=None):
captured.append(list(messages))
return await original(
messages, model=model, temperature=temperature, response_format=response_format
)
provider.chat = recording_chat
messages = [Message(role="user", content="grade")]
await structured_completion(provider, messages, model="m", schema=Score, schema_hint=HINT)
assert captured, "provider was never called"
last_user = next(m for m in reversed(captured[0]) if m.role == "user")
assert HINT in last_user.content
assert "ONLY" in last_user.content # JSON-only instruction present
# Determinism: same request yields same reply
# Determinism check: same request yields same reply
again = await structured_completion(
provider, messages, model="m", schema=Score, schema_hint=HINT
)
@@ -1,76 +0,0 @@
"""Tutor agent tests — persona, Socratic structure, distinctness vs Coach (REQ-2-006)."""
from ai_service.agents.coach import CoachAgent
from ai_service.agents.tutor import TutorAgent
from ai_service.config import Settings
from ai_service.corpus.learner_context import get_learner_context
from ai_service.llm.mock import MockProvider
from ai_service.llm.types import Message
def make_tutor() -> TutorAgent:
return TutorAgent(MockProvider(), Settings(provider="mock"))
def test_system_prompt_includes_learner_context():
tutor = make_tutor()
prompt = tutor.system_prompt(get_learner_context())
assert "Alex Rivera" in prompt
assert "Socratic" in prompt
assert "ONE concept" in prompt
def test_system_prompt_marks_current_focus():
tutor = make_tutor()
prompt = tutor.system_prompt(get_learner_context())
assert "Multi-agent communication patterns" in prompt
def test_build_messages_system_history_user():
tutor = make_tutor()
history = [Message(role="user", content="q1"), Message(role="assistant", content="a1")]
messages = tutor.build_messages(history, "explain again", get_learner_context())
assert messages[0].role == "system"
assert "Tutor" in messages[0].content
assert [m.content for m in messages[1:]] == ["q1", "a1", "explain again"]
async def test_stream_reply_yields_deltas():
tutor = make_tutor()
tokens = [
t
async for t in tutor.stream_reply(
user_input="hi", learner_context=get_learner_context()
)
]
assert tokens
def test_coach_and_tutor_personas_are_distinct():
"""Distinct system prompts (P3 must-have)."""
ctx = get_learner_context()
coach_prompt = CoachAgent(MockProvider(), Settings(provider="mock")).system_prompt(ctx)
tutor_prompt = TutorAgent(MockProvider(), Settings(provider="mock")).system_prompt(ctx)
assert coach_prompt != tutor_prompt
assert "retrieval practice" in coach_prompt.lower()
assert "Socratic" in tutor_prompt
async def test_coach_and_tutor_stream_outputs_are_distinct():
"""Mock outputs differ because system prompts differ (hash-seeded on content)."""
ctx = get_learner_context()
settings = Settings(provider="mock")
async def full_reply(agent):
tokens = [
t async for t in agent.stream_reply(user_input="stuck", learner_context=ctx)
]
return "".join(tokens)
coach_tokens = await full_reply(CoachAgent(MockProvider(), settings))
tutor_tokens = await full_reply(TutorAgent(MockProvider(), settings))
assert coach_tokens != tutor_tokens
def test_agent_name():
assert make_tutor().name == "tutor"
@@ -1,87 +0,0 @@
"""Assessment API tests — stored-grade coaching contract (REQ-3-007)."""
from __future__ import annotations
import json
from datetime import UTC, datetime
import pytest
from fastapi.testclient import TestClient
from ai_service.config import Settings
from ai_service.grading.store import GradeRecord, SQLiteGradeStore
from ai_service.llm.mock import MockProvider
from ai_service.main import create_app
from ai_service.telemetry.ingest import TraceIntegrityMap
from ai_service.telemetry.store import SQLiteTraceStore
COACHING_JSON = json.dumps(
{
"summary": "Good iterative work.",
"strengths": ["Tests after changes."],
"gaps": ["Missing edge cases."],
"next_steps": ["Add an edge-case test."],
}
)
class CoachingMock(MockProvider):
def _reply_for(self, messages, response_format):
if response_format is not None and response_format.get("type") == "json_object":
return COACHING_JSON
return super()._reply_for(messages, response_format)
@pytest.fixture()
def client(tmp_path) -> TestClient:
app = create_app(Settings(provider="mock"))
app.state.provider = CoachingMock()
app.state.trace_store = SQLiteTraceStore(db_path=tmp_path / "t.db")
app.state.grade_store = SQLiteGradeStore(db_path=tmp_path / "g.db")
app.state.trace_integrity = TraceIntegrityMap()
with TestClient(app) as c:
yield c
def _seed_grade(client: TestClient) -> None:
app = client.app
store: SQLiteGradeStore = app.state.grade_store
store.save(
GradeRecord(
learner_id="api-learner",
task_id="api-task",
variant_seed=None,
digest={"error_fix_cycles": 1},
scores={"criteria": {"process_quality": 3}, "verdict": "developing"},
verdict="GRADED",
model="gemma4:31b",
created_at=datetime.now(UTC),
)
)
def test_evaluate_renders_stored_grade_as_coaching(client) -> None:
_seed_grade(client)
resp = client.post(
"/v1/assessment/evaluate",
json={"learner_id": "api-learner", "task_id": "api-task"},
)
assert resp.status_code == 200, resp.text
body = resp.json()
assert body["grade_verdict"] == "GRADED"
assert body["coaching"]["summary"]
def test_evaluate_without_grade_404(client) -> None:
resp = client.post(
"/v1/assessment/evaluate",
json={"learner_id": "nobody", "task_id": "nothing"},
)
assert resp.status_code == 404
assert "grade first" in resp.json()["detail"]
def test_missing_fields_422(client) -> None:
resp = client.post("/v1/assessment/evaluate", json={"learner_id": "x"})
assert resp.status_code == 422
@@ -94,16 +94,12 @@ def test_second_turn_replays_history_to_provider(client):
assert len(captured) == 2
turn1, turn2 = captured
# Agent routing (P3): provider now receives [system, ...history, user turn]
assert turn1[0].role == "system"
assert turn1[-1].content == "turn one"
assert len(turn1) == 2 # system + first user turn
assert [m.content for m in turn1] == ["turn one"]
turn2_contents = [m.content for m in turn2]
assert "turn one" in turn2_contents
assert "turn two" in turn2_contents
assert any(m.role == "assistant" for m in turn2) # persisted reply replayed
assert "turn two" == turn2_contents[-1] # new user turn last
assert turn2[0].role == "system" # every routed call starts with the persona
def test_history_replay_is_windowed(client):
@@ -135,60 +131,12 @@ def test_history_replay_is_windowed(client):
last_input = captured[-1]
contents = [m.content for m in last_input]
# system prompt + window(20) + the new user message
assert last_input[0].role == "system"
assert len(last_input) == 1 + 20 + 1
assert len(last_input) == 20 + 1 # window(20) + the new user message
assert "turn 0" not in contents # oldest messages trimmed out of replay
assert "turn 14" in contents
assert contents[-1] == "turn 14"
def test_replayed_system_prompt_carries_routed_persona(client):
"""History replay puts the routed agent's persona at position 0 (P3).
The system prompt on every turn including replays must match the
routed agent: Coach calls get the Coach persona, Tutor calls the Tutor
persona, and the two are observably distinct.
"""
from ai_service.llm.types import Message
def capture_two_turns(agent_name: str, session: str) -> list[list[Message]]:
provider = client.app.state.provider
original = provider.stream_chat
captured: list[list[Message]] = []
async def recording_stream(messages, *, model, temperature=0.7, response_format=None):
captured.append(list(messages))
async for t in original(
messages, model=model, temperature=temperature, response_format=response_format
):
yield t
provider.stream_chat = recording_stream
try:
for content in ("turn one", "turn two"):
stream_events(client, {
"agent": agent_name, "session_id": session,
"messages": [{"role": "user", "content": content}],
})
finally:
provider.stream_chat = original
return captured
coach_calls = capture_two_turns("coach", "persona-coach")
tutor_calls = capture_two_turns("tutor", "persona-tutor")
coach_replay_system = coach_calls[1][0]
tutor_replay_system = tutor_calls[1][0]
assert coach_replay_system.role == "system"
assert tutor_replay_system.role == "system"
assert "Coach" in coach_replay_system.content
assert "Tutor" in tutor_replay_system.content
assert coach_replay_system.content != tutor_replay_system.content
# persona is stable across turns within one session
assert coach_calls[0][0].content == coach_replay_system.content
def test_session_agent_scoped(client):
payload = {
"agent": "tutor",
@@ -205,37 +153,3 @@ def test_session_agent_scoped(client):
session = asyncio.run(check())
assert session.agent == "tutor"
def test_client_retry_does_not_duplicate_user_turn(client):
"""P1 fix (final review): resending the same user turn after a failure
must not double-append it to session history."""
import asyncio
payload = {
"agent": "coach",
"session_id": "sess-retry",
"messages": [{"role": "user", "content": "same question"}],
}
# First attempt fails mid-stream (turn was already appended before streaming)
provider = client.app.state.provider
provider.fail_mid_stream_at_index = 0
stream_events_retry(client, payload)
provider.fail_mid_stream_at_index = None
# Client retry: identical payload
stream_events_retry(client, payload)
store = client.app.state.session_store
async def check():
return await store.history_window("sess-retry")
contents = [m.content for m in asyncio.run(check())]
user_turns = [c for c in contents if c == "same question"]
assert len(user_turns) == 1, f"expected exactly 1 stored user turn, got {len(user_turns)}"
def stream_events_retry(client, payload):
with client.stream("POST", "/v1/chat/stream", json=payload) as response:
for _ in response.iter_lines():
pass
@@ -75,57 +75,4 @@ def test_pre_first_byte_failure_yields_provider_unavailable(client):
error_events = [e for e in events if e["type"] == "error"]
assert len(error_events) == 1
assert error_events[0]["code"] == "provider_unavailable"
def test_unknown_agent_rejected_422(client):
response = client.post(
"/v1/chat/stream",
json={
"agent": "oracle",
"session_id": "s",
"messages": [{"role": "user", "content": "hi"}],
},
)
assert response.status_code == 422
assert "oracle" in response.json()["detail"]
def test_coach_routes_and_meta_names_agent(client):
payload = {
"agent": "coach",
"session_id": "route-coach",
"messages": [{"role": "user", "content": "pace me"}],
}
events = parse_events(stream_lines(client, payload))
assert events[0]["type"] == "meta"
assert events[0]["agent"] == "coach"
assert events[-1]["type"] == "[DONE]"
def test_tutor_routes_and_meta_names_agent(client):
payload = {
"agent": "tutor",
"session_id": "route-tutor",
"messages": [{"role": "user", "content": "teach me"}],
}
events = parse_events(stream_lines(client, payload))
assert events[0]["type"] == "meta"
assert events[0]["agent"] == "tutor"
assert events[-1]["type"] == "[DONE]"
def test_coach_and_tutor_streams_are_distinct(client):
"""Agent routing selects the right agent: distinct system prompts →
distinct hash-seeded mock outputs for the same user input."""
same_message = [{"role": "user", "content": "same question"}]
coach = parse_events(stream_lines(client, {
"agent": "coach", "session_id": "d1", "messages": same_message,
}))
tutor = parse_events(stream_lines(client, {
"agent": "tutor", "session_id": "d2", "messages": same_message,
}))
coach_text = "".join(e["content"] for e in coach if e["type"] == "delta")
tutor_text = "".join(e["content"] for e in tutor if e["type"] == "delta")
assert coach_text and tutor_text
assert coach_text != tutor_text
assert coach[-1]["type"] == "[DONE]"
-82
View File
@@ -1,82 +0,0 @@
"""CORS policy tests (A-008, D-038 network mode).
v0.3 initially shipped `allow_methods` WITHOUT "PUT" while the learner
build surface writes workspace files with PUT (engine-client writeFile)
every cross-origin Save failed preflight. These tests pin the policy so a
future method-list edit fails loudly instead of silently breaking the
headline flow.
v0.3.5 network mode (D-038): the default AI_CORS_ORIGINS='*' admits any
origin (safe ONLY because credentials are never enabled); an explicit list
restricts. Both modes are pinned here:
- wildcard: remote origin gets the grant; credentials still never sent;
- explicit: unlisted origins get no grant.
"""
from __future__ import annotations
from fastapi.testclient import TestClient
ALLOWED_ORIGIN = "http://localhost:3000"
REMOTE_ORIGIN = "http://nextcraft-1:3000"
ALL_CLIENT_METHODS = ("GET", "POST", "PUT", "DELETE")
def _allow_origin(resp) -> str | None:
return resp.headers.get("access-control-allow-origin")
def test_preflight_allows_every_method_the_web_client_uses(client: TestClient) -> None:
for method in ALL_CLIENT_METHODS:
resp = client.options(
"/v1/sandboxes",
headers={
"Origin": ALLOWED_ORIGIN,
"Access-Control-Request-Method": method,
},
)
assert resp.status_code == 200, f"preflight {method} failed: {resp.status_code}"
assert _allow_origin(resp) in ("*", ALLOWED_ORIGIN)
allowed = resp.headers["access-control-allow-methods"].split(", ")
assert method in allowed, f"{method} missing from CORS methods: {allowed}"
def test_cross_origin_get_echoes_allow_origin(client: TestClient) -> None:
resp = client.get("/v1/sandboxes", headers={"Origin": ALLOWED_ORIGIN})
assert resp.status_code == 200
assert _allow_origin(resp) in ("*", ALLOWED_ORIGIN)
def test_wildcard_mode_grants_remote_origins(client: TestClient) -> None:
"""D-038 default: '*' grants any origin — remote browsers work zero-config."""
resp = client.get("/v1/sandboxes", headers={"Origin": REMOTE_ORIGIN})
assert resp.status_code == 200
assert _allow_origin(resp) in ("*", REMOTE_ORIGIN)
def test_explicit_list_mode_denies_unlisted_origins(
settings, monkeypatch, tmp_path
) -> None:
"""Explicit AI_CORS_ORIGINS restricts to the listed origins only."""
from fastapi.testclient import TestClient as TC
from ai_service.main import create_app
restricted = settings.model_copy(update={"cors_origins": "http://localhost:3000"})
app = create_app(restricted)
with TC(app) as c:
resp = c.get("/v1/sandboxes", headers={"Origin": "https://evil.example"})
assert resp.status_code == 200 # non-CORS requests still serve
assert resp.headers.get("access-control-allow-origin") is None
def test_credentials_never_allowed(client: TestClient) -> None:
resp = client.options(
"/v1/sandboxes",
headers={
"Origin": ALLOWED_ORIGIN,
"Access-Control-Request-Method": "PUT",
"Access-Control-Request-Headers": "Content-Type",
},
)
assert resp.headers.get("access-control-allow-credentials") != "true"
-234
View File
@@ -1,234 +0,0 @@
"""Defense endpoint tests (Task 5-3-01, REQ-3-006) — mock voice + mock LLM."""
from __future__ import annotations
import json
from pathlib import Path
import pytest
from fastapi.testclient import TestClient
from ai_service.agents.examiner import ExaminerAgent
from ai_service.config import Settings
from ai_service.grading.store import SQLiteGradeStore
from ai_service.llm.mock import MockProvider
from ai_service.main import create_app
from ai_service.telemetry.ingest import TraceIntegrityMap
from ai_service.telemetry.store import SQLiteTraceStore
from ai_service.variants.store import SQLiteVariantStore
from ai_service.voice.defense_store import SQLiteDefenseStore
from ai_service.voice.mock import MockVoiceProvider
VERDICT = {
"verdict": "developing",
"understanding": "Explains the build clearly.",
"process_justification": "Justifies choices.",
"communication": "Clear and specific.",
"strengths": ["Grounded answers in the digest."],
"gaps": ["Did not address the edge cases."],
}
class ScriptedLLM(MockProvider):
"""Question-mode calls get a question; verdict-mode calls get D-020 JSON.
Discriminator: the verdict prompt contains "final verdict JSON" the
question prompt says "next question".
"""
def __init__(self) -> None:
super().__init__()
async def chat(self, messages, *, model, temperature=0.7, response_format=None):
all_text = "\n".join(m.content for m in messages)
if "final verdict JSON" in all_text:
return json.dumps(VERDICT)
return "Why did you structure the fix that way?"
@pytest.fixture()
def app(tmp_path: Path):
application = create_app(Settings(provider="mock", voice_provider="mock"))
llm = ScriptedLLM()
application.state.provider = llm
application.state.trace_store = SQLiteTraceStore(db_path=tmp_path / "t.db")
application.state.variant_store = SQLiteVariantStore(db_path=tmp_path / "v.db")
application.state.grade_store = SQLiteGradeStore(db_path=tmp_path / "g.db")
application.state.trace_integrity = TraceIntegrityMap()
application.state.defense_store = SQLiteDefenseStore(db_path=tmp_path / "d.db")
application.state.voice_provider = MockVoiceProvider(
["the fix was in the retry loop"]
)
settings = Settings(provider="mock", voice_provider="mock")
application.state.examiner_agent = ExaminerAgent(llm, settings)
return application
@pytest.fixture()
def client(app) -> TestClient:
with TestClient(app) as c:
yield c
def _start(client: TestClient) -> dict:
resp = client.post(
"/v1/defense/start", json={"learner_id": "defense-learner", "task_id": "defense-task"}
)
assert resp.status_code == 200, resp.text
return resp.json()
class TestStart:
def test_start_returns_first_question_and_descriptor(self, client) -> None:
body = _start(client)
assert body["first_question"]
assert body["defense_id"]
assert body["voice_descriptor"]["mode"] == "mock"
assert body["trace_complete"] is True
stored = client.get(f"/v1/defense/{body['defense_id']}")
assert stored.status_code == 200
turns = stored.json()["turns"]
assert turns and turns[0]["role"] == "examiner"
def test_start_with_unknown_trace_is_complete_flag(self, client) -> None:
body = _start(client)
assert body["trace_complete"] is True
class TestBrowserFallback:
def test_browser_mode_serves_browser_descriptor(self, tmp_path: Path) -> None:
"""Must-Have #6: AI_VOICE_PROVIDER=browser → start returns the
browser-native SR/TTS fallback descriptor (D-030), not 'mock'."""
application = create_app(Settings(provider="mock", voice_provider="browser"))
application.state.provider = ScriptedLLM()
application.state.trace_store = SQLiteTraceStore(db_path=tmp_path / "t.db")
application.state.variant_store = SQLiteVariantStore(db_path=tmp_path / "v.db")
application.state.grade_store = SQLiteGradeStore(db_path=tmp_path / "g.db")
application.state.trace_integrity = TraceIntegrityMap()
application.state.defense_store = SQLiteDefenseStore(db_path=tmp_path / "d.db")
application.state.voice_provider = MockVoiceProvider(["answer"])
llm = ScriptedLLM()
application.state.examiner_agent = ExaminerAgent(
llm, Settings(provider="mock", voice_provider="browser")
)
with TestClient(application) as c:
body = _start(c)
assert body["voice_descriptor"]["mode"] == "browser"
assert body["voice_descriptor"]["sr_available"]
assert "SpeechRecognition" in body["voice_descriptor"]["hint"]
class TestAnswer:
def test_typed_answer_yields_followup_with_latency(self, client) -> None:
defense_id = _start(client)["defense_id"]
resp = client.post(
f"/v1/defense/{defense_id}/answer", data={"text": "I fixed the loop."}
)
assert resp.status_code == 200, resp.text
body = resp.json()
assert body["question"]
assert body["turn_latency"]["llm_ms"] is not None
def test_audio_answer_transcribed_and_recorded(self, client) -> None:
defense_id = _start(client)["defense_id"]
wav_bytes = b"RIFF" + b"\x00" * 64
resp = client.post(
f"/v1/defense/{defense_id}/answer",
files={"audio": ("answer.wav", wav_bytes, "audio/wav")},
)
assert resp.status_code == 200, resp.text
stored = client.get(f"/v1/defense/{defense_id}").json()
learner_turns = [t for t in stored["turns"] if t["role"] == "learner"]
assert learner_turns, "learner turn missing after audio answer"
assert learner_turns[0]["text"] == "the fix was in the retry loop"
def test_empty_audio_is_422_not_500(self, client) -> None:
"""Zero-byte upload must 422 before the provider call (a real
provider would raise the same way the mock does validate first)."""
defense_id = _start(client)["defense_id"]
resp = client.post(
f"/v1/defense/{defense_id}/answer",
files={"audio": ("answer.wav", b"", "audio/wav")},
)
assert resp.status_code == 422, resp.text
stored = client.get(f"/v1/defense/{defense_id}").json()
assert len(stored["turns"]) == 1 # nothing appended
def test_answer_after_finish_is_409(self, client) -> None:
"""A sealed transcript is append-only-no-more: the endpoints own
turn-vs-finalize sequencing (defense_store contract)."""
defense_id = _start(client)["defense_id"]
client.post(f"/v1/defense/{defense_id}/answer", data={"text": "a"})
assert client.post(f"/v1/defense/{defense_id}/finish").status_code == 200
resp = client.post(f"/v1/defense/{defense_id}/answer", data={"text": "late"})
assert resp.status_code == 409, resp.text
stored = client.get(f"/v1/defense/{defense_id}").json()
assert len(stored["turns"]) == 3 # ex, lrn, ex — no post-finish turns
def test_neither_text_nor_audio_422(self, client) -> None:
defense_id = _start(client)["defense_id"]
resp = client.post(f"/v1/defense/{defense_id}/answer")
assert resp.status_code == 422
def test_unknown_defense_404(self, client) -> None:
resp = client.post("/v1/defense/dfn-nope/answer", data={"text": "hi"})
assert resp.status_code == 404
class TestAudioEndpoint:
def test_examiner_turn_streams_wav(self, client) -> None:
defense_id = _start(client)["defense_id"]
resp = client.get(f"/v1/defense/{defense_id}/audio/0")
assert resp.status_code == 200
assert resp.content
assert resp.headers["content-type"].startswith("audio/")
def test_unknown_turn_404(self, client) -> None:
defense_id = _start(client)["defense_id"]
assert client.get(f"/v1/defense/{defense_id}/audio/42").status_code == 404
class TestFinishAndGet:
def test_full_loop_verdict_and_signals(self, client) -> None:
defense_id = _start(client)["defense_id"]
client.post(f"/v1/defense/{defense_id}/answer", data={"text": "answer one"})
finish = client.post(f"/v1/defense/{defense_id}/finish")
assert finish.status_code == 200, finish.text
body = finish.json()
assert body["verdict"]["verdict"] == "developing"
assert body["integrity_signals"]["pause_threshold_ms"]
stored = client.get(f"/v1/defense/{defense_id}").json()
assert stored["status"] == "finished"
assert stored["integrity_signals"]
# Must-Have #1: "verdict + transcript persisted" — the verdict must
# be retrievable from GET after finish, not only in the finish body.
assert stored["integrity_signals"]["verdict"]["verdict"] == "developing"
def test_long_pause_flagged(self, client, app) -> None:
from datetime import UTC, datetime
from ai_service.voice.defense_store import DefenseTurn
defense_id = _start(client)["defense_id"]
# inject a slow learner turn directly (simulated latency)
store = app.state.defense_store
store.append_turn(
defense_id,
DefenseTurn(
defense_id=defense_id,
seq=99,
role="learner",
text="slow reply",
ts=datetime.now(UTC),
latency_ms=30_000,
created_at=datetime.now(UTC),
),
)
finish = client.post(f"/v1/defense/{defense_id}/finish")
assert finish.status_code == 200
signals = finish.json()["integrity_signals"]
assert any(p["turn"] == 99 for p in signals["long_pauses"])
def test_unknown_defense_404_on_all(self, client) -> None:
assert client.post("/v1/defense/dfn-nope/finish").status_code == 404
assert client.get("/v1/defense/dfn-nope").status_code == 404
@@ -1,216 +0,0 @@
"""Full credential-flow E2E over real engines (Task 6-5-01, REQ-3-007/008).
Endpoint-level end-to-end with mock LLM/voice providers (G-2 precedent:
real engine plumbing over real endpoints; provider choice is
service-internal): variant -> telemetry-wired sandbox -> real in-sandbox
exec -> trace -> grade -> oral defense -> verdict/signals -> proctor.
No corpus fixture anywhere in the flow.
Probe-guarded for user namespaces (the in-sandbox exec needs them).
"""
from __future__ import annotations
import asyncio
import contextlib
import json
import socket
import time
from pathlib import Path
import pytest
import uvicorn
from ai_service.config import Settings
from ai_service.grading.store import SQLiteGradeStore
from ai_service.llm.mock import MockProvider
from ai_service.main import create_app
from ai_service.telemetry.ingest import TraceIntegrityMap
from ai_service.telemetry.store import SQLiteTraceStore
from ai_service.variants.store import SQLiteVariantStore
from ai_service.voice.defense_store import SQLiteDefenseStore
from tests.sandbox.test_isolation import USERSNS_AVAILABLE
COACHING_JSON = json.dumps(
{
"summary": "Strong iteration.",
"strengths": ["Tests early."],
"gaps": ["One edge case missing."],
"next_steps": ["Add it."],
}
)
VERDICT_JSON = json.dumps(
{
"verdict": "developing",
"understanding": "Explains the build clearly.",
"process_justification": "Choices defended.",
"communication": "Clear.",
"strengths": ["Grounded in the digest."],
"gaps": ["Missed one edge case."],
}
)
PROCTOR_JSON = json.dumps(
{
"signals": [
{"signal_type": "idle_gap", "severity": "low", "note": "A short pause."}
],
"intervention": "Keep momentum.",
"summary": "Healthy session.",
}
)
class FlowLLM(MockProvider):
"""Prompt-discriminated: verdict vs coaching vs question vs proctor JSON."""
def _reply_for(self, messages, response_format):
all_text = "\n".join(m.content for m in messages)
if response_format is not None and response_format.get("type") == "json_object":
if "final verdict JSON" in all_text:
return VERDICT_JSON
if "rubric" in all_text.lower() and "Score this build session" in all_text:
return json.dumps(
{
"criteria": {
"process_quality": 4,
"correctness": 3,
"debugging_discipline": 3,
"test_usage": 4,
},
"strengths": ["Iterated with tests."],
"gaps": ["One edge case missing."],
"verdict": "developing",
}
)
if "Explain it as coaching" in all_text:
return COACHING_JSON
if "integrity signals supportively" in all_text:
return PROCTOR_JSON
return json.dumps(
{
"statement": (
"Build a judge for code-review answers scoring factual "
"accuracy with 3 edge cases and 5 test examples."
)
}
)
return "Walk me through your last fix — what changed and why?"
def _free_port() -> int:
with socket.socket() as s:
s.bind(("127.0.0.1", 0))
return s.getsockname()[1]
@pytest.mark.asyncio
async def test_full_credential_flow(tmp_path: Path) -> None:
if not USERSNS_AVAILABLE:
pytest.skip("user namespaces unavailable on this host (probe)")
import httpx
port = _free_port()
settings = Settings(provider="mock", voice_provider="mock", port=port)
app = create_app(settings)
app.state.provider = FlowLLM()
app.state.trace_store = SQLiteTraceStore(db_path=tmp_path / "t.db")
app.state.grade_store = SQLiteGradeStore(db_path=tmp_path / "g.db")
app.state.variant_store = SQLiteVariantStore(db_path=tmp_path / "v.db")
app.state.defense_store = SQLiteDefenseStore(db_path=tmp_path / "d.db")
app.state.trace_integrity = TraceIntegrityMap()
server = uvicorn.Server(uvicorn.Config(app, host="127.0.0.1", port=port, log_level="warning"))
serve_task = asyncio.get_running_loop().create_task(server.serve())
try:
for _ in range(100):
if server.started:
break
await asyncio.sleep(0.1)
assert server.started
base = f"http://127.0.0.1:{port}"
async with httpx.AsyncClient(base_url=base, timeout=30.0) as client:
# 1. Variant (real seeded generation, mock-rendered).
var = (await client.post("/v1/variants", json={
"learner_id": "pilot-learner", "competency_id": "stack-orchestration-c007",
})).json()
assert var["task_id"] and var["statement"] and var["starter_files"]
task_id = var["task_id"]
# 2. Telemetry-wired sandbox (real namespaces; agent joins via loopback).
sbx = (await client.post("/v1/sandboxes", json={
"learner_id": "pilot-learner", "task_id": task_id,
}))
assert sbx.status_code == 201, sbx.text
sandbox_id = sbx.json()["id"]
# 3. Real in-sandbox exec: run the starter test (capture agent streams
# the workspace effects into the trace).
exec_resp = await client.post(f"/v1/sandboxes/{sandbox_id}/exec", json={
"cmd": ["pytest", "-q"],
})
assert exec_resp.status_code == 200, exec_resp.text
# 4. Trace: events landed in order (the capture agent runs async).
trace: list = []
deadline = time.monotonic() + 20.0
while time.monotonic() < deadline:
tr = await client.get(f"/v1/telemetry/traces/pilot-learner/{task_id}")
if tr.status_code == 200:
trace = tr.json().get("events", [])
if trace:
break
await asyncio.sleep(0.25)
assert trace, "no telemetry events arrived from the real sandbox"
seqs = [e["seq"] for e in trace]
assert seqs == sorted(seqs)
# 5. Grade: rubric from the real digest (G-4 gate passed: no gaps).
grade = (await client.post("/v1/assessment/grade", json={
"learner_id": "pilot-learner", "task_id": task_id,
})).json()
assert grade["verdict"] == "GRADED", grade
assert grade["scores"]["criteria"]["process_quality"] == 4
assert grade["variant_seed"] == var["seed"] # D-029 stamped
# 6. Assessor coaching FROM the stored grade.
coaching = (await client.post("/v1/assessment/evaluate", json={
"learner_id": "pilot-learner", "task_id": task_id,
})).json()
assert coaching["coaching"]["summary"]
# 7. Oral defense: start -> typed answers -> finish (mock voice).
defense = (await client.post("/v1/defense/start", json={
"learner_id": "pilot-learner", "task_id": task_id,
})).json()
assert defense["first_question"]
did = defense["defense_id"]
ans = await client.post(f"/v1/defense/{did}/answer", data={"text": "I fixed the loop."})
assert ans.status_code == 200, ans.text
finish = (await client.post(f"/v1/defense/{did}/finish")).json()
assert finish["verdict"]["verdict"] == "developing"
assert finish["integrity_signals"]
# 8. Proctor over the real digest + defense signals.
proctor = (await client.post("/v1/proctor/signals", json={
"learner_id": "pilot-learner", "task_id": task_id,
})).json()
assert proctor["intervention"]
# 9. Sandbox destroyed; no leaks.
destroy = await client.delete(f"/v1/sandboxes/{sandbox_id}")
assert destroy.status_code == 204
listed = (await client.get("/v1/sandboxes")).json()
assert all(s["id"] != sandbox_id for s in (listed.get("sandboxes") or []))
# 10. No corpus fixtures anywhere in this flow's payloads.
corpus_markers = ("lab-scenario", "proctor-scenario", "artifact-")
for payload in (var, grade, defense, finish, proctor):
assert not any(
m in json.dumps(payload) for m in corpus_markers
), "corpus fixture leaked into the learner path"
finally:
server.should_exit = True
with contextlib.suppress(Exception):
await asyncio.wait_for(serve_task, timeout=10.0)
-397
View File
@@ -1,397 +0,0 @@
"""Assessment grade endpoint tests — grading engine over HTTP (Task 3-3-01).
Contract under test (api/assessment.py, REQ-3-004):
POST /v1/assessment/grade {learner_id, task_id}
GRADED 200, rubric scores + digest summary
UNGRADABLE_TRACE_INCOMPLETE 200, gate record (missing_seqs /
integrity_flag surfaced in scores)
UNGRADABLE_EMPTY_TRACE 200, gate record (documented choice: an
unknown task is ALSO an empty pair; the
engine cannot distinguish, and the gate
outcome is a durable first-class result)
StructuredOutputError 502 (provider exhausted the D-020 budget)
GET /v1/assessment/grade/{learner_id}/{task_id}
stored latest grade 200 (same scores as the POST)
never-graded pair 404
Wiring: per-test tmp-path SQLite stores + a pre-set GradingEngine
(state-injection override the lifespan adopts trace_store/grade_store/
trace_integrity/grading_engine from app.state instead of constructing
them; same pattern as test_telemetry_ingest.py / test_sandboxes.py). The
engine binds a scripted provider so each test controls the LLM exactly,
and counts calls so gate tests can assert the LLM was never reached
(G-4 holds through the whole HTTP stack).
Zero network: providers are MockProvider subclasses only (conftest rule).
"""
from __future__ import annotations
from collections.abc import Iterator
from datetime import UTC, datetime, timedelta
from pathlib import Path
import pytest
from fastapi.testclient import TestClient
from ai_service.config import Settings
from ai_service.corpus.trace_fixtures import STRONG_BUILDER
from ai_service.grading.engine import GradingEngine
from ai_service.grading.store import SQLiteGradeStore
from ai_service.llm.mock import MockProvider, ScriptedJSONProvider
from ai_service.main import create_app
from ai_service.telemetry.ingest import TraceIntegrityMap
from ai_service.telemetry.models import TelemetryEvent
from ai_service.telemetry.store import SQLiteTraceStore
LEARNER = "grade-learner"
TASK = "task-grade-1"
T0 = datetime(2026, 9, 12, 3, 0, 0, tzinfo=UTC)
#: Canonical well-formed rubric payload (matches grading.engine.RubricScore).
RUBRIC_PAYLOAD: dict = {
"criteria": {
"process_quality": 4,
"correctness": 4,
"debugging_discipline": 3,
"test_usage": 4,
},
"strengths": ["tight edit-test loops throughout"],
"gaps": ["final commit discipline loose"],
"verdict": "mastered",
}
class CountingScriptedProvider(ScriptedJSONProvider):
"""ScriptedJSONProvider that counts chat() calls.
Gate tests assert calls == 0 (the LLM is never reached through the
whole HTTP stack G-4); happy-path tests sanity-check calls >= 1.
"""
def __init__(self, payload: dict) -> None:
super().__init__(payload)
self.calls = 0
async def chat(self, messages, *, model, temperature=0.7, response_format=None):
self.calls += 1
return await super().chat(
messages, model=model, temperature=temperature, response_format=response_format
)
def _event(seq: int, kind: str, payload: dict | None = None, offset_s: float = 0.0):
return TelemetryEvent(
learner_id=LEARNER,
task_id=TASK,
seq=seq,
kind=kind,
payload=payload or {},
ts=T0 + timedelta(seconds=offset_s),
sandbox_id="sbx-grade",
)
def _complete_trace() -> list[TelemetryEvent]:
"""Contiguous seq 0..6 — a gradeable trace (gap-free, unflagged)."""
return [
_event(0, "activity", {"state": "starting"}, 0.0),
_event(1, "file_diff", {"path": "a.py", "added": 12}, 10.0),
_event(2, "command", {"cmd": "pytest -q"}, 20.0),
_event(3, "test_result", {"passed": False, "exit_code": 1}, 25.0),
_event(4, "file_diff", {"path": "a.py", "added": 4, "removed": 2}, 40.0),
_event(5, "command", {"cmd": "pytest -q"}, 60.0),
_event(6, "test_result", {"passed": True, "exit_code": 0}, 65.0),
]
@pytest.fixture()
def trace_store(tmp_path: Path) -> Iterator[SQLiteTraceStore]:
store = SQLiteTraceStore(db_path=tmp_path / "traces.db")
yield store
store.close()
@pytest.fixture()
def grade_store(tmp_path: Path) -> Iterator[SQLiteGradeStore]:
store = SQLiteGradeStore(db_path=tmp_path / "grades.db")
yield store
store.close()
@pytest.fixture()
def integrity() -> TraceIntegrityMap:
return TraceIntegrityMap()
def _make_client(
tmp_path: Path,
trace_store: SQLiteTraceStore,
grade_store: SQLiteGradeStore,
integrity: TraceIntegrityMap,
provider: MockProvider,
) -> TestClient:
"""App + TestClient with stores, integrity map and a pre-set engine.
The lifespan adopts every pre-set service (state-injection override);
the engine binds OUR provider, so the fixture not Settings scripts
the LLM. Cloud-free guard: the provider must be a MockProvider family
member (conftest rule, enforced here because this module builds its
own client rather than consuming the conftest one).
"""
assert isinstance(provider, MockProvider)
settings = Settings(
provider="mock",
db_path=tmp_path / "grading-test.db",
sandbox_dir=tmp_path / "sandboxes",
)
app = create_app(settings)
app.state.trace_store = trace_store
app.state.grade_store = grade_store
app.state.trace_integrity = integrity
app.state.grading_engine = GradingEngine(
trace_store, grade_store, integrity, provider, model="gemma4:31b"
)
return TestClient(app)
def _seed(store: SQLiteTraceStore, events: list[TelemetryEvent]) -> None:
for e in events:
store.append(e)
def _post(client: TestClient, learner: str = LEARNER, task: str = TASK):
return client.post("/v1/assessment/grade", json={"learner_id": learner, "task_id": task})
def _get(client: TestClient, learner: str = LEARNER, task: str = TASK):
return client.get(f"/v1/assessment/grade/{learner}/{task}")
# -- happy path: complete trace → GRADED -----------------------------------------
class TestGraded:
def test_post_complete_trace_returns_rubric_scores_and_digest(
self, tmp_path, trace_store, grade_store, integrity
):
_seed(trace_store, _complete_trace())
provider = CountingScriptedProvider(RUBRIC_PAYLOAD)
with _make_client(tmp_path, trace_store, grade_store, integrity, provider) as client:
response = _post(client)
assert response.status_code == 200
body = response.json()
assert body["verdict"] == "GRADED"
assert body["learner_id"] == LEARNER
assert body["task_id"] == TASK
assert body["scores"] == RUBRIC_PAYLOAD # rubric verdict rides in scores
assert body["scores"]["verdict"] == "mastered"
assert body["variant_seed"] is None # null until P4 (D-029)
assert body["model"] == "gemma4:31b" # provenance travels
assert provider.calls == 1 # happy path: exactly one LLM call
# digest summary rides along (D-028 reproducible input)
assert body["digest"]["event_count"] == 7
assert body["digest"]["final_test_status"] == "pass"
def test_corpus_fixture_trace_grades(
self, tmp_path, trace_store, grade_store, integrity
):
"""The strong-builder corpus fixture (real calibration trace) is
gradeable over HTTP end to end."""
_seed(trace_store, STRONG_BUILDER.events)
provider = CountingScriptedProvider(RUBRIC_PAYLOAD)
with _make_client(tmp_path, trace_store, grade_store, integrity, provider) as client:
response = _post(
client, learner=STRONG_BUILDER.learner_id, task=STRONG_BUILDER.task_id
)
assert response.status_code == 200
body = response.json()
assert body["verdict"] == "GRADED"
assert body["digest"]["event_count"] == len(STRONG_BUILDER.event_specs)
def test_get_after_post_returns_same_stored_scores(
self, tmp_path, trace_store, grade_store, integrity
):
_seed(trace_store, _complete_trace())
provider = CountingScriptedProvider(RUBRIC_PAYLOAD)
with _make_client(tmp_path, trace_store, grade_store, integrity, provider) as client:
posted = _post(client)
assert posted.status_code == 200
fetched = _get(client)
assert fetched.status_code == 200
stored = fetched.json()
assert stored["scores"] == posted.json()["scores"]
assert stored["verdict"] == "GRADED"
assert stored["digest"] == posted.json()["digest"]
assert stored["created_at"] == posted.json()["created_at"]
def test_post_regrade_upserts_get_returns_latest(
self, tmp_path, trace_store, grade_store, integrity
):
"""POST twice → one row holding the LATEST grade (GradeStore upsert)."""
_seed(trace_store, _complete_trace())
provider = CountingScriptedProvider(RUBRIC_PAYLOAD)
with _make_client(tmp_path, trace_store, grade_store, integrity, provider) as client:
first = _post(client)
assert first.status_code == 200
assert first.json()["scores"]["verdict"] == "mastered"
# the same provider now scripts a different (weaker) rubric
updated = {
**RUBRIC_PAYLOAD,
"criteria": {**RUBRIC_PAYLOAD["criteria"], "process_quality": 1},
"verdict": "not_yet",
}
provider.payload = updated
second = _post(client)
assert second.status_code == 200
assert second.json()["scores"] == updated
# GET returns the LATEST grade, not the first
fetched = _get(client)
assert fetched.json()["scores"] == updated
assert fetched.json()["scores"]["verdict"] == "not_yet"
# one row, latest wins (upsert, not append)
grades = grade_store.list_for_learner(LEARNER)
assert len(grades) == 1
assert grades[0].scores == updated
# -- G-4 gate outcomes over HTTP: 200 with the ungradable record -------------------
class TestGateOutcomes:
def test_gapped_trace_200_ungradable_incomplete_with_missing_seqs(
self, tmp_path, trace_store, grade_store, integrity
):
_seed(trace_store, [e for e in _complete_trace() if e.seq != 2]) # seq 2 missing
provider = CountingScriptedProvider(RUBRIC_PAYLOAD)
with _make_client(tmp_path, trace_store, grade_store, integrity, provider) as client:
response = _post(client)
assert response.status_code == 200 # ungradable IS a valid result — not 5xx
body = response.json()
assert body["verdict"] == "UNGRADABLE_TRACE_INCOMPLETE"
assert body["scores"]["missing_seqs"] == [2]
assert body["scores"]["integrity_flag"] is None
assert body["digest"] == {} # nothing was graded
assert body["model"] == "none" # no LLM involved (honest provenance)
assert provider.calls == 0, "LLM was called despite a gapped trace (G-4)"
def test_flooded_trace_200_ungradable_incomplete_with_integrity_reason(
self, tmp_path, trace_store, grade_store, integrity
):
"""Integrity-flagged trace (G-3 INCOMPLETE_FLOODED) surfaces the flag
reason in scores the same verdict as gaps, different detail."""
_seed(trace_store, _complete_trace())
integrity.mark(LEARNER, TASK, "INCOMPLETE_FLOODED")
provider = CountingScriptedProvider(RUBRIC_PAYLOAD)
with _make_client(tmp_path, trace_store, grade_store, integrity, provider) as client:
response = _post(client)
assert response.status_code == 200
body = response.json()
assert body["verdict"] == "UNGRADABLE_TRACE_INCOMPLETE"
assert body["scores"]["integrity_flag"] == "INCOMPLETE_FLOODED"
assert body["scores"]["missing_seqs"] == [] # rows complete but untrusted
assert provider.calls == 0, "LLM was called despite an integrity flag (G-4)"
def test_empty_trace_200_ungradable_empty(
self, tmp_path, trace_store, grade_store, integrity
):
"""Zero events for the pair (including a task that never had a
trace) UNGRADABLE_EMPTY_TRACE, persisted like any grade."""
provider = CountingScriptedProvider(RUBRIC_PAYLOAD)
with _make_client(tmp_path, trace_store, grade_store, integrity, provider) as client:
response = _post(client, learner="ghost-learner", task="never-started")
assert response.status_code == 200
body = response.json()
assert body["verdict"] == "UNGRADABLE_EMPTY_TRACE"
assert body["scores"] == {"integrity_flag": None, "missing_seqs": []}
assert body["digest"] == {}
assert body["model"] == "none"
assert provider.calls == 0
# the gate record is durable: GET now returns it (no longer 404)
with _make_client(tmp_path, trace_store, grade_store, integrity, provider) as client:
fetched = _get(client, learner="ghost-learner", task="never-started")
assert fetched.status_code == 200
assert fetched.json()["verdict"] == "UNGRADABLE_EMPTY_TRACE"
# -- provider failure → 502 --------------------------------------------------------
class TestProviderFailure:
def test_persistent_malformed_llm_output_502(
self, tmp_path, trace_store, grade_store, integrity
):
"""Stock MockProvider: its json_object reply is wrong-shaped, so the
D-020 defense exhausts both attempts endpoint maps to 502 and
nothing is persisted."""
_seed(trace_store, _complete_trace())
provider = MockProvider() # always malformed for the rubric schema
with _make_client(tmp_path, trace_store, grade_store, integrity, provider) as client:
response = _post(client)
assert response.status_code == 502
assert "grading failed" in response.json()["detail"]
# no fabricated/partial record was persisted
assert grade_store.get(LEARNER, TASK) is None
def test_502_leaves_earlier_grade_intact(
self, tmp_path, trace_store, grade_store, integrity
):
"""A failed regrade must not clobber the previously stored grade."""
_seed(trace_store, _complete_trace())
scripted = CountingScriptedProvider(RUBRIC_PAYLOAD)
with _make_client(tmp_path, trace_store, grade_store, integrity, scripted) as client:
assert _post(client).status_code == 200
# regrade attempt hits a provider that now always fails
broken = MockProvider()
with _make_client(tmp_path, trace_store, grade_store, integrity, broken) as client:
assert _post(client).status_code == 502
fetched = _get(client)
assert fetched.status_code == 200 # earlier grade still readable
assert fetched.json()["scores"] == RUBRIC_PAYLOAD
# -- stored-grade reads ------------------------------------------------------------
class TestStoredGradeReads:
def test_get_unknown_pair_404(self, tmp_path, trace_store, grade_store, integrity):
provider = CountingScriptedProvider(RUBRIC_PAYLOAD)
with _make_client(tmp_path, trace_store, grade_store, integrity, provider) as client:
response = _get(client, learner="nobody", task="never-graded")
assert response.status_code == 404
assert "no stored grade" in response.json()["detail"]
def test_missing_body_fields_422(self, tmp_path, trace_store, grade_store, integrity):
provider = CountingScriptedProvider(RUBRIC_PAYLOAD)
with _make_client(tmp_path, trace_store, grade_store, integrity, provider) as client:
assert client.post("/v1/assessment/grade", json={}).status_code == 422
assert (
client.post(
"/v1/assessment/grade", json={"learner_id": LEARNER}
).status_code
== 422
)
-86
View File
@@ -1,86 +0,0 @@
"""Lab endpoint tests — LIVE trace contract (REQ-3-007).
v0.3 re-grounding: POST /v1/lab/feedback takes {learner_id, task_id}; the
digest is computed from the learner's real TraceStore events. No corpus
scenarios; empty trace is a valid "no telemetry yet" coaching path.
"""
from __future__ import annotations
from datetime import UTC, datetime, timedelta
from pathlib import Path
import pytest
from fastapi.testclient import TestClient
from ai_service.config import Settings
from ai_service.grading.store import SQLiteGradeStore
from ai_service.llm.mock import MockProvider
from ai_service.main import create_app
from ai_service.telemetry.ingest import TraceIntegrityMap
from ai_service.telemetry.models import TelemetryEvent
from ai_service.telemetry.store import SQLiteTraceStore
T0 = datetime(2026, 9, 12, tzinfo=UTC)
def _event(seq: int, kind: str, payload: dict, offset_s: float):
return TelemetryEvent(
learner_id="lab-learner",
task_id="lab-task",
seq=seq,
kind=kind,
payload=payload,
ts=T0 + timedelta(seconds=offset_s),
sandbox_id="sbx-lab",
)
class StreamingMock(MockProvider):
"""Deterministic token stream for the SSE path."""
def _reply_for(self, messages, response_format):
return "Feedback grounded in your live session digest."
@pytest.fixture()
def client(tmp_path: Path) -> TestClient:
app = create_app(Settings(provider="mock"))
app.state.provider = StreamingMock()
app.state.trace_store = SQLiteTraceStore(db_path=tmp_path / "t.db")
app.state.grade_store = SQLiteGradeStore(db_path=tmp_path / "g.db")
app.state.trace_integrity = TraceIntegrityMap()
with TestClient(app) as c:
yield c
def test_live_trace_streams_full_envelope(client: TestClient) -> None:
store: SQLiteTraceStore = client.app.state.trace_store
for e in [
_event(0, "file_diff", {"path": "x.py"}, 0),
_event(1, "command", {"cmd": "pytest -q"}, 10),
_event(2, "test_result", {"passed": False, "exit_code": 1}, 20),
_event(3, "test_result", {"passed": True, "exit_code": 0}, 40),
]:
store.append(e)
with client.stream(
"POST", "/v1/lab/feedback", json={"learner_id": "lab-learner", "task_id": "lab-task"}
) as resp:
assert resp.status_code == 200
body = "".join(chunk.decode() for chunk in resp.iter_raw())
assert '"agent": "lab"' in body or '"agent":"lab"' in body
assert '"task_id": "lab-task"' in body
assert '"type": "delta"' in body or '"type":"delta"' in body
def test_empty_trace_coaches_the_baseline(client: TestClient) -> None:
"""No telemetry is NOT an error — Lab coaches 'run the starter test'."""
with client.stream(
"POST", "/v1/lab/feedback", json={"learner_id": "lab-learner", "task_id": "no-events"}
) as resp:
assert resp.status_code == 200
def test_missing_fields_422(client: TestClient) -> None:
resp = client.post("/v1/lab/feedback", json={"learner_id": "x"})
assert resp.status_code == 422
-56
View File
@@ -1,56 +0,0 @@
"""Mentor narrative endpoint tests — SSE, session-backed (REQ-2-010)."""
import json
def stream_events(client, payload) -> list[dict]:
with client.stream("POST", "/v1/mentor/narrative", json=payload) as response:
assert response.status_code == 200
events = []
for line in response.iter_lines():
if line.startswith("data:"):
d = line.removeprefix("data:").strip()
if d == "[DONE]":
events.append({"type": "[DONE]"})
else:
events.append(json.loads(d))
return events
def test_narrative_streams_full_envelope(client):
events = stream_events(client, {"session_id": "mentor-1"})
assert events[0]["type"] == "meta"
assert events[0]["agent"] == "mentor"
deltas = [e for e in events if e["type"] == "delta"]
assert len(deltas) >= 1
assert any(e["type"] == "done" for e in events)
assert events[-1]["type"] == "[DONE]"
def test_narrative_is_session_backed(client):
"""Second call replays history: provider input grows; distinct mock output."""
first = stream_events(client, {"session_id": "mentor-2", "prompt": "narrate my path"})
second = stream_events(client, {"session_id": "mentor-2", "prompt": "what next?"})
first_text = "".join(e["content"] for e in first if e["type"] == "delta")
second_text = "".join(e["content"] for e in second if e["type"] == "delta")
assert first_text != second_text
def test_narrative_persists_turns(client):
import asyncio
store = client.app.state.session_store
stream_events(client, {"session_id": "mentor-3", "prompt": "hello trajectory"})
async def check():
return await store.history_window("mentor-3")
contents = [m.content for m in asyncio.run(check())]
assert "hello trajectory" in contents
assert len(contents) >= 2 # user + assistant persisted
def test_missing_session_id_422(client):
response = client.post("/v1/mentor/narrative", json={"prompt": "hi"})
assert response.status_code == 422
-121
View File
@@ -1,121 +0,0 @@
"""Proctor signals endpoint tests — REAL inputs contract (REQ-3-007).
v0.3 re-grounding: POST /v1/proctor/signals takes {learner_id, task_id}
and gathers digest + defense signals + variant context server-side.
"""
from __future__ import annotations
import json
from datetime import UTC, datetime, timedelta
from pathlib import Path
import pytest
from fastapi.testclient import TestClient
from ai_service.config import Settings
from ai_service.grading.store import SQLiteGradeStore
from ai_service.llm.mock import MockProvider
from ai_service.main import create_app
from ai_service.telemetry.ingest import TraceIntegrityMap
from ai_service.telemetry.models import TelemetryEvent
from ai_service.telemetry.store import SQLiteTraceStore
from ai_service.variants.store import SQLiteVariantStore
from ai_service.voice.defense_store import SQLiteDefenseStore
VALID = {
"signals": [
{"signal_type": "idle_gap", "severity": "low", "note": "One long pause."},
],
"intervention": "Offer a short break, then restate the plan",
"summary": "Coaching-shaped session note",
}
T0 = datetime(2026, 9, 12, tzinfo=UTC)
class ProctorJSON(MockProvider):
def _reply_for(self, messages, response_format):
if response_format is not None and response_format.get("type") == "json_object":
return json.dumps(VALID)
return super()._reply_for(messages, response_format)
class BrokenJSON(MockProvider):
def _reply_for(self, messages, response_format):
if response_format is not None and response_format.get("type") == "json_object":
return "not json ever"
return super()._reply_for(messages, response_format)
@pytest.fixture()
def client(tmp_path: Path) -> TestClient:
app = create_app(Settings(provider="mock"))
app.state.trace_store = SQLiteTraceStore(db_path=tmp_path / "t.db")
app.state.grade_store = SQLiteGradeStore(db_path=tmp_path / "g.db")
app.state.variant_store = SQLiteVariantStore(db_path=tmp_path / "v.db")
app.state.defense_store = SQLiteDefenseStore(db_path=tmp_path / "d.db")
app.state.trace_integrity = TraceIntegrityMap()
with TestClient(app) as c:
yield c
def _seed_trace(client: TestClient) -> None:
store: SQLiteTraceStore = client.app.state.trace_store
for e in [
TelemetryEvent(
learner_id="p-learner",
task_id="p-task",
seq=0,
kind="activity",
payload={"state": "idle"},
ts=T0,
sandbox_id="sbx-p",
),
TelemetryEvent(
learner_id="p-learner",
task_id="p-task",
seq=1,
kind="activity",
payload={"state": "idle"},
ts=T0 + timedelta(seconds=400),
sandbox_id="sbx-p",
),
]:
store.append(e)
def test_signals_returns_validated_json(client: TestClient) -> None:
client.app.state.provider = ProctorJSON()
_seed_trace(client)
resp = client.post(
"/v1/proctor/signals", json={"learner_id": "p-learner", "task_id": "p-task"}
)
assert resp.status_code == 200, resp.text
body = resp.json()
assert body["intervention"]
assert body["signals"][0]["signal_type"] == "idle_gap"
def test_empty_trace_is_valid_not_404(client: TestClient) -> None:
"""No telemetry → the proctor still assesses (nothing to flag)."""
client.app.state.provider = ProctorJSON()
resp = client.post(
"/v1/proctor/signals", json={"learner_id": "nobody", "task_id": "nothing"}
)
assert resp.status_code == 200
def test_unparseable_provider_502(client: TestClient) -> None:
client.app.state.provider = BrokenJSON()
_seed_trace(client)
resp = client.post(
"/v1/proctor/signals", json={"learner_id": "p-learner", "task_id": "p-task"}
)
assert resp.status_code == 502
def test_missing_fields_422(client: TestClient) -> None:
client.app.state.provider = ProctorJSON()
resp = client.post("/v1/proctor/signals", json={"learner_id": "x"})
assert resp.status_code == 422
-449
View File
@@ -1,449 +0,0 @@
"""/v1/sandboxes API tests — lifecycle, pool guard, G-5 abuse control (REQ-3-001).
Every test here runs against `StubBackend` (workdir layout without namespace
spawning) EXCEPT the final probe-guarded real-backend path; Wave 1/2 already
covers boundary isolation, so the API layer only needs the manager contract.
DI/lifespan wiring note: tests build the manager themselves, pre-set
`app.state.sandbox_manager` BEFORE the TestClient lifespan runs, and the
lifespan adopts that instance (state-injection override) instead of
constructing the real UnshareBackend one. The lifespan still `start()`s it,
runs the periodic reaper, and `destroy_all()`s on shutdown the same
production path, just with a stub at the port.
Snapshot/destroy content plumbing for stub tests uses the explicit test seam
`app.state.sandbox_test_layout` `X-Snapshot-Copy` / `X-Workspace-Copy`
response headers (production never sets them).
"""
from __future__ import annotations
import shutil
import time
from collections.abc import Iterator
from datetime import UTC, datetime
from pathlib import Path
import pytest
from fastapi.testclient import TestClient
from ai_service.config import Settings
from ai_service.main import create_app
from ai_service.sandbox import SandboxHandle, SandboxManager, UnshareBackend
from ai_service.sandbox.backend import ExecResult, SandboxSpec
from ai_service.sandbox.workdir import SandboxDir, create_layout
from ai_service.sandbox.workdir import snapshot as workdir_snapshot
from tests.sandbox.test_isolation import requires_userns
PILOT = "pilot-learner"
OTHER = "pilot-learner-2"
BASE_SETTINGS: dict = {
"provider": "mock",
"sandbox_max_concurrent": 5,
"sandbox_max_per_learner": 1,
"sandbox_creates_per_min": 10,
}
class StubBackend:
"""Structural SandboxBackend: lays out the workdir, spawns nothing.
Tracks spawn/destroy calls so lifespan shutdown behaviour is assertable.
"""
def __init__(self) -> None:
self.spawned_specs: list[SandboxSpec] = []
self.destroyed_ids: list[str] = []
async def spawn(self, spec: SandboxSpec) -> SandboxHandle:
create_layout(spec)
self.spawned_specs.append(spec)
return SandboxHandle(
id=spec.sandbox_id, pid=None, workdir=spec.workdir,
created_at=datetime.now(UTC),
)
async def exec(self, handle: SandboxHandle, cmd: list[str]) -> ExecResult:
raise NotImplementedError("API tests never exec")
async def snapshot(self, handle: SandboxHandle) -> Path:
return workdir_snapshot(handle.workdir)
async def destroy(self, handle: SandboxHandle) -> None:
self.destroyed_ids.append(handle.id)
handle.pid = None
def _seed_workspace(layout: SandboxDir, name: str, content: str) -> None:
layout.workspace.mkdir(parents=True, exist_ok=True)
(layout.workspace / name).write_text(content)
@pytest.fixture(autouse=True)
def _clear_create_window() -> Iterator[None]:
"""Module-global rate window must start empty in every test."""
from ai_service.api import sandboxes as sandboxes_api
sandboxes_api._CREATE_TIMES.clear()
yield
sandboxes_api._CREATE_TIMES.clear()
@pytest.fixture()
def stub_backend() -> StubBackend:
return StubBackend()
@pytest.fixture()
def client(
tmp_path: Path, stub_backend: StubBackend, monkeypatch: pytest.MonkeyPatch
) -> Iterator[TestClient]:
"""App + adopted stub manager, sandbox dir rooted in a per-test tmp_path."""
monkeypatch.setenv("AI_SANDBOX_CREATES_PER_MIN", "150") # shared-window headroom
settings = Settings(**{**BASE_SETTINGS, "sandbox_dir": tmp_path / "sandboxes"})
app = create_app(settings)
app.state.sandbox_manager = SandboxManager(backend=stub_backend, settings=settings)
with TestClient(app) as c:
yield c
# -- lifecycle roundtrip --------------------------------------------------------
def test_create_list_get_snapshot_delete_roundtrip(client: TestClient) -> None:
created = client.post("/v1/sandboxes", json={"learner_id": PILOT})
assert created.status_code == 201
body = created.json()
assert body["id"].startswith("sbx-")
assert body["learner_id"] == PILOT
assert body["created_at"]
assert "pid" not in body # host-process detail, excluded from the contract
sandbox_id = body["id"]
listed = client.get("/v1/sandboxes")
assert listed.status_code == 200
rows = listed.json()["sandboxes"]
assert [r["id"] for r in rows] == [sandbox_id]
assert rows[0]["learner_id"] == PILOT
got = client.get(f"/v1/sandboxes/{sandbox_id}")
assert got.status_code == 200
assert got.json()["id"] == sandbox_id
# Seed workspace content on the host side, then snapshot it away.
layout = SandboxDir(
root=Path(body["workdir"]),
workspace=Path(body["workdir"]) / "workspace",
snapshots=Path(body["workdir"]) / "snapshots",
)
_seed_workspace(layout, "solution.py", "print(42)\n")
snap = client.post(f"/v1/sandboxes/{sandbox_id}/snapshot")
assert snap.status_code == 200
snap_body = snap.json()
assert snap_body["sandbox_id"] == sandbox_id
assert snap_body["files"] == ["solution.py"]
snapshot_path = Path(snap_body["snapshot_path"])
assert (snapshot_path / "solution.py").read_text() == "print(42)\n"
# Destroy: 204, registry empties, but the workdir (and snapshot) survives.
deleted = client.delete(f"/v1/sandboxes/{sandbox_id}")
assert deleted.status_code == 204
assert deleted.content == b""
assert client.get("/v1/sandboxes").json()["sandboxes"] == []
assert client.get(f"/v1/sandboxes/{sandbox_id}").status_code == 404
assert layout.root.is_dir() # kept for restore; not purged (API semantic)
assert (snapshot_path / "solution.py").is_file()
def test_create_validates_body_422(client: TestClient) -> None:
assert client.post("/v1/sandboxes", json={}).status_code == 422
assert client.post("/v1/sandboxes", json={"learner_id": ""}).status_code == 422
def test_snapshot_unknown_id_404(client: TestClient) -> None:
response = client.post("/v1/sandboxes/sbx-nope/snapshot")
assert response.status_code == 404
assert "sbx-nope" in response.json()["detail"]
def test_delete_unknown_id_404(client: TestClient) -> None:
response = client.delete("/v1/sandboxes/sbx-nope")
assert response.status_code == 404
assert "sbx-nope" in response.json()["detail"]
def test_get_unknown_id_404(client: TestClient) -> None:
assert client.get("/v1/sandboxes/sbx-nope").status_code == 404
# -- D-032 pool guard → 503 ------------------------------------------------------
def test_pool_full_returns_503(tmp_path: Path, stub_backend: StubBackend) -> None:
settings = Settings(
provider="mock",
sandbox_dir=tmp_path / "sandboxes",
sandbox_max_concurrent=2,
sandbox_max_per_learner=5, # cap disabled for this test
sandbox_creates_per_min=150,
)
app = create_app(settings)
app.state.sandbox_manager = SandboxManager(backend=stub_backend, settings=settings)
with TestClient(app) as client:
ids = [
client.post("/v1/sandboxes", json={"learner_id": PILOT}).json()["id"]
for _ in range(2)
]
overflow = client.post("/v1/sandboxes", json={"learner_id": PILOT})
assert overflow.status_code == 503
assert "pool full" in overflow.json()["detail"]
# Capacity frees on destroy: the next create succeeds.
client.delete(f"/v1/sandboxes/{ids[0]}")
refilled = client.post("/v1/sandboxes", json={"learner_id": PILOT})
assert refilled.status_code == 201
# -- G-5 abuse control ------------------------------------------------------------
def test_non_allowlisted_learner_403(client: TestClient) -> None:
response = client.post("/v1/sandboxes", json={"learner_id": "intruder-7"})
assert response.status_code == 403
assert "allowlist" in response.json()["detail"]
assert client.get("/v1/sandboxes").json()["sandboxes"] == [] # nothing spawned
def test_second_active_sandbox_for_same_learner_429(
tmp_path: Path, stub_backend: StubBackend
) -> None:
settings = Settings(
provider="mock",
sandbox_dir=tmp_path / "sandboxes",
learner_allowlist=[PILOT, OTHER],
sandbox_max_concurrent=10,
sandbox_max_per_learner=1,
sandbox_creates_per_min=150,
)
app = create_app(settings)
app.state.sandbox_manager = SandboxManager(backend=stub_backend, settings=settings)
with TestClient(app) as client:
first = client.post("/v1/sandboxes", json={"learner_id": PILOT})
assert first.status_code == 201
second = client.post("/v1/sandboxes", json={"learner_id": PILOT})
assert second.status_code == 429
assert "per-learner cap" in second.json()["detail"]
# Cap counts ACTIVE sandboxes: after destroy the learner can create again.
client.delete(f"/v1/sandboxes/{first.json()['id']}")
again = client.post("/v1/sandboxes", json={"learner_id": PILOT})
assert again.status_code == 201
# A different allowlisted learner was never blocked by the pilot's sandbox.
other = client.post("/v1/sandboxes", json={"learner_id": OTHER})
assert other.status_code == 201
def test_burst_over_global_create_rate_429(tmp_path: Path, stub_backend: StubBackend) -> None:
settings = Settings(
provider="mock",
sandbox_dir=tmp_path / "sandboxes",
sandbox_max_concurrent=10,
sandbox_max_per_learner=10, # per-learner cap disabled for this test
sandbox_creates_per_min=3,
)
app = create_app(settings)
manager = SandboxManager(backend=stub_backend, settings=settings)
app.state.sandbox_manager = manager
with TestClient(app) as client:
codes = [
client.post("/v1/sandboxes", json={"learner_id": PILOT}).status_code
for _ in range(3)
]
assert codes == [201, 201, 201]
burst = client.post("/v1/sandboxes", json={"learner_id": PILOT})
assert burst.status_code == 429
assert "create rate" in burst.json()["detail"]
assert manager.active_count == 3 # the 429 spawned nothing
# Rejected creates never consume budget; still full one moment later.
assert client.post("/v1/sandboxes", json={"learner_id": PILOT}).status_code == 429
def test_create_rate_window_slides(monkeypatch: pytest.MonkeyPatch) -> None:
"""The window is monotonic-time based; fake the clock, skip real sleeping."""
from fastapi import HTTPException
from ai_service.api import sandboxes as sandboxes_api
sandboxes_api._CREATE_TIMES.clear()
fake_now = 1_000.0
monkeypatch.setattr(time, "monotonic", lambda: fake_now)
settings = Settings(provider="mock", sandbox_creates_per_min=2)
sandboxes_api._check_global_create_rate(settings)
sandboxes_api._check_global_create_rate(settings)
with pytest.raises(HTTPException) as excinfo:
sandboxes_api._check_global_create_rate(settings)
assert excinfo.value.status_code == 429
fake_now += 61.0 # window slides: oldest entries age out
sandboxes_api._check_global_create_rate(settings) # admitted again
# -- lifespan wiring --------------------------------------------------------------
def test_lifespan_start_and_shutdown_destroy(
tmp_path: Path, stub_backend: StubBackend
) -> None:
"""a-1 at boot: an orphan workdir with a dead recorded pid is reaped;
at shutdown every live sandbox is destroyed (no orphans outlive us)."""
orphan_root = tmp_path / "sandboxes" / "sbx-orphaned"
(orphan_root / "workspace").mkdir(parents=True)
(orphan_root / "sandbox.json").write_text(
'{"sandbox_id": "sbx-orphaned", "learner_id": "?", "pid": 4194304}'
)
settings = Settings(
provider="mock",
sandbox_dir=tmp_path / "sandboxes",
sandbox_creates_per_min=150,
)
manager = SandboxManager(backend=stub_backend, settings=settings)
app = create_app(settings)
app.state.sandbox_manager = manager
with TestClient(app) as client:
assert not orphan_root.exists() # startup reaper ran during lifespan boot
assert any(e.kind == "orphan_reaped" for e in manager.integrity_events)
created = client.post("/v1/sandboxes", json={"learner_id": PILOT})
sandbox_id = created.json()["id"]
assert manager.active_count == 1
# Exiting the TestClient runs the shutdown half of the lifespan.
assert stub_backend.destroyed_ids == [sandbox_id]
assert manager.active_count == 0
def test_lifespan_constructs_real_manager_when_not_overridden(tmp_path: Path) -> None:
"""No override → the lifespan builds the production UnshareBackend manager."""
settings = Settings(provider="mock", sandbox_dir=tmp_path / "sandboxes")
app = create_app(settings)
with TestClient(app):
manager = app.state.sandbox_manager
assert isinstance(manager, SandboxManager)
# -- real-backend path (probe-guarded; Wave 1/2 owns deep isolation) --------------
#: Repo-anchored, gitignored scratch root (tests/sandboxes/ like isolation
#: tests) — /tmp is off-limits: the in-namespace tmpfs shadows host /tmp.
REAL_SANDBOXES_ROOT = Path(__file__).resolve().parents[1] / "sandboxes"
@requires_userns
def test_real_backend_create_path_runs(tmp_path: Path) -> None:
"""One cheap real-namespace pass through the API (create only, no exec)."""
sandbox_root = REAL_SANDBOXES_ROOT / f"api-{tmp_path.name}"
settings = Settings(
provider="mock",
sandbox_dir=sandbox_root,
sandbox_creates_per_min=150,
)
app = create_app(settings) # no override → lifespan wires UnshareBackend
try:
with TestClient(app) as client:
created = client.post("/v1/sandboxes", json={"learner_id": PILOT})
assert created.status_code == 201
sandbox_id = created.json()["id"]
assert (sandbox_root / sandbox_id / "workspace").is_dir()
assert isinstance(app.state.sandbox_manager._backend, UnshareBackend)
# Shutdown destroyed the live sandbox; the kept workdir still exists.
assert app.state.sandbox_manager.active_count == 0
assert (sandbox_root / sandbox_id / "workspace").is_dir()
finally:
shutil.rmtree(sandbox_root, ignore_errors=True)
class TestFilesAndExecRoutes:
"""Workspace CRUD + Run/Test exec (Phase 6, REQ-3-008, CUT-2)."""
def test_file_write_read_list_roundtrip(self, client):
handle = client.post(
"/v1/sandboxes", json={"learner_id": "pilot-learner"}
).json()
sbx = handle["id"]
put = client.put(
f"/v1/sandboxes/{sbx}/files/main.py",
json={"path": "main.py", "content": "print('hi')"},
)
assert put.status_code == 200, put.text
got = client.get(f"/v1/sandboxes/{sbx}/files/main.py")
assert got.status_code == 200
assert "print('hi')" in got.json()["content"]
listed = client.get(f"/v1/sandboxes/{sbx}/files")
assert "main.py" in listed.json()["files"]
client.delete(f"/v1/sandboxes/{sbx}")
def test_traversal_rejected(self, client):
handle = client.post(
"/v1/sandboxes", json={"learner_id": "pilot-learner"}
).json()
sbx = handle["id"]
bad = client.put(
f"/v1/sandboxes/{sbx}/files/..%2Fescape.txt",
json={"path": "../escape.txt", "content": "x"},
)
assert bad.status_code == 422
client.delete(f"/v1/sandboxes/{sbx}")
def test_unknown_file_404(self, client):
handle = client.post(
"/v1/sandboxes", json={"learner_id": "pilot-learner"}
).json()
sbx = handle["id"]
assert client.get(f"/v1/sandboxes/{sbx}/files/ghost.py").status_code == 404
client.delete(f"/v1/sandboxes/{sbx}")
def test_exec_unknown_sandbox_404(self, client):
resp = client.post(
"/v1/sandboxes/sbx-nope/exec", json={"cmd": ["echo", "hi"]}
)
assert resp.status_code == 404
def test_unknown_sandbox_file_routes_404_not_500(self, client):
"""P7: read/write on an unknown sandbox must 404 (SandboxNotFoundError
previously escaped _workspace_dir as an unhandled 500)."""
assert (
client.get("/v1/sandboxes/sbx-nope/files/whatever.py").status_code == 404
)
put = client.put(
"/v1/sandboxes/sbx-nope/files/whatever.py",
json={"path": "whatever.py", "content": "x"},
)
assert put.status_code == 404
def test_symlink_escape_rejected(self, client):
"""P7: an exec-planted symlink in the workspace must not let the
file routes read/write OUTSIDE the bind (lexical traversal checks
cannot see symlinks resolve + containment re-check is the gate)."""
handle = client.post(
"/v1/sandboxes", json={"learner_id": "pilot-learner"}
).json()
sbx = handle["id"]
workspace = Path(handle["workdir"]) / "workspace"
outside = workspace.parent / "secret.txt"
outside.write_text("host secret") # a host file OUTSIDE the bind
try:
(workspace / "leak.txt").symlink_to(outside)
read = client.get(f"/v1/sandboxes/{sbx}/files/leak.txt")
assert read.status_code == 422, (
f"symlink escape read must 422, got {read.status_code}: {read.text}"
)
write = client.put(
f"/v1/sandboxes/{sbx}/files/leak.txt",
json={"path": "leak.txt", "content": "pwned"},
)
assert write.status_code == 422, (
f"symlink escape write must 422, got {write.status_code}: {write.text}"
)
assert outside.read_text() == "host secret" # untouched
finally:
client.delete(f"/v1/sandboxes/{sbx}")
outside.unlink(missing_ok=True)
@@ -1,93 +0,0 @@
"""Client-disconnect regression tests — SSE generators must tolerate aclose().
A `yield` inside `finally` re-raises "async generator ignored GeneratorExit"
when sse-starlette closes the iterator on client disconnect (P0 finding,
final review). These tests reproduce the close path directly against each
endpoint's event_stream generator shape.
"""
import json
from ai_service.corpus.learner_context import get_learner_context
from ai_service.llm.types import Message
async def _chat_event_stream(client, session_id="close-chat", content="hi"):
"""Rebuild the chat endpoint's event_stream generator exactly as
chat.py builds it (same code shape, same session flow)."""
app = client.app
settings = app.state.settings
provider = app.state.provider
registry = app.state.agent_registry
sessions = app.state.session_store
agent = registry.get(provider, settings, "tutor")
learner_context = get_learner_context(None)
if await sessions.get(session_id) is None:
await sessions.create(session_id, agent="tutor", learner_id="learner-001")
user_turn = Message(role="user", content=content)
history = await sessions.history_window(session_id)
await sessions.append(session_id, user_turn)
async def event_stream():
yield {"event": "message", "data": json.dumps({
"type": "meta", "agent": "tutor", "session_id": session_id,
"model": settings.model,
})}
first_byte = True
reply_parts: list[str] = []
try:
async for token in agent.stream_reply(
history=history, user_input=user_turn.content,
learner_context=learner_context,
):
first_byte = False
reply_parts.append(token)
yield {"event": "message", "data": json.dumps({
"type": "delta", "content": token
})}
full_reply = "".join(reply_parts)
if full_reply:
await sessions.append(
session_id, Message(role="assistant", content=full_reply)
)
yield {"event": "message", "data": json.dumps({
"type": "done", "finish_reason": "stop"
})}
yield {"event": "message", "data": "[DONE]"}
except Exception as exc:
code = "provider_unavailable" if first_byte else "provider_error"
yield {"event": "message", "data": json.dumps({
"type": "error", "code": code, "message": str(exc)
})}
yield {"event": "message", "data": "[DONE]"}
return event_stream()
async def test_chat_stream_generator_survives_aclose(client):
"""Partial consumption then aclose() must not raise
'async generator ignored GeneratorExit' (yield-in-finally regression)."""
gen = await _chat_event_stream(client, session_id="close-chat")
meta = await gen.__anext__()
assert json.loads(meta["data"])["type"] == "meta"
delta = await gen.__anext__()
assert json.loads(delta["data"])["type"] == "delta"
# The critical assertion: closing mid-stream must be clean (no raise).
await gen.aclose()
async def test_chat_stream_survives_close_at_different_points(client):
"""Close right after meta, and right after done — all must be clean."""
gen = await _chat_event_stream(client, session_id="close-early")
await gen.__anext__() # meta only
await gen.aclose()
gen2 = await _chat_event_stream(client, session_id="close-late")
events = []
async for ev in gen2:
events.append(json.loads(ev["data"]))
if len(events) == 2:
break
await gen2.aclose()
assert events[0]["type"] == "meta"

Some files were not shown because too many files have changed in this diff Show More