Compare commits
1 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| e58b027a57 |
+25
-145
@@ -4,19 +4,7 @@
|
||||
|
||||
Nextcraft is a TypeScript monorepo (pnpm workspaces + turborepo) with a Next.js application hosting four surfaces (Learner, Marketplace, Employer Dashboard, Admin), a shared component library, typed mock data layer, and shared types package. As of v0.2, a Python FastAPI application (`apps/ai-service`) hosts six AI tutor agents backed by a provider-agnostic LLM layer.
|
||||
|
||||
**v0.3 additions (Credential Engines):** real credential engines replace v0.2 mock inputs — a sandbox fabric (isolated per-learner coding environments via Linux user/mount/pid/net namespaces), a live build-telemetry pipeline (WebSocket ingest + SQLite-ordered event log), a process-trace grading engine, seeded per-learner variant task generation, and a voice-based oral defense (STT/TTS via a new provider-agnostic voice layer). **First real persistence introduced: SQLite** (`ai_service/telemetry/`, grading, variant, defense stores). Lab/Assessor/Proctor agents are re-grounded onto real telemetry/traces. **Identity/age-gating (KYC) deferred per founder directive** — no security engineer persona; secrets-hygiene checklist only.
|
||||
|
||||
**v0.4 additions (Distribution & Bootstrap CLI, founder directive D-016):** a new `apps/cli` package — the `nextcraft` bootstrap CLI (`doctor`/`bootstrap`/`verify`/`dev`) compiled to a self-contained linux x64 binary via **Node SEA** (probe-verified: Go/Rust absent, node v24.15.0 SEA-capable), installed by a repo-served one-liner script that resolves the latest Gitea release, downloads binary + sha256 sidecar, verifies, and installs to `~/.local/bin`. Every release from v0.4 onward attaches the binary + checksum as release assets (the "ongoing binaries" requirement). The CLI is a thin wrapper: all orchestration logic stays in `apps/ai-service/scripts/` (bootstrap.sh/dev.sh) — the CLI composes them via subprocess (A-202), duplicating nothing. Previously-planned v0.4 seams (real STT/TTS, KYC, design/sim envs, seq-lease) move to v0.5.
|
||||
|
||||
### v0.5 Research Conclusions (Phase 0 RESEARCH)
|
||||
|
||||
25. **D-040 Voice real path = `OpenAIAudioProvider` on the shared httpx pool (D-R01)** — implements the D-030 protocol: `transcribe()` = multipart `POST {AI_VOICE_BASE_URL}/audio/transcriptions` (`file` + `model`, response_format=json → `TranscriptSegment`), `synthesize()` = streaming `POST /audio/speech` (JSON body, raw byte stream, `input` ≤4096 chars). Carries `descriptor = VoiceDescriptor(mode="server", ...)` which defense.py:143 already prefers — mode selection drops in with zero API changes. Constructor takes the lifespan `httpx.AsyncClient` (D-017; read=300s already tolerates multi-minute clips). Factory signature becomes `voice_provider_from_settings(settings, http_client)`. Errors sanitized with key redaction (mirror `llm/openai_compat.py:_sanitize`).
|
||||
26. **D-041 Voice route fixes (D-R02)** — defense.py:189 fmt derivation must strip codec params (`"webm;codecs=opus"` → `webm`, else real STT 400s); ~10 MB audio guard (422/413) before provider call; TTS route media_type becomes format-aware (or force `response_format=wav`). Web: defense-session POSTs the recorded blob (FormData), engine-client gains the audio-answer variant.
|
||||
27. **D-042 Identity = 5th D-027 store + provider protocol (D-R03)** — `ai_service/identity/` (D-031): `IdentityProvider` protocol (submit/poll/verify), deterministic mock, SQLite store modeled on DefenseStore (WAL, FK on, portable columns). Stores **derived age_band** (`16-17`/`18+`, never raw DOB) + document **refs** (never raw docs); mock verdicts carry a `mock` marker so downstream never displays them as production-verified (A-304). PII hygiene pinned by caplog sentinel test. **Negative finding:** no age-gate UI or enrollment API exists today — the flow is built, not swapped (D-010's "visual flow" is vestigial: one mock field).
|
||||
28. **D-043 Gate composition (D-R04)** — api/-layer dependencies in order: G-5 allowlist (403 pilot guard, retained) → identity verdict (403 + verify-CTA payload) → rate caps (429). School 16+ gates variant generation + sandbox create + defense start; marketplace 18+ via reusable `require_verified_adult` demonstrated on one minimal gated route (marketplace has no backend today — the thin route proves the contract end-to-end). All middleware-layer, never in the manager.
|
||||
29. **D-044 Environment registry at the template/variant layer (D-R05)** — `TaskTemplate.environment: Literal["build","design","simulation"]` → carried through VariantRecord → API → TS `TaskVariant`; surfaces the existing-but-dead `test_command` field (Run/Test buttons stop hardcoding pytest). Manager/backend/workdir unchanged (D-024 untouched — A-307: an environment is a typing over starter files + command policy). Per-kind exec command policy at the api/ exec route (template-declared harness + generic file/nav commands, 422 on violation). Grading digest stays kind-agnostic by construction (features derive from event kinds — pinned with a digest test over a synthetic design-kind trace). Starter contents: design = SVG/HTML/schematic artifacts + validate/render harness (stack-designer c001/c002 already sanctioned); simulation = parameterized benchmark script + dataset files (stack-science, stack-operator).
|
||||
30. **D-045 Seq-ack protocol (D-R06/D-R07)** — new WS frame `{"type":"seq_ack","seq":N}` emitted per successful append from `IngestSession._append` (durable `latest_seq`; O(1); advisory — gap detection stays authoritative, G-3/G-4 unchanged). The capture agent's supervisor loop (which today discards all non-close frames) parses text frames and trims spool+pending to `seq > ack` under `_emit_lock` via atomic `Spool.rewrite` — closing the one-line replay-margin gap (`_flush_locked` pops N in-flight frames but `replay_margin()` requeues only `_last_sent`). An explicit spool bound is added (A-309's "spool cap exists" premise was factually wrong on disk). Regression test: mid-burst kill in the test_durability.py real-server harness (uvicorn + KillableProxy) — the exact scenario the P07 de-flake documented as uncovered.
|
||||
31. **D-046 v0.5 wave order (D-R08)** — P1 seq-lease (smallest, protocol-only, fixes transport before env phases add reconnecting producers), P2 voice, P3 identity, P4 environments, P5 final. Phase numbers renumbered accordingly (was 1=voice..4=seq-lease in the initial ROADMAP draft).
|
||||
**v0.2 additions:** real AI services (streaming chat), agent framework, mock engine inputs for Lab/Assessor/Proctor. **Still no database, no auth** — in-memory session store; real engines (sandbox fabric, assessment engine, identity) are v0.3+.
|
||||
|
||||
### Confirmed Technology Stack (v0.2)
|
||||
|
||||
@@ -38,8 +26,6 @@ Nextcraft is a TypeScript monorepo (pnpm workspaces + turborepo) with a Next.js
|
||||
| pydantic | 2.13.x | Request/response models, structured outputs |
|
||||
| pydantic-settings | 2.15.x | Settings + env-file loading (replaces python-dotenv) |
|
||||
| httpx | 0.28.x | Async LLM HTTP client (ollama-cloud + local providers) |
|
||||
| sqlmodel / sqlalchemy | 0.0.24 / 2.x | Typed SQLite persistence for the v0.3 engine stores (D-027) |
|
||||
| python-multipart | 0.0.x | Multipart audio upload for the defense answer route (REQ-3-006) |
|
||||
| sse-starlette | 3.4.x | SSE framing, ping keep-alive |
|
||||
| pytest | 9.x | Test runner |
|
||||
| pytest-asyncio | 1.4.x | Async tests (auto mode) |
|
||||
@@ -59,28 +45,6 @@ Deliberately **not** used: openai-python SDK (the `LLMProvider` protocol is the
|
||||
7. **D-022 Monorepo integration** — zero-dependency shim `package.json` in apps/ai-service + `ai#*` turbo passthrough tasks (`cache:false, outputs:[]`) + root `ai:dev`/`ai:test` scripts + idempotent venv bootstrap.
|
||||
8. **D-023 Testing** — pytest-asyncio auto mode; TestClient `client.stream()` for SSE; httpx MockTransport for byte-exact provider parser tests; scripted mock provider incl. failure modes. Tests never call the cloud.
|
||||
|
||||
### v0.4 Architecture Decisions (from Research — Distribution & Bootstrap CLI)
|
||||
|
||||
18. **D-033 Binary toolchain = Node SEA (probe-verified)** — Go and Rust are absent from this box; node v24.15.0 ships SEA support (`--experimental-sea-config`, postject-free on linux via `cp node nextcraft && node sea-config` … blob injection with the system `dd`/`npx postject` if needed). CLI source lives in `apps/cli` (TypeScript, compiled to a single CJS bundle by esbuild, then SEA-injected into a copy of the node binary → `nextcraft-linux-x64`). Fallback if SEA breaks: python3 `zipapp` (3.11.2 available). No new toolchain deps beyond dev-scoped esbuild.
|
||||
19. **D-034 CLI = thin wrapper, orchestration stays in scripts/** — `nextcraft` composes `apps/ai-service/scripts/bootstrap.sh` and `scripts/dev.sh` equivalents via `spawn` with inherited stdio and timeout guards (A-202/A-209). doctor/bootstrap/verify implement only *checking* logic (prereqs, env template, health) — never re-implement installs. This keeps one source of truth for bootstrap semantics.
|
||||
20. **D-035 Install path = repo raw `install.sh` + Gitea latest-release API** — the one-liner `curl -fsSL <forge>/coreci/nextcraft/raw/main/scripts/install.sh | bash` resolves `GET /api/v1/repos/coreci/nextcraft/releases/latest`, downloads the `nextcraft-linux-x64` + `nextcraft-linux-x64.sha256` assets, verifies sha256 (`shasum -a 256`), installs to `~/.local/bin` (PATH hint), and degrades to printed source-bootstrap instructions when no binary asset exists or the platform mismatches (A-203/A-204/A-206).
|
||||
21. **D-036 Ongoing binaries = ship-workflow asset step** — the release pipeline (v0.3's `ShipWorkflow.createRelease` equivalent, executed as the ship step's asset stage) builds the binary + checksum and attaches both to every Gitea release from v0.4 onward (A-205). Token resolution stays `.env*`-only (D-006/D-014); binaries are linux x64 only for v0.4 (macOS arm64 deferred — unverifiable on this box).
|
||||
22. **D-037 CLI package layout** — `apps/cli` is a pnpm workspace package (`@nextcraft/cli`): `src/` (entry, commands/, checks/, lib/), `scripts/build-binary.mjs` (esbuild bundle → SEA inject), unit tests runnable via `pnpm --filter @nextcraft/cli test` (node:test, no new test framework). Root `package.json` gains `cli:*` passthrough scripts mirroring the `ai:*` pattern (D-022).
|
||||
23. **D-038 Network mode (v0.3.5)** — dev binds 0.0.0.0 (`AI_HOST`, default 0.0.0.0, revert via 127.0.0.1); CORS + WS-origin gates read `AI_CORS_ORIGINS` (default `*` — any origin, safe only because credentials are never enabled; explicit comma list restricts); the web client derives the API base URL from the browser hostname at runtime (`engine-base-url.ts`: `NEXT_PUBLIC_AI_SERVICE_URL` override → `http://${window.location.hostname}:8420` → `localhost` server-side). Hotfix also fixes: SEA direct-run detection (`require("node:sea").isSea()` — argv shape differs by invocation style), installer honesty gate (silent `--version` = hard fail), bootstrap venv recovery (poisoned partial `.venv` removal + distro-specific `apt install python3.XX-venv` hint), and doctor venv-capability probe with bootstrap preflight.
|
||||
24. **D-039 Single-port deploy + unattended dev (v0.3.6)** — only :8420 is reachable behind HAProxy, so the web app ships as a **static export** (`output: 'export'`, `NEXT_PUBLIC_AI_SERVICE_URL=self` → relative same-origin fetches) served by the ai-service itself (`AI_WEB_STATIC_DIR` StaticFiles mount, default off; `nextcraft dev` auto-wires it when `apps/web/out` exists). Daemon surface: `dev -d` (detached, `~/.nextcraft/run/<clone-hash>/dev.{pid,log}`), `stop`, `log [-n N|-f]`. Durable state (DB `~/.nextcraft/data/`, sandbox workdirs `~/.nextcraft/sandboxes/`) moves **out of the repo** (founder directive; `expanduser` validator makes `AI_DB_PATH=~/...` env overrides work). API routes beat the static mount; unknown paths serve the export's 404.html.
|
||||
|
||||
### v0.3 Architecture Decisions (from Research — Credential Engines)
|
||||
|
||||
9. **D-024 Sandbox isolation = Linux namespaces via `unshare`** — per-learner sandbox runs as a subprocess entered into fresh user+mount+pid+network namespaces (`unshare --user --map-root-user --mount --pid --fork --net`). Probe-verified on this box: in-namespace uid=0, **network fully isolated** (0 interfaces), learner writes land in a per-sandbox directory; proc-remount not permitted here but not required. Chosen because no container runtime (docker/podman/bwrap/firejail) exists on the box and there is no sudo. A `SandboxBackend` protocol abstracts the spawner so a future containerd/runc backend can replace namespace-spawning without touching callers.
|
||||
10. **D-025 Sandbox scope = coding IDE only (v0.3)** — the sandbox fabric provisions a single build environment (shell + filesystem + run/test). REQ-F-021's design-tool and simulation environments are deferred to v0.4; one real build path proves the full credential pipeline (telemetry → trace → grade → defense).
|
||||
11. **D-026 Telemetry = WebSocket ingest + SQLite ordered event log** — in-sandbox capture agent streams structured events over WebSocket to `ai_service` (`/v1/telemetry/ingest`); events persisted to SQLite with a per-(learner,task) monotonic `seq` for gap detection, giving durability + at-least-once delivery + replay without a message broker.
|
||||
12. **D-027 First persistence = SQLite, protocol-wrapped** — introduces a real DB (`ai_service/data/*.db`) for telemetry traces, grades, variants, and defenses. Access via SQLModel. Every store is a protocol (`TraceStore`, `VariantStore`, `DefenseStore`, `GradeStore`) with a SQLite implementation — Postgres-migration-ready, mirroring D-019's SessionStore pattern.
|
||||
13. **D-028 Process-trace grading = hybrid deterministic + LLM** — deterministic features (test pass/fail, edit count, error/fix cycles, idle gaps, command categories) computed in code into a compact trace digest; the digest feeds an Assessor-style rubric prompt and returns structured scores via D-020 JSON defense. The LLM never sees the raw trace — only the digest.
|
||||
14. **D-029 Variant generation = seeded template instantiation** — task templates with typed parameter slots; an LLM instantiates a unique variant per learner from a seed; seed+parameters persisted (D-027) for grading fairness and proctoring cross-check.
|
||||
15. **D-030 Voice = provider-agnostic, mock-first, browser-fallback** — a `VoiceProvider` protocol (mirror of `LLMProvider`) with STT (OpenAI-compatible `/audio/transcriptions`) + TTS (`/audio/speech`) against a configurable endpoint, a deterministic mock (canned transcript/audio) for tests, and browser-native `SpeechRecognition`/`speechSynthesis` as a no-key fallback. Voice defense reuses `BaseAgent` + the existing SSE pipeline (new `Examiner` agent).
|
||||
16. **D-031 No new apps — extend ai-service** — telemetry, grading, variant, voice, and sandbox orchestration are new modules inside `apps/ai-service` (sharing the LLM pool, config, and session infra). Only the in-sandbox capture agent is a separate tiny Python process shipped into the namespace. No new top-level `apps/` entry.
|
||||
17. **D-032 Capacity = single-box, 1–5 concurrent sandboxes** — concurrency guard returns 503 when the sandbox pool is full. No queueing, no horizontal scaling in v0.3 (solo-founder/pilot scale).
|
||||
|
||||
---
|
||||
|
||||
## Components
|
||||
@@ -89,69 +53,38 @@ Deliberately **not** used: openai-python SDK (the `LLMProvider` protocol is the
|
||||
|
||||
| Component | Description | Boundaries | Depends On |
|
||||
|-----------|-------------|------------|------------|
|
||||
| `ai_service/main.py` | FastAPI app factory, lifespan (httpx client pool, provider factory, SandboxManager + reaper loop, SQLite engine stores on app.state), CORS (localhost only, incl. PUT for file writes), /health; v0.3.6: optional StaticFiles mount of the exported web app when `AI_WEB_STATIC_DIR` is set (single-port deploy, D-039) | App entry | config, llm, agents, api, engines |
|
||||
| `ai_service/main.py` | FastAPI app factory, lifespan (httpx client pool, provider factory), CORS (localhost only), /health | App entry | config, llm, agents, api |
|
||||
| `ai_service/config.py` | pydantic-settings Settings (env_prefix="AI_", env_file, SecretStr key) | Configuration only | None |
|
||||
| `ai_service/api/` | Endpoints: chat.py (POST /v1/chat/stream, SSE), lab.py, assessment.py (POST /v1/assessment/evaluate v0.2 + POST /v1/assessment/grade v0.3), proctor.py, mentor.py, sandboxes.py (lifecycle + files/exec routes, G-5 abuse gates), telemetry.py (WS ingest + trace/gaps reads), variants.py (seeded per-learner variants), defense.py (defense loop, REQ-3-006); deps.py (DI) | Composes agents + sessions + engines; never imported by llm/ or agents/ | agents, llm, sandbox, telemetry, grading, variants, voice |
|
||||
| `ai_service/llm/` | types.py (Message; ChatDelta/ChoiceDelta removed in P3 — no consumers), base.py (LLMProvider protocol), openai_compat.py (ollama-cloud + local), mock.py (deterministic), factory.py | Never imports agents/ or api/ | config |
|
||||
| `ai_service/agents/` | base.py (BaseAgent ABC), registry.py, session.py (SessionStore), structured.py (JSON defense), coach/tutor/lab/assessor/proctor/mentor.py + examiner.py (seventh agent, v0.3) | Never imports api/ | llm, prompts, corpus, telemetry |
|
||||
| `ai_service/api/` | Endpoints: chat.py (POST /v1/chat/stream, SSE), lab.py, assessment.py, proctor.py, mentor.py; deps.py (DI) | Composes agents + sessions; never imported by llm/ or agents/ | agents, llm |
|
||||
| `ai_service/llm/` | types.py (Message, ChatDelta), base.py (LLMProvider protocol), openai_compat.py (ollama-cloud + local), mock.py (deterministic), factory.py | Never imports agents/ or api/ | config |
|
||||
| `ai_service/agents/` | base.py (BaseAgent ABC), registry.py, session.py (SessionStore), structured.py (JSON defense), coach/tutor/lab/assessor/proctor/mentor.py | Never imports api/ | llm, prompts, corpus |
|
||||
| `ai_service/prompts/` | Per-agent system prompt constants + render_context functions (str.format_map) | Data only | None |
|
||||
| `ai_service/corpus/` | Mock engine inputs: learner_context.py, telemetry.py (Lab/Proctor scenarios), artifacts.py (pre-baked artifacts, rubrics, transcripts). Since v0.3 P6 these are DORMANT, test-only fixtures (dormant-header noted) — the live learner path uses real engine inputs | Pydantic-typed; aligned with TS packages/mock-data by convention | None |
|
||||
| `scripts/` | bootstrap.sh (venv + pip install idempotent), dev.sh (exports keys from .ciagent/.env.secrets → uvicorn), test.sh (pytest), lint.sh (ruff check, v0.2 G-3) | Dev entry points | pyproject.toml |
|
||||
| `ai_service/corpus/` | Mock engine inputs: learner_context.py, telemetry.py (Lab/Proctor scenarios), artifacts.py (pre-baked artifacts, rubrics, transcripts) | Pydantic-typed; aligned with TS packages/mock-data by convention | None |
|
||||
| `scripts/` | bootstrap.sh (venv + pip install idempotent), dev.sh (exports keys from .ciagent/.env.secrets → uvicorn), test.sh (pytest) | Dev entry points | pyproject.toml |
|
||||
| `tests/` | conftest.py (mock provider, settings override, TestClient), health, llm (MockTransport parser), agents (framework + per-agent), api (SSE stream tests) | Mock provider only — no cloud | all |
|
||||
|
||||
**Module boundary rules:** `llm/` never imports `agents/` or `api/`; `agents/` never imports `api/`; `api/` composes both via DI. `corpus/` is the only home of mock engine data. Prompts are code — versioned and reviewed in git.
|
||||
|
||||
### apps/ai-service — v0.3 Credential Engine modules (NEW)
|
||||
|
||||
| Component | Description | Boundaries | Depends On |
|
||||
|-----------|-------------|------------|------------|
|
||||
| `ai_service/sandbox/` | `backend.py` (SandboxBackend protocol), `unshare_backend.py` (userns/mount/pid/net spawner, D-024), `manager.py` (lifecycle: create/list/snapshot/destroy + concurrency guard D-032), `workdir.py` (per-sandbox fs layout) | Never imports api/ or agents/; spawns subprocesses only | config |
|
||||
| `ai_service/telemetry/` | `models.py` (TelemetryEvent, TraceSpan), `store.py` (TraceStore protocol + SQLite impl D-027), `ingest.py` (WebSocket /v1/telemetry/ingest, seq gap detection D-026; v0.5 REQ-5-007/D-045: emits advisory `seq_ack` frames — highest-contiguous received seq per successful append; capture agent trims its spool to the ack, closing the replay-margin gap) | Persistence; never imports agents/ | config |
|
||||
| `ai_service/grading/` | `features.py` (deterministic trace digest D-028 — kind-agnostic by construction, pinned over design/sim traces in v0.5), `engine.py` (rubric scoring orchestration), `store.py` (GradeStore) | LLM only via digest; never sees raw trace | llm, telemetry, prompts |
|
||||
| `ai_service/variants/` | `templates.py` (task template library; v0.5 REQ-5-005/D-044: `environment: Literal[build,design,simulation]` registry + per-kind starter files + harness/test commands, G-15 shlex-roundtrip validation), `generator.py` (seeded LLM instantiation D-029), `store.py` (VariantStore; idempotent `_ensure_v05_columns` backfill for pre-v0.5 DBs) | LLM via structured output | llm, grading |
|
||||
| `ai_service/voice/` | `base.py` (VoiceProvider protocol D-030), `browser.py` (native SR/TTS fallback descriptor), `mock.py` (deterministic), `openai_audio.py` (v0.5 REQ-5-001, D-040: real server STT/TTS against OpenAI-compatible `/audio/transcriptions` + `/audio/speech` on the shared httpx pool — CUT-1/G-7 seam CLOSED), `factory.py` (provider selection by `AI_VOICE_PROVIDER`, G-11 boot-safe fallback to mock), `defense_store.py` (DefenseStore: transcripts + integrity signals, D-027) | Never imports agents/ or api/ | config |
|
||||
| `ai_service/identity/` | **v0.5 NEW (REQ-5-003/004, D-042/43):** `base.py` (IdentityProvider protocol: submit/poll/verify), `mock.py` (deterministic approve-on-policy mock; verdicts carry a `mock` marker, A-304), `store.py` (5th D-027 store: identity_record table — derived `age_band`, document **refs**, PII never stored raw); age-gate dependencies `require_verified_age`/`require_verified_adult` (gate composition D-043: G-5 allowlist → identity verdict → rate caps; mounted on variants/sandbox-create/defense-start + the G-18 marketplace stub), exposed via `api/identity.py` (`/v1/identity/*`) | Never imports agents/; api/ composes it via DI | config |
|
||||
| `ai_service/agents/examiner.py` | Seventh agent: oral defense examiner; streams over existing SSE, consumes process traces + emits integrity signals | reuses BaseAgent (D-018) | llm, prompts, telemetry |
|
||||
| `ai_service/data/*.db` | SQLite databases (telemetry/grades/variants/defenses) — **v0.3.6: default moved to `~/.nextcraft/data/nextcraft.db` (state out of the repo; AI_DB_PATH overrides, ~ expanded)** | outside repo (home) | — |
|
||||
| `scripts/sandbox-agent.py` | Tiny in-namespace capture process shipped into the sandbox; streams telemetry to ingest | standalone | stdlib only |
|
||||
|
||||
**Boundary additions:** `sandbox/`, `telemetry/`, `grading/`, `variants/`, `voice/` are engine modules — they never import `api/` (which composes them via DI) and never import `agents/` (agents call engines through narrow interfaces, not vice versa).
|
||||
|
||||
### apps/cli — Nextcraft Bootstrap CLI (v0.4 NEW)
|
||||
|
||||
| Component | Description | Boundaries | Depends On |
|
||||
|-----------|-------------|------------|------------|
|
||||
| `src/index.ts` | Entry: arg parsing (no deps beyond node stdlib at runtime), command dispatch, `--help`/`--version`, exit-code contract (0 ok / 1 failure / 2 usage) | CLI surface only | commands/ |
|
||||
| `src/commands/` | `doctor.ts` (prereq checks + actionable errors), `bootstrap.ts` (pnpm install + scripts/bootstrap.sh wrapper + env template copy + key validation), `verify.ts` (health: venv imports, ports, env, build readiness), `dev.ts` (thin passthrough to scripts/dev.sh) | Compose checks/ + lib/; spawn scripts — never re-implement them | checks/, lib/ |
|
||||
| `src/checks/` | Pure check functions: `check-command.ts` (binary-on-PATH + version compare), `check-env.ts` (template diff, required/optional key classification) | Pure logic, unit-testable, no fs side effects at import | None |
|
||||
| `src/lib/` | `spawn.ts` (subprocess with timeout + inherited stdio), `log.ts` (✓/✗/warn output formatter) | Shared utilities | None |
|
||||
| `scripts/build-binary.mjs` | esbuild → CJS bundle → Node SEA injection → `dist/nextcraft-linux-x64` + sha256 sidecar | Build-time only | esbuild (dev dep) |
|
||||
| `scripts/install.sh` | The one-liner install script served from repo raw: Gitea latest-release resolve → download + checksum verify → ~/.local/bin; source-bootstrap fallback | Standalone POSIX sh | forge API |
|
||||
| `tests/` | node:test unit tests: command dispatch, check logic, env template diff, install-script shellcheck-style assertions | Fixtures only — never mutate repo state | src/ |
|
||||
|
||||
**Boundary rules:** the CLI never imports from `apps/web`, `packages/*`, or `ai_service` Python modules — it orchestrates them exclusively via subprocess/filesystem. Runtime deps: node stdlib only (no runtime npm deps; esbuild is dev-only). The binary embeds the bundle; `scripts/bootstrap.sh` remains the single source of bootstrap truth (D-034).
|
||||
|
||||
### apps/web — Next.js Application
|
||||
|
||||
| Component | Description | Boundaries | Depends On |
|
||||
|-----------|-------------|------------|------------|
|
||||
| `app/(learner)/` | Learner surface route group: landing, catalog, competency stack, dashboard, byte viewer, build surface (`/build/[competencyId]` — real in-browser build), defense surface (`/defend/[competencyId]` — live oral defense + grading) | Learner-only routes and layouts | packages/ui, packages/mock-data, packages/types |
|
||||
| `app/(learner)/` | Learner surface route group: landing, catalog, competency stack, dashboard, byte viewer, sandbox mockup, assessment mockup | Learner-only routes and layouts | packages/ui, packages/mock-data, packages/types |
|
||||
| `app/(marketplace)/` | Marketplace surface route group: job board, job detail, employer profile, search/filter, pricing | Marketplace-only routes and layouts | packages/ui, packages/mock-data, packages/types |
|
||||
| `app/(employer)/` | Employer dashboard route group: overview, talent search, candidate profile, posting management | Employer-only routes and layouts | packages/ui, packages/mock-data, packages/types |
|
||||
| `app/(admin)/` | Admin surface route group: overview, learner management, competency graph viewer, moderation | Admin-only routes and layouts | packages/ui, packages/mock-data, packages/types |
|
||||
| `app/layout.tsx` | Root layout: theme provider, navigation shell, responsive container | All routes | packages/ui |
|
||||
| `components/` | Surface-specific components (learner/, marketplace/, employer/, admin/) plus shared chrome (navigation-shell, header/footer, role-switcher, theme-provider, breadcrumbs, dark-mode-toggle); v0.3 learner: build-surface, sandbox-terminal (read-only exec output), defense-session | App-level components | packages/ui |
|
||||
| `hooks/` | use-chat-stream.ts — SSE client hook: fetch + ReadableStream, byte buffering + frame reassembly, idempotent AbortController cleanup; use-sandbox-session.ts (v0.3) — sandbox lifecycle for the build session: create on task open, destroy on unmount, mid-start failure cleanup, 503/403/429 honest surfaces | Client components only | ai-service SSE / engine API |
|
||||
| `lib/` | sse.ts (shared SSE frame parser — CRLF normalization + `: ping` immunity, v0.2 G-1), breadcrumbs.ts, format.ts, engine-base-url.ts (v0.3.5: runtime API base — env override → browser hostname → localhost), engine-client.ts (v0.3: typed fetch client for /v1/sandboxes, files/exec, variants, grade, defense, traces) | Pure utilities | None |
|
||||
| `components/` | Surface-specific components (not shared across surfaces) | Per-surface only | packages/ui |
|
||||
|
||||
### packages/ui — Shared Component Library
|
||||
|
||||
| Component | Description | Boundaries | Depends On |
|
||||
|-----------|-------------|------------|------------|
|
||||
| `tokens/` | Design tokens as TS constants: colors, spacing, radii, shadows, breakpoints (mirrored as Tailwind v4 `@theme` tokens in apps/web globals.css) | Foundation layer — no dependencies | None |
|
||||
| `primitives/` | Button, Input, Card, Badge, Avatar (v0.1) + TerminalFrame, TelemetryStatus, MicControl, GradeBadge, TranscriptViewer (v0.3 build/defense surfaces) — each with a Storybook story | Atomic UI components | tokens, packages/types |
|
||||
|
||||
Composite/layout/theme components (navigation shell, tables, chat panels, graph viewer, theme provider) live in `apps/web/components/` as app-level components, not in packages/ui.
|
||||
| `design-tokens/` | CSS custom properties: color palette, typography scale, spacing system, breakpoints, shadows, radii | Foundation layer — no dependencies | None |
|
||||
| `primitives/` | Button, Input, Card, Badge, Avatar, Dialog, Tabs, Progress, Tooltip, Skeleton, Toast | Atomic UI components | design-tokens |
|
||||
| `composites/` | Navigation, Table, SearchBar, FilterPanel, ChatInterface, GraphViewer, ArtifactCard, CompetencyBadge, JobCard, CandidateCard, MetricCard | Composite components built from primitives | primitives, packages/types |
|
||||
| `layouts/` | Container, Grid, Sidebar, SplitPanel, DashboardLayout | Layout components | primitives, design-tokens |
|
||||
| `theme/` | Theme provider, CSS variable overrides per surface (learner, marketplace, employer, admin) | Theme context | design-tokens |
|
||||
|
||||
### packages/mock-data — Mock Data Layer
|
||||
|
||||
@@ -162,8 +95,7 @@ Composite/layout/theme components (navigation shell, tables, chat panels, graph
|
||||
| `candidates.ts` | 15+ mock candidate profiles with artifacts, process traces, defense scores, microcredentials | Typed mock data | packages/types |
|
||||
| `employers.ts` | 10+ mock employer profiles with logos, descriptions, open positions | Typed mock data | packages/types |
|
||||
| `learner-progress.ts` | Mock learner progress data: active competencies, completion percentages, recent artifacts | Typed mock data | packages/types |
|
||||
| `admin.ts` | Admin surface mock data: platform metrics, activity feed, system health, learner roster (admin view), moderation queues | Typed mock data | packages/types |
|
||||
| `ai-scenarios.ts` | AI engine-input scenario IDs + display metadata for the learner agent panels; IDs string-identical to `ai_service/corpus/` (D-021) | Typed mock data | packages/types |
|
||||
| `ai-tutor-responses.ts` | Pre-scripted AI tutor chat responses for Coach and Tutor agent mockups | Typed mock data | packages/types |
|
||||
|
||||
### packages/types — Shared Types
|
||||
|
||||
@@ -173,42 +105,11 @@ Composite/layout/theme components (navigation shell, tables, chat panels, graph
|
||||
| `marketplace.ts` | Job, Employer, Candidate, JobPosting, TalentMatch, SearchFilter | Marketplace types | None |
|
||||
| `user.ts` | Learner, Admin, EmployerUser, AgeGroup, Role | User types | None |
|
||||
| `ui.ts` | Component props, theme config, breakpoint definitions | UI types | None |
|
||||
| `telemetry.ts` | TelemetryEvent/ExecResult wire shapes for the live build surface (v0.3) | Engine types | None |
|
||||
| `variants.ts` | Variant/TaskTemplate shapes for per-learner task statements (v0.3) | Engine types | None |
|
||||
| `grading.ts` | GradeRecord/RubricScore shapes for live grading display (v0.3) | Engine types | None |
|
||||
| `defense.ts` | DefenseSession/transcript/integrity-signal shapes for the defense surface (v0.3) | Engine types | None |
|
||||
|
||||
---
|
||||
|
||||
## Data Flow
|
||||
|
||||
### v0.3 credential flow (current)
|
||||
|
||||
```
|
||||
[learner build surface /build/*] [learner defense surface /defend/*]
|
||||
file CRUD + Run/Test (HTTP) mic MediaRecorder / typed + TTS playback
|
||||
│ │
|
||||
▼ ▼
|
||||
[api/sandboxes files/exec] ──exec──▶ [namespace sandbox] [api/defense start/answer/finish]
|
||||
│ │ capture agent │
|
||||
│ ▼ (WS telemetry) ▼
|
||||
│ [api/telemetry ingest] [DefenseStore (SQLite)]
|
||||
│ │ SQLite │ transcript + integrity signals
|
||||
│ ▼ │
|
||||
│ [TraceStore] ────▶ [GradingEngine: digest (grading/features)
|
||||
│ │ + rubric LLM (D-028)] ──▶ [GradeStore]
|
||||
│ ▼ ▼
|
||||
└──▶ Lab agent (live digest) Assessor (grade output) / Proctor (integrity)
|
||||
Examiner agent (SSE) ◀── defense sessions
|
||||
variants: [api/variants] ◀── [VariantStore (seeded, D-029)] ── per-learner task statements
|
||||
```
|
||||
|
||||
- Lab consumes the live trace digest; Assessor consumes grading-engine output; Proctor consumes telemetry + defense integrity signals (REQ-3-007) — no mock fallback in the learner path (v0.2 corpus scenarios are dormant test-only fixtures).
|
||||
- The learner's path is: variant task → in-sandbox build (telemetry streams to SQLite) → grade My Work (rubric scores from the real trace) → oral defense → verdict.
|
||||
- Flooded/gapped traces are terminal: ingest closes 1008 and marks INCOMPLETE_FLOODED (G-3); the grader returns UNGRADABLE_TRACE_INCOMPLETE (G-4) — no credential from an incomplete trace.
|
||||
|
||||
### v0.2 chat flow (complete, still live)
|
||||
|
||||
```
|
||||
[packages/mock-data + packages/types] [ai_service/corpus]
|
||||
│ (TS, web surfaces) │ (Python, agent inputs)
|
||||
@@ -223,31 +124,14 @@ Composite/layout/theme components (navigation shell, tables, chat panels, graph
|
||||
(https://ollama.com/v1)
|
||||
```
|
||||
|
||||
- Web surfaces remain server-component-first; client components (chat, filters, graph viewer, dark mode toggle) fetch directly from ai-service over SSE (A-002: no Next.js API-route proxy).
|
||||
- The LLM provider layer is a dumb pipe — OpenAI-compatible chunks pass through byte-identically; envelope logic (meta/done/error) lives only in the API layer (D-016).
|
||||
- The v0.2 corpus scenarios (`ai_service/corpus/`) are retained as dormant, test-only fixtures (dormant-header noted); they are no longer inputs to the live learner path.
|
||||
- Web surfaces remain server-component-first; client components (chat, filters, graph viewer, dark mode toggle) fetch directly from ai-service over SSE (A-002: no Next.js API-route proxy in v0.2).
|
||||
- The LLM provider layer is a dumb pipe — OpenAI-compatible chunks pass through byte-identical; envelope logic (meta/done/error) lives only in the API layer (D-016).
|
||||
- Lab/Assessor/Proctor read mock scenarios from `ai_service/corpus/` — real engines are v0.3+.
|
||||
- All automated tests use the deterministic mock provider; the cloud is for manual probes only.
|
||||
|
||||
---
|
||||
|
||||
## Build Order (v0.4)
|
||||
|
||||
1. **Bootstrap CLI core** — apps/cli package: doctor checks (node/pnpm/python3/git/unshare), bootstrap wrapper (pnpm install + scripts/bootstrap.sh + .env template + key validation), verify health check, dev passthrough; unit tests
|
||||
2. **Binary build + release pipeline** — esbuild bundle → Node SEA binary (`nextcraft-linux-x64`) + sha256 sidecar; install.sh one-liner (Gitea latest-release resolve + checksum verify + PATH install); release-asset upload wired into the ship flow (ongoing binaries from v0.4 onward)
|
||||
3. **Install docs + fresh-clone E2E** — README quickstart (one-liner → doctor → bootstrap → dev), CLI reference, fresh-clone end-to-end test proving a clean clone reaches a running stack
|
||||
|
||||
## Build Order (v0.3 — complete)
|
||||
|
||||
1. **Sandbox fabric** — SandboxBackend protocol + unshare namespace spawner + lifecycle manager (create/list/snapshot/destroy) + concurrency guard + per-sandbox workdir; isolation + resource-limit probes
|
||||
2. **Live build telemetry** — TelemetryEvent models + SQLite TraceStore + WebSocket ingest endpoint + seq gap detection + in-sandbox capture agent
|
||||
3. **Process-trace grading engine** — deterministic feature/digest computation + rubric scoring via LLM structured output + GradeStore; calibrated against v0.2 mock corpora
|
||||
4. **Variant task generation** — template library + seeded LLM instantiation + VariantStore + difficulty normalization anchors
|
||||
5. **Oral / voice defense** — VoiceProvider protocol + STT/TTS + mock + browser fallback + Examiner agent + transcript/integrity-signal capture
|
||||
6. **Agent re-grounding + learner surface integration** — Lab/Assessor/Proctor consume real telemetry/grades/defense signals; learner sandbox mockup → real in-browser build/run (Run/Test buttons executing in a namespace sandbox, read-only exec-output panel — no interactive shell, CUT-2/G-8); assessment mockup → live defense + live grading
|
||||
|
||||
---
|
||||
|
||||
## Build Order (v0.2 — complete)
|
||||
## Build Order (v0.2)
|
||||
|
||||
1. **AI service scaffolding** — apps/ai-service: FastAPI app, config, provider layer (ollama-cloud/local/mock), SSE chat endpoint, pytest harness, turbo integration
|
||||
2. **Agent framework** — BaseAgent, registry, session store, structured output, prompts scaffolding, learner-context corpus
|
||||
@@ -260,19 +144,15 @@ The v0.1 build order (monorepo → types → mock data → tokens → primitives
|
||||
|
||||
---
|
||||
|
||||
## Future Architecture (Post-v0.4, for reference)
|
||||
## Future Architecture (Post-v0.2, for reference)
|
||||
|
||||
v0.4 delivers distribution (CLI + binary releases); later milestones fill in the remaining platform:
|
||||
v0.2 delivers the ai-service skeleton that later milestones fill in:
|
||||
|
||||
- **In-memory sessions → PostgreSQL + Drizzle/SQLModel** — SessionStore + v0.3 TraceStore/GradeStore/VariantStore/DefenseStore protocols swap SQLite→Postgres with no API changes
|
||||
- **userns subprocess sandboxes → containerd/runc backend** — D-024 `SandboxBackend` protocol swap; same lifecycle API
|
||||
- **Coding-IDE sandbox → design tool + simulation environments** — REQ-F-021 full scope (v0.5)
|
||||
- **Mock corpus → real engines** — Lab consumes real sandbox telemetry (v0.3 sandbox fabric); Assessor grades real process traces (v0.3 assessment engine); Proctor consumes real identity/attention signals (v0.3 identity verification)
|
||||
- **In-memory sessions → PostgreSQL + Drizzle ORM** — SessionStore protocol swap, no API changes
|
||||
- **Mock provider → per-agent model routing** — provider factory already selects by config; per-agent `AI_<AGENT>_MODEL` overrides
|
||||
- **No auth → real KYC + sessions** — **deferred per founder directive; moved to v0.5 with D-016**; REQ-F-017 identity/age-gating lands post-v0.4. Age-gating remains the v0.1 visual flow mockup
|
||||
- **Mock voice → real server STT/TTS (openai-audio provider)** — CUT-1/G-7 seam moved to v0.5 per D-016; VoiceProvider protocol is the drop-in point
|
||||
- **linux x64 binary → macOS arm64 + auto-update** — D-036 defers non-linux targets (unverifiable on this box); `nextcraft upgrade` (self-replace from latest release) is the natural v0.5+ follow-up
|
||||
- **Exec-telemetry seq-lease / replay-margin fix** — the P6-lesson one-line ACK gap moves to v0.5 per D-016
|
||||
- **No auth → real KYC + sessions** — A-008 dropped in v0.3 when identity verification lands
|
||||
- **No search → Semantic vector search (pgvector)** — Filter UI replaced with vector similarity search
|
||||
- **No payments → Payment processing** — Pricing page replaced with real subscription/payment flows
|
||||
|
||||
The monorepo structure (apps/web + apps/ai-service + apps/cli + packages/*) accommodates further apps without restructuring.
|
||||
The monorepo structure (apps/web + apps/ai-service + packages/*) accommodates further apps without restructuring.
|
||||
@@ -0,0 +1,8 @@
|
||||
{
|
||||
"phase": 0,
|
||||
"stage": "mvp_ux_check",
|
||||
"milestone": "v0.2",
|
||||
"phase_role": "pre_execution",
|
||||
"attempts": 0,
|
||||
"updated_at": "2026-09-11T15:55:00Z"
|
||||
}
|
||||
+23
-103
@@ -1,116 +1,36 @@
|
||||
# Nextcraft v0.3 — GRILL.md (Adversarial Review Verdict)
|
||||
# Nextcraft v0.2 — GRILL.md (Adversarial Review Verdict)
|
||||
|
||||
**Stage:** GRILL, Phase 0 pre-execution · **Verdict:** GO-WITH-CHANGES · **Confidence:** 0.72
|
||||
|
||||
## Summary
|
||||
|
||||
The credential-pipeline architecture (telemetry → trace → grade → defense) is sound and correctly sequenced. Three plan claims did NOT survive contact with this box and were correct before execution. The central problem: the plan **overstated sandbox resource-limit enforcement** and deferred KYC **without closing the resulting no-auth local abuse vector**. Fixed via binding decisions G-1..G-6 + scope cuts CUT-1/CUT-2 — no redesign required.
|
||||
**Stage:** GRILL, Phase 0 pre-execution · **Verdict:** GO-WITH-CHANGES · **Confidence:** 0.82
|
||||
|
||||
## Per-Axis Findings
|
||||
|
||||
| Axis | Verdict | Rationale |
|
||||
|------|---------|-----------|
|
||||
| Feasibility | CONCERN | Core `unshare` userns/mount/pid/net isolation probe-verified (uid=0 in-ns, network isolated, writes contained). But rlimit enforcement is partial: `RLIMIT_NPROC` scopes to the real host uid (5 sandboxes share one pids budget) and no disk-quota tool exists on the box. |
|
||||
| Over-scoping | CONCERN | 39 tasks across 5 new subsystems + learner-surface rewrite for a solo founder. Voice-real-path and the interactive terminal relay are separable from the pipeline proof → cut (CUT-1, CUT-2). |
|
||||
| Architecture risk | CONCERN | SandboxBackend/TraceStore protocols are the right seams. Overclaimed "limits enforced" + WS ingest "drop-oldest on unbounded growth" contradicted the at-least-once grading guarantee. |
|
||||
| Phase sequencing | PASS | P1 sandbox → P2 telemetry → P3 grading → P4 variants → P5 voice → P6 integration is a correct dependency DAG. |
|
||||
| Verification honesty | FAIL (fixed) | Must-Haves asserted "resource limits enforced + observable" (CPU/memory/disk/time quotas) the named mechanism cannot satisfy; probe tests would pass while the guarantee was false. Corrected by G-1. |
|
||||
| Cost/quota | CONCERN | LLM/voice mock-gated (good). No-auth sandbox creation + unbounded disk + shared NPROC let one learner starve others at zero cost → closed by G-5. |
|
||||
| Milestone honesty | CONCERN (fixed) | Release note disclosed KYC deferral + IDE-only but was silent on partial resource-limit enforcement → G-6. |
|
||||
| Security | FAIL (fixed) | `POST /v1/sandboxes {learner_id}` client-supplied over localhost CORS let any local process mint sandboxes/flood/exhaust shared resources → G-5 abuse control ships despite KYC deferral. |
|
||||
| Operability | CONCERN (advisory) | In-memory sandbox registry loses handles on restart (orphaned namespaces) → a-1 startup reaper. |
|
||||
| Feasibility | PASS | Environment claims verified (venv works, all 8 pinned packages on PyPI for py3.11, port 8420 free, secrets gitignored). Stall points pre-mitigated (idempotent bootstrap, 300s read timeout, idempotent abort; reactStrictMode confirmed in next.config). |
|
||||
| Over-scoping | CONCERN (mild) | 4-layer JSON defense justified; 500-cap LRU is over-spec but cheap — keep, extend no further. Real gaps: no Python lint (fixed via G-3), cache dirs in gitignore (fixed), script path resolution (advisory c). |
|
||||
| Architecture risk | CONCERN | sse-starlette `: ping` frames invisible to short TestClient streams — parser gap fixed via G-1. Agent registration had two contradictory patterns — fixed via G-4. Cloud outage blast radius contained by mock-only tests. |
|
||||
| Phase sequencing | PASS | Integration-last correct: P1 freezes the SSE contract before client code exists; A-002 eliminates dev-server buffering trap. P4 heaviest but mechanical. |
|
||||
| Verification honesty | CONCERN | Persona distinctness circular against self-authored mocks (disclosed; cloud probe optional). P6 must-haves manual-only. D-021 ID alignment unmechanized (advisory b). Disclosed honestly. |
|
||||
| Cost/quota | PASS | Cloud burn bounded (~<100K tokens milestone-wide, manual probes only). Mock-only rule structurally enforced; mechanical guard via advisory (a). |
|
||||
| Milestone honesty | PASS (conditional) | Mock-engine caveats present everywhere that matters. Release note content now bound by G-5; dead `aiTutorResponses` disposal bound by G-5. |
|
||||
|
||||
## Binding Decisions (applied to PLAN.md/ARCHITECTURE-adjacent docs/REQUIREMENTS.md/ROADMAP.md/PROJECT.md)
|
||||
## Binding Decisions (applied to PLAN.md/ROADMAP.md)
|
||||
|
||||
- **G-1 (BINDING) — Resource-limit claims match the deliverable mechanism.** Memory (RLIMIT_AS) + CPU (RLIMIT_CPU) + single-file (RLIMIT_FSIZE) + wall-clock reaper are kernel-enforced; per-sandbox pids and hard disk quota are NOT kernel-enforceable without cgroup delegation/sudo → documented as accepted v0.3 risk. Applied to REQ-3-002, PLAN P1 Must-Haves + Task 1-2-02, ROADMAP P1 criteria.
|
||||
- **G-2 (BINDING) — Disk cap via manager workdir-size sweep.** `AI_SANDBOX_MAX_WORKDIR_MB` (default 512MB); sweep snapshots+destroys over-cap sandboxes and logs an integrity signal; closes the unbounded-`dd` hole. Applied to PLAN Task 1-2-01 + 1-2-02(e) + P1 Must-Haves.
|
||||
- **G-3 (BINDING) — Telemetry flood control WITHOUT silent drop.** Bounded queue; on overflow or >`AI_TELEMETRY_MAX_EVENTS_PER_TASK` (default 50k) → WS close 1008 + trace marked `INCOMPLETE_FLOODED` (Proctor signal). Silent drop-oldest forbidden (corrupts grading). Applied to PLAN Task 2-2-02 + P2 Must-Haves.
|
||||
- **G-4 (BINDING) — Grader refuses incomplete/gapped traces.** `grade()` gates on `TraceStore.gaps()` + `INCOMPLETE_FLOODED` → returns `verdict=UNGRADABLE_TRACE_INCOMPLETE`; no credential from a gapped trace. Applied to PLAN Task 3-2-01 + P3 Must-Haves.
|
||||
- **G-5 (BINDING) — No-auth abuse control at MVP scale.** Per-learner sandbox cap + global create-rate cap (429) + server-side `learner_id` allowlist (403) so the unauthenticated surface can't exhaust shared NPROC/disk. Ships WITH the milestone even though KYC is deferred. Applied to PLAN Task 1-3-01 + PROJECT A-110.
|
||||
- **G-6 (BINDING) — Disclose partial enforcement in the P7 release note.** Item (e): which limits are kernel-enforced vs best-effort, and that full enforcement is deferred to the post-MVP containerd backend. Applied to PLAN P7 release-note honesty block.
|
||||
- **G-1 (BINDING):** SSE frame parser in Task 6-1-01 must ignore frames with no `data:` lines (sse-starlette ping keep-alive). Added to Action + Phase 6 must-have.
|
||||
- **G-2 (BINDING):** P6 end-to-end verification may run with `AI_PROVIDER=mock` fallback — real ai-service over HTTP is the requirement; provider choice is service-internal. Prevents cloud outage blocking P6.
|
||||
- **G-3 (BINDING):** ruff (check-only) added: Task 1-1-04, pyproject dev extra, scripts/lint.sh, root `ai:lint`, turbo `ai#lint`, Phase 1 must-have.
|
||||
- **G-4 (BINDING):** Mentor/Proctor registered centrally in `registry.py` (single registration pattern), matching P3/P4.
|
||||
- **G-5 (BINDING):** v0.2.0 release note must state Lab/Assessor/Proctor run on mock engine inputs (v0.3+ for real); dead `aiTutorResponses` export disposed of in P7.
|
||||
|
||||
## Scope Cuts (accepted — preserve the end-to-end credential pipeline)
|
||||
## Advisory (non-binding; applied where cheap)
|
||||
|
||||
- **CUT-1 (G-7) — Real server STT/TTS (`OpenAIAudioProvider`) deferred to v0.4.** Voice is mock-first (D-030); the `/audio/*` real path can never run in CI and was the least-verifiable surface. v0.3 proves the full defense *dialogue* + integrity-signal pipeline over mock + browser-native fallback; the `VoiceProvider` protocol is the future drop-in seam. Applied to PLAN Phase 5 Goal + Task 5-1-01/5-4-01 + P5 Must-Haves.
|
||||
- **CUT-2 (G-8) — Interactive xterm.js shell relay deferred to v0.4.** The credential pipeline needs *process events* (Run/Test + file edits), not a live keystroke-level shell — the most fragile real-time piece, unverifiable without a real terminal. The build panel becomes Run/Test buttons + read-only exec output render; `@xterm/*` is NOT a v0.3 dependency. Applied to PLAN Env-facts, P6 Goal, Task 6-2-01/6-2-02/6-3-01, P6 Must-Haves, MVP/UX sections.
|
||||
- Variants (Phase 4) and the Examiner agent dialogue **KEPT** — both are on the credential critical path (anti-collusion + the defense dialogue).
|
||||
|
||||
## Advisory (applied)
|
||||
|
||||
- **a-1** Startup reaper: on lifespan boot, scan `AI_SANDBOX_DIR`, reap workdirs whose recorded pid is dead, log a warning. → Task 1-2-01.
|
||||
- **a-2** `RLIMIT_FSIZE` (~50MB) as a cheap partial single-file disk guard in the spawner's `preexec_fn`. → Task 1-2-01 + 1-2-02(c).
|
||||
- **a-3** SQLite `PRAGMA journal_mode=WAL` + `synchronous=NORMAL` at engine creation (avoids `database is locked` under concurrent ingest + grader reads). → Task 2-1-02.
|
||||
- **a-4** Grading prompt note: treat high edit/command churn with no test-progress as a process-quality negative (softens digest-gaming naivety). → Task 3-2-01.
|
||||
- **a-5** Variant fairness envelope: two variants of one template must compute digests within the template's expected feature envelope ("same bar" is testable). → Task 4-2-01 Verify.
|
||||
- (a) conftest asserts provider is MockProvider — **applied in P1 implementation**
|
||||
- (b) pytest reads packages/mock-data as text, asserts corpus IDs appear — **applied in P4 implementation**
|
||||
- (c) scripts resolve repo root via script-relative dirname — **folded into Task 1-1-04/1-1-03**
|
||||
- (d) `.pytest_cache/` + `.ruff_cache/` in .gitignore — **applied**
|
||||
- (e) keep LRU as-is; no further session-store sophistication in v0.2
|
||||
- (f) Task 6-2-02 REQ tag fixed to REQ-2-011 (routing) — **applied**
|
||||
|
||||
## Outcome
|
||||
|
||||
**GO** — all six binding decisions and both scope cuts applied to PLAN.md / REQUIREMENTS.md / ROADMAP.md / PROJECT.md before Phase 1 execution. No axis requires escalation (all resolvable at confidence ≥ 0.85). The milestone no longer claims resource enforcement it cannot deliver, and the no-auth abuse vector is closed at MVP scale.
|
||||
|
||||
---
|
||||
# Nextcraft v0.4 — GRILL.md (Adversarial Review Verdict)
|
||||
|
||||
**Stage:** GRILL, Phase 0 pre-execution · **Verdict:** GO-WITH-CHANGES · **Confidence:** 0.83
|
||||
|
||||
## Summary
|
||||
|
||||
The distribution milestone is small, founder-directed (D-016, confidence 0.99), and additive (zero changes to the running credential pipeline). The plan's central risk: **Node SEA was probe-verified as a flag, not as a working build** — the v0.3 lesson (A-101: probe the mechanism, not the existence) applies. Second gap: a binary whose `--version` lies (stale package.json) would poison the "ongoing binaries" contract. Third: sed-based JSON parsing in install.sh is a fragility + integrity risk. Fourth: "ongoing binaries" has no enforcement mechanism beyond prose. All four closed by binding decisions G-101..G-104 below. No scope cuts required — the milestone is already minimal.
|
||||
|
||||
## Per-Axis Findings
|
||||
|
||||
| Axis | Verdict | Rationale |
|
||||
|------|---------|-----------|
|
||||
| Business case | PASS | Founder directive explicit + recorded (D-016). Evidence of need: live Gitea probe shows latest release v0.2.8 with ZERO assets; bootstrap requires repo archaeology (scripts found only via package.json spelunking). |
|
||||
| Scope | PASS | 5 REQs, 3 execution phases, one focused surface (apps/cli + scripts). Smallest milestone yet. macOS arm64 already cut (D-036, unverifiable here). |
|
||||
| Feasibility | CONCERN (fixed) | SEA flag exists on node v24.15.0, but no end-to-end SEA binary was built during RESEARCH. postject availability assumed (`npx postject` — needs npm registry reachability, unproven). Zipapp fallback requires python3 on target — an honest-degradation ladder, not a silent downgrade. → G-101. |
|
||||
| Honest versioning | CONCERN (fixed) | `--version` from package.json would print a stale hardcoded version inside a per-release binary — breaks upgrade detection + the one-liner's re-run-to-upgrade promise. → G-102. |
|
||||
| Install integrity | CONCERN (fixed) | sed/grep JSON parsing is brittle; a parse failure must never fall through to installing an unverified artifact. Exact asset-name matching + hard-degrade to source instructions. → G-103. |
|
||||
| Sequencing | PASS | P1 CLI (source-runnable) → P2 binary+pipeline → P3 docs+E2E matches dependency order; each phase ships independently. |
|
||||
| Cost/quota | PASS | Zero new paid infra; binaries built on-box; Gitea releases free. Dev-only esbuild dep. |
|
||||
| Risks | CONCERN (fixed) | Top 3: SEA end-to-end (→ G-101 live probe FIRST in P2), npm registry reachability for esbuild (→ proven by P1's pnpm install must-have), Gitea asset-upload token scope (→ live-proven at the v0.3.2 ship itself). |
|
||||
| Adoption/operability | PASS | Consumer = founder + future pilots; one command replaces README archaeology. Rollback trivial (rm ~/.local/bin/nextcraft). No server changes. |
|
||||
|
||||
## Binding Decisions (applied to PLAN.md)
|
||||
|
||||
- **G-101 (BINDING) — SEA live-build probe is the FIRST P2 action.** Task 2-1-01 builds a real binary before anything depends on it; the build script encodes the fallback ladder explicitly (SEA → zipapp with "requires python3" honesty). If SEA fails on this box, zipapp becomes primary with the docs stating the requirement — no silent claim of node-less operation.
|
||||
- **G-102 (BINDING) — Version stamping at build time.** `build-binary` accepts the shipping tag and stamps it into the bundle (`NEXTCRAFT_VERSION` replace); `--version` prints it; install E2E asserts the installed binary reports the tag it was downloaded from. A binary may never report a version it was not built as.
|
||||
- **G-103 (BINDING) — Install-script integrity hard-degrade.** install.sh matches assets by EXACT name (`nextcraft-linux-x64`, `nextcraft-linux-x64.sha256`); any parse/lookup/download failure degrades to source-bootstrap instructions (exit 0) — never installs unverified or name-approximate artifacts. Checksum mismatch = hard stop, exit 1, explicit do-not-run message. dash-safe POSIX sh, no jq.
|
||||
- **G-104 (BINDING) — Ongoing-binaries enforcement.** Every ship from v0.3.2 onward MUST run `scripts/release-assets.sh <tag>` after tag+merge (best-effort, non-blocking, `release_pending` escalation on failure — but attempted + logged every release). The final-phase audit gate includes "milestone release carries both assets" as a check. This makes the founder's "ongoing binaries" directive a pipeline property, not prose.
|
||||
|
||||
## Escalations
|
||||
|
||||
None. All four concerns resolved at confidence ≥ 0.85. No axis requires founder escalation (directive already explicit).
|
||||
|
||||
## Outcome
|
||||
|
||||
**GO** — G-101..G-104 applied to PLAN.md before Phase 1 execution. The milestone claims only what its probes prove, and the ongoing-binaries contract has an enforcement mechanism.
|
||||
|
||||
---
|
||||
|
||||
# v0.5 GRILL (Phase 0, 2026-09-13) — Real Server Voice, KYC/Identity, Sandbox Environments, Seq-Lease
|
||||
|
||||
**Verdict: GO-WITH-CHANGES · Confidence: 0.78** — CUT-3 + G-9..G-18 binding; applied to PLAN.md before P1 execution.
|
||||
|
||||
**Evidence:** every load-bearing plan claim verified against ground-truth code and held (fmt bug defense.py:189, descriptor flow defense.py:143, supervisor frame discard sandbox-agent.py:487-496, no spool cap, dead test_command field, one-line replay margin sandbox-agent.py:416-431, factory rejection). The research layer (D-040..D-046) sold no vapor; the misses were all *interaction* surface.
|
||||
|
||||
## Binding items (all applied to PLAN.md)
|
||||
|
||||
- **CUT-3**: MH-3e identity E2E downgraded from real-app uvicorn harness to `TestClient(create_app(...))` — identity is plain JSON HTTP; the uvicorn harness stays where transport matters (P1 durability, P4 in-ns exec).
|
||||
- **G-9**: `verified_pilot` conftest fixture MUST land with the identity gates (Wave 3-2) — variants/defense are ungated today; without the fixture the existing 413-test suite 403s en masse. MH-M6 reworded: regressions zero *for verified allowlisted pilot learners*.
|
||||
- **G-10**: engine-client must discriminate 403 reasons (`verify_cta` → new `VerifyRequiredError`; allowlist detail → existing `NotAllowlistedError`) — today every 403 collapses into NotAllowlistedError, which would render a verify-CTA as an allowlist lie.
|
||||
- **G-11**: misconfigured `openai-audio` must never crash the boot — lifespan catches the factory error, logs loudly, falls back to mock (descriptor honestly reads mock). Unattended deploy (D-039) survival rule.
|
||||
- **G-12**: client recording bound — auto-stop at 180s default + visible timer + timeslice; 413 renders "re-record" honestly (never silent answer loss).
|
||||
- **G-13**: identity submit caps — one active pending per learner (409 on resubmit) + per-learner rate cap (429); each submission becomes vendor money later.
|
||||
- **G-14**: spool-bound overflow is by-design gap creation — pinned: dropped-counter > 0 → replayed trace gapped → ungradable (never silently-truncated-but-gradable); worst-case spool ≈256MB documented against the 512MB G-2 sweep.
|
||||
- **G-15**: argv contract — templates validate shlex-roundtrip at definition (no quotes/globs); TS splits whitespace-only (no shlex in browsers); exec policy matches EXACT argv tokens (never prefix); `sh -c` passthrough disallowed for design/sim kinds (the digest-gaming vector).
|
||||
- **G-16**: `voice_tts_format` is a `Literal["mp3","wav","opus"]` enum, not a free string — it feeds a Content-Type.
|
||||
- **G-17**: MH-4e "seq-ack intact" was hope-shaped; merged into MH-4d as concrete assertions (stored seqs contiguous 0..N exactly once; spool ≤ ack margin; or cite the P1 suite where a fake agent runs).
|
||||
- **G-18**: the marketplace gated stub proves the gate then returns 501 + `stub: true` + mock markers — never a fabricated "applied" outcome (A-304 honesty house rule).
|
||||
|
||||
## Advisories (recorded)
|
||||
|
||||
a-6 ack on dedup'd appends too (tight margin); a-7 MH-1d 3× runs = one-time ship validation, not per-CI; a-8 manual voice probe must be executable-by-anyone-with-keys (fixture wav + fixed phrase, falsifiable asserts); a-9 413 Content-Length fast path before buffering; a-10 ROADMAP P4 "Depends On: 4" self-typo → fix to 0; a-11 TS `TaskVariant` fields required on the wire; a-12 identity insert-only growth fine at pilot scale; a-13 `recorder.start(1000)`; a-14 do NOT build ack batching unless a flood test shows drainer stall; a-15 provider must always carry `descriptor` (missing it would badge server as mock).
|
||||
|
||||
## Escalations
|
||||
|
||||
None. All four seams remain founder-locked via D-016 / REQ-5-001..007; every finding resolved at confidence ≥ 0.65.
|
||||
GO — all five binding decisions applied to PLAN.md/ROADMAP.md/.gitignore/ARCHITECTURE.md before Phase 1 execution. No axis requires escalation.
|
||||
+70
-155
@@ -2,13 +2,11 @@
|
||||
|
||||
## Persona Roster
|
||||
|
||||
> **v0.5 update (RESEARCH, lead-developer assessment):** milestone = Real Server Voice, KYC/Identity, Sandbox Environments, Seq-Lease. **Reactivated:** voice-engineer (real server STT/TTS, its v0.3 territory), sandbox-engineer (design/sim environment kinds + the seq-ack protocol on the fabric/agent it owns), frontend-engineer (defense audio upload, identity enrollment flow, kind-aware build surface). **New custom persona: identity-engineer** (KYC/identity domain — 5th store, provider protocol, age-gate dependencies, PII hygiene). **security-auditor re-activated (phase-specific)** for identity PII + age-gate bypass + audio upload attack surface (phases 1-4 review, final phase). ai-engineer light-touch (no LLM-facing work this milestone). backend-engineer retains settings/factory wiring + turbo/root scripts. cli-engineer inactive (v0.3.6 hotfix shipped; no CLI work planned). design-system-engineer/data-engineer inactive (one TS types extension only — data-engineer light-touch for variants/identity types).
|
||||
|
||||
### lead-developer
|
||||
```yaml
|
||||
active: true
|
||||
phase_specific: false
|
||||
reason: Coordinates the four v0.5 seams across voice/identity/sandbox/telemetry territories; resolves wave-order (D-046: seq-lease first) and identity-gate composition (D-043) boundaries
|
||||
reason: Coordinates task decomposition across web, AI service, and data territories; resolves conflicts between frontend, backend, and AI personas
|
||||
domain: coordination
|
||||
frameworks:
|
||||
- next.js
|
||||
@@ -27,143 +25,85 @@ territory:
|
||||
- "apps/ai-service/pyproject.toml"
|
||||
```
|
||||
|
||||
### voice-engineer
|
||||
### frontend-engineer
|
||||
```yaml
|
||||
active: true
|
||||
phase_specific: false
|
||||
reason: v0.3 territory reactivated — owns REQ-5-001/002: OpenAIAudioProvider (D-040) on the shared httpx pool, factory signature change, defense route fixes (fmt strip, size guard, media_type), descriptor mode=server, web audio POST
|
||||
domain: ai-media
|
||||
reason: Phase 6 learner-surface integration — useChatStream hook, agent switcher, streaming/error/loading states, Lab/Assessor/Proctor output panels. Owns all page components, layouts, and surface-specific UI.
|
||||
domain: frontend
|
||||
frameworks:
|
||||
- httpx
|
||||
- fastapi
|
||||
- pytest
|
||||
- react
|
||||
- next.js
|
||||
- tailwindcss
|
||||
- lucide-react
|
||||
- recharts
|
||||
- react-flow
|
||||
constraints:
|
||||
- provider-agnostic-protocol (D-030 drop-in; descriptor wins selection)
|
||||
- never-call-cloud-in-tests (MockTransport byte-contract pins)
|
||||
- key-redaction (mirror openai_compat _sanitize)
|
||||
- bounded-audio-in-memory (10MB guard before provider call)
|
||||
- component-first
|
||||
- server-components-default
|
||||
- minimal-client-js
|
||||
- sse-client-buffering (buffer bytes, split frames on \n\n, join data: lines)
|
||||
- abortcontroller-cleanup (idempotent abort in effect cleanup)
|
||||
- responsive-all-breakpoints
|
||||
- dark-mode-support
|
||||
territory:
|
||||
- "apps/ai-service/ai_service/voice/**"
|
||||
- "apps/ai-service/tests/voice/**"
|
||||
- "apps/web/components/learner/defense-session.tsx"
|
||||
- "apps/web/**"
|
||||
- "packages/ui/**"
|
||||
- "packages/mock-data/**"
|
||||
- "packages/types/**"
|
||||
```
|
||||
|
||||
### identity-engineer
|
||||
### data-engineer
|
||||
```yaml
|
||||
active: true
|
||||
phase_specific: false
|
||||
reason: v0.5 custom persona (RESEARCH, D-042/43) — owns the new identity module: IdentityProvider protocol + mock, 5th SQLite store (DefenseStore pattern), /v1/identity router, age-gate dependencies (allowlist → identity → rate caps), derived age_band + document refs (PII minimal), caplog sentinel scrub test
|
||||
domain: identity
|
||||
reason: Owns TS mock data layer schema and typed definitions. Does NOT own the Python corpus (that is ai-engineer territory) — the two are aligned by documented convention (D-021).
|
||||
domain: data
|
||||
frameworks:
|
||||
- fastapi
|
||||
- sqlmodel
|
||||
- pytest
|
||||
- typescript
|
||||
constraints:
|
||||
- pii-never-stored-raw (refs + derived bands only, A-305)
|
||||
- pii-never-logged (caplog sentinel pin)
|
||||
- mock-verdict-honesty (mock marker rides every response, A-304)
|
||||
- gate-composition-order (G-5 allowlist first, identity second, 429 caps last)
|
||||
- schema-first
|
||||
- type-safe
|
||||
- migration-ready
|
||||
- mock-data-only
|
||||
territory:
|
||||
- "apps/ai-service/ai_service/identity/**"
|
||||
- "apps/ai-service/tests/identity/**"
|
||||
```
|
||||
|
||||
### sandbox-engineer
|
||||
```yaml
|
||||
active: true
|
||||
phase_specific: false
|
||||
reason: v0.3 territory reactivated — owns REQ-5-005/006 (template-layer env registry, starter contents, exec command policy) AND REQ-5-007 (seq-ack frame in ingest + agent spool trim, mid-burst regression test) — both live on the fabric/agent surfaces it built
|
||||
domain: infra
|
||||
frameworks:
|
||||
- python
|
||||
- linux-namespaces
|
||||
- pytest
|
||||
constraints:
|
||||
- stdlib-only-agent (AST-pinned sandbox-agent.py imports)
|
||||
- acks-are-advisory (gap detection stays authoritative; G-3/G-4 unchanged)
|
||||
- thread-safe-spool-trim (under _emit_lock, atomic Spool.rewrite)
|
||||
- digest-kind-agnostic (features derive from event kinds — pinned)
|
||||
territory:
|
||||
- "apps/ai-service/ai_service/sandbox/**"
|
||||
- "apps/ai-service/ai_service/telemetry/**"
|
||||
- "apps/ai-service/scripts/sandbox-agent.py"
|
||||
- "apps/ai-service/ai_service/variants/templates.py"
|
||||
- "apps/ai-service/tests/sandbox/**"
|
||||
- "apps/ai-service/tests/telemetry/**"
|
||||
- "packages/types/**"
|
||||
- "packages/mock-data/**"
|
||||
```
|
||||
|
||||
### backend-engineer
|
||||
```yaml
|
||||
active: true
|
||||
phase_specific: false
|
||||
reason: Owns the settings surface both new providers hang off (voice_base_url/key/models, identity provider selection), factory + main.py lifespan wiring (voice factory gains http_client), .env.example documentation
|
||||
reason: Reactivated for v0.2 — owns apps/ai-service infrastructure: FastAPI app, settings, SSE plumbing, endpoints, monorepo/turbo integration, test harness.
|
||||
domain: backend
|
||||
frameworks:
|
||||
- fastapi
|
||||
- pydantic-settings
|
||||
- bash
|
||||
- turborepo
|
||||
- pnpm
|
||||
constraints:
|
||||
- secrets-via-env-only (D-014; keys never in code or commits)
|
||||
- idempotent-scripts
|
||||
territory:
|
||||
- "apps/ai-service/ai_service/config.py"
|
||||
- "apps/ai-service/ai_service/main.py"
|
||||
- "apps/ai-service/ai_service/voice/factory.py"
|
||||
- "apps/ai-service/scripts/**"
|
||||
- "apps/ai-service/.env.example"
|
||||
- "package.json"
|
||||
- "turbo.json"
|
||||
```
|
||||
|
||||
### frontend-engineer
|
||||
```yaml
|
||||
active: true
|
||||
phase_specific: false
|
||||
reason: Reactivated for v0.5 UX: defense-session audio POST (recorded blob upload), identity enrollment flow (submit → pending → verified states + verify-CTA surfaces), build-surface kind-awareness (variant environment + test_command), engine-client identity + audio functions
|
||||
domain: frontend
|
||||
frameworks:
|
||||
- react
|
||||
- next.js
|
||||
- tailwindcss
|
||||
constraints:
|
||||
- component-first
|
||||
- server-components-default
|
||||
- honest-state-surfaces (provider badge: mock vs browser vs server; unverified labels)
|
||||
territory:
|
||||
- "apps/web/**"
|
||||
- "packages/ui/**"
|
||||
- "packages/types/**"
|
||||
```
|
||||
|
||||
### security-auditor
|
||||
```yaml
|
||||
active: true
|
||||
phase_specific: true
|
||||
reason: v0.5 re-activated — identity PII (storage + logs), age-gate bypass review, audio upload attack surface (10MB guard, format confusion), exec command policy (per-kind allowlist), TTS media_type. Phases 1-4 reviews + final phase
|
||||
domain: security
|
||||
frameworks:
|
||||
- pytest
|
||||
- uvicorn
|
||||
- pydantic
|
||||
- httpx
|
||||
- pytest
|
||||
constraints:
|
||||
- STRIDE-classified
|
||||
- pii-never-stored-raw
|
||||
- pii-never-logged
|
||||
- bounded-uploads
|
||||
- provider-agnostic-boundaries (llm/ imports nothing from agents/ or api/)
|
||||
- streaming-first
|
||||
- no-database-v0.2
|
||||
- secrets-via-env-only
|
||||
- mock-provider-in-tests
|
||||
territory:
|
||||
- "apps/ai-service/ai_service/identity/**"
|
||||
- "apps/ai-service/ai_service/api/defense.py"
|
||||
- "apps/ai-service/ai_service/api/sandboxes.py"
|
||||
- "apps/ai-service/ai_service/voice/**"
|
||||
- "apps/ai-service/ai_service/main.py"
|
||||
- "apps/ai-service/ai_service/config.py"
|
||||
- "apps/ai-service/ai_service/api/**"
|
||||
- "apps/ai-service/scripts/**"
|
||||
- "apps/ai-service/package.json"
|
||||
- "apps/ai-service/tests/api/**"
|
||||
- "turbo.json"
|
||||
```
|
||||
|
||||
### ai-engineer
|
||||
```yaml
|
||||
active: true
|
||||
phase_specific: false
|
||||
reason: Light-touch v0.5 — no LLM-facing work (voice STT/TTS is media plumbing, not model work; examiner agent unchanged); guards the agent/engine boundaries the new surfaces touch
|
||||
reason: Custom persona for v0.2 — owns the LLM provider layer, agent framework, prompt library, structured outputs, and mock corpora for the six tutor agents.
|
||||
domain: ai
|
||||
frameworks:
|
||||
- pydantic
|
||||
@@ -171,84 +111,59 @@ frameworks:
|
||||
- pytest
|
||||
constraints:
|
||||
- provider-agnostic-protocol
|
||||
- prompts-are-code
|
||||
- json-defensive-parsing
|
||||
- never-call-cloud-in-tests
|
||||
- delta-passthrough
|
||||
territory:
|
||||
- "apps/ai-service/ai_service/llm/**"
|
||||
- "apps/ai-service/ai_service/agents/**"
|
||||
- "apps/ai-service/ai_service/prompts/**"
|
||||
```
|
||||
|
||||
### cli-engineer
|
||||
```yaml
|
||||
active: false
|
||||
phase_specific: false
|
||||
reason: v0.3 persona — CLI shipped complete (v0.3.6 daemon surface); no v0.5 CLI work planned. Reactivated if milestone work touches apps/cli.
|
||||
domain: cli
|
||||
frameworks:
|
||||
- node
|
||||
- typescript
|
||||
- node:test
|
||||
- esbuild
|
||||
- node-sea
|
||||
- posix-sh
|
||||
constraints:
|
||||
- stdlib-only-runtime
|
||||
- thin-wrapper
|
||||
- timeout-every-spawn
|
||||
- fail-loud-exit-codes
|
||||
territory:
|
||||
- "apps/cli/**"
|
||||
- "scripts/install.sh"
|
||||
- "scripts/release-assets.sh"
|
||||
- "apps/ai-service/ai_service/corpus/**"
|
||||
- "apps/ai-service/tests/llm/**"
|
||||
- "apps/ai-service/tests/agents/**"
|
||||
```
|
||||
|
||||
### design-system-engineer
|
||||
```yaml
|
||||
active: false
|
||||
active: true
|
||||
phase_specific: false
|
||||
reason: No design-token or primitive work planned in v0.5 (existing primitives — MicControl, GradeBadge, TranscriptViewer — cover the voice surfaces); roster retained.
|
||||
reason: Owns the shared component library, design tokens, and visual consistency. Light duty in v0.2: Phase 6 may need new primitives (agent-switcher control, stream-status indicator, error toast variant).
|
||||
domain: frontend
|
||||
frameworks:
|
||||
- tailwindcss
|
||||
- storybook
|
||||
- lucide-react
|
||||
constraints:
|
||||
- design-token-driven
|
||||
- wcag-aa-contrast
|
||||
- dark-mode-required
|
||||
- consistent-across-surfaces
|
||||
territory:
|
||||
- "packages/ui/**"
|
||||
```
|
||||
|
||||
### data-engineer
|
||||
### security-auditor
|
||||
```yaml
|
||||
active: false
|
||||
phase_specific: false
|
||||
reason: Light-touch via frontend-engineer territory (variants/identity TS type extensions); no schema/mock-data work beyond the two typed additions.
|
||||
domain: data
|
||||
frameworks:
|
||||
- typescript
|
||||
constraints:
|
||||
- schema-first
|
||||
- type-safe
|
||||
- dual-schema-sync (TS/Python changes made in both places)
|
||||
territory:
|
||||
- "packages/types/**"
|
||||
- "packages/mock-data/**"
|
||||
reason: No auth in v0.2 (A-008); CORS is localhost-only; no real user data. Security review handled by verifier's STRIDE analysis layer plus a Phase 7 checklist item: secrets hygiene (key absent from code/logs/commits/errors), localhost-only CORS, no PII in prompts.
|
||||
domain: security
|
||||
frameworks: []
|
||||
constraints: []
|
||||
territory: []
|
||||
```
|
||||
|
||||
## Phase-Specific Personas
|
||||
|
||||
| Persona | Phases | Removed After |
|
||||
|---------|--------|---------------|
|
||||
| security-auditor | 1 (seq-ack/agent), 2 (audio upload), 3 (identity PII), 4 (exec policy), 5 (final review) | milestone complete |
|
||||
|
||||
All other personas span the milestone. Deactivated personas receive no tasks.
|
||||
None for v0.2. All active personas span the entire milestone. data-engineer and design-system-engineer are light-touch after Phase 1.
|
||||
|
||||
## Territory Conflict Resolution
|
||||
|
||||
| Conflict | Resolution |
|
||||
|----------|------------|
|
||||
| voice-engineer vs backend-engineer (factory.py) | backend-engineer owns factory.py settings wiring + signature; voice-engineer owns openai_audio.py + its factory branch (provider construction), consumes config via DI |
|
||||
| sandbox-engineer vs frontend-engineer (build-surface) | sandbox-engineer owns engine-side kinds/templates/policy; frontend-engineer owns the web surface + client session hook; wire contract = VariantResponse TS types |
|
||||
| identity-engineer vs sandbox-engineer (gates) | identity-engineer owns the IdentityGate dependencies; sandbox-engineer owns the sandbox route they mount on (gate order D-043 is a joint review) |
|
||||
| security-auditor vs identity-engineer | identity-engineer implements; security-auditor reviews + may patch security defects directly in identity/ + api/defense.py (its territory) |
|
||||
| lead-developer vs any | lead-developer coordinates only, does not directly modify code files |
|
||||
| frontend-engineer vs data-engineer (packages/types, packages/mock-data) | data-engineer owns type definitions and mock data schema; frontend-engineer consumes them. If changes needed, data-engineer updates types first. |
|
||||
| frontend-engineer vs design-system-engineer (packages/ui) | design-system-engineer owns design tokens and primitive components; frontend-engineer owns composite components and page-level UI. |
|
||||
| ai-engineer vs data-engineer (mock data duplication) | ai-engineer owns `ai_service/corpus/` (Python); data-engineer owns `packages/mock-data` (TS). Shared entity IDs and shapes kept aligned by documented convention (D-021): cross-referencing file headers, identical `comp-*` ID strings. |
|
||||
| backend-engineer vs ai-engineer (apps/ai-service) | backend-engineer owns app shell, config, API endpoints, scripts, and test harness; ai-engineer owns llm/, agents/, prompts/, corpus/. Boundary: `ai_service/api/` (backend) composes `ai_service/agents/` (AI) via DI — agents never import api/. |
|
||||
| lead-developer vs any | lead-developer coordinates only, does not directly modify code files. |
|
||||
+405
-165
@@ -1,182 +1,422 @@
|
||||
# Nextcraft v0.5 — PLAN.md
|
||||
# Nextcraft v0.2 — PLAN.md
|
||||
|
||||
## Overview
|
||||
|
||||
**Milestone:** v0.5 — Real Server Voice, KYC/Identity, Sandbox Environments, Seq-Lease
|
||||
**Tag line:** v0.4.x patches (P0 → v0.4.0 … P5 → v0.4.5 = milestone release)
|
||||
**Branch:** milestone/v0.5-real-voice-identity-envs
|
||||
This plan covers execution phases 1-6 of milestone v0.2 (AI Tutor Architecture): the six AI tutor agents (Coach, Tutor, Lab, Assessor, Proctor, Mentor) as real LLM-backed services in a new `apps/ai-service` Python FastAPI application, wired into the existing v0.1 learner surface with streaming responses. Phases are strictly sequential (P1→P6); within each phase, Wave 1 tasks are parallelizable (no cross-file dependencies) and later waves depend on earlier ones.
|
||||
|
||||
**Environment facts (probe-verified, apply throughout):** python 3.11.2 (venv at apps/ai-service/.venv), node v24.15.0, pnpm 12.3.4, no docker/podman/sudo, `unshare` userns verified working, SQLite is the only persistence (D-027), tests NEVER call the cloud (mock providers + httpx MockTransport), secrets only in gitignored `.ciagent/.env.secrets` / `.env*` (D-006/D-014). Runtime state lives in `~/.nextcraft/` (D-039) — tests pin to tmp dirs; sandbox workdir fixtures use the bind-mount-safe `sandbox_dir` fixture (overlayfs /tmp breaks userns binds).
|
||||
**Environment facts (apply throughout):** Python via `python3 -m venv` (no system pip, no uv); pnpm 12.3.4 via corepack; ai-service port **8420**; default model `gemma4:31b` (config via `AI_TUTOR_MODEL`); ollama-cloud base `https://ollama.com/v1` (OpenAI-compatible, Bearer auth); API keys live only in gitignored `.ciagent/.env.secrets`, exported by `scripts/dev.sh` — never in code, commits, or logs; all automated tests use the deterministic mock provider and **never call the cloud**.
|
||||
|
||||
**Wave order (D-046, binding):** P1 seq-lease → P2 voice → P3 identity → P4 environments → P5 final. Seq-lease lands first: it fixes the transport before environment phases add reconnecting telemetry producers; it touches no stores, no web, no settings.
|
||||
| Phase | Name | Requirements | Waves | Personas |
|
||||
|-------|------|-------------|-------|----------|
|
||||
| 1 | AI service scaffolding | REQ-2-001, 002, 003 | 3 | backend-engineer, ai-engineer |
|
||||
| 2 | Agent framework | REQ-2-004 | 2 | ai-engineer, backend-engineer |
|
||||
| 3 | Coach + Tutor agents | REQ-2-005, 006 | 3 | ai-engineer, backend-engineer |
|
||||
| 4 | Lab + Assessor agents | REQ-2-007, 008 | 4 | ai-engineer, backend-engineer, data-engineer |
|
||||
| 5 | Proctor + Mentor agents | REQ-2-009, 010 | 3 | ai-engineer, backend-engineer |
|
||||
| 6 | Learner surface integration | REQ-2-011, 012 | 3 | frontend-engineer, design-system-engineer |
|
||||
|
||||
**Research decisions D-040..D-046 are binding design contracts.** This plan operationalizes them; it does not re-litigate them.
|
||||
---
|
||||
|
||||
## Phase 1: AI Service Scaffolding
|
||||
|
||||
**Requirements:** REQ-2-001, REQ-2-002, REQ-2-003
|
||||
**Goal:** apps/ai-service runs under uvicorn, /health responds, provider-agnostic LLM layer with ollama-cloud/local/mock providers, SSE chat streaming verified, pytest suite green with mock provider, turbo integration wired
|
||||
|
||||
### Wave 1: Service shell + LLM core (parallel — no shared files)
|
||||
|
||||
#### Task 1-1-01: FastAPI app scaffolding
|
||||
- **Persona:** backend-engineer — **REQ:** REQ-2-001
|
||||
- **Files:** `apps/ai-service/pyproject.toml`, `apps/ai-service/ai_service/main.py`, `apps/ai-service/ai_service/config.py`, `apps/ai-service/ai_service/__init__.py`, `apps/ai-service/tests/conftest.py`, `apps/ai-service/tests/test_health.py`, `apps/ai-service/.env.example`, `apps/ai-service/README.md`
|
||||
- **Action:** pydantic-settings `Settings` (env_prefix `AI_`, env_file, `SecretStr` key, port 8420, provider select, `AI_TUTOR_MODEL` default `gemma4:31b`). FastAPI app factory in `main.py` with lifespan stub (httpx client pool comes in Wave 2), CORS localhost-only (A-008), `GET /health`. pyproject with pinned deps (fastapi, uvicorn, pydantic, pydantic-settings, httpx, sse-starlette, pytest, pytest-asyncio) and dev extra. conftest: settings override + TestClient fixture. `.env.example` documents all `AI_*` vars; README documents venv setup and dev workflow.
|
||||
- **Verify:** `scripts/bootstrap.sh && scripts/test.sh` — test_health passes; `curl localhost:8420/health` returns 200
|
||||
|
||||
#### Task 1-1-02: LLM types, protocol, mock provider
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-2-002
|
||||
- **Files:** `apps/ai-service/ai_service/llm/types.py`, `apps/ai-service/ai_service/llm/base.py`, `apps/ai-service/ai_service/llm/mock.py`, `apps/ai-service/ai_service/llm/__init__.py`
|
||||
- **Action:** `types.py`: pydantic `Message` (role/content), `ChatDelta` (OpenAI-compatible chunk shape). `base.py`: `LLMProvider` protocol — async `stream_chat(messages, model, response_format=None) -> AsyncIterator[ChatDelta]`; the provider is a dumb pipe, no envelope logic (D-016 keeps envelope in API layer). `mock.py`: deterministic scripted provider (hash-seeded token streams, scripted failure modes: connect error, mid-stream error, malformed JSON) for tests and CI.
|
||||
- **Verify:** mock provider importable and deterministic; two identical calls yield identical streams
|
||||
|
||||
#### Task 1-1-03: Monorepo integration (shim + turbo + scripts)
|
||||
- **Persona:** backend-engineer — **REQ:** REQ-2-001
|
||||
- **Files:** `apps/ai-service/package.json`, `apps/ai-service/scripts/bootstrap.sh`, `apps/ai-service/scripts/dev.sh`, `apps/ai-service/scripts/test.sh`, `turbo.json` (update), `package.json` (root, update)
|
||||
- **Action:** zero-dependency shim `package.json` in apps/ai-service with `dev`/`test`/`bootstrap` script entries. Turbo passthrough tasks `ai#dev`, `ai#test`, `ai#bootstrap` (`cache: false`, `outputs: []`). Root scripts `ai:dev`, `ai:test`, `ai:bootstrap`. `bootstrap.sh`: idempotent `python3 -m venv .venv` + pip install. `dev.sh`: exports keys from `.ciagent/.env.secrets` → uvicorn on 8420. `test.sh`: pytest via venv.
|
||||
- **Verify:** `corepack pnpm install && pnpm ai:bootstrap && pnpm ai:test` runs pytest through turbo; re-running bootstrap is a no-op
|
||||
|
||||
#### Task 1-1-04: Python lint (ruff, check-only)
|
||||
- **Persona:** backend-engineer — **REQ:** REQ-2-001
|
||||
- **Files:** `apps/ai-service/pyproject.toml` (update: `[tool.ruff]` config + `ruff` in dev extra), `apps/ai-service/scripts/lint.sh`, `package.json` (root, update), `turbo.json` (update)
|
||||
- **Action:** Add `ruff` (check-only, no formatter) to the dev extra; `[tool.ruff]` with line-length 100, target py311. `scripts/lint.sh`: `.venv/bin/ruff check .` with repo-root path resolution via script-relative dirname (not CWD). Root script `ai:lint`, turbo passthrough `ai#lint` (`cache: false`, `outputs: []`). Run over the entire ai-service tree; fix all findings before phase ship (G-3).
|
||||
- **Verify:** `pnpm ai:lint` exits 0 on the Phase 1 codebase
|
||||
|
||||
### Wave 2: Real providers + SSE endpoint (depends on Wave 1)
|
||||
|
||||
#### Task 1-2-01: OpenAI-compatible provider + factory
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-2-002
|
||||
- **Files:** `apps/ai-service/ai_service/llm/openai_compat.py`, `apps/ai-service/ai_service/llm/factory.py`
|
||||
- **Action:** `openai_compat.py`: single `OpenAICompatProvider` for ollama-cloud (`https://ollama.com/v1`, Bearer) and local endpoints (base URL from settings); raw httpx against `/v1/chat/completions` with `stream: true`, byte-identical delta passthrough, tolerant of ollama-cloud quirks. Uses the lifespan-managed `httpx.AsyncClient` (10s connect / 300s read, D-017) — no openai SDK. `factory.py`: select provider from settings (`ollama-cloud` | `local` | `mock`).
|
||||
- **Verify:** provider constructs from settings for all 3 names; manual probe against ollama-cloud streams tokens (documented in README, not a test)
|
||||
|
||||
#### Task 1-2-02: Lifespan wiring + SSE chat endpoint
|
||||
- **Persona:** backend-engineer — **REQ:** REQ-2-003
|
||||
- **Files:** `apps/ai-service/ai_service/main.py` (update), `apps/ai-service/ai_service/api/deps.py`, `apps/ai-service/ai_service/api/chat.py`, `apps/ai-service/ai_service/api/__init__.py`
|
||||
- **Action:** Lifespan creates the shared `httpx.AsyncClient` and provider factory; deps.py provides provider via DI. `POST /v1/chat/stream` in chat.py implements the D-016 envelope: `meta` event (agent/session/model) flushed before first token → raw OpenAI chunks passed through as `data: {json}` → `done` event → `error` event before `[DONE]` on mid-stream failure; pre-first-byte failures return proper HTTP status codes. Headers `Cache-Control: no-cache`, `X-Accel-Buffering: no`; sse-starlette ping keep-alive.
|
||||
- **Verify:** `curl -N -X POST localhost:8420/v1/chat/stream` with mock provider shows meta event, token deltas, done, `[DONE]`
|
||||
|
||||
### Wave 3: Provider + endpoint test suites (depends on Wave 2)
|
||||
|
||||
#### Task 1-3-01: LLM provider tests (byte-exact, no cloud)
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-2-002
|
||||
- **Files:** `apps/ai-service/tests/llm/test_openai_compat.py`, `apps/ai-service/tests/llm/test_mock.py`, `apps/ai-service/tests/llm/__init__.py`
|
||||
- **Action:** httpx `MockTransport` tests parsing byte-exact fixture streams (happy path, empty delta, `[DONE]`, malformed line, mid-stream disconnect). Mock provider tests: determinism, scripted failure modes, response_format echo.
|
||||
- **Verify:** `pnpm ai:test` — llm suite green; zero network calls in tests
|
||||
|
||||
#### Task 1-3-02: SSE stream endpoint tests
|
||||
- **Persona:** backend-engineer — **REQ:** REQ-2-003
|
||||
- **Files:** `apps/ai-service/tests/api/test_chat_stream.py`, `apps/ai-service/tests/api/__init__.py`
|
||||
- **Action:** TestClient `client.stream()` tests: meta-first ordering, delta passthrough, done + `[DONE]` sentinel, mid-stream error event, pre-first-byte failure → HTTP status, required headers. pytest-asyncio auto mode (D-023).
|
||||
- **Verify:** `pnpm ai:test` — api suite green
|
||||
|
||||
### Must-Haves (Phase 1)
|
||||
- [ ] `scripts/bootstrap.sh` is idempotent; creates venv + installs deps without system pip
|
||||
- [ ] `pnpm ai:dev` starts uvicorn; `curl localhost:8420/health` returns 200
|
||||
- [ ] `pnpm ai:lint` exits 0 (ruff check over the ai-service tree) (G-3)
|
||||
- [ ] `pnpm ai:test` runs the full pytest suite via turbo and passes (mock provider only — no network)
|
||||
- [ ] SSE stream delivers tokens: meta event, incremental deltas, done, `[DONE]` observed via `curl -N`
|
||||
- [ ] Mid-stream failure emits `error` event before `[DONE]`; pre-first-byte failure returns HTTP error status
|
||||
- [ ] Provider factory resolves ollama-cloud / local / mock from settings; manual ollama-cloud probe documented in README
|
||||
- [ ] `llm/` imports nothing from `agents/` or `api/` (boundary rule holds)
|
||||
|
||||
---
|
||||
|
||||
## Phase 2: Agent Framework
|
||||
|
||||
**Requirements:** REQ-2-004
|
||||
**Goal:** Shared framework all six agents use: BaseAgent contract, session store, prompt library, registry, structured outputs — all tested against the mock provider
|
||||
|
||||
### Wave 1: Framework primitives (parallel — no shared files)
|
||||
|
||||
#### Task 2-1-01: BaseAgent ABC
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-2-004
|
||||
- **Files:** `apps/ai-service/ai_service/agents/base.py`, `apps/ai-service/tests/agents/test_base.py`, `apps/ai-service/ai_service/agents/__init__.py`, `apps/ai-service/tests/agents/__init__.py`
|
||||
- **Action:** `BaseAgent` ABC (D-018): `name`, `system_prompt`, `build_messages(history, learner_context)`, `stream_reply(...) -> AsyncIterator[ChatDelta]` (delegates to provider), `structured_reply(...)` (delegates to structured module, landed Wave 2). Subclass contract tested with a stub agent + mock provider.
|
||||
- **Verify:** `pnpm ai:test` — test_base green
|
||||
|
||||
#### Task 2-1-02: SessionStore protocol + in-memory implementation
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-2-004
|
||||
- **Files:** `apps/ai-service/ai_service/agents/session.py`, `apps/ai-service/tests/test_session.py`
|
||||
- **Action:** `SessionStore` protocol + `InMemorySessionStore` (D-019): asyncio.Lock-guarded dict, agent-scoped session keys, 20-message rolling window, 500-cap LRU eviction. Protocol shape is DB-migration-ready (A-003).
|
||||
- **Verify:** test_session covers create/append/window-trim/LRU-eviction/agent scoping
|
||||
|
||||
#### Task 2-1-03: Prompt library scaffolding
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-2-004
|
||||
- **Files:** `apps/ai-service/ai_service/prompts/coach.py`, `apps/ai-service/ai_service/prompts/tutor.py`, `apps/ai-service/ai_service/prompts/lab.py`, `apps/ai-service/ai_service/prompts/assessor.py`, `apps/ai-service/ai_service/prompts/proctor.py`, `apps/ai-service/ai_service/prompts/mentor.py`, `apps/ai-service/ai_service/prompts/__init__.py`
|
||||
- **Action:** Per-agent module with a versioned `SYSTEM_PROMPT` constant + `render_context(learner_context) -> dict` using `str.format_map` for learner-context injection (D-018: prompts are code, versioned in git). Initial drafts for all six; final personas land in Phases 3-5.
|
||||
- **Verify:** all six prompt modules import; render_context fills placeholders without KeyError
|
||||
|
||||
#### Task 2-1-04: Learner context corpus
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-2-004
|
||||
- **Files:** `apps/ai-service/ai_service/corpus/learner_context.py`, `apps/ai-service/ai_service/corpus/__init__.py`
|
||||
- **Action:** Pydantic-typed learner context (active stack, competencies, progress, recent artifacts) mirroring TS `packages/mock-data` IDs per D-021 convention (cross-referencing header comment, identical `comp-*`/`stack-*` ID strings).
|
||||
- **Verify:** context renders into prompt placeholders; IDs match packages/mock-data strings
|
||||
|
||||
### Wave 2: Composition layers (depends on Wave 1)
|
||||
|
||||
#### Task 2-2-01: Structured output defense
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-2-004
|
||||
- **Files:** `apps/ai-service/ai_service/agents/structured.py`, `apps/ai-service/tests/test_structured.py`
|
||||
- **Action:** 4-layer defense (D-020): (1) `response_format` request with auto-degrade on provider 400; (2) prompt-embedded JSON schema; (3) parse: strip code fences → first balanced JSON object; (4) single bounded retry with validation-error feedback. Returns pydantic-validated model or raises `StructuredOutputError`.
|
||||
- **Verify:** test_structured covers fenced/unfenced/invalid JSON, retry path, degrade path — all against mock provider
|
||||
|
||||
#### Task 2-2-02: Agent registry
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-2-004
|
||||
- **Files:** `apps/ai-service/ai_service/agents/registry.py`, `apps/ai-service/tests/test_registry.py`
|
||||
- **Action:** Explicit registry: name → agent factory map with `register(name, factory)` / `get(name)`; raises on unknown agent. Agents are registered in their own phases (P3-P5).
|
||||
- **Verify:** test_registry: register/get round-trip, unknown-agent error, duplicate registration error
|
||||
|
||||
#### Task 2-2-03: Session + agent DI wiring into API layer
|
||||
- **Persona:** backend-engineer — **REQ:** REQ-2-004
|
||||
- **Files:** `apps/ai-service/ai_service/api/deps.py` (update), `apps/ai-service/ai_service/api/chat.py` (update)
|
||||
- **Action:** deps.py exposes SessionStore and provider singletons via DI. chat.py persists turn history through the session store (agent-scoped) and includes session ID in the meta event. API composes agents via DI — agents never import api/.
|
||||
- **Verify:** chat request appends to and replays windowed history; `pnpm ai:test` green
|
||||
|
||||
### Must-Haves (Phase 2)
|
||||
- [ ] BaseAgent unit tests pass (stub agent streams via mock provider)
|
||||
- [ ] Session store tested: create/append, 20-message window trim, 500-cap LRU eviction, agent-scoped keys
|
||||
- [ ] Structured output parsing tested against mock provider: fence-strip, first-balanced-object, invalid JSON, one bounded retry, response_format auto-degrade
|
||||
- [ ] Registry tested: register/get/unknown/duplicate
|
||||
- [ ] Six prompt modules render learner context without errors
|
||||
- [ ] Module boundaries hold: `agents/` never imports `api/`; `llm/` never imports `agents/` or `api/`
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: Coach + Tutor Agents
|
||||
|
||||
**Requirements:** REQ-2-005, REQ-2-006
|
||||
**Goal:** Both learner-facing conversational agents fully implemented with distinct personas, registered, routed through the chat streaming endpoint
|
||||
|
||||
### Wave 1: Agent implementations (parallel — no shared files)
|
||||
|
||||
#### Task 3-1-01: Coach agent
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-2-005
|
||||
- **Files:** `apps/ai-service/ai_service/prompts/coach.py` (finalize), `apps/ai-service/ai_service/agents/coach.py`, `apps/ai-service/tests/test_coach.py`
|
||||
- **Action:** Final Coach persona: pacing guidance, motivation, retrieval practice prompts; system prompt injects learner context (active stack, progress). `CoachAgent(BaseAgent)` streams replies. Mock provider scripts a distinct coach-voice response for tests.
|
||||
- **Verify:** test_coach: build_messages includes system prompt + windowed history; stream_reply yields deltas; on-persona content asserted against mock script
|
||||
|
||||
#### Task 3-1-02: Tutor agent
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-2-006
|
||||
- **Files:** `apps/ai-service/ai_service/prompts/tutor.py` (finalize), `apps/ai-service/ai_service/agents/tutor.py`, `apps/ai-service/tests/test_tutor.py`
|
||||
- **Action:** Final Tutor persona: concept delivery, Socratic questioning, worked examples. `TutorAgent(BaseAgent)` streams replies; mock scripts a distinct tutor-voice response.
|
||||
- **Verify:** test_tutor mirrors test_coach; Coach and Tutor mock outputs are observably distinct
|
||||
|
||||
### Wave 2: Registration + persona verification (depends on Wave 1)
|
||||
|
||||
#### Task 3-2-01: Register Coach + Tutor; document cloud probe
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-2-005, REQ-2-006
|
||||
- **Files:** `apps/ai-service/ai_service/agents/registry.py` (update), `apps/ai-service/tests/test_registry.py` (update), `apps/ai-service/README.md` (update)
|
||||
- **Action:** Register both agents in the explicit registry. Extend test_registry to assert both resolve. Document the manual ollama-cloud persona probe in README (curl commands with `AI_PROVIDER=ollama-cloud`): Coach and Tutor produce distinct on-persona responses; tests remain cloud-free.
|
||||
- **Verify:** `pnpm ai:test` green; manual probe against ollama-cloud shows distinct personas (documented, not automated)
|
||||
|
||||
### Wave 3: Chat endpoint agent routing (depends on Wave 2)
|
||||
|
||||
#### Task 3-3-01: Agent routing on /v1/chat/stream
|
||||
- **Persona:** backend-engineer — **REQ:** REQ-2-005, REQ-2-006
|
||||
- **Files:** `apps/ai-service/ai_service/api/chat.py` (update), `apps/ai-service/ai_service/api/deps.py` (update), `apps/ai-service/tests/api/test_chat_stream.py` (update)
|
||||
- **Action:** Chat request gains `agent` field (validated against the registry; unknown agent → 422). Endpoint resolves the agent via DI, persists to the agent-scoped session, meta event carries the agent name. No autonomous routing in v0.2 (A-007).
|
||||
- **Verify:** TestClient tests: `agent=coach` and `agent=tutor` route correctly, session scoped per agent, unknown agent rejected
|
||||
|
||||
### Must-Haves (Phase 3)
|
||||
- [ ] Both agents produce distinct, on-persona responses (mock-asserted; manual ollama-cloud probe documented in README)
|
||||
- [ ] Agent routing tested: coach/tutor resolve via registry; unknown agent returns 422
|
||||
- [ ] Both agents exposed end-to-end via `POST /v1/chat/stream` with agent-scoped session history
|
||||
- [ ] `pnpm ai:test` green; no cloud calls in tests
|
||||
|
||||
---
|
||||
|
||||
## Phase 4: Lab + Assessor Agents
|
||||
|
||||
**Requirements:** REQ-2-007, REQ-2-008
|
||||
**Goal:** Lab consumes simulated sandbox telemetry and streams in-flow feedback; Assessor applies rubrics to pre-baked artifacts and returns structured scores — both over mock engine inputs
|
||||
|
||||
### Wave 1: Mock engine inputs (parallel — no shared files)
|
||||
|
||||
#### Task 4-1-01: Simulated sandbox telemetry corpus
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-2-007
|
||||
- **Files:** `apps/ai-service/ai_service/corpus/telemetry.py`
|
||||
- **Action:** Pydantic-typed Lab telemetry scenarios: scripted build-session event streams (keystrokes, commits, test runs, errors, idle gaps) keyed by scenario ID, aligned with packages/mock-data IDs (D-021).
|
||||
- **Verify:** scenarios import, validate, and are addressable by ID
|
||||
|
||||
#### Task 4-1-02: Pre-baked artifacts + rubrics corpus
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-2-008
|
||||
- **Files:** `apps/ai-service/ai_service/corpus/artifacts.py`
|
||||
- **Action:** Pre-baked artifacts (code, design, simulation), assessment rubrics (criteria, levels, weights), and defense transcripts keyed by ID — the Assessor's mock inputs (real engines are v0.3+).
|
||||
- **Verify:** rubric/artifact/transcript fixtures validate; IDs align with TS mock data
|
||||
|
||||
#### Task 4-1-03: TS mock-data alignment for engine inputs
|
||||
- **Persona:** data-engineer — **REQ:** REQ-2-012
|
||||
- **Files:** `packages/mock-data/ai-scenarios.ts` (new), `packages/mock-data/index.ts` (update)
|
||||
- **Action:** Export scenario/artifact ID constants + display metadata used by the Phase 6 learner panels, mirroring `ai_service/corpus/` IDs exactly (D-021). Data-engineer owns the TS side; headers cross-reference the Python corpus.
|
||||
- **Verify:** `pnpm typecheck` passes; IDs string-equal to corpus IDs
|
||||
|
||||
### Wave 2: Agent implementations (depends on Wave 1)
|
||||
|
||||
#### Task 4-2-01: Lab agent
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-2-007
|
||||
- **Files:** `apps/ai-service/ai_service/prompts/lab.py` (finalize), `apps/ai-service/ai_service/agents/lab.py`, `apps/ai-service/tests/test_lab.py`
|
||||
- **Action:** `LabAgent(BaseAgent)` consumes a telemetry scenario, builds messages summarizing the event stream, streams concrete in-flow feedback (what happened, what to adjust, next step). No session chat — scenario-driven.
|
||||
- **Verify:** test_lab: given a mock scenario, feedback references scenario events (mock-scripted assertions)
|
||||
|
||||
#### Task 4-2-02: Assessor agent
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-2-008
|
||||
- **Files:** `apps/ai-service/ai_service/prompts/assessor.py` (finalize), `apps/ai-service/ai_service/agents/assessor.py`, `apps/ai-service/tests/test_assessor.py`
|
||||
- **Action:** `AssessorAgent(BaseAgent)` applies a rubric to an artifact + defense transcript via `structured_reply`, returning a pydantic-validated rubric score model (per-criterion scores, strengths, gaps, verdict).
|
||||
- **Verify:** test_assessor: structured output validates against the rubric model; failure modes exercise the 4-layer defense
|
||||
|
||||
### Wave 3: Registration (depends on Wave 2)
|
||||
|
||||
#### Task 4-3-01: Register Lab + Assessor
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-2-007, REQ-2-008
|
||||
- **Files:** `apps/ai-service/ai_service/agents/registry.py` (update), `apps/ai-service/tests/test_registry.py` (update)
|
||||
- **Action:** Register both agents; extend registry tests.
|
||||
- **Verify:** registry resolves coach/tutor/lab/assessor; `pnpm ai:test` green
|
||||
|
||||
### Wave 4: Endpoints (depends on Wave 3)
|
||||
|
||||
#### Task 4-4-01: Lab feedback endpoint
|
||||
- **Persona:** backend-engineer — **REQ:** REQ-2-007
|
||||
- **Files:** `apps/ai-service/ai_service/api/lab.py`, `apps/ai-service/tests/api/test_lab.py`
|
||||
- **Action:** `POST /v1/lab/feedback` with scenario ID → resolves corpus scenario + Lab agent → SSE stream using the D-016 envelope (meta names agent=lab). Unknown scenario → 404.
|
||||
- **Verify:** TestClient streams meta + deltas + done + `[DONE]`; unknown scenario 404
|
||||
|
||||
#### Task 4-4-02: Assessment evaluate endpoint
|
||||
- **Persona:** backend-engineer — **REQ:** REQ-2-008
|
||||
- **Files:** `apps/ai-service/ai_service/api/assessment.py`, `apps/ai-service/tests/api/test_assessment.py`
|
||||
- **Action:** `POST /v1/assessment/evaluate` with artifact ID → resolves corpus artifact/rubric/transcript + Assessor agent → JSON response (validated rubric score model). Unknown artifact → 404.
|
||||
- **Verify:** TestClient returns validated rubric JSON; unknown artifact 404; `pnpm ai:test` green
|
||||
|
||||
### Must-Haves (Phase 4)
|
||||
- [ ] Lab produces scenario-relevant in-flow feedback for mock telemetry scenarios (mock provider, tested)
|
||||
- [ ] Assessor returns structured rubric scores (pydantic-validated JSON) for pre-baked artifacts/transcripts
|
||||
- [ ] `POST /v1/lab/feedback` streams (meta → deltas → done → `[DONE]`); `POST /v1/assessment/evaluate` returns validated JSON
|
||||
- [ ] Unknown scenario/artifact IDs return 404
|
||||
- [ ] Corpus IDs align with packages/mock-data (D-021); `pnpm ai:test` and `pnpm typecheck` green
|
||||
|
||||
---
|
||||
|
||||
## Phase 5: Proctor + Mentor Agents
|
||||
|
||||
**Requirements:** REQ-2-009, REQ-2-010
|
||||
**Goal:** Proctor classifies integrity signals with coaching interventions from mock telemetry; Mentor generates long-horizon career narrative; both exposed via endpoints
|
||||
|
||||
### Wave 1: Proctor scenarios + Mentor agent (parallel — no shared files)
|
||||
|
||||
#### Task 5-1-01: Proctor telemetry scenarios
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-2-009
|
||||
- **Files:** `apps/ai-service/ai_service/corpus/telemetry.py` (update)
|
||||
- **Action:** Add proctor scenarios: tab switches, idle time, paste events, focus loss — scripted integrity-relevant event sets keyed by scenario ID.
|
||||
- **Verify:** proctor scenarios validate; distinguishable from lab scenarios by type
|
||||
|
||||
#### Task 5-1-02: Mentor agent (+ registration)
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-2-010
|
||||
- **Files:** `apps/ai-service/ai_service/prompts/mentor.py` (finalize), `apps/ai-service/ai_service/agents/mentor.py`, `apps/ai-service/tests/test_mentor.py`, `apps/ai-service/ai_service/agents/registry.py` (update)
|
||||
- **Action:** `MentorAgent(BaseAgent)`: long-horizon career narrative — trajectory story, competency-stack progression guidance, market positioning — streaming, session-backed. Registered centrally in `registry.py` (single registration pattern, G-4).
|
||||
- **Verify:** test_mentor: narrative references learner context (mock-scripted); registry resolves mentor
|
||||
|
||||
### Wave 2: Proctor agent + Mentor endpoint (depends on Wave 1)
|
||||
|
||||
#### Task 5-2-01: Proctor agent (+ registration)
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-2-009
|
||||
- **Files:** `apps/ai-service/ai_service/prompts/proctor.py` (finalize), `apps/ai-service/ai_service/agents/proctor.py`, `apps/ai-service/tests/test_proctor.py`, `apps/ai-service/ai_service/agents/registry.py` (update)
|
||||
- **Action:** `ProctorAgent(BaseAgent)`: consumes proctor scenario → `structured_reply` returns pydantic-validated signal classification (severity, signal type) + recommended coaching intervention (supportive, not punitive). Registered centrally in `registry.py` (single registration pattern, G-4).
|
||||
- **Verify:** test_proctor: classified signals + interventions validate for each mock scenario; registry resolves all six agents
|
||||
|
||||
#### Task 5-2-02: Mentor narrative endpoint
|
||||
- **Persona:** backend-engineer — **REQ:** REQ-2-010
|
||||
- **Files:** `apps/ai-service/ai_service/api/mentor.py`, `apps/ai-service/tests/api/test_mentor.py`
|
||||
- **Action:** `POST /v1/mentor/narrative` → Mentor agent → SSE stream with D-016 envelope, session-backed.
|
||||
- **Verify:** TestClient streams meta (agent=mentor) → deltas → done → `[DONE]`
|
||||
|
||||
### Wave 3: Proctor endpoint (depends on Wave 2)
|
||||
|
||||
#### Task 5-3-01: Proctor signals endpoint
|
||||
- **Persona:** backend-engineer — **REQ:** REQ-2-009
|
||||
- **Files:** `apps/ai-service/ai_service/api/proctor.py`, `apps/ai-service/tests/api/test_proctor.py`
|
||||
- **Action:** `POST /v1/proctor/signals` with scenario ID → resolves corpus scenario + Proctor agent → JSON response (validated signals + interventions). Unknown scenario → 404.
|
||||
- **Verify:** TestClient returns classified signals JSON; `pnpm ai:test` green — full suite (all six agents registered)
|
||||
|
||||
### Must-Haves (Phase 5)
|
||||
- [ ] Proctor produces classified signals with recommended coaching interventions for each mock scenario (structured JSON, validated)
|
||||
- [ ] Mentor produces coherent long-horizon career narrative (streaming, session-backed)
|
||||
- [ ] `POST /v1/proctor/signals` returns validated JSON; `POST /v1/mentor/narrative` streams
|
||||
- [ ] Registry resolves all six agents; full ai-service test suite green, cloud-free
|
||||
|
||||
---
|
||||
|
||||
## Phase 6: Learner Surface Integration
|
||||
|
||||
**Requirements:** REQ-2-011, REQ-2-012
|
||||
**Goal:** v0.1 learner surfaces wired to the real ai-service: streaming chat with agent switcher, Lab/Assessor/Proctor/Mentor outputs surfaced, error/loading states, build + typecheck green
|
||||
**Note (G-2):** End-to-end verification may run with `AI_PROVIDER=mock` as a fallback — the requirement is the real ai-service over HTTP (not canned client-side responses); provider choice is service-internal. This prevents an ollama-cloud outage from blocking P6 verification. Cloud persona probes remain separate (Task 3-2-01).
|
||||
|
||||
### Wave 1: Client plumbing + primitives (parallel — no shared files)
|
||||
|
||||
#### Task 6-1-01: useChatStream hook
|
||||
- **Persona:** frontend-engineer — **REQ:** REQ-2-011
|
||||
- **Files:** `apps/web/hooks/use-chat-stream.ts`, `apps/web/.env.example` (update)
|
||||
- **Action:** `useChatStream(agent)` hook: `fetch` POST to `${NEXT_PUBLIC_AI_SERVICE_URL}/v1/chat/stream` (default `http://localhost:8420`, A-002 — no API-route proxy); consumes `ReadableStream` with byte buffering, frame split on `\n\n`, joined `data:` lines; **ignores frames containing no `data:` lines (sse-starlette `: ping` keep-alive comment frames)** — TestClient streams are too short to surface pings, but real cloud delta gaps emit them (G-1); handles meta / delta / done / error events and `[DONE]` sentinel; idempotent `AbortController.abort()` in effect cleanup; exposes `{messages, isStreaming, error, send, retry, abort}`. `.env.example` gains `NEXT_PUBLIC_AI_SERVICE_URL`.
|
||||
- **Verify:** hook unit-tested or exercised via the chat UI; unmount mid-stream aborts cleanly (no state updates after unmount)
|
||||
|
||||
#### Task 6-1-02: Agent switcher + streaming primitives
|
||||
- **Persona:** design-system-engineer — **REQ:** REQ-2-011
|
||||
- **Files:** `packages/ui/src/primitives/agent-switcher.tsx`, `packages/ui/src/primitives/stream-status.tsx`, `packages/ui/src/primitives/toast.tsx`, `packages/ui/src/primitives/index.ts` (update), `packages/ui/src/index.ts` (update)
|
||||
- **Action:** Token-driven primitives: AgentSwitcher (segmented coach/tutor control with active state), StreamStatus (idle/streaming/error indicator), Toast with error variant. Dark mode + WCAG AA contrast; exported from `@nextcraft/ui`.
|
||||
- **Verify:** primitives import from `@nextcraft/ui`; storybook stories render (dark + light)
|
||||
|
||||
### Wave 2: Chat rewrite + Byte viewer panel (depends on Wave 1)
|
||||
|
||||
#### Task 6-2-01: Real streaming learner chat
|
||||
- **Persona:** frontend-engineer — **REQ:** REQ-2-011
|
||||
- **Files:** `apps/web/components/learner/ai-tutor-chat.tsx` (rewrite), `apps/web/components/learner/agent-switcher.tsx` (composition wrapper)
|
||||
- **Action:** Replace the canned `aiTutorResponses` behavior with `useChatStream`: agent switcher (Coach/Tutor per A-007), token-by-token rendering, streaming cursor + loading state, error state with retry button when ai-service is down (A-010), suggested-action chips from the meta event. Seed welcome message stays static.
|
||||
- **Verify:** with ai-service running, messages stream visibly token-by-token; with ai-service stopped, error state + retry appears (no crash, no console errors)
|
||||
|
||||
#### Task 6-2-02: Byte viewer Tutor panel
|
||||
- **Persona:** frontend-engineer — **REQ:** REQ-2-011 (agent routing: byte viewer always uses Tutor, A-007)
|
||||
- **Files:** `apps/web/app/(learner)/learn/[competencyId]/page.tsx` (update), `apps/web/components/learner/byte-tutor-panel.tsx` (new)
|
||||
- **Action:** Byte viewer gains a Tutor explanation panel (byte viewer always uses Tutor, A-007): "Explain this byte" streams a Socratic concept walkthrough for the current competency via useChatStream (agent fixed to tutor).
|
||||
- **Verify:** on a byte page, the panel streams a Tutor explanation; error state when service down
|
||||
|
||||
### Wave 3: Lab / Assessment / Mentor panels (depends on Wave 1 hook; parallel — no shared files)
|
||||
|
||||
#### Task 6-3-01: Sandbox Lab feedback panel
|
||||
- **Persona:** frontend-engineer — **REQ:** REQ-2-012
|
||||
- **Files:** `apps/web/app/(learner)/build/[competencyId]/page.tsx` (update), `apps/web/components/learner/lab-feedback-panel.tsx` (new)
|
||||
- **Action:** Sandbox telemetry sidebar gains a Lab feedback panel: posts the scenario ID (from `packages/mock-data/ai-scenarios`) to `/v1/lab/feedback`, streams in-flow feedback into the panel; loading + error states.
|
||||
- **Verify:** sandbox page streams Lab feedback for the mock scenario; error state when service down
|
||||
|
||||
#### Task 6-3-02: Assessment Assessor + Proctor surfaces
|
||||
- **Persona:** frontend-engineer — **REQ:** REQ-2-012
|
||||
- **Files:** `apps/web/app/(learner)/defend/[competencyId]/page.tsx` (update), `apps/web/components/learner/assessor-results-panel.tsx` (new), `apps/web/components/learner/proctor-banner.tsx` (new)
|
||||
- **Action:** Assessment mockup: AI reviewer panel calls `/v1/assessment/evaluate` with the artifact ID and renders the structured rubric scores (per-criterion bars, strengths, gaps, verdict); a Proctor integrity banner surfaces `/v1/proctor/signals` classifications with coaching tone. Loading skeletons + error states.
|
||||
- **Verify:** defend page renders real Assessor rubric output + Proctor banner; error states when service down
|
||||
|
||||
#### Task 6-3-03: Dashboard Mentor panel
|
||||
- **Persona:** frontend-engineer — **REQ:** REQ-2-012
|
||||
- **Files:** `apps/web/app/(learner)/dashboard/page.tsx` (update), `apps/web/components/learner/mentor-panel.tsx` (new)
|
||||
- **Action:** Learner dashboard gains a Mentor panel: streams career narrative from `/v1/mentor/narrative` (learner progress context), with regenerate button, loading + error states. Sits alongside the existing AI tutor chat.
|
||||
- **Verify:** dashboard shows streaming Mentor narrative; error state when service down
|
||||
|
||||
### Must-Haves (Phase 6)
|
||||
- [ ] With ai-service running: learner chat at http://localhost:3000/dashboard streams real responses token-by-token (mock provider fallback allowed per G-2 — service over HTTP is the requirement)
|
||||
- [ ] Agent switcher flips Coach ↔ Tutor and the response persona changes accordingly
|
||||
- [ ] Hook tolerates keep-alive comment frames (`: ping`, no data lines) during live streams (G-1)
|
||||
- [ ] With ai-service stopped: all chat/panels show error states with retry — no crashes, no unhandled promise rejections, no console errors
|
||||
- [ ] Byte viewer, sandbox, and assessment mockups surface Tutor/Lab/Assessor/Proctor outputs; dashboard shows Mentor narrative
|
||||
- [ ] Unmounting/navigating mid-stream aborts cleanly (no post-unmount state updates)
|
||||
- [ ] `pnpm build` and `pnpm typecheck` pass; `pnpm ai:test` still green
|
||||
|
||||
---
|
||||
|
||||
## Phase 7: Final Review + Ship (no planned tasks)
|
||||
|
||||
Orchestrated by the SHIP stage, not this plan: multi-persona code review (correctness, testing, secrets hygiene — key absent from code/logs/commits/errors, localhost-only CORS, no PII in prompts), project health audit (reconstruction test, .ciagent/ discipline, branch/commit hygiene), then merge milestone → main, tag v0.2.0, create Gitea release, mark all 12 v0.2 requirements complete.
|
||||
|
||||
**Release-note honesty (G-5):** the v0.2.0 release note must explicitly state that Lab/Assessor/Proctor operate on mock engine inputs (real engines are v0.3+), per D-015.
|
||||
|
||||
**Dead-code disposal (G-5):** review must dispose of `aiTutorResponses` (packages/mock-data/ai-tutor-responses.ts) — its only consumer is rewritten in Task 6-2-01; remove the export or mark it deprecated.
|
||||
|
||||
---
|
||||
|
||||
## User-Facing Surface
|
||||
|
||||
1. **Voice defense with a provider badge** (`/defend/[competencyId]`): the learner's spoken answers upload as audio and are transcribed server-side when `AI_VOICE_PROVIDER=openai-audio` is configured; examiner questions play as server TTS audio; the mic control shows an honest badge — `server voice`, `browser voice`, or `mock` — derived from the provider descriptor (`VoiceDescriptor.mode`), and degrades visibly (browser/mock fallback) with keys absent.
|
||||
2. **Identity enrollment flow** (new `/enroll` learner route + marketplace surfaces): submit verification → pending state → verified/rejected state; verified learners proceed to variants/sandboxes/defense; unverified learners hitting gated routes see a structured verify-CTA (403 payload rendered as an actionable prompt, not a dead error). Mock verdicts are labeled `mock` everywhere they surface (A-304 honesty).
|
||||
3. **Environment-typed build flows** (`/build/[competencyId]`): design competencies open a design environment (SVG/HTML/schematic artifact starter files, Run = validator harness), simulation competencies open a simulation environment (benchmark script + dataset starter files, Run = bounded harness execution); the Run/Test buttons use the variant's real `test_command` instead of hardcoded pytest; the surface, file tree, editor, read-only output panel (CUT-2), and telemetry pulse are unchanged across kinds.
|
||||
4. **Invisible durability**: mid-connection kill of a build session loses nothing on reconnect (seq-ack protocol) — no visible UI, proven by tests.
|
||||
The primary user-facing surface is the **learner dashboard and learning flow** at `http://localhost:3000`, backed by the real ai-service at `http://localhost:8420`:
|
||||
|
||||
- `/dashboard` — AI tutor chat (Coach/Tutor switcher, streaming) + Mentor career-narrative panel
|
||||
- `/learn/[competencyId]` — byte viewer with streaming Tutor explanations
|
||||
- `/build/[competencyId]` — sandbox with Lab in-flow feedback panel (mock telemetry)
|
||||
- `/defend/[competencyId]` — assessment with live Assessor rubric scores + Proctor integrity banner
|
||||
|
||||
The marketplace, employer, and admin surfaces are unchanged from v0.1.
|
||||
|
||||
## Happy Path
|
||||
|
||||
**Defense with real voice:** learner opens `/defend/cmp-*` → DefenseSession starts → examiner question streams (SSE) → learner speaks → MediaRecorder captures webm → POST `/v1/defense/{id}/answer` (multipart audio) → server strips codec param (`webm;codecs=opus` → `webm`), enforces ≤10MB, calls `OpenAIAudioProvider.transcribe` → transcript turn stored (STT latency recorded) → learner clicks the speaker icon on an examiner turn → GET `/v1/defense/{id}/audio/{turn_id}` → server TTS bytes stream back with correct media_type → verdict + integrity signals render unchanged.
|
||||
|
||||
**Identity-gated build:** learner completes `/enroll` (submit → mock provider verdict `verified`, age band 18+) → opens a design competency → POST `/v1/variants` passes the identity gate (verified ≥16) → variant carries `environment: "design"` + `test_command` → sandbox created (allowlist ✓ → identity ✓ → rate cap ✓) → starter files written → Run executes the validator harness in the namespace sandbox → telemetry streams with seq-acks → grade digest renders.
|
||||
1. Learner opens `/dashboard` → chat shows welcome message; meta event confirms coach/model in the stream
|
||||
2. Learner types "I'm stuck on multi-agent communication" → reply streams token-by-token with pacing guidance + a retrieval-practice prompt
|
||||
3. Learner switches to **Tutor** → asks the same question → gets a Socratic concept walkthrough instead
|
||||
4. Learner opens a byte tutorial → Tutor panel streams an explanation of the current competency
|
||||
5. Learner opens the build sandbox → Lab panel streams feedback on the simulated telemetry scenario
|
||||
6. Learner opens the defense mockup → Assessor panel shows structured rubric scores; Proctor banner shows integrity signals in coaching tone
|
||||
7. Back on `/dashboard`, the Mentor panel streams a career narrative tied to the learner's progress
|
||||
8. Learner kills ai-service (or it crashes) → next message shows an inline error state with **Retry**; restarting the service and retrying resumes streaming
|
||||
|
||||
## UX Acceptance Criteria
|
||||
|
||||
1. The mic control badge always tells the truth about which voice path is live (server/browser/mock) — never claims server when mock is wired.
|
||||
2. Unverified/under-age callers on gated routes get a 403 with an actionable verify-CTA payload (rendered as a prompt with a link to enrollment) — never a bare JSON error in the UI.
|
||||
3. Mock identity verdicts are visibly labeled `mock` in every surface that shows verification state.
|
||||
4. Design/sim environments are indistinguishable from build environments in surface mechanics (file tree, editor, Run/Test, output panel) — only starter contents and the Run command differ; no route changes, no new navigation.
|
||||
5. Run/Test buttons reflect the variant's `test_command` (no hardcoded pytest on a design competency).
|
||||
6. No durable state is written inside the repo (state in `~/.nextcraft/`; tests in tmp dirs).
|
||||
7. All existing accessibility baselines hold (WCAG AA contrast on new badge/CTA states).
|
||||
|
||||
---
|
||||
|
||||
## Phase 1: Seq-Lease + Replay Margin (REQ-5-007)
|
||||
|
||||
**Goal:** Close the one-line replay-margin/ACK gap (P07-documented): the capture agent requeues only `_last_sent` on detected disconnect while N frames may be in TCP flight — frames 1..N-1 are lost. Server seq-acks close it.
|
||||
|
||||
### Wave 1-1: ingest ack emission
|
||||
|
||||
- **Task 1-1-01** (sandbox-engineer): `telemetry/ingest.py` — emit `{"type":"seq_ack","seq":N}` after every successful `store.append` in `IngestSession._append` (N = durable `latest_seq` snapshot already computed; O(1), advisory — gap detection stays authoritative; G-3 flood semantics untouched). Unit test: appended seq N → ack frame N observed on the WS (TestClient portal pattern, `tests/api/test_telemetry_ingest.py`).
|
||||
- Files: `apps/ai-service/ai_service/telemetry/ingest.py`, `apps/ai-service/tests/telemetry/`
|
||||
- **MH-1a**: a test proves each successful append emits an ack carrying the post-append `latest_seq`.
|
||||
|
||||
### Wave 1-2: agent ack consumption + spool trim
|
||||
|
||||
- **Task 1-2-01** (sandbox-engineer): `scripts/sandbox-agent.py` — supervisor loop (currently discards all non-close frames) parses text frames; on `type == "seq_ack"` trims every spool/pending line with `seq <= ack` under `_emit_lock` via atomic `Spool.rewrite`; clears `_last_sent` if its seq ≤ ack; tolerates any frame interleaving (acks/gap_warning/rejected); keepalive pings (binary) unaffected. **Stdlib-only (AST-pinned).**
|
||||
- **Task 1-2-02** (sandbox-engineer): add an explicit spool bound (max lines, e.g. 4096 — documented; D-R07 correction: no cap existed) — oldest-beyond-bound dropped with a counter; **G-14: overflow is by-design gap creation — unit test proves dropped-counter > 0 → replayed trace exhibits gaps → grader/gap path marks it ungradable (never a silently-truncated-but-gradable trace); document the worst-case arithmetic (~64KB diff cap × 4096 lines ≈ 256MB, under but HALF the 512MB G-2 budget — the spool lives inside the swept workdir)**; document G-2+flood-cap as the outer bound.
|
||||
- Files: `apps/ai-service/scripts/sandbox-agent.py`, `apps/ai-service/tests/sandbox/test_sandbox_agent.py`
|
||||
- **MH-1b**: unit test — agent with a scripted WS that acks mid-drain trims its spool to `seq > ack` exactly (no over-trim, no under-trim), stays within the explicit bound, and overflow drops create honest gaps (ungradable, G-14).
|
||||
- **MH-1c**: `replay_margin()` behavior after acks: requeue window is bounded by unacked in-flight only (repeated reconnect/ack cycles never lose or duplicate a spooled line).
|
||||
|
||||
### Wave 1-3: mid-burst regression test (the real proof)
|
||||
|
||||
- **Task 1-3-01** (sandbox-engineer): extend `tests/telemetry/test_durability.py` with `test_midburst_disconnect_loses_nothing` — real uvicorn + KillableProxy; sever the connection **immediately after a rapid multi-frame send, WITHOUT waiting for server observation** (the exact scenario the P07 de-flake documented as uncovered); revive; assert every emitted seq stored exactly once, in order; assert spool trimmed to ≤ ack margin; keep `_await_events`/portal patterns (deterministic, no socket surgery).
|
||||
- Files: `apps/ai-service/tests/telemetry/test_durability.py`
|
||||
- **MH-1d**: the mid-burst test passes repeatedly (≥3 consecutive runs) with zero loss/dup outside the acked margin.
|
||||
|
||||
**Verification strategy P1:** `pnpm ai:test` (413+green), ruff, no web/TS changes, no settings changes. Existing durability + reconnect-flush suites stay green.
|
||||
|
||||
**Risks:** mid-burst determinism (mitigated: ack protocol is the fix; wait for server-side observation of the ack itself); `_emit_lock` reentrancy from supervisor thread (trim under the same lock as flush); frame-order tolerance.
|
||||
|
||||
---
|
||||
|
||||
## Phase 2: Real Server Voice (REQ-5-001, REQ-5-002)
|
||||
|
||||
**Goal:** The `openai-audio` VoiceProvider — server STT/TTS for voice defense end-to-end when keys exist; mock/browser unchanged.
|
||||
|
||||
### Wave 2-1: settings + factory + provider
|
||||
|
||||
- **Task 2-1-01** (backend-engineer): `config.py` — add `voice_base_url: str = ""`, `voice_api_key: str = ""` (never logged), `voice_stt_model: str = "whisper-1"`, `voice_tts_model: str = "tts-1"`, `voice_tts_voice: str = "alloy"`, `voice_tts_format: Literal["mp3","wav","opus"] = "mp3"` (G-16: enum, not free string — it feeds a Content-Type), `voice_max_audio_mb: int = 10`; `.env.example` documents them (keys never in code/commits); unknown format value → fall back to default with a loud log (consistent with G-11).
|
||||
- **Task 2-1-02** (backend-engineer): `voice/factory.py` — `voice_provider_from_settings(settings, http_client)` (signature change); `openai-audio` branch constructs `OpenAIAudioProvider`; rejects selection when `voice_base_url`/`voice_api_key` empty with an actionable error for direct callers; **G-11: `main.py` lifespan catches the factory error, logs loudly (naming the missing vars), and falls back to the mock provider — a typo'd env in an unattended deploy must never crash the boot (D-039); descriptor then honestly reads `mock`**; lifespan passes `app.state.http_client` (state-injection preserved).
|
||||
- **Task 2-1-03** (voice-engineer): `voice/openai_audio.py` — `transcribe()`: multipart `POST {base}/audio/transcriptions` (`file` + `model`, response_format=json → `TranscriptSegment`); `synthesize()`: streaming `POST /audio/speech` (JSON body `model/input/voice/response_format`, raw byte chunks); `descriptor = VoiceDescriptor(mode="server", sr_available=True, tts_available=True, hint=...)`; errors sanitized with key redaction (mirror `llm/openai_compat.py:_sanitize`); reuses the shared httpx client (D-017; read=300s).
|
||||
- Files: `apps/ai-service/ai_service/config.py`, `voice/factory.py`, `voice/openai_audio.py`, `main.py`, `.env.example`
|
||||
- **MH-2a**: MockTransport byte-contract tests — STT: multipart fields + response parse → `TranscriptSegment`; TTS: JSON body + byte stream → concatenated chunks; failure pins 413/400/429/timeout → sanitized errors, NO key leak (pinned).
|
||||
- **MH-2b**: factory: `openai-audio` selected + configured → server-mode provider; selected + unconfigured → actionable rejection for direct callers AND app boot survives with mock fallback + loud log naming the fix (G-11); mock/browser unchanged; provider always carries a `descriptor` (a-15 — a missing descriptor would badge the server path as mock). Invert the v0.4 rejection test (`test_real_server_stt_tts_rejected_as_v04_seam`).
|
||||
|
||||
### Wave 2-2: defense route fixes + audio upload
|
||||
|
||||
- **Task 2-2-01** (voice-engineer): `api/defense.py` — fmt derivation strips codec params (`"webm;codecs=opus"` → `webm`; else real STT 400s); enforce `voice_max_audio_mb` BEFORE provider call (422 empty / 413 oversize; a-9: Content-Length fast path before buffering); TTS route media_type mapped from the `voice_tts_format` enum (G-16); descriptor now comes from the provider (mode=server flows to the client untouched).
|
||||
- **Task 2-2-02** (frontend-engineer): `engine-client.ts` — `answerDefense` audio variant (FormData: blob + filename + content-type); `defense-session.tsx` — POST the recorded blob instead of discarding it; **G-12: recording bound — auto-stop at a max duration (default 180s) with a visible timer, `recorder.start(timeslice)` for observable size (a-13); a 413 response renders as an honest "answer too long — re-record" prompt, never silent loss**; provider badge from descriptor (`server voice`/`browser voice`/`mock`) with visible degradation states.
|
||||
- Files: `apps/ai-service/ai_service/api/defense.py`, `apps/web/lib/engine-client.ts`, `apps/web/components/learner/defense-session.tsx`, `apps/web/tests/`
|
||||
- **MH-2c**: API test — webm;codecs=opus content-type reaches the provider as clean `webm`; oversize audio → 413 without provider call; empty → 422 (existing).
|
||||
- **MH-2d**: web tests — audio POST path builds correct FormData; badge reflects descriptor mode; auto-stop fires at the bound; 413 renders the re-record prompt (G-12).
|
||||
|
||||
**Verification strategy P2:** `pnpm ai:test`, ruff, `pnpm typecheck`, web tests, `pnpm build` (static export still emits). Manual cloud probe recipe documented in `.env.example` comments (never CI). Defense flow tests stay mock-only (cloud-free rule).
|
||||
|
||||
**Risks:** full audio bytes buffered in memory (bounded by the 10MB guard); TTS `input` ≤4096 chars (examiner questions are short — enforced with a guard + truncation error); factory signature change touches main.py lifespan (state-injection pattern preserved).
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: Identity + Age-Gating (REQ-5-003, REQ-5-004)
|
||||
|
||||
**Goal:** Identity verification backend behind a provider protocol (mock-first); backend-enforced age gates composed with G-5.
|
||||
|
||||
### Wave 3-1: identity module core
|
||||
|
||||
- **Task 3-1-01** (identity-engineer): `ai_service/identity/base.py` — `IdentityProvider` protocol: `submit(learner_id, submission) -> submission_id`, `poll(submission_id) -> verdict {status, age_band, provider, mock, refs}`; `mock.py` — deterministic mock (approve-on-policy: age band from a scripted DOB field, reject scripted-bad); verdicts carry `mock: true` marker (A-304).
|
||||
- **Task 3-1-02** (identity-engineer): `ai_service/identity/store.py` — 5th D-027 store, DefenseStore pattern (WAL, `foreign_keys=ON`, portable columns, `@validates`): `identity_record` table — `id` (submission id, minted once), `learner_id` (indexed), `status: pending|verified|rejected`, `provider`, `provider_verdict` (JSON, mock-marked), `age_band` (derived `16-17`|`18+`, NEVER raw DOB), `document_refs` (JSON refs — raw documents NEVER stored), `submitted_at`, `verified_at`; insert-only + latest-per-learner lookup.
|
||||
- **Task 3-1-03** (identity-engineer): `main.py` lifespan — `app.state.identity_store` + `app.state.identity_provider` (state-injection overrides preserved); `config.py` — `identity_provider: str = "mock"`, identity store rides the same db_path.
|
||||
- **MH-3a**: store tests — insert/poll/latest/verdict provenance; constraints fire on invalid bands; same SQLite file (additive table, D-027 family).
|
||||
|
||||
### Wave 3-2: API surface + gates
|
||||
|
||||
- **Task 3-2-01** (identity-engineer): `api/identity.py` — `/v1/identity/submit` (submission → pending; G-13: one active pending per learner — resubmit while pending → 409 echoing the pending state; per-learner submit rate cap → 429), `/v1/identity/status/{learner_id}` (latest record + mock marker), `/v1/identity/verify/{submission_id}` (poll provider → verified/rejected transition); router mounted in main.py.
|
||||
- **Task 3-2-02** (identity-engineer): gate dependencies — `require_verified_age(min_age)` FastAPI dependencies; **composition order binding (D-043)**: G-5 allowlist (403 pilot guard) → identity verdict (403 + verify-CTA payload `{reason, min_age, current_status, verify_cta}`) → rate caps (429). Apply: school 16+ on variant generation (`api/variants.py`), sandbox create (`api/sandboxes.py`), defense start (`api/defense.py`); marketplace 18+ via `require_verified_adult` on ONE minimal gated route (`POST /v1/marketplace/apply` — G-18: honest stub; passes the gate composition then returns 501 with `stub: true` + mock-verdict markers, never a fabricated "applied" outcome).
|
||||
- **Task 3-2-03** (identity-engineer, G-9): conftest `verified_pilot` fixture — seeds an identity record (mock provider, band 18+ or 16+) for test learner ids + allowlist coverage in test settings; MUST land in the same wave as the gates or the existing variant/sandbox/defense suites 403 en masse (those routes are ungated today).
|
||||
- **Task 3-2-04** (security-auditor, phase-specific): PII review — sentinel scrub test: submit identity with sentinel PII strings → assert they appear in NO log record (caplog) and NO stored raw form (store inspection); API responses expose verdict + mock marker only.
|
||||
- Files: `apps/ai-service/ai_service/identity/**`, `api/identity.py`, `api/variants.py`, `api/sandboxes.py`, `api/defense.py`, `main.py`, `config.py`, `tests/identity/`, `tests/api/`
|
||||
- **MH-3b**: gate tests — verified 18+ passes all school gates; 16-17 passes school gates but 403s the marketplace route (honest stub response beyond the gate, G-18); unverified → 403 with verify-CTA payload; under-16 → 403 everywhere gated; allowlist rejection (403) still fires FIRST for non-pilot learner ids; identity submit: pending-resubmit → 409, rate cap → 429 (G-13); **all pre-existing variant/sandbox/defense API suites remain green under the `verified_pilot` fixture (G-9)**.
|
||||
- **MH-3c**: caplog sentinel test green (PII never logged/stored).
|
||||
|
||||
### Wave 3-3: web enrollment flow
|
||||
|
||||
- **Task 3-3-01** (frontend-engineer): `/enroll` route — submit → pending → verified/rejected states (honest, mock-labeled); engine-client identity functions; **G-10: client 403 discrimination — `verify_cta` present in the 403 payload → new `VerifyRequiredError` `{reason, min_age, current_status, verify_cta}`; allowlist detail → existing `NotAllowlistedError` (today engine-client.ts collapses every 403 into NotAllowlistedError — a verify-CTA would render as an allowlist lie)**; gated-route 403 CTA rendered as actionable prompt (link to `/enroll`); dashboard learner age badge reflects verified state.
|
||||
- Files: `apps/web/app/(learner)/enroll/`, `apps/web/lib/engine-client.ts`, `apps/web/components/`, `packages/types/`
|
||||
- **MH-3d**: web tests — identity client functions; CTA payload shape; enrollment states render; **403 discrimination: verify-CTA → VerifyRequiredError, allowlist detail → NotAllowlistedError (G-10)**. `pnpm build` emits the new route.
|
||||
- **MH-3e** (CUT-3): identity flow test via `TestClient` against real `create_app` (routers mounted, gates composed, real stores, mock providers) — unverified learner → variants POST → 403 verify-CTA → submit + verify (mock) → variants POST 200. No uvicorn harness (identity is plain JSON; the real-server harness stays where transport matters — P1/P4).
|
||||
|
||||
**Verification strategy P3:** `pnpm ai:test`, ruff, typecheck, web tests, `pnpm build`. PII caplog test is release-blocking (security-auditor sign-off).
|
||||
|
||||
**Risks:** self-asserted `learner_id` trust level (documented as pilot-scale — same as G-5 today; real auth is post-v0.5); shared-SQLite additive table (safe); mock-verdict honesty must ride every response (pinned).
|
||||
|
||||
---
|
||||
|
||||
## Phase 4: Design/Sim Environments (REQ-5-005, REQ-5-006)
|
||||
|
||||
**Goal:** Environment kinds at the template/variant layer; per-kind starter contents + exec policy; kind flows through telemetry; digest untouched.
|
||||
|
||||
### Wave 4-1: registry + wire types
|
||||
|
||||
- **Task 4-1-01** (sandbox-engineer): `variants/templates.py` — `TaskTemplate.environment: Literal["build","design","simulation"] = "build"` + per-kind `starter_files` + `harness_command`/`test_command` policy fields; **G-15: command fields validated at definition time (Python, where shlex exists) — must roundtrip `shlex.split` → whitespace-join → `shlex.split` identically (no quotes/globs/metachars; violation is a template-authoring bug caught in tests)**; add one design template (stack-designer c001/c002 — already sanctioned) + one simulation template (stack-science or stack-operator); `variants/generator.py` + `store.py` carry the field through `VariantRecord`.
|
||||
- **Task 4-1-02** (frontend-engineer): `api/variants.py` — `VariantResponse` gains `environment` + `test_command` (closes the dead-field gap); TS `packages/types/variants.ts` + engine-client types sync (dual-schema rule: both places, same change).
|
||||
- **MH-4a**: variant tests — design/sim templates generate kind-tagged variants with correct starter files + commands; command fields roundtrip the shlex validator (G-15); wire response carries both fields (a-11: required on the wire, TS required-field parity); TS types match Python field-for-field.
|
||||
|
||||
### Wave 4-2: exec command policy
|
||||
|
||||
- **Task 4-2-01** (sandbox-engineer + security-auditor): `api/sandboxes.py` exec route — per-kind command policy: **G-15: EXACT argv-token matching** against {template-declared harness/test argv[0]} ∪ a small generic file/nav set — never prefix/substring (trivially bypassed via flags/`-c` passthrough); `sh -c` passthrough DISALLOWED for design/sim kinds (the gaming vector: faking build-style test cycles into a kind-agnostic digest); violation → 422 naming the allowed set; policy table is code (reviewable, versioned); build-kind flows do not regress (existing tests green).
|
||||
- **MH-4b**: exec tests — design kind rejects pytest-style arbitrary commands not in policy (422); simulation kind accepts its declared harness; build kind flows unchanged.
|
||||
|
||||
### Wave 4-3: learner surface kind-awareness
|
||||
|
||||
- **Task 4-3-01** (frontend-engineer): `use-sandbox-session.ts` — `test()` uses `variant.test_command`; **G-15: TS splits on whitespace ONLY (no shlex in the browser — safe because templates validated quote-free at authoring, Task 4-1-01)**; `build-surface.tsx` — RunControls commands from the variant; starter-file materialization loop already kind-agnostic (verify against design/sim starter sets); honest busy/denied/error states carry over.
|
||||
- **MH-4c**: web tests — test command comes from the variant; Run button label/command per kind.
|
||||
|
||||
### Wave 4-4: telemetry + grading proof
|
||||
|
||||
- **Task 4-4-01** (sandbox-engineer): digest pin test — `compute_digest` over a synthetic design-kind trace (validator harness events) → same feature classes as build traces (kind-agnostic by construction — now pinned); telemetry capture agent unchanged (content-agnostic `_EVENT_KINDS` verified).
|
||||
- Files: `apps/ai-service/ai_service/variants/templates.py`, `generator.py`, `store.py`, `api/variants.py`, `api/sandboxes.py`, `grading/features.py` (tests only), `apps/web/hooks/use-sandbox-session.ts`, `apps/web/components/learner/build-surface.tsx`, `packages/types/variants.ts`
|
||||
- **MH-4d** (absorbs former MH-4e per G-17 — no hope-shaped must-haves): design-kind E2E in the real-server harness — design variant → sandbox → starter files → Run validator harness in-ns → telemetry flows → digest computes (mock LLM, real stores) with concrete assertions: stored seqs contiguous 0..N exactly once AND the agent's spool ends ≤ the ack margin (real agent; where a fake agent is used, cite the P1 suite as the ack/trim coverage instead of asserting).
|
||||
|
||||
**Verification strategy P4:** `pnpm ai:test`, ruff, typecheck, web tests, `pnpm build`. Grading digest diff vs build traces = zero behavioral drift (pinned).
|
||||
|
||||
**Risks:** starter-file materialization is client-driven (missing-file hazard — mitigate: starter sets are small + template-authored; document server-side materialization as a future seam); `test_command` is wire-visible (template-authored, not learner-authored — documented); dual TS/Python schema sync (rule enforced in review).
|
||||
|
||||
---
|
||||
|
||||
## Phase 5: Final Review + Ship (milestone release v0.4.5)
|
||||
|
||||
- **Task 5-1-01** (lead-developer → ci-review personas): multi-persona review of all v0.5 changes (correctness, testing, security, performance, maintainability) — P0s fixed in-phase.
|
||||
- **Task 5-2-01** (ci-audit): reconstruction test (git log ↔ .ciagent/ files), file/branch/commit discipline, tag hygiene; critical fixes in-phase.
|
||||
- **Task 5-3-01** (lead-developer → ci-ship): merge phase/05 → milestone → main; tag **v0.4.5** (milestone release); Gitea release with full summary + `nextcraft-linux-x64` + `.sha256` assets; delete milestone branches; mark requirements complete; clear checkpoint.
|
||||
|
||||
---
|
||||
|
||||
## Must-Haves (Milestone)
|
||||
|
||||
- **MH-M1**: Mid-burst disconnect loses zero events outside the acked margin (P1 regression test, repeatable).
|
||||
- **MH-M2**: `openai-audio` provider passes byte-contract STT/TTS tests with key-redaction pins; defense flow works server-side end-to-end (manual probe documented); mock/browser paths unchanged.
|
||||
- **MH-M3**: Identity flow (submit → pending → verified) works under mock; gates enforce 16+/18+ in the binding composition order with verify-CTA payloads; PII never stored raw or logged (caplog sentinel green).
|
||||
- **MH-M4**: Design + simulation environment kinds provisionable with per-kind starter contents + exec policy; grading digest proven kind-agnostic.
|
||||
- **MH-M5**: All gates green at every phase ship: `pnpm ai:test`, ruff, `pnpm typecheck`, `pnpm build`, web + cli tests; binary assets on every release (D-036).
|
||||
- **MH-M6**: v0.1 surface regressions: zero **for verified allowlisted pilot learners** (existing learner/marketplace/employer/admin flows unchanged except the additive enroll route + badges/CTAs; gating unverified/under-age ids is REQ-5-004's purpose, not a regression).
|
||||
|
||||
**Out of scope (guarded):** real KYC vendor integration (protocol only), real auth sessions (self-asserted learner_id documented as pilot-scale), marketplace backend beyond the one gated stub route, new isolation tech (D-024 unchanged), server-side starter materialization (documented as future seam).
|
||||
1. Streaming is visibly incremental — tokens appear as they arrive, not as one blob
|
||||
2. Agent switcher shows the active agent (Coach/Tutor) and the response persona visibly changes
|
||||
3. Loading state during connection (streaming cursor / skeleton) before first token
|
||||
4. When ai-service is unreachable: inline error state + retry action on every chat/panel — no crashes, no console errors, no blank UI
|
||||
5. `[DONE]` reliably ends the stream (input re-enables, no stuck "typing" state)
|
||||
6. Navigating away mid-stream aborts cleanly — no leaked requests or post-unmount updates
|
||||
7. All new UI uses design tokens, supports dark mode, meets WCAG AA contrast
|
||||
8. Responsive at 375px, 768px, 1280px
|
||||
9. No hardcoded model names or URLs in UI code — all via `NEXT_PUBLIC_AI_SERVICE_URL` and server meta events
|
||||
10. `pnpm build` and `pnpm typecheck` pass with zero errors
|
||||
+23
-83
@@ -8,97 +8,35 @@ Nextcraft is an AI-native outcome school where graduates prove what they can bui
|
||||
|
||||
---
|
||||
|
||||
## Milestone v0.5 — Real Server Voice, KYC/Identity, Sandbox Environments, Seq-Lease (COMPLETE, shipped as v0.4.5)
|
||||
## Current Milestone: v0.2 — AI Tutor Architecture
|
||||
|
||||
**Scope (the four seams D-016 deferred out of v0.4, locked at v0.5 Phase 0 SPECIFY):**
|
||||
**Scope:** The six AI tutor agents (Coach, Tutor, Lab, Assessor, Proctor, Mentor) as real LLM-backed services in a new `apps/ai-service` Python FastAPI application, wired into the existing v0.1 learner surface chat UI with streaming responses. Provider-agnostic LLM layer (ollama-cloud default, local endpoint + deterministic mock for tests). Lab/Assessor/Proctor operate on mock engine inputs (simulated telemetry, pre-baked artifacts) — their real engines (sandbox fabric, assessment engine, identity verification) are v0.3+.
|
||||
|
||||
1. **Real server STT/TTS** (CUT-1/G-7 seam): implement the `openai-audio` VoiceProvider against OpenAI-compatible `/audio/transcriptions` (STT) + `/audio/speech` (TTS) — the D-030 protocol drop-in; mock stays first-class for tests; browser fallback stays for the no-key path. Voice defense (Examiner flows) gains the real server path end-to-end.
|
||||
2. **KYC/identity + age-gating backend** (REQ-F-017): real identity verification behind a provider protocol (mock-first, D-014 pattern), age-gate enforcement (school 16+, marketplace 18+ with verified identity) replacing the v0.1 visual-only flow, session/learner identity wired into the API surface (the G-5 allowlist evolves toward real identity).
|
||||
3. **Design/simulation sandbox environments** (REQ-F-021 remainder): extend the D-024/D-025 namespace fabric beyond the coding IDE — design-tool and simulation environment types alongside the existing build environment, one lifecycle, per-type telemetry.
|
||||
4. **Exec-telemetry seq-lease / replay-margin fix**: the one-line ACK gap documented in the P6 lesson (P07 review): reconnect replay margin so at-least-once ingest acknowledges received seqs and the capture agent resumes from the ack — closing the documented gap.
|
||||
**Status of v0.1:** Complete and shipped (v0.1.0). Founder agreement recorded (D-013).
|
||||
|
||||
**Success (milestone-level):** voice defense runs on real server STT/TTS when keys exist; age-gating is enforced by the backend (not a mockup page); the sandbox fabric provisions design/sim environments; the reconnect-replay gap is closed with a regression test.
|
||||
|
||||
## Prior Milestone: v0.4 — Distribution & Bootstrap CLI (COMPLETE, shipped as v0.3.4; hotfixes v0.3.5 fresh-box experience, v0.3.6 single-port unattended deploy)
|
||||
|
||||
**Scope (founder directive, 2026-09-12):** Streamline installing Nextcraft. Ship a bootstrap CLI with a single-liner install script, and publish release binaries on an ongoing basis for every release going forward.
|
||||
|
||||
**Delivered:** `nextcraft` CLI — `doctor` (prerequisite checks), `bootstrap` (deps + venv + env from templates + key validation), `verify` (health check), `dev` (thin passthrough to scripts/dev.sh); one-liner install script downloading the linux x64 binary from the latest Gitea release with sha256 + version integrity gates; binary release pipeline attached to every ship from v0.3.2 onward; install/quickstart documentation backed by a fresh-clone E2E test. All 5 requirements (REQ-4-001..005) complete.
|
||||
|
||||
**Status of v0.3:** Complete and shipped (v0.2.8). Credential engines live: namespace-isolated sandbox fabric, live build telemetry, process-trace grading, seeded variants, oral defense; real learner build/defense/grading surfaces.
|
||||
|
||||
**Status of v0.2:** Complete and shipped (v0.2.0). Six AI tutor agents live over mockengine inputs (D-015).
|
||||
|
||||
**Deferred from earlier plan:** REQ-F-017 (identity verification + age-gating KYC), real server STT/TTS, design/simulation sandbox environments, and the exec-telemetry seq-lease are all deferred to v0.5. Age-gating remains the v0.1-style visual flow mockup.
|
||||
|
||||
**Tech stack:** v0.1 TS monorepo (pnpm/turborepo, Next.js) + v0.2 Python FastAPI ai-service + new credential-engine services (sandbox fabric orchestrator, telemetry ingest, grading engine) in Python/TypeScript as determined at RESEARCH.
|
||||
**Tech stack:** v0.1 TS monorepo (pnpm/turborepo, Next.js) + new Python FastAPI service (`apps/ai-service`) with pydantic, SSE streaming, and an OpenAI-compatible provider client.
|
||||
|
||||
---
|
||||
|
||||
## v0.4 Requirements (Complete)
|
||||
## Requirements (Validated)
|
||||
|
||||
All 5 v0.4 requirements (REQ-4-001..005) are complete and shipped as v0.3.4:
|
||||
1. Bootstrap CLI — `nextcraft` executable with `doctor` / `bootstrap` / `verify` / `dev` (REQ-4-001, REQ-4-002)
|
||||
2. One-liner install — `curl | sh` fetching the linux x64 binary from the latest Gitea release with sha256 + version integrity verification (REQ-4-003)
|
||||
3. Ongoing release binaries — every release from v0.3.2 onward ships the CLI binary + checksum as release assets (REQ-4-004)
|
||||
4. Install documentation — README quickstart + CLI reference verified by a fresh-clone E2E test (REQ-4-005)
|
||||
The following requirements have been validated during specification and are locked for milestone v0.2 (REQ-F-001..006 activated from the deferred pool):
|
||||
|
||||
## v0.3 Requirements (Complete)
|
||||
|
||||
All 8 v0.3 requirements (REQ-3-001..008) are complete and shipped as v0.2.8. See REQUIREMENTS.md traceability matrix.
|
||||
1. AI tutor service infrastructure — `apps/ai-service` FastAPI application, provider-agnostic LLM client, SSE streaming, session/state handling
|
||||
2. Agent framework — base agent contracts, prompt management, streaming pipeline, structured outputs
|
||||
3. Coach agent — pacing, motivation, retrieval practice (REQ-F-001)
|
||||
4. Tutor agent — concept delivery, Socratic questioning (REQ-F-002)
|
||||
5. Lab agent — in-flow feedback over simulated sandbox telemetry (REQ-F-003, mock inputs)
|
||||
6. Assessor agent — rubric application to pre-baked artifacts and defenses (REQ-F-004, mock inputs)
|
||||
7. Proctor agent — integrity signals from mock telemetry, coaching interventions (REQ-F-005, mock inputs)
|
||||
8. Mentor agent — long-horizon career narrative (REQ-F-006)
|
||||
9. Learner surface integration — streaming chat UI wired to the real service, error/loading states
|
||||
|
||||
## v0.1 Requirements (Complete)
|
||||
|
||||
All 28 v0.1 requirements (REQ-001..028) are complete and shipped as v0.1.0. See REQUIREMENTS.md traceability matrix.
|
||||
|
||||
## Clarified Assumptions (v0.5 CLARIFY stage, full autonomy — auto-resolved)
|
||||
|
||||
| # | Ambiguity | Resolution | Confidence |
|
||||
|---|-----------|------------|------------|
|
||||
| A-301 | Which STT/TTS endpoint for `openai-audio`? | **OpenAI-compatible `/audio/transcriptions` + `/audio/speech`** via `AI_VOICE_BASE_URL` (ollama-cloud does not expose audio endpoints; any OpenAI-compatible audio API works — the provider is endpoint-agnostic by config, D-014 pattern). Model via `AI_VOICE_MODEL` (default `whisper-1` STT / `tts-1` TTS class names, configurable). Keys env-only, never committed. | 0.75 |
|
||||
| A-302 | Does real voice replace browser fallback? | **No — layered.** `openai-audio` when configured, browser-native second-class fallback (CUT-1 keeps it first-class for the no-key path), deterministic mock for tests. The defense surface probes provider capability and degrades with a visible badge (mic mock vs browser vs server). | 0.85 |
|
||||
| A-303 | KYC vendor in v0.5? | **No vendor — protocol + mock-first backend.** `IdentityProvider` protocol (submit/poll/verify), deterministic mock (approve-on-policy), SQLite store. A real vendor (Stripe Identity/Persona/Onfido class) drops in later without API changes. Solo-founder economics: no vendor spend before pilot. | 0.85 |
|
||||
| A-304 | What counts as "verified 18+" for marketplace? | **Provider verdict = DOB-verified 18+; until a real vendor exists the mock verdict is explicit and marked `mock` in API responses** so downstream surfaces can label unverified state honestly (never display mock-verified as production-verified). | 0.80 |
|
||||
| A-305 | PII storage? | **Server-side only, minimal**: document refs + provider verdicts in SQLite; raw documents NEVER stored, NEVER logged (payload scrubbing pinned by test). Enrollment DOB stored as derived age-band, not raw DOB where possible. | 0.85 |
|
||||
| A-306 | Age-gate enforcement points? | **School enrollment (16+)**: identity submit → verify → age check at enrollment API. **Marketplace (18+ verified)**: gated marketplace routes check verified-identity verdict; unverified → 403 with verify-CTA payload. G-5 learner allowlist REMAINS as a pilot guard (identity doesn't replace sandbox rate caps). | 0.82 |
|
||||
| A-307 | What are "design" and "simulation" environments concretely? | **Same namespace fabric, typed starter contents + allowed commands.** `design`: canvas-style artifacts (files the learner edits: SVG/HTML/schematic text), Run = validator/renderer harness command; `simulation`: parameterized run harness (benchmark scripts + dataset files), Run = bounded harness execution. No new isolation tech — D-024 unchanged; the ENVIRONMENT is a typing over starter files + command policy. | 0.78 |
|
||||
| A-308 | Does the learner UI need new surfaces for design/sim? | **Reuse the build surface.** The existing /build flow accepts `environment` kind; the file tree, editor, Run/Test buttons, and read-only output panel (CUT-2) all carry over; only starter contents and command policy differ per kind. No new route groups; variants carry the kind. | 0.80 |
|
||||
| A-309 | Seq-lease protocol shape? | **Ingest acks highest-contiguous received seq per (learner, task) on the existing WS (JSON ack frame); the capture agent trims its spool to the ack and replays from there on reconnect.** Replay margin bounded (spool cap already exists). No new endpoint; G-3 flood semantics unchanged; acks are advisory hints, gap detection stays authoritative. | 0.80 |
|
||||
| A-310 | Where does identity live in the app? | **New `ai_service/identity/` module (D-031 pattern — extend ai-service, no new apps)**: provider protocol + mock, SQLite store (D-027 family), `/v1/identity/*` router. Web enrollment + marketplace surfaces call it via engine-client. | 0.85 |
|
||||
|
||||
## Clarified Assumptions (v0.4 CLARIFY stage, full autonomy — auto-resolved)
|
||||
|
||||
| # | Ambiguity | Resolution | Confidence |
|
||||
|---|-----------|------------|-------------|
|
||||
| A-201 | CLI language/toolchain for the binary? | **Probe-driven at RESEARCH** — Go → Rust → Node SEA → Python zipapp fallback chain; spec stays toolchain-agnostic so PLAN locks the probe-verified toolchain | 0.70 |
|
||||
| A-202 | Does bootstrap replace scripts/bootstrap.sh? | **No — reuse it.** CLI wraps existing `scripts/bootstrap.sh` + `scripts/dev.sh` via subprocess; zero orchestration logic duplicated in the CLI (thin passthrough pattern) | 0.85 |
|
||||
| A-203 | Where does the one-liner fetch the binary? | **Gitea latest-release API** (`/repos/{owner}/{repo}/releases/latest`) → download `nextcraft-linux-x64` + `.sha256` asset; repo raw serves `install.sh` as the stable URL | 0.80 |
|
||||
| A-204 | Install target + PATH? | **~/.local/bin** (XDG-style, no sudo), PATH hint printed when missing; `--dest` override flag | 0.85 |
|
||||
| A-205 | Binary "ongoing releases" scope? | **Every ship from v0.4 onward** attaches `nextcraft-linux-x64` + sha256 sidecar to the Gitea release — the ship workflow gains an asset step; retroactive binaries for old releases NOT required | 0.90 |
|
||||
| A-206 | No binary available yet / non-linux? | **Graceful degradation**: install script prints source-bootstrap instructions (git clone + scripts/bootstrap.sh) — never a hard fail | 0.88 |
|
||||
| A-207 | Checksum trust root? | **sha256 sidecar shipped as a release asset next to the binary** (same release, same channel); script verifies download against it. Signature/PKI out of scope for v0.4 (single forge, TLS transport) | 0.75 |
|
||||
| A-208 | Which prerequisites does doctor check? | node ≥18, pnpm ≥8, python3 ≥3.11, git, `unshare` availability (sandbox fabric needs it) — versions from the existing bootstrap tooling, not invented | 0.85 |
|
||||
| A-209 | Does `dev` manage multiple processes? | **No.** Thin passthrough to scripts/dev.sh only — the CLI stays bootstrap-scoped (D-016); orchestration remains in dev.sh | 0.82 |
|
||||
| A-210 | `.env.secrets` handling by bootstrap? | **Template copy only for `.env.example` → `.env`; secrets NEVER generated, NEVER committed; bootstrap validates presence of optional keys and warns (not blocks) when missing — mock-first providers keep the stack runnable** | 0.90 |
|
||||
|
||||
## Clarified Assumptions (v0.3 CLARIFY stage, full autonomy — auto-resolved)
|
||||
|
||||
| # | Ambiguity | Resolution | Confidence |
|
||||
|---|-----------|------------|-------------|
|
||||
| A-101 | Sandbox isolation technology? | **`unshare` user+mount+pid+net namespace subprocess isolation** per sandbox (probe-verified: in-ns uid=0, network fully isolated with 0 interfaces, writes land in an isolated bind-mounted workdir; proc-remount is not permitted in this context but is not required). No Docker/Podman/VMs — none present on the box; no sudo. A `SandboxBackend` protocol keeps a future containerd swap possible. Falls back further to a plain chroot-free subprocess with a cwd-jail if userns ever unavailable (tested path is userns). | 0.8 |
|
||||
| A-102 | Sandbox scope in v0.3? | **Coding IDE only** (web terminal + file tree + run/test). The "design tool" and "simulation" environments specified in REQ-F-021 are deferred to v0.4 — a single real build environment is enough to prove the credential pipeline end-to-end (telemetry → trace → grade → defense). | 0.75 |
|
||||
| A-103 | Live in-browser build UX? | **Run/Test buttons executing in the namespace sandbox + HTTP file-tree/CRUD + read-only exec-output panel** (CUT-2/G-8 — the interactive xterm.js shell relay is deferred to v0.4; `@xterm/*` is not a v0.3 dependency). No full Monaco LSP in v0.3 — a code editor with syntax highlight (existing) is sufficient and far cheaper. | 0.72 |
|
||||
| A-104 | Telemetry transport? | **WebSocket** from sandbox to a new ingestion endpoint on ai-service for live events; **SQLite-backed** ordered event log (`ai_service/telemetry/`) gives durability + at-least-once delivery + replay. Events carry monotonic `seq` per (learner,task) so gaps are detectable. | 0.8 |
|
||||
| A-105 | Where do traces live? | **SQLite** (`ai_service` data dir), introducing the first real persistence. SQLModel/SQLAlchemy for typed access. Chosen over Postgres because solo-founder + single box + low write volume; the `TraceStore` protocol is Postgres-migration-ready like SessionStore was. | 0.75 |
|
||||
| A-106 | Process-trace grading model? | **LLM-based grader**: structure the trace into a compact timeline digest (command categories, error/fix cycles, idle gaps, test passes) → Assessor-style rubric prompt → structured score via existing D-020 JSON defense. Deterministic features (test pass/fail, edit count) computed in code, not left to the LLM. | 0.7 |
|
||||
| A-107 | Variant generation mechanism? | **Parameterized task templates + LLM instantiation**, seeded per learner. Generator fills typed parameter slots (scenario, constraints, data) from a template library; variant seed + parameters persisted to SQLite for grading fairness and proctoring cross-check. Difficulty normalized by template-level rubric anchors. | 0.72 |
|
||||
| A-108 | Voice defense — STT/TTS providers? | **Provider-agnostic, mock-first like the LLM layer (D-014).** Real path: browser `MediaRecorder` → audio to ai-service → **OpenAI-compatible `/audio/transcriptions`** (Whisper STT) and **`/audio/speech`** (TTS) against ollama-cloud or a compatible endpoint; fallbacks: browser `SpeechRecognition`/`speechSynthesis` when no server keys. `VoiceProvider` protocol + deterministic mock (returns canned transcript) so tests never call a voice API. | 0.62 |
|
||||
| A-109 | Defense dialogue shape? | Reuse BaseAgent: an `Examiner` agent (seventh agent) streams examiner questions over the existing SSE pipeline; integrity signals (long pauses, off-scope answers, reading-from-notes cadence) emitted alongside the transcript to Proctor. | 0.8 |
|
||||
| A-110 | KYC / age-gating in v0.3? | **Deferred per founder directive.** No real identity backend. Age-gating stays the v0.1 visual flow mockup. Personas omit a security-engineer; security review via verifier + Phase 7 secrets-hygiene checklist. **Abuse control is NOT deferred with KYC (G-5):** v0.3 ships per-learner sandbox caps (`AI_SANDBOX_MAX_PER_LEARNER`), a global create-rate cap, and a server-side `learner_id` allowlist (`AI_LEARNER_ALLOWLIST`) so the unauthenticated surface cannot exhaust shared NPROC/disk. Documented in the release note. | 0.98 |
|
||||
| A-111 | New services vs extend ai-service? | **Extend ai-service**, don't fork new Python apps. Telemetry ingestion, trace grading, variant generation, voice, and sandbox orchestration all live as new modules in `apps/ai-service` (they share the LLM provider pool + config + session infra). Only the in-sandbox capture agent is a separate tiny Python process shipped into the namespace sandbox. | 0.82 |
|
||||
| A-112 | Sandbox on a single dev/school box — capacity? | v0.3 targets **1–5 concurrent sandboxes** (founder + pilot learners). No horizontal scaling, no queue. Concurrency guard returns 503 when full. Scaling is post-MVP. | 0.8 |
|
||||
|
||||
## Clarified Assumptions (v0.2 CLARIFY stage, full autonomy — auto-resolved)
|
||||
## Clarified Assumptions (CLARIFY stage, full autonomy — auto-resolved)
|
||||
|
||||
| # | Ambiguity | Resolution | Confidence |
|
||||
|---|-----------|------------|-------------|
|
||||
@@ -115,10 +53,14 @@ All 28 v0.1 requirements (REQ-001..028) are complete and shipped as v0.1.0. See
|
||||
|
||||
## Requirements (Active — Future Milestones)
|
||||
|
||||
The following remain deferred beyond v0.3 and will be activated in subsequent milestones:
|
||||
The following remain deferred beyond v0.2 and will be activated in subsequent milestones:
|
||||
|
||||
The following remain deferred beyond v0.2 and will be activated in subsequent milestones:
|
||||
|
||||
- Identity verification and age-gating logic (16+/18+) — the real KYC backend (**deferred from v0.3 per founder directive**; visual flow already exists in v0.1)
|
||||
- Competency graph engine and adaptive pathways
|
||||
- Assessment engine (process-trace grading, oral defense, per-learner variant tasks) — v0.3+
|
||||
- Sandbox fabric (sandboxed IDE, design tool, simulation) — v0.3+
|
||||
- Identity verification and age-gating logic (16+/18+) — the real KYC backend (v0.3+; visual flow already exists in v0.1)
|
||||
- Marketplace job aggregation pipeline (3M+ jobs from 120K companies)
|
||||
- AI-powered tagging, semantic vector search, company enrichment
|
||||
- AI resume parsing and job matching
|
||||
@@ -164,7 +106,7 @@ The following remain deferred beyond v0.3 and will be activated in subsequent mi
|
||||
| D-003 | TypeScript monorepo (pnpm/turborepo) + Next.js | Unified codebase for all 4 surfaces. Shared component library, types, mock data. Next.js App Router for route-based surface separation. Python AI services deferred to later milestones. | Monorepo structure with apps/web + packages/* |
|
||||
| D-004 | All 4 surfaces in v0.1 (Learner, Marketplace, Employer, Admin) | Founder selected all 4 surfaces for the prototype. Complete product visualization before any backend work. | 24 REQ-IDs covering all surfaces + shared infrastructure |
|
||||
| D-005 | High-fidelity interactive prototype | Founder selected high-fidelity over wireframes. Realistic mock data, navigation flows, responsive layouts, component library. No backend calls. | Clickable prototype with realistic content |
|
||||
| D-006 | Release forge = Gitea @ git.coreci.dev, owner=coreci, repo=nextcraft | Founder-provided Gitea instance for release management. Token stored in .ciagent/.env.secrets. Forge migrated 2026-09-12 from git.cloudinit.dev → git.coreci.dev (old host decommissioned; all releases migrated, IDs preserved). | Ship workflow creates tags + releases on Gitea |
|
||||
| D-006 | Release forge = Gitea @ git.cloudinit.dev, owner=coreci, repo=nextcraft | Founder-provided Gitea instance for release management. Token stored in .ciagent/.env.secrets. | Ship workflow creates tags + releases on Gitea |
|
||||
| D-007 | Full autonomy for CIAgent pipeline | Founder selected full autonomy. No HITL after clarify. Auto-decide above confidence 0.60. Escalation hooks: deploy, delete_data, merge_to_main. | Rapid autonomous building with kill criteria |
|
||||
| D-008 | Shared component library in packages/ui/ | All surfaces share a unified design system with surface-specific theming via CSS variables. Promotes consistency and reduces duplication. | packages/ui, packages/mock-data, packages/types |
|
||||
| D-009 | AI tutor UI as chat interface mockup with pre-scripted responses | The learner surface includes an AI tutor chat UI mockup. No real AI backend — pre-scripted responses simulate the Coach and Tutor agents. | Mockup only in v0.1, real agents in future milestone |
|
||||
@@ -174,8 +116,6 @@ The following remain deferred beyond v0.3 and will be activated in subsequent mi
|
||||
| D-013 | v0.1 prototype founder-agreed; D-001 business-logic gate unlocked | Founder approved starting v0.2 with AI Tutor Architecture, which constitutes agreement of the v0.1 prototype per D-001. Recorded at v0.2 SPECIFY. | Business logic authorized from v0.2 onward |
|
||||
| D-014 | Provider-agnostic LLM layer; ollama-cloud as initial provider | OpenAI-compatible client abstraction with pluggable providers: ollama-cloud (https://ollama.com/v1, default), local OpenAI-compatible endpoint, deterministic mock (tests/CI). Keys in gitignored .ciagent/.env.secrets, never in code or commits. | apps/ai-service llm package with 3 providers; default=ollama-cloud |
|
||||
| D-015 | All six agents implemented as real LLM services; engines mocked | Coach/Tutor/Mentor fully real. Lab/Assessor/Proctor are real LLM logic over mock inputs (simulated telemetry, pre-baked artifacts) since sandbox fabric, assessment engine, and identity verification are v0.3+. Consistent with v0.1's mock-data approach. | REQ-F-001..006 complete in v0.2; real engines deferred to v0.3+ |
|
||||
| D-016 | v0.4 = Distribution & Bootstrap CLI (founder directive supersedes previously-named v0.4 seams) | Founder directive 2026-09-12: focus this milestone on streamlining install, a bootstrap CLI with a one-liner install script, ongoing release binaries. Real server STT/TTS, KYC, design/simulation envs, seq-lease move to v0.5. | Milestone scope locked at SPECIFY; binary = CLI-only, linux x64 |
|
||||
| D-039 | Single-port deploy + unattended dev (v0.3.6 hotfix, founder directive) | Only :8420 reachable behind HAProxy; site was down (domain root hit the API 404). Web app = static export served same-origin by the ai-service (`AI_WEB_STATIC_DIR`); `dev -d`/`stop`/`log` daemonize ops; durable state (DB, sandboxes, pid/log) moves to `~/.nextcraft/` out of the repo. | One port serves UI+API; unattended deploy recipe in deploy/README.md |
|
||||
|
||||
---
|
||||
|
||||
|
||||
+59
-167
@@ -1,106 +1,33 @@
|
||||
# Nextcraft — REQUIREMENTS.md
|
||||
|
||||
## v0.5 Requirements (Complete — Real Server Voice, KYC/Identity, Sandbox Environments, Seq-Lease; shipped as v0.4.5)
|
||||
|
||||
### Real Server Voice
|
||||
|
||||
| ID | Description | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-5-001 | `openai-audio` VoiceProvider: server STT (`/audio/transcriptions`) + TTS (`/audio/speech`) against an OpenAI-compatible endpoint via the existing D-030 protocol; provider selection by `AI_VOICE_PROVIDER` (+ base URL/key from env, never committed); deterministic mock stays first-class; browser fallback unchanged | critical | 2 | complete |
|
||||
| REQ-5-002 | Voice defense real path end-to-end: examiner dialogue answers transcribed server-side (audio upload → transcript), examiner questions spoken via server TTS (audio returned to the client); transcripts + integrity signals unchanged | critical | 2 | complete |
|
||||
|
||||
### Identity & Age-Gating
|
||||
|
||||
| ID | Description | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-5-003 | Identity provider protocol + mock-first backend (REQ-F-017): verify-identity flow (submit → pending → verified/rejected with document refs), provider-agnostic (mock default; a real KYC vendor drops in later), PII stored server-side only, never logged | critical | 3 | complete |
|
||||
| REQ-5-004 | Age-gating enforced by the backend: school floor 16+ verified at enrollment, marketplace 18+ with verified identity — API surfaces reject under-age/unverified callers on gated routes (replaces the v0.1 visual-only flow; G-5 allowlist evolves toward real identity, allowlist remains as pilot guard) | critical | 3 | complete |
|
||||
|
||||
### Sandbox Environments
|
||||
|
||||
| ID | Description | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-5-005 | Design + simulation sandbox environment types (REQ-F-021 remainder): extend the namespace fabric with environment kinds beyond the coding IDE (design-tool: canvas/editor surfaces with file artifacts; simulation: run/benchmark harnesses) — one lifecycle, one telemetry path, per-type starter contents + allowed commands | high | 4 | complete |
|
||||
| REQ-5-006 | Environment-typed learner surface: the build/defend flow accepts environment kind, telemetry captures per-kind events, grading digest stays kind-agnostic | high | 4 | complete |
|
||||
|
||||
### Telemetry Durability
|
||||
|
||||
| ID | Description | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-5-007 | Exec-telemetry seq-lease/replay-margin fix: WS ingest acknowledges received seqs; capture agent resumes from the ack on reconnect (bounded replay margin) — closes the P6-lesson one-line ACK gap with a real-server regression test | high | 1 | complete |
|
||||
|
||||
## v0.4 Requirements (Complete — Distribution & Bootstrap CLI, shipped as v0.3.4)
|
||||
|
||||
### Bootstrap CLI
|
||||
|
||||
| ID | Description | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-4-001 | `nextcraft` CLI (linux x64 binary): `doctor` command checking prerequisites (node, pnpm, python3, git, unshare) with actionable error messages | critical | 1 | complete |
|
||||
| REQ-4-002 | `bootstrap` command: pnpm install, ai-service venv + pinned deps, .env from templates, key validation, .env.secrets handling; `verify` health check (ports, imports, builds); `dev` thin passthrough to scripts/dev.sh | critical | 1 | complete |
|
||||
|
||||
### Distribution
|
||||
|
||||
| ID | Description | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-4-003 | One-liner install script (`curl -fsSL <url> \| bash`): detects linux x64, resolves latest release from Gitea API, downloads binary + checksum, verifies sha256, installs to ~/.local/bin (PATH hint), degrades to source-bootstrap instructions when no binary | critical | 2 | complete |
|
||||
| REQ-4-004 | Binary release pipeline: reproducible linux x64 build script, sha256 checksum sidecar, upload as release assets on every ship from v0.4 onward (ongoing binaries requirement) | critical | 2 | complete |
|
||||
| REQ-4-005 | Install + quickstart documentation: README one-liner quickstart, CLI command reference, fresh-clone-to-running-stack end-to-end verification | high | 3 | complete |
|
||||
|
||||
## v0.3 Requirements (Credential Engines)
|
||||
|
||||
### Sandbox & Telemetry
|
||||
|
||||
| ID | Description | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-3-001 | Sandbox fabric: isolated per-learner execution environments (sandboxed IDE, design tool, simulation) with lifecycle management | critical | 1 | complete |
|
||||
| REQ-3-002 | Sandbox isolation + resource limits: per-learner isolation boundary, CPU/memory quotas (rlimits), wall-clock time quota, disk-quota via per-sandbox workdir usage sweep (best-effort, not kernel-enforced), no cross-tenant access, snapshot support. **Known gap (v0.3): per-sandbox pids and hard disk caps are NOT kernel-enforceable without cgroup delegation/sudo — documented as accepted risk** | critical | 1 | complete |
|
||||
| REQ-3-003 | Live build telemetry: in-environment capture of process events (commands, file diffs, run/test results, activity) streamed reliably to ai-service with per-learner trace persistence | critical | 2 | complete |
|
||||
|
||||
### Credential Engines
|
||||
|
||||
| ID | Description | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-3-004 | Process-trace grading engine: grade artifacts from their full process traces; rubric-aligned structured scores; feeds Assessor real inputs | critical | 3 | complete |
|
||||
| REQ-3-005 | Variant task generation: per-learner task variants (no two learners get identical prompts); variant seed registry; difficulty normalization | high | 4 | complete |
|
||||
| REQ-3-006 | Oral/voice defense: AI examiner conducts spoken defense (STT → dialogue → TTS); transcript + integrity signals captured; feeds Proctor/Mentor | high | 5 | complete |
|
||||
|
||||
### Agent Re-grounding & Integration
|
||||
|
||||
| ID | Description | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-3-007 | Agent re-grounding: Lab consumes live telemetry; Assessor consumes grading-engine output; Proctor consumes telemetry + defense integrity signals (replace v0.2 mocks) | critical | 6 | complete |
|
||||
| REQ-3-008 | Learner surface integration: sandbox mockup → real in-browser build/run with live telemetry; assessment mockup → live defense + live grading | critical | 6 | complete |
|
||||
|
||||
---
|
||||
|
||||
## v0.2 Requirements (Complete — AI Tutor Architecture)
|
||||
## v0.2 Requirements (AI Tutor Architecture)
|
||||
|
||||
### AI Service Infrastructure
|
||||
|
||||
| ID | Description | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-2-001 | apps/ai-service scaffolding: Python FastAPI app, pydantic settings, uvicorn, health endpoint, CORS, pytest setup, pnpm/turbo integration scripts | critical | 1 | complete |
|
||||
| REQ-2-002 | Provider-agnostic LLM client: OpenAI-compatible provider interface with ollama-cloud (default), local-endpoint, and deterministic mock providers; key resolution from env files | critical | 1 | complete |
|
||||
| REQ-2-003 | SSE streaming endpoint plumbing: chat completion streaming from provider through FastAPI to the Next.js client | critical | 1 | complete |
|
||||
| REQ-2-004 | Agent framework: base agent contracts, session/state store, prompt management, streaming pipeline, structured output support | critical | 2 | complete |
|
||||
| REQ-2-001 | apps/ai-service scaffolding: Python FastAPI app, pydantic settings, uvicorn, health endpoint, CORS, pytest setup, pnpm/turbo integration scripts | critical | 1 | pending |
|
||||
| REQ-2-002 | Provider-agnostic LLM client: OpenAI-compatible provider interface with ollama-cloud (default), local-endpoint, and deterministic mock providers; key resolution from env files | critical | 1 | pending |
|
||||
| REQ-2-003 | SSE streaming endpoint plumbing: chat completion streaming from provider through FastAPI to the Next.js client | critical | 1 | pending |
|
||||
| REQ-2-004 | Agent framework: base agent contracts, session/state store, prompt management, streaming pipeline, structured output support | critical | 2 | pending |
|
||||
|
||||
### AI Tutor Agents
|
||||
|
||||
| ID | Description | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-2-005 | Coach agent (REQ-F-001): pacing, motivation, retrieval practice — full LLM implementation | critical | 3 | complete |
|
||||
| REQ-2-006 | Tutor agent (REQ-F-002): concept delivery, Socratic questioning — full LLM implementation | critical | 3 | complete |
|
||||
| REQ-2-007 | Lab agent (REQ-F-003): in-flow feedback over simulated sandbox telemetry (mock inputs) | high | 4 | complete |
|
||||
| REQ-2-008 | Assessor agent (REQ-F-004): rubric application to pre-baked artifacts and defense transcripts (mock inputs) | high | 4 | complete |
|
||||
| REQ-2-009 | Proctor agent (REQ-F-005): integrity signals from mock telemetry, coaching interventions | high | 5 | complete |
|
||||
| REQ-2-010 | Mentor agent (REQ-F-006): long-horizon career narrative | high | 5 | complete |
|
||||
| REQ-2-005 | Coach agent (REQ-F-001): pacing, motivation, retrieval practice — full LLM implementation | critical | 3 | pending |
|
||||
| REQ-2-006 | Tutor agent (REQ-F-002): concept delivery, Socratic questioning — full LLM implementation | critical | 3 | pending |
|
||||
| REQ-2-007 | Lab agent (REQ-F-003): in-flow feedback over simulated sandbox telemetry (mock inputs) | high | 4 | pending |
|
||||
| REQ-2-008 | Assessor agent (REQ-F-004): rubric application to pre-baked artifacts and defense transcripts (mock inputs) | high | 4 | pending |
|
||||
| REQ-2-009 | Proctor agent (REQ-F-005): integrity signals from mock telemetry, coaching interventions | high | 5 | pending |
|
||||
| REQ-2-010 | Mentor agent (REQ-F-006): long-horizon career narrative | high | 5 | pending |
|
||||
|
||||
### Learner Surface Integration
|
||||
|
||||
| ID | Description | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-2-011 | Learner chat UI wired to real service: streaming responses, agent routing, error and loading states | critical | 6 | complete |
|
||||
| REQ-2-012 | Byte tutorial viewer, build sandbox, and assessment mockups surface Lab/Assessor/Proctor outputs (mock engine inputs) | high | 6 | complete |
|
||||
| REQ-2-011 | Learner chat UI wired to real service: streaming responses, agent routing, error and loading states | critical | 6 | pending |
|
||||
| REQ-2-012 | Byte tutorial viewer, build sandbox, and assessment mockups surface Lab/Assessor/Proctor outputs (mock engine inputs) | high | 6 | pending |
|
||||
|
||||
---
|
||||
|
||||
@@ -111,58 +38,58 @@
|
||||
| ID | Description | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-001 | Monorepo scaffolding: pnpm workspaces, turborepo, Next.js app, TypeScript config, ESLint, Prettier | critical | 1 | complete |
|
||||
| REQ-002 | Shared component library: design tokens, Button, Input, Card, Badge, Avatar, Dialog, Navigation, Table, Tabs, Progress, Tooltip, Skeleton, Toast | critical | 1 | complete |
|
||||
| REQ-003 | Mock data layer: typed mock data for competency stacks, job listings, candidate profiles, employers, learner progress | critical | 1 | complete |
|
||||
| REQ-004 | Routing and navigation: App Router route groups for (learner), (marketplace), (employer), (admin); shared layout components; cross-surface navigation | critical | 1 | complete |
|
||||
| REQ-005 | Responsive layout system: mobile, tablet, desktop breakpoints; container components; grid system | high | 1 | complete |
|
||||
| REQ-002 | Shared component library: design tokens, Button, Input, Card, Badge, Avatar, Dialog, Navigation, Table, Tabs, Progress, Tooltip, Skeleton, Toast | critical | 1 | pending |
|
||||
| REQ-003 | Mock data layer: typed mock data for competency stacks, job listings, candidate profiles, employers, learner progress | critical | 1 | pending |
|
||||
| REQ-004 | Routing and navigation: App Router route groups for (learner), (marketplace), (employer), (admin); shared layout components; cross-surface navigation | critical | 1 | pending |
|
||||
| REQ-005 | Responsive layout system: mobile, tablet, desktop breakpoints; container components; grid system | high | 1 | pending |
|
||||
|
||||
### Learner Surface
|
||||
|
||||
| ID | Description | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-006 | Landing page: hero, value proposition, program highlights, how-it-works (Byte→Build→Demonstrate→Defend), testimonials mockup, CTA to program catalog | critical | 2 | complete |
|
||||
| REQ-007 | Program catalog: grid of competency stacks (AI Orchestration Engineer, AI Safety & Governance Lead, Human-AI Product Designer, AI-Augmented Field Operator, Computational Sciences Practitioner); stack cards with role descriptions | critical | 2 | complete |
|
||||
| REQ-008 | Competency stack view: selected stack detail with 12-18 competencies listed, progress indicators, microcredential badges, mastery status | critical | 2 | complete |
|
||||
| REQ-009 | Learner dashboard: active competencies, progress graph, recent artifacts, upcoming defenses, AI tutor chat mockup, milestone tracker | critical | 2 | complete |
|
||||
| REQ-010 | Byte tutorial viewer: 3-7 minute micro-tutorial layout with concept panel, worked example panel, code/design/simulation viewer mockup | high | 2 | complete |
|
||||
| REQ-011 | Build sandbox mockup: sandboxed IDE/design tool/simulation UI mockup with toolbar, file explorer, editor area, telemetry sidebar (process capture indicators) | high | 2 | complete |
|
||||
| REQ-012 | Assessment/defense mockup: rubric display, AI reviewer panel, oral defense interface (voice/mic mockup), process trace timeline, artifact viewer | high | 2 | complete |
|
||||
| REQ-006 | Landing page: hero, value proposition, program highlights, how-it-works (Byte→Build→Demonstrate→Defend), testimonials mockup, CTA to program catalog | critical | 2 | pending |
|
||||
| REQ-007 | Program catalog: grid of competency stacks (AI Orchestration Engineer, AI Safety & Governance Lead, Human-AI Product Designer, AI-Augmented Field Operator, Computational Sciences Practitioner); stack cards with role descriptions | critical | 2 | pending |
|
||||
| REQ-008 | Competency stack view: selected stack detail with 12-18 competencies listed, progress indicators, microcredential badges, mastery status | critical | 2 | pending |
|
||||
| REQ-009 | Learner dashboard: active competencies, progress graph, recent artifacts, upcoming defenses, AI tutor chat mockup, milestone tracker | critical | 2 | pending |
|
||||
| REQ-010 | Byte tutorial viewer: 3-7 minute micro-tutorial layout with concept panel, worked example panel, code/design/simulation viewer mockup | high | 2 | pending |
|
||||
| REQ-011 | Build sandbox mockup: sandboxed IDE/design tool/simulation UI mockup with toolbar, file explorer, editor area, telemetry sidebar (process capture indicators) | high | 2 | pending |
|
||||
| REQ-012 | Assessment/defense mockup: rubric display, AI reviewer panel, oral defense interface (voice/mic mockup), process trace timeline, artifact viewer | high | 2 | pending |
|
||||
|
||||
### Marketplace Surface
|
||||
|
||||
| ID | Description | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-013 | Job board listing: searchable grid of AI-era job listings, filter sidebar (skills, seniority, location, salary), result cards with match score | critical | 3 | complete |
|
||||
| REQ-014 | Job detail page: full job description, required competencies, employer info, application CTA, related jobs, AI-matched skills breakdown | critical | 3 | complete |
|
||||
| REQ-015 | Employer profile: company overview, logo, description, social links, open positions, company culture mockup | high | 3 | complete |
|
||||
| REQ-016 | Search/filter UI: semantic search bar, skill tags, category filters, seniority filter, remote/on-site toggle, saved searches mockup | critical | 3 | complete |
|
||||
| REQ-017 | Pricing page: job posting packages (single, bundle, enterprise), talent access plans, subscription tiers, feature comparison table | high | 3 | complete |
|
||||
| REQ-013 | Job board listing: searchable grid of AI-era job listings, filter sidebar (skills, seniority, location, salary), result cards with match score | critical | 3 | pending |
|
||||
| REQ-014 | Job detail page: full job description, required competencies, employer info, application CTA, related jobs, AI-matched skills breakdown | critical | 3 | pending |
|
||||
| REQ-015 | Employer profile: company overview, logo, description, social links, open positions, company culture mockup | high | 3 | pending |
|
||||
| REQ-016 | Search/filter UI: semantic search bar, skill tags, category filters, seniority filter, remote/on-site toggle, saved searches mockup | critical | 3 | pending |
|
||||
| REQ-017 | Pricing page: job posting packages (single, bundle, enterprise), talent access plans, subscription tiers, feature comparison table | high | 3 | pending |
|
||||
|
||||
### Employer Dashboard
|
||||
|
||||
| ID | Description | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-018 | Employer dashboard overview: active postings, applicant pipeline, talent matches, analytics mockup (charts, placement stats) | critical | 4 | complete |
|
||||
| REQ-019 | Talent search: searchable candidate database with AI-matched filters, candidate cards showing competency stack, microcredentials, artifact count, defense score | critical | 4 | complete |
|
||||
| REQ-020 | Candidate profile view: full candidate profile with artifact gallery, process trace summary, oral defense transcripts, competency graph, microcredential verification | high | 4 | complete |
|
||||
| REQ-021 | Posting management: create/edit/delete job postings, posting status tracking, applicant list per posting, interview pipeline mockup | high | 4 | complete |
|
||||
| REQ-018 | Employer dashboard overview: active postings, applicant pipeline, talent matches, analytics mockup (charts, placement stats) | critical | 4 | pending |
|
||||
| REQ-019 | Talent search: searchable candidate database with AI-matched filters, candidate cards showing competency stack, microcredentials, artifact count, defense score | critical | 4 | pending |
|
||||
| REQ-020 | Candidate profile view: full candidate profile with artifact gallery, process trace summary, oral defense transcripts, competency graph, microcredential verification | high | 4 | pending |
|
||||
| REQ-021 | Posting management: create/edit/delete job postings, posting status tracking, applicant list per posting, interview pipeline mockup | high | 4 | pending |
|
||||
|
||||
### Admin Surface
|
||||
|
||||
| ID | Description | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-022 | Admin overview: platform metrics (learners, employers, placements, completion rate, NPS), recent activity feed, system health mockup | critical | 5 | complete |
|
||||
| REQ-023 | Learner management: searchable learner table, learner detail view, progress tracking, competency completion status, credential issuance log | high | 5 | complete |
|
||||
| REQ-024 | Competency graph viewer: interactive visualization of competency stacks and their relationships, node/edge graph using react-flow, stack details on node click | high | 5 | complete |
|
||||
| REQ-025 | Marketplace moderation: job posting review queue, employer verification queue, flagged content, content moderation tools mockup | high | 5 | complete |
|
||||
| REQ-022 | Admin overview: platform metrics (learners, employers, placements, completion rate, NPS), recent activity feed, system health mockup | critical | 5 | pending |
|
||||
| REQ-023 | Learner management: searchable learner table, learner detail view, progress tracking, competency completion status, credential issuance log | high | 5 | pending |
|
||||
| REQ-024 | Competency graph viewer: interactive visualization of competency stacks and their relationships, node/edge graph using react-flow, stack details on node click | high | 5 | pending |
|
||||
| REQ-025 | Marketplace moderation: job posting review queue, employer verification queue, flagged content, content moderation tools mockup | high | 5 | pending |
|
||||
|
||||
### Polish & Integration
|
||||
|
||||
| ID | Description | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-026 | Cross-surface navigation: role switcher (learner/employer/admin), breadcrumbs, consistent header/footer across all surfaces | critical | 6 | complete |
|
||||
| REQ-027 | Visual consistency audit: typography scale, color palette, spacing system, dark mode toggle, accessibility baseline (WCAG AA contrast) | high | 6 | complete |
|
||||
| REQ-028 | Component library finalization: Storybook setup, component documentation, prop tables, usage examples | medium | 6 | complete |
|
||||
| REQ-026 | Cross-surface navigation: role switcher (learner/employer/admin), breadcrumbs, consistent header/footer across all surfaces | critical | 6 | pending |
|
||||
| REQ-027 | Visual consistency audit: typography scale, color palette, spacing system, dark mode toggle, accessibility baseline (WCAG AA contrast) | high | 6 | pending |
|
||||
| REQ-028 | Component library finalization: Storybook setup, component documentation, prop tables, usage examples | medium | 6 | pending |
|
||||
|
||||
---
|
||||
|
||||
@@ -172,11 +99,11 @@
|
||||
|
||||
| ID | Description | Priority | Milestone | Status |
|
||||
|----|-------------|----------|-----------|--------|
|
||||
| REQ-F-007 | Process-trace grading engine → activated as REQ-3-004 | high | v0.3 | activated |
|
||||
| REQ-F-008 | Per-learner variant task generation → activated as REQ-3-005 | high | v0.3 | activated |
|
||||
| REQ-F-009 | Oral/voice defense with AI examiner → activated as REQ-3-006 | high | v0.3 | activated |
|
||||
| REQ-F-010 | Live in-environment build with telemetry → activated as REQ-3-003 | high | v0.3 | activated |
|
||||
| REQ-F-021 | Sandbox fabric: sandboxed IDE, design tool, simulation → activated as REQ-3-001/002 | high | v0.3 | activated |
|
||||
| REQ-F-007 | Process-trace grading engine | high | v0.3+ | deferred |
|
||||
| REQ-F-008 | Per-learner variant task generation | high | v0.3+ | deferred |
|
||||
| REQ-F-009 | Oral/voice defense with AI examiner | high | v0.3+ | deferred |
|
||||
| REQ-F-010 | Live in-environment build with telemetry | high | v0.3+ | deferred |
|
||||
| REQ-F-021 | Sandbox fabric: sandboxed IDE, design tool, simulation | high | v0.3+ | deferred |
|
||||
|
||||
### Marketplace Engine
|
||||
|
||||
@@ -193,7 +120,7 @@
|
||||
|
||||
| ID | Description | Priority | Milestone | Status |
|
||||
|----|-------------|----------|-----------|--------|
|
||||
| REQ-F-017 | Identity verification and age-gating (16+/18+) — real KYC backend | high | v0.5 | activated → complete (REQ-5-003/004) |
|
||||
| REQ-F-017 | Identity verification and age-gating (16+/18+) — real KYC backend | high | v0.3+ | deferred |
|
||||
| REQ-F-018 | Payment processing and subscription management | high | v0.3+ | deferred |
|
||||
| REQ-F-019 | Human tutor marketplace (third-party courses) | medium | v0.4+ | deferred |
|
||||
| REQ-F-020 | CIRR-style placement tracking and audit | medium | v0.5+ | deferred |
|
||||
@@ -216,57 +143,22 @@
|
||||
|
||||
## Traceability Matrix
|
||||
|
||||
### v0.5 (complete)
|
||||
### v0.2 (current milestone)
|
||||
|
||||
| Requirement | Phase | Status |
|
||||
|-------------|-------|--------|
|
||||
| REQ-5-007 | 1 | complete |
|
||||
| REQ-5-001 | 2 | complete |
|
||||
| REQ-5-002 | 2 | complete |
|
||||
| REQ-5-003 | 3 | complete |
|
||||
| REQ-5-004 | 3 | complete |
|
||||
| REQ-5-005 | 4 | complete |
|
||||
| REQ-5-006 | 4 | complete |
|
||||
|
||||
### v0.4 (complete)
|
||||
|
||||
| Requirement | Phase | Status |
|
||||
|-------------|-------|--------|
|
||||
| REQ-4-001 | 1 | complete |
|
||||
| REQ-4-002 | 1 | complete |
|
||||
| REQ-4-003 | 2 | complete |
|
||||
| REQ-4-004 | 2 | complete |
|
||||
| REQ-4-005 | 3 | complete |
|
||||
|
||||
### v0.3 (complete)
|
||||
|
||||
| Requirement | Phase | Status |
|
||||
|-------------|-------|--------|
|
||||
| REQ-3-001 | 1 | complete |
|
||||
| REQ-3-002 | 1 | complete |
|
||||
| REQ-3-003 | 2 | complete |
|
||||
| REQ-3-004 | 3 | complete |
|
||||
| REQ-3-005 | 4 | complete |
|
||||
| REQ-3-006 | 5 | complete |
|
||||
| REQ-3-007 | 6 | complete |
|
||||
| REQ-3-008 | 6 | complete |
|
||||
|
||||
### v0.2 (complete)
|
||||
|
||||
| Requirement | Phase | Status |
|
||||
|-------------|-------|--------|
|
||||
| REQ-2-001 | 1 | complete |
|
||||
| REQ-2-002 | 1 | complete |
|
||||
| REQ-2-003 | 1 | complete |
|
||||
| REQ-2-004 | 2 | complete |
|
||||
| REQ-2-005 | 3 | complete |
|
||||
| REQ-2-006 | 3 | complete |
|
||||
| REQ-2-007 | 4 | complete |
|
||||
| REQ-2-008 | 4 | complete |
|
||||
| REQ-2-009 | 5 | complete |
|
||||
| REQ-2-010 | 5 | complete |
|
||||
| REQ-2-011 | 6 | complete |
|
||||
| REQ-2-012 | 6 | complete |
|
||||
| REQ-2-001 | 1 | pending |
|
||||
| REQ-2-002 | 1 | pending |
|
||||
| REQ-2-003 | 1 | pending |
|
||||
| REQ-2-004 | 2 | pending |
|
||||
| REQ-2-005 | 3 | pending |
|
||||
| REQ-2-006 | 3 | pending |
|
||||
| REQ-2-007 | 4 | pending |
|
||||
| REQ-2-008 | 4 | pending |
|
||||
| REQ-2-009 | 5 | pending |
|
||||
| REQ-2-010 | 5 | pending |
|
||||
| REQ-2-011 | 6 | pending |
|
||||
| REQ-2-012 | 6 | pending |
|
||||
|
||||
### v0.1 (complete)
|
||||
|
||||
|
||||
+141
-82
@@ -2,19 +2,13 @@
|
||||
|
||||
## Overview
|
||||
|
||||
**Milestone v0.5 — COMPLETE (shipped as v0.4.5, 2026-09-13).** Real Server Voice, KYC/Identity, Sandbox Environments, Seq-Lease: the four seams D-016 deferred out of v0.4. Server STT/TTS for voice defense (`openai-audio` VoiceProvider), identity verification + backend-enforced age-gating (REQ-F-017), design/simulation sandbox environments (REQ-F-021 remainder), and the exec-telemetry seq-lease/replay-margin fix.
|
||||
**Milestone v0.2** — AI Tutor Architecture: The six AI tutor agents (Coach, Tutor, Lab, Assessor, Proctor, Mentor) as real LLM-backed services in a new `apps/ai-service` Python FastAPI application, wired into the existing v0.1 learner surface with streaming responses. Provider-agnostic LLM layer (ollama-cloud default). Lab/Assessor/Proctor operate on mock engine inputs — their real engines are v0.3+.
|
||||
|
||||
**Milestone v0.4 — COMPLETE (shipped as v0.3.4, 2026-09-13; hotfixes v0.3.5 fresh-box, v0.3.6 single-port unattended deploy).** Distribution & Bootstrap CLI: `nextcraft` CLI (doctor/bootstrap/verify/dev) shipped as a linux x64 SEA binary, one-liner install script with checksum + version integrity gates, and binaries published on **every ongoing release** (v0.3.2 onward). v0.3.6 adds the single-port same-origin deploy (static export served by the ai-service on :8420) and unattended ops (`dev -d`/`stop`/`log`), with runtime state moved to `~/.nextcraft/`.
|
||||
**Prior milestone:** v0.1 (nextcraft-ui-prototype) — complete, shipped as v0.1.0, founder-agreed (D-013).
|
||||
|
||||
**Milestone v0.3** — Credential Engines: complete, shipped as v0.2.8 (2026-09-12). Real sandbox fabric, live build telemetry, process-trace grading, per-learner variants, oral defense, real learner surfaces.
|
||||
|
||||
**Deferred per founder directive (D-016):** REQ-F-017 identity verification + age-gating (real KYC backend), real server STT/TTS, design/simulation sandbox environments, and the exec-telemetry seq-lease are deferred to v0.5. Age-gating remains the v0.1 visual flow mockup.
|
||||
|
||||
**Prior milestone:** v0.2 (ai-tutor-architecture) — complete, shipped as v0.2.0, six tutor agents live over mock engine inputs (D-015).
|
||||
|
||||
**Milestone type:** Feature (voice provider + identity backend + new env types + ingest protocol fix)
|
||||
**Tag line:** v0.4.x (patches on the v0.4 line; milestone release as the final v0.4.x patch)
|
||||
**Branch:** milestone/v0.5-real-voice-identity-envs
|
||||
**Milestone type:** Feature (new AI service + real agent capabilities)
|
||||
**Tag line:** v0.1.x (patches on the v0.1 line; milestone release as v0.2.0)
|
||||
**Branch:** milestone/v0.2-ai-tutor-architecture
|
||||
|
||||
---
|
||||
|
||||
@@ -22,12 +16,14 @@
|
||||
|
||||
| # | Name | Status | Depends On | Requirements | Success Criteria |
|
||||
|---|------|--------|------------|--------------|------------------|
|
||||
| 0 | Pre-execution | complete | — | — | Specification, clarify, research, plan, grill complete; .ciagent/ files updated for v0.5 |
|
||||
| 1 | Seq-lease + replay margin | complete | 0 | REQ-5-007 | Ingest acks received seqs (`seq_ack` frame); capture agent trims spool to ack on reconnect (bounded margin); real-server mid-burst regression test closes the P07-documented gap |
|
||||
| 2 | Real server voice | complete | 0 | REQ-5-001, REQ-5-002 | `openai-audio` provider passes STT/TTS contract tests (MockTransport); voice defense runs server-side end-to-end when keys exist; mock/browser paths unchanged; suite green |
|
||||
| 3 | Identity + age-gating | complete | 0 | REQ-5-003, REQ-5-004 | Identity protocol + mock backend + verification flow API; gated routes enforce 16+/18+ (allowlist → identity → rate caps); PII hygiene pinned by caplog test |
|
||||
| 4 | Design/sim environments | complete | 0 | REQ-5-005, REQ-5-006 | Template-layer env registry (build/design/simulation); per-kind starter contents + exec policy; test_command surfaced; grading digest kind-agnostic (pinned) |
|
||||
| 5 | Final review + ship | complete | 1-4 | — | Code review clean; audit passes; milestone tagged (final v0.4.x patch); release with binary assets on Gitea |
|
||||
| 0 | Pre-execution | in_progress | — | — | Specification, clarify, research, plan complete; .ciagent/ files updated for v0.2 |
|
||||
| 1 | AI service scaffolding | pending | 0 | REQ-2-001, REQ-2-002, REQ-2-003 | apps/ai-service runs (uvicorn), health endpoint responds, provider-agnostic LLM client with 3 providers (ollama-cloud/local/mock), SSE streaming verified, pytest suite passes with mock provider, turbo scripts wired |
|
||||
| 2 | Agent framework | pending | 1 | REQ-2-004 | Base agent contract, session/state store, prompt templates, streaming pipeline, structured outputs; all tested |
|
||||
| 3 | Coach + Tutor agents | pending | 2 | REQ-2-005, REQ-2-006 | Coach (pacing/motivation/retrieval practice) and Tutor (concept delivery/Socratic questioning) fully implemented with system prompts, tested against mock provider, wired to chat endpoint |
|
||||
| 4 | Lab + Assessor agents | pending | 2 | REQ-2-007, REQ-2-008 | Lab consumes simulated sandbox telemetry (mock); Assessor applies rubrics to pre-baked artifacts/defense transcripts (mock); both tested |
|
||||
| 5 | Proctor + Mentor agents | pending | 2 | REQ-2-009, REQ-2-010 | Proctor produces integrity signals + coaching interventions from mock telemetry; Mentor generates long-horizon career narrative; both tested |
|
||||
| 6 | Learner surface integration | pending | 3, 4, 5 | REQ-2-011, REQ-2-012 | Learner chat streams real responses; agent routing works; byte viewer/sandbox/assessment mockups surface agent outputs; error/loading states; pnpm build + typecheck pass |
|
||||
| 7 | Final review + ship | pending | 6 | — | Code review clean; audit passes; milestone tagged v0.2.0; release created on Gitea |
|
||||
|
||||
---
|
||||
|
||||
@@ -35,86 +31,149 @@
|
||||
|
||||
### Phase 0: Pre-execution
|
||||
|
||||
**Goal:** Lock the v0.5 specification (four deferred seams), clarify ambiguities, research the audio endpoint contract + KYC provider landscape + env-type design + the seq-lease protocol, plan waves, grill adversarially.
|
||||
**Goal:** Establish v0.2 specification, clarify ambiguities, research AI service architecture, create detailed plans.
|
||||
|
||||
**Stages:** SPECIFY → CLARIFY → RESEARCH → PLAN → GRILL → MVP/UX CHECK → SHIP
|
||||
**Stages:** SPECIFY → CLARIFY → RESEARCH → PLAN → GRILL → SHIP
|
||||
|
||||
**Deliverables:**
|
||||
- Updated .ciagent/config.json, PROJECT.md, REQUIREMENTS.md, ROADMAP.md, ARCHITECTURE.md, PERSONAS.md, PLAN.md
|
||||
|
||||
**Success criteria:** All .ciagent/ files updated for v0.5; phase 0 shipped as v0.4.0.
|
||||
|
||||
### Phase 1: Seq-Lease + Replay Margin (REQ-5-007)
|
||||
|
||||
**Goal:** Close the reconnect-replay ACK gap from the v0.3 P6 lesson.
|
||||
|
||||
**Key deliverables:**
|
||||
- Ingest WS: ack frames carrying highest-contiguous received seq per (learner, task); capture agent tracks ack and resumes spool replay from acked position (bounded replay margin)
|
||||
- Real-server regression test: mid-burst kill → reconnect → no loss/dup exactly-once within margin (deterministic — waits for server-side observation like the P07 fix)
|
||||
|
||||
**Success criteria:** regression test green on repeated runs; margin documented; flood/gap semantics (G-3/G-4) unchanged.
|
||||
|
||||
### Phase 2: Real Server Voice (REQ-5-001, REQ-5-002)
|
||||
|
||||
**Goal:** The `openai-audio` VoiceProvider against OpenAI-compatible STT/TTS endpoints; voice defense real path end-to-end.
|
||||
|
||||
**Key deliverables:**
|
||||
- `ai_service/voice/openai_audio.py` — provider per D-030 protocol (STT: multipart upload → transcript; TTS: text → audio bytes); httpx via the app pool; timeouts; error mapping
|
||||
- `AI_VOICE_BASE_URL`/`AI_VOICE_API_KEY`/`AI_VOICE_MODEL` settings resolution (env-only, never committed); `AI_VOICE_PROVIDER=openai-audio|browser|mock` selection in factory
|
||||
- Defense answer route: real STT path (audio upload transcribed server-side when provider=openai-audio; browser/mock paths unchanged); examiner TTS question audio
|
||||
- Tests: MockTransport byte-contract tests (multipart shape, response parse, error modes); defense-flow tests stay mock-only (never call the cloud)
|
||||
|
||||
**Success criteria:** provider contract pinned by tests; voice defense works against a scripted audio endpoint; `pnpm ai:test` green; manual cloud probe documented.
|
||||
|
||||
### Phase 3: Identity + Age-Gating (REQ-5-003, REQ-5-004)
|
||||
|
||||
**Goal:** Real identity verification backend behind a provider protocol; backend-enforced age gates.
|
||||
|
||||
**Key deliverables:**
|
||||
- `ai_service/identity/` — IdentityProvider protocol (submit/poll/verify), mock provider (deterministic), store (SQLite, D-027 family), API router (`/v1/identity/*`)
|
||||
- Age-gate enforcement: school 16+ (enrollment), marketplace 18+ verified (gated routes reject under-age/unverified with 403) — dependency-injected gate, allowlist (G-5) retained as pilot guard
|
||||
- PII hygiene: document refs stored, never raw docs in logs; secrets-hygiene checklist extended
|
||||
|
||||
**Success criteria:** verification flow API green under mock; gated routes enforce ages in tests; zero PII in captured logs (pinned by test).
|
||||
|
||||
### Phase 4: Design/Sim Environments (REQ-5-005, REQ-5-006)
|
||||
|
||||
**Goal:** Environment kinds beyond the coding IDE on the existing namespace fabric.
|
||||
|
||||
**Key deliverables:**
|
||||
- Environment-type registry in the sandbox fabric: `build` (existing), `design` (canvas/file artifacts), `simulation` (run/benchmark harness) — per-type starter contents + allowed exec commands, one lifecycle, one telemetry path
|
||||
- Learner surface + engine-client accept environment kind; telemetry events carry kind; grading digest stays kind-agnostic (D-028 unchanged)
|
||||
|
||||
**Success criteria:** all three kinds provisionable + provable isolation; per-kind starter files land; kind flows through telemetry to the digest.
|
||||
|
||||
### Phase 5: Final Review + Ship
|
||||
|
||||
**Goal:** Code review, audit, milestone release with binary assets.
|
||||
|
||||
**Success criteria:** review P0s fixed in-phase; audit reconstruction clean; final v0.4.x patch tagged; Gitea release with `nextcraft-linux-x64` + `.sha256`; milestone merged to main; branches deleted.
|
||||
**Success criteria:** All .ciagent/ files updated for v0.2; phase 0 shipped as v0.1.1.
|
||||
|
||||
---
|
||||
|
||||
## v0.5 (Complete — Shipped as v0.4.5)
|
||||
### Phase 1: AI Service Scaffolding
|
||||
|
||||
Real Server Voice, KYC/Identity, Sandbox Environments, Seq-Lease: real server STT/TTS (`openai-audio` VoiceProvider with boot-safe fallback), the identity module (5th D-027 store, provider protocol + mock, D-043 age-gate composition on variants/sandboxes/defense/marketplace), design/simulation environment kinds at the template layer with per-kind exec policy, and the seq-ack protocol closing the P07 replay-margin gap. 6 phases (P0–P5). All 7 requirements (REQ-5-001..007) complete. Tags v0.4.0–v0.4.4 per phase, milestone release v0.4.5.
|
||||
**Goal:** Stand up apps/ai-service with the provider-agnostic LLM layer and SSE streaming.
|
||||
|
||||
## v0.4 (Complete — Shipped as v0.3.4)
|
||||
**Requirements:** REQ-2-001, REQ-2-002, REQ-2-003
|
||||
|
||||
Distribution & Bootstrap CLI: `nextcraft` CLI (doctor/bootstrap/verify/dev) as a self-contained linux x64 SEA binary, one-liner install with sha256 + version integrity gates, release-asset pipeline attaching binaries to every ongoing release (v0.3.2 onward), install/quickstart docs backed by a fresh-clone E2E test. 5 phases (P0–P4). All 5 requirements (REQ-4-001..005) complete. Tags v0.3.0–v0.3.3 per phase, milestone release v0.3.4.
|
||||
**Key deliverables:**
|
||||
- apps/ai-service: FastAPI app, pydantic-settings, uvicorn, /health, CORS for localhost
|
||||
- llm package: provider interface + ollama-cloud/local/mock providers; key resolution from .ciagent/.env.secrets via env
|
||||
- SSE streaming: /v1/chat/stream endpoint streaming provider deltas
|
||||
- pytest suite with mock provider; root scripts: ai:dev, ai:test; turbo integration
|
||||
|
||||
## v0.3 (Complete — Shipped as v0.2.8)
|
||||
**Success criteria:**
|
||||
- `python -m uvicorn` starts the service; /health returns 200
|
||||
- Provider unit tests pass (mock); ollama-cloud integration probe works (manual)
|
||||
- SSE stream delivers tokens to an HTTP client
|
||||
|
||||
Credential Engines: real sandbox fabric (Linux namespaces), live build telemetry
|
||||
(at-least-once/exactly-once), process-trace grading (G-4 gated), seeded per-learner
|
||||
variants (fairness anchors wired to grading), oral defense with integrity signals
|
||||
(mock-first voice, browser fallback), and real learner build/defense/grading surfaces.
|
||||
8 phases. All 8 requirements (REQ-3-001..008) complete. Tags v0.2.1–v0.2.7 per phase,
|
||||
milestone release v0.2.8.
|
||||
---
|
||||
|
||||
## v0.2 (Complete — Shipped as v0.2.0)
|
||||
### Phase 2: Agent Framework
|
||||
|
||||
AI Tutor Architecture: Six AI tutor agents (Coach, Tutor, Lab, Assessor, Proctor, Mentor) as real LLM-backed services over mock engine inputs, wired into the learner surface with streaming. 7 phases. All 12 requirements complete. Milestone release v0.2.0.
|
||||
**Goal:** Build the shared framework all six agents use.
|
||||
|
||||
**Requirements:** REQ-2-004
|
||||
|
||||
**Key deliverables:**
|
||||
- BaseAgent contract: system prompt, message history, streaming completion, structured output
|
||||
- Session/state store: in-memory per-learner session with message history
|
||||
- Prompt management: per-agent system prompt templates with learner context injection
|
||||
- Streaming pipeline: agent → provider → SSE with agent identification
|
||||
- Structured outputs: JSON-schema outputs for Assessor rubric scores, Proctor signals
|
||||
|
||||
**Success criteria:**
|
||||
- BaseAgent unit tests pass
|
||||
- Session store tested (create/append/persist in-memory)
|
||||
- Structured output parsing tested against mock provider
|
||||
|
||||
---
|
||||
|
||||
### Phase 3: Coach + Tutor Agents
|
||||
|
||||
**Goal:** Implement the two learner-facing conversational agents.
|
||||
|
||||
**Requirements:** REQ-2-005, REQ-2-006
|
||||
|
||||
**Key deliverables:**
|
||||
- Coach agent: pacing guidance, motivation, retrieval practice prompts; distinct persona
|
||||
- Tutor agent: concept delivery, Socratic questioning, worked examples
|
||||
- Agent registry: route chat messages to the correct agent by context/selection
|
||||
- Per-agent system prompts with competency-stack context injection from packages/mock-data
|
||||
|
||||
**Success criteria:**
|
||||
- Both agents produce distinct, on-persona responses (verified against mock + ollama-cloud)
|
||||
- Agent routing tested
|
||||
- Both agents exposed via the chat streaming endpoint
|
||||
|
||||
---
|
||||
|
||||
### Phase 4: Lab + Assessor Agents
|
||||
|
||||
**Goal:** Implement the two build/assessment agents over mock engine inputs.
|
||||
|
||||
**Requirements:** REQ-2-007, REQ-2-008
|
||||
|
||||
**Key deliverables:**
|
||||
- Lab agent: consumes simulated sandbox telemetry (mock event streams), produces in-flow feedback
|
||||
- Assessor agent: applies rubrics to pre-baked artifacts and defense transcripts, returns structured scores + feedback
|
||||
- Mock engine inputs: simulated telemetry generator, pre-baked artifact corpus in packages/mock-data
|
||||
- Endpoints: /v1/lab/feedback, /v1/assessment/evaluate
|
||||
|
||||
**Success criteria:**
|
||||
- Lab produces relevant feedback for mock telemetry scenarios
|
||||
- Assessor returns structured rubric scores (JSON) for pre-baked artifacts
|
||||
- Both tested against mock provider
|
||||
|
||||
---
|
||||
|
||||
### Phase 5: Proctor + Mentor Agents
|
||||
|
||||
**Goal:** Implement the integrity and narrative agents.
|
||||
|
||||
**Requirements:** REQ-2-009, REQ-2-010
|
||||
|
||||
**Key deliverables:**
|
||||
- Proctor agent: integrity signals from mock telemetry (tab switches, idle time, paste events), coaching interventions
|
||||
- Mentor agent: long-horizon career narrative, competency-stack progression guidance
|
||||
- Endpoints: /v1/proctor/signals, /v1/mentor/narrative
|
||||
|
||||
**Success criteria:**
|
||||
- Proctor produces classified signals with recommended interventions for mock scenarios
|
||||
- Mentor produces coherent career-narrative responses
|
||||
- Both tested against mock provider
|
||||
|
||||
---
|
||||
|
||||
### Phase 6: Learner Surface Integration
|
||||
|
||||
**Goal:** Wire the v0.1 learner surface to the real AI service.
|
||||
|
||||
**Requirements:** REQ-2-011, REQ-2-012
|
||||
|
||||
**Key deliverables:**
|
||||
- Learner dashboard chat: real streaming via SSE, agent switcher (Coach/Tutor), error/loading states
|
||||
- Byte tutorial viewer: Tutor concept explanations
|
||||
- Build sandbox: Lab feedback panel fed by mock telemetry + Lab agent
|
||||
- Assessment mockup: Assessor rubric output display, Proctor integrity banner
|
||||
- Mentor panel on learner dashboard
|
||||
|
||||
**Success criteria:**
|
||||
- Streaming chat works end-to-end with ai-service running
|
||||
- All four learner surfaces surface agent outputs
|
||||
- Graceful degradation when ai-service is down (error states, not crashes)
|
||||
- `pnpm build` and `pnpm typecheck` pass
|
||||
|
||||
---
|
||||
|
||||
### Phase 7: Final Review + Ship
|
||||
|
||||
**Goal:** Code review, audit, milestone release.
|
||||
|
||||
**Key deliverables:**
|
||||
- Multi-persona code review (correctness, testing, security, performance, maintainability)
|
||||
- Project health audit (reconstruction test, .ciagent/ file discipline, branch hygiene, commit discipline)
|
||||
- Milestone ship: merge milestone → main, tag v0.2.0, create Gitea release
|
||||
|
||||
**Success criteria:**
|
||||
- Code review: P0 fixes applied, P1+ documented
|
||||
- Audit: all checks pass, project state reconstructable from git log
|
||||
- Ship: v0.2.0 tagged, milestone branch merged to main, Gitea release created — **release note explicitly states Lab/Assessor/Proctor operate on mock engine inputs (real engines v0.3+)** (G-5); dead `aiTutorResponses` export disposed of (G-5)
|
||||
- All 12 v0.2 requirements marked complete
|
||||
|
||||
---
|
||||
|
||||
## v0.1 (Complete — Shipped as v0.1.0)
|
||||
|
||||
|
||||
@@ -9,7 +9,7 @@
|
||||
},
|
||||
"release": {
|
||||
"forge": "gitea",
|
||||
"base_url": "https://git.coreci.dev",
|
||||
"base_url": "https://git.cloudinit.dev",
|
||||
"owner": "coreci",
|
||||
"repo": "nextcraft"
|
||||
},
|
||||
@@ -46,9 +46,9 @@
|
||||
"projects": [],
|
||||
"active_project": null,
|
||||
"milestone": {
|
||||
"version": "v0.5",
|
||||
"name": "real-voice-identity-envs",
|
||||
"version": "v0.2",
|
||||
"name": "ai-tutor-architecture",
|
||||
"type": "feature",
|
||||
"branch": "milestone/v0.5-real-voice-identity-envs"
|
||||
"branch": "milestone/v0.2-ai-tutor-architecture"
|
||||
}
|
||||
}
|
||||
+1
-14
@@ -20,8 +20,6 @@ dist/
|
||||
.next/
|
||||
.turbo/
|
||||
*.tsbuildinfo
|
||||
# Next.js static export (v0.3.6 build output, served by the ai-service)
|
||||
apps/web/out/
|
||||
|
||||
# Storybook
|
||||
storybook-static/
|
||||
@@ -48,15 +46,4 @@ coverage/
|
||||
|
||||
# Python tooling caches
|
||||
.pytest_cache/
|
||||
.ruff_cache/
|
||||
*.egg-info/
|
||||
.ciagent/bin/
|
||||
|
||||
# v0.3 engine runtime data (SQLite + sandboxes)
|
||||
apps/ai-service/ai_service/data/
|
||||
apps/ai-service/**/sandboxes/
|
||||
*.db
|
||||
*.db-journal
|
||||
|
||||
# in-sandbox capture agent runtime spool
|
||||
.nc-agent/
|
||||
.ruff_cache/
|
||||
@@ -2,86 +2,8 @@
|
||||
|
||||
AI-native outcome school + marketplace — graduates prove what they can build, not what they can write.
|
||||
|
||||
## Quickstart
|
||||
|
||||
One-liner install (linux x64):
|
||||
|
||||
```sh
|
||||
curl -fsSL https://git.coreci.dev/coreci/nextcraft/raw/main/scripts/install.sh | sh
|
||||
```
|
||||
|
||||
That downloads the latest release's `nextcraft` CLI binary, verifies its sha256 checksum, and installs it to `~/.local/bin` (PATH hint printed if needed). Every release ships fresh binaries — re-run the one-liner to upgrade.
|
||||
|
||||
Then, from a clone of this repo:
|
||||
|
||||
```sh
|
||||
nextcraft doctor # check prerequisites: node >= 18, pnpm >= 8, python3 >= 3.11, git, unshare
|
||||
nextcraft bootstrap # pnpm install + ai-service venv + .env from template (idempotent)
|
||||
nextcraft verify # health check: venv imports, uvicorn, ports, env
|
||||
nextcraft dev # run the ai-service dev server on :8420 (web dev server: pnpm dev)
|
||||
```
|
||||
|
||||
The E2E test (`apps/cli/tests/fresh-clone-e2e.test.ts`) proves this exact sequence on a fresh clone.
|
||||
|
||||
### No binary / non-linux?
|
||||
|
||||
The installer degrades to printed source instructions. Manual equivalent:
|
||||
|
||||
```sh
|
||||
git clone https://git.coreci.dev/coreci/nextcraft.git && cd nextcraft
|
||||
pnpm install
|
||||
bash apps/ai-service/scripts/bootstrap.sh
|
||||
cp apps/ai-service/.env.example apps/ai-service/.env
|
||||
pnpm ai:dev
|
||||
```
|
||||
|
||||
## CLI reference (`nextcraft`)
|
||||
|
||||
| Command | What it does | Exit codes |
|
||||
|---------|--------------|------------|
|
||||
| `doctor` | Checks prerequisites on PATH: node >= 18, pnpm >= 8, python3 >= 3.11, git, unshare (sandbox fabric). Every ✗ prints a fix hint. | 0 all pass, 1 any fail |
|
||||
| `bootstrap` | Sets up the monorepo from a fresh clone: (1) locates the repo root, (2) `pnpm install`, (3) ai-service venv via `apps/ai-service/scripts/bootstrap.sh`, (4) copies `.env.example` → `.env` if absent, (5) warns on missing optional keys. Idempotent — safe to re-run. | 0 ok, 1 step failed |
|
||||
| `verify` | Health check: ai-service venv + `import ai_service`, uvicorn importable, `.env` present (warn-only), `AI_PORT` (default 8420) free, workspace `node_modules` present. | 0 ok, 1 failures |
|
||||
| `dev` | Thin passthrough to `apps/ai-service/scripts/dev.sh` (exports secrets from `.ciagent/.env.secrets` if present, runs uvicorn on :8420). Ctrl+C stops it. The web dev server is separate: `pnpm dev`. | child's exit code |
|
||||
| `dev -d` / `dev --detach` | Same, but as an **unattended daemon**: detached, output appended to `~/.nextcraft/run/<clone>/dev.log`, pidfile beside it. Refuses to double-start. Auto-serves the built web UI same-origin when `apps/web/out` exists. | 0 started, 1 already running |
|
||||
| `stop` | Stops the daemon started by `dev -d` (SIGTERM, SIGKILL after 5s), cleans the pidfile. | 0 stopped/clean, 1 kill failed |
|
||||
| `log` | Tails the daemon log: last 50 lines by default, `-n N` for more, `-f`/`--follow` to stream. | 0 ok, 1 no log yet |
|
||||
| `--help` / `-h` | Usage for the CLI or any command. | 0 |
|
||||
| `--version` | Prints the version this binary was built as (matches the release tag). | 0 |
|
||||
|
||||
Exit-code contract: `0` success, `1` check/step failure (hint printed), `2` usage error.
|
||||
|
||||
### Remote server / single-port deployment
|
||||
|
||||
`nextcraft dev` binds the API on **0.0.0.0:8420** (and `pnpm dev` serves the web app on all interfaces), so the stack works from other machines out of the box:
|
||||
|
||||
- Browse `http://<your-host>:3000` — the web app targets `http://<your-host>:8420` automatically (derived from the browser's hostname).
|
||||
- CORS admits any origin (`AI_CORS_ORIGINS=*` in `apps/ai-service/.env`). This is safe **only** because credentials are never enabled; to restrict, set an explicit list: `AI_CORS_ORIGINS=http://<your-host>:3000`.
|
||||
- To revert to loopback-only: `AI_HOST=127.0.0.1` in `apps/ai-service/.env`.
|
||||
- Security note: this is an unauthenticated dev API reachable from any network the box exposes. Mitigations that still apply: per-learner sandbox caps + global rate caps + learner allowlist (G-5), telemetry flood control (traces marked `INCOMPLETE_FLOODED` are refused by the grader). Expose only on trusted networks until identity/KYC lands (v0.5).
|
||||
|
||||
**Production behind one port (v0.3.6):** when only one port is reachable (e.g. behind HAProxy), build the web app as a static export and let the ai-service serve it same-origin on :8420 — UI and API on one port, no CORS, no mixed content:
|
||||
|
||||
```sh
|
||||
curl -fsSL https://git.coreci.dev/coreci/nextcraft/raw/main/scripts/install.sh | sh # fresh binary
|
||||
git pull --ff-only && pnpm install
|
||||
NEXT_PUBLIC_AI_SERVICE_URL=self pnpm build # emits apps/web/out
|
||||
nextcraft dev -d # daemon; auto-serves the UI from :8420
|
||||
nextcraft log -f # tail the daemon log
|
||||
```
|
||||
|
||||
Durable state (SQLite DB, sandbox workdirs, daemon pid/log) lives in `~/.nextcraft/` — never inside the repo. Full recipe incl. HAProxy timeouts and a systemd unit: [deploy/README.md](deploy/README.md).
|
||||
|
||||
## Docs
|
||||
|
||||
- [apps/cli/README.md](apps/cli/README.md) — CLI internals: build, binary pipeline, troubleshooting
|
||||
- [.ciagent/PROJECT.md](.ciagent/PROJECT.md) — product spec and milestone history
|
||||
- [.ciagent/ARCHITECTURE.md](.ciagent/ARCHITECTURE.md) — system architecture
|
||||
|
||||
## Status
|
||||
|
||||
**Milestone v0.5** — Real Server Voice, KYC/Identity, Sandbox Environments, Seq-Lease (shipped as v0.4.5)
|
||||
|
||||
Prior: v0.3 Credential Engines (shipped v0.2.8) · v0.2 AI Tutor Architecture (v0.2.0) · v0.1 UI/UX Prototype (v0.1.0)
|
||||
**Milestone v0.1** — UI/UX Prototype (high-fidelity interactive, all mock data)
|
||||
|
||||
Initialized via CIAgent v0.7.0
|
||||
@@ -1,76 +0,0 @@
|
||||
# Nextcraft AI Service — environment template (copy values, never commit real keys)
|
||||
# Real keys live in .ciagent/.env.secrets (gitignored) and are exported by scripts/dev.sh
|
||||
|
||||
AI_PORT=8420
|
||||
# Network mode (v0.3.5, D-038): dev server binds 0.0.0.0 so remote machines can
|
||||
# reach the stack. Set to 127.0.0.1 to revert to loopback-only.
|
||||
AI_HOST=0.0.0.0
|
||||
# CORS + WS-origin policy: '*' (default) admits any origin — safe because
|
||||
# credentials are never enabled. Restrict with a comma list, e.g.:
|
||||
# AI_CORS_ORIGINS=http://nextcraft-1:3000
|
||||
AI_CORS_ORIGINS=*
|
||||
AI_PROVIDER=ollama-cloud
|
||||
AI_MODEL=gemma4:31b
|
||||
AI_OLLAMA_CLOUD_BASE_URL=https://ollama.com/v1
|
||||
AI_OLLAMA_CLOUD_API_KEY=
|
||||
AI_LOCAL_BASE_URL=http://localhost:11434/v1
|
||||
AI_JSON_MODE=auto
|
||||
|
||||
# Sandbox fabric (v0.3). v0.3.6: durable state defaults to ~/.nextcraft/ —
|
||||
# OUTSIDE the repo (the old repo-relative 'sandboxes' default polluted the
|
||||
# git tree). Set an absolute path (~/ works) or a repo-relative one only for
|
||||
# throwaway dev clones.
|
||||
AI_SANDBOX_MAX_CONCURRENT=5
|
||||
AI_SANDBOX_TIMEOUT_S=900
|
||||
AI_SANDBOX_MAX_WORKDIR_MB=512
|
||||
|
||||
# G-5 abuse control (NOT auth — KYC/identity deferred):
|
||||
# comma-separated learner allowlist; unknown ids can't create sandboxes (403)
|
||||
AI_LEARNER_ALLOWLIST=pilot-learner
|
||||
# max ACTIVE sandboxes per learner → 429 when exceeded
|
||||
AI_SANDBOX_MAX_PER_LEARNER=1
|
||||
# global creates per rolling 60s window (in-memory) → 429 when exceeded
|
||||
AI_SANDBOX_CREATES_PER_MIN=10
|
||||
|
||||
# Persistence (SQLite). v0.3.6: default moved out of the repo to
|
||||
# ~/.nextcraft/data/nextcraft.db (~/ paths in env overrides are expanded).
|
||||
# Uncomment + edit ONLY to relocate:
|
||||
# AI_DB_PATH=~/.nextcraft/data/nextcraft.db
|
||||
|
||||
# Single-port deploy (v0.3.6): directory of the exported web app
|
||||
# (apps/web/out). When set, the UI is served from this same service on :8420
|
||||
# — build it with `NEXT_PUBLIC_AI_SERVICE_URL=self pnpm build`. The
|
||||
# `nextcraft dev` command auto-sets this when apps/web/out exists.
|
||||
# AI_WEB_STATIC_DIR=../../web/out
|
||||
# --- Identity (REQ-5-003, D-042) ---
|
||||
# 'mock' (default — deterministic, no vendor spend pre-pilot; verdicts carry
|
||||
# mock=True forever per A-304). A real KYC vendor drops in via the
|
||||
# IdentityProvider protocol without API changes.
|
||||
AI_IDENTITY_PROVIDER=mock
|
||||
# G-13: identity submit caps — one active pending per learner (409), and a
|
||||
# per-learner submit rate ceiling (429 over a rolling 60s window).
|
||||
AI_IDENTITY_SUBMITS_PER_MIN=3
|
||||
|
||||
# --- Voice (REQ-3-006 D-030; real server path REQ-5-001, D-040) ---
|
||||
# 'mock' (default; no key needed — tests/dev), 'browser' (client-native
|
||||
# SR/TTS), or 'openai-audio' (real server STT/TTS, live since v0.5).
|
||||
AI_VOICE_PROVIDER=mock
|
||||
# openai-audio requires BOTH (unconfigured → app boots, voice falls back to
|
||||
# mock with a loud log — G-11; the badge then honestly reports mock):
|
||||
# AI_VOICE_BASE_URL=https://your-audio-endpoint/v1
|
||||
# AI_VOICE_API_KEY=
|
||||
# Optional model/voice/format knobs (defaults shown):
|
||||
# AI_VOICE_STT_MODEL=whisper-1
|
||||
# AI_VOICE_TTS_MODEL=tts-1
|
||||
# AI_VOICE_TTS_VOICE=alloy
|
||||
# AI_VOICE_TTS_FORMAT=mp3 (enum: mp3 | wav | opus)
|
||||
# AI_VOICE_MAX_AUDIO_MB=10
|
||||
# Manual probe recipe (executable by anyone with the keys; never CI):
|
||||
# STT: curl -sS $AI_VOICE_BASE_URL/audio/transcriptions \
|
||||
# -H "Authorization: Bearer $AI_VOICE_API_KEY" \
|
||||
# -F file=@test/fixtures/answer.wav -F model=whisper-1 | jq -e '.text'
|
||||
# TTS: curl -sS $AI_VOICE_BASE_URL/audio/speech \
|
||||
# -H "Authorization: Bearer $AI_VOICE_API_KEY" \
|
||||
# -H 'Content-Type: application/json' \
|
||||
# -d '{"model":"tts-1","input":"Nextcraft","voice":"alloy"}' \
|
||||
# -o /tmp/probe.mp3 && file /tmp/probe.mp3 | grep -i audio
|
||||
@@ -1,268 +0,0 @@
|
||||
# Nextcraft AI Service (`apps/ai-service`)
|
||||
|
||||
Python FastAPI service hosting the six AI tutor agents (Coach, Tutor, Lab, Assessor, Proctor, Mentor) behind a provider-agnostic LLM layer. Port **8420**.
|
||||
|
||||
## Quickstart
|
||||
|
||||
```bash
|
||||
# 1. Bootstrap (idempotent): venv + deps
|
||||
bash scripts/bootstrap.sh
|
||||
|
||||
# 2. Run tests (mock provider only — zero network calls)
|
||||
bash scripts/test.sh
|
||||
|
||||
# 3. Lint
|
||||
bash scripts/lint.sh
|
||||
|
||||
# 4. Dev server (exports keys from .ciagent/.env.secrets if present)
|
||||
bash scripts/dev.sh
|
||||
```
|
||||
|
||||
Or via the monorepo root (`corepack pnpm install` first):
|
||||
|
||||
```bash
|
||||
pnpm ai:bootstrap
|
||||
pnpm ai:test
|
||||
pnpm ai:lint
|
||||
pnpm ai:dev
|
||||
```
|
||||
|
||||
## Configuration
|
||||
|
||||
All settings use the `AI_` env prefix (pydantic-settings; see `.env.example`).
|
||||
|
||||
| Var | Default | Purpose |
|
||||
|-----|---------|---------|
|
||||
| `AI_PORT` | 8420 | Listen port |
|
||||
| `AI_PROVIDER` | mock | `ollama-cloud` \| `local` \| `mock` |
|
||||
| `AI_MODEL` | gemma4:31b | Model for all agents |
|
||||
| `AI_OLLAMA_CLOUD_BASE_URL` | https://ollama.com/v1 | Cloud base URL |
|
||||
| `AI_OLLAMA_CLOUD_API_KEY` | (empty) | Bearer key — **never commit** |
|
||||
| `AI_JSON_MODE` | auto | `auto` sends response_format, degrades on 400; `off` never sends |
|
||||
|
||||
Tests run with `AI_PROVIDER=mock` (enforced in `tests/conftest.py` by an instance assertion) — the suite never calls the cloud.
|
||||
|
||||
## Endpoints
|
||||
|
||||
- `GET /health` — status, configured provider, model (no cloud call)
|
||||
- `POST /v1/chat/stream` — SSE chat stream. Body: `{"agent": "coach"|"tutor", "session_id": "...", "messages": [{"role":"user","content":"..."}]}`. Unknown agents are rejected with 422.
|
||||
|
||||
SSE envelope (D-016): `meta` event first (agent/session/model), then `delta` events (incremental content), then `done`; on mid-stream failure an `error` event precedes the terminal `[DONE]` sentinel. sse-starlette emits `: ping` keep-alive comment lines on idle connections — clients must ignore frames without `data:`.
|
||||
|
||||
## Manual ollama-cloud persona probe (Phase 3, documented — not automated)
|
||||
|
||||
With the real provider, Coach and Tutor must produce distinct on-persona
|
||||
responses to the same prompt:
|
||||
|
||||
```bash
|
||||
# start with the cloud provider (keys exported from .ciagent/.env.secrets)
|
||||
AI_PROVIDER=ollama-cloud .venv/bin/uvicorn ai_service.main:app --port 8420
|
||||
|
||||
# Coach: expect pacing + one concrete next action + a retrieval-practice question
|
||||
curl -sN -X POST localhost:8420/v1/chat/stream -H 'Content-Type: application/json' \
|
||||
-d '{"agent":"coach","session_id":"probe-coach","messages":[{"role":"user","content":"I am stuck on multi-agent communication patterns"}]}' \
|
||||
| grep '^data:'
|
||||
|
||||
# Tutor: expect ONE concept + a worked example + a Socratic check question
|
||||
curl -sN -X POST localhost:8420/v1/chat/stream -H 'Content-Type: application/json' \
|
||||
-d '{"agent":"tutor","session_id":"probe-tutor","messages":[{"role":"user","content":"I am stuck on multi-agent communication patterns"}]}' \
|
||||
| grep '^data:'
|
||||
```
|
||||
|
||||
Verify: the two responses have visibly different voice/structure (Coach:
|
||||
action + accountability; Tutor: concept + example + question). The
|
||||
automated suite never calls the cloud — distinctness is enforced against
|
||||
the deterministic mock (distinct system prompts → distinct hash-seeded
|
||||
outputs).
|
||||
|
||||
## Sandbox isolation (v0.3)
|
||||
|
||||
The v0.3 code-execution sandbox runs learner/agent code in a Linux **user
|
||||
namespace** (`unshare --user --map-root-user --mount --pid --fork --net`): the
|
||||
child is uid 0 *inside* the userns (mapped to the unprivileged host uid), gets
|
||||
a private mount + PID + network namespace, and uses `RLIMIT_*` for resource
|
||||
caps. No containers, no sudo — see "Why not containers" below.
|
||||
|
||||
### A-101 / D-024 isolation probe transcript
|
||||
|
||||
Verbatim output captured on the CI box (Linux, uid 1001 `opencode`, no
|
||||
docker/podman/bwrap; `iproute2` absent so interface state is read from
|
||||
kernel sockets + `/proc/net/dev`).
|
||||
|
||||
**1. Root-in-userns, uid 0, isolated namespaces:**
|
||||
|
||||
```console
|
||||
$ unshare --user --map-root-user --mount --pid --fork --net id -u
|
||||
0
|
||||
```
|
||||
|
||||
**2. Fresh netns has exactly one interface: `lo` only (no eth0, no route
|
||||
out).** Host baseline for contrast:
|
||||
|
||||
```console
|
||||
$ unshare --user --map-root-user --mount --pid --fork --net \
|
||||
python3 -c "import socket; print(socket.if_nameindex())"
|
||||
[(1, 'lo')]
|
||||
|
||||
$ unshare --user --map-root-user --mount --pid --fork --net \
|
||||
awk 'NR>2{print $1}' /proc/net/dev
|
||||
lo:
|
||||
|
||||
$ python3 -c "import socket; print(socket.if_nameindex())" # host
|
||||
[(1, 'lo'), (2, 'eth0')]
|
||||
```
|
||||
|
||||
(Note: `/sys/class/net` shows host interfaces even inside the netns because
|
||||
`sysfs` here is not netns-aware — the socket-level view above is the
|
||||
authoritative kernel evidence: 1 interface, loopback only, zero rx bytes, no
|
||||
carrier to any external link.)
|
||||
|
||||
**3. Write containment — writes inside the sandbox workdir are visible on the
|
||||
host under the sandbox dir, owned by the real (unprivileged) host uid:**
|
||||
|
||||
```console
|
||||
$ unshare --user --map-root-user --mount --pid --fork --net bash -c "
|
||||
mkdir -p /tmp/demo-work && cd /tmp/demo-work
|
||||
echo 'hello-from-inside-sandbox (uid=0 in-ns)' > contained.txt
|
||||
id -u"
|
||||
0
|
||||
|
||||
$ cat /tmp/demo-work/contained.txt # host
|
||||
hello-from-inside-sandbox (uid=0 in-ns)
|
||||
$ ls -la /tmp/demo-work/contained.txt # host
|
||||
-rw-r--r-- 1 opencode opencode 40 ... /tmp/demo-work/contained.txt
|
||||
```
|
||||
|
||||
The in-userns "root" writes land on the host filesystem as uid 1001
|
||||
(`opencode`) — the uid-mapping is doing the confinement; nothing escapes the
|
||||
sandbox workdir as any other identity.
|
||||
|
||||
**4. `/proc` remount is NOT permitted in this context — probe + exact error:**
|
||||
|
||||
```console
|
||||
$ unshare --user --map-root-user --mount --pid --fork --net \
|
||||
bash -c "mount -t proc proc /proc"
|
||||
mount: /proc: permission denied.
|
||||
dmesg(1) may have more information after failed mount system call.
|
||||
(exit 32)
|
||||
```
|
||||
|
||||
`mount -t proc` fails even with in-ns "root" because `/proc` is owned by a
|
||||
userns that does not contain our uid mapping (the box's `/` is itself
|
||||
owned by `nobody:nogroup` — we're already inside a container). **This is
|
||||
acceptable for v0.3**: the sandbox does not depend on a custom `/proc` view;
|
||||
the child sees the host `/proc` read-only-ish view which is already filtered
|
||||
by the pid namespace (only in-ns pids are visible). The pidns itself is what
|
||||
provides process isolation, not the proc remount.
|
||||
|
||||
### Locked resource-limit mechanism (G-1 / G-2)
|
||||
|
||||
Resource enforcement is settled for v0.3 — this is the locked decision:
|
||||
|
||||
| Resource | Mechanism | Notes |
|
||||
|----------|-----------|-------|
|
||||
| **Memory** | `RLIMIT_AS` (address space) | setrlimit in the child pre-exec; deterministic, no cgroup needed |
|
||||
| **CPU** | `RLIMIT_CPU` | kernel SIGKILL at the cpu-seconds ceiling |
|
||||
| **Single-file size** | `RLIMIT_FSIZE` | catches runaway single-file writes |
|
||||
| **Wall clock** | **manager reaper kill** (parent watchdog) | RLIMIT_CPU doesn't cover sleeping/idle children; the manager kills the sandbox on wall-clock timeout |
|
||||
| **Per-sandbox process count** | `RLIMIT_NPROC` | ⚠️ **SHARED at the host uid, not per-sandbox** — the counter is per-real-uid across all of that uid's process trees, so two concurrent sandboxes share the same NPROC budget. Accepted v0.3 gap: without cgroup delegation there's no per-sandbox pid cap; mitigations are (a) the manager serializes sandbox runs and (b) NPROC is still a hard fork-bomb ceiling. |
|
||||
| **Hard disk quota** | **NOT kernel-enforceable** | ⚠️ without cgroup delegation or sudo (`quotactl`, project quotas) there is no kernel-enforced per-sandbox disk cap. Accepted v0.3 gap. **Mitigation: a manager-side workdir-size sweep** — after each run (and on a periodic reaper pass) the manager walks the sandbox workdir and enforces `AI_SANDBOX_MAX_WORKDIR_MB` (**default 512 MB**); oversized dirs are reaped. Combined with `RLIMIT_FSIZE` this bounds disk growth between sweeps. |
|
||||
|
||||
Both accepted gaps (shared NPROC, no kernel disk quota) are documented here as
|
||||
v0.3 scope boundaries; closing them requires cgroup v2 delegation or sudo,
|
||||
neither of which is available in the target environment.
|
||||
|
||||
### Why not containers
|
||||
|
||||
Container runtimes / privileged wrapper tools are probed-and-absent on the
|
||||
box, and we have no `sudo`:
|
||||
|
||||
```console
|
||||
$ for cmd in docker podman bwrap firejail; do
|
||||
printf '%-8s: ' "$cmd"; command -v "$cmd" || echo MISSING
|
||||
done; printf '%-8s: ' sudo; command -v sudo || echo MISSING
|
||||
docker : MISSING
|
||||
podman : MISSING
|
||||
bwrap : MISSING
|
||||
firejail: MISSING
|
||||
sudo : MISSING
|
||||
$ id -u
|
||||
1001
|
||||
```
|
||||
|
||||
Unprivileged user namespaces are on the box's kernel and need neither a
|
||||
daemon, nor suid helpers, nor network access — they are the only isolation
|
||||
primitive that works here, so that's what v0.3 uses.
|
||||
|
||||
## Telemetry delivery semantics (v0.3, REQ-3-003)
|
||||
|
||||
Delivery is **at-least-once**; storage is **exactly-once** — the two compose:
|
||||
|
||||
- The in-sandbox capture agent (stdlib-only, `scripts/sandbox-agent.py`)
|
||||
spools every event to a durable JSONL file (fsync per append) BEFORE any
|
||||
send attempt, so no event can be lost to a dead socket or a SIGKILL.
|
||||
- The WS ingest endpoint (`WS /v1/telemetry/ingest?learner_id&task_id`,
|
||||
D-026) dedups server-side on the `(learner_id, task_id, seq)` primary key:
|
||||
re-sends (reconnect flushes, replay margin) are collapsed, never upserted.
|
||||
- On disconnect the agent reconnects with exponential backoff and flushes
|
||||
the spool in `seq` order; a transient outage therefore loses nothing and
|
||||
stores each event exactly once (`tests/telemetry/test_durability.py`
|
||||
proves this end-to-end against a real namespace sandbox + live server).
|
||||
- Replay/read path: `GET /v1/telemetry/traces/{learner}/{task}` returns the
|
||||
complete ordered trace; `GET /v1/telemetry/gaps/{learner}/{task}` returns
|
||||
missing seqs for gap detection.
|
||||
- Flood boundary (G-3): a connection exceeding `AI_TELEMETRY_MAX_EVENTS_PER_TASK`
|
||||
(default 50,000) is closed with WS code 1008 and its trace is marked
|
||||
`INCOMPLETE_FLOODED` — a terminal integrity flag the grader refuses to
|
||||
grade. Silent event dropping is forbidden: it would corrupt grading input.
|
||||
|
||||
## Voice defense (v0.3, REQ-3-006)
|
||||
|
||||
Voice is **mock-first** (D-030): the defense pipeline is fully proven over
|
||||
the deterministic `MockVoiceProvider` + browser-native fallback — no task
|
||||
requires a real voice key. Real server STT/TTS (`OpenAIAudioProvider` over
|
||||
OpenAI-compatible `/audio/transcriptions` + `/audio/speech`) is **deferred
|
||||
to v0.4** together with KYC (GRILL CUT-1 / G-7): it could never be exercised
|
||||
in CI, so v0.3 ships the protocol seam instead of an unverifiable claim.
|
||||
|
||||
- `AI_VOICE_PROVIDER=mock` (default) — deterministic canned STT/TTS
|
||||
- `AI_VOICE_PROVIDER=browser` — the web client uses SpeechRecognition +
|
||||
speechSynthesis; the server keeps text-turn persistence
|
||||
- Conversational budget: a defense turn should complete in **< 4s**
|
||||
(`DEFENSE_TURN_BUDGET_MS` in `tests/voice/test_latency.py`). v0.3
|
||||
asserts instrumentation (stt_ms/llm_ms/tts_ms populated per turn); the
|
||||
wall-clock acceptance probe against a real voice endpoint is a v0.4
|
||||
criterion, run manually with `AI_VOICE_PROVIDER` set to the real
|
||||
provider and keys in `.ciagent/.env.secrets` (never in code/commits).
|
||||
|
||||
## End-to-end credential flow (v0.3, REQ-3-007/008)
|
||||
|
||||
`tests/api/test_e2e_credential_flow.py` runs the full pipeline against a REAL
|
||||
uvicorn server with REAL namespace sandboxes (mock LLM/voice per G-2
|
||||
precedent): variant -> telemetry-wired sandbox -> in-sandbox exec -> trace
|
||||
persistence -> process-trace grade (variant seed stamped) -> assessor
|
||||
coaching -> oral defense -> verdict + integrity signals -> proctor. It
|
||||
asserts no corpus fixture appears anywhere in the learner path.
|
||||
|
||||
Manual browser pass (documented, not automated): `pnpm ai:dev` + `pnpm dev`,
|
||||
then open `/build/stack-orchestration-c007` — variant statement + starter
|
||||
files load, edit a file, Run/Test execute in the sandbox with output in the
|
||||
read-only panel, the telemetry status pulses, Lab streams feedback from the
|
||||
live digest; then `/defend/stack-orchestration-c007` — Start Defense, typed
|
||||
answers (mic path needs permission), Finish, Grade My Work renders the real
|
||||
rubric bars. Navigating away destroys the sandbox
|
||||
(`curl localhost:8420/v1/sandboxes` shows the count drop).
|
||||
|
||||
## Layout
|
||||
|
||||
```
|
||||
ai_service/
|
||||
main.py app factory, lifespan (httpx pool), CORS, /health
|
||||
config.py pydantic-settings
|
||||
api/ endpoints (SSE envelope lives here, D-016)
|
||||
llm/ provider layer — dumb pipe, no envelope logic
|
||||
scripts/ bootstrap.sh dev.sh test.sh lint.sh
|
||||
tests/ pytest — mock provider only
|
||||
```
|
||||
|
||||
Boundary rules: `llm/` imports nothing from `agents/` or `api/`; `agents/` imports nothing from `api/`.
|
||||
@@ -1,3 +0,0 @@
|
||||
"""Nextcraft AI tutor service — six LLM agents behind a provider-agnostic layer."""
|
||||
|
||||
__version__ = "0.2.0"
|
||||
@@ -1,18 +0,0 @@
|
||||
"""Agent framework — BaseAgent ABC (D-018), registry, sessions, structured outputs.
|
||||
|
||||
Boundary rule: agents/ imports from llm/, prompts/, corpus/ — never from api/.
|
||||
"""
|
||||
|
||||
from .base import BaseAgent
|
||||
from .registry import AgentRegistry
|
||||
from .session import InMemorySessionStore, SessionStore
|
||||
from .structured import StructuredOutputError, extract_json_object
|
||||
|
||||
__all__ = [
|
||||
"AgentRegistry",
|
||||
"BaseAgent",
|
||||
"InMemorySessionStore",
|
||||
"SessionStore",
|
||||
"StructuredOutputError",
|
||||
"extract_json_object",
|
||||
]
|
||||
@@ -1,63 +0,0 @@
|
||||
"""AssessorAgent — rubric coaching over REAL grading output (REQ-3-007).
|
||||
|
||||
v0.3 re-grounding: the Assessor no longer invents scores from corpus
|
||||
artifacts — the process-trace grading engine (Phase 3) computes and
|
||||
persists the validated RubricScore. This agent now renders the STORED
|
||||
grade as rubric-anchored coaching: explains the criteria, cites strengths
|
||||
and gaps, and frames next steps. Corpus artifacts are retired from this
|
||||
path (corpus dormancy, Task 6-1-04).
|
||||
"""
|
||||
|
||||
from pydantic import BaseModel, Field
|
||||
|
||||
from ..corpus.learner_context import LearnerContext, get_learner_context
|
||||
from ..grading.store import GradeRecord
|
||||
from ..prompts.assessor import SYSTEM_PROMPT, render_context
|
||||
from .base import BaseAgent
|
||||
|
||||
|
||||
class GradeCoaching(BaseModel):
|
||||
"""Rubric-anchored coaching rendered FROM the stored grade (not invented)."""
|
||||
|
||||
summary: str = Field(min_length=1)
|
||||
strengths: list[str] = Field(min_length=1, max_length=3)
|
||||
gaps: list[str] = Field(min_length=1, max_length=3)
|
||||
next_steps: list[str] = Field(min_length=1, max_length=3)
|
||||
|
||||
|
||||
GRADE_COACHING_SCHEMA_HINT = (
|
||||
'{"summary": "<two sentences on the grade>", '
|
||||
'"strengths": ["<one sentence>"], "gaps": ["<one sentence>"], '
|
||||
'"next_steps": ["<one sentence>"]}'
|
||||
)
|
||||
|
||||
|
||||
class AssessorAgent(BaseAgent):
|
||||
name = "assessor"
|
||||
|
||||
def system_prompt(self, learner_context: LearnerContext | None = None) -> str:
|
||||
ctx = learner_context or get_learner_context()
|
||||
return SYSTEM_PROMPT.format_map(render_context(ctx))
|
||||
|
||||
async def coach_grade(
|
||||
self,
|
||||
grade: GradeRecord,
|
||||
learner_context: LearnerContext | None = None,
|
||||
) -> GradeCoaching:
|
||||
"""Render the STORED grade as coaching via the D-020 defense."""
|
||||
grade_json = {
|
||||
"verdict": grade.verdict,
|
||||
"scores": grade.scores,
|
||||
"digest": grade.digest,
|
||||
}
|
||||
coaching: GradeCoaching = await self.structured_reply(
|
||||
history=None,
|
||||
user_input=(
|
||||
"The learner's process-trace grade (computed by the grading "
|
||||
f"engine) is:\n{grade_json!r}\nExplain it as coaching."
|
||||
),
|
||||
learner_context=learner_context,
|
||||
schema=GradeCoaching,
|
||||
schema_hint=GRADE_COACHING_SCHEMA_HINT,
|
||||
)
|
||||
return coaching
|
||||
@@ -1,77 +0,0 @@
|
||||
"""BaseAgent ABC — the contract all six tutor agents implement (D-018).
|
||||
|
||||
Subclasses set `name`, override `system_prompt()`, and rarely `stream_reply()`.
|
||||
The default pipeline: build_messages() → provider.stream_chat()/chat().
|
||||
"""
|
||||
|
||||
from abc import ABC, abstractmethod
|
||||
from collections.abc import AsyncIterator
|
||||
|
||||
from pydantic import BaseModel
|
||||
|
||||
from ..config import Settings
|
||||
from ..corpus.learner_context import LearnerContext
|
||||
from ..llm.base import LLMProvider
|
||||
from ..llm.types import Message
|
||||
from .structured import structured_completion
|
||||
|
||||
|
||||
class BaseAgent(ABC):
|
||||
"""A tutor agent: system prompt + message assembly + provider delegation."""
|
||||
|
||||
name: str = "base"
|
||||
|
||||
def __init__(self, provider: LLMProvider, settings: Settings) -> None:
|
||||
self.provider = provider
|
||||
self.settings = settings
|
||||
|
||||
@abstractmethod
|
||||
def system_prompt(self, learner_context: LearnerContext | None = None) -> str:
|
||||
"""Return the agent's system prompt, learner-context-aware."""
|
||||
|
||||
def build_messages(
|
||||
self,
|
||||
history: list[Message] | None = None,
|
||||
user_input: str = "",
|
||||
learner_context: LearnerContext | None = None,
|
||||
) -> list[Message]:
|
||||
"""Compose the full message list: system prompt + history + user turn."""
|
||||
messages: list[Message] = [
|
||||
Message(role="system", content=self.system_prompt(learner_context))
|
||||
]
|
||||
for m in history or []:
|
||||
messages.append(m)
|
||||
if user_input:
|
||||
messages.append(Message(role="user", content=user_input))
|
||||
return messages
|
||||
|
||||
async def stream_reply(
|
||||
self,
|
||||
history: list[Message] | None = None,
|
||||
user_input: str = "",
|
||||
learner_context: LearnerContext | None = None,
|
||||
response_format: dict | None = None,
|
||||
) -> AsyncIterator[str]:
|
||||
"""Stream incremental content deltas for a conversational reply."""
|
||||
messages = self.build_messages(history, user_input, learner_context)
|
||||
async for token in self.provider.stream_chat(
|
||||
messages, model=self.settings.model, response_format=response_format
|
||||
):
|
||||
yield token
|
||||
|
||||
async def structured_reply(
|
||||
self,
|
||||
history: list[Message] | None = None,
|
||||
user_input: str = "",
|
||||
learner_context: LearnerContext | None = None,
|
||||
schema: type[BaseModel] | None = None,
|
||||
schema_hint: str = "",
|
||||
) -> BaseModel:
|
||||
"""Non-streaming completion parsed into a pydantic model (D-020 defense)."""
|
||||
if schema is None:
|
||||
raise ValueError("structured_reply requires a schema")
|
||||
messages = self.build_messages(history, user_input, learner_context)
|
||||
return await structured_completion(
|
||||
self.provider, messages, model=self.settings.model,
|
||||
schema=schema, schema_hint=schema_hint,
|
||||
)
|
||||
@@ -1,13 +0,0 @@
|
||||
"""CoachAgent — pacing, motivation, retrieval practice (REQ-2-005)."""
|
||||
|
||||
from ..corpus.learner_context import LearnerContext, get_learner_context
|
||||
from ..prompts.coach import SYSTEM_PROMPT, render_context
|
||||
from .base import BaseAgent
|
||||
|
||||
|
||||
class CoachAgent(BaseAgent):
|
||||
name = "coach"
|
||||
|
||||
def system_prompt(self, learner_context: LearnerContext | None = None) -> str:
|
||||
ctx = learner_context or get_learner_context()
|
||||
return SYSTEM_PROMPT.format_map(render_context(ctx))
|
||||
@@ -1,104 +0,0 @@
|
||||
"""ExaminerAgent — the seventh agent: oral-defense examiner (REQ-3-006, A-109).
|
||||
|
||||
BOUNDARY DECISION (PERSONAS conflict rule, honored by construction): the
|
||||
examiner is a TEXT agent. It composes the LLM provider through BaseAgent and
|
||||
consumes defense transcript turns; it NEVER imports voice/ — STT/TTS belong
|
||||
to the API endpoints (they move audio bytes; the agent moves question text).
|
||||
Integrity signals (long pauses, off-scope cadence) are computed by the
|
||||
endpoint layer from turn metadata (latency_ms etc.), not by the agent.
|
||||
|
||||
Digest discipline (D-028 mirror): questions are grounded in the compact
|
||||
TraceDigest + variant statement — never the raw trace, never learner ids.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from pydantic import BaseModel, ConfigDict, Field
|
||||
|
||||
from ..grading.features import TraceDigest
|
||||
from ..llm.types import Message
|
||||
from ..prompts.examiner import SYSTEM_PROMPT, VERDICT_SCHEMA_HINT, render_digest_context
|
||||
from .base import BaseAgent
|
||||
|
||||
|
||||
class DefenseVerdict(BaseModel):
|
||||
"""D-20-validated final defense verdict (structured mode)."""
|
||||
|
||||
model_config = ConfigDict(extra="forbid")
|
||||
|
||||
verdict: str = Field(pattern="^(mastered|developing|not_yet)$")
|
||||
understanding: str = Field(min_length=1)
|
||||
process_justification: str = Field(min_length=1)
|
||||
communication: str = Field(min_length=1)
|
||||
strengths: list[str] = Field(min_length=1, max_length=2)
|
||||
gaps: list[str] = Field(min_length=1, max_length=2)
|
||||
|
||||
|
||||
class ExaminerAgent(BaseAgent):
|
||||
"""Conducts the oral defense: next_question + final_verdict."""
|
||||
|
||||
name = "examiner"
|
||||
|
||||
def system_prompt(self, learner_context=None) -> str: # noqa: ANN001
|
||||
"""Examiner is context-free (digest-anonymous, D-028 mirror)."""
|
||||
return SYSTEM_PROMPT
|
||||
|
||||
def build_defense_messages(
|
||||
self,
|
||||
trace_digest: TraceDigest | None = None,
|
||||
variant_statement: str | None = None,
|
||||
history: list[Message] | None = None,
|
||||
) -> list[Message]:
|
||||
"""System + grounding + defense transcript (no learner id — D-028)."""
|
||||
digest_json = (
|
||||
trace_digest.model_dump_json() if trace_digest is not None else "{}"
|
||||
)
|
||||
messages: list[Message] = [
|
||||
Message(role="system", content=SYSTEM_PROMPT),
|
||||
Message(role="user", content=render_digest_context(digest_json, variant_statement)),
|
||||
Message(
|
||||
role="assistant",
|
||||
content="Understood. I will question the learner about this build session.",
|
||||
),
|
||||
]
|
||||
for m in history or []:
|
||||
messages.append(m)
|
||||
return messages
|
||||
|
||||
async def next_question(
|
||||
self,
|
||||
history: list[Message],
|
||||
trace_digest: TraceDigest | None = None,
|
||||
variant_statement: str | None = None,
|
||||
) -> str:
|
||||
"""One examiner question (streamed over SSE by the endpoints)."""
|
||||
messages = self.build_defense_messages(trace_digest, variant_statement, history)
|
||||
messages.append(
|
||||
Message(role="user", content="Ask the learner your next question now.")
|
||||
)
|
||||
reply = await self.provider.chat(messages, model=self.settings.model)
|
||||
return reply
|
||||
|
||||
async def final_verdict(
|
||||
self,
|
||||
history: list[Message],
|
||||
trace_digest: TraceDigest | None = None,
|
||||
variant_statement: str | None = None,
|
||||
) -> DefenseVerdict:
|
||||
"""Structured verdict via the D-020 4-layer defense."""
|
||||
from .structured import structured_completion # module-direct (G-4)
|
||||
|
||||
messages = self.build_defense_messages(trace_digest, variant_statement, history)
|
||||
messages.append(
|
||||
Message(
|
||||
role="user",
|
||||
content="The defense is finished. Return the final verdict JSON now.",
|
||||
)
|
||||
)
|
||||
return await structured_completion(
|
||||
self.provider,
|
||||
messages,
|
||||
model=self.settings.model,
|
||||
schema=DefenseVerdict,
|
||||
schema_hint=VERDICT_SCHEMA_HINT,
|
||||
)
|
||||
@@ -1,39 +0,0 @@
|
||||
"""LabAgent — in-flow feedback over LIVE sandbox telemetry (REQ-3-007).
|
||||
|
||||
v0.3 re-grounding: consumes a TraceDigest computed from the learner's real
|
||||
trace (grading/features.compute_digest over TraceStore events) — the v0.2
|
||||
corpus scenarios are retired from this path (corpus dormancy, Task 6-1-04).
|
||||
No session chat — each request is one live-trace read.
|
||||
"""
|
||||
|
||||
from collections.abc import AsyncIterator
|
||||
|
||||
from ..config import Settings
|
||||
from ..corpus.learner_context import LearnerContext, get_learner_context
|
||||
from ..grading.features import TraceDigest
|
||||
from ..llm.base import LLMProvider
|
||||
from ..prompts.lab import SYSTEM_PROMPT, render_context, render_digest_timeline
|
||||
from .base import BaseAgent
|
||||
|
||||
|
||||
class LabAgent(BaseAgent):
|
||||
name = "lab"
|
||||
|
||||
def __init__(self, provider: LLMProvider, settings: Settings) -> None:
|
||||
super().__init__(provider, settings)
|
||||
|
||||
def system_prompt(self, learner_context: LearnerContext | None = None) -> str:
|
||||
ctx = learner_context or get_learner_context()
|
||||
return SYSTEM_PROMPT.format_map(render_context(ctx))
|
||||
|
||||
async def stream_feedback(
|
||||
self,
|
||||
digest: TraceDigest | None,
|
||||
learner_context: LearnerContext | None = None,
|
||||
) -> AsyncIterator[str]:
|
||||
"""Feedback grounded in the learner's live trace digest."""
|
||||
timeline = render_digest_timeline(digest)
|
||||
async for token in self.stream_reply(
|
||||
history=None, user_input=timeline, learner_context=learner_context
|
||||
):
|
||||
yield token
|
||||
@@ -1,17 +0,0 @@
|
||||
"""MentorAgent — long-horizon career narrative (REQ-2-010).
|
||||
|
||||
Streaming, session-backed conversational agent: the learner can ask
|
||||
follow-up questions about their trajectory and the Mentor keeps context.
|
||||
"""
|
||||
|
||||
from ..corpus.learner_context import LearnerContext, get_learner_context
|
||||
from ..prompts.mentor import SYSTEM_PROMPT, render_context
|
||||
from .base import BaseAgent
|
||||
|
||||
|
||||
class MentorAgent(BaseAgent):
|
||||
name = "mentor"
|
||||
|
||||
def system_prompt(self, learner_context: LearnerContext | None = None) -> str:
|
||||
ctx = learner_context or get_learner_context()
|
||||
return SYSTEM_PROMPT.format_map(render_context(ctx))
|
||||
@@ -1,79 +0,0 @@
|
||||
"""ProctorAgent — integrity signals + coaching over REAL inputs (REQ-3-007).
|
||||
|
||||
v0.3 re-grounding: consumes the learner's live trace digest (idle gaps,
|
||||
command cadence), the DefenseStore integrity signals (long pauses from the
|
||||
oral defense), and the variant seed cross-check — NOT v0.2 corpus
|
||||
scenarios. The proctor COACHES: it classifies signals supportively and
|
||||
recommends one intervention; it never punishes and never accuses.
|
||||
|
||||
Integrity inputs (computed server-side, passed in by the API layer):
|
||||
- trace digest: idle_gap_count/total, command_categories histogram,
|
||||
error/fix cycles, huge-burst indicators (edit_count vs test runs)
|
||||
- defense signals: long_pauses list from the finished defense (A-109)
|
||||
- variant: seed + params when the task is variant-derived (off-template
|
||||
work is a cross-check input, not an accusation)
|
||||
"""
|
||||
|
||||
from pydantic import BaseModel, Field
|
||||
|
||||
from ..corpus.learner_context import LearnerContext, get_learner_context
|
||||
from ..grading.features import TraceDigest
|
||||
from ..prompts.proctor import SYSTEM_PROMPT, render_context
|
||||
from .base import BaseAgent
|
||||
|
||||
|
||||
class IntegritySignal(BaseModel):
|
||||
signal_type: str # "idle_gap" | "long_pause" | "burst_edit" | "off_template"
|
||||
severity: str # "low" | "medium" | "high"
|
||||
note: str
|
||||
|
||||
|
||||
class ProctorAssessment(BaseModel):
|
||||
signals: list[IntegritySignal] = Field(min_length=0)
|
||||
intervention: str # ONE supportive coaching recommendation
|
||||
summary: str
|
||||
|
||||
|
||||
PROCTOR_ASSESSMENT_SCHEMA_HINT = (
|
||||
'{"signals": [{"signal_type": "<type>", '
|
||||
'"severity": "low"|"medium"|"high", "note": "<one sentence>"}], '
|
||||
'"intervention": "<one supportive recommendation>", '
|
||||
'"summary": "<one sentence>"}'
|
||||
)
|
||||
|
||||
|
||||
class ProctorAgent(BaseAgent):
|
||||
name = "proctor"
|
||||
|
||||
def system_prompt(self, learner_context: LearnerContext | None = None) -> str:
|
||||
ctx = learner_context or get_learner_context()
|
||||
return SYSTEM_PROMPT.format_map(render_context(ctx))
|
||||
|
||||
async def assess(
|
||||
self,
|
||||
digest: TraceDigest | None,
|
||||
defense_signals: dict | None = None,
|
||||
variant_context: dict | None = None,
|
||||
learner_context: LearnerContext | None = None,
|
||||
) -> ProctorAssessment:
|
||||
"""Classify REAL integrity inputs into supportive signals + coaching."""
|
||||
parts: list[str] = []
|
||||
if digest is not None:
|
||||
parts.append(f"Build-session digest:\n{digest.model_dump_json()}")
|
||||
else:
|
||||
parts.append("No build telemetry recorded for this task yet.")
|
||||
if defense_signals:
|
||||
parts.append(f"Oral-defense integrity signals:\n{defense_signals}")
|
||||
if variant_context:
|
||||
parts.append(f"Variant audit context (seed + params):\n{variant_context}")
|
||||
assessment: ProctorAssessment = await self.structured_reply(
|
||||
history=None,
|
||||
user_input=(
|
||||
"Assess this learner's integrity signals supportively.\n\n"
|
||||
+ "\n\n".join(parts)
|
||||
),
|
||||
learner_context=learner_context,
|
||||
schema=ProctorAssessment,
|
||||
schema_hint=PROCTOR_ASSESSMENT_SCHEMA_HINT,
|
||||
)
|
||||
return assessment
|
||||
@@ -1,74 +0,0 @@
|
||||
"""Agent registry — explicit name → agent factory map (D-018, G-4).
|
||||
|
||||
Agents are registered centrally in their own phases (P3-P5) via
|
||||
`registry.register(name, factory)`. One registration pattern, one registry.
|
||||
"""
|
||||
|
||||
from collections.abc import Callable
|
||||
|
||||
from ..config import Settings
|
||||
from ..llm.base import LLMProvider
|
||||
from .base import BaseAgent
|
||||
|
||||
AgentFactory = Callable[[LLMProvider, Settings], BaseAgent]
|
||||
|
||||
|
||||
def register_builtin_agents(registry: "AgentRegistry") -> None:
|
||||
"""Central registration of all seven shipped agents (G-4: one pattern).
|
||||
|
||||
coach, tutor, lab, assessor, proctor, mentor, examiner (Phase 5).
|
||||
New agents register here in their landing phase.
|
||||
"""
|
||||
from .assessor import AssessorAgent
|
||||
from .coach import CoachAgent
|
||||
from .examiner import ExaminerAgent
|
||||
from .lab import LabAgent
|
||||
from .mentor import MentorAgent
|
||||
from .proctor import ProctorAgent
|
||||
from .tutor import TutorAgent
|
||||
|
||||
registry.register("coach", lambda provider, settings: CoachAgent(provider, settings))
|
||||
registry.register("tutor", lambda provider, settings: TutorAgent(provider, settings))
|
||||
registry.register("lab", lambda provider, settings: LabAgent(provider, settings))
|
||||
registry.register(
|
||||
"assessor", lambda provider, settings: AssessorAgent(provider, settings)
|
||||
)
|
||||
registry.register(
|
||||
"proctor", lambda provider, settings: ProctorAgent(provider, settings)
|
||||
)
|
||||
registry.register(
|
||||
"mentor", lambda provider, settings: MentorAgent(provider, settings)
|
||||
)
|
||||
registry.register(
|
||||
"examiner", lambda provider, settings: ExaminerAgent(provider, settings)
|
||||
)
|
||||
|
||||
|
||||
class UnknownAgentError(KeyError):
|
||||
"""Raised when resolving an agent name that was never registered."""
|
||||
|
||||
|
||||
class DuplicateAgentError(ValueError):
|
||||
"""Raised when registering an agent name that already exists."""
|
||||
|
||||
|
||||
class AgentRegistry:
|
||||
def __init__(self) -> None:
|
||||
self._factories: dict[str, AgentFactory] = {}
|
||||
|
||||
def register(self, name: str, factory: AgentFactory) -> None:
|
||||
if name in self._factories:
|
||||
raise DuplicateAgentError(f"agent {name!r} already registered")
|
||||
self._factories[name] = factory
|
||||
|
||||
def names(self) -> list[str]:
|
||||
return sorted(self._factories)
|
||||
|
||||
def get(self, provider: LLMProvider, settings: Settings, name: str) -> BaseAgent:
|
||||
try:
|
||||
factory = self._factories[name]
|
||||
except KeyError:
|
||||
raise UnknownAgentError(
|
||||
f"unknown agent {name!r}; registered: {self.names()}"
|
||||
) from None
|
||||
return factory(provider, settings)
|
||||
@@ -1,98 +0,0 @@
|
||||
"""SessionStore — protocol + in-memory implementation (D-019).
|
||||
|
||||
Protocol is DB-migration-ready (A-003): swap InMemorySessionStore for a
|
||||
Redis/PG-backed implementation without touching the API layer.
|
||||
|
||||
Sessions are agent-scoped: switching agents starts a new session ID (avoids
|
||||
persona bleed, A-007). History windowing happens here (last N messages),
|
||||
controlling token growth per session.
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
from collections import OrderedDict
|
||||
from dataclasses import dataclass, field
|
||||
from typing import Protocol
|
||||
|
||||
from ..llm.types import Message
|
||||
|
||||
DEFAULT_WINDOW = 20
|
||||
DEFAULT_MAX_SESSIONS = 500
|
||||
|
||||
|
||||
@dataclass
|
||||
class AgentSession:
|
||||
session_id: str
|
||||
agent: str
|
||||
learner_id: str = "seed-learner-1"
|
||||
messages: list[Message] = field(default_factory=list)
|
||||
|
||||
|
||||
class SessionStore(Protocol):
|
||||
def get(self, session_id: str) -> AgentSession | None: ...
|
||||
def create(
|
||||
self, session_id: str, agent: str, learner_id: str = "seed-learner-1"
|
||||
) -> AgentSession: ...
|
||||
def append(self, session_id: str, message: Message) -> None: ...
|
||||
def history_window(
|
||||
self, session_id: str, max_messages: int = DEFAULT_WINDOW
|
||||
) -> list[Message]: ...
|
||||
def delete(self, session_id: str) -> None: ...
|
||||
|
||||
|
||||
class InMemorySessionStore:
|
||||
"""asyncio.Lock-guarded dict with 20-message windows and 500-cap LRU eviction."""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
window: int = DEFAULT_WINDOW,
|
||||
max_sessions: int = DEFAULT_MAX_SESSIONS,
|
||||
) -> None:
|
||||
self._sessions: OrderedDict[str, AgentSession] = OrderedDict()
|
||||
self._lock = asyncio.Lock()
|
||||
self._window = window
|
||||
self._max_sessions = max_sessions
|
||||
|
||||
async def get(self, session_id: str) -> AgentSession | None:
|
||||
async with self._lock:
|
||||
session = self._sessions.get(session_id)
|
||||
if session is not None:
|
||||
self._sessions.move_to_end(session_id) # LRU touch
|
||||
return session
|
||||
|
||||
async def create(
|
||||
self, session_id: str, agent: str, learner_id: str = "seed-learner-1"
|
||||
) -> AgentSession:
|
||||
async with self._lock:
|
||||
session = AgentSession(session_id=session_id, agent=agent, learner_id=learner_id)
|
||||
self._sessions[session_id] = session
|
||||
self._evict_locked()
|
||||
return session
|
||||
|
||||
async def append(self, session_id: str, message: Message) -> None:
|
||||
async with self._lock:
|
||||
session = self._sessions.get(session_id)
|
||||
if session is None:
|
||||
raise KeyError(f"unknown session {session_id!r}")
|
||||
session.messages.append(message)
|
||||
# Bound stored history too (window bounds replay, not storage):
|
||||
# keep at most 2x window so retries/recent context survive.
|
||||
if len(session.messages) > self._window * 2:
|
||||
del session.messages[: len(session.messages) - self._window * 2]
|
||||
self._sessions.move_to_end(session_id)
|
||||
|
||||
async def history_window(
|
||||
self, session_id: str, max_messages: int = DEFAULT_WINDOW
|
||||
) -> list[Message]:
|
||||
async with self._lock:
|
||||
session = self._sessions.get(session_id)
|
||||
if session is None:
|
||||
raise KeyError(f"unknown session {session_id!r}")
|
||||
return list(session.messages[-max_messages:])
|
||||
|
||||
async def delete(self, session_id: str) -> None:
|
||||
async with self._lock:
|
||||
self._sessions.pop(session_id, None)
|
||||
|
||||
def _evict_locked(self) -> None:
|
||||
while len(self._sessions) > self._max_sessions:
|
||||
self._sessions.popitem(last=False) # evict least-recently-used
|
||||
@@ -1,117 +0,0 @@
|
||||
"""Structured output defense — 4 layers (D-020).
|
||||
|
||||
Layer 1: response_format={"type":"json_object"} request (auto-degrades on 400
|
||||
inside the provider).
|
||||
Layer 2: prompt-embedded schema hint ("Respond with ONLY valid JSON...").
|
||||
Layer 3: defensive parse — strip markdown fences, extract first balanced
|
||||
JSON object, pydantic model_validate.
|
||||
Layer 4: single bounded retry with the validation error fed back.
|
||||
"""
|
||||
|
||||
from typing import TypeVar
|
||||
|
||||
from pydantic import BaseModel, ValidationError
|
||||
|
||||
from ..llm.base import LLMProvider
|
||||
from ..llm.types import Message
|
||||
|
||||
T = TypeVar("T", bound=BaseModel)
|
||||
|
||||
|
||||
class StructuredOutputError(Exception):
|
||||
"""Raised when the model output cannot be validated after one retry."""
|
||||
|
||||
|
||||
def extract_json_object(text: str) -> str:
|
||||
"""Strip fences and return the first balanced {...} block from text."""
|
||||
stripped = text.strip()
|
||||
if stripped.startswith("```"):
|
||||
first_newline = stripped.find("\n")
|
||||
if first_newline != -1:
|
||||
stripped = stripped[first_newline + 1:]
|
||||
if stripped.rstrip().endswith("```"):
|
||||
stripped = stripped.rstrip()[:-3]
|
||||
stripped = stripped.strip()
|
||||
start = stripped.find("{")
|
||||
if start == -1:
|
||||
raise StructuredOutputError("no JSON object found in model output")
|
||||
depth = 0
|
||||
in_string = False
|
||||
escape = False
|
||||
for i, ch in enumerate(stripped[start:], start=start):
|
||||
if escape:
|
||||
escape = False
|
||||
continue
|
||||
if ch == "\\":
|
||||
escape = True
|
||||
continue
|
||||
if ch == '"' and not escape:
|
||||
in_string = not in_string
|
||||
continue
|
||||
if in_string:
|
||||
continue
|
||||
if ch == "{":
|
||||
depth += 1
|
||||
elif ch == "}":
|
||||
depth -= 1
|
||||
if depth == 0:
|
||||
return stripped[start:i + 1]
|
||||
raise StructuredOutputError("unbalanced JSON object in model output")
|
||||
|
||||
|
||||
def parse_structured(text: str, schema: type[T]) -> T:
|
||||
"""Layer 3: fence-strip + first-balanced-object + pydantic validation."""
|
||||
candidate = extract_json_object(text)
|
||||
try:
|
||||
return schema.model_validate_json(candidate)
|
||||
except ValidationError as exc:
|
||||
raise StructuredOutputError(f"schema validation failed: {exc}") from exc
|
||||
|
||||
|
||||
def schema_instruction(schema_hint: str) -> str:
|
||||
"""Layer 2: prompt-side schema text."""
|
||||
return (
|
||||
"Respond with ONLY a valid JSON object matching this schema — "
|
||||
"no markdown fences, no prose outside the JSON. "
|
||||
f"Schema: {schema_hint}"
|
||||
)
|
||||
|
||||
|
||||
async def structured_completion(
|
||||
provider: LLMProvider,
|
||||
messages: list[Message],
|
||||
*,
|
||||
model: str,
|
||||
schema: type[T],
|
||||
schema_hint: str,
|
||||
retry_feedback: str | None = None,
|
||||
) -> T:
|
||||
"""Full 4-layer pipeline. One bounded retry (layer 4), then raise."""
|
||||
# Build request: append schema instruction to the last user message (layer 2).
|
||||
request = list(messages)
|
||||
last_user = next((m for m in reversed(request) if m.role == "user"), None)
|
||||
if last_user is not None:
|
||||
request = [
|
||||
Message(role=m.role, content=(m.content + "\n\n" + schema_instruction(schema_hint)))
|
||||
if m is last_user else m
|
||||
for m in request
|
||||
]
|
||||
response_format = {"type": "json_object"}
|
||||
raw = await provider.chat(request, model=model, response_format=response_format)
|
||||
try:
|
||||
return parse_structured(raw, schema) # layers 1+2+3
|
||||
except StructuredOutputError as exc:
|
||||
# Layer 4: single bounded retry with error feedback
|
||||
retry_prompt = (
|
||||
f"Your previous response was invalid: {exc}. "
|
||||
f"Return ONLY the corrected JSON matching: {schema_hint}"
|
||||
)
|
||||
request2 = list(messages)
|
||||
request2.append(Message(role="user", content=retry_prompt))
|
||||
raw2 = await provider.chat(request2, model=model, response_format=response_format)
|
||||
try:
|
||||
return parse_structured(raw2, schema)
|
||||
except StructuredOutputError as exc2:
|
||||
raise StructuredOutputError(
|
||||
f"structured output failed after retry: {exc2}"
|
||||
) from exc2
|
||||
@@ -1,13 +0,0 @@
|
||||
"""TutorAgent — concept delivery, Socratic questioning (REQ-2-006)."""
|
||||
|
||||
from ..corpus.learner_context import LearnerContext, get_learner_context
|
||||
from ..prompts.tutor import SYSTEM_PROMPT, render_context
|
||||
from .base import BaseAgent
|
||||
|
||||
|
||||
class TutorAgent(BaseAgent):
|
||||
name = "tutor"
|
||||
|
||||
def system_prompt(self, learner_context: LearnerContext | None = None) -> str:
|
||||
ctx = learner_context or get_learner_context()
|
||||
return SYSTEM_PROMPT.format_map(render_context(ctx))
|
||||
@@ -1,26 +0,0 @@
|
||||
"""API package — composes providers, sessions, and agents via DI.
|
||||
|
||||
Boundary rule: api/ composes agents/ and llm/; they never import api/.
|
||||
"""
|
||||
|
||||
from .assessment import router as assessment_router
|
||||
from .chat import router as chat_router
|
||||
from .defense import router as defense_router
|
||||
from .lab import router as lab_router
|
||||
from .mentor import router as mentor_router
|
||||
from .proctor import router as proctor_router
|
||||
from .sandboxes import router as sandboxes_router
|
||||
from .telemetry import router as telemetry_router
|
||||
from .variants import router as variants_router
|
||||
|
||||
__all__ = [
|
||||
"assessment_router",
|
||||
"chat_router",
|
||||
"lab_router",
|
||||
"mentor_router",
|
||||
"proctor_router",
|
||||
"sandboxes_router",
|
||||
"telemetry_router",
|
||||
"variants_router",
|
||||
"defense_router",
|
||||
]
|
||||
@@ -1,190 +0,0 @@
|
||||
"""/v1/assessment — rubric evaluation + trace grading endpoints (REQ-2-008, REQ-3-004).
|
||||
|
||||
Two endpoint families share this router:
|
||||
|
||||
POST /v1/assessment/evaluate (v0.2, REQ-2-008) — corpus
|
||||
artifact evaluation through
|
||||
the Assessor agent.
|
||||
POST /v1/assessment/grade (v0.3, REQ-3-004) — grade a
|
||||
REAL process trace through
|
||||
the GradingEngine.
|
||||
GET /v1/assessment/grade/{learner_id}/{task_id} — stored latest grade.
|
||||
|
||||
Grading status-code mapping (the engine's outcomes are CONTRACT, not errors):
|
||||
|
||||
GradeRecord(verdict=GRADED) → 200 — rubric scores +
|
||||
verdict (in scores.verdict)
|
||||
+ digest summary.
|
||||
GradeRecord(UNGRADABLE_TRACE_INCOMPLETE) → 200 — the ungradable
|
||||
record IS a valid result:
|
||||
the trace cannot be graded,
|
||||
and the gate surfaces WHY
|
||||
(scores.missing_seqs +
|
||||
scores.integrity_flag).
|
||||
Persisted like any grade.
|
||||
GradeRecord(UNGRADABLE_EMPTY_TRACE) → 200 — no events stored for
|
||||
the pair. This covers BOTH
|
||||
a known pair whose trace
|
||||
ended up empty AND a task
|
||||
that never had a trace at
|
||||
all: the engine cannot
|
||||
distinguish them (zero
|
||||
stored events is zero
|
||||
events), and grading an
|
||||
absent trace genuinely has
|
||||
the empty-trace outcome —
|
||||
a 404 here would erase the
|
||||
durable gate record the
|
||||
engine persists for the
|
||||
pair. PLAN's 404 applies to
|
||||
GET of a never-graded pair.
|
||||
StructuredOutputError → 502 — the trace was
|
||||
gradable but the provider
|
||||
failed the D-020 budget;
|
||||
provider failure (bad
|
||||
gateway to the model), same
|
||||
mapping as evaluate.
|
||||
GET of an unknown (never-graded) pair → 404.
|
||||
|
||||
DI (D-027/D-032 house pattern): engine + store arrive via deps.get_grading_engine
|
||||
/ get_grade_store from app.state; this module owns all FastAPI wiring — the
|
||||
engine knows nothing of HTTP. UNGRADABLE_* bodies are rendered by the same
|
||||
GradeResponse model as GRADED ones (a gate record's `scores` holds the gate
|
||||
detail instead of rubric scores), so consumers read ONE shape.
|
||||
"""
|
||||
|
||||
from datetime import datetime
|
||||
from typing import Any
|
||||
|
||||
from fastapi import APIRouter, Depends, HTTPException
|
||||
from pydantic import BaseModel, Field
|
||||
|
||||
from ..agents.registry import AgentRegistry
|
||||
from ..agents.structured import StructuredOutputError
|
||||
from ..config import Settings
|
||||
from ..corpus.learner_context import get_learner_context
|
||||
from ..grading.engine import GradingEngine
|
||||
from ..grading.store import GradeRecord, GradeStore
|
||||
from .deps import (
|
||||
get_agent_registry,
|
||||
get_grade_store,
|
||||
get_grading_engine,
|
||||
get_provider,
|
||||
get_settings,
|
||||
)
|
||||
|
||||
router = APIRouter(prefix="/v1")
|
||||
|
||||
|
||||
# --- v0.2 artifact evaluation (REQ-2-008) --------------------------------------
|
||||
|
||||
|
||||
class EvaluateRequest(BaseModel):
|
||||
learner_id: str = Field(min_length=1)
|
||||
task_id: str = Field(min_length=1)
|
||||
|
||||
|
||||
@router.post("/assessment/evaluate")
|
||||
async def assessment_evaluate(
|
||||
body: EvaluateRequest,
|
||||
registry: AgentRegistry = Depends(get_agent_registry),
|
||||
settings: Settings = Depends(get_settings),
|
||||
provider=Depends(get_provider),
|
||||
grade_store=Depends(get_grade_store),
|
||||
) -> dict:
|
||||
"""Assessor coaching rendered FROM the learner's stored grade (REQ-3-007).
|
||||
|
||||
The grading engine computes the scores (POST /assessment/grade); this
|
||||
endpoint explains them. No stored grade yet -> 404 (grade first).
|
||||
"""
|
||||
grade = grade_store.get(body.learner_id, body.task_id)
|
||||
if grade is None:
|
||||
raise HTTPException(
|
||||
status_code=404,
|
||||
detail=f"no stored grade for {body.learner_id}/{body.task_id} - grade first",
|
||||
)
|
||||
agent = registry.get(provider, settings, "assessor")
|
||||
learner_context = get_learner_context(body.learner_id)
|
||||
try:
|
||||
coaching = await agent.coach_grade(grade, learner_context)
|
||||
except Exception as exc:
|
||||
raise HTTPException(
|
||||
status_code=502,
|
||||
detail=f"assessment evaluation failed: {exc}",
|
||||
) from exc
|
||||
return {
|
||||
"learner_id": grade.learner_id,
|
||||
"task_id": grade.task_id,
|
||||
"grade_verdict": grade.verdict,
|
||||
"grade_scores": grade.scores,
|
||||
"coaching": coaching.model_dump(),
|
||||
}
|
||||
|
||||
|
||||
# --- # --- v0.3 trace grading (REQ-3-004) ---------------------------------------------
|
||||
|
||||
|
||||
class GradeRequest(BaseModel):
|
||||
learner_id: str = Field(min_length=1)
|
||||
task_id: str = Field(min_length=1)
|
||||
|
||||
|
||||
class GradeResponse(BaseModel):
|
||||
"""GradeRecord over HTTP — one shape for GRADED and UNGRADABLE_* alike.
|
||||
|
||||
`scores` holds the validated rubric (criteria 0-4, strengths, gaps,
|
||||
rubric verdict) for a GRADED record, or the gate detail
|
||||
({integrity_flag, missing_seqs}) for an UNGRADABLE_* record — never both.
|
||||
`digest` is the compact trace summary that fed the rubric prompt (empty
|
||||
for gate records: nothing was graded).
|
||||
"""
|
||||
|
||||
learner_id: str
|
||||
task_id: str
|
||||
variant_seed: str | None
|
||||
digest: dict[str, Any]
|
||||
scores: dict[str, Any]
|
||||
verdict: str
|
||||
model: str
|
||||
created_at: datetime
|
||||
|
||||
|
||||
def _grade_response(record: GradeRecord) -> GradeResponse:
|
||||
return GradeResponse.model_validate(record, from_attributes=True)
|
||||
|
||||
|
||||
@router.post("/assessment/grade", response_model=GradeResponse)
|
||||
async def assessment_grade(
|
||||
body: GradeRequest,
|
||||
engine: GradingEngine = Depends(get_grading_engine),
|
||||
) -> GradeResponse:
|
||||
"""Run the grading engine for one (learner_id, task_id) trace.
|
||||
|
||||
Gate outcomes (UNGRADABLE_*) are 200s — they are first-class results the
|
||||
engine persists, not failures. Only a provider that exhausts the D-020
|
||||
budget turns into a 502; nothing is persisted on that path.
|
||||
"""
|
||||
try:
|
||||
record = await engine.grade(body.learner_id, body.task_id)
|
||||
except StructuredOutputError as exc:
|
||||
raise HTTPException(
|
||||
status_code=502,
|
||||
detail=f"grading failed: {exc}",
|
||||
) from exc
|
||||
return _grade_response(record)
|
||||
|
||||
|
||||
@router.get("/assessment/grade/{learner_id}/{task_id}", response_model=GradeResponse)
|
||||
async def assessment_get_grade(
|
||||
learner_id: str,
|
||||
task_id: str,
|
||||
store: GradeStore = Depends(get_grade_store),
|
||||
) -> GradeResponse:
|
||||
"""Latest stored grade for the pair; 404 when none was ever stored."""
|
||||
record = store.get(learner_id, task_id)
|
||||
if record is None:
|
||||
raise HTTPException(
|
||||
status_code=404,
|
||||
detail=f"no stored grade for {learner_id!r}/{task_id!r}",
|
||||
)
|
||||
return _grade_response(record)
|
||||
@@ -1,128 +0,0 @@
|
||||
"""POST /v1/chat/stream — SSE chat with the D-016 envelope + agent routing.
|
||||
|
||||
Envelope: meta event first (flushed before first token), then raw content
|
||||
deltas, then done; error event before [DONE] on mid-stream failure.
|
||||
Pre-first-byte provider failures surface as in-band `provider_unavailable`
|
||||
error events (SSE 200 headers are already committed once meta flushes).
|
||||
|
||||
Agent routing (A-007): the request names its agent; unknown agents are
|
||||
rejected with 422. No autonomous routing in v0.2.
|
||||
"""
|
||||
|
||||
import json
|
||||
from collections.abc import AsyncIterator
|
||||
|
||||
from fastapi import APIRouter, Depends, HTTPException
|
||||
from pydantic import BaseModel, Field
|
||||
from sse_starlette.sse import EventSourceResponse
|
||||
|
||||
from ..agents.registry import AgentRegistry, UnknownAgentError
|
||||
from ..agents.session import SessionStore
|
||||
from ..config import Settings
|
||||
from ..corpus.learner_context import get_learner_context
|
||||
from ..llm.base import LLMProvider
|
||||
from ..llm.types import Message
|
||||
from .deps import get_agent_registry, get_provider, get_session_store, get_settings
|
||||
|
||||
router = APIRouter(prefix="/v1")
|
||||
|
||||
|
||||
class ChatStreamRequest(BaseModel):
|
||||
agent: str = Field(min_length=1)
|
||||
session_id: str = Field(min_length=1)
|
||||
learner_id: str | None = None
|
||||
messages: list[Message] = Field(min_length=1)
|
||||
|
||||
|
||||
@router.post("/chat/stream")
|
||||
async def chat_stream(
|
||||
body: ChatStreamRequest,
|
||||
registry: AgentRegistry = Depends(get_agent_registry),
|
||||
sessions: SessionStore = Depends(get_session_store),
|
||||
settings: Settings = Depends(get_settings),
|
||||
provider: LLMProvider = Depends(get_provider),
|
||||
) -> EventSourceResponse:
|
||||
# Route to the named agent (A-007); unknown → 422 before any streaming.
|
||||
try:
|
||||
agent = registry.get(provider, settings, body.agent)
|
||||
except UnknownAgentError as exc:
|
||||
raise HTTPException(
|
||||
status_code=422, detail=str(exc)
|
||||
) from None
|
||||
|
||||
learner_context = get_learner_context(body.learner_id)
|
||||
|
||||
# Agent-scoped session (A-007/G-4): persisted turn history, windowed replay.
|
||||
session = await sessions.get(body.session_id)
|
||||
if session is None:
|
||||
session = await sessions.create(
|
||||
body.session_id, agent=body.agent, learner_id=body.learner_id or "learner-001"
|
||||
)
|
||||
# The new user turn is the last message of the request.
|
||||
user_turn = body.messages[-1]
|
||||
history = await sessions.history_window(body.session_id)
|
||||
# Retry dedupe (P1 from final review): a client retry resends the same
|
||||
# turn after a provider failure — don't double-append it to history.
|
||||
last_stored = history[-1] if history else None
|
||||
is_retry = (
|
||||
last_stored is not None
|
||||
and last_stored.role == "user"
|
||||
and last_stored.content == user_turn.content
|
||||
)
|
||||
if not is_retry:
|
||||
await sessions.append(body.session_id, user_turn)
|
||||
else:
|
||||
# On retry the history replay should exclude the stored duplicate.
|
||||
history = history[:-1]
|
||||
|
||||
async def event_stream() -> AsyncIterator[dict]:
|
||||
yield {"event": "message", "data": json.dumps({
|
||||
"type": "meta",
|
||||
"agent": body.agent,
|
||||
"session_id": body.session_id,
|
||||
"model": settings.model,
|
||||
})}
|
||||
first_byte = True
|
||||
reply_parts: list[str] = []
|
||||
try:
|
||||
async for token in agent.stream_reply(
|
||||
history=history,
|
||||
user_input=user_turn.content,
|
||||
learner_context=learner_context,
|
||||
):
|
||||
first_byte = False
|
||||
reply_parts.append(token)
|
||||
yield {"event": "message", "data": json.dumps({
|
||||
"type": "delta", "content": token
|
||||
})}
|
||||
full_reply = "".join(reply_parts)
|
||||
if full_reply:
|
||||
await sessions.append(
|
||||
body.session_id, Message(role="assistant", content=full_reply)
|
||||
)
|
||||
yield {"event": "message", "data": json.dumps({
|
||||
"type": "done", "finish_reason": "stop"
|
||||
})}
|
||||
yield {"event": "message", "data": "[DONE]"}
|
||||
except Exception as exc: # CancelledError is BaseException — passes through
|
||||
message = str(exc)
|
||||
if first_byte:
|
||||
yield {"event": "message", "data": json.dumps({
|
||||
"type": "error", "code": "provider_unavailable", "message": message
|
||||
})}
|
||||
else:
|
||||
yield {"event": "message", "data": json.dumps({
|
||||
"type": "error", "code": "provider_error", "message": message
|
||||
})}
|
||||
# [DONE] is yielded from the except branch, NEVER from finally:
|
||||
# a yield inside finally would re-raise after GeneratorExit when the
|
||||
# client disconnects ("async generator ignored GeneratorExit").
|
||||
yield {"event": "message", "data": "[DONE]"}
|
||||
|
||||
return EventSourceResponse(
|
||||
event_stream(),
|
||||
headers={
|
||||
"Cache-Control": "no-cache",
|
||||
"X-Accel-Buffering": "no",
|
||||
},
|
||||
)
|
||||
@@ -1,378 +0,0 @@
|
||||
"""Oral-defense endpoints (Task 5-3-01, REQ-3-006, A-109).
|
||||
|
||||
Full defense loop over HTTP with mock-first voice (D-030) and the seventh
|
||||
Examiner agent (SSE question streaming happens through the chat pipeline;
|
||||
these endpoints are the session orchestration + transcript persistence):
|
||||
|
||||
POST /v1/defense/start {learner_id, task_id}
|
||||
POST /v1/defense/{id}/answer {text} | multipart audio (STT)
|
||||
GET /v1/defense/{id}/audio/{turn_id} TTS bytes (streaming)
|
||||
POST /v1/defense/{id}/finish verdict + integrity signals
|
||||
GET /v1/defense/{id} transcript + signals
|
||||
|
||||
Integrity signals (A-109) are computed server-side from turn metadata:
|
||||
long pauses = learner turns whose latency_ms exceeds PAUSE_THRESHOLD_MS.
|
||||
The defense does NOT gate on trace completeness (the grader does, G-4);
|
||||
an incomplete trace is surfaced as `trace_complete: false` so the UI can
|
||||
disclose it before the learner defends.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import time
|
||||
from datetime import UTC, datetime
|
||||
|
||||
from fastapi import APIRouter, Depends, File, Form, HTTPException, Request, UploadFile
|
||||
from fastapi.responses import StreamingResponse
|
||||
from pydantic import BaseModel, Field
|
||||
|
||||
from ..agents.examiner import ExaminerAgent
|
||||
from ..grading.features import TraceDigest, compute_digest
|
||||
from ..llm.types import Message
|
||||
from ..voice.base import VoiceDescriptor
|
||||
from ..voice.browser import BROWSER_FALLBACK_DESCRIPTOR
|
||||
from ..voice.defense_store import DefenseRecord, DefenseStore, DefenseTurn
|
||||
from .deps import (
|
||||
get_examiner,
|
||||
get_settings,
|
||||
get_trace_store,
|
||||
get_variant_store,
|
||||
get_voice_provider,
|
||||
get_voice_store,
|
||||
)
|
||||
from .identity import require_verified_age # noqa: E402 (gate dep, D-043)
|
||||
|
||||
router = APIRouter(prefix="/v1/defense", tags=["defense"])
|
||||
|
||||
#: A-109: learner turns slower than this are flagged as long pauses (ms).
|
||||
PAUSE_THRESHOLD_MS = 15_000
|
||||
|
||||
_ROLE_EXAMINER = "examiner"
|
||||
_ROLE_LEARNER = "learner"
|
||||
|
||||
|
||||
class StartRequest(BaseModel):
|
||||
learner_id: str = Field(min_length=1)
|
||||
task_id: str = Field(min_length=1)
|
||||
|
||||
|
||||
class StartResponse(BaseModel):
|
||||
defense_id: str
|
||||
voice_descriptor: dict
|
||||
trace_complete: bool
|
||||
first_question: str
|
||||
|
||||
|
||||
class AnswerResponse(BaseModel):
|
||||
question: str
|
||||
turn_latency: dict[str, int | None]
|
||||
|
||||
|
||||
class FinishResponse(BaseModel):
|
||||
verdict: dict
|
||||
integrity_signals: dict
|
||||
|
||||
|
||||
async def _digest_for_task(
|
||||
trace_store, learner_id: str, task_id: str
|
||||
) -> tuple[TraceDigest | None, bool]:
|
||||
"""Digest of the learner's trace for this task + completeness flag."""
|
||||
if not trace_store.list_tasks(learner_id) or task_id not in trace_store.list_tasks(
|
||||
learner_id
|
||||
):
|
||||
return None, True # no trace at all is "complete" for defense purposes
|
||||
trace = trace_store.get_trace(learner_id, task_id)
|
||||
gaps = trace_store.gaps(learner_id, task_id)
|
||||
return (compute_digest(trace) if trace else None), (len(gaps) == 0)
|
||||
|
||||
|
||||
def _voice_descriptor(settings) -> VoiceDescriptor:
|
||||
"""The capability descriptor for the configured voice mode (D-030).
|
||||
|
||||
Must-Have #6: browser mode returns BROWSER_FALLBACK_DESCRIPTOR so the
|
||||
web client selects native SpeechRecognition/speechSynthesis; mock mode
|
||||
returns the mock descriptor. (A real server provider returns
|
||||
mode="server" — the protocol seam.)
|
||||
"""
|
||||
if (settings.voice_provider or "mock").strip().lower() == "browser":
|
||||
return BROWSER_FALLBACK_DESCRIPTOR
|
||||
return VoiceDescriptor(
|
||||
mode="mock", sr_available=True, tts_available=True, hint=""
|
||||
)
|
||||
|
||||
|
||||
@router.post("/start", response_model=StartResponse)
|
||||
async def start_defense(
|
||||
body: StartRequest,
|
||||
request: Request,
|
||||
examiner: ExaminerAgent = Depends(get_examiner),
|
||||
voice_store: DefenseStore = Depends(get_voice_store),
|
||||
voice_provider=Depends(get_voice_provider),
|
||||
trace_store=Depends(get_trace_store),
|
||||
variant_store=Depends(get_variant_store),
|
||||
settings=Depends(get_settings),
|
||||
) -> StartResponse:
|
||||
# v0.5 identity gate (D-043): allowlist first (G-5), then the school
|
||||
# 16+ verified verdict, before any defense machinery runs.
|
||||
if body.learner_id not in settings.learner_allowlist:
|
||||
raise HTTPException(
|
||||
status_code=403,
|
||||
detail=(
|
||||
f"learner_id {body.learner_id!r} is not on the sandbox "
|
||||
"allowlist (G-5)"
|
||||
),
|
||||
)
|
||||
await require_verified_age(
|
||||
16, body.learner_id, request.app.state.identity_store
|
||||
)
|
||||
record = voice_store.start(
|
||||
DefenseRecord(
|
||||
id=f"dfn-{int(time.time() * 1000):x}-{body.learner_id[:8]}",
|
||||
learner_id=body.learner_id,
|
||||
task_id=body.task_id,
|
||||
status="in_progress",
|
||||
created_at=datetime.now(UTC),
|
||||
)
|
||||
)
|
||||
digest, trace_complete = await _digest_for_task(trace_store, body.learner_id, body.task_id)
|
||||
variant = variant_store.get_by_task(body.task_id)
|
||||
statement = variant.statement if variant is not None else None
|
||||
|
||||
started = time.perf_counter()
|
||||
question = await examiner.next_question(
|
||||
history=[], trace_digest=digest, variant_statement=statement
|
||||
)
|
||||
llm_ms = int((time.perf_counter() - started) * 1000)
|
||||
voice_store.append_turn(
|
||||
record.id,
|
||||
DefenseTurn(
|
||||
defense_id=record.id,
|
||||
seq=0,
|
||||
role=_ROLE_EXAMINER,
|
||||
text=question,
|
||||
ts=datetime.now(UTC),
|
||||
latency_ms=llm_ms,
|
||||
created_at=datetime.now(UTC),
|
||||
),
|
||||
)
|
||||
descriptor = getattr(voice_provider, "descriptor", None) or _voice_descriptor(
|
||||
settings
|
||||
)
|
||||
return StartResponse(
|
||||
defense_id=record.id,
|
||||
voice_descriptor=descriptor.model_dump(),
|
||||
trace_complete=trace_complete,
|
||||
first_question=question,
|
||||
)
|
||||
|
||||
|
||||
@router.post("/{defense_id}/answer", response_model=AnswerResponse)
|
||||
async def answer_defense(
|
||||
defense_id: str,
|
||||
text: str | None = Form(default=None),
|
||||
audio: UploadFile | None = File(default=None),
|
||||
voice_store: DefenseStore = Depends(get_voice_store),
|
||||
voice_provider=Depends(get_voice_provider),
|
||||
examiner: ExaminerAgent = Depends(get_examiner),
|
||||
trace_store=Depends(get_trace_store),
|
||||
variant_store=Depends(get_variant_store),
|
||||
settings=Depends(get_settings),
|
||||
) -> AnswerResponse:
|
||||
record = voice_store.get(defense_id)
|
||||
if record is None:
|
||||
raise HTTPException(status_code=404, detail=f"no defense {defense_id!r}")
|
||||
if record.status == "finished":
|
||||
# The store owns the finished transition but does NOT police turn
|
||||
# sequencing (defense_store.py: "turns after finalize are a sequencing
|
||||
# bug for the endpoints to prevent") — this is the endpoint half of
|
||||
# that contract: a sealed transcript is append-only-no-more.
|
||||
raise HTTPException(
|
||||
status_code=409, detail="defense is finished; start a new defense"
|
||||
)
|
||||
if text is None and audio is None:
|
||||
raise HTTPException(status_code=422, detail="provide {text} or audio")
|
||||
|
||||
# STT (typed fallback bypasses the voice provider entirely).
|
||||
stt_ms: int | None = None
|
||||
if audio is not None:
|
||||
stt_started = time.perf_counter()
|
||||
raw = await audio.read()
|
||||
if not raw:
|
||||
# Empty upload is a client error (422), not a provider crash
|
||||
# (500): validate before the provider call so every provider
|
||||
# sees the same contract.
|
||||
raise HTTPException(status_code=422, detail="audio upload is empty")
|
||||
max_bytes = settings.voice_max_audio_mb * 1024 * 1024
|
||||
if len(raw) > max_bytes:
|
||||
# D-041/G-12: bounded audio BEFORE the provider call — the
|
||||
# client renders this as an honest re-record prompt.
|
||||
raise HTTPException(
|
||||
status_code=413,
|
||||
detail=(
|
||||
f"audio exceeds {settings.voice_max_audio_mb}MB "
|
||||
"— re-record a shorter answer"
|
||||
),
|
||||
)
|
||||
# D-041: strip codec params — MediaRecorder sends
|
||||
# 'audio/webm;codecs=opus'; the bare extension is the provider
|
||||
# contract ('webm'), else real STT endpoints reject the multipart.
|
||||
fmt = (audio.content_type or "audio/wav").split("/")[-1].split(";")[0].strip()
|
||||
try:
|
||||
segment = await voice_provider.transcribe(raw, fmt)
|
||||
except RuntimeError as exc:
|
||||
# Provider failure is the 502 house pattern (assessment.py /
|
||||
# proctor.py), not a 500: a real endpoint outage (or the
|
||||
# default mock's unscripted queue — final-review cross-phase
|
||||
# P0) must surface as an honest upstream error. Both providers
|
||||
# raise RuntimeError with sanitized text (mock: MockVoiceFailure;
|
||||
# openai-audio: key-redacted _sanitize).
|
||||
raise HTTPException(
|
||||
status_code=502, detail=f"voice transcription failed: {exc}"
|
||||
) from exc
|
||||
stt_ms = int((time.perf_counter() - stt_started) * 1000)
|
||||
text = segment.text
|
||||
|
||||
turns = record.turns if hasattr(record, "turns") else []
|
||||
history = [
|
||||
Message(role="assistant" if t.role == _ROLE_EXAMINER else "user", content=t.text)
|
||||
for t in turns
|
||||
]
|
||||
next_seq = len(turns)
|
||||
|
||||
voice_store.append_turn(
|
||||
defense_id,
|
||||
DefenseTurn(
|
||||
defense_id=defense_id,
|
||||
seq=next_seq,
|
||||
role=_ROLE_LEARNER,
|
||||
text=text or "",
|
||||
ts=datetime.now(UTC),
|
||||
latency_ms=stt_ms,
|
||||
created_at=datetime.now(UTC),
|
||||
),
|
||||
)
|
||||
|
||||
digest, _ = await _digest_for_task(trace_store, record.learner_id, record.task_id)
|
||||
variant = variant_store.get_by_task(record.task_id)
|
||||
|
||||
llm_started = time.perf_counter()
|
||||
question = await examiner.next_question(
|
||||
history=history + [Message(role="user", content=text or "")],
|
||||
trace_digest=digest,
|
||||
variant_statement=variant.statement if variant is not None else None,
|
||||
)
|
||||
llm_ms = int((time.perf_counter() - llm_started) * 1000)
|
||||
|
||||
voice_store.append_turn(
|
||||
defense_id,
|
||||
DefenseTurn(
|
||||
defense_id=defense_id,
|
||||
seq=next_seq + 1,
|
||||
role=_ROLE_EXAMINER,
|
||||
text=question,
|
||||
ts=datetime.now(UTC),
|
||||
latency_ms=llm_ms,
|
||||
created_at=datetime.now(UTC),
|
||||
),
|
||||
)
|
||||
return AnswerResponse(
|
||||
question=question,
|
||||
turn_latency={"stt_ms": stt_ms, "llm_ms": llm_ms, "tts_ms": None},
|
||||
)
|
||||
|
||||
|
||||
@router.get("/{defense_id}/audio/{turn_id}")
|
||||
async def defense_audio(
|
||||
defense_id: str,
|
||||
turn_id: int,
|
||||
voice_store: DefenseStore = Depends(get_voice_store),
|
||||
voice_provider=Depends(get_voice_provider),
|
||||
settings=Depends(get_settings),
|
||||
):
|
||||
record = voice_store.get(defense_id)
|
||||
if record is None:
|
||||
raise HTTPException(status_code=404, detail=f"no defense {defense_id!r}")
|
||||
turn = next((t for t in record.turns if t.seq == turn_id), None)
|
||||
if turn is None or turn.role != _ROLE_EXAMINER:
|
||||
raise HTTPException(status_code=404, detail=f"no examiner turn {turn_id!r}")
|
||||
|
||||
async def stream():
|
||||
async for chunk in voice_provider.synthesize(turn.text):
|
||||
yield chunk
|
||||
|
||||
# G-16/D-041: the TTS format is a settings enum; the media_type maps
|
||||
# from it (was hardcoded audio/wav — wrong for every real format).
|
||||
return StreamingResponse(
|
||||
stream(), media_type=f"audio/{settings.voice_tts_format}"
|
||||
)
|
||||
|
||||
|
||||
@router.post("/{defense_id}/finish", response_model=FinishResponse)
|
||||
async def finish_defense(
|
||||
defense_id: str,
|
||||
voice_store: DefenseStore = Depends(get_voice_store),
|
||||
examiner: ExaminerAgent = Depends(get_examiner),
|
||||
trace_store=Depends(get_trace_store),
|
||||
variant_store=Depends(get_variant_store),
|
||||
) -> FinishResponse:
|
||||
record = voice_store.get(defense_id)
|
||||
if record is None:
|
||||
raise HTTPException(status_code=404, detail=f"no defense {defense_id!r}")
|
||||
|
||||
turns = record.turns if hasattr(record, "turns") else []
|
||||
history = [
|
||||
Message(role="assistant" if t.role == _ROLE_EXAMINER else "user", content=t.text)
|
||||
for t in turns
|
||||
]
|
||||
digest, _ = await _digest_for_task(trace_store, record.learner_id, record.task_id)
|
||||
variant = variant_store.get_by_task(record.task_id)
|
||||
verdict = await examiner.final_verdict(
|
||||
history=history,
|
||||
trace_digest=digest,
|
||||
variant_statement=variant.statement if variant is not None else None,
|
||||
)
|
||||
|
||||
signals: dict = {
|
||||
"long_pauses": [
|
||||
{"turn": t.seq, "latency_ms": t.latency_ms}
|
||||
for t in turns
|
||||
if t.role == _ROLE_LEARNER and (t.latency_ms or 0) > PAUSE_THRESHOLD_MS
|
||||
],
|
||||
"pause_threshold_ms": PAUSE_THRESHOLD_MS,
|
||||
# Must-Have #1: "verdict + transcript persisted" — the verdict is
|
||||
# stored INSIDE integrity_signals so GET /{id} after finish can
|
||||
# re-serve it (the finish response alone would lose it). Signals
|
||||
# are a JSON object dict (DefenseStore.finalize contract), so the
|
||||
# verdict nests under the "verdict" key alongside the A-109
|
||||
# markers the Proctor/Mentor feeds read.
|
||||
"verdict": verdict.model_dump(),
|
||||
}
|
||||
voice_store.finalize(defense_id, signals)
|
||||
return FinishResponse(verdict=verdict.model_dump(), integrity_signals=signals)
|
||||
|
||||
|
||||
@router.get("/{defense_id}")
|
||||
async def get_defense(
|
||||
defense_id: str,
|
||||
voice_store: DefenseStore = Depends(get_voice_store),
|
||||
):
|
||||
record = voice_store.get(defense_id)
|
||||
if record is None:
|
||||
raise HTTPException(status_code=404, detail=f"no defense {defense_id!r}")
|
||||
return {
|
||||
"defense_id": record.id,
|
||||
"learner_id": record.learner_id,
|
||||
"task_id": record.task_id,
|
||||
"status": record.status,
|
||||
"turns": [
|
||||
{
|
||||
"seq": t.seq,
|
||||
"role": t.role,
|
||||
"text": t.text,
|
||||
"ts": t.ts,
|
||||
"latency_ms": t.latency_ms,
|
||||
}
|
||||
for t in record.turns
|
||||
],
|
||||
"integrity_signals": record.integrity_signals or {},
|
||||
}
|
||||
@@ -1,79 +0,0 @@
|
||||
"""FastAPI dependencies — provider, settings, sessions, agents via app.state (DI)."""
|
||||
|
||||
from fastapi import Request
|
||||
|
||||
from ..agents.examiner import ExaminerAgent
|
||||
from ..agents.registry import AgentRegistry
|
||||
from ..agents.session import SessionStore
|
||||
from ..config import Settings
|
||||
from ..grading.engine import GradingEngine
|
||||
from ..grading.store import GradeStore
|
||||
from ..llm.base import LLMProvider
|
||||
from ..sandbox.manager import SandboxManager
|
||||
from ..sandbox.workdir import SandboxDir
|
||||
from ..telemetry.ingest import TraceIntegrityMap
|
||||
from ..telemetry.store import TraceStore
|
||||
from ..variants.generator import VariantGenerator
|
||||
from ..variants.store import VariantStore
|
||||
from ..voice.base import VoiceProvider
|
||||
from ..voice.defense_store import DefenseStore
|
||||
|
||||
|
||||
def get_settings(request: Request) -> Settings:
|
||||
return request.app.state.settings
|
||||
|
||||
|
||||
def get_provider(request: Request) -> LLMProvider:
|
||||
return request.app.state.provider
|
||||
|
||||
|
||||
def get_session_store(request: Request) -> SessionStore:
|
||||
return request.app.state.session_store
|
||||
|
||||
|
||||
def get_agent_registry(request: Request) -> AgentRegistry:
|
||||
return request.app.state.agent_registry
|
||||
|
||||
|
||||
def get_sandbox_manager(request: Request) -> SandboxManager:
|
||||
return request.app.state.sandbox_manager
|
||||
|
||||
|
||||
def get_sandbox_test_layout(request: Request) -> SandboxDir | None:
|
||||
"""Optional test seam (app.state.sandbox_test_layout); always None in prod."""
|
||||
return getattr(request.app.state, "sandbox_test_layout", None)
|
||||
|
||||
|
||||
def get_trace_store(request: Request) -> TraceStore:
|
||||
return request.app.state.trace_store
|
||||
|
||||
|
||||
def get_trace_integrity(request: Request) -> TraceIntegrityMap:
|
||||
return request.app.state.trace_integrity
|
||||
|
||||
|
||||
def get_grade_store(request: Request) -> GradeStore:
|
||||
return request.app.state.grade_store
|
||||
|
||||
|
||||
def get_grading_engine(request: Request) -> GradingEngine:
|
||||
return request.app.state.grading_engine
|
||||
|
||||
|
||||
def get_variant_generator(request: Request) -> VariantGenerator:
|
||||
return request.app.state.variant_generator
|
||||
|
||||
|
||||
def get_variant_store(request: Request) -> VariantStore:
|
||||
return request.app.state.variant_store
|
||||
|
||||
def get_voice_store(request: Request) -> DefenseStore:
|
||||
return request.app.state.defense_store
|
||||
|
||||
|
||||
def get_voice_provider(request: Request) -> VoiceProvider:
|
||||
return request.app.state.voice_provider
|
||||
|
||||
|
||||
def get_examiner(request: Request) -> ExaminerAgent:
|
||||
return request.app.state.examiner_agent
|
||||
@@ -1,294 +0,0 @@
|
||||
"""Identity API + age-gate dependencies (REQ-5-003/004, D-042/43).
|
||||
|
||||
Flow (all under mock provider by default; A-304 mock markers ride every
|
||||
response so downstream surfaces never treat mock-verified as real):
|
||||
POST /v1/identity/submit submission → pending (G-13 caps first)
|
||||
GET /v1/identity/status/{lid} latest record + mock marker
|
||||
POST /v1/identity/verify/{sid} poll provider → terminal transition
|
||||
|
||||
Gate dependencies (D-043 binding composition order, mounted by the gated
|
||||
routes — variants/sandbox-create/defense-start for school 16+; one
|
||||
marketplace route for 18+ verified):
|
||||
allowlist (403, G-5 pilot guard) → identity verdict (403 + verify-CTA)
|
||||
→ rate caps (429, owned by the calling routes)
|
||||
|
||||
PII (A-305): raw DOB enters via the submission, is used to derive the
|
||||
band, and is NEVER stored or logged (caplog sentinel test pins it).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import time
|
||||
from datetime import UTC, datetime
|
||||
|
||||
from fastapi import APIRouter, Depends, HTTPException, Request
|
||||
from pydantic import BaseModel, Field, field_validator
|
||||
|
||||
from ..config import Settings
|
||||
from ..identity.base import IdentityProvider, IdentitySubmission
|
||||
from ..identity.store import IdentityRecord, IdentityStore
|
||||
from .deps import get_settings
|
||||
|
||||
router = APIRouter(prefix="/v1/identity", tags=["identity"])
|
||||
|
||||
|
||||
# -- DI ------------------------------------------------------------------------
|
||||
|
||||
|
||||
def get_identity_store(request: Request) -> IdentityStore:
|
||||
return request.app.state.identity_store
|
||||
|
||||
|
||||
def get_identity_provider(request: Request) -> IdentityProvider:
|
||||
return request.app.state.identity_provider
|
||||
|
||||
|
||||
# -- models ----------------------------------------------------------------------
|
||||
|
||||
|
||||
class SubmitBody(BaseModel):
|
||||
learner_id: str = Field(min_length=1)
|
||||
#: ISO date. Validated AT THE BOUNDARY (D1 verifier fix): a malformed
|
||||
#: value would otherwise blow up as a 500 inside derive_age_band on the
|
||||
#: verify path — echoing the raw DOB into the traceback (A-305) and
|
||||
#: leaving a poisoned pending record that G-13 turns into a permanent
|
||||
#: learner lockout.
|
||||
date_of_birth: str = Field(description="ISO date; never stored or logged")
|
||||
document_refs: list[str] = Field(default_factory=list)
|
||||
|
||||
@field_validator("date_of_birth")
|
||||
@classmethod
|
||||
def _validate_dob(cls, value: str) -> str:
|
||||
try:
|
||||
datetime.fromisoformat(value)
|
||||
except ValueError as exc:
|
||||
# 422 with the input scrubbed (A-305/D1 — the app-level
|
||||
# RequestValidationError handler redacts PII field inputs).
|
||||
raise ValueError("date_of_birth must be an ISO date (YYYY-MM-DD)") from exc
|
||||
return value
|
||||
|
||||
|
||||
class SubmitResponse(BaseModel):
|
||||
submission_id: str
|
||||
status: str
|
||||
#: A-304: honesty marker — a mock verdict is NEVER production-verified.
|
||||
mock: bool = True
|
||||
detail: str = ""
|
||||
|
||||
|
||||
class StatusResponse(BaseModel):
|
||||
learner_id: str
|
||||
status: str
|
||||
age_band: str | None
|
||||
mock: bool
|
||||
verified_at: str | None
|
||||
|
||||
|
||||
class VerifyResponse(SubmitResponse):
|
||||
age_band: str | None
|
||||
|
||||
|
||||
# -- G-13 submit caps ---------------------------------------------------------------
|
||||
|
||||
|
||||
class _SubmitRateLimiter:
|
||||
"""Per-learner submit rate cap (in-memory, process-local — the G-5
|
||||
creates-per-min pattern from the sandboxes route)."""
|
||||
|
||||
def __init__(self) -> None:
|
||||
self._window: dict[str, list[float]] = {}
|
||||
|
||||
def check(self, learner_id: str, per_min: int) -> None:
|
||||
now = time.monotonic()
|
||||
window = self._window.setdefault(learner_id, [])
|
||||
window[:] = [t for t in window if now - t < 60.0]
|
||||
if len(window) >= per_min:
|
||||
raise HTTPException(
|
||||
status_code=429,
|
||||
detail="identity submit rate exceeded — wait a minute",
|
||||
)
|
||||
window.append(now)
|
||||
|
||||
|
||||
_rate_limiter = _SubmitRateLimiter()
|
||||
|
||||
|
||||
# -- verification flow ------------------------------------------------------------
|
||||
|
||||
|
||||
@router.post("/submit", response_model=SubmitResponse)
|
||||
async def submit_identity(
|
||||
body: SubmitBody,
|
||||
store: IdentityStore = Depends(get_identity_store),
|
||||
provider: IdentityProvider = Depends(get_identity_provider),
|
||||
settings: Settings = Depends(get_settings),
|
||||
) -> SubmitResponse:
|
||||
# G-13: one active pending submission per learner — resubmit while
|
||||
# pending echoes the pending state (409), not a second submission.
|
||||
if store.count_pending_for_learner(body.learner_id) > 0:
|
||||
latest = store.latest_for_learner(body.learner_id)
|
||||
raise HTTPException(
|
||||
status_code=409,
|
||||
detail={
|
||||
"reason": "submission_pending",
|
||||
"submission_id": latest.id if latest else None,
|
||||
"status": "pending",
|
||||
},
|
||||
)
|
||||
_rate_limiter.check(body.learner_id, settings.identity_submits_per_min)
|
||||
|
||||
submission = IdentitySubmission(
|
||||
learner_id=body.learner_id,
|
||||
date_of_birth=body.date_of_birth,
|
||||
document_refs=body.document_refs,
|
||||
)
|
||||
submission_id = await provider.submit(submission)
|
||||
|
||||
record = store.insert(
|
||||
IdentityRecord(
|
||||
id=submission_id,
|
||||
learner_id=body.learner_id,
|
||||
status="pending",
|
||||
provider="mock" if settings.identity_provider == "mock" else settings.identity_provider,
|
||||
document_refs=body.document_refs, # A-305: opaque handles, never contents
|
||||
submitted_at=datetime.now(UTC),
|
||||
)
|
||||
)
|
||||
return SubmitResponse(
|
||||
submission_id=record.id, status=record.status, mock=record.mock
|
||||
)
|
||||
|
||||
|
||||
@router.get("/status/{learner_id}", response_model=StatusResponse)
|
||||
async def identity_status(
|
||||
learner_id: str,
|
||||
store: IdentityStore = Depends(get_identity_store),
|
||||
) -> StatusResponse:
|
||||
record = store.latest_for_learner(learner_id)
|
||||
if record is None:
|
||||
raise HTTPException(status_code=404, detail="no identity record")
|
||||
return StatusResponse(
|
||||
learner_id=learner_id,
|
||||
status=record.status,
|
||||
age_band=record.age_band,
|
||||
mock=record.mock,
|
||||
verified_at=record.verified_at.isoformat() if record.verified_at else None,
|
||||
)
|
||||
|
||||
|
||||
@router.post("/verify/{submission_id}", response_model=VerifyResponse)
|
||||
async def verify_identity(
|
||||
submission_id: str,
|
||||
store: IdentityStore = Depends(get_identity_store),
|
||||
provider: IdentityProvider = Depends(get_identity_provider),
|
||||
) -> VerifyResponse:
|
||||
record = store.get(submission_id)
|
||||
if record is None:
|
||||
raise HTTPException(status_code=404, detail="no such submission")
|
||||
if record.status != "pending":
|
||||
raise HTTPException(
|
||||
status_code=409,
|
||||
detail=f"submission already {record.status}",
|
||||
)
|
||||
|
||||
verdict = await provider.poll(submission_id)
|
||||
# A-305: derive the band from the provider's verdict; the raw DOB never
|
||||
# entered the store and is not logged here.
|
||||
updated = store.mark_verified(
|
||||
submission_id, verdict.model_dump(), verdict.age_band
|
||||
)
|
||||
assert updated is not None # record existed a moment ago
|
||||
return VerifyResponse(
|
||||
submission_id=updated.id,
|
||||
status=updated.status,
|
||||
age_band=updated.age_band,
|
||||
mock=updated.mock,
|
||||
detail=verdict.detail,
|
||||
)
|
||||
|
||||
|
||||
# -- age-gate dependencies (D-043 composition) -------------------------------------
|
||||
|
||||
|
||||
def _verify_cta_payload(
|
||||
reason: str, min_age: int, record: IdentityRecord | None
|
||||
) -> dict:
|
||||
"""A-306/UX acceptance #2: an actionable 403 — never a bare error."""
|
||||
return {
|
||||
"reason": reason,
|
||||
"min_age": min_age,
|
||||
"current_status": record.status if record else "none",
|
||||
"verify_cta": "/enroll",
|
||||
}
|
||||
|
||||
|
||||
async def require_verified_age(
|
||||
min_age: int,
|
||||
learner_id: str,
|
||||
store: IdentityStore,
|
||||
) -> IdentityRecord:
|
||||
"""The identity half of the D-043 composition (allowlist runs FIRST in
|
||||
the calling routes; this is the second gate; caps come after).
|
||||
|
||||
School 16+ → min_age=16; marketplace 18+ verified → min_age=18.
|
||||
"""
|
||||
record = store.latest_for_learner(learner_id)
|
||||
if record is None or record.status != "verified":
|
||||
raise HTTPException(
|
||||
status_code=403,
|
||||
detail=_verify_cta_payload("identity_verification_required", min_age, record),
|
||||
)
|
||||
# D2 (verifier): FAIL CLOSED. Only canonical bands can pass — None,
|
||||
# unknown, or under-16 bands reject (the gate is the security boundary
|
||||
# for the future vendor and direct store writes; it never trusts a
|
||||
# band it does not recognize).
|
||||
if record.age_band not in ("16-17", "18+"):
|
||||
raise HTTPException(
|
||||
status_code=403,
|
||||
detail=_verify_cta_payload("identity_verification_required", min_age, record),
|
||||
)
|
||||
if min_age > 16 and record.age_band != "18+":
|
||||
raise HTTPException(
|
||||
status_code=403,
|
||||
detail=_verify_cta_payload("age_gate_18_plus", 18, record),
|
||||
)
|
||||
return record
|
||||
|
||||
|
||||
#: Convenience alias for the marketplace 18+ composition (D-043's
|
||||
#: `require_verified_adult` — direct calls use require_verified_age(18, ...)).
|
||||
require_verified_adult = require_verified_age
|
||||
|
||||
|
||||
# -- marketplace 18+ gated stub (G-18, REQ-5-004) ------------------------------------
|
||||
|
||||
marketplace_router = APIRouter(prefix="/v1/marketplace", tags=["marketplace"])
|
||||
|
||||
|
||||
class MarketplaceApplyBody(BaseModel):
|
||||
learner_id: str = Field(min_length=1)
|
||||
job_id: str = Field(min_length=1)
|
||||
|
||||
|
||||
@marketplace_router.post("/apply", status_code=501)
|
||||
async def marketplace_apply_stub(
|
||||
body: MarketplaceApplyBody,
|
||||
request: Request,
|
||||
settings: Settings = Depends(get_settings),
|
||||
) -> dict:
|
||||
"""The ONE gated marketplace route (D-043): proves the 18+ verified
|
||||
composition end-to-end. G-18 honesty: after passing the gate it returns
|
||||
501 with explicit stub + mock markers — the marketplace backend does
|
||||
not exist yet; this route never fabricates an 'applied' outcome.
|
||||
"""
|
||||
if body.learner_id not in settings.learner_allowlist:
|
||||
raise HTTPException(
|
||||
status_code=403,
|
||||
detail=f"learner_id {body.learner_id!r} is not on the sandbox allowlist (G-5)",
|
||||
)
|
||||
await require_verified_age(18, body.learner_id, get_identity_store(request))
|
||||
return {
|
||||
"detail": "marketplace applications are not live yet",
|
||||
"stub": True,
|
||||
"mock": True,
|
||||
}
|
||||
@@ -1,83 +0,0 @@
|
||||
"""POST /v1/lab/feedback — SSE stream of Lab in-flow feedback (REQ-3-007).
|
||||
|
||||
v0.3 re-grounding: LIVE trace digest. Request carries {learner_id, task_id};
|
||||
the digest is computed from the learner's real TraceStore events (D-028)
|
||||
and handed to the Lab agent. No corpus scenarios. Empty/unknown trace is NOT
|
||||
an error — Lab gets a "no telemetry yet" timeline and coaches the baseline.
|
||||
D-016 envelope with agent=lab.
|
||||
"""
|
||||
|
||||
import json
|
||||
from collections.abc import AsyncIterator
|
||||
|
||||
from fastapi import APIRouter, Depends
|
||||
from pydantic import BaseModel, Field
|
||||
from sse_starlette.sse import EventSourceResponse
|
||||
|
||||
from ..agents.registry import AgentRegistry
|
||||
from ..config import Settings
|
||||
from ..corpus.learner_context import get_learner_context
|
||||
from ..grading.features import compute_digest
|
||||
from .deps import (
|
||||
get_agent_registry,
|
||||
get_provider,
|
||||
get_settings,
|
||||
get_trace_store,
|
||||
)
|
||||
|
||||
router = APIRouter(prefix="/v1")
|
||||
|
||||
|
||||
class LabFeedbackRequest(BaseModel):
|
||||
learner_id: str = Field(min_length=1)
|
||||
task_id: str = Field(min_length=1)
|
||||
|
||||
|
||||
@router.post("/lab/feedback")
|
||||
async def lab_feedback(
|
||||
body: LabFeedbackRequest,
|
||||
registry: AgentRegistry = Depends(get_agent_registry),
|
||||
settings: Settings = Depends(get_settings),
|
||||
provider=Depends(get_provider),
|
||||
trace_store=Depends(get_trace_store),
|
||||
) -> EventSourceResponse:
|
||||
agent = registry.get(provider, settings, "lab")
|
||||
learner_context = get_learner_context(body.learner_id)
|
||||
trace = (
|
||||
trace_store.get_trace(body.learner_id, body.task_id)
|
||||
if body.task_id in trace_store.list_tasks(body.learner_id)
|
||||
else []
|
||||
)
|
||||
digest = compute_digest(trace) if trace else None
|
||||
|
||||
async def event_stream() -> AsyncIterator[dict]:
|
||||
yield {"event": "message", "data": json.dumps({
|
||||
"type": "meta",
|
||||
"agent": "lab",
|
||||
"task_id": body.task_id,
|
||||
"model": settings.model,
|
||||
})}
|
||||
first_byte = True
|
||||
try:
|
||||
async for token in agent.stream_feedback(digest, learner_context):
|
||||
first_byte = False
|
||||
yield {"event": "message", "data": json.dumps({
|
||||
"type": "delta", "content": token
|
||||
})}
|
||||
yield {"event": "message", "data": json.dumps({
|
||||
"type": "done", "finish_reason": "stop"
|
||||
})}
|
||||
yield {"event": "message", "data": "[DONE]"}
|
||||
except Exception as exc:
|
||||
code = "provider_unavailable" if first_byte else "provider_error"
|
||||
yield {"event": "message", "data": json.dumps({
|
||||
"type": "error", "code": code, "message": str(exc)
|
||||
})}
|
||||
# [DONE] from except, not finally — a yield in finally would
|
||||
# re-raise after GeneratorExit on client disconnect.
|
||||
yield {"event": "message", "data": "[DONE]"}
|
||||
|
||||
return EventSourceResponse(
|
||||
event_stream(),
|
||||
headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"},
|
||||
)
|
||||
@@ -1,95 +0,0 @@
|
||||
"""POST /v1/mentor/narrative — SSE career narrative stream (REQ-2-010).
|
||||
|
||||
D-016 envelope with agent=mentor. Session-backed: the client supplies a
|
||||
session_id; the Mentor keeps conversation context across follow-ups.
|
||||
"""
|
||||
|
||||
import json
|
||||
from collections.abc import AsyncIterator
|
||||
|
||||
from fastapi import APIRouter, Depends, HTTPException
|
||||
from pydantic import BaseModel, Field
|
||||
from sse_starlette.sse import EventSourceResponse
|
||||
|
||||
from ..agents.registry import AgentRegistry, UnknownAgentError
|
||||
from ..agents.session import SessionStore
|
||||
from ..config import Settings
|
||||
from ..corpus.learner_context import get_learner_context
|
||||
from ..llm.types import Message
|
||||
from .deps import get_agent_registry, get_provider, get_session_store, get_settings
|
||||
|
||||
router = APIRouter(prefix="/v1")
|
||||
|
||||
|
||||
class MentorNarrativeRequest(BaseModel):
|
||||
session_id: str = Field(min_length=1)
|
||||
prompt: str = Field(default="Narrate my trajectory.")
|
||||
learner_id: str | None = None
|
||||
|
||||
|
||||
@router.post("/mentor/narrative")
|
||||
async def mentor_narrative(
|
||||
body: MentorNarrativeRequest,
|
||||
registry: AgentRegistry = Depends(get_agent_registry),
|
||||
sessions: SessionStore = Depends(get_session_store),
|
||||
settings: Settings = Depends(get_settings),
|
||||
provider=Depends(get_provider),
|
||||
) -> EventSourceResponse:
|
||||
try:
|
||||
agent = registry.get(provider, settings, "mentor")
|
||||
except UnknownAgentError as exc:
|
||||
raise HTTPException(status_code=422, detail=str(exc)) from None
|
||||
|
||||
learner_context = get_learner_context(body.learner_id)
|
||||
|
||||
session = await sessions.get(body.session_id)
|
||||
if session is None:
|
||||
session = await sessions.create(
|
||||
body.session_id, agent="mentor", learner_id=body.learner_id or "learner-001"
|
||||
)
|
||||
history = await sessions.history_window(body.session_id)
|
||||
user_message = Message(role="user", content=body.prompt)
|
||||
await sessions.append(body.session_id, user_message)
|
||||
|
||||
async def event_stream() -> AsyncIterator[dict]:
|
||||
yield {"event": "message", "data": json.dumps({
|
||||
"type": "meta",
|
||||
"agent": "mentor",
|
||||
"session_id": body.session_id,
|
||||
"model": settings.model,
|
||||
})}
|
||||
first_byte = True
|
||||
reply_parts: list[str] = []
|
||||
try:
|
||||
async for token in agent.stream_reply(
|
||||
history=history,
|
||||
user_input=body.prompt,
|
||||
learner_context=learner_context,
|
||||
):
|
||||
first_byte = False
|
||||
reply_parts.append(token)
|
||||
yield {"event": "message", "data": json.dumps({
|
||||
"type": "delta", "content": token
|
||||
})}
|
||||
full_reply = "".join(reply_parts)
|
||||
if full_reply:
|
||||
await sessions.append(
|
||||
body.session_id, Message(role="assistant", content=full_reply)
|
||||
)
|
||||
yield {"event": "message", "data": json.dumps({
|
||||
"type": "done", "finish_reason": "stop"
|
||||
})}
|
||||
yield {"event": "message", "data": "[DONE]"}
|
||||
except Exception as exc:
|
||||
code = "provider_unavailable" if first_byte else "provider_error"
|
||||
yield {"event": "message", "data": json.dumps({
|
||||
"type": "error", "code": code, "message": str(exc)
|
||||
})}
|
||||
# [DONE] from except, not finally — a yield in finally would
|
||||
# re-raise after GeneratorExit on client disconnect.
|
||||
yield {"event": "message", "data": "[DONE]"}
|
||||
|
||||
return EventSourceResponse(
|
||||
event_stream(),
|
||||
headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"},
|
||||
)
|
||||
@@ -1,76 +0,0 @@
|
||||
"""POST /v1/proctor/signals — integrity signals over REAL inputs (REQ-3-007).
|
||||
|
||||
v0.3 re-grounding: live trace digest + DefenseStore long-pause signals +
|
||||
variant seed cross-check, no corpus scenarios. The proctor coaches:
|
||||
a pydantic-validated ProctorAssessment (JSON response, not SSE).
|
||||
"""
|
||||
|
||||
from fastapi import APIRouter, Depends, HTTPException
|
||||
from pydantic import BaseModel, Field
|
||||
|
||||
from ..agents.proctor import ProctorAssessment
|
||||
from ..agents.registry import AgentRegistry
|
||||
from ..config import Settings
|
||||
from ..corpus.learner_context import get_learner_context
|
||||
from ..grading.features import compute_digest
|
||||
from .deps import (
|
||||
get_agent_registry,
|
||||
get_provider,
|
||||
get_settings,
|
||||
get_trace_store,
|
||||
get_variant_store,
|
||||
get_voice_store,
|
||||
)
|
||||
|
||||
router = APIRouter(prefix="/v1")
|
||||
|
||||
|
||||
class ProctorRequest(BaseModel):
|
||||
learner_id: str = Field(min_length=1)
|
||||
task_id: str = Field(min_length=1)
|
||||
|
||||
|
||||
@router.post("/proctor/signals", response_model=ProctorAssessment)
|
||||
async def proctor_signals(
|
||||
body: ProctorRequest,
|
||||
registry: AgentRegistry = Depends(get_agent_registry),
|
||||
settings: Settings = Depends(get_settings),
|
||||
provider=Depends(get_provider),
|
||||
trace_store=Depends(get_trace_store),
|
||||
variant_store=Depends(get_variant_store),
|
||||
voice_store=Depends(get_voice_store),
|
||||
) -> ProctorAssessment:
|
||||
"""Real integrity inputs: live digest + defense signals + variant context."""
|
||||
agent = registry.get(provider, settings, "proctor")
|
||||
learner_context = get_learner_context(body.learner_id)
|
||||
|
||||
trace = (
|
||||
trace_store.get_trace(body.learner_id, body.task_id)
|
||||
if body.task_id in trace_store.list_tasks(body.learner_id)
|
||||
else []
|
||||
)
|
||||
digest = compute_digest(trace) if trace else None
|
||||
|
||||
defense_signals = None
|
||||
for record in voice_store.list_for_learner(body.learner_id):
|
||||
if record.task_id == body.task_id and record.status == "finished":
|
||||
defense_signals = record.integrity_signals or None
|
||||
break
|
||||
|
||||
variant = variant_store.get_by_task(body.task_id)
|
||||
variant_context = (
|
||||
{"template_id": variant.template_id, "seed": variant.seed, "params": variant.params}
|
||||
if variant is not None
|
||||
else None
|
||||
)
|
||||
try:
|
||||
return await agent.assess(
|
||||
digest,
|
||||
defense_signals=defense_signals,
|
||||
variant_context=variant_context,
|
||||
learner_context=learner_context,
|
||||
)
|
||||
except Exception as exc:
|
||||
raise HTTPException(
|
||||
status_code=502, detail=f"proctor assessment failed: {exc}"
|
||||
) from exc
|
||||
@@ -1,421 +0,0 @@
|
||||
"""/v1/sandboxes — sandbox lifecycle API with G-5 abuse control (REQ-3-001).
|
||||
|
||||
The sandbox fabric (`sandbox/manager.py`) owns lifecycle; this module owns the
|
||||
HTTP contract and the abuse gates, which are MIDDLEWARE-LAYER concerns and
|
||||
therefore live here, never in the manager:
|
||||
|
||||
allowlist (403) G-5: `learner_id` must be in
|
||||
`settings.learner_allowlist`. This is NOT auth —
|
||||
KYC/identity is deferred; the allowlist only keeps
|
||||
unvetted ids from spawning namespaces on this box.
|
||||
per-learner cap (429) `settings.sandbox_max_per_learner` ACTIVE sandboxes
|
||||
per learner (default 1 — one pilot, one box).
|
||||
global create cap (429) `settings.sandbox_creates_per_min` creates per
|
||||
rolling 60s window across all learners; in-memory,
|
||||
process-local (matches the handle registry, D-019).
|
||||
pool full (503) D-032 capacity guard (`PoolFullError`), no queue.
|
||||
|
||||
Response models are local to the api/ surface. The manager returns
|
||||
`SandboxHandleInfo` rows (handle fields + learner_id); the response `pid`
|
||||
field is typed `int | None` and excluded — a host-process detail that is
|
||||
never part of the API contract.
|
||||
"""
|
||||
|
||||
import time
|
||||
from collections import deque
|
||||
from pathlib import Path
|
||||
|
||||
from fastapi import APIRouter, Depends, HTTPException, Request, Response
|
||||
from pydantic import BaseModel, ConfigDict, Field
|
||||
|
||||
from ..config import Settings
|
||||
from ..sandbox.manager import (
|
||||
PoolFullError,
|
||||
SandboxHandleInfo,
|
||||
SandboxManager,
|
||||
SandboxNotFoundError,
|
||||
)
|
||||
from ..sandbox.workdir import SandboxDir
|
||||
from .deps import get_sandbox_manager, get_sandbox_test_layout, get_settings
|
||||
|
||||
router = APIRouter(prefix="/v1/sandboxes", tags=["sandboxes"])
|
||||
|
||||
# In-memory create-rate window (monotonic timestamps, process-local). Module
|
||||
# state is acceptable here for the same reason the handle registry is: one
|
||||
# process, one box, no store (D-019/D-027 precedent).
|
||||
_CREATE_TIMES: deque[float] = deque()
|
||||
|
||||
|
||||
# -- contracts ----------------------------------------------------------------
|
||||
|
||||
|
||||
class SandboxCreateRequest(BaseModel):
|
||||
learner_id: str = Field(min_length=1)
|
||||
# Optional task key: when set, the sandbox is telemetry-wired (REQ-3-003)
|
||||
# — the in-sandbox capture agent streams workspace events to the ingest.
|
||||
task_id: str | None = None
|
||||
|
||||
|
||||
class SandboxResponse(BaseModel):
|
||||
"""Public sandbox handle. `workdir` is the absolute host path."""
|
||||
|
||||
model_config = ConfigDict(frozen=True)
|
||||
|
||||
id: str
|
||||
learner_id: str
|
||||
workdir: str
|
||||
created_at: str
|
||||
pid: int | None = Field(
|
||||
default=None,
|
||||
exclude=True, # host-process detail; never part of the API contract
|
||||
)
|
||||
|
||||
|
||||
def _to_response(info: SandboxHandleInfo) -> SandboxResponse:
|
||||
return SandboxResponse(
|
||||
id=info.id,
|
||||
learner_id=info.learner_id,
|
||||
workdir=str(info.workdir),
|
||||
created_at=info.created_at.isoformat(),
|
||||
pid=info.pid,
|
||||
)
|
||||
|
||||
|
||||
class SandboxListResponse(BaseModel):
|
||||
sandboxes: list[SandboxResponse]
|
||||
|
||||
|
||||
class SnapshotResponse(BaseModel):
|
||||
sandbox_id: str
|
||||
snapshot_path: str
|
||||
files: list[str]
|
||||
|
||||
|
||||
# -- abuse control (G-5; middleware layer, not auth) ---------------------------
|
||||
|
||||
|
||||
from ..variants.templates import get_template # noqa: E402
|
||||
from .identity import require_verified_age # noqa: E402 (gate dep, D-043)
|
||||
|
||||
|
||||
def _enforce_allowlist(learner_id: str, settings: Settings) -> None:
|
||||
if learner_id not in settings.learner_allowlist:
|
||||
raise HTTPException(
|
||||
status_code=403,
|
||||
detail=f"learner_id {learner_id!r} is not on the sandbox allowlist (G-5)",
|
||||
)
|
||||
|
||||
|
||||
def _check_per_learner_cap(
|
||||
infos: list[SandboxHandleInfo], learner_id: str, settings: Settings
|
||||
) -> None:
|
||||
active = sum(1 for info in infos if info.learner_id == learner_id)
|
||||
if active >= settings.sandbox_max_per_learner:
|
||||
raise HTTPException(
|
||||
status_code=429,
|
||||
detail=(
|
||||
f"learner {learner_id!r} already has {active} active "
|
||||
f"sandbox(es); per-learner cap is {settings.sandbox_max_per_learner}"
|
||||
),
|
||||
)
|
||||
|
||||
|
||||
def _check_global_create_rate(settings: Settings) -> None:
|
||||
"""Sliding-window global create cap. Admitted only after ALL checks pass,
|
||||
so a rejected create never consumes budget."""
|
||||
now = time.monotonic()
|
||||
while _CREATE_TIMES and now - _CREATE_TIMES[0] > 60.0:
|
||||
_CREATE_TIMES.popleft()
|
||||
if len(_CREATE_TIMES) >= settings.sandbox_creates_per_min:
|
||||
raise HTTPException(
|
||||
status_code=429,
|
||||
detail=(
|
||||
f"global sandbox create rate exceeded "
|
||||
f"({settings.sandbox_creates_per_min}/min); retry shortly"
|
||||
),
|
||||
)
|
||||
_CREATE_TIMES.append(now)
|
||||
|
||||
|
||||
# -- endpoints ---------------------------------------------------------------
|
||||
|
||||
|
||||
@router.post("", status_code=201, response_model=SandboxResponse)
|
||||
async def create_sandbox(
|
||||
body: SandboxCreateRequest,
|
||||
request: Request,
|
||||
manager: SandboxManager = Depends(get_sandbox_manager),
|
||||
settings: Settings = Depends(get_settings),
|
||||
) -> SandboxResponse:
|
||||
# v0.5 identity gate (D-043, REQ-5-004): allowlist (G-5) → identity
|
||||
# verdict (403 + verify-CTA) → caps (429) — the binding composition.
|
||||
_enforce_allowlist(body.learner_id, settings)
|
||||
await require_verified_age(
|
||||
16, body.learner_id, request.app.state.identity_store
|
||||
)
|
||||
_check_per_learner_cap(await manager.list(), body.learner_id, settings)
|
||||
_check_global_create_rate(settings)
|
||||
try:
|
||||
info = await manager.create(body.learner_id, task_id=body.task_id)
|
||||
except PoolFullError as exc:
|
||||
raise HTTPException(status_code=503, detail=str(exc)) from exc
|
||||
return _to_response(info)
|
||||
|
||||
|
||||
@router.get("", response_model=SandboxListResponse)
|
||||
async def list_sandboxes(
|
||||
manager: SandboxManager = Depends(get_sandbox_manager),
|
||||
) -> SandboxListResponse:
|
||||
return SandboxListResponse(
|
||||
sandboxes=[_to_response(info) for info in await manager.list()]
|
||||
)
|
||||
|
||||
|
||||
@router.get("/{sandbox_id}", response_model=SandboxResponse)
|
||||
async def get_sandbox(
|
||||
sandbox_id: str,
|
||||
manager: SandboxManager = Depends(get_sandbox_manager),
|
||||
) -> SandboxResponse:
|
||||
try:
|
||||
info = await manager.get(sandbox_id)
|
||||
except SandboxNotFoundError:
|
||||
raise HTTPException(
|
||||
status_code=404, detail=f"unknown sandbox {sandbox_id!r}"
|
||||
) from None
|
||||
return _to_response(info)
|
||||
|
||||
|
||||
@router.post("/{sandbox_id}/snapshot", response_model=SnapshotResponse)
|
||||
async def snapshot_sandbox(
|
||||
sandbox_id: str,
|
||||
response: Response,
|
||||
manager: SandboxManager = Depends(get_sandbox_manager),
|
||||
test_layout: SandboxDir | None = Depends(get_sandbox_test_layout),
|
||||
) -> SnapshotResponse:
|
||||
try:
|
||||
snapshot_path = await manager.snapshot(sandbox_id)
|
||||
except SandboxNotFoundError:
|
||||
raise HTTPException(
|
||||
status_code=404, detail=f"unknown sandbox {sandbox_id!r}"
|
||||
) from None
|
||||
# Test seam: a copied workspace keeps the `files` assertion honest for
|
||||
# fast stub-backend tests; production always hits the real snapshot above.
|
||||
if test_layout is not None and test_layout.snapshots.is_dir():
|
||||
copies = sorted(test_layout.snapshots.iterdir())
|
||||
if copies:
|
||||
response.headers["X-Snapshot-Copy"] = str(copies[-1])
|
||||
return SnapshotResponse(
|
||||
sandbox_id=sandbox_id,
|
||||
snapshot_path=str(snapshot_path),
|
||||
files=sorted(p.name for p in snapshot_path.iterdir()),
|
||||
)
|
||||
|
||||
|
||||
@router.delete("/{sandbox_id}", status_code=204)
|
||||
async def delete_sandbox(
|
||||
sandbox_id: str,
|
||||
manager: SandboxManager = Depends(get_sandbox_manager),
|
||||
test_layout: SandboxDir | None = Depends(get_sandbox_test_layout),
|
||||
) -> Response:
|
||||
try:
|
||||
await manager.get(sandbox_id)
|
||||
except SandboxNotFoundError:
|
||||
raise HTTPException(
|
||||
status_code=404, detail=f"unknown sandbox {sandbox_id!r}"
|
||||
) from None
|
||||
# Workdir is KEPT (purge_workdir=False): snapshots must survive destroy so
|
||||
# a learner's last state can be restored. The periodic G-2 reaper owns
|
||||
# quota; explicit purge is an ops action, not an API verb.
|
||||
await manager.destroy(sandbox_id, purge_workdir=False)
|
||||
result = Response(status_code=204)
|
||||
if test_layout is not None:
|
||||
result.headers["X-Workspace-Copy"] = str(test_layout.workspace)
|
||||
return result
|
||||
|
||||
|
||||
# -- workspace files + exec (Phase 6, REQ-3-008; CUT-2) -------------------------
|
||||
#
|
||||
# The build surface reads/writes/list workspace files and runs Run/Test
|
||||
# commands through the manager's backend. NO interactive shell relay (CUT-2:
|
||||
# keystroke-level stdin/stdout is v0.4) — each exec is a bounded command with
|
||||
# captured output. Paths are WORKSPACE-RELATIVE; traversal outside the
|
||||
# workspace is rejected (the workdir bind is the boundary, but the API adds
|
||||
# its own containment check — defense in depth).
|
||||
|
||||
|
||||
class FileWriteRequest(BaseModel):
|
||||
path: str = Field(min_length=1)
|
||||
content: str
|
||||
|
||||
|
||||
class ExecRequest(BaseModel):
|
||||
cmd: list[str] = Field(min_length=1)
|
||||
|
||||
|
||||
class ExecResponse(BaseModel):
|
||||
cmd: list[str]
|
||||
returncode: int
|
||||
stdout: str
|
||||
stderr: str
|
||||
duration_s: float
|
||||
|
||||
|
||||
async def _workspace_dir(manager: SandboxManager, sandbox_id: str):
|
||||
"""Resolve the sandbox workspace (tracked layout or shell layout)."""
|
||||
info = await manager.get(sandbox_id) # raises SandboxNotFoundError -> 404
|
||||
backend = manager._backend # noqa: SLF001 - API owns the composition seam
|
||||
tracked = getattr(backend, "_tracked", {}).get(sandbox_id)
|
||||
if tracked is not None:
|
||||
return tracked.workspace, info
|
||||
return info.workdir / "workspace", info
|
||||
|
||||
|
||||
def _safe_rel_path(raw: str) -> Path:
|
||||
"""Workspace-relative path; reject absolute/traversal paths."""
|
||||
candidate = Path(raw)
|
||||
if candidate.is_absolute() or ".." in candidate.parts:
|
||||
raise HTTPException(status_code=422, detail=f"invalid workspace path {raw!r}")
|
||||
return candidate
|
||||
|
||||
|
||||
def _resolve_in_workspace(workspace: Path, rel: Path) -> Path:
|
||||
"""Resolve `rel` under `workspace`, refusing symlink escapes (P7).
|
||||
|
||||
The lexical check in `_safe_rel_path` cannot see symlinks: an exec can
|
||||
plant `ln -s /etc target` in the workspace and a follow-up read/write
|
||||
would follow it OUT of the bind. Resolve with the workspace as the
|
||||
anchor (strict: a symlink chain escaping raises) and confirm the
|
||||
normalized target still sits inside the workspace — defense in depth
|
||||
for both read_file and write_file.
|
||||
"""
|
||||
try:
|
||||
target = (workspace / rel).resolve(strict=False)
|
||||
target.relative_to(workspace.resolve(strict=False))
|
||||
except ValueError:
|
||||
raise HTTPException(
|
||||
status_code=422, detail=f"path escapes the workspace: {rel.as_posix()!r}"
|
||||
) from None
|
||||
return target
|
||||
|
||||
|
||||
@router.get("/{sandbox_id}/files")
|
||||
async def list_files(
|
||||
sandbox_id: str,
|
||||
manager: SandboxManager = Depends(get_sandbox_manager),
|
||||
) -> dict:
|
||||
try:
|
||||
workspace, _ = await _workspace_dir(manager, sandbox_id)
|
||||
except SandboxNotFoundError:
|
||||
raise HTTPException(status_code=404, detail=f"no sandbox {sandbox_id!r}") from None
|
||||
return {"files": sorted(p.name for p in workspace.iterdir()) if workspace.is_dir() else []}
|
||||
|
||||
|
||||
@router.get("/{sandbox_id}/files/{path:path}")
|
||||
async def read_file(
|
||||
sandbox_id: str,
|
||||
path: str,
|
||||
manager: SandboxManager = Depends(get_sandbox_manager),
|
||||
) -> dict:
|
||||
try:
|
||||
workspace, _ = await _workspace_dir(manager, sandbox_id)
|
||||
except SandboxNotFoundError:
|
||||
raise HTTPException(status_code=404, detail=f"no sandbox {sandbox_id!r}") from None
|
||||
rel = _safe_rel_path(path)
|
||||
target = _resolve_in_workspace(workspace, rel)
|
||||
if not target.is_file():
|
||||
raise HTTPException(status_code=404, detail=f"no file {path!r}")
|
||||
return {"path": path, "content": target.read_text(errors="replace")}
|
||||
|
||||
|
||||
@router.put("/{sandbox_id}/files/{path:path}")
|
||||
async def write_file(
|
||||
sandbox_id: str,
|
||||
path: str,
|
||||
body: FileWriteRequest,
|
||||
manager: SandboxManager = Depends(get_sandbox_manager),
|
||||
) -> dict:
|
||||
try:
|
||||
workspace, _ = await _workspace_dir(manager, sandbox_id)
|
||||
except SandboxNotFoundError:
|
||||
raise HTTPException(status_code=404, detail=f"no sandbox {sandbox_id!r}") from None
|
||||
rel = _safe_rel_path(body.path)
|
||||
target = _resolve_in_workspace(workspace, rel)
|
||||
target.parent.mkdir(parents=True, exist_ok=True)
|
||||
target.write_text(body.content)
|
||||
return {"path": body.path, "written": True}
|
||||
|
||||
|
||||
#: G-15 (REQ-5-005): per-kind exec command policy — EXACT argv[0] token
|
||||
#: matching, never prefix/substring (trivially bypassed via flags/-c
|
||||
#: passthrough). 'sh -c' passthrough is DISALLOWED for design/simulation
|
||||
#: kinds: the gaming vector would be faking build-style test cycles into a
|
||||
#: kind-agnostic digest. Build kinds keep v0.3 behavior (any command —
|
||||
#: the CUT-2 surface is Run/Test buttons, not a shell relay).
|
||||
#: python (bare) is deliberately absent — in-ns PATH resolves only python3
|
||||
#: (verifier P1); pip is absent (no network in the namespace).
|
||||
_GENERIC_FIRST_TOKENS = frozenset(
|
||||
{"ls", "cat", "pwd", "echo", "python3", "pytest"}
|
||||
)
|
||||
|
||||
|
||||
def _enforce_exec_policy(
|
||||
cmd: list[str], environment: str | None, allowed: set[str] | None = None
|
||||
) -> None:
|
||||
"""422 with the allowed set when a design/sim command is out of policy.
|
||||
|
||||
`allowed` defaults to the generic set; the exec route unions in the
|
||||
template's DECLARED harness argv[0] (a future non-python harness
|
||||
template must not reject its own Run command)."""
|
||||
if environment not in ("design", "simulation"):
|
||||
return # build kind: unchanged v0.3 semantics
|
||||
allowed = set(allowed) if allowed is not None else set(_GENERIC_FIRST_TOKENS)
|
||||
first = cmd[0] if cmd else ""
|
||||
if first in ("sh", "bash"):
|
||||
raise HTTPException(
|
||||
status_code=422,
|
||||
detail=(
|
||||
f"shell passthrough is not allowed in a {environment} "
|
||||
f"environment; allowed commands: {sorted(allowed)}"
|
||||
),
|
||||
)
|
||||
if first not in allowed:
|
||||
raise HTTPException(
|
||||
status_code=422,
|
||||
detail=(
|
||||
f"command {first!r} is not allowed in a {environment} "
|
||||
f"environment; allowed commands: {sorted(allowed)}"
|
||||
),
|
||||
)
|
||||
|
||||
|
||||
@router.post("/{sandbox_id}/exec", response_model=ExecResponse)
|
||||
async def exec_command(
|
||||
sandbox_id: str,
|
||||
body: ExecRequest,
|
||||
request: Request,
|
||||
manager: SandboxManager = Depends(get_sandbox_manager),
|
||||
) -> ExecResponse:
|
||||
try:
|
||||
await manager.get(sandbox_id)
|
||||
except SandboxNotFoundError:
|
||||
raise HTTPException(status_code=404, detail=f"no sandbox {sandbox_id!r}") from None
|
||||
backend = manager._backend # noqa: SLF001 - API owns the composition seam
|
||||
handle = manager._handles.get(sandbox_id) # noqa: SLF001
|
||||
if handle is None:
|
||||
raise HTTPException(status_code=404, detail=f"no live handle {sandbox_id!r}")
|
||||
# REQ-5-005 (G-15): resolve the sandbox's variant environment by its
|
||||
# task_id (manager side-table) and enforce the per-kind command policy
|
||||
# BEFORE execution.
|
||||
task_id = manager._task_ids.get(sandbox_id) # noqa: SLF001 - composition seam
|
||||
if task_id:
|
||||
variant = request.app.state.variant_store.get_by_task(task_id)
|
||||
if variant is not None:
|
||||
template = get_template(variant.template_id)
|
||||
declared = (
|
||||
{template.run_command.split()[0]} if template is not None else set()
|
||||
)
|
||||
_enforce_exec_policy(
|
||||
body.cmd, variant.environment, _GENERIC_FIRST_TOKENS | declared
|
||||
)
|
||||
result = await backend.exec(handle, body.cmd)
|
||||
return ExecResponse(**result.model_dump())
|
||||
@@ -1,175 +0,0 @@
|
||||
"""/v1/telemetry — WS ingest + trace read endpoints (REQ-3-003, D-026, G-3).
|
||||
|
||||
The router composes the telemetry engine via DI: `telemetry/ingest.py` owns
|
||||
the WS protocol (frame contract + flood control + keepalive) and this module
|
||||
only wires `app.state.trace_store` / `app.state.trace_integrity` /
|
||||
`app.state.settings` into it, plus the two HTTP read faces:
|
||||
|
||||
WS /v1/telemetry/ingest?learner_id&task_id[&sandbox_id] (D-026)
|
||||
GET /v1/telemetry/traces/{learner_id}/{task_id} ordered trace; 404 unknown
|
||||
GET /v1/telemetry/gaps/{learner_id}/{task_id} missing seqs ; 404 unknown
|
||||
|
||||
The WS route is a thin DI shell: it validates the query-param identity and
|
||||
the Origin (browser pages are gated to the localhost dev origins — CORS
|
||||
middleware does not cover WS upgrades; the stdlib capture agent sends no
|
||||
Origin and is unaffected), pulls store/integrity/settings from `app.state`,
|
||||
and calls `telemetry_ingest_endpoint(...)` — the engine stays
|
||||
FastAPI-DI-free so it's testable without a router and the api/ layer owns
|
||||
all composition.
|
||||
|
||||
Unknown-trace contract: a trace is KNOWN when it has >=1 stored event OR
|
||||
carries an integrity flag — a flooded trace with zero stored rows still 200s
|
||||
so Proctor/grader can read WHY it's unusable (G-4 consumes
|
||||
`integrity_reason`). `TraceResponse.incomplete` / `.integrity_reason` mirror
|
||||
the map so HTTP consumers never touch process internals.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from fastapi import APIRouter, Depends, HTTPException, WebSocket
|
||||
from pydantic import BaseModel
|
||||
|
||||
from ..telemetry.ingest import (
|
||||
TraceIntegrityMap,
|
||||
telemetry_ingest_endpoint,
|
||||
)
|
||||
from ..telemetry.models import TelemetryEvent
|
||||
from ..telemetry.store import TraceStore
|
||||
from .deps import get_trace_integrity, get_trace_store
|
||||
|
||||
router = APIRouter(prefix="/v1/telemetry", tags=["telemetry"])
|
||||
|
||||
#: Browser Origins allowed to open the ingest socket (A-008 mirror, D-038).
|
||||
#: The stdlib capture agent sends NO Origin header (it is not a browser) and
|
||||
#: stays allowed; a malicious page loaded in the learner's browser would
|
||||
#: carry an Origin and must not be able to poison/flood the trace. CORS
|
||||
#: middleware does NOT cover WebSocket upgrades, so this gate is explicit.
|
||||
#: In network mode the configured CORS list governs (default '*' — any
|
||||
#: origin, since credentials are never used); an explicit list still rejects
|
||||
#: unlisted origins with 1008.
|
||||
_LOCAL_WS_ORIGINS = frozenset(
|
||||
{"http://localhost:3000", "http://127.0.0.1:3000", "http://localhost:8420"}
|
||||
)
|
||||
|
||||
|
||||
def _allowed_ws_origins(settings: object) -> frozenset[str]:
|
||||
configured = getattr(settings, "cors_origin_list", None)
|
||||
if configured is None:
|
||||
return _LOCAL_WS_ORIGINS
|
||||
if configured == ["*"]:
|
||||
return frozenset() # empty = wildcard = every Origin passes
|
||||
return frozenset(configured) | _LOCAL_WS_ORIGINS
|
||||
|
||||
|
||||
# --- WS ingest (D-026) ---------------------------------------------------------
|
||||
|
||||
|
||||
@router.websocket("/ingest")
|
||||
async def telemetry_ingest_ws(websocket: WebSocket) -> None:
|
||||
"""DI shell: resolve app.state services, then hand the socket to the engine.
|
||||
|
||||
The engine's session + flood logic is fully typed and testable without
|
||||
FastAPI; this shim is the only place the two layers meet.
|
||||
"""
|
||||
origin = (websocket.headers.get("origin") or "").strip()
|
||||
allowed = _allowed_ws_origins(getattr(websocket.app.state, "settings", None))
|
||||
if origin and allowed and origin not in allowed:
|
||||
# Same-origin dev pages (Next.js on :3000, the service itself on
|
||||
# :8420) pass; anything else is refused pre-accept. Non-browser
|
||||
# producers (the capture agent, tests) send no Origin and pass.
|
||||
# Wildcard (empty frozenset) passes every Origin in network mode.
|
||||
await websocket.close(
|
||||
code=1008, reason=f"origin {origin!r} not allowed for telemetry ingest"
|
||||
)
|
||||
return
|
||||
query = websocket.query_params
|
||||
learner_id = query.get("learner_id", "")
|
||||
task_id = query.get("task_id", "")
|
||||
if not learner_id or not task_id:
|
||||
# Reject BEFORE accept: closing the pre-accept handshake is the
|
||||
# cheapest denial and unambiguous for the stdlib capture agent.
|
||||
await websocket.close(code=1008, reason="learner_id/task_id query params required")
|
||||
return
|
||||
await telemetry_ingest_endpoint(
|
||||
websocket=websocket,
|
||||
learner_id=learner_id,
|
||||
task_id=task_id,
|
||||
sandbox_id=query.get("sandbox_id", ""),
|
||||
store=websocket.app.state.trace_store,
|
||||
integrity=websocket.app.state.trace_integrity,
|
||||
settings=websocket.app.state.settings,
|
||||
)
|
||||
|
||||
|
||||
# --- HTTP reads -----------------------------------------------------------------
|
||||
|
||||
|
||||
class TraceResponse(BaseModel):
|
||||
"""Ordered trace + integrity signal (G-4 reads incomplete/reason)."""
|
||||
|
||||
learner_id: str
|
||||
task_id: str
|
||||
events: list[TelemetryEvent]
|
||||
incomplete: bool
|
||||
integrity_reason: str | None
|
||||
|
||||
|
||||
class GapsResponse(BaseModel):
|
||||
"""Missing seqs + integrity signal."""
|
||||
|
||||
learner_id: str
|
||||
task_id: str
|
||||
gaps: list[int]
|
||||
incomplete: bool
|
||||
integrity_reason: str | None
|
||||
|
||||
|
||||
def _is_known_trace(
|
||||
store: TraceStore, integrity: TraceIntegrityMap, learner_id: str, task_id: str
|
||||
) -> bool:
|
||||
"""Known = has stored events OR carries an integrity flag (a flooded trace
|
||||
with zero rows must still be readable — Proctor needs the reason)."""
|
||||
return (
|
||||
store.latest_seq(learner_id, task_id) >= 0
|
||||
or integrity.is_incomplete(learner_id, task_id)
|
||||
)
|
||||
|
||||
|
||||
@router.get("/traces/{learner_id}/{task_id}", response_model=TraceResponse)
|
||||
async def get_trace(
|
||||
learner_id: str,
|
||||
task_id: str,
|
||||
store: TraceStore = Depends(get_trace_store),
|
||||
integrity: TraceIntegrityMap = Depends(get_trace_integrity),
|
||||
) -> TraceResponse:
|
||||
if not _is_known_trace(store, integrity, learner_id, task_id):
|
||||
raise HTTPException(
|
||||
status_code=404, detail=f"unknown trace {learner_id!r}/{task_id!r}"
|
||||
)
|
||||
return TraceResponse(
|
||||
learner_id=learner_id,
|
||||
task_id=task_id,
|
||||
events=store.get_trace(learner_id, task_id),
|
||||
incomplete=integrity.is_incomplete(learner_id, task_id),
|
||||
integrity_reason=integrity.reason(learner_id, task_id),
|
||||
)
|
||||
|
||||
|
||||
@router.get("/gaps/{learner_id}/{task_id}", response_model=GapsResponse)
|
||||
async def get_gaps(
|
||||
learner_id: str,
|
||||
task_id: str,
|
||||
store: TraceStore = Depends(get_trace_store),
|
||||
integrity: TraceIntegrityMap = Depends(get_trace_integrity),
|
||||
) -> GapsResponse:
|
||||
if not _is_known_trace(store, integrity, learner_id, task_id):
|
||||
raise HTTPException(
|
||||
status_code=404, detail=f"unknown trace {learner_id!r}/{task_id!r}"
|
||||
)
|
||||
return GapsResponse(
|
||||
learner_id=learner_id,
|
||||
task_id=task_id,
|
||||
gaps=store.gaps(learner_id, task_id),
|
||||
incomplete=integrity.is_incomplete(learner_id, task_id),
|
||||
integrity_reason=integrity.reason(learner_id, task_id),
|
||||
)
|
||||
@@ -1,216 +0,0 @@
|
||||
"""/v1/variants — seeded per-learner task variant endpoints (REQ-3-005, D-029).
|
||||
|
||||
Three faces over the variant engine (the generator + store stay FastAPI-free;
|
||||
this module owns all HTTP wiring — the D-027/D-032 house pattern):
|
||||
|
||||
POST /v1/variants {learner_id, template_id | competency_id}
|
||||
→ 200 the learner's variant — GENERATED on the first request,
|
||||
CACHED (store read, zero LLM calls) on every repeat: D-029
|
||||
reproducibility means one (learner_id, template_id) is ONE
|
||||
variant forever, so a regenerate is always a 200 of the SAME
|
||||
variant, never a second render.
|
||||
→ 404 unknown template_id, or competency_id with no bound template.
|
||||
→ 422 neither template_id nor competency_id given.
|
||||
GET /v1/variants/{task_id}
|
||||
→ 200 the stored variant owning the task key (the grading and
|
||||
telemetry join path); 404 when no variant was ever generated
|
||||
for the task.
|
||||
GET /v1/variants?learner_id=...
|
||||
→ 200 the learner's variants, chronological; [] when none.
|
||||
|
||||
Template resolution: an explicit `template_id` wins; without it the FIRST
|
||||
template bound to `competency_id` is used (`template_for_competency`,
|
||||
D-021 corpus alignment). The response carries `competency_id` resolved
|
||||
from the template library at read time — an enrichment, not a persisted
|
||||
column (the seed re-derives the whole variant, D-029) — so the learner
|
||||
surface can bind a variant to its competency without a library round-trip.
|
||||
|
||||
Distinctness (REQ-3-005): different learners on the same template draw
|
||||
different seeded params and receive distinct statements and task_ids;
|
||||
tests/api/test_variants.py asserts this end-to-end through the API.
|
||||
"""
|
||||
|
||||
from datetime import datetime
|
||||
from typing import Self
|
||||
|
||||
from fastapi import APIRouter, Depends, HTTPException, Request
|
||||
from pydantic import BaseModel, Field, model_validator
|
||||
|
||||
from ..config import Settings
|
||||
from ..identity.store import IdentityStore
|
||||
from ..variants.generator import VariantGenerator
|
||||
from ..variants.store import VariantRecord, VariantStore
|
||||
from ..variants.templates import TaskTemplate, get_template, template_for_competency
|
||||
from .deps import get_settings, get_variant_generator, get_variant_store
|
||||
from .identity import require_verified_age
|
||||
|
||||
|
||||
def _identity_store(request: Request) -> IdentityStore:
|
||||
return request.app.state.identity_store
|
||||
|
||||
|
||||
def _enforce_allowlist(learner_id: str, settings: Settings) -> None:
|
||||
"""G-5 pilot guard — allowlist runs FIRST in the composition (D-043)."""
|
||||
if learner_id not in settings.learner_allowlist:
|
||||
raise HTTPException(
|
||||
status_code=403,
|
||||
detail=f"learner_id {learner_id!r} is not on the sandbox allowlist (G-5)",
|
||||
)
|
||||
|
||||
router = APIRouter(prefix="/v1/variants", tags=["variants"])
|
||||
|
||||
|
||||
# -- contracts ------------------------------------------------------------------
|
||||
|
||||
|
||||
class VariantGenerateRequest(BaseModel):
|
||||
"""One variant identity: an explicit `template_id`, or the first
|
||||
template bound to a `competency_id` (D-021). `template_id` wins when
|
||||
both are given (explicit identity beats derived); at least one is
|
||||
required — 422 otherwise.
|
||||
"""
|
||||
|
||||
learner_id: str = Field(min_length=1)
|
||||
template_id: str | None = None
|
||||
competency_id: str | None = None
|
||||
|
||||
@model_validator(mode="after")
|
||||
def _require_template_or_competency(self) -> Self:
|
||||
if self.template_id is None and self.competency_id is None:
|
||||
raise ValueError("template_id or competency_id is required")
|
||||
return self
|
||||
|
||||
|
||||
class VariantResponse(BaseModel):
|
||||
"""VariantRecord over HTTP, plus the `competency_id` enrichment.
|
||||
|
||||
Every field except `competency_id` mirrors `VariantRecord` exactly
|
||||
(snake_case; `created_at` is an ISO 8601 UTC datetime) — the wire shape
|
||||
typed as `TaskVariant` in packages/types/variants.ts.
|
||||
"""
|
||||
|
||||
learner_id: str
|
||||
task_id: str
|
||||
template_id: str
|
||||
competency_id: str
|
||||
seed: str
|
||||
params: dict[str, str | int]
|
||||
statement: str
|
||||
starter_files: dict[str, str]
|
||||
#: REQ-5-005 (a-11): REQUIRED on the wire — always emitted.
|
||||
environment: str
|
||||
test_command: str
|
||||
created_at: datetime
|
||||
|
||||
|
||||
class VariantListResponse(BaseModel):
|
||||
variants: list[VariantResponse]
|
||||
|
||||
|
||||
# -- resolution + rendering -----------------------------------------------------
|
||||
|
||||
|
||||
def _resolve_template(body: VariantGenerateRequest) -> TaskTemplate:
|
||||
"""Template for the request: the explicit id, else the first template
|
||||
bound to the competency; 404 when neither resolves."""
|
||||
if body.template_id is not None:
|
||||
template = get_template(body.template_id)
|
||||
if template is None:
|
||||
raise HTTPException(
|
||||
status_code=404,
|
||||
detail=f"no task template with id {body.template_id!r}",
|
||||
)
|
||||
return template
|
||||
# The request validator guarantees the disjunction, so reaching here
|
||||
# means a competency_id was given (never None).
|
||||
assert body.competency_id is not None
|
||||
templates = template_for_competency(body.competency_id)
|
||||
if not templates:
|
||||
raise HTTPException(
|
||||
status_code=404,
|
||||
detail=f"no task template for competency {body.competency_id!r}",
|
||||
)
|
||||
return templates[0]
|
||||
|
||||
|
||||
def _competency_for(template_id: str) -> str:
|
||||
"""competency_id enrichment for stored records (read paths)."""
|
||||
template = get_template(template_id)
|
||||
if template is None:
|
||||
# Integrity guard: a stored variant referencing a template that is
|
||||
# no longer in the library cannot be enriched; fail loudly rather
|
||||
# than fabricate a competency binding.
|
||||
raise HTTPException(
|
||||
status_code=500,
|
||||
detail=f"stored variant references unknown template {template_id!r}",
|
||||
)
|
||||
return template.competency_id
|
||||
|
||||
|
||||
def _to_response(record: VariantRecord, competency_id: str) -> VariantResponse:
|
||||
return VariantResponse(
|
||||
learner_id=record.learner_id,
|
||||
task_id=record.task_id,
|
||||
template_id=record.template_id,
|
||||
competency_id=competency_id,
|
||||
seed=record.seed,
|
||||
params=dict(record.params),
|
||||
statement=record.statement,
|
||||
starter_files=dict(record.starter_files),
|
||||
# a-11: REQUIRED on the wire — the server always emits both (v0.5).
|
||||
environment=record.environment or "build",
|
||||
test_command=record.test_command or "pytest -q",
|
||||
created_at=record.created_at,
|
||||
)
|
||||
|
||||
|
||||
# -- endpoints ------------------------------------------------------------------
|
||||
|
||||
|
||||
@router.post("", response_model=VariantResponse)
|
||||
async def generate_variant(
|
||||
body: VariantGenerateRequest,
|
||||
request: Request,
|
||||
generator: VariantGenerator = Depends(get_variant_generator),
|
||||
settings: Settings = Depends(get_settings),
|
||||
) -> VariantResponse:
|
||||
"""The learner's variant for the resolved template — generated on the
|
||||
first request, cached (no LLM call) on every repeat: D-029 makes a
|
||||
regenerate a 200 of the SAME stored variant.
|
||||
|
||||
v0.5 identity gate (D-043, REQ-5-004): the school floor is 16+ verified.
|
||||
Composition order: G-5 allowlist (403) → identity verdict (403 + CTA).
|
||||
"""
|
||||
_enforce_allowlist(body.learner_id, settings)
|
||||
await require_verified_age(16, body.learner_id, _identity_store(request))
|
||||
template = _resolve_template(body)
|
||||
record = await generator.generate(body.learner_id, template.id)
|
||||
return _to_response(record, competency_id=template.competency_id)
|
||||
|
||||
|
||||
@router.get("", response_model=VariantListResponse)
|
||||
async def list_variants(
|
||||
learner_id: str,
|
||||
store: VariantStore = Depends(get_variant_store),
|
||||
) -> VariantListResponse:
|
||||
"""All stored variants for the learner, chronological; [] when none."""
|
||||
variants = [
|
||||
_to_response(record, competency_id=_competency_for(record.template_id))
|
||||
for record in store.list_for_learner(learner_id)
|
||||
]
|
||||
return VariantListResponse(variants=variants)
|
||||
|
||||
|
||||
@router.get("/{task_id}", response_model=VariantResponse)
|
||||
async def get_variant(
|
||||
task_id: str,
|
||||
store: VariantStore = Depends(get_variant_store),
|
||||
) -> VariantResponse:
|
||||
"""The stored variant owning the task key — the grading and telemetry
|
||||
join path; 404 when no variant was ever generated for the task."""
|
||||
record = store.get_by_task(task_id)
|
||||
if record is None:
|
||||
raise HTTPException(
|
||||
status_code=404, detail=f"no stored variant for task {task_id!r}"
|
||||
)
|
||||
return _to_response(record, competency_id=_competency_for(record.template_id))
|
||||
@@ -1,174 +0,0 @@
|
||||
"""Service settings — pydantic-settings, env prefix AI_, .env support."""
|
||||
|
||||
import logging
|
||||
from pathlib import Path
|
||||
from typing import Annotated
|
||||
|
||||
from pydantic import field_validator
|
||||
from pydantic_settings import BaseSettings, NoDecode, SettingsConfigDict
|
||||
|
||||
# apps/ai-service/ (sandbox dir default is relative to the app, not the CWD)
|
||||
_SERVICE_ROOT = Path(__file__).resolve().parent.parent
|
||||
|
||||
# v0.3.6: durable runtime state lives OUTSIDE the repo (founder directive —
|
||||
# the repo is a git tree, not a data dir; old defaults under apps/ai-service/
|
||||
# polluted the working copy). ~/.nextcraft is the state root for DB + sandbox
|
||||
# workdirs; env overrides may use ~/ paths (expanded by the validator below).
|
||||
_STATE_ROOT = Path.home() / ".nextcraft"
|
||||
|
||||
# G-16: TTS response-format whitelist (feeds the TTS route's Content-Type).
|
||||
_TTS_FORMATS = ("mp3", "wav", "opus")
|
||||
|
||||
|
||||
class Settings(BaseSettings):
|
||||
model_config = SettingsConfigDict(env_prefix="AI_", env_file=".env", extra="ignore")
|
||||
|
||||
port: int = 8420
|
||||
provider: str = "mock"
|
||||
model: str = "gemma4:31b"
|
||||
|
||||
ollama_cloud_base_url: str = "https://ollama.com/v1"
|
||||
ollama_cloud_api_key: str = "" # SecretStr adds friction here; never logged, never echoed
|
||||
local_base_url: str = "http://localhost:11434/v1"
|
||||
|
||||
# "auto" sends response_format and degrades on 400; "off" never sends it
|
||||
json_mode: str = "auto"
|
||||
|
||||
# v0.3 sandbox fabric (REQ-3-001): root holding per-sandbox workdirs.
|
||||
# Relative paths resolve against the app dir (apps/ai-service/), not the CWD.
|
||||
sandbox_dir: Path = _STATE_ROOT / "sandboxes"
|
||||
|
||||
# D-032: single-box capacity, no queue — pool full → API maps to 503.
|
||||
sandbox_max_concurrent: int = 5
|
||||
|
||||
# Wall-clock ceiling per sandbox; the manager's async reaper destroys
|
||||
# sandboxes idle past this age (same timer runs the G-2 workdir sweep).
|
||||
sandbox_timeout_s: float = 900.0
|
||||
|
||||
# G-2: soft disk cap per sandbox workdir, enforced best-effort by the
|
||||
# manager sweep (NOT kernel-enforced — no cgroup delegation/sudo here).
|
||||
sandbox_max_workdir_mb: int = 512
|
||||
|
||||
# G-5 abuse control (NOT auth — KYC/auth is deferred; these keep the
|
||||
# single-box pilot from melting down before identity lands):
|
||||
#
|
||||
# Server-side learner allowlist. Env form is a COMMA-SEPARATED string
|
||||
# (e.g. AI_LEARNER_ALLOWLIST="pilot-learner,learner-2"); NoDecode skips
|
||||
# pydantic-settings' JSON decoding of complex types and the validator
|
||||
# below splits/strips/drops empties. Default: the single mock pilot id.
|
||||
learner_allowlist: Annotated[list[str], NoDecode] = ["pilot-learner"]
|
||||
|
||||
# Max ACTIVE sandboxes per learner → API maps excess to 429.
|
||||
sandbox_max_per_learner: int = 1
|
||||
|
||||
# Global create-rate ceiling (creates per rolling 60s window, shared
|
||||
# across learners) → API maps excess to 429. In-memory, process-local.
|
||||
sandbox_creates_per_min: int = 10
|
||||
|
||||
@field_validator("learner_allowlist", mode="before")
|
||||
@classmethod
|
||||
def _split_allowlist_csv(cls, value: object) -> object:
|
||||
if isinstance(value, str):
|
||||
return [item.strip() for item in value.split(",") if item.strip()]
|
||||
return value
|
||||
|
||||
# D-027: SQLite path for telemetry/grades/variants/defenses stores.
|
||||
# v0.3.6: default moved out of the repo to the home state root.
|
||||
db_path: Path = _STATE_ROOT / "data" / "nextcraft.db"
|
||||
|
||||
# v0.3.6 single-port deploy: directory of the exported Next.js app
|
||||
# (apps/web/out). When set, the app serves it at / (StaticFiles) so the
|
||||
# whole site — UI + API — answers on one port behind HAProxy; the web
|
||||
# build bakes NEXT_PUBLIC_AI_SERVICE_URL=self (relative same-origin
|
||||
# fetches). Empty default → no mount; dev/tests unchanged.
|
||||
web_static_dir: Path = Path("")
|
||||
|
||||
@field_validator("db_path", "sandbox_dir", "web_static_dir", mode="before")
|
||||
@classmethod
|
||||
def _expanduser_paths(cls, value: object) -> object:
|
||||
# Env-set paths may use ~ (e.g. AI_DB_PATH=~/.nextcraft/x.db);
|
||||
# pydantic Path does not expand it natively.
|
||||
if isinstance(value, str):
|
||||
return Path(value).expanduser()
|
||||
return value
|
||||
|
||||
@field_validator("voice_tts_format")
|
||||
@classmethod
|
||||
def _validate_tts_format(cls, value: str) -> str:
|
||||
# G-16 + G-11 consistency: unknown values NEVER crash the boot —
|
||||
# fall back to the default with a loud warning (the boot-survival
|
||||
# log lives in main.py's voice fallback; this validator normalizes).
|
||||
v = value.strip().lower()
|
||||
if v not in _TTS_FORMATS:
|
||||
logging.getLogger(__name__).warning(
|
||||
"AI_VOICE_TTS_FORMAT=%r is not one of %s — falling back to 'mp3'",
|
||||
value,
|
||||
_TTS_FORMATS,
|
||||
)
|
||||
return "mp3"
|
||||
return v
|
||||
|
||||
# G-3 flood control (NOT backpressure-by-silence): max events ingested per
|
||||
# (learner_id, task_id) trace before the WS endpoint closes the connection
|
||||
# with 1008 and marks the trace INCOMPLETE_FLOODED. Drop-oldest is
|
||||
# FORBIDDEN — it corrupts grading input (GRILL G-3).
|
||||
telemetry_max_events_per_task: int = 50000
|
||||
|
||||
# Sandbox telemetry wiring (REQ-3-003): loopback host the in-sandbox capture
|
||||
# agent dials to reach this service's WS ingest (the agent joins the sandbox
|
||||
# mount ns but NOT the net ns — exec namespaces are offline, so the agent
|
||||
# shares the host network and reaches the app over loopback). Port reuses
|
||||
# `port` (A-004); only the host is configurable — never a second port.
|
||||
telemetry_ingest_host: str = "127.0.0.1"
|
||||
|
||||
# v0.3.5 network mode (D-038): dev.sh binds 0.0.0.0 so remote browsers can
|
||||
# reach the stack; '*' (default) lets any origin call the API (safe ONLY
|
||||
# because credentials are never enabled — A-008). Set a comma-separated
|
||||
# origin list (e.g. 'http://nextcraft-1:3000') to restrict instead.
|
||||
# NOTE (v0.3.6): with the UI served same-origin from 8420 this is moot in
|
||||
# production (same-origin requests skip CORS); it stays for the two-port
|
||||
# dev topology.
|
||||
cors_origins: str = "*"
|
||||
|
||||
@property
|
||||
def cors_origin_list(self) -> list[str]:
|
||||
value = self.cors_origins.strip()
|
||||
if value == "*":
|
||||
return ["*"]
|
||||
return [o.strip() for o in value.split(",") if o.strip()]
|
||||
|
||||
# Identity provider selection (REQ-5-003, A-303): 'mock' (default —
|
||||
# deterministic, no vendor spend pre-pilot; verdicts carry mock=True
|
||||
# forever per A-304). A real KYC vendor drops in via the
|
||||
# IdentityProvider protocol without API changes.
|
||||
identity_provider: str = "mock"
|
||||
# G-13: identity submit caps — one active pending per learner (409 on
|
||||
# resubmit) and a per-learner submit rate ceiling.
|
||||
identity_submits_per_min: int = 3
|
||||
|
||||
# Voice provider selection (REQ-5-001, D-040): 'mock' (default — the
|
||||
# no-key path is first-class; tests never call a real voice API),
|
||||
# 'browser' (client-native SR/TTS; the descriptor tells the web client),
|
||||
# or 'openai-audio' (real server STT/TTS against an OpenAI-compatible
|
||||
# audio endpoint). openai-audio requires voice_base_url + voice_api_key;
|
||||
# when unconfigured the lifespan falls back to mock with a loud log
|
||||
# (G-11 — a typo'd env must never crash the unattended boot).
|
||||
voice_provider: str = "mock"
|
||||
|
||||
# Real server voice (D-040, A-301): endpoint-agnostic by config (D-014
|
||||
# pattern) — any OpenAI-compatible audio API works. Keys env-only,
|
||||
# never committed, never logged (mirrors ollama_cloud_api_key).
|
||||
voice_base_url: str = ""
|
||||
voice_api_key: str = ""
|
||||
voice_stt_model: str = "whisper-1"
|
||||
voice_tts_model: str = "tts-1"
|
||||
voice_tts_voice: str = "alloy"
|
||||
# G-16: whitelist, not free string — this feeds the TTS route's
|
||||
# Content-Type. A str + mode-after validator (NOT a pydantic Literal):
|
||||
# a Literal would raise ValidationError at Settings construction, before
|
||||
# main.py's G-11 fallback could catch it — crashing the unattended boot
|
||||
# on a typo'd env. Invalid values fall back to the default LOUDLY.
|
||||
voice_tts_format: str = "mp3"
|
||||
# A-302/D-041: upload guard before the provider call (webm/opus is
|
||||
# ~0.5-1MB/min, so 10MB tolerates very long answers).
|
||||
voice_max_audio_mb: int = 10
|
||||
@@ -1,15 +0,0 @@
|
||||
"""Mock engine inputs — pydantic-typed corpus (D-021).
|
||||
|
||||
Convention-aligned with the TS `packages/mock-data` layer: identical ID
|
||||
strings (stack-*, comp-*, learner-*, art-*, mc-*), cross-referenced by the
|
||||
counterpart files. No codegen in v0.2 — alignment is by documented
|
||||
convention; revisit codegen only if drift bites (v0.3).
|
||||
"""
|
||||
|
||||
from .learner_context import LEARNER_CONTEXTS, LearnerContext, get_learner_context
|
||||
|
||||
__all__ = [
|
||||
"LEARNER_CONTEXTS",
|
||||
"LearnerContext",
|
||||
"get_learner_context",
|
||||
]
|
||||
@@ -1,214 +0,0 @@
|
||||
"""Pre-baked artifacts + rubrics + defense transcripts — Assessor mock inputs (REQ-2-008).
|
||||
|
||||
v0.2 mock engine inputs (pre-baked artifacts/rubrics/transcripts) — DORMANT as of v0.3 re-
|
||||
grounding (Task 6-1-04): no production code path imports this module. Retained as Phase-3
|
||||
calibration history.
|
||||
|
||||
|
||||
Counterpart: packages/mock-data/ai-scenarios.ts (artifact IDs string-identical,
|
||||
D-021). Real process-trace grading is a v0.3+ engine (assessment engine);
|
||||
these pre-baked submissions stand in for artifact + defense evaluation.
|
||||
"""
|
||||
|
||||
from pydantic import BaseModel
|
||||
|
||||
|
||||
class RubricCriterion(BaseModel):
|
||||
criterion_id: str
|
||||
name: str
|
||||
weight: float
|
||||
description: str
|
||||
|
||||
|
||||
class AssessmentRubric(BaseModel):
|
||||
rubric_id: str
|
||||
competency_id: str
|
||||
criteria: list[RubricCriterion]
|
||||
|
||||
|
||||
class ArtifactSubmission(BaseModel):
|
||||
artifact_id: str
|
||||
name: str
|
||||
artifact_type: str # "code" | "design" | "simulation"
|
||||
competency_id: str
|
||||
description: str
|
||||
evidence_excerpt: str # what the grader sees of the artifact itself
|
||||
|
||||
|
||||
class DefenseTranscript(BaseModel):
|
||||
transcript_id: str
|
||||
artifact_id: str
|
||||
turns: list[dict] # {"speaker": "examiner"|"learner", "text": "..."}
|
||||
|
||||
|
||||
_RUBRIC_ORCHESTRATION = AssessmentRubric(
|
||||
rubric_id="rubric-orchestration-c002",
|
||||
competency_id="stack-orchestration-c002",
|
||||
criteria=[
|
||||
RubricCriterion(
|
||||
criterion_id="rc-architecture",
|
||||
name="Agent architecture soundness",
|
||||
weight=0.3,
|
||||
description="State boundaries and responsibilities are clearly separated",
|
||||
),
|
||||
RubricCriterion(
|
||||
criterion_id="rc-communication",
|
||||
name="Inter-agent communication design",
|
||||
weight=0.3,
|
||||
description="Message contracts are explicit, typed, and failure-aware",
|
||||
),
|
||||
RubricCriterion(
|
||||
criterion_id="rc-reliability",
|
||||
name="Reliability engineering",
|
||||
weight=0.25,
|
||||
description="Retries, timeouts, and degradation paths handled",
|
||||
),
|
||||
RubricCriterion(
|
||||
criterion_id="rc-process",
|
||||
name="Process trace quality",
|
||||
weight=0.15,
|
||||
description="Telemetry shows iterative building with real checkpoints",
|
||||
),
|
||||
],
|
||||
)
|
||||
|
||||
_RUBRIC_TOOL_USE = AssessmentRubric(
|
||||
rubric_id="rubric-orchestration-c003",
|
||||
competency_id="stack-orchestration-c003",
|
||||
criteria=[
|
||||
RubricCriterion(
|
||||
criterion_id="rc-eval-design",
|
||||
name="Evaluation design rigor",
|
||||
weight=0.35,
|
||||
description="Hypotheses, controls, and metrics are explicit and defensible",
|
||||
),
|
||||
RubricCriterion(
|
||||
criterion_id="rc-eval-robustness",
|
||||
name="Harness robustness",
|
||||
weight=0.35,
|
||||
description="Error handling, variance awareness, and reproducibility",
|
||||
),
|
||||
RubricCriterion(
|
||||
criterion_id="rc-eval-insight",
|
||||
name="Insight extraction",
|
||||
weight=0.3,
|
||||
description="Results are interpreted into concrete engineering decisions",
|
||||
),
|
||||
],
|
||||
)
|
||||
|
||||
_ARTIFACT_RESEARCH_ASSISTANT = ArtifactSubmission(
|
||||
artifact_id="art-eval-research-assistant",
|
||||
name="Multi-agent research assistant (eval build)",
|
||||
artifact_type="code",
|
||||
competency_id="stack-orchestration-c002",
|
||||
description=(
|
||||
"LangGraph-based assistant planning, retrieving, drafting cited reviews."
|
||||
),
|
||||
evidence_excerpt=(
|
||||
"planner.py defines state schema with explicit fields "
|
||||
"(plan, findings, draft); tool_node.py wraps retrieval with a "
|
||||
"3-retry loop and typed ToolMessage responses; tests cover "
|
||||
"planner->tool->writer handoffs; README shows graph diagram"
|
||||
),
|
||||
)
|
||||
|
||||
_TRANSCRIPT_RESEARCH_ASSISTANT = DefenseTranscript(
|
||||
transcript_id="defense-art-eval-research-assistant",
|
||||
artifact_id="art-eval-research-assistant",
|
||||
turns=[
|
||||
{"speaker": "examiner",
|
||||
"text": "Why did you give the planner sole write access to the plan field?"},
|
||||
{"speaker": "learner",
|
||||
"text": "So worker nodes can't mutate each other's inputs — "
|
||||
"the state stays predictable and the graph is debuggable"},
|
||||
{"speaker": "examiner",
|
||||
"text": "What happens when the retrieval tool times out three times?"},
|
||||
{"speaker": "learner",
|
||||
"text": "The tool node degrades to a no-op ToolMessage with a "
|
||||
"retry flag so the writer can fall back to existing findings"},
|
||||
{"speaker": "examiner", "text": "How would you extend this to a third agent?"},
|
||||
{"speaker": "learner",
|
||||
"text": "Add a reviewer node with its own typed messages, same pattern"},
|
||||
],
|
||||
)
|
||||
|
||||
_ARTIFACT_RAG_DASHBOARD = ArtifactSubmission(
|
||||
artifact_id="art-eval-rag-dashboard",
|
||||
name="RAG retrieval quality dashboard (eval build)",
|
||||
artifact_type="code",
|
||||
competency_id="stack-orchestration-c003",
|
||||
description="Dashboard comparing chunking strategies/rerankers across 800 queries.",
|
||||
evidence_excerpt=(
|
||||
"eval harness sweeps 4 chunk sizes x 3 rerankers; results table auto-generated; "
|
||||
"no error handling on the query loader; tests only cover the happy path"
|
||||
),
|
||||
)
|
||||
|
||||
_TRANSCRIPT_RAG_DASHBOARD = DefenseTranscript(
|
||||
transcript_id="defense-art-eval-rag-dashboard",
|
||||
artifact_id="art-eval-rag-dashboard",
|
||||
turns=[
|
||||
{"speaker": "examiner", "text": "How did you control for query difficulty across runs?"},
|
||||
{"speaker": "learner", "text": "I, um, used the same query set each time"},
|
||||
{"speaker": "examiner", "text": "What happens if the query loader hits a malformed row?"},
|
||||
{"speaker": "learner", "text": "I didn't handle that. It would probably crash."},
|
||||
{"speaker": "examiner", "text": "What would you improve first?"},
|
||||
{"speaker": "learner",
|
||||
"text": "Probably add the error handling, then look at variance between runs"},
|
||||
],
|
||||
)
|
||||
|
||||
RUBRICS: dict[str, AssessmentRubric] = {
|
||||
_RUBRIC_ORCHESTRATION.rubric_id: _RUBRIC_ORCHESTRATION,
|
||||
_RUBRIC_TOOL_USE.rubric_id: _RUBRIC_TOOL_USE,
|
||||
}
|
||||
|
||||
ARTIFACTS: dict[str, ArtifactSubmission] = {
|
||||
a.artifact_id: a
|
||||
for a in (_ARTIFACT_RESEARCH_ASSISTANT, _ARTIFACT_RAG_DASHBOARD)
|
||||
}
|
||||
|
||||
TRANSCRIPTS: dict[str, DefenseTranscript] = {
|
||||
t.transcript_id: t
|
||||
for t in (_TRANSCRIPT_RESEARCH_ASSISTANT, _TRANSCRIPT_RAG_DASHBOARD)
|
||||
}
|
||||
|
||||
|
||||
def rubric_for_competency(competency_id: str) -> AssessmentRubric | None:
|
||||
for rubric in RUBRICS.values():
|
||||
if rubric.competency_id == competency_id:
|
||||
return rubric
|
||||
return None
|
||||
|
||||
|
||||
def get_artifact_bundle(artifact_id: str) -> tuple[ArtifactSubmission, AssessmentRubric] | None:
|
||||
"""Resolve (artifact, rubric) for an artifact ID; None if unknown."""
|
||||
artifact = ARTIFACTS.get(artifact_id)
|
||||
if artifact is None:
|
||||
return None
|
||||
rubric = rubric_for_competency(artifact.competency_id)
|
||||
if rubric is None:
|
||||
return None
|
||||
return artifact, rubric
|
||||
|
||||
|
||||
def get_transcript_for_artifact(artifact_id: str) -> DefenseTranscript | None:
|
||||
for transcript in TRANSCRIPTS.values():
|
||||
if transcript.artifact_id == artifact_id:
|
||||
return transcript
|
||||
return None
|
||||
|
||||
|
||||
def render_rubric(rubric: AssessmentRubric) -> str:
|
||||
lines = [f"Rubric: {rubric.rubric_id} (competency {rubric.competency_id})"]
|
||||
for c in rubric.criteria:
|
||||
lines.append(f"- {c.criterion_id} ({c.weight:.2f}): {c.name} — {c.description}")
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def render_transcript(transcript: DefenseTranscript) -> str:
|
||||
lines = [f"Defense transcript: {transcript.transcript_id}"]
|
||||
for turn in transcript.turns:
|
||||
lines.append(f"{turn['speaker']}: {turn['text']}")
|
||||
return "\n".join(lines)
|
||||
@@ -1,108 +0,0 @@
|
||||
"""Learner context corpus — pydantic mirror of TS learner-progress.ts (D-021).
|
||||
|
||||
Counterpart: packages/mock-data/src/learner-progress.ts (or learner-progress.ts
|
||||
at package root). IDs are string-identical: learner-001, stack-orchestration,
|
||||
stack-safety, stack-orchestration-c00N, art-*, mc-*.
|
||||
"""
|
||||
|
||||
from pydantic import BaseModel
|
||||
|
||||
|
||||
class CompetencyProgress(BaseModel):
|
||||
competency_id: str
|
||||
title: str
|
||||
status: str # "mastered" | "in_progress" | "not_started"
|
||||
|
||||
|
||||
class StackProgress(BaseModel):
|
||||
stack_id: str
|
||||
title: str
|
||||
percent: int
|
||||
|
||||
|
||||
class LearnerContext(BaseModel):
|
||||
learner_id: str
|
||||
name: str
|
||||
active_stacks: list[StackProgress]
|
||||
active_competencies: list[CompetencyProgress]
|
||||
microcredential_count: int
|
||||
recent_artifacts: list[str] # artifact names
|
||||
|
||||
|
||||
_STACK_ORCHESTRATION = StackProgress(
|
||||
stack_id="stack-orchestration", title="AI Orchestration Engineer", percent=62
|
||||
)
|
||||
_STACK_SAFETY = StackProgress(
|
||||
stack_id="stack-safety", title="AI Safety & Governance Lead", percent=41
|
||||
)
|
||||
|
||||
_LEARNER_1 = LearnerContext(
|
||||
learner_id="learner-001",
|
||||
name="Alex Rivera",
|
||||
active_stacks=[_STACK_ORCHESTRATION, _STACK_SAFETY],
|
||||
active_competencies=[
|
||||
CompetencyProgress(
|
||||
competency_id="stack-orchestration-c001",
|
||||
title="Agent architecture fundamentals",
|
||||
status="mastered",
|
||||
),
|
||||
CompetencyProgress(
|
||||
competency_id="stack-orchestration-c002",
|
||||
title="Multi-agent communication patterns",
|
||||
status="in_progress",
|
||||
),
|
||||
CompetencyProgress(
|
||||
competency_id="stack-orchestration-c003",
|
||||
title="Tool use and function calling",
|
||||
status="in_progress",
|
||||
),
|
||||
CompetencyProgress(
|
||||
competency_id="stack-safety-c021",
|
||||
title="Red-team basics for agent systems",
|
||||
status="in_progress",
|
||||
),
|
||||
],
|
||||
microcredential_count=4,
|
||||
recent_artifacts=[
|
||||
"Multi-agent research assistant",
|
||||
"RAG retrieval quality dashboard",
|
||||
],
|
||||
)
|
||||
|
||||
_LEARNER_2 = LearnerContext(
|
||||
learner_id="learner-002",
|
||||
name="Priya Chen",
|
||||
active_stacks=[
|
||||
StackProgress(
|
||||
stack_id="stack-designer", title="Human-AI Product Designer", percent=55
|
||||
),
|
||||
],
|
||||
active_competencies=[
|
||||
CompetencyProgress(
|
||||
competency_id="stack-designer-c001",
|
||||
title="Prompt-to-prototype workflows",
|
||||
status="mastered",
|
||||
),
|
||||
CompetencyProgress(
|
||||
competency_id="stack-designer-c002",
|
||||
title="Evaluating AI UX patterns",
|
||||
status="in_progress",
|
||||
),
|
||||
],
|
||||
microcredential_count=2,
|
||||
recent_artifacts=["AI onboarding flow concept test"],
|
||||
)
|
||||
|
||||
LEARNER_CONTEXTS: dict[str, LearnerContext] = {
|
||||
_LEARNER_1.learner_id: _LEARNER_1,
|
||||
_LEARNER_2.learner_id: _LEARNER_2,
|
||||
}
|
||||
|
||||
DEFAULT_LEARNER_ID = "learner-001"
|
||||
|
||||
|
||||
def get_learner_context(learner_id: str | None = None) -> LearnerContext:
|
||||
"""Resolve a learner context by ID, falling back to the default seed."""
|
||||
if learner_id is None:
|
||||
return LEARNER_CONTEXTS[DEFAULT_LEARNER_ID]
|
||||
return LEARNER_CONTEXTS.get(learner_id, LEARNER_CONTEXTS[DEFAULT_LEARNER_ID])
|
||||
@@ -1,183 +0,0 @@
|
||||
"""Simulated sandbox telemetry corpus — Lab agent mock engine inputs (REQ-2-007).
|
||||
|
||||
v0.2 mock engine inputs (Lab/Proctor scenarios) — DORMANT as of v0.3 re-grounding (Task 6-1-04):
|
||||
no production code path imports this module. Retained as Phase-3 calibration history
|
||||
(corpus/trace_fixtures.py references it from TESTS only).
|
||||
|
||||
|
||||
Counterpart: packages/mock-data/ai-scenarios.ts (scenario IDs string-identical,
|
||||
D-021). Real sandbox telemetry is a v0.3+ engine (sandbox fabric); these
|
||||
scripted event streams stand in for the build-session process trace.
|
||||
"""
|
||||
|
||||
from pydantic import BaseModel
|
||||
|
||||
|
||||
class TelemetryEvent(BaseModel):
|
||||
timestamp: int # seconds since session start
|
||||
kind: str # "keystroke_burst" | "file_save" | "run_tests" | "test_pass"
|
||||
# | "test_fail" | "console_error" | "idle" | "paste" | "commit"
|
||||
|
||||
|
||||
detail: str = ""
|
||||
|
||||
|
||||
class LabTelemetryScenario(BaseModel):
|
||||
scenario_id: str
|
||||
title: str
|
||||
competency_id: str
|
||||
events: list[TelemetryEvent]
|
||||
|
||||
|
||||
class ProctorEvent(BaseModel):
|
||||
timestamp: int # seconds since session start
|
||||
kind: str # "tab_switch" | "idle" | "paste_large" | "focus_lost" | "keystroke_burst"
|
||||
detail: str = ""
|
||||
|
||||
|
||||
class ProctorScenario(BaseModel):
|
||||
scenario_id: str
|
||||
title: str
|
||||
competency_id: str
|
||||
events: list[ProctorEvent]
|
||||
|
||||
|
||||
_PROCTOR_SCENARIO_HEALTHY = ProctorScenario(
|
||||
scenario_id="proctor-scenario-healthy",
|
||||
title="Healthy defense session — focused throughout",
|
||||
competency_id="stack-orchestration-c002",
|
||||
events=[
|
||||
ProctorEvent(timestamp=0, kind="keystroke_burst", detail="session begins"),
|
||||
ProctorEvent(timestamp=310, kind="keystroke_burst", detail="long answer in progress"),
|
||||
ProctorEvent(timestamp=640, kind="keystroke_burst", detail="revision pass"),
|
||||
ProctorEvent(timestamp=900, kind="keystroke_burst", detail="final answer"),
|
||||
],
|
||||
)
|
||||
|
||||
_PROCTOR_SCENARIO_DISTRACTED = ProctorScenario(
|
||||
scenario_id="proctor-scenario-distracted",
|
||||
title="Distracted defense session — tab switches and idle gaps",
|
||||
competency_id="stack-orchestration-c002",
|
||||
events=[
|
||||
ProctorEvent(timestamp=0, kind="keystroke_burst", detail="session begins"),
|
||||
ProctorEvent(timestamp=120, kind="tab_switch", detail="to docs.nextjs.org"),
|
||||
ProctorEvent(timestamp=125, kind="focus_lost", detail="window blur 40s"),
|
||||
ProctorEvent(timestamp=300, kind="idle", detail="no activity for 5 minutes"),
|
||||
ProctorEvent(timestamp=600, kind="tab_switch", detail="to github.com"),
|
||||
ProctorEvent(timestamp=605, kind="focus_lost", detail="window blur 2m"),
|
||||
ProctorEvent(timestamp=720, kind="keystroke_burst", detail="resumes typing"),
|
||||
],
|
||||
)
|
||||
|
||||
_PROCTOR_SCENARIO_FLAGGED = ProctorScenario(
|
||||
scenario_id="proctor-scenario-flagged",
|
||||
title="Flagged defense session — large paste during exam",
|
||||
competency_id="stack-orchestration-c003",
|
||||
events=[
|
||||
ProctorEvent(timestamp=0, kind="keystroke_burst", detail="short intro typed"),
|
||||
ProctorEvent(timestamp=85, kind="paste_large", detail="3,100 chars pasted in 2s"),
|
||||
ProctorEvent(timestamp=90, kind="idle", detail="no activity for 4 minutes"),
|
||||
ProctorEvent(timestamp=330, kind="paste_large", detail="2,800 chars pasted in 2s"),
|
||||
],
|
||||
)
|
||||
|
||||
PROCTOR_SCENARIOS: dict[str, ProctorScenario] = {
|
||||
s.scenario_id: s
|
||||
for s in (
|
||||
_PROCTOR_SCENARIO_HEALTHY,
|
||||
_PROCTOR_SCENARIO_DISTRACTED,
|
||||
_PROCTOR_SCENARIO_FLAGGED,
|
||||
)
|
||||
}
|
||||
|
||||
|
||||
def get_proctor_scenario(scenario_id: str) -> ProctorScenario | None:
|
||||
return PROCTOR_SCENARIOS.get(scenario_id)
|
||||
|
||||
|
||||
def summarize_proctor_scenario(scenario: ProctorScenario) -> str:
|
||||
"""Render the proctor event timeline as compact text for prompt injection."""
|
||||
lines = [f"Defense session: {scenario.title} (competency {scenario.competency_id})"]
|
||||
for event in scenario.events:
|
||||
lines.append(f"t+{event.timestamp}s {event.kind}: {event.detail}".rstrip(": "))
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
_LAB_SCENARIO_STRONG = LabTelemetryScenario(
|
||||
scenario_id="lab-scenario-strong",
|
||||
title="Strong build session — multi-agent research assistant",
|
||||
competency_id="stack-orchestration-c002",
|
||||
events=[
|
||||
TelemetryEvent(timestamp=0, kind="keystroke_burst", detail="planner.py"),
|
||||
TelemetryEvent(timestamp=95, kind="file_save", detail="planner.py"),
|
||||
TelemetryEvent(timestamp=120, kind="run_tests", detail="3 tests"),
|
||||
TelemetryEvent(timestamp=126, kind="test_pass",
|
||||
detail="3/3 passed"),
|
||||
TelemetryEvent(timestamp=180, kind="keystroke_burst", detail="tool_node.py"),
|
||||
TelemetryEvent(timestamp=260, kind="file_save", detail="tool_node.py"),
|
||||
TelemetryEvent(timestamp=275, kind="run_tests", detail="4 tests"),
|
||||
TelemetryEvent(timestamp=281, kind="test_pass",
|
||||
detail="4/4 passed"),
|
||||
TelemetryEvent(timestamp=340, kind="commit",
|
||||
detail="add tool node with retries"),
|
||||
],
|
||||
)
|
||||
|
||||
_LAB_SCENARIO_STRUGGLING = LabTelemetryScenario(
|
||||
scenario_id="lab-scenario-struggling",
|
||||
title="Struggling build session — repeated failures, no checkpoints",
|
||||
competency_id="stack-orchestration-c002",
|
||||
events=[
|
||||
TelemetryEvent(timestamp=0, kind="keystroke_burst", detail="main.py"),
|
||||
TelemetryEvent(timestamp=210, kind="run_tests", detail="2 tests"),
|
||||
TelemetryEvent(timestamp=215, kind="test_fail",
|
||||
detail="ImportError: no module named 'tools'"),
|
||||
TelemetryEvent(timestamp=216, kind="console_error", detail="traceback dumped"),
|
||||
TelemetryEvent(timestamp=300, kind="keystroke_burst",
|
||||
detail="main.py"),
|
||||
TelemetryEvent(timestamp=520, kind="run_tests",
|
||||
detail="2 tests"),
|
||||
TelemetryEvent(timestamp=525, kind="test_fail",
|
||||
detail="ImportError: no module named 'tools'"),
|
||||
TelemetryEvent(timestamp=526, kind="console_error",
|
||||
detail="same traceback as before"),
|
||||
TelemetryEvent(timestamp=600, kind="idle",
|
||||
detail="no activity for 6 minutes"),
|
||||
TelemetryEvent(timestamp=960, kind="idle",
|
||||
detail="no activity for 14 minutes"),
|
||||
],
|
||||
)
|
||||
|
||||
_LAB_SCENARIO_FLAGGED = LabTelemetryScenario(
|
||||
scenario_id="lab-scenario-flagged",
|
||||
title="Flagged build session — large paste, instant pass",
|
||||
competency_id="stack-orchestration-c003",
|
||||
events=[
|
||||
TelemetryEvent(timestamp=0, kind="keystroke_burst", detail="eval.py"),
|
||||
TelemetryEvent(timestamp=30, kind="paste",
|
||||
detail="2,400 chars pasted into eval.py"),
|
||||
TelemetryEvent(timestamp=45, kind="run_tests", detail="6 tests"),
|
||||
TelemetryEvent(timestamp=47, kind="test_pass", detail="6/6 passed"),
|
||||
TelemetryEvent(timestamp=48, kind="commit",
|
||||
detail="finish eval harness"),
|
||||
],
|
||||
)
|
||||
|
||||
LAB_SCENARIOS: dict[str, LabTelemetryScenario] = {
|
||||
s.scenario_id: s
|
||||
for s in (_LAB_SCENARIO_STRONG, _LAB_SCENARIO_STRUGGLING, _LAB_SCENARIO_FLAGGED)
|
||||
}
|
||||
|
||||
DEFAULT_LAB_SCENARIO_ID = "lab-scenario-strong"
|
||||
|
||||
|
||||
def get_lab_scenario(scenario_id: str) -> LabTelemetryScenario | None:
|
||||
return LAB_SCENARIOS.get(scenario_id)
|
||||
|
||||
|
||||
def summarize_scenario(scenario: LabTelemetryScenario) -> str:
|
||||
"""Render the event timeline as compact text for prompt injection."""
|
||||
lines = [f"Session: {scenario.title} (competency {scenario.competency_id})"]
|
||||
for event in scenario.events:
|
||||
lines.append(f"t+{event.timestamp}s {event.kind}: {event.detail}".rstrip(": "))
|
||||
return "\n".join(lines)
|
||||
@@ -1,230 +0,0 @@
|
||||
"""Synthetic trace fixtures for grading calibration (Task 3-2-02, REQ-3-004).
|
||||
|
||||
Three builder archetypes as REAL `TelemetryEvent` traces (the grading
|
||||
engine's native input — these are NOT the v0.2 corpus's simplified
|
||||
`{timestamp, kind, detail}` event shapes):
|
||||
|
||||
strong builder — iterative debugging: small edits, tests early,
|
||||
failed runs closed by targeted fixes, eventual pass.
|
||||
lazy builder — one large paste, a single late test run, pass.
|
||||
(v0.2 alignment: the `lab-scenario-flagged` paste-
|
||||
and-run archetype; D-021.)
|
||||
struggling builder — many edit/test cycles, failures never close,
|
||||
never reaches a pass.
|
||||
|
||||
ID convention (D-021 alignment, documented in each fixture):
|
||||
v0.2 corpus scenario IDs are `<domain>-scenario-<slug>` (`lab-scenario-strong`,
|
||||
`lab-scenario-struggling`, `lab-scenario-flagged`, `proctor-scenario-*` — see
|
||||
corpus/telemetry.py). Grading fixtures carry ids string-aligned to that
|
||||
convention:
|
||||
|
||||
fixture id = "<scenario id>::<archetype>-trace"
|
||||
learner id = "learner-003" (a member of the corpus learner-00N id space;
|
||||
learner-001/002 exist in learner_context.py)
|
||||
|
||||
so a fixture is greppable against its v0.2 scenario counterpart while staying
|
||||
a distinct id space (a grading trace is a real event stream, not the v0.2
|
||||
mock scenario timeline — same convention, richer event kind set).
|
||||
|
||||
Instance hygiene: fixtures store event SPECS (plain tuples) and materialize
|
||||
FRESH `TelemetryEvent` instances on every `fixture.events` access. SQLModel
|
||||
rows carry SQLAlchemy instance state once a session has flushed them —
|
||||
re-adding the SAME instance to another store is a silent no-op, which would
|
||||
poison sequential test runs (a fixture ingested by test N would vanish for
|
||||
test N+1). Materializing per access keeps every consumer independent.
|
||||
|
||||
These fixtures exist for the CALIBRATION CONTRACT (tests/grading/
|
||||
test_calibration.py): the mock provider scripts archetype-mapped scores and
|
||||
the test asserts the ORDERING the rubric must eventually enforce. They are
|
||||
NOT an LLM quality benchmark — see the test module docstring for the honest
|
||||
scope statement.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from datetime import UTC, datetime, timedelta
|
||||
from typing import Final
|
||||
|
||||
from pydantic import BaseModel, Field
|
||||
|
||||
from ..telemetry.models import TelemetryEvent
|
||||
|
||||
_T0: Final = datetime(2026, 9, 12, 0, 0, 0, tzinfo=UTC)
|
||||
|
||||
#: Event spec: (seq, kind, payload, seconds-since-session-start).
|
||||
EventSpec = tuple[int, str, dict, float]
|
||||
|
||||
|
||||
class TraceFixture(BaseModel):
|
||||
"""One named synthetic trace + its v0.2 scenario alignment (D-021).
|
||||
|
||||
`event_specs` is the durable, session-state-free description; `events`
|
||||
materializes fresh TelemetryEvent rows from it on every access.
|
||||
"""
|
||||
|
||||
model_config = {"frozen": True}
|
||||
|
||||
fixture_id: str # "<v0.2 scenario id>::<archetype>-trace"
|
||||
archetype: str # "strong" | "lazy" | "struggling"
|
||||
aligned_scenario_id: str # the v0.2 corpus scenario this fixture mirrors
|
||||
competency_id: str # corpus competency id space (stack-orchestration-c00N)
|
||||
task_id: str # trace task id (grading operates on (learner, task))
|
||||
learner_id: str
|
||||
event_specs: tuple[EventSpec, ...] = Field(default=())
|
||||
|
||||
@property
|
||||
def events(self) -> list[TelemetryEvent]:
|
||||
"""FRESH TelemetryEvent instances — safe to ingest into any store.
|
||||
|
||||
Never cache these: an instance flushed by one SQLite session
|
||||
carries persistent identity, and re-appending it elsewhere no-ops.
|
||||
"""
|
||||
return [
|
||||
TelemetryEvent(
|
||||
learner_id=self.learner_id,
|
||||
task_id=self.task_id,
|
||||
seq=seq,
|
||||
kind=kind,
|
||||
payload=payload,
|
||||
ts=_T0 + timedelta(seconds=offset_s),
|
||||
sandbox_id="sbx-calibration",
|
||||
)
|
||||
for seq, kind, payload, offset_s in self.event_specs
|
||||
]
|
||||
|
||||
def description(self) -> str:
|
||||
return (
|
||||
f"{self.fixture_id} (archetype={self.archetype}, aligned="
|
||||
f"{self.aligned_scenario_id}, competency={self.competency_id})"
|
||||
)
|
||||
|
||||
|
||||
class _Builder:
|
||||
"""Seq-accurate event-spec builder for one (learner, task) pair."""
|
||||
|
||||
def __init__(self) -> None:
|
||||
self.specs: list[EventSpec] = []
|
||||
self._seq = 0
|
||||
|
||||
def add(self, kind: str, payload: dict | None, at: float) -> None:
|
||||
self.specs.append((self._seq, kind, payload or {}, at))
|
||||
self._seq += 1
|
||||
|
||||
|
||||
def _strong_builder_specs() -> list[EventSpec]:
|
||||
"""Iterative debugging: tests early, tight edit→test loops, eventual pass.
|
||||
|
||||
Mirrors `lab-scenario-strong` (v0.2: keystrokes → file_save → run_tests →
|
||||
test_pass, c002) at full telemetry fidelity — every failed cycle is
|
||||
closed by a targeted edit followed by a re-run that passes.
|
||||
"""
|
||||
b = _Builder()
|
||||
b.add("activity", {"state": "starting"}, 0)
|
||||
b.add("file_diff", {"path": "planner.py", "added": 14}, 95) # small edit
|
||||
b.add("command", {"cmd": "pytest -q tests/test_planner.py"}, 120) # tests EARLY
|
||||
b.add("test_result", {"passed": True, "exit_code": 0}, 126)
|
||||
b.add("file_diff", {"path": "tool_node.py", "added": 22}, 180)
|
||||
b.add("command", {"cmd": "pytest -q"}, 275)
|
||||
b.add("test_result", {"passed": False, "exit_code": 1}, 281) # honest failure
|
||||
b.add("file_diff", {"path": "tool_node.py", "added": 6, "removed": 2}, 340) # targeted fix
|
||||
b.add("command", {"cmd": "pytest -q"}, 430)
|
||||
b.add("test_result", {"passed": True, "exit_code": 0}, 436) # cycle CLOSED
|
||||
b.add("command", {"cmd": "git commit -m 'tool node with retries'"}, 500)
|
||||
return b.specs
|
||||
|
||||
|
||||
def _lazy_builder_specs() -> list[EventSpec]:
|
||||
"""Paste-and-run: one large paste, a single LATE test run, instant pass.
|
||||
|
||||
Mirrors `lab-scenario-flagged` (v0.2: paste of 2,400 chars → run_tests →
|
||||
instant 6/6 pass, c003) at full telemetry fidelity — zero iteration, zero
|
||||
verification during construction, one terminal test run only.
|
||||
"""
|
||||
b = _Builder()
|
||||
b.add("activity", {"state": "starting"}, 0)
|
||||
b.add("file_diff", {"path": "eval.py", "added": 240, "removed": 0}, 30) # one bulk paste
|
||||
b.add("file_diff", {"path": "README.md", "added": 12}, 40)
|
||||
b.add("command", {"cmd": "npm run build"}, 45)
|
||||
b.add("run_result", {"exit_code": 0, "ok": True}, 60)
|
||||
b.add("command", {"cmd": "pytest -q"}, 520) # single LATE test run
|
||||
b.add("test_result", {"passed": True, "exit_code": 0}, 540) # instant pass
|
||||
return b.specs
|
||||
|
||||
|
||||
def _struggling_builder_specs() -> list[EventSpec]:
|
||||
"""Many cycles, none close: repeated failures, no eventual pass.
|
||||
|
||||
Mirrors `lab-scenario-struggling` (v0.2: repeated identical ImportErrors,
|
||||
idle gaps, no checkpoint, c002) at full telemetry fidelity — edits happen
|
||||
between failures, but the same failure recurs; no pass is ever reached.
|
||||
"""
|
||||
b = _Builder()
|
||||
b.add("activity", {"state": "starting"}, 0)
|
||||
b.add("file_diff", {"path": "main.py", "added": 40}, 20)
|
||||
b.add("command", {"cmd": "pytest -q"}, 210)
|
||||
b.add("test_result", {"passed": False, "exit_code": 1}, 215) # ImportError
|
||||
b.add("file_diff", {"path": "main.py", "added": 8, "removed": 3}, 300)
|
||||
b.add("command", {"cmd": "pytest -q"}, 520)
|
||||
b.add("test_result", {"passed": False, "exit_code": 1}, 525) # SAME error
|
||||
b.add("file_diff", {"path": "main.py", "added": 5}, 610)
|
||||
b.add("command", {"cmd": "pytest -q"}, 960)
|
||||
b.add("test_result", {"passed": False, "exit_code": 1}, 965) # STILL failing
|
||||
b.add("activity", {"state": "idle"}, 1500) # long idle
|
||||
b.add("activity", {"state": "idle"}, 2200)
|
||||
return b.specs
|
||||
|
||||
|
||||
#: Calibration learner — a member of the corpus learner-00N id space (D-021;
|
||||
#: learner-001/002 live in corpus/learner_context.py; grading fixtures use
|
||||
#: a third id so calibration traces never collide with mock-context reads).
|
||||
CALIBRATION_LEARNER_ID: Final = "learner-003"
|
||||
|
||||
|
||||
STRONG_BUILDER: Final = TraceFixture(
|
||||
fixture_id="lab-scenario-strong::strong-trace",
|
||||
archetype="strong",
|
||||
aligned_scenario_id="lab-scenario-strong",
|
||||
competency_id="stack-orchestration-c002",
|
||||
task_id="task-calibration-strong",
|
||||
learner_id=CALIBRATION_LEARNER_ID,
|
||||
event_specs=tuple(_strong_builder_specs()),
|
||||
)
|
||||
|
||||
LAZY_BUILDER: Final = TraceFixture(
|
||||
# v0.2's paste-and-run archetype is the "flagged" lab scenario (D-021):
|
||||
# large paste → instant test pass. "lazy builder" is that behavior
|
||||
# without the proctor flag; the alignment is behavioral, documented here.
|
||||
fixture_id="lab-scenario-flagged::lazy-trace",
|
||||
archetype="lazy",
|
||||
aligned_scenario_id="lab-scenario-flagged",
|
||||
competency_id="stack-orchestration-c003",
|
||||
task_id="task-calibration-lazy",
|
||||
learner_id=CALIBRATION_LEARNER_ID,
|
||||
event_specs=tuple(_lazy_builder_specs()),
|
||||
)
|
||||
|
||||
STRUGGLING_BUILDER: Final = TraceFixture(
|
||||
fixture_id="lab-scenario-struggling::struggling-trace",
|
||||
archetype="struggling",
|
||||
aligned_scenario_id="lab-scenario-struggling",
|
||||
competency_id="stack-orchestration-c002",
|
||||
task_id="task-calibration-struggling",
|
||||
learner_id=CALIBRATION_LEARNER_ID,
|
||||
event_specs=tuple(_struggling_builder_specs()),
|
||||
)
|
||||
|
||||
TRACE_FIXTURES: Final[dict[str, TraceFixture]] = {
|
||||
f.fixture_id: f
|
||||
for f in (STRONG_BUILDER, LAZY_BUILDER, STRUGGLING_BUILDER)
|
||||
}
|
||||
|
||||
|
||||
def get_trace_fixture(fixture_id: str) -> TraceFixture | None:
|
||||
return TRACE_FIXTURES.get(fixture_id)
|
||||
|
||||
|
||||
def digest_of(fixture: TraceFixture):
|
||||
"""Compute the digest for a fixture (pure compute; test-side helper)."""
|
||||
from ..grading.features import compute_digest
|
||||
|
||||
return compute_digest(fixture.events)
|
||||
@@ -1,27 +0,0 @@
|
||||
"""Process-trace grading — trace digest, rubric engine, GradeStore (REQ-3-004).
|
||||
|
||||
Boundary rule (D-027): grading/ is an engine module — it never imports
|
||||
api/; the single sanctioned agents/ dependency is the shared D-020
|
||||
structured defense (agents/structured.py), imported module-direct in
|
||||
engine.py (see its docstring for why). features.py imports telemetry/;
|
||||
store.py imports config only; engine.py composes llm/ + telemetry/ +
|
||||
prompts/ + agents.structured.
|
||||
|
||||
Wave status: features.py (TraceDigest, compute_digest) + store.py
|
||||
(GradeRecord, GradeStore, SQLiteGradeStore) landed in Wave 1 (3-1-01 +
|
||||
3-1-02); engine.py (GradingEngine, RubricScore) is Wave 2 (3-2-01).
|
||||
"""
|
||||
|
||||
from .engine import GradingEngine, RubricScore
|
||||
from .features import TraceDigest, compute_digest
|
||||
from .store import GradeRecord, GradeStore, SQLiteGradeStore
|
||||
|
||||
__all__ = [
|
||||
"GradeRecord",
|
||||
"GradeStore",
|
||||
"GradingEngine",
|
||||
"RubricScore",
|
||||
"SQLiteGradeStore",
|
||||
"TraceDigest",
|
||||
"compute_digest",
|
||||
]
|
||||
@@ -1,313 +0,0 @@
|
||||
"""GradingEngine — rubric scoring over real process traces (REQ-3-004).
|
||||
|
||||
The Wave-2 composition of the grading stack:
|
||||
|
||||
trace completeness gate (G-4, FIRST — nothing is sent to any LLM
|
||||
when the gate trips) → TraceStore.get_trace → compute_digest (D-028)
|
||||
→ prompts.grading.render_trace_digest → D-020 4-layer structured
|
||||
defense (agents/structured.py — REUSED, composed, never duplicated)
|
||||
→ RubricScore validation → GradeStore persistence → GradeRecord.
|
||||
|
||||
Gate verdicts are FIRST-CLASS RESULTS, not exceptions (G-4 is binding):
|
||||
UNGRADABLE_TRACE_INCOMPLETE — seq gaps in the store OR the trace is
|
||||
flagged by TraceIntegrityMap (INCOMPLETE_FLOODED). The gap list /
|
||||
flag reason is surfaced in `scores` for API rendering, and the
|
||||
record is PERSISTED like any grade so a learner sees why no
|
||||
credential can be issued for this trace — the gate outcome is
|
||||
durable and auditable, not a transient error string.
|
||||
UNGRADABLE_EMPTY_TRACE — no events stored for the pair.
|
||||
On either verdict `scores.criteria` is empty and the LLM is never called.
|
||||
|
||||
DI (D-027/D-032 house pattern): the engine receives trace_store,
|
||||
grade_store, integrity and provider through the constructor and knows
|
||||
NOTHING of FastAPI — api/ composes it (Task 3-3-01). `model` is injected
|
||||
alongside the provider so tests script the mock against the production
|
||||
wiring without touching Settings.
|
||||
|
||||
Boundary (D-027): grading/ imports llm/ (provider protocol + Message
|
||||
type), telemetry/ (store + integrity map), prompts/ (rubric text) and
|
||||
ONLY agents.structured — the sanctioned shared D-020 defense. We import
|
||||
the MODULE directly (`from ..agents.structured import structured_completion`)
|
||||
rather than the `agents` package, mirroring how agents/base.py consumes
|
||||
it (same direct-module import): that keeps the dependency surface to
|
||||
exactly the two names the engine needs (structured_completion,
|
||||
StructuredOutputError) and avoids executing agents/__init__ re-exports
|
||||
(BaseAgent, registry, session store) that grading has no business
|
||||
loading — a side-effect-hygiene choice that keeps this import line
|
||||
grep-auditable as "the one agents dependency". grading/ never imports api/.
|
||||
|
||||
RubricScore placement (documented decision): the validated output model
|
||||
lives HERE, not in prompts/. The pydantic model is the engine's return
|
||||
CONTRACT (the shape GradeStore.scores must hold), while prompts/grading.py
|
||||
is pure prompt text + its mirror schema HINT string — the same split as
|
||||
agents/assessor.py (model + hint) but with the model owned by the engine
|
||||
module that validates it. Prompt files hold text, per the prompts/ house
|
||||
style; engine files hold typed contracts.
|
||||
"""
|
||||
|
||||
import logging
|
||||
from datetime import UTC, datetime
|
||||
from typing import TYPE_CHECKING, Final
|
||||
|
||||
from pydantic import BaseModel, ConfigDict, Field, field_validator
|
||||
|
||||
from ..agents.structured import StructuredOutputError, structured_completion
|
||||
from ..llm.base import LLMProvider
|
||||
from ..llm.types import Message
|
||||
from ..prompts.grading import (
|
||||
RUBRIC_CRITERIA,
|
||||
RUBRIC_SCORE_SCHEMA_HINT,
|
||||
SYSTEM_PROMPT,
|
||||
render_trace_digest,
|
||||
)
|
||||
from ..telemetry.ingest import TraceIntegrityMap
|
||||
from ..telemetry.store import TraceStore
|
||||
from .features import compute_digest
|
||||
from .store import GradeRecord, GradeStore
|
||||
|
||||
if TYPE_CHECKING: # pragma: no cover - protocol-only import for the optional
|
||||
from ..variants.store import VariantStore # noqa: TC001 (variant-blind without it)
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
#: First-class gate verdicts (G-4). GradeRecord.verdict values; non-empty
|
||||
#: by store contract. Rubric verdicts (mastered/developing/not_yet) ride in
|
||||
#: `scores.verdict` — `record.verdict` stays the machine-readable outcome.
|
||||
VERDICT_UNGRADABLE_INCOMPLETE: Final = "UNGRADABLE_TRACE_INCOMPLETE"
|
||||
VERDICT_UNGRADABLE_EMPTY: Final = "UNGRADABLE_EMPTY_TRACE"
|
||||
#: record.verdict for a successfully LLM-graded trace (the rubric verdict
|
||||
#: travels inside scores); keeps verdict non-empty for every persisted row.
|
||||
VERDICT_GRADED: Final = "GRADED"
|
||||
|
||||
#: Gate-detail keys surfaced in GradeRecord.scores (tests assert on these).
|
||||
_INTEGRITY_FLAG_KEY: Final = "integrity_flag"
|
||||
|
||||
|
||||
def _anchors_context(variant) -> str: # noqa: ANN001 - VariantRecord (duck-typed)
|
||||
"""Render the variant template's difficulty anchors for the grader prompt.
|
||||
|
||||
Contains only the template id + anchor numbers — no learner-identifying
|
||||
material (D-028 anonymity preserved; the digest-leak tests keep holding).
|
||||
Lazy template import: grading must not import variants/ at module load
|
||||
(variants/prompts import-cycle safety mirrors llm/ rules).
|
||||
"""
|
||||
from ..variants.templates import get_template
|
||||
|
||||
template = get_template(variant.template_id)
|
||||
if template is None:
|
||||
return f"template={variant.template_id} (anchors unavailable)"
|
||||
a = template.rubric_anchors
|
||||
return (
|
||||
f"template={template.id}; "
|
||||
f"expected_edit_count_band={list(a.expected_edit_count_band)}; "
|
||||
f"expected_min_test_runs={a.expected_min_test_runs}; "
|
||||
f"expected_error_fix_cycles_band={list(a.expected_error_fix_cycles_band)}"
|
||||
)
|
||||
_MISSING_SEQS_KEY: Final = "missing_seqs"
|
||||
|
||||
|
||||
class RubricScore(BaseModel):
|
||||
"""Validated LLM output: per-criterion 0-4 scores + strengths + gaps + verdict.
|
||||
|
||||
The D-020 schema for the grader: `structured_completion` parses the
|
||||
model reply into THIS shape (layer 3), retrying once with the
|
||||
validation error fed back (layer 4). Exact criteria set + 0-4 ranges
|
||||
are enforced here, so the scores dict persisted to GradeStore is
|
||||
always rubric-shaped no matter what the model produced.
|
||||
"""
|
||||
|
||||
model_config = ConfigDict(extra="forbid")
|
||||
|
||||
criteria: dict[str, int]
|
||||
strengths: list[str] = Field(min_length=1, max_length=2)
|
||||
gaps: list[str] = Field(min_length=1, max_length=2)
|
||||
verdict: str # "mastered" | "developing" | "not_yet"
|
||||
|
||||
@field_validator("criteria")
|
||||
@classmethod
|
||||
def _criteria_rubric_shaped(cls, value: dict[str, int]) -> dict[str, int]:
|
||||
"""Exact criteria keys (no extras, no omissions) and 0-4 scores."""
|
||||
expected = set(RUBRIC_CRITERIA)
|
||||
got = set(value)
|
||||
if got != expected:
|
||||
raise ValueError(
|
||||
f"criteria keys must be exactly {sorted(expected)}, got {sorted(got)}"
|
||||
)
|
||||
for key, score in value.items():
|
||||
if not 0 <= score <= 4:
|
||||
raise ValueError(f"criterion {key!r} must be within 0-4, got {score}")
|
||||
return value
|
||||
|
||||
@field_validator("verdict")
|
||||
@classmethod
|
||||
def _verdict_known(cls, value: str) -> str:
|
||||
allowed = {"mastered", "developing", "not_yet"}
|
||||
if value not in allowed:
|
||||
raise ValueError(f"verdict must be one of {sorted(allowed)}, got {value!r}")
|
||||
return value
|
||||
|
||||
|
||||
class GradingEngine:
|
||||
"""Scores a (learner_id, task_id) trace into a persisted GradeRecord."""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
trace_store: TraceStore,
|
||||
grade_store: GradeStore,
|
||||
integrity: TraceIntegrityMap,
|
||||
provider: LLMProvider,
|
||||
*,
|
||||
model: str = "gemma4:31b",
|
||||
variant_store: "VariantStore | None" = None,
|
||||
) -> None:
|
||||
self._trace_store = trace_store
|
||||
self._grade_store = grade_store
|
||||
self._integrity = integrity
|
||||
self._provider = provider
|
||||
self._model = model
|
||||
# Phase 4 (MH#4): optional variant lookup — when the graded task
|
||||
# derives from a generated variant, its template's difficulty anchors
|
||||
# ship to the grader prompt (same bar for every variant of the
|
||||
# template, a-5) and the variant seed is stamped on the record.
|
||||
# Optional so engine tests stay decoupled; main.py lifespan wires it.
|
||||
self._variant_store = variant_store
|
||||
|
||||
async def grade(self, learner_id: str, task_id: str) -> GradeRecord:
|
||||
"""Grade one trace; persist latest-state (GradeStore upserts); return it.
|
||||
|
||||
Gate FIRST (G-4): the LLM is only ever reached from the fully
|
||||
guarded path — no gate state can be masked by an LLM error.
|
||||
"""
|
||||
record = await self._grade(learner_id, task_id)
|
||||
self._grade_store.save(record)
|
||||
return record
|
||||
|
||||
# ------------------------------------------------------------------ core
|
||||
|
||||
async def _grade(self, learner_id: str, task_id: str) -> GradeRecord:
|
||||
# -- G-4 gate FIRST: integrity flag OR seq gaps. Ordering matters:
|
||||
# gaps() returns [] for an EMPTY trace, so the empty check below is
|
||||
# reachable only when no rows exist at all; a gapped or flooded
|
||||
# trace can never fall through to the LLM path.
|
||||
if self._integrity.is_incomplete(learner_id, task_id):
|
||||
reason = self._integrity.reason(learner_id, task_id) or "unknown"
|
||||
gaps = self._trace_store.gaps(learner_id, task_id)
|
||||
logger.info(
|
||||
"grade gate (G-4): %s/%s integrity-flagged (%s) — ungradable",
|
||||
learner_id,
|
||||
task_id,
|
||||
reason,
|
||||
)
|
||||
return self._ungradable(
|
||||
learner_id,
|
||||
task_id,
|
||||
detail={_INTEGRITY_FLAG_KEY: reason, _MISSING_SEQS_KEY: gaps},
|
||||
verdict=VERDICT_UNGRADABLE_INCOMPLETE,
|
||||
)
|
||||
gaps = self._trace_store.gaps(learner_id, task_id)
|
||||
if gaps:
|
||||
logger.info(
|
||||
"grade gate (G-4): %s/%s seq gaps %s — ungradable",
|
||||
learner_id,
|
||||
task_id,
|
||||
gaps,
|
||||
)
|
||||
return self._ungradable(
|
||||
learner_id,
|
||||
task_id,
|
||||
detail={_INTEGRITY_FLAG_KEY: None, _MISSING_SEQS_KEY: gaps},
|
||||
verdict=VERDICT_UNGRADABLE_INCOMPLETE,
|
||||
)
|
||||
|
||||
trace = self._trace_store.get_trace(learner_id, task_id)
|
||||
if not trace:
|
||||
logger.info(
|
||||
"grade gate: %s/%s empty trace — ungradable", learner_id, task_id
|
||||
)
|
||||
return self._ungradable(
|
||||
learner_id,
|
||||
task_id,
|
||||
detail={_INTEGRITY_FLAG_KEY: None, _MISSING_SEQS_KEY: []},
|
||||
verdict=VERDICT_UNGRADABLE_EMPTY,
|
||||
)
|
||||
|
||||
# -- Guarded path: digest (D-028) → prompt → D-020 4-layer defense.
|
||||
digest = compute_digest(trace)
|
||||
variant = self._lookup_variant(task_id)
|
||||
anchors_context = (
|
||||
_anchors_context(variant) if variant is not None else None
|
||||
)
|
||||
messages = [
|
||||
Message(role="system", content=SYSTEM_PROMPT),
|
||||
Message(
|
||||
role="user",
|
||||
content=render_trace_digest(digest, anchors_context=anchors_context),
|
||||
),
|
||||
]
|
||||
try:
|
||||
rubric = await structured_completion(
|
||||
self._provider,
|
||||
messages,
|
||||
model=self._model,
|
||||
schema=RubricScore,
|
||||
schema_hint=RUBRIC_SCORE_SCHEMA_HINT,
|
||||
)
|
||||
except StructuredOutputError as exc:
|
||||
# The trace was gradable but the model failed to produce valid
|
||||
# JSON within the D-020 budget (two attempts). Raise — the API
|
||||
# layer maps this to a 502 (assessor precedent). Persisting a
|
||||
# fabricated or partial grade here would violate the no-silent-
|
||||
# fallback rule: no credential-worthy record without a validated
|
||||
# RubricScore.
|
||||
raise StructuredOutputError(f"grading LLM failed for {task_id}: {exc}") from exc
|
||||
|
||||
logger.debug(
|
||||
"graded %s/%s: %s (model=%s)",
|
||||
learner_id,
|
||||
task_id,
|
||||
rubric.verdict,
|
||||
self._model,
|
||||
)
|
||||
return GradeRecord(
|
||||
learner_id=learner_id,
|
||||
task_id=task_id,
|
||||
variant_seed=variant.seed if variant is not None else None, # D-029
|
||||
digest=digest.model_dump(),
|
||||
scores=rubric.model_dump(),
|
||||
verdict=VERDICT_GRADED,
|
||||
model=self._model,
|
||||
created_at=datetime.now(tz=UTC),
|
||||
)
|
||||
|
||||
def _lookup_variant(self, task_id: str): # noqa: ANN202 - VariantRecord | None
|
||||
"""MH#4: resolve the graded task's variant (None when not variant-derived)."""
|
||||
if self._variant_store is None:
|
||||
return None
|
||||
return self._variant_store.get_by_task(task_id)
|
||||
|
||||
# ------------------------------------------------------------ gate record
|
||||
|
||||
@staticmethod
|
||||
def _ungradable(
|
||||
learner_id: str,
|
||||
task_id: str,
|
||||
*,
|
||||
detail: dict,
|
||||
verdict: str,
|
||||
) -> GradeRecord:
|
||||
"""Build a gate record: no digest (nothing was graded), gate detail
|
||||
surfaced in `scores` (the store allows an empty scores dict, but
|
||||
G-4 requires the gap list / flag reason surfaced — the detail IS the
|
||||
verdict's payload), model="none" (no LLM was involved; provenance
|
||||
stays honest).
|
||||
"""
|
||||
return GradeRecord(
|
||||
learner_id=learner_id,
|
||||
task_id=task_id,
|
||||
variant_seed=None,
|
||||
digest={},
|
||||
scores=detail,
|
||||
verdict=verdict,
|
||||
model="none",
|
||||
created_at=datetime.now(tz=UTC),
|
||||
)
|
||||
@@ -1,246 +0,0 @@
|
||||
"""Deterministic process-trace digest (D-028, REQ-3-004).
|
||||
|
||||
Pure compute — no LLM, no I/O. `compute_digest` reduces an ordered
|
||||
TelemetryEvent trace to a compact, bounded `TraceDigest` that is safe to
|
||||
embed in a grading prompt:
|
||||
|
||||
- FIXED fields + small histograms only; NO raw commands, NO file contents,
|
||||
NO payloads — the raw trace NEVER reaches the LLM (D-028), which also
|
||||
bounds the prompt-injection surface.
|
||||
- Tolerant to both live trace mixes: daemon-topology traces carry
|
||||
`activity` + `file_diff` kinds (workspace watcher), while REPL-driven
|
||||
traces carry `command`/`stdin`/`stdout`/`run_result`/`test_result`
|
||||
(P2 verification P1). Features derive from whatever kinds are present and
|
||||
never crash on absent kinds.
|
||||
|
||||
Feature semantics (conservative, deterministic):
|
||||
- test pass/fail counts + final status derive from `test_result` payloads
|
||||
when present, falling back to `run_result` exit codes (0 = pass).
|
||||
- an error/fix CYCLE = a failing run/test followed by >= 1 edit and then a
|
||||
later run/test (pass or fail) — the next observed result closes the cycle.
|
||||
- idle gaps = wall-clock gaps between consecutive events exceeding
|
||||
`idle_threshold_s` (default 120s): count + total seconds.
|
||||
- command category histogram classifies `command`-kind payloads: build /
|
||||
test / file / nav / debug / other.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from collections import Counter
|
||||
|
||||
from pydantic import BaseModel, Field
|
||||
|
||||
from ..telemetry.models import TelemetryEvent
|
||||
|
||||
_IDLE_DEFAULT_S: float = 120.0
|
||||
|
||||
_TEST_HINTS = ("test", "pytest", "vitest", "jest", "mocha", "unittest", "go test", "npm test")
|
||||
_BUILD_HINTS = ("make", "npm run build", "pip install", "pnpm", "cargo build", "gcc", "tsc")
|
||||
_DEBUG_HINTS = ("gdb", "pdb", "print(", "debug", "strace", "ltrace", "curl", "ping")
|
||||
_NAV_HINTS = ("ls", "cd", "pwd", "cat ", "grep ", "find", "rg ", "tree", "head", "tail", "less")
|
||||
_FILE_HINTS = ("mv ", "cp ", "rm ", "mkdir", "touch", "chmod", "nano", "vim", "sed -i", "tee ")
|
||||
|
||||
|
||||
class TraceDigest(BaseModel):
|
||||
"""Compact, bounded, LLM-safe summary of a process trace (D-028).
|
||||
|
||||
Fixed fields + small histograms. Serializes well under 4 KB; contains no
|
||||
raw commands, file contents, or event payloads.
|
||||
"""
|
||||
|
||||
model_config = {"frozen": True}
|
||||
|
||||
event_count: int = Field(ge=0)
|
||||
session_duration_s: float = Field(ge=0.0)
|
||||
edit_count: int = Field(ge=0)
|
||||
command_count: int = Field(ge=0)
|
||||
run_count: int = Field(ge=0)
|
||||
test_pass_count: int = Field(ge=0)
|
||||
test_fail_count: int = Field(ge=0)
|
||||
final_test_status: str = Field(pattern="^(pass|fail|none)$")
|
||||
first_test_pass_offset_s: float | None = None
|
||||
error_fix_cycles: int = Field(ge=0)
|
||||
mean_fix_latency_s: float | None = None
|
||||
idle_gap_count: int = Field(ge=0)
|
||||
idle_gap_total_s: float = Field(ge=0.0)
|
||||
command_categories: dict[str, int] = Field(default_factory=dict)
|
||||
kind_histogram: dict[str, int] = Field(default_factory=dict)
|
||||
|
||||
|
||||
def _event_pass_status(event: TelemetryEvent) -> bool | None:
|
||||
"""True (pass) / False (fail) / None (not a result event) for one event."""
|
||||
payload = event.payload or {}
|
||||
if event.kind == "test_result":
|
||||
if "passed" in payload:
|
||||
return bool(payload["passed"])
|
||||
if "exit_code" in payload:
|
||||
return int(payload["exit_code"]) == 0
|
||||
if "status" in payload:
|
||||
return str(payload["status"]).lower() in ("pass", "passed", "ok", "success")
|
||||
return None
|
||||
if event.kind == "run_result":
|
||||
if "exit_code" in payload:
|
||||
return int(payload["exit_code"]) == 0
|
||||
if "ok" in payload:
|
||||
return bool(payload["ok"])
|
||||
return None
|
||||
return None
|
||||
|
||||
|
||||
def _classify_command(text: str) -> str:
|
||||
lowered = text.lower()
|
||||
if any(h in lowered for h in _TEST_HINTS):
|
||||
return "test"
|
||||
if any(h in lowered for h in _BUILD_HINTS):
|
||||
return "build"
|
||||
if any(h in lowered for h in _DEBUG_HINTS):
|
||||
return "debug"
|
||||
if any(h in lowered for h in _NAV_HINTS):
|
||||
return "nav"
|
||||
if any(h in lowered for h in _FILE_HINTS):
|
||||
return "file"
|
||||
return "other"
|
||||
|
||||
|
||||
def _command_text(event: TelemetryEvent) -> str:
|
||||
payload = event.payload or {}
|
||||
return str(payload.get("cmd") or payload.get("command") or payload.get("line") or "")
|
||||
|
||||
|
||||
def compute_digest(
|
||||
trace: list[TelemetryEvent], *, idle_threshold_s: float = _IDLE_DEFAULT_S
|
||||
) -> TraceDigest:
|
||||
"""Reduce an ordered trace to a bounded digest. Never raises on odd input."""
|
||||
events = sorted(trace, key=lambda e: (e.seq, e.ts))
|
||||
if not events:
|
||||
return TraceDigest(
|
||||
event_count=0,
|
||||
session_duration_s=0.0,
|
||||
edit_count=0,
|
||||
command_count=0,
|
||||
run_count=0,
|
||||
test_pass_count=0,
|
||||
test_fail_count=0,
|
||||
final_test_status="none",
|
||||
first_test_pass_offset_s=None,
|
||||
error_fix_cycles=0,
|
||||
mean_fix_latency_s=None,
|
||||
idle_gap_count=0,
|
||||
idle_gap_total_s=0.0,
|
||||
command_categories={},
|
||||
kind_histogram={},
|
||||
)
|
||||
|
||||
kind_histogram = Counter(e.kind for e in events)
|
||||
start_ts = events[0].ts
|
||||
end_ts = events[-1].ts
|
||||
duration = max(0.0, (end_ts - start_ts).total_seconds())
|
||||
|
||||
edit_count = kind_histogram.get("file_diff", 0)
|
||||
command_count = kind_histogram.get("command", 0)
|
||||
run_count = kind_histogram.get("run_result", 0)
|
||||
|
||||
# Tests: prefer test_result events; fall back to run_result exit codes.
|
||||
test_statuses: list[tuple[TelemetryEvent, bool]] = []
|
||||
for e in events:
|
||||
if e.kind == "test_result":
|
||||
ok = _event_pass_status(e)
|
||||
if ok is not None:
|
||||
test_statuses.append((e, ok))
|
||||
if not test_statuses:
|
||||
for e in events:
|
||||
if e.kind == "run_result":
|
||||
ok = _event_pass_status(e)
|
||||
if ok is not None:
|
||||
test_statuses.append((e, ok))
|
||||
|
||||
test_pass_count = sum(1 for _, ok in test_statuses if ok)
|
||||
test_fail_count = len(test_statuses) - test_pass_count
|
||||
if not test_statuses:
|
||||
final_test_status = "none"
|
||||
else:
|
||||
final_test_status = "pass" if test_statuses[-1][1] else "fail"
|
||||
first_pass = next((e for e, ok in test_statuses if ok), None)
|
||||
first_pass_offset = (
|
||||
max(0.0, (first_pass.ts - start_ts).total_seconds()) if first_pass is not None else None
|
||||
)
|
||||
|
||||
# Error/fix cycles: a failing result starts a pending cycle; the NEXT
|
||||
# observed result closes it (regardless of outcome) — a fix attempt that
|
||||
# fails again is itself another iteration of debugging, so it closes the
|
||||
# previous cycle and opens a new one. Edits since the fail mark the
|
||||
# close as a genuine fix attempt; latency = first edit -> closing result.
|
||||
cycles = 0
|
||||
fix_latencies: list[float] = []
|
||||
pending_fail_ts: float | None = None # seconds since start
|
||||
edits_since_fail = 0
|
||||
first_edit_ts: float | None = None
|
||||
for e in events:
|
||||
t = max(0.0, (e.ts - start_ts).total_seconds())
|
||||
if e.kind == "file_diff":
|
||||
if pending_fail_ts is not None:
|
||||
if edits_since_fail == 0:
|
||||
first_edit_ts = t
|
||||
edits_since_fail += 1
|
||||
continue
|
||||
ok = _event_pass_status(e)
|
||||
if ok is None:
|
||||
continue
|
||||
if ok is False:
|
||||
if pending_fail_ts is not None and edits_since_fail > 0 and first_edit_ts is not None:
|
||||
# failed fix attempt: closes the previous cycle, opens a new one
|
||||
cycles += 1
|
||||
fix_latencies.append(t - first_edit_ts)
|
||||
pending_fail_ts = t
|
||||
edits_since_fail = 0
|
||||
first_edit_ts = None
|
||||
continue
|
||||
if ok is True and pending_fail_ts is not None:
|
||||
if edits_since_fail > 0 and first_edit_ts is not None:
|
||||
cycles += 1
|
||||
fix_latencies.append(t - first_edit_ts)
|
||||
pending_fail_ts = None
|
||||
edits_since_fail = 0
|
||||
first_edit_ts = None
|
||||
|
||||
mean_fix_latency = (
|
||||
sum(fix_latencies) / len(fix_latencies) if fix_latencies else None
|
||||
)
|
||||
|
||||
# Idle gaps between consecutive events.
|
||||
idle_gap_count = 0
|
||||
idle_gap_total = 0.0
|
||||
prev_ts = None
|
||||
for e in events:
|
||||
if prev_ts is not None:
|
||||
gap = (e.ts - prev_ts).total_seconds()
|
||||
if gap > idle_threshold_s:
|
||||
idle_gap_count += 1
|
||||
idle_gap_total += gap
|
||||
prev_ts = e.ts
|
||||
|
||||
# Command category histogram (command-kind events only).
|
||||
categories: Counter[str] = Counter()
|
||||
for e in events:
|
||||
if e.kind == "command":
|
||||
categories[_classify_command(_command_text(e))] += 1
|
||||
|
||||
return TraceDigest(
|
||||
event_count=len(events),
|
||||
session_duration_s=round(duration, 3),
|
||||
edit_count=edit_count,
|
||||
command_count=command_count,
|
||||
run_count=run_count,
|
||||
test_pass_count=test_pass_count,
|
||||
test_fail_count=test_fail_count,
|
||||
final_test_status=final_test_status,
|
||||
first_test_pass_offset_s=(
|
||||
round(first_pass_offset, 3) if first_pass_offset is not None else None
|
||||
),
|
||||
error_fix_cycles=cycles,
|
||||
mean_fix_latency_s=(round(mean_fix_latency, 3) if mean_fix_latency is not None else None),
|
||||
idle_gap_count=idle_gap_count,
|
||||
idle_gap_total_s=round(idle_gap_total, 3),
|
||||
command_categories=dict(sorted(categories.items())),
|
||||
kind_histogram=dict(sorted(kind_histogram.items())),
|
||||
)
|
||||
@@ -1,237 +0,0 @@
|
||||
"""GradeStore — grade persistence protocol + SQLite implementation (REQ-3-004, D-027).
|
||||
|
||||
Postgres-migration-ready (D-027): the protocol is the only surface the
|
||||
grading engine and API layers touch; swapping SQLiteGradeStore for a
|
||||
Postgres-backed implementation must not change call sites. The
|
||||
`grade_record` table uses only portable column types (str / JSON /
|
||||
datetime), so the same SQLModel schema stands up unchanged on Postgres.
|
||||
|
||||
Upsert, NOT append: (learner_id, task_id) is the grade identity — one row
|
||||
per learner per task holding the LATEST grade. `save` overwrites the whole
|
||||
row when the pair already exists, so a regrade replaces scores, verdict,
|
||||
created_at, digest, model and variant_seed wholesale. That is deliberately
|
||||
the opposite of TraceStore.append's dedup-keep-first contract: a trace is an
|
||||
append-only event log, a grade is latest-state, so the engine can re-grade
|
||||
a task idempotently as its rubric or input evolves.
|
||||
|
||||
Concurrency (a-3): the engine enables WAL + synchronous=NORMAL and a busy
|
||||
timeout at connection time, so a regrade writer and API readers do not hit
|
||||
`database is locked` on the single-box pilot.
|
||||
|
||||
`created_at` contract: callers stamp UTC (datetime.now(UTC)); SQLite stores
|
||||
it naive and the read paths re-label it tz-aware UTC (same boundary
|
||||
normalization as TelemetryEvent.ts, so the contract holds on any backend).
|
||||
|
||||
Boundary (D-027): `grading/` never imports `agents/` / `api/`; this module
|
||||
imports config only.
|
||||
"""
|
||||
|
||||
import logging
|
||||
import sqlite3
|
||||
from collections.abc import Iterator
|
||||
from contextlib import contextmanager
|
||||
from datetime import UTC, datetime
|
||||
from pathlib import Path
|
||||
from typing import Any, Protocol
|
||||
|
||||
import sqlalchemy as sa
|
||||
from sqlalchemy import JSON, Index
|
||||
from sqlalchemy.orm import validates
|
||||
from sqlmodel import Field, Session, SQLModel, create_engine, select
|
||||
|
||||
from ..config import Settings
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class GradeRecord(SQLModel, table=True):
|
||||
"""A persisted grade; (learner_id, task_id) is the PK — latest wins.
|
||||
|
||||
Written by the grading engine (one save per grade attempt), read by the
|
||||
API layer through the GradeStore protocol. Constraint enforcement
|
||||
mirrors TelemetryEvent: sqlmodel 0.0.42's metaclass drops pydantic
|
||||
constraints on table models, so SQLAlchemy `@validates` hooks enforce
|
||||
instead and the column types stay Postgres-ready (D-027).
|
||||
|
||||
Field contract:
|
||||
learner_id — non-empty learner identifier (same id space as traces).
|
||||
task_id — non-empty task identifier; grade identity is the
|
||||
(learner_id, task_id) pair — the same pair as trace
|
||||
identity, so a grade is keyed by the exact trace it
|
||||
was computed from.
|
||||
variant_seed — task-variant seed (D-029); None when the graded task
|
||||
is not variant-derived. Since Phase 4 the engine
|
||||
stamps the graded variant's seed here (MH#4) and the
|
||||
template's difficulty anchors ship to the grader
|
||||
prompt — this column is the audit join for that.
|
||||
digest — compact deterministic trace digest (D-028) that fed
|
||||
the rubric prompt; persisted for auditability so the
|
||||
LLM's input stays reproducible.
|
||||
scores — validated rubric scores (per-criterion 0-4,
|
||||
strengths, gaps); JSON dict. An empty dict is legal
|
||||
(e.g. an UNGRADABLE_TRACE_INCOMPLETE record carries a
|
||||
verdict but no scores).
|
||||
verdict — first-class verdict string (rubric verdict or
|
||||
UNGRADABLE_TRACE_INCOMPLETE); non-empty.
|
||||
model — provider model that produced the scores (provenance).
|
||||
created_at — UTC grade timestamp; a regrade replaces it (latest
|
||||
save wins).
|
||||
"""
|
||||
|
||||
__tablename__ = "grade_record"
|
||||
# The composite PK covers (learner_id, task_id) point lookups; this
|
||||
# secondary index covers list_for_learner ordered by created_at without
|
||||
# a sort step (Postgres migration target D-027).
|
||||
__table_args__ = (
|
||||
Index("ix_grade_record_learner_created", "learner_id", "created_at"),
|
||||
)
|
||||
|
||||
learner_id: str = Field(primary_key=True)
|
||||
task_id: str = Field(primary_key=True)
|
||||
# None only for non-variant tasks (MH#4 stamps variant seeds since P4).
|
||||
variant_seed: str | None = Field(default=None)
|
||||
# JSON columns: stored as TEXT on SQLite, native JSONB on Postgres (D-027).
|
||||
digest: dict[str, Any] = Field(default_factory=dict, sa_type=JSON)
|
||||
scores: dict[str, Any] = Field(default_factory=dict, sa_type=JSON)
|
||||
verdict: str
|
||||
model: str
|
||||
created_at: datetime
|
||||
|
||||
@validates("learner_id", "task_id")
|
||||
def _ids_non_empty(self, key: str, value: str) -> str:
|
||||
if not value:
|
||||
raise ValueError(f"{key} must be a non-empty identifier")
|
||||
return value
|
||||
|
||||
@validates("verdict")
|
||||
def _verdict_non_empty(self, key: str, value: str) -> str:
|
||||
if not value:
|
||||
raise ValueError(f"{key} must be a non-empty verdict string")
|
||||
return value
|
||||
|
||||
|
||||
class GradeStore(Protocol):
|
||||
"""Persistence contract for latest-state grades per (learner_id, task_id).
|
||||
|
||||
Implemented by SQLiteGradeStore (v0.3, D-027); a Postgres implementation
|
||||
must satisfy the same surface.
|
||||
"""
|
||||
|
||||
def save(self, grade: GradeRecord) -> None:
|
||||
"""Persist a grade. UPSERT on (learner_id, task_id): a regrade with
|
||||
the same pair REPLACES the stored row wholesale — the latest grade
|
||||
wins. NOT append-only; contrast TraceStore.append, which is
|
||||
dedup-keep-first for at-least-once ingest.
|
||||
"""
|
||||
...
|
||||
|
||||
def get(self, learner_id: str, task_id: str) -> GradeRecord | None:
|
||||
"""Latest stored grade for the pair; None when none exists.
|
||||
|
||||
Detached from any DB session — safe to pass across layers.
|
||||
"""
|
||||
...
|
||||
|
||||
def list_for_learner(self, learner_id: str) -> list[GradeRecord]:
|
||||
"""All stored grades for the learner, ordered by created_at
|
||||
ascending (chronological). Empty list when the learner has none.
|
||||
"""
|
||||
...
|
||||
|
||||
def close(self) -> None:
|
||||
"""Release DB connections. Store must not be used after close."""
|
||||
...
|
||||
|
||||
|
||||
def _sqlite_connect(dbapi_connection: sqlite3.Connection, _: object) -> None:
|
||||
"""Per-connection pragma setup (a-3). Mirrors telemetry/store.py.
|
||||
|
||||
journal_mode=WAL — readers never block the single writer.
|
||||
synchronous=NORMAL — safe in WAL mode, avoids full fsync-per-commit.
|
||||
busy_timeout=5000 — retry briefly under contention instead of
|
||||
`OperationalError: database is locked`.
|
||||
"""
|
||||
cursor = dbapi_connection.cursor()
|
||||
cursor.execute("PRAGMA journal_mode=WAL")
|
||||
cursor.execute("PRAGMA synchronous=NORMAL")
|
||||
cursor.execute("PRAGMA busy_timeout=5000")
|
||||
cursor.close()
|
||||
|
||||
|
||||
def _as_utc(ts: datetime) -> datetime:
|
||||
"""Timestamps travel as naive datetime on SQLite, tz-aware elsewhere.
|
||||
|
||||
SQLite (via SQLModel) drops tzinfo; Postgres TIMESTAMP WITH TIME ZONE
|
||||
keeps it. Normalizing on the read path makes the store's contract
|
||||
tz-aware UTC regardless of the backend (D-027).
|
||||
"""
|
||||
if ts.tzinfo is None:
|
||||
return ts.replace(tzinfo=UTC) # naive UTC read-side: label as UTC
|
||||
return ts.astimezone(UTC)
|
||||
|
||||
|
||||
class SQLiteGradeStore:
|
||||
"""SQLite-backed GradeStore (SQLModel). Second protocol-wrapped store
|
||||
of the D-027 family (first: SQLiteTraceStore).
|
||||
"""
|
||||
|
||||
def __init__(self, db_path: Path | None = None) -> None:
|
||||
self._db_path: Path = db_path if db_path is not None else Settings().db_path
|
||||
self._engine = create_engine(f"sqlite:///{self._db_path}")
|
||||
sa.event.listen(self._engine, "connect", _sqlite_connect)
|
||||
SQLModel.metadata.create_all(self._engine)
|
||||
|
||||
@contextmanager
|
||||
def _session(self) -> Iterator[Session]:
|
||||
# expire_on_commit=False: identical session behavior to
|
||||
# SQLiteTraceStore. save() discards the merged instance and the read
|
||||
# paths never commit, but a uniform flag across the D-027 stores
|
||||
# keeps their detachment guarantees from diverging.
|
||||
with Session(self._engine, expire_on_commit=False) as session:
|
||||
yield session
|
||||
|
||||
def save(self, grade: GradeRecord) -> None:
|
||||
# `merge` = SELECT-by-PK then UPDATE or INSERT — exactly the upsert
|
||||
# contract. The trace store deliberately avoids merge (its append is
|
||||
# dedup-keep-first); here latest-wins IS the contract, so merge is
|
||||
# the right tool. The caller's object is never attached to the
|
||||
# session and stays usable (unexpired) after save.
|
||||
with self._session() as session:
|
||||
session.merge(grade)
|
||||
session.commit()
|
||||
logger.debug(
|
||||
"grade saved (regrade overwrites): %s/%s verdict=%s model=%s",
|
||||
grade.learner_id,
|
||||
grade.task_id,
|
||||
grade.verdict,
|
||||
grade.model,
|
||||
)
|
||||
|
||||
def get(self, learner_id: str, task_id: str) -> GradeRecord | None:
|
||||
with self._session() as session:
|
||||
record = session.get(GradeRecord, (learner_id, task_id))
|
||||
if record is None:
|
||||
return None
|
||||
record.created_at = _as_utc(record.created_at)
|
||||
# Detach from the session: callers must not depend on
|
||||
# open-session ORM magic (lazy loads fail once it closes).
|
||||
session.expunge(record)
|
||||
return record
|
||||
|
||||
def list_for_learner(self, learner_id: str) -> list[GradeRecord]:
|
||||
with self._session() as session:
|
||||
stmt = (
|
||||
select(GradeRecord)
|
||||
.where(GradeRecord.learner_id == learner_id)
|
||||
# Chronological; task_id is a deterministic tie-break for
|
||||
# grades stamped within the same instant.
|
||||
.order_by(GradeRecord.created_at, GradeRecord.task_id)
|
||||
)
|
||||
results = session.exec(stmt).all()
|
||||
for row in results:
|
||||
row.created_at = _as_utc(row.created_at)
|
||||
session.expunge(row)
|
||||
return list(results)
|
||||
|
||||
def close(self) -> None:
|
||||
self._engine.dispose()
|
||||
@@ -1,24 +0,0 @@
|
||||
"""Identity engine package — provider protocol + store (REQ-5-003, D-042)."""
|
||||
|
||||
from .base import (
|
||||
AgeBand,
|
||||
IdentityProvider,
|
||||
IdentityStatus,
|
||||
IdentitySubmission,
|
||||
IdentityVerdict,
|
||||
)
|
||||
from .mock import MockIdentityProvider, derive_age_band
|
||||
from .store import IdentityRecord, IdentityStore, SQLiteIdentityStore
|
||||
|
||||
__all__ = [
|
||||
"AgeBand",
|
||||
"IdentityStatus",
|
||||
"IdentitySubmission",
|
||||
"IdentityVerdict",
|
||||
"IdentityProvider",
|
||||
"IdentityRecord",
|
||||
"IdentityStore",
|
||||
"SQLiteIdentityStore",
|
||||
"MockIdentityProvider",
|
||||
"derive_age_band",
|
||||
]
|
||||
@@ -1,64 +0,0 @@
|
||||
"""Identity verification protocol (REQ-5-003, D-042, A-303).
|
||||
|
||||
Provider-agnostic like LLMProvider/VoiceProvider (D-014/D-030): a narrow
|
||||
protocol the identity API composes via DI, a deterministic mock, and a
|
||||
future real KYC vendor (Stripe Identity / Persona / Onfido class) that
|
||||
drops in without API changes. PII rules (A-305): the provider sees
|
||||
document REFERENCES, never raw documents; verdicts carry a mock marker
|
||||
(A-304) so downstream surfaces never display mock-verified as
|
||||
production-verified.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Literal, Protocol, runtime_checkable
|
||||
|
||||
from pydantic import BaseModel, ConfigDict, Field
|
||||
|
||||
AgeBand = Literal["16-17", "18+"]
|
||||
IdentityStatus = Literal["pending", "verified", "rejected"]
|
||||
|
||||
|
||||
class IdentitySubmission(BaseModel):
|
||||
"""What a learner submits: derived data + document refs only.
|
||||
|
||||
`date_of_birth` is a REAL date (the provider derives the age band) but
|
||||
raw DOB is NEVER persisted — only the derived band (A-305). Document
|
||||
refs are opaque handles (upload ids), never contents.
|
||||
"""
|
||||
|
||||
model_config = ConfigDict(frozen=True)
|
||||
|
||||
learner_id: str = Field(min_length=1)
|
||||
date_of_birth: str = Field(description="ISO date; used to derive age_band, never stored")
|
||||
document_refs: list[str] = Field(
|
||||
default_factory=list,
|
||||
description="Opaque upload handles; raw documents are never stored",
|
||||
)
|
||||
|
||||
|
||||
class IdentityVerdict(BaseModel):
|
||||
"""Provider verdict — what gets stored + surfaced."""
|
||||
|
||||
model_config = ConfigDict(frozen=True)
|
||||
|
||||
status: IdentityStatus
|
||||
age_band: AgeBand | None = None
|
||||
provider: str
|
||||
#: A-304 honesty: mock verdicts carry mock=True forever — downstream
|
||||
#: surfaces must never treat a mock verdict as production-verified.
|
||||
mock: bool = True
|
||||
detail: str = ""
|
||||
|
||||
|
||||
@runtime_checkable
|
||||
class IdentityProvider(Protocol):
|
||||
"""The KYC port: submit → (poll) → verdict. Never imports api/."""
|
||||
|
||||
async def submit(self, submission: IdentitySubmission) -> str:
|
||||
"""Start verification; returns a submission id (minted once)."""
|
||||
...
|
||||
|
||||
async def poll(self, submission_id: str) -> IdentityVerdict:
|
||||
"""Fetch the (possibly pending) verdict for a submission."""
|
||||
...
|
||||
@@ -1,94 +0,0 @@
|
||||
"""Deterministic mock identity provider (REQ-5-003, A-303).
|
||||
|
||||
Approve-on-policy: every submission verifies unless the caller scripts a
|
||||
rejection (by learner id) or the derived age band fails the floor
|
||||
(under-16 → rejected with an age detail). Verdicts are mock-marked (A-304)
|
||||
— the marker rides every verdict so no downstream surface can ever
|
||||
display mock-verified as production-verified. Never calls the network.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import itertools
|
||||
import secrets
|
||||
from datetime import UTC, datetime
|
||||
|
||||
from .base import IdentitySubmission, IdentityVerdict
|
||||
|
||||
|
||||
def derive_age_band(date_of_birth: str, today: datetime | None = None) -> str:
|
||||
"""Derive the age band from an ISO date. Under-16 returns "under-16".
|
||||
|
||||
Pure + deterministic; used by the store test and the API alike.
|
||||
"""
|
||||
dob = datetime.fromisoformat(date_of_birth)
|
||||
now = today or datetime.now(UTC)
|
||||
age = now.year - dob.year - (
|
||||
(now.month, now.day) < (dob.month, dob.day)
|
||||
)
|
||||
if age < 16:
|
||||
return "under-16"
|
||||
if age < 18:
|
||||
return "16-17"
|
||||
return "18+"
|
||||
|
||||
|
||||
class MockIdentityProvider:
|
||||
"""Scriptable, deterministic; no network, no vendor calls.
|
||||
|
||||
Submission ids are minted UNIQUELY per submit() call (a monotonic
|
||||
counter + the per-process seed from `secrets`): the id is the PK of
|
||||
the insert-only IdentityStore, and a deterministic id derived from
|
||||
(learner_id, date_of_birth) collides on any resubmit-after-terminal
|
||||
(e.g. a rejected learner retrying with the same DOB) — the store
|
||||
surfaces IntegrityError and the API would 500 (cross-phase P0,
|
||||
final review). Uniqueness per call is the contract; determinism of
|
||||
VERDICTS (what tests actually pin) is preserved — poll() derives the
|
||||
band purely from the stored submission.
|
||||
"""
|
||||
|
||||
def __init__(self, reject_learners: set[str] | None = None) -> None:
|
||||
self._submissions: dict[str, IdentitySubmission] = {}
|
||||
self._reject_learners = reject_learners or set()
|
||||
# Per-process nonce: ids are opaque handles (A-305) — never
|
||||
# derived from PII. Counter + nonce keeps ids unique within and
|
||||
# across provider instances on one box.
|
||||
self._nonce = secrets.randbits(32)
|
||||
self._counter = itertools.count()
|
||||
|
||||
async def submit(self, submission: IdentitySubmission) -> str:
|
||||
submission_id = f"idc-{self._nonce:08x}{next(self._counter):08x}"
|
||||
self._submissions[submission_id] = submission
|
||||
return submission_id
|
||||
|
||||
async def poll(self, submission_id: str) -> IdentityVerdict:
|
||||
submission = self._submissions.get(submission_id)
|
||||
if submission is None:
|
||||
return IdentityVerdict(
|
||||
status="rejected",
|
||||
provider="mock",
|
||||
mock=True,
|
||||
detail="unknown submission id",
|
||||
)
|
||||
if submission.learner_id in self._reject_learners:
|
||||
return IdentityVerdict(
|
||||
status="rejected",
|
||||
provider="mock",
|
||||
mock=True,
|
||||
detail="scripted rejection (test)",
|
||||
)
|
||||
band = derive_age_band(submission.date_of_birth)
|
||||
if band == "under-16":
|
||||
return IdentityVerdict(
|
||||
status="rejected",
|
||||
provider="mock",
|
||||
mock=True,
|
||||
detail="under 16 — the AI school floor is 16+ (COPPA avoidance)",
|
||||
)
|
||||
return IdentityVerdict(
|
||||
status="verified",
|
||||
age_band=band, # type: ignore[arg-type]
|
||||
provider="mock",
|
||||
mock=True,
|
||||
detail="mock verdict — not production verification",
|
||||
)
|
||||
@@ -1,214 +0,0 @@
|
||||
"""IdentityStore — verification records (REQ-5-003, D-042, D-027 FIFTH store).
|
||||
|
||||
Insert-only + latest-per-learner lookup, modeled on the DefenseStore
|
||||
conventions: WAL + synchronous=NORMAL + busy_timeout + foreign_keys=ON
|
||||
pragmas at connect time, portable column types (str/datetime/JSON) for
|
||||
Postgres parity, @validates hooks for constraints sqlmodel's metaclass
|
||||
drops, tz-aware→naive→tz-aware boundary normalization.
|
||||
|
||||
PII contract (A-305): stores the DERIVED age_band (16-17 | 18+), NEVER a
|
||||
raw date of birth; document_refs are opaque handles, NEVER contents.
|
||||
Verdict provenance is audit data: every record carries provider + the
|
||||
mock marker (A-304) so downstream surfaces can label unverified state
|
||||
honestly.
|
||||
|
||||
Insert-only growth is fine at pilot scale (a-12): learner_id indexed,
|
||||
latest-per-learner lookup, no compaction pre-vendor.
|
||||
|
||||
Boundary (D-027): `identity/` never imports `agents/` / `api/`; this
|
||||
module imports config only.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
from datetime import UTC, datetime
|
||||
from pathlib import Path
|
||||
from typing import Any, Protocol
|
||||
|
||||
from sqlalchemy import event, text
|
||||
from sqlalchemy.types import JSON, String
|
||||
from sqlmodel import Field, Session, SQLModel, create_engine
|
||||
|
||||
from ..config import Settings
|
||||
from .base import IdentityStatus
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class IdentityRecord(SQLModel, table=True):
|
||||
"""One verification submission's lifecycle + verdict provenance."""
|
||||
|
||||
__tablename__ = "identity_record"
|
||||
|
||||
#: PK = the submission id minted once by the provider's submit().
|
||||
id: str = Field(primary_key=True)
|
||||
learner_id: str = Field(index=True)
|
||||
# Bare Literal annotations crash sqlmodel's column inference; explicit
|
||||
# sa_type + the validates hook below give the same contract
|
||||
# (VARCHAR column, Literal-rejected values — DefenseStore pattern).
|
||||
status: IdentityStatus = Field(default="pending", sa_type=String)
|
||||
provider: str
|
||||
#: Verdict provenance: the provider's raw verdict (mock-marked).
|
||||
verdict: dict[str, Any] = Field(default_factory=dict, sa_type=JSON)
|
||||
#: DERIVED band only — raw DOB is never persisted (A-305).
|
||||
age_band: str | None = Field(default=None, sa_type=String)
|
||||
#: Opaque document handles — raw documents never stored (A-305).
|
||||
document_refs: list[str] = Field(default_factory=list, sa_type=JSON)
|
||||
submitted_at: datetime
|
||||
verified_at: datetime | None = Field(default=None)
|
||||
|
||||
@property
|
||||
def mock(self) -> bool:
|
||||
"""A-304: the mock marker rides every surface (record + API)."""
|
||||
return bool(self.verdict.get("mock", True))
|
||||
|
||||
def _validate(self) -> None:
|
||||
if not self.id or not self.learner_id:
|
||||
raise ValueError("id and learner_id must be non-empty")
|
||||
if self.status not in ("pending", "verified", "rejected"):
|
||||
raise ValueError(f"invalid identity status {self.status!r}")
|
||||
# D3 (verifier): a stored band must be canonical or None (pending).
|
||||
# The gate fails closed on anything else; the store refuses to
|
||||
# create it in the first place.
|
||||
if self.age_band is not None and self.age_band not in (
|
||||
"16-17",
|
||||
"18+",
|
||||
"under-16",
|
||||
):
|
||||
raise ValueError(f"invalid age_band {self.age_band!r}")
|
||||
|
||||
def _normalize(self) -> None:
|
||||
self.submitted_at = _as_utc(self.submitted_at)
|
||||
if self.verified_at is not None:
|
||||
self.verified_at = _as_utc(self.verified_at)
|
||||
|
||||
|
||||
def _as_utc(ts: datetime) -> datetime:
|
||||
"""SQLite stores naive; read paths re-label tz-aware UTC (D-027 pattern)."""
|
||||
if ts.tzinfo is None:
|
||||
return ts.replace(tzinfo=UTC)
|
||||
return ts
|
||||
|
||||
|
||||
def _sqlite_connect(dbapi_connection: object, _: object) -> None:
|
||||
"""Per-connection pragmas — mirrors the other D-027 stores."""
|
||||
cursor = dbapi_connection.cursor() # type: ignore[attr-defined]
|
||||
cursor.execute("PRAGMA journal_mode=WAL")
|
||||
cursor.execute("PRAGMA synchronous=NORMAL")
|
||||
cursor.execute("PRAGMA busy_timeout=5000")
|
||||
cursor.execute("PRAGMA foreign_keys=ON")
|
||||
cursor.close() # type: ignore[attr-defined]
|
||||
|
||||
|
||||
class IdentityStore(Protocol):
|
||||
"""Persistence contract for identity records."""
|
||||
|
||||
def insert(self, record: IdentityRecord) -> IdentityRecord:
|
||||
"""INSERT-ONLY: a duplicate id raises IntegrityError (surfaced, not
|
||||
swallowed — a submission id is minted once)."""
|
||||
...
|
||||
|
||||
def get(self, submission_id: str) -> IdentityRecord | None:
|
||||
"""Point lookup by submission id."""
|
||||
...
|
||||
|
||||
def latest_for_learner(self, learner_id: str) -> IdentityRecord | None:
|
||||
"""Newest record for the learner (or None)."""
|
||||
...
|
||||
|
||||
def count_pending_for_learner(self, learner_id: str) -> int:
|
||||
"""G-13: active pending submissions (cap = 1)."""
|
||||
...
|
||||
|
||||
def mark_verified(
|
||||
self, submission_id: str, verdict: dict[str, Any], age_band: str | None
|
||||
) -> IdentityRecord | None:
|
||||
"""Terminal transition (verified or rejected): stamp + store the
|
||||
provider verdict + derived band. Unknown id → None."""
|
||||
...
|
||||
|
||||
|
||||
class SQLiteIdentityStore:
|
||||
"""SQLite implementation of IdentityStore (D-027)."""
|
||||
|
||||
def __init__(self, db_path: Path | None = None) -> None:
|
||||
self._db_path: Path = db_path if db_path is not None else Settings().db_path
|
||||
self._engine = create_engine(
|
||||
f"sqlite:///{self._db_path}",
|
||||
connect_args={"check_same_thread": False},
|
||||
)
|
||||
event.listen(self._engine, "connect", _sqlite_connect)
|
||||
SQLModel.metadata.create_all(self._engine)
|
||||
|
||||
def insert(self, record: IdentityRecord) -> IdentityRecord:
|
||||
record._validate()
|
||||
record._normalize()
|
||||
with Session(self._engine) as session:
|
||||
session.add(record)
|
||||
session.commit() # IntegrityError SURFACES (insert-only, minted-once)
|
||||
session.refresh(record)
|
||||
return record
|
||||
|
||||
def get(self, submission_id: str) -> IdentityRecord | None:
|
||||
with Session(self._engine) as session:
|
||||
rec = session.get(IdentityRecord, submission_id)
|
||||
if rec is None:
|
||||
return None
|
||||
session.refresh(rec)
|
||||
rec.submitted_at = _as_utc(rec.submitted_at)
|
||||
if rec.verified_at is not None:
|
||||
rec.verified_at = _as_utc(rec.verified_at)
|
||||
return rec
|
||||
|
||||
def latest_for_learner(self, learner_id: str) -> IdentityRecord | None:
|
||||
with Session(self._engine) as session:
|
||||
# D4 (verifier): submitted_at alone can tie at microsecond
|
||||
# resolution — sqlite rowid breaks the tie deterministically
|
||||
# (last inserted wins, mirroring insert-only chronology).
|
||||
# rowid is a SQLite physical column, not a SQLModel field — it
|
||||
# rides the query as raw text.
|
||||
rec = (
|
||||
session.query(IdentityRecord)
|
||||
.filter(IdentityRecord.learner_id == learner_id)
|
||||
.order_by(
|
||||
IdentityRecord.submitted_at.desc(),
|
||||
text("rowid DESC"),
|
||||
)
|
||||
.first()
|
||||
)
|
||||
if rec is not None:
|
||||
rec.submitted_at = _as_utc(rec.submitted_at)
|
||||
if rec.verified_at is not None:
|
||||
rec.verified_at = _as_utc(rec.verified_at)
|
||||
return rec
|
||||
|
||||
def count_pending_for_learner(self, learner_id: str) -> int:
|
||||
with Session(self._engine) as session:
|
||||
return (
|
||||
session.query(IdentityRecord)
|
||||
.filter(
|
||||
IdentityRecord.learner_id == learner_id,
|
||||
IdentityRecord.status == "pending",
|
||||
)
|
||||
.count()
|
||||
)
|
||||
|
||||
def mark_verified(
|
||||
self, submission_id: str, verdict: dict[str, Any], age_band: str | None
|
||||
) -> IdentityRecord | None:
|
||||
with Session(self._engine) as session:
|
||||
rec = session.get(IdentityRecord, submission_id)
|
||||
if rec is None:
|
||||
return None
|
||||
rec.verdict = verdict
|
||||
rec.age_band = age_band
|
||||
rec.verified_at = datetime.now(UTC)
|
||||
rec.status = "verified" if verdict.get("status") == "verified" else "rejected"
|
||||
rec._validate() # D3: transitions validate like inserts
|
||||
session.commit()
|
||||
session.refresh(rec)
|
||||
return rec
|
||||
|
||||
def close(self) -> None:
|
||||
self._engine.dispose()
|
||||
@@ -1,15 +0,0 @@
|
||||
"""LLM package — provider-agnostic layer (D-017)."""
|
||||
|
||||
from .base import LLMProvider
|
||||
from .factory import create_provider
|
||||
from .mock import MockProvider
|
||||
from .openai_compat import OpenAICompatProvider
|
||||
from .types import Message
|
||||
|
||||
__all__ = [
|
||||
"LLMProvider",
|
||||
"Message",
|
||||
"MockProvider",
|
||||
"OpenAICompatProvider",
|
||||
"create_provider",
|
||||
]
|
||||
@@ -1,35 +0,0 @@
|
||||
"""LLMProvider protocol — the port all agents depend on (D-017).
|
||||
|
||||
Implementations: openai_compat.OpenAICompatProvider (ollama-cloud + local),
|
||||
mock.MockProvider (deterministic, tests/CI). Providers are dumb pipes:
|
||||
no envelope logic here — the API layer owns meta/done/error events (D-016).
|
||||
"""
|
||||
|
||||
from collections.abc import AsyncIterator
|
||||
from typing import Protocol
|
||||
|
||||
from .types import Message
|
||||
|
||||
|
||||
class LLMProvider(Protocol):
|
||||
async def stream_chat(
|
||||
self,
|
||||
messages: list[Message],
|
||||
*,
|
||||
model: str,
|
||||
temperature: float = 0.7,
|
||||
response_format: dict | None = None,
|
||||
) -> AsyncIterator[str]:
|
||||
"""Yield incremental content deltas (plain text chunks)."""
|
||||
...
|
||||
|
||||
async def chat(
|
||||
self,
|
||||
messages: list[Message],
|
||||
*,
|
||||
model: str,
|
||||
temperature: float = 0.7,
|
||||
response_format: dict | None = None,
|
||||
) -> str:
|
||||
"""Non-streaming completion — returns the full reply text."""
|
||||
...
|
||||
@@ -1,34 +0,0 @@
|
||||
"""Provider factory — selects the LLM provider from settings (D-014)."""
|
||||
|
||||
import httpx
|
||||
|
||||
from ..config import Settings
|
||||
from .mock import MockProvider
|
||||
from .openai_compat import OpenAICompatProvider
|
||||
|
||||
PROVIDER_NAMES = ("ollama-cloud", "local", "mock")
|
||||
|
||||
|
||||
def create_provider(settings: Settings, http_client: httpx.AsyncClient):
|
||||
"""Return the provider instance for settings.provider.
|
||||
|
||||
Raises ValueError for unknown provider names.
|
||||
"""
|
||||
if settings.provider == "ollama-cloud":
|
||||
return OpenAICompatProvider(
|
||||
http_client=http_client,
|
||||
base_url=settings.ollama_cloud_base_url,
|
||||
api_key=settings.ollama_cloud_api_key,
|
||||
json_mode=settings.json_mode,
|
||||
)
|
||||
if settings.provider == "local":
|
||||
return OpenAICompatProvider(
|
||||
http_client=http_client,
|
||||
base_url=settings.local_base_url,
|
||||
json_mode=settings.json_mode,
|
||||
)
|
||||
if settings.provider == "mock":
|
||||
return MockProvider()
|
||||
raise ValueError(
|
||||
f"unknown provider {settings.provider!r}; expected one of {PROVIDER_NAMES}"
|
||||
)
|
||||
@@ -1,94 +0,0 @@
|
||||
"""Deterministic mock provider — tests and CI. NEVER calls the network.
|
||||
|
||||
Determinism: the reply text is seeded from the message content hash, so
|
||||
identical inputs always produce identical outputs. Supports scripted
|
||||
failure modes for error-path coverage (D-023, A-010).
|
||||
"""
|
||||
|
||||
import hashlib
|
||||
import json
|
||||
from collections.abc import AsyncIterator
|
||||
|
||||
from .types import Message
|
||||
|
||||
_REPLIES = [
|
||||
"Great question — let's break this down step by step and see where it leads.",
|
||||
"Here is the key idea: small, verified moves compound into mastery over time.",
|
||||
"Think about it this way: what would the simplest working version look like?",
|
||||
"You are closer than you think. Try restating the goal in one sentence first.",
|
||||
"Let me offer a different angle before we move to the next step.",
|
||||
]
|
||||
|
||||
_JSON_REPLY = '{"summary": "mock structured reply", "confidence": 0.87}'
|
||||
|
||||
|
||||
class MockProvider:
|
||||
"""Scripted provider: deterministic streams, no network, failure injection."""
|
||||
|
||||
def __init__(self) -> None:
|
||||
self.fail_before_first_token: bool = False
|
||||
self.fail_mid_stream_at_index: int | None = None
|
||||
self.abort_recorded: bool = False # set in stream finally-block (cancellation test)
|
||||
|
||||
def _reply_for(self, messages: list[Message], response_format: dict | None) -> str:
|
||||
seed_src = "|".join(f"{m.role}:{m.content}" for m in messages)
|
||||
if response_format is not None and response_format.get("type") == "json_object":
|
||||
return _JSON_REPLY
|
||||
digest = hashlib.sha256(seed_src.encode()).hexdigest()
|
||||
base = _REPLIES[int(digest[:2], 16) % len(_REPLIES)]
|
||||
# Deterministic seed tag guarantees distinct inputs → distinct replies
|
||||
return f"{base} [#{digest[:8]}]"
|
||||
|
||||
def _tokenize(self, text: str) -> list[str]:
|
||||
words = text.split(" ")
|
||||
tokens: list[str] = []
|
||||
for i, word in enumerate(words):
|
||||
suffix = " " if i < len(words) - 1 else ""
|
||||
tokens.append(word + suffix)
|
||||
return tokens
|
||||
|
||||
async def stream_chat(
|
||||
self,
|
||||
messages: list[Message],
|
||||
*,
|
||||
model: str,
|
||||
temperature: float = 0.7,
|
||||
response_format: dict | None = None,
|
||||
) -> AsyncIterator[str]:
|
||||
if self.fail_before_first_token:
|
||||
raise RuntimeError("mock provider: scripted failure before first token")
|
||||
reply = self._reply_for(messages, response_format)
|
||||
tokens = self._tokenize(reply)
|
||||
try:
|
||||
for i, token in enumerate(tokens):
|
||||
if self.fail_mid_stream_at_index is not None and i == self.fail_mid_stream_at_index:
|
||||
raise RuntimeError("mock provider: scripted mid-stream failure")
|
||||
yield token
|
||||
finally:
|
||||
# Cancellation (GeneratorExit/CancelledError) lands here — tests assert this.
|
||||
self.abort_recorded = True
|
||||
|
||||
async def chat(
|
||||
self,
|
||||
messages: list[Message],
|
||||
*,
|
||||
model: str,
|
||||
temperature: float = 0.7,
|
||||
response_format: dict | None = None,
|
||||
) -> str:
|
||||
if self.fail_before_first_token:
|
||||
raise RuntimeError("mock provider: scripted failure before completion")
|
||||
return self._reply_for(messages, response_format)
|
||||
|
||||
|
||||
class ScriptedJSONProvider(MockProvider):
|
||||
"""Mock variant returning a fixed JSON payload for structured tests."""
|
||||
|
||||
def __init__(self, payload: dict) -> None:
|
||||
super().__init__()
|
||||
self.payload = payload
|
||||
|
||||
def _reply_for(self, messages: list[Message], response_format: dict | None) -> str:
|
||||
if response_format is not None and response_format.get("type") == "json_object":
|
||||
return json.dumps(self.payload)
|
||||
return super()._reply_for(messages, response_format)
|
||||
@@ -1,125 +0,0 @@
|
||||
"""OpenAI-compatible provider — one implementation serves ollama-cloud AND local
|
||||
endpoints (they differ only in base_url/key). Raw httpx, no SDK (D-017).
|
||||
|
||||
Boundary rules:
|
||||
- llm/ imports nothing from agents/ or api/
|
||||
- api_key NEVER appears in exceptions, logs, or error messages
|
||||
"""
|
||||
|
||||
import json
|
||||
from collections.abc import AsyncIterator
|
||||
|
||||
import httpx
|
||||
|
||||
from .types import Message
|
||||
|
||||
|
||||
class OpenAICompatProvider:
|
||||
def __init__(
|
||||
self,
|
||||
http_client: httpx.AsyncClient,
|
||||
base_url: str,
|
||||
api_key: str = "",
|
||||
json_mode: str = "auto",
|
||||
) -> None:
|
||||
self._client = http_client
|
||||
self._base_url = base_url.rstrip("/")
|
||||
self._api_key = api_key
|
||||
self._json_mode = json_mode
|
||||
|
||||
def _headers(self) -> dict[str, str]:
|
||||
headers = {"Content-Type": "application/json"}
|
||||
if self._api_key:
|
||||
headers["Authorization"] = f"Bearer {self._api_key}"
|
||||
return headers
|
||||
|
||||
def _payload(
|
||||
self,
|
||||
messages: list[Message],
|
||||
model: str,
|
||||
temperature: float,
|
||||
response_format: dict | None,
|
||||
stream: bool,
|
||||
) -> dict:
|
||||
payload: dict = {
|
||||
"model": model,
|
||||
"messages": [{"role": m.role, "content": m.content} for m in messages],
|
||||
"temperature": temperature,
|
||||
}
|
||||
if stream:
|
||||
payload["stream"] = True
|
||||
else:
|
||||
payload["stream"] = False
|
||||
# json_mode="auto": send response_format and degrade on 400; "off": never send
|
||||
if response_format is not None and self._json_mode == "auto":
|
||||
payload["response_format"] = response_format
|
||||
return payload
|
||||
|
||||
def _sanitize(self, exc: Exception) -> RuntimeError:
|
||||
text = str(exc)
|
||||
if self._api_key and self._api_key in text:
|
||||
text = text.replace(self._api_key, "[REDACTED]")
|
||||
return RuntimeError(f"llm provider error: {text}")
|
||||
|
||||
async def stream_chat(
|
||||
self,
|
||||
messages: list[Message],
|
||||
*,
|
||||
model: str,
|
||||
temperature: float = 0.7,
|
||||
response_format: dict | None = None,
|
||||
) -> AsyncIterator[str]:
|
||||
payload = self._payload(messages, model, temperature, response_format, stream=True)
|
||||
try:
|
||||
async with self._client.stream(
|
||||
"POST", f"{self._base_url}/chat/completions",
|
||||
json=payload, headers=self._headers(),
|
||||
) as response:
|
||||
response.raise_for_status()
|
||||
async for line in response.aiter_lines():
|
||||
if not line or not line.startswith("data:"):
|
||||
continue # keep-alive comments (": ping"), empty lines
|
||||
data = line.removeprefix("data:").strip()
|
||||
if data == "[DONE]":
|
||||
return
|
||||
try:
|
||||
chunk = json.loads(data)
|
||||
except json.JSONDecodeError:
|
||||
continue # malformed line — tolerate (ollama-cloud quirks)
|
||||
choices = chunk.get("choices") or []
|
||||
if not choices:
|
||||
continue
|
||||
content = (choices[0].get("delta") or {}).get("content")
|
||||
if content:
|
||||
yield content
|
||||
except httpx.HTTPError as exc:
|
||||
raise self._sanitize(exc) from exc
|
||||
|
||||
async def chat(
|
||||
self,
|
||||
messages: list[Message],
|
||||
*,
|
||||
model: str,
|
||||
temperature: float = 0.7,
|
||||
response_format: dict | None = None,
|
||||
) -> str:
|
||||
payload = self._payload(messages, model, temperature, response_format, stream=False)
|
||||
try:
|
||||
response = await self._client.post(
|
||||
f"{self._base_url}/chat/completions",
|
||||
json=payload, headers=self._headers(),
|
||||
)
|
||||
if response.status_code == 400 and "response_format" in payload:
|
||||
# json_mode auto-degrade (D-020 layer 1): retry once without it
|
||||
payload.pop("response_format")
|
||||
response = await self._client.post(
|
||||
f"{self._base_url}/chat/completions",
|
||||
json=payload, headers=self._headers(),
|
||||
)
|
||||
response.raise_for_status()
|
||||
data = response.json()
|
||||
return (data["choices"][0]["message"]["content"]) or ""
|
||||
except httpx.HTTPError as exc:
|
||||
raise self._sanitize(exc) from exc
|
||||
except (KeyError, ValueError) as exc:
|
||||
raise self._sanitize(exc) from exc
|
||||
@@ -1,16 +0,0 @@
|
||||
"""LLM layer types — messages.
|
||||
|
||||
Boundary rule: nothing in llm/ imports from agents/ or api/.
|
||||
|
||||
Providers yield plain str deltas (providers-as-pipes, D-016/D-017);
|
||||
the OpenAI chunk shape lives only at the wire level inside
|
||||
openai_compat.py. ChatDelta/ChoiceDelta were removed in Phase 3 after
|
||||
two verification cycles confirmed no consumers (P2-a finding).
|
||||
"""
|
||||
|
||||
from pydantic import BaseModel
|
||||
|
||||
|
||||
class Message(BaseModel):
|
||||
role: str
|
||||
content: str
|
||||
@@ -1,306 +0,0 @@
|
||||
"""FastAPI app factory — lifespan, CORS, health, routers."""
|
||||
|
||||
import asyncio
|
||||
import contextlib
|
||||
import logging
|
||||
from contextlib import asynccontextmanager
|
||||
|
||||
import httpx
|
||||
from fastapi import FastAPI
|
||||
from fastapi.middleware.cors import CORSMiddleware
|
||||
|
||||
from .agents.registry import AgentRegistry, register_builtin_agents
|
||||
from .agents.session import InMemorySessionStore
|
||||
from .api import (
|
||||
assessment_router,
|
||||
chat_router,
|
||||
defense_router,
|
||||
lab_router,
|
||||
mentor_router,
|
||||
proctor_router,
|
||||
sandboxes_router,
|
||||
telemetry_router,
|
||||
variants_router,
|
||||
)
|
||||
from .config import Settings
|
||||
from .grading.engine import GradingEngine
|
||||
from .grading.store import SQLiteGradeStore
|
||||
from .identity.mock import MockIdentityProvider
|
||||
from .identity.store import SQLiteIdentityStore
|
||||
from .llm import create_provider
|
||||
from .sandbox import SandboxManager, UnshareBackend
|
||||
from .telemetry.ingest import TraceIntegrityMap
|
||||
from .telemetry.store import SQLiteTraceStore
|
||||
from .variants.generator import VariantGenerator
|
||||
from .variants.store import SQLiteVariantStore
|
||||
from .voice.defense_store import SQLiteDefenseStore
|
||||
from .voice.factory import voice_provider_from_settings
|
||||
from .voice.mock import MockVoiceProvider
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
#: Interval between wall-clock/G-2 reaper passes (the manager owns the pass;
|
||||
#: the lifespan owns the loop). 60s against a 900s default timeout → ≤6.7% lag.
|
||||
REAPER_INTERVAL_S = 60.0
|
||||
|
||||
|
||||
def create_app(settings: Settings | None = None) -> FastAPI:
|
||||
settings = settings or Settings()
|
||||
|
||||
@asynccontextmanager
|
||||
async def lifespan(app: FastAPI):
|
||||
# Shared HTTP client pool (D-017): 10s connect / 300s read for cloud TTFT
|
||||
timeout = httpx.Timeout(connect=10.0, read=300.0, write=30.0, pool=10.0)
|
||||
app.state.http_client = httpx.AsyncClient(timeout=timeout)
|
||||
app.state.settings = settings
|
||||
# State-injection override (same pattern as the stores): tests may
|
||||
# pre-set app.state.provider with a scripted mock; only construct the
|
||||
# configured provider when none is present.
|
||||
if getattr(app.state, "provider", None) is None:
|
||||
app.state.provider = create_provider(settings, app.state.http_client)
|
||||
app.state.session_store = InMemorySessionStore()
|
||||
app.state.agent_registry = AgentRegistry()
|
||||
register_builtin_agents(app.state.agent_registry)
|
||||
|
||||
# v0.3 sandbox fabric (REQ-3-001): singleton manager, DI'd via
|
||||
# app.state. Tests may pre-set app.state.sandbox_manager (dependency
|
||||
# override by state injection) to swap the backend; the lifespan then
|
||||
# adopts it instead of constructing the real UnshareBackend one.
|
||||
manager = getattr(app.state, "sandbox_manager", None)
|
||||
if manager is None:
|
||||
manager = SandboxManager(backend=UnshareBackend(), settings=settings)
|
||||
app.state.sandbox_manager = manager
|
||||
await manager.start() # a-1: reap on-disk orphans from a previous process
|
||||
|
||||
# Telemetry persistence (REQ-3-003, D-027): TraceStore wired through
|
||||
# app.state. Tests may pre-set app.state.trace_store (state-injection
|
||||
# override, same pattern as sandbox_manager) — the lifespan adopts it.
|
||||
trace_store = getattr(app.state, "trace_store", None)
|
||||
if trace_store is None:
|
||||
settings.db_path.parent.mkdir(parents=True, exist_ok=True)
|
||||
trace_store = SQLiteTraceStore(db_path=settings.db_path)
|
||||
app.state.trace_store = trace_store
|
||||
# Trace-integrity flags (G-3 INCOMPLETE_FLOODED): process-local map is
|
||||
# intentional (D-019 registry precedent); the lifespan owns it so the
|
||||
# grader and the ingest endpoint share one instance.
|
||||
if getattr(app.state, "trace_integrity", None) is None:
|
||||
app.state.trace_integrity = TraceIntegrityMap()
|
||||
|
||||
# Variant generation (REQ-3-005): VariantStore from the same
|
||||
# SQLite file as traces/grades (D-027), one VariantGenerator singleton
|
||||
# wired through app.state — the generator receives store + provider
|
||||
# via constructor DI and knows nothing of FastAPI (api/ composes it,
|
||||
# same pattern as GradingEngine). Tests may pre-set
|
||||
# app.state.variant_store / app.state.variant_generator (the same
|
||||
# state-injection override); the lifespan adopts a pre-set store but
|
||||
# NEVER rebuilds a pre-set generator (its provider binding is part
|
||||
# of the test fixture).
|
||||
# ORDER NOTE: built BEFORE the grading engine — the engine takes the
|
||||
# variant store (Phase 4 MH#4: variant anchors ship to the grader
|
||||
# prompt; variant_seed stamped on graded records).
|
||||
variant_store = getattr(app.state, "variant_store", None)
|
||||
if variant_store is None:
|
||||
settings.db_path.parent.mkdir(parents=True, exist_ok=True)
|
||||
variant_store = SQLiteVariantStore(db_path=settings.db_path)
|
||||
app.state.variant_store = variant_store
|
||||
if getattr(app.state, "variant_generator", None) is None:
|
||||
app.state.variant_generator = VariantGenerator(
|
||||
variant_store,
|
||||
app.state.provider,
|
||||
model=settings.model,
|
||||
)
|
||||
|
||||
# Oral defense (REQ-3-006): DefenseStore (same SQLite file) + the
|
||||
# mock-first voice provider (D-030) + the seventh Examiner agent.
|
||||
# Tests may pre-set app.state.defense_store / voice_provider /
|
||||
# examiner_agent (state-injection override; never rebuilt if pre-set).
|
||||
defense_store = getattr(app.state, "defense_store", None)
|
||||
if defense_store is None:
|
||||
defense_store = SQLiteDefenseStore(db_path=settings.db_path)
|
||||
app.state.defense_store = defense_store
|
||||
if getattr(app.state, "voice_provider", None) is None:
|
||||
# G-11 (boot survival): a misconfigured real provider must never
|
||||
# crash the unattended deploy — fall back to mock loudly. The
|
||||
# mock provider's descriptor honestly reports mode='mock' so the
|
||||
# UI badge cannot lie about which path is live.
|
||||
try:
|
||||
app.state.voice_provider = voice_provider_from_settings(
|
||||
settings, app.state.http_client
|
||||
)
|
||||
except Exception as exc:
|
||||
logger.warning(
|
||||
"voice provider %r unavailable (%s); falling back to mock "
|
||||
"— fix the AI_VOICE_* settings and restart",
|
||||
settings.voice_provider,
|
||||
exc,
|
||||
)
|
||||
app.state.voice_provider = MockVoiceProvider()
|
||||
if getattr(app.state, "examiner_agent", None) is None:
|
||||
from .agents.examiner import ExaminerAgent
|
||||
|
||||
app.state.examiner_agent = ExaminerAgent(app.state.provider, settings)
|
||||
|
||||
# Identity verification (REQ-5-003): 5th D-027 store (same SQLite
|
||||
# file) + mock-first provider (A-303). State-injection overrides
|
||||
# preserved — tests may pre-set either.
|
||||
identity_store = getattr(app.state, "identity_store", None)
|
||||
if identity_store is None:
|
||||
identity_store = SQLiteIdentityStore(db_path=settings.db_path)
|
||||
app.state.identity_store = identity_store
|
||||
if getattr(app.state, "identity_provider", None) is None:
|
||||
app.state.identity_provider = MockIdentityProvider()
|
||||
|
||||
# Grading persistence + engine (REQ-3-004): GradeStore from the same
|
||||
# SQLite file as traces (D-027), one GradingEngine singleton wired
|
||||
# through app.state — the engine receives its stores via constructor
|
||||
# DI and knows nothing of FastAPI (api/ owns composition). Tests may
|
||||
# pre-set app.state.grade_store / app.state.grading_engine (the same
|
||||
# state-injection override as sandbox_manager/trace_store) to swap
|
||||
# either; the lifespan adopts a pre-set store but NEVER rebuilds a
|
||||
# pre-set engine (its provider binding is part of the test fixture).
|
||||
grade_store = getattr(app.state, "grade_store", None)
|
||||
if grade_store is None:
|
||||
grade_store = SQLiteGradeStore(db_path=settings.db_path)
|
||||
app.state.grade_store = grade_store
|
||||
if getattr(app.state, "grading_engine", None) is None:
|
||||
app.state.grading_engine = GradingEngine(
|
||||
trace_store,
|
||||
grade_store,
|
||||
app.state.trace_integrity,
|
||||
app.state.provider,
|
||||
model=settings.model,
|
||||
variant_store=variant_store, # MH#4: anchors + seed (D-029)
|
||||
)
|
||||
|
||||
async def _reaper_loop() -> None:
|
||||
# Wall-clock timeout + G-2 workdir-size sweep, one pass per tick.
|
||||
while True:
|
||||
await asyncio.sleep(REAPER_INTERVAL_S)
|
||||
try:
|
||||
await manager.reap_expired()
|
||||
except Exception:
|
||||
logger.exception("sandbox reaper pass failed; retrying next tick")
|
||||
|
||||
reaper = asyncio.create_task(_reaper_loop())
|
||||
try:
|
||||
yield
|
||||
finally:
|
||||
reaper.cancel()
|
||||
with contextlib.suppress(asyncio.CancelledError):
|
||||
await reaper
|
||||
# No orphans outlive the process (a-1, shutdown half): destroy
|
||||
# everything live; workdirs stay on disk for snapshot restore.
|
||||
await manager.destroy_all()
|
||||
trace_store.close()
|
||||
grade_store.close()
|
||||
variant_store.close()
|
||||
defense_store.close()
|
||||
identity_store.close()
|
||||
await app.state.http_client.aclose()
|
||||
|
||||
app = FastAPI(title="Nextcraft AI Service", version="0.3.0", lifespan=lifespan)
|
||||
|
||||
# A-008 + D-038: no-credentials CORS. Default '*' admits remote-browser
|
||||
# origins in network mode (safe only because allow_credentials stays
|
||||
# False — never enable credentials with a wildcard). AI_CORS_ORIGINS
|
||||
# restricts to an explicit list. PUT is CONTRACT, not trivia: the learner
|
||||
# build surface writes workspace files with PUT (engine-client writeFile)
|
||||
# — v0.3 initially shipped without it and every cross-origin Save failed
|
||||
# preflight (caught in P7 review; tests/api/test_cors.py pins the policy).
|
||||
app.add_middleware(
|
||||
CORSMiddleware,
|
||||
allow_origins=settings.cors_origin_list,
|
||||
allow_methods=["GET", "POST", "PUT", "DELETE", "OPTIONS"],
|
||||
allow_headers=["Content-Type"],
|
||||
allow_credentials=False,
|
||||
)
|
||||
|
||||
@app.get("/health")
|
||||
async def health() -> dict:
|
||||
return {
|
||||
"status": "ok",
|
||||
"provider": settings.provider,
|
||||
"model": settings.model,
|
||||
}
|
||||
|
||||
app.include_router(chat_router)
|
||||
app.include_router(lab_router)
|
||||
app.include_router(assessment_router)
|
||||
app.include_router(mentor_router)
|
||||
app.include_router(proctor_router)
|
||||
app.include_router(sandboxes_router)
|
||||
app.include_router(telemetry_router)
|
||||
app.include_router(variants_router)
|
||||
app.include_router(defense_router)
|
||||
|
||||
# v0.5 identity (REQ-5-003/004): verification flow + age-gate deps
|
||||
# + the one marketplace 18+ gated stub (G-18).
|
||||
from .api.identity import marketplace_router
|
||||
from .api.identity import router as identity_router
|
||||
|
||||
app.include_router(identity_router)
|
||||
app.include_router(marketplace_router)
|
||||
|
||||
# A-305/D1: PII-safe 422s — FastAPI echoes the offending `input` in
|
||||
# validation errors by default; for the identity submit body that
|
||||
# leaks the raw DOB into responses + client logs. The handler scrubs
|
||||
# PII field inputs (scoped app-wide; harmless elsewhere).
|
||||
from fastapi.exceptions import RequestValidationError
|
||||
from fastapi.responses import JSONResponse
|
||||
|
||||
def _scrub_pii_validation(request, exc): # type: ignore[unused-arg]
|
||||
pii_fields = frozenset({"date_of_birth"})
|
||||
scrubbed = []
|
||||
for err in exc.errors():
|
||||
err = dict(err)
|
||||
if err.get("loc") and err["loc"][-1] in pii_fields:
|
||||
# A-305/D1: never echo the submitted value (input) or the
|
||||
# ctx payload (carries the ValueError); the msg is the
|
||||
# constraint text and is safe.
|
||||
err["input"] = "[redacted]"
|
||||
err.pop("ctx", None)
|
||||
else:
|
||||
# FastAPI's default 422s are JSON-safe EXCEPT ctx payloads
|
||||
# carrying raw ValueError objects (pydantic model_validator
|
||||
# errors); strip ctx body-wide so non-PII routes keep their
|
||||
# 422 shape (msg + loc carry the meaning).
|
||||
ctx = err.get("ctx")
|
||||
if isinstance(ctx, dict):
|
||||
err["ctx"] = {
|
||||
k: v for k, v in ctx.items() if isinstance(v, (str, int, float, bool))
|
||||
}
|
||||
scrubbed.append(err)
|
||||
return JSONResponse(status_code=422, content={"detail": scrubbed})
|
||||
|
||||
app.add_exception_handler(RequestValidationError, _scrub_pii_validation)
|
||||
|
||||
# v0.3.6 single-port deploy: serve the exported web app (apps/web/out)
|
||||
# from the SAME origin as the API when AI_WEB_STATIC_DIR is set. Mounted
|
||||
# AFTER all routers, so /v1/*, /health, /docs win; StaticFiles(html=True)
|
||||
# then resolves / → index.html, /dashboard/ → dashboard/index.html. A
|
||||
# 404 handler below serves the export's 404.html for unknown paths so
|
||||
# browsers see the site's not-found page instead of FastAPI's JSON.
|
||||
# Default unset → no mount, dev/tests see the plain API app.
|
||||
if str(settings.web_static_dir) not in ("", "."):
|
||||
from fastapi.responses import FileResponse
|
||||
from fastapi.staticfiles import StaticFiles
|
||||
|
||||
web_dir = settings.web_static_dir
|
||||
if not web_dir.is_dir():
|
||||
raise RuntimeError(
|
||||
f"AI_WEB_STATIC_DIR is set but {web_dir} does not exist — "
|
||||
"build the web app first (pnpm build) or unset the setting"
|
||||
)
|
||||
|
||||
@app.exception_handler(404)
|
||||
async def _spa_404(request, exc): # type: ignore[unused]
|
||||
not_found_page = web_dir / "404.html"
|
||||
if not_found_page.is_file():
|
||||
return FileResponse(not_found_page, status_code=404)
|
||||
raise exc
|
||||
|
||||
app.mount("/", StaticFiles(directory=web_dir, html=True), name="web")
|
||||
return app
|
||||
|
||||
|
||||
app = create_app()
|
||||
@@ -1,22 +0,0 @@
|
||||
"""Prompt library — prompts are code: versioned in git, reviewed like code (D-018).
|
||||
|
||||
Each module exposes a `versioned SYSTEM_PROMPT` constant and a
|
||||
`render_context(learner_context) -> dict` for str.format_map injection.
|
||||
Final personas land in Phases 3-5; these are the initial drafts.
|
||||
"""
|
||||
|
||||
from .coach import SYSTEM_PROMPT as COACH_PROMPT
|
||||
from .coach import render_context as render_coach
|
||||
from .mentor import SYSTEM_PROMPT as MENTOR_PROMPT
|
||||
from .mentor import render_context as render_mentor
|
||||
from .tutor import SYSTEM_PROMPT as TUTOR_PROMPT
|
||||
from .tutor import render_context as render_tutor
|
||||
|
||||
__all__ = [
|
||||
"COACH_PROMPT",
|
||||
"MENTOR_PROMPT",
|
||||
"TUTOR_PROMPT",
|
||||
"render_coach",
|
||||
"render_mentor",
|
||||
"render_tutor",
|
||||
]
|
||||
@@ -1,29 +0,0 @@
|
||||
"""Assessor agent prompt — rubric coaching over REAL grades (REQ-3-007).
|
||||
|
||||
v0.3 re-grounding: the grading engine (Phase 3) computes the rubric scores
|
||||
from the process trace; Assessor EXPLAINS the stored grade as coaching —
|
||||
it never invents or re-scores. Rigorous, fair, actionable.
|
||||
Version: assessor-v3 (v0.3 live).
|
||||
"""
|
||||
|
||||
SYSTEM_PROMPT = """You are Assessor, the grading agent of Nextcraft, an AI-native competency school.
|
||||
Learner: {learner_name}.
|
||||
|
||||
You receive the learner's STORED process-trace grade (verdict, per-criterion
|
||||
scores, and the build digest) computed by the grading engine. Your job:
|
||||
- Explain what the grade means in plain language (summary).
|
||||
- Strengths: cite what the digest + scores show the learner did well.
|
||||
- Gaps: name the missed opportunities the scores point to.
|
||||
- Next steps: concrete, buildable actions that would move the weakest
|
||||
criterion up one level.
|
||||
|
||||
Rules:
|
||||
- Rigorous but fair. A polished artifact with a weak defense is NOT mastery.
|
||||
- Respond with ONLY a valid JSON object matching the provided schema —
|
||||
no markdown fences, no prose outside the JSON."""
|
||||
|
||||
PROMPT_VERSION = "assessor-v2"
|
||||
|
||||
|
||||
def render_context(learner_context) -> dict:
|
||||
return {"learner_name": learner_context.name}
|
||||
@@ -1,46 +0,0 @@
|
||||
"""Coach agent prompt — pacing, motivation, retrieval practice (REQ-2-005).
|
||||
|
||||
Final persona (Phase 3). Coach is an accountability partner: warm,
|
||||
action-oriented, allergic to fluff. Always ends with exactly one next action
|
||||
and weaves retrieval practice into every reply.
|
||||
Version: coach-v2 (final for v0.2).
|
||||
"""
|
||||
|
||||
SYSTEM_PROMPT = """You are Coach, the pacing and motivation agent of Nextcraft,
|
||||
an AI-native competency school.
|
||||
|
||||
Learner: {learner_name}
|
||||
Active stack: {stacks}
|
||||
Current focus: {progress}
|
||||
|
||||
Your style:
|
||||
- Warm, direct, allergic to fluff. Two short paragraphs maximum.
|
||||
- Pacing: name the learner's next concrete step in their current competency.
|
||||
- Motivation: tie effort to their trajectory — what this unlocks, specifically.
|
||||
- Retrieval practice: before introducing anything new, ask the learner to
|
||||
recall or apply something they already covered (one pointed question).
|
||||
|
||||
Rules:
|
||||
- End with exactly ONE clear next action phrased as a command ("Post your
|
||||
plan for the orchestrator retry loop before starting").
|
||||
- Never lecture; never list more than two options.
|
||||
- If the learner is stuck or frustrated, slow down and shrink the step."""
|
||||
|
||||
PROMPT_VERSION = "coach-v2"
|
||||
|
||||
|
||||
def render_context(learner_context) -> dict:
|
||||
stacks = ", ".join(f"{s.title} ({s.percent}%)" for s in learner_context.active_stacks)
|
||||
in_progress = [
|
||||
c for c in learner_context.active_competencies if c.status == "in_progress"
|
||||
]
|
||||
progress = (
|
||||
f"{in_progress[0].title} ({in_progress[0].competency_id})"
|
||||
if in_progress
|
||||
else "no competency currently in progress"
|
||||
)
|
||||
return {
|
||||
"learner_name": learner_context.name,
|
||||
"stacks": stacks or "none yet",
|
||||
"progress": progress,
|
||||
}
|
||||
@@ -1,60 +0,0 @@
|
||||
"""Examiner agent prompt — oral defense questioning + final verdict (REQ-3-006).
|
||||
|
||||
The examiner is the seventh agent (Phase 5). It conducts a Socratic oral
|
||||
defense of the learner's submitted work: probes understanding, challenges
|
||||
process choices grounded in the trace digest ("why did you take that
|
||||
approach at that point?"), one question per turn, adapting to answers.
|
||||
It never reveals rubric internals; tone is rigorous but supportive.
|
||||
|
||||
Digest discipline (D-028 mirror): the examiner's variable inputs are the
|
||||
compact TraceDigest JSON, the variant task statement, and the defense
|
||||
transcript — never the raw trace, never learner-identifying material.
|
||||
|
||||
Version: examiner-v1.
|
||||
"""
|
||||
|
||||
SYSTEM_PROMPT = """You are Examiner, the oral-defense agent of Nextcraft,
|
||||
an AI-native competency school.
|
||||
|
||||
You receive: (a) a compact build-process digest (deterministic counters of the
|
||||
learner's build session), (b) the learner's task statement, and (c) the defense
|
||||
transcript so far. Your job:
|
||||
- Ask ONE question per turn: probe understanding and challenge process
|
||||
choices, grounded in the digest facts ("you hit N failed runs before
|
||||
passing — walk me through what changed") or the task statement.
|
||||
- Adapt: follow up on the learner's answers; drill into vague responses.
|
||||
- Never reveal rubric details or scoring internals.
|
||||
- Tone: rigorous, precise, supportive. A defense is a conversation, not an
|
||||
interrogation.
|
||||
|
||||
When asked for a FINAL VERDICT (the structured mode), judge:
|
||||
- understanding: can the learner explain their own work?
|
||||
- process_justification: are the build-session choices defensible from the
|
||||
digest facts and the answers?
|
||||
- communication: are answers clear, specific, and on-topic?
|
||||
Score honestly; a weak defense of strong work is NOT mastery.
|
||||
|
||||
Rules:
|
||||
- Respond with ONLY what the turn requires: a single question (question mode)
|
||||
or a valid JSON object matching the provided schema (verdict mode).
|
||||
- If the digest shows error_fix_cycles > 0, at least one question should ask
|
||||
about the debugging path.
|
||||
- If the learner's answer is off-topic, redirect once, then move on.
|
||||
"""
|
||||
|
||||
VERDICT_SCHEMA_HINT = (
|
||||
'{"verdict": "mastered" | "developing" | "not_yet", '
|
||||
'"understanding": "<one sentence>", '
|
||||
'"process_justification": "<one sentence>", '
|
||||
'"communication": "<one sentence>", '
|
||||
'"strengths": ["<one sentence>"], '
|
||||
'"gaps": ["<one sentence>"]}'
|
||||
)
|
||||
|
||||
|
||||
def render_digest_context(digest_json: str, statement: str | None) -> str:
|
||||
"""The examiner's per-session grounding: digest JSON + task statement."""
|
||||
parts = [f"Build-process digest:\n{digest_json}"]
|
||||
if statement:
|
||||
parts.append(f"Learner's task statement:\n{statement}")
|
||||
return "\n\n".join(parts)
|
||||
@@ -1,140 +0,0 @@
|
||||
"""Grading rubric prompt — criteria, level anchors, digest render (REQ-3-004).
|
||||
|
||||
The grading prompt is deliberately learner-anonymous and trace-bare: the
|
||||
model receives ONLY the fixed rubric text and the compact numeric digest
|
||||
(TraceDigest JSON, D-028) — never a raw command, file path, payload
|
||||
string, learner id, or task id. Everything variable the LLM sees is
|
||||
deterministic counters, which both bounds the prompt-injection surface
|
||||
and makes "no raw trace reaches the prompt" assert-able in tests (plant
|
||||
a distinctive marker in a command payload; assert it absent from every
|
||||
message the provider received).
|
||||
|
||||
Rubric (four criteria, each scored 0-4 — the ids are the validated
|
||||
RubricScore keys enforced by grading/engine.py):
|
||||
process_quality — iterative building in small, verified steps.
|
||||
correctness — where the session ended (test/run outcomes).
|
||||
debugging_discipline — how failures were handled.
|
||||
test_usage — when and how often tests were run.
|
||||
|
||||
Advisory a-4 (embedded in the process_quality anchors): high edit/command
|
||||
churn with NO test progress is a process-quality NEGATIVE — churn is not
|
||||
work. A session with many edits/commands whose test state never moves is
|
||||
thrashing, not iterating, and must score low on process quality.
|
||||
|
||||
House-style deviation, documented: unlike the tutor prompts, this module
|
||||
has no SYSTEM_PROMPT placeholders and no render_context(learner_context)
|
||||
— grading is context-free by design (learner anonymity; the digest is the
|
||||
only variable input). Runtime imports are TYPE_CHECKING-only so this
|
||||
module stays pure text and can never import-cycle with grading/engine.py
|
||||
(engine imports this module; if this module imported grading.* at runtime
|
||||
while grading/__init__ pulls engine, the package init would deadlock on a
|
||||
partially-initialized module).
|
||||
|
||||
Version: grader-v1.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import TYPE_CHECKING, Final
|
||||
|
||||
if TYPE_CHECKING: # pragma: no cover - typing only; keeps this module pure text
|
||||
from ..grading.features import TraceDigest
|
||||
|
||||
PROMPT_VERSION = "grader-v1"
|
||||
|
||||
#: Canonical criterion ids. The engine validates RubricScore criteria keys
|
||||
#: against this tuple; the schema hint and anchors below speak the same ids.
|
||||
RUBRIC_CRITERIA: Final[tuple[str, ...]] = (
|
||||
"process_quality",
|
||||
"correctness",
|
||||
"debugging_discipline",
|
||||
"test_usage",
|
||||
)
|
||||
|
||||
#: Sentinel line the engine's user turn is rendered around. Tests (and the
|
||||
#: calibration mock) split on it to locate the digest JSON in the prompt.
|
||||
DIGEST_MARKER: Final = "PROCESS TRACE DIGEST (JSON):"
|
||||
|
||||
SYSTEM_PROMPT = """You are the Grader of Nextcraft, an AI-native competency school.
|
||||
You score a learner's build session from a compact numeric digest of their
|
||||
process trace. You NEVER see the raw trace — commands, file contents, and
|
||||
payloads do not exist on your side; every number you need is in the digest.
|
||||
|
||||
Rubric — score each criterion 0-4:
|
||||
|
||||
process_quality — iterative building in small, verified steps.
|
||||
4: tight edit→test loops throughout; small verified increments; healthy pacing.
|
||||
3: steady small edits with regular runs; progress mostly verified.
|
||||
2: some iteration, but large unverified leaps or long idle stretches.
|
||||
1: a single bulk change (e.g. one large paste) then a single run; no iteration.
|
||||
0: no meaningful work visible.
|
||||
ADVISORY: high edit/command churn with NO test progress (no runs, no
|
||||
movement in pass counts) is a process-quality NEGATIVE — churn is not
|
||||
work. Cap such a session at 1 on this criterion no matter how many
|
||||
edits or commands were counted.
|
||||
|
||||
correctness — where the session ended up.
|
||||
4: final test status pass, with tests passing early and consistently.
|
||||
3: final pass, reached through fail→fix→pass cycles that closed.
|
||||
2: final pass, but preceded by a long unresolved failure streak.
|
||||
1: final fail, but partial passes observed along the way.
|
||||
0: final fail, or no test/run evidence at all.
|
||||
|
||||
debugging_discipline — how failures were handled.
|
||||
4: every failure cycle closes; targeted fixes with low mean fix latency.
|
||||
3: most fail→edit→re-run cycles close with a pass.
|
||||
2: failures followed by edits, but cycles rarely close.
|
||||
1: repeated failures with no targeted edits between runs (flailing).
|
||||
0: failures with no fix attempts at all.
|
||||
|
||||
test_usage — when and how often tests were run.
|
||||
4: tests run early (small first-pass offset) and throughout the session.
|
||||
3: regular test runs interleaved with edits.
|
||||
2: sparse tests; long stretches of unverified edits.
|
||||
1: a single late test run only.
|
||||
0: no test or run evidence.
|
||||
|
||||
Rules:
|
||||
- Judge STRICTLY from the digest numbers; cite the fields you used.
|
||||
- Strengths: the two strongest digest observations, one sentence each.
|
||||
- Gaps: the two most important missed opportunities, one sentence each
|
||||
(a clean session names its next-level improvement instead).
|
||||
- Be rigorous but fair: a session that ends green was not necessarily
|
||||
well built, and a struggling session that never passed may still show
|
||||
real debugging discipline.
|
||||
- Respond with ONLY a valid JSON object matching the provided schema —
|
||||
no markdown fences, no prose outside the JSON."""
|
||||
|
||||
RUBRIC_SCORE_SCHEMA_HINT = (
|
||||
'{"criteria": {"process_quality": <0-4 int>, "correctness": <0-4 int>, '
|
||||
'"debugging_discipline": <0-4 int>, "test_usage": <0-4 int>}, '
|
||||
'"strengths": ["<one sentence>"], "gaps": ["<one sentence>"], '
|
||||
'"verdict": "mastered" | "developing" | "not_yet"}'
|
||||
)
|
||||
|
||||
|
||||
def render_trace_digest(
|
||||
digest: TraceDigest,
|
||||
anchors_context: str | None = None,
|
||||
) -> str:
|
||||
"""Render the grader's user turn: a marker line + the digest JSON — nothing else.
|
||||
|
||||
This is the ONLY per-session content that ever reaches the LLM (D-028):
|
||||
the engine composes [system: SYSTEM_PROMPT, user: render_trace_digest(digest)]
|
||||
and the D-020 defense appends its generic schema instruction to this
|
||||
user turn at request time. No learner id, task id, or raw trace material
|
||||
is injected — assert-able by tests.
|
||||
|
||||
`anchors_context` (Phase 4, MH#4): when the graded task derives from a
|
||||
variant, the engine passes the template's difficulty-normalization
|
||||
anchors (the expected effort envelope) so the rubric is applied against
|
||||
the SAME bar for every variant of that template (a-5). It contains only
|
||||
the anchor numbers + the template id — no learner-identifying material.
|
||||
"""
|
||||
base = (
|
||||
"Score this build session against the rubric.\n"
|
||||
f"{DIGEST_MARKER}\n{digest.model_dump_json()}"
|
||||
)
|
||||
if anchors_context:
|
||||
base = f"{base}\n\nExpected effort envelope for this task variant:\n{anchors_context}"
|
||||
return base
|
||||
@@ -1,48 +0,0 @@
|
||||
"""Lab agent prompt — in-flow feedback over LIVE telemetry (REQ-3-007).
|
||||
|
||||
v0.3 re-grounding: the timeline is the learner's real TraceDigest (D-028
|
||||
compact counters — commands, test outcomes, idle gaps, edit cadence), not
|
||||
v0.2 corpus scenarios. Lab is a pragmatic build partner: reads the live
|
||||
digest, names the one most useful adjustment, gives one concrete next
|
||||
step. No session chat.
|
||||
Version: lab-v3 (v0.3 live).
|
||||
"""
|
||||
|
||||
SYSTEM_PROMPT = """You are Lab, the in-flow feedback agent watching a learner
|
||||
build in the Nextcraft sandbox.
|
||||
Learner: {learner_name}. Active stack: {stacks}.
|
||||
|
||||
You receive a telemetry timeline of the learner's build session below.
|
||||
Your job, in order:
|
||||
1. Say what the telemetry shows — name the specific events that matter.
|
||||
2. Name the single most useful adjustment (one thing, not a list).
|
||||
3. Give one concrete next step phrased as a command.
|
||||
|
||||
Rules:
|
||||
- Be specific to the events you see. If tests failed twice with the same
|
||||
error, say so. If there is a long idle gap, name it.
|
||||
- If the session looks healthy, say so briefly and set the next challenge.
|
||||
- If something looks off (e.g., a huge paste followed by instant success),
|
||||
treat it as a coaching moment, not an accusation — suggest a quick
|
||||
self-check that would prove understanding.
|
||||
- Three short paragraphs maximum. No headers, no bullet lists."""
|
||||
|
||||
PROMPT_VERSION = "lab-v3"
|
||||
|
||||
|
||||
def render_digest_timeline(digest) -> str:
|
||||
"""Live-trace timeline: the compact TraceDigest JSON (D-028)."""
|
||||
if digest is None:
|
||||
return (
|
||||
"No telemetry yet for this build session. Ask the learner to run "
|
||||
"the task's starter test to establish a baseline."
|
||||
)
|
||||
return f"Live build-session digest:\n{digest.model_dump_json()}"
|
||||
|
||||
|
||||
def render_context(learner_context) -> dict:
|
||||
stacks = ", ".join(f"{s.title} ({s.percent}%)" for s in learner_context.active_stacks)
|
||||
return {
|
||||
"learner_name": learner_context.name,
|
||||
"stacks": stacks or "none yet",
|
||||
}
|
||||
@@ -1,49 +0,0 @@
|
||||
"""Mentor agent prompt — long-horizon career narrative (REQ-2-010).
|
||||
|
||||
Final persona (Phase 5). Mentor is a wise career guide: connects today's
|
||||
competencies and artifacts to a long-horizon AI-era trajectory.
|
||||
Version: mentor-v2 (final for v0.2).
|
||||
"""
|
||||
|
||||
SYSTEM_PROMPT = """You are Mentor, the long-horizon career agent of Nextcraft,
|
||||
an AI-native competency school.
|
||||
|
||||
Learner: {learner_name}
|
||||
Active stacks: {stacks}
|
||||
Current focus: {progress}
|
||||
Microcredentials earned: {microcredentials}
|
||||
Recent artifacts: {artifacts}
|
||||
|
||||
Your job: narrate the learner's trajectory in two to three paragraphs:
|
||||
1. Where they are now — what their competency progress and artifacts say
|
||||
about them as a builder (specific, evidence-based).
|
||||
2. What their current stack unlocks next — name the next competency or
|
||||
microcredential worth chasing and the role it points toward.
|
||||
3. How they position in the AI-era labor market — which employer problems
|
||||
their profile already answers.
|
||||
|
||||
Rules:
|
||||
- Forward-looking and concrete. No fortune-telling, no flattery.
|
||||
- Reference their artifacts by name at least once.
|
||||
- Write like a mentor writing to one person, not a career-services brochure."""
|
||||
|
||||
PROMPT_VERSION = "mentor-v2"
|
||||
|
||||
|
||||
def render_context(learner_context) -> dict:
|
||||
stacks = ", ".join(f"{s.title} ({s.percent}%)" for s in learner_context.active_stacks)
|
||||
in_progress = [
|
||||
c for c in learner_context.active_competencies if c.status == "in_progress"
|
||||
]
|
||||
progress = (
|
||||
f"{in_progress[0].title} ({in_progress[0].competency_id})"
|
||||
if in_progress
|
||||
else "no competency currently in progress"
|
||||
)
|
||||
return {
|
||||
"learner_name": learner_context.name,
|
||||
"stacks": stacks or "none yet",
|
||||
"progress": progress,
|
||||
"microcredentials": str(learner_context.microcredential_count),
|
||||
"artifacts": ", ".join(learner_context.recent_artifacts) or "none yet",
|
||||
}
|
||||
@@ -1,38 +0,0 @@
|
||||
"""Proctor agent prompt — integrity signals with coaching interventions (REQ-3-007).
|
||||
|
||||
v0.3 re-grounding: inputs are REAL — the live trace digest (idle gaps,
|
||||
command cadence, edit bursts), the oral-defense integrity signals (long
|
||||
pauses), and the variant audit context (seed + params). Proctor is a
|
||||
supportive observer, never punitive: classifies signals, recommends ONE
|
||||
coaching intervention. Assume good faith.
|
||||
Version: proctor-v3 (v0.3 live).
|
||||
"""
|
||||
|
||||
SYSTEM_PROMPT = """You are Proctor, the integrity-support agent of Nextcraft,
|
||||
an AI-native competency school.
|
||||
Learner: {learner_name}.
|
||||
|
||||
You receive the learner's REAL build-session digest (idle gaps, command
|
||||
categories, edit/test cadence), oral-defense integrity signals (long
|
||||
pauses), and — when the task is variant-derived — the variant seed context.
|
||||
Your job:
|
||||
- Classify EACH notable signal: type ("idle_gap" | "long_pause" |
|
||||
"burst_edit" | "off_template"), severity ("low" | "medium" | "high"),
|
||||
and a one-sentence note citing the numbers.
|
||||
- Recommend exactly ONE supportive coaching intervention for the session
|
||||
overall — never punitive, never accusatory. Frame around helping the
|
||||
learner succeed.
|
||||
|
||||
Rules:
|
||||
- Assume good faith. Tab switches to documentation are normal engineering.
|
||||
- Idle gaps are often thinking. Only unusual patterns deserve higher severity.
|
||||
- A large paste during an assessment deserves "high" severity but the
|
||||
intervention stays coaching-shaped: verification, not punishment.
|
||||
- Respond with ONLY a valid JSON object matching the provided schema —
|
||||
no markdown fences, no prose outside the JSON."""
|
||||
|
||||
PROMPT_VERSION = "proctor-v3"
|
||||
|
||||
|
||||
def render_context(learner_context) -> dict:
|
||||
return {"learner_name": learner_context.name}
|
||||
@@ -1,46 +0,0 @@
|
||||
"""Tutor agent prompt — concept delivery, Socratic questioning (REQ-2-006).
|
||||
|
||||
Final persona (Phase 3). Tutor is a patient expert teacher: one concept at
|
||||
a time, worked example first, Socratic check before moving on.
|
||||
Version: tutor-v2 (final for v0.2).
|
||||
"""
|
||||
|
||||
SYSTEM_PROMPT = """You are Tutor, the concept-delivery agent of Nextcraft,
|
||||
an AI-native competency school.
|
||||
|
||||
Learner: {learner_name}
|
||||
Active stack: {stacks}
|
||||
Current focus: {progress}
|
||||
|
||||
Your style:
|
||||
- Teach exactly ONE concept per reply. Never more.
|
||||
- Structure: (1) name the concept in one sentence, (2) give a short worked
|
||||
example (5-8 lines) the learner can trace, (3) ask ONE Socratic question
|
||||
that checks whether they can apply it to a slightly different case.
|
||||
|
||||
Rules:
|
||||
- Never dump walls of text. If the concept needs more than ~150 words, teach
|
||||
only its first slice and promise the rest after the learner answers.
|
||||
- If the learner's last message reveals a misconception, correct it gently
|
||||
before teaching.
|
||||
- If the learner answers your question, evaluate the answer explicitly
|
||||
(right / partly right / not yet) before the next concept."""
|
||||
|
||||
PROMPT_VERSION = "tutor-v2"
|
||||
|
||||
|
||||
def render_context(learner_context) -> dict:
|
||||
stacks = ", ".join(f"{s.title} ({s.percent}%)" for s in learner_context.active_stacks)
|
||||
in_progress = [
|
||||
c for c in learner_context.active_competencies if c.status == "in_progress"
|
||||
]
|
||||
progress = (
|
||||
f"{in_progress[0].title} ({in_progress[0].competency_id})"
|
||||
if in_progress
|
||||
else "no competency currently in progress"
|
||||
)
|
||||
return {
|
||||
"learner_name": learner_context.name,
|
||||
"stacks": stacks or "none yet",
|
||||
"progress": progress,
|
||||
}
|
||||
@@ -1,46 +0,0 @@
|
||||
"""Variant instantiation prompt (D-029, REQ-3-005).
|
||||
|
||||
The model's ONLY job is to render already-sampled slot values into a task
|
||||
statement — it never invents parameters (the seeded sampler is pure code)
|
||||
and never changes difficulty. Prompt-injection surface is bounded: the
|
||||
variable inputs are the skeleton text, the seeded slot values, and the
|
||||
template title — nothing from the learner's environment.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import TYPE_CHECKING
|
||||
|
||||
from ..llm.types import Message
|
||||
|
||||
if TYPE_CHECKING: # pragma: no cover - keeps this module pure text
|
||||
from ..variants.templates import TaskTemplate
|
||||
|
||||
VARIANT_SYSTEM_PROMPT = (
|
||||
"You instantiate per-learner task variants for a competency-based AI school. "
|
||||
"You receive a task statement skeleton and ALREADY-SAMPLED slot values. "
|
||||
"Render the slot values into the skeleton, producing a complete, unambiguous "
|
||||
"task statement a learner can build against. Rules:\n"
|
||||
"- Use EXACTLY the given slot values; do not invent, rename, or add parameters.\n"
|
||||
"- Keep the engineering depth IDENTICAL across draws: slot values change the "
|
||||
"scenario, never the difficulty or scope.\n"
|
||||
"- Keep the statement in the same language and register as the skeleton.\n"
|
||||
"- Output STRICT JSON only: {\"statement\": \"<rendered statement>\"}.\n"
|
||||
)
|
||||
|
||||
VARIANT_SCHEMA_HINT = '{"statement": "<complete rendered task statement string>"}'
|
||||
|
||||
|
||||
def render_variant_prompt(template: TaskTemplate, params: dict[str, str | int]) -> list[Message]:
|
||||
"""Messages for one seeded instantiation (D-020 defense drives the call)."""
|
||||
slot_lines = "\n".join(f" {{{slot.name}}} = {params[slot.name]!r}" for slot in template.slots)
|
||||
user = (
|
||||
f"Template: {template.title} (id={template.id})\n"
|
||||
f"Statement skeleton:\n{template.statement_skeleton}\n\n"
|
||||
f"Seeded slot values (use EXACTLY these):\n{slot_lines}\n\n"
|
||||
"Render the complete task statement now."
|
||||
)
|
||||
return [
|
||||
Message(role="system", content=VARIANT_SYSTEM_PROMPT),
|
||||
Message(role="user", content=user),
|
||||
]
|
||||
@@ -1,44 +0,0 @@
|
||||
"""Sandbox fabric (v0.3, REQ-3-001) — learner code-execution isolation via Linux namespaces.
|
||||
|
||||
Public surface:
|
||||
SandboxSpec / SandboxHandle / ResourceLimits / ExecResult — pydantic contracts.
|
||||
SandboxBackend — the protocol every backend implements (D-024 port).
|
||||
UnshareBackend — util-linux `unshare` backend (D-024 backend).
|
||||
SandboxManager — lifecycle + pool guard (D-032) + reapers (G-2, a-1).
|
||||
SandboxHandleInfo — manager return row: handle fields + learner_id.
|
||||
PoolFullError / SandboxNotFoundError / SandboxIntegrityEvent — manager surface.
|
||||
SandboxUnavailableError — raised when namespaces are not usable on this host.
|
||||
SandboxDir / workspace_path / create_layout / snapshot — per-sandbox workdir layout.
|
||||
|
||||
Boundary rule: `sandbox/` never imports `api/` or `agents/`; it owns subprocess spawning only.
|
||||
"""
|
||||
|
||||
from .backend import ExecResult, ResourceLimits, SandboxBackend, SandboxHandle, SandboxSpec
|
||||
from .manager import (
|
||||
PoolFullError,
|
||||
SandboxHandleInfo,
|
||||
SandboxIntegrityEvent,
|
||||
SandboxManager,
|
||||
SandboxNotFoundError,
|
||||
)
|
||||
from .unshare_backend import SandboxUnavailableError, UnshareBackend
|
||||
from .workdir import SandboxDir, create_layout, snapshot, workspace_path
|
||||
|
||||
__all__ = [
|
||||
"ExecResult",
|
||||
"PoolFullError",
|
||||
"ResourceLimits",
|
||||
"SandboxBackend",
|
||||
"SandboxDir",
|
||||
"SandboxHandle",
|
||||
"SandboxHandleInfo",
|
||||
"SandboxIntegrityEvent",
|
||||
"SandboxManager",
|
||||
"SandboxNotFoundError",
|
||||
"SandboxSpec",
|
||||
"SandboxUnavailableError",
|
||||
"UnshareBackend",
|
||||
"create_layout",
|
||||
"snapshot",
|
||||
"workspace_path",
|
||||
]
|
||||
@@ -1,95 +0,0 @@
|
||||
"""SandboxBackend protocol + pydantic contracts (REQ-3-001, D-024).
|
||||
|
||||
`SandboxBackend` is the port the sandbox fabric depends on. The only
|
||||
implementation in v0.3 is `unshare_backend.UnshareBackend`; a future
|
||||
firecracker/bwrap backend must satisfy this same surface.
|
||||
|
||||
Contracts are plain pydantic models so API/agent layers can construct and
|
||||
validate them at the request boundary without importing the backend itself.
|
||||
"""
|
||||
|
||||
from datetime import datetime
|
||||
from pathlib import Path
|
||||
from typing import Protocol, runtime_checkable
|
||||
|
||||
from pydantic import BaseModel, ConfigDict, Field
|
||||
|
||||
|
||||
class ResourceLimits(BaseModel):
|
||||
"""Per-sandbox rlimits, applied via `preexec_fn` immediately before exec.
|
||||
|
||||
- memory_bytes → RLIMIT_AS (address space; hard OOM ceiling)
|
||||
- cpu_seconds → RLIMIT_CPU (CPU-seconds; SIGKILL on hard expiry)
|
||||
- file_size_bytes → RLIMIT_FSIZE (~50 MB single-file cap)
|
||||
|
||||
RLIMIT_NPROC is NOT set: the counter is shared per host UID across all
|
||||
namespaces, so it cannot isolate one sandbox from another on this host.
|
||||
Total disk usage is enforced by the manager sweep (G-2), not here.
|
||||
"""
|
||||
|
||||
model_config = ConfigDict(frozen=True)
|
||||
|
||||
memory_bytes: int = Field(default=256 * 1024 * 1024, gt=0)
|
||||
cpu_seconds: int = Field(default=30, gt=0)
|
||||
file_size_bytes: int = Field(default=50 * 1024 * 1024, gt=0)
|
||||
|
||||
|
||||
class SandboxHandle(BaseModel):
|
||||
"""A live (or reaped) sandbox. `pid` is None once `destroy()` completes."""
|
||||
|
||||
model_config = ConfigDict(arbitrary_types_allowed=True)
|
||||
|
||||
id: str
|
||||
pid: int | None
|
||||
workdir: Path
|
||||
created_at: datetime
|
||||
|
||||
|
||||
class SandboxSpec(BaseModel):
|
||||
"""Immutable description of the sandbox to lay out on disk.
|
||||
|
||||
`capture_env` (REQ-3-003): when non-empty, the backend starts a persistent
|
||||
telemetry-wired sandbox — helper + inner namespaces + the stdlib capture
|
||||
agent, launched with these env vars (NC_LEARNER_ID, NC_TASK_ID,
|
||||
NC_INGEST_URL, NC_SANDBOX_ID). When None (default), spawn keeps the pure
|
||||
shell semantics (REQ-3-001): disk layout only, fresh namespaces per exec.
|
||||
"""
|
||||
|
||||
model_config = ConfigDict(arbitrary_types_allowed=True)
|
||||
|
||||
sandbox_id: str
|
||||
learner_id: str
|
||||
workdir: Path
|
||||
limits: ResourceLimits = ResourceLimits()
|
||||
capture_env: dict[str, str] | None = None
|
||||
|
||||
|
||||
class ExecResult(BaseModel):
|
||||
"""One namespaced execution: cwd = the bind-mounted workspace (`/work`)."""
|
||||
|
||||
cmd: list[str]
|
||||
returncode: int
|
||||
stdout: str
|
||||
stderr: str
|
||||
duration_s: float
|
||||
|
||||
|
||||
@runtime_checkable
|
||||
class SandboxBackend(Protocol):
|
||||
"""The sandbox port. Backends spawn subprocesses; they never touch HTTP."""
|
||||
|
||||
async def spawn(self, spec: SandboxSpec) -> SandboxHandle:
|
||||
"""Create the sandbox from `spec` and return its handle."""
|
||||
...
|
||||
|
||||
async def exec(self, handle: SandboxHandle, cmd: list[str]) -> ExecResult:
|
||||
"""Run `cmd` inside the sandbox workspace; capture stdout/stderr."""
|
||||
...
|
||||
|
||||
async def snapshot(self, handle: SandboxHandle) -> Path:
|
||||
"""Copy the workspace into `<workdir>/snapshots/<utc-ts>/`; return the path."""
|
||||
...
|
||||
|
||||
async def destroy(self, handle: SandboxHandle) -> None:
|
||||
"""Tear the sandbox down. Must be idempotent."""
|
||||
...
|
||||
@@ -1,461 +0,0 @@
|
||||
"""SandboxManager — lifecycle, concurrency guard, reapers (REQ-3-001, REQ-3-002).
|
||||
|
||||
Responsibilities (D-032, G-2, a-1):
|
||||
|
||||
- create/list/get/snapshot/destroy over a `SandboxBackend` port. create/list/
|
||||
get return `SandboxHandleInfo` rows — the handle fields plus the owning
|
||||
`learner_id` — so the API layer never re-asks "who owns this id?".
|
||||
- Capacity guard (D-032): `create` raises `PoolFullError` when the active
|
||||
count reaches `settings.sandbox_max_concurrent`. No queue — the API layer
|
||||
maps this to 503.
|
||||
- Wall-clock reaper: `reap_expired()` destroys sandboxes older than
|
||||
`settings.sandbox_timeout_s`. Run it on an async timer owned by the caller
|
||||
(app lifespan wires the loop; the manager owns only the pass).
|
||||
- Workdir-size sweep (G-2): the same timer pass also measures each sandbox's
|
||||
`workspace/` tree; anything over `settings.sandbox_max_workdir_mb` is
|
||||
snapshotted (evidence preserved), destroyed, and recorded as an integrity
|
||||
signal. SOFT CAP, best-effort, NOT kernel-enforced — without cgroup
|
||||
delegation or sudo there is no hard per-sandbox disk quota on this host.
|
||||
RLIMIT_FSIZE bounds a single file; this sweep bounds aggregate growth
|
||||
between passes.
|
||||
- Startup reaper (a-1): `start()` scans `settings.sandbox_dir` for workdirs
|
||||
whose recorded pid is dead (marker file `sandbox.json` beside workspace/)
|
||||
and reaps them, logging a warning. Handles are IN-MEMORY and process-local
|
||||
(D-019 precedent): on process restart every handle is orphaned, so boot
|
||||
must recover disk state.
|
||||
|
||||
Registry: plain dict guarded by an `asyncio.Lock`, process-local, explicitly
|
||||
NOT a store. Swapping in persistence (D-027) must not change this interface.
|
||||
|
||||
Resource limits: enforced at exec time by the spawner's in-namespace shim
|
||||
(RLIMIT_AS / RLIMIT_CPU / RLIMIT_FSIZE — see UnshareBackend), never here;
|
||||
the manager's enforcement surface is lifecycle (capacity, wall-clock, disk
|
||||
sweep).
|
||||
|
||||
BOUNDARY: this module NEVER imports `api/` or `agents/`.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import json
|
||||
import logging
|
||||
import os
|
||||
import shutil
|
||||
import uuid
|
||||
from collections.abc import Callable
|
||||
from datetime import UTC, datetime
|
||||
from pathlib import Path
|
||||
|
||||
from pydantic import BaseModel, ConfigDict, Field
|
||||
|
||||
from ..config import Settings
|
||||
from . import workdir as workdir_mod
|
||||
from .backend import SandboxBackend, SandboxHandle
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
#: Marker file beside workspace/ recording the owning process + metadata.
|
||||
#: It is what the startup reaper uses after this process's in-memory
|
||||
#: registry is lost (crash/restart → orphan detection, a-1).
|
||||
PID_MARKER = "sandbox.json"
|
||||
|
||||
|
||||
class PoolFullError(RuntimeError):
|
||||
"""D-032: active sandbox count reached `settings.sandbox_max_concurrent`.
|
||||
|
||||
The API layer maps this to 503. There is deliberately NO queue.
|
||||
"""
|
||||
|
||||
|
||||
class SandboxNotFoundError(KeyError):
|
||||
"""No live sandbox with that id in this process's registry."""
|
||||
|
||||
|
||||
class SandboxIntegrityEvent(BaseModel):
|
||||
"""One manager-observed integrity signal (G-2).
|
||||
|
||||
Recorded in-process (`SandboxManager.integrity_events`, for the proctor
|
||||
pipeline to drain) AND logged at WARNING (durable trail) — the same
|
||||
dual-sink pattern a DB-backed store will keep behind D-027.
|
||||
"""
|
||||
|
||||
model_config = ConfigDict(frozen=True)
|
||||
|
||||
kind: str = Field(description="e.g. 'workdir_size_cap' (G-2), 'orphan_reaped' (a-1)")
|
||||
sandbox_id: str
|
||||
learner_id: str
|
||||
detail: str
|
||||
observed_at: datetime = Field(default_factory=lambda: datetime.now(UTC))
|
||||
|
||||
|
||||
class SandboxHandleInfo(BaseModel):
|
||||
"""A `SandboxHandle` plus its owning `learner_id` (manager return row).
|
||||
|
||||
Handles alone don't carry the learner — the registry side-table does —
|
||||
and every API read/list needs it, so the manager joins the two ONCE here
|
||||
instead of exposing `_learner_ids` internals.
|
||||
"""
|
||||
|
||||
model_config = ConfigDict(arbitrary_types_allowed=True)
|
||||
|
||||
id: str
|
||||
learner_id: str
|
||||
workdir: Path
|
||||
created_at: datetime
|
||||
pid: int | None = None
|
||||
|
||||
|
||||
class SandboxManager:
|
||||
"""Lifecycle owner for learner sandboxes.
|
||||
|
||||
Dependencies are injected (D-017 style): the backend port, settings, and
|
||||
a wall clock. Single-process only; the registry is in-memory.
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
*,
|
||||
backend: SandboxBackend,
|
||||
settings: Settings,
|
||||
clock: Callable[[], datetime] | None = None,
|
||||
) -> None:
|
||||
self._backend = backend
|
||||
self._settings = settings
|
||||
self._clock = clock or (lambda: datetime.now(UTC))
|
||||
self._handles: dict[str, SandboxHandle] = {}
|
||||
self._learner_ids: dict[str, str] = {} # sandbox_id -> learner_id
|
||||
#: REQ-5-005 (G-15): sandbox_id -> task_id side-table — the exec
|
||||
#: policy resolves the variant's environment kind by task_id.
|
||||
self._task_ids: dict[str, str | None] = {}
|
||||
self._lock = asyncio.Lock()
|
||||
self._integrity_events: list[SandboxIntegrityEvent] = []
|
||||
self._started = False
|
||||
|
||||
# -- introspection ------------------------------------------------------
|
||||
|
||||
@property
|
||||
def active_count(self) -> int:
|
||||
return len(self._handles)
|
||||
|
||||
@property
|
||||
def integrity_events(self) -> list[SandboxIntegrityEvent]:
|
||||
"""Drainable view of recorded integrity signals (G-2, a-1)."""
|
||||
return list(self._integrity_events)
|
||||
|
||||
# -- lifecycle ----------------------------------------------------------
|
||||
|
||||
async def create(
|
||||
self, learner_id: str, task_id: str | None = None
|
||||
) -> SandboxHandleInfo:
|
||||
"""Spawn a sandbox for `learner_id`, or raise `PoolFullError` (D-032).
|
||||
|
||||
`task_id` (REQ-3-003): when set, the sandbox is telemetry-wired — the
|
||||
backend copies `scripts/sandbox-agent.py` into the workdir and starts
|
||||
the stdlib capture agent inside the sandbox with the `NC_*` env baked
|
||||
here (identity + WS ingest URL). The agent's lifecycle is tied to the
|
||||
sandbox: `destroy()` reaps it (agent → inner → helper). When `task_id`
|
||||
is None the sandbox is a pure shell sandbox (no capture).
|
||||
"""
|
||||
async with self._lock:
|
||||
if len(self._handles) >= self._settings.sandbox_max_concurrent:
|
||||
raise PoolFullError(
|
||||
f"sandbox pool full "
|
||||
f"({len(self._handles)}/{self._settings.sandbox_max_concurrent}); "
|
||||
"no queue (D-032) — retry later"
|
||||
)
|
||||
sandbox_id = f"sbx-{uuid.uuid4().hex[:12]}"
|
||||
spec = workdir_mod.spec_for(sandbox_id, learner_id, self._settings)
|
||||
if task_id is not None:
|
||||
spec = spec.model_copy(
|
||||
update={
|
||||
"capture_env": self._capture_env(sandbox_id, learner_id, task_id)
|
||||
}
|
||||
)
|
||||
handle = await self._backend.spawn(spec)
|
||||
self._handles[handle.id] = handle
|
||||
self._learner_ids[handle.id] = learner_id
|
||||
self._task_ids[handle.id] = task_id
|
||||
self._write_pid_marker(handle, learner_id)
|
||||
logger.info(
|
||||
"sandbox created: id=%s learner=%s task=%s",
|
||||
handle.id,
|
||||
learner_id,
|
||||
task_id or "-",
|
||||
)
|
||||
return self._info_for(handle)
|
||||
|
||||
def _capture_env(self, sandbox_id: str, learner_id: str, task_id: str) -> dict[str, str]:
|
||||
"""Env baked for the in-sandbox capture agent (REQ-3-003).
|
||||
|
||||
The agent joins the sandbox mount namespace but NOT its (offline)
|
||||
network namespace, so it reaches this service over loopback
|
||||
(`telemetry_ingest_host`, A-004 port).
|
||||
"""
|
||||
from urllib.parse import urlencode
|
||||
|
||||
query = urlencode(
|
||||
{
|
||||
"learner_id": learner_id,
|
||||
"task_id": task_id,
|
||||
"sandbox_id": sandbox_id,
|
||||
}
|
||||
)
|
||||
ingest_url = (
|
||||
f"ws://{self._settings.telemetry_ingest_host}:{self._settings.port}"
|
||||
f"/v1/telemetry/ingest?{query}"
|
||||
)
|
||||
return {
|
||||
"NC_LEARNER_ID": learner_id,
|
||||
"NC_TASK_ID": task_id,
|
||||
"NC_SANDBOX_ID": sandbox_id,
|
||||
"NC_INGEST_URL": ingest_url,
|
||||
}
|
||||
|
||||
async def list(self) -> list[SandboxHandleInfo]:
|
||||
"""All live sandboxes (idle + busy; the backend has no busy flag)."""
|
||||
async with self._lock:
|
||||
return [self._info_for(h) for h in self._handles.values()]
|
||||
|
||||
async def get(self, sandbox_id: str) -> SandboxHandleInfo:
|
||||
async with self._lock:
|
||||
handle = self._handles.get(sandbox_id)
|
||||
learner_id = self._learner_ids.get(sandbox_id, "unknown")
|
||||
if handle is None:
|
||||
raise SandboxNotFoundError(sandbox_id)
|
||||
return SandboxHandleInfo(
|
||||
id=handle.id,
|
||||
learner_id=learner_id,
|
||||
workdir=handle.workdir,
|
||||
created_at=handle.created_at,
|
||||
pid=handle.pid,
|
||||
)
|
||||
|
||||
async def snapshot(self, sandbox_id: str) -> Path:
|
||||
"""Copy the workspace into `<workdir>/snapshots/<utc-ts>/`; return it."""
|
||||
async with self._lock:
|
||||
handle = self._handles.get(sandbox_id)
|
||||
if handle is None:
|
||||
raise SandboxNotFoundError(sandbox_id)
|
||||
return await self._backend.snapshot(handle)
|
||||
|
||||
async def destroy(self, sandbox_id: str, *, purge_workdir: bool = False) -> None:
|
||||
"""Tear down one sandbox (idempotent).
|
||||
|
||||
`purge_workdir=False` keeps the workdir on disk — snapshots must
|
||||
survive destroy so a learner's last state can be restored (this is
|
||||
also why UnshareBackend.destroy intentionally leaves the tree alone).
|
||||
`purge_workdir=True` removes the whole workdir.
|
||||
"""
|
||||
async with self._lock:
|
||||
handle = self._handles.pop(sandbox_id, None)
|
||||
learner_id = self._learner_ids.pop(sandbox_id, "unknown")
|
||||
self._task_ids.pop(sandbox_id, None)
|
||||
if handle is not None:
|
||||
await self._backend.destroy(handle)
|
||||
logger.info(
|
||||
"sandbox destroyed: id=%s learner=%s purge=%s",
|
||||
sandbox_id,
|
||||
learner_id,
|
||||
purge_workdir,
|
||||
)
|
||||
root = handle.workdir
|
||||
else:
|
||||
# Idempotent destroy of an unknown id: resolve the on-disk root so
|
||||
# an explicit purge still works (e.g. cleanup of orphan leftovers).
|
||||
root = workdir_mod.resolve_sandbox_dir(self._settings) / sandbox_id
|
||||
if purge_workdir:
|
||||
shutil.rmtree(root, ignore_errors=True)
|
||||
|
||||
# -- reapers ------------------------------------------------------------
|
||||
|
||||
async def reap_expired(self) -> list[str]:
|
||||
"""One reaper pass: wall-clock timeout + G-2 workdir-size sweep.
|
||||
|
||||
Destroys sandboxes older than `settings.sandbox_timeout_s`, then sweeps
|
||||
every remaining sandbox whose `workspace/` exceeds
|
||||
`settings.sandbox_max_workdir_mb` (snapshot → destroy → integrity
|
||||
signal). Returns the ids destroyed this pass. Invoke on an async timer
|
||||
(the app lifespan owns the loop interval); both checks deliberately
|
||||
share one pass so the periodic work is O(live sandboxes) once.
|
||||
"""
|
||||
now = self._clock()
|
||||
destroyed: list[str] = []
|
||||
timeout_s = float(self._settings.sandbox_timeout_s)
|
||||
cap_bytes = int(self._settings.sandbox_max_workdir_mb) * 1024 * 1024
|
||||
|
||||
async with self._lock:
|
||||
rows = [
|
||||
(handle, self._learner_ids.get(handle.id, "unknown"), handle.created_at)
|
||||
for handle in self._handles.values()
|
||||
]
|
||||
|
||||
for handle, learner_id, created_at in rows:
|
||||
age_s = (now - created_at).total_seconds()
|
||||
if age_s > timeout_s:
|
||||
await self.destroy(handle.id)
|
||||
destroyed.append(handle.id)
|
||||
logger.warning(
|
||||
"sandbox reaped (timeout): id=%s age=%.0fs > %.0fs",
|
||||
handle.id,
|
||||
age_s,
|
||||
timeout_s,
|
||||
)
|
||||
continue # already gone; no size sweep needed on a dead handle
|
||||
size = _tree_size_bytes(workdir_mod.workspace_path_from_workdir(handle.workdir))
|
||||
if size > cap_bytes:
|
||||
await self._reap_oversized(handle, learner_id, size, cap_bytes)
|
||||
destroyed.append(handle.id)
|
||||
return destroyed
|
||||
|
||||
async def start(self) -> None:
|
||||
"""Boot hook (a-1): reap on-disk orphans left by a previous process.
|
||||
|
||||
The handle registry is in-memory and process-local (D-019 precedent):
|
||||
after a restart nothing here remembers old sandboxes, so we scan
|
||||
`settings.sandbox_dir` for workdirs whose pid marker names a dead
|
||||
process and purge them, logging a warning. Idempotent; safe to call
|
||||
once per process lifetime.
|
||||
"""
|
||||
if self._started:
|
||||
return
|
||||
self._started = True
|
||||
root = workdir_mod.resolve_sandbox_dir(self._settings)
|
||||
if not root.is_dir():
|
||||
return
|
||||
for entry in sorted(root.iterdir()):
|
||||
if not entry.is_dir():
|
||||
continue
|
||||
marker = entry / PID_MARKER
|
||||
pid = _read_marker_pid(marker)
|
||||
if pid is not None and _pid_alive(pid):
|
||||
continue # live sandbox owned by another live process — leave it
|
||||
logger.warning(
|
||||
"startup reaper (a-1): reaping orphaned workdir %s "
|
||||
"(recorded pid %s is dead or marker missing)",
|
||||
entry,
|
||||
pid,
|
||||
)
|
||||
shutil.rmtree(entry, ignore_errors=True)
|
||||
self._record_integrity(
|
||||
SandboxIntegrityEvent(
|
||||
kind="orphan_reaped",
|
||||
sandbox_id=entry.name,
|
||||
learner_id="unknown",
|
||||
detail=f"workdir {entry} reaped at boot; recorded pid={pid} dead",
|
||||
)
|
||||
)
|
||||
|
||||
async def destroy_all(self) -> None:
|
||||
"""Shutdown hook: destroy every live sandbox (no orphans on exit).
|
||||
|
||||
Workdirs (and their snapshots) are kept on disk — destroy semantics
|
||||
here match `destroy(purge_workdir=False)`; the next boot's startup
|
||||
reaper (a-1) decides what to clean based on pid markers.
|
||||
"""
|
||||
async with self._lock:
|
||||
handles = list(self._handles.values())
|
||||
for handle in handles:
|
||||
await self.destroy(handle.id)
|
||||
|
||||
# -- internals ------------------------------------------------------------
|
||||
|
||||
def _info_for(self, handle: SandboxHandle) -> SandboxHandleInfo:
|
||||
# Caller holds the lock (create/list) — the side-table read is atomic.
|
||||
return SandboxHandleInfo(
|
||||
id=handle.id,
|
||||
learner_id=self._learner_ids.get(handle.id, "unknown"),
|
||||
workdir=handle.workdir,
|
||||
created_at=handle.created_at,
|
||||
pid=handle.pid,
|
||||
)
|
||||
|
||||
def _write_pid_marker(self, handle: SandboxHandle, learner_id: str) -> None:
|
||||
marker = handle.workdir / PID_MARKER
|
||||
try:
|
||||
marker.write_text(
|
||||
json.dumps(
|
||||
{
|
||||
"sandbox_id": handle.id,
|
||||
"learner_id": learner_id,
|
||||
"pid": os.getpid(),
|
||||
"created_at": handle.created_at.isoformat(),
|
||||
}
|
||||
)
|
||||
)
|
||||
except OSError: # marker is advisory; spawning must not fail on it
|
||||
logger.warning("could not write pid marker %s", marker)
|
||||
|
||||
async def _reap_oversized(
|
||||
self,
|
||||
handle: SandboxHandle,
|
||||
learner_id: str,
|
||||
size_bytes: int,
|
||||
cap_bytes: int,
|
||||
) -> None:
|
||||
"""G-2 sweep step: snapshot evidence → destroy → record the signal."""
|
||||
snapshot_path: Path | None = None
|
||||
try:
|
||||
snapshot_path = await self._backend.snapshot(handle)
|
||||
except (OSError, RuntimeError):
|
||||
logger.exception(
|
||||
"G-2 sweep: snapshot failed for over-cap sandbox %s; destroying anyway",
|
||||
handle.id,
|
||||
)
|
||||
await self.destroy(handle.id)
|
||||
event = SandboxIntegrityEvent(
|
||||
kind="workdir_size_cap",
|
||||
sandbox_id=handle.id,
|
||||
learner_id=learner_id,
|
||||
detail=(
|
||||
f"workspace {size_bytes}B exceeded soft cap {cap_bytes}B; "
|
||||
f"snapshot={snapshot_path} then destroyed (G-2, best-effort, "
|
||||
"NOT kernel-enforced)"
|
||||
),
|
||||
)
|
||||
self._record_integrity(event)
|
||||
logger.warning(
|
||||
"G-2 workdir sweep: sandbox %s (learner=%s) destroyed over soft disk cap",
|
||||
handle.id,
|
||||
learner_id,
|
||||
)
|
||||
|
||||
def _record_integrity(self, event: SandboxIntegrityEvent) -> None:
|
||||
self._integrity_events.append(event)
|
||||
|
||||
|
||||
# -- module helpers ---------------------------------------------------------
|
||||
|
||||
|
||||
def _tree_size_bytes(root: Path) -> int:
|
||||
"""Total bytes under `root` (best-effort; unreadable entries count 0)."""
|
||||
if not root.is_dir():
|
||||
return 0
|
||||
total = 0
|
||||
for dirpath, _dirnames, filenames in os.walk(root):
|
||||
for name in filenames:
|
||||
try:
|
||||
total += (Path(dirpath) / name).lstat().st_size
|
||||
except OSError:
|
||||
continue
|
||||
return total
|
||||
|
||||
|
||||
def _read_marker_pid(marker: Path) -> int | None:
|
||||
try:
|
||||
data = json.loads(marker.read_text())
|
||||
except (OSError, json.JSONDecodeError):
|
||||
return None
|
||||
pid = data.get("pid")
|
||||
return pid if isinstance(pid, int) else None
|
||||
|
||||
|
||||
def _pid_alive(pid: int) -> bool:
|
||||
"""True if `pid` exists on this host (signal 0 probe; no signal sent)."""
|
||||
try:
|
||||
os.kill(pid, 0)
|
||||
except ProcessLookupError:
|
||||
return False
|
||||
except PermissionError:
|
||||
return True # exists, owned by another user
|
||||
return True
|
||||
@@ -1,548 +0,0 @@
|
||||
"""UnshareBackend — D-024 Linux-namespace sandboxing via util-linux `unshare`.
|
||||
|
||||
Two execution modes share one backend:
|
||||
|
||||
1. Pure shell sandbox (`task_id is None`, REQ-3-001): isolation is established
|
||||
PER-EXEC — every `exec` spawns a fresh namespace:
|
||||
|
||||
unshare --user --map-root-user --mount --pid --fork --net sh -c '<shim>'
|
||||
|
||||
There is no persistent process; the in-namespace shim is:
|
||||
|
||||
mount -t tmpfs tmpfs /tmp # private scratch, discarded on exit
|
||||
mkdir -p /tmp/work
|
||||
mount --bind <host workspace> /tmp/work
|
||||
cd /tmp/work
|
||||
ulimit -v/-t/-f … # applied AFTER the bind, so rlimits
|
||||
exec <cmd> # constrain the PAYLOAD, not unshare
|
||||
|
||||
2. Telemetry-wired task sandbox (REQ-3-003, `capture_env` set): a PERSISTENT,
|
||||
TRACKED topology so the stdlib capture agent can live inside the sandbox and
|
||||
still stream events to ai-service. Per exec a fresh OFFLINE namespace would
|
||||
leave the agent nowhere to run and (on this host, where a userns can't
|
||||
bring `lo` up) no loopback to reach `ws://127.0.0.1`. So spawn creates a
|
||||
long-lived helper (outer user+mount ns, ONLINE) and an inner sandbox
|
||||
(mount+pid+fork+net — OFFLINE), both rooted at a private `ns/` subtree:
|
||||
|
||||
helper : unshare --user --map-root-user --mount (mounts ns/ private)
|
||||
inner : unshare --mount --pid --fork --net (tmpfs on ns/, bind
|
||||
<workdir>/host/workspace -> <ns>/work) <- the sandbox
|
||||
exec : nsenter -t <inner sleep> -m -- sh -c … (joins inner mount ns;
|
||||
offline + pid-isolated, uid 0, writes land on the host workspace)
|
||||
agent : nsenter -t <inner sleep> -m -- python3 <agent> (joins the inner
|
||||
MOUNT ns only — NOT pid/net — so it watches the live workspace
|
||||
and stays ONLINE, reaching the app's WS ingest on loopback)
|
||||
|
||||
The agent is deliberately pid/net-exempt from the sandbox: it is OUR trusted
|
||||
capture process, and isolating its network would cut the very link it needs.
|
||||
`destroy` reaps agent → inner → helper (in that order). The handle's `pid`
|
||||
is the agent's host pid (None for a pure shell sandbox).
|
||||
|
||||
Why rlimits are applied in the shim, not Python's preexec_fn: setting
|
||||
RLIMIT_AS on the *unshare* process itself can trip the memory ceiling on the
|
||||
post-fork Python parent (whose interpreter image already exceeds the sandbox
|
||||
budget). Applying them in the innermost child — just before exec'ing the
|
||||
payload — keeps `unshare`/`mount` unconstrained and limits the learner code.
|
||||
|
||||
Containment honesty (D-024 / G-1): a user namespace is NOT a write barrier.
|
||||
Writes made OUTSIDE the bind fall through to host paths, and because inner
|
||||
uid 0 maps to the invoking host uid, a sandboxed process can write anywhere
|
||||
that host uid can write. Isolation here is: private PIDs/MNT/NET/UTS, tmpfs
|
||||
scratch, payload rlimits, and a uid map yielding no privilege the host uid
|
||||
did not already have. A per-sandbox runtime uid (D-025) is the follow-up that
|
||||
hardens DAC.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import os
|
||||
import shlex
|
||||
import shutil
|
||||
import signal
|
||||
import time
|
||||
from datetime import UTC, datetime
|
||||
from pathlib import Path
|
||||
|
||||
from .backend import ExecResult, ResourceLimits, SandboxHandle, SandboxSpec
|
||||
from .workdir import create_layout
|
||||
from .workdir import snapshot as workdir_snapshot
|
||||
|
||||
#: Args shared by every PURE shell namespace we spawn (D-024). No
|
||||
#: `unshare --bind` on util-linux 2.38 — the bind is done from inside instead.
|
||||
UNSHARE_ARGS: tuple[str, ...] = (
|
||||
"--user", # new user namespace …
|
||||
"--map-root-user", # … in which we are uid 0 (mapped to host uid outside)
|
||||
"--mount", # private mount table
|
||||
"--pid", # private PID table
|
||||
"--fork", # child is PID 1 in its namespace (reaps zombies, gets signals)
|
||||
"--net", # fresh net namespace: no usable route → effectively offline
|
||||
)
|
||||
|
||||
IN_NS_WORKDIR = "/tmp/work" # where the workspace is bound inside a pure shell ns
|
||||
|
||||
#: Sentinels the long-lived namespace supervisors print once their mounts are
|
||||
#: laid out. exec()/the manager must not run before the bind exists.
|
||||
_HELPER_READY = "NC_HELPER_READY"
|
||||
_INNER_READY = "NC_INNER_READY"
|
||||
|
||||
|
||||
class SandboxUnavailableError(RuntimeError):
|
||||
"""`unshare`/`nsenter` missing or user namespaces blocked on this host."""
|
||||
|
||||
|
||||
def _build_shim(workspace: Path, limits: ResourceLimits, cmd: list[str]) -> str:
|
||||
"""Compose the single POSIX string executed by the in-namespace /bin/sh.
|
||||
|
||||
The shim runs under `unshare`'s forked child → does the bind mounts →
|
||||
forks a subshell that applies rlimits → and `exec`s the payload. Applying
|
||||
rlimits in the subshell (last hop) keeps the memory/tools unconstrained
|
||||
and constrains only the learner process.
|
||||
"""
|
||||
quoted_cmd = " ".join(shlex.quote(part) for part in cmd)
|
||||
rlimit_prefix = (
|
||||
f"ulimit -v {limits.memory_bytes // 1024}; " # RLIMIT_AS, KB
|
||||
f"ulimit -t {limits.cpu_seconds}; " # RLIMIT_CPU, s
|
||||
f"ulimit -f {limits.file_size_bytes // 512}; " # RLIMIT_FSIZE, 512 blocks
|
||||
)
|
||||
return (
|
||||
"set -eu; "
|
||||
"mount -t tmpfs tmpfs /tmp; "
|
||||
f"mkdir -p {IN_NS_WORKDIR}; "
|
||||
f"mount --bind {shlex.quote(str(workspace))} {IN_NS_WORKDIR}; "
|
||||
f"cd {IN_NS_WORKDIR}; "
|
||||
f"exec sh -c {shlex.quote(rlimit_prefix + 'exec ' + quoted_cmd)}"
|
||||
)
|
||||
|
||||
|
||||
class _Tracked:
|
||||
"""The process tree + paths for one persistent (telemetry-wired) sandbox."""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
*,
|
||||
helper: asyncio.subprocess.Process,
|
||||
inner: asyncio.subprocess.Process,
|
||||
agent: asyncio.subprocess.Process | None,
|
||||
inner_pid: int, # host pid of the SANDBOXED init (sleep) — ns enter target
|
||||
host_dir: Path,
|
||||
workspace: Path,
|
||||
ns_root: Path,
|
||||
ns_workdir: Path,
|
||||
) -> None:
|
||||
self.helper = helper
|
||||
self.inner = inner
|
||||
self.agent = agent
|
||||
self.inner_pid = inner_pid
|
||||
self.host_dir = host_dir
|
||||
self.workspace = workspace
|
||||
self.ns_root = ns_root
|
||||
self.ns_workdir = ns_workdir
|
||||
|
||||
|
||||
class UnshareBackend: # satisfies SandboxBackend structurally (Protocol)
|
||||
"""D-024 backend: namespace subprocesses; persistent tree for task sandboxes."""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
unshare_path: str | None = None,
|
||||
limits: ResourceLimits | None = None, # per-spec override lands in 1-04
|
||||
nsenter_path: str | None = None,
|
||||
agent_script: Path | None = None,
|
||||
) -> None:
|
||||
self._unshare = unshare_path or shutil.which("unshare") or "unshare"
|
||||
self._nsenter = nsenter_path or shutil.which("nsenter") or "nsenter"
|
||||
self._limits = limits or ResourceLimits()
|
||||
# The stdlib-only capture agent script, copied into each tracked
|
||||
# workdir's host/ tree so nsenter can reach it inside the sandbox.
|
||||
# ai_service/sandbox/unshare_backend.py -> parents[2] = apps/ai-service.
|
||||
self._agent_script = agent_script or (
|
||||
Path(__file__).resolve().parents[2] / "scripts" / "sandbox-agent.py"
|
||||
)
|
||||
# Tracked (persistent) sandboxes by id; pure shell sandboxes are absent.
|
||||
self._tracked: dict[str, _Tracked] = {}
|
||||
|
||||
# -- spawn ------------------------------------------------------------------
|
||||
|
||||
async def spawn(self, spec: SandboxSpec) -> SandboxHandle:
|
||||
"""Lay out the workdir; if `spec.capture_env` is set, start the sandbox.
|
||||
|
||||
A spec WITHOUT capture_env keeps REQ-3-001 semantics: spawn only
|
||||
prepares disk state and each exec forks a fresh (offline) namespace.
|
||||
A spec WITH capture_env starts the persistent helper/inner tree and the
|
||||
capture agent, and `handle.pid` carries the agent's host pid.
|
||||
"""
|
||||
create_layout(spec)
|
||||
if not spec.capture_env:
|
||||
return SandboxHandle(
|
||||
id=spec.sandbox_id,
|
||||
pid=None, # no persistent process; each exec forks short-lived PIDs
|
||||
workdir=spec.workdir,
|
||||
created_at=datetime.now(UTC),
|
||||
)
|
||||
tracked = await self._spawn_tracked(spec)
|
||||
self._tracked[spec.sandbox_id] = tracked
|
||||
return SandboxHandle(
|
||||
id=spec.sandbox_id,
|
||||
pid=tracked.agent.pid if tracked.agent is not None else tracked.inner_pid,
|
||||
workdir=spec.workdir,
|
||||
created_at=datetime.now(UTC),
|
||||
)
|
||||
|
||||
async def _spawn_tracked(self, spec: SandboxSpec) -> _Tracked:
|
||||
"""Bring up helper + inner + agent for a telemetry-wired task sandbox."""
|
||||
host_dir = spec.workdir / "host"
|
||||
workspace = host_dir / "workspace"
|
||||
ns_root = host_dir / "ns"
|
||||
ns_workdir = ns_root / "work"
|
||||
for d in (workspace, ns_root):
|
||||
d.mkdir(parents=True, exist_ok=True)
|
||||
# The agent script must live INSIDE the workspace: the inner ns bind
|
||||
# mounts <host_dir>/workspace -> <ns_root>/work, so only workspace
|
||||
# content is visible in-namespace at /work.
|
||||
agent_host_path = workspace / "sandbox-agent.py"
|
||||
shutil.copyfile(self._agent_script, agent_host_path)
|
||||
|
||||
helper = await self._launch_ns(
|
||||
[
|
||||
self._unshare,
|
||||
"--user",
|
||||
"--map-root-user",
|
||||
"--mount",
|
||||
"sh",
|
||||
"-c",
|
||||
(
|
||||
# Isolate ns/ so the inner tmpfs never propagates back to the
|
||||
# host mount table (make-private is best-effort on this host).
|
||||
f"mount --bind {shlex.quote(str(ns_root))} {shlex.quote(str(ns_root))}; "
|
||||
f"mount --make-private {shlex.quote(str(ns_root))} 2>/dev/null; "
|
||||
f"echo {_HELPER_READY}; exec sleep 3600"
|
||||
),
|
||||
],
|
||||
sentinel=_HELPER_READY,
|
||||
label="helper",
|
||||
)
|
||||
try:
|
||||
inner = await self._launch_ns(
|
||||
[
|
||||
*self._helper_join_argv(helper),
|
||||
self._unshare,
|
||||
"--mount",
|
||||
"--pid",
|
||||
"--fork",
|
||||
"--net",
|
||||
"sh",
|
||||
"-c",
|
||||
(
|
||||
f"mount -t tmpfs tmpfs {shlex.quote(str(ns_root))}; "
|
||||
f"mkdir -p {shlex.quote(str(ns_workdir))}; "
|
||||
f"mount --bind {shlex.quote(str(workspace))} "
|
||||
f"{shlex.quote(str(ns_workdir))}; "
|
||||
f"echo {_INNER_READY}; exec sleep 3600"
|
||||
),
|
||||
],
|
||||
sentinel=_INNER_READY,
|
||||
label="inner",
|
||||
)
|
||||
except Exception:
|
||||
await self._reap(helper)
|
||||
raise
|
||||
|
||||
await asyncio.sleep(0) # let the inner child's sleep fork settle
|
||||
inner_pid = await asyncio.to_thread(self._find_child_pid, inner.pid)
|
||||
if inner_pid is None:
|
||||
await self._reap(inner)
|
||||
await self._reap(helper)
|
||||
raise SandboxUnavailableError(
|
||||
f"could not resolve sandboxed init pid for {spec.sandbox_id}"
|
||||
)
|
||||
|
||||
tracked = _Tracked(
|
||||
helper=helper,
|
||||
inner=inner,
|
||||
agent=None,
|
||||
inner_pid=inner_pid,
|
||||
host_dir=host_dir,
|
||||
workspace=workspace,
|
||||
ns_root=ns_root,
|
||||
ns_workdir=ns_workdir,
|
||||
)
|
||||
if spec.capture_env:
|
||||
tracked.agent = await self._launch_agent(spec, tracked, agent_host_path)
|
||||
return tracked
|
||||
|
||||
# -- process launch helpers --------------------------------------------------
|
||||
|
||||
def _helper_join_argv(self, helper: asyncio.subprocess.Process) -> list[str]:
|
||||
"""nsenter argv that runs a command inside the helper's user+mount ns."""
|
||||
if helper.pid is None:
|
||||
raise SandboxUnavailableError("helper namespace process is not running")
|
||||
return [
|
||||
self._nsenter,
|
||||
"-t",
|
||||
str(helper.pid),
|
||||
"-m",
|
||||
"-U",
|
||||
"--preserve-credentials",
|
||||
"--",
|
||||
]
|
||||
|
||||
def _sandbox_join_argv(self, tracked: _Tracked) -> list[str]:
|
||||
"""nsenter argv that joins the inner sandbox MOUNT namespace (uid 0)."""
|
||||
return [self._nsenter, "-t", str(tracked.inner_pid), "-m", "--"]
|
||||
|
||||
async def _launch_agent(
|
||||
self, spec: SandboxSpec, tracked: _Tracked, agent_host_path: Path
|
||||
) -> asyncio.subprocess.Process:
|
||||
"""Launch the capture agent: joins the sandbox mount ns, NOT pid/net.
|
||||
|
||||
The nsenter chain swaps the mount table under the process, so a HOST
|
||||
cwd/relative path is invalid after the join (observed: python3
|
||||
resolved ``sandbox-agent.py`` against a stale root → ``//…`` and
|
||||
exited rc=2). The launch therefore happens through ``sh -c`` INSIDE
|
||||
the joined namespace, using only in-namespace absolute paths: the
|
||||
workspace is bind-mounted at ``<ns_root>/work``, the agent script was
|
||||
copied into the host workspace, so ``/work/sandbox-agent.py`` exists
|
||||
after the join. The agent runs ONLINE (joins mount ns only, not the
|
||||
offline net ns) so it can dial the ai-service WS ingest loopback.
|
||||
"""
|
||||
env = dict(spec.capture_env or {})
|
||||
in_ns_script = f"{tracked.ns_root / 'work' / 'sandbox-agent.py'}"
|
||||
in_ns_cwd = f"{tracked.ns_root / 'work'}"
|
||||
launch = f"cd {shlex.quote(in_ns_cwd)} && exec python3 {shlex.quote(in_ns_script)}"
|
||||
try:
|
||||
return await asyncio.create_subprocess_exec(
|
||||
*self._helper_join_argv(tracked.helper),
|
||||
*self._sandbox_join_argv(tracked),
|
||||
"sh",
|
||||
"-c",
|
||||
launch,
|
||||
stdin=asyncio.subprocess.DEVNULL,
|
||||
stdout=asyncio.subprocess.DEVNULL,
|
||||
stderr=asyncio.subprocess.DEVNULL,
|
||||
env=env,
|
||||
)
|
||||
except FileNotFoundError as exc: # pragma: no cover - env-dependent
|
||||
raise SandboxUnavailableError("python3 unavailable for capture agent") from exc
|
||||
|
||||
async def _launch_ns(
|
||||
self, argv: list[str], *, sentinel: str, label: str
|
||||
) -> asyncio.subprocess.Process:
|
||||
"""Spawn a namespace supervisor and wait for its `sentinel` line."""
|
||||
proc = await asyncio.create_subprocess_exec(
|
||||
*argv,
|
||||
stdout=asyncio.subprocess.PIPE,
|
||||
stderr=asyncio.subprocess.STDOUT,
|
||||
)
|
||||
|
||||
async def _wait_ready() -> None:
|
||||
if proc.stdout is None: # pragma: no cover (stdout is a PIPE)
|
||||
raise SandboxUnavailableError(f"{label} namespace missing stdout pipe")
|
||||
async for raw in proc.stdout:
|
||||
if raw.decode(errors="replace").strip() == sentinel:
|
||||
return
|
||||
raise SandboxUnavailableError(
|
||||
f"{label} namespace exited before signalling readiness: {argv[:3]}"
|
||||
)
|
||||
|
||||
try:
|
||||
await asyncio.wait_for(_wait_ready(), timeout=10.0)
|
||||
except TimeoutError as exc:
|
||||
await self._reap(proc)
|
||||
raise SandboxUnavailableError(
|
||||
f"{label} namespace never became ready (timeout): {argv[:3]}"
|
||||
) from exc
|
||||
except SandboxUnavailableError:
|
||||
await self._reap(proc)
|
||||
raise
|
||||
return proc
|
||||
|
||||
@staticmethod
|
||||
def _find_child_pid(parent_pid: int | None) -> int | None:
|
||||
"""First direct child of `parent_pid` (the pid-namespaced `sleep`).
|
||||
|
||||
The helper→unshare shim is inner.pid's parent chain head, but the
|
||||
SANDBOXED mount/pid namespaces belong to its forked child (the
|
||||
`sleep`). nsenter must target THAT pid to land inside the sandbox.
|
||||
Reads /proc directly — best-effort, host-local, no subprocess.
|
||||
"""
|
||||
if parent_pid is None:
|
||||
return None
|
||||
for entry in os.listdir("/proc"):
|
||||
if not entry.isdigit():
|
||||
continue
|
||||
try:
|
||||
with open(f"/proc/{entry}/stat") as fh:
|
||||
# ppid is field 4; comm (field 2) may contain spaces, so
|
||||
# parse relative to the LAST ')'.
|
||||
rest = fh.read().rsplit(") ", 1)[1].split()
|
||||
if int(rest[1]) == parent_pid: # state=rest[0], ppid=rest[1]
|
||||
return int(entry)
|
||||
except (OSError, IndexError, ValueError):
|
||||
continue
|
||||
return None
|
||||
|
||||
# -- exec --------------------------------------------------------------------
|
||||
|
||||
async def exec(self, handle: SandboxHandle, cmd: list[str]) -> ExecResult:
|
||||
"""Run `cmd` in the sandbox workspace (cwd = the bound workspace)."""
|
||||
if not cmd:
|
||||
raise ValueError("exec requires a non-empty cmd")
|
||||
tracked = self._tracked.get(handle.id)
|
||||
if tracked is not None:
|
||||
return await self._exec_tracked(handle, tracked, cmd)
|
||||
return await self._exec_fresh(handle, cmd)
|
||||
|
||||
async def _exec_fresh(self, handle: SandboxHandle, cmd: list[str]) -> ExecResult:
|
||||
"""Pure shell sandbox: spawn one fresh offline namespace per exec."""
|
||||
workspace = handle.workdir / "workspace"
|
||||
if not workspace.is_dir():
|
||||
raise SandboxUnavailableError(f"spawn() first: no workspace at {workspace}")
|
||||
argv = [
|
||||
self._unshare,
|
||||
*UNSHARE_ARGS,
|
||||
"sh",
|
||||
"-c",
|
||||
_build_shim(workspace, self._limits, cmd),
|
||||
]
|
||||
started = time.monotonic()
|
||||
proc = await asyncio.create_subprocess_exec(
|
||||
*argv,
|
||||
stdin=asyncio.subprocess.DEVNULL,
|
||||
stdout=asyncio.subprocess.PIPE,
|
||||
stderr=asyncio.subprocess.PIPE,
|
||||
)
|
||||
out, err = await proc.communicate()
|
||||
return ExecResult(
|
||||
cmd=cmd,
|
||||
returncode=proc.returncode if proc.returncode is not None else -1,
|
||||
stdout=out.decode(errors="replace"),
|
||||
stderr=err.decode(errors="replace"),
|
||||
duration_s=time.monotonic() - started,
|
||||
)
|
||||
|
||||
async def _exec_tracked(
|
||||
self, handle: SandboxHandle, tracked: _Tracked, cmd: list[str]
|
||||
) -> ExecResult:
|
||||
"""Task sandbox: join the persistent inner namespace (offline, uid 0).
|
||||
|
||||
rlimits apply in the joining subshell so only the payload is limited;
|
||||
cwd is the bound workspace (`<ns>/work`).
|
||||
"""
|
||||
if tracked.inner.returncode is not None:
|
||||
raise SandboxUnavailableError(
|
||||
f"sandbox {handle.id} is not running (inner namespace exited)"
|
||||
)
|
||||
quoted_cmd = " ".join(shlex.quote(part) for part in cmd)
|
||||
rlimit_prefix = (
|
||||
f"ulimit -v {self._limits.memory_bytes // 1024}; "
|
||||
f"ulimit -t {self._limits.cpu_seconds}; "
|
||||
f"ulimit -f {self._limits.file_size_bytes // 512}; "
|
||||
)
|
||||
shell = (
|
||||
f"cd {shlex.quote(str(tracked.ns_workdir))}; "
|
||||
f"{rlimit_prefix}"
|
||||
f"exec {quoted_cmd}"
|
||||
)
|
||||
argv = [
|
||||
*self._helper_join_argv(tracked.helper),
|
||||
*self._sandbox_join_argv(tracked),
|
||||
"sh",
|
||||
"-c",
|
||||
shell,
|
||||
]
|
||||
started = time.monotonic()
|
||||
proc = await asyncio.create_subprocess_exec(
|
||||
*argv,
|
||||
stdin=asyncio.subprocess.DEVNULL,
|
||||
stdout=asyncio.subprocess.PIPE,
|
||||
stderr=asyncio.subprocess.PIPE,
|
||||
)
|
||||
out, err = await proc.communicate()
|
||||
return ExecResult(
|
||||
cmd=cmd,
|
||||
returncode=proc.returncode if proc.returncode is not None else -1,
|
||||
stdout=out.decode(errors="replace"),
|
||||
stderr=err.decode(errors="replace"),
|
||||
duration_s=time.monotonic() - started,
|
||||
)
|
||||
|
||||
# -- snapshot / destroy --------------------------------------------------------
|
||||
|
||||
async def snapshot(self, handle: SandboxHandle) -> Path:
|
||||
workspace_root = handle.workdir
|
||||
tracked = self._tracked.get(handle.id)
|
||||
if tracked is not None:
|
||||
# Copy the tracked workspace, not the legacy <workdir>/workspace.
|
||||
dest_parent = handle.workdir / "snapshots"
|
||||
dest_parent.mkdir(parents=True, exist_ok=True)
|
||||
return workdir_snapshot_from_workspace(tracked.workspace, dest_parent)
|
||||
return workdir_snapshot(workspace_root)
|
||||
|
||||
async def destroy(self, handle: SandboxHandle) -> None:
|
||||
"""Best-effort teardown. Pure shell sandboxes die with their exec; for a
|
||||
tracked task sandbox reap AGENT → INNER → HELPER so no capture process
|
||||
or namespace supervisor outlives the handle (REQ-3-003 lifecycle).
|
||||
|
||||
Keeping the workdir is deliberate: snapshots must survive destroy so a
|
||||
learner's last state can be restored by the manager layer.
|
||||
"""
|
||||
tracked = self._tracked.pop(handle.id, None)
|
||||
if tracked is not None:
|
||||
# Agent first (it must not flush a "stopped" event into a dead
|
||||
# sandbox), then the namespace tree. The inner `unshare --fork`
|
||||
# shim is NOT the namespace init: killing it orphans its child
|
||||
# (the `sleep` that is PID 1 of the sandbox pid+mnt+net ns),
|
||||
# which reparents to host init and holds the tmpfs + bind for
|
||||
# a full hour (observed: ~30 leaked `sleep 3600` after a test
|
||||
# run). `--kill-child` does not reach it either (util-linux
|
||||
# 2.38 leaks the same child under this flag combo — the child
|
||||
# is reparented before unshare's signal handler runs). The
|
||||
# deterministic kill is SIGKILL on the ns-init's HOST pid,
|
||||
# which we already track as `tracked.inner_pid` (nsenter uses
|
||||
# it for exec); the kernel then tears down the namespace with
|
||||
# its init (no processes remain).
|
||||
for proc in (tracked.agent, tracked.inner, tracked.helper):
|
||||
if proc is not None:
|
||||
await self._reap(proc)
|
||||
self._kill_pid(tracked.inner_pid)
|
||||
handle.pid = None
|
||||
|
||||
@staticmethod
|
||||
def _kill_pid(pid: int | None, sig: int = signal.SIGKILL) -> None:
|
||||
"""Best-effort host-side signal; pid recycled or gone is not an error."""
|
||||
if pid is None:
|
||||
return
|
||||
try:
|
||||
os.kill(pid, sig)
|
||||
except (ProcessLookupError, PermissionError):
|
||||
pass # already dead, or not ours — nothing to do
|
||||
|
||||
@staticmethod
|
||||
async def _reap(proc: asyncio.subprocess.Process) -> None:
|
||||
"""SIGTERM then SIGKILL, tolerant of an already-dead process."""
|
||||
if proc.returncode is not None:
|
||||
return
|
||||
try:
|
||||
proc.terminate()
|
||||
except ProcessLookupError:
|
||||
return
|
||||
try:
|
||||
await asyncio.wait_for(proc.wait(), timeout=5.0)
|
||||
except TimeoutError:
|
||||
try:
|
||||
proc.kill()
|
||||
except ProcessLookupError:
|
||||
return
|
||||
try:
|
||||
await asyncio.wait_for(proc.wait(), timeout=5.0)
|
||||
except TimeoutError: # pragma: no cover - SIGKILL always wins
|
||||
pass
|
||||
|
||||
|
||||
def workdir_snapshot_from_workspace(workspace: Path, snapshots_dir: Path) -> Path:
|
||||
"""Snapshot helper for tracked sandboxes whose workspace is `<workdir>/host/workspace`
|
||||
instead of the legacy `<workdir>/workspace` layout."""
|
||||
dest = snapshots_dir / datetime.now(UTC).strftime("%Y%m%dT%H%M%SZ")
|
||||
shutil.copytree(workspace, dest, symlinks=False)
|
||||
return dest
|
||||
@@ -1,81 +0,0 @@
|
||||
"""Per-sandbox workdir layout and snapshots (REQ-3-001).
|
||||
|
||||
Layout, rooted under `settings.sandbox_dir` (default `apps/ai-service/sandboxes/`):
|
||||
|
||||
<sandbox_dir>/<sandbox_id>/
|
||||
workspace/ bind-mounted into the namespace at /work (learner-writable)
|
||||
snapshots/ host-side timestamped copies produced by snapshot()
|
||||
|
||||
The workspace is the ONLY directory the namespaced process can write that is
|
||||
also visible on the host. Everything else either stays on the host (see
|
||||
UnshareBackend's DAC note) or lands in a discarded tmpfs.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import shutil
|
||||
from datetime import UTC, datetime
|
||||
from pathlib import Path
|
||||
|
||||
from pydantic import BaseModel, ConfigDict
|
||||
|
||||
from ..config import Settings
|
||||
from .backend import SandboxSpec
|
||||
|
||||
|
||||
class SandboxDir(BaseModel):
|
||||
"""Concrete paths for one sandbox's on-disk layout."""
|
||||
|
||||
model_config = ConfigDict(arbitrary_types_allowed=True)
|
||||
|
||||
root: Path
|
||||
workspace: Path
|
||||
snapshots: Path
|
||||
|
||||
|
||||
def workspace_path(spec: SandboxSpec) -> Path:
|
||||
"""Return the workspace path for a sandbox laid out under `spec.workdir`."""
|
||||
return workspace_path_from_workdir(spec.workdir)
|
||||
|
||||
|
||||
def workspace_path_from_workdir(workdir: Path) -> Path:
|
||||
"""Workspace path given a sandbox workdir root."""
|
||||
return workdir / "workspace"
|
||||
|
||||
|
||||
def create_layout(spec: SandboxSpec) -> SandboxDir:
|
||||
"""Create `<workdir>/{workspace,snapshots}` (parents included, idempotent)."""
|
||||
layout = SandboxDir(
|
||||
root=spec.workdir,
|
||||
workspace=spec.workdir / "workspace",
|
||||
snapshots=spec.workdir / "snapshots",
|
||||
)
|
||||
layout.workspace.mkdir(parents=True, exist_ok=True)
|
||||
layout.snapshots.mkdir(parents=True, exist_ok=True)
|
||||
return layout
|
||||
|
||||
|
||||
def snapshot(workdir: Path) -> Path:
|
||||
"""Recursively copy `<workdir>/workspace` to `<workdir>/snapshots/<utc-ts>/`.
|
||||
|
||||
Symlinks are never followed or recreated (`symlinks=False`); a symlink in
|
||||
the workspace is replaced by the file it points at, so a snapshot can
|
||||
never retain a host-escape link. Returns the new snapshot directory.
|
||||
"""
|
||||
dest = workdir / "snapshots" / datetime.now(UTC).strftime("%Y%m%dT%H%M%SZ")
|
||||
shutil.copytree(workdir / "workspace", dest, symlinks=False)
|
||||
return dest
|
||||
|
||||
|
||||
def resolve_sandbox_dir(settings: Settings) -> Path:
|
||||
"""Resolve `settings.sandbox_dir` (relative → anchored at the app dir)."""
|
||||
sandbox_dir = settings.sandbox_dir
|
||||
if sandbox_dir.is_absolute():
|
||||
return sandbox_dir
|
||||
return (Path(__file__).resolve().parent.parent / sandbox_dir).resolve()
|
||||
|
||||
|
||||
def spec_for(sandbox_id: str, learner_id: str, settings: Settings) -> SandboxSpec:
|
||||
"""Build a `SandboxSpec` rooted under the configured sandbox dir."""
|
||||
workdir = resolve_sandbox_dir(settings) / sandbox_id
|
||||
return SandboxSpec(sandbox_id=sandbox_id, learner_id=learner_id, workdir=workdir)
|
||||
@@ -1,21 +0,0 @@
|
||||
"""Live build telemetry — event models and the TraceStore protocol (REQ-3-003).
|
||||
|
||||
Boundary rule (D-027): telemetry/ imports from config only — never from
|
||||
agents/ or api/ (agents call engines through narrow interfaces, never
|
||||
the reverse; api/ composes stores via DI).
|
||||
"""
|
||||
|
||||
from .ingest import IngestSession, TraceIntegrityMap, telemetry_ingest_endpoint
|
||||
from .models import EventKind, TelemetryEvent, TraceSpan
|
||||
from .store import SQLiteTraceStore, TraceStore
|
||||
|
||||
__all__ = [
|
||||
"EventKind",
|
||||
"IngestSession",
|
||||
"SQLiteTraceStore",
|
||||
"TelemetryEvent",
|
||||
"TraceIntegrityMap",
|
||||
"TraceSpan",
|
||||
"TraceStore",
|
||||
"telemetry_ingest_endpoint",
|
||||
]
|
||||
@@ -1,458 +0,0 @@
|
||||
"""WS ingest protocol for learner telemetry (REQ-3-003, D-026, G-3).
|
||||
|
||||
Frame contract — trace identity travels as QUERY PARAMS on the WS upgrade
|
||||
(`WS /v1/telemetry/ingest?learner_id=...&task_id=...&sandbox_id=...`), NOT as
|
||||
a first init frame. Rationale: the in-sandbox capture agent (Task 2-2-01) is a
|
||||
stdlib-only RFC6455 client where the URL is the cheapest thing to parametrize
|
||||
(`NC_INGEST_URL` carries the query string); identity is also visible to the
|
||||
server BEFORE accept(), so a malformed handshake can be rejected without an
|
||||
accept/close round-trip. Client messages are then ONE event per JSON text
|
||||
frame — no envelope:
|
||||
|
||||
{"seq": 0, "kind": "command", "payload": {...}, "ts": "...",
|
||||
"sandbox_id": "..."} # learner_id / task_id forbidden (URL owns them)
|
||||
|
||||
Server → client frames are typed status envelopes:
|
||||
|
||||
{"type": "ack_total", "count": N} — final flush summary, then close 1000
|
||||
{"type": "seq_ack", "seq": N} — advisory: durable latest_seq after
|
||||
each successful append (D-045; the
|
||||
agent trims its spool to seq > ack)
|
||||
{"type": "gap_warning", "missing_seqs": [...]} — seq skipped ahead
|
||||
{"type": "event_rejected", "detail": "..."} — one frame failed validation
|
||||
(seq echoed when parseable)
|
||||
{"type": "event_rejected", "seq": N, "detail": "..."} — stored-field rejected (bad kind)
|
||||
{"type": "flooded", "reason": "cap_exceeded"|"queue_overflow",
|
||||
"count": N} — sent before close(1008)
|
||||
|
||||
Keepalive: the server sends an opaque ping frame every `PING_INTERVAL_S` (the
|
||||
capture agent auto-pongs at the frame layer); a peer that is silent past
|
||||
`PONG_TIMEOUT_S` is assumed wedged, but the keepalive half only LOGS — the
|
||||
receiver half owns disconnect detection (single-box pilot: TCP EOF is
|
||||
reliable; an aggressive pong-watchdog would false-positive on loaded boxes).
|
||||
|
||||
Flood control (GRILL G-3, BINDING — silent drop-oldest is FORBIDDEN):
|
||||
* per-connection inbound queue bounded at `INBOUND_QUEUE_MAX` frames; on
|
||||
overflow → close code 1008 (policy violation) + trace marked
|
||||
INCOMPLETE_FLOODED via `TraceIntegrityMap`.
|
||||
* total events for the (learner, task) exceeding
|
||||
`Settings.telemetry_max_events_per_task` → same 1008 + INCOMPLETE_FLOODED.
|
||||
`INCOMPLETE_FLOODED` is an integrity signal Proctor/Phase-3 grader read via
|
||||
`TraceIntegrityMap.is_incomplete()` (the G-4 gate): a flooded trace can
|
||||
never yield a credential.
|
||||
|
||||
Boundary (D-027): telemetry/ never imports agents/ or api/. This module
|
||||
imports only `fastapi.WebSocket` for the socket type (a protocol surface, not
|
||||
a DI framework); the session engine below depends only on the TraceStore
|
||||
protocol + Settings, and api/telemetry.py injects both through plain
|
||||
parameters.
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import contextlib
|
||||
import json
|
||||
import logging
|
||||
from datetime import datetime
|
||||
from typing import Any, Final
|
||||
|
||||
from fastapi import WebSocket, WebSocketDisconnect
|
||||
from pydantic import BaseModel, ConfigDict, Field, ValidationError
|
||||
|
||||
from ..config import Settings
|
||||
from .models import TelemetryEvent
|
||||
from .store import TraceStore
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
#: WebSocket close code 1008 — policy violation (RFC 6455 §7.4.1).
|
||||
WS_CLOSE_POLICY_VIOLATION: Final = 1008
|
||||
|
||||
#: Bounded inbound queue depth per connection (G-3). Sized for burst-tolerance
|
||||
#: well above the capture agent's emission rate; overflow is a flood signal,
|
||||
#: not a backpressure knob.
|
||||
INBOUND_QUEUE_MAX: Final = 256
|
||||
|
||||
PING_INTERVAL_S: Final = 20.0
|
||||
|
||||
|
||||
class InboundEventFrame(BaseModel):
|
||||
"""Client → server event frame (one TelemetryEvent minus URL-owned ids).
|
||||
|
||||
`extra="forbid"`: learner_id/task_id arriving in the frame body is a
|
||||
contract violation — identity comes from the query params only, so a
|
||||
replayed frame can never lie about which trace it belongs to.
|
||||
"""
|
||||
|
||||
model_config = ConfigDict(extra="forbid")
|
||||
|
||||
seq: int = Field(ge=0)
|
||||
kind: str = Field(min_length=1)
|
||||
payload: dict[str, Any] = Field(default_factory=dict)
|
||||
ts: datetime
|
||||
sandbox_id: str = ""
|
||||
|
||||
|
||||
class TraceIntegrityMap:
|
||||
"""Integrity flags for traces that can never be graded (G-3/G-4).
|
||||
|
||||
Process-local and deliberately small: v0.3 runs ONE ai-service process per
|
||||
box, and the Phase-3 grader reads this flag through the same DI container
|
||||
— D-019-style in-memory registry precedent (the sandbox handle registry is
|
||||
the same shape). The flag is terminal within the process: a reconnect
|
||||
sending legal events does NOT clear it — the trace is already untrusted as
|
||||
grading input. Restarting ai-service resets flags; grading runs against a
|
||||
live service, and the SQLite trace rows themselves are durable.
|
||||
|
||||
All methods are sync: mutation is a dict write, reads are dict lookups —
|
||||
no await needed, so callers from any layer (API handlers, the grader)
|
||||
don't inherit an async surface for a nanosecond operation.
|
||||
"""
|
||||
|
||||
def __init__(self) -> None:
|
||||
# (learner_id, task_id) -> machine-readable reason (INCOMPLETE_FLOODED)
|
||||
self._flags: dict[tuple[str, str], str] = {}
|
||||
|
||||
def mark(self, learner_id: str, task_id: str, reason: str) -> None:
|
||||
"""Set an integrity flag. Presence of the flag is the signal; the
|
||||
reason is informational (last write wins)."""
|
||||
self._flags[(learner_id, task_id)] = reason
|
||||
|
||||
def clear(self, learner_id: str, task_id: str) -> None:
|
||||
"""Test seam: reset a flag (production ingest never clears)."""
|
||||
self._flags.pop((learner_id, task_id), None)
|
||||
|
||||
def is_incomplete(self, learner_id: str, task_id: str) -> bool:
|
||||
"""True when the trace carries ANY terminal integrity flag."""
|
||||
return (learner_id, task_id) in self._flags
|
||||
|
||||
def reason(self, learner_id: str, task_id: str) -> str | None:
|
||||
"""The flag's reason (INCOMPLETE_FLOODED), or None when unflagged."""
|
||||
return self._flags.get((learner_id, task_id))
|
||||
|
||||
|
||||
class IngestSession:
|
||||
"""One WebSocket ingest connection: receive → queue → drain → store.
|
||||
|
||||
Two tasks per connection:
|
||||
* `_receiver` — reads frames, validates shape, enqueues (bounded queue,
|
||||
G-3). Receives never block on SQLite.
|
||||
* `_drainer` — pops frames in arrival order, appends via TraceStore
|
||||
(idempotent on (learner,task,seq)), emits gap warnings, enforces the
|
||||
per-trace event cap.
|
||||
Either task detecting a flood closes the WS with 1008 and marks the trace
|
||||
INCOMPLETE_FLOODED. The events queue carries `None` as the client-
|
||||
disconnect sentinel.
|
||||
|
||||
- `telemetry_max_events_per_task` is consulted at connect and re-checked
|
||||
per append against the DURABLE row count (cap compares against stored
|
||||
events, so a skipped-ahead seq cannot burn budget that was never sent).
|
||||
Durable count via `TraceStore.count()` (COUNT(*)) — a single aggregate
|
||||
per append, never materializing trace rows (the pre-P7 code read
|
||||
`len(get_trace(...))` which was O(trace) per event / O(n²) per session).
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
websocket: WebSocket,
|
||||
store: TraceStore,
|
||||
integrity: TraceIntegrityMap,
|
||||
settings: Settings,
|
||||
learner_id: str,
|
||||
task_id: str,
|
||||
sandbox_id: str,
|
||||
) -> None:
|
||||
self._ws = websocket
|
||||
self._store = store
|
||||
self._integrity = integrity
|
||||
# Snapshot of the one setting ingest consults: read once at connect so
|
||||
# a hot-reloaded Settings object mid-session can't move the cap.
|
||||
self._max_events = settings.telemetry_max_events_per_task
|
||||
self.learner_id = learner_id
|
||||
self.task_id = task_id
|
||||
self.sandbox_id = sandbox_id
|
||||
|
||||
self._queue: asyncio.Queue[InboundEventFrame | None] = asyncio.Queue(
|
||||
maxsize=INBOUND_QUEUE_MAX
|
||||
)
|
||||
self._seen: set[int] = set()
|
||||
self._next_expected: int | None = None # in-connection monotonic hint
|
||||
self._received = 0
|
||||
self._stored = 0
|
||||
self._deduped = 0
|
||||
self._rejected = 0
|
||||
self._flooded = False
|
||||
self._flood_reason = ""
|
||||
|
||||
# -- receive half ----------------------------------------------------------
|
||||
|
||||
async def run(self) -> None:
|
||||
"""Accept, run receiver+drainer, close cleanly. Owns the WS lifecycle."""
|
||||
await self._ws.accept()
|
||||
logger.info(
|
||||
"telemetry ingest connected: %s/%s sandbox=%s",
|
||||
self.learner_id,
|
||||
self.task_id,
|
||||
self.sandbox_id or "(none)",
|
||||
)
|
||||
pinger = asyncio.create_task(self._keepalive())
|
||||
receiver = asyncio.create_task(self._receiver())
|
||||
drainer = asyncio.create_task(self._drainer())
|
||||
# First terminal outcome shuts the session down: client disconnect
|
||||
# (receiver ends) → drainer flushes; drainer ended (clean close after
|
||||
# flush or a 1008 flood close) → receiver must not linger.
|
||||
pending: set[asyncio.Task[None]] = {receiver, drainer}
|
||||
try:
|
||||
done, pending = await asyncio.wait(
|
||||
pending, return_when=asyncio.FIRST_COMPLETED
|
||||
)
|
||||
if receiver in done and drainer in pending:
|
||||
try:
|
||||
await drainer # final flush → sends ack_total, close 1000
|
||||
finally:
|
||||
pending.discard(drainer)
|
||||
finally:
|
||||
for task in (pinger, *pending):
|
||||
task.cancel()
|
||||
with contextlib.suppress(asyncio.CancelledError):
|
||||
await task
|
||||
|
||||
async def _receiver(self) -> None:
|
||||
"""Read frames; parse+enqueue. Overflow → flood shutdown (G-3).
|
||||
|
||||
RuntimeError from receive_text is benign here: it fires when the
|
||||
socket was closed by the drainer (1008 flood close) while this task
|
||||
was parked in receive — a terminal condition, not a bug.
|
||||
"""
|
||||
try:
|
||||
while True:
|
||||
raw = await self._ws.receive_text()
|
||||
frame = self._parse(raw)
|
||||
if frame is None:
|
||||
# Rejected frame — keep the connection open; the producer
|
||||
# gets an event_rejected status frame so a malformed batch
|
||||
# is visible (and its seq is never stored). Yield so the
|
||||
# status frame flushes before we block on the next receive.
|
||||
await self._reject_frame(raw)
|
||||
await asyncio.sleep(0)
|
||||
continue
|
||||
self._received += 1
|
||||
try:
|
||||
self._queue.put_nowait(frame)
|
||||
except asyncio.QueueFull:
|
||||
# Bounded queue — overflow is a flood, never drop-oldest.
|
||||
# _trigger_flood closes the socket; fall through to the
|
||||
# tail so the disconnect sentinel is still enqueued — the
|
||||
# drainer is never left parked on an empty queue after a
|
||||
# flood (P7 review: the pre-fix code `return`ed from the
|
||||
# QueueFull branch WITHOUT the sentinel, leaking the
|
||||
# session task set — one per flooded trace).
|
||||
await self._trigger_flood("queue_overflow")
|
||||
return
|
||||
except WebSocketDisconnect:
|
||||
pass
|
||||
except RuntimeError:
|
||||
logger.debug(
|
||||
"ingest receiver: socket already closed (flood path) %s/%s",
|
||||
self.learner_id,
|
||||
self.task_id,
|
||||
)
|
||||
# Client gone (clean close, drop, or flood close): sentinel unblocks
|
||||
# the drainer for a final flush. put_nowait can only fail under flood,
|
||||
# which already terminated the session.
|
||||
with contextlib.suppress(asyncio.QueueFull):
|
||||
self._queue.put_nowait(None)
|
||||
|
||||
def _parse(self, raw: str) -> InboundEventFrame | None:
|
||||
"""Validate one frame; None means malformed (caller rejects it)."""
|
||||
try:
|
||||
return InboundEventFrame.model_validate_json(raw)
|
||||
except ValidationError:
|
||||
return None
|
||||
|
||||
async def _reject_frame(self, raw: str) -> None:
|
||||
"""Malformed envelope: log + event_rejected status frame (never stored)."""
|
||||
self._rejected += 1
|
||||
detail = "invalid event frame"
|
||||
try:
|
||||
InboundEventFrame.model_validate_json(raw)
|
||||
except ValidationError as exc:
|
||||
detail = exc.errors()[0].get("msg", "validation error")
|
||||
logger.warning(
|
||||
"telemetry frame rejected: %s/%s: %s", self.learner_id, self.task_id, detail
|
||||
)
|
||||
seq: int | None = None
|
||||
with contextlib.suppress(Exception):
|
||||
seq = int(json.loads(raw).get("seq")) # best-effort echo for the producer
|
||||
payload: dict[str, Any] = {"type": "event_rejected", "detail": detail}
|
||||
if seq is not None:
|
||||
payload["seq"] = seq
|
||||
await self._send_json(payload)
|
||||
|
||||
# -- drain half --------------------------------------------------------------
|
||||
|
||||
async def _drainer(self) -> None:
|
||||
"""Pop queued frames, append to the store, then close 1000 + summary."""
|
||||
while True:
|
||||
frame = await self._queue.get()
|
||||
if frame is None: # disconnect sentinel → flush complete
|
||||
await self._send_json(
|
||||
{
|
||||
"type": "ack_total",
|
||||
"count": self._stored,
|
||||
"deduped": self._deduped,
|
||||
"rejected": self._rejected,
|
||||
}
|
||||
)
|
||||
with contextlib.suppress(RuntimeError, WebSocketDisconnect):
|
||||
await self._ws.close(code=1000)
|
||||
return
|
||||
await self._append(frame)
|
||||
|
||||
async def _append(self, frame: InboundEventFrame) -> None:
|
||||
# Precedence: a trace already flagged INCOMPLETE_FLOODED is terminal —
|
||||
# the connection that triggered it is being torn down, and any stray
|
||||
# queued frames must not resurrect the trace's intake.
|
||||
if self._integrity.is_incomplete(self.learner_id, self.task_id):
|
||||
await self._trigger_flood("already_flagged")
|
||||
return
|
||||
|
||||
# Per-trace cap (G-3): checked against the DURABLE row count so a
|
||||
# reconnect resumes the budget instead of resetting it, and a
|
||||
# skipped-ahead seq cannot burn budget that was never sent.
|
||||
if self._flood_breached():
|
||||
await self._trigger_flood("cap_exceeded")
|
||||
return
|
||||
|
||||
# TelemetryEvent's @validates hooks fire on CONSTRUCTION (setattr), so
|
||||
# the try must wrap building the model too — an unknown kind raises
|
||||
# before `store.append` is ever reached.
|
||||
event: TelemetryEvent
|
||||
before = self._store.latest_seq(self.learner_id, self.task_id)
|
||||
try:
|
||||
event = TelemetryEvent(
|
||||
learner_id=self.learner_id,
|
||||
task_id=self.task_id,
|
||||
seq=frame.seq,
|
||||
kind=frame.kind,
|
||||
payload=frame.payload,
|
||||
ts=frame.ts,
|
||||
sandbox_id=frame.sandbox_id or self.sandbox_id,
|
||||
)
|
||||
self._store.append(event)
|
||||
except ValueError as exc: # unknown kind / invalid field
|
||||
self._rejected += 1
|
||||
await self._send_json(
|
||||
{"type": "event_rejected", "seq": frame.seq, "detail": str(exc)}
|
||||
)
|
||||
return
|
||||
after = self._store.latest_seq(self.learner_id, self.task_id)
|
||||
|
||||
if after == before and frame.seq in self._seen:
|
||||
self._deduped += 1 # at-least-once retry; stored once (idempotent)
|
||||
else:
|
||||
self._stored += 1
|
||||
self._seen.add(frame.seq)
|
||||
# Seq-ack (D-045, REQ-5-007): advisory hint carrying the durable
|
||||
# latest_seq AFTER this append — the capture agent trims its spool to
|
||||
# seq > ack on receipt, bounding the replay margin to the in-flight
|
||||
# window. Emitted on dedup'd appends too (a-6) so a replay flush
|
||||
# tightens the margin immediately. Gap detection stays authoritative
|
||||
# (_check_gap below); G-3 flood semantics untouched.
|
||||
await self._send_json({"type": "seq_ack", "seq": after})
|
||||
await self._check_gap(frame.seq)
|
||||
# SQLite appends are sync and fast; on a burst the drainer can hold
|
||||
# the loop between receives. Yield so the WS writer flushes the close
|
||||
# and the pinger/interleave stay live under the eventlet-free portal.
|
||||
await asyncio.sleep(0)
|
||||
|
||||
def _flood_breached(self) -> bool:
|
||||
"""True when this append would exceed the per-trace event budget."""
|
||||
# Durable count (NOT latest_seq+1 — a skipped-ahead seq must not burn
|
||||
# un-sent events' budget) via COUNT(*): never materialize the trace
|
||||
# per append (P7 review — the old len(get_trace(...)) built every row
|
||||
# object per event, O(trace) per append / O(n²) per session).
|
||||
durable = self._store.count(learner_id=self.learner_id, task_id=self.task_id)
|
||||
return durable >= self._max_events
|
||||
|
||||
async def _check_gap(self, incoming_seq: int) -> None:
|
||||
"""Seq skipped ahead → log + per-connection gap_warning status frame."""
|
||||
if self._next_expected is not None and incoming_seq > self._next_expected:
|
||||
missing = list(range(self._next_expected, incoming_seq))
|
||||
logger.warning(
|
||||
"telemetry gap: %s/%s missing seqs %s (arrived seq=%d)",
|
||||
self.learner_id,
|
||||
self.task_id,
|
||||
missing,
|
||||
incoming_seq,
|
||||
)
|
||||
await self._send_json({"type": "gap_warning", "missing_seqs": missing})
|
||||
if self._next_expected is None or incoming_seq >= self._next_expected:
|
||||
self._next_expected = incoming_seq + 1
|
||||
|
||||
# -- flood + keepalive ------------------------------------------------------
|
||||
|
||||
async def _trigger_flood(self, reason: str) -> None:
|
||||
"""G-3: 1008 close + INCOMPLETE_FLOODED mark. Exactly once."""
|
||||
if self._flooded:
|
||||
return
|
||||
self._flooded = True
|
||||
self._flood_reason = reason
|
||||
self._integrity.mark(self.learner_id, self.task_id, "INCOMPLETE_FLOODED")
|
||||
logger.warning(
|
||||
"telemetry flood: %s/%s reason=%s — closing 1008, trace marked "
|
||||
"INCOMPLETE_FLOODED (G-3; Proctor/grade gate will refuse it)",
|
||||
self.learner_id,
|
||||
self.task_id,
|
||||
reason,
|
||||
)
|
||||
await self._send_json(
|
||||
{"type": "flooded", "reason": reason, "count": self._received}
|
||||
)
|
||||
with contextlib.suppress(RuntimeError, WebSocketDisconnect):
|
||||
await self._ws.close(
|
||||
code=WS_CLOSE_POLICY_VIOLATION,
|
||||
reason=f"telemetry flood control (G-3): {reason}",
|
||||
)
|
||||
|
||||
async def _keepalive(self) -> None:
|
||||
"""Protocol-level ping on an interval (agent auto-pongs at frame level).
|
||||
|
||||
A send failure means the socket is already gone — the receiver half
|
||||
independently surfaces the disconnect; we just stop pinging.
|
||||
"""
|
||||
while True:
|
||||
await asyncio.sleep(PING_INTERVAL_S)
|
||||
try:
|
||||
await self._ws.send_bytes(b"\x89ping-nextcraft")
|
||||
except (RuntimeError, WebSocketDisconnect):
|
||||
return
|
||||
|
||||
async def _send_json(self, payload: dict[str, Any]) -> None:
|
||||
"""Best-effort status frame; the socket may already be gone."""
|
||||
with contextlib.suppress(RuntimeError, WebSocketDisconnect):
|
||||
await self._ws.send_json(payload)
|
||||
|
||||
|
||||
async def telemetry_ingest_endpoint(
|
||||
websocket: WebSocket,
|
||||
learner_id: str,
|
||||
task_id: str,
|
||||
store: TraceStore,
|
||||
integrity: TraceIntegrityMap,
|
||||
settings: Settings,
|
||||
sandbox_id: str = "",
|
||||
) -> None:
|
||||
"""Engine entry: build the session and run it. api/telemetry.py calls this
|
||||
with query params + app.state services already resolved — this signature
|
||||
is deliberately Depends-free (telemetry/ never knows FastAPI DI exists).
|
||||
"""
|
||||
session = IngestSession(
|
||||
websocket=websocket,
|
||||
store=store,
|
||||
integrity=integrity,
|
||||
settings=settings,
|
||||
learner_id=learner_id,
|
||||
task_id=task_id,
|
||||
sandbox_id=sandbox_id,
|
||||
)
|
||||
await session.run()
|
||||
@@ -1,108 +0,0 @@
|
||||
"""Telemetry event record — the row the trace store persists (REQ-3-003, D-027).
|
||||
|
||||
One model serves both as the JSON payload sent by producers and as the SQLite
|
||||
row schema. `payload` is stored as a JSON column (native JSONB on Postgres —
|
||||
no migration-time shape change, D-027).
|
||||
|
||||
Field contract (consumed by the trace store and the grader):
|
||||
learner_id — non-empty learner identifier.
|
||||
task_id — non-empty task/session identifier; trace identity is the
|
||||
(learner_id, task_id) pair.
|
||||
seq — sequence number per trace, >= 0. Monotonicity per
|
||||
(learner, task) is enforced by the store (Task 2-1-02);
|
||||
this model only rejects negative seqs.
|
||||
kind — event discriminator: command | file_diff | run_result |
|
||||
test_result | activity | stdin | stdout.
|
||||
payload — free-form JSON detail blob.
|
||||
ts — envelope timestamp (UTC); monotonicity enforced at ingest.
|
||||
sandbox_id — originating sandbox ("" for non-sandbox sources).
|
||||
|
||||
Boundary (D-027): telemetry/ never imports agents/ or api/.
|
||||
"""
|
||||
|
||||
from datetime import datetime
|
||||
from typing import Any, Literal
|
||||
|
||||
from pydantic import BaseModel, ConfigDict, Field
|
||||
from sqlalchemy import JSON, Index, String
|
||||
from sqlalchemy.orm import validates
|
||||
from sqlmodel import Field as SQLField
|
||||
from sqlmodel import SQLModel
|
||||
|
||||
EventKind = Literal[
|
||||
"command",
|
||||
"file_diff",
|
||||
"run_result",
|
||||
"test_result",
|
||||
"activity",
|
||||
"stdin",
|
||||
"stdout",
|
||||
]
|
||||
_EVENT_KINDS: frozenset[str] = frozenset(EventKind.__args__)
|
||||
|
||||
|
||||
class TelemetryEvent(SQLModel, table=True):
|
||||
"""A single durable telemetry event; (learner_id, task_id, seq) is PK.
|
||||
|
||||
Constraint enforcement uses SQLAlchemy `@validates` hooks: sqlmodel
|
||||
0.0.42's metaclass drops pydantic `Field(ge=...)`/`field_validator`
|
||||
constraints for table models (the decorators register but never make it
|
||||
into the core schema), while `@validates` fires on every attribute set —
|
||||
construction included — and raises ValueError on violation. seq >= 0 plus
|
||||
a VARCHAR kind column keep the DB shape Postgres-ready (D-027).
|
||||
"""
|
||||
|
||||
__tablename__ = "telemetry_event"
|
||||
# PK columns already produce a unique index; this secondary index covers
|
||||
# trace reads ordered by seq without depending on the PK column order
|
||||
# (Postgres migration target D-027).
|
||||
__table_args__ = (Index("ix_telemetry_event_trace", "learner_id", "task_id"),)
|
||||
|
||||
learner_id: str = SQLField(primary_key=True)
|
||||
task_id: str = SQLField(primary_key=True)
|
||||
seq: int = SQLField(primary_key=True)
|
||||
# Bare Literal annotations crash sqlmodel<=0.0.42's column inference
|
||||
# (issubclass(TypeAlias, Enum)); an explicit sa_type + the validates hook
|
||||
# below gives the same contract: VARCHAR column, Literal-rejected values.
|
||||
kind: EventKind = SQLField(sa_type=String)
|
||||
# JSON column: stored as TEXT on SQLite, native JSONB on Postgres (D-027).
|
||||
payload: dict[str, Any] = SQLField(default_factory=dict, sa_type=JSON)
|
||||
ts: datetime
|
||||
sandbox_id: str = SQLField(default="")
|
||||
|
||||
@validates("learner_id", "task_id")
|
||||
def _ids_non_empty(self, key: str, value: str) -> str:
|
||||
if not value:
|
||||
raise ValueError(f"{key} must be a non-empty identifier")
|
||||
return value
|
||||
|
||||
@validates("seq")
|
||||
def _seq_non_negative(self, key: str, value: int) -> int:
|
||||
if value < 0:
|
||||
raise ValueError("seq must be >= 0 (monotonicity is the store's job)")
|
||||
return value
|
||||
|
||||
@validates("kind")
|
||||
def _kind_is_known(self, key: str, value: str) -> str:
|
||||
if value not in _EVENT_KINDS:
|
||||
raise ValueError(f"unknown event kind: {value!r}")
|
||||
return value
|
||||
|
||||
|
||||
class TraceSpan(BaseModel):
|
||||
"""Derived view: the ordered event trace for one (learner_id, task_id).
|
||||
|
||||
NOT a table — materialized by the store from persisted TelemetryEvents
|
||||
(grader/Lab consume this shape; replay order is the seq column).
|
||||
"""
|
||||
|
||||
model_config = ConfigDict(frozen=True)
|
||||
|
||||
learner_id: str = Field(min_length=1)
|
||||
task_id: str = Field(min_length=1)
|
||||
events: tuple[TelemetryEvent, ...] = ()
|
||||
|
||||
@property
|
||||
def latest_seq(self) -> int:
|
||||
"""Highest seq in the span; -1 when empty (store convention)."""
|
||||
return self.events[-1].seq if self.events else -1
|
||||
@@ -1,218 +0,0 @@
|
||||
"""TraceStore — telemetry persistence protocol + SQLite implementation (REQ-3-003, D-027).
|
||||
|
||||
Postgres-migration-ready (D-027): the protocol is the only surface the API /
|
||||
grader layers touch; swapping SQLiteTraceStore for a Postgres-backed
|
||||
implementation must not change call sites. The `telemetry_event` table uses
|
||||
only portable column types (str / int / datetime / JSON), so the same SQLModel
|
||||
schema stands up unchanged on Postgres.
|
||||
|
||||
Ingest is at-least-once: duplicates carry the same (learner_id, task_id, seq)
|
||||
idempotency key, so `append` with a triplet that is already stored is a no-op.
|
||||
The pair (learner_id, task_id) identifies a trace; `seq` numbers events in it
|
||||
starting at 0.
|
||||
|
||||
Concurrency (a-3): the engine enables WAL + synchronous=NORMAL and a busy
|
||||
timeout at connection time, so the ingest writer and grader readers do not hit
|
||||
`database is locked` on the single-box pilot.
|
||||
|
||||
Boundary: `telemetry/` never imports `agents/` / `api/` and has no FastAPI
|
||||
dependency.
|
||||
"""
|
||||
|
||||
import logging
|
||||
import sqlite3
|
||||
from collections.abc import Iterator
|
||||
from contextlib import contextmanager
|
||||
from datetime import UTC, datetime
|
||||
from pathlib import Path
|
||||
from typing import Any, Protocol
|
||||
|
||||
import sqlalchemy as sa
|
||||
from sqlmodel import Session, SQLModel, create_engine, select
|
||||
|
||||
from ..config import Settings
|
||||
from .models import TelemetryEvent
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class TraceStore(Protocol):
|
||||
"""Persistence contract for ordered per-learner task trace streams.
|
||||
|
||||
Implemented by SQLiteTraceStore (v0.3, D-027); a Postgres implementation
|
||||
must satisfy the same surface.
|
||||
"""
|
||||
|
||||
def append(self, event: TelemetryEvent) -> None:
|
||||
"""Store one event. IDEMPOTENT on (learner_id, task_id, seq):
|
||||
|
||||
at-least-once ingest retries with the same triplet are deduped
|
||||
(stored once), not rejected. Later events must not overwrite an
|
||||
existing row.
|
||||
"""
|
||||
...
|
||||
|
||||
def get_trace(self, learner_id: str, task_id: str) -> list[TelemetryEvent]:
|
||||
"""All stored events for the trace, ordered by seq ascending.
|
||||
|
||||
Detached from any DB session — safe to pass across layers. Empty list
|
||||
when the trace has no events.
|
||||
"""
|
||||
...
|
||||
|
||||
def gaps(self, learner_id: str, task_id: str) -> list[int]:
|
||||
"""Missing seqs in 0..latest for the trace ([0,2,3] stored -> [1])."""
|
||||
...
|
||||
|
||||
def latest_seq(self, learner_id: str, task_id: str) -> int:
|
||||
"""Highest stored seq for the trace; -1 when no events exist."""
|
||||
...
|
||||
|
||||
def count(self, learner_id: str, task_id: str) -> int:
|
||||
"""Number of stored events for the trace (COUNT(*), never
|
||||
materializes rows — the ingest cap consults this per append, so
|
||||
an O(trace) implementation would make ingest O(n²) per session).
|
||||
"""
|
||||
...
|
||||
|
||||
def list_tasks(self, learner_id: str) -> list[str]:
|
||||
"""Distinct task_ids with at least one event for the learner."""
|
||||
...
|
||||
|
||||
def close(self) -> None:
|
||||
"""Release DB connections. Store must not be used after close."""
|
||||
...
|
||||
|
||||
|
||||
def _sqlite_connect(dbapi_connection: sqlite3.Connection, _: object) -> None:
|
||||
"""Per-connection pragma setup (a-3).
|
||||
|
||||
journal_mode=WAL — readers never block the single writer.
|
||||
synchronous=NORMAL — safe in WAL mode, avoids full fsync-per-commit.
|
||||
busy_timeout=5000 — retry briefly under contention instead of
|
||||
`OperationalError: database is locked`.
|
||||
"""
|
||||
cursor = dbapi_connection.cursor()
|
||||
cursor.execute("PRAGMA journal_mode=WAL")
|
||||
cursor.execute("PRAGMA synchronous=NORMAL")
|
||||
cursor.execute("PRAGMA busy_timeout=5000")
|
||||
cursor.close()
|
||||
|
||||
|
||||
def _as_utc(ts: datetime) -> datetime:
|
||||
"""Timestamps travel as naive datetime on SQLite, tz-aware elsewhere.
|
||||
|
||||
SQLite (via SQLModel) drops tzinfo; Postgres TIMESTAMP WITH TIME ZONE
|
||||
keeps it. Normalizing on the read/write boundary makes the store's
|
||||
contract tz-aware UTC regardless of the backend (D-027).
|
||||
"""
|
||||
if ts.tzinfo is None:
|
||||
return ts.replace(tzinfo=UTC) # naive UTC read-side: label as UTC
|
||||
return ts.astimezone(UTC)
|
||||
|
||||
|
||||
class SQLiteTraceStore:
|
||||
"""SQLite-backed TraceStore (SQLModel). First real persistence (D-027)."""
|
||||
|
||||
def __init__(self, db_path: Path | None = None) -> None:
|
||||
self._db_path: Path = db_path if db_path is not None else Settings().db_path
|
||||
self._engine = create_engine(f"sqlite:///{self._db_path}")
|
||||
sa.event.listen(self._engine, "connect", _sqlite_connect)
|
||||
SQLModel.metadata.create_all(self._engine)
|
||||
|
||||
@contextmanager
|
||||
def _session(self) -> Iterator[Session]:
|
||||
# expire_on_commit=False: ORM objects returned from `append`'s
|
||||
# IntegrityError path stay usable without a refresh round-trip.
|
||||
with Session(self._engine, expire_on_commit=False) as session:
|
||||
yield session
|
||||
|
||||
def append(self, event: TelemetryEvent) -> None:
|
||||
# INSERT-if-absent via PK: sqlite3 raises IntegrityError on a
|
||||
# duplicate (learner_id, task_id, seq); swallow it — the row is
|
||||
# already stored, which is the dedup contract for at-least-once
|
||||
# ingest. `session.merge` would upsert instead; wrong semantics here.
|
||||
with self._session() as session:
|
||||
try:
|
||||
session.add(event)
|
||||
session.commit()
|
||||
except sa.exc.IntegrityError:
|
||||
session.rollback()
|
||||
logger.debug(
|
||||
"trace event dedup: %s/%s seq=%d already stored",
|
||||
event.learner_id,
|
||||
event.task_id,
|
||||
event.seq,
|
||||
)
|
||||
|
||||
def get_trace(self, learner_id: str, task_id: str) -> list[TelemetryEvent]:
|
||||
with self._session() as session:
|
||||
stmt = (
|
||||
select(TelemetryEvent)
|
||||
.where(TelemetryEvent.learner_id == learner_id)
|
||||
.where(TelemetryEvent.task_id == task_id)
|
||||
.order_by(TelemetryEvent.seq)
|
||||
)
|
||||
results = session.exec(stmt).all()
|
||||
# Detach from the session: callers must not depend on open-session
|
||||
# ORM magic (lazy loads fail once the session is closed).
|
||||
for row in results:
|
||||
row.ts = _as_utc(row.ts)
|
||||
session.expunge(row)
|
||||
return list(results)
|
||||
|
||||
def _stored_seqs(self, learner_id: str, task_id: str) -> list[int]:
|
||||
with self._session() as session:
|
||||
stmt = (
|
||||
select(TelemetryEvent.seq)
|
||||
.where(TelemetryEvent.learner_id == learner_id)
|
||||
.where(TelemetryEvent.task_id == task_id)
|
||||
.order_by(TelemetryEvent.seq)
|
||||
)
|
||||
# sqlmodel scalar select: rows are plain ints, not 1-tuples.
|
||||
return [int(seq) for seq in session.exec(stmt).all()]
|
||||
|
||||
def gaps(self, learner_id: str, task_id: str) -> list[int]:
|
||||
seqs = self._stored_seqs(learner_id, task_id)
|
||||
if not seqs:
|
||||
return []
|
||||
present = set(seqs)
|
||||
# seq numbering starts at 0; a gap is any seq in 0..latest not stored.
|
||||
return [seq for seq in range(seqs[-1] + 1) if seq not in present]
|
||||
|
||||
def latest_seq(self, learner_id: str, task_id: str) -> int:
|
||||
with self._session() as session:
|
||||
stmt = (
|
||||
select(sa.func.max(TelemetryEvent.seq))
|
||||
.where(TelemetryEvent.learner_id == learner_id)
|
||||
.where(TelemetryEvent.task_id == task_id)
|
||||
)
|
||||
latest: Any = session.exec(stmt).one()
|
||||
return -1 if latest is None else int(latest)
|
||||
|
||||
def count(self, learner_id: str, task_id: str) -> int:
|
||||
# COUNT(*) at the DB — no row materialization. The ingest flood cap
|
||||
# calls this per append (telemetry/ingest._flood_breached); the
|
||||
# docstring-free body keeps it obvious what the query shape is.
|
||||
with self._session() as session:
|
||||
stmt = (
|
||||
select(sa.func.count(TelemetryEvent.seq))
|
||||
.where(TelemetryEvent.learner_id == learner_id)
|
||||
.where(TelemetryEvent.task_id == task_id)
|
||||
)
|
||||
total: Any = session.exec(stmt).one()
|
||||
return int(total or 0)
|
||||
|
||||
def list_tasks(self, learner_id: str) -> list[str]:
|
||||
with self._session() as session:
|
||||
stmt = (
|
||||
select(TelemetryEvent.task_id)
|
||||
.where(TelemetryEvent.learner_id == learner_id)
|
||||
.distinct()
|
||||
.order_by(TelemetryEvent.task_id)
|
||||
)
|
||||
# sqlmodel scalar select: rows are plain strs, not 1-tuples.
|
||||
return [str(task_id) for task_id in session.exec(stmt).all()]
|
||||
|
||||
def close(self) -> None:
|
||||
self._engine.dispose()
|
||||
@@ -1,31 +0,0 @@
|
||||
"""Per-learner variant task generation — templates, generator, VariantStore (REQ-3-005).
|
||||
|
||||
Boundary rule (D-027): variants/ is an engine module — it never imports
|
||||
api/; its ONLY agents/ dependency is the module-direct
|
||||
agents.structured import in generator.py (the sanctioned shared D-020
|
||||
structured defense, same exception as grading/engine.py). api/ composes
|
||||
the generator and store via DI; store.py imports config only.
|
||||
|
||||
CO-ORDINATION NOTE (ADD, don't REMOVE — same convention as grading/):
|
||||
This __init__.py is a minimal placeholder created by the VariantStore
|
||||
task (4-1-02). The templates task (4-1-01) owns this file's final shape
|
||||
— when templates.py lands, ADD its exports alongside these; do not
|
||||
remove the store exports below.
|
||||
|
||||
Wave status: store.py (VariantRecord, VariantStore, SQLiteVariantStore)
|
||||
landed in Wave 1 (task 4-1-02); templates.py is Wave 1 task 4-1-01;
|
||||
generator.py is Wave 2 (4-2-01).
|
||||
"""
|
||||
|
||||
from .store import SQLiteVariantStore, VariantRecord, VariantStore
|
||||
from .templates import TEMPLATES, TaskTemplate, get_template, template_for_competency
|
||||
|
||||
__all__ = [
|
||||
"SQLiteVariantStore",
|
||||
"TEMPLATES",
|
||||
"TaskTemplate",
|
||||
"VariantRecord",
|
||||
"VariantStore",
|
||||
"get_template",
|
||||
"template_for_competency",
|
||||
]
|
||||
@@ -1,138 +0,0 @@
|
||||
"""Seeded per-learner variant generator (D-029, REQ-3-005).
|
||||
|
||||
Contract (binding, from GRILL + PLAN Must-Haves):
|
||||
- REPRODUCIBLE: seed = sha256(template_id|learner_id|milestone); the same
|
||||
(template, learner) re-derives the same seed, params, task_id — and the
|
||||
second generate() call is a cache hit with NO LLM call.
|
||||
- DISTINCT: different learners on the same template draw different params
|
||||
(the sampler is seeded per-learner) and receive distinct statements.
|
||||
- NEVER BLOCKS ON THE LLM: the deterministic skeleton render
|
||||
(`template.render(params)`) is a complete, valid statement; if the D-020
|
||||
LLM render fails after its bounded retry, the fallback is used — and
|
||||
because the fallback is exactly `template.render(seed-params)`, it is
|
||||
auditable from the persisted seed + params without a provenance column.
|
||||
- AUDITABLE: seed + params + statement persist via VariantStore
|
||||
(insert-only first-wins) — the proctoring cross-check path.
|
||||
- FAIR (a-5): slot draws change the scenario, never the difficulty; the
|
||||
template's rubric anchors bound the expected effort envelope, so every
|
||||
variant of one template is held to the same bar.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
from datetime import UTC, datetime
|
||||
from typing import TYPE_CHECKING
|
||||
|
||||
from pydantic import BaseModel, ConfigDict, Field
|
||||
|
||||
from ..agents.structured import StructuredOutputError, structured_completion
|
||||
from ..llm.types import Message
|
||||
from ..prompts.variant import VARIANT_SCHEMA_HINT, render_variant_prompt
|
||||
from .store import VariantRecord
|
||||
from .templates import TaskTemplate, get_template
|
||||
|
||||
if TYPE_CHECKING: # pragma: no cover
|
||||
from ..llm.base import LLMProvider
|
||||
from .store import VariantStore
|
||||
|
||||
MILESTONE = "v0.3"
|
||||
|
||||
|
||||
class RenderedVariant(BaseModel):
|
||||
"""D-20-validated LLM render output (statement only — files come from the template)."""
|
||||
|
||||
model_config = ConfigDict(extra="forbid")
|
||||
|
||||
statement: str = Field(min_length=20)
|
||||
|
||||
|
||||
class UnknownTemplateError(ValueError):
|
||||
"""Raised when generate() is asked for a template id not in the library."""
|
||||
|
||||
|
||||
def derive_seed(template_id: str, learner_id: str, milestone: str = MILESTONE) -> str:
|
||||
"""Reproducible per-(template, learner, milestone) seed (D-029)."""
|
||||
return hashlib.sha256(f"{template_id}|{learner_id}|{milestone}".encode()).hexdigest()
|
||||
|
||||
|
||||
def derive_task_id(seed: str) -> str:
|
||||
"""Deterministic grading/telemetry task key from the seed (16 hex chars)."""
|
||||
return f"task-{seed[:16]}"
|
||||
|
||||
|
||||
class VariantGenerator:
|
||||
"""Seeded instantiation over the template library. DI: store + provider."""
|
||||
|
||||
def __init__(self, store: VariantStore, provider: LLMProvider, model: str) -> None:
|
||||
self._store = store
|
||||
self._provider = provider
|
||||
self._model = model
|
||||
|
||||
async def generate(self, learner_id: str, template_id: str) -> VariantRecord:
|
||||
template = get_template(template_id)
|
||||
if template is None:
|
||||
raise UnknownTemplateError(f"no task template with id {template_id!r}")
|
||||
|
||||
# Cache: D-029 reproducibility — same (learner, template) is served
|
||||
# from the store with no LLM call.
|
||||
cached = self._store.get(learner_id, template_id)
|
||||
if cached is not None:
|
||||
return cached
|
||||
|
||||
seed_hex = derive_seed(template_id, learner_id)
|
||||
task_id = derive_task_id(seed_hex)
|
||||
params = template.sample_params(_seed_int(seed_hex))
|
||||
_validate_params(template, params)
|
||||
|
||||
statement = await self._render(template, params)
|
||||
record = VariantRecord(
|
||||
learner_id=learner_id,
|
||||
task_id=task_id,
|
||||
template_id=template_id,
|
||||
seed=seed_hex,
|
||||
params=dict(params),
|
||||
statement=statement,
|
||||
starter_files=dict(template.starter_files),
|
||||
environment=template.environment,
|
||||
test_command=template.test_command,
|
||||
created_at=datetime.now(UTC),
|
||||
)
|
||||
self._store.save(record)
|
||||
return record
|
||||
|
||||
async def _render(self, template: TaskTemplate, params: dict[str, str | int]) -> str:
|
||||
"""LLM render via D-020; deterministic fallback never blocks task work.
|
||||
|
||||
Provenance note: unlike grades, variants carry no `model` column —
|
||||
the deterministic fallback is exactly `template.render(params)`,
|
||||
re-derivable from the persisted seed + params, so a fallback render is
|
||||
auditable without storing provenance (the seed IS the provenance).
|
||||
"""
|
||||
messages: list[Message] = render_variant_prompt(template, params)
|
||||
try:
|
||||
rendered = await structured_completion(
|
||||
self._provider,
|
||||
messages,
|
||||
model=self._model,
|
||||
schema=RenderedVariant,
|
||||
schema_hint=VARIANT_SCHEMA_HINT,
|
||||
)
|
||||
except StructuredOutputError:
|
||||
# Deterministic fallback: the skeleton + seeded slots is already a
|
||||
# complete statement, re-derivable from the persisted seed.
|
||||
return template.render(params)
|
||||
return rendered.statement
|
||||
|
||||
|
||||
def _seed_int(seed_hex: str) -> int:
|
||||
"""Stable int for random.Random from the hex seed."""
|
||||
return int(seed_hex[:16], 16)
|
||||
|
||||
|
||||
def _validate_params(template: TaskTemplate, params: dict[str, str | int]) -> None:
|
||||
"""Defense in depth: every sampled value must be schema-valid (a-5)."""
|
||||
for slot in template.slots:
|
||||
value = params.get(slot.name)
|
||||
if value is None or not slot.validate_value(value):
|
||||
raise ValueError(f"sampled params invalid for slot {slot.name!r}: {value!r}")
|
||||
@@ -1,351 +0,0 @@
|
||||
"""VariantStore — variant persistence protocol + SQLite implementation (REQ-3-005, D-027).
|
||||
|
||||
Postgres-migration-ready (D-027): the protocol is the only surface the
|
||||
variant generator and API layers touch; swapping SQLiteVariantStore for a
|
||||
Postgres-backed implementation must not change call sites. The
|
||||
`variant_record` table uses only portable column types (str / JSON /
|
||||
datetime), so the same SQLModel schema stands up unchanged on Postgres.
|
||||
|
||||
Insert-only, NOT upsert: (learner_id, template_id) is the variant identity
|
||||
and the FIRST generation is authoritative — reproducibility (D-029) means
|
||||
the seed re-derives the same variant, so the generator's cache path serves
|
||||
`get` instead of saving again. `save` is a plain INSERT; a duplicate pair
|
||||
raises sqlalchemy.exc.IntegrityError to the caller (documented behavior).
|
||||
`task_id` is unique too — it is the grading/telemetry trace key, so a
|
||||
trace or grade can never silently join to a different variant. Both
|
||||
rejections are deliberate: overwriting a stored variant would swap a
|
||||
learner's graded task underneath its trace and grade (audit corruption).
|
||||
Contrast TraceStore.append (dedup-keep-first, swallowed — at-least-once
|
||||
ingest) and GradeStore.save (upsert-latest-wins — a regrade is
|
||||
latest-state); this store is the third contract of the D-027 family.
|
||||
|
||||
Concurrency (a-3): the store enables WAL + synchronous=NORMAL and a busy
|
||||
timeout at connection time, so a generation writer and API readers do not
|
||||
hit `database is locked` on the single-box pilot.
|
||||
|
||||
`created_at` contract: callers stamp UTC (datetime.now(UTC)); SQLite
|
||||
stores it naive and the read paths re-label it tz-aware UTC (same
|
||||
boundary normalization as TelemetryEvent.ts / GradeRecord.created_at, so
|
||||
the contract holds on any backend).
|
||||
|
||||
Boundary (D-027): `variants/` never imports `agents/` / `api/`; this
|
||||
module imports config only.
|
||||
"""
|
||||
|
||||
import logging
|
||||
import sqlite3
|
||||
from collections.abc import Iterator
|
||||
from contextlib import contextmanager
|
||||
from datetime import UTC, datetime
|
||||
from pathlib import Path
|
||||
from typing import Any, Protocol
|
||||
|
||||
import sqlalchemy as sa
|
||||
from sqlalchemy import JSON, Index, String, UniqueConstraint
|
||||
from sqlalchemy.orm import validates
|
||||
from sqlmodel import Field, Session, SQLModel, create_engine, select
|
||||
|
||||
from ..config import Settings
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class VariantRecord(SQLModel, table=True):
|
||||
"""A persisted task variant; (learner_id, template_id) is the PK — first wins.
|
||||
|
||||
Written once by the variant generator (Task 4-2-01), read by the API
|
||||
layer and proctoring cross-checks through the VariantStore protocol.
|
||||
Constraint enforcement mirrors TelemetryEvent / GradeRecord: sqlmodel
|
||||
0.0.42's metaclass drops pydantic constraints on table models, so
|
||||
SQLAlchemy `@validates` hooks enforce instead and the column types
|
||||
stay Postgres-ready (D-027).
|
||||
|
||||
Field contract:
|
||||
learner_id — non-empty learner identifier (same id space as
|
||||
traces and grades).
|
||||
template_id — non-empty task template identifier; variant
|
||||
identity is the (learner_id, template_id) pair —
|
||||
the pair the generator caches on (exactly one
|
||||
variant per learner per template).
|
||||
task_id — non-empty, GLOBALLY unique task identifier; the
|
||||
grading/telemetry trace key (the (learner_id,
|
||||
task_id) pair TraceStore / GradeStore key on),
|
||||
stamped at generation so a variant's trace and
|
||||
grade join back to it exactly once.
|
||||
seed — non-empty variant seed (D-029); derived from
|
||||
(template_id, learner_id, milestone) so the
|
||||
variant is reproducible and auditable.
|
||||
params — typed parameter-slot values the generator filled;
|
||||
JSON dict. An empty dict is legal (a slotless
|
||||
template).
|
||||
statement — non-empty rendered task statement shown to the
|
||||
learner (distinct per learner by construction,
|
||||
REQ-3-005).
|
||||
starter_files — workspace scaffold: filename -> file content;
|
||||
JSON dict. An empty dict is legal (no scaffold).
|
||||
created_at — UTC generation timestamp.
|
||||
"""
|
||||
|
||||
__tablename__ = "variant_record"
|
||||
# The composite PK covers (learner_id, template_id) point lookups; the
|
||||
# unique task_id covers get_by_task (the grading/telemetry join path);
|
||||
# the two secondary indexes cover list_for_learner / list_by_template
|
||||
# ordered by created_at without a sort step (Postgres target D-027).
|
||||
__table_args__ = (
|
||||
UniqueConstraint("task_id", name="uq_variant_record_task_id"),
|
||||
Index("ix_variant_record_learner_created", "learner_id", "created_at"),
|
||||
Index("ix_variant_record_template_created", "template_id", "created_at"),
|
||||
)
|
||||
|
||||
learner_id: str = Field(primary_key=True)
|
||||
template_id: str = Field(primary_key=True)
|
||||
task_id: str
|
||||
seed: str
|
||||
# JSON columns: stored as TEXT on SQLite, native JSONB on Postgres (D-027).
|
||||
params: dict[str, Any] = Field(default_factory=dict, sa_type=JSON)
|
||||
statement: str
|
||||
starter_files: dict[str, Any] = Field(default_factory=dict, sa_type=JSON)
|
||||
#: REQ-5-005 (D-044): the environment kind rides the record to the API
|
||||
#: and TS client; defaults keep pre-v0.5 rows 'build'.
|
||||
environment: str = Field(default="build", sa_type=String)
|
||||
#: The variant's real test command (was a dead template field — v0.5
|
||||
#: surfaces it so the Run/Test buttons stop hardcoding pytest).
|
||||
test_command: str = Field(default="", sa_type=String)
|
||||
created_at: datetime
|
||||
|
||||
@validates("learner_id", "template_id", "task_id")
|
||||
def _ids_non_empty(self, key: str, value: str) -> str:
|
||||
if not value:
|
||||
raise ValueError(f"{key} must be a non-empty identifier")
|
||||
return value
|
||||
|
||||
@validates("seed")
|
||||
def _seed_non_empty(self, key: str, value: str) -> str:
|
||||
if not value:
|
||||
raise ValueError(f"{key} must be a non-empty seed string")
|
||||
return value
|
||||
|
||||
@validates("statement")
|
||||
def _statement_non_empty(self, key: str, value: str) -> str:
|
||||
if not value:
|
||||
raise ValueError(f"{key} must be a non-empty statement string")
|
||||
return value
|
||||
|
||||
|
||||
class VariantStore(Protocol):
|
||||
"""Persistence contract for reproducible per-learner task variants.
|
||||
|
||||
Implemented by SQLiteVariantStore (v0.3, D-027); a Postgres
|
||||
implementation must satisfy the same surface.
|
||||
"""
|
||||
|
||||
def save(self, variant: VariantRecord) -> None:
|
||||
"""Persist a new variant. INSERT-ONLY on (learner_id, template_id):
|
||||
the FIRST generated variant is authoritative (reproducibility,
|
||||
D-029); a duplicate pair raises sqlalchemy.exc.IntegrityError to
|
||||
the caller — the generator serves cached variants via `get`
|
||||
instead of saving again. `task_id` is unique too: claiming an
|
||||
existing trace key for a different variant is equally rejected.
|
||||
NOT upsert; contrast GradeStore.save (latest-wins) and
|
||||
TraceStore.append (dedup-keep-first, swallowed).
|
||||
"""
|
||||
...
|
||||
|
||||
def get(self, learner_id: str, template_id: str) -> VariantRecord | None:
|
||||
"""The learner's stored variant for the template; None when none
|
||||
exists. Detached from any DB session — safe to pass across layers.
|
||||
"""
|
||||
...
|
||||
|
||||
def get_by_task(self, task_id: str) -> VariantRecord | None:
|
||||
"""The variant owning the task key (the grading/telemetry join
|
||||
path); None when none exists. Detached from any DB session.
|
||||
"""
|
||||
...
|
||||
|
||||
def list_for_learner(self, learner_id: str) -> list[VariantRecord]:
|
||||
"""All stored variants for the learner, ordered by created_at
|
||||
ascending (chronological; task_id breaks same-instant ties).
|
||||
Empty list when the learner has none.
|
||||
"""
|
||||
...
|
||||
|
||||
def list_by_template(self, template_id: str) -> list[VariantRecord]:
|
||||
"""All stored variants generated from the template — one row per
|
||||
learner — ordered by created_at ascending (chronological;
|
||||
learner_id breaks same-instant ties). Empty list when the
|
||||
template has none. The proctoring cross-check path (seed params
|
||||
per learner) reads through this.
|
||||
"""
|
||||
...
|
||||
|
||||
def close(self) -> None:
|
||||
"""Release DB connections. Store must not be used after close."""
|
||||
...
|
||||
|
||||
|
||||
def _sqlite_connect(dbapi_connection: sqlite3.Connection, _: object) -> None:
|
||||
"""Per-connection pragma setup (a-3). Mirrors telemetry/grading stores.
|
||||
|
||||
journal_mode=WAL — readers never block the single writer.
|
||||
synchronous=NORMAL — safe in WAL mode, avoids full fsync-per-commit.
|
||||
busy_timeout=5000 — retry briefly under contention instead of
|
||||
`OperationalError: database is locked`.
|
||||
"""
|
||||
cursor = dbapi_connection.cursor()
|
||||
cursor.execute("PRAGMA journal_mode=WAL")
|
||||
cursor.execute("PRAGMA synchronous=NORMAL")
|
||||
cursor.execute("PRAGMA busy_timeout=5000")
|
||||
cursor.close()
|
||||
|
||||
|
||||
def _as_utc(ts: datetime) -> datetime:
|
||||
"""Timestamps travel as naive datetime on SQLite, tz-aware elsewhere.
|
||||
|
||||
SQLite (via SQLModel) drops tzinfo; Postgres TIMESTAMP WITH TIME ZONE
|
||||
keeps it. Normalizing on the read path makes the store's contract
|
||||
tz-aware UTC regardless of the backend (D-027).
|
||||
"""
|
||||
if ts.tzinfo is None:
|
||||
return ts.replace(tzinfo=UTC) # naive UTC read-side: label as UTC
|
||||
return ts.astimezone(UTC)
|
||||
|
||||
|
||||
class SQLiteVariantStore:
|
||||
"""SQLite-backed VariantStore (SQLModel). Third protocol-wrapped store
|
||||
of the D-027 family (first: SQLiteTraceStore, second: SQLiteGradeStore).
|
||||
"""
|
||||
|
||||
def __init__(self, db_path: Path | None = None) -> None:
|
||||
self._db_path: Path = db_path if db_path is not None else Settings().db_path
|
||||
self._engine = create_engine(f"sqlite:///{self._db_path}")
|
||||
sa.event.listen(self._engine, "connect", _sqlite_connect)
|
||||
SQLModel.metadata.create_all(self._engine)
|
||||
# v0.5 schema (D-044) added two columns to an existing table.
|
||||
# create_all does NOT ALTER existing tables: on a box with a
|
||||
# pre-v0.5 ~/.nextcraft/data/nextcraft.db, every variant read/write
|
||||
# would raise OperationalError("no such column: variant_record.
|
||||
# environment") — a silent total breakage of the variant path
|
||||
# (final-review P0, verified empirically). Backfill the missing
|
||||
# columns with the model defaults ('build' keeps pre-v0.5 rows
|
||||
# build-kind per the field contract; '' falls back to pytest at
|
||||
# the API seam, api/variants._to_response). Idempotent: the
|
||||
# PRAGMA table_info check makes re-runs no-ops.
|
||||
self._ensure_v05_columns()
|
||||
|
||||
def _ensure_v05_columns(self) -> None:
|
||||
"""Add v0.5 columns to a pre-v0.5 variant_record table (idempotent)."""
|
||||
from sqlalchemy import text
|
||||
|
||||
with self._engine.begin() as conn:
|
||||
columns = {row[1] for row in conn.execute(text("PRAGMA table_info(variant_record)"))}
|
||||
if "environment" not in columns:
|
||||
conn.execute(
|
||||
text(
|
||||
"ALTER TABLE variant_record ADD COLUMN environment "
|
||||
"VARCHAR DEFAULT 'build' NOT NULL"
|
||||
)
|
||||
)
|
||||
logger.info("variant store: backfilled 'environment' (pre-v0.5 schema)")
|
||||
if "test_command" not in columns:
|
||||
conn.execute(
|
||||
text(
|
||||
"ALTER TABLE variant_record ADD COLUMN test_command "
|
||||
"VARCHAR DEFAULT '' NOT NULL"
|
||||
)
|
||||
)
|
||||
logger.info("variant store: backfilled 'test_command' (pre-v0.5 schema)")
|
||||
|
||||
@contextmanager
|
||||
def _session(self) -> Iterator[Session]:
|
||||
# expire_on_commit=False: identical session behavior to the other
|
||||
# D-027 stores. save() never commits on the error path and the read
|
||||
# paths never commit, but a uniform flag across the family keeps
|
||||
# their detachment guarantees from diverging.
|
||||
with Session(self._engine, expire_on_commit=False) as session:
|
||||
yield session
|
||||
|
||||
def save(self, variant: VariantRecord) -> None:
|
||||
# Plain INSERT, no merge: overwriting a stored variant would swap a
|
||||
# learner's graded task underneath its trace and grade (audit
|
||||
# corruption), so a duplicate identity is a race or bug to SURFACE,
|
||||
# not paper over. The generator's cache path (get before generate)
|
||||
# makes duplicate saves a programming error, not a normal flow.
|
||||
# The trace store swallows its IntegrityError (dedup is the
|
||||
# contract there); the grade store merges (latest-wins is the
|
||||
# contract there); this store re-raises (first-wins is the
|
||||
# contract here).
|
||||
with self._session() as session:
|
||||
try:
|
||||
session.add(variant)
|
||||
session.commit()
|
||||
except sa.exc.IntegrityError:
|
||||
session.rollback()
|
||||
logger.debug(
|
||||
"variant insert rejected (identity already stored): "
|
||||
"learner=%s template=%s task=%s",
|
||||
variant.learner_id,
|
||||
variant.template_id,
|
||||
variant.task_id,
|
||||
)
|
||||
raise
|
||||
logger.debug(
|
||||
"variant saved: %s/%s task=%s seed=%s",
|
||||
variant.learner_id,
|
||||
variant.template_id,
|
||||
variant.task_id,
|
||||
variant.seed,
|
||||
)
|
||||
|
||||
def get(self, learner_id: str, template_id: str) -> VariantRecord | None:
|
||||
with self._session() as session:
|
||||
record = session.get(VariantRecord, (learner_id, template_id))
|
||||
if record is None:
|
||||
return None
|
||||
record.created_at = _as_utc(record.created_at)
|
||||
# Detach from the session: callers must not depend on
|
||||
# open-session ORM magic (lazy loads fail once it closes).
|
||||
session.expunge(record)
|
||||
return record
|
||||
|
||||
def get_by_task(self, task_id: str) -> VariantRecord | None:
|
||||
with self._session() as session:
|
||||
stmt = select(VariantRecord).where(VariantRecord.task_id == task_id)
|
||||
record = session.exec(stmt).first()
|
||||
if record is None:
|
||||
return None
|
||||
record.created_at = _as_utc(record.created_at)
|
||||
session.expunge(record)
|
||||
return record
|
||||
|
||||
def list_for_learner(self, learner_id: str) -> list[VariantRecord]:
|
||||
with self._session() as session:
|
||||
stmt = (
|
||||
select(VariantRecord)
|
||||
.where(VariantRecord.learner_id == learner_id)
|
||||
# Chronological; task_id is a deterministic tie-break for
|
||||
# variants stamped within the same instant.
|
||||
.order_by(VariantRecord.created_at, VariantRecord.task_id)
|
||||
)
|
||||
results = session.exec(stmt).all()
|
||||
for row in results:
|
||||
row.created_at = _as_utc(row.created_at)
|
||||
session.expunge(row)
|
||||
return list(results)
|
||||
|
||||
def list_by_template(self, template_id: str) -> list[VariantRecord]:
|
||||
with self._session() as session:
|
||||
stmt = (
|
||||
select(VariantRecord)
|
||||
.where(VariantRecord.template_id == template_id)
|
||||
# Chronological; learner_id is a deterministic tie-break.
|
||||
.order_by(VariantRecord.created_at, VariantRecord.learner_id)
|
||||
)
|
||||
results = session.exec(stmt).all()
|
||||
for row in results:
|
||||
row.created_at = _as_utc(row.created_at)
|
||||
session.expunge(row)
|
||||
return list(results)
|
||||
|
||||
def close(self) -> None:
|
||||
self._engine.dispose()
|
||||
@@ -1,474 +0,0 @@
|
||||
"""Task template library for seeded variant generation (D-029, REQ-3-005).
|
||||
|
||||
A `TaskTemplate` binds a competency (D-021-aligned corpus ID), a statement
|
||||
skeleton with `{slot}` placeholders, typed `ParameterSlot`s, difficulty-
|
||||
normalization rubric anchors (the expected feature envelope that bounds
|
||||
variant fairness in the a-5 envelope test — grader-prompt shipment is the
|
||||
tracked P4 follow-up; grading is variant-blind today), and starter-file
|
||||
scaffolds served into the sandbox workdir (wired in P6).
|
||||
|
||||
Slot sampling is PURE CODE: `random.Random(seed)` over typed slots — fully
|
||||
reproducible for a given seed, independent of the LLM. The LLM only renders
|
||||
the seeded slot values into the statement skeleton (D-020 defense).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import random
|
||||
import re
|
||||
from typing import Literal
|
||||
|
||||
from pydantic import BaseModel, ConfigDict, Field, field_validator
|
||||
|
||||
#: REQ-5-005 environment kinds (D-044): one namespace fabric, typed starter
|
||||
#: contents + command policy. 'build' = the v0.3 coding IDE; 'design' =
|
||||
#: artifact editing with a validator/renderer harness; 'simulation' = a
|
||||
#: parameterized run harness (benchmark scripts + datasets).
|
||||
EnvironmentKind = Literal["build", "design", "simulation"]
|
||||
|
||||
|
||||
def validate_simple_argv(value: str) -> str:
|
||||
"""G-15: command fields must roundtrip shlex.split → join → split.
|
||||
|
||||
No quotes, no shell metachars — the TS client splits on whitespace only
|
||||
(no shlex in browsers), so anything quote-aware would split differently
|
||||
on the two sides. A violation is a template-AUTHORING bug caught here,
|
||||
at definition time, in Python where shlex exists.
|
||||
"""
|
||||
import shlex
|
||||
|
||||
parts = shlex.split(value)
|
||||
if not parts:
|
||||
raise ValueError("command must not be empty")
|
||||
joined = " ".join(parts)
|
||||
if shlex.split(joined) != parts:
|
||||
raise ValueError(f"command is not whitespace-joinable: {value!r}")
|
||||
return joined
|
||||
|
||||
SlotType = Literal["enum", "int_range", "string_set"]
|
||||
|
||||
|
||||
class ParameterSlot(BaseModel):
|
||||
"""One typed fill-in for a statement skeleton."""
|
||||
|
||||
model_config = ConfigDict(frozen=True)
|
||||
|
||||
name: str = Field(min_length=1)
|
||||
type: SlotType
|
||||
values: list[str] = Field(default_factory=list) # enum/string_set options
|
||||
lo: int | None = None # int_range bounds
|
||||
hi: int | None = None
|
||||
|
||||
@field_validator("values")
|
||||
@classmethod
|
||||
def _values_nonempty_for_enums(cls, v: list[str], info) -> list[str]:
|
||||
if info.data.get("type") in ("enum", "string_set") and not v:
|
||||
raise ValueError(f"slot {info.data.get('name')!r} needs values")
|
||||
return v
|
||||
|
||||
def sample(self, rng: random.Random) -> str | int:
|
||||
"""Deterministic sample from the seeded RNG. Validated after sampling."""
|
||||
if self.type == "enum" or self.type == "string_set":
|
||||
return rng.choice(self.values)
|
||||
if self.type == "int_range":
|
||||
lo = self.lo if self.lo is not None else 0
|
||||
hi = self.hi if self.hi is not None else lo
|
||||
if hi < lo:
|
||||
raise ValueError(f"slot {self.name!r}: hi < lo")
|
||||
return rng.randint(lo, hi)
|
||||
raise ValueError(f"unsupported slot type: {self.type!r}")
|
||||
|
||||
def validate_value(self, value: str | int) -> bool:
|
||||
"""Is `value` schema-valid for this slot? (params JSON gate, a-5.)"""
|
||||
if self.type in ("enum", "string_set"):
|
||||
return isinstance(value, str) and value in self.values
|
||||
if self.type == "int_range":
|
||||
lo = self.lo if self.lo is not None else 0
|
||||
hi = self.hi if self.hi is not None else lo
|
||||
return isinstance(value, int) and lo <= value <= hi
|
||||
return False
|
||||
|
||||
|
||||
class RubricAnchors(BaseModel):
|
||||
"""Difficulty-normalization anchors for the grader (a-5).
|
||||
|
||||
Expected FEATURE ENVELOPE (digest-space): the expected effort band
|
||||
for this template, so two variants of one template are held to the
|
||||
same bar regardless of which slot values a learner drew. The a-5
|
||||
envelope test (tests/variants/test_generator.py) binds variants to
|
||||
these bands in code, and — since Phase 4 (MH#4) — the grading engine
|
||||
ships this envelope into the grader prompt
|
||||
(grading/engine._anchors_context) and stamps the variant seed on the
|
||||
GradeRecord, so the anchors gate variant fairness in BOTH tests and
|
||||
the live rubric.
|
||||
"""
|
||||
|
||||
model_config = ConfigDict(frozen=True)
|
||||
|
||||
expected_edit_count_band: tuple[int, int]
|
||||
expected_min_test_runs: int
|
||||
expected_error_fix_cycles_band: tuple[int, int]
|
||||
notes: str = ""
|
||||
|
||||
|
||||
class TaskTemplate(BaseModel):
|
||||
"""A reusable task shape; variants instantiate it per learner."""
|
||||
|
||||
model_config = ConfigDict(frozen=True)
|
||||
|
||||
id: str = Field(min_length=1)
|
||||
competency_id: str = Field(min_length=1) # D-021 corpus alignment
|
||||
title: str
|
||||
statement_skeleton: str = Field(min_length=1) # {slot} placeholders
|
||||
slots: list[ParameterSlot] = Field(min_length=1)
|
||||
rubric_anchors: RubricAnchors
|
||||
starter_files: dict[str, str] = Field(default_factory=dict) # path -> content
|
||||
test_command: str
|
||||
#: REQ-5-005 (D-044): the environment kind rides the variant through
|
||||
#: the API to the TS client; 'build' default keeps v0.3 behavior.
|
||||
environment: EnvironmentKind = "build"
|
||||
#: The kind's Run harness (design: validator/renderer; simulation:
|
||||
#: benchmark script). Defaults to the test_command for build kinds.
|
||||
harness_command: str = ""
|
||||
|
||||
@field_validator("statement_skeleton")
|
||||
@classmethod
|
||||
def _skeleton_placeholders(cls, v: str) -> str:
|
||||
if "{" not in v or "}" not in v:
|
||||
raise ValueError("statement_skeleton needs at least one {slot}")
|
||||
return v
|
||||
|
||||
@field_validator("test_command", "harness_command")
|
||||
@classmethod
|
||||
def _simple_argv(cls, v: str) -> str:
|
||||
return validate_simple_argv(v) if v else v
|
||||
|
||||
@property
|
||||
def run_command(self) -> str:
|
||||
"""The Run button's command: kind harness when declared, else tests."""
|
||||
return self.harness_command or self.test_command
|
||||
|
||||
def render(self, params: dict[str, str | int]) -> str:
|
||||
"""Fill the skeleton with validated params."""
|
||||
for slot in self.slots:
|
||||
if slot.name not in params:
|
||||
raise ValueError(f"missing param for slot {slot.name!r}")
|
||||
if not slot.validate_value(params[slot.name]):
|
||||
raise ValueError(f"invalid value for slot {slot.name!r}: {params[slot.name]!r}")
|
||||
return self.statement_skeleton.format(**params)
|
||||
|
||||
def sample_params(self, seed: int) -> dict[str, str | int]:
|
||||
"""Seeded, reproducible, schema-valid slot values (pure code)."""
|
||||
rng = random.Random(seed)
|
||||
return {slot.name: slot.sample(rng) for slot in self.slots}
|
||||
|
||||
|
||||
# --- Template library (v0.3 initial set) --------------------------------------
|
||||
# Competency IDs are D-021-aligned with the Python corpus
|
||||
# (ai_service/corpus/learner_context.py) and the TS mock-data layer
|
||||
# (packages/mock-data/competency-stacks.ts: deterministic cid() scheme).
|
||||
|
||||
TEMPLATES: dict[str, TaskTemplate] = {
|
||||
"tpl-llm-judge": TaskTemplate(
|
||||
id="tpl-llm-judge",
|
||||
competency_id="stack-orchestration-c007",
|
||||
title="Build an LLM-as-Judge Evaluator",
|
||||
statement_skeleton=(
|
||||
"Build a small LLM-as-judge evaluator for {domain} answers. "
|
||||
"The judge must score each answer on {criterion} using a 0-4 scale, "
|
||||
"return structured JSON, and handle at least {edge_cases} edge-case "
|
||||
"answer classes (empty, off-topic, adversarial). Include a tiny "
|
||||
"repro test set of at least {test_size} examples and print a summary "
|
||||
"table of scores."
|
||||
),
|
||||
slots=[
|
||||
ParameterSlot(
|
||||
name="domain",
|
||||
type="enum",
|
||||
values=["customer-support", "code-review", "summarization", "tutoring"],
|
||||
),
|
||||
ParameterSlot(
|
||||
name="criterion",
|
||||
type="enum",
|
||||
values=["factual-accuracy", "helpfulness", "safety", "completeness"],
|
||||
),
|
||||
ParameterSlot(name="edge_cases", type="int_range", lo=2, hi=4),
|
||||
ParameterSlot(name="test_size", type="int_range", lo=3, hi=8),
|
||||
],
|
||||
rubric_anchors=RubricAnchors(
|
||||
expected_edit_count_band=(3, 25),
|
||||
expected_min_test_runs=2,
|
||||
expected_error_fix_cycles_band=(0, 4),
|
||||
notes="Slot draw changes the SCENARIO, not the engineering depth.",
|
||||
),
|
||||
starter_files={
|
||||
"README.md": (
|
||||
"# LLM-as-Judge Evaluator\n\n"
|
||||
"Implement `judge.py`:\n"
|
||||
"- `score(answer: str) -> dict` — 0-4 on the named criterion\n"
|
||||
"- structured JSON output (schema below)\n"
|
||||
"- edge-case classes handled explicitly\n"
|
||||
"- `pytest` must pass\n"
|
||||
),
|
||||
"judge.py": "def score(answer: str) -> dict:\n raise NotImplementedError\n",
|
||||
"test_judge.py": "def test_placeholder():\n assert True\n",
|
||||
},
|
||||
test_command="pytest -q",
|
||||
),
|
||||
"tpl-guardrail-schema": TaskTemplate(
|
||||
id="tpl-guardrail-schema",
|
||||
competency_id="stack-orchestration-c008",
|
||||
title="Schema Guardrail Pipeline",
|
||||
statement_skeleton=(
|
||||
"Implement an output-validation guardrail for a model returning "
|
||||
"{entity} records. Validate against a typed schema with {field_count} "
|
||||
"required fields, coerce or reject {failure_mode} failures, and emit "
|
||||
"a fallback response for invalid payloads. Cover with at least "
|
||||
"{test_size} unit tests including malformed JSON."
|
||||
),
|
||||
slots=[
|
||||
ParameterSlot(
|
||||
name="entity",
|
||||
type="enum",
|
||||
values=["user-profile", "job-posting", "candidate", "invoice"],
|
||||
),
|
||||
ParameterSlot(
|
||||
name="failure_mode",
|
||||
type="enum",
|
||||
values=["strict-reject", "coerce-when-safe"],
|
||||
),
|
||||
ParameterSlot(name="field_count", type="int_range", lo=4, hi=8),
|
||||
ParameterSlot(name="test_size", type="int_range", lo=4, hi=10),
|
||||
],
|
||||
rubric_anchors=RubricAnchors(
|
||||
expected_edit_count_band=(3, 30),
|
||||
expected_min_test_runs=2,
|
||||
expected_error_fix_cycles_band=(0, 5),
|
||||
notes="All slot draws land in the same engineering band.",
|
||||
),
|
||||
starter_files={
|
||||
"README.md": (
|
||||
"# Schema Guardrail\n\nImplement `guardrail.py`:\n"
|
||||
"- `validate(payload: dict) -> dict | Fallback`\n"
|
||||
"- required-field checks, failure policy, fallback emission\n"
|
||||
),
|
||||
"guardrail.py": "def validate(payload: dict):\n raise NotImplementedError\n",
|
||||
"test_guardrail.py": "def test_placeholder():\n assert True\n",
|
||||
},
|
||||
test_command="pytest -q",
|
||||
),
|
||||
"tpl-rag-chunker": TaskTemplate(
|
||||
id="tpl-rag-chunker",
|
||||
competency_id="stack-orchestration-c005",
|
||||
title="RAG Chunking Strategy",
|
||||
statement_skeleton=(
|
||||
"Implement a document chunker for {doc_type} retrieval. Support "
|
||||
"{strategy} chunking with a target size of ~{chunk_size} tokens, "
|
||||
"preserve {invariant} across chunk boundaries, and evaluate overlap "
|
||||
"quality with at least {test_size} fixture documents."
|
||||
),
|
||||
slots=[
|
||||
ParameterSlot(
|
||||
name="doc_type",
|
||||
type="enum",
|
||||
values=["technical-docs", "legal-contracts", "transcripts"],
|
||||
),
|
||||
ParameterSlot(
|
||||
name="strategy",
|
||||
type="enum",
|
||||
values=["fixed-window", "semantic-boundary", "hybrid"],
|
||||
),
|
||||
ParameterSlot(
|
||||
name="invariant",
|
||||
type="enum",
|
||||
values=["code-block-integrity", "section-headers", "sentence-completeness"],
|
||||
),
|
||||
ParameterSlot(name="chunk_size", type="int_range", lo=200, hi=800),
|
||||
ParameterSlot(name="test_size", type="int_range", lo=3, hi=6),
|
||||
],
|
||||
rubric_anchors=RubricAnchors(
|
||||
expected_edit_count_band=(4, 35),
|
||||
expected_min_test_runs=2,
|
||||
expected_error_fix_cycles_band=(0, 6),
|
||||
notes="Strategy draw changes implementation shape, not depth.",
|
||||
),
|
||||
starter_files={
|
||||
"README.md": (
|
||||
"# RAG Chunker\n\nImplement `chunker.py`:\n"
|
||||
"- `chunk(text: str) -> list[str]`\n- invariant preserved\n- tests green\n"
|
||||
),
|
||||
"chunker.py": "def chunk(text: str) -> list[str]:\n raise NotImplementedError\n",
|
||||
"test_chunker.py": "def test_placeholder():\n assert True\n",
|
||||
},
|
||||
test_command="pytest -q",
|
||||
),
|
||||
# -- REQ-5-005 design environment (D-044): artifact editing with a
|
||||
# -- validator/renderer harness - same fabric, typed starter contents.
|
||||
"tpl-conversation-flow-design": TaskTemplate(
|
||||
id="tpl-conversation-flow-design",
|
||||
competency_id="stack-designer-c001",
|
||||
title="Conversational Flow Artifact",
|
||||
statement_skeleton=(
|
||||
"Design a conversational flow for a {persona} assistant helping "
|
||||
"users accomplish {goal}. Author the flow as a structured artifact "
|
||||
"with at least {turn_count} conversation turns, explicit fallback "
|
||||
"paths for misunderstandings, and an AI-transparency disclosure "
|
||||
"pattern. The flow must render validly (the harness validates "
|
||||
"structure) and read naturally end to end."
|
||||
),
|
||||
slots=[
|
||||
ParameterSlot(
|
||||
name="persona",
|
||||
type="enum",
|
||||
values=["travel-planner", "homework-tutor", "fitness-coach", "recipe-guide"],
|
||||
),
|
||||
ParameterSlot(
|
||||
name="goal",
|
||||
type="enum",
|
||||
values=["book-a-trip", "master-a-concept", "start-a-routine", "cook-a-meal"],
|
||||
),
|
||||
ParameterSlot(name="turn_count", type="int_range", lo=6, hi=12),
|
||||
],
|
||||
rubric_anchors=RubricAnchors(
|
||||
expected_edit_count_band=(3, 20),
|
||||
expected_min_test_runs=1,
|
||||
expected_error_fix_cycles_band=(0, 3),
|
||||
notes="Design kind: artifact quality + iteration cadence, not code depth.",
|
||||
),
|
||||
starter_files={
|
||||
"README.md": (
|
||||
"# Conversational Flow Design\n\n"
|
||||
"Edit flow.md - the structured flow artifact. python3 "
|
||||
"validate_flow.py checks structure (turn headings, fallback "
|
||||
"sections, a transparency disclosure) and reports issues.\n"
|
||||
),
|
||||
"flow.md": (
|
||||
"# Flow: your persona here\n\n"
|
||||
"## Turn 1\n- **AI:** (opening)\n- **User (expected):** ...\n\n"
|
||||
"## Fallback\n- (misunderstanding handling)\n\n"
|
||||
"## AI Transparency Disclosure\n- (disclosure pattern)\n"
|
||||
),
|
||||
"validate_flow.py": (
|
||||
"import re, sys\n"
|
||||
"text = open('flow.md').read()\n"
|
||||
"issues = []\n"
|
||||
"turns = len(re.findall(r'^## Turn', text, re.M))\n"
|
||||
"if turns < 3:\n"
|
||||
" issues.append(f'expected at least 3 turn sections, found {turns}')\n"
|
||||
"if not re.search(r'^## Fallback', text, re.M):\n"
|
||||
" issues.append('missing Fallback section')\n"
|
||||
"if not re.search(r'^## AI Transparency', text, re.M):\n"
|
||||
" issues.append('missing AI Transparency Disclosure')\n"
|
||||
"print('VALID' if not issues else 'ISSUES: ' + '; '.join(issues))\n"
|
||||
"sys.exit(0 if not issues else 1)\n"
|
||||
),
|
||||
},
|
||||
test_command="python3 validate_flow.py",
|
||||
environment="design",
|
||||
harness_command="python3 validate_flow.py",
|
||||
),
|
||||
# -- REQ-5-005 simulation environment (D-044): parameterized benchmark
|
||||
# -- harness with dataset generation.
|
||||
"tpl-sensor-benchmark": TaskTemplate(
|
||||
id="tpl-sensor-benchmark",
|
||||
competency_id="stack-orchestration-c011",
|
||||
title="Sensor Data Simulation Harness",
|
||||
statement_skeleton=(
|
||||
"Build a simulation harness for {sensor} readings over {duration_min} "
|
||||
"minutes at {sample_hz} Hz. Generate a synthetic dataset with a "
|
||||
"realistic noise profile, run the analysis pipeline, and print a "
|
||||
"metrics summary (mean, p95, anomaly count at {anomaly_sigma} sigma). "
|
||||
"The harness must be reproducible from the committed seed."
|
||||
),
|
||||
slots=[
|
||||
ParameterSlot(
|
||||
name="sensor",
|
||||
type="enum",
|
||||
values=["temperature", "vibration", "luminosity", "pressure"],
|
||||
),
|
||||
ParameterSlot(name="duration_min", type="int_range", lo=5, hi=60),
|
||||
ParameterSlot(name="sample_hz", type="int_range", lo=1, hi=10),
|
||||
ParameterSlot(name="anomaly_sigma", type="int_range", lo=2, hi=4),
|
||||
],
|
||||
rubric_anchors=RubricAnchors(
|
||||
expected_edit_count_band=(3, 25),
|
||||
expected_min_test_runs=2,
|
||||
expected_error_fix_cycles_band=(0, 3),
|
||||
notes="Simulation kind: pipeline correctness + reproducibility.",
|
||||
),
|
||||
starter_files={
|
||||
"README.md": (
|
||||
"# Sensor Simulation Harness\n\n"
|
||||
"Edit simulate.py - python3 simulate.py runs the full pipeline: "
|
||||
"generate, analyze, print metrics. pytest covers the analysis "
|
||||
"functions.\n"
|
||||
),
|
||||
"simulate.py": (
|
||||
"import random, statistics\n\n"
|
||||
"def generate(n=300, seed=42):\n"
|
||||
" rng = random.Random(seed)\n"
|
||||
" return [rng.gauss(20.0, 1.5) for _ in range(n)]\n\n"
|
||||
"def analyze(samples, sigma=3):\n"
|
||||
" mean = statistics.fmean(samples)\n"
|
||||
" stdev = statistics.pstdev(samples)\n"
|
||||
" anomalies = [s for s in samples if abs(s - mean) > sigma * stdev]\n"
|
||||
" p95 = sorted(samples)[int(0.95 * len(samples))]\n"
|
||||
" return {'mean': mean, 'p95': p95, 'anomalies': len(anomalies)}\n\n"
|
||||
"if __name__ == '__main__':\n"
|
||||
" print(analyze(generate()))\n"
|
||||
),
|
||||
"test_simulate.py": (
|
||||
"from simulate import generate, analyze\n\n"
|
||||
"def test_reproducible():\n"
|
||||
" assert generate() == generate()\n\n"
|
||||
"def test_metrics_shape():\n"
|
||||
" m = analyze(generate())\n"
|
||||
" assert set(m) == {'mean', 'p95', 'anomalies'}\n"
|
||||
),
|
||||
},
|
||||
test_command="pytest -q",
|
||||
environment="simulation",
|
||||
harness_command="python3 simulate.py",
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
_KNOWN_COMPETENCY_IDS: set[str] = {
|
||||
# D-021: mirrored from ai_service/corpus/learner_context.py — the Python
|
||||
# source of truth for stack-orchestration competencies used by v0.2 agents.
|
||||
"stack-orchestration-c001",
|
||||
"stack-orchestration-c002",
|
||||
"stack-orchestration-c003",
|
||||
"stack-orchestration-c004",
|
||||
"stack-orchestration-c005",
|
||||
"stack-orchestration-c007",
|
||||
"stack-orchestration-c008",
|
||||
"stack-orchestration-c011",
|
||||
"stack-designer-c001",
|
||||
"stack-designer-c002",
|
||||
"stack-safety-c021",
|
||||
}
|
||||
|
||||
|
||||
def get_template(template_id: str) -> TaskTemplate | None:
|
||||
return TEMPLATES.get(template_id)
|
||||
|
||||
|
||||
def template_for_competency(competency_id: str) -> list[TaskTemplate]:
|
||||
return [t for t in TEMPLATES.values() if t.competency_id == competency_id]
|
||||
|
||||
|
||||
def validate_competency_binding() -> None:
|
||||
"""All templates must bind to known D-021 corpus competency IDs."""
|
||||
for t in TEMPLATES.values():
|
||||
if t.competency_id not in _KNOWN_COMPETENCY_IDS:
|
||||
raise ValueError(
|
||||
f"template {t.id!r} binds unknown competency {t.competency_id!r}"
|
||||
)
|
||||
|
||||
|
||||
def slots_pattern_ok(skeleton: str, slots: list[ParameterSlot]) -> bool:
|
||||
"""Every {placeholder} in the skeleton has a matching slot and vice versa."""
|
||||
placeholders = set(re.findall(r"\{([a-z_][a-z0-9_]*)\}", skeleton))
|
||||
slot_names = {s.name for s in slots}
|
||||
return placeholders == slot_names
|
||||
@@ -1,35 +0,0 @@
|
||||
"""VoiceProvider protocol (D-030, REQ-3-006) — mirrors the LLMProvider seam.
|
||||
|
||||
Two implementations in v0.3:
|
||||
- MockVoiceProvider — deterministic canned transcripts + canned tone WAV
|
||||
chunks + scripted failure modes (tests + no-key default; tests NEVER call
|
||||
a real voice API).
|
||||
- browser descriptor — not a provider but a FALLBACK HINT: the web client
|
||||
selects browser-native SpeechRecognition/speechSynthesis when the server
|
||||
reports no real voice backend.
|
||||
|
||||
OpenAIAudioProvider (real server STT/TTS over OpenAI-compatible
|
||||
/audio/transcriptions + /audio/speech) is INTENTIONALLY NOT BUILT in v0.3 —
|
||||
deferred to v0.4 with KYC, when there is a real key and real users
|
||||
(GRILL CUT-1 / G-7). This protocol is its future drop-in seam.
|
||||
|
||||
Boundary: `voice/` never imports `agents/` or `api/`.
|
||||
"""
|
||||
|
||||
from ai_service.voice.base import (
|
||||
TranscriptSegment,
|
||||
VoiceDescriptor,
|
||||
VoiceProvider,
|
||||
)
|
||||
from ai_service.voice.browser import BROWSER_FALLBACK_DESCRIPTOR
|
||||
from ai_service.voice.factory import voice_provider_from_settings
|
||||
from ai_service.voice.mock import MockVoiceProvider
|
||||
|
||||
__all__ = [
|
||||
"BROWSER_FALLBACK_DESCRIPTOR",
|
||||
"MockVoiceProvider",
|
||||
"TranscriptSegment",
|
||||
"VoiceDescriptor",
|
||||
"VoiceProvider",
|
||||
"voice_provider_from_settings",
|
||||
]
|
||||
@@ -1,62 +0,0 @@
|
||||
"""VoiceProvider protocol + shared voice contracts (D-030, REQ-3-006).
|
||||
|
||||
Mirrors the LLMProvider seam (D-014 pattern): a narrow protocol the Examiner
|
||||
agent and the defense API compose via DI, with a deterministic mock and a
|
||||
browser-fallback descriptor. No network in this module — concrete providers
|
||||
live in their own modules and are selected by factory/config.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from collections.abc import AsyncIterator
|
||||
from typing import Literal, Protocol, runtime_checkable
|
||||
|
||||
from pydantic import BaseModel, ConfigDict, Field
|
||||
|
||||
VoiceRole = Literal["examiner", "learner"]
|
||||
|
||||
|
||||
class TranscriptSegment(BaseModel):
|
||||
"""One STT result: the transcribed text + timing metadata."""
|
||||
|
||||
model_config = ConfigDict(frozen=True)
|
||||
|
||||
text: str = Field(min_length=1)
|
||||
language: str = "en"
|
||||
duration_ms: int | None = None
|
||||
confidence: float | None = Field(default=None, ge=0.0, le=1.0)
|
||||
|
||||
|
||||
class VoiceDescriptor(BaseModel):
|
||||
"""Capability descriptor served to the web client (D-030).
|
||||
|
||||
The assessment UI reads this to decide HOW the learner speaks/hears:
|
||||
- `mode="server"` → server-side STT/TTS (openai-audio, live since v0.5)
|
||||
- `mode="browser"` → browser-native SpeechRecognition/speechSynthesis
|
||||
- `mode="mock"` → deterministic no-op path (tests / no-key dev)
|
||||
The descriptor never contains secrets — only capability hints.
|
||||
"""
|
||||
|
||||
model_config = ConfigDict(frozen=True)
|
||||
|
||||
mode: Literal["server", "browser", "mock"]
|
||||
sr_available: bool
|
||||
tts_available: bool
|
||||
hint: str = ""
|
||||
|
||||
|
||||
@runtime_checkable
|
||||
class VoiceProvider(Protocol):
|
||||
"""The voice port (D-030): STT in, TTS out. Never imports agents/api."""
|
||||
|
||||
async def transcribe(self, audio: bytes, fmt: str) -> TranscriptSegment:
|
||||
"""STT: audio bytes (fmt: 'wav' | 'webm' | 'mp3') → transcript."""
|
||||
...
|
||||
|
||||
def synthesize(self, text: str, voice: str = "default") -> AsyncIterator[bytes]:
|
||||
"""TTS: text -> async byte chunks (audio stream).
|
||||
|
||||
Implementations may be async generators (async-def + yield) — the
|
||||
consumer contract is `async for chunk in provider.synthesize(text)`.
|
||||
"""
|
||||
...
|
||||
@@ -1,33 +0,0 @@
|
||||
"""Browser-native fallback descriptor (D-030, CUT-1 / G-7, REQ-3-006).
|
||||
|
||||
Browser-native SR/TTS is the no-key CLIENT-side path. When the factory
|
||||
selects `browser` mode, the defense endpoints return this descriptor and the
|
||||
WEB CLIENT performs SpeechRecognition + speechSynthesis natively; the server
|
||||
persists text turns as usual. Real server STT/TTS (openai-audio) is live
|
||||
since v0.5 — this descriptor is the no-key fallback.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from .base import VoiceDescriptor
|
||||
|
||||
BROWSER_FALLBACK_DESCRIPTOR = VoiceDescriptor(
|
||||
mode="browser",
|
||||
sr_available=True,
|
||||
tts_available=True,
|
||||
hint=(
|
||||
"No server voice backend configured. Use browser-native "
|
||||
"SpeechRecognition for STT and speechSynthesis for TTS; send the "
|
||||
"transcribed text to POST /v1/defense/{id}/answer ({text} form)."
|
||||
),
|
||||
)
|
||||
|
||||
MOCK_DESCRIPTOR = VoiceDescriptor(
|
||||
mode="mock",
|
||||
sr_available=True,
|
||||
tts_available=True,
|
||||
hint=(
|
||||
"Deterministic mock voice (tests / no-key dev). Real server "
|
||||
"STT/TTS is live since v0.5 (AI_VOICE_PROVIDER=openai-audio)."
|
||||
),
|
||||
)
|
||||
@@ -1,506 +0,0 @@
|
||||
"""DefenseStore — oral-defense persistence: protocol + SQLite impl (REQ-3-006, D-027).
|
||||
|
||||
FOURTH protocol-wrapped store of the D-027 family and the first spanning
|
||||
TWO related tables: `defense_record` (the defense session + integrity
|
||||
signals) and `defense_turn` (the ordered examiner/learner transcript,
|
||||
FK → defense_record.id).
|
||||
|
||||
Postgres-migration-ready (D-027): the protocol is the only surface the
|
||||
Examiner pipeline (task 5-2-01) and the defense endpoints (task 5-3-01)
|
||||
touch; swapping SQLiteDefenseStore for a Postgres implementation must
|
||||
not change call sites. Both tables use only portable column types
|
||||
(str / int / datetime / JSON), so the same SQLModel schema stands up
|
||||
unchanged on Postgres.
|
||||
|
||||
Save semantics — where this sits among the D-027 stores (each has a
|
||||
deliberately different contract):
|
||||
TraceStore.append dedup-keep-first; IntegrityError SWALLOWED
|
||||
(at-least-once event ingest).
|
||||
GradeStore.save upsert-latest-wins (a regrade is latest-state).
|
||||
VariantStore.save insert-only first-wins; IntegrityError RAISED
|
||||
(reproducibility; a duplicate is a bug).
|
||||
DefenseStore a LIFECYCLE store:
|
||||
start() insert-only; a duplicate id raises
|
||||
(a defense id is minted once per session).
|
||||
append_turn() insert-only per (defense_id, seq); a duplicate
|
||||
seq raises AND an unknown defense_id raises (FK
|
||||
enforced) — a transcript turn must never silently
|
||||
vanish (it is the integrity/grading input) nor
|
||||
attach to a defense that does not exist.
|
||||
finalize() targeted UPDATE (status → finished; finished_at +
|
||||
integrity_signals JSON). Unknown id → None
|
||||
(documented below). Re-finalize overwrites
|
||||
signals + finished_at — latest-wins, mirroring
|
||||
GradeStore.save: a recomputed verdict replaces
|
||||
the previous one wholesale.
|
||||
|
||||
append_turn does NOT police status (turns after finalize are a
|
||||
sequencing bug for the endpoints to prevent, task 5-3-01): the store
|
||||
enforces DATA integrity (FK + PK + non-empty), not workflow.
|
||||
|
||||
integrity_signals (A-109): JSON dict on the record — long pauses,
|
||||
off-scope cadence markers and friends, computed by the Examiner over
|
||||
turn metadata and persisted by finalize for the Proctor/Mentor feed.
|
||||
An empty dict is legal (defense not finished, or a clean defense).
|
||||
|
||||
Concurrency (a-3): WAL + synchronous=NORMAL + busy timeout at
|
||||
connection time (mirrors the other D-027 stores), PLUS foreign_keys=ON
|
||||
— this is the family's first real foreign key and it is actually
|
||||
enforced on SQLite, matching Postgres's native behavior (D-027 parity).
|
||||
|
||||
`created_at` / `ts` contract: callers stamp UTC (datetime.now(UTC));
|
||||
SQLite stores them naive and read paths re-label tz-aware UTC (same
|
||||
boundary normalization as TelemetryEvent.ts / GradeRecord.created_at,
|
||||
so the contract holds on any backend).
|
||||
|
||||
Boundary (D-027): `voice/` never imports `agents/` / `api/`; this
|
||||
module imports config only.
|
||||
"""
|
||||
|
||||
import logging
|
||||
import sqlite3
|
||||
from collections.abc import Iterator
|
||||
from contextlib import contextmanager
|
||||
from datetime import UTC, datetime
|
||||
from pathlib import Path
|
||||
from typing import Any, Literal, Protocol
|
||||
|
||||
import sqlalchemy as sa
|
||||
from sqlalchemy import JSON, Index, String
|
||||
from sqlalchemy.orm import validates
|
||||
from sqlmodel import Field, Session, SQLModel, create_engine, select
|
||||
|
||||
from ..config import Settings
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
DefenseStatus = Literal["in_progress", "finished"]
|
||||
_DEFENSE_STATUSES: frozenset[str] = frozenset(DefenseStatus.__args__)
|
||||
|
||||
TurnRole = Literal["examiner", "learner"]
|
||||
_TURN_ROLES: frozenset[str] = frozenset(TurnRole.__args__)
|
||||
|
||||
|
||||
class DefenseRecord(SQLModel, table=True):
|
||||
"""A persisted oral-defense session; id is the PK.
|
||||
|
||||
Written by the defense endpoints (task 5-3-01) through the
|
||||
DefenseStore protocol; read back by the endpoints, the Examiner
|
||||
pipeline and the Proctor/Mentor feeds. Constraint enforcement
|
||||
mirrors TelemetryEvent / GradeRecord / VariantRecord: sqlmodel
|
||||
0.0.42's metaclass drops pydantic constraints on table models, so
|
||||
SQLAlchemy `@validates` hooks enforce instead and the column types
|
||||
stay Postgres-ready (D-027).
|
||||
|
||||
Field contract:
|
||||
id — non-empty defense identifier, minted once
|
||||
per session (a duplicate start raises).
|
||||
learner_id — non-empty learner identifier (same id space
|
||||
as traces, grades and variants).
|
||||
task_id — non-empty task identifier; the defense
|
||||
defends the submitted work for this trace
|
||||
key ((learner_id, task_id) joins to the
|
||||
trace/grade/variant the defense is about).
|
||||
status — in_progress | finished; the STORE owns the
|
||||
transition: start() forces in_progress,
|
||||
finalize() sets finished. Validated.
|
||||
integrity_signals — A-109 signal dict (long pauses, off-scope
|
||||
cadence markers, ...); {} until finalize;
|
||||
JSON column. An empty dict is legal.
|
||||
created_at — UTC start timestamp.
|
||||
finished_at — UTC finalize timestamp; None while in
|
||||
progress.
|
||||
|
||||
`turns` (property): the seq-ordered DefenseTurn transcript, attached
|
||||
ONLY by DefenseStore.get(); records from list_for_learner carry
|
||||
turns == [] — call get() for a full transcript.
|
||||
"""
|
||||
|
||||
__tablename__ = "defense_record"
|
||||
# The id PK covers point lookups; this secondary index covers
|
||||
# list_for_learner ordered by created_at without a sort step
|
||||
# (Postgres migration target D-027).
|
||||
__table_args__ = (
|
||||
Index("ix_defense_record_learner_created", "learner_id", "created_at"),
|
||||
)
|
||||
|
||||
id: str = Field(primary_key=True)
|
||||
learner_id: str
|
||||
task_id: str
|
||||
# Bare Literal annotations crash sqlmodel<=0.0.42's column inference
|
||||
# (issubclass(TypeAlias, Enum)); an explicit sa_type + the validates
|
||||
# hook below give the same contract: VARCHAR column, Literal-rejected
|
||||
# values (same pattern as TelemetryEvent.kind).
|
||||
status: DefenseStatus = Field(default="in_progress", sa_type=String)
|
||||
# JSON column: stored as TEXT on SQLite, native JSONB on Postgres (D-027).
|
||||
integrity_signals: dict[str, Any] = Field(default_factory=dict, sa_type=JSON)
|
||||
created_at: datetime
|
||||
finished_at: datetime | None = Field(default=None)
|
||||
|
||||
@property
|
||||
def turns(self) -> list["DefenseTurn"]:
|
||||
"""Seq-ordered transcript; [] unless attached by get().
|
||||
|
||||
Table models reject ad-hoc attributes (pydantic __setattr__
|
||||
raises on non-fields), so the store stashes the detached turn
|
||||
list via object.__setattr__ and this read-only property surfaces
|
||||
it. The returned list is a copy — caller mutations cannot
|
||||
corrupt the stash.
|
||||
"""
|
||||
return list(self.__dict__.get("_turns", []))
|
||||
|
||||
@validates("id", "learner_id", "task_id")
|
||||
def _ids_non_empty(self, key: str, value: str) -> str:
|
||||
if not value:
|
||||
raise ValueError(f"{key} must be a non-empty identifier")
|
||||
return value
|
||||
|
||||
@validates("status")
|
||||
def _status_is_known(self, key: str, value: str) -> str:
|
||||
if value not in _DEFENSE_STATUSES:
|
||||
raise ValueError(f"unknown defense status: {value!r}")
|
||||
return value
|
||||
|
||||
|
||||
class DefenseTurn(SQLModel, table=True):
|
||||
"""One examiner/learner dialogue turn; (defense_id, seq) is the PK.
|
||||
|
||||
Rows are append-only transcript entries written through
|
||||
DefenseStore.append_turn. seq numbers the dialogue within one
|
||||
defense starting at 0; monotonic assignment is the endpoints' job
|
||||
(task 5-3-01), this model only rejects negatives — the same split
|
||||
as TelemetryEvent.seq (model rejects < 0, store owns ordering).
|
||||
|
||||
Field contract:
|
||||
defense_id — non-empty; FK → defense_record.id. ENFORCED on
|
||||
SQLite via foreign_keys=ON (first real FK in the
|
||||
D-027 family; Postgres enforces FKs natively, so
|
||||
this keeps the backends equivalent, D-027).
|
||||
seq — turn index within the defense, >= 0. (defense_id,
|
||||
seq) is the PK: a duplicate raises instead of
|
||||
silently overwriting — the transcript is the
|
||||
integrity/grading input, a vanishing turn is
|
||||
audit corruption.
|
||||
role — examiner | learner (who spoke). Validated.
|
||||
text — non-empty utterance text (examiner question, or
|
||||
STT output for learner answers).
|
||||
ts — UTC utterance timestamp.
|
||||
latency_ms — per-turn pipeline latency in ms (STT + LLM TTFT +
|
||||
TTS, A-109); int or None. Populated by the
|
||||
endpoints (task 5-4-01); None allowed here — the
|
||||
store persists, it does not measure.
|
||||
created_at — UTC row-write timestamp.
|
||||
"""
|
||||
|
||||
__tablename__ = "defense_turn"
|
||||
# The composite PK (defense_id, seq) doubles as the covering index
|
||||
# for the per-defense seq-ordered read in get() — no secondary index
|
||||
# needed (contrast defense_record's learner-listing index).
|
||||
|
||||
defense_id: str = Field(foreign_key="defense_record.id", primary_key=True)
|
||||
seq: int = Field(primary_key=True)
|
||||
role: TurnRole = Field(sa_type=String)
|
||||
text: str
|
||||
ts: datetime
|
||||
latency_ms: int | None = Field(default=None)
|
||||
created_at: datetime
|
||||
|
||||
@validates("defense_id")
|
||||
def _defense_id_non_empty(self, key: str, value: str) -> str:
|
||||
if not value:
|
||||
raise ValueError(f"{key} must be a non-empty identifier")
|
||||
return value
|
||||
|
||||
@validates("seq")
|
||||
def _seq_non_negative(self, key: str, value: int) -> int:
|
||||
if value < 0:
|
||||
raise ValueError("seq must be >= 0 (ordering is the endpoints' job)")
|
||||
return value
|
||||
|
||||
@validates("role")
|
||||
def _role_is_known(self, key: str, value: str) -> str:
|
||||
if value not in _TURN_ROLES:
|
||||
raise ValueError(f"unknown turn role: {value!r}")
|
||||
return value
|
||||
|
||||
@validates("text")
|
||||
def _text_non_empty(self, key: str, value: str) -> str:
|
||||
if not value:
|
||||
raise ValueError(f"{key} must be a non-empty utterance string")
|
||||
return value
|
||||
|
||||
@validates("latency_ms")
|
||||
def _latency_non_negative(self, key: str, value: int | None) -> int | None:
|
||||
# None is legal (not yet instrumented); a NEGATIVE latency is
|
||||
# nonsense and surfaces as a construction error.
|
||||
if value is not None and value < 0:
|
||||
raise ValueError("latency_ms must be >= 0 or None")
|
||||
return value
|
||||
|
||||
|
||||
class DefenseStore(Protocol):
|
||||
"""Persistence contract for oral-defense sessions + transcripts.
|
||||
|
||||
Implemented by SQLiteDefenseStore (v0.3, D-027); a Postgres
|
||||
implementation must satisfy the same surface.
|
||||
"""
|
||||
|
||||
def start(self, defense: DefenseRecord) -> DefenseRecord:
|
||||
"""Insert a new defense. INSERT-ONLY: a duplicate id raises
|
||||
sqlalchemy.exc.IntegrityError (a defense id is minted once per
|
||||
session — surfacing, not swallowing, mirrors VariantStore).
|
||||
The store owns the lifecycle: status is forced to "in_progress"
|
||||
and finished_at to None, whatever the caller passed — only
|
||||
finalize() may move a defense to finished. Returns the stored
|
||||
record, detached from any DB session.
|
||||
"""
|
||||
...
|
||||
|
||||
def append_turn(self, defense_id: str, turn: DefenseTurn) -> DefenseTurn:
|
||||
"""Insert one transcript turn, ordered by (defense_id, seq).
|
||||
turn.defense_id MUST equal the defense_id argument — a mismatch
|
||||
raises ValueError (the defense identity must never be
|
||||
ambiguous). A duplicate (defense_id, seq) raises
|
||||
IntegrityError; an unknown defense_id raises IntegrityError
|
||||
(FK enforced). Does NOT police status — sequencing turns vs
|
||||
finalize is the endpoints' job (task 5-3-01). Returns the
|
||||
stored turn, detached.
|
||||
"""
|
||||
...
|
||||
|
||||
def finalize(
|
||||
self, defense_id: str, integrity_signals: dict[str, Any]
|
||||
) -> DefenseRecord | None:
|
||||
"""Seal the defense: status → "finished", finished_at = now(UTC),
|
||||
integrity_signals stored as JSON. UNKNOWN defense_id → None
|
||||
(documented choice: the API layer maps it to 404 without an
|
||||
exception dance; contrast start/append_turn where IntegrityError
|
||||
IS the contract — those are inserts, this is an update on a key
|
||||
the caller may legitimately not hold). Re-finalize overwrites
|
||||
signals + finished_at: latest-wins, mirroring GradeStore.save
|
||||
(a recomputed verdict replaces the previous one wholesale).
|
||||
Returns the updated record, detached, WITHOUT turns — get() is
|
||||
the with-turns path.
|
||||
"""
|
||||
...
|
||||
|
||||
def get(self, defense_id: str) -> DefenseRecord | None:
|
||||
"""The defense with its FULL transcript (turns in seq order,
|
||||
detached) and integrity signals; None when it does not exist.
|
||||
Safe to pass across layers — no open-session ORM magic.
|
||||
"""
|
||||
...
|
||||
|
||||
def list_for_learner(self, learner_id: str) -> list[DefenseRecord]:
|
||||
"""All stored defenses for the learner, ordered by created_at
|
||||
ascending (chronological; id breaks same-instant ties), WITHOUT
|
||||
turns — records carry turns == []; call get() for a transcript.
|
||||
Empty list when the learner has none.
|
||||
"""
|
||||
...
|
||||
|
||||
def close(self) -> None:
|
||||
"""Release DB connections. Store must not be used after close."""
|
||||
...
|
||||
|
||||
|
||||
def _sqlite_connect(dbapi_connection: sqlite3.Connection, _: object) -> None:
|
||||
"""Per-connection pragma setup (a-3). Mirrors the other D-027 stores.
|
||||
|
||||
journal_mode=WAL — readers never block the single writer.
|
||||
synchronous=NORMAL — safe in WAL mode, avoids full fsync-per-commit.
|
||||
busy_timeout=5000 — retry briefly under contention instead of
|
||||
`OperationalError: database is locked`.
|
||||
foreign_keys=ON — NEW vs the family: defense_turn is the first
|
||||
real FK among the D-027 stores; SQLite leaves
|
||||
FKs OFF by default while Postgres enforces them
|
||||
natively, so the pragma keeps the backends
|
||||
equivalent (D-027 parity).
|
||||
"""
|
||||
cursor = dbapi_connection.cursor()
|
||||
cursor.execute("PRAGMA journal_mode=WAL")
|
||||
cursor.execute("PRAGMA synchronous=NORMAL")
|
||||
cursor.execute("PRAGMA busy_timeout=5000")
|
||||
cursor.execute("PRAGMA foreign_keys=ON")
|
||||
cursor.close()
|
||||
|
||||
|
||||
def _as_utc(ts: datetime) -> datetime:
|
||||
"""Timestamps travel as naive datetime on SQLite, tz-aware elsewhere.
|
||||
|
||||
SQLite (via SQLModel) drops tzinfo; Postgres TIMESTAMP WITH TIME ZONE
|
||||
keeps it. Normalizing on the read path makes the store's contract
|
||||
tz-aware UTC regardless of the backend (D-027).
|
||||
"""
|
||||
if ts.tzinfo is None:
|
||||
return ts.replace(tzinfo=UTC) # naive UTC read-side: label as UTC
|
||||
return ts.astimezone(UTC)
|
||||
|
||||
|
||||
class SQLiteDefenseStore:
|
||||
"""SQLite-backed DefenseStore (SQLModel). Fourth protocol-wrapped
|
||||
store of the D-027 family (first: SQLiteTraceStore, second:
|
||||
SQLiteGradeStore, third: SQLiteVariantStore) and the first spanning
|
||||
two related tables.
|
||||
"""
|
||||
|
||||
def __init__(self, db_path: Path | None = None) -> None:
|
||||
self._db_path: Path = db_path if db_path is not None else Settings().db_path
|
||||
self._engine = create_engine(f"sqlite:///{self._db_path}")
|
||||
sa.event.listen(self._engine, "connect", _sqlite_connect)
|
||||
SQLModel.metadata.create_all(self._engine)
|
||||
|
||||
@contextmanager
|
||||
def _session(self) -> Iterator[Session]:
|
||||
# expire_on_commit=False: identical session behavior to the other
|
||||
# D-027 stores. start/append_turn return the caller's instance
|
||||
# after commit and get/finalize return rows expunged mid-session;
|
||||
# a uniform flag across the family keeps their detachment
|
||||
# guarantees from diverging.
|
||||
with Session(self._engine, expire_on_commit=False) as session:
|
||||
yield session
|
||||
|
||||
def start(self, defense: DefenseRecord) -> DefenseRecord:
|
||||
# The store owns the lifecycle: a defense is BORN in_progress and
|
||||
# only finalize() may move it to finished. A smuggled "finished"
|
||||
# status is normalized away, not rejected — the insert itself
|
||||
# stays insert-only, and a duplicate id raises to the caller
|
||||
# (mirroring VariantStore: the id is minted once per session).
|
||||
defense.status = "in_progress"
|
||||
defense.finished_at = None
|
||||
with self._session() as session:
|
||||
try:
|
||||
session.add(defense)
|
||||
session.commit()
|
||||
except sa.exc.IntegrityError:
|
||||
session.rollback()
|
||||
logger.debug("defense start rejected (id already stored): %s", defense.id)
|
||||
raise
|
||||
logger.debug(
|
||||
"defense started: %s learner=%s task=%s",
|
||||
defense.id,
|
||||
defense.learner_id,
|
||||
defense.task_id,
|
||||
)
|
||||
return defense
|
||||
|
||||
def append_turn(self, defense_id: str, turn: DefenseTurn) -> DefenseTurn:
|
||||
# The explicit defense_id argument is the defense identity for
|
||||
# this write; a turn object claiming another defense is a
|
||||
# programming error — surface it before touching the DB.
|
||||
if turn.defense_id != defense_id:
|
||||
raise ValueError(
|
||||
f"turn.defense_id {turn.defense_id!r} does not match the "
|
||||
f"defense_id argument {defense_id!r}"
|
||||
)
|
||||
with self._session() as session:
|
||||
try:
|
||||
session.add(turn)
|
||||
session.commit()
|
||||
except sa.exc.IntegrityError:
|
||||
# Two possible causes, both surfaced, neither swallowed:
|
||||
# duplicate (defense_id, seq) PK — a transcript turn must
|
||||
# never silently vanish; unknown defense_id — the FK
|
||||
# (foreign_keys=ON) rejects the orphan.
|
||||
session.rollback()
|
||||
logger.debug(
|
||||
"defense turn rejected (duplicate (defense_id, seq) "
|
||||
"or unknown defense_id): defense=%s seq=%s",
|
||||
defense_id,
|
||||
turn.seq,
|
||||
)
|
||||
raise
|
||||
logger.debug(
|
||||
"defense turn appended: %s seq=%d role=%s",
|
||||
defense_id,
|
||||
turn.seq,
|
||||
turn.role,
|
||||
)
|
||||
return turn
|
||||
|
||||
def finalize(
|
||||
self, defense_id: str, integrity_signals: dict[str, Any]
|
||||
) -> DefenseRecord | None:
|
||||
# A None signals blob would break the read contract (signals are
|
||||
# a dict, {} until finalize); reject before writing.
|
||||
if not isinstance(integrity_signals, dict):
|
||||
raise ValueError(
|
||||
"integrity_signals must be a JSON-object dict, got "
|
||||
f"{type(integrity_signals).__name__}"
|
||||
)
|
||||
with self._session() as session:
|
||||
record = session.get(DefenseRecord, defense_id)
|
||||
if record is None:
|
||||
# Documented unknown-id behavior: None, not a raise — the
|
||||
# defense endpoints map this to 404. Contrast start() /
|
||||
# append_turn(), where IntegrityError IS the contract.
|
||||
return None
|
||||
# Latest-wins re-finalize, mirroring GradeStore.save: a
|
||||
# recomputed verdict (fresh signals) replaces the stored one
|
||||
# wholesale; status just stays finished.
|
||||
record.status = "finished"
|
||||
record.finished_at = datetime.now(UTC)
|
||||
record.integrity_signals = integrity_signals
|
||||
session.commit()
|
||||
record.created_at = _as_utc(record.created_at)
|
||||
if record.finished_at is not None:
|
||||
record.finished_at = _as_utc(record.finished_at)
|
||||
# Detach from the session: callers must not depend on
|
||||
# open-session ORM magic (lazy loads fail once it closes).
|
||||
session.expunge(record)
|
||||
logger.debug(
|
||||
"defense finalized: %s signals=%s", defense_id, sorted(integrity_signals)
|
||||
)
|
||||
return record
|
||||
|
||||
def get(self, defense_id: str) -> DefenseRecord | None:
|
||||
with self._session() as session:
|
||||
record = session.get(DefenseRecord, defense_id)
|
||||
if record is None:
|
||||
return None
|
||||
record.created_at = _as_utc(record.created_at)
|
||||
if record.finished_at is not None:
|
||||
record.finished_at = _as_utc(record.finished_at)
|
||||
stmt = (
|
||||
select(DefenseTurn)
|
||||
.where(DefenseTurn.defense_id == defense_id)
|
||||
.order_by(DefenseTurn.seq)
|
||||
)
|
||||
turns = session.exec(stmt).all()
|
||||
for turn in turns:
|
||||
turn.ts = _as_utc(turn.ts)
|
||||
turn.created_at = _as_utc(turn.created_at)
|
||||
# Detach each turn: the transcript must be usable once
|
||||
# the session closes (no lazy-load magic).
|
||||
session.expunge(turn)
|
||||
session.expunge(record)
|
||||
# Table models reject ad-hoc attributes (pydantic __setattr__
|
||||
# raises on non-fields), so the seq-ordered transcript is
|
||||
# stashed via object.__setattr__ and surfaced through the
|
||||
# read-only `turns` property. Rows are detached either way —
|
||||
# safe to pass across layers.
|
||||
object.__setattr__(record, "_turns", list(turns))
|
||||
return record
|
||||
|
||||
def list_for_learner(self, learner_id: str) -> list[DefenseRecord]:
|
||||
with self._session() as session:
|
||||
stmt = (
|
||||
select(DefenseRecord)
|
||||
.where(DefenseRecord.learner_id == learner_id)
|
||||
# Chronological; id is a deterministic tie-break for
|
||||
# defenses stamped within the same instant.
|
||||
.order_by(DefenseRecord.created_at, DefenseRecord.id)
|
||||
)
|
||||
results = session.exec(stmt).all()
|
||||
for row in results:
|
||||
row.created_at = _as_utc(row.created_at)
|
||||
if row.finished_at is not None:
|
||||
row.finished_at = _as_utc(row.finished_at)
|
||||
# Turns are deliberately NOT loaded here: the list feed
|
||||
# (Proctor/Mentor) needs session headers, not full
|
||||
# transcripts — get() is the with-turns path.
|
||||
session.expunge(row)
|
||||
return list(results)
|
||||
|
||||
def close(self) -> None:
|
||||
self._engine.dispose()
|
||||
@@ -1,67 +0,0 @@
|
||||
"""Voice provider factory (D-030; REQ-5-001 real path, D-040).
|
||||
|
||||
`AI_VOICE_PROVIDER = mock | browser | openai-audio` (default: mock — the
|
||||
no-key path is first-class). `openai-audio` requires voice_base_url +
|
||||
voice_api_key: the factory raises `UnknownVoiceProviderError` with an
|
||||
actionable message for direct callers (tests), while the lifespan in
|
||||
main.py CATCHES it and falls back to mock with a loud log — a typo'd env
|
||||
must never crash the unattended boot (G-11), and the mock provider's
|
||||
descriptor then honestly reports mode='mock' so the UI badge cannot lie.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import httpx
|
||||
|
||||
from ..config import Settings
|
||||
from .base import VoiceProvider
|
||||
from .mock import MockVoiceProvider
|
||||
from .openai_audio import OpenAIAudioProvider
|
||||
|
||||
|
||||
class UnknownVoiceProviderError(ValueError):
|
||||
"""Raised for a provider name outside the contract, or a real provider
|
||||
selected without its required configuration."""
|
||||
|
||||
|
||||
def voice_provider_from_settings(
|
||||
settings: Settings, http_client: httpx.AsyncClient | None = None
|
||||
) -> VoiceProvider:
|
||||
"""Select the voice provider by settings (env `AI_VOICE_PROVIDER`).
|
||||
|
||||
`http_client` is required for the `openai-audio` branch (D-017 shared
|
||||
pool); mock/browser ignore it.
|
||||
"""
|
||||
name = (settings.voice_provider or "mock").strip().lower()
|
||||
if name == "mock":
|
||||
return MockVoiceProvider()
|
||||
if name == "browser":
|
||||
# Browser mode is a CLIENT-side capability: the server composes the
|
||||
# same MockVoiceProvider (typed fallback answers still work; the UI
|
||||
# uses the descriptor for mic/speech). See browser.py.
|
||||
return MockVoiceProvider()
|
||||
if name in ("openai-audio", "openai", "server"):
|
||||
if not settings.voice_base_url or not settings.voice_api_key:
|
||||
raise UnknownVoiceProviderError(
|
||||
"AI_VOICE_PROVIDER=openai-audio requires AI_VOICE_BASE_URL "
|
||||
"and AI_VOICE_API_KEY — set both, or use 'mock'/'browser'. "
|
||||
"(main.py falls back to mock when these are missing; the "
|
||||
"voice badge then honestly reports mock — G-11)"
|
||||
)
|
||||
if http_client is None:
|
||||
raise UnknownVoiceProviderError(
|
||||
"openai-audio requires the shared httpx client "
|
||||
"(voice_provider_from_settings(settings, http_client))"
|
||||
)
|
||||
return OpenAIAudioProvider(
|
||||
http_client=http_client,
|
||||
base_url=settings.voice_base_url,
|
||||
api_key=settings.voice_api_key,
|
||||
stt_model=settings.voice_stt_model,
|
||||
tts_model=settings.voice_tts_model,
|
||||
tts_voice=settings.voice_tts_voice,
|
||||
tts_format=settings.voice_tts_format,
|
||||
)
|
||||
raise UnknownVoiceProviderError(
|
||||
f"unknown AI_VOICE_PROVIDER {name!r}: use 'mock', 'browser', or 'openai-audio'"
|
||||
)
|
||||
@@ -1,96 +0,0 @@
|
||||
"""Deterministic MockVoiceProvider (D-030, REQ-3-006).
|
||||
|
||||
Canned transcripts (scripted per test via queue) + canned 1kHz-tone WAV bytes
|
||||
+ scripted failure modes. Two identical transcribe calls yield identical
|
||||
segments; tests NEVER touch a real voice API (conftest cloud-free rule).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import io
|
||||
import math
|
||||
import struct
|
||||
import wave
|
||||
from collections.abc import AsyncIterator
|
||||
|
||||
from .base import TranscriptSegment
|
||||
|
||||
|
||||
def _tone_wav(duration_ms: int = 250, freq_hz: float = 1000.0) -> bytes:
|
||||
"""A small, deterministic 16-bit mono WAV: a sine tone (stdlib only)."""
|
||||
rate = 8000
|
||||
n_samples = max(1, int(rate * duration_ms / 1000))
|
||||
buf = io.BytesIO()
|
||||
with wave.open(buf, "wb") as w:
|
||||
w.setnchannels(1)
|
||||
w.setsampwidth(2)
|
||||
w.setframerate(rate)
|
||||
for i in range(n_samples):
|
||||
sample = int(12000 * math.sin(2 * math.pi * freq_hz * i / rate))
|
||||
w.writeframes(struct.pack("<h", sample))
|
||||
return buf.getvalue()
|
||||
|
||||
|
||||
class MockVoiceFailure(RuntimeError):
|
||||
"""Scripted failure mode for tests."""
|
||||
|
||||
|
||||
class MockVoiceProvider:
|
||||
"""Deterministic voice provider: scripted STT, canned-tone TTS.
|
||||
|
||||
- `transcribe`: pops the next scripted transcript from a queue (or a
|
||||
default); two identical calls with the same queue state are identical.
|
||||
Failure mode: raise MockVoiceFailure when the queue holds a failure
|
||||
marker (the string "FAIL") or `audio` is empty.
|
||||
- `synthesize`: yields the canned tone WAV in fixed-size chunks; failure
|
||||
mode: empty text raises MockVoiceFailure.
|
||||
"""
|
||||
|
||||
def __init__(self, transcripts: list[str] | None = None) -> None:
|
||||
self._transcripts = list(transcripts or [])
|
||||
self._cursor = 0
|
||||
self.transcribe_calls = 0
|
||||
self.synthesize_calls = 0
|
||||
|
||||
def script(self, transcripts: list[str]) -> None:
|
||||
"""Replace the scripted queue (tests set expectations up front)."""
|
||||
self._transcripts = list(transcripts)
|
||||
self._cursor = 0
|
||||
|
||||
async def transcribe(self, audio: bytes, fmt: str) -> TranscriptSegment:
|
||||
self.transcribe_calls += 1
|
||||
if not audio:
|
||||
raise MockVoiceFailure("no audio bytes provided")
|
||||
if not self._transcripts:
|
||||
raise MockVoiceFailure("transcript queue exhausted — script it")
|
||||
item = self._transcripts[self._cursor]
|
||||
self._cursor = (self._cursor + 1) % len(self._transcripts)
|
||||
if item == "FAIL":
|
||||
raise MockVoiceFailure("scripted STT failure")
|
||||
return TranscriptSegment(
|
||||
text=item,
|
||||
duration_ms=max(1, len(audio) // 32), # deterministic pseudo-duration
|
||||
)
|
||||
|
||||
async def synthesize(self, text: str, voice: str = "default") -> AsyncIterator[bytes]: # noqa: ASYNC109 (protocol parity)
|
||||
# NOTE: protocol parity matters more than the async-generator purity
|
||||
# lint; the real provider (openai_audio.py) streams over HTTP.
|
||||
self.synthesize_calls += 1
|
||||
if not text:
|
||||
raise MockVoiceFailure("cannot synthesize empty text")
|
||||
wav = _tone_wav(duration_ms=min(2000, max(120, len(text) * 12)))
|
||||
for i in range(0, len(wav), 1024):
|
||||
yield wav[i : i + 1024]
|
||||
await asyncio.sleep(0) # yield to the loop like a network stream
|
||||
|
||||
|
||||
# Protocol-shape parity guard (mock must satisfy the D-030 port).
|
||||
from .base import VoiceProvider # noqa: E402
|
||||
|
||||
|
||||
def _assert_protocol() -> None:
|
||||
assert isinstance(MockVoiceProvider(), VoiceProvider)
|
||||
|
||||
|
||||
_assert_protocol()
|
||||
@@ -1,133 +0,0 @@
|
||||
"""OpenAI-compatible audio provider — real server STT/TTS (D-040, REQ-5-001).
|
||||
|
||||
One implementation serves any OpenAI-compatible audio endpoint (base_url is
|
||||
config; A-301 endpoint-agnostic by config, D-014 pattern). Raw httpx on the
|
||||
shared lifespan client (D-017; read=300s tolerates multi-minute clips).
|
||||
|
||||
Boundary rules (mirror llm/openai_compat.py):
|
||||
- voice/ imports nothing from agents/ or api/
|
||||
- api_key NEVER appears in exceptions, logs, or error messages
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from collections.abc import AsyncIterator
|
||||
from typing import Literal
|
||||
|
||||
import httpx
|
||||
|
||||
from .base import TranscriptSegment, VoiceDescriptor
|
||||
|
||||
|
||||
class OpenAIAudioProvider:
|
||||
"""Server STT (`/audio/transcriptions`) + TTS (`/audio/speech`)."""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
http_client: httpx.AsyncClient,
|
||||
base_url: str,
|
||||
api_key: str,
|
||||
stt_model: str = "whisper-1",
|
||||
tts_model: str = "tts-1",
|
||||
tts_voice: str = "alloy",
|
||||
tts_format: Literal["mp3", "wav", "opus"] = "mp3",
|
||||
) -> None:
|
||||
self._client = http_client
|
||||
self._base_url = base_url.rstrip("/")
|
||||
self._api_key = api_key
|
||||
self._stt_model = stt_model
|
||||
self._tts_model = tts_model
|
||||
self._tts_voice = tts_voice
|
||||
self._tts_format = tts_format
|
||||
# a-15: the descriptor is what defense.py prefers; a missing one
|
||||
# would badge the real server path as "mock".
|
||||
self.descriptor = VoiceDescriptor(
|
||||
mode="server",
|
||||
sr_available=True,
|
||||
tts_available=True,
|
||||
hint="server STT/TTS via AI_VOICE_BASE_URL",
|
||||
)
|
||||
|
||||
def _headers(self) -> dict[str, str]:
|
||||
headers: dict[str, str] = {}
|
||||
if self._api_key:
|
||||
headers["Authorization"] = f"Bearer {self._api_key}"
|
||||
return headers
|
||||
|
||||
def _sanitize(self, exc: Exception) -> RuntimeError:
|
||||
text = str(exc)
|
||||
if self._api_key and self._api_key in text:
|
||||
text = text.replace(self._api_key, "[REDACTED]")
|
||||
return RuntimeError(f"voice provider error: {text}")
|
||||
|
||||
async def transcribe(self, audio: bytes, fmt: str) -> TranscriptSegment:
|
||||
"""STT: multipart upload (`file` + `model`) → TranscriptSegment.
|
||||
|
||||
`fmt` is a bare extension ('wav' | 'webm' | 'mp3') — the defense
|
||||
route strips codec params before this call (D-041).
|
||||
"""
|
||||
files = {"file": (f"answer.{fmt}", audio, f"audio/{fmt}")}
|
||||
data = {"model": self._stt_model, "response_format": "json"}
|
||||
try:
|
||||
resp = await self._client.post(
|
||||
f"{self._base_url}/audio/transcriptions",
|
||||
files=files,
|
||||
data=data,
|
||||
headers=self._headers(),
|
||||
)
|
||||
resp.raise_for_status()
|
||||
body = resp.json()
|
||||
except httpx.HTTPError as exc:
|
||||
raise self._sanitize(exc) from exc
|
||||
if not isinstance(body, dict):
|
||||
# P2 (verifier): a non-dict 200 body is a contract break —
|
||||
# context-wrap instead of a raw AttributeError.
|
||||
raise RuntimeError(
|
||||
"voice provider error: unexpected transcription response shape"
|
||||
)
|
||||
text = str(body.get("text", "")).strip()
|
||||
if not text:
|
||||
# 200 with an empty transcript is a provider contract break —
|
||||
# TranscriptSegment(min_length=1) would raise a bare pydantic
|
||||
# error; wrap it with provider context instead.
|
||||
raise RuntimeError("voice provider error: empty transcription")
|
||||
return TranscriptSegment(text=text)
|
||||
|
||||
def synthesize(
|
||||
self, text: str, voice: str = "default"
|
||||
) -> AsyncIterator[bytes]:
|
||||
"""TTS: JSON body → raw audio byte stream.
|
||||
|
||||
OpenAI's TTS caps `input` at 4096 chars; examiner questions are
|
||||
short, but enforce the guard so a long question fails loudly at the
|
||||
seam instead of as an opaque provider 400.
|
||||
"""
|
||||
return self._synthesize_stream(text, voice)
|
||||
|
||||
async def _synthesize_stream(
|
||||
self, text: str, voice: str
|
||||
) -> AsyncIterator[bytes]:
|
||||
if len(text) > 4096:
|
||||
raise RuntimeError(
|
||||
f"voice provider error: TTS input exceeds 4096 chars ({len(text)})"
|
||||
)
|
||||
payload = {
|
||||
"model": self._tts_model,
|
||||
"input": text,
|
||||
"voice": voice if voice != "default" else self._tts_voice,
|
||||
"response_format": self._tts_format,
|
||||
}
|
||||
try:
|
||||
async with self._client.stream(
|
||||
"POST",
|
||||
f"{self._base_url}/audio/speech",
|
||||
content=json.dumps(payload),
|
||||
headers={**self._headers(), "Content-Type": "application/json"},
|
||||
) as resp:
|
||||
resp.raise_for_status()
|
||||
async for chunk in resp.aiter_bytes():
|
||||
if chunk:
|
||||
yield chunk
|
||||
except httpx.HTTPError as exc:
|
||||
raise self._sanitize(exc) from exc
|
||||
@@ -1,11 +0,0 @@
|
||||
{
|
||||
"name": "@nextcraft/ai-service",
|
||||
"private": true,
|
||||
"version": "0.2.0",
|
||||
"scripts": {
|
||||
"dev": "bash scripts/dev.sh",
|
||||
"test": "bash scripts/test.sh",
|
||||
"bootstrap": "bash scripts/bootstrap.sh",
|
||||
"lint": "bash scripts/lint.sh"
|
||||
}
|
||||
}
|
||||
@@ -1,47 +0,0 @@
|
||||
[build-system]
|
||||
requires = ["setuptools>=68"]
|
||||
build-backend = "setuptools.build_meta"
|
||||
|
||||
[project]
|
||||
name = "nextcraft-ai-service"
|
||||
version = "0.2.0"
|
||||
description = "Nextcraft AI tutor service — six LLM agents behind a provider-agnostic layer"
|
||||
requires-python = ">=3.11"
|
||||
dependencies = [
|
||||
"fastapi>=0.141,<0.142",
|
||||
"uvicorn>=0.52,<0.53",
|
||||
"pydantic>=2.13,<2.14",
|
||||
"pydantic-settings>=2.15,<2.16",
|
||||
"httpx>=0.28,<0.29",
|
||||
"sse-starlette>=3.4,<3.5",
|
||||
"sqlmodel>=0.0.24,<0.1",
|
||||
"sqlalchemy>=2.0,<2.1",
|
||||
"websockets>=13,<16",
|
||||
"aiofiles>=24.1,<26",
|
||||
# POST /v1/defense/{id}/answer multipart audio (REQ-3-006): FastAPI
|
||||
# form/File parsing requires python-multipart at runtime.
|
||||
"python-multipart>=0.0.32,<0.1",
|
||||
]
|
||||
|
||||
[project.optional-dependencies]
|
||||
dev = [
|
||||
"pytest>=9.1,<10",
|
||||
"pytest-asyncio>=1.4,<2",
|
||||
"ruff>=0.14",
|
||||
]
|
||||
|
||||
[tool.setuptools.packages.find]
|
||||
include = ["ai_service*"]
|
||||
|
||||
[tool.pytest.ini_options]
|
||||
asyncio_mode = "auto"
|
||||
testpaths = ["tests"]
|
||||
|
||||
[tool.ruff]
|
||||
line-length = 100
|
||||
target-version = "py311"
|
||||
|
||||
[tool.ruff.lint]
|
||||
select = ["E", "F", "W", "I", "UP", "B"]
|
||||
# B008: Depends() in argument defaults is the idiomatic FastAPI DI pattern
|
||||
ignore = ["B008"]
|
||||
@@ -1,75 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Idempotent bootstrap: create venv + install deps.
|
||||
# Handles Debian/Ubuntu systems without python3-venv/ensurepip via --without-pip + get-pip.
|
||||
# v2 (v0.3.5): recovers from a poisoned partial .venv left by a failed earlier
|
||||
# attempt, cleans before each retry, and dies with a distro-specific fix hint
|
||||
# when venv creation is impossible (e.g. missing python3.XX-venv package).
|
||||
set -euo pipefail
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
APP_DIR="$(cd "$SCRIPT_DIR/.." && pwd)"
|
||||
VENV="$APP_DIR/.venv"
|
||||
|
||||
venv_usable() {
|
||||
[ -x "$VENV/bin/python3" ]
|
||||
}
|
||||
|
||||
rm_broken_venv() {
|
||||
echo "bootstrap: removing broken partial .venv from a failed earlier attempt" >&2
|
||||
rm -rf "$VENV"
|
||||
}
|
||||
|
||||
mkdir -p "$HOME/.cache/ciagent"
|
||||
|
||||
if venv_usable && [ ! -x "$VENV/bin/pip" ]; then
|
||||
# A usable python3 without pip means the --without-pip fallback half-ran and
|
||||
# the get-pip step never completed: start over cleanly.
|
||||
rm_broken_venv
|
||||
fi
|
||||
|
||||
if ! venv_usable; then
|
||||
if [ -d "$VENV" ]; then
|
||||
# Directory exists but no working python3: remains of a crashed venv create.
|
||||
rm_broken_venv
|
||||
fi
|
||||
if python3 -m venv "$VENV" 2>/tmp/venv-create.err; then
|
||||
:
|
||||
else
|
||||
rm -rf "$VENV"
|
||||
if python3 -m venv --without-pip "$VENV" 2>>/tmp/venv-create.err; then
|
||||
:
|
||||
else
|
||||
rm -rf "$VENV"
|
||||
PYVER="$(python3 -c 'import sys; print("%d.%d" % sys.version_info[:2])' 2>/dev/null || true)"
|
||||
PKG="python3-venv"
|
||||
[ -n "$PYVER" ] && PKG="python${PYVER}-venv"
|
||||
echo "bootstrap: could not create a virtual environment." >&2
|
||||
echo " python3 reported:" >&2
|
||||
sed 's/^/ /' /tmp/venv-create.err >&2 || true
|
||||
echo " fix (Debian/Ubuntu): install the venv support package, then re-run nextcraft bootstrap:" >&2
|
||||
echo " apt install $PKG" >&2
|
||||
exit 1
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
|
||||
if [ ! -x "$VENV/bin/pip" ]; then
|
||||
GET_PIP="$HOME/.cache/ciagent/get-pip.py"
|
||||
if [ ! -f "$GET_PIP" ]; then
|
||||
if ! curl -sSf --max-time 60 https://bootstrap.pypa.io/get-pip.py -o "$GET_PIP"; then
|
||||
rm -rf "$VENV"
|
||||
echo "bootstrap: get-pip.py download failed (no network?)." >&2
|
||||
echo " fix: restore network access and re-run nextcraft bootstrap" >&2
|
||||
exit 1
|
||||
fi
|
||||
fi
|
||||
if ! "$VENV/bin/python3" "$GET_PIP" --quiet; then
|
||||
rm -rf "$VENV"
|
||||
echo "bootstrap: pip installation into the venv failed." >&2
|
||||
echo " fix: re-run nextcraft bootstrap (the venv was cleaned; this retry is safe)" >&2
|
||||
exit 1
|
||||
fi
|
||||
fi
|
||||
|
||||
"$VENV/bin/pip" install --quiet --upgrade pip
|
||||
"$VENV/bin/pip" install --quiet -e "$APP_DIR[dev]"
|
||||
echo "bootstrap complete: $VENV"
|
||||
@@ -1,40 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Dev server: export secrets (if present) then run uvicorn.
|
||||
# Binds 0.0.0.0 by default so the stack is reachable from other machines
|
||||
# (v0.3.5 network mode) — set AI_HOST=127.0.0.1 in .env to revert to loopback.
|
||||
set -euo pipefail
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
APP_DIR="$(cd "$SCRIPT_DIR/.." && pwd)"
|
||||
REPO_ROOT="$(cd "$APP_DIR/../.." && pwd)"
|
||||
VENV="$APP_DIR/.venv"
|
||||
|
||||
if [ ! -x "$VENV/bin/uvicorn" ]; then
|
||||
echo "venv missing — run nextcraft bootstrap first (or: bash scripts/bootstrap.sh)" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
SECRETS="$REPO_ROOT/.ciagent/.env.secrets"
|
||||
if [ -f "$SECRETS" ]; then
|
||||
while IFS='=' read -r key value; do
|
||||
case "$key" in
|
||||
OLLAMA_API_KEY) export AI_OLLAMA_CLOUD_API_KEY="$value" ;;
|
||||
OLLAMA_BASE_URL) export AI_OLLAMA_CLOUD_BASE_URL="$value" ;;
|
||||
AI_TUTOR_MODEL) export AI_MODEL="$value" ;;
|
||||
esac
|
||||
done < "$SECRETS"
|
||||
fi
|
||||
|
||||
ENV_FILE="$APP_DIR/.env"
|
||||
if [ -f "$ENV_FILE" ]; then
|
||||
while IFS='=' read -r key value; do
|
||||
case "$key" in
|
||||
AI_HOST|AI_PORT|AI_CORS_ORIGINS) export "$key=$value" ;;
|
||||
esac
|
||||
done < "$ENV_FILE"
|
||||
fi
|
||||
|
||||
HOST="${AI_HOST:-0.0.0.0}"
|
||||
PORT="${AI_PORT:-8420}"
|
||||
|
||||
cd "$APP_DIR"
|
||||
exec "$VENV/bin/uvicorn" ai_service.main:app --reload --host "$HOST" --port "$PORT"
|
||||
@@ -1,14 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Lint: ruff check over the ai-service tree.
|
||||
set -euo pipefail
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
APP_DIR="$(cd "$SCRIPT_DIR/.." && pwd)"
|
||||
VENV="$APP_DIR/.venv"
|
||||
|
||||
if [ ! -x "$VENV/bin/ruff" ]; then
|
||||
echo "venv missing — run scripts/bootstrap.sh first" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
cd "$APP_DIR"
|
||||
exec "$VENV/bin/ruff" check .
|
||||
@@ -1,803 +0,0 @@
|
||||
#!/usr/bin/env python3
|
||||
"""sandbox-agent — stdlib-only in-sandbox telemetry capture agent (REQ-3-003).
|
||||
|
||||
D-031: this file is copied into the sandbox namespace and runs against the
|
||||
system Python — no third-party packages are importable there, so this module
|
||||
depends on the standard library ONLY (the test suite enforces this with an
|
||||
AST scan of the file).
|
||||
|
||||
What it does:
|
||||
* wraps a non-interactive `/bin/sh` REPL: each stdin line is executed via
|
||||
`sh -c` inside the workspace and reported as `stdin` -> `command` ->
|
||||
`stdout` -> `run_result`/`test_result` events;
|
||||
* polls the workspace tree (~250 ms) and emits `file_diff` events
|
||||
(created/modified/deleted with unified diffs) plus periodic `activity`
|
||||
heartbeats;
|
||||
* streams events to ai-service as TelemetryEvent-shaped JSON frames over a
|
||||
raw-socket RFC 6455 WebSocket client (no `websockets` package exists in
|
||||
the namespace — the client handshake + frame codec is implemented here);
|
||||
* at-least-once delivery (D-026): every event is appended to an fsync'd
|
||||
JSONL spool file inside the workdir BEFORE any send attempt; on
|
||||
disconnect the spool grows; after reconnect (exponential backoff) the
|
||||
spool is flushed oldest-first. The server dedups on (learner, task, seq)
|
||||
so replayed duplicates are harmless — loss is not tolerated.
|
||||
|
||||
Configured entirely through env baked at spawn time:
|
||||
NC_LEARNER_ID / NC_TASK_ID / NC_INGEST_URL / NC_SANDBOX_ID (required)
|
||||
NC_WORKSPACE workspace root to watch/run in (default: cwd)
|
||||
NC_SPOOL spool path (default: <workspace>/.nc-agent/spool.jsonl)
|
||||
NC_POLL_INTERVAL_S / NC_ACTIVITY_INTERVAL_S / NC_COMMAND_TIMEOUT_S
|
||||
NC_BACKOFF_BASE_S / NC_BACKOFF_MAX_S (optional knobs)
|
||||
|
||||
Sequencing survives process restarts (incl. SIGKILL): on boot the spool is
|
||||
replayed into the pending queue and `seq` resumes at max(spooled seq) + 1.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import base64
|
||||
import difflib
|
||||
import hashlib
|
||||
import json
|
||||
import os
|
||||
import secrets
|
||||
import socket
|
||||
import ssl
|
||||
import struct
|
||||
import subprocess
|
||||
import sys
|
||||
import threading
|
||||
import time
|
||||
import urllib.parse
|
||||
from collections import deque
|
||||
from collections.abc import Mapping
|
||||
from dataclasses import dataclass
|
||||
from datetime import UTC, datetime
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
_WS_GUID = "258EAFA5-E914-47DA-95CA-C5AB0DC85B11"
|
||||
_EVENT_KINDS = frozenset(
|
||||
{"command", "file_diff", "run_result", "test_result", "activity", "stdin", "stdout"}
|
||||
)
|
||||
_AGENT_DIR_PREFIX = ".nc-" # agent-private paths (spool) are excluded from watching
|
||||
_MAX_DIFF_BYTES = 64 * 1024 # files larger than this are reported truncated, no diff
|
||||
_MAX_OUTPUT_CHARS = 64 * 1024 # captured stdout/stderr tail cap per command
|
||||
_HANDSHAKE_MAX_BYTES = 64 * 1024
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- config
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class AgentConfig:
|
||||
"""Runtime configuration, normally built from `NC_*` env baked at spawn."""
|
||||
|
||||
learner_id: str
|
||||
task_id: str
|
||||
ingest_url: str
|
||||
sandbox_id: str
|
||||
workspace: Path
|
||||
spool_path: Path
|
||||
poll_interval_s: float = 0.25
|
||||
activity_interval_s: float = 5.0
|
||||
command_timeout_s: float = 30.0
|
||||
backoff_base_s: float = 0.25
|
||||
backoff_max_s: float = 8.0
|
||||
#: Spool bound (G-14): explicit cap where none existed. Worst case
|
||||
#: ~64KB/line (diff cap) * SPOOL_MAX_LINES must stay well under the
|
||||
#: G-2 512MB workdir sweep: 4096 * 64KB = 256MB (half the budget).
|
||||
spool_max_lines: int = 4096
|
||||
|
||||
def __post_init__(self) -> None:
|
||||
for name in ("learner_id", "task_id", "ingest_url", "sandbox_id"):
|
||||
if not getattr(self, name):
|
||||
raise ValueError(f"missing required config: NC_{name.upper()}")
|
||||
|
||||
@classmethod
|
||||
def from_env(cls, env: Mapping[str, str] | None = None) -> AgentConfig:
|
||||
src = os.environ if env is None else env
|
||||
workspace = Path(src.get("NC_WORKSPACE") or os.getcwd()).resolve()
|
||||
return cls(
|
||||
learner_id=src.get("NC_LEARNER_ID", ""),
|
||||
task_id=src.get("NC_TASK_ID", ""),
|
||||
ingest_url=src.get("NC_INGEST_URL", ""),
|
||||
sandbox_id=src.get("NC_SANDBOX_ID", ""),
|
||||
workspace=workspace,
|
||||
spool_path=Path(
|
||||
src.get("NC_SPOOL") or (workspace / ".nc-agent" / "spool.jsonl")
|
||||
),
|
||||
poll_interval_s=float(src.get("NC_POLL_INTERVAL_S", "0.25")),
|
||||
activity_interval_s=float(src.get("NC_ACTIVITY_INTERVAL_S", "5.0")),
|
||||
command_timeout_s=float(src.get("NC_COMMAND_TIMEOUT_S", "30.0")),
|
||||
backoff_base_s=float(src.get("NC_BACKOFF_BASE_S", "0.25")),
|
||||
backoff_max_s=float(src.get("NC_BACKOFF_MAX_S", "8.0")),
|
||||
)
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- spool
|
||||
|
||||
|
||||
class Spool:
|
||||
"""Append-only JSONL spool with per-append fsync (survives SIGKILL).
|
||||
|
||||
`rewrite` swaps in a compacted file atomically (tmp file + os.replace).
|
||||
Lines are stored without trailing newlines in memory, one per line on disk.
|
||||
"""
|
||||
|
||||
def __init__(self, path: Path) -> None:
|
||||
self._path = path
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
@property
|
||||
def path(self) -> Path:
|
||||
return self._path
|
||||
|
||||
def append(self, line: str) -> None:
|
||||
with self._path.open("a", encoding="utf-8") as fh:
|
||||
fh.write(line + "\n")
|
||||
fh.flush()
|
||||
os.fsync(fh.fileno())
|
||||
|
||||
def read_all(self) -> list[str]:
|
||||
if not self._path.exists():
|
||||
return []
|
||||
with self._path.open("r", encoding="utf-8") as fh:
|
||||
return [line.rstrip("\n") for line in fh if line.strip()]
|
||||
|
||||
def rewrite(self, lines: list[str]) -> None:
|
||||
tmp = self._path.with_name(self._path.name + ".tmp")
|
||||
with tmp.open("w", encoding="utf-8") as fh:
|
||||
for line in lines:
|
||||
fh.write(line + "\n")
|
||||
fh.flush()
|
||||
os.fsync(fh.fileno())
|
||||
os.replace(tmp, self._path)
|
||||
|
||||
|
||||
def _line_seq(line: str) -> int | None:
|
||||
"""Best-effort seq extraction from a spool line (None when unparseable).
|
||||
|
||||
The event's seq is a top-level wire field (`_next_event`). Used only for
|
||||
ack trimming; an unparseable line is retained (never dropped by the ack
|
||||
path — the overflow bound is the only dropper).
|
||||
"""
|
||||
try:
|
||||
seq = json.loads(line).get("seq")
|
||||
return seq if isinstance(seq, int) else None
|
||||
except (ValueError, AttributeError):
|
||||
return None
|
||||
|
||||
|
||||
# --------------------------------------------------------------- websocket codec
|
||||
|
||||
|
||||
def _encode_frame(opcode: int, payload: bytes) -> bytes:
|
||||
"""RFC 6455 client frame: FIN set, always masked (servers require it)."""
|
||||
header = bytearray([0x80 | opcode])
|
||||
n = len(payload)
|
||||
if n < 126:
|
||||
header.append(0x80 | n)
|
||||
elif n < 65536:
|
||||
header.append(0x80 | 126)
|
||||
header += struct.pack("!H", n)
|
||||
else:
|
||||
header.append(0x80 | 127)
|
||||
header += struct.pack("!Q", n)
|
||||
mask = secrets.token_bytes(4)
|
||||
header += mask
|
||||
masked = bytes(b ^ mask[i % 4] for i, b in enumerate(payload))
|
||||
return bytes(header) + masked
|
||||
|
||||
|
||||
class WsConnection:
|
||||
"""Minimal blocking RFC 6455 client over a raw socket (stdlib only)."""
|
||||
|
||||
def __init__(self, sock: socket.socket) -> None:
|
||||
self._sock = sock
|
||||
self._write_lock = threading.Lock()
|
||||
|
||||
@classmethod
|
||||
def connect(cls, url: str, timeout_s: float = 5.0) -> WsConnection:
|
||||
parts = urllib.parse.urlsplit(url)
|
||||
if parts.scheme not in ("ws", "wss"):
|
||||
raise ValueError(f"unsupported scheme in NC_INGEST_URL: {parts.scheme!r}")
|
||||
host = parts.hostname or "localhost"
|
||||
port = parts.port or (443 if parts.scheme == "wss" else 80)
|
||||
path = parts.path or "/"
|
||||
if parts.query:
|
||||
path += "?" + parts.query
|
||||
|
||||
sock = socket.create_connection((host, port), timeout=timeout_s)
|
||||
if parts.scheme == "wss":
|
||||
sock = ssl.create_default_context().wrap_socket(sock, server_hostname=host)
|
||||
|
||||
key = base64.b64encode(secrets.token_bytes(16)).decode("ascii")
|
||||
request = (
|
||||
f"GET {path} HTTP/1.1\r\n"
|
||||
f"Host: {host}:{port}\r\n"
|
||||
"Upgrade: websocket\r\n"
|
||||
"Connection: Upgrade\r\n"
|
||||
f"Sec-WebSocket-Key: {key}\r\n"
|
||||
"Sec-WebSocket-Version: 13\r\n\r\n"
|
||||
)
|
||||
sock.sendall(request.encode("ascii"))
|
||||
response = cls._read_http_response(sock)
|
||||
cls._validate_handshake(response, key)
|
||||
return cls(sock)
|
||||
|
||||
@staticmethod
|
||||
def _read_http_response(sock: socket.socket) -> bytes:
|
||||
buf = b""
|
||||
while b"\r\n\r\n" not in buf:
|
||||
chunk = sock.recv(4096)
|
||||
if not chunk:
|
||||
raise ConnectionError("server closed during WebSocket handshake")
|
||||
buf += chunk
|
||||
if len(buf) > _HANDSHAKE_MAX_BYTES:
|
||||
raise ConnectionError("handshake response exceeded size cap")
|
||||
return buf.split(b"\r\n\r\n", 1)[0]
|
||||
|
||||
@staticmethod
|
||||
def _validate_handshake(response: bytes, key: str) -> None:
|
||||
head = response.decode("latin-1")
|
||||
lines = head.split("\r\n")
|
||||
if not lines or " 101" not in lines[0]:
|
||||
raise ConnectionError(f"handshake rejected: {lines[0] if lines else '<empty>'}")
|
||||
headers = {}
|
||||
for line in lines[1:]:
|
||||
if ":" in line:
|
||||
name, _, value = line.partition(":")
|
||||
headers[name.strip().lower()] = value.strip()
|
||||
expect = base64.b64encode(
|
||||
hashlib.sha1((key + _WS_GUID).encode("ascii")).digest()
|
||||
).decode("ascii")
|
||||
if headers.get("sec-websocket-accept") != expect:
|
||||
raise ConnectionError("bad Sec-WebSocket-Accept in handshake response")
|
||||
|
||||
# -- send ------------------------------------------------------------
|
||||
def send_text(self, text: str) -> None:
|
||||
with self._write_lock:
|
||||
self._sock.sendall(_encode_frame(0x1, text.encode("utf-8")))
|
||||
|
||||
def _send_frame(self, opcode: int, payload: bytes) -> None:
|
||||
with self._write_lock:
|
||||
self._sock.sendall(_encode_frame(opcode, payload))
|
||||
|
||||
# -- receive ---------------------------------------------------------
|
||||
def recv_message(self, timeout_s: float) -> tuple[int, bytes] | None:
|
||||
"""Return (opcode, payload) for a data/close frame, or None on timeout.
|
||||
|
||||
Ping frames are answered with pong internally and never surfaced;
|
||||
pongs are swallowed. Fragmented messages are reassembled. Raises
|
||||
ConnectionError/OSError when the socket breaks.
|
||||
"""
|
||||
deadline = time.monotonic() + timeout_s
|
||||
fragments = bytearray()
|
||||
frag_opcode = 0
|
||||
while True:
|
||||
frame = self._recv_one_frame(deadline)
|
||||
if frame is None:
|
||||
return None
|
||||
fin, opcode, payload = frame
|
||||
if opcode == 0x9: # ping
|
||||
self._send_frame(0xA, payload)
|
||||
continue
|
||||
if opcode == 0xA: # pong
|
||||
continue
|
||||
if opcode == 0x0: # continuation
|
||||
fragments += payload
|
||||
else:
|
||||
fragments = bytearray(payload)
|
||||
frag_opcode = opcode
|
||||
if fin:
|
||||
return frag_opcode, bytes(fragments)
|
||||
|
||||
def _recv_one_frame(self, deadline: float) -> tuple[bool, int, bytes] | None:
|
||||
header = self._read_exact(2, deadline)
|
||||
if header is None:
|
||||
return None
|
||||
b0, b1 = header[0], header[1]
|
||||
fin = bool(b0 & 0x80)
|
||||
opcode = b0 & 0x0F
|
||||
length = b1 & 0x7F
|
||||
if length == 126:
|
||||
ext = self._read_exact(2, deadline)
|
||||
if ext is None:
|
||||
return None
|
||||
length = struct.unpack("!H", ext)[0]
|
||||
elif length == 127:
|
||||
ext = self._read_exact(8, deadline)
|
||||
if ext is None:
|
||||
return None
|
||||
length = struct.unpack("!Q", ext)[0]
|
||||
mask = self._read_exact(4, deadline) if (b1 & 0x80) else b""
|
||||
if mask is None:
|
||||
return None
|
||||
payload = self._read_exact(length, deadline) if length else b""
|
||||
if payload is None:
|
||||
return None
|
||||
if mask:
|
||||
payload = bytes(b ^ mask[i % 4] for i, b in enumerate(payload))
|
||||
return fin, opcode, payload
|
||||
|
||||
def _read_exact(self, n: int, deadline: float) -> bytes | None:
|
||||
buf = bytearray()
|
||||
while len(buf) < n:
|
||||
remaining = deadline - time.monotonic()
|
||||
if remaining <= 0:
|
||||
return None
|
||||
self._sock.settimeout(remaining)
|
||||
try:
|
||||
chunk = self._sock.recv(n - len(buf))
|
||||
except TimeoutError:
|
||||
return None
|
||||
if not chunk:
|
||||
raise ConnectionError("peer closed the WebSocket connection")
|
||||
buf += chunk
|
||||
return bytes(buf)
|
||||
|
||||
def close(self) -> None:
|
||||
try:
|
||||
self._sock.close()
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- agent
|
||||
|
||||
|
||||
class Agent:
|
||||
"""Wires capture (shell + workspace watcher) to the framed event stream."""
|
||||
|
||||
def __init__(self, config: AgentConfig) -> None:
|
||||
self.config = config
|
||||
self._spool = Spool(config.spool_path)
|
||||
self._pending: deque[str] = deque()
|
||||
self._seq = 0
|
||||
# RLock (D-3 verifier fix): emit() holds this across sends; a send
|
||||
# failure drops the conn, and _drop_conn -> replay_margin re-enters
|
||||
# the same lock. A plain Lock deadlocked the emitting thread there.
|
||||
self._emit_lock = threading.RLock() # serializes seq + spool + flush
|
||||
self._conn_lock = threading.Lock() # guards _conn swaps
|
||||
self._conn: WsConnection | None = None
|
||||
self._last_sent: str | None = None # one-line replay margin, see below
|
||||
self._stop = threading.Event()
|
||||
self._threads: list[threading.Thread] = []
|
||||
self._baseline: dict[str, tuple[int, int, str | None]] = {}
|
||||
# Spool bound (G-14): explicit cap where none existed. Worst case
|
||||
# ~64KB/line (diff cap) x SPOOL_MAX_LINES must stay well under the
|
||||
# G-2 512MB workdir sweep; overflow drops OLDEST lines with a
|
||||
# counter — by-design gap creation, so the trace goes ungradable
|
||||
# (G-4) instead of silently truncated-but-gradable.
|
||||
self._dropped_overflow = 0
|
||||
self._enforce_spool_bound_locked()
|
||||
self._resume_from_spool()
|
||||
|
||||
# -- durability ------------------------------------------------------
|
||||
def _resume_from_spool(self) -> None:
|
||||
highest = -1
|
||||
for line in self._spool.read_all():
|
||||
self._pending.append(line)
|
||||
try:
|
||||
seq = int(json.loads(line).get("seq", -1))
|
||||
except (ValueError, AttributeError):
|
||||
continue
|
||||
highest = max(highest, seq)
|
||||
self._seq = highest + 1
|
||||
|
||||
# -- event construction ---------------------------------------------
|
||||
def _wire_frame(self, spooled_line: str) -> str:
|
||||
"""Spool format -> wire format: strip URL-owned identity fields.
|
||||
|
||||
The spool keeps full events (local durability + restart recovery).
|
||||
The ingest endpoint binds identity at the WS handshake (query params)
|
||||
and rejects frames carrying learner_id/task_id (`extra="forbid"`
|
||||
anti-spoofing), so the wire frame carries only seq/kind/payload/ts.
|
||||
"""
|
||||
import json as _json
|
||||
|
||||
full = _json.loads(spooled_line)
|
||||
wire = {
|
||||
k: full[k]
|
||||
for k in ("seq", "kind", "payload", "ts")
|
||||
}
|
||||
if full.get("sandbox_id"):
|
||||
wire["sandbox_id"] = full["sandbox_id"]
|
||||
return _json.dumps(wire)
|
||||
|
||||
def _next_event(self, kind: str, payload: dict[str, Any]) -> dict[str, Any]:
|
||||
if kind not in _EVENT_KINDS:
|
||||
raise ValueError(f"unknown event kind: {kind!r}")
|
||||
event = {
|
||||
"learner_id": self.config.learner_id,
|
||||
"task_id": self.config.task_id,
|
||||
"seq": self._seq,
|
||||
"kind": kind,
|
||||
"payload": payload,
|
||||
"ts": datetime.now(UTC).isoformat(),
|
||||
"sandbox_id": self.config.sandbox_id,
|
||||
}
|
||||
self._seq += 1
|
||||
return event
|
||||
|
||||
# -- emission / flush (D-026) ----------------------------------------
|
||||
def emit(self, kind: str, payload: dict[str, Any]) -> dict[str, Any]:
|
||||
"""Spool-then-send. Never blocks on reconnect; loss is impossible."""
|
||||
with self._emit_lock:
|
||||
line = json.dumps(self._next_event(kind, payload))
|
||||
self._spool.append(line) # durable BEFORE any send attempt
|
||||
self._pending.append(line)
|
||||
self._enforce_spool_bound_locked()
|
||||
self._flush_locked()
|
||||
return json.loads(line)
|
||||
|
||||
def _enforce_spool_bound_locked(self) -> None:
|
||||
"""Drop OLDEST spooled lines past `spool_max_lines` (G-14).
|
||||
|
||||
Overflow is by-design gap creation: the dropped seqs become
|
||||
permanent gaps server-side, the gap path flags the trace, and the
|
||||
grader refuses it (G-4) — never a silently-truncated-but-gradable
|
||||
trace. Caller must hold `_emit_lock`. `_dropped_overflow` is the
|
||||
observable counter (surfaced in the final stop status event).
|
||||
"""
|
||||
lines = self._spool.read_all()
|
||||
overflow = len(lines) - self.config.spool_max_lines
|
||||
if overflow <= 0:
|
||||
return
|
||||
self._dropped_overflow += overflow
|
||||
self._spool.rewrite(lines[overflow:])
|
||||
# Pending may reference dropped lines; they replay as no-ops (server
|
||||
# dedup) but trimming them keeps the replay window honest.
|
||||
dropped = set(lines[:overflow])
|
||||
self._pending = deque(ln for ln in self._pending if ln not in dropped)
|
||||
|
||||
def _flush_locked(self) -> None:
|
||||
conn = self._current_conn()
|
||||
while self._pending and conn is not None:
|
||||
line = self._pending[0]
|
||||
try:
|
||||
conn.send_text(self._wire_frame(line))
|
||||
except (ConnectionError, OSError):
|
||||
self._drop_conn()
|
||||
return
|
||||
self._pending.popleft()
|
||||
self._last_sent = line
|
||||
# D-045/D-1 (verifier fix): the spool is NEVER compacted below the
|
||||
# unacked window. A send into a silently-dead socket "succeeds" at
|
||||
# TCP level — those lines may be lost in flight — so they stay in
|
||||
# the spool until the server's seq_ack proves durable storage
|
||||
# (trim_to_ack is the ONLY spool shrinker besides the overflow
|
||||
# bound). Replays are harmless: the server dedups on
|
||||
# (learner, task, seq).
|
||||
|
||||
def replay_margin(self) -> None:
|
||||
"""Requeue every unacked spooled line after a detected disconnect.
|
||||
|
||||
D-045/D-1 (verifier fix): the pre-ack one-line margin could not cover
|
||||
a multi-frame in-flight window — a burst accepted by a dying socket
|
||||
popped N lines from pending while the spool had been compacted to the
|
||||
last one, permanently losing lines 1..N-1. The spool now retains
|
||||
everything unacked, so replay requeues the full unacked window;
|
||||
server-side dedup absorbs the duplicates.
|
||||
"""
|
||||
with self._emit_lock:
|
||||
if self._pending:
|
||||
return # mid-flush caller holds the pending queue intact
|
||||
self._pending = deque(self._spool.read_all())
|
||||
self._last_sent = None
|
||||
|
||||
def trim_to_ack(self, ack_seq: int) -> None:
|
||||
"""Drop every spooled/pending line with seq <= ack_seq (D-045).
|
||||
|
||||
Advisory server hint: `seq_ack` carries the durable latest_seq, so
|
||||
everything up to and including it is stored server-side and dedup
|
||||
absorbs nothing on replay. Runs under `_emit_lock` — ordered against
|
||||
concurrent emit()/flush, and the rewrite is atomic (Spool.rewrite).
|
||||
Lines without a parseable seq are retained; the overflow bound is
|
||||
the only dropper of unparseable lines.
|
||||
"""
|
||||
with self._emit_lock:
|
||||
spooled = self._spool.read_all()
|
||||
kept_pending = [
|
||||
ln for ln in self._pending if (s := _line_seq(ln)) is None or s > ack_seq
|
||||
]
|
||||
kept_spooled = [
|
||||
ln for ln in spooled if (s := _line_seq(ln)) is None or s > ack_seq
|
||||
]
|
||||
if len(kept_pending) != len(self._pending) or len(kept_spooled) != len(spooled):
|
||||
self._pending = deque(kept_pending)
|
||||
self._spool.rewrite(kept_spooled)
|
||||
if self._last_sent is not None:
|
||||
s = _line_seq(self._last_sent)
|
||||
if s is not None and s <= ack_seq:
|
||||
self._last_sent = None
|
||||
|
||||
# -- connection supervision ------------------------------------------
|
||||
def _current_conn(self) -> WsConnection | None:
|
||||
with self._conn_lock:
|
||||
return self._conn
|
||||
|
||||
def _set_conn(self, conn: WsConnection | None) -> None:
|
||||
with self._conn_lock:
|
||||
self._conn = conn
|
||||
|
||||
def _drop_conn(self) -> None:
|
||||
conn = self._current_conn()
|
||||
self._set_conn(None)
|
||||
if conn is not None:
|
||||
conn.close()
|
||||
self.replay_margin()
|
||||
|
||||
def is_connected(self) -> bool:
|
||||
return self._current_conn() is not None
|
||||
|
||||
def wait_connected(self, timeout_s: float) -> bool:
|
||||
return self._wait_for(lambda: self.is_connected(), timeout_s)
|
||||
|
||||
def wait_disconnected(self, timeout_s: float) -> bool:
|
||||
return self._wait_for(lambda: not self.is_connected(), timeout_s)
|
||||
|
||||
def _wait_for(self, pred: Any, timeout_s: float) -> bool:
|
||||
deadline = time.monotonic() + timeout_s
|
||||
while time.monotonic() < deadline:
|
||||
if pred():
|
||||
return True
|
||||
time.sleep(0.02)
|
||||
return pred()
|
||||
|
||||
def _supervisor_loop(self) -> None:
|
||||
"""Maintain the WS connection: connect, flush backlog, read, backoff."""
|
||||
backoff = self.config.backoff_base_s
|
||||
while not self._stop.is_set():
|
||||
if self._current_conn() is None:
|
||||
try:
|
||||
conn = WsConnection.connect(self.config.ingest_url)
|
||||
except (ConnectionError, OSError, ValueError, TimeoutError):
|
||||
self._stop.wait(backoff)
|
||||
backoff = min(self.config.backoff_max_s, backoff * 2)
|
||||
continue
|
||||
self._set_conn(conn)
|
||||
self._last_sent = None
|
||||
backoff = self.config.backoff_base_s
|
||||
with self._emit_lock: # ordered against concurrent emit()s
|
||||
self._flush_locked()
|
||||
else:
|
||||
conn = self._current_conn()
|
||||
if conn is None:
|
||||
continue
|
||||
try:
|
||||
frame = conn.recv_message(timeout_s=1.0)
|
||||
except (ConnectionError, OSError):
|
||||
self._drop_conn()
|
||||
continue
|
||||
if frame is None:
|
||||
continue
|
||||
opcode, payload = frame
|
||||
if opcode == 0x1: # server text frame — parse advisory envelopes
|
||||
self._handle_server_text(payload)
|
||||
continue
|
||||
if opcode == 0x8: # server close frame
|
||||
self._drop_conn()
|
||||
|
||||
def _handle_server_text(self, payload: bytes) -> None:
|
||||
"""Consume server->agent envelopes (advisory; never fatal).
|
||||
|
||||
`seq_ack` (D-045): the server's durable latest_seq — trims the
|
||||
spool/pending to `seq > ack`, bounding the replay margin to the
|
||||
in-flight window (REQ-5-007). Unknown/malformed frames are ignored:
|
||||
acks are hints; gap detection and flood semantics stay authoritative
|
||||
server-side.
|
||||
"""
|
||||
try:
|
||||
envelope = json.loads(payload.decode("utf-8"))
|
||||
except (ValueError, UnicodeDecodeError):
|
||||
return
|
||||
if not isinstance(envelope, dict):
|
||||
return
|
||||
if envelope.get("type") == "seq_ack":
|
||||
ack_seq = envelope.get("seq")
|
||||
if isinstance(ack_seq, int) and ack_seq >= 0:
|
||||
self.trim_to_ack(ack_seq)
|
||||
|
||||
# -- workspace watcher ------------------------------------------------
|
||||
def _snapshot_workspace(self) -> dict[str, tuple[int, int, str | None]]:
|
||||
"""Map rel path -> (mtime_ns, size, text-or-None-if-too-large)."""
|
||||
snap: dict[str, tuple[int, int, str | None]] = {}
|
||||
root = self.config.workspace
|
||||
if not root.is_dir():
|
||||
return snap
|
||||
for dirpath, dirnames, filenames in os.walk(root):
|
||||
dirnames[:] = sorted(
|
||||
d for d in dirnames if not d.startswith(_AGENT_DIR_PREFIX)
|
||||
)
|
||||
for name in sorted(filenames):
|
||||
if name.startswith(_AGENT_DIR_PREFIX):
|
||||
continue
|
||||
path = Path(dirpath) / name
|
||||
try:
|
||||
st = path.stat()
|
||||
except OSError:
|
||||
continue
|
||||
rel = path.relative_to(root).as_posix()
|
||||
text: str | None = None
|
||||
if st.st_size <= _MAX_DIFF_BYTES:
|
||||
try:
|
||||
text = path.read_text(encoding="utf-8", errors="replace")
|
||||
except OSError:
|
||||
pass
|
||||
snap[rel] = (st.st_mtime_ns, st.st_size, text)
|
||||
return snap
|
||||
|
||||
def _file_diff_payload(self, rel: str, change: str, old: str | None, new: str | None) -> dict:
|
||||
payload: dict[str, Any] = {"path": rel, "change": change}
|
||||
if old is None and new is None:
|
||||
payload["truncated"] = True
|
||||
return payload
|
||||
diff = "".join(
|
||||
difflib.unified_diff(
|
||||
(old or "").splitlines(keepends=True),
|
||||
(new or "").splitlines(keepends=True),
|
||||
fromfile=f"a/{rel}",
|
||||
tofile=f"b/{rel}",
|
||||
)
|
||||
)
|
||||
payload["diff"] = diff
|
||||
payload["size"] = len(new or "")
|
||||
return payload
|
||||
|
||||
def _watcher_loop(self) -> None:
|
||||
# Baseline is taken in start() before it returns, so any change made
|
||||
# after start() completes is guaranteed to be observed.
|
||||
baseline = self._baseline
|
||||
last_heartbeat = time.monotonic()
|
||||
while not self._stop.wait(self.config.poll_interval_s):
|
||||
current = self._snapshot_workspace()
|
||||
for rel in sorted(current.keys() | baseline.keys()):
|
||||
if rel not in baseline and rel in current:
|
||||
self.emit(
|
||||
"file_diff",
|
||||
self._file_diff_payload(rel, "created", None, current[rel][2]),
|
||||
)
|
||||
elif rel in baseline and rel not in current:
|
||||
self.emit(
|
||||
"file_diff",
|
||||
self._file_diff_payload(rel, "deleted", baseline[rel][2], None),
|
||||
)
|
||||
else:
|
||||
old_stat, new_stat = baseline[rel], current[rel]
|
||||
if old_stat[:2] != new_stat[:2] and old_stat[2] != new_stat[2]:
|
||||
self.emit(
|
||||
"file_diff",
|
||||
self._file_diff_payload(
|
||||
rel, "modified", old_stat[2], new_stat[2]
|
||||
),
|
||||
)
|
||||
baseline = current
|
||||
if time.monotonic() - last_heartbeat >= self.config.activity_interval_s:
|
||||
self.emit("activity", {"state": "idle", "spooled": len(self._pending)})
|
||||
last_heartbeat = time.monotonic()
|
||||
|
||||
# -- shell wrapper ------------------------------------------------------
|
||||
@staticmethod
|
||||
def _is_test_command(cmd: str) -> bool:
|
||||
return "test" in cmd.lower()
|
||||
|
||||
def run_command(self, line: str) -> dict[str, Any] | None:
|
||||
"""Run one REPL line; emits stdin/command/stdout/run|test_result."""
|
||||
line = line.strip()
|
||||
if not line:
|
||||
return None
|
||||
self.emit("stdin", {"line": line})
|
||||
self.emit("activity", {"state": "command", "spooled": len(self._pending)})
|
||||
self.emit("command", {"cmd": line})
|
||||
started = time.monotonic()
|
||||
timed_out = False
|
||||
exit_code: int | None = None
|
||||
out: str | bytes = ""
|
||||
err: str | bytes = ""
|
||||
try:
|
||||
proc = subprocess.run(
|
||||
["sh", "-c", line],
|
||||
cwd=self.config.workspace,
|
||||
capture_output=True,
|
||||
timeout=self.config.command_timeout_s,
|
||||
text=True,
|
||||
errors="replace",
|
||||
)
|
||||
exit_code, out, err = proc.returncode, proc.stdout, proc.stderr
|
||||
except subprocess.TimeoutExpired as exc:
|
||||
timed_out = True
|
||||
# TimeoutExpired output attrs are always bytes (even in text mode).
|
||||
out = exc.stdout or b""
|
||||
err = exc.stderr or b""
|
||||
duration = time.monotonic() - started
|
||||
for stream, data in (("stdout", out), ("stderr", err)):
|
||||
if isinstance(data, bytes):
|
||||
data = data.decode(errors="replace")
|
||||
if data:
|
||||
self.emit("stdout", {"stream": stream, "data": data[-_MAX_OUTPUT_CHARS:]})
|
||||
kind = "test_result" if self._is_test_command(line) else "run_result"
|
||||
result = self.emit(
|
||||
kind,
|
||||
{
|
||||
"cmd": line,
|
||||
"exit_code": exit_code,
|
||||
"duration_s": round(duration, 6),
|
||||
"timed_out": timed_out,
|
||||
},
|
||||
)
|
||||
self.emit("activity", {"state": "idle", "spooled": len(self._pending)})
|
||||
return result
|
||||
|
||||
# -- lifecycle ----------------------------------------------------------
|
||||
def start(self) -> None:
|
||||
self.config.workspace.mkdir(parents=True, exist_ok=True)
|
||||
self._baseline = self._snapshot_workspace()
|
||||
self.emit("activity", {"state": "starting", "spooled": len(self._pending)})
|
||||
self._threads = [
|
||||
threading.Thread(target=self._supervisor_loop, daemon=True, name="nc-ws"),
|
||||
threading.Thread(target=self._watcher_loop, daemon=True, name="nc-watch"),
|
||||
]
|
||||
for thread in self._threads:
|
||||
thread.start()
|
||||
|
||||
def stop(self) -> None:
|
||||
if self._stop.is_set():
|
||||
return
|
||||
try:
|
||||
self.emit(
|
||||
"activity",
|
||||
{
|
||||
"state": "stopped",
|
||||
"spooled": len(self._pending),
|
||||
"dropped_overflow": self._dropped_overflow,
|
||||
},
|
||||
)
|
||||
finally:
|
||||
self._stop.set()
|
||||
self._drop_conn()
|
||||
for thread in self._threads:
|
||||
thread.join(timeout=3)
|
||||
for thread in self._threads:
|
||||
thread.join(timeout=3)
|
||||
with self._emit_lock:
|
||||
self._spool.rewrite(list(self._pending))
|
||||
|
||||
|
||||
def main() -> int:
|
||||
try:
|
||||
config = AgentConfig.from_env()
|
||||
except ValueError as exc:
|
||||
print(f"sandbox-agent: {exc}", file=sys.stderr)
|
||||
return 2
|
||||
agent = Agent(config)
|
||||
agent.start()
|
||||
try:
|
||||
# REPL mode: each stdin line is executed and reported (interactive use).
|
||||
# Daemon mode: when stdin is closed/absent (the sandbox backend spawns
|
||||
# the agent with stdin=DEVNULL), keep streaming workspace diffs +
|
||||
# activity until SIGTERM/SIGINT so the agent's lifecycle is tied to
|
||||
# the sandbox (destroy() reaps it) rather than to stdin EOF.
|
||||
if sys.stdin is None or sys.stdin.closed: # pragma: no cover - defensive
|
||||
agent._stop.wait() # noqa: SLF001 - daemon block
|
||||
else:
|
||||
line = sys.stdin.readline()
|
||||
while line:
|
||||
agent.run_command(line)
|
||||
line = sys.stdin.readline()
|
||||
if not agent._stop.is_set() and not sys.stdin.isatty(): # noqa: SLF001
|
||||
# EOF on a pipe (DEVNULL): daemonize — watch + stream until killed.
|
||||
import signal
|
||||
|
||||
signal.signal(signal.SIGTERM, lambda *_: agent._stop.set()) # noqa: SLF001
|
||||
agent._stop.wait() # noqa: SLF001
|
||||
except KeyboardInterrupt:
|
||||
pass
|
||||
finally:
|
||||
agent.stop()
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -1,14 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Test runner: pytest via venv — mock provider only, zero network calls.
|
||||
set -euo pipefail
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
APP_DIR="$(cd "$SCRIPT_DIR/.." && pwd)"
|
||||
VENV="$APP_DIR/.venv"
|
||||
|
||||
if [ ! -x "$VENV/bin/pytest" ]; then
|
||||
echo "venv missing — run scripts/bootstrap.sh first" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
cd "$APP_DIR"
|
||||
exec "$VENV/bin/pytest" -q "$@"
|
||||
@@ -1,118 +0,0 @@
|
||||
"""Assessor agent tests — live-grade coaching contract (REQ-3-007).
|
||||
|
||||
v0.3 re-grounding: the Assessor renders coaching FROM the stored grade
|
||||
(GradeRecord) — it never invents scores (the grading engine owns that).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from datetime import UTC, datetime
|
||||
|
||||
import pytest
|
||||
|
||||
from ai_service.agents.assessor import AssessorAgent, GradeCoaching
|
||||
from ai_service.config import Settings
|
||||
from ai_service.grading.store import GradeRecord, SQLiteGradeStore
|
||||
from ai_service.llm.mock import MockProvider
|
||||
from ai_service.llm.types import Message
|
||||
|
||||
COACHING_JSON = json.dumps(
|
||||
{
|
||||
"summary": "Solid iterative build; tests drove the fixes.",
|
||||
"strengths": ["Ran tests after each change."],
|
||||
"gaps": ["Did not cover the empty-input case."],
|
||||
"next_steps": ["Add one edge-case test."],
|
||||
}
|
||||
)
|
||||
|
||||
|
||||
class ScriptedProvider(MockProvider):
|
||||
def __init__(self) -> None:
|
||||
super().__init__()
|
||||
self.requests: list[list[Message]] = []
|
||||
self.replies: list[str] = []
|
||||
|
||||
async def chat(self, messages, *, model, temperature=0.7, response_format=None):
|
||||
self.requests.append([Message(role=m.role, content=m.content) for m in messages])
|
||||
if self.replies:
|
||||
return self.replies.pop(0)
|
||||
return COACHING_JSON
|
||||
|
||||
|
||||
def _grade() -> GradeRecord:
|
||||
return GradeRecord(
|
||||
learner_id="assessor-learner",
|
||||
task_id="assessor-task",
|
||||
variant_seed=None,
|
||||
digest={"error_fix_cycles": 2, "final_test_status": "pass"},
|
||||
scores={
|
||||
"criteria": {
|
||||
"process_quality": 4,
|
||||
"correctness": 3,
|
||||
"debugging_discipline": 4,
|
||||
"test_usage": 3,
|
||||
},
|
||||
"strengths": ["s"],
|
||||
"gaps": ["g"],
|
||||
"verdict": "developing",
|
||||
},
|
||||
verdict="GRADED",
|
||||
model="gemma4:31b",
|
||||
created_at=datetime.now(UTC),
|
||||
)
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def provider() -> ScriptedProvider:
|
||||
return ScriptedProvider()
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def agent(provider) -> AssessorAgent:
|
||||
return AssessorAgent(provider, Settings(provider="mock"))
|
||||
|
||||
|
||||
class TestCoachGrade:
|
||||
async def test_prompt_contains_stored_grade_not_learner_id(self, agent, provider) -> None:
|
||||
await agent.coach_grade(_grade())
|
||||
all_text = "\n".join(
|
||||
m.content for request in provider.requests for m in request
|
||||
)
|
||||
assert "process_quality" in all_text # stored scores rendered
|
||||
assert "GRADED" in all_text
|
||||
assert "assessor-learner" not in all_text # D-028 anonymity
|
||||
|
||||
async def test_coaching_validates_via_d020(self, agent, provider) -> None:
|
||||
coaching = await agent.coach_grade(_grade())
|
||||
assert isinstance(coaching, GradeCoaching)
|
||||
assert coaching.summary
|
||||
assert coaching.next_steps
|
||||
|
||||
async def test_malformed_then_good_exercises_retry(self, agent, provider) -> None:
|
||||
provider.replies = ["garbage", COACHING_JSON]
|
||||
coaching = await agent.coach_grade(_grade())
|
||||
assert coaching.summary
|
||||
assert len(provider.requests) == 2
|
||||
|
||||
async def test_no_corpus_artifact_imports(self) -> None:
|
||||
import ast
|
||||
from pathlib import Path
|
||||
|
||||
py = Path(__file__).parents[2] / "ai_service" / "agents" / "assessor.py"
|
||||
tree = ast.parse(py.read_text())
|
||||
for node in ast.walk(tree):
|
||||
if isinstance(node, ast.ImportFrom) and node.module:
|
||||
assert "corpus.artifacts" not in node.module
|
||||
assert "corpus.telemetry" not in node.module
|
||||
|
||||
|
||||
class TestStoreRoundtrip:
|
||||
def test_grade_store_roundtrip(self, tmp_path) -> None:
|
||||
store = SQLiteGradeStore(db_path=tmp_path / "g.db")
|
||||
record = _grade()
|
||||
store.save(record)
|
||||
fetched = store.get("assessor-learner", "assessor-task")
|
||||
assert fetched is not None
|
||||
assert fetched.scores["criteria"]["process_quality"] == 4
|
||||
store.close()
|
||||
@@ -1,55 +0,0 @@
|
||||
"""BaseAgent contract tests — stub agent + mock provider."""
|
||||
|
||||
|
||||
from ai_service.agents.base import BaseAgent
|
||||
from ai_service.config import Settings
|
||||
from ai_service.corpus.learner_context import get_learner_context
|
||||
from ai_service.llm.mock import MockProvider
|
||||
from ai_service.llm.types import Message
|
||||
|
||||
|
||||
class StubAgent(BaseAgent):
|
||||
name = "stub"
|
||||
|
||||
def system_prompt(self, learner_context=None) -> str:
|
||||
return "You are Stub. Answer briefly."
|
||||
|
||||
|
||||
def make_agent() -> StubAgent:
|
||||
return StubAgent(MockProvider(), Settings(provider="mock"))
|
||||
|
||||
|
||||
async def test_build_messages_composition():
|
||||
agent = make_agent()
|
||||
history = [Message(role="user", content="earlier"), Message(role="assistant", content="reply")]
|
||||
messages = agent.build_messages(history, "new question")
|
||||
assert messages[0].role == "system"
|
||||
assert messages[0].content == "You are Stub. Answer briefly."
|
||||
assert [m.content for m in messages[1:]] == ["earlier", "reply", "new question"]
|
||||
|
||||
|
||||
async def test_stream_reply_yields_deltas():
|
||||
agent = make_agent()
|
||||
tokens = [t async for t in agent.stream_reply(user_input="hello")]
|
||||
assert len(tokens) >= 1
|
||||
assert all(isinstance(t, str) for t in tokens)
|
||||
|
||||
|
||||
async def test_stream_reply_with_history_and_context():
|
||||
agent = make_agent()
|
||||
ctx = get_learner_context()
|
||||
history = [Message(role="user", content="earlier"), Message(role="assistant", content="ok")]
|
||||
tokens = [t async for t in agent.stream_reply(history, "next", ctx)]
|
||||
assert tokens
|
||||
|
||||
|
||||
async def test_structured_reply_requires_schema():
|
||||
import pytest
|
||||
|
||||
agent = make_agent()
|
||||
with pytest.raises(ValueError):
|
||||
await agent.structured_reply(user_input="x", schema=None)
|
||||
|
||||
|
||||
async def test_name_defaults():
|
||||
assert make_agent().name == "stub"
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user