Compare commits
51 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 135ea21a61 | |||
| 1648467828 | |||
| 1a68d808fb | |||
| 2c68b44c1a | |||
| fb2db4ee97 | |||
| 485d86e117 | |||
| e798e1a6da | |||
| cd45097e52 | |||
| 1d03f0c8f5 | |||
| 12b2300f6f | |||
| e1460aed3c | |||
| b7a56d35bc | |||
| 16fb52d8f7 | |||
| a905eb8c67 | |||
| ed243594d2 | |||
| c760f9af2b | |||
| b4ae388f22 | |||
| 925ab096fb | |||
| 82ae839cd4 | |||
| b6c1bc9d54 | |||
| f281eeaf62 | |||
| 22d4fa212c | |||
| 007865a5a1 | |||
| 04bdccf189 | |||
| f3071e4b79 | |||
| 3d72fd28ec | |||
| 97893f2386 | |||
| e63b996361 | |||
| 0b34255855 | |||
| 6ab0ae2c0a | |||
| 9ff86f9cd0 | |||
| 4acffac71e | |||
| 82b9de382a | |||
| 430b4a727d | |||
| b52bef93e5 | |||
| e798c52edb | |||
| b303a41a45 | |||
| 3322c1dab9 | |||
| 0bde9cbf2e | |||
| 85028678c7 | |||
| 1a46606827 | |||
| 0fccb8d250 | |||
| 2474678da6 | |||
| 3f58fc3454 | |||
| edf586b03a | |||
| 4439486858 | |||
| b49b9189fe | |||
| f75352d0f0 | |||
| 26b4a5be60 | |||
| b6850fc036 | |||
| cfdceac17a |
+88
-21
@@ -6,6 +6,8 @@ Nextcraft is a TypeScript monorepo (pnpm workspaces + turborepo) with a Next.js
|
||||
|
||||
**v0.3 additions (Credential Engines):** real credential engines replace v0.2 mock inputs — a sandbox fabric (isolated per-learner coding environments via Linux user/mount/pid/net namespaces), a live build-telemetry pipeline (WebSocket ingest + SQLite-ordered event log), a process-trace grading engine, seeded per-learner variant task generation, and a voice-based oral defense (STT/TTS via a new provider-agnostic voice layer). **First real persistence introduced: SQLite** (`ai_service/telemetry/`, grading, variant, defense stores). Lab/Assessor/Proctor agents are re-grounded onto real telemetry/traces. **Identity/age-gating (KYC) deferred per founder directive** — no security engineer persona; secrets-hygiene checklist only.
|
||||
|
||||
**v0.4 additions (Distribution & Bootstrap CLI, founder directive D-016):** a new `apps/cli` package — the `nextcraft` bootstrap CLI (`doctor`/`bootstrap`/`verify`/`dev`) compiled to a self-contained linux x64 binary via **Node SEA** (probe-verified: Go/Rust absent, node v24.15.0 SEA-capable), installed by a repo-served one-liner script that resolves the latest Gitea release, downloads binary + sha256 sidecar, verifies, and installs to `~/.local/bin`. Every release from v0.4 onward attaches the binary + checksum as release assets (the "ongoing binaries" requirement). The CLI is a thin wrapper: all orchestration logic stays in `apps/ai-service/scripts/` (bootstrap.sh/dev.sh) — the CLI composes them via subprocess (A-202), duplicating nothing. Previously-planned v0.4 seams (real STT/TTS, KYC, design/sim envs, seq-lease) move to v0.5.
|
||||
|
||||
### Confirmed Technology Stack (v0.2)
|
||||
|
||||
| Technology | Version | Purpose |
|
||||
@@ -26,6 +28,8 @@ Nextcraft is a TypeScript monorepo (pnpm workspaces + turborepo) with a Next.js
|
||||
| pydantic | 2.13.x | Request/response models, structured outputs |
|
||||
| pydantic-settings | 2.15.x | Settings + env-file loading (replaces python-dotenv) |
|
||||
| httpx | 0.28.x | Async LLM HTTP client (ollama-cloud + local providers) |
|
||||
| sqlmodel / sqlalchemy | 0.0.24 / 2.x | Typed SQLite persistence for the v0.3 engine stores (D-027) |
|
||||
| python-multipart | 0.0.x | Multipart audio upload for the defense answer route (REQ-3-006) |
|
||||
| sse-starlette | 3.4.x | SSE framing, ping keep-alive |
|
||||
| pytest | 9.x | Test runner |
|
||||
| pytest-asyncio | 1.4.x | Async tests (auto mode) |
|
||||
@@ -45,6 +49,15 @@ Deliberately **not** used: openai-python SDK (the `LLMProvider` protocol is the
|
||||
7. **D-022 Monorepo integration** — zero-dependency shim `package.json` in apps/ai-service + `ai#*` turbo passthrough tasks (`cache:false, outputs:[]`) + root `ai:dev`/`ai:test` scripts + idempotent venv bootstrap.
|
||||
8. **D-023 Testing** — pytest-asyncio auto mode; TestClient `client.stream()` for SSE; httpx MockTransport for byte-exact provider parser tests; scripted mock provider incl. failure modes. Tests never call the cloud.
|
||||
|
||||
### v0.4 Architecture Decisions (from Research — Distribution & Bootstrap CLI)
|
||||
|
||||
18. **D-033 Binary toolchain = Node SEA (probe-verified)** — Go and Rust are absent from this box; node v24.15.0 ships SEA support (`--experimental-sea-config`, postject-free on linux via `cp node nextcraft && node sea-config` … blob injection with the system `dd`/`npx postject` if needed). CLI source lives in `apps/cli` (TypeScript, compiled to a single CJS bundle by esbuild, then SEA-injected into a copy of the node binary → `nextcraft-linux-x64`). Fallback if SEA breaks: python3 `zipapp` (3.11.2 available). No new toolchain deps beyond dev-scoped esbuild.
|
||||
19. **D-034 CLI = thin wrapper, orchestration stays in scripts/** — `nextcraft` composes `apps/ai-service/scripts/bootstrap.sh` and `scripts/dev.sh` equivalents via `spawn` with inherited stdio and timeout guards (A-202/A-209). doctor/bootstrap/verify implement only *checking* logic (prereqs, env template, health) — never re-implement installs. This keeps one source of truth for bootstrap semantics.
|
||||
20. **D-035 Install path = repo raw `install.sh` + Gitea latest-release API** — the one-liner `curl -fsSL <forge>/coreci/nextcraft/raw/main/scripts/install.sh | bash` resolves `GET /api/v1/repos/coreci/nextcraft/releases/latest`, downloads the `nextcraft-linux-x64` + `nextcraft-linux-x64.sha256` assets, verifies sha256 (`shasum -a 256`), installs to `~/.local/bin` (PATH hint), and degrades to printed source-bootstrap instructions when no binary asset exists or the platform mismatches (A-203/A-204/A-206).
|
||||
21. **D-036 Ongoing binaries = ship-workflow asset step** — the release pipeline (v0.3's `ShipWorkflow.createRelease` equivalent, executed as the ship step's asset stage) builds the binary + checksum and attaches both to every Gitea release from v0.4 onward (A-205). Token resolution stays `.env*`-only (D-006/D-014); binaries are linux x64 only for v0.4 (macOS arm64 deferred — unverifiable on this box).
|
||||
22. **D-037 CLI package layout** — `apps/cli` is a pnpm workspace package (`@nextcraft/cli`): `src/` (entry, commands/, checks/, lib/), `scripts/build-binary.mjs` (esbuild bundle → SEA inject), unit tests runnable via `pnpm --filter @nextcraft/cli test` (node:test, no new test framework). Root `package.json` gains `cli:*` passthrough scripts mirroring the `ai:*` pattern (D-022).
|
||||
23. **D-038 Network mode (v0.3.5)** — dev binds 0.0.0.0 (`AI_HOST`, default 0.0.0.0, revert via 127.0.0.1); CORS + WS-origin gates read `AI_CORS_ORIGINS` (default `*` — any origin, safe only because credentials are never enabled; explicit comma list restricts); the web client derives the API base URL from the browser hostname at runtime (`engine-base-url.ts`: `NEXT_PUBLIC_AI_SERVICE_URL` override → `http://${window.location.hostname}:8420` → `localhost` server-side). Hotfix also fixes: SEA direct-run detection (`require("node:sea").isSea()` — argv shape differs by invocation style), installer honesty gate (silent `--version` = hard fail), bootstrap venv recovery (poisoned partial `.venv` removal + distro-specific `apt install python3.XX-venv` hint), and doctor venv-capability probe with bootstrap preflight.
|
||||
|
||||
### v0.3 Architecture Decisions (from Research — Credential Engines)
|
||||
|
||||
9. **D-024 Sandbox isolation = Linux namespaces via `unshare`** — per-learner sandbox runs as a subprocess entered into fresh user+mount+pid+network namespaces (`unshare --user --map-root-user --mount --pid --fork --net`). Probe-verified on this box: in-namespace uid=0, **network fully isolated** (0 interfaces), learner writes land in a per-sandbox directory; proc-remount not permitted here but not required. Chosen because no container runtime (docker/podman/bwrap/firejail) exists on the box and there is no sudo. A `SandboxBackend` protocol abstracts the spawner so a future containerd/runc backend can replace namespace-spawning without touching callers.
|
||||
@@ -65,14 +78,14 @@ Deliberately **not** used: openai-python SDK (the `LLMProvider` protocol is the
|
||||
|
||||
| Component | Description | Boundaries | Depends On |
|
||||
|-----------|-------------|------------|------------|
|
||||
| `ai_service/main.py` | FastAPI app factory, lifespan (httpx client pool, provider factory), CORS (localhost only), /health | App entry | config, llm, agents, api |
|
||||
| `ai_service/main.py` | FastAPI app factory, lifespan (httpx client pool, provider factory, SandboxManager + reaper loop, SQLite engine stores on app.state), CORS (localhost only, incl. PUT for file writes), /health | App entry | config, llm, agents, api, engines |
|
||||
| `ai_service/config.py` | pydantic-settings Settings (env_prefix="AI_", env_file, SecretStr key) | Configuration only | None |
|
||||
| `ai_service/api/` | Endpoints: chat.py (POST /v1/chat/stream, SSE), lab.py, assessment.py, proctor.py, mentor.py; deps.py (DI) | Composes agents + sessions; never imported by llm/ or agents/ | agents, llm |
|
||||
| `ai_service/api/` | Endpoints: chat.py (POST /v1/chat/stream, SSE), lab.py, assessment.py (POST /v1/assessment/evaluate v0.2 + POST /v1/assessment/grade v0.3), proctor.py, mentor.py, sandboxes.py (lifecycle + files/exec routes, G-5 abuse gates), telemetry.py (WS ingest + trace/gaps reads), variants.py (seeded per-learner variants), defense.py (defense loop, REQ-3-006); deps.py (DI) | Composes agents + sessions + engines; never imported by llm/ or agents/ | agents, llm, sandbox, telemetry, grading, variants, voice |
|
||||
| `ai_service/llm/` | types.py (Message; ChatDelta/ChoiceDelta removed in P3 — no consumers), base.py (LLMProvider protocol), openai_compat.py (ollama-cloud + local), mock.py (deterministic), factory.py | Never imports agents/ or api/ | config |
|
||||
| `ai_service/agents/` | base.py (BaseAgent ABC), registry.py, session.py (SessionStore), structured.py (JSON defense), coach/tutor/lab/assessor/proctor/mentor.py | Never imports api/ | llm, prompts, corpus |
|
||||
| `ai_service/agents/` | base.py (BaseAgent ABC), registry.py, session.py (SessionStore), structured.py (JSON defense), coach/tutor/lab/assessor/proctor/mentor.py + examiner.py (seventh agent, v0.3) | Never imports api/ | llm, prompts, corpus, telemetry |
|
||||
| `ai_service/prompts/` | Per-agent system prompt constants + render_context functions (str.format_map) | Data only | None |
|
||||
| `ai_service/corpus/` | Mock engine inputs: learner_context.py, telemetry.py (Lab/Proctor scenarios), artifacts.py (pre-baked artifacts, rubrics, transcripts) | Pydantic-typed; aligned with TS packages/mock-data by convention | None |
|
||||
| `scripts/` | bootstrap.sh (venv + pip install idempotent), dev.sh (exports keys from .ciagent/.env.secrets → uvicorn), test.sh (pytest), lint.sh (ruff check, G-3) | Dev entry points | pyproject.toml |
|
||||
| `ai_service/corpus/` | Mock engine inputs: learner_context.py, telemetry.py (Lab/Proctor scenarios), artifacts.py (pre-baked artifacts, rubrics, transcripts). Since v0.3 P6 these are DORMANT, test-only fixtures (dormant-header noted) — the live learner path uses real engine inputs | Pydantic-typed; aligned with TS packages/mock-data by convention | None |
|
||||
| `scripts/` | bootstrap.sh (venv + pip install idempotent), dev.sh (exports keys from .ciagent/.env.secrets → uvicorn), test.sh (pytest), lint.sh (ruff check, v0.2 G-3) | Dev entry points | pyproject.toml |
|
||||
| `tests/` | conftest.py (mock provider, settings override, TestClient), health, llm (MockTransport parser), agents (framework + per-agent), api (SSE stream tests) | Mock provider only — no cloud | all |
|
||||
|
||||
**Module boundary rules:** `llm/` never imports `agents/` or `api/`; `agents/` never imports `api/`; `api/` composes both via DI. `corpus/` is the only home of mock engine data. Prompts are code — versioned and reviewed in git.
|
||||
@@ -85,32 +98,46 @@ Deliberately **not** used: openai-python SDK (the `LLMProvider` protocol is the
|
||||
| `ai_service/telemetry/` | `models.py` (TelemetryEvent, TraceSpan), `store.py` (TraceStore protocol + SQLite impl D-027), `ingest.py` (WebSocket /v1/telemetry/ingest, seq gap detection D-026) | Persistence; never imports agents/ | config |
|
||||
| `ai_service/grading/` | `features.py` (deterministic trace digest D-028), `engine.py` (rubric scoring orchestration), `store.py` (GradeStore) | LLM only via digest; never sees raw trace | llm, telemetry, prompts |
|
||||
| `ai_service/variants/` | `templates.py` (task template library), `generator.py` (seeded LLM instantiation D-029), `store.py` (VariantStore) | LLM via structured output | llm, grading |
|
||||
| `ai_service/voice/` | `base.py` (VoiceProvider protocol D-030), `openai_audio.py` (STT/TTS vs compatible endpoint), `browser.py` (native SR/TTS fallback descriptor), `mock.py` (deterministic) | Never imports agents/ or api/ | config |
|
||||
| `ai_service/voice/` | `base.py` (VoiceProvider protocol D-030), `browser.py` (native SR/TTS fallback descriptor), `mock.py` (deterministic; the real server STT/TTS provider is the v0.4 seam — GRILL CUT-1/G-7), `factory.py` (provider selection), `defense_store.py` (DefenseStore: transcripts + integrity signals, D-027) | Never imports agents/ or api/ | config |
|
||||
| `ai_service/agents/examiner.py` | Seventh agent: oral defense examiner; streams over existing SSE, consumes process traces + emits integrity signals | reuses BaseAgent (D-018) | llm, prompts, telemetry |
|
||||
| `ai_service/data/*.db` | SQLite databases (telemetry/grades/variants/defenses) | gitignored | — |
|
||||
| `scripts/sandbox-agent.py` | Tiny in-namespace capture process shipped into the sandbox; streams telemetry to ingest | standalone | stdlib only |
|
||||
|
||||
**Boundary additions:** `sandbox/`, `telemetry/`, `grading/`, `variants/`, `voice/` are engine modules — they never import `api/` (which composes them via DI) and never import `agents/` (agents call engines through narrow interfaces, not vice versa).
|
||||
|
||||
### apps/cli — Nextcraft Bootstrap CLI (v0.4 NEW)
|
||||
|
||||
| Component | Description | Boundaries | Depends On |
|
||||
|-----------|-------------|------------|------------|
|
||||
| `src/index.ts` | Entry: arg parsing (no deps beyond node stdlib at runtime), command dispatch, `--help`/`--version`, exit-code contract (0 ok / 1 failure / 2 usage) | CLI surface only | commands/ |
|
||||
| `src/commands/` | `doctor.ts` (prereq checks + actionable errors), `bootstrap.ts` (pnpm install + scripts/bootstrap.sh wrapper + env template copy + key validation), `verify.ts` (health: venv imports, ports, env, build readiness), `dev.ts` (thin passthrough to scripts/dev.sh) | Compose checks/ + lib/; spawn scripts — never re-implement them | checks/, lib/ |
|
||||
| `src/checks/` | Pure check functions: `check-command.ts` (binary-on-PATH + version compare), `check-env.ts` (template diff, required/optional key classification) | Pure logic, unit-testable, no fs side effects at import | None |
|
||||
| `src/lib/` | `spawn.ts` (subprocess with timeout + inherited stdio), `log.ts` (✓/✗/warn output formatter) | Shared utilities | None |
|
||||
| `scripts/build-binary.mjs` | esbuild → CJS bundle → Node SEA injection → `dist/nextcraft-linux-x64` + sha256 sidecar | Build-time only | esbuild (dev dep) |
|
||||
| `scripts/install.sh` | The one-liner install script served from repo raw: Gitea latest-release resolve → download + checksum verify → ~/.local/bin; source-bootstrap fallback | Standalone POSIX sh | forge API |
|
||||
| `tests/` | node:test unit tests: command dispatch, check logic, env template diff, install-script shellcheck-style assertions | Fixtures only — never mutate repo state | src/ |
|
||||
|
||||
**Boundary rules:** the CLI never imports from `apps/web`, `packages/*`, or `ai_service` Python modules — it orchestrates them exclusively via subprocess/filesystem. Runtime deps: node stdlib only (no runtime npm deps; esbuild is dev-only). The binary embeds the bundle; `scripts/bootstrap.sh` remains the single source of bootstrap truth (D-034).
|
||||
|
||||
### apps/web — Next.js Application
|
||||
|
||||
| Component | Description | Boundaries | Depends On |
|
||||
|-----------|-------------|------------|------------|
|
||||
| `app/(learner)/` | Learner surface route group: landing, catalog, competency stack, dashboard, byte viewer, sandbox mockup, assessment mockup | Learner-only routes and layouts | packages/ui, packages/mock-data, packages/types |
|
||||
| `app/(learner)/` | Learner surface route group: landing, catalog, competency stack, dashboard, byte viewer, build surface (`/build/[competencyId]` — real in-browser build), defense surface (`/defend/[competencyId]` — live oral defense + grading) | Learner-only routes and layouts | packages/ui, packages/mock-data, packages/types |
|
||||
| `app/(marketplace)/` | Marketplace surface route group: job board, job detail, employer profile, search/filter, pricing | Marketplace-only routes and layouts | packages/ui, packages/mock-data, packages/types |
|
||||
| `app/(employer)/` | Employer dashboard route group: overview, talent search, candidate profile, posting management | Employer-only routes and layouts | packages/ui, packages/mock-data, packages/types |
|
||||
| `app/(admin)/` | Admin surface route group: overview, learner management, competency graph viewer, moderation | Admin-only routes and layouts | packages/ui, packages/mock-data, packages/types |
|
||||
| `app/layout.tsx` | Root layout: theme provider, navigation shell, responsive container | All routes | packages/ui |
|
||||
| `components/` | Surface-specific components (learner/, marketplace/, employer/, admin/) plus shared chrome (navigation-shell, header/footer, role-switcher, theme-provider, breadcrumbs, dark-mode-toggle) | App-level components | packages/ui |
|
||||
| `hooks/` | use-chat-stream.ts — SSE client hook: fetch + ReadableStream, byte buffering + frame reassembly, idempotent AbortController cleanup | Client components only | ai-service SSE |
|
||||
| `lib/` | sse.ts (shared SSE frame parser — CRLF normalization + `: ping` immunity, G-1), breadcrumbs.ts, format.ts | Pure utilities | None |
|
||||
| `components/` | Surface-specific components (learner/, marketplace/, employer/, admin/) plus shared chrome (navigation-shell, header/footer, role-switcher, theme-provider, breadcrumbs, dark-mode-toggle); v0.3 learner: build-surface, sandbox-terminal (read-only exec output), defense-session | App-level components | packages/ui |
|
||||
| `hooks/` | use-chat-stream.ts — SSE client hook: fetch + ReadableStream, byte buffering + frame reassembly, idempotent AbortController cleanup; use-sandbox-session.ts (v0.3) — sandbox lifecycle for the build session: create on task open, destroy on unmount, mid-start failure cleanup, 503/403/429 honest surfaces | Client components only | ai-service SSE / engine API |
|
||||
| `lib/` | sse.ts (shared SSE frame parser — CRLF normalization + `: ping` immunity, v0.2 G-1), breadcrumbs.ts, format.ts, engine-base-url.ts (v0.3.5: runtime API base — env override → browser hostname → localhost), engine-client.ts (v0.3: typed fetch client for /v1/sandboxes, files/exec, variants, grade, defense, traces) | Pure utilities | None |
|
||||
|
||||
### packages/ui — Shared Component Library
|
||||
|
||||
| Component | Description | Boundaries | Depends On |
|
||||
|-----------|-------------|------------|------------|
|
||||
| `tokens/` | Design tokens as TS constants: colors, spacing, radii, shadows, breakpoints (mirrored as Tailwind v4 `@theme` tokens in apps/web globals.css) | Foundation layer — no dependencies | None |
|
||||
| `primitives/` | Button, Input, Card, Badge, Avatar — each with a Storybook story | Atomic UI components | tokens, packages/types |
|
||||
| `primitives/` | Button, Input, Card, Badge, Avatar (v0.1) + TerminalFrame, TelemetryStatus, MicControl, GradeBadge, TranscriptViewer (v0.3 build/defense surfaces) — each with a Storybook story | Atomic UI components | tokens, packages/types |
|
||||
|
||||
Composite/layout/theme components (navigation shell, tables, chat panels, graph viewer, theme provider) live in `apps/web/components/` as app-level components, not in packages/ui.
|
||||
|
||||
@@ -134,11 +161,42 @@ Composite/layout/theme components (navigation shell, tables, chat panels, graph
|
||||
| `marketplace.ts` | Job, Employer, Candidate, JobPosting, TalentMatch, SearchFilter | Marketplace types | None |
|
||||
| `user.ts` | Learner, Admin, EmployerUser, AgeGroup, Role | User types | None |
|
||||
| `ui.ts` | Component props, theme config, breakpoint definitions | UI types | None |
|
||||
| `telemetry.ts` | TelemetryEvent/ExecResult wire shapes for the live build surface (v0.3) | Engine types | None |
|
||||
| `variants.ts` | Variant/TaskTemplate shapes for per-learner task statements (v0.3) | Engine types | None |
|
||||
| `grading.ts` | GradeRecord/RubricScore shapes for live grading display (v0.3) | Engine types | None |
|
||||
| `defense.ts` | DefenseSession/transcript/integrity-signal shapes for the defense surface (v0.3) | Engine types | None |
|
||||
|
||||
---
|
||||
|
||||
## Data Flow
|
||||
|
||||
### v0.3 credential flow (current)
|
||||
|
||||
```
|
||||
[learner build surface /build/*] [learner defense surface /defend/*]
|
||||
file CRUD + Run/Test (HTTP) mic MediaRecorder / typed + TTS playback
|
||||
│ │
|
||||
▼ ▼
|
||||
[api/sandboxes files/exec] ──exec──▶ [namespace sandbox] [api/defense start/answer/finish]
|
||||
│ │ capture agent │
|
||||
│ ▼ (WS telemetry) ▼
|
||||
│ [api/telemetry ingest] [DefenseStore (SQLite)]
|
||||
│ │ SQLite │ transcript + integrity signals
|
||||
│ ▼ │
|
||||
│ [TraceStore] ────▶ [GradingEngine: digest (grading/features)
|
||||
│ │ + rubric LLM (D-028)] ──▶ [GradeStore]
|
||||
│ ▼ ▼
|
||||
└──▶ Lab agent (live digest) Assessor (grade output) / Proctor (integrity)
|
||||
Examiner agent (SSE) ◀── defense sessions
|
||||
variants: [api/variants] ◀── [VariantStore (seeded, D-029)] ── per-learner task statements
|
||||
```
|
||||
|
||||
- Lab consumes the live trace digest; Assessor consumes grading-engine output; Proctor consumes telemetry + defense integrity signals (REQ-3-007) — no mock fallback in the learner path (v0.2 corpus scenarios are dormant test-only fixtures).
|
||||
- The learner's path is: variant task → in-sandbox build (telemetry streams to SQLite) → grade My Work (rubric scores from the real trace) → oral defense → verdict.
|
||||
- Flooded/gapped traces are terminal: ingest closes 1008 and marks INCOMPLETE_FLOODED (G-3); the grader returns UNGRADABLE_TRACE_INCOMPLETE (G-4) — no credential from an incomplete trace.
|
||||
|
||||
### v0.2 chat flow (complete, still live)
|
||||
|
||||
```
|
||||
[packages/mock-data + packages/types] [ai_service/corpus]
|
||||
│ (TS, web surfaces) │ (Python, agent inputs)
|
||||
@@ -153,21 +211,27 @@ Composite/layout/theme components (navigation shell, tables, chat panels, graph
|
||||
(https://ollama.com/v1)
|
||||
```
|
||||
|
||||
- Web surfaces remain server-component-first; client components (chat, filters, graph viewer, dark mode toggle) fetch directly from ai-service over SSE (A-002: no Next.js API-route proxy in v0.2).
|
||||
- The LLM provider layer is a dumb pipe — OpenAI-compatible chunks pass through byte-identical; envelope logic (meta/done/error) lives only in the API layer (D-016).
|
||||
- Lab/Assessor/Proctor read mock scenarios from `ai_service/corpus/` — real engines are v0.3+.
|
||||
- Web surfaces remain server-component-first; client components (chat, filters, graph viewer, dark mode toggle) fetch directly from ai-service over SSE (A-002: no Next.js API-route proxy).
|
||||
- The LLM provider layer is a dumb pipe — OpenAI-compatible chunks pass through byte-identically; envelope logic (meta/done/error) lives only in the API layer (D-016).
|
||||
- The v0.2 corpus scenarios (`ai_service/corpus/`) are retained as dormant, test-only fixtures (dormant-header noted); they are no longer inputs to the live learner path.
|
||||
- All automated tests use the deterministic mock provider; the cloud is for manual probes only.
|
||||
|
||||
---
|
||||
|
||||
## Build Order (v0.3)
|
||||
## Build Order (v0.4)
|
||||
|
||||
1. **Bootstrap CLI core** — apps/cli package: doctor checks (node/pnpm/python3/git/unshare), bootstrap wrapper (pnpm install + scripts/bootstrap.sh + .env template + key validation), verify health check, dev passthrough; unit tests
|
||||
2. **Binary build + release pipeline** — esbuild bundle → Node SEA binary (`nextcraft-linux-x64`) + sha256 sidecar; install.sh one-liner (Gitea latest-release resolve + checksum verify + PATH install); release-asset upload wired into the ship flow (ongoing binaries from v0.4 onward)
|
||||
3. **Install docs + fresh-clone E2E** — README quickstart (one-liner → doctor → bootstrap → dev), CLI reference, fresh-clone end-to-end test proving a clean clone reaches a running stack
|
||||
|
||||
## Build Order (v0.3 — complete)
|
||||
|
||||
1. **Sandbox fabric** — SandboxBackend protocol + unshare namespace spawner + lifecycle manager (create/list/snapshot/destroy) + concurrency guard + per-sandbox workdir; isolation + resource-limit probes
|
||||
2. **Live build telemetry** — TelemetryEvent models + SQLite TraceStore + WebSocket ingest endpoint + seq gap detection + in-sandbox capture agent
|
||||
3. **Process-trace grading engine** — deterministic feature/digest computation + rubric scoring via LLM structured output + GradeStore; calibrated against v0.2 mock corpora
|
||||
4. **Variant task generation** — template library + seeded LLM instantiation + VariantStore + difficulty normalization anchors
|
||||
5. **Oral / voice defense** — VoiceProvider protocol + STT/TTS + mock + browser fallback + Examiner agent + transcript/integrity-signal capture
|
||||
6. **Agent re-grounding + learner surface integration** — Lab/Assessor/Proctor consume real telemetry/grades/defense signals; learner sandbox mockup → real in-browser xterm.js build/run; assessment mockup → live defense + live grading
|
||||
6. **Agent re-grounding + learner surface integration** — Lab/Assessor/Proctor consume real telemetry/grades/defense signals; learner sandbox mockup → real in-browser build/run (Run/Test buttons executing in a namespace sandbox, read-only exec-output panel — no interactive shell, CUT-2/G-8); assessment mockup → live defense + live grading
|
||||
|
||||
---
|
||||
|
||||
@@ -184,16 +248,19 @@ The v0.1 build order (monorepo → types → mock data → tokens → primitives
|
||||
|
||||
---
|
||||
|
||||
## Future Architecture (Post-v0.3, for reference)
|
||||
## Future Architecture (Post-v0.4, for reference)
|
||||
|
||||
v0.3 delivers the real credential engines; later milestones fill in the remaining platform:
|
||||
v0.4 delivers distribution (CLI + binary releases); later milestones fill in the remaining platform:
|
||||
|
||||
- **In-memory sessions → PostgreSQL + Drizzle/SQLModel** — SessionStore + v0.3 TraceStore/GradeStore/VariantStore/DefenseStore protocols swap SQLite→Postgres with no API changes
|
||||
- **userns subprocess sandboxes → containerd/runc backend** — D-024 `SandboxBackend` protocol swap; same lifecycle API
|
||||
- **Coding-IDE sandbox → design tool + simulation environments** — REQ-F-021 full scope (v0.4)
|
||||
- **Coding-IDE sandbox → design tool + simulation environments** — REQ-F-021 full scope (v0.5)
|
||||
- **Mock provider → per-agent model routing** — provider factory already selects by config; per-agent `AI_<AGENT>_MODEL` overrides
|
||||
- **No auth → real KYC + sessions** — **deferred per founder directive**; REQ-F-017 identity/age-gating lands post-v0.3 (v0.4+). Age-gating remains the v0.1 visual flow mockup
|
||||
- **No auth → real KYC + sessions** — **deferred per founder directive; moved to v0.5 with D-016**; REQ-F-017 identity/age-gating lands post-v0.4. Age-gating remains the v0.1 visual flow mockup
|
||||
- **Mock voice → real server STT/TTS (openai-audio provider)** — CUT-1/G-7 seam moved to v0.5 per D-016; VoiceProvider protocol is the drop-in point
|
||||
- **linux x64 binary → macOS arm64 + auto-update** — D-036 defers non-linux targets (unverifiable on this box); `nextcraft upgrade` (self-replace from latest release) is the natural v0.5+ follow-up
|
||||
- **Exec-telemetry seq-lease / replay-margin fix** — the P6-lesson one-line ACK gap moves to v0.5 per D-016
|
||||
- **No search → Semantic vector search (pgvector)** — Filter UI replaced with vector similarity search
|
||||
- **No payments → Payment processing** — Pricing page replaced with real subscription/payment flows
|
||||
|
||||
The monorepo structure (apps/web + apps/ai-service + packages/*) accommodates further apps without restructuring.
|
||||
The monorepo structure (apps/web + apps/ai-service + apps/cli + packages/*) accommodates further apps without restructuring.
|
||||
@@ -1,8 +0,0 @@
|
||||
{
|
||||
"phase": 1,
|
||||
"stage": "complete",
|
||||
"milestone": "v0.3",
|
||||
"phase_role": "execution",
|
||||
"attempts": 0,
|
||||
"updated_at": "2026-09-11T18:43:22Z"
|
||||
}
|
||||
@@ -46,3 +46,41 @@ The credential-pipeline architecture (telemetry → trace → grade → defense)
|
||||
## Outcome
|
||||
|
||||
**GO** — all six binding decisions and both scope cuts applied to PLAN.md / REQUIREMENTS.md / ROADMAP.md / PROJECT.md before Phase 1 execution. No axis requires escalation (all resolvable at confidence ≥ 0.85). The milestone no longer claims resource enforcement it cannot deliver, and the no-auth abuse vector is closed at MVP scale.
|
||||
|
||||
---
|
||||
# Nextcraft v0.4 — GRILL.md (Adversarial Review Verdict)
|
||||
|
||||
**Stage:** GRILL, Phase 0 pre-execution · **Verdict:** GO-WITH-CHANGES · **Confidence:** 0.83
|
||||
|
||||
## Summary
|
||||
|
||||
The distribution milestone is small, founder-directed (D-016, confidence 0.99), and additive (zero changes to the running credential pipeline). The plan's central risk: **Node SEA was probe-verified as a flag, not as a working build** — the v0.3 lesson (A-101: probe the mechanism, not the existence) applies. Second gap: a binary whose `--version` lies (stale package.json) would poison the "ongoing binaries" contract. Third: sed-based JSON parsing in install.sh is a fragility + integrity risk. Fourth: "ongoing binaries" has no enforcement mechanism beyond prose. All four closed by binding decisions G-101..G-104 below. No scope cuts required — the milestone is already minimal.
|
||||
|
||||
## Per-Axis Findings
|
||||
|
||||
| Axis | Verdict | Rationale |
|
||||
|------|---------|-----------|
|
||||
| Business case | PASS | Founder directive explicit + recorded (D-016). Evidence of need: live Gitea probe shows latest release v0.2.8 with ZERO assets; bootstrap requires repo archaeology (scripts found only via package.json spelunking). |
|
||||
| Scope | PASS | 5 REQs, 3 execution phases, one focused surface (apps/cli + scripts). Smallest milestone yet. macOS arm64 already cut (D-036, unverifiable here). |
|
||||
| Feasibility | CONCERN (fixed) | SEA flag exists on node v24.15.0, but no end-to-end SEA binary was built during RESEARCH. postject availability assumed (`npx postject` — needs npm registry reachability, unproven). Zipapp fallback requires python3 on target — an honest-degradation ladder, not a silent downgrade. → G-101. |
|
||||
| Honest versioning | CONCERN (fixed) | `--version` from package.json would print a stale hardcoded version inside a per-release binary — breaks upgrade detection + the one-liner's re-run-to-upgrade promise. → G-102. |
|
||||
| Install integrity | CONCERN (fixed) | sed/grep JSON parsing is brittle; a parse failure must never fall through to installing an unverified artifact. Exact asset-name matching + hard-degrade to source instructions. → G-103. |
|
||||
| Sequencing | PASS | P1 CLI (source-runnable) → P2 binary+pipeline → P3 docs+E2E matches dependency order; each phase ships independently. |
|
||||
| Cost/quota | PASS | Zero new paid infra; binaries built on-box; Gitea releases free. Dev-only esbuild dep. |
|
||||
| Risks | CONCERN (fixed) | Top 3: SEA end-to-end (→ G-101 live probe FIRST in P2), npm registry reachability for esbuild (→ proven by P1's pnpm install must-have), Gitea asset-upload token scope (→ live-proven at the v0.3.2 ship itself). |
|
||||
| Adoption/operability | PASS | Consumer = founder + future pilots; one command replaces README archaeology. Rollback trivial (rm ~/.local/bin/nextcraft). No server changes. |
|
||||
|
||||
## Binding Decisions (applied to PLAN.md)
|
||||
|
||||
- **G-101 (BINDING) — SEA live-build probe is the FIRST P2 action.** Task 2-1-01 builds a real binary before anything depends on it; the build script encodes the fallback ladder explicitly (SEA → zipapp with "requires python3" honesty). If SEA fails on this box, zipapp becomes primary with the docs stating the requirement — no silent claim of node-less operation.
|
||||
- **G-102 (BINDING) — Version stamping at build time.** `build-binary` accepts the shipping tag and stamps it into the bundle (`NEXTCRAFT_VERSION` replace); `--version` prints it; install E2E asserts the installed binary reports the tag it was downloaded from. A binary may never report a version it was not built as.
|
||||
- **G-103 (BINDING) — Install-script integrity hard-degrade.** install.sh matches assets by EXACT name (`nextcraft-linux-x64`, `nextcraft-linux-x64.sha256`); any parse/lookup/download failure degrades to source-bootstrap instructions (exit 0) — never installs unverified or name-approximate artifacts. Checksum mismatch = hard stop, exit 1, explicit do-not-run message. dash-safe POSIX sh, no jq.
|
||||
- **G-104 (BINDING) — Ongoing-binaries enforcement.** Every ship from v0.3.2 onward MUST run `scripts/release-assets.sh <tag>` after tag+merge (best-effort, non-blocking, `release_pending` escalation on failure — but attempted + logged every release). The final-phase audit gate includes "milestone release carries both assets" as a check. This makes the founder's "ongoing binaries" directive a pipeline property, not prose.
|
||||
|
||||
## Escalations
|
||||
|
||||
None. All four concerns resolved at confidence ≥ 0.85. No axis requires founder escalation (directive already explicit).
|
||||
|
||||
## Outcome
|
||||
|
||||
**GO** — G-101..G-104 applied to PLAN.md before Phase 1 execution. The milestone claims only what its probes prove, and the ongoing-binaries contract has an enforcement mechanism.
|
||||
|
||||
+107
-138
@@ -2,19 +2,19 @@
|
||||
|
||||
## Persona Roster
|
||||
|
||||
> **v0.3 update (RESEARCH, lead-developer assessment):** backend-engineer territory extended to the new engine modules (telemetry/grading/variants persistence + APIs). New phase-relevant custom personas added: **sandbox-engineer** (Linux-namespace isolation infra) and **voice-engineer** (STT/TTS + Examiner agent audio pipeline). ai-engineer re-scoped to LLM/agents/prompts + grading/variant/voice *model-facing* logic. **security-auditor stays inactive** (KYC deferred per founder directive). frontend-engineer gains real-sandbox (xterm.js), live-telemetry, and live-defense surfaces.
|
||||
> **v0.4 update (RESEARCH, lead-developer assessment):** milestone pivoted to Distribution & Bootstrap CLI (founder directive D-016). New custom persona **cli-engineer** (domain `cli`) owns apps/cli end-to-end: doctor/bootstrap/verify/dev commands, checks, spawn wrappers, the SEA binary build, the one-liner install script, and the release-asset pipeline. backend-engineer retains the scripts/ + turbo/root-package integration surface. **sandbox-engineer and voice-engineer deactivated** (their v0.3 code is complete and untouched this milestone — reason fields below). ai-engineer light-touch (no model-facing work in v0.4). **security-auditor re-activated (phase-specific)** for the install pipeline: curl|bash attack surface, checksum trust, PATH writes, secrets handling in the release flow. frontend-engineer/design-system-engineer/data-engineer inactive (zero UI/data-scope tasks in v0.4 — retained below with reasons).
|
||||
|
||||
### lead-developer
|
||||
```yaml
|
||||
active: true
|
||||
phase_specific: false
|
||||
reason: Coordinates task decomposition across web, AI service, engine, and sandbox territories; resolves conflicts between frontend, backend, AI, sandbox, and voice personas
|
||||
reason: Coordinates task decomposition across CLI, scripts, release-pipeline, and docs territories; resolves cli-engineer/backend-engineer boundary (scripts vs CLI)
|
||||
domain: coordination
|
||||
frameworks:
|
||||
- next.js
|
||||
- turborepo
|
||||
- pnpm
|
||||
- fastapi
|
||||
- node
|
||||
constraints:
|
||||
- pragmatic
|
||||
- battle-tested defaults
|
||||
@@ -27,99 +27,83 @@ territory:
|
||||
- "apps/ai-service/pyproject.toml"
|
||||
```
|
||||
|
||||
### frontend-engineer
|
||||
### cli-engineer
|
||||
```yaml
|
||||
active: true
|
||||
phase_specific: false
|
||||
reason: Phase 6 real-engine learner-surface integration — xterm.js in-browser terminal, file-tree/run/test controls, live telemetry panels, live voice defense UI, live grading display. Owns all page components, layouts, surface-specific UI.
|
||||
domain: frontend
|
||||
frameworks:
|
||||
- react
|
||||
- next.js
|
||||
- tailwindcss
|
||||
- lucide-react
|
||||
- recharts
|
||||
- react-flow
|
||||
- "@xterm/xterm"
|
||||
- "@xterm/addon-fit"
|
||||
constraints:
|
||||
- component-first
|
||||
- server-components-default
|
||||
- minimal-client-js
|
||||
- sse-client-buffering (buffer bytes, split frames on \n\n, join data: lines)
|
||||
- abortcontroller-cleanup (idempotent abort in effect cleanup)
|
||||
- websocket-lifecycle (typed messages, reconnect backoff, cleanup)
|
||||
- mediarecorder-permission-ux (mic consent, graceful no-mic fallback)
|
||||
- responsive-all-breakpoints
|
||||
- dark-mode-support
|
||||
territory:
|
||||
- "apps/web/**"
|
||||
- "packages/ui/**"
|
||||
- "packages/mock-data/**"
|
||||
- "packages/types/**"
|
||||
```
|
||||
|
||||
### data-engineer
|
||||
```yaml
|
||||
active: true
|
||||
phase_specific: false
|
||||
reason: Owns TS mock data layer schema and typed definitions + TS types for telemetry/trace/grade/variant/defense shapes the web surfaces consume. Does NOT own the Python corpus or engine stores — aligned by convention (D-021).
|
||||
domain: data
|
||||
reason: v0.4 custom persona (RESEARCH) — owns the distribution milestone core: nextcraft CLI (doctor/bootstrap/verify/dev), pure check logic, spawn wrappers with timeouts, Node SEA binary build (D-033), one-liner install.sh (D-035), checksum sidecar, and Gitea release-asset upload (D-036)
|
||||
domain: cli
|
||||
frameworks:
|
||||
- node
|
||||
- typescript
|
||||
- node:test
|
||||
- esbuild
|
||||
- node-sea
|
||||
- posix-sh
|
||||
constraints:
|
||||
- schema-first
|
||||
- type-safe
|
||||
- migration-ready
|
||||
- mock-data-only
|
||||
- stdlib-only-runtime (no runtime npm deps; esbuild dev-only)
|
||||
- thin-wrapper (never re-implement scripts/bootstrap.sh or dev.sh — compose via spawn, A-202/A-209)
|
||||
- timeout-every-spawn (no unbounded subprocess)
|
||||
- actionable-errors (every failed check tells the user how to fix it)
|
||||
- graceful-degradation (install never hard-fails; source-bootstrap fallback, A-206)
|
||||
- checksum-before-install (sha256 verify before chmod+install, A-207)
|
||||
- secrets-never-in-cli (no key generation; .env.example -> .env copy only, A-210)
|
||||
- fail-loud-exit-codes (0 ok / 1 failure / 2 usage)
|
||||
territory:
|
||||
- "packages/types/**"
|
||||
- "packages/mock-data/**"
|
||||
- "apps/cli/**"
|
||||
- "scripts/install.sh"
|
||||
- "scripts/release-assets.sh"
|
||||
```
|
||||
|
||||
### backend-engineer
|
||||
```yaml
|
||||
active: true
|
||||
phase_specific: false
|
||||
reason: Owns apps/ai-service app shell, config, API endpoints (incl. WebSocket telemetry ingest), engine persistence (SQLite stores), scripts, and test harness. Extended for v0.3 engine modules.
|
||||
reason: Owns the script + monorepo integration surface the CLI composes: apps/ai-service/scripts/*, root package.json cli:* passthrough scripts, turbo task wiring (D-037/D-022). Python ai-service itself is untouched this milestone (v0.3 complete).
|
||||
domain: backend
|
||||
frameworks:
|
||||
- fastapi
|
||||
- uvicorn
|
||||
- pydantic
|
||||
- pydantic-settings
|
||||
- httpx
|
||||
- pytest
|
||||
- sqlmodel
|
||||
- sqlalchemy
|
||||
- websockets
|
||||
- aiofiles
|
||||
- bash
|
||||
- turborepo
|
||||
- pnpm
|
||||
constraints:
|
||||
- provider-agnostic-boundaries (engine modules import nothing from agents/ or api/)
|
||||
- streaming-first
|
||||
- sqlite-first-persistence (protocol-wrapped stores, Postgres-ready, D-027)
|
||||
- secrets-via-env-only
|
||||
- mock-provider-in-tests
|
||||
- websocket-contract (typed envelopes, seq gap detection, D-026)
|
||||
- scripts-are-truth (bootstrap.sh/dev.sh stay the single source of bootstrap orchestration; CLI only wraps)
|
||||
- idempotent-scripts (re-runnable without side effects)
|
||||
- secrets-via-env-only (D-014; dev.sh exports from .ciagent/.env.secrets)
|
||||
territory:
|
||||
- "apps/ai-service/ai_service/main.py"
|
||||
- "apps/ai-service/ai_service/config.py"
|
||||
- "apps/ai-service/ai_service/api/**"
|
||||
- "apps/ai-service/ai_service/telemetry/store.py"
|
||||
- "apps/ai-service/ai_service/telemetry/ingest.py"
|
||||
- "apps/ai-service/ai_service/grading/store.py"
|
||||
- "apps/ai-service/ai_service/variants/store.py"
|
||||
- "apps/ai-service/scripts/**"
|
||||
- "apps/ai-service/package.json"
|
||||
- "apps/ai-service/tests/api/**"
|
||||
- "package.json"
|
||||
- "turbo.json"
|
||||
- ".gitignore"
|
||||
```
|
||||
|
||||
### security-auditor
|
||||
```yaml
|
||||
active: true
|
||||
phase_specific: true
|
||||
reason: v0.4 re-activated (phase-specific) — the install pipeline is the first externally-consumed attack surface: curl|bash piping, latest-release resolution, checksum trust root, PATH writes to ~/.local/bin, download tempdir hygiene, release-asset upload token handling. No KYC/PII work (still v0.5).
|
||||
domain: security
|
||||
frameworks:
|
||||
- posix-sh
|
||||
- curl
|
||||
- sha256sum
|
||||
constraints:
|
||||
- STRIDE-classified
|
||||
- no-pipe-to-shell-without-checksum (download -> verify -> install order)
|
||||
- tmpdir-safe (mktemp, no predictable paths, trap cleanup)
|
||||
- token-never-echoed (release upload resolves .env* only, never logs)
|
||||
territory:
|
||||
- "scripts/install.sh"
|
||||
- "scripts/release-assets.sh"
|
||||
- "apps/cli/src/lib/spawn.ts"
|
||||
```
|
||||
|
||||
### ai-engineer
|
||||
```yaml
|
||||
active: true
|
||||
phase_specific: false
|
||||
reason: Owns the LLM provider layer, agent framework, prompt library, structured outputs, and the model-facing logic of v0.3 engines — trace-digest→rubric grading prompts (grading/features.py+engine.py), variant instantiation (variants/templates.py+generator.py), and the Examiner agent. Owns the deterministic-mock corpora.
|
||||
reason: Light-touch v0.4 — no model-facing work in the distribution milestone; retained to guard the CLI against touching agent/engine boundaries and to keep territory mappings accurate for v0.5 (voice real-path, seq-lease).
|
||||
domain: ai
|
||||
frameworks:
|
||||
- pydantic
|
||||
@@ -127,116 +111,101 @@ frameworks:
|
||||
- pytest
|
||||
constraints:
|
||||
- provider-agnostic-protocol
|
||||
- prompts-are-code
|
||||
- json-defensive-parsing
|
||||
- never-call-cloud-in-tests
|
||||
- delta-passthrough
|
||||
- llm-sees-digest-not-raw-trace (D-028)
|
||||
- seeded-variant-reproducibility (D-029)
|
||||
territory:
|
||||
- "apps/ai-service/ai_service/llm/**"
|
||||
- "apps/ai-service/ai_service/agents/**"
|
||||
- "apps/ai-service/ai_service/prompts/**"
|
||||
- "apps/ai-service/ai_service/corpus/**"
|
||||
- "apps/ai-service/ai_service/grading/features.py"
|
||||
- "apps/ai-service/ai_service/grading/engine.py"
|
||||
- "apps/ai-service/ai_service/variants/templates.py"
|
||||
- "apps/ai-service/ai_service/variants/generator.py"
|
||||
- "apps/ai-service/tests/llm/**"
|
||||
- "apps/ai-service/tests/agents/**"
|
||||
```
|
||||
|
||||
### sandbox-engineer
|
||||
### frontend-engineer
|
||||
```yaml
|
||||
active: true
|
||||
phase_specific: true
|
||||
reason: v0.3 custom persona (RESEARCH) — owns the sandbox fabric: SandboxBackend protocol, unshare-based Linux user/mount/pid/net namespace spawner, per-sandbox workdir, resource limits, lifecycle manager, concurrency guard, and the in-sandbox capture agent. Probe-verified isolation on this box (D-024).
|
||||
domain: infra
|
||||
active: false
|
||||
phase_specific: false
|
||||
reason: v0.4 has zero UI-scope work (no web/pages/components changes planned in the distribution milestone); v0.3 surfaces are complete. Reactivated at v0.5 when deferred UX work resumes.
|
||||
domain: frontend
|
||||
frameworks:
|
||||
- python
|
||||
- linux-namespaces
|
||||
- asyncio
|
||||
- pytest
|
||||
- react
|
||||
- next.js
|
||||
- tailwindcss
|
||||
constraints:
|
||||
- isolation-verified (probe must show in-ns uid=0, network isolated, writes to workdir only)
|
||||
- backend-protocol-swap (no containerd assumption; D-024)
|
||||
- resource-limits-enforced (cpu/mem/time quotas observable)
|
||||
- no-daemon (subprocess-only; no docker/containerd service)
|
||||
- capacity-guard (1-5 concurrent; 503 when full, D-032)
|
||||
- component-first
|
||||
- server-components-default
|
||||
territory:
|
||||
- "apps/ai-service/ai_service/sandbox/**"
|
||||
- "apps/ai-service/scripts/sandbox-agent.py"
|
||||
- "apps/ai-service/tests/sandbox/**"
|
||||
```
|
||||
|
||||
### voice-engineer
|
||||
```yaml
|
||||
active: true
|
||||
phase_specific: true
|
||||
reason: v0.3 custom persona (RESEARCH) — owns the voice layer: VoiceProvider protocol, STT/TTS against a compatible endpoint, deterministic mock (tests never call a voice API), browser-native fallback, and the media-path wiring consumed by the Examiner agent and assessment UI.
|
||||
domain: ai-media
|
||||
frameworks:
|
||||
- pydantic
|
||||
- httpx
|
||||
- pytest
|
||||
- web-mediarecorder
|
||||
constraints:
|
||||
- provider-agnostic-protocol (D-030)
|
||||
- never-call-voice-api-in-tests
|
||||
- browser-native-fallback (no-key path still functions)
|
||||
- bounded-turn-latency (conversational feel budget)
|
||||
territory:
|
||||
- "apps/ai-service/ai_service/voice/**"
|
||||
- "apps/ai-service/tests/voice/**"
|
||||
- "apps/web/**"
|
||||
- "packages/ui/**"
|
||||
```
|
||||
|
||||
### design-system-engineer
|
||||
```yaml
|
||||
active: true
|
||||
active: false
|
||||
phase_specific: false
|
||||
reason: Owns the shared component library, design tokens, and visual consistency. v0.3 duty: new primitives for the real build/assessment surfaces (terminal frame, telemetry status indicator, mic/record control, grade badge, defense transcript viewer).
|
||||
reason: No design-token or primitive work in v0.4; roster retained for v0.5.
|
||||
domain: frontend
|
||||
frameworks:
|
||||
- tailwindcss
|
||||
- storybook
|
||||
- lucide-react
|
||||
constraints:
|
||||
- design-token-driven
|
||||
- wcag-aa-contrast
|
||||
- dark-mode-required
|
||||
- consistent-across-surfaces
|
||||
territory:
|
||||
- "packages/ui/**"
|
||||
```
|
||||
|
||||
### security-auditor
|
||||
### data-engineer
|
||||
```yaml
|
||||
active: false
|
||||
phase_specific: false
|
||||
reason: Identity/age-gating (KYC) deferred beyond v0.3 per founder directive (A-110) — no real auth or PII backend lands this milestone. Security coverage remains: verifier's STRIDE layer + Phase 7 secrets-hygiene checklist (keys absent from code/logs/commits/errors, localhost-only CORS, no PII in prompts). Sandbox isolation safety is owned by sandbox-engineer's probe-verified constraint.
|
||||
domain: security
|
||||
reason: No schema/mock-data work in v0.4; types packages untouched. Reactivated if CLI surfaces need shared types (not planned — CLI is self-contained).
|
||||
domain: data
|
||||
frameworks:
|
||||
- typescript
|
||||
constraints:
|
||||
- schema-first
|
||||
- type-safe
|
||||
territory:
|
||||
- "packages/types/**"
|
||||
- "packages/mock-data/**"
|
||||
```
|
||||
|
||||
### sandbox-engineer
|
||||
```yaml
|
||||
active: false
|
||||
phase_specific: false
|
||||
reason: v0.3 persona — sandbox fabric shipped complete (v0.2.x series); v0.4 touches no sandbox code. doctor only *checks* unshare availability; no sandbox logic changes. Reactivated at v0.5 (design/sim environments).
|
||||
domain: infra
|
||||
frameworks:
|
||||
- python
|
||||
- linux-namespaces
|
||||
constraints: []
|
||||
territory:
|
||||
- "apps/ai-service/ai_service/sandbox/**"
|
||||
```
|
||||
|
||||
### voice-engineer
|
||||
```yaml
|
||||
active: false
|
||||
phase_specific: false
|
||||
reason: v0.3 persona — voice defense shipped complete (mock-first, CUT-1); real server STT/TTS moved to v0.5 per D-016. No v0.4 voice work.
|
||||
domain: ai-media
|
||||
frameworks: []
|
||||
constraints: []
|
||||
territory: []
|
||||
territory:
|
||||
- "apps/ai-service/ai_service/voice/**"
|
||||
```
|
||||
|
||||
## Phase-Specific Personas
|
||||
|
||||
| Persona | Phases | Removed After |
|
||||
|---------|--------|---------------|
|
||||
| sandbox-engineer | 1 (primary), 2, 6 | persists while sandbox fabric exists |
|
||||
| voice-engineer | 5 (primary), 6 | persists while voice defense exists |
|
||||
| security-auditor | 2 (primary: install pipeline), 3, 4 (final review) | milestone complete |
|
||||
|
||||
All other active personas span the entire milestone. data-engineer and design-system-engineer are light-touch outside their phases.
|
||||
All other personas span the milestone. Deactivated personas receive no tasks.
|
||||
|
||||
## Territory Conflict Resolution
|
||||
|
||||
| Conflict | Resolution |
|
||||
|----------|------------|
|
||||
| frontend-engineer vs data-engineer (packages/types, packages/mock-data) | data-engineer owns type definitions and mock data schema (incl. new telemetry/grade/variant/defense TS types); frontend-engineer consumes them. |
|
||||
| frontend-engineer vs design-system-engineer (packages/ui) | design-system-engineer owns design tokens and primitive components (terminal frame, mic control, grade badge); frontend-engineer owns composite components and page-level UI. |
|
||||
| ai-engineer vs backend-engineer (grading/variants) | ai-engineer owns the model-facing files (features/engine/templates/generator = LLM logic + prompts); backend-engineer owns the persistence stores + API endpoints. Boundary: stores are pure SQLite; engine logic is pure compute. |
|
||||
| sandbox-engineer vs backend-engineer (sandbox/) | sandbox-engineer owns `ai_service/sandbox/**` + capture agent; backend-engineer owns the API route that composes `sandbox/manager.py` via DI. manager.py has a narrow typed interface consumed by api/. |
|
||||
| voice-engineer vs ai-engineer (Examiner agent) | ai-engineer owns `agents/examiner.py` + its prompt; voice-engineer owns `voice/**` (audio in/out). Examiner calls `voice/` through the `VoiceProvider` protocol — never imports concrete providers. |
|
||||
| ai-engineer vs data-engineer (mock duplication) | ai-engineer owns `ai_service/corpus/` (Python); data-engineer owns `packages/mock-data` (TS). Shared IDs/shapes aligned by convention (D-021). |
|
||||
| lead-developer vs any | lead-developer coordinates only, does not directly modify code files. |
|
||||
| cli-engineer vs backend-engineer (scripts/) | backend-engineer owns `apps/ai-service/scripts/**` + root `package.json`/`turbo.json` wiring; cli-engineer owns `apps/cli/**` + top-level `scripts/install.sh` + `scripts/release-assets.sh` and *consumes* backend scripts via spawn — never edits them |
|
||||
| cli-engineer vs security-auditor (install.sh) | cli-engineer implements; security-auditor reviews + may patch security defects directly in install.sh/spawn.ts (its territory) |
|
||||
| lead-developer vs any | lead-developer coordinates only, does not directly modify code files |
|
||||
+194
-422
@@ -1,438 +1,210 @@
|
||||
# Nextcraft v0.3 — PLAN.md
|
||||
# Nextcraft v0.4 — PLAN.md
|
||||
|
||||
## Overview
|
||||
|
||||
This plan covers execution phases 1-6 of milestone v0.3 (Credential Engines): the real engines that replace v0.2's mock inputs — a namespace-isolated sandbox fabric, live build telemetry over WebSocket + SQLite, a process-trace grading engine, seeded per-learner variant task generation, and an oral/voice defense with a seventh Examiner agent — plus re-grounding the Lab/Assessor/Proctor agents onto real engine inputs and wiring the v0.1 learner surfaces to the real build/defense/grading paths. Phases are strictly sequential (P1→P6); within each phase, Wave 1 tasks are parallelizable (no cross-file dependencies) and later waves depend on earlier ones.
|
||||
This plan covers execution phases 1–3 of milestone v0.4 (Distribution & Bootstrap CLI) plus the final phase (P4 review+ship). The milestone delivers the founder directive (D-016): a streamlined install for Nextcraft — a `nextcraft` bootstrap CLI shipped as a linux x64 binary, installed via a one-liner script, with binaries published on **every ongoing release** from v0.4 onward. Phases are strictly sequential (P1→P3); within each phase, Wave 1 tasks are parallelizable (no cross-file dependencies) and later waves depend on earlier ones.
|
||||
|
||||
**Environment facts (apply throughout):** Python 3.11.2 via `python3 -m venv` (no uv, no system pip); pnpm 12.3.4 via corepack; turborepo; ai-service port **8420**; default model `gemma4:31b` (config via `AI_TUTOR_MODEL`); ollama-cloud base `https://ollama.com/v1` (OpenAI-compatible, Bearer auth); keys live only in gitignored `.ciagent/.env.secrets` (exported by `scripts/dev.sh`) — never in code, commits, or logs; all automated tests use the deterministic mock LLM and mock voice providers and **never call cloud or voice APIs**. New ai-service deps this milestone: `sqlmodel`, `sqlalchemy`, `websockets`, `aiofiles` (all PyPI-verified; added in Task 1-1-02). SQLite data (`apps/ai-service/ai_service/data/`) and sandbox dirs (`apps/ai-service/sandboxes/`) are already gitignored. **KYC/age-gating is deferred per founder directive (A-110)** — no identity work; age-gating stays the v0.1 visual flow mockup. **GRILL scope decisions (this revision):** real server STT/TTS (`OpenAIAudioProvider`) deferred to v0.4 — voice defense is mock+browser-first (CUT-1/G-7); the interactive xterm.js shell relay is deferred to v0.4 — the build panel is run/test-buttons + read-only exec output (CUT-2/G-8), so `@xterm/*` is NOT a v0.3 dependency; sandbox abuse control (per-learner caps + allowlist, G-5) ships even though KYC is deferred; sandbox resource limits are partially enforced (memory/CPU/wall-clock + workdir-size sweep; per-sandbox pids and hard disk quota are accepted gaps, G-1/G-2).
|
||||
**Environment facts (probe-verified, apply throughout):** Go MISSING, Rust MISSING, gcc 12.2 present, **node v24.15.0 x64 linux (SEA-capable)**, python3 3.11.2, `shasum` 6.02, pnpm 12.3.4 via corepack, turborepo 2.3.3, tsx 4.23 in root devDeps path. Gitea API verified live at `https://git.coreci.dev/api/v1` (latest release v0.2.8, **zero assets** — the gap this milestone closes). Existing orchestration: `apps/ai-service/scripts/bootstrap.sh` (idempotent venv+pip incl. the no-ensurepip get-pip path), `apps/ai-service/scripts/dev.sh` (secrets export → uvicorn :8420), `apps/ai-service/.env.example` (full AI_* template). Root scripts: `ai:dev/ai:test/ai:bootstrap/ai:lint` turbo passthroughs (D-022 pattern to mirror as `cli:*`). Secrets live only in gitignored `.ciagent/.env.secrets` (GITEA_TOKEN, OLLAMA_API_KEY, OLLAMA_BASE_URL) — never in code, commits, or logs; tests never call the cloud or the forge (mocks/fixtures only).
|
||||
|
||||
**Milestone type:** feature. Tags: phase 0 → **v0.3.0**, P1 → v0.3.1, P2 → v0.3.2, P3 → v0.3.3, final phase P4 → **v0.3.4 = milestone release**. **GRILL binding decisions (this revision):** G-101 — SEA live-build probe is the FIRST P2 action (mechanism, not flag, must be proven); fallback ladder encoded honestly (zipapp requires python3 on target). G-102 — binary `--version` stamped from the shipping tag at build time (never a stale package.json version); install E2E asserts the installed binary reports its release tag. G-103 — install.sh matches assets by exact name; any parse/download failure degrades to source-bootstrap instructions (exit 0), never installs unverified artifacts; checksum mismatch = hard stop exit 1. G-104 — every ship from v0.3.2 onward runs `scripts/release-assets.sh <tag>` (best-effort, logged, non-blocking); the P4 audit gate checks the milestone release carries both assets.
|
||||
|
||||
| Phase | Name | Requirements | Waves | Personas |
|
||||
|-------|------|-------------|-------|----------|
|
||||
| 1 | Sandbox fabric | REQ-3-001, 002 | 3 | sandbox-engineer, backend-engineer, ai-engineer (W1 lint only) |
|
||||
| 2 | Live build telemetry | REQ-3-003 | 4 | sandbox-engineer, backend-engineer, data-engineer |
|
||||
| 3 | Process-trace grading engine | REQ-3-004 | 3 | ai-engineer, backend-engineer |
|
||||
| 4 | Variant task generation | REQ-3-005 | 3 | ai-engineer, backend-engineer, data-engineer |
|
||||
| 5 | Oral / voice defense | REQ-3-006 | 4 | voice-engineer, ai-engineer, backend-engineer |
|
||||
| 6 | Agent re-grounding + learner surface integration | REQ-3-007, 008 | 5 | ai-engineer, frontend-engineer, design-system-engineer, data-engineer, backend-engineer, lead-developer |
|
||||
|
||||
---
|
||||
|
||||
## Phase 1: Sandbox Fabric
|
||||
|
||||
**Requirements:** REQ-3-001, REQ-3-002
|
||||
**Goal:** `SandboxBackend` protocol + `unshare`-based namespace spawner (D-024) + lifecycle manager with concurrency guard (D-032) + per-sandbox workdir; isolation and resources probe-verified on this box; `/v1/sandboxes` API live; deps + gitignore landed
|
||||
|
||||
### Wave 1: Foundations (parallel — no shared files)
|
||||
|
||||
#### Task 1-1-01: SandboxBackend protocol + unshare spawner + probe test
|
||||
- **Persona:** sandbox-engineer — **REQ:** REQ-3-001, REQ-3-002
|
||||
- **Files:** `apps/ai-service/ai_service/sandbox/__init__.py`, `apps/ai-service/ai_service/sandbox/backend.py`, `apps/ai-service/ai_service/sandbox/workdir.py`, `apps/ai-service/ai_service/sandbox/unshare_backend.py`, `apps/ai-service/tests/sandbox/__init__.py`, `apps/ai-service/tests/sandbox/test_isolation.py`
|
||||
- **Action:** `backend.py`: `SandboxBackend` protocol + `SandboxSpec` (sandbox_id, learner_id, workdir, resource limits) + `SandboxHandle` (id, pid, workdir, created_at); `spawn(spec)`, `exec(handle, cmd)`, `snapshot(handle) -> Path`, `destroy(handle)`. `workdir.py`: per-sandbox layout under `apps/ai-service/sandboxes/<id>/` (workspace/ writable, snapshot() = recursive copy to `snapshots/<ts>/`) — no symlinks as the snapshot mechanism. `unshare_backend.py`: subprocess spawner — `unshare --user --map-root-user --mount --pid --fork --net` with the per-sandbox dir bind-mounted (`--bind <dir> /work`) and `chdir /work` (D-024); pipes for stdout/stderr; async wrappers. `test_isolation.py` — **re-verify box isolation properties (runs on this box, guarded by probe skip):** (a) `id -u` inside namespace prints `0`; (b) `ip link` inside namespace shows 0 usable interfaces (loopback-only/no carrier) — network isolated; (c) file written to `/work/inside.txt` lands at `sandboxes/<id>/workspace/inside.txt` on the host; (d) attempt to write outside the mount (e.g. host tmp path via bind) does not escape the per-sandbox dir; (e) `/proc` visibility degraded (proc-remount not permitted per A-101 — assert the probe documents this, not that it fails).
|
||||
- **Verify:** `pnpm ai:test` — `tests/sandbox/test_isolation.py` green on this box (probe-gated: skips with an explicit reason if userns unavailable); `lint` clean
|
||||
|
||||
#### Task 1-1-02: v0.3 dependencies + gitignore + config additions
|
||||
- **Persona:** backend-engineer — **REQ:** REQ-3-001
|
||||
- **Files:** `apps/ai-service/pyproject.toml` (update), `apps/ai-service/ai_service/config.py` (update), `apps/ai-service/.env.example` (update), `apps/ai-service/scripts/bootstrap.sh` (update if needed), root `package.json` (no change), `turbo.json` (no change)
|
||||
- **Action:** Add pinned deps: `sqlmodel`, `sqlalchemy`, `websockets`, `aiofiles` to pyproject. `config.py` additions (env_prefix `AI_`): `AI_SANDBOX_DIR` (default `apps/ai-service/sandboxes`), `AI_SANDBOX_MAX_CONCURRENT` (default 5, D-032), `AI_SANDBOX_CPU_LIMIT` (default 1 core / cpu.max), `AI_SANDBOX_MEM_LIMIT_MB` (default 512), `AI_SANDBOX_PIDS_LIMIT` (default 256), `AI_SANDBOX_TIMEOUT_S` (default 1800), `AI_DB_PATH` (default `ai_service/data/nextcraft.db`), `AI_VOICE_BASE_URL` / `AI_VOICE_API_KEY` / `AI_VOICE_STT_MODEL` / `AI_VOICE_TTS_MODEL` (all **optional**, default empty — mock-first, D-030; documented in `.env.example` and README). Confirm `.gitignore` already covers `ai_service/data/` + `sandboxes/` (it does — v0.3 block present). Re-run bootstrap idempotently to install new deps.
|
||||
- **Verify:** `pnpm ai:bootstrap` re-installs cleanly (no-op venv, new wheels land); `python -c "import sqlmodel, sqlalchemy, websockets, aiofiles"` succeeds in the venv; settings parse with new keys unset
|
||||
|
||||
#### Task 1-1-03: Resource-limit probe documentation + harness ruff pass
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-3-001
|
||||
- **Files:** `apps/ai-service/README.md` (update: sandbox section + probe transcript), `apps/ai-service/pyproject.toml` (no change), `apps/ai-service/tests/sandbox/test_isolation.py` (no change)
|
||||
- **Action:** Record the A-101 probe transcript in README (verbatim commands + observed output from this box: `unshare --user --map-root-user --mount --pid --fork --net id -u` → `0`; `ip link` → loopback only; write containment). Document the v0.3 resource-limit mechanism choice: cgroup-v2 delegation via per-sandbox scope files is **not available** on this box without sudo → enforcement = subprocess-level (`ulimit`-equivalent via `preexec_fn`: RLIMIT_AS for memory, RLIMIT_CPU for CPU-seconds, RLIMIT_NPROC for pids) + hard wall-clock timeout kill in the manager. This is the locked v0.3 mechanism (D-024 + no-sudo constraint). Run ruff over the new tree; fix all findings.
|
||||
- **Verify:** `pnpm ai:lint` exits 0; README shows the probe transcript and the rlimit mechanism note
|
||||
|
||||
### Wave 2: Lifecycle manager (depends on Wave 1)
|
||||
|
||||
#### Task 1-2-01: Sandbox manager + concurrency guard + snapshots
|
||||
- **Persona:** sandbox-engineer — **REQ:** REQ-3-001, REQ-3-002
|
||||
- **Files:** `apps/ai-service/ai_service/sandbox/manager.py`, `apps/ai-service/tests/sandbox/test_manager.py`
|
||||
- **Action:** `SandboxManager`: `create(learner_id) -> SandboxHandle` (guard: active count ≥ `AI_SANDBOX_MAX_CONCURRENT` → raise `PoolFullError` → API maps to **503**, D-032; no queue); `list() -> list[SandboxHandle]`; `get(id)`; `snapshot(id) -> Path` (delegates to workdir); `destroy(id)` (kill process tree, keep or purge workdir per flag); `reap_expired()` background hook for `AI_SANDBOX_TIMEOUT_S` which also performs a **workdir-size sweep**: any sandbox whose `workdir` exceeds `AI_SANDBOX_MAX_WORKDIR_MB` (new config, default 512) is snapshotted-then-destroyed and the event logged as an integrity signal (G-2 — soft disk cap, best-effort, not kernel-enforced); the sweep runs on the same timer as the timeout reaper. Enforce rlimits per spawner (Task 1-1-03: RLIMIT_AS + RLIMIT_CPU + RLIMIT_FSIZE=50MB as a cheap single-file disk guard (a-2); RLIMIT_NPROC noted as shared-per-host-uid, not relied on (G-1)) at exec time. Handle registry persisted **in-memory** (v0.3, single process; not a store — see D-019 precedent) with a clear note that handles are process-local. Startup reaper (a-1): on lifespan boot, scan `AI_SANDBOX_DIR` for workdirs whose recorded pid is dead and reap them, logging a warning. Narrow typed interface only — manager never imports api/ (boundary rule).
|
||||
- **Verify:** `pnpm ai:test` — test_manager covers create/list/destroy/snapshot, 6th create raises PoolFullError (503 path), destroy kills the namespace process (pid gone), snapshot dir exists with workspace contents, timeout reaper removes a stale handle
|
||||
|
||||
#### Task 1-2-02: Resource-limit enforcement test
|
||||
- **Persona:** sandbox-engineer — **REQ:** REQ-3-002
|
||||
- **Files:** `apps/ai-service/tests/sandbox/test_resource_limits.py`
|
||||
- **Action:** Concrete enforcement probes (guarded like isolation tests): (a) spawn a process that allocates > `RLIMIT_AS` → assert it dies with MemoryError/killed within a bound; (b) spawn a CPU spinner past `RLIMIT_CPU` → assert SIGXCPU/kill; (c) single huge file > `RLIMIT_FSIZE` → assert write failure (a-2 partial disk guard); (d) wall-clock: spawn `sleep 9999` with a small manager timeout → reaper destroys it; (e) **disk sweep (G-2)**: write > `AI_SANDBOX_MAX_WORKDIR_MB` across many files → assert the manager sweep destroys the sandbox and logs the integrity signal. Assert limits are observable (handle reports its limit set). NOTE (G-1): per-sandbox `RLIMIT_NPROC` is shared at the host uid — the fork-bomb probe is documented as shared-budget behavior, NOT asserted as per-sandbox isolation.
|
||||
- **Verify:** `pnpm ai:test` — test_resource_limits green; limits proven enforced and observable
|
||||
|
||||
### Wave 3: API exposure (depends on Wave 2)
|
||||
|
||||
#### Task 1-3-01: Sandboxes API module
|
||||
- **Persona:** backend-engineer — **REQ:** REQ-3-001, REQ-3-002
|
||||
- **Files:** `apps/ai-service/ai_service/api/sandboxes.py`, `apps/ai-service/ai_service/main.py` (update: include router + lifespan manager), `apps/ai-service/ai_service/api/deps.py` (update), `apps/ai-service/tests/api/test_sandboxes.py`
|
||||
- **Action:** DI exposes a singleton `SandboxManager`. Endpoints: `POST /v1/sandboxes {learner_id}` → 201 handle (503 when pool full); `GET /v1/sandboxes` → list; `GET /v1/sandboxes/{id}` → handle; `POST /v1/sandboxes/{id}/snapshot` → snapshot path; `DELETE /v1/sandboxes/{id}` → 204. All behind localhost CORS (A-008). Lifespan creates/destroys the manager; on shutdown destroys any live sandboxes (no orphans). **Abuse control (G-5, NOT KYC):** even in no-auth v0.3 the sandbox API enforces per-`learner_id` rate limiting (`AI_SANDBOX_MAX_PER_LEARNER`, default 1 active → 429) and a global create-rate cap (`AI_SANDBOX_CREATES_PER_MIN`, default 10 → 429); `learner_id` is validated against a server-side allowlist from config (`AI_LEARNER_ALLOWLIST`, default the single mock pilot id → unknown ids rejected 403). This ships in the no-auth milestone so a rogue local process can't exhaust shared NPROC/disk.
|
||||
- **Verify:** `pnpm ai:test` — test_sandboxes green (create→list→snapshot→delete roundtrip via TestClient; 6th create → 503; delete of unknown id → 404; abuse control (G-5): non-allowlisted learner_id → 403; >1 active sandbox for one learner → 429; burst of >10 creates/min → 429); manual probe: `curl -X POST localhost:8420/v1/sandboxes -d '{"learner_id":"l1"}'` returns a handle id
|
||||
|
||||
### Must-Haves (Phase 1)
|
||||
- [ ] Isolation probe test green on this box: in-namespace uid=0, network isolated (0 usable interfaces), host writes confined to the per-sandbox bind dir (A-101 re-verified as an automated test, not just research notes)
|
||||
- [ ] Resource limits enforced + observable: memory (RLIMIT_AS) + CPU (RLIMIT_CPU) rlimits kill violating processes; wall-clock reaper destroys stale sandboxes; disk usage capped by a periodic workdir-size sweep in the manager (soft cap, configurable `AI_SANDBOX_MAX_WORKDIR_MB`, default 512MB — NOT kernel-enforced); per-sandbox NPROC is shared across sandboxes at the host uid — documented, not relied on for isolation (G-1, G-2 — test_resource_limits green)
|
||||
- [ ] No cross-tenant access: sandbox A cannot read sandbox B's workdir (isolation test asserts containment)
|
||||
- [ ] Lifecycle API works end-to-end: create/list/snapshot/destroy via TestClient; pool full → **503** (D-032, no queue)
|
||||
- [ ] Snapshot produces a restorable directory copy under the sandbox's own snapshots/ dir
|
||||
- [ ] `pnpm ai:test` and `pnpm ai:lint` green; new deps (`sqlmodel`, `sqlalchemy`, `websockets`, `aiofiles`) installed via idempotent bootstrap
|
||||
- [ ] Boundary rules hold: `sandbox/` imports nothing from `api/` or `agents/`; only `api/sandboxes.py` composes the manager via DI
|
||||
- [ ] No docker/podman/sudo anywhere in the spawner path (D-024); `SandboxBackend` protocol is the only coupling to the spawner (containerd swap possible later)
|
||||
|
||||
---
|
||||
|
||||
## Phase 2: Live Build Telemetry
|
||||
|
||||
**Requirements:** REQ-3-003
|
||||
**Goal:** First real persistence (D-027: SQLite via SQLModel) with a `TraceStore` protocol; `TelemetryEvent` model with per-(learner,task) monotonic `seq`; WebSocket ingest endpoint (D-026) with gap detection; stdlib-only in-sandbox capture agent streams real sandbox activity into ai-service; trace retrievable by learner+task
|
||||
|
||||
### Wave 1: Models + stores + TS types (parallel — no shared files)
|
||||
|
||||
#### Task 2-1-01: Telemetry event models
|
||||
- **Persona:** backend-engineer — **REQ:** REQ-3-003
|
||||
- **Files:** `apps/ai-service/ai_service/telemetry/__init__.py`, `apps/ai-service/ai_service/telemetry/models.py`, `apps/ai-service/tests/telemetry/__init__.py`, `apps/ai-service/tests/telemetry/test_models.py`
|
||||
- **Action:** Pydantic/SQLModel `TelemetryEvent`: `learner_id`, `task_id`, `seq` (int, monotonic per (learner,task)), `kind` (`command` | `file_diff` | `run_result` | `test_result` | `activity` | `stdin` | `stdout`), `payload` (JSON), `ts` (datetime, monotonic-envelope), `sandbox_id`. `TraceSpan` derived view (ordered events for one (learner,task)). Validation: seq ≥ 0, kind enum, non-empty ids.
|
||||
- **Verify:** `pnpm ai:test` — test_models green (validation rules enforced, JSON payload roundtrip)
|
||||
|
||||
#### Task 2-1-02: TraceStore protocol + SQLite implementation
|
||||
- **Persona:** backend-engineer — **REQ:** REQ-3-003
|
||||
- **Files:** `apps/ai-service/ai_service/telemetry/store.py`, `apps/ai-service/tests/telemetry/test_store.py`, `apps/ai-service/ai_service/data/.gitkeep`
|
||||
- **Action:** `TraceStore` protocol (D-027, Postgres-migration-ready): `append(event) -> None` (idempotent on (learner,task,seq) — at-least-once dedup), `get_trace(learner_id, task_id) -> list[TelemetryEvent]` (ordered by seq), `gaps(learner_id, task_id) -> list[int]` (missing seqs), `latest_seq(learner_id, task_id) -> int`, `list_tasks(learner_id) -> list[str]`, `close()`. `SQLiteTraceStore(SQLModel)`: single `telemetry_event` table, composite PK ((learner_id, task_id, seq)), indexes on (learner_id, task_id). Engine creation from `AI_DB_PATH`; `SQLModel.metadata.create_all` at app lifespan. Enable `PRAGMA journal_mode=WAL` + `synchronous=NORMAL` at engine creation (a-3) so concurrent ingest (writer) and trace reads (grader) don't hit `database is locked` under concurrent sandboxes.
|
||||
- **Verify:** `pnpm ai:test` — test_store green (append/ordered-get/dedup-on-retry/gap detection/latest_seq; tmp-path SQLite per test)
|
||||
|
||||
#### Task 2-1-03: TS types for telemetry/traces
|
||||
- **Persona:** data-engineer — **REQ:** REQ-3-003
|
||||
- **Files:** `packages/types/telemetry.ts` (new), `packages/types/index.ts` (update)
|
||||
- **Action:** TS `TelemetryEvent`, `TraceSpan`, `TelemetryKind` mirroring the Python model field-for-field (cross-referencing header, same string enums). Consumed by Phase 6 web surfaces; no runtime code.
|
||||
- **Verify:** `pnpm typecheck` passes; TS type keys match Python model keys exactly
|
||||
|
||||
### Wave 2: Capture agent + ingest (depends on Wave 1)
|
||||
|
||||
#### Task 2-2-01: In-sandbox capture agent (stdlib-only)
|
||||
- **Persona:** sandbox-engineer — **REQ:** REQ-3-003
|
||||
- **Files:** `apps/ai-service/scripts/sandbox-agent.py`, `apps/ai-service/tests/sandbox/test_sandbox_agent.py`
|
||||
- **Action:** Tiny standalone process (D-031, **stdlib only** — no deps shipped into the namespace): wraps a shell inside the sandbox; captures commands, file diffs (mtime/content polling of `workspace/` at 250ms), run/test results, activity; assigns per-(learner,task) `seq`; buffers to a local spool file on disconnect (at-least-once, D-026); reconnects with **exponential backoff** and flushes spool in order; small WebSocket client implemented over raw `socket` (RFC6455 client handshake + frames — stdlib only, no `websockets` in-namespace). Configured via env baked at spawn (`NC_LEARNER_ID`, `NC_TASK_ID`, `NC_INGEST_URL`).
|
||||
- **Verify:** `pnpm ai:test` — unit tests with a loopback fake WS server: ordered seq emission, spool-on-disconnect, reconnect flush preserves order (no loss, dupes deduped server-side), no third-party imports in the file (asserted by AST scan)
|
||||
|
||||
#### Task 2-2-02: WebSocket ingest endpoint
|
||||
- **Persona:** backend-engineer — **REQ:** REQ-3-003
|
||||
- **Files:** `apps/ai-service/ai_service/telemetry/ingest.py`, `apps/ai-service/ai_service/api/sandboxes.py` (update: register WS route), `apps/ai-service/ai_service/main.py` (update), `apps/ai-service/tests/api/test_telemetry_ingest.py`
|
||||
- **Action:** `WS /v1/telemetry/ingest` (D-026): accepts connections carrying learner_id/task_id/sandbox_id; validates + appends events to `TraceStore` (idempotent — server-side dedup on (learner,task,seq)); emits **gap warnings** when seq skips (logged + surfaced in a per-connection status); ping/pong keepalive. **Backpressure / flood control (G-3 — replaces silent drop):** bounded inbound queue; on overflow OR when a per-connection cap `AI_TELEMETRY_MAX_EVENTS_PER_TASK` (default 50000) is exceeded → **reject with a 1008 policy-violation close and mark the (learner,task) trace `INCOMPLETE_FLOODED`** (an integrity signal consumed by Proctor). Silent drop-oldest is FORBIDDEN because it corrupts grading input and is indistinguishable from trace-gaming. `GET /v1/telemetry/traces/{learner_id}/{task_id}` returns the ordered trace; `GET /v1/telemetry/gaps/{learner_id}/{task_id}` returns missing seqs.
|
||||
- **Verify:** `pnpm ai:test` — test_telemetry_ingest green (TestClient websocket: connect → send 3 events → trace retrievable ordered; resend event 2 → deduped; skip seq 5 → gap reported; unknown sandbox tolerated in v0.3 no-auth mode)
|
||||
|
||||
### Wave 3: Sandbox telemetry wiring (depends on Wave 2)
|
||||
|
||||
#### Task 2-3-01: Spawn sandboxes with the capture agent
|
||||
- **Persona:** sandbox-engineer — **REQ:** REQ-3-003
|
||||
- **Files:** `apps/ai-service/ai_service/sandbox/manager.py` (update), `apps/ai-service/ai_service/sandbox/unshare_backend.py` (update), `apps/ai-service/tests/sandbox/test_telemetry_wiring.py`
|
||||
- **Action:** `create()` gains optional `task_id`; when set the spawner copies `scripts/sandbox-agent.py` into the sandbox workdir, injects `NC_*` env, and launches the agent as a child of the namespace process (agent lifecycle tied to sandbox lifecycle; destroy kills the agent). No capture when task_id absent (pure shell sandbox).
|
||||
- **Verify:** `pnpm ai:test` — end-to-end on this box: create sandbox with task_id → run 2 commands via exec → events arrive at the ingest endpoint and land in SQLite in order
|
||||
|
||||
### Wave 4: Reliability probe (depends on Wave 3)
|
||||
|
||||
#### Task 2-4-01: Dropped-connection durability probe
|
||||
- **Persona:** backend-engineer — **REQ:** REQ-3-003
|
||||
- **Files:** `apps/ai-service/tests/telemetry/test_durability.py`
|
||||
- **Action:** Integration probe: run a capture agent against ingest, kill the WS connection mid-stream (simulate network failure), keep generating events, reconnect, assert the SQLite trace contains **every** event exactly once in order (spool + dedup). Document at-least-once semantics + replay path in README.
|
||||
- **Verify:** `pnpm ai:test` — test_durability green; README documents semantics
|
||||
|
||||
### Must-Haves (Phase 2)
|
||||
- [ ] Real telemetry from a live sandbox arrives at ai-service: shell commands, file diffs, run/test results appear as ordered events in SQLite (end-to-end, no mocks)
|
||||
- [ ] Per-(learner,task) monotonic `seq`; gap detection reports missing seqs; replay yields the complete ordered trace
|
||||
- [ ] At-least-once proven: transient disconnect + reconnect loses no events; duplicates deduped server-side (durability probe green)
|
||||
- [ ] Capture agent is stdlib-only (AST-verified) and its lifecycle is tied to the sandbox (destroy kills it)
|
||||
- [ ] Trace retrievable by learner+task via `GET /v1/telemetry/traces/...`; unknown trace → 404
|
||||
- [ ] `TraceStore` protocol respected: no api/ code touches SQLite directly; `telemetry/` never imports `agents/` (D-027)
|
||||
- [ ] Flood control (G-3): burst past `AI_TELEMETRY_MAX_EVENTS_PER_TASK` → connection closed 1008 + trace marked `INCOMPLETE_FLOODED`; no silent event drop on overflow
|
||||
- [ ] `pnpm ai:test` green; `packages/types` telemetry TS types compile (`pnpm typecheck`)
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: Process-Trace Grading Engine
|
||||
|
||||
**Requirements:** REQ-3-004
|
||||
**Goal:** Deterministic feature computation over traces (D-028) → compact digest → LLM rubric scoring via existing D-020 JSON defense → validated structured scores stored in `GradeStore`; LLM never sees the raw trace; engine calibrated against v0.2 mock corpora so process quality separates paste-and-run from iterative debugging
|
||||
|
||||
### Wave 1: Features + grades store (parallel — no shared files)
|
||||
|
||||
#### Task 3-1-01: Deterministic trace digest (features)
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-3-004
|
||||
- **Files:** `apps/ai-service/ai_service/grading/__init__.py`, `apps/ai-service/ai_service/grading/features.py`, `apps/ai-service/tests/grading/__init__.py`, `apps/ai-service/tests/grading/test_features.py`
|
||||
- **Action:** Pure compute module (D-028): `compute_digest(trace: list[TelemetryEvent]) -> TraceDigest`. Deterministic features: test pass/fail counts + final status; edit count; error/fix cycle count + mean fix latency; idle gaps (>Ns, count + total); command category histogram (build/test/file/nav/debug/other); session duration; first-test-pass offset. `TraceDigest` pydantic model — compact (bounded size, no raw commands), LLM-safe.
|
||||
- **Verify:** `pnpm ai:test` — test_features green over synthetic traces: paste-and-run trace (0 error/fix cycles, single test pass at end) vs iterative trace (many cycles) produce observably different digests
|
||||
|
||||
#### Task 3-1-02: GradeStore protocol + SQLite implementation
|
||||
- **Persona:** backend-engineer — **REQ:** REQ-3-004
|
||||
- **Files:** `apps/ai-service/ai_service/grading/store.py`, `apps/ai-service/tests/grading/test_store.py`
|
||||
- **Action:** `GradeStore` protocol (D-027): `save(grade)`, `get(learner_id, task_id)`, `list_for_learner(learner_id)`, `close()`. SQLModel `GradeRecord`: learner_id, task_id, variant_seed (null until P4), digest (JSON), scores (JSON), verdict, model, created_at. PK (learner_id, task_id). Postgres-migration-ready.
|
||||
- **Verify:** `pnpm ai:test` — test_store green (save/get/list roundtrip, overwrite-on-regrade documented)
|
||||
|
||||
### Wave 2: Grading engine + calibration (depends on Wave 1)
|
||||
|
||||
#### Task 3-2-01: Rubric scoring engine
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-3-004
|
||||
- **Files:** `apps/ai-service/ai_service/grading/engine.py`, `apps/ai-service/ai_service/prompts/grading.py` (new), `apps/ai-service/tests/grading/test_engine.py`
|
||||
- **Action:** `GradingEngine.grade(learner_id, task_id) -> GradeRecord`: **trace-completeness gate (G-4)** — first call `TraceStore.gaps()` + check the trace `INCOMPLETE_FLOODED` flag; if gaps are non-empty OR the trace is flagged incomplete → return `verdict=UNGRADABLE_TRACE_INCOMPLETE` (a first-class verdict, not an exception) surfacing the gap list; a credential is NEVER issued from a gapped/incomplete trace. Otherwise: load trace via `TraceStore` → `compute_digest` → render rubric prompt (`prompts/grading.py`: criteria + level anchors for process quality, correctness, debugging discipline, test usage; a-4: treat high edit/command churn with no test-progress as a process-quality negative) → LLM structured output through the **existing D-020 4-layer defense** (`agents/structured.py` reused — engine composes it, never duplicates it) → validate `RubricScore` model (per-criterion 0-4 + strengths + gaps + verdict) → persist via `GradeStore`. Grading depends on `llm/` + `telemetry/` + `prompts/` only (boundary). Mock provider scripts deterministic rubric JSON for tests, including the INCOMPLETE path.
|
||||
- **Verify:** `pnpm ai:test` — test_engine green (mock provider: digest-only prompt asserted — **raw trace string absent from prompt**; validated scores returned; malformed JSON exercises D-020 retry; unknown trace → error)
|
||||
|
||||
#### Task 3-2-02: Calibration against v0.2 mock corpora
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-3-004
|
||||
- **Files:** `apps/ai-service/ai_service/corpus/trace_fixtures.py` (new), `apps/ai-service/tests/grading/test_calibration.py`
|
||||
- **Action:** Synthetic trace fixtures aligned with v0.2 `corpus/telemetry.py` + `corpus/artifacts.py` scenario IDs (strong/lazy/struggling builder archetypes). Assert grading separates them: strong archetype scores ≥ lazy archetype on process-quality criterion (mock provider maps digest shape → scripted scores; test asserts the ordering contract + that fixture IDs align with existing corpus IDs, D-021).
|
||||
- **Verify:** `pnpm ai:test` — test_calibration enforces the ordering contract
|
||||
|
||||
### Wave 3: Grading endpoint (depends on Wave 2)
|
||||
|
||||
#### Task 3-3-01: Assessment grade endpoint (real traces)
|
||||
- **Persona:** backend-engineer — **REQ:** REQ-3-004
|
||||
- **Files:** `apps/ai-service/ai_service/api/assessment.py` (update), `apps/ai-service/tests/api/test_grading.py`
|
||||
- **Action:** `POST /v1/assessment/grade {learner_id, task_id}` → `GradingEngine` → validated `RubricScore` JSON; `GET /v1/assessment/grade/{learner_id}/{task_id}` → stored grade; unknown trace → 404. Composed via DI (api/ owns wiring; engine knows nothing of FastAPI).
|
||||
- **Verify:** `pnpm ai:test` — test_grading green (grade roundtrip via TestClient with mock provider; 404 on unknown; GET after POST returns same scores)
|
||||
|
||||
### Must-Haves (Phase 3)
|
||||
- [ ] Engine emits structured rubric-aligned scores from a **real process trace** (not pre-baked input) — TestClient roundtrip green
|
||||
- [ ] Deterministic features computed in code (test pass/fail, edit count, error/fix cycles, idle gaps, command categories); LLM receives the **digest only** — test asserts the raw trace never reaches the prompt (D-028)
|
||||
- [ ] Scores distinguish process quality: iterative-debugging archetype out-scores paste-and-run on the process criterion (calibration contract test)
|
||||
- [ ] Grades persisted + retrievable by learner+task via GradeStore (SQLite, protocol-wrapped, D-027)
|
||||
- [ ] Boundary rules hold: `grading/` imports no `api/`/`agents/` internals except the shared D-020 structured defense; `pnpm ai:test` + `pnpm ai:lint` green
|
||||
- [ ] Incomplete-trace gate (G-4): gapped or `INCOMPLETE_FLOODED` trace → `verdict=UNGRADABLE_TRACE_INCOMPLETE` with the gap list; no credential issued from an incomplete trace (test green)
|
||||
|
||||
---
|
||||
|
||||
## Phase 4: Variant Task Generation
|
||||
|
||||
**Requirements:** REQ-3-005
|
||||
**Goal:** Template library with typed parameter slots (D-029) + seeded LLM instantiation + per-learner variant registry (SQLite) with difficulty-normalization anchors; two learners on the same competency get provably distinct, reproducible, auditable tasks
|
||||
|
||||
### Wave 1: Templates + store (parallel — no shared files)
|
||||
|
||||
#### Task 4-1-01: Task template library
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-3-005
|
||||
- **Files:** `apps/ai-service/ai_service/variants/__init__.py`, `apps/ai-service/ai_service/variants/templates.py`, `apps/ai-service/tests/variants/__init__.py`, `apps/ai-service/tests/variants/test_templates.py`
|
||||
- **Action:** ≥3 initial task templates bound to existing competency IDs (D-021 alignment). `TaskTemplate`: id, competency_id, statement skeleton with `{slot}` placeholders, `ParameterSlot[]` (name, type: enum/int-range/string-set, allowed values), `rubric anchors` (difficulty normalization: expected feature envelope — e.g. expected edit-count band — used by grading context), starter-file scaffolds served to the sandbox. Seeded slot sampler is pure code (`random.Random(seed)`), fully reproducible.
|
||||
- **Verify:** `pnpm ai:test` — test_templates green (slot validation: bad value rejected; seeded sampling reproducible across runs; all templates bind to real competency IDs)
|
||||
|
||||
#### Task 4-1-02: VariantStore protocol + SQLite implementation
|
||||
- **Persona:** backend-engineer — **REQ:** REQ-3-005
|
||||
- **Files:** `apps/ai-service/ai_service/variants/store.py`, `apps/ai-service/tests/variants/test_store.py`
|
||||
- **Action:** `VariantStore` protocol (D-027): `save(variant)`, `get(learner_id, template_id_or_task_id)`, `list_for_learner(learner_id)`, `list_by_template(template_id)`, `close()`. SQLModel `VariantRecord`: learner_id, task_id (the grading/telemetry task key), template_id, seed, params (JSON), statement (rendered), created_at. Unique (learner_id, template_id).
|
||||
- **Verify:** `pnpm ai:test` — test_store green (roundtrip, unique constraint, audit listing)
|
||||
|
||||
### Wave 2: Generator (depends on Wave 1)
|
||||
|
||||
#### Task 4-2-01: Seeded LLM variant generator
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-3-005
|
||||
- **Files:** `apps/ai-service/ai_service/variants/generator.py`, `apps/ai-service/ai_service/prompts/variant.py` (new), `apps/ai-service/tests/variants/test_generator.py`
|
||||
- **Action:** `VariantGenerator.generate(learner_id, template_id) -> VariantRecord`: derive seed (`sha256(template_id|learner_id|milestone)` — reproducible, D-029); sample typed slots in code; render a fill prompt (statement skeleton + concrete slot values) → LLM via D-020 structured defense → unique task statement + starter files → validate → persist (seed + params + statement) via `VariantStore`. Cache: existing (learner,template) returns the stored variant (no duplicate work). Mock provider scripts deterministic statements per seed for tests.
|
||||
- **Verify:** `pnpm ai:test` — test_generator green: two different learner_ids → distinct statements for the same template; same learner twice → identical stored variant (reproducible); params JSON contains only schema-valid slot values; **fairness envelope (a-5):** two variants of one template compute digests within the template's expected feature envelope (comparable slot complexity/difficulty features) — "same bar" is testable, not asserted
|
||||
|
||||
### Wave 3: Variant endpoint + TS types (depends on Wave 2)
|
||||
|
||||
#### Task 4-3-01: Variant task endpoint
|
||||
- **Persona:** backend-engineer — **REQ:** REQ-3-005
|
||||
- **Files:** `apps/ai-service/ai_service/api/variants.py` (new), `apps/ai-service/ai_service/main.py` (update), `apps/ai-service/tests/api/test_variants.py`
|
||||
- **Action:** `POST /v1/variants {learner_id, template_id or competency_id}` → generated (or cached) variant: statement, starter files, task_id, seed; `GET /v1/variants/{task_id}` → stored variant; `GET /v1/variants?learner_id=` → learner's variants. DI wiring in api/ only.
|
||||
- **Verify:** `pnpm ai:test` — test_variants green (generate→get roundtrip; cache hit on regenerate; distinct learners → distinct statements asserted at the API layer)
|
||||
|
||||
#### Task 4-3-02: TS types for variants (+ grades)
|
||||
- **Persona:** data-engineer — **REQ:** REQ-3-005
|
||||
- **Files:** `packages/types/variants.ts` (new), `packages/types/grading.ts` (new), `packages/types/index.ts` (update)
|
||||
- **Action:** TS `TaskVariant`, `VariantParams`, `RubricScore`, `GradeRecord` matching Python models (cross-referencing headers; same field names). Consumed by Phase 6 surfaces.
|
||||
- **Verify:** `pnpm typecheck` passes
|
||||
|
||||
### Must-Haves (Phase 4)
|
||||
- [ ] Two learners requesting the same competency receive **provably distinct** task variants (API-level test)
|
||||
- [ ] Seed derivation reproducible: same (template, learner) → same variant, served from cache without a second LLM call (D-029)
|
||||
- [ ] Variant seed + typed params persisted + auditable (VariantStore listing; proctoring cross-check path exists)
|
||||
- [ ] Difficulty normalization anchors present per template and shipped to the grader prompt context
|
||||
- [ ] Starter-file scaffolds defined per template (P6 wires them into the sandbox workdir)
|
||||
- [ ] `pnpm ai:test` + `pnpm typecheck` green; `variants/` imports no api/ (boundary)
|
||||
|
||||
---
|
||||
|
||||
## Phase 5: Oral / Voice Defense
|
||||
|
||||
**Requirements:** REQ-3-006
|
||||
**Goal:** `VoiceProvider` protocol with mock + browser fallback (D-030) + seventh `Examiner` agent streaming over existing SSE + `DefenseStore` persisting transcript + integrity signals. **Real server STT/TTS (`OpenAIAudioProvider`) is DEFERRED to v0.4 (with KYC, when there's a real key + real users)** — voice is mock-first (D-030) and the `/audio/*` real path could never be exercised in CI, so v0.3 proves the full defense *dialogue* + integrity-signal pipeline over mock + browser-native fallback only; the protocol seam keeps the real provider a drop-in later.
|
||||
|
||||
### Wave 1: Voice provider layer (parallel — no shared files)
|
||||
|
||||
#### Task 5-1-01: VoiceProvider protocol + mock provider + browser fallback
|
||||
- **Persona:** voice-engineer — **REQ:** REQ-3-006
|
||||
- **Files:** `apps/ai-service/ai_service/voice/__init__.py`, `apps/ai-service/ai_service/voice/base.py`, `apps/ai-service/ai_service/voice/mock.py`, `apps/ai-service/ai_service/voice/browser.py`, `apps/ai-service/ai_service/voice/factory.py`, `apps/ai-service/tests/voice/__init__.py`, `apps/ai-service/tests/voice/test_mock.py`, `apps/ai-service/tests/voice/test_factory.py`
|
||||
- **Action:** `VoiceProvider` protocol mirroring `LLMProvider` (D-030): `transcribe(audio: bytes, fmt) -> TranscriptSegment` + `synthesize(text, voice) -> AsyncIterator[bytes]`. `MockVoiceProvider`: deterministic canned transcript (scripted per test), canned 1kHz-tone WAV bytes, scripted failure modes. `browser.py`: fallback **descriptor** (`sr_available: true`, endpoint hints) the web client uses to select browser-native `SpeechRecognition`/`speechSynthesis` when no server provider. `factory.py`: `AI_VOICE_PROVIDER=browser | mock` (default mock when no key). **`OpenAIAudioProvider` (real server STT/TTS) intentionally NOT built in v0.3 — deferred to v0.4**; the protocol is its future seam. `voice/` never imports `agents/` or `api/`.
|
||||
- **Verify:** `pnpm ai:test` — test_mock + test_factory green (deterministic transcribe/synthesize; failure modes; factory selects mock with empty key, browser when provider=browser; zero network calls)
|
||||
|
||||
#### Task 5-1-03: DefenseStore protocol + SQLite implementation
|
||||
- **Persona:** backend-engineer — **REQ:** REQ-3-006
|
||||
- **Files:** `apps/ai-service/ai_service/voice/defense_store.py`, `apps/ai-service/tests/voice/test_defense_store.py`
|
||||
- **Action:** `DefenseStore` protocol (D-027): `start(defense)`, `append_turn(defense_id, turn)`, `finalize(defense_id, integrity_signals)`, `get(defense_id)`, `list_for_learner(learner_id)`, `close()`. SQLModel `DefenseRecord` (id, learner_id, task_id, status, created/finished_at) + `DefenseTurn` (defense_id FK, turn seq, role examiner|learner, text, ts, latency_ms) + integrity signals JSON on the record (long pauses, off-scope cadence markers — A-109).
|
||||
- **Verify:** `pnpm ai:test` — test_defense_store green (start→append turns→finalize→get roundtrip; ordered turns by seq)
|
||||
|
||||
### Wave 2: Examiner agent (depends on Wave 1)
|
||||
|
||||
#### Task 5-2-01: Examiner agent (seventh agent)
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-3-006
|
||||
- **Files:** `apps/ai-service/ai_service/agents/examiner.py`, `apps/ai-service/ai_service/prompts/examiner.py` (new), `apps/ai-service/ai_service/agents/registry.py` (update: central registration, G-4 pattern), `apps/ai-service/tests/agents/test_examiner.py`
|
||||
- **Action:** `ExaminerAgent(BaseAgent)` (D-030/A-109): builds questions from learner transcript + trace digest + (P4) variant statement; probes understanding + challenges process choices ("why did you choose X at step N?"); streams questions over the existing SSE pipeline; `structured` verdict mode returns verdict + per-answer integrity signal list (long pause flags, off-scope answers) computed from turn metadata; calls voice **only through the `VoiceProvider` protocol** (PERSONAS conflict rule — never concrete providers). Session-scoped history reused from v0.2.
|
||||
- **Verify:** `pnpm ai:test` — test_examiner green (question stream references trace-digest facts; verdict structured output validates via D-020 defense; registry resolves all seven agents; mock-VoiceProvider wiring through protocol only — asserted by import scan in test)
|
||||
|
||||
### Wave 3: Defense endpoints (depends on Wave 2)
|
||||
|
||||
#### Task 5-3-01: Defense session + audio endpoints
|
||||
- **Persona:** backend-engineer — **REQ:** REQ-3-006
|
||||
- **Files:** `apps/ai-service/ai_service/api/defense.py` (new), `apps/ai-service/ai_service/main.py` (update), `apps/ai-service/tests/api/test_defense.py`
|
||||
- **Action:** `POST /v1/defense/start {learner_id, task_id}` → creates DefenseRecord + streams the first examiner question (SSE, agent=examiner); `POST /v1/defense/{id}/answer` (multipart audio from browser MediaRecorder, or `{text}` for typed fallback) → STT via VoiceProvider → append learner turn → stream examiner follow-up (SSE) → TTS audio chunks over `GET /v1/defense/{id}/audio/{turn_id}`; `POST /v1/defense/{id}/finish` → verdict + integrity signals persisted; `GET /v1/defense/{id}` → full transcript + signals. Browser-fallback mode: when provider=browser, start returns the fallback descriptor instead of server audio.
|
||||
- **Verify:** `pnpm ai:test` — test_defense green (full loop with mock voice + mock LLM: start → answer(text) → answer(audio bytes) → finish → transcript retrievable with per-turn latency; unknown id → 404; `GET` signals present after finish)
|
||||
|
||||
### Wave 4: Examiner latency instrumentation (depends on Wave 3)
|
||||
|
||||
#### Task 5-4-01: Per-turn latency instrumentation (mock-based)
|
||||
- **Persona:** voice-engineer — **REQ:** REQ-3-006
|
||||
- **Files:** `apps/ai-service/tests/voice/test_latency.py`, `apps/ai-service/README.md` (update: conversational-budget doc + v0.4 voice note)
|
||||
- **Action:** Instrument per-turn latency (STT ms + LLM TTFT ms + TTS ms) recorded on each DefenseTurn; deterministic test over mock providers asserts instrumentation presence + `latency_ms` populated + budget constant defined (mock runs are near-instant — wall-clock asserted in v0.4 against a real endpoint). README documents the conversational-latency target as a v0.4 acceptance criterion (real STT/TTS deferred per CUT-1/G-7).
|
||||
- **Verify:** `pnpm ai:test` — test_latency green (latency_ms fields populated on every turn; budget constant defined); README documents the deferred real-voice acceptance probe
|
||||
|
||||
### Must-Haves (Phase 5)
|
||||
- [ ] Spoken defense runs end-to-end over HTTP with mock providers: start → answer (audio + typed fallback) → examiner follow-up streams → verdict + transcript persisted (automated)
|
||||
- [ ] Examiner is the seventh registered agent; streams over the existing SSE envelope (meta agent=examiner)
|
||||
- [ ] Instrumented per-turn latency fields populated on every DefenseTurn (STT ms + LLM TTFT ms + TTS ms); conversational budget named (A-109)
|
||||
- [ ] `VoiceProvider` protocol respected: examiner + api touch voice only via the protocol; mock-first — **no task requires a real voice key to pass**
|
||||
- [ ] Browser-native SR/TTS fallback descriptor returned when no server voice provider configured (mock/browser are first-class, D-030)
|
||||
- [ ] Real server STT/TTS (`OpenAIAudioProvider`) explicitly deferred to v0.4 (with real keys/users); the defense pipeline is fully proven over mock+browser — documented in README + release note
|
||||
- [ ] Optional future voice config noted for v0.4 (`AI_VOICE_BASE_URL` / `AI_VOICE_API_KEY`) in `.env.example` + README; keys only in gitignored `.ciagent/.env.secrets`; tests never call a voice API
|
||||
- [ ] Boundary rules hold: `voice/` imports no `agents/`/`api/`; `pnpm ai:test` + `pnpm ai:lint` green
|
||||
|
||||
---
|
||||
|
||||
## Phase 6: Agent Re-grounding + Learner Surface Integration
|
||||
|
||||
**Requirements:** REQ-3-007, REQ-3-008
|
||||
**Goal:** Lab/Assessor/Proctor consume real engine inputs (live telemetry, grading output, defense signals) with **no mock fallback in the learner path**; the v0.1 sandbox + assessment mockups become real — in-browser build/run (Run/Test buttons + read-only exec output, CUT-2 — no interactive shell), live telemetry panel, browser/typed voice defense, live grading; `pnpm build` + `pnpm typecheck` + full `pnpm ai:test` green
|
||||
**Note:** E2E verification runs against the real engines over HTTP with `AI_PROVIDER=mock` + `AI_VOICE_PROVIDER=mock` permitted (G-2 precedent) — the requirement is real engine plumbing (sandbox/telemetry/grading/defense over real endpoints, no corpus mocks in the learner path); a cloud outage must not block P6.
|
||||
|
||||
### Wave 1: Agent re-grounding (parallel — no shared files)
|
||||
|
||||
#### Task 6-1-01: Lab agent on live telemetry
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-3-007
|
||||
- **Files:** `apps/ai-service/ai_service/agents/lab.py` (update), `apps/ai-service/ai_service/prompts/lab.py` (update: render real trace digest), `apps/ai-service/tests/agents/test_lab_live.py` (new)
|
||||
- **Action:** Lab consumes a **live trace digest** (grading/features `compute_digest` over `TraceStore` events) instead of `corpus/telemetry.py`. build_messages renders digest facts (recent commands, failing tests, idle). Mock-provider scripts assert digest-derived content. v0.2 corpus path removed from the agent (dormant corpus retained until Task 6-1-04 check).
|
||||
- **Verify:** `pnpm ai:test` — test_lab_live green: feedback references events actually present in a seeded SQLite trace (not corpus fixtures); no `corpus.telemetry` import in `agents/lab.py` (AST-asserted)
|
||||
|
||||
#### Task 6-1-02: Assessor agent on grading output
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-3-007
|
||||
- **Files:** `apps/ai-service/ai_service/agents/assessor.py` (update), `apps/ai-service/ai_service/prompts/assessor.py` (update), `apps/ai-service/tests/agents/test_assessor_live.py` (new)
|
||||
- **Action:** Assessor consumes `GradeStore` output (validated `RubricScore` + digest) for learner+task instead of pre-baked artifacts/transcripts; renders strengths/gaps/verdict with rubric-anchored coaching framing. Structured output unchanged (D-020).
|
||||
- **Verify:** `pnpm ai:test` — test_assessor_live green: given a real stored grade, Assessor output reflects its scores; corpus artifact path gone from the agent (AST-asserted)
|
||||
|
||||
#### Task 6-1-03: Proctor agent on telemetry + defense signals
|
||||
- **Persona:** ai-engineer — **REQ:** REQ-3-007
|
||||
- **Files:** `apps/ai-service/ai_service/agents/proctor.py` (update), `apps/ai-service/ai_service/prompts/proctor.py` (update), `apps/ai-service/tests/agents/test_proctor_live.py` (new)
|
||||
- **Action:** Proctor consumes real integrity inputs: idle gaps + command cadence from the trace digest + defense integrity signals from `DefenseStore` → classified signals + coaching interventions (supportive tone retained). Cross-checks variant seed params (P4) for off-template work.
|
||||
- **Verify:** `pnpm ai:test` — test_proctor_live green: signals derived from seeded real trace + defense records; corpus proctor scenarios no longer imported (AST-asserted)
|
||||
|
||||
#### Task 6-1-04: Corpus dormancy + mockup removal verification
|
||||
- **Persona:** lead-developer — **REQ:** REQ-3-007
|
||||
- **Files:** `apps/ai-service/ai_service/corpus/telemetry.py`, `apps/ai-service/ai_service/corpus/artifacts.py`, `apps/ai-service/ai_service/corpus/` (README note), `apps/ai-service/tests/test_corpus_dormancy.py` (new)
|
||||
- **Action:** Verify no production code path imports `corpus/telemetry.py` or `corpus/artifacts.py` anymore (test scans imports across `agents/`, `api/`, engines). Retain files as Phase-3 calibration history with a header note marking them **dormant — v0.2 mocks, not used at runtime**; learner-context corpus stays (agents still need learner context). Disposes v0.2's G-5-class dead-code risk deliberately.
|
||||
- **Verify:** `pnpm ai:test` — test_corpus_dormancy green (zero runtime importers); suite otherwise unchanged
|
||||
|
||||
### Wave 2: Client plumbing + design primitives (parallel — no shared files)
|
||||
|
||||
#### Task 6-2-01: Sandbox build-panel engine client (run/test-only — no raw shell relay)
|
||||
- **Persona:** frontend-engineer — **REQ:** REQ-3-008
|
||||
- **Files:** `apps/web/hooks/use-sandbox-session.ts` (new), `apps/web/lib/engine-client.ts` (new), `apps/web/.env.example` (update)
|
||||
- **Action:** `engine-client.ts`: typed fetch client for `/v1/sandboxes` (create/destroy — learner_id from the v0.3 mock session constant, allowlisted server-side per G-5), `/v1/sandboxes/{id}/files` (workspace CRUD), `/v1/sandboxes/{id}/exec` (run/test), `/v1/variants`, `/v1/assessment/grade`, `/v1/telemetry/traces`, `/v1/defense/*`; base `NEXT_PUBLIC_AI_SERVICE_URL`. `use-sandbox-session.ts`: create sandbox+variant on task open → destroy on unmount (idempotent cleanup, AbortController pattern); 503 pool-full → user-facing "environment busy, retry" (D-032); 403/429 abuse-control surfaced honestly (G-5). **CUT-2 (G-8): NO raw interactive WS terminal relay (keystroke-level stdin/stdout) in v0.3** — the credential pipeline needs *process events* (from Run/Test + file edits), not a live shell; the interactive xterm relay is the most fragile real-time piece and is deferred to v0.4. The build panel is a **Run/Test output viewer** (exec results + telemetry pulse render), not an interactive shell. `@xterm/xterm` is therefore NOT a dependency in v0.3.
|
||||
- **Verify:** `pnpm install && pnpm typecheck` pass; hook unmount destroys the sandbox (manual probe: `curl localhost:8420/v1/sandboxes` shows count drop after navigation); RUN/TEST buttons produce streamed output + telemetry events in the trace
|
||||
|
||||
#### Task 6-2-02: New design primitives
|
||||
- **Persona:** design-system-engineer — **REQ:** REQ-3-008
|
||||
- **Files:** `packages/ui/src/primitives/terminal-frame.tsx` (new), `packages/ui/src/primitives/mic-control.tsx` (new), `packages/ui/src/primitives/grade-badge.tsx` (new), `packages/ui/src/primitives/telemetry-status.tsx` (new), `packages/ui/src/primitives/transcript-viewer.tsx` (new), `packages/ui/src/primitives/index.ts` (update), `packages/ui/src/index.ts` (update)
|
||||
- **Action:** Token-driven primitives: TerminalFrame (CUT-2: a read-only exec-output viewer chrome — streams Run/Test results, NOT an interactive shell), MicControl (record/stop with consent state + no-mic fallback state, MediaRecorder permission UX), GradeBadge (verdict rendering), TelemetryStatus (live event pulse / disconnected indicator), TranscriptViewer (examiner/learner turn list). Dark mode + WCAG AA; stories for each.
|
||||
- **Verify:** primitives import from `@nextcraft/ui`; Storybook stories render dark + light; `pnpm build` (ui package) passes
|
||||
|
||||
#### Task 6-2-03: Defense TS types
|
||||
- **Persona:** data-engineer — **REQ:** REQ-3-008
|
||||
- **Files:** `packages/types/defense.ts` (new), `packages/types/index.ts` (update)
|
||||
- **Action:** TS `DefenseSession`, `DefenseTurn`, `IntegritySignal`, `Verdict` mirroring P5 Python models (cross-referencing header).
|
||||
- **Verify:** `pnpm typecheck` passes
|
||||
|
||||
### Wave 3: Real build surface (depends on Wave 2)
|
||||
|
||||
#### Task 6-3-01: Sandbox mockup → real in-browser IDE
|
||||
- **Persona:** frontend-engineer — **REQ:** REQ-3-008
|
||||
- **Files:** `apps/web/app/(learner)/build/[competencyId]/page.tsx` (rewrite), `apps/web/components/learner/sandbox-terminal.tsx` (new), `apps/web/components/learner/file-tree.tsx` (new), `apps/web/components/learner/run-controls.tsx` (new), `apps/web/components/learner/lab-feedback-panel.tsx` (update: live trace)
|
||||
- **Action:** Replace the mockup with the real build environment (A-103, CUT-2): file tree (HTTP CRUD into the sandbox workdir via `/v1/sandboxes/{id}/files` routes added to api/sandboxes — read/write/list workspace files), syntax-highlight editor (existing), **Run**/**Test** buttons (exec in sandbox; results stream to a read-only TerminalFrame output panel — no interactive shell), starter files from the P4 variant scaffold. Lab panel posts `learner_id+task_id` → streams Lab feedback over the **live** trace (no scenario IDs). Telemetry sidebar shows live TelemetryStatus. Pool-full 503 → busy state with retry; 403/429 surfaced.
|
||||
- **Verify:** with ai-service up: open `/build/comp-01` → variant statement + starter files load → edit a file → **Run** executes the command in-sandbox and output renders in the panel → **Test** runs the test suite in-sandbox → Lab panel streams digest-derived feedback → telemetry status shows live events. Manual probe documented; `pnpm typecheck` green
|
||||
|
||||
### Wave 4: Live defense + grading surfaces (depends on Waves 2-3)
|
||||
|
||||
#### Task 6-4-01: Assessment mockup → live defense + live grading
|
||||
- **Persona:** frontend-engineer — **REQ:** REQ-3-008
|
||||
- **Files:** `apps/web/app/(learner)/defend/[competencyId]/page.tsx` (rewrite), `apps/web/components/learner/defense-session.tsx` (new), `apps/web/components/learner/assessor-results-panel.tsx` (update), `apps/web/components/learner/proctor-banner.tsx` (update), `apps/web/components/learner/oral-defense-interface.tsx` (rewrite or remove)
|
||||
- **Action:** Real assessment flow: **Start Defense** → POST `/v1/defense/start` → examiner question streams → learner answers via MicControl (MediaRecorder webm/opus → multipart POST) with typed fallback when mic denied or `provider=browser` (native `SpeechRecognition`/`speechSynthesis` path per fallback descriptor) → follow-ups stream → **Finish** → verdict + integrity signals panel (TranscriptViewer, GradeBadge) + **Grade My Work** → POST `/v1/assessment/grade` → structured rubric bars render. Proctor banner reads signals for task_id. Loading + error states throughout; no mock defense data remains in the learner path.
|
||||
- **Verify:** with ai-service up: full defense loop runs in-browser (typed fallback acceptable in CI-less manual probe; mic path exercised manually with permission granted); grade panel renders real rubric scores; `pnpm typecheck` green
|
||||
|
||||
### Wave 5: End-to-end verification (depends on Waves 3-4)
|
||||
|
||||
#### Task 6-5-01: Full learner-path E2E probe + green builds
|
||||
- **Persona:** backend-engineer — **REQ:** REQ-3-007, REQ-3-008
|
||||
- **Files:** `apps/ai-service/tests/api/test_e2e_credential_flow.py` (new), `apps/ai-service/README.md` (update: E2E probe doc)
|
||||
- **Action:** Endpoint-level E2E (mock LLM/voice providers, real engines): create variant → create sandbox with task_id → ingest trace events via real sandbox exec → POST grade → start defense (typed answers) → finish → assert: trace persisted, grade stored + digest-linked, defense transcript + signals stored, Proctor/Assessor endpoints serve them. No corpus fixture anywhere in the flow (AST-asserted). Run `pnpm build` + `pnpm typecheck` + full `pnpm ai:test` at repo root; fix all failures before phase ship.
|
||||
- **Verify:** `pnpm ai:test` green incl. test_e2e_credential_flow; `pnpm build` + `pnpm typecheck` green; README E2E probe section documents the manual browser pass
|
||||
|
||||
### Must-Haves (Phase 6)
|
||||
- [ ] Lab/Assessor/Proctor operate on real inputs with **no mock fallback in the learner path** (AST-verified: no corpus telemetry/artifact/proctor imports in production paths)
|
||||
- [ ] Learner builds in-browser for real (CUT-2): Run/Test buttons execute in a namespace sandbox and stream output to a read-only panel; file tree CRUD works; starter files come from the learner's variant scaffold
|
||||
- [ ] Live telemetry: build activity streams to ai-service and the sidebar shows live status (TelemetryStatus); trace persisted in SQLite
|
||||
- [ ] Live defense: start → answer (mic or typed fallback) → examiner follow-ups → finish → transcript + integrity signals + verdict rendered; browser-native path works with no server voice key
|
||||
- [ ] Live grading: grade request returns structured rubric scores computed from the real trace digest; grade panel renders them
|
||||
- [ ] 503 pool-full surfaced honestly in UI; navigating away destroys the sandbox (no leaked sandboxes — `GET /v1/sandboxes` manual probe)
|
||||
- [ ] Examiner remains protocol-clean (voice only via `VoiceProvider`); module boundary rules hold across all new code
|
||||
- [ ] `pnpm build` and `pnpm typecheck` pass; full `pnpm ai:test` green; no cloud/voice calls in any automated test
|
||||
- [ ] Release-note input (for P7): v0.3 ships real engines; **identity/age-gating (KYC) remains deferred — age-gating is still a visual mockup** (A-110); sandbox scope is coding-IDE only (design tool/simulation deferred to v0.4, D-025)
|
||||
|
||||
---
|
||||
|
||||
## Phase 7: Final Review + Ship (no planned tasks)
|
||||
|
||||
Orchestrated by the SHIP stage, not this plan: multi-persona code review (correctness, testing, module boundaries, secrets hygiene — keys absent from code/logs/commits/errors, localhost-only CORS, no PII in prompts), project health audit (reconstruction, .ciagent/ discipline, branch/commit hygiene), then merge milestone → main, tag the final v0.2.x patch, create the Gitea release, mark all 8 v0.3 requirements complete.
|
||||
|
||||
**Release-note honesty:** the release note must state (a) Lab/Assessor/Proctor now run on real engine inputs (v0.2 mock-input caveat retired), (b) sandbox scope = coding IDE only — design tool + simulation environments deferred to v0.4 (D-025), (c) identity/age-gating (KYC) deferred per founder directive — age-gating remains the v0.1 visual flow mockup; abuse control (per-learner sandbox caps + server-side learner allowlist, G-5) ships in place of auth (A-110), (d) voice runs mock-first with browser-native fallback — real server STT/TTS deferred to v0.4 (CUT-1), (e) sandbox resource limits are partially enforced (memory/CPU/wall-clock kernel-enforced via rlimits; per-sandbox pids + hard disk quota are NOT — mitigated by a workdir-size sweep and per-learner caps; full enforcement requires cgroup delegation, deferred to the post-MVP containerd backend, D-024/G-1/G-2).
|
||||
|
||||
**Secrets-hygiene checklist (P7):** `.ciagent/.env.secrets` gitignored and never committed; `AI_VOICE_API_KEY` / `AI_TUTOR_API_KEY` referenced only via env; SQLite DBs + sandbox dirs gitignored; no keys in logs, error messages, or test fixtures.
|
||||
|
||||
**Disposal checks (G-5 class):** v0.2 dormant corpus files carry the dormant-header note (Task 6-1-04); any now-unused mock-data exports for the old sandbox/defense mockups (e.g. `aiTutorResponses`-class leftovers) must be removed or deprecated by review.
|
||||
|
||||
---
|
||||
| 1 | Bootstrap CLI core | REQ-4-001, REQ-4-002 | 3 | cli-engineer, backend-engineer, security-auditor (W3 review) |
|
||||
| 2 | Binary build + release pipeline | REQ-4-003, REQ-4-004 | 3 | cli-engineer, backend-engineer, security-auditor |
|
||||
| 3 | Install docs + fresh-clone E2E | REQ-4-005 | 2 | cli-engineer, backend-engineer |
|
||||
| 4 | Final review + ship | — | 1 | all reviewers |
|
||||
|
||||
## User-Facing Surface
|
||||
|
||||
The primary user-facing surface is the **learner build + defend flow** at `http://localhost:3000`, backed by the real engines in ai-service at `http://localhost:8420`:
|
||||
|
||||
- `/dashboard` — AI tutor chat (Coach/Tutor, streaming) + Mentor panel (unchanged from v0.2)
|
||||
- `/learn/[competencyId]` — byte viewer with streaming Tutor explanations (unchanged)
|
||||
- `/build/[competencyId]` — **real in-browser build environment**: file tree, syntax editor, **Run/Test buttons that execute in a namespace sandbox and stream output to a read-only panel** (CUT-2 — no interactive shell), live telemetry status, live Lab feedback, per-learner variant task statement
|
||||
- `/defend/[competencyId]` — **live oral defense + live grading**: Examiner voice/typed dialogue, transcript + integrity signals, real rubric scores from the process trace
|
||||
|
||||
The marketplace, employer, and admin surfaces are unchanged from v0.1/v0.2.
|
||||
1. **One-liner install (README quickstart):** `curl -fsSL https://git.coreci.dev/coreci/nextcraft/raw/main/scripts/install.sh | bash` — downloads the latest release's `nextcraft-linux-x64` binary, verifies its sha256, installs to `~/.local/bin`, prints a PATH hint if needed.
|
||||
2. **CLI commands:** `nextcraft doctor` (prereq checks), `nextcraft bootstrap` (fresh clone → runnable stack), `nextcraft verify` (health check), `nextcraft dev` (dev server passthrough), plus `--help`/`--version`.
|
||||
3. **Release surface:** every Gitea release from v0.3.2 onward carries `nextcraft-linux-x64` + `nextcraft-linux-x64.sha256` assets.
|
||||
|
||||
## Happy Path
|
||||
|
||||
1. Learner opens `/build/comp-01` → a per-learner **variant statement** and starter files load; a namespace sandbox is created for the session
|
||||
2. Learner edits files in the tree and clicks **Run**/**Test** → commands execute in the sandbox and real output renders in the build panel; the telemetry sidebar pulses as events stream to ai-service and persist in SQLite
|
||||
3. The Lab panel streams feedback derived from the **live trace digest** (real commands, real failures)
|
||||
4. Learner opens `/defend/comp-01` → **Start Defense**: the Examiner streams an opening question ("Walk me through your build — why did you structure it this way?")
|
||||
5. Learner answers by voice (mic consent → MediaRecorder → STT) or typed fallback → examiner follow-ups probe the trace ("You hit three test failures before passing — what changed?"); TTS plays examiner audio (or browser speech in fallback)
|
||||
6. Learner finishes the defense → transcript + integrity signals appear; verdict renders in a GradeBadge
|
||||
7. Learner clicks **Grade My Work** → the grading engine computes the digest from the real trace, rubric-scores it, and the panel renders per-criterion bars + strengths/gaps/verdict
|
||||
8. Proctor banner shows integrity signals from the live trace + defense in coaching tone; Mentor panel on `/dashboard` can narrate the real outcome
|
||||
9. Navigating away destroys the sandbox (pool slot freed); killing ai-service shows inline error + retry states on every panel, with no crashes
|
||||
Before execution, the end-to-end scenario this milestone must make true:
|
||||
|
||||
1. A consumer on a linux x64 box runs the one-liner; `nextcraft` lands in `~/.local/bin`.
|
||||
2. They clone the repo (or the CLI detects the repo root), run `nextcraft doctor` — all prerequisites report ✓ with actionable messages for any gap.
|
||||
3. `nextcraft bootstrap` — pnpm install, ai-service venv via the existing bootstrap.sh, `.env` created from `.env.example`, optional-key warnings (not blockers), mock providers keep the stack runnable keyless.
|
||||
4. `nextcraft verify` — venv imports, ports, env presence, build readiness all ✓.
|
||||
5. `nextcraft dev` — the dev stack runs; Ctrl+C stops it (passthrough semantics).
|
||||
6. On every ship, the Gitea release shows the binary + checksum assets; re-running the one-liner upgrades to the latest binary.
|
||||
|
||||
## UX Acceptance Criteria
|
||||
|
||||
1. Run/Test output visibly reflects the real sandbox execution (command round-trip to the sandbox, real stdout/stderr), not a replay animation
|
||||
2. The learner path contains **no mock engine data** — scenarios, canned artifacts, and scripted defense transcripts from v0.2 are gone from runtime
|
||||
3. Variant statements visibly differ between two learner sessions on the same competency
|
||||
4. Mic permission flow is graceful: consent prompt, recording indicator, no-mic/typed fallback, and browser-native speech path when no server voice key is configured
|
||||
5. Defense transcript renders turn-by-turn with latency shown; integrity signals render in coaching (supportive) tone
|
||||
6. Grade results render as structured per-criterion bars with verdict, from the real trace — not from final-output-only heuristics
|
||||
7. Pool-full (503) shows an honest "environment busy — retry" state; navigation/unmount destroys sandboxes with no leaks
|
||||
8. When ai-service is unreachable: inline error + retry on every panel — no crashes, no console errors, no blank UI
|
||||
9. All new UI uses design tokens, supports dark mode, meets WCAG AA contrast, responsive at 375px / 768px / 1280px
|
||||
10. `pnpm build` and `pnpm typecheck` pass with zero errors; `pnpm ai:test` green cloud-free and voice-key-free
|
||||
- `doctor` output lists every prerequisite with ✓/✗ and a **fix hint** on every ✗; exit code 1 if any ✗, 0 otherwise.
|
||||
- `bootstrap` is **idempotent** — running twice produces the same end state, second run fast (no reinstalls where avoidable).
|
||||
- `bootstrap` never writes secrets, never blocks on missing optional keys — warns with the exact key names and where to set them.
|
||||
- `verify` gives a single-glance green/red summary; every red item names the failing command it ran.
|
||||
- Every command supports `--help`; unknown command/flag exits 2 with usage.
|
||||
- The one-liner **never hard-fails silently**: any error path (no release, no binary asset, checksum mismatch, platform mismatch) prints a specific message + the source-bootstrap alternative.
|
||||
- Checksum mismatch = hard stop + explicit "do not run this binary" message.
|
||||
- PATH hint: if `~/.local/bin` is not on PATH, the installer prints the exact export line to add.
|
||||
- Binary runs standalone on a box with node NOT installed (SEA self-containment) — `./nextcraft-linux-x64 --version` works.
|
||||
|
||||
---
|
||||
|
||||
## Phase 1: Bootstrap CLI Core
|
||||
|
||||
**Requirements:** REQ-4-001, REQ-4-002
|
||||
**Goal:** `apps/cli` package with doctor/bootstrap/verify/dev fully working from source (`node dist` + pnpm bin), unit-tested, wired into the monorepo (turbo + root scripts), composing — not duplicating — the existing scripts.
|
||||
|
||||
### Wave 1: Package foundation (parallel)
|
||||
|
||||
#### Task 1-1-01: CLI package scaffold + entry + dispatch
|
||||
- **Persona:** cli-engineer — **REQ:** REQ-4-001
|
||||
- **Files:** `apps/cli/package.json`, `apps/cli/tsconfig.json`, `apps/cli/src/index.ts`, `apps/cli/src/commands/help.ts` (usage text), `apps/cli/tests/dispatch.test.ts`
|
||||
- **Action:** pnpm workspace package `@nextcraft/cli` (private, `"bin": {"nextcraft": "dist/index.js"}`). Entry: parse argv (hand-rolled, no runtime deps), dispatch to commands, `--help`/`-h`, `--version` (from package.json version), unknown → exit 2 with usage. Exit-code contract: 0 ok / 1 failure / 2 usage. shebang `#!/usr/bin/env node` on the built entry (esbuild banner in P2; for P1 `tsx` runs in dev via package script `"dev": "tsx src/index.ts"`).
|
||||
- **Verify:** `pnpm --filter @nextcraft/cli test` green (dispatch: routes doctor/bootstrap/verify/dev; unknown exits 2; --help exits 0; --version prints package version); `pnpm typecheck` green.
|
||||
|
||||
#### Task 1-1-02: Checks library (pure logic)
|
||||
- **Persona:** cli-engineer — **REQ:** REQ-4-001, REQ-4-002
|
||||
- **Files:** `apps/cli/src/checks/check-command.ts`, `apps/cli/src/checks/check-env.ts`, `apps/cli/src/lib/log.ts`, `apps/cli/tests/checks.test.ts`
|
||||
- **Action:** `check-command`: given a name + optional `--version` probe + a min-version parser, resolve binary on PATH (`which`), semver-ish compare (major.minor tolerant), return `CheckResult {name, ok, found, version, hint}`. `check-env`: diff `.env.example` template keys vs an existing `.env` (missing keys → warn-classified; required-vs-optional classification table from the template's own comments + a static required list of zero keys — all optional per A-210), return per-key results. `log.ts`: `ok(msg)`, `fail(msg, hint)`, `warn(msg)`, `info(msg)` formatters with symbols and consistent alignment. Pure functions — no side effects at import; fs access injected as parameters for testability.
|
||||
- **Verify:** unit tests green: version compare (>= boundaries), missing binary → ok:false + hint, env diff missing/new/extra keys, required-optional classification.
|
||||
|
||||
#### Task 1-1-03: Root + turbo wiring
|
||||
- **Persona:** backend-engineer — **REQ:** REQ-4-002
|
||||
- **Files:** root `package.json` (update), `turbo.json` (update), `pnpm-workspace.yaml` (verify apps/* already covered — no change expected)
|
||||
- **Action:** Add `cli:dev`, `cli:test`, `cli:build`, `cli:typecheck`, `cli:lint` root scripts mirroring the `ai:*` passthrough pattern (D-022/D-037). Turbo tasks for the CLI package: `build` (dependsOn `^build`, outputs `dist/**`), `test`, `typecheck`, `lint` (cache:false, outputs:[] for test — same shape as ai-service). No changes to existing ai:* tasks.
|
||||
- **Verify:** `pnpm cli:test` + `pnpm cli:typecheck` green from repo root; `pnpm build` still green for web+ai-service (turbo graph unaffected); `pnpm ai:test` still green.
|
||||
|
||||
### Wave 2: Commands (depends on Wave 1)
|
||||
|
||||
#### Task 1-2-01: doctor command
|
||||
- **Persona:** cli-engineer — **REQ:** REQ-4-001
|
||||
- **Files:** `apps/cli/src/commands/doctor.ts`, `apps/cli/tests/doctor.test.ts`
|
||||
- **Action:** Checks (each with actionable hint): node ≥18 (`process.version`), pnpm ≥8 on PATH (`pnpm --version`), python3 ≥3.11 (`python3 --version` parse), git (`git --version`), corepack available-or-pnpm-present nuance folded into pnpm check, `unshare` binary on PATH (`which unshare` — sandbox fabric needs it; hint explains what breaks without it). Sequential execution with per-check timeout; summary line; exit 1 if any ✗. Runs from any cwd (no repo required — pure environment check).
|
||||
- **Verify:** unit tests with injected spawn results: all-pass → exit 0 + summary; missing pnpm → ✗ + hint + exit 1; missing unshare → ✗ with sandbox-specific hint.
|
||||
|
||||
#### Task 1-2-02: bootstrap command
|
||||
- **Persona:** cli-engineer — **REQ:** REQ-4-002
|
||||
- **Files:** `apps/cli/src/commands/bootstrap.ts`, `apps/cli/src/lib/spawn.ts`, `apps/cli/tests/bootstrap.test.ts`
|
||||
- **Action:** `spawn.ts`: `run(cmd, args, {timeoutMs, cwd, env})` — promisified child_process.spawn, inherited stdio, timeout kill (SIGTERM→SIGKILL escalation), returns `{code}`; throws never (codes always returned). `bootstrap.ts` steps (each logged before/after): (1) locate repo root (walk up for pnpm-workspace.yaml; error with hint if not in a clone); (2) `pnpm install` at root; (3) delegate ai-service venv to `apps/ai-service/scripts/bootstrap.sh` via spawn with generous timeout (10 min) — **zero pip/venv logic in the CLI** (A-202); (4) copy `.env.example` → `.env` if absent (preserve existing; report created vs kept); (5) validate optional keys in `.env` vs template — warn-only (A-210); never touch `.ciagent/.env.secrets`; (6) print next-steps (`nextcraft verify`, `nextcraft dev`). Idempotent: every step safe to re-run.
|
||||
- **Verify:** unit tests with stub spawn: step order, env copy semantics (absent → create, present → keep), timeout path returns failure code, secrets file never written; `pnpm cli:test` green.
|
||||
|
||||
#### Task 1-2-03: verify + dev commands
|
||||
- **Persona:** cli-engineer — **REQ:** REQ-4-002
|
||||
- **Files:** `apps/cli/src/commands/verify.ts`, `apps/cli/src/commands/dev.ts`, `apps/cli/tests/verify.test.ts`
|
||||
- **Action:** `verify.ts` health checks (each runnable + reported): ai-service venv python imports (`import ai_service` via venv python), uvicorn present in venv, ports 3000/8420 free (net stat via node), `.env` exists with AI_PORT parseable, `pnpm build` dry readiness (turbo graph parses — run `turbo build --dry=json` cheap check or typecheck-only default; choose the cheap one). Summary + exit code. `dev.ts`: locate repo root, exec passthrough to `apps/ai-service/scripts/dev.sh` with **inherited stdio and signals** (Ctrl+C semantics), no timeout (long-running); document that web dev server runs via `pnpm dev` separately (dev.sh owns ai-service only).
|
||||
- **Verify:** unit tests: verify aggregates check results → exit codes; dev spawns dev.sh with signal passthrough assertions (mock spawn).
|
||||
|
||||
### Wave 3: Integration review (depends on Wave 2)
|
||||
|
||||
#### Task 1-3-01: CLI security + integration review pass
|
||||
- **Persona:** security-auditor — **REQ:** REQ-4-001, REQ-4-002
|
||||
- **Files:** `apps/cli/src/lib/spawn.ts` (review; patch if defect), `apps/cli/src/commands/bootstrap.ts` (review), `apps/cli/tests/**` (add regression if defect found)
|
||||
- **Action:** STRIDE pass on the CLI surface: spawn injection (args never through shell string — array form only), timeout enforcement, secrets never logged, env template copy doesn't overwrite user edits, no shell=true anywhere, PATH resolution honest errors. Findings → P0 patches now with regression tests; P1+ noted for final-phase review.
|
||||
- **Verify:** `pnpm cli:test` green incl. any added regressions; `grep -rn "shell: *true" apps/cli/src` returns nothing.
|
||||
|
||||
### Must-Haves (Phase 1)
|
||||
- [ ] `pnpm --filter @nextcraft/cli test` green; `pnpm typecheck` green; `pnpm build` green
|
||||
- [ ] doctor: every prerequisite reported with ✓/✗ + actionable hint; exit 1 on any ✗; runs outside a repo clone
|
||||
- [ ] bootstrap: composes scripts/bootstrap.sh (no pip/venv logic in CLI); idempotent; .env created from template only when absent; optional-key warnings, never blocks; never writes secrets
|
||||
- [ ] verify: venv import + uvicorn + ports + env checks with single-glance summary and named failing commands
|
||||
- [ ] dev: passthrough with signal inheritance (Ctrl+C stops the stack)
|
||||
- [ ] Exit-code contract: 0/1/2; --help everywhere; unknown command → 2
|
||||
- [ ] No runtime npm dependencies in apps/cli (dev deps only)
|
||||
- [ ] Root `cli:*` scripts work from repo root; ai:* scripts unaffected
|
||||
|
||||
---
|
||||
|
||||
## Phase 2: Binary Build + Release Pipeline
|
||||
|
||||
**Requirements:** REQ-4-003, REQ-4-004
|
||||
**Goal:** `nextcraft-linux-x64` SEA binary + sha256 sidecar built reproducibly from the CLI package; one-liner `install.sh` verified end-to-end against a real release; release-asset upload wired so **every ship from now on carries binaries**.
|
||||
|
||||
### Wave 1: Binary build (parallel)
|
||||
|
||||
#### Task 2-1-01: SEA binary build script (G-101: live-build probe FIRST — mechanism must be proven before the pipeline depends on it)
|
||||
- **Persona:** cli-engineer — **REQ:** REQ-4-004
|
||||
- **Files:** `apps/cli/scripts/build-binary.mjs`, `apps/cli/package.json` (add `build:binary` script), `apps/cli/.sea-config.json` (or generated in-script)
|
||||
- **Action:** **First action of this task: build one real SEA binary end-to-end and run it** (`--version` + `doctor` smoke) before writing the polished script.** Pipeline: esbuild bundle `src/index.ts` → `dist/bundle.cjs` (platform node, target node18, banner shebang, SEA config: `{main: "dist/bundle.cjs", output: "dist/sea-prep.blob", disableExperimentalSEAWarning: true}`) → `node --experimental-sea-config` → copy system node binary → inject blob (`npx postject` with sentinel `NODE_SEA_BLOB_FUSE` fuse, or `dd` fallback) → chmod +x → `dist/nextcraft-linux-x64` → **stamp version from the shipping tag argument** (`NEXTCRAFT_VERSION` injected via esbuild `define`, G-102 — `--version` prints it; absent arg → dev stamp `0.0.0-dev`) → `shasum -a 256` → `dist/nextcraft-linux-x64.sha256`. Fallback (documented, scripted, honest): if SEA injection fails, python3 zipapp builds `nextcraft-linux-x64.pyz` (requires python3 on target — install.sh handles both asset shapes and the docs say so; NO silent claim of node-less operation, G-101).
|
||||
- **Verify:** `pnpm --filter @nextcraft/cli build:binary` produces the binary; `./dist/nextcraft-linux-x64 --version` runs **with node absent from PATH** (test via `env -i /bin/sh -c 'PATH=/usr/bin:/bin ...'` sandbox or by temporarily stripping PATH in a subprocess test); sha256 file matches `shasum -c`.
|
||||
|
||||
#### Task 2-1-02: Release-asset upload helper
|
||||
- **Persona:** cli-engineer — **REQ:** REQ-4-004
|
||||
- **Files:** `scripts/release-assets.sh`, `apps/cli/tests/release-assets.test.ts` (fixture-level)
|
||||
- **Action:** Given a tag: build binary (Task 2-1-01), resolve GITEA_TOKEN from `.env`/`.env.secrets`/`.env.*` **via the secrets loader only** (never shell env — v1.8 root cause), create/locate the Gitea release via API, upload both assets (`POST /api/v1/repos/{owner}/{repo}/releases/{id}/assets?name=...` multipart). Bounded retry (3) per config.ship.max_release_retries; token never echoed; failure = non-blocking escalation message (release_pending semantics) — tag+merge already complete the ship.
|
||||
- **Verify:** fixture test: token resolution order (.env.secrets wins over .env; shell env NEVER consulted — assert with a poisoned env var fixture); dry-run mode prints the exact curl-multipart it would send (no net in tests).
|
||||
|
||||
#### Task 2-1-03: install.sh one-liner
|
||||
- **Persona:** cli-engineer — **REQ:** REQ-4-003
|
||||
- **Files:** `scripts/install.sh`, `apps/cli/tests/install-script.test.ts`
|
||||
- **Action:** POSIX sh (no bashisms — dash-safe): `set -eu`; platform check (uname linux + x86_64; else print source-bootstrap path + exit 0 — a graceful no-op, not an error); resolve latest release via Gitea API (`curl -fsSL .../releases/latest`, parse `tag_name` + asset `browser_download_url`s with sed/grep — no jq dependency); **match assets by EXACT name** (`nextcraft-linux-x64`, `nextcraft-linux-x64.sha256` — any parse/lookup miss = degrade to source-bootstrap instructions, exit 0, G-103 — never a name-approximate install); handle the zipapp asset shape (`nextcraft-linux-x64.pyz` + sidecar) when the binary is absent, printing the python3 requirement honestly; download both assets to `mktemp -d` (trap cleanup EXIT); **verify sha256 before anything else** (`shasum -a 256 -c` or sha256sum); on mismatch → hard stop, explicit "do not run" message, exit 1; install to `~/.local/bin` (mkdir -p; `--dest` override); PATH hint when missing (print exact export line); print the binary's own `--version` output (G-102: must equal the resolved release tag — mismatch = install-time integrity stop) + `nextcraft doctor` next-step. No-binary-asset path: print the git-clone + scripts/bootstrap.sh instructions + exit 0. Zero secrets required (public release assets).
|
||||
- **Verify:** unit tests over the script's pure helpers extracted where feasible; **live E2E in Task 2-3-01**. `sh -n scripts/install.sh` syntax-clean; `dash scripts/install.sh --help` safe if dash present.
|
||||
|
||||
### Wave 2: Ship-flow integration (depends on Wave 1)
|
||||
|
||||
#### Task 2-2-01: Wire binaries into every ship (G-104: enforcement, not prose)
|
||||
- **Persona:** backend-engineer — **REQ:** REQ-4-004
|
||||
- **Files:** `.ciagent/config.json` (no schema change needed — release section already configured), this repo's ship procedure notes (update `.ciagent/ARCHITECTURE.md` Build Order note if needed), `scripts/release-assets.sh` (finalize from 2-1-02)
|
||||
- **Action:** Establish the ship-time contract going forward: after every phase ship (tag + merge complete = ship gate per config.ship), run `scripts/release-assets.sh <tag>` to attach binary + checksum to the freshly created release. **G-104:** this run is MANDATORY-ATTEMPTED on every release from v0.3.2 onward — best-effort/non-blocking like release creation (release_pending escalation on exhaustion), logged in the ship commit, and the P4 final audit gate includes "milestone release carries both assets" as an explicit check. This makes "ongoing binaries" a property of the pipeline, not a one-off.
|
||||
- **Verify:** The P2 ship itself executes the step against tag v0.3.2 (live validation — see Ship).
|
||||
|
||||
### Wave 3: End-to-end validation (depends on Wave 2)
|
||||
|
||||
#### Task 2-3-01: Install E2E against the live release
|
||||
- **Persona:** security-auditor — **REQ:** REQ-4-003
|
||||
- **Files:** `apps/cli/tests/install-e2e.test.ts` (marked slow/e2e), `apps/cli/README.md` (install internals section)
|
||||
- **Action:** Live E2E after the v0.3.2 release exists (run post-ship, documented as the verify gate for this phase's asset path): fresh HOME tmpdir → run install.sh → assert binary at `$HOME/.local/bin/nextcraft`, `--version` output equals the release tag (G-102 integrity assertion), checksum verified path taken (tamper test: flip a byte in a local fixture download → script refuses + exits 1). Record the transcript in the phase verify commit. If the live release isn't reachable at verify time, run the full local equivalent (serve assets from a fixture dir via `python3 -m http.server` + FORGE_BASE override) and mark live re-check as a P1 follow-up.
|
||||
- **Verify:** E2E green locally (fixture server path mandatory in tests — no test depends on the live forge); tamper-rejection proven; transcript recorded.
|
||||
|
||||
### Must-Haves (Phase 2)
|
||||
- [ ] **G-101:** a real SEA binary built + smoke-run BEFORE the pipeline depends on it; if SEA fails, zipapp is primary and docs state the python3 requirement
|
||||
- [ ] **G-102:** binary `--version` reports the shipping tag (stamped at build); install E2E asserts version == release tag
|
||||
- [ ] `pnpm --filter @nextcraft/cli build:binary` produces `nextcraft-linux-x64` + `.sha256`; binary runs without node on PATH (`--version`, `doctor` smoke)
|
||||
- [ ] **G-103:** `sh -n scripts/install.sh` clean; dash-safe; exact-name asset matching; platform mismatch → graceful source-bootstrap path (exit 0)
|
||||
- [ ] Checksum verified before install; tamper → hard stop with explicit warning (E2E-proven)
|
||||
- [ ] install.sh resolves latest release + assets from the Gitea API with zero secrets and no jq
|
||||
- [ ] release-assets.sh resolves GITEA_TOKEN from .env* files only (never shell env — tested with poisoned env)
|
||||
- [ ] **G-104:** v0.3.2 release carries both assets (live validation at ship); upload failure is non-blocking escalation, attempted + logged every release
|
||||
- [ ] `pnpm build`, `pnpm typecheck`, `pnpm cli:test` all green
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: Install Docs + Fresh-Clone E2E
|
||||
|
||||
**Requirements:** REQ-4-005
|
||||
**Goal:** README quickstart + CLI reference matching the tested reality exactly, plus a fresh-clone E2E test proving the happy path end-to-end.
|
||||
|
||||
### Wave 1: Fresh-clone E2E (drives doc accuracy)
|
||||
|
||||
#### Task 3-1-01: Fresh-clone bootstrap E2E
|
||||
- **Persona:** cli-engineer — **REQ:** REQ-4-005
|
||||
- **Files:** `apps/cli/tests/fresh-clone-e2e.test.ts` (slow/e2e-marked)
|
||||
- **Action:** In a `mktemp -d` sandbox: `git clone` the repo locally (file:// clone of HEAD — no network), run `pnpm --filter @nextcraft/cli dev -- doctor` (or the built binary from P2) → then `bootstrap` → then `verify`, asserting each step's exit codes and key output markers. Skips gracefully when network-dependent steps are unavailable (CI marker). Documents the exact happy path the README will state.
|
||||
- **Verify:** E2E green locally (clone of the working tree); output transcript matches README claims (cross-checked in 3-2-01).
|
||||
|
||||
### Wave 2: Documentation (depends on Wave 1 transcript)
|
||||
|
||||
#### Task 3-2-01: README quickstart + CLI reference
|
||||
- **Persona:** backend-engineer — **REQ:** REQ-4-005
|
||||
- **Files:** root `README.md` (update quickstart section), `apps/cli/README.md` (CLI reference)
|
||||
- **Action:** Root README quickstart: the one-liner (exact tested URL), then doctor → bootstrap → verify → dev sequence with expected outputs; source-bootstrap alternative documented (clone + scripts). apps/cli README: every command, flags, exit codes, the env-template copy semantics, optional-key warning semantics, secrets policy (never generated/committed; .ciagent/.env.secrets location), binary install internals, troubleshooting table keyed to actual failure modes observed in E2E.
|
||||
- **Verify:** Every command line in both READMEs is copy-paste runnable — verified against the 3-1-01 transcript; doc drift check: no references to commands/flags that don't exist in `--help` output.
|
||||
|
||||
### Must-Haves (Phase 3)
|
||||
- [ ] Fresh-clone E2E green: doctor → bootstrap → verify sequence from a clean clone
|
||||
- [ ] README quickstart matches the E2E transcript exactly (no aspirational docs)
|
||||
- [ ] CLI reference covers all 4 commands + --help/--version + exit codes
|
||||
- [ ] `pnpm build`, `pnpm typecheck`, `pnpm test` (all suites) green
|
||||
|
||||
---
|
||||
|
||||
## Phase 4: Final Review + Ship (milestone release v0.3.4)
|
||||
|
||||
1. Branch gate → `phase/04-final-review-ship`.
|
||||
2. Multi-persona review across the milestone (correctness, testing, security, performance, maintainability, adversarial) — P0 auto-fixed, P1+ fixed in this phase.
|
||||
3. Audit: reconstruction test (.ciagent files ↔ git log), file discipline, branch hygiene, commit discipline, P0-review flags resolved, **G-104 gate: milestone release v0.3.4 carries `nextcraft-linux-x64` + `.sha256` assets**.
|
||||
4. Milestone ship: merge phase/04 → milestone/v0.4-distribution; merge milestone → main; tag **v0.3.4** (= milestone release); attach binary + checksum assets (the ongoing-binaries contract); release notes with full milestone summary (all phases, all REQ-4-001..005, the "ongoing binaries from now on" statement, v0.5 deferral list per D-016); delete all milestone/phase branches.
|
||||
5. Complete: REQUIREMENTS.md REQ-4-001..005 → complete; ROADMAP.md v0.4 → complete; checkpoint cleared.
|
||||
|
||||
## Must-Haves (Milestone)
|
||||
- [ ] One-liner installs a working binary from the live Gitea release (E2E-proven, tamper-tested)
|
||||
- [ ] Fresh clone → doctor → bootstrap → verify → dev: the full happy path green from a clean environment
|
||||
- [ ] Every release from v0.3.2 onward carries `nextcraft-linux-x64` + `.sha256` assets
|
||||
- [ ] Zero runtime npm deps in the CLI; secrets only ever from .env* files; never in code/logs/commits
|
||||
- [ ] All suites green: `pnpm build`, `pnpm typecheck`, `pnpm ai:test`, `pnpm cli:test`
|
||||
+37
-17
@@ -8,45 +8,64 @@ Nextcraft is an AI-native outcome school where graduates prove what they can bui
|
||||
|
||||
---
|
||||
|
||||
## Current Milestone: v0.3 — Credential Engines
|
||||
## Current Milestone: v0.5 — Real Server Voice, KYC/Identity, Sandbox Environments, Seq-Lease
|
||||
|
||||
**Scope:** Replace v0.2's mock engine inputs with real credential engines. Build the sandbox fabric (sandboxed IDE / design tool / simulation), the live in-environment build-telemetry pipeline, the process-trace grading engine, per-learner variant task generation, and the oral/voice defense with AI examiner. Lab/Assessor/Proctor agents move from mock inputs to real engine inputs; the six tutor agents operate on authentic telemetry and artifacts.
|
||||
**Scope (deferred seams from D-016, to be specified at v0.5 Phase 0):** real server STT/TTS (`openai-audio` voice provider, CUT-1/G-7 seam), KYC/identity verification + age-gating backend (REQ-F-017), design/simulation sandbox environments (REQ-F-021 remainder), exec-telemetry seq-lease/replay-margin fix.
|
||||
|
||||
## Prior Milestone: v0.4 — Distribution & Bootstrap CLI (COMPLETE, shipped as v0.3.4)
|
||||
|
||||
**Scope (founder directive, 2026-09-12):** Streamline installing Nextcraft. Ship a bootstrap CLI with a single-liner install script, and publish release binaries on an ongoing basis for every release going forward.
|
||||
|
||||
**Delivered:** `nextcraft` CLI — `doctor` (prerequisite checks), `bootstrap` (deps + venv + env from templates + key validation), `verify` (health check), `dev` (thin passthrough to scripts/dev.sh); one-liner install script downloading the linux x64 binary from the latest Gitea release with sha256 + version integrity gates; binary release pipeline attached to every ship from v0.3.2 onward; install/quickstart documentation backed by a fresh-clone E2E test. All 5 requirements (REQ-4-001..005) complete.
|
||||
|
||||
**Status of v0.3:** Complete and shipped (v0.2.8). Credential engines live: namespace-isolated sandbox fabric, live build telemetry, process-trace grading, seeded variants, oral defense; real learner build/defense/grading surfaces.
|
||||
|
||||
**Status of v0.2:** Complete and shipped (v0.2.0). Six AI tutor agents live over mockengine inputs (D-015).
|
||||
|
||||
**Deferred from earlier plan:** REQ-F-017 (identity verification + 16+/18+ age-gating) is explicitly deferred to a later milestone per founder directive. Age-gating remains the v0.1-style visual flow mockup; no real KYC backend is built in v0.3.
|
||||
**Deferred from earlier plan:** REQ-F-017 (identity verification + age-gating KYC), real server STT/TTS, design/simulation sandbox environments, and the exec-telemetry seq-lease are all deferred to v0.5. Age-gating remains the v0.1-style visual flow mockup.
|
||||
|
||||
**Tech stack:** v0.1 TS monorepo (pnpm/turborepo, Next.js) + v0.2 Python FastAPI ai-service + new credential-engine services (sandbox fabric orchestrator, telemetry ingest, grading engine) in Python/TypeScript as determined at RESEARCH.
|
||||
|
||||
---
|
||||
|
||||
## Requirements (Validated)
|
||||
## v0.4 Requirements (Complete)
|
||||
|
||||
The following requirements have been validated during specification and are locked for milestone v0.3 (REQ-F-007..010 and REQ-F-021 activated from the deferred pool; REQ-F-017 deferred per founder directive):
|
||||
All 5 v0.4 requirements (REQ-4-001..005) are complete and shipped as v0.3.4:
|
||||
1. Bootstrap CLI — `nextcraft` executable with `doctor` / `bootstrap` / `verify` / `dev` (REQ-4-001, REQ-4-002)
|
||||
2. One-liner install — `curl | sh` fetching the linux x64 binary from the latest Gitea release with sha256 + version integrity verification (REQ-4-003)
|
||||
3. Ongoing release binaries — every release from v0.3.2 onward ships the CLI binary + checksum as release assets (REQ-4-004)
|
||||
4. Install documentation — README quickstart + CLI reference verified by a fresh-clone E2E test (REQ-4-005)
|
||||
|
||||
1. Sandbox fabric — sandboxed IDE, design tool, and simulation environments with isolated execution and lifecycle management (REQ-F-021)
|
||||
2. Live build telemetry — in-environment capture of process events (keystrokes, commands, file diffs, run/test results) streamed to ai-service (REQ-F-010)
|
||||
3. Process-trace grading engine — grade artifacts from their process traces, not just final output (REQ-F-007); feeds the Assessor agent real inputs
|
||||
4. Variant task generation — per-learner task variants so no two learners receive identical prompts (REQ-F-008)
|
||||
5. Oral/voice defense — AI examiner conducts spoken defense of submitted work (REQ-F-009); feeds the Proctor/Mentor agents
|
||||
6. Agent re-grounding — Lab/Assessor/Proctor consume real engine inputs (telemetry, traces, defenses) instead of v0.2 mocks
|
||||
7. Learner surface integration — wire the v0.1 sandbox + assessment mockups to the real engines (build/run in-browser, live telemetry, live defense)
|
||||
## v0.3 Requirements (Complete)
|
||||
|
||||
## v0.2 Requirements (Complete)
|
||||
|
||||
All 12 v0.2 requirements (REQ-2-001..012) are complete and shipped as v0.2.0. See REQUIREMENTS.md traceability matrix.
|
||||
All 8 v0.3 requirements (REQ-3-001..008) are complete and shipped as v0.2.8. See REQUIREMENTS.md traceability matrix.
|
||||
|
||||
## v0.1 Requirements (Complete)
|
||||
|
||||
All 28 v0.1 requirements (REQ-001..028) are complete and shipped as v0.1.0. See REQUIREMENTS.md traceability matrix.
|
||||
|
||||
## Clarified Assumptions (v0.4 CLARIFY stage, full autonomy — auto-resolved)
|
||||
|
||||
| # | Ambiguity | Resolution | Confidence |
|
||||
|---|-----------|------------|-------------|
|
||||
| A-201 | CLI language/toolchain for the binary? | **Probe-driven at RESEARCH** — Go → Rust → Node SEA → Python zipapp fallback chain; spec stays toolchain-agnostic so PLAN locks the probe-verified toolchain | 0.70 |
|
||||
| A-202 | Does bootstrap replace scripts/bootstrap.sh? | **No — reuse it.** CLI wraps existing `scripts/bootstrap.sh` + `scripts/dev.sh` via subprocess; zero orchestration logic duplicated in the CLI (thin passthrough pattern) | 0.85 |
|
||||
| A-203 | Where does the one-liner fetch the binary? | **Gitea latest-release API** (`/repos/{owner}/{repo}/releases/latest`) → download `nextcraft-linux-x64` + `.sha256` asset; repo raw serves `install.sh` as the stable URL | 0.80 |
|
||||
| A-204 | Install target + PATH? | **~/.local/bin** (XDG-style, no sudo), PATH hint printed when missing; `--dest` override flag | 0.85 |
|
||||
| A-205 | Binary "ongoing releases" scope? | **Every ship from v0.4 onward** attaches `nextcraft-linux-x64` + sha256 sidecar to the Gitea release — the ship workflow gains an asset step; retroactive binaries for old releases NOT required | 0.90 |
|
||||
| A-206 | No binary available yet / non-linux? | **Graceful degradation**: install script prints source-bootstrap instructions (git clone + scripts/bootstrap.sh) — never a hard fail | 0.88 |
|
||||
| A-207 | Checksum trust root? | **sha256 sidecar shipped as a release asset next to the binary** (same release, same channel); script verifies download against it. Signature/PKI out of scope for v0.4 (single forge, TLS transport) | 0.75 |
|
||||
| A-208 | Which prerequisites does doctor check? | node ≥18, pnpm ≥8, python3 ≥3.11, git, `unshare` availability (sandbox fabric needs it) — versions from the existing bootstrap tooling, not invented | 0.85 |
|
||||
| A-209 | Does `dev` manage multiple processes? | **No.** Thin passthrough to scripts/dev.sh only — the CLI stays bootstrap-scoped (D-016); orchestration remains in dev.sh | 0.82 |
|
||||
| A-210 | `.env.secrets` handling by bootstrap? | **Template copy only for `.env.example` → `.env`; secrets NEVER generated, NEVER committed; bootstrap validates presence of optional keys and warns (not blocks) when missing — mock-first providers keep the stack runnable** | 0.90 |
|
||||
|
||||
## Clarified Assumptions (v0.3 CLARIFY stage, full autonomy — auto-resolved)
|
||||
|
||||
| # | Ambiguity | Resolution | Confidence |
|
||||
|---|-----------|------------|-------------|
|
||||
| A-101 | Sandbox isolation technology? | **`unshare` user+mount+pid+net namespace subprocess isolation** per sandbox (probe-verified: in-ns uid=0, network fully isolated with 0 interfaces, writes land in an isolated bind-mounted workdir; proc-remount is not permitted in this context but is not required). No Docker/Podman/VMs — none present on the box; no sudo. A `SandboxBackend` protocol keeps a future containerd swap possible. Falls back further to a plain chroot-free subprocess with a cwd-jail if userns ever unavailable (tested path is userns). | 0.8 |
|
||||
| A-102 | Sandbox scope in v0.3? | **Coding IDE only** (web terminal + file tree + run/test). The "design tool" and "simulation" environments specified in REQ-F-021 are deferred to v0.4 — a single real build environment is enough to prove the credential pipeline end-to-end (telemetry → trace → grade → defense). | 0.75 |
|
||||
| A-103 | Live in-browser build UX? | **WebSocket xterm.js terminal** attached to the bwrap sandbox shell + HTTP file-tree/CRUD + run/test buttons. No full Monaco LSP in v0.3 — a code editor with syntax highlight (existing) + real shell is sufficient and far cheaper. | 0.72 |
|
||||
| A-103 | Live in-browser build UX? | **Run/Test buttons executing in the namespace sandbox + HTTP file-tree/CRUD + read-only exec-output panel** (CUT-2/G-8 — the interactive xterm.js shell relay is deferred to v0.4; `@xterm/*` is not a v0.3 dependency). No full Monaco LSP in v0.3 — a code editor with syntax highlight (existing) is sufficient and far cheaper. | 0.72 |
|
||||
| A-104 | Telemetry transport? | **WebSocket** from sandbox to a new ingestion endpoint on ai-service for live events; **SQLite-backed** ordered event log (`ai_service/telemetry/`) gives durability + at-least-once delivery + replay. Events carry monotonic `seq` per (learner,task) so gaps are detectable. | 0.8 |
|
||||
| A-105 | Where do traces live? | **SQLite** (`ai_service` data dir), introducing the first real persistence. SQLModel/SQLAlchemy for typed access. Chosen over Postgres because solo-founder + single box + low write volume; the `TraceStore` protocol is Postgres-migration-ready like SessionStore was. | 0.75 |
|
||||
| A-106 | Process-trace grading model? | **LLM-based grader**: structure the trace into a compact timeline digest (command categories, error/fix cycles, idle gaps, test passes) → Assessor-style rubric prompt → structured score via existing D-020 JSON defense. Deterministic features (test pass/fail, edit count) computed in code, not left to the LLM. | 0.7 |
|
||||
@@ -54,7 +73,7 @@ All 28 v0.1 requirements (REQ-001..028) are complete and shipped as v0.1.0. See
|
||||
| A-108 | Voice defense — STT/TTS providers? | **Provider-agnostic, mock-first like the LLM layer (D-014).** Real path: browser `MediaRecorder` → audio to ai-service → **OpenAI-compatible `/audio/transcriptions`** (Whisper STT) and **`/audio/speech`** (TTS) against ollama-cloud or a compatible endpoint; fallbacks: browser `SpeechRecognition`/`speechSynthesis` when no server keys. `VoiceProvider` protocol + deterministic mock (returns canned transcript) so tests never call a voice API. | 0.62 |
|
||||
| A-109 | Defense dialogue shape? | Reuse BaseAgent: an `Examiner` agent (seventh agent) streams examiner questions over the existing SSE pipeline; integrity signals (long pauses, off-scope answers, reading-from-notes cadence) emitted alongside the transcript to Proctor. | 0.8 |
|
||||
| A-110 | KYC / age-gating in v0.3? | **Deferred per founder directive.** No real identity backend. Age-gating stays the v0.1 visual flow mockup. Personas omit a security-engineer; security review via verifier + Phase 7 secrets-hygiene checklist. **Abuse control is NOT deferred with KYC (G-5):** v0.3 ships per-learner sandbox caps (`AI_SANDBOX_MAX_PER_LEARNER`), a global create-rate cap, and a server-side `learner_id` allowlist (`AI_LEARNER_ALLOWLIST`) so the unauthenticated surface cannot exhaust shared NPROC/disk. Documented in the release note. | 0.98 |
|
||||
| A-111 | New services vs extend ai-service? | **Extend ai-service**, don't fork new Python apps. Telemetry ingestion, trace grading, variant generation, voice, and sandbox orchestration all live as new modules in `apps/ai-service` (they share the LLM provider pool + config + session infra). Only the in-sandbox capture agent is a separate tiny Python process shipped into the bwrap environment. | 0.82 |
|
||||
| A-111 | New services vs extend ai-service? | **Extend ai-service**, don't fork new Python apps. Telemetry ingestion, trace grading, variant generation, voice, and sandbox orchestration all live as new modules in `apps/ai-service` (they share the LLM provider pool + config + session infra). Only the in-sandbox capture agent is a separate tiny Python process shipped into the namespace sandbox. | 0.82 |
|
||||
| A-112 | Sandbox on a single dev/school box — capacity? | v0.3 targets **1–5 concurrent sandboxes** (founder + pilot learners). No horizontal scaling, no queue. Concurrency guard returns 503 when full. Scaling is post-MVP. | 0.8 |
|
||||
|
||||
## Clarified Assumptions (v0.2 CLARIFY stage, full autonomy — auto-resolved)
|
||||
@@ -133,6 +152,7 @@ The following remain deferred beyond v0.3 and will be activated in subsequent mi
|
||||
| D-013 | v0.1 prototype founder-agreed; D-001 business-logic gate unlocked | Founder approved starting v0.2 with AI Tutor Architecture, which constitutes agreement of the v0.1 prototype per D-001. Recorded at v0.2 SPECIFY. | Business logic authorized from v0.2 onward |
|
||||
| D-014 | Provider-agnostic LLM layer; ollama-cloud as initial provider | OpenAI-compatible client abstraction with pluggable providers: ollama-cloud (https://ollama.com/v1, default), local OpenAI-compatible endpoint, deterministic mock (tests/CI). Keys in gitignored .ciagent/.env.secrets, never in code or commits. | apps/ai-service llm package with 3 providers; default=ollama-cloud |
|
||||
| D-015 | All six agents implemented as real LLM services; engines mocked | Coach/Tutor/Mentor fully real. Lab/Assessor/Proctor are real LLM logic over mock inputs (simulated telemetry, pre-baked artifacts) since sandbox fabric, assessment engine, and identity verification are v0.3+. Consistent with v0.1's mock-data approach. | REQ-F-001..006 complete in v0.2; real engines deferred to v0.3+ |
|
||||
| D-016 | v0.4 = Distribution & Bootstrap CLI (founder directive supersedes previously-named v0.4 seams) | Founder directive 2026-09-12: focus this milestone on streamlining install, a bootstrap CLI with a one-liner install script, ongoing release binaries. Real server STT/TTS, KYC, design/simulation envs, seq-lease move to v0.5. | Milestone scope locked at SPECIFY; binary = CLI-only, linux x64 |
|
||||
|
||||
---
|
||||
|
||||
|
||||
+67
-40
@@ -1,5 +1,22 @@
|
||||
# Nextcraft — REQUIREMENTS.md
|
||||
|
||||
## v0.4 Requirements (Complete — Distribution & Bootstrap CLI, shipped as v0.3.4)
|
||||
|
||||
### Bootstrap CLI
|
||||
|
||||
| ID | Description | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-4-001 | `nextcraft` CLI (linux x64 binary): `doctor` command checking prerequisites (node, pnpm, python3, git, unshare) with actionable error messages | critical | 1 | complete |
|
||||
| REQ-4-002 | `bootstrap` command: pnpm install, ai-service venv + pinned deps, .env from templates, key validation, .env.secrets handling; `verify` health check (ports, imports, builds); `dev` thin passthrough to scripts/dev.sh | critical | 1 | complete |
|
||||
|
||||
### Distribution
|
||||
|
||||
| ID | Description | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-4-003 | One-liner install script (`curl -fsSL <url> \| bash`): detects linux x64, resolves latest release from Gitea API, downloads binary + checksum, verifies sha256, installs to ~/.local/bin (PATH hint), degrades to source-bootstrap instructions when no binary | critical | 2 | complete |
|
||||
| REQ-4-004 | Binary release pipeline: reproducible linux x64 build script, sha256 checksum sidecar, upload as release assets on every ship from v0.4 onward (ongoing binaries requirement) | critical | 2 | complete |
|
||||
| REQ-4-005 | Install + quickstart documentation: README one-liner quickstart, CLI command reference, fresh-clone-to-running-stack end-to-end verification | high | 3 | complete |
|
||||
|
||||
## v0.3 Requirements (Credential Engines)
|
||||
|
||||
### Sandbox & Telemetry
|
||||
@@ -8,22 +25,22 @@
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-3-001 | Sandbox fabric: isolated per-learner execution environments (sandboxed IDE, design tool, simulation) with lifecycle management | critical | 1 | complete |
|
||||
| REQ-3-002 | Sandbox isolation + resource limits: per-learner isolation boundary, CPU/memory quotas (rlimits), wall-clock time quota, disk-quota via per-sandbox workdir usage sweep (best-effort, not kernel-enforced), no cross-tenant access, snapshot support. **Known gap (v0.3): per-sandbox pids and hard disk caps are NOT kernel-enforceable without cgroup delegation/sudo — documented as accepted risk** | critical | 1 | complete |
|
||||
| REQ-3-003 | Live build telemetry: in-environment capture of process events (commands, file diffs, run/test results, activity) streamed reliably to ai-service with per-learner trace persistence | critical | 2 | pending |
|
||||
| REQ-3-003 | Live build telemetry: in-environment capture of process events (commands, file diffs, run/test results, activity) streamed reliably to ai-service with per-learner trace persistence | critical | 2 | complete |
|
||||
|
||||
### Credential Engines
|
||||
|
||||
| ID | Description | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-3-004 | Process-trace grading engine: grade artifacts from their full process traces; rubric-aligned structured scores; feeds Assessor real inputs | critical | 3 | pending |
|
||||
| REQ-3-005 | Variant task generation: per-learner task variants (no two learners get identical prompts); variant seed registry; difficulty normalization | high | 4 | pending |
|
||||
| REQ-3-006 | Oral/voice defense: AI examiner conducts spoken defense (STT → dialogue → TTS); transcript + integrity signals captured; feeds Proctor/Mentor | high | 5 | pending |
|
||||
| REQ-3-004 | Process-trace grading engine: grade artifacts from their full process traces; rubric-aligned structured scores; feeds Assessor real inputs | critical | 3 | complete |
|
||||
| REQ-3-005 | Variant task generation: per-learner task variants (no two learners get identical prompts); variant seed registry; difficulty normalization | high | 4 | complete |
|
||||
| REQ-3-006 | Oral/voice defense: AI examiner conducts spoken defense (STT → dialogue → TTS); transcript + integrity signals captured; feeds Proctor/Mentor | high | 5 | complete |
|
||||
|
||||
### Agent Re-grounding & Integration
|
||||
|
||||
| ID | Description | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-3-007 | Agent re-grounding: Lab consumes live telemetry; Assessor consumes grading-engine output; Proctor consumes telemetry + defense integrity signals (replace v0.2 mocks) | critical | 6 | pending |
|
||||
| REQ-3-008 | Learner surface integration: sandbox mockup → real in-browser build/run with live telemetry; assessment mockup → live defense + live grading | critical | 6 | pending |
|
||||
| REQ-3-007 | Agent re-grounding: Lab consumes live telemetry; Assessor consumes grading-engine output; Proctor consumes telemetry + defense integrity signals (replace v0.2 mocks) | critical | 6 | complete |
|
||||
| REQ-3-008 | Learner surface integration: sandbox mockup → real in-browser build/run with live telemetry; assessment mockup → live defense + live grading | critical | 6 | complete |
|
||||
|
||||
---
|
||||
|
||||
@@ -65,58 +82,58 @@
|
||||
| ID | Description | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-001 | Monorepo scaffolding: pnpm workspaces, turborepo, Next.js app, TypeScript config, ESLint, Prettier | critical | 1 | complete |
|
||||
| REQ-002 | Shared component library: design tokens, Button, Input, Card, Badge, Avatar, Dialog, Navigation, Table, Tabs, Progress, Tooltip, Skeleton, Toast | critical | 1 | pending |
|
||||
| REQ-003 | Mock data layer: typed mock data for competency stacks, job listings, candidate profiles, employers, learner progress | critical | 1 | pending |
|
||||
| REQ-004 | Routing and navigation: App Router route groups for (learner), (marketplace), (employer), (admin); shared layout components; cross-surface navigation | critical | 1 | pending |
|
||||
| REQ-005 | Responsive layout system: mobile, tablet, desktop breakpoints; container components; grid system | high | 1 | pending |
|
||||
| REQ-002 | Shared component library: design tokens, Button, Input, Card, Badge, Avatar, Dialog, Navigation, Table, Tabs, Progress, Tooltip, Skeleton, Toast | critical | 1 | complete |
|
||||
| REQ-003 | Mock data layer: typed mock data for competency stacks, job listings, candidate profiles, employers, learner progress | critical | 1 | complete |
|
||||
| REQ-004 | Routing and navigation: App Router route groups for (learner), (marketplace), (employer), (admin); shared layout components; cross-surface navigation | critical | 1 | complete |
|
||||
| REQ-005 | Responsive layout system: mobile, tablet, desktop breakpoints; container components; grid system | high | 1 | complete |
|
||||
|
||||
### Learner Surface
|
||||
|
||||
| ID | Description | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-006 | Landing page: hero, value proposition, program highlights, how-it-works (Byte→Build→Demonstrate→Defend), testimonials mockup, CTA to program catalog | critical | 2 | pending |
|
||||
| REQ-007 | Program catalog: grid of competency stacks (AI Orchestration Engineer, AI Safety & Governance Lead, Human-AI Product Designer, AI-Augmented Field Operator, Computational Sciences Practitioner); stack cards with role descriptions | critical | 2 | pending |
|
||||
| REQ-008 | Competency stack view: selected stack detail with 12-18 competencies listed, progress indicators, microcredential badges, mastery status | critical | 2 | pending |
|
||||
| REQ-009 | Learner dashboard: active competencies, progress graph, recent artifacts, upcoming defenses, AI tutor chat mockup, milestone tracker | critical | 2 | pending |
|
||||
| REQ-010 | Byte tutorial viewer: 3-7 minute micro-tutorial layout with concept panel, worked example panel, code/design/simulation viewer mockup | high | 2 | pending |
|
||||
| REQ-011 | Build sandbox mockup: sandboxed IDE/design tool/simulation UI mockup with toolbar, file explorer, editor area, telemetry sidebar (process capture indicators) | high | 2 | pending |
|
||||
| REQ-012 | Assessment/defense mockup: rubric display, AI reviewer panel, oral defense interface (voice/mic mockup), process trace timeline, artifact viewer | high | 2 | pending |
|
||||
| REQ-006 | Landing page: hero, value proposition, program highlights, how-it-works (Byte→Build→Demonstrate→Defend), testimonials mockup, CTA to program catalog | critical | 2 | complete |
|
||||
| REQ-007 | Program catalog: grid of competency stacks (AI Orchestration Engineer, AI Safety & Governance Lead, Human-AI Product Designer, AI-Augmented Field Operator, Computational Sciences Practitioner); stack cards with role descriptions | critical | 2 | complete |
|
||||
| REQ-008 | Competency stack view: selected stack detail with 12-18 competencies listed, progress indicators, microcredential badges, mastery status | critical | 2 | complete |
|
||||
| REQ-009 | Learner dashboard: active competencies, progress graph, recent artifacts, upcoming defenses, AI tutor chat mockup, milestone tracker | critical | 2 | complete |
|
||||
| REQ-010 | Byte tutorial viewer: 3-7 minute micro-tutorial layout with concept panel, worked example panel, code/design/simulation viewer mockup | high | 2 | complete |
|
||||
| REQ-011 | Build sandbox mockup: sandboxed IDE/design tool/simulation UI mockup with toolbar, file explorer, editor area, telemetry sidebar (process capture indicators) | high | 2 | complete |
|
||||
| REQ-012 | Assessment/defense mockup: rubric display, AI reviewer panel, oral defense interface (voice/mic mockup), process trace timeline, artifact viewer | high | 2 | complete |
|
||||
|
||||
### Marketplace Surface
|
||||
|
||||
| ID | Description | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-013 | Job board listing: searchable grid of AI-era job listings, filter sidebar (skills, seniority, location, salary), result cards with match score | critical | 3 | pending |
|
||||
| REQ-014 | Job detail page: full job description, required competencies, employer info, application CTA, related jobs, AI-matched skills breakdown | critical | 3 | pending |
|
||||
| REQ-015 | Employer profile: company overview, logo, description, social links, open positions, company culture mockup | high | 3 | pending |
|
||||
| REQ-016 | Search/filter UI: semantic search bar, skill tags, category filters, seniority filter, remote/on-site toggle, saved searches mockup | critical | 3 | pending |
|
||||
| REQ-017 | Pricing page: job posting packages (single, bundle, enterprise), talent access plans, subscription tiers, feature comparison table | high | 3 | pending |
|
||||
| REQ-013 | Job board listing: searchable grid of AI-era job listings, filter sidebar (skills, seniority, location, salary), result cards with match score | critical | 3 | complete |
|
||||
| REQ-014 | Job detail page: full job description, required competencies, employer info, application CTA, related jobs, AI-matched skills breakdown | critical | 3 | complete |
|
||||
| REQ-015 | Employer profile: company overview, logo, description, social links, open positions, company culture mockup | high | 3 | complete |
|
||||
| REQ-016 | Search/filter UI: semantic search bar, skill tags, category filters, seniority filter, remote/on-site toggle, saved searches mockup | critical | 3 | complete |
|
||||
| REQ-017 | Pricing page: job posting packages (single, bundle, enterprise), talent access plans, subscription tiers, feature comparison table | high | 3 | complete |
|
||||
|
||||
### Employer Dashboard
|
||||
|
||||
| ID | Description | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-018 | Employer dashboard overview: active postings, applicant pipeline, talent matches, analytics mockup (charts, placement stats) | critical | 4 | pending |
|
||||
| REQ-019 | Talent search: searchable candidate database with AI-matched filters, candidate cards showing competency stack, microcredentials, artifact count, defense score | critical | 4 | pending |
|
||||
| REQ-020 | Candidate profile view: full candidate profile with artifact gallery, process trace summary, oral defense transcripts, competency graph, microcredential verification | high | 4 | pending |
|
||||
| REQ-021 | Posting management: create/edit/delete job postings, posting status tracking, applicant list per posting, interview pipeline mockup | high | 4 | pending |
|
||||
| REQ-018 | Employer dashboard overview: active postings, applicant pipeline, talent matches, analytics mockup (charts, placement stats) | critical | 4 | complete |
|
||||
| REQ-019 | Talent search: searchable candidate database with AI-matched filters, candidate cards showing competency stack, microcredentials, artifact count, defense score | critical | 4 | complete |
|
||||
| REQ-020 | Candidate profile view: full candidate profile with artifact gallery, process trace summary, oral defense transcripts, competency graph, microcredential verification | high | 4 | complete |
|
||||
| REQ-021 | Posting management: create/edit/delete job postings, posting status tracking, applicant list per posting, interview pipeline mockup | high | 4 | complete |
|
||||
|
||||
### Admin Surface
|
||||
|
||||
| ID | Description | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-022 | Admin overview: platform metrics (learners, employers, placements, completion rate, NPS), recent activity feed, system health mockup | critical | 5 | pending |
|
||||
| REQ-023 | Learner management: searchable learner table, learner detail view, progress tracking, competency completion status, credential issuance log | high | 5 | pending |
|
||||
| REQ-024 | Competency graph viewer: interactive visualization of competency stacks and their relationships, node/edge graph using react-flow, stack details on node click | high | 5 | pending |
|
||||
| REQ-025 | Marketplace moderation: job posting review queue, employer verification queue, flagged content, content moderation tools mockup | high | 5 | pending |
|
||||
| REQ-022 | Admin overview: platform metrics (learners, employers, placements, completion rate, NPS), recent activity feed, system health mockup | critical | 5 | complete |
|
||||
| REQ-023 | Learner management: searchable learner table, learner detail view, progress tracking, competency completion status, credential issuance log | high | 5 | complete |
|
||||
| REQ-024 | Competency graph viewer: interactive visualization of competency stacks and their relationships, node/edge graph using react-flow, stack details on node click | high | 5 | complete |
|
||||
| REQ-025 | Marketplace moderation: job posting review queue, employer verification queue, flagged content, content moderation tools mockup | high | 5 | complete |
|
||||
|
||||
### Polish & Integration
|
||||
|
||||
| ID | Description | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-026 | Cross-surface navigation: role switcher (learner/employer/admin), breadcrumbs, consistent header/footer across all surfaces | critical | 6 | pending |
|
||||
| REQ-027 | Visual consistency audit: typography scale, color palette, spacing system, dark mode toggle, accessibility baseline (WCAG AA contrast) | high | 6 | pending |
|
||||
| REQ-028 | Component library finalization: Storybook setup, component documentation, prop tables, usage examples | medium | 6 | pending |
|
||||
| REQ-026 | Cross-surface navigation: role switcher (learner/employer/admin), breadcrumbs, consistent header/footer across all surfaces | critical | 6 | complete |
|
||||
| REQ-027 | Visual consistency audit: typography scale, color palette, spacing system, dark mode toggle, accessibility baseline (WCAG AA contrast) | high | 6 | complete |
|
||||
| REQ-028 | Component library finalization: Storybook setup, component documentation, prop tables, usage examples | medium | 6 | complete |
|
||||
|
||||
---
|
||||
|
||||
@@ -170,18 +187,28 @@
|
||||
|
||||
## Traceability Matrix
|
||||
|
||||
### v0.3 (current milestone)
|
||||
### v0.4 (current milestone)
|
||||
|
||||
| Requirement | Phase | Status |
|
||||
|-------------|-------|--------|
|
||||
| REQ-4-001 | 1 | complete |
|
||||
| REQ-4-002 | 1 | complete |
|
||||
| REQ-4-003 | 2 | complete |
|
||||
| REQ-4-004 | 2 | complete |
|
||||
| REQ-4-005 | 3 | complete |
|
||||
|
||||
### v0.3 (complete)
|
||||
|
||||
| Requirement | Phase | Status |
|
||||
|-------------|-------|--------|
|
||||
| REQ-3-001 | 1 | complete |
|
||||
| REQ-3-002 | 1 | complete |
|
||||
| REQ-3-003 | 2 | pending |
|
||||
| REQ-3-004 | 3 | pending |
|
||||
| REQ-3-005 | 4 | pending |
|
||||
| REQ-3-006 | 5 | pending |
|
||||
| REQ-3-007 | 6 | pending |
|
||||
| REQ-3-008 | 6 | pending |
|
||||
| REQ-3-003 | 2 | complete |
|
||||
| REQ-3-004 | 3 | complete |
|
||||
| REQ-3-005 | 4 | complete |
|
||||
| REQ-3-006 | 5 | complete |
|
||||
| REQ-3-007 | 6 | complete |
|
||||
| REQ-3-008 | 6 | complete |
|
||||
|
||||
### v0.2 (complete)
|
||||
|
||||
|
||||
+65
-107
@@ -2,15 +2,17 @@
|
||||
|
||||
## Overview
|
||||
|
||||
**Milestone v0.3** — Credential Engines: Replace v0.2's mock engine inputs with real credential engines. Build the sandbox fabric (sandboxed IDE / design tool / simulation), the live in-environment build-telemetry pipeline, the process-trace grading engine, per-learner variant task generation, and the oral/voice defense with AI examiner. Lab/Assessor/Proctor agents move from mock inputs to real engine inputs.
|
||||
**Milestone v0.4 — COMPLETE (shipped as v0.3.4, 2026-09-13).** Distribution & Bootstrap CLI: `nextcraft` CLI (doctor/bootstrap/verify/dev) shipped as a linux x64 SEA binary, one-liner install script with checksum + version integrity gates, and binaries published on **every ongoing release** (v0.3.2 onward). Next milestone: v0.5 (real server STT/TTS + KYC/identity + design/simulation sandbox environments + exec-telemetry seq-lease — the seams deferred out of v0.4 by founder directive D-016).
|
||||
|
||||
**Deferred per founder directive:** REQ-F-017 identity verification + age-gating (real KYC backend) is deferred beyond v0.3. Age-gating remains the v0.1 visual flow mockup.
|
||||
**Milestone v0.3** — Credential Engines: complete, shipped as v0.2.8 (2026-09-12). Real sandbox fabric, live build telemetry, process-trace grading, per-learner variants, oral defense, real learner surfaces.
|
||||
|
||||
**Deferred per founder directive (D-016):** REQ-F-017 identity verification + age-gating (real KYC backend), real server STT/TTS, design/simulation sandbox environments, and the exec-telemetry seq-lease are deferred to v0.5. Age-gating remains the v0.1 visual flow mockup.
|
||||
|
||||
**Prior milestone:** v0.2 (ai-tutor-architecture) — complete, shipped as v0.2.0, six tutor agents live over mock engine inputs (D-015).
|
||||
|
||||
**Milestone type:** Feature (new credential-engine services + real agent inputs)
|
||||
**Tag line:** v0.2.x (patches on the v0.2 line; milestone release as the final v0.2.x patch)
|
||||
**Branch:** milestone/v0.3-credential-engines
|
||||
**Milestone type:** Feature (new CLI + distribution pipeline)
|
||||
**Tag line:** v0.3.x (patches on the v0.3 line; milestone release as the final v0.3.x patch)
|
||||
**Branch:** milestone/v0.4-distribution
|
||||
|
||||
---
|
||||
|
||||
@@ -18,14 +20,11 @@
|
||||
|
||||
| # | Name | Status | Depends On | Requirements | Success Criteria |
|
||||
|---|------|--------|------------|--------------|------------------|
|
||||
| 0 | Pre-execution | in-progress | — | — | Specification, clarify, research, plan complete; .ciagent/ files updated for v0.3 |
|
||||
| 1 | Sandbox fabric | complete | 0 | REQ-3-001, REQ-3-002 | Isolated per-learner sandbox environments provisioned (IDE / design / simulation); lifecycle API (create/destroy/snapshot); resource limits enforced; no cross-tenant access |
|
||||
| 2 | Live build telemetry | pending | 1 | REQ-3-003 | In-environment capture of process events (commands, file diffs, run/test results, keystroke-level activity) streamed to ai-service; reliable transport; per-learner trace persistence |
|
||||
| 3 | Process-trace grading engine | pending | 2 | REQ-3-004 | Grades artifacts from their full process traces (not just final output); emits structured rubric-aligned scores; feeds Assessor real inputs |
|
||||
| 4 | Variant task generation | pending | 1 | REQ-3-005 | Per-learner task variants generated so no two learners receive identical prompts; variant seed recorded for grading fairness |
|
||||
| 5 | Oral / voice defense | pending | 3 | REQ-3-006 | AI examiner conducts spoken defense of submitted work; STT → dialogue → TTS; transcript + integrity signals captured; feeds Proctor/Mentor |
|
||||
| 6 | Agent re-grounding + learner surface integration | pending | 2,3,4,5 | REQ-3-007, REQ-3-008 | Lab/Assessor/Proctor consume real engine inputs; v0.1 sandbox + assessment mockups wired to real engines (in-browser build/run, live telemetry, live defense) |
|
||||
| 7 | Final review + ship | pending | 6 | — | Code review clean; audit passes; milestone tagged (v0.2.x final patch); release created on Gitea |
|
||||
| 0 | Pre-execution | complete | — | — | Specification, clarify, research, plan complete; .ciagent/ files updated for v0.4 |
|
||||
| 1 | Bootstrap CLI core | complete | 0 | REQ-4-001, REQ-4-002 | `nextcraft doctor/bootstrap/verify/dev` work against a fresh clone; unit tests green |
|
||||
| 2 | Binary build + release pipeline | complete | 1 | REQ-4-003, REQ-4-004 | Reproducible linux x64 binary + sha256 checksum; one-liner install script; assets uploaded to the Gitea release |
|
||||
| 3 | Install docs + fresh-clone E2E | complete | 2 | REQ-4-005 | README quickstart verified end-to-end from a clean environment; fresh clone reaches running stack |
|
||||
| 4 | Final review + ship | complete | 3 | — | Code review clean; audit passes; milestone tagged (v0.3.x final patch); release with binary assets created on Gitea |
|
||||
|
||||
---
|
||||
|
||||
@@ -33,148 +32,107 @@
|
||||
|
||||
### Phase 0: Pre-execution
|
||||
|
||||
**Goal:** Establish v0.3 specification, clarify ambiguities, research credential-engine architecture (sandbox isolation, telemetry transport, trace grading, variant generation, voice IO), create detailed plans.
|
||||
**Goal:** Establish v0.4 specification (founder directive D-016), clarify ambiguities, research the binary toolchain + Gitea release-asset API + existing bootstrap scripts, create detailed plans, grill adversarially.
|
||||
|
||||
**Stages:** SPECIFY → CLARIFY → RESEARCH → PLAN → GRILL → MVP/UX CHECK → SHIP
|
||||
|
||||
**Deliverables:**
|
||||
- Updated .ciagent/config.json, PROJECT.md, REQUIREMENTS.md, ROADMAP.md, ARCHITECTURE.md, PERSONAS.md, PLAN.md
|
||||
|
||||
**Success criteria:** All .ciagent/ files updated for v0.3; phase 0 shipped as first v0.2.x patch.
|
||||
**Success criteria:** All .ciagent/ files updated for v0.4; phase 0 shipped as v0.3.0.
|
||||
|
||||
---
|
||||
|
||||
### Phase 1: Sandbox Fabric
|
||||
### Phase 1: Bootstrap CLI Core
|
||||
|
||||
**Goal:** Provision and manage isolated per-learner execution environments.
|
||||
**Goal:** A working `nextcraft` CLI with doctor/bootstrap/verify/dev commands, unit-tested against the real monorepo.
|
||||
|
||||
**Requirements:** REQ-3-001, REQ-3-002
|
||||
**Requirements:** REQ-4-001, REQ-4-002
|
||||
|
||||
**Key deliverables:**
|
||||
- Sandbox orchestrator service: create/list/destroy/snapshot sandbox instances (IDE, design tool, simulation)
|
||||
- Isolation boundary: per-learner containerization or VM-grade isolation; no cross-tenant filesystem/network access
|
||||
- Resource limits: CPU/memory/disk/time quotas per sandbox
|
||||
- Sandbox lifecycle API consumed by ai-service and the web learner surface
|
||||
- `apps/cli` package: `nextcraft` executable (source-runnable in dev, binary-built in P2)
|
||||
- `doctor`: checks node ≥18, pnpm, python3 ≥3.11, git, unshare availability — actionable errors, exit codes
|
||||
- `bootstrap`: idempotent — pnpm install, ai-service venv + pinned deps (reuses scripts/bootstrap.sh logic), .env from .env.example templates, key validation (warnings not blockers for optional keys), .env.secrets handling
|
||||
- `verify`: health check — venv imports, pnpm build readiness, ports free, env vars present
|
||||
- `dev`: thin passthrough to scripts/dev.sh (no orchestration logic duplicated)
|
||||
- Unit tests: doctor/bootstrap parsing + command dispatch, against fixtures (never modifying the real repo state)
|
||||
|
||||
**Success criteria:**
|
||||
- A sandbox can be created, written to, snapshotted, and destroyed via API
|
||||
- Isolation verified: a sandbox cannot read another learner's data
|
||||
- Resource limits enforced and observable — enforcement mechanism: rlimits (memory/CPU) + wall-clock reaper + workdir-size sweep; per-sandbox pids and hard-disk-quota are accepted v0.3 gaps (no cgroup delegation/sudo on this box, G-1)
|
||||
- `nextcraft doctor` reports each prerequisite with actionable guidance
|
||||
- `nextcraft bootstrap` on a fresh clone reaches a state where `verify` passes
|
||||
- All commands have `--help`, exit non-zero on failure, no shell-out without timeout
|
||||
- `pnpm build`, `pnpm typecheck`, `pnpm ai:test` green
|
||||
|
||||
---
|
||||
|
||||
### Phase 2: Live Build Telemetry
|
||||
### Phase 2: Binary Build + Release Pipeline
|
||||
|
||||
**Goal:** Capture in-environment process events and stream them to ai-service reliably.
|
||||
**Goal:** Reproducible linux x64 binary + one-liner install + release-asset upload wired into the ship flow.
|
||||
|
||||
**Requirements:** REQ-3-003
|
||||
**Requirements:** REQ-4-003, REQ-4-004
|
||||
|
||||
**Key deliverables:**
|
||||
- Telemetry capture agent (in-sandbox): commands, file diffs, run/test results, keystroke-level/activity events
|
||||
- Telemetry transport: durable, ordered, resumable stream to ai-service ingestion endpoint
|
||||
- Trace persistence: per-learner, per-task process traces stored for grading and proctoring
|
||||
- Transport hardening: retries, backpressure, exactly-once-or-at-least-once semantics documented
|
||||
- Build script producing `nextcraft-linux-x64` + `nextcraft-linux-x64.sha256` (toolchain probe-verified at RESEARCH; embedded script assets)
|
||||
- One-liner install script `install.sh` served from the repo: detect linux x64, resolve latest release via Gitea API, download + verify checksum, install to `~/.local/bin`, PATH hint, source-bootstrap fallback when no binary/asset
|
||||
- Ship integration: every release from v0.4 onward attaches the binary + checksum as release assets (the "ongoing binaries" requirement)
|
||||
- Asset-upload helper using the Gitea token from `.env*` files only (never shell env)
|
||||
|
||||
**Success criteria:**
|
||||
- Sandbox activity produces a complete ordered process trace in ai-service
|
||||
- Stream survives transient network failure without trace loss
|
||||
- Trace retrievable by learner+task ID for grading
|
||||
- Binary runs on this box: `./nextcraft-linux-x64 doctor` green against the repo
|
||||
- Install script verified end-to-end against the real Gitea release (or local dry-run if release pending)
|
||||
- Checksum verification rejects a corrupted download (tested)
|
||||
- Release assets present on the phase ship
|
||||
|
||||
---
|
||||
|
||||
### Phase 3: Process-Trace Grading Engine
|
||||
### Phase 3: Install Docs + Fresh-Clone E2E
|
||||
|
||||
**Goal:** Grade learner artifacts from their full process traces.
|
||||
**Goal:** Documentation and end-to-end proof that a fresh consumer reaches a running stack via the one-liner.
|
||||
|
||||
**Requirements:** REQ-3-004
|
||||
**Requirements:** REQ-4-005
|
||||
|
||||
**Key deliverables:**
|
||||
- Trace analyzer: reconstructs build/decision timeline from a process trace
|
||||
- Grading engine: rubric-aligned scoring over the trace (process quality, not just final artifact)
|
||||
- Structured score output consumable by the Assessor agent
|
||||
- Calibration against v0.2 mock corpora to validate grading dimensions
|
||||
- README quickstart: one-liner → `nextcraft doctor` → `nextcraft bootstrap` → `nextcraft dev`
|
||||
- CLI command reference (all flags, exit codes)
|
||||
- Fresh-clone E2E test: clean temp clone → doctor → bootstrap → verify → build green (sandboxed; no network beyond package registries already used)
|
||||
- Install-script docs: prerequisites, offline/manual install, troubleshooting
|
||||
|
||||
**Success criteria:**
|
||||
- Engine emits structured rubric-aligned scores from a real process trace
|
||||
- Scores distinguish process quality (e.g., iterative debugging vs. paste-and-run)
|
||||
- Output feeds Assessor; replaces pre-baked artifact corpus inputs
|
||||
- A fresh clone bootstraps to a passing `verify` with one command sequence
|
||||
- README quickstart matches the actual tested flow exactly
|
||||
- E2E test green in CI-equivalent local run
|
||||
|
||||
---
|
||||
|
||||
### Phase 4: Variant Task Generation
|
||||
### Phase 4: Final Review + Ship
|
||||
|
||||
**Goal:** Generate per-learner task variants so no two learners receive identical prompts.
|
||||
|
||||
**Requirements:** REQ-3-005
|
||||
|
||||
**Key deliverables:**
|
||||
- Variant generator: parameterized task templates → unique per-learner instances
|
||||
- Variant seed registry: record variant parameters for grading fairness and proctoring
|
||||
- Difficulty normalization: variants calibrated to equivalent difficulty
|
||||
|
||||
**Success criteria:**
|
||||
- Two learners requesting the same competency receive distinct task variants
|
||||
- Variant parameters persisted and auditable
|
||||
- Grading engine scores variants equitably
|
||||
|
||||
---
|
||||
|
||||
### Phase 5: Oral / Voice Defense
|
||||
|
||||
**Goal:** AI examiner conducts a spoken defense of the learner's submitted work.
|
||||
|
||||
**Requirements:** REQ-3-006
|
||||
|
||||
**Key deliverables:**
|
||||
- Voice pipeline: STT → defense dialogue (LLM examiner) → TTS
|
||||
- Examiner agent: probes understanding, challenges process choices from the trace
|
||||
- Transcript + integrity signals captured for Proctor/Mentor
|
||||
- Latency budget: defense feels conversational (bounded turn latency)
|
||||
|
||||
**Success criteria:**
|
||||
- A spoken defense runs end-to-end (speak → examiner question → learner response → verdict)
|
||||
- Transcript + integrity signals persisted and consumable by Proctor
|
||||
- Turn latency within the documented budget
|
||||
|
||||
---
|
||||
|
||||
### Phase 6: Agent Re-grounding + Learner Surface Integration
|
||||
|
||||
**Goal:** Move Lab/Assessor/Proctor to real engine inputs; wire learner surface to the real engines.
|
||||
|
||||
**Requirements:** REQ-3-007, REQ-3-008
|
||||
|
||||
**Key deliverables:**
|
||||
- Lab agent consumes live sandbox telemetry (replaces v0.2 mock telemetry)
|
||||
- Assessor agent consumes grading-engine output (replaces pre-baked artifacts)
|
||||
- Proctor consumes telemetry + defense integrity signals (replaces mock telemetry)
|
||||
- Learner sandbox mockup → real in-browser build/run; assessment mockup → live defense + live grading
|
||||
|
||||
**Success criteria:**
|
||||
- Lab/Assessor/Proctor operate on real inputs with no mock fallback in the learner path
|
||||
- Learner can build in-browser and see live telemetry + live feedback
|
||||
- Assessment surface runs a live defense and shows live grading
|
||||
- `pnpm build` and `pnpm typecheck` pass
|
||||
|
||||
---
|
||||
|
||||
### Phase 7: Final Review + Ship
|
||||
|
||||
**Goal:** Code review, audit, milestone release.
|
||||
**Goal:** Code review, audit, milestone release with binary assets.
|
||||
|
||||
**Key deliverables:**
|
||||
- Multi-persona code review (correctness, testing, security, performance, maintainability)
|
||||
- Project health audit (reconstruction test, .ciagent/ file discipline, branch hygiene, commit discipline)
|
||||
- Milestone ship: merge milestone → main, tag final v0.2.x patch, create Gitea release
|
||||
- Milestone ship: merge milestone → main, tag final v0.3.x patch, create Gitea release WITH binary + checksum assets, verify assets downloadable
|
||||
|
||||
**Success criteria:**
|
||||
- Code review: P0 fixes applied, P1+ documented
|
||||
- Audit: all checks pass, project state reconstructable from git log
|
||||
- Ship: milestone tagged, branch merged to main, Gitea release created — release note states identity/age-gating (KYC) is deferred and age-gating remains a visual mockup
|
||||
- All v0.3 requirements marked complete
|
||||
- Ship: milestone tagged, branch merged to main, Gitea release created with `nextcraft-linux-x64` + `.sha256` assets attached — the first of the ongoing binary releases
|
||||
|
||||
---
|
||||
|
||||
## v0.4 (Complete — Shipped as v0.3.4)
|
||||
|
||||
Distribution & Bootstrap CLI: `nextcraft` CLI (doctor/bootstrap/verify/dev) as a self-contained linux x64 SEA binary, one-liner install with sha256 + version integrity gates, release-asset pipeline attaching binaries to every ongoing release (v0.3.2 onward), install/quickstart docs backed by a fresh-clone E2E test. 5 phases (P0–P4). All 5 requirements (REQ-4-001..005) complete. Tags v0.3.0–v0.3.3 per phase, milestone release v0.3.4.
|
||||
|
||||
## v0.3 (Complete — Shipped as v0.2.8)
|
||||
|
||||
Credential Engines: real sandbox fabric (Linux namespaces), live build telemetry
|
||||
(at-least-once/exactly-once), process-trace grading (G-4 gated), seeded per-learner
|
||||
variants (fairness anchors wired to grading), oral defense with integrity signals
|
||||
(mock-first voice, browser fallback), and real learner build/defense/grading surfaces.
|
||||
8 phases. All 8 requirements (REQ-3-001..008) complete. Tags v0.2.1–v0.2.7 per phase,
|
||||
milestone release v0.2.8.
|
||||
|
||||
## v0.2 (Complete — Shipped as v0.2.0)
|
||||
|
||||
AI Tutor Architecture: Six AI tutor agents (Coach, Tutor, Lab, Assessor, Proctor, Mentor) as real LLM-backed services over mock engine inputs, wired into the learner surface with streaming. 7 phases. All 12 requirements complete. Milestone release v0.2.0.
|
||||
|
||||
@@ -46,9 +46,9 @@
|
||||
"projects": [],
|
||||
"active_project": null,
|
||||
"milestone": {
|
||||
"version": "v0.3",
|
||||
"name": "credential-engines",
|
||||
"version": "v0.4",
|
||||
"name": "distribution",
|
||||
"type": "feature",
|
||||
"branch": "milestone/v0.3-credential-engines"
|
||||
"branch": "milestone/v0.4-distribution"
|
||||
}
|
||||
}
|
||||
@@ -55,3 +55,6 @@ apps/ai-service/ai_service/data/
|
||||
apps/ai-service/**/sandboxes/
|
||||
*.db
|
||||
*.db-journal
|
||||
|
||||
# in-sandbox capture agent runtime spool
|
||||
.nc-agent/
|
||||
|
||||
@@ -2,8 +2,71 @@
|
||||
|
||||
AI-native outcome school + marketplace — graduates prove what they can build, not what they can write.
|
||||
|
||||
## Quickstart
|
||||
|
||||
One-liner install (linux x64):
|
||||
|
||||
```sh
|
||||
curl -fsSL https://git.coreci.dev/coreci/nextcraft/raw/main/scripts/install.sh | sh
|
||||
```
|
||||
|
||||
That downloads the latest release's `nextcraft` CLI binary, verifies its sha256 checksum, and installs it to `~/.local/bin` (PATH hint printed if needed). Every release ships fresh binaries — re-run the one-liner to upgrade.
|
||||
|
||||
Then, from a clone of this repo:
|
||||
|
||||
```sh
|
||||
nextcraft doctor # check prerequisites: node >= 18, pnpm >= 8, python3 >= 3.11, git, unshare
|
||||
nextcraft bootstrap # pnpm install + ai-service venv + .env from template (idempotent)
|
||||
nextcraft verify # health check: venv imports, uvicorn, ports, env
|
||||
nextcraft dev # run the ai-service dev server on :8420 (web dev server: pnpm dev)
|
||||
```
|
||||
|
||||
The E2E test (`apps/cli/tests/fresh-clone-e2e.test.ts`) proves this exact sequence on a fresh clone.
|
||||
|
||||
### No binary / non-linux?
|
||||
|
||||
The installer degrades to printed source instructions. Manual equivalent:
|
||||
|
||||
```sh
|
||||
git clone https://git.coreci.dev/coreci/nextcraft.git && cd nextcraft
|
||||
pnpm install
|
||||
bash apps/ai-service/scripts/bootstrap.sh
|
||||
cp apps/ai-service/.env.example apps/ai-service/.env
|
||||
pnpm ai:dev
|
||||
```
|
||||
|
||||
## CLI reference (`nextcraft`)
|
||||
|
||||
| Command | What it does | Exit codes |
|
||||
|---------|--------------|------------|
|
||||
| `doctor` | Checks prerequisites on PATH: node >= 18, pnpm >= 8, python3 >= 3.11, git, unshare (sandbox fabric). Every ✗ prints a fix hint. | 0 all pass, 1 any fail |
|
||||
| `bootstrap` | Sets up the monorepo from a fresh clone: (1) locates the repo root, (2) `pnpm install`, (3) ai-service venv via `apps/ai-service/scripts/bootstrap.sh`, (4) copies `.env.example` → `.env` if absent, (5) warns on missing optional keys. Idempotent — safe to re-run. | 0 ok, 1 step failed |
|
||||
| `verify` | Health check: ai-service venv + `import ai_service`, uvicorn importable, `.env` present (warn-only), `AI_PORT` (default 8420) free, workspace `node_modules` present. | 0 ok, 1 failures |
|
||||
| `dev` | Thin passthrough to `apps/ai-service/scripts/dev.sh` (exports secrets from `.ciagent/.env.secrets` if present, runs uvicorn on :8420). Ctrl+C stops it. The web dev server is separate: `pnpm dev`. | child's exit code |
|
||||
| `--help` / `-h` | Usage for the CLI or any command. | 0 |
|
||||
| `--version` | Prints the version this binary was built as (matches the release tag). | 0 |
|
||||
|
||||
Exit-code contract: `0` success, `1` check/step failure (hint printed), `2` usage error.
|
||||
|
||||
### Remote server
|
||||
|
||||
`nextcraft dev` binds the API on **0.0.0.0:8420** (and `pnpm dev` serves the web app on all interfaces), so the stack works from other machines out of the box:
|
||||
|
||||
- Browse `http://<your-host>:3000` — the web app targets `http://<your-host>:8420` automatically (derived from the browser's hostname).
|
||||
- CORS admits any origin (`AI_CORS_ORIGINS=*` in `apps/ai-service/.env`). This is safe **only** because credentials are never enabled; to restrict, set an explicit list: `AI_CORS_ORIGINS=http://<your-host>:3000`.
|
||||
- To revert to loopback-only: `AI_HOST=127.0.0.1` in `apps/ai-service/.env`.
|
||||
- Security note: this is an unauthenticated dev API reachable from any network the box exposes. Mitigations that still apply: per-learner sandbox caps + global rate caps + learner allowlist (G-5), telemetry flood control (traces marked `INCOMPLETE_FLOODED` are refused by the grader). Expose only on trusted networks until identity/KYC lands (v0.5).
|
||||
|
||||
## Docs
|
||||
|
||||
- [apps/cli/README.md](apps/cli/README.md) — CLI internals: build, binary pipeline, troubleshooting
|
||||
- [.ciagent/PROJECT.md](.ciagent/PROJECT.md) — product spec and milestone history
|
||||
- [.ciagent/ARCHITECTURE.md](.ciagent/ARCHITECTURE.md) — system architecture
|
||||
|
||||
## Status
|
||||
|
||||
**Milestone v0.1** — UI/UX Prototype (high-fidelity interactive, all mock data)
|
||||
**Milestone v0.4** — Distribution & Bootstrap CLI (one-liner install, `nextcraft` binary releases on every ship)
|
||||
|
||||
Prior: v0.3 Credential Engines (shipped v0.2.8) · v0.2 AI Tutor Architecture (v0.2.0) · v0.1 UI/UX Prototype (v0.1.0)
|
||||
|
||||
Initialized via CIAgent v0.7.0
|
||||
@@ -2,6 +2,13 @@
|
||||
# Real keys live in .ciagent/.env.secrets (gitignored) and are exported by scripts/dev.sh
|
||||
|
||||
AI_PORT=8420
|
||||
# Network mode (v0.3.5, D-038): dev server binds 0.0.0.0 so remote machines can
|
||||
# reach the stack. Set to 127.0.0.1 to revert to loopback-only.
|
||||
AI_HOST=0.0.0.0
|
||||
# CORS + WS-origin policy: '*' (default) admits any origin — safe because
|
||||
# credentials are never enabled. Restrict with a comma list, e.g.:
|
||||
# AI_CORS_ORIGINS=http://nextcraft-1:3000
|
||||
AI_CORS_ORIGINS=*
|
||||
AI_PROVIDER=ollama-cloud
|
||||
AI_MODEL=gemma4:31b
|
||||
AI_OLLAMA_CLOUD_BASE_URL=https://ollama.com/v1
|
||||
@@ -24,4 +31,9 @@ AI_SANDBOX_MAX_PER_LEARNER=1
|
||||
AI_SANDBOX_CREATES_PER_MIN=10
|
||||
|
||||
# Persistence (SQLite)
|
||||
AI_DB_PATH=ai_service/data/nextcraft.db
|
||||
AI_DB_PATH=ai_service/data/nextcraft.db
|
||||
# --- v0.3 Voice (REQ-3-006, D-030) ---
|
||||
# 'mock' (default; no key needed — tests/dev) or 'browser' (client-native SR/TTS).
|
||||
# Real server STT/TTS ('openai-audio' + AI_VOICE_BASE_URL/AI_VOICE_API_KEY)
|
||||
# is deferred to v0.4 per GRILL CUT-1/G-7 — keys never in code or commits.
|
||||
AI_VOICE_PROVIDER=mock
|
||||
|
||||
@@ -194,6 +194,65 @@ Unprivileged user namespaces are on the box's kernel and need neither a
|
||||
daemon, nor suid helpers, nor network access — they are the only isolation
|
||||
primitive that works here, so that's what v0.3 uses.
|
||||
|
||||
## Telemetry delivery semantics (v0.3, REQ-3-003)
|
||||
|
||||
Delivery is **at-least-once**; storage is **exactly-once** — the two compose:
|
||||
|
||||
- The in-sandbox capture agent (stdlib-only, `scripts/sandbox-agent.py`)
|
||||
spools every event to a durable JSONL file (fsync per append) BEFORE any
|
||||
send attempt, so no event can be lost to a dead socket or a SIGKILL.
|
||||
- The WS ingest endpoint (`WS /v1/telemetry/ingest?learner_id&task_id`,
|
||||
D-026) dedups server-side on the `(learner_id, task_id, seq)` primary key:
|
||||
re-sends (reconnect flushes, replay margin) are collapsed, never upserted.
|
||||
- On disconnect the agent reconnects with exponential backoff and flushes
|
||||
the spool in `seq` order; a transient outage therefore loses nothing and
|
||||
stores each event exactly once (`tests/telemetry/test_durability.py`
|
||||
proves this end-to-end against a real namespace sandbox + live server).
|
||||
- Replay/read path: `GET /v1/telemetry/traces/{learner}/{task}` returns the
|
||||
complete ordered trace; `GET /v1/telemetry/gaps/{learner}/{task}` returns
|
||||
missing seqs for gap detection.
|
||||
- Flood boundary (G-3): a connection exceeding `AI_TELEMETRY_MAX_EVENTS_PER_TASK`
|
||||
(default 50,000) is closed with WS code 1008 and its trace is marked
|
||||
`INCOMPLETE_FLOODED` — a terminal integrity flag the grader refuses to
|
||||
grade. Silent event dropping is forbidden: it would corrupt grading input.
|
||||
|
||||
## Voice defense (v0.3, REQ-3-006)
|
||||
|
||||
Voice is **mock-first** (D-030): the defense pipeline is fully proven over
|
||||
the deterministic `MockVoiceProvider` + browser-native fallback — no task
|
||||
requires a real voice key. Real server STT/TTS (`OpenAIAudioProvider` over
|
||||
OpenAI-compatible `/audio/transcriptions` + `/audio/speech`) is **deferred
|
||||
to v0.4** together with KYC (GRILL CUT-1 / G-7): it could never be exercised
|
||||
in CI, so v0.3 ships the protocol seam instead of an unverifiable claim.
|
||||
|
||||
- `AI_VOICE_PROVIDER=mock` (default) — deterministic canned STT/TTS
|
||||
- `AI_VOICE_PROVIDER=browser` — the web client uses SpeechRecognition +
|
||||
speechSynthesis; the server keeps text-turn persistence
|
||||
- Conversational budget: a defense turn should complete in **< 4s**
|
||||
(`DEFENSE_TURN_BUDGET_MS` in `tests/voice/test_latency.py`). v0.3
|
||||
asserts instrumentation (stt_ms/llm_ms/tts_ms populated per turn); the
|
||||
wall-clock acceptance probe against a real voice endpoint is a v0.4
|
||||
criterion, run manually with `AI_VOICE_PROVIDER` set to the real
|
||||
provider and keys in `.ciagent/.env.secrets` (never in code/commits).
|
||||
|
||||
## End-to-end credential flow (v0.3, REQ-3-007/008)
|
||||
|
||||
`tests/api/test_e2e_credential_flow.py` runs the full pipeline against a REAL
|
||||
uvicorn server with REAL namespace sandboxes (mock LLM/voice per G-2
|
||||
precedent): variant -> telemetry-wired sandbox -> in-sandbox exec -> trace
|
||||
persistence -> process-trace grade (variant seed stamped) -> assessor
|
||||
coaching -> oral defense -> verdict + integrity signals -> proctor. It
|
||||
asserts no corpus fixture appears anywhere in the learner path.
|
||||
|
||||
Manual browser pass (documented, not automated): `pnpm ai:dev` + `pnpm dev`,
|
||||
then open `/build/stack-orchestration-c007` — variant statement + starter
|
||||
files load, edit a file, Run/Test execute in the sandbox with output in the
|
||||
read-only panel, the telemetry status pulses, Lab streams feedback from the
|
||||
live digest; then `/defend/stack-orchestration-c007` — Start Defense, typed
|
||||
answers (mic path needs permission), Finish, Grade My Work renders the real
|
||||
rubric bars. Navigating away destroys the sandbox
|
||||
(`curl localhost:8420/v1/sandboxes` shows the count drop).
|
||||
|
||||
## Layout
|
||||
|
||||
```
|
||||
|
||||
@@ -1,53 +1,34 @@
|
||||
"""AssessorAgent — rubric application to pre-baked artifacts (REQ-2-008).
|
||||
"""AssessorAgent — rubric coaching over REAL grading output (REQ-3-007).
|
||||
|
||||
Structured-output showcase: applies the 4-layer defense (D-020) to return
|
||||
a pydantic-validated rubric score. Mock engine inputs (corpus artifacts +
|
||||
transcripts); real process-trace grading is v0.3+.
|
||||
v0.3 re-grounding: the Assessor no longer invents scores from corpus
|
||||
artifacts — the process-trace grading engine (Phase 3) computes and
|
||||
persists the validated RubricScore. This agent now renders the STORED
|
||||
grade as rubric-anchored coaching: explains the criteria, cites strengths
|
||||
and gaps, and frames next steps. Corpus artifacts are retired from this
|
||||
path (corpus dormancy, Task 6-1-04).
|
||||
"""
|
||||
|
||||
from pydantic import BaseModel, Field
|
||||
|
||||
from ..corpus.artifacts import (
|
||||
ArtifactSubmission,
|
||||
AssessmentRubric,
|
||||
DefenseTranscript,
|
||||
render_rubric,
|
||||
render_transcript,
|
||||
)
|
||||
from ..corpus.learner_context import LearnerContext, get_learner_context
|
||||
from ..grading.store import GradeRecord
|
||||
from ..prompts.assessor import SYSTEM_PROMPT, render_context
|
||||
from .base import BaseAgent
|
||||
|
||||
|
||||
class CriterionScore(BaseModel):
|
||||
criterion_id: str
|
||||
name: str
|
||||
score: int = Field(ge=0, le=100)
|
||||
evidence: str
|
||||
class GradeCoaching(BaseModel):
|
||||
"""Rubric-anchored coaching rendered FROM the stored grade (not invented)."""
|
||||
|
||||
summary: str = Field(min_length=1)
|
||||
strengths: list[str] = Field(min_length=1, max_length=3)
|
||||
gaps: list[str] = Field(min_length=1, max_length=3)
|
||||
next_steps: list[str] = Field(min_length=1, max_length=3)
|
||||
|
||||
|
||||
class RubricScore(BaseModel):
|
||||
rubric_id: str
|
||||
artifact_id: str
|
||||
competency_id: str
|
||||
scores: list[CriterionScore]
|
||||
strengths: list[str] = Field(min_length=1, max_length=2)
|
||||
gaps: list[str] = Field(min_length=1, max_length=2)
|
||||
verdict: str # "mastered" | "developing" | "not_yet"
|
||||
|
||||
def weighted_total(self, rubric: AssessmentRubric) -> float:
|
||||
by_id = {c.criterion_id: c for c in rubric.criteria}
|
||||
total = 0.0
|
||||
for s in self.scores:
|
||||
total += s.score * by_id[s.criterion_id].weight
|
||||
return total
|
||||
|
||||
|
||||
RUBRIC_SCORE_SCHEMA_HINT = (
|
||||
'{"rubric_id": "<id>", "artifact_id": "<id>", "competency_id": "<id>", '
|
||||
'"scores": [{"criterion_id": "<id>", "name": "<name>", "score": <0-100>, '
|
||||
'"evidence": "<one sentence>"}], "strengths": ["<one sentence>"], '
|
||||
'"gaps": ["<one sentence>"], "verdict": "mastered"|"developing"|"not_yet"}'
|
||||
GRADE_COACHING_SCHEMA_HINT = (
|
||||
'{"summary": "<two sentences on the grade>", '
|
||||
'"strengths": ["<one sentence>"], "gaps": ["<one sentence>"], '
|
||||
'"next_steps": ["<one sentence>"]}'
|
||||
)
|
||||
|
||||
|
||||
@@ -58,35 +39,25 @@ class AssessorAgent(BaseAgent):
|
||||
ctx = learner_context or get_learner_context()
|
||||
return SYSTEM_PROMPT.format_map(render_context(ctx))
|
||||
|
||||
def build_evaluation_input(
|
||||
async def coach_grade(
|
||||
self,
|
||||
artifact: ArtifactSubmission,
|
||||
rubric: AssessmentRubric,
|
||||
transcript: DefenseTranscript | None,
|
||||
) -> str:
|
||||
parts = [
|
||||
f"ARTIFACT: {artifact.name} ({artifact.artifact_id})",
|
||||
f"Evidence excerpt: {artifact.evidence_excerpt}",
|
||||
"",
|
||||
render_rubric(rubric),
|
||||
]
|
||||
if transcript is not None:
|
||||
parts += ["", render_transcript(transcript)]
|
||||
return "\n".join(parts)
|
||||
|
||||
async def evaluate(
|
||||
self,
|
||||
artifact: ArtifactSubmission,
|
||||
rubric: AssessmentRubric,
|
||||
transcript: DefenseTranscript | None,
|
||||
grade: GradeRecord,
|
||||
learner_context: LearnerContext | None = None,
|
||||
) -> RubricScore:
|
||||
evaluation_input = self.build_evaluation_input(artifact, rubric, transcript)
|
||||
score: RubricScore = await self.structured_reply(
|
||||
) -> GradeCoaching:
|
||||
"""Render the STORED grade as coaching via the D-020 defense."""
|
||||
grade_json = {
|
||||
"verdict": grade.verdict,
|
||||
"scores": grade.scores,
|
||||
"digest": grade.digest,
|
||||
}
|
||||
coaching: GradeCoaching = await self.structured_reply(
|
||||
history=None,
|
||||
user_input=evaluation_input,
|
||||
user_input=(
|
||||
"The learner's process-trace grade (computed by the grading "
|
||||
f"engine) is:\n{grade_json!r}\nExplain it as coaching."
|
||||
),
|
||||
learner_context=learner_context,
|
||||
schema=RubricScore,
|
||||
schema_hint=RUBRIC_SCORE_SCHEMA_HINT,
|
||||
schema=GradeCoaching,
|
||||
schema_hint=GRADE_COACHING_SCHEMA_HINT,
|
||||
)
|
||||
return score
|
||||
return coaching
|
||||
|
||||
@@ -0,0 +1,104 @@
|
||||
"""ExaminerAgent — the seventh agent: oral-defense examiner (REQ-3-006, A-109).
|
||||
|
||||
BOUNDARY DECISION (PERSONAS conflict rule, honored by construction): the
|
||||
examiner is a TEXT agent. It composes the LLM provider through BaseAgent and
|
||||
consumes defense transcript turns; it NEVER imports voice/ — STT/TTS belong
|
||||
to the API endpoints (they move audio bytes; the agent moves question text).
|
||||
Integrity signals (long pauses, off-scope cadence) are computed by the
|
||||
endpoint layer from turn metadata (latency_ms etc.), not by the agent.
|
||||
|
||||
Digest discipline (D-028 mirror): questions are grounded in the compact
|
||||
TraceDigest + variant statement — never the raw trace, never learner ids.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from pydantic import BaseModel, ConfigDict, Field
|
||||
|
||||
from ..grading.features import TraceDigest
|
||||
from ..llm.types import Message
|
||||
from ..prompts.examiner import SYSTEM_PROMPT, VERDICT_SCHEMA_HINT, render_digest_context
|
||||
from .base import BaseAgent
|
||||
|
||||
|
||||
class DefenseVerdict(BaseModel):
|
||||
"""D-20-validated final defense verdict (structured mode)."""
|
||||
|
||||
model_config = ConfigDict(extra="forbid")
|
||||
|
||||
verdict: str = Field(pattern="^(mastered|developing|not_yet)$")
|
||||
understanding: str = Field(min_length=1)
|
||||
process_justification: str = Field(min_length=1)
|
||||
communication: str = Field(min_length=1)
|
||||
strengths: list[str] = Field(min_length=1, max_length=2)
|
||||
gaps: list[str] = Field(min_length=1, max_length=2)
|
||||
|
||||
|
||||
class ExaminerAgent(BaseAgent):
|
||||
"""Conducts the oral defense: next_question + final_verdict."""
|
||||
|
||||
name = "examiner"
|
||||
|
||||
def system_prompt(self, learner_context=None) -> str: # noqa: ANN001
|
||||
"""Examiner is context-free (digest-anonymous, D-028 mirror)."""
|
||||
return SYSTEM_PROMPT
|
||||
|
||||
def build_defense_messages(
|
||||
self,
|
||||
trace_digest: TraceDigest | None = None,
|
||||
variant_statement: str | None = None,
|
||||
history: list[Message] | None = None,
|
||||
) -> list[Message]:
|
||||
"""System + grounding + defense transcript (no learner id — D-028)."""
|
||||
digest_json = (
|
||||
trace_digest.model_dump_json() if trace_digest is not None else "{}"
|
||||
)
|
||||
messages: list[Message] = [
|
||||
Message(role="system", content=SYSTEM_PROMPT),
|
||||
Message(role="user", content=render_digest_context(digest_json, variant_statement)),
|
||||
Message(
|
||||
role="assistant",
|
||||
content="Understood. I will question the learner about this build session.",
|
||||
),
|
||||
]
|
||||
for m in history or []:
|
||||
messages.append(m)
|
||||
return messages
|
||||
|
||||
async def next_question(
|
||||
self,
|
||||
history: list[Message],
|
||||
trace_digest: TraceDigest | None = None,
|
||||
variant_statement: str | None = None,
|
||||
) -> str:
|
||||
"""One examiner question (streamed over SSE by the endpoints)."""
|
||||
messages = self.build_defense_messages(trace_digest, variant_statement, history)
|
||||
messages.append(
|
||||
Message(role="user", content="Ask the learner your next question now.")
|
||||
)
|
||||
reply = await self.provider.chat(messages, model=self.settings.model)
|
||||
return reply
|
||||
|
||||
async def final_verdict(
|
||||
self,
|
||||
history: list[Message],
|
||||
trace_digest: TraceDigest | None = None,
|
||||
variant_statement: str | None = None,
|
||||
) -> DefenseVerdict:
|
||||
"""Structured verdict via the D-020 4-layer defense."""
|
||||
from .structured import structured_completion # module-direct (G-4)
|
||||
|
||||
messages = self.build_defense_messages(trace_digest, variant_statement, history)
|
||||
messages.append(
|
||||
Message(
|
||||
role="user",
|
||||
content="The defense is finished. Return the final verdict JSON now.",
|
||||
)
|
||||
)
|
||||
return await structured_completion(
|
||||
self.provider,
|
||||
messages,
|
||||
model=self.settings.model,
|
||||
schema=DefenseVerdict,
|
||||
schema_hint=VERDICT_SCHEMA_HINT,
|
||||
)
|
||||
@@ -1,17 +1,18 @@
|
||||
"""LabAgent — in-flow feedback over simulated sandbox telemetry (REQ-2-007).
|
||||
"""LabAgent — in-flow feedback over LIVE sandbox telemetry (REQ-3-007).
|
||||
|
||||
Scenario-driven: consumes a LabTelemetryScenario from the corpus, renders
|
||||
the event timeline into the conversation, streams concrete feedback.
|
||||
No session chat — each request is one scenario read.
|
||||
v0.3 re-grounding: consumes a TraceDigest computed from the learner's real
|
||||
trace (grading/features.compute_digest over TraceStore events) — the v0.2
|
||||
corpus scenarios are retired from this path (corpus dormancy, Task 6-1-04).
|
||||
No session chat — each request is one live-trace read.
|
||||
"""
|
||||
|
||||
from collections.abc import AsyncIterator
|
||||
|
||||
from ..config import Settings
|
||||
from ..corpus.learner_context import LearnerContext, get_learner_context
|
||||
from ..corpus.telemetry import LabTelemetryScenario, summarize_scenario
|
||||
from ..grading.features import TraceDigest
|
||||
from ..llm.base import LLMProvider
|
||||
from ..prompts.lab import SYSTEM_PROMPT, render_context
|
||||
from ..prompts.lab import SYSTEM_PROMPT, render_context, render_digest_timeline
|
||||
from .base import BaseAgent
|
||||
|
||||
|
||||
@@ -27,10 +28,11 @@ class LabAgent(BaseAgent):
|
||||
|
||||
async def stream_feedback(
|
||||
self,
|
||||
scenario: LabTelemetryScenario,
|
||||
digest: TraceDigest | None,
|
||||
learner_context: LearnerContext | None = None,
|
||||
) -> AsyncIterator[str]:
|
||||
timeline = summarize_scenario(scenario)
|
||||
"""Feedback grounded in the learner's live trace digest."""
|
||||
timeline = render_digest_timeline(digest)
|
||||
async for token in self.stream_reply(
|
||||
history=None, user_input=timeline, learner_context=learner_context
|
||||
):
|
||||
|
||||
@@ -1,33 +1,41 @@
|
||||
"""ProctorAgent — integrity signals + coaching interventions (REQ-2-009).
|
||||
"""ProctorAgent — integrity signals + coaching over REAL inputs (REQ-3-007).
|
||||
|
||||
Consumes a ProctorScenario from the corpus, returns pydantic-validated
|
||||
signal classifications via structured_reply (4-layer defense).
|
||||
Mock engine inputs; real identity/attention signals are v0.3+.
|
||||
v0.3 re-grounding: consumes the learner's live trace digest (idle gaps,
|
||||
command cadence), the DefenseStore integrity signals (long pauses from the
|
||||
oral defense), and the variant seed cross-check — NOT v0.2 corpus
|
||||
scenarios. The proctor COACHES: it classifies signals supportively and
|
||||
recommends one intervention; it never punishes and never accuses.
|
||||
|
||||
Integrity inputs (computed server-side, passed in by the API layer):
|
||||
- trace digest: idle_gap_count/total, command_categories histogram,
|
||||
error/fix cycles, huge-burst indicators (edit_count vs test runs)
|
||||
- defense signals: long_pauses list from the finished defense (A-109)
|
||||
- variant: seed + params when the task is variant-derived (off-template
|
||||
work is a cross-check input, not an accusation)
|
||||
"""
|
||||
|
||||
from pydantic import BaseModel, Field
|
||||
|
||||
from ..corpus.learner_context import LearnerContext, get_learner_context
|
||||
from ..corpus.telemetry import ProctorScenario, summarize_proctor_scenario
|
||||
from ..grading.features import TraceDigest
|
||||
from ..prompts.proctor import SYSTEM_PROMPT, render_context
|
||||
from .base import BaseAgent
|
||||
|
||||
|
||||
class IntegritySignal(BaseModel):
|
||||
signal_type: str # e.g. "context_switch" | "idle_gap" | "large_paste"
|
||||
signal_type: str # "idle_gap" | "long_pause" | "burst_edit" | "off_template"
|
||||
severity: str # "low" | "medium" | "high"
|
||||
note: str
|
||||
|
||||
|
||||
class ProctorAssessment(BaseModel):
|
||||
scenario_id: str
|
||||
signals: list[IntegritySignal] = Field(min_length=0)
|
||||
intervention: str # ONE supportive coaching recommendation
|
||||
summary: str
|
||||
|
||||
|
||||
PROCTOR_ASSESSMENT_SCHEMA_HINT = (
|
||||
'{"scenario_id": "<id>", "signals": [{"signal_type": "<type>", '
|
||||
'{"signals": [{"signal_type": "<type>", '
|
||||
'"severity": "low"|"medium"|"high", "note": "<one sentence>"}], '
|
||||
'"intervention": "<one supportive recommendation>", '
|
||||
'"summary": "<one sentence>"}'
|
||||
@@ -43,13 +51,27 @@ class ProctorAgent(BaseAgent):
|
||||
|
||||
async def assess(
|
||||
self,
|
||||
scenario: ProctorScenario,
|
||||
digest: TraceDigest | None,
|
||||
defense_signals: dict | None = None,
|
||||
variant_context: dict | None = None,
|
||||
learner_context: LearnerContext | None = None,
|
||||
) -> ProctorAssessment:
|
||||
timeline = summarize_proctor_scenario(scenario)
|
||||
"""Classify REAL integrity inputs into supportive signals + coaching."""
|
||||
parts: list[str] = []
|
||||
if digest is not None:
|
||||
parts.append(f"Build-session digest:\n{digest.model_dump_json()}")
|
||||
else:
|
||||
parts.append("No build telemetry recorded for this task yet.")
|
||||
if defense_signals:
|
||||
parts.append(f"Oral-defense integrity signals:\n{defense_signals}")
|
||||
if variant_context:
|
||||
parts.append(f"Variant audit context (seed + params):\n{variant_context}")
|
||||
assessment: ProctorAssessment = await self.structured_reply(
|
||||
history=None,
|
||||
user_input=timeline,
|
||||
user_input=(
|
||||
"Assess this learner's integrity signals supportively.\n\n"
|
||||
+ "\n\n".join(parts)
|
||||
),
|
||||
learner_context=learner_context,
|
||||
schema=ProctorAssessment,
|
||||
schema_hint=PROCTOR_ASSESSMENT_SCHEMA_HINT,
|
||||
|
||||
@@ -14,13 +14,14 @@ AgentFactory = Callable[[LLMProvider, Settings], BaseAgent]
|
||||
|
||||
|
||||
def register_builtin_agents(registry: "AgentRegistry") -> None:
|
||||
"""Central registration of all six shipped tutor agents (G-4: one pattern).
|
||||
"""Central registration of all seven shipped agents (G-4: one pattern).
|
||||
|
||||
coach, tutor, lab, assessor, proctor, mentor. New agents register here
|
||||
in their landing phase.
|
||||
coach, tutor, lab, assessor, proctor, mentor, examiner (Phase 5).
|
||||
New agents register here in their landing phase.
|
||||
"""
|
||||
from .assessor import AssessorAgent
|
||||
from .coach import CoachAgent
|
||||
from .examiner import ExaminerAgent
|
||||
from .lab import LabAgent
|
||||
from .mentor import MentorAgent
|
||||
from .proctor import ProctorAgent
|
||||
@@ -38,6 +39,9 @@ def register_builtin_agents(registry: "AgentRegistry") -> None:
|
||||
registry.register(
|
||||
"mentor", lambda provider, settings: MentorAgent(provider, settings)
|
||||
)
|
||||
registry.register(
|
||||
"examiner", lambda provider, settings: ExaminerAgent(provider, settings)
|
||||
)
|
||||
|
||||
|
||||
class UnknownAgentError(KeyError):
|
||||
|
||||
@@ -5,10 +5,13 @@ Boundary rule: api/ composes agents/ and llm/; they never import api/.
|
||||
|
||||
from .assessment import router as assessment_router
|
||||
from .chat import router as chat_router
|
||||
from .defense import router as defense_router
|
||||
from .lab import router as lab_router
|
||||
from .mentor import router as mentor_router
|
||||
from .proctor import router as proctor_router
|
||||
from .sandboxes import router as sandboxes_router
|
||||
from .telemetry import router as telemetry_router
|
||||
from .variants import router as variants_router
|
||||
|
||||
__all__ = [
|
||||
"assessment_router",
|
||||
@@ -17,4 +20,7 @@ __all__ = [
|
||||
"mentor_router",
|
||||
"proctor_router",
|
||||
"sandboxes_router",
|
||||
"telemetry_router",
|
||||
"variants_router",
|
||||
"defense_router",
|
||||
]
|
||||
|
||||
@@ -1,47 +1,190 @@
|
||||
"""POST /v1/assessment/evaluate — structured rubric scores (REQ-2-008).
|
||||
"""/v1/assessment — rubric evaluation + trace grading endpoints (REQ-2-008, REQ-3-004).
|
||||
|
||||
JSON response (not SSE): a pydantic-validated RubricScore. Unknown
|
||||
artifact → 404. The Assessor's structured output IS the payload.
|
||||
Two endpoint families share this router:
|
||||
|
||||
POST /v1/assessment/evaluate (v0.2, REQ-2-008) — corpus
|
||||
artifact evaluation through
|
||||
the Assessor agent.
|
||||
POST /v1/assessment/grade (v0.3, REQ-3-004) — grade a
|
||||
REAL process trace through
|
||||
the GradingEngine.
|
||||
GET /v1/assessment/grade/{learner_id}/{task_id} — stored latest grade.
|
||||
|
||||
Grading status-code mapping (the engine's outcomes are CONTRACT, not errors):
|
||||
|
||||
GradeRecord(verdict=GRADED) → 200 — rubric scores +
|
||||
verdict (in scores.verdict)
|
||||
+ digest summary.
|
||||
GradeRecord(UNGRADABLE_TRACE_INCOMPLETE) → 200 — the ungradable
|
||||
record IS a valid result:
|
||||
the trace cannot be graded,
|
||||
and the gate surfaces WHY
|
||||
(scores.missing_seqs +
|
||||
scores.integrity_flag).
|
||||
Persisted like any grade.
|
||||
GradeRecord(UNGRADABLE_EMPTY_TRACE) → 200 — no events stored for
|
||||
the pair. This covers BOTH
|
||||
a known pair whose trace
|
||||
ended up empty AND a task
|
||||
that never had a trace at
|
||||
all: the engine cannot
|
||||
distinguish them (zero
|
||||
stored events is zero
|
||||
events), and grading an
|
||||
absent trace genuinely has
|
||||
the empty-trace outcome —
|
||||
a 404 here would erase the
|
||||
durable gate record the
|
||||
engine persists for the
|
||||
pair. PLAN's 404 applies to
|
||||
GET of a never-graded pair.
|
||||
StructuredOutputError → 502 — the trace was
|
||||
gradable but the provider
|
||||
failed the D-020 budget;
|
||||
provider failure (bad
|
||||
gateway to the model), same
|
||||
mapping as evaluate.
|
||||
GET of an unknown (never-graded) pair → 404.
|
||||
|
||||
DI (D-027/D-032 house pattern): engine + store arrive via deps.get_grading_engine
|
||||
/ get_grade_store from app.state; this module owns all FastAPI wiring — the
|
||||
engine knows nothing of HTTP. UNGRADABLE_* bodies are rendered by the same
|
||||
GradeResponse model as GRADED ones (a gate record's `scores` holds the gate
|
||||
detail instead of rubric scores), so consumers read ONE shape.
|
||||
"""
|
||||
|
||||
from datetime import datetime
|
||||
from typing import Any
|
||||
|
||||
from fastapi import APIRouter, Depends, HTTPException
|
||||
from pydantic import BaseModel, Field
|
||||
|
||||
from ..agents.assessor import RubricScore
|
||||
from ..agents.registry import AgentRegistry
|
||||
from ..agents.structured import StructuredOutputError
|
||||
from ..config import Settings
|
||||
from ..corpus.artifacts import get_artifact_bundle, get_transcript_for_artifact
|
||||
from ..corpus.learner_context import get_learner_context
|
||||
from .deps import get_agent_registry, get_provider, get_settings
|
||||
from ..grading.engine import GradingEngine
|
||||
from ..grading.store import GradeRecord, GradeStore
|
||||
from .deps import (
|
||||
get_agent_registry,
|
||||
get_grade_store,
|
||||
get_grading_engine,
|
||||
get_provider,
|
||||
get_settings,
|
||||
)
|
||||
|
||||
router = APIRouter(prefix="/v1")
|
||||
|
||||
|
||||
class AssessmentRequest(BaseModel):
|
||||
artifact_id: str = Field(min_length=1)
|
||||
learner_id: str | None = None
|
||||
# --- v0.2 artifact evaluation (REQ-2-008) --------------------------------------
|
||||
|
||||
|
||||
@router.post("/assessment/evaluate", response_model=RubricScore)
|
||||
class EvaluateRequest(BaseModel):
|
||||
learner_id: str = Field(min_length=1)
|
||||
task_id: str = Field(min_length=1)
|
||||
|
||||
|
||||
@router.post("/assessment/evaluate")
|
||||
async def assessment_evaluate(
|
||||
body: AssessmentRequest,
|
||||
body: EvaluateRequest,
|
||||
registry: AgentRegistry = Depends(get_agent_registry),
|
||||
settings: Settings = Depends(get_settings),
|
||||
provider=Depends(get_provider),
|
||||
) -> RubricScore:
|
||||
bundle = get_artifact_bundle(body.artifact_id)
|
||||
if bundle is None:
|
||||
grade_store=Depends(get_grade_store),
|
||||
) -> dict:
|
||||
"""Assessor coaching rendered FROM the learner's stored grade (REQ-3-007).
|
||||
|
||||
The grading engine computes the scores (POST /assessment/grade); this
|
||||
endpoint explains them. No stored grade yet -> 404 (grade first).
|
||||
"""
|
||||
grade = grade_store.get(body.learner_id, body.task_id)
|
||||
if grade is None:
|
||||
raise HTTPException(
|
||||
status_code=404, detail=f"unknown artifact {body.artifact_id!r}"
|
||||
status_code=404,
|
||||
detail=f"no stored grade for {body.learner_id}/{body.task_id} - grade first",
|
||||
)
|
||||
artifact, rubric = bundle
|
||||
transcript = get_transcript_for_artifact(artifact.artifact_id)
|
||||
agent = registry.get(provider, settings, "assessor")
|
||||
learner_context = get_learner_context(body.learner_id)
|
||||
try:
|
||||
return await agent.evaluate(artifact, rubric, transcript, learner_context)
|
||||
coaching = await agent.coach_grade(grade, learner_context)
|
||||
except Exception as exc:
|
||||
raise HTTPException(
|
||||
status_code=502,
|
||||
detail=f"assessment evaluation failed: {exc}",
|
||||
) from exc
|
||||
return {
|
||||
"learner_id": grade.learner_id,
|
||||
"task_id": grade.task_id,
|
||||
"grade_verdict": grade.verdict,
|
||||
"grade_scores": grade.scores,
|
||||
"coaching": coaching.model_dump(),
|
||||
}
|
||||
|
||||
|
||||
# --- # --- v0.3 trace grading (REQ-3-004) ---------------------------------------------
|
||||
|
||||
|
||||
class GradeRequest(BaseModel):
|
||||
learner_id: str = Field(min_length=1)
|
||||
task_id: str = Field(min_length=1)
|
||||
|
||||
|
||||
class GradeResponse(BaseModel):
|
||||
"""GradeRecord over HTTP — one shape for GRADED and UNGRADABLE_* alike.
|
||||
|
||||
`scores` holds the validated rubric (criteria 0-4, strengths, gaps,
|
||||
rubric verdict) for a GRADED record, or the gate detail
|
||||
({integrity_flag, missing_seqs}) for an UNGRADABLE_* record — never both.
|
||||
`digest` is the compact trace summary that fed the rubric prompt (empty
|
||||
for gate records: nothing was graded).
|
||||
"""
|
||||
|
||||
learner_id: str
|
||||
task_id: str
|
||||
variant_seed: str | None
|
||||
digest: dict[str, Any]
|
||||
scores: dict[str, Any]
|
||||
verdict: str
|
||||
model: str
|
||||
created_at: datetime
|
||||
|
||||
|
||||
def _grade_response(record: GradeRecord) -> GradeResponse:
|
||||
return GradeResponse.model_validate(record, from_attributes=True)
|
||||
|
||||
|
||||
@router.post("/assessment/grade", response_model=GradeResponse)
|
||||
async def assessment_grade(
|
||||
body: GradeRequest,
|
||||
engine: GradingEngine = Depends(get_grading_engine),
|
||||
) -> GradeResponse:
|
||||
"""Run the grading engine for one (learner_id, task_id) trace.
|
||||
|
||||
Gate outcomes (UNGRADABLE_*) are 200s — they are first-class results the
|
||||
engine persists, not failures. Only a provider that exhausts the D-020
|
||||
budget turns into a 502; nothing is persisted on that path.
|
||||
"""
|
||||
try:
|
||||
record = await engine.grade(body.learner_id, body.task_id)
|
||||
except StructuredOutputError as exc:
|
||||
raise HTTPException(
|
||||
status_code=502,
|
||||
detail=f"grading failed: {exc}",
|
||||
) from exc
|
||||
return _grade_response(record)
|
||||
|
||||
|
||||
@router.get("/assessment/grade/{learner_id}/{task_id}", response_model=GradeResponse)
|
||||
async def assessment_get_grade(
|
||||
learner_id: str,
|
||||
task_id: str,
|
||||
store: GradeStore = Depends(get_grade_store),
|
||||
) -> GradeResponse:
|
||||
"""Latest stored grade for the pair; 404 when none was ever stored."""
|
||||
record = store.get(learner_id, task_id)
|
||||
if record is None:
|
||||
raise HTTPException(
|
||||
status_code=404,
|
||||
detail=f"no stored grade for {learner_id!r}/{task_id!r}",
|
||||
)
|
||||
return _grade_response(record)
|
||||
|
||||
@@ -0,0 +1,332 @@
|
||||
"""Oral-defense endpoints (Task 5-3-01, REQ-3-006, A-109).
|
||||
|
||||
Full defense loop over HTTP with mock-first voice (D-030) and the seventh
|
||||
Examiner agent (SSE question streaming happens through the chat pipeline;
|
||||
these endpoints are the session orchestration + transcript persistence):
|
||||
|
||||
POST /v1/defense/start {learner_id, task_id}
|
||||
POST /v1/defense/{id}/answer {text} | multipart audio (STT)
|
||||
GET /v1/defense/{id}/audio/{turn_id} TTS bytes (streaming)
|
||||
POST /v1/defense/{id}/finish verdict + integrity signals
|
||||
GET /v1/defense/{id} transcript + signals
|
||||
|
||||
Integrity signals (A-109) are computed server-side from turn metadata:
|
||||
long pauses = learner turns whose latency_ms exceeds PAUSE_THRESHOLD_MS.
|
||||
The defense does NOT gate on trace completeness (the grader does, G-4);
|
||||
an incomplete trace is surfaced as `trace_complete: false` so the UI can
|
||||
disclose it before the learner defends.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import time
|
||||
from datetime import UTC, datetime
|
||||
|
||||
from fastapi import APIRouter, Depends, File, Form, HTTPException, UploadFile
|
||||
from fastapi.responses import StreamingResponse
|
||||
from pydantic import BaseModel, Field
|
||||
|
||||
from ..agents.examiner import ExaminerAgent
|
||||
from ..grading.features import TraceDigest, compute_digest
|
||||
from ..llm.types import Message
|
||||
from ..voice.base import VoiceDescriptor
|
||||
from ..voice.browser import BROWSER_FALLBACK_DESCRIPTOR
|
||||
from ..voice.defense_store import DefenseRecord, DefenseStore, DefenseTurn
|
||||
from .deps import (
|
||||
get_examiner,
|
||||
get_settings,
|
||||
get_trace_store,
|
||||
get_variant_store,
|
||||
get_voice_provider,
|
||||
get_voice_store,
|
||||
)
|
||||
|
||||
router = APIRouter(prefix="/v1/defense", tags=["defense"])
|
||||
|
||||
#: A-109: learner turns slower than this are flagged as long pauses (ms).
|
||||
PAUSE_THRESHOLD_MS = 15_000
|
||||
|
||||
_ROLE_EXAMINER = "examiner"
|
||||
_ROLE_LEARNER = "learner"
|
||||
|
||||
|
||||
class StartRequest(BaseModel):
|
||||
learner_id: str = Field(min_length=1)
|
||||
task_id: str = Field(min_length=1)
|
||||
|
||||
|
||||
class StartResponse(BaseModel):
|
||||
defense_id: str
|
||||
voice_descriptor: dict
|
||||
trace_complete: bool
|
||||
first_question: str
|
||||
|
||||
|
||||
class AnswerResponse(BaseModel):
|
||||
question: str
|
||||
turn_latency: dict[str, int | None]
|
||||
|
||||
|
||||
class FinishResponse(BaseModel):
|
||||
verdict: dict
|
||||
integrity_signals: dict
|
||||
|
||||
|
||||
async def _digest_for_task(
|
||||
trace_store, learner_id: str, task_id: str
|
||||
) -> tuple[TraceDigest | None, bool]:
|
||||
"""Digest of the learner's trace for this task + completeness flag."""
|
||||
if not trace_store.list_tasks(learner_id) or task_id not in trace_store.list_tasks(
|
||||
learner_id
|
||||
):
|
||||
return None, True # no trace at all is "complete" for defense purposes
|
||||
trace = trace_store.get_trace(learner_id, task_id)
|
||||
gaps = trace_store.gaps(learner_id, task_id)
|
||||
return (compute_digest(trace) if trace else None), (len(gaps) == 0)
|
||||
|
||||
|
||||
def _voice_descriptor(settings) -> VoiceDescriptor:
|
||||
"""The capability descriptor for the configured voice mode (D-030).
|
||||
|
||||
Must-Have #6: browser mode returns BROWSER_FALLBACK_DESCRIPTOR so the
|
||||
web client selects native SpeechRecognition/speechSynthesis; mock mode
|
||||
returns the mock descriptor. (A v0.4 server provider would return
|
||||
mode="server" — the protocol seam.)
|
||||
"""
|
||||
if (settings.voice_provider or "mock").strip().lower() == "browser":
|
||||
return BROWSER_FALLBACK_DESCRIPTOR
|
||||
return VoiceDescriptor(
|
||||
mode="mock", sr_available=True, tts_available=True, hint=""
|
||||
)
|
||||
|
||||
|
||||
@router.post("/start", response_model=StartResponse)
|
||||
async def start_defense(
|
||||
body: StartRequest,
|
||||
examiner: ExaminerAgent = Depends(get_examiner),
|
||||
voice_store: DefenseStore = Depends(get_voice_store),
|
||||
voice_provider=Depends(get_voice_provider),
|
||||
trace_store=Depends(get_trace_store),
|
||||
variant_store=Depends(get_variant_store),
|
||||
settings=Depends(get_settings),
|
||||
) -> StartResponse:
|
||||
record = voice_store.start(
|
||||
DefenseRecord(
|
||||
id=f"dfn-{int(time.time() * 1000):x}-{body.learner_id[:8]}",
|
||||
learner_id=body.learner_id,
|
||||
task_id=body.task_id,
|
||||
status="in_progress",
|
||||
created_at=datetime.now(UTC),
|
||||
)
|
||||
)
|
||||
digest, trace_complete = await _digest_for_task(trace_store, body.learner_id, body.task_id)
|
||||
variant = variant_store.get_by_task(body.task_id)
|
||||
statement = variant.statement if variant is not None else None
|
||||
|
||||
started = time.perf_counter()
|
||||
question = await examiner.next_question(
|
||||
history=[], trace_digest=digest, variant_statement=statement
|
||||
)
|
||||
llm_ms = int((time.perf_counter() - started) * 1000)
|
||||
voice_store.append_turn(
|
||||
record.id,
|
||||
DefenseTurn(
|
||||
defense_id=record.id,
|
||||
seq=0,
|
||||
role=_ROLE_EXAMINER,
|
||||
text=question,
|
||||
ts=datetime.now(UTC),
|
||||
latency_ms=llm_ms,
|
||||
created_at=datetime.now(UTC),
|
||||
),
|
||||
)
|
||||
descriptor = getattr(voice_provider, "descriptor", None) or _voice_descriptor(
|
||||
settings
|
||||
)
|
||||
return StartResponse(
|
||||
defense_id=record.id,
|
||||
voice_descriptor=descriptor.model_dump(),
|
||||
trace_complete=trace_complete,
|
||||
first_question=question,
|
||||
)
|
||||
|
||||
|
||||
@router.post("/{defense_id}/answer", response_model=AnswerResponse)
|
||||
async def answer_defense(
|
||||
defense_id: str,
|
||||
text: str | None = Form(default=None),
|
||||
audio: UploadFile | None = File(default=None),
|
||||
voice_store: DefenseStore = Depends(get_voice_store),
|
||||
voice_provider=Depends(get_voice_provider),
|
||||
examiner: ExaminerAgent = Depends(get_examiner),
|
||||
trace_store=Depends(get_trace_store),
|
||||
variant_store=Depends(get_variant_store),
|
||||
) -> AnswerResponse:
|
||||
record = voice_store.get(defense_id)
|
||||
if record is None:
|
||||
raise HTTPException(status_code=404, detail=f"no defense {defense_id!r}")
|
||||
if record.status == "finished":
|
||||
# The store owns the finished transition but does NOT police turn
|
||||
# sequencing (defense_store.py: "turns after finalize are a sequencing
|
||||
# bug for the endpoints to prevent") — this is the endpoint half of
|
||||
# that contract: a sealed transcript is append-only-no-more.
|
||||
raise HTTPException(
|
||||
status_code=409, detail="defense is finished; start a new defense"
|
||||
)
|
||||
if text is None and audio is None:
|
||||
raise HTTPException(status_code=422, detail="provide {text} or audio")
|
||||
|
||||
# STT (typed fallback bypasses the voice provider entirely).
|
||||
stt_ms: int | None = None
|
||||
if audio is not None:
|
||||
stt_started = time.perf_counter()
|
||||
raw = await audio.read()
|
||||
if not raw:
|
||||
# Empty upload is a client error (422), not a provider crash
|
||||
# (500): validate before the provider call so every provider —
|
||||
# mock today, the v0.4 real one — sees the same contract.
|
||||
raise HTTPException(status_code=422, detail="audio upload is empty")
|
||||
fmt = (audio.content_type or "audio/wav").split("/")[-1]
|
||||
segment = await voice_provider.transcribe(raw, fmt)
|
||||
stt_ms = int((time.perf_counter() - stt_started) * 1000)
|
||||
text = segment.text
|
||||
|
||||
turns = record.turns if hasattr(record, "turns") else []
|
||||
history = [
|
||||
Message(role="assistant" if t.role == _ROLE_EXAMINER else "user", content=t.text)
|
||||
for t in turns
|
||||
]
|
||||
next_seq = len(turns)
|
||||
|
||||
voice_store.append_turn(
|
||||
defense_id,
|
||||
DefenseTurn(
|
||||
defense_id=defense_id,
|
||||
seq=next_seq,
|
||||
role=_ROLE_LEARNER,
|
||||
text=text or "",
|
||||
ts=datetime.now(UTC),
|
||||
latency_ms=stt_ms,
|
||||
created_at=datetime.now(UTC),
|
||||
),
|
||||
)
|
||||
|
||||
digest, _ = await _digest_for_task(trace_store, record.learner_id, record.task_id)
|
||||
variant = variant_store.get_by_task(record.task_id)
|
||||
|
||||
llm_started = time.perf_counter()
|
||||
question = await examiner.next_question(
|
||||
history=history + [Message(role="user", content=text or "")],
|
||||
trace_digest=digest,
|
||||
variant_statement=variant.statement if variant is not None else None,
|
||||
)
|
||||
llm_ms = int((time.perf_counter() - llm_started) * 1000)
|
||||
|
||||
voice_store.append_turn(
|
||||
defense_id,
|
||||
DefenseTurn(
|
||||
defense_id=defense_id,
|
||||
seq=next_seq + 1,
|
||||
role=_ROLE_EXAMINER,
|
||||
text=question,
|
||||
ts=datetime.now(UTC),
|
||||
latency_ms=llm_ms,
|
||||
created_at=datetime.now(UTC),
|
||||
),
|
||||
)
|
||||
return AnswerResponse(
|
||||
question=question,
|
||||
turn_latency={"stt_ms": stt_ms, "llm_ms": llm_ms, "tts_ms": None},
|
||||
)
|
||||
|
||||
|
||||
@router.get("/{defense_id}/audio/{turn_id}")
|
||||
async def defense_audio(
|
||||
defense_id: str,
|
||||
turn_id: int,
|
||||
voice_store: DefenseStore = Depends(get_voice_store),
|
||||
voice_provider=Depends(get_voice_provider),
|
||||
):
|
||||
record = voice_store.get(defense_id)
|
||||
if record is None:
|
||||
raise HTTPException(status_code=404, detail=f"no defense {defense_id!r}")
|
||||
turn = next((t for t in record.turns if t.seq == turn_id), None)
|
||||
if turn is None or turn.role != _ROLE_EXAMINER:
|
||||
raise HTTPException(status_code=404, detail=f"no examiner turn {turn_id!r}")
|
||||
|
||||
async def stream():
|
||||
async for chunk in voice_provider.synthesize(turn.text):
|
||||
yield chunk
|
||||
|
||||
return StreamingResponse(stream(), media_type="audio/wav")
|
||||
|
||||
|
||||
@router.post("/{defense_id}/finish", response_model=FinishResponse)
|
||||
async def finish_defense(
|
||||
defense_id: str,
|
||||
voice_store: DefenseStore = Depends(get_voice_store),
|
||||
examiner: ExaminerAgent = Depends(get_examiner),
|
||||
trace_store=Depends(get_trace_store),
|
||||
variant_store=Depends(get_variant_store),
|
||||
) -> FinishResponse:
|
||||
record = voice_store.get(defense_id)
|
||||
if record is None:
|
||||
raise HTTPException(status_code=404, detail=f"no defense {defense_id!r}")
|
||||
|
||||
turns = record.turns if hasattr(record, "turns") else []
|
||||
history = [
|
||||
Message(role="assistant" if t.role == _ROLE_EXAMINER else "user", content=t.text)
|
||||
for t in turns
|
||||
]
|
||||
digest, _ = await _digest_for_task(trace_store, record.learner_id, record.task_id)
|
||||
variant = variant_store.get_by_task(record.task_id)
|
||||
verdict = await examiner.final_verdict(
|
||||
history=history,
|
||||
trace_digest=digest,
|
||||
variant_statement=variant.statement if variant is not None else None,
|
||||
)
|
||||
|
||||
signals: dict = {
|
||||
"long_pauses": [
|
||||
{"turn": t.seq, "latency_ms": t.latency_ms}
|
||||
for t in turns
|
||||
if t.role == _ROLE_LEARNER and (t.latency_ms or 0) > PAUSE_THRESHOLD_MS
|
||||
],
|
||||
"pause_threshold_ms": PAUSE_THRESHOLD_MS,
|
||||
# Must-Have #1: "verdict + transcript persisted" — the verdict is
|
||||
# stored INSIDE integrity_signals so GET /{id} after finish can
|
||||
# re-serve it (the finish response alone would lose it). Signals
|
||||
# are a JSON object dict (DefenseStore.finalize contract), so the
|
||||
# verdict nests under the "verdict" key alongside the A-109
|
||||
# markers the Proctor/Mentor feeds read.
|
||||
"verdict": verdict.model_dump(),
|
||||
}
|
||||
voice_store.finalize(defense_id, signals)
|
||||
return FinishResponse(verdict=verdict.model_dump(), integrity_signals=signals)
|
||||
|
||||
|
||||
@router.get("/{defense_id}")
|
||||
async def get_defense(
|
||||
defense_id: str,
|
||||
voice_store: DefenseStore = Depends(get_voice_store),
|
||||
):
|
||||
record = voice_store.get(defense_id)
|
||||
if record is None:
|
||||
raise HTTPException(status_code=404, detail=f"no defense {defense_id!r}")
|
||||
return {
|
||||
"defense_id": record.id,
|
||||
"learner_id": record.learner_id,
|
||||
"task_id": record.task_id,
|
||||
"status": record.status,
|
||||
"turns": [
|
||||
{
|
||||
"seq": t.seq,
|
||||
"role": t.role,
|
||||
"text": t.text,
|
||||
"ts": t.ts,
|
||||
"latency_ms": t.latency_ms,
|
||||
}
|
||||
for t in record.turns
|
||||
],
|
||||
"integrity_signals": record.integrity_signals or {},
|
||||
}
|
||||
@@ -2,12 +2,21 @@
|
||||
|
||||
from fastapi import Request
|
||||
|
||||
from ..agents.examiner import ExaminerAgent
|
||||
from ..agents.registry import AgentRegistry
|
||||
from ..agents.session import SessionStore
|
||||
from ..config import Settings
|
||||
from ..grading.engine import GradingEngine
|
||||
from ..grading.store import GradeStore
|
||||
from ..llm.base import LLMProvider
|
||||
from ..sandbox.manager import SandboxManager
|
||||
from ..sandbox.workdir import SandboxDir
|
||||
from ..telemetry.ingest import TraceIntegrityMap
|
||||
from ..telemetry.store import TraceStore
|
||||
from ..variants.generator import VariantGenerator
|
||||
from ..variants.store import VariantStore
|
||||
from ..voice.base import VoiceProvider
|
||||
from ..voice.defense_store import DefenseStore
|
||||
|
||||
|
||||
def get_settings(request: Request) -> Settings:
|
||||
@@ -33,3 +42,38 @@ def get_sandbox_manager(request: Request) -> SandboxManager:
|
||||
def get_sandbox_test_layout(request: Request) -> SandboxDir | None:
|
||||
"""Optional test seam (app.state.sandbox_test_layout); always None in prod."""
|
||||
return getattr(request.app.state, "sandbox_test_layout", None)
|
||||
|
||||
|
||||
def get_trace_store(request: Request) -> TraceStore:
|
||||
return request.app.state.trace_store
|
||||
|
||||
|
||||
def get_trace_integrity(request: Request) -> TraceIntegrityMap:
|
||||
return request.app.state.trace_integrity
|
||||
|
||||
|
||||
def get_grade_store(request: Request) -> GradeStore:
|
||||
return request.app.state.grade_store
|
||||
|
||||
|
||||
def get_grading_engine(request: Request) -> GradingEngine:
|
||||
return request.app.state.grading_engine
|
||||
|
||||
|
||||
def get_variant_generator(request: Request) -> VariantGenerator:
|
||||
return request.app.state.variant_generator
|
||||
|
||||
|
||||
def get_variant_store(request: Request) -> VariantStore:
|
||||
return request.app.state.variant_store
|
||||
|
||||
def get_voice_store(request: Request) -> DefenseStore:
|
||||
return request.app.state.defense_store
|
||||
|
||||
|
||||
def get_voice_provider(request: Request) -> VoiceProvider:
|
||||
return request.app.state.voice_provider
|
||||
|
||||
|
||||
def get_examiner(request: Request) -> ExaminerAgent:
|
||||
return request.app.state.examiner_agent
|
||||
|
||||
@@ -1,27 +1,36 @@
|
||||
"""POST /v1/lab/feedback — SSE stream of Lab in-flow feedback (REQ-2-007).
|
||||
"""POST /v1/lab/feedback — SSE stream of Lab in-flow feedback (REQ-3-007).
|
||||
|
||||
D-016 envelope with agent=lab. Unknown scenario → 404 before streaming.
|
||||
v0.3 re-grounding: LIVE trace digest. Request carries {learner_id, task_id};
|
||||
the digest is computed from the learner's real TraceStore events (D-028)
|
||||
and handed to the Lab agent. No corpus scenarios. Empty/unknown trace is NOT
|
||||
an error — Lab gets a "no telemetry yet" timeline and coaches the baseline.
|
||||
D-016 envelope with agent=lab.
|
||||
"""
|
||||
|
||||
import json
|
||||
from collections.abc import AsyncIterator
|
||||
|
||||
from fastapi import APIRouter, Depends, HTTPException
|
||||
from fastapi import APIRouter, Depends
|
||||
from pydantic import BaseModel, Field
|
||||
from sse_starlette.sse import EventSourceResponse
|
||||
|
||||
from ..agents.registry import AgentRegistry
|
||||
from ..config import Settings
|
||||
from ..corpus.learner_context import get_learner_context
|
||||
from ..corpus.telemetry import get_lab_scenario
|
||||
from .deps import get_agent_registry, get_provider, get_settings
|
||||
from ..grading.features import compute_digest
|
||||
from .deps import (
|
||||
get_agent_registry,
|
||||
get_provider,
|
||||
get_settings,
|
||||
get_trace_store,
|
||||
)
|
||||
|
||||
router = APIRouter(prefix="/v1")
|
||||
|
||||
|
||||
class LabFeedbackRequest(BaseModel):
|
||||
scenario_id: str = Field(min_length=1)
|
||||
learner_id: str | None = None
|
||||
learner_id: str = Field(min_length=1)
|
||||
task_id: str = Field(min_length=1)
|
||||
|
||||
|
||||
@router.post("/lab/feedback")
|
||||
@@ -30,25 +39,27 @@ async def lab_feedback(
|
||||
registry: AgentRegistry = Depends(get_agent_registry),
|
||||
settings: Settings = Depends(get_settings),
|
||||
provider=Depends(get_provider),
|
||||
trace_store=Depends(get_trace_store),
|
||||
) -> EventSourceResponse:
|
||||
scenario = get_lab_scenario(body.scenario_id)
|
||||
if scenario is None:
|
||||
raise HTTPException(
|
||||
status_code=404, detail=f"unknown scenario {body.scenario_id!r}"
|
||||
)
|
||||
agent = registry.get(provider, settings, "lab")
|
||||
learner_context = get_learner_context(body.learner_id)
|
||||
trace = (
|
||||
trace_store.get_trace(body.learner_id, body.task_id)
|
||||
if body.task_id in trace_store.list_tasks(body.learner_id)
|
||||
else []
|
||||
)
|
||||
digest = compute_digest(trace) if trace else None
|
||||
|
||||
async def event_stream() -> AsyncIterator[dict]:
|
||||
yield {"event": "message", "data": json.dumps({
|
||||
"type": "meta",
|
||||
"agent": "lab",
|
||||
"scenario_id": body.scenario_id,
|
||||
"task_id": body.task_id,
|
||||
"model": settings.model,
|
||||
})}
|
||||
first_byte = True
|
||||
try:
|
||||
async for token in agent.stream_feedback(scenario, learner_context):
|
||||
async for token in agent.stream_feedback(digest, learner_context):
|
||||
first_byte = False
|
||||
yield {"event": "message", "data": json.dumps({
|
||||
"type": "delta", "content": token
|
||||
|
||||
@@ -1,7 +1,8 @@
|
||||
"""POST /v1/proctor/signals — structured integrity signals (REQ-2-009).
|
||||
"""POST /v1/proctor/signals — integrity signals over REAL inputs (REQ-3-007).
|
||||
|
||||
JSON response (not SSE): a pydantic-validated ProctorAssessment.
|
||||
Unknown scenario → 404. Coaching-shaped interventions only.
|
||||
v0.3 re-grounding: live trace digest + DefenseStore long-pause signals +
|
||||
variant seed cross-check, no corpus scenarios. The proctor coaches:
|
||||
a pydantic-validated ProctorAssessment (JSON response, not SSE).
|
||||
"""
|
||||
|
||||
from fastapi import APIRouter, Depends, HTTPException
|
||||
@@ -11,15 +12,22 @@ from ..agents.proctor import ProctorAssessment
|
||||
from ..agents.registry import AgentRegistry
|
||||
from ..config import Settings
|
||||
from ..corpus.learner_context import get_learner_context
|
||||
from ..corpus.telemetry import get_proctor_scenario
|
||||
from .deps import get_agent_registry, get_provider, get_settings
|
||||
from ..grading.features import compute_digest
|
||||
from .deps import (
|
||||
get_agent_registry,
|
||||
get_provider,
|
||||
get_settings,
|
||||
get_trace_store,
|
||||
get_variant_store,
|
||||
get_voice_store,
|
||||
)
|
||||
|
||||
router = APIRouter(prefix="/v1")
|
||||
|
||||
|
||||
class ProctorRequest(BaseModel):
|
||||
scenario_id: str = Field(min_length=1)
|
||||
learner_id: str | None = None
|
||||
learner_id: str = Field(min_length=1)
|
||||
task_id: str = Field(min_length=1)
|
||||
|
||||
|
||||
@router.post("/proctor/signals", response_model=ProctorAssessment)
|
||||
@@ -28,16 +36,40 @@ async def proctor_signals(
|
||||
registry: AgentRegistry = Depends(get_agent_registry),
|
||||
settings: Settings = Depends(get_settings),
|
||||
provider=Depends(get_provider),
|
||||
trace_store=Depends(get_trace_store),
|
||||
variant_store=Depends(get_variant_store),
|
||||
voice_store=Depends(get_voice_store),
|
||||
) -> ProctorAssessment:
|
||||
scenario = get_proctor_scenario(body.scenario_id)
|
||||
if scenario is None:
|
||||
raise HTTPException(
|
||||
status_code=404, detail=f"unknown scenario {body.scenario_id!r}"
|
||||
)
|
||||
"""Real integrity inputs: live digest + defense signals + variant context."""
|
||||
agent = registry.get(provider, settings, "proctor")
|
||||
learner_context = get_learner_context(body.learner_id)
|
||||
|
||||
trace = (
|
||||
trace_store.get_trace(body.learner_id, body.task_id)
|
||||
if body.task_id in trace_store.list_tasks(body.learner_id)
|
||||
else []
|
||||
)
|
||||
digest = compute_digest(trace) if trace else None
|
||||
|
||||
defense_signals = None
|
||||
for record in voice_store.list_for_learner(body.learner_id):
|
||||
if record.task_id == body.task_id and record.status == "finished":
|
||||
defense_signals = record.integrity_signals or None
|
||||
break
|
||||
|
||||
variant = variant_store.get_by_task(body.task_id)
|
||||
variant_context = (
|
||||
{"template_id": variant.template_id, "seed": variant.seed, "params": variant.params}
|
||||
if variant is not None
|
||||
else None
|
||||
)
|
||||
try:
|
||||
return await agent.assess(scenario, learner_context)
|
||||
return await agent.assess(
|
||||
digest,
|
||||
defense_signals=defense_signals,
|
||||
variant_context=variant_context,
|
||||
learner_context=learner_context,
|
||||
)
|
||||
except Exception as exc:
|
||||
raise HTTPException(
|
||||
status_code=502, detail=f"proctor assessment failed: {exc}"
|
||||
|
||||
@@ -23,6 +23,7 @@ never part of the API contract.
|
||||
|
||||
import time
|
||||
from collections import deque
|
||||
from pathlib import Path
|
||||
|
||||
from fastapi import APIRouter, Depends, HTTPException, Response
|
||||
from pydantic import BaseModel, ConfigDict, Field
|
||||
@@ -50,6 +51,9 @@ _CREATE_TIMES: deque[float] = deque()
|
||||
|
||||
class SandboxCreateRequest(BaseModel):
|
||||
learner_id: str = Field(min_length=1)
|
||||
# Optional task key: when set, the sandbox is telemetry-wired (REQ-3-003)
|
||||
# — the in-sandbox capture agent streams workspace events to the ingest.
|
||||
task_id: str | None = None
|
||||
|
||||
|
||||
class SandboxResponse(BaseModel):
|
||||
@@ -142,7 +146,7 @@ async def create_sandbox(
|
||||
_check_per_learner_cap(await manager.list(), body.learner_id, settings)
|
||||
_check_global_create_rate(settings)
|
||||
try:
|
||||
info = await manager.create(body.learner_id)
|
||||
info = await manager.create(body.learner_id, task_id=body.task_id)
|
||||
except PoolFullError as exc:
|
||||
raise HTTPException(status_code=503, detail=str(exc)) from exc
|
||||
return _to_response(info)
|
||||
@@ -217,3 +221,133 @@ async def delete_sandbox(
|
||||
if test_layout is not None:
|
||||
result.headers["X-Workspace-Copy"] = str(test_layout.workspace)
|
||||
return result
|
||||
|
||||
|
||||
# -- workspace files + exec (Phase 6, REQ-3-008; CUT-2) -------------------------
|
||||
#
|
||||
# The build surface reads/writes/list workspace files and runs Run/Test
|
||||
# commands through the manager's backend. NO interactive shell relay (CUT-2:
|
||||
# keystroke-level stdin/stdout is v0.4) — each exec is a bounded command with
|
||||
# captured output. Paths are WORKSPACE-RELATIVE; traversal outside the
|
||||
# workspace is rejected (the workdir bind is the boundary, but the API adds
|
||||
# its own containment check — defense in depth).
|
||||
|
||||
|
||||
class FileWriteRequest(BaseModel):
|
||||
path: str = Field(min_length=1)
|
||||
content: str
|
||||
|
||||
|
||||
class ExecRequest(BaseModel):
|
||||
cmd: list[str] = Field(min_length=1)
|
||||
|
||||
|
||||
class ExecResponse(BaseModel):
|
||||
cmd: list[str]
|
||||
returncode: int
|
||||
stdout: str
|
||||
stderr: str
|
||||
duration_s: float
|
||||
|
||||
|
||||
async def _workspace_dir(manager: SandboxManager, sandbox_id: str):
|
||||
"""Resolve the sandbox workspace (tracked layout or shell layout)."""
|
||||
info = await manager.get(sandbox_id) # raises SandboxNotFoundError -> 404
|
||||
backend = manager._backend # noqa: SLF001 - API owns the composition seam
|
||||
tracked = getattr(backend, "_tracked", {}).get(sandbox_id)
|
||||
if tracked is not None:
|
||||
return tracked.workspace, info
|
||||
return info.workdir / "workspace", info
|
||||
|
||||
|
||||
def _safe_rel_path(raw: str) -> Path:
|
||||
"""Workspace-relative path; reject absolute/traversal paths."""
|
||||
candidate = Path(raw)
|
||||
if candidate.is_absolute() or ".." in candidate.parts:
|
||||
raise HTTPException(status_code=422, detail=f"invalid workspace path {raw!r}")
|
||||
return candidate
|
||||
|
||||
|
||||
def _resolve_in_workspace(workspace: Path, rel: Path) -> Path:
|
||||
"""Resolve `rel` under `workspace`, refusing symlink escapes (P7).
|
||||
|
||||
The lexical check in `_safe_rel_path` cannot see symlinks: an exec can
|
||||
plant `ln -s /etc target` in the workspace and a follow-up read/write
|
||||
would follow it OUT of the bind. Resolve with the workspace as the
|
||||
anchor (strict: a symlink chain escaping raises) and confirm the
|
||||
normalized target still sits inside the workspace — defense in depth
|
||||
for both read_file and write_file.
|
||||
"""
|
||||
try:
|
||||
target = (workspace / rel).resolve(strict=False)
|
||||
target.relative_to(workspace.resolve(strict=False))
|
||||
except ValueError:
|
||||
raise HTTPException(
|
||||
status_code=422, detail=f"path escapes the workspace: {rel.as_posix()!r}"
|
||||
) from None
|
||||
return target
|
||||
|
||||
|
||||
@router.get("/{sandbox_id}/files")
|
||||
async def list_files(
|
||||
sandbox_id: str,
|
||||
manager: SandboxManager = Depends(get_sandbox_manager),
|
||||
) -> dict:
|
||||
try:
|
||||
workspace, _ = await _workspace_dir(manager, sandbox_id)
|
||||
except SandboxNotFoundError:
|
||||
raise HTTPException(status_code=404, detail=f"no sandbox {sandbox_id!r}") from None
|
||||
return {"files": sorted(p.name for p in workspace.iterdir()) if workspace.is_dir() else []}
|
||||
|
||||
|
||||
@router.get("/{sandbox_id}/files/{path:path}")
|
||||
async def read_file(
|
||||
sandbox_id: str,
|
||||
path: str,
|
||||
manager: SandboxManager = Depends(get_sandbox_manager),
|
||||
) -> dict:
|
||||
try:
|
||||
workspace, _ = await _workspace_dir(manager, sandbox_id)
|
||||
except SandboxNotFoundError:
|
||||
raise HTTPException(status_code=404, detail=f"no sandbox {sandbox_id!r}") from None
|
||||
rel = _safe_rel_path(path)
|
||||
target = _resolve_in_workspace(workspace, rel)
|
||||
if not target.is_file():
|
||||
raise HTTPException(status_code=404, detail=f"no file {path!r}")
|
||||
return {"path": path, "content": target.read_text(errors="replace")}
|
||||
|
||||
|
||||
@router.put("/{sandbox_id}/files/{path:path}")
|
||||
async def write_file(
|
||||
sandbox_id: str,
|
||||
path: str,
|
||||
body: FileWriteRequest,
|
||||
manager: SandboxManager = Depends(get_sandbox_manager),
|
||||
) -> dict:
|
||||
try:
|
||||
workspace, _ = await _workspace_dir(manager, sandbox_id)
|
||||
except SandboxNotFoundError:
|
||||
raise HTTPException(status_code=404, detail=f"no sandbox {sandbox_id!r}") from None
|
||||
rel = _safe_rel_path(body.path)
|
||||
target = _resolve_in_workspace(workspace, rel)
|
||||
target.parent.mkdir(parents=True, exist_ok=True)
|
||||
target.write_text(body.content)
|
||||
return {"path": body.path, "written": True}
|
||||
|
||||
|
||||
@router.post("/{sandbox_id}/exec", response_model=ExecResponse)
|
||||
async def exec_command(
|
||||
sandbox_id: str,
|
||||
body: ExecRequest,
|
||||
manager: SandboxManager = Depends(get_sandbox_manager),
|
||||
) -> ExecResponse:
|
||||
try:
|
||||
await manager.get(sandbox_id)
|
||||
except SandboxNotFoundError:
|
||||
raise HTTPException(status_code=404, detail=f"no sandbox {sandbox_id!r}") from None
|
||||
backend = manager._backend # noqa: SLF001 - API owns the composition seam
|
||||
handle = manager._handles.get(sandbox_id) # noqa: SLF001
|
||||
if handle is None:
|
||||
raise HTTPException(status_code=404, detail=f"no live handle {sandbox_id!r}")
|
||||
result = await backend.exec(handle, body.cmd)
|
||||
return ExecResponse(**result.model_dump())
|
||||
|
||||
@@ -0,0 +1,175 @@
|
||||
"""/v1/telemetry — WS ingest + trace read endpoints (REQ-3-003, D-026, G-3).
|
||||
|
||||
The router composes the telemetry engine via DI: `telemetry/ingest.py` owns
|
||||
the WS protocol (frame contract + flood control + keepalive) and this module
|
||||
only wires `app.state.trace_store` / `app.state.trace_integrity` /
|
||||
`app.state.settings` into it, plus the two HTTP read faces:
|
||||
|
||||
WS /v1/telemetry/ingest?learner_id&task_id[&sandbox_id] (D-026)
|
||||
GET /v1/telemetry/traces/{learner_id}/{task_id} ordered trace; 404 unknown
|
||||
GET /v1/telemetry/gaps/{learner_id}/{task_id} missing seqs ; 404 unknown
|
||||
|
||||
The WS route is a thin DI shell: it validates the query-param identity and
|
||||
the Origin (browser pages are gated to the localhost dev origins — CORS
|
||||
middleware does not cover WS upgrades; the stdlib capture agent sends no
|
||||
Origin and is unaffected), pulls store/integrity/settings from `app.state`,
|
||||
and calls `telemetry_ingest_endpoint(...)` — the engine stays
|
||||
FastAPI-DI-free so it's testable without a router and the api/ layer owns
|
||||
all composition.
|
||||
|
||||
Unknown-trace contract: a trace is KNOWN when it has >=1 stored event OR
|
||||
carries an integrity flag — a flooded trace with zero stored rows still 200s
|
||||
so Proctor/grader can read WHY it's unusable (G-4 consumes
|
||||
`integrity_reason`). `TraceResponse.incomplete` / `.integrity_reason` mirror
|
||||
the map so HTTP consumers never touch process internals.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from fastapi import APIRouter, Depends, HTTPException, WebSocket
|
||||
from pydantic import BaseModel
|
||||
|
||||
from ..telemetry.ingest import (
|
||||
TraceIntegrityMap,
|
||||
telemetry_ingest_endpoint,
|
||||
)
|
||||
from ..telemetry.models import TelemetryEvent
|
||||
from ..telemetry.store import TraceStore
|
||||
from .deps import get_trace_integrity, get_trace_store
|
||||
|
||||
router = APIRouter(prefix="/v1/telemetry", tags=["telemetry"])
|
||||
|
||||
#: Browser Origins allowed to open the ingest socket (A-008 mirror, D-038).
|
||||
#: The stdlib capture agent sends NO Origin header (it is not a browser) and
|
||||
#: stays allowed; a malicious page loaded in the learner's browser would
|
||||
#: carry an Origin and must not be able to poison/flood the trace. CORS
|
||||
#: middleware does NOT cover WebSocket upgrades, so this gate is explicit.
|
||||
#: In network mode the configured CORS list governs (default '*' — any
|
||||
#: origin, since credentials are never used); an explicit list still rejects
|
||||
#: unlisted origins with 1008.
|
||||
_LOCAL_WS_ORIGINS = frozenset(
|
||||
{"http://localhost:3000", "http://127.0.0.1:3000", "http://localhost:8420"}
|
||||
)
|
||||
|
||||
|
||||
def _allowed_ws_origins(settings: object) -> frozenset[str]:
|
||||
configured = getattr(settings, "cors_origin_list", None)
|
||||
if configured is None:
|
||||
return _LOCAL_WS_ORIGINS
|
||||
if configured == ["*"]:
|
||||
return frozenset() # empty = wildcard = every Origin passes
|
||||
return frozenset(configured) | _LOCAL_WS_ORIGINS
|
||||
|
||||
|
||||
# --- WS ingest (D-026) ---------------------------------------------------------
|
||||
|
||||
|
||||
@router.websocket("/ingest")
|
||||
async def telemetry_ingest_ws(websocket: WebSocket) -> None:
|
||||
"""DI shell: resolve app.state services, then hand the socket to the engine.
|
||||
|
||||
The engine's session + flood logic is fully typed and testable without
|
||||
FastAPI; this shim is the only place the two layers meet.
|
||||
"""
|
||||
origin = (websocket.headers.get("origin") or "").strip()
|
||||
allowed = _allowed_ws_origins(getattr(websocket.app.state, "settings", None))
|
||||
if origin and allowed and origin not in allowed:
|
||||
# Same-origin dev pages (Next.js on :3000, the service itself on
|
||||
# :8420) pass; anything else is refused pre-accept. Non-browser
|
||||
# producers (the capture agent, tests) send no Origin and pass.
|
||||
# Wildcard (empty frozenset) passes every Origin in network mode.
|
||||
await websocket.close(
|
||||
code=1008, reason=f"origin {origin!r} not allowed for telemetry ingest"
|
||||
)
|
||||
return
|
||||
query = websocket.query_params
|
||||
learner_id = query.get("learner_id", "")
|
||||
task_id = query.get("task_id", "")
|
||||
if not learner_id or not task_id:
|
||||
# Reject BEFORE accept: closing the pre-accept handshake is the
|
||||
# cheapest denial and unambiguous for the stdlib capture agent.
|
||||
await websocket.close(code=1008, reason="learner_id/task_id query params required")
|
||||
return
|
||||
await telemetry_ingest_endpoint(
|
||||
websocket=websocket,
|
||||
learner_id=learner_id,
|
||||
task_id=task_id,
|
||||
sandbox_id=query.get("sandbox_id", ""),
|
||||
store=websocket.app.state.trace_store,
|
||||
integrity=websocket.app.state.trace_integrity,
|
||||
settings=websocket.app.state.settings,
|
||||
)
|
||||
|
||||
|
||||
# --- HTTP reads -----------------------------------------------------------------
|
||||
|
||||
|
||||
class TraceResponse(BaseModel):
|
||||
"""Ordered trace + integrity signal (G-4 reads incomplete/reason)."""
|
||||
|
||||
learner_id: str
|
||||
task_id: str
|
||||
events: list[TelemetryEvent]
|
||||
incomplete: bool
|
||||
integrity_reason: str | None
|
||||
|
||||
|
||||
class GapsResponse(BaseModel):
|
||||
"""Missing seqs + integrity signal."""
|
||||
|
||||
learner_id: str
|
||||
task_id: str
|
||||
gaps: list[int]
|
||||
incomplete: bool
|
||||
integrity_reason: str | None
|
||||
|
||||
|
||||
def _is_known_trace(
|
||||
store: TraceStore, integrity: TraceIntegrityMap, learner_id: str, task_id: str
|
||||
) -> bool:
|
||||
"""Known = has stored events OR carries an integrity flag (a flooded trace
|
||||
with zero rows must still be readable — Proctor needs the reason)."""
|
||||
return (
|
||||
store.latest_seq(learner_id, task_id) >= 0
|
||||
or integrity.is_incomplete(learner_id, task_id)
|
||||
)
|
||||
|
||||
|
||||
@router.get("/traces/{learner_id}/{task_id}", response_model=TraceResponse)
|
||||
async def get_trace(
|
||||
learner_id: str,
|
||||
task_id: str,
|
||||
store: TraceStore = Depends(get_trace_store),
|
||||
integrity: TraceIntegrityMap = Depends(get_trace_integrity),
|
||||
) -> TraceResponse:
|
||||
if not _is_known_trace(store, integrity, learner_id, task_id):
|
||||
raise HTTPException(
|
||||
status_code=404, detail=f"unknown trace {learner_id!r}/{task_id!r}"
|
||||
)
|
||||
return TraceResponse(
|
||||
learner_id=learner_id,
|
||||
task_id=task_id,
|
||||
events=store.get_trace(learner_id, task_id),
|
||||
incomplete=integrity.is_incomplete(learner_id, task_id),
|
||||
integrity_reason=integrity.reason(learner_id, task_id),
|
||||
)
|
||||
|
||||
|
||||
@router.get("/gaps/{learner_id}/{task_id}", response_model=GapsResponse)
|
||||
async def get_gaps(
|
||||
learner_id: str,
|
||||
task_id: str,
|
||||
store: TraceStore = Depends(get_trace_store),
|
||||
integrity: TraceIntegrityMap = Depends(get_trace_integrity),
|
||||
) -> GapsResponse:
|
||||
if not _is_known_trace(store, integrity, learner_id, task_id):
|
||||
raise HTTPException(
|
||||
status_code=404, detail=f"unknown trace {learner_id!r}/{task_id!r}"
|
||||
)
|
||||
return GapsResponse(
|
||||
learner_id=learner_id,
|
||||
task_id=task_id,
|
||||
gaps=store.gaps(learner_id, task_id),
|
||||
incomplete=integrity.is_incomplete(learner_id, task_id),
|
||||
integrity_reason=integrity.reason(learner_id, task_id),
|
||||
)
|
||||
@@ -0,0 +1,187 @@
|
||||
"""/v1/variants — seeded per-learner task variant endpoints (REQ-3-005, D-029).
|
||||
|
||||
Three faces over the variant engine (the generator + store stay FastAPI-free;
|
||||
this module owns all HTTP wiring — the D-027/D-032 house pattern):
|
||||
|
||||
POST /v1/variants {learner_id, template_id | competency_id}
|
||||
→ 200 the learner's variant — GENERATED on the first request,
|
||||
CACHED (store read, zero LLM calls) on every repeat: D-029
|
||||
reproducibility means one (learner_id, template_id) is ONE
|
||||
variant forever, so a regenerate is always a 200 of the SAME
|
||||
variant, never a second render.
|
||||
→ 404 unknown template_id, or competency_id with no bound template.
|
||||
→ 422 neither template_id nor competency_id given.
|
||||
GET /v1/variants/{task_id}
|
||||
→ 200 the stored variant owning the task key (the grading and
|
||||
telemetry join path); 404 when no variant was ever generated
|
||||
for the task.
|
||||
GET /v1/variants?learner_id=...
|
||||
→ 200 the learner's variants, chronological; [] when none.
|
||||
|
||||
Template resolution: an explicit `template_id` wins; without it the FIRST
|
||||
template bound to `competency_id` is used (`template_for_competency`,
|
||||
D-021 corpus alignment). The response carries `competency_id` resolved
|
||||
from the template library at read time — an enrichment, not a persisted
|
||||
column (the seed re-derives the whole variant, D-029) — so the learner
|
||||
surface can bind a variant to its competency without a library round-trip.
|
||||
|
||||
Distinctness (REQ-3-005): different learners on the same template draw
|
||||
different seeded params and receive distinct statements and task_ids;
|
||||
tests/api/test_variants.py asserts this end-to-end through the API.
|
||||
"""
|
||||
|
||||
from datetime import datetime
|
||||
from typing import Self
|
||||
|
||||
from fastapi import APIRouter, Depends, HTTPException
|
||||
from pydantic import BaseModel, Field, model_validator
|
||||
|
||||
from ..variants.generator import VariantGenerator
|
||||
from ..variants.store import VariantRecord, VariantStore
|
||||
from ..variants.templates import TaskTemplate, get_template, template_for_competency
|
||||
from .deps import get_variant_generator, get_variant_store
|
||||
|
||||
router = APIRouter(prefix="/v1/variants", tags=["variants"])
|
||||
|
||||
|
||||
# -- contracts ------------------------------------------------------------------
|
||||
|
||||
|
||||
class VariantGenerateRequest(BaseModel):
|
||||
"""One variant identity: an explicit `template_id`, or the first
|
||||
template bound to a `competency_id` (D-021). `template_id` wins when
|
||||
both are given (explicit identity beats derived); at least one is
|
||||
required — 422 otherwise.
|
||||
"""
|
||||
|
||||
learner_id: str = Field(min_length=1)
|
||||
template_id: str | None = None
|
||||
competency_id: str | None = None
|
||||
|
||||
@model_validator(mode="after")
|
||||
def _require_template_or_competency(self) -> Self:
|
||||
if self.template_id is None and self.competency_id is None:
|
||||
raise ValueError("template_id or competency_id is required")
|
||||
return self
|
||||
|
||||
|
||||
class VariantResponse(BaseModel):
|
||||
"""VariantRecord over HTTP, plus the `competency_id` enrichment.
|
||||
|
||||
Every field except `competency_id` mirrors `VariantRecord` exactly
|
||||
(snake_case; `created_at` is an ISO 8601 UTC datetime) — the wire shape
|
||||
typed as `TaskVariant` in packages/types/variants.ts.
|
||||
"""
|
||||
|
||||
learner_id: str
|
||||
task_id: str
|
||||
template_id: str
|
||||
competency_id: str
|
||||
seed: str
|
||||
params: dict[str, str | int]
|
||||
statement: str
|
||||
starter_files: dict[str, str]
|
||||
created_at: datetime
|
||||
|
||||
|
||||
class VariantListResponse(BaseModel):
|
||||
variants: list[VariantResponse]
|
||||
|
||||
|
||||
# -- resolution + rendering -----------------------------------------------------
|
||||
|
||||
|
||||
def _resolve_template(body: VariantGenerateRequest) -> TaskTemplate:
|
||||
"""Template for the request: the explicit id, else the first template
|
||||
bound to the competency; 404 when neither resolves."""
|
||||
if body.template_id is not None:
|
||||
template = get_template(body.template_id)
|
||||
if template is None:
|
||||
raise HTTPException(
|
||||
status_code=404,
|
||||
detail=f"no task template with id {body.template_id!r}",
|
||||
)
|
||||
return template
|
||||
# The request validator guarantees the disjunction, so reaching here
|
||||
# means a competency_id was given (never None).
|
||||
assert body.competency_id is not None
|
||||
templates = template_for_competency(body.competency_id)
|
||||
if not templates:
|
||||
raise HTTPException(
|
||||
status_code=404,
|
||||
detail=f"no task template for competency {body.competency_id!r}",
|
||||
)
|
||||
return templates[0]
|
||||
|
||||
|
||||
def _competency_for(template_id: str) -> str:
|
||||
"""competency_id enrichment for stored records (read paths)."""
|
||||
template = get_template(template_id)
|
||||
if template is None:
|
||||
# Integrity guard: a stored variant referencing a template that is
|
||||
# no longer in the library cannot be enriched; fail loudly rather
|
||||
# than fabricate a competency binding.
|
||||
raise HTTPException(
|
||||
status_code=500,
|
||||
detail=f"stored variant references unknown template {template_id!r}",
|
||||
)
|
||||
return template.competency_id
|
||||
|
||||
|
||||
def _to_response(record: VariantRecord, competency_id: str) -> VariantResponse:
|
||||
return VariantResponse(
|
||||
learner_id=record.learner_id,
|
||||
task_id=record.task_id,
|
||||
template_id=record.template_id,
|
||||
competency_id=competency_id,
|
||||
seed=record.seed,
|
||||
params=dict(record.params),
|
||||
statement=record.statement,
|
||||
starter_files=dict(record.starter_files),
|
||||
created_at=record.created_at,
|
||||
)
|
||||
|
||||
|
||||
# -- endpoints ------------------------------------------------------------------
|
||||
|
||||
|
||||
@router.post("", response_model=VariantResponse)
|
||||
async def generate_variant(
|
||||
body: VariantGenerateRequest,
|
||||
generator: VariantGenerator = Depends(get_variant_generator),
|
||||
) -> VariantResponse:
|
||||
"""The learner's variant for the resolved template — generated on the
|
||||
first request, cached (no LLM call) on every repeat: D-029 makes a
|
||||
regenerate a 200 of the SAME stored variant.
|
||||
"""
|
||||
template = _resolve_template(body)
|
||||
record = await generator.generate(body.learner_id, template.id)
|
||||
return _to_response(record, competency_id=template.competency_id)
|
||||
|
||||
|
||||
@router.get("", response_model=VariantListResponse)
|
||||
async def list_variants(
|
||||
learner_id: str,
|
||||
store: VariantStore = Depends(get_variant_store),
|
||||
) -> VariantListResponse:
|
||||
"""All stored variants for the learner, chronological; [] when none."""
|
||||
variants = [
|
||||
_to_response(record, competency_id=_competency_for(record.template_id))
|
||||
for record in store.list_for_learner(learner_id)
|
||||
]
|
||||
return VariantListResponse(variants=variants)
|
||||
|
||||
|
||||
@router.get("/{task_id}", response_model=VariantResponse)
|
||||
async def get_variant(
|
||||
task_id: str,
|
||||
store: VariantStore = Depends(get_variant_store),
|
||||
) -> VariantResponse:
|
||||
"""The stored variant owning the task key — the grading and telemetry
|
||||
join path; 404 when no variant was ever generated for the task."""
|
||||
record = store.get_by_task(task_id)
|
||||
if record is None:
|
||||
raise HTTPException(
|
||||
status_code=404, detail=f"no stored variant for task {task_id!r}"
|
||||
)
|
||||
return _to_response(record, competency_id=_competency_for(record.template_id))
|
||||
@@ -64,3 +64,37 @@ class Settings(BaseSettings):
|
||||
|
||||
# D-027: SQLite path for telemetry/grades/variants/defenses stores.
|
||||
db_path: Path = _SERVICE_ROOT / "ai_service" / "data" / "nextcraft.db"
|
||||
|
||||
# G-3 flood control (NOT backpressure-by-silence): max events ingested per
|
||||
# (learner_id, task_id) trace before the WS endpoint closes the connection
|
||||
# with 1008 and marks the trace INCOMPLETE_FLOODED. Drop-oldest is
|
||||
# FORBIDDEN — it corrupts grading input (GRILL G-3).
|
||||
telemetry_max_events_per_task: int = 50000
|
||||
|
||||
# Sandbox telemetry wiring (REQ-3-003): loopback host the in-sandbox capture
|
||||
# agent dials to reach this service's WS ingest (the agent joins the sandbox
|
||||
# mount ns but NOT the net ns — exec namespaces are offline, so the agent
|
||||
# shares the host network and reaches the app over loopback). Port reuses
|
||||
# `port` (A-004); only the host is configurable — never a second port.
|
||||
telemetry_ingest_host: str = "127.0.0.1"
|
||||
|
||||
# v0.3.5 network mode (D-038): dev.sh binds 0.0.0.0 so remote browsers can
|
||||
# reach the stack; '*' (default) lets any origin call the API (safe ONLY
|
||||
# because credentials are never enabled — A-008). Set a comma-separated
|
||||
# origin list (e.g. 'http://nextcraft-1:3000') to restrict instead.
|
||||
cors_origins: str = "*"
|
||||
|
||||
@property
|
||||
def cors_origin_list(self) -> list[str]:
|
||||
value = self.cors_origins.strip()
|
||||
if value == "*":
|
||||
return ["*"]
|
||||
return [o.strip() for o in value.split(",") if o.strip()]
|
||||
|
||||
# Voice provider selection (REQ-3-006, D-030): 'mock' (default — the
|
||||
# no-key path is first-class; tests never call a real voice API) or
|
||||
# 'browser' (browser-native SpeechRecognition/speechSynthesis fallback;
|
||||
# the descriptor tells the web client). The real server STT/TTS
|
||||
# ('openai-audio') is a v0.4 seam (GRILL CUT-1 / G-7) — AI_VOICE_BASE_URL
|
||||
# and AI_VOICE_API_KEY are documented in .env.example for that future.
|
||||
voice_provider: str = "mock"
|
||||
|
||||
@@ -1,5 +1,10 @@
|
||||
"""Pre-baked artifacts + rubrics + defense transcripts — Assessor mock inputs (REQ-2-008).
|
||||
|
||||
v0.2 mock engine inputs (pre-baked artifacts/rubrics/transcripts) — DORMANT as of v0.3 re-
|
||||
grounding (Task 6-1-04): no production code path imports this module. Retained as Phase-3
|
||||
calibration history.
|
||||
|
||||
|
||||
Counterpart: packages/mock-data/ai-scenarios.ts (artifact IDs string-identical,
|
||||
D-021). Real process-trace grading is a v0.3+ engine (assessment engine);
|
||||
these pre-baked submissions stand in for artifact + defense evaluation.
|
||||
|
||||
@@ -1,5 +1,10 @@
|
||||
"""Simulated sandbox telemetry corpus — Lab agent mock engine inputs (REQ-2-007).
|
||||
|
||||
v0.2 mock engine inputs (Lab/Proctor scenarios) — DORMANT as of v0.3 re-grounding (Task 6-1-04):
|
||||
no production code path imports this module. Retained as Phase-3 calibration history
|
||||
(corpus/trace_fixtures.py references it from TESTS only).
|
||||
|
||||
|
||||
Counterpart: packages/mock-data/ai-scenarios.ts (scenario IDs string-identical,
|
||||
D-021). Real sandbox telemetry is a v0.3+ engine (sandbox fabric); these
|
||||
scripted event streams stand in for the build-session process trace.
|
||||
|
||||
@@ -0,0 +1,230 @@
|
||||
"""Synthetic trace fixtures for grading calibration (Task 3-2-02, REQ-3-004).
|
||||
|
||||
Three builder archetypes as REAL `TelemetryEvent` traces (the grading
|
||||
engine's native input — these are NOT the v0.2 corpus's simplified
|
||||
`{timestamp, kind, detail}` event shapes):
|
||||
|
||||
strong builder — iterative debugging: small edits, tests early,
|
||||
failed runs closed by targeted fixes, eventual pass.
|
||||
lazy builder — one large paste, a single late test run, pass.
|
||||
(v0.2 alignment: the `lab-scenario-flagged` paste-
|
||||
and-run archetype; D-021.)
|
||||
struggling builder — many edit/test cycles, failures never close,
|
||||
never reaches a pass.
|
||||
|
||||
ID convention (D-021 alignment, documented in each fixture):
|
||||
v0.2 corpus scenario IDs are `<domain>-scenario-<slug>` (`lab-scenario-strong`,
|
||||
`lab-scenario-struggling`, `lab-scenario-flagged`, `proctor-scenario-*` — see
|
||||
corpus/telemetry.py). Grading fixtures carry ids string-aligned to that
|
||||
convention:
|
||||
|
||||
fixture id = "<scenario id>::<archetype>-trace"
|
||||
learner id = "learner-003" (a member of the corpus learner-00N id space;
|
||||
learner-001/002 exist in learner_context.py)
|
||||
|
||||
so a fixture is greppable against its v0.2 scenario counterpart while staying
|
||||
a distinct id space (a grading trace is a real event stream, not the v0.2
|
||||
mock scenario timeline — same convention, richer event kind set).
|
||||
|
||||
Instance hygiene: fixtures store event SPECS (plain tuples) and materialize
|
||||
FRESH `TelemetryEvent` instances on every `fixture.events` access. SQLModel
|
||||
rows carry SQLAlchemy instance state once a session has flushed them —
|
||||
re-adding the SAME instance to another store is a silent no-op, which would
|
||||
poison sequential test runs (a fixture ingested by test N would vanish for
|
||||
test N+1). Materializing per access keeps every consumer independent.
|
||||
|
||||
These fixtures exist for the CALIBRATION CONTRACT (tests/grading/
|
||||
test_calibration.py): the mock provider scripts archetype-mapped scores and
|
||||
the test asserts the ORDERING the rubric must eventually enforce. They are
|
||||
NOT an LLM quality benchmark — see the test module docstring for the honest
|
||||
scope statement.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from datetime import UTC, datetime, timedelta
|
||||
from typing import Final
|
||||
|
||||
from pydantic import BaseModel, Field
|
||||
|
||||
from ..telemetry.models import TelemetryEvent
|
||||
|
||||
_T0: Final = datetime(2026, 9, 12, 0, 0, 0, tzinfo=UTC)
|
||||
|
||||
#: Event spec: (seq, kind, payload, seconds-since-session-start).
|
||||
EventSpec = tuple[int, str, dict, float]
|
||||
|
||||
|
||||
class TraceFixture(BaseModel):
|
||||
"""One named synthetic trace + its v0.2 scenario alignment (D-021).
|
||||
|
||||
`event_specs` is the durable, session-state-free description; `events`
|
||||
materializes fresh TelemetryEvent rows from it on every access.
|
||||
"""
|
||||
|
||||
model_config = {"frozen": True}
|
||||
|
||||
fixture_id: str # "<v0.2 scenario id>::<archetype>-trace"
|
||||
archetype: str # "strong" | "lazy" | "struggling"
|
||||
aligned_scenario_id: str # the v0.2 corpus scenario this fixture mirrors
|
||||
competency_id: str # corpus competency id space (stack-orchestration-c00N)
|
||||
task_id: str # trace task id (grading operates on (learner, task))
|
||||
learner_id: str
|
||||
event_specs: tuple[EventSpec, ...] = Field(default=())
|
||||
|
||||
@property
|
||||
def events(self) -> list[TelemetryEvent]:
|
||||
"""FRESH TelemetryEvent instances — safe to ingest into any store.
|
||||
|
||||
Never cache these: an instance flushed by one SQLite session
|
||||
carries persistent identity, and re-appending it elsewhere no-ops.
|
||||
"""
|
||||
return [
|
||||
TelemetryEvent(
|
||||
learner_id=self.learner_id,
|
||||
task_id=self.task_id,
|
||||
seq=seq,
|
||||
kind=kind,
|
||||
payload=payload,
|
||||
ts=_T0 + timedelta(seconds=offset_s),
|
||||
sandbox_id="sbx-calibration",
|
||||
)
|
||||
for seq, kind, payload, offset_s in self.event_specs
|
||||
]
|
||||
|
||||
def description(self) -> str:
|
||||
return (
|
||||
f"{self.fixture_id} (archetype={self.archetype}, aligned="
|
||||
f"{self.aligned_scenario_id}, competency={self.competency_id})"
|
||||
)
|
||||
|
||||
|
||||
class _Builder:
|
||||
"""Seq-accurate event-spec builder for one (learner, task) pair."""
|
||||
|
||||
def __init__(self) -> None:
|
||||
self.specs: list[EventSpec] = []
|
||||
self._seq = 0
|
||||
|
||||
def add(self, kind: str, payload: dict | None, at: float) -> None:
|
||||
self.specs.append((self._seq, kind, payload or {}, at))
|
||||
self._seq += 1
|
||||
|
||||
|
||||
def _strong_builder_specs() -> list[EventSpec]:
|
||||
"""Iterative debugging: tests early, tight edit→test loops, eventual pass.
|
||||
|
||||
Mirrors `lab-scenario-strong` (v0.2: keystrokes → file_save → run_tests →
|
||||
test_pass, c002) at full telemetry fidelity — every failed cycle is
|
||||
closed by a targeted edit followed by a re-run that passes.
|
||||
"""
|
||||
b = _Builder()
|
||||
b.add("activity", {"state": "starting"}, 0)
|
||||
b.add("file_diff", {"path": "planner.py", "added": 14}, 95) # small edit
|
||||
b.add("command", {"cmd": "pytest -q tests/test_planner.py"}, 120) # tests EARLY
|
||||
b.add("test_result", {"passed": True, "exit_code": 0}, 126)
|
||||
b.add("file_diff", {"path": "tool_node.py", "added": 22}, 180)
|
||||
b.add("command", {"cmd": "pytest -q"}, 275)
|
||||
b.add("test_result", {"passed": False, "exit_code": 1}, 281) # honest failure
|
||||
b.add("file_diff", {"path": "tool_node.py", "added": 6, "removed": 2}, 340) # targeted fix
|
||||
b.add("command", {"cmd": "pytest -q"}, 430)
|
||||
b.add("test_result", {"passed": True, "exit_code": 0}, 436) # cycle CLOSED
|
||||
b.add("command", {"cmd": "git commit -m 'tool node with retries'"}, 500)
|
||||
return b.specs
|
||||
|
||||
|
||||
def _lazy_builder_specs() -> list[EventSpec]:
|
||||
"""Paste-and-run: one large paste, a single LATE test run, instant pass.
|
||||
|
||||
Mirrors `lab-scenario-flagged` (v0.2: paste of 2,400 chars → run_tests →
|
||||
instant 6/6 pass, c003) at full telemetry fidelity — zero iteration, zero
|
||||
verification during construction, one terminal test run only.
|
||||
"""
|
||||
b = _Builder()
|
||||
b.add("activity", {"state": "starting"}, 0)
|
||||
b.add("file_diff", {"path": "eval.py", "added": 240, "removed": 0}, 30) # one bulk paste
|
||||
b.add("file_diff", {"path": "README.md", "added": 12}, 40)
|
||||
b.add("command", {"cmd": "npm run build"}, 45)
|
||||
b.add("run_result", {"exit_code": 0, "ok": True}, 60)
|
||||
b.add("command", {"cmd": "pytest -q"}, 520) # single LATE test run
|
||||
b.add("test_result", {"passed": True, "exit_code": 0}, 540) # instant pass
|
||||
return b.specs
|
||||
|
||||
|
||||
def _struggling_builder_specs() -> list[EventSpec]:
|
||||
"""Many cycles, none close: repeated failures, no eventual pass.
|
||||
|
||||
Mirrors `lab-scenario-struggling` (v0.2: repeated identical ImportErrors,
|
||||
idle gaps, no checkpoint, c002) at full telemetry fidelity — edits happen
|
||||
between failures, but the same failure recurs; no pass is ever reached.
|
||||
"""
|
||||
b = _Builder()
|
||||
b.add("activity", {"state": "starting"}, 0)
|
||||
b.add("file_diff", {"path": "main.py", "added": 40}, 20)
|
||||
b.add("command", {"cmd": "pytest -q"}, 210)
|
||||
b.add("test_result", {"passed": False, "exit_code": 1}, 215) # ImportError
|
||||
b.add("file_diff", {"path": "main.py", "added": 8, "removed": 3}, 300)
|
||||
b.add("command", {"cmd": "pytest -q"}, 520)
|
||||
b.add("test_result", {"passed": False, "exit_code": 1}, 525) # SAME error
|
||||
b.add("file_diff", {"path": "main.py", "added": 5}, 610)
|
||||
b.add("command", {"cmd": "pytest -q"}, 960)
|
||||
b.add("test_result", {"passed": False, "exit_code": 1}, 965) # STILL failing
|
||||
b.add("activity", {"state": "idle"}, 1500) # long idle
|
||||
b.add("activity", {"state": "idle"}, 2200)
|
||||
return b.specs
|
||||
|
||||
|
||||
#: Calibration learner — a member of the corpus learner-00N id space (D-021;
|
||||
#: learner-001/002 live in corpus/learner_context.py; grading fixtures use
|
||||
#: a third id so calibration traces never collide with mock-context reads).
|
||||
CALIBRATION_LEARNER_ID: Final = "learner-003"
|
||||
|
||||
|
||||
STRONG_BUILDER: Final = TraceFixture(
|
||||
fixture_id="lab-scenario-strong::strong-trace",
|
||||
archetype="strong",
|
||||
aligned_scenario_id="lab-scenario-strong",
|
||||
competency_id="stack-orchestration-c002",
|
||||
task_id="task-calibration-strong",
|
||||
learner_id=CALIBRATION_LEARNER_ID,
|
||||
event_specs=tuple(_strong_builder_specs()),
|
||||
)
|
||||
|
||||
LAZY_BUILDER: Final = TraceFixture(
|
||||
# v0.2's paste-and-run archetype is the "flagged" lab scenario (D-021):
|
||||
# large paste → instant test pass. "lazy builder" is that behavior
|
||||
# without the proctor flag; the alignment is behavioral, documented here.
|
||||
fixture_id="lab-scenario-flagged::lazy-trace",
|
||||
archetype="lazy",
|
||||
aligned_scenario_id="lab-scenario-flagged",
|
||||
competency_id="stack-orchestration-c003",
|
||||
task_id="task-calibration-lazy",
|
||||
learner_id=CALIBRATION_LEARNER_ID,
|
||||
event_specs=tuple(_lazy_builder_specs()),
|
||||
)
|
||||
|
||||
STRUGGLING_BUILDER: Final = TraceFixture(
|
||||
fixture_id="lab-scenario-struggling::struggling-trace",
|
||||
archetype="struggling",
|
||||
aligned_scenario_id="lab-scenario-struggling",
|
||||
competency_id="stack-orchestration-c002",
|
||||
task_id="task-calibration-struggling",
|
||||
learner_id=CALIBRATION_LEARNER_ID,
|
||||
event_specs=tuple(_struggling_builder_specs()),
|
||||
)
|
||||
|
||||
TRACE_FIXTURES: Final[dict[str, TraceFixture]] = {
|
||||
f.fixture_id: f
|
||||
for f in (STRONG_BUILDER, LAZY_BUILDER, STRUGGLING_BUILDER)
|
||||
}
|
||||
|
||||
|
||||
def get_trace_fixture(fixture_id: str) -> TraceFixture | None:
|
||||
return TRACE_FIXTURES.get(fixture_id)
|
||||
|
||||
|
||||
def digest_of(fixture: TraceFixture):
|
||||
"""Compute the digest for a fixture (pure compute; test-side helper)."""
|
||||
from ..grading.features import compute_digest
|
||||
|
||||
return compute_digest(fixture.events)
|
||||
@@ -0,0 +1,27 @@
|
||||
"""Process-trace grading — trace digest, rubric engine, GradeStore (REQ-3-004).
|
||||
|
||||
Boundary rule (D-027): grading/ is an engine module — it never imports
|
||||
api/; the single sanctioned agents/ dependency is the shared D-020
|
||||
structured defense (agents/structured.py), imported module-direct in
|
||||
engine.py (see its docstring for why). features.py imports telemetry/;
|
||||
store.py imports config only; engine.py composes llm/ + telemetry/ +
|
||||
prompts/ + agents.structured.
|
||||
|
||||
Wave status: features.py (TraceDigest, compute_digest) + store.py
|
||||
(GradeRecord, GradeStore, SQLiteGradeStore) landed in Wave 1 (3-1-01 +
|
||||
3-1-02); engine.py (GradingEngine, RubricScore) is Wave 2 (3-2-01).
|
||||
"""
|
||||
|
||||
from .engine import GradingEngine, RubricScore
|
||||
from .features import TraceDigest, compute_digest
|
||||
from .store import GradeRecord, GradeStore, SQLiteGradeStore
|
||||
|
||||
__all__ = [
|
||||
"GradeRecord",
|
||||
"GradeStore",
|
||||
"GradingEngine",
|
||||
"RubricScore",
|
||||
"SQLiteGradeStore",
|
||||
"TraceDigest",
|
||||
"compute_digest",
|
||||
]
|
||||
@@ -0,0 +1,313 @@
|
||||
"""GradingEngine — rubric scoring over real process traces (REQ-3-004).
|
||||
|
||||
The Wave-2 composition of the grading stack:
|
||||
|
||||
trace completeness gate (G-4, FIRST — nothing is sent to any LLM
|
||||
when the gate trips) → TraceStore.get_trace → compute_digest (D-028)
|
||||
→ prompts.grading.render_trace_digest → D-020 4-layer structured
|
||||
defense (agents/structured.py — REUSED, composed, never duplicated)
|
||||
→ RubricScore validation → GradeStore persistence → GradeRecord.
|
||||
|
||||
Gate verdicts are FIRST-CLASS RESULTS, not exceptions (G-4 is binding):
|
||||
UNGRADABLE_TRACE_INCOMPLETE — seq gaps in the store OR the trace is
|
||||
flagged by TraceIntegrityMap (INCOMPLETE_FLOODED). The gap list /
|
||||
flag reason is surfaced in `scores` for API rendering, and the
|
||||
record is PERSISTED like any grade so a learner sees why no
|
||||
credential can be issued for this trace — the gate outcome is
|
||||
durable and auditable, not a transient error string.
|
||||
UNGRADABLE_EMPTY_TRACE — no events stored for the pair.
|
||||
On either verdict `scores.criteria` is empty and the LLM is never called.
|
||||
|
||||
DI (D-027/D-032 house pattern): the engine receives trace_store,
|
||||
grade_store, integrity and provider through the constructor and knows
|
||||
NOTHING of FastAPI — api/ composes it (Task 3-3-01). `model` is injected
|
||||
alongside the provider so tests script the mock against the production
|
||||
wiring without touching Settings.
|
||||
|
||||
Boundary (D-027): grading/ imports llm/ (provider protocol + Message
|
||||
type), telemetry/ (store + integrity map), prompts/ (rubric text) and
|
||||
ONLY agents.structured — the sanctioned shared D-020 defense. We import
|
||||
the MODULE directly (`from ..agents.structured import structured_completion`)
|
||||
rather than the `agents` package, mirroring how agents/base.py consumes
|
||||
it (same direct-module import): that keeps the dependency surface to
|
||||
exactly the two names the engine needs (structured_completion,
|
||||
StructuredOutputError) and avoids executing agents/__init__ re-exports
|
||||
(BaseAgent, registry, session store) that grading has no business
|
||||
loading — a side-effect-hygiene choice that keeps this import line
|
||||
grep-auditable as "the one agents dependency". grading/ never imports api/.
|
||||
|
||||
RubricScore placement (documented decision): the validated output model
|
||||
lives HERE, not in prompts/. The pydantic model is the engine's return
|
||||
CONTRACT (the shape GradeStore.scores must hold), while prompts/grading.py
|
||||
is pure prompt text + its mirror schema HINT string — the same split as
|
||||
agents/assessor.py (model + hint) but with the model owned by the engine
|
||||
module that validates it. Prompt files hold text, per the prompts/ house
|
||||
style; engine files hold typed contracts.
|
||||
"""
|
||||
|
||||
import logging
|
||||
from datetime import UTC, datetime
|
||||
from typing import TYPE_CHECKING, Final
|
||||
|
||||
from pydantic import BaseModel, ConfigDict, Field, field_validator
|
||||
|
||||
from ..agents.structured import StructuredOutputError, structured_completion
|
||||
from ..llm.base import LLMProvider
|
||||
from ..llm.types import Message
|
||||
from ..prompts.grading import (
|
||||
RUBRIC_CRITERIA,
|
||||
RUBRIC_SCORE_SCHEMA_HINT,
|
||||
SYSTEM_PROMPT,
|
||||
render_trace_digest,
|
||||
)
|
||||
from ..telemetry.ingest import TraceIntegrityMap
|
||||
from ..telemetry.store import TraceStore
|
||||
from .features import compute_digest
|
||||
from .store import GradeRecord, GradeStore
|
||||
|
||||
if TYPE_CHECKING: # pragma: no cover - protocol-only import for the optional
|
||||
from ..variants.store import VariantStore # noqa: TC001 (variant-blind without it)
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
#: First-class gate verdicts (G-4). GradeRecord.verdict values; non-empty
|
||||
#: by store contract. Rubric verdicts (mastered/developing/not_yet) ride in
|
||||
#: `scores.verdict` — `record.verdict` stays the machine-readable outcome.
|
||||
VERDICT_UNGRADABLE_INCOMPLETE: Final = "UNGRADABLE_TRACE_INCOMPLETE"
|
||||
VERDICT_UNGRADABLE_EMPTY: Final = "UNGRADABLE_EMPTY_TRACE"
|
||||
#: record.verdict for a successfully LLM-graded trace (the rubric verdict
|
||||
#: travels inside scores); keeps verdict non-empty for every persisted row.
|
||||
VERDICT_GRADED: Final = "GRADED"
|
||||
|
||||
#: Gate-detail keys surfaced in GradeRecord.scores (tests assert on these).
|
||||
_INTEGRITY_FLAG_KEY: Final = "integrity_flag"
|
||||
|
||||
|
||||
def _anchors_context(variant) -> str: # noqa: ANN001 - VariantRecord (duck-typed)
|
||||
"""Render the variant template's difficulty anchors for the grader prompt.
|
||||
|
||||
Contains only the template id + anchor numbers — no learner-identifying
|
||||
material (D-028 anonymity preserved; the digest-leak tests keep holding).
|
||||
Lazy template import: grading must not import variants/ at module load
|
||||
(variants/prompts import-cycle safety mirrors llm/ rules).
|
||||
"""
|
||||
from ..variants.templates import get_template
|
||||
|
||||
template = get_template(variant.template_id)
|
||||
if template is None:
|
||||
return f"template={variant.template_id} (anchors unavailable)"
|
||||
a = template.rubric_anchors
|
||||
return (
|
||||
f"template={template.id}; "
|
||||
f"expected_edit_count_band={list(a.expected_edit_count_band)}; "
|
||||
f"expected_min_test_runs={a.expected_min_test_runs}; "
|
||||
f"expected_error_fix_cycles_band={list(a.expected_error_fix_cycles_band)}"
|
||||
)
|
||||
_MISSING_SEQS_KEY: Final = "missing_seqs"
|
||||
|
||||
|
||||
class RubricScore(BaseModel):
|
||||
"""Validated LLM output: per-criterion 0-4 scores + strengths + gaps + verdict.
|
||||
|
||||
The D-020 schema for the grader: `structured_completion` parses the
|
||||
model reply into THIS shape (layer 3), retrying once with the
|
||||
validation error fed back (layer 4). Exact criteria set + 0-4 ranges
|
||||
are enforced here, so the scores dict persisted to GradeStore is
|
||||
always rubric-shaped no matter what the model produced.
|
||||
"""
|
||||
|
||||
model_config = ConfigDict(extra="forbid")
|
||||
|
||||
criteria: dict[str, int]
|
||||
strengths: list[str] = Field(min_length=1, max_length=2)
|
||||
gaps: list[str] = Field(min_length=1, max_length=2)
|
||||
verdict: str # "mastered" | "developing" | "not_yet"
|
||||
|
||||
@field_validator("criteria")
|
||||
@classmethod
|
||||
def _criteria_rubric_shaped(cls, value: dict[str, int]) -> dict[str, int]:
|
||||
"""Exact criteria keys (no extras, no omissions) and 0-4 scores."""
|
||||
expected = set(RUBRIC_CRITERIA)
|
||||
got = set(value)
|
||||
if got != expected:
|
||||
raise ValueError(
|
||||
f"criteria keys must be exactly {sorted(expected)}, got {sorted(got)}"
|
||||
)
|
||||
for key, score in value.items():
|
||||
if not 0 <= score <= 4:
|
||||
raise ValueError(f"criterion {key!r} must be within 0-4, got {score}")
|
||||
return value
|
||||
|
||||
@field_validator("verdict")
|
||||
@classmethod
|
||||
def _verdict_known(cls, value: str) -> str:
|
||||
allowed = {"mastered", "developing", "not_yet"}
|
||||
if value not in allowed:
|
||||
raise ValueError(f"verdict must be one of {sorted(allowed)}, got {value!r}")
|
||||
return value
|
||||
|
||||
|
||||
class GradingEngine:
|
||||
"""Scores a (learner_id, task_id) trace into a persisted GradeRecord."""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
trace_store: TraceStore,
|
||||
grade_store: GradeStore,
|
||||
integrity: TraceIntegrityMap,
|
||||
provider: LLMProvider,
|
||||
*,
|
||||
model: str = "gemma4:31b",
|
||||
variant_store: "VariantStore | None" = None,
|
||||
) -> None:
|
||||
self._trace_store = trace_store
|
||||
self._grade_store = grade_store
|
||||
self._integrity = integrity
|
||||
self._provider = provider
|
||||
self._model = model
|
||||
# Phase 4 (MH#4): optional variant lookup — when the graded task
|
||||
# derives from a generated variant, its template's difficulty anchors
|
||||
# ship to the grader prompt (same bar for every variant of the
|
||||
# template, a-5) and the variant seed is stamped on the record.
|
||||
# Optional so engine tests stay decoupled; main.py lifespan wires it.
|
||||
self._variant_store = variant_store
|
||||
|
||||
async def grade(self, learner_id: str, task_id: str) -> GradeRecord:
|
||||
"""Grade one trace; persist latest-state (GradeStore upserts); return it.
|
||||
|
||||
Gate FIRST (G-4): the LLM is only ever reached from the fully
|
||||
guarded path — no gate state can be masked by an LLM error.
|
||||
"""
|
||||
record = await self._grade(learner_id, task_id)
|
||||
self._grade_store.save(record)
|
||||
return record
|
||||
|
||||
# ------------------------------------------------------------------ core
|
||||
|
||||
async def _grade(self, learner_id: str, task_id: str) -> GradeRecord:
|
||||
# -- G-4 gate FIRST: integrity flag OR seq gaps. Ordering matters:
|
||||
# gaps() returns [] for an EMPTY trace, so the empty check below is
|
||||
# reachable only when no rows exist at all; a gapped or flooded
|
||||
# trace can never fall through to the LLM path.
|
||||
if self._integrity.is_incomplete(learner_id, task_id):
|
||||
reason = self._integrity.reason(learner_id, task_id) or "unknown"
|
||||
gaps = self._trace_store.gaps(learner_id, task_id)
|
||||
logger.info(
|
||||
"grade gate (G-4): %s/%s integrity-flagged (%s) — ungradable",
|
||||
learner_id,
|
||||
task_id,
|
||||
reason,
|
||||
)
|
||||
return self._ungradable(
|
||||
learner_id,
|
||||
task_id,
|
||||
detail={_INTEGRITY_FLAG_KEY: reason, _MISSING_SEQS_KEY: gaps},
|
||||
verdict=VERDICT_UNGRADABLE_INCOMPLETE,
|
||||
)
|
||||
gaps = self._trace_store.gaps(learner_id, task_id)
|
||||
if gaps:
|
||||
logger.info(
|
||||
"grade gate (G-4): %s/%s seq gaps %s — ungradable",
|
||||
learner_id,
|
||||
task_id,
|
||||
gaps,
|
||||
)
|
||||
return self._ungradable(
|
||||
learner_id,
|
||||
task_id,
|
||||
detail={_INTEGRITY_FLAG_KEY: None, _MISSING_SEQS_KEY: gaps},
|
||||
verdict=VERDICT_UNGRADABLE_INCOMPLETE,
|
||||
)
|
||||
|
||||
trace = self._trace_store.get_trace(learner_id, task_id)
|
||||
if not trace:
|
||||
logger.info(
|
||||
"grade gate: %s/%s empty trace — ungradable", learner_id, task_id
|
||||
)
|
||||
return self._ungradable(
|
||||
learner_id,
|
||||
task_id,
|
||||
detail={_INTEGRITY_FLAG_KEY: None, _MISSING_SEQS_KEY: []},
|
||||
verdict=VERDICT_UNGRADABLE_EMPTY,
|
||||
)
|
||||
|
||||
# -- Guarded path: digest (D-028) → prompt → D-020 4-layer defense.
|
||||
digest = compute_digest(trace)
|
||||
variant = self._lookup_variant(task_id)
|
||||
anchors_context = (
|
||||
_anchors_context(variant) if variant is not None else None
|
||||
)
|
||||
messages = [
|
||||
Message(role="system", content=SYSTEM_PROMPT),
|
||||
Message(
|
||||
role="user",
|
||||
content=render_trace_digest(digest, anchors_context=anchors_context),
|
||||
),
|
||||
]
|
||||
try:
|
||||
rubric = await structured_completion(
|
||||
self._provider,
|
||||
messages,
|
||||
model=self._model,
|
||||
schema=RubricScore,
|
||||
schema_hint=RUBRIC_SCORE_SCHEMA_HINT,
|
||||
)
|
||||
except StructuredOutputError as exc:
|
||||
# The trace was gradable but the model failed to produce valid
|
||||
# JSON within the D-020 budget (two attempts). Raise — the API
|
||||
# layer maps this to a 502 (assessor precedent). Persisting a
|
||||
# fabricated or partial grade here would violate the no-silent-
|
||||
# fallback rule: no credential-worthy record without a validated
|
||||
# RubricScore.
|
||||
raise StructuredOutputError(f"grading LLM failed for {task_id}: {exc}") from exc
|
||||
|
||||
logger.debug(
|
||||
"graded %s/%s: %s (model=%s)",
|
||||
learner_id,
|
||||
task_id,
|
||||
rubric.verdict,
|
||||
self._model,
|
||||
)
|
||||
return GradeRecord(
|
||||
learner_id=learner_id,
|
||||
task_id=task_id,
|
||||
variant_seed=variant.seed if variant is not None else None, # D-029
|
||||
digest=digest.model_dump(),
|
||||
scores=rubric.model_dump(),
|
||||
verdict=VERDICT_GRADED,
|
||||
model=self._model,
|
||||
created_at=datetime.now(tz=UTC),
|
||||
)
|
||||
|
||||
def _lookup_variant(self, task_id: str): # noqa: ANN202 - VariantRecord | None
|
||||
"""MH#4: resolve the graded task's variant (None when not variant-derived)."""
|
||||
if self._variant_store is None:
|
||||
return None
|
||||
return self._variant_store.get_by_task(task_id)
|
||||
|
||||
# ------------------------------------------------------------ gate record
|
||||
|
||||
@staticmethod
|
||||
def _ungradable(
|
||||
learner_id: str,
|
||||
task_id: str,
|
||||
*,
|
||||
detail: dict,
|
||||
verdict: str,
|
||||
) -> GradeRecord:
|
||||
"""Build a gate record: no digest (nothing was graded), gate detail
|
||||
surfaced in `scores` (the store allows an empty scores dict, but
|
||||
G-4 requires the gap list / flag reason surfaced — the detail IS the
|
||||
verdict's payload), model="none" (no LLM was involved; provenance
|
||||
stays honest).
|
||||
"""
|
||||
return GradeRecord(
|
||||
learner_id=learner_id,
|
||||
task_id=task_id,
|
||||
variant_seed=None,
|
||||
digest={},
|
||||
scores=detail,
|
||||
verdict=verdict,
|
||||
model="none",
|
||||
created_at=datetime.now(tz=UTC),
|
||||
)
|
||||
@@ -0,0 +1,246 @@
|
||||
"""Deterministic process-trace digest (D-028, REQ-3-004).
|
||||
|
||||
Pure compute — no LLM, no I/O. `compute_digest` reduces an ordered
|
||||
TelemetryEvent trace to a compact, bounded `TraceDigest` that is safe to
|
||||
embed in a grading prompt:
|
||||
|
||||
- FIXED fields + small histograms only; NO raw commands, NO file contents,
|
||||
NO payloads — the raw trace NEVER reaches the LLM (D-028), which also
|
||||
bounds the prompt-injection surface.
|
||||
- Tolerant to both live trace mixes: daemon-topology traces carry
|
||||
`activity` + `file_diff` kinds (workspace watcher), while REPL-driven
|
||||
traces carry `command`/`stdin`/`stdout`/`run_result`/`test_result`
|
||||
(P2 verification P1). Features derive from whatever kinds are present and
|
||||
never crash on absent kinds.
|
||||
|
||||
Feature semantics (conservative, deterministic):
|
||||
- test pass/fail counts + final status derive from `test_result` payloads
|
||||
when present, falling back to `run_result` exit codes (0 = pass).
|
||||
- an error/fix CYCLE = a failing run/test followed by >= 1 edit and then a
|
||||
later run/test (pass or fail) — the next observed result closes the cycle.
|
||||
- idle gaps = wall-clock gaps between consecutive events exceeding
|
||||
`idle_threshold_s` (default 120s): count + total seconds.
|
||||
- command category histogram classifies `command`-kind payloads: build /
|
||||
test / file / nav / debug / other.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from collections import Counter
|
||||
|
||||
from pydantic import BaseModel, Field
|
||||
|
||||
from ..telemetry.models import TelemetryEvent
|
||||
|
||||
_IDLE_DEFAULT_S: float = 120.0
|
||||
|
||||
_TEST_HINTS = ("test", "pytest", "vitest", "jest", "mocha", "unittest", "go test", "npm test")
|
||||
_BUILD_HINTS = ("make", "npm run build", "pip install", "pnpm", "cargo build", "gcc", "tsc")
|
||||
_DEBUG_HINTS = ("gdb", "pdb", "print(", "debug", "strace", "ltrace", "curl", "ping")
|
||||
_NAV_HINTS = ("ls", "cd", "pwd", "cat ", "grep ", "find", "rg ", "tree", "head", "tail", "less")
|
||||
_FILE_HINTS = ("mv ", "cp ", "rm ", "mkdir", "touch", "chmod", "nano", "vim", "sed -i", "tee ")
|
||||
|
||||
|
||||
class TraceDigest(BaseModel):
|
||||
"""Compact, bounded, LLM-safe summary of a process trace (D-028).
|
||||
|
||||
Fixed fields + small histograms. Serializes well under 4 KB; contains no
|
||||
raw commands, file contents, or event payloads.
|
||||
"""
|
||||
|
||||
model_config = {"frozen": True}
|
||||
|
||||
event_count: int = Field(ge=0)
|
||||
session_duration_s: float = Field(ge=0.0)
|
||||
edit_count: int = Field(ge=0)
|
||||
command_count: int = Field(ge=0)
|
||||
run_count: int = Field(ge=0)
|
||||
test_pass_count: int = Field(ge=0)
|
||||
test_fail_count: int = Field(ge=0)
|
||||
final_test_status: str = Field(pattern="^(pass|fail|none)$")
|
||||
first_test_pass_offset_s: float | None = None
|
||||
error_fix_cycles: int = Field(ge=0)
|
||||
mean_fix_latency_s: float | None = None
|
||||
idle_gap_count: int = Field(ge=0)
|
||||
idle_gap_total_s: float = Field(ge=0.0)
|
||||
command_categories: dict[str, int] = Field(default_factory=dict)
|
||||
kind_histogram: dict[str, int] = Field(default_factory=dict)
|
||||
|
||||
|
||||
def _event_pass_status(event: TelemetryEvent) -> bool | None:
|
||||
"""True (pass) / False (fail) / None (not a result event) for one event."""
|
||||
payload = event.payload or {}
|
||||
if event.kind == "test_result":
|
||||
if "passed" in payload:
|
||||
return bool(payload["passed"])
|
||||
if "exit_code" in payload:
|
||||
return int(payload["exit_code"]) == 0
|
||||
if "status" in payload:
|
||||
return str(payload["status"]).lower() in ("pass", "passed", "ok", "success")
|
||||
return None
|
||||
if event.kind == "run_result":
|
||||
if "exit_code" in payload:
|
||||
return int(payload["exit_code"]) == 0
|
||||
if "ok" in payload:
|
||||
return bool(payload["ok"])
|
||||
return None
|
||||
return None
|
||||
|
||||
|
||||
def _classify_command(text: str) -> str:
|
||||
lowered = text.lower()
|
||||
if any(h in lowered for h in _TEST_HINTS):
|
||||
return "test"
|
||||
if any(h in lowered for h in _BUILD_HINTS):
|
||||
return "build"
|
||||
if any(h in lowered for h in _DEBUG_HINTS):
|
||||
return "debug"
|
||||
if any(h in lowered for h in _NAV_HINTS):
|
||||
return "nav"
|
||||
if any(h in lowered for h in _FILE_HINTS):
|
||||
return "file"
|
||||
return "other"
|
||||
|
||||
|
||||
def _command_text(event: TelemetryEvent) -> str:
|
||||
payload = event.payload or {}
|
||||
return str(payload.get("cmd") or payload.get("command") or payload.get("line") or "")
|
||||
|
||||
|
||||
def compute_digest(
|
||||
trace: list[TelemetryEvent], *, idle_threshold_s: float = _IDLE_DEFAULT_S
|
||||
) -> TraceDigest:
|
||||
"""Reduce an ordered trace to a bounded digest. Never raises on odd input."""
|
||||
events = sorted(trace, key=lambda e: (e.seq, e.ts))
|
||||
if not events:
|
||||
return TraceDigest(
|
||||
event_count=0,
|
||||
session_duration_s=0.0,
|
||||
edit_count=0,
|
||||
command_count=0,
|
||||
run_count=0,
|
||||
test_pass_count=0,
|
||||
test_fail_count=0,
|
||||
final_test_status="none",
|
||||
first_test_pass_offset_s=None,
|
||||
error_fix_cycles=0,
|
||||
mean_fix_latency_s=None,
|
||||
idle_gap_count=0,
|
||||
idle_gap_total_s=0.0,
|
||||
command_categories={},
|
||||
kind_histogram={},
|
||||
)
|
||||
|
||||
kind_histogram = Counter(e.kind for e in events)
|
||||
start_ts = events[0].ts
|
||||
end_ts = events[-1].ts
|
||||
duration = max(0.0, (end_ts - start_ts).total_seconds())
|
||||
|
||||
edit_count = kind_histogram.get("file_diff", 0)
|
||||
command_count = kind_histogram.get("command", 0)
|
||||
run_count = kind_histogram.get("run_result", 0)
|
||||
|
||||
# Tests: prefer test_result events; fall back to run_result exit codes.
|
||||
test_statuses: list[tuple[TelemetryEvent, bool]] = []
|
||||
for e in events:
|
||||
if e.kind == "test_result":
|
||||
ok = _event_pass_status(e)
|
||||
if ok is not None:
|
||||
test_statuses.append((e, ok))
|
||||
if not test_statuses:
|
||||
for e in events:
|
||||
if e.kind == "run_result":
|
||||
ok = _event_pass_status(e)
|
||||
if ok is not None:
|
||||
test_statuses.append((e, ok))
|
||||
|
||||
test_pass_count = sum(1 for _, ok in test_statuses if ok)
|
||||
test_fail_count = len(test_statuses) - test_pass_count
|
||||
if not test_statuses:
|
||||
final_test_status = "none"
|
||||
else:
|
||||
final_test_status = "pass" if test_statuses[-1][1] else "fail"
|
||||
first_pass = next((e for e, ok in test_statuses if ok), None)
|
||||
first_pass_offset = (
|
||||
max(0.0, (first_pass.ts - start_ts).total_seconds()) if first_pass is not None else None
|
||||
)
|
||||
|
||||
# Error/fix cycles: a failing result starts a pending cycle; the NEXT
|
||||
# observed result closes it (regardless of outcome) — a fix attempt that
|
||||
# fails again is itself another iteration of debugging, so it closes the
|
||||
# previous cycle and opens a new one. Edits since the fail mark the
|
||||
# close as a genuine fix attempt; latency = first edit -> closing result.
|
||||
cycles = 0
|
||||
fix_latencies: list[float] = []
|
||||
pending_fail_ts: float | None = None # seconds since start
|
||||
edits_since_fail = 0
|
||||
first_edit_ts: float | None = None
|
||||
for e in events:
|
||||
t = max(0.0, (e.ts - start_ts).total_seconds())
|
||||
if e.kind == "file_diff":
|
||||
if pending_fail_ts is not None:
|
||||
if edits_since_fail == 0:
|
||||
first_edit_ts = t
|
||||
edits_since_fail += 1
|
||||
continue
|
||||
ok = _event_pass_status(e)
|
||||
if ok is None:
|
||||
continue
|
||||
if ok is False:
|
||||
if pending_fail_ts is not None and edits_since_fail > 0 and first_edit_ts is not None:
|
||||
# failed fix attempt: closes the previous cycle, opens a new one
|
||||
cycles += 1
|
||||
fix_latencies.append(t - first_edit_ts)
|
||||
pending_fail_ts = t
|
||||
edits_since_fail = 0
|
||||
first_edit_ts = None
|
||||
continue
|
||||
if ok is True and pending_fail_ts is not None:
|
||||
if edits_since_fail > 0 and first_edit_ts is not None:
|
||||
cycles += 1
|
||||
fix_latencies.append(t - first_edit_ts)
|
||||
pending_fail_ts = None
|
||||
edits_since_fail = 0
|
||||
first_edit_ts = None
|
||||
|
||||
mean_fix_latency = (
|
||||
sum(fix_latencies) / len(fix_latencies) if fix_latencies else None
|
||||
)
|
||||
|
||||
# Idle gaps between consecutive events.
|
||||
idle_gap_count = 0
|
||||
idle_gap_total = 0.0
|
||||
prev_ts = None
|
||||
for e in events:
|
||||
if prev_ts is not None:
|
||||
gap = (e.ts - prev_ts).total_seconds()
|
||||
if gap > idle_threshold_s:
|
||||
idle_gap_count += 1
|
||||
idle_gap_total += gap
|
||||
prev_ts = e.ts
|
||||
|
||||
# Command category histogram (command-kind events only).
|
||||
categories: Counter[str] = Counter()
|
||||
for e in events:
|
||||
if e.kind == "command":
|
||||
categories[_classify_command(_command_text(e))] += 1
|
||||
|
||||
return TraceDigest(
|
||||
event_count=len(events),
|
||||
session_duration_s=round(duration, 3),
|
||||
edit_count=edit_count,
|
||||
command_count=command_count,
|
||||
run_count=run_count,
|
||||
test_pass_count=test_pass_count,
|
||||
test_fail_count=test_fail_count,
|
||||
final_test_status=final_test_status,
|
||||
first_test_pass_offset_s=(
|
||||
round(first_pass_offset, 3) if first_pass_offset is not None else None
|
||||
),
|
||||
error_fix_cycles=cycles,
|
||||
mean_fix_latency_s=(round(mean_fix_latency, 3) if mean_fix_latency is not None else None),
|
||||
idle_gap_count=idle_gap_count,
|
||||
idle_gap_total_s=round(idle_gap_total, 3),
|
||||
command_categories=dict(sorted(categories.items())),
|
||||
kind_histogram=dict(sorted(kind_histogram.items())),
|
||||
)
|
||||
@@ -0,0 +1,237 @@
|
||||
"""GradeStore — grade persistence protocol + SQLite implementation (REQ-3-004, D-027).
|
||||
|
||||
Postgres-migration-ready (D-027): the protocol is the only surface the
|
||||
grading engine and API layers touch; swapping SQLiteGradeStore for a
|
||||
Postgres-backed implementation must not change call sites. The
|
||||
`grade_record` table uses only portable column types (str / JSON /
|
||||
datetime), so the same SQLModel schema stands up unchanged on Postgres.
|
||||
|
||||
Upsert, NOT append: (learner_id, task_id) is the grade identity — one row
|
||||
per learner per task holding the LATEST grade. `save` overwrites the whole
|
||||
row when the pair already exists, so a regrade replaces scores, verdict,
|
||||
created_at, digest, model and variant_seed wholesale. That is deliberately
|
||||
the opposite of TraceStore.append's dedup-keep-first contract: a trace is an
|
||||
append-only event log, a grade is latest-state, so the engine can re-grade
|
||||
a task idempotently as its rubric or input evolves.
|
||||
|
||||
Concurrency (a-3): the engine enables WAL + synchronous=NORMAL and a busy
|
||||
timeout at connection time, so a regrade writer and API readers do not hit
|
||||
`database is locked` on the single-box pilot.
|
||||
|
||||
`created_at` contract: callers stamp UTC (datetime.now(UTC)); SQLite stores
|
||||
it naive and the read paths re-label it tz-aware UTC (same boundary
|
||||
normalization as TelemetryEvent.ts, so the contract holds on any backend).
|
||||
|
||||
Boundary (D-027): `grading/` never imports `agents/` / `api/`; this module
|
||||
imports config only.
|
||||
"""
|
||||
|
||||
import logging
|
||||
import sqlite3
|
||||
from collections.abc import Iterator
|
||||
from contextlib import contextmanager
|
||||
from datetime import UTC, datetime
|
||||
from pathlib import Path
|
||||
from typing import Any, Protocol
|
||||
|
||||
import sqlalchemy as sa
|
||||
from sqlalchemy import JSON, Index
|
||||
from sqlalchemy.orm import validates
|
||||
from sqlmodel import Field, Session, SQLModel, create_engine, select
|
||||
|
||||
from ..config import Settings
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class GradeRecord(SQLModel, table=True):
|
||||
"""A persisted grade; (learner_id, task_id) is the PK — latest wins.
|
||||
|
||||
Written by the grading engine (one save per grade attempt), read by the
|
||||
API layer through the GradeStore protocol. Constraint enforcement
|
||||
mirrors TelemetryEvent: sqlmodel 0.0.42's metaclass drops pydantic
|
||||
constraints on table models, so SQLAlchemy `@validates` hooks enforce
|
||||
instead and the column types stay Postgres-ready (D-027).
|
||||
|
||||
Field contract:
|
||||
learner_id — non-empty learner identifier (same id space as traces).
|
||||
task_id — non-empty task identifier; grade identity is the
|
||||
(learner_id, task_id) pair — the same pair as trace
|
||||
identity, so a grade is keyed by the exact trace it
|
||||
was computed from.
|
||||
variant_seed — task-variant seed (D-029); None when the graded task
|
||||
is not variant-derived. Since Phase 4 the engine
|
||||
stamps the graded variant's seed here (MH#4) and the
|
||||
template's difficulty anchors ship to the grader
|
||||
prompt — this column is the audit join for that.
|
||||
digest — compact deterministic trace digest (D-028) that fed
|
||||
the rubric prompt; persisted for auditability so the
|
||||
LLM's input stays reproducible.
|
||||
scores — validated rubric scores (per-criterion 0-4,
|
||||
strengths, gaps); JSON dict. An empty dict is legal
|
||||
(e.g. an UNGRADABLE_TRACE_INCOMPLETE record carries a
|
||||
verdict but no scores).
|
||||
verdict — first-class verdict string (rubric verdict or
|
||||
UNGRADABLE_TRACE_INCOMPLETE); non-empty.
|
||||
model — provider model that produced the scores (provenance).
|
||||
created_at — UTC grade timestamp; a regrade replaces it (latest
|
||||
save wins).
|
||||
"""
|
||||
|
||||
__tablename__ = "grade_record"
|
||||
# The composite PK covers (learner_id, task_id) point lookups; this
|
||||
# secondary index covers list_for_learner ordered by created_at without
|
||||
# a sort step (Postgres migration target D-027).
|
||||
__table_args__ = (
|
||||
Index("ix_grade_record_learner_created", "learner_id", "created_at"),
|
||||
)
|
||||
|
||||
learner_id: str = Field(primary_key=True)
|
||||
task_id: str = Field(primary_key=True)
|
||||
# None only for non-variant tasks (MH#4 stamps variant seeds since P4).
|
||||
variant_seed: str | None = Field(default=None)
|
||||
# JSON columns: stored as TEXT on SQLite, native JSONB on Postgres (D-027).
|
||||
digest: dict[str, Any] = Field(default_factory=dict, sa_type=JSON)
|
||||
scores: dict[str, Any] = Field(default_factory=dict, sa_type=JSON)
|
||||
verdict: str
|
||||
model: str
|
||||
created_at: datetime
|
||||
|
||||
@validates("learner_id", "task_id")
|
||||
def _ids_non_empty(self, key: str, value: str) -> str:
|
||||
if not value:
|
||||
raise ValueError(f"{key} must be a non-empty identifier")
|
||||
return value
|
||||
|
||||
@validates("verdict")
|
||||
def _verdict_non_empty(self, key: str, value: str) -> str:
|
||||
if not value:
|
||||
raise ValueError(f"{key} must be a non-empty verdict string")
|
||||
return value
|
||||
|
||||
|
||||
class GradeStore(Protocol):
|
||||
"""Persistence contract for latest-state grades per (learner_id, task_id).
|
||||
|
||||
Implemented by SQLiteGradeStore (v0.3, D-027); a Postgres implementation
|
||||
must satisfy the same surface.
|
||||
"""
|
||||
|
||||
def save(self, grade: GradeRecord) -> None:
|
||||
"""Persist a grade. UPSERT on (learner_id, task_id): a regrade with
|
||||
the same pair REPLACES the stored row wholesale — the latest grade
|
||||
wins. NOT append-only; contrast TraceStore.append, which is
|
||||
dedup-keep-first for at-least-once ingest.
|
||||
"""
|
||||
...
|
||||
|
||||
def get(self, learner_id: str, task_id: str) -> GradeRecord | None:
|
||||
"""Latest stored grade for the pair; None when none exists.
|
||||
|
||||
Detached from any DB session — safe to pass across layers.
|
||||
"""
|
||||
...
|
||||
|
||||
def list_for_learner(self, learner_id: str) -> list[GradeRecord]:
|
||||
"""All stored grades for the learner, ordered by created_at
|
||||
ascending (chronological). Empty list when the learner has none.
|
||||
"""
|
||||
...
|
||||
|
||||
def close(self) -> None:
|
||||
"""Release DB connections. Store must not be used after close."""
|
||||
...
|
||||
|
||||
|
||||
def _sqlite_connect(dbapi_connection: sqlite3.Connection, _: object) -> None:
|
||||
"""Per-connection pragma setup (a-3). Mirrors telemetry/store.py.
|
||||
|
||||
journal_mode=WAL — readers never block the single writer.
|
||||
synchronous=NORMAL — safe in WAL mode, avoids full fsync-per-commit.
|
||||
busy_timeout=5000 — retry briefly under contention instead of
|
||||
`OperationalError: database is locked`.
|
||||
"""
|
||||
cursor = dbapi_connection.cursor()
|
||||
cursor.execute("PRAGMA journal_mode=WAL")
|
||||
cursor.execute("PRAGMA synchronous=NORMAL")
|
||||
cursor.execute("PRAGMA busy_timeout=5000")
|
||||
cursor.close()
|
||||
|
||||
|
||||
def _as_utc(ts: datetime) -> datetime:
|
||||
"""Timestamps travel as naive datetime on SQLite, tz-aware elsewhere.
|
||||
|
||||
SQLite (via SQLModel) drops tzinfo; Postgres TIMESTAMP WITH TIME ZONE
|
||||
keeps it. Normalizing on the read path makes the store's contract
|
||||
tz-aware UTC regardless of the backend (D-027).
|
||||
"""
|
||||
if ts.tzinfo is None:
|
||||
return ts.replace(tzinfo=UTC) # naive UTC read-side: label as UTC
|
||||
return ts.astimezone(UTC)
|
||||
|
||||
|
||||
class SQLiteGradeStore:
|
||||
"""SQLite-backed GradeStore (SQLModel). Second protocol-wrapped store
|
||||
of the D-027 family (first: SQLiteTraceStore).
|
||||
"""
|
||||
|
||||
def __init__(self, db_path: Path | None = None) -> None:
|
||||
self._db_path: Path = db_path if db_path is not None else Settings().db_path
|
||||
self._engine = create_engine(f"sqlite:///{self._db_path}")
|
||||
sa.event.listen(self._engine, "connect", _sqlite_connect)
|
||||
SQLModel.metadata.create_all(self._engine)
|
||||
|
||||
@contextmanager
|
||||
def _session(self) -> Iterator[Session]:
|
||||
# expire_on_commit=False: identical session behavior to
|
||||
# SQLiteTraceStore. save() discards the merged instance and the read
|
||||
# paths never commit, but a uniform flag across the D-027 stores
|
||||
# keeps their detachment guarantees from diverging.
|
||||
with Session(self._engine, expire_on_commit=False) as session:
|
||||
yield session
|
||||
|
||||
def save(self, grade: GradeRecord) -> None:
|
||||
# `merge` = SELECT-by-PK then UPDATE or INSERT — exactly the upsert
|
||||
# contract. The trace store deliberately avoids merge (its append is
|
||||
# dedup-keep-first); here latest-wins IS the contract, so merge is
|
||||
# the right tool. The caller's object is never attached to the
|
||||
# session and stays usable (unexpired) after save.
|
||||
with self._session() as session:
|
||||
session.merge(grade)
|
||||
session.commit()
|
||||
logger.debug(
|
||||
"grade saved (regrade overwrites): %s/%s verdict=%s model=%s",
|
||||
grade.learner_id,
|
||||
grade.task_id,
|
||||
grade.verdict,
|
||||
grade.model,
|
||||
)
|
||||
|
||||
def get(self, learner_id: str, task_id: str) -> GradeRecord | None:
|
||||
with self._session() as session:
|
||||
record = session.get(GradeRecord, (learner_id, task_id))
|
||||
if record is None:
|
||||
return None
|
||||
record.created_at = _as_utc(record.created_at)
|
||||
# Detach from the session: callers must not depend on
|
||||
# open-session ORM magic (lazy loads fail once it closes).
|
||||
session.expunge(record)
|
||||
return record
|
||||
|
||||
def list_for_learner(self, learner_id: str) -> list[GradeRecord]:
|
||||
with self._session() as session:
|
||||
stmt = (
|
||||
select(GradeRecord)
|
||||
.where(GradeRecord.learner_id == learner_id)
|
||||
# Chronological; task_id is a deterministic tie-break for
|
||||
# grades stamped within the same instant.
|
||||
.order_by(GradeRecord.created_at, GradeRecord.task_id)
|
||||
)
|
||||
results = session.exec(stmt).all()
|
||||
for row in results:
|
||||
row.created_at = _as_utc(row.created_at)
|
||||
session.expunge(row)
|
||||
return list(results)
|
||||
|
||||
def close(self) -> None:
|
||||
self._engine.dispose()
|
||||
@@ -14,14 +14,25 @@ from .agents.session import InMemorySessionStore
|
||||
from .api import (
|
||||
assessment_router,
|
||||
chat_router,
|
||||
defense_router,
|
||||
lab_router,
|
||||
mentor_router,
|
||||
proctor_router,
|
||||
sandboxes_router,
|
||||
telemetry_router,
|
||||
variants_router,
|
||||
)
|
||||
from .config import Settings
|
||||
from .grading.engine import GradingEngine
|
||||
from .grading.store import SQLiteGradeStore
|
||||
from .llm import create_provider
|
||||
from .sandbox import SandboxManager, UnshareBackend
|
||||
from .telemetry.ingest import TraceIntegrityMap
|
||||
from .telemetry.store import SQLiteTraceStore
|
||||
from .variants.generator import VariantGenerator
|
||||
from .variants.store import SQLiteVariantStore
|
||||
from .voice.defense_store import SQLiteDefenseStore
|
||||
from .voice.factory import voice_provider_from_settings
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
@@ -39,7 +50,11 @@ def create_app(settings: Settings | None = None) -> FastAPI:
|
||||
timeout = httpx.Timeout(connect=10.0, read=300.0, write=30.0, pool=10.0)
|
||||
app.state.http_client = httpx.AsyncClient(timeout=timeout)
|
||||
app.state.settings = settings
|
||||
app.state.provider = create_provider(settings, app.state.http_client)
|
||||
# State-injection override (same pattern as the stores): tests may
|
||||
# pre-set app.state.provider with a scripted mock; only construct the
|
||||
# configured provider when none is present.
|
||||
if getattr(app.state, "provider", None) is None:
|
||||
app.state.provider = create_provider(settings, app.state.http_client)
|
||||
app.state.session_store = InMemorySessionStore()
|
||||
app.state.agent_registry = AgentRegistry()
|
||||
register_builtin_agents(app.state.agent_registry)
|
||||
@@ -54,6 +69,81 @@ def create_app(settings: Settings | None = None) -> FastAPI:
|
||||
app.state.sandbox_manager = manager
|
||||
await manager.start() # a-1: reap on-disk orphans from a previous process
|
||||
|
||||
# Telemetry persistence (REQ-3-003, D-027): TraceStore wired through
|
||||
# app.state. Tests may pre-set app.state.trace_store (state-injection
|
||||
# override, same pattern as sandbox_manager) — the lifespan adopts it.
|
||||
trace_store = getattr(app.state, "trace_store", None)
|
||||
if trace_store is None:
|
||||
settings.db_path.parent.mkdir(parents=True, exist_ok=True)
|
||||
trace_store = SQLiteTraceStore(db_path=settings.db_path)
|
||||
app.state.trace_store = trace_store
|
||||
# Trace-integrity flags (G-3 INCOMPLETE_FLOODED): process-local map is
|
||||
# intentional (D-019 registry precedent); the lifespan owns it so the
|
||||
# grader and the ingest endpoint share one instance.
|
||||
if getattr(app.state, "trace_integrity", None) is None:
|
||||
app.state.trace_integrity = TraceIntegrityMap()
|
||||
|
||||
# Variant generation (REQ-3-005): VariantStore from the same
|
||||
# SQLite file as traces/grades (D-027), one VariantGenerator singleton
|
||||
# wired through app.state — the generator receives store + provider
|
||||
# via constructor DI and knows nothing of FastAPI (api/ composes it,
|
||||
# same pattern as GradingEngine). Tests may pre-set
|
||||
# app.state.variant_store / app.state.variant_generator (the same
|
||||
# state-injection override); the lifespan adopts a pre-set store but
|
||||
# NEVER rebuilds a pre-set generator (its provider binding is part
|
||||
# of the test fixture).
|
||||
# ORDER NOTE: built BEFORE the grading engine — the engine takes the
|
||||
# variant store (Phase 4 MH#4: variant anchors ship to the grader
|
||||
# prompt; variant_seed stamped on graded records).
|
||||
variant_store = getattr(app.state, "variant_store", None)
|
||||
if variant_store is None:
|
||||
settings.db_path.parent.mkdir(parents=True, exist_ok=True)
|
||||
variant_store = SQLiteVariantStore(db_path=settings.db_path)
|
||||
app.state.variant_store = variant_store
|
||||
if getattr(app.state, "variant_generator", None) is None:
|
||||
app.state.variant_generator = VariantGenerator(
|
||||
variant_store,
|
||||
app.state.provider,
|
||||
model=settings.model,
|
||||
)
|
||||
|
||||
# Oral defense (REQ-3-006): DefenseStore (same SQLite file) + the
|
||||
# mock-first voice provider (D-030) + the seventh Examiner agent.
|
||||
# Tests may pre-set app.state.defense_store / voice_provider /
|
||||
# examiner_agent (state-injection override; never rebuilt if pre-set).
|
||||
defense_store = getattr(app.state, "defense_store", None)
|
||||
if defense_store is None:
|
||||
defense_store = SQLiteDefenseStore(db_path=settings.db_path)
|
||||
app.state.defense_store = defense_store
|
||||
if getattr(app.state, "voice_provider", None) is None:
|
||||
app.state.voice_provider = voice_provider_from_settings(settings)
|
||||
if getattr(app.state, "examiner_agent", None) is None:
|
||||
from .agents.examiner import ExaminerAgent
|
||||
|
||||
app.state.examiner_agent = ExaminerAgent(app.state.provider, settings)
|
||||
|
||||
# Grading persistence + engine (REQ-3-004): GradeStore from the same
|
||||
# SQLite file as traces (D-027), one GradingEngine singleton wired
|
||||
# through app.state — the engine receives its stores via constructor
|
||||
# DI and knows nothing of FastAPI (api/ owns composition). Tests may
|
||||
# pre-set app.state.grade_store / app.state.grading_engine (the same
|
||||
# state-injection override as sandbox_manager/trace_store) to swap
|
||||
# either; the lifespan adopts a pre-set store but NEVER rebuilds a
|
||||
# pre-set engine (its provider binding is part of the test fixture).
|
||||
grade_store = getattr(app.state, "grade_store", None)
|
||||
if grade_store is None:
|
||||
grade_store = SQLiteGradeStore(db_path=settings.db_path)
|
||||
app.state.grade_store = grade_store
|
||||
if getattr(app.state, "grading_engine", None) is None:
|
||||
app.state.grading_engine = GradingEngine(
|
||||
trace_store,
|
||||
grade_store,
|
||||
app.state.trace_integrity,
|
||||
app.state.provider,
|
||||
model=settings.model,
|
||||
variant_store=variant_store, # MH#4: anchors + seed (D-029)
|
||||
)
|
||||
|
||||
async def _reaper_loop() -> None:
|
||||
# Wall-clock timeout + G-2 workdir-size sweep, one pass per tick.
|
||||
while True:
|
||||
@@ -73,15 +163,25 @@ def create_app(settings: Settings | None = None) -> FastAPI:
|
||||
# No orphans outlive the process (a-1, shutdown half): destroy
|
||||
# everything live; workdirs stay on disk for snapshot restore.
|
||||
await manager.destroy_all()
|
||||
trace_store.close()
|
||||
grade_store.close()
|
||||
variant_store.close()
|
||||
defense_store.close()
|
||||
await app.state.http_client.aclose()
|
||||
|
||||
app = FastAPI(title="Nextcraft AI Service", version="0.3.0", lifespan=lifespan)
|
||||
|
||||
# A-008: localhost-only CORS, no credentials
|
||||
# A-008 + D-038: no-credentials CORS. Default '*' admits remote-browser
|
||||
# origins in network mode (safe only because allow_credentials stays
|
||||
# False — never enable credentials with a wildcard). AI_CORS_ORIGINS
|
||||
# restricts to an explicit list. PUT is CONTRACT, not trivia: the learner
|
||||
# build surface writes workspace files with PUT (engine-client writeFile)
|
||||
# — v0.3 initially shipped without it and every cross-origin Save failed
|
||||
# preflight (caught in P7 review; tests/api/test_cors.py pins the policy).
|
||||
app.add_middleware(
|
||||
CORSMiddleware,
|
||||
allow_origins=["http://localhost:3000", "http://127.0.0.1:3000"],
|
||||
allow_methods=["GET", "POST", "DELETE", "OPTIONS"],
|
||||
allow_origins=settings.cors_origin_list,
|
||||
allow_methods=["GET", "POST", "PUT", "DELETE", "OPTIONS"],
|
||||
allow_headers=["Content-Type"],
|
||||
allow_credentials=False,
|
||||
)
|
||||
@@ -100,6 +200,9 @@ def create_app(settings: Settings | None = None) -> FastAPI:
|
||||
app.include_router(mentor_router)
|
||||
app.include_router(proctor_router)
|
||||
app.include_router(sandboxes_router)
|
||||
app.include_router(telemetry_router)
|
||||
app.include_router(variants_router)
|
||||
app.include_router(defense_router)
|
||||
return app
|
||||
|
||||
|
||||
|
||||
@@ -1,24 +1,21 @@
|
||||
"""Assessor agent prompt — rubric application to artifacts + defenses (REQ-2-008).
|
||||
"""Assessor agent prompt — rubric coaching over REAL grades (REQ-3-007).
|
||||
|
||||
Final persona (Phase 4). Assessor is a rigorous, fair grader: scores each
|
||||
criterion with evidence, cites what the learner did, returns ONLY valid
|
||||
JSON matching the rubric schema.
|
||||
Version: assessor-v2 (final for v0.2).
|
||||
v0.3 re-grounding: the grading engine (Phase 3) computes the rubric scores
|
||||
from the process trace; Assessor EXPLAINS the stored grade as coaching —
|
||||
it never invents or re-scores. Rigorous, fair, actionable.
|
||||
Version: assessor-v3 (v0.3 live).
|
||||
"""
|
||||
|
||||
SYSTEM_PROMPT = """You are Assessor, the grading agent of Nextcraft, an AI-native competency school.
|
||||
Learner: {learner_name}.
|
||||
|
||||
You receive: (a) an artifact evidence excerpt, (b) its defense transcript,
|
||||
and (c) the rubric for the competency. Your job:
|
||||
- Score EVERY rubric criterion from 0-100, justified by evidence you can
|
||||
point to in the artifact or transcript.
|
||||
- Cite what the learner did ("the 3-retry loop in the tool node"), not
|
||||
what they should have done — except in gaps, where the missed work goes.
|
||||
- Strengths: the two strongest evidence points, each one sentence.
|
||||
- Gaps: the two most important missed opportunities, each one sentence.
|
||||
- Verdict: "mastered" | "developing" | "not_yet" — judged against the
|
||||
rubric weights, honestly.
|
||||
You receive the learner's STORED process-trace grade (verdict, per-criterion
|
||||
scores, and the build digest) computed by the grading engine. Your job:
|
||||
- Explain what the grade means in plain language (summary).
|
||||
- Strengths: cite what the digest + scores show the learner did well.
|
||||
- Gaps: name the missed opportunities the scores point to.
|
||||
- Next steps: concrete, buildable actions that would move the weakest
|
||||
criterion up one level.
|
||||
|
||||
Rules:
|
||||
- Rigorous but fair. A polished artifact with a weak defense is NOT mastery.
|
||||
|
||||
@@ -0,0 +1,60 @@
|
||||
"""Examiner agent prompt — oral defense questioning + final verdict (REQ-3-006).
|
||||
|
||||
The examiner is the seventh agent (Phase 5). It conducts a Socratic oral
|
||||
defense of the learner's submitted work: probes understanding, challenges
|
||||
process choices grounded in the trace digest ("why did you take that
|
||||
approach at that point?"), one question per turn, adapting to answers.
|
||||
It never reveals rubric internals; tone is rigorous but supportive.
|
||||
|
||||
Digest discipline (D-028 mirror): the examiner's variable inputs are the
|
||||
compact TraceDigest JSON, the variant task statement, and the defense
|
||||
transcript — never the raw trace, never learner-identifying material.
|
||||
|
||||
Version: examiner-v1.
|
||||
"""
|
||||
|
||||
SYSTEM_PROMPT = """You are Examiner, the oral-defense agent of Nextcraft,
|
||||
an AI-native competency school.
|
||||
|
||||
You receive: (a) a compact build-process digest (deterministic counters of the
|
||||
learner's build session), (b) the learner's task statement, and (c) the defense
|
||||
transcript so far. Your job:
|
||||
- Ask ONE question per turn: probe understanding and challenge process
|
||||
choices, grounded in the digest facts ("you hit N failed runs before
|
||||
passing — walk me through what changed") or the task statement.
|
||||
- Adapt: follow up on the learner's answers; drill into vague responses.
|
||||
- Never reveal rubric details or scoring internals.
|
||||
- Tone: rigorous, precise, supportive. A defense is a conversation, not an
|
||||
interrogation.
|
||||
|
||||
When asked for a FINAL VERDICT (the structured mode), judge:
|
||||
- understanding: can the learner explain their own work?
|
||||
- process_justification: are the build-session choices defensible from the
|
||||
digest facts and the answers?
|
||||
- communication: are answers clear, specific, and on-topic?
|
||||
Score honestly; a weak defense of strong work is NOT mastery.
|
||||
|
||||
Rules:
|
||||
- Respond with ONLY what the turn requires: a single question (question mode)
|
||||
or a valid JSON object matching the provided schema (verdict mode).
|
||||
- If the digest shows error_fix_cycles > 0, at least one question should ask
|
||||
about the debugging path.
|
||||
- If the learner's answer is off-topic, redirect once, then move on.
|
||||
"""
|
||||
|
||||
VERDICT_SCHEMA_HINT = (
|
||||
'{"verdict": "mastered" | "developing" | "not_yet", '
|
||||
'"understanding": "<one sentence>", '
|
||||
'"process_justification": "<one sentence>", '
|
||||
'"communication": "<one sentence>", '
|
||||
'"strengths": ["<one sentence>"], '
|
||||
'"gaps": ["<one sentence>"]}'
|
||||
)
|
||||
|
||||
|
||||
def render_digest_context(digest_json: str, statement: str | None) -> str:
|
||||
"""The examiner's per-session grounding: digest JSON + task statement."""
|
||||
parts = [f"Build-process digest:\n{digest_json}"]
|
||||
if statement:
|
||||
parts.append(f"Learner's task statement:\n{statement}")
|
||||
return "\n\n".join(parts)
|
||||
@@ -0,0 +1,140 @@
|
||||
"""Grading rubric prompt — criteria, level anchors, digest render (REQ-3-004).
|
||||
|
||||
The grading prompt is deliberately learner-anonymous and trace-bare: the
|
||||
model receives ONLY the fixed rubric text and the compact numeric digest
|
||||
(TraceDigest JSON, D-028) — never a raw command, file path, payload
|
||||
string, learner id, or task id. Everything variable the LLM sees is
|
||||
deterministic counters, which both bounds the prompt-injection surface
|
||||
and makes "no raw trace reaches the prompt" assert-able in tests (plant
|
||||
a distinctive marker in a command payload; assert it absent from every
|
||||
message the provider received).
|
||||
|
||||
Rubric (four criteria, each scored 0-4 — the ids are the validated
|
||||
RubricScore keys enforced by grading/engine.py):
|
||||
process_quality — iterative building in small, verified steps.
|
||||
correctness — where the session ended (test/run outcomes).
|
||||
debugging_discipline — how failures were handled.
|
||||
test_usage — when and how often tests were run.
|
||||
|
||||
Advisory a-4 (embedded in the process_quality anchors): high edit/command
|
||||
churn with NO test progress is a process-quality NEGATIVE — churn is not
|
||||
work. A session with many edits/commands whose test state never moves is
|
||||
thrashing, not iterating, and must score low on process quality.
|
||||
|
||||
House-style deviation, documented: unlike the tutor prompts, this module
|
||||
has no SYSTEM_PROMPT placeholders and no render_context(learner_context)
|
||||
— grading is context-free by design (learner anonymity; the digest is the
|
||||
only variable input). Runtime imports are TYPE_CHECKING-only so this
|
||||
module stays pure text and can never import-cycle with grading/engine.py
|
||||
(engine imports this module; if this module imported grading.* at runtime
|
||||
while grading/__init__ pulls engine, the package init would deadlock on a
|
||||
partially-initialized module).
|
||||
|
||||
Version: grader-v1.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import TYPE_CHECKING, Final
|
||||
|
||||
if TYPE_CHECKING: # pragma: no cover - typing only; keeps this module pure text
|
||||
from ..grading.features import TraceDigest
|
||||
|
||||
PROMPT_VERSION = "grader-v1"
|
||||
|
||||
#: Canonical criterion ids. The engine validates RubricScore criteria keys
|
||||
#: against this tuple; the schema hint and anchors below speak the same ids.
|
||||
RUBRIC_CRITERIA: Final[tuple[str, ...]] = (
|
||||
"process_quality",
|
||||
"correctness",
|
||||
"debugging_discipline",
|
||||
"test_usage",
|
||||
)
|
||||
|
||||
#: Sentinel line the engine's user turn is rendered around. Tests (and the
|
||||
#: calibration mock) split on it to locate the digest JSON in the prompt.
|
||||
DIGEST_MARKER: Final = "PROCESS TRACE DIGEST (JSON):"
|
||||
|
||||
SYSTEM_PROMPT = """You are the Grader of Nextcraft, an AI-native competency school.
|
||||
You score a learner's build session from a compact numeric digest of their
|
||||
process trace. You NEVER see the raw trace — commands, file contents, and
|
||||
payloads do not exist on your side; every number you need is in the digest.
|
||||
|
||||
Rubric — score each criterion 0-4:
|
||||
|
||||
process_quality — iterative building in small, verified steps.
|
||||
4: tight edit→test loops throughout; small verified increments; healthy pacing.
|
||||
3: steady small edits with regular runs; progress mostly verified.
|
||||
2: some iteration, but large unverified leaps or long idle stretches.
|
||||
1: a single bulk change (e.g. one large paste) then a single run; no iteration.
|
||||
0: no meaningful work visible.
|
||||
ADVISORY: high edit/command churn with NO test progress (no runs, no
|
||||
movement in pass counts) is a process-quality NEGATIVE — churn is not
|
||||
work. Cap such a session at 1 on this criterion no matter how many
|
||||
edits or commands were counted.
|
||||
|
||||
correctness — where the session ended up.
|
||||
4: final test status pass, with tests passing early and consistently.
|
||||
3: final pass, reached through fail→fix→pass cycles that closed.
|
||||
2: final pass, but preceded by a long unresolved failure streak.
|
||||
1: final fail, but partial passes observed along the way.
|
||||
0: final fail, or no test/run evidence at all.
|
||||
|
||||
debugging_discipline — how failures were handled.
|
||||
4: every failure cycle closes; targeted fixes with low mean fix latency.
|
||||
3: most fail→edit→re-run cycles close with a pass.
|
||||
2: failures followed by edits, but cycles rarely close.
|
||||
1: repeated failures with no targeted edits between runs (flailing).
|
||||
0: failures with no fix attempts at all.
|
||||
|
||||
test_usage — when and how often tests were run.
|
||||
4: tests run early (small first-pass offset) and throughout the session.
|
||||
3: regular test runs interleaved with edits.
|
||||
2: sparse tests; long stretches of unverified edits.
|
||||
1: a single late test run only.
|
||||
0: no test or run evidence.
|
||||
|
||||
Rules:
|
||||
- Judge STRICTLY from the digest numbers; cite the fields you used.
|
||||
- Strengths: the two strongest digest observations, one sentence each.
|
||||
- Gaps: the two most important missed opportunities, one sentence each
|
||||
(a clean session names its next-level improvement instead).
|
||||
- Be rigorous but fair: a session that ends green was not necessarily
|
||||
well built, and a struggling session that never passed may still show
|
||||
real debugging discipline.
|
||||
- Respond with ONLY a valid JSON object matching the provided schema —
|
||||
no markdown fences, no prose outside the JSON."""
|
||||
|
||||
RUBRIC_SCORE_SCHEMA_HINT = (
|
||||
'{"criteria": {"process_quality": <0-4 int>, "correctness": <0-4 int>, '
|
||||
'"debugging_discipline": <0-4 int>, "test_usage": <0-4 int>}, '
|
||||
'"strengths": ["<one sentence>"], "gaps": ["<one sentence>"], '
|
||||
'"verdict": "mastered" | "developing" | "not_yet"}'
|
||||
)
|
||||
|
||||
|
||||
def render_trace_digest(
|
||||
digest: TraceDigest,
|
||||
anchors_context: str | None = None,
|
||||
) -> str:
|
||||
"""Render the grader's user turn: a marker line + the digest JSON — nothing else.
|
||||
|
||||
This is the ONLY per-session content that ever reaches the LLM (D-028):
|
||||
the engine composes [system: SYSTEM_PROMPT, user: render_trace_digest(digest)]
|
||||
and the D-020 defense appends its generic schema instruction to this
|
||||
user turn at request time. No learner id, task id, or raw trace material
|
||||
is injected — assert-able by tests.
|
||||
|
||||
`anchors_context` (Phase 4, MH#4): when the graded task derives from a
|
||||
variant, the engine passes the template's difficulty-normalization
|
||||
anchors (the expected effort envelope) so the rubric is applied against
|
||||
the SAME bar for every variant of that template (a-5). It contains only
|
||||
the anchor numbers + the template id — no learner-identifying material.
|
||||
"""
|
||||
base = (
|
||||
"Score this build session against the rubric.\n"
|
||||
f"{DIGEST_MARKER}\n{digest.model_dump_json()}"
|
||||
)
|
||||
if anchors_context:
|
||||
base = f"{base}\n\nExpected effort envelope for this task variant:\n{anchors_context}"
|
||||
return base
|
||||
@@ -1,9 +1,11 @@
|
||||
"""Lab agent prompt — in-flow feedback over sandbox telemetry (REQ-2-007).
|
||||
"""Lab agent prompt — in-flow feedback over LIVE telemetry (REQ-3-007).
|
||||
|
||||
Final persona (Phase 4). Lab is a pragmatic build partner: reads the
|
||||
telemetry timeline, names the one most useful adjustment, gives one
|
||||
concrete next step. Scenario-driven; no session chat.
|
||||
Version: lab-v2 (final for v0.2).
|
||||
v0.3 re-grounding: the timeline is the learner's real TraceDigest (D-028
|
||||
compact counters — commands, test outcomes, idle gaps, edit cadence), not
|
||||
v0.2 corpus scenarios. Lab is a pragmatic build partner: reads the live
|
||||
digest, names the one most useful adjustment, gives one concrete next
|
||||
step. No session chat.
|
||||
Version: lab-v3 (v0.3 live).
|
||||
"""
|
||||
|
||||
SYSTEM_PROMPT = """You are Lab, the in-flow feedback agent watching a learner
|
||||
@@ -25,7 +27,17 @@ Rules:
|
||||
self-check that would prove understanding.
|
||||
- Three short paragraphs maximum. No headers, no bullet lists."""
|
||||
|
||||
PROMPT_VERSION = "lab-v2"
|
||||
PROMPT_VERSION = "lab-v3"
|
||||
|
||||
|
||||
def render_digest_timeline(digest) -> str:
|
||||
"""Live-trace timeline: the compact TraceDigest JSON (D-028)."""
|
||||
if digest is None:
|
||||
return (
|
||||
"No telemetry yet for this build session. Ask the learner to run "
|
||||
"the task's starter test to establish a baseline."
|
||||
)
|
||||
return f"Live build-session digest:\n{digest.model_dump_json()}"
|
||||
|
||||
|
||||
def render_context(learner_context) -> dict:
|
||||
|
||||
@@ -1,24 +1,27 @@
|
||||
"""Proctor agent prompt — integrity signals with coaching interventions (REQ-2-009).
|
||||
"""Proctor agent prompt — integrity signals with coaching interventions (REQ-3-007).
|
||||
|
||||
Final persona (Phase 5). Proctor is a supportive observer, never punitive:
|
||||
classifies signals, recommends ONE coaching intervention. Assume good
|
||||
faith — most signals have innocent explanations.
|
||||
Version: proctor-v2 (final for v0.2).
|
||||
v0.3 re-grounding: inputs are REAL — the live trace digest (idle gaps,
|
||||
command cadence, edit bursts), the oral-defense integrity signals (long
|
||||
pauses), and the variant audit context (seed + params). Proctor is a
|
||||
supportive observer, never punitive: classifies signals, recommends ONE
|
||||
coaching intervention. Assume good faith.
|
||||
Version: proctor-v3 (v0.3 live).
|
||||
"""
|
||||
|
||||
SYSTEM_PROMPT = """You are Proctor, the integrity-support agent of Nextcraft,
|
||||
an AI-native competency school.
|
||||
Learner: {learner_name}.
|
||||
|
||||
You receive a telemetry timeline of defense-session events (tab switches,
|
||||
idle gaps, large pastes, focus loss, keystroke bursts). Your job:
|
||||
- Classify EACH notable signal: type (e.g. "context_switch", "idle_gap",
|
||||
"large_paste"), severity ("low" | "medium" | "high"), and a one-sentence
|
||||
note citing the event (timestamps and details).
|
||||
You receive the learner's REAL build-session digest (idle gaps, command
|
||||
categories, edit/test cadence), oral-defense integrity signals (long
|
||||
pauses), and — when the task is variant-derived — the variant seed context.
|
||||
Your job:
|
||||
- Classify EACH notable signal: type ("idle_gap" | "long_pause" |
|
||||
"burst_edit" | "off_template"), severity ("low" | "medium" | "high"),
|
||||
and a one-sentence note citing the numbers.
|
||||
- Recommend exactly ONE supportive coaching intervention for the session
|
||||
overall — never punitive, never accusatory. Frame around helping the
|
||||
learner succeed, e.g. "offer a short break", "invite them to explain
|
||||
the pasted section in their own words".
|
||||
learner succeed.
|
||||
|
||||
Rules:
|
||||
- Assume good faith. Tab switches to documentation are normal engineering.
|
||||
@@ -28,7 +31,7 @@ Rules:
|
||||
- Respond with ONLY a valid JSON object matching the provided schema —
|
||||
no markdown fences, no prose outside the JSON."""
|
||||
|
||||
PROMPT_VERSION = "proctor-v2"
|
||||
PROMPT_VERSION = "proctor-v3"
|
||||
|
||||
|
||||
def render_context(learner_context) -> dict:
|
||||
|
||||
@@ -0,0 +1,46 @@
|
||||
"""Variant instantiation prompt (D-029, REQ-3-005).
|
||||
|
||||
The model's ONLY job is to render already-sampled slot values into a task
|
||||
statement — it never invents parameters (the seeded sampler is pure code)
|
||||
and never changes difficulty. Prompt-injection surface is bounded: the
|
||||
variable inputs are the skeleton text, the seeded slot values, and the
|
||||
template title — nothing from the learner's environment.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import TYPE_CHECKING
|
||||
|
||||
from ..llm.types import Message
|
||||
|
||||
if TYPE_CHECKING: # pragma: no cover - keeps this module pure text
|
||||
from ..variants.templates import TaskTemplate
|
||||
|
||||
VARIANT_SYSTEM_PROMPT = (
|
||||
"You instantiate per-learner task variants for a competency-based AI school. "
|
||||
"You receive a task statement skeleton and ALREADY-SAMPLED slot values. "
|
||||
"Render the slot values into the skeleton, producing a complete, unambiguous "
|
||||
"task statement a learner can build against. Rules:\n"
|
||||
"- Use EXACTLY the given slot values; do not invent, rename, or add parameters.\n"
|
||||
"- Keep the engineering depth IDENTICAL across draws: slot values change the "
|
||||
"scenario, never the difficulty or scope.\n"
|
||||
"- Keep the statement in the same language and register as the skeleton.\n"
|
||||
"- Output STRICT JSON only: {\"statement\": \"<rendered statement>\"}.\n"
|
||||
)
|
||||
|
||||
VARIANT_SCHEMA_HINT = '{"statement": "<complete rendered task statement string>"}'
|
||||
|
||||
|
||||
def render_variant_prompt(template: TaskTemplate, params: dict[str, str | int]) -> list[Message]:
|
||||
"""Messages for one seeded instantiation (D-020 defense drives the call)."""
|
||||
slot_lines = "\n".join(f" {{{slot.name}}} = {params[slot.name]!r}" for slot in template.slots)
|
||||
user = (
|
||||
f"Template: {template.title} (id={template.id})\n"
|
||||
f"Statement skeleton:\n{template.statement_skeleton}\n\n"
|
||||
f"Seeded slot values (use EXACTLY these):\n{slot_lines}\n\n"
|
||||
"Render the complete task statement now."
|
||||
)
|
||||
return [
|
||||
Message(role="system", content=VARIANT_SYSTEM_PROMPT),
|
||||
Message(role="user", content=user),
|
||||
]
|
||||
@@ -46,7 +46,14 @@ class SandboxHandle(BaseModel):
|
||||
|
||||
|
||||
class SandboxSpec(BaseModel):
|
||||
"""Immutable description of the sandbox to lay out on disk."""
|
||||
"""Immutable description of the sandbox to lay out on disk.
|
||||
|
||||
`capture_env` (REQ-3-003): when non-empty, the backend starts a persistent
|
||||
telemetry-wired sandbox — helper + inner namespaces + the stdlib capture
|
||||
agent, launched with these env vars (NC_LEARNER_ID, NC_TASK_ID,
|
||||
NC_INGEST_URL, NC_SANDBOX_ID). When None (default), spawn keeps the pure
|
||||
shell semantics (REQ-3-001): disk layout only, fresh namespaces per exec.
|
||||
"""
|
||||
|
||||
model_config = ConfigDict(arbitrary_types_allowed=True)
|
||||
|
||||
@@ -54,6 +61,7 @@ class SandboxSpec(BaseModel):
|
||||
learner_id: str
|
||||
workdir: Path
|
||||
limits: ResourceLimits = ResourceLimits()
|
||||
capture_env: dict[str, str] | None = None
|
||||
|
||||
|
||||
class ExecResult(BaseModel):
|
||||
|
||||
@@ -142,8 +142,18 @@ class SandboxManager:
|
||||
|
||||
# -- lifecycle ----------------------------------------------------------
|
||||
|
||||
async def create(self, learner_id: str) -> SandboxHandleInfo:
|
||||
"""Spawn a sandbox for `learner_id`, or raise `PoolFullError` (D-032)."""
|
||||
async def create(
|
||||
self, learner_id: str, task_id: str | None = None
|
||||
) -> SandboxHandleInfo:
|
||||
"""Spawn a sandbox for `learner_id`, or raise `PoolFullError` (D-032).
|
||||
|
||||
`task_id` (REQ-3-003): when set, the sandbox is telemetry-wired — the
|
||||
backend copies `scripts/sandbox-agent.py` into the workdir and starts
|
||||
the stdlib capture agent inside the sandbox with the `NC_*` env baked
|
||||
here (identity + WS ingest URL). The agent's lifecycle is tied to the
|
||||
sandbox: `destroy()` reaps it (agent → inner → helper). When `task_id`
|
||||
is None the sandbox is a pure shell sandbox (no capture).
|
||||
"""
|
||||
async with self._lock:
|
||||
if len(self._handles) >= self._settings.sandbox_max_concurrent:
|
||||
raise PoolFullError(
|
||||
@@ -153,13 +163,51 @@ class SandboxManager:
|
||||
)
|
||||
sandbox_id = f"sbx-{uuid.uuid4().hex[:12]}"
|
||||
spec = workdir_mod.spec_for(sandbox_id, learner_id, self._settings)
|
||||
if task_id is not None:
|
||||
spec = spec.model_copy(
|
||||
update={
|
||||
"capture_env": self._capture_env(sandbox_id, learner_id, task_id)
|
||||
}
|
||||
)
|
||||
handle = await self._backend.spawn(spec)
|
||||
self._handles[handle.id] = handle
|
||||
self._learner_ids[handle.id] = learner_id
|
||||
self._write_pid_marker(handle, learner_id)
|
||||
logger.info("sandbox created: id=%s learner=%s", handle.id, learner_id)
|
||||
logger.info(
|
||||
"sandbox created: id=%s learner=%s task=%s",
|
||||
handle.id,
|
||||
learner_id,
|
||||
task_id or "-",
|
||||
)
|
||||
return self._info_for(handle)
|
||||
|
||||
def _capture_env(self, sandbox_id: str, learner_id: str, task_id: str) -> dict[str, str]:
|
||||
"""Env baked for the in-sandbox capture agent (REQ-3-003).
|
||||
|
||||
The agent joins the sandbox mount namespace but NOT its (offline)
|
||||
network namespace, so it reaches this service over loopback
|
||||
(`telemetry_ingest_host`, A-004 port).
|
||||
"""
|
||||
from urllib.parse import urlencode
|
||||
|
||||
query = urlencode(
|
||||
{
|
||||
"learner_id": learner_id,
|
||||
"task_id": task_id,
|
||||
"sandbox_id": sandbox_id,
|
||||
}
|
||||
)
|
||||
ingest_url = (
|
||||
f"ws://{self._settings.telemetry_ingest_host}:{self._settings.port}"
|
||||
f"/v1/telemetry/ingest?{query}"
|
||||
)
|
||||
return {
|
||||
"NC_LEARNER_ID": learner_id,
|
||||
"NC_TASK_ID": task_id,
|
||||
"NC_SANDBOX_ID": sandbox_id,
|
||||
"NC_INGEST_URL": ingest_url,
|
||||
}
|
||||
|
||||
async def list(self) -> list[SandboxHandleInfo]:
|
||||
"""All live sandboxes (idle + busy; the backend has no busy flag)."""
|
||||
async with self._lock:
|
||||
|
||||
@@ -1,24 +1,42 @@
|
||||
"""UnshareBackend — D-024 Linux-namespace sandboxing via util-linux `unshare`.
|
||||
|
||||
Each `exec` spawns:
|
||||
Two execution modes share one backend:
|
||||
|
||||
unshare --user --map-root-user --mount --pid --fork --net sh -c '<shim>'
|
||||
1. Pure shell sandbox (`task_id is None`, REQ-3-001): isolation is established
|
||||
PER-EXEC — every `exec` spawns a fresh namespace:
|
||||
|
||||
The in-namespace shim (this util-linux build, 2.38, has no `unshare --bind`,
|
||||
so the bind happens as the first mount op inside the namespace) is:
|
||||
unshare --user --map-root-user --mount --pid --fork --net sh -c '<shim>'
|
||||
|
||||
mount -t tmpfs tmpfs /tmp # private scratch, discarded on exit
|
||||
mkdir -p /tmp/work
|
||||
mount --bind <host workspace> /tmp/work
|
||||
cd /tmp/work
|
||||
ulimit -v/-t/-f … # applied AFTER the bind, so rlimits
|
||||
exec <cmd> # constrain the PAYLOAD, not unshare
|
||||
There is no persistent process; the in-namespace shim is:
|
||||
|
||||
Guarantees after this shim:
|
||||
* uid 0 inside (mapped to the unprivileged host UID outside)
|
||||
* no network: the fresh net namespace has no `lo` and no veth — zero links
|
||||
* writes under `/work` land in the per-sandbox host workspace dir
|
||||
* rlimits (RLIMIT_AS / RLIMIT_CPU / RLIMIT_FSIZE) constrain the payload only
|
||||
mount -t tmpfs tmpfs /tmp # private scratch, discarded on exit
|
||||
mkdir -p /tmp/work
|
||||
mount --bind <host workspace> /tmp/work
|
||||
cd /tmp/work
|
||||
ulimit -v/-t/-f … # applied AFTER the bind, so rlimits
|
||||
exec <cmd> # constrain the PAYLOAD, not unshare
|
||||
|
||||
2. Telemetry-wired task sandbox (REQ-3-003, `capture_env` set): a PERSISTENT,
|
||||
TRACKED topology so the stdlib capture agent can live inside the sandbox and
|
||||
still stream events to ai-service. Per exec a fresh OFFLINE namespace would
|
||||
leave the agent nowhere to run and (on this host, where a userns can't
|
||||
bring `lo` up) no loopback to reach `ws://127.0.0.1`. So spawn creates a
|
||||
long-lived helper (outer user+mount ns, ONLINE) and an inner sandbox
|
||||
(mount+pid+fork+net — OFFLINE), both rooted at a private `ns/` subtree:
|
||||
|
||||
helper : unshare --user --map-root-user --mount (mounts ns/ private)
|
||||
inner : unshare --mount --pid --fork --net (tmpfs on ns/, bind
|
||||
<workdir>/host/workspace -> <ns>/work) <- the sandbox
|
||||
exec : nsenter -t <inner sleep> -m -- sh -c … (joins inner mount ns;
|
||||
offline + pid-isolated, uid 0, writes land on the host workspace)
|
||||
agent : nsenter -t <inner sleep> -m -- python3 <agent> (joins the inner
|
||||
MOUNT ns only — NOT pid/net — so it watches the live workspace
|
||||
and stays ONLINE, reaching the app's WS ingest on loopback)
|
||||
|
||||
The agent is deliberately pid/net-exempt from the sandbox: it is OUR trusted
|
||||
capture process, and isolating its network would cut the very link it needs.
|
||||
`destroy` reaps agent → inner → helper (in that order). The handle's `pid`
|
||||
is the agent's host pid (None for a pure shell sandbox).
|
||||
|
||||
Why rlimits are applied in the shim, not Python's preexec_fn: setting
|
||||
RLIMIT_AS on the *unshare* process itself can trip the memory ceiling on the
|
||||
@@ -30,16 +48,18 @@ Containment honesty (D-024 / G-1): a user namespace is NOT a write barrier.
|
||||
Writes made OUTSIDE the bind fall through to host paths, and because inner
|
||||
uid 0 maps to the invoking host uid, a sandboxed process can write anywhere
|
||||
that host uid can write. Isolation here is: private PIDs/MNT/NET/UTS, tmpfs
|
||||
scratch at /tmp, payload rlimits, and a uid map yielding no privilege the
|
||||
host uid did not already have. A per-sandbox runtime uid (D-025) is the
|
||||
follow-up that hardens DAC.
|
||||
scratch, payload rlimits, and a uid map yielding no privilege the host uid
|
||||
did not already have. A per-sandbox runtime uid (D-025) is the follow-up that
|
||||
hardens DAC.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import os
|
||||
import shlex
|
||||
import shutil
|
||||
import signal
|
||||
import time
|
||||
from datetime import UTC, datetime
|
||||
from pathlib import Path
|
||||
@@ -48,22 +68,27 @@ from .backend import ExecResult, ResourceLimits, SandboxHandle, SandboxSpec
|
||||
from .workdir import create_layout
|
||||
from .workdir import snapshot as workdir_snapshot
|
||||
|
||||
#: Args shared by every namespace we spawn (D-024). No `unshare --bind` on
|
||||
#: util-linux 2.38 — the bind is done from inside the namespace instead.
|
||||
#: Args shared by every PURE shell namespace we spawn (D-024). No
|
||||
#: `unshare --bind` on util-linux 2.38 — the bind is done from inside instead.
|
||||
UNSHARE_ARGS: tuple[str, ...] = (
|
||||
"--user", # new user namespace …
|
||||
"--map-root-user", # … in which we are uid 0 (mapped to host uid outside)
|
||||
"--mount", # private mount table
|
||||
"--pid", # private PID table
|
||||
"--fork", # child is PID 1 in its namespace (reaps zombies, gets signals)
|
||||
"--net", # fresh net namespace: no lo, no veth → fully offline
|
||||
"--net", # fresh net namespace: no usable route → effectively offline
|
||||
)
|
||||
|
||||
IN_NS_WORKDIR = "/tmp/work" # where the workspace is bound inside the namespace
|
||||
IN_NS_WORKDIR = "/tmp/work" # where the workspace is bound inside a pure shell ns
|
||||
|
||||
#: Sentinels the long-lived namespace supervisors print once their mounts are
|
||||
#: laid out. exec()/the manager must not run before the bind exists.
|
||||
_HELPER_READY = "NC_HELPER_READY"
|
||||
_INNER_READY = "NC_INNER_READY"
|
||||
|
||||
|
||||
class SandboxUnavailableError(RuntimeError):
|
||||
"""`unshare` missing or user namespaces blocked on this host."""
|
||||
"""`unshare`/`nsenter` missing or user namespaces blocked on this host."""
|
||||
|
||||
|
||||
def _build_shim(workspace: Path, limits: ResourceLimits, cmd: list[str]) -> str:
|
||||
@@ -90,35 +115,286 @@ def _build_shim(workspace: Path, limits: ResourceLimits, cmd: list[str]) -> str:
|
||||
)
|
||||
|
||||
|
||||
class _Tracked:
|
||||
"""The process tree + paths for one persistent (telemetry-wired) sandbox."""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
*,
|
||||
helper: asyncio.subprocess.Process,
|
||||
inner: asyncio.subprocess.Process,
|
||||
agent: asyncio.subprocess.Process | None,
|
||||
inner_pid: int, # host pid of the SANDBOXED init (sleep) — ns enter target
|
||||
host_dir: Path,
|
||||
workspace: Path,
|
||||
ns_root: Path,
|
||||
ns_workdir: Path,
|
||||
) -> None:
|
||||
self.helper = helper
|
||||
self.inner = inner
|
||||
self.agent = agent
|
||||
self.inner_pid = inner_pid
|
||||
self.host_dir = host_dir
|
||||
self.workspace = workspace
|
||||
self.ns_root = ns_root
|
||||
self.ns_workdir = ns_workdir
|
||||
|
||||
|
||||
class UnshareBackend: # satisfies SandboxBackend structurally (Protocol)
|
||||
"""D-024 backend: subprocess-per-exec inside fresh Linux namespaces."""
|
||||
"""D-024 backend: namespace subprocesses; persistent tree for task sandboxes."""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
unshare_path: str | None = None,
|
||||
limits: ResourceLimits | None = None, # per-spec override lands in 1-04
|
||||
nsenter_path: str | None = None,
|
||||
agent_script: Path | None = None,
|
||||
) -> None:
|
||||
self._unshare = unshare_path or shutil.which("unshare") or "unshare"
|
||||
self._nsenter = nsenter_path or shutil.which("nsenter") or "nsenter"
|
||||
self._limits = limits or ResourceLimits()
|
||||
# The stdlib-only capture agent script, copied into each tracked
|
||||
# workdir's host/ tree so nsenter can reach it inside the sandbox.
|
||||
# ai_service/sandbox/unshare_backend.py -> parents[2] = apps/ai-service.
|
||||
self._agent_script = agent_script or (
|
||||
Path(__file__).resolve().parents[2] / "scripts" / "sandbox-agent.py"
|
||||
)
|
||||
# Tracked (persistent) sandboxes by id; pure shell sandboxes are absent.
|
||||
self._tracked: dict[str, _Tracked] = {}
|
||||
|
||||
# -- spawn ------------------------------------------------------------------
|
||||
|
||||
async def spawn(self, spec: SandboxSpec) -> SandboxHandle:
|
||||
"""Lay out the workdir and return a handle.
|
||||
"""Lay out the workdir; if `spec.capture_env` is set, start the sandbox.
|
||||
|
||||
Isolation is established per-`exec` (each exec = fresh namespaces), so
|
||||
spawn only prepares on-disk state; there is no long-lived init process.
|
||||
A spec WITHOUT capture_env keeps REQ-3-001 semantics: spawn only
|
||||
prepares disk state and each exec forks a fresh (offline) namespace.
|
||||
A spec WITH capture_env starts the persistent helper/inner tree and the
|
||||
capture agent, and `handle.pid` carries the agent's host pid.
|
||||
"""
|
||||
create_layout(spec)
|
||||
if not spec.capture_env:
|
||||
return SandboxHandle(
|
||||
id=spec.sandbox_id,
|
||||
pid=None, # no persistent process; each exec forks short-lived PIDs
|
||||
workdir=spec.workdir,
|
||||
created_at=datetime.now(UTC),
|
||||
)
|
||||
tracked = await self._spawn_tracked(spec)
|
||||
self._tracked[spec.sandbox_id] = tracked
|
||||
return SandboxHandle(
|
||||
id=spec.sandbox_id,
|
||||
pid=None, # no persistent process; each exec forks short-lived PIDs
|
||||
pid=tracked.agent.pid if tracked.agent is not None else tracked.inner_pid,
|
||||
workdir=spec.workdir,
|
||||
created_at=datetime.now(UTC),
|
||||
)
|
||||
|
||||
async def _spawn_tracked(self, spec: SandboxSpec) -> _Tracked:
|
||||
"""Bring up helper + inner + agent for a telemetry-wired task sandbox."""
|
||||
host_dir = spec.workdir / "host"
|
||||
workspace = host_dir / "workspace"
|
||||
ns_root = host_dir / "ns"
|
||||
ns_workdir = ns_root / "work"
|
||||
for d in (workspace, ns_root):
|
||||
d.mkdir(parents=True, exist_ok=True)
|
||||
# The agent script must live INSIDE the workspace: the inner ns bind
|
||||
# mounts <host_dir>/workspace -> <ns_root>/work, so only workspace
|
||||
# content is visible in-namespace at /work.
|
||||
agent_host_path = workspace / "sandbox-agent.py"
|
||||
shutil.copyfile(self._agent_script, agent_host_path)
|
||||
|
||||
helper = await self._launch_ns(
|
||||
[
|
||||
self._unshare,
|
||||
"--user",
|
||||
"--map-root-user",
|
||||
"--mount",
|
||||
"sh",
|
||||
"-c",
|
||||
(
|
||||
# Isolate ns/ so the inner tmpfs never propagates back to the
|
||||
# host mount table (make-private is best-effort on this host).
|
||||
f"mount --bind {shlex.quote(str(ns_root))} {shlex.quote(str(ns_root))}; "
|
||||
f"mount --make-private {shlex.quote(str(ns_root))} 2>/dev/null; "
|
||||
f"echo {_HELPER_READY}; exec sleep 3600"
|
||||
),
|
||||
],
|
||||
sentinel=_HELPER_READY,
|
||||
label="helper",
|
||||
)
|
||||
try:
|
||||
inner = await self._launch_ns(
|
||||
[
|
||||
*self._helper_join_argv(helper),
|
||||
self._unshare,
|
||||
"--mount",
|
||||
"--pid",
|
||||
"--fork",
|
||||
"--net",
|
||||
"sh",
|
||||
"-c",
|
||||
(
|
||||
f"mount -t tmpfs tmpfs {shlex.quote(str(ns_root))}; "
|
||||
f"mkdir -p {shlex.quote(str(ns_workdir))}; "
|
||||
f"mount --bind {shlex.quote(str(workspace))} "
|
||||
f"{shlex.quote(str(ns_workdir))}; "
|
||||
f"echo {_INNER_READY}; exec sleep 3600"
|
||||
),
|
||||
],
|
||||
sentinel=_INNER_READY,
|
||||
label="inner",
|
||||
)
|
||||
except Exception:
|
||||
await self._reap(helper)
|
||||
raise
|
||||
|
||||
await asyncio.sleep(0) # let the inner child's sleep fork settle
|
||||
inner_pid = await asyncio.to_thread(self._find_child_pid, inner.pid)
|
||||
if inner_pid is None:
|
||||
await self._reap(inner)
|
||||
await self._reap(helper)
|
||||
raise SandboxUnavailableError(
|
||||
f"could not resolve sandboxed init pid for {spec.sandbox_id}"
|
||||
)
|
||||
|
||||
tracked = _Tracked(
|
||||
helper=helper,
|
||||
inner=inner,
|
||||
agent=None,
|
||||
inner_pid=inner_pid,
|
||||
host_dir=host_dir,
|
||||
workspace=workspace,
|
||||
ns_root=ns_root,
|
||||
ns_workdir=ns_workdir,
|
||||
)
|
||||
if spec.capture_env:
|
||||
tracked.agent = await self._launch_agent(spec, tracked, agent_host_path)
|
||||
return tracked
|
||||
|
||||
# -- process launch helpers --------------------------------------------------
|
||||
|
||||
def _helper_join_argv(self, helper: asyncio.subprocess.Process) -> list[str]:
|
||||
"""nsenter argv that runs a command inside the helper's user+mount ns."""
|
||||
if helper.pid is None:
|
||||
raise SandboxUnavailableError("helper namespace process is not running")
|
||||
return [
|
||||
self._nsenter,
|
||||
"-t",
|
||||
str(helper.pid),
|
||||
"-m",
|
||||
"-U",
|
||||
"--preserve-credentials",
|
||||
"--",
|
||||
]
|
||||
|
||||
def _sandbox_join_argv(self, tracked: _Tracked) -> list[str]:
|
||||
"""nsenter argv that joins the inner sandbox MOUNT namespace (uid 0)."""
|
||||
return [self._nsenter, "-t", str(tracked.inner_pid), "-m", "--"]
|
||||
|
||||
async def _launch_agent(
|
||||
self, spec: SandboxSpec, tracked: _Tracked, agent_host_path: Path
|
||||
) -> asyncio.subprocess.Process:
|
||||
"""Launch the capture agent: joins the sandbox mount ns, NOT pid/net.
|
||||
|
||||
The nsenter chain swaps the mount table under the process, so a HOST
|
||||
cwd/relative path is invalid after the join (observed: python3
|
||||
resolved ``sandbox-agent.py`` against a stale root → ``//…`` and
|
||||
exited rc=2). The launch therefore happens through ``sh -c`` INSIDE
|
||||
the joined namespace, using only in-namespace absolute paths: the
|
||||
workspace is bind-mounted at ``<ns_root>/work``, the agent script was
|
||||
copied into the host workspace, so ``/work/sandbox-agent.py`` exists
|
||||
after the join. The agent runs ONLINE (joins mount ns only, not the
|
||||
offline net ns) so it can dial the ai-service WS ingest loopback.
|
||||
"""
|
||||
env = dict(spec.capture_env or {})
|
||||
in_ns_script = f"{tracked.ns_root / 'work' / 'sandbox-agent.py'}"
|
||||
in_ns_cwd = f"{tracked.ns_root / 'work'}"
|
||||
launch = f"cd {shlex.quote(in_ns_cwd)} && exec python3 {shlex.quote(in_ns_script)}"
|
||||
try:
|
||||
return await asyncio.create_subprocess_exec(
|
||||
*self._helper_join_argv(tracked.helper),
|
||||
*self._sandbox_join_argv(tracked),
|
||||
"sh",
|
||||
"-c",
|
||||
launch,
|
||||
stdin=asyncio.subprocess.DEVNULL,
|
||||
stdout=asyncio.subprocess.DEVNULL,
|
||||
stderr=asyncio.subprocess.DEVNULL,
|
||||
env=env,
|
||||
)
|
||||
except FileNotFoundError as exc: # pragma: no cover - env-dependent
|
||||
raise SandboxUnavailableError("python3 unavailable for capture agent") from exc
|
||||
|
||||
async def _launch_ns(
|
||||
self, argv: list[str], *, sentinel: str, label: str
|
||||
) -> asyncio.subprocess.Process:
|
||||
"""Spawn a namespace supervisor and wait for its `sentinel` line."""
|
||||
proc = await asyncio.create_subprocess_exec(
|
||||
*argv,
|
||||
stdout=asyncio.subprocess.PIPE,
|
||||
stderr=asyncio.subprocess.STDOUT,
|
||||
)
|
||||
|
||||
async def _wait_ready() -> None:
|
||||
if proc.stdout is None: # pragma: no cover (stdout is a PIPE)
|
||||
raise SandboxUnavailableError(f"{label} namespace missing stdout pipe")
|
||||
async for raw in proc.stdout:
|
||||
if raw.decode(errors="replace").strip() == sentinel:
|
||||
return
|
||||
raise SandboxUnavailableError(
|
||||
f"{label} namespace exited before signalling readiness: {argv[:3]}"
|
||||
)
|
||||
|
||||
try:
|
||||
await asyncio.wait_for(_wait_ready(), timeout=10.0)
|
||||
except TimeoutError as exc:
|
||||
await self._reap(proc)
|
||||
raise SandboxUnavailableError(
|
||||
f"{label} namespace never became ready (timeout): {argv[:3]}"
|
||||
) from exc
|
||||
except SandboxUnavailableError:
|
||||
await self._reap(proc)
|
||||
raise
|
||||
return proc
|
||||
|
||||
@staticmethod
|
||||
def _find_child_pid(parent_pid: int | None) -> int | None:
|
||||
"""First direct child of `parent_pid` (the pid-namespaced `sleep`).
|
||||
|
||||
The helper→unshare shim is inner.pid's parent chain head, but the
|
||||
SANDBOXED mount/pid namespaces belong to its forked child (the
|
||||
`sleep`). nsenter must target THAT pid to land inside the sandbox.
|
||||
Reads /proc directly — best-effort, host-local, no subprocess.
|
||||
"""
|
||||
if parent_pid is None:
|
||||
return None
|
||||
for entry in os.listdir("/proc"):
|
||||
if not entry.isdigit():
|
||||
continue
|
||||
try:
|
||||
with open(f"/proc/{entry}/stat") as fh:
|
||||
# ppid is field 4; comm (field 2) may contain spaces, so
|
||||
# parse relative to the LAST ')'.
|
||||
rest = fh.read().rsplit(") ", 1)[1].split()
|
||||
if int(rest[1]) == parent_pid: # state=rest[0], ppid=rest[1]
|
||||
return int(entry)
|
||||
except (OSError, IndexError, ValueError):
|
||||
continue
|
||||
return None
|
||||
|
||||
# -- exec --------------------------------------------------------------------
|
||||
|
||||
async def exec(self, handle: SandboxHandle, cmd: list[str]) -> ExecResult:
|
||||
"""Run `cmd` in a fresh namespace rooted at the sandbox workspace."""
|
||||
"""Run `cmd` in the sandbox workspace (cwd = the bound workspace)."""
|
||||
if not cmd:
|
||||
raise ValueError("exec requires a non-empty cmd")
|
||||
tracked = self._tracked.get(handle.id)
|
||||
if tracked is not None:
|
||||
return await self._exec_tracked(handle, tracked, cmd)
|
||||
return await self._exec_fresh(handle, cmd)
|
||||
|
||||
async def _exec_fresh(self, handle: SandboxHandle, cmd: list[str]) -> ExecResult:
|
||||
"""Pure shell sandbox: spawn one fresh offline namespace per exec."""
|
||||
workspace = handle.workdir / "workspace"
|
||||
if not workspace.is_dir():
|
||||
raise SandboxUnavailableError(f"spawn() first: no workspace at {workspace}")
|
||||
@@ -129,7 +405,6 @@ class UnshareBackend: # satisfies SandboxBackend structurally (Protocol)
|
||||
"-c",
|
||||
_build_shim(workspace, self._limits, cmd),
|
||||
]
|
||||
|
||||
started = time.monotonic()
|
||||
proc = await asyncio.create_subprocess_exec(
|
||||
*argv,
|
||||
@@ -146,13 +421,128 @@ class UnshareBackend: # satisfies SandboxBackend structurally (Protocol)
|
||||
duration_s=time.monotonic() - started,
|
||||
)
|
||||
|
||||
async def _exec_tracked(
|
||||
self, handle: SandboxHandle, tracked: _Tracked, cmd: list[str]
|
||||
) -> ExecResult:
|
||||
"""Task sandbox: join the persistent inner namespace (offline, uid 0).
|
||||
|
||||
rlimits apply in the joining subshell so only the payload is limited;
|
||||
cwd is the bound workspace (`<ns>/work`).
|
||||
"""
|
||||
if tracked.inner.returncode is not None:
|
||||
raise SandboxUnavailableError(
|
||||
f"sandbox {handle.id} is not running (inner namespace exited)"
|
||||
)
|
||||
quoted_cmd = " ".join(shlex.quote(part) for part in cmd)
|
||||
rlimit_prefix = (
|
||||
f"ulimit -v {self._limits.memory_bytes // 1024}; "
|
||||
f"ulimit -t {self._limits.cpu_seconds}; "
|
||||
f"ulimit -f {self._limits.file_size_bytes // 512}; "
|
||||
)
|
||||
shell = (
|
||||
f"cd {shlex.quote(str(tracked.ns_workdir))}; "
|
||||
f"{rlimit_prefix}"
|
||||
f"exec {quoted_cmd}"
|
||||
)
|
||||
argv = [
|
||||
*self._helper_join_argv(tracked.helper),
|
||||
*self._sandbox_join_argv(tracked),
|
||||
"sh",
|
||||
"-c",
|
||||
shell,
|
||||
]
|
||||
started = time.monotonic()
|
||||
proc = await asyncio.create_subprocess_exec(
|
||||
*argv,
|
||||
stdin=asyncio.subprocess.DEVNULL,
|
||||
stdout=asyncio.subprocess.PIPE,
|
||||
stderr=asyncio.subprocess.PIPE,
|
||||
)
|
||||
out, err = await proc.communicate()
|
||||
return ExecResult(
|
||||
cmd=cmd,
|
||||
returncode=proc.returncode if proc.returncode is not None else -1,
|
||||
stdout=out.decode(errors="replace"),
|
||||
stderr=err.decode(errors="replace"),
|
||||
duration_s=time.monotonic() - started,
|
||||
)
|
||||
|
||||
# -- snapshot / destroy --------------------------------------------------------
|
||||
|
||||
async def snapshot(self, handle: SandboxHandle) -> Path:
|
||||
return workdir_snapshot(handle.workdir)
|
||||
workspace_root = handle.workdir
|
||||
tracked = self._tracked.get(handle.id)
|
||||
if tracked is not None:
|
||||
# Copy the tracked workspace, not the legacy <workdir>/workspace.
|
||||
dest_parent = handle.workdir / "snapshots"
|
||||
dest_parent.mkdir(parents=True, exist_ok=True)
|
||||
return workdir_snapshot_from_workspace(tracked.workspace, dest_parent)
|
||||
return workdir_snapshot(workspace_root)
|
||||
|
||||
async def destroy(self, handle: SandboxHandle) -> None:
|
||||
"""Best-effort teardown. Namespaces die with their process; nothing to kill.
|
||||
"""Best-effort teardown. Pure shell sandboxes die with their exec; for a
|
||||
tracked task sandbox reap AGENT → INNER → HELPER so no capture process
|
||||
or namespace supervisor outlives the handle (REQ-3-003 lifecycle).
|
||||
|
||||
Keeping the workdir is deliberate: snapshots must survive destroy so a
|
||||
learner's last state can be restored by the manager layer.
|
||||
"""
|
||||
tracked = self._tracked.pop(handle.id, None)
|
||||
if tracked is not None:
|
||||
# Agent first (it must not flush a "stopped" event into a dead
|
||||
# sandbox), then the namespace tree. The inner `unshare --fork`
|
||||
# shim is NOT the namespace init: killing it orphans its child
|
||||
# (the `sleep` that is PID 1 of the sandbox pid+mnt+net ns),
|
||||
# which reparents to host init and holds the tmpfs + bind for
|
||||
# a full hour (observed: ~30 leaked `sleep 3600` after a test
|
||||
# run). `--kill-child` does not reach it either (util-linux
|
||||
# 2.38 leaks the same child under this flag combo — the child
|
||||
# is reparented before unshare's signal handler runs). The
|
||||
# deterministic kill is SIGKILL on the ns-init's HOST pid,
|
||||
# which we already track as `tracked.inner_pid` (nsenter uses
|
||||
# it for exec); the kernel then tears down the namespace with
|
||||
# its init (no processes remain).
|
||||
for proc in (tracked.agent, tracked.inner, tracked.helper):
|
||||
if proc is not None:
|
||||
await self._reap(proc)
|
||||
self._kill_pid(tracked.inner_pid)
|
||||
handle.pid = None
|
||||
|
||||
@staticmethod
|
||||
def _kill_pid(pid: int | None, sig: int = signal.SIGKILL) -> None:
|
||||
"""Best-effort host-side signal; pid recycled or gone is not an error."""
|
||||
if pid is None:
|
||||
return
|
||||
try:
|
||||
os.kill(pid, sig)
|
||||
except (ProcessLookupError, PermissionError):
|
||||
pass # already dead, or not ours — nothing to do
|
||||
|
||||
@staticmethod
|
||||
async def _reap(proc: asyncio.subprocess.Process) -> None:
|
||||
"""SIGTERM then SIGKILL, tolerant of an already-dead process."""
|
||||
if proc.returncode is not None:
|
||||
return
|
||||
try:
|
||||
proc.terminate()
|
||||
except ProcessLookupError:
|
||||
return
|
||||
try:
|
||||
await asyncio.wait_for(proc.wait(), timeout=5.0)
|
||||
except TimeoutError:
|
||||
try:
|
||||
proc.kill()
|
||||
except ProcessLookupError:
|
||||
return
|
||||
try:
|
||||
await asyncio.wait_for(proc.wait(), timeout=5.0)
|
||||
except TimeoutError: # pragma: no cover - SIGKILL always wins
|
||||
pass
|
||||
|
||||
|
||||
def workdir_snapshot_from_workspace(workspace: Path, snapshots_dir: Path) -> Path:
|
||||
"""Snapshot helper for tracked sandboxes whose workspace is `<workdir>/host/workspace`
|
||||
instead of the legacy `<workdir>/workspace` layout."""
|
||||
dest = snapshots_dir / datetime.now(UTC).strftime("%Y%m%dT%H%M%SZ")
|
||||
shutil.copytree(workspace, dest, symlinks=False)
|
||||
return dest
|
||||
|
||||
@@ -0,0 +1,21 @@
|
||||
"""Live build telemetry — event models and the TraceStore protocol (REQ-3-003).
|
||||
|
||||
Boundary rule (D-027): telemetry/ imports from config only — never from
|
||||
agents/ or api/ (agents call engines through narrow interfaces, never
|
||||
the reverse; api/ composes stores via DI).
|
||||
"""
|
||||
|
||||
from .ingest import IngestSession, TraceIntegrityMap, telemetry_ingest_endpoint
|
||||
from .models import EventKind, TelemetryEvent, TraceSpan
|
||||
from .store import SQLiteTraceStore, TraceStore
|
||||
|
||||
__all__ = [
|
||||
"EventKind",
|
||||
"IngestSession",
|
||||
"SQLiteTraceStore",
|
||||
"TelemetryEvent",
|
||||
"TraceIntegrityMap",
|
||||
"TraceSpan",
|
||||
"TraceStore",
|
||||
"telemetry_ingest_endpoint",
|
||||
]
|
||||
@@ -0,0 +1,448 @@
|
||||
"""WS ingest protocol for learner telemetry (REQ-3-003, D-026, G-3).
|
||||
|
||||
Frame contract — trace identity travels as QUERY PARAMS on the WS upgrade
|
||||
(`WS /v1/telemetry/ingest?learner_id=...&task_id=...&sandbox_id=...`), NOT as
|
||||
a first init frame. Rationale: the in-sandbox capture agent (Task 2-2-01) is a
|
||||
stdlib-only RFC6455 client where the URL is the cheapest thing to parametrize
|
||||
(`NC_INGEST_URL` carries the query string); identity is also visible to the
|
||||
server BEFORE accept(), so a malformed handshake can be rejected without an
|
||||
accept/close round-trip. Client messages are then ONE event per JSON text
|
||||
frame — no envelope:
|
||||
|
||||
{"seq": 0, "kind": "command", "payload": {...}, "ts": "...",
|
||||
"sandbox_id": "..."} # learner_id / task_id forbidden (URL owns them)
|
||||
|
||||
Server → client frames are typed status envelopes:
|
||||
|
||||
{"type": "ack_total", "count": N} — final flush summary, then close 1000
|
||||
{"type": "gap_warning", "missing_seqs": [...]} — seq skipped ahead
|
||||
{"type": "event_rejected", "detail": "..."} — one frame failed validation
|
||||
(seq echoed when parseable)
|
||||
{"type": "event_rejected", "seq": N, "detail": "..."} — stored-field rejected (bad kind)
|
||||
{"type": "flooded", "reason": "cap_exceeded"|"queue_overflow",
|
||||
"count": N} — sent before close(1008)
|
||||
|
||||
Keepalive: the server sends an opaque ping frame every `PING_INTERVAL_S` (the
|
||||
capture agent auto-pongs at the frame layer); a peer that is silent past
|
||||
`PONG_TIMEOUT_S` is assumed wedged, but the keepalive half only LOGS — the
|
||||
receiver half owns disconnect detection (single-box pilot: TCP EOF is
|
||||
reliable; an aggressive pong-watchdog would false-positive on loaded boxes).
|
||||
|
||||
Flood control (GRILL G-3, BINDING — silent drop-oldest is FORBIDDEN):
|
||||
* per-connection inbound queue bounded at `INBOUND_QUEUE_MAX` frames; on
|
||||
overflow → close code 1008 (policy violation) + trace marked
|
||||
INCOMPLETE_FLOODED via `TraceIntegrityMap`.
|
||||
* total events for the (learner, task) exceeding
|
||||
`Settings.telemetry_max_events_per_task` → same 1008 + INCOMPLETE_FLOODED.
|
||||
`INCOMPLETE_FLOODED` is an integrity signal Proctor/Phase-3 grader read via
|
||||
`TraceIntegrityMap.is_incomplete()` (the G-4 gate): a flooded trace can
|
||||
never yield a credential.
|
||||
|
||||
Boundary (D-027): telemetry/ never imports agents/ or api/. This module
|
||||
imports only `fastapi.WebSocket` for the socket type (a protocol surface, not
|
||||
a DI framework); the session engine below depends only on the TraceStore
|
||||
protocol + Settings, and api/telemetry.py injects both through plain
|
||||
parameters.
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import contextlib
|
||||
import json
|
||||
import logging
|
||||
from datetime import datetime
|
||||
from typing import Any, Final
|
||||
|
||||
from fastapi import WebSocket, WebSocketDisconnect
|
||||
from pydantic import BaseModel, ConfigDict, Field, ValidationError
|
||||
|
||||
from ..config import Settings
|
||||
from .models import TelemetryEvent
|
||||
from .store import TraceStore
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
#: WebSocket close code 1008 — policy violation (RFC 6455 §7.4.1).
|
||||
WS_CLOSE_POLICY_VIOLATION: Final = 1008
|
||||
|
||||
#: Bounded inbound queue depth per connection (G-3). Sized for burst-tolerance
|
||||
#: well above the capture agent's emission rate; overflow is a flood signal,
|
||||
#: not a backpressure knob.
|
||||
INBOUND_QUEUE_MAX: Final = 256
|
||||
|
||||
PING_INTERVAL_S: Final = 20.0
|
||||
|
||||
|
||||
class InboundEventFrame(BaseModel):
|
||||
"""Client → server event frame (one TelemetryEvent minus URL-owned ids).
|
||||
|
||||
`extra="forbid"`: learner_id/task_id arriving in the frame body is a
|
||||
contract violation — identity comes from the query params only, so a
|
||||
replayed frame can never lie about which trace it belongs to.
|
||||
"""
|
||||
|
||||
model_config = ConfigDict(extra="forbid")
|
||||
|
||||
seq: int = Field(ge=0)
|
||||
kind: str = Field(min_length=1)
|
||||
payload: dict[str, Any] = Field(default_factory=dict)
|
||||
ts: datetime
|
||||
sandbox_id: str = ""
|
||||
|
||||
|
||||
class TraceIntegrityMap:
|
||||
"""Integrity flags for traces that can never be graded (G-3/G-4).
|
||||
|
||||
Process-local and deliberately small: v0.3 runs ONE ai-service process per
|
||||
box, and the Phase-3 grader reads this flag through the same DI container
|
||||
— D-019-style in-memory registry precedent (the sandbox handle registry is
|
||||
the same shape). The flag is terminal within the process: a reconnect
|
||||
sending legal events does NOT clear it — the trace is already untrusted as
|
||||
grading input. Restarting ai-service resets flags; grading runs against a
|
||||
live service, and the SQLite trace rows themselves are durable.
|
||||
|
||||
All methods are sync: mutation is a dict write, reads are dict lookups —
|
||||
no await needed, so callers from any layer (API handlers, the grader)
|
||||
don't inherit an async surface for a nanosecond operation.
|
||||
"""
|
||||
|
||||
def __init__(self) -> None:
|
||||
# (learner_id, task_id) -> machine-readable reason (INCOMPLETE_FLOODED)
|
||||
self._flags: dict[tuple[str, str], str] = {}
|
||||
|
||||
def mark(self, learner_id: str, task_id: str, reason: str) -> None:
|
||||
"""Set an integrity flag. Presence of the flag is the signal; the
|
||||
reason is informational (last write wins)."""
|
||||
self._flags[(learner_id, task_id)] = reason
|
||||
|
||||
def clear(self, learner_id: str, task_id: str) -> None:
|
||||
"""Test seam: reset a flag (production ingest never clears)."""
|
||||
self._flags.pop((learner_id, task_id), None)
|
||||
|
||||
def is_incomplete(self, learner_id: str, task_id: str) -> bool:
|
||||
"""True when the trace carries ANY terminal integrity flag."""
|
||||
return (learner_id, task_id) in self._flags
|
||||
|
||||
def reason(self, learner_id: str, task_id: str) -> str | None:
|
||||
"""The flag's reason (INCOMPLETE_FLOODED), or None when unflagged."""
|
||||
return self._flags.get((learner_id, task_id))
|
||||
|
||||
|
||||
class IngestSession:
|
||||
"""One WebSocket ingest connection: receive → queue → drain → store.
|
||||
|
||||
Two tasks per connection:
|
||||
* `_receiver` — reads frames, validates shape, enqueues (bounded queue,
|
||||
G-3). Receives never block on SQLite.
|
||||
* `_drainer` — pops frames in arrival order, appends via TraceStore
|
||||
(idempotent on (learner,task,seq)), emits gap warnings, enforces the
|
||||
per-trace event cap.
|
||||
Either task detecting a flood closes the WS with 1008 and marks the trace
|
||||
INCOMPLETE_FLOODED. The events queue carries `None` as the client-
|
||||
disconnect sentinel.
|
||||
|
||||
- `telemetry_max_events_per_task` is consulted at connect and re-checked
|
||||
per append against the DURABLE row count (cap compares against stored
|
||||
events, so a skipped-ahead seq cannot burn budget that was never sent).
|
||||
Durable count via `TraceStore.count()` (COUNT(*)) — a single aggregate
|
||||
per append, never materializing trace rows (the pre-P7 code read
|
||||
`len(get_trace(...))` which was O(trace) per event / O(n²) per session).
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
websocket: WebSocket,
|
||||
store: TraceStore,
|
||||
integrity: TraceIntegrityMap,
|
||||
settings: Settings,
|
||||
learner_id: str,
|
||||
task_id: str,
|
||||
sandbox_id: str,
|
||||
) -> None:
|
||||
self._ws = websocket
|
||||
self._store = store
|
||||
self._integrity = integrity
|
||||
# Snapshot of the one setting ingest consults: read once at connect so
|
||||
# a hot-reloaded Settings object mid-session can't move the cap.
|
||||
self._max_events = settings.telemetry_max_events_per_task
|
||||
self.learner_id = learner_id
|
||||
self.task_id = task_id
|
||||
self.sandbox_id = sandbox_id
|
||||
|
||||
self._queue: asyncio.Queue[InboundEventFrame | None] = asyncio.Queue(
|
||||
maxsize=INBOUND_QUEUE_MAX
|
||||
)
|
||||
self._seen: set[int] = set()
|
||||
self._next_expected: int | None = None # in-connection monotonic hint
|
||||
self._received = 0
|
||||
self._stored = 0
|
||||
self._deduped = 0
|
||||
self._rejected = 0
|
||||
self._flooded = False
|
||||
self._flood_reason = ""
|
||||
|
||||
# -- receive half ----------------------------------------------------------
|
||||
|
||||
async def run(self) -> None:
|
||||
"""Accept, run receiver+drainer, close cleanly. Owns the WS lifecycle."""
|
||||
await self._ws.accept()
|
||||
logger.info(
|
||||
"telemetry ingest connected: %s/%s sandbox=%s",
|
||||
self.learner_id,
|
||||
self.task_id,
|
||||
self.sandbox_id or "(none)",
|
||||
)
|
||||
pinger = asyncio.create_task(self._keepalive())
|
||||
receiver = asyncio.create_task(self._receiver())
|
||||
drainer = asyncio.create_task(self._drainer())
|
||||
# First terminal outcome shuts the session down: client disconnect
|
||||
# (receiver ends) → drainer flushes; drainer ended (clean close after
|
||||
# flush or a 1008 flood close) → receiver must not linger.
|
||||
pending: set[asyncio.Task[None]] = {receiver, drainer}
|
||||
try:
|
||||
done, pending = await asyncio.wait(
|
||||
pending, return_when=asyncio.FIRST_COMPLETED
|
||||
)
|
||||
if receiver in done and drainer in pending:
|
||||
try:
|
||||
await drainer # final flush → sends ack_total, close 1000
|
||||
finally:
|
||||
pending.discard(drainer)
|
||||
finally:
|
||||
for task in (pinger, *pending):
|
||||
task.cancel()
|
||||
with contextlib.suppress(asyncio.CancelledError):
|
||||
await task
|
||||
|
||||
async def _receiver(self) -> None:
|
||||
"""Read frames; parse+enqueue. Overflow → flood shutdown (G-3).
|
||||
|
||||
RuntimeError from receive_text is benign here: it fires when the
|
||||
socket was closed by the drainer (1008 flood close) while this task
|
||||
was parked in receive — a terminal condition, not a bug.
|
||||
"""
|
||||
try:
|
||||
while True:
|
||||
raw = await self._ws.receive_text()
|
||||
frame = self._parse(raw)
|
||||
if frame is None:
|
||||
# Rejected frame — keep the connection open; the producer
|
||||
# gets an event_rejected status frame so a malformed batch
|
||||
# is visible (and its seq is never stored). Yield so the
|
||||
# status frame flushes before we block on the next receive.
|
||||
await self._reject_frame(raw)
|
||||
await asyncio.sleep(0)
|
||||
continue
|
||||
self._received += 1
|
||||
try:
|
||||
self._queue.put_nowait(frame)
|
||||
except asyncio.QueueFull:
|
||||
# Bounded queue — overflow is a flood, never drop-oldest.
|
||||
# _trigger_flood closes the socket; fall through to the
|
||||
# tail so the disconnect sentinel is still enqueued — the
|
||||
# drainer is never left parked on an empty queue after a
|
||||
# flood (P7 review: the pre-fix code `return`ed from the
|
||||
# QueueFull branch WITHOUT the sentinel, leaking the
|
||||
# session task set — one per flooded trace).
|
||||
await self._trigger_flood("queue_overflow")
|
||||
return
|
||||
except WebSocketDisconnect:
|
||||
pass
|
||||
except RuntimeError:
|
||||
logger.debug(
|
||||
"ingest receiver: socket already closed (flood path) %s/%s",
|
||||
self.learner_id,
|
||||
self.task_id,
|
||||
)
|
||||
# Client gone (clean close, drop, or flood close): sentinel unblocks
|
||||
# the drainer for a final flush. put_nowait can only fail under flood,
|
||||
# which already terminated the session.
|
||||
with contextlib.suppress(asyncio.QueueFull):
|
||||
self._queue.put_nowait(None)
|
||||
|
||||
def _parse(self, raw: str) -> InboundEventFrame | None:
|
||||
"""Validate one frame; None means malformed (caller rejects it)."""
|
||||
try:
|
||||
return InboundEventFrame.model_validate_json(raw)
|
||||
except ValidationError:
|
||||
return None
|
||||
|
||||
async def _reject_frame(self, raw: str) -> None:
|
||||
"""Malformed envelope: log + event_rejected status frame (never stored)."""
|
||||
self._rejected += 1
|
||||
detail = "invalid event frame"
|
||||
try:
|
||||
InboundEventFrame.model_validate_json(raw)
|
||||
except ValidationError as exc:
|
||||
detail = exc.errors()[0].get("msg", "validation error")
|
||||
logger.warning(
|
||||
"telemetry frame rejected: %s/%s: %s", self.learner_id, self.task_id, detail
|
||||
)
|
||||
seq: int | None = None
|
||||
with contextlib.suppress(Exception):
|
||||
seq = int(json.loads(raw).get("seq")) # best-effort echo for the producer
|
||||
payload: dict[str, Any] = {"type": "event_rejected", "detail": detail}
|
||||
if seq is not None:
|
||||
payload["seq"] = seq
|
||||
await self._send_json(payload)
|
||||
|
||||
# -- drain half --------------------------------------------------------------
|
||||
|
||||
async def _drainer(self) -> None:
|
||||
"""Pop queued frames, append to the store, then close 1000 + summary."""
|
||||
while True:
|
||||
frame = await self._queue.get()
|
||||
if frame is None: # disconnect sentinel → flush complete
|
||||
await self._send_json(
|
||||
{
|
||||
"type": "ack_total",
|
||||
"count": self._stored,
|
||||
"deduped": self._deduped,
|
||||
"rejected": self._rejected,
|
||||
}
|
||||
)
|
||||
with contextlib.suppress(RuntimeError, WebSocketDisconnect):
|
||||
await self._ws.close(code=1000)
|
||||
return
|
||||
await self._append(frame)
|
||||
|
||||
async def _append(self, frame: InboundEventFrame) -> None:
|
||||
# Precedence: a trace already flagged INCOMPLETE_FLOODED is terminal —
|
||||
# the connection that triggered it is being torn down, and any stray
|
||||
# queued frames must not resurrect the trace's intake.
|
||||
if self._integrity.is_incomplete(self.learner_id, self.task_id):
|
||||
await self._trigger_flood("already_flagged")
|
||||
return
|
||||
|
||||
# Per-trace cap (G-3): checked against the DURABLE row count so a
|
||||
# reconnect resumes the budget instead of resetting it, and a
|
||||
# skipped-ahead seq cannot burn budget that was never sent.
|
||||
if self._flood_breached():
|
||||
await self._trigger_flood("cap_exceeded")
|
||||
return
|
||||
|
||||
# TelemetryEvent's @validates hooks fire on CONSTRUCTION (setattr), so
|
||||
# the try must wrap building the model too — an unknown kind raises
|
||||
# before `store.append` is ever reached.
|
||||
event: TelemetryEvent
|
||||
before = self._store.latest_seq(self.learner_id, self.task_id)
|
||||
try:
|
||||
event = TelemetryEvent(
|
||||
learner_id=self.learner_id,
|
||||
task_id=self.task_id,
|
||||
seq=frame.seq,
|
||||
kind=frame.kind,
|
||||
payload=frame.payload,
|
||||
ts=frame.ts,
|
||||
sandbox_id=frame.sandbox_id or self.sandbox_id,
|
||||
)
|
||||
self._store.append(event)
|
||||
except ValueError as exc: # unknown kind / invalid field
|
||||
self._rejected += 1
|
||||
await self._send_json(
|
||||
{"type": "event_rejected", "seq": frame.seq, "detail": str(exc)}
|
||||
)
|
||||
return
|
||||
after = self._store.latest_seq(self.learner_id, self.task_id)
|
||||
|
||||
if after == before and frame.seq in self._seen:
|
||||
self._deduped += 1 # at-least-once retry; stored once (idempotent)
|
||||
else:
|
||||
self._stored += 1
|
||||
self._seen.add(frame.seq)
|
||||
await self._check_gap(frame.seq)
|
||||
# SQLite appends are sync and fast; on a burst the drainer can hold
|
||||
# the loop between receives. Yield so the WS writer flushes the close
|
||||
# and the pinger/interleave stay live under the eventlet-free portal.
|
||||
await asyncio.sleep(0)
|
||||
|
||||
def _flood_breached(self) -> bool:
|
||||
"""True when this append would exceed the per-trace event budget."""
|
||||
# Durable count (NOT latest_seq+1 — a skipped-ahead seq must not burn
|
||||
# un-sent events' budget) via COUNT(*): never materialize the trace
|
||||
# per append (P7 review — the old len(get_trace(...)) built every row
|
||||
# object per event, O(trace) per append / O(n²) per session).
|
||||
durable = self._store.count(learner_id=self.learner_id, task_id=self.task_id)
|
||||
return durable >= self._max_events
|
||||
|
||||
async def _check_gap(self, incoming_seq: int) -> None:
|
||||
"""Seq skipped ahead → log + per-connection gap_warning status frame."""
|
||||
if self._next_expected is not None and incoming_seq > self._next_expected:
|
||||
missing = list(range(self._next_expected, incoming_seq))
|
||||
logger.warning(
|
||||
"telemetry gap: %s/%s missing seqs %s (arrived seq=%d)",
|
||||
self.learner_id,
|
||||
self.task_id,
|
||||
missing,
|
||||
incoming_seq,
|
||||
)
|
||||
await self._send_json({"type": "gap_warning", "missing_seqs": missing})
|
||||
if self._next_expected is None or incoming_seq >= self._next_expected:
|
||||
self._next_expected = incoming_seq + 1
|
||||
|
||||
# -- flood + keepalive ------------------------------------------------------
|
||||
|
||||
async def _trigger_flood(self, reason: str) -> None:
|
||||
"""G-3: 1008 close + INCOMPLETE_FLOODED mark. Exactly once."""
|
||||
if self._flooded:
|
||||
return
|
||||
self._flooded = True
|
||||
self._flood_reason = reason
|
||||
self._integrity.mark(self.learner_id, self.task_id, "INCOMPLETE_FLOODED")
|
||||
logger.warning(
|
||||
"telemetry flood: %s/%s reason=%s — closing 1008, trace marked "
|
||||
"INCOMPLETE_FLOODED (G-3; Proctor/grade gate will refuse it)",
|
||||
self.learner_id,
|
||||
self.task_id,
|
||||
reason,
|
||||
)
|
||||
await self._send_json(
|
||||
{"type": "flooded", "reason": reason, "count": self._received}
|
||||
)
|
||||
with contextlib.suppress(RuntimeError, WebSocketDisconnect):
|
||||
await self._ws.close(
|
||||
code=WS_CLOSE_POLICY_VIOLATION,
|
||||
reason=f"telemetry flood control (G-3): {reason}",
|
||||
)
|
||||
|
||||
async def _keepalive(self) -> None:
|
||||
"""Protocol-level ping on an interval (agent auto-pongs at frame level).
|
||||
|
||||
A send failure means the socket is already gone — the receiver half
|
||||
independently surfaces the disconnect; we just stop pinging.
|
||||
"""
|
||||
while True:
|
||||
await asyncio.sleep(PING_INTERVAL_S)
|
||||
try:
|
||||
await self._ws.send_bytes(b"\x89ping-nextcraft")
|
||||
except (RuntimeError, WebSocketDisconnect):
|
||||
return
|
||||
|
||||
async def _send_json(self, payload: dict[str, Any]) -> None:
|
||||
"""Best-effort status frame; the socket may already be gone."""
|
||||
with contextlib.suppress(RuntimeError, WebSocketDisconnect):
|
||||
await self._ws.send_json(payload)
|
||||
|
||||
|
||||
async def telemetry_ingest_endpoint(
|
||||
websocket: WebSocket,
|
||||
learner_id: str,
|
||||
task_id: str,
|
||||
store: TraceStore,
|
||||
integrity: TraceIntegrityMap,
|
||||
settings: Settings,
|
||||
sandbox_id: str = "",
|
||||
) -> None:
|
||||
"""Engine entry: build the session and run it. api/telemetry.py calls this
|
||||
with query params + app.state services already resolved — this signature
|
||||
is deliberately Depends-free (telemetry/ never knows FastAPI DI exists).
|
||||
"""
|
||||
session = IngestSession(
|
||||
websocket=websocket,
|
||||
store=store,
|
||||
integrity=integrity,
|
||||
settings=settings,
|
||||
learner_id=learner_id,
|
||||
task_id=task_id,
|
||||
sandbox_id=sandbox_id,
|
||||
)
|
||||
await session.run()
|
||||
@@ -0,0 +1,108 @@
|
||||
"""Telemetry event record — the row the trace store persists (REQ-3-003, D-027).
|
||||
|
||||
One model serves both as the JSON payload sent by producers and as the SQLite
|
||||
row schema. `payload` is stored as a JSON column (native JSONB on Postgres —
|
||||
no migration-time shape change, D-027).
|
||||
|
||||
Field contract (consumed by the trace store and the grader):
|
||||
learner_id — non-empty learner identifier.
|
||||
task_id — non-empty task/session identifier; trace identity is the
|
||||
(learner_id, task_id) pair.
|
||||
seq — sequence number per trace, >= 0. Monotonicity per
|
||||
(learner, task) is enforced by the store (Task 2-1-02);
|
||||
this model only rejects negative seqs.
|
||||
kind — event discriminator: command | file_diff | run_result |
|
||||
test_result | activity | stdin | stdout.
|
||||
payload — free-form JSON detail blob.
|
||||
ts — envelope timestamp (UTC); monotonicity enforced at ingest.
|
||||
sandbox_id — originating sandbox ("" for non-sandbox sources).
|
||||
|
||||
Boundary (D-027): telemetry/ never imports agents/ or api/.
|
||||
"""
|
||||
|
||||
from datetime import datetime
|
||||
from typing import Any, Literal
|
||||
|
||||
from pydantic import BaseModel, ConfigDict, Field
|
||||
from sqlalchemy import JSON, Index, String
|
||||
from sqlalchemy.orm import validates
|
||||
from sqlmodel import Field as SQLField
|
||||
from sqlmodel import SQLModel
|
||||
|
||||
EventKind = Literal[
|
||||
"command",
|
||||
"file_diff",
|
||||
"run_result",
|
||||
"test_result",
|
||||
"activity",
|
||||
"stdin",
|
||||
"stdout",
|
||||
]
|
||||
_EVENT_KINDS: frozenset[str] = frozenset(EventKind.__args__)
|
||||
|
||||
|
||||
class TelemetryEvent(SQLModel, table=True):
|
||||
"""A single durable telemetry event; (learner_id, task_id, seq) is PK.
|
||||
|
||||
Constraint enforcement uses SQLAlchemy `@validates` hooks: sqlmodel
|
||||
0.0.42's metaclass drops pydantic `Field(ge=...)`/`field_validator`
|
||||
constraints for table models (the decorators register but never make it
|
||||
into the core schema), while `@validates` fires on every attribute set —
|
||||
construction included — and raises ValueError on violation. seq >= 0 plus
|
||||
a VARCHAR kind column keep the DB shape Postgres-ready (D-027).
|
||||
"""
|
||||
|
||||
__tablename__ = "telemetry_event"
|
||||
# PK columns already produce a unique index; this secondary index covers
|
||||
# trace reads ordered by seq without depending on the PK column order
|
||||
# (Postgres migration target D-027).
|
||||
__table_args__ = (Index("ix_telemetry_event_trace", "learner_id", "task_id"),)
|
||||
|
||||
learner_id: str = SQLField(primary_key=True)
|
||||
task_id: str = SQLField(primary_key=True)
|
||||
seq: int = SQLField(primary_key=True)
|
||||
# Bare Literal annotations crash sqlmodel<=0.0.42's column inference
|
||||
# (issubclass(TypeAlias, Enum)); an explicit sa_type + the validates hook
|
||||
# below gives the same contract: VARCHAR column, Literal-rejected values.
|
||||
kind: EventKind = SQLField(sa_type=String)
|
||||
# JSON column: stored as TEXT on SQLite, native JSONB on Postgres (D-027).
|
||||
payload: dict[str, Any] = SQLField(default_factory=dict, sa_type=JSON)
|
||||
ts: datetime
|
||||
sandbox_id: str = SQLField(default="")
|
||||
|
||||
@validates("learner_id", "task_id")
|
||||
def _ids_non_empty(self, key: str, value: str) -> str:
|
||||
if not value:
|
||||
raise ValueError(f"{key} must be a non-empty identifier")
|
||||
return value
|
||||
|
||||
@validates("seq")
|
||||
def _seq_non_negative(self, key: str, value: int) -> int:
|
||||
if value < 0:
|
||||
raise ValueError("seq must be >= 0 (monotonicity is the store's job)")
|
||||
return value
|
||||
|
||||
@validates("kind")
|
||||
def _kind_is_known(self, key: str, value: str) -> str:
|
||||
if value not in _EVENT_KINDS:
|
||||
raise ValueError(f"unknown event kind: {value!r}")
|
||||
return value
|
||||
|
||||
|
||||
class TraceSpan(BaseModel):
|
||||
"""Derived view: the ordered event trace for one (learner_id, task_id).
|
||||
|
||||
NOT a table — materialized by the store from persisted TelemetryEvents
|
||||
(grader/Lab consume this shape; replay order is the seq column).
|
||||
"""
|
||||
|
||||
model_config = ConfigDict(frozen=True)
|
||||
|
||||
learner_id: str = Field(min_length=1)
|
||||
task_id: str = Field(min_length=1)
|
||||
events: tuple[TelemetryEvent, ...] = ()
|
||||
|
||||
@property
|
||||
def latest_seq(self) -> int:
|
||||
"""Highest seq in the span; -1 when empty (store convention)."""
|
||||
return self.events[-1].seq if self.events else -1
|
||||
@@ -0,0 +1,218 @@
|
||||
"""TraceStore — telemetry persistence protocol + SQLite implementation (REQ-3-003, D-027).
|
||||
|
||||
Postgres-migration-ready (D-027): the protocol is the only surface the API /
|
||||
grader layers touch; swapping SQLiteTraceStore for a Postgres-backed
|
||||
implementation must not change call sites. The `telemetry_event` table uses
|
||||
only portable column types (str / int / datetime / JSON), so the same SQLModel
|
||||
schema stands up unchanged on Postgres.
|
||||
|
||||
Ingest is at-least-once: duplicates carry the same (learner_id, task_id, seq)
|
||||
idempotency key, so `append` with a triplet that is already stored is a no-op.
|
||||
The pair (learner_id, task_id) identifies a trace; `seq` numbers events in it
|
||||
starting at 0.
|
||||
|
||||
Concurrency (a-3): the engine enables WAL + synchronous=NORMAL and a busy
|
||||
timeout at connection time, so the ingest writer and grader readers do not hit
|
||||
`database is locked` on the single-box pilot.
|
||||
|
||||
Boundary: `telemetry/` never imports `agents/` / `api/` and has no FastAPI
|
||||
dependency.
|
||||
"""
|
||||
|
||||
import logging
|
||||
import sqlite3
|
||||
from collections.abc import Iterator
|
||||
from contextlib import contextmanager
|
||||
from datetime import UTC, datetime
|
||||
from pathlib import Path
|
||||
from typing import Any, Protocol
|
||||
|
||||
import sqlalchemy as sa
|
||||
from sqlmodel import Session, SQLModel, create_engine, select
|
||||
|
||||
from ..config import Settings
|
||||
from .models import TelemetryEvent
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class TraceStore(Protocol):
|
||||
"""Persistence contract for ordered per-learner task trace streams.
|
||||
|
||||
Implemented by SQLiteTraceStore (v0.3, D-027); a Postgres implementation
|
||||
must satisfy the same surface.
|
||||
"""
|
||||
|
||||
def append(self, event: TelemetryEvent) -> None:
|
||||
"""Store one event. IDEMPOTENT on (learner_id, task_id, seq):
|
||||
|
||||
at-least-once ingest retries with the same triplet are deduped
|
||||
(stored once), not rejected. Later events must not overwrite an
|
||||
existing row.
|
||||
"""
|
||||
...
|
||||
|
||||
def get_trace(self, learner_id: str, task_id: str) -> list[TelemetryEvent]:
|
||||
"""All stored events for the trace, ordered by seq ascending.
|
||||
|
||||
Detached from any DB session — safe to pass across layers. Empty list
|
||||
when the trace has no events.
|
||||
"""
|
||||
...
|
||||
|
||||
def gaps(self, learner_id: str, task_id: str) -> list[int]:
|
||||
"""Missing seqs in 0..latest for the trace ([0,2,3] stored -> [1])."""
|
||||
...
|
||||
|
||||
def latest_seq(self, learner_id: str, task_id: str) -> int:
|
||||
"""Highest stored seq for the trace; -1 when no events exist."""
|
||||
...
|
||||
|
||||
def count(self, learner_id: str, task_id: str) -> int:
|
||||
"""Number of stored events for the trace (COUNT(*), never
|
||||
materializes rows — the ingest cap consults this per append, so
|
||||
an O(trace) implementation would make ingest O(n²) per session).
|
||||
"""
|
||||
...
|
||||
|
||||
def list_tasks(self, learner_id: str) -> list[str]:
|
||||
"""Distinct task_ids with at least one event for the learner."""
|
||||
...
|
||||
|
||||
def close(self) -> None:
|
||||
"""Release DB connections. Store must not be used after close."""
|
||||
...
|
||||
|
||||
|
||||
def _sqlite_connect(dbapi_connection: sqlite3.Connection, _: object) -> None:
|
||||
"""Per-connection pragma setup (a-3).
|
||||
|
||||
journal_mode=WAL — readers never block the single writer.
|
||||
synchronous=NORMAL — safe in WAL mode, avoids full fsync-per-commit.
|
||||
busy_timeout=5000 — retry briefly under contention instead of
|
||||
`OperationalError: database is locked`.
|
||||
"""
|
||||
cursor = dbapi_connection.cursor()
|
||||
cursor.execute("PRAGMA journal_mode=WAL")
|
||||
cursor.execute("PRAGMA synchronous=NORMAL")
|
||||
cursor.execute("PRAGMA busy_timeout=5000")
|
||||
cursor.close()
|
||||
|
||||
|
||||
def _as_utc(ts: datetime) -> datetime:
|
||||
"""Timestamps travel as naive datetime on SQLite, tz-aware elsewhere.
|
||||
|
||||
SQLite (via SQLModel) drops tzinfo; Postgres TIMESTAMP WITH TIME ZONE
|
||||
keeps it. Normalizing on the read/write boundary makes the store's
|
||||
contract tz-aware UTC regardless of the backend (D-027).
|
||||
"""
|
||||
if ts.tzinfo is None:
|
||||
return ts.replace(tzinfo=UTC) # naive UTC read-side: label as UTC
|
||||
return ts.astimezone(UTC)
|
||||
|
||||
|
||||
class SQLiteTraceStore:
|
||||
"""SQLite-backed TraceStore (SQLModel). First real persistence (D-027)."""
|
||||
|
||||
def __init__(self, db_path: Path | None = None) -> None:
|
||||
self._db_path: Path = db_path if db_path is not None else Settings().db_path
|
||||
self._engine = create_engine(f"sqlite:///{self._db_path}")
|
||||
sa.event.listen(self._engine, "connect", _sqlite_connect)
|
||||
SQLModel.metadata.create_all(self._engine)
|
||||
|
||||
@contextmanager
|
||||
def _session(self) -> Iterator[Session]:
|
||||
# expire_on_commit=False: ORM objects returned from `append`'s
|
||||
# IntegrityError path stay usable without a refresh round-trip.
|
||||
with Session(self._engine, expire_on_commit=False) as session:
|
||||
yield session
|
||||
|
||||
def append(self, event: TelemetryEvent) -> None:
|
||||
# INSERT-if-absent via PK: sqlite3 raises IntegrityError on a
|
||||
# duplicate (learner_id, task_id, seq); swallow it — the row is
|
||||
# already stored, which is the dedup contract for at-least-once
|
||||
# ingest. `session.merge` would upsert instead; wrong semantics here.
|
||||
with self._session() as session:
|
||||
try:
|
||||
session.add(event)
|
||||
session.commit()
|
||||
except sa.exc.IntegrityError:
|
||||
session.rollback()
|
||||
logger.debug(
|
||||
"trace event dedup: %s/%s seq=%d already stored",
|
||||
event.learner_id,
|
||||
event.task_id,
|
||||
event.seq,
|
||||
)
|
||||
|
||||
def get_trace(self, learner_id: str, task_id: str) -> list[TelemetryEvent]:
|
||||
with self._session() as session:
|
||||
stmt = (
|
||||
select(TelemetryEvent)
|
||||
.where(TelemetryEvent.learner_id == learner_id)
|
||||
.where(TelemetryEvent.task_id == task_id)
|
||||
.order_by(TelemetryEvent.seq)
|
||||
)
|
||||
results = session.exec(stmt).all()
|
||||
# Detach from the session: callers must not depend on open-session
|
||||
# ORM magic (lazy loads fail once the session is closed).
|
||||
for row in results:
|
||||
row.ts = _as_utc(row.ts)
|
||||
session.expunge(row)
|
||||
return list(results)
|
||||
|
||||
def _stored_seqs(self, learner_id: str, task_id: str) -> list[int]:
|
||||
with self._session() as session:
|
||||
stmt = (
|
||||
select(TelemetryEvent.seq)
|
||||
.where(TelemetryEvent.learner_id == learner_id)
|
||||
.where(TelemetryEvent.task_id == task_id)
|
||||
.order_by(TelemetryEvent.seq)
|
||||
)
|
||||
# sqlmodel scalar select: rows are plain ints, not 1-tuples.
|
||||
return [int(seq) for seq in session.exec(stmt).all()]
|
||||
|
||||
def gaps(self, learner_id: str, task_id: str) -> list[int]:
|
||||
seqs = self._stored_seqs(learner_id, task_id)
|
||||
if not seqs:
|
||||
return []
|
||||
present = set(seqs)
|
||||
# seq numbering starts at 0; a gap is any seq in 0..latest not stored.
|
||||
return [seq for seq in range(seqs[-1] + 1) if seq not in present]
|
||||
|
||||
def latest_seq(self, learner_id: str, task_id: str) -> int:
|
||||
with self._session() as session:
|
||||
stmt = (
|
||||
select(sa.func.max(TelemetryEvent.seq))
|
||||
.where(TelemetryEvent.learner_id == learner_id)
|
||||
.where(TelemetryEvent.task_id == task_id)
|
||||
)
|
||||
latest: Any = session.exec(stmt).one()
|
||||
return -1 if latest is None else int(latest)
|
||||
|
||||
def count(self, learner_id: str, task_id: str) -> int:
|
||||
# COUNT(*) at the DB — no row materialization. The ingest flood cap
|
||||
# calls this per append (telemetry/ingest._flood_breached); the
|
||||
# docstring-free body keeps it obvious what the query shape is.
|
||||
with self._session() as session:
|
||||
stmt = (
|
||||
select(sa.func.count(TelemetryEvent.seq))
|
||||
.where(TelemetryEvent.learner_id == learner_id)
|
||||
.where(TelemetryEvent.task_id == task_id)
|
||||
)
|
||||
total: Any = session.exec(stmt).one()
|
||||
return int(total or 0)
|
||||
|
||||
def list_tasks(self, learner_id: str) -> list[str]:
|
||||
with self._session() as session:
|
||||
stmt = (
|
||||
select(TelemetryEvent.task_id)
|
||||
.where(TelemetryEvent.learner_id == learner_id)
|
||||
.distinct()
|
||||
.order_by(TelemetryEvent.task_id)
|
||||
)
|
||||
# sqlmodel scalar select: rows are plain strs, not 1-tuples.
|
||||
return [str(task_id) for task_id in session.exec(stmt).all()]
|
||||
|
||||
def close(self) -> None:
|
||||
self._engine.dispose()
|
||||
@@ -0,0 +1,31 @@
|
||||
"""Per-learner variant task generation — templates, generator, VariantStore (REQ-3-005).
|
||||
|
||||
Boundary rule (D-027): variants/ is an engine module — it never imports
|
||||
api/; its ONLY agents/ dependency is the module-direct
|
||||
agents.structured import in generator.py (the sanctioned shared D-020
|
||||
structured defense, same exception as grading/engine.py). api/ composes
|
||||
the generator and store via DI; store.py imports config only.
|
||||
|
||||
CO-ORDINATION NOTE (ADD, don't REMOVE — same convention as grading/):
|
||||
This __init__.py is a minimal placeholder created by the VariantStore
|
||||
task (4-1-02). The templates task (4-1-01) owns this file's final shape
|
||||
— when templates.py lands, ADD its exports alongside these; do not
|
||||
remove the store exports below.
|
||||
|
||||
Wave status: store.py (VariantRecord, VariantStore, SQLiteVariantStore)
|
||||
landed in Wave 1 (task 4-1-02); templates.py is Wave 1 task 4-1-01;
|
||||
generator.py is Wave 2 (4-2-01).
|
||||
"""
|
||||
|
||||
from .store import SQLiteVariantStore, VariantRecord, VariantStore
|
||||
from .templates import TEMPLATES, TaskTemplate, get_template, template_for_competency
|
||||
|
||||
__all__ = [
|
||||
"SQLiteVariantStore",
|
||||
"TEMPLATES",
|
||||
"TaskTemplate",
|
||||
"VariantRecord",
|
||||
"VariantStore",
|
||||
"get_template",
|
||||
"template_for_competency",
|
||||
]
|
||||
@@ -0,0 +1,136 @@
|
||||
"""Seeded per-learner variant generator (D-029, REQ-3-005).
|
||||
|
||||
Contract (binding, from GRILL + PLAN Must-Haves):
|
||||
- REPRODUCIBLE: seed = sha256(template_id|learner_id|milestone); the same
|
||||
(template, learner) re-derives the same seed, params, task_id — and the
|
||||
second generate() call is a cache hit with NO LLM call.
|
||||
- DISTINCT: different learners on the same template draw different params
|
||||
(the sampler is seeded per-learner) and receive distinct statements.
|
||||
- NEVER BLOCKS ON THE LLM: the deterministic skeleton render
|
||||
(`template.render(params)`) is a complete, valid statement; if the D-020
|
||||
LLM render fails after its bounded retry, the fallback is used — and
|
||||
because the fallback is exactly `template.render(seed-params)`, it is
|
||||
auditable from the persisted seed + params without a provenance column.
|
||||
- AUDITABLE: seed + params + statement persist via VariantStore
|
||||
(insert-only first-wins) — the proctoring cross-check path.
|
||||
- FAIR (a-5): slot draws change the scenario, never the difficulty; the
|
||||
template's rubric anchors bound the expected effort envelope, so every
|
||||
variant of one template is held to the same bar.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
from datetime import UTC, datetime
|
||||
from typing import TYPE_CHECKING
|
||||
|
||||
from pydantic import BaseModel, ConfigDict, Field
|
||||
|
||||
from ..agents.structured import StructuredOutputError, structured_completion
|
||||
from ..llm.types import Message
|
||||
from ..prompts.variant import VARIANT_SCHEMA_HINT, render_variant_prompt
|
||||
from .store import VariantRecord
|
||||
from .templates import TaskTemplate, get_template
|
||||
|
||||
if TYPE_CHECKING: # pragma: no cover
|
||||
from ..llm.base import LLMProvider
|
||||
from .store import VariantStore
|
||||
|
||||
MILESTONE = "v0.3"
|
||||
|
||||
|
||||
class RenderedVariant(BaseModel):
|
||||
"""D-20-validated LLM render output (statement only — files come from the template)."""
|
||||
|
||||
model_config = ConfigDict(extra="forbid")
|
||||
|
||||
statement: str = Field(min_length=20)
|
||||
|
||||
|
||||
class UnknownTemplateError(ValueError):
|
||||
"""Raised when generate() is asked for a template id not in the library."""
|
||||
|
||||
|
||||
def derive_seed(template_id: str, learner_id: str, milestone: str = MILESTONE) -> str:
|
||||
"""Reproducible per-(template, learner, milestone) seed (D-029)."""
|
||||
return hashlib.sha256(f"{template_id}|{learner_id}|{milestone}".encode()).hexdigest()
|
||||
|
||||
|
||||
def derive_task_id(seed: str) -> str:
|
||||
"""Deterministic grading/telemetry task key from the seed (16 hex chars)."""
|
||||
return f"task-{seed[:16]}"
|
||||
|
||||
|
||||
class VariantGenerator:
|
||||
"""Seeded instantiation over the template library. DI: store + provider."""
|
||||
|
||||
def __init__(self, store: VariantStore, provider: LLMProvider, model: str) -> None:
|
||||
self._store = store
|
||||
self._provider = provider
|
||||
self._model = model
|
||||
|
||||
async def generate(self, learner_id: str, template_id: str) -> VariantRecord:
|
||||
template = get_template(template_id)
|
||||
if template is None:
|
||||
raise UnknownTemplateError(f"no task template with id {template_id!r}")
|
||||
|
||||
# Cache: D-029 reproducibility — same (learner, template) is served
|
||||
# from the store with no LLM call.
|
||||
cached = self._store.get(learner_id, template_id)
|
||||
if cached is not None:
|
||||
return cached
|
||||
|
||||
seed_hex = derive_seed(template_id, learner_id)
|
||||
task_id = derive_task_id(seed_hex)
|
||||
params = template.sample_params(_seed_int(seed_hex))
|
||||
_validate_params(template, params)
|
||||
|
||||
statement = await self._render(template, params)
|
||||
record = VariantRecord(
|
||||
learner_id=learner_id,
|
||||
task_id=task_id,
|
||||
template_id=template_id,
|
||||
seed=seed_hex,
|
||||
params=dict(params),
|
||||
statement=statement,
|
||||
starter_files=dict(template.starter_files),
|
||||
created_at=datetime.now(UTC),
|
||||
)
|
||||
self._store.save(record)
|
||||
return record
|
||||
|
||||
async def _render(self, template: TaskTemplate, params: dict[str, str | int]) -> str:
|
||||
"""LLM render via D-020; deterministic fallback never blocks task work.
|
||||
|
||||
Provenance note: unlike grades, variants carry no `model` column —
|
||||
the deterministic fallback is exactly `template.render(params)`,
|
||||
re-derivable from the persisted seed + params, so a fallback render is
|
||||
auditable without storing provenance (the seed IS the provenance).
|
||||
"""
|
||||
messages: list[Message] = render_variant_prompt(template, params)
|
||||
try:
|
||||
rendered = await structured_completion(
|
||||
self._provider,
|
||||
messages,
|
||||
model=self._model,
|
||||
schema=RenderedVariant,
|
||||
schema_hint=VARIANT_SCHEMA_HINT,
|
||||
)
|
||||
except StructuredOutputError:
|
||||
# Deterministic fallback: the skeleton + seeded slots is already a
|
||||
# complete statement, re-derivable from the persisted seed.
|
||||
return template.render(params)
|
||||
return rendered.statement
|
||||
|
||||
|
||||
def _seed_int(seed_hex: str) -> int:
|
||||
"""Stable int for random.Random from the hex seed."""
|
||||
return int(seed_hex[:16], 16)
|
||||
|
||||
|
||||
def _validate_params(template: TaskTemplate, params: dict[str, str | int]) -> None:
|
||||
"""Defense in depth: every sampled value must be schema-valid (a-5)."""
|
||||
for slot in template.slots:
|
||||
value = params.get(slot.name)
|
||||
if value is None or not slot.validate_value(value):
|
||||
raise ValueError(f"sampled params invalid for slot {slot.name!r}: {value!r}")
|
||||
@@ -0,0 +1,311 @@
|
||||
"""VariantStore — variant persistence protocol + SQLite implementation (REQ-3-005, D-027).
|
||||
|
||||
Postgres-migration-ready (D-027): the protocol is the only surface the
|
||||
variant generator and API layers touch; swapping SQLiteVariantStore for a
|
||||
Postgres-backed implementation must not change call sites. The
|
||||
`variant_record` table uses only portable column types (str / JSON /
|
||||
datetime), so the same SQLModel schema stands up unchanged on Postgres.
|
||||
|
||||
Insert-only, NOT upsert: (learner_id, template_id) is the variant identity
|
||||
and the FIRST generation is authoritative — reproducibility (D-029) means
|
||||
the seed re-derives the same variant, so the generator's cache path serves
|
||||
`get` instead of saving again. `save` is a plain INSERT; a duplicate pair
|
||||
raises sqlalchemy.exc.IntegrityError to the caller (documented behavior).
|
||||
`task_id` is unique too — it is the grading/telemetry trace key, so a
|
||||
trace or grade can never silently join to a different variant. Both
|
||||
rejections are deliberate: overwriting a stored variant would swap a
|
||||
learner's graded task underneath its trace and grade (audit corruption).
|
||||
Contrast TraceStore.append (dedup-keep-first, swallowed — at-least-once
|
||||
ingest) and GradeStore.save (upsert-latest-wins — a regrade is
|
||||
latest-state); this store is the third contract of the D-027 family.
|
||||
|
||||
Concurrency (a-3): the store enables WAL + synchronous=NORMAL and a busy
|
||||
timeout at connection time, so a generation writer and API readers do not
|
||||
hit `database is locked` on the single-box pilot.
|
||||
|
||||
`created_at` contract: callers stamp UTC (datetime.now(UTC)); SQLite
|
||||
stores it naive and the read paths re-label it tz-aware UTC (same
|
||||
boundary normalization as TelemetryEvent.ts / GradeRecord.created_at, so
|
||||
the contract holds on any backend).
|
||||
|
||||
Boundary (D-027): `variants/` never imports `agents/` / `api/`; this
|
||||
module imports config only.
|
||||
"""
|
||||
|
||||
import logging
|
||||
import sqlite3
|
||||
from collections.abc import Iterator
|
||||
from contextlib import contextmanager
|
||||
from datetime import UTC, datetime
|
||||
from pathlib import Path
|
||||
from typing import Any, Protocol
|
||||
|
||||
import sqlalchemy as sa
|
||||
from sqlalchemy import JSON, Index, UniqueConstraint
|
||||
from sqlalchemy.orm import validates
|
||||
from sqlmodel import Field, Session, SQLModel, create_engine, select
|
||||
|
||||
from ..config import Settings
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class VariantRecord(SQLModel, table=True):
|
||||
"""A persisted task variant; (learner_id, template_id) is the PK — first wins.
|
||||
|
||||
Written once by the variant generator (Task 4-2-01), read by the API
|
||||
layer and proctoring cross-checks through the VariantStore protocol.
|
||||
Constraint enforcement mirrors TelemetryEvent / GradeRecord: sqlmodel
|
||||
0.0.42's metaclass drops pydantic constraints on table models, so
|
||||
SQLAlchemy `@validates` hooks enforce instead and the column types
|
||||
stay Postgres-ready (D-027).
|
||||
|
||||
Field contract:
|
||||
learner_id — non-empty learner identifier (same id space as
|
||||
traces and grades).
|
||||
template_id — non-empty task template identifier; variant
|
||||
identity is the (learner_id, template_id) pair —
|
||||
the pair the generator caches on (exactly one
|
||||
variant per learner per template).
|
||||
task_id — non-empty, GLOBALLY unique task identifier; the
|
||||
grading/telemetry trace key (the (learner_id,
|
||||
task_id) pair TraceStore / GradeStore key on),
|
||||
stamped at generation so a variant's trace and
|
||||
grade join back to it exactly once.
|
||||
seed — non-empty variant seed (D-029); derived from
|
||||
(template_id, learner_id, milestone) so the
|
||||
variant is reproducible and auditable.
|
||||
params — typed parameter-slot values the generator filled;
|
||||
JSON dict. An empty dict is legal (a slotless
|
||||
template).
|
||||
statement — non-empty rendered task statement shown to the
|
||||
learner (distinct per learner by construction,
|
||||
REQ-3-005).
|
||||
starter_files — workspace scaffold: filename -> file content;
|
||||
JSON dict. An empty dict is legal (no scaffold).
|
||||
created_at — UTC generation timestamp.
|
||||
"""
|
||||
|
||||
__tablename__ = "variant_record"
|
||||
# The composite PK covers (learner_id, template_id) point lookups; the
|
||||
# unique task_id covers get_by_task (the grading/telemetry join path);
|
||||
# the two secondary indexes cover list_for_learner / list_by_template
|
||||
# ordered by created_at without a sort step (Postgres target D-027).
|
||||
__table_args__ = (
|
||||
UniqueConstraint("task_id", name="uq_variant_record_task_id"),
|
||||
Index("ix_variant_record_learner_created", "learner_id", "created_at"),
|
||||
Index("ix_variant_record_template_created", "template_id", "created_at"),
|
||||
)
|
||||
|
||||
learner_id: str = Field(primary_key=True)
|
||||
template_id: str = Field(primary_key=True)
|
||||
task_id: str
|
||||
seed: str
|
||||
# JSON columns: stored as TEXT on SQLite, native JSONB on Postgres (D-027).
|
||||
params: dict[str, Any] = Field(default_factory=dict, sa_type=JSON)
|
||||
statement: str
|
||||
starter_files: dict[str, Any] = Field(default_factory=dict, sa_type=JSON)
|
||||
created_at: datetime
|
||||
|
||||
@validates("learner_id", "template_id", "task_id")
|
||||
def _ids_non_empty(self, key: str, value: str) -> str:
|
||||
if not value:
|
||||
raise ValueError(f"{key} must be a non-empty identifier")
|
||||
return value
|
||||
|
||||
@validates("seed")
|
||||
def _seed_non_empty(self, key: str, value: str) -> str:
|
||||
if not value:
|
||||
raise ValueError(f"{key} must be a non-empty seed string")
|
||||
return value
|
||||
|
||||
@validates("statement")
|
||||
def _statement_non_empty(self, key: str, value: str) -> str:
|
||||
if not value:
|
||||
raise ValueError(f"{key} must be a non-empty statement string")
|
||||
return value
|
||||
|
||||
|
||||
class VariantStore(Protocol):
|
||||
"""Persistence contract for reproducible per-learner task variants.
|
||||
|
||||
Implemented by SQLiteVariantStore (v0.3, D-027); a Postgres
|
||||
implementation must satisfy the same surface.
|
||||
"""
|
||||
|
||||
def save(self, variant: VariantRecord) -> None:
|
||||
"""Persist a new variant. INSERT-ONLY on (learner_id, template_id):
|
||||
the FIRST generated variant is authoritative (reproducibility,
|
||||
D-029); a duplicate pair raises sqlalchemy.exc.IntegrityError to
|
||||
the caller — the generator serves cached variants via `get`
|
||||
instead of saving again. `task_id` is unique too: claiming an
|
||||
existing trace key for a different variant is equally rejected.
|
||||
NOT upsert; contrast GradeStore.save (latest-wins) and
|
||||
TraceStore.append (dedup-keep-first, swallowed).
|
||||
"""
|
||||
...
|
||||
|
||||
def get(self, learner_id: str, template_id: str) -> VariantRecord | None:
|
||||
"""The learner's stored variant for the template; None when none
|
||||
exists. Detached from any DB session — safe to pass across layers.
|
||||
"""
|
||||
...
|
||||
|
||||
def get_by_task(self, task_id: str) -> VariantRecord | None:
|
||||
"""The variant owning the task key (the grading/telemetry join
|
||||
path); None when none exists. Detached from any DB session.
|
||||
"""
|
||||
...
|
||||
|
||||
def list_for_learner(self, learner_id: str) -> list[VariantRecord]:
|
||||
"""All stored variants for the learner, ordered by created_at
|
||||
ascending (chronological; task_id breaks same-instant ties).
|
||||
Empty list when the learner has none.
|
||||
"""
|
||||
...
|
||||
|
||||
def list_by_template(self, template_id: str) -> list[VariantRecord]:
|
||||
"""All stored variants generated from the template — one row per
|
||||
learner — ordered by created_at ascending (chronological;
|
||||
learner_id breaks same-instant ties). Empty list when the
|
||||
template has none. The proctoring cross-check path (seed params
|
||||
per learner) reads through this.
|
||||
"""
|
||||
...
|
||||
|
||||
def close(self) -> None:
|
||||
"""Release DB connections. Store must not be used after close."""
|
||||
...
|
||||
|
||||
|
||||
def _sqlite_connect(dbapi_connection: sqlite3.Connection, _: object) -> None:
|
||||
"""Per-connection pragma setup (a-3). Mirrors telemetry/grading stores.
|
||||
|
||||
journal_mode=WAL — readers never block the single writer.
|
||||
synchronous=NORMAL — safe in WAL mode, avoids full fsync-per-commit.
|
||||
busy_timeout=5000 — retry briefly under contention instead of
|
||||
`OperationalError: database is locked`.
|
||||
"""
|
||||
cursor = dbapi_connection.cursor()
|
||||
cursor.execute("PRAGMA journal_mode=WAL")
|
||||
cursor.execute("PRAGMA synchronous=NORMAL")
|
||||
cursor.execute("PRAGMA busy_timeout=5000")
|
||||
cursor.close()
|
||||
|
||||
|
||||
def _as_utc(ts: datetime) -> datetime:
|
||||
"""Timestamps travel as naive datetime on SQLite, tz-aware elsewhere.
|
||||
|
||||
SQLite (via SQLModel) drops tzinfo; Postgres TIMESTAMP WITH TIME ZONE
|
||||
keeps it. Normalizing on the read path makes the store's contract
|
||||
tz-aware UTC regardless of the backend (D-027).
|
||||
"""
|
||||
if ts.tzinfo is None:
|
||||
return ts.replace(tzinfo=UTC) # naive UTC read-side: label as UTC
|
||||
return ts.astimezone(UTC)
|
||||
|
||||
|
||||
class SQLiteVariantStore:
|
||||
"""SQLite-backed VariantStore (SQLModel). Third protocol-wrapped store
|
||||
of the D-027 family (first: SQLiteTraceStore, second: SQLiteGradeStore).
|
||||
"""
|
||||
|
||||
def __init__(self, db_path: Path | None = None) -> None:
|
||||
self._db_path: Path = db_path if db_path is not None else Settings().db_path
|
||||
self._engine = create_engine(f"sqlite:///{self._db_path}")
|
||||
sa.event.listen(self._engine, "connect", _sqlite_connect)
|
||||
SQLModel.metadata.create_all(self._engine)
|
||||
|
||||
@contextmanager
|
||||
def _session(self) -> Iterator[Session]:
|
||||
# expire_on_commit=False: identical session behavior to the other
|
||||
# D-027 stores. save() never commits on the error path and the read
|
||||
# paths never commit, but a uniform flag across the family keeps
|
||||
# their detachment guarantees from diverging.
|
||||
with Session(self._engine, expire_on_commit=False) as session:
|
||||
yield session
|
||||
|
||||
def save(self, variant: VariantRecord) -> None:
|
||||
# Plain INSERT, no merge: overwriting a stored variant would swap a
|
||||
# learner's graded task underneath its trace and grade (audit
|
||||
# corruption), so a duplicate identity is a race or bug to SURFACE,
|
||||
# not paper over. The generator's cache path (get before generate)
|
||||
# makes duplicate saves a programming error, not a normal flow.
|
||||
# The trace store swallows its IntegrityError (dedup is the
|
||||
# contract there); the grade store merges (latest-wins is the
|
||||
# contract there); this store re-raises (first-wins is the
|
||||
# contract here).
|
||||
with self._session() as session:
|
||||
try:
|
||||
session.add(variant)
|
||||
session.commit()
|
||||
except sa.exc.IntegrityError:
|
||||
session.rollback()
|
||||
logger.debug(
|
||||
"variant insert rejected (identity already stored): "
|
||||
"learner=%s template=%s task=%s",
|
||||
variant.learner_id,
|
||||
variant.template_id,
|
||||
variant.task_id,
|
||||
)
|
||||
raise
|
||||
logger.debug(
|
||||
"variant saved: %s/%s task=%s seed=%s",
|
||||
variant.learner_id,
|
||||
variant.template_id,
|
||||
variant.task_id,
|
||||
variant.seed,
|
||||
)
|
||||
|
||||
def get(self, learner_id: str, template_id: str) -> VariantRecord | None:
|
||||
with self._session() as session:
|
||||
record = session.get(VariantRecord, (learner_id, template_id))
|
||||
if record is None:
|
||||
return None
|
||||
record.created_at = _as_utc(record.created_at)
|
||||
# Detach from the session: callers must not depend on
|
||||
# open-session ORM magic (lazy loads fail once it closes).
|
||||
session.expunge(record)
|
||||
return record
|
||||
|
||||
def get_by_task(self, task_id: str) -> VariantRecord | None:
|
||||
with self._session() as session:
|
||||
stmt = select(VariantRecord).where(VariantRecord.task_id == task_id)
|
||||
record = session.exec(stmt).first()
|
||||
if record is None:
|
||||
return None
|
||||
record.created_at = _as_utc(record.created_at)
|
||||
session.expunge(record)
|
||||
return record
|
||||
|
||||
def list_for_learner(self, learner_id: str) -> list[VariantRecord]:
|
||||
with self._session() as session:
|
||||
stmt = (
|
||||
select(VariantRecord)
|
||||
.where(VariantRecord.learner_id == learner_id)
|
||||
# Chronological; task_id is a deterministic tie-break for
|
||||
# variants stamped within the same instant.
|
||||
.order_by(VariantRecord.created_at, VariantRecord.task_id)
|
||||
)
|
||||
results = session.exec(stmt).all()
|
||||
for row in results:
|
||||
row.created_at = _as_utc(row.created_at)
|
||||
session.expunge(row)
|
||||
return list(results)
|
||||
|
||||
def list_by_template(self, template_id: str) -> list[VariantRecord]:
|
||||
with self._session() as session:
|
||||
stmt = (
|
||||
select(VariantRecord)
|
||||
.where(VariantRecord.template_id == template_id)
|
||||
# Chronological; learner_id is a deterministic tie-break.
|
||||
.order_by(VariantRecord.created_at, VariantRecord.learner_id)
|
||||
)
|
||||
results = session.exec(stmt).all()
|
||||
for row in results:
|
||||
row.created_at = _as_utc(row.created_at)
|
||||
session.expunge(row)
|
||||
return list(results)
|
||||
|
||||
def close(self) -> None:
|
||||
self._engine.dispose()
|
||||
@@ -0,0 +1,305 @@
|
||||
"""Task template library for seeded variant generation (D-029, REQ-3-005).
|
||||
|
||||
A `TaskTemplate` binds a competency (D-021-aligned corpus ID), a statement
|
||||
skeleton with `{slot}` placeholders, typed `ParameterSlot`s, difficulty-
|
||||
normalization rubric anchors (the expected feature envelope that bounds
|
||||
variant fairness in the a-5 envelope test — grader-prompt shipment is the
|
||||
tracked P4 follow-up; grading is variant-blind today), and starter-file
|
||||
scaffolds served into the sandbox workdir (wired in P6).
|
||||
|
||||
Slot sampling is PURE CODE: `random.Random(seed)` over typed slots — fully
|
||||
reproducible for a given seed, independent of the LLM. The LLM only renders
|
||||
the seeded slot values into the statement skeleton (D-020 defense).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import random
|
||||
import re
|
||||
from typing import Literal
|
||||
|
||||
from pydantic import BaseModel, ConfigDict, Field, field_validator
|
||||
|
||||
SlotType = Literal["enum", "int_range", "string_set"]
|
||||
|
||||
|
||||
class ParameterSlot(BaseModel):
|
||||
"""One typed fill-in for a statement skeleton."""
|
||||
|
||||
model_config = ConfigDict(frozen=True)
|
||||
|
||||
name: str = Field(min_length=1)
|
||||
type: SlotType
|
||||
values: list[str] = Field(default_factory=list) # enum/string_set options
|
||||
lo: int | None = None # int_range bounds
|
||||
hi: int | None = None
|
||||
|
||||
@field_validator("values")
|
||||
@classmethod
|
||||
def _values_nonempty_for_enums(cls, v: list[str], info) -> list[str]:
|
||||
if info.data.get("type") in ("enum", "string_set") and not v:
|
||||
raise ValueError(f"slot {info.data.get('name')!r} needs values")
|
||||
return v
|
||||
|
||||
def sample(self, rng: random.Random) -> str | int:
|
||||
"""Deterministic sample from the seeded RNG. Validated after sampling."""
|
||||
if self.type == "enum" or self.type == "string_set":
|
||||
return rng.choice(self.values)
|
||||
if self.type == "int_range":
|
||||
lo = self.lo if self.lo is not None else 0
|
||||
hi = self.hi if self.hi is not None else lo
|
||||
if hi < lo:
|
||||
raise ValueError(f"slot {self.name!r}: hi < lo")
|
||||
return rng.randint(lo, hi)
|
||||
raise ValueError(f"unsupported slot type: {self.type!r}")
|
||||
|
||||
def validate_value(self, value: str | int) -> bool:
|
||||
"""Is `value` schema-valid for this slot? (params JSON gate, a-5.)"""
|
||||
if self.type in ("enum", "string_set"):
|
||||
return isinstance(value, str) and value in self.values
|
||||
if self.type == "int_range":
|
||||
lo = self.lo if self.lo is not None else 0
|
||||
hi = self.hi if self.hi is not None else lo
|
||||
return isinstance(value, int) and lo <= value <= hi
|
||||
return False
|
||||
|
||||
|
||||
class RubricAnchors(BaseModel):
|
||||
"""Difficulty-normalization anchors for the grader (a-5).
|
||||
|
||||
Expected FEATURE ENVELOPE (digest-space): the expected effort band
|
||||
for this template, so two variants of one template are held to the
|
||||
same bar regardless of which slot values a learner drew. The a-5
|
||||
envelope test (tests/variants/test_generator.py) binds variants to
|
||||
these bands in code, and — since Phase 4 (MH#4) — the grading engine
|
||||
ships this envelope into the grader prompt
|
||||
(grading/engine._anchors_context) and stamps the variant seed on the
|
||||
GradeRecord, so the anchors gate variant fairness in BOTH tests and
|
||||
the live rubric.
|
||||
"""
|
||||
|
||||
model_config = ConfigDict(frozen=True)
|
||||
|
||||
expected_edit_count_band: tuple[int, int]
|
||||
expected_min_test_runs: int
|
||||
expected_error_fix_cycles_band: tuple[int, int]
|
||||
notes: str = ""
|
||||
|
||||
|
||||
class TaskTemplate(BaseModel):
|
||||
"""A reusable task shape; variants instantiate it per learner."""
|
||||
|
||||
model_config = ConfigDict(frozen=True)
|
||||
|
||||
id: str = Field(min_length=1)
|
||||
competency_id: str = Field(min_length=1) # D-021 corpus alignment
|
||||
title: str
|
||||
statement_skeleton: str = Field(min_length=1) # {slot} placeholders
|
||||
slots: list[ParameterSlot] = Field(min_length=1)
|
||||
rubric_anchors: RubricAnchors
|
||||
starter_files: dict[str, str] = Field(default_factory=dict) # path -> content
|
||||
test_command: str
|
||||
|
||||
@field_validator("statement_skeleton")
|
||||
@classmethod
|
||||
def _skeleton_placeholders(cls, v: str) -> str:
|
||||
if "{" not in v or "}" not in v:
|
||||
raise ValueError("statement_skeleton needs at least one {slot}")
|
||||
return v
|
||||
|
||||
def render(self, params: dict[str, str | int]) -> str:
|
||||
"""Fill the skeleton with validated params."""
|
||||
for slot in self.slots:
|
||||
if slot.name not in params:
|
||||
raise ValueError(f"missing param for slot {slot.name!r}")
|
||||
if not slot.validate_value(params[slot.name]):
|
||||
raise ValueError(f"invalid value for slot {slot.name!r}: {params[slot.name]!r}")
|
||||
return self.statement_skeleton.format(**params)
|
||||
|
||||
def sample_params(self, seed: int) -> dict[str, str | int]:
|
||||
"""Seeded, reproducible, schema-valid slot values (pure code)."""
|
||||
rng = random.Random(seed)
|
||||
return {slot.name: slot.sample(rng) for slot in self.slots}
|
||||
|
||||
|
||||
# --- Template library (v0.3 initial set) --------------------------------------
|
||||
# Competency IDs are D-021-aligned with the Python corpus
|
||||
# (ai_service/corpus/learner_context.py) and the TS mock-data layer
|
||||
# (packages/mock-data/competency-stacks.ts: deterministic cid() scheme).
|
||||
|
||||
TEMPLATES: dict[str, TaskTemplate] = {
|
||||
"tpl-llm-judge": TaskTemplate(
|
||||
id="tpl-llm-judge",
|
||||
competency_id="stack-orchestration-c007",
|
||||
title="Build an LLM-as-Judge Evaluator",
|
||||
statement_skeleton=(
|
||||
"Build a small LLM-as-judge evaluator for {domain} answers. "
|
||||
"The judge must score each answer on {criterion} using a 0-4 scale, "
|
||||
"return structured JSON, and handle at least {edge_cases} edge-case "
|
||||
"answer classes (empty, off-topic, adversarial). Include a tiny "
|
||||
"repro test set of at least {test_size} examples and print a summary "
|
||||
"table of scores."
|
||||
),
|
||||
slots=[
|
||||
ParameterSlot(
|
||||
name="domain",
|
||||
type="enum",
|
||||
values=["customer-support", "code-review", "summarization", "tutoring"],
|
||||
),
|
||||
ParameterSlot(
|
||||
name="criterion",
|
||||
type="enum",
|
||||
values=["factual-accuracy", "helpfulness", "safety", "completeness"],
|
||||
),
|
||||
ParameterSlot(name="edge_cases", type="int_range", lo=2, hi=4),
|
||||
ParameterSlot(name="test_size", type="int_range", lo=3, hi=8),
|
||||
],
|
||||
rubric_anchors=RubricAnchors(
|
||||
expected_edit_count_band=(3, 25),
|
||||
expected_min_test_runs=2,
|
||||
expected_error_fix_cycles_band=(0, 4),
|
||||
notes="Slot draw changes the SCENARIO, not the engineering depth.",
|
||||
),
|
||||
starter_files={
|
||||
"README.md": (
|
||||
"# LLM-as-Judge Evaluator\n\n"
|
||||
"Implement `judge.py`:\n"
|
||||
"- `score(answer: str) -> dict` — 0-4 on the named criterion\n"
|
||||
"- structured JSON output (schema below)\n"
|
||||
"- edge-case classes handled explicitly\n"
|
||||
"- `pytest` must pass\n"
|
||||
),
|
||||
"judge.py": "def score(answer: str) -> dict:\n raise NotImplementedError\n",
|
||||
"test_judge.py": "def test_placeholder():\n assert True\n",
|
||||
},
|
||||
test_command="pytest -q",
|
||||
),
|
||||
"tpl-guardrail-schema": TaskTemplate(
|
||||
id="tpl-guardrail-schema",
|
||||
competency_id="stack-orchestration-c008",
|
||||
title="Schema Guardrail Pipeline",
|
||||
statement_skeleton=(
|
||||
"Implement an output-validation guardrail for a model returning "
|
||||
"{entity} records. Validate against a typed schema with {field_count} "
|
||||
"required fields, coerce or reject {failure_mode} failures, and emit "
|
||||
"a fallback response for invalid payloads. Cover with at least "
|
||||
"{test_size} unit tests including malformed JSON."
|
||||
),
|
||||
slots=[
|
||||
ParameterSlot(
|
||||
name="entity",
|
||||
type="enum",
|
||||
values=["user-profile", "job-posting", "candidate", "invoice"],
|
||||
),
|
||||
ParameterSlot(
|
||||
name="failure_mode",
|
||||
type="enum",
|
||||
values=["strict-reject", "coerce-when-safe"],
|
||||
),
|
||||
ParameterSlot(name="field_count", type="int_range", lo=4, hi=8),
|
||||
ParameterSlot(name="test_size", type="int_range", lo=4, hi=10),
|
||||
],
|
||||
rubric_anchors=RubricAnchors(
|
||||
expected_edit_count_band=(3, 30),
|
||||
expected_min_test_runs=2,
|
||||
expected_error_fix_cycles_band=(0, 5),
|
||||
notes="All slot draws land in the same engineering band.",
|
||||
),
|
||||
starter_files={
|
||||
"README.md": (
|
||||
"# Schema Guardrail\n\nImplement `guardrail.py`:\n"
|
||||
"- `validate(payload: dict) -> dict | Fallback`\n"
|
||||
"- required-field checks, failure policy, fallback emission\n"
|
||||
),
|
||||
"guardrail.py": "def validate(payload: dict):\n raise NotImplementedError\n",
|
||||
"test_guardrail.py": "def test_placeholder():\n assert True\n",
|
||||
},
|
||||
test_command="pytest -q",
|
||||
),
|
||||
"tpl-rag-chunker": TaskTemplate(
|
||||
id="tpl-rag-chunker",
|
||||
competency_id="stack-orchestration-c005",
|
||||
title="RAG Chunking Strategy",
|
||||
statement_skeleton=(
|
||||
"Implement a document chunker for {doc_type} retrieval. Support "
|
||||
"{strategy} chunking with a target size of ~{chunk_size} tokens, "
|
||||
"preserve {invariant} across chunk boundaries, and evaluate overlap "
|
||||
"quality with at least {test_size} fixture documents."
|
||||
),
|
||||
slots=[
|
||||
ParameterSlot(
|
||||
name="doc_type",
|
||||
type="enum",
|
||||
values=["technical-docs", "legal-contracts", "transcripts"],
|
||||
),
|
||||
ParameterSlot(
|
||||
name="strategy",
|
||||
type="enum",
|
||||
values=["fixed-window", "semantic-boundary", "hybrid"],
|
||||
),
|
||||
ParameterSlot(
|
||||
name="invariant",
|
||||
type="enum",
|
||||
values=["code-block-integrity", "section-headers", "sentence-completeness"],
|
||||
),
|
||||
ParameterSlot(name="chunk_size", type="int_range", lo=200, hi=800),
|
||||
ParameterSlot(name="test_size", type="int_range", lo=3, hi=6),
|
||||
],
|
||||
rubric_anchors=RubricAnchors(
|
||||
expected_edit_count_band=(4, 35),
|
||||
expected_min_test_runs=2,
|
||||
expected_error_fix_cycles_band=(0, 6),
|
||||
notes="Strategy draw changes implementation shape, not depth.",
|
||||
),
|
||||
starter_files={
|
||||
"README.md": (
|
||||
"# RAG Chunker\n\nImplement `chunker.py`:\n"
|
||||
"- `chunk(text: str) -> list[str]`\n- invariant preserved\n- tests green\n"
|
||||
),
|
||||
"chunker.py": "def chunk(text: str) -> list[str]:\n raise NotImplementedError\n",
|
||||
"test_chunker.py": "def test_placeholder():\n assert True\n",
|
||||
},
|
||||
test_command="pytest -q",
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
_KNOWN_COMPETENCY_IDS: set[str] = {
|
||||
# D-021: mirrored from ai_service/corpus/learner_context.py — the Python
|
||||
# source of truth for stack-orchestration competencies used by v0.2 agents.
|
||||
"stack-orchestration-c001",
|
||||
"stack-orchestration-c002",
|
||||
"stack-orchestration-c003",
|
||||
"stack-orchestration-c004",
|
||||
"stack-orchestration-c005",
|
||||
"stack-orchestration-c007",
|
||||
"stack-orchestration-c008",
|
||||
"stack-orchestration-c011",
|
||||
"stack-designer-c001",
|
||||
"stack-designer-c002",
|
||||
"stack-safety-c021",
|
||||
}
|
||||
|
||||
|
||||
def get_template(template_id: str) -> TaskTemplate | None:
|
||||
return TEMPLATES.get(template_id)
|
||||
|
||||
|
||||
def template_for_competency(competency_id: str) -> list[TaskTemplate]:
|
||||
return [t for t in TEMPLATES.values() if t.competency_id == competency_id]
|
||||
|
||||
|
||||
def validate_competency_binding() -> None:
|
||||
"""All templates must bind to known D-021 corpus competency IDs."""
|
||||
for t in TEMPLATES.values():
|
||||
if t.competency_id not in _KNOWN_COMPETENCY_IDS:
|
||||
raise ValueError(
|
||||
f"template {t.id!r} binds unknown competency {t.competency_id!r}"
|
||||
)
|
||||
|
||||
|
||||
def slots_pattern_ok(skeleton: str, slots: list[ParameterSlot]) -> bool:
|
||||
"""Every {placeholder} in the skeleton has a matching slot and vice versa."""
|
||||
placeholders = set(re.findall(r"\{([a-z_][a-z0-9_]*)\}", skeleton))
|
||||
slot_names = {s.name for s in slots}
|
||||
return placeholders == slot_names
|
||||
@@ -0,0 +1,35 @@
|
||||
"""VoiceProvider protocol (D-030, REQ-3-006) — mirrors the LLMProvider seam.
|
||||
|
||||
Two implementations in v0.3:
|
||||
- MockVoiceProvider — deterministic canned transcripts + canned tone WAV
|
||||
chunks + scripted failure modes (tests + no-key default; tests NEVER call
|
||||
a real voice API).
|
||||
- browser descriptor — not a provider but a FALLBACK HINT: the web client
|
||||
selects browser-native SpeechRecognition/speechSynthesis when the server
|
||||
reports no real voice backend.
|
||||
|
||||
OpenAIAudioProvider (real server STT/TTS over OpenAI-compatible
|
||||
/audio/transcriptions + /audio/speech) is INTENTIONALLY NOT BUILT in v0.3 —
|
||||
deferred to v0.4 with KYC, when there is a real key and real users
|
||||
(GRILL CUT-1 / G-7). This protocol is its future drop-in seam.
|
||||
|
||||
Boundary: `voice/` never imports `agents/` or `api/`.
|
||||
"""
|
||||
|
||||
from ai_service.voice.base import (
|
||||
TranscriptSegment,
|
||||
VoiceDescriptor,
|
||||
VoiceProvider,
|
||||
)
|
||||
from ai_service.voice.browser import BROWSER_FALLBACK_DESCRIPTOR
|
||||
from ai_service.voice.factory import voice_provider_from_settings
|
||||
from ai_service.voice.mock import MockVoiceProvider
|
||||
|
||||
__all__ = [
|
||||
"BROWSER_FALLBACK_DESCRIPTOR",
|
||||
"MockVoiceProvider",
|
||||
"TranscriptSegment",
|
||||
"VoiceDescriptor",
|
||||
"VoiceProvider",
|
||||
"voice_provider_from_settings",
|
||||
]
|
||||
@@ -0,0 +1,62 @@
|
||||
"""VoiceProvider protocol + shared voice contracts (D-030, REQ-3-006).
|
||||
|
||||
Mirrors the LLMProvider seam (D-014 pattern): a narrow protocol the Examiner
|
||||
agent and the defense API compose via DI, with a deterministic mock and a
|
||||
browser-fallback descriptor. No network in this module — concrete providers
|
||||
live in their own modules and are selected by factory/config.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from collections.abc import AsyncIterator
|
||||
from typing import Literal, Protocol, runtime_checkable
|
||||
|
||||
from pydantic import BaseModel, ConfigDict, Field
|
||||
|
||||
VoiceRole = Literal["examiner", "learner"]
|
||||
|
||||
|
||||
class TranscriptSegment(BaseModel):
|
||||
"""One STT result: the transcribed text + timing metadata."""
|
||||
|
||||
model_config = ConfigDict(frozen=True)
|
||||
|
||||
text: str = Field(min_length=1)
|
||||
language: str = "en"
|
||||
duration_ms: int | None = None
|
||||
confidence: float | None = Field(default=None, ge=0.0, le=1.0)
|
||||
|
||||
|
||||
class VoiceDescriptor(BaseModel):
|
||||
"""Capability descriptor served to the web client (D-030).
|
||||
|
||||
The assessment UI reads this to decide HOW the learner speaks/hears:
|
||||
- `mode="server"` → server-side STT/TTS (v0.4 real provider seam)
|
||||
- `mode="browser"` → browser-native SpeechRecognition/speechSynthesis
|
||||
- `mode="mock"` → deterministic no-op path (tests / no-key dev)
|
||||
The descriptor never contains secrets — only capability hints.
|
||||
"""
|
||||
|
||||
model_config = ConfigDict(frozen=True)
|
||||
|
||||
mode: Literal["server", "browser", "mock"]
|
||||
sr_available: bool
|
||||
tts_available: bool
|
||||
hint: str = ""
|
||||
|
||||
|
||||
@runtime_checkable
|
||||
class VoiceProvider(Protocol):
|
||||
"""The voice port (D-030): STT in, TTS out. Never imports agents/api."""
|
||||
|
||||
async def transcribe(self, audio: bytes, fmt: str) -> TranscriptSegment:
|
||||
"""STT: audio bytes (fmt: 'wav' | 'webm' | 'mp3') → transcript."""
|
||||
...
|
||||
|
||||
def synthesize(self, text: str, voice: str = "default") -> AsyncIterator[bytes]:
|
||||
"""TTS: text -> async byte chunks (audio stream).
|
||||
|
||||
Implementations may be async generators (async-def + yield) — the
|
||||
consumer contract is `async for chunk in provider.synthesize(text)`.
|
||||
"""
|
||||
...
|
||||
@@ -0,0 +1,32 @@
|
||||
"""Browser-native fallback descriptor (D-030, CUT-1 / G-7, REQ-3-006).
|
||||
|
||||
v0.3 has NO real server STT/TTS (deferred to v0.4 with KYC/keys — GRILL
|
||||
CUT-1). When the factory selects `browser` mode, the defense endpoints return
|
||||
this descriptor and the WEB CLIENT performs SpeechRecognition + speechSynthesis
|
||||
natively; the server persists text turns as usual.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from .base import VoiceDescriptor
|
||||
|
||||
BROWSER_FALLBACK_DESCRIPTOR = VoiceDescriptor(
|
||||
mode="browser",
|
||||
sr_available=True,
|
||||
tts_available=True,
|
||||
hint=(
|
||||
"No server voice backend configured. Use browser-native "
|
||||
"SpeechRecognition for STT and speechSynthesis for TTS; send the "
|
||||
"transcribed text to POST /v1/defense/{id}/answer ({text} form)."
|
||||
),
|
||||
)
|
||||
|
||||
MOCK_DESCRIPTOR = VoiceDescriptor(
|
||||
mode="mock",
|
||||
sr_available=True,
|
||||
tts_available=True,
|
||||
hint=(
|
||||
"Deterministic mock voice (tests / no-key dev). Server STT/TTS "
|
||||
"endpoints serve canned responses; real server STT/TTS lands in v0.4."
|
||||
),
|
||||
)
|
||||
@@ -0,0 +1,506 @@
|
||||
"""DefenseStore — oral-defense persistence: protocol + SQLite impl (REQ-3-006, D-027).
|
||||
|
||||
FOURTH protocol-wrapped store of the D-027 family and the first spanning
|
||||
TWO related tables: `defense_record` (the defense session + integrity
|
||||
signals) and `defense_turn` (the ordered examiner/learner transcript,
|
||||
FK → defense_record.id).
|
||||
|
||||
Postgres-migration-ready (D-027): the protocol is the only surface the
|
||||
Examiner pipeline (task 5-2-01) and the defense endpoints (task 5-3-01)
|
||||
touch; swapping SQLiteDefenseStore for a Postgres implementation must
|
||||
not change call sites. Both tables use only portable column types
|
||||
(str / int / datetime / JSON), so the same SQLModel schema stands up
|
||||
unchanged on Postgres.
|
||||
|
||||
Save semantics — where this sits among the D-027 stores (each has a
|
||||
deliberately different contract):
|
||||
TraceStore.append dedup-keep-first; IntegrityError SWALLOWED
|
||||
(at-least-once event ingest).
|
||||
GradeStore.save upsert-latest-wins (a regrade is latest-state).
|
||||
VariantStore.save insert-only first-wins; IntegrityError RAISED
|
||||
(reproducibility; a duplicate is a bug).
|
||||
DefenseStore a LIFECYCLE store:
|
||||
start() insert-only; a duplicate id raises
|
||||
(a defense id is minted once per session).
|
||||
append_turn() insert-only per (defense_id, seq); a duplicate
|
||||
seq raises AND an unknown defense_id raises (FK
|
||||
enforced) — a transcript turn must never silently
|
||||
vanish (it is the integrity/grading input) nor
|
||||
attach to a defense that does not exist.
|
||||
finalize() targeted UPDATE (status → finished; finished_at +
|
||||
integrity_signals JSON). Unknown id → None
|
||||
(documented below). Re-finalize overwrites
|
||||
signals + finished_at — latest-wins, mirroring
|
||||
GradeStore.save: a recomputed verdict replaces
|
||||
the previous one wholesale.
|
||||
|
||||
append_turn does NOT police status (turns after finalize are a
|
||||
sequencing bug for the endpoints to prevent, task 5-3-01): the store
|
||||
enforces DATA integrity (FK + PK + non-empty), not workflow.
|
||||
|
||||
integrity_signals (A-109): JSON dict on the record — long pauses,
|
||||
off-scope cadence markers and friends, computed by the Examiner over
|
||||
turn metadata and persisted by finalize for the Proctor/Mentor feed.
|
||||
An empty dict is legal (defense not finished, or a clean defense).
|
||||
|
||||
Concurrency (a-3): WAL + synchronous=NORMAL + busy timeout at
|
||||
connection time (mirrors the other D-027 stores), PLUS foreign_keys=ON
|
||||
— this is the family's first real foreign key and it is actually
|
||||
enforced on SQLite, matching Postgres's native behavior (D-027 parity).
|
||||
|
||||
`created_at` / `ts` contract: callers stamp UTC (datetime.now(UTC));
|
||||
SQLite stores them naive and read paths re-label tz-aware UTC (same
|
||||
boundary normalization as TelemetryEvent.ts / GradeRecord.created_at,
|
||||
so the contract holds on any backend).
|
||||
|
||||
Boundary (D-027): `voice/` never imports `agents/` / `api/`; this
|
||||
module imports config only.
|
||||
"""
|
||||
|
||||
import logging
|
||||
import sqlite3
|
||||
from collections.abc import Iterator
|
||||
from contextlib import contextmanager
|
||||
from datetime import UTC, datetime
|
||||
from pathlib import Path
|
||||
from typing import Any, Literal, Protocol
|
||||
|
||||
import sqlalchemy as sa
|
||||
from sqlalchemy import JSON, Index, String
|
||||
from sqlalchemy.orm import validates
|
||||
from sqlmodel import Field, Session, SQLModel, create_engine, select
|
||||
|
||||
from ..config import Settings
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
DefenseStatus = Literal["in_progress", "finished"]
|
||||
_DEFENSE_STATUSES: frozenset[str] = frozenset(DefenseStatus.__args__)
|
||||
|
||||
TurnRole = Literal["examiner", "learner"]
|
||||
_TURN_ROLES: frozenset[str] = frozenset(TurnRole.__args__)
|
||||
|
||||
|
||||
class DefenseRecord(SQLModel, table=True):
|
||||
"""A persisted oral-defense session; id is the PK.
|
||||
|
||||
Written by the defense endpoints (task 5-3-01) through the
|
||||
DefenseStore protocol; read back by the endpoints, the Examiner
|
||||
pipeline and the Proctor/Mentor feeds. Constraint enforcement
|
||||
mirrors TelemetryEvent / GradeRecord / VariantRecord: sqlmodel
|
||||
0.0.42's metaclass drops pydantic constraints on table models, so
|
||||
SQLAlchemy `@validates` hooks enforce instead and the column types
|
||||
stay Postgres-ready (D-027).
|
||||
|
||||
Field contract:
|
||||
id — non-empty defense identifier, minted once
|
||||
per session (a duplicate start raises).
|
||||
learner_id — non-empty learner identifier (same id space
|
||||
as traces, grades and variants).
|
||||
task_id — non-empty task identifier; the defense
|
||||
defends the submitted work for this trace
|
||||
key ((learner_id, task_id) joins to the
|
||||
trace/grade/variant the defense is about).
|
||||
status — in_progress | finished; the STORE owns the
|
||||
transition: start() forces in_progress,
|
||||
finalize() sets finished. Validated.
|
||||
integrity_signals — A-109 signal dict (long pauses, off-scope
|
||||
cadence markers, ...); {} until finalize;
|
||||
JSON column. An empty dict is legal.
|
||||
created_at — UTC start timestamp.
|
||||
finished_at — UTC finalize timestamp; None while in
|
||||
progress.
|
||||
|
||||
`turns` (property): the seq-ordered DefenseTurn transcript, attached
|
||||
ONLY by DefenseStore.get(); records from list_for_learner carry
|
||||
turns == [] — call get() for a full transcript.
|
||||
"""
|
||||
|
||||
__tablename__ = "defense_record"
|
||||
# The id PK covers point lookups; this secondary index covers
|
||||
# list_for_learner ordered by created_at without a sort step
|
||||
# (Postgres migration target D-027).
|
||||
__table_args__ = (
|
||||
Index("ix_defense_record_learner_created", "learner_id", "created_at"),
|
||||
)
|
||||
|
||||
id: str = Field(primary_key=True)
|
||||
learner_id: str
|
||||
task_id: str
|
||||
# Bare Literal annotations crash sqlmodel<=0.0.42's column inference
|
||||
# (issubclass(TypeAlias, Enum)); an explicit sa_type + the validates
|
||||
# hook below give the same contract: VARCHAR column, Literal-rejected
|
||||
# values (same pattern as TelemetryEvent.kind).
|
||||
status: DefenseStatus = Field(default="in_progress", sa_type=String)
|
||||
# JSON column: stored as TEXT on SQLite, native JSONB on Postgres (D-027).
|
||||
integrity_signals: dict[str, Any] = Field(default_factory=dict, sa_type=JSON)
|
||||
created_at: datetime
|
||||
finished_at: datetime | None = Field(default=None)
|
||||
|
||||
@property
|
||||
def turns(self) -> list["DefenseTurn"]:
|
||||
"""Seq-ordered transcript; [] unless attached by get().
|
||||
|
||||
Table models reject ad-hoc attributes (pydantic __setattr__
|
||||
raises on non-fields), so the store stashes the detached turn
|
||||
list via object.__setattr__ and this read-only property surfaces
|
||||
it. The returned list is a copy — caller mutations cannot
|
||||
corrupt the stash.
|
||||
"""
|
||||
return list(self.__dict__.get("_turns", []))
|
||||
|
||||
@validates("id", "learner_id", "task_id")
|
||||
def _ids_non_empty(self, key: str, value: str) -> str:
|
||||
if not value:
|
||||
raise ValueError(f"{key} must be a non-empty identifier")
|
||||
return value
|
||||
|
||||
@validates("status")
|
||||
def _status_is_known(self, key: str, value: str) -> str:
|
||||
if value not in _DEFENSE_STATUSES:
|
||||
raise ValueError(f"unknown defense status: {value!r}")
|
||||
return value
|
||||
|
||||
|
||||
class DefenseTurn(SQLModel, table=True):
|
||||
"""One examiner/learner dialogue turn; (defense_id, seq) is the PK.
|
||||
|
||||
Rows are append-only transcript entries written through
|
||||
DefenseStore.append_turn. seq numbers the dialogue within one
|
||||
defense starting at 0; monotonic assignment is the endpoints' job
|
||||
(task 5-3-01), this model only rejects negatives — the same split
|
||||
as TelemetryEvent.seq (model rejects < 0, store owns ordering).
|
||||
|
||||
Field contract:
|
||||
defense_id — non-empty; FK → defense_record.id. ENFORCED on
|
||||
SQLite via foreign_keys=ON (first real FK in the
|
||||
D-027 family; Postgres enforces FKs natively, so
|
||||
this keeps the backends equivalent, D-027).
|
||||
seq — turn index within the defense, >= 0. (defense_id,
|
||||
seq) is the PK: a duplicate raises instead of
|
||||
silently overwriting — the transcript is the
|
||||
integrity/grading input, a vanishing turn is
|
||||
audit corruption.
|
||||
role — examiner | learner (who spoke). Validated.
|
||||
text — non-empty utterance text (examiner question, or
|
||||
STT output for learner answers).
|
||||
ts — UTC utterance timestamp.
|
||||
latency_ms — per-turn pipeline latency in ms (STT + LLM TTFT +
|
||||
TTS, A-109); int or None. Populated by the
|
||||
endpoints (task 5-4-01); None allowed here — the
|
||||
store persists, it does not measure.
|
||||
created_at — UTC row-write timestamp.
|
||||
"""
|
||||
|
||||
__tablename__ = "defense_turn"
|
||||
# The composite PK (defense_id, seq) doubles as the covering index
|
||||
# for the per-defense seq-ordered read in get() — no secondary index
|
||||
# needed (contrast defense_record's learner-listing index).
|
||||
|
||||
defense_id: str = Field(foreign_key="defense_record.id", primary_key=True)
|
||||
seq: int = Field(primary_key=True)
|
||||
role: TurnRole = Field(sa_type=String)
|
||||
text: str
|
||||
ts: datetime
|
||||
latency_ms: int | None = Field(default=None)
|
||||
created_at: datetime
|
||||
|
||||
@validates("defense_id")
|
||||
def _defense_id_non_empty(self, key: str, value: str) -> str:
|
||||
if not value:
|
||||
raise ValueError(f"{key} must be a non-empty identifier")
|
||||
return value
|
||||
|
||||
@validates("seq")
|
||||
def _seq_non_negative(self, key: str, value: int) -> int:
|
||||
if value < 0:
|
||||
raise ValueError("seq must be >= 0 (ordering is the endpoints' job)")
|
||||
return value
|
||||
|
||||
@validates("role")
|
||||
def _role_is_known(self, key: str, value: str) -> str:
|
||||
if value not in _TURN_ROLES:
|
||||
raise ValueError(f"unknown turn role: {value!r}")
|
||||
return value
|
||||
|
||||
@validates("text")
|
||||
def _text_non_empty(self, key: str, value: str) -> str:
|
||||
if not value:
|
||||
raise ValueError(f"{key} must be a non-empty utterance string")
|
||||
return value
|
||||
|
||||
@validates("latency_ms")
|
||||
def _latency_non_negative(self, key: str, value: int | None) -> int | None:
|
||||
# None is legal (not yet instrumented); a NEGATIVE latency is
|
||||
# nonsense and surfaces as a construction error.
|
||||
if value is not None and value < 0:
|
||||
raise ValueError("latency_ms must be >= 0 or None")
|
||||
return value
|
||||
|
||||
|
||||
class DefenseStore(Protocol):
|
||||
"""Persistence contract for oral-defense sessions + transcripts.
|
||||
|
||||
Implemented by SQLiteDefenseStore (v0.3, D-027); a Postgres
|
||||
implementation must satisfy the same surface.
|
||||
"""
|
||||
|
||||
def start(self, defense: DefenseRecord) -> DefenseRecord:
|
||||
"""Insert a new defense. INSERT-ONLY: a duplicate id raises
|
||||
sqlalchemy.exc.IntegrityError (a defense id is minted once per
|
||||
session — surfacing, not swallowing, mirrors VariantStore).
|
||||
The store owns the lifecycle: status is forced to "in_progress"
|
||||
and finished_at to None, whatever the caller passed — only
|
||||
finalize() may move a defense to finished. Returns the stored
|
||||
record, detached from any DB session.
|
||||
"""
|
||||
...
|
||||
|
||||
def append_turn(self, defense_id: str, turn: DefenseTurn) -> DefenseTurn:
|
||||
"""Insert one transcript turn, ordered by (defense_id, seq).
|
||||
turn.defense_id MUST equal the defense_id argument — a mismatch
|
||||
raises ValueError (the defense identity must never be
|
||||
ambiguous). A duplicate (defense_id, seq) raises
|
||||
IntegrityError; an unknown defense_id raises IntegrityError
|
||||
(FK enforced). Does NOT police status — sequencing turns vs
|
||||
finalize is the endpoints' job (task 5-3-01). Returns the
|
||||
stored turn, detached.
|
||||
"""
|
||||
...
|
||||
|
||||
def finalize(
|
||||
self, defense_id: str, integrity_signals: dict[str, Any]
|
||||
) -> DefenseRecord | None:
|
||||
"""Seal the defense: status → "finished", finished_at = now(UTC),
|
||||
integrity_signals stored as JSON. UNKNOWN defense_id → None
|
||||
(documented choice: the API layer maps it to 404 without an
|
||||
exception dance; contrast start/append_turn where IntegrityError
|
||||
IS the contract — those are inserts, this is an update on a key
|
||||
the caller may legitimately not hold). Re-finalize overwrites
|
||||
signals + finished_at: latest-wins, mirroring GradeStore.save
|
||||
(a recomputed verdict replaces the previous one wholesale).
|
||||
Returns the updated record, detached, WITHOUT turns — get() is
|
||||
the with-turns path.
|
||||
"""
|
||||
...
|
||||
|
||||
def get(self, defense_id: str) -> DefenseRecord | None:
|
||||
"""The defense with its FULL transcript (turns in seq order,
|
||||
detached) and integrity signals; None when it does not exist.
|
||||
Safe to pass across layers — no open-session ORM magic.
|
||||
"""
|
||||
...
|
||||
|
||||
def list_for_learner(self, learner_id: str) -> list[DefenseRecord]:
|
||||
"""All stored defenses for the learner, ordered by created_at
|
||||
ascending (chronological; id breaks same-instant ties), WITHOUT
|
||||
turns — records carry turns == []; call get() for a transcript.
|
||||
Empty list when the learner has none.
|
||||
"""
|
||||
...
|
||||
|
||||
def close(self) -> None:
|
||||
"""Release DB connections. Store must not be used after close."""
|
||||
...
|
||||
|
||||
|
||||
def _sqlite_connect(dbapi_connection: sqlite3.Connection, _: object) -> None:
|
||||
"""Per-connection pragma setup (a-3). Mirrors the other D-027 stores.
|
||||
|
||||
journal_mode=WAL — readers never block the single writer.
|
||||
synchronous=NORMAL — safe in WAL mode, avoids full fsync-per-commit.
|
||||
busy_timeout=5000 — retry briefly under contention instead of
|
||||
`OperationalError: database is locked`.
|
||||
foreign_keys=ON — NEW vs the family: defense_turn is the first
|
||||
real FK among the D-027 stores; SQLite leaves
|
||||
FKs OFF by default while Postgres enforces them
|
||||
natively, so the pragma keeps the backends
|
||||
equivalent (D-027 parity).
|
||||
"""
|
||||
cursor = dbapi_connection.cursor()
|
||||
cursor.execute("PRAGMA journal_mode=WAL")
|
||||
cursor.execute("PRAGMA synchronous=NORMAL")
|
||||
cursor.execute("PRAGMA busy_timeout=5000")
|
||||
cursor.execute("PRAGMA foreign_keys=ON")
|
||||
cursor.close()
|
||||
|
||||
|
||||
def _as_utc(ts: datetime) -> datetime:
|
||||
"""Timestamps travel as naive datetime on SQLite, tz-aware elsewhere.
|
||||
|
||||
SQLite (via SQLModel) drops tzinfo; Postgres TIMESTAMP WITH TIME ZONE
|
||||
keeps it. Normalizing on the read path makes the store's contract
|
||||
tz-aware UTC regardless of the backend (D-027).
|
||||
"""
|
||||
if ts.tzinfo is None:
|
||||
return ts.replace(tzinfo=UTC) # naive UTC read-side: label as UTC
|
||||
return ts.astimezone(UTC)
|
||||
|
||||
|
||||
class SQLiteDefenseStore:
|
||||
"""SQLite-backed DefenseStore (SQLModel). Fourth protocol-wrapped
|
||||
store of the D-027 family (first: SQLiteTraceStore, second:
|
||||
SQLiteGradeStore, third: SQLiteVariantStore) and the first spanning
|
||||
two related tables.
|
||||
"""
|
||||
|
||||
def __init__(self, db_path: Path | None = None) -> None:
|
||||
self._db_path: Path = db_path if db_path is not None else Settings().db_path
|
||||
self._engine = create_engine(f"sqlite:///{self._db_path}")
|
||||
sa.event.listen(self._engine, "connect", _sqlite_connect)
|
||||
SQLModel.metadata.create_all(self._engine)
|
||||
|
||||
@contextmanager
|
||||
def _session(self) -> Iterator[Session]:
|
||||
# expire_on_commit=False: identical session behavior to the other
|
||||
# D-027 stores. start/append_turn return the caller's instance
|
||||
# after commit and get/finalize return rows expunged mid-session;
|
||||
# a uniform flag across the family keeps their detachment
|
||||
# guarantees from diverging.
|
||||
with Session(self._engine, expire_on_commit=False) as session:
|
||||
yield session
|
||||
|
||||
def start(self, defense: DefenseRecord) -> DefenseRecord:
|
||||
# The store owns the lifecycle: a defense is BORN in_progress and
|
||||
# only finalize() may move it to finished. A smuggled "finished"
|
||||
# status is normalized away, not rejected — the insert itself
|
||||
# stays insert-only, and a duplicate id raises to the caller
|
||||
# (mirroring VariantStore: the id is minted once per session).
|
||||
defense.status = "in_progress"
|
||||
defense.finished_at = None
|
||||
with self._session() as session:
|
||||
try:
|
||||
session.add(defense)
|
||||
session.commit()
|
||||
except sa.exc.IntegrityError:
|
||||
session.rollback()
|
||||
logger.debug("defense start rejected (id already stored): %s", defense.id)
|
||||
raise
|
||||
logger.debug(
|
||||
"defense started: %s learner=%s task=%s",
|
||||
defense.id,
|
||||
defense.learner_id,
|
||||
defense.task_id,
|
||||
)
|
||||
return defense
|
||||
|
||||
def append_turn(self, defense_id: str, turn: DefenseTurn) -> DefenseTurn:
|
||||
# The explicit defense_id argument is the defense identity for
|
||||
# this write; a turn object claiming another defense is a
|
||||
# programming error — surface it before touching the DB.
|
||||
if turn.defense_id != defense_id:
|
||||
raise ValueError(
|
||||
f"turn.defense_id {turn.defense_id!r} does not match the "
|
||||
f"defense_id argument {defense_id!r}"
|
||||
)
|
||||
with self._session() as session:
|
||||
try:
|
||||
session.add(turn)
|
||||
session.commit()
|
||||
except sa.exc.IntegrityError:
|
||||
# Two possible causes, both surfaced, neither swallowed:
|
||||
# duplicate (defense_id, seq) PK — a transcript turn must
|
||||
# never silently vanish; unknown defense_id — the FK
|
||||
# (foreign_keys=ON) rejects the orphan.
|
||||
session.rollback()
|
||||
logger.debug(
|
||||
"defense turn rejected (duplicate (defense_id, seq) "
|
||||
"or unknown defense_id): defense=%s seq=%s",
|
||||
defense_id,
|
||||
turn.seq,
|
||||
)
|
||||
raise
|
||||
logger.debug(
|
||||
"defense turn appended: %s seq=%d role=%s",
|
||||
defense_id,
|
||||
turn.seq,
|
||||
turn.role,
|
||||
)
|
||||
return turn
|
||||
|
||||
def finalize(
|
||||
self, defense_id: str, integrity_signals: dict[str, Any]
|
||||
) -> DefenseRecord | None:
|
||||
# A None signals blob would break the read contract (signals are
|
||||
# a dict, {} until finalize); reject before writing.
|
||||
if not isinstance(integrity_signals, dict):
|
||||
raise ValueError(
|
||||
"integrity_signals must be a JSON-object dict, got "
|
||||
f"{type(integrity_signals).__name__}"
|
||||
)
|
||||
with self._session() as session:
|
||||
record = session.get(DefenseRecord, defense_id)
|
||||
if record is None:
|
||||
# Documented unknown-id behavior: None, not a raise — the
|
||||
# defense endpoints map this to 404. Contrast start() /
|
||||
# append_turn(), where IntegrityError IS the contract.
|
||||
return None
|
||||
# Latest-wins re-finalize, mirroring GradeStore.save: a
|
||||
# recomputed verdict (fresh signals) replaces the stored one
|
||||
# wholesale; status just stays finished.
|
||||
record.status = "finished"
|
||||
record.finished_at = datetime.now(UTC)
|
||||
record.integrity_signals = integrity_signals
|
||||
session.commit()
|
||||
record.created_at = _as_utc(record.created_at)
|
||||
if record.finished_at is not None:
|
||||
record.finished_at = _as_utc(record.finished_at)
|
||||
# Detach from the session: callers must not depend on
|
||||
# open-session ORM magic (lazy loads fail once it closes).
|
||||
session.expunge(record)
|
||||
logger.debug(
|
||||
"defense finalized: %s signals=%s", defense_id, sorted(integrity_signals)
|
||||
)
|
||||
return record
|
||||
|
||||
def get(self, defense_id: str) -> DefenseRecord | None:
|
||||
with self._session() as session:
|
||||
record = session.get(DefenseRecord, defense_id)
|
||||
if record is None:
|
||||
return None
|
||||
record.created_at = _as_utc(record.created_at)
|
||||
if record.finished_at is not None:
|
||||
record.finished_at = _as_utc(record.finished_at)
|
||||
stmt = (
|
||||
select(DefenseTurn)
|
||||
.where(DefenseTurn.defense_id == defense_id)
|
||||
.order_by(DefenseTurn.seq)
|
||||
)
|
||||
turns = session.exec(stmt).all()
|
||||
for turn in turns:
|
||||
turn.ts = _as_utc(turn.ts)
|
||||
turn.created_at = _as_utc(turn.created_at)
|
||||
# Detach each turn: the transcript must be usable once
|
||||
# the session closes (no lazy-load magic).
|
||||
session.expunge(turn)
|
||||
session.expunge(record)
|
||||
# Table models reject ad-hoc attributes (pydantic __setattr__
|
||||
# raises on non-fields), so the seq-ordered transcript is
|
||||
# stashed via object.__setattr__ and surfaced through the
|
||||
# read-only `turns` property. Rows are detached either way —
|
||||
# safe to pass across layers.
|
||||
object.__setattr__(record, "_turns", list(turns))
|
||||
return record
|
||||
|
||||
def list_for_learner(self, learner_id: str) -> list[DefenseRecord]:
|
||||
with self._session() as session:
|
||||
stmt = (
|
||||
select(DefenseRecord)
|
||||
.where(DefenseRecord.learner_id == learner_id)
|
||||
# Chronological; id is a deterministic tie-break for
|
||||
# defenses stamped within the same instant.
|
||||
.order_by(DefenseRecord.created_at, DefenseRecord.id)
|
||||
)
|
||||
results = session.exec(stmt).all()
|
||||
for row in results:
|
||||
row.created_at = _as_utc(row.created_at)
|
||||
if row.finished_at is not None:
|
||||
row.finished_at = _as_utc(row.finished_at)
|
||||
# Turns are deliberately NOT loaded here: the list feed
|
||||
# (Proctor/Mentor) needs session headers, not full
|
||||
# transcripts — get() is the with-turns path.
|
||||
session.expunge(row)
|
||||
return list(results)
|
||||
|
||||
def close(self) -> None:
|
||||
self._engine.dispose()
|
||||
@@ -0,0 +1,37 @@
|
||||
"""Voice provider factory (D-030, REQ-3-006).
|
||||
|
||||
`AI_VOICE_PROVIDER = browser | mock` (default: mock — the no-key path is
|
||||
first-class). The real server provider (`openai-audio`) is a v0.4 seam and
|
||||
is REJECTED here with a clear error naming the deferral, so a stale env var
|
||||
can't silently pretend a real backend exists.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from ..config import Settings
|
||||
from .base import VoiceProvider
|
||||
from .mock import MockVoiceProvider
|
||||
|
||||
|
||||
class UnknownVoiceProviderError(ValueError):
|
||||
"""Raised for a provider name outside the v0.3 contract."""
|
||||
|
||||
|
||||
def voice_provider_from_settings(settings: Settings) -> VoiceProvider:
|
||||
"""Select the voice provider by settings (env `AI_VOICE_PROVIDER`)."""
|
||||
name = (settings.voice_provider or "mock").strip().lower()
|
||||
if name == "mock":
|
||||
return MockVoiceProvider()
|
||||
if name == "browser":
|
||||
# Browser mode is a CLIENT-side capability: the server composes the
|
||||
# same MockVoiceProvider (typed fallback answers still work; the UI
|
||||
# uses the descriptor for mic/speech). See browser.py.
|
||||
return MockVoiceProvider()
|
||||
if name in ("openai-audio", "openai", "server"):
|
||||
raise UnknownVoiceProviderError(
|
||||
"real server STT/TTS (OpenAIAudioProvider) is deferred to v0.4 "
|
||||
"(GRILL CUT-1 / G-7): set AI_VOICE_PROVIDER=mock or browser"
|
||||
)
|
||||
raise UnknownVoiceProviderError(
|
||||
f"unknown AI_VOICE_PROVIDER {name!r}: use 'mock' or 'browser'"
|
||||
)
|
||||
@@ -0,0 +1,96 @@
|
||||
"""Deterministic MockVoiceProvider (D-030, REQ-3-006).
|
||||
|
||||
Canned transcripts (scripted per test via queue) + canned 1kHz-tone WAV bytes
|
||||
+ scripted failure modes. Two identical transcribe calls yield identical
|
||||
segments; tests NEVER touch a real voice API (conftest cloud-free rule).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import io
|
||||
import math
|
||||
import struct
|
||||
import wave
|
||||
from collections.abc import AsyncIterator
|
||||
|
||||
from .base import TranscriptSegment
|
||||
|
||||
|
||||
def _tone_wav(duration_ms: int = 250, freq_hz: float = 1000.0) -> bytes:
|
||||
"""A small, deterministic 16-bit mono WAV: a sine tone (stdlib only)."""
|
||||
rate = 8000
|
||||
n_samples = max(1, int(rate * duration_ms / 1000))
|
||||
buf = io.BytesIO()
|
||||
with wave.open(buf, "wb") as w:
|
||||
w.setnchannels(1)
|
||||
w.setsampwidth(2)
|
||||
w.setframerate(rate)
|
||||
for i in range(n_samples):
|
||||
sample = int(12000 * math.sin(2 * math.pi * freq_hz * i / rate))
|
||||
w.writeframes(struct.pack("<h", sample))
|
||||
return buf.getvalue()
|
||||
|
||||
|
||||
class MockVoiceFailure(RuntimeError):
|
||||
"""Scripted failure mode for tests."""
|
||||
|
||||
|
||||
class MockVoiceProvider:
|
||||
"""Deterministic voice provider: scripted STT, canned-tone TTS.
|
||||
|
||||
- `transcribe`: pops the next scripted transcript from a queue (or a
|
||||
default); two identical calls with the same queue state are identical.
|
||||
Failure mode: raise MockVoiceFailure when the queue holds a failure
|
||||
marker (the string "FAIL") or `audio` is empty.
|
||||
- `synthesize`: yields the canned tone WAV in fixed-size chunks; failure
|
||||
mode: empty text raises MockVoiceFailure.
|
||||
"""
|
||||
|
||||
def __init__(self, transcripts: list[str] | None = None) -> None:
|
||||
self._transcripts = list(transcripts or [])
|
||||
self._cursor = 0
|
||||
self.transcribe_calls = 0
|
||||
self.synthesize_calls = 0
|
||||
|
||||
def script(self, transcripts: list[str]) -> None:
|
||||
"""Replace the scripted queue (tests set expectations up front)."""
|
||||
self._transcripts = list(transcripts)
|
||||
self._cursor = 0
|
||||
|
||||
async def transcribe(self, audio: bytes, fmt: str) -> TranscriptSegment:
|
||||
self.transcribe_calls += 1
|
||||
if not audio:
|
||||
raise MockVoiceFailure("no audio bytes provided")
|
||||
if not self._transcripts:
|
||||
raise MockVoiceFailure("transcript queue exhausted — script it")
|
||||
item = self._transcripts[self._cursor]
|
||||
self._cursor = (self._cursor + 1) % len(self._transcripts)
|
||||
if item == "FAIL":
|
||||
raise MockVoiceFailure("scripted STT failure")
|
||||
return TranscriptSegment(
|
||||
text=item,
|
||||
duration_ms=max(1, len(audio) // 32), # deterministic pseudo-duration
|
||||
)
|
||||
|
||||
async def synthesize(self, text: str, voice: str = "default") -> AsyncIterator[bytes]: # noqa: ASYNC109 (protocol parity)
|
||||
# NOTE: protocol parity matters more than the async-generator purity
|
||||
# lint; the real provider seam (v0.4) will stream over HTTP.
|
||||
self.synthesize_calls += 1
|
||||
if not text:
|
||||
raise MockVoiceFailure("cannot synthesize empty text")
|
||||
wav = _tone_wav(duration_ms=min(2000, max(120, len(text) * 12)))
|
||||
for i in range(0, len(wav), 1024):
|
||||
yield wav[i : i + 1024]
|
||||
await asyncio.sleep(0) # yield to the loop like a network stream
|
||||
|
||||
|
||||
# Protocol-shape parity guard (mock must satisfy the D-030 port).
|
||||
from .base import VoiceProvider # noqa: E402
|
||||
|
||||
|
||||
def _assert_protocol() -> None:
|
||||
assert isinstance(MockVoiceProvider(), VoiceProvider)
|
||||
|
||||
|
||||
_assert_protocol()
|
||||
@@ -18,6 +18,9 @@ dependencies = [
|
||||
"sqlalchemy>=2.0,<2.1",
|
||||
"websockets>=13,<16",
|
||||
"aiofiles>=24.1,<26",
|
||||
# POST /v1/defense/{id}/answer multipart audio (REQ-3-006): FastAPI
|
||||
# form/File parsing requires python-multipart at runtime.
|
||||
"python-multipart>=0.0.32,<0.1",
|
||||
]
|
||||
|
||||
[project.optional-dependencies]
|
||||
|
||||
@@ -1,28 +1,73 @@
|
||||
#!/usr/bin/env bash
|
||||
# Idempotent bootstrap: create venv + install deps.
|
||||
# Handles Debian systems without python3-venv/ensurepip via --without-pip + get-pip.
|
||||
# Handles Debian/Ubuntu systems without python3-venv/ensurepip via --without-pip + get-pip.
|
||||
# v2 (v0.3.5): recovers from a poisoned partial .venv left by a failed earlier
|
||||
# attempt, cleans before each retry, and dies with a distro-specific fix hint
|
||||
# when venv creation is impossible (e.g. missing python3.XX-venv package).
|
||||
set -euo pipefail
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
APP_DIR="$(cd "$SCRIPT_DIR/.." && pwd)"
|
||||
VENV="$APP_DIR/.venv"
|
||||
|
||||
venv_usable() {
|
||||
[ -x "$VENV/bin/python3" ]
|
||||
}
|
||||
|
||||
rm_broken_venv() {
|
||||
echo "bootstrap: removing broken partial .venv from a failed earlier attempt" >&2
|
||||
rm -rf "$VENV"
|
||||
}
|
||||
|
||||
mkdir -p "$HOME/.cache/ciagent"
|
||||
|
||||
if [ ! -x "$VENV/bin/python3" ]; then
|
||||
if python3 -m venv "$VENV" 2>/dev/null; then
|
||||
if venv_usable && [ ! -x "$VENV/bin/pip" ]; then
|
||||
# A usable python3 without pip means the --without-pip fallback half-ran and
|
||||
# the get-pip step never completed: start over cleanly.
|
||||
rm_broken_venv
|
||||
fi
|
||||
|
||||
if ! venv_usable; then
|
||||
if [ -d "$VENV" ]; then
|
||||
# Directory exists but no working python3: remains of a crashed venv create.
|
||||
rm_broken_venv
|
||||
fi
|
||||
if python3 -m venv "$VENV" 2>/tmp/venv-create.err; then
|
||||
:
|
||||
else
|
||||
# No ensurepip available — create bare venv and bootstrap pip separately.
|
||||
python3 -m venv --without-pip "$VENV"
|
||||
rm -rf "$VENV"
|
||||
if python3 -m venv --without-pip "$VENV" 2>>/tmp/venv-create.err; then
|
||||
:
|
||||
else
|
||||
rm -rf "$VENV"
|
||||
PYVER="$(python3 -c 'import sys; print("%d.%d" % sys.version_info[:2])' 2>/dev/null || true)"
|
||||
PKG="python3-venv"
|
||||
[ -n "$PYVER" ] && PKG="python${PYVER}-venv"
|
||||
echo "bootstrap: could not create a virtual environment." >&2
|
||||
echo " python3 reported:" >&2
|
||||
sed 's/^/ /' /tmp/venv-create.err >&2 || true
|
||||
echo " fix (Debian/Ubuntu): install the venv support package, then re-run nextcraft bootstrap:" >&2
|
||||
echo " apt install $PKG" >&2
|
||||
exit 1
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
|
||||
if [ ! -x "$VENV/bin/pip" ]; then
|
||||
GET_PIP="$HOME/.cache/ciagent/get-pip.py"
|
||||
if [ ! -f "$GET_PIP" ]; then
|
||||
curl -sSf --max-time 60 https://bootstrap.pypa.io/get-pip.py -o "$GET_PIP"
|
||||
if ! curl -sSf --max-time 60 https://bootstrap.pypa.io/get-pip.py -o "$GET_PIP"; then
|
||||
rm -rf "$VENV"
|
||||
echo "bootstrap: get-pip.py download failed (no network?)." >&2
|
||||
echo " fix: restore network access and re-run nextcraft bootstrap" >&2
|
||||
exit 1
|
||||
fi
|
||||
fi
|
||||
if ! "$VENV/bin/python3" "$GET_PIP" --quiet; then
|
||||
rm -rf "$VENV"
|
||||
echo "bootstrap: pip installation into the venv failed." >&2
|
||||
echo " fix: re-run nextcraft bootstrap (the venv was cleaned; this retry is safe)" >&2
|
||||
exit 1
|
||||
fi
|
||||
"$VENV/bin/python3" "$GET_PIP" --quiet
|
||||
fi
|
||||
|
||||
"$VENV/bin/pip" install --quiet --upgrade pip
|
||||
|
||||
@@ -1,5 +1,7 @@
|
||||
#!/usr/bin/env bash
|
||||
# Dev server: export secrets (if present) then run uvicorn on :8420.
|
||||
# Dev server: export secrets (if present) then run uvicorn.
|
||||
# Binds 0.0.0.0 by default so the stack is reachable from other machines
|
||||
# (v0.3.5 network mode) — set AI_HOST=127.0.0.1 in .env to revert to loopback.
|
||||
set -euo pipefail
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
APP_DIR="$(cd "$SCRIPT_DIR/.." && pwd)"
|
||||
@@ -7,7 +9,7 @@ REPO_ROOT="$(cd "$APP_DIR/../.." && pwd)"
|
||||
VENV="$APP_DIR/.venv"
|
||||
|
||||
if [ ! -x "$VENV/bin/uvicorn" ]; then
|
||||
echo "venv missing — run scripts/bootstrap.sh first" >&2
|
||||
echo "venv missing — run nextcraft bootstrap first (or: bash scripts/bootstrap.sh)" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
@@ -22,5 +24,17 @@ if [ -f "$SECRETS" ]; then
|
||||
done < "$SECRETS"
|
||||
fi
|
||||
|
||||
ENV_FILE="$APP_DIR/.env"
|
||||
if [ -f "$ENV_FILE" ]; then
|
||||
while IFS='=' read -r key value; do
|
||||
case "$key" in
|
||||
AI_HOST|AI_PORT|AI_CORS_ORIGINS) export "$key=$value" ;;
|
||||
esac
|
||||
done < "$ENV_FILE"
|
||||
fi
|
||||
|
||||
HOST="${AI_HOST:-0.0.0.0}"
|
||||
PORT="${AI_PORT:-8420}"
|
||||
|
||||
cd "$APP_DIR"
|
||||
exec "$VENV/bin/uvicorn" ai_service.main:app --reload --port 8420
|
||||
exec "$VENV/bin/uvicorn" ai_service.main:app --reload --host "$HOST" --port "$PORT"
|
||||
@@ -0,0 +1,691 @@
|
||||
#!/usr/bin/env python3
|
||||
"""sandbox-agent — stdlib-only in-sandbox telemetry capture agent (REQ-3-003).
|
||||
|
||||
D-031: this file is copied into the sandbox namespace and runs against the
|
||||
system Python — no third-party packages are importable there, so this module
|
||||
depends on the standard library ONLY (the test suite enforces this with an
|
||||
AST scan of the file).
|
||||
|
||||
What it does:
|
||||
* wraps a non-interactive `/bin/sh` REPL: each stdin line is executed via
|
||||
`sh -c` inside the workspace and reported as `stdin` -> `command` ->
|
||||
`stdout` -> `run_result`/`test_result` events;
|
||||
* polls the workspace tree (~250 ms) and emits `file_diff` events
|
||||
(created/modified/deleted with unified diffs) plus periodic `activity`
|
||||
heartbeats;
|
||||
* streams events to ai-service as TelemetryEvent-shaped JSON frames over a
|
||||
raw-socket RFC 6455 WebSocket client (no `websockets` package exists in
|
||||
the namespace — the client handshake + frame codec is implemented here);
|
||||
* at-least-once delivery (D-026): every event is appended to an fsync'd
|
||||
JSONL spool file inside the workdir BEFORE any send attempt; on
|
||||
disconnect the spool grows; after reconnect (exponential backoff) the
|
||||
spool is flushed oldest-first. The server dedups on (learner, task, seq)
|
||||
so replayed duplicates are harmless — loss is not tolerated.
|
||||
|
||||
Configured entirely through env baked at spawn time:
|
||||
NC_LEARNER_ID / NC_TASK_ID / NC_INGEST_URL / NC_SANDBOX_ID (required)
|
||||
NC_WORKSPACE workspace root to watch/run in (default: cwd)
|
||||
NC_SPOOL spool path (default: <workspace>/.nc-agent/spool.jsonl)
|
||||
NC_POLL_INTERVAL_S / NC_ACTIVITY_INTERVAL_S / NC_COMMAND_TIMEOUT_S
|
||||
NC_BACKOFF_BASE_S / NC_BACKOFF_MAX_S (optional knobs)
|
||||
|
||||
Sequencing survives process restarts (incl. SIGKILL): on boot the spool is
|
||||
replayed into the pending queue and `seq` resumes at max(spooled seq) + 1.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import base64
|
||||
import difflib
|
||||
import hashlib
|
||||
import json
|
||||
import os
|
||||
import secrets
|
||||
import socket
|
||||
import ssl
|
||||
import struct
|
||||
import subprocess
|
||||
import sys
|
||||
import threading
|
||||
import time
|
||||
import urllib.parse
|
||||
from collections import deque
|
||||
from collections.abc import Mapping
|
||||
from dataclasses import dataclass
|
||||
from datetime import UTC, datetime
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
_WS_GUID = "258EAFA5-E914-47DA-95CA-C5AB0DC85B11"
|
||||
_EVENT_KINDS = frozenset(
|
||||
{"command", "file_diff", "run_result", "test_result", "activity", "stdin", "stdout"}
|
||||
)
|
||||
_AGENT_DIR_PREFIX = ".nc-" # agent-private paths (spool) are excluded from watching
|
||||
_MAX_DIFF_BYTES = 64 * 1024 # files larger than this are reported truncated, no diff
|
||||
_MAX_OUTPUT_CHARS = 64 * 1024 # captured stdout/stderr tail cap per command
|
||||
_HANDSHAKE_MAX_BYTES = 64 * 1024
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- config
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class AgentConfig:
|
||||
"""Runtime configuration, normally built from `NC_*` env baked at spawn."""
|
||||
|
||||
learner_id: str
|
||||
task_id: str
|
||||
ingest_url: str
|
||||
sandbox_id: str
|
||||
workspace: Path
|
||||
spool_path: Path
|
||||
poll_interval_s: float = 0.25
|
||||
activity_interval_s: float = 5.0
|
||||
command_timeout_s: float = 30.0
|
||||
backoff_base_s: float = 0.25
|
||||
backoff_max_s: float = 8.0
|
||||
|
||||
def __post_init__(self) -> None:
|
||||
for name in ("learner_id", "task_id", "ingest_url", "sandbox_id"):
|
||||
if not getattr(self, name):
|
||||
raise ValueError(f"missing required config: NC_{name.upper()}")
|
||||
|
||||
@classmethod
|
||||
def from_env(cls, env: Mapping[str, str] | None = None) -> AgentConfig:
|
||||
src = os.environ if env is None else env
|
||||
workspace = Path(src.get("NC_WORKSPACE") or os.getcwd()).resolve()
|
||||
return cls(
|
||||
learner_id=src.get("NC_LEARNER_ID", ""),
|
||||
task_id=src.get("NC_TASK_ID", ""),
|
||||
ingest_url=src.get("NC_INGEST_URL", ""),
|
||||
sandbox_id=src.get("NC_SANDBOX_ID", ""),
|
||||
workspace=workspace,
|
||||
spool_path=Path(
|
||||
src.get("NC_SPOOL") or (workspace / ".nc-agent" / "spool.jsonl")
|
||||
),
|
||||
poll_interval_s=float(src.get("NC_POLL_INTERVAL_S", "0.25")),
|
||||
activity_interval_s=float(src.get("NC_ACTIVITY_INTERVAL_S", "5.0")),
|
||||
command_timeout_s=float(src.get("NC_COMMAND_TIMEOUT_S", "30.0")),
|
||||
backoff_base_s=float(src.get("NC_BACKOFF_BASE_S", "0.25")),
|
||||
backoff_max_s=float(src.get("NC_BACKOFF_MAX_S", "8.0")),
|
||||
)
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- spool
|
||||
|
||||
|
||||
class Spool:
|
||||
"""Append-only JSONL spool with per-append fsync (survives SIGKILL).
|
||||
|
||||
`rewrite` swaps in a compacted file atomically (tmp file + os.replace).
|
||||
Lines are stored without trailing newlines in memory, one per line on disk.
|
||||
"""
|
||||
|
||||
def __init__(self, path: Path) -> None:
|
||||
self._path = path
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
@property
|
||||
def path(self) -> Path:
|
||||
return self._path
|
||||
|
||||
def append(self, line: str) -> None:
|
||||
with self._path.open("a", encoding="utf-8") as fh:
|
||||
fh.write(line + "\n")
|
||||
fh.flush()
|
||||
os.fsync(fh.fileno())
|
||||
|
||||
def read_all(self) -> list[str]:
|
||||
if not self._path.exists():
|
||||
return []
|
||||
with self._path.open("r", encoding="utf-8") as fh:
|
||||
return [line.rstrip("\n") for line in fh if line.strip()]
|
||||
|
||||
def rewrite(self, lines: list[str]) -> None:
|
||||
tmp = self._path.with_name(self._path.name + ".tmp")
|
||||
with tmp.open("w", encoding="utf-8") as fh:
|
||||
for line in lines:
|
||||
fh.write(line + "\n")
|
||||
fh.flush()
|
||||
os.fsync(fh.fileno())
|
||||
os.replace(tmp, self._path)
|
||||
|
||||
|
||||
# --------------------------------------------------------------- websocket codec
|
||||
|
||||
|
||||
def _encode_frame(opcode: int, payload: bytes) -> bytes:
|
||||
"""RFC 6455 client frame: FIN set, always masked (servers require it)."""
|
||||
header = bytearray([0x80 | opcode])
|
||||
n = len(payload)
|
||||
if n < 126:
|
||||
header.append(0x80 | n)
|
||||
elif n < 65536:
|
||||
header.append(0x80 | 126)
|
||||
header += struct.pack("!H", n)
|
||||
else:
|
||||
header.append(0x80 | 127)
|
||||
header += struct.pack("!Q", n)
|
||||
mask = secrets.token_bytes(4)
|
||||
header += mask
|
||||
masked = bytes(b ^ mask[i % 4] for i, b in enumerate(payload))
|
||||
return bytes(header) + masked
|
||||
|
||||
|
||||
class WsConnection:
|
||||
"""Minimal blocking RFC 6455 client over a raw socket (stdlib only)."""
|
||||
|
||||
def __init__(self, sock: socket.socket) -> None:
|
||||
self._sock = sock
|
||||
self._write_lock = threading.Lock()
|
||||
|
||||
@classmethod
|
||||
def connect(cls, url: str, timeout_s: float = 5.0) -> WsConnection:
|
||||
parts = urllib.parse.urlsplit(url)
|
||||
if parts.scheme not in ("ws", "wss"):
|
||||
raise ValueError(f"unsupported scheme in NC_INGEST_URL: {parts.scheme!r}")
|
||||
host = parts.hostname or "localhost"
|
||||
port = parts.port or (443 if parts.scheme == "wss" else 80)
|
||||
path = parts.path or "/"
|
||||
if parts.query:
|
||||
path += "?" + parts.query
|
||||
|
||||
sock = socket.create_connection((host, port), timeout=timeout_s)
|
||||
if parts.scheme == "wss":
|
||||
sock = ssl.create_default_context().wrap_socket(sock, server_hostname=host)
|
||||
|
||||
key = base64.b64encode(secrets.token_bytes(16)).decode("ascii")
|
||||
request = (
|
||||
f"GET {path} HTTP/1.1\r\n"
|
||||
f"Host: {host}:{port}\r\n"
|
||||
"Upgrade: websocket\r\n"
|
||||
"Connection: Upgrade\r\n"
|
||||
f"Sec-WebSocket-Key: {key}\r\n"
|
||||
"Sec-WebSocket-Version: 13\r\n\r\n"
|
||||
)
|
||||
sock.sendall(request.encode("ascii"))
|
||||
response = cls._read_http_response(sock)
|
||||
cls._validate_handshake(response, key)
|
||||
return cls(sock)
|
||||
|
||||
@staticmethod
|
||||
def _read_http_response(sock: socket.socket) -> bytes:
|
||||
buf = b""
|
||||
while b"\r\n\r\n" not in buf:
|
||||
chunk = sock.recv(4096)
|
||||
if not chunk:
|
||||
raise ConnectionError("server closed during WebSocket handshake")
|
||||
buf += chunk
|
||||
if len(buf) > _HANDSHAKE_MAX_BYTES:
|
||||
raise ConnectionError("handshake response exceeded size cap")
|
||||
return buf.split(b"\r\n\r\n", 1)[0]
|
||||
|
||||
@staticmethod
|
||||
def _validate_handshake(response: bytes, key: str) -> None:
|
||||
head = response.decode("latin-1")
|
||||
lines = head.split("\r\n")
|
||||
if not lines or " 101" not in lines[0]:
|
||||
raise ConnectionError(f"handshake rejected: {lines[0] if lines else '<empty>'}")
|
||||
headers = {}
|
||||
for line in lines[1:]:
|
||||
if ":" in line:
|
||||
name, _, value = line.partition(":")
|
||||
headers[name.strip().lower()] = value.strip()
|
||||
expect = base64.b64encode(
|
||||
hashlib.sha1((key + _WS_GUID).encode("ascii")).digest()
|
||||
).decode("ascii")
|
||||
if headers.get("sec-websocket-accept") != expect:
|
||||
raise ConnectionError("bad Sec-WebSocket-Accept in handshake response")
|
||||
|
||||
# -- send ------------------------------------------------------------
|
||||
def send_text(self, text: str) -> None:
|
||||
with self._write_lock:
|
||||
self._sock.sendall(_encode_frame(0x1, text.encode("utf-8")))
|
||||
|
||||
def _send_frame(self, opcode: int, payload: bytes) -> None:
|
||||
with self._write_lock:
|
||||
self._sock.sendall(_encode_frame(opcode, payload))
|
||||
|
||||
# -- receive ---------------------------------------------------------
|
||||
def recv_message(self, timeout_s: float) -> tuple[int, bytes] | None:
|
||||
"""Return (opcode, payload) for a data/close frame, or None on timeout.
|
||||
|
||||
Ping frames are answered with pong internally and never surfaced;
|
||||
pongs are swallowed. Fragmented messages are reassembled. Raises
|
||||
ConnectionError/OSError when the socket breaks.
|
||||
"""
|
||||
deadline = time.monotonic() + timeout_s
|
||||
fragments = bytearray()
|
||||
frag_opcode = 0
|
||||
while True:
|
||||
frame = self._recv_one_frame(deadline)
|
||||
if frame is None:
|
||||
return None
|
||||
fin, opcode, payload = frame
|
||||
if opcode == 0x9: # ping
|
||||
self._send_frame(0xA, payload)
|
||||
continue
|
||||
if opcode == 0xA: # pong
|
||||
continue
|
||||
if opcode == 0x0: # continuation
|
||||
fragments += payload
|
||||
else:
|
||||
fragments = bytearray(payload)
|
||||
frag_opcode = opcode
|
||||
if fin:
|
||||
return frag_opcode, bytes(fragments)
|
||||
|
||||
def _recv_one_frame(self, deadline: float) -> tuple[bool, int, bytes] | None:
|
||||
header = self._read_exact(2, deadline)
|
||||
if header is None:
|
||||
return None
|
||||
b0, b1 = header[0], header[1]
|
||||
fin = bool(b0 & 0x80)
|
||||
opcode = b0 & 0x0F
|
||||
length = b1 & 0x7F
|
||||
if length == 126:
|
||||
ext = self._read_exact(2, deadline)
|
||||
if ext is None:
|
||||
return None
|
||||
length = struct.unpack("!H", ext)[0]
|
||||
elif length == 127:
|
||||
ext = self._read_exact(8, deadline)
|
||||
if ext is None:
|
||||
return None
|
||||
length = struct.unpack("!Q", ext)[0]
|
||||
mask = self._read_exact(4, deadline) if (b1 & 0x80) else b""
|
||||
if mask is None:
|
||||
return None
|
||||
payload = self._read_exact(length, deadline) if length else b""
|
||||
if payload is None:
|
||||
return None
|
||||
if mask:
|
||||
payload = bytes(b ^ mask[i % 4] for i, b in enumerate(payload))
|
||||
return fin, opcode, payload
|
||||
|
||||
def _read_exact(self, n: int, deadline: float) -> bytes | None:
|
||||
buf = bytearray()
|
||||
while len(buf) < n:
|
||||
remaining = deadline - time.monotonic()
|
||||
if remaining <= 0:
|
||||
return None
|
||||
self._sock.settimeout(remaining)
|
||||
try:
|
||||
chunk = self._sock.recv(n - len(buf))
|
||||
except TimeoutError:
|
||||
return None
|
||||
if not chunk:
|
||||
raise ConnectionError("peer closed the WebSocket connection")
|
||||
buf += chunk
|
||||
return bytes(buf)
|
||||
|
||||
def close(self) -> None:
|
||||
try:
|
||||
self._sock.close()
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- agent
|
||||
|
||||
|
||||
class Agent:
|
||||
"""Wires capture (shell + workspace watcher) to the framed event stream."""
|
||||
|
||||
def __init__(self, config: AgentConfig) -> None:
|
||||
self.config = config
|
||||
self._spool = Spool(config.spool_path)
|
||||
self._pending: deque[str] = deque()
|
||||
self._seq = 0
|
||||
self._emit_lock = threading.Lock() # serializes seq + spool + flush
|
||||
self._conn_lock = threading.Lock() # guards _conn swaps
|
||||
self._conn: WsConnection | None = None
|
||||
self._last_sent: str | None = None # one-line replay margin, see below
|
||||
self._stop = threading.Event()
|
||||
self._threads: list[threading.Thread] = []
|
||||
self._baseline: dict[str, tuple[int, int, str | None]] = {}
|
||||
self._resume_from_spool()
|
||||
|
||||
# -- durability ------------------------------------------------------
|
||||
def _resume_from_spool(self) -> None:
|
||||
highest = -1
|
||||
for line in self._spool.read_all():
|
||||
self._pending.append(line)
|
||||
try:
|
||||
seq = int(json.loads(line).get("seq", -1))
|
||||
except (ValueError, AttributeError):
|
||||
continue
|
||||
highest = max(highest, seq)
|
||||
self._seq = highest + 1
|
||||
|
||||
# -- event construction ---------------------------------------------
|
||||
def _wire_frame(self, spooled_line: str) -> str:
|
||||
"""Spool format -> wire format: strip URL-owned identity fields.
|
||||
|
||||
The spool keeps full events (local durability + restart recovery).
|
||||
The ingest endpoint binds identity at the WS handshake (query params)
|
||||
and rejects frames carrying learner_id/task_id (`extra="forbid"`
|
||||
anti-spoofing), so the wire frame carries only seq/kind/payload/ts.
|
||||
"""
|
||||
import json as _json
|
||||
|
||||
full = _json.loads(spooled_line)
|
||||
wire = {
|
||||
k: full[k]
|
||||
for k in ("seq", "kind", "payload", "ts")
|
||||
}
|
||||
if full.get("sandbox_id"):
|
||||
wire["sandbox_id"] = full["sandbox_id"]
|
||||
return _json.dumps(wire)
|
||||
|
||||
def _next_event(self, kind: str, payload: dict[str, Any]) -> dict[str, Any]:
|
||||
if kind not in _EVENT_KINDS:
|
||||
raise ValueError(f"unknown event kind: {kind!r}")
|
||||
event = {
|
||||
"learner_id": self.config.learner_id,
|
||||
"task_id": self.config.task_id,
|
||||
"seq": self._seq,
|
||||
"kind": kind,
|
||||
"payload": payload,
|
||||
"ts": datetime.now(UTC).isoformat(),
|
||||
"sandbox_id": self.config.sandbox_id,
|
||||
}
|
||||
self._seq += 1
|
||||
return event
|
||||
|
||||
# -- emission / flush (D-026) ----------------------------------------
|
||||
def emit(self, kind: str, payload: dict[str, Any]) -> dict[str, Any]:
|
||||
"""Spool-then-send. Never blocks on reconnect; loss is impossible."""
|
||||
with self._emit_lock:
|
||||
line = json.dumps(self._next_event(kind, payload))
|
||||
self._spool.append(line) # durable BEFORE any send attempt
|
||||
self._pending.append(line)
|
||||
self._flush_locked()
|
||||
return json.loads(line)
|
||||
|
||||
def _flush_locked(self) -> None:
|
||||
conn = self._current_conn()
|
||||
while self._pending and conn is not None:
|
||||
line = self._pending[0]
|
||||
try:
|
||||
conn.send_text(self._wire_frame(line))
|
||||
except (ConnectionError, OSError):
|
||||
self._drop_conn()
|
||||
return
|
||||
self._pending.popleft()
|
||||
self._last_sent = line # kept until a later send proves delivery
|
||||
if not self._pending and self._last_sent is not None:
|
||||
# Compact, but retain the most recently sent line: a send into a
|
||||
# silently-dead socket "succeeds" once at TCP level, so the last
|
||||
# line is only confirmed-sent once a later write works. Retention
|
||||
# is cheap; the server dedups on (learner, task, seq).
|
||||
self._spool.rewrite([self._last_sent])
|
||||
|
||||
def replay_margin(self) -> None:
|
||||
"""Requeue the last-sent line after a detected disconnect."""
|
||||
with self._emit_lock:
|
||||
if self._last_sent is not None and (
|
||||
not self._pending or self._pending[0] != self._last_sent
|
||||
):
|
||||
self._pending.appendleft(self._last_sent)
|
||||
self._spool.rewrite(list(self._pending))
|
||||
self._last_sent = None
|
||||
|
||||
# -- connection supervision ------------------------------------------
|
||||
def _current_conn(self) -> WsConnection | None:
|
||||
with self._conn_lock:
|
||||
return self._conn
|
||||
|
||||
def _set_conn(self, conn: WsConnection | None) -> None:
|
||||
with self._conn_lock:
|
||||
self._conn = conn
|
||||
|
||||
def _drop_conn(self) -> None:
|
||||
conn = self._current_conn()
|
||||
self._set_conn(None)
|
||||
if conn is not None:
|
||||
conn.close()
|
||||
self.replay_margin()
|
||||
|
||||
def is_connected(self) -> bool:
|
||||
return self._current_conn() is not None
|
||||
|
||||
def wait_connected(self, timeout_s: float) -> bool:
|
||||
return self._wait_for(lambda: self.is_connected(), timeout_s)
|
||||
|
||||
def wait_disconnected(self, timeout_s: float) -> bool:
|
||||
return self._wait_for(lambda: not self.is_connected(), timeout_s)
|
||||
|
||||
def _wait_for(self, pred: Any, timeout_s: float) -> bool:
|
||||
deadline = time.monotonic() + timeout_s
|
||||
while time.monotonic() < deadline:
|
||||
if pred():
|
||||
return True
|
||||
time.sleep(0.02)
|
||||
return pred()
|
||||
|
||||
def _supervisor_loop(self) -> None:
|
||||
"""Maintain the WS connection: connect, flush backlog, read, backoff."""
|
||||
backoff = self.config.backoff_base_s
|
||||
while not self._stop.is_set():
|
||||
if self._current_conn() is None:
|
||||
try:
|
||||
conn = WsConnection.connect(self.config.ingest_url)
|
||||
except (ConnectionError, OSError, ValueError, TimeoutError):
|
||||
self._stop.wait(backoff)
|
||||
backoff = min(self.config.backoff_max_s, backoff * 2)
|
||||
continue
|
||||
self._set_conn(conn)
|
||||
self._last_sent = None
|
||||
backoff = self.config.backoff_base_s
|
||||
with self._emit_lock: # ordered against concurrent emit()s
|
||||
self._flush_locked()
|
||||
else:
|
||||
conn = self._current_conn()
|
||||
if conn is None:
|
||||
continue
|
||||
try:
|
||||
frame = conn.recv_message(timeout_s=1.0)
|
||||
except (ConnectionError, OSError):
|
||||
self._drop_conn()
|
||||
continue
|
||||
if frame is None:
|
||||
continue
|
||||
opcode, _payload = frame
|
||||
if opcode == 0x8: # server close frame
|
||||
self._drop_conn()
|
||||
|
||||
# -- workspace watcher ------------------------------------------------
|
||||
def _snapshot_workspace(self) -> dict[str, tuple[int, int, str | None]]:
|
||||
"""Map rel path -> (mtime_ns, size, text-or-None-if-too-large)."""
|
||||
snap: dict[str, tuple[int, int, str | None]] = {}
|
||||
root = self.config.workspace
|
||||
if not root.is_dir():
|
||||
return snap
|
||||
for dirpath, dirnames, filenames in os.walk(root):
|
||||
dirnames[:] = sorted(
|
||||
d for d in dirnames if not d.startswith(_AGENT_DIR_PREFIX)
|
||||
)
|
||||
for name in sorted(filenames):
|
||||
if name.startswith(_AGENT_DIR_PREFIX):
|
||||
continue
|
||||
path = Path(dirpath) / name
|
||||
try:
|
||||
st = path.stat()
|
||||
except OSError:
|
||||
continue
|
||||
rel = path.relative_to(root).as_posix()
|
||||
text: str | None = None
|
||||
if st.st_size <= _MAX_DIFF_BYTES:
|
||||
try:
|
||||
text = path.read_text(encoding="utf-8", errors="replace")
|
||||
except OSError:
|
||||
pass
|
||||
snap[rel] = (st.st_mtime_ns, st.st_size, text)
|
||||
return snap
|
||||
|
||||
def _file_diff_payload(self, rel: str, change: str, old: str | None, new: str | None) -> dict:
|
||||
payload: dict[str, Any] = {"path": rel, "change": change}
|
||||
if old is None and new is None:
|
||||
payload["truncated"] = True
|
||||
return payload
|
||||
diff = "".join(
|
||||
difflib.unified_diff(
|
||||
(old or "").splitlines(keepends=True),
|
||||
(new or "").splitlines(keepends=True),
|
||||
fromfile=f"a/{rel}",
|
||||
tofile=f"b/{rel}",
|
||||
)
|
||||
)
|
||||
payload["diff"] = diff
|
||||
payload["size"] = len(new or "")
|
||||
return payload
|
||||
|
||||
def _watcher_loop(self) -> None:
|
||||
# Baseline is taken in start() before it returns, so any change made
|
||||
# after start() completes is guaranteed to be observed.
|
||||
baseline = self._baseline
|
||||
last_heartbeat = time.monotonic()
|
||||
while not self._stop.wait(self.config.poll_interval_s):
|
||||
current = self._snapshot_workspace()
|
||||
for rel in sorted(current.keys() | baseline.keys()):
|
||||
if rel not in baseline and rel in current:
|
||||
self.emit(
|
||||
"file_diff",
|
||||
self._file_diff_payload(rel, "created", None, current[rel][2]),
|
||||
)
|
||||
elif rel in baseline and rel not in current:
|
||||
self.emit(
|
||||
"file_diff",
|
||||
self._file_diff_payload(rel, "deleted", baseline[rel][2], None),
|
||||
)
|
||||
else:
|
||||
old_stat, new_stat = baseline[rel], current[rel]
|
||||
if old_stat[:2] != new_stat[:2] and old_stat[2] != new_stat[2]:
|
||||
self.emit(
|
||||
"file_diff",
|
||||
self._file_diff_payload(
|
||||
rel, "modified", old_stat[2], new_stat[2]
|
||||
),
|
||||
)
|
||||
baseline = current
|
||||
if time.monotonic() - last_heartbeat >= self.config.activity_interval_s:
|
||||
self.emit("activity", {"state": "idle", "spooled": len(self._pending)})
|
||||
last_heartbeat = time.monotonic()
|
||||
|
||||
# -- shell wrapper ------------------------------------------------------
|
||||
@staticmethod
|
||||
def _is_test_command(cmd: str) -> bool:
|
||||
return "test" in cmd.lower()
|
||||
|
||||
def run_command(self, line: str) -> dict[str, Any] | None:
|
||||
"""Run one REPL line; emits stdin/command/stdout/run|test_result."""
|
||||
line = line.strip()
|
||||
if not line:
|
||||
return None
|
||||
self.emit("stdin", {"line": line})
|
||||
self.emit("activity", {"state": "command", "spooled": len(self._pending)})
|
||||
self.emit("command", {"cmd": line})
|
||||
started = time.monotonic()
|
||||
timed_out = False
|
||||
exit_code: int | None = None
|
||||
out: str | bytes = ""
|
||||
err: str | bytes = ""
|
||||
try:
|
||||
proc = subprocess.run(
|
||||
["sh", "-c", line],
|
||||
cwd=self.config.workspace,
|
||||
capture_output=True,
|
||||
timeout=self.config.command_timeout_s,
|
||||
text=True,
|
||||
errors="replace",
|
||||
)
|
||||
exit_code, out, err = proc.returncode, proc.stdout, proc.stderr
|
||||
except subprocess.TimeoutExpired as exc:
|
||||
timed_out = True
|
||||
# TimeoutExpired output attrs are always bytes (even in text mode).
|
||||
out = exc.stdout or b""
|
||||
err = exc.stderr or b""
|
||||
duration = time.monotonic() - started
|
||||
for stream, data in (("stdout", out), ("stderr", err)):
|
||||
if isinstance(data, bytes):
|
||||
data = data.decode(errors="replace")
|
||||
if data:
|
||||
self.emit("stdout", {"stream": stream, "data": data[-_MAX_OUTPUT_CHARS:]})
|
||||
kind = "test_result" if self._is_test_command(line) else "run_result"
|
||||
result = self.emit(
|
||||
kind,
|
||||
{
|
||||
"cmd": line,
|
||||
"exit_code": exit_code,
|
||||
"duration_s": round(duration, 6),
|
||||
"timed_out": timed_out,
|
||||
},
|
||||
)
|
||||
self.emit("activity", {"state": "idle", "spooled": len(self._pending)})
|
||||
return result
|
||||
|
||||
# -- lifecycle ----------------------------------------------------------
|
||||
def start(self) -> None:
|
||||
self.config.workspace.mkdir(parents=True, exist_ok=True)
|
||||
self._baseline = self._snapshot_workspace()
|
||||
self.emit("activity", {"state": "starting", "spooled": len(self._pending)})
|
||||
self._threads = [
|
||||
threading.Thread(target=self._supervisor_loop, daemon=True, name="nc-ws"),
|
||||
threading.Thread(target=self._watcher_loop, daemon=True, name="nc-watch"),
|
||||
]
|
||||
for thread in self._threads:
|
||||
thread.start()
|
||||
|
||||
def stop(self) -> None:
|
||||
if self._stop.is_set():
|
||||
return
|
||||
try:
|
||||
self.emit("activity", {"state": "stopped", "spooled": len(self._pending)})
|
||||
finally:
|
||||
self._stop.set()
|
||||
self._drop_conn()
|
||||
for thread in self._threads:
|
||||
thread.join(timeout=3)
|
||||
for thread in self._threads:
|
||||
thread.join(timeout=3)
|
||||
with self._emit_lock:
|
||||
self._spool.rewrite(list(self._pending))
|
||||
|
||||
|
||||
def main() -> int:
|
||||
try:
|
||||
config = AgentConfig.from_env()
|
||||
except ValueError as exc:
|
||||
print(f"sandbox-agent: {exc}", file=sys.stderr)
|
||||
return 2
|
||||
agent = Agent(config)
|
||||
agent.start()
|
||||
try:
|
||||
# REPL mode: each stdin line is executed and reported (interactive use).
|
||||
# Daemon mode: when stdin is closed/absent (the sandbox backend spawns
|
||||
# the agent with stdin=DEVNULL), keep streaming workspace diffs +
|
||||
# activity until SIGTERM/SIGINT so the agent's lifecycle is tied to
|
||||
# the sandbox (destroy() reaps it) rather than to stdin EOF.
|
||||
if sys.stdin is None or sys.stdin.closed: # pragma: no cover - defensive
|
||||
agent._stop.wait() # noqa: SLF001 - daemon block
|
||||
else:
|
||||
line = sys.stdin.readline()
|
||||
while line:
|
||||
agent.run_command(line)
|
||||
line = sys.stdin.readline()
|
||||
if not agent._stop.is_set() and not sys.stdin.isatty(): # noqa: SLF001
|
||||
# EOF on a pipe (DEVNULL): daemonize — watch + stream until killed.
|
||||
import signal
|
||||
|
||||
signal.signal(signal.SIGTERM, lambda *_: agent._stop.set()) # noqa: SLF001
|
||||
agent._stop.wait() # noqa: SLF001
|
||||
except KeyboardInterrupt:
|
||||
pass
|
||||
finally:
|
||||
agent.stop()
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -1,103 +1,118 @@
|
||||
"""Assessor agent tests — structured rubric scores (REQ-2-008).
|
||||
"""Assessor agent tests — live-grade coaching contract (REQ-3-007).
|
||||
|
||||
The Assessor is the structured-output showcase: tests use ScriptedJSONProvider
|
||||
for valid payloads and exercise the 4-layer defense failure modes.
|
||||
v0.3 re-grounding: the Assessor renders coaching FROM the stored grade
|
||||
(GradeRecord) — it never invents scores (the grading engine owns that).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from datetime import UTC, datetime
|
||||
|
||||
import pytest
|
||||
|
||||
from ai_service.agents.assessor import AssessorAgent, RubricScore
|
||||
from ai_service.agents.structured import StructuredOutputError
|
||||
from ai_service.agents.assessor import AssessorAgent, GradeCoaching
|
||||
from ai_service.config import Settings
|
||||
from ai_service.corpus.artifacts import (
|
||||
get_artifact_bundle,
|
||||
get_transcript_for_artifact,
|
||||
render_rubric,
|
||||
from ai_service.grading.store import GradeRecord, SQLiteGradeStore
|
||||
from ai_service.llm.mock import MockProvider
|
||||
from ai_service.llm.types import Message
|
||||
|
||||
COACHING_JSON = json.dumps(
|
||||
{
|
||||
"summary": "Solid iterative build; tests drove the fixes.",
|
||||
"strengths": ["Ran tests after each change."],
|
||||
"gaps": ["Did not cover the empty-input case."],
|
||||
"next_steps": ["Add one edge-case test."],
|
||||
}
|
||||
)
|
||||
from ai_service.corpus.learner_context import get_learner_context
|
||||
from ai_service.llm.mock import MockProvider, ScriptedJSONProvider
|
||||
|
||||
VALID_SCORE = {
|
||||
"rubric_id": "rubric-orchestration-c002",
|
||||
"artifact_id": "art-eval-research-assistant",
|
||||
"competency_id": "stack-orchestration-c002",
|
||||
"scores": [
|
||||
{"criterion_id": "rc-architecture", "name": "Agent architecture soundness",
|
||||
"score": 92, "evidence": "Explicit state schema with planner-only write access"},
|
||||
{"criterion_id": "rc-communication", "name": "Inter-agent communication design",
|
||||
"score": 88, "evidence": "Typed ToolMessage responses with retry flags"},
|
||||
{"criterion_id": "rc-reliability", "name": "Reliability engineering",
|
||||
"score": 85, "evidence": "3-retry loop with degradation path"},
|
||||
{"criterion_id": "rc-process", "name": "Process trace quality",
|
||||
"score": 90, "evidence": "Iterative saves with passing test checkpoints"},
|
||||
],
|
||||
"strengths": ["Clean state boundaries", "Failure-aware tool wrapping"],
|
||||
"gaps": ["No reviewer node yet", "Graph diagram only in README"],
|
||||
"verdict": "mastered",
|
||||
}
|
||||
|
||||
|
||||
def make_assessor(provider=None) -> AssessorAgent:
|
||||
return AssessorAgent(provider or MockProvider(), Settings(provider="mock"))
|
||||
class ScriptedProvider(MockProvider):
|
||||
def __init__(self) -> None:
|
||||
super().__init__()
|
||||
self.requests: list[list[Message]] = []
|
||||
self.replies: list[str] = []
|
||||
|
||||
async def chat(self, messages, *, model, temperature=0.7, response_format=None):
|
||||
self.requests.append([Message(role=m.role, content=m.content) for m in messages])
|
||||
if self.replies:
|
||||
return self.replies.pop(0)
|
||||
return COACHING_JSON
|
||||
|
||||
|
||||
def get_bundle(artifact_id="art-eval-research-assistant"):
|
||||
bundle = get_artifact_bundle(artifact_id)
|
||||
assert bundle is not None
|
||||
return bundle
|
||||
|
||||
|
||||
async def test_evaluate_returns_validated_rubric_score():
|
||||
provider = ScriptedJSONProvider(VALID_SCORE)
|
||||
assessor = make_assessor(provider)
|
||||
artifact, rubric = get_bundle()
|
||||
transcript = get_transcript_for_artifact(artifact.artifact_id)
|
||||
result = await assessor.evaluate(artifact, rubric, transcript)
|
||||
assert isinstance(result, RubricScore)
|
||||
assert result.verdict == "mastered"
|
||||
assert len(result.scores) == 4
|
||||
assert result.weighted_total(rubric) == pytest.approx(
|
||||
92 * 0.3 + 88 * 0.3 + 85 * 0.25 + 90 * 0.15
|
||||
def _grade() -> GradeRecord:
|
||||
return GradeRecord(
|
||||
learner_id="assessor-learner",
|
||||
task_id="assessor-task",
|
||||
variant_seed=None,
|
||||
digest={"error_fix_cycles": 2, "final_test_status": "pass"},
|
||||
scores={
|
||||
"criteria": {
|
||||
"process_quality": 4,
|
||||
"correctness": 3,
|
||||
"debugging_discipline": 4,
|
||||
"test_usage": 3,
|
||||
},
|
||||
"strengths": ["s"],
|
||||
"gaps": ["g"],
|
||||
"verdict": "developing",
|
||||
},
|
||||
verdict="GRADED",
|
||||
model="gemma4:31b",
|
||||
created_at=datetime.now(UTC),
|
||||
)
|
||||
|
||||
|
||||
async def test_evaluate_rejects_invalid_schema_after_retry():
|
||||
"""Plain MockProvider returns non-rubric JSON → 4-layer defense exhausts
|
||||
its single retry and raises StructuredOutputError."""
|
||||
assessor = make_assessor(MockProvider())
|
||||
artifact, rubric = get_bundle()
|
||||
transcript = get_transcript_for_artifact(artifact.artifact_id)
|
||||
with pytest.raises(StructuredOutputError):
|
||||
await assessor.evaluate(artifact, rubric, transcript)
|
||||
@pytest.fixture()
|
||||
def provider() -> ScriptedProvider:
|
||||
return ScriptedProvider()
|
||||
|
||||
|
||||
def test_build_evaluation_input_carries_all_inputs():
|
||||
assessor = make_assessor()
|
||||
artifact, rubric = get_bundle()
|
||||
transcript = get_transcript_for_artifact(artifact.artifact_id)
|
||||
text = assessor.build_evaluation_input(artifact, rubric, transcript)
|
||||
assert artifact.name in text
|
||||
assert artifact.evidence_excerpt in text
|
||||
assert "rc-architecture" in text # rubric rendered
|
||||
assert "examiner:" in text # transcript rendered
|
||||
@pytest.fixture()
|
||||
def agent(provider) -> AssessorAgent:
|
||||
return AssessorAgent(provider, Settings(provider="mock"))
|
||||
|
||||
|
||||
def test_build_evaluation_input_without_transcript():
|
||||
assessor = make_assessor()
|
||||
artifact, rubric = get_bundle()
|
||||
text = assessor.build_evaluation_input(artifact, rubric, None)
|
||||
assert artifact.name in text
|
||||
assert "examiner:" not in text
|
||||
class TestCoachGrade:
|
||||
async def test_prompt_contains_stored_grade_not_learner_id(self, agent, provider) -> None:
|
||||
await agent.coach_grade(_grade())
|
||||
all_text = "\n".join(
|
||||
m.content for request in provider.requests for m in request
|
||||
)
|
||||
assert "process_quality" in all_text # stored scores rendered
|
||||
assert "GRADED" in all_text
|
||||
assert "assessor-learner" not in all_text # D-028 anonymity
|
||||
|
||||
async def test_coaching_validates_via_d020(self, agent, provider) -> None:
|
||||
coaching = await agent.coach_grade(_grade())
|
||||
assert isinstance(coaching, GradeCoaching)
|
||||
assert coaching.summary
|
||||
assert coaching.next_steps
|
||||
|
||||
async def test_malformed_then_good_exercises_retry(self, agent, provider) -> None:
|
||||
provider.replies = ["garbage", COACHING_JSON]
|
||||
coaching = await agent.coach_grade(_grade())
|
||||
assert coaching.summary
|
||||
assert len(provider.requests) == 2
|
||||
|
||||
async def test_no_corpus_artifact_imports(self) -> None:
|
||||
import ast
|
||||
from pathlib import Path
|
||||
|
||||
py = Path(__file__).parents[2] / "ai_service" / "agents" / "assessor.py"
|
||||
tree = ast.parse(py.read_text())
|
||||
for node in ast.walk(tree):
|
||||
if isinstance(node, ast.ImportFrom) and node.module:
|
||||
assert "corpus.artifacts" not in node.module
|
||||
assert "corpus.telemetry" not in node.module
|
||||
|
||||
|
||||
def test_system_prompt_names_assessor_persona():
|
||||
prompt = make_assessor().system_prompt(get_learner_context())
|
||||
assert "Assessor" in prompt
|
||||
assert "ONLY" in prompt # JSON-only instruction
|
||||
|
||||
|
||||
def test_rubric_render_in_prompt_is_complete():
|
||||
"""The rubric passed to the model lists every criterion (fair grading)."""
|
||||
artifact, rubric = get_bundle()
|
||||
text = render_rubric(rubric)
|
||||
assert text.count("rc-") == len(rubric.criteria)
|
||||
class TestStoreRoundtrip:
|
||||
def test_grade_store_roundtrip(self, tmp_path) -> None:
|
||||
store = SQLiteGradeStore(db_path=tmp_path / "g.db")
|
||||
record = _grade()
|
||||
store.save(record)
|
||||
fetched = store.get("assessor-learner", "assessor-task")
|
||||
assert fetched is not None
|
||||
assert fetched.scores["criteria"]["process_quality"] == 4
|
||||
store.close()
|
||||
|
||||
@@ -0,0 +1,159 @@
|
||||
"""Examiner agent tests (Task 5-2-01, REQ-3-006)."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
|
||||
import pytest
|
||||
|
||||
from ai_service.agents.examiner import DefenseVerdict, ExaminerAgent
|
||||
from ai_service.agents.registry import AgentRegistry, register_builtin_agents
|
||||
from ai_service.config import Settings
|
||||
from ai_service.grading.features import compute_digest
|
||||
from ai_service.llm.mock import MockProvider
|
||||
from ai_service.llm.types import Message
|
||||
|
||||
VERDICT_JSON = json.dumps(
|
||||
{
|
||||
"verdict": "developing",
|
||||
"understanding": "Explains the retry loop clearly.",
|
||||
"process_justification": "Justifies the edit-then-test cadence from the digest.",
|
||||
"communication": "Answers are specific and on-topic.",
|
||||
"strengths": ["Grounded the fix in a failed test."],
|
||||
"gaps": ["Did not justify the chunk-size choice."],
|
||||
}
|
||||
)
|
||||
|
||||
|
||||
class RecordingProvider(MockProvider):
|
||||
"""Mock provider that records every message list (prompt assertions)."""
|
||||
|
||||
def __init__(self, replies: list[str] | None = None) -> None:
|
||||
super().__init__()
|
||||
self.replies = list(replies or [])
|
||||
self.requests: list[list[Message]] = []
|
||||
|
||||
async def chat(self, messages, *, model, temperature=0.7, response_format=None):
|
||||
self.requests.append([Message(role=m.role, content=m.content) for m in messages])
|
||||
if self.replies:
|
||||
return self.replies.pop(0)
|
||||
return "Tell me about your build."
|
||||
|
||||
|
||||
def _digest():
|
||||
from datetime import UTC, datetime, timedelta
|
||||
|
||||
from ai_service.telemetry.models import TelemetryEvent
|
||||
|
||||
t0 = datetime(2026, 9, 12, tzinfo=UTC)
|
||||
events = [
|
||||
TelemetryEvent(
|
||||
learner_id="examiner-learner",
|
||||
task_id="examiner-task",
|
||||
seq=n,
|
||||
kind=kind,
|
||||
payload=payload,
|
||||
ts=t0 + timedelta(seconds=n * 10),
|
||||
sandbox_id="sbx-examiner",
|
||||
)
|
||||
for n, (kind, payload) in enumerate(
|
||||
[
|
||||
("file_diff", {"path": "a.py"}),
|
||||
("command", {"cmd": "pytest -q"}),
|
||||
("test_result", {"passed": False, "exit_code": 1}),
|
||||
("file_diff", {"path": "a.py"}),
|
||||
("test_result", {"passed": True, "exit_code": 0}),
|
||||
]
|
||||
)
|
||||
]
|
||||
return compute_digest(events)
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def provider() -> RecordingProvider:
|
||||
return RecordingProvider()
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def settings() -> Settings:
|
||||
return Settings(provider="mock")
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def examiner(provider, settings) -> ExaminerAgent:
|
||||
return ExaminerAgent(provider, settings)
|
||||
|
||||
|
||||
class TestNextQuestion:
|
||||
async def test_prompt_contains_digest_but_no_learner_id(
|
||||
self, examiner, provider
|
||||
) -> None:
|
||||
await examiner.next_question(
|
||||
history=[Message(role="assistant", content="First question?")],
|
||||
trace_digest=_digest(),
|
||||
variant_statement="Build a chunker.",
|
||||
)
|
||||
all_content = "\n".join(
|
||||
m.content for request in provider.requests for m in request
|
||||
)
|
||||
assert "error_fix_cycles" in all_content # digest JSON grounded
|
||||
assert "examiner-learner" not in all_content # D-028 anonymity
|
||||
assert "Build a chunker." in all_content # variant statement grounded
|
||||
assert all_content.count('"examiner-learner"') == 0
|
||||
|
||||
async def test_question_returned_from_provider(self, examiner) -> None:
|
||||
question = await examiner.next_question(
|
||||
history=[], trace_digest=_digest()
|
||||
)
|
||||
assert isinstance(question, str)
|
||||
|
||||
|
||||
class TestFinalVerdict:
|
||||
async def test_verdict_validates_via_d020(self, examiner, provider) -> None:
|
||||
provider.replies = [VERDICT_JSON]
|
||||
verdict = await examiner.final_verdict(
|
||||
history=[Message(role="assistant", content="Q?")],
|
||||
trace_digest=_digest(),
|
||||
)
|
||||
assert isinstance(verdict, DefenseVerdict)
|
||||
assert verdict.verdict == "developing"
|
||||
assert verdict.strengths and verdict.gaps
|
||||
|
||||
async def test_malformed_then_good_exercises_retry(self, examiner, provider) -> None:
|
||||
provider.replies = ["not json", VERDICT_JSON]
|
||||
verdict = await examiner.final_verdict(history=[], trace_digest=_digest())
|
||||
assert verdict.verdict == "developing"
|
||||
assert len(provider.requests) == 2 # D-020 bounded retry
|
||||
|
||||
|
||||
class TestRegistry:
|
||||
def test_all_seven_agents_resolve(self, provider, settings) -> None:
|
||||
registry = AgentRegistry()
|
||||
register_builtin_agents(registry)
|
||||
assert registry.names() == [
|
||||
"assessor",
|
||||
"coach",
|
||||
"examiner",
|
||||
"lab",
|
||||
"mentor",
|
||||
"proctor",
|
||||
"tutor",
|
||||
]
|
||||
agent = registry.get(provider, settings, "examiner")
|
||||
assert isinstance(agent, ExaminerAgent)
|
||||
|
||||
|
||||
class TestBoundary:
|
||||
def test_examiner_never_imports_voice_or_api(self) -> None:
|
||||
import ast
|
||||
from pathlib import Path
|
||||
|
||||
py = Path(__file__).parents[2] / "ai_service" / "agents" / "examiner.py"
|
||||
tree = ast.parse(py.read_text())
|
||||
for node in ast.walk(tree):
|
||||
if isinstance(node, ast.ImportFrom) and node.module:
|
||||
assert "voice" not in node.module, "examiner must not import voice/"
|
||||
assert not node.module.startswith("ai_service.api")
|
||||
if isinstance(node, ast.Import):
|
||||
for alias in node.names:
|
||||
assert alias.name != "fastapi"
|
||||
@@ -97,10 +97,11 @@ def test_proctor_and_mentor_resolve_via_registry():
|
||||
assert isinstance(mentor, MentorAgent)
|
||||
|
||||
|
||||
def test_registry_resolves_all_six_agents():
|
||||
"""Must-Have (Phase 5): the full roster — coach/tutor/lab/assessor/proctor/mentor."""
|
||||
def test_registry_resolves_all_seven_agents():
|
||||
"""Must-Have (Phase 5): the full roster — six tutors + the Examiner."""
|
||||
from ai_service.agents.assessor import AssessorAgent
|
||||
from ai_service.agents.coach import CoachAgent
|
||||
from ai_service.agents.examiner import ExaminerAgent
|
||||
from ai_service.agents.lab import LabAgent
|
||||
from ai_service.agents.mentor import MentorAgent
|
||||
from ai_service.agents.proctor import ProctorAgent
|
||||
@@ -108,7 +109,9 @@ def test_registry_resolves_all_six_agents():
|
||||
|
||||
registry = AgentRegistry()
|
||||
register_builtin_agents(registry)
|
||||
assert registry.names() == ["assessor", "coach", "lab", "mentor", "proctor", "tutor"]
|
||||
assert registry.names() == [
|
||||
"assessor", "coach", "examiner", "lab", "mentor", "proctor", "tutor",
|
||||
]
|
||||
settings = Settings(provider="mock")
|
||||
expected = {
|
||||
"coach": CoachAgent,
|
||||
@@ -117,6 +120,7 @@ def test_registry_resolves_all_six_agents():
|
||||
"assessor": AssessorAgent,
|
||||
"proctor": ProctorAgent,
|
||||
"mentor": MentorAgent,
|
||||
"examiner": ExaminerAgent,
|
||||
}
|
||||
for name, cls in expected.items():
|
||||
agent = registry.get(MockProvider(), settings, name)
|
||||
|
||||
@@ -0,0 +1,207 @@
|
||||
"""Live re-grounding tests: Lab, Assessor, Proctor on REAL inputs (REQ-3-007)."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import tempfile
|
||||
from datetime import UTC, datetime, timedelta
|
||||
from pathlib import Path
|
||||
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from ai_service.agents.assessor import AssessorAgent, GradeCoaching
|
||||
from ai_service.agents.lab import LabAgent
|
||||
from ai_service.agents.proctor import ProctorAgent, ProctorAssessment
|
||||
from ai_service.config import Settings
|
||||
from ai_service.grading.store import GradeRecord, SQLiteGradeStore
|
||||
from ai_service.llm.mock import MockProvider
|
||||
from ai_service.llm.types import Message
|
||||
from ai_service.telemetry.models import TelemetryEvent
|
||||
from ai_service.telemetry.store import SQLiteTraceStore
|
||||
|
||||
COACHING_JSON = json.dumps(
|
||||
{
|
||||
"summary": "Iterative build with test discipline.",
|
||||
"strengths": ["Tested after changes."],
|
||||
"gaps": ["Missing edge cases."],
|
||||
"next_steps": ["Add an edge-case test."],
|
||||
}
|
||||
)
|
||||
PROCTOR_JSON = json.dumps(
|
||||
{
|
||||
"signals": [
|
||||
{"signal_type": "idle_gap", "severity": "low", "note": "One 400s gap."}
|
||||
],
|
||||
"intervention": "Offer a short break.",
|
||||
"summary": "Healthy session overall.",
|
||||
}
|
||||
)
|
||||
T0 = datetime(2026, 9, 12, tzinfo=UTC)
|
||||
|
||||
|
||||
class RecordingProvider(MockProvider):
|
||||
def __init__(self, structured_json: str) -> None:
|
||||
super().__init__()
|
||||
self._structured_json = structured_json
|
||||
self.requests: list[list[Message]] = []
|
||||
|
||||
def _reply_for(self, messages, response_format):
|
||||
self.requests.append([Message(role=m.role, content=m.content) for m in messages])
|
||||
if response_format is not None and response_format.get("type") == "json_object":
|
||||
return self._structured_json
|
||||
return "Coaching feedback referencing your latest test run."
|
||||
|
||||
|
||||
def _event(
|
||||
seq: int, kind: str, payload: dict, offset_s: float,
|
||||
learner="live-learner", task="live-task",
|
||||
):
|
||||
return TelemetryEvent(
|
||||
learner_id=learner,
|
||||
task_id=task,
|
||||
seq=seq,
|
||||
kind=kind,
|
||||
payload=payload,
|
||||
ts=T0 + timedelta(seconds=offset_s),
|
||||
sandbox_id="sbx-live",
|
||||
)
|
||||
|
||||
|
||||
def _seed_trace(store: SQLiteTraceStore) -> None:
|
||||
events = [
|
||||
_event(0, "file_diff", {"path": "a.py"}, 0),
|
||||
_event(1, "command", {"cmd": "pytest -q"}, 10),
|
||||
_event(2, "test_result", {"passed": False, "exit_code": 1}, 15),
|
||||
_event(3, "file_diff", {"path": "a.py"}, 30),
|
||||
_event(4, "test_result", {"passed": True, "exit_code": 0}, 45),
|
||||
_event(5, "activity", {"state": "idle"}, 500), # >120s gap -> idle
|
||||
]
|
||||
for e in events:
|
||||
store.append(e)
|
||||
|
||||
|
||||
class TestLabLive:
|
||||
async def test_lab_prompt_contains_digest_not_corpus(self, tmp_path) -> None:
|
||||
store = SQLiteTraceStore(db_path=tmp_path / "t.db")
|
||||
_seed_trace(store)
|
||||
provider = RecordingProvider("feedback")
|
||||
agent = LabAgent(provider, Settings(provider="mock"))
|
||||
|
||||
from ai_service.grading.features import compute_digest
|
||||
|
||||
digest = compute_digest(store.get_trace("live-learner", "live-task"))
|
||||
tokens = [t async for t in agent.stream_feedback(digest)]
|
||||
assert tokens
|
||||
all_text = "\n".join(
|
||||
m.content for request in provider.requests for m in request
|
||||
)
|
||||
assert "error_fix_cycles" in all_text
|
||||
assert "lab-scenario" not in all_text # no corpus fixture ids
|
||||
store.close()
|
||||
|
||||
def test_no_corpus_telemetry_import_in_lab(self) -> None:
|
||||
import ast
|
||||
from pathlib import Path
|
||||
|
||||
py = Path(__file__).parents[2] / "ai_service" / "agents" / "lab.py"
|
||||
tree = ast.parse(py.read_text())
|
||||
for node in ast.walk(tree):
|
||||
if isinstance(node, ast.ImportFrom) and node.module:
|
||||
assert "corpus.telemetry" not in node.module
|
||||
|
||||
|
||||
class TestAssessorLive:
|
||||
async def test_assessor_prompt_contains_stored_scores(self) -> None:
|
||||
provider = RecordingProvider(COACHING_JSON)
|
||||
agent = AssessorAgent(provider, Settings(provider="mock"))
|
||||
grade = GradeRecord(
|
||||
learner_id="live-learner",
|
||||
task_id="live-task",
|
||||
variant_seed=None,
|
||||
digest={"error_fix_cycles": 1},
|
||||
scores={"criteria": {"process_quality": 3}},
|
||||
verdict="GRADED",
|
||||
model="mock",
|
||||
created_at=datetime.now(UTC),
|
||||
)
|
||||
coaching = await agent.coach_grade(grade)
|
||||
assert isinstance(coaching, GradeCoaching)
|
||||
all_text = "\n".join(
|
||||
m.content for request in provider.requests for m in request
|
||||
)
|
||||
assert "process_quality" in all_text
|
||||
assert "live-learner" not in all_text
|
||||
|
||||
def test_no_corpus_artifact_import_in_assessor(self) -> None:
|
||||
import ast
|
||||
from pathlib import Path
|
||||
|
||||
py = Path(__file__).parents[2] / "ai_service" / "agents" / "assessor.py"
|
||||
tree = ast.parse(py.read_text())
|
||||
for node in ast.walk(tree):
|
||||
if isinstance(node, ast.ImportFrom) and node.module:
|
||||
assert "corpus.artifacts" not in node.module
|
||||
assert "corpus.telemetry" not in node.module
|
||||
|
||||
|
||||
class TestProctorLive:
|
||||
async def test_proctor_receives_real_digest_and_defense_signals(self) -> None:
|
||||
provider = RecordingProvider(PROCTOR_JSON)
|
||||
agent = ProctorAgent(provider, Settings(provider="mock"))
|
||||
from ai_service.grading.features import compute_digest
|
||||
|
||||
store = SQLiteTraceStore(db_path=Path(tempfile.mkdtemp()) / "proctor-t.db")
|
||||
_seed_trace(store)
|
||||
digest = compute_digest(store.get_trace("live-learner", "live-task"))
|
||||
store.close()
|
||||
|
||||
assessment = await agent.assess(
|
||||
digest,
|
||||
defense_signals={"long_pauses": [{"turn": 3, "latency_ms": 30000}]},
|
||||
variant_context={"template_id": "tpl-llm-judge", "seed": "cafe", "params": {}},
|
||||
)
|
||||
assert isinstance(assessment, ProctorAssessment)
|
||||
all_text = "\n".join(
|
||||
m.content for request in provider.requests for m in request
|
||||
)
|
||||
assert "idle_gap" in all_text or "idle_gap_count" in all_text
|
||||
assert "long_pauses" in all_text
|
||||
assert "tpl-llm-judge" in all_text
|
||||
assert "live-learner" not in all_text
|
||||
|
||||
def test_no_corpus_scenario_import_in_proctor(self) -> None:
|
||||
import ast
|
||||
from pathlib import Path
|
||||
|
||||
py = Path(__file__).parents[2] / "ai_service" / "agents" / "proctor.py"
|
||||
tree = ast.parse(py.read_text())
|
||||
for node in ast.walk(tree):
|
||||
if isinstance(node, ast.ImportFrom) and node.module:
|
||||
assert "corpus.telemetry" not in node.module
|
||||
|
||||
|
||||
class TestProctorEndpoint:
|
||||
def test_signals_endpoint_serves_real_inputs(self, tmp_path) -> None:
|
||||
from ai_service.main import create_app
|
||||
from ai_service.telemetry.ingest import TraceIntegrityMap
|
||||
from ai_service.variants.store import SQLiteVariantStore
|
||||
from ai_service.voice.defense_store import SQLiteDefenseStore
|
||||
|
||||
app = create_app(Settings(provider="mock"))
|
||||
provider = RecordingProvider(PROCTOR_JSON)
|
||||
app.state.provider = provider
|
||||
app.state.trace_store = SQLiteTraceStore(db_path=tmp_path / "t.db")
|
||||
app.state.grade_store = SQLiteGradeStore(db_path=tmp_path / "g.db")
|
||||
app.state.variant_store = SQLiteVariantStore(db_path=tmp_path / "v.db")
|
||||
app.state.defense_store = SQLiteDefenseStore(db_path=tmp_path / "d.db")
|
||||
app.state.trace_integrity = TraceIntegrityMap()
|
||||
_seed_trace(app.state.trace_store)
|
||||
with TestClient(app) as client:
|
||||
resp = client.post(
|
||||
"/v1/proctor/signals",
|
||||
json={"learner_id": "live-learner", "task_id": "live-task"},
|
||||
)
|
||||
assert resp.status_code == 200, resp.text
|
||||
body = resp.json()
|
||||
assert body["signals"] is not None
|
||||
assert "intervention" in body
|
||||
@@ -1,79 +1,87 @@
|
||||
"""Assessment evaluate endpoint tests — validated JSON, 404s (REQ-2-008)."""
|
||||
"""Assessment API tests — stored-grade coaching contract (REQ-3-007)."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from datetime import UTC, datetime
|
||||
|
||||
import pytest
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from ai_service.config import Settings
|
||||
from ai_service.grading.store import GradeRecord, SQLiteGradeStore
|
||||
from ai_service.llm.mock import MockProvider
|
||||
from ai_service.main import create_app
|
||||
from ai_service.telemetry.ingest import TraceIntegrityMap
|
||||
from ai_service.telemetry.store import SQLiteTraceStore
|
||||
|
||||
COACHING_JSON = json.dumps(
|
||||
{
|
||||
"summary": "Good iterative work.",
|
||||
"strengths": ["Tests after changes."],
|
||||
"gaps": ["Missing edge cases."],
|
||||
"next_steps": ["Add an edge-case test."],
|
||||
}
|
||||
)
|
||||
|
||||
|
||||
from ai_service.agents.assessor import RubricScore
|
||||
from ai_service.llm.mock import ScriptedJSONProvider
|
||||
|
||||
VALID_SCORE = {
|
||||
"rubric_id": "rubric-orchestration-c002",
|
||||
"artifact_id": "art-eval-research-assistant",
|
||||
"competency_id": "stack-orchestration-c002",
|
||||
"scores": [
|
||||
{"criterion_id": "rc-architecture", "name": "Agent architecture soundness",
|
||||
"score": 92, "evidence": "Explicit state schema"},
|
||||
{"criterion_id": "rc-communication", "name": "Inter-agent communication design",
|
||||
"score": 88, "evidence": "Typed ToolMessage responses"},
|
||||
{"criterion_id": "rc-reliability", "name": "Reliability engineering",
|
||||
"score": 85, "evidence": "3-retry loop"},
|
||||
{"criterion_id": "rc-process", "name": "Process trace quality",
|
||||
"score": 90, "evidence": "Iterative checkpoints"},
|
||||
],
|
||||
"strengths": ["Clean state boundaries", "Failure-aware tools"],
|
||||
"gaps": ["No reviewer node", "Diagram only in README"],
|
||||
"verdict": "mastered",
|
||||
}
|
||||
class CoachingMock(MockProvider):
|
||||
def _reply_for(self, messages, response_format):
|
||||
if response_format is not None and response_format.get("type") == "json_object":
|
||||
return COACHING_JSON
|
||||
return super()._reply_for(messages, response_format)
|
||||
|
||||
|
||||
def test_evaluate_returns_validated_rubric_json(client):
|
||||
# Swap the app provider for a scripted-JSON provider for this test
|
||||
original = client.app.state.provider
|
||||
client.app.state.provider = ScriptedJSONProvider(VALID_SCORE)
|
||||
try:
|
||||
response = client.post(
|
||||
"/v1/assessment/evaluate",
|
||||
json={"artifact_id": "art-eval-research-assistant"},
|
||||
@pytest.fixture()
|
||||
def client(tmp_path) -> TestClient:
|
||||
app = create_app(Settings(provider="mock"))
|
||||
app.state.provider = CoachingMock()
|
||||
app.state.trace_store = SQLiteTraceStore(db_path=tmp_path / "t.db")
|
||||
app.state.grade_store = SQLiteGradeStore(db_path=tmp_path / "g.db")
|
||||
app.state.trace_integrity = TraceIntegrityMap()
|
||||
with TestClient(app) as c:
|
||||
yield c
|
||||
|
||||
|
||||
def _seed_grade(client: TestClient) -> None:
|
||||
|
||||
app = client.app
|
||||
store: SQLiteGradeStore = app.state.grade_store
|
||||
store.save(
|
||||
GradeRecord(
|
||||
learner_id="api-learner",
|
||||
task_id="api-task",
|
||||
variant_seed=None,
|
||||
digest={"error_fix_cycles": 1},
|
||||
scores={"criteria": {"process_quality": 3}, "verdict": "developing"},
|
||||
verdict="GRADED",
|
||||
model="gemma4:31b",
|
||||
created_at=datetime.now(UTC),
|
||||
)
|
||||
finally:
|
||||
client.app.state.provider = original
|
||||
assert response.status_code == 200
|
||||
data = response.json()
|
||||
validated = RubricScore.model_validate(data) # response contract holds
|
||||
assert validated.verdict == "mastered"
|
||||
assert len(validated.scores) == 4
|
||||
|
||||
|
||||
def test_unknown_artifact_404(client):
|
||||
response = client.post("/v1/assessment/evaluate", json={"artifact_id": "ghost"})
|
||||
assert response.status_code == 404
|
||||
assert "ghost" in response.json()["detail"]
|
||||
|
||||
|
||||
def test_unparseable_provider_502(client):
|
||||
"""Plain MockProvider yields non-rubric JSON → structured defense exhausts
|
||||
retry → endpoint translates to 502 (bad gateway to the model)."""
|
||||
# default mock already returns non-rubric JSON
|
||||
response = client.post(
|
||||
"/v1/assessment/evaluate",
|
||||
json={"artifact_id": "art-eval-research-assistant"},
|
||||
)
|
||||
assert response.status_code == 502
|
||||
assert "failed" in response.json()["detail"].lower()
|
||||
|
||||
|
||||
def test_missing_artifact_id_422(client):
|
||||
response = client.post("/v1/assessment/evaluate", json={})
|
||||
assert response.status_code == 422
|
||||
def test_evaluate_renders_stored_grade_as_coaching(client) -> None:
|
||||
_seed_grade(client)
|
||||
resp = client.post(
|
||||
"/v1/assessment/evaluate",
|
||||
json={"learner_id": "api-learner", "task_id": "api-task"},
|
||||
)
|
||||
assert resp.status_code == 200, resp.text
|
||||
body = resp.json()
|
||||
assert body["grade_verdict"] == "GRADED"
|
||||
assert body["coaching"]["summary"]
|
||||
|
||||
|
||||
def test_second_artifact_also_evaluates(client):
|
||||
original = client.app.state.provider
|
||||
payload = dict(VALID_SCORE, artifact_id="art-eval-rag-dashboard")
|
||||
client.app.state.provider = ScriptedJSONProvider(payload)
|
||||
try:
|
||||
response = client.post(
|
||||
"/v1/assessment/evaluate", json={"artifact_id": "art-eval-rag-dashboard"}
|
||||
)
|
||||
finally:
|
||||
client.app.state.provider = original
|
||||
assert response.status_code == 200
|
||||
assert response.json()["artifact_id"] == "art-eval-rag-dashboard"
|
||||
def test_evaluate_without_grade_404(client) -> None:
|
||||
resp = client.post(
|
||||
"/v1/assessment/evaluate",
|
||||
json={"learner_id": "nobody", "task_id": "nothing"},
|
||||
)
|
||||
assert resp.status_code == 404
|
||||
assert "grade first" in resp.json()["detail"]
|
||||
|
||||
|
||||
def test_missing_fields_422(client) -> None:
|
||||
resp = client.post("/v1/assessment/evaluate", json={"learner_id": "x"})
|
||||
assert resp.status_code == 422
|
||||
|
||||
@@ -0,0 +1,82 @@
|
||||
"""CORS policy tests (A-008, D-038 network mode).
|
||||
|
||||
v0.3 initially shipped `allow_methods` WITHOUT "PUT" while the learner
|
||||
build surface writes workspace files with PUT (engine-client writeFile) —
|
||||
every cross-origin Save failed preflight. These tests pin the policy so a
|
||||
future method-list edit fails loudly instead of silently breaking the
|
||||
headline flow.
|
||||
|
||||
v0.3.5 network mode (D-038): the default AI_CORS_ORIGINS='*' admits any
|
||||
origin (safe ONLY because credentials are never enabled); an explicit list
|
||||
restricts. Both modes are pinned here:
|
||||
- wildcard: remote origin gets the grant; credentials still never sent;
|
||||
- explicit: unlisted origins get no grant.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
ALLOWED_ORIGIN = "http://localhost:3000"
|
||||
REMOTE_ORIGIN = "http://nextcraft-1:3000"
|
||||
ALL_CLIENT_METHODS = ("GET", "POST", "PUT", "DELETE")
|
||||
|
||||
|
||||
def _allow_origin(resp) -> str | None:
|
||||
return resp.headers.get("access-control-allow-origin")
|
||||
|
||||
|
||||
def test_preflight_allows_every_method_the_web_client_uses(client: TestClient) -> None:
|
||||
for method in ALL_CLIENT_METHODS:
|
||||
resp = client.options(
|
||||
"/v1/sandboxes",
|
||||
headers={
|
||||
"Origin": ALLOWED_ORIGIN,
|
||||
"Access-Control-Request-Method": method,
|
||||
},
|
||||
)
|
||||
assert resp.status_code == 200, f"preflight {method} failed: {resp.status_code}"
|
||||
assert _allow_origin(resp) in ("*", ALLOWED_ORIGIN)
|
||||
allowed = resp.headers["access-control-allow-methods"].split(", ")
|
||||
assert method in allowed, f"{method} missing from CORS methods: {allowed}"
|
||||
|
||||
|
||||
def test_cross_origin_get_echoes_allow_origin(client: TestClient) -> None:
|
||||
resp = client.get("/v1/sandboxes", headers={"Origin": ALLOWED_ORIGIN})
|
||||
assert resp.status_code == 200
|
||||
assert _allow_origin(resp) in ("*", ALLOWED_ORIGIN)
|
||||
|
||||
|
||||
def test_wildcard_mode_grants_remote_origins(client: TestClient) -> None:
|
||||
"""D-038 default: '*' grants any origin — remote browsers work zero-config."""
|
||||
resp = client.get("/v1/sandboxes", headers={"Origin": REMOTE_ORIGIN})
|
||||
assert resp.status_code == 200
|
||||
assert _allow_origin(resp) in ("*", REMOTE_ORIGIN)
|
||||
|
||||
|
||||
def test_explicit_list_mode_denies_unlisted_origins(
|
||||
settings, monkeypatch, tmp_path
|
||||
) -> None:
|
||||
"""Explicit AI_CORS_ORIGINS restricts to the listed origins only."""
|
||||
from fastapi.testclient import TestClient as TC
|
||||
|
||||
from ai_service.main import create_app
|
||||
|
||||
restricted = settings.model_copy(update={"cors_origins": "http://localhost:3000"})
|
||||
app = create_app(restricted)
|
||||
with TC(app) as c:
|
||||
resp = c.get("/v1/sandboxes", headers={"Origin": "https://evil.example"})
|
||||
assert resp.status_code == 200 # non-CORS requests still serve
|
||||
assert resp.headers.get("access-control-allow-origin") is None
|
||||
|
||||
|
||||
def test_credentials_never_allowed(client: TestClient) -> None:
|
||||
resp = client.options(
|
||||
"/v1/sandboxes",
|
||||
headers={
|
||||
"Origin": ALLOWED_ORIGIN,
|
||||
"Access-Control-Request-Method": "PUT",
|
||||
"Access-Control-Request-Headers": "Content-Type",
|
||||
},
|
||||
)
|
||||
assert resp.headers.get("access-control-allow-credentials") != "true"
|
||||
@@ -0,0 +1,234 @@
|
||||
"""Defense endpoint tests (Task 5-3-01, REQ-3-006) — mock voice + mock LLM."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from ai_service.agents.examiner import ExaminerAgent
|
||||
from ai_service.config import Settings
|
||||
from ai_service.grading.store import SQLiteGradeStore
|
||||
from ai_service.llm.mock import MockProvider
|
||||
from ai_service.main import create_app
|
||||
from ai_service.telemetry.ingest import TraceIntegrityMap
|
||||
from ai_service.telemetry.store import SQLiteTraceStore
|
||||
from ai_service.variants.store import SQLiteVariantStore
|
||||
from ai_service.voice.defense_store import SQLiteDefenseStore
|
||||
from ai_service.voice.mock import MockVoiceProvider
|
||||
|
||||
VERDICT = {
|
||||
"verdict": "developing",
|
||||
"understanding": "Explains the build clearly.",
|
||||
"process_justification": "Justifies choices.",
|
||||
"communication": "Clear and specific.",
|
||||
"strengths": ["Grounded answers in the digest."],
|
||||
"gaps": ["Did not address the edge cases."],
|
||||
}
|
||||
|
||||
|
||||
class ScriptedLLM(MockProvider):
|
||||
"""Question-mode calls get a question; verdict-mode calls get D-020 JSON.
|
||||
|
||||
Discriminator: the verdict prompt contains "final verdict JSON" — the
|
||||
question prompt says "next question".
|
||||
"""
|
||||
|
||||
def __init__(self) -> None:
|
||||
super().__init__()
|
||||
|
||||
async def chat(self, messages, *, model, temperature=0.7, response_format=None):
|
||||
all_text = "\n".join(m.content for m in messages)
|
||||
if "final verdict JSON" in all_text:
|
||||
return json.dumps(VERDICT)
|
||||
return "Why did you structure the fix that way?"
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def app(tmp_path: Path):
|
||||
application = create_app(Settings(provider="mock", voice_provider="mock"))
|
||||
llm = ScriptedLLM()
|
||||
application.state.provider = llm
|
||||
application.state.trace_store = SQLiteTraceStore(db_path=tmp_path / "t.db")
|
||||
application.state.variant_store = SQLiteVariantStore(db_path=tmp_path / "v.db")
|
||||
application.state.grade_store = SQLiteGradeStore(db_path=tmp_path / "g.db")
|
||||
application.state.trace_integrity = TraceIntegrityMap()
|
||||
application.state.defense_store = SQLiteDefenseStore(db_path=tmp_path / "d.db")
|
||||
application.state.voice_provider = MockVoiceProvider(
|
||||
["the fix was in the retry loop"]
|
||||
)
|
||||
settings = Settings(provider="mock", voice_provider="mock")
|
||||
application.state.examiner_agent = ExaminerAgent(llm, settings)
|
||||
return application
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def client(app) -> TestClient:
|
||||
with TestClient(app) as c:
|
||||
yield c
|
||||
|
||||
|
||||
def _start(client: TestClient) -> dict:
|
||||
resp = client.post(
|
||||
"/v1/defense/start", json={"learner_id": "defense-learner", "task_id": "defense-task"}
|
||||
)
|
||||
assert resp.status_code == 200, resp.text
|
||||
return resp.json()
|
||||
|
||||
|
||||
class TestStart:
|
||||
def test_start_returns_first_question_and_descriptor(self, client) -> None:
|
||||
body = _start(client)
|
||||
assert body["first_question"]
|
||||
assert body["defense_id"]
|
||||
assert body["voice_descriptor"]["mode"] == "mock"
|
||||
assert body["trace_complete"] is True
|
||||
stored = client.get(f"/v1/defense/{body['defense_id']}")
|
||||
assert stored.status_code == 200
|
||||
turns = stored.json()["turns"]
|
||||
assert turns and turns[0]["role"] == "examiner"
|
||||
|
||||
def test_start_with_unknown_trace_is_complete_flag(self, client) -> None:
|
||||
body = _start(client)
|
||||
assert body["trace_complete"] is True
|
||||
|
||||
|
||||
class TestBrowserFallback:
|
||||
def test_browser_mode_serves_browser_descriptor(self, tmp_path: Path) -> None:
|
||||
"""Must-Have #6: AI_VOICE_PROVIDER=browser → start returns the
|
||||
browser-native SR/TTS fallback descriptor (D-030), not 'mock'."""
|
||||
application = create_app(Settings(provider="mock", voice_provider="browser"))
|
||||
application.state.provider = ScriptedLLM()
|
||||
application.state.trace_store = SQLiteTraceStore(db_path=tmp_path / "t.db")
|
||||
application.state.variant_store = SQLiteVariantStore(db_path=tmp_path / "v.db")
|
||||
application.state.grade_store = SQLiteGradeStore(db_path=tmp_path / "g.db")
|
||||
application.state.trace_integrity = TraceIntegrityMap()
|
||||
application.state.defense_store = SQLiteDefenseStore(db_path=tmp_path / "d.db")
|
||||
application.state.voice_provider = MockVoiceProvider(["answer"])
|
||||
llm = ScriptedLLM()
|
||||
application.state.examiner_agent = ExaminerAgent(
|
||||
llm, Settings(provider="mock", voice_provider="browser")
|
||||
)
|
||||
with TestClient(application) as c:
|
||||
body = _start(c)
|
||||
assert body["voice_descriptor"]["mode"] == "browser"
|
||||
assert body["voice_descriptor"]["sr_available"]
|
||||
assert "SpeechRecognition" in body["voice_descriptor"]["hint"]
|
||||
|
||||
|
||||
class TestAnswer:
|
||||
def test_typed_answer_yields_followup_with_latency(self, client) -> None:
|
||||
defense_id = _start(client)["defense_id"]
|
||||
resp = client.post(
|
||||
f"/v1/defense/{defense_id}/answer", data={"text": "I fixed the loop."}
|
||||
)
|
||||
assert resp.status_code == 200, resp.text
|
||||
body = resp.json()
|
||||
assert body["question"]
|
||||
assert body["turn_latency"]["llm_ms"] is not None
|
||||
|
||||
def test_audio_answer_transcribed_and_recorded(self, client) -> None:
|
||||
defense_id = _start(client)["defense_id"]
|
||||
wav_bytes = b"RIFF" + b"\x00" * 64
|
||||
resp = client.post(
|
||||
f"/v1/defense/{defense_id}/answer",
|
||||
files={"audio": ("answer.wav", wav_bytes, "audio/wav")},
|
||||
)
|
||||
assert resp.status_code == 200, resp.text
|
||||
stored = client.get(f"/v1/defense/{defense_id}").json()
|
||||
learner_turns = [t for t in stored["turns"] if t["role"] == "learner"]
|
||||
assert learner_turns, "learner turn missing after audio answer"
|
||||
assert learner_turns[0]["text"] == "the fix was in the retry loop"
|
||||
|
||||
def test_empty_audio_is_422_not_500(self, client) -> None:
|
||||
"""Zero-byte upload must 422 before the provider call (a real
|
||||
provider would raise the same way the mock does — validate first)."""
|
||||
defense_id = _start(client)["defense_id"]
|
||||
resp = client.post(
|
||||
f"/v1/defense/{defense_id}/answer",
|
||||
files={"audio": ("answer.wav", b"", "audio/wav")},
|
||||
)
|
||||
assert resp.status_code == 422, resp.text
|
||||
stored = client.get(f"/v1/defense/{defense_id}").json()
|
||||
assert len(stored["turns"]) == 1 # nothing appended
|
||||
|
||||
def test_answer_after_finish_is_409(self, client) -> None:
|
||||
"""A sealed transcript is append-only-no-more: the endpoints own
|
||||
turn-vs-finalize sequencing (defense_store contract)."""
|
||||
defense_id = _start(client)["defense_id"]
|
||||
client.post(f"/v1/defense/{defense_id}/answer", data={"text": "a"})
|
||||
assert client.post(f"/v1/defense/{defense_id}/finish").status_code == 200
|
||||
resp = client.post(f"/v1/defense/{defense_id}/answer", data={"text": "late"})
|
||||
assert resp.status_code == 409, resp.text
|
||||
stored = client.get(f"/v1/defense/{defense_id}").json()
|
||||
assert len(stored["turns"]) == 3 # ex, lrn, ex — no post-finish turns
|
||||
|
||||
def test_neither_text_nor_audio_422(self, client) -> None:
|
||||
defense_id = _start(client)["defense_id"]
|
||||
resp = client.post(f"/v1/defense/{defense_id}/answer")
|
||||
assert resp.status_code == 422
|
||||
|
||||
def test_unknown_defense_404(self, client) -> None:
|
||||
resp = client.post("/v1/defense/dfn-nope/answer", data={"text": "hi"})
|
||||
assert resp.status_code == 404
|
||||
|
||||
|
||||
class TestAudioEndpoint:
|
||||
def test_examiner_turn_streams_wav(self, client) -> None:
|
||||
defense_id = _start(client)["defense_id"]
|
||||
resp = client.get(f"/v1/defense/{defense_id}/audio/0")
|
||||
assert resp.status_code == 200
|
||||
assert resp.content
|
||||
assert resp.headers["content-type"].startswith("audio/")
|
||||
|
||||
def test_unknown_turn_404(self, client) -> None:
|
||||
defense_id = _start(client)["defense_id"]
|
||||
assert client.get(f"/v1/defense/{defense_id}/audio/42").status_code == 404
|
||||
|
||||
|
||||
class TestFinishAndGet:
|
||||
def test_full_loop_verdict_and_signals(self, client) -> None:
|
||||
defense_id = _start(client)["defense_id"]
|
||||
client.post(f"/v1/defense/{defense_id}/answer", data={"text": "answer one"})
|
||||
finish = client.post(f"/v1/defense/{defense_id}/finish")
|
||||
assert finish.status_code == 200, finish.text
|
||||
body = finish.json()
|
||||
assert body["verdict"]["verdict"] == "developing"
|
||||
assert body["integrity_signals"]["pause_threshold_ms"]
|
||||
stored = client.get(f"/v1/defense/{defense_id}").json()
|
||||
assert stored["status"] == "finished"
|
||||
assert stored["integrity_signals"]
|
||||
# Must-Have #1: "verdict + transcript persisted" — the verdict must
|
||||
# be retrievable from GET after finish, not only in the finish body.
|
||||
assert stored["integrity_signals"]["verdict"]["verdict"] == "developing"
|
||||
|
||||
def test_long_pause_flagged(self, client, app) -> None:
|
||||
from datetime import UTC, datetime
|
||||
|
||||
from ai_service.voice.defense_store import DefenseTurn
|
||||
|
||||
defense_id = _start(client)["defense_id"]
|
||||
# inject a slow learner turn directly (simulated latency)
|
||||
store = app.state.defense_store
|
||||
store.append_turn(
|
||||
defense_id,
|
||||
DefenseTurn(
|
||||
defense_id=defense_id,
|
||||
seq=99,
|
||||
role="learner",
|
||||
text="slow reply",
|
||||
ts=datetime.now(UTC),
|
||||
latency_ms=30_000,
|
||||
created_at=datetime.now(UTC),
|
||||
),
|
||||
)
|
||||
finish = client.post(f"/v1/defense/{defense_id}/finish")
|
||||
assert finish.status_code == 200
|
||||
signals = finish.json()["integrity_signals"]
|
||||
assert any(p["turn"] == 99 for p in signals["long_pauses"])
|
||||
|
||||
def test_unknown_defense_404_on_all(self, client) -> None:
|
||||
assert client.post("/v1/defense/dfn-nope/finish").status_code == 404
|
||||
assert client.get("/v1/defense/dfn-nope").status_code == 404
|
||||
@@ -0,0 +1,216 @@
|
||||
"""Full credential-flow E2E over real engines (Task 6-5-01, REQ-3-007/008).
|
||||
|
||||
Endpoint-level end-to-end with mock LLM/voice providers (G-2 precedent:
|
||||
real engine plumbing over real endpoints; provider choice is
|
||||
service-internal): variant -> telemetry-wired sandbox -> real in-sandbox
|
||||
exec -> trace -> grade -> oral defense -> verdict/signals -> proctor.
|
||||
No corpus fixture anywhere in the flow.
|
||||
|
||||
Probe-guarded for user namespaces (the in-sandbox exec needs them).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import contextlib
|
||||
import json
|
||||
import socket
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
import uvicorn
|
||||
|
||||
from ai_service.config import Settings
|
||||
from ai_service.grading.store import SQLiteGradeStore
|
||||
from ai_service.llm.mock import MockProvider
|
||||
from ai_service.main import create_app
|
||||
from ai_service.telemetry.ingest import TraceIntegrityMap
|
||||
from ai_service.telemetry.store import SQLiteTraceStore
|
||||
from ai_service.variants.store import SQLiteVariantStore
|
||||
from ai_service.voice.defense_store import SQLiteDefenseStore
|
||||
from tests.sandbox.test_isolation import USERSNS_AVAILABLE
|
||||
|
||||
COACHING_JSON = json.dumps(
|
||||
{
|
||||
"summary": "Strong iteration.",
|
||||
"strengths": ["Tests early."],
|
||||
"gaps": ["One edge case missing."],
|
||||
"next_steps": ["Add it."],
|
||||
}
|
||||
)
|
||||
VERDICT_JSON = json.dumps(
|
||||
{
|
||||
"verdict": "developing",
|
||||
"understanding": "Explains the build clearly.",
|
||||
"process_justification": "Choices defended.",
|
||||
"communication": "Clear.",
|
||||
"strengths": ["Grounded in the digest."],
|
||||
"gaps": ["Missed one edge case."],
|
||||
}
|
||||
)
|
||||
PROCTOR_JSON = json.dumps(
|
||||
{
|
||||
"signals": [
|
||||
{"signal_type": "idle_gap", "severity": "low", "note": "A short pause."}
|
||||
],
|
||||
"intervention": "Keep momentum.",
|
||||
"summary": "Healthy session.",
|
||||
}
|
||||
)
|
||||
|
||||
|
||||
class FlowLLM(MockProvider):
|
||||
"""Prompt-discriminated: verdict vs coaching vs question vs proctor JSON."""
|
||||
|
||||
def _reply_for(self, messages, response_format):
|
||||
all_text = "\n".join(m.content for m in messages)
|
||||
if response_format is not None and response_format.get("type") == "json_object":
|
||||
if "final verdict JSON" in all_text:
|
||||
return VERDICT_JSON
|
||||
if "rubric" in all_text.lower() and "Score this build session" in all_text:
|
||||
return json.dumps(
|
||||
{
|
||||
"criteria": {
|
||||
"process_quality": 4,
|
||||
"correctness": 3,
|
||||
"debugging_discipline": 3,
|
||||
"test_usage": 4,
|
||||
},
|
||||
"strengths": ["Iterated with tests."],
|
||||
"gaps": ["One edge case missing."],
|
||||
"verdict": "developing",
|
||||
}
|
||||
)
|
||||
if "Explain it as coaching" in all_text:
|
||||
return COACHING_JSON
|
||||
if "integrity signals supportively" in all_text:
|
||||
return PROCTOR_JSON
|
||||
return json.dumps(
|
||||
{
|
||||
"statement": (
|
||||
"Build a judge for code-review answers scoring factual "
|
||||
"accuracy with 3 edge cases and 5 test examples."
|
||||
)
|
||||
}
|
||||
)
|
||||
return "Walk me through your last fix — what changed and why?"
|
||||
|
||||
|
||||
def _free_port() -> int:
|
||||
with socket.socket() as s:
|
||||
s.bind(("127.0.0.1", 0))
|
||||
return s.getsockname()[1]
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_full_credential_flow(tmp_path: Path) -> None:
|
||||
if not USERSNS_AVAILABLE:
|
||||
pytest.skip("user namespaces unavailable on this host (probe)")
|
||||
|
||||
import httpx
|
||||
|
||||
port = _free_port()
|
||||
settings = Settings(provider="mock", voice_provider="mock", port=port)
|
||||
app = create_app(settings)
|
||||
app.state.provider = FlowLLM()
|
||||
app.state.trace_store = SQLiteTraceStore(db_path=tmp_path / "t.db")
|
||||
app.state.grade_store = SQLiteGradeStore(db_path=tmp_path / "g.db")
|
||||
app.state.variant_store = SQLiteVariantStore(db_path=tmp_path / "v.db")
|
||||
app.state.defense_store = SQLiteDefenseStore(db_path=tmp_path / "d.db")
|
||||
app.state.trace_integrity = TraceIntegrityMap()
|
||||
|
||||
server = uvicorn.Server(uvicorn.Config(app, host="127.0.0.1", port=port, log_level="warning"))
|
||||
serve_task = asyncio.get_running_loop().create_task(server.serve())
|
||||
try:
|
||||
for _ in range(100):
|
||||
if server.started:
|
||||
break
|
||||
await asyncio.sleep(0.1)
|
||||
assert server.started
|
||||
|
||||
base = f"http://127.0.0.1:{port}"
|
||||
async with httpx.AsyncClient(base_url=base, timeout=30.0) as client:
|
||||
# 1. Variant (real seeded generation, mock-rendered).
|
||||
var = (await client.post("/v1/variants", json={
|
||||
"learner_id": "pilot-learner", "competency_id": "stack-orchestration-c007",
|
||||
})).json()
|
||||
assert var["task_id"] and var["statement"] and var["starter_files"]
|
||||
task_id = var["task_id"]
|
||||
|
||||
# 2. Telemetry-wired sandbox (real namespaces; agent joins via loopback).
|
||||
sbx = (await client.post("/v1/sandboxes", json={
|
||||
"learner_id": "pilot-learner", "task_id": task_id,
|
||||
}))
|
||||
assert sbx.status_code == 201, sbx.text
|
||||
sandbox_id = sbx.json()["id"]
|
||||
|
||||
# 3. Real in-sandbox exec: run the starter test (capture agent streams
|
||||
# the workspace effects into the trace).
|
||||
exec_resp = await client.post(f"/v1/sandboxes/{sandbox_id}/exec", json={
|
||||
"cmd": ["pytest", "-q"],
|
||||
})
|
||||
assert exec_resp.status_code == 200, exec_resp.text
|
||||
|
||||
# 4. Trace: events landed in order (the capture agent runs async).
|
||||
trace: list = []
|
||||
deadline = time.monotonic() + 20.0
|
||||
while time.monotonic() < deadline:
|
||||
tr = await client.get(f"/v1/telemetry/traces/pilot-learner/{task_id}")
|
||||
if tr.status_code == 200:
|
||||
trace = tr.json().get("events", [])
|
||||
if trace:
|
||||
break
|
||||
await asyncio.sleep(0.25)
|
||||
assert trace, "no telemetry events arrived from the real sandbox"
|
||||
seqs = [e["seq"] for e in trace]
|
||||
assert seqs == sorted(seqs)
|
||||
|
||||
# 5. Grade: rubric from the real digest (G-4 gate passed: no gaps).
|
||||
grade = (await client.post("/v1/assessment/grade", json={
|
||||
"learner_id": "pilot-learner", "task_id": task_id,
|
||||
})).json()
|
||||
assert grade["verdict"] == "GRADED", grade
|
||||
assert grade["scores"]["criteria"]["process_quality"] == 4
|
||||
assert grade["variant_seed"] == var["seed"] # D-029 stamped
|
||||
|
||||
# 6. Assessor coaching FROM the stored grade.
|
||||
coaching = (await client.post("/v1/assessment/evaluate", json={
|
||||
"learner_id": "pilot-learner", "task_id": task_id,
|
||||
})).json()
|
||||
assert coaching["coaching"]["summary"]
|
||||
|
||||
# 7. Oral defense: start -> typed answers -> finish (mock voice).
|
||||
defense = (await client.post("/v1/defense/start", json={
|
||||
"learner_id": "pilot-learner", "task_id": task_id,
|
||||
})).json()
|
||||
assert defense["first_question"]
|
||||
did = defense["defense_id"]
|
||||
ans = await client.post(f"/v1/defense/{did}/answer", data={"text": "I fixed the loop."})
|
||||
assert ans.status_code == 200, ans.text
|
||||
finish = (await client.post(f"/v1/defense/{did}/finish")).json()
|
||||
assert finish["verdict"]["verdict"] == "developing"
|
||||
assert finish["integrity_signals"]
|
||||
|
||||
# 8. Proctor over the real digest + defense signals.
|
||||
proctor = (await client.post("/v1/proctor/signals", json={
|
||||
"learner_id": "pilot-learner", "task_id": task_id,
|
||||
})).json()
|
||||
assert proctor["intervention"]
|
||||
|
||||
# 9. Sandbox destroyed; no leaks.
|
||||
destroy = await client.delete(f"/v1/sandboxes/{sandbox_id}")
|
||||
assert destroy.status_code == 204
|
||||
listed = (await client.get("/v1/sandboxes")).json()
|
||||
assert all(s["id"] != sandbox_id for s in (listed.get("sandboxes") or []))
|
||||
|
||||
# 10. No corpus fixtures anywhere in this flow's payloads.
|
||||
corpus_markers = ("lab-scenario", "proctor-scenario", "artifact-")
|
||||
for payload in (var, grade, defense, finish, proctor):
|
||||
assert not any(
|
||||
m in json.dumps(payload) for m in corpus_markers
|
||||
), "corpus fixture leaked into the learner path"
|
||||
finally:
|
||||
server.should_exit = True
|
||||
with contextlib.suppress(Exception):
|
||||
await asyncio.wait_for(serve_task, timeout=10.0)
|
||||
@@ -0,0 +1,397 @@
|
||||
"""Assessment grade endpoint tests — grading engine over HTTP (Task 3-3-01).
|
||||
|
||||
Contract under test (api/assessment.py, REQ-3-004):
|
||||
|
||||
POST /v1/assessment/grade {learner_id, task_id}
|
||||
GRADED → 200, rubric scores + digest summary
|
||||
UNGRADABLE_TRACE_INCOMPLETE → 200, gate record (missing_seqs /
|
||||
integrity_flag surfaced in scores)
|
||||
UNGRADABLE_EMPTY_TRACE → 200, gate record (documented choice: an
|
||||
unknown task is ALSO an empty pair; the
|
||||
engine cannot distinguish, and the gate
|
||||
outcome is a durable first-class result)
|
||||
StructuredOutputError → 502 (provider exhausted the D-020 budget)
|
||||
GET /v1/assessment/grade/{learner_id}/{task_id}
|
||||
stored latest grade → 200 (same scores as the POST)
|
||||
never-graded pair → 404
|
||||
|
||||
Wiring: per-test tmp-path SQLite stores + a pre-set GradingEngine
|
||||
(state-injection override — the lifespan adopts trace_store/grade_store/
|
||||
trace_integrity/grading_engine from app.state instead of constructing
|
||||
them; same pattern as test_telemetry_ingest.py / test_sandboxes.py). The
|
||||
engine binds a scripted provider so each test controls the LLM exactly,
|
||||
and counts calls so gate tests can assert the LLM was never reached
|
||||
(G-4 holds through the whole HTTP stack).
|
||||
|
||||
Zero network: providers are MockProvider subclasses only (conftest rule).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from collections.abc import Iterator
|
||||
from datetime import UTC, datetime, timedelta
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from ai_service.config import Settings
|
||||
from ai_service.corpus.trace_fixtures import STRONG_BUILDER
|
||||
from ai_service.grading.engine import GradingEngine
|
||||
from ai_service.grading.store import SQLiteGradeStore
|
||||
from ai_service.llm.mock import MockProvider, ScriptedJSONProvider
|
||||
from ai_service.main import create_app
|
||||
from ai_service.telemetry.ingest import TraceIntegrityMap
|
||||
from ai_service.telemetry.models import TelemetryEvent
|
||||
from ai_service.telemetry.store import SQLiteTraceStore
|
||||
|
||||
LEARNER = "grade-learner"
|
||||
TASK = "task-grade-1"
|
||||
T0 = datetime(2026, 9, 12, 3, 0, 0, tzinfo=UTC)
|
||||
|
||||
#: Canonical well-formed rubric payload (matches grading.engine.RubricScore).
|
||||
RUBRIC_PAYLOAD: dict = {
|
||||
"criteria": {
|
||||
"process_quality": 4,
|
||||
"correctness": 4,
|
||||
"debugging_discipline": 3,
|
||||
"test_usage": 4,
|
||||
},
|
||||
"strengths": ["tight edit-test loops throughout"],
|
||||
"gaps": ["final commit discipline loose"],
|
||||
"verdict": "mastered",
|
||||
}
|
||||
|
||||
|
||||
class CountingScriptedProvider(ScriptedJSONProvider):
|
||||
"""ScriptedJSONProvider that counts chat() calls.
|
||||
|
||||
Gate tests assert calls == 0 (the LLM is never reached through the
|
||||
whole HTTP stack — G-4); happy-path tests sanity-check calls >= 1.
|
||||
"""
|
||||
|
||||
def __init__(self, payload: dict) -> None:
|
||||
super().__init__(payload)
|
||||
self.calls = 0
|
||||
|
||||
async def chat(self, messages, *, model, temperature=0.7, response_format=None):
|
||||
self.calls += 1
|
||||
return await super().chat(
|
||||
messages, model=model, temperature=temperature, response_format=response_format
|
||||
)
|
||||
|
||||
|
||||
def _event(seq: int, kind: str, payload: dict | None = None, offset_s: float = 0.0):
|
||||
return TelemetryEvent(
|
||||
learner_id=LEARNER,
|
||||
task_id=TASK,
|
||||
seq=seq,
|
||||
kind=kind,
|
||||
payload=payload or {},
|
||||
ts=T0 + timedelta(seconds=offset_s),
|
||||
sandbox_id="sbx-grade",
|
||||
)
|
||||
|
||||
|
||||
def _complete_trace() -> list[TelemetryEvent]:
|
||||
"""Contiguous seq 0..6 — a gradeable trace (gap-free, unflagged)."""
|
||||
return [
|
||||
_event(0, "activity", {"state": "starting"}, 0.0),
|
||||
_event(1, "file_diff", {"path": "a.py", "added": 12}, 10.0),
|
||||
_event(2, "command", {"cmd": "pytest -q"}, 20.0),
|
||||
_event(3, "test_result", {"passed": False, "exit_code": 1}, 25.0),
|
||||
_event(4, "file_diff", {"path": "a.py", "added": 4, "removed": 2}, 40.0),
|
||||
_event(5, "command", {"cmd": "pytest -q"}, 60.0),
|
||||
_event(6, "test_result", {"passed": True, "exit_code": 0}, 65.0),
|
||||
]
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def trace_store(tmp_path: Path) -> Iterator[SQLiteTraceStore]:
|
||||
store = SQLiteTraceStore(db_path=tmp_path / "traces.db")
|
||||
yield store
|
||||
store.close()
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def grade_store(tmp_path: Path) -> Iterator[SQLiteGradeStore]:
|
||||
store = SQLiteGradeStore(db_path=tmp_path / "grades.db")
|
||||
yield store
|
||||
store.close()
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def integrity() -> TraceIntegrityMap:
|
||||
return TraceIntegrityMap()
|
||||
|
||||
|
||||
def _make_client(
|
||||
tmp_path: Path,
|
||||
trace_store: SQLiteTraceStore,
|
||||
grade_store: SQLiteGradeStore,
|
||||
integrity: TraceIntegrityMap,
|
||||
provider: MockProvider,
|
||||
) -> TestClient:
|
||||
"""App + TestClient with stores, integrity map and a pre-set engine.
|
||||
|
||||
The lifespan adopts every pre-set service (state-injection override);
|
||||
the engine binds OUR provider, so the fixture — not Settings — scripts
|
||||
the LLM. Cloud-free guard: the provider must be a MockProvider family
|
||||
member (conftest rule, enforced here because this module builds its
|
||||
own client rather than consuming the conftest one).
|
||||
"""
|
||||
assert isinstance(provider, MockProvider)
|
||||
settings = Settings(
|
||||
provider="mock",
|
||||
db_path=tmp_path / "grading-test.db",
|
||||
sandbox_dir=tmp_path / "sandboxes",
|
||||
)
|
||||
app = create_app(settings)
|
||||
app.state.trace_store = trace_store
|
||||
app.state.grade_store = grade_store
|
||||
app.state.trace_integrity = integrity
|
||||
app.state.grading_engine = GradingEngine(
|
||||
trace_store, grade_store, integrity, provider, model="gemma4:31b"
|
||||
)
|
||||
return TestClient(app)
|
||||
|
||||
|
||||
def _seed(store: SQLiteTraceStore, events: list[TelemetryEvent]) -> None:
|
||||
for e in events:
|
||||
store.append(e)
|
||||
|
||||
|
||||
def _post(client: TestClient, learner: str = LEARNER, task: str = TASK):
|
||||
return client.post("/v1/assessment/grade", json={"learner_id": learner, "task_id": task})
|
||||
|
||||
|
||||
def _get(client: TestClient, learner: str = LEARNER, task: str = TASK):
|
||||
return client.get(f"/v1/assessment/grade/{learner}/{task}")
|
||||
|
||||
|
||||
# -- happy path: complete trace → GRADED -----------------------------------------
|
||||
|
||||
|
||||
class TestGraded:
|
||||
def test_post_complete_trace_returns_rubric_scores_and_digest(
|
||||
self, tmp_path, trace_store, grade_store, integrity
|
||||
):
|
||||
_seed(trace_store, _complete_trace())
|
||||
provider = CountingScriptedProvider(RUBRIC_PAYLOAD)
|
||||
|
||||
with _make_client(tmp_path, trace_store, grade_store, integrity, provider) as client:
|
||||
response = _post(client)
|
||||
|
||||
assert response.status_code == 200
|
||||
body = response.json()
|
||||
assert body["verdict"] == "GRADED"
|
||||
assert body["learner_id"] == LEARNER
|
||||
assert body["task_id"] == TASK
|
||||
assert body["scores"] == RUBRIC_PAYLOAD # rubric verdict rides in scores
|
||||
assert body["scores"]["verdict"] == "mastered"
|
||||
assert body["variant_seed"] is None # null until P4 (D-029)
|
||||
assert body["model"] == "gemma4:31b" # provenance travels
|
||||
assert provider.calls == 1 # happy path: exactly one LLM call
|
||||
# digest summary rides along (D-028 reproducible input)
|
||||
assert body["digest"]["event_count"] == 7
|
||||
assert body["digest"]["final_test_status"] == "pass"
|
||||
|
||||
def test_corpus_fixture_trace_grades(
|
||||
self, tmp_path, trace_store, grade_store, integrity
|
||||
):
|
||||
"""The strong-builder corpus fixture (real calibration trace) is
|
||||
gradeable over HTTP end to end."""
|
||||
_seed(trace_store, STRONG_BUILDER.events)
|
||||
provider = CountingScriptedProvider(RUBRIC_PAYLOAD)
|
||||
|
||||
with _make_client(tmp_path, trace_store, grade_store, integrity, provider) as client:
|
||||
response = _post(
|
||||
client, learner=STRONG_BUILDER.learner_id, task=STRONG_BUILDER.task_id
|
||||
)
|
||||
|
||||
assert response.status_code == 200
|
||||
body = response.json()
|
||||
assert body["verdict"] == "GRADED"
|
||||
assert body["digest"]["event_count"] == len(STRONG_BUILDER.event_specs)
|
||||
|
||||
def test_get_after_post_returns_same_stored_scores(
|
||||
self, tmp_path, trace_store, grade_store, integrity
|
||||
):
|
||||
_seed(trace_store, _complete_trace())
|
||||
provider = CountingScriptedProvider(RUBRIC_PAYLOAD)
|
||||
|
||||
with _make_client(tmp_path, trace_store, grade_store, integrity, provider) as client:
|
||||
posted = _post(client)
|
||||
assert posted.status_code == 200
|
||||
|
||||
fetched = _get(client)
|
||||
assert fetched.status_code == 200
|
||||
stored = fetched.json()
|
||||
|
||||
assert stored["scores"] == posted.json()["scores"]
|
||||
assert stored["verdict"] == "GRADED"
|
||||
assert stored["digest"] == posted.json()["digest"]
|
||||
assert stored["created_at"] == posted.json()["created_at"]
|
||||
|
||||
def test_post_regrade_upserts_get_returns_latest(
|
||||
self, tmp_path, trace_store, grade_store, integrity
|
||||
):
|
||||
"""POST twice → one row holding the LATEST grade (GradeStore upsert)."""
|
||||
_seed(trace_store, _complete_trace())
|
||||
provider = CountingScriptedProvider(RUBRIC_PAYLOAD)
|
||||
|
||||
with _make_client(tmp_path, trace_store, grade_store, integrity, provider) as client:
|
||||
first = _post(client)
|
||||
assert first.status_code == 200
|
||||
assert first.json()["scores"]["verdict"] == "mastered"
|
||||
|
||||
# the same provider now scripts a different (weaker) rubric
|
||||
updated = {
|
||||
**RUBRIC_PAYLOAD,
|
||||
"criteria": {**RUBRIC_PAYLOAD["criteria"], "process_quality": 1},
|
||||
"verdict": "not_yet",
|
||||
}
|
||||
provider.payload = updated
|
||||
second = _post(client)
|
||||
assert second.status_code == 200
|
||||
assert second.json()["scores"] == updated
|
||||
|
||||
# GET returns the LATEST grade, not the first
|
||||
fetched = _get(client)
|
||||
assert fetched.json()["scores"] == updated
|
||||
assert fetched.json()["scores"]["verdict"] == "not_yet"
|
||||
|
||||
# one row, latest wins (upsert, not append)
|
||||
grades = grade_store.list_for_learner(LEARNER)
|
||||
assert len(grades) == 1
|
||||
assert grades[0].scores == updated
|
||||
|
||||
|
||||
# -- G-4 gate outcomes over HTTP: 200 with the ungradable record -------------------
|
||||
|
||||
|
||||
class TestGateOutcomes:
|
||||
def test_gapped_trace_200_ungradable_incomplete_with_missing_seqs(
|
||||
self, tmp_path, trace_store, grade_store, integrity
|
||||
):
|
||||
_seed(trace_store, [e for e in _complete_trace() if e.seq != 2]) # seq 2 missing
|
||||
provider = CountingScriptedProvider(RUBRIC_PAYLOAD)
|
||||
|
||||
with _make_client(tmp_path, trace_store, grade_store, integrity, provider) as client:
|
||||
response = _post(client)
|
||||
|
||||
assert response.status_code == 200 # ungradable IS a valid result — not 5xx
|
||||
body = response.json()
|
||||
assert body["verdict"] == "UNGRADABLE_TRACE_INCOMPLETE"
|
||||
assert body["scores"]["missing_seqs"] == [2]
|
||||
assert body["scores"]["integrity_flag"] is None
|
||||
assert body["digest"] == {} # nothing was graded
|
||||
assert body["model"] == "none" # no LLM involved (honest provenance)
|
||||
assert provider.calls == 0, "LLM was called despite a gapped trace (G-4)"
|
||||
|
||||
def test_flooded_trace_200_ungradable_incomplete_with_integrity_reason(
|
||||
self, tmp_path, trace_store, grade_store, integrity
|
||||
):
|
||||
"""Integrity-flagged trace (G-3 INCOMPLETE_FLOODED) surfaces the flag
|
||||
reason in scores — the same verdict as gaps, different detail."""
|
||||
_seed(trace_store, _complete_trace())
|
||||
integrity.mark(LEARNER, TASK, "INCOMPLETE_FLOODED")
|
||||
provider = CountingScriptedProvider(RUBRIC_PAYLOAD)
|
||||
|
||||
with _make_client(tmp_path, trace_store, grade_store, integrity, provider) as client:
|
||||
response = _post(client)
|
||||
|
||||
assert response.status_code == 200
|
||||
body = response.json()
|
||||
assert body["verdict"] == "UNGRADABLE_TRACE_INCOMPLETE"
|
||||
assert body["scores"]["integrity_flag"] == "INCOMPLETE_FLOODED"
|
||||
assert body["scores"]["missing_seqs"] == [] # rows complete but untrusted
|
||||
assert provider.calls == 0, "LLM was called despite an integrity flag (G-4)"
|
||||
|
||||
def test_empty_trace_200_ungradable_empty(
|
||||
self, tmp_path, trace_store, grade_store, integrity
|
||||
):
|
||||
"""Zero events for the pair (including a task that never had a
|
||||
trace) → UNGRADABLE_EMPTY_TRACE, persisted like any grade."""
|
||||
provider = CountingScriptedProvider(RUBRIC_PAYLOAD)
|
||||
|
||||
with _make_client(tmp_path, trace_store, grade_store, integrity, provider) as client:
|
||||
response = _post(client, learner="ghost-learner", task="never-started")
|
||||
|
||||
assert response.status_code == 200
|
||||
body = response.json()
|
||||
assert body["verdict"] == "UNGRADABLE_EMPTY_TRACE"
|
||||
assert body["scores"] == {"integrity_flag": None, "missing_seqs": []}
|
||||
assert body["digest"] == {}
|
||||
assert body["model"] == "none"
|
||||
assert provider.calls == 0
|
||||
|
||||
# the gate record is durable: GET now returns it (no longer 404)
|
||||
with _make_client(tmp_path, trace_store, grade_store, integrity, provider) as client:
|
||||
fetched = _get(client, learner="ghost-learner", task="never-started")
|
||||
assert fetched.status_code == 200
|
||||
assert fetched.json()["verdict"] == "UNGRADABLE_EMPTY_TRACE"
|
||||
|
||||
|
||||
# -- provider failure → 502 --------------------------------------------------------
|
||||
|
||||
|
||||
class TestProviderFailure:
|
||||
def test_persistent_malformed_llm_output_502(
|
||||
self, tmp_path, trace_store, grade_store, integrity
|
||||
):
|
||||
"""Stock MockProvider: its json_object reply is wrong-shaped, so the
|
||||
D-020 defense exhausts both attempts → endpoint maps to 502 and
|
||||
nothing is persisted."""
|
||||
_seed(trace_store, _complete_trace())
|
||||
provider = MockProvider() # always malformed for the rubric schema
|
||||
|
||||
with _make_client(tmp_path, trace_store, grade_store, integrity, provider) as client:
|
||||
response = _post(client)
|
||||
|
||||
assert response.status_code == 502
|
||||
assert "grading failed" in response.json()["detail"]
|
||||
# no fabricated/partial record was persisted
|
||||
assert grade_store.get(LEARNER, TASK) is None
|
||||
|
||||
def test_502_leaves_earlier_grade_intact(
|
||||
self, tmp_path, trace_store, grade_store, integrity
|
||||
):
|
||||
"""A failed regrade must not clobber the previously stored grade."""
|
||||
_seed(trace_store, _complete_trace())
|
||||
scripted = CountingScriptedProvider(RUBRIC_PAYLOAD)
|
||||
with _make_client(tmp_path, trace_store, grade_store, integrity, scripted) as client:
|
||||
assert _post(client).status_code == 200
|
||||
|
||||
# regrade attempt hits a provider that now always fails
|
||||
broken = MockProvider()
|
||||
with _make_client(tmp_path, trace_store, grade_store, integrity, broken) as client:
|
||||
assert _post(client).status_code == 502
|
||||
fetched = _get(client)
|
||||
|
||||
assert fetched.status_code == 200 # earlier grade still readable
|
||||
assert fetched.json()["scores"] == RUBRIC_PAYLOAD
|
||||
|
||||
|
||||
# -- stored-grade reads ------------------------------------------------------------
|
||||
|
||||
|
||||
class TestStoredGradeReads:
|
||||
def test_get_unknown_pair_404(self, tmp_path, trace_store, grade_store, integrity):
|
||||
provider = CountingScriptedProvider(RUBRIC_PAYLOAD)
|
||||
with _make_client(tmp_path, trace_store, grade_store, integrity, provider) as client:
|
||||
response = _get(client, learner="nobody", task="never-graded")
|
||||
|
||||
assert response.status_code == 404
|
||||
assert "no stored grade" in response.json()["detail"]
|
||||
|
||||
def test_missing_body_fields_422(self, tmp_path, trace_store, grade_store, integrity):
|
||||
provider = CountingScriptedProvider(RUBRIC_PAYLOAD)
|
||||
with _make_client(tmp_path, trace_store, grade_store, integrity, provider) as client:
|
||||
assert client.post("/v1/assessment/grade", json={}).status_code == 422
|
||||
assert (
|
||||
client.post(
|
||||
"/v1/assessment/grade", json={"learner_id": LEARNER}
|
||||
).status_code
|
||||
== 422
|
||||
)
|
||||
@@ -1,47 +1,86 @@
|
||||
"""Lab feedback endpoint tests — SSE envelope with agent=lab (REQ-2-007)."""
|
||||
"""Lab endpoint tests — LIVE trace contract (REQ-3-007).
|
||||
|
||||
import json
|
||||
v0.3 re-grounding: POST /v1/lab/feedback takes {learner_id, task_id}; the
|
||||
digest is computed from the learner's real TraceStore events. No corpus
|
||||
scenarios; empty trace is a valid "no telemetry yet" coaching path.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from datetime import UTC, datetime, timedelta
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from ai_service.config import Settings
|
||||
from ai_service.grading.store import SQLiteGradeStore
|
||||
from ai_service.llm.mock import MockProvider
|
||||
from ai_service.main import create_app
|
||||
from ai_service.telemetry.ingest import TraceIntegrityMap
|
||||
from ai_service.telemetry.models import TelemetryEvent
|
||||
from ai_service.telemetry.store import SQLiteTraceStore
|
||||
|
||||
T0 = datetime(2026, 9, 12, tzinfo=UTC)
|
||||
|
||||
|
||||
def stream_events(client, payload) -> list[dict]:
|
||||
with client.stream("POST", "/v1/lab/feedback", json=payload) as response:
|
||||
assert response.status_code == 200
|
||||
events = []
|
||||
for line in response.iter_lines():
|
||||
if line.startswith("data:"):
|
||||
d = line.removeprefix("data:").strip()
|
||||
if d == "[DONE]":
|
||||
events.append({"type": "[DONE]"})
|
||||
else:
|
||||
events.append(json.loads(d))
|
||||
return events
|
||||
def _event(seq: int, kind: str, payload: dict, offset_s: float):
|
||||
return TelemetryEvent(
|
||||
learner_id="lab-learner",
|
||||
task_id="lab-task",
|
||||
seq=seq,
|
||||
kind=kind,
|
||||
payload=payload,
|
||||
ts=T0 + timedelta(seconds=offset_s),
|
||||
sandbox_id="sbx-lab",
|
||||
)
|
||||
|
||||
|
||||
def test_lab_feedback_streams_full_envelope(client):
|
||||
events = stream_events(client, {"scenario_id": "lab-scenario-strong"})
|
||||
assert events[0]["type"] == "meta"
|
||||
assert events[0]["agent"] == "lab"
|
||||
assert events[0]["scenario_id"] == "lab-scenario-strong"
|
||||
deltas = [e for e in events if e["type"] == "delta"]
|
||||
assert len(deltas) >= 1
|
||||
assert any(e["type"] == "done" for e in events)
|
||||
assert events[-1]["type"] == "[DONE]"
|
||||
class StreamingMock(MockProvider):
|
||||
"""Deterministic token stream for the SSE path."""
|
||||
|
||||
def _reply_for(self, messages, response_format):
|
||||
return "Feedback grounded in your live session digest."
|
||||
|
||||
|
||||
def test_unknown_scenario_404(client):
|
||||
response = client.post("/v1/lab/feedback", json={"scenario_id": "nope"})
|
||||
assert response.status_code == 404
|
||||
assert "nope" in response.json()["detail"]
|
||||
@pytest.fixture()
|
||||
def client(tmp_path: Path) -> TestClient:
|
||||
app = create_app(Settings(provider="mock"))
|
||||
app.state.provider = StreamingMock()
|
||||
app.state.trace_store = SQLiteTraceStore(db_path=tmp_path / "t.db")
|
||||
app.state.grade_store = SQLiteGradeStore(db_path=tmp_path / "g.db")
|
||||
app.state.trace_integrity = TraceIntegrityMap()
|
||||
with TestClient(app) as c:
|
||||
yield c
|
||||
|
||||
|
||||
def test_distinct_scenarios_distinct_replies(client):
|
||||
strong = stream_events(client, {"scenario_id": "lab-scenario-strong"})
|
||||
struggling = stream_events(client, {"scenario_id": "lab-scenario-struggling"})
|
||||
strong_text = "".join(e["content"] for e in strong if e["type"] == "delta")
|
||||
struggling_text = "".join(e["content"] for e in struggling if e["type"] == "delta")
|
||||
assert strong_text != struggling_text
|
||||
def test_live_trace_streams_full_envelope(client: TestClient) -> None:
|
||||
store: SQLiteTraceStore = client.app.state.trace_store
|
||||
for e in [
|
||||
_event(0, "file_diff", {"path": "x.py"}, 0),
|
||||
_event(1, "command", {"cmd": "pytest -q"}, 10),
|
||||
_event(2, "test_result", {"passed": False, "exit_code": 1}, 20),
|
||||
_event(3, "test_result", {"passed": True, "exit_code": 0}, 40),
|
||||
]:
|
||||
store.append(e)
|
||||
with client.stream(
|
||||
"POST", "/v1/lab/feedback", json={"learner_id": "lab-learner", "task_id": "lab-task"}
|
||||
) as resp:
|
||||
assert resp.status_code == 200
|
||||
body = "".join(chunk.decode() for chunk in resp.iter_raw())
|
||||
assert '"agent": "lab"' in body or '"agent":"lab"' in body
|
||||
assert '"task_id": "lab-task"' in body
|
||||
assert '"type": "delta"' in body or '"type":"delta"' in body
|
||||
|
||||
|
||||
def test_missing_scenario_id_422(client):
|
||||
response = client.post("/v1/lab/feedback", json={})
|
||||
assert response.status_code == 422
|
||||
def test_empty_trace_coaches_the_baseline(client: TestClient) -> None:
|
||||
"""No telemetry is NOT an error — Lab coaches 'run the starter test'."""
|
||||
with client.stream(
|
||||
"POST", "/v1/lab/feedback", json={"learner_id": "lab-learner", "task_id": "no-events"}
|
||||
) as resp:
|
||||
assert resp.status_code == 200
|
||||
|
||||
|
||||
def test_missing_fields_422(client: TestClient) -> None:
|
||||
resp = client.post("/v1/lab/feedback", json={"learner_id": "x"})
|
||||
assert resp.status_code == 422
|
||||
|
||||
@@ -1,49 +1,121 @@
|
||||
"""Proctor signals endpoint tests — validated JSON, 404s (REQ-2-009)."""
|
||||
"""Proctor signals endpoint tests — REAL inputs contract (REQ-3-007).
|
||||
|
||||
from ai_service.agents.proctor import ProctorAssessment
|
||||
from ai_service.llm.mock import ScriptedJSONProvider
|
||||
v0.3 re-grounding: POST /v1/proctor/signals takes {learner_id, task_id}
|
||||
and gathers digest + defense signals + variant context server-side.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from datetime import UTC, datetime, timedelta
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from ai_service.config import Settings
|
||||
from ai_service.grading.store import SQLiteGradeStore
|
||||
from ai_service.llm.mock import MockProvider
|
||||
from ai_service.main import create_app
|
||||
from ai_service.telemetry.ingest import TraceIntegrityMap
|
||||
from ai_service.telemetry.models import TelemetryEvent
|
||||
from ai_service.telemetry.store import SQLiteTraceStore
|
||||
from ai_service.variants.store import SQLiteVariantStore
|
||||
from ai_service.voice.defense_store import SQLiteDefenseStore
|
||||
|
||||
VALID = {
|
||||
"scenario_id": "proctor-scenario-distracted",
|
||||
"signals": [
|
||||
{"signal_type": "context_switch", "severity": "low",
|
||||
"note": "Docs tab at t+120s is normal"},
|
||||
{"signal_type": "idle_gap", "severity": "medium",
|
||||
"note": "5-minute idle at t+300s"},
|
||||
{"signal_type": "idle_gap", "severity": "low", "note": "One long pause."},
|
||||
],
|
||||
"intervention": "Offer a short break, then restate the plan",
|
||||
"summary": "Coaching-shaped session note",
|
||||
}
|
||||
|
||||
|
||||
def test_signals_returns_validated_json(client):
|
||||
original = client.app.state.provider
|
||||
client.app.state.provider = ScriptedJSONProvider(VALID)
|
||||
try:
|
||||
response = client.post(
|
||||
"/v1/proctor/signals", json={"scenario_id": "proctor-scenario-distracted"}
|
||||
)
|
||||
finally:
|
||||
client.app.state.provider = original
|
||||
assert response.status_code == 200
|
||||
validated = ProctorAssessment.model_validate(response.json())
|
||||
assert validated.scenario_id == "proctor-scenario-distracted"
|
||||
assert validated.intervention
|
||||
T0 = datetime(2026, 9, 12, tzinfo=UTC)
|
||||
|
||||
|
||||
def test_unknown_scenario_404(client):
|
||||
response = client.post("/v1/proctor/signals", json={"scenario_id": "ghost"})
|
||||
assert response.status_code == 404
|
||||
class ProctorJSON(MockProvider):
|
||||
def _reply_for(self, messages, response_format):
|
||||
if response_format is not None and response_format.get("type") == "json_object":
|
||||
return json.dumps(VALID)
|
||||
return super()._reply_for(messages, response_format)
|
||||
|
||||
|
||||
def test_unparseable_provider_502(client):
|
||||
response = client.post(
|
||||
"/v1/proctor/signals", json={"scenario_id": "proctor-scenario-healthy"}
|
||||
class BrokenJSON(MockProvider):
|
||||
def _reply_for(self, messages, response_format):
|
||||
if response_format is not None and response_format.get("type") == "json_object":
|
||||
return "not json ever"
|
||||
return super()._reply_for(messages, response_format)
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def client(tmp_path: Path) -> TestClient:
|
||||
app = create_app(Settings(provider="mock"))
|
||||
app.state.trace_store = SQLiteTraceStore(db_path=tmp_path / "t.db")
|
||||
app.state.grade_store = SQLiteGradeStore(db_path=tmp_path / "g.db")
|
||||
app.state.variant_store = SQLiteVariantStore(db_path=tmp_path / "v.db")
|
||||
app.state.defense_store = SQLiteDefenseStore(db_path=tmp_path / "d.db")
|
||||
app.state.trace_integrity = TraceIntegrityMap()
|
||||
with TestClient(app) as c:
|
||||
yield c
|
||||
|
||||
|
||||
def _seed_trace(client: TestClient) -> None:
|
||||
store: SQLiteTraceStore = client.app.state.trace_store
|
||||
for e in [
|
||||
TelemetryEvent(
|
||||
learner_id="p-learner",
|
||||
task_id="p-task",
|
||||
seq=0,
|
||||
kind="activity",
|
||||
payload={"state": "idle"},
|
||||
ts=T0,
|
||||
sandbox_id="sbx-p",
|
||||
),
|
||||
TelemetryEvent(
|
||||
learner_id="p-learner",
|
||||
task_id="p-task",
|
||||
seq=1,
|
||||
kind="activity",
|
||||
payload={"state": "idle"},
|
||||
ts=T0 + timedelta(seconds=400),
|
||||
sandbox_id="sbx-p",
|
||||
),
|
||||
]:
|
||||
store.append(e)
|
||||
|
||||
|
||||
def test_signals_returns_validated_json(client: TestClient) -> None:
|
||||
client.app.state.provider = ProctorJSON()
|
||||
_seed_trace(client)
|
||||
resp = client.post(
|
||||
"/v1/proctor/signals", json={"learner_id": "p-learner", "task_id": "p-task"}
|
||||
)
|
||||
assert response.status_code == 502
|
||||
assert "failed" in response.json()["detail"].lower()
|
||||
assert resp.status_code == 200, resp.text
|
||||
body = resp.json()
|
||||
assert body["intervention"]
|
||||
assert body["signals"][0]["signal_type"] == "idle_gap"
|
||||
|
||||
|
||||
def test_missing_scenario_id_422(client):
|
||||
response = client.post("/v1/proctor/signals", json={})
|
||||
assert response.status_code == 422
|
||||
def test_empty_trace_is_valid_not_404(client: TestClient) -> None:
|
||||
"""No telemetry → the proctor still assesses (nothing to flag)."""
|
||||
client.app.state.provider = ProctorJSON()
|
||||
resp = client.post(
|
||||
"/v1/proctor/signals", json={"learner_id": "nobody", "task_id": "nothing"}
|
||||
)
|
||||
assert resp.status_code == 200
|
||||
|
||||
|
||||
def test_unparseable_provider_502(client: TestClient) -> None:
|
||||
client.app.state.provider = BrokenJSON()
|
||||
_seed_trace(client)
|
||||
resp = client.post(
|
||||
"/v1/proctor/signals", json={"learner_id": "p-learner", "task_id": "p-task"}
|
||||
)
|
||||
assert resp.status_code == 502
|
||||
|
||||
|
||||
def test_missing_fields_422(client: TestClient) -> None:
|
||||
client.app.state.provider = ProctorJSON()
|
||||
resp = client.post("/v1/proctor/signals", json={"learner_id": "x"})
|
||||
assert resp.status_code == 422
|
||||
|
||||
@@ -359,3 +359,91 @@ def test_real_backend_create_path_runs(tmp_path: Path) -> None:
|
||||
assert (sandbox_root / sandbox_id / "workspace").is_dir()
|
||||
finally:
|
||||
shutil.rmtree(sandbox_root, ignore_errors=True)
|
||||
|
||||
|
||||
class TestFilesAndExecRoutes:
|
||||
"""Workspace CRUD + Run/Test exec (Phase 6, REQ-3-008, CUT-2)."""
|
||||
|
||||
def test_file_write_read_list_roundtrip(self, client):
|
||||
handle = client.post(
|
||||
"/v1/sandboxes", json={"learner_id": "pilot-learner"}
|
||||
).json()
|
||||
sbx = handle["id"]
|
||||
put = client.put(
|
||||
f"/v1/sandboxes/{sbx}/files/main.py",
|
||||
json={"path": "main.py", "content": "print('hi')"},
|
||||
)
|
||||
assert put.status_code == 200, put.text
|
||||
got = client.get(f"/v1/sandboxes/{sbx}/files/main.py")
|
||||
assert got.status_code == 200
|
||||
assert "print('hi')" in got.json()["content"]
|
||||
listed = client.get(f"/v1/sandboxes/{sbx}/files")
|
||||
assert "main.py" in listed.json()["files"]
|
||||
client.delete(f"/v1/sandboxes/{sbx}")
|
||||
|
||||
def test_traversal_rejected(self, client):
|
||||
handle = client.post(
|
||||
"/v1/sandboxes", json={"learner_id": "pilot-learner"}
|
||||
).json()
|
||||
sbx = handle["id"]
|
||||
bad = client.put(
|
||||
f"/v1/sandboxes/{sbx}/files/..%2Fescape.txt",
|
||||
json={"path": "../escape.txt", "content": "x"},
|
||||
)
|
||||
assert bad.status_code == 422
|
||||
client.delete(f"/v1/sandboxes/{sbx}")
|
||||
|
||||
def test_unknown_file_404(self, client):
|
||||
handle = client.post(
|
||||
"/v1/sandboxes", json={"learner_id": "pilot-learner"}
|
||||
).json()
|
||||
sbx = handle["id"]
|
||||
assert client.get(f"/v1/sandboxes/{sbx}/files/ghost.py").status_code == 404
|
||||
client.delete(f"/v1/sandboxes/{sbx}")
|
||||
|
||||
def test_exec_unknown_sandbox_404(self, client):
|
||||
resp = client.post(
|
||||
"/v1/sandboxes/sbx-nope/exec", json={"cmd": ["echo", "hi"]}
|
||||
)
|
||||
assert resp.status_code == 404
|
||||
|
||||
def test_unknown_sandbox_file_routes_404_not_500(self, client):
|
||||
"""P7: read/write on an unknown sandbox must 404 (SandboxNotFoundError
|
||||
previously escaped _workspace_dir as an unhandled 500)."""
|
||||
assert (
|
||||
client.get("/v1/sandboxes/sbx-nope/files/whatever.py").status_code == 404
|
||||
)
|
||||
put = client.put(
|
||||
"/v1/sandboxes/sbx-nope/files/whatever.py",
|
||||
json={"path": "whatever.py", "content": "x"},
|
||||
)
|
||||
assert put.status_code == 404
|
||||
|
||||
def test_symlink_escape_rejected(self, client):
|
||||
"""P7: an exec-planted symlink in the workspace must not let the
|
||||
file routes read/write OUTSIDE the bind (lexical traversal checks
|
||||
cannot see symlinks — resolve + containment re-check is the gate)."""
|
||||
handle = client.post(
|
||||
"/v1/sandboxes", json={"learner_id": "pilot-learner"}
|
||||
).json()
|
||||
sbx = handle["id"]
|
||||
workspace = Path(handle["workdir"]) / "workspace"
|
||||
outside = workspace.parent / "secret.txt"
|
||||
outside.write_text("host secret") # a host file OUTSIDE the bind
|
||||
try:
|
||||
(workspace / "leak.txt").symlink_to(outside)
|
||||
read = client.get(f"/v1/sandboxes/{sbx}/files/leak.txt")
|
||||
assert read.status_code == 422, (
|
||||
f"symlink escape read must 422, got {read.status_code}: {read.text}"
|
||||
)
|
||||
write = client.put(
|
||||
f"/v1/sandboxes/{sbx}/files/leak.txt",
|
||||
json={"path": "leak.txt", "content": "pwned"},
|
||||
)
|
||||
assert write.status_code == 422, (
|
||||
f"symlink escape write must 422, got {write.status_code}: {write.text}"
|
||||
)
|
||||
assert outside.read_text() == "host secret" # untouched
|
||||
finally:
|
||||
client.delete(f"/v1/sandboxes/{sbx}")
|
||||
outside.unlink(missing_ok=True)
|
||||
|
||||
@@ -0,0 +1,451 @@
|
||||
"""WS telemetry ingest API tests (REQ-3-003, D-026, G-3).
|
||||
|
||||
Frame contract under test (telemetry/ingest.py):
|
||||
|
||||
WS /v1/telemetry/ingest?learner_id=L&task_id=T[&sandbox_id=S]
|
||||
client → server: one event per JSON text frame
|
||||
{"seq", "kind", "payload", "ts", "sandbox_id"?}
|
||||
server → client: {"type": "gap_warning"|"event_rejected"|"flooded"|"ack_total"}
|
||||
close: 1000 clean flush · 1008 policy violation (G-3 flood)
|
||||
|
||||
Each test runs against a per-test SQLite file in tmp_path via the same
|
||||
state-injection override the sandbox API tests use: build a real
|
||||
SQLiteTraceStore, pre-set `app.state.trace_store`, let the lifespan adopt it.
|
||||
|
||||
Covered:
|
||||
- connect → send 3 events → GET trace returns them ordered
|
||||
- resend event 2 → deduped (no double-store; at-least-once contract)
|
||||
- skip to seq 5 → gap_warning frame + GET /gaps reports missing seqs
|
||||
- burst past a lowered cap → 1008 close + INCOMPLETE_FLOODED +
|
||||
is_incomplete(...) true
|
||||
- queue-overflow path → same 1008 + INCOMPLETE_FLOODED
|
||||
- reconnect after flood still cannot append (flag is terminal)
|
||||
- unknown trace GET → 404 (traces + gaps)
|
||||
- missing identity query params → handshake rejected
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import time
|
||||
from collections.abc import Iterator
|
||||
from datetime import UTC, datetime
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
from fastapi.testclient import TestClient
|
||||
from starlette.websockets import WebSocketDisconnect
|
||||
|
||||
from ai_service.config import Settings
|
||||
from ai_service.main import create_app
|
||||
from ai_service.telemetry import ingest as ingest_mod
|
||||
from ai_service.telemetry.store import SQLiteTraceStore
|
||||
|
||||
LEARNER = "learner-9"
|
||||
TASK = "task-9"
|
||||
_BASE_TS = datetime(2026, 9, 11, 12, 0, 0, tzinfo=UTC)
|
||||
|
||||
KNOWN_KINDS = ("command", "file_diff", "run_result", "test_result", "activity")
|
||||
|
||||
|
||||
def _frame(seq: int, kind: str = "command", sandbox_id: str = "") -> str:
|
||||
return json.dumps(
|
||||
{
|
||||
"seq": seq,
|
||||
"kind": kind,
|
||||
"payload": {"n": seq},
|
||||
"ts": _BASE_TS.isoformat(),
|
||||
**({"sandbox_id": sandbox_id} if sandbox_id else {}),
|
||||
}
|
||||
)
|
||||
|
||||
|
||||
def _drain(ws) -> None:
|
||||
"""Give the portal a beat to deliver queued frames.
|
||||
|
||||
Starlette's TestClient runs the app on a portal thread; frames produced in
|
||||
the app loop still need a portal round-trip to reach the test thread. A
|
||||
real (tiny) sleep flushes them deterministically without hardcoding
|
||||
protocol timing.
|
||||
"""
|
||||
time.sleep(0.05)
|
||||
|
||||
|
||||
def _recv_status(ws) -> dict:
|
||||
"""Next JSON status frame, skipping keepalive pings; re-raise close.
|
||||
|
||||
TestClient delivers close as a `websocket.close` message dict, not by
|
||||
itself raising — `_recv_status` must surface it as WebSocketDisconnect so
|
||||
`pytest.raises(WebSocketDisconnect)` around a read loop terminates.
|
||||
"""
|
||||
while True:
|
||||
msg = ws.receive()
|
||||
if msg.get("bytes") is not None: # keepalive ping — not a status frame
|
||||
continue
|
||||
if msg["type"] == "websocket.close":
|
||||
raise WebSocketDisconnect(msg.get("code", 1000), msg.get("reason", ""))
|
||||
return json.loads(msg["text"])
|
||||
|
||||
|
||||
def _ingest_url(learner_id: str = LEARNER, task_id: str = TASK, sandbox_id: str = "") -> str:
|
||||
url = f"/v1/telemetry/ingest?learner_id={learner_id}&task_id={task_id}"
|
||||
return f"{url}&sandbox_id={sandbox_id}" if sandbox_id else url
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def store(tmp_path: Path) -> Iterator[SQLiteTraceStore]:
|
||||
s = SQLiteTraceStore(db_path=tmp_path / "telemetry-test.db")
|
||||
yield s
|
||||
s.close()
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def app(store: SQLiteTraceStore, tmp_path: Path, monkeypatch: pytest.MonkeyPatch) -> Iterator:
|
||||
# Keep the keepalive interval comfortably above the per-test read window:
|
||||
# a ping that fires mid-assertion interleaves with status frames and the
|
||||
# reader must skip it; at 2s a single test's reads all land ping-free
|
||||
# while a *blocked* read still fails fast on the next ping tick.
|
||||
monkeypatch.setattr(ingest_mod, "PING_INTERVAL_S", 2.0)
|
||||
settings = Settings(
|
||||
provider="mock",
|
||||
db_path=tmp_path / "telemetry-test.db",
|
||||
sandbox_dir=tmp_path / "sandboxes",
|
||||
telemetry_max_events_per_task=5, # lowered G-3 cap for the flood test
|
||||
)
|
||||
application = create_app(settings)
|
||||
application.state.trace_store = store # state-injection override
|
||||
yield application
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def client(app) -> Iterator[TestClient]:
|
||||
with TestClient(app) as c:
|
||||
yield c
|
||||
|
||||
|
||||
# -- happy path: ordered trace retrieval --------------------------------------
|
||||
|
||||
|
||||
def test_connect_send_three_get_trace_ordered(client: TestClient) -> None:
|
||||
with client.websocket_connect(_ingest_url()) as ws:
|
||||
# Arrive out of arrival order? No — seq arrival order here; ordering is
|
||||
# by seq at read. Send 0,1,2 then close cleanly.
|
||||
for seq in (0, 1, 2):
|
||||
ws.send_text(_frame(seq))
|
||||
|
||||
response = client.get(f"/v1/telemetry/traces/{LEARNER}/{TASK}")
|
||||
assert response.status_code == 200
|
||||
body = response.json()
|
||||
assert [e["seq"] for e in body["events"]] == [0, 1, 2]
|
||||
assert [e["kind"] for e in body["events"]] == ["command"] * 3
|
||||
assert all(e["learner_id"] == LEARNER for e in body["events"])
|
||||
assert body["incomplete"] is False
|
||||
assert body["integrity_reason"] is None
|
||||
|
||||
|
||||
def test_different_kinds_and_sandbox_id_roundtrip(client: TestClient) -> None:
|
||||
with client.websocket_connect(_ingest_url(sandbox_id="sbx-42")) as ws:
|
||||
for seq, kind in enumerate(KNOWN_KINDS[:3]):
|
||||
ws.send_text(_frame(seq, kind))
|
||||
|
||||
body = client.get(f"/v1/telemetry/traces/{LEARNER}/{TASK}").json()
|
||||
assert [e["kind"] for e in body["events"]] == list(KNOWN_KINDS[:3])
|
||||
# Frame had no sandbox_id → the connection's query-param sandbox_id applies.
|
||||
assert all(e["sandbox_id"] == "sbx-42" for e in body["events"])
|
||||
|
||||
|
||||
# -- idempotent dedup (at-least-once, server-side) ------------------------------
|
||||
|
||||
|
||||
def test_resend_event_deduped_no_double_store(client: TestClient) -> None:
|
||||
with client.websocket_connect(_ingest_url()) as ws:
|
||||
for seq in (0, 1, 2):
|
||||
ws.send_text(_frame(seq))
|
||||
ws.send_text(_frame(1, kind="file_diff")) # at-least-once retry of seq 1
|
||||
|
||||
body = client.get(f"/v1/telemetry/traces/{LEARNER}/{TASK}").json()
|
||||
assert [e["seq"] for e in body["events"]] == [0, 1, 2] # stored exactly once
|
||||
# First write wins — the retry's different body must NOT overwrite.
|
||||
assert body["events"][1]["kind"] == "command"
|
||||
|
||||
|
||||
# -- gap detection --------------------------------------------------------------
|
||||
|
||||
|
||||
def test_skip_to_seq_5_reports_gap_warning_and_gaps_endpoint(
|
||||
client: TestClient,
|
||||
) -> None:
|
||||
with client.websocket_connect(_ingest_url()) as ws:
|
||||
for seq in (0, 1, 2):
|
||||
ws.send_text(_frame(seq))
|
||||
ws.send_text(_frame(5)) # skips 3,4 → gap_warning frame
|
||||
_drain(ws)
|
||||
frames = [_recv_status(ws)]
|
||||
assert frames == [{"type": "gap_warning", "missing_seqs": [3, 4]}]
|
||||
|
||||
gaps = client.get(f"/v1/telemetry/gaps/{LEARNER}/{TASK}")
|
||||
assert gaps.status_code == 200
|
||||
assert gaps.json()["gaps"] == [3, 4]
|
||||
assert gaps.json()["incomplete"] is False
|
||||
|
||||
|
||||
# -- G-3 flood control: cap breach → 1008 + INCOMPLETE_FLOODED --------------------
|
||||
|
||||
|
||||
def test_burst_past_cap_closes_1008_and_marks_incomplete_flooded(
|
||||
client: TestClient, app, store: SQLiteTraceStore
|
||||
) -> None:
|
||||
# Cap is 5 (fixture); the 6th event trips cap_exceeded.
|
||||
with client.websocket_connect(_ingest_url()) as ws:
|
||||
for seq in range(6):
|
||||
ws.send_text(_frame(seq))
|
||||
_drain(ws) # let the app flush the flooded frame + 1008 close
|
||||
frames: list[dict] = []
|
||||
with pytest.raises(WebSocketDisconnect) as excinfo:
|
||||
while True:
|
||||
frames.append(_recv_status(ws))
|
||||
assert excinfo.value.code == 1008
|
||||
assert frames == [{"type": "flooded", "reason": "cap_exceeded", "count": 6}]
|
||||
|
||||
# Trace is durably intact up to the cap — no silent drop corrupted it.
|
||||
assert [e.seq for e in store.get_trace(LEARNER, TASK)] == [0, 1, 2, 3, 4]
|
||||
|
||||
# The integrity signal Proctor/Phase-3 grader consume (G-4).
|
||||
assert app.state.trace_integrity.is_incomplete(LEARNER, TASK) is True
|
||||
assert app.state.trace_integrity.reason(LEARNER, TASK) == "INCOMPLETE_FLOODED"
|
||||
|
||||
# ...and it is readable over HTTP (trace with events + flag → 200, not 404).
|
||||
body = client.get(f"/v1/telemetry/traces/{LEARNER}/{TASK}").json()
|
||||
assert body["incomplete"] is True
|
||||
assert body["integrity_reason"] == "INCOMPLETE_FLOODED"
|
||||
|
||||
|
||||
def test_queue_overflow_also_floods(
|
||||
client: TestClient, app, monkeypatch: pytest.MonkeyPatch
|
||||
) -> None:
|
||||
"""G-3: inbound-queue overflow → same 1008 + INCOMPLETE_FLOODED (never
|
||||
drop-oldest). Lower the queue bound so one burst trips it without needing
|
||||
256 frames."""
|
||||
monkeypatch.setattr(ingest_mod, "INBOUND_QUEUE_MAX", 1)
|
||||
with client.websocket_connect(_ingest_url(learner_id="L-q", task_id="T-q")) as ws:
|
||||
for seq in range(64): # receiver out-drains the drainer with bound=1
|
||||
ws.send_text(_frame(seq))
|
||||
_drain(ws)
|
||||
with pytest.raises(WebSocketDisconnect) as excinfo:
|
||||
while True:
|
||||
_recv_status(ws)
|
||||
assert excinfo.value.code == 1008
|
||||
assert app.state.trace_integrity.is_incomplete("L-q", "T-q") is True
|
||||
assert app.state.trace_integrity.reason("L-q", "T-q") == "INCOMPLETE_FLOODED"
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_queue_overflow_flood_session_task_terminates(tmp_path, monkeypatch):
|
||||
"""P7 regression: the queue-overflow flood path must not LEAK the
|
||||
session coroutine. v0.3's receiver returned from its QueueFull branch
|
||||
without the disconnect sentinel, so the drainer parked on an empty queue
|
||||
forever and IngestSession.run() never returned — one leaked
|
||||
(pinger+drainer) task-set per flooded trace, unbounded over a long-lived
|
||||
process. A REAL uvicorn server (TestClient teardown hides the leak) is
|
||||
stopped after the flood; the session tasks must be gone shortly after.
|
||||
"""
|
||||
import asyncio
|
||||
import contextlib
|
||||
import socket as socket_mod
|
||||
|
||||
import uvicorn
|
||||
|
||||
from ai_service.main import create_app
|
||||
from ai_service.telemetry.ingest import TraceIntegrityMap
|
||||
from ai_service.telemetry.store import SQLiteTraceStore
|
||||
|
||||
monkeypatch.setattr(ingest_mod, "INBOUND_QUEUE_MAX", 1)
|
||||
store = SQLiteTraceStore(db_path=tmp_path / "leak.db")
|
||||
app = create_app(
|
||||
Settings(
|
||||
provider="mock",
|
||||
db_path=tmp_path / "leak.db",
|
||||
sandbox_dir=tmp_path / "sandboxes",
|
||||
)
|
||||
)
|
||||
app.state.trace_store = store
|
||||
app.state.trace_integrity = TraceIntegrityMap()
|
||||
|
||||
with socket_mod.socket() as s:
|
||||
s.bind(("127.0.0.1", 0))
|
||||
port = s.getsockname()[1]
|
||||
server = uvicorn.Server(
|
||||
uvicorn.Config(app, host="127.0.0.1", port=port, log_level="warning")
|
||||
)
|
||||
serve_task = asyncio.get_running_loop().create_task(server.serve())
|
||||
leaked = True
|
||||
try:
|
||||
for _ in range(100):
|
||||
if server.started:
|
||||
break
|
||||
await asyncio.sleep(0.1)
|
||||
assert server.started
|
||||
|
||||
import websockets
|
||||
|
||||
uri = (
|
||||
f"ws://127.0.0.1:{port}/v1/telemetry/ingest"
|
||||
f"?learner_id=L-leak&task_id=T-leak"
|
||||
)
|
||||
async with websockets.connect(uri) as ws:
|
||||
for seq in range(64): # bound=1 → guaranteed overflow
|
||||
await ws.send(_frame(seq, sandbox_id=""))
|
||||
# The flood close (1008) reaches the client.
|
||||
try:
|
||||
await asyncio.wait_for(ws.recv(), timeout=10.0)
|
||||
await asyncio.wait_for(ws.recv(), timeout=10.0)
|
||||
except (websockets.exceptions.ConnectionClosed, TimeoutError, OSError):
|
||||
pass
|
||||
assert app.state.trace_integrity.is_incomplete("L-leak", "T-leak")
|
||||
|
||||
# The session's run() must have returned: no lingering nc-* tasks
|
||||
# holding the socket open. Poll briefly — teardown is async.
|
||||
deadline = asyncio.get_running_loop().time() + 5.0
|
||||
while asyncio.get_running_loop().time() < deadline:
|
||||
names = {
|
||||
t.get_name()
|
||||
for t in asyncio.all_tasks()
|
||||
if t is not asyncio.current_task()
|
||||
}
|
||||
if not any("ingest" in n.lower() for n in names):
|
||||
leaked = False
|
||||
break
|
||||
await asyncio.sleep(0.1)
|
||||
finally:
|
||||
server.should_exit = True
|
||||
with contextlib.suppress(Exception):
|
||||
await asyncio.wait_for(serve_task, timeout=10.0)
|
||||
store.close()
|
||||
assert not leaked, "IngestSession task set leaked after queue-overflow flood"
|
||||
|
||||
|
||||
def test_reconnect_after_flood_cannot_resurrect_trace(
|
||||
client: TestClient, app, store: SQLiteTraceStore
|
||||
) -> None:
|
||||
"""INCOMPLETE_FLOODED is terminal: a fresh connection for the same trace
|
||||
is immediately closed 1008 — intake never resumes on a flagged trace."""
|
||||
with client.websocket_connect(_ingest_url()) as ws:
|
||||
for seq in range(6):
|
||||
ws.send_text(_frame(seq))
|
||||
_drain(ws)
|
||||
with pytest.raises(WebSocketDisconnect):
|
||||
while True:
|
||||
_recv_status(ws)
|
||||
assert store.latest_seq(LEARNER, TASK) == 4
|
||||
|
||||
with client.websocket_connect(_ingest_url()) as ws:
|
||||
ws.send_text(_frame(5))
|
||||
_drain(ws)
|
||||
with pytest.raises(WebSocketDisconnect) as excinfo:
|
||||
while True:
|
||||
_recv_status(ws)
|
||||
assert excinfo.value.code == 1008
|
||||
assert store.latest_seq(LEARNER, TASK) == 4 # nothing appended post-flag
|
||||
|
||||
|
||||
# -- frame validation ------------------------------------------------------------
|
||||
|
||||
|
||||
def test_frame_with_learner_id_in_body_rejected(client: TestClient) -> None:
|
||||
"""Identity lives in the URL (extra=forbid): a frame carrying learner_id
|
||||
is a contract violation → event_rejected, nothing stored."""
|
||||
with client.websocket_connect(_ingest_url()) as ws:
|
||||
bad = json.loads(_frame(0))
|
||||
bad["learner_id"] = "spoofed"
|
||||
ws.send_text(json.dumps(bad))
|
||||
_drain(ws)
|
||||
frames = [_recv_status(ws)]
|
||||
assert frames[0]["type"] == "event_rejected"
|
||||
assert client.get(f"/v1/telemetry/traces/{LEARNER}/{TASK}").status_code == 404
|
||||
|
||||
|
||||
def test_unknown_event_kind_rejected_nothing_stored(client: TestClient) -> None:
|
||||
with client.websocket_connect(_ingest_url()) as ws:
|
||||
ws.send_text(_frame(0, kind="rm_rf_everything"))
|
||||
_drain(ws)
|
||||
frames = [_recv_status(ws)]
|
||||
assert frames[0]["type"] == "event_rejected"
|
||||
assert "unknown event kind" in frames[0]["detail"]
|
||||
assert client.get(f"/v1/telemetry/traces/{LEARNER}/{TASK}").status_code == 404
|
||||
|
||||
|
||||
# -- unknown trace reads ---------------------------------------------------------
|
||||
|
||||
|
||||
def test_unknown_trace_get_returns_404(client: TestClient) -> None:
|
||||
assert client.get("/v1/telemetry/traces/nobody/never").status_code == 404
|
||||
assert client.get("/v1/telemetry/gaps/nobody/never").status_code == 404
|
||||
|
||||
|
||||
# -- handshake contract ------------------------------------------------------------
|
||||
|
||||
|
||||
def test_missing_identity_query_params_rejected_at_handshake(
|
||||
client: TestClient,
|
||||
) -> None:
|
||||
with pytest.raises(WebSocketDisconnect) as excinfo:
|
||||
with client.websocket_connect("/v1/telemetry/ingest"):
|
||||
pass
|
||||
assert excinfo.value.code == 1008
|
||||
|
||||
|
||||
def test_browser_origin_rejected_in_explicit_list_mode(
|
||||
settings, monkeypatch: pytest.MonkeyPatch
|
||||
) -> None:
|
||||
"""D-038: with an explicit AI_CORS_ORIGINS list, a page loaded in the
|
||||
learner's browser (unlisted Origin) must not be able to open the ingest
|
||||
socket and poison/flood the trace. The stdlib capture agent sends no
|
||||
Origin and is unaffected (see the no-origin test below)."""
|
||||
from fastapi.testclient import TestClient as TC
|
||||
|
||||
restricted = settings.model_copy(update={"cors_origins": "http://localhost:3000"})
|
||||
app = create_app(restricted)
|
||||
with TC(app) as c:
|
||||
with pytest.raises(WebSocketDisconnect) as excinfo:
|
||||
with c.websocket_connect(
|
||||
_ingest_url(), headers={"Origin": "https://evil.example"}
|
||||
):
|
||||
pass
|
||||
assert excinfo.value.code == 1008
|
||||
|
||||
|
||||
def test_wildcard_mode_admits_any_browser_origin(client: TestClient) -> None:
|
||||
"""D-038 default ('*'): remote-browser origins open the ingest socket —
|
||||
the remote build surface streams telemetry from the learner's browser."""
|
||||
with client.websocket_connect(
|
||||
_ingest_url(), headers={"Origin": "http://nextcraft-1:3000"}
|
||||
) as ws:
|
||||
ws.send_text(_frame(0))
|
||||
|
||||
|
||||
def test_dev_origin_and_no_origin_both_allowed(client: TestClient) -> None:
|
||||
"""The same-origin dev page (Next.js :3000) opens fine, and so does the
|
||||
capture-agent path (no Origin header at all)."""
|
||||
for headers in ({"Origin": "http://localhost:3000"}, {}):
|
||||
with client.websocket_connect(_ingest_url(), headers=headers) as ws:
|
||||
ws.send_text(_frame(0))
|
||||
body = client.get(f"/v1/telemetry/traces/{LEARNER}/{TASK}").json()
|
||||
assert [e["seq"] for e in body["events"]] == [0]
|
||||
# Unique trace per iteration would collide on (LEARNER, TASK) PK —
|
||||
# seq 0 re-sent is deduped, so one row is the invariant either way.
|
||||
assert len(body["events"]) == 1
|
||||
|
||||
|
||||
# -- keepalive ---------------------------------------------------------------------
|
||||
|
||||
|
||||
def test_server_sends_periodic_ping(
|
||||
client: TestClient, monkeypatch: pytest.MonkeyPatch
|
||||
) -> None:
|
||||
"""Keepalive (D-026): the server emits a ping on PING_INTERVAL_S so the
|
||||
stdlib capture agent's auto-pong keeps the path warm. Frame content is
|
||||
opaque to the agent; the contract is that a ping arrives while idle."""
|
||||
monkeypatch.setattr(ingest_mod, "PING_INTERVAL_S", 0.05)
|
||||
with client.websocket_connect(_ingest_url()) as ws:
|
||||
ping = ws.receive_bytes() # first server→client frame while idle
|
||||
assert ping # opaque payload; contract is "a ping frame arrives"
|
||||
@@ -0,0 +1,277 @@
|
||||
"""Variant API tests — generation, cache, distinctness over HTTP (Task 4-3-01).
|
||||
|
||||
Contract under test (api/variants.py, REQ-3-005):
|
||||
|
||||
POST /v1/variants {learner_id, template_id}
|
||||
first request → 200 generated variant (task_id, seed, params,
|
||||
statement, starter_files, competency_id)
|
||||
repeat request → 200 the SAME cached variant with ZERO LLM
|
||||
calls (D-029 reproducibility through the
|
||||
whole HTTP stack)
|
||||
unknown template → 404
|
||||
POST /v1/variants {learner_id, competency_id}
|
||||
bound competency → 200 the first template for that competency
|
||||
unknown competency → 404
|
||||
neither id given → 422
|
||||
GET /v1/variants/{task_id}
|
||||
generated earlier → 200 the stored variant (generate→get roundtrip)
|
||||
unknown task → 404
|
||||
GET /v1/variants?learner_id=...
|
||||
→ 200 that learner's variants only (scoping); [] for a learner
|
||||
with none.
|
||||
|
||||
Distinctness (REQ-3-005, API level): two learners POSTing the same
|
||||
template receive distinct statements, seeds and task_ids, and GET by
|
||||
task_id hands each back their own variant.
|
||||
|
||||
Wiring: per-test tmp-path SQLiteVariantStore + a pre-set VariantGenerator
|
||||
(state-injection override — the lifespan adopts variant_store /
|
||||
variant_generator from app.state instead of constructing them; same
|
||||
pattern as test_grading.py / test_telemetry_ingest.py). The generator
|
||||
binds a ScriptedRenderProvider so each test controls the render LLM
|
||||
exactly, and counts calls so the cache test can assert the LLM was never
|
||||
reached on the second POST.
|
||||
|
||||
Zero network: providers are MockProvider family members only (conftest
|
||||
rule, enforced in _make_client).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from collections.abc import Iterator
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from ai_service.config import Settings
|
||||
from ai_service.llm.mock import MockProvider
|
||||
from ai_service.main import create_app
|
||||
from ai_service.variants.generator import VariantGenerator
|
||||
from ai_service.variants.store import SQLiteVariantStore
|
||||
from ai_service.variants.templates import get_template
|
||||
|
||||
LEARNER_A = "variant-learner-a"
|
||||
LEARNER_B = "variant-learner-b"
|
||||
TEMPLATE = "tpl-llm-judge"
|
||||
COMPETENCY = "stack-orchestration-c007" # tpl-llm-judge's D-021 binding
|
||||
|
||||
|
||||
class ScriptedRenderProvider(MockProvider):
|
||||
"""Deterministic render whose statement embeds the params (distinct per
|
||||
draw); counts calls so the cache test asserts zero LLM calls on the
|
||||
second POST. Mirrors tests/variants/test_generator.py."""
|
||||
|
||||
def __init__(self) -> None:
|
||||
super().__init__()
|
||||
self.calls = 0
|
||||
|
||||
async def chat(self, messages, *, model, temperature=0.7, response_format=None): # noqa: ANN001
|
||||
self.calls += 1
|
||||
import json
|
||||
|
||||
user = next(m.content for m in reversed(messages) if m.role == "user")
|
||||
# Distinct per distinct params: hash the seeded slot lines.
|
||||
fingerprint = abs(hash(user)) % 10_000
|
||||
return json.dumps({"statement": f"Scripted variant #{fingerprint} — build it."})
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def store(tmp_path: Path) -> Iterator[SQLiteVariantStore]:
|
||||
s = SQLiteVariantStore(db_path=tmp_path / "variants.db")
|
||||
yield s
|
||||
s.close()
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def provider() -> ScriptedRenderProvider:
|
||||
return ScriptedRenderProvider()
|
||||
|
||||
|
||||
def _make_client(
|
||||
tmp_path: Path, store: SQLiteVariantStore, provider: MockProvider
|
||||
) -> TestClient:
|
||||
"""App + TestClient with a pre-set store + generator.
|
||||
|
||||
The lifespan adopts both (state-injection override); the generator
|
||||
binds OUR provider, so the fixture — not Settings — scripts the LLM.
|
||||
Cloud-free guard: the provider must be a MockProvider family member
|
||||
(conftest rule, enforced here because this module builds its own
|
||||
client rather than consuming the conftest one).
|
||||
"""
|
||||
assert isinstance(provider, MockProvider)
|
||||
settings = Settings(
|
||||
provider="mock",
|
||||
db_path=tmp_path / "variant-test.db",
|
||||
sandbox_dir=tmp_path / "sandboxes",
|
||||
)
|
||||
app = create_app(settings)
|
||||
app.state.variant_store = store
|
||||
if getattr(app.state, "variant_generator", None) is None:
|
||||
app.state.variant_generator = VariantGenerator(
|
||||
store, provider, model="gemma4:31b"
|
||||
)
|
||||
return TestClient(app)
|
||||
|
||||
|
||||
def _post(
|
||||
client: TestClient,
|
||||
learner_id: str,
|
||||
template_id: str | None = TEMPLATE,
|
||||
competency_id: str | None = None,
|
||||
):
|
||||
return client.post(
|
||||
"/v1/variants",
|
||||
json={
|
||||
"learner_id": learner_id,
|
||||
**({"template_id": template_id} if template_id is not None else {}),
|
||||
**({"competency_id": competency_id} if competency_id is not None else {}),
|
||||
},
|
||||
)
|
||||
|
||||
|
||||
# -- POST: generation + response shape --------------------------------------------
|
||||
|
||||
|
||||
class TestGenerate:
|
||||
def test_post_returns_full_variant_payload(self, tmp_path, store, provider):
|
||||
with _make_client(tmp_path, store, provider) as client:
|
||||
response = _post(client, LEARNER_A)
|
||||
|
||||
assert response.status_code == 200
|
||||
body = response.json()
|
||||
assert body["learner_id"] == LEARNER_A
|
||||
assert body["template_id"] == TEMPLATE
|
||||
assert body["competency_id"] == COMPETENCY # enrichment from get_template
|
||||
assert body["task_id"].startswith("task-") and len(body["task_id"]) == len("task-") + 16
|
||||
assert body["seed"] # non-empty D-029 seed
|
||||
assert body["statement"] # rendered, non-empty
|
||||
assert body["starter_files"] == get_template(TEMPLATE).starter_files
|
||||
assert body["params"] # seeded slot values
|
||||
assert body["created_at"] # ISO timestamp travels
|
||||
assert provider.calls == 1 # first POST renders exactly once
|
||||
|
||||
def test_post_then_get_roundtrip(self, tmp_path, store, provider):
|
||||
with _make_client(tmp_path, store, provider) as client:
|
||||
posted = _post(client, LEARNER_A)
|
||||
assert posted.status_code == 200
|
||||
|
||||
fetched = client.get(f"/v1/variants/{posted.json()['task_id']}")
|
||||
|
||||
assert fetched.status_code == 200
|
||||
assert fetched.json() == posted.json() # identical stored variant
|
||||
|
||||
def test_post_regenerate_is_cache_hit_no_llm_call(self, tmp_path, store, provider):
|
||||
with _make_client(tmp_path, store, provider) as client:
|
||||
first = _post(client, LEARNER_A)
|
||||
assert first.status_code == 200
|
||||
assert provider.calls == 1
|
||||
|
||||
second = _post(client, LEARNER_A) # same (learner, template) → cache
|
||||
assert second.status_code == 200
|
||||
assert second.json() == first.json() # D-029: the SAME variant
|
||||
assert provider.calls == 1, "cached regenerate must not render again"
|
||||
|
||||
|
||||
# -- POST: distinctness (REQ-3-005, API level) ------------------------------------
|
||||
|
||||
|
||||
class TestDistinctLearners:
|
||||
def test_two_learners_distinct_statements_and_task_ids(
|
||||
self, tmp_path, store, provider
|
||||
):
|
||||
with _make_client(tmp_path, store, provider) as client:
|
||||
a = _post(client, LEARNER_A)
|
||||
b = _post(client, LEARNER_B)
|
||||
assert a.status_code == 200 and b.status_code == 200
|
||||
|
||||
a_body, b_body = a.json(), b.json()
|
||||
|
||||
# ...and each learner GETs back exactly their own variant
|
||||
a_fetched = client.get(f"/v1/variants/{a_body['task_id']}")
|
||||
b_fetched = client.get(f"/v1/variants/{b_body['task_id']}")
|
||||
|
||||
assert a_body["statement"] != b_body["statement"]
|
||||
assert a_body["seed"] != b_body["seed"]
|
||||
assert a_body["task_id"] != b_body["task_id"]
|
||||
assert a_body["competency_id"] == b_body["competency_id"] # same bar (a-5)
|
||||
assert a_fetched.status_code == 200 and a_fetched.json() == a_body
|
||||
assert b_fetched.status_code == 200 and b_fetched.json() == b_body
|
||||
|
||||
|
||||
# -- POST: template / competency resolution ---------------------------------------
|
||||
|
||||
|
||||
class TestTemplateResolution:
|
||||
def test_unknown_template_404(self, tmp_path, store, provider):
|
||||
with _make_client(tmp_path, store, provider) as client:
|
||||
response = _post(client, LEARNER_A, template_id="tpl-does-not-exist")
|
||||
|
||||
assert response.status_code == 404
|
||||
assert "no task template" in response.json()["detail"]
|
||||
|
||||
def test_competency_lookup_uses_first_bound_template(self, tmp_path, store, provider):
|
||||
with _make_client(tmp_path, store, provider) as client:
|
||||
response = _post(
|
||||
client, LEARNER_A, template_id=None, competency_id=COMPETENCY
|
||||
)
|
||||
|
||||
assert response.status_code == 200
|
||||
body = response.json()
|
||||
assert body["template_id"] == TEMPLATE # the competency's first template
|
||||
assert body["competency_id"] == COMPETENCY
|
||||
|
||||
def test_unknown_competency_404(self, tmp_path, store, provider):
|
||||
with _make_client(tmp_path, store, provider) as client:
|
||||
response = _post(
|
||||
client, LEARNER_A, template_id=None, competency_id="stack-none-c999"
|
||||
)
|
||||
|
||||
assert response.status_code == 404
|
||||
assert "no task template for competency" in response.json()["detail"]
|
||||
|
||||
def test_neither_id_given_422(self, tmp_path, store, provider):
|
||||
with _make_client(tmp_path, store, provider) as client:
|
||||
response = _post(client, LEARNER_A, template_id=None, competency_id=None)
|
||||
|
||||
assert response.status_code == 422
|
||||
|
||||
|
||||
# -- GET: stored reads ------------------------------------------------------------
|
||||
|
||||
|
||||
class TestStoredReads:
|
||||
def test_get_unknown_task_404(self, tmp_path, store, provider):
|
||||
with _make_client(tmp_path, store, provider) as client:
|
||||
response = client.get("/v1/variants/task-0000000000000000")
|
||||
|
||||
assert response.status_code == 404
|
||||
assert "no stored variant" in response.json()["detail"]
|
||||
|
||||
def test_list_scoped_to_one_learner(self, tmp_path, store, provider):
|
||||
with _make_client(tmp_path, store, provider) as client:
|
||||
assert _post(client, LEARNER_A).status_code == 200
|
||||
assert _post(
|
||||
client, LEARNER_A, template_id="tpl-guardrail-schema"
|
||||
).status_code == 200
|
||||
assert _post(client, LEARNER_B).status_code == 200
|
||||
|
||||
listed = client.get("/v1/variants", params={"learner_id": LEARNER_A})
|
||||
empty = client.get("/v1/variants", params={"learner_id": "nobody"})
|
||||
|
||||
assert listed.status_code == 200
|
||||
variants = listed.json()["variants"]
|
||||
assert len(variants) == 2 # LEARNER_A's two, LEARNER_B's excluded
|
||||
assert {v["learner_id"] for v in variants} == {LEARNER_A}
|
||||
assert {v["template_id"] for v in variants} == {TEMPLATE, "tpl-guardrail-schema"}
|
||||
# chronological ordering; both carry the competency enrichment
|
||||
assert all(v["competency_id"] for v in variants)
|
||||
|
||||
assert empty.status_code == 200
|
||||
assert empty.json()["variants"] == []
|
||||
|
||||
def test_list_missing_learner_param_422(self, tmp_path, store, provider):
|
||||
with _make_client(tmp_path, store, provider) as client:
|
||||
response = client.get("/v1/variants")
|
||||
|
||||
assert response.status_code == 422
|
||||
@@ -0,0 +1 @@
|
||||
"""Process-trace grading tests (REQ-3-004)."""
|
||||
@@ -0,0 +1,361 @@
|
||||
"""Grading calibration CONTRACT test (Task 3-2-02, REQ-3-004, D-021).
|
||||
|
||||
HONEST SCOPE — what this test is and is NOT:
|
||||
|
||||
This is a CALIBRATION CONTRACT test, not an LLM quality test. The mock
|
||||
provider scripts archetype-mapped rubric scores for each fixture; the
|
||||
assertions verify that:
|
||||
|
||||
1. the ENGINE pipeline (gate → digest → prompt → D-020 → validate →
|
||||
persist) runs each archetype end-to-end and preserves the score
|
||||
ordering the scripts impose;
|
||||
2. the fixtures' digest FEATURES actually separate the archetypes
|
||||
deterministically (the non-LLM half of the calibration: paste-and-run
|
||||
vs iterative debugging vs never-passing MUST produce observably
|
||||
different digests — otherwise no rubric could ever separate them);
|
||||
3. fixture IDs are valid corpus-aligned strings (D-021).
|
||||
|
||||
It does NOT prove the real LLM scores archetypes in this order — that
|
||||
requires live-model evaluation, which is out of scope for v0.3 tests
|
||||
(tests NEVER call the cloud; conftest enforces). What the contract pins
|
||||
is the ORDERING the rubric + digest must eventually enforce:
|
||||
|
||||
strong.process_quality >= lazy.process_quality
|
||||
strong.correctness > struggling.correctness
|
||||
|
||||
If a future prompt/digest change makes the scripted mapping unachievable
|
||||
by ANY model (e.g. the digest stops separating the archetypes), test 2
|
||||
fails — that is the calibration signal.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
from ai_service.corpus.telemetry import LAB_SCENARIOS
|
||||
from ai_service.corpus.trace_fixtures import (
|
||||
CALIBRATION_LEARNER_ID,
|
||||
LAZY_BUILDER,
|
||||
STRONG_BUILDER,
|
||||
STRUGGLING_BUILDER,
|
||||
TRACE_FIXTURES,
|
||||
digest_of,
|
||||
)
|
||||
from ai_service.grading.engine import GradingEngine
|
||||
from ai_service.grading.store import SQLiteGradeStore
|
||||
from ai_service.telemetry.ingest import TraceIntegrityMap
|
||||
from ai_service.telemetry.store import SQLiteTraceStore
|
||||
|
||||
#: Archetype-mapped rubric scripts (the ORDERING CONTRACT in numbers).
|
||||
#: strong: iterative verified building — highest process quality.
|
||||
#: lazy: ends pass with zero iteration — the a-4 negative (churn with no
|
||||
#: test progress caps process_quality at 1; here it's paste-churn +
|
||||
#: one late test, so 1). correctness lands high (final pass, fast).
|
||||
#: struggling: never passes — lowest correctness.
|
||||
ARCHETYPE_SCORES: dict[str, dict] = {
|
||||
"strong": {
|
||||
"criteria": {
|
||||
"process_quality": 4,
|
||||
"correctness": 4,
|
||||
"debugging_discipline": 4,
|
||||
"test_usage": 4,
|
||||
},
|
||||
"strengths": ["tight edit-test loops; every failure cycle closed"],
|
||||
"gaps": ["session ended without a final full-suite run"],
|
||||
"verdict": "mastered",
|
||||
},
|
||||
"lazy": {
|
||||
"criteria": {
|
||||
"process_quality": 1, # a-4: bulk paste + churn, one late test
|
||||
"correctness": 3,
|
||||
"debugging_discipline": 1,
|
||||
"test_usage": 1,
|
||||
},
|
||||
"strengths": ["final tests pass on the first run"],
|
||||
"gaps": ["one bulk paste with zero verified increments"],
|
||||
"verdict": "developing",
|
||||
},
|
||||
"struggling": {
|
||||
"criteria": {
|
||||
"process_quality": 2,
|
||||
"correctness": 0, # never passes
|
||||
"debugging_discipline": 1, # cycles never close
|
||||
"test_usage": 3,
|
||||
},
|
||||
"strengths": ["persisted through repeated failures"],
|
||||
"gaps": ["the same failure recurred three times"],
|
||||
"verdict": "not_yet",
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
class ArchetypeProvider:
|
||||
"""Mock provider scripting archetype-mapped scores keyed by fixture.
|
||||
|
||||
Keying on the digest embedded in the user turn (not on call order)
|
||||
keeps the mapping robust to gate-path noise and asserts the prompt
|
||||
really carries the fixture's digest. Records the last user turn so
|
||||
tests can re-assert D-028 (digest in prompt; raw events out).
|
||||
"""
|
||||
|
||||
def __init__(self) -> None:
|
||||
self.digest_to_archetype: dict[str, str] = {}
|
||||
self.calls = 0
|
||||
self.last_user_prompt: str = ""
|
||||
|
||||
def register(self, fixture) -> None:
|
||||
self.digest_to_archetype[digest_of(fixture).model_dump_json()] = (
|
||||
fixture.archetype
|
||||
)
|
||||
|
||||
async def chat(self, messages, *, model, temperature=0.7, response_format=None):
|
||||
self.calls += 1
|
||||
user = next((m for m in reversed(messages) if m.role == "user"), None)
|
||||
assert user is not None
|
||||
self.last_user_prompt = user.content
|
||||
archetype = None
|
||||
for digest_json, arch in self.digest_to_archetype.items():
|
||||
if digest_json in user.content:
|
||||
archetype = arch
|
||||
break
|
||||
assert archetype is not None, "fixture digest not found in prompt"
|
||||
return json.dumps(ARCHETYPE_SCORES[archetype])
|
||||
|
||||
async def stream_chat(self, messages, *, model, temperature=0.7, response_format=None):
|
||||
yield await self.chat(
|
||||
messages, model=model, temperature=temperature, response_format=response_format
|
||||
)
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def trace_store(tmp_path: Path) -> SQLiteTraceStore:
|
||||
store = SQLiteTraceStore(db_path=tmp_path / "traces.db")
|
||||
yield store
|
||||
store.close()
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def grade_store(tmp_path: Path) -> SQLiteGradeStore:
|
||||
store = SQLiteGradeStore(db_path=tmp_path / "grades.db")
|
||||
yield store
|
||||
store.close()
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def provider() -> ArchetypeProvider:
|
||||
p = ArchetypeProvider()
|
||||
for fixture in (STRONG_BUILDER, LAZY_BUILDER, STRUGGLING_BUILDER):
|
||||
p.register(fixture)
|
||||
return p
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def engine(trace_store, grade_store, provider) -> GradingEngine:
|
||||
return GradingEngine(
|
||||
trace_store, grade_store, TraceIntegrityMap(), provider, model="gemma4:31b"
|
||||
)
|
||||
|
||||
|
||||
async def _grade_fixture(engine, trace_store, fixture) -> dict:
|
||||
for event in fixture.events:
|
||||
trace_store.append(event)
|
||||
record = await engine.grade(fixture.learner_id, fixture.task_id)
|
||||
assert record.verdict == "GRADED", (
|
||||
f"fixture {fixture.fixture_id} did not grade cleanly: {record.verdict}"
|
||||
)
|
||||
return record.scores
|
||||
|
||||
|
||||
class TestFixtureAlignment:
|
||||
"""D-021: fixture IDs align with the v0.2 corpus scenario IDs."""
|
||||
|
||||
def test_every_fixture_id_is_corpus_scenario_scoped(self):
|
||||
for fixture in TRACE_FIXTURES.values():
|
||||
assert fixture.fixture_id.count("::") == 1
|
||||
scenario_part, trace_part = fixture.fixture_id.split("::")
|
||||
assert scenario_part in LAB_SCENARIOS, (
|
||||
f"{scenario_part!r} is not a v0.2 lab scenario id"
|
||||
)
|
||||
assert trace_part.endswith("-trace")
|
||||
# double-brace-free single slug
|
||||
assert "-" in trace_part
|
||||
|
||||
def test_every_fixture_aligns_to_an_existing_scenario(self):
|
||||
assert set(f.aligned_scenario_id for f in TRACE_FIXTURES.values()) <= set(
|
||||
LAB_SCENARIOS
|
||||
)
|
||||
|
||||
def test_fixture_ids_are_valid_strings(self):
|
||||
for fixture in TRACE_FIXTURES.values():
|
||||
assert fixture.fixture_id and isinstance(fixture.fixture_id, str)
|
||||
assert fixture.learner_id == CALIBRATION_LEARNER_ID
|
||||
assert fixture.learner_id.startswith("learner-00")
|
||||
|
||||
def test_lazy_alignment_is_the_flagged_scenario(self):
|
||||
# v0.2's paste-and-run archetype is the flagged lab scenario (D-021).
|
||||
assert LAZY_BUILDER.aligned_scenario_id == "lab-scenario-flagged"
|
||||
assert LAZY_BUILDER.archetype == "lazy"
|
||||
|
||||
def test_all_three_archetypes_present(self):
|
||||
assert {f.archetype for f in TRACE_FIXTURES.values()} == {
|
||||
"strong",
|
||||
"lazy",
|
||||
"struggling",
|
||||
}
|
||||
|
||||
def test_events_are_valid_telemetry_events(self):
|
||||
from ai_service.telemetry.models import TelemetryEvent
|
||||
|
||||
for fixture in TRACE_FIXTURES.values():
|
||||
assert len(fixture.events) > 0
|
||||
for e in fixture.events:
|
||||
assert isinstance(e, TelemetryEvent)
|
||||
assert e.seq >= 0
|
||||
# contiguous seqs: a grading trace must clear the G-4 gate
|
||||
seqs = [e.seq for e in fixture.events]
|
||||
assert seqs == list(range(len(seqs))), (
|
||||
f"{fixture.fixture_id} has seq gaps — it would trip the gate"
|
||||
)
|
||||
|
||||
|
||||
class TestDigestSeparatesArchetypes:
|
||||
"""The non-LLM half of calibration: digests differ in the RIGHT direction.
|
||||
|
||||
If these deterministic assertions ever fail, no rubric — however good —
|
||||
could separate the archetypes: the digest would be lossy in the wrong
|
||||
places. These are the features the a-4 advisory and the rubric anchors
|
||||
speak about, computed over REAL trace fixtures.
|
||||
"""
|
||||
|
||||
def test_strong_shows_iterative_cycles_lazy_shows_none(self):
|
||||
strong = digest_of(STRONG_BUILDER)
|
||||
lazy = digest_of(LAZY_BUILDER)
|
||||
assert strong.error_fix_cycles >= 1
|
||||
assert lazy.error_fix_cycles == 0
|
||||
|
||||
def test_strong_tests_early_lazy_tests_late(self):
|
||||
strong = digest_of(STRONG_BUILDER)
|
||||
lazy = digest_of(LAZY_BUILDER)
|
||||
assert strong.first_test_pass_offset_s is not None
|
||||
assert lazy.first_test_pass_offset_s is not None
|
||||
assert strong.first_test_pass_offset_s < lazy.first_test_pass_offset_s
|
||||
|
||||
def test_struggling_never_passes(self):
|
||||
struggling = digest_of(STRUGGLING_BUILDER)
|
||||
assert struggling.final_test_status == "fail"
|
||||
assert struggling.test_pass_count == 0
|
||||
assert struggling.test_fail_count >= 3
|
||||
|
||||
def test_strong_and_struggling_both_end(self):
|
||||
strong = digest_of(STRONG_BUILDER)
|
||||
assert strong.final_test_status == "pass"
|
||||
assert strong.test_fail_count >= 1 # honest failure, then closed
|
||||
|
||||
def test_lazy_churns_without_test_progress(self):
|
||||
"""a-4 anchor: bulk edit + command churn with (almost) no test signal."""
|
||||
lazy = digest_of(LAZY_BUILDER)
|
||||
assert lazy.edit_count >= 2
|
||||
assert lazy.test_fail_count == 0
|
||||
assert lazy.command_count >= 2
|
||||
assert lazy.final_test_status == "pass" # ...and still one late test
|
||||
|
||||
|
||||
class TestOrderingContract:
|
||||
"""The scripted ORDERING CONTRACT through the full engine pipeline."""
|
||||
|
||||
async def test_strong_at_least_lazy_on_process_quality(
|
||||
self, engine, trace_store
|
||||
):
|
||||
strong = await _grade_fixture(engine, trace_store, STRONG_BUILDER)
|
||||
lazy = await _grade_fixture(engine, trace_store, LAZY_BUILDER)
|
||||
assert (
|
||||
strong["criteria"]["process_quality"]
|
||||
>= lazy["criteria"]["process_quality"]
|
||||
)
|
||||
|
||||
async def test_strong_above_struggling_on_correctness(self, engine, trace_store):
|
||||
strong = await _grade_fixture(engine, trace_store, STRONG_BUILDER)
|
||||
struggling = await _grade_fixture(engine, trace_store, STRUGGLING_BUILDER)
|
||||
assert strong["criteria"]["correctness"] > struggling["criteria"]["correctness"]
|
||||
|
||||
async def test_full_ordering_across_all_three_archetypes(
|
||||
self, engine, trace_store, grade_store
|
||||
):
|
||||
scores = {
|
||||
arch: await _grade_fixture(engine, trace_store, f)
|
||||
for arch, f in (
|
||||
("strong", STRONG_BUILDER),
|
||||
("lazy", LAZY_BUILDER),
|
||||
("struggling", STRUGGLING_BUILDER),
|
||||
)
|
||||
}
|
||||
# the required ordering contract (task 3-2-02):
|
||||
assert (
|
||||
scores["strong"]["criteria"]["process_quality"]
|
||||
>= scores["lazy"]["criteria"]["process_quality"]
|
||||
)
|
||||
assert (
|
||||
scores["strong"]["criteria"]["correctness"]
|
||||
> scores["struggling"]["criteria"]["correctness"]
|
||||
)
|
||||
# Deliberate NON-ordering, asserted to document the calibration
|
||||
# intent: struggling > lazy on process quality. Process quality is
|
||||
# not outcome — the struggling builder iterated (edits between
|
||||
# runs, three test runs), the lazy builder only pasted once and
|
||||
# verified once (a-4 caps that at 1). A rubric that scores a
|
||||
# never-passing iterator BELOW a paste-and-runner on process is
|
||||
# miscalibrated on its own anchors.
|
||||
assert (
|
||||
scores["struggling"]["criteria"]["process_quality"]
|
||||
> scores["lazy"]["criteria"]["process_quality"]
|
||||
)
|
||||
# verdicts also ordered: mastered >= developing > not_yet
|
||||
assert scores["strong"]["verdict"] == "mastered"
|
||||
assert scores["lazy"]["verdict"] == "developing"
|
||||
assert scores["struggling"]["verdict"] == "not_yet"
|
||||
|
||||
# all three persisted under the calibration learner, latest-state
|
||||
grades = grade_store.list_for_learner(CALIBRATION_LEARNER_ID)
|
||||
assert len(grades) == 3
|
||||
assert {g.task_id for g in grades} == {
|
||||
"task-calibration-strong",
|
||||
"task-calibration-lazy",
|
||||
"task-calibration-struggling",
|
||||
}
|
||||
|
||||
async def test_prompt_carries_fixture_digest_not_raw_events(
|
||||
self, engine, trace_store, provider
|
||||
):
|
||||
"""D-028 on the calibration path: digest in the prompt, raw events
|
||||
never — the lazy fixture plants distinctive payload strings
|
||||
("eval.py", the bulk 240-line paste) that must NOT appear."""
|
||||
for event in LAZY_BUILDER.events:
|
||||
trace_store.append(event)
|
||||
await engine.grade(LAZY_BUILDER.learner_id, LAZY_BUILDER.task_id)
|
||||
|
||||
assert provider.calls == 1
|
||||
prompt = provider.last_user_prompt
|
||||
# the fixture's digest reached the LLM...
|
||||
assert "PROCESS TRACE DIGEST (JSON):" in prompt
|
||||
assert '"error_fix_cycles"' in prompt
|
||||
# ...but no raw trace material or identity did (D-028)
|
||||
assert "eval.py" not in prompt
|
||||
assert LAZY_BUILDER.learner_id not in prompt
|
||||
assert LAZY_BUILDER.task_id not in prompt
|
||||
|
||||
async def test_struggling_digests_reach_engine_without_gate(
|
||||
self, engine, trace_store, provider
|
||||
):
|
||||
for event in STRUGGLING_BUILDER.events:
|
||||
trace_store.append(event)
|
||||
record = await engine.grade(
|
||||
STRUGGLING_BUILDER.learner_id, STRUGGLING_BUILDER.task_id
|
||||
)
|
||||
assert record.verdict == "GRADED" # honest struggle is gradable
|
||||
assert record.scores["criteria"]["correctness"] == 0
|
||||
# a never-passing session still yields a real verdict — the grade
|
||||
# is low, not missing
|
||||
assert record.scores["verdict"] == "not_yet"
|
||||
@@ -0,0 +1,528 @@
|
||||
"""GradingEngine tests (Task 3-2-01, REQ-3-004, G-4 binding).
|
||||
|
||||
Contract under test:
|
||||
- G-4 gate FIRST: seq gaps OR integrity-flagged trace →
|
||||
verdict=UNGRADABLE_TRACE_INCOMPLETE with the gap list surfaced in
|
||||
scores; the LLM is NEVER called on those paths (asserted). Empty
|
||||
trace → UNGRADABLE_EMPTY_TRACE. All are first-class GradeRecords
|
||||
(returned + persisted), not exceptions.
|
||||
- Complete trace: prompt carries the digest JSON but NO raw trace
|
||||
material (D-028) — a distinctive marker planted in a command payload
|
||||
must be absent from every message the provider received.
|
||||
- Malformed LLM JSON → the D-020 retry path recovers (scripted bad
|
||||
first, good second) → graded.
|
||||
- Re-grade overwrites the stored record (GradeStore upsert).
|
||||
|
||||
Zero network: everything runs against the mock provider (conftest rule).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from datetime import UTC, datetime, timedelta
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
from pydantic import BaseModel, ValidationError
|
||||
|
||||
from ai_service.grading.engine import GradingEngine, RubricScore
|
||||
from ai_service.grading.store import SQLiteGradeStore
|
||||
from ai_service.llm.types import Message
|
||||
from ai_service.telemetry.ingest import TraceIntegrityMap
|
||||
from ai_service.telemetry.models import TelemetryEvent
|
||||
from ai_service.telemetry.store import SQLiteTraceStore
|
||||
from ai_service.variants.store import SQLiteVariantStore
|
||||
|
||||
T0 = datetime(2026, 9, 12, 1, 0, 0, tzinfo=UTC)
|
||||
RAW_MARKER = "SECRET-COMMAND-MARKER-7f3a"
|
||||
|
||||
|
||||
class RubricJSON:
|
||||
"""Canonical well-formed rubric payload (matches engine.RubricScore)."""
|
||||
|
||||
PAYLOAD: dict = {
|
||||
"criteria": {
|
||||
"process_quality": 3,
|
||||
"correctness": 4,
|
||||
"debugging_discipline": 3,
|
||||
"test_usage": 4,
|
||||
},
|
||||
"strengths": ["tight edit-test loops throughout"],
|
||||
"gaps": ["final commit discipline loose"],
|
||||
"verdict": "mastered",
|
||||
}
|
||||
|
||||
|
||||
class RecordingProvider:
|
||||
"""Mock LLMProvider: scripted replies + captured requests (no network).
|
||||
|
||||
Subclass-composes ai_service.llm.mock.MockProvider so the conftest
|
||||
cloud-free guard philosophy holds — but records every message list
|
||||
and counts calls, which the stock mocks don't do.
|
||||
"""
|
||||
|
||||
def __init__(self) -> None:
|
||||
from ai_service.llm.mock import MockProvider
|
||||
|
||||
self._mock = MockProvider()
|
||||
self.replies: list[str] = [] # consumed left-to-right; last repeats
|
||||
self.requests: list[list[Message]] = []
|
||||
self.calls = 0
|
||||
|
||||
async def chat(self, messages, *, model, temperature=0.7, response_format=None):
|
||||
self.calls += 1
|
||||
self.requests.append([Message(role=m.role, content=m.content) for m in messages])
|
||||
if self.replies:
|
||||
reply = self.replies.pop(0)
|
||||
else:
|
||||
reply = json.dumps(RubricJSON.PAYLOAD)
|
||||
return reply
|
||||
|
||||
async def stream_chat(self, messages, *, model, temperature=0.7, response_format=None):
|
||||
yield await self.chat(
|
||||
messages, model=model, temperature=temperature, response_format=response_format
|
||||
)
|
||||
|
||||
|
||||
def _event(
|
||||
seq: int,
|
||||
kind: str,
|
||||
payload: dict | None = None,
|
||||
offset_s: float = 0.0,
|
||||
*,
|
||||
learner: str = "engine-learner",
|
||||
task: str = "engine-task",
|
||||
) -> TelemetryEvent:
|
||||
return TelemetryEvent(
|
||||
learner_id=learner,
|
||||
task_id=task,
|
||||
seq=seq,
|
||||
kind=kind,
|
||||
payload=payload or {},
|
||||
ts=T0 + timedelta(seconds=offset_s),
|
||||
sandbox_id="sbx-engine",
|
||||
)
|
||||
|
||||
|
||||
def _complete_trace(marker_command: bool = True) -> list[TelemetryEvent]:
|
||||
"""A healthy iterative trace; seq 0..N contiguous (no gaps).
|
||||
|
||||
`marker_command` plants a distinctive string inside a command payload
|
||||
— the D-028 assertion target: raw trace material must never reach the
|
||||
provider, so the marker must not appear in any captured request.
|
||||
"""
|
||||
if marker_command:
|
||||
cmd_payload = {"cmd": f"echo {RAW_MARKER} && pytest -q"}
|
||||
else:
|
||||
cmd_payload = {"cmd": "pytest -q"}
|
||||
return [
|
||||
_event(0, "activity", {"state": "starting"}, 0.0),
|
||||
_event(1, "file_diff", {"path": "a.py", "added": 12}, 10.0),
|
||||
_event(2, "command", cmd_payload, 20.0),
|
||||
_event(3, "test_result", {"passed": False, "exit_code": 1}, 25.0),
|
||||
_event(4, "file_diff", {"path": "a.py", "added": 4, "removed": 2}, 40.0),
|
||||
_event(5, "command", {"cmd": "pytest -q"}, 60.0),
|
||||
_event(6, "test_result", {"passed": True, "exit_code": 0}, 65.0),
|
||||
]
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def trace_store(tmp_path: Path) -> SQLiteTraceStore:
|
||||
store = SQLiteTraceStore(db_path=tmp_path / "traces.db")
|
||||
yield store
|
||||
store.close()
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def grade_store(tmp_path: Path) -> SQLiteGradeStore:
|
||||
store = SQLiteGradeStore(db_path=tmp_path / "grades.db")
|
||||
yield store
|
||||
store.close()
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def integrity() -> TraceIntegrityMap:
|
||||
return TraceIntegrityMap()
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def provider() -> RecordingProvider:
|
||||
return RecordingProvider()
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def engine(trace_store, grade_store, integrity, provider) -> GradingEngine:
|
||||
return GradingEngine(
|
||||
trace_store, grade_store, integrity, provider, model="gemma4:31b"
|
||||
)
|
||||
|
||||
|
||||
def _ingest(store: SQLiteTraceStore, events: list[TelemetryEvent]) -> None:
|
||||
for e in events:
|
||||
store.append(e)
|
||||
|
||||
|
||||
class TestRubricScoreModel:
|
||||
"""The D-020 schema contract — engine-side validation rules."""
|
||||
|
||||
def _rubric(self, **overrides) -> dict:
|
||||
payload = json.loads(json.dumps(RubricJSON.PAYLOAD)) # deep copy
|
||||
payload.update(overrides)
|
||||
return payload
|
||||
|
||||
def test_well_formed_payload_validates(self):
|
||||
score = RubricScore.model_validate(RubricJSON.PAYLOAD)
|
||||
assert score.criteria["process_quality"] == 3
|
||||
assert score.verdict == "mastered"
|
||||
|
||||
def test_unknown_criterion_rejected(self):
|
||||
bad = self._rubric()
|
||||
bad["criteria"]["extra_criterion"] = 2
|
||||
with pytest.raises(ValidationError, match="criteria keys must be exactly"):
|
||||
RubricScore.model_validate(bad)
|
||||
|
||||
def test_missing_criterion_rejected(self):
|
||||
bad = self._rubric()
|
||||
del bad["criteria"]["test_usage"]
|
||||
with pytest.raises(ValidationError, match="criteria keys must be exactly"):
|
||||
RubricScore.model_validate(bad)
|
||||
|
||||
def test_out_of_range_score_rejected(self):
|
||||
bad = self._rubric()
|
||||
bad["criteria"]["process_quality"] = 5
|
||||
with pytest.raises(ValidationError, match="must be within 0-4"):
|
||||
RubricScore.model_validate(bad)
|
||||
|
||||
def test_unknown_verdict_rejected(self):
|
||||
bad = self._rubric()
|
||||
bad["verdict"] = "excellent"
|
||||
with pytest.raises(ValidationError, match="verdict must be one of"):
|
||||
RubricScore.model_validate(bad)
|
||||
|
||||
|
||||
class TestCompleteTraceGrading:
|
||||
async def test_graded_returns_validated_rubric_persisted_and_returned(
|
||||
self, engine, trace_store, grade_store, provider
|
||||
):
|
||||
_ingest(trace_store, _complete_trace())
|
||||
|
||||
record = await engine.grade("engine-learner", "engine-task")
|
||||
|
||||
assert record.verdict == "GRADED"
|
||||
assert record.model == "gemma4:31b"
|
||||
assert record.scores == RubricJSON.PAYLOAD
|
||||
assert record.scores["criteria"]["process_quality"] == 3
|
||||
assert record.digest.get("error_fix_cycles") == 1
|
||||
assert record.digest.get("final_test_status") == "pass"
|
||||
# digest persisted for auditability (D-028 reproducible input)
|
||||
assert record.digest.get("event_count") == 7
|
||||
|
||||
stored = grade_store.get("engine-learner", "engine-task")
|
||||
assert stored is not None
|
||||
assert stored.scores == record.scores
|
||||
assert stored.verdict == "GRADED"
|
||||
|
||||
async def test_prompt_contains_digest_but_no_raw_trace(self, engine, trace_store, provider):
|
||||
"""D-028: the provider saw the digest JSON but NEVER the raw trace.
|
||||
|
||||
The marker was planted inside a command payload (seq 2); it must be
|
||||
absent from every message of every request, while the digest marker
|
||||
line + a JSON object with the digest's field names must be present.
|
||||
"""
|
||||
_ingest(trace_store, _complete_trace(marker_command=True))
|
||||
|
||||
await engine.grade("engine-learner", "engine-task")
|
||||
|
||||
assert provider.calls == 1 # happy path: exactly one LLM call
|
||||
all_messages = [m for req in provider.requests for m in req]
|
||||
prompt_text = "\n".join(m.content for m in all_messages)
|
||||
|
||||
assert RAW_MARKER not in prompt_text, (
|
||||
"raw trace material reached the LLM prompt (D-028 violation)"
|
||||
)
|
||||
# digest presence: marker line + a JSON object carrying digest fields
|
||||
assert "PROCESS TRACE DIGEST (JSON):" in prompt_text
|
||||
assert '"error_fix_cycles"' in prompt_text
|
||||
assert '"final_test_status"' in prompt_text
|
||||
# learner anonymity: identity strings never reach the LLM either
|
||||
assert "engine-learner" not in prompt_text
|
||||
assert "engine-task" not in prompt_text
|
||||
|
||||
async def test_system_prompt_has_rubric_and_a4_advisory(
|
||||
self, engine, trace_store, provider
|
||||
):
|
||||
_ingest(trace_store, _complete_trace())
|
||||
|
||||
await engine.grade("engine-learner", "engine-task")
|
||||
|
||||
system = provider.requests[0][0]
|
||||
assert system.role == "system"
|
||||
assert "process_quality" in system.content
|
||||
assert "correctness" in system.content
|
||||
assert "debugging_discipline" in system.content
|
||||
assert "test_usage" in system.content
|
||||
# a-4: churn-without-test-progress is a process-quality NEGATIVE
|
||||
assert "churn is not" in system.content
|
||||
# 0-4 level anchors present for each criterion
|
||||
assert "ADVISORY" in system.content
|
||||
|
||||
|
||||
class TestGateFirst:
|
||||
"""G-4 binding: the gate fires BEFORE any grading work."""
|
||||
|
||||
async def test_gapped_trace_ungradable_llm_not_called(
|
||||
self, engine, trace_store, provider, grade_store
|
||||
):
|
||||
# seqs 0,1,3 stored — seq 2 missing (mirrors the task spec example)
|
||||
events = _complete_trace()
|
||||
gapped = [e for e in events if e.seq != 2]
|
||||
_ingest(trace_store, gapped)
|
||||
|
||||
record = await engine.grade("engine-learner", "engine-task")
|
||||
|
||||
assert record.verdict == "UNGRADABLE_TRACE_INCOMPLETE"
|
||||
assert provider.calls == 0, "LLM was called despite a gapped trace (G-4)"
|
||||
# gap list surfaced in scores
|
||||
assert record.scores["missing_seqs"] == [2]
|
||||
assert record.scores["integrity_flag"] is None
|
||||
# first-class record: persisted, retrievable, empty digest
|
||||
stored = grade_store.get("engine-learner", "engine-task")
|
||||
assert stored is not None
|
||||
assert stored.verdict == "UNGRADABLE_TRACE_INCOMPLETE"
|
||||
assert stored.scores["missing_seqs"] == [2]
|
||||
assert stored.digest == {}
|
||||
assert stored.model == "none"
|
||||
|
||||
async def test_integrity_flagged_trace_ungradable_llm_not_called(
|
||||
self, engine, trace_store, integrity, provider
|
||||
):
|
||||
_ingest(trace_store, _complete_trace())
|
||||
integrity.mark("engine-learner", "engine-task", "INCOMPLETE_FLOODED")
|
||||
|
||||
record = await engine.grade("engine-learner", "engine-task")
|
||||
|
||||
assert record.verdict == "UNGRADABLE_TRACE_INCOMPLETE"
|
||||
assert provider.calls == 0, "LLM was called despite integrity flag (G-4)"
|
||||
assert record.scores["integrity_flag"] == "INCOMPLETE_FLOODED"
|
||||
assert record.scores["missing_seqs"] == [] # complete rows, but untrusted
|
||||
|
||||
async def test_flag_wins_even_with_gaps(self, engine, trace_store, integrity, provider):
|
||||
"""Both gate signals set: the flag reason is surfaced (both listed)."""
|
||||
events = _complete_trace()
|
||||
_ingest(trace_store, [e for e in events if e.seq != 2])
|
||||
integrity.mark("engine-learner", "engine-task", "INCOMPLETE_FLOODED")
|
||||
|
||||
record = await engine.grade("engine-learner", "engine-task")
|
||||
|
||||
assert record.verdict == "UNGRADABLE_TRACE_INCOMPLETE"
|
||||
assert provider.calls == 0
|
||||
assert record.scores["integrity_flag"] == "INCOMPLETE_FLOODED"
|
||||
assert record.scores["missing_seqs"] == [2]
|
||||
|
||||
async def test_empty_trace_ungradable(self, engine, provider, grade_store):
|
||||
record = await engine.grade("engine-learner", "engine-task")
|
||||
|
||||
assert record.verdict == "UNGRADABLE_EMPTY_TRACE"
|
||||
assert provider.calls == 0, "LLM was called for an empty trace"
|
||||
assert record.scores == {"integrity_flag": None, "missing_seqs": []}
|
||||
assert record.digest == {}
|
||||
stored = grade_store.get("engine-learner", "engine-task")
|
||||
assert stored is not None
|
||||
assert stored.verdict == "UNGRADABLE_EMPTY_TRACE"
|
||||
|
||||
|
||||
class TestD020RetryPath:
|
||||
async def test_malformed_first_response_recovers_via_retry(
|
||||
self, engine, trace_store, provider
|
||||
):
|
||||
"""First reply invalid (wrong shape, fenced), second valid → graded.
|
||||
|
||||
This exercises the D-020 layer-4 path INSIDE the engine through the
|
||||
real agents/structured.py — the engine does not reimplement it.
|
||||
"""
|
||||
_ingest(trace_store, _complete_trace())
|
||||
provider.replies = [
|
||||
'```json\n{"summary": "wrong shape"}\n```', # layer 3 rejects
|
||||
json.dumps(RubricJSON.PAYLOAD), # retry recovers
|
||||
]
|
||||
|
||||
record = await engine.grade("engine-learner", "engine-task")
|
||||
|
||||
assert provider.calls == 2 # bounded to exactly one retry (layer 4)
|
||||
assert record.verdict == "GRADED"
|
||||
assert record.scores == RubricJSON.PAYLOAD
|
||||
# the retry feedback must carry the validation error back
|
||||
retry_request = provider.requests[1]
|
||||
retry_text = " ".join(m.content for m in retry_request)
|
||||
assert "previous response was invalid" in retry_text
|
||||
|
||||
async def test_persistently_malformed_raises_after_bounded_retry(
|
||||
self, engine, trace_store, provider
|
||||
):
|
||||
from ai_service.agents.structured import StructuredOutputError
|
||||
|
||||
_ingest(trace_store, _complete_trace())
|
||||
provider.replies = ["not json at all", "still not json"]
|
||||
|
||||
with pytest.raises(StructuredOutputError):
|
||||
await engine.grade("engine-learner", "engine-task")
|
||||
assert provider.calls == 2 # bounded: never more than one retry
|
||||
|
||||
|
||||
class TestRegrade:
|
||||
async def test_regrade_overwrites_stored_record(
|
||||
self, engine, trace_store, grade_store, provider
|
||||
):
|
||||
_ingest(trace_store, _complete_trace())
|
||||
|
||||
first = await engine.grade("engine-learner", "engine-task")
|
||||
assert first.verdict == "GRADED"
|
||||
|
||||
# second grade: the provider now scripts a DIFFERENT rubric outcome
|
||||
updated = json.loads(json.dumps(RubricJSON.PAYLOAD))
|
||||
updated["criteria"]["process_quality"] = 1
|
||||
updated["verdict"] = "not_yet"
|
||||
provider.replies = [json.dumps(updated)]
|
||||
|
||||
second = await engine.grade("engine-learner", "engine-task")
|
||||
|
||||
assert second.scores == updated
|
||||
assert second.created_at >= first.created_at
|
||||
# upsert, not append: exactly one row, holding the LATEST grade
|
||||
grades = grade_store.list_for_learner("engine-learner")
|
||||
assert len(grades) == 1
|
||||
assert grades[0].scores == updated
|
||||
assert grades[0].verdict == "GRADED"
|
||||
|
||||
async def test_regrade_after_gate_outcome_replaces_it(
|
||||
self, engine, trace_store, grade_store, provider, integrity
|
||||
):
|
||||
"""A later successful regrade wholesale-replaces a gate record —
|
||||
the GradeStore upsert contract applied across verdict kinds."""
|
||||
# first: gapped → gate record persisted
|
||||
events = _complete_trace()
|
||||
_ingest(trace_store, [e for e in events if e.seq != 2])
|
||||
gated = await engine.grade("engine-learner", "engine-task")
|
||||
assert gated.verdict == "UNGRADABLE_TRACE_INCOMPLETE"
|
||||
|
||||
# gap healed (late delivery): same pair regrades cleanly
|
||||
trace_store.append(next(e for e in events if e.seq == 2))
|
||||
record = await engine.grade("engine-learner", "engine-task")
|
||||
|
||||
assert record.verdict == "GRADED"
|
||||
stored = grade_store.get("engine-learner", "engine-task")
|
||||
assert stored is not None
|
||||
assert stored.verdict == "GRADED"
|
||||
assert stored.scores == RubricJSON.PAYLOAD
|
||||
assert len(grade_store.list_for_learner("engine-learner")) == 1
|
||||
|
||||
|
||||
class TestBoundary:
|
||||
def test_engine_module_imports_no_fastapi(self):
|
||||
"""D-027: the engine module must not import fastapi (transitively
|
||||
beyond the sanctioned agents/structured + llm + telemetry + prompts)."""
|
||||
import sys
|
||||
|
||||
import ai_service.grading.engine as engine_mod
|
||||
|
||||
code = open(engine_mod.__file__).read()
|
||||
assert "fastapi" not in code
|
||||
assert "from ..api" not in code and "from ..api." not in code
|
||||
# the one sanctioned agents import is the module-direct structured defense
|
||||
assert "from ..agents.structured import" in code
|
||||
assert "from ..agents import" not in code
|
||||
# sanity: the module really is loaded (no import cycle surprises)
|
||||
assert engine_mod.__name__ in sys.modules
|
||||
|
||||
|
||||
class TestModelValidationSmoke:
|
||||
"""RubricScore as a plain pydantic model (D-020 layer-3 target type)."""
|
||||
|
||||
def test_rubric_score_is_frozen_shape_for_scores_dict(self):
|
||||
score = RubricScore.model_validate(RubricJSON.PAYLOAD)
|
||||
dumped = score.model_dump()
|
||||
assert set(dumped["criteria"]) == {
|
||||
"process_quality",
|
||||
"correctness",
|
||||
"debugging_discipline",
|
||||
"test_usage",
|
||||
}
|
||||
assert isinstance(score, BaseModel)
|
||||
|
||||
|
||||
class TestVariantAnchorsShipment:
|
||||
"""Phase 4 MH#4: variant anchors + seed ship into grading (a-5 same bar)."""
|
||||
|
||||
@pytest.fixture
|
||||
def variant_engine(self, trace_store, grade_store, integrity, provider, tmp_path):
|
||||
store = SQLiteVariantStore(db_path=tmp_path / "variants.db")
|
||||
yield GradingEngine(
|
||||
trace_store,
|
||||
grade_store,
|
||||
integrity,
|
||||
provider,
|
||||
model="gemma4:31b",
|
||||
variant_store=store,
|
||||
), store
|
||||
store.close()
|
||||
|
||||
@pytest.fixture
|
||||
def plain_engine(self, trace_store, grade_store, integrity, provider):
|
||||
return GradingEngine(
|
||||
trace_store, grade_store, integrity, provider, model="gemma4:31b"
|
||||
)
|
||||
|
||||
async def test_variant_task_grades_with_anchors_and_seed(
|
||||
self, variant_engine, trace_store, provider
|
||||
):
|
||||
engine, vstore = variant_engine
|
||||
_ingest(trace_store, _complete_trace())
|
||||
|
||||
# A stored variant whose task_id matches the graded trace.
|
||||
from datetime import UTC
|
||||
from datetime import datetime as dt
|
||||
|
||||
from ai_service.variants.store import VariantRecord
|
||||
|
||||
vstore.save(
|
||||
VariantRecord(
|
||||
learner_id="engine-learner",
|
||||
task_id="engine-task",
|
||||
template_id="tpl-llm-judge",
|
||||
seed="cafe" * 16,
|
||||
params={"domain": "tutoring"},
|
||||
statement="Scripted statement long enough to be legal.",
|
||||
starter_files={"README.md": "x"},
|
||||
created_at=dt.now(UTC),
|
||||
)
|
||||
)
|
||||
|
||||
record = await engine.grade("engine-learner", "engine-task")
|
||||
|
||||
assert record.verdict == "GRADED"
|
||||
assert record.variant_seed == "cafe" * 16 # D-029 stamped
|
||||
user_msgs = [m.content for req in provider.requests for m in req if m.role == "user"]
|
||||
assert any("Expected effort envelope" in u for u in user_msgs)
|
||||
assert any("tpl-llm-judge" in u for u in user_msgs)
|
||||
assert any("expected_min_test_runs" in u for u in user_msgs)
|
||||
|
||||
async def test_non_variant_task_has_no_anchors(
|
||||
self, variant_engine, trace_store, provider
|
||||
):
|
||||
engine, _ = variant_engine
|
||||
_ingest(trace_store, _complete_trace())
|
||||
|
||||
record = await engine.grade("engine-learner", "engine-task")
|
||||
|
||||
assert record.verdict == "GRADED"
|
||||
assert record.variant_seed is None
|
||||
user_msgs = [m.content for req in provider.requests for m in req if m.role == "user"]
|
||||
assert not any("Expected effort envelope" in u for u in user_msgs)
|
||||
|
||||
async def test_plain_engine_stays_variant_blind(
|
||||
self, plain_engine, trace_store
|
||||
):
|
||||
_ingest(trace_store, _complete_trace())
|
||||
record = await plain_engine.grade("engine-learner", "engine-task")
|
||||
assert record.verdict == "GRADED"
|
||||
assert record.variant_seed is None
|
||||
@@ -0,0 +1,170 @@
|
||||
"""Trace digest tests (Task 3-1-01, REQ-3-004).
|
||||
|
||||
Contract: `compute_digest` is deterministic, bounded (< 4 KB), and carries
|
||||
NO raw trace material (commands, file contents, payload strings) — the raw
|
||||
trace never reaches the LLM (D-028).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from datetime import UTC, datetime, timedelta
|
||||
|
||||
import pytest
|
||||
|
||||
from ai_service.grading.features import TraceDigest, compute_digest
|
||||
from ai_service.telemetry.models import TelemetryEvent
|
||||
|
||||
T0 = datetime(2026, 9, 12, 1, 0, 0, tzinfo=UTC)
|
||||
|
||||
|
||||
def _event(
|
||||
seq: int, kind: str, payload: dict | None = None, offset_s: float = 0.0
|
||||
) -> TelemetryEvent:
|
||||
return TelemetryEvent(
|
||||
learner_id="features-learner",
|
||||
task_id="features-task",
|
||||
seq=seq,
|
||||
kind=kind,
|
||||
payload=payload or {},
|
||||
ts=T0 + timedelta(seconds=offset_s),
|
||||
sandbox_id="sbx-features",
|
||||
)
|
||||
|
||||
|
||||
def _paste_and_run_trace() -> list[TelemetryEvent]:
|
||||
"""One big file dump, a single passing test at the very end (no cycles)."""
|
||||
return [
|
||||
_event(0, "activity", {"state": "starting"}, 0.0),
|
||||
_event(1, "file_diff", {"path": "main.py", "added": 180, "removed": 0}, 5.0),
|
||||
_event(2, "command", {"cmd": "python -m pytest -q"}, 30.0),
|
||||
_event(3, "test_result", {"passed": True, "exit_code": 0}, 45.0),
|
||||
_event(4, "activity", {"state": "idle"}, 50.0),
|
||||
]
|
||||
|
||||
|
||||
def _iterative_trace() -> list[TelemetryEvent]:
|
||||
"""Many small edits; failed runs interleaved; eventual pass (>= 2 cycles)."""
|
||||
events: list[TelemetryEvent] = [
|
||||
_event(0, "activity", {"state": "starting"}, 0.0),
|
||||
_event(1, "file_diff", {"path": "a.py", "added": 12}, 10.0),
|
||||
_event(2, "command", {"cmd": "pytest -q"}, 20.0),
|
||||
_event(3, "test_result", {"passed": False, "exit_code": 1}, 25.0),
|
||||
_event(4, "file_diff", {"path": "a.py", "added": 4, "removed": 2}, 40.0),
|
||||
_event(5, "file_diff", {"path": "b.py", "added": 6}, 50.0),
|
||||
_event(6, "command", {"cmd": "pytest -q"}, 60.0),
|
||||
_event(7, "test_result", {"passed": False, "exit_code": 2}, 65.0),
|
||||
_event(8, "file_diff", {"path": "a.py", "added": 3, "removed": 1}, 80.0),
|
||||
_event(9, "command", {"cmd": "pytest -q tests/"}, 90.0),
|
||||
_event(10, "test_result", {"passed": True, "exit_code": 0}, 95.0),
|
||||
_event(11, "activity", {"state": "idle"}, 100.0),
|
||||
]
|
||||
return events
|
||||
|
||||
|
||||
def test_paste_and_run_vs_iterative_produce_observably_different_digests() -> None:
|
||||
paste = compute_digest(_paste_and_run_trace())
|
||||
iterative = compute_digest(_iterative_trace())
|
||||
assert paste.error_fix_cycles == 0
|
||||
assert iterative.error_fix_cycles == 2
|
||||
assert paste.test_fail_count == 0
|
||||
assert iterative.test_fail_count == 2
|
||||
assert paste.edit_count < iterative.edit_count
|
||||
assert iterative.first_test_pass_offset_s is not None
|
||||
assert paste.first_test_pass_offset_s is not None
|
||||
# Iterative debugs longer before the first pass.
|
||||
assert iterative.first_test_pass_offset_s > paste.first_test_pass_offset_s
|
||||
|
||||
|
||||
def test_digest_is_deterministic() -> None:
|
||||
trace = _iterative_trace()
|
||||
assert compute_digest(trace) == compute_digest(list(reversed(trace))) # seq sort normalizes
|
||||
|
||||
|
||||
def test_empty_trace_yields_valid_zeroed_digest() -> None:
|
||||
digest = compute_digest([])
|
||||
assert digest.event_count == 0
|
||||
assert digest.final_test_status == "none"
|
||||
assert digest.first_test_pass_offset_s is None
|
||||
assert digest.mean_fix_latency_s is None
|
||||
assert isinstance(digest, TraceDigest)
|
||||
|
||||
|
||||
def test_digest_json_is_bounded_under_4kb() -> None:
|
||||
big = [
|
||||
_event(i, "command", {"cmd": f"grep PATTERN-{i} file-{i}.py"}, i * 1.0)
|
||||
for i in range(200)
|
||||
]
|
||||
serialized = compute_digest(big).model_dump_json()
|
||||
assert len(serialized.encode()) < 4096, f"digest too large: {len(serialized)}B"
|
||||
|
||||
|
||||
def test_no_raw_command_string_leaks_into_digest() -> None:
|
||||
marker = "SECRET-COMMAND-MARKER-7f3a"
|
||||
trace = [
|
||||
_event(0, "command", {"cmd": f"echo {marker} && cat /etc/hostname"}, 0.0),
|
||||
_event(1, "file_diff", {"path": marker + ".py"}, 1.0),
|
||||
_event(2, "run_result", {"exit_code": 0, "stdout": marker}, 2.0),
|
||||
]
|
||||
digest_json = compute_digest(trace).model_dump_json()
|
||||
assert marker not in digest_json, "raw payload material leaked into digest"
|
||||
|
||||
|
||||
def test_idle_gaps_computed_over_threshold() -> None:
|
||||
trace = [
|
||||
_event(0, "activity", {"state": "starting"}, 0.0),
|
||||
_event(1, "activity", {"state": "idle"}, 400.0), # > 120s gap
|
||||
_event(2, "activity", {"state": "idle"}, 500.0), # 100s gap (below)
|
||||
_event(3, "activity", {"state": "stopped"}, 800.0), # > 120s gap
|
||||
]
|
||||
digest = compute_digest(trace)
|
||||
assert digest.idle_gap_count == 2
|
||||
assert digest.idle_gap_total_s == pytest.approx(400.0 + 300.0, rel=1e-6)
|
||||
|
||||
|
||||
def test_command_category_histogram() -> None:
|
||||
trace = [
|
||||
_event(0, "command", {"cmd": "npm run build"}, 0.0),
|
||||
_event(1, "command", {"cmd": "pytest -q"}, 1.0),
|
||||
_event(2, "command", {"cmd": "ls -la"}, 2.0),
|
||||
_event(3, "command", {"cmd": "rm -rf build/"}, 3.0),
|
||||
_event(4, "command", {"cmd": "curl localhost:8420/health"}, 4.0),
|
||||
_event(5, "command", {"cmd": "python mystery.py"}, 5.0),
|
||||
]
|
||||
digest = compute_digest(trace)
|
||||
assert digest.command_categories == {
|
||||
"build": 1,
|
||||
"debug": 1,
|
||||
"file": 1,
|
||||
"nav": 1,
|
||||
"other": 1,
|
||||
"test": 1,
|
||||
}
|
||||
|
||||
|
||||
def test_daemon_topology_mix_is_tolerated() -> None:
|
||||
"""P2-verify P1: live traces carry activity+file_diff only — no crash."""
|
||||
trace = [
|
||||
_event(0, "activity", {"state": "starting"}, 0.0),
|
||||
_event(1, "file_diff", {"path": "made-by-exec.txt", "added": 1}, 2.0),
|
||||
_event(2, "activity", {"state": "idle"}, 3.0),
|
||||
]
|
||||
digest = compute_digest(trace)
|
||||
assert digest.final_test_status == "none"
|
||||
assert digest.edit_count == 1
|
||||
assert digest.test_pass_count == 0
|
||||
assert digest.kind_histogram.get("file_diff") == 1
|
||||
|
||||
|
||||
def test_run_result_exit_codes_fall_back_for_test_status() -> None:
|
||||
"""No test_result events: run_result exit codes decide pass/fail."""
|
||||
trace = [
|
||||
_event(0, "file_diff", {"path": "x.py"}, 0.0),
|
||||
_event(1, "run_result", {"exit_code": 1}, 10.0),
|
||||
_event(2, "file_diff", {"path": "x.py"}, 20.0),
|
||||
_event(3, "run_result", {"exit_code": 0}, 30.0),
|
||||
]
|
||||
digest = compute_digest(trace)
|
||||
assert digest.test_fail_count == 1
|
||||
assert digest.test_pass_count == 1
|
||||
assert digest.final_test_status == "pass"
|
||||
assert digest.error_fix_cycles == 1
|
||||
@@ -0,0 +1,268 @@
|
||||
"""SQLiteGradeStore tests (REQ-3-004, D-027).
|
||||
|
||||
Each test gets its own tmp-path SQLite file — no shared disk state. Covers:
|
||||
- save / get / list_for_learner roundtrip (all fields survive,
|
||||
including nested JSON dicts and the tz-aware created_at contract)
|
||||
- upsert-on-regrade: a second save with the same (learner_id, task_id)
|
||||
REPLACES the row wholesale — scores, verdict, created_at, digest,
|
||||
model and variant_seed all reflect the latest save (documented
|
||||
contract; deliberately opposite of TraceStore.append's dedup)
|
||||
- unknown (learner, task) pair -> None; unknown learner -> empty list
|
||||
- scoping: grades for other learners/tasks are never returned
|
||||
- variant_seed stays None until P4 (D-029) and survives a roundtrip
|
||||
- rows are detached: usable after the store is closed
|
||||
- WAL + synchronous=NORMAL pragmas actually applied to the DB file
|
||||
- concurrent writer + reader against the same DB file (a-3 smoke test)
|
||||
"""
|
||||
|
||||
import concurrent.futures
|
||||
import threading
|
||||
from datetime import UTC, datetime, timedelta
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
import pytest
|
||||
import sqlalchemy as sa
|
||||
|
||||
from ai_service.grading.store import GradeRecord, SQLiteGradeStore
|
||||
|
||||
_BASE_TS = datetime(2026, 9, 12, 12, 0, 0, tzinfo=UTC)
|
||||
|
||||
|
||||
def make_grade(
|
||||
task_id: str = "task-1",
|
||||
learner_id: str = "learner-1",
|
||||
variant_seed: str | None = None,
|
||||
digest: dict[str, Any] | None = None,
|
||||
scores: dict[str, Any] | None = None,
|
||||
verdict: str = "STRONG",
|
||||
model: str = "gemma4:31b",
|
||||
created_at: datetime | None = None,
|
||||
) -> GradeRecord:
|
||||
"""Canonical kwargs builder — tests override only what they assert on."""
|
||||
return GradeRecord(
|
||||
learner_id=learner_id,
|
||||
task_id=task_id,
|
||||
variant_seed=variant_seed,
|
||||
digest=digest
|
||||
if digest is not None
|
||||
else {"error_fix_cycles": 3, "command_categories": {"build": 2}},
|
||||
scores=scores
|
||||
if scores is not None
|
||||
else {
|
||||
"criteria": {"process_quality": 3, "correctness": 4},
|
||||
"strengths": ["iterative debugging"],
|
||||
"gaps": ["no final test pass"],
|
||||
},
|
||||
verdict=verdict,
|
||||
model=model,
|
||||
created_at=created_at if created_at is not None else _BASE_TS,
|
||||
)
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def store(tmp_path: Path) -> SQLiteGradeStore:
|
||||
s = SQLiteGradeStore(db_path=tmp_path / "grades.db")
|
||||
yield s
|
||||
s.close()
|
||||
|
||||
|
||||
def test_save_and_get_roundtrip(store: SQLiteGradeStore) -> None:
|
||||
grade = make_grade()
|
||||
store.save(grade)
|
||||
|
||||
fetched = store.get("learner-1", "task-1")
|
||||
assert fetched is not None
|
||||
assert fetched.learner_id == "learner-1"
|
||||
assert fetched.task_id == "task-1"
|
||||
assert fetched.variant_seed is None # null until P4 (D-029)
|
||||
assert fetched.digest == grade.digest
|
||||
assert fetched.scores == grade.scores
|
||||
assert fetched.verdict == "STRONG"
|
||||
assert fetched.model == "gemma4:31b"
|
||||
assert fetched.created_at == _BASE_TS
|
||||
assert fetched.created_at.tzinfo is UTC # tz-normalized on read
|
||||
|
||||
|
||||
def test_get_unknown_pair_returns_none(store: SQLiteGradeStore) -> None:
|
||||
store.save(make_grade())
|
||||
|
||||
assert store.get("learner-1", "task-missing") is None
|
||||
assert store.get("learner-missing", "task-1") is None
|
||||
assert store.get("nobody", "nothing") is None
|
||||
|
||||
|
||||
def test_list_for_learner_roundtrip(store: SQLiteGradeStore) -> None:
|
||||
# Created out of insertion order; list must come back chronological.
|
||||
store.save(make_grade(task_id="task-c", created_at=_BASE_TS + timedelta(hours=2)))
|
||||
store.save(make_grade(task_id="task-a", created_at=_BASE_TS))
|
||||
store.save(make_grade(task_id="task-b", created_at=_BASE_TS + timedelta(hours=1)))
|
||||
|
||||
grades = store.list_for_learner("learner-1")
|
||||
assert [g.task_id for g in grades] == ["task-a", "task-b", "task-c"]
|
||||
assert all(g.learner_id == "learner-1" for g in grades)
|
||||
hours = (timedelta(hours=0), timedelta(hours=1), timedelta(hours=2))
|
||||
assert all(
|
||||
g.created_at == _BASE_TS + offset
|
||||
for g, offset in zip(grades, hours, strict=True)
|
||||
)
|
||||
assert all(g.created_at.tzinfo is UTC for g in grades)
|
||||
|
||||
|
||||
def test_list_for_learner_unknown_learner_returns_empty_list(
|
||||
store: SQLiteGradeStore,
|
||||
) -> None:
|
||||
assert store.list_for_learner("nobody") == []
|
||||
|
||||
|
||||
def test_lists_are_scoped_to_the_learner(store: SQLiteGradeStore) -> None:
|
||||
store.save(make_grade(learner_id="learner-1", task_id="task-1"))
|
||||
store.save(make_grade(learner_id="learner-2", task_id="task-1"))
|
||||
|
||||
assert [g.task_id for g in store.list_for_learner("learner-1")] == ["task-1"]
|
||||
assert [g.learner_id for g in store.list_for_learner("learner-2")] == ["learner-2"]
|
||||
# (learner-2, task-1) is a distinct row: same task_id, different grade.
|
||||
grades_2 = store.list_for_learner("learner-2")
|
||||
assert len(grades_2) == 1
|
||||
assert grades_2[0].learner_id == "learner-2"
|
||||
|
||||
|
||||
def test_regrade_overwrites_the_stored_row(store: SQLiteGradeStore) -> None:
|
||||
"""THE contract of this store (upsert on the PK pair, latest wins).
|
||||
|
||||
The grading engine re-grades a task as its rubric or input evolves;
|
||||
the second save replaces scores, verdict, created_at, digest, model
|
||||
and variant_seed wholesale — exactly one row survives per pair.
|
||||
"""
|
||||
first = make_grade(
|
||||
verdict="DEVELOPING",
|
||||
model="gemma4:31b",
|
||||
scores={"criteria": {"process_quality": 1}},
|
||||
digest={"error_fix_cycles": 0},
|
||||
created_at=_BASE_TS,
|
||||
)
|
||||
store.save(first)
|
||||
|
||||
second = make_grade(
|
||||
verdict="EXEMPLARY",
|
||||
model="gemma4:31b-p2",
|
||||
variant_seed="seed-77",
|
||||
scores={"criteria": {"process_quality": 4}},
|
||||
digest={"error_fix_cycles": 6},
|
||||
created_at=_BASE_TS + timedelta(hours=1),
|
||||
)
|
||||
store.save(second)
|
||||
|
||||
fetched = store.get("learner-1", "task-1")
|
||||
assert fetched is not None
|
||||
# The regrade replaced every field of the first save.
|
||||
assert fetched.verdict == "EXEMPLARY"
|
||||
assert fetched.model == "gemma4:31b-p2"
|
||||
assert fetched.variant_seed == "seed-77"
|
||||
assert fetched.scores == {"criteria": {"process_quality": 4}}
|
||||
assert fetched.digest == {"error_fix_cycles": 6}
|
||||
assert fetched.created_at == _BASE_TS + timedelta(hours=1)
|
||||
|
||||
# Latest-wins also holds in list_for_learner — one row, not two.
|
||||
grades = store.list_for_learner("learner-1")
|
||||
assert len(grades) == 1
|
||||
assert grades[0].verdict == "EXEMPLARY"
|
||||
|
||||
|
||||
def test_regrade_preserves_other_pairs(store: SQLiteGradeStore) -> None:
|
||||
# An upsert on (learner-1, task-1) must not touch (learner-1, task-2).
|
||||
store.save(make_grade(task_id="task-1", verdict="STRONG"))
|
||||
store.save(make_grade(task_id="task-2", verdict="DEVELOPING"))
|
||||
store.save(make_grade(task_id="task-1", verdict="EXEMPLARY"))
|
||||
|
||||
other = store.get("learner-1", "task-2")
|
||||
assert other is not None
|
||||
assert other.verdict == "DEVELOPING" # untouched by the task-1 regrade
|
||||
assert len(store.list_for_learner("learner-1")) == 2
|
||||
|
||||
|
||||
def test_empty_scores_dict_roundtrips(store: SQLiteGradeStore) -> None:
|
||||
# Legal shape: an UNGRADABLE_TRACE_INCOMPLETE record carries a verdict
|
||||
# but no scores (and here, no digest either).
|
||||
store.save(make_grade(verdict="UNGRADABLE_TRACE_INCOMPLETE", scores={}, digest={}))
|
||||
|
||||
fetched = store.get("learner-1", "task-1")
|
||||
assert fetched is not None
|
||||
assert fetched.scores == {}
|
||||
assert fetched.digest == {}
|
||||
assert fetched.verdict == "UNGRADABLE_TRACE_INCOMPLETE"
|
||||
|
||||
|
||||
def test_rows_are_detached_after_save(store: SQLiteGradeStore, tmp_path: Path) -> None:
|
||||
# The engine hands GradeRecords across layers; rows must survive the
|
||||
# store that produced them being closed (no open-session ORM magic).
|
||||
store.save(make_grade())
|
||||
fetched = store.get("learner-1", "task-1")
|
||||
store.close()
|
||||
|
||||
assert fetched is not None
|
||||
assert fetched.verdict == "STRONG"
|
||||
assert fetched.scores["criteria"]["process_quality"] == 3
|
||||
|
||||
# A fresh store on the same file sees the same row (durability).
|
||||
reopened = SQLiteGradeStore(db_path=tmp_path / "grades.db")
|
||||
try:
|
||||
again = reopened.get("learner-1", "task-1")
|
||||
assert again is not None
|
||||
assert again.verdict == "STRONG"
|
||||
finally:
|
||||
reopened.close()
|
||||
|
||||
|
||||
def test_pragmas_are_applied(store: SQLiteGradeStore) -> None:
|
||||
# Pragmas are per-connection; query through the store's engine so the
|
||||
# connect hook (not a default sqlite3 connection) is what we inspect.
|
||||
with store._engine.connect() as conn:
|
||||
(journal_mode,) = conn.execute(sa.text("PRAGMA journal_mode")).one()
|
||||
(synchronous,) = conn.execute(sa.text("PRAGMA synchronous")).one()
|
||||
assert journal_mode == "wal"
|
||||
# synchronous=NORMAL is 1 in SQLite's pragma numbering.
|
||||
assert synchronous == 1
|
||||
|
||||
|
||||
def test_concurrent_writer_and_reader_no_database_is_locked(tmp_path: Path) -> None:
|
||||
"""One thread saves while another reads in a tight loop (a-3).
|
||||
|
||||
Without WAL + busy_timeout this pattern reliably produces
|
||||
`OperationalError: database is locked` on SQLite. The assertion is
|
||||
that every reader call completes and the latest write lands intact.
|
||||
"""
|
||||
db_path = tmp_path / "grades.db"
|
||||
n_pairs = 60 # distinct tasks -> distinct PK pairs, one regrade each
|
||||
stop_writing = threading.Event()
|
||||
|
||||
writer = SQLiteGradeStore(db_path=db_path)
|
||||
reader = SQLiteGradeStore(db_path=db_path)
|
||||
try:
|
||||
|
||||
def write_grades() -> None:
|
||||
for seq in range(n_pairs):
|
||||
writer.save(
|
||||
make_grade(
|
||||
task_id=f"task-{seq}",
|
||||
created_at=_BASE_TS + timedelta(seconds=seq),
|
||||
)
|
||||
)
|
||||
stop_writing.set()
|
||||
|
||||
def read_grades() -> None:
|
||||
while not stop_writing.is_set():
|
||||
reader.list_for_learner("learner-1")
|
||||
# Final read after the writer is done.
|
||||
assert len(reader.list_for_learner("learner-1")) == n_pairs
|
||||
|
||||
with concurrent.futures.ThreadPoolExecutor(max_workers=2) as pool:
|
||||
futures = [pool.submit(write_grades), pool.submit(read_grades)]
|
||||
for future in futures:
|
||||
future.result(timeout=30)
|
||||
|
||||
# Every save was an upsert on its own pair: nothing lost, none doubled.
|
||||
assert len(writer.list_for_learner("learner-1")) == n_pairs
|
||||
finally:
|
||||
reader.close()
|
||||
writer.close()
|
||||
@@ -0,0 +1,632 @@
|
||||
"""Unit tests for scripts/sandbox-agent.py (Task 2-2-01, REQ-3-003).
|
||||
|
||||
Live-agent tests run the real agent (no sandbox namespace required — it is a
|
||||
plain stdlib process) against a loopback fake WebSocket server implemented
|
||||
with raw socket frames, mirroring the agent's client codec. Covers:
|
||||
|
||||
* ordered seq emission over a live connection;
|
||||
* spool-on-disconnect (socket drop leaves events durable in the spool);
|
||||
* reconnect + flush preserves order (no loss; one-line replay margin may
|
||||
duplicate the last pre-drop event — the server dedups on
|
||||
(learner, task, seq));
|
||||
* SIGKILL resilience (boot replays the spool, `seq` resumes monotonically);
|
||||
* file_diff polling capture (created / modified / deleted);
|
||||
* an AST scan asserting the agent file is stdlib-only (D-031).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import ast
|
||||
import base64
|
||||
import hashlib
|
||||
import json
|
||||
import socket
|
||||
import struct
|
||||
import sys
|
||||
import threading
|
||||
import time
|
||||
from collections.abc import Iterator
|
||||
from datetime import datetime
|
||||
from importlib.util import module_from_spec, spec_from_file_location
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
AGENT_PATH = (
|
||||
Path(__file__).resolve().parents[2] / "scripts" / "sandbox-agent.py"
|
||||
).resolve()
|
||||
|
||||
_WS_GUID = "258EAFA5-E914-47DA-95CA-C5AB0DC85B11"
|
||||
|
||||
# Python 3.11.2+ stdlib top-level names (incl. platform leaves removed in 3.13
|
||||
# — the AST scan asserts policy, i.e. no third-party imports, not 3.13 import
|
||||
# viability). Kept explicit because sys.stdlib_module_names lacks them.
|
||||
_PY311_EXTRA_STDLIB = {
|
||||
"aifc",
|
||||
"asynchat",
|
||||
"asyncore",
|
||||
"audioop",
|
||||
"cgi",
|
||||
"cgitb",
|
||||
"chunk",
|
||||
"crypt",
|
||||
"imghdr",
|
||||
"mailcap",
|
||||
"msilib",
|
||||
"nis",
|
||||
"nntplib",
|
||||
"ossaudiodev",
|
||||
"pipes",
|
||||
"sndhdr",
|
||||
"spwd",
|
||||
"sunau",
|
||||
"uu",
|
||||
"xdrlib",
|
||||
"__future__",
|
||||
}
|
||||
_ALLOWED_MODULES = frozenset(set(sys.stdlib_module_names) | _PY311_EXTRA_STDLIB)
|
||||
|
||||
|
||||
def _load_agent():
|
||||
"""Load the hyphenated script as a module (it is not importable by name)."""
|
||||
spec = spec_from_file_location("sandbox_agent", AGENT_PATH)
|
||||
assert spec is not None and spec.loader is not None
|
||||
mod = module_from_spec(spec)
|
||||
sys.modules[spec.name] = mod
|
||||
spec.loader.exec_module(mod)
|
||||
return mod
|
||||
|
||||
|
||||
agent = _load_agent()
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Raw-socket frame helpers (test-harness mirror of the agent's codec)
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def _read_exact(conn: socket.socket, n: int) -> bytes:
|
||||
buf = bytearray()
|
||||
while len(buf) < n:
|
||||
chunk = conn.recv(n - len(buf))
|
||||
if not chunk:
|
||||
raise ConnectionError("unexpected EOF")
|
||||
buf += chunk
|
||||
return bytes(buf)
|
||||
|
||||
|
||||
def _server_read_frame(conn: socket.socket) -> tuple[int, bytes]:
|
||||
"""Read one client frame (client frames are always masked per RFC 6455)."""
|
||||
b0, b1 = _read_exact(conn, 2)
|
||||
opcode = b0 & 0x0F
|
||||
length = b1 & 0x7F
|
||||
if length == 126:
|
||||
length = struct.unpack("!H", _read_exact(conn, 2))[0]
|
||||
elif length == 127:
|
||||
length = struct.unpack("!Q", _read_exact(conn, 8))[0]
|
||||
mask = _read_exact(conn, 4) if (b1 & 0x80) else b""
|
||||
payload = _read_exact(conn, length) if length else b""
|
||||
if mask:
|
||||
payload = bytes(b ^ mask[i % 4] for i, b in enumerate(payload))
|
||||
return opcode, payload
|
||||
|
||||
|
||||
def _server_send_frame(conn: socket.socket, opcode: int, payload: bytes = b"") -> None:
|
||||
header = bytearray([0x80 | opcode])
|
||||
n = len(payload)
|
||||
if n < 126:
|
||||
header.append(n)
|
||||
elif n < 65536:
|
||||
header.append(126)
|
||||
header += struct.pack("!H", n)
|
||||
else:
|
||||
header.append(127)
|
||||
header += struct.pack("!Q", n)
|
||||
conn.sendall(bytes(header) + payload)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Loopback fake WS ingest server (thread-per-connection, no dependencies)
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
class _WSConn:
|
||||
def __init__(self, sock: socket.socket) -> None:
|
||||
self.sock = sock
|
||||
self.alive = True
|
||||
|
||||
|
||||
class FakeWSServer:
|
||||
"""Accepts RFC 6455 handshakes and collects text frames in arrival order."""
|
||||
|
||||
def __init__(self) -> None:
|
||||
self._listener = socket.socket()
|
||||
self._listener.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
|
||||
self._listener.bind(("127.0.0.1", 0))
|
||||
self._listener.listen(8)
|
||||
self._listener.settimeout(0.5)
|
||||
self.port: int = self._listener.getsockname()[1]
|
||||
self._connections: list[_WSConn] = []
|
||||
self._frames: list[tuple[int, bytes]] = []
|
||||
self._lock = threading.Lock()
|
||||
self._stop = threading.Event()
|
||||
self._thread = threading.Thread(target=self._run, daemon=True)
|
||||
|
||||
@property
|
||||
def url(self) -> str:
|
||||
return f"ws://127.0.0.1:{self.port}/v1/telemetry/ingest"
|
||||
|
||||
@property
|
||||
def events(self) -> list[dict]:
|
||||
with self._lock:
|
||||
return [json.loads(payload) for op, payload in self._frames if op == 0x1]
|
||||
|
||||
def has_live_connection(self) -> bool:
|
||||
with self._lock:
|
||||
return any(c.alive for c in self._connections)
|
||||
|
||||
def count_connections(self) -> int:
|
||||
with self._lock:
|
||||
return len(self._connections)
|
||||
|
||||
def kill_all_connections(self) -> None:
|
||||
"""Abruptly close the socket without a WS close frame (network drop)."""
|
||||
with self._lock:
|
||||
conns = list(self._connections)
|
||||
for conn in conns:
|
||||
conn.alive = False
|
||||
try:
|
||||
conn.sock.shutdown(socket.SHUT_RDWR)
|
||||
except OSError:
|
||||
pass
|
||||
try:
|
||||
conn.sock.close()
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
# -- internals ------------------------------------------------------
|
||||
def start(self) -> None:
|
||||
self._thread.start()
|
||||
|
||||
def stop(self) -> None:
|
||||
self._stop.set()
|
||||
self.kill_all_connections()
|
||||
self._thread.join(timeout=3)
|
||||
self._listener.close()
|
||||
|
||||
def _run(self) -> None:
|
||||
while not self._stop.is_set():
|
||||
try:
|
||||
sock, _ = self._listener.accept()
|
||||
except TimeoutError:
|
||||
continue
|
||||
except OSError:
|
||||
return
|
||||
threading.Thread(target=self._serve, args=(sock,), daemon=True).start()
|
||||
|
||||
def _serve(self, sock: socket.socket) -> None:
|
||||
conn = _WSConn(sock)
|
||||
with self._lock:
|
||||
self._connections.append(conn)
|
||||
try:
|
||||
self._handshake(sock)
|
||||
while conn.alive and not self._stop.is_set():
|
||||
opcode, payload = _server_read_frame(sock)
|
||||
if opcode == 0x8: # close
|
||||
_server_send_frame(sock, 0x8, payload)
|
||||
break
|
||||
with self._lock:
|
||||
self._frames.append((opcode, payload))
|
||||
except (ConnectionError, OSError):
|
||||
pass
|
||||
finally:
|
||||
conn.alive = False
|
||||
try:
|
||||
sock.close()
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
@staticmethod
|
||||
def _handshake(sock: socket.socket) -> None:
|
||||
buf = b""
|
||||
while b"\r\n\r\n" not in buf:
|
||||
chunk = sock.recv(4096)
|
||||
if not chunk:
|
||||
raise ConnectionError("EOF during handshake")
|
||||
buf += chunk
|
||||
headers: dict[str, str] = {}
|
||||
for line in buf.decode("latin-1").split("\r\n")[1:]:
|
||||
if ":" in line:
|
||||
name, _, value = line.partition(":")
|
||||
headers[name.strip().lower()] = value.strip()
|
||||
accept = base64.b64encode(
|
||||
hashlib.sha1((headers["sec-websocket-key"] + _WS_GUID).encode()).digest()
|
||||
).decode()
|
||||
sock.sendall(
|
||||
(
|
||||
"HTTP/1.1 101 Switching Protocols\r\n"
|
||||
"Upgrade: websocket\r\n"
|
||||
"Connection: Upgrade\r\n"
|
||||
f"Sec-WebSocket-Accept: {accept}\r\n\r\n"
|
||||
).encode("ascii")
|
||||
)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Killable TCP proxy: lets a test hold the network down deterministically.
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
class KillableProxy:
|
||||
"""Forwards bytes both ways; `kill()` severs + closes the listener so the
|
||||
client can neither talk nor reconnect until `resume()` re-opens it."""
|
||||
|
||||
def __init__(self, target_port: int) -> None:
|
||||
self._target = ("127.0.0.1", target_port)
|
||||
self._listener = socket.socket()
|
||||
self._listener.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
|
||||
self._listener.bind(("127.0.0.1", 0))
|
||||
self._listener.listen(8)
|
||||
self._listener.settimeout(0.2)
|
||||
self.port: int = self._listener.getsockname()[1]
|
||||
self._up = threading.Event()
|
||||
self._up.set()
|
||||
self._dead = threading.Event()
|
||||
self._conns: list[socket.socket] = []
|
||||
self._lock = threading.Lock()
|
||||
self._thread = threading.Thread(target=self._run, daemon=True)
|
||||
|
||||
@property
|
||||
def url(self) -> str:
|
||||
return f"ws://127.0.0.1:{self.port}/v1/telemetry/ingest"
|
||||
|
||||
def start(self) -> None:
|
||||
self._thread.start()
|
||||
|
||||
def kill(self) -> None:
|
||||
"""Sever live sockets AND stop accepting until resumed (network down)."""
|
||||
self._up.clear()
|
||||
with self._lock:
|
||||
conns = list(self._conns)
|
||||
for s in conns:
|
||||
try:
|
||||
s.shutdown(socket.SHUT_RDWR)
|
||||
except OSError:
|
||||
pass
|
||||
try:
|
||||
s.close()
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
def resume(self) -> None:
|
||||
self._up.set()
|
||||
|
||||
def stop(self) -> None:
|
||||
self._dead.set()
|
||||
self.kill()
|
||||
self._thread.join(timeout=3)
|
||||
self._listener.close()
|
||||
|
||||
def _run(self) -> None:
|
||||
while not self._dead.is_set():
|
||||
try:
|
||||
client, _ = self._listener.accept()
|
||||
except TimeoutError:
|
||||
continue
|
||||
except OSError:
|
||||
return
|
||||
if not self._up.is_set():
|
||||
client.close()
|
||||
continue
|
||||
try:
|
||||
upstream = socket.create_connection(self._target, timeout=2)
|
||||
except OSError:
|
||||
client.close()
|
||||
continue
|
||||
with self._lock:
|
||||
self._conns.extend([client, upstream])
|
||||
threading.Thread(
|
||||
target=self._pump, args=(client, upstream), daemon=True
|
||||
).start()
|
||||
threading.Thread(
|
||||
target=self._pump, args=(upstream, client), daemon=True
|
||||
).start()
|
||||
|
||||
def _pump(self, src: socket.socket, dst: socket.socket) -> None:
|
||||
try:
|
||||
while self._up.is_set() and not self._dead.is_set():
|
||||
data = src.recv(65536)
|
||||
if not data:
|
||||
break
|
||||
dst.sendall(data)
|
||||
except OSError:
|
||||
pass
|
||||
finally:
|
||||
for s in (src, dst):
|
||||
try:
|
||||
s.shutdown(socket.SHUT_RDWR)
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Fixtures / helpers
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def fake_server() -> Iterator[FakeWSServer]:
|
||||
server = FakeWSServer()
|
||||
server.start()
|
||||
yield server
|
||||
server.stop()
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def held_link(fake_server: FakeWSServer) -> Iterator[KillableProxy]:
|
||||
"""A killable proxy in front of the fake server (network-down control)."""
|
||||
proxy = KillableProxy(fake_server.port)
|
||||
proxy.start()
|
||||
yield proxy
|
||||
proxy.stop()
|
||||
|
||||
def _config(tmp_path: Path, url: str, spool_path: Path | None = None) -> agent.AgentConfig:
|
||||
workspace = tmp_path / "workspace"
|
||||
return agent.AgentConfig(
|
||||
learner_id="learner-1",
|
||||
task_id="task-1",
|
||||
ingest_url=url,
|
||||
sandbox_id="sb-test",
|
||||
workspace=workspace,
|
||||
spool_path=spool_path or (tmp_path / "spool" / "spool.jsonl"),
|
||||
poll_interval_s=0.05,
|
||||
activity_interval_s=30.0,
|
||||
backoff_base_s=0.05,
|
||||
backoff_max_s=0.2,
|
||||
)
|
||||
|
||||
|
||||
def _make_agent(tmp_path: Path, url: str, spool_path: Path | None = None) -> agent.Agent:
|
||||
return agent.Agent(_config(tmp_path, url, spool_path))
|
||||
|
||||
|
||||
def _wait_until(pred, timeout_s: float = 8.0, interval_s: float = 0.02) -> bool:
|
||||
deadline = time.monotonic() + timeout_s
|
||||
while time.monotonic() < deadline:
|
||||
if pred():
|
||||
return True
|
||||
time.sleep(interval_s)
|
||||
return pred()
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Tests
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
class TestOrderedEmission:
|
||||
def test_command_produces_ordered_seq_chain(
|
||||
self, tmp_path: Path, fake_server: FakeWSServer
|
||||
) -> None:
|
||||
test_agent = _make_agent(tmp_path, fake_server.url)
|
||||
test_agent.start()
|
||||
try:
|
||||
assert test_agent.wait_connected(5), "agent never connected"
|
||||
test_agent.run_command("echo hello-agent")
|
||||
finally:
|
||||
test_agent.stop()
|
||||
assert _wait_until(lambda: len(fake_server.events) >= 6)
|
||||
|
||||
events = fake_server.events
|
||||
delivered_seqs = [e["seq"] for e in events]
|
||||
assert delivered_seqs == sorted(delivered_seqs), "wire order must be seq order"
|
||||
assert min(delivered_seqs) == 0
|
||||
|
||||
kinds = [e["kind"] for e in events if e["kind"] != "activity"]
|
||||
assert kinds == ["stdin", "command", "stdout", "run_result"], kinds
|
||||
stdout = next(e for e in events if e["kind"] == "stdout")
|
||||
assert "hello-agent" in stdout["payload"]["data"]
|
||||
run_result = next(e for e in events if e["kind"] == "run_result")
|
||||
assert run_result["payload"]["exit_code"] == 0
|
||||
|
||||
# Wire contract (D-026): identity is bound at the WS handshake (URL
|
||||
# query params) — the frame body must NOT carry learner_id/task_id
|
||||
# (the ingest endpoint rejects them with extra="forbid").
|
||||
required = {"seq", "kind", "payload", "ts"}
|
||||
forbidden = {"learner_id", "task_id"}
|
||||
for event in events:
|
||||
assert required <= set(event), event
|
||||
assert not (forbidden & set(event)), f"URL-owned ids leaked in frame: {event}"
|
||||
assert event.get("sandbox_id", "sb-test") == "sb-test"
|
||||
datetime.fromisoformat(event["ts"]) # must parse
|
||||
|
||||
def test_test_command_classified_as_test_result(
|
||||
self, tmp_path: Path, fake_server: FakeWSServer
|
||||
) -> None:
|
||||
test_agent = _make_agent(tmp_path, fake_server.url)
|
||||
test_agent.start()
|
||||
try:
|
||||
assert test_agent.wait_connected(5)
|
||||
test_agent.run_command("pytest -q")
|
||||
finally:
|
||||
test_agent.stop()
|
||||
kinds = [e["kind"] for e in fake_server.events]
|
||||
assert "test_result" in kinds
|
||||
assert "run_result" not in kinds
|
||||
|
||||
|
||||
class TestSpoolOnDisconnect:
|
||||
def test_disconnect_spools_events_without_loss(
|
||||
self, tmp_path: Path, fake_server: FakeWSServer, held_link: KillableProxy
|
||||
) -> None:
|
||||
test_agent = _make_agent(tmp_path, held_link.url)
|
||||
test_agent.start()
|
||||
assert test_agent.wait_connected(5)
|
||||
first = test_agent.run_command("echo before-drop")
|
||||
assert _wait_until(
|
||||
lambda: any(
|
||||
e["kind"] == "run_result" and e["seq"] == first["seq"]
|
||||
for e in fake_server.events
|
||||
)
|
||||
)
|
||||
|
||||
held_link.kill() # network hard-down: sockets severed, listener closed
|
||||
assert test_agent.wait_disconnected(5), "agent never noticed the drop"
|
||||
seq_before = test_agent._seq # noqa: SLF001 — white-box unit test
|
||||
spooled = [
|
||||
test_agent.emit("activity", {"state": "offline", "n": n}) for n in range(3)
|
||||
]
|
||||
spooled_seqs = {e["seq"] for e in spooled}
|
||||
assert min(spooled_seqs) == seq_before, "seqs must continue monotonically"
|
||||
test_agent.stop()
|
||||
|
||||
# Network stayed down for the whole window → nothing was delivered.
|
||||
spool_lines = test_agent._spool.read_all()
|
||||
spool_seqs = [json.loads(line)["seq"] for line in spool_lines]
|
||||
assert spooled_seqs <= set(spool_seqs), "dropped events must be spool-durable"
|
||||
received_seqs = [e["seq"] for e in fake_server.events]
|
||||
assert spooled_seqs.isdisjoint(received_seqs), (
|
||||
"events emitted during the outage must NOT have reached the server"
|
||||
)
|
||||
# seq is monotonic across the whole run: offline events + the final
|
||||
# stop marker account for every value allocated after the drop.
|
||||
assert test_agent._seq == seq_before + 4 # 3 offline + "stopped"
|
||||
|
||||
|
||||
class TestReconnectFlush:
|
||||
def test_reconnect_flushes_spool_in_order_no_loss(
|
||||
self, tmp_path: Path, fake_server: FakeWSServer, held_link: KillableProxy
|
||||
) -> None:
|
||||
test_agent = _make_agent(tmp_path, held_link.url)
|
||||
test_agent.start()
|
||||
assert test_agent.wait_connected(5)
|
||||
first = test_agent.run_command("echo first")
|
||||
# P7 de-flake: wait for the pre-kill burst to be OBSERVED at the
|
||||
# server before severing (the sibling TestSpoolOnDisconnect test
|
||||
# already had this discipline). Killing mid-burst exercises a
|
||||
# DIFFERENT, documented limitation — the agent's one-line replay
|
||||
# margin cannot cover a multi-frame TCP in-flight window (an
|
||||
# ACK-protocol gap tracked for v0.4) — which made this test
|
||||
# nondeterministic under load instead of testing what its name
|
||||
# says: the reconnect flush of OFFLINE-spooled events.
|
||||
assert _wait_until(
|
||||
lambda: any(
|
||||
e["kind"] == "run_result" and e["seq"] == first["seq"]
|
||||
for e in fake_server.events
|
||||
),
|
||||
timeout_s=10.0,
|
||||
), "pre-kill burst never reached the server"
|
||||
|
||||
held_link.kill() # outage begins: no traffic, no reconnect possible
|
||||
assert test_agent.wait_disconnected(5)
|
||||
buffered = [test_agent.emit("activity", {"state": "offline", "n": n}) for n in range(3)]
|
||||
buffered_seqs = [e["seq"] for e in buffered]
|
||||
assert set(buffered_seqs).isdisjoint(e["seq"] for e in fake_server.events)
|
||||
|
||||
held_link.resume() # outage ends → supervisor reconnects, flushes spool
|
||||
# 20s deadline: under full-suite load the supervisor thread can be
|
||||
# starved past its normal sub-second reconnect; 8s was flaky.
|
||||
assert _wait_until(
|
||||
lambda: all(
|
||||
seq in {e["seq"] for e in fake_server.events} for seq in buffered_seqs
|
||||
),
|
||||
timeout_s=20.0,
|
||||
), "spooled events never flushed after reconnect"
|
||||
test_agent.stop()
|
||||
|
||||
replayed_wire_seqs = [e["seq"] for e in fake_server.events]
|
||||
# No loss: every emitted seq arrived, in ascending wire order (dedup
|
||||
# unique seqs — the one-line replay margin may repeat the pre-drop
|
||||
# tail event; the server dedups on (learner, task, seq)).
|
||||
unique_seqs = sorted(set(replayed_wire_seqs))
|
||||
assert unique_seqs == list(range(test_agent._seq)), (
|
||||
f"loss or gap: wire had {unique_seqs}, expected 0..{test_agent._seq - 1}"
|
||||
)
|
||||
positions = {seq: replayed_wire_seqs.index(seq) for seq in unique_seqs}
|
||||
assert [positions[s] for s in unique_seqs] == sorted(positions.values())
|
||||
# Dupes (if any) are confined to the replay margin: at or before the
|
||||
# first buffered (outage) seq.
|
||||
duped = {s for s in replayed_wire_seqs if replayed_wire_seqs.count(s) > 1}
|
||||
assert all(s <= min(buffered_seqs) for s in duped), f"unexpected dupes: {duped}"
|
||||
|
||||
|
||||
class TestSpoolResumeAfterRestart:
|
||||
def test_restart_recovers_seq_and_pending(self, tmp_path: Path) -> None:
|
||||
"""Simulates SIGKILL: no stop(), spool must carry state into next boot."""
|
||||
spool_path = tmp_path / "spool.jsonl"
|
||||
cfg = _config(tmp_path, "ws://127.0.0.1:1/x", spool_path) # unreachable
|
||||
first = agent.Agent(cfg)
|
||||
e0 = first.emit("activity", {"state": "boot"})
|
||||
e1 = first.emit("stdin", {"line": "ls"})
|
||||
del first # "killed" — no stop(), nothing flushed beyond per-append fsync
|
||||
|
||||
second = agent.Agent(cfg)
|
||||
try:
|
||||
assert second._seq == 2, f"seq must resume at 2, got {second._seq}"
|
||||
pending_seqs = [json.loads(line)["seq"] for line in second._pending]
|
||||
assert pending_seqs == [e0["seq"], e1["seq"]]
|
||||
finally:
|
||||
second.stop()
|
||||
# stop() with no connection preserves pending events on disk, plus the
|
||||
# final "stopped" activity marker.
|
||||
on_disk = agent.Spool(spool_path).read_all()
|
||||
assert len(on_disk) == 3
|
||||
assert json.loads(on_disk[-1])["payload"]["state"] == "stopped"
|
||||
|
||||
|
||||
class TestFileDiffWatcher:
|
||||
def test_workspace_changes_emitted(self, tmp_path: Path) -> None:
|
||||
test_agent = _make_agent(tmp_path, "ws://127.0.0.1:1/x") # offline: pure capture
|
||||
try:
|
||||
test_agent.start()
|
||||
target = test_agent.config.workspace / "hello.txt"
|
||||
target.write_text("v1\n")
|
||||
assert _wait_until(lambda: _find_diff(test_agent, "hello.txt", "created"))
|
||||
target.write_text("v1\nv2\n")
|
||||
assert _wait_until(lambda: _find_diff(test_agent, "hello.txt", "modified"))
|
||||
target.unlink()
|
||||
assert _wait_until(lambda: _find_diff(test_agent, "hello.txt", "deleted"))
|
||||
finally:
|
||||
test_agent.stop()
|
||||
|
||||
created = _find_diff(test_agent, "hello.txt", "created")
|
||||
assert created["payload"]["diff"] # unified diff present
|
||||
assert "v1" in created["payload"]["diff"]
|
||||
|
||||
# the spool dir itself must never be reported as a diff
|
||||
paths = [
|
||||
json.loads(line)["payload"].get("path")
|
||||
for line in test_agent._spool.read_all() # noqa: SLF001
|
||||
if json.loads(line)["kind"] == "file_diff"
|
||||
]
|
||||
assert all(p and not p.startswith(".nc-") for p in paths)
|
||||
|
||||
|
||||
def _find_diff(test_agent: agent.Agent, path: str, change: str) -> dict | None:
|
||||
for line in test_agent._spool.read_all(): # noqa: SLF001
|
||||
event = json.loads(line)
|
||||
if (
|
||||
event["kind"] == "file_diff"
|
||||
and event["payload"].get("path") == path
|
||||
and event["payload"].get("change") == change
|
||||
):
|
||||
return event
|
||||
return None
|
||||
|
||||
|
||||
class TestStdlibOnly:
|
||||
def test_no_third_party_imports(self) -> None:
|
||||
"""D-031: the agent runs in-namespace where only stdlib exists."""
|
||||
tree = ast.parse(AGENT_PATH.read_text(encoding="utf-8"))
|
||||
imported: set[str] = set()
|
||||
for node in ast.walk(tree):
|
||||
if isinstance(node, ast.Import):
|
||||
imported.update(alias.name.split(".")[0] for alias in node.names)
|
||||
elif isinstance(node, ast.ImportFrom) and node.level == 0:
|
||||
if node.module:
|
||||
imported.add(node.module.split(".")[0])
|
||||
third_party = imported - _ALLOWED_MODULES
|
||||
assert not third_party, (
|
||||
f"sandbox-agent.py must be stdlib-only (D-031); found: {sorted(third_party)}"
|
||||
)
|
||||
# sanity: the scan really saw the agent's core imports
|
||||
assert {"socket", "json", "threading", "ssl", "subprocess"} <= imported
|
||||
@@ -0,0 +1,203 @@
|
||||
"""End-to-end telemetry wiring (Task 2-3-01, REQ-3-003).
|
||||
|
||||
Real-namespace probe: a telemetry-wired sandbox (``create(..., task_id=...)``)
|
||||
spawns the stdlib capture agent, which dials a REAL uvicorn server over
|
||||
loopback (a TestClient app has no listening socket, so a true subprocess
|
||||
cannot reach it — this test must boot the actual ASGI server). Commands
|
||||
executed via the sandbox exec path produce ordered telemetry events that
|
||||
arrive at the live WS ingest endpoint and persist in SQLite.
|
||||
|
||||
Probe-guarded: skips (not fails) on hosts without user namespaces.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import contextlib
|
||||
import socket
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
import uvicorn
|
||||
|
||||
from ai_service.config import Settings
|
||||
from ai_service.main import create_app
|
||||
from ai_service.sandbox import SandboxManager
|
||||
from ai_service.sandbox.unshare_backend import UnshareBackend
|
||||
from ai_service.telemetry.ingest import TraceIntegrityMap
|
||||
from ai_service.telemetry.store import SQLiteTraceStore
|
||||
|
||||
from .test_isolation import USERSNS_AVAILABLE
|
||||
|
||||
|
||||
def _userns_probe() -> None:
|
||||
if not USERSNS_AVAILABLE:
|
||||
pytest.skip("user namespaces unavailable on this host (probe)")
|
||||
|
||||
|
||||
def _free_port() -> int:
|
||||
with socket.socket() as s:
|
||||
s.bind(("127.0.0.1", 0))
|
||||
return s.getsockname()[1]
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_task_sandbox_streams_events_to_ingest(tmp_path: Path) -> None:
|
||||
"""create(task_id=...) -> exec -> events land in SQLite in order (e2e)."""
|
||||
_userns_probe()
|
||||
|
||||
learner = "wiring-learner"
|
||||
task = "task-e2e-1"
|
||||
db = tmp_path / "wiring.db"
|
||||
store = SQLiteTraceStore(db_path=db)
|
||||
app = create_app()
|
||||
app.state.trace_store = store
|
||||
app.state.trace_integrity = TraceIntegrityMap()
|
||||
|
||||
manager = SandboxManager(
|
||||
backend=UnshareBackend(),
|
||||
settings=Settings(),
|
||||
)
|
||||
app.state.sandbox_manager = manager
|
||||
|
||||
port = _free_port()
|
||||
config = uvicorn.Config(app, host="127.0.0.1", port=port, log_level="warning")
|
||||
server = uvicorn.Server(config)
|
||||
serve_task = asyncio.get_running_loop().create_task(server.serve())
|
||||
try:
|
||||
deadline = time.monotonic() + 10.0
|
||||
while not server.started and time.monotonic() < deadline:
|
||||
await asyncio.sleep(0.05)
|
||||
assert server.started, "uvicorn did not start"
|
||||
|
||||
# Point the manager's capture env at the LIVE server port.
|
||||
manager._settings = Settings(telemetry_ingest_host="127.0.0.1") # noqa: SLF001
|
||||
orig_capture_env = manager._capture_env # noqa: SLF001
|
||||
|
||||
def _capture_env(sandbox_id: str, learner_id: str, task_id: str):
|
||||
env = orig_capture_env(sandbox_id, learner_id, task_id)
|
||||
env["NC_INGEST_URL"] = (
|
||||
f"ws://127.0.0.1:{port}/v1/telemetry/ingest"
|
||||
f"?learner_id={learner}&task_id={task}&sandbox_id={sandbox_id}"
|
||||
)
|
||||
return env
|
||||
|
||||
manager._capture_env = _capture_env # noqa: SLF001
|
||||
|
||||
handle = await manager.create(learner, task_id=task)
|
||||
try:
|
||||
live = manager._handles[handle.id] # noqa: SLF001
|
||||
result = await manager._backend.exec( # noqa: SLF001
|
||||
live, ["sh", "-c", "echo hello-telemetry && echo second-line"]
|
||||
)
|
||||
assert result.returncode == 0, result.stderr
|
||||
|
||||
# In the persistent-sandbox topology the agent observes execs
|
||||
# through its workspace watcher (the exec path is a separate
|
||||
# nsenter join), so the streaming kinds are activity + file_diff;
|
||||
# command/stdout kinds belong to the agent REPL (interactive use).
|
||||
events = await _await_events(store, learner, task, minimum=2)
|
||||
seqs = [e.seq for e in events]
|
||||
assert seqs == sorted(seqs), f"events out of order: {seqs}"
|
||||
kinds = {e.kind for e in events}
|
||||
assert kinds & {"activity", "file_diff", "command"}, kinds
|
||||
assert all(e.sandbox_id == handle.id for e in events), "sandbox_id mismatch"
|
||||
|
||||
# The learner path stays scoped: nothing for a different task.
|
||||
assert store.get_trace(learner, "some-other-task") == []
|
||||
finally:
|
||||
await manager.destroy(handle.id)
|
||||
finally:
|
||||
server.should_exit = True
|
||||
with contextlib.suppress(Exception):
|
||||
await asyncio.wait_for(serve_task, timeout=10.0)
|
||||
store.close()
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_no_task_id_means_no_capture(tmp_path: Path) -> None:
|
||||
"""Pure shell sandbox (task_id=None) spawns no capture agent (REQ-3-001 path)."""
|
||||
_userns_probe()
|
||||
|
||||
store = SQLiteTraceStore(db_path=tmp_path / "shell.db")
|
||||
backend = UnshareBackend()
|
||||
manager = SandboxManager(backend=backend, settings=Settings())
|
||||
|
||||
handle = await manager.create("shell-learner")
|
||||
try:
|
||||
assert handle.id not in backend._tracked # noqa: SLF001
|
||||
live = manager._handles[handle.id] # noqa: SLF001
|
||||
result = await backend.exec(live, ["sh", "-c", "echo plain"])
|
||||
assert "plain" in result.stdout
|
||||
finally:
|
||||
await manager.destroy(handle.id)
|
||||
store.close()
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_destroy_kills_inner_namespace_not_just_the_shim() -> None:
|
||||
"""Destroy must reap the ns-init, not only the `unshare --fork` shim.
|
||||
|
||||
Regression: `_reap(inner)` kills the unshare PARENT, but its forked child
|
||||
(PID 1 of the sandbox pid/mnt/net namespace, the `sleep`) reparents to
|
||||
host init and holds the tmpfs + workspace bind for the full sleep
|
||||
duration — a leaked sandbox per destroy (observed ~30 orphaned
|
||||
`sleep 3600` processes after one suite run; `--kill-child` did not reach
|
||||
it under this flag combo). The ns-init's host pid is `tracked.inner_pid`.
|
||||
"""
|
||||
_userns_probe()
|
||||
|
||||
backend = UnshareBackend()
|
||||
manager = SandboxManager(backend=backend, settings=Settings())
|
||||
|
||||
handle = await manager.create("lifecycle-learner", task_id="lifecycle-task")
|
||||
try:
|
||||
tracked = backend._tracked[handle.id] # noqa: SLF001
|
||||
assert tracked.agent is not None
|
||||
assert _proc_alive(tracked.agent.pid), "agent should live with the sandbox"
|
||||
assert _proc_alive(tracked.inner_pid), "ns-init should live with the sandbox"
|
||||
finally:
|
||||
await manager.destroy(handle.id)
|
||||
|
||||
# Reaping the unshare shim is asynchronous up to _reap's timeouts; poll
|
||||
# until both the agent AND the ns-init host pid are gone.
|
||||
deadline = time.monotonic() + 20.0
|
||||
while time.monotonic() < deadline and (
|
||||
_proc_alive(tracked.agent.pid) or _proc_alive(tracked.inner_pid)
|
||||
):
|
||||
await asyncio.sleep(0.25)
|
||||
assert not _proc_alive(tracked.agent.pid), "agent pid survived destroy (lifecycle)"
|
||||
assert not _proc_alive(tracked.inner_pid), (
|
||||
"inner namespace init survived destroy — leaked sandbox (tmpfs + bind "
|
||||
"held for a full hour; every destroy leaked one namespace process)"
|
||||
)
|
||||
|
||||
|
||||
def _proc_alive(pid: int | None) -> bool:
|
||||
if pid is None:
|
||||
return False
|
||||
try:
|
||||
with open(f"/proc/{pid}/stat") as fh:
|
||||
# state is the first field after comm: Z (zombie) counts as dead
|
||||
# for our purpose — it holds no namespace and is reaped next.
|
||||
return fh.read().rsplit(") ", 1)[1].split()[0] != "Z"
|
||||
except OSError:
|
||||
return False
|
||||
|
||||
|
||||
async def _await_events(
|
||||
store: SQLiteTraceStore, learner: str, task: str, *, minimum: int, timeout_s: float = 20.0
|
||||
) -> list:
|
||||
"""Poll the store until `minimum` events arrive (agent streams async)."""
|
||||
deadline = time.monotonic() + timeout_s
|
||||
while time.monotonic() < deadline:
|
||||
events = store.get_trace(learner, task)
|
||||
if len(events) >= minimum:
|
||||
return events
|
||||
await asyncio.sleep(0.25)
|
||||
events = store.get_trace(learner, task)
|
||||
pytest.fail(
|
||||
f"only {len(events)}/{minimum} telemetry events arrived within {timeout_s}s: "
|
||||
f"{[(e.seq, e.kind) for e in events]}"
|
||||
)
|
||||
@@ -0,0 +1,239 @@
|
||||
"""Dropped-connection durability probe (Task 2-4-01, REQ-3-003).
|
||||
|
||||
At-least-once delivery + server-side (learner,task,seq) dedup = every event
|
||||
stored exactly once, in order — proven against a REAL uvicorn ingest and a
|
||||
REAL namespace sandbox, with the connection severed mid-stream by a killable
|
||||
TCP proxy (deterministic "network down" window).
|
||||
|
||||
Probe-guarded: skips (not fails) on hosts without user namespaces.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import contextlib
|
||||
import socket
|
||||
import struct
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
import uvicorn
|
||||
|
||||
from ai_service.config import Settings
|
||||
from ai_service.main import create_app
|
||||
from ai_service.sandbox import SandboxManager
|
||||
from ai_service.sandbox.unshare_backend import UnshareBackend
|
||||
from ai_service.telemetry.ingest import TraceIntegrityMap
|
||||
from ai_service.telemetry.store import SQLiteTraceStore
|
||||
|
||||
from ..sandbox.test_isolation import USERSNS_AVAILABLE
|
||||
|
||||
|
||||
def _userns_probe() -> None:
|
||||
if not USERSNS_AVAILABLE:
|
||||
pytest.skip("user namespaces unavailable on this host (probe)")
|
||||
|
||||
|
||||
def _free_port() -> int:
|
||||
with socket.socket() as s:
|
||||
s.bind(("127.0.0.1", 0))
|
||||
return s.getsockname()[1]
|
||||
|
||||
|
||||
class KillableProxy:
|
||||
"""TCP forwarder whose data path dies on ``kill()`` — the listener stays.
|
||||
|
||||
Accepts are preserved so the agent's reconnects fail fast (connection
|
||||
reset) instead of hanging on a black-holed socket, keeping the outage
|
||||
window deterministic. ``revive()`` restores forwarding.
|
||||
"""
|
||||
|
||||
def __init__(self, target_port: int) -> None:
|
||||
self.target_port = target_port
|
||||
self.listen_port = _free_port()
|
||||
self._listener: socket.socket | None = None
|
||||
self._killed = False
|
||||
self._conns: list[socket.socket] = []
|
||||
|
||||
async def start(self) -> None:
|
||||
self._listener = socket.socket()
|
||||
self._listener.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
|
||||
self._listener.bind(("127.0.0.1", self.listen_port))
|
||||
self._listener.listen(16)
|
||||
self._listener.setblocking(False)
|
||||
loop = asyncio.get_running_loop()
|
||||
loop.create_task(self._accept_loop())
|
||||
|
||||
async def _accept_loop(self) -> None:
|
||||
loop = asyncio.get_running_loop()
|
||||
assert self._listener is not None
|
||||
while True:
|
||||
try:
|
||||
client, _ = await loop.sock_accept(self._listener)
|
||||
except OSError:
|
||||
return # listener closed
|
||||
client.setblocking(False) # accepted sockets default to blocking
|
||||
if self._killed:
|
||||
with contextlib.suppress(OSError):
|
||||
# RST on accept: reconnects fail fast while "down"
|
||||
client.setsockopt(
|
||||
socket.SOL_SOCKET, socket.SO_LINGER,
|
||||
struct.pack("<ii", 1, 0),
|
||||
)
|
||||
client.close()
|
||||
continue
|
||||
upstream = socket.socket()
|
||||
upstream.setblocking(False)
|
||||
try:
|
||||
await loop.sock_connect(upstream, ("127.0.0.1", self.target_port))
|
||||
except OSError:
|
||||
with contextlib.suppress(OSError):
|
||||
client.close()
|
||||
continue
|
||||
self._conns.extend((client, upstream))
|
||||
loop.create_task(self._pump(client, upstream))
|
||||
loop.create_task(self._pump(upstream, client))
|
||||
|
||||
async def _pump(self, src: socket.socket, dst: socket.socket) -> None:
|
||||
"""Forward until EOF, error, or the killed flag is observed.
|
||||
|
||||
Only ``dst`` is closed here — the reverse-direction pump owns ``src``
|
||||
(closing the peer's socket from this task causes EBADF cascades).
|
||||
"""
|
||||
loop = asyncio.get_running_loop()
|
||||
try:
|
||||
while True:
|
||||
data = await loop.sock_recv(src, 65536)
|
||||
if not data:
|
||||
return
|
||||
if self._killed:
|
||||
return # outage: drop the relay, peer sees EOF on dst close
|
||||
await loop.sock_sendall(dst, data)
|
||||
except OSError:
|
||||
return
|
||||
finally:
|
||||
with contextlib.suppress(OSError):
|
||||
dst.close()
|
||||
|
||||
def kill(self) -> None:
|
||||
"""Sever the data path — flag only.
|
||||
|
||||
The pumps and the accept loop observe ``_killed`` themselves and close
|
||||
THEIR OWN sockets from inside the event loop. Closing sockets that the
|
||||
loop is currently awaiting from the outside deadlocks the loop's
|
||||
child-watcher in this environment (observed: post-kill subprocess
|
||||
spawns hung forever) — so no external socket surgery, ever.
|
||||
"""
|
||||
|
||||
def revive(self) -> None:
|
||||
self._killed = False
|
||||
|
||||
|
||||
async def _await_events(
|
||||
store: SQLiteTraceStore, learner: str, task: str, *, minimum: int, timeout_s: float = 30.0
|
||||
) -> list:
|
||||
deadline = time.monotonic() + timeout_s
|
||||
while time.monotonic() < deadline:
|
||||
events = store.get_trace(learner, task)
|
||||
if len(events) >= minimum:
|
||||
return events
|
||||
await asyncio.sleep(0.25)
|
||||
return store.get_trace(learner, task)
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_disconnect_reconnect_loses_nothing(tmp_path: Path) -> None:
|
||||
"""Sever the agent's WS mid-stream; every event lands exactly once, ordered."""
|
||||
_userns_probe()
|
||||
|
||||
learner = "durability-learner"
|
||||
task = "task-durability"
|
||||
|
||||
store = SQLiteTraceStore(db_path=tmp_path / "durability.db")
|
||||
app = create_app()
|
||||
app.state.trace_store = store
|
||||
app.state.trace_integrity = TraceIntegrityMap()
|
||||
|
||||
server_port = _free_port()
|
||||
server = uvicorn.Server(
|
||||
uvicorn.Config(app, host="127.0.0.1", port=server_port, log_level="warning")
|
||||
)
|
||||
serve_task = asyncio.get_running_loop().create_task(server.serve())
|
||||
|
||||
proxy = KillableProxy(target_port=server_port)
|
||||
await proxy.start()
|
||||
|
||||
manager = SandboxManager(backend=UnshareBackend(), settings=Settings())
|
||||
app.state.sandbox_manager = manager
|
||||
|
||||
def capture_env(sandbox_id: str, learner_id: str, task_id: str) -> dict[str, str]:
|
||||
return {
|
||||
"NC_LEARNER_ID": learner_id,
|
||||
"NC_TASK_ID": task_id,
|
||||
"NC_SANDBOX_ID": sandbox_id,
|
||||
# Agent dials the PROXY; the proxy forwards to the real server.
|
||||
"NC_INGEST_URL": (
|
||||
f"ws://127.0.0.1:{proxy.listen_port}/v1/telemetry/ingest"
|
||||
f"?learner_id={learner_id}&task_id={task_id}&sandbox_id={sandbox_id}"
|
||||
),
|
||||
# Fast reconnect so the probe stays bounded.
|
||||
"NC_BACKOFF_BASE_S": "0.1",
|
||||
"NC_BACKOFF_MAX_S": "0.5",
|
||||
}
|
||||
|
||||
manager._capture_env = capture_env # noqa: SLF001 - test seam
|
||||
|
||||
try:
|
||||
for _ in range(100):
|
||||
if server.started:
|
||||
break
|
||||
await asyncio.sleep(0.1)
|
||||
assert server.started, "uvicorn did not start"
|
||||
|
||||
handle = await manager.create(learner, task_id=task)
|
||||
try:
|
||||
live = manager._handles[handle.id] # noqa: SLF001 - test seam
|
||||
backend: UnshareBackend = manager._backend # noqa: SLF001 - test seam
|
||||
|
||||
# Phase A — connected: a write streams through.
|
||||
result = await backend.exec(live, ["sh", "-c", "echo one > a.txt"])
|
||||
assert result.returncode == 0, result.stderr
|
||||
events = await _await_events(store, learner, task, minimum=1)
|
||||
assert events, "no events arrived before the outage"
|
||||
|
||||
# Phase B — network down: sever mid-stream, keep generating.
|
||||
proxy.kill()
|
||||
for n in ("two", "three", "four"):
|
||||
result = await backend.exec(live, ["sh", "-c", f"echo {n} > {n}.txt"])
|
||||
assert result.returncode == 0, result.stderr
|
||||
await asyncio.sleep(0.3)
|
||||
|
||||
# Phase C — revive: the agent reconnects (fast backoff) and
|
||||
# flushes the spool. Server dedups on (learner,task,seq).
|
||||
proxy.revive()
|
||||
await asyncio.sleep(1.5)
|
||||
for n in ("five", "six"):
|
||||
result = await backend.exec(live, ["sh", "-c", f"echo {n} > {n}.txt"])
|
||||
assert result.returncode == 0, result.stderr
|
||||
|
||||
events = await _await_events(
|
||||
store, learner, task, minimum=6, timeout_s=30.0
|
||||
)
|
||||
seqs = [e.seq for e in events]
|
||||
assert seqs == sorted(seqs), f"out of order after reconnect: {seqs}"
|
||||
assert len(set(seqs)) == len(seqs), f"duplicates stored: {seqs}"
|
||||
assert len(events) >= 6, f"lost events across the outage: {seqs}"
|
||||
# Exactly-once storage despite at-least-once delivery: every spool
|
||||
# flush may re-send the replay-margin line; dedup collapses it.
|
||||
stored_files = {
|
||||
e.payload.get("path") for e in events if e.kind == "file_diff"
|
||||
}
|
||||
assert stored_files, "file_diff events missing from the stored trace"
|
||||
finally:
|
||||
await manager.destroy(handle.id)
|
||||
finally:
|
||||
server.should_exit = True
|
||||
with contextlib.suppress(Exception):
|
||||
await asyncio.wait_for(serve_task, timeout=10.0)
|
||||
store.close()
|
||||
@@ -0,0 +1,126 @@
|
||||
"""TelemetryEvent model contract tests (REQ-3-003, D-027).
|
||||
|
||||
Pure validation tests — no store, no network, no fixtures beyond a
|
||||
canonical kwargs builder.
|
||||
"""
|
||||
|
||||
import json
|
||||
from datetime import UTC, datetime
|
||||
|
||||
import pytest
|
||||
from pydantic import ValidationError
|
||||
|
||||
from ai_service.telemetry import EventKind, TelemetryEvent, TraceSpan
|
||||
|
||||
|
||||
def make_event_kwargs(**overrides) -> dict:
|
||||
base = {
|
||||
"learner_id": "learner-1",
|
||||
"task_id": "task-1",
|
||||
"seq": 0,
|
||||
"kind": "command",
|
||||
"payload": {"argv": ["pytest"], "cwd": "/workspace"},
|
||||
"ts": datetime(2026, 9, 11, tzinfo=UTC),
|
||||
"sandbox_id": "sbx-1",
|
||||
}
|
||||
base.update(overrides)
|
||||
return base
|
||||
|
||||
|
||||
def make_event(**overrides) -> TelemetryEvent:
|
||||
return TelemetryEvent(**make_event_kwargs(**overrides))
|
||||
|
||||
|
||||
def test_valid_event_constructs_with_all_kinds():
|
||||
for kind in (
|
||||
"command",
|
||||
"file_diff",
|
||||
"run_result",
|
||||
"test_result",
|
||||
"activity",
|
||||
"stdin",
|
||||
"stdout",
|
||||
):
|
||||
event = make_event(kind=kind)
|
||||
assert event.kind == kind
|
||||
|
||||
|
||||
def test_zero_seq_accepted():
|
||||
assert make_event(seq=0).seq == 0
|
||||
|
||||
|
||||
def test_negative_seq_rejected():
|
||||
# @validates hooks raise ValueError (not pydantic ValidationError) at
|
||||
# construction — same for kind and empty-id checks below.
|
||||
with pytest.raises(ValueError, match="seq must be >= 0"):
|
||||
make_event(seq=-1)
|
||||
|
||||
|
||||
def test_bad_kind_rejected():
|
||||
with pytest.raises(ValueError, match="unknown event kind"):
|
||||
make_event(kind="keystroke") # not in the EventKind literal
|
||||
|
||||
|
||||
def test_empty_learner_id_rejected():
|
||||
with pytest.raises(ValueError, match="non-empty identifier"):
|
||||
make_event(learner_id="")
|
||||
|
||||
|
||||
def test_empty_task_id_rejected():
|
||||
with pytest.raises(ValueError, match="non-empty identifier"):
|
||||
make_event(task_id="")
|
||||
|
||||
|
||||
def test_payload_json_roundtrip():
|
||||
payload = {
|
||||
"command": "pytest -q",
|
||||
"exit_code": 1,
|
||||
"durations": [0.12, 3.4],
|
||||
"nested": {"passed": 7, "failed": 2},
|
||||
"unicode": "héllo",
|
||||
}
|
||||
event = make_event(payload=payload)
|
||||
assert event.payload == payload
|
||||
# JSON-serializable payloads survive a full dumps/loads roundtrip.
|
||||
assert json.loads(json.dumps(event.payload)) == payload
|
||||
|
||||
|
||||
def test_event_json_roundtrip():
|
||||
event = make_event(payload={"stdout": "ok", "n": 3})
|
||||
restored = TelemetryEvent.model_validate_json(event.model_dump_json())
|
||||
# pydantic leaves a JSON-str ts as str; compare field-by-field instead of
|
||||
# dataclass equality, which distinguishes 'Z' string vs parsed datetime.
|
||||
assert restored.learner_id == event.learner_id
|
||||
assert restored.task_id == event.task_id
|
||||
assert restored.seq == event.seq
|
||||
assert restored.kind == event.kind
|
||||
assert restored.payload == event.payload
|
||||
assert restored.sandbox_id == event.sandbox_id
|
||||
# SQLModel leaves a JSON-serialized ts as its str form; parse to compare.
|
||||
restored_ts = datetime.fromisoformat(str(restored.ts).replace("Z", "+00:00"))
|
||||
assert restored_ts == event.ts
|
||||
|
||||
|
||||
def test_tracespan_holds_ordered_events():
|
||||
events = tuple(make_event(seq=seq, kind="file_diff") for seq in range(3))
|
||||
span = TraceSpan(learner_id="learner-1", task_id="task-1", events=events)
|
||||
assert [e.seq for e in span.events] == [0, 1, 2]
|
||||
assert span.latest_seq == 2
|
||||
|
||||
|
||||
def test_tracespan_empty_has_no_latest_seq():
|
||||
span = TraceSpan(learner_id="learner-1", task_id="task-1")
|
||||
assert span.events == ()
|
||||
assert span.latest_seq == -1
|
||||
|
||||
|
||||
def test_tracespan_rejects_empty_ids():
|
||||
with pytest.raises(ValidationError):
|
||||
TraceSpan(learner_id="", task_id="task-1")
|
||||
with pytest.raises(ValidationError):
|
||||
TraceSpan(learner_id="learner-1", task_id="")
|
||||
|
||||
|
||||
def test_event_kind_type_is_exported_literal():
|
||||
# The kind discriminator is part of the public module surface.
|
||||
assert "command" in EventKind.__args__
|
||||
@@ -0,0 +1,206 @@
|
||||
"""SQLiteTraceStore tests (REQ-3-003, D-027).
|
||||
|
||||
Each test gets its own tmp-path SQLite file — no shared disk state. Covers:
|
||||
- append / ordered get (events may arrive out of order; reads come back
|
||||
ordered by seq)
|
||||
- dedup on retry: the same (learner_id, task_id, seq) appended twice is
|
||||
stored exactly once (at-least-once ingest contract)
|
||||
- retry dedup must not overwrite the originally stored row
|
||||
- gap detection, latest_seq, list_tasks
|
||||
- WAL + synchronous=NORMAL pragmas actually applied to the DB file
|
||||
- concurrent writer + reader against the same DB file (a-3 smoke test)
|
||||
"""
|
||||
|
||||
import concurrent.futures
|
||||
import threading
|
||||
from datetime import UTC, datetime, timedelta
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
import pytest
|
||||
import sqlalchemy as sa
|
||||
|
||||
from ai_service.telemetry.models import EventKind, TelemetryEvent
|
||||
from ai_service.telemetry.store import SQLiteTraceStore
|
||||
|
||||
_BASE_TS = datetime(2026, 9, 11, 12, 0, 0, tzinfo=UTC)
|
||||
|
||||
|
||||
def make_event(
|
||||
seq: int,
|
||||
learner_id: str = "learner-1",
|
||||
task_id: str = "task-1",
|
||||
kind: EventKind = "command",
|
||||
payload: dict[str, Any] | None = None,
|
||||
sandbox_id: str = "sbx-1",
|
||||
) -> TelemetryEvent:
|
||||
return TelemetryEvent(
|
||||
learner_id=learner_id,
|
||||
task_id=task_id,
|
||||
seq=seq,
|
||||
kind=kind,
|
||||
payload=payload if payload is not None else {"seq": seq},
|
||||
ts=_BASE_TS + timedelta(seconds=seq),
|
||||
sandbox_id=sandbox_id,
|
||||
)
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def store(tmp_path: Path):
|
||||
s = SQLiteTraceStore(db_path=tmp_path / "telemetry.db")
|
||||
yield s
|
||||
s.close()
|
||||
|
||||
|
||||
def test_append_and_get_trace_orders_by_seq(store: SQLiteTraceStore) -> None:
|
||||
# Append out of order; reads must come back ordered by seq.
|
||||
store.append(make_event(2, kind="stdout"))
|
||||
store.append(make_event(0, kind="stdin"))
|
||||
store.append(make_event(1, kind="run_result"))
|
||||
|
||||
trace = store.get_trace("learner-1", "task-1")
|
||||
assert [e.seq for e in trace] == [0, 1, 2]
|
||||
assert [e.kind for e in trace] == ["stdin", "run_result", "stdout"]
|
||||
assert all(e.learner_id == "learner-1" for e in trace)
|
||||
assert all(e.task_id == "task-1" for e in trace)
|
||||
|
||||
|
||||
def test_get_trace_round_trips_fields(store: SQLiteTraceStore) -> None:
|
||||
event = make_event(0, payload={"file": "a.py", "nested": {"ok": True}})
|
||||
store.append(event)
|
||||
|
||||
(row,) = store.get_trace("learner-1", "task-1")
|
||||
assert row.learner_id == event.learner_id
|
||||
assert row.task_id == event.task_id
|
||||
assert row.seq == 0
|
||||
assert row.kind == "command"
|
||||
assert row.payload == {"file": "a.py", "nested": {"ok": True}}
|
||||
assert row.ts == _BASE_TS
|
||||
assert row.sandbox_id == "sbx-1"
|
||||
|
||||
|
||||
def test_get_trace_is_scoped_to_the_task_pair(store: SQLiteTraceStore) -> None:
|
||||
store.append(make_event(0, learner_id="learner-1", task_id="task-1"))
|
||||
store.append(make_event(0, learner_id="learner-1", task_id="task-2"))
|
||||
store.append(make_event(0, learner_id="learner-2", task_id="task-1"))
|
||||
|
||||
assert [e.seq for e in store.get_trace("learner-1", "task-1")] == [0]
|
||||
assert len(store.get_trace("learner-1", "task-2")) == 1
|
||||
assert len(store.get_trace("learner-2", "task-1")) == 1
|
||||
assert store.get_trace("learner-1", "task-missing") == []
|
||||
|
||||
|
||||
def test_append_is_idempotent_on_retry(store: SQLiteTraceStore) -> None:
|
||||
# At-least-once ingest re-delivers the same event (same idempotency key).
|
||||
# It must be stored exactly once and the retry must be a no-op success.
|
||||
original = make_event(0, kind="command", payload={"attempt": 1})
|
||||
store.append(original)
|
||||
store.append(original)
|
||||
# A distinct-but-conflicting retry (same key, different body) is also
|
||||
# deduped — the first stored row wins, no overwrite.
|
||||
retry = make_event(0, kind="file_diff", payload={"attempt": 2})
|
||||
store.append(retry)
|
||||
|
||||
trace = store.get_trace("learner-1", "task-1")
|
||||
assert len(trace) == 1
|
||||
assert trace[0].kind == "command"
|
||||
assert trace[0].payload == {"attempt": 1}
|
||||
assert store.latest_seq("learner-1", "task-1") == 0
|
||||
|
||||
|
||||
def test_gaps_reports_missing_seqs(store: SQLiteTraceStore) -> None:
|
||||
for seq in (0, 2, 3, 7):
|
||||
store.append(make_event(seq))
|
||||
|
||||
assert store.gaps("learner-1", "task-1") == [1, 4, 5, 6]
|
||||
# Gaps are per-trace: an empty trace has no gaps at all.
|
||||
assert store.gaps("learner-1", "task-unknown") == []
|
||||
|
||||
|
||||
def test_latest_seq(store: SQLiteTraceStore) -> None:
|
||||
assert store.latest_seq("learner-1", "task-1") == -1
|
||||
|
||||
store.append(make_event(0))
|
||||
assert store.latest_seq("learner-1", "task-1") == 0
|
||||
|
||||
store.append(make_event(5)) # gaps do not move latest_seq
|
||||
assert store.latest_seq("learner-1", "task-1") == 5
|
||||
# Scoped to the trace pair.
|
||||
assert store.latest_seq("learner-1", "task-2") == -1
|
||||
|
||||
|
||||
def test_count_is_durable_row_count_not_latest_seq(store: SQLiteTraceStore) -> None:
|
||||
"""count() backs the ingest flood cap (P7): it must reflect stored ROWS
|
||||
(a skipped-ahead seq must not burn un-sent budget) and stay O(1)-ish
|
||||
(COUNT(*), never materialize the trace per append)."""
|
||||
assert store.count("learner-1", "task-1") == 0
|
||||
store.append(make_event(0))
|
||||
store.append(make_event(2)) # skipped 1 — count is rows, not latest+1
|
||||
assert store.count("learner-1", "task-1") == 2
|
||||
# Dedup retries do not inflate the count (at-least-once contract).
|
||||
store.append(make_event(2))
|
||||
assert store.count("learner-1", "task-1") == 2
|
||||
# Scoped to the trace pair.
|
||||
assert store.count("learner-1", "task-2") == 0
|
||||
|
||||
|
||||
def test_list_tasks(store: SQLiteTraceStore) -> None:
|
||||
assert store.list_tasks("learner-1") == []
|
||||
|
||||
store.append(make_event(0, task_id="task-b"))
|
||||
store.append(make_event(1, task_id="task-b"))
|
||||
store.append(make_event(0, task_id="task-a"))
|
||||
|
||||
assert store.list_tasks("learner-1") == ["task-a", "task-b"]
|
||||
# Scoped per learner.
|
||||
assert store.list_tasks("learner-2") == []
|
||||
|
||||
|
||||
def test_pragmas_are_applied(store: SQLiteTraceStore) -> None:
|
||||
# Pragmas are per-connection; query through the store's engine so the
|
||||
# connect hook (not a default sqlite3 connection) is what we inspect.
|
||||
with store._engine.connect() as conn:
|
||||
(journal_mode,) = conn.execute(sa.text("PRAGMA journal_mode")).one()
|
||||
(synchronous,) = conn.execute(sa.text("PRAGMA synchronous")).one()
|
||||
assert journal_mode == "wal"
|
||||
# synchronous=NORMAL is 1 in SQLite's pragma numbering.
|
||||
assert synchronous == 1
|
||||
|
||||
|
||||
def test_concurrent_writer_and_reader_no_database_is_locked(tmp_path: Path) -> None:
|
||||
"""One thread appends while another reads in a tight loop (a-3).
|
||||
|
||||
Without WAL + busy_timeout this pattern reliably produces
|
||||
`OperationalError: database is locked` on SQLite. The assertion is that
|
||||
every reader call completes and the final trace is complete.
|
||||
"""
|
||||
db_path = tmp_path / "telemetry.db"
|
||||
n_events = 60
|
||||
stop_writing = threading.Event()
|
||||
|
||||
writer = SQLiteTraceStore(db_path=db_path)
|
||||
reader = SQLiteTraceStore(db_path=db_path)
|
||||
try:
|
||||
|
||||
def write_events() -> None:
|
||||
for seq in range(n_events):
|
||||
writer.append(make_event(seq))
|
||||
stop_writing.set()
|
||||
|
||||
def read_trace() -> None:
|
||||
while not stop_writing.is_set():
|
||||
reader.get_trace("learner-1", "task-1")
|
||||
# Final read after the writer is done.
|
||||
assert len(reader.get_trace("learner-1", "task-1")) == n_events
|
||||
|
||||
with concurrent.futures.ThreadPoolExecutor(max_workers=2) as pool:
|
||||
futures = [pool.submit(write_events), pool.submit(read_trace)]
|
||||
for future in futures:
|
||||
future.result(timeout=30)
|
||||
|
||||
assert [e.seq for e in writer.get_trace("learner-1", "task-1")] == list(
|
||||
range(n_events)
|
||||
)
|
||||
finally:
|
||||
reader.close()
|
||||
writer.close()
|
||||
@@ -0,0 +1,79 @@
|
||||
"""Corpus dormancy verification (Task 6-1-04, REQ-3-007).
|
||||
|
||||
The v0.2 mock engine inputs (`corpus/telemetry.py`, `corpus/artifacts.py`)
|
||||
must have ZERO production importers after the v0.3 re-grounding: Lab/
|
||||
Assessor/Proctor run on real engine inputs with no mock fallback in the
|
||||
learner path. The files stay on disk (Phase-3 calibration history) but are
|
||||
not imported by any production module. `learner_context` remains ACTIVE
|
||||
(agents still need learner context). Test-only references (e.g.
|
||||
`corpus/trace_fixtures.py` in grading calibration tests) are allowed.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import ast
|
||||
from pathlib import Path
|
||||
|
||||
REPO = Path(__file__).parents[1]
|
||||
|
||||
#: Production trees whose imports of dormant corpus modules are forbidden.
|
||||
PRODUCTION_PATHS = [
|
||||
REPO / "ai_service" / "agents",
|
||||
REPO / "ai_service" / "api",
|
||||
REPO / "ai_service" / "grading",
|
||||
REPO / "ai_service" / "telemetry",
|
||||
REPO / "ai_service" / "variants",
|
||||
REPO / "ai_service" / "voice",
|
||||
REPO / "ai_service" / "sandbox",
|
||||
REPO / "ai_service" / "llm",
|
||||
REPO / "ai_service" / "main.py",
|
||||
]
|
||||
|
||||
DORMANT_MODULES = ("corpus.telemetry", "corpus.artifacts")
|
||||
|
||||
|
||||
def _module_targets(node: ast.AST, *, level: int, module: str | None) -> set[str]:
|
||||
"""Resolve relative + absolute import targets to dotted ai_service paths."""
|
||||
targets: set[str] = set()
|
||||
if module and ("corpus" in module):
|
||||
targets.add(module)
|
||||
return targets
|
||||
|
||||
|
||||
def test_no_production_imports_of_dormant_corpus() -> None:
|
||||
violators: list[str] = []
|
||||
for path in PRODUCTION_PATHS:
|
||||
files = [path] if path.suffix == ".py" else sorted(path.rglob("*.py"))
|
||||
for py in files:
|
||||
tree = ast.parse(py.read_text())
|
||||
for node in ast.walk(tree):
|
||||
if isinstance(node, ast.ImportFrom):
|
||||
module = node.module or ""
|
||||
level = node.level
|
||||
if level: # relative: resolve against corpus
|
||||
if module.endswith("telemetry") and "corpus" in module:
|
||||
violators.append(f"{py}: {module}")
|
||||
if module.endswith("artifacts") and "corpus" in module:
|
||||
violators.append(f"{py}: {module}")
|
||||
# bare `from . import telemetry` inside corpus/ itself is fine
|
||||
else:
|
||||
for target in _module_targets(node, level=level, module=module):
|
||||
violators.append(f"{py}: {target}")
|
||||
elif isinstance(node, ast.Import):
|
||||
for alias in node.names:
|
||||
if any(alias.name.startswith(m) for m in DORMANT_MODULES):
|
||||
violators.append(f"{py}: {alias.name}")
|
||||
assert not violators, f"dormant corpus imports in production: {violators}"
|
||||
|
||||
|
||||
def test_learner_context_stays_active() -> None:
|
||||
"""Learner context corpus is NOT dormant — agents still use it."""
|
||||
agents_lab = (REPO / "ai_service" / "agents" / "lab.py").read_text()
|
||||
assert "corpus.learner_context" in agents_lab
|
||||
(REPO / "ai_service" / "corpus" / "learner_context.py").exists()
|
||||
|
||||
|
||||
def test_dormancy_headers_present() -> None:
|
||||
for fname in ("telemetry.py", "artifacts.py"):
|
||||
src = (REPO / "ai_service" / "corpus" / fname).read_text()
|
||||
assert "DORMANT" in src, f"{fname} missing dormancy header"
|
||||
@@ -0,0 +1 @@
|
||||
"""Variant task generation tests (REQ-3-005)."""
|
||||
@@ -0,0 +1,179 @@
|
||||
"""Variant generator tests (Task 4-2-01, REQ-3-005) — D-029 + a-5 binding."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from datetime import UTC, datetime, timedelta
|
||||
|
||||
import pytest
|
||||
|
||||
from ai_service.grading.features import compute_digest
|
||||
from ai_service.llm.mock import MockProvider
|
||||
from ai_service.telemetry.models import TelemetryEvent
|
||||
from ai_service.variants.generator import (
|
||||
MILESTONE,
|
||||
VariantGenerator,
|
||||
derive_seed,
|
||||
derive_task_id,
|
||||
)
|
||||
from ai_service.variants.store import SQLiteVariantStore, VariantRecord
|
||||
from ai_service.variants.templates import TEMPLATES, get_template
|
||||
|
||||
|
||||
class ScriptedRenderProvider(MockProvider):
|
||||
"""Deterministic render: the statement embeds the params (distinct per draw)."""
|
||||
|
||||
def __init__(self) -> None:
|
||||
super().__init__()
|
||||
self.calls = 0
|
||||
|
||||
async def chat(self, messages, model, response_format=None): # noqa: ANN001
|
||||
self.calls += 1
|
||||
import json
|
||||
|
||||
user = next(m.content for m in reversed(messages) if m.role == "user")
|
||||
# Distinct per distinct params: hash the seeded slot lines.
|
||||
fingerprint = abs(hash(user)) % 10_000
|
||||
return json.dumps({"statement": f"Scripted variant #{fingerprint} — build it."})
|
||||
|
||||
|
||||
class FailingRenderProvider(MockProvider):
|
||||
"""Always fails D-020 validation -> deterministic fallback path."""
|
||||
|
||||
def __init__(self) -> None:
|
||||
super().__init__()
|
||||
self.calls = 0
|
||||
|
||||
async def chat(self, messages, model, response_format=None): # noqa: ANN001
|
||||
self.calls += 1
|
||||
return "this is not json at all"
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def store(tmp_path): # noqa: ANN001
|
||||
s = SQLiteVariantStore(db_path=tmp_path / "variants.db")
|
||||
yield s
|
||||
s.close()
|
||||
|
||||
|
||||
async def test_two_learners_distinct_statements(store) -> None: # noqa: ANN001
|
||||
provider = ScriptedRenderProvider()
|
||||
gen = VariantGenerator(store, provider, model="mock")
|
||||
a = await gen.generate("learner-a", "tpl-llm-judge")
|
||||
b = await gen.generate("learner-b", "tpl-llm-judge")
|
||||
assert a.statement != b.statement
|
||||
assert a.seed != b.seed
|
||||
assert a.task_id != b.task_id
|
||||
|
||||
|
||||
async def test_same_learner_is_cached_no_second_llm_call(store) -> None: # noqa: ANN001
|
||||
provider = ScriptedRenderProvider()
|
||||
gen = VariantGenerator(store, provider, model="mock")
|
||||
first = await gen.generate("learner-a", "tpl-llm-judge")
|
||||
calls_after_first = provider.calls
|
||||
second = await gen.generate("learner-a", "tpl-llm-judge")
|
||||
assert first == second # identical stored variant (D-029 reproducible)
|
||||
assert provider.calls == calls_after_first # cache hit: NO LLM call
|
||||
|
||||
|
||||
async def test_seed_derivation_reproducible() -> None:
|
||||
s1 = derive_seed("tpl-llm-judge", "learner-a", MILESTONE)
|
||||
s2 = derive_seed("tpl-llm-judge", "learner-a", MILESTONE)
|
||||
assert s1 == s2
|
||||
assert derive_seed("tpl-llm-judge", "learner-b", MILESTONE) != s1
|
||||
assert derive_task_id(s1).startswith("task-")
|
||||
assert len(derive_task_id(s1)) == len("task-") + 16
|
||||
|
||||
|
||||
async def test_params_are_schema_valid(store) -> None: # noqa: ANN001
|
||||
gen = VariantGenerator(store, ScriptedRenderProvider(), model="mock")
|
||||
record = await gen.generate("learner-a", "tpl-guardrail-schema")
|
||||
template = get_template("tpl-guardrail-schema")
|
||||
for slot in template.slots: # type: ignore[union-attr]
|
||||
value = record.params[slot.name]
|
||||
assert slot.validate_value(value), f"slot {slot.name} drew invalid {value!r}"
|
||||
|
||||
|
||||
async def test_llm_failure_falls_back_deterministically(store) -> None: # noqa: ANN001
|
||||
provider = FailingRenderProvider()
|
||||
gen = VariantGenerator(store, provider, model="mock")
|
||||
record = await gen.generate("learner-a", "tpl-rag-chunker")
|
||||
template = get_template("tpl-rag-chunker")
|
||||
expected = template.render({k: v for k, v in record.params.items()}) # type: ignore
|
||||
assert record.statement == expected # skeleton render, seed-auditable
|
||||
assert provider.calls == 2 # D-020 bounded retry, then fallback
|
||||
|
||||
|
||||
async def test_unknown_template_raises(store) -> None: # noqa: ANN001
|
||||
gen = VariantGenerator(store, ScriptedRenderProvider(), model="mock")
|
||||
with pytest.raises(ValueError, match="no task template"):
|
||||
await gen.generate("learner-a", "tpl-does-not-exist")
|
||||
|
||||
|
||||
def test_fairness_envelope_same_bar_per_template() -> None:
|
||||
"""a-5 (BINDING): every legal variant of one template fits the anchors.
|
||||
|
||||
For 10 different learners: draw the seeded params, then synthesize a
|
||||
trace whose edit count is sampled INSIDE the template's anchor band and
|
||||
whose test runs meet the anchor minimum — the resulting digests must
|
||||
all sit within the template's expected feature envelope. That is the
|
||||
testable form of "same bar": no slot draw can push a variant outside
|
||||
the effort band the grader context assumes.
|
||||
"""
|
||||
import random
|
||||
|
||||
for template in TEMPLATES.values():
|
||||
anchors = template.rubric_anchors
|
||||
lo_edits, hi_edits = anchors.expected_edit_count_band
|
||||
t0 = datetime(2026, 9, 12, tzinfo=UTC)
|
||||
for i in range(10):
|
||||
params = template.sample_params(seed=10_000 + i)
|
||||
rng = random.Random(i)
|
||||
n_edits = rng.randint(lo_edits, hi_edits)
|
||||
events = [
|
||||
TelemetryEvent(
|
||||
learner_id=f"fair-learner-{i}",
|
||||
task_id=f"fair-task-{i}",
|
||||
seq=n,
|
||||
kind="file_diff",
|
||||
payload={"path": f"f{n}.py"},
|
||||
ts=t0 + timedelta(seconds=n * 10),
|
||||
sandbox_id="sbx-fair",
|
||||
)
|
||||
for n in range(n_edits)
|
||||
]
|
||||
# Meet the anchor's minimum test-run expectation.
|
||||
for t in range(anchors.expected_min_test_runs):
|
||||
events.append(
|
||||
TelemetryEvent(
|
||||
learner_id=f"fair-learner-{i}",
|
||||
task_id=f"fair-task-{i}",
|
||||
seq=len(events),
|
||||
kind="test_result",
|
||||
payload={"passed": t == anchors.expected_min_test_runs - 1},
|
||||
ts=t0 + timedelta(seconds=(n_edits + t) * 10),
|
||||
sandbox_id="sbx-fair",
|
||||
)
|
||||
)
|
||||
digest = compute_digest(events)
|
||||
assert lo_edits <= digest.edit_count <= hi_edits
|
||||
assert digest.test_pass_count + digest.test_fail_count >= (
|
||||
anchors.expected_min_test_runs
|
||||
)
|
||||
# Slot values never appear in the digest (no scenario leakage into
|
||||
# grading features — difficulty stays scenario-independent).
|
||||
digest_json = digest.model_dump_json()
|
||||
for value in params.values():
|
||||
assert str(value) not in digest_json or isinstance(value, int)
|
||||
|
||||
|
||||
async def test_variant_record_roundtrips_through_store(store) -> None: # noqa: ANN001
|
||||
gen = VariantGenerator(store, ScriptedRenderProvider(), model="mock")
|
||||
record = await gen.generate("learner-a", "tpl-llm-judge")
|
||||
fetched = store.get("learner-a", "tpl-llm-judge")
|
||||
assert fetched is not None
|
||||
assert fetched.statement == record.statement
|
||||
assert fetched.seed == record.seed
|
||||
by_task = store.get_by_task(record.task_id)
|
||||
assert by_task is not None
|
||||
assert by_task.learner_id == "learner-a"
|
||||
assert isinstance(record, VariantRecord)
|
||||
@@ -0,0 +1,358 @@
|
||||
"""SQLiteVariantStore tests (REQ-3-005, D-027).
|
||||
|
||||
Each test gets its own tmp-path SQLite file — no shared disk state. Covers:
|
||||
- save / get roundtrip by (learner_id, template_id) AND by task_id
|
||||
(all fields survive, including nested JSON params, starter_files
|
||||
filename->content map, and the tz-aware created_at contract)
|
||||
- insert-only: a second save for the same (learner_id, template_id)
|
||||
raises IntegrityError (first-wins, documented choice); the original
|
||||
row is untouched — NOT upsert, NOT swallowed
|
||||
- unique task_id: a second variant claiming an existing trace key is
|
||||
rejected even under a different (learner, template) pair
|
||||
- list_for_learner / list_by_template scoped + chronological + auditable
|
||||
(seed + params readable back — the proctoring cross-check path)
|
||||
- unknown learner / template / task -> None / empty lists
|
||||
- scoping: rows for other learners/templates never leak
|
||||
- rows are detached: usable after the store is closed
|
||||
- WAL + synchronous=NORMAL pragmas actually applied to the DB file
|
||||
- concurrent writer + reader against the same DB file (a-3 smoke test)
|
||||
"""
|
||||
|
||||
import concurrent.futures
|
||||
import threading
|
||||
from datetime import UTC, datetime, timedelta
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
import pytest
|
||||
import sqlalchemy as sa
|
||||
|
||||
from ai_service.variants.store import SQLiteVariantStore, VariantRecord
|
||||
|
||||
_BASE_TS = datetime(2026, 9, 12, 12, 0, 0, tzinfo=UTC)
|
||||
|
||||
|
||||
def make_variant(
|
||||
task_id: str = "task-1",
|
||||
learner_id: str = "learner-1",
|
||||
template_id: str = "template-1",
|
||||
seed: str = "seed-a1b2c3",
|
||||
params: dict[str, Any] | None = None,
|
||||
statement: str | None = None,
|
||||
starter_files: dict[str, Any] | None = None,
|
||||
created_at: datetime | None = None,
|
||||
) -> VariantRecord:
|
||||
"""Canonical kwargs builder — tests override only what they assert on."""
|
||||
return VariantRecord(
|
||||
learner_id=learner_id,
|
||||
template_id=template_id,
|
||||
task_id=task_id,
|
||||
seed=seed,
|
||||
params=params
|
||||
if params is not None
|
||||
else {
|
||||
"scenario": "cache invalidation",
|
||||
"constraints": ["no external deps", "streaming"],
|
||||
"data_shape": {"rows": 10_000, "columns": ["ts", "event", "user"]},
|
||||
},
|
||||
statement=statement
|
||||
if statement is not None
|
||||
else (
|
||||
"Implement a cache layer for the event stream service that "
|
||||
"survives restarts without losing buffered rows, learner "
|
||||
f"{learner_id} edition."
|
||||
),
|
||||
starter_files=starter_files
|
||||
if starter_files is not None
|
||||
else {
|
||||
"src/stream_cache.py": "class StreamCache:\n pass\n",
|
||||
"tests/test_stream_cache.py": "def test_roundtrip():\n pass\n",
|
||||
},
|
||||
created_at=created_at if created_at is not None else _BASE_TS,
|
||||
)
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def store(tmp_path: Path) -> SQLiteVariantStore:
|
||||
s = SQLiteVariantStore(db_path=tmp_path / "variants.db")
|
||||
yield s
|
||||
s.close()
|
||||
|
||||
|
||||
def test_save_and_get_roundtrip(store: SQLiteVariantStore) -> None:
|
||||
variant = make_variant()
|
||||
store.save(variant)
|
||||
|
||||
# Point lookup by variant identity (learner_id, template_id).
|
||||
fetched = store.get("learner-1", "template-1")
|
||||
assert fetched is not None
|
||||
assert fetched.learner_id == "learner-1"
|
||||
assert fetched.template_id == "template-1"
|
||||
assert fetched.task_id == "task-1"
|
||||
assert fetched.seed == "seed-a1b2c3"
|
||||
assert fetched.params == variant.params
|
||||
assert fetched.statement == variant.statement
|
||||
assert fetched.starter_files == variant.starter_files
|
||||
assert fetched.created_at == _BASE_TS
|
||||
assert fetched.created_at.tzinfo is UTC # tz-normalized on read
|
||||
|
||||
# Same row via the grading/telemetry trace key — the join path a
|
||||
# grade or proctor check uses with only a task_id in hand.
|
||||
by_task = store.get_by_task("task-1")
|
||||
assert by_task is not None
|
||||
assert by_task.learner_id == "learner-1"
|
||||
assert by_task.template_id == "template-1"
|
||||
assert by_task.seed == "seed-a1b2c3"
|
||||
assert by_task.statement == variant.statement
|
||||
assert by_task.starter_files == variant.starter_files
|
||||
assert by_task.params == variant.params
|
||||
assert by_task.created_at.tzinfo is UTC
|
||||
|
||||
|
||||
def test_get_unknown_returns_none(store: SQLiteVariantStore) -> None:
|
||||
store.save(make_variant())
|
||||
|
||||
assert store.get("learner-1", "template-missing") is None
|
||||
assert store.get("learner-missing", "template-1") is None
|
||||
assert store.get("nobody", "nothing") is None
|
||||
assert store.get_by_task("task-missing") is None
|
||||
|
||||
|
||||
def test_duplicate_learner_template_pair_raises_integrity_error(
|
||||
store: SQLiteVariantStore,
|
||||
) -> None:
|
||||
"""THE contract of this store (insert-only, first wins).
|
||||
|
||||
The first generated variant is authoritative (D-029 reproducibility:
|
||||
the seed re-derives the same variant); a duplicate save is a
|
||||
programming error or lost race, not a normal flow — the generator's
|
||||
cache path serves `get` instead. So the IntegrityError surfaces to
|
||||
the caller and the stored row is left untouched.
|
||||
"""
|
||||
first = make_variant(statement="first authoritative statement")
|
||||
store.save(first)
|
||||
|
||||
second = make_variant(
|
||||
task_id="task-2", # distinct task key; the PAIR collides
|
||||
seed="seed-999",
|
||||
statement="an impostor statement",
|
||||
created_at=_BASE_TS + timedelta(hours=1),
|
||||
)
|
||||
with pytest.raises(sa.exc.IntegrityError):
|
||||
store.save(second)
|
||||
|
||||
# First save survived intact — nothing was overwritten.
|
||||
fetched = store.get("learner-1", "template-1")
|
||||
assert fetched is not None
|
||||
assert fetched.task_id == "task-1"
|
||||
assert fetched.statement == "first authoritative statement"
|
||||
assert fetched.seed == "seed-a1b2c3"
|
||||
assert fetched.created_at == _BASE_TS
|
||||
|
||||
# The rejected save left no row behind under its task_id either.
|
||||
assert store.get_by_task("task-2") is None
|
||||
|
||||
|
||||
def test_duplicate_task_id_raises_integrity_error(store: SQLiteVariantStore) -> None:
|
||||
# task_id is the grading/telemetry trace key — globally unique: a
|
||||
# second variant may never claim an existing trace key, even under a
|
||||
# different (learner_id, template_id) pair.
|
||||
store.save(make_variant(learner_id="learner-1", template_id="template-1"))
|
||||
|
||||
with pytest.raises(sa.exc.IntegrityError):
|
||||
store.save(
|
||||
make_variant(
|
||||
learner_id="learner-2",
|
||||
template_id="template-2",
|
||||
task_id="task-1", # collides with learner-1's trace key
|
||||
)
|
||||
)
|
||||
|
||||
# Rejected row not partially stored under either identity.
|
||||
assert store.get("learner-2", "template-2") is None
|
||||
assert len(store.list_by_template("template-2")) == 0
|
||||
|
||||
|
||||
def test_list_for_learner_roundtrip_and_audit(store: SQLiteVariantStore) -> None:
|
||||
# Created out of insertion order; list must come back chronological.
|
||||
store.save(make_variant(template_id="template-c", task_id="task-c",
|
||||
created_at=_BASE_TS + timedelta(hours=2)))
|
||||
store.save(make_variant(template_id="template-a", task_id="task-a",
|
||||
created_at=_BASE_TS))
|
||||
store.save(make_variant(template_id="template-b", task_id="task-b",
|
||||
created_at=_BASE_TS + timedelta(hours=1)))
|
||||
|
||||
variants = store.list_for_learner("learner-1")
|
||||
assert [v.template_id for v in variants] == [
|
||||
"template-a",
|
||||
"template-b",
|
||||
"template-c",
|
||||
]
|
||||
assert all(v.learner_id == "learner-1" for v in variants)
|
||||
hours = (timedelta(hours=0), timedelta(hours=1), timedelta(hours=2))
|
||||
assert all(
|
||||
v.created_at == _BASE_TS + offset
|
||||
for v, offset in zip(variants, hours, strict=True)
|
||||
)
|
||||
assert all(v.created_at.tzinfo is UTC for v in variants)
|
||||
|
||||
# Auditable: every stored variant reads back its seed and typed params
|
||||
# (the proctoring cross-check path reads exactly this).
|
||||
for v in variants:
|
||||
assert v.seed.startswith("seed-")
|
||||
assert v.params["scenario"] == "cache invalidation"
|
||||
assert "streaming" in v.params["constraints"]
|
||||
assert v.params["data_shape"]["rows"] == 10_000
|
||||
assert v.starter_files["tests/test_stream_cache.py"].count("\n") >= 1
|
||||
|
||||
|
||||
def test_list_by_template_roundtrip_and_audit(store: SQLiteVariantStore) -> None:
|
||||
# Three learners on the same template: distinct, auditable variants.
|
||||
store.save(make_variant(learner_id="learner-b", task_id="task-b",
|
||||
seed="seed-222",
|
||||
created_at=_BASE_TS + timedelta(hours=1)))
|
||||
store.save(make_variant(learner_id="learner-a", task_id="task-a",
|
||||
seed="seed-111", created_at=_BASE_TS))
|
||||
store.save(make_variant(learner_id="learner-c", task_id="task-c",
|
||||
seed="seed-333",
|
||||
created_at=_BASE_TS + timedelta(hours=2)))
|
||||
|
||||
variants = store.list_by_template("template-1")
|
||||
assert [v.learner_id for v in variants] == ["learner-a", "learner-b", "learner-c"]
|
||||
assert all(v.template_id == "template-1" for v in variants)
|
||||
|
||||
# Every learner's variant carries its own seed + params (auditable,
|
||||
# REQ-3-005: variant parameters persisted and auditable).
|
||||
seeds = {v.learner_id: v.seed for v in variants}
|
||||
assert seeds == {
|
||||
"learner-a": "seed-111",
|
||||
"learner-b": "seed-222",
|
||||
"learner-c": "seed-333",
|
||||
}
|
||||
assert all(v.params["scenario"] == "cache invalidation" for v in variants)
|
||||
assert all(v.created_at.tzinfo is UTC for v in variants)
|
||||
|
||||
|
||||
def test_list_unknown_returns_empty_lists(store: SQLiteVariantStore) -> None:
|
||||
store.save(make_variant())
|
||||
|
||||
assert store.list_for_learner("nobody") == []
|
||||
assert store.list_by_template("no-template") == []
|
||||
|
||||
|
||||
def test_lists_are_scoped(store: SQLiteVariantStore) -> None:
|
||||
store.save(make_variant(learner_id="learner-1", template_id="template-1",
|
||||
task_id="task-1"))
|
||||
store.save(make_variant(learner_id="learner-2", template_id="template-1",
|
||||
task_id="task-2"))
|
||||
store.save(make_variant(learner_id="learner-1", template_id="template-2",
|
||||
task_id="task-3"))
|
||||
|
||||
# learner lists see only that learner's rows.
|
||||
assert [v.template_id for v in store.list_for_learner("learner-1")] == [
|
||||
"template-1",
|
||||
"template-2",
|
||||
]
|
||||
assert [v.template_id for v in store.list_for_learner("learner-2")] == ["template-1"]
|
||||
|
||||
# template lists see one row per learner, none from other templates.
|
||||
template_rows = store.list_by_template("template-1")
|
||||
assert sorted(v.learner_id for v in template_rows) == ["learner-1", "learner-2"]
|
||||
assert all(v.template_id == "template-1" for v in template_rows)
|
||||
|
||||
# get stays a pair-scoped point lookup: same template, other learner.
|
||||
assert store.get("learner-1", "template-2") is not None
|
||||
assert store.get("learner-2", "template-2") is None
|
||||
|
||||
|
||||
def test_empty_params_and_starter_files_roundtrip(store: SQLiteVariantStore) -> None:
|
||||
# Legal shapes: a slotless template carries no params; a variant may
|
||||
# ship without a workspace scaffold.
|
||||
store.save(make_variant(params={}, starter_files={}))
|
||||
|
||||
fetched = store.get("learner-1", "template-1")
|
||||
assert fetched is not None
|
||||
assert fetched.params == {}
|
||||
assert fetched.starter_files == {}
|
||||
|
||||
|
||||
def test_rows_are_detached_after_save(store: SQLiteVariantStore, tmp_path: Path) -> None:
|
||||
# The API layer hands VariantRecords across layers; rows must survive
|
||||
# the store that produced them being closed (no open-session ORM magic).
|
||||
store.save(make_variant())
|
||||
fetched = store.get("learner-1", "template-1")
|
||||
by_task = store.get_by_task("task-1")
|
||||
store.close()
|
||||
|
||||
assert fetched is not None
|
||||
assert fetched.statement == fetched.statement # usable post-close
|
||||
assert fetched.starter_files["src/stream_cache.py"] == "class StreamCache:\n pass\n"
|
||||
assert by_task is not None
|
||||
assert by_task.seed == "seed-a1b2c3"
|
||||
|
||||
# A fresh store on the same file sees the same row (durability).
|
||||
reopened = SQLiteVariantStore(db_path=tmp_path / "variants.db")
|
||||
try:
|
||||
again = reopened.get("learner-1", "template-1")
|
||||
assert again is not None
|
||||
assert again.starter_files["src/stream_cache.py"].startswith("class StreamCache")
|
||||
assert again.created_at.tzinfo is UTC
|
||||
finally:
|
||||
reopened.close()
|
||||
|
||||
|
||||
def test_pragmas_are_applied(store: SQLiteVariantStore) -> None:
|
||||
# Pragmas are per-connection; query through the store's engine so the
|
||||
# connect hook (not a default sqlite3 connection) is what we inspect.
|
||||
with store._engine.connect() as conn:
|
||||
(journal_mode,) = conn.execute(sa.text("PRAGMA journal_mode")).one()
|
||||
(synchronous,) = conn.execute(sa.text("PRAGMA synchronous")).one()
|
||||
assert journal_mode == "wal"
|
||||
# synchronous=NORMAL is 1 in SQLite's pragma numbering.
|
||||
assert synchronous == 1
|
||||
|
||||
|
||||
def test_concurrent_writer_and_reader_no_database_is_locked(tmp_path: Path) -> None:
|
||||
"""One thread saves while another reads in a tight loop (a-3).
|
||||
|
||||
Without WAL + busy_timeout this pattern reliably produces
|
||||
`OperationalError: database is locked` on SQLite. The assertion is
|
||||
that every reader call completes and every distinct-row write lands.
|
||||
"""
|
||||
db_path = tmp_path / "variants.db"
|
||||
n_variants = 60 # distinct (learner, template) pairs, one save each
|
||||
stop_writing = threading.Event()
|
||||
|
||||
writer = SQLiteVariantStore(db_path=db_path)
|
||||
reader = SQLiteVariantStore(db_path=db_path)
|
||||
try:
|
||||
|
||||
def write_variants() -> None:
|
||||
for seq in range(n_variants):
|
||||
writer.save(
|
||||
make_variant(
|
||||
learner_id=f"learner-{seq}",
|
||||
template_id="template-1",
|
||||
task_id=f"task-{seq}",
|
||||
seed=f"seed-{seq:03d}",
|
||||
created_at=_BASE_TS + timedelta(seconds=seq),
|
||||
)
|
||||
)
|
||||
stop_writing.set()
|
||||
|
||||
def read_variants() -> None:
|
||||
while not stop_writing.is_set():
|
||||
reader.list_by_template("template-1")
|
||||
# Final read after the writer is done.
|
||||
assert len(reader.list_by_template("template-1")) == n_variants
|
||||
|
||||
with concurrent.futures.ThreadPoolExecutor(max_workers=2) as pool:
|
||||
futures = [pool.submit(write_variants), pool.submit(read_variants)]
|
||||
for future in futures:
|
||||
future.result(timeout=30)
|
||||
|
||||
# Every save was a distinct row: nothing lost, none doubled.
|
||||
assert len(writer.list_by_template("template-1")) == n_variants
|
||||
finally:
|
||||
reader.close()
|
||||
writer.close()
|
||||
@@ -0,0 +1,102 @@
|
||||
"""Task template tests (Task 4-1-01, REQ-3-005)."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import random
|
||||
|
||||
import pytest
|
||||
|
||||
from ai_service.variants.templates import (
|
||||
TEMPLATES,
|
||||
ParameterSlot,
|
||||
TaskTemplate,
|
||||
get_template,
|
||||
slots_pattern_ok,
|
||||
template_for_competency,
|
||||
validate_competency_binding,
|
||||
)
|
||||
|
||||
|
||||
def test_all_templates_bind_to_real_competency_ids() -> None:
|
||||
validate_competency_binding() # raises on any unknown binding
|
||||
assert len(TEMPLATES) >= 3
|
||||
|
||||
|
||||
def test_slot_validation_rejects_bad_values() -> None:
|
||||
slot = ParameterSlot(name="domain", type="enum", values=["a", "b"])
|
||||
assert slot.validate_value("a")
|
||||
assert not slot.validate_value("c")
|
||||
assert not slot.validate_value(3)
|
||||
|
||||
rng = random.Random(42)
|
||||
assert slot.sample(rng) in {"a", "b"}
|
||||
|
||||
|
||||
def test_int_range_slot_bounds() -> None:
|
||||
slot = ParameterSlot(name="n", type="int_range", lo=2, hi=4)
|
||||
rng = random.Random(0)
|
||||
for _ in range(20):
|
||||
assert 2 <= slot.sample(rng) <= 4
|
||||
assert slot.validate_value(3)
|
||||
assert not slot.validate_value(5)
|
||||
assert not slot.validate_value("3")
|
||||
|
||||
|
||||
def test_seeded_sampling_is_reproducible() -> None:
|
||||
tpl = TEMPLATES["tpl-llm-judge"]
|
||||
first = tpl.sample_params(seed=1234)
|
||||
second = tpl.sample_params(seed=1234)
|
||||
other = tpl.sample_params(seed=1235)
|
||||
assert first == second # D-029: same seed -> identical params
|
||||
assert first != other # different seed -> (near-certainly) different draw
|
||||
|
||||
|
||||
def test_skeleton_placeholders_match_slots() -> None:
|
||||
for tpl in TEMPLATES.values():
|
||||
assert slots_pattern_ok(tpl.statement_skeleton, tpl.slots), tpl.id
|
||||
|
||||
|
||||
def test_render_validates_params_and_fills() -> None:
|
||||
tpl = TEMPLATES["tpl-llm-judge"]
|
||||
params = tpl.sample_params(seed=7)
|
||||
rendered = tpl.render(params)
|
||||
for value in params.values():
|
||||
assert str(value) in rendered
|
||||
with pytest.raises(ValueError, match="invalid value"):
|
||||
tpl.render({**params, "edge_cases": 99}) # out of band
|
||||
|
||||
|
||||
def test_rubric_anchors_present_per_template() -> None:
|
||||
for tpl in TEMPLATES.values():
|
||||
anchors = tpl.rubric_anchors
|
||||
assert anchors.expected_edit_count_band[0] <= anchors.expected_edit_count_band[1]
|
||||
assert anchors.expected_min_test_runs >= 1
|
||||
cycles = anchors.expected_error_fix_cycles_band
|
||||
assert cycles[0] <= cycles[1]
|
||||
|
||||
|
||||
def test_starter_files_defined_per_template() -> None:
|
||||
for tpl in TEMPLATES.values():
|
||||
assert tpl.starter_files, f"{tpl.id} missing starter scaffolds"
|
||||
assert "README.md" in tpl.starter_files
|
||||
assert tpl.test_command
|
||||
|
||||
|
||||
def test_competency_lookup() -> None:
|
||||
tpls = template_for_competency("stack-orchestration-c005")
|
||||
assert len(tpls) == 1
|
||||
assert get_template("nope-xyz") is None
|
||||
|
||||
|
||||
def test_bad_skeleton_rejected() -> None:
|
||||
with pytest.raises(ValueError, match="slot"):
|
||||
TaskTemplate(
|
||||
id="tpl-bad",
|
||||
competency_id="stack-orchestration-c001",
|
||||
title="Bad",
|
||||
statement_skeleton="no placeholders at all",
|
||||
slots=[ParameterSlot(name="x", type="enum", values=["a"])],
|
||||
rubric_anchors=TEMPLATES["tpl-llm-judge"].rubric_anchors,
|
||||
starter_files={},
|
||||
test_command="pytest -q",
|
||||
)
|
||||
@@ -0,0 +1 @@
|
||||
"""Voice layer tests — provider mock/fallback + DefenseStore (REQ-3-006)."""
|
||||
@@ -0,0 +1,469 @@
|
||||
"""SQLiteDefenseStore tests (REQ-3-006, D-027).
|
||||
|
||||
Each test gets its own tmp-path SQLite file — no shared disk state. Covers:
|
||||
- start → append_turn (examiner + learner interleaved) → finalize →
|
||||
get roundtrip: every field survives (including nested JSON
|
||||
integrity signals, per-turn latency_ms, and the tz-aware
|
||||
created_at/ts contract) and turns come back ordered by seq
|
||||
- lifecycle ownership: start forces in_progress + finished_at=None
|
||||
even if the caller smuggles a finished status
|
||||
- insert-only start: a duplicate id raises IntegrityError (the id is
|
||||
minted once per session); the original row is untouched
|
||||
- append_turn validation: duplicate (defense_id, seq) raises
|
||||
(a transcript turn must never silently vanish); unknown
|
||||
defense_id raises (FK enforced); mismatched defense_id argument
|
||||
raises ValueError before touching the DB; negative seq rejected
|
||||
- finalize: unknown id → None (documented behavior — the API maps
|
||||
it to 404); re-finalize is latest-wins on signals + finished_at
|
||||
- get: unknown id → None; turns attached only by get() —
|
||||
list_for_learner records carry turns == []
|
||||
- list_for_learner: scoped per learner, chronological
|
||||
- rows detached: usable after the store is closed; durability across
|
||||
a fresh store on the same file
|
||||
- WAL + synchronous=NORMAL + foreign_keys=ON pragmas actually applied
|
||||
"""
|
||||
|
||||
from datetime import UTC, datetime, timedelta
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
import sqlalchemy as sa
|
||||
|
||||
from ai_service.voice.defense_store import (
|
||||
DefenseRecord,
|
||||
DefenseTurn,
|
||||
SQLiteDefenseStore,
|
||||
)
|
||||
|
||||
_BASE_TS = datetime(2026, 9, 12, 12, 0, 0, tzinfo=UTC)
|
||||
|
||||
|
||||
def make_defense(
|
||||
id: str = "defense-1", # shadows builtin deliberately: DefenseRecord field name
|
||||
learner_id: str = "learner-1",
|
||||
task_id: str = "task-1",
|
||||
status: str = "in_progress",
|
||||
created_at: datetime | None = None,
|
||||
) -> DefenseRecord:
|
||||
"""Canonical kwargs builder — tests override only what they assert on."""
|
||||
return DefenseRecord(
|
||||
id=id,
|
||||
learner_id=learner_id,
|
||||
task_id=task_id,
|
||||
status=status,
|
||||
created_at=created_at if created_at is not None else _BASE_TS,
|
||||
)
|
||||
|
||||
|
||||
def make_turn(
|
||||
defense_id: str = "defense-1",
|
||||
seq: int = 0,
|
||||
role: str = "examiner",
|
||||
text: str = "Walk me through your cache invalidation strategy.",
|
||||
ts: datetime | None = None,
|
||||
latency_ms: int | None = 240,
|
||||
created_at: datetime | None = None,
|
||||
) -> DefenseTurn:
|
||||
return DefenseTurn(
|
||||
defense_id=defense_id,
|
||||
seq=seq,
|
||||
role=role,
|
||||
text=text,
|
||||
ts=ts if ts is not None else _BASE_TS + timedelta(seconds=seq),
|
||||
latency_ms=latency_ms,
|
||||
created_at=created_at
|
||||
if created_at is not None
|
||||
else _BASE_TS + timedelta(seconds=seq),
|
||||
)
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def store(tmp_path: Path) -> SQLiteDefenseStore:
|
||||
s = SQLiteDefenseStore(db_path=tmp_path / "defenses.db")
|
||||
yield s
|
||||
s.close()
|
||||
|
||||
|
||||
def test_start_append_finalize_get_roundtrip(store: SQLiteDefenseStore) -> None:
|
||||
"""THE roundtrip of the defense lifecycle (REQ-3-006)."""
|
||||
# start: a defense is born in_progress with empty signals.
|
||||
started = store.start(make_defense())
|
||||
assert started.status == "in_progress"
|
||||
assert started.finished_at is None
|
||||
assert started.integrity_signals == {}
|
||||
assert started.created_at == _BASE_TS
|
||||
assert started.created_at.tzinfo is UTC
|
||||
|
||||
# append: examiner + learner turns interleaved — append them OUT of
|
||||
# seq order to prove get() orders by seq, not by insertion.
|
||||
learner_a = make_turn(
|
||||
seq=1, role="learner", text="I invalidate on write-ahead flush.", latency_ms=980
|
||||
)
|
||||
examiner_b = make_turn(
|
||||
seq=2, role="examiner", text="Why not invalidate on read?", latency_ms=180
|
||||
)
|
||||
learner_c = make_turn(
|
||||
seq=3, role="learner", text="Read-path misses were rare in my trace.", latency_ms=1100
|
||||
)
|
||||
store.append_turn("defense-1", make_turn(seq=0)) # first examiner question
|
||||
store.append_turn("defense-1", learner_a)
|
||||
store.append_turn("defense-1", examiner_b)
|
||||
store.append_turn("defense-1", learner_c)
|
||||
|
||||
# finalize: seal with A-109 integrity signals (long pauses, off-scope).
|
||||
finalize_started = datetime.now(UTC) # real wall clock, not the fixture
|
||||
signals = {
|
||||
"long_pauses": {"count": 2, "threshold_ms": 2000, "turns": [1, 3]},
|
||||
"off_scope": {"count": 1, "turns": [3], "markers": ["unrelated tangent"]},
|
||||
"verdict": "PASS_WITH_NOTES",
|
||||
}
|
||||
finalized = store.finalize("defense-1", signals)
|
||||
assert finalized is not None
|
||||
assert finalized.status == "finished"
|
||||
assert finalized.finished_at is not None
|
||||
assert finalized.finished_at.tzinfo is UTC
|
||||
# finalize() does not attach turns — get() is the with-turns path.
|
||||
assert finalized.turns == []
|
||||
|
||||
# get: full transcript in seq order + signals persisted.
|
||||
fetched = store.get("defense-1")
|
||||
assert fetched is not None
|
||||
assert fetched.id == "defense-1"
|
||||
assert fetched.learner_id == "learner-1"
|
||||
assert fetched.task_id == "task-1"
|
||||
assert fetched.status == "finished"
|
||||
assert fetched.integrity_signals == signals
|
||||
assert fetched.integrity_signals["long_pauses"]["turns"] == [1, 3] # nested JSON survives
|
||||
assert fetched.created_at == _BASE_TS
|
||||
assert fetched.created_at.tzinfo is UTC
|
||||
assert fetched.finished_at is not None
|
||||
assert fetched.finished_at >= finalize_started # stamped at finalize time
|
||||
|
||||
# THE ordered-transcript assertion.
|
||||
assert [t.seq for t in fetched.turns] == [0, 1, 2, 3]
|
||||
assert [t.role for t in fetched.turns] == [
|
||||
"examiner",
|
||||
"learner",
|
||||
"examiner",
|
||||
"learner",
|
||||
]
|
||||
assert fetched.turns[0].text == "Walk me through your cache invalidation strategy."
|
||||
assert fetched.turns[1].latency_ms == 980
|
||||
assert fetched.turns[2].latency_ms == 180
|
||||
assert fetched.turns[3].latency_ms == 1100
|
||||
# Per-turn ts contract: tz-aware UTC on read regardless of backend.
|
||||
assert all(t.ts.tzinfo is UTC for t in fetched.turns)
|
||||
assert all(t.created_at.tzinfo is UTC for t in fetched.turns)
|
||||
assert [t.defense_id for t in fetched.turns] == ["defense-1"] * 4
|
||||
|
||||
|
||||
def test_start_forces_in_progress_lifecycle(store: SQLiteDefenseStore) -> None:
|
||||
"""The store owns the lifecycle: a smuggled finished status is
|
||||
normalized away at birth — only finalize() may move a defense to
|
||||
finished."""
|
||||
smuggled = make_defense(status="finished")
|
||||
smuggled.finished_at = _BASE_TS + timedelta(hours=1)
|
||||
started = store.start(smuggled)
|
||||
|
||||
assert started.status == "in_progress"
|
||||
assert started.finished_at is None
|
||||
|
||||
fetched = store.get("defense-1")
|
||||
assert fetched is not None
|
||||
assert fetched.status == "in_progress"
|
||||
assert fetched.finished_at is None
|
||||
|
||||
|
||||
def test_duplicate_defense_id_raises_integrity_error(
|
||||
store: SQLiteDefenseStore,
|
||||
) -> None:
|
||||
"""Insert-only start (first-wins): the id is minted once per session;
|
||||
a duplicate is a bug or lost race to surface, not swallow."""
|
||||
store.start(make_defense(learner_id="learner-1"))
|
||||
store.start(make_defense(id="defense-2", learner_id="learner-2"))
|
||||
|
||||
with pytest.raises(sa.exc.IntegrityError):
|
||||
store.start(make_defense(learner_id="learner-3")) # same id again
|
||||
|
||||
# The original rows survived intact — nothing was overwritten.
|
||||
first = store.get("defense-1")
|
||||
assert first is not None
|
||||
assert first.learner_id == "learner-1"
|
||||
assert first.status == "in_progress"
|
||||
second = store.get("defense-2")
|
||||
assert second is not None
|
||||
assert second.learner_id == "learner-2"
|
||||
|
||||
|
||||
def test_append_turn_duplicate_seq_raises_integrity_error(
|
||||
store: SQLiteDefenseStore,
|
||||
) -> None:
|
||||
"""A transcript turn must never silently vanish: (defense_id, seq) is
|
||||
the PK, so a duplicate raises instead of overwriting."""
|
||||
store.start(make_defense())
|
||||
store.append_turn("defense-1", make_turn(seq=0))
|
||||
store.append_turn("defense-1", make_turn(seq=1))
|
||||
|
||||
with pytest.raises(sa.exc.IntegrityError):
|
||||
store.append_turn(
|
||||
"defense-1",
|
||||
make_turn(seq=1, role="learner", text="an impostor answer"),
|
||||
)
|
||||
|
||||
# The stored turn is untouched — NOT upsert.
|
||||
fetched = store.get("defense-1")
|
||||
assert fetched is not None
|
||||
assert len(fetched.turns) == 2
|
||||
assert fetched.turns[1].role == "examiner"
|
||||
assert fetched.turns[1].text != "an impostor answer"
|
||||
|
||||
|
||||
def test_append_turn_unknown_defense_raises_integrity_error(
|
||||
store: SQLiteDefenseStore,
|
||||
) -> None:
|
||||
"""FK enforced (foreign_keys=ON): an orphan turn is rejected, not
|
||||
silently attached to a defense that does not exist."""
|
||||
store.start(make_defense())
|
||||
|
||||
with pytest.raises(sa.exc.IntegrityError):
|
||||
store.append_turn(
|
||||
"defense-missing",
|
||||
make_turn(defense_id="defense-missing", seq=0),
|
||||
)
|
||||
|
||||
# And the turn did not land on the existing defense either.
|
||||
fetched = store.get("defense-1")
|
||||
assert fetched is not None
|
||||
assert fetched.turns == []
|
||||
|
||||
|
||||
def test_append_turn_identity_mismatch_raises_value_error(
|
||||
store: SQLiteDefenseStore,
|
||||
) -> None:
|
||||
"""The defense_id argument is the write identity: a turn object
|
||||
claiming another defense is a programming error — surfaced BEFORE
|
||||
any DB round-trip."""
|
||||
store.start(make_defense())
|
||||
|
||||
with pytest.raises(ValueError, match="does not match"):
|
||||
store.append_turn(
|
||||
"defense-1",
|
||||
make_turn(defense_id="defense-other", seq=0),
|
||||
)
|
||||
|
||||
fetched = store.get("defense-1")
|
||||
assert fetched is not None
|
||||
assert fetched.turns == []
|
||||
|
||||
|
||||
def test_turn_model_rejects_negative_seq_and_unknown_role(store: SQLiteDefenseStore) -> None:
|
||||
"""Model-level contracts (@validates hooks fire at construction)."""
|
||||
with pytest.raises(ValueError, match="seq"):
|
||||
make_turn(seq=-1)
|
||||
with pytest.raises(ValueError, match="role"):
|
||||
make_turn(role="proctor")
|
||||
with pytest.raises(ValueError, match="text"):
|
||||
make_turn(text="")
|
||||
with pytest.raises(ValueError, match="latency_ms"):
|
||||
make_turn(latency_ms=-5)
|
||||
with pytest.raises(ValueError, match="defense_id"):
|
||||
make_turn(defense_id="")
|
||||
|
||||
|
||||
def test_append_turn_allows_missing_latency(store: SQLiteDefenseStore) -> None:
|
||||
"""latency_ms is None until the endpoints instrument it (task
|
||||
5-4-01) — None must roundtrip cleanly."""
|
||||
store.start(make_defense())
|
||||
store.append_turn(
|
||||
"defense-1",
|
||||
make_turn(seq=0, role="examiner", latency_ms=None),
|
||||
)
|
||||
|
||||
fetched = store.get("defense-1")
|
||||
assert fetched is not None
|
||||
assert fetched.turns[0].latency_ms is None
|
||||
|
||||
|
||||
def test_finalize_unknown_id_returns_none(store: SQLiteDefenseStore) -> None:
|
||||
"""Documented unknown-id behavior: None, not a raise — the defense
|
||||
endpoints (task 5-3-01) map this straight to 404."""
|
||||
assert store.finalize("nobody", {"verdict": "PASS"}) is None
|
||||
|
||||
|
||||
def test_finalize_is_latest_wins_on_refinalize(store: SQLiteDefenseStore) -> None:
|
||||
"""A recomputed verdict replaces the stored one wholesale (mirrors
|
||||
GradeStore.save): signals + finished_at are overwritten, status
|
||||
just stays finished."""
|
||||
store.start(make_defense())
|
||||
first_signals = {"verdict": "FAIL", "long_pauses": {"count": 5}}
|
||||
store.finalize("defense-1", first_signals)
|
||||
|
||||
better_signals = {
|
||||
"verdict": "PASS",
|
||||
"long_pauses": {"count": 1},
|
||||
"off_scope": {"count": 0},
|
||||
}
|
||||
refinalized = store.finalize("defense-1", better_signals)
|
||||
assert refinalized is not None
|
||||
|
||||
fetched = store.get("defense-1")
|
||||
assert fetched is not None
|
||||
assert fetched.status == "finished"
|
||||
assert fetched.integrity_signals == better_signals # wholesale replace
|
||||
assert fetched.integrity_signals["verdict"] == "PASS"
|
||||
assert refinalized.finished_at is not None
|
||||
assert fetched.finished_at == refinalized.finished_at # stamped anew
|
||||
|
||||
|
||||
def test_finalize_rejects_non_dict_signals(store: SQLiteDefenseStore) -> None:
|
||||
"""The signals column contract is a JSON OBJECT dict; None/str/list
|
||||
would break every reader (Examiner feed, Proctor/Mentor)."""
|
||||
store.start(make_defense())
|
||||
|
||||
with pytest.raises(ValueError, match="integrity_signals"):
|
||||
store.finalize("defense-1", None) # type: ignore[arg-type]
|
||||
|
||||
fetched = store.get("defense-1")
|
||||
assert fetched is not None
|
||||
assert fetched.status == "in_progress" # untouched by the rejected call
|
||||
|
||||
|
||||
def test_get_unknown_returns_none(store: SQLiteDefenseStore) -> None:
|
||||
store.start(make_defense())
|
||||
|
||||
assert store.get("defense-1") is not None
|
||||
assert store.get("defense-missing") is None
|
||||
|
||||
|
||||
def test_empty_signals_until_finalize(store: SQLiteDefenseStore) -> None:
|
||||
"""A-109: signals are {} until finalize — the in-progress transcript
|
||||
is readable without any verdict present."""
|
||||
store.start(make_defense())
|
||||
store.append_turn("defense-1", make_turn(seq=0))
|
||||
|
||||
fetched = store.get("defense-1")
|
||||
assert fetched is not None
|
||||
assert fetched.status == "in_progress"
|
||||
assert fetched.integrity_signals == {}
|
||||
assert fetched.finished_at is None
|
||||
assert len(fetched.turns) == 1
|
||||
|
||||
|
||||
def test_list_for_learner_scoped_and_chronological(store: SQLiteDefenseStore) -> None:
|
||||
# Created out of insertion order; list must come back chronological.
|
||||
store.start(make_defense(id="defense-c", learner_id="learner-1",
|
||||
created_at=_BASE_TS + timedelta(hours=2)))
|
||||
store.start(make_defense(id="defense-a", learner_id="learner-1"))
|
||||
store.start(make_defense(id="defense-b", learner_id="learner-2"))
|
||||
# Turns on learner-1's defenses prove the list carries NONE of them.
|
||||
store.append_turn("defense-c", make_turn(defense_id="defense-c", seq=0))
|
||||
store.append_turn("defense-a", make_turn(defense_id="defense-a", seq=0))
|
||||
|
||||
defenses = store.list_for_learner("learner-1")
|
||||
assert [d.id for d in defenses] == ["defense-a", "defense-c"]
|
||||
assert all(d.learner_id == "learner-1" for d in defenses)
|
||||
# WITHOUT turns: the list feed carries session headers only — get()
|
||||
# is the with-turns path.
|
||||
assert all(d.turns == [] for d in defenses)
|
||||
# Headers intact: status + signals readable for the Proctor/Mentor feed.
|
||||
assert all(d.status == "in_progress" for d in defenses)
|
||||
assert all(d.integrity_signals == {} for d in defenses)
|
||||
assert all(d.created_at.tzinfo is UTC for d in defenses)
|
||||
|
||||
# Other learners never leak.
|
||||
assert [d.id for d in store.list_for_learner("learner-2")] == ["defense-b"]
|
||||
assert store.list_for_learner("learner-missing") == []
|
||||
|
||||
|
||||
def test_list_includes_finalized_with_signals(store: SQLiteDefenseStore) -> None:
|
||||
"""The learner feed must surface finished defenses WITH their sealed
|
||||
signals (the Proctor/Mentor cross-check reads exactly this)."""
|
||||
store.start(make_defense())
|
||||
signals = {"verdict": "PASS", "long_pauses": {"count": 0}}
|
||||
store.finalize("defense-1", signals)
|
||||
|
||||
defenses = store.list_for_learner("learner-1")
|
||||
assert len(defenses) == 1
|
||||
assert defenses[0].status == "finished"
|
||||
assert defenses[0].integrity_signals == signals
|
||||
assert defenses[0].finished_at is not None
|
||||
assert defenses[0].finished_at.tzinfo is UTC
|
||||
|
||||
|
||||
def test_rows_are_detached_and_durable(store: SQLiteDefenseStore, tmp_path: Path) -> None:
|
||||
"""Detached from any session: the API layer hands DefenseRecords
|
||||
across layers; rows must survive the store that produced them
|
||||
being closed, and a fresh store must see the same rows."""
|
||||
store.start(make_defense())
|
||||
store.append_turn("defense-1", make_turn(seq=0))
|
||||
store.append_turn("defense-1", make_turn(seq=1, role="learner", text="My answer."))
|
||||
store.finalize("defense-1", {"verdict": "PASS"})
|
||||
fetched = store.get("defense-1")
|
||||
assert fetched is not None
|
||||
store.close()
|
||||
|
||||
# Usable post-close — no open-session ORM magic.
|
||||
assert fetched.status == "finished"
|
||||
assert fetched.integrity_signals["verdict"] == "PASS"
|
||||
assert [t.text for t in fetched.turns] == [
|
||||
"Walk me through your cache invalidation strategy.",
|
||||
"My answer.",
|
||||
]
|
||||
|
||||
# A fresh store on the same file sees the same rows (durability).
|
||||
reopened = SQLiteDefenseStore(db_path=tmp_path / "defenses.db")
|
||||
try:
|
||||
again = reopened.get("defense-1")
|
||||
assert again is not None
|
||||
assert again.status == "finished"
|
||||
assert again.integrity_signals == {"verdict": "PASS"}
|
||||
assert [t.seq for t in again.turns] == [0, 1]
|
||||
assert again.turns[1].latency_ms == 240
|
||||
assert all(t.ts.tzinfo is UTC for t in again.turns)
|
||||
finally:
|
||||
reopened.close()
|
||||
|
||||
|
||||
def test_pragmas_are_applied(store: SQLiteDefenseStore) -> None:
|
||||
# Pragmas are per-connection; query through the store's engine so the
|
||||
# connect hook (not a default sqlite3 connection) is what we inspect.
|
||||
with store._engine.connect() as conn:
|
||||
(journal_mode,) = conn.execute(sa.text("PRAGMA journal_mode")).one()
|
||||
(synchronous,) = conn.execute(sa.text("PRAGMA synchronous")).one()
|
||||
(foreign_keys,) = conn.execute(sa.text("PRAGMA foreign_keys")).one()
|
||||
assert journal_mode == "wal"
|
||||
# synchronous=NORMAL is 1 in SQLite's pragma numbering.
|
||||
assert synchronous == 1
|
||||
# The FK is actually enforced on SQLite (Postgres parity, D-027).
|
||||
assert foreign_keys == 1
|
||||
|
||||
|
||||
def test_two_defenses_same_learner_independent_transcripts(
|
||||
store: SQLiteDefenseStore,
|
||||
) -> None:
|
||||
"""Scoped transcripts: two defenses never see each other's turns."""
|
||||
store.start(make_defense(id="defense-a", task_id="task-1"))
|
||||
store.start(make_defense(id="defense-b", task_id="task-2"))
|
||||
# Both defenses reuse seq 0,1 — per-defense numbering.
|
||||
for defense_id in ("defense-a", "defense-b"):
|
||||
store.append_turn(defense_id, make_turn(defense_id=defense_id, seq=0))
|
||||
store.append_turn(
|
||||
defense_id,
|
||||
make_turn(
|
||||
defense_id=defense_id, seq=1, role="learner",
|
||||
text=f"answer for {defense_id}",
|
||||
),
|
||||
)
|
||||
|
||||
a = store.get("defense-a")
|
||||
b = store.get("defense-b")
|
||||
assert a is not None and b is not None
|
||||
assert [t.seq for t in a.turns] == [0, 1]
|
||||
assert [t.text for t in a.turns] == [
|
||||
"Walk me through your cache invalidation strategy.",
|
||||
"answer for defense-a",
|
||||
]
|
||||
assert [t.text for t in b.turns] == [
|
||||
"Walk me through your cache invalidation strategy.",
|
||||
"answer for defense-b",
|
||||
]
|
||||
@@ -0,0 +1,101 @@
|
||||
"""Per-turn latency instrumentation tests (Task 5-4-01, REQ-3-006, A-109).
|
||||
|
||||
Mock-based: asserts instrumentation PRESENCE and population (stt_ms / llm_ms /
|
||||
tts_ms fields, per-turn latency_ms persisted, the budget constant defined) —
|
||||
wall-clock against a real voice endpoint is a v0.4 acceptance criterion
|
||||
(real STT/TTS deferred per GRILL CUT-1 / G-7).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from ai_service.agents.examiner import ExaminerAgent
|
||||
from ai_service.config import Settings
|
||||
from ai_service.grading.store import SQLiteGradeStore
|
||||
from ai_service.llm.mock import MockProvider
|
||||
from ai_service.main import create_app
|
||||
from ai_service.telemetry.ingest import TraceIntegrityMap
|
||||
from ai_service.telemetry.store import SQLiteTraceStore
|
||||
from ai_service.variants.store import SQLiteVariantStore
|
||||
from ai_service.voice.defense_store import SQLiteDefenseStore
|
||||
from ai_service.voice.mock import MockVoiceProvider
|
||||
|
||||
#: A-109: the documented conversational budget (acceptance criterion for the
|
||||
#: v0.4 real-voice probe; mock turns are near-instant so v0.3 asserts
|
||||
#: instrumentation, not wall-clock).
|
||||
DEFENSE_TURN_BUDGET_MS = 4_000
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def client(tmp_path: Path) -> TestClient:
|
||||
llm = MockProvider()
|
||||
app = create_app(Settings(provider="mock", voice_provider="mock"))
|
||||
app.state.provider = llm
|
||||
app.state.trace_store = SQLiteTraceStore(db_path=tmp_path / "t.db")
|
||||
app.state.variant_store = SQLiteVariantStore(db_path=tmp_path / "v.db")
|
||||
app.state.grade_store = SQLiteGradeStore(db_path=tmp_path / "g.db")
|
||||
app.state.trace_integrity = TraceIntegrityMap()
|
||||
app.state.defense_store = SQLiteDefenseStore(db_path=tmp_path / "d.db")
|
||||
app.state.voice_provider = MockVoiceProvider(["my answer"])
|
||||
app.state.examiner_agent = ExaminerAgent(llm, Settings(provider="mock"))
|
||||
with TestClient(app) as c:
|
||||
yield c
|
||||
|
||||
|
||||
class TestLatencyInstrumentation:
|
||||
def test_budget_constant_defined(self) -> None:
|
||||
"""The conversational budget is a named, documented constant (A-109)."""
|
||||
assert DEFENSE_TURN_BUDGET_MS > 0
|
||||
assert DEFENSE_TURN_BUDGET_MS <= 5_000 # conversational feel target
|
||||
|
||||
def test_answer_reports_per_phase_latency(self, client: TestClient) -> None:
|
||||
start = client.post(
|
||||
"/v1/defense/start",
|
||||
json={"learner_id": "lat-learner", "task_id": "lat-task"},
|
||||
).json()
|
||||
resp = client.post(
|
||||
f"/v1/defense/{start['defense_id']}/answer", data={"text": "answer"}
|
||||
)
|
||||
assert resp.status_code == 200
|
||||
latency = resp.json()["turn_latency"]
|
||||
assert latency["llm_ms"] is not None and latency["llm_ms"] >= 0
|
||||
assert "stt_ms" in latency and "tts_ms" in latency
|
||||
|
||||
def test_audio_answer_populates_stt_ms(self, client: TestClient) -> None:
|
||||
start = client.post(
|
||||
"/v1/defense/start",
|
||||
json={"learner_id": "lat-learner", "task_id": "lat-task"},
|
||||
).json()
|
||||
resp = client.post(
|
||||
f"/v1/defense/{start['defense_id']}/answer",
|
||||
files={"audio": ("a.wav", b"RIFF" + b"\x00" * 32, "audio/wav")},
|
||||
)
|
||||
latency = resp.json()["turn_latency"]
|
||||
assert latency["stt_ms"] is not None and latency["stt_ms"] >= 0
|
||||
|
||||
def test_every_turn_persists_latency_ms(self, client: TestClient) -> None:
|
||||
start = client.post(
|
||||
"/v1/defense/start",
|
||||
json={"learner_id": "lat-learner", "task_id": "lat-task"},
|
||||
).json()
|
||||
client.post(f"/v1/defense/{start['defense_id']}/answer", data={"text": "a"})
|
||||
transcript = client.get(f"/v1/defense/{start['defense_id']}").json()
|
||||
assert transcript["turns"]
|
||||
for turn in transcript["turns"]:
|
||||
assert "latency_ms" in turn
|
||||
assert turn["latency_ms"] is not None or turn["role"] == "learner"
|
||||
|
||||
def test_mock_turns_within_budget(self, client: TestClient) -> None:
|
||||
"""Mock turns must be near-instant — the budget holds trivially."""
|
||||
start = client.post(
|
||||
"/v1/defense/start",
|
||||
json={"learner_id": "lat-learner", "task_id": "lat-task"},
|
||||
).json()
|
||||
resp = client.post(
|
||||
f"/v1/defense/{start['defense_id']}/answer", data={"text": "a"}
|
||||
).json()
|
||||
assert resp["turn_latency"]["llm_ms"] < DEFENSE_TURN_BUDGET_MS
|
||||
@@ -0,0 +1,118 @@
|
||||
"""Voice layer tests (Task 5-1-01, REQ-3-006) — D-030 mock-first, zero network."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import pytest
|
||||
|
||||
from ai_service.config import Settings
|
||||
from ai_service.voice.browser import BROWSER_FALLBACK_DESCRIPTOR, MOCK_DESCRIPTOR
|
||||
from ai_service.voice.factory import UnknownVoiceProviderError, voice_provider_from_settings
|
||||
from ai_service.voice.mock import MockVoiceFailure, MockVoiceProvider, _tone_wav
|
||||
|
||||
|
||||
class TestMockVoiceProvider:
|
||||
async def test_transcribe_deterministic(self) -> None:
|
||||
provider = MockVoiceProvider(["hello defense"])
|
||||
a = await provider.transcribe(b"x" * 3200, "wav")
|
||||
b = await provider.transcribe(b"x" * 3200, "wav")
|
||||
assert a.text == b.text == "hello defense"
|
||||
assert a.duration_ms == b.duration_ms
|
||||
|
||||
async def test_transcribe_canned_queue(self) -> None:
|
||||
provider = MockVoiceProvider(["first answer", "second answer"])
|
||||
first = await provider.transcribe(b"audio", "webm")
|
||||
second = await provider.transcribe(b"audio", "webm")
|
||||
assert first.text == "first answer"
|
||||
assert second.text == "second answer"
|
||||
|
||||
async def test_transcribe_empty_audio_fails(self) -> None:
|
||||
provider = MockVoiceProvider(["x"])
|
||||
with pytest.raises(MockVoiceFailure, match="no audio"):
|
||||
await provider.transcribe(b"", "wav")
|
||||
|
||||
async def test_transcribe_scripted_failure_mode(self) -> None:
|
||||
provider = MockVoiceProvider(["FAIL"])
|
||||
with pytest.raises(MockVoiceFailure, match="scripted STT failure"):
|
||||
await provider.transcribe(b"audio", "wav")
|
||||
|
||||
async def test_synthesize_yields_nonempty_chunks(self) -> None:
|
||||
provider = MockVoiceProvider()
|
||||
chunks = [chunk async for chunk in provider.synthesize("question text")]
|
||||
assert chunks
|
||||
assert all(isinstance(c, bytes) and c for c in chunks)
|
||||
assert provider.synthesize_calls == 1
|
||||
|
||||
async def test_synthesize_empty_text_fails(self) -> None:
|
||||
provider = MockVoiceProvider()
|
||||
with pytest.raises(MockVoiceFailure):
|
||||
async for _ in provider.synthesize(""):
|
||||
pass
|
||||
|
||||
def test_tone_wav_is_real_wav(self) -> None:
|
||||
import io
|
||||
import wave
|
||||
|
||||
raw = _tone_wav(duration_ms=100)
|
||||
with wave.open(io.BytesIO(raw)) as w:
|
||||
assert w.getnchannels() == 1
|
||||
assert w.getsampwidth() == 2
|
||||
assert w.getframerate() == 8000
|
||||
|
||||
async def test_identical_synthesize_calls_identical_bytes(self) -> None:
|
||||
p1, p2 = MockVoiceProvider(), MockVoiceProvider()
|
||||
c1 = b"".join([c async for c in p1.synthesize("same text")])
|
||||
c2 = b"".join([c async for c in p2.synthesize("same text")])
|
||||
assert c1 == c2
|
||||
|
||||
|
||||
class TestFactory:
|
||||
def test_default_is_mock(self) -> None:
|
||||
provider = voice_provider_from_settings(Settings())
|
||||
assert isinstance(provider, MockVoiceProvider)
|
||||
|
||||
def test_explicit_mock(self) -> None:
|
||||
provider = voice_provider_from_settings(Settings(voice_provider="mock"))
|
||||
assert isinstance(provider, MockVoiceProvider)
|
||||
|
||||
def test_browser_mode_selects_server_side_mock_for_text_fallback(self) -> None:
|
||||
# Browser mode composes the same deterministic provider server-side;
|
||||
# the descriptor tells the CLIENT to use native SR/TTS.
|
||||
provider = voice_provider_from_settings(Settings(voice_provider="browser"))
|
||||
assert isinstance(provider, MockVoiceProvider)
|
||||
|
||||
def test_real_server_stt_tts_rejected_as_v04_seam(self) -> None:
|
||||
with pytest.raises(UnknownVoiceProviderError, match="v0.4"):
|
||||
voice_provider_from_settings(Settings(voice_provider="openai-audio"))
|
||||
|
||||
def test_unknown_provider_rejected(self) -> None:
|
||||
with pytest.raises(UnknownVoiceProviderError, match="unknown"):
|
||||
voice_provider_from_settings(Settings(voice_provider="watson"))
|
||||
|
||||
|
||||
class TestDescriptors:
|
||||
def test_browser_fallback_descriptor(self) -> None:
|
||||
assert BROWSER_FALLBACK_DESCRIPTOR.mode == "browser"
|
||||
assert BROWSER_FALLBACK_DESCRIPTOR.sr_available
|
||||
assert BROWSER_FALLBACK_DESCRIPTOR.tts_available
|
||||
assert "SpeechRecognition" in BROWSER_FALLBACK_DESCRIPTOR.hint
|
||||
|
||||
def test_mock_descriptor(self) -> None:
|
||||
assert MOCK_DESCRIPTOR.mode == "mock"
|
||||
assert "v0.4" in MOCK_DESCRIPTOR.hint
|
||||
|
||||
|
||||
class TestZeroNetwork:
|
||||
def test_voice_package_never_imports_agents_or_api(self) -> None:
|
||||
import ast
|
||||
from pathlib import Path
|
||||
|
||||
pkg = Path(__file__).parents[2] / "ai_service" / "voice"
|
||||
for py in pkg.glob("*.py"):
|
||||
tree = ast.parse(py.read_text())
|
||||
for node in ast.walk(tree):
|
||||
if isinstance(node, ast.ImportFrom) and node.module:
|
||||
assert not node.module.startswith("ai_service.agents"), py
|
||||
assert not node.module.startswith("ai_service.api"), py
|
||||
if node.level and node.module:
|
||||
assert node.module.split(".")[-1] != "agents", py
|
||||
assert node.module.split(".")[-1] != "api", py
|
||||
@@ -0,0 +1,48 @@
|
||||
# @nextcraft/cli — nextcraft
|
||||
|
||||
The bootstrap CLI for the Nextcraft monorepo, shipped as a self-contained linux x64 binary (Node SEA) on every release.
|
||||
|
||||
## Commands
|
||||
|
||||
See the [root README quickstart](../../README.md) for the user-facing flow. Internals:
|
||||
|
||||
- `src/index.ts` — argv dispatch, exit-code contract (0 ok / 1 failure / 2 usage), direct-run guard (`argv[0] === argv[1]` detects SEA context — the installer renames the binary, so filename matching is unreliable)
|
||||
- `src/commands/` — doctor / bootstrap / verify / dev; all orchestration delegates to `apps/ai-service/scripts/*.sh` via `src/lib/spawn.ts` (array-args only, SIGTERM→SIGKILL timeout ladder)
|
||||
- `src/checks/` — pure logic: version compare, `.env` template diff
|
||||
- `tests/` — node:test suites: dispatch, checks, spawn, command stubs, real-box doctor integration, install.sh fixture-server E2E (tamper rejection, degradation), release-assets token isolation, fresh-clone E2E
|
||||
|
||||
## Build
|
||||
|
||||
```sh
|
||||
pnpm cli:typecheck # tsc --noEmit
|
||||
pnpm cli:test # node:test suites
|
||||
pnpm cli:build # tsc -p tsconfig.build.json -> dist/
|
||||
pnpm --filter @nextcraft/cli build:binary <tag> # SEA binary + sha256 sidecar
|
||||
```
|
||||
|
||||
`build:binary <tag>`: esbuild bundle (CJS, node18 target, version stamped via `NEXTCRAFT_VERSION_STAMP` define — `--version` reports the tag it was built as) → `node --experimental-sea-config` → postject injection into a copy of the system node binary → `dist/nextcraft-linux-x64` + `dist/nextcraft-linux-x64.sha256`. The binary runs without node on PATH (runtime embedded, ~117 MB).
|
||||
|
||||
## Release pipeline
|
||||
|
||||
Every ship from v0.3.2 onward runs `scripts/release-assets.sh <tag>` after tag+merge:
|
||||
|
||||
1. Builds the binary stamped with the tag
|
||||
2. Resolves `GITEA_TOKEN` from `.env*` files ONLY (`.ciagent/.env.secrets` first) — never from shell env
|
||||
3. Attaches `nextcraft-linux-x64` + `nextcraft-linux-x64.sha256` to the Gitea release (bounded retry, best-effort — never blocks the ship)
|
||||
|
||||
`scripts/install.sh` (POSIX sh, dash-safe): platform gate → Gitea latest-release API resolve → exact-name asset match → sha256 verify BEFORE install (mismatch = hard stop) → `~/.local/bin` install → PATH hint. Any failure degrades to printed source-bootstrap instructions.
|
||||
|
||||
## Secrets policy
|
||||
|
||||
The CLI never generates, writes, or echoes secrets. `bootstrap` copies `.env.example` → `.env` only when absent and warns on missing optional keys (mock providers keep the stack runnable keyless). Real keys live only in gitignored `.ciagent/.env.secrets`, exported by `apps/ai-service/scripts/dev.sh`.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
| Symptom | Cause / fix |
|
||||
|---------|-------------|
|
||||
| `pnpm not found` in doctor | `corepack enable pnpm` (installs to ~/.local/bin — ensure PATH includes it) |
|
||||
| doctor passes but verify fails on venv | re-run `nextcraft bootstrap` (venv/pip resolution is idempotent) |
|
||||
| `port 8420 busy` in verify | stop the process on :8420 (`kill $(lsof -t -i:8420)`) or set `AI_PORT` |
|
||||
| install.sh says "no binary assets yet" | release predates the binary pipeline (pre-v0.3.2); use source bootstrap |
|
||||
| Binary silent after rename | fixed since v0.3.2 (SEA argv detection); re-download the latest release |
|
||||
| Checksum mismatch on install | do NOT run the download; delete it and retry — report if it persists |
|
||||
@@ -0,0 +1,22 @@
|
||||
{
|
||||
"name": "@nextcraft/cli",
|
||||
"version": "0.0.0",
|
||||
"private": true,
|
||||
"type": "module",
|
||||
"bin": {
|
||||
"nextcraft": "dist/index.js"
|
||||
},
|
||||
"scripts": {
|
||||
"dev": "tsx src/index.ts",
|
||||
"test": "tsx --test tests/*.test.ts",
|
||||
"typecheck": "tsc --noEmit",
|
||||
"build": "tsc -p tsconfig.build.json",
|
||||
"build:binary": "node scripts/build-binary.mjs"
|
||||
},
|
||||
"devDependencies": {
|
||||
"@types/node": "^24.0.0",
|
||||
"tsx": "^4.23.0",
|
||||
"typescript": "^5.7.2",
|
||||
"esbuild": "0.28.2"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,63 @@
|
||||
#!/usr/bin/env node
|
||||
import { execFileSync } from "node:child_process";
|
||||
import { copyFileSync, existsSync, readFileSync, rmSync, statSync, writeFileSync } from "node:fs";
|
||||
import { createHash } from "node:crypto";
|
||||
import { dirname, join } from "node:path";
|
||||
import { fileURLToPath } from "node:url";
|
||||
|
||||
const pkgDir = dirname(fileURLToPath(import.meta.url)) + "/..";
|
||||
const dist = join(pkgDir, "dist");
|
||||
const bundle = join(dist, "bundle.cjs");
|
||||
const blob = join(dist, "sea-prep.blob");
|
||||
const config = join(dist, "sea-config.json");
|
||||
const out = join(dist, "nextcraft-linux-x64");
|
||||
const checksum = out + ".sha256";
|
||||
const version = process.argv[2] ?? "0.0.0-dev";
|
||||
|
||||
if (version !== "0.0.0-dev" && !/^v?\d/.test(version)) {
|
||||
console.error(`refusing to stamp implausible version: ${version}`);
|
||||
process.exit(1);
|
||||
}
|
||||
|
||||
rmSync(dist, { recursive: true, force: true });
|
||||
execFileSync(
|
||||
join(pkgDir, "node_modules/.bin/esbuild"),
|
||||
[
|
||||
join(pkgDir, "src/index.ts"),
|
||||
"--bundle",
|
||||
"--platform=node",
|
||||
"--format=cjs",
|
||||
"--target=node18",
|
||||
`--define:NEXTCRAFT_VERSION_STAMP=${JSON.stringify(version)}`,
|
||||
"--outfile=" + bundle,
|
||||
],
|
||||
{ stdio: "inherit" },
|
||||
);
|
||||
|
||||
const seaConfig = {
|
||||
main: bundle,
|
||||
output: blob,
|
||||
disableExperimentalSEAWarning: true,
|
||||
};
|
||||
writeFileSync(config, JSON.stringify(seaConfig));
|
||||
execFileSync(process.execPath, ["--experimental-sea-config", config], { stdio: "inherit" });
|
||||
|
||||
const nodeBin = process.execPath;
|
||||
copyFileSync(nodeBin, out);
|
||||
execFileSync(
|
||||
"npx",
|
||||
["--yes", "postject", out, "NODE_SEA_BLOB", blob, "--sentinel-fuse", "NODE_SEA_FUSE_fce680ab2cc467b6e072b8b5df1996b2"],
|
||||
{ stdio: "inherit" },
|
||||
);
|
||||
execFileSync("chmod", ["+x", out]);
|
||||
|
||||
const size = statSync(out).size;
|
||||
const hash = createHash("sha256").update(readFileSync(out)).digest("hex");
|
||||
writeFileSync(checksum, `${hash} nextcraft-linux-x64\n`);
|
||||
|
||||
console.log(`built ${out} (${(size / 1024 / 1024).toFixed(1)} MB) stamped ${version}`);
|
||||
console.log(`checksum ${checksum}: ${hash}`);
|
||||
if (!existsSync(out) || !existsSync(checksum)) {
|
||||
console.error("expected artifacts missing");
|
||||
process.exit(1);
|
||||
}
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user