ARCHITECTURE.md: - D-024 sandbox isolation via unshare Linux user/mount/pid/net namespaces (probe-verified: in-ns uid=0, network isolated, no container runtime/sudo on box) - D-025 sandbox scope = coding IDE only (design/sim deferred to v0.4) - D-026 telemetry = WebSocket ingest + SQLite ordered event log (seq gap detection) - D-027 first persistence = SQLite, protocol-wrapped stores (Trace/Grade/Variant/Defense) - D-028 trace grading = deterministic digest features + LLM rubric (LLM never sees raw trace) - D-029 variant generation = seeded template instantiation, seed persisted - D-030 voice = provider-agnostic mock-first STT/TTS + browser fallback (D-014 mirror) - D-031 extend ai-service (no new apps); D-032 single-box 1-5 concurrent sandboxes - New engine component tables + v0.3 build order PERSONAS.md (lead-developer assessment): - new sandbox-engineer + voice-engineer custom personas (phase-specific) - backend-engineer territory extended to engine persistence/APIs; ai-engineer re-scoped to model-facing grading/variant/voice logic; security-auditor stays inactive (KYC deferred) A-101 corrected post-probe: bwrap absent -> unshare namespaces (verified) .gitignore: ignore SQLite data + sandbox runtime dirs ---ci--- phase: 0 milestone: v0.3 status: research ---/ci---
20 KiB
Nextcraft — ARCHITECTURE.md
Overview
Nextcraft is a TypeScript monorepo (pnpm workspaces + turborepo) with a Next.js application hosting four surfaces (Learner, Marketplace, Employer Dashboard, Admin), a shared component library, typed mock data layer, and shared types package. As of v0.2, a Python FastAPI application (apps/ai-service) hosts six AI tutor agents backed by a provider-agnostic LLM layer.
v0.3 additions (Credential Engines): real credential engines replace v0.2 mock inputs — a sandbox fabric (isolated per-learner coding environments via Linux user/mount/pid/net namespaces), a live build-telemetry pipeline (WebSocket ingest + SQLite-ordered event log), a process-trace grading engine, seeded per-learner variant task generation, and a voice-based oral defense (STT/TTS via a new provider-agnostic voice layer). First real persistence introduced: SQLite (ai_service/telemetry/, grading, variant, defense stores). Lab/Assessor/Proctor agents are re-grounded onto real telemetry/traces. Identity/age-gating (KYC) deferred per founder directive — no security engineer persona; secrets-hygiene checklist only.
Confirmed Technology Stack (v0.2)
| Technology | Version | Purpose |
|---|---|---|
| Node.js | v24.15.0 | Runtime (web) |
| pnpm | 12.3.4 | Package manager + workspaces |
| turborepo | 2.3.3 | Build orchestration |
| Next.js | 15 (App Router) | Web application framework |
| React | 19+ | UI library |
| TypeScript | 5.x | Type system |
| Tailwind CSS | v4 | Utility-first CSS |
| lucide-react | latest | Icons |
| recharts | latest | Charts |
| @xyflow/react | latest | Competency graph viewer |
| Python | 3.11.2 | Runtime (ai-service) |
| FastAPI | 0.141.x | AI service framework |
| uvicorn | 0.52.x | ASGI server |
| pydantic | 2.13.x | Request/response models, structured outputs |
| pydantic-settings | 2.15.x | Settings + env-file loading (replaces python-dotenv) |
| httpx | 0.28.x | Async LLM HTTP client (ollama-cloud + local providers) |
| sse-starlette | 3.4.x | SSE framing, ping keep-alive |
| pytest | 9.x | Test runner |
| pytest-asyncio | 1.4.x | Async tests (auto mode) |
| ruff | latest | Python lint (check-only, no formatter) — pnpm ai:lint |
| ollama-cloud | https://ollama.com/v1 | Default LLM provider (OpenAI-compatible, Bearer auth) |
Deliberately not used: openai-python SDK (the LLMProvider protocol is the port; raw httpx keeps delta passthrough and ollama-cloud quirk tolerance), python-dotenv (pydantic-settings reads .env natively), respx (httpx MockTransport is built in).
v0.2 Architecture Decisions (from Research)
- D-016 SSE envelope —
metaevent (agent/session/model, flushed before first token) → raw OpenAI-compatible chunk passthrough → optionaldone→errorevent before[DONE]on mid-stream failure. Pre-first-byte failures use proper HTTP status. Headers:Cache-Control: no-cache,X-Accel-Buffering: no. - D-017 httpx direct client — lifespan-managed
httpx.AsyncClient(10s connect / 300s read), shared by ollama-cloud and local providers; no SDK. - D-018 Agent framework —
BaseAgentABC (system_prompt/build_messages/stream_reply/structured_reply) + explicit registry; prompts are versioned code inprompts/. - D-019 Session store —
SessionStoreprotocol +InMemorySessionStore(asyncio.Lock, 20-message window, 500-cap LRU, agent-scoped sessions). DB-migration-ready. - D-020 Structured outputs — 4-layer defense:
response_format(auto-degrade) → prompt-embedded schema → fence-strip/first-balanced-object parse → single bounded retry. - D-021 Mock corpus in Python —
ai_service/corpus/pydantic-typed, convention-aligned with TSpackages/mock-data(shared IDs, cross-referencing headers); no codegen in v0.2. - D-022 Monorepo integration — zero-dependency shim
package.jsonin apps/ai-service +ai#*turbo passthrough tasks (cache:false, outputs:[]) + rootai:dev/ai:testscripts + idempotent venv bootstrap. - D-023 Testing — pytest-asyncio auto mode; TestClient
client.stream()for SSE; httpx MockTransport for byte-exact provider parser tests; scripted mock provider incl. failure modes. Tests never call the cloud.
v0.3 Architecture Decisions (from Research — Credential Engines)
- D-024 Sandbox isolation = Linux namespaces via
unshare— per-learner sandbox runs as a subprocess entered into fresh user+mount+pid+network namespaces (unshare --user --map-root-user --mount --pid --fork --net). Probe-verified on this box: in-namespace uid=0, network fully isolated (0 interfaces), learner writes land in a per-sandbox directory; proc-remount not permitted here but not required. Chosen because no container runtime (docker/podman/bwrap/firejail) exists on the box and there is no sudo. ASandboxBackendprotocol abstracts the spawner so a future containerd/runc backend can replace namespace-spawning without touching callers. - D-025 Sandbox scope = coding IDE only (v0.3) — the sandbox fabric provisions a single build environment (shell + filesystem + run/test). REQ-F-021's design-tool and simulation environments are deferred to v0.4; one real build path proves the full credential pipeline (telemetry → trace → grade → defense).
- D-026 Telemetry = WebSocket ingest + SQLite ordered event log — in-sandbox capture agent streams structured events over WebSocket to
ai_service(/v1/telemetry/ingest); events persisted to SQLite with a per-(learner,task) monotonicseqfor gap detection, giving durability + at-least-once delivery + replay without a message broker. - D-027 First persistence = SQLite, protocol-wrapped — introduces a real DB (
ai_service/data/*.db) for telemetry traces, grades, variants, and defenses. Access via SQLModel. Every store is a protocol (TraceStore,VariantStore,DefenseStore,GradeStore) with a SQLite implementation — Postgres-migration-ready, mirroring D-019's SessionStore pattern. - D-028 Process-trace grading = hybrid deterministic + LLM — deterministic features (test pass/fail, edit count, error/fix cycles, idle gaps, command categories) computed in code into a compact trace digest; the digest feeds an Assessor-style rubric prompt and returns structured scores via D-020 JSON defense. The LLM never sees the raw trace — only the digest.
- D-029 Variant generation = seeded template instantiation — task templates with typed parameter slots; an LLM instantiates a unique variant per learner from a seed; seed+parameters persisted (D-027) for grading fairness and proctoring cross-check.
- D-030 Voice = provider-agnostic, mock-first, browser-fallback — a
VoiceProviderprotocol (mirror ofLLMProvider) with STT (OpenAI-compatible/audio/transcriptions) + TTS (/audio/speech) against a configurable endpoint, a deterministic mock (canned transcript/audio) for tests, and browser-nativeSpeechRecognition/speechSynthesisas a no-key fallback. Voice defense reusesBaseAgent+ the existing SSE pipeline (newExamineragent). - D-031 No new apps — extend ai-service — telemetry, grading, variant, voice, and sandbox orchestration are new modules inside
apps/ai-service(sharing the LLM pool, config, and session infra). Only the in-sandbox capture agent is a separate tiny Python process shipped into the namespace. No new top-levelapps/entry. - D-032 Capacity = single-box, 1–5 concurrent sandboxes — concurrency guard returns 503 when the sandbox pool is full. No queueing, no horizontal scaling in v0.3 (solo-founder/pilot scale).
Components
apps/ai-service — AI Tutor Service (v0.2 NEW)
| Component | Description | Boundaries | Depends On |
|---|---|---|---|
ai_service/main.py |
FastAPI app factory, lifespan (httpx client pool, provider factory), CORS (localhost only), /health | App entry | config, llm, agents, api |
ai_service/config.py |
pydantic-settings Settings (env_prefix="AI_", env_file, SecretStr key) | Configuration only | None |
ai_service/api/ |
Endpoints: chat.py (POST /v1/chat/stream, SSE), lab.py, assessment.py, proctor.py, mentor.py; deps.py (DI) | Composes agents + sessions; never imported by llm/ or agents/ | agents, llm |
ai_service/llm/ |
types.py (Message; ChatDelta/ChoiceDelta removed in P3 — no consumers), base.py (LLMProvider protocol), openai_compat.py (ollama-cloud + local), mock.py (deterministic), factory.py | Never imports agents/ or api/ | config |
ai_service/agents/ |
base.py (BaseAgent ABC), registry.py, session.py (SessionStore), structured.py (JSON defense), coach/tutor/lab/assessor/proctor/mentor.py | Never imports api/ | llm, prompts, corpus |
ai_service/prompts/ |
Per-agent system prompt constants + render_context functions (str.format_map) | Data only | None |
ai_service/corpus/ |
Mock engine inputs: learner_context.py, telemetry.py (Lab/Proctor scenarios), artifacts.py (pre-baked artifacts, rubrics, transcripts) | Pydantic-typed; aligned with TS packages/mock-data by convention | None |
scripts/ |
bootstrap.sh (venv + pip install idempotent), dev.sh (exports keys from .ciagent/.env.secrets → uvicorn), test.sh (pytest), lint.sh (ruff check, G-3) | Dev entry points | pyproject.toml |
tests/ |
conftest.py (mock provider, settings override, TestClient), health, llm (MockTransport parser), agents (framework + per-agent), api (SSE stream tests) | Mock provider only — no cloud | all |
Module boundary rules: llm/ never imports agents/ or api/; agents/ never imports api/; api/ composes both via DI. corpus/ is the only home of mock engine data. Prompts are code — versioned and reviewed in git.
apps/ai-service — v0.3 Credential Engine modules (NEW)
| Component | Description | Boundaries | Depends On |
|---|---|---|---|
ai_service/sandbox/ |
backend.py (SandboxBackend protocol), unshare_backend.py (userns/mount/pid/net spawner, D-024), manager.py (lifecycle: create/list/snapshot/destroy + concurrency guard D-032), workdir.py (per-sandbox fs layout) |
Never imports api/ or agents/; spawns subprocesses only | config |
ai_service/telemetry/ |
models.py (TelemetryEvent, TraceSpan), store.py (TraceStore protocol + SQLite impl D-027), ingest.py (WebSocket /v1/telemetry/ingest, seq gap detection D-026) |
Persistence; never imports agents/ | config |
ai_service/grading/ |
features.py (deterministic trace digest D-028), engine.py (rubric scoring orchestration), store.py (GradeStore) |
LLM only via digest; never sees raw trace | llm, telemetry, prompts |
ai_service/variants/ |
templates.py (task template library), generator.py (seeded LLM instantiation D-029), store.py (VariantStore) |
LLM via structured output | llm, grading |
ai_service/voice/ |
base.py (VoiceProvider protocol D-030), openai_audio.py (STT/TTS vs compatible endpoint), browser.py (native SR/TTS fallback descriptor), mock.py (deterministic) |
Never imports agents/ or api/ | config |
ai_service/agents/examiner.py |
Seventh agent: oral defense examiner; streams over existing SSE, consumes process traces + emits integrity signals | reuses BaseAgent (D-018) | llm, prompts, telemetry |
ai_service/data/*.db |
SQLite databases (telemetry/grades/variants/defenses) | gitignored | — |
scripts/sandbox-agent.py |
Tiny in-namespace capture process shipped into the sandbox; streams telemetry to ingest | standalone | stdlib only |
Boundary additions: sandbox/, telemetry/, grading/, variants/, voice/ are engine modules — they never import api/ (which composes them via DI) and never import agents/ (agents call engines through narrow interfaces, not vice versa).
apps/web — Next.js Application
| Component | Description | Boundaries | Depends On |
|---|---|---|---|
app/(learner)/ |
Learner surface route group: landing, catalog, competency stack, dashboard, byte viewer, sandbox mockup, assessment mockup | Learner-only routes and layouts | packages/ui, packages/mock-data, packages/types |
app/(marketplace)/ |
Marketplace surface route group: job board, job detail, employer profile, search/filter, pricing | Marketplace-only routes and layouts | packages/ui, packages/mock-data, packages/types |
app/(employer)/ |
Employer dashboard route group: overview, talent search, candidate profile, posting management | Employer-only routes and layouts | packages/ui, packages/mock-data, packages/types |
app/(admin)/ |
Admin surface route group: overview, learner management, competency graph viewer, moderation | Admin-only routes and layouts | packages/ui, packages/mock-data, packages/types |
app/layout.tsx |
Root layout: theme provider, navigation shell, responsive container | All routes | packages/ui |
components/ |
Surface-specific components (learner/, marketplace/, employer/, admin/) plus shared chrome (navigation-shell, header/footer, role-switcher, theme-provider, breadcrumbs, dark-mode-toggle) | App-level components | packages/ui |
hooks/ |
use-chat-stream.ts — SSE client hook: fetch + ReadableStream, byte buffering + frame reassembly, idempotent AbortController cleanup | Client components only | ai-service SSE |
lib/ |
sse.ts (shared SSE frame parser — CRLF normalization + : ping immunity, G-1), breadcrumbs.ts, format.ts |
Pure utilities | None |
packages/ui — Shared Component Library
| Component | Description | Boundaries | Depends On |
|---|---|---|---|
tokens/ |
Design tokens as TS constants: colors, spacing, radii, shadows, breakpoints (mirrored as Tailwind v4 @theme tokens in apps/web globals.css) |
Foundation layer — no dependencies | None |
primitives/ |
Button, Input, Card, Badge, Avatar — each with a Storybook story | Atomic UI components | tokens, packages/types |
Composite/layout/theme components (navigation shell, tables, chat panels, graph viewer, theme provider) live in apps/web/components/ as app-level components, not in packages/ui.
packages/mock-data — Mock Data Layer
| Component | Description | Boundaries | Depends On |
|---|---|---|---|
competency-stacks.ts |
5 competency stacks (AI Orchestration Engineer, AI Safety & Governance, Human-AI Product Designer, AI-Augmented Field Operator, Computational Sciences), each with 12-18 competencies | Typed mock data | packages/types |
jobs.ts |
20+ mock AI-era job listings with skills, seniority, salary, match scores | Typed mock data | packages/types |
candidates.ts |
15+ mock candidate profiles with artifacts, process traces, defense scores, microcredentials | Typed mock data | packages/types |
employers.ts |
10+ mock employer profiles with logos, descriptions, open positions | Typed mock data | packages/types |
learner-progress.ts |
Mock learner progress data: active competencies, completion percentages, recent artifacts | Typed mock data | packages/types |
admin.ts |
Admin surface mock data: platform metrics, activity feed, system health, learner roster (admin view), moderation queues | Typed mock data | packages/types |
ai-scenarios.ts |
AI engine-input scenario IDs + display metadata for the learner agent panels; IDs string-identical to ai_service/corpus/ (D-021) |
Typed mock data | packages/types |
packages/types — Shared Types
| Component | Description | Boundaries | Depends On |
|---|---|---|---|
domain.ts |
Competency, CompetencyStack, Microcredential, Artifact, ProcessTrace, OralDefense, AssessmentRubric | Domain types | None |
marketplace.ts |
Job, Employer, Candidate, JobPosting, TalentMatch, SearchFilter | Marketplace types | None |
user.ts |
Learner, Admin, EmployerUser, AgeGroup, Role | User types | None |
ui.ts |
Component props, theme config, breakpoint definitions | UI types | None |
Data Flow
[packages/mock-data + packages/types] [ai_service/corpus]
│ (TS, web surfaces) │ (Python, agent inputs)
▼ ▼
[Next.js Route Groups] [ai-service agents]
(learner)/ (marketplace)/ coach tutor lab assessor proctor mentor
(employer)/ (admin)/ │
│ ▼
│ SSE (fetch + ReadableStream) [LLMProvider]
└────── client components ◄───────────────────┤
http://localhost:8420 ollama-cloud / local / mock
(https://ollama.com/v1)
- Web surfaces remain server-component-first; client components (chat, filters, graph viewer, dark mode toggle) fetch directly from ai-service over SSE (A-002: no Next.js API-route proxy in v0.2).
- The LLM provider layer is a dumb pipe — OpenAI-compatible chunks pass through byte-identical; envelope logic (meta/done/error) lives only in the API layer (D-016).
- Lab/Assessor/Proctor read mock scenarios from
ai_service/corpus/— real engines are v0.3+. - All automated tests use the deterministic mock provider; the cloud is for manual probes only.
Build Order (v0.3)
- Sandbox fabric — SandboxBackend protocol + unshare namespace spawner + lifecycle manager (create/list/snapshot/destroy) + concurrency guard + per-sandbox workdir; isolation + resource-limit probes
- Live build telemetry — TelemetryEvent models + SQLite TraceStore + WebSocket ingest endpoint + seq gap detection + in-sandbox capture agent
- Process-trace grading engine — deterministic feature/digest computation + rubric scoring via LLM structured output + GradeStore; calibrated against v0.2 mock corpora
- Variant task generation — template library + seeded LLM instantiation + VariantStore + difficulty normalization anchors
- Oral / voice defense — VoiceProvider protocol + STT/TTS + mock + browser fallback + Examiner agent + transcript/integrity-signal capture
- Agent re-grounding + learner surface integration — Lab/Assessor/Proctor consume real telemetry/grades/defense signals; learner sandbox mockup → real in-browser xterm.js build/run; assessment mockup → live defense + live grading
Build Order (v0.2 — complete)
- AI service scaffolding — apps/ai-service: FastAPI app, config, provider layer (ollama-cloud/local/mock), SSE chat endpoint, pytest harness, turbo integration
- Agent framework — BaseAgent, registry, session store, structured output, prompts scaffolding, learner-context corpus
- Coach + Tutor agents — full implementations, chat endpoint agent routing
- Lab + Assessor agents — telemetry scenarios + pre-baked artifacts corpus, /v1/lab/feedback + /v1/assessment/evaluate
- Proctor + Mentor agents — proctor scenarios, /v1/proctor/signals + /v1/mentor/narrative
- Learner surface integration — useChatStream hook, agent switcher, streaming/error/loading states, agent output panels across the four learner surfaces
The v0.1 build order (monorepo → types → mock data → tokens → primitives → layout → composites → surfaces → polish) is complete and preserved in git history (tags v0.0.1–v0.1.0).
Future Architecture (Post-v0.3, for reference)
v0.3 delivers the real credential engines; later milestones fill in the remaining platform:
- In-memory sessions → PostgreSQL + Drizzle/SQLModel — SessionStore + v0.3 TraceStore/GradeStore/VariantStore/DefenseStore protocols swap SQLite→Postgres with no API changes
- userns subprocess sandboxes → containerd/runc backend — D-024
SandboxBackendprotocol swap; same lifecycle API - Coding-IDE sandbox → design tool + simulation environments — REQ-F-021 full scope (v0.4)
- Mock provider → per-agent model routing — provider factory already selects by config; per-agent
AI_<AGENT>_MODELoverrides - No auth → real KYC + sessions — deferred per founder directive; REQ-F-017 identity/age-gating lands post-v0.3 (v0.4+). Age-gating remains the v0.1 visual flow mockup
- No search → Semantic vector search (pgvector) — Filter UI replaced with vector similarity search
- No payments → Payment processing — Pricing page replaced with real subscription/payment flows
The monorepo structure (apps/web + apps/ai-service + packages/*) accommodates further apps without restructuring.