G-1 resource-limit claims match mechanism (mem/cpu/fsize/wallclock kernel-enforced; per-sandbox pids + hard disk quota documented as accepted v0.3 gap, no cgroup/sudo) G-2 disk cap via manager workdir-size sweep (AI_SANDBOX_MAX_WORKDIR_MB, default 512MB) G-3 telemetry flood control: WS 1008 + INCOMPLETE_FLOODED trace (no silent drop-oldest) G-4 grader refuses gapped/incomplete traces (UNGRADABLE_TRACE_INCOMPLETE verdict) G-5 no-auth abuse control: per-learner caps + learner allowlist (ships despite KYC defer) G-6 release note discloses partial resource enforcement (P7 honesty) CUT-1 real server STT/TTS deferred to v0.4 (voice mock+browser-first) CUT-2 interactive xterm shell relay deferred to v0.4 (build panel = run/test + output) Advisories a-1..a-5 applied (startup reaper, RLIMIT_FSIZE, SQLite WAL, digest-gaming prompt note, variant fairness envelope) ---ci--- phase: 0 milestone: v0.3 status: grill ---/ci---
6.6 KiB
Nextcraft v0.3 — GRILL.md (Adversarial Review Verdict)
Stage: GRILL, Phase 0 pre-execution · Verdict: GO-WITH-CHANGES · Confidence: 0.72
Summary
The credential-pipeline architecture (telemetry → trace → grade → defense) is sound and correctly sequenced. Three plan claims did NOT survive contact with this box and were correct before execution. The central problem: the plan overstated sandbox resource-limit enforcement and deferred KYC without closing the resulting no-auth local abuse vector. Fixed via binding decisions G-1..G-6 + scope cuts CUT-1/CUT-2 — no redesign required.
Per-Axis Findings
| Axis | Verdict | Rationale |
|---|---|---|
| Feasibility | CONCERN | Core unshare userns/mount/pid/net isolation probe-verified (uid=0 in-ns, network isolated, writes contained). But rlimit enforcement is partial: RLIMIT_NPROC scopes to the real host uid (5 sandboxes share one pids budget) and no disk-quota tool exists on the box. |
| Over-scoping | CONCERN | 39 tasks across 5 new subsystems + learner-surface rewrite for a solo founder. Voice-real-path and the interactive terminal relay are separable from the pipeline proof → cut (CUT-1, CUT-2). |
| Architecture risk | CONCERN | SandboxBackend/TraceStore protocols are the right seams. Overclaimed "limits enforced" + WS ingest "drop-oldest on unbounded growth" contradicted the at-least-once grading guarantee. |
| Phase sequencing | PASS | P1 sandbox → P2 telemetry → P3 grading → P4 variants → P5 voice → P6 integration is a correct dependency DAG. |
| Verification honesty | FAIL (fixed) | Must-Haves asserted "resource limits enforced + observable" (CPU/memory/disk/time quotas) the named mechanism cannot satisfy; probe tests would pass while the guarantee was false. Corrected by G-1. |
| Cost/quota | CONCERN | LLM/voice mock-gated (good). No-auth sandbox creation + unbounded disk + shared NPROC let one learner starve others at zero cost → closed by G-5. |
| Milestone honesty | CONCERN (fixed) | Release note disclosed KYC deferral + IDE-only but was silent on partial resource-limit enforcement → G-6. |
| Security | FAIL (fixed) | POST /v1/sandboxes {learner_id} client-supplied over localhost CORS let any local process mint sandboxes/flood/exhaust shared resources → G-5 abuse control ships despite KYC deferral. |
| Operability | CONCERN (advisory) | In-memory sandbox registry loses handles on restart (orphaned namespaces) → a-1 startup reaper. |
Binding Decisions (applied to PLAN.md/ARCHITECTURE-adjacent docs/REQUIREMENTS.md/ROADMAP.md/PROJECT.md)
- G-1 (BINDING) — Resource-limit claims match the deliverable mechanism. Memory (RLIMIT_AS) + CPU (RLIMIT_CPU) + single-file (RLIMIT_FSIZE) + wall-clock reaper are kernel-enforced; per-sandbox pids and hard disk quota are NOT kernel-enforceable without cgroup delegation/sudo → documented as accepted v0.3 risk. Applied to REQ-3-002, PLAN P1 Must-Haves + Task 1-2-02, ROADMAP P1 criteria.
- G-2 (BINDING) — Disk cap via manager workdir-size sweep.
AI_SANDBOX_MAX_WORKDIR_MB(default 512MB); sweep snapshots+destroys over-cap sandboxes and logs an integrity signal; closes the unbounded-ddhole. Applied to PLAN Task 1-2-01 + 1-2-02(e) + P1 Must-Haves. - G-3 (BINDING) — Telemetry flood control WITHOUT silent drop. Bounded queue; on overflow or >
AI_TELEMETRY_MAX_EVENTS_PER_TASK(default 50k) → WS close 1008 + trace markedINCOMPLETE_FLOODED(Proctor signal). Silent drop-oldest forbidden (corrupts grading). Applied to PLAN Task 2-2-02 + P2 Must-Haves. - G-4 (BINDING) — Grader refuses incomplete/gapped traces.
grade()gates onTraceStore.gaps()+INCOMPLETE_FLOODED→ returnsverdict=UNGRADABLE_TRACE_INCOMPLETE; no credential from a gapped trace. Applied to PLAN Task 3-2-01 + P3 Must-Haves. - G-5 (BINDING) — No-auth abuse control at MVP scale. Per-learner sandbox cap + global create-rate cap (429) + server-side
learner_idallowlist (403) so the unauthenticated surface can't exhaust shared NPROC/disk. Ships WITH the milestone even though KYC is deferred. Applied to PLAN Task 1-3-01 + PROJECT A-110. - G-6 (BINDING) — Disclose partial enforcement in the P7 release note. Item (e): which limits are kernel-enforced vs best-effort, and that full enforcement is deferred to the post-MVP containerd backend. Applied to PLAN P7 release-note honesty block.
Scope Cuts (accepted — preserve the end-to-end credential pipeline)
- CUT-1 (G-7) — Real server STT/TTS (
OpenAIAudioProvider) deferred to v0.4. Voice is mock-first (D-030); the/audio/*real path can never run in CI and was the least-verifiable surface. v0.3 proves the full defense dialogue + integrity-signal pipeline over mock + browser-native fallback; theVoiceProviderprotocol is the future drop-in seam. Applied to PLAN Phase 5 Goal + Task 5-1-01/5-4-01 + P5 Must-Haves. - CUT-2 (G-8) — Interactive xterm.js shell relay deferred to v0.4. The credential pipeline needs process events (Run/Test + file edits), not a live keystroke-level shell — the most fragile real-time piece, unverifiable without a real terminal. The build panel becomes Run/Test buttons + read-only exec output render;
@xterm/*is NOT a v0.3 dependency. Applied to PLAN Env-facts, P6 Goal, Task 6-2-01/6-2-02/6-3-01, P6 Must-Haves, MVP/UX sections. - Variants (Phase 4) and the Examiner agent dialogue KEPT — both are on the credential critical path (anti-collusion + the defense dialogue).
Advisory (applied)
- a-1 Startup reaper: on lifespan boot, scan
AI_SANDBOX_DIR, reap workdirs whose recorded pid is dead, log a warning. → Task 1-2-01. - a-2
RLIMIT_FSIZE(~50MB) as a cheap partial single-file disk guard in the spawner'spreexec_fn. → Task 1-2-01 + 1-2-02(c). - a-3 SQLite
PRAGMA journal_mode=WAL+synchronous=NORMALat engine creation (avoidsdatabase is lockedunder concurrent ingest + grader reads). → Task 2-1-02. - a-4 Grading prompt note: treat high edit/command churn with no test-progress as a process-quality negative (softens digest-gaming naivety). → Task 3-2-01.
- a-5 Variant fairness envelope: two variants of one template must compute digests within the template's expected feature envelope ("same bar" is testable). → Task 4-2-01 Verify.
Outcome
GO — all six binding decisions and both scope cuts applied to PLAN.md / REQUIREMENTS.md / ROADMAP.md / PROJECT.md before Phase 1 execution. No axis requires escalation (all resolvable at confidence ≥ 0.85). The milestone no longer claims resource enforcement it cannot deliver, and the no-auth abuse vector is closed at MVP scale.