"""Grading rubric prompt — criteria, level anchors, digest render (REQ-3-004). The grading prompt is deliberately learner-anonymous and trace-bare: the model receives ONLY the fixed rubric text and the compact numeric digest (TraceDigest JSON, D-028) — never a raw command, file path, payload string, learner id, or task id. Everything variable the LLM sees is deterministic counters, which both bounds the prompt-injection surface and makes "no raw trace reaches the prompt" assert-able in tests (plant a distinctive marker in a command payload; assert it absent from every message the provider received). Rubric (four criteria, each scored 0-4 — the ids are the validated RubricScore keys enforced by grading/engine.py): process_quality — iterative building in small, verified steps. correctness — where the session ended (test/run outcomes). debugging_discipline — how failures were handled. test_usage — when and how often tests were run. Advisory a-4 (embedded in the process_quality anchors): high edit/command churn with NO test progress is a process-quality NEGATIVE — churn is not work. A session with many edits/commands whose test state never moves is thrashing, not iterating, and must score low on process quality. House-style deviation, documented: unlike the tutor prompts, this module has no SYSTEM_PROMPT placeholders and no render_context(learner_context) — grading is context-free by design (learner anonymity; the digest is the only variable input). Runtime imports are TYPE_CHECKING-only so this module stays pure text and can never import-cycle with grading/engine.py (engine imports this module; if this module imported grading.* at runtime while grading/__init__ pulls engine, the package init would deadlock on a partially-initialized module). Version: grader-v1. """ from __future__ import annotations from typing import TYPE_CHECKING, Final if TYPE_CHECKING: # pragma: no cover - typing only; keeps this module pure text from ..grading.features import TraceDigest PROMPT_VERSION = "grader-v1" #: Canonical criterion ids. The engine validates RubricScore criteria keys #: against this tuple; the schema hint and anchors below speak the same ids. RUBRIC_CRITERIA: Final[tuple[str, ...]] = ( "process_quality", "correctness", "debugging_discipline", "test_usage", ) #: Sentinel line the engine's user turn is rendered around. Tests (and the #: calibration mock) split on it to locate the digest JSON in the prompt. DIGEST_MARKER: Final = "PROCESS TRACE DIGEST (JSON):" SYSTEM_PROMPT = """You are the Grader of Nextcraft, an AI-native competency school. You score a learner's build session from a compact numeric digest of their process trace. You NEVER see the raw trace — commands, file contents, and payloads do not exist on your side; every number you need is in the digest. Rubric — score each criterion 0-4: process_quality — iterative building in small, verified steps. 4: tight edit→test loops throughout; small verified increments; healthy pacing. 3: steady small edits with regular runs; progress mostly verified. 2: some iteration, but large unverified leaps or long idle stretches. 1: a single bulk change (e.g. one large paste) then a single run; no iteration. 0: no meaningful work visible. ADVISORY: high edit/command churn with NO test progress (no runs, no movement in pass counts) is a process-quality NEGATIVE — churn is not work. Cap such a session at 1 on this criterion no matter how many edits or commands were counted. correctness — where the session ended up. 4: final test status pass, with tests passing early and consistently. 3: final pass, reached through fail→fix→pass cycles that closed. 2: final pass, but preceded by a long unresolved failure streak. 1: final fail, but partial passes observed along the way. 0: final fail, or no test/run evidence at all. debugging_discipline — how failures were handled. 4: every failure cycle closes; targeted fixes with low mean fix latency. 3: most fail→edit→re-run cycles close with a pass. 2: failures followed by edits, but cycles rarely close. 1: repeated failures with no targeted edits between runs (flailing). 0: failures with no fix attempts at all. test_usage — when and how often tests were run. 4: tests run early (small first-pass offset) and throughout the session. 3: regular test runs interleaved with edits. 2: sparse tests; long stretches of unverified edits. 1: a single late test run only. 0: no test or run evidence. Rules: - Judge STRICTLY from the digest numbers; cite the fields you used. - Strengths: the two strongest digest observations, one sentence each. - Gaps: the two most important missed opportunities, one sentence each (a clean session names its next-level improvement instead). - Be rigorous but fair: a session that ends green was not necessarily well built, and a struggling session that never passed may still show real debugging discipline. - Respond with ONLY a valid JSON object matching the provided schema — no markdown fences, no prose outside the JSON.""" RUBRIC_SCORE_SCHEMA_HINT = ( '{"criteria": {"process_quality": <0-4 int>, "correctness": <0-4 int>, ' '"debugging_discipline": <0-4 int>, "test_usage": <0-4 int>}, ' '"strengths": [""], "gaps": [""], ' '"verdict": "mastered" | "developing" | "not_yet"}' ) def render_trace_digest(digest: TraceDigest) -> str: """Render the grader's user turn: a marker line + the digest JSON — nothing else. This is the ONLY per-session content that ever reaches the LLM (D-028): the engine composes [system: SYSTEM_PROMPT, user: render_trace_digest(digest)] and the D-020 defense appends its generic schema instruction to this user turn at request time. No learner id, task id, or raw trace material is injected — assert-able by tests. """ return ( "Score this build session against the rubric.\n" f"{DIGEST_MARKER}\n{digest.model_dump_json()}" )