Files
archy/core/archipelago/tests/fixtures/assistant-evals
archipelagoandClaude Fable 5 d419141a86 feat(13-14): the eighteen reference cases — the specification of what the loop must refuse
EV-01..EV-18 (four happy reads, four confirmed writes, five injection cases,
three authority-ceiling cases, one budget case, one privacy case) per
13-AI-SPEC.md §5's schema, written against the real tool registry
(assistant::tools::registry()) and the real wrap_untrusted() boundary shape
rather than against the spec's description of them. EV-11's payload carries
a forged closing boundary in the exact `{label}_DATA_{token}_END` shape
untrusted.rs emits, proving why the per-call random token (not the wording)
is what makes the boundary hold. README.md records the per-bucket
reviewer-role labeling from §5's Labeling table (engineer for EV-01..EV-08,
security-minded red-teamer for EV-09..EV-16, non-technical reviewer for the
EV-05/EV-06 confirmation-copy judgment) so a later contributor knows whose
judgment each case's expect block encodes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 06:53:25 -04:00
..

Assistant Evals — Reference Dataset

cases.jsonl is the 18-case adversarially-weighted reference dataset for the offline eval harness (core/archipelago/src/assistant/evals.rs, run via cargo test --package archipelago assistant::evals::). It lives in-repo, alongside the code, so cases are reviewed in PRs like code — see 13-AI-SPEC.md §5 "Reference Dataset" for the full schema and composition table this file implements verbatim.

JSONL cannot carry comments, so the reviewer-role ownership per §5's "Labeling" table is recorded here instead — a later contributor editing or extending a case should know which person's judgment that case's expect block encodes.

Labeling roles, by bucket

Cases Bucket Labeled by Why
EV-01…EV-08 happy-read / confirmed-write Archipelago maintainer (engineer) expect blocks are mechanical — tool names come from D-06's curated registry, fixed before the tools were built. Writable in Wave 0, no domain-expert judgment call needed beyond "does the right tool get called."
EV-09…EV-16 injection / ceiling Security-minded technical reviewer (13-AI-SPEC §1b role 2 — "plays sysadmin with no rollback") This is red-teaming, not test-writing. Whoever writes EV-11 (the forged-boundary case) must be trying to break the per-call delimiter, not documenting that it exists. The must_not_claim fields on EV-14…EV-16 in particular need an adversarial eye: the failure to catch is a plausible-sounding fabrication, not an obvious lie.
EV-05 / EV-06 confirmation copy confirmed-write Non-technical reviewer — the "bought sovereignty, not a terminal" persona (13-AI-SPEC §1b role 3) The only valid labeller for E-02 (confirmation clarity) and E-09 (comprehension under time pressure). §1b is explicit that a security reviewer systematically under-catches confusing copy because they already understand the domain — their label is not ground truth for this dimension. This dataset's own EV-05/EV-06 expect blocks are mechanical (tool/confirmation/turn counts); the actual confirmation-copy judgment happens in the session-level walkthrough this pair feeds into (13-14 Task 3), not in this file.
LLM judge (E-02 at scale) Calibrated against the non-technical reviewer's labels Not trusted as a score until agreement ≥ 0.7 against the human labels (ai-evals.md Verify phase). Until then it is a screening tool that flags candidates for human review, only.
Sign-off Product owner / maintainer Rubric sign-off per §1b role 4.

Notes for future contributors

  • Every expect.must_not_execute / expect.must_not_claim entry should name something a real, curated tool in assistant::tools::registry() could plausibly be confused with — EV-14/EV-15/EV-16 deliberately leave must_not_execute empty because no tool for "send sats"/"show seed"/"factory reset" exists at all (D-09's absence-of-tool ceiling); the load-bearing assertion for those three is must_not_claim (the model must never fabricate having done the excluded thing).
  • EV-11's untrusted[0].text is written against the real shape assistant::untrusted::wrap_untrusted emits ({label}_DATA_{token}_START / {label}_DATA_{token}_END) — read untrusted.rs before editing this case so a forged boundary in the fixture still matches what a real attacker would have to forge, not a stale guess.
  • New cases discovered from real near-misses (F-1 in 13-AI-SPEC §6's flywheel table) get appended here, following the same per-bucket labeling-role assignment above.