Kid mode 🧒

This page is Tarot V3 Gate Results. It is one of the projects Brandon built and put online. The grown-up details are below, tap Grown-up mode to read them the normal way.

Gate Verdicts

FAIL

PRE-CONTINUATION-FIX (567 words)

v3_rag_final_marie.PRE-CONTINUATION-FIX-2026-06-04.md
0.580
threshold: 0.65  |  gap: -0.070
9 runs: [0.15, 0.50, 0.55, 0.55, 0.58, 0.60, 0.70, 0.70, 0.90]
range: 0.75  |  IQR: 0.15
Hard fails: 0 Anti-AI: PASS Voice score: FAIL
Truncated mid-sentence at 567 words. No structural conclusion or Three of Pentacles outcome section.
FAIL

FIXED (792 words, complete)

v3_rag_final_marie.md
0.600
threshold: 0.65  |  gap: -0.050
9 runs: [0.00, 0.30, 0.40, 0.45, 0.60, 0.67, 0.70, 0.71, 0.75]
range: 0.75  |  IQR: 0.30
Hard fails: 0 Anti-AI: PASS Voice score: FAIL
Complete reading with 3 action items, Three of Pentacles outcome section, and disclaimer.

Score Delta: Fixed vs PRE-FIX

+0.020
The continuation fix moved the Kate median score by +0.02 (0.58 to 0.60). With per-run variance of 0.30-0.75, this delta is within noise and is not statistically meaningful. Both files fail the 0.65 gate. The fix did not demonstrably move quality in the Kate judge's view.
File (a) not tested: /tmp/v3_marie_chiron_fix.md (the chiron file-set postfix from q-2026-06-05-83de14) was ephemeral and not found at gate time. The chiron leg of this validation cannot be compared without re-running the chiron file-set generator.

Variance Analysis

MetricPRE-FIX (567w)FIXED (792w)Note
Median (9 runs)0.5800.600Delta: +0.020
Min0.150.00Fixed has a 0.0 outlier (judge saw no match)
Max0.900.75PRE-FIX had more high-end outliers
Range0.750.75Identical noise floor
Runs above threshold (0.65)3 / 94 / 9Fixed slightly more consistent above gate
Hard fails00Both clean on em-dash, ellipsis, correctio, etc.
Anti-AI gatePASSPASSBoth clean
VerdictFAILFAILBoth below 0.65 median threshold

Interpretation

What this means for V3

  1. The variance problem is not solved by n=9. The 9-run spread is 0.75 wide for both files. The n=9 CI is tighter than n=3 by sqrt(3) as expected, but the underlying judge variance is high enough that even the tighter median cannot detect a 0.02 shift as signal vs. noise.
  2. The continuation fix did not move quality detectably. From the Kate judge's perspective, a truncated 567-word reading and a complete 792-word reading score the same. The quality ceiling is not in the ending structure, it is in the RAG retrieval + generation quality throughout.
  3. The fixed file has 4/9 runs above threshold vs. 3/9. This is a marginal improvement, but not reliable. True gate passage needs the median above 0.65 consistently.
  4. Gap to 0.65 is -0.05 for the fixed file. That gap is within the variance range, meaning on any given gate run the file might pass or fail depending on DeepSeek's sampling. A genuine pass would require the true score to be around 0.72+ so the median reliably exceeds 0.65.
  5. Next move: The RAG grounding or the generation temperature/prompt needs adjustment to push the median above 0.70. The structure and voice-rule compliance are already clean (zero hard fails, anti-AI passing). The shortfall is in the Kate-specific phrasing patterns that the LLM judge scores on.

Queue Action

Queue ItemAction
q-2026-06-05-bbe398 (this task)Marked DONE. Both (b) and (c) gated; (a) missing.
q-2026-06-04-ca8c2e (V3 Phase 2D deliverable bundler)Still BLOCKED. Gate must pass before bundle run.
q-2026-06-04-a9cb77 / c9cc2f / d1c059 / 8acf2e (V2 Phase 1C batch)Still BLOCKED on 1B gate (separate from this validation).