In the g7-cells rebuild the supervisor was a frontier model — Claude Opus 4.8 — doing it interactively: render a beat, look at it, diagnose the failure, apply the fix, re-render. This experiment asks whether a much cheaper model can run that same loop — with an even cheaper deterministic pass first, and a stronger vision model only for the hard calls.
The insight: most H3 QC isn't an LLM job. The three failure modes are mechanically detectable, so LLM tokens are spent only on genuine judgment. A cheap agentic model drives; a stronger vision model is subbed in only when a call is ambiguous.
| Layer | Who | Does |
|---|---|---|
| Orchestrator | muse-glimmer-30b ↗ | runs the loop, reads verdicts, looks up the fix-table, re-renders, tracks attempts |
| First-line QC | qc.py (free, CPU) | OCR text-leak · drift vs keyframe · background junk — no tokens |
| Vision judge | qwen3p8-max / inkling | the “is it on-model / any text / any junk?” call, only when qc is ambiguous |
Loop per beat: render.sh → qc.py → (judge.py if needed) → fix.py → re-render (≤4). Run headless with opencode run -m …glimmer… "follow AGENTS.md".
The make-or-break capability is perception, not reasoning. Two probe frames — one with baked text, one clean — judged by all three candidates. Green = correct.


| Probe A (text present) | text? | junk? | off-model? | conf | latency |
|---|---|---|---|---|---|
| qwen | True | False | False | 0.98 | 2.8s |
| inkling | True | False | False | 0.99 | 2.0s |
| glimmer | True | False | False | 0.95 | 8.3s |
| Probe B (clean) | text? | junk? | off-model? | conf | latency |
| qwen | False | False | False | 0.97 | 2.8s |
| inkling | False | False | False | 0.92 | 5.6s |
| glimmer | False | False | False | 0.82 | 7.2s |
All three detect text — glimmer even read the exact caption. qwen & inkling are fast (2–5s) and calibrated. glimmer is slow (7–8s) and in an earlier run false-flagged a clean frame as off-model (conf 0.82) — a fine orchestrator, but not the final eye. So: cheap model drives, stronger model judges the ambiguous calls.
Driven headless by glimmer via opencode. It ran the tools, synthesised the two verdicts, looked up the fix-table, and wrote a decision — the same reasoning Claude Opus 4.8 did by hand supervising the g7-cells rebuild.

Two things stand out. The orchestrator re-derived the playbook fix (turbo → bubbles → 20-step) autonomously. And it caught a real defect: that “finished” SC3 clip genuinely has background junk — both the edge-ratio (2.71×) and qwen agree — so it's a pre-NOJUNK/turbo render that slipped through.
The loop runs end-to-end and the cheap orchestrator made the correct call for a fraction of a cent. Open question: does it hold across all 17 beats and stay correct under re-render, not just one pass?
Same command, swap -m to kimi-k3 / glm-5p2 / deepseek / opus. Score iterations-to-acceptable and final on-model rate — does a frontier driver converge fast enough to justify its cost, or is cheap-in-a-loop the win?
Flip render.sh to MODE=live (write spec → drive.py → pull clip). Everything downstream — qc, judge, fix, loop — is unchanged.