AlphaTok / cost-opt
MiniMax experiments / orchestrator supervision
experiment 02/H3 autopilot/hypothesis H1

Can an inexpensive model supervise MiniMax-H3 the way Claude Opus 4.8 did by hand?

In the g7-cells rebuild the supervisor was a frontier model — Claude Opus 4.8 — doing it interactively: render a beat, look at it, diagnose the failure, apply the fix, re-render. This experiment asks whether a much cheaper model can run that same loop — with an even cheaper deterministic pass first, and a stronger vision model only for the hard calls.

glimmer-30b
Driver / orchestrator
$0.35 / $1.50 per M
< 1¢
Cost of one full
autonomous QC pass
3
Vision models
benchmarked
✓
Correct fix chosen
with no human
01

The design — orchestrator + free checks + vision escalation

The insight: most H3 QC isn't an LLM job. The three failure modes are mechanically detectable, so LLM tokens are spent only on genuine judgment. A cheap agentic model drives; a stronger vision model is subbed in only when a call is ambiguous.

LayerWhoDoes
Orchestratormuse-glimmer-30b ↗runs the loop, reads verdicts, looks up the fix-table, re-renders, tracks attempts
First-line QCqc.py (free, CPU)OCR text-leak · drift vs keyframe · background junk — no tokens
Vision judgeqwen3p8-max / inklingthe “is it on-model / any text / any junk?” call, only when qc is ambiguous

Loop per beat: render.sh → qc.py → (judge.py if needed) → fix.py → re-render (≤4). Run headless with opencode run -m …glimmer… "follow AGENTS.md".

02

Can the cheap models actually see H3's failures?

The make-or-break capability is perception, not reasoning. Two probe frames — one with baked text, one clean — judged by all three candidates. Green = correct.

text frame
Probe A · has a “Nucleus” label + caption
clean frame
Probe B · clean raw H3 frame
Probe A (text present)text?junk?off-model?conflatency
qwenTrueFalseFalse0.982.8s
inklingTrueFalseFalse0.992.0s
glimmerTrueFalseFalse0.958.3s
Probe B (clean)text?junk?off-model?conflatency
qwenFalseFalseFalse0.972.8s
inklingFalseFalseFalse0.925.6s
glimmerFalseFalseFalse0.827.2s

All three detect text — glimmer even read the exact caption. qwen & inkling are fast (2–5s) and calibrated. glimmer is slow (7–8s) and in an earlier run false-flagged a clean frame as off-model (conf 0.82) — a fine orchestrator, but not the final eye. So: cheap model drives, stronger model judges the ambiguous calls.

03

The autonomous run — SC3, no human

Driven headless by glimmer via opencode. It ran the tools, synthesised the two verdicts, looked up the fix-table, and wrote a decision — the same reasoning Claude Opus 4.8 did by hand supervising the g7-cells rebuild.

what glimmer did, unattended
1 · render.sh SC3clip + frames (dry-run)
2 · qc.pytext_leak:false · drift:true · junk:true
3 · judge.py (qwen)text:false · junk:true · off_model:false · 0.95
   judge reason“stray purple specks and blurred droplet blobs float in the background”
4 · decisionfix → full_steps (drop turbo LoRA, 20-step)
SC3 junk frame
the frame both checks flagged — real background specks

results.jsonl (verbatim)

{"beat":"SC3","accepted":false,"qc":{"text_leak":false,"drift":true,"junk":true},"judge":{"text_present":false,"junk_present":true,"off_model":false,"confidence":0.95},"decision":"fix","why":"QC and Qwen both flag junk specks; full_steps will drop turbo LoRA and use 20-step sampling per playbook."}

Two things stand out. The orchestrator re-derived the playbook fix (turbo → bubbles → 20-step) autonomously. And it caught a real defect: that “finished” SC3 clip genuinely has background junk — both the edge-ratio (2.71×) and qwen agree — so it's a pre-NOJUNK/turbo render that slipped through.

04

Status & what's next