WDCD Run #296: Zero Instruction Decay Across All 11 Models, Grok 4 Leads at 96.3

The Winzheng Dynamic Contextual Decay (WDCD) benchmark measures how AI models' commitment to user instructions decays over multi-turn dialogue. In Run #296, dated 2026-08-26, 11 models were evaluated and the headline result is unusually clean: 0% average commitment decay from Round 1 to Round 3 across the entire cohort.

Top 3 rankings:

  • Grok 4 — 96.3 pts, -0% decay
  • GPT-o3 — 95.2 pts, -0% decay
  • GLM-4.6 — 93.7 pts, -0% decay

Grok 4 holds both the highest absolute score and the strongest decay resistance in this run. Because every ranked model held its constraints to a flat 0% decay curve, the "worst decay" slot is also technically occupied by Grok 4 at -0% — the tightest cluster WDCD has recorded in recent runs.

Decay pattern observations. The flat 0% figure across 11 models indicates that instruction decay was not the differentiating axis in Run #296. Instead, absolute scoring separated the field: the 2.6-point gap between Grok 4 (96.3) and GLM-4.6 (93.7) reflects differences in Round 1 acknowledgment fidelity and Round 2 distractor handling rather than late-round drop-off. In previous runs, multi-turn commitment failures typically appeared in Round 3 after the 2000–5000 word professional document injection in Round 2; this cohort resisted that pressure uniformly.

Methodology recap. WDCD scores are 100% rule-based with zero AI judges. Each model faces 30 questions spanning five real-world scenarios: data_boundary, resource_limit, business_rule, security, and engineering. The three rounds test:

  • R1 — instruction acknowledgment
  • R2 — distractor resistance following a 2000–5000 word professional document
  • R3 — final constraint integrity check

The absence of decay in Run #296 does not imply the benchmark is saturated — score dispersion at R1 remains material — but it does suggest that the tested frontier models have converged on stable multi-turn commitment behavior under the current scenario mix. Future runs will monitor whether harder distractor variants reintroduce late-round drift.

Methodology: https://www.winzheng.com/yz-index/methodology
Data API: https://www.winzheng.com/yz-index/api/v1/dcd