WDCD Run #242: Grok 4 and GLM-4.6 Hold Zero Instruction Decay as Gemini 3.1 Pro Collapses at -100%

The Winzheng Dynamic Contextual Decay (WDCD) benchmark measures how strongly AI models preserve user instructions across multi-turn dialogue. In Run #242 (2026-07-22), 11 models were evaluated and the fleet-wide average instruction decay from Round 1 to Round 3 reached 18.2%, with a sharp bifurcation between models that held perfectly and one that fully collapsed.

Top 3 rankings:

  • Grok 4 — 93.8 pts, 0% decay
  • GLM-4.6 — 92.0 pts, 0% decay
  • DeepSeek V4 Pro — 90.9 pts, recorded as the run's best decay-resistance profile

At the other end of the distribution, Gemini 3.1 Pro registered the run's worst outcome with a -100% decay reading, indicating a total loss of Round-1 constraint adherence by Round 3. This is the sharpest single-model collapse observed against the Run #242 cohort average of 18.2%.

Decay patterns. The three-round protocol isolates where multi-turn commitment breaks down: R1 tests initial instruction acknowledgment, R2 injects 2,000–5,000 word professional distractor documents to stress context prioritization, and R3 performs a final constraint integrity check. In Run #242, the leading models absorbed the R2 distractor payload without shedding any R1 constraints, while the lowest-ranked model failed to carry any measurable constraint into R3.

Scoring integrity. WDCD uses 100% rule-based scoring with zero AI judges. The 30-question set spans five real-world scenario families: data_boundary, resource_limit, business_rule, security, and engineering. Because verification is deterministic, decay figures reflect measurable rule violations rather than stylistic judgments.

Notable observations from this run:

  • The gap between the top tier (0% decay) and the bottom (-100%) is the defining feature of Run #242 — decay is not gradual across the cohort but polarized.
  • The top three models cluster within a 2.9-point band (90.9–93.8), suggesting the frontier is now competitive on raw score, with decay resistance as the primary differentiator.
  • The 18.2% mean decay indicates that a meaningful portion of the cohort still loses instruction adherence once long-form distractor context is introduced in R2.

Full methodology, scoring rubrics, and scenario definitions: https://www.winzheng.com/yz-index/methodology

Structured data endpoint for Run #242 and historical runs: https://www.winzheng.com/yz-index/api/v1/dcd