WDCD Run #247: Grok 4 Leads with Negative Decay as Average Instruction Decay Narrows to -1.8%

The Winzheng Dynamic Contextual Decay (WDCD) benchmark measures how AI models' commitment to user instructions decays across multi-turn dialogue, using 100% rule-based scoring with zero AI judges. In Run #247, executed on 2026-07-26 across 11 models, the average instruction decay from Round 1 to Round 3 was -1.8%, with top-ranked models actually strengthening compliance rather than losing it.

Ranking highlights. The top three placements in Run #247 were:

  • Grok 4 — 94.2 points, --63% decay (best decay resistance in the run)
  • DeepSeek V4 Pro — 87.0 points, --25% decay
  • GLM-4.6 — 83.9 points, -0% decay (fully stable across all three rounds)

Decay patterns. The three-round protocol probes different failure modes: R1 measures initial instruction acknowledgment, R2 tests distractor resistance after injecting 2,000–5,000 word professional documents, and R3 performs a final constraint integrity check. Negative decay values — as seen for Grok 4 and DeepSeek V4 Pro — indicate that models tightened compliance under distractor pressure rather than drifting, a pattern consistent with stronger internal representation of user constraints under context load. GLM-4.6's flat 0% decay represents pure stability: identical constraint adherence in R1 and R3.

At the opposite end of the spectrum, Qwen3 Max recorded the worst result of the run at -50% decay, losing roughly half of its Round 1 multi-turn commitment by Round 3. This gap between top and bottom — a spread of over 100 percentage points in decay behavior — underscores that headline single-turn accuracy figures do not predict how a model holds a constraint under sustained context.

Scenario coverage. Run #247 used the standard WDCD set of 30 questions distributed across five real-world scenarios: data_boundary, resource_limit, business_rule, security, and engineering. All scoring is deterministic and rule-based, meaning results are fully reproducible from the raw dialogue transcripts.

Read from previous run. The narrowing of the fleet-wide average decay to -1.8% suggests that top-tier models are converging toward stable or improving multi-turn commitment, while a long tail — represented by Qwen3 Max in this run — continues to exhibit substantial drift under distractor load.

For full protocol details, see the WDCD methodology. Raw per-model, per-round data is available via the WDCD data API.