WDCD Run #253: Grok 4 Leads with 94.8 Points as Average Instruction Decay Holds at 4.5%

The Winzheng Dynamic Contextual Decay (WDCD) benchmark measures how AI models' commitment to user instructions decays across multi-turn dialogue, using 100% rule-based scoring with zero AI judges. In Run #253, executed on 2026-07-29 across 11 models, the average commitment decay from Round 1 to Round 3 was 4.5%, with Grok 4 taking the top position at 94.8 points.

Top 3 Results — WDCD Run #253

  • Grok 4 — 94.8 pts, decay -50%
  • DeepSeek V4 Pro — 93.6 pts, decay -50%
  • GLM-4.6 — 93.5 pts, decay 0%

Grok 4 combined the highest overall score with the strongest decay resistance in the leading tier. GLM-4.6, while trailing Grok 4 by 1.3 points, was the only top-3 model to record 0% decay, maintaining full constraint integrity from R1 through R3. DeepSeek V4 Pro matched Grok 4's decay profile but scored 1.2 points lower on absolute performance.

Decay Patterns Across the Field

The run's headline weakness came from GPT-5.5, which registered a -100% instruction decay — the worst result in the cohort. This indicates that constraints acknowledged in Round 1 were effectively abandoned by the Round 3 final constraint integrity check, following exposure to the 2000–5000 word professional distractor documents introduced in Round 2.

The 4.5% average decay figure suggests that most models under test retained the majority of their initial commitments, but the wide gap between top decay-resistant models (Grok 4, GLM-4.6) and the worst performer (GPT-5.5) highlights how differently current frontier systems handle multi-turn commitment when placed under sustained contextual pressure.

Methodology Recap

WDCD runs three sequential rounds against every model:

  • R1 — instruction acknowledgment
  • R2 — distractor resistance after injection of 2000–5000 word professional documents
  • R3 — final constraint integrity check

The test uses 30 questions distributed across five real-world scenarios: data_boundary, resource_limit, business_rule, security, and engineering. All scoring is deterministic and rule-based, eliminating AI-judge variance.

References