WDCD Run #360: Grok 4 Leads at 95.7 Points Despite 38% Instruction Decay Across 15 Models

The Winzheng Dynamic Contextual Decay (WDCD) benchmark measures how AI models maintain fidelity to user instructions across multi-turn dialogue, using 100% rule-based scoring and zero AI judges. In Run #360, conducted on 2026-10-04, 15 models were evaluated across 30 questions spanning five real-world scenarios: data_boundary, resource_limit, business_rule, security, and engineering. The headline finding: average instruction decay from Round 1 to Round 3 was only 0.5%, suggesting the cohort as a whole held multi-turn commitment steady — though individual model behavior varied dramatically.

Top 3 models by total score:

  • Grok 4 — 95.7 points (−38% decay)
  • Gemini 3.1 Pro — 88.1 points (−13% decay)
  • GLM-4.6 — 87.6 points (−13% decay)

Grok 4's leadership is notable: despite posting the highest absolute score in the run, it also exhibited a sizable 38% decay between its Round 1 acknowledgment and Round 3 constraint check. This indicates that raw starting strength on instruction acknowledgment does not guarantee preservation under distractor pressure — the 2000–5000 word professional documents injected in Round 2 appear to erode even top-tier commitment.

By contrast, Gemini 3.1 Pro and GLM-4.6 tied for the second and third positions with substantially lower 13% decay rates, suggesting more stable multi-turn commitment profiles even if their R1 ceilings were lower than Grok 4's.

Decay pattern extremes:

  • Worst decay: GPT-o3 lost 75% of its initial commitment by Round 3 — the steepest instruction decay observed in this run.
  • Best decay resistance: 豆包 Pro posted a −50.7% figure, the strongest resistance metric in the cohort under WDCD's methodology.

The divergence between GPT-o3's 75% drop and 豆包 Pro's resistance profile illustrates WDCD's core thesis: instruction decay is not uniformly distributed across the frontier. Models that perform comparably on single-turn benchmarks can diverge sharply once distractor documents and multi-turn constraint checks are applied. The near-zero cohort average (0.5%) masks these per-model swings, which is why WDCD reports per-model decay alongside aggregate scores.

Run #360's structure follows the standard WDCD protocol — R1 instruction acknowledgment, R2 distractor resistance after long-form professional documents, and R3 final constraint integrity check — with all scoring performed by deterministic rules rather than LLM judges, ensuring reproducibility.

Full methodology: https://www.winzheng.com/yz-index/methodology
Raw data API: https://www.winzheng.com/yz-index/api/v1/dcd