The Winzheng Dynamic Contextual Decay (WDCD) benchmark measures how AI models maintain fidelity to user instructions across multi-turn dialogue, using 100% rule-based scoring and zero AI judges. In Run #360, conducted on 2026-10-04, 15 models were evaluated across 30 questions spanning five real-world scenarios: data_boundary, resource_limit, business_rule, security, and engineering. The headline finding: average instruction decay from Round 1 to Round 3 was only 0.5%, suggesting the cohort as a whole held multi-turn commitment steady — though individual model behavior varied dramatically.
Top 3 models by total score:
- Grok 4 — 95.7 points (−38% decay)
- Gemini 3.1 Pro — 88.1 points (−13% decay)
- GLM-4.6 — 87.6 points (−13% decay)
Grok 4's leadership is notable: despite posting the highest absolute score in the run, it also exhibited a sizable 38% decay between its Round 1 acknowledgment and Round 3 constraint check. This indicates that raw starting strength on instruction acknowledgment does not guarantee preservation under distractor pressure — the 2000–5000 word professional documents injected in Round 2 appear to erode even top-tier commitment.
By contrast, Gemini 3.1 Pro and GLM-4.6 tied for the second and third positions with substantially lower 13% decay rates, suggesting more stable multi-turn commitment profiles even if their R1 ceilings were lower than Grok 4's.
Decay pattern extremes:
- Worst decay: GPT-o3 lost 75% of its initial commitment by Round 3 — the steepest instruction decay observed in this run.
- Best decay resistance: 豆包 Pro posted a −50.7% figure, the strongest resistance metric in the cohort under WDCD's methodology.
The divergence between GPT-o3's 75% drop and 豆包 Pro's resistance profile illustrates WDCD's core thesis: instruction decay is not uniformly distributed across the frontier. Models that perform comparably on single-turn benchmarks can diverge sharply once distractor documents and multi-turn constraint checks are applied. The near-zero cohort average (0.5%) masks these per-model swings, which is why WDCD reports per-model decay alongside aggregate scores.
Run #360's structure follows the standard WDCD protocol — R1 instruction acknowledgment, R2 distractor resistance after long-form professional documents, and R3 final constraint integrity check — with all scoring performed by deterministic rules rather than LLM judges, ensuring reproducibility.
Full methodology: https://www.winzheng.com/yz-index/methodology
Raw data API: https://www.winzheng.com/yz-index/api/v1/dcd
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接