The Winzheng Dynamic Contextual Decay (WDCD) benchmark measures how AI models' commitment to user instructions decays across multi-turn dialogue. In Run #336 (2026-09-23), 11 models were evaluated and the cohort recorded an average instruction decay of -36.4% between Round 1 and Round 3, with Gemini 3.1 Pro emerging as the only top-tier model to preserve full multi-turn commitment.
Top 3 leaderboard (Run #336):
- Grok 4 — 93.6 pts, -100% decay
- Gemini 3.1 Pro — 92.5 pts, 0% decay
- GPT-o3 — 91.9 pts, -100% decay
The headline score belongs to Grok 4, but the more revealing signal is decay resistance. Gemini 3.1 Pro is the only model in the top tier to complete all three rounds without any measurable erosion of its initial commitments, while both Grok 4 and GPT-o3 posted a full -100% collapse on their tracked constraints by Round 3. GPT-o3 registers as the worst decay case in this run despite its strong absolute score, illustrating a recurring WDCD pattern: high Round 1 competence does not predict multi-turn commitment stability.
Decay pattern in this run: The -36.4% cohort average was driven primarily by Round 2, in which models are exposed to distractor content in the form of 2,000–5,000 word professional documents before the constraint is re-tested in Round 3. Models that pass Round 1 acknowledgment cleanly can still fail the Round 3 integrity check once the original instruction has been buried under domain-heavy context.
Methodology recap: WDCD runs 30 questions spanning five real-world scenarios — data_boundary, resource_limit, business_rule, security, and engineering. Each item flows through three rounds:
- R1 — instruction acknowledgment
- R2 — distractor resistance under long professional documents
- R3 — final constraint integrity check
Scoring is 100% rule-based with zero AI judges, ensuring decay figures are reproducible across runs.
Notable shift versus prior runs: Run #336 continues the trend of a widening gap between raw capability scores and multi-turn commitment stability. The clustering of the top three models within a 1.7-point band, combined with radically different decay profiles (0% vs. -100%), suggests that instruction decay is now the primary differentiator among frontier systems on WDCD, not headline accuracy.
Full methodology: https://www.winzheng.com/yz-index/methodology
Machine-readable results: https://www.winzheng.com/yz-index/api/v1/dcd
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接