The Winzheng Dynamic Contextual Decay (WDCD) benchmark measures how reliably AI models retain their commitment to user instructions across multi-turn dialogue. In Run #271 (2026-08-09), 11 models were evaluated, with the group-wide average instruction decay from Round 1 to Round 3 registering at only 0.9% — one of the tightest spreads recorded to date.
Top 3 Results
- Grok 4 — 91 pts (decay: −25%)
- DeepSeek V4 Pro — 89 pts (decay: −50%)
- Claude Sonnet 4.6 — 83.7 pts (decay: 0%)
Grok 4 took the top slot on absolute score, but the run's most notable pattern is the divergence between raw score and decay resistance. Claude Sonnet 4.6 was the only model in the top three to post 0% decay, sustaining full constraint integrity from R1 through R3 despite ranking third on total points. DeepSeek V4 Pro placed second overall but recorded a −50% decay, indicating that its high R1 acknowledgment score partially offset weaker R3 constraint retention.
At the other end of the distribution, GPT-5.5 posted the run's worst decay at −50%, matching DeepSeek V4 Pro's decay magnitude but from a lower base — a reminder that decay percentage alone does not translate directly into ranking position. The gap between best and worst decay-resistant models highlights that multi-turn commitment remains an uneven capability across current frontier systems, even when single-turn instruction following appears comparable.
Methodology recap. WDCD scores each model across three rounds:
- R1 — instruction acknowledgment
- R2 — distractor resistance following a 2,000–5,000 word professional document
- R3 — final constraint integrity check
The benchmark uses 30 questions across 5 real-world scenarios: data_boundary, resource_limit, business_rule, security, and engineering. Scoring is 100% rule-based with zero AI judges, ensuring that results reflect deterministic constraint checks rather than subjective evaluation.
The narrow 0.9% average decay in Run #271 suggests overall improvement in cross-turn stability at the population level, but the −50% outliers at both DeepSeek V4 Pro and GPT-5.5 indicate that specific failure modes — particularly under long-document distractor conditions in R2 — persist. Subsequent runs will track whether the population-level tightening holds or whether it reflects run-specific scenario weighting.
Full methodology: https://www.winzheng.com/yz-index/methodology
Data API: https://www.winzheng.com/yz-index/api/v1/dcd
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接