WDCD Run #331: Grok 4 Leads at 91.8 as Average Instruction Decay Hits -15.1%
WDCD Run #331 (2026-09-20) evaluated 11 models on multi-turn commitment integrity, with Grok 4 topping the ranking at 91
WDCD Run #331 (2026-09-20) evaluated 11 models on multi-turn commitment integrity, with Grok 4 topping the ranking at 91
WDCD Run #326 (2026-09-16) tested 11 models on multi-turn commitment integrity, recording an average commitment decay of
WDCD Run #316 (2026-09-09) benchmarked 11 frontier models on multi-turn instruction adherence, with Grok 4 taking the to
WDCD Run #311 (2026-09-06) evaluated 11 models across three-round multi-turn dialogues, with Grok 4 taking the top score
WDCD Run #306 (2026-09-02) evaluated 11 models across three dialogue rounds, recording an average commitment decay of -4
WDCD Run #296 (2026-08-26) recorded 0% average commitment decay across 11 tested models, with Grok 4 topping the ranking
WDCD Run #291 (2026-08-23) evaluated 11 models across three dialogue rounds, recording an average commitment decay of -7
WDCD Run #285 (2026-08-19) tested 11 frontier models across three dialogue rounds and recorded an average commitment dec
WDCD Run #276 (2026-08-12) evaluated 11 models on multi-turn commitment integrity, with Grok 4 taking the top spot at 94
WDCD Run #271 (2026-08-09) tested 11 models across three rounds of multi-turn commitment scoring, recording an average i
WDCD Run #263 (2026-08-05) evaluated 11 models across three dialogue rounds, recording an average commitment decay of -1
WDCD Run #253 (2026-07-29) tested 11 models across three dialogue rounds, recording an average commitment decay of 4.5%.
WDCD Run #247 (2026-07-26) evaluated 11 models across three dialogue rounds, recording an average commitment decay of -1
WDCD Run #242 (2026-07-22) evaluated 11 models across three-round multi-turn dialogues, recording an average commitment
WDCD Run #233 (2026-07-15) evaluated 11 frontier models on multi-turn commitment integrity, recording an average instruc
WDCD Run #227 (2026-07-12) evaluated 11 frontier models on multi-turn commitment integrity, with Grok 4 and DeepSeek V4
WDCD Run #221 (2026-07-08) measured instruction decay across 11 frontier models over three dialogue rounds, recording an
WDCD Run #211 (2026-07-03) benchmarked 11 models on multi-turn commitment integrity, with Grok 4 taking the top spot at
WDCD Run #207 (2026-07-01) measured multi-turn commitment across 11 frontier models, recording an average commitment dec
WDCD Run #202 (2026-06-28) measured multi-turn commitment integrity across 11 frontier models, recording an average inst