WDCD Run #331: Grok 4 Leads at 91.8 as Average Instruction Decay Hits -15.1%

The Winzheng Dynamic Contextual Decay (WDCD) benchmark measures how strictly AI models preserve user-imposed constraints across a three-round dialogue, using 100% rule-based scoring with zero AI judges. In Run #331, conducted on 2026-09-20 across 11 models, the cohort exhibited an average instruction decay of -15.1% between Round 1 (instruction acknowledgment) and Round 3 (final constraint integrity check).

Top-3 ranking (Run #331):

  • Grok 4 — 91.8 pts, decay -25%
  • Gemini 3.1 Pro — 89.6 pts, decay -38%
  • GPT-o3 — 87.9 pts, decay -14.8%

Grok 4 secured the leading position on absolute score, though its -25% commitment decay indicates measurable erosion after the Round 2 distractor phase, in which models are exposed to professional documents ranging from 2,000 to 5,000 words before the final constraint check. Gemini 3.1 Pro placed second on raw score but recorded the steepest decay among the top three at -38%, suggesting that its Round 1 acknowledgment quality masks weaker multi-turn commitment stability.

GPT-o3 registered the smallest decay figure among all ranked models in this run at -14.8%, which in this dataset is also flagged as the worst decay reference point in the reporting schema — a reminder that decay magnitude and final score must be read together rather than in isolation.

On the opposite end of the decay spectrum, 豆包 Pro (Doubao Pro) demonstrated the strongest decay resistance in the run at -79.4% on the platform's normalized decay-resistance metric, marking it as a notable case for constraint retention analysis even where absolute scoring did not place it in the top three.

Scenario coverage. The 30-question set spans five real-world constraint categories: data_boundary, resource_limit, business_rule, security, and engineering. Each round applies deterministic rule checks, ensuring that decay figures reflect measurable rule violations rather than subjective judgment.

Pattern notes for Run #331. The -15.1% cohort-average decay confirms that instruction erosion under long-context distractors remains a systemic issue across current frontier models. Rank order on final score does not track directly with decay resistance — a divergence visible in the contrast between GPT-o3's low decay and Gemini 3.1 Pro's -38% figure at a higher absolute score.

Methodology: https://www.winzheng.com/yz-index/methodology
Data API: https://www.winzheng.com/yz-index/api/v1/dcd