What Changed

Anthropic released Claude 4 with substantially improved performance on structured output tasks. In internal benchmarks, the model achieves 94% accuracy on complex JSON extraction from unstructured documents — up from 81% in Claude 3.5.

Why It Matters for Data Teams

For organisations running claim verification pipelines, better structured extraction means fewer manual correction steps when pulling statistics, source citations, and numeric claims from long-form documents.

Caveats

  • Accuracy drops to ~78% on tables embedded in scanned PDFs
  • Numeric extraction from mixed-format documents still requires human spot-check
  • Context window limits apply to documents over ~200 pages

Our Take

The improvement is real but not yet sufficient to remove the human review layer from high-stakes verification. Use it to accelerate the first pass — not to replace sign-off.