STAYCHARTED AMT / INSIGHTS
AI classification benchmarks on real data
Compare measured text and image classification results, the data behind them, and the review work that remains.
Measured results at a glance
| Dataset | Task · categories | Approach | Accuracy | Review share and projection | Report date / method |
|---|---|---|---|---|---|
| BANKING77 | Customer messages · 77 | Dedicated AI Model | 93.7% | 9.2% → projected 98.0% | 3,080 held-out test messages |
| LEDGAR | Contract clauses · 100 | Dedicated AI Model | 84.9% | 41.7% → projected 97.6% | Separate test set; four approaches compared |
| Catalog product photos | Pictures · 40 | AI Image Classifier Model (SigLIP 2) | 91.1% | Not measured | All 7,966 training pictures; see article split |
| Cat and dog breeds | Pictures · 37 | AI Image Classifier Model (SigLIP 2) | 95.2% | Not measured | All 3,680 training pictures; see article split |
Review figures are projections assuming flagged errors are corrected without introducing new errors. They are not measured human-review results. Datasets, model sizes, and test conditions differ; compare the full reports before applying these results to your task.
Read a benchmark in context
A score describes a model on a particular dataset and split. Check the categories, number of examples, model type, and errors before applying the result to your own task. A high overall score does not mean every category performs equally well.
Accuracy and review are different measurements
Held-out accuracy measures the model’s answers before a person changes them. Review projections estimate what could happen if flagged mistakes were corrected. Keep those figures separate when comparing approaches or planning a workflow.
EXPLORE NEXT
Try it on your own examples.
Start with a Free workspace for text classification. Compare plans for more AI capacity, Vision, and API access.