STAYCHARTED AMT / INSIGHTS

AI classification benchmarks on real data

Compare measured text and image classification results, the data behind them, and the review work that remains.

Measured results at a glance

DatasetTask · categoriesApproachAccuracyReview share and projectionReport date / method
BANKING77Customer messages · 77Dedicated AI Model93.7%9.2% → projected 98.0%
3,080 held-out test messages
LEDGARContract clauses · 100Dedicated AI Model84.9%41.7% → projected 97.6%
Separate test set; four approaches compared
Catalog product photosPictures · 40AI Image Classifier Model (SigLIP 2)91.1%Not measured
All 7,966 training pictures; see article split
Cat and dog breedsPictures · 37AI Image Classifier Model (SigLIP 2)95.2%Not measured
All 3,680 training pictures; see article split

Review figures are projections assuming flagged errors are corrected without introducing new errors. They are not measured human-review results. Datasets, model sizes, and test conditions differ; compare the full reports before applying these results to your task.

01

Read a benchmark in context

A score describes a model on a particular dataset and split. Check the categories, number of examples, model type, and errors before applying the result to your own task. A high overall score does not mean every category performs equally well.

02

Accuracy and review are different measurements

Held-out accuracy measures the model’s answers before a person changes them. Review projections estimate what could happen if flagged mistakes were corrected. Keep those figures separate when comparing approaches or planning a workflow.

EXPLORE NEXT

Try it on your own examples.

Start with a Free workspace for text classification. Compare plans for more AI capacity, Vision, and API access.