← All insights

Developing AMT

Asking an assistant vs. training a classifier: three benchmarks

Compare Claude Sonnet 5.5 with StayCharted trained classifiers on banking messages, contract clauses and medical abstracts. See all results, method and limits.

By Manoj Mohandas · Measurements: October 6, 2026 · Published · Updated

A trained model was clearly better on two datasets. On the third, they were close.

Should you ask an assistant to categorize your text, or train a model on labeled examples? We compared both workflows on three public datasets using StayCharted AI Model Trainer. The answer depended on the dataset.

AI Classifier Model and Dedicated AI Model are training options in StayCharted AI Model Trainer (AMT). Claude Sonnet 5.5 is the external assistant used for comparison.

Accuracy on the same held-out test rows per dataset
EvaluationExternal assistantStayChartedAI Model TrainerDifference
Dataset · taskTest rowsClaude Sonnet 5.5AI Classifier ModelPre-TrainedDedicated AI ModelStayCharted Dedicated vs. Claude
BANKING77 · 77 support intents3,08084.7%92.6%93.7%+9.0 points
LEDGAR · 100 clause types4,00076.9%82.7%84.9%+8.0 points
Medical Abstracts · 5 condition groups2,88865.7%60.0%63.6%−2.1 points

The AI Classifier Model tries meaning alone and words plus meaning, and keeps whichever scores better on held-back validation rows. In the Medical Abstracts benchmark, meaning alone scored 60.0% against 52.0% for words plus meaning, so 60.0% is shown. These were separately benchmarked approaches; the product’s automatic selection workflow was not run end to end on this dataset.

Put as mistakes per 100 rows: on BANKING77, the Dedicated AI Model made 6.3 against Claude’s 15.3, fewer than half. On LEDGAR it made 15.1 against 23.1, about a third fewer. On Medical Abstracts, Claude made 34.3 and the trained model 36.4.

How much weight to put on a gap depends on the test size and method. Approximate 95% binomial margins for each accuracy range from ±0.9 to ±1.8 percentage points across these tests. Those estimates describe sampling uncertainty under independent-row assumptions, not variation between training runs or performance on new data. The banking and legal gaps are much larger. The medical results are close, but overlapping individual intervals do not establish a tie: a paired analysis of the same rows is needed to assess the 2.1-point difference. We have not performed that analysis.

These are test-set measurements, not guarantees for your data. Percentages and error comparisons are rounded.

Same test rows, different ways to learn the task

Claude received category names and a one-line definition of each, with no labeled examples. StayCharted’s models learned from about 10,000 labeled training rows per dataset: 10,003 for BANKING77 and 10,000 each for LEDGAR and Medical Abstracts. Test rows were held out of StayCharted training.

This compares a definition-based assistant workflow with supervised training. It does not isolate the effect of model architecture or compare equal access to examples.

  • Assistant: Claude Sonnet 5.5 through Amazon Bedrock, recorded in the reports as us.anthropic.claude-sonnet-5-5, with default settings and a request for exactly one category per row.
  • Output limit and retries: the initial limit was 40 output tokens. Empty or cut-off replies were retried with the same prompt and 1,024 tokens; the final category was used. This affected 408 banking, 138 legal and 857 medical replies.
  • Invalid answers: after retries, zero banking, two legal and zero medical replies named no valid category. These were counted as wrong.
  • Trained approaches: the AI Classifier Model runs on CPU; the Dedicated AI Model used a fine-tuned Qwen3.5-4B model on GPU. Their existing reports used the same frozen test rows.
  • Date: the assistant reports are dated October 6, 2026. Results are from this setup, not an average across multiple complete benchmark runs.

Examples help most when the categories are your conventions

The banking and legal errors show why category boundaries deserve attention. On BANKING77, Claude labeled 32 of the 40 get_physical_card messages as change_pin. On LEDGAR, it labeled 34 “Waivers” clauses as “No Waivers” and 34 “Definitions” clauses as “Defined Terms.”

Labeled examples can demonstrate distinctions that a short definition leaves unclear. These results are consistent with that explanation, but do not prove the cause of each error. Prompt definitions, dataset conventions and training inputs can all affect the comparison.

These errors concern labeling conventions as well as general knowledge. Knowing what a waiver clause is does not necessarily settle how a particular dataset distinguishes “Waivers” from “No Waivers,” or which of two plausible banking intents a team uses. Labeled examples make those boundaries concrete. This is a plausible explanation for the pattern, not a cause isolated by our experiment.

Medical Abstracts shows the other side. Its five groups—cardiovascular, digestive, nervous system, neoplasms and general pathological conditions—draw on general medical knowledge. There, Claude’s 65.7% and the trained model’s 63.6% were close. Both struggled with the broad “General pathological conditions” group, where Claude scored 40.2%; overlap with the other groups may help explain that difficulty. Categories grounded in general knowledge are a reason to try an assistant. Categories shaped by your team’s conventions are a reason to compare training on your own examples.

Why not teach Claude your categories?

The natural next step is to give the assistant your examples. You can, and this benchmark did not test it. That workflow still needs a way to manage examples, evaluate results and apply corrections:

  • Examples need to be available in context. A prompt-based workflow supplies examples or retrieves relevant ones when classifying. Five examples for each of 77 or 100 categories means 385 or 500 examples if you include every category. Saved context and prompt caching can reduce repetition; they do not turn those examples into a model trained on your categories.
  • Choosing examples is part of the work. Which categories need examples, and which examples explain the boundaries? Retrieving similar past rows is an option, but that selection process needs configuration and evaluation too.
  • Corrections need a path back into the workflow. Update the shared prompts, examples or retrieval source when someone fixes a label. This can be automated. In StayCharted, reviewed corrections can become training examples for the next run; they do not change the published model immediately.
  • Accuracy still needs measurement. Test either workflow on held-out rows with known answers, including the categories it confuses. StayCharted provides that evaluation when you train a model.
  • Confidence needs testing on your rows. A confidence value stated by an assistant is not automatically calibrated to your task. StayCharted shows how confidence cutoffs behave on evaluated rows, helping you choose what to review. Confident mistakes can still pass through.
  • Version changes need re-evaluation. Assistant model updates or retirements can change classification behavior. Shared, versioned prompts help you control the workflow. StayCharted keeps the published trained version in use until you publish another.
  • Teams need one maintained definition of the task. Different prompts can produce different interpretations. You can standardize an assistant’s instructions; a shared trained model is another way to apply the same category rules across your team.
  • Volume changes the engineering tradeoffs. Assistant APIs can batch rows, cache prompts and retrieve examples. Those approaches still consume model input and output capacity. An AI Classifier Model runs on CPU without sending a prompt and examples with each prediction. This benchmark did not measure the cost or speed difference.

None of this makes an assistant the wrong tool. For a one-off sort, or categories that follow general knowledge, it may be enough. For recurring work in your own categories, StayCharted puts the examples, measurements and corrections in one place. You can also connect your assistant to a model you have trained and tested.

Claude was mostly consistent when asked twice

We repeated the identical request for the first 500 rows of each dataset. Claude returned the same category on 99.2% of banking rows, 98.8% of legal rows and 97.0% of medical rows.

That is agreement across two attempts on a subset, not accuracy or a test of different prompts. Consistency alone is not the strongest reason to train a model here; matching the labels and measuring mistakes are more useful considerations.

How to choose between an assistant and a trained classifier →

Choose based on your examples and review needs

Ask yourself where your categories come from. If they follow general knowledge and you have no labeled examples, start with an assistant and check a sample against the answers you expect. If they are your team’s own conventions and you already have categorized examples, train a model on them and compare it on rows it never saw. Either way, look at the mistakes per category, not just the overall score.

StayCharted can use prediction confidence to help prioritize human review. Confidence is not a guarantee: some incorrect predictions are confident, and new data can differ from the test set.

This benchmark did not measure cost or end-to-end speed. It did not test Claude’s chat interface, ChatGPT, connectors, or Sheets. It also did not test giving Claude labeled examples, alternative definitions or prompt optimization. Public datasets may have appeared in a pretrained model’s training material; we did not audit that exposure.

You can also combine the workflows: ask your assistant to use a model you have evaluated, with the workspace and model permissions you choose.

To try it on your own examples, see the models or start free.

Sources and further results

StayCharted’s comparison reports provide the measurements above. The dataset sources below describe the public corpora; they are not independent verification of our results.

Frequently asked questions

Is a trained classifier more accurate than Claude?

On two of three datasets, the trained model had a clear accuracy lead in this test. StayCharted’s Dedicated AI Model scored 93.7% against Claude’s 84.7% on BANKING77, and 84.9% against 76.9% on LEDGAR. On Medical Abstracts the two were close: Claude 65.7%, the trained model 63.6%. The largest gains occurred on tasks with dataset-specific category conventions; these results do not prove a universal rule.

Is a 2-point difference meaningful?

It depends on the test design. With 2,888 medical test rows, approximate 95% binomial margins are ±1.7 points for Claude and ±1.8 for the Dedicated AI Model, assuming independent rows. These are individual sampling intervals, not a paired test of the difference or a measure of run-to-run variation. The medical scores were close; a paired analysis is needed before claiming significance or equivalence. The banking and legal gaps were substantially larger at 9.0 and 8.0 points.

Did Claude receive the same training examples?

No. Claude received category names and one-line definitions without labeled examples. StayCharted’s models learned from about 10,000 labeled rows per dataset. This compares two workflows, not models with equal access to examples. Giving an assistant examples in context is possible, including through shared prompts or retrieval, but was not tested here. Those examples and their selection still need to be maintained and evaluated; see “Why not teach Claude your categories?” in this article.

Was this tested inside Claude or ChatGPT?

No. Claude Sonnet 5.5 was called through Amazon Bedrock. The Claude and ChatGPT interfaces, their connectors and spreadsheet integrations were not evaluated.

Does confidence guarantee a correct answer?

No. Confidence can help prioritize review, but confident predictions can still be wrong. Evaluate your model on representative examples and check performance on new data.

Put the findings into practice

Compare StayCharted models →

Connect Claude or ChatGPT →

Compare plans →

Explore more Insights

Compare AMT’s design decisions and training reports →

STAYCHARTED AI MODEL TRAINER

Train AI to categorize the way your team does.

Start with examples you already have, see how often it’s right on items it never saw, and use confidence to prioritize the ones a person should review.

Try it on your own data