← All insights

Developing AMT

Why Our AI Classifier Reads Both Words and Meaning

Why StayCharted combines words and meaning for AI text classification, with banking and contract-clause benchmarks and human-review tradeoffs.

By StayCharted · Measurements: September 24, 2026 · Published

Most text that businesses label by hand is short and repetitive: support messages, complaint reasons, transaction descriptions, contract clauses, product titles. Our AI Classifier Model is built for that work. It uses a small pre-trained AI to understand what each row means, learns your categories from your own examples, and can train on short text in about a minute, with no GPU. Longer text takes more time.

We did not start there. We measured six ways of doing the job on two public benchmarks that pull in opposite directions, and let the numbers decide. This note explains what we measured, what we chose, and where the choice has limits.

The short version

  • Meaning wins on short, everyday text. On 77 banking intents with 10 examples each, reading meaning scored 81.3% and matching words scored 62–68%.
  • Words win on formulaic text. On 100 types of contract clause, word matching beat meaning at every tested training size, by roughly 3–4 points.
  • So the AI Classifier fits both, and keeps whichever does better on your data. It computes the meaning of every row once and fits two small classifiers: meaning alone, and words plus meaning together. It then scores both on rows it held back from your file and keeps the winner. You don't choose; your data does.
  • At full size it comes within about a point of a fine-tuned language model on short text: 92.6% against 93.7%. It does this on an ordinary CPU in about a minute, where the fine-tuned model took more than two hours on a GPU.
  • The fine-tuned model required a smaller review queue in this benchmark. Projected accuracy was about 98% after reviewing roughly 18% of rows with words + meaning, versus 9% with the fine-tuned model, assuming all flagged errors were corrected without new errors. Human-review outcomes and review time were not measured.

Two benchmarks that disagree

We picked two public datasets that stress different things.

BANKING77 is 13,083 real messages to a retail bank's customer service, each labelled with one of 77 intents. The messages are one or two sentences, and many intents are near neighbours: card_arrival against card_delivery_estimate, and top_up_failed against top_up_reverted. People say the same thing in many different words. We trained on the dataset's 10,003-message training file and scored on its separate 3,080-message test file.

LEDGAR is contract provisions from US securities filings, labelled with 100 clause types. A typical provision is 87 words long, one in ten runs past 230, and the categories are largely named by their language: "governed by the laws of", "may be executed in counterparts". We trained on 10,000 provisions and scored on 4,000 others.

On both, we also trained on only the first 5, 10, 20 and 50 examples of each category. That is the situation of a team trying automation for the first time with the examples it already has.

What we measured

Six approaches to text classification
ApproachHow it decides
Word matching, nearest neighboursFinds the five most similar training rows by shared words and lets them vote.
Word matching, logistic regressionLearns a weight for every word and two-word phrase in every category.
Meaning (multilingual-e5-small)A pre-trained model of about 100 million parameters turns each row into 384 numbers that place similar meanings close together; a logistic regression learns where each category sits.
Meaning (Qwen3-Embedding-0.6B)The same approach with a model six times larger.
Words + meaningBoth feature sets side by side in one logistic regression.
Fine-tuned language modelQwen3.5-4B, adapted to your categories with LoRA on a GPU (our Dedicated AI Model). Measured on BANKING77.

The results

These are benchmark results for individual approaches, not a measured accuracy guarantee for the product’s automatic selection. The fine-tuned BANKING77 model used 8,500 training rows and held 1,503 back for validation; the other full-size approaches used 10,003 training rows.

BANKING77: short customer messages, 77 intents

BANKING77: accuracy across 77 intents
training examplesWords (neighbours)Words (regression)MeaningWords + meaningFine-tuned model
5 per intent51.7%54.4%72.1%65.0%—
10 per intent62.0%68.2%81.3%77.6%—
20 per intent67.5%76.1%85.7%83.8%—
50 per intent75.7%85.6%89.1%90.1%—
100 per intent79.6%88.7%90.1%91.9%—
all (10,003)81.1%89.4%91.0%92.6%93.7%

LEDGAR: contract clauses, 100 types

LEDGAR: accuracy across 100 clause types
training examplesWords (neighbours)Words (regression)MeaningWords + meaning
5 per type54.3%61.1%57.8%62.7%
10 per type60.3%69.0%64.8%70.3%
20 per type64.9%74.5%70.7%75.2%
50 per type69.3%77.9%74.3%78.1%
all (10,000)74.5%82.3%78.9%82.7%

Uniform random guessing scores about 1.3% on BANKING77 and 1% on LEDGAR.

The four decisions, and the evidence for each

1. Read meaning, not only words

On short text, meaning is worth 13–20 points over word matching when there are only 5 or 10 examples per category. At 10 examples per intent, the meaning-based classifier reached 81.3% and the best word method 68.2%. To a word counter, "my card never came" and "the card hasn't arrived" have almost nothing in common. A model that reads meaning puts them side by side.

2. But keep the words, and let the data choose

LEDGAR reverses the result: words beat meaning at every size. Contract language is formulaic, the exact phrase is the category, and a general-purpose model with no legal training blurs distinctions that a lawyer's vocabulary makes sharp.

Putting both in one classifier was the best choice on LEDGAR at every size, and on BANKING77 from 50 examples per intent up. With very few short examples, meaning alone was still ahead, because the word half has too little to learn from.

No single option wins everywhere, so the AI Classifier fits the two that can win, meaning and words + meaning, and keeps whichever scores better on validation rows held back from your own file. Both classifiers reuse the same text embeddings. Each fit takes seconds.

On these benchmarks, picking the better of the two at each size would give 72.1 → 92.6% on BANKING77 and 62.7 → 82.7% on LEDGAR. In practice the choice is made on your validation rows, not on a test set, so it will occasionally pick the second-best.

3. A small, fast embedding model, not a larger one

Embedding models: accuracy and measured CPU throughput
ApproachBANKING77, all dataBANKING77, 10 per intentrows per second on 2 CPU threads
multilingual-e5-small91.0%81.3%427
Qwen3-Embedding-0.6B88.6%75.4%2

The larger model was less accurate at every size in our setup and about 200 times slower on a CPU. We did not tune its prompt or pooling.

multilingual-e5-small is MIT-licensed, reads about 100 languages, and runs on ordinary CPUs. The AI Classifier can run on CPUs; these measurements do not guarantee production latency or availability.

4. Logistic regression, not nearest neighbours

The same word counts scored 81.1% with a nearest-neighbour vote and 89.4% with logistic regression on BANKING77. On LEDGAR the gap was 74.5% against 82.3%. Changing the classifier improved accuracy by about eight percentage points before adding meaning-based features.

How much checking it takes

Every model reports how confident it is in each answer. Send the unsure rows to a person, and what you deliver is better than the model's raw accuracy. The real question is how many rows that takes.

Review volume and projected accuracy after correction
Approachaccuracyrows a person reviewsprojected after review
BANKING77, words + meaning92.6%18.4%98.3%
BANKING77, fine-tuned model93.7%9.2%98.0%
LEDGAR, words + meaning82.7%44.4%97.8%

Projected accuracy assumes every flagged error is corrected without introducing new errors. Human-review outcomes were not measured.

Two things follow:

  • On BANKING77, the two approaches reached similar projected post-review accuracy, with about twice the review for words + meaning. On 10,000 messages a month, that is roughly 1,800 rows to check against 900. For many teams that trade is worth a model that trains in a minute on a CPU.
  • Long, specialised text is harder for every inexpensive approach. About 44% of contract clauses needed checking to reach 98%. A Dedicated AI Model is another approach to evaluate for such work; a fine-tuned LEDGAR result is not included in this comparison.

What we built around the choice

  • Pinned, never swapped silently. Each AI Classifier Model records exactly which embedding model and version it was trained with. If that model ever changes, your model says it needs retraining instead of quietly giving different answers.
  • Checked on your data. Every model is scored on examples held back from your own file before you use it. The accuracy you see is on rows it never trained on.
  • The pre-trained model never learns from your data. It is used as-is; only the small classifier on top is trained, and that is yours.

What these results do and do not show

  • One run of each approach, on two benchmarks. Small gaps are not rankings.
  • Review cutoffs were read from the test results. In practice they should be chosen on separate validation rows.
  • English only. The embedding model is multilingual; these benchmarks don't test that.
  • Long text is cut. The embedding model reads the first 512 word-pieces of a row. The word half reads everything.
  • Two measurement paths. The fine-tuned model was measured end to end through our product. The other approaches were measured with a benchmark script on the same rows and the same scoring.
  • Laptop speeds. Throughput was measured on a laptop's CPU, and cloud servers are typically somewhat slower.

Method details

  • BANKING77: the dataset's own train (10,003) and test (3,080) files. Smaller sets are the first N per intent after a fixed shuffle (seed 17), identical for every approach.
  • LEDGAR: LexGLUE's ledgar configuration; 10,000 training and 4,000 test provisions after a fixed shuffle (seed 17).
  • Words: TF-IDF over words and two-word phrases, sublinear term frequency; nearest neighbours with k = 5 and similarity-weighted voting; logistic regression with C = 10, fixed in advance.
  • Meaning: intfloat/multilingual-e5-small, inputs prefixed query: , normalised 384-dimension vectors, 512-token limit; logistic regression, C = 10.
  • Words + meaning: both feature sets, each L2-normalised per row, concatenated; logistic regression, C = 10. The product caps the word vocabulary at 20,000 terms, which cost at most 0.8 points in our tests.
  • Fine-tuned model: Qwen3.5-4B, LoRA r = 16, 3 epochs, one NVIDIA L4 GPU, 2 h 16 min of training.
  • Confidence: the highest class probability (for nearest neighbours, the winning category's share of the weighted vote).

Sources and attribution

StayCharted’s measurements and analysis are independent of the datasets’ and models’ authors.

Explore more Insights

Compare AMT’s design decisions and training reports →

STAYCHARTED AI MODEL TRAINER

Put your business knowledge to work.

Build your Business-Specific AI with AMT. Start with examples you already have, test your model’s results, and put it into your workflow.

Try AI Model Trainer