← All articles

AI Model Training Report

Training AI for Banking Support: Accuracy, Confidence, and Human Review

How can banking teams turn customer messages into the right categories—and decide which predictions need a person? A training run with StayCharted’s AI Model Trainer (AMT) explores both accuracy and the work of checking results.

By StayCharted · Published · Updated

93.7%Measured model accuracy
3,080Held-out test examples
77Banking support intents

Can AI help categorize banking support requests?

Yes. StayCharted’s AI Model Trainer (AMT) can train a private model on labeled examples to predict request categories. In this single BANKING77 run, a fine-tuned model achieved 93.7% accuracy across 77 banking-support intents. Confidence scores helped prioritize human review; performance on a bank’s own requests must be evaluated separately.

Measured versus projected: 93.7% is the model’s measured accuracy. The 98.0% figure is a projection assuming every flagged error is corrected, with no new errors, after reviewing 9.2% of messages. It is not a measured human-review outcome.

A customer writes a message. Someone reads it, chooses a category, and decides where it should go. For banking support teams, automating part of that work requires a model that understands the distinctions in their categories—and a way to decide what still needs checking.

We tested StayCharted AMT on BANKING77, a public banking-support dataset. A fine-tuned model correctly labeled 2,885 of 3,080 held-out messages: 93.7% accuracy. Its confidence scores also helped identify mistakes. Reviewing 9.2% of messages would surface 67.7% of errors in this run; correcting all those errors would bring accuracy to a projected 98.0%.

The banking workflow behind the benchmark

Many financial-services workflows have the same starting point: a column of short text and a category assigned by a person. The categories may overlap, so a small wording difference can change how a request should be handled.

Potential applications of text categorization. Only support-intent classification was measured in this run.
WorkflowIncoming textCategory or outcome
Support routingChat and email requestsIntent and destination queue
Complaint categorizationComplaint narrativesProduct and issue category
Transaction categorizationMerchant and payment descriptionsSpending or accounting code
Dispute triageCustomer dispute descriptionsSuggested reason category for review
Back-office triageOperational notesWork queue

BANKING77 tests the first kind of decision: identifying the intent of a banking-support message. The other examples have a similar structure, but would need their own training data, evaluation, and review process. This run did not measure complaint, transaction, or dispute classification.

The measured result: 93.7% across 77 banking intents

BANKING77 contains 13,083 customer-service queries across 77 intents. We used its published split: 10,003 training examples and 3,080 test examples. The distinctions include card_arrival versus card_delivery_estimate, and top_up_reverted versus top_up_failed.

Using StayCharted AMT, we fine-tuned Qwen3.5-4B with a LoRA adapter. On the separate test file, the model correctly categorized 2,885 of 3,080 messages, with 195 errors. All 3,080 predictions were valid category names: no blank cells, invented labels, or explanatory prose.

Valid output matters when a prediction fills a spreadsheet column or feeds another workflow. It is separate from correctness: the 195 mistakes were valid categories assigned to the wrong messages.

For context, Table 3 of the dataset’s original paper reports 93.66% for fine-tuned BERT and 93.36% for USE+ConveRT with the full training set. Our 93.7% result is numerically similar. These are different experimental setups, not a controlled head-to-head comparison, and one run does not establish superiority.

Accuracy is only part of the operational question

A support team also needs to know which predictions deserve attention. StayCharted AMT flags low-confidence predictions for review. At the 0.40 cutoff used in this run, only 0.4% of messages were flagged.

Of those flagged messages, 84.6% were wrong. The queue was concentrated with errors, but it contained only 11 of the model’s 195 mistakes. That left 184 errors unflagged.

Raising the cutoff to 0.90 in a post-hoc analysis sent 9.2% of messages to review and caught 67.7% of errors. That is approximately one message in eleven checked, with two-thirds of mistakes surfaced. It demonstrates a useful tradeoff in this dataset, not a universal cutoff or a change to the product default.

Confidence is not a calibrated probability of correctness. A score of 0.90 does not automatically mean a prediction has a 90% chance of being right, and some confident predictions will still be wrong.

How much review buys how much improvement?

Share of the test set sent for review
Cutoff 0.400.4% reviewed
Cutoff 0.806.3% reviewed
Cutoff 0.909.2% reviewed
Cutoff 0.9921.1% reviewed
Bars use a 0–100% scale. A higher cutoff sends more rows to review.

Measured model accuracy: 93.7%. Projected accuracy after review: 98.0% at cutoff 0.90. The projection assumes a person corrects every flagged error and introduces no new errors. We did not measure a human-review trial.

Same test set; rounded percentages. Post-review accuracy is a projection.
CutoffRows reviewedReviewed rows that were wrongErrors caughtProjected accuracy after review
0.400.4%84.6%5.6%94.0%
0.703.9%56.3%34.4%95.8%
0.806.3%50.8%50.8%96.9%
0.909.2%46.6%67.7%98.0%
0.9511.9%40.4%75.9%98.5%
0.9921.1%27.1%90.3%99.4%

The calculation is: correct predictions plus flagged errors successfully corrected, divided by all test rows. At 0.90, correcting about 132 of 195 errors would leave about 63 errors, yielding approximately 98.0% accuracy.

This changes the workflow, not the model’s underlying accuracy. Real outcomes depend on reviewer accuracy, available context, and the cost of each mistake. Review volume also does not directly measure time saved: difficult rows can take longer to check.

How we tested

We used the dataset’s original train/test split across all 77 intents. Of the 10,003 training-split rows, 8,500 were used for fine-tuning and 1,503 for internal validation. The separate 3,080-row test split was used for the evaluation reported here.

Configuration reported for this run
Base modelQwen3.5-4B; pinned revision 851bf6e
Training methodSupervised fine-tuning with LoRA; 0.71% of parameters trained
LoRA settingsRank 16; alpha 32; dropout 0.05
Training3 epochs; learning rate 0.0002
Batch size4 per device × 4 accumulation steps = 16
Hardware and precisionOne NVIDIA L4 (24 GB); bfloat16
Adapted layers32 of 32 decoder layers
Training time8,160 seconds (about 2 hours 16 minutes)

Each example paired a request with its category name. The model generated the category as text rather than using a separate classification head. An automated harness exercised StayCharted AMT’s customer workflow: upload a spreadsheet, map the columns, train and publish a model, then upload a second spreadsheet with its label column blank and download the predictions. The run took place on September 23, 2026.

Measuring through that workflow includes the product’s prompting and prediction path. It does not constitute a usability study or a measurement of human review speed.

How to interpret this result

  • One training run. Run-to-run variation is unmeasured. We make no ranking or superiority claim.
  • Post-hoc review analysis. The cutoff sweep was computed on the reported test results. Choose a deployment cutoff on validation data, then check it on a fresh evaluation set.
  • Short banking-support text. The result does not establish performance on long documents, other languages, or a bank’s own requests. Nor does it demonstrate fraud detection, credit decisions, or regulatory compliance.
  • Separate fine-tuning test data. The published test split was excluded from this fine-tune. We cannot establish whether the base model encountered this public dataset during pretraining.
  • Reproduction details. The base-model revision was pinned to commit 851bf6e; training settings and dataset splits are listed above. Per-row predictions are not published with this report.

What this means for banks and fintechs

Choose the checking based on what a mistake costs

A misrouted support message and a wrongly categorized dispute can have different consequences. Use validation results to decide which predictions can move forward and which need a person. For higher-impact work, consider mandatory review as well as confidence thresholds. The benchmark helps frame that decision; it does not make it for you.

Review the category definitions as well as the model

Confusions included top_up_reverted and top_up_failed, compromised_card and lost_or_stolen_card, and the two card-delivery categories. No single directional confusion occurred more than four times. These pairs are worth examining with the people who define your categories. The counts alone do not prove ambiguity.

Start with decisions your team has already made

Resolved tickets and their assigned categories can provide examples for training. StayCharted AMT turns those examples into a private model you can test against held-out requests. For a support team, the potential benefit is more consistent categorization and a focused review queue; actual time savings and resolution improvements need to be measured in the team’s workflow.

Keep people and traceability in the workflow

StayCharted AMT is designed with human review and an audit trail recording who changed a result and when. Models are trained for your use rather than pooled across customers, and your data can be exported and deleted. These product capabilities support oversight; they do not by themselves establish regulatory compliance or satisfy an institution’s model-risk requirements.

Apply the lesson to your own workflow

  1. Start with representative examples and the labels your team actually assigned.
  2. Hold back separate data for validation and final evaluation.
  3. Measure both accuracy and the errors your review queue catches.
  4. Choose the cutoff based on the consequences of mistakes and your team’s review capacity.

For support ticket routing, start with suggested categories and review before assignment. For AI inside your application, design the human-review or fallback path alongside the prediction.

The practical question is: how much work can this model take on, and what checking is needed to trust the result?

Sources and attribution

BANKING77 was introduced by Iñigo Casanueva, Tadas Temčinas, Daniela Gerz, Matthew Henderson, and Ivan Vulić in Efficient Intent Detection with Dual Sentence Encoders (2020). The PolyAI dataset repository provides the original splits under CC BY 4.0. Our training split was further divided for validation; the test split was retained as supplied.

Base model: Qwen3.5-4B. StayCharted’s run measurements and threshold analysis are reported above; the dataset authors and model provider do not endorse these results.

STAYCHARTED AI MODEL TRAINER

Put your business knowledge to work.

Build your Business-Specific AI with AMT. Start with examples you already have, test your model’s results, and put it into your workflow.

Try AI Model Trainer