← All insights

Developing AMT

Grading food by photo: fruit quality and cashew defects, measured

Food grading by photo with StayCharted AMT: compare fruit quality accuracy, cashew defect detection, and projected human review on FruitNet and VisA.

By Manoj Mohandas · Measurements: October 4, 2026 · Published

The short version

Can a model trained on your own sorted photos tell good produce from bad, and a sound cashew from a damaged one? We trained both of StayCharted AMT's picture models on two public datasets and scored them on photos they had never seen.

  • Fruit, good or bad (six fruits):
    • Both models were right on 98.7% of 600 test photos.
    • The AI Image Classifier Model trained in under a minute.
    • The Dedicated AI Image Model flags the photos it is unsure about. Correcting every error among the 2.2% it flagged projects 99.5% accuracy.
  • Cashews, sound or defective:
    • The Dedicated AI Image Model was right on 93.8% of 240 test photos.
    • It found only 26 of the 40 defective cashews (65%).
    • When it did say "defect", it was right 96% of the time.
    • The AI Image Classifier Model was right on 90.0%, but found only 16 of the 40 defects (40%). Every defect it did call was a real one.
    • Small, rare defects are the hard part, and this is where training the AI itself on your photos pays: the Dedicated model found 10 more defects.
    • Overall accuracy hides that, so the defect catch rate is the number to look at.

1. The two tests

FruitNet VisA cashews
What it is Photos of six fruits (apple, banana, guava, lime, orange, pomegranate), each labelled good or bad quality Close-up photos of single cashews, labelled OK or defective
Source Meshram & Patil, 2021 (Mendeley Data) Amazon's VisA dataset, Zou et al., ECCV 2022
Licence CC BY 4.0 CC BY 4.0
Training photos 1,800 (900 good, 900 bad) 360 (300 OK, 60 defective)
Test photos 600 (100 per fruit, half good and half bad) 240 (200 OK, 40 defective)
How the test set was chosen 50 per fruit and quality, by a fixed shuffle; the rest trained Amazon's own published split, unchanged

Preparation:

  • Every photo was resized so its longest side is at most 512 pixels, which is what AMT keeps and what both models read.
  • Training photos went in as a ZIP with one folder per category, the way a customer uploads them.
  • Test photos were renamed to neutral file names, so neither model could read the answer from a folder or file name.

2. The two models

  • AI Image Classifier Model (Vision add-on, Essentials and Business):
    • a ready-made vision AI turns each photo into a numerical representation of its visual features;
    • a small classifier learns your categories from those representations;
    • it trains in seconds to a minute and needs no GPU.
  • Dedicated AI Image Model (Business with Vision):
    • the AI itself is trained on your photos, on a GPU;
    • it takes longer to train, and is built for small differences a ready-made model may not separate well.

3. Results

Measured predictions, projected review outcomes. Accuracy and defect detection were measured on held-out photos. All after-review figures below are projections assuming every flagged error is corrected without introducing new errors. They are not measured human-review results.

Accuracy Balanced accuracy* Flagged for review at the product's cutoff Projected accuracy after review Training time
Fruit: AI Image Classifier Model 98.7% 98.7% 0.0% 98.7% 0.3 min
Fruit: Dedicated AI Image Model 98.7% 98.7% 2.2% 99.5% 41.5 min
Cashews: AI Image Classifier Model 90.0% 70.0% 0.0% 90.0% 0.3 min
Cashews: Dedicated AI Image Model 93.8% 82.3% 3.8% 95.8% 18.3 min

* Balanced accuracy is the average of the accuracy on each category. On the cashews, a model that answered "OK" for every photo would be 83% accurate (200 of 240) and catch no defects at all. Balanced accuracy would show that model at 50%. That is why it, and the defect catch rate below, matter more than overall accuracy for inspection work.

Fruit: per fruit

Fruit AI Image Classifier Model Dedicated AI Image Model
Apple 98% 98%
Banana 99% 100%
Guava 100% 100%
Lime 100% 99%
Orange 95% 95%
Pomegranate 100% 100%

Oranges were the hardest fruit for both models (95%).

Cashews: where the mistakes are

Photos AI Image Classifier Model Dedicated AI Image Model
Defective cashews found 40 16 (40%) 26 (65%)
Sound cashews called sound 200 200 (100%) 199 (99.5%)
When it says "defect", it is right 100% 96%

Almost every mistake, for both models, is a defect that was missed, not a false alarm. That matters, because a missed defect goes into the bag.

4. How much checking it takes

Each answer comes with a confidence. AMT can send the less confident answers to a person for checking, and the customer picks the cutoff. These are the Dedicated AI Image Model's results at different cutoffs:

Cutoff Fruit: photos checked Fruit: projected accuracy after review Cashews: photos checked Cashews: share of mistakes caught Cashews: projected accuracy after review
0.80 1.3% 99.2% 2.5% 20% 95.0%
0.90 (product default) 2.2% 99.5% 3.8% 33% 95.8%
0.95 3.2% 99.7% 4.2% 33% 95.8%
0.99 6.7% 99.8% 24.2% 80% 98.8%

Fruit: reviewing about 1 photo in 15 could remove nearly every remaining mistake, if each flagged error is corrected without introducing new errors.

Cashews:

  • At the default cutoff the model is confident even about most of the defects it misses, so checking catches only a third of them.
  • Raising the cutoff to 0.99 sends about 1 photo in 4 to a person and catches 80% of the mistakes, giving 98.8% projected accuracy after review.
  • Whether that trade is acceptable depends on the cost of missed defects; a person would review a quarter of the photos, with some errors remaining.

The AI Image Classifier Model on cashews: it needs a cutoff well above its default of 0.5, because at 0.5 it flags nothing.

Cutoff Photos checked Share of mistakes caught Projected accuracy after review
0.7 10.8% 50% 95.0%
0.8 18.8% 71% 97.1%
0.9 51.7% 96% 99.6%

So with review it can reach a similar result to the Dedicated model (97.1% at 18.8% checked, against 98.8% at 24.2%), but it starts from fewer defects found, and anything that slips past the review goes into the bag.

5. What this means for food and produce teams

  • Clear quality grades work on the inexpensive model. Good against bad fruit was learned from 150 photos per fruit and grade, in under a minute, at 98.7%. Start there.
  • Rare, small defects are harder, and need a review step.
    • With only 60 defective examples to learn from, the Dedicated model missed a third of the defects at its default cutoff, and the AI Image Classifier Model more than half.
    • Here the Dedicated AI Image Model is worth its longer training: it found 65% of defects against 40%.
    • Use a high review cutoff, and add more examples of each defect type. The defects it misses are the ones to add.
  • Judge an inspection model on the defects it catches. Overall accuracy looked good on cashews (93.8%) while a third of defects got through.
    • AMT's validation report shows accuracy by category for this reason. Read the defect row first.

6. What these results do and do not show

  • One run per model, on two public datasets. Your photos, lighting and defect types will give different numbers. Validate on your own held-out photos before relying on a model.
  • FruitNet's photo sessions.
    • For apples, guavas and pomegranates, the good and bad photos were taken in different years, so a model could partly learn the session instead of the fruit.
    • The per-fruit table does not rule out this session effect. Treat these results as potentially optimistic; test on photos from separate sessions.
  • Small training sets: 60 defective cashews is few. More defect examples usually help most.
  • Photo size: photos were at most 512 pixels on the longest side. A defect smaller than a few pixels at that size may not be visible to either model.
  • Speed is not measured here. The fill times in our logs include waiting for servers to start on a development environment, so they say nothing about speed in production.

Method

  • Run through StayCharted AMT's development environment on 4 October 2026, on the Business plan with Vision, exactly as a customer would:
    1. upload a ZIP of sorted photos;
    2. review data quality;
    3. train;
    4. publish;
    5. fill a ZIP of test photos;
    6. score the filled answers against the held-out labels.
  • AI Image Classifier Model: SigLIP 2 image embeddings plus a linear classifier.
  • Dedicated AI Image Model: Qwen3.5-4B fine-tuned on the training photos, on a SageMaker GPU.
  • Flagged for review means answers below the review cutoff (0.5 for the AI Image Classifier Model, 0.9 for the Dedicated AI Image Model, the product defaults). Projected accuracy after review assumes every flagged mistake is corrected and no new errors are introduced. Human-review outcomes were not measured.

Sources and attribution

Put the findings into practice

Visual inspection and review workflows →

Retail and e-commerce workflows →

Compare plans →

Explore more Insights

Compare AMT’s design decisions and training reports →

STAYCHARTED AI MODEL TRAINER

Put your business knowledge to work.

Build your Business-Specific AI with AMT. Start with examples you already have, test your model’s results, and put it into your workflow.

Try it on your own data