← All insights

Developing AMT

Why Our Image Model Is Built on SigLIP 2

How StayCharted compared SigLIP 2 and DINOv2 for product photos and pet breeds, and chose an image model for learning business categories.

By StayCharted · Measurements: September 25, 2026 · Published

Many businesses sort pictures by hand: product photos into catalogue categories, damage photos into claim types, site photos into inspection outcomes, scanned items into bins. Our Image Model automates that from your own examples. You give it pictures you have already sorted, as a ZIP with one folder per category or a spreadsheet of picture links. A pre-trained vision model reads what is in each picture, and a small classifier learns your categories on top.

The most important choice in that design is which pre-trained vision model reads the pictures. We measured three leading open models on two public datasets, chosen because they reward different strengths, and let the numbers decide. This note explains what we found and why we chose SigLIP 2.

The short version

  • On catalogue product photos, SigLIP 2 won clearly: 91.1% against 88.3–88.8% with all the data. With 5 examples per category it led by 6–10 points (83.5% against 73.2–77.7%).
  • On fine-grained natural photos (37 cat and dog breeds), DINOv2 base was ahead: 96.0% against 95.2% with all the data, and by more with very few examples.
  • We chose SigLIP 2. We prioritized catalogue-style categorization for this design. SigLIP 2 led on that dataset at every tested training size, including 5 to 50 examples per category. Other business imagery may favor a different model.
  • It runs on ordinary CPUs: about 20 pictures a second on two laptop CPU threads, including decoding and resizing. No GPU is needed to train or to fill a file.
  • The faster model was less accurate on product photos. DINOv2 small is about 2.4 times faster, but gave up 3–10 points on product photos.

The two datasets

Catalogue product photos. A public set of fashion catalogue images, each labelled with its product type: shirts, watches, handbags, sports shoes, sunglasses and so on. We kept the 40 product types with at least 60 pictures each: 7,966 pictures to train on and 1,989 to test on. The photos are clean, well lit and on plain backgrounds, and the categories are kinds of thing. This is the closest public stand-in we found for what businesses sort.

Oxford-IIIT Pet. Photos of 37 cat and dog breeds, taken in ordinary settings: 3,680 to train on and 1,850 to test on. This is the opposite kind of task. Every picture is a cat or a dog, and the categories differ in fine visual detail such as coat, ears and face shape.

On both, we also trained on only the first 5, 10, 20 and 50 examples of each category. Most teams start with the examples they already have, and that is rarely thousands.

The three models

Vision encoders and feature vectors used in this run
ApproachFromWhat it learned fromVector size
SigLIP 2 base (google/siglip2-base-patch16-224)Google, 2025Pictures paired with text describing them768
DINOv2 small (facebook/dinov2-small)Meta, 2023Pictures alone (self-supervised)768
DINOv2 base (facebook/dinov2-base)Meta, 2023Pictures alone (self-supervised)1,536

All three are open and run without a GPU. Each turns a picture into a list of numbers (a vector), and pictures with similar content get similar vectors. We trained the same classifier on each model's vectors: a multinomial logistic regression with its settings fixed in advance. So the comparison is between the models alone.

The results

Catalogue product photos: 40 product types

Catalogue product photos: accuracy across 40 categories
training picturesSigLIP 2 baseDINOv2 smallDINOv2 base
5 per category (200)83.5%73.2%77.7%
10 per category (400)85.9%79.0%81.9%
20 per category (800)87.5%81.0%83.9%
50 per category (2,000)89.8%84.7%86.4%
all (7,966)91.1%88.3%88.8%

Oxford-IIIT Pet: 37 breeds

Oxford-IIIT Pet subset: accuracy across 37 breeds
training picturesSigLIP 2 baseDINOv2 smallDINOv2 base
5 per category (185)85.5%87.3%90.2%
10 per category (370)90.6%91.1%93.9%
20 per category (740)92.9%93.3%94.6%
50 per category (1,850)94.2%94.2%95.7%
all (3,680)95.2%95.2%96.0%

Random guessing scores 2.5% on the first and 2.7% on the second.

Why the two datasets disagree. This is our reading; we did not test it directly. SigLIP 2 learned from pictures paired with descriptions, so it has absorbed what things are called: a watch, a clutch, a flip-flop. That is exactly what separates catalogue categories. DINOv2 learned from pictures alone and is known for its fine visual detail, which is what separates a Birman from a Ragdoll. That interpretation informed our choice for catalogue-style tasks; it is not a claim about every business image workflow.

Why SigLIP 2

  1. It led on the product-photo benchmark. At every training size on product photos, SigLIP 2 led the next-best model by 2.3 to 5.8 points, and DINOv2 small by up to 10.3. On the breed task it trailed DINOv2 base by 0.8 to 4.7 points.
  2. It wins most when examples are few. With 5 per category, SigLIP 2 reached 83.5% on products, where the others reached 73–78%.
  3. Measured processing speed. It ran at about 20 pictures a second on two laptop CPU threads, similar to DINOv2 base in our setup, with half the output-vector size (768 against 1,536). DINOv2 small is faster, at about 48 a second, but its accuracy on products was lower at every size.
  4. The model is pinned and self-hosted. We use one exact version of the pre-trained model and run it within our service, rather than sending pictures to a model provider for inference.

How much checking it takes

Each Image Model reports how confident it is in every answer, so low-confidence pictures can go to a person first. With all the training data:

SigLIP 2: review volume and projected accuracy
Approachcutoffpictures a person reviewserrors caughtprojected after review
Products, SigLIP 20.7018.0%75.7%97.8%
Products, SigLIP 20.8024.7%87.0%98.8%
Breeds, SigLIP 20.507.4%48.9%97.6%
Breeds, SigLIP 20.6013.0%70.5%98.6%

Projected accuracy assumes every flagged error is corrected without introducing new errors. Human-review outcomes were not measured.

The review needed to reach about 98% was broadly similar across the three models, and it followed accuracy. SigLIP 2 needed somewhat less than the other two on products; DINOv2 base needed less on breeds, where it reached 98.3% while flagging 7.2%. So review didn't change the choice; it came from the accuracy above.

What we built around the choice

  • It learns from the picture, never from the file name. Letting a folder name such as damaged leak into the features would teach a model the wrong lesson. It would score well on pictures from the same folder and fail on the next delivery. The Image Model is trained on the picture's vector alone.
  • Every picture is cleaned before it is stored. Location data, camera serial numbers and other metadata are removed, rotated phone photos are turned upright, and each picture is resized and kept with a thumbnail you can see while checking your data.
  • Checking your pictures before training. Before you train, the Image Model finds and shows you:
  • pictures that look like they belong in another category;
  • the same picture filed under two answers;
  • near-identical copies;
  • links that could not be read.

Each finding is one click to fix.

  • Pinned, never swapped silently. Each Image Model records exactly which vision model and version produced its vectors. If that model ever changes, your model says it needs retraining instead of quietly giving different answers.
  • The pre-trained model never learns from your pictures. It is used as-is; only the small classifier on top is trained, and that is yours.

What these results do and do not show

  • Two datasets, one run each. Small gaps are not rankings, and neither dataset is your data. Check the accuracy on your own held-back pictures, which the Image Model shows you after training.
  • 224 pixels. All three models read a picture at 224 × 224 pixels. Very fine detail, such as a hairline scratch or small printed text, may not survive that, and pictures that are mostly text (documents, receipts) are a different problem.
  • Review cutoffs were read from the test results. In practice they should be chosen on separate validation pictures.
  • Laptop speeds. Throughput was measured on a laptop's CPU. These are not measurements or throughput guarantees for a production cloud server.
  • No fine-tuned vision model yet. We compared pre-trained models with a classifier on top. Fine-tuning a vision model on specialized imagery is future work; any improvement remains to be measured.

Method details

  • Split: the run used separate training and test files with the counts above. The construction of the product split and the 1,850-image Pet test subset is not fully documented, so these results should not be treated as an official full-test-set score. Smaller training sets are the first N pictures per category after a fixed shuffle (seed 17), identical for all three models. For the product photos, the 40 product types with at least 60 pictures were kept.
  • Pre-processing: each model's own published image processor, with the reported software versions (transformers 4.57.1, Pillow 12.2.0).
  • Vectors: SigLIP 2, the image embedding; DINOv2, the class token and the mean of the patch tokens side by side (the linear-evaluation recipe from the DINOv2 paper), which is why DINOv2 base's vector is 1,536 long.
  • Classifier: multinomial logistic regression on the vectors, regularisation fixed in advance, not tuned on the test pictures.
  • Confidence: the highest class probability.
  • Hardware: a laptop, with its GPU for the bulk embedding and two CPU threads for the throughput figures.

Sources and attribution

StayCharted’s measurements and analysis are independent of the datasets’ and models’ authors.

Explore more Insights

Compare AMT’s design decisions and training reports →

STAYCHARTED AI MODEL TRAINER

Put your business knowledge to work.

Build your Business-Specific AI with AMT. Start with examples you already have, test your model’s results, and put it into your workflow.

Try AI Model Trainer