← All insights

Developing AMT

Sorting medical research abstracts by condition: four approaches, measured

Medical abstract classification for life sciences literature triage: word models, embeddings and a fine-tuned model compared on published abstracts.

By Manoj Mohandas · Measurements: 5 October 2026 (CPU approaches) and 5–6 October 2026 (fine-tuned model) · Published

The short version

Research, medical-affairs and pharmacovigilance teams sort scientific literature all the time: screening new publications, building evidence libraries, routing papers to the right therapeutic-area team. We measured four ways to automate the first pass on the Medical Abstracts corpus, 14,438 published medical research abstracts, each labeled with one of five condition groups.

Word matching Meaning (embeddings) Words + meaning Fine-tuned language model
Accuracy, 10,000 training abstracts 46.5% / 49.9%¹ 60.0% 52.0% 63.6%
Accuracy, 10 examples per condition group 35.5% / 36.8%¹ 49.8% 49.4% not measured
Share reviewed for a projected ~98% 98% not reached² / 92.9%¹ 91.7% 98% not reached² 90.1% (97.5% at 78.7%)
Time to train, 10,000 abstracts seconds to two minutes about 4 minutes about 5½ minutes about 8¼ hours on a GPU
Hardware any CPU a small CPU server a small CPU server a GPU

¹ Two word-matching methods: nearest neighbors / logistic regression; see below. ² The highest cutoff measured (0.95) projects 97.6% for both.

The fine-tuned model led, narrowly. 63.6% against 60.0% for embeddings: 3.6 points, or about one fewer mistake in every 28 abstracts. From a single run, that is a lead, not a firm ranking.

Meaning carried more of the signal than words. Embeddings beat both word-matching methods at every training size, by 3 to 14 points with 5 to 100 abstracts per group and by 10 points at full size. Medical vocabulary is specific, but abstracts describe the same condition in very different prose. Adding words to meaning did not help here: the combination matched embeddings up to 500 abstracts, then fell to 52.0% at 10,000.

This is a hard task, and overlapping labels may help explain why. Every approach ended between 46% and 64%. The same words-plus-meaning method scored 92.6% on banking messages and 82.7% on contract clauses. Of the fine-tuned model's 1,051 mistakes, 821 (78%) involve "General pathological conditions", the broadest group and a third of the test set.

Review is expensive when categories overlap. The fine-tuned model needs 78.7% of abstracts checked for a projected 97.5%, against 9.2% for banking messages and 41.7% for contract clauses. When the labels themselves overlap, the model is often confidently wrong, so flagging by confidence catches mistakes slowly.

1. The benchmark

The Medical Abstracts corpus was published by Tim Schopf, Daniel Braun and Florian Matthes (Technical University of Munich) with their paper Evaluating Unsupervised Text Classification: Zero-Shot and Similarity-Based Approaches (NLPIR 2022). It contains 14,438 abstracts of published medical research, each assigned to one of five condition groups, with a fixed split of 11,550 training and 2,888 test abstracts. It is released under CC BY-SA 3.0.

Condition group Training Test Share of test
General pathological conditions 3,844 961 33.3%
Neoplasms 2,530 633 21.9%
Cardiovascular diseases 2,441 610 21.1%
Nervous system diseases 1,540 385 13.3%
Digestive system diseases 1,195 299 10.4%
Total 11,550 2,888

Two baselines to read every number against: guessing at random over five groups scores 20%, and always answering the largest group ("General pathological conditions") scores 33.3%.

A typical abstract is 177 words long; one in ten runs past 275, and the longest is 596.

What we used. 10,000 abstracts from the training split, drawn in a fixed shuffled order, and all 2,888 test abstracts: the same rows for every approach.

  • The fine-tuned model trains on the 10,000; that keeps its run to a working day.
  • To see how the inexpensive approaches cope with little data, we also trained them on the first 5, 10, 20, 50 and 100 abstracts of each condition group.

What it is, and is not. This is published research text. It contains no patient records, and nothing here involves classifying a patient, a diagnosis or a treatment decision. The task is sorting literature.

2. The approaches

The same four approaches as our contract clause benchmark and banking messages benchmark, to explore how their performance changes across tasks. The datasets and category definitions differ, so these are not controlled industry comparisons.

1. Word matching (TF-IDF). Each abstract becomes a weighted count of the words and two-word phrases it contains. Two ways of deciding a label:

  • Nearest neighbors: find the five most similar training abstracts and let them vote.
  • Logistic regression: learn a weight for every word and phrase per condition group.

2. Meaning (embeddings). A small pre-trained multilingual model (multilingual-e5-small, about 100 million parameters) turns each abstract into 384 numbers that place similar meanings close together. A logistic regression learns where each condition group sits. It reads at most the first 512 word-pieces of an abstract.

3. Words + meaning. Both of the above, side by side, in one logistic regression that can use whichever carries the signal.

4. A fine-tuned language model. Qwen3.5-4B, a 4-billion-parameter open model, adapted with LoRA to write the condition group for each abstract. It was trained on one NVIDIA L4 GPU and measured end to end through StayCharted AI Model Trainer: we uploaded a spreadsheet of labeled abstracts, trained, then filled a second spreadsheet whose label column was blank.

3. Accuracy, by how much training data you have

training abstracts Word matching (nearest neighbors) Word matching (logistic regression) Meaning (embeddings) Words + meaning Fine-tuned model
5 per group (25) 30.4% 30.7% 41.4% 40.4% —
10 per group (50) 35.5% 36.8% 49.8% 49.4% —
20 per group (100) 42.4% 45.0% 51.7% 51.8% —
50 per group (250) 44.4% 51.7% 56.4% 56.4% —
100 per group (500) 47.3% 54.0% 57.3% 57.9% —
all (10,000)² 46.5% 49.9% 60.0% 52.0% 63.6%

Random guessing scores 20%; always answering the largest group, 33.3%.

² The fine-tuned model trained on about 85% of these and held the rest back for its own validation; the other approaches used all 10,000. The fine-tuned model was measured at full size only.

What stands out:

  • Meaning first. Embeddings beat both word-matching methods at every size. With 10 abstracts per group they reached 49.8%, ahead of both word-matching methods with 20 per group (42.4% and 45.0%).
  • Most of the gain comes early. Embeddings rose 16 points from 5 to 100 per group (41.4% → 57.3%), then only 2.7 more with 10,000 abstracts. For a first model on overlapping categories, a hundred good examples per group goes most of the way.
  • More data did not always help. Both logistic-regression models scored lower at 10,000 than at 500 (words 54.0% → 49.9%; words + meaning 57.9% → 52.0%). One likely reason: the 500-row sets are balanced across groups, while the full set is not, and a third of it is the overlapping general group. Settings were fixed in advance and not re-tuned for size.
  • The fine-tuned model adds 3.6 points over the best inexpensive approach: about one fewer mistake in every 28 abstracts.

Where the mistakes are. The fine-tuned model's five most frequent mistakes, out of 1,051:

the abstract was labeled the model said times
General pathological conditions Cardiovascular diseases 171
General pathological conditions Neoplasms 128
General pathological conditions Digestive system diseases 119
Nervous system diseases General pathological conditions 99
Cardiovascular diseases General pathological conditions 96

All five involve "General pathological conditions", and so do 821 of the 1,051 mistakes (78%). It is the broadest group, and many abstracts about a general condition also concern a specific organ system, so the labels overlap the way a real taxonomy's often do. Between the four organ-system groups the model rarely confused one for another: the most frequent such pair, neoplasms read as digestive diseases, happened 33 times.

4. Against the published work

The corpus's authors built it to evaluate classification without labeled training data: zero-shot and similarity-based methods that are given only the names of the condition groups. This benchmark measures the opposite case, training on labeled examples. The numbers answer a different question and are not comparable one-to-one with the paper's.

The useful contrast is the kind of information each uses: the paper's methods know only the names of the five groups, while every approach here learned from labeled examples. If your categories already have hundreds of labeled examples, use them.

5. How much checking?

Each approach reports its confidence. Flag low-confidence abstracts for a person, and the delivered result rises above the model's own accuracy. The question is how much of the file that takes.

model accuracy review this share… …to project
Word matching, nearest neighbors 46.5% 91.0% (cutoff 0.95) 97.6%
Word matching, logistic regression 49.9% 92.9% (cutoff 0.95) 99.1%
Meaning (embeddings) 60.0% 91.7% (cutoff 0.90) 98.4%
Words + meaning 52.0% 88.3% (cutoff 0.95) 97.6%
Fine-tuned model 63.6% 78.7% (cutoff 0.95) 97.5%

Projected accuracy assumes every flagged abstract is checked and corrected without introducing a new error. It is not a measurement of reviewers.

Fine-tuned model, cutoff 0.80 0.90 0.95 0.99
abstracts flagged 55.6% 69.9% 78.7% 90.1%
share of its mistakes among them 76.5% 88.3% 93.1% 98.3%
projected after checking them 91.4% 95.7% 97.5% 99.4%

Compare the same fine-tuned approach on banking messages, where a 0.90 cutoff flagged 9.2% of messages for a projected 98%, and on contract clauses, where 41.7% at 0.99 projected 97.6%. Here a projected 97.5% takes 78.7% of the file. When categories overlap, a model is often confidently wrong: at a 0.80 cutoff, only half of the flagged abstracts were mistakes, against a 36% error rate overall. Review still helps, but on a taxonomy like this one the bigger gain comes from sharpening the categories first (see section 7).

The cutoff that suits one kind of text can be far too low or too high for another, so it has to be set on your own data. StayCharted AI Model Trainer lets each customer choose it per model, from these same numbers measured on their own held-out rows, and sends every answer below it to a Review Queue, where a person confirms or corrects it and the next version learns from the correction.

6. What it costs to run

train on 10,000 abstracts label new abstracts hardware
Word matching seconds to two minutes hundreds to thousands per second any CPU
Meaning, or words + meaning about 3–4 minutes to embed 12,888 abstracts on a laptop, plus 1 second (meaning) to 2½ minutes (words + meaning) to fit about 12–17 a second on 2 CPU cores a small CPU server
Fine-tuned model about 8¼ hours on one NVIDIA L4, from submission to finish about 2.3 a second through the product, start-up included a GPU

We are deliberately not quoting prices: they depend on your provider, region and volume.

7. What this means for life sciences teams

Start from what you have already sorted. Literature-monitoring, medical-information and evidence teams often hold thousands of papers already tagged by therapeutic area, product or topic. That is a training file. Hold some back to score the model before trusting it.

Name your groups so they don't overlap, or decide how overlaps are labeled. In this benchmark, 78% of the fine-tuned model's mistakes came from one broad group that overlaps the others. If an abstract can belong to both "general" and a specific system, write down which wins. A model cannot learn a distinction your labels do not make consistently.

Set the cutoff from your own test, and decide where review happens. Score a sample whose answers you already know, look at what each cutoff would have caught, and send everything below it to a person.

Where it fits:

  • first-pass sorting of new publications by therapeutic area or topic;
  • routing papers, inquiries or internal documents to the right team;
  • building and cleaning a tagged literature library;
  • prioritizing what a specialist reads first.

Where it does not:

  • it sorts text into the categories you define; it does not interpret findings, assess evidence quality or give medical advice;
  • it is not a medical device and makes no diagnostic or treatment decision;
  • it is not a substitute for the human review your quality system or regulator requires, for example in adverse-event case processing.

What teams ask, answered plainly:

  • A person reviews what the model is unsure about.
  • Changes are recorded in an activity log.
  • Each model is trained on your examples for your use, not pooled with anyone else's.
  • Your data can be exported and deleted.
  • Use published or de-identified text; this benchmark used no patient records.

8. What these results do, and do not, show

  • One run of each, on one test set. Run-to-run variation is unmeasured; small gaps are not firm rankings.
  • Five broad condition groups. Real literature taxonomies are often finer (dozens of therapeutic areas or MeSH-level topics) and harder.
  • Abstracts only, in English, from the corpus's source; full texts, other languages and other document types are untested.
  • Cutoffs chosen after the fact on the test results; in practice, choose them on a separate validation set.
  • Long abstracts are cut by the embedding model at 512 word-pieces. The median abstract is 177 words and only 2.1% run past 350. Word counts do not establish how much text is truncated; token lengths vary.
  • Two measurement paths. The fine-tuned model is measured through our product end to end. The other approaches are measured with a benchmark script on the same rows, with the same scoring.

Method details

  • Data: Medical Abstracts TC corpus, train and test files as published (medical_tc_train.csv, medical_tc_test.csv). 10,000 training abstracts after a fixed shuffle (seed 17); all 2,888 test abstracts. Smaller training sets are the first N per condition group in that order, identical across approaches.
  • Word matching: TF-IDF over words and two-word phrases (sublinear term frequency); nearest neighbors with k = 5 and similarity-weighted voting; logistic regression with C = 10, fixed in advance.
  • Embeddings: intfloat/multilingual-e5-small, inputs prefixed query: , normalized 384-dimension vectors, 512-token limit; logistic regression, C = 10.
  • Words + meaning: the two feature sets above, each L2-normalized per abstract, concatenated; logistic regression, C = 10.
  • Fine-tuned model: Qwen3.5-4B; LoRA r = 16; 3 epochs; one NVIDIA L4, bfloat16; 29,863 billed training seconds (8 h 18 min) on one ml.g6.xlarge, after a 6-minute wait for a GPU. A stratified 15% of the training abstracts was held back for validation.
  • Confidence: the highest class probability for the inexpensive approaches (nearest neighbors: the winning group's share of the weighted vote); for the fine-tuned model, its own probability for its answer.
  • Hardware: the inexpensive approaches ran on a laptop; throughput is measured on two CPU threads.

Sources and attribution

  • Medical Abstracts corpus: Schopf, Braun & Matthes, Evaluating Unsupervised Text Classification: Zero-Shot and Similarity-Based Approaches, NLPIR 2022 (ACM publication, 2023; paper). Dataset, licensed CC BY-SA 3.0. We report measurements on the data; we do not redistribute it.
  • multilingual-e5-small: Wang et al., Multilingual E5 Text Embeddings (intfloat), MIT license.
  • Qwen3.5-4B: the Qwen team, via Hugging Face.

StayCharted's measurements and analysis are reported independently of the dataset's and models' authors. Nothing in this report is medical advice.

Frequently asked questions

What is the best way to classify medical abstracts by condition?

On this benchmark, a fine-tuned language model was the most accurate at 63.6% across five condition groups, against 60.0% for the best inexpensive approach (embeddings with logistic regression). The fine-tuned model took about 8¼ hours on a GPU to train; the embeddings approach took minutes on a laptop.

How many labeled examples do I need to train a medical text classifier?

With 10 labeled abstracts per condition group, the best inexpensive approach reached 49.8%; with 100 per group, 57.9%; with 10,000 in total, 60.0%. Your own categories and text will differ, so hold back a sample and measure.

Can AI sort research papers without anyone checking them?

Not on overlapping categories like these. Even the most accurate approach here left more than one abstract in three wrong before review. Flagging the answers it is least sure of for a person projects 97.5% when 78.7% are checked, assuming every flagged error is corrected. On clearer categories far less checking is needed: 9.2% on our banking benchmark.

Was patient data used?

No. The benchmark uses abstracts of published research from a public corpus (CC BY-SA 3.0). It contains no patient records.

Is this a medical device or a diagnostic tool?

No. It sorts text into categories you define, such as therapeutic areas or topics. It does not diagnose, recommend treatment or interpret findings.

Do word-based models or embeddings work better on medical text?

Embeddings, on this benchmark: 60.0% against 49.9% for the better word-based method with 10,000 abstracts, and 49.8% against 36.8% with 10 per group. Combining the two did not help at full size (52.0%).

How is the review cutoff chosen?

From a test on rows whose answers you already know: each cutoff shows how many answers it would flag and how many mistakes it would catch. In StayCharted AI Model Trainer each model has its own cutoff, and the answers below it go to a Review Queue for a person to confirm or correct.

Put the findings into practice

Explore classification workflows →

Compare the contract clause benchmark →

Compare plans →

Explore more Insights

Compare AMT’s design decisions and training reports →

STAYCHARTED AI MODEL TRAINER

Put your business knowledge to work.

Build your Business-Specific AI with AMT. Start with examples you already have, test your model’s results, and put it into your workflow.

Try it on your own data