← All insights

Developing AMT

Four ways to label contract clauses, measured on LEDGAR

Legal clause classification on LEDGAR: compare words, embeddings, hybrid classifiers, and fine-tuning across accuracy, training time, and projected review effort.

By Manoj Mohandas · Measurements: September 24 and October 3, 2026 · Published

The short version

Legal teams sort contract text by clause type all the time: building clause libraries, reviewing third-party paper, due diligence, renewals. We measured four ways to automate that on LEDGAR, a public benchmark of real contract provisions from US securities filings, labeled with 100 clause types.

Word matching Meaning (embeddings) Words + meaning Fine-tuned language model
Accuracy, 10,000 training clauses 74.5% – 82.3%¹ 78.9% 82.7% 84.9%
Accuracy, 10 examples per clause type 60.3% – 69.0%¹ 64.8% 70.3% not measured
Share reviewed for projected ~98% accuracy 49% – 54%¹ 53% 44% **42%**²
Time to train, 10,000 clauses seconds about 1½ minutes on a laptop GPU about 1½ minutes on a laptop GPU about 6 hours on a GPU
Hardware any CPU a small CPU server a small CPU server a GPU

¹ Two word-matching methods; see below. ² Projects 97.6% after review at the highest tested cutoff (0.99). See §5.

The fine-tuned model had the highest measured accuracy in this run. It leads words + meaning by 2.2 percentage points on contract clauses, while taking about six hours to train. Small gaps from one run are not firm rankings.

On contract language, words carry most of the signal. The exact vocabulary lawyers use names most clause types, so the inexpensive approaches get close.

Each tested approach projects roughly 98% accuracy after reviewing about two clauses in five or more. The fine-tuned model's checking share is the lowest, but barely: 42% against 44%. On banking messages it cut the checking to half of the next best.

Many mistakes involve closely related clause labels. "Definitions" for "Defined Terms", "Waivers" for "No Waivers", "Applicable Laws" for "Governing Laws". The benchmark counts these as wrong. Whether the labels are interchangeable depends on the clause and the legal team’s taxonomy.

1. The benchmark

LEDGAR is a set of contract provisions (paragraphs) from agreements filed with the US Securities and Exchange Commission, each labeled with its clause type. It was introduced by Tuggener, von Däniken, Peetz and Cieliebak (LREC 2020). It is distributed as part of the LexGLUE legal-language benchmark (Chalkidis et al., ACL 2022), under CC BY 4.0. LexGLUE fixes the split: 60,000 training, 10,000 validation and 10,000 test provisions, 100 clause types.

The most common types are Governing Laws, Counterparts, Notices, Entire Agreements, Severability, Survival, Expenses, Assignments, Amendments and Terminations. The rarest occur a handful of times: Anti-Corruption Laws, Venues, Qualifications. A typical provision is 87 words long; one in ten runs past 230; the longest is nearly a thousand.

What we used. 10,000 provisions from the training split and 4,000 from the test split, drawn in a fixed shuffled order, the same rows for every approach.

  • The smaller training size keeps the fine-tuned model's run to a working day. It also makes the task harder than the full benchmark (see §4).
  • To see how the inexpensive approaches cope with little data, we also trained them on the first 5, 10, 20 and 50 examples of each clause type.

2. The approaches

1. Word matching (TF-IDF). Each clause becomes a weighted count of the words and two-word phrases it contains. There are two ways of deciding a label:

  • Nearest neighbors: find the five most similar training clauses and let them vote.
  • Logistic regression: learn a weight for every word and phrase per clause type.

2. Meaning (embeddings). A small pre-trained multilingual model (multilingual-e5-small, about 100 million parameters) turns each clause into 384 numbers that place similar meanings close together. A logistic regression learns where each clause type sits. The model reads at most the first 512 word-pieces of a clause, which is enough for most clauses but not all.

3. Words + meaning. Both of the above, side by side, in one logistic regression that can use whichever carries the signal.

4. A fine-tuned language model. Qwen3.5-4B, a 4-billion-parameter open model, adapted with LoRA to write the clause type for each provision.

  • It was trained on one NVIDIA L4 GPU and measured end to end through StayCharted AI Model Trainer. We uploaded a spreadsheet, trained, then filled a second spreadsheet whose label column was blank.
  • This is the same setup as our banking report.

3. Accuracy, by how much training data you have

training clauses Word matching (nearest neighbors) Word matching (logistic regression) Meaning (embeddings) Words + meaning Fine-tuned model
5 per type (496) 54.3% 61.1% 57.8% 62.7% —
10 per type (987) 60.3% 69.0% 64.8% 70.3% —
20 per type (1,947) 64.9% 74.5% 70.7% 75.2% —
50 per type (4,538) 69.3% 77.9% 74.3% 78.1% —
all (10,000)² 74.5% 82.3% 78.9% 82.7% 84.9%

Random guessing over 100 clause types scores about 1%.

² The fine-tuned model trained on about 85% of these and held the rest back for its own validation; the other approaches used all 10,000. The fine-tuned model was measured at full size only.

What stands out:

  • Words beat meaning here, at every size. This is the opposite of what we found on short banking messages. Clause types are largely named by their language: "governed by and construed in accordance with the laws of", "may be executed in counterparts", "shall survive the termination". A general-purpose embedding model has no legal training and reads only the start of long clauses; a word model learns the exact formulas.
  • Combining both helps a little, everywhere: 0.2 to 1.6 points over the best single approach at each size.
  • How you decide matters more than what you measure. The same word counts score 74.5% with nearest neighbors and 82.3% with logistic regression.
  • The fine-tuned model adds 2.2 points to the best inexpensive approach. Its error rate falls from 17.3% to 15.1%: about one fewer mistake in every eight.

What the fine-tuned model gets wrong. Its most frequent mistakes are between closely related clause labels:

the clause was labeled the model said times
Definitions Defined Terms 11
Interpretations Construction 10
Waivers No Waivers 10
Applicable Laws Governing Laws 10
Integration Entire Agreements 9
Tax Withholdings Withholdings 9
Modifications Amendments 9

Most of these pairs also go the other way (Withholdings for Tax Withholdings 8 times, No Waivers for Waivers 8, Defined Terms for Definitions 8).

  • LEDGAR's labels come from the headings the drafting lawyers chose, and different firms head the same provision differently.
  • In the fifteen most frequent confusions alone, 114 of the model's 603 mistakes are between such pairs. Treating those 114 disagreements as equivalent would produce about 88% accuracy. This is a hypothetical relabeling calculation, not the measured benchmark score.
  • Review overlapping category names with your legal team. Consistent labels may reduce ambiguity; any improvement requires retraining and evaluation.

4. Against the published results

With all 60,000 training provisions, the LexGLUE paper reports a micro-F1 of 87.0 for a TF-IDF + support-vector-machine baseline. Its fine-tuned transformer models score 87.6–88.3, including models pre-trained on legal text (Legal-BERT 88.2, CaseLaw-BERT 88.3) (Table 3).

Two readings, both useful:

  • On this task a word model gets within about a point of the best transformers when it has plenty of data. Our results agree: the fine-tuned model's lead over the inexpensive approaches is small here.
  • Our setup used a sixth of the training data (10,000 vs 60,000) and scored on a 4,000-clause sample of the test set. They are not comparable one-to-one with the paper's, and we don't present them as such. (For single-label classification, micro-F1 equals accuracy.)

5. How much checking?

Each approach reports its confidence. Flag low-confidence clauses for a person, and the delivered result rises above the model's own accuracy. The question is how much of the file that takes.

model accuracy review this share… …projected accuracy after review
Word matching, nearest neighbors 74.5% 54.1% 98.0%
Word matching, logistic regression 82.3% 48.6% 98.1%
Meaning (embeddings) 78.9% 52.9% 98.0%
Words + meaning 82.7% 44.4% 97.8%
Fine-tuned model 84.9% 41.7% 97.6%

Projected accuracy assumes every flagged error is corrected without introducing new errors. Human-review outcomes were not measured. No tested cutoff for the fine-tuned model projects a full 98%; the highest tested cutoff projects 97.6%.

On contract text, the fine-tuned model saves little checking.

  • It projects 97.6% after reviewing 41.7% of clauses at a cutoff of 0.99. None of the tested cutoffs reaches 98%: even at 0.99, 16% of its mistakes are answers it was sure of.
  • At 0.90, the cutoff that projected 98% after review on banking messages, it flags 18.7% of clauses, catches 57% of its mistakes, and projects 93.5% after review.
  • On banking messages, the same model at 0.90 flagged 9.2% and projected 98% after review.

Related clause labels and formulaic contract language may help explain the difference. These are possible explanations, not causes isolated by this benchmark. The cutoff that suits one kind of text can be far too low for another, so it has to be set on your own data. Choose review cutoffs on separate validation data before relying on them.

Fine-tuned model, cutoff 0.80 0.90 0.95 0.99
clauses flagged 13.1% 18.7% 26.3% 41.7%
share of its mistakes among them 44% 57% 68% 84%
projected accuracy after review 91.6% 93.5% 95.2% 97.6%

For comparison, on short banking messages the inexpensive approaches projected approximately 98% accuracy after reviewing 19% to 44% of rows. On contract text they need 44% to 54%.

Confidence scores are not comparable across methods, so each approach needs its own cutoff. For the inexpensive approaches here, the ~98% points sit at cutoffs between 0.70 and 0.90.

6. What it costs to run

train on 10,000 clauses label new clauses hardware
Word matching seconds hundreds to thousands per second any CPU
Meaning, or words + meaning about 1½ minutes on a laptop GPU (embedding the clauses) plus seconds to fit ~45 a second on 2 CPU cores a small CPU server
Fine-tuned model 5 h 58 min 3.4 a second (4,000 in about 20 minutes) a GPU

Long text is slow to train on.

  • The same model trained on the same number of short banking messages in 2 h 16 min. Contract clauses took 2.6 times as long, because each one is several times longer.
  • Embedding preparation was measured on a laptop GPU; the 45-clauses-per-second figure was measured on two CPU cores. Training time is not a measured CPU-only result.
  • We are deliberately not quoting prices: they depend on your provider, region and volume.

8. What these results do, and do not, show

  • One run of each, on one sample of the test set. Run-to-run variation is unmeasured; small gaps, including the 2.2 points between the two best approaches, are not firm rankings.
  • A sixth of the published training data, deliberately. Numbers are not comparable one-to-one with LexGLUE's.
  • Cutoffs chosen after the fact on the test results; in practice, choose them on a separate validation set.
  • English, US securities filings. Other jurisdictions, languages and contract types (NDAs, leases, employment) are untested.
  • Long clauses are cut by the embedding model at 512 word-pieces.
    • The fine-tuned model trains on about the first 440 word-pieces of each clause and reads up to 4,096 when labeling.
    • The word models read the whole clause.
  • Two measurement paths. The fine-tuned model is measured through our product end to end. The other approaches are measured with a benchmark script on the same rows, with the same scoring.

Method details

  • Data: LexGLUE ledgar configuration, train and test splits as published. 10,000 training and 4,000 test provisions after a fixed shuffle (seed 17). Smaller sets are the first N per clause type in that order, identical across approaches.
  • Word matching: TF-IDF over words and two-word phrases (sublinear term frequency); nearest neighbors with k = 5 and similarity-weighted voting; logistic regression with C = 10, fixed in advance.
  • Embeddings: intfloat/multilingual-e5-small, inputs prefixed query: , normalised 384-dimension vectors, 512-token limit; logistic regression, C = 10.
  • Words + meaning: the two feature sets above, each L2-normalised per clause, concatenated; logistic regression, C = 10.
  • Fine-tuned model: Qwen3.5-4B; LoRA r = 16; 3 epochs; one NVIDIA L4, bfloat16; 21,508 s of billed training time. A stratified 15% of the training clauses was held back for validation.
  • Confidence: the highest class probability for the inexpensive approaches (nearest neighbors: the winning type's share of the weighted vote); for the fine-tuned model, its own probability for its answer.
  • Hardware: the inexpensive approaches ran on a laptop CPU/GPU. Embedding 14,000 clauses took about 1.5 minutes on the laptop's GPU and runs at roughly 45 clauses a second on two CPU cores.

Sources and attribution

  • LEDGAR: Tuggener, von Däniken, Peetz & Cieliebak, LEDGAR: A Large-Scale Multi-label Corpus for Text Classification of Legal Provisions in Contracts, LREC 2020.
  • LexGLUE: Chalkidis, Jana, Hartung, Bommarito, Androutsopoulos, Katz & Aletras, LexGLUE: A Benchmark Dataset for Legal Language Understanding in English, ACL 2022 (arXiv:2110.00976). CC BY 4.0.
  • multilingual-e5-small: Wang et al., Multilingual E5 Text Embeddings (intfloat), MIT licence.
  • Qwen3.5-4B: the Qwen team, via Hugging Face.

StayCharted's measurements and analysis are reported independently of the dataset's and models' authors.

See our Text Classifier Benchmark and Banking Intent Classification report, or explore document classification.

Put the findings into practice

Contract clause classification for legal teams →

Compare the banking benchmark →

Compare plans →

Explore more Insights

Compare AMT’s design decisions and training reports →

STAYCHARTED AI MODEL TRAINER

Put your business knowledge to work.

Build your Business-Specific AI with AMT. Start with examples you already have, test your model’s results, and put it into your workflow.

Try it on your own data