Developing AMT
Four ways to label contract clauses, measured on LEDGAR
Legal clause classification on LEDGAR: compare words, embeddings, hybrid classifiers, and fine-tuning across accuracy, training time, and projected review effort.
The short version
Legal teams sort contract text by clause type all the time: building clause libraries, reviewing third-party paper, due diligence, renewals. We measured four ways to automate that on LEDGAR, a public benchmark of real contract provisions from US securities filings, labeled with 100 clause types.
| Word matching | Meaning (embeddings) | Words + meaning | Fine-tuned language model | |
|---|---|---|---|---|
| Accuracy, 10,000 training clauses | 74.5% – 82.3%¹ | 78.9% | 82.7% | 84.9% |
| Accuracy, 10 examples per clause type | 60.3% – 69.0%¹ | 64.8% | 70.3% | not measured |
| Share reviewed for projected ~98% accuracy | 49% – 54%¹ | 53% | 44% | **42%**² |
| Time to train, 10,000 clauses | seconds | about 1½ minutes on a laptop GPU | about 1½ minutes on a laptop GPU | about 6 hours on a GPU |
| Hardware | any CPU | a small CPU server | a small CPU server | a GPU |
¹ Two word-matching methods; see below. ² Projects 97.6% after review at the highest tested cutoff (0.99). See §5.
The fine-tuned model had the highest measured accuracy in this run. It leads words + meaning by 2.2 percentage points on contract clauses, while taking about six hours to train. Small gaps from one run are not firm rankings.
On contract language, words carry most of the signal. The exact vocabulary lawyers use names most clause types, so the inexpensive approaches get close.
Each tested approach projects roughly 98% accuracy after reviewing about two clauses in five or more. The fine-tuned model's checking share is the lowest, but barely: 42% against 44%. On banking messages it cut the checking to half of the next best.
Many mistakes involve closely related clause labels. "Definitions" for "Defined Terms", "Waivers" for "No Waivers", "Applicable Laws" for "Governing Laws". The benchmark counts these as wrong. Whether the labels are interchangeable depends on the clause and the legal team’s taxonomy.
1. The benchmark
LEDGAR is a set of contract provisions (paragraphs) from agreements filed with the US Securities and Exchange Commission, each labeled with its clause type. It was introduced by Tuggener, von Däniken, Peetz and Cieliebak (LREC 2020). It is distributed as part of the LexGLUE legal-language benchmark (Chalkidis et al., ACL 2022), under CC BY 4.0. LexGLUE fixes the split: 60,000 training, 10,000 validation and 10,000 test provisions, 100 clause types.
The most common types are Governing Laws, Counterparts, Notices, Entire Agreements, Severability, Survival, Expenses, Assignments, Amendments and Terminations. The rarest occur a handful of times: Anti-Corruption Laws, Venues, Qualifications. A typical provision is 87 words long; one in ten runs past 230; the longest is nearly a thousand.
What we used. 10,000 provisions from the training split and 4,000 from the test split, drawn in a fixed shuffled order, the same rows for every approach.
- The smaller training size keeps the fine-tuned model's run to a working day. It also makes the task harder than the full benchmark (see §4).
- To see how the inexpensive approaches cope with little data, we also trained them on the first 5, 10, 20 and 50 examples of each clause type.
2. The approaches
1. Word matching (TF-IDF). Each clause becomes a weighted count of the words and two-word phrases it contains. There are two ways of deciding a label:
- Nearest neighbors: find the five most similar training clauses and let them vote.
- Logistic regression: learn a weight for every word and phrase per clause type.
2. Meaning (embeddings). A small pre-trained multilingual model (multilingual-e5-small, about 100 million parameters) turns each clause into 384 numbers that place similar meanings close together. A logistic regression learns where each clause type sits. The model reads at most the first 512 word-pieces of a clause, which is enough for most clauses but not all.
3. Words + meaning. Both of the above, side by side, in one logistic regression that can use whichever carries the signal.
4. A fine-tuned language model. Qwen3.5-4B, a 4-billion-parameter open model, adapted with LoRA to write the clause type for each provision.
- It was trained on one NVIDIA L4 GPU and measured end to end through StayCharted AI Model Trainer. We uploaded a spreadsheet, trained, then filled a second spreadsheet whose label column was blank.
- This is the same setup as our banking report.
3. Accuracy, by how much training data you have
| training clauses | Word matching (nearest neighbors) | Word matching (logistic regression) | Meaning (embeddings) | Words + meaning | Fine-tuned model |
|---|---|---|---|---|---|
| 5 per type (496) | 54.3% | 61.1% | 57.8% | 62.7% | — |
| 10 per type (987) | 60.3% | 69.0% | 64.8% | 70.3% | — |
| 20 per type (1,947) | 64.9% | 74.5% | 70.7% | 75.2% | — |
| 50 per type (4,538) | 69.3% | 77.9% | 74.3% | 78.1% | — |
| all (10,000)² | 74.5% | 82.3% | 78.9% | 82.7% | 84.9% |
Random guessing over 100 clause types scores about 1%.
² The fine-tuned model trained on about 85% of these and held the rest back for its own validation; the other approaches used all 10,000. The fine-tuned model was measured at full size only.
What stands out:
- Words beat meaning here, at every size. This is the opposite of what we found on short banking messages. Clause types are largely named by their language: "governed by and construed in accordance with the laws of", "may be executed in counterparts", "shall survive the termination". A general-purpose embedding model has no legal training and reads only the start of long clauses; a word model learns the exact formulas.
- Combining both helps a little, everywhere: 0.2 to 1.6 points over the best single approach at each size.
- How you decide matters more than what you measure. The same word counts score 74.5% with nearest neighbors and 82.3% with logistic regression.
- The fine-tuned model adds 2.2 points to the best inexpensive approach. Its error rate falls from 17.3% to 15.1%: about one fewer mistake in every eight.
What the fine-tuned model gets wrong. Its most frequent mistakes are between closely related clause labels:
| the clause was labeled | the model said | times |
|---|---|---|
| Definitions | Defined Terms | 11 |
| Interpretations | Construction | 10 |
| Waivers | No Waivers | 10 |
| Applicable Laws | Governing Laws | 10 |
| Integration | Entire Agreements | 9 |
| Tax Withholdings | Withholdings | 9 |
| Modifications | Amendments | 9 |
Most of these pairs also go the other way (Withholdings for Tax Withholdings 8 times, No Waivers for Waivers 8, Defined Terms for Definitions 8).
- LEDGAR's labels come from the headings the drafting lawyers chose, and different firms head the same provision differently.
- In the fifteen most frequent confusions alone, 114 of the model's 603 mistakes are between such pairs. Treating those 114 disagreements as equivalent would produce about 88% accuracy. This is a hypothetical relabeling calculation, not the measured benchmark score.
- Review overlapping category names with your legal team. Consistent labels may reduce ambiguity; any improvement requires retraining and evaluation.
4. Against the published results
With all 60,000 training provisions, the LexGLUE paper reports a micro-F1 of 87.0 for a TF-IDF + support-vector-machine baseline. Its fine-tuned transformer models score 87.6–88.3, including models pre-trained on legal text (Legal-BERT 88.2, CaseLaw-BERT 88.3) (Table 3).
Two readings, both useful:
- On this task a word model gets within about a point of the best transformers when it has plenty of data. Our results agree: the fine-tuned model's lead over the inexpensive approaches is small here.
- Our setup used a sixth of the training data (10,000 vs 60,000) and scored on a 4,000-clause sample of the test set. They are not comparable one-to-one with the paper's, and we don't present them as such. (For single-label classification, micro-F1 equals accuracy.)
5. How much checking?
Each approach reports its confidence. Flag low-confidence clauses for a person, and the delivered result rises above the model's own accuracy. The question is how much of the file that takes.
| model accuracy | review this share… | …projected accuracy after review | |
|---|---|---|---|
| Word matching, nearest neighbors | 74.5% | 54.1% | 98.0% |
| Word matching, logistic regression | 82.3% | 48.6% | 98.1% |
| Meaning (embeddings) | 78.9% | 52.9% | 98.0% |
| Words + meaning | 82.7% | 44.4% | 97.8% |
| Fine-tuned model | 84.9% | 41.7% | 97.6% |
Projected accuracy assumes every flagged error is corrected without introducing new errors. Human-review outcomes were not measured. No tested cutoff for the fine-tuned model projects a full 98%; the highest tested cutoff projects 97.6%.
On contract text, the fine-tuned model saves little checking.
- It projects 97.6% after reviewing 41.7% of clauses at a cutoff of 0.99. None of the tested cutoffs reaches 98%: even at 0.99, 16% of its mistakes are answers it was sure of.
- At 0.90, the cutoff that projected 98% after review on banking messages, it flags 18.7% of clauses, catches 57% of its mistakes, and projects 93.5% after review.
- On banking messages, the same model at 0.90 flagged 9.2% and projected 98% after review.
Related clause labels and formulaic contract language may help explain the difference. These are possible explanations, not causes isolated by this benchmark. The cutoff that suits one kind of text can be far too low for another, so it has to be set on your own data. Choose review cutoffs on separate validation data before relying on them.
| Fine-tuned model, cutoff | 0.80 | 0.90 | 0.95 | 0.99 |
|---|---|---|---|---|
| clauses flagged | 13.1% | 18.7% | 26.3% | 41.7% |
| share of its mistakes among them | 44% | 57% | 68% | 84% |
| projected accuracy after review | 91.6% | 93.5% | 95.2% | 97.6% |
For comparison, on short banking messages the inexpensive approaches projected approximately 98% accuracy after reviewing 19% to 44% of rows. On contract text they need 44% to 54%.
Confidence scores are not comparable across methods, so each approach needs its own cutoff. For the inexpensive approaches here, the ~98% points sit at cutoffs between 0.70 and 0.90.
6. What it costs to run
| train on 10,000 clauses | label new clauses | hardware | |
|---|---|---|---|
| Word matching | seconds | hundreds to thousands per second | any CPU |
| Meaning, or words + meaning | about 1½ minutes on a laptop GPU (embedding the clauses) plus seconds to fit | ~45 a second on 2 CPU cores | a small CPU server |
| Fine-tuned model | 5 h 58 min | 3.4 a second (4,000 in about 20 minutes) | a GPU |
Long text is slow to train on.
- The same model trained on the same number of short banking messages in 2 h 16 min. Contract clauses took 2.6 times as long, because each one is several times longer.
- Embedding preparation was measured on a laptop GPU; the 45-clauses-per-second figure was measured on two CPU cores. Training time is not a measured CPU-only result.
- We are deliberately not quoting prices: they depend on your provider, region and volume.
7. What this means for legal teams
Start from the clause library you already have. Most legal-ops and contract teams hold thousands of classified provisions: playbooks, precedent banks, a CLM system's tags. That is a training file. Hold some back to score the model before trusting it.
Test on clauses you already know the answers to, and set the cutoff from that. Contract language makes models confident. Score a sample where you already know the right clause type, and look at what each cutoff would have caught. Do that before deciding what to route straight through.
Use consistent clause labels. Have your legal team review potentially overlapping categories before training. Merge them only when the underlying clauses and your labeling rules justify it; similar names can encode meaningful differences. On this benchmark, nearly one in five of the fine-tuned model's mistakes came from pairs like these.
Expect to review, and decide where. Even the most accurate approach here leaves about one clause in seven wrong before review. Set the cutoff by what a misclassified clause costs you: a misfiled "Counterparts" clause is not a misfiled "Indemnification".
Where it fits: first-pass sorting of third-party paper, building or cleaning a clause library, finding every governing-law or assignment clause across a portfolio, grouping provisions by predicted clause type for a lawyer to review.
Where it does not: it classifies clauses. It does not interpret them, assess risk or give legal advice, and it does not replace a lawyer's review of what a clause says.
What legal teams ask, answered plainly:
- A person reviews what the model is unsure about.
- Changes are recorded.
- The model is trained on your clauses for your use, not pooled with anyone else's.
- Export and deletion options are described in our Terms of Service and Privacy Policy.
8. What these results do, and do not, show
- One run of each, on one sample of the test set. Run-to-run variation is unmeasured; small gaps, including the 2.2 points between the two best approaches, are not firm rankings.
- A sixth of the published training data, deliberately. Numbers are not comparable one-to-one with LexGLUE's.
- Cutoffs chosen after the fact on the test results; in practice, choose them on a separate validation set.
- English, US securities filings. Other jurisdictions, languages and contract types (NDAs, leases, employment) are untested.
- Long clauses are cut by the embedding model at 512 word-pieces.
- The fine-tuned model trains on about the first 440 word-pieces of each clause and reads up to 4,096 when labeling.
- The word models read the whole clause.
- Two measurement paths. The fine-tuned model is measured through our product end to end. The other approaches are measured with a benchmark script on the same rows, with the same scoring.
Method details
- Data: LexGLUE
ledgarconfiguration, train and test splits as published. 10,000 training and 4,000 test provisions after a fixed shuffle (seed 17). Smaller sets are the first N per clause type in that order, identical across approaches. - Word matching: TF-IDF over words and two-word phrases (sublinear term frequency); nearest neighbors with k = 5 and similarity-weighted voting; logistic regression with C = 10, fixed in advance.
- Embeddings:
intfloat/multilingual-e5-small, inputs prefixedquery:, normalised 384-dimension vectors, 512-token limit; logistic regression, C = 10. - Words + meaning: the two feature sets above, each L2-normalised per clause, concatenated; logistic regression, C = 10.
- Fine-tuned model: Qwen3.5-4B; LoRA r = 16; 3 epochs; one NVIDIA L4, bfloat16; 21,508 s of billed training time. A stratified 15% of the training clauses was held back for validation.
- Confidence: the highest class probability for the inexpensive approaches (nearest neighbors: the winning type's share of the weighted vote); for the fine-tuned model, its own probability for its answer.
- Hardware: the inexpensive approaches ran on a laptop CPU/GPU. Embedding 14,000 clauses took about 1.5 minutes on the laptop's GPU and runs at roughly 45 clauses a second on two CPU cores.
Sources and attribution
- LEDGAR: Tuggener, von Däniken, Peetz & Cieliebak, LEDGAR: A Large-Scale Multi-label Corpus for Text Classification of Legal Provisions in Contracts, LREC 2020.
- LexGLUE: Chalkidis, Jana, Hartung, Bommarito, Androutsopoulos, Katz & Aletras, LexGLUE: A Benchmark Dataset for Legal Language Understanding in English, ACL 2022 (arXiv:2110.00976). CC BY 4.0.
- multilingual-e5-small: Wang et al., Multilingual E5 Text Embeddings (intfloat), MIT licence.
- Qwen3.5-4B: the Qwen team, via Hugging Face.
StayCharted's measurements and analysis are reported independently of the dataset's and models' authors.
See our Text Classifier Benchmark and Banking Intent Classification report, or explore document classification.
Put the findings into practice
Contract clause classification for legal teams →