← All insights

Developing AMT

Can a model learn who should have access?

StayCharted’s Classifier Model scored 97.8% on 72,700 held-out group-access decisions, with 10% fewer errors than a team-based rule. Compare errors and review workload.

By Manoj Mohandas · Measurements: October 8, 2026 · Published

The data

There's no public export of a real company's SharePoint permissions, for good reason. The closest public record of real access decisions is the Amazon Access Samples dataset: 36,063 Amazon employees, published anonymised on the UCI Machine Learning Repository.

For each person it records:

  • 11 attributes: department, business title, title detail, job code, job family, location, company, three levels of the organisation, and manager;
  • every access group they hold.

Every value is an anonymised ID number, so the model never sees a real department name or title.

We picked the 10 groups held by the most people (between 5.5% and 10.8% of employees each) and asked one question per group: should this person be in it?

We set aside 7,270 people (20%, chosen by a fixed hash of their ID) before anything was trained. Every figure below is measured on them. The model learned from the other 28,793.

What we compared

StayCharted's Classifier Model learned each group from the people already in it. It used the same training engine customers use in the app. Each group trained in about five seconds on an ordinary CPU, with no GPU and no rules written by anyone.

We compared it with the rules an administrator would set up by hand:

  • Grant nobody. The cautious default.
  • Same as most of the department.
  • Same as most people with the same job in the department.
  • Same as most of the team: the people under the same manager. For a manager it hasn't seen before, it falls back to the department.

Why accuracy alone misleads

"Grant nobody" is right 92.3% of the time, because most people aren't in any given group. It is also useless: it leaves out every single person who needs access.

So we report the two errors separately, because they aren't the same error:

  • Given access wrongly: of the people who shouldn't have access, the share who would get it. This is oversharing.
  • Left out: of the people who should have access, the share who wouldn't get it. This is a support ticket and a frustrated new starter.

The results

All 10 groups, 72,700 decisions on people the model never saw:

Right Given access wrongly Left out Wrong answers
Grant nobody 92.3% 0% 100% 5,609
Same as most of the department 94.4% 0.8% 62.9% 4,064
Same as same job in the department 94.5% 0.9% 61.7% 4,028
Same as most of the team 97.6% 0.9% 21.1% 1,772
StayCharted Classifier Model 97.8% 0.9% 17.7% 1,593

Three things stand out.

The team decides access, not the job title. Rules built on department or job leave out about 62% of the people who need each group. Copying the team does far better, because access follows the work a team does.

The model found that on its own. Nobody told it that the manager mattered. It learned from the decisions, and it made 10% fewer mistakes than the team rule (1,593 against 1,772). It left out fewer people (17.7% against 21.1%) and gave access wrongly no more often (0.9%).

For new teams, it does better than the fallback rule. For the 123 held-back people whose manager had nobody in the training data, the team rule falls back to the department. There the model left out 31.5% of the people who needed access, against 39.8% for the rule. That slice is small, so treat it as a direction rather than a result.

Knowing when it isn't sure

A model that is right 97.8% of the time is still wrong on thousands of decisions at scale. What matters is whether it can point to them.

Every answer comes with a confidence. Answers below a chosen cutoff go to a person. Here's what that does, measured on the same held-back people:

Check answers below Share sent to a person Mistakes among them
0.60 0.9% 295 of 1,593
0.70 1.9% 537
0.80 3.4% 838
0.90 5.9% 1,139 (72%)
0.95 8.6% 1,308 (82%)

At 0.90, a person looks at 5.9% of the decisions, and that 5.9% contains 72% of all the model's mistakes.

If those checks are done and the wrong answers corrected, 99.4% of decisions would be right. That's a projection: it assumes every flagged answer is checked and corrected. The 97.8% above is what the model got right on its own.

Could a rule flag its own unsure cases? Partly. We ranked the team rule's answers by how split the team was, and gave it the same budget. Checking the same 5.9% catches 1,083 of its 1,772 mistakes, against the model's 1,139 of 1,593. A carefully built rule comes close. The difference is that nobody has to build it, or rebuild it for every group, or keep it current as teams change.

How much history is enough

We retrained on smaller slices of the history, measured on the same 7,270 people:

People in the history Right Given access wrongly Left out Sent to a person at 0.90 Right after checking (projected)
500 95.7% 1.9% 33.0% 11.0% 99.0%
1,000 95.8% 1.8% 32.7% 10.5% 99.1%
2,000 96.6% 1.5% 27.2% 9.0% 99.2%
5,000 97.2% 1.2% 21.5% 7.2% 99.2%
10,000 97.6% 1.0% 19.3% 6.5% 99.3%
28,793 97.8% 0.9% 17.7% 5.9% 99.4%

It's useful from a few hundred past decisions. More history mostly means fewer people left out and fewer answers to check. Each group is held by 5–11% of people, so the smallest slice had roughly 25 to 55 members per group to learn from.

Where it fails

Two of the 10 groups couldn't be predicted from these attributes by any method we tried. Even the model left out half or more of the people who needed them (49.0% and 62.7%). Membership of those groups depends on something this data doesn't record: a project, a one-off request, a person's history.

These errors make per-group evaluation important. Confidence can help prioritize review, but it does not catch every incorrect access decision.

What this does and doesn't show

  • This is people to groups, not files. It shows a model can learn an organisation's own access decisions from its history. It doesn't measure file or site permissions, where the inputs are names, folders and owners rather than job attributes. We'd expect the same pattern, but we haven't measured it.
  • The attributes are anonymised codes. A real directory has department names and titles, which give a model more to work with.
  • One company, from 2011. And the held-out people are a random sample of current employees, not future hires.
  • The 99.4% is a projection. It assumes every flagged answer is checked and corrected.

Doing this with your own data

The same approach works for any access decision your team already makes. For SharePoint, an administrator can:

  1. Export current permissions (from Microsoft Graph, PowerShell or a SharePoint report).
  2. Upload the export to StayCharted and train.
  3. Ask for suggestions for each new site, file or starter, through a file or the API.
  4. Apply the confident ones, and check the rest.

StayCharted doesn't connect to your tenant or change permissions. It suggests, and you decide.

Start with one department's export, and see how often it agrees with the decisions your team already made before you use it on anything new.


Method notes

  • Data: Amazon Access Samples, UCI Machine Learning Repository, dataset 216 (doi:10.24432/C5JW2K), licensed CC BY 4.0. File amzn-anon-access-samples-2.0.csv: 36,063 people after dropping 716,063 empty trailing rows.
  • Groups: the 10 most-held groups held by 5–60% of people (5.5% to 10.8%).
  • Split: a fixed hash of the person ID (seed staycharted-access-bench-2026-10-08): 28,793 to learn from, 7,270 held out. The split is reproducible from that seed and the source IDs.
  • Model:
    • StayCharted Classifier Model (words), one model per group.
    • Each person is one row of "Field: value" lines.
    • Each value is tagged with its field, so the same ID number in two fields stays two different words.
    • The engine holds back part of the training rows to choose its own settings, as it does in the app.
    • Run from the app's own training code (8 October 2026), outside the app's screens.
  • Team rule: the majority of training people with the same manager, falling back to the department majority, then to "no".
  • Review figures: confidence below the cutoff, checked against the truth. "Right after checking" assumes each flagged answer is corrected.

StayCharted AI Model Trainer learns how your team categorises things from your own past decisions, and sends the uncertain ones to a person.

Model your file and SharePoint permissions →

Frequently asked questions

Can AI learn permission decisions from past assignments?

In this test, StayCharted’s Classifier Model learned historical membership for 10 groups and achieved 97.8% accuracy across 72,700 held-out decisions. It reproduced recorded assignments; the test did not establish whether those assignments were appropriate security policy.

Was 99.4% accuracy measured?

No. It is a projection assuming a person checks the least-confident 5.9% of decisions and corrects every flagged error. Unassisted accuracy was 97.8%.

Does this benchmark test SharePoint files?

No. It tests group membership for people in Amazon Access Samples. File and site permissioning requires its own evaluation.

Explore more Insights

Compare AMT’s design decisions and training reports →

STAYCHARTED AI MODEL TRAINER

Train AI to categorize the way your team does.

Start with examples you already have, see how often it’s right on items it never saw, and use confidence to prioritize the ones a person should review.

Start free