Developing AMT
How StayCharted AMT Checks Sensitive Training Data—and Why It’s Included on Every Plan
How StayCharted AMT detects and masks recognizable personal information before training, applies masking to new text, and handles the limits of pattern-based checks.
When you train on business data, sensitive values can become part of what a model learns. Email addresses, phone numbers, or card numbers may be retained in a vocabulary, an exported model, or learned patterns. Deleting the original spreadsheet does not undo the training.
StayCharted AMT checks text examples for recognizable personal information and secrets before training. You can mask detected values, exclude affected rows, or keep them. This check is included on every plan, including Free. It reduces exposure; it does not guarantee that all sensitive information is found.
What the personal-information check does
In StayCharted AMT, the check is the first step under Data quality → Personal information for models that learn from text.
| Type | How it is checked |
|---|---|
| Email addresses | Recognizable address structure. |
| Payment card numbers | Luhn checksum and recognized card-network prefixes. These checks do not prove a card is active or belongs to someone. |
| Bank account numbers (IBAN) | The IBAN mod-97 checksum. |
| US Social Security numbers | Dashed format with rules to exclude invalid number ranges. |
| Phone numbers | 10–15 digits in phone-like layouts, with rules to reduce matches on dates, times, and amounts. |
| IP addresses | IPv4 and IPv6 patterns, with rules to reduce false matches on software versions. |
| Passwords and API keys | Recognized provider key formats, private-key blocks, and login-token patterns—not every possible password or secret. |
| Credentials inside links | Usernames and passwords embedded in web addresses. |
The findings show the number of affected rows and partially hidden examples such as j•••@acme.com or •••• 4242. The findings preview does not display the full detected value.
You choose what to do with each detected type
- Mask it—the recommended action. Replace detected values with placeholders such as
[MASKED-EMAIL]. The row keeps its category, so the model can learn that an email address appeared without learning that address. - Exclude affected rows. Leave those examples out of training.
- Keep the values. Retain them when appropriate for your task. Your decision is recorded with who made it and when.
The check offers a decision point rather than an automatic training block. Review its findings in the context of your data and task.
The model reads new text the same way
A StayCharted AMT model trained on masked text applies the same masking before reading new text in the Try box, file fills, and API calls. Masking also travels with exported models so the same rules can run outside the application.
This keeps the model’s input consistent with its training examples. Your filled files retain their original text; masking changes what the model reads. It does not delete or anonymize the original upload.
Why we built the detectors inside StayCharted AMT
StayCharted AMT checks structured values with recognizable formats. Here is how that helps you prepare your training data.
1. Keep the check where the data already lives
Your text is not sent to an additional provider for this check.
2. Reduce false alarms
A long order number can resemble a payment card number, and a date can resemble a phone number. Checksum, prefix, and exclusion rules help reduce those mistakes. We test values that should match and examples that should not. Passing a format check is evidence for a finding, not proof of someone’s identity.
3. Include the check on every plan
StayCharted AMT includes personal-information checks on every plan, including Free, with no separate inspection charge.
4. Use consistent rules across the workflow
The same masking rules apply during training and prediction, including exported models.
5. Check structure without relying on sentence language
Email addresses, card numbers, and IBANs have recognizable structures even when the surrounding text changes language. This helps with structured identifiers; it does not mean all languages, local identifier formats, or kinds of personal information are covered.
6. Make findings explainable
A reason such as “matches a recognized card prefix and passes the checksum” gives you something concrete to inspect when deciding whether to mask a value.
What the check does not catch yet
StayCharted AMT does not currently detect people’s names, street addresses, or health details written in free text. Those require more context than a fixed-format pattern supplies. Remove these details yourself before training when they should not be included.
National identity numbers outside the US are not comprehensively covered. Pictures are not scanned for personal information written within them, although location and other hidden image metadata are removed on upload.
No automated detector finds everything. Presidio’s documentation also describes this limitation. Masking detected values is a risk-reduction measure, not a guarantee that the remaining data cannot identify a person.
Why it is included on every plan
The risk from sensitive training data does not depend on a subscription tier. StayCharted AMT includes this first layer of detection on every plan, including Free, and recommends masking as the default choice.
Reducing sensitive values in models can reduce exposure for your team and the people represented in your data. The structured-value check is included on every plan.
Sources and further reading
- Presidio: detection approaches and limitations
- Amazon Comprehend: PII real-time analysis
- Amazon Comprehend: DetectPiiEntities API
- Google Cloud Sensitive Data Protection: pricing
These sources describe alternative detection tools. This article explains the design of StayCharted AMT’s check; it does not compare detection accuracy across products.