Desert Ant Labs

Redact

Multilingual on-device PII detection and redaction.

Redact model page
Platforms
iOS, macOS, tvOS, visionOS, Android, Linux, Windows, Browser, Node
Languages
27
Weights
v0.4.0

Install

requirements

Swift
.package(url: "https://github.com/Desert-Ant-Labs/desert-ant-core.git", from: "3.5.0")

Then add the Redact product to your target.

requirements

Kotlin
implementation("ai.desertant:redact:3.5.0")

requirements

Terminal
npm i @desert-ant-labs/redact @litertjs/core   # browser
npm i @desert-ant-labs/redact                  # Node, prebuilt native core

Usage

Redaction is reversible. Mask personal data before sending text to an LLM, then restore the originals in the reply, on device.

Swift
import Redact

let redact = Redact()
let result = try await redact.redaction(of: "Email Anna Kovács at anna@example.hu.")

print(result.redactedText)
// Email [GIVEN_NAME_1] [SURNAME_1] at [EMAIL_1].

for item in result.items {
    print(item.label.displayName, item.original, item.placeholder, item.confidence)
}

let reply = try await myLLM.rewrite(result.redactedText)
let restored = result.restore(reply)

Filter by category, or raise the confidence floor:

Swift
let options = Options(minimumConfidence: 0.7, labels: [.email, .phone, .creditCard])
let contactOnly = try await redact.redaction(of: text, options: options)
Kotlin
import ai.desertant.redact.Redact

Redact(context).use { redact ->
    val result = redact.redaction("Email Anna Kovács at anna@example.hu.")
    println(result.redactedText)                 // Email [GIVEN_NAME_1] [SURNAME_1] at [EMAIL_1].
    val restored = result.restore(llmReply)
}
TypeScript
import { Redact } from "@desert-ant-labs/redact";     // browser
// import { Redact } from "@desert-ant-labs/redact/native"; // server-side Node

const redact = await Redact.load();
const result = await redact.redaction("Email Anna Kovács at anna@example.hu.");
console.log(result.redactedText);   // Email [GIVEN_NAME_1] [SURNAME_1] at [EMAIL_1].
const restored = result.restore(llmReply);
redact.dispose();

Loading the model

The weights are fetched from the Hub on first use and cached. See model downloads and caching to prefetch them or to ship them with your app.

Runtime settings

Recommended defaults: min_score = 0.6, max_length = 256, stride = 64.

Taxonomy (20 public labels, plus ORG)

GIVEN_NAME, SURNAME, STREET_NAME, BUILDING_NUMBER, SECONDARY_ADDRESS, CITY, STATE, ZIP_CODE, EMAIL, PHONE, CREDIT_CARD, BANK_ACCOUNT, ROUTING_NUMBER, IP_ADDRESS, URL, GOVERNMENT_ID, PASSPORT, DRIVERS_LICENSE, TAX_ID, SSN.

ORG (organisation / company name) is detected but not redacted by default: a company is not a natural person. It exists so that Silverfin, Odoo or Visma Nova are recognised as organisations instead of being mislabelled as a SURNAME. Opt in by passing it explicitly in the SDK's labels option.

The deterministic layer additionally emits IMEI (device identifier), a deterministic-only label outside the neural head.

How it compares

Every system below was scored by the same harness on the same rows, each at its own operating point, so the comparison measures the models rather than the plumbing.

SystemRecallPrecisionSizeParams
redact88.899.611.6 MB23M
GLiNER-PII91.190.42.3 GB570M
Rampart61.497.214.7 MB18.5M
OpenAI privacy filter60.293.53 GB1.5B

Recall is the share of personal data fully masked (leak-safe), macro-averaged over WikiANN, MultiNERD and a format-valid structured-PII set across 24 EU languages. Precision is the share of masked spans that were really personal data, on the structured set. Size is the Apple build; the Android and web build is 24.5 MB.

Not masking ordinary words matters as much as catching real ones, because a false positive corrupts the text a downstream model receives. On an 11,528-row negative set across 27 languages, built to provoke exactly that (sentence-initial capitals, ALL-CAPS input, month and weekday names, UI vocabulary, bare numbers, company names), 94.1% of rows come back untouched.

AWS Comprehend, English only

Comprehend is the other service teams weigh, and it is not in the table above because its PII API only accepts English, and every other language code is refused outright, so there is no way to run it on the other 23. Scored on the same English rows:

SystemNames (WikiANN)Names (MultiNERD)StructuredEnglish composite
redact69.594.995.086.5
AWS Comprehend84.398.591.991.6

Leak-safe recall; precision is the same for both (99.8 against 100.0). On English names Comprehend is ahead of us. It also runs in the cloud, bills per call, and covers one of the 27 languages listed below.

Languages

27 languages: every official EU language, plus 3 more. Latin, Greek and Cyrillic scripts.

The 24 EU languages

CodeLanguage
bgBulgarian
hrCroatian
csCzech
daDanish
nlDutch
enEnglish
etEstonian
fiFinnish
frFrench
deGerman
elGreek
huHungarian
gaIrish
itItalian
lvLatvian
ltLithuanian
mtMaltese
plPolish
ptPortuguese
roRomanian
skSlovak
slSlovenian
esSpanish
svSwedish

Beyond the EU

CodeLanguage
nbNorwegian Bokmål
nnNorwegian Nynorsk
isIcelandic

Coverage is not uniform: the largest EU languages are the strongest, and Maltese and Irish are the weakest of the 24. The per-language detection numbers are in the benchmark data.