Benchmarks

We measure instead of claiming. Every number on this page comes from a logged measurement run on the same German gold corpus, and every number carries its measurement date. If we haven't measured it, it isn't here.

Last updated: 2026-08-16 · Diese Seite auf Deutsch →

Free PII engines compared (standalone)

Four freely available detection systems and Saklam Bridge, each standalone on the same corpus with the same metric. We deliberately don't name commercial vendors here — we only compare what anyone can reproduce for free.

Systemdetected (of 489)scoremeasured on
Saklam Bridge v0.2.1748298.6 %2026-08-16
Saklam Bridge v0.2.1046695.3 %2026-07-17
OpenAI privacy filter39981.6 %2026-07-17
GLiNER2-PII (Fastino, threshold 0.3)34871.2 %2026-07-17
Microsoft Presidio + spaCy de_core_news_lg34570.6 %2026-07-17
Kiji (Dataiku)22245.4 %2026-07-29 *

* Follow-up measurement with identical corpus and identical metric. Presidio and Kiji each measured in their standard setup — both products allow custom recognizers / regex rules (at which point you own the pattern maintenance yourself).
Intermediate state of the same series: Bridge v0.2.12 scores 467/489 = 95.5 % on the same corpus (2026-07-22). The gain up to v0.2.17 comes from a deterministic additional layer for reference and case numbers, organization names and title chains — no new model, no tuning on the corpus.

In fairness: what the others do well

  • GLiNER2-PII delivered the field's best address result (98 % vs. Saklam's 94 %) and was strong on employee and customer IDs.
  • The OpenAI filter was strong on person names and secrets — in the first run it beat our then-current version on person names. Those findings went into v0.2.10 the same day (person names at 93 % afterwards).
  • Kiji (Apache 2.0) was flawless on email addresses (53/53) and strong on tax IDs (90 %) and IBANs (88 %) — its gaps stem from a label schema that doesn't know German domain identifiers (case numbers, courts, client numbers, registers).
  • The shared gap of the free engines is German structural reality: organizations, courts, case numbers, tax IDs, grouped IBANs.
  • Saklam has gaps too — 23 of the 489 values (as of 2026-07-17) were deliberately not tuned away on the corpus: context-free short IDs belong in a customer's custom configuration, and rare foreign name forms are a model limit we keep working on. No benchmark overfitting. As of 2026-08-16 there are 7 — all of them person names without a context anchor: bare surnames, user handles, rare foreign name forms.

Ongoing series — latest measurements

We continuously measure newly released models as the NER component inside the Saklam pipeline: same corpus, same regex layer, identical thresholds. Numbers within this series are comparable to each other — but not to the standalone table above: there, e.g. GLiNER2 runs without our pipeline (71.2 %), here as a component inside it (95.1 %). That's why every row carries its measurement date, and why we never mix numbers from different measurement states in one comparison row.

Model (as NER component in the Saklam pipeline)detected (of 489)scoremeasured on
Saklam Engine (production state) — reference46795.5 %2026-08-07
SauerkrautLM-LFM2.5-GLiNER (VAGO Solutions, 411M)46695.3 %2026-08-09
GLiNER2-PII (Fastino)46595.1 %2026-08-07
gliner-pii-edge v1.0 (Knowledgator, 86 MB fp16)39881.4 %2026-08-07
  • SauerkrautLM-LFM2.5-GLiNER (Aug 9): just below the reference, with a genuine strength — the model is format-sensitive on structured IDs (customer, employee and client numbers) and catches formats other candidates in the series miss.
  • GLiNER2-PII (Aug 7): close to the reference inside the pipeline; a few additional gaps on person names.
  • gliner-pii-edge (Aug 7): very small and roughly 7× faster — but with systematic gaps on person names. Speed is cheap, recall isn't.

Methodology

  • Corpus: German-language gold corpus, 489 annotated PII values across 48 synthetic, realistic documents — client letters, HR cases, invoices, support tickets, log excerpts. Fully synthetic: no real personal data.
  • Scoring: strictly value-based — an annotated value only counts as protected if it no longer appears in the masked output, not even partially as a substring. Over-masking doesn't count as an error; a partial leak counts as a leak.
  • Comparability: identical pipeline, identical thresholds and identical corpus within each measurement series. No tuning against the corpus (our "no benchmark overfitting" rule).
  • Saklam detection stack: 350+ validated PII patterns (EU-27 + EEA) combined with an NER model — the numbers above always measure the combination, never a single piece.

Why the corpus isn't public

The obvious question: "Show us the corpus." The honest answer: a gold corpus full of realistic German PII patterns — names, IBANs, case numbers, employee IDs inside real-looking documents — would itself be exactly the kind of artifact this engine protects against. Publishing it would also let every engine (ours included) optimize against the test set — at which point it measures nothing. Instead: the documents follow the pattern of the examples below (newly constructed, not copied from the corpus), and the live demos at the bottom let anyone measure with their own text.

Example 1 — client letter (law firm):

Before
Sehr geehrte Frau Kollegin Dr. Wenkstern, in der Sache unseres Mandanten
Halil Demirbaş (Az. 4 O 2231/26) übersenden wir den Vergleichsvorschlag.
Die Zahlung über 4.850 € geht auf das Konto DE21 3704 0044 0532 0130 87.
Rückfragen gern an h.demirbas@web.de.
After (masked)
Sehr geehrte Frau Kollegin Dr. [PER_9c41f2ab], in der Sache unseres Mandanten
[PER_5e17d803] (Az. [CAS_7d21e0a3]) übersenden wir den Vergleichsvorschlag.
Die Zahlung über 4.850 € geht auf das Konto [IBA_2f90c6d4].
Rückfragen gern an [EMA_b34a9e01].

Example 2 — HR case:

Before
Mitarbeiterin Annika Brendel hat zum 31.10. gekündigt. Bitte formuliere
ein wohlwollendes Zwischenzeugnis, Abteilung Vertrieb Süd,
Vorgesetzter: Timo Kaltenbach.
After (masked)
Mitarbeiterin [PER_41d8e07f] hat zum 31.10. gekündigt. Bitte formuliere
ein wohlwollendes Zwischenzeugnis, Abteilung Vertrieb Süd,
Vorgesetzter: [PER_c290ab54].

Example 3 — invoice:

Before
Rechnung an Brandt & Cie. Steuerberatung, z. Hd. Frau Melisa Yavuz,
USt-IdNr. DE314592807 — bitte prüfe die Positionen auf Plausibilität.
After (masked)
Rechnung an [ORG_88f1c3d2], z. Hd. Frau [PER_67ab04e9],
USt-IdNr. [VAT_d10c58fa] — bitte prüfe die Positionen auf Plausibilität.

All examples newly constructed in the style of the corpus; all names, IBANs and identifiers are invented, placeholder IDs illustrative. Note what does not get masked: amounts, deadlines, the actual matter — that's the payload the model needs to work with. What gets masked is whatever makes the request attributable to a person or company.

Measure it yourself

You don't have to take our word for it. The same detection runs in our live demos right in your browser — paste your own text and inspect the result line by line:

Or right inside the chat window: Saklam browser extension for ChatGPT & Co. →