Know what leaves. Block what shouldn't.
DLP detects entities — PII, credentials, and regulated data — across every channel through a 3-tier DLP pipeline at under 2ms p99 latency. Content Categories classify topics — classifying prompts by topic across 8 domains and 26 sub-categories. Together, they form two complementary detection surfaces in a single inspection pipeline, enabling governance rules that fire on both what data is in the request and what the request is about.
Every request and response is inspected in real time — flagged events appear in the detection log with severity, matched pattern, and the enforcement action taken.
Capabilities
80+ Pattern Detectors, with checksum validation where applicable
Detect credit card numbers, IBANs, Social Security numbers, and government IDs across the regex layer. Where the entity type carries a check digit or structural format spec — IBAN (MOD-97), ABA routing, NPI (Luhn), DEA, ITIN, EIN, Canadian SIN, IMEI, SWIFT/BIC, and similar regulated identifiers — a validator runs after the regex match to confirm the data is structurally real, eliminating false positives on invoice numbers and arbitrary digit strings. Patterns that lack a published checksum spec rely on context keywords and downstream tier-2/tier-3 confirmation. The credential detection layer covers API keys, database connection strings, OAuth tokens, and cloud provider secrets across every major platform.
ML-Driven Named Entity Recognition with Contextual Validation
ML-driven named entity recognition runs inference to catch names, addresses, medical record numbers, financial account identifiers, and more. An ML-driven contextual validation model (under 2ms p99 latency) evaluates each match against surrounding text to reduce false positives. The result: fewer false alarms and faster review cycles for security teams.
12 Compliance Frameworks, Pre-Mapped
Enforce PCI-DSS, HIPAA, GDPR, GLBA, SOX, CCPA, BSA/AML, SEC Reg FD, FERPA, EU AI Act, NIST AI RMF, ISO/IEC 42001 with pre-built compliance bundles. Each bundle activates the correct detectors across all 3 DLP layers and maps enforcement actions — block, redact, or log — to the specific data types each regulation covers. Activate a framework and Arbitex handles the mapping.
Content Category Detection
Topic-based classification as a policy engine condition. Content Categories classify prompts across 8 content domains and 26 sub-categories — legal, financial, medical, security, competitive, and more. The detection surface extends beyond named topics to include code detection, language identification, and prompt safety classifiers (jailbreak, prompt leak, off-topic) — all keyword and rule-based. All classifications fire as content_category conditions in the same policy rule chain as DLP entity findings.
How Detection Quality Is Built
Detection quality does not come from any single detector. It comes from layering — pattern rules, ML entity recognition, and contextual validation each inspecting the same content, so that a match one stage surfaces, another can confirm or reject.
Payment card numbers, IBANs, and other identifiers carrying a check digit are arithmetically validated, not just pattern-matched. When the format is unambiguous, a validated match is close to unambiguous too.
ML-driven contextual validation analyzes surrounding text to distinguish genuine sensitive data from false positives. An identifier in a code fixture is treated differently from the same identifier in a patient record.
Detection quality is evaluated per entity type against a labeled corpus rather than as a single blended score, because an aggregate hides the weak spots that matter most when tuning detectors.
Per-entity evaluation is an internal engineering practice. Arbitex does not currently publish per-entity accuracy figures — the evaluation corpus is not yet large enough for those numbers to be meaningful.
How It Works
Define your compliance scope
Select which regulatory frameworks, data types, and content categories to enforce. Compliance bundles activate the correct detectors across regex, entity recognition, and contextual validation layers. Content Categories add topic classification — 8 domains and 26 sub-categories — as a separate configuration surface alongside entity detection. Custom rules extend coverage for organization-specific patterns.
Inspect on input and output
Requests pass through all 3 detection layers. Response inspection available on select plans. Regex detectors validate format and checksum. ML-driven entity recognition extracts entities from free text. A context-aware language model confirms or rejects each detection. Compliance bundles enforce framework-specific rules. Each match triggers a configurable action: block the request, redact the sensitive content, or log the detection and allow it through.
Manage from the DLP console
Build custom detection rules in the rule builder. Test patterns against sample payloads with the preview endpoint. Monitor detection trends and false positive rates from the analytics dashboard. Bulk import and export rules across environments.
Credential Intelligence
Detect known compromised credentials before they reach any AI model.
Pattern detection catches what looks like a credential. Credential Intelligence catches what is one — and tells you how dangerous it is.
When a user sends text through the AI gateway, Credential Intelligence checks it against a compromised credential dataset using secure comparison. No cleartext is ever stored or transmitted. The check runs in parallel with the DLP pipeline, adding negligible latency to each request.
What distinguishes this from generic credential detection is the frequency signal. A credential seen in millions of breaches carries a different risk profile than one seen once. Credential Intelligence assigns each detection to a risk level — Critical, High, Medium, or Low — so security teams can act on the highest-risk exposures immediately.
Monthly dataset refresh
Secure comparison against a curated compromised credential dataset. Frequency-weighted risk levels: Critical, High, Medium, Low.
Credential Intelligence
Checks every request against a compromised credential dataset. Frequency-weighted risk levels: Critical, High, Medium, Low.
Related Resources
Rules-based governance for every AI request
Content CategoriesTopic classification as a policy engine condition
Compliance FrameworksPre-built policy packs for regulatory requirements
Audit LogTamper-proof activity trail
HealthcareHIPAA compliance and PHI detection
Financial ServicesPCI-DSS and SOX compliance