Siberson
Partnership Contact Request a Demo

How does the Siberson Veriket Data Classification engine detect sensitive content?

Siberson Veriket Data Classification employs a multi-layered detection architecture that combines AI-powered automated classification with configurable rule-based detection to deliver comprehensive, accurate coverage across diverse content types and languages:

Detection Method How It Works Best Applied To
AI-Powered Automated Classification Machine learning models analyze document content, context, structure, and metadata to automatically assign the most appropriate sensitivity classification — without requiring predefined patterns or keyword lists for every document type Unstructured documents, mixed-content files, legacy data estates where manual rules cannot scale
Keyword Matching Predefined and custom keyword lists trigger classification when specified terms, phrases, or combinations are detected in document content or metadata Domain-specific terminology, regulatory identifiers, internal project codenames
Regular Expressions (Regex) Pattern matching engine detects structured sensitive data — credit card numbers, national identity numbers, IBAN codes, passport numbers — based on format patterns rather than exact values Structured PII, financial data, reference numbers, regulated data formats
Contextual Analysis Evaluates the surrounding context of detected patterns — validating that a detected number sequence is structurally consistent with a credit card number and appears in a financial context — reducing false positive classification High-precision classification in document types where patterns may appear innocuously
OCR (Optical Character Recognition) Extracts machine-readable text from image-based content — scanned documents, screenshots, image-embedded PDFs — enabling the full detection stack to be applied to content that is not natively machine-readable Scanned contracts, photographed identity documents, legacy paper-based records
Document Fingerprinting Creates a structural fingerprint of known sensitive document templates or content sets, enabling detection of documents derived from or similar to protected originals — even when specific patterns have been modified Template-based documents, proprietary forms, standard agreements

✅ Real-Time Guidance

Siberson Veriket Data Classification guides users in real time while they work on documents — automatically detecting personal data and special category data, and displaying the active policy and triggered keywords directly within the user's working environment.

Last updated: 2026-04-12