How does the Siberson Veriket Data Classification engine detect sensitive content?
Siberson Veriket Data Classification employs a multi-layered detection architecture that combines AI-powered automated classification with configurable rule-based detection to deliver comprehensive, accurate coverage across diverse content types and languages:
| Detection Method | How It Works | Best Applied To |
|---|---|---|
| AI-Powered Automated Classification | Machine learning models analyze document content, context, structure, and metadata to automatically assign the most appropriate sensitivity classification — without requiring predefined patterns or keyword lists for every document type | Unstructured documents, mixed-content files, legacy data estates where manual rules cannot scale |
| Keyword Matching | Predefined and custom keyword lists trigger classification when specified terms, phrases, or combinations are detected in document content or metadata | Domain-specific terminology, regulatory identifiers, internal project codenames |
| Regular Expressions (Regex) | Pattern matching engine detects structured sensitive data — credit card numbers, national identity numbers, IBAN codes, passport numbers — based on format patterns rather than exact values | Structured PII, financial data, reference numbers, regulated data formats |
| Contextual Analysis | Evaluates the surrounding context of detected patterns — validating that a detected number sequence is structurally consistent with a credit card number and appears in a financial context — reducing false positive classification | High-precision classification in document types where patterns may appear innocuously |
| OCR (Optical Character Recognition) | Extracts machine-readable text from image-based content — scanned documents, screenshots, image-embedded PDFs — enabling the full detection stack to be applied to content that is not natively machine-readable | Scanned contracts, photographed identity documents, legacy paper-based records |
| Document Fingerprinting | Creates a structural fingerprint of known sensitive document templates or content sets, enabling detection of documents derived from or similar to protected originals — even when specific patterns have been modified | Template-based documents, proprietary forms, standard agreements |
✅ Real-Time Guidance
Siberson Veriket Data Classification guides users in real time while they work on documents — automatically detecting personal data and special category data, and displaying the active policy and triggered keywords directly within the user's working environment.
Last updated: 2026-04-12