Security
Content analysis: files and email
Last updated · October 7, 2026
Two optional modules read content rather than metadata.
Content analysis: files and email
- What triggers a file scan. The file exposure module runs only if you granted the Drive or SharePoint scopes. It then runs nightly, and on demand from the application.
- What is read. We download the file and parse it in memory: Office documents, PDFs and plain text, capped at 512 KB of extracted text and 8 MB per document. We open the file; we do not keep it.
- What is written. A sensitivity level and a count per detected category, and nothing else. Never the matched value, never an excerpt. A file we could not read is recorded as unreadable rather than clean, so an unopened document is never reported as safe.
- Who classifies. Self-hosted deterministic detectors: Luhn with an issuer table for card numbers, mod-97 for IBAN and French social security numbers, and a vendored gitleaks ruleset for credentials. No third-party classification API and no language model. Your file content is read by code running on a machine in France and is sent nowhere.
- Keyword triage. An administrator can search up to 30 keywords across at most 200 files. Matches are returned to that administrator's browser and are never stored.
- Email discovery is off by default. It is enabled per workspace, and it additionally requires mailbox access to be granted out of band in your own admin console, so it cannot be switched on by accident from our side.
- What email discovery reads. Three headers from Gmail, three fields from Microsoft Graph, narrowed to signup-shaped subjects. No call we make can return a body. The subject line is classified in memory and discarded, and only the vendor's domain is kept.
- For prompt protection the extension reports the AI tool's hostname even when that host is absent from our SaaS catalogue, because the catalogue gate would otherwise drop the event. The content script runs only on the three AI origins declared in its manifest.