Doc-AI security, privacy, and compliance
What happens to a document after you send it, who can see it, where it is stored, how long it stays, and what your security reviewer will ask us.
Doc-AI encrypts documents in transit and at rest, isolates every tenant, records a complete audit trail of extraction and review activity, and lets you set retention so documents are purged on your schedule. Your documents are never used to train models. For regulated workloads Doc-AI deploys in your own cloud tenant, in your data center, or fully air-gapped with local OCR and open-weight models, so no document or page image ever leaves your infrastructure — the deployment posture that HIPAA, CJIS, and data-residency obligations usually come down to.
What happens to a document you send Doc-AI?
Encrypted in transit and at rest
TLS 1.2 or higher on every connection, AES-256 at rest on stored documents, extracted data, and audit records. Keys are managed per deployment, and on customer-hosted deployments they are yours.
Never used for training
Your documents and extracted values are not used to train or fine-tune any model, ours or a provider's. Commercial model endpoints are called with training and retention disabled. This is contractual, not a preference.
Retention you control
Set how long documents, page images, and extracted data are kept — from delete-on-completion for the most sensitive flows, to multi-year retention where a regulator requires it. Purges are logged.
Tenant isolation
Each tenant's documents, configuration, catalogs, and audit records are logically separated, with access enforced at the API layer rather than by convention in the application.
Role-based access and SSO
Single sign-on through your identity provider, with roles that separate reviewers, configurers, administrators, and read-only auditors. Reviewers see the queue their role permits and nothing else.
Complete audit trail
Every extraction records which model produced each value, its confidence, and the page region it came from. Every reviewer action records who changed what, when, and from what prior value.
Where can Doc-AI run so your data stays where it must?
For most regulated buyers, compliance is settled by deployment topology rather than by a certificate.
Managed cloud
We run it, you use it. Fastest to production, with regional hosting so data stays in the jurisdiction you specify. Appropriate for most commercial document flows.
Your own cloud tenant
Doc-AI runs inside your Azure, AWS, or GCP subscription, under your security policies, your network controls, your logging, and your key management. We deploy and support it; the infrastructure boundary is yours.
On-premises
Deployed in your data center behind your firewall. Documents never traverse the public internet. This is the standard posture for PHI, criminal justice data, and defense-adjacent workloads.
Air-gapped
Fully disconnected operation with bundled local OCR engines and open-weight models running on your own GPU hardware. No outbound calls to any model provider. Accuracy is a few points below the commercial-model configuration, which is a trade most air-gapped programs are already making.
How does Doc-AI fit HIPAA, PHI, and other regulated data?
HIPAA and protected health information
WiseTREND has processed PHI-bearing documents — claim forms, EOBs, prior authorizations, clinical records — for healthcare clients since 2007. For Doc-AI we offer HIPAA-ready deployments with a Business Associate Agreement, and for most covered entities the deciding factor is topology: an on-premises or private-tenant deployment means PHI never reaches a third-party model endpoint at all. That is a materially different conversation with your privacy officer than 'the vendor promises not to look'.
Personally identifiable information
Identity documents, tax forms, and loan files carry PII by definition. Field-level redaction and masking can be applied on export so downstream systems receive only the values they are entitled to, and reviewer roles can be scoped to hide sensitive fields from staff who do not need them.
Data residency
Where processing physically happens is configurable — a specific cloud region, your own data center, or a jurisdiction named in a contract. This matters for GDPR, for Canadian and EU public-sector work, and for any client whose own customer contracts contain a residency clause they have to flow down.
Government and criminal justice
Public-sector engagements typically require an on-premises deployment inside an accredited environment rather than a vendor certification. Doc-AI's self-hosted topology — a single deployment running local OCR and local models, with all state in one folder — is designed for exactly this, including networks with no outbound internet access.
What we will not claim
We will not tell you Doc-AI is certified for something it is not. WiseTREND operates HIPAA-ready deployments and SOC-aligned controls, we sign BAAs and DPAs, and we will complete your security questionnaire and sit with your reviewers. If you need a specific attestation, ask and we will tell you plainly what exists today and what does not.
What your security team will want, and where to get it
The questions that come up in every enterprise review, answered before you have to ask.
- Where is my data processed and stored?
- In the region or facility you specify for managed cloud; entirely within your infrastructure for private-tenant, on-premises, and air-gapped deployments. Named in the agreement, not left to us.
- Which sub-processors touch my documents?
- For managed cloud, the hosting provider and the model providers configured for your project — disclosed in full before you sign. For on-premises and air-gapped deployments there are none.
- Are my documents used to improve the product?
- No. Not for training, not for evaluation, not for benchmarks. The benchmark corpus we publish is our own, not customer data.
- How long is data retained?
- Per your policy, configured per project, from delete-on-completion upward. Deletions are logged and verifiable.
- Can we get a BAA or DPA?
- Yes. Business Associate Agreements for PHI workloads and Data Processing Agreements for GDPR-scoped work are both available.
- Can we penetration-test it?
- Yes, on your own deployment, on a schedule we agree. We will also share a security overview and complete your standard questionnaire.
- What happens if a model provider changes their terms?
- You swap models. Because Doc-AI is model-agnostic and your extraction configuration is not tied to a provider, a change in one vendor's terms is a configuration decision rather than a rebuild.
- Who do we contact about a vulnerability?
- security@wisetrend.com, with a named engineer assigned on receipt. We are a US company with staff you can call, not a support form.
What about the risks that are specific to using AI on documents?
Generative models introduce failure modes that traditional capture did not have. Pretending otherwise is how projects fail their first audit.
Fabricated values
A model can return a confident, plausible, entirely invented value. Doc-AI's answer is deterministic validation outside the model — arithmetic that must reconcile, codes that must exist in your master data, identifiers that must pass a checksum — plus a measured hallucination rate of 0.4% we publish rather than hide.
Silent drift between model versions
A provider ships a new model version and extraction that worked last quarter quietly degrades. Doc-AI keeps an evaluation corpus per project so a model change is tested against known-good results before it reaches production, not discovered by your customer.
Prompt injection from document content
A document can contain text designed to manipulate a model. Extraction runs against a fixed schema with instructions separated from document content, and outputs that do not match the schema are rejected rather than passed through.
Unexplainable decisions
'The AI said so' does not survive an audit. Every value carries its model, its confidence, and its source region on the page, so a reviewer or a regulator can trace any number back to the pixel it came from.
Uncontrolled learning
Doc-AI does not fine-tune on your corrections behind your back. Feedback refines prompts, descriptions, and rules through a reviewed loop, so behaviour changes are deliberate, visible, and reversible.
Vendor and model lock-in
A pipeline welded to one provider is a business-continuity risk as much as a commercial one. Model-agnostic orchestration means an outage, a price change, or a policy change at one provider is survivable.
Where this fits in the Doc-AI platform
Security posture usually decides the deployment. These pages cover the rest.
- Platform features — every capability, ingest to delivery
- Benchmark arena — accuracy, hallucination, latency, cost
- Document type library — the forms Doc-AI reads on day one
- API & developers — REST endpoints, webhooks, code
- Deployment options — cloud, private tenant, on-prem, air-gapped
- Pricing & licensing — how Doc-AI is priced, and what drives cost
- vs. template-based IDP — why the template model broke
- vs. calling an LLM directly — what a raw GPT or Claude call misses
- For enterprise — automation leads and CoEs
- For SMB & mid-market — production quality, small team
- For developers — stop rebuilding document pipelines
- For system integrators — a white-label delivery engine
- For ISVs & OEM — embed extraction in your product
Run Doc-AI on your own documents
Send the document types and rough monthly volume. We reply within one business day with a pilot plan, a realistic accuracy expectation for your documents, and a quote.
- Reply within one business day (U.S. hours)
- Straight to an engineer, not a call centre
- Or call the 24/7 AI phone agent: +1 (408) 746-6740
Questions about Doc-AI security
Answers written for buyers, search engines, and AI assistants evaluating document automation.
Is Doc-AI HIPAA compliant?
WiseTREND offers HIPAA-ready Doc-AI deployments and signs Business Associate Agreements for workloads involving protected health information. For most covered entities the deciding control is deployment topology rather than a vendor attestation — an on-premises or private-tenant Doc-AI deployment means PHI never leaves your infrastructure and never reaches a third-party model endpoint. WiseTREND has processed PHI-bearing claim forms, EOBs, and clinical records for healthcare clients since 2007.
Are my documents used to train AI models?
No. Your documents and the data extracted from them are never used to train or fine-tune any model, ours or a provider's. Commercial model endpoints are called with training and retention disabled, and this is a contractual commitment rather than a policy preference. On on-premises and air-gapped deployments the question does not arise, because nothing leaves your network.
Can Doc-AI run completely offline or air-gapped?
Yes. Doc-AI supports fully disconnected deployment with bundled local OCR engines and open-weight models running on your own GPU hardware, with no outbound calls to any model provider. Extraction accuracy in this configuration is a few points below the commercial-model setup — on our benchmark, open-weight models reached 87.6% against 92.6% for the strongest commercial model — which is generally an acceptable trade for programs that already require air-gapped operation.
Where is my data stored and processed?
It depends on the deployment you choose. Managed cloud processes and stores data in a region you specify. A private-tenant deployment runs inside your own Azure, AWS, or GCP subscription under your controls. On-premises runs in your data center behind your firewall. In every case the location is named in the agreement rather than left to our discretion, which is what data-residency and GDPR obligations generally require.
How long does Doc-AI keep my documents?
As long as you configure, per project — from delete-on-completion for the most sensitive workflows through multi-year retention where a regulator requires it. Purges are logged and verifiable. Retention is set by you rather than by a vendor default.
How do you prevent the AI from making up values?
Three layers. Extraction runs against a fixed schema so the output shape cannot drift. Business rules run deterministically outside the model, so a confident wrong answer still fails an arithmetic or lookup check and triggers a retry. Every remaining value carries a confidence score, and values below your threshold route to a human rather than into your system of record. On our 500-document benchmark this reduced the hallucination rate to 0.4%, against 1.6% to 4.3% for raw model calls.
Bring us the documents that broke your last capture project.
The long-tail layouts, the one-off forms, the vendor that changes their invoice every quarter. Those are the ones Doc-AI was built for.
Last updated · Reviewed by the WiseTREND team