Legal AI Governance

Legal AI Vendor Security Checklist

Evaluate legal AI vendors across data flows, model use, tenancy, access, encryption, logging, retention, testing, resilience, and contract evidence.

Direct answer

A legal AI vendor security checklist maps every data flow, model and provider dependency, training-use rule, tenant boundary, identity path, security control, retention and deletion process, incident response commitment, integration, test result, vulnerability process, continuity plan, and contract obligation. Require configuration-specific evidence and a remediation owner. Certifications can support diligence but do not guarantee security, confidentiality, legal privilege, performance, or compliance for a particular deployment.

Definitions

Legal AI service

A hosted, embedded, or managed system that applies artificial intelligence to legal information, workflows, documents, communications, decisions, or operational records.

Data flow

A documented movement of data into, through, between, and out of a service, including users, applications, storage, models, providers, subprocessors, regions, logs, backups, and support channels.

Model provider

The organization that develops, hosts, fine-tunes, serves, or otherwise supplies a model used by the legal AI service, whether directly or through another provider.

Subprocessor

A downstream service provider engaged to process customer data or support the service, including cloud, model, analytics, support, security, storage, and infrastructure providers.

Training use

A vendor-defined use of customer prompts, files, outputs, telemetry, feedback, or derived data to train, fine-tune, evaluate, improve, or otherwise develop a model or service.

Tenant boundary

The technical and administrative separation that prevents one customer, matter, workspace, identity, or authorization context from reaching another customer's data or operations.

Configuration-specific evidence

Evidence that applies to the purchased service, region, plan, feature set, tenant configuration, model route, integration, and contract rather than only to a broad corporate program.

Deletion evidence

A record showing what was deleted, from which systems and copies, under what request or lifecycle rule, when the action completed, what exceptions remain, and who verified the result.

Adversarial testing

Structured testing intended to identify security, privacy, reliability, misuse, prompt-injection, data-exposure, or unsafe-behavior weaknesses in an AI system and its surrounding controls.

Certification or attestation

An external or internal statement about a defined scope, period, criteria, or control program; it is evidence to evaluate, not a guarantee that a specific legal AI deployment is secure or suitable.

Field definitions

Service scope and data flows

use_case_scope
Approved legal workflow, users, entities, matters, jurisdictions, data classes, outputs, decisions, prohibited uses, and accountable owner.
Type: Structured risk record
Requiredness: Always required
Validation: Reject an approval that does not name the workflow, data classes, human review, and prohibited actions.
Owner: Legal technology owner
data_flow_register
Every input, output, retrieval, embedding, log, backup, support, telemetry, export, deletion, and provider path with purpose and location.
Type: Linked flow records
Requiredness: Required before production approval
Validation: Each flow must identify direction, data type, recipient, region, retention, access, encryption, and optional or required status.
Owner: Security architect
provider_dependency_register
Models, model providers, cloud regions, subprocessors, storage, observability, support, and fallback dependencies for each feature.
Type: Provider inventory
Requiredness: Required before production approval
Validation: Match provider and model records to the current contract, feature, region, and customer configuration.
Owner: Vendor manager
training_use_rule
Allowed and prohibited uses of prompts, files, outputs, feedback, telemetry, identifiers, and derived data for training or improvement.
Type: Policy and contract control
Requiredness: Always required
Validation: Record defaults, feature exceptions, opt-out behavior, human review, retention, and the controlling contract language.
Owner: Privacy and legal owner

Isolation, identity, and cryptography

tenant_boundary
Technical and administrative controls separating tenants, matters, workspaces, identities, indexes, caches, logs, backups, and support access.
Type: Architecture and test evidence
Requiredness: Required before production approval
Validation: Require negative tests for search, retrieval, exports, APIs, notifications, integrations, support, and administrator paths.
Owner: Security architect
identity_authorization_model
SSO, MFA, lifecycle automation, roles, attributes, matter restrictions, service accounts, API keys, sessions, support access, and privileged actions.
Type: Control matrix
Requiredness: Required before production approval
Validation: Show effective authorization for users and background jobs and define how access is revoked and reviewed.
Owner: Identity owner
encryption_key_control
Encryption coverage, protocols, algorithms, key ownership, rotation, access, recovery, revocation, and customer-managed-key options.
Type: Security control record
Requiredness: Required before production approval
Validation: Identify exclusions for prompts, outputs, embeddings, temporary data, queues, logs, backups, exports, and support channels.
Owner: Security architect

Evidence, lifecycle, and response

audit_event_catalog
Logged actors, tenants, resources, prompts, retrieval, outputs, exports, configurations, support access, deletions, and security events.
Type: Event catalog
Requiredness: Required before production approval
Validation: Capture timestamp, correlation, version, region, outcome, export method, retention, tamper protection, and customer access.
Owner: Security operations
retention_deletion_schedule
Retention and deletion rules for live data, inputs, outputs, embeddings, indexes, logs, backups, support, abuse monitoring, and derived data.
Type: Lifecycle matrix
Requiredness: Always required
Validation: Test deletion, backup expiry, subprocessor propagation, restore behavior, legal holds, exceptions, and verification evidence.
Owner: Records and privacy owner
incident_vulnerability_terms
Incident definitions, notification, cooperation, evidence, vulnerability disclosure, patching, dependency monitoring, and remediation duties.
Type: Contract and response record
Requiredness: Required before production approval
Validation: Map duties to contacts, severity, clock start, channels, customer obligations, and unresolved exceptions.
Owner: Security and legal owners
assurance_evidence_register
Configuration-specific certifications, attestations, test reports, architecture evidence, exceptions, remediation, and review dates.
Type: Evidence register
Requiredness: Required before production approval
Validation: Record scope, period, system, region, exclusions, control owner, finding status, and whether evidence covers the purchased configuration.
Owner: Third-party risk owner

Continuity and exit

continuity_exit_plan
Availability, recovery, fallback, degraded mode, support, portability, export, deletion, transition, and termination assistance controls.
Type: Resilience and exit plan
Requiredness: Required before production approval
Validation: Test retrieval of source records, outputs, citations, audit evidence, configurations, permissions, and deletion attestations.
Owner: Service owner

Controlled vocabulary guidance

data_class
Examples: Public, Internal, Confidential, Client-confidential, Privileged or restricted, Personal data, Sensitive personal data, Security-sensitive
Governance: Apply the organization and client-approved classification before enabling a feature. Do not infer a lower classification because a vendor calls a field metadata or telemetry.
flow_status
Examples: Proposed, Verified, Approved, Exception, Blocked, Retired
Governance: A flow is production-approved only when its recipient, purpose, region, retention, access, security controls, and contract basis are verified.
training_use
Examples: Prohibited, Contractually prohibited, Opt-out available, Opt-in only, Allowed for defined purpose, Unknown
Governance: Use Unknown until the vendor answers data types, feature exceptions, human review, derived data, retention, and controlling terms.
tenant_test_result
Examples: Pass, Pass with finding, Fail, Not tested, Not applicable
Governance: Record the test population, configuration, feature, model route, date, evidence, and remediation. A Pass is not a guarantee outside the tested boundary.
evidence_status
Examples: Current and configuration-specific, Current but scope-limited, Expired, Pending validation, Remediation open, Unavailable
Governance: Record scope, period, exclusions, control mapping, reviewer, and open findings for every attestation, certification, test, or vendor answer.
incident_severity
Examples: Critical, High, Medium, Low, Unknown
Governance: Use organization-designed severity anchors based on data exposure, tenant scope, legal or operational impact, ongoing risk, evidence integrity, and urgency. Severity is not a legal conclusion.
lifecycle_state
Examples: Active, Archived, Deletion requested, Deletion verified, Held, Restored, Terminated
Governance: Do not mark deletion verified without evidence from the relevant live, backup, index, log, support, and subprocessor paths or a documented exception.

Practical workflow

  1. Define the use case and risk boundary

    Name the legal workflow, users, entities, clients, matters, jurisdictions, data classes, output decisions, human reviewers, prohibited uses, and business consequence of failure. Separate assistive drafting or retrieval from actions that can disclose information, alter records, communicate externally, set deadlines, approve transactions, or influence a legal determination. Record the intended model, features, integrations, environment, and accountable owner.

  2. Draw the complete data-flow map

    Trace prompts, files, metadata, embeddings, retrieved passages, outputs, feedback, telemetry, support tickets, abuse signals, audit logs, backups, exports, caches, and deletion requests from the customer boundary through the vendor, model providers, cloud regions, subprocessors, administrators, and return paths. Record data type, purpose, direction, protocol, storage location, retention, access role, encryption state, and whether the flow is optional or required. Ask the vendor to validate the map for the exact plan and configuration.

  3. Identify every model and provider dependency

    List each foundation model, hosted model, fine-tuned model, embedding model, reranker, classifier, OCR engine, speech or translation service, vector store, cloud platform, observability service, support tool, and subprocessors used by each feature. Capture provider legal name, service role, version or model family, region, change notice, fallback route, data-use terms, access boundary, and failure behavior. Require a process for notifying customers before material provider or model changes.

  4. Set the training and improvement rule

    Obtain an explicit answer for whether customer prompts, files, outputs, feedback, usage metadata, identifiers, support content, or derived representations are used for training, fine-tuning, evaluation, abuse detection, product improvement, or human review. Confirm default and opt-in settings, account-level controls, feature exceptions, de-identification claims, retention of opt-out data, and the contract language that prevails. Do not accept a marketing statement such as "not used for training" without defining every data type and processing purpose.

  5. Verify tenant and matter isolation

    Ask how the service separates tenants, workspaces, matters, clients, users, administrator roles, retrieval indexes, caches, embeddings, logs, backups, exports, and support tooling. Request architecture diagrams, authorization rules, negative test results, cross-tenant test evidence, incident history, and controls for shared infrastructure. Test that search, previews, citations, APIs, notifications, bulk exports, integrations, and support access cannot cross an approved matter or tenant boundary.

  6. Review regions, residency, and transfer paths

    Record processing and storage regions for live data, temporary files, logs, backups, disaster-recovery copies, model services, support access, telemetry, and subprocessors. Ask whether routing can change by feature, incident, load, or provider fallback. Confirm the vendor can identify where a specific customer record was processed and can support the organization's privacy, client, contractual, records, and cross-border review. Residency is a deployment and contract question, not an inference from the vendor headquarters.

  7. Inspect encryption and key management

    Document encryption in transit, at rest, backups, queues, temporary storage, exports, and administrative channels; protocol and algorithm choices; key ownership; key generation, rotation, access, separation, recovery, and revocation; and the treatment of model inputs, outputs, embeddings, and logs. Ask whether customer-managed keys, tenant-specific keys, bring-your-own-key, or regional key controls are supported and what data or features are excluded. Require evidence for the purchased service, not only a corporate policy.

  8. Check identity, authorization, and privileged access

    Verify SSO, federation, MFA, SCIM or lifecycle automation, role-based and attribute-based authorization, matter or workspace restrictions, service accounts, API keys, session controls, support impersonation, break-glass access, administrator separation, least privilege, and joiner-mover-leaver handling. Ask how the system evaluates authorization for retrieval, prompts, outputs, exports, citations, integrations, and background jobs. Require attributable access and review evidence for vendor personnel and privileged operations.

  9. Specify logging and audit evidence

    Define which events are logged for sign-in, authorization, prompt receipt, source retrieval, model route, output, citation, export, integration, configuration, administrator action, support access, deletion, policy change, and security event. Ask for timestamps, actor and tenant IDs, correlation IDs, model or feature version, region, outcome, tamper protection, customer access, export format, retention, and clock handling. Confirm that logging does not silently create a new sensitive-data copy outside the approved boundary.

  10. Test retention, deletion, and legal holds

    Map default and configurable retention for prompts, files, outputs, embeddings, indexes, feedback, logs, backups, snapshots, support records, abuse-monitoring data, and derived data. Define deletion triggers, user and administrator controls, API behavior, backup expiration, subprocessor propagation, verification evidence, exceptions for security or legal obligations, and handling of legal holds. Test a representative record through active, archived, deleted, restored, and backup-expiry states and retain the result.

  11. Evaluate incident and vulnerability response

    Obtain the vendor's incident definitions, detection and escalation process, customer notification commitments, contact channels, evidence preservation, containment, root-cause reporting, remediation tracking, vulnerability disclosure process, patch timelines, dependency monitoring, and cooperation with customer investigation. Ask how the vendor handles prompt injection, data exfiltration, model abuse, credential compromise, cross-tenant exposure, poisoned retrieval content, unsafe integrations, and provider incidents. Contractual notice periods should be confirmed against the organization's obligations and response needs.

  12. Review testing and model assurance

    Request security testing scope and dates, penetration tests, vulnerability scans, code review, dependency and supply-chain controls, red-team or adversarial testing, privacy testing, authorization tests, data-leakage tests, prompt-injection tests, retrieval poisoning tests, abuse controls, availability tests, and regression evidence. Require the test population, version, configuration, exclusions, severity method, unresolved findings, retest results, and independent reviewer. For legal workflows, also test citation support, source boundaries, hallucination handling, human review, and output labeling.

  13. Assess integrations and downstream exposure

    Inventory browsers, email, document stores, case and contract systems, identity providers, APIs, webhooks, plugins, connectors, exports, search indexes, storage targets, and automation actions. For each integration, document permissions, scopes, token storage, data minimization, synchronization direction, cache and failure behavior, retry and replay risk, approval gates, tenant mapping, logging, revocation, and downstream retention. Disable actions that can send, change, delete, or share records unless the authorization, preview, confirmation, and audit path is explicit.

  14. Check business continuity and exit

    Review availability commitments, regions and redundancy, recovery time and recovery point objectives, backup protection, dependency failure modes, model or provider fallback, degraded mode, support coverage, disaster exercises, customer communications, portability, export format, deletion on termination, transition assistance, and orderly service shutdown. Test that the organization can retrieve source records, prompts, outputs, citations, audit evidence, configurations, permissions, and deletion attestations without relying on a vendor dashboard that may no longer be available.

  15. Convert findings into contract controls

    Attach the approved data-flow map, service and model scope, training-use rule, region list, subprocessor process, security controls, access model, logging, retention and deletion schedule, incident and vulnerability duties, testing evidence, continuity terms, audit rights, cooperation obligations, confidentiality, intellectual-property treatment, data return, deletion certification, change notice, suspension rights, and termination assistance to the procurement record. Assign each requirement an owner, evidence source, review date, exception, and acceptance decision.

Comparison

Review areaControlled diligenceWeak diligence
Data useThe customer maps prompts, files, outputs, embeddings, telemetry, support, logs, backups, providers, and deletion paths for each feature.A privacy page or sales statement is treated as the complete data-flow map.
TrainingThe contract defines each data type, default, feature exception, opt-out, human review, derived data, and retention rule.The vendor says data is not used to train models without defining improvement, abuse, feedback, or downstream-provider uses.
Tenant isolationArchitecture, authorization, negative tests, retrieval, exports, APIs, integrations, support, and logs are reviewed for the purchased boundary.A shared-cloud statement or certification is treated as proof that every matter and feature is isolated.
AssuranceCertifications and attestations are checked for scope, period, exclusions, findings, configuration, region, and contract relevance.A badge or report is treated as a guarantee of confidentiality, privilege, security, performance, or legal compliance.
DeletionLifecycle rules cover active data, indexes, embeddings, logs, backups, support records, derived data, holds, exceptions, and verification.A delete button is treated as proof that every copy and downstream provider has removed the record.
TestingTesting includes authorization, leakage, prompt injection, retrieval poisoning, integrations, model changes, reliability, and human review with evidence.A generic penetration test or benchmark is treated as proof that legal workflows are safe.
ContractThe approved data flows, provider list, controls, response duties, audit evidence, continuity, exit, and change notices are contract-linked.Security requirements remain in a questionnaire and are not tied to the service configuration or remedies.

Limitations and exceptions

  • This is an organization-designed procurement and third-party-risk checklist, not a universal security standard, legal opinion, privilege determination, privacy impact assessment, or certification of any vendor.
  • A certification, attestation, audit report, penetration test, encryption statement, or security questionnaire describes a defined scope, period, method, and evidence set. It does not guarantee the security, confidentiality, availability, accuracy, legal suitability, privilege, or compliance of a particular deployment.
  • Vendor answers can be incomplete, stale, feature-specific, plan-specific, or inconsistent with the actual tenant configuration. Validate material claims with architecture, testing, contract language, operational evidence, and customer-controlled settings.
  • No checklist proves that a model will avoid hallucinations, prompt injection, data leakage, unsafe outputs, biased behavior, unauthorized retrieval, or every future vulnerability. Use human review, restricted actions, monitoring, testing, incident response, and change control.
  • Residency, deletion, retention, confidentiality, privilege, records, and notification requirements depend on the data, entity, client, contract, jurisdiction, system, and facts. Qualified legal, privacy, security, records, and business owners must make the applicable decisions.
  • Provider, model, subprocessor, feature, region, and integration changes can alter risk after approval. Reassess on material change, incident, vulnerability, contract change, new data class, new use case, or failed control and preserve the prior decision.

Primary sources

NIST AI Risk Management Framework 1.0NIST voluntary framework for managing AI risks across design, development, deployment, use, and evaluation. It provides governance and risk-management context but does not certify a legal AI vendor or prescribe one procurement checklist.NIST AI 600-1: Generative AI ProfileNIST profile for applying AI RMF concepts to generative AI risks, including lifecycle, evaluation, governance, and trustworthiness considerations relevant to vendor and use-case diligence.NIST SP 800-53 Rev. 5, Update 1: Security and Privacy ControlsNIST control catalog that informs questions about access, least privilege, auditability, configuration, system and communications protection, incident response, contingency, privacy, and assessment evidence.NIST SP 800-61 Rev. 3: Incident ResponseNIST incident-response guidance for integrating preparation, detection, response, recovery, and improvement into cybersecurity risk management, useful for vendor incident and cooperation requirements.CISA Artificial IntelligenceCISA AI resource hub emphasizing that AI systems should be secure by design and linking guidance for providers, developers, and organizations adopting externally developed AI systems.CISA Secure by DesignCISA secure-by-design guidance for putting security responsibility into product design, development, maintenance, transparency, and customer outcomes rather than relying only on downstream users.CISA Secure by Demand GuideCISA procurement-oriented guidance for customers evaluating whether software providers build security into products and services, useful for translating vendor claims into questions and evidence requests.

Methodology

Use this checklist as a versioned third-party-risk record. Start with the approved use case, data classes, matters, entities, jurisdictions, users, prohibited actions, human-review points, and consequence of failure. Build a complete data-flow register for prompts, files, metadata, retrieval, embeddings, outputs, feedback, telemetry, logs, support, backups, exports, deletion, models, cloud regions, and subprocessors. Record each model and provider dependency, training-use rule, tenant boundary, region, encryption and key control, identity path, privileged access, logging and audit events, retention and deletion rule, incident and vulnerability duty, testing result, integration scope, continuity plan, exit evidence, and contract term. Treat organization-designed thresholds as internal controls, not universal NIST or CISA requirements. For each requirement, record owner, evidence, scope, period, configuration, model or feature version, exception, remediation, approval, review date, and change trigger. Validate negative cases for cross-tenant retrieval, unauthorized exports, prompt injection, poisoned sources, integration overreach, support access, deletion, restore, and provider failure. Certifications and attestations can support assurance, but they are not guarantees; compare their scope and exclusions to the actual service and supplement them with configuration-specific evidence, testing, contract rights, monitoring, and human review.

Contact

Make legal AI vendor diligence evidence-driven

Reach out and learn more about our offerings and how CaseDocker can help you

Built for legal operations teams

Share your use case and we will connect you with the right team for product guidance, pricing, and rollout planning.

Clear next steps

Expect a response from our team with the most relevant next step for your inquiry.

Get in Touch

Get in Touch

We usually reply quickly

FAQs

Start with the use case, data classes, people and matters in scope, prohibited actions, human review, and consequence of failure. Then map every data flow and dependency, including prompts, files, outputs, embeddings, logs, support, backups, model providers, cloud regions, subprocessors, integrations, and deletion paths. The vendor should validate the map for the exact plan and configuration being purchased.

No. A report or certification can be useful evidence for a defined scope and period, but it does not guarantee the security, confidentiality, privilege, availability, accuracy, or legal suitability of a specific legal AI deployment. Check scope, exclusions, findings, region, feature, model, tenant configuration, and contract terms, then add configuration-specific testing and operational evidence.

Define whether prompts, files, outputs, feedback, telemetry, identifiers, support content, or derived data may be used for training, fine-tuning, evaluation, abuse monitoring, human review, or product improvement. State defaults, feature exceptions, opt-out or opt-in behavior, retention, downstream-provider use, deletion, and the controlling contractual language. Do not rely on an undefined statement that customer data is not used to train models.

Use a controlled test population with separate tenants, matters, workspaces, identities, roles, indexes, and integrations. Test search, retrieval, citations, prompts, outputs, previews, exports, APIs, notifications, support access, administrator paths, logs, caches, and background jobs with both allowed and denied cases. Record the feature, model route, configuration, date, evidence, defects, and retest result.

Ask for lifecycle rules and evidence covering live records, prompts, outputs, embeddings, indexes, logs, backups, snapshots, support records, abuse-monitoring data, derived data, subprocessors, restores, and legal holds. Test a representative record through active, archived, deletion-requested, deleted, restored, and backup-expiry states. Require a verification record and document any security, legal, or operational exception.

Review document stores, case and contract systems, email, identity providers, browser extensions, APIs, webhooks, plugins, search indexes, storage targets, and automations. Check scopes, token storage, data minimization, synchronization direction, cache and retry behavior, tenant mapping, downstream retention, revocation, logging, approval gates, and whether the integration can send, change, delete, or share records without a human confirmation.

Request and supplement security tests for authorization, tenant leakage, prompt injection, retrieval poisoning, sensitive-data exposure, insecure plugins, unsafe tool use, model or provider changes, dependencies, vulnerabilities, availability, deletion, and incident recovery. For legal workflows, also test source grounding, citation support, output labeling, hallucination handling, human review, and restricted actions using representative but controlled data.

Repeat it when the model, provider, subprocessor, region, feature, integration, contract, data class, use case, tenant configuration, retention rule, or security control changes. Reassess after an incident, vulnerability, failed test, deletion exception, provider outage, material finding, or new client requirement. Preserve the prior approval and record the change, evidence, owner, decision, and any temporary exception.

Related CaseDocker capabilities

Case management

Keep AI vendor assessments, use cases, matter scope, evidence, approvals, incidents, exceptions, and review history connected to the legal work they protect.

Explore

Compliance management

Track security requirements, control owners, evidence requests, remediation, attestations, review dates, and vendor obligations in a governed workspace.

Explore

Contract management

Link data-use, security, incident, subprocessor, deletion, audit, continuity, change, and exit requirements to enforceable vendor contract terms.

Explore

Playbooks

Turn the security checklist into repeatable intake, review, approval, exception, incident, reassessment, and termination workflows.

Explore

Turn this guide into an operating plan

Share your current legal workflow and CaseDocker can map the right modules, integrations, controls, and rollout sequence.

Book a walkthrough