Legal AI Governance
Legal AI Pilot Success Metrics
Define denominator-aware legal AI pilot metrics for quality, citations, review time, adoption, incidents, cost, latency, support, and go/no-go decisions.
Direct answer
Legal AI pilot success metrics should connect a declared eligible population to benchmark tasks, quality and error rubrics, appropriate abstention, citation validity, review time, rework, adoption, security and privacy incidents, cost, latency, and support demand. Publish numerator, denominator, unit, period, exclusions, unknowns, and every failure state. Use organization-approved stop or go criteria as governance gates, not as a guarantee of causation, ROI, legal accuracy, or production safety.
Definitions
Eligible population
The versioned set of users, roles, matters, documents, requests, or tasks that meet the pilot inclusion rule during the declared period and could receive the measured workflow.
Benchmark task
A repeatable, scoped evaluation scenario with a task ID, input set, expected evidence or rubric, risk tier, evaluator instructions, and a declared treatment of unavailable or disputed answers.
Reference answer
The approved answer, range, classification, source set, or decision rubric used to evaluate a benchmark task, including uncertainty where more than one answer can be defensible.
Material error
An error category defined in advance that could change a legal, operational, financial, privacy, security, client, or compliance decision within the pilot scope.
Appropriate abstention
A refusal, escalation, or request for more information that occurs when the task is unsupported, ambiguous, unauthorized, outside scope, or otherwise requires a declared human route.
Citation validity
The proportion of sampled citations that resolve to the claimed source and support the stated proposition at the recorded location, version, and access boundary.
Review time
The measured reviewer effort or elapsed review duration for a defined output population, reported with the start event, stop event, pause rule, unit, and distribution.
Rework
A material correction, rerun, escalation, or replacement needed before an output can satisfy the declared workflow acceptance condition.
Adoption
A qualifying use measure tied to the eligible population and intended workflow, such as an eligible user completing a permitted task, not merely signing in or opening a page.
Security or privacy incident
A confirmed or suspected event involving unauthorized access, disclosure, retention, integrity loss, policy violation, unsafe handling, or other defined security or privacy impact.
Stop or go criterion
An organization-approved decision gate that triggers stop, contain, revise, expand, or hold action based on declared evidence, thresholds, risk, authority, and unresolved limitations.
Failure register
The preserved record of failed, rejected, incomplete, disputed, unavailable, unsafe, or out-of-scope pilot cases and the resulting owner, action, status, and disposition.
Field definitions
Population and benchmark design
- pilot_id
- Stable identifier for the pilot, scope, approved use case, model or configuration version, and measurement version.
- Type: Versioned record
- Requiredness: Always required
- Validation: Do not reuse an identifier after a material change to scope, model, source corpus, rubric, or decision authority.
- Owner: Pilot owner
- eligible_population_definition
- Reproducible rule for users, roles, matters, documents, requests, or tasks eligible for each metric.
- Type: Query, roster, or cohort record
- Requiredness: Always required
- Validation: Store source, effective date, jurisdiction, inclusion, exclusion, unknown, and not-applicable counts.
- Owner: Metric steward
- benchmark_task_id
- Stable identifier for a benchmark scenario, input set, risk tier, expected evidence, and evaluation rubric.
- Type: Controlled task record
- Requiredness: Required for benchmark measures
- Validation: Preserve task inputs, source versions, expected answer or range, evaluator instructions, and representativeness rationale.
- Owner: Evaluation lead
- reference_and_rubric_version
- Versioned reference answer, acceptable range, citation requirements, material-error taxonomy, and abstention rule.
- Type: Versioned rubric
- Requiredness: Always required for quality measures
- Validation: Permit not assessable, disputed, and multiple-defensible-answer outcomes with adjudication notes.
- Owner: Subject-matter reviewer
Quality and evidence
- quality_result
- Task-level evaluation state and rubric result for the model output before and after any human correction.
- Type: Pass, partial, fail, not assessable, or disputed
- Requiredness: Required for every evaluated task
- Validation: Retain original output, reviewer, review time, failure categories, severity, and adjudication; never overwrite a failure with the corrected answer.
- Owner: Evaluation reviewer
- material_error_category
- Controlled category for errors such as unsupported conclusion, missing issue, wrong extraction, unsafe instruction, privacy exposure, or citation defect.
- Type: Controlled value plus evidence
- Requiredness: Required when an error is present
- Validation: Allow multiple categories, record task and proposition scope, and distinguish observed error from reviewer disagreement.
- Owner: Evaluation lead
- abstention_result
- The requested task state, model response state, appropriateness judgment, escalation, and final disposition.
- Type: Structured decision record
- Requiredness: Required for abstention-eligible tasks
- Validation: Separate appropriate abstention from inappropriate abstention, unsupported answer, and human escalation.
- Owner: Workflow owner
- citation_check
- Citation-level result for resolution, source identity, source version, proposition support, location, and authorized access.
- Type: Citation evidence record
- Requiredness: Required when output cites sources
- Validation: Classify valid, unsupported, fabricated, stale, inaccessible, ambiguous, or not checked, and preserve the cited location.
- Owner: Evidence reviewer
- review_time
- Measured reviewer duration from declared start to stop events with pauses, role, time zone, and unit.
- Type: Seconds or minutes
- Requiredness: Required for review-effort measures
- Validation: Keep queue delay, model latency, active review time, and rework time as separate clocks.
- Owner: Measurement steward
Adoption, operations, and risk
- qualifying_adoption_event
- Meaningful permitted action that demonstrates the eligible role used the pilot workflow for its intended purpose.
- Type: Event with actor and workflow
- Requiredness: Required for adoption measures
- Validation: Do not count login, page view, training attendance, or test traffic unless the pilot definition explicitly makes it the decision event.
- Owner: Workflow owner
- incident_record
- Confirmed or suspected security, privacy, access, retention, integrity, or policy event with impact and response history.
- Type: Incident or near-miss record
- Requiredness: Required when an event occurs
- Validation: Preserve suspected events, severity, affected population, containment, notification decision, owner, and closure evidence.
- Owner: Security or privacy owner
- cost_record
- Direct pilot spend and usage quantities by declared currency, period, task, user, model, infrastructure, review, support, and remediation category.
- Type: Currency and usage ledger
- Requiredness: Required for cost measures
- Validation: Keep direct observed cost separate from avoided cost, opportunity cost, forecast, benefit, or ROI assumptions.
- Owner: Finance or pilot owner
- latency_measurement
- Request and output timestamps, clock definition, unit, timeout state, and percentile calculation population.
- Type: Milliseconds, seconds, or minutes
- Requiredness: Required for latency measures
- Validation: Publish p50 and p95 with count, timeout count, clock source, and missing-timestamp count.
- Owner: Technical owner
- support_case
- Support request, incident, question, defect, training need, or enhancement request linked to a pilot user, task, release, or workflow.
- Type: Support record
- Requiredness: Required when support is measured
- Validation: Record reason, severity, affected population, first response, resolution, repeat contact, and unresolved risk.
- Owner: Support owner
Decision and failure preservation
- failure_register_id
- Stable reference for every failed, rejected, disputed, incomplete, unsafe, unavailable, or out-of-scope case.
- Type: Linked issue record
- Requiredness: Always required when applicable
- Validation: Do not delete or collapse failures into a successful corrected output; link remediation and the original evidence.
- Owner: Pilot governance owner
- stop_go_decision
- Decision, scope, evidence version, criteria results, residual risks, authority, conditions, expiry, and next review.
- Type: Governance decision record
- Requiredness: Required before expansion or closure
- Validation: Use stop, contain, revise, hold, limited go, or go within scope; record unmet criteria and accepted exceptions.
- Owner: Decision authority
Controlled vocabulary guidance
- Quality result
- Examples: Pass, Partial, Fail, Not assessable, Disputed, Withdrawn
- Governance: Define each state in the rubric, preserve the original output, and require adjudication for disputed or multiple-defensible answers.
- Error severity
- Examples: Critical, High, Moderate, Low, Informational
- Governance: Use organization-approved anchors tied to decision impact, affected population, exposure, detectability, and required containment; do not infer severity from frequency alone.
- Abstention result
- Examples: Appropriate abstention, Inappropriate abstention, Unsupported answer, Human escalation, Not an abstention opportunity
- Governance: Maintain a task-level expected action and explain why the observed response did or did not satisfy the boundary.
- Citation status
- Examples: Valid, Unsupported, Fabricated, Stale, Inaccessible, Ambiguous, Not checked
- Governance: Check source identity, version, location, proposition support, and reviewer access. Keep unchecked citations out of the valid numerator.
- Pilot decision
- Examples: Stop, Contain, Revise, Hold, Limited go, Go within scope
- Governance: Tie each decision to versioned criteria, evidence, residual risk, authority, conditions, expiry, and next review; a go decision is not a performance guarantee.
- Failure disposition
- Examples: Open, Contained, Remediation planned, Accepted with owner, Re-test required, Closed with evidence
- Governance: Preserve the original failure and require an owner, due date or trigger, compensating control, and closure evidence.
Practical workflow
State the pilot decision
Write the decision the pilot must support: refine the workflow, expand to a named cohort, contain use, stop the test, or hold for more evidence. Identify the accountable decision owner and the authority for access, client use, data handling, and production release.
Declare the eligible population
Define eligible users, roles, matters, document sets, requests, or benchmark tasks with a stable query or roster, effective date, jurisdiction, access rule, reporting period, and owner. Count eligible, ineligible, unknown, unavailable, and not-applicable records before interpreting a rate.
Set the pilot boundary
Record permitted use cases, excluded use cases, data classes, jurisdictions, matter types, model or configuration version, human-review requirement, retention rule, escalation route, and prohibited decisions. A pilot boundary is part of the denominator and interpretation contract.
Build representative benchmark tasks
Create task IDs across routine, difficult, rare, high-impact, ambiguous, multilingual, incomplete, and adversarial examples that are relevant to the declared workflow. Preserve source references, expected evidence, risk tier, evaluator instructions, and the reason each task is representative or intentionally stress-oriented.
Define the reference and rubric
For every task, state the expected answer or acceptable range, required citations, material-error categories, abstention conditions, privacy or security checks, and reviewer adjudication rule. Permit a documented unresolved or multiple-defensible-answer state instead of forcing false certainty.
Freeze instrumentation and versions
Record the model, prompt or template, retrieval source, index or knowledge-base version, configuration, user role, data snapshot, timestamp, latency clock, evaluator version, and metric version. Do not compare outputs across silent changes in task mix, source corpus, model, or rubric.
Measure quality and errors
Evaluate each eligible task against the rubric and preserve pass, partial, fail, not assessable, and disputed states. Report quality pass rate, material-error rate, error categories, severity, affected task IDs, reviewer, and adjudication. Do not hide a failed task because a reviewer corrected it.
Measure abstention and escalation
Separate appropriate abstentions, inappropriate abstentions, unsupported answers, human escalations, and tasks that should not have been attempted. Use the eligible abstention-opportunity population for the abstention measure and retain the task-level reason, expected action, and final disposition.
Validate citations and source grounding
Sample or census citations at the citation level. Check that each citation resolves, refers to the intended source and version, supports the proposition, and is available to the authorized reviewer. Record missing, stale, inaccessible, contradictory, fabricated, or overbroad citations as failures or limitations.
Measure reviewer effort and rework
Define review start and stop events, pause rules, reviewer role, and unit. Track median and percentile review time, outputs requiring material correction, correction categories, reruns, escalations, and unresolved work. Keep model output time separate from human review time and queue delay.
Measure adoption and workflow completion
Count eligible users who perform the declared qualifying action and eligible workflow instances that reach the accepted terminal state. Report access, activation, repeat use, completion, and non-use separately by role, cohort, matter type, office, and configuration version.
Monitor security, privacy, and support
Record confirmed and suspected incidents, near misses, unauthorized data paths, access failures, retention exceptions, policy violations, support contacts, severity, response time, affected population, and unresolved risk. Report counts and rate bases without treating a zero observed incident count as proof of zero risk.
Measure cost and latency with units
Publish direct pilot spend, usage or inference units, cost per completed qualifying task, and cost categories separately. Measure request-to-first-output and request-to-final-output latency in seconds or minutes using p50, p95, maximum, timeout, and missing-clock counts. Do not convert these measures into promised savings.
Review the failure register
Reconcile all benchmark tasks, production-like tasks, rejected inputs, abstentions, incidents, support cases, timeouts, citation defects, reviewer corrections, unknowns, and excluded records. Every failure needs a stable ID, category, severity, owner, containment or remediation, status, and decision impact.
Apply stop or go governance
Compare the complete evidence set with approved gates. Stop or contain for a defined critical failure, unresolved privacy or security issue, unsafe material-error pattern, unsupported use, or missing authority. Go only for the stated scope when evidence meets the approved criteria and residual risks, limitations, owners, review cadence, and rollback route are accepted.
Comparison
| Metric family | Denominator and unit | Interpretation boundary |
|---|---|---|
| Quality and material error | Evaluated eligible benchmark tasks; percent of tasks, plus task count | Shows rubric performance for the declared task set; does not prove legal correctness outside scope. |
| Abstention | Eligible abstention-opportunity tasks; percent of opportunities, plus task count | Shows boundary behavior; a high or low rate needs task mix and appropriateness review. |
| Citation validity | Citations checked or sampled; percent of citations, plus citation count | Shows source support in the checked set; it does not prove every uncited proposition. |
| Review time and rework | Reviewed outputs or completed tasks; minutes per output and percent requiring material correction | Separates human effort from model latency and should include distributions and missing clocks. |
| Adoption and completion | Eligible users or eligible workflow instances; percent of users or instances | Shows qualifying use or completion, not value creation or safe production readiness. |
| Incidents and support | Confirmed or suspected events and support cases; count and rate per declared users or tasks | A zero observed event count does not establish zero risk when detection or coverage is limited. |
| Cost and latency | Direct currency per completed qualifying task and seconds or minutes per request | Reports observed pilot operations; it is not a savings, ROI, or service-level guarantee. |
Limitations and exceptions
- Pilot results are conditional on the declared population, tasks, sources, model, configuration, reviewer rubric, period, and operating controls. They do not establish performance for every legal workflow or future release.
- Quality, citation, adoption, cost, and time measures can be distorted by task mix, reviewer disagreement, missing events, access limits, learning effects, concurrent process changes, and selective participation.
- A measured association between the pilot and an outcome does not prove that the AI system caused the outcome. Do not claim causal improvement, avoided cost, savings, or ROI without an appropriate design and validated accounting treatment.
- A low observed incident count can reflect limited exposure, weak monitoring, incomplete reporting, or a short pilot. Preserve suspected events and near misses, and state detection and coverage limits.
- Benchmark tasks can become stale, overfit, or unrepresentative. Version task sets, refresh them after workflow or source changes, and include difficult, ambiguous, excluded, and failure cases.
- A valid citation does not by itself make an output legally correct, complete, privileged to disclose, or suitable for a client, court, regulator, or business decision.
- Stop or go criteria are governance controls designed by the organization. They are not universal legal thresholds, certifications, warranties, or permission to remove qualified human review.
Primary sources
Methodology
Use a versioned metric contract for every pilot measure. The contract must state the decision purpose, unit of analysis, eligible population, reporting period, time zone, source systems, event definitions, exclusions, unknown handling, reviewer or adjudicator, formula, owner, and interpretation limits. Freeze the eligible population before calculating rates. Keep population counts, task counts, citation counts, output counts, user counts, incident counts, support counts, currency, and time units visible. Quality pass rate = benchmark tasks with a rubric result of Pass / eligible benchmark tasks evaluated x 100. Publish Partial, Fail, Not assessable, Disputed, Withdrawn, and missing-rubric counts separately; do not move corrected failures into the Pass numerator. Material error rate = evaluated tasks with at least one declared material error / eligible benchmark tasks evaluated x 100. Also report error count by category and severity because a task-level rate can hide multiple errors. If the decision is citation-sensitive, citation validity rate = citations verified as resolving to the claimed source, version, location, and supported proposition / citations checked x 100. A citation that is not checked is not valid for the numerator. Report unsupported, fabricated, stale, inaccessible, ambiguous, and not-checked citations separately. Appropriate abstention rate = tasks with an expected abstention or escalation where the model abstained or escalated in the approved manner / eligible abstention-opportunity tasks evaluated x 100. Inappropriate abstention rate = abstentions that occurred where the task was supported and within scope / eligible supported tasks attempted x 100. Keep unsupported answers and tasks that should not have been attempted as separate failure categories. Review time per output = reviewer active minutes / outputs reviewed, with queue delay, model latency, pause time, reruns, and missing clocks reported separately. Publish median, p75, p95, maximum, count, and unit when the sample supports those summaries. Rework rate = outputs requiring one or more material corrections or reruns before acceptance / outputs reviewed x 100. Preserve every correction and count repeated rework separately from first-pass failure. Qualifying adoption rate = eligible pilot users who complete at least one declared permitted action / eligible pilot users provisioned and in scope x 100. Workflow completion rate = eligible workflow instances reaching the accepted terminal event with required evidence / eligible applicable workflow instances x 100. Report non-use, blocked access, canceled work, unknown actor, and not-applicable workflow counts rather than treating them as successful or failed completion. Incident rate may be reported as confirmed or suspected incidents / eligible pilot users, eligible tasks, or pilot days, multiplied by a declared scale such as 1,000; publish the raw count, coverage, detection method, severity, and unresolved cases because a rate is not evidence that unobserved risk is absent. Direct cost per completed qualifying task = observed direct pilot currency / completed qualifying tasks, with currency, period, cost categories, and excluded or unallocated spend stated. Do not subtract forecasts, hypothetical avoided cost, internal time estimates, or benefits unless they are separately labeled and supported; this guide does not establish ROI or causation. Latency is measured from a declared request event to a declared output event in milliseconds, seconds, or minutes; report p50, p95, maximum, timeout count, and missing-clock count. Support contact rate = support contacts linked to the pilot / eligible active pilot users or completed qualifying tasks, using the denominator that matches the question and publishing both raw contacts and repeat contacts. Define organization-designed stop or go gates before inspecting the result. Example gates may require no unresolved Critical security or privacy incident, no prohibited-data use, citation validity above an approved local threshold for citation-required tasks, material-error and appropriate-abstention results within the approved risk-tier bands, complete failure-register ownership, review-time evidence for the intended workflow, and an approved rollback or containment route. These examples are not universal thresholds. Apply stricter gates to high-impact or client-facing tasks, retain every failed or disputed case, and require authorized review of exceptions. A limited go means only the named population, task family, data class, configuration, human-review step, and period may proceed. Reassess after model, prompt, retrieval, source, policy, workflow, user, jurisdiction, or data changes.
Turn an AI pilot into evidence your legal team can govern
Reach out and learn more about our offerings and how CaseDocker can help you
Built for legal operations teams
Share your use case and we will connect you with the right team for product guidance, pricing, and rollout planning.
Clear next steps
Expect a response from our team with the most relevant next step for your inquiry.
Get in Touch
Get in Touch
FAQs
Related CaseDocker capabilities
Case management
Connect AI-assisted matter workflows to ownership, deadlines, evidence, permissions, review history, and controlled human escalation.
ExploreContract management
Evaluate AI-assisted contract intake, review, obligations, approvals, and post-signature work with source-linked records and accountable owners.
ExploreCompliance management
Route AI pilot findings, incidents, controls, evidence, exceptions, owners, and remediation through a governed compliance workflow.
ExplorePlaybooks
Turn approved pilot boundaries, review gates, escalation rules, and repeatable legal workflows into controlled operating playbooks.
ExploreTurn this guide into an operating plan
Share your current legal workflow and CaseDocker can map the right modules, integrations, controls, and rollout sequence.
