What an AI Impact Assessment Should Measure Before a Public-Sector Deployment
A practical measurement framework for municipalities, universities, and public agencies evaluating high-impact AI systems before launch

Public agencies are moving artificial intelligence from controlled pilots into services that affect residents, students, employees, and public resources. That transition changes the central governance question. It is no longer enough to ask whether an AI product works. An AI impact assessment for government must measure what the system changes, who carries the risk, how performance varies across real operating conditions, and whether the agency can detect and correct harm after deployment.
A strong assessment is not a generic questionnaire completed at the end of procurement. It is a documented decision process that connects a specific use case to evidence, approval conditions, monitoring thresholds, and an accountable owner. The output should tell an agency whether to proceed, proceed with controls, redesign the use case, or stop deployment.
Why Public-Sector AI Impact Assessments Need Measurable Evidence
Public-sector AI operates inside legal duties, administrative procedures, records obligations, accessibility requirements, cybersecurity controls, and public expectations of due process. A system that summarizes internal notes presents a different risk profile from one that prioritizes housing inspections, flags benefit applications, scores job candidates, or recommends disciplinary action. The assessment must therefore begin with the decision context, not the vendor’s model description.
The NIST AI Risk Management Framework organizes AI risk work around four functions: Govern, Map, Measure, and Manage. That sequence is useful because measurement is meaningful only after the agency defines the system’s context and governance. Management then turns the findings into controls, monitoring, and response plans.
The U.S. Government Accountability Office AI Accountability Framework similarly organizes oversight around governance, data, performance, and monitoring. Canada’s public-sector Algorithmic Impact Assessment offers another concrete model: a structured questionnaire that determines an impact level and links that level to required mitigation measures.
What an AI Impact Assessment for Government Should Measure

1. Purpose, authority, and decision scope
Define the exact task the AI system will perform and the decision it can influence. Record the program authority, the business owner, the affected service, the intended users, and any prohibited uses. The assessment should distinguish between assistance, recommendation, ranking, prediction, and automated decision-making. It should also document whether a human reviewer can change the outcome and what information that reviewer receives.
- Decision or workflow being changed
- Legal and policy authority for the use
- Population and service boundaries
- Degree of automation and human discretion
- Uses that are explicitly out of scope
2. People, rights, and severity of potential harm
Map every group that can be affected, including people who do not directly interact with the system. Estimate both the probability and severity of harm. Relevant consequences may include denial or delay of services, financial loss, loss of privacy, discrimination, physical danger, reputational damage, or reduced ability to challenge a government action. Severity should be evaluated at the individual and population levels because a small error rate can still produce many harmful outcomes at scale.
3. Data fitness and provenance
Measure whether the data is appropriate for the public purpose. An assessment should document sources, collection authority, consent or notice where applicable, time period, missingness, labeling methods, known proxies, retention, and restrictions on reuse. Training data quality alone is not enough. Agencies also need to evaluate live input data, vendor-supplied data, and feedback data created during operation.
- Coverage of the intended population and operating conditions
- Rates and patterns of missing, stale, or incorrect data
- Provenance and rights to use each dataset
- Sensitive attributes and proxy variables
- Data drift between evaluation and production
4. Performance under real operating conditions
Overall accuracy can hide the failures that matter most. Select metrics that match the task and the cost of different errors. A fraud-screening tool may require separate analysis of false positives and false negatives. A generative assistant may require factuality, citation accuracy, refusal behavior, and completion rate. A forecasting model may require calibration and error distributions, not a single average score.
Test with representative agency data, realistic edge cases, accessibility scenarios, low-quality inputs, language variations, and operational constraints. Record the baseline process so reviewers can determine whether the AI system actually improves the service rather than merely appearing more sophisticated.
5. Bias, equity, and distribution of errors
Public-sector AI bias testing should compare outcomes and error rates across legally and operationally relevant groups. The goal is not to produce one universal fairness score. It is to identify who experiences false matches, missed detections, delays, or lower-quality service, and whether those differences are justified by the public purpose. Small groups require careful treatment because unstable estimates can conceal risk or produce misleading conclusions.
- Selection, approval, referral, or escalation rates by group
- False-positive and false-negative rates by group
- Service quality across language, disability, geography, and access conditions
- Intersectional effects where sample sizes support analysis
- Mitigation results and any residual disparity
6. Privacy, security, and misuse resistance
The assessment should model both ordinary failures and adversarial behavior. Measure exposure of personal or confidential information, prompt-injection susceptibility, unauthorized access, data leakage, model extraction, malicious inputs, and dependence on external services. Document what the vendor retains, where data is processed, whether agency data can train other models, and how logs support investigation without creating unnecessary surveillance risk.
7. Human oversight, contestability, and operational readiness
Human review must be designed, staffed, and tested. Simply placing a person after an algorithm does not create meaningful oversight. The reviewer needs authority, sufficient time, understandable evidence, and a documented path to override the output. Affected people need notice when appropriate, a channel to challenge consequential outcomes, and a process that does not require technical expertise to use.
Operational readiness also includes incident ownership, escalation times, rollback capability, vendor support, change control, staff training, records retention, and a decommissioning plan. These controls determine whether the agency can respond when the system changes or fails.
Turn the Assessment Into a Deployment Decision
A completed algorithmic impact assessment template should not end with an unqualified score. It should produce a decision record with evidence, conditions, and named accountability. A useful record includes:
- Risk classification and rationale
- Evidence reviewed and tests performed
- Unresolved limitations and affected groups
- Required controls and contract obligations
- Approval authority and system owner
- Monitoring metrics, thresholds, and review cadence
- Triggers for suspension, rollback, or reassessment
High-severity use cases should require stronger evidence before launch. If the agency cannot test subgroup performance, explain model behavior sufficiently for the decision context, or reverse harmful outcomes, the appropriate result may be a limited pilot or no deployment.
An AI Impact Assessment Is a Lifecycle Control

Pre-deployment review creates a baseline, not a permanent approval. AI systems and their operating environments change. Models are updated, vendors alter services, data distributions shift, staff develop workarounds, and agencies expand use cases. Each material change can invalidate earlier evidence.
NIST’s AI RMF Manage guidance calls for post-deployment monitoring, feedback, appeal and override mechanisms, incident response, recovery, and change management. Agencies should connect these activities to measurable triggers. Examples include a rise in error rates, a widening subgroup disparity, a security incident, a vendor model change, a new data source, a new affected population, or use of the output in a more consequential decision.
Assign each metric an owner, data source, review interval, acceptable range, and response. Monitoring that cannot trigger action is reporting, not control.
How to Use NIST AI RMF in a Municipal Assessment
NIST AI RMF implementation does not require a municipality to reproduce every framework statement in a procurement file. A practical approach is to map the assessment record to the four core functions:
| AI RMF function | Assessment question | Required evidence |
| Govern | Who is accountable and what rules apply? | Owner, authority, policy, procurement and escalation controls |
| Map | What is the context and who can be affected? | Use case, workflow, populations, impacts and dependencies |
| Measure | How does the system perform and fail? | Data tests, task metrics, subgroup results, security and usability evidence |
| Manage | What will the agency do about identified risk? | Conditions, monitoring, incident response, change control, and retirement plan |
Questions Public Agencies Should Ask Vendors
- What exact model version and configuration will the agency receive?
- Which evaluation data reflects the agency’s population and operating conditions?
- Can the agency independently test outputs before and after updates?
- What subgroup and accessibility testing has been completed?
- What data is retained, reused, or disclosed to subprocessors?
- What events require vendor notice, and how quickly will the vendor respond?
- Can the agency export logs, preserve records, suspend the system, and terminate without losing access to essential data?
- Which contractual remedies apply when performance or risk controls fail?
Vendor documentation can support an assessment, but it cannot replace the agency’s evaluation of its own use case. The same model can create different impacts when used in different programs, with different data, or at different points in a decision process.
The Minimum Viable Public-Sector AI Assessment
For a lower-risk pilot, a concise assessment can still be rigorous. At minimum, document the use and owner, affected people, data sources, baseline process, task-specific tests, known limitations, human review, security and privacy controls, monitoring plan, and stop conditions. Increase the depth of evidence as consequences, scale, opacity, or irreversibility increase.
The purpose of an AI impact assessment is not to make deployment frictionless. It is to make the agency’s decision traceable. When the assessment identifies evidence gaps before procurement or launch, it has done useful work. Those gaps become test requirements, contract terms, pilot boundaries, or reasons not to proceed.
Evidence links included in the article:
NIST Artificial Intelligence Risk Management Framework
NIST AI RMF 1.0
NIST AI RMF Manage Playbook
U.S. GAO Artificial Intelligence Accountability Framework
Government of Canada Algorithmic Impact Assessment Tool
Government of Canada Directive on Automated Decision-Making