Where AI delivers measurable results in government operations
Agencies evaluating where to apply artificial intelligence commonly begin with public facing services. Those applications are visible, they demonstrate effectively, and they support a clear narrative.
The stronger initial candidates are generally internal. Correspondence handling, records classification, case triage, quality review, and backlog processing are high volume, procedurally consistent, and largely invisible outside the organization. They present conditions materially more favorable to a demonstrable result.
Conditions favoring internal operations
Volume and repeatability. Value from automation scales with recurrence. Operations work recurs continuously and follows documented procedure, which establishes a definable correct answer against which performance is measured.
Established baselines. Operations functions track throughput, cycle time, and backlog because they are managed against those measures. The advantage is more consequential than it appears. Programs unable to demonstrate a prior state encounter difficulty justifying continuation irrespective of the improvement achieved.
Contained risk. Internal work subject to staff review before any output becomes final presents a manageable failure mode. Errors are identified by a reviewer rather than experienced by a member of the public.
Receptive users. Staff managing a persistent backlog are generally receptive to assistance. Adoption resistance declines substantially where capability addresses a condition staff already regard as burdensome.
Applications performing consistently
Correspondence and FOIA triage. Categorizing requests, routing to the appropriate office, identifying likely responsive record sets, and flagging material requiring review. Determinations remain with staff while search and sort burden declines.
Records classification and retention. Applying retention schedules to unstructured content at scale is work agencies rarely have capacity to perform to standard. The compliance benefit is direct.
Case triage. Sorting incoming cases by complexity and completeness so straightforward matters progress efficiently and complex matters reach experienced staff earlier.
Quality review sampling. Identifying transactions carrying characteristics associated with error, in place of random sampling, concentrating limited review capacity where error is most probable.
Knowledge retrieval for staff. Frontline staff searching extensive policy documentation constitutes a strong retrieval application, provided output cites its source to permit verification.
Selecting the initial use case
Candidates should be scored across four dimensions.
Volume. Frequency of occurrence.
Consistency. Presence of documented procedure and a definable correct answer.
Review tolerance. Whether output can be verified before it becomes consequential.
Baseline availability. Whether the function is already measured.
Candidates scoring well across all four are strong. Candidates scoring poorly on review tolerance, where output becomes consequential without verification, belong in a later phase requiring materially greater rigor and likely fall under M-25-21 high impact requirements.
Measurement
Baseline before deployment. Measured rather than estimated: cycle time, throughput, error rate under existing review, and cost per transaction. This step is routinely omitted, and its omission is why many pilots cannot demonstrate value they in fact produced.
Quality reported with speed. Increased processing rate at reduced accuracy is not an improvement. Both measures should be reported together.
Review burden accounted for. Where output requires substantial correction, net gain may be limited. Review time constitutes part of the cost.
Unsuccessful pilots reported. Programs reporting only successful outcomes lose internal credibility. Pilots that did not succeed constitute useful information and establish the credibility of those that did.
Strategic position
Beginning with internal operations establishes what an agency requires before undertaking higher visibility work: demonstrated competence, an exercised governance process, staff experienced in operating these capabilities, and results measured in the agency’s established terms.
That foundation materially improves the probability that subsequent citizen facing work succeeds, and substantially improves its prospects for funding.
Resistance in this category of work originates in an unexpected location. It is rarely the staff, who are generally receptive. It originates more frequently in the reporting function, because a backlog explained consistently for several years acquires a different explanation, and accountability for that change must be assigned. The condition is worth anticipating. The technical pilot is frequently the more straightforward component; the more difficult component is establishing what the revised measures indicate.
The most valuable artificial intelligence work in government is frequently the work that receives no external visibility.