The 73% You Never See
Riverside County's AI form-filling tool delivered a 73% cut in application processing time — a clean institutional metric that says nothing about whether applicants left with every programme they qualified for. When a county measures throughput and error rate rather than referral breadth or benefit value per case, the AI ships the efficiency, the institution keeps it, and the resident encounters an already-incomplete system running faster.
A 73% reduction in application completion times shipped this year in Riverside County, California, and nobody outside the casework chain knows they are now inside it. The Riverside County Children and Families Commission piloted a form-filling assistant with staff earlier this year, a generative AI tool that autonomously navigates multiple databases and benefit portals and then completes application forms, leaving caseworkers to review and submit. The project received a $1.5 million grant from Google's Generative AI Accelerator program in June 2025, and 20 staff members tested the tool in a 4-month pilot running February to June 2026. The families they serve saw less waiting, fewer repeated questions, faster benefit enrollment. They did not choose any of it, and the measurement that mattered to the institution is not the one that tells them whether the system now works for them.
You sit across from a caseworker who is helping you apply for CalWORKs, Medi-Cal, SNAP, WIC. Typically the caseworker has to manually input your information into the application, field by field, and in Riverside County, caseworkers were asking clients for the same personal information multiple times because their systems weren't connected. The AI agent changes that. It can automatically pull existing data from the case management system and fill out the application for the caseworker to review. The caseworker is no longer typing while you talk. While the AI agent automatically fills the application, the caseworker has more time to talk to their client, possibly uncovering more needs and connecting them to resources. That is the theory, and it is a plausible one, and whether it happens depends on what the caseworker is measured on next.
An estimated $227 billion in benefits go unclaimed annually due to barriers such as lack of information and administrative burdens in the application process, according to Nava Labs Director Genevieve Gaudet. The throughput number is real. The 73% is real. The question is what it buys for the person on the other side of the desk. If the county measures the caseworker on applications closed per week, you get fewer repeat visits and the county celebrates a capacity unlock. If the county measures quality of referral or breadth of programme match, you might get the thing Gaudet mentioned: uncovering more needs. If nobody measures those, the AI ships the efficiency, the institution keeps it, and you leave with one benefit enrolled when you qualified for three.
Federal rule changes to Medicaid and SNAP will impact how applications are designed and administered, creating increasing workloads for caseworkers. States' share of SNAP programme costs will increase based on their payment error rate to beneficiaries. The federal government will reduce its cost-sharing contribution from 50% to 25% in October 2026. That fiscal squeeze is the occasion for the tool. The vendor pitch is administrative relief; the budget case is error reduction and cost transfer. The caseworker feels both, and the resident sits downstream of whichever one the county prioritises in the performance rubric.
The tool does not submit the application autonomously. Nava did not enable the AI agent to submit the application even though it's technically possible, preserving that final step for the caseworker and their client to maintain trust and accountability. That is a defensible design decision, but it also means the measurement split begins right there. Time saved is an institution metric. Errors caught is an institution metric. The caseworker sees a pre-filled form they must review, knowing they are liable for what it says, and that review step is either quality assurance or a production bottleneck depending on how the supervisor sees it. You, the applicant, experience it as the caseworker reading a screen instead of asking you questions. Whether that shift produces better outcomes for you depends entirely on what happens in the time the AI bought back.
The broader Nava Labs toolkit contains four tools. A referral generator that matches clients with relevant benefits and community resources and produces step-by-step action plans; an assistive chatbot that provides real-time, plain-language responses to policy questions with source citations; a document processing tool that identifies document types, captures required information and flags issues during review; and the form-filling capability. The chatbot improved caseworker accuracy by an average of 40%, with the most significant gains occurring in the most difficult client scenarios, based on a 14-week pilot in Los Angeles with 61 caseworkers. Accuracy is good. Accuracy protects you from a denial you should not have received or a benefit calculation that underpays. But the measurement does not tell you if the caseworker now has the margin to notice you also qualify for a local childcare subsidy the chatbot was not trained on, or whether accuracy became the performance ceiling and referral breadth fell off the scorecard.
Nava Labs built, tested, and iterated on benefit navigation tools with governments and community partners in California, Texas, Pennsylvania, and Maryland for over two years, and early results show these tools offer real promise to help caseworkers be more efficient and effective. The institution running the measurement decided what effective means. The resident on the receiving end cannot appeal that choice, cannot see the rubric, and in most cases will not know the tool exists. The 73% is locked in at the county level. What you experience is second-order: whether the county chose to spend that time reduction on seeing more people per day, on deeper intake conversations, or on lowering the casework budget line and keeping throughput flat. All three are rational institutional responses to a 73% process improvement. Only one of them makes you better off in a way you would recognise.
The AI is open source. The tools are built on Nava's open-source Strata platform, designed to provide a standardized architecture for modernising public service delivery systems, and releasing the toolkit as open source is intended to give government agencies greater control over implementation and reduce reliance on proprietary systems. That architectural choice is important for procurement sovereignty and for avoiding vendor lock. It does nothing for the resident unless the agency that deploys it also publishes what it measures and whether those measures include anything the resident cares about. A well-run agency will track referral completeness, benefit dollar value per case, time to first payment, and whether families come back because something was missed. A badly run agency will track applications processed, average handle time, and case closure rate. The AI works identically in both, and you would not know which one you walked into until you either get the call about three more programs you qualify for or you do not.
The Silver Tsunami of Baby Boomers retiring and the struggle to attract a younger workforce to public sector jobs compounds the system's burden, and bringing in computing automation can help optimise how governments assess eligibility, make enrollment decisions and manage benefit delivery. The workforce argument is real. The optimisation is real. The thing that does not follow automatically is that optimisation for institutional capacity produces optimisation for resident outcomes. It can, if the institution aligns the two. The technology does not enforce that alignment. The measurement system does, and you are stuck with the one the county chose before the tool went live.
Theform-filling assistant is in production. The chatbot ran a controlled pilot and published results. Riverside and Los Angeles are replicable templates. The tool will spread, because the business case is unambiguous and the fiscal pressure is acute. What spreads with it is the measurement frame the first deployments locked in. If early adopters measured throughput and error rate, later adopters will inherit those KPIs because they are the ones attached to the case studies and the grant reporting. If nobody measured referral breadth, time to first payment, or whether multi-benefit households got everything they qualified for in one visit, those outcomes stay invisible, and the AI optimises a system that was already incomplete.
The 73% is the number Riverside can defend. It is a clean, auditable measure of a process that used to take longer and now does not. It says nothing about whether you left whole, whether the caseworker had time to ask the question that would have surfaced your eligibility for something else, or whether the error rate the county is now penalised for has actually moved. Those are harder to measure, slower to validate, and not the number that justified the $1.5 million. You never see the 73%. You see what the county decided to do with it, and that decision was made when they picked what to measure, which happened before you walked in.
Tarry Singh is the founder and CEO of Real AI (realai.eu), an enterprise AI advisory and deployment firm working with global enterprises on production agent systems, model risk, and AI sovereignty strategy. He also leads Earthscan (earthscan.io) for Energy AI, and is a founding contributor to the EU-funded HCAIM and PANORAIMA programmes for responsible AI education across European universities. He writes at tarrysingh.com.