Are we ready for this system?

By Thomas Byrnes
• • 14

Welfare AI, human sign-off, and the judgement we hand over with it

A widow in Hyderabad was denied subsidised food for more than seven years because an algorithm confused her late husband, a rickshaw puller, with a car owner. When she brought the evidence, officials agreed the algorithm had got it wrong. Then they turned her down on income grounds that, according to Al Jazeera, weren't true either.

In Al Jazeera's words, "Once excluded, the onus is on the removed beneficiaries to prove to government agencies that they were entitled to the subsidised food."

Once the system had decided, the household had to disprove it. Being right wasn't enough.

The system is Samagra Vedika. Telangana piloted it in its food security scheme in 2016, and by 2018 was using it for most of its welfare schemes. It links records across government databases to decide who qualifies. Between 2014 and 2019, the state cancelled more than 1.86 million food security cards and rejected 142,086 new applications without notice. Those figures start two years before the pilot, so they can't all be the platform's doing, and Al Jazeera doesn't claim they are. What it did find is that "several thousands of these exclusions were done wrongfully, owing to faulty data and bad algorithmic decisions". A Supreme Court-ordered re-verification later suggested that at least 7.5% of the cards it checked had been wrongly rejected.

AI risk assessment starts with three questions. What judgement are we delegating? What authority are we granting? And can we supervise the work?

These aren't hypothetical questions. Telangana and the other cases here come from AI Adoption in Social Protection: A Global Evidence Review (2020-2025), which Avril James and I wrote for the DCI AI Hub. GIZ published it in August. It documents 99 cases in 48 countries. Those are the countries where we could verify a documented case, not a map of everywhere AI is in use.

In the UK, the Department for Work and Pensions uses a machine learning model to flag Universal Credit advance requests for fraud review. The department says "it is always a human [who] makes the decision". Its own fairness analysis found statistically significant referral disparities by age, disability, marital status and nationality, and Computer Weekly reported that the assessment "did not mention anything about the role or prevalence of automation bias".

The same department has tested generative AI too. One tool, AIgent, supports Personal Independence Payment staff "by summarising evidence for inclusion in decision letters", according to DWP's 2023-24 annual report as reported by PublicTechnology.

Now picture a summary tool of that kind used somewhere else: a fraud review in a means-tested programme. It reads a household's case file, highlights a gap in the income records and leaves out the explanation the claimant gave. The officer is busy, the summary reads well and the file is long. She accepts the recommendation. That's an illustration, not a documented case.

Telangana's platform links records. DWP's fraud model scores risk. Neither is generative AI. What they share with a summary tool is selection before human review: the system decides what the officer gets to look at. Generative tools take that a step further, from flagging into interpreting. Telangana shows where selection failures end up.

A human makes the final decision. The AI decides which evidence reaches her, and how it looks when it gets there.

For the rest of this piece I'll work through a programme like these. I'll call the AI tool "the assistant" and the institution running it "the programme".

The judgement hides inside ordinary tasks

The first step is to spot the judgement inside a task that looks routine.

Ask an assistant to summarise a document and it has to pick what matters, decide what to leave out and interpret material that may be incomplete or contradictory. Ask it to prioritise cases and it's deciding which differences deserve scarce attention.

That's the useful part of generative AI. It's also the part we tend to wave away as "producing text".

By judgement I mean selection and interpretation within a task. The AI has no duty of care and no professional accountability. Those stay with us, which is exactly why the selection matters.

A summary changes what a manager notices. A recommendation becomes the default unless someone has the time and the confidence to push back.

So delegation starts well before the assistant gets permission to send a message, change a record or suspend a payment. The professional's job is deciding which judgement can be handed over, giving the assistant what it needs, setting the boundaries and checking whether the result is fit for use. Writing the prompt is only one part of that.

The model sits inside an arrangement

In The Model Was Never the Whole Story, I argued that model performance can't carry the governance burden on its own. These cases show why.

Testing the model might tell you how reliably it spots a particular pattern. Useful. It won't tell you whether the records are current, whether the summary keeps the counter-evidence, whether officers can get to the source file, or whether a household can challenge the result.

Now give the assistant permission to suspend payments automatically. Its analytical performance hasn't moved. The cost of an error has, and it lands faster.

Or keep human approval but double the officers' review queue with the same staff. The safeguard is still on the organogram. Whether it works on a busy Tuesday afternoon is another matter.

The risk belongs to the whole arrangement: model, data, workflow, permissions, people and institution. When I ask "are we ready for this system?", that's the system I mean. The AI is one component of it.

That's the problem Dr Simona Dobre, Dr Laura van der Erve and I set out to address in the AI Hub Risk Assessment Framework for Social Protection, developed for the DCI AI Hub for Social Protection and published by GIZ in August. Its central question is the title of this piece.

A note on scope. The framework is tailored to non-contributory social assistance and can be adapted to other parts of social protection with expert judgement. In its own words, it is not a substitute for legal compliance, a data protection impact assessment or procurement due diligence.

Start by defining the delegation

"Use AI to improve fraud detection" isn't something anyone can assess. It doesn't say what the AI will actually do.

So write the job down. The same assistant, working on the same case file, could be asked to:

  • flag records that look unusual
  • read the claimant's explanations and judge whether they hold up
  • rank cases for investigation
  • recommend that a payment be suspended

Each step down that list hands over more judgement, and each needs a different assessment. A tool that flags records for a person to check is a different risk from one whose recommendation usually becomes the decision.

Then write down what a good result looks like, and what the assistant must never drop. For a fraud review, a tightly bounded brief might say: organise the evidence, point out discrepancies and missing records, and suggest questions for the officer. Always include the claimant's own explanation. Label every point as either a suspicion or an established finding.

Even then, check it. Writing a rule into the brief doesn't guarantee the assistant follows it.

It also needs a reason to exist. Has the programme shown the problem needs AI at all? Would better records, clearer procedures or a simpler method get the same benefit? Tool 1 asks exactly this before any scoring starts: was a non-AI approach considered?

Consequences first, then readiness

Tool 1, the Intrinsic Risk Assessment Worksheet, scores likelihood, severity and scale. If any one of them is high, the risk is high. A serious risk can't hide inside a reassuring average.

For our programme, that means the chance of a misleading finding, what it costs a household, and how far errors could spread. A small trial cuts the number of people exposed. For each household caught by an error, the harm is the same size.

Sweden shows what happens when scale and severity are both high. Its social insurance agency, Försäkringskassan, used a model whose risk scores automatically triggered fraud investigations into parents claiming temporary benefit to care for sick children. In 2017 the benefit drew around 977,730 applications, and the model picked 5,082 of them for investigation. That's about half of one percent. But it's half a percent of a national scheme, and for the families picked, Lighthouse Reports found benefit payments could be delayed while investigators worked. Journalists found it disproportionately flagged women, people with immigrant backgrounds, low earners and people without university degrees. The agency took it out of use during an inspection by the Swedish data protection authority, saying it wanted to assess whether it complied with the new European AI regulation.

Tool 2, the Context Risk Multiplier, then looks at the legal framework, institutional capacity, the data ecosystem, and public trust and transparency.

Here I'd want evidence that officers understand the assistant's limits and have enough time to check its output against the case file. I'd want to see how the programme handles disagreement, corrects records and responds when a household challenges a finding.

Affected people need a role in judging those arrangements, and the checklist makes consultation with affected communities or their representatives a minimum safeguard. An appeal channel that looks accessible to the people who designed it can be unusable for someone without reliable connectivity, literacy, documents, or confidence that complaining is safe.

Readiness shows up in how the service works for the people who depend on it.

Instructions, access and authority fail in different ways

In our AidGPT sessions on connected systems, we separate written instructions, technical permissions and organisational authority.

"Do not suspend payments" is an instruction. Removing access to the payment controls is a technical restriction. Organisational authority decides who may approve a suspension, and on what conditions.

An AI assistant can have technical access well beyond its task. A staff member can connect a tool without the authority to expose the records it reaches. And the assistant itself can follow its instructions perfectly while doing something it should never have been asked to do.

So I'd ask whether the programme's assistant could do its useful work with less access. Then I'd ask for evidence that the restrictions actually hold. A permission in a planning document is a proposal until someone has configured it and tested it.

Oversight has to name the work

I've written before that a framework can't sit at the keyboard. It can't exercise judgement for the person using it.

Appointing a reviewer is the start. That person needs the relevant evidence, enough time, the competence to examine the result and the authority to disagree. Without those, the institution keeps a human step and quietly gives away the judgement. As the framework puts it, human-in-the-loop then "becomes a compliance label rather than a safeguard".

That's why the DWP gap matters. A human decision that nobody has tested for automation bias is a safeguard on paper.

It helps to separate three activities that often get blurred. Review asks whether the work is suitable and spots problems. Verification checks the consequential claims against the evidence. Approval authorises a particular use or action.

An officer can find a summary well written without checking its account of the household's circumstances. A second AI tool can critique that summary without ever opening the original records. Neither gets the recommendation to approval.

I'd want the programme to walk me through its hardest cases. Show me how omitted evidence gets caught. Show me what happens when the reviewer disagrees. Show me who can stop an adverse action, and how a wrong decision gets put right.

The evidence has to travel with the work

Supervision gets harder when work passes through several hands.

One AI tool extracts the information. Another writes the summary. A colleague forwards it. The officer gets a polished document and can't see what was uncertain at the start.

That's how a tentative finding turns into a confident recommendation without anyone deciding the evidence got stronger.

So ask every stage what it receives, what it produces and what record goes with it. Sources, qualifications, checks and the current version all matter.

In our programme, the summary mustn't blur missing evidence into evidence of wrongdoing. And a check on an earlier draft doesn't count as approval of the revised recommendation.

Disclosure belongs here too. The next person needs to know what AI contributed, what was checked and what's still open, or they can't supervise it.

A pilot has to earn permission too

Tool 3, the Risk Register, records risks, mitigations and responsible owners. Tool 4, the Final Go / Pilot / No-Go Checklist, then tests whether the essential safeguards are in place.

Calling something a pilot doesn't let you skip them. Under the checklist, a single "No" on the minimum safeguards means an automatic No-Go. A pilot also needs defined boundaries, enhanced monitoring, automatic stop conditions and a fully human-led pathway for anyone who wants one.

The checklist requires a named person to review and approve eligibility, exclusion, sanction and appeal decisions before they take effect. For pilots, that review extends to every decision point.

Those are the toolkit's requirements. What the applicable law demands is a separate question, and you still have to answer it.

So the programme can't justify a poorly governed deployment by calling it experimental. It has to show the proposed use is acceptable within its real boundaries. If it can't, the options are redesign, a narrower task, more preparation, or not using AI for this at all.

Useful delegation needs a workable route

The point of all this is defensible decisions about useful work.

AI could help an officer work through a long record, find inconsistencies or surface questions worth asking. Those benefits deserve to be weighed against the cost of supervision, correction and failure. A good design makes the work possible and keeps the protections people rely on.

So start with one proposed use case. Write down the judgement being delegated, the evidence available, the access needed and what happens if it goes wrong. Work through the framework and toolkit. Name what's unresolved and what would count as resolving it.

Then make the deployment decision explicit, with its limits and the circumstances that would trigger a reassessment.

For the programme, the test is whether it can explain and defend the whole path from a case record to an action that affects a household. For the rest of us, it's the same question.

What judgement are we delegating? What authority are we granting? And can we supervise the work?

Practising this before it's live

This is what we work through in AidGPT, the responsible AI training programme I run for humanitarian and development professionals. Six live 90-minute sessions, practising on synthetic cases, so people learn to spot the judgement inside a task, set its boundaries and check the output before anything real depends on it.

The next two cohorts start on 27 October (Tuesdays and Thursdays) and 27 November (Fridays and Saturdays), each with a morning and an afternoon group, Cyprus time. Places are €350, or €280 at the reduced rate. Dates and applications are at aidgpt.org/open-cohorts.

Over to you

  • Where have you seen a human sign-off stay on paper while the real judgement moved to an AI tool?
  • If you run a review step, does your reviewer have the time and the authority to disagree? How do you know?
  • Has anyone here taken an AI use case through a Go/Pilot/No-Go decision? What made it a No?

Share what you're seeing. I'd especially like to hear from people running these tools day to day.

Tom


Thomas Byrnes is CEO of MarketImpact Digital Solutions Ltd, leads the AidGPT responsible AI training programme, co-authored the DCI AI Hub Risk Assessment Framework for Social Protection with Dr Simona Dobre and Dr Laura van der Erve, and the Hub's global evidence review on AI adoption in social protection with Avril James.

How I used AI on this piece. OpenAI Codex helped with source review, comparison with my earlier articles and AidGPT teaching material, structural revision and drafting. Claude (Anthropic) helped select the cases, checked the facts against the published review, framework and toolkit and the reporting cited, and helped revise the wording. No beneficiary data went into either tool. The summary scenario is illustrative, not a claim about any named system. The links between delegation practice and the framework are my interpretation, separate from the toolkit's formal requirements. I set the argument, reviewed each revision and am responsible for the final text. If you spot an error, tell me and I'll correct it publicly. More on how we use AI: marketimpact.org/how-we-use-ai.

Enjoyed this article?

This post is from Aid and Dev Dispatches, a LinkedIn newsletter with expert analysis on humanitarian reform, AI adoption, crisis economics, and the politics of aid. Join 9,000+ subscribers.

Subscribe on LinkedIn

About the Author

Thomas Byrnes is a Humanitarian & Digital Social Protection Expert and CEO of MarketImpact.