InfoGov

The Ada Lovelace Institute has called on the government to require local authorities to record their use of AI transcription tools under the Algorithmic Transparency Recording Standard and to introduce mandatory markers on AI-generated care records, after research across 17 councils found hallucinated and inaccurate content entering statutory social care documentation with social workers acting as the sole safeguard.

Lara Groves, senior researcher at the institute and co-author of the report, Scribe and prejudice?, said policymakers should be looking at how technology can help public sector workers, but that the risks the tools introduce were not being fully assessed or mitigated and were being left to frontline staff to manage.

Imogen Parker, the institute's associate director for society, justice and public services, said structures were needed to assess the tools in practice rather than leaving individual social workers to work out their responsible use on their own.

The report, written by Groves and Oliver Bruff, is based on 39 interviews conducted between March and October 2025 with social workers and senior digital, IT and service delivery managers at 14 English and three Scottish local authorities, alongside interviews with vendors and a natural language processing expert.

One tool, Beam's Magic Notes, was already in use by 85 councils in early 2025 and its developer reports working with around 100. Many interviewees were instead using Microsoft Copilot through their council's existing licence, often outside any dedicated pilot. The institute's recommendations are:

  • Government should extend the mandatory scope of the Algorithmic Transparency Recording Standard to require local authorities to report their use of AI transcription tools, and develop mechanisms such as mandatory "watermarks" so decision-makers can identify when a care record was produced with AI
  • Government should fund coordinated pilots of AI transcription tools across multiple sites and public sector contexts, using qualitative as well as quantitative evidence gathering
  • Government should establish a What Works Centre for AI in Public Services to synthesise evidence from pilots and set best practice for evaluation
  • Researchers, policymakers, civil society and community groups should collaborate on research into the systemic impacts of the tools, with ring-fenced funding
  • Social care regulators and local authorities should produce guidance on the use of the tools in statutory processes and formal proceedings, with clear accountability structures and an advisory board of people with lived experience of care
  • Government should produce context-specific evidence on AI's contribution to cost savings and productivity rather than extrapolating across settings
  • Local authorities should specify their theory of change when procuring the tools so that staff and service users share an understanding of what they are for

The UK GDPR accuracy principle at Article 5(1)(d) requires personal data to be accurate and, where necessary, kept up to date, and Article 35 requires a data protection impact assessment where processing is likely to result in high risk to individuals.

The report noted that managers cited DPIAs as a piloting commitment but that AI-specific or equalities impact assessments were far less commonly reported. The ATRS is currently mandatory only for central government departments and arm's-length bodies delivering frontline services, not for local authorities, so councils' use of these tools is not systematically recorded anywhere.

Interviewees described a range of inaccuracies in AI outputs. One social worker recounted a summary that stated a client had expressed suicidal ideation when no such discussion had taken place, and said that had it entered the case note unchecked it could have affected the person's care.

Others reported transcripts that substituted unrelated words for descriptions of domestic conflict, frequently misspelt names, produced "gibberish" for regional accents and dialects, and generated summaries in academic, formal language that was less person-centred than manually written records.

One social worker said the shift in register mattered because a record has to be readable and relatable for the person it concerns in case they make a subject access request. The report also cites a study of foundation models summarising long-term care records from an English council, which found some models consistently downplayed women's health needs relative to men's.

The institute found that vendors and local authorities had designed rollouts on the basis that the social worker is fully accountable for anything entered into the case management system, with one manager saying whatever is submitted is audited against the worker's name. Because that accountability is assumed from the outset, the report says, few councils had put additional quality assurance in place, and management of the tools' technical risks rested on individuals with varying training and widely differing perceptions of risk.

Some social workers described extensive checking regimes; a minority said there was nothing to lose from using the tools. Managers were similarly divided, with one describing a transcription product as "light-touch generative AI" that should not be inventing words.

Evaluations were focused almost entirely on efficiency. Every manager interviewed prioritised measuring time saved or cases completed, and one said their pilot's metrics had shifted to usage frequency, which told them nothing about the quality of work. Only two councils examined the effect on the quality of documentation; one, applying its normal dip-sampling quality assurance, found the tool often did not raise overall quality and may have introduced unchecked hallucinations.

Only one council involved people who draw on care in its pilot. Managers said they had generally not used the government's published guidance on AI evaluation, and the report found no evidence of councils running the counterfactual-based tests that guidance recommends.

The report recorded specific concerns about general-purpose tools. One manager said Copilot's content filtering refused to transcribe descriptions of abuse, material social workers deal with daily, and another said very few tools offered a level of data protection suitable for the sensitive data social workers process, judging that Magic Notes met that bar but Copilot did not.

Social workers using Copilot without council guidance said they did not know who they would turn to if something went wrong, and one said the introduction of the tool had never been raised in regular meetings with the director of children's services.

There was no consensus across councils on where the tools should not be used. Some had banned their use for statutory documents such as Care Act assessments while others permitted it, and in the absence of a common position individual social workers were deciding for themselves, particularly where they were experimenting outside an approved pilot. The report also notes that AI transcription tools used in healthcare are classified as Class I medical devices by the MHRA, but the same tool deployed in social care attracts no equivalent classification and falls only within the general remit of the Care Quality Commission or Care Inspectorate.

Almost all social workers interviewed reported meaningful benefits, including being more present in conversations, capturing more detail, improving work-life balance and reducing waiting lists, but the report found these were unevenly distributed. Some said time saved was simply absorbed by additional allocated work, some found their council's complex forms could not accommodate AI outputs, and one said managers now expected longer, more detailed assessments, so that correcting AI drafts took more time than transcription saved. One team manager said they intended to raise monthly visit targets on the strength of observed time savings and performance-manage staff who did not meet them.

Scribe and prejudice? Exploring the use of AI transcription tools in social care: https://www.adalovelaceinstitute.org/report/scribe-and-prejudice/

InfoGov Masthead Newsletter 800