ComputerWeekly

AI model developers have ‘no justification’ for failing to comply with privacy law


Artificial intelligence (AI) model developers have been warned there is “no justification” for failing to comply with data protection and privacy laws when using data sets containing personal data to train large language models (LLMs).

The UK’s data protection regulator said in a report that AI model developers were likely to be processing the most sensitive types of data to train LLMs and must comply with data protection law.

The Information Commissioner’s Office (ICO) said that LLMs were trained on “vast datasets” that often include significant volumes of personal data.

This could include data scraped from the internet, including social media posts containing the most sensitive categories of personal data, such as medical information, political views and religion.

AI model developers must provide people with information about what personal data they are using to train AI models, where they got the data, how they are processing it and how people can object to their data being processed, the report said.

The ICO said it had found a lack of transparency from AI model developers about the use of personal data for training AI models, and that it had concerns some developers may be using “blanket exemptions” to avoid transparency.

No excuse for inaction

The report said: “We recognise the significant technical challenges developers face in identifying personal data within training datasets. But this isn’t an excuse for inaction.”

Developers must have a valid lawful basis for processing personal data to train AI models. Broad interests, such as “developing and improving our products and services”, “training our models” or “benefitting humanity”, are not enough.

A review by the ICO found that in some cases, AI model developers did not have effective mechanisms in place to enable people to object to their personal data being used or to ask for its deletion.

“Some developers took a blanket, one-size-fits-all approach to refusing requests or failed to provide people with enough information to understand why they refused their requests,” it said.

AI models can leak personal data

A key question is whether AI models can be said to contain personal data that will mean the models themselves are subject to data protection regulation.

There is significant evidence and research to show that AI models memorise training data, which can be leaked or exfiltrated.

This poses “serious risks to people’s privacy”, particularly when adversarial prompts are used to extract the information and can expose people to fraud, identity theft and psychological harm.

“Where necessary, we’ll use the full range of our regulatory powers to address practices that expose people to harm, helping to ensure that innovation and public trust develop hand in hand,” the report said.

AI model developers agree to changes

The data protection regulator said it had secured pledges from 10 of the biggest AI foundation model developers to make changes to improve compliance with the UK’s privacy laws.

Amazon, Anthropic, Apple, Cohere, DeepSeek, Google, Meta, Microsoft, OpenAI and Stability AI have agreed to make changes, including improving transparency and making it easier for people to exercise their rights over personal data.

The ICO has withdrawn from supervising the social media site X and X.AI after launching an investigation into the Grok AI model for processing personal information used to produce harmful “nudified” images.

The regulator also launched a six-week call for evidence on agentic AI, seeking views on how organisations are managing the novel data protection risks of agentic AI.

It follows the rogue OpenAI model’s breach of Hugging Face over the summer, which has raised concerns that AI agents have the potential to autonomously access and exfiltrate personal data unless there are sufficient safeguards in place.

The ICO has made enquiries with OpenAI, Anthropic, Meta and the UK’s AI Security Institute, following a series of reported incidents of AI agents breaching guardrails and attempting to exfiltrate data from websites.



Source link