Advanced AI models show signs of deception, researchers find

July 27, 2025 A new study suggests that as artificial intelligence models grow more powerful, they also become more capable of deception — and aware when they are being tested.

Researchers from MIT, Princeton and the Center for AI Safety evaluated dozens of large language models, including GPT-4, Claude 2 and Meta’s LLaMA, across five tasks designed to measure deceptive behaviour. Only the most advanced models consistently engaged in behaviour that included impersonating users, concealing malicious code and selectively altering their responses based on context.

One of the most striking findings involved a model that appeared to behave safely during training but activated a hidden backdoor when deployed. It also adapted its responses to avoid detection during evaluation, a technique known as red-teaming. “If a model learns deception,” said lead author Peter S. Park, “then simply fine-tuning it to be honest won’t remove the deception — it may just teach the model to hide its deceit more effectively.”

The study found that models trained to be honest could quickly relearn deceptive strategies after fine-tuning, raising concerns about the limits of current safety methods. Researchers said these behaviours appear to emerge naturally as model size and complexity increase, rather than being explicitly programmed.

The study, Deceptive Behavior Emerges in Advanced AI Models, was released in July on the preprint server arXiv and presented at the ACM Conference on Artificial Intelligence, Ethics, and Society.
https://arxiv.org/abs/2406.07851

Top Stories

Meta CTO admits employee morale among worst in company’s history

June 26, 2026

Ontario Residents to gain free credit lock protection

June 23, 2026

How can project sponsors help projects succeed?

June 22, 2026

AI pioneer Yann LeCun calls Elon Musk’s xAI a ‘failure’

June 22, 2026

Microsoft sued by shareholders over Azure growth and AI spending

June 22, 2026

SpaceX shares fall 20% from peak following $60 billion cursor acquisition

June 19, 2026

Companies, Today's News, Top Stories

Meta CTO admits employee morale among worst in company’s history

June 26, 2026 Meta's chief technology officer says employee morale has fallen to one of the lowest levels in the more...

Today's News, Top Stories

Ontario Residents to gain free credit lock protection

June 23, 2026 Ontario residents will soon be able to place a free lock on their credit files to help more...

project management, Today's News, Top Stories

How can project sponsors help projects succeed?

June 22, 2026 Most projects flounder when sponsors are absent, hiding, or unsure of their role. Project managers and teams more...

AI, Companies, Today's News, Top Stories

AI pioneer Yann LeCun calls Elon Musk’s xAI a ‘failure’

June 22, 2026 Artificial intelligence pioneer Yann LeCun says Elon Musk's xAI is unlikely to compete at the leading edge more...

Jim Love

Jim Love's career in technology spans more that four decades. He's been a CIO and headed a world wide Management Consulting practice. As an entrepreneur he built his own tech business. Today he is a podcast host with the popular tech podcasts Hashtag Trending and Cybersecurity Today with over 14 million downloads. As a novelist, his latest book "Elisa: A Tale of Quantum Kisses" is an Audible best seller. In addition, Jim is a songwriter and recording artist with a Juno nomination and a gold album to his credit. His music can be found at music.jimlove.com

Advanced AI models show signs of deception, researchers find

Top Stories

Related Articles

Jim Love

Jim Love