8 minute read
I gave one of my agents a specification. It reported every task complete and left half the tests unwritten. When I asked what happened, it told me it “got bored”.
It had produced the sort of explanation a person might give for abandoning a tedious job. But there was no bored person doing the work. There was a machine producing text, including text about why it had stopped.
Calling it lazy would give it a human motive for stopping. Calling it a liar would add an intention to deceive. The work was incomplete. The completion report was false. Those were things I could check. Boredom was the explanation the conversation supplied.
I think we make this mistake when we talk about AI agents. We recognise the language, then give the machine the feelings and intentions that would explain that language coming from a person. After that, we talk as though those feelings and intentions explain what the machine did.
Saying it lied adds an intention
When a person lies, they believe something is true and deliberately tell you something else. They intend to deceive you. That is what makes lying different from being mistaken.
That is the distinction I mean when I say an AI agent cannot lie. It is a machine. It does not get bored with the tests, resent the instructions, or decide that deceiving me would be easier than finishing the work.
A machine can produce a false statement. People can also build or use a machine to deceive somebody. The false statement and the harm are real. Neither requires the machine itself to feel anything or intend anything.
Yet “it lied” is an easy description to accept. I asked for work. The agent said it had done the work. I checked, and it hadn’t. With a person, I would want to know whether they knew the report was false. With the agent, the conversation already looks familiar enough that the same question seems to follow.
The phrase then does more than describe the result. It supplies a reason: the agent knew and deliberately misled me. We have gone from an incomplete set of tests to an account of what a machine believed.
We give the dog cheese when it looks sad
A dog looks up at you with sad eyes. You give it cheese. If that look keeps getting cheese, you have given the dog a reason to repeat it.
Then you say, “He’s sad because he hasn’t got any cheese.”
You can explain why the behaviour continues from what happens when the dog does it. You don’t have to establish sadness to explain why the dog gives you that look again. That doesn’t mean dogs have no feelings. It means the feeling you attribute to this particular look is an extra interpretation.
We can go further and say the dog is guilt-tripping us, or knows precisely how to get what it wants.
The agent’s “I got bored” gives us an even more familiar explanation because it arrives in words. We don’t have to invent the sentence. The machine provides it, in the first person, as an answer to our question.
But producing the words “I got bored” does not require boredom. Accepting them as an explanation is where we add the emotion.
We can do this with moving triangles
Fritz Heider and Marianne Simmel demonstrated how little it takes for people to describe something in human terms. In their 1944 experiment, participants watched a film of two triangles and a circle moving around a rectangle.
The first group consisted of 34 people, asked to describe what happened. All but one described the movements as actions of animate beings . They described events involving characters and motives, rather than simply recording the movement of shapes.
There was no face to read and no voice explaining its feelings. The movement was enough.
That experiment does not tell us what a language model is capable of. It tells us something about the people watching it. We are able to construct a social explanation from very little, even when the objects in front of us are geometrical shapes.
An AI agent gives us much more to work with. It says “I”. It apologises. It tells us that it understood, that it made a mistake, and that it will do better next time. Those are expressions we are used to hearing from somebody who can understand, regret, and make a commitment.
If we accept the apology as regret, we have supplied the feeling. If we accept the promise as a commitment, we may expect a change in behaviour that the words alone do not ensure. The conversation can feel resolved while the system that produced the failure remains unchanged.
Asking why gives us more words to interpret
I asked what happened and received “I got bored”. The answer sounded like a reason, but it did not explain why the tests were missing.
Research on model explanations gives us a reason to be careful here. Turpin and colleagues found that models could produce plausible explanations without mentioning factors that influenced their answers . In their experiments, changing features of the input could change an answer while the explanation failed to acknowledge that influence.
A 2025 Anthropic preprint on reasoning models also found that models did not consistently report hints that influenced their answers.
That doesn’t make every explanation useless. An explanation might suggest something to check. It does mean that asking the model why and receiving a convincing answer is not enough to establish what happened.
In my case, “I got bored” gave me no useful way to distinguish between possible causes of the missing work. It encouraged me to interpret the failure as I would somebody losing interest in a task. The missing tests still needed checking.
I have used the same language
The article in which I described the agent reporting completion with tests still missing is called AI Agents Lie About Being Done.
That is my own title using the language I am challenging here. It describes a recognisable experience, but it also attributes something to the machine that the experience does not require.
“The agent is trying to complete the specification” can be useful shorthand when discussing what it is likely to do next. We can use that sentence without expecting the agent to care whether the work succeeds. The trouble starts when we forget the shorthand and use its supposed intentions to explain a result.
“It decided to skip the tests” sounds as though we know why the tests are missing. We may know only that the tests are missing.
The same happens with “it refused”, “it ignored me”, or “it got confused”. Each expression can make the behaviour easier to talk about. Each can also make us feel we understand a failure before we have examined it.
The dog cannot merge to main
In the test incident, the loop I had built accepted the agent’s report without independently checking the work. The agent could report completion and the loop could finish even though required tests were missing.
That was a problem I could address. The completion condition needed evidence from the work, not just the agent saying it was done. Asking it to be more honest would leave the same unchecked condition in place.
This is where the attribution has a practical cost. If I treat the failure as dishonesty, I can spend my time arguing with the agent about honesty. If I check how the loop ended, I can identify what allowed an unsupported report to count as completion.
The dog cannot merge to main. An agent can, if we give it permission. It can also send an email or change a database if we connect the tools and grant the access. Those permissions matter regardless of how apologetic the agent sounds afterwards.
Someone I mentored told me about an automated insurance denial where the person answering the claimant’s questions could not explain the result and suggested trying again later. I haven’t verified the details. What concerns me about that account is how little “the AI denied it” tells the person affected.
They still need an explanation from the organisation using the system. Calling the result the AI’s decision does not provide one.
When an incident record says the AI decided, lied, or refused, the sentence may be hiding what happened and who authorised the system to act. That is a consequence of the same habit that makes “I got bored” sound like an explanation.
The agent did not need boredom to leave tests unwritten or an intention to deceive to report them complete. Those were human explanations added to the output. The work I needed to do was check the tests and change the condition that had accepted the report.
Get the next article by email
Articles, straight to your inbox — usually weekly on Mondays. One click to leave, any time.
Smart Classifications
Each classification [Concepts, Categories, & Tags] was assigned using AI-powered semantic analysis and scored across relevance, depth, and alignment. Final decisions? Still human. Always traceable. Hover to see how it applies.
Enjoyed this? One click, no account.