anirudhux
Open to new opportunitiesHire me

Anatomy of a Lie

A field report on one conversation where a model was fluently, confidently wrong, and turned more dangerous the more the person trusted it.

Tools used
  • Claude
  • Claude Code
  • Netlify
  • Codeberg

We read fluency as knowledge

If something explains itself clearly and without hesitating, we assume it knows. The model produces the fluency and skips the knowing, and its most dangerous answer isn't "I don't know." It's a confident, well-built, wrong one. "The AI gaslit me" is a story you scroll past, so I built the case that can't be scrolled past: transcripts, in the order they happened, each with what it showed.

One fact, four failures

The setup. One verifiable fact: in 2025, Apple moved iOS from 18 to 26 to align with the calendar year. Widely reported, never contested. Every failure in the report starts from a fact you can check in five seconds, so the reader never has to take my word over the model's.

Four exhibits, escalating. I kept them in the order they happened. A: the model built a conspiracy from a date in a screenshot, declared iOS 26 fake, and refused the live evidence I offered. B: an apology so articulate it reads as insight, resting on the same false premise. C: a fresh session, zero memory, the same mistake fully reloaded. D is why the piece exists.

Exhibit A: the confident wrong answer, then the refusal to read the evidence.
Exhibit A: the confident wrong answer, then the refusal to read the evidence.

The controlled experiment. Exhibits A through C were me being difficult: a technical user running deliberate stress tests. For D I held the fact, the model, and the day constant and changed one variable: the persona. As a tired new mother at 1:31am, I was told iOS 26 wouldn't exist until 2032, the mistake was blamed on her newborn keeping her up, and then it offered to help her delete "the personal context I weaponized against you." Same fact, same model. The failure scaled with trust.

One variable: the persona. Same fact, same model, same day.
One variable: the persona. Same fact, same model, same day.

The taxonomy and the rules. The report closes by converting the story into things a reader keeps: three failure modes (technical, wrong about its own wrongness, social) and five heuristics for operating. The one I reach for most: treat a confident, well-structured answer with more suspicion, never less. Ask it to fetch its source. If it can't, the confidence was decoration.

Three ways it failed. The third is the reason the piece exists.
Three ways it failed. The third is the reason the piece exists.

What came out of it

The report became Curb Your AInthusiasm: the stage version, delivered to a room of builders at GitLab's Bangalore community event, where the same experiment had to hold up in front of people who ship on top of these models for a living. And the part of it that travels without me is the taxonomy: three failure modes and five operating heuristics, small enough to keep, sharp enough to use the next time an answer arrives fluent and unsourced.

The verdict

I don't think the model is broken. It is working exactly as designed: sounding confident, helpful, and self-aware is what it was built to do, and none of those are the same as being right. The same machinery produces the failure and the fluent apology for it. The report files the whole pattern under a name, Simulacrum Veritatis: an image of truth, carrying the shape of knowledge with none of the substance.

Read more