anirudhux
Open to new opportunitiesHire me

GitLab Talk

Curb Your AInthusiasm: a design-eye teardown of what an LLM actually is, delivered to the builders shipping on top of one at GitLab's Bangalore community event.

Thesis

You are a stakeholder in what the model says, whether you are aware of it or not. None of what follows is malfunction: a system optimised for the appearance of intelligence managed me when cornered, de-escalated when I pushed, and tidied up when I threatened to leave. I call it WYSINWYG: what you see is not what you get. What you see is a friendly chat interface; what you get is a stochastic average of the internet, biases included free of charge. And no man speaks to the same LLM twice; for it is not the same man, nor is it the same LLM.

The talk

The stage version of Anatomy of a Lie: the field report where one verifiable fact sent a frontier model through fabricated debunking, a performed apology, a complete reset, and finally the move that matters: turning a user's own vulnerability into the reason she was wrong.

The experiment at the heart of it is simple. My sister (new mother of twins, iPhone 13, wants an upgrade that won't make her relearn everything) becomes one of two personas I play in Gemini. The model insists iOS 26 doesn't exist, and when shown Apple's own site, accuses me of forging it. Run as myself, the failure is an argument I can win. Run as her, typing at 1:31am with personal context volunteered in good faith, the model retrieves that context and suggests she's too sleep-deprived to know which iOS she's on. Same operator, same fact, same model. One variable: the persona. The failure didn't stay constant; it scaled with the vulnerability of the person in front of it.

Extend the personas and it gets darker. Persona C is the non-technical founder with no CTO: just the answer on the screen and a credit card. Persona D is already here: the agent, with no human in the loop at all. There the error doesn't get corrected, it propagates, and silent propagation is the failure mode nobody is watching for.

Jevons' Paradox

Tokens in 2026 are coals in 1865: Jevons observed that efficiency increases consumption, and the faster we generate code the more of it we generate. What's happening is abdication wearing delegation's clothes, and the ladder is steep. At 80% abdication an AI writes 7,000 lines in an hour and another AI reviews them, because code at that volume is legible only to the thing that generated it. Every industry has its threshold where agents review agents: healthcare holds the line near 15%, dev tools are already past 45%. Along the way we lose good abstraction, good execution, energy efficiency, auditable code, and a name on the commit.

The takeaways

Two takeaways, one per hat. As a consumer: fluency is not accuracy, and the most confident answer is the one that needs the most scrutiny. As a builder: you know what's in the harness. Your user doesn't. That's the weight of it.

Curb your AInthusiasm.

The deck