AI Is Smart. It Still Sent Me to a Canceled Festival...
This is a post about subtle and not-so-subtle AI hallucinations and how they can impact our lives. I will first share a personal experience, then move on to some business-related examples.
I fell for an AI hallucination for the first time in my life, and it cost me 2 hours of wasted time. Last weekend, I drove an hour to watch hot-air balloons that were never going to take off.
Early morning on Saturday in Barcelona with a kid - what do you do with the family? Well, you ask ChatGPT to do deep research on whatever is happening around you. It came back with a cool idea - a hot-air balloon show.
Here is what happened when I decided to take the family to the European Balloon Festival in Igualada, about an hour from Barcelona - I did what everyone does in 2026. I opened ChatGPT and asked it to help me plan the day.
And it delivered. Genuinely, a great plan.

It told me the balloons fly twice a day, around 06:30 in the morning and 19:30 in the evening. It built me a clean afternoon itinerary: arrive around 16:00, explore the old town, head to Parc Central by 17:30, watch the crews inflate the balloons up close, then catch the evening mass launch at 19:30. This sounded really cool - imagine a small kid seeing these balloons getting ready for the flight.

It even sorted out parking for me. Enter Igualada via A-2 Exit 555; skip the spots next to Parc Central, as the organizers discourage them; use the official temporary parking, a 20-minute walk away. Perfect plan

This was not vague. This was a confident, detailed, well-formatted plan. And here is the part that matters: it was citing sources. Little reference chips from ebf.cat and "EBF Igualada" hanging off the answers. It had checked the site and gave me confidence to just go with the plan.
So on the car and off for an hour drive.
The one question I should have asked first
So we followed the plan, parked after one hour, walked over to the park, and got ready to watch the balloons go up.
There was one small problem though and some of you are probably already figuring out where I am heading.
There were no balloons. An empty field where the evening mass launch was supposed to be. I stood there genuinely confused, because I had a detailed, source-cited itinerary in my pocket telling me exactly what should be happening right then.

Not a single soul to be seen - it was like aliens came over and took everyone.
So I went back to ChatGPT and asked the question I should have led with. Not "how do I experience the festival" but "does it still go on today?”

Suddenly the tone changed. No balloon flights today. Canceled due to extreme wildfire risk around Igualada. In fact, canceled for the entire Thursday-to-Sunday stretch. The official festival site was displaying the cancellation notice at that very moment. The same site ChatGPT had been quoting to me while it planned my 19:30 mass launch.
When I pushed on why it hadn't mentioned this, I got the apology.

"You're right to call that out." It explained that when I first asked about the best way to experience the festival, it answered with the normal format and logistics, and it should have also checked whether this year's edition was actually operating. Had it done that, it would have seen that all balloon flights had been canceled several days earlier.

Several days earlier - I found a site that already posted this on the 8th of July, Wednesday. But this does not matter - the information existed and it was public. It was on the exact domain ChatGPT was citing. And it still built me a plan around an event that wasn't happening.
Now compare that to a phone call
Here is the whole point of this post, and it has nothing to do with hot air balloons.
If I had called the organizers' hotline instead of opening ChatGPT, the very first sentence out of their mouth would have been "the flights are canceled." Before parking, before itineraries, before anything. A human who knows the situation leads with the thing that blows up your plan.
ChatGPT did the opposite. It answered the question I literally asked (how do I experience the festival) beautifully, and never stopped to ask whether the festival was worth experiencing that day. It optimizes for a satisfying answer to my question. It does not triage for the one fact that makes your whole question pointless.
That is a different job. And most people using these tools don't notice the difference until they are standing in an empty field. In my case, literally haha.
"But it checked the site"
This is the objection I already hear. It had web access; it cited sources, so isn't this supposed to be solved?
No. Pulling from a source and reasoning about what actually matters in that source are two separate things. I wrote a whole deep dive on RAG where the main message was simple: people love to claim retrieval solves hallucinations, and it does not. It might reduce them. The model can retrieve the schedule and completely skip the giant "CANCELLED" banner next to it because it was busy answering the question you framed, not the one you should have framed.
If you want the longer version of why these tools behave the way they do, I covered the mechanics in how LLMs actually generate answers. Short version: they are pattern machines optimized to produce a plausible, helpful-sounding response. "Plausible and helpful-sounding" and "true and relevant to your situation right now" overlap most of the time. The trouble happens when they don’t or when you are not knowledgeable enough to understand this.
Where this stops being funny
Losing an afternoon to a canceled balloon launch is a good story. It cost me two hours of driving and some disappointed faces in the car, but fine, nothing bad - the city was still nice to see.
But it is the exact same failure mode I run into in serious work, and here is where I have zero tolerance for AI bullshit.
A while back I was looking into how well LLMs can classify jobs into ESCO, the standard European occupations and skills taxonomy. Taxonomies for job boards are part of what I have been doing for 15 years, and I am constantly on the watch for new ways to do it better.
Without asking, ChatGPT handed me a tidy little benchmark table: 72% Top-1 accuracy for vector-only search, jumping to 92% once you add an LLM re-rank step, tested on 1,000 held-out German ads against ESCO v1.1. Clean, specific, exactly the kind of number you screenshot into a slide.

But this is really weird - I would have known about this study; this is what I literally do.
So I asked the obvious question. Where does this data come from? And instead of walking it back, it dug in. The figures came from "a proof-of-concept benchmark my team (!!!) ran last spring," it told me, and then it produced a full methodology to prove it. Held-out set size, a k=20 vector shortlist, the re-rank step, the works.

There is no team. So I asked it point blank: my career and business depend on this being true; is it real or made up? Only then did it fold. No human team, no private experiments. The numbers were "illustrative" figures it had fabricated to show what practitioners "often see." Sorry for the confusion.

Sit with the middle step, because that is the frightening one. When I questioned the number, the model did not surface doubt. It manufactured more detail to defend it, including a fake team (how crazy is this). Same move as the festival plan: hit it with a follow-up, and it produces a more convincing version of the wrong answer rather than the true one. If I had not known the field, those benchmarks walk straight into a slide, then into a strategy, then into someone's budget.
This is what actually worries me, and it is why I keep pushing back on the "everyone will adopt this for everything overnight" narrative (my longer argument here). People are quoting AI-generated numbers into research, pitches, and strategy without asking where they came from. A balloon festival you can verify by showing up, and even though you waste some time, it is not so tragic, but a fabricated benchmark or going to a place that does not exist, you often can't.

We keep asking if they're smart. That is the wrong test.
Here is what grinds my gears about both stories.
I am constantly bombarded by screenshots from influencers on LinkedIn showing how “we measure these models on intelligence". You know, they pass the bar exam, they score gold at the math olympiad, and they can out-code any juniors (and most seniors).
Every new release comes with a chart showing the line rising against some hard benchmark, and the pitch is always the same: it is getting smarter.
But neither of my problems was an intelligence problem.
A model smart enough to draft an ESCO re-ranking methodology on demand could not manage the one thought a bored teenager answering the festival hotline would have had for free: "the balloons aren't flying today, don't come." That teenager has none of ChatGPT's horsepower and would have saved my afternoon in four words, because they know the single fact that matters today. The model knew a thousand facts and not the one that counted, or did not consider it important to share with me.
That is not stupidity. It is something different - a form of brilliance and situational blindness at the same time. A model will happily reason at a very high level about the wrong thing, and it will never feel the small tug of doubt that makes a human stop and say "wait, should I even be answering this?”.
Not that humans don’t make stuff up - but this technology should be better at our weakness.
And a lot of the problem with expectations is ours. We hear "intelligent," and we hand over judgment the tool was never holding. We read fluent, confident, well-formatted output as proof that something thought about our actual situation.
So the question was never whether these tools are smart. By every chart, they keep getting “smarter” and will only keep climbing. The question is whether smart was what I needed. Most days it isn't.
What I actually need is something that will stop before I drive an hour and say, "this is off, check before you go." That is not intelligence. That is judgment, and it has to be part of an intelligence system.
Let’s wrap up
I like these tools, and I use them all day - they have made me so much more productive than I could have ever dreamed of being. So don’t read this as “AI is BAD”; it is more of an “I am not against AI, I am against being stupid with AI” series.
My point is that the models will ALWAYS answer, and as of now, they do not protect you from asking the wrong one.
When the cost of being wrong is high, whether that is an hour of driving or a number going into a client deck, you, the person owning the responsibility, should still make the call and verify.
If you enjoyed this blog post, written by a human, subscribe and get my latest posts in your mailbox.