5 questions your AI booking agent CAN’T answer — and how to catch them before your patient does

Subtitle: In 40,000 simulations, humans let 33% of malicious AI agent actions slip through. These 5 questions reveal whether your booking system has the same blind spot.

If you only have 1 minute, read this

A study published this week on Hacker News analyzed 40,000 simulations where AI agents executed real actions. The humans supervising approved 1 in 3 malicious actions without noticing. If your clinic, workshop, or accounting firm uses an AI agent for booking, reminders, or automated replies, you need to know whether your system passes these 5 questions. Because the problem isn’t that AI fails — it’s that nobody checks when it does.

The stat that should make you uncomfortable

During the week of August 6, 2026, a thread on Hacker News gathered hundreds of comments debating a Scalex study: 40,000 executions of AI agents with human oversight. The result was that humans caught fraud or error in only 67% of cases. A third slipped through without anyone noticing.

What does this have to do with your dental clinic or auto repair shop? Everything. Because more and more Spanish SMEs use AI agents — WhatsApp bots, appointment confirmation systems, automated query responses — without a single protocol to verify what that agent does when nobody’s watching.

We’re not talking science fiction. We’re talking about the bot that confirms an appointment for Tuesday when the doctor only works Monday and Thursday. The reminder that sends the old consultation price. The automated reply that says “yes, we accept your insurance” when they stopped accepting it six months ago.

Question 1: What happens when your agent doesn’t know the answer?

This is the most important question and the one fewest people ask. When a patient writes on WhatsApp “do you offer pediatric physiotherapy?” and your AI agent doesn’t have that information in its database, what does it do?

There are three possible behaviors, and only one is correct:

  • Invents an answer. Says yes, they offer pediatric physiotherapy, when that’s not true. The patient arrives, gets frustrated, never comes back. This failure is silent — you never find out.
  • Doesn’t respond. The message is left on read. The patient assumes you’re unprofessional and calls someone else. Also silent.
  • Hands off to a human. “I don’t have that information, but I’ll connect you with María who can help in 5 minutes.” This is the correct behavior.

If you don’t know which of the three your agent is configured to do, you already have your first red flag.

Question 2: Does your agent know what it can NOT say?

An AI agent managing appointments at a clinic handles health data — the category most protected by GDPR. If your bot confirms a dermatology appointment in a WhatsApp message that also includes the patient’s full name and time, it’s transmitting health data to Meta’s servers.

The question isn’t whether your agent is useful. The question is whether it knows what information it shouldn’t include in an open message. Does it say “your oncology appointment” or just “your Tuesday appointment”? Does it mention the treatment or just the date?

«The processing of special categories of personal data, such as health data, is only permitted when the data subject has given explicit consent to the processing.» — GDPR, Art. 9

If your agent doesn’t have clear rules about what data it can mention on each channel, you’re one AEPD audit away from finding out.

Question 3: Did anyone review what the agent did yesterday?

In the Scalex study, 33% of incorrect actions went through because nobody reviewed them. In your business, the equivalent is: does your team review every morning which appointments the bot confirmed, who it replied to, and what it said?

Most SMEs using AI agents have a dashboard that nobody opens. The bot works “on its own” — until it confirms a double booking, sends a reminder with the wrong price, or responds to a patient who’s no longer yours.

A system without daily review isn’t an automated system. It’s an abandoned system. The difference between a useful agent and a dangerous one isn’t the technology — it’s whether someone opens the report every morning before the first appointment.

Question 4: What does your agent do when two patients want the same slot?

The most advanced booking systems have a window of between 2 and 15 seconds where two patients are booking the same slot. If your agent doesn’t have a priority rule — first to confirm gets it, or VIP patients get preference — you’ll end up with double bookings.

The problem isn’t technical. It’s that nobody defined the rule before building the agent. And when the conflict arrives, the bot doesn’t know what to do, so it does the worst thing: confirms both appointments and tells nobody.

Question 5: How much money do you lose when the agent fails?

If your agent confirms 40 appointments a week and fails in 5% of cases (2 appointments per week), and each unmanaged appointment costs an average of €45 in lost revenue, you’re losing €360 per month — more than €4,300 per year — in failures you don’t even know are happening.

The calculation is simple. The hard part is doing it. Because most SME owners don’t know how many appointments their agent handles per month, how many it fails at, or how much each failure costs. If you don’t know, your agent is operating blind.

Your Quick Win for today

Open your AI agent’s dashboard (or ask your provider to send you last month’s report). Note 3 numbers: (1) how many interactions it handled, (2) how many were handed off to a human, and (3) how many went unanswered. If you don’t have those three numbers, your agent is operating unsupervised. You have 10 minutes to request them.

Does your AI agent answer these 5 questions? If you’re unsure about any of them, it’s best to review it before a patient makes the failure visible.

Request Free Diagnosis