You have fifteen people. Support is whoever notices the inbox. Someone on the team has said “we should just put an AI on this,” and they're probably right. But most of what you'll read on how to do it comes from two places: vendors who want you to buy their bot, and security firms who want you to buy their scanner. Neither is writing for you.
This is the operator's version. It's what I'd tell a founder over coffee, not what I'd charge for. If you get to the end and want help with the part that's actually hard, you know where I am.
The one thing to understand before anything else
In 2024, a Canadian tribunal ordered Air Canada to honor a refund its chatbot had promised. The bot told a grieving customer he could apply for a bereavement fare after flying. The real policy said the opposite. Air Canada argued the chatbot was “a separate legal entity responsible for its own actions.” The tribunal called that argument remarkable, in the way you don't want a judge to call your argument remarkable, and ruled that a chatbot is just part of your website. You're responsible for what it says.
Every decision in this guide follows from that. The bot is you. It speaks with your authority. When it invents a policy, you invented a policy.
Air Canada could absorb that. You probably can't. Which is why the controls matter more at your size, not less.
First question: should you deploy one at all?
Not yet, if any of these are true:
- You get fewer than 200 support conversations a month. At that volume a human answering in under an hour beats a bot at any resolution rate, and the setup cost isn't worth it.
- Your help center is thin, stale, or contradicts itself. Every serious vendor will tell you this privately: the bot is only as good as the documentation it reads. If your docs are bad, you're automating the delivery of bad answers.
- Most of your tickets are judgment calls. Billing disputes, edge-case bugs, “is this the right plan for me.” Bots handle repetitive questions. If your questions aren't repetitive, you don't have a bot problem, you have a product clarity problem.
Deploy when: you're past 200 conversations a month, at least 40 percent of them are variations on the same ten questions, and you have current documentation that answers those ten questions correctly.
Don't automate what you haven't learned yet
Here's the thing most guides skip. At fifteen people, you don't have a support playbook. You have a few people answering emails and slowly figuring out what customers actually ask, what the right answer is, and where your policy has holes. That figuring-out process is the most valuable thing happening in your support inbox. It's how you learn what your product is confusing about, what your terms don't cover, and what your customers assume that you never told them.
A bot that answers those questions for you doesn't just risk answering wrong. It removes you from the conversation where you'd have learned the right answer.
So until your patterns stabilize, the bot's job is not to resolve tickets. It's to handle the small set of things that are already settled and route everything else to a human who's still learning.
Early on, confine the bot to two things:
- Your terms of service and published policies. Refund windows, cancellation terms, what plans include. Things that are written down, legally reviewed, and not going to change based on a conversation.
- Information that already lives on your website. Pricing, feature descriptions, how-to steps you've documented. If it's on a page you'd link to, the bot can say it. If it isn't, the bot shouldn't.
Everything else becomes an email. Any question the bot can't answer from those two sources, it says so and prompts the customer to email support. Not a chat handoff, not a queue. An actual email, so a human reads it, answers it, and adds what they learned to the docs. That loop is how the playbook gets written.
This feels conservative. It is. But it means the bot never invents policy, because it's only allowed to repeat policy that exists. And it means every question the bot can't handle becomes a data point for what you need to document next.
You expand the bot's scope when a question has been asked enough times, answered consistently enough, and written down clearly enough that automating it is safe. Not before.
The resolution numbers vendors don't lead with
Vendor marketing says 67 to 86 percent of conversations resolved autonomously. Independent testing of production deployments at small businesses lands closer to 38 to 50 percent. B2B runs another 17 to 25 points below B2C, because B2B questions are harder.
Plan around 40 percent. If you hit 60, celebrate. If you built your budget on 80, you'll be explaining a miss.
There is also a billing trick worth knowing about. Some platforms count a conversation as “resolved” when the customer goes silent for 24 hours after the bot responds. That includes the customer who gave up, the customer who emailed you directly instead, and the customer who's now posting about you. You're paying for those as successes. Ask any vendor exactly what counts as a resolution before you sign.
What actually breaks
Five failure modes account for most of the damage. Each one has happened publicly to a company with far more resources than you.
1. It invents policy. The Air Canada case. The bot fills a gap in its knowledge with something plausible, states it confidently, and a customer acts on it. This is the most expensive failure because it creates obligations.
2. It gets talked into things. A Chevrolet dealership deployed a bot that a customer convinced to agree to sell a $58,000 truck for one dollar and to call the offer legally binding. The customer wasn't a hacker. He just asked nicely and persistently. Any bot that can be steered by conversation can be steered by a motivated stranger.
3. It changes behavior after an update. DPD's delivery bot started swearing at customers and writing poems about how bad DPD was, after a software update removed a filter nobody realized was load-bearing. The screenshot got 800,000 views in a day. Your bot's behavior is not fixed. Every model update, prompt change, or knowledge base edit can shift it.
4. It handles data badly. A customer pastes their card number asking the bot to look something up. An unprotected bot may echo it back, log it in plain text, or store it somewhere it shouldn't. This is a compliance problem before it's a customer problem.
5. It never hands off, or hands off badly. The bot can't help, doesn't know it can't help, and keeps trying. Or it hands off to a human queue nobody's watching. The customer's experience is talking to a wall, then talking to a different wall.
The controls, cheapest first
You don't need all of these on day one. You need them roughly in this order.
Fix the documentation before the bot exists. This is free and it's the single highest-leverage thing you can do. Audit your help center. Delete anything stale. Resolve anything that contradicts something else. Write down the ten questions you actually get and make sure each has one correct, current answer. A bot reading clean docs is dramatically safer than a bot reading a mess.
Scope it narrow. Early on, that means terms and website content only, as above. Later it means your product and nothing else. Not the weather, not competitor comparisons, not legal advice, not what it thinks about your pricing. Most platforms let you set this. Set it aggressively. A bot that says “I can't help with that, but a person can” is doing its job.
Give it a hard rule on money. No refunds, credits, discounts, cancellations, or anything involving a dollar amount without a human confirming. The bot can explain the policy. It cannot execute it. This one rule would have prevented both the Air Canada and Chevrolet incidents.
Make the escalation path real. “Talk to a human” has to go somewhere a human will see it within your stated response window. Test it yourself. Then test it at 11pm on a Saturday. If the path is broken, the bot is worse than no bot, because it's advertising help you're not providing.
Log everything, and read it weekly. Every conversation, every escalation, every place the bot said “I don't know.” The weekly read is where you find the question it's getting wrong, the topic it should be refusing, and the gap in your docs. Thirty minutes a week. Non-negotiable.
Name an owner. One person is responsible for the bot the way one person is responsible for the website. When it says something wrong, they hear about it first and fix it. If nobody owns it, it will drift, and you'll find out from a customer.
Re-test after every change. Model update, prompt change, new help article, new integration. Run your ten core questions through it and check the answers. Ten minutes. The DPD incident was a change nobody tested.
What to actually run at your size
As of late 2026, roughly:
Under 300 conversations a month, mostly FAQ: A flat-rate tool that reads your help center and answers from it. Tidio Lyro sits around $79 a month flat. Chatbase around $49 for pure FAQ deflection. These are limited, which at your stage is a feature.
Already on a CRM or helpdesk: Use the AI bundled with it before buying anything else. HubSpot's Breeze comes with Service Hub. Freshdesk includes Freddy on its Growth plan for about $15 a seat. You already have the data there, the integration is done, and the incremental cost is low.
300 or more conversations a month with real routing needs: This is where Fin (the platform formerly called Intercom) and Zendesk start to make sense, and where the pricing gets serious. Conversation-based pricing is cheap at low volume and steep at high volume. Documented bills have jumped from $4,000 to $9,000 a month as usage grew. Model it before you commit.
Building your own on Claude or another model: Only if you have an engineer who wants to own it long term. It's more flexible and can be cheaper at scale. It's also a product you now maintain forever. Most companies under fifty people should not.
The pattern across all of these: start with what you already pay for, move up only when you hit a wall you can name.
A 30-day starter plan
Week 1. Audit and fix the help center. Write down your ten most common questions and the correct answer to each. Separate them into two lists: the ones already answered by your terms or your website, and the ones you're still working out. Pick an owner.
Week 2. Deploy on the cheapest option that fits your volume. Scope it to the first list only: terms, published policy, and information that's already on a page. For everything else, the bot says it can't help with that and prompts an email to support. Set the hard rule on money. Test the email path.
Week 3. Run it on 25 percent of traffic. Read every transcript. Read every email it generated. The emails are your curriculum: they tell you what customers ask that you haven't documented yet. Fix what the bot gets wrong. Start writing down answers to the second list.
Week 4. Expand to full traffic on the same narrow scope. Set the weekly review on the calendar. Set the re-test checklist for any future change. Do not widen the bot's scope yet.
At the end of 30 days you'll know your real resolution rate on the settled questions, you'll have a list of the unsettled ones with real answers forming, and you'll know whether the tool you picked is the right one. Widen scope one question at a time, after each one has a stable written answer.
When to get help
This guide covers the version that works for a small team with a clear product and a clean help center. It gets harder when:
- You're running multiple products or a marketplace with two sides
- Support touches billing systems, and the bot needs to look up account state
- You're B2B with long relationships where a bad bot interaction costs a renewal
- You're about to scale support headcount and need to decide what the bot handles versus what people handle
- You've deployed something and it's producing incidents you can't diagnose
That's the operations problem, not the tooling problem. It's the part I do.
Nicholas Martin is an operating partner who works with Series A to C founders on the infrastructure that makes companies scale. He has built support and BI systems at companies from seed through acquisition. Start a conversation.
How this was made: the ideas, positions, and examples come from my operating experience. I worked with AI to research, draft, and organize, then edited and verified the result. That's how I'd expect any operator to use the tool, and it's how I use it in client work.
Start a conversation