The tools were not the problem. That is the part that still sits oddly with us.
We ran AI chatbots. We ran AI callers. They did roughly what the vendors said they would do. Enquiries got handled, contact got made, nothing broke, nothing embarrassed us. On any dashboard you care to build, deflection rate, response time, cost per contact, the numbers looked good.
We pulled them out anyway. The phone side now goes to a call centre in Australia and New Zealand, staffed by people, at a cost we could have avoided entirely.
If you are looking purely at unit economics, that is an irrational decision. We think the economics are being measured in the wrong place, and we think most businesses currently signing AI contracts are making the same measurement error.
The cost landed in a ledger nobody was keeping
What changed our minds was not a metric. It was one sentence.
A customer pointed out, plainly, that the AI was their first point of contact with our company. Not a page on our website, not an ad, not a conversation with someone on the team. A machine. And they did not like it.
There is a comfortable response available in that moment, which is to explain the efficiency case. There is a less comfortable one, which is to accept that they have described exactly what happened, and then ask whether that is how you want to introduce yourself to someone who is about to decide whether you are worth their money.
Then there was a hotel. A week-long stay, reception answered by AI. Same question asked five times across the week, five different answers. Not wrong answers, exactly. Just inconsistent enough that by day three you stop believing anything it tells you, and by day four you have formed a view about the hotel that has nothing to do with the room, the bed or the breakfast.
Nothing in that hotel's reporting would have shown a problem. The system answered every call. Average handling time was probably excellent.
That is the gap. The system performed and the relationship degraded, and those are two entirely separate ledgers. Almost every business we talk to is keeping the first one obsessively and the second one not at all.
What most businesses are actually measuring
Here is the standard evaluation for a customer-facing AI tool. Deflection rate: what percentage of enquiries never reached a human. Cost per contact: what each interaction now costs versus before. Response time: how fast the first reply goes out. Availability: whether the thing runs at 11pm on a Sunday.
Every one of those metrics improves when you install AI. That is not a coincidence, it is the point. Those are the metrics the category was built to move.
Now notice what is missing. Not one of them tells you what the interaction did to the customer's opinion of you. Not one of them distinguishes between a customer who got their answer and felt served, and a customer who got their answer and felt processed. Those two people show up identically in the reporting and behave completely differently eighteen months later.
The AI decision gets run through operations, because it looks like an operations decision. It is a brand positioning decision wearing an operations costume.
The research caught up after we had already decided
We made this call on instinct and one piece of customer feedback, which is not a methodology. The data has since arrived and it is blunter than we expected.
Gartner research reported in 2026 found that around half of US consumers would prefer to give their business to brands that do not use generative AI in customer-facing messages, advertising or content. Not neutral about. Prefer to avoid.
A Harris Poll reported in June 2026 found roughly 63 per cent of consumers said they would be less likely to buy from a brand using AI-generated advertising, and about 73 per cent said they would be less likely to trust an ad they suspected had been made with AI.
Research published by the IAB in 2026 measured the distance between how advertisers feel about AI-generated advertising and how consumers feel about it. That gap widened from 32 points in 2024 to 37 points in 2026. The people producing this work are drifting further from their audience over time, not closer.
DoubleVerify's 2026 EMEA study found 42 per cent of consumers would feel negatively towards a brand whose advertising appeared alongside low-quality AI content, against 24 per cent who would feel positively.
Two caveats, because they matter and because leaving them out would be exactly the sort of thing this article is arguing against.
First, all of that research is American, British or European. None of it is New Zealand data. We have not seen a credible NZ-specific study on this yet, and if one exists we would genuinely like to read it.
Second, none of it says consumers are anti-AI. What it consistently says is that consumers object to AI where it feels like the company saved money at their expense. That distinction is the whole game.
People can tell, and they are getting better at it
Consumer research from Klaviyo found the two most common ways people identify an AI interaction are replies that arrive too fast and language that reads as too formal or robotic. Around half of consumers named each one.
Which means detection does not get harder as the models improve. If anything it sharpens, because people are calibrating in real time against a rising volume of examples.
The counter-argument, from someone who thinks we are wrong
Darren Pratley does not agree with us on this, and his position deserves to be put properly rather than nodded at on the way past.
His argument runs roughly like this. As a customer, what he wants is efficiency and a good outcome. He has dealt with plenty of businesses where the human was the weak link, where the person on the other end failed to ask the right questions and failed to get him what he needed. If an AI were configured to ask better questions and produce a better result, he would take the AI without hesitation. He is careful to say we are not there yet. He also thinks that is the direction of travel, and that businesses betting permanently against it will look silly.
That is not a soft position and we are not going to pretend it is. It is probably correct about a large category of interaction.
Think about the last time you tried to change a booking, check an order, or find out whether something was in stock. In that moment you do not want a relationship. You want the answer, and you want it in eleven seconds rather than after four minutes of hold music and an apology for unusually high call volumes.
Darren is describing a real failure mode that AI genuinely fixes. Our disagreement is not about whether that failure mode exists. It is about which interactions belong in that category.
The line that has held up for us
The test we now use is simple enough to apply in a meeting.
Does the AI remove friction the customer never chose, or does it replace an interaction the customer came for?
Automated scheduling that ends a five-email exchange about meeting times removes friction. Nobody wanted that exchange. An AI answering your main line replaces an interaction. Order tracking that fires automatically removes friction. An AI-generated video ad replaces the part where a human made something for another human to watch.
A first conversation with a prospective client is not friction. That is the product. You are asking someone to hand you money on the strength of how you make them feel in the first ninety seconds, and then automating the first ninety seconds.
Where we still use AI constantly
We should be clear, because an agency claiming to be AI-free would be both dishonest and slightly ridiculous in 2026.
Internally, we use it heavily. Research, transcripts, turning a mapped-out process into a working system, drafting, analysis, building financial dashboards off our accounting data. There is enormous value there and none of it touches a customer relationship.
The one place we deliberately overspend is video. We keep paying human videographers when AI video is cheaper and faster, because we do not believe it is what our clients' customers want to watch. We might be wrong about that in three years. We are comfortable being wrong on that side of the line rather than the other one.
Why the downside is not symmetrical
Trust does not move evenly in both directions, and that asymmetry is what makes this decision bigger than its cost suggests.
It builds through small unglamorous acts. You pay the invoice on the day you said you would. You turn up on time. You send the email you promised to send. Individually none of those is worth anything. Repeated across two years, they are worth more than your advertising budget.
It breaks in a single interaction. And where a bad experience used to reach ten people, it now reaches hundreds, quickly, and stays searchable long after everyone involved has moved on.
So the calculation is not "does this save us money." It is "what does the downside distribution look like." A system that works nine times out of ten and produces one uncanny, cost-cutting-flavoured interaction is not a 90 per cent success rate. It is one damaged relationship per ten customers, compounding in the wrong direction, in a channel where you have no visibility of the damage.
What this changes for New Zealand businesses specifically
New Zealand is small. Everyone says it. Almost nobody adjusts their decision-making for it.
In a market of five million people, reputation does not travel through review platforms. It travels through networks. The person who has a poor first interaction with your business is statistically likely to know someone who knows you, or to sit two seats away from your next prospective client at an industry lunch. There is far less room to absorb a bad experience quietly here than there is in a market of three hundred million.
That does not make AI wrong for New Zealand businesses. It does mean the downside is concentrated in a way none of the international research above captures, and that the real cost of getting first contact wrong here is higher than those numbers suggest.
It also means the opposite is available. In a market where everyone is quietly automating the same touchpoints, being reachable by an actual person is becoming a differentiator rather than a baseline. That is an odd thing to be able to write in 2026, and we do not think it stays true forever. Right now it is true.
If someone is pitching you an AI phone or chat solution this quarter, the question is not whether it works. It almost certainly works. The question is whether the thing it replaces was ever the inefficiency you assumed it was.
Want the full conversation? Listen to "21 Business Lessons I Have Been Saving Up For This Conversation" with Darren Pratley on the Marketing 4 Business podcast, available on your favourite podcast streaming service, or watch it now on YouTube.
.jpg)

