AI voice agents: how to set them up to work in Europe
Latency, local speech and three layers of law at once. Notes from setting phone assistants up, and the order of work that makes them succeed in Europe.
The first thing we measure when a phone assistant goes live at a new client is not whether it answers correctly. It is how long it takes from the moment the caller stops talking until the assistant starts. Everything else in the setup is easier to fix afterwards than that number is.
It seems like a strange thing to spend a week on, right up until you have listened to enough recordings. Then it stops seeming strange.
These are notes from setting them up rather than a product description. Technical, linguistic and legal together, because a European call has to satisfy all three at once.
Latency decides everything else
There are figures from real traffic to work against now. A review of twelve production deployments published in 2026 found a median of 680 milliseconds from the caller finishing to the assistant responding, and 1,180 milliseconds for the slowest five percent of responses.
The thresholds in that material match what we hear in the recordings. Below 600 milliseconds people cannot tell a system from a person. Between 600 and 800 it sounds natural. Above 1,000 callers start talking over the assistant, and above 1,200 you get the sentence nobody wants on the recording: "hello, are you there?"
There is a price for getting down there. An architecture where speech goes straight to speech lands around 560 milliseconds and costs meaningfully more per minute than the arrangement where speech becomes text, text becomes an answer, and the answer becomes speech again. The second one lands around 700 milliseconds and is cheaper to run at volume. Both are usable. Choosing between them is a decision about what a call is worth, not a technical detail.
A call that goes silent for two seconds is lost. The customer does not think the system is thinking. The customer thinks the line dropped.
Which is why the first principle in our setups is that the system is never silent. A lookup in the booking system that takes three seconds has to be covered by a sentence saying somebody is checking. Without that sentence people hang up while the answer is on its way.
Danish is not English with an accent
Language is where European deployments get underestimated most often.
The large providers write their language lists with English, Spanish and German at the centre. Danish is rarely on the list, and when it is, that is not the same as it having been tried on a Jutland accent, on an address with a street name and a house number, or on a time said as "half past three" the Danish way rather than as 15.30. It has to be measured, not assumed. We test with background noise, with numbers and with names, because that is where it breaks.
Language is not only a quality question in Europe either. In several countries it is a requirement. France demands French in consumer facing systems, Belgium demands the language of the region, Poland demands Polish in business dealings, Spain demands the co official languages in the relevant territories, and Lithuania has required customer service in Lithuanian since 1 January 2026. If the assistant will take calls in more than one country, the language decision belongs in the setup from the start.
The three rulebooks a European call satisfies at once
A single call touches three layers of law, and they are not stacked neatly on top of each other.
Disclosure comes first. Article 50 of the AI Act applies from 2 August 2026 and means the call has to open by saying the person is speaking to a system. Not in small print somewhere. Audibly. The ceiling is €15 million or 3 percent of turnover.
Then there are the recordings, and this is where it turns national. Denmark is among the countries where it is enough that one party to the conversation knows. Germany is not: recording without the consent of every party is a criminal offence under §201 of the German criminal code. Austria, Portugal, Belgium and a number of others sit in the same group. The practical rule is to build for the strictest country you call into, and put an audible consent gate at the start of the call.
Outbound calling carries its own layer on top. Article 13 of the ePrivacy directive requires prior consent for automated calling machines, and an assistant is a calling machine in that sense. France tightened further from 11 August 2026 with a strict opt in requirement for consumer contact.
One last item we get asked about surprisingly often: no, the system may not assess employees' emotions from their voice. That has been prohibited under Article 5 since February 2025.
What a minute costs
Pricing usually arrives as a single number from a supplier, and that number rarely covers the whole call. The same material from twelve production deployments puts the all in cost between 0.07 and 0.21 dollars per connected minute.
The spread comes from two things. One is architecture, as above. The other is volume: at around a thousand minutes a month you sit at the expensive end, at ten thousand it drops, and at a hundred thousand you are at the bottom of the range. A typical minute divides across telephony, speech to text, the language model itself and speech back out, and the language model is the part that swings most.
The calculation that matters for a smaller business is rarely the price per minute. It is what an unanswered call costs. A hairdresser, a dentist or a plumbing firm that misses the phone at the busy hour loses a booking, not a minute. Put those two numbers next to each other before negotiating over fractions of a cent.
What actually makes them fail
The deployments we have seen fail rarely fail on the model. They fail at the edges.
The most common one is that nobody decided when the assistant should give up. In the published production material a median deployment handles 78 percent of calls on its own, and the best ones 84 to 88 percent. The best ones are also the ones with the narrowest job: a slot in the calendar, an order status, a confirmation. None of them try to do everything.
The second most common is names and addresses that get guessed instead of confirmed. A system that says "I have you as Kirsten Andersen, is that right?" is slower and far better than one that simply carries on. Read address changes back. Every time.
The third is that nobody wrote down what happens when it goes wrong. A number to transfer to, a message that gets created, an email that gets sent. There has to be an exit, and it has to be chosen in advance.
How to set one up so it works
This is the order we use.
- Pick one job first. Book an appointment, give an order status, take an order. Not three at once.
- Write the opening sentence word for word, including the disclosure that this is a system, and do not bury it in politeness.
- Decide what happens on uncertainty: which number it transfers to, who gets the message, how long the system may try on its own.
- Set a retention limit for recordings and transcripts, and hold to it. Shorter is easier to defend than longer.
- Make sure customer data is stored in the EU, and describe it in your privacy policy with words that actually match your setup.
- Measure every call: latency, interruptions, how many were handled without a person. Without numbers a phone assistant is a feeling.
- Test with real voices in real conditions before it takes a call from a customer. Dialect, noise, numbers, names.
Legal and technical sit closer together here than almost anywhere else. We came at the same tension in the piece on GDPR safe AI and where it stops, and a phone assistant is the place where that theory becomes audible to your customer.
If you want to hear what it sounds like once it is set up, there is a recorded conversation on the phone assistant page. It is faster to listen to than to read about.
One last note from the recordings. The calls that go well are rarely the ones where the system is impressive. They are the ones where the customer got an appointment, never thought about who picked up, and hung up after ninety seconds.
What’s your brand’s score?
Find out where your brand stands. Book a discovery call and we’ll run Signal on your brand together.
ChatGPT ads arrive in the EU: what the launch means
Ads inside ChatGPT reached 31 European countries on 24 August. Here is what can be bought, what cannot, and why that difference decides your budget.
Agent to agent: what the A2A protocol means for SMEs
A2A lets one company's AI agent deal directly with another's. Here is what that changes for small and medium businesses in Europe, and what to do this year.
Claude Epic: what to expect from the next model series
Anthropic has announced no series called Epic. We use the name as a placeholder and ask what the pattern of 2026 says about the generation that follows.