Once a month our product team shows the rest of the company what actually shipped. The past month had two items worth writing up. The on-premise version of our voice platform is finished and moving into client deployments. And we have added a German speech synthesis engine trained and hosted inside the European Union.
Why clients want voice AI inside their own perimeter
Cloud is the default for a reason, and most of our traffic runs there. But roughly three situations keep coming back in conversations with large organisations, and they rarely arrive one at a time.
Internal security policy. Banks, large healthcare providers and government organisations often have rules that forbid running third-party software outside their own network. There is no feature that wins this argument. Either the software installs inside the perimeter or the project does not happen.
Data residency law. Several countries prohibit citizen data from leaving national borders. Picking a nearer cloud region does not answer the legal question.
Latency. Voice is unforgiving. When audio has to travel a long distance to a foreign server and back, the delay becomes audible. The customer hears a gap before the AI Agent answers, and the conversation stops feeling like a conversation. Moving processing close to the traffic removes the problem rather than masking it.
What on-premise means in practice now
Until this year, every on-premise project was custom work and months of engineering behind it. That part is done. The delivery package has four components.
- A prebuilt distribution of the voice platform, ready to install
- Integration documentation covering the APIs
- Two sizing models, one for cloud infrastructure and one for physical hardware, calculated from call volume and how long the client needs to retain data
- An 11-section readiness checklist with several dozen questions
The checklist is the part that changes the economics. A client completes it together with our team in one or two days. The output is a list of infrastructure gaps and a decision on who closes each one, including anything custom such as a specific monitoring stack or database. From there, a typical implementation runs three to eight weeks.
Regional installations, for clients who do not need their own server room
Full on-premise is not the only answer to a data residency question. For those who do not want a public cloud but also do not want to operate the platform themselves, we run country-level installations and give each client a separate tenant.
That already exists in Germany and Poland. Installations in Kazakhstan and Uzbekistan are planned. Each one can then be reused to connect the next client in the same region without reopening the legal or latency discussion.
One clarification that came up previously. The voice platform and the AI engines behind it (the language model, speech synthesis and speech recognition) are separate projects. A client can host the platform locally and still call external engines, or host the engines locally too. The second option is possible and has been delivered, but it carries its own infrastructure requirements and is scoped separately.
First deployments
Development finished on schedule in mid-August. We deployed it ourselves first, inside our own perimeter, before offering it to anyone else.
The first external commercial deployment is a national mobile and broadband operator in Uzbekistan, serving a market of around 37 million people. Their contact centre handles about 1.1 million inbound calls a month against 1.5 million conversation minutes, with queue waits of eight to ten minutes. Roughly 400,000 of those calls are abandoned before anyone answers.
The plan is to automate about 30% of traffic, which is close to 500,000 conversation minutes a month, by reusing a scenario we already run for another operator in the region. The point of the project is not headcount. Their 2,000 operators stay where they are. The target is the 400,000 conversations that currently never happen at all.
A German voice that sounds German
The second item is smaller but landed well. We have integrated a speech synthesis engine from a Berlin company that trains and hosts native German models inside Europe.
Three things make it interesting.
It can run inside a customer perimeter, which makes it a natural fit for the on-premise work above, where relevant. It returns the first audio in tens of milliseconds, an order of magnitude below what we are used to from the widely used commercial engines, and in voice that difference is heard rather than measured. And it is fully European-owned and compliant.
Our team members who listened to it first said the same thing independently: it sounds emphatic, like a native speaker rather than a rendering of one.
The through line across both topics is the same one we keep coming back to. Where the software runs, where the data sits and where the voice is generated are not implementation details for regulated clients. They are the decision.
Apifonica runs on one platform operated under an ISO/IEC 27001:2022-certified Information Security Management System. Data stays within EU/EEA infrastructure by default, and client data is never used for AI model training.
