A practical guide / Fonix.AI
On-premise voice AI: architecture, capacity and acceptance testing
Workflows, exceptions and the questions to ask before deployment.
Define what must stay inside the deployment
On-premise voice AI runs the agreed conversation components on infrastructure an organisation controls. The important purchasing question is which components are included and what dependencies remain outside that boundary. Hosting a language model locally does not by itself establish where telephony, speech processing, transcripts, tool calls or monitoring run.
Fonix.AI offers cloud and on-premise deployment. AntEngage's technology overview describes its conversation architectures. This guide provides a deployment worksheet and acceptance plan so your infrastructure and operations teams can evaluate the actual proposed system.
Start with one business task and a diagram of its entire path. State the required boundary in concrete terms: model execution, audio processing, transcript storage, business systems, operator access and permitted external connections. Have the relevant internal teams agree that scope before hardware sizing or procurement.
Map the whole call path, including dependencies
Separate the customer-facing phone network from the application and model environment. A system answering ordinary telephone calls has a telephony path; describe how that path reaches the deployment instead of using “offline” as an undefined promise.
For each component, record location, data received, data retained, permitted outbound connections and its owner. Include less visible services such as logs, crash reports, licence checks, package downloads, backup storage and remote support. These can matter as much as the primary model endpoint.
| Component | Questions to answer | Acceptance evidence |
|---|---|---|
| Telephony ingress and egress | Which provider, trunk and network route carry audio? | Documented route and a tested call |
| Speech and conversation models | Which stages execute locally? Is any hosted inference used? | Runtime inventory and observed connections |
| Business tools | Which CRM, calendar or account APIs are called? | Allowed endpoints, permissions and test results |
| Record storage | Where are audio, transcripts and outcomes stored? | Storage locations, access roles and deletion process |
| Observability | Where do logs and diagnostics go? | Logging configuration and inspected outbound destinations |
| Updates and support | How do releases and support sessions enter the boundary? | Approved update and access process |
Distinguish required external business services from accidental dependencies. If a CRM is hosted elsewhere, a locally hosted voice system may still call it. The system design should make that connection explicit rather than imply that local inference moves every business record on-premise.
Choose the conversation architecture for the task
Discuss whether the deployment uses a speech-to-text, reasoning and speech-synthesis pipeline, a direct speech-to-speech path or another supported configuration. Ask how each option handles business actions, language coverage, transcript creation, interruption and recovery.
For a transactional task, the architecture needs more than natural audio. It must obtain the right system record, perform an authorised action, detect its result and preserve enough context for a handoff. Review where those decisions happen and which component owns state.
If a third-party or internally selected model is proposed, document its hosting and resource requirements separately. Do not assume that changing the reasoning model leaves language quality, latency or capacity unchanged. Run the target workflow on the exact configuration being considered.
The multilingual evaluation guide supplies a task-based test design. Use it alongside infrastructure tests so a deployment that meets a network boundary also proves that it can perform the intended conversation.
Size for concurrent conversations and peak bursts
Monthly call minutes help estimate usage, but they do not establish the required real-time capacity. Capture peak simultaneous calls, arrival bursts, typical duration, language mix and the business actions each conversation requires.
A simple illustrative load calculation is: if 120 calls arrive uniformly over ten minutes and each lasts two minutes, the average offered concurrency is about 24 calls. That is a planning example, not a hardware recommendation or a Fonix capacity claim. Real arrivals are uneven, and provisioning requires measured performance and headroom.
Ask for benchmarks on the proposed hardware, model versions and telephony configuration. Review end-to-end response timing, dropped connections, tool latency, resource use and capacity at the target peak. Record what happens when load exceeds the agreed limit: queue, busy treatment, transfer or another defined fallback.
Confirm whether capacity is shared with chat, campaigns or other workloads. A successful quiet-hour demonstration does not establish behaviour while outbound campaigns and inbound support calls run together.
Assign operating responsibilities before launch
Agree who owns application releases, model updates, operating-system patches, certificates, backups, capacity monitoring and incident response. List the vendor's responsibilities and the customer's responsibilities in terms operators can use during an outage.
Define access roles for recordings, transcripts, configuration and business-tool credentials. Decide how changes are reviewed, how access is removed and which audit records are available. Specify retention and deletion processes for each stored artefact rather than treating every log like a call recording.
For restricted networks, agree how software packages and model artefacts arrive, how they are verified, and how an unsuccessful update is rolled back. Check whether setup, licensing or normal operation needs connectivity beyond the permitted routes.
Local hosting is an architectural choice. It does not automatically establish compliance with every law, certification or sector policy. Evaluate those requirements against the proposed data flow and operating process with the appropriate internal reviewers; avoid accepting a product badge as the whole assessment.
Write a failure and recovery plan you can demonstrate
| Failure | Required decision | Evidence from the test |
|---|---|---|
| Model worker stops during a call | Continue through the supported recovery path or provide fallback | Customer experience and final record |
| Business-system lookup is unavailable | Avoid an invented answer or action result | Tool error and explicit unresolved outcome |
| Storage becomes unavailable | Follow the agreed recording and task policy | Alert, handling behaviour and recoverability |
| Telephony connection drops | Determine whether the action completed before another attempt | Call reference and source-system state |
| Demand exceeds capacity | Apply the defined admission or fallback behaviour | Queue, transfer or busy response under load |
| Release fails | Roll back through the agreed process | Restored version and repeat acceptance result |
Review the cases with operations staff. A recovery plan should state who receives the alert, what they can inspect and how they restore service. “Contact support” is incomplete if support cannot reach the environment under its access policy.
Where an action result is unknown, keep it unknown until the source system is checked. The same principle applies to local and hosted deployment: a timed-out booking write must not become a confident spoken confirmation.
Make acceptance a joint review
Use a controlled evaluation environment and the actual intended hardware and network rules. With your team's approved monitoring tools, observe the configured destinations during normal calls, failures, restarts and updates. Record the model and application versions alongside the results.
Combine the dependency inventory, peak-load test, representative language calls, tool permissions and recovery demonstrations into one acceptance record. Set thresholds and evidence requirements before the demonstration, based on the business task and operational risks.
Compare deployment costs over the same period and scope. Include hardware, redundancy, infrastructure staff, support, telephony, updates and integration operation. Keep assumptions explicit; a smaller hosted subscription and an on-premise first-year purchase are not directly comparable without a shared workload model.
Scope your boundary with the Fonix team
Bring your data-flow diagram, telephony setup, peak concurrency, required languages, existing infrastructure and recovery expectations. Ask for a concrete architecture and an acceptance demonstration rather than a generic statement about local hosting.
Discuss an on-premise Fonix deployment. The calling software guide covers workflow and integration evaluation; this page focuses on the hosting boundary, capacity and ongoing operation of that workflow.
