On-premise and private cloud
Custom AI agents on your own servers and GPUs. Installed, integrated and run by Agentifire.
The same voice, chat and WhatsApp agents as our cloud, deployed inside your data center on Agentifire's own speech recognition, dialogue model and voices. Nothing leaves your network.
- 100%
- of audio and data stays on your network
- 3
- proprietary models: STT, LLM, TTS
- 1–3 weeks
- install and integrate, hardware ready
- 40+
- languages, including Urdu and Arabic
How a call moves through the stack.
Audio arrives from a phone line, WhatsApp or the browser, is transcribed, understood and answered on your GPUs, and goes back to the caller in under a second. Every stage is Agentifire's own.
Agentifire STT
Streaming speech recognition built for telephony audio: noisy lines, code-switching between Urdu and English, Gulf and South Asian accents. Partial transcripts stream to the model as the caller speaks.
Agentifire LLM
A dialogue model tuned for business calls: follows the script, uses your knowledge base and CRM tools, handles interruptions and knows when to hand off. Runs quantised on a single GPU for the default tier.
Agentifire TTS
Streaming speech synthesis with 30 voices across languages and accents. First audio in a few hundred milliseconds so turns feel natural, with custom voices available for enterprise.
What Agentifire does, and what you provide.
You provide the rack. We provide everything that runs in it.
Agentifire delivers
Installation
Kubernetes or Docker deployment, model serving, GPU scheduling, storage, backups and monitoring, on your hardware or private cloud.
Integration
SIP trunks and PBX, DID numbers, WhatsApp Business API, WebRTC for the website, CRM/ERP/helpdesk connectors, SSO and your call center or BPO for handoff.
Deployment and tuning
Agent design, scripts, knowledge base ingestion, voice selection, pilot with real calls, then go-live with your team trained on the inbox and SIP client.
Operations under SLA
Model and security updates, capacity reviews, 24/7 monitoring and a named engineer. Zero-downtime updates on highly available deployments.
You provide
- GPU servers per the calculator below, or a private-cloud GPU tenancy
- Network access to your PBX / SIP trunk and, for WhatsApp, an outbound path to Meta
- Access to the systems the agent should read and write: CRM, ticketing, calendars
- A WhatsApp Business account (we handle Meta verification with you)
- A technical contact for firewall, DNS and certificates
- Sign-off on scripts, voices and handoff rules before go-live
Reference hardware
A single 2U server with one NVIDIA L40S, 32 vCPU, 128 GB RAM and 2 TB NVMe runs about 20 concurrent voice agents plus chat. Add a second server for high availability.
How much hardware do you need?
Pick how many agents run at the same time. The calculator sizes GPUs, servers, memory and storage for the full Agentifire stack — STT, LLM and TTS — in your own rack.
Phone, WhatsApp and web voice calls happening at once.
Website and WhatsApp text conversations (no speech models needed).
Recommended for 10 voice + 10 chat agents on the 8B dialogue model
2× NVIDIA L40S
PCIe, 350 W — best value for 10–40 agents
2
servers (incl. spare)
48 vCPU
192 GB RAM
1.7 TB
NVMe, 6-month recordings
| GPU option | GPUs | Servers | RAM | Storage |
|---|---|---|---|---|
| NVIDIA L424 GB · 6 voice agents per GPU | 6 | 2 | 448 GB | 1.7 TB |
| NVIDIA A1024 GB · 8 voice agents per GPU | 4 | 2 | 320 GB | 1.7 TB |
| NVIDIA L40S48 GB · 20 voice agents per GPU | 2 | 2 | 192 GB | 1.7 TB |
| NVIDIA A100 80GB80 GB · 35 voice agents per GPU | 2 | 2 | 192 GB | 1.7 TB |
| NVIDIA H100 80GB80 GB · 50 voice agents per GPU | 2 | 2 | 192 GB | 1.7 TB |
Planning estimates for the full pipeline at low latency, with headroom for peaks. Agentifire engineers confirm the final bill of materials after reviewing your call volumes, languages and integrations.
Get a sizing review for 10 agentsOn-premise questions
- What does an on-premise Agentifire deployment include?
- The complete platform: Agentifire speech recognition (STT), the dialogue language model (LLM), text to speech (TTS), the media gateway for SIP, WebRTC and WhatsApp, the orchestrator, the team inbox and the analytics stack. Agentifire installs it on your servers, integrates it with your PBX, WhatsApp Business account and CRM, and operates it under an SLA.
- Which servers and GPUs do I need?
- It depends on how many conversations run at the same time and which dialogue model you choose. As a planning figure, one NVIDIA L40S handles roughly 20 concurrent voice agents on the default 8B model; 5 agents fit on a single NVIDIA L4. Use the calculator on this page for your numbers; Agentifire confirms the final bill of materials after a workload review.
- Does any data leave my environment?
- No. Audio, transcripts, recordings and customer records stay on your network. Agentifire STT, LLM and TTS run locally on your GPUs; no third-party AI API is called. Remote support access, if you allow it, is through a channel you control and audit.
- Can I run it in a private cloud instead of my own hardware?
- Yes. The same stack deploys to a private cloud tenancy or a dedicated GPU host in your chosen region, including data-residency-restricted regions.
- How long does an on-premise deployment take?
- With hardware in place, installation and integration typically take one to three weeks depending on the number of channels and systems to connect, followed by a short pilot before going live.
- Who maintains the models and updates?
- Agentifire does, under the support agreement: model updates, security patches, monitoring and capacity reviews, scheduled with your team and applied without downtime where the deployment is highly available.
- Can I bring my own LLM?
- The platform ships with Agentifire models tuned for business calls. Enterprise deployments can add an approved open-weight model for specific tasks; talk to us about your requirements.
Bring us your call volumes. We'll bring the bill of materials.
A 45-minute review covers channels, languages, integrations and sizing.