Mihu AI On-Premise

Your agents. Your models. Your voices. Your infrastructure. Your data.

Mihu runs the complete AI agent stack — from speech and reasoning to telephony and actions — inside your infrastructure. Customer data does not have to leave your environment.

7
On-premise speech languages
CPU
Mihu STT & TTS inference
99.9%
Availability target
Deployment boundary
Inside your network
Telephony
SIP infrastructure, broadcast services
Local
Speech
Mihu STT, Mihu TTS and voice cloning on CPU
Local
Reasoning
Open-weight LLM, self-hosted
Local
Actions
Action Server, integrations, workflows
Local
Data
Database, Redis, recordings, transcripts
Local
Overview

Private AI voice infrastructure, deployed entirely inside your environment.

Mihu AI On-Premise is an enterprise offering that allows organizations to run the complete Mihu AI agent stack within their own infrastructure.

Speech processing, AI inference, orchestration, telephony, integrations, application services and data storage can operate without customer data leaving the organization's environment.

The entire deployment is delivered as containerized Docker services and runs on customer-managed infrastructure, deployed on Kubernetes with Rancher or Helm.

Speech
On-premise, CPU inference
LLM
Open weights, self-hosted
Telephony
SIP inside your network
Data
Customer-controlled storage
Speech Infrastructure

Mihu STT & Mihu TTS, completely on-premise.

Mihu provides its own speech models — Mihu STT (Speech-to-Text) and Mihu TTS (Text-to-Speech) — completely on-premise. For supported deployments:

  • Mihu STT (Speech-to-Text) runs locally.
  • Mihu TTS (Text-to-Speech) runs locally.
  • CPU-only inference is supported for the speech layer.
  • No external speech API is required.
  • Voice cloning can be deployed inside the customer's infrastructure.
  • Customer audio and training data can remain entirely inside the customer's environment.
Current on-premise speech stack · 7 languages
English German French Spanish Italian Russian Turkish

These 7 languages run entirely on Mihu's own STT and TTS models, inside your infrastructure. The Mihu cloud platform supports 40+ languages through additional speech models.

5–9%
Typical Word Error Rate (WER) of current Mihu STT (Speech-to-Text) deployments, depending on the language and evaluation dataset.

Lightweight CPU speech models

Mihu's speech infrastructure is designed to keep inference requirements small. Depending on language and model configuration, deployed speech models are approximately 66M – 143M parameters.

This makes it possible to operate STT and TTS workloads on CPU infrastructure without requiring dedicated GPUs for every voice workload.

Actual capacity depends on
  • concurrent conversations
  • selected language
  • selected STT/TTS model
  • audio configuration
  • latency target
  • redundancy requirements

Concurrency figures are established per deployment through benchmarking on the target hardware, not quoted as a fixed CPU-to-call ratio.

Customization

Private fine-tuning and voice cloning

Proprietary speech data and brand voices stay on infrastructure you control.

Private Fine-Tuning

Organizations with high-quality proprietary speech data can further customize the speech layer. Mihu can fine-tune supported STT/TTS models directly using infrastructure controlled by the customer, so proprietary datasets do not need to be transferred to a third-party inference provider.

Custom models can be
  • fine-tuned using customer-provided datasets
  • optimized for industry-specific terminology
  • adapted for accents and domain-specific speech
  • deployed exclusively inside the customer's infrastructure
  • kept private, entirely within your own infrastructure

Particularly useful for organizations with specialized vocabulary, regional accents or strict data-residency requirements.

Voice Cloning

Mihu On-Premise also supports private voice cloning. Provided voice samples can be used to create organization-specific voices, which are then hosted inside the customer's own infrastructure.

Both the source recordings and the resulting voice model can remain private.

This allows enterprises to create consistent branded voices without relying on an external TTS provider during production inference.

Voice sample requirement: approximately 30 seconds of clean speech, recorded without long pauses or gaps between sentences.

Multi-Language Architecture

Language-specific on CPU, or multilingual on GPU

The lightweight CPU speech models are primarily language-specific rather than a single multilingual model. For AI agents that need to switch languages dynamically during the same interaction, Mihu supports multiple deployment strategies.

CPU efficiency

Model Orchestration

Multiple language-specific speech models run simultaneously. Mihu's orchestration layer detects or receives the active language and dynamically routes speech processing to the appropriate model.

System-level overrides can also control model selection during an interaction.

Maximum multilingual flexibility

GPU Multilingual Models

For scenarios requiring highly dynamic multilingual conversations, larger multilingual speech models can be deployed on GPU infrastructure instead.

The architecture is optimized for either CPU efficiency with language-specific models, or maximum multilingual flexibility with GPU-based models, depending on the workload.

Note: multilingual STT and multilingual TTS both require GPU infrastructure. The CPU-only speech path applies to the language-specific Mihu STT and Mihu TTS models.

Local LLM Infrastructure

The reasoning layer runs on open-weight models inside your environment.

The reasoning layer runs entirely within your environment using open-weight LLMs: models whose weights are publicly released and can run fully offline on your own hardware, with no connection to the model vendor. Mihu continuously evaluates new open-weight models instead of permanently tying the on-premise stack to a single model family.

High-performance LLM deployments generally require GPU infrastructure. The recommended model may change as new models become available, and Mihu can update or replace the reasoning model independently from the rest of the agent infrastructure.

The infrastructure is model-agnostic. The best available model can change; your architecture does not have to.

Recommended configuration as of deployment date

The reasoning model is selected at deployment time from the current evaluation set. As of September 2026, one of our preferred high-performance configurations is Kimi K2.6 Instant. It runs on dedicated GPU infrastructure provisioned in addition to the 128 vCPU baseline. Open-weight alternatives from other vendors and regions, such as Llama or Mistral, can be selected where procurement or model-origin requirements apply. The selection is confirmed in each deployment proposal.

Mihu Core Infrastructure

What a full deployment consists of

A full Mihu On-Premise deployment consists of several independent services. All core components are delivered as containerized Docker services and deployed on Kubernetes with Rancher or Helm.

01

Mihu Core

Core AI agent runtime and business logic.

02

Mihu Orchestration

Coordinates speech models, LLMs, agents and runtime services.

03

Mihu Relay

Real-time communication layer between services and agent sessions.

04

SIP Infrastructure

Handles enterprise telephony and SIP connectivity.

05

Broadcast Services

Supports outbound and high-volume communication workloads.

06

Action Server

Executes tools, integrations and agent actions.

07

Management / Web Application

Administration, configuration and operational management interface.

08

Database Infrastructure

Persistent application and operational data storage.

09

Redis Infrastructure

Real-time state, caching and distributed runtime coordination.

Optional infrastructure

Additional Mihu capabilities can also be deployed on-premise. These components require additional server capacity depending on usage.

Mihu Builder

For generating and running custom services, code, integrations and workflows. Running Builder requires additional GPU capacity on the server.

Mihu MCP Server

For exposing the Mihu workspace and capabilities to MCP-compatible AI systems.

PBX Connector Server

Connects your existing PBX or telephony system to the Mihu SIP infrastructure for inbound and outbound calls.

SMTP Server

Required if you want to connect the email channel. Agent and notification emails are sent through your own SMTP server.

Observability & Logging

Dedicated observability for production deployments

Production deployments should include dedicated observability infrastructure. Mihu recommends separate infrastructure for centralized logging and monitoring.

Centralized logging

Logs from Mihu Core, orchestration, SIP, relay, model servers and application services are collected centrally.

Monitoring

Infrastructure and application metrics are monitored continuously. Alerts can be generated when predefined operational thresholds are exceeded.

Monitored metrics
  • CPU utilization
  • Memory utilization
  • GPU utilization
  • Disk capacity
  • Network health
  • Service health
  • Model inference latency
  • Queue depth
  • Call infrastructure
  • Application errors
Network Architecture

No external connectivity required

All Mihu services run on servers inside the customer's own network. No VPN, tunnel or persistent connection to Mihu is required for production operation.

Internal AI services do not need public exposure. Telephony reaches the SIP infrastructure through the customer's existing carrier or PBX connectivity, and network rules are defined entirely by the customer's infrastructure and security policies.

Customer Network
Entry pointsSIP trunk / PBXWeb appAPI
Core / Orchestration / SIP
STTLLMTTSActions
DB / Caching / Storage
OperationsLogsNotificationsMonitoring
Infrastructure Requirements

Production baseline

For a production Mihu On-Premise installation, the starting infrastructure configuration is determined according to expected concurrency and workloads.

Compute

128 vCPU
Minimum recommended · core services and CPU speech layer
+ GPU
Depending on the selected model · e.g. Kimi K2.6 Instant, Qwen, Llama or Mistral; the recommended model may change over time

The 128 vCPU baseline covers the Mihu Core services and CPU-based STT/TTS. It does not cover the reasoning layer: a high-performance self-hosted LLM such as Kimi K2.6 Instant requires dedicated GPU infrastructure on top of this baseline, sized for the selected model and expected concurrency. Selecting a smaller open-weight model reduces the GPU footprint accordingly. GPU capacity is also added when multilingual GPU speech models are required. The exact allocation depends on concurrent AI conversations, STT/TTS workload, languages deployed, LLM selection, redundancy configuration and additional services such as Builder and MCP.

Physical core versus vCPU allocation is confirmed in the deployment proposal.

Storage

calls × average duration × recording format × retention period
Recording capacity

Storage requirements primarily depend on whether call recordings and other media are retained locally. Mihu application databases require persistent storage, while recording capacity is calculated separately. Customers retaining months or years of recordings should provision a dedicated storage volume sized according to their retention policy.

Customers define their own
  • recording retention
  • transcript retention
  • backup policy
  • archival policy
  • deletion policy

High Availability

99.9%
Operational availability target

Production architecture can be configured with redundancy across critical Mihu services. Achieving this target requires the agreed production architecture, redundancy and infrastructure requirements to be maintained.

Dedicated operational ownership is assigned for enterprise on-premise installations. In the absence of underlying infrastructure or hardware failures outside Mihu's operational responsibility, the deployment is designed toward the agreed 99.9% service availability target. Layer-by-layer responsibilities are defined in the SLA document.

Deployment

Pre-configured, Dockerized, delivered together with your infrastructure team

The complete Mihu stack is provided as pre-configured Dockerized services. Mihu's engineering team handles deployment configuration and production readiness together with the customer's infrastructure team. A typical on-premise deployment is delivered in 3–6 months, from infrastructure preparation to production, and is available to enterprise customers.

  1. Infrastructure preparation
  2. Network configuration
  3. Mihu services
  4. Speech models
  5. LLM
  6. Telephony
  7. Storage
  8. Observability
  9. Validation
  10. Production

Kubernetes: the stack can be deployed with Helm charts. Mihu's standard and fastest path is a Rancher-managed Kubernetes deployment.

Versioning & releases

Mihu On-Premise uses branch-based versioning: each customer deployment runs on its own release branch, and updates are delivered as monthly releases. Every release is validated before rollout and applied in coordination with the customer's infrastructure team, so upgrades are scheduled rather than pushed. Release validation covers the standard test suite; some unit tests may differ depending on customer-specific customizations.

Dedicated support team

For each on-premise environment, Mihu assigns a customer success team and a DevOps team to work alongside your team. Support runs through a shared Slack group, so operational questions, release coordination and incidents are handled directly with the people who know your deployment.

Everything stays inside.

A one-page view of what runs where in a Mihu On-Premise deployment.

SpeechOn-premise
Mihu STTCPU
Mihu TTSCPU
Voice cloningPrivate
LLMOpen weights / self-hosted
Fine-tuningOn customer infrastructure
TelephonySIP
DataCustomer-controlled
RecordingsCustomer-controlled
DatabaseOn-premise
IntegrationsOn-premise Action Server
DeploymentDockerized · Kubernetes (Rancher / Helm)
ConnectivityInside the customer network
ObservabilityDedicated
SLA target99.9%
Delivery3–6 months to production
AvailabilityEnterprise customers
SupportCustomer success + DevOps team, shared Slack group
ReleasesMonthly, branch-based versioning
FAQ

Mihu AI On-Premise — Frequently Asked Questions

What actually runs inside our infrastructure?

The complete stack: Speech-to-Text, Text-to-Speech, voice cloning, the LLM reasoning layer, orchestration, SIP telephony, the Action Server for integrations, the management application, and the database and Redis layers. Customer data does not have to leave your environment.

Do we need GPUs?

Not for the speech layer. STT and TTS run on CPU with models of roughly 66M to 143M parameters. High-performance open-weight LLMs generally require GPU infrastructure, and GPU capacity is also added if you choose larger multilingual speech models.

Which languages does the on-premise speech stack support?

The current on-premise CPU speech stack supports seven languages: English, German, French, Spanish, Italian, Russian and Turkish. These language-specific Mihu STT and Mihu TTS models run on CPU, without GPUs. Mihu STT typically achieves around 5–9% Word Error Rate (WER); the exact figure is based on the language and the evaluation dataset.

Which LLM is used, and can it be changed later?

The reasoning layer is model-agnostic and runs open-weight models. Mihu recommends a model as of the deployment date, and alternatives from different vendors and regions, for example Llama or Mistral alongside Kimi K2.6 Instant, can be selected to meet procurement or model-origin requirements. The model can be updated or replaced independently from the rest of the agent infrastructure as better open-weight models become available.

Can we fine-tune the speech models on our own data?

Yes. Supported STT and TTS models can be fine-tuned on customer-controlled infrastructure using your own datasets, so proprietary audio never goes to a third-party provider. The resulting models stay private to your organization.

What infrastructure do we need to start?

A typical production baseline starts at 128 vCPU for the core Mihu services and the CPU-based speech layer. The self-hosted LLM is sized separately and requires dedicated GPU infrastructure on top of that baseline; GPU capacity is also added for multilingual GPU speech models. Storage is sized from your recording and transcript retention policy. Exact sizing is confirmed in the deployment proposal.

How is the deployment delivered and connected?

Everything ships as pre-configured Docker services and runs on the customer's own servers. No VPN or connection back to Mihu is required for production operation, and no internal AI service needs public exposure. Mihu's engineering team handles deployment configuration together with your infrastructure team.

Can you train STT and TTS for a language outside the current seven?

Yes. New languages can be added on request. Training effort differs by language group, and STT and TTS are trained separately. Mihu can supply or purchase the training dataset for you, or work with data you provide. Training runs on machines you supply or on infrastructure Mihu provides, and the resulting Mihu STT and Mihu TTS models are then deployed inside your environment exactly like the standard ones.

Planning an on-premise deployment?

Share your expected concurrency, languages and data-residency requirements. Our engineering team will come back with a sizing proposal and deployment plan.