Case Study: High-Availability SIP Infrastructure for a Live Voice Service
Automatic failover for a multi-component SIP service stack
The client needed standard SIP desk phones and softphones to work with an established mobile and fixed-line voice platform. We engineered a paired SIP service stack that combined routing, call handling, media anchoring, and real-time event translation, with automatic failover and provisioning tied to a stable virtual endpoint.
The requirement went beyond making a SIP proxy available
The client provides business communications across mobile and fixed-line services. A user could handle calls on a mobile device or a SIP endpoint, while call activity and controls remained visible through the existing backend PBX. Standard SIP endpoints therefore had to work within call and session behaviour that had not been designed around them.
The service also needed to support more than one endpoint for the same user. A desk phone and softphone had to register with the same credentials and ring concurrently, while transfers and three-way calls still had to respect a backend that permitted only one active channel for each user.
Availability depended on more than SIP routing
- Each customer and its endpoints had to be assigned to the correct SIP service instance through one stable virtual address.
- Kamailio had to register multiple endpoints against one user and offer incoming calls to each active device.
- Transfer handling had to keep the backend call leg active while the endpoint-facing side changed channels.
- Incoming-call events, call-control commands, presence, and Busy Lamp Field state had to move between the backend PBX and the SIP service.
- Media anchoring and SDP rewriting had to remain separate from SIP registration and routing.
- Failure detection had to cover the complete service stack, not only whether an individual process was running.
An active/standby stack with automatic service recovery
We separated registration and SIP routing, call handling, media anchoring and real-time event translation within each SIP service instance. Kamailio handled registration and SIP routing. FreeSWITCH acted as a back-to-back user agent and anchored the call legs needed for transfer handling. rtpengine provided media anchoring and SDP rewriting, while a separate real-time translation component interpreted call and presence events from the backend PBX.
Each instance ran as an active/standby pair. keepalived managed separate virtual endpoints for Kamailio and rtpengine, and bidirectionally replicated MariaDB held Kamailio state across the pair. The services ran in containers with explicit startup ordering, bringing the database up before the SIP service and enabling the virtual endpoints afterward.
Custom checks monitored process health, SIP traffic processing, active SIP requests, and sustained processing errors. Threshold and debounce logic prevented an isolated error from causing an unnecessary handover. A confirmed fault automatically promoted the standby node to the active role; once repaired, the preferred node could resume that role. Endpoints then re-registered against the same virtual address according to their normal registration timers.
The same checks covered the real-time translation component. It maintained long-running sessions for registered users, retained the presence state received from both sides, and resumed its sessions after a failover. This allowed the SIP service to recover as a coordinated stack rather than as a collection of unrelated containers.
Call handling worked around the backend constraint
Calls followed a deliberately layered path through Kamailio and FreeSWITCH. The endpoint-facing side could maintain the additional channel needed during a transfer, while FreeSWITCH kept the single backend PBX leg active. The real-time translation component then applied the corresponding transfer or three-way-call command through the backend control path and changed the active channel when required.
This arrangement addressed interoperability that had previously limited standard SIP devices. Kamailio registered a desk phone and softphone against the same user and offered incoming calls to both. Users retained the mobile and backend call-control behaviour already in place rather than moving to a separate voice platform.
A separately developed real-time event gateway relayed CometD messages between the backend PBX, desktop applications, and the SIP service over MQTT. Message interpretation and translation into SIP call-control and presence behaviour remained within the SIP service stack.
Provisioning tied each endpoint to the right service instance
The client's device-provisioning platform assigned each customer to a SIP service instance and rendered the corresponding desk-phone and softphone configuration. Every endpoint for that customer used the instance's virtual address. Kamailio passed registrations through to the existing PBX, while a provisioning-platform hook used registration activity to initialise the corresponding real-time session. The FreeSWITCH-based translation component then retrieved the account details it needed from the provisioning API to establish and maintain that backend session.
Keeping assignment and configuration in the provisioning platform made the distributed deployment manageable without moving operational responsibility into the SIP routing layer. Registration could continue during a temporary provisioning-platform outage because Kamailio still passed registrations to the PBX. Provisioning changes and the initialisation of new real-time sessions still required the corresponding integrations.
The service evolved without replacing the backend PBX
The SIP service was deployed across paired infrastructure in two data centres, with independently assigned instances serving defined groups of customer endpoints. The same active/standby pattern allowed each instance to fail and recover independently.
Over several years, the platform was extended to support multiple SIP endpoints per user, simultaneous ringing, transfer handling, presence and BLF translation, and changes in the surrounding network. Later introduction of session border controllers moved external NAT and security responsibilities to the network edge while rtpengine remained responsible for media anchoring and SDP rewriting.
Related capability:
- Voice & Protocol Engineering covers SIP, RTP, call-control, media, and integration requirements around established voice platforms.
Related case study:
- Unified Access and Real-Time Integration Across Telecoms Platforms explains how the real-time event gateway handled desktop sessions and relayed events to this SIP service.
Client names and identifying details have been withheld to protect confidentiality.
Discuss your SIP infrastructure
If you need an HA design around an existing voice platform, talk to us about the state, routing, media, and operational constraints.