Skip to content

Debt collection · 2025

Moving a 1,200-agent floor off self-hosted VICIdial without dropping a call

Cutover of the largest single floor I have run — 1,200 concurrent agents — from ageing on-premise hardware to a managed cluster, executed live between shifts.

A North American collections BPO6 weeksDelivered by Callix
Outcomes
Downtime during cutover
0 min
No shift lost dialing time
Concurrent agents sustained
1,200
Peak load, post-migration
Answer-to-agent latency
−38%
Against the previous platform
Unplanned outages since
0
Across the following 9 months

The situation

Twelve hundred agents across three shifts on hardware that was seven years old and out of warranty. Two disk failures in the preceding quarter, each one costing a full day of dialing.

The floor could not go dark. Collections targets were contractual, and every hour offline was measurable revenue.

The incumbent build had drifted years away from stock VICIdial through undocumented edits, so a straight reinstall would have lost behaviour nobody could fully describe.

What I did

  1. 01

    Audited the running system and diffed it against stock VICIdial to recover every local modification as a documented patch set.

  2. 02

    Built a parallel cluster — separate dialers, database and web tier — and replayed production call load against it for a fortnight before touching anything live.

  3. 03

    Migrated carrier trunks one at a time so each route could be verified under real traffic and rolled back independently.

  4. 04

    Moved agents campaign by campaign between shift changes, keeping both platforms live and reconcilable throughout.

  5. 05

    Handed over runbooks and trained the client’s own operations team on the cluster before stepping back.