Debt collection · 2025
Moving a 1,200-agent floor off self-hosted VICIdial without dropping a call
Cutover of the largest single floor I have run — 1,200 concurrent agents — from ageing on-premise hardware to a managed cluster, executed live between shifts.
- Downtime during cutover
- 0 min
- No shift lost dialing time
- Concurrent agents sustained
- 1,200
- Peak load, post-migration
- Answer-to-agent latency
- −38%
- Against the previous platform
- Unplanned outages since
- 0
- Across the following 9 months
The situation
Twelve hundred agents across three shifts on hardware that was seven years old and out of warranty. Two disk failures in the preceding quarter, each one costing a full day of dialing.
The floor could not go dark. Collections targets were contractual, and every hour offline was measurable revenue.
The incumbent build had drifted years away from stock VICIdial through undocumented edits, so a straight reinstall would have lost behaviour nobody could fully describe.
What I did
- 01
Audited the running system and diffed it against stock VICIdial to recover every local modification as a documented patch set.
- 02
Built a parallel cluster — separate dialers, database and web tier — and replayed production call load against it for a fortnight before touching anything live.
- 03
Migrated carrier trunks one at a time so each route could be verified under real traffic and rolled back independently.
- 04
Moved agents campaign by campaign between shift changes, keeping both platforms live and reconcilable throughout.
- 05
Handed over runbooks and trained the client’s own operations team on the cluster before stepping back.