VMware to Proxmox Migration: 300 VMs in 72 Hours War Story
Migrating a virtual infrastructure is a complex operation demanding meticulous planning and steady nerves. When you’re moving 300 Virtual Machines (VMs) from a VMware vSphere environment to a Proxmox VE cluster, the complexity escalates exponentially. This is the account of a large-scale migration that severely tested the team but ultimately succeeded, providing invaluable lessons in managing production crises.
In the initial stages of a migration of this magnitude, optimism often runs high. We analyze requirements, define procedures, and set timelines. However, the true critical issues only emerge in the heat of battle. This narrative is not merely a chronicle of events; it’s an in-depth analysis of the mistakes made and the strategies adopted to correct course. My goal is to offer practical insights to anyone facing similar challenges in enterprise environments with hundreds of VMs.
Prerequisites / Test Environment
The starting environment consisted of a VMware vSphere 7.x cluster with approximately 300 VMs. These ranged from critical application servers to Oracle and PostgreSQL databases, Active Directory services, and file servers. The target environment was a new Proxmox VE 8.x cluster with shared storage based on Ceph/ZFS, designed to offer greater flexibility and reduce licensing costs. The migration was planned for a long weekend, aiming for completion within 72 hours, with no perceived downtime for end-users.
The Context: 300 Production VMs, Immovable Deadline
The organization had decided to consolidate its virtual infrastructure by moving from VMware to Proxmox VE to optimize costs and increase operational agility. The deadline was tight: 72 hours over an Easter weekend to minimize business impact. The 300 VMs supported critical services for over 2,000 users, and any prolonged interruption would have significant repercussions. The plan involved a cold migration for most VMs, with exceptions for the most sensitive services, which were to be live-migrated where possible.
Hours 0-8: Everything Seemed Fine (But It Wasn’t)
The migration began as planned. The first, less critical VMs were shut down on VMware, exported, and imported into the Proxmox cluster. Initial performance metrics were encouraging. The team worked at full pace, following a predefined checklist. However, an alarm signal, underestimated at the time, was the slow data transfer rate between the two environments, initially attributed to suboptimal network load. What we hadn’t considered was that the target storage, while passing preliminary tests, had never been subjected to a prolonged and intensive stress test comparable to the simultaneous migration of tens of terabytes of data.
Pre-migration tests focused primarily on functionality and basic configuration, neglecting an in-depth analysis of I/O performance under extreme load. This is a common, yet fatal, error in this context. According to a 2025 Enterprise Strategy Group report, 68% of IT migration failures are attributable to insufficient planning and testing, especially regarding storage performance.
The Critical Moment: Storage Latency at 4,800ms During Business Hours
On Friday afternoon, at the peak of migration operations, the situation deteriorated rapidly. Applications began to slow down, and users reported service disruptions. The Proxmox monitoring dashboards showed an alarming figure: storage latency had spiked to 4,800ms (4.8 seconds) for the virtual disks of active VMs on the new cluster. Diagnostic commands left no doubt:
iostat -xz 1 10
Executing this command on one of the Proxmox nodes revealed extremely high peaks in await (average I/O wait time) and disk utilization approaching 100%. The problem was clear: the Ceph/ZFS storage could not handle the combined load of already migrated VMs and ongoing import/export operations. The technical team was under immense pressure, with communications intensifying and the CEO demanding updates every twenty minutes. Production was at real risk.
Checking VMs in a problematic state became a priority:
qm list | grep -v running
This allowed us to quickly identify VMs that had not restarted correctly or were stuck in an anomalous boot state due to storage latency. VMware logs on the source cluster also showed errors related to VM exports, though less frequent:
tail -f /var/log/vmware/hostd.log | grep -i error
The Impossible Decision: Rollback or Push Forward
With storage latency through the roof and production crippled, we reached a critical crossroads. A full rollback would mean undoing all migrations, restoring services on VMware, and starting from scratch, delaying the overall operation by at least 24-36 hours. This would result in an unacceptable business impact. On the other hand, continuing to push the migration meant risking total failure and data corruption. The choice was between a certain, lesser evil and a potential, but more severe, one.
The decision was to push forward, but with a radically modified strategy. The objective became not just completing the migration but stabilizing the environment in real-time. This involved immediately suspending all ongoing cold migrations and focusing on stabilizing the Proxmox cluster.
How We Recovered Without Losing a Minute of Production
The first step was to identify and isolate the most I/O-demanding VMs, moving them to secondary storage or slowing down their operations. Simultaneously, we initiated an in-depth analysis of the Ceph/ZFS configuration to optimize performance parameters. We also leveraged Proxmox’s flexibility to distribute the load across less saturated nodes, even if it meant temporarily deviating from the initial plan.
Communication with the business was handled with maximum transparency, informing them of the critical issues but guaranteeing a constant commitment to restoration. Fortunately, the organization had an internal IT team that understood the complexity of the operation. This bought us valuable time. Despite the difficulties, we managed to stabilize storage latency within an hour, bringing it down to acceptable values (below 50ms). Migrations resumed at a more controlled pace, prioritizing critical VMs and constantly monitoring network bandwidth and I/O latencies. To monitor network bandwidth during the migration, a tool like iperf3 proved essential:
iperf3 -c <host> -t 30 -P 4
This allowed us to confirm that the network was not the primary bottleneck, reinforcing that the problem was rooted in storage. By the end of the 72 hours, all 300 VMs were operational on Proxmox, with no prolonged downtime perceived by end-users, only brief slowdowns resolved in real-time.
What We Would Have Done Differently (Post-Mortem Analysis)
The experience provided invaluable lessons. The main one is the importance of stress testing storage in a pre-production environment that replicates the real load of a massive migration, not just daily operational load. We should have simulated the simultaneous transfer of hundreds of virtual disks, not just the normal operation of VMs. Furthermore, a more detailed and tested rollback plan, with clear procedures for each failure scenario, would have reduced stress and accelerated decision-making. A 2024 SANS Institute report indicates that 40% of organizations do not adequately test their disaster recovery plans, a statistic that extends to complex migrations.
3 Operational Takeaways for Managing Similar Migrations
- Realistic Storage Stress Tests: Don’t limit yourself to testing storage performance with standard workloads. Simulate the I/O peak that a massive migration can generate. This includes simultaneous transfers of large data volumes and mixed workloads (intensive read/write). Use tools like
fioorddto generate significant I/O loads. - Detailed Contingency and Rollback Plans: Every migration must have a clear and tested rollback plan for each phase. What happens if storage fails? What if the network saturates? Having predefined procedures for these eventualities can save the entire operation. The time spent planning multiple failure scenarios is an investment, not a cost.
- Proactive Monitoring and Constant Communication: Implement granular monitoring on all critical components (CPU, RAM, disk I/O, network latency) and establish realistic alert thresholds. Transparent and constant communication with the business and the team is fundamental for managing expectations and coordinating actions in an emergency. An informed team and an understanding business can make the difference between success and failure.
Common Errors and Troubleshooting
Common errors in migrations of this scale include underestimating available network bandwidth, failing to optimize target storage parameters, and insufficient attention to file system permission details. In cases of high latency, in addition to iostat, it’s useful to check storage logs (journalctl -u ceph-osd@*.service for Ceph) and kernel logs (dmesg). Also verify network bonding configuration and MTU to avoid fragmentation that slows transfers.
Conclusions with Operational Takeaways
This VMware to Proxmox migration war story demonstrates that even the most accurate planning can encounter unforeseen issues in production. The key to success lies in the ability to react, the team’s experience, and the willingness to adapt strategy in real-time. The lessons learned – the importance of realistic storage stress tests, detailed rollback plans, and impeccable communication – are operational takeaways every IT professional should consider for their future migrations. Infrastructure resilience depends not only on technology but on the preparation and responsiveness of the people who manage it. For more insights into planning your infrastructure, you can consult the official Proxmox VE documentation here.
Updated: May 2026
Read also: Proxmox VE: Complete Installation & Configuration Guide 2026
Read also: Proxmox VE: Guía Completa de Instalación y Configuración (2026)
Read also: PostgreSQL Streaming Replication: High Availability Guide (2026)