← All missions

Proxmox works well when the surrounding design is solid: storage, quorum, networking, backups, upgrades, and failure handling.

This mission covers new clusters, migrations from vSphere or older platforms, and reviews of existing Proxmox/Ceph setups before they become production-critical.

When this is NOT a fit

  • You need a fully managed hypervisor service and do not want to operate storage, quorum, backups, and hardware lifecycle internally.
  • The available network cannot provide predictable low-latency paths for Corosync, migration traffic, and Ceph replication.
  • The hardware mix is too uneven to make capacity planning, failure domains, or performance baselines reliable.
  • There is no maintenance window or test environment for validating failover, restore, and node replacement procedures.

Common failure modes we’ve seen

  • Corosync shares congested production networks, leading to quorum instability and split-brain risk during ordinary switch or host issues.
  • Ceph is sized from raw disk capacity instead of usable capacity, recovery bandwidth, OSD memory, and failure-domain math.
  • HA is enabled before fencing, backups, and restore tests are proven, so an automated restart can make a storage or network incident worse.
  • Cluster upgrades are treated like package updates instead of coordinated platform changes with migration, rollback, and workload placement rules.

What's included

  • Architecture design: cluster sizing, Ceph topology (replication/erasure coding), network plane separation, LACP/EVPN/VPC
  • Proxmox cluster setup: Corosync configuration, quorum tuning, live migration, HA groups and fencing
  • Storage backends: ZFS, NFS, SAN, NVMe over Fabric, Ceph (managed by Proxmox or cephadm)
  • PXE automation: bare-metal provisioning pipeline for rapid node deployment and replacement
  • Backup strategy: Proxmox Backup Server integration, snapshot scheduling, off-site replication
  • Migration execution: planned VM migration from vSphere with minimal downtime, or Proxmox version upgrades

Deliverables

  • Production Proxmox HA cluster with verified failover testing
  • PXE boot infrastructure for zero-touch node provisioning
  • Documented network architecture with VLAN map and switch configuration
  • Backup and recovery runbook with tested restore procedures
  • Storage performance baseline report per backend (IOPS, throughput, latency)
  • Operational runbook covering common maintenance and failure scenarios

Tech stack

ProxmoxCephPXEZFSNVMe-oFHAProxyCorosync

Want to scope this?

Think this fits your needs? Let's scope it together.