Mission 03
Proxmox / Ceph HA Platform
Design or harden Proxmox/Ceph clusters for migration, HA, storage, failure handling, and day-2 operations.
Proxmox works well when the surrounding design is solid: storage, quorum, networking, backups, upgrades, and failure handling.
This mission covers new clusters, migrations from vSphere or older platforms, and reviews of existing Proxmox/Ceph setups before they become production-critical.
When this is NOT a fit
- You need a fully managed hypervisor service and do not want to operate storage, quorum, backups, and hardware lifecycle internally.
- The available network cannot provide predictable low-latency paths for Corosync, migration traffic, and Ceph replication.
- The hardware mix is too uneven to make capacity planning, failure domains, or performance baselines reliable.
- There is no maintenance window or test environment for validating failover, restore, and node replacement procedures.
Common failure modes we’ve seen
- Corosync shares congested production networks, leading to quorum instability and split-brain risk during ordinary switch or host issues.
- Ceph is sized from raw disk capacity instead of usable capacity, recovery bandwidth, OSD memory, and failure-domain math.
- HA is enabled before fencing, backups, and restore tests are proven, so an automated restart can make a storage or network incident worse.
- Cluster upgrades are treated like package updates instead of coordinated platform changes with migration, rollback, and workload placement rules.
What's included
- Architecture design: cluster sizing, Ceph topology (replication/erasure coding), network plane separation, LACP/EVPN/VPC
- Proxmox cluster setup: Corosync configuration, quorum tuning, live migration, HA groups and fencing
- Storage backends: ZFS, NFS, SAN, NVMe over Fabric, Ceph (managed by Proxmox or cephadm)
- PXE automation: bare-metal provisioning pipeline for rapid node deployment and replacement
- Backup strategy: Proxmox Backup Server integration, snapshot scheduling, off-site replication
- Migration execution: planned VM migration from vSphere with minimal downtime, or Proxmox version upgrades
Deliverables
- Production Proxmox HA cluster with verified failover testing
- PXE boot infrastructure for zero-touch node provisioning
- Documented network architecture with VLAN map and switch configuration
- Backup and recovery runbook with tested restore procedures
- Storage performance baseline report per backend (IOPS, throughput, latency)
- Operational runbook covering common maintenance and failure scenarios
Tech stack
ProxmoxCephPXEZFSNVMe-oFHAProxyCorosync
Want to scope this?
Think this fits your needs? Let's scope it together.