Proxmox VE Clustering and Shared Storage

Proxmox VE Clustering and Shared Storage

Adding another node does not automatically make an environment resilient. This follow-on course explains how cluster decisions, networking, storage and managed workloads fit together, and where their dependencies can fail.

It is designed for multi-node infrastructure administrators who need to understand both planned maintenance and unexpected outages. We distinguish control-plane availability, access to data and the actual operation of an application.

Join official Proxmox VE training delivered by initMAX, an authorized Proxmox training partner. Your instructor is Tomáš Heřmánek, and the course includes a Certificate of Completion.

Explore the foundation course · Explore the full Bundle

Information about the course

Designed for the product: Proxmox VE
Focus: Clusters, shared storage and high availability
Who it is for: Multi-node environment administrators and infrastructure specialists
Prerequisites: Completion of Proxmox VE Deployment and Management
Practical work: Guest migration and simulated node, network and storage failures
Arrangements: Official training with an authorized Proxmox training partner
Course language: English or Czech — your choice for requested dates
Training format: Online · Hybrid · On-site
Previous course: Deployment and Management
Combined package: Proxmox VE Bundle
Course length: 14 h
Course price: 36 000,-€ 1,490$1730 excluding VAT

Course content

Designing a multi-node environment

 

Before creating a cluster, assess node dependencies, network paths and data placement. Separate centralized administration from surviving a failure. Include the capacity needed when one server is taken out of service.

In practice: Identify components whose failure can affect several nodes simultaneously.

Corosync, pmxcfs and quorum

 

Understand cluster-member communication and configuration distribution. Quorum protects cluster decision-making; it does not vote on whether an application is running. Assess loss of quorum against node state and communication paths.

In practice: Read cluster status and distinguish membership, configuration and guest-availability issues.

Redundant networking, VLANs and SDN

 

Design for management traffic, cluster communication, guest migration and storage traffic. Consider latency and concurrent transfers as well as addressing. Check whether supposedly redundant paths depend on the same component.

In practice: Assess a simulated link failure and the resulting availability of each traffic type.

Shared storage and data availability

 

Compare approaches to shared data and their relationship to migration. Distinguish storage configuration visible in cluster management from actual disk accessibility on the target node. Consider differences between file and block backends.

In practice: Verify the prerequisites for moving a guest between two nodes.

Ceph services, disks and hyperconvergence

 

Distinguish the roles of MON, MGR and OSD services and understand storage running alongside virtualization on the same servers. Plan disk capacity, CPU, memory and networking for both normal operation and recovery after a failure.

In practice: Read a Ceph health overview and relate each service to its purpose.

Ceph RBD, CephFS and operational maintenance

 

Distinguish RBD block devices from file access through CephFS. Explore the relationship between pools, data placement and resilience. Assess disk and node maintenance against remaining capacity and recovery status.

In practice: Determine what to check before planned storage maintenance using a sample cluster state.

Migrating virtual machines and containers

 

Assess a guest move against CPU compatibility, disk access, networking and locally attached devices. Evaluate VMs and containers separately: do not assume the same migration process or uninterrupted operation for every workload.

In practice: Prepare a migration preflight checklist and verify the service on the destination node.

Asynchronous ZFS replication

 

Understand replication-job scheduling, monitoring and its relationship to guest migration. A replica is not necessarily an up-to-the-moment copy: after a failure, recovery depends on the last successful transfer.

In practice: Assess a delayed replica and identify the data state that can be expected during recovery.

High availability, fencing and workload placement

 

Relate HA resources to node and data availability. Understand why the same workload must not start simultaneously on isolated nodes. Distinguish restarting a guest after failure from the availability of the application inside it.

In practice: Walk through HA decisions and check the prerequisites for safe workload recovery.

Failure scenarios and return to normal operation

 

Assess node, network and storage faults separately. Observe service availability and data state as well as the cluster response. Restored connectivity does not necessarily mean that storage recovery has completed.

In practice: Record a sample incident, expected behaviour and the conditions for concluding the intervention.

Two nodes and an external QDevice

 

Explore a small cluster with an external quorum vote. A QDevice assists quorum decisions; it adds neither compute capacity nor another copy of data. Consider the independence of its location and its connectivity.

In practice: Assess what happens when a node, interconnect or external vote becomes unavailable.

Automation and focused diagnostics

 

Understand how the web interface relates to the REST API and CLI tools. Separate authentication, authorization, task submission and result verification in automation. Focus diagnostics on the specific failing layer.

In practice: Design restricted automation access and a check that verifies the result of its request.

Tomáš Heřmánek

Tomáš Heřmánek

CEO & Zabbix Certified Trainer
Tomáš is a huge fan of OpenSource technologies, including Zabbix. He specializes in application servers, automation, and monitoring. Over the last ten years, he has been involved in several large-scale projects that have received extremely positive feedback. Tomáš holds a unique Zabbix Certified Trainer certificate from Zabbix SIA. Additionally, he enjoys sharing his expertise through training sessions and community events.

Additional information

What do I need before the follow-on course? 

The official prerequisite is completion of Deployment and Management. The follow-on course does not repeat the entire foundation and focuses on multi-node environments. If you have extensive experience but have not completed the first course, discuss a suitable route with us before registering; we do not promise an automatic exception.

Does HA guarantee operation without any outage? 

No. HA manages the availability of nodes and guests; it does not remove every application dependency. Cluster decisions, access to data and service restart can take time. The course therefore distinguishes planned moves, unexpected failures and actual verification of application functionality.

How do I choose a date and book training? 

Choose a date in the registration form on the course page. For individual or team training, contact us with the number of participants, their experience and your preferred period. We will prepare a proposal and agree on the dates, teaching language and delivery format to suit your needs.

Do I need my own production server for the course? 

Failure exercises should not be performed on your production systems. Practical work is intended for safely checking procedures in an isolated training environment. Computer, connectivity and lab-access requirements will be provided for each scheduled course. Please do not send production credentials in advance.

Does the number of hours mean a fixed number of days? 

The stated duration describes teaching time, not a fixed calendar schedule. Full-day or shorter blocks, breaks and the time zone will be specified for each scheduled course. Until the timetable is available, course duration alone is not enough to plan travel or book accommodation.

Will training also deliver our migration or architecture design? 

Training builds knowledge and provides practice. Assessment of your specific infrastructure, an implementation plan, downtime and responsibility for deployment are separate services. We can follow up with consulting, an environment review or a migration.

Which Proxmox VE version will be used? 

The training-environment version will be stated alongside each scheduled course. Explanations will distinguish general principles from version-specific behaviour. The version number is therefore not part of the course name or intended URL.

Certificate

Complete your official Proxmox VE training and receive a Certificate of Completion confirming your attendance and completion of the course.

Training is delivered by initMAX, an authorized Proxmox training partner. Your instructor is Tomáš Heřmánek.

initMAX - certificate
Other training courses

Skills and knowledge gained

Understanding cluster decisions

Distinguish cluster status, quorum and individual service availability, and know which evidence to compare during a fault.

Assessing independent paths

Check whether redundant networking or storage shares a hidden point of failure.

Reading Ceph health

Relate individual services and warnings to the layer that needs further investigation.

Preparing a guest move

Check the destination node, networking and data access before migration, then verify the service afterwards.

Separating HA from data protection

Distinguish guest recovery, replicas and backups. Each addresses a different part of the risk.

Verifying a failure scenario

Record what the cluster did, the service state that followed and what still needs checking before closing the incident.

What you will learn

Organizational impact

Your team is better prepared to plan maintenance and assess failure scenarios. A shared view of the cluster, storage and application operation helps define responsibilities and prepare verification tests.

Individual impact

Administrators gain the context needed to manage multiple nodes. They learn to distinguish communication, quorum, data-access and application problems and select focused checks.

Registration

"(Required)" indicates required fields

This field is for validation purposes and should be left unchanged.

Tell us your preferred training format and time zone.
Privacy(Required)

×Shopping Cart

Your cart is empty.