Volume Capacity and Label Versions
Explore this Page
- Overview
- How the Cluster Determines the Label Version
- How Capacity Is Calculated
- What This Means for Usable Capacity
- Resizing a Volume
- Checking the Label Version of a Volume
- Considerations and Limitations
- Benefits of the V2 Label
Overview
DataCore Puls8 Replicated PV Mayastor uses a versioned on-disk layout for volume replicas, referred to as V1 and V2. V2 is the later layout, and a cluster uses it once every node supports it.
The layout version determines how much pool capacity each replica consumes, and how much of the requested capacity is available to a workload. Under V1 the usable capacity of a volume is slightly smaller than the size requested. Under V2 the requested capacity is delivered in full, at the cost of a small amount of additional pool space for each replica.
This document describes the two layout versions, how the cluster determines which version to use, the effect of V2 on pool capacity, and how to plan capacity for volumes that use it.
How the Cluster Determines the Label Version
Each io-engine node reports the label version it supports when it registers with the control plane. The cluster adopts the lowest version reported by any registered node, and recalculates it whenever a node registers. A single node that supports only V1 therefore holds the entire cluster at V1, which keeps a cluster consistent while an upgrade is in progress.
A cluster whose nodes all support V2 uses V2, which is the case for a new installation. A cluster upgraded from a release that predates V2 continues to use V1 until every node supports V2.
A volume is assigned the label version of the cluster at the moment it is created, and keeps that version for its lifetime. A volume created from a snapshot, or as a clone of an existing volume, inherits the label version of its source instead. Only a volume created from new takes the version currently negotiated by the cluster.
An existing V1 volume, and anything restored or cloned from it, therefore remains V1 after the cluster advances to V2. Upgrading a cluster does not change the capacity consumed by volumes that already exist.
The size requested in the PersistentVolumeClaim is preserved in the metadata of the volume, so the original request is retained even where the reported volume size is larger.
The label version is not a StorageClass parameter. It is negotiated by the cluster and cannot be selected for an individual volume.
The Cluster Version Only Advances
Once the label version of a cluster has advanced it is not lowered again. An attempt to lower it is rejected and recorded in the control plane log. Adding a node that supports only V1 to a cluster already running V2 therefore does not return the cluster to V1.
How Capacity Is Calculated
The layout version determines the size of a volume and the size of each of its replicas.
| Property | V1 | V2 |
|---|---|---|
| Volume size | The requested size, unchanged | The requested size rounded up to 1 MiB |
| Replica size | The volume size | The volume size plus 8 MiB of metadata overhead, rounded up to the pool cluster size |
| Block device size | Not reported | The aligned volume size |
The default pool cluster size is 4 MiB. The calculation applies to every replica independently, so multiply the replica size by the replica count of the volume to obtain the total pool capacity it consumes.
Example: a 10 MiB V2 Volume at the Default Pool Cluster Size
- Volume size. 10 MiB is already a multiple of 1 MiB, so it is unchanged.
- Replica size. 10 MiB plus 8 MiB of metadata overhead is 18 MiB, which rounds up to 20 MiB at a 4 MiB cluster size.
- Block device size. 10 MiB.
The volume presents a 10 MiB device to the workload, and each of its replicas consumes 20 MiB of pool space.
A size expressed in decimal units aligns differently from one expressed in binary units. A request for 10 MB is 9.54 MiB, which rounds up to a volume size of 10 MiB. Express sizes in binary units to keep the consumed capacity predictable.
What This Means for Usable Capacity
Under V1 the space for metadata was taken from the requested size, so a workload saw approximately 6 MiB less than it asked for. Under V2 that space is added to the replica instead, so the full requested capacity is available.
When sizing pools for an upgrade, allow up to 8 MiB more for each replica. The reported replica size increases by more than that, because a V1 replica was already cluster-aligned on disk without reporting it.
Resizing a Volume
A resize applies the same two steps to the new size. The capacity a resized volume consumes is therefore determined by its label version and the cluster size of the pool, and not by the size requested in the claim alone. Refer to the Volume Resize documentation for the procedure.
Checking the Label Version of a Volume
The label version of a volume is reported in its specification. Retrieve the volume in YAML and read the value from the spec section. A value of 1 indicates the V1 layout and a value of 2 indicates V2.
Considerations and Limitations
- The metadata overhead applies to each replica. A volume with three replicas consumes three times the replica size across the cluster.
- Small volumes carry a disproportionate overhead. Plan capacity from the replica size rather than the requested size when provisioning a large number of small volumes.
- Existing volumes are not converted when a cluster advances to V2, so a cluster can hold volumes of both versions at the same time.
- Thin provisioning does not avoid the metadata overhead. The overhead is reserved when the replica is created, irrespective of how much of the volume is written. Refer to the Thin Provisioning documentation for more information.
Benefits of the V2 Label
- Requested Capacity Is Delivered: The usable capacity of a volume is equal to or greater than the size requested, where under V1 it was slightly smaller.
- Predictable Alignment: Volume sizes align to megabyte boundaries, so replica layout is consistent regardless of the size requested.
- Accurate Capacity Reporting: The reported replica size reflects what a replica actually occupies on the backend.
Learn More