Known Issues in DataCore Puls8 v4.6
This section documents the known issues that may impact this release, including any available workarounds or mitigation steps.
The items below describe the known issues:
XFS Volumes Fail to Mount on Nodes Running Kernels Earlier Than 5.19
mkfs.xfs version 6.5 and later enable the nrext64 filesystem feature by default, and only Linux kernel 5.19 and later can mount a filesystem that has it. On a cluster where any node runs an older kernel, including RHEL 8 and 9, Ubuntu 20.04 and 22.04, and SLES 15, an XFS volume formatted by a newer mkfs.xfs fails to mount on that node with the error Superblock has unknown incompatible features (0x20) enabled. This affects Local PV LVM and Local PV ZFS volumes.
Workaround: Set the XFS default on the node agent so that volumes are formatted without the feature, or set formatOptions on the StorageClass. Refer to StorageClass Parameters. Existing volumes are unaffected, and clusters where every node runs kernel 5.19 or later require no change.
Kerberos NFS Mounts Fail with "Invalid argument" on Unprepared Nodes
Mounting an NFS RWX volume that uses krb5, krb5i, or krb5p fails on a node that does not have rpc-gssd running or does not have a host keytab installed. The kernel reports the failure as EINVAL ("Invalid argument"), which does not obviously indicate a Kerberos problem.
Workaround: Ensure every node that will mount Kerberos-secured NFS volumes has rpc-gssd running and a valid host keytab installed, as described in NFS RWX Using the Puls8 NFS Driver. DataCore Puls8 surfaces a Kerberos-specific hint on the PVC events for this failure.
NFS Server Pod Restarts Reset the NFSv4 Grace Period
When an NFS server pod is rescheduled or restarted, NFSv4 clients enter a grace period during which they reclaim their locks. Clients that do not reclaim within the grace period may observe a brief pause in I/O while the server completes recovery.
Workaround: No action is required; the grace period lifts as soon as all known clients have reclaimed. For workloads sensitive to this pause, assign a priority class to the NFS server pods through the StorageClass server.priorityClassName parameter so they are less likely to be evicted.
Snapshot Creation Fails After a Node Reboot for Pods Without a Controller
If a node hosting a pod reboots and that pod is not managed by a controller such as a Deployment or StatefulSet, the volume unpublish operation may not be triggered. The control plane continues to treat the volume as published, so the FIFREEZE operation fails during snapshot creation and the snapshot is retried without succeeding.
Workaround: Recreate or rebind the pod so that the volume is mounted and recognized correctly by the control plane. Alternatively, set quiesceFs to none in the VolumeSnapshotClass to take the snapshot without filesystem quiescing, at the cost of filesystem consistency. Refer to Volume Snapshots.
Slow Recovery for Large DiskPools After an Unclean Shutdown
A DiskPool in the 10-20 TiB range may take a considerable time to recover after an unclean shutdown of the node hosting its io-engine. The pool remains unavailable until recovery completes.
Workaround: No action is required. Allow the recovery to finish before returning the node to service, and account for this when planning maintenance windows on clusters with large pools.
Kernel Crash on Oracle Linux 9 During NVMe-TCP Volume Detach Operations (CVE-2024-53170)
When running Replicated PV Mayastor on Oracle Linux 9 (kernel 5.14.x), servers may unexpectedly reboot during volume detach operations due to a kernel bug (CVE-2024-53170) in the block layer. This issue is not caused by Replicated PV Mayastor but is triggered more frequently because of its NVMe-TCP connection lifecycle.
Workaround: Upgrade to kernel 6.11.11, 6.12.2, or later, which includes the fix.
Velero Uninstall Removes DataCore Puls8 Namespace and Components
When Velero is enabled via kubectl puls8, using the velero uninstall -n puls8 command is not recommended. This command attempts to delete the entire Puls8 namespace, expecting it to be used exclusively by Velero, and removes all Puls8 components.
Workaround: Instead of using the Velero CLI uninstall command, run the following command to safely remove only Velero-specific components:
kubectl puls8 velero uninstall --release-name puls8 --chart-name oci://registry-1.docker.io/datacoresoftware/puls8 -n puls8
Backup Hangs When Namespace Contains Replicated Volumes and Local Volumes
When backing up a namespace with a mix of Replicated PV Mayastor and Local PV Hostpath volumes, the Velero backup job includes all PVCs.
Workaround: Avoid creating Local PV Hostpath PVCs in the same namespace as Replicated PV Mayastor volumes. Alternatively, use a label selector during backup (Example: --selector backup=true) to limit the backup to specific resources.