Troubleshooting
Explore this Page
- Overview
- Storage Provisioning and Mounting Issues
- Backup Failures
- Restore Failures
- Infrastructure and Platform Issues
- Air-Gapped Installation Issues
- NFS ReadWriteMany (RWX) Issues
- Eventing and Observability Issues
Overview
This document provides guidance for identifying and resolving common issues encountered when deploying and operating DataCore Puls8 storage solutions, including Local Storage and Replicated Storage. It covers scenarios ranging from PVC provisioning failures to system-level incompatibilities, kernel constraints, and known behavioral limitations. You are encouraged to follow the documented workarounds and resolutions to ensure a stable and consistent experience in production and development environments.
Ensure that all system and platform prerequisites are met before troubleshooting. Refer to the Product Installation and Configuration documentation for environment-specific instructions.
Storage Provisioning and Mounting Issues
PVC Stuck in Pending State
Problem
A Persistent Volume Claim (PVC) created using the localpv-hostpath StorageClass remains in the Pending state, and no corresponding Persistent Volume (PV) is created.
Cause
The default Local PV StorageClasses use volumeBindingMode: WaitForFirstConsumer, which delays PV provisioning until the application pod is scheduled. If the pod specification includes a nodeName, the Kubernetes scheduler is bypassed, preventing volume provisioning.
Resolution
- Deploy the application that uses the PVC to trigger volume provisioning.
- Avoid setting the
nodeNamein the pod spec. Use a node selector instead:
Once the pod is scheduled, the PVC will be bound, and the PV will be created automatically.
All SCSI Devices Claimed in OpenShift
Problem
All SCSI devices on the node are claimed by the multipathd service, potentially disrupting volume device access.
Cause
The /etc/multipath.conf file is missing either the find_multipaths directive or an appropriate blacklist, causing multipathd to claim all available SCSI devices.
Resolution
Add the following to etc/multipath.conf:
Then run the following command to refresh the multipath configuration:
Replace /dev/sdc with the appropriate device name.
Unable to Mount XFS File System
Problem
A volume formatted with the XFS filesystem fails to mount when used by an application.
Cause
Nodes running Linux kernel versions earlier than 5.10 may not support certain options used by newer versions of xfsprogs, resulting in mount failures.
Resolution
Upgrade the kernel on affected nodes to version 5.10 or later to ensure compatibility with newer XFS filesystem features.
Backup Failures
DataUploader Pod or Node Fails Mid-Upload
Problem
A backup operation fails partially when the datauploader pod or its node becomes unavailable while uploading snapshot data to S3.
Cause
During a namespace backup, volume snapshots are created and restored to temporary volumes. These are mounted by the datauploader pod, which uploads them to S3 using Kopia. If the pod or node goes down during upload:
- Velero does not recreate the
datauploaderpod. - The temporary volume is deleted.
- The
DataUploadcustom resource transitions toFailed. - The overall backup is marked as
PartiallyFailed.
Resolution
There is no automatic recovery for this scenario. To recover:
- Manually re-trigger the backup.
- If the backup was created as part of a scheduled backup, the next scheduled job will attempt the backup again.
CSI / REST API / Core Agent Unavailable During Backup
Problem
Backup remains stuck in InProgress or fails after timeout due to snapshot creation failure. No datauploader pod is created.
Cause
If any of the CSI components, the REST API (app=api-rest), or the core agent are unavailable:
- VolumeSnapshots may not be created.
snapshot.status.readyToUse = false- The
datauploadpod is never scheduled. - After the default
csi-snapshot-timeoutof 10 minutes, the backup moves toPartiallyFailed.
Resolution
- Verify availability of CSI controller, REST server, and core agent.
- Restore CSI operations before the timeout (default 10 minutes) to allow the backup to proceed.
- If timeout is exceeded, re-trigger the backup.
- Optionally, increase the
csi-snapshot-timeoutwhen creating the backup to accommodate temporary delays.
Backup Fails Due to Insufficient Pool Capacity
Problem
Backup operation fails for one or more volumes due to lack of available storage capacity in the underlying pool.
Cause
Thick-provisioned volumes and high replica counts increase space requirements. If there is no enough capacity in the pool, snapshot creation fails, causing the backup to partially fail.
Test scenario (3-node cluster with 10 GB pool and 3-replica volumes):
| Volumes | Size (GB) | Result |
|---|---|---|
| 2 | 4 | Pass |
| 1 | 7 | Pass |
| 1 | 8 | Fail |
| 1 | 9 | Fail |
Resolution
- Increase pool capacity before triggering the backup again.
- In some cases, freeing space by deleting volumes before the 10-minute timeout can allow snapshots to succeed and the backup to complete.
- Analyze pool usage patterns and snapshot failure thresholds for better capacity planning.
Backup Fails When All Replicas Are Not Online
Problem
Backup operation fails for a volume when not all of its replicas are healthy or online.
Cause
Snapshots are only created if all replicas are available. If one or more child replicas are faulted, snapshot creation fails with the following error:
The number of healthy replicas does not match the expected replica count of volume '<volume-uuid>'
After the csi-snapshot-timeout (default 10 minutes), the backup enters PartiallyFailed state.
Resolution
- Ensure all replicas are online before initiating backup.
- If replica rebuild fails due to node count or topology constraints, consider:
- Scaling down the volume (Example: Reducing replica count).
- Resolving topology issues to allow rebuild.
- Increase snapshot timeout if you expect replicas to recover shortly.
Restore Failures
DataDownloader Pod Fails During Restore
Problem
Restore operation fails when the datadownloader pod is deleted or interrupted during data download from S3.
Cause
During restore:
- The PVC is recreated and mounted to a temporary volume.
- The
datadownloaderpod retrieves volume data from the S3 backup. - If the pod fails (Example: Due to node issue or eviction), Velero does not recreate it or resume download.
Resolution
- Clean up stale resources (Example: Partially restored volumes, PVCs, or custom resources).
- Re-trigger the restore operation manually.
- Ensure node stability and availability during restore to prevent interruptions.
Infrastructure and Platform Issues
IO Engine Fails to Start Due to IOVA Limit Error
Problem
The io-engine fails to start with the following error message:
Cause
The host node is likely to have IOMMU enabled, which might make default DMA mask width insufficient for IOVA supported address ranges.
Resolution
To resolve the issue, configure the IOVA mode to use physical addressing by setting the following Helm variable during the installation or upgrade:
This setting ensures that the io-engine operates in a mode compatible with the system’s DMA mask constraints.
Air-Gapped Installation Issues
MinIO Pods for Loki Fail to Start with ImagePullBackOff
Problem
The MinIO pods that provide object storage for Loki remain in ImagePullBackOff, and the image reference points at quay.io/minio.
Cause
DataCore Puls8 4.6.0 and earlier pulled the MinIO images from quay.io/minio, which can fail to serve them. From 4.6.1 the images are served from docker.io/openebs, with the same image tags.
Resolution
Upgrade to DataCore Puls8 4.6.1 or later, which moves the repositories as part of the upgrade. If you cannot upgrade yet, override the two repositories on your existing release.
helm upgrade puls8 datacore/puls8 --namespace puls8 --reuse-values \
--version <your-current-version> \
--set loki.minio.image.repository=docker.io/openebs/minio \
--set loki.minio.mcImage.repository=docker.io/openebs/mc
--reuse-values keeps your existing settings, including the previous repositories, so the two --set flags are required. An upgrade performed with the DataCore Puls8 plugin moves the repositories for you.
ImagePullBackOff on Grafana, Loki, NATS, MinIO, Velero, or External Secrets Pods
Problem
After an air-gapped installation, a pod belonging to Grafana Alloy, Loki, NATS, MinIO, Velero, kubectl, or External Secrets is stuck in ImagePullBackOff, even though other pods pulled successfully from the private registry.
Cause
These sub-charts do not honor the top-level global.imageRegistry setting and require their own explicit image overrides.
Resolution
Confirm that the corresponding per-sub-chart override is present in airgap-values.yaml, as described in DataCore Puls8 on Air-Gapped Environments, then apply the change with helm upgrade.
ImagePullBackOff with an HTTP 401 or 403
Problem
A pod fails to pull its image with an authentication error (HTTP 401 or 403) even though the private registry is reachable.
Cause
Registry credentials are not reaching the pod. Some sub-charts do not honor global.imagePullSecrets and instead need the pull secret attached directly to the service account their workload uses.
Resolution
Identify the service account used by the failing pod, attach the pull secret to it, then restart the workload.
# Identify the service account used by the failing pod
kubectl get pod <pod> -n puls8 -o jsonpath='{.spec.serviceAccountName}{"\n"}'
# Attach the pull secret to that service account
kubectl patch serviceaccount <service-account> -n puls8 \
-p '{"imagePullSecrets":[{"name":"puls8-regcred"}]}'
# Restart the workload so new pods pick up the secret
kubectl rollout restart deploy/<name> -n puls8
manifest unknown or an Unexpected Image Tag
Problem
A pod fails to pull its image with a manifest unknown error, or an unexpected image tag appears in the cluster.
Cause
The Helm chart, the images.txt file, and the image tags mirrored to the private registry came from different Puls8 releases.
Resolution
Repeat the download and mirroring steps in DataCore Puls8 on Air-Gapped Environments using a single, matched release for the chart, the image list, and the mirrored tags.
Pod Fails to Start on a Node of a Different Architecture
Problem
A pod fails to start on a node whose CPU architecture differs from the host used to mirror images (for example, amd64 images on an arm64 node).
Cause
docker save and docker load carry only the architecture that was pulled, so a single-architecture mirror cannot satisfy nodes of a different architecture.
Resolution
Re-mirror the images using skopeo copy --all to preserve full multi-architecture manifests, or pull with --platform matching your node architecture.
NFS ReadWriteMany (RWX) Issues
The following issues are specific to volumes provisioned with the DataCore Puls8 NFS driver (provisioner com.datacore.puls8.nfs).
PVC Stuck in Pending State
Problem
An NFS RWX PVC (or a snapshot) remains in the Pending state and no volume is bound.
Resolution
Inspect the PVC and cluster events, then match the event message against the causes below.
kubectl describe pvc <pvc> -n <ns>
kubectl get events -n <ns> --sort-by=.lastTimestamp
valid license required: No active license. Install or activate a valid license (see License Activation).backendStorageClass not found / missing: The backend StorageClass is wrong or missing. Correct thebackendStorageClassparameter.missing kerberos: Akrb5*security mode is set without akerberosblock. Add the block.keytabSecret and kadminSecret both set: Both Kerberos keytab sources are provided. Keep exactly one.backend PVC not found / not Bound / not RWO / capacity: The adoption fitness check failed. Fix the backend PVC (it must be in the driver namespace, Bound, RWO, and large enough).already claimed by: The backend PVC already backs another NFS volume. Use a different backend PVC.snapshot not ready: The restore source is not yetreadyToUse. Wait; the PVC binds automatically when the snapshot is ready.
Mount Hangs or Fails from the Application Pod
Problem
An application pod cannot mount the NFS volume, or the mount hangs.
Resolution
Inspect the NFS server pod for the volume in the driver namespace.
kubectl get pods -n puls8 | grep nfs-server
kubectl logs -n puls8 <nfs-server-pod>
kubectl describe pod -n puls8 <nfs-server-pod>
- FailedMount on a Kerberos volume: The node is missing its machine keytab or
rpc-gssd/rpc_pipefs. Also check clock skew, as Kerberos is time-sensitive. - Server pod Pending or Unschedulable: The
server.nodeSelectorortolerationsmatch no node. - Server pod not Ready with LDAP: SSSD cannot reach LDAP. A malformed
ldapblock or an unreachable server surfaces at the first mount, not at provisioning.
Permission Denied on Files
Problem
Files return "permission denied", or ownership displays as nobody/65534.
Resolution
- sys mode: Run the pod as the UID/GID that owns the files.
- Root squash active: A root container is mapped to anonymous (65534). Either run non-root or use
No_Root_Squash. - Kerberos + LDAP: If files show as
nobody/65534, LDAP identity resolution is down and the server cannot map principals to UIDs. Check LDAP reachability and the bind Secret. - Display only: NFSv4 may show
nobody:nogroupwhen it cannot resolve a number to a name; the numeric permission check can still be correct.
The group is squashed on the same rule, and a pod hits this without running as root: leave runAsGroup out and the primary group is 0, which squashes to anonymousGid even though runAsUser is an ordinary UID.
Files are readable by pods that should not have access. This is the opposite symptom, and the more dangerous one, because nothing fails. Check the group the files were actually created with:
kubectl exec -n puls8 <nfs-server-pod> -c nfs-server -- ls -ln /data/export/<dir>
A group of 0 or 65534 on files you expected to belong to a real group means the writing pod had no runAsGroup; every pod defaults to that same group, so the group permission bits grant access to all of them. Fix the pod spec and then re-create or chgrp the affected files - changing the spec alone does not relabel what is already on disk.
Kerberos Authentication Failures
Problem
Mounting a krb5* volume fails to authenticate.
Resolution
- Confirm a valid ticket in the client pod (
klist). - Confirm the node keytab and
rpc-gssdare present and active. - Check clock synchronization across the nodes and the KDC.
LDAP / SSSD Diagnostics
When LDAP identity mapping is not working - files show as nobody, or the SSSD sidecar will not start - check SSSD from inside the NFS server pod. The SSSD sidecar carries sssd and sss_cache but no ldapsearch or getent, so run directory queries (ldapsearch, getent) from a separate pod that has LDAP client tools, not from the sidecar.
# The socket the server talks to SSSD through - absent means SSSD never came up
kubectl exec -n puls8 <nfs-server-pod> -c nfs-server -- ls -l /var/lib/sss/pipes/nss
# What SSSD itself is saying; bind and TLS failures appear here
kubectl logs -n puls8 <nfs-server-pod> -c sssd-sidecar
# Drop cached lookups after fixing a directory entry
kubectl exec -n puls8 <nfs-server-pod> -c sssd-sidecar -- sss_cache -E
Whether resolution is working shows up in file ownership: a file written by an LDAP user should carry that user's UID and GID from the directory. Files owned by 65534/nobody when you expected a real user mean the server could not resolve the principal - check the socket and the sidecar log, in that order.
SSSD logs a SELinux warning on every start (SELINUX_getpeercon failed). On a cluster without SELinux this is expected and is not the cause of a failure.
If SSSD will not start, verify the bind Secret has the keys bindDn and password spelled exactly, that the LDAP server is reachable from inside the cluster, and that tlsEnabled matches what your LDAP server offers.
Collecting Information for a Support Ticket
When raising an issue, attach the following. Run it with your NFS PVC's namespace and name.
NS=my-app; PVC=shared-data
PV=$(kubectl -n $NS get pvc $PVC -o jsonpath='{.spec.volumeName}')
# The volume and its server pod
kubectl -n $NS describe pvc $PVC
kubectl get po -A -l app.kubernetes.io/instance=$PV -o wide
# Server logs, including the previous container if it restarted
kubectl -n puls8 logs <server-pod> --all-containers --tail=-1
kubectl -n puls8 logs <server-pod> --all-containers --previous
# What the clients are doing
kubectl -n $NS describe pod <your-pod>
The server pod's settings live in a ConfigMap named nfs-config-<volume-id> in the driver namespace, which support may request. Do not edit it - the driver owns and rewrites it.
Cleaning Up a Stuck Volume
If a volume is stuck during deletion (a finalizer prevents removal), the driver normally removes finalizers during DeleteVolume. If the controller is down, remove them manually, then delete any per-volume resources that remain in the driver namespace.
# Check what is holding the finalizer
kubectl get pvc <pvc-name> -n <namespace> -o jsonpath='{.metadata.finalizers}'
# Remove finalizers manually (only if the controller is down)
kubectl patch pvc <pvc-name> -n <namespace> \
-p '{"metadata":{"finalizers":null}}' --type=merge
# Remove per-volume resources if they remain (driver namespace)
kubectl delete statefulset nfs-server-<vol> -n puls8
kubectl delete service nfs-svc-<vol> -n puls8
kubectl delete configmap nfs-config-<vol> -n puls8
kubectl delete secret nfs-sssd-<vol> -n puls8
kubectl delete pvc nfs-backend-<vol> -n puls8
Eventing and Observability Issues
The issues in this section relate to the Eventing Aggregator and to the retrieval of cluster events through the DataCore Puls8 kubectl plugin.
Eventing Aggregator Pod Remains in the Init State
Problem
The Eventing Aggregator pod does not progress past initialization and remains in the Init state.
Cause
The aggregator pod runs an init container that waits for NATS to become reachable before the main container starts. NATS is either not running or not yet ready.
Resolution
Verify that the NATS pods are running in the namespace where DataCore Puls8 is installed. Once NATS is ready, the init container completes and the aggregator starts without further intervention.
No Events Are Returned
Problem
The command to retrieve events completes successfully but returns no events.
Cause
Events are collected only when both eventing and the Eventing Aggregator are enabled. An empty result can also indicate that the requested time window contains no events, or that the applied filters are too specific.
Resolution
Confirm that eventing and the aggregator are both enabled, and that the aggregator pod is running. Then widen the time window and remove all filters, to establish whether any events exist.
When the filters exclude every record, the plugin reports the number of events retrieved before filtering. This distinguishes an empty dataset from a filter that is too specific.
Events Older Than the Requested Time Window Are Missing
Problem
A query for a longer time window returns only recent events.
Cause
When Loki is not deployed, only the events held on the volume of the aggregator pod are available. That volume is ephemeral, so it is cleared whenever the aggregator pod restarts, and it holds a bounded number of events governed by the dirSizeLimit value.
Resolution
Deploy Loki for event history that survives a pod restart. Alternatively, increase the size limit of the aggregator volume to retain a longer window between restarts.
--set openebs.mayastor.eventing.aggregator.dirSizeLimit=500Mi
Querying Fails When a Loki Endpoint Is Specified
Problem
The command returns an error when --loki-endpoint is supplied, although it succeeds when the option is omitted.
Cause
Supplying --loki-endpoint signals that Loki is expected to be used, so the plugin does not fall back to reading the volume of the aggregator pod. An unreachable endpoint therefore produces an immediate error rather than results from another source.
Resolution
Confirm that the specified endpoint is reachable, or omit the option so that the plugin selects its source automatically.
Fewer Events Are Returned Than the Requested Time Window
Problem
A query with a long --since window returns fewer events than expected, even though Loki is deployed.
Cause
Loki enforces both a retention period and a maximum query range, the latter defaulting to 30 days. A request for a window beyond either limit returns only the events that fall within it.
Resolution
Adjust the retention_period value and the maximum query range in the Loki limits_config as required. Setting the maximum query range to 0 removes the limit.
Learn More