Troubleshooting

Explore this Page

Overview

This document provides guidance for identifying and resolving common issues encountered when deploying and operating DataCore Puls8 storage solutions, including Local Storage and Replicated Storage. It covers scenarios ranging from PVC provisioning failures to system-level incompatibilities, kernel constraints, and known behavioral limitations. You are encouraged to follow the documented workarounds and resolutions to ensure a stable and consistent experience in production and development environments.

Ensure that all system and platform prerequisites are met before troubleshooting. Refer to the Product Installation and Configuration documentation for environment-specific instructions.

Storage Provisioning and Mounting Issues

PVC Stuck in Pending State

Problem

A Persistent Volume Claim (PVC) created using the localpv-hostpath StorageClass remains in the Pending state, and no corresponding Persistent Volume (PV) is created.

Cause

The default Local PV StorageClasses use volumeBindingMode: WaitForFirstConsumer, which delays PV provisioning until the application pod is scheduled. If the pod specification includes a nodeName, the Kubernetes scheduler is bypassed, preventing volume provisioning.

Resolution

  • Deploy the application that uses the PVC to trigger volume provisioning.
  • Avoid setting the nodeName in the pod spec. Use a node selector instead:
  • Copy
    YAML
    nodeSelector:
      kubernetes.io/hostname: <desired-node-name>

Once the pod is scheduled, the PVC will be bound, and the PV will be created automatically.

All SCSI Devices Claimed in OpenShift

Problem

All SCSI devices on the node are claimed by the multipathd service, potentially disrupting volume device access.

Cause

The /etc/multipath.conf file is missing either the find_multipaths directive or an appropriate blacklist, causing multipathd to claim all available SCSI devices.

Resolution

Add the following to etc/multipath.conf:

Copy
conf
defaults {
    user_friendly_names yes
    find_multipaths yes
}

Then run the following command to refresh the multipath configuration:

Copy
Refresh Multipath Configuration
multipath -w /dev/sdc

Replace /dev/sdc with the appropriate device name.

Unable to Mount XFS File System

Problem

A volume formatted with the XFS filesystem fails to mount when used by an application.

Cause

Nodes running Linux kernel versions earlier than 5.10 may not support certain options used by newer versions of xfsprogs, resulting in mount failures.

Resolution

Upgrade the kernel on affected nodes to version 5.10 or later to ensure compatibility with newer XFS filesystem features.

Backup Failures

DataUploader Pod or Node Fails Mid-Upload

Problem

A backup operation fails partially when the datauploader pod or its node becomes unavailable while uploading snapshot data to S3.

Cause

During a namespace backup, volume snapshots are created and restored to temporary volumes. These are mounted by the datauploader pod, which uploads them to S3 using Kopia. If the pod or node goes down during upload:

  • Velero does not recreate the datauploader pod.
  • The temporary volume is deleted.
  • The DataUpload custom resource transitions to Failed.
  • The overall backup is marked as PartiallyFailed.

Resolution

There is no automatic recovery for this scenario. To recover:

  • Manually re-trigger the backup.
  • If the backup was created as part of a scheduled backup, the next scheduled job will attempt the backup again.

CSI / REST API / Core Agent Unavailable During Backup

Problem

Backup remains stuck in InProgress or fails after timeout due to snapshot creation failure. No datauploader pod is created.

Cause

If any of the CSI components, the REST API (app=api-rest), or the core agent are unavailable:

  • VolumeSnapshots may not be created.
  • snapshot.status.readyToUse = false
  • The dataupload pod is never scheduled.
  • After the default csi-snapshot-timeout of 10 minutes, the backup moves to PartiallyFailed.

Resolution

  • Verify availability of CSI controller, REST server, and core agent.
  • Restore CSI operations before the timeout (default 10 minutes) to allow the backup to proceed.
  • If timeout is exceeded, re-trigger the backup.
  • Optionally, increase the csi-snapshot-timeout when creating the backup to accommodate temporary delays.

Backup Fails Due to Insufficient Pool Capacity

Problem

Backup operation fails for one or more volumes due to lack of available storage capacity in the underlying pool.

Cause

Thick-provisioned volumes and high replica counts increase space requirements. If there is no enough capacity in the pool, snapshot creation fails, causing the backup to partially fail.

Test scenario (3-node cluster with 10 GB pool and 3-replica volumes):

Volumes Size (GB) Result
2 4 Pass
1 7 Pass
1 8 Fail
1 9 Fail

Resolution

  • Increase pool capacity before triggering the backup again.
  • In some cases, freeing space by deleting volumes before the 10-minute timeout can allow snapshots to succeed and the backup to complete.
  • Analyze pool usage patterns and snapshot failure thresholds for better capacity planning.

Backup Fails When All Replicas Are Not Online

Problem

Backup operation fails for a volume when not all of its replicas are healthy or online.

Cause

Snapshots are only created if all replicas are available. If one or more child replicas are faulted, snapshot creation fails with the following error:

Copy
Error
The number of healthy replicas does not match the expected replica count of volume '<volume-uuid>'

After the csi-snapshot-timeout (default 10 minutes), the backup enters PartiallyFailed state.

Resolution

  • Ensure all replicas are online before initiating backup.
  • If replica rebuild fails due to node count or topology constraints, consider:
    • Scaling down the volume (Example: Reducing replica count).
    • Resolving topology issues to allow rebuild.
  • Increase snapshot timeout if you expect replicas to recover shortly.

Restore Failures

DataDownloader Pod Fails During Restore

Problem

Restore operation fails when the datadownloader pod is deleted or interrupted during data download from S3.

Cause

During restore:

  • The PVC is recreated and mounted to a temporary volume.
  • The datadownloader pod retrieves volume data from the S3 backup.
  • If the pod fails (Example: Due to node issue or eviction), Velero does not recreate it or resume download.

Resolution

  • Clean up stale resources (Example: Partially restored volumes, PVCs, or custom resources).
  • Re-trigger the restore operation manually.
  • Ensure node stability and availability during restore to prevent interruptions.

Infrastructure and Platform Issues

IO Engine Fails to Start Due to IOVA Limit Error

Problem

The io-engine fails to start with the following error message:

Copy
Error
Couldn't allocate memory due to IOVA exceeding limits of current DMA mask

Cause

The host node is likely to have IOMMU enabled, which might make default DMA mask width insufficient for IOVA supported address ranges.

Resolution

To resolve the issue, configure the IOVA mode to use physical addressing by setting the following Helm variable during the installation or upgrade:

Copy
Set IOVA mode to 'pa'
--set openebs.mayastor.io_engine.envcontext=iova-mode=pa

This setting ensures that the io-engine operates in a mode compatible with the system’s DMA mask constraints.

Air-Gapped Installation Issues

MinIO Pods for Loki Fail to Start with ImagePullBackOff

Problem

The MinIO pods that provide object storage for Loki remain in ImagePullBackOff, and the image reference points at quay.io/minio.

Cause

DataCore Puls8 4.6.0 and earlier pulled the MinIO images from quay.io/minio, which can fail to serve them. From 4.6.1 the images are served from docker.io/openebs, with the same image tags.

Resolution

Upgrade to DataCore Puls8 4.6.1 or later, which moves the repositories as part of the upgrade. If you cannot upgrade yet, override the two repositories on your existing release.

Copy
Override the MinIO Repositories on an Existing Release
helm upgrade puls8 datacore/puls8 --namespace puls8 --reuse-values \
  --version <your-current-version> \
  --set loki.minio.image.repository=docker.io/openebs/minio \
  --set loki.minio.mcImage.repository=docker.io/openebs/mc

--reuse-values keeps your existing settings, including the previous repositories, so the two --set flags are required. An upgrade performed with the DataCore Puls8 plugin moves the repositories for you.

ImagePullBackOff on Grafana, Loki, NATS, MinIO, Velero, or External Secrets Pods

Problem

After an air-gapped installation, a pod belonging to Grafana Alloy, Loki, NATS, MinIO, Velero, kubectl, or External Secrets is stuck in ImagePullBackOff, even though other pods pulled successfully from the private registry.

Cause

These sub-charts do not honor the top-level global.imageRegistry setting and require their own explicit image overrides.

Resolution

Confirm that the corresponding per-sub-chart override is present in airgap-values.yaml, as described in DataCore Puls8 on Air-Gapped Environments, then apply the change with helm upgrade.

ImagePullBackOff with an HTTP 401 or 403

Problem

A pod fails to pull its image with an authentication error (HTTP 401 or 403) even though the private registry is reachable.

Cause

Registry credentials are not reaching the pod. Some sub-charts do not honor global.imagePullSecrets and instead need the pull secret attached directly to the service account their workload uses.

Resolution

Identify the service account used by the failing pod, attach the pull secret to it, then restart the workload.

Copy
Attach the Pull Secret to a Workload's Service Account
# Identify the service account used by the failing pod
kubectl get pod <pod> -n puls8 -o jsonpath='{.spec.serviceAccountName}{"\n"}'

# Attach the pull secret to that service account
kubectl patch serviceaccount <service-account> -n puls8 \
  -p '{"imagePullSecrets":[{"name":"puls8-regcred"}]}'

# Restart the workload so new pods pick up the secret
kubectl rollout restart deploy/<name> -n puls8

manifest unknown or an Unexpected Image Tag

Problem

A pod fails to pull its image with a manifest unknown error, or an unexpected image tag appears in the cluster.

Cause

The Helm chart, the images.txt file, and the image tags mirrored to the private registry came from different Puls8 releases.

Resolution

Repeat the download and mirroring steps in DataCore Puls8 on Air-Gapped Environments using a single, matched release for the chart, the image list, and the mirrored tags.

Pod Fails to Start on a Node of a Different Architecture

Problem

A pod fails to start on a node whose CPU architecture differs from the host used to mirror images (for example, amd64 images on an arm64 node).

Cause

docker save and docker load carry only the architecture that was pulled, so a single-architecture mirror cannot satisfy nodes of a different architecture.

Resolution

Re-mirror the images using skopeo copy --all to preserve full multi-architecture manifests, or pull with --platform matching your node architecture.

NFS ReadWriteMany (RWX) Issues

The following issues are specific to volumes provisioned with the DataCore Puls8 NFS driver (provisioner com.datacore.puls8.nfs).

PVC Stuck in Pending State

Problem

An NFS RWX PVC (or a snapshot) remains in the Pending state and no volume is bound.

Resolution

Inspect the PVC and cluster events, then match the event message against the causes below.

Copy
Inspect the PVC and Events
kubectl describe pvc <pvc> -n <ns>
kubectl get events -n <ns> --sort-by=.lastTimestamp
  • valid license required: No active license. Install or activate a valid license (see License Activation).
  • backendStorageClass not found / missing: The backend StorageClass is wrong or missing. Correct the backendStorageClass parameter.
  • missing kerberos: A krb5* security mode is set without a kerberos block. Add the block.
  • keytabSecret and kadminSecret both set: Both Kerberos keytab sources are provided. Keep exactly one.
  • backend PVC not found / not Bound / not RWO / capacity: The adoption fitness check failed. Fix the backend PVC (it must be in the driver namespace, Bound, RWO, and large enough).
  • already claimed by: The backend PVC already backs another NFS volume. Use a different backend PVC.
  • snapshot not ready: The restore source is not yet readyToUse. Wait; the PVC binds automatically when the snapshot is ready.

Mount Hangs or Fails from the Application Pod

Problem

An application pod cannot mount the NFS volume, or the mount hangs.

Resolution

Inspect the NFS server pod for the volume in the driver namespace.

Copy
Inspect the NFS Server Pod
kubectl get pods -n puls8 | grep nfs-server
kubectl logs -n puls8 <nfs-server-pod>
kubectl describe pod -n puls8 <nfs-server-pod>
  • FailedMount on a Kerberos volume: The node is missing its machine keytab or rpc-gssd/rpc_pipefs. Also check clock skew, as Kerberos is time-sensitive.
  • Server pod Pending or Unschedulable: The server.nodeSelector or tolerations match no node.
  • Server pod not Ready with LDAP: SSSD cannot reach LDAP. A malformed ldap block or an unreachable server surfaces at the first mount, not at provisioning.

Permission Denied on Files

Problem

Files return "permission denied", or ownership displays as nobody/65534.

Resolution

  • sys mode: Run the pod as the UID/GID that owns the files.
  • Root squash active: A root container is mapped to anonymous (65534). Either run non-root or use No_Root_Squash.
  • Kerberos + LDAP: If files show as nobody/65534, LDAP identity resolution is down and the server cannot map principals to UIDs. Check LDAP reachability and the bind Secret.
  • Display only: NFSv4 may show nobody:nogroup when it cannot resolve a number to a name; the numeric permission check can still be correct.

The group is squashed on the same rule, and a pod hits this without running as root: leave runAsGroup out and the primary group is 0, which squashes to anonymousGid even though runAsUser is an ordinary UID.

Files are readable by pods that should not have access. This is the opposite symptom, and the more dangerous one, because nothing fails. Check the group the files were actually created with:

Copy
Check the Group on Files
kubectl exec -n puls8 <nfs-server-pod> -c nfs-server -- ls -ln /data/export/<dir>

A group of 0 or 65534 on files you expected to belong to a real group means the writing pod had no runAsGroup; every pod defaults to that same group, so the group permission bits grant access to all of them. Fix the pod spec and then re-create or chgrp the affected files - changing the spec alone does not relabel what is already on disk.

Kerberos Authentication Failures

Problem

Mounting a krb5* volume fails to authenticate.

Resolution

  • Confirm a valid ticket in the client pod (klist).
  • Confirm the node keytab and rpc-gssd are present and active.
  • Check clock synchronization across the nodes and the KDC.

LDAP / SSSD Diagnostics

When LDAP identity mapping is not working - files show as nobody, or the SSSD sidecar will not start - check SSSD from inside the NFS server pod. The SSSD sidecar carries sssd and sss_cache but no ldapsearch or getent, so run directory queries (ldapsearch, getent) from a separate pod that has LDAP client tools, not from the sidecar.

Copy
Check SSSD from the NFS Server Pod
# The socket the server talks to SSSD through - absent means SSSD never came up
kubectl exec -n puls8 <nfs-server-pod> -c nfs-server -- ls -l /var/lib/sss/pipes/nss

# What SSSD itself is saying; bind and TLS failures appear here
kubectl logs -n puls8 <nfs-server-pod> -c sssd-sidecar

# Drop cached lookups after fixing a directory entry
kubectl exec -n puls8 <nfs-server-pod> -c sssd-sidecar -- sss_cache -E

Whether resolution is working shows up in file ownership: a file written by an LDAP user should carry that user's UID and GID from the directory. Files owned by 65534/nobody when you expected a real user mean the server could not resolve the principal - check the socket and the sidecar log, in that order.

SSSD logs a SELinux warning on every start (SELINUX_getpeercon failed). On a cluster without SELinux this is expected and is not the cause of a failure.

If SSSD will not start, verify the bind Secret has the keys bindDn and password spelled exactly, that the LDAP server is reachable from inside the cluster, and that tlsEnabled matches what your LDAP server offers.

Collecting Information for a Support Ticket

When raising an issue, attach the following. Run it with your NFS PVC's namespace and name.

Copy
Collect NFS Volume Diagnostics
NS=my-app; PVC=shared-data
PV=$(kubectl -n $NS get pvc $PVC -o jsonpath='{.spec.volumeName}')

# The volume and its server pod
kubectl -n $NS describe pvc $PVC
kubectl get po -A -l app.kubernetes.io/instance=$PV -o wide

# Server logs, including the previous container if it restarted
kubectl -n puls8 logs <server-pod> --all-containers --tail=-1
kubectl -n puls8 logs <server-pod> --all-containers --previous

# What the clients are doing
kubectl -n $NS describe pod <your-pod>

The server pod's settings live in a ConfigMap named nfs-config-<volume-id> in the driver namespace, which support may request. Do not edit it - the driver owns and rewrites it.

Cleaning Up a Stuck Volume

If a volume is stuck during deletion (a finalizer prevents removal), the driver normally removes finalizers during DeleteVolume. If the controller is down, remove them manually, then delete any per-volume resources that remain in the driver namespace.

Copy
Clear a Stuck NFS Volume and Its Finalizers
# Check what is holding the finalizer
kubectl get pvc <pvc-name> -n <namespace> -o jsonpath='{.metadata.finalizers}'

# Remove finalizers manually (only if the controller is down)
kubectl patch pvc <pvc-name> -n <namespace> \
  -p '{"metadata":{"finalizers":null}}' --type=merge

# Remove per-volume resources if they remain (driver namespace)
kubectl delete statefulset nfs-server-<vol>  -n puls8
kubectl delete service     nfs-svc-<vol>     -n puls8
kubectl delete configmap   nfs-config-<vol>  -n puls8
kubectl delete secret      nfs-sssd-<vol>    -n puls8
kubectl delete pvc         nfs-backend-<vol> -n puls8

Eventing and Observability Issues

The issues in this section relate to the Eventing Aggregator and to the retrieval of cluster events through the DataCore Puls8 kubectl plugin.

Eventing Aggregator Pod Remains in the Init State

Problem

The Eventing Aggregator pod does not progress past initialization and remains in the Init state.

Cause

The aggregator pod runs an init container that waits for NATS to become reachable before the main container starts. NATS is either not running or not yet ready.

Resolution

Verify that the NATS pods are running in the namespace where DataCore Puls8 is installed. Once NATS is ready, the init container completes and the aggregator starts without further intervention.

Copy
List the Pods in the Puls8 Namespace
kubectl get pods -n puls8

No Events Are Returned

Problem

The command to retrieve events completes successfully but returns no events.

Cause

Events are collected only when both eventing and the Eventing Aggregator are enabled. An empty result can also indicate that the requested time window contains no events, or that the applied filters are too specific.

Resolution

Confirm that eventing and the aggregator are both enabled, and that the aggregator pod is running. Then widen the time window and remove all filters, to establish whether any events exist.

Copy
Query Events Across a Wider Time Window
kubectl puls8 mayastor get events -n puls8 --since 7d

When the filters exclude every record, the plugin reports the number of events retrieved before filtering. This distinguishes an empty dataset from a filter that is too specific.

Events Older Than the Requested Time Window Are Missing

Problem

A query for a longer time window returns only recent events.

Cause

When Loki is not deployed, only the events held on the volume of the aggregator pod are available. That volume is ephemeral, so it is cleared whenever the aggregator pod restarts, and it holds a bounded number of events governed by the dirSizeLimit value.

Resolution

Deploy Loki for event history that survives a pod restart. Alternatively, increase the size limit of the aggregator volume to retain a longer window between restarts.

Copy
Flag to Increase the Aggregator Volume Size Limit
--set openebs.mayastor.eventing.aggregator.dirSizeLimit=500Mi

Querying Fails When a Loki Endpoint Is Specified

Problem

The command returns an error when --loki-endpoint is supplied, although it succeeds when the option is omitted.

Cause

Supplying --loki-endpoint signals that Loki is expected to be used, so the plugin does not fall back to reading the volume of the aggregator pod. An unreachable endpoint therefore produces an immediate error rather than results from another source.

Resolution

Confirm that the specified endpoint is reachable, or omit the option so that the plugin selects its source automatically.

Fewer Events Are Returned Than the Requested Time Window

Problem

A query with a long --since window returns fewer events than expected, even though Loki is deployed.

Cause

Loki enforces both a retention period and a maximum query range, the latter defaulting to 30 days. A request for a window beyond either limit returns only the events that fall within it.

Resolution

Adjust the retention_period value and the maximum query range in the Loki limits_config as required. Setting the maximum query range to 0 removes the limit.

Learn More