NFS RWX Using the Puls8 NFS Driver

Explore This Page

Overview

DataCore Puls8 replicated block storage (Replicated PV Mayastor) delivers high-performance block volumes, but a block volume can be attached to only one node at a time (ReadWriteOnce, RWO). Many workloads require the opposite: multiple pods, distributed across multiple nodes, reading from and writing to the same data simultaneously (ReadWriteMany, RWX).

The DataCore Puls8 NFS driver addresses this requirement. When an RWX volume is requested, the driver provisions a Replicated PV Mayastor block volume, wraps it in a dedicated in-cluster NFS server, and exports it over the network. Every pod that mounts the claim communicates with that single export and sees the same files. Enterprise-grade Kerberos encryption and LDAP identity resolution can be layered on top when moving to production, without changes to the application.

This document describes how to provision, secure, and manage NFS ReadWriteMany (RWX) volumes in DataCore Puls8, including requirements, StorageClass configuration, the available security modes (sys, Kerberos, and LDAP), snapshots and restore, online expansion, high availability, and using shared volumes in your workloads.

How It Works

  • One NFS Server per Volume: Each NFS volume is served by a dedicated NFS server pod managed entirely by Kubernetes; no manual administration is required.
  • Lazy Start: When the PVC is created, the backend volume and the server's supporting resources are created, but the server itself remains at zero replicas. The server starts on the first pod mount and then stays running to serve all clients, including through unmounts, so a later pod re-attaches to the same server and data.
  • Thin Delegation Layer: The driver owns the NFS export and its security model (sys, Kerberos, or LDAP) and delegates all storage operations to the backend Replicated PV Mayastor CSI driver. Provisioning, snapshots, and resize are performed on the backend volume, which is why snapshot, restore, and expansion are surfaced through NFS as backend features.

When to Use NFS RWX

Use NFS RWX when:

  • Multiple pods must share a single filesystem (RWX). Typical cases include a web CMS with a shared uploads directory, machine-learning jobs whose workers share a dataset, a CI system with a shared artifact cache, or a legacy application that expects a POSIX shared folder.
  • You already run DataCore Puls8 / Replicated PV Mayastor and want shared storage without a separate NAS appliance.
  • You want a simple development configuration now and enterprise authentication (Kerberos/LDAP) later, without modifying the application.

Do not use NFS RWX when:

  • A single pod owns the volume; plain RWO Replicated PV Mayastor is faster and simpler.
  • Object storage or raw block semantics are required; this feature provides a POSIX file share.

Requirements

Ensure the following requirements are met before provisioning NFS RWX volumes.

Common (All Modes)

  • DataCore Puls8 / Replicated PV Mayastor: Must be installed and healthy, with a usable backend StorageClass (for example, mayastor-3). Each NFS volume is backed by a Replicated PV Mayastor RWO PVC.
Copy
Verify DiskPools and Backend StorageClass
kubectl get diskpools -n mayastor
kubectl get storageclass          # confirm the backend StorageClass exists
  • NFS Client Tools: Must be installed on every worker node that will mount NFS volumes.
Copy
Install NFS Client Tools
# Ubuntu / Debian
apt-get install -y nfs-common

# RHEL/CentOS/Fedora
dnf install -y nfs-utils
  • NFS Driver and License: The NFS driver must be installed (controller plus node DaemonSet in Running state), and a valid license must be active (see Licensing).

Additional for LDAP

  • LDAP Server: A reachable, standards-compliant LDAP server (OpenLDAP, 389-DS, or Active Directory) whose users carry POSIX attributes (uidNumber, gidNumber, objectClass: posixAccount) and whose groups carry objectClass: posixGroup and memberUid.
  • Bind Account: A read-only bind account, and a Secret holding its credentials in the driver namespace (see StorageClass Configuration Reference).

Additional for Kerberos (krb5 / krb5i / krb5p)

  • KDC: MIT Kerberos, FreeIPA, or Active Directory.
  • Kerberos Configuration: /etc/krb5.conf on every node, pointing at your realm/KDC. The NFS server pod reads this file from the host via a HostPath mount; distributing it to every node is the operator's responsibility (Ansible, a DaemonSet, or node configuration management).
Copy
Example /etc/krb5.conf (Every Node)
[libdefaults]
    default_realm = EXAMPLE.COM
    dns_lookup_realm = false
    dns_lookup_kdc = false
    rdns = false

[realms]
    EXAMPLE.COM = {
        kdc = kdc.example.com
        admin_server = kdc.example.com
    }

[domain_realm]
    .example.com = EXAMPLE.COM
    example.com = EXAMPLE.COM

Set rdns = false. With reverse DNS enabled, the client builds a different service principal than the one the keytab holds, and the mount fails with no obvious cause.

  • rpc-gssd: Running on every node with rpc_pipefs mounted. This is a kernel-level requirement of the NFS client and cannot be containerized.
Copy
Enable rpc-gssd on Every Node
apt-get install -y nfs-common krb5-user
mount -t rpc_pipefs sunrpc /var/lib/nfs/rpc_pipefs
echo "sunrpc /var/lib/nfs/rpc_pipefs rpc_pipefs defaults 0 0" >> /etc/fstab
systemctl enable --now rpc-gssd
  • Node Machine Keytab: /etc/krb5.keytab (host/<fqdn>@REALM) on every node that will mount a Kerberos volume. The mount authenticates the node via this keytab; a node without a keytab cannot mount a krb5* export.
  • NFS-Server Keytab: One per volume, provided by one of two methods (see StorageClass Configuration Reference).

Licensing

The NFS RWX feature is license-gated. A valid DataCore Puls8 license (trial or full) must be active on the cluster to provision new NFS RWX volumes. For instructions on installing and activating a license, refer to the License Activation documentation.

Exactly two operations are license-gated: creating a volume and creating a snapshot. Mount (publish), unmount (unpublish), expansion, and deletion are not gated.

  • No License: Provisioning is refused. A PVC or snapshot created against the NFS RWX driver remains Pending, with events carrying the message valid license required to perform operation. The rejection uses a retryable status, so the external provisioner continues to retry and the PVC binds automatically once a valid license is installed; no manual re-apply is required.
  • Trial License: A trial activates the feature for 30 days. If the trial elapses without a full license, new provisioning is refused again; existing volumes continue running.
  • Running Volumes Are Never Interrupted: Licensing is enforced only on the create and create-snapshot control path, never on the data path. Already-bound volumes continue serving I/O even if the license is deactivated, expires, or the cluster falls out of compliance. Only new volume and snapshot creation is blocked.
  • Agent-Unreachable Grace: License state is cached and refreshed approximately every 60 seconds. If the license agent is briefly unreachable, provisioning continues on the last-known-compliant state for a 180-second grace window; beyond that, the check fails closed (provisioning is refused) until the agent is reachable again.

Keep a valid license active before provisioning. Losing the license later does not take data offline; it only prevents the creation of new NFS RWX volumes.

Security and Authentication Model

Two independent settings determine identity and access. Review both before selecting a mode.

securityMode: How the Server Verifies Identity

Mode Meaning Wire Protection
sys Honor system; the client's raw UID/GID are trusted as sent None
krb5 Kerberos ticket required (verified identity) Authenticated
krb5i Kerberos plus integrity checksums (tamper-evident) Authenticated + integrity
krb5p Kerberos plus encryption (private) Authenticated + integrity + encrypted

sys requires no additional infrastructure and is appropriate for trusted or development clusters. krb5p is the strongest option and the recommended production channel.

export.squash: Whether the Server Preserves or Flattens Identity

Setting Effect on Root (UID 0) Effect on Normal Users
No_Root_Squash (default) Stays root on the server Unchanged
Root_Squash Mapped to anonymous (default 65534) Unchanged
All_Squash Mapped to anonymous Everyone mapped to anonymous

Anonymous_uid / Anonymous_gid (default 65534, the conventional nobody:nogroup) is the identity that squashed clients assume.

LDAP Identity Resolution

Adding an ldap block enables server-side group resolution. The server ignores the group list sent by the client and instead queries the directory (via SSSD to LDAP) for the user's actual group membership. This makes the server, not the client, the authority on group membership, so a pod cannot spoof its groups, and it removes the AUTH_SYS 16-supplementary-group limit. With LDAP, ls -la shows real usernames, and the same username maps to the same UID/GID on every node.

Kerberos and LDAP serve different purposes. Kerberos authenticates who you are; LDAP resolves your UID/GID and groups. With Kerberos but no LDAP, authenticated users have no name mapping and appear as numeric UIDs (or nobody/65534). Use LDAP when named, consistent identities are required.

Combining Modes

Mode Authentication Identity Mapping Typical Use
sys None None Development / trusted cluster
sys + LDAP None LDAP to consistent UID/GID/groups Development/test needing identity consistency
krb5p Kerberos, encrypted None (numeric) Encrypted transport, no directory
krb5p + LDAP Kerberos, encrypted LDAP to named users Recommended for production

All_Squash is rejected in combination with Kerberos or LDAP (validation error on export.squash), because flattening every user to a single anonymous identity discards the identity those systems compute. Use Root_Squash or No_Root_Squash instead.

StorageClass Configuration Reference

A StorageClass for this driver uses provisioner: com.datacore.puls8.nfs. Parameters are flat key-to-string values; nested settings (export, kerberos, ldap, server) are YAML-encoded strings under their respective keys.

Parameter Table

Parameter Required Default Description
backendStorageClass Yes - Backend (Replicated PV Mayastor) StorageClass from which the real RWO volume is provisioned. Must be a Replicated PV Mayastor StorageClass; see the note below the table.
securityMode No sys sys, krb5, krb5i, or krb5p
export No see export Block Squash, access, anonymous IDs (YAML string)
kerberos Required if securityMode is not sys - Realm and keytab source (YAML string)
ldap No - LDAP/SSSD identity configuration (YAML string)
server No see server Block Server pod resources and scheduling (YAML string)

The Replicated PV Mayastor CSI driver (io.openebs.csi-mayastor) is the only backend that is tested and accepted. No other driver is supported.

The driver applies this restriction to every backend object it is handed, so it governs more than provisioning. Each of the following fails, reporting that the NFS backend must be provisioned by one of the accepted drivers:

  • The provisioner of the backendStorageClass. The NFS PVC does not bind.
  • The driver behind a volume adopted through the nfs.puls8.datacore.com/backend-pvc annotation. The adopted PVC is left unclaimed.
  • The driver behind the backendSnapshotClass used for snapshot delegation. The VolumeSnapshot never reaches readyToUse.

A backendStorageClass or backendSnapshotClass that does not exist reports a different error, stating that the object was not found.

Unknown top-level parameters are ignored with a warning; unknown keys inside the export, kerberos, ldap, or server blocks are rejected.

imagePullPolicy is not a StorageClass parameter. The image pull policy for the NFS server pods is set once at install time through the Helm values (image.pullPolicy, or global.imagePullPolicy for all components); placing it in a StorageClass has no effect and is ignored with a warning.

The following standard Kubernetes StorageClass fields also apply and are significant here:

Field Recommended Reason
allowVolumeExpansion true Required to resize volumes
reclaimPolicy Delete Cleans up the backend volume on PVC deletion (dynamic mode)
volumeBindingMode Immediate Provisions as soon as the PVC is created
mountOptions see Mount Options NFS client options applied at mount time

PVC Annotations

Some behavior is controlled by annotations rather than StorageClass parameters. Placement matters: an annotation set on the wrong object is silently ignored.

Annotation Set On Description
nfs.puls8.datacore.com/backend-pvc NFS PVC Adopts a pre-existing RWO PVC instead of provisioning a new backend volume. Value format: <driver-namespace>:<pvc-name>.
nfs.puls8.datacore.com/export-sub-path Backend PVC Subdirectory of an adopted volume to export. Default: the volume root. No StorageClass or NFS-PVC equivalent.
nfs.puls8.datacore.com/recovery-sub-path Backend PVC Directory where lock-recovery data is kept on an adopted volume. Default: .ganesha. Must be a directory (not the root), and not the exported directory or a parent of it.
puls8.datacore.com/lifeline-grace-period NFS PVC Overrides the failover grace period for this volume (see High Availability).
puls8.datacore.com/nfs Backend PVC (driver-set) Stamped by the driver on an adopted backend PVC to claim it; its value is the NFS volume ID. You never set this - read it to confirm an adoption took effect.

The two sub-path annotations are read only from the backend PVC, never from the NFS PVC or the StorageClass. Setting them on the wrong object has no effect.

Mount Options

The driver injects sensible defaults at mount time, so a working mount requires no mountOptions:

Injected Default Meaning
vers=4.1 NFS protocol version
hard Retry indefinitely on server timeout (protects data integrity)
sec=<securityMode> GSS flavor derived from the StorageClass securityMode; you need not set sec yourself

The mount layer also supplies the server address and port automatically, and adds ro for read-only mounts. Any option added to the StorageClass mountOptions list is merged on top; if an option belongs to the same either/or group as a default (for example soft vs hard, ac vs noac), the specified option replaces the default. Common additions:

Option Reason
noac No attribute caching; every client sees fresh data/metadata immediately. Essential for RWX coherence across pods; trades some performance for consistency.
nconnect=8 Opens 8 TCP connections per mount for higher throughput.

export Block

Copy
export Block
export: |
  squash: No_Root_Squash    # No_Root_Squash | Root_Squash | All_Squash
  access: RW                # RW | RO
  anonymousUid: 65534       # identity used for squashed access
  anonymousGid: 65534
  rootMode: "0777"          # permissions on the top of the share (or "none" to leave adopted volumes untouched)

rootMode sets the permission bits on the top directory of the share. On a volume the driver creates, the default is 0777, because a brand-new filesystem root is otherwise writable only by root; on an adopted volume the driver leaves the existing root untouched unless rootMode is set. Use any octal mode, or none to leave it alone. If the identity that arrives (after squash and any LDAP mapping) cannot write to the root, clients mount successfully and then fail on the first write; the server logs the export root's mode and owner at start-up.

Choosing a Squash Setting:

Scenario Recommended
Development / trusted cluster, non-root containers No_Root_Squash
Containers may run as root, but root should not be granted on the share Root_Squash
Public/scratch share, identity irrelevant All_Squash (not allowed with Kerberos/LDAP)
Production (krb5p + LDAP) Root_Squash

kerberos Block

Required when securityMode is krb5, krb5i, or krb5p.

Copy
kerberos Block
kerberos: |
  realm: EXAMPLE.COM
  # provide exactly ONE of:
  keytabSecret: nfs-server-keytab     # manual: pre-create the keytab Secret
  # kadminSecret: kdc-admin-creds     # auto: driver mints the service keytab itself
  • Manual (keytabSecret): Create the NFS-server principal and keytab on the KDC and store it as a Secret (key nfs.keytab) in the driver namespace.
  • Auto (kadminSecret): Provide the driver with kadmin credentials (keys principal and password); it creates the principal and keytab itself via an init step before the server starts.

keytabSecret and kadminSecret are mutually exclusive.

For the manual route, the server principal name is derived by the driver and must be exactly:

Copy
NFS Server Principal Name (Manual Keytab)
nfs/nfs-svc-<volume-id>.<driver-namespace>.svc.<cluster-domain>@<REALM>

The <volume-id> is not the PVC name. It is pvc- followed by the PVC's UID (for example pvc-4f98e010-adeb-4e49-a888-6288083012e0), and this holds for adopted volumes too. Because that UID exists only once the PVC does, the manual route is a two-pass task: create the PVC against a StorageClass whose keytabSecret does not exist yet, let provisioning fail (the PVC stays Pending and keeps retrying), read the UID, build the keytab, and create the Secret under the name the StorageClass already references. The next retry finds it and the volume binds.

Copy
Derive the Principal and Create the Keytab Secret
DRIVER_NS="puls8"
CLUSTER_DOMAIN="cluster.local"
REALM="EXAMPLE.COM"

# Volume ID = "pvc-" + the PVC's UID (not its name)
VOLUME_ID="pvc-$(kubectl get pvc <your-pvc> -n <your-ns> -o jsonpath='{.metadata.uid}')"
PRINCIPAL="nfs/nfs-svc-${VOLUME_ID}.${DRIVER_NS}.svc.${CLUSTER_DOMAIN}@${REALM}"

# On the KDC:
kadmin.local -q "addprinc -randkey ${PRINCIPAL}"
kadmin.local -q "ktadd -k /tmp/nfs-${VOLUME_ID}.keytab ${PRINCIPAL}"

# Store under the exact name the StorageClass references:
kubectl create secret generic <the-keytabSecret-name> -n $DRIVER_NS \
  --from-file=nfs.keytab=/tmp/nfs-${VOLUME_ID}.keytab

The keytabSecret value is used verbatim - the driver does not expand a <volume-id> placeholder - so a keytab is tied to one volume's principal, and a StorageClass using keytabSecret serves exactly one volume. For more than one, use a StorageClass per volume, or use kadminSecret. The <cluster-domain> is usually cluster.local, and the realm comes from the kerberos block. The kadminSecret route needs none of this, as the driver builds the principal itself.

ldap Block

Copy
ldap Block
ldap: |
  servers: ldap://ldap.example.com   # space-separated for failover
  baseDn: dc=example,dc=com
  userSearchBase: ou=users,dc=example,dc=com
  groupSearchBase: ou=groups,dc=example,dc=com
  bindSecret: ldap-bind-creds        # Secret with keys bindDn + password
  tlsEnabled: true                   # use TLS (set false only on isolated networks)
  caCertSecret: ldap-ca-cert         # optional: CA cert Secret for strict TLS

The bind Secret must exist in the driver namespace and use the fixed keys bindDn and password.

server Block

All fields are optional; omit the block to use defaults.

Copy
server Block
server: |
  cpu: { request: 200m, limit: 2000m }
  memory: { request: 512Mi, limit: 2Gi }
  nodeSelector: { disktype: ssd }
  tolerations:
    - { key: dedicated, operator: Equal, value: nfs, effect: NoSchedule }
  labels: { team: storage }
  priorityClassName: high-priority
  hostNetwork: false           # true: server uses host network, Service becomes headless
  port: 2049                   # NFS port
  lifelineGracePeriod: 3m      # default HA rescue grace for this SC's volumes

Secrets Summary (All in the Driver Namespace)

Copy
Create Required Secrets in the Driver Namespace
DRIVER_NS=puls8

# LDAP bind (when ldap: is set)
kubectl create secret generic ldap-bind-creds -n $DRIVER_NS \
  --from-literal=bindDn="cn=nfs-service,ou=Services,dc=example,dc=com" \
  --from-literal=password="<password>"

# NFS server keytab, manual Kerberos mode
kubectl create secret generic nfs-server-keytab -n $DRIVER_NS \
  --from-file=nfs.keytab=/path/to/nfs.keytab

# kadmin credentials, auto Kerberos mode
kubectl create secret generic kdc-admin-creds -n $DRIVER_NS \
  --from-literal=principal="admin/admin@EXAMPLE.COM" \
  --from-literal=password="<kadmin-password>"

# LDAP CA certificate, optional strict TLS
kubectl create secret generic ldap-ca-cert -n $DRIVER_NS \
  --from-file=ca.crt=/path/to/ca.crt

Provisioning NFS RWX Volumes

Every example uses provisioner: com.datacore.puls8.nfs and an RWX PVC. Replace mayastor-3 and the driver namespace (puls8) with the values for your cluster.

Shared RWX Volume (sys)

The simplest configuration: multiple pods share one volume, raw UID/GID on the wire, no Kerberos or LDAP.

Copy
StorageClass for Shared RWX (sys)
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: nfs-sys
provisioner: com.datacore.puls8.nfs
reclaimPolicy: Delete
volumeBindingMode: Immediate
allowVolumeExpansion: true
parameters:
  backendStorageClass: mayastor-3
  securityMode: sys
  export: |
    squash: No_Root_Squash
    access: RW
mountOptions: ["noac", "hard", "nconnect=8"]
Copy
PersistentVolumeClaim for Shared RWX
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: shared-data
  namespace: my-app
spec:
  accessModes: [ReadWriteMany]
  storageClassName: nfs-sys
  resources:
    requests:
      storage: 10Gi

Point a multi-replica Deployment at shared-data (see Using the Volume in Your Pods) and every replica, on any node, shares the files.

Server-Side Identity with LDAP (sys + LDAP)

Add an ldap block. The server resolves each user's actual groups from the directory, so a pod that presents no group information can still obtain group-based access, and file ownership displays consistent named users.

Copy
StorageClass for sys with LDAP
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: nfs-sys-ldap
provisioner: com.datacore.puls8.nfs
reclaimPolicy: Delete
volumeBindingMode: Immediate
allowVolumeExpansion: true
parameters:
  backendStorageClass: mayastor-3
  securityMode: sys
  export: |
    squash: No_Root_Squash
    access: RW
  ldap: |
    servers: ldap://ldap.example.com
    baseDn: dc=example,dc=com
    userSearchBase: ou=users,dc=example,dc=com
    groupSearchBase: ou=groups,dc=example,dc=com
    bindSecret: ldap-bind-creds
    tlsEnabled: true
mountOptions: ["noac", "hard"]

Create the ldap-bind-creds Secret in the driver namespace before creating the PVC.

Encrypted NFS with Kerberos (krb5p)

Every RPC is authenticated and encrypted. This requires the Kerberos node prerequisites (see Requirements). Without an ldap block, authenticated users have no name mapping and may appear as nobody (65534), which is acceptable when only access enforcement, not per-user identity, is required.

Copy
StorageClass for Encrypted NFS (krb5p)
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: nfs-krb5p
provisioner: com.datacore.puls8.nfs
reclaimPolicy: Delete
volumeBindingMode: Immediate
parameters:
  backendStorageClass: mayastor-3
  securityMode: krb5p
  export: |
    squash: No_Root_Squash
    access: RW
  kerberos: |
    realm: EXAMPLE.COM
    kadminSecret: kdc-admin-creds     # or keytabSecret: nfs-server-keytab
mountOptions: ["sec=krb5p", "noac", "hard"]

Full Enterprise (krb5p + LDAP)

The recommended production configuration: encrypted transport with named, consistent identities resolved server-side.

Copy
StorageClass for Full Enterprise (krb5p + LDAP)
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: nfs-krb5p-ldap
provisioner: com.datacore.puls8.nfs
reclaimPolicy: Delete
volumeBindingMode: Immediate
allowVolumeExpansion: true
parameters:
  backendStorageClass: mayastor-3
  securityMode: krb5p
  export: |
    squash: Root_Squash
    access: RW
  kerberos: |
    realm: EXAMPLE.COM
    kadminSecret: kdc-admin-creds
  ldap: |
    servers: ldap://ldap.example.com
    baseDn: dc=example,dc=com
    userSearchBase: ou=users,dc=example,dc=com
    groupSearchBase: ou=groups,dc=example,dc=com
    bindSecret: ldap-bind-creds
    tlsEnabled: true
mountOptions: ["sec=krb5p", "noac", "hard"]

Adopt an Existing Backend Volume

An existing RWO Replicated PV Mayastor PVC can be exposed as RWX without re-provisioning or copying data. Set an annotation on the NFS PVC naming the backend PVC as <driver-namespace>:<pvc-name>. The driver wraps that PVC in an export instead of creating a new backend volume.

Copy
PersistentVolumeClaim to Adopt an Existing Backend Volume
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: shared-data
  namespace: my-app
  annotations:
    nfs.puls8.datacore.com/backend-pvc: "puls8:legacy-data"   # <ns>:<name>
spec:
  accessModes: [ReadWriteMany]
  storageClassName: nfs-sys
  resources:
    requests:
      storage: 5Gi        # must be <= the backend PVC's capacity

Requirements for the backend PVC: it must reside in the driver namespace, be Bound, and be ReadWriteOnce; its capacity must be greater than or equal to the size requested by the NFS PVC; and it must not already be claimed by another NFS volume (each backend PVC backs one NFS volume).

After the NFS PVC binds, confirm the adoption took effect by reading the claim annotation the driver stamps on the backend PVC. Its value is the NFS volume ID.

Copy
Confirm the Adoption Claim on the Backend PVC
kubectl get pvc <backend-pvc> -n puls8 -o jsonpath='{.metadata.annotations.puls8\.datacore\.com/nfs}'

An adopted backend PVC is never deleted by the driver. Deleting the NFS PVC leaves the backend PVC and its data intact and re-adoptable; you own its lifecycle. In default dynamic mode, the driver creates and deletes the backend PVC for you.

When adopting a volume, the recovery-sub-path annotation on the backend PVC is validated at provisioning. Three mistakes are rejected outright, so the NFS PVC stays Pending with the message on it:

  • Naming the directory you are exporting: is the exported directory; every client would be denied.
  • Naming the volume root (. or empty): is the backend volume root; name a directory instead.
  • An absolute path (leading /): must be relative to the backend volume root, not absolute.

One case is not caught at provisioning: if recovery-sub-path names an existing file, the volume binds and then the server pod crashloops with exists but is not a directory. Look at the server pod rather than the PVC. Point the annotation at a directory (or remove the file) and recreate the volume, since the exported and recovery paths are fixed at creation.

Snapshot and Restore

Snapshots are delegated to the backend CSI driver. Create two VolumeSnapshotClass objects: an NFS class that points at a backend class.

Copy
VolumeSnapshotClass for NFS and Backend
# Backend snapshot class; driver MUST match your backend CSI provisioner
apiVersion: snapshot.storage.k8s.io/v1
kind: VolumeSnapshotClass
metadata:
  name: nfs-backend-snapclass
driver: io.openebs.csi-mayastor
deletionPolicy: Delete
---
# NFS snapshot class; delegates to the backend class above
apiVersion: snapshot.storage.k8s.io/v1
kind: VolumeSnapshotClass
metadata:
  name: nfs-snapclass
driver: com.datacore.puls8.nfs
deletionPolicy: Delete
parameters:
  backendSnapshotClass: nfs-backend-snapclass
Copy
Create a Snapshot
apiVersion: snapshot.storage.k8s.io/v1
kind: VolumeSnapshot
metadata:
  name: shared-data-snap
  namespace: my-app
spec:
  volumeSnapshotClassName: nfs-snapclass
  source:
    persistentVolumeClaimName: shared-data

Request exactly the size the snapshot reports. The backend requires a clone to match its snapshot size, so a larger request is rejected and the PVC stays Pending with a Cloned snapshot volume must match the snapshot size error on the backend claim. Read the size from the snapshot, then grow the volume after it binds if you need more room.

Copy
Read the Snapshot's Restore Size
kubectl -n my-app get volumesnapshot shared-data-snap -o jsonpath='{.status.restoreSize}'
Copy
Restore into a New RWX PVC
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: shared-data-restored
  namespace: my-app
spec:
  accessModes: [ReadWriteMany]
  storageClassName: nfs-sys
  resources:
    requests:
      storage: 10Gi         # must equal the snapshot's restoreSize exactly
  dataSource:
    apiGroup: snapshot.storage.k8s.io
    kind: VolumeSnapshot
    name: shared-data-snap

The restore PVC remains Pending until the snapshot is readyToUse, then binds automatically. PVC-to-PVC cloning is not supported; creating a PVC with a dataSource of kind PersistentVolumeClaim fails with volume cloning is not supported. Restore from a snapshot instead.

Online Volume Expansion

A volume can be grown in place, even while pods are mounted and performing I/O. The expansion is delegated to the backend resizer. This requires allowVolumeExpansion: true on both the NFS StorageClass and the backend StorageClass; if the backend forbids it, expansion is rejected.

Copy
Expand a Volume In Place
kubectl -n my-app patch pvc shared-data \
  -p '{"spec":{"resources":{"requests":{"storage":"20Gi"}}}}' 

Mounted pods observe the larger filesystem without remount or restart. If no pod has ever mounted the volume, it expands offline: the backend filesystem is resized when the server next starts, not at request time. Because that is also the first moment a client can reach the export, no client ever observes the old size.

High Availability (Lifeline)

NFS server pods are rescue targets of the DataCore Puls8 Lifeline controller (they automatically carry the puls8.datacore.com/lifeline-target: true annotation). If the node running a volume's NFS server becomes NotReady, Lifeline waits the grace period, then force-deletes the stranded server pod and clears its stale volume attachments, so the StatefulSet reschedules it onto a healthy node. The backend volume re-attaches there and hard-mounted clients resume I/O without data loss.

The grace period (a duration such as 30s or 2m) can be tuned:

  • Per Volume: Annotation on the PVC, puls8.datacore.com/lifeline-grace-period: 2m.
  • Per StorageClass: server.lifelineGracePeriod (see the StorageClass Configuration Reference), used as the default when the PVC has no annotation.

The PVC annotation takes precedence over the StorageClass default. An unparseable value fails provisioning with an event naming the invalid value; validate the format before applying.

What Clients See: With the default hard mount, clients pause and then resume during a failover - no I/O errors and no lost writes. A pod doing continuous I/O appears to freeze for the recovery window and then continues. A server pod restart or move on a healthy cluster resumes in roughly 30 seconds; a full node failure is dominated by the grace period above (about 6 to 7 minutes with the shipped 5-minute default), after which recovery takes seconds.

Kerberos and LDAP are on the restart path. If the NFS server restarts while the KDC is unreachable, it cannot fetch its keytab and the volume stays unavailable until the KDC returns; existing mounts and their locks are unaffected, only a restart blocks. If it restarts while LDAP is unreachable, it comes up but cannot resolve users until LDAP returns, so ownership and group permissions behave as if the directory were empty. Plan KDC and LDAP availability alongside the cluster's.

Lock Reclaim: File locks (used by most databases and many queues) are reclaimed automatically after the server restarts or moves; applications do not re-take them. Briefly after a restart the server accepts reclaims from clients that held locks and rejects new lock requests until that window closes - typically under a minute and never more than 90 seconds. Clients only reading and writing see nothing. This requires every client machine to have a unique hostname; Kubernetes satisfies this by default (each node's hostname is its node name), but duplicate hostnames break lock reclaim.

Where Recovery Data Lives: The server keeps its lock-recovery state on the volume itself, in a reserved directory (.ganesha at the volume root by default) that is hidden from directory listings and refuses ordinary access. You do not manage, back up, or size for it.

Treat the hidden recovery directory as tidiness, not a security boundary; it prevents accidental access, not a client that deliberately goes looking. For hostile-tenant isolation, use separate volumes. On an adopted volume the directory can be relocated by setting nfs.puls8.datacore.com/recovery-sub-path on the backend PVC before the NFS volume is created; it must be a directory, not the root, and not the exported directory or a parent of it.

Moving Your Own Pods Off a Dead Node

The failover above brings the NFS server back. There is a second problem on your side of the mount: an application pod that was running on the dead node stays scheduled there, because Kubernetes will not move a pod off an unreachable node while its volume is still attached. The replica can sit Terminating, or simply stop making progress, long after the volume is healthy elsewhere.

The Lifeline controller can rescue such pods - when a node stops responding it force-deletes the pods that opted in and releases their volume attachments, so the replacement schedules on a healthy node. The NFS server pod is enrolled automatically; your application pods are opt-in, because force-deleting a pod is not safe for every workload. Opt a pod in by annotating the pod template (not the Deployment, and not the PVC):

Copy
Opt an Application Pod Into Lifeline Rescue
apiVersion: apps/v1
kind: Deployment
metadata:
  name: my-app
  namespace: my-app
spec:
  template:
    metadata:
      annotations:
        puls8.datacore.com/lifeline-target: "true"
        # optional; how long to give the node before acting on this pod
        puls8.datacore.com/lifeline-grace-period: 60s
    spec:
      containers:
        - name: app
          image: my-app:latest
          volumeMounts:
            - name: shared
              mountPath: /data
      volumes:
        - name: shared
          persistentVolumeClaim:
            claimName: shared-data
  • It force-deletes: The pod is removed without a graceful shutdown, so if your application depends on its preStop hook or a clean flush, leave this off and let the workload's own controller handle the failure.
  • Give the replacement somewhere to go: A rescued pod helps only if the scheduler can place it. requiredDuringScheduling anti-affinity across nodes, or a replica count equal to the node count, leaves the replacement Pending; use preferred anti-affinity instead.

Using the Volume in Your Pods

Multi-Replica Deployment (Any Mode)

Copy
Multi-Replica Deployment Sharing One RWX Volume
apiVersion: apps/v1
kind: Deployment
metadata:
  name: my-app
  namespace: my-app
spec:
  replicas: 3                    # all replicas share the one volume
  selector:
    matchLabels: { app: my-app }
  template:
    metadata:
      labels: { app: my-app }
    spec:
      # Spread replicas across nodes to exercise real multi-node RWX
      affinity:
        podAntiAffinity:
          preferredDuringSchedulingIgnoredDuringExecution:
            - weight: 100
              podAffinityTerm:
                topologyKey: kubernetes.io/hostname
                labelSelector: { matchLabels: { app: my-app } }
      containers:
        - name: app
          image: my-app:latest
          volumeMounts:
            - { name: shared, mountPath: /data }
      volumes:
        - name: shared
          persistentVolumeClaim:
            claimName: shared-data

Matching Identity (sys Mode)

In sys mode, file ownership is by numeric UID/GID. Set the pod's identity so that it owns or can access the files. On any shared share - anything with ldap:, or any use of group permissions - set both runAsUser and runAsGroup:

Copy
Pod Security Context (sys Mode)
spec:
  securityContext:
    runAsUser: 10001
    runAsGroup: 10001

runAsGroup is not stylistic. Omit it and the pod's primary group defaults to 0 (or 65534 under Root_Squash), and that group is stamped on every file the pod creates. Because every pod without runAsGroup defaults to the same group, the group permission bits stop distinguishing anyone and a 0660 file effectively behaves like 0666: it looks correctly restricted, but every pod on the share can open it. With LDAP, group membership for reading is authoritative on the server (it rebuilds the group list from the UID via SSSD), but that check is only as good as the group the file was created with, so runAsGroup still matters.

Do not use fsGroup on these shares. It is not another way to set the group - it asks the kubelet to walk and re-own the entire volume at mount time. On a share with squash: Root_Squash the kubelet's root is squashed, the walk is refused, and the pod never starts (applyFSGroup failed ... permission denied); fsGroupChangePolicy: OnRootMismatch does not avoid it. Use runAsGroup instead - it is a plain process credential and never touches the volume.

supplementalGroups behaves differently depending on whether the class sets ldap:. With LDAP it is ignored for access decisions - the server rebuilds the group list from the UID, so declared extra groups are discarded; put the membership in the directory instead. Without LDAP the server trusts the credential as sent, so supplementalGroups is how a pod joins additional groups, up to the 16 that AUTH_SYS carries. Either way it never sets the group on a new file - that always comes from runAsGroup.

Kerberos-Authenticated Pods

In any krb5* mode the mount works only while the accessing user has a valid Kerberos ticket, and one detail drives everything else: the ticket is used by the node, not by the pod. Each node runs rpc.gssd, which looks for the ticket cache on the node's filesystem at /tmp/krb5cc_<uid> for the UID doing the I/O. A ticket obtained only inside the container, in a path only the container can see, is never found and the mount is refused. There are two workable approaches.

Option A - One identity per node (no per-pod setup): Give each node a machine keytab (with host, root, and nfs principals) and every pod on that node shares that identity. The pod spec needs nothing extra - no kinit, Secret, or environment variables. Use this when every workload on the share can act as the same identity.

Option B - Per-user credentials (real per-user identity): Use this when different applications must appear as different users, which is the point of running with LDAP. An init container runs kinit and writes the ticket cache to the node's /tmp through a hostPath; the application container runs as the matching UID. The user must exist in both the KDC (for authentication) and LDAP (for the UID/GID), or authentication succeeds and files come out owned by nobody.

Copy
Per-User Kerberos Credential via an Init Container
spec:
  securityContext:
    runAsUser: 10001          # the user's UID in LDAP
    runAsGroup: 10001
  initContainers:
    - name: kinit
      image: my-krb5-client:latest    # any image with kinit
      command: ["sh", "-c"]
      args:
        - |
          set -e
          export KRB5_CONFIG=/krb5/krb5.conf
          export KRB5CCNAME=FILE:/host-tmp/krb5cc_$(id -u)
          kinit -kt /etc/krb5-user/alice.keytab alice@EXAMPLE.COM
      volumeMounts:
        - { name: user-keytab, mountPath: /etc/krb5-user, readOnly: true }
        - { name: krb5-conf,   mountPath: /krb5,          readOnly: true }
        - { name: host-tmp,    mountPath: /host-tmp }
  containers:
    - name: app
      image: my-app:latest
      volumeMounts:
        - { name: shared,   mountPath: /data }
        - { name: host-tmp, mountPath: /host-tmp }
  volumes:
    - name: shared
      persistentVolumeClaim: { claimName: shared-data-krb5 }
    - name: user-keytab
      secret: { secretName: alice-keytab }
    - name: krb5-conf
      configMap: { name: krb5-conf }
    - name: host-tmp
      hostPath: { path: /tmp, type: Directory }
  • KRB5CCNAME must point at the host-visible path, and the file name must end in the UID doing the I/O (krb5cc_10001 for UID 10001).
  • runAsUser must match the UID LDAP gives the user; a ticket for one user under a different UID is refused.
  • The pod also needs your realm's krb5.conf, published as a ConfigMap in the application namespace.
  • Keep the ticket alive: Kerberos tickets expire. For long-running workloads, re-run kinit on a schedule (a sidecar looping kinit then sleep at well under the ticket lifetime); it refreshes the same cache in place, with no restart.

The ticket cache is per node, keyed by UID: it lives at /tmp/krb5cc_<uid> on the node, so every pod on that node running as that UID uses whichever ticket is there, including one a different pod placed there. Two workloads that must not act as each other should not share a UID, and a pod that stops needing its identity keeps it until the ticket expires or is removed. Give each identity its own UID where that matters.

Without LDAP (krb5, krb5i, or krb5p alone) there is no directory to resolve principals against, so authentication still works but files appear owned by the anonymous user. Use Kerberos with LDAP when file ownership must mean something.

Diagnosing Kerberos Mount Failures

A mount that fails in a krb5, krb5i, or krb5p mode reports which part of the Kerberos chain did not complete, so that the node prerequisites and the credential can be told apart. Each message carries a reason, which is the part to read first.

Reason Meaning What to check
KerberosPrereqMissing The kernel could not set up Kerberos for the mount at all. Confirm that rpc-gssd is running on the node, that rpc_pipefs is mounted, and that /etc/krb5.keytab is present. This reason can also indicate an invalid mount option.
KerberosKeytabStale Kerberos was set up, but rpc.gssd obtained no usable credential. The host keytab is most likely stale. Refresh the node credential and retry the mount.
Copy
Refresh the Node Credential from the Host Keytab
kinit -k host/$(hostname)@<REALM>

Both reasons point at the node, not at the StorageClass or the export. A volume whose StorageClass is correct still fails to mount on a node that was never provisioned for Kerberos, so confirm the node prerequisites before revisiting the security configuration.

Read-Only Mount

Copy
Read-Only Volume Mount
volumes:
  - name: shared
    persistentVolumeClaim:
      claimName: shared-data
      readOnly: true

Alternatively, make the whole export read-only with export: | access: RO on the StorageClass.

Best Practices

  • Start Simple, Harden Later: Develop on sys, then switch the StorageClass to krb5p + LDAP for production. Application manifests do not change.
  • Always set noac and hard for shared RWX workloads. noac provides cross-pod read-after-write coherence; hard protects against data loss during transient server unavailability.
  • Use nconnect=8 for throughput-sensitive workloads.
  • Enable allowVolumeExpansion: true on both the NFS and backend StorageClasses up front; volumes cannot be resized without it.
  • Prefer auto Kerberos (kadminSecret) unless policy forbids granting the driver kadmin rights; it avoids minting and rotating a keytab Secret per volume.
  • Use Root_Squash in multi-tenant and production shares so that a rogue root container cannot bypass ownership.
  • Keep TLS enabled for LDAP (tlsEnabled: true) outside isolated test networks.
  • Spread client replicas across nodes (pod anti-affinity) to genuinely exercise and test multi-node RWX.
  • Keep a valid license active before provisioning. Data remains online without it, but new volumes cannot be created.
  • Size the backend correctly for adoption: the NFS request must be less than or equal to the backend PVC capacity, and the backend PVC must be RWO and Bound.

Considerations and Limitations

  • RWX via NFS Semantics Only: This is a POSIX file share (NFSv4), not block or object storage. Expect NFS consistency and locking behavior.
  • One NFS Server Pod per Volume: That server is a per-volume component; high availability is provided by Lifeline rescue, not by running multiple server replicas.
  • PVC-to-PVC Cloning Is Not Supported: Use snapshot and restore.
  • All_Squash Restriction: Cannot be combined with Kerberos or LDAP.
  • Kerberos Requires Node-Level Setup: rpc-gssd, rpc_pipefs, /etc/krb5.conf, and a machine keytab cannot be containerized; a node missing these cannot mount krb5* volumes.
  • Adopted Backend PVCs Are Not Garbage-Collected: You own their lifecycle.
  • NFSv4 Only: NFSv3 clients are not served.
  • Resilience Follows the Backend: A volume is only as resilient as its backend StorageClass. A single-replica backend class means a node failure takes the data with it; use a multi-replica backend for redundancy.
  • Exported Directory Is Fixed at Creation: The exported directory cannot be changed after the volume exists; editing it has no effect. To change it, recreate the NFS PVC (adoption preserves the backend) or move the data on the backend volume.
  • Recovery Directory Is Fixed at Creation: The reserved lock-recovery directory is likewise fixed at creation. Both it and the exported directory are read only from the backend PVC of an adopted volume, never from the NFS PVC or the StorageClass.
  • Grow-Only: Volumes can expand but never shrink; a smaller requested size is rejected.
  • Reserved .ganesha Name: Do not create a file or directory named .ganesha (or your chosen recovery-directory name) at the top of a volume - it is in use by the server.
  • Released Adopted Volumes Retain the Reserved Directory: Releasing an adopted volume leaves its reserved recovery directory behind; remove it yourself if you need the volume pristine.
  • Kerberos Principals Are Not Deleted: The driver creates a service principal per volume and never removes it, so deleting a volume leaves a stale nfs/nfs-svc-<volume-id> entry in the KDC. Nothing breaks, but prune them yourself if the KDC is audited or volumes are created in bulk.

Benefits of NFS ReadWriteMany (RWX)

  • Shared Multi-Node Access: Multiple pods across multiple nodes read and write the same filesystem concurrently through a single RWX volume, which is not possible with plain RWO block storage.
  • No Separate NAS Required: Shared storage is delivered on your existing DataCore Puls8 / Replicated PV Mayastor cluster, using the same pools and StorageClasses.
  • Enterprise Security When You Need It: Layer Kerberos encryption and LDAP identity on top without changing application manifests; develop on sys and promote to krb5p + LDAP for production.
  • Backend-Grade Data Services: Snapshots, restore, and online expansion are delegated to Replicated PV Mayastor, so they work through NFS with no extra tooling.
  • Built-In High Availability: NFS servers are automatically protected by Lifeline rescue, so a node failure reschedules the server and clients resume I/O without data loss.
  • Zero Server Administration: Kubernetes manages the per-volume NFS server end to end; there is nothing to deploy or maintain manually.

Learn More