Eventing Aggregator

Explore this Page

Overview

Replicated PV Mayastor emits an event whenever a significant change occurs in the cluster, such as the creation of a volume, a change of state on a pool, or the start of a rebuild. These events describe what the storage layer did and when, which makes them the primary record for reconstructing the sequence of a failure or for confirming that an operation completed.

The Eventing Aggregator is the component that collects those events and makes them available for querying through the DataCore Puls8 kubectl plugin. It is deployed by default and requires no configuration for standard deployments.

This document describes how events are collected, where they are stored and for how long, how to configure the aggregator, and how to query events from a running cluster or from a support bundle collected earlier.

For the catalogue of events that Replicated PV Mayastor emits, refer to the Eventing documentation. This document covers the collection and retrieval of events rather than their content.

How It Works

Each Replicated PV Mayastor component publishes its events to the NATS message bus. The Eventing Aggregator subscribes to that bus and writes the events it receives to a destination that the kubectl plugin can read.

Two characteristics of this pipeline are visible during normal operation:

  • Events appear in batches. The aggregator writes events in groups rather than individually, so a new event can take up to 10 seconds to appear in query results. An event that has just occurred may therefore be absent from a query issued immediately afterwards.
  • A restart does not lose events. Events that the aggregator had received but not yet written are redelivered when it restarts.

The aggregator runs as a single deployment. It is not a data-path component, so its availability affects the visibility of events only, and never volume access or I/O.

Requirements

To collect and query events, ensure the following requirements are fulfilled:

  • Eventing is enabled. This is the default. No events are produced while eventing is disabled.
  • The Eventing Aggregator is deployed. This is the default. Events are still produced while the aggregator is disabled, but they are not collected and cannot be queried.
  • NATS is deployed with JetStream enabled. This is the default. The aggregator does not start until NATS is reachable.
  • Loki is deployed. This is optional. Without Loki, only recent events are available.

Event Sources and Retention

Events are held in one of two places, and the choice determines both how long they remain available and which source the plugin reads from. The plugin selects a source automatically, so no action is required in either case.

Source Used When Retention Governed By Survives a Pod Restart
Loki Loki is deployed The retention_period value in the Loki limits_config Yes
Aggregator volume Loki is not deployed openebs.mayastor.eventing.aggregator.dirSizeLimit No

A third source, the NATS message bus, can be read directly. This is useful when the aggregator itself is unavailable, because the events remain on the bus.

  • Deploy Loki if you require event history that survives a pod restart. The aggregator volume is ephemeral and is cleared whenever the aggregator pod restarts.
  • Loki also applies a maximum query range, which defaults to 30 days. A request for a longer window returns only the events that fall within that range, even when the retention period is longer.

Configuration

The aggregator is configured through the dependent Replicated PV Mayastor chart.

Copy
Eventing Aggregator Helm Values
openebs:
  mayastor:
    eventing:
      enabled: true
      aggregator:
        enabled: true
        logLevel: "info"
        dirSizeLimit: "100Mi"
Value Default Description
openebs.mayastor.eventing.enabled true Enables event generation. No events are produced when this is set to false.
openebs.mayastor.eventing.aggregator.enabled true Deploys the Eventing Aggregator.
openebs.mayastor.eventing.aggregator.logLevel info Log verbosity of the aggregator container.
openebs.mayastor.eventing.aggregator.dirSizeLimit 100Mi Size limit for the events held on the aggregator volume. Applies only when Loki is not deployed.

The aggregator also accepts the standard resources, tolerations, nodeSelector and priorityClassName values. It is deployed with CPU and memory limits of 100m and 32Mi, and requests of 50m and 16Mi.

Disabling Event Collection

To stop collecting events while leaving event generation enabled, disable the aggregator during installation or upgrade. Disable eventing itself only when no events are to be produced at all.

Copy
Flag to Disable the Eventing Aggregator
--set openebs.mayastor.eventing.aggregator.enabled=false

Retaining More Events Without Loki

If Loki is not deployed and a longer window of event history is required, increase the size limit of the aggregator volume.

Copy
Flag to Increase the Aggregator Volume Size Limit
--set openebs.mayastor.eventing.aggregator.dirSizeLimit=500Mi

This increases the number of events held between restarts. It does not make them survive a restart. Deploy Loki for durable event history.

Verifying the Deployment

  1. Confirm that the aggregator pod is running.

    Copy
    List the Eventing Aggregator Pod
    kubectl get pods -n puls8 -l app=eventing-aggregator
    Copy
    Sample Output
    NAME                                          READY   STATUS    RESTARTS   AGE
    puls8-eventing-aggregator-7c9f4b6d8c-x2mkq    1/1     Running   0          11m
  2. Confirm that events are being collected by querying them. A populated result confirms the full path from the message bus through to the plugin.

    Copy
    Confirm Event Collection
    kubectl puls8 mayastor get events -n puls8 --limit 5

A pod that remains in the Init state is waiting for NATS. Refer to the Troubleshooting documentation.

Querying Events

Use the kubectl plugin to retrieve events. The plugin resolves its own source, so the command is the same whether the events are held in Loki or on the aggregator volume.

Copy
Query Cluster Events
kubectl puls8 mayastor get events -n puls8
Copy
Sample Output
TIMESTAMP             CATEGORY  ACTION         TARGET                                NODE      COMPONENT
2026-08-12T09:14:02Z  volume    create         18e30e83-b106-4e0d-9fb6-2b04e761e18a  worker-1  CoreAgent
2026-08-12T09:14:03Z  replica   create         c0f9a1d2-77b3-4a51-9e0c-1b2a3c4d5e6f  worker-2  IoEngine
2026-08-12T09:18:47Z  nexus     rebuild_begin  18e30e83-b106-4e0d-9fb6-2b04e761e18a  worker-1  IoEngine

By default, the command returns events from the last 24 hours, up to a maximum of 1000 records. Narrow the result with the filters below rather than raising the limit, because a filtered result is quicker to interpret.

Filters

Option Description
--category Event category, such as volume or pool
--action Event action, such as create or state_change
--component Component that produced the event
--node Node name
--target Target resource ID
--pool Pool name, matched as a substring of the target
--volume Volume UUID
--replica Replica UUID
--rebuild-status Rebuild outcome
--state Previous or next state of a state-change event, matched as a substring
--filter <path=value> Any field of the event payload, addressed by a dot path such as metadata.source.component=IoEngine. Leading and trailing * wildcards are supported.

Filters are combined using AND logic, so each additional filter narrows the result. The --category, --action, --component and --rebuild-status options accept comma-separated or repeated values.

An event is excluded by --filter when the specified path is absent from its payload. A filter that returns nothing may therefore indicate a path that does not exist on the events concerned, rather than an absence of matching events.

Selecting a Source

Option Description
--loki-endpoint <URL> Reads from the specified Loki instance instead of discovering one
--nats-endpoint [URL] Reads directly from the NATS message bus. Omit the value to discover the service automatically.
--from-file <path> Reads from a local file instead of a cluster

These three options are mutually exclusive, and specifying more than one is rejected. Supplying --loki-endpoint also disables the automatic fallback to the aggregator volume, so an unreachable endpoint returns an error rather than results from another source.

Additional Options

Option Default Description
-o, --output <format> table Output format: table, json or yaml. The JSON and YAML formats include the full event payload.
--since <duration> 24h Events from the last duration, such as 1h or 7d
--limit <N> 1000 Maximum number of events to return. A value of 0 means unlimited.
--tenant-id <ID> puls8 The Loki X-Scope-OrgID header. Applies to the Loki source only, and is required when Loki runs with authentication enabled.

Examples

Copy
Retrieve All Events for a Volume
kubectl puls8 mayastor get events -n puls8 --volume 18e30e83-b106-4e0d-9fb6-2b04e761e18a
Copy
Retrieve Nexus Events from the Last Six Hours
kubectl puls8 mayastor get events -n puls8 --category nexus --since 6h
Copy
Retrieve Pool Events from a Node in JSON Format
kubectl puls8 mayastor get events -n puls8 --category pool --node worker-1 -o json

Collecting Events in a Support Bundle

Events are included in the support bundle, so the record of what the storage layer did is available to an engineer who has no access to the cluster.

Copy
Generate a Support Bundle with all System Information
kubectl puls8 dump system -n puls8 -d <output_directory_path>

The archive contains events.ndjson, holding the collected events, and events-source.txt, recording the source they were collected from. Refer to the Supportability documentation for the structure of the archive.

Analysing Events Offline

An extracted events file supports the same filters as a live cluster and requires no cluster connection, so a bundle can be examined on any machine that has the plugin installed.

Copy
Query Events from an Extracted File
kubectl puls8 mayastor get events --from-file ./events.ndjson --component io-engine

Best Practices

  • Deploy Loki where event history matters. Without it, events do not survive a restart of the aggregator pod, which is exactly when a record of the preceding events is most valuable.
  • Filter rather than raise the limit. A query narrowed by category, node or resource is quicker to interpret than a larger unfiltered result.
  • Widen the time window before concluding that no events exist. The default window is 24 hours, and an empty result is often a window or a filter that is too narrow.
  • Collect a support bundle before remedial action. The bundle preserves the events leading up to an incident, which a restart of the aggregator can otherwise discard when Loki is not deployed.

Considerations and Limitations

  • A new event can take up to 10 seconds to appear in query results.
  • When Loki is not deployed, events are held on an ephemeral volume and are cleared when the aggregator pod restarts.
  • Loki applies a maximum query range, which defaults to 30 days, independently of its retention period.
  • The aggregator does not start until NATS is reachable, and collects no events while it is unavailable. Events that have already been published remain on the message bus.
  • Disabling eventing stops event production for all consumers, not only for the aggregator.

Benefits of the Eventing Aggregator

  • Cluster-Wide Visibility: Collects events from every Replicated PV Mayastor component into a single queryable record.
  • Targeted Investigation: Narrows results by category, action, component, node or resource, so that a single volume or pool can be followed through its lifecycle.
  • Offline Analysis: Applies the same filters to a support bundle, so that events can be reviewed without access to the cluster.
  • Operational Simplicity: Deployed and enabled by default, with no configuration required for standard deployments.

Learn More