Skip to content

feat: add klio.plugin.wal.restore_duration end-to-end metric - #206

Open
hh24k wants to merge 11 commits into
cloudnative-pg:mainfrom
hh24k:dev/139
Open

hh24k wants to merge 11 commits into
cloudnative-pg:mainfrom
hh24k:dev/139

Conversation

@hh24k

@hh24k hh24k commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

RESTORE_WAL timing only went to the log, so restore latency couldn't be graphed
or alerted on. That latency sets the replica lag when a replica cluster
replicates from an external Klio source. See #139 for the background, and for why
klio.server.wal.get_duration doesn't cover it.

This adds klio.plugin.wal.restore_duration, a histogram in nanoseconds. It uses
the same per-WAL-file buckets as the server's get and upload histograms, so the
three read alike. It is tagged with outcome, cache_hit, tier and
cluster_name.

cache_hit and tier are only known deep in the restore path, so a small
restoreOutcome carries them back up through restoreWAL. Restore records
from a defer, so failures are measured too, including the early ones that
return before any tier is tried.

A hit means the prefetch finished before PostgreSQL asked for the file. That
check is now named walEntry.isReadyPrefetch and has a table test. If the
prefetch is still running the caller waits for it, so it counts as a miss — which
is what makes the hit rate show whether prefetch is keeping up with replay.

Three dashboard panels in the Client / Plugin section: p95 latency split by
cache_hit, rate by outcome, and the prefetch hit ratio. The hit ratio is the
one to watch over time — it falling means prefetch is no longer keeping pace with
replay, which is fixable through the prefetch settings.
klio-dashboard.json is regenerated from the builder.

Closes #139

@hh24k
hh24k marked this pull request as ready for review September 2, 2026 07:22
@hh24k
hh24k force-pushed the dev/139 branch 2 times, most recently from fb72f09 to 7b09794 Compare September 2, 2026 07:30
@YanniHu1996
YanniHu1996 force-pushed the dev/139 branch 4 times, most recently from 87c4228 to c3e02a3 Compare September 7, 2026 07:46
@hh24k
hh24k requested review from a team and jlong49 as code owners September 8, 2026 03:10
@hh24k
hh24k force-pushed the dev/139 branch 4 times, most recently from 02b1d1e to 119aadb Compare September 8, 2026 05:49
@gabriele-wolfox
gabriele-wolfox force-pushed the dev/139 branch 4 times, most recently from d8202de to b7b8f2d Compare October 6, 2026 13:53
hh24k and others added 10 commits October 6, 2026 16:34
The CNPG-I plugin's RESTORE_WAL path measured its end-to-end duration but only
logged it. Expose it as an OTel histogram (klio.plugin.wal.restore_duration, ns,
per-file buckets) tagged with outcome, cache_hit, tier and cluster_name.

This is the latency PostgreSQL actually experiences when it asks for a WAL
segment. In a replica cluster whose designated primary replicates from an
external Klio source, it is the speed of the replication path itself and so
drives replica lag; during recovery it drives how fast a cluster catches up.
Neither is visible today, and the server-side get_duration cannot stand in for
it: that times only one WAL.Get call, while prefetch cache hits are served from
the local spool and never reach the server at all.

cache_hit and tier are threaded up from the prefetcher through restoreWAL via a
small restoreOutcome; Restore records on every exit path via defer so failures
are measured too.

Signed-off-by: Hai He <hai.he@enterprisedb.com>
Add two panels to the Client / Plugin section for the new
klio.plugin.wal.restore_duration metric: p95 end-to-end restore latency split by
cache_hit and restore rate by outcome.

A prefetch hit is a local rename while a miss waits on a download, so the two
are orders of magnitude apart and are shown split rather than pooled; a falling
hit ratio means prefetch is not keeping ahead of replay. Regenerated
klio-dashboard.json via the builder.

Assisted-by: Claude

Signed-off-by: Hai He <hai.he@enterprisedb.com>
Keep WAL restore metrics aligned with the real restore path by retaining the last attempted tier on miss-all-tiers failures and expose a direct prefetch hit-ratio panel in the Grafana dashboard.

Signed-off-by: Hai He <hai.he@enterprisedb.com>
Signed-off-by: huyantian <yantian.hu@enterprisedb.com>
The restoreResult call and the errors.Is immediately after it classified the
same error twice. Deriving the status from the classification keeps
restoreResult the single definition of "not found": if another not-found
sentinel is added there and this call site is not updated alongside it, the
metric would report not_found while CloudNativePG receives Internal, the
divergence this attribute exists to prevent.

Behaviour is unchanged: restoreResult returns OutcomeNotFound exactly when
errors.Is(err, errWALNotFound) holds.

Signed-off-by: Hai He <hai.he@enterprisedb.com>
…tions

The WAL restore panels predate the multi-cluster rework of the dashboard and
still followed its old layout. Rebuild them on the same patterns the other
panels use.

The duration is now a p50/p90/p99 pair built with the shared quantile
helpers, a rolling-window (rate) panel next to a since-restart (total) one,
like the WAL block send and WAL file get percentiles. The restore counts get
the same rate/total pair as the WAL files written panels and are split by
tier and outcome, so the tier that served a restore is visible on the
dashboard as cloudnative-pg#139 requires. Every panel, the hit ratio included, now groups
by cluster_name like the rest of the Client / Plugin section, and uses the
named unit constants and panel sizes.

The section description and the user documentation are updated to match,
and klio-dashboard.json is regenerated from the builder.

Assisted-by: Claude

Signed-off-by: Gabriele Quaresima <gabriele.quaresima@enterprisedb.com>
…tric

Add the WAL restore duration histogram to the OpenTelemetry metrics
reference, and list the new cache_hit attribute, the not_found outcome and
the unknown tier value in the attribute table.

Assisted-by: Claude

Signed-off-by: Gabriele Quaresima <gabriele.quaresima@enterprisedb.com>
The "unknown" value only exists as a metric attribute, but it was declared
as a tier, so getClient had to handle it as a target to connect to.
restoreWAL now returns an empty tier when none was tried, and
recordWalRestore maps it to "unknown" like it already does for the cluster
name.

Assisted-by: Claude

Signed-off-by: Gabriele Quaresima <gabriele.quaresima@enterprisedb.com>
Add a check to the OTEL scenario that covers
klio.plugin.wal.restore_duration.

Run the restore_command in the primary twice: once for a WAL Klio holds,
retrying until it is available, and once for a WAL that was never
archived. Then check that the collector receives a success data point
with explicit buckets and a not_found one.

Assisted-by: Claude

Signed-off-by: Gabriele Quaresima <gabriele.quaresima@enterprisedb.com>
The cache_hit attribute comes from getCompleteWAL, but no test called it:
only the isReadyPrefetch check was covered. Add a table test for entries
already tracked by the prefetcher: a ready prefetch is a hit, while an
in-flight prefetch, a direct download and a failed download are misses.

Also update the TestRestoreResult comment, which no longer holds now that
the e2e scenario covers a successful restore.

Assisted-by: Claude

Signed-off-by: Gabriele Quaresima <gabriele.quaresima@enterprisedb.com>
Refresh the Client / Plugin row screenshot so it shows the new WAL restore
panels, with data for both tier1 and tier2:

- WAL restore duration percentiles (rate)
- WAL restore duration percentiles (total)
- WAL restores (rate)
- WAL restores (total)
- WAL restore prefetch hit ratio

Signed-off-by: Gabriele Quaresima <gabriele.quaresima@enterprisedb.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

WAL restore: add an end-to-end restore time metric

3 participants