Repository navigation
Conversation
hh24k
marked this pull request as ready for review
September 2, 2026 07:22
hh24k
force-pushed
the
dev/139
branch
2 times, most recently
from
September 2, 2026 07:30
fb72f09 to
7b09794
Compare
YanniHu1996
force-pushed
the
dev/139
branch
4 times, most recently
from
September 7, 2026 07:46
87c4228 to
c3e02a3
Compare
hh24k
force-pushed
the
dev/139
branch
4 times, most recently
from
September 8, 2026 05:49
02b1d1e to
119aadb
Compare
YanniHu1996
approved these changes
Sep 16, 2026
gabriele-wolfox
force-pushed
the
dev/139
branch
4 times, most recently
from
October 6, 2026 13:53
d8202de to
b7b8f2d
Compare
The CNPG-I plugin's RESTORE_WAL path measured its end-to-end duration but only logged it. Expose it as an OTel histogram (klio.plugin.wal.restore_duration, ns, per-file buckets) tagged with outcome, cache_hit, tier and cluster_name. This is the latency PostgreSQL actually experiences when it asks for a WAL segment. In a replica cluster whose designated primary replicates from an external Klio source, it is the speed of the replication path itself and so drives replica lag; during recovery it drives how fast a cluster catches up. Neither is visible today, and the server-side get_duration cannot stand in for it: that times only one WAL.Get call, while prefetch cache hits are served from the local spool and never reach the server at all. cache_hit and tier are threaded up from the prefetcher through restoreWAL via a small restoreOutcome; Restore records on every exit path via defer so failures are measured too. Signed-off-by: Hai He <hai.he@enterprisedb.com>
Add two panels to the Client / Plugin section for the new klio.plugin.wal.restore_duration metric: p95 end-to-end restore latency split by cache_hit and restore rate by outcome. A prefetch hit is a local rename while a miss waits on a download, so the two are orders of magnitude apart and are shown split rather than pooled; a falling hit ratio means prefetch is not keeping ahead of replay. Regenerated klio-dashboard.json via the builder. Assisted-by: Claude Signed-off-by: Hai He <hai.he@enterprisedb.com>
Keep WAL restore metrics aligned with the real restore path by retaining the last attempted tier on miss-all-tiers failures and expose a direct prefetch hit-ratio panel in the Grafana dashboard. Signed-off-by: Hai He <hai.he@enterprisedb.com>
Signed-off-by: huyantian <yantian.hu@enterprisedb.com>
The restoreResult call and the errors.Is immediately after it classified the same error twice. Deriving the status from the classification keeps restoreResult the single definition of "not found": if another not-found sentinel is added there and this call site is not updated alongside it, the metric would report not_found while CloudNativePG receives Internal, the divergence this attribute exists to prevent. Behaviour is unchanged: restoreResult returns OutcomeNotFound exactly when errors.Is(err, errWALNotFound) holds. Signed-off-by: Hai He <hai.he@enterprisedb.com>
…tions The WAL restore panels predate the multi-cluster rework of the dashboard and still followed its old layout. Rebuild them on the same patterns the other panels use. The duration is now a p50/p90/p99 pair built with the shared quantile helpers, a rolling-window (rate) panel next to a since-restart (total) one, like the WAL block send and WAL file get percentiles. The restore counts get the same rate/total pair as the WAL files written panels and are split by tier and outcome, so the tier that served a restore is visible on the dashboard as cloudnative-pg#139 requires. Every panel, the hit ratio included, now groups by cluster_name like the rest of the Client / Plugin section, and uses the named unit constants and panel sizes. The section description and the user documentation are updated to match, and klio-dashboard.json is regenerated from the builder. Assisted-by: Claude Signed-off-by: Gabriele Quaresima <gabriele.quaresima@enterprisedb.com>
…tric Add the WAL restore duration histogram to the OpenTelemetry metrics reference, and list the new cache_hit attribute, the not_found outcome and the unknown tier value in the attribute table. Assisted-by: Claude Signed-off-by: Gabriele Quaresima <gabriele.quaresima@enterprisedb.com>
The "unknown" value only exists as a metric attribute, but it was declared as a tier, so getClient had to handle it as a target to connect to. restoreWAL now returns an empty tier when none was tried, and recordWalRestore maps it to "unknown" like it already does for the cluster name. Assisted-by: Claude Signed-off-by: Gabriele Quaresima <gabriele.quaresima@enterprisedb.com>
Add a check to the OTEL scenario that covers klio.plugin.wal.restore_duration. Run the restore_command in the primary twice: once for a WAL Klio holds, retrying until it is available, and once for a WAL that was never archived. Then check that the collector receives a success data point with explicit buckets and a not_found one. Assisted-by: Claude Signed-off-by: Gabriele Quaresima <gabriele.quaresima@enterprisedb.com>
The cache_hit attribute comes from getCompleteWAL, but no test called it: only the isReadyPrefetch check was covered. Add a table test for entries already tracked by the prefetcher: a ready prefetch is a hit, while an in-flight prefetch, a direct download and a failed download are misses. Also update the TestRestoreResult comment, which no longer holds now that the e2e scenario covers a successful restore. Assisted-by: Claude Signed-off-by: Gabriele Quaresima <gabriele.quaresima@enterprisedb.com>
Refresh the Client / Plugin row screenshot so it shows the new WAL restore panels, with data for both tier1 and tier2: - WAL restore duration percentiles (rate) - WAL restore duration percentiles (total) - WAL restores (rate) - WAL restores (total) - WAL restore prefetch hit ratio Signed-off-by: Gabriele Quaresima <gabriele.quaresima@enterprisedb.com>
gabriele-wolfox
force-pushed
the
dev/139
branch
from
October 6, 2026 14:34
b7b8f2d to
b295f95
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
RESTORE_WALtiming only went to the log, so restore latency couldn't be graphedor alerted on. That latency sets the replica lag when a replica cluster
replicates from an external Klio source. See #139 for the background, and for why
klio.server.wal.get_durationdoesn't cover it.This adds
klio.plugin.wal.restore_duration, a histogram in nanoseconds. It usesthe same per-WAL-file buckets as the server's get and upload histograms, so the
three read alike. It is tagged with
outcome,cache_hit,tierandcluster_name.cache_hitandtierare only known deep in the restore path, so a smallrestoreOutcomecarries them back up throughrestoreWAL.Restorerecordsfrom a
defer, so failures are measured too, including the early ones thatreturn before any tier is tried.
A hit means the prefetch finished before PostgreSQL asked for the file. That
check is now named
walEntry.isReadyPrefetchand has a table test. If theprefetch is still running the caller waits for it, so it counts as a miss — which
is what makes the hit rate show whether prefetch is keeping up with replay.
Three dashboard panels in the Client / Plugin section: p95 latency split by
cache_hit, rate byoutcome, and the prefetch hit ratio. The hit ratio is theone to watch over time — it falling means prefetch is no longer keeping pace with
replay, which is fixable through the prefetch settings.
klio-dashboard.jsonis regenerated from the builder.Closes #139