Repository navigation
feat(purchases): flip RecoverStrandedApprovals to idempotent re-drive (now that #636 makes it safe) #639
Description
Activity
- addedenhancementNew feature or requestNew feature or requesttriagedItem has been triagedItem has been triagedpriority/p3Polish / idea / may never shipPolish / idea / may never shipseverity/lowMinor harmMinor harmurgency/eventuallyNo deadlineNo deadlineimpact/all-usersAffects every userAffects every usereffort/sHoursHourstype/featNew capabilityNew capability
on May 21, 2026 Blocked by #641. The purchase-workflow trace found that idempotency (from #636) is honored ONLY by EC2 RI + Savings Plans. Every other executor — AWS RDS/ElastiCache/MemoryDB/OpenSearch/Redshift, all Azure reservations, and GCP — silently ignores
opts.IdempotencyToken. FlippingRecoverStrandedApprovalsto auto-re-drive (this issue) before #641 closes would double-purchase on every non-EC2/SP commitment. Do not implement #639 until #641 lands.- added 4 commits that reference this issue
on May 21, 2026 Blocker: Azure two-step reservation purchases are not idempotent — re-drive will double-buy
Implementation paused. Investigation surfaced an unresolved gap in the #641 prerequisite chain that makes auto-re-drive unsafe today for all 7 Azure reservation services.
What is OK
Spot-checked one executor per "honoring
opts.IdempotencyToken" claim — these are all good:- AWS SP (
providers/aws/services/savingsplans/client.go:205-206): setsCreateSavingsPlanInput.ClientToken = opts.IdempotencyToken— server-side dedup. - AWS EC2 RI (
providers/aws/services/ec2/client.go:126-137):findRIByIdempotencyTokenlookup viatag:cudly-idempotency-token, short-circuits on hit, tags fresh purchases with the token. - AWS RDS / MemoryDB / OpenSearch / Redshift / ElastiCache:
IdempotentReservationID("<prefix>", opts.IdempotencyToken)(e.g.providers/aws/services/rds/client.go:203,providers/aws/services/memorydb/client.go:135) plusidempotencyGuard+recoverAlreadyExists— token-derived reservation ID, server rejects duplicate. - GCP Compute (
providers/gcp/services/computeengine/client.go:522-548): commitment name +RequestIdboth derived fromopts.IdempotencyToken(regression testclient_test.go:530-549pins same-token-same-RequestId).
DeriveIdempotencyToken(execID, recIndex)itself is a puresha256("<execID>:<recIndex>")(pkg/common/tokens.go:41-44) — fully reproducible across runs.What is broken
All 7 Azure reservation executors lost client-side idempotency in PR #680 (commit
65ecbf81c, merged after #653/#641):cache,compute,cosmosdb,database,search,managedredis,synapse— every one now delegates the purchase toproviders/azure/services/internal/reservations/DoPurchaseTwoStep(purchase.go:106-133). That helper:- POSTs
calculatePrice— Azure mints a freshreservationOrderIdserver-side every call (doCalculatePrice,purchase.go:137-164). - POSTs
reservationOrders/{azure-id}/purchase.
opts.IdempotencyTokenis not threaded into either step. The reservation order ID that #653 derived from the token viaIdempotencyGUIDis no longer used by these services. The PR #680 commit body explicitly acknowledges this and promises:"Idempotency (Option B): Azure mints the order ID at calculatePrice time so client-supplied IDs are no longer feasible. Caller-level deduplication uses the purchase-automation tag already attached to each reservation request."
The caller-level guard does not exist.
internal/purchase/execution.go:710 executeSinglePurchasecallsserviceClient.PurchaseCommitment(recCtx, recommendation, opts)directly — no pre-flight check for an already-existing reservation tagged with the(executionID, recIndex)identity. The only tag attached iscudly-purchase-source: web|cli(PurchaseTagKey, e.g.providers/azure/services/compute/client.go:395), which is not unique per execution and so cannot deduplicate a re-drive.A stranded Azure approval re-driven by
RecoverStrandedApprovalswill create a second reservation for every rec whose original attempt landed at Azure but didn't finalize the row. The compute client's own header comment (compute/client.go:409-412) describes the missing caller-level guard as if it already exists, which it doesn't.Required to unblock #639
Either:
- Restore client-side idempotency in the two-step flow — most surgical option is to derive a stable reservation order ID from
opts.IdempotencyTokenand PUT directly toreservationOrders/{derived-id}oncecalculatePricevalidates the body, skipping Azure's server-minted ID. Requires Azure-side validation that PUT on a pre-known ID still works on the SKU families that motivated fix(providers/azure): direct PUT to reservationOrders/{id} rejected by Azure ("Session timed out - Call CalculatePrice again") — switch to two-step CalculatePrice→Purchase flow (P0) #677/fix(azure): switch 7 reservation clients to two-step calculatePrice->purchase flow (closes #677) #680. - Implement the promised caller-level guard — before
executeSinglePurchasecalls into an Azure service, list existing reservations tagged withcudly-idempotency-token: <token>(i.e. extend Azure executors to attach a per-rec identity tag, mirroringPurchaseTagKey) and short-circuit if found, mirroringfindRIByIdempotencyTokenon the EC2 RI side.
Either path needs its own issue + PR before #639 can land. AWS + GCP are ready; Azure is the lone hold-out.
Why this matters operationally
RecoverStrandedApprovalstriggers on a 15-minapproved-status threshold. Any Azure rec whose original execution hung past 15 min after the commitment landed (Lambda timeout, network partition between Azure and the row save, the partial-failure paths that #642 records) is a guaranteed double-buy under the proposed flip — which is exactly the scenario #639 is supposed to make safer.Branch
feat/639-recover-strands-auto-re-driveis parked (no code changes, only investigation) until the Azure path is closed.- AWS SP (
Verification: cannot-verify-autopilot - Partially shipped. PR #728 (merged 2026-05-26) delivers AWS-only idempotent re-drive in
RecoverStrandedApprovals(allRecsAWSpredicate +claimAndRedrive). PR #729 (merged 2026-05-26) restores idempotency token threading throughDoPurchaseTwoStepfor Azure. However,allRecsAWScomment onfeat/multicloud-web-frontendstill explicitly says 'Azure re-drive idempotency is blocked' and Azure/GCP strands still fall through to safe-fail. Full idempotent re-drive for all providers is not yet shipped.- added a commit that references this issue
on Jun 5, 2026
Final step of the stranded-approval story. Depends on both PR #635 (recovery sweep that currently safe-FAILS stranded
approvedrows) and PR #638 (#636 idempotent commitment creation) merging.Once both are in: change
RecoverStrandedApprovalsto re-drive a strandedapprovedexecution instead of failing it, since a re-drive is now double-buy-safe:ClientToken(deterministic per feat(purchases): idempotent commitment creation (ClientToken/dedupe) (closes #636) #638) dedupes server-side.tag:cudly-idempotency-token+ active/payment-pending lookup) short-circuits a repeat, fails-loud on lookup error. Residual succeed-then-tag-fail window is irreducible and acceptable for an automated re-drive (the operator still sees the row).Add a regression test: a stranded row is re-driven, and a re-drive that hits an already-existing commitment does NOT create a second one.
Note: #636's body and #632's analysis incorrectly assumed SP has no native idempotency; #638 corrected that. This issue reflects the corrected, safer reality.