Python: Fix/7859 discard pending state - #8819
Patel Namraa (Namraa310806) wants to merge 23 commits into
Conversation
|
I have fixed the failing workflows caused by my changes |
|
/review |
| await iteration_task | ||
| except Exception: | ||
| # Discard pending state writes from the failed superstep | ||
| self._state.discard() |
There was a problem hiding this comment.
Patel Namraa (@Namraa310806) This still misses cancellation: asyncio.CancelledError inherits from BaseException, so cancellation and generator-close paths skip this except Exception block and can leave pending state for the next run to commit. Please discard pending state on cancellation as well, with a regression that stages a write before cancelling the run task.
| # Non-success path: still publish headers diagnostically, then raise. | ||
| self._assign_response_headers(state, result) | ||
| # Commit the state before raising so headers are persisted even on error | ||
| ctx.state.commit() |
There was a problem hiding this comment.
Patel Namraa (@Namraa310806) This mid-executor ctx.state.commit() defeats the failed-superstep isolation fix by committing the entire pending workflow-state buffer before raising. The runner can only discard what remains pending, so state from the failed HTTP action—and unrelated writes from the same superstep—can become durable. Please avoid committing from this error path; if non-2xx headers must remain observable, preserve them without committing failed-superstep state.
There was a problem hiding this comment.
MAF Automated Review — Iteration 1
Result: No findings
Scope: full PR (23 commit(s)): 956f43f046c7, 59a705a704d7, b5e6ca071232, da0154c74884, 2848ad2594ec, debc9ba0040b, 8bfce51295b0, f6f7535399fb, 018f570668ef, 8dd250b2bc8e, 729818d192f8, aa573c740807, d5a32232b729, eda1814275fb, 547db3cb6fae, 7182915ee7dd, d074130e7f4f, 89e43bb0a329, f809d8ea6171, c70e117ed9dc, c52fdd2f7a0c, 5e856024baac, e2e23ef16bb3
Model: gpt-5.6-sol-fast
Overview
The PR adds failed-superstep pending-state cleanup and routes concurrent delivery through a helper that cancels and awaits siblings before propagating errors. Regression coverage confirms ordinary executor failures do not leak staged State writes into a later successful run, and existing tests cover active-task cancellation and checkpoint restoration races. No additional publishable finding remains after deduplicating against the supplied unresolved feedback and separating pre-existing lifecycle behavior from PR-introduced defects.
Reviewed the supplied pull-request change set across correctness, security/reliability, architecture, and failure behavior.
No publishable findings remained after source verification for this scope.
Motivation & Context
Fixes a state isolation bug where pending
Statewrites from a failed or cancelled workflow superstep can leak into a later successfulWorkflow.run()on the sameWorkflowinstance.State.set()stages writes in a pending buffer, which is committed only when a superstep completes successfully. However, when a superstep fails or is cancelled, the pending writes were not discarded.Because workflow state persists across
run()calls, a later successful superstep could callState.commit()and unintentionally commit stale writes from the previous failed run.This results in silent state corruption: state written by a failed run can become part of the committed state of a later successful run.
Description & Review Guide
What are the major changes?
What is the impact of these changes?
State.commit()behavior.What do you want reviewers to focus on?
Related Issue
Fixes #7859
Contribution Checklist
breaking changelabel (or add "[BREAKING]" to the title prefix) — a workflow keeps the label and title prefix in sync automatically.Extending #7880 and fixing CI failing workflows