Skip to content

fix: user opcode handlers under PHP 8.6's tail-call VM (macOS arm64) - #281

Draft
lisachenko wants to merge 10 commits into
masterfrom
claude/zdebug-php-8.6-testing-2pwowy
Draft

lisachenko wants to merge 10 commits into
masterfrom
claude/zdebug-php-8.6-testing-2pwowy

Conversation

@lisachenko

@lisachenko lisachenko commented Aug 29, 2026 •

Copy link
Copy Markdown
Owner

What

Fixes #280 on z-engine's side: on PHP 8.6 builds using the new tail-call VM (ZEND_VM_KIND_TAILCALL — clang without global-register support, notably Apple Silicon macOS), user opcode handlers mis-resume execution inside the engine and corrupt the process. The diagnostics on this branch reduced it to a php-src bug (pure ext-ffi repro, no z-engine code — see the upstream-ready report on #280): the generated ZEND_USER_OPCODE_SPEC_TAILCALL_HANDLER returns the single-step dispatch result up the musttail chain and execute_ex() resumes it against its stale entry frame.

Until php-src resolves it, z-engine fails fast instead of corrupting the debuggee:

  • Core::vmKind() — reports zend_vm_kind() through a dedicated one-symbol FFI binding (usable before init(), no generated-header changes), with VM_KIND_* constants mirroring Zend/zend_vm_opcodes.h
  • OpCodeHook::install() throws OpCodeHookException::tailCallVmUnsupported() on VM_KIND_TAILCALL, naming the issue
  • OpCodeHookVmKindGuardTest runs in every CI leg (deliberately not in the internal group — the existing opcode-hook tests are internal-only, which is why the macOS legs never caught this): on tail-call builds it asserts the refusal, elsewhere the unchanged install/uninstall lifecycle
  • README/AGENTS document the platform caveat

Diagnostics kept as the upstream repro harness

tools/diagnostics/issue-280/ (probe ladder + pure-ffi-repro.php) and the now dispatch-only Diagnose issue 280 workflow stay in-tree: re-dispatching it against a newer PHP build answers "is the php-src bug fixed yet"; both go away with the guard once upstream resolves it.

Verification

  • macOS arm64 (tail-call VM): full CI green with the guard test taking the refusal branch; the pure-FFI repro crashes there (vm_kind=5, SIGSEGV) proving the underlying bug is engine-level — run 33257997984
  • macOS x64 / linux (hybrid VM): all probes and the repro pass; full suite, PHPStan level max, cs — green
  • Probe evidence for the mechanism (stage-marked ladder, lldb, crash reports): run 33256998594

Downstream: lisachenko/zdebug#24 keeps its macOS arm64 + 8.6 leg experimental until upstream is fixed; with this guard, consumers get a clear OpCodeHookException instead of corrupted debuggees.


🤖 Generated with Claude Code

https://claude.ai/code/session_01U4cKi87mVQ896uSrw592fG

claude added 4 commits August 29, 2026 14:06
PHP 8.6 on macOS arm64 (the clang/aarch64 tail-call VM build) corrupts VM
state when a user opcode handler re-enters PHP (#280, found via
lisachenko/zdebug#24). The probes install a raw zend_set_user_opcode_handler
callback - no OpCodeHook, no ExecutionData - and climb from a handler that
touches nothing to a per-fire dump of EG(current_execute_data) chaining,
vm_stack_top/end and the interrupted frame, with an ADD-without-EXT_STMT
baseline. The temporary diagnose-280 workflow runs the ladder on
macos-latest (arm64, failing) and macos-15-intel (x64, control), plus an
lldb backtrace of the failing shape. All modes pass on linux-x64 8.6
(hybrid VM). Probes and workflow are removed once the fix lands.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U4cKi87mVQ896uSrw592fG
…cktrace

The first arm64 run segfaulted in every mode with zero output, so the
crash point is unknown (lldb's -k commands produced no backtrace and the
nearest-symbol frame zend_class_init_statics is unreliable for the
static TAILCALL handlers). Stderr stage markers now bracket the crash
(autoload / init / options / installed / payload-first-statement /
payload-done / uninstalled), install-only separates handler installation
from the first dispatch, the probe loop tolerates crashes so all modes
report, and lldb uses -o so bt/registers/disassembly actually print.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U4cKi87mVQ896uSrw592fG
The generated ZEND_USER_OPCODE_SPEC_TAILCALL_HANDLER dispatches through
the single-step (fastcall) handler table and returns its next-opline
result straight up the musttail chain; execute_ex() only refreshes its
execute_data local in the ZEND_VM_ENTER_BIT branch, so a plain returned
opline resumes execution against the frame execute_ex was ENTERED with.
Any user opcode handler firing in a deeper frame (an include, any
function call) then executes the following oplines against the wrong
frame: wrong run-time cache (calls dispatch to the wrong function),
wrong CVs, and eventually SIGSEGV. Hybrid/call VM builds are immune
(opline and frame live in global registers / per-op returns), which is
why only clang-without-global-regs builds - notably Apple Silicon
macOS - are affected. Diagnosed on the arm64 runner via the layered
probes in tools/diagnostics/issue-280 (a payload fwrite executed as the
outer frame's cached unlink is the smoking gun); this is a php-src
bug to be reported upstream.

Until php-src resolves it, z-engine fails fast instead of corrupting
the debuggee:

- Core::vmKind() reports zend_vm_kind() through a dedicated one-symbol
  FFI binding (usable before init(), no generated-header changes), with
  the VM_KIND_* constants mirroring Zend/zend_vm_opcodes.h
- OpCodeHook::install() throws OpCodeHookException::tailCallVmUnsupported()
  on VM_KIND_TAILCALL with a message naming the issue
- OpCodeHookVmKindGuardTest runs in every CI leg (deliberately not in
  the internal group): on tail-call builds it asserts the refusal, on
  every other build the unchanged install/uninstall lifecycle
- README/AGENTS document the platform caveat

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U4cKi87mVQ896uSrw592fG
…VM bug

Reduces the arm64 corruption to ext-ffi against the engine's exported
API alone (an ADD user handler returning DISPATCH, fired inside a
function frame) - the shape the upstream php-src report needs. Green on
hybrid-VM builds (linux 8.5/8.6); the diagnose workflow verifies the
crash on the arm64 tail-call build.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U4cKi87mVQ896uSrw592fG
The arm64 verification is done (probe run 33256998594, pure-FFI repro
run 33257997984): the corruption is php-src's tail-call VM bug and the
OpCodeHook guard covers consumers. Keep the probes and the workflow as
the upstream repro harness, but stop running them on every pull request
- a manual dispatch against a new PHP build answers 'is it fixed yet'.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U4cKi87mVQ896uSrw592fG
@lisachenko

Copy link
Copy Markdown
Owner Author

@claude Look for the relevant bug report for PHP at https://github.com/php/php-src/ (issues + PRs) and link it here if you will find anything relevant

Copy link
Copy Markdown
Owner Author

Searched php/php-src issues and PRs (queries: zend_set_user_opcode_handler, ZEND_USER_OPCODE_DISPATCH, USER_OPCODE, "user opcode handler" + tailcall/tail-call, plus web search) — no existing report covers this bug. Nobody has filed the user-opcode-handler mis-resume under the tail-call VM yet, so the report drafted on #280 is still worth submitting upstream.

Closest related upstream items (same bug class — the tail-call VM returning a next-opline/tagged value that the caller mishandles — but all about vm_interrupt, not user opcode handlers):

  • php/php-src#17849 — "Tail call VM" by arnaud-lb: the PR that introduced the tail-call VM (handlers return the next opline up the musttail chain), useful context and the right people to CC (arnaud-lb, morrisonlevi).
  • php/php-src#21922 — "Fix tailcall helper dispatch after VM interrupt" (merged, PHP-8.5): ZEND_VM_DISPATCH_TO_HELPER() dereferenced an opline tagged with ZEND_VM_ENTER_BIT returned by a helper during an interrupt — precisely the same tagged/next-opline mishandled by the resuming code pattern our diagnosis found in ZEND_USER_OPCODE_SPEC_TAILCALL_HANDLER's ZEND_USER_OPCODE_DISPATCH path.
  • php/php-src#22265 — "Fix another vm_interrupt bug for tailcall VM" (merged, PHP-8.5): a follow-up crash in the same area (missing SAVE_OPLINE before ZEND_VM_LOOP_INTERRUPT).

Those two merged fixes show the pattern is a known weak spot, which should make the upstream report easy to triage — it can reference them, the pure-FFI repro (tools/diagnostics/issue-280/pure-ffi-repro.php, no z-engine code involved), and the analysis in the #280 comment. I can't file it from this session (no access to php/php-src), so that comment is ready to copy-paste as the issue body.


Generated by Claude Code

@lisachenko

Copy link
Copy Markdown
Owner Author

let's recheck if php has fixed this or not?

Copy link
Copy Markdown
Owner Author

Re-checked against the current php-src build (run 36771134832, dispatched on this branch): not fixed yet.

Tested: PHP 8.6.0-dev @ dc1145e (built Sep 29, 2026, NTS) on macos-latest / arm64 — vm_kind=5 (ZEND_VM_KIND_TAILCALL), and the failure signatures are byte-for-byte the ones from the original diagnosis:

  • install-only passes; every mode that actually dispatches an opcode still corrupts the frame: the probe's next statement executes as the outer frame's cached call — Uncaught TypeError: unlink(): Argument #2 ($context) must be of type resource or null, string given at probe.php:17 (exit 255) in noop/log-const/globals/use-ref/diag
  • add-baseline: SIGSEGV right after STAGE payload-first-statement (exit 139)
  • pure-FFI repro (no z-engine): vm_kind=5 then SIGSEGV, EXC_BAD_ACCESS / KERN_INVALID_ADDRESS at 0xffffffffffffffe8 (exit 139)

I also checked the php-src side directly: no commit touching ZEND_USER_OPCODE/tailcall dispatch has landed on the PHP-8.6 branch since the diagnosis (the recent zend_vm_execute.h changes are unrelated — GH-23980 include/eval assert and 8.4/8.5 merges), and there's still no upstream issue covering this — only the merged vm_interrupt fixes in the same area (php/php-src#21922, php/php-src#22265). So the upstream-ready report on #280 remains unfiled and worth submitting.

(The macos-15-intel control leg was still compiling PHP from source when I read the results; it's only the hybrid-VM control and doesn't affect the verdict. Re-dispatching this workflow stays the way to re-test future builds.)


Generated by Claude Code

setup-php currently has no PHP 8.6 build for macos-15-intel: the Set up
PHP step gives up after ~25 minutes ('Could not setup PHP 8.6'),
reproducibly (run 36771134832, both attempts), which fails the whole
dispatch and paints the PR red. The x64 leg is only the hybrid-VM
control - the arm64 leg alone answers 'is the php-src bug fixed yet'
(it still crashes on 8.6.0-dev@dc1145e, see PR #281) - so mark it
continue-on-error until setup-php ships 8.6 for Intel macOS again.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U4cKi87mVQ896uSrw592fG

Copy link
Copy Markdown
Owner Author

The Probes (x64) failure on the re-dispatched diagnostics run is not a diagnostic signal — it's setup-php infrastructure: on macos-15-intel the Set up PHP step gives up after ~25 minutes with "Could not setup PHP 8.6" (no 8.6 build currently available for Intel macOS there), reproducibly on both attempts of run 36771134832. The probes never ran against 8.6 on that leg; the always() repro step fell through to the runner's stock PHP 8.5.9 (vm_kind=1, prints OK — the expected control result). The arm64 verdict above is unaffected.

Pushed d88d693 marking the x64 leg continue-on-error in the dispatch-only Diagnose issue 280 workflow: it's only the hybrid-VM control, and a control leg that can't obtain PHP shouldn't paint the run (and this PR) red. It starts failing loudly again the moment it gets a real 8.6 result to report — and can be flipped back to required once setup-php ships 8.6 for Intel macOS.


Generated by Claude Code

…ookTest

PHPStan 2.2.16 (resolved fresh in CI - no lockfile) infers an empty
new ArrayObject() as ArrayObject<*NEVER*, *NEVER*> and rejects the
$log[0] read. Declare the generic type the closure actually appends,
and assert the entry's presence instead of overriding the offsetGet
null with an inline @var - assertNotNull() narrows it for real.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U4cKi87mVQ896uSrw592fG
…ache

php-src's 8.6 branch grew four zend_executor_globals fields since the
committed snapshot (vm_stack_page_cache{,_count} and their fiber
counterparts; ZEND_MODULE_API_NO 20250926 -> 20260924, EG size
2336 -> 2368 on linux-x64), so every 'Generated headers up to date'
check went red against the current 8.6.0-dev build. This commit applies
the exact diff the generator produced in CI run 36777824429 (docker is
unavailable in this environment; the check re-derives and re-diffs these
files, so a mismatch would fail it again). The darwin artifacts are
refreshed by the 'Generate darwin headers' workflow dispatched right
after this push.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U4cKi87mVQ896uSrw592fG
… page cache

Same zend_executor_globals drift as the linux refresh (vm_stack page
cache fields, API 20250926 -> 20260924; EG 2168 -> 2200 nts). These are
the artifacts the 'Generate darwin headers' run 36779982938 produced on
the arm64 runners; its commit job was skipped because the x64 legs
cannot obtain PHP 8.6 from setup-php on macos-15-intel at the moment,
so the arm64 artifacts are committed from the run's uploads instead.
The darwin-x64 definitions stay at the previous snapshot until
setup-php ships 8.6 for Intel macOS again.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U4cKi87mVQ896uSrw592fG
…n 80600

Resolves the TODO left when 2.2.9 still capped phpVersion at 80599;
2.2.16 accepts 80600 and the analysis is clean at level max.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U4cKi87mVQ896uSrw592fG

Copy link
Copy Markdown
Owner Author

Recovery complete — CI on 0b80ea6 is green everywhere except two upstream-dependent groups (details below). What happened and what was pushed:

php-src's 8.6 branch moved under this branch in two ways today, and PHPStan released 2.2.16 (no lockfile, so CI resolves fresh):

  • 837c2b6 — PHPStan 2.2.16 now infers empty new ArrayObject() as ArrayObject<*NEVER*, *NEVER*>; typed the log container in ErrorCallbackHookTest and replaced the inline @var with a real assertNotNull() narrowing.
  • 542d13a / f8cd6bf — PHP 8.6 grew four zend_executor_globals vm_stack page-cache fields (API 20250926→20260924), so the committed 8.6 headers were stale. Refreshed linux from the generator's own CI diff (docker unavailable in my environment; the check re-derives and confirms) and darwin arm64 from the artifacts the dispatched Generate darwin headers run produced — its commit job was skipped because the Intel legs can't obtain PHP 8.6, so the arm64 artifacts were committed directly.
  • 0b80ea6 — phpstan.dist.neon phpVersion 80500→80600, resolving the in-tree TODO (2.2.16 accepts it; level max clean).
  • d88d693 — the dispatch-only diagnose-280 workflow's x64 control leg no longer fails the run while setup-php lacks Intel 8.6.

Remaining red, both outside this PR's control:

  1. Tests (macOS x64, nts/zts) + Generated darwin headers up to date (x64, nts/zts) — setup-php currently has no PHP 8.6 build for macos-15-intel (every attempt dies after ~27 min with "Could not setup PHP 8.6"). Clears when setup-php ships Intel 8.6 builds again.
  2. Generated headers up to date (linux) — a moving-target race: this check compiles php-src from source at check time, and 8.6 HEAD gained the NTS/ZTS module-ABI unification fields (max_execution_timer_timer/pid/oldact, copied_functions_count, globals_ptr→globals_id_ptr) between today's 21:12 and 22:17 check runs. The committed headers deliberately match setup-php's dc1145e build — the PHP the test legs actually run — because regenerating to HEAD would instead trip the test legs' ZENGINE_STRICT_LAYOUT_CHECK. Once setup-php picks up a build with the new ABI, one composer gen-headers / darwin-workflow rerun re-syncs everything.

If you'd like, these two groups can be made non-blocking (same continue-on-error treatment as the diagnose workflow) while upstream catches up — your call as it changes what the CI matrix enforces.


Generated by Claude Code

@lisachenko

Copy link
Copy Markdown
Owner Author

Reported the bug to PHP itself: php/php-src#24081

@lisachenko

Copy link
Copy Markdown
Owner Author

Should be resolved once php/php-src#24109 will be merged into the main branches

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants