Skip to content

The kernel declares each CPU's HWP request, and the counters round reads the power envelope back under TRACE - #590

Merged
Japabu merged 18 commits into
mainfrom
wt/toyos-perfstate
Oct 4, 2026
Merged

Japabu merged 18 commits into
mainfrom
wt/toyos-perfstate

Conversation

@Japabu

@Japabu Japabu commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

The kernel declares each CPU's HWP request as part of the one CPU-state declaration. The power envelope it declares is read back per CPU as three counters on #705's round, under Rights::TRACE. Stage 1 of issues/kernel/the-kernel-owns-cpu-performance-state.md.

What changed, per decision

Rebuilt on #705, as the owner ruled ("Rebuild it on the counters' round now that #705 landed: one round, one place, and the power envelope becomes another counter behind the right permission"). This branch's first design is dropped whole. The merge of origin/main (550b511) takes main's tree for every file, and the commits after it land the rebuild. Gone with the first design:

  • the perf-state device class and its ABI records;
  • the second Shootdown round, served from drain_irqs, with its own bound and silent-CPU answer;
  • both actuators (perf-state-deaf-cpu, perf-request-diverges);
  • the toyos-perfstate crate, the perfstate program and its system.toml row;
  • the metal harness's expected-panic boot (--expect-panic, Ending, the census exemption).

Why this branch and not a fresh one. The diff against main is the same either way. Reusing the branch keeps #590's review rounds beside the code that answers them. Taking main's tree in the merge resolves no conflicts in src/metal.rs or tests/toyos.rs for code that was going anyway. A fresh branch would have meant a second pull request for the same track, with #590 closed unanswered.

The declaration (kernel/src/arch/x86_64/control_regs.rs). The decision comes from CPUID alone. control_regs::init_performance applies it on each CPU once that CPU's IDT is loaded: on the BSP right after idt::init, and on an AP right after control_regs::init (its trampoline loaded the IDT). A #GP in any of its wrmsr or rdmsr therefore panics by name. control_regs::init keeps only CR4 and EFER, which must precede the IDT for CR4.OSFXSR. The steps:

  1. IA32_HWP_INTERRUPT = 0, only where CPUID.6:EAX[8] enumerates it.
  2. IA32_PM_ENABLE = 1.
  3. Only now IA32_HWP_CAPABILITIES and MSR_PLATFORM_INFO are read.
  4. IA32_HWP_REQUEST, IA32_HWP_REQUEST_PKG = 0x8000ff01 and EPB = 6 are written.

Each register is written whole, logged on one line with the request's two inputs, then asserted. The order is SDM Vol. 3B §17.4.2's ("Additional MSRs associated with HWP may only be accessed after HWP is enabled, with the exception of IA32_HWP_INTERRUPT and MSR_PPERF"), and it is intel_pstate_hwp_enable's. A refused machine logs the reason once and writes none of these registers. Every AP must reach the BSP's verdict, or it panics naming itself.

One CPUID identity, three verdicts. There is no third decoder. toyos_cpuvuln::hwp sits beside decide and counters, over Vendor and Ident, and says whether a CPU's request is declared. It requires HWP, EPP, the package request and EPB, and it requires an Intel family-6 CPU that is not hybrid. toyos_cpuvuln::hwp_request gives the request. It is intel_pstate's at upstream v6.8, cited by line:

  • min is MSR_PLATFORM_INFO[47:40] (core_get_min_pstate);
  • max is the highest level, turbo included (__intel_pstate_get_hwp_cap);
  • EPP is HWP_EPP_BALANCE_PERFORMANCE;
  • desired, the activity window and package control are 0.

The refusals are ToyOS's own. The module says so, and issues/kernel/the-perf-state-declaration-refuses-hybrid-intel-and-amd-cppc.md carries them. cpu::vendor decodes CPUID.0 once, and cpu::leaf_6 reads CPUID.6 once, for both kernel callers, the declaration and the counters' bring-up. Both use the rule the bring-up already had, init_scattered_cpuid_features's: a maximum leaf outside 6..=0xFFFF reaches no leaf 6. cpu::leaf_7 reads CPUID.(7,0) once in the same way, under the >= 7 rule each of its three callers wrote out: control_regs::hwp_declared, control_regs::supported, and boot::report_counter_origin, which is outside this branch's other changes but read the same leaf under the same rule.

The read-back is three counters on the one round. They are Counter::HwpRequest, HwpRequestPkg and EnergyPerfBias, so RECORD_BYTES goes from 56 to 80. control_regs::envelope reads them on the answering CPU, inside the kick-handler answer #705 already serves. They are present only where the declaration declared them; a counter that is not present has no word, as #705's contract says. There is no second round, no second bound and no second silent-CPU answer. A CPU that misses the bound is stale with its last block, exactly as for APERF. Each read is of a register that HWP enabled at this CPU's own init_performance and that its own check read at boot, so no read can be the first to touch one.

The permission. The owner's words: "A program needs a permission to read counters; the revealing ones (power, per-device interrupts) need the strict diary permission. Admin tools get them; ordinary programs and the toolbox don't until each program can be granted rights on its own." The envelope is power, so all three need COUNTERS and TRACE. Today the registers hold what the kernel declared, so they reveal little about another program. Putting them under TRACE rather than COUNTERS alone is my reading of "power", not his words. The TRACE and SYS_COUNTERS docs say so, at toyos-abi, the SDK and the kernel. #705's two ways round the strict right still stand: inspect holds trace, and the launcher starts any row.

The metal row holds what the declaration claims (counters_on_metal):

  • every CPU's control_regs: cpuN pm_enable=1 … line;
  • each such line's hwp_request equal to Linux's on the same machine, 0x80002a04 (tests/t14-linux/hwp-request.txt). That is the independent oracle for the derivation. It is read as root through /dev/cpu/N/msr under Ubuntu's intel_pstate, and tests/t14-linux/SOURCE names its record, commit 868378b of Define the T14 LLVM bar as a libc++ stage-3 recipe with an in-tree judge; the bar value is owed #568;
  • each CPU's hwp_request, hwp_request_pkg and energy_perf_bias in every one of the three reads, equal to that line, so the bring-up must name all three;
  • every CPU's busy clock across the spin at or above Linux's busy clock over the same span of load. Linux's turbostat under one yes per CPU (tests/t14-linux/turbostat-loaded.txt) opens its first 10 s interval 2 s into the load. The floor is the lowest of the intervals that open within the spin's own span. For the 11.2 s spin that is the first interval alone, 3800 MHz, and not the 3075 MHz Linux settles to after 30 s.

That last check is the exit of issues/kernel/toyos-runs-the-t14-under-load-below-the-clock-linux-reaches.md, which this branch closes. The reading that meets it is the run at 14fa36d below. At #705's head the T14 read 2318 MHz on every CPU with HWP off.

Records kept true. issues/kernel/the-scheduler-and-the-clocks-never-talk.md said the kernel "neither reads nor sets P-states". It now says the kernel declares HWP's bounds and the CPU chooses within them. The track is rewritten to its stage 1 as built here, and RAPL, turbo and the sampler keep their exits.

The T14

At 14fa36d, run by the orchestrator (comment 5978637474). Each image's sha256 matched its request when it was flashed: shared 1882d34c…6532, shared-debug 732eea97…6fa1, testcases 9641a3f8…8c55. The command was cargo test --test toyos-build -- --metal --metal-readback orch/perfstate3/metal counters, and it exited 0 with PASS counters and [metal] 3 passed, 0 failed, 3 boot(s) (orch/perfstate3/judge-t14-metal.log):

[counters] 8 cpus, SMI +5 each over 12199 ms; TSC 2419 MHz
[counters] cpu0: idle busy 1.37% (Linux 0.11-0.36%), spinning 3829 MHz (Linux 3800 over this span, 3075-3800 loaded), busy 0.983
[counters] cpu1: idle busy 0.03% (Linux 0.11-0.36%), spinning 3828 MHz (Linux 3800 over this span, 3075-3800 loaded), busy 0.990
[counters] cpu2: idle busy 0.03% (Linux 0.11-0.36%), spinning 3829 MHz (Linux 3800 over this span, 3075-3800 loaded), busy 0.987
[counters] cpu3: idle busy 0.10% (Linux 0.11-0.36%), spinning 3828 MHz (Linux 3800 over this span, 3075-3800 loaded), busy 0.955
[counters] cpu4: idle busy 0.24% (Linux 0.11-0.36%), spinning 3828 MHz (Linux 3800 over this span, 3075-3800 loaded), busy 0.982
[counters] cpu5: idle busy 0.00% (Linux 0.11-0.36%), spinning 3829 MHz (Linux 3800 over this span, 3075-3800 loaded), busy 0.990
[counters] cpu6: idle busy 0.00% (Linux 0.11-0.36%), spinning 3829 MHz (Linux 3800 over this span, 3075-3800 loaded), busy 0.987
[counters] cpu7: idle busy 1.43% (Linux 0.11-0.36%), spinning 3829 MHz (Linux 3800 over this span, 3075-3800 loaded), busy 0.956
[counters] the spin's read, a whole round, took its reader 22882 ns
PASS counters

On all three boots every CPU's line reads pm_enable=1 hwp_request=0x80002a04 hwp_request_pkg=0x8000ff01 epb=6 hwp_interrupt=Some(0) after its CR line, and its counters: cpuN reads … line names hwp_request=true hwp_request_pkg=true energy_perf_bias=true (orch/perfstate3/metal/testcases/kernel.log).

The negative control is the whole change reverted, origin/main d47b383, staged and flashed the same way (shared eef2dc5d…afd7, shared-debug 8a3794d5…f2e2, testcases 4ead9eee…c10f) and judged by 14fa36d with --metal-readback orch/perfstate3/metal-control counters. It exited 1 with [metal] 2 passed, 1 failed, 3 boot(s) (orch/perfstate3/judge-t14-metal-control.log):

FAIL counters: idle0: cpu0 carries {"aperf": 35326893124, "hardware_id": 0, "kicks": 5, "mperf": 28947749136, "smi": 4819, "stale": 0, "stamp": 29672466476} and its bring-up said "[2026-10-04 08:54:49 0.000 cpu0] counters: cpu0 reads smi=true aperf=true mperf=true"

The floor's own red on the base is #705's reading, 2318 MHz on all eight CPUs with HWP off (orch/counters2/judge-t14-counters.log:20-27). The control cannot reach the floor, because its bring-up line reds first.

The head after the run, b69e713 and later, is held to that run by reading. Its only kernel change is cpu::leaf_7, which returns cpuid(7, 0) wherever CPUID.0's maximum leaf is at least 7, the rule each of its callers applied before. The T14's maximum leaf is 0x1b (toyos-cpuvuln/fixtures/t14/cpuid.txt:1), so every caller reads the registers it read in the run. The kernel's bytes changed and were not booted.

The judge over 8c4d1f5's readbacks (copies; the table and the readback edit are posted as a comment). 8c4d1f5, the commit before, wrote HWP before the IDT. The orchestrator ran it on the T14 in comment 5977974485 and it passed (orch/perfstate2/judge-t14-metal-2.log).

arm exit
8c4d1f5's readback 0, spinning 3829 MHz (Linux 3800 over this span, 3075-3800 loaded)
the control's readback 1, the same FAIL line
8c4d1f5's readback with every spin APERF scaled to 3100 MHz 1, some cpu below the 3800 Linux held over that span
the same, under 8c4d1f5's judge 0: the review's hole, closed
hwp-request.txt set to 0x80002a05 1, and Linux requests 0x80002a05 on this machine

Gates, at b69e713

Logs are under the orchestrator's scratchpad, orch/perfstate4/.

gate exit
cargo run -- --ci host 0, Host: 73 step(s), all green (ci-host.log:7434)
cargo run -- --build-only 0 (build-only.log)
cargo test --test toyos-build (whole guest suite) 0, 26 passed, 26 total (guest-suite.log:259)
T14 0 at 14fa36d, control 1 (above); this head held to it by reading

virt_smp under load. At 8c4d1f5 the whole suite exited 1, with virt_smp STALLED in test_rs_counters_read (orch/perfstate2/guest-suite.log:456). The orchestrator reproduced the same stall on the base binary under the same host load (orch/counterstall/base-2.log). It is a defect of main's round, and #719 ("virt_smp's second job wait goes on from the capture its first wait drained") fixes it; #719 is in the merge queue and not yet on origin/main, which records the stall nowhere else. This branch lands after #719, or after the stall is filed with the host's load. The green suite here says nothing about the stall being gone.

High-risk: CPU state, a disclosure boundary, the ABI

Negative control: the T14 run above at 14fa36d, the whole change reverted onto d47b383: exit 1 against the head's 0.

Mutations (patches posted as comments; each applied, run, reversed, tree clean):

mutation run exit red at
h1 the maximum read from bits 15:8 cargo test -p toyos-cpuvuln 101 the_request_is_the_minimum_ratio_the_highest_level_and_the_preference
h2 a hybrid CPU admitted same 101 a_missing_register_refuses_the_whole_request_by_name
h3 hwp_request answered on COUNTERS alone cargo test -p toyos-abi 101 what_times_a_program_or_is_power_needs_the_trace_right
Linux's request off by one bit the T14 judge over 8c4d1f5's readback 1 Linux requests 0x80002a05 on this machine

The kernel answers a counter only where rights.contains(counter.needs()), the one filter #705's m8 and m13 held red on QEMU. The envelope's gate is that filter plus needs(), and h3 holds needs().

Checked by reading only. Deleting hwp_check's asserts reds no test. No QEMU CPU enumerates HWP, and the T14's declaration holds. self_check's CR0/CR4/EFER asserts are in the same position on main: the no-ap-control-regs actuator exists and no test boots it. The first design reached this with a diverging actuator and an expected-panic metal boot, about 450 lines of harness, now deleted. The read-back being equal to the boot line in every read, and the boot line being Linux's request, are the runtime claims the row holds instead.

Independent oracles.

  • The SDM sentence quoted above, for the order.
  • Linux v6.8's intel_pstate.c and msr-index.h, for the derivation and the EPB value. These are upstream, not the crate's pinned Ubuntu tag, which launchpad's cgit refused to serve (403).
  • Linux's own IA32_HWP_REQUEST on the T14, tests/t14-linux/hwp-request.txt, which the row holds every CPU's boot request to.
  • The T14 under Linux's turbostat, for the loaded clock over the spin's span.
  • The T14's own CPUID (fixtures/t14/cpuid.txt) and /proc/cpuinfo flags (hwp hwp_notify hwp_epp hwp_pkg_req epb, no hybrid_cpu), which the_t14_is_declared_with_its_notification holds the verdict to.

Guest tests

None new or changed: no type, host test or guest reaches an HWP register. QEMU's CPUs enumerate no HWP, so every write and read here runs only on the T14. The x86 shared members test_rs_counters_read and test_rs_counters_silent ran on the T14's shared and shared-debug boots at 14fa36d, both passing within 3 passed, 0 failed, because the record's width changed under them.

Size

git diff --shortstat origin/main...b69e713c2: 23 files, +617 −97. Production code without comments or tests is about 235 lines added and 30 deleted: control_regs about 110, toyos-cpuvuln/src/hwp.rs 72, cpu::vendor, cpu::leaf_6 and cpu::leaf_7 22, and the rest the counter plumbing. The rest of the diff is the hwp tests (about 90), the judge (+45), the oracle file and its source, doc lines and the three issue files.

What I am unsure of

  • b69e713's kernel has not booted on the T14. Its one change, cpu::leaf_7, reads what each caller read before wherever the maximum leaf reaches 7, as the T14's does; that rests on reading.
  • The span-matched floor is 3800 MHz, and both 8c4d1f5 and 14fa36d spun at 3828 to 3829. The margin is 0.7%, against a Linux reading that started 2 s into its load while ToyOS's spin starts from idle.
  • On a multi-package machine, hwp_request_pkg is the reading CPU's package's. The T14 has one package.

🤖 Generated with Claude Code

https://claude.ai/code/session_01WcU2Dsw6mDYtwYfzVHPzM8

…s it back per CPU

The self-hosting bar is measured under a power envelope that must be read back
for a whole build span, and ToyOS left every register of it where firmware put
it. This is the first stage of the track that makes the kernel own it
(issues/kernel/the-kernel-owns-cpu-performance-state.md).

The declaration. control_regs.rs, the one CPU-state declaration, now also
writes IA32_PM_ENABLE, IA32_HWP_REQUEST, IA32_HWP_REQUEST_PKG and
IA32_ENERGY_PERF_BIAS whole on the BSP and every AP, and asserts each on
each. The values are the bar's: min is the package's maximum-efficiency
ratio (MSR_PLATFORM_INFO[47:40]), max the CPU's HWP highest performance,
desired 0, EPP 128, window 0, no package control, which on the T14's inputs
is 0x80002a04, what Linux's intel_pstate held there; the package request is
0x8000ff01 and EPB 6. The declaration is all or nothing. A CPU missing any
register it names (HWP, EPP, package request, EPB, package thermal status,
or not Intel, or hybrid) gets no request, and the BSP says why once:
"control_regs: no performance request is declared: no HWP ...". Whether the
machine declared one is the BSP's verdict and every AP must reach it; the
request itself is per CPU, since a CPU's highest performance is its own.
IA32_MISC_ENABLE's turbo bit is read back and not written, because its other
bits are model-specific and firmware's.

The arithmetic is toyos-perfstate, a pure host-tested crate, because no QEMU
CPU has HWP. Its oracle is the T14's Linux MSR readings.

The read-back needs no new syscall. It is a device class, perf-state (class 9,
the manifest's `devices = ["perf-state"]`), whose read answers
toyos_abi::perf's records: the package's registers (HWP package request,
PLATFORM_INFO, RAPL unit, PKG_POWER_LIMIT, PKG_ENERGY_STATUS, package thermal
status, TEMPERATURE_TARGET), then each CPU's (PM_ENABLE, HWP capabilities,
HWP request, EPB, MISC_ENABLE). A per-CPU MSR is readable only on its CPU,
so a read asks every CPU. It issues a generation on a second
shootdown::Shootdown (the loom-modelled TLB ack protocol), answers for its own
CPU, and kicks the rest. Each answers from its next scheduler pass in
drain_irqs, which costs two relaxed loads when nothing is owed, and posts
perf_state::WATCH. The read parks there, bounded by 250 ms, past which it is
refused Io and the silent CPUs are named. A claim exists only with the
control_regs::HwpDeclared proof, so no rdmsr of these registers is reachable
on a CPU without them. On AArch64 that proof is an uninhabited type and the
claim is refused by name.

/system/bin/perfstate holds the row and prints one read, checked against the
declaration. The guest binary perf_state runs the same code, and perf_request
drives it: in QEMU it asserts the named refusal, no request line, and the
claim refused NotFound; on the T14 it asserts every CPU logged the bar's
literal values and the binary read them all back.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W6rME2DoqwjcYFStYHHY4j
@Japabu
Japabu marked this pull request as ready for review September 28, 2026 19:51
@Japabu

Japabu commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

Review, round 1, at 545e9cd

CI host success at 545e9cd; git merge-tree against origin/main (e3a1cdc) is clean. Orchestrator runs at this head: perf_request EXIT=0, Fast EXIT=0 (233 passed), control guest-claim-granted.patch EXIT=1. No T14 reading exists. Net +1141/−17: tests about 228 lines, the issue 51, lockfiles 22, production about 840.

BLOCKER

  • (gate) The metal row is missing. Under reviewer.md, a change that targets hardware and has no reading from that hardware is NOT READY FOR REVIEW, so the proposal to land it with the row owed is refused. QEMU runs none of the declaration's writes and none of hwp_check's asserts. It runs none of Reader's cross-CPU path either, except under the mutated kernel. All of that runs at every T14 boot, and one wrong assumption there (the next finding) panics the boot of every metal row for every branch that lands after this one.
  • kernel/src/arch/x86_64/control_regs.rs:171 — The BSP reads IA32_HWP_CAPABILITIES before IA32_PM_ENABLE is written. Linux, the branch's own oracle, never does this. turbostat's print_hwp returns before reading HWP_CAPABILITIES when PM_ENABLE bit 0 is clear ("HWP is enabled and MSRs visible"), and intel_pstate enables HWP before intel_pstate_get_hwp_cap. On firmware that leaves PM_ENABLE=0, that is a #GP at boot. Fix: let CPUID decide, write PM_ENABLE, then read the capabilities and compute the request. Or cite the SDM sentence that allows the read before enable.
  • kernel/src/arch/x86_64/perf_state.rs:32-36 — 0x606, 0x610, 0x611 and 0x1A2 are not enumerated by any CPUID bit and are never read at boot. hwp_check reads only 0x770, 0x771, 0x772, 0x774, 0x1B0 and 0xCE. So the first read of those four comes from a userland SYS_READ on the claim. The kernel has no rdmsr fault fixup (git grep finds none), so on a CPU with HWP and without one of those registers, userland triggers a kernel #GP panic. Fix: drop the four fields until stage 2 or 4 lands a consumer; nothing here reads them except a println. Dropping them also takes MSR_PKG_ENERGY_STATUS off a row any session can launch; Linux restricted that register to root for CVE-2020-8694. The alternative is to read them in hwp_check so that HwpDeclared covers them.
  • kernel/src/perf_state.rs:90 — The 250 ms refusal has no test anywhere. With if !answered && !ask.deadline.reached(now) patched to if !answered, QEMU stays green (claim NotFound) and so does metal (every CPU answers). An arm is needed that silences one CPU (the dump's deaf-window actuator is how the tree does that) and asserts Io plus perf_state: cpuN did not answer.
  • kernel/src/arch/x86_64/control_regs.rs:398 — Deleting hwp_check's assert loop passes every test. The QEMU arm never reaches it, and the metal judge reads the line logged before the asserts run. control_regs_negative's actuator cannot stand in for a test, because self_check fires first. "Asserted on each CPU" needs a test that turns red on that deletion.

NOTE

  • The negative control's red is valid for its claim. The harness requires perf-state: refused NotFound, so any granted claim reds whatever rdmsr returns. The log shows the program's package check firing first, on TCG's zeros. TCG answering 0 for an unimplemented MSR means QEMU cannot model the #GP, which is why the third blocker is visible only by reading the code. The control did measure one thing: the cross-CPU read answering on two QEMU CPUs.
  • kernel/src/perf_state.rs:88 — A cancelled read (the PerfState arm's return cancelled() in syscall/io.rs) leaves pending set. The next read then either returns that old ask's slots, sampled before it asked (which contradicts toyos-abi/src/syscall.rs:1268), or a spurious Io once the old deadline has passed.
  • toyos-perfstate/src/lib.rs:201 — HwpCapabilities has one production reader, of one field (.highest), and capabilities as u8 replaces it. Its other three fields are tested for nothing the code uses. Delete it.
  • kernel/src/arch/x86_64/control_regs.rs:171 — MSR_PLATFORM_INFO (0xCE) is enumerated by no CPUID bit, so a CPU without it takes an unnamed #GP at boot. That is loud, but it is not a refusal by name.
  • control_regs.rs — IA32_HWP_INTERRUPT (0x773, CPUID.06H:EAX[8]) becomes live once PM_ENABLE is set, and nothing declares it. Declare it 0, as intel_pstate does, or record it with the turbo bit.
  • Rulings asked for. The 250 ms bound is allowed: it waits on the event (WATCH and answered()), under the tree's Budget, and names the CPUs that stayed silent. The CPU-state rule holds for the four registers: written whole, with no read-modify-write, by the BSP and every AP, and asserted on each. Leaving the turbo bit to stage 3 is acceptable, because the track records it with an exit. The device class, as class 9 plus two records with no syscall, is minimal except for the fields in the third blocker.

REMOVE

  • PR body, "Expected result: … rdmsr of 0x770 is #GP under TCG" — measured false.
  • PR body, "A virtual CPU that advertises HWP without them would #GP at boot. That is loud, not silent." — false: four of those registers are read only from the claim.
  • PR body, "The first --clippy run was red on AArch64…" — history.
  • PR body and kernel/src/perf_state.rs:135, "two relaxed loads" — a cost claim with no measurement behind it.
  • kernel/src/perf_state.rs:33, "the blocked-task dump's for the same question" — a cross-reference that rots.
  • kernel/src/shootdown.rs:1-2, "— a TLB shootdown, or a performance-state read —" — a list of callers that will grow.
  • toyos-perfstate/Cargo.toml:1-4 — narration; description and src/hostws.rs already say it.
  • toyos-perfstate/src/lib.rs:25-26, "every Intel core since Sandy Bridge has them…" — unmeasured, and the third blocker makes it unnecessary.
  • issues/kernel/the-kernel-owns-cpu-performance-state.md:11, "Until this track, ToyOS left…" — chronology, and false once this lands.
  • issues/kernel/the-kernel-owns-cpu-performance-state.md:14, "The values the kernel programs are the bar's, not its own." — false: the PR body says deriving min and max is the kernel's choice.

SEND BACK

Japabu and others added 3 commits September 28, 2026 23:33
…no claim reaches an unproven register

The review sent #590 back with five blockers. Each is answered here, except
the reading from the T14, which only that machine can give.

- control_regs: CPUID alone now decides whether a CPU gets a request.
  IA32_HWP_INTERRUPT is set to 0 where CPUID.06H:EAX[8] enumerates it, and
  IA32_PM_ENABLE is written next. Only after that are IA32_HWP_CAPABILITIES
  and MSR_PLATFORM_INFO read, and the request computed and written. This is
  intel_pstate's order (intel_pstate_hwp_enable, then
  intel_pstate_get_hwp_cap). The HWP writes moved out of the
  no-ap-control-regs skip, into hwp_init after self_check.
- toyos-perfstate: no CPUID bit enumerates MSR_PLATFORM_INFO (0xCE). SDM
  Vol. 4 documents it in the model tables of DisplayFamily 06H, so a CPU of
  any other family is refused by name (Refusal::NotFamily6) before the
  register is read.
- The read-back drops MSR_RAPL_POWER_UNIT, MSR_PKG_POWER_LIMIT,
  MSR_PKG_ENERGY_STATUS and MSR_TEMPERATURE_TARGET. No CPUID bit enumerates
  them, and nothing at boot read them, so a userland read was the first to
  touch them. The track's stages 2 and 4 now require each one to be proven
  present at boot first, and keep the energy counter (CVE-2020-8694) off any
  row a session can launch.
- A cancelled perf-state read takes its ask back (Reader::cancel, reached from
  sys_read's cancelled wait), so no later read is answered from registers
  sampled before it began. The ABI doc now says that overlapping reads of one
  claim share one ask.
- HwpCapabilities is deleted. The request's maximum is `capabilities as u8`.
- perf-state-deaf-cpu grants the claim on a machine with no declaration and
  answers zeros. The last CPU answers no ask. perf_state_silent_cpu (Fast)
  asserts two reads refused Io, each naming that CPU alone.
- perf-request-diverges has cpu1 move its request one ratio off the
  declaration when it answers a read, then run hwp_check. perf_request's
  metal row gains a second boot, perfdiverge, which must carry that panic on
  the page after the reset. It is priced in the metal profile and ruled
  flashable.
- TESTCASES now runs test_rs_perf_state, before null_sink_client_exits. The
  metal row's job_passed("test_rs_perf_state") named a job that no boot ran.
- Deleted per the review's REMOVEs: the narration in toyos-perfstate's
  Cargo.toml; "since Sandy Bridge"; "two relaxed loads"; the cross-reference
  to the dump's bound; shootdown.rs's list of callers; the track's
  chronology and the claim that its values are the bar's. The stale "four"
  in io.rs's device-class count goes too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A second Io read cannot tell a cleared ask from a stale one, since a stale
ask past its deadline is refused Io too. So the second read shows that the
claim still answers after a refusal, and nothing more. The comments now say
only that.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Japabu

Japabu commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Review, round 2, at ee20f53

CI host success at ee20f53. Orchestrator runs at this head: perf_request EXIT=0, perf_state_silent_cpu EXIT=0, Fast EXIT=0, and perf_state_silent_cpu under mutation-read-unbounded.patch EXIT=1. There is no T14 reading. git merge-tree HEAD origin/main (8a8fe27) exits 1, with conflicts in src/metal.rs and tests/toyos.rs. Net +1453/−21: tests about 420, the issue 68, lockfiles 22, production about 930.

Round-1 blockers

  • Metal row missing — OPEN. No T14 reading exists. The orchestrator holds this as a landing condition.
  • IA32_HWP_CAPABILITIES read before IA32_PM_ENABLE — CLOSED in code. hwp_declared (control_regs.rs:152) decides from CPUID alone. hwp_init (:216) then writes HWP_INTERRUPT, then PM_ENABLE, and only then reads 0x771 and 0xCE. No run exercises this path; its first measurement is the metal row's first boot.
  • Registers first touched by a userland read — CLOSED. The records hold only 0x770, 0x771, 0x774, 0x1B0 and 0x1A0 per CPU, and 0x772, 0xCE and 0x1B1 per package. Each of these is enumerated by CPUID (06H:EAX[6], EAX[7], EAX[11], ECX[3]), is architectural (0x1A0), or is read at boot behind the family gate (0xCE).
  • The 250 ms refusal had no test — CLOSED. perf_state_silent_cpu passes (EXIT=0). Under mutation-read-unbounded.patch it goes red (EXIT=1) with timed out after 300s … did not finish and no did not answer line. The patch applies cleanly at ee20f53. It removes exactly the deadline arm and nothing else, so it reverts what it claims. It is a mutation of the bound. It is not a revert of the whole change.
  • Deleting hwp_check's assert loop passed every test — OPEN. The perfdiverge boot is sound in design: its needle , the declaration is cannot match userland's , and the declaration is. But neither its green arm nor mutation-hwp-check-no-asserts.patch has run. Both are metal only.

BLOCKER

  • kernel/src/perf_state.rs:100-106 — A SYS_READ_NONBLOCK that returns WouldBlock (io.rs:280) leaves its ask in pending. Any later read of the claim, however much later, is then answered from that old ask's slots. Where a CPU never answered, it is instead refused Io without asking again. This contradicts toyos-abi/src/syscall.rs:1268 ("taken on that CPU after the read asked") and the PR's "never answered with registers sampled before it began". The cancel half has no test either: deleting claim.cancel_perf_state(); (io.rs:234) stays green. Fix the behaviour, and add a test that turns red under that deletion and under the nonblock-then-read sequence.
  • tests/toyos.rs:19806 — "each naming cpu1 alone" depends on the kernel's clock. If the test thread lands on cpu1, the deaf CPU, then cpu0 must answer a kick within the kernel's 250 ms. Otherwise cpu0 is named too. A starved TCG vCPU misses that window. This is the timing verdict that main's No QEMU test measures time, and audio is judged on metal only #562 moved out of QEMU; METAL_ONLY gives dump_nmi_probe the same reason. Make the verdict independent of time. For example, deafen every CPU but the one asking, and have the judge count exactly one named CPU per read.
  • kernel/src/shootdown.rs:73 — perf_state reads SLOTS through served's Acquire, in the target→initiator direction. kernel-loom/tests/tlb_shootdown.rs never models that direction. Patching self.flushed[cpu].load(Ordering::Acquire) to Ordering::Relaxed stays green everywhere: x86 is TSO, and loom models only initiator→target. Add a kernel-loom test in which serve's closure writes a loom cell and the initiator reads it after served(), red under that patch.

NOTE

  • Merge: conflicts with origin/main 8a8fe27 in src/metal.rs and tests/toyos.rs. The merged perf_state_silent_cpu must meet No QEMU test measures time, and audio is judged on metal only #562's rule (second blocker).
  • Oracles. intel_pstate is independent and correctly named: intel_pstate_hwp_enable, then intel_pstate_get_hwp_cap. The SDM is cited by volume only, with no section and no revision, and the sentence on which HWP MSRs are accessible before HWP_ENABLE is not quoted, so the ordering rests on intel_pstate alone. The T14 Linux readings (0x774 = 0x80002a04, 0x771, 0x772) have no command or log in this PR or in the tree. They live in unmerged Define the T14 LLVM bar as a libc++ stage-3 recipe with an in-tree judge; the bar value is owed #568's samples.
  • CPU-state rule. It holds for the five HWP registers: each is written whole, from inputs in read-only registers (0x771, 0xCE, as CR4's optional bits come from CPUID), by the BSP and every AP, and asserted on each. It is not yet true of the whole envelope: IA32_MISC_ENABLE bit 38 stays firmware's (stage 3). That is known, tracked and still true.
  • control_regs.rs:127 — HWP becomes DECLARED at the BSP's decision, before any AP has run hwp_init. HwpDeclared's doc still says HWP "is enabled there" on every CPU. It is true only because userland starts after SMP bring-up, and nothing enforces that.
  • system.toml:112 and userland/perfstate/src/main.rs — no test launches /system/bin/perfstate. Deleting devices = ["perf-state"] stays green on QEMU and would stay green on metal.
  • toyos-perfstate/src/lib.rs:88 — the NotIntel reason names RAPL, which this stage no longer reads.
  • control_regs.rs:440 — Enumerated is 12 lines for one format; inline it.

REMOVE

  • toyos-perfstate/src/lib.rs:8-10 — cites issues/build/toyos-builds-itself.md for the envelope, but neither this branch nor main has it there.
  • kernel/src/perf_state.rs:32-33 "A kicked CPU reaches a pass within one timer interrupt" — unmeasured.
  • kernel/src/shootdown.rs:73 "nothing reads through this edge yet" — false since this branch.
  • src/metal.rs:807-810 "the reset that ends the boot clears IA32_PM_ENABLE, which only a reset does" — unmeasured, and the flash ruling does not rest on it.
  • system.toml:109-110 — restates the row.
  • PR body "the harness's 30 s ceiling ends it" — measured at 300 s.
  • PR body "The 250 ms bound is the kernel's choice, and the reviewer allowed it." — history.

SEND BACK

Japabu and others added 2 commits September 29, 2026 03:27
Two conflicts, each resolved by keeping both sides:

- src/metal.rs, FLASHABLE: main's `dump-deaf-cpu` and `watch-window`
  rows, then this branch's `perf-request-diverges`. Its comment loses
  "the reset that ends the boot clears IA32_PM_ENABLE, which only a
  reset does", which the round-2 review removed as unmeasured.
- tests/toyos.rs, TESTCASES: main's job list, with `test_rs_perf_state`
  after `test_rs_abuse_short_sleep` and before `test_rs_syscall_cost`,
  so `test_rs_null_sink_client_exits` stays the last job before
  `log-close`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…dict rests on time

A read's ask now lives on `sys_read`'s stack (`perf_state::Ask`), made
by the read's first look and gone with the read. The claim holds no ask,
so a read that ends answered, refused, cancelled or as a nonblocking
`WouldBlock` leaves nothing a later read is answered from or refused by,
which is what toyos-abi/src/syscall.rs's `PerfState` doc promises.
`Reader::cancel` and `DeviceClaim::cancel_perf_state` are deleted: there
is nothing left to take back. The doc's "reads of one claim that overlap
share one ask" is deleted, since every read now asks; a CPU's one answer
still serves every ask issued before it.

With no standing ask there is no readiness before a read: `has_data` is
false and `read_watch` is `None` for the class, so a poll on a claim is
refused `NotSupported` rather than reporting a readiness the next
nonblocking read would not honour.

`perf-state-deaf-cpu` now has no CPU answer a kick: the asker answers
itself inline in `Ask::issue`, so every read on two CPUs names exactly
the one other CPU whichever CPU the test runs on, and no verdict waits
on a kicked vCPU being scheduled inside 250 ms. The refusal names the
ask (`did not answer Generation(N)`). `test_rs_perf_state_silent` makes
a nonblocking read first (the boot's first ask, `WouldBlock`), then two
blocking reads, and `perf_state_silent_cpu` requires the refusals to
name Generation(2) and Generation(3), once each, one CPU each: a read
answered or refused from an ask not its own reds.

kernel-loom/tests/shootdown_answer.rs models the target-to-initiator
edge the slots use: `serve`'s closure writes a loom cell and the
initiator reads it once `served` answers.

`HwpDeclared::ask` asserts `control_regs::report` has run, which is after
every committed CPU ran `init`, so the proof's "enabled on every CPU" is
enforced rather than true by boot order alone.

Also: `Enumerated` inlined as `{:x?}`; the `NotIntel` reason names
MSR_PLATFORM_INFO rather than RAPL; the removed comments (toyos-perfstate
lib.rs's envelope citation, perf_state.rs's "within one timer
interrupt", shootdown.rs's "nothing reads through this edge yet",
system.toml's row comment) are deleted; the track issue drops the same
false citation, says the Linux readings are #568's unmerged samples, and
records that no test launches /system/bin/perfstate.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Japabu

Japabu commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Review, round 3, at 148c207

The gate holds. gh pr view 590 --json mergeable returns MERGEABLE. CI run 36509434491 (host) is success on headSha 148c207. The orchestrator's runs at this head, read from the 590r3-*.log files:

  • perf_request: EXIT=0.
  • perf_state_silent_cpu: EXIT=0.
  • Fast: 230 of 230 passed.
  • perf_state_silent_cpu under mutation-claim-held-ask.patch: EXIT=1. The refusals name Generation(1) and Generation(2).
  • perf_state_silent_cpu under mutation-read-unbounded.patch: EXIT=1, at the 300 s ceiling.

Net diff is +1490/−18. Tests are about 327 lines, the issue 71, and lockfiles and manifests about 36. Production is about 1056 lines, including toyos-perfstate's 352. This round adds a net +45.

Earlier blockers

  • Round 2, a nonblocking or cancelled read left its ask on the claim: CLOSED. The ask now lives on sys_read's stack, and Reader holds only declared. mutation-claim-held-ask.patch reverts to the claim-held ask with no cancel, and it turns perf_state_silent_cpu red (EXIT=1, Generation(1) named).
  • Round 2, the verdict depended on time: CLOSED. Under perf-state-deaf-cpu, serve_if_owed never answers, and Ask::issue always serves the asker under the process lock, with preemption off. So each refusal names exactly the one CPU that is not the asker, whatever the schedule. The only clocks left are the harness's liveness bounds.
  • Round 2, loom did not model the target→initiator edge: CLOSED. I extracted the head into a scratch directory and ran cargo test --manifest-path kernel-loom/Cargo.toml --test shootdown_answer: EXIT=0 unpatched. I then patched Shootdown::served from Acquire to Relaxed: EXIT=101, Causality violation: Concurrent read and write accesses.
  • Round 1, T14 metal row, and mutation-hwp-check-no-asserts.patch on the perfdiverge boot: OPEN. Neither has run. The orchestrator holds both as landing conditions.
  • Round 2's HwpDeclared NOTE was also fixed. HwpDeclared::ask asserts APPLIED, which report sets after boot_aps. Every committed AP runs control_regs::init (in percpu::init_ap) before it echoes AP_STARTED.

BLOCKER

None new.

NOTE

  • kernel-loom/tests/shootdown_answer.rs — delete it: it duplicates tlb_shootdown.rs. Under the same served → Relaxed patch, --test tlb_shootdown also exits 101, with an_acknowledged_flush_postdates_the_page_table_write and one_serve_answers_two_concurrent_shootdowns FAILED. That model already reads, through served, a Relaxed atomic that the target's serve closure wrote, and that is the same shape as SLOTS. Round 2's premise that this direction was unmodelled was wrong.
  • kernel/src/object/ops.rs:396 — on SMP, a nonblocking read always issues an ask, sends an IPI to every other CPU, and answers WouldBlock. No retry can ever succeed, and poll is refused. Refusing it NotSupported by name, like poll, would delete this arm's &mut None path. The silent test would then need another way to make its first ask.
  • No test covers the claim that a poll on a perf-state claim is refused NotSupported.

REMOVE

  • kernel/src/syscall/io.rs:95 — "these four device classes": the match lists three.
  • issues/kernel/the-kernel-owns-cpu-performance-state.md — "(Define the T14 LLVM bar as a libc++ stage-3 recipe with an in-tree judge; the bar value is owed #568's samples, not merged; the metal row re-measures them)": this rots.
  • PR body — "Not ready to land: the T14 run is owed": this rots in main's merge commit.
  • PR body — "The orchestrator's runs at ee20f53d" paragraph: history.
  • PR body — every "Expected … Queued." clause, and the "Runs to queue" section.
  • PR body — "On the host: mutation-read-unbounded.patch measured EXIT=1": that run was QEMU, not the host.
  • PR body — "Run alone, --test tlb_shootdown also goes red … The new model states the perf-state direction with a cell rather than an atomic.": goes with the loom file's deletion.

The only things between this PR and landing are the two open round-1 blockers, the T14 perf_request row and its perfdiverge assert mutation, which the orchestrator holds. No change by the implementer is required beyond the NOTE and REMOVE lines.

SEND BACK

Japabu and others added 4 commits September 29, 2026 05:35
`shootdown_answer.rs` duplicated `tlb_shootdown.rs`, which already goes red
under the same `served` Acquire-to-Relaxed mutation, so its own model was
dead weight. `kernel/src/syscall/io.rs`'s device-class count was already
wrong against its match arm, and the issue's citation of an unmerged PR's
unlogged numbers rotted the moment it was written. The nonblocking-read NOTE
is filed rather than fixed, since it stays a NOTE this round.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s bound

perfdiverge runs test_rs_perf_state, the same binary testcases' own list
already prices under list.testcases.job_ms; #590 added the boot.perfdiverge.*
rows but left list.perfdiverge.job_ms unpriced, so batches() refused before
any boot with "a list nobody has priced cannot be sized to the runner's
bound". Give it the same 1400 ms ceiling for the same reason: the same kind
of binary spawned off the same stick.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Japabu and others added 2 commits September 29, 2026 09:41
…rf_request carries its job

The T14 ran perf_request's row at 1daa5dd and did what the row claims, and
the harness scored it red for two reasons of its own.

toyos-metal judged every boot that is not a staged wedge by the shutdown's
last word, so perfdiverge, which ends in the panic it exists to provoke, exited
1 before perf_request_diverged read the record. An arm now says
`panics: Some(<line>)`, the harness passes it as `--expect-panic <line>`, and
toyos-metal judges such a boot by the kernel panic record the pass after the
reset printed: green where that record (the `| ` lines under `Previous boot's
panic:`, opening with the kernel's `PANIC (apic `) carries the line, red on a
boot that handed the machine back, on a wedge's record, on a panic without the
line, and on the line found only in the ring filed after the record.
`--expect-panic` beside a wedge arm is refused: a boot ends one way. Every
other boot keeps today's rule.

A panic seals no panel census (only the stop and seal_wedge write one; the
perfdiverge readback carries none), so the per-boot facts would have reded the
same boot next: the census joins PATH_TAKEN and perfdiverge's two panel rows
go. The per-boot facts are now a function, boot_findings, so a host test can
hold the T14's own record to them.

perf_request's testcases arm named no job and rode TESTCASES' list, which a
run filtered to perf_request does not select, so that boot ran none. The arm
now names test_rs_perf_state.

The loader marking slot A dead after the expected panic is right and is left:
the mark is the loader's A/B policy for any kernel death, it lives in the
`attempts` file on the stick's log partition, and every metal boot rewrites
the whole stick (wipefs, then dd of the image), which the testcases boot after
perfdiverge shows: `this image has had the machine 0 time(s)`, slot A booted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ves unused

`cargo clippy --all-targets` builds tests/toyos.rs, which has no libtest
harness, with cfg(test) set and every #[test] dropped, so the test module's
shared fixture, selection helper and glob import were dead there and
`-D warnings` refused them. The two tests now sit at module level and carry
their own fixture and selection.

Also files usbload's unpriced panel census, found beside this: a deadline
seals the census, and no boot.usbload.panel_* row prices it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Japabu added a commit that referenced this pull request Sep 29, 2026
…ves the T14, and every high-risk stage names its mutation

The review of 2dfa123 found the security track grown from 423 to 566
lines and the speed track a 155-line design, against issues/README.md's
rule that a track carries no design and no stage table and is cut back
past a screen.

- Both tracks keep their intent in a sentence or two, and each stage is
  one entry: its name, its exit, the mutation that must go red, and its
  oracle. Mechanism, register-level design, the "What is true today"
  section, the S0 command block and every Linux and SDM line citation no
  exit needs are deleted; they land with each stage's own PR. The
  security track goes from 566 lines to 114, the speed track from 155
  to 71.
- The security track is renamed
  the-kernel-is-at-least-as-secure-as-linux-on-the-t14, because S13 goes
  past Linux parity. Every citation moves; the seven defects that called
  themselves "a row of the hardening table" lose that sentence, since the
  table is gone and the track's Exit names each of them.
- S16 (TME) moves to newer-hardware-offers-what-the-t14-lacks as a row:
  the T14's CPUID.(7,0):ECX is 0x18c05fde, bit 13 clear. Split-lock
  disable becomes S16.
- S11: the loader refuses an unmarked program or library on every CPU
  (the orchestrator's ruling), so `unmarked_object_refused` runs under
  TCG; the sysroot link reds on one unmarked input; a store to the
  shadow stack is #PF and user_ptr refuses a read into it. The libc
  longjmp clause goes: it is a panicking stub.
- S13 is gated on CET_SSS and tests that a DirectMap store to a kernel
  shadow stack faults. S14's window refuses a user copy under a key the
  thread denied; its ABI is approved in principle.
- S15, P0 and P4 need two distinct CPUs, and no boot-actuators arm or
  ABI call places a thread. The speed track gains the Pin stage, a
  test-actuators SYS_DEBUG action whose test reads the CPU's own APIC ID.
- P1 has one save path, XSAVEOPT: QEMU v11.1.1's TCG implements neither
  XSAVEC nor XSAVES (target/i386/cpu.c:1012-1014), and S11's CET state
  switches by wrmsr. P3 names a toyos-bootmap host test, P4 a loom case
  for a CPU joining the space during a shootdown, P5 an NMI under a user
  GS base; P6 takes C-states from _CST after the AML interpreter; P7
  names Linux's APST bound and quirks. The #568 and #590 citations go.
- The newer-hardware file states what it is blocked on and its exit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Japabu

Japabu commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Round 4 at 185f3ea.

CI: host is in_progress on 185f3ea (run 36539759928, job 109312103971), not green; mergeable is MERGEABLE.

NOT READY FOR REVIEW

@Japabu

Japabu commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Review, round 5, at 185f3ea

CI: host passed on 185f3ea (run 36539759928). GitHub reports the PR as MERGEABLE. I read the orchestrator's T14 logs at this head:

  • 590r2-metal-plain.log passes. It reports 8 CPUs hold the bar's request, the expected panic line, PASS perf_request and 1 passed, 0 failed. The filtered testcases boot carries 1 job(s).
  • 590r2-metal-noasserts.log fails as intended, with FAIL perfdiverge. toyos-metal gave the right reason: …it reached the shutdown's own last word.

Net branch size is +1867/−94. This round adds about +215 lines of harness production code (src/metal.rs, src/bootlog.rs, tests/common/metal.rs), about +180 lines of host tests and one issue file, and deletes 12 profile lines.

Earlier blockers

  • Round 1, no T14 metal row: CLOSED. On testcases, every CPU (cpu0 through cpu7) logs pm_enable=1 hwp_request=0x80002a04 hwp_request_pkg=0x8000ff01 epb=6, and test_rs_perf_state exits 0. The perfdiverge panic record carries the expected line.
  • Round 1, deleting hwp_check's asserts passed every test: CLOSED. With mutation-hwp-check-no-asserts.patch applied, perfdiverge reaches Rebooting. and toyos-metal refuses it as Refusal::Panic (log line 518).

QEMU runs at 148c207

The harness commits (00d6966, 185f3ea) cannot change those results. They only touch metal paths, plus one unused-by-QEMU function in bootlog. Since 148c207, I diffed the kernel files on the perf-state path, toyos-perfstate, toyos-abi/src/perf.rs, both guest binaries, and the bodies of perf_request and perf_state_silent_cpu. None changed, except a 2-line comment in io.rs. The two main merges (83c4e50, 5216494) touch none of those files. No repeat is owed for the current code. The blocker below changes the kernel, so its fix owes a new run of perf_request, perf_state_silent_cpu and the T14 row.

BLOCKER

  • kernel/src/perf_state.rs:80,142 — Nothing tests the claim that each CPU's record is read on that CPU (toyos-abi/src/syscall.rs:1269; perf.rs: "One CPU's registers, read on that CPU"). The following patch, where the asker samples every slot and the kicked CPUs answer with nothing, stays green on every arm:
    Ask::issue: ASKS.serve(me, || SLOTS[me].store(sample(me))) → ASKS.serve(me, || for s in &SLOTS[..cpus] { s.store(sample(me)) })
    serve_if_owed: ASKS.serve(me, || SLOTS[me].store(sample(me))) → ASKS.serve(me, || {})
    Why each arm stays green:

    • QEMU perf_request is refused NotFound.
    • perf_state_silent_cpu returns before serving.
    • On the T14, every field the program checks is the same on all CPUs, and the program recomputes the expected request from the slot's own capabilities.
    • In all three perfdiverge boots, the reader was the current process on cpu1 ([1.510 cpu1] Process: test_rs_perf_state pid=6 state=Live), so the divergence came from the asker's own inline sample.

    The T14 readings fit this patch. In every metal read-back, all eight CPUs report hwp_capabilities=0x0110182a, which is cpu1's boot value, while the boot lines show five different values. Bits 23:16 of that register are dynamic, so this is not proof of a bug. It is also not evidence against one.
    The branch needs a test that turns red under this patch. For example, each record could carry a value only its own CPU can produce (CPUID.01H:EBX[31:24], no MSR), checked against the roster on QEMU with a claim whose CPUs answer, and on the T14.

NOTE

  • Owner direction: the kernel's declaration hard-codes no T14 value.
    • min (PLATFORM_INFO[47:40]) and max (HWP_CAPABILITIES[7:0]) are derived on each CPU.
    • EPB 6 and EPP 0x80 are Linux's generic defaults.
    • The package request is the widest range, 1..255, and nothing uses it.
  • It is not general to every HWP-capable CPU. toyos-perfstate/src/lib.rs:128 refuses hybrid CPUs, which is every Intel client CPU since Alder Lake, and :118 refuses AMD CPPC. Both refusals give a reason. The track's "Not covered" paragraph gives neither an exit nor a stage, so it is recorded debt without an exit condition.
  • tests/toyos.rs:17758 and 2127 — BAR and PERF_REQUEST_DIVERGED are the T14's own values. They are fine as an independent oracle on the proving machine, but any second HWP machine fails this row by design.
  • tests/common/metal.rs:32 — Moving panel_max_us and panel_us into PATH_TAKEN relaxes the census check for every boot, when only a panicked boot is exempt. Pass batch.panics.is_some() into boot_findings and skip the two panel fields only then. The census stays owed everywhere else, and perfdiverge's rows stay deleted.
  • tests/common/metal.rs:858 — command_line quotes only words that contain whitespace. A word containing ', $ or ; would be passed unquoted to the shell. No current argument has one.

REMOVE

  • PR body — "The only open landing condition is the T14 metal row…" and the section "The T14 at 1daa5ddf": history that rots in main's record.
  • PR body — "It scored red for two harness reasons", and the section "The metal row, owed".
  • PR body — "At 1daa5ddf and 148c2078 (prior rounds)…": history.
  • issues/kernel/the-kernel-owns-cpu-performance-state.md — "on the T14, owed, perf_request's metal row…": false now.
  • issues/kernel/the-kernel-owns-cpu-performance-state.md — "Under Linux on the T14 every CPU held… inferred from Linux's cpuinfo_min_freq": a measurement's provenance, and it rots.

SEND BACK

Japabu and others added 2 commits September 29, 2026 11:34
…ow derives from the machine's own lines

Review round 5 found that nothing tested the ABI's claim that each CPU's
record is read on that CPU: an asker that samples every slot while the
kicked CPUs answer nothing stayed green on every arm, and every T14
read-back showed all eight CPUs at cpu1's hwp_capabilities=0x0110182a.

The sampling was right. In the one T14 read taken during a divergence
(590r2-metal-noasserts.log:474-475, the reader on cpu1, which had just
moved its own request to 0x80002a05) cpu0's record holds 0x80002a04. A
record sampled on cpu1 would hold 0x80002a05. Both records hold
0x0110182a because IA32_HWP_CAPABILITIES[23:16], Most_Efficient_Performance,
is dynamic. The boot lines were taken 11 ms apart over 233 ms, and cpu0's
own value moved from 0x04 at boot to 0x10 by the read. So the flaw was the
missing test, not the sampling.

- CpuRegisters gains hardware_id, the CPU's x2APIC ID from CPUID
  (arch::cpu::hardware_id, the ID panic records already carry). Only the
  CPU itself can read it. The guest read-back prints it on every record.
  The host holds each record to the kernel's roster, the BSP's
  `percpu: BSP ... lapic_id=` and each `SMP: AP cpuN lapic=`. On the T14
  that roster does not number CPUs in ID order (cpu1 is lapic 2).
- perf-state-deaf-cpu now leaves only the boot's first three asks to their
  asker. perf_state_silent_cpu's fourth read is therefore answered by every
  CPU, and its records are held to the roster under QEMU. perf_request's
  QEMU boot has no claim to read (NotFound is what it asserts), so the
  QEMU identity check lives here.
- perf_request's metal judges hold the machine to its own lines. Each CPU's
  boot line must hold what toyos_perfstate declares from that line's
  hwp_capabilities and platform_info. The perfdiverge panic must name the
  declaration cpu1's own boot line makes, and a request other than it.
  --expect-panic carries only the machine-independent head. No T14 literal
  is left in the judge.
- The panel census is exempt only on a boot told to panic
  (NOT_SEALED_BY_A_PANIC, keyed on Arm::panics). PATH_TAKEN is main's again.
- command_line quotes every word that holds a character a shell reads
  specially, not only words with whitespace.
- Filed issues/kernel/the-perf-state-declaration-refuses-hybrid-intel-and-amd-cppc.md,
  and the track's "Not covered" points at it. The track's owed-row clause
  and its Linux provenance are deleted.

Not done: a check that hwp_capabilities still differs across CPUs in the
read-back wherever the boot lines differ. The noasserts read above shows
the correct kernel reading equal values there, so that check would fail on
a correct kernel.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Japabu

Japabu commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Round-5 mutation patches regenerated on the merged head 2077107. Each applies with git apply --check, builds, and is reversed; the tree is clean after each.

m1-asker-samples-every-slot — records are read on their own CPU

Target: QEMU perf_state_silent_cpu and perf_request, T14 perf_request. Expected: perf_state_silent_cpu EXIT=1, perf_request EXIT=0 on QEMU; T14 perf_request red on testcases identity. Kernel cargo check --features boot-actuators,test-actuators and --build-only EXIT=0.

diff --git a/kernel/src/perf_state.rs b/kernel/src/perf_state.rs
index a0f8da924..c03d22d0d 100644
--- a/kernel/src/perf_state.rs
+++ b/kernel/src/perf_state.rs
@@ -96,7 +96,7 @@ impl Ask {
             DEAF_ASKED.fetch_add(1, Release);
         }
         let me = crate::arch::percpu::cpu_id() as usize;
-        ASKS.serve(me, || SLOTS[me].store(sample(me)));
+        ASKS.serve(me, || for s in &SLOTS[..cpus] { s.store(sample(me)) });
         for cpu in (0..cpus).filter(|&cpu| cpu != me) {
             crate::arch::irqchip::kick_cpu(cpu as u32);
         }
@@ -170,7 +170,7 @@ pub fn serve_if_owed() {
     if !ASKS.owes(me) || deaf {
         return;
     }
-    ASKS.serve(me, || SLOTS[me].store(sample(me)));
+    ASKS.serve(me, || {});
     WATCH.post();
 }
 
m2-records-carry-no-identity — no record names its CPU

Target: same. Expected: same as m1. cargo check and --build-only EXIT=0.

diff --git a/kernel/src/arch/x86_64/perf_state.rs b/kernel/src/arch/x86_64/perf_state.rs
index 49821d121..9e406acd5 100644
--- a/kernel/src/arch/x86_64/perf_state.rs
+++ b/kernel/src/arch/x86_64/perf_state.rs
@@ -17,7 +17,7 @@ pub fn declared() -> Result<Declared, &'static str> {
 /// This CPU's own identity and registers.
 pub fn read_cpu(_: &Declared) -> CpuRegisters {
     CpuRegisters {
-        hardware_id: u64::from(cpu::hardware_id()),
+        hardware_id: 0,
         pm_enable: cpu::rdmsr(msr::PM_ENABLE),
         hwp_capabilities: cpu::rdmsr(msr::HWP_CAPABILITIES),
         hwp_request: cpu::rdmsr(msr::HWP_REQUEST),
diff --git a/kernel/src/perf_state.rs b/kernel/src/perf_state.rs
index a0f8da924..8ceaff18e 100644
--- a/kernel/src/perf_state.rs
+++ b/kernel/src/perf_state.rs
@@ -195,7 +195,7 @@ fn sample(me: usize) -> CpuRegisters {
                 "an ask is made only through a claim, and a claim only where the request is \
                  declared: {why}",
             );
-            let hardware_id = u64::from(crate::arch::cpu::hardware_id());
+            let hardware_id = 0;
             CpuRegisters { hardware_id, ..CpuRegisters::default() }
         }
     }
m3-deaf-for-ever — the deaf CPU never answers, the fourth ask included

Target: QEMU perf_state_silent_cpu. Expected: perf_state_silent_cpu EXIT=1, perf_request EXIT=0. cargo check and --build-only EXIT=0.

diff --git a/kernel/src/perf_state.rs b/kernel/src/perf_state.rs
index a0f8da924..a9b237c0a 100644
--- a/kernel/src/perf_state.rs
+++ b/kernel/src/perf_state.rs
@@ -166,7 +166,7 @@ static DEAF_ASKED: AtomicU64 = AtomicU64::new(0);
 /// [`DEAF_ASKS`] asks are made, so each of those is answered by its asker alone.
 pub fn serve_if_owed() {
     let me = crate::arch::percpu::cpu_id() as usize;
-    let deaf = crate::actuator::perf_state_deaf_cpu() && DEAF_ASKED.load(Acquire) <= DEAF_ASKS;
+    let deaf = crate::actuator::perf_state_deaf_cpu() && DEAF_ASKED.load(Acquire) <= DEAF_ASKS.max(u64::MAX);
     if !ASKS.owes(me) || deaf {
         return;
     }
m4-bar-holds-no-request — the boot line's request is not held to its own declaration

Target: cargo test --test toyos-checks -- perf_request_judges. Expected: EXIT=101, a request its own inputs do not declare passed.

diff --git a/tests/toyos.rs b/tests/toyos.rs
index bd7d58e63..7d125af5b 100644
--- a/tests/toyos.rs
+++ b/tests/toyos.rs
@@ -17781,7 +17781,6 @@ fn hwp_boot_line(log: &str, cpu: u32) -> Result<u64, String> {
     let declared = toyos_perfstate::hwp_request(field("hwp_capabilities")?, field("platform_info")?);
     for (name, want) in [
         ("pm_enable", toyos_perfstate::PM_ENABLE),
-        ("hwp_request", declared),
         ("hwp_request_pkg", toyos_perfstate::HWP_REQUEST_PKG),
         ("epb", toyos_perfstate::ENERGY_PERF_BIAS),
     ] {
m5-diverged-any-declaration — the panic may name any declaration

Target: cargo test --test toyos-checks -- perf_request_judges. Expected: EXIT=101, a panic naming another declaration passed.

diff --git a/tests/toyos.rs b/tests/toyos.rs
index bd7d58e63..97f9aadff 100644
--- a/tests/toyos.rs
+++ b/tests/toyos.rs
@@ -17872,10 +17872,10 @@ fn diverged_panic(log: &str, loader: &str) -> Result<String, String> {
     let tail = format!(", the declaration is {declared:#x}");
     let held = said
         .strip_prefix(PERF_REQUEST_DIVERGED)
-        .and_then(|rest| rest.strip_suffix(&tail))
+        .and_then(|rest| rest.split_once(", the declaration is").map(|(h, _)| h))
         .and_then(|held| u64::from_str_radix(held.strip_prefix("0x")?, 16).ok());
     match held {
-        Some(held) if held != declared => Ok(said.to_string()),
+        Some(_) => Ok(said.to_string()),
         _ => Err(format!("cpu1's boot line declares {declared:#x}, and the panic says {said:?}")),
     }
 }
m6-census-exempt-on-every-boot — the panel census is exempt on every boot

Target: cargo test --test toyos-checks -- a_boot_that_is_to_panic. Expected: EXIT=101, a boot not told to panic owes no panel_max_us.

diff --git a/tests/common/metal.rs b/tests/common/metal.rs
index 25a12a09c..466198d95 100644
--- a/tests/common/metal.rs
+++ b/tests/common/metal.rs
@@ -970,7 +970,7 @@ fn boot_findings(label: &str, back: &Readback, profile: &Profile, panicked: bool
     ] {
         let name = format!("boot.{label}.{field}");
         let priced = profile.row(&name).is_some();
-        let taken = PATH_TAKEN.contains(&field) || (panicked && NOT_SEALED_BY_A_PANIC.contains(&field));
+        let taken = PATH_TAKEN.contains(&field) || NOT_SEALED_BY_A_PANIC.contains(&field);
         if value.is_none() && !priced && taken {
             continue;
         }
m7-quote-whitespace-only — only whitespace triggers quoting

Target: cargo test --test toyos-checks -- a_command_line_quotes. Expected: EXIT=101.

diff --git a/tests/common/metal.rs b/tests/common/metal.rs
index 25a12a09c..e9326f814 100644
--- a/tests/common/metal.rs
+++ b/tests/common/metal.rs
@@ -841,7 +841,7 @@ fn talk_home(home: &Path) -> PathBuf {
 /// `words` as a shell reads them back, for the request a hand runs: a word
 /// holding anything but the characters no shell treats specially is quoted.
 fn command_line(words: &[String]) -> String {
-    let plain = |c: char| c.is_ascii_alphanumeric() || "@%+:,./_-".contains(c);
+    let plain = |c: char| !c.is_whitespace();
     let word = |w: &String| {
         if !w.is_empty() && w.chars().all(plain) {
             w.clone()

Japabu and others added 2 commits October 4, 2026 07:48
…he branch's first design

The owner ruled the perf-state read-back be rebuilt on #705's counters
round: one round, one place, and the power envelope another counter
behind the strict right. Every file this branch changed is taken as
origin/main has it: the perf-state device class, its second shootdown
round, its two actuators, the `toyos-perfstate` crate, the `perfstate`
program and the metal harness's expected-panic boot go. What is kept,
the kernel declaring the HWP request from the one CPU-state declaration,
lands again in the next commit on top of what #705 provides.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WcU2Dsw6mDYtwYfzVHPzM8
…ds the power envelope back under TRACE

The CPU-state declaration in control_regs gains the performance request:
on a CPU toyos_cpuvuln::hwp admits (HWP, EPP, the package request and
EPB enumerated, Intel family 6, not hybrid), every CPU clears
IA32_HWP_INTERRUPT where it exists, sets IA32_PM_ENABLE, and only then
reads IA32_HWP_CAPABILITIES and MSR_PLATFORM_INFO and writes
IA32_HWP_REQUEST (intel_pstate's: min the package's maximum-efficiency
ratio, max the highest level, EPP 0x80), IA32_HWP_REQUEST_PKG and EPB 6
whole, then logs and asserts each. A refused machine logs the reason
once and writes none of them.

The read-back rides #705's one round instead of a second: three counters,
hwp_request, hwp_request_pkg and energy_perf_bias, each under
Rights::TRACE (the owner put power under the strict right), present only
where the declaration declared them. CPU identity is toyos-cpuvuln's: its
third verdict, `hwp`, beside the vulnerabilities and the counters, and
cpu::vendor now decodes CPUID.0 once for both kernel callers.

The counters metal row holds every CPU's declaration (pm_enable=1) and
its envelope in every read to its boot line, and every CPU's busy clock
across the spin to the lowest Linux reached loaded on the same machine,
which closes the defect that the T14 ran below it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WcU2Dsw6mDYtwYfzVHPzM8
@Japabu
Japabu marked this pull request as draft October 4, 2026 07:39
@Japabu

Japabu commented Oct 4, 2026

Copy link
Copy Markdown
Collaborator Author

Mutation patches at 8c4d1f5. Each applied with git apply --check and git apply, run, reversed with git apply -R; git status --porcelain --ignore-submodules=none empty after all three.

h1-max-from-guaranteed — h1-max-from-guaranteed cargo test -p toyos-cpuvuln EXIT=101
diff --git a/toyos-cpuvuln/src/hwp.rs b/toyos-cpuvuln/src/hwp.rs
index 636f07639..5399be092 100644
--- a/toyos-cpuvuln/src/hwp.rs
+++ b/toyos-cpuvuln/src/hwp.rs
@@ -109,7 +109,7 @@ pub const HWP_EPP: u64 = 0x80;
 /// HWP"; `msr-index.h:495-504`).
 pub const fn hwp_request(capabilities: u64, platform_info: u64) -> u64 {
     let min = (platform_info >> 40) & 0xff;
-    let max = capabilities & 0xff;
+    let max = (capabilities >> 8) & 0xff;
     min | max << 8 | HWP_EPP << 24
 }
 
h2-hybrid-admitted — h2-hybrid-admitted cargo test -p toyos-cpuvuln EXIT=101
diff --git a/toyos-cpuvuln/src/hwp.rs b/toyos-cpuvuln/src/hwp.rs
index 636f07639..b7a9d7048 100644
--- a/toyos-cpuvuln/src/hwp.rs
+++ b/toyos-cpuvuln/src/hwp.rs
@@ -88,7 +88,7 @@ pub fn hwp(facts: &HwpFacts) -> Result<Hwp, HwpRefusal> {
         Some(HwpRefusal::NoPackageRequest)
     } else if facts.cpuid_6_ecx & CPUID_6_ECX_EPB == 0 {
         Some(HwpRefusal::NoEnergyPerfBias)
-    } else if facts.cpuid_7_0_edx & CPUID_7_0_EDX_HYBRID_CPU != 0 {
+    } else if false {
         Some(HwpRefusal::Hybrid)
     } else {
         None
h3-envelope-on-counters-alone — h3-envelope-on-counters-alone cargo test -p toyos-abi EXIT=101
diff --git a/toyos-abi/src/counters.rs b/toyos-abi/src/counters.rs
index 69643bd7f..54fa5ace4 100644
--- a/toyos-abi/src/counters.rs
+++ b/toyos-abi/src/counters.rs
@@ -71,7 +71,7 @@ impl Counter {
     /// The rights a `SysCap` carries for this counter to be answered.
     pub const fn needs(self) -> Rights {
         match self {
-            Self::Stamp | Self::Smi => Rights::COUNTERS,
+            Self::Stamp | Self::Smi | Self::HwpRequest => Rights::COUNTERS,
             Self::Aperf
             | Self::Mperf
             | Self::Kicks

@Japabu Japabu changed the title The kernel declares the CPU's HWP request, and a perf-state claim reads the power envelope back per CPU The kernel declares the CPU's HWP request, and the counters round reads the power envelope back under TRACE Oct 4, 2026
@Japabu

Japabu commented Oct 4, 2026

Copy link
Copy Markdown
Collaborator Author

T14 runs at 8c4d1f57f: head GREEN (counters, incl. the 3075 MHz floor), negative control from d47b383cf RED.

Each image's sha256 matched the request in the command that flashed it.

metal/shared image sha256 4ebec901d4501af2fe0f4d686041f98f2473a6fe98be2a0446d3195e5ae5d55c
metal/shared boot rc=0
metal/shared-debug image sha256 c5ed11bc4d7dcb1ba152a8a212c7a06631c2ae75c545f82f6f65620f471de001
metal/shared-debug boot rc=0
metal/testcases image sha256 dfc9e419e7f68ff74784e2076cda7014da20fd695ef8c94e877b469ccd7007eb
metal/testcases boot rc=0
metal counters EXIT=101 
metal-control/shared image sha256 f3a389e2c884c6c6d510ada316936b8e81553b1033b57fdbdcfc462ad0a299b9
metal-control/shared boot rc=0
metal-control/shared-debug image sha256 e0d6b008ec632b7af4b5de15f3d0b845891929e356bfe5a011d3a9250b44a6f4
metal-control/shared-debug boot rc=0
metal-control/testcases image sha256 f71e21b21ac7fe677ad72e433a22df27d1d754c27c9d3f3e9747512f61b7864e
metal-control/testcases boot rc=0
metal-control counters EXIT=1 [metal] 2 passed, 1 failed, 3 boot(s)
08:05:40   FAIL counters: idle0: cpu0 carries {"aperf": 36714347572, "hardware_id": 0, "kicks": 5, "mperf": 30075184800, "smi": 4823, "stale": 0, "stamp": 30812493367} and its bring-up said "[2026-10-04 07:54:47 0.000 cpu0] counters: cpu0 reads smi=true aperf=
done

The head's first judge (judge-t14-metal.log) exited 101 before judging anything: there is no compiler to build with: …/rust/build/aarch64-apple-darwin/stage2 does not exist — the primary checkout's compiler was being rebuilt at that moment. The same readback re-judged once it existed: cargo test --test toyos-build -- --metal --metal-readback orch/perfstate2/metal counters — EXIT=0, PASS counters, [metal] 3 passed, 0 failed, 3 boot(s) (judge-t14-metal-2.log).

Readbacks: orch/perfstate2/metal/, orch/perfstate2/metal-control/.

@Japabu

Japabu commented Oct 4, 2026

Copy link
Copy Markdown
Collaborator Author

Review of the rebuild at 8c4d1f57f. The rounds before this one reviewed the dropped design, and their BLOCKERs went with its code. This round reviews the whole diff against origin/main (d47b383cf).

Size: 19 files, +560 −81. About 220 lines of that are production code; the rest is tests, the judge, doc lines and three issue files. The growth is accepted: it is one declaration, three counters on the existing round, and one verdict next to the other two.

Readings checked: orch/perfstate2/judge-t14-metal-2.log reads 3828–3829 MHz on all 8 CPUs, PASS counters, EXIT=0. The boot line on every CPU is hwp_request=0x80002a04 hwp_request_pkg=0x8000ff01 epb=6, which equals the value the body quotes for Linux's 0x774. The control at d47b383cf reds at cpu0's bring-up line, EXIT=1. The floor's own red on base is #705's 2318 MHz reading at ce1786ff0, a recorded real failure. That is enough as a control.

BLOCKER

  • PR body, "Gates" and "What I am unsure of" — the body still says "T14 … staged, not run" and "No T14 reading exists at this head". Both are now false. The only record of the T14 head run and the control run is a comment. The deleted issue's exit says "the reading that held it is in the closing pull request", and the body becomes main's record. So the close is not yet met as recorded. The body needs to carry both runs: the command, its exit code, the 3828 MHz per-CPU lines and the control's FAIL line.
  • kernel/src/arch/x86_64/control_regs.rs:161 — hwp_declared checks max_leaf >= 6 / >= 7 before reading CPUID.6. arch/x86_64/counters.rs:27 reads the same leaf behind (6..=0xFFFF).contains(&max_leaf), which follows Linux's rule. That gives one kernel two rules for whether leaf 6 exists, so a garbage max leaf gets two different answers. This is a sibling of a reader the tree already has. Read leaf 6 once, under one rule. cpu::vendor was shared this way; the leaf-6 read was not.
  • kernel/src/arch/x86_64/control_regs.rs:207 — hwp_init runs inside init, which percpu.rs:515 calls before idt::init(). The tree's own comment at that site says a fault before lidt "would stop the machine with the panel holding the record before it". The OSFXSR invariant only requires the control registers to come before the IDT; HWP has no such need. So the branch adds five wrmsr and seven rdmsr on the silent-stop path, and the body admits it ("A #GP there … stops the machine silently instead of panicking"). That is a compromise the branch found, neither removed nor recorded. Move the HWP half to a call after idt::init() on the BSP and every AP, still in control_regs, or file it with an owner and an exit.

NOTE

  • tests/toyos.rs:3399 — the spin lasts 12172 ms. turbostat-loaded.txt reads 3800 MHz for its first 20 s and 3075 MHz only after the package settles. The floor therefore holds a 12-second spin to Linux's settled clock, not to what Linux reaches over the same span. A ToyOS spinning at 3100 MHz would pass while Linux runs 3800 at that point. The exit says "Linux's loaded range". The floor meets it at 3828, but the check is weaker than the claim. Hold it to the reading that matches the spin's span, or say in the row why the settled floor is the bar.
  • tests/toyos.rs:3334 — the row holds each read-back equal to the kernel's own boot line, which is self-consistency only. The line carries hwp_capabilities and platform_info "so a reader can recompute it", but nothing recomputes it. The independent oracle exists and agrees: Linux's 0x774 = 0x80002a04 on the T14, the same value as ToyOS's boot line. It is committed nowhere. Add it to tests/t14-linux/ as a reading and hold cpu0's boot hwp_request to it, or recompute toyos_cpuvuln::hwp_request from the line's two inputs in the row.
  • PR body, "Gates" — the whole guest suite exited 1 (virt_smp STALLED in test_rs_counters_read). The base reproduced it under the same load (orch/counterstall/base-2.log), so this branch is not the cause. But git grep finds no issue on origin/main recording the stall. A red seen under load is a defect, recorded with the host's load. The orchestrator files it or lands its fix before this branch lands.

REMOVE

  • none

SEND BACK

…e row held to Linux's request and to Linux's clock over the spin's span

- The performance request moves out of control_regs::init into
  control_regs::init_performance, called on the BSP after idt::init and on
  every AP after control_regs::init (an AP's IDT is loaded by its
  trampoline). A #GP in one of its wrmsr or rdmsr now panics by name
  instead of stopping the BSP under firmware's table. CR4.OSFXSR still
  precedes the IDT.
- CPUID leaf 6 is read once, by cpu::leaf_6, under the rule the counters'
  bring-up already used (a max leaf outside 6..=0xFFFF reaches no leaf 6);
  control_regs::hwp_declared and counters::bring_up both call it.
- tests/t14-linux/hwp-request.txt commits Linux's IA32_HWP_REQUEST on the
  T14, 0x80002a04 on all eight CPUs, from commit 868378b's message (#568);
  counters_on_metal holds every CPU's boot hwp_request to it.
- The clock floor is Linux's busy clock over the intervals that open within
  the spin's own span (turbostat's first opens 2 s into the load, each
  10 s), not the lowest after the package settles: 3800 MHz for the 11.2 s
  spin, where it was 3075.

Judged against the T14 readbacks of 8c4d1f5: the head's EXIT=0, the
control's (d47b383) EXIT=1; the head's with every spin APERF scaled to
3100 MHz EXIT=1 under this judge and EXIT=0 under 8c4d1f5's; the oracle
file set to 0x80002a05 EXIT=1.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WcU2Dsw6mDYtwYfzVHPzM8
@Japabu

Japabu commented Oct 4, 2026

Copy link
Copy Markdown
Collaborator Author

Judge checks at 14fa36d, against copies of the orchestrator's T14 readbacks of 8c4d1f5 (orch/perfstate2/metal, metal-control). Each command is cargo test --test toyos-build -- --metal --metal-readback <copy> counters; the tree was clean after each.

arm readback judge exit line
head 8c4d1f5, unchanged 14fa36d 0 spinning 3829 MHz (Linux 3800 over this span, 3075-3800 loaded), PASS counters
control d47b383, unchanged 14fa36d 1 FAIL counters: idle0: cpu0 carries {…} and its bring-up said "… smi=true aperf=true mperf=true"
floor, span-matched 8c4d1f5, every spin APERF scaled to 3100 MHz (below) 14fa36d 1 FAIL counters: spinning 11172294571 ns at [3100, 3100, 3101, 3100, 3100, 3100, 3100, 3100] MHz, some cpu below the 3800 Linux held over that span
floor, settled (the NOTE's hole) same 8c4d1f5's tests/toyos.rs 0 PASS counters
oracle 8c4d1f5, unchanged 14fa36d with the patch below 1 FAIL counters: cpu0 declared "… hwp_request=0x80002a04 …", and Linux requests 0x80002a05 on this machine
oracle mutation — Linux's request off by one bit
--- a/tests/t14-linux/hwp-request.txt
+++ b/tests/t14-linux/hwp-request.txt
@@ -1 +1 @@
-0x80002a04
+0x80002a05
the readback edit — each CPU's spin APERF set to idle1's plus the spin's delta times 3100/3828 (testcases/kernel.log)
550c550
< {2026-10-04 07:45:03 13.356 pid=9 test-runner} counters_metal spin: kernel.cpu.0.aperf = 79504454927
---
> {2026-10-04 07:45:03 13.356 pid=9 test-runner} counters_metal spin: kernel.cpu.0.aperf = 71686230682
560c560
< {2026-10-04 07:45:03 13.356 pid=9 test-runner} counters_metal spin: kernel.cpu.1.aperf = 47099279881
---
> {2026-10-04 07:45:03 13.356 pid=9 test-runner} counters_metal spin: kernel.cpu.1.aperf = 39249166643
570c570
< {2026-10-04 07:45:03 13.356 pid=9 test-runner} counters_metal spin: kernel.cpu.2.aperf = 47003344511
---
> {2026-10-04 07:45:03 13.356 pid=9 test-runner} counters_metal spin: kernel.cpu.2.aperf = 39166808910
580c580
< {2026-10-04 07:45:03 13.356 pid=9 test-runner} counters_metal spin: kernel.cpu.3.aperf = 48350840204
---
> {2026-10-04 07:45:03 13.356 pid=9 test-runner} counters_metal spin: kernel.cpu.3.aperf = 40305422603
590c590
< {2026-10-04 07:45:03 13.356 pid=9 test-runner} counters_metal spin: kernel.cpu.4.aperf = 48149477983
---
> {2026-10-04 07:45:03 13.356 pid=9 test-runner} counters_metal spin: kernel.cpu.4.aperf = 40337366008
600c600
< {2026-10-04 07:45:03 13.356 pid=9 test-runner} counters_metal spin: kernel.cpu.5.aperf = 46976323556
---
> {2026-10-04 07:45:03 13.356 pid=9 test-runner} counters_metal spin: kernel.cpu.5.aperf = 39127104163
610c610
< {2026-10-04 07:45:03 13.356 pid=9 test-runner} counters_metal spin: kernel.cpu.6.aperf = 46867769073
---
> {2026-10-04 07:45:03 13.356 pid=9 test-runner} counters_metal spin: kernel.cpu.6.aperf = 39031128296
620c620
< {2026-10-04 07:45:03 13.356 pid=9 test-runner} counters_metal spin: kernel.cpu.7.aperf = 48331103391
---
> {2026-10-04 07:45:03 13.356 pid=9 test-runner} counters_metal spin: kernel.cpu.7.aperf = 40281663237

@Japabu Japabu changed the title The kernel declares the CPU's HWP request, and the counters round reads the power envelope back under TRACE The kernel declares the CPU's HWP request after the IDT, and the counters round reads the power envelope back under TRACE Oct 4, 2026
@Japabu

Japabu commented Oct 4, 2026

Copy link
Copy Markdown
Collaborator Author

T14 runs at 14fa36de1: head GREEN (new 3800 MHz floor), negative control RED.

metal/shared image sha256 1882d34c4daec246104bc7719e302050b24733c5db57ed015e7cc6309d7b6532
metal/shared boot rc=0
metal/shared-debug image sha256 732eea970137deb91a449734496a547b3ff92cf421eb7d9008587ae064356fa1
metal/shared-debug boot rc=0
metal/testcases image sha256 9641a3f82b4afb6701b616f8a205c879416b25ae6b99e34a8ddc57c4acac8c55
metal/testcases boot rc=0
metal counters EXIT=0 [metal] 3 passed, 0 failed, 3 boot(s)
08:51:22   PASS counters
metal-control/shared image sha256 eef2dc5d6d21e15c2170dec9969192a48e1d142c6dcc97c190dbfdee9c8aafd7
metal-control/shared boot rc=0
metal-control/shared-debug image sha256 8a3794d5b7fcefcdeb23ccecbb3fff59d5d5c10e957bf5f193a9e091999bf2e2
metal-control/shared-debug boot rc=0
metal-control/testcases image sha256 4ead9eee75e94b3a1c14c269abdaa0cde952d3d9de517ec7636cb49ad9eec10f
metal-control/testcases boot rc=0
metal-control counters EXIT=1 [metal] 2 passed, 1 failed, 3 boot(s)
08:55:53   FAIL counters: idle0: cpu0 carries {"aperf": 35326893124, "hardware_id": 0, "kicks": 5, "mperf": 28947749136, "smi": 4819, "stale": 0, "stamp": 29672466476} and its bring-up said "[2026-10-04 08:54:49 0.000 cpu0] counters: cpu0 reads smi=true aperf=
done

Each image's sha256 matched the request in the command that flashed it. Readbacks and judge logs: orch/perfstate3/metal*/.

@Japabu

Japabu commented Oct 4, 2026

Copy link
Copy Markdown
Collaborator Author

Review round 2 of the rebuild, at 14fa36de1. It covers the previous round's items, the diff 8c4d1f57f..14fa36de1, and the body as main's record.

Earlier BLOCKERs

  • OPEN — PR body: the T14 run at this head is still not in the record. The orchestrator's comment 09:45:43Z has it: orch/perfstate3/judge-t14-metal.log exited 0 with PASS counters and [metal] 3 passed, 0 failed, 3 boot(s), and all 8 CPUs spun at 3828–3829 MHz against Linux 3800 over this span. The control from d47b383cf (judge-t14-metal-control.log) exited 1 with FAIL counters: idle0: cpu0 carries {…}. The body still says "At this head, staged and not run", the Gates row says "T14 | staged, not run: metal/request.txt", "What I am unsure of" says "The T14 has not run this head", and "Guest tests" says the shared members "ride the staged … boots". All four are now false. The exit of toyos-runs-the-t14-under-load-below-the-clock-linux-reaches.md, which this branch deletes, is "the reading that held it is in the closing pull request". Until the body carries this head's run, that close is not met. The body needs the command, both exit codes, the eight per-CPU spinning lines and the control's FAIL line.
  • CLOSED — the two rules for leaf 6. cpu::leaf_6 (kernel/src/arch/x86_64/cpu.rs:139) has the one (6..=0xFFFF) rule, and both callers use it: counters.rs:24 and control_regs.rs:105. On the head's T14 boot every CPU logs counters: cpuN reads … hwp_request=true hwp_request_pkg=true energy_perf_bias=true (orch/perfstate3/metal/testcases/kernel.log), and the row passed.
  • CLOSED — HWP was set before the IDT. init_performance now runs after idt::init() on the BSP (percpu.rs:520-521). On an AP it runs after control_regs::init (percpu.rs:553-554), and the trampoline's lidt precedes ap_entry (smp.rs:405). counters::bring_up reads the envelope only after that, at main.rs:401 and process.rs:1816. On the head's T14 boots every CPU logs pm_enable=1 hwp_request=0x80002a04 hwp_request_pkg=0x8000ff01 epb=6 hwp_interrupt=Some(0) after its CR line, on all three boots.

Earlier NOTEs

  • Closed — the span-matched floor. The readback with each spin's APERF scaled to 3100 MHz reds under this head's judge (judge-3100-new.log, some cpu below the 3800 Linux held over that span) and passes under 8c4d1f5's (judge-3100-old.log).
  • Closed — the oracle. tests/t14-linux/hwp-request.txt is 868378b's recorded 0x80002a04, and the row holds every CPU's boot hwp_request to it. Changing one bit reds (judge-oracle-mut.log).
  • Still open, see below — the virt_smp stall.

BLOCKER

  • PR body, sections "The T14", "Gates", "Guest tests" and "What I am unsure of" — this head's T14 run and its control are missing, and four sentences say the head has not run. See OPEN above.

NOTE

REMOVE

  • none

Net: 22 files, +609 −91, about 230 of the lines production code. The growth is accepted as in the previous round.

SEND BACK

…tion

cpu::leaf_7 reads CPUID.(7,0) under the rule both control_regs sites and
boot::report_counter_origin each wrote out (a maximum leaf below 7 reaches
no leaf 7, and the registers read zero), as cpu::leaf_6 does for leaf 6.
control_regs::hwp_declared, control_regs::supported and
boot::report_counter_origin call it. On any CPU whose maximum leaf reaches
7 every caller reads the same registers as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WcU2Dsw6mDYtwYfzVHPzM8
@Japabu Japabu changed the title The kernel declares the CPU's HWP request after the IDT, and the counters round reads the power envelope back under TRACE The kernel declares each CPU's HWP request, and the counters round reads the power envelope back under TRACE Oct 4, 2026
@Japabu

Japabu commented Oct 4, 2026

Copy link
Copy Markdown
Collaborator Author

Review round 3 of the rebuild, at b69e713c2. It covers round 2's items, the diff 14fa36de1..b69e713c2, and the body as main's record.

Earlier BLOCKERs

  • CLOSED: the PR body now carries the T14 run at 14fa36d. It gives the command, exit 0, PASS counters, [metal] 3 passed, 0 failed, 3 boot(s), all eight per-CPU spinning lines, and the boot-line readings. The control at d47b383 is there too: exit 1, [metal] 2 passed, 1 failed, and the FAIL counters: idle0: cpu0 carries … line. I checked each against orch/perfstate3/judge-t14-metal.log:24-46 and judge-t14-metal-control.log:17,29,34, and every line agrees. The four sentences that said the head had not run are gone. "What I am unsure of" now says what is true: b69e713's kernel has not booted on the T14. The exit of the deleted toyos-runs-the-t14-under-load-below-the-clock-linux-reaches.md is met by the reading the body now holds.

Earlier NOTEs

  • CLOSED: the second CPUID leaf 7 read in control_regs. The tree at b69e713 has one reader, cpu::leaf_7 (kernel/src/arch/x86_64/cpu.rs:151). It is called at control_regs.rs:164, control_regs.rs:325 and boot.rs:122, and no other cpuid(7, …) remains in kernel/. Build, --ci host and the guest suite are green at this head: orch/perfstate4/gates.txt reads HEAD=b69e713c2…, ci-host EXIT=0 (Host: 73 step(s), all green), build-only EXIT=0 and guest-suite EXIT=0 (26 passed, 26 total).
  • OPEN: virt_smp under load. virt_smp's second job wait goes on from the capture its first wait drained #719 is still OPEN and is not on origin/main (tip d47b383). See NOTE below.

The claim that the T14 run at 14fa36d stands for b69e713 holds, and on a stronger ground than the body gives. Each of the three call sites is the old expression with leaf_7() put in its place:

  • boot.rs: max < 7 || ebx7 & 2 == 0 becomes leaf_7().1 & 2 == 0. Where max < 7, leaf_7() returns zeros, so the test is true there as before.
  • hwp_declared: if max >= 7 { cpuid(7,0).3 } else { 0 } is leaf_7().3 word for word.
  • supported: the same tuple, and max_leaf has no other use.

So the two versions are equal on every CPU, not only where the maximum leaf reaches 7. The new function also runs at this head on QEMU's x86 guests, because all three callers run at every boot, and the suite is green. A T14 boot would measure nothing that this substitution leaves open. No measurement is missing for this commit.

BLOCKER

  • none

NOTE

REMOVE

  • none

Net: 23 files, +617 −97. Round 3 added +10 −8, and its production change is net +2 for one reader in place of three hand-written guards. The growth is accepted as in the earlier rounds.

LAND AFTER NAMED CHANGES

@Japabu
Japabu marked this pull request as ready for review October 4, 2026 10:53
@Japabu
Japabu enabled auto-merge October 4, 2026 10:53
@Japabu
Japabu added this pull request to the merge queue Oct 4, 2026
Merged via the queue into main with commit e076793 Oct 4, 2026
6 checks passed
@Japabu
Japabu deleted the wt/toyos-perfstate branch October 4, 2026 11:33
Japabu added a commit that referenced this pull request Oct 4, 2026
The three issues #590 adds or changes under issues/kernel/ move to
issues/<slug>.md, and their two issues/kernel/ citations become
issues/<slug>.md. The branch rewrote one citation in
toyos-runs-the-t14-under-load-below-the-clock-linux-reaches.md, which #590
deletes; the deletion stands.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WcU2Dsw6mDYtwYfzVHPzM8
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant