Repository navigation
Commit 7af50cc
authored
perf: inline the fused binop and compare helpers (#58)
Second of a three-PR stack on `next` (`785be0e`):
1. #57 perf: grow the value stack out of line
2. #58 **perf: inline the fused binop and compare helpers** (this PR)
3. #59 perf: reserve each function's operand stack on entry
This branch includes #57 below it; the last commit is new here.
`#[inline(always)]` on `exec_binop_32/64` and `exec_cmp_32/64`. LLVM
kept them out of line, so every fused `BinOp*` / compare handler called
out for the operator `match` and spilled registers around the call.
Inlined, the `match` runs inside the handler. The binary grows by about
16 KB, and I didn't see a notable rise in L1 cache misses on Apple A14
efficiency cores in my benchmarks.
No behavior change.
**iPhone 12 E-cores**, same setup as the first PR; this PR against the
first:
| | cycles | instructions | rows faster |
|---|---:|---:|---:|
| this PR vs the first | **−1.3 %** | −2.0 % | 12/16 |
| first two PRs vs `next` | −5.3 % | −5.5 % | 16/16 |
The four rows that didn't get faster moved by 0.5 % or less.
**Checks** on this branch: tinywasm's CI matrix run on my local M4 Mac:
`cargo test --workspace` and `cargo test --workspace --examples` on rust
1.98 and with `--features tinywasm/nightly-tail-calls`, each with and
without default features.1 parent fe27adf commit 7af50cc
1 file changed
Lines changed: 2 additions & 0 deletions
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
273 | 273 | | |
274 | 274 | | |
275 | 275 | | |
| 276 | + | |
276 | 277 | | |
277 | 278 | | |
278 | 279 | | |
| |||
296 | 297 | | |
297 | 298 | | |
298 | 299 | | |
| 300 | + | |
299 | 301 | | |
300 | 302 | | |
301 | 303 | | |
| |||
0 commit comments