configure: disable Apple clang size optimizations that hurt runtime speed - #23080
configure: disable Apple clang size optimizations that hurt runtime speed#23080staabm wants to merge 1 commit into
Conversation
…peed Apple clang enables the AArch64 machine outliner and hot/cold code splitting by default at -O2. Both trade speed for size: the outliner inserts extra bl/ret pairs into hot paths (measured inside the inlined zval destructor loops of zend_array_destroy, among others), and cold-split fragments (4700+ in a default cli build) are compiled for size, which mispredicted branch hints turn into slow hot code. Disabling both speeds up a CPU-bound static analysis workload (PHPStan analysing its own codebase, no opcache) by ~4% wall time on an Apple M-series machine. The flags are added only when the compiler accepts them (-Werror guards against clang's unknown -m flag warnings), so non-Apple toolchains are unaffected.
|
I played with this PR locally on my M4 Max and can confirm the speedup:
With this setup:
One interesting caveat: When compiling PHP with
|
|
@staabm @realFlowControl if you have experience with profiling on MacOS, are you able to see if the speed up is localized to some functions, or is this a more general speed up? It would be interesting to find which code paths are mislabeled as cold, too. If we can identify those we could make adjustment to the code. I would assume that at least |
|
I have no experience with profilling at the php-src level. Just today I was talking with ondrej, and he showed me the impact of pessimized PHPStan performance because of the way how PGO is trained in the homebrew-php, see shivammathur/homebrew-php#5605 (maybe there is the information you just asked for) |
|
This isn't related to PGO, if you build PHP manually on macOS it's not PGO-optimized, you need some extra steps for that. |
That's a good idea, I may be able to do a run tonight to find where the speed up is coming from |
|
So few things., and as @ondrejmirtes meant, this has nothing to do with PGO, these are "optimizations" Apple Clang is using by default. On
|
|
Thank you! About PGO: I know that we are not talking about a PGO build, but I was implying that in a PGO build the compiler could possibly make better decisions regarding what is hot or cold code, which relates directly to
Interesting. I was half expecting that Given that these flags impact only Apple Clang, and may become obsolete if these Clang optimizations become profitable in the future, it may be more relevant to apply them only in shivammathur/homebrew-php? WDYT? |
Yeah I though that too, but playing with it showed that outliner generation is completely ignoring those. |
disclaimer: this change was generated by claude opus. I have little experience with php-src development
Apple clang enables the AArch64 machine outliner and hot/cold code splitting by default at -O2. Both trade speed for size: the outliner inserts extra bl/ret pairs into hot paths (measured inside the inlined zval destructor loops of zend_array_destroy, among others), and cold-split fragments (4700+ in a default cli build) are compiled for size, which mispredicted branch hints turn into slow hot code.
Disabling both speeds up a CPU-bound static analysis workload (PHPStan analysing its own codebase, no opcache) by ~4% wall time on an Apple M-series machine. The flags are added only when the compiler accepts them (-Werror guards against clang's unknown -m flag warnings), so non-Apple toolchains are unaffected.
benchmark (3 runs after 2 warmup runs):
run
/Users/staabm/workspace/php-src/sapi/cli/php bin/phpstan clear-result-cache -q && /Users/staabm/workspace/php-src/sapi/cli/php -d memory_limit=450M bin/phpstan -vafter checking out phpstan/phpstan-src#5942 from the git root folder.(on first time phpstan-src checkout you need
composer installandmaketo prepare the codebase)before this PR:
14,47s
14,87s
14,60s
after this PR:
13,91s
13,68s
13,86s
on M4-Pro with macOS 26.6 (25G72)