build_release_files: predict translation sizes instead of building them on PRs - #11326
Conversation
|
This PR is a follow up to #11314 |
|
The |
…em on PRs On pull requests the translations other than en_US are built only to prove that they still fit in flash. The compiled code is identical for every translation; only three generated data files differ: the compressed strings, the compression dictionary and the terminal font. When the first language leaves less than 10 KB of headroom, run only the generators (the make targets for translations-<lang>.c and autogen_display_resources-<lang>.c, about 4 s, no compiler), count the bytes of the generated tables and predict the flash usage. Only translations predicted within LANGUAGE_MARGIN (1 KB) of the limit, or whose build configuration differs (clean build), are really built. LANGUAGE_PREDICT=dryrun builds everything and prints the prediction error; LANGUAGE_PREDICT=off restores the old behaviour. Measured on feather_m4_can, pybadge and metro_m4_express with every language (41 predictions): error -5..+174 bytes, almost always an over-estimate. On metro_m4_express the prediction flagged fr as not fitting; the real build then overflowed by 56 bytes. feather_m4_can board job: 443 s -> 228 s; the remainder is the two clean builds. Push and release builds are unchanged.
41a5731 to
a09cc62
Compare
dhalbert
left a comment
There was a problem hiding this comment.
Thanks for doing this. It seems to work well.
he basic structure seems fine but the variable names are too alike, and it's hard to follow: prediction, predict_flash, usage. Claude is not so good at naming, as I think you mentioned. Also I keep saying "predict what?" in my head, so could you add "flash size" or the equivalent in the printed message?
I would make changes myself but I cannot push edits.
prediction, predict_flash and usage were three near-identical names for
three different things, and two of them were tuples read by index. Name
them by their role instead: baseline_flash and flash_region for what the
first language's build establishes, baseline_translation_bytes and
translation_growth for the translation data, predicted_flash and
actual_flash for the two sides of the check. flash_usage() returns two
values so every caller unpacks them.
Say flash size in the messages too:
Predicted flash size for itsybitsy_m0_express es: 253527 of 253696
bytes (169 free, +1479 vs en_US) -> build
Flash size check itsybitsy_m0_express es: predicted 253527,
actual 253504, error +23
|
Agree 🙂 I gave him a hints for a better naming. |
dhalbert
left a comment
There was a problem hiding this comment.
Thanks - save even more time!
The translations other than en_US are built on a pull request only to check that they still fit. After the en_US build:
translations-<lang>.candautogen_display_resources-<lang>.c(~4 s, no compile), and add their size difference to the en_US figureChecked by building everything and comparing: 1552 predictions on 100 tight boards, the worst under-estimate -113 bytes, so 1KB is 9x the worst case. Board job time per pull request 23.2 h -> 19.3 h, atmel-samd -40 %, nordic -30 %. At 600 B the remaining 68 relinks would be 15. I kept 1 KB for now, but you can tune it 🙂.
The final code is AI generated.