Conversation
src/fx is a second Rust engine that produces the same bytes as the existing one, built for speed: - Characters are struct-of-arrays indexed by u32 slot; scenes, paths, names and visuals are u32 handles into flat tables, so the per-frame path neither allocates, hashes strings nor refcounts. - Visuals are formatted once and interned. - The active set is a bitmap; characters whose next ticks only count down doze and cost nothing until they are due. - Pure path steps are computed 8 (AVX-512) or 4 (AVX2) lanes wide ahead of the tick, and ticks that change nothing are skipped. - The cell grid is maintained incrementally from a per-frame change log; only dirty rows are re-emitted, patched in place, and a frame goes out as one iovec per row. With two or more CPUs the renderer runs on a thread of its own. - The RNG generates in batches (8 lanes side by side with AVX-512 via a precomputed jump matrix), with the same sequence. All 37 effects run on it. Effects with out-of-range duration options fall back to the existing engine. The existing engine picks up the interned visual pool, a faster renderer and input parser. Merging onto 0.4.0: - The existing engine keeps the interned (never freed) visual pool rather than omacom#31's Rc share table: its emitter compares visuals by address. - Rng::state/from_state report the stream position of the batched RNG. - The asm engine still gets the first offer; TTFX_ASM=0 runs fx. Output matches master on every effect (parity at two geometries and seeds, threaded and single-threaded); cargo test passes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The fx engine is ahead of the assembly engine on every effect, so it now takes the run first; the assembly engine only gets the runs fx declines (TTFX_FX=0, or out-of-range duration options), then the old engine. TTFX_ASM=1 (or on) restores the assembly engine's first offer for comparisons; TTFX_ASM=force does too and still exits 3 on a decline, and TTFX_ASM=0 still skips it, so the asm tooling keeps working unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
src/utils/simd.rs holds the unchecked array and vector loads and stores (debug-checked, one SAFETY each, as fx::At). The motion batch, matrix, RNG and render kernels become safe target_feature fns, so their bodies are no longer unsafe contexts; unsafe stays at the runtime dispatch and the batch gathers. The batch kernels' repeated lane masks go through small helpers. spotlights fills its visual table with resize. Output is byte-identical; instruction counts are unchanged on every effect on both render tiers. emit and its copies keep raw pointers: every slice form moved register allocation in its hot loop. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
On aarch64 the motion batch was off (tier 0), so every step took the scalar walk and no idle tick was skipped. It now has a NEON kernel, motion_batch_avx2 two lanes at a time: loads for the gather, fcvtns for cvtpd2dq, and NaN lanes left to the scalar step with the out-of-range ones. The NEON kernel pays for itself only on some effects (swarm 1.3x, spray and rings 5-18% slower), and no batch statistic predicted which. So on aarch64 whether update runs it is timed as it goes: every 256 updates the first 16 run in pairs with and without the batch, each pair a vote, and the winner runs the rest (Adapt). The first update is not timed. TTFX_NEON=on/off/flip overrides; the x86 kernels always run. update is compiled with and without the batch (update_with, not inlined): a batch-less update keeps the tick loop the compiler had when the batch could only return 0, which was 5-8% faster on aarch64. Also on aarch64: round_half_even rounds inline (frintn + fcvtzs, +inf still takes the routine) instead of always calling the cold path; wake_mask and chunk_blocks use NEON; prefetch issues prfm. simd.rs gets an aarch64 module beside the x86 one (NEON access through the same macro, and lane-mask helpers for what x86 gets from movemask); Plain and fits are shared. Tests check the mask helpers, wake_mask and chunk_blocks against scalar references. On an M1 at 200x50 (best of 7, vs 14a4615): geomean 1.029x across the 37 effects, total 1421.6 -> 1364.0 ms; swarm 1.27x, expand 1.24x, slice 1.14x, scattered 1.11x; worst fireworks 0.976x. Output matches: parity 148/148 with the batch timed, on and flipping every update, threaded and single-threaded; cargo test passes in debug and release. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
spotlights ran 14% slower than the asm engine. Its illumination code cost about the same cycles; the difference was memory: fx stalled on memory for 49% of pipeline slots (asm 16.5%), and nearly all of its DRAM traffic was first writes to fresh memory as ~140K visuals were made. The visual pool now keeps a visual's 24-byte key instead of its ~70-byte VisualInfo (info() rebuilds it; Color::from_key), and finds keys through a table of one word per entry (hash and handle, 32K entries, doubling at half full, growing without reading keys) instead of a HashMap pre-sized to 229K buckets and written all over. Visuals are formatted in place in the pool rather than on the stack and copied, and Color::from_rgb builds its hex digits in one word; both stalled loads on the stores just made. The renderer's copy of the pool is reserved to the pool's capacity (Render::fit) instead of doubling its way up. spotlights takes the asm's path: the character's last factor, then a (bright visual, factor) memo, then the pair adjusted and interned. The dense factor table and the adjusted-pair, pair-index and per-symbol tables are gone. Output is byte-identical: parity against 0.4.0 in all five modes and on an SSE2-only CPU (qemu64); cargo test passes. vs f4a1610, one core at a time: spotlights -12% cycles, -14% wall; all 37 effects geomean -1.0% cycles, -1.3% wall; instructions are at or below f4a1610 on every effect. print and laseretch run ~1% more cycles with instruction- identical hot functions that only moved (code layout). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
spotlights still ran behind the asm engine. Its illumination code cost no more than the asm's; the rest of the gap was VisualPool::make and Table::grow. 85% of make's samples sat on its first load: the fg color's key read as a word back from the caller's VisualInfo, which Color::from_rgb had just stored in pieces (a store-forwarding stall). VisualPool::make_rgb takes the sym and the adjusted colors as `1 << 24 | rgb` and builds the key in registers (Color::hex_key, shared with from_rgb); only a new visual makes the VisualInfo, in the out-of-line `add`. The table quadruples when it grows, as the asm's does, instead of doubling: spotlights rehashes twice instead of four times. Output is byte-identical: parity against 0.4.0 in all five modes and on an SSE2-only CPU (qemu64); cargo test passes. vs 7939e0b, one core at a time: spotlights -5.9% cycles and wall, colorshift -2.2%; all 37 effects geomean -0.5% cycles, -0.65% wall. randomsequence and waves run ~0.5% more cycles with instruction-identical hot functions that only moved (code layout). spotlights against the asm engine: 1 core -2.8% cycles, -4.7% wall; 2 cores -4.5% wall. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
emoon
marked this pull request as ready for review
September 27, 2026 10:20
Merged
3 tasks
52 tasks
rustfmt (edition 2021 defaults) and clippy, on the 66 files this PR changes only; the rest of the tree is left as it is. clippy, all 26 warnings in these files: - clone_on_copy: dropped .clone() on Color, ColorPair and Option<Color> - unnecessary_unwrap (engine/ctx.rs): if let instead of is_some + unwrap - needless_range_loop (bubbles, slice): iterate with enumerate - map_entry (fireworks build): the Entry API - manual_is_multiple_of (pour), type_complexity (randomsequence, unstable): is_multiple_of, local type aliases - manual_clamp (engine/animation.rs, spotlights): allowed, not changed. clamp() keeps a NaN where .min(1.0).max(0.0) gives 1.0 Output is unchanged: the quick oracle passes on all 37 effects against the previous commit (13,726 cases), the old engine (TTFX_FX=0) matches byte for byte on all 37, and cargo test --release passes (65 tests). Only the 13 edited functions differ in the disassembly; cycles on one core are even with the previous commit (geomean 0.999; a few effects move by 1-2% either way from code layout, with the same instruction counts). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Contributor
|
The optimization should be to use as little energy (Wh) as possible, longer battery time, less heat. If the context switching with several threads make it faster, but uses more energy, then it is not a win. |
|
lgtm |
The window test compared unsigned offsets against top - bottom and right - left; when the terminal reported a window with its top below its bottom (LINES=0 or negative) those spans went negative, every coordinate matched, and cells were written past the grid, crashing the run. The test now uses the window's row and column counts, clamped at 0. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MxJcVYpBxFDam6czK4bGjp
On 32-bit targets `hasher.finish() as usize >> 52` first truncates the hash to 32 bits and then shifts it by 52, more than its width: debug builds panic on the overflow, and release builds mask the shift to 20 and index the memo with other bits than on 64-bit. Shift the u64 and then narrow: the same slot on 64-bit, and the top 12 bits everywhere. (cargo check --target i686-unknown-linux-gnu passes either way; the overflow is a runtime one.) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Only TTFX_THREADS=1 kept the renderer on the effect's thread; 0 ran it threaded, the opposite of what it reads as. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The step loop trusted the motion batch's result for a slot once the epoch matched, and checked that the slot still walks the batched path only in a debug_assert. The full check (two loads) now decides: a slot whose path changed takes the scalar step instead. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
from_key was a safe pub fn that copied any word into an RgbString, whose deref trusts its length and UTF-8 without checking: a bad key read past the bytes or made an invalid str. It is now pub(crate) unsafe with the contract written down (a color_arg.key() of a Color) and debug-asserted; its one caller, VisualPool's key decoding, only holds such keys. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
src/utils is shared with the Rust engine, the oracle fx is checked against, and it should not lean on fx's unchecked At. fill_below and fill_below_pairs now index plainly; refilling on `pos >= BATCH` rather than `== BATCH` leaves pos < BATCH plain to the compiler, so neither index keeps a check. Timed against the At version (60 interleaved runs on one core): smoke, decrypt, crumble, sweep and synthgrid within noise, so fx needs no unchecked copy. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The full path compare in release cost fireworks 1.8% (100 interleaved runs on one core); the other motion effects were flat. A batched result goes stale within an epoch only when another character's tick drops the slot's mirror, and dropping it clears the slot's mirror bit, in the word the batch has just read. Release checks that bit (one hot load, within noise on fireworks); the full check stays a debug_assert. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
set_visual, set_visuals and layer_changed indexed the character tables with the effect's slot unchecked, as did step_animation and activate_scene (through doze_wake and the scene steps), and binarypath, smoke, thunderstorm and errorcorrect with their own slot-indexed tables. A bad slot there was undefined behaviour instead of a panic. Each entry point now checks the effect's slot once, and the reads after it stay unchecked with a note of the bound: every Chars column has len() entries. set_visible and coordinate_changed already checked it (flags, is_visible). The scene steps, which take their slots from the active set, show visuals through an unchecked show_visual: checking there cost colorshift 5% and beams and smoke 3%; with it, all three are within noise. binarypath checks each representation's bit slots once for its shift loop (every bit, every tick) and indexes a path's last segment checked (a path without segments would wrap). Loop indices bounded by their own range stay unchecked. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
LINES of -5, 0, 1 and 2 and COLUMNS of -3, 0, 1 and 2 shrink the visible window to a row or a column, or to nothing: with the canvas sized by the terminal, and with a 30x10 canvas larger than it from three canvas anchors and three text anchors, in the output and in the parity dump. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Contributor
|
Thank you. This is now on master via #47, with your commits intact, and fx is the engine from ttfx 0.5 on. Those measurements held up here: on a quiet Zen 5, fx was 1.20× faster than the assembly engine on one core and 1.25× on two, with byte-identical output on all 37 effects. So we removed the assembly engine. #47 also adds fixes from the review:
|
This was referenced Sep 28, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds
src/fx, a second Rust engine built for speed, running all 37 effects with byte-identical output. 11.0× faster than the Rust engine on two cores (8.2× on one), and 1.26× faster than the assembly engine from #35 (1.22× on one core) (geometric means). It is plain Rust: no NASM, and every CPU architecture builds it.Draft until the follow-ups below are done, mainly replacing the older engines with fx.
How it fits in
TTFX_FX=0skips fx.TTFX_THREADS=1forces one thread.TTFX_ASM=1(oron) gives asm the first offer again.TTFX_ASM=forcedoes too, and still exits 3 on a decline.TTFX_ASM=0skips asm.TTFX_NO_AVX512=1andTTFX_NO_AVX2=1switch off the wider motion-batch and RNG kernels, for testing.TTFX_NEON=on,offorflipoverrides the timed choice of the NEON motion batch on aarch64, for testing.Rcshare table, because its emitter compares visuals by address.Verification
TTFX_ASM=0). This holds threaded and withTTFX_THREADS=1.TTFX_NO_AVX512=1andTTFX_NO_AVX2=1, and on an emulated SSE2-only CPU (qemu-user-cpu qemu64).cargo test --releasepasses (65 tests, includingasm_diff).TTFX_THREADS=1, with the NEON motion batch timed, forced on and forced off (TTFX_NEON=on/off).cargo test --releasepasses (43 tests; the others build only on x86-64).Performance
Machine: AMD Ryzen 9 9950X3D (Zen 5), Arch Linux (kernel 7.2.3), rustc 1.93.0. Runs were pinned to the CCD without 3D V-cache (32 MB L3, like the 9955HX in #35): 1 core is CPU 14, 2 cores is CPUs 14 and 15 (two physical cores). The machine was not idle (load average about 2).
Workload (the same as #35): 200×50 canvas, 190×46 text,
--frame-rate 0,--virtual-clock, seed 1, output to/dev/null, best of 5 runs.What each column measures:
702b630(0.4.0). Rust runs withTTFX_ASM=0and is single-threaded. asm runs withTTFX_ASM=force, and withTTFX_ASM_THREADS=1for its 1-core ratio.9dc59ba). Its 1-core ratios useTTFX_THREADS=1.aarch64 (Apple M1): there is no asm engine on aarch64, so fx is compared with the Rust engine only.
Machine: Mac mini, Apple M1 (4 performance and 4 efficiency cores), macOS 26.6.2, rustc 1.98.0-nightly. macOS cannot pin threads to cores, so runs were not pinned.
Workload: the same as above (200×50 canvas, 190×46 text,
--frame-rate 0,--virtual-clockfor matrix and thunderstorm, seed 1, output to/dev/null, best of 5 runs).What each column measures:
702b630(0.4.0), single-threaded.f4a1610). It uses two threads, and its 1-thread ratios useTTFX_THREADS=1.14a4615), the geometric mean is 1.03×. swarm gains 1.27×, expand 1.24×, slice 1.13× and scattered 1.11×. No effect is slower beyond run-to-run noise (worst fireworks, 0.98×).Where the speed comes from:
u32slot. Scenes, paths, names and visuals areu32handles into flat tables, so the per-frame path neither allocates, hashes strings nor refcounts. Visuals are formatted once and interned.tools/asm/speed.pyon this branch compares fx (TTFX_ASM=0now runs fx) with asm on the same workload.Before this leaves draft
utils. The fallback for out-of-range duration options needs a home first.skip_missesin matrix (about 2%).src/fx/effects.🤖 Generated with Claude Code