Skip to content

Add the fx engine: 11.0x faster than Rust, 1.26x faster than asm - #44

Closed
emoon wants to merge 16 commits into
omacom:masterfrom
emoon:perf/fx-engine
Closed

emoon wants to merge 16 commits into
omacom:masterfrom
emoon:perf/fx-engine

Conversation

@emoon

@emoon emoon commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Adds src/fx, a second Rust engine built for speed, running all 37 effects with byte-identical output. 11.0× faster than the Rust engine on two cores (8.2× on one), and 1.26× faster than the assembly engine from #35 (1.22× on one core) (geometric means). It is plain Rust: no NASM, and every CPU architecture builds it.

Draft until the follow-ups below are done, mainly replacing the older engines with fx.

How it fits in

  • Each run is offered to fx first. The asm engine (Add an x86-64 assembly engine: 9.8x faster than Rust, 322x faster than Python #35) only gets the runs fx declines, and the existing Rust engine comes last. fx declines only on out-of-range duration options.
  • CPU support: on x86-64 the motion batch (8 lanes with AVX-512, 4 with AVX2), the renderer (v3/v4 builds) and the RNG (8 lanes with AVX-512) pick their SIMD kernels at runtime. Every kernel has a scalar fallback, so any x86-64 CPU (SSE2) and other architectures run fx.
  • aarch64: the motion batch has a NEON kernel (2 lanes). It only pays off on some effects, so whether a run uses it is timed as the run goes. The wake mask, chunk scan and rounding use NEON too. The RNG is still scalar there.
  • Threads: unpaced runs with two or more CPUs render and write frame N on a second thread while the main thread computes frame N+1. Paced runs stay single-threaded.
  • Environment switches:
    • TTFX_FX=0 skips fx.
    • TTFX_THREADS=1 forces one thread.
    • TTFX_ASM=1 (or on) gives asm the first offer again. TTFX_ASM=force does too, and still exits 3 on a decline. TTFX_ASM=0 skips asm.
    • TTFX_NO_AVX512=1 and TTFX_NO_AVX2=1 switch off the wider motion-batch and RNG kernels, for testing.
    • TTFX_NEON=on, off or flip overrides the timed choice of the NEON motion batch on aarch64, for testing.
  • Existing Rust engine: it gains an interned visual pool, a faster renderer and a faster input parser. It keeps the never-freed visual pool rather than Reduce memory between 2%-77% depending on the effect #31's Rc share table, because its emitter compares visuals by address.

Verification

  • Parity against master (0.4.0): output is byte-identical on every effect, at seeds 1 and 7, 200×50 and 80×24, against both master's asm engine and its Rust engine (TTFX_ASM=0). This holds threaded and with TTFX_THREADS=1.
  • Lower SIMD tiers: the same holds with TTFX_NO_AVX512=1 and TTFX_NO_AVX2=1, and on an emulated SSE2-only CPU (qemu-user -cpu qemu64).
  • cargo test --release passes (65 tests, including asm_diff).
  • aarch64 (Apple M1): output is byte-identical to master's Rust engine on the same effects, seeds and sizes, threaded and with TTFX_THREADS=1, with the NEON motion batch timed, forced on and forced off (TTFX_NEON=on/off). cargo test --release passes (43 tests; the others build only on x86-64).

Performance

Machine: AMD Ryzen 9 9950X3D (Zen 5), Arch Linux (kernel 7.2.3), rustc 1.93.0. Runs were pinned to the CCD without 3D V-cache (32 MB L3, like the 9955HX in #35): 1 core is CPU 14, 2 cores is CPUs 14 and 15 (two physical cores). The machine was not idle (load average about 2).

Workload (the same as #35): 200×50 canvas, 190×46 text, --frame-rate 0, --virtual-clock, seed 1, output to /dev/null, best of 5 runs.

What each column measures:

  • Rust and asm are master 702b630 (0.4.0). Rust runs with TTFX_ASM=0 and is single-threaded. asm runs with TTFX_ASM=force, and with TTFX_ASM_THREADS=1 for its 1-core ratio.
  • fx is this branch (9dc59ba). Its 1-core ratios use TTFX_THREADS=1.
effect Rust ms asm ms (2 cores) fx ms (2 cores) vs Rust, 1 core vs Rust, 2 cores vs asm, 1 core vs asm, 2 cores
beams 147.3 10.1 8.0 12.9× 18.5× 1.24× 1.26×
binarypath 584.3 120.8 52.6 6.2× 11.1× 2.03× 2.30×
blackhole 310.8 42.9 37.4 6.1× 8.3× 1.15× 1.15×
bouncyballs 213.1 23.1 18.8 7.8× 11.4× 1.21× 1.23×
bubbles 328.7 30.5 24.4 7.8× 13.5× 1.23× 1.25×
burn 230.3 16.6 11.2 12.3× 20.6× 1.45× 1.48×
colorshift 241.8 16.9 13.5 14.8× 17.9× 1.12× 1.25×
crumble 202.9 39.1 32.6 4.5× 6.2× 1.18× 1.20×
decrypt 276.3 14.8 11.9 18.7× 23.2× 1.21× 1.25×
errorcorrect 165.0 16.0 11.1 9.7× 14.9× 1.32× 1.45×
expand 84.6 13.7 11.1 5.3× 7.6× 1.17× 1.23×
fireworks 235.6 50.6 38.4 4.6× 6.1× 1.29× 1.32×
highlight 34.4 3.4 2.9 11.6× 11.9× 1.12× 1.16×
laseretch 349.7 36.6 25.2 7.8× 13.9× 1.39× 1.45×
matrix 141.4 17.0 14.7 6.7× 9.6× 1.21× 1.16×
middleout 60.6 8.5 7.0 7.6× 8.7× 1.19× 1.22×
orbittingvolley 54.1 12.8 10.4 3.4× 5.2× 1.17× 1.23×
overflow 89.6 11.3 9.8 7.0× 9.1× 1.12× 1.15×
pour 142.5 14.7 11.7 8.3× 12.2× 1.14× 1.25×
print 196.9 7.3 5.8 28.4× 33.8× 1.07× 1.25×
rain 143.7 15.8 13.1 7.6× 11.0× 1.18× 1.21×
randomsequence 33.9 3.4 3.1 8.8× 10.9× 1.08× 1.10×
rings 486.0 85.4 67.4 5.4× 7.2× 1.21× 1.27×
scattered 97.8 18.2 14.2 4.3× 6.9× 1.19× 1.29×
slice 47.8 6.5 5.6 7.6× 8.5× 1.13× 1.15×
slide 60.3 9.5 8.2 5.2× 7.4× 1.14× 1.16×
smoke 117.2 10.3 6.2 14.7× 19.0× 1.59× 1.67×
spotlights 237.9 27.2 26.5 8.2× 9.0× 1.06× 1.03×
spray 119.0 21.1 18.5 4.6× 6.4× 1.14× 1.14×
swarm 438.7 70.8 62.2 4.8× 7.1× 1.15× 1.14×
sweep 40.0 4.4 3.6 9.9× 11.2× 1.28× 1.25×
synthgrid 75.4 6.7 5.4 12.9× 14.0× 1.13× 1.24×
thunderstorm 260.1 19.3 12.5 18.6× 20.8× 1.58× 1.55×
unstable 121.8 24.7 21.8 3.5× 5.6× 1.10× 1.14×
vhstape 141.5 22.9 17.8 7.1× 8.0× 1.25× 1.29×
waves 313.6 14.0 11.3 21.4× 27.7× 1.18× 1.23×
wipe 29.9 3.3 2.9 9.6× 10.3× 1.08× 1.14×
geometric mean 8.17× 10.98× 1.22× 1.26×
  • Instructions vs cycles: against asm, fx executes about as many instructions (0.99×) but needs 17% fewer cycles and 18% less CPU time.
  • spotlights, the last effect behind asm, is now ahead of it: 1.06× on one core, 1.03× on two.
  • binarypath, smoke and thunderstorm gain the most over asm: 1.5–2.2×.

aarch64 (Apple M1): there is no asm engine on aarch64, so fx is compared with the Rust engine only.

Machine: Mac mini, Apple M1 (4 performance and 4 efficiency cores), macOS 26.6.2, rustc 1.98.0-nightly. macOS cannot pin threads to cores, so runs were not pinned.

Workload: the same as above (200×50 canvas, 190×46 text, --frame-rate 0, --virtual-clock for matrix and thunderstorm, seed 1, output to /dev/null, best of 5 runs).

What each column measures:

  • Rust is master 702b630 (0.4.0), single-threaded.
  • fx is this branch (f4a1610). It uses two threads, and its 1-thread ratios use TTFX_THREADS=1.
effect Rust ms fx ms (2 threads) vs Rust, 1 thread vs Rust, 2 threads
beams 197.4 15.1 9.2× 13.1×
binarypath 978.2 84.9 7.4× 11.5×
blackhole 658.6 96.6 5.7× 6.8×
bouncyballs 348.3 39.6 5.4× 8.8×
bubbles 554.3 50.4 6.2× 11.0×
burn 316.6 24.8 8.3× 12.8×
colorshift 455.8 27.5 14.1× 16.6×
crumble 310.9 58.1 4.1× 5.4×
decrypt 483.4 25.9 16.3× 18.7×
errorcorrect 275.9 26.9 7.9× 10.3×
expand 193.9 28.0 5.7× 6.9×
fireworks 403.8 87.6 3.8× 4.6×
highlight 50.7 6.4 7.5× 7.9×
laseretch 548.9 60.4 5.7× 9.1×
matrix 175.7 52.3 2.7× 3.4×
middleout 121.3 18.6 6.0× 6.5×
orbittingvolley 73.0 18.9 2.6× 3.9×
overflow 135.3 16.3 6.5× 8.3×
pour 261.6 28.0 6.2× 9.3×
print 326.2 18.9 15.7× 17.3×
rain 218.8 26.7 5.1× 8.2×
randomsequence 46.2 6.8 5.9× 6.8×
rings 608.6 110.1 4.4× 5.5×
scattered 203.4 36.1 4.2× 5.6×
slice 87.9 18.6 4.5× 4.7×
slide 95.2 21.0 3.6× 4.5×
smoke 157.7 12.2 9.6× 12.9×
spotlights 299.3 39.6 6.8× 7.6×
spray 164.6 31.8 4.0× 5.2×
swarm 557.7 137.1 3.0× 4.1×
sweep 63.4 7.7 7.1× 8.2×
synthgrid 100.0 11.2 7.8× 8.9×
thunderstorm 421.3 22.5 16.6× 18.7×
unstable 264.4 35.5 5.0× 7.4×
vhstape 226.9 31.9 6.3× 7.1×
waves 494.9 19.8 19.3× 25.0×
wipe 49.8 6.2 7.4× 8.0×
geometric mean 6.37× 8.19×
  • Geometric mean: 6.37× on one thread and 8.19× on two, against 8.17× and 10.98× on the Zen 5.
  • waves, decrypt, thunderstorm and print gain the most: 16–25× on two threads.
  • matrix, orbittingvolley and swarm gain the least: 3.4–4.1×.
  • NEON: against the branch without the NEON kernels (14a4615), the geometric mean is 1.03×. swarm gains 1.27×, expand 1.24×, slice 1.13× and scattered 1.11×. No effect is slower beyond run-to-run noise (worst fireworks, 0.98×).

Where the speed comes from:

  1. Data layout: characters are struct-of-arrays indexed by u32 slot. Scenes, paths, names and visuals are u32 handles into flat tables, so the per-frame path neither allocates, hashes strings nor refcounts. Visuals are formatted once and interned.
  2. Skipping work:
    • The active set is a bitmap.
    • Characters whose next ticks only count down "doze" and cost nothing until they are due.
    • Ticks that change nothing are skipped.
  3. Incremental output: the cell grid is updated from a per-frame change log. Only dirty rows are re-emitted, patched in place, and a frame goes out as one iovec per row.
  4. SIMD and threads:
    • Path steps are computed 8 or 4 lanes wide ahead of the tick.
    • The RNG fills 8-lane batches via a precomputed jump matrix.
    • The render thread.

tools/asm/speed.py on this branch compares fx (TTFX_ASM=0 now runs fx) with asm on the same workload.

Before this leaves draft

  • Replace the older engines with fx (asm and/or the existing Rust engine), then simplify utils. The fallback for out-of-range duration options needs a home first.
  • aarch64: a NEON RNG refill (the RNG is about 6% of the RNG-heavy effects' time) and a NEON skip_misses in matrix (about 2%).
  • Performance on more hardware: Intel, AVX2-only, no AVX2, 1–2 core machines, and aarch64 Linux.
  • Clippy style warnings in src/fx/effects.

🤖 Generated with Claude Code

emoon and others added 6 commits September 27, 2026 07:43
src/fx is a second Rust engine that produces the same bytes as the
existing one, built for speed:

- Characters are struct-of-arrays indexed by u32 slot; scenes, paths,
  names and visuals are u32 handles into flat tables, so the per-frame
  path neither allocates, hashes strings nor refcounts.
- Visuals are formatted once and interned.
- The active set is a bitmap; characters whose next ticks only count
  down doze and cost nothing until they are due.
- Pure path steps are computed 8 (AVX-512) or 4 (AVX2) lanes wide ahead
  of the tick, and ticks that change nothing are skipped.
- The cell grid is maintained incrementally from a per-frame change log;
  only dirty rows are re-emitted, patched in place, and a frame goes out
  as one iovec per row. With two or more CPUs the renderer runs on a
  thread of its own.
- The RNG generates in batches (8 lanes side by side with AVX-512 via a
  precomputed jump matrix), with the same sequence.

All 37 effects run on it. Effects with out-of-range duration options
fall back to the existing engine. The existing engine picks up the
interned visual pool, a faster renderer and input parser.

Merging onto 0.4.0:
- The existing engine keeps the interned (never freed) visual pool rather
  than omacom#31's Rc share table: its emitter compares visuals by address.
- Rng::state/from_state report the stream position of the batched RNG.
- The asm engine still gets the first offer; TTFX_ASM=0 runs fx.

Output matches master on every effect (parity at two geometries and
seeds, threaded and single-threaded); cargo test passes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The fx engine is ahead of the assembly engine on every effect, so it now
takes the run first; the assembly engine only gets the runs fx declines
(TTFX_FX=0, or out-of-range duration options), then the old engine.

TTFX_ASM=1 (or on) restores the assembly engine's first offer for
comparisons; TTFX_ASM=force does too and still exits 3 on a decline, and
TTFX_ASM=0 still skips it, so the asm tooling keeps working unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
src/utils/simd.rs holds the unchecked array and vector loads and stores
(debug-checked, one SAFETY each, as fx::At). The motion batch, matrix,
RNG and render kernels become safe target_feature fns, so their bodies
are no longer unsafe contexts; unsafe stays at the runtime dispatch and
the batch gathers. The batch kernels' repeated lane masks go through
small helpers. spotlights fills its visual table with resize.

Output is byte-identical; instruction counts are unchanged on every
effect on both render tiers. emit and its copies keep raw pointers:
every slice form moved register allocation in its hot loop.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
On aarch64 the motion batch was off (tier 0), so every step took the
scalar walk and no idle tick was skipped. It now has a NEON kernel,
motion_batch_avx2 two lanes at a time: loads for the gather, fcvtns for
cvtpd2dq, and NaN lanes left to the scalar step with the out-of-range
ones.

The NEON kernel pays for itself only on some effects (swarm 1.3x, spray
and rings 5-18% slower), and no batch statistic predicted which. So on
aarch64 whether update runs it is timed as it goes: every 256 updates the
first 16 run in pairs with and without the batch, each pair a vote, and
the winner runs the rest (Adapt). The first update is not timed.
TTFX_NEON=on/off/flip overrides; the x86 kernels always run.

update is compiled with and without the batch (update_with, not inlined):
a batch-less update keeps the tick loop the compiler had when the batch
could only return 0, which was 5-8% faster on aarch64.

Also on aarch64: round_half_even rounds inline (frintn + fcvtzs, +inf
still takes the routine) instead of always calling the cold path;
wake_mask and chunk_blocks use NEON; prefetch issues prfm.

simd.rs gets an aarch64 module beside the x86 one (NEON access through
the same macro, and lane-mask helpers for what x86 gets from movemask);
Plain and fits are shared. Tests check the mask helpers, wake_mask and
chunk_blocks against scalar references.

On an M1 at 200x50 (best of 7, vs 14a4615): geomean 1.029x across the 37
effects, total 1421.6 -> 1364.0 ms; swarm 1.27x, expand 1.24x, slice
1.14x, scattered 1.11x; worst fireworks 0.976x. Output matches: parity
148/148 with the batch timed, on and flipping every update, threaded
and single-threaded; cargo test passes in debug and release.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
spotlights ran 14% slower than the asm engine. Its illumination code
cost about the same cycles; the difference was memory: fx stalled on
memory for 49% of pipeline slots (asm 16.5%), and nearly all of its DRAM
traffic was first writes to fresh memory as ~140K visuals were made.

The visual pool now keeps a visual's 24-byte key instead of its ~70-byte
VisualInfo (info() rebuilds it; Color::from_key), and finds keys through
a table of one word per entry (hash and handle, 32K entries, doubling at
half full, growing without reading keys) instead of a HashMap pre-sized
to 229K buckets and written all over. Visuals are formatted in place in
the pool rather than on the stack and copied, and Color::from_rgb builds
its hex digits in one word; both stalled loads on the stores just made.
The renderer's copy of the pool is reserved to the pool's capacity
(Render::fit) instead of doubling its way up.

spotlights takes the asm's path: the character's last factor, then a
(bright visual, factor) memo, then the pair adjusted and interned. The
dense factor table and the adjusted-pair, pair-index and per-symbol
tables are gone.

Output is byte-identical: parity against 0.4.0 in all five modes and on
an SSE2-only CPU (qemu64); cargo test passes. vs f4a1610, one core at a
time: spotlights -12% cycles, -14% wall; all 37 effects geomean -1.0%
cycles, -1.3% wall; instructions are at or below f4a1610 on every
effect. print and laseretch run ~1% more cycles with instruction-
identical hot functions that only moved (code layout).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
spotlights still ran behind the asm engine. Its illumination code cost
no more than the asm's; the rest of the gap was VisualPool::make and
Table::grow.

85% of make's samples sat on its first load: the fg color's key read as
a word back from the caller's VisualInfo, which Color::from_rgb had just
stored in pieces (a store-forwarding stall). VisualPool::make_rgb takes
the sym and the adjusted colors as `1 << 24 | rgb` and builds the key in
registers (Color::hex_key, shared with from_rgb); only a new visual
makes the VisualInfo, in the out-of-line `add`.

The table quadruples when it grows, as the asm's does, instead of
doubling: spotlights rehashes twice instead of four times.

Output is byte-identical: parity against 0.4.0 in all five modes and on
an SSE2-only CPU (qemu64); cargo test passes. vs 7939e0b, one core at a
time: spotlights -5.9% cycles and wall, colorshift -2.2%; all 37 effects
geomean -0.5% cycles, -0.65% wall. randomsequence and waves run ~0.5%
more cycles with instruction-identical hot functions that only moved
(code layout). spotlights against the asm engine: 1 core -2.8% cycles,
-4.7% wall; 2 cores -4.5% wall.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
rustfmt (edition 2021 defaults) and clippy, on the 66 files this PR
changes only; the rest of the tree is left as it is.

clippy, all 26 warnings in these files:
- clone_on_copy: dropped .clone() on Color, ColorPair and Option<Color>
- unnecessary_unwrap (engine/ctx.rs): if let instead of is_some + unwrap
- needless_range_loop (bubbles, slice): iterate with enumerate
- map_entry (fireworks build): the Entry API
- manual_is_multiple_of (pour), type_complexity (randomsequence,
  unstable): is_multiple_of, local type aliases
- manual_clamp (engine/animation.rs, spotlights): allowed, not changed.
  clamp() keeps a NaN where .min(1.0).max(0.0) gives 1.0

Output is unchanged: the quick oracle passes on all 37 effects against
the previous commit (13,726 cases), the old engine (TTFX_FX=0) matches
byte for byte on all 37, and cargo test --release passes (65 tests).
Only the 13 edited functions differ in the disassembly; cycles on one
core are even with the previous commit (geomean 0.999; a few effects
move by 1-2% either way from code layout, with the same instruction
counts).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@hegjon

hegjon commented Sep 27, 2026

Copy link
Copy Markdown
Contributor

The optimization should be to use as little energy (Wh) as possible, longer battery time, less heat. If the context switching with several threads make it faster, but uses more energy, then it is not a win.

@emailschaden

Copy link
Copy Markdown

lgtm

dhh and others added 9 commits September 27, 2026 15:41
The window test compared unsigned offsets against top - bottom and right -
left; when the terminal reported a window with its top below its bottom
(LINES=0 or negative) those spans went negative, every coordinate matched,
and cells were written past the grid, crashing the run. The test now uses
the window's row and column counts, clamped at 0.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MxJcVYpBxFDam6czK4bGjp
On 32-bit targets `hasher.finish() as usize >> 52` first truncates the
hash to 32 bits and then shifts it by 52, more than its width: debug
builds panic on the overflow, and release builds mask the shift to 20
and index the memo with other bits than on 64-bit. Shift the u64 and
then narrow: the same slot on 64-bit, and the top 12 bits everywhere.
(cargo check --target i686-unknown-linux-gnu passes either way; the
overflow is a runtime one.)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Only TTFX_THREADS=1 kept the renderer on the effect's thread; 0 ran it
threaded, the opposite of what it reads as.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The step loop trusted the motion batch's result for a slot once the
epoch matched, and checked that the slot still walks the batched path
only in a debug_assert. The full check (two loads) now decides: a slot
whose path changed takes the scalar step instead.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
from_key was a safe pub fn that copied any word into an RgbString, whose
deref trusts its length and UTF-8 without checking: a bad key read past
the bytes or made an invalid str. It is now pub(crate) unsafe with the
contract written down (a color_arg.key() of a Color) and debug-asserted;
its one caller, VisualPool's key decoding, only holds such keys.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
src/utils is shared with the Rust engine, the oracle fx is checked
against, and it should not lean on fx's unchecked At. fill_below and
fill_below_pairs now index plainly; refilling on `pos >= BATCH` rather
than `== BATCH` leaves pos < BATCH plain to the compiler, so neither
index keeps a check. Timed against the At version (60 interleaved runs
on one core): smoke, decrypt, crumble, sweep and synthgrid within
noise, so fx needs no unchecked copy.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The full path compare in release cost fireworks 1.8% (100 interleaved
runs on one core); the other motion effects were flat. A batched result
goes stale within an epoch only when another character's tick drops the
slot's mirror, and dropping it clears the slot's mirror bit, in the word
the batch has just read. Release checks that bit (one hot load, within
noise on fireworks); the full check stays a debug_assert.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
set_visual, set_visuals and layer_changed indexed the character tables
with the effect's slot unchecked, as did step_animation and
activate_scene (through doze_wake and the scene steps), and binarypath,
smoke, thunderstorm and errorcorrect with their own slot-indexed
tables. A bad slot there was undefined behaviour instead of a panic.

Each entry point now checks the effect's slot once, and the reads after
it stay unchecked with a note of the bound: every Chars column has
len() entries. set_visible and coordinate_changed already checked it
(flags, is_visible). The scene steps, which take their slots from the
active set, show visuals through an unchecked show_visual: checking
there cost colorshift 5% and beams and smoke 3%; with it, all three are
within noise. binarypath checks each representation's bit slots once
for its shift loop (every bit, every tick) and indexes a path's last
segment checked (a path without segments would wrap). Loop indices
bounded by their own range stay unchecked.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
LINES of -5, 0, 1 and 2 and COLUMNS of -3, 0, 1 and 2 shrink the
visible window to a row or a column, or to nothing: with the canvas
sized by the terminal, and with a 30x10 canvas larger than it from
three canvas anchors and three text anchors, in the output and in the
parity dump.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@dhh

dhh commented Sep 27, 2026

Copy link
Copy Markdown
Contributor

Thank you. This is now on master via #47, with your commits intact, and fx is the engine from ttfx 0.5 on. Those measurements held up here: on a quiet Zen 5, fx was 1.20× faster than the assembly engine on one core and 1.25× on two, with byte-identical output on all 37 effects. So we removed the assembly engine.

#47 also adds fixes from the review:

  • LINES=0 or negative made the visible-window test match every cell and write past the grid, which crashed the run. The test now uses clamped row and column counts.
  • The shared RNG no longer uses fx's unchecked indexing.
  • Color::from_key is now an unsafe, crate-private constructor.
  • The appearance hash is shifted before it's narrowed, so 32-bit builds behave the same.
  • The batched step is checked against its mirror bit in release builds.
  • Effect-supplied slots are checked at the engine's entry points.
  • TTFX_THREADS=0 also means single-threaded.
  • The oracle tools moved to tools/fx and compare fx with the original engine.
  • New oracle cases for tiny and negative terminals.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants