Skip to content

[tmva][sofie] Drop the BLAS dependency of the generated code - #23600

Open
guitargeek wants to merge 2 commits into
root-project:masterfrom
guitargeek:sofie-drop-blas
Open

guitargeek wants to merge 2 commits into
root-project:masterfrom
guitargeek:sofie-drop-blas

Conversation

@guitargeek

Copy link
Copy Markdown
Contributor

The inference code SOFIE emits no longer calls any BLAS routines.

Matrix multiplications are now done by Gemm_Ref, a self-contained reference implementation of the sgemm contract emitted into the generated header next to the other inference helpers, and bias accumulations by the corresponding Axpy_Ref (saxpy). Both keep the pointer-based BLAS calling convention, so the emission sites in the Gemm, Conv, ConvTranspose, Einsum, GRU, LSTM and RNN operators only change the function name. Gemm_Ref walks C column by column with unit stride so the compiler can auto-vectorize the innermost loop; this is close to BLAS performance for the per-event (gemv-shaped) inference SOFIE is used for in RDataFrame loops and RooFit, while removing

  • the link-time BLAS dependency of every generated model,
  • symbol collisions with whatever sgemm_ happens to be loaded in the process (e.g. a conda OpenBLAS in PyROOT sessions),
  • the implicit, environment-dependent multithreading of the BLAS implementation picked up at runtime, which interferes with RDataFrame implicit MT.

Consequently deleted: the BLAS namespace emission in RModel::GenerateHeaderInfo, fNeededBlasRoutines / AddBlasRoutines / GetBlasRoutines, the extern "C" sgemm_/sgemv_/saxpy_ declarations (dead Gemv), the commented-out BLAS call sites in Conv/ConvTranspose, and the hand-rolled CMake BLAS handling of the SOFIE tests (BLAS_FOUND guards, BLAS::BLAS link libraries, the hard configure error in SearchInstalledSoftware).

The hand-written Clad pullbacks/pushforwards for Gemm_Call are kept: they are pure C++ built on Gemm_Call itself, and are both more accurate and much cheaper than the tape Clad would generate for the inner loops.

Also remove the test_tmva_sofie configuration option: with the BLAS requirement gone the tests need nothing beyond what any other ROOT test needs, so they now follow the global 'testing' option like all other subsystems.

Verified: all SOFIE tests (TestSofieModels, TestCustomModelsFromONNX, TestCladAutodiff, TestGemmDerivative, TestSofieParser, ...), testRooONNXFunc and the SOFIE tutorials pass; a generated LSTM and ConvTranspose model header compiles and runs with plain g++ -std=c++17 with no ROOT and no BLAS.

Performance (Zen5-class 12-core machine, emitted code compiled -O2, vs OpenBLAS 0.3.33 single-threaded -- ROOT's BLAS dependency here):

Per-event inference (batch=1 GEMV shapes, the RDataFrame / RooFit SBI pattern) is at parity up to layer size ~512, both at kernel level (gemv 512x1x512: 3.5 vs 3.6 GFLOP/s) and end-to-end through Session::infer of a 4-layer MLP (D32: 2.1 vs 2.1 us, D128: 31 vs 27 us, D512: 464 vs 440 us per event). A regression sets in around layer size 1024 (weights ~4 MB): 3-5x kernel, 2.9x end-to-end (1899 vs 649 us).

The transposed-A case (which is what per-event inference of row-major weights emits) needs the dot-product loop nest in Gemm_Ref: with the plain accumulation nest the strided A reads make it 27x slower than BLAS at D=1024 instead of 2.9x.

Large batched evaluation regresses 20-100x end-to-end from D=128/batch=1000 up (D1024 B1000: 1.8 s vs 18 ms), matching the kernel gap for square GEMMs (~48x at 512^3, ~64-116x at 2048^3). Threaded OpenBLAS (12 threads) does not change the picture for typical SOFIE shapes: it is equal or slower than its single-threaded self below ~2048^3, and the shapes where it wins are exactly the ones SOFIE-emitted code does not target.

🤖 Done with the help of AI

@dpiparo dpiparo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, this PR arrived so quickly!!

@github-actions

github-actions Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Test Results

    24 files      24 suites   3d 18h 12m 23s ⏱️
 3 880 tests  3 876 ✅ 0 💤 4 ❌
79 535 runs  79 531 ✅ 0 💤 4 ❌

For more details on these failures, see this check.

Results for commit c938f8d.

♻️ This comment has been updated with latest results.

The SOFIE tests are currently disabled on Windows CI
(test_tmva_sofie=OFF in windows10.txt); enabling them surfaces two
Windows-only failures:

* RModelParser_ONNX::Parse split the model path at backslashes only,
  so a forward-slash path like dir/model.onnx (which the Windows file
  APIs happily accept) yielded an empty model directory, and the
  external weight data of such a model was looked up in the working
  directory instead of next to the model (TestSofieParser.
  ExternalDataLocationRelativeToModelDirectory). On Windows, split at
  the last of either separator; on POSIX a backslash stays a filename
  character.

* Catching an exception thrown by cling-JIT code does not work on
  Windows: the runtime finds no handler and terminates the process
  with 0xE06D7363, so the safetensors tests whose whole point is
  checking that the generated session constructors throw can only run
  elsewhere (TestSafetensorsWeights dies in MissingFileThrows, and all
  the other throw-probing tests would follow). Skip them on MSVC; the
  success-path construction (ModerateNestingAccepted) stays enabled.

🤖 Done with the help of AI
The inference code SOFIE emits no longer calls any BLAS routines.

Matrix multiplications are now done by Gemm_Ref, a self-contained
reference implementation of the sgemm contract emitted into the
generated header next to the other inference helpers, and bias
accumulations by the corresponding Axpy_Ref (saxpy). Both keep the
pointer-based BLAS calling convention, so the emission sites in the
Gemm, Conv, ConvTranspose, Einsum, GRU, LSTM and RNN operators only
change the function name. Gemm_Ref walks C column by column with unit
stride so the compiler can auto-vectorize the innermost loop; this is
close to BLAS performance for the per-event (gemv-shaped) inference
SOFIE is used for in RDataFrame loops and RooFit, while removing

- the link-time BLAS dependency of every generated model,
- symbol collisions with whatever sgemm_ happens to be loaded in the
  process (e.g. a conda OpenBLAS in PyROOT sessions),
- the implicit, environment-dependent multithreading of the BLAS
  implementation picked up at runtime, which interferes with RDataFrame
  implicit MT.

Consequently deleted: the BLAS namespace emission in
RModel::GenerateHeaderInfo, fNeededBlasRoutines / AddBlasRoutines /
GetBlasRoutines, the extern "C" sgemm_/sgemv_/saxpy_ declarations
(dead Gemv), the commented-out BLAS call sites in Conv/ConvTranspose,
and the hand-rolled CMake BLAS handling of the SOFIE tests
(BLAS_FOUND guards, BLAS::BLAS link libraries, the hard configure
error in SearchInstalledSoftware).

The hand-written Clad pullbacks/pushforwards for Gemm_Call are kept:
they are pure C++ built on Gemm_Call itself, and are both more
accurate and much cheaper than the tape Clad would generate for the
inner loops.

Also remove the test_tmva_sofie configuration option: with the BLAS
requirement gone the tests need nothing beyond what any other ROOT
test needs, so they now follow the global 'testing' option like all
other subsystems.

Verified: all SOFIE tests (TestSofieModels, TestCustomModelsFromONNX,
TestCladAutodiff, TestGemmDerivative, TestSofieParser, ...),
testRooONNXFunc and the SOFIE tutorials pass; a generated LSTM and
ConvTranspose model header compiles and runs with plain g++ -std=c++17
with no ROOT and no BLAS.

Performance (Zen5-class 12-core machine, emitted code compiled -O2,
vs OpenBLAS 0.3.33 single-threaded -- ROOT's BLAS dependency here):

Per-event inference (batch=1 GEMV shapes, the RDataFrame / RooFit SBI
pattern) is at parity up to layer size ~512, both at kernel level
(gemv 512x1x512: 3.5 vs 3.6 GFLOP/s) and end-to-end through
Session::infer of a 4-layer MLP (D32: 2.1 vs 2.1 us, D128: 31 vs
27 us, D512: 464 vs 440 us per event). A regression sets in around
layer size 1024 (weights ~4 MB): 3-5x kernel, 2.9x end-to-end
(1899 vs 649 us).

The transposed-A case (which is what per-event inference of row-major
weights emits) needs the dot-product loop nest in Gemm_Ref: with the
plain accumulation nest the strided A reads make it 27x slower than
BLAS at D=1024 instead of 2.9x.

Large *batched* evaluation regresses 20-100x end-to-end from
D=128/batch=1000 up (D1024 B1000: 1.8 s vs 18 ms), matching the
kernel gap for square GEMMs (~48x at 512^3, ~64-116x at 2048^3).
Threaded OpenBLAS (12 threads) does not change the picture for
typical SOFIE shapes: it is equal or slower than its single-threaded
self below ~2048^3, and the shapes where it wins are exactly the ones
SOFIE-emitted code does not target.

🤖 Done with the help of AI
@GiacomoXT

Copy link
Copy Markdown
Contributor

Sorry for jumping on this, but as you might have noticed (and discussed privately with @dpiparo), I am rather interested to SOFIE and I am keeping an eye on the development, since I still want to push upstream few improvements. I tested this on top of my current SOFIE branch with a Belle II GNN track-finding model (https://gitlab.desy.de/giacomo.depietro/catfinder-sofie) with dynamic input size (number of hits), single-threaded. Machine is an AMD EPYC 7452 (Zen 2, AVX2, no AVX-512), 512 KiB L2 per core, generated code compiled with gcc 11.5 -O3. Baseline is the same tree without this PR, linking OpenBLAS 0.3.29.

Numerically everything agrees — outputs match ONNX Runtime to ~7 significant digits before and after, no NaN. But end-to-end inference gets substantially slower:

┌──────┬──────────────┬──────────────────┬──────────────────┬──────────┐
│ hits │ ONNX Runtime │ SOFIE + OpenBLAS │ SOFIE + Gemm_Ref │ slowdown │
├──────┼──────────────┼──────────────────┼──────────────────┼──────────┤
│ 256  │ 23.1 ms      │ 22.1 ms          │ 45.4 ms          │ 2.1x     │
├──────┼──────────────┼──────────────────┼──────────────────┼──────────┤
│ 512  │ 54.7 ms      │ 53.7 ms          │ 103.3 ms         │ 1.9x     │
├──────┼──────────────┼──────────────────┼──────────────────┼──────────┤
│ 1024 │ 134.9 ms     │ 136.1 ms         │ 230.5 ms         │ 1.7x     │
├──────┼──────────────┼──────────────────┼──────────────────┼──────────┤
│ 2048 │ 292.0 ms     │ 351.1 ms         │ 532.5 ms         │ 1.5x     │
└──────┴──────────────┴──────────────────┴──────────────────┴──────────┘

With another model where GEMM is the dominating operator the slowdown is even larger. So SOFIE goes from being at parity with ONNX Runtime to being about 2x slower than it.

The model emits 42 Gemm_Call sites per inference, and 38 of them are transa='t', transb='n', so they all take the dot-product branch. In these calls m is the layer's output width, k its input width, and n is the batch dimension, i.e. the number of hits: A is the m×k weight matrix (transposed), B the k×n activations, C the m×n result. Here is Gemm_Ref against OpenBLAS on the shapes that carry the model, at n = 2048, single-threaded:

┌─────┬─────┬─────┬──────────────┬──────────────┬───────┐
│  ×  │  m  │  k  │   OpenBLAS   │   Gemm_Ref   │ ratio │
├─────┼─────┼─────┼──────────────┼──────────────┼───────┤
│ 4   │ 252 │ 252 │ 92.6 GFLOP/s │ 15.9 GFLOP/s │ 5.8x  │
├─────┼─────┼─────┼──────────────┼──────────────┼───────┤
│ 12  │ 126 │ 126 │ 82.4 GFLOP/s │ 13.4 GFLOP/s │ 6.1x  │
├─────┼─────┼─────┼──────────────┼──────────────┼───────┤
│ 3   │ 126 │ 504 │ 82.7 GFLOP/s │ 17.4 GFLOP/s │ 4.8x  │
├─────┼─────┼─────┼──────────────┼──────────────┼───────┤
│ 4   │ 252 │ 126 │ 91.3 GFLOP/s │ 13.4 GFLOP/s │ 6.8x  │
├─────┼─────┼─────┼──────────────┼──────────────┼───────┤
│ 1   │ 126 │ 14  │ 52.0 GFLOP/s │ 4.1 GFLOP/s  │ 12.6x │
└─────┴─────┴─────┴──────────────┴──────────────┴───────┘

Those 24 calls are about 90% of the model's GEMM work.

I understand the motivations for dropping the link-time BLAS dependency, the symbol collisions and the implicit threading. But batched inference with hundreds to thousands of rows is the normal case for us, not an edge case, and a 1.5–2x end-to-end regression there is hard to absorb. Would it be possible to keep an optional BLAS path, like emitting the sgemm_ call when a flag or generation option is set, falling back to Gemm_Ref otherwise, so that the default build gains the self-containment while large-batch users can still get BLAS performance?

@guitargeek

Copy link
Copy Markdown
Contributor Author

Thanks for following this! And thanks for the systematic study, it's exactly the kind of input we need!

Short answer first: yes, I'll add an opt-in BLAS path in this PR. The default generated code will be self-contained (no external symbols, as it is in this PR), and a generation option will make it emit the sgemm_ calls for users who link BLAS themselves. The BLAS path will carry lighter testing guarantees (a few GEMM-heavy models rather than the full suite, as we don't want to duplicate all tests).

Some context on why the change happens at all, since it explains where SOFIE in ROOT is heading. It currently serves two mandates:

  1. A friction-less, dependency-free ONNX inference runtime inside ROOT's analysis libraries (RooFit, RDataFrame), callable from Python and C++. Here SOFIE is an implementation detail of ROOT. The user shouldn't care whether inference is done by code generation plus JIT or something else. What matters is that it works in-process, without missing-library failures or symbol collisions. The generated C++ is internal, so its stability across releases doesn't matter much.
  2. Standalone ONNX-to-C++ code generation for reconstruction frameworks, like your use case. Here the generated code is the product: its interface stability matters, performance matters more, and framework users accept more setup cost in exchange. This is also the home for hardware acceleration and quantization, which is what the R&D fork (https://github.com/ml4ep/sofie) is for.

These mandates pull in opposite directions on performance, usability, and interface stability, and they require different interfaces to begin with. For the second usecase, it could even be a Python package with just an executable!

The direction we plan to follow is:

  • Base SOFIE stays in ROOT for mandate 1 with the minimal required interface.
  • We think mandate 2 is better served by a dedicated standalone tool (a sofie-emit/onnx2cpp executable) rather than a code generator buried in a library. We haven't found a maintainer for that tool yet, which is why ROOT covers mandate 2 on a best-effort basis for now. The suggestion to set up a plan for sharing developments and bug fixes between ROOT's SOFIE and the standalone R&D one is very much part of that conversation.

The Belle II response to our Q3 questionnaire said SOFIE would "likely soon" be used for offline reconstruction given the GNN runtime performance on CPU, which is why the PR has stayed unmerged rather than landing the regression. But if and when the standalone tool materializes, the long-term future of mandate-2 usage patterns in ROOT itself is likely to evolve. The Q2/Q3 meeting notes ([1], [2]) have the details, and your input there (as the one experiment with an actual GNN about to go into production) is very important.

What's your take on this? What part of ONNX tooling do you expect to be in ROOT, and what do you expect to be separate?

[1] https://indico.cern.ch/event/1699702/
[2] https://indico.cern.ch/event/1725596/

@guitargeek

Copy link
Copy Markdown
Contributor Author

By the way, have you tried https://github.com/onnx/onnx-mlir ? That would make a great addition to your comparison table! And maybe the fact that it's also a code generation tool makes it easy to integrate in your pipeline? I have always been curious what are the differences in performance and user experience between SOFIE and ONNX-MLIR

@GiacomoXT

Copy link
Copy Markdown
Contributor

I need more time to formulate proper answers to your previous questions. Regarding onnx-mlir, yes, I am aware of it. The engineer we work with uses it to deploy a simplified version of our models on an FPGA. I haven't had time to install and run it myself yet, but it's pretty high on my to-do list for adding another column to my benchmark results. :)

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants