Skip to content

benchmarks: the instrument changes for the 26.10.1 re-pin campaign, with a columnar insert in the bindings (#211) - #212

Merged
tae898 merged 116 commits into
mainfrom
repin-prep-2
Oct 6, 2026
Merged

tae898 merged 116 commits into
mainfrom
repin-prep-2

Conversation

@tae898

@tae898 tae898 commented Oct 6, 2026

Copy link
Copy Markdown

Fixes #211.

DO NOT MERGE while the old chain is running. Every benchmark stage starts with git pull --ff-only, so this change would reach the stage that is measuring and split its rows. The order is: cut the old chain, merge this, place the official pair on the benchmark machine, regenerate and lint the stages, launch the first stage. This is a draft so that CI runs while that is decided.

What it is: the instrument changes for the re-pin campaign on the official 26.10.1 release, 105 commits. 56 of the 70 checklist rows of benchmarks/experiments/CAMPAIGN.md section 7 are implemented and tested here; the other 12 are measurements on the benchmark machine or open items between campaigns, and none of them is needed before the first stage. It also holds the official pair (the image and the wheel the stages bake, with verify_pair_c25.sh, which prints PAIR VERIFIED d36b4ca3a, JVM 25), a queue linter that knows the wheel, regenerated stages, and insert_columns in the bindings (#150, with 19 new tests that all fail on the old wheel).

Found by running the new questions against the real engines and a DuckDB reference on LDBC data, which a test that only reads the query texts could not have found: SurrealQL integer division in the undirected mean age, SurrealQL closures that hide a query parameter and an outer variable (the 3-hop exclusions removed nothing, off by one for 26 of 96 starts), FalkorDB and LadybugDB reusing a relationship in a chain (the 2-hop and 3-hop reads counted the start), and a MongoDB $lookup localField that made every interactive Mongo graph cell fail. Each is fixed here with a test. Analytics agree on 12 of 12 engines and LSQB on seven; interactive reads agree on 11 of 12 (PostgreSQL with AGE: its 3-hop read is censored at the budget, so unverified).

Checked: 400 benchmark tests including the real-engine cases, the generated stages lint clean (queue_lint 15 scripts, 0 problems) and pass bash -n, the bindings suite with the new tests, bandit, the docs build, and the wheel-size gate (72.8 MB). The documents, time-series, and sparse rows ran through the real runner on the official pair at micro scale, with answers equal to DuckDB's and SQLite's. Nothing was run at campaign scale; timings from the development machine are mechanics, never a measurement.

Not in this change: the SurrealDB served interactive cell with the final 3-hop text was not run through the runner (its text was checked directly against the pinned 3.2.4 server for every start of a small graph), LSQB q6 and q9 stay excluded at sf1full until a first-touch probe on the benchmark machine, and no comparator pin moves.

🤖 Generated with Claude Code

https://claude.ai/code/session_01Ubz3m6QXHESerCkehC1qgi

tae898 and others added 30 commits October 2, 2026 14:30
…ugh the bulk path (ArcadeData#8287; BUGS F123)

Embedded: graph_batch(use_wal=True, expected_edge_count) for persons+KNOWS, graph_batch(use_wal=True)
for the message half with endpoints from returned RIDs. Served: POST /api/v1/batch?wal=true
(&expectedEdgeCount for KNOWS; idMapping=true so the message request references Persons by RID),
streamed JSONL. Laptop micro smoke on 5b040a4: all 20 answer digests/counts identical to the
current loaders; build 1.75 -> 1.03 s embedded, 7.38 -> 1.21 s served. Message half not smoked
(needs the LDBC corpus): smoke on mini after ALL-DONE.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8NdUMbUoCLwNJEWPir4YL
…tch as INSERT ... CONTENT :rows (ArcadeData#8337, DECISIONS #116 item 4)

Replaces a sqlscript of 2,000 INSERT ... SET statements with the values in the
text. The batch size is BENCH_SERVED_LOAD_BATCH (default 2000), recorded in the
row as served_load_batch, for the 2k/5k/10k sweep on the bench host.

Laptop smoke, SF0.1, served arm on 7effa09, main's loader against this one:
600,572 line items both, all five answer digests identical, ingest 52.1 s -> 43.4 s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8NdUMbUoCLwNJEWPir4YL
…ds INSERT ... CONTENT :rows; ingest sentences follow (ArcadeData#8337, DECISIONS #116 item 4)

Replaces 500-statement sqlscript batches with each vector spelled out as text.
From this Python harness CONTENT is the fastest served vector path (laptop, 50k x
96, batch 2,000: sqlscript literals 2,754 rows/s, psycopg over the Postgres wire
3,990, CONTENT 4,980; every path stores the float32 values exactly), so it is used
instead of the Postgres wire ArcadeData#8337 found fastest from Java. Batch size is
ArcadeServer.load_batch (BENCH_SERVED_LOAD_BATCH, default 2000).

The October ingest sentences for l3d and the documents tables now describe the
bound batches, pinned to the load_batch attributes; SERVER_BATCH is retired.
Smoke, micro, served arm on 7effa09: recall@10 1.0 both, ingest 2.71 -> 1.36 s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8NdUMbUoCLwNJEWPir4YL
… query vector, index name, k and ef (DECISIONS #116 item 2)

As the embedded arm already does. Laptop probe, 20k SIFT, 300 queries: identical
top-10 on all 300, p50 5.40 -> 4.37 ms (repros/vector-query/served_bound_vector_probe.py).
Micro smoke with the index name bound too: recall@10 1.0.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8NdUMbUoCLwNJEWPir4YL
…LTP statements bind their values (DECISIONS #116 item 2)

ArcadeServerTPC: new-order and payment stay one sqlscript request each, with
named parameters shared across the script's statements; the four CRUD
operations bind :c/:pk. SurrealTPC (and the served twin, which inherits it):
$vars with RecordID objects, so record-id addressing is kept and nothing is
written into the SurrealQL text.

Laptop smoke at SF0.01 (BENCH_CRUD_OPS=500), main against this branch:
answer digests identical for arcadedb_server and surrealdb_tpc. Timing is
sub-ms on the laptop and not reliable there; mini measures at the re-pin.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8NdUMbUoCLwNJEWPir4YL
…eir values on every engine that can (DECISIONS #116 item 2; BUGS F125, F130)

The shared Cypher template takes $id, $new_id and $name, passed through each
engine's own driver: ArcadeDB embedded (query/command with a params map) and
served (HTTP params), Neo4j and Memgraph (session.run params), FalkorDB
(query params), LadybugDB (prepared once per text per connection, then
executed with params). SurrealDB binds $vars with RecordID objects, and its
write is now one BEGIN/COMMIT transaction: it was two statements in one
request, which SurrealDB runs as two transactions (F130). DuckPGQ keeps its
values in the text and says so: it rejects any parameter in a GRAPH_TABLE
query (cwida/duckpgq-extension#75); its writes and scan already bind.

Laptop micro smoke, 9 engines, October head against this branch: all 8
answer digests identical on every engine. Timings are laptop numbers, for
direction only: Neo4j point 30.2 -> 7.2 ms, ArcadeDB embedded 2.17 -> 0.32,
served 5.9 -> 3.5, Memgraph/FalkorDB/LadybugDB about level, SurrealDB served
write 14.3 -> 7.6 (one transaction). SurrealDB embedded hop1 read 5.1 -> 9.8
under CPU contention from a probe; being checked with an in-process A/B.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8NdUMbUoCLwNJEWPir4YL
…g timed values; one set-based e2 update; served sparse and mutate loads through CONTENT (DECISIONS #116 items 2 and 4; BUGS F130, F131)

e2: both ArcadeDB arms bind pid and the id lists (`pid IN :ids` still reads
the unique index, EXPLAIN: FETCH FROM INDEX) and send ONE set-based update
instead of one statement per product (F131; served: one request instead of up
to six); SurrealDB binds $q and RecordID lists and updates with one
`UPDATE $ids`; PG+AGE's single-pid read binds through cypher()'s third
argument. l3d: the deletes bind (ArcadeDB `IN :ids`, DuckDB `IN (SELECT
unnest(?))`), Milvus deletes by ids=, SurrealDB binds its query vector and
deletes one `DELETE $ids` per batch (it was 200 statements, 200 transactions:
F130); LanceDB's predicate-only API is declared. The served mutate insert and
the served sparse build load through `INSERT ... CONTENT :rows`. Page: the
graph ingest sentence now describes item 1's bulk path, the sparse one the
bound batches, each pinned.

Laptop smokes, before/after, answer digests identical everywhere they exist:
ArcadeDB served arms on the upstream-main pair (7effa09; a first pass
without the image export ran 26.8.1 and is discarded): graph point 5.75 ->
3.30 ms, e2 transaction 26.7 -> 14.7 ms, dense recall 1.0 with mutate insert
0.655 -> 0.265 ms/vector, sparse recall 1.0. Comparators: every e2 digest and
atomicity result unchanged; dense recall and mutate results unchanged. One
regression, fixed in the next commit: PG+AGE loses its pid expression index
for a bound LIST.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8NdUMbUoCLwNJEWPir4YL
… loses the expression index (2.0 -> 26.4 ms), the single pid stays bound (DECISIONS #116 item 2)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8NdUMbUoCLwNJEWPir4YL
…candidates, on top of main's F132 fix (rebase onto #117)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8NdUMbUoCLwNJEWPir4YL
…rough graph_batch with the WAL on, one Java float[] per vector (DECISIONS #116 item 3)

Replaces a SQL INSERT per vector in 10,000-row transactions with the engine's
bulk loader, WAL kept on as the maintainers recommend (ArcadeData#8287), fed 10,000
vectors per commit so the live Java arrays stay bounded at deep10m. The
October ingest sentence says so, pinned to BATCH.

Laptop, pinned pair 7effa09, the tier's 20,000 SIFT vectors: ingest 1.84 ->
1.46 s fp32 and 1.82 -> 1.49 s int8, recall 0.9996 -> 0.9998 and 0.9953 ->
0.9951 (HNSW build noise), query p50 level. The 1.4x of the 2026-09-24 audit
was measured at 200k; mini measures it at the re-pin.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8NdUMbUoCLwNJEWPir4YL
…the peak memory for JVM engines (DECISIONS #116 item 7)

A `JVM heap` column right after `peak memory GiB` on every table that has
one, from the rows behind each cell at its own size: ArcadeDB, Neo4j (and
the composed stack's Neo4j half), Elasticsearch, and QuestDB record a heap;
every other engine prints a dash. A JVM engine's peak sits close to the heap
it was given, so the decision keeps the peak (what an operator provisions)
and puts the setting next to it; no new instrumentation. The lookup is shared
with the memory note (_entry_heaps), so the two cannot name different heaps.

Dry landing from this branch (e2 scoped, October rows): every gate passes;
l2 shows 4g/12g beside ArcadeDB's and Neo4j's peaks, e2 6g/16g.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8NdUMbUoCLwNJEWPir4YL
…d loads through /batch with the WAL on (DECISIONS #116 items 1 and 4)

Products and RELATED edges stream as JSONL to POST /api/v1/batch
(?wal=true&expectedEdgeCount=N), each embedding as the float32 values'
exact doubles, as the graph lane's served arm loads (ArcadeData#8287); the counts the
server returns are checked. It replaces CREATE VERTEX / CREATE EDGE ... FROM
(SELECT ...) statements with every value pasted into sqlscript, where the
embeddings crossed at six significant digits. The October e2 ingest sentence
says so; September's keeps what September ran.

Laptop, 7effa09 pair, 50k products: build 40.5 -> 17.4 s (ingest 28.7 ->
9.2 s), all three answer digests equal the embedded arm's, retrieval recall
0.785 -> 0.7855, atomicity 0 torn of 40 unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8NdUMbUoCLwNJEWPir4YL
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8NdUMbUoCLwNJEWPir4YL
…nt takes the query vector as a Java float[] (BUGS F145)

_vec_topk and hybrid_op already pass it that way; _rank_candidates passed a
Python list, which crosses JPype one element at a time. Laptop, 50,000
products, 39 candidates: 1.42 -> 1.00 ms p50 with the October statement,
0.80 ms with the ids bound too; the same top ten on 300 of 300 queries.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QHcrtLcfHdr5UhbCdnpM77
…hat filters, a per-read budget, and a digest that survives six-digit ties (BUGS F146, DECISIONS #119, #120)

Ages: ldbc_snb read the corpus's epoch-millisecond birthday as a date
string, so every age was 0 and hop3f's x.age > 30 matched nobody. Parsed
as milliseconds (ISO kept as the fallback), LDBC ages span 36-46, so the
threshold moves to 41 (keeps 49.7% at SF1, 49.3% at SF10), one constant,
graph_common.HOP3F_MIN_AGE, read by all five spellings.

Per-read budget (#120): each transactional read gets one budget across
both passes (budget_lookup, lane default OLTP_READ_BUDGET_S = 1,800 s,
clamped to 540 s at SF1). A read that spends it stops, keeps percentiles
over the starts it answered, records <read>_censored/_iters/_budget_s,
and declares its answer with bench_common.record_censored_answer; the
gate lists it under E4b, the page says it is not compared (never
'cannot express') and names it in the per-engine budget sentence, now
worded right for a single cell. Found because SurrealDB embedded's 2.3.10
core dedups quadratically and its SF1 cell timed out whole.

Digest: floats round twice (12 significant digits, then 6), because exact
six-digit ties (41.15625, a mean of 32 ages) split Neo4j's running mean
from eight engines on hop1. Tie and declaration tests added.

Laptop smoke, SF1, ten engines at this tree: 0 equivalence failures, 8
groups agreeing (hop3f on real, non-zero answers), SurrealDB embedded's
hop3f censored after 111 of 500 starts with its other reads and writes
measured, the cell done in 996 s. PAGE-SPEC and the fairness comment
updated.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QHcrtLcfHdr5UhbCdnpM77
… settle until the engine reports no uncompacted sample (DECISIONS #121)

settle() polls SELECT FROM schema:types until every shard's mutableSamples
is 0 (bounded, BENCH_TS_COMPACT_WAIT_S, default 300 s), outside the ingest
timer, as QuestDB's arm waits for its WAL; the driver records
ts_mutable_at_query and ts_compaction_wait_s at the first timed query, so
a row proves its compaction state. Why: ArcadeData#8574, the
newest reading is 21 ms on the mutable tail and 0.13 ms compacted.

Also removes an asymmetry: the embedded native arm slept the lane's
BENCH_TS_SETTLE_S itself and the driver then slept it again (180 s at
October's 90 s) while the served arm's settle was empty. The symmetric
floor stays the driver's. The shard list is read as a map or its JSON
text (the October wheel's to_json_list renders a nested result as a
string); anything unreadable counts as -1, never 0.

Laptop smoke at ts100, BENCH_TS_SETTLE_S=0: served arm waited 55.5 s,
embedded 5.1 s, both ts_mutable_at_query 0.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QHcrtLcfHdr5UhbCdnpM77
…tart is recorded, a raising cell no longer ends its batch (F150); the served sparse rows are the generator's lists (F151)

F150: run_cell's finally reads _therm0, which is taken only once the client
starts, so since dea7115 every early return before that (server_not_ready,
the server-heap mismatch) raised UnboundLocalError instead of returning its
row. And worker() had no except: the exception killed the thread, the cells
still pending on it never ran, and with no error row the runner exited 0.
Now _therm0 is initialised beside cli_cid, and a raising cell is printed as
RAISED <cell>, the batch continues, and the exit is 1. Not fired on mini in
October (no crash in STATUS.txt). Laptop: the JDK 21 served cell that
crashed now records server_not_ready and exits 1.

F151: item 4's served sparse loader passed every weight through
np.asarray(..., float32) and back although both corpus sources already yield
Python lists (rows compare equal without it): 5.30 s of client CPU at 100k
on BigANN, 0.08 s without. Laptop, BigANN 100k, c44f660 pair: served
build 56.3 / 62.0 s against October's loader 63.2 / 70.0 s (embedded
23.6 / 26.4 s); the re-pin loader stores every sampled weight exactly
(October's 9-decimal text does not, 80 of 101 records off by up to 3.4e-5).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QHcrtLcfHdr5UhbCdnpM77
…s bulk path for LineItem, measured, with a bucket-count control (CAMPAIGN item 10, ingest)

The answer on ArcadeData#8478: async createRecord on a type with
buckets = k x the executor's parallel level (k 1 or 2), set at CREATE TYPE.
Applied to the embedded documents loader: LineItem created with BUCKETS =
writers x BENCH_ARCADE_BUCKETS_PER_WRITER (default 1), loaded through
insert_many(parallel=True) with writers x commitEvery rows per call
(waitCompletion() commits every writer's open batch, so smaller calls mean
smaller commits), the executor's WAL flush set per class (its writers stamp
their own and ignore txWalFlush: yes_full at strict, via
bench_common.arcade_async_sync), the stored count checked after the load (an
older wheel drops rejected records silently, F153), and the row records
lineitem_buckets, async_writers, async_sync and load_call_rows.
BENCH_ARCADE_LINEITEM_BUCKETS overrides the count for control runs; both
variables are on the runner's container allowlist.

Laptop, SF0.1, cores 0-3 (3 writers), engine c44f660: ingest 1.6x faster
at both classes, every answer digest identical. But a control on this same
harness with only the bucket count varied puts the five analytics queries
17-47% slower at 3 buckets than at 1. That contradicts the answer's "no
longer costs on scans", so this is prepared, not adopted: the question goes
back upstream with a Java repro before the re-pin decides.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QHcrtLcfHdr5UhbCdnpM77
…ables (DECISIONS #128, #131)

A pgage_graph adapter on dbbench:pg-age: the lane's own Cypher through
AGE's cypher(), values bound as one agtype map, rows parsed back from
agtype after the timed call. Loaded by COPY into AGE's label tables with
the graphids AGE's sequences would assign (20,000 persons in 0.12 s and
418,599 KNOWS in 2.9 s on the laptop); Message is a PostgreSQL
inheritance parent of Post and Comment, so LSQB's (:Message) finds both.
One added index, a btree on the Person id property, by measurement
(#112; point 58x, hop1 16x, hop2 3.2x, hop3f 2.1x, same answers).

Smoked on the laptop: micro oltp and olap, the LDBC SF1 slice, and the
capped sf1full slice; every digest matches Neo4j's and LadybugDB's except
the two CRUD read-backs at micro, where AGE's btree scan starting at a
>= bound skips the equal key (apache/age#2587, filed today, every AGE
release since 1.6.0). equivalence_check declares it as KNOWN.

Also: Dockerfile.pgage pins postgresql-18-age=1.8.0~rc0-2.pgdg12+1 (the
build the cross-model rows carry; CAMPAIGN section 7 row 26), and both
laptop smoke scripts run their LSQB stage at sf1full: at sf1 the message
half never loads, so no smoke since 2026-09-19 asked LSQB's nine.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr
…pools fitted to the cell like every other engine's (FAIRNESS F3/F6)

The user: "make its resources usage (cpu, memory, heap, etc) fair as we
do the other dbs". Both AGE arms (pgage_graph, pg_age_e2) now get
parallel workers per query = cpuset - 1 (the leader is the last core),
max_parallel_workers = cpuset, and work_mem = (cap x 0.75) / (cpuset x
16), beside the cap-sized shared_buffers, effective_cache_size and
maintenance_work_mem they already had. PostgreSQL's fixed defaults (2
workers, 4 MB) reached 3 cores of 12 and spilled instead of using the
cap, where every other engine's pools reach the cell.

Measured on the full LDBC SF1 network (3.16M message vertices, 13.6M
edges), 8 cores and a 16 GiB cap, repros/age-dialect/
resource_fit_probe.py: workers fitted made LSQB q2 1.8x faster and q4,
q5, q7 1.1-1.25x; the fitted 96 MB work_mem cut temp-file spill from
42 GB to 0.7 GB per analytics pass at no net time cost; 384 MB bought
nothing more; no OOM kill; answers identical across all four settings.
The graph adapter reads every setting back onto the row (pg_*), the
cross-model rows carry them in server_cmd. Smoked through the runner:
l2 micro oltp+olap (digests unchanged) and e2 hybrid.

FAIRNESS F6 gains the PG+AGE row; COMPARATORS names the graph arm and
the pinned package.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr
…oin the dense tables; DuckDB VSS and ArangoDB split ingest from index (DECISIONS #131 item 3, #132)

Five new arms, each at the matched operating point in its own unit and
with its pools fitted to the cell and read back onto the row:
elasticsearch_dense (hnsw) and elasticsearch_dense_int8 (int8_hnsw, top
k re-scored), on the sparse arm's 9.5.4 image and heap; memgraph_dense
(3.13.1, the graph arm's image and flags; its USearch config is fixed at
connectivity 16 / expansion_add 128 and any key we pass is ignored);
falkordb_dense (6.0.1, since 4.x crashes on a write in a vector
statement; the background-built index is waited for to OPERATIONAL,
recall 0.05 without the wait); ladybug_dense (the official vector
extension, ml = 2M base degree, COPY from Arrow). DuckDB VSS and ArangoDB
now time their load and their index separately, so
PHASE_SPLIT_DISCLOSED is empty and the two Elasticsearch arms are
declared (F14c). l3d_dense and dense_multipass_driver record each
adapter's row_extra.

Smoked through the runner at micro (mutation forced on) and tiny: recall
0.98-0.9997 beside Qdrant's 0.9999, deleted hits 0, F14c ok, base degree
32 everywhere. Two engine defects found by the mutation phase: LadybugDB
0.20.4 crashes when searching after re-inserting deleted keys (0.21.1,
the re-pin's, is clean); Memgraph keeps committed deletes in its vector
index until GC (memgraph/memgraph#4975), so its delete runs FREE MEMORY
inside the timer. NOTES-dense-arms-20261002.md has each arm's status and
the decisions left.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr
…ONS #131 item 2, CAMPAIGN rows 27 and 28)

hop2 seeds $graphLookup at the person with maxDepth 1 and keeps the depth-1 edges; hop3f and the three-hop visited probe seed it at each first edge's end and keep the depth-1 edges that are not the first edge. Both are exact when the graph has no self-loops, which build() now counts onto the row (mongodb_knows_self_loops) and refuses. A probe found 0 disagreements with the chained $lookup form over 200 ids on micro and the LDBC SF1 slice; seeding three hops at the person disagreed on 92 and 12 of 200. hop1, the writes, and the analytics stay $lookup, and QUERY_LANGUAGE says so.

Found by the capped SF1 slice smoke: run_write inserted the new person and edge even when the anchor person was not loaded, where the Cypher engines' MATCH creates nothing, so the write and update digests differed. It now looks the anchor up first, in the same transaction. main() also reads row_extra after the build.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr
…ONS #131 item 4, CAMPAIGN row 39)

Each runs vector top-k, the hop, and the update in one transaction, and the atomicity probe by rollback, like the existing e2 arms. memgraph_e2: served on the graph arm's image and pools, the vector index at its defaults (Memgraph takes no HNSW parameters; disclosed), search breadth 100 then keep k (asking for k gave recall 0.584). ladybug_e2: embedded, the official vector extension, HNSW matched to the table (mu 16, ml 32, efc 100, efs 100), threads and buffer pool fitted (F160), rollback on any exception. duckdb_e2: tested and added, vss HNSW (M 16, ef_construction 100, ef_search 100) under the experimental persistence flag, a DuckPGQ hop with the pid list inlined (DuckPGQ binds no parameter), threads fitted from the affinity mask and memory_limit from the cgroup. Durability is read back for all three. e2_hybrid.main() records each adapter's row_extra.

Laptop smoke at e2 scale: all four graph-capable arms (these three and pg_age_e2) agree on every digest; atomicity 0 torn of 40 for each. Registered in runner (BACKENDS, LANES e2), export_web display names, page_check (optional until measured), and fairness_check (e2 index map; LadybugDB and DuckDB strict-only). COMPARATORS and FAIRNESS F6 rows name the arms.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr
 item 9, CAMPAIGN row 43)

MULTIMODEL_ENGINES gains "PostgreSQL + AGE". MULTIMODEL_ALIASES folds the cross-model arm's name ("PostgreSQL + pgvector + AGE") into that row, because both arms run the dbbench:pg-age image; PostgreSQL's documents, pgvector, and TimescaleDB arms run other images and stay their own engines. page_check's roster uses the same fold. PAGE-SPEC says five engines.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr
…reSQL pool fit on every arm but the defaults arm, the quantization survey (DECISIONS #131 items 6 and 8, FAIRNESS F6)

postgres_ts: COPY into a plain table, the (host, ts DESC) btree TimescaleDB
carries (measured at ts100: q_last 108.95 -> 0.48 ms, q_range 92.91 ->
1.08 ms; the 12 h aggregate pays 136 -> 217 ms, declared), date_trunc
buckets in UTC; smoke digests equal TimescaleDB's and SQLite's on all six.

runner.PG_FIT_CMD factors out the AGE fit and applies it to
postgres_tuned (its constant 64MB work_mem removed), timescaledb,
pgvector_dense, pgvector_sparse, and postgres_ts; not to `postgres`, the
defaults arm. BENCH_PG_FIT=off formats PostgreSQL's defaults for a
same-run A/B: answers identical in every pair, no OOM kill (FAIRNESS F6
row). The time-series PostgreSQL arms read every setting back (pg_*).

QUANTIZATION.md: what each dense and sparse comparator ships against our
DDL. Neo4j 2026.08.1 defaults its vector index to BINARY, so the dense and
cross-model Neo4j arms record fp32 while running binary (verified on the
pinned image); Elasticsearch 9.5.4 defaults to int8_hnsw at our dims;
LanceDB 0.39.0 does build an unquantized IVF_HNSW_FLAT.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr
…DECISIONS #131 item 5, queued last by #133)

l5_lifecycle_sql.py runs the lane's situations, sizes, session, and mode
set on SQLite and DuckDB in-process, streaming the build in batches
(F161); what a model lacks is declared with the engine's own error.
DuckDB builds vector (VSS HNSW) and graph (DuckPGQ property graph, read
with GRAPH_TABLE); SQLite builds the document and time-series situations.
Read digests equal ArcadeDB's and SurrealDB's at lc10k (doc, doc_idx10,
ts, graph), and the lifecycle table builds with every declaration noted.
l5_lifecycle.main now routes SurrealDB before the generic _server branch,
which would have sent a served SurrealDB arm to ArcadeDB's served module.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr
…pdate and delete are no-ops when the person is absent

The graph and cross-model fork found MongoDB writing a person whose
anchor the capped LDBC slice had not loaded, where every Cypher
engine's MATCH creates nothing; ArangoDB's write had the same
unconditional INSERTs. It now reads the anchor with DOCUMENT() and
FILTERs, still one AQL query. That made the update and delete reach an
absent person, where `UPDATE {_key}` and `REMOVE {_key}` raise "document
not found" and Cypher does nothing, so both now FILTER on _key (the
primary index) instead. Smoked at micro and on the capped SF1 slice:
all eight digests equal LadybugDB's on both.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr
…arm, Memgraph's delete disclosed, SurrealDB served off lifecycle (DECISIONS #135)

Neo4j 2026.08.1 builds a vector index BINARY-quantized when the
definition names none, and both Neo4j arms named none while recording
fp32 (BUGS F164). Both definitions now set `vector.quantization.type`
(NONE on the fp32 arms, SCALAR on the new neo4j_dense_int8), read the
applied index configuration back onto the row (neo4j_vector_quantization,
neo4j_vector_index_config), and refuse the cell on a mismatch. The e2
lane now merges row_extra after build() as well as at connect, since the
read-back exists only once the index does. Smoked on the laptop: dense
micro NONE recall 0.9986 and SCALAR 0.991, e2 hybrid NONE.

The dense table says under it that Memgraph's timed delete includes the
FREE MEMORY it needs (its vector index keeps deleted points until GC);
COMPARATORS lists SurrealDB served under "Not added, and why" (its server
has no open or close of a database to time) and records the Neo4j
quantization history.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr
…k (F165); LadybugDB 0.21.2, pinned harness libraries, FalkorDB graph arm at 6.0.1 (#129)

SurrealDB 3.2.4 served takes its sync mode on the storage path; the October
campaign ran its default, a sync at every commit, in both durability classes
under a "not verified" label (BUGS F165, DECISIONS #136). The five served arms
start with ?sync=never, the strict class rewrites it to ?sync=every, the runner
reads "Sync mode:" back from the startup log before the client starts and
refuses the cell on a mismatch (surreal_sync_mode), bench_common moves the arm
from NO_DURABILITY_SETTING to STRICT_OF, and UNVERIFIED_ALLOWED is empty.
Laptop smoke through the runner, TPC OLTP micro: relaxed reads back never,
strict every transaction commit, reads unchanged between the two.

build_images.sh pins numpy, pandas, pyarrow, requests, and psycopg (bare in
every image; dbbench:duckdb had drifted to pandas 3.0.5) and moves ladybug to
0.21.2 (0.20.4 crashes the dense mutation phase). The FalkorDB graph arm moves
to the dense arm's 6.0.1 digest. Smoke at micro: graph OLTP in both classes,
graph analytics, and LadybugDB dense with the mutation phase, 10 of 10 ok,
13 answer groups agreeing with Neo4j. CAMPAIGN section 7 rows 47 and 48.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr
…y for green

The Qdrant sparse arm sends its batches with wait=False and only the last with wait=True, then waited for a green status. On the laptop neither meant the collection held the corpus: at the Big-ANN 100k slice the exact count at the first green was 66,064 of 100,000 (81,064 on the uint8 arm), and the lane searched it, recall@10 0.93 against exact ground truth (brute force over the same files agrees with the ground-truth file, 1.0). The same ingest with wait=True on every batch returns 1.0.

post_build now polls the exact count until it equals what was sent, with green, inside the build timer, bounded by BENCH_QDRANT_SETTLE_S (3600 s; a short collection refuses the cell). The row records qdrant_points_at_first_green and qdrant_settle_after_green_s. After the fix the arm returns 1.0 at that slice. The sparse lane and its multipass driver now merge an adapter's row_extra onto the row, as the dense lane does.

The mini rows at the October pin record 1.0 at 100k and 1M and 0.9998 to 0.9999 at 8.84M, so the race cost no recall there that a row shows; whether their build_s stopped before the last points were applied is not on any row.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WWcDr7TwvcfMsQkeYQVHZr
tae898 and others added 26 commits October 6, 2026 13:18
…row 57, DECISIONS #151 item 3, BUGS F170)

TPC-H Q1 returns ten columns. The October lane's Q1 returned seven on every engine: l_tax was never loaded, so it
could not compute sum_charge, and it left out avg_price and avg_disc too, while the page called it TPC-H's own Q1.

l_tax joins LI_COLS (last, so every positional table keeps its order) and every engine's load (SQLite and PostgreSQL
DDL, both ArcadeDB inserts and the served DDL, the decimal-to-double cast; MongoDB, SurrealDB, and ArangoDB take
the column from LI_COLS). Every Q1 text computes sum_charge = sum(price * (1 - discount) * (1 + tax)), avg_price,
and avg_disc beside the five it had (DuckDB, SQLite, PostgreSQL, ArcadeDB, SurrealQL, AQL, the MongoDB pipeline);
OLAP_DIGEST['q1'] declares all ten columns with every measure compared as a number; every documents analytics row
records tpch_q1 = full, which export_web._partial_q1 retires from. Q1's digests and timings are re-baselined.

test_tpch_q1_full.py: the load and every text, the digest declaration, the stamp, the real SQLite and DuckDB
adapters against the specification in pandas on twelve line items and against each other (one digest), and both
against DuckDB's bundled official SF0.01 answer on all ten columns. The DuckDB adapter also equals the official SF1
answer on 32 of 32 cells (laptop, ad hoc, recorded in REPIN-REHEARSAL).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ubz3m6QXHESerCkehC1qgi
…warm-up for every engine (CAMPAIGN row 65, DECISIONS #157)

New-order, payment, and the four single-record operations were timed from the first operation after the load, each
leaving out only its first 20 from the percentiles, so on the JVM engines the statements compiled inside the timed
window: on the October rows ArcadeDB embedded's new-order is 2.34x (SF1) and 2.56x (SF10) its own payment, timed right
after it, against 0.94x to 1.30x for every engine not on a JVM; on the laptop the window is 2.90x to 3.79x the warm
median and an untimed warm-up of 2,000 operations on disjoint keys cuts it to 0.38x.

Every engine now runs OLTP_WARMUP (2,000, the count measured on the laptop; the bench host may choose another through
BENCH_DOCS_OLTP_WARMUP) untimed operations of each kind before the timed ones: new-orders and payments on order keys
0..W-1, the four single-record operations on keys 0..W-1 (deleting their own rows), the timed keys starting at W so the
timed inserts still land at the end of the order index as before; part keys come from the same distribution on their
own seeded stream, so the timed stream is the one it was. The count is the same in both durability classes. The first
operation of the session is timed and kept as the cold column; the row records oltp_warmup and oltp_warmup_s, and
no longer stamps cold_warm_na (already warm by construction is not true of this lane), which
export_web._docs_warmup_note retires from. Not e2: its transaction runs after 400 timed reads and its window was not
measured (the row says so).

test_docs_oltp_warmup.py runs the whole lane in-process on the real SQLite adapter over a tiny TPC-H-shaped parquet:
the keys the warm-up wrote and the timed ones, the tables left as the timed phases expect, the cold number, the
digests, and the old protocol when the count is 0.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ubz3m6QXHESerCkehC1qgi
…t to SQLite's FULL (CAMPAIGN row 5, the user's decision of 2026-10-04)

txWalFlush=2 is FileChannel.force(true): the log's data and metadata at every commit, one fsync. SQLite's synchronous=FULL
in WAL mode is one fdatasync, which is txWalFlush=1 (force(false)); upstream's transactions documentation (docs commit
c52695ba66) calls 1 'Safe against power loss. Recommended for production' and says 2 has 'No additional recovery value
over 1' with no measurable performance difference, and the server's production mode defaults to 1. So 1 is the setting
matched by effect to SQLite's, not the flattering choice the knob-free-for-them rule guards against. Before this change
every strict-class cell of every ArcadeDB arm would have been measured at the setting the decision replaced, which the
page would then have had to withdraw: all the strict cells of documents OLTP, graph OLTP, and the cross-model transaction.

bench_common.ARCADE_STRICT_TX_WAL_FLUSH = 1 is the one place the value lives: the embedded JVM argument, the async
executor's own spelling (yes_nometadata; its writers ignore txWalFlush and stamp their own flush, so the documents load
is told separately), the served container's JAVA_OPTS (runner.durability_server_patch, recorded as
durability_server_flags), the read-back from GlobalConfiguration.TX_WAL_FLUSH, and the strict string with its
classification. The October string is kept (DURABILITY_ARCADEDB_STRICT_FULL_SYNC) for the rows measured at 2 and still
classifies as strict; export_web's sentence that names the change stays keyed on that text, so it goes when the last such
row does. PROTOCOL.md section 7 gains the row. Smoked on the laptop with the official 26.10.1 wheel: the JVM argument,
the engine's read-back, and the executor all say 1.

test_strict_durability.py holds every route and both strings; test_overrides follows the served flag.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ubz3m6QXHESerCkehC1qgi
… host and commits it (the 26.10.1 chain's qRJ, CAMPAIGN row 29)

The Python-cost stage is host-side, not runner cells: it writes $HOME/pycost/mini_results_<pin>.csv outside the tree (a file
the laptop later commits must not sit untracked on the bench host, BUGS F19), and its header says the landing copies it to
benchmarks/python-bindings/jpype_overhead/results/mini_results_<pin>.csv, which export_web prefers over the tracked file.
land_stage had no such step, so a qRJ landing would have published the table from the September file with every gate green.

land_stage.pull_pycost scp's the pin's file when the landing covers it (no --only-lanes, or pycost named), refuses an empty one,
says when the host has none yet, counts the RESULT lines, and the commit step adds the file it pulled to the bindings commit.

test_land_stage_pycost.py: scoped away, pulled by pin and counted, none on the host, empty, and the commit and read paths.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ubz3m6QXHESerCkehC1qgi
…y MongoDB's graph hops are chained lookups again (CAMPAIGN rows 31, 27, 28)

Row 31: check_close_cost printed its section as F11, which FAIRNESS.md numbers F13 (F11 there is equivalent queries); a
reader following a finding to FAIRNESS.md landed on the wrong invariant. The three headings read F13 now and
test_fairness_headers.py holds the numbering.

Rows 27 and 28: DECISIONS #131 item 2 says to use $graphLookup wherever it expresses the question and to keep and disclose
$lookup where it cannot. With every graph question undirected over one stored edge per friendship (row 56), a
single-direction $graphLookup cannot state the walk exactly, so the hops are chained $lookup, which PROTOCOL.md section 7
now says with the reason (the old row carried a performance figure from the directed question), and the exclusion of edges
already on the path is stated at every hop. QUERY_LANGUAGE already says what runs (row 56). CAMPAIGN rows 27 and 28 record both.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ubz3m6QXHESerCkehC1qgi
…ge (issue #150, CAMPAIGN row 63)

insert_many hands the engine one JSON text per batch, so every value is formatted in Python, parsed in Java, and converted
again. insert_columns takes the data the way an analytics user already holds it, one array per property (numpy arrays, a
pandas frame, or plain sequences), and crosses the bridge once per column: a long[], double[], or boolean[] copied from the
numpy buffer, a String[] built from a sequence. The documents are built in Java (DocumentBatcher.insertColumns), in the
caller's transaction when there is one and otherwise in commit_every batches, or through the async executor's bucket writers
with parallel=True, where a record a writer rejects is reported, not dropped.

The failure contract is insert_many's: a transaction the call opened is rolled back whole on a failure, a caller's own is left
alone. Bad input is refused before anything is written: ragged columns, no columns, a non-string name, a two-dimensional
column, a dtype that does not cross natively (a TypeError that names insert_many), and an unsigned value past the signed
range. A pandas nullable column and a sequence with None cross as nulls.

tests/test_insert_columns.py (19 tests) holds typed values landing with their declared types, nulls, lists and frames,
the rows being the ones insert_many stores, commit_every batching, the parallel mode and its reported rejection, the failure
contract in all four shapes, and each refusal. The docs name it the recommended path for column data (api/database.md,
guide/import.md, the bridge and testing pages). Run against the official 26.10.1 jars in the official embedded image with
the patched bridge jar and core.py: 19 passed; the same file against the wheel as built without it fails on every test.

The benchmark lane that uses it follows in its own commit.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ubz3m6QXHESerCkehC1qgi
…s' columnar insert (CAMPAIGN row 63, the user's choice of 2026-10-04)

Each parquet batch crosses the bridge as whole columns (a long[] or double[] from the numpy buffer, a String[] from a list) and
the documents are built in Java, instead of one JSON payload per batch. The transport is the only change: the same async
writers, the same bucket rule, the same executor flush class, parallel=True. A wheel without Database.insert_columns keeps
insert_many and the row records nothing, so a measurement meant to be columnar cannot silently become the other one;
`columnar_insert` is written only when the call was made, and the page's list of changes already keys on it.

The documents tables' ingest sentence is now chosen from the rows behind the table (_docs_ingest_note): the columnar sentence
when every ArcadeDB embedded row carries the field, the October insert_many sentence (kept beside it as
OCT_PROSE[...]["ingest_insert_many"], with its pins) while any row does not, so no table says "insert_columns" about rows
that loaded through insert_many. test_docs_columnar_load.py holds both (6 tests; 3 fail and 2 error against the sources
before this commit). The October preview payload is unchanged.

The bindings change this depends on is 7a36dc6. CAMPAIGN row 63 records both, and that the measured wheel must contain it.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ubz3m6QXHESerCkehC1qgi
… against the DuckDB reference on the SF1 slice, CAMPAIGN row 56)

The undirected spelling of friend_age_by_city computes the per-city mean as a sum over a count (S / N) because the arrow
traversals cannot ask for the mean of a concatenation per row. SurrealQL divides two integers as integers (5 / 2 is 2), so the
adapter answered 41 where the reference and every other engine answered 41.79, and the answer digest, which rounds at the sixth
digit, split. The sum is cast: `<float> S / N`. The served twin inherits the text.

test_graph_undirected.py gains two tests: the text carries the cast on both adapters, and the real engine (the SDK's in-memory
store, skipped when surrealdb is not installed; run with --with surrealdb==2.0.0) returns 41.25 for four friend ages that sum to
165, where the old text returned 41.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ubz3m6QXHESerCkehC1qgi
…y by design (CAMPAIGN rows 58 and 59)

Row 58 added graph_insert_edges and graph_delete_edges, the KNOWS edges into the written persons read back after the insert and
after the delete. The second is empty because the delete removed them, which is the point of reading it, but the check reported
it as an EMPTY answer, and a flag is meant to be a defect or a sentence someone wrote down. Both states join WRITE_STATES (the
same answer at two scales is expected of a write state) and the delete read-back is declared by design with its reason; the insert
read-back is still flagged when it is empty, because an insert that left no edge behind is a defect.

Found by running the check over the verification rows of this re-pin.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ubz3m6QXHESerCkehC1qgi
Three commits: #200 and #202 are bindings docs (the SIGTERM known issue, the release procedure); #204 adds a gate that fails the
bindings build and the release workflow for a wheel at or over 100 MB (scripts/verify_wheel_size.py, a step in build.sh and in
release-python-packages.yml, tests/test_wheel_size.py). No benchmark file is touched. The official-engine wheel for the re-pin is
72 MB, under the new limit.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ubz3m6QXHESerCkehC1qgi
…ss it is declared to run (CAMPAIGN row 69)

F10 requires every timed write cell in both durability classes. The image-defaults arm is declared in runner.ARM_RUNS to run the
documents OLTP cell in the relaxed class only (the generator stages it that way and the user confirmed it), so the gate failed it
for a strict cell nobody queued, which would have failed the landing of the chain's last stage (qRO), the one that carries the arm.
The required classes now come from ARM_RUNS (_durability_classes_required); every other arm still needs both, and another lane of
the same arm still needs both.

Found by feeding the gate the laptop rows of this re-pin's rehearsal; no earlier rehearsal had a row of this arm to judge.
test_f10_restricted_arm.py (4 tests; 2 fail against the gate before this commit).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ubz3m6QXHESerCkehC1qgi
… an expression (found by their first run on mongod, CAMPAIGN row 56)

The undirected hops are built by _step, which joined the walk so far to the knows collection with localField set to "$f" or "$f2",
the expression form the callers also use elsewhere. $lookup's localField is a field name, and mongod refuses the other with
"FieldPath field names may not start with '$'". Every interactive Mongo graph cell would have died at its first 2-hop read. The
analytics pipelines never use _step and had passed on the capped slice, and the unit tests read only the pipelines' structure.
The sigil is dropped inside _step. test_graph_undirected.py checks every $lookup the hop builders emit.

Found by the interactive reads on the capped SF1 slice, the first time this text ran on a real mongod.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ubz3m6QXHESerCkehC1qgi
… where the engine can see them (found against the DuckDB reference, CAMPAIGN row 56)

Comparing every start person of the SF1 slice with a DuckDB reference found 26 of the 96 present starts off by one or more on the
filtered 3-hop read and on the visited count (the 2-hop and the point and 1-hop reads agreed). Probing the embedded core
(2.3.10) found why: inside a closure (`|$a| ...`) a query parameter such as `$p` reads as NONE, and an outer closure's variable is
not visible to an inner one, so `complement(..., [$p])` and `complement(..., [$a])`, both written inside the inner closures,
removed nothing, and the walks p-a-p-x and p-a-b-a were counted. What is visible inside a closure is the current record, so the
start is excluded as `[id]`, and the first friend is excluded from the union over b in the OUTER closure, where `$a` is its own
variable. For one fixed a that union minus {a} is exactly the ends of the walks p-a-b-x with b not p and x not a. The 2-hop text,
whose exclusion sits outside any closure, was right and is unchanged. All 96 present starts now equal the reference.

test_graph_undirected.py gains a real-engine test (the SDK's in-memory store; skipped without surrealdb) that compares every start
of a small graph with triangles and a pendant person against a brute-force enumeration of distinct-edge walks, and a static one
that holds the shape of the text; four tests fail against the text before this commit. The served arm inherits the text and runs
a different engine version (3.2.4): its cell is the check.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ubz3m6QXHESerCkehC1qgi
…h the exclusions on the vertices (found against the DuckDB reference, CAMPAIGN row 56)

Cypher's MATCH forbids reusing a relationship within one pattern, and that rule is what makes the shared 2-hop, filtered 3-hop, and
visited reads the question "walks of k distinct friendships". FalkorDB and LadybugDB do not apply it to a chain of relationships:
on the SF1 slice both answered those three reads differently from the DuckDB reference (the point and 1-hop reads, with one
relationship, agreed), because the walk along a friendship and straight back along it was counted, so the start person was its own
friend of a friend. A four-person graph shows it on LadybugDB in the test: the shared text answers 3 for the friends of friends of
person 1, the right answer is 2.

graph_common gains OLTP_READS_BY_VERTEX and HOP3_VISITED_BY_VERTEX (`fof <> p`; `m2 <> p AND x <> m1`, the exclusions DuckPGQ's
and SurrealQL's spellings already carry), and Base.REPEATS_RELATIONSHIPS, True for exactly those two adapters, selects them. In a
simple graph, which LDBC's knows is (0 friendships stored twice at SF1), the walks the exclusions remove are exactly the walks
that reuse a relationship, so the two spellings answer alike. PROTOCOL records the spelling beside the MongoDB row.
test_graph_undirected.py: which engines use it, that the others are sent the shared text, and the LadybugDB behaviour itself.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ubz3m6QXHESerCkehC1qgi
…verified and found (REPIN-REHEARSAL, round 5)

Row 56 records the DuckDB-reference comparison (analytics 12 of 12 engines, interactive reads 11 of 12 on the capped SF1 slice) and
the four defects it found in the row's spellings, each fixed in its own commit. Row 5 records that the strict class was run through
the runner at micro on both ArcadeDB documents arms. Row 55 records that a warm-up overshoots its cap by up to one read per
operation for a slow engine. Row 69 records that F10 asked the image-defaults arm for a class it never runs, now fixed.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ubz3m6QXHESerCkehC1qgi
…ction (lean_http), requests restorable (levers, CAMPAIGN 7 row 72)

About three quarters of a served ArcadeDB statement was the Python client: one bound one-row read costs 468 to 628 us through requests.Session and
119 to 150 us through a persistent http.client connection against the same 26.10.1 server. lean_http.Session() has the part of requests.Session's
surface the lanes use (.auth, .get, .post(json=, data=, headers=, timeout=), .close, and a response with status_code, reason, headers, content, text,
json(), ok, url, raise_for_status() with requests' own message), sends the same headers, Basic auth and JSON bytes, decodes gzip and deflate, streams an
iterator body chunked, sets TCP_NODELAY, replaces a connection the server closed, retries a read once on a reused connection and a write never.

The nine ArcadeDB served call sites change by one import (l1_tabular, l1_tpc, l2_graph, l3_sparse, l3d_dense, l4_tsbs x2, e2_hybrid,
l5_lifecycle_server, l6_restart's re-attach). BENCH_ARCADEDB_HTTP_CLIENT=requests restores the October client (default lean, anything else refused).
Every served row records arcadedb_http_client from the session that ran, the manifest records the resolved choice, runner.run_cell forwards the
variable to the lane container, and page_check declares the field in NOT_PRINTED. No comparator client moves.

test_lean_http.py (70 tests): response parsing and raise_for_status text equal to requests' for 200, 4xx, 5xx, empty, large, gzip and deflate; the
request each client sends; reconnect; retry only for reads; the timeout class; the switch; the row, manifest, allowlist and page-gate records.
test_lean_http_live.py (10 tests, skipped unless LEAN_HTTP_LIVE names a server): the lanes' own adapters against arcadedb-c25:26.10.1 with both
clients, identical answer digests. Laptop A/B on cores 6-7, 10 kept runs per arm, two runs: requests/lean 3.05 to 3.07x on a bound read, 2.5 to
2.6x on a one-statement write, 1.3 to 1.4x on a session transaction, 1.06x on an 8,000-row result, 1.0x on large bodies.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ubz3m6QXHESerCkehC1qgi
…s and names them (CAMPAIGN 7 row 72 addendum)

The e4 table said what the client/server split costs, timed through one Python client. The e4 cell's image (dbbench:arcadedb) has no requests, so that
client was e4_decomp.py's urllib shim, which opens a connection for every call, and it no longer described anything a served cell pays now that the
served arms use lean_http's persistent connection.

lean_http takes the per-call auth= the probe passes. deployment_decomp_probe.py measures every HTTP arm under both clients in one interleave that
rotates its start each round (inproc_http and docker_http through the probe's own session, the _lean arms through lean_http.LeanSession), checks every
path returns the same rows, names the client of each arm (meta.arm_clients, meta.client_names), records it on each measurement
(arcadedb_http_client) and derives the client rows of the decomposition. e4_decomp.py stamps arcadedb_http_clients on the row; the shim names itself.
export_web prints each HTTP column once per client and generates the sentence that names them; an artifact without arm_clients (October's) keeps its
three columns, and repetitions that disagree about the clients refuse to build. page_check declares the new row field; PAGE-SPEC and CAMPAIGN row 72
say what changed and what is owed at the re-pin (the e4 stage re-runs; the page prose around the table is rewritten).

test_e4_client_axis.py (19 tests, 15 fail on the previous commit): per-call auth against requests, the rotating interleave, the arm and client naming,
the probe's main() end to end against stubs for both the in-process server and the container with requests and with the shim as the legacy client,
the row stamp, the page table with and without the axis, the refusal on mixed repetitions. test_lean_http_live.py gains two live tests that run
e4_decomp.py against a real server with requests importable and blocked: five arms, identical answer digests across all of them.
Laptop A/B, cores 6-7, 10 runs per legacy client: the client is 0.4 to 2.0 ms of an HTTP call at 1 to 10,000 rows and not separable at 100,000.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ubz3m6QXHESerCkehC1qgi
…d per row (CAMPAIGN 7 row 73)

The wheel declares jpype1>=1.5.0 with no upper bound and Dockerfile.bench resolved the newest JPype at image build time, so an image rebuilt on the day of the next JPype release would change the instrument in the middle of a campaign and split its rows. A laptop A/B of JPype master (240 commits ahead of 1.7.1, unreleased) against 1.7.1 under the 26.10.1 wheel moves raw JPype call costs by 15 to 25 percent and removes the leak of one Python object per boxed number returned from Java (jpype#1379).

Pin: Dockerfile.bench takes ARG JPYPE_VERSION=1.7.1 and installs jpype1==${JPYPE_VERSION} in the same pip command as the local wheel (and whenever PIP_PACKAGES names arcadedb-embedded), then imports it and fails the build on any other version. build_images.sh reads the default from the Dockerfile and passes it down. verify_pair_c25.sh checks the embedded image's JPype as its fifth check. bench_common.JPYPE_PIN is the same number; make_2610_stages.py --check (and every emit) refuses a Dockerfile whose default differs or that installs jpype1 any way but a hard ==. Every generated stage reads JPype back out of dbbench:arcadedb after the build and aborts naming both versions; the host-side Python-cost stage does the same for the repo venv.

Record: every row carries jpype_version, read from jpype.__version__ of the process that ran the cell (bench_common.run_conditions, and the five lane writers that do not call it); an arm that never imports JPype records an empty value and asking never imports it. The run manifest records the version the ArcadeDB bench image holds when the batch has an embedded ArcadeDB arm. page_check.NOT_PRINTED declares the field.

test_jpype_pin.py (41 tests, 40 fail on the base): the Dockerfile RUN block run with a recording pip and a fake JPype, the generator refusing a moved default or a non-hard install, the stage checks run against stub docker and a stub venv python for the abort message, the recording path with and without JPype loaded, the manifest, the gate declaration, row 73. Also run for real: the Dockerfile built with the 26.10.1 wheel gives JPype 1.7.1 and the stage check passes, a 1.7.0 build is refused with the message. Gates: pytest benchmarks/experiments 523 passed and 19 skipped in a clean uv env (base 482 and 19); queue_lint on the regenerated stages and fairness_check and page_check --preview output identical to the base.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ubz3m6QXHESerCkehC1qgi
…ient decodes, not by which codings the installed urllib3 advertises

Six request-equality tests asserted that lean_http's Accept-Encoding equals requests' own. requests advertises "gzip, deflate, br, zstd" when its urllib3 has the optional brotli and zstd libraries (the repo .venv) and "gzip, deflate" without them (a clean env), so the expectation depended on the environment the test ran in, and mini's repo .venv is where these tests may run.

The equality loop no longer compares accept-encoding. New tests: the lean client's header is stable across calls, every coding it names round-trips through a stub server that encodes in that coding (a coding the stub cannot encode counts as undecodable), and it is a subset of what the installed requests sends. The check is run against "br" and "zstd" headers and against a client patched to advertise them, and both are reported by name; with lean_http.py itself changed to advertise br the real test fails ("advertises ['br'] and does not decode it"). The behavioural decoding checks stay and gain identity and raw deflate, each equal to requests.

pytest benchmarks/experiments: clean uv env 533 passed, 19 skipped, 0 failed; repo .venv 532 passed, 19 skipped, 1 failed (the known test_sparse_mp_versions).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ubz3m6QXHESerCkehC1qgi
The bindings 26.10.1 release (629b242), #214 (fewer JPype crossings on a short read: RowAccess.nextRows and RowBatcher.nextJsonBatch, scalar
parameters bound without the conversion walk), #216 (docs, CI summary, scripts), the pruned foreign .github file, and upstream engine syncs.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ubz3m6QXHESerCkehC1qgi
…n whether the wheel is installed (CAMPAIGN row 64)

The test asserted that a comparator pass's engine_version never starts with "26.". engine_version on a comparator pass is whatever
run_conditions() stamps from the package in the container: the ArcadeDB wheel's version where one is installed (a repo venv) and
"unknown" where none is (a bench client image), which is exactly why the pass also records lib_version, the field the page reads
(export_web). So the assertion passed vacuously in a clean environment and failed in a venv holding the wheel. The test now holds
what the row promises: lib_version names the adapter's own version, and no sparse_result is claimed for a comparator.

Passes in the clean uv environment and in the repo venv.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ubz3m6QXHESerCkehC1qgi
…of bindings #214; row 63 names the wheel by the tree it is built from

Bindings #214 (fewer JPype crossings on a short read, scalar parameters bound without the conversion walk) landed on main after the
re-pin wheel was first built, and the bindings 26.10.1 release was cut before it and before the columnar insert. The wheel the stages
bake is therefore a laptop build from this tree over the official jars, version 26.10.1, rebuilt after the merge of origin/main
(abd7319): row 74 records what it carries, its size and sha256, the commit it was built from, that it is not the PyPI file, and that
the re-pin's ArcadeDB embedded rows are not comparable with October's at the call-overhead level.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ubz3m6QXHESerCkehC1qgi
…ement without .tolist() (the JVM payload guard, CAMPAIGN row 63)

tests/test_jvm_payload.py, which came in from main after the columnar insert was written, flags a .tolist() result that reaches
JArray(): core.py:107, the path for a string, bytes, or object column. That path is correct (a text column has no native array and
converts one value at a time, reusing one Java String per distinct value), but the guard reads the shape of the code, and the
single failure of the bindings suite on the rebuilt wheel was this one. The elements of such an array are already Python objects, so
list(values.astype(object, copy=False)) gives the same list without the flagged call.

test_insert_columns.py gains a test for a NumPy unicode array and an object array holding None, which the file did not cover
(the strings it used were Python lists). The wheel the stages bake is rebuilt from this commit.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ubz3m6QXHESerCkehC1qgi
… payload guard on the columnar insert (sha256, size, commit)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ubz3m6QXHESerCkehC1qgi
…FAILS = True), CAMPAIGN row 68 (DECISIONS 161 D3)

An unstamped row fails the gate from now on; the tests that assumed the report-only default set it explicitly, and one that reset the switch to False in a finally block restores the previous value.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ubz3m6QXHESerCkehC1qgi
@tae898
tae898 marked this pull request as ready for review October 6, 2026 10:43
@tae898
tae898 merged commit 11cdf96 into main Oct 6, 2026
26 checks passed
@tae898
tae898 deleted the repin-prep-2 branch October 6, 2026 11:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

benchmarks: the instrument changes for the 26.10.1 re-pin campaign (CAMPAIGN rows 1 to 70), with a columnar insert in the bindings (#150)

1 participant