Skip to content

Bump lucene.version from 8.9.0 to 8.10.0 - #91

Merged
lvca merged 1 commit into
mainfrom
dependabot/maven/lucene.version-8.10.0
Oct 4, 2021
Merged

lvca merged 1 commit into
mainfrom
dependabot/maven/lucene.version-8.10.0

Conversation

@dependabot

@dependabot dependabot Bot commented on behalf of github Oct 4, 2021

Copy link
Copy Markdown
Contributor

Bumps lucene.version from 8.9.0 to 8.10.0.
Updates lucene-analyzers-common from 8.9.0 to 8.10.0

Updates lucene-queryparser from 8.9.0 to 8.10.0

Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting @dependabot rebase.


Dependabot commands and options

You can trigger Dependabot actions by commenting on this PR:

  • @dependabot rebase will rebase this PR
  • @dependabot recreate will recreate this PR, overwriting any edits that have been made to it
  • @dependabot merge will merge this PR after your CI passes on it
  • @dependabot squash and merge will squash and merge this PR after your CI passes on it
  • @dependabot cancel merge will cancel a previously requested merge and block automerging
  • @dependabot reopen will reopen this PR if it is closed
  • @dependabot close will close this PR and stop Dependabot recreating it. You can achieve the same result by closing it manually
  • @dependabot ignore this major version will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself)
  • @dependabot ignore this minor version will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself)
  • @dependabot ignore this dependency will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself)

Bumps `lucene.version` from 8.9.0 to 8.10.0.

Updates `lucene-analyzers-common` from 8.9.0 to 8.10.0

Updates `lucene-queryparser` from 8.9.0 to 8.10.0

---
updated-dependencies:
- dependency-name: org.apache.lucene:lucene-analyzers-common
  dependency-type: direct:production
  update-type: version-update:semver-minor
- dependency-name: org.apache.lucene:lucene-queryparser
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
@dependabot dependabot Bot added dependencies Pull requests that update a dependency file java labels Oct 4, 2021
@lvca lvca self-assigned this Oct 4, 2021
@lvca lvca added this to the 21.10.1 milestone Oct 4, 2021
@lvca
lvca merged commit f051598 into main Oct 4, 2021
@dependabot
dependabot Bot deleted the dependabot/maven/lucene.version-8.10.0 branch October 4, 2021 14:33
tae898 pushed a commit to humemai/arcadedb-embedded-python that referenced this pull request Jun 28, 2026
…ene.version-8.10.0

Bump lucene.version from 8.9.0 to 8.10.0
tae898 added a commit to humemai/arcadedb-embedded-python that referenced this pull request Sep 14, 2026
…found

DECISIONS ArcadeData#86: every lane and every backend at the smallest size each has, one
repetition, both durability classes on every timed write, 138 rows into
results/runs_skeleton_laptop.jsonl, 120 of them frozen, published to the
preview route with every gate green. Seven cells exceeded their budget and are
censored observations rather than gaps: SurrealDB embedded on document
analytics, graph analytics, and the cross-model lane in both classes, and both
ArcadeDB document-path time-series arms.

THE EXPORTER DID NOT PUBLISH THE OCTOBER COLUMNS. The rows have carried them
since the instrument branch landed; the page's column lists had not moved, so
the document tables printed neither the four single-record operations nor three
of the five analytical queries, the graph tables printed neither update nor the
two new analytics, the dense table printed one ingest timer instead of two and
neither mutation pass, the cross-model table printed neither read path, and no
table printed a cold column. OCT_TABLE_METRICS now REPLACES a lane's September
column list under the 2026-10 instrument rather than adding to it, which is
ArcadeData#89's own arithmetic: one warm median per query, one ninety-ninth percentile on
the table's headline query, one cold column, throughput, recall, memory, disk,
and the vector lanes' two build timers.

FIVE MORE THINGS THE RUN FOUND, each of which would have shipped:

  * The document split was keyed on the RETIRED synthetic lanes being present,
    so under October, where they are not run, the page would have gone back to
    one raw table with every column in it. Keyed on the TPC lane now.
  * The graph analytics table was keyed to the campaign's two LDBC tiers, so a
    skeleton silently drew ten tables instead of eleven and nothing said which
    one was missing.
  * The time-series table's size column was a typed constant, so a laptop slice
    printed the campaign's corpus over it.
  * `_censored_cells` read runs.jsonl whatever file the publish was built from,
    so a laptop timeout left an engine off a table with no note at all.
  * A skeleton export overwrote results/web_benchmarks.json, the tracked record
    of what the LIVE page serves. It writes web_benchmarks_skeleton.json now,
    the same split make_paper_tables already applies to the freeze.

THE ANSWER CHECK EARNED ITS KEEP. It refused the publish over the per-host
hourly time-series query: ArcadeDB's served arm returns one bucket per host
where its own embedded twin, on the same build, and DuckDB, SQLite, MongoDB,
QuestDB and TimescaleDB all return one per host and hour. The single-key form
of the same expression agrees on both arms, so it is the two-key group-by that
is wrong. equivalence_check gains KNOWN_DISAGREEMENTS, which prints it on every
run and stops failing the publish, and export_web gains WITHHELD_CELLS, which
takes that one cell off the page and says why under the table. A latency for a
query that answered something else is not a measurement of that query.

Also: the graph lane stamped `graph_source` as "ldbc-<scale>" on every run,
synthetic ones included, because the line sat inside the gen_edges wrapper; the
laptop has no LDBC corpus and its rows claimed one. F11's known-regression
exception matched on engine_commit alone, which a skeleton does not stamp, so a
documented upstream-fixed cost failed the one publish that could do nothing
about it; it matches the releases now, which is evidence rather than a claim.
page_check found the site by assuming the two repositories are siblings, which
a worktree is not, and then died in a traceback instead of reporting it.

The skeleton's own guards, unchanged and exercised: every frozen row is the
laptop at sweep tier, F1 and F3 are waived by name in the payload, every other
gate runs as it will in October, and the banner and every table's conditions
say in plain words that the machine was running other work while these were
timed, that each cell ran once rather than five times, and that the numbers are
not comparable between engines or against the live page.

AND DECISIONS ArcadeData#91, WHICH ARRIVED MID-SWEEP AND IS IN THIS COMMIT. Both served
SurrealDB dense cells at ten million vectors on the bench host built their
index and then died inside the timed queries with "no close frame received or
sent", on a server that was still up. Read out of the SDK rather than guessed:
it opens its socket with the websockets library's defaults, whose keepalive
pings every twenty seconds and closes the connection when no pong arrives
within twenty more, so a server busy answering a long query gets hung up on by
its own client. surreal_common now sets those options explicitly and records
them on the row, wraps the served client so a dropped socket is rebuilt,
re-authenticated and re-selected once, and DISCARDS the sample taken across the
drop rather than publishing the reconnect's cost as a query's latency. Every
SurrealDB row carries `reconnects`, zero when nothing happened.

Proved on the laptop, since the ten-million case cannot be: the server
container was restarted underneath a running micro cell, and the cell finished
rc=0 with reconnects=1, 1,959 timed operations against every other engine's
1,960 (the discarded sample), and new-order and insert digests identical to
DuckDB's and ArcadeDB's. test_surreal_reconnect.py covers the machinery itself,
fourteen checks with no container.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JB6Hg77dQVqABoTJmiUnV2
tae898 added a commit to humemai/arcadedb-embedded-python that referenced this pull request Sep 14, 2026
…us hid

The two wrong answers found on 2026-09-14 (BUGS F42, F43) were both caught at
scale factor 0.01. Answer equivalence is scale-sensitive in both directions, so
the document lane ran again on the laptop at the campaign's own SF1: 6,001,215
line items and 200,000 parts, ten engines, both workloads, one repetition,
BENCH_OLAP_ITER=2 because the point is the answers rather than the latencies.
Nine of ten engines answered the analytical workload and all ten answered the
transactional one, into results/runs_sf1_equiv.jsonl. SurrealDB embedded loaded
SF1 in 1,832 s and was killed inside its first Q1 at the tier's 2,700 s budget,
a censored observation with the phase on the row.

THE PRICING SUMMARY SPLIT SEVEN ENGINES TO ONE, AND NOBODY WAS WRONG. ArangoDB
returned sum_qty as 37734107 where every other engine returned 37734107.0. AQL's
SUM over a column VelocyPack stored as integers returns an integer; SQL sum(),
MongoDB's $sum and SurrealQL's math::sum return a double. The canonical form
prints an int exactly and a double to six significant digits, so the same number
became "37734107" and "3.77341e+07". In the same row n was an int on both sides
and identical, avg_qty was bit-identical, and avg_qty * n is sum_qty: the two
sides held one number and the digest split on its spelling. The two spellings
coincide below 10**6, which is why the SF0.01 skeleton passed; at SF1 three of
Q1's four groups cross the boundary and at SF10 all four do. That gate would
have refused an October publish over arithmetic.

The fix is a DECLARATION, not a looser hash. Every summed or averaged measure is
declared `num` and compared as a number whatever type the driver returns, while
counts and identifiers stay exact, because rounding a count to six significant
digits would let 1,234,567 and 1,234,568 agree. Applied to l1_tpc's five
analytical queries, to the graph lane's two averages, and to every value column
of the time-series lane, whose TSBS fields are line-protocol integers and are
latent for the same reason. Declared per query and never per engine. The five
SF1 queries now agree across all eight engines that produced a clean row, five
of them re-measured end to end and every one of the eight re-digested from
stored full-precision answers.

AND SIX SIGNIFICANT DIGITS IS THE RIGHT ROUNDING, MEASURED RATHER THAN ASSUMED.
tpc_answer_probe.py records each engine's answers at seventeen digits through
the lane's own adapters; tpc_rounding_audit.py re-digests them at 6, 8, 10, 12
and 17 and reports the observed spread. At SF1, over 111 float cells and eight
engines, the worst cross-engine gap is 1.45e-14 relative against a six-digit
tolerance of 4.5e-6, which is 3e8 of headroom, and six independent summation
strategies over the real column span 6.1e-14 (8.1e-13 at ten times the addends).
It is not too tight. Nor is tightening the fix: at TEN digits the real SF1 data
splits, PostgreSQL alone on top_parts, because five engines return exactly
2659862.7955 while DuckDB and PostgreSQL land one double ulp either side of it
and TPC-H money sums at that magnitude have exactly eleven significant digits,
so ten-digit rounding cuts on the last one and the noise decides the bucket.
Straddles per 111 cells: none at 6, none at 8, one at 10, none at 12. The
exposure that remains is at SF10, where a single missing row becomes invisible
in Q6 and by_month, the two queries carrying no count column; closing it wants a
count on those queries, not a different rounding, and that is a query-text
change for the campaign owner rather than something to land here.

TWO MORE DEFECTS, EACH FOUND ONLY BECAUSE THE CORPUS WAS BIGGER:

  * EVERY MONGODB CELL IN THIS LANE DIED at the final json.dump with "Object of
    type Collection is not JSON serializable", after the load and after every
    query. A pymongo Database answers ANY attribute name with a Collection, so
    stamp_reconnects' getattr(client, "reconnects", None) was never None and
    wrote that object onto the row; sample_kept had the same duck-typing and
    returned True by luck. Not scale-dependent -- DECISIONS ArcadeData#91 landed mid-sweep
    and the l1tpc MongoDB cells had already run, so SF1 is simply the first size
    at which they ran again. A reconnect counter is an int now, and the test
    grows an adapter whose client answers every attribute name.

  * AN INDEXED `>=` RANGE ON A STRING COLUMN LOSES ROWS. Both ArcadeDB arms
    answered Q6 at SF0.1 with 11,801,684.4174 against eight engines'
    11,803,420.2534 -- short by 1,735.836, which is exactly one line item of the
    qualifying set. Isolated through the lane's own adapter: the window
    `l_shipdate >= '1994-01-01' AND l_shipdate < '1995-01-01'` returns 92,037
    rows where `> '1993-12-31'` and a substring predicate no index can serve both
    return 92,040, deterministically over ten repeats and two builds, and 908,652
    against 909,455 at SF1. It is NOT registered as a known disagreement: at SF1
    Q6's other predicates happen to exclude every lost row, so the gate is
    correctly failing at the size where the loss lands on a row that qualifies.
    It wants a Java repro against the index range scan before it goes upstream.

Also measured and written down: the ArcadeDB SQL parser narrows a decimal
literal that needs more than single precision, even into a DOUBLE property,
while a bound parameter does not (arcadedb_literal_precision_probe.py, six of
fifteen literals). TPC-H money is unaffected and the served arm's SF1 corpus is
bit-identical to the embedded arm's, so the served TPC rows stand.

equivalence_check grows --list, which prints every group with the engines that
answered it and the digest each gave. The counts alone cannot be audited: "ok,
32 groups agreeing" reads the same whether thirty-two questions were put to ten
engines or to two.

The time-series lane's ts100 IS its campaign size and was re-measured here
independently: seven engines, six queries, five unanimous and q_groupby splitting
with the served native arm alone, which reproduces the documented known
disagreement exactly. A complete SF0.1 pass covers the one size at which all ten
engines, SurrealDB embedded included, answer together.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JB6Hg77dQVqABoTJmiUnV2
tae898 added a commit to humemai/arcadedb-embedded-python that referenced this pull request Sep 15, 2026
…, ts) index

Three comparator arms for the time-series lane, on SQLite's footing: no
time-series type, one table with a datetime field, the tag, the three
metrics, and a composite (host, ts) index defined before the load.

surrealdb_ts (SDK 2.0.0, core 2.3.10, SurrealKV) and surrealdb_ts_server
(3.2.4 on RocksDB, the pinned digest) run one SurrealQL text: time::floor
for the minute and hour buckets, d'...' datetime literals for the ranges,
and the grouped-ordered-limited query as a subquery, because core 2.3.10
sorts by the group key ascending after GROUP BY and would take the wrong
five buckets (BUGS F31 again); 3.2.4 answers both spellings in the same
time. The served arm goes through the reconnecting client (DECISIONS ArcadeData#91),
the timed loop drops a sample taken across a reconnect, and every SurrealDB
row records `reconnects`.

arangodb_ts (3.12.11, the pinned digest) holds ts as epoch milliseconds,
a persistent index over ["host", "ts"], DATE_TRUNC for the buckets, and
the collection is created with waitForSync at the cell's class and read
back (DECISIONS ArcadeData#90, ArcadeData#81).

Laptop smoke through runner.py (ts100, one rep, BENCH_QITER=3, cpuset
0-11, results/runs_ts_plain.jsonl, untracked): all three cells ok, every
query returned the expected shape (1 / 60 / 12 / 1200 / 32,944 / 5 rows,
100 hosts x 12 buckets), and equivalence_check over these rows plus the
skeleton's reports every digest equal to the engines already on the table
(32 groups agreeing, 0 failures; the one KNOWN entry is the pre-existing
arcadedb_ts_native_server q_groupby). Laptop numbers are smoke, not record.

Registries: runner BACKENDS and LANES["l4"], export_web display names,
L4_CANON_LABELS, order and deployment, fairness_check UNVERIFIED_ALLOWED
for the served SurrealDB arm, and one COMPARATORS.md role per engine.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JB6Hg77dQVqABoTJmiUnV2
tae898 added a commit to humemai/arcadedb-embedded-python that referenced this pull request Sep 15, 2026
…es what it censors

DECISIONS ArcadeData#100. Embedded SurrealDB 2.3.10 scans an all-host time range at
about 35 us per row, so at ts100 each of the four scanning queries is a
minute per iteration and a hundred iterations is eight to eleven hours
against the lane's one-hour cell timeout: the cell would time out whole and
leave nothing. The lane now takes graph_common.OLAP_BUDGET_S's mechanism
(#82b): the clock starts before the cold pass, the cold pass always runs,
the loop stops at the first iteration that would start after the budget,
and the row records <q>_budget_s, <q>_iters, <q>_elapsed_s and
<q>_censored, counted on iterations run rather than samples kept so a ArcadeData#91
reconnect drop cannot read as a budget hit.

QUERY_BUDGET_S = 600 s, a property of the lane and the same for every
engine. Derived from the slowest legitimate served engine at ts100,
ArcadeDB's own served document path: 1.55-1.65 s per twelve-hour aggregate
on mini in September (runs_paper.csv, five reps), so about 165 s for a
hundred iterations, and the three October scans of the same shape sit
within 8% of it on the laptop skeleton, so about 180 s on mini; 600 s is
3.3x that, the largest value at which the slowest engine measured so far
fits its four scans, ingest, last-point iterations and shape check inside
the 3,600 s cell timeout, and above the laptop skeleton's slowest served
arm (531 s), so the preview censors nothing the bench host would not.
BENCH_TS_QUERY_BUDGET_S overrides it for a laptop probe only.

export_web: the table now NAMES a per-query censored cell. _censored_notes
covered the whole-cell timeout; the per-query fields the graph lane has
recorded since #82b reached no sentence, so three skeleton triangle counts
stood at 16, 5 and 18 iterations under a note that said 100.
_query_budget_notes reads <q>_censored/<q>_budget_s/<q>_iters (and
_elapsed_s) off the frozen rows for l2olap and l4 and prints, beside the
counts note, which engine's which query exceeded which budget after how
many iterations, reaching what time, with the lane's own reading of what
the count includes (the graph lane's cold pass is outside its warm count;
the time-series lane's first iteration is the cold pass). Not a declared
absence: the cell keeps its p50 over the iterations taken, so the coverage
gate sees a value. page_check A2 declares <q>_elapsed_s measured-not-printed
beside <q>_budget_s. PROTOCOL.md section 2 carries the rule beside the
query-set bullet.

Proved on the laptop: surrealdb_ts at ts100 under BENCH_TS_QUERY_BUDGET_S=120
(QITER 100) ran the newest reading and the one-hour range to 100 iterations
and censored the four scans at 2 of 100 (134, 210, 124 and 123 s), every
digest still equal to the table's; a skeleton publish from those rows plus
the skeleton's printed four sentences on the l4 table and three on l2olap,
and equivalence_check, provenance_check, fairness_check and page_check all
passed. The probe rows were not folded into the skeleton and the preview
payload was restored.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JB6Hg77dQVqABoTJmiUnV2
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

dependencies Pull requests that update a dependency file

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant