fix(index): prune stale label list entries during updates - #9352
Merged
Xuanwo merged 1 commit intoSep 30, 2026
Merged
Conversation
Apply old_data_filter to existing label postings and list-level null rows before merging replacement data. This prevents updated rows with stable row IDs from matching their previous labels after index optimization. Add a two-row Python regression and Rust coverage for row ID and fragment filters, including null-list and null-element transitions. Validation: 11 Python LabelList tests, 18 Rust LabelList tests, the 32-case local comparison matrix, and all eight original core probe cases passed.
Contributor
There was a problem hiding this comment.
The update now applies the live-row filter to both per-label postings and the separate null-list bitmap before replacement rows are merged. This is the right incremental fix, with coverage for stable row IDs, physical row addresses, null/list edge cases, and the optimize path.
Already-stale LABEL_LIST indexes are not repaired by upgrading. Deployments that may have optimized such an index after updates should rebuild it once after adopting this fix; otherwise stale matches can persist.
Contributor
Author
|
@westonpace Can you review this PR? I made it for related #8840 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
With stable row IDs enabled, a
LABEL_LISTindex can match an updated row against its previous label after index optimization, even though the stored value is correct. This change appliesold_data_filterto existing label postings and list-level null rows before merging replacement data.Original scenario and impact
This was found while evaluating stable row IDs for
_row_last_updated_at_versionin an embedded LanceDB application that uses alist<string>column for access-control labels. The reporting application's production tables do not enable stable row IDs; the reproductions use synthetic data.In the original 50-row example, changing one row's
aclfrom["Anonymous"]to["Manager"]correctly reduced the results forarray_has_any(acl, ['Anonymous'])from 50 to 49. Aftertable.optimize(), the count returned to 50, including the updated row with its current["Manager"]value. This occurred with bothupdateandmerge_insert, and with both filtered scans and vector searches using prefiltering.Applications that rely on this predicate to enforce access restrictions, without independently checking the current permissions, could therefore return a row after its previous access label has been revoked. This is a query-correctness defect with potential application-level authorization impact. The reproductions do not establish a bypass of built-in authentication or a production disclosure.
The behavior was reproduced with LanceDB 0.37.1 and 0.38.0, and with source-built Lance at
2602724cf6256ff8f55571805fdc5b0614d706b8.Minimal reproduction
The problem reduces to two rows, one list column, one
LABEL_LISTindex, and an update to one row. A BTREE index, vector search, and file compaction are not required;optimize_indices()alone reproduces it.Before the fix, the final assertion fails with
1 == 0: the indexed query returns the row containing["new"], while the forced scan correctly returns no rows. The corresponding test with stable row IDs disabled passes.Cause and fix
LabelListIndex::updatereceived but discardedold_data_filter, allowing obsolete memberships to survive index updates. Pass that filter through to the existing bitmap builder and also apply it to the oldlist_nullsbitmap. Prune old state before adding replacements, since replacement rows may reuse the same stable row IDs.This addresses the existing correctness limitation documented in #8840. It preserves the streaming build approach and does not change the public API or index file format.
The Python regression covers stable row IDs enabled and disabled across two fragments. The Rust regression covers row-ID and fragment filters, null elements, null lists, empty lists, and preservation of unchanged rows.
Validation
The full workspace test suite, workspace-wide Clippy, and full Python lint target have not been run. Post-fix validation used the locally built Lance library; the released LanceDB wheels were not modified.
Existing indices
This prevents stale memberships from being retained during future updates. It does not automatically repair an index that already contains them. In the affected synthetic dataset, opening it with the patched library still returned the stale match; rebuilding the index with
create_scalar_index("acl", "LABEL_LIST", replace=True)restored the correct result.Please consider whether the conditional authorization impact warrants
critical-fixtreatment and a backport.Reported and fixed by Manabu TERADA (@terapyon).