Sequence numbers (`_seq_no`) are a key mechanism for making Elasticsearch's distributed system work.
The primary shard assigns sequence numbers of write operations, and replicas use them to make sure that they have "caught up" to the primary.
They also help to power concurrency
Elasticsearch 9.5 ships with an improved doc-values codec - ES95.
Since we have no way of knowing what the name comes from, let's talk about the codec itself.
It is an adaptive codec that chooses its own adventure (i.e. compression path).
Monotonic timestamp and long-counter
Inverted indexes are one of the core pillars of information retrieval.
It is literally THE thing that powers keyword (lexical) searches, by mapping terms against matching document IDs.
But, does every job and every data type require an inverted index? As it turns out, maybe
We indexed 138M vectors in under 10 minutes.
Benchmarked on MS MARCO: 138M vectors at 1,024 dims, from ~4 TB of documents, on 8 NVIDIA RTX PRO 6000 GPUs at 95% target recall:
Built with NVIDIA cuVS @NVIDIAAI
- Indexing throughput: 38K docs/s on CPU, 281K on GPU
- 6x lower p90