A Nokogiri-compatible Ruby binding for
libleptris, a pure-C99 XML 1.0 parser
with full XPath 1.0,
XML Namespaces 1.0, SAX, and C14N (1.0 / 1.1 / Exclusive), plus
XSLT 1.0–3.0 transforms, an XQuery 1.0 core face, and a growing
XPath 2/3.1 expression subset available standalone from #xpath
itself.
The C DOM is the single source of truth — Ruby objects are thin FFI handles over the C pointers, so every Ruby method maps to one FFI call. No tree hydration, no parallel Ruby-side model.
Add to your Gemfile:
gem "leptris"Then bundle install.
leptris shells out to the native libleptris shared library via FFI.
You need libleptris.{dylib,so,dll} installed on the host. Options:
-
Homebrew (macOS, easiest):
brew install lutaml/tap/libleptris(if packaged) or build from source (see below). -
Build from source (Linux/macOS/Windows):
git clone https://github.com/leptris/leptris.git cd leptris cmake -B build -S . \ -DCMAKE_BUILD_TYPE=Release \ -DLEPTRIS_BUILD_SHARED=ON \ -DLEPTRIS_BUILD_STATIC=OFF \ -DCMAKE_WINDOWS_EXPORT_ALL_SYMBOLS=ON cmake --build build -j sudo cmake --install build # optional, system-wide
-
Point Leptris at a specific path by setting
LEPTRIS_LIB_PATH:export LEPTRIS_LIB_PATH=/usr/local/lib/libleptris.dylib
If Leptris can’t find the library at startup, every parse call raises
LoadError.
The ruby-platform variant (what TruffleRuby and JRuby resolve)
vendors precompiled libleptris + libutf8proc binaries for the
common engine platforms under lib/leptris/vendor/:
-
arm64-darwin,x86_64-darwin -
x86_64-linux,aarch64-linux(glibc) andx86_64-linux-musl,aarch64-linux-musl(tried as fallback for Alpine hosts) -
arm-linux,arm-linux-musl(32-bit ARM EABI5) -
ppc64le-linux,s390x-linux,s390x-linux-musl(POWER8+ LE and IBM Z, qemu-built)
gem install leptris just works on both engines — the FFI layer
selects the matching vendored binary at require time (the binding
is FFI-based; no MRI C extension anywhere). LEPTRIS_LIB_PATH
still overrides, and a system library remains the fallback for
OSes the variant does not carry.
The top-level entry point is Leptris::XML. Parse a string or an IO:
require "leptris"
doc = Leptris::XML.parse(<<~XML)
<library xmlns="http://example.org/ns">
<book id="b1" lang="en">
<title>Refactoring</title>
<author>Martin Fowler</author>
</book>
<book id="b2" lang="fr">
<title>Programmer en Ruby</title>
</book>
</library>
XML
doc.root.name # => "library"
doc.root.children.size # => 5 (2 element children + 3 whitespace text nodes)Or a file:
doc = Leptris::XML.parse_file("books.xml")This is the direct Nokogiri equivalent of Nokogiri::XML(…). The
returned object is a Leptris::XML::Document.
Malformed input raises Leptris::XML::ParseError. Parse with
recover: true to get libxml2’s XML_PARSE_RECOVER semantics
instead — an empty document back, with the failure recorded on the
thread-global last error:
doc = Leptris::XML.parse("<broken", recover: true)
doc.root # => nil
Leptris::XML::FFI.leptris_last_error # => "..."Leptris::XML.parse_html(html) parses tolerant HTML4/5 into a
standard Document — the same nodes, pool, serializer, and
XPath/XSLT/XQuery machinery as XML (libleptris 1.9.75, the last
Nokogiri capability gap). The Leptris::HTML, Leptris::HTML4,
and Leptris::HTML5 module facades provide the
Nokogiri-shaped entry points over the same engine
(Leptris::HTML5.parse(html), Leptris::HTML(html)):
doc = Leptris::XML.parse_html(%(<ul><li>a<li>b</ul>))
doc.at_css("body").inner_html # => "<ul><li>a</li><li>b</li></ul>"Implied end tags (p/li/td/tr/…), void elements, raw-text
<script>/<style>, case-insensitive lowercased names,
minimized/unquoted attributes, and the HTML named entities all
parse; html/body are synthesized (an empty head is not) and
tbody is never implied. Malformed markup degrades to text rather
than raising.
Document#root
|
root |
Node#name
|
element name (e.g. |
Node#content (alias #text, #inner_text)
|
all descendant text concatenated. |
Node#[] (alias #attr, #get_attribute)
|
attribute value by name. |
Node#attributes
|
hash of |
Node#key? (alias #has_attribute?)
|
attribute presence. |
Node#children
|
|
Node#element_children
|
|
Node#first_element_child, #last_element_child
|
first/last element child (skip text nodes). |
Node#next_element, #previous_element
|
next/prev sibling element. |
Node#parent, #next_sibling, #previous_sibling
|
tree navigation. |
Node#line
|
1-based source line number. |
Node#type (alias #node_type)
|
integer type code. Element predicates: |
Example — walk all book titles:
doc.root.children.select(&:element?).each do |book|
title = book.children.find { |c| c.element? && c.name == "title" }
puts "#{book[:id]}: #{title&.content}"
end
# b1: Refactoring
# b2: Programmer en RubyNode#visit walks the subtree with ONE C call and C-tracked depth —
elements yield (node, entering, depth) enter/leave pairs, every other
kind yields once; the leanest full-subtree iteration the binding offers:
doc.root.visit do |node, entering, depth|
puts "#{' ' * depth}#{node.name} #{entering ? 'enter' : 'leave'}"
endNode#traverse walks the subtree in post-order via a single C-side
callback (one FFI call for the whole traversal, not one per node):
doc.root.traverse do |node|
case node
when Leptris::XML::Element then puts "E #{node.name}"
when Leptris::XML::Text then puts "T #{node.content.inspect}"
when Leptris::XML::Comment then puts "C #{node.content.inspect}"
end
endDocument#node is the document-level navigation head (the libxml2
model): Document#children reads [prolog comments/PIs, the root,
epilog comments/PIs] in document order, with whitespace kept per
libxml2’s exact rule. Document-level PIs are first-class:
Document#add_pi(target, data) appends one, Document#remove_pi(target_or_index)
removes by target or index, and PI#target=/PI#data=/PI#unlink
mutate and detach. Document#processing_instructions and
Document#comments remain the memoized readers.
For the dominant parse-query-serialize workload, parse with
readonly: true (or call Document#readonly!, one-way):
doc = Leptris::XML.parse(xml, readonly: true)
doc.root.children.first["id"] # reads are memoized
doc.root.name # plain ivar after first call
doc.root << doc.create_element("x") # raises ReadOnlyErrorReads (name, content, children, attributes) memoize
aggressively — they cannot go stale because mutation is forbidden.
Every mutator raises Leptris::XML::ReadOnlyError. Detached factories
(create_element and friends) still work: building a new tree
against a readonly document is legal; mutating the frozen one is not.
Document, Element, and DocumentFragment (via
Leptris::XML::Searchable) support:
#xpath(*exprs)
|
evaluate XPath; returns |
#at_xpath(*exprs)
|
first match (or scalar), like |
#css(*selectors)
|
minimal CSS-to-XPath translation, then |
#at_css(*selectors)
|
first match of |
#search(*exprs)
|
dispatches on syntax — path-prefixed expressions ( |
#at(*exprs)
|
first match of |
doc.xpath("//book") # => NodeSet of both <book>
doc.xpath("count(//book)") # => 2.0
doc.xpath("//book[@lang='fr']/title") # => NodeSet[<title>Programmer en Ruby</title>]
doc.at_xpath("//book[@id='b1']") # => <book id="b1" ...>
doc.at_xpath("string(//book[1]/@id)") # => "b1"
doc.css("book[lang='en'] title") # => NodeSet[<title>Refactoring</title>]
doc.at_css("book#b1 title") # => <title>Refactoring</title> (id selector)
doc.css("book:first-child") # first <book>
# css is receiver-relative: scoped to the receiver, not the document
doc.root.at_css("book").css("title") # titles under THAT book only
frag = doc.fragment("<a x='1'><n/></a>")
frag.css("a > n") # searches the fragmentXPath result type follows XPath 1.0 semantics:
count(…) → Float, boolean(…) → true/false,
string(…) → String, otherwise a Leptris::XML::NodeSet.
Beyond XPath 1.0, the engine accepts a growing XPath 2/3.1 expression subset standalone — no stylesheet required:
doc.xpath("let $n := count(//book) return $n + 1") # => 3.0
doc.xpath("for $b in //book return count($b/*)") # sequence
doc.xpath("if (count(//book) = 2) then 'two' else '?'")
doc.xpath("count(1 to 4)") # => 4.0
doc.xpath("//book => count()") # 3.1 arrow
doc.xpath("//title ! string(.)") # 3.1 simple map
doc.xpath("'v' || count(//book)") # 3.1 concat
doc.xpath("'42' castable as xs:integer") # 2.0 type ops
doc.xpath("1.9 cast as xs:integer") # => 1.0 (truncates)
doc.xpath("//title instance of node()+") # 2.0 sequence types
doc.xpath("map { 'b': 'beta' }?b") # 3.1 maps
doc.xpath("[10, 20, 30]?2") # 3.1 arrays
doc.xpath(%q{serialize(parse-json('{"b":"beta"}'), map { 'method': 'json' })})
doc.xpath("function($x) { $x + 1 }(41)") # 3.1 function items
doc.xpath("fold-left(1 to 4, 0, function($a, $b) { $a + $b })")
doc.xpath("count(//a) eq 2") # 2.0 value comparators
doc.xpath("some $a in //a satisfies $a/@v > 1") # 2.0 quantifiers
doc.xpath("count(//a except //a[@v = 1])") # 2.0 set algebraSequence and constructor items arrive as Leptris::XML::ResultText
objects — #content serves the value directly:
doc.xpath("for $w in //a return string($w/@v)").map(&:content) # => ["1", "2"]Remaining grammar gaps (tracked upstream): 3.1 string templates, XQuery beyond the 1.0 core.
Minimal subset (translated to XPath via Leptris::XML::CssToXPath):
-
Type/universal:
book,* -
Class/ID:
.highlight,#b1 -
Attribute presence:
[lang] -
Attribute value:
[lang='en'],[lang~='en'],[lang^='en'],[lang$='en'],[lang*='en'] -
Combinators: descendant (space), child (
>), comma (multi-selector) -
Pseudo-classes:
:first-child,:last-child,:only-child,:empty,:root,:not(…)
For anything more sophisticated, drop down to xpath.
Documents expose factory methods; elements expose mutation methods:
doc = Leptris::XML.parse("<root/>")
book = doc.create_element("book")
book[:id] = "b3"
book.add_child(doc.create_element("title")).content = "New book"
doc.root.add_child(book)
puts doc.to_xml
# <?xml version="1.0"?>
# <root><book id="b3"><title>New book</title></book></root>
Document#create_element(name)
|
detached element owned by the document. |
Document#create_text_node(str), #create_comment(str), #create_cdata(str)
|
text-class factories. |
Document#create_processing_instruction(target, data)
|
PI factory. |
Document#fragment(markup)
|
parse a markup fragment (multiple top-level children allowed). |
Element#name=, #content=
|
rename / replace inner text. |
Element#[]= (alias #set_attribute)
|
add/update an attribute. |
Element#remove_attribute (alias #delete)
|
drop an attribute. |
Element#add_child(node_or_markup) (alias #<<)
|
append a Node, or parse+append a markup String. |
Element#prepend_child(node)
|
insert as the first child. |
Element#add_next_sibling(node), #add_previous_sibling(node)
|
sibling insertion. |
Element#remove_child(node)
|
detach (does not free). |
Element#children=
|
replace all children. |
Element#replace(node) / #swap(node)
|
replace in parent. |
Element#wrap(node_or_markup)
|
wrap this element in a new one. |
Node#unlink
|
detach from the tree. |
Element#attribute_pairs
|
bulk read-only attribute listing — |
Document diffing (Leptris::XML.diff(a, b)) yields an op list with
#ops, per-kind counts via #summary
({update_attr: 1, insert: 1}, {} when identical), and #to_json
for the op list — the leptris diff --summary/--json CLI modes'
library face.
Document.create starts an empty document; build up from there
(document-level prolog parts included — PIs, comments, the XML
declaration, and the DOCTYPE are all first-class):
doc = Leptris::XML::Document.create
book = doc.create_element("book")
book[:id] = "b3"
book << doc.create_element("title")
doc.root = doc.create_element("catalog")
doc.root << book
doc.add_comment("built programmatically") # epilog comment
doc.set_doctype("catalog", system_id: "catalog.dtd")
puts doc.to_xml
Document.create
|
empty document (no root yet). |
Document#root=
|
attach the root element (bottom-up construction). |
Document#set_doctype(name, public_id:, system_id:)
|
programmatic DOCTYPE. |
Document#doctype (alias #internal_subset)
|
the |
Document#clear_declaration
|
un-set the XML declaration (serialize as if the input had none; idempotent). |
Document#remove_doctype
|
un-set the DOCTYPE (true when removed; the DocType stays readable until |
Document#add_pi(target, data) / #remove_pi(target_or_index)
|
document-level processing instructions. |
Document#add_comment(str)
|
epilog comment. |
Document#create_entity_reference(name)
|
|
Element#namespace
|
the element’s in-scope namespace as a |
Element#namespaces
|
all in-scope namespaces (inherited from ancestors) as a |
Element#namespace_definitions
|
only namespaces declared directly on this element. |
Element#add_namespace_definition(prefix, href) (alias #add_namespace)
|
declare |
Element#default_namespace=(href)
|
declare/replace |
Element#namespace=(uri)
|
Nokogiri |
Element#remove_namespace_definition(prefix)
|
drop a declaration. |
Element#attribute_ns(uri, local)
|
attribute value by expanded name (URI + local); nil URI matches no-namespace attributes only. |
Element#has_attribute_ns?(uri, local)
|
presence by expanded name. |
Attr#prefix
|
the attribute’s prefix as written ( |
Attr#namespace_uri
|
resolved through the owning element’s in-scope declarations at read time ( |
root = doc.root
root.add_namespace_definition("t", "https://example.org/types")
puts root.namespaces
# {"xmlns"=>"http://example.org/ns", "xmlns:t"=>"https://example.org/types"}
# XPath with prefixes is dispatched straight to libleptris, which resolves
# prefixes using the in-scope namespace declarations.
doc.xpath("//t:title")
Document#to_xml(indent: 0, no_decl: false, encoding: nil, indent_text: false) (aliases #to_s, #serialize)
|
serialize the whole document. |
Element#to_xml(…)
|
serialize a subtree (also takes |
indent_text carries two meanings: a STRING is the indent unit with
Nokogiri’s semantics (the unit replaces the default spaces, repeated
indent times per depth level — byte-identical to Nokogiri’s output);
true selects the display form, which also indents text and mixed
content:
doc.to_xml(indent: 2, indent_text: "\t") # tab-indented
doc.to_xml(indent: 2, indent_text: true) # display form (documents only)Element#inner_html serializes the children with correct escaping —
well-formed by construction (a re-parse spec pins it).
Element#to_xml(expand_empty: true) emits <a></a> instead of <a/>
for empty elements — libxml2’s XML_SAVE_NO_EMPTY_TAGS parity.
Node#digest(drop_ws: false) answers a content-defined 64-bit Merkle
hash of the subtree: equal flags and equal digests imply structural
equivalence (names, namespaces, sorted attributes, document-order
children); inequality implies nothing — descend and decide. Stable
across processes; drop_ws: true skips whitespace-only text nodes.
XSLT 1.0–3.0 transforms run through Leptris::XML::XSLT:
style = Leptris::XML::XSLT.parse(stylesheet_xml) # or .parse_file
result = style.apply_to(doc) # => Document
style.serialize(doc) # => StringThe engine dispatches on the stylesheet’s declared version. XSLT
1.0 is complete; the 3.0 instruction set has grown through
libleptris 1.9.36 — grouping, xsl:accumulator,
xsl:analyze-string, xsl:mode/@on-no-match dispositions,
xsl:sequence, xsl:perform-sort, and fn:format-integer — with
the XPath 3.1 core expressions available inside transforms (see
the subset list under Searching: XPath and CSS; value
comparators and sequence types are still out,
tracked upstream).
Known engine bug: the shallow-skip and text-only-copy
dispositions drop unmatched subtrees in the built-in initial
descent (#705) —
the other four dispositions are spec-correct. XQuery has no
entry point yet.
== Descriptor materialization
Compile a schema descriptor once, then materialize a whole subtree against it in ONE native pass — no per-element Ruby calls (frameworks rebuilding typed models from XML; libleptris 1.9.162):
descriptor = Leptris::XML::Descriptor.build(
name: "catalog",
attributes: [{ name: "version", kind: :scalar }],
children: [
{ name: "item", kind: :nested, plan: {
name: "item",
attributes: [{ name: "id", kind: :scalar }],
children: [
{ name: "name", kind: :scalar },
{ name: "price", kind: :scalar },
{ name: "opt", kind: :collection },
] } },
])
tree = descriptor.walk(doc.root).to_ruby
# { kind: :element, attributes: { "version" => "2.0" }, children: [
# { kind: :element, name: "item", attributes: { "id" => "1" },
# children: ["first", "1.99", ["a", "b"]] }, ...] }Row kinds: :scalar, :collection, :nested (via plan:), :raw
(serialized subtree), :content (mixed-content text runs — pair with
flags: [:mixed_content]), :callback (raw value + document byte
offset + type_tag echo, matching Node#byte_offset). Namespace
binding lives on plans: ns: is :none (default), :any, or
{ exact: "urn:…" }; flags: also accepts :cdata, :ordered,
:ns_lenient (#754 out-of-namespace adoption), and :order_spine
(#1273 — unmatched sibling text/comment/PI runs emit as positioned
scalars so ordered hosts rebuild element order without re-parsing).
Children the plan does not describe are skipped; undescribed document
order within a row is preserved. #walk returns a lazy PlanValue
tree (the result outlives the document); #to_ruby materializes it.
Plan rows accept type: :integer | :float | :boolean — the walk
parses the values natively (in-pass, libleptris #1269a) and
#to_ruby / the bulk face return Integer / Float / true /
false. Unparseable values fall back to the raw String;
PlanValue#string_value stays the raw escape, and
#int_value / #float_value / #bool_value expose the parsed
accessors directly.
Descriptor#materialize(source) is the fused entry: source bytes →
typed rows in ONE C call — parse, walk, and free happen natively
(#1269b); the standalone PlanValue tree is all that surfaces.
descriptor = Leptris::XML::Descriptor.build(
name: "iso",
children: [{ name: "row", kind: :collection, plan: {
name: "row",
attributes: [{ name: "id", kind: :scalar, type: :integer }],
children: [
{ name: "price", kind: :scalar, type: :float },
{ name: "active", kind: :scalar, type: :boolean },
] } },
])
rows = descriptor.materialize(xml_source) # one C call
rows.as_kwarg_hash[:children]["row"] # typed, hydrator-readySame-wire-name rows can partition on attribute values with when: —
AND across pairs, exclusive routing (first matching row wins per
occurrence):
Leptris::XML::Descriptor.build(
name: "r",
children: [
{ name: "item", kind: :collection, when: { "kind" => "a" },
plan: { name: "item", children: [{ name: "name", kind: :scalar }] } },
{ name: "item", kind: :collection, when: { "kind" => "b" },
plan: { name: "item", children: [{ name: "name", kind: :scalar }] } },
])PlanValue#node_kind and #order_index carry document-order
identity for matched rows (and, under :order_spine, for the
interleaved text runs) — ordered/mixed content reconstructs
without fragment re-parsing. PlanValue#as_kwarg_hash is the
bulk hydrator face: one Ruby call per node returns the typed
attribute/children hash (the crossings floor on the 5k-row ISO
fixture is ~4 FFI crossings per row, pinned by spec).
XQuery 1.0 core through Leptris::XML::XQuery (compile once,
evaluate many):
query = Leptris::XML::XQuery.parse(<<~XQ)
declare variable $min := 3;
for $i in //item
where number($i/@qty) > $min
order by $i/@qty descending
return <big>{$i/name/text()}</big>
XQ
query.eval(doc) # plain expressions keep their XPath result type;
# FLWOR results arrive as the sequence channel —
# read them through an aggregate until the engine
# materializes readable sequence itemsSupported: the prolog (declare variable / declare namespace /
declare function local:*), nested for with at positions,
let, where, stable multi-key order by, group by, direct and
computed constructors with attribute value templates, and plain
XPath expression bodies. Known grammar gaps are
tracked upstream.
Leptris::XML.buffer_has_nonstandard_entity?(string) is a
one-pass C pre-scan (for adapter layers): true when the buffer
contains a named entity outside the five predefined ones (or
numeric) — ~3x faster than the equivalent Ruby regex and no
false positives on bare &.
Document#save(path, **opts)-
serialize to a file.
Document#canonicalize(version, inclusive_ns, with_comments:, exclusive:, mode:)(alias#c14n)-
canonical XML.
Element#canonicalize(…)-
subtree canonicalization.
doc.to_xml # one-line, no indent
doc.to_xml(indent: 2) # pretty-printed
doc.canonicalize # C14N 1.0
doc.canonicalize(Leptris::XML::FFI::C14N_1_1) # C14N 1.1
doc.canonicalize(exclusive: true) # Exclusive C14N
doc.canonicalize(with_comments: true) # keep comments
doc.canonicalize(exclusive: true, inclusive_namespaces: ["ds"]) # InclusiveNamespacesFor very large documents, use the streaming SAX parser. Subclass
Leptris::XML::SAX::Document and override the events you care about:
class Counter < Leptris::XML::SAX::Document
attr_reader :elements, :depth
def initialize
@elements = 0
@depth = 0
end
def start_element(name, attrs = [])
@elements += 1
@depth += 1
puts " " * (@depth - 1) + "<#{name}>"
end
def end_element(name)
@depth -= 1
end
def characters(str)
puts " " * @depth + "text: #{str.inspect}" unless str.strip.empty?
end
end
parser = Leptris::XML::SAX::Parser.new(Counter.new)
parser.parse(File.open("huge.xml")) # streams in 4 KB chunksSAX::Parser#parse accepts a String, an IO, or any object responding
to #read. The handler callbacks are:
start_document, end_document
|
document boundaries. |
xmldecl(version, encoding, standalone)
|
XML declaration. |
start_element(name, attrs), end_element(name)
|
element events; |
characters(str), comment(str), cdata_block(str)
|
text-class events. |
processing_instruction(name, content)
|
PI event. |
start_prefix_mapping(prefix, uri), end_prefix_mapping(prefix)
|
namespace events. |
warning(str), error(msg, line, col)
|
recoverable parser messages. |
The parser picks its transport by what your handler overrides. Overriding
one hot kind (say characters) attaches only that callback — the engine
skips C-side emission for the rest entirely. Overriding several rides the
bulk recorder (one C call stages the whole document, then a lean dispatch
loop) — measured on a 1.9 MB document: text-only 21 ms, all-events 119 ms
vs Nokogiri’s 130/131 ms. The handler cannot tell the transports apart;
xmlns declarations ride the attribute pairs and prefix-mapping events
fire alongside.
For raw bulk event streams, SAX::Recorder.parse(xml, kinds:) drains
buffered events per chunk — unwanted kinds cost one array read — and
Recorder#reset reuses one recorder across documents.
Leptris::XML::Pull is the StAX-style cursor: Parser#each delivers
Event structs with types including :start_prefix/:end_prefix (the
default namespace’s prefix is ""); Parser#each_batch(max) delivers
events in bulk with a corruption guard that fails loudly rather than
delivering garbage. Leptris::XML::Iterparse.parse(xml, mode:) yields
completed subtrees top-level (:top_level) or every element post-order
(:full_document) with bounded memory, plus #namespace_uri on the last
yielded element and an #error channel for truncated input. Yielded
elements carry an internal lifetime scope: #document answers nil (the
documented contract) while memoization and liveness guards engage —
using an element after the iteration raises UseAfterFreeError instead
of crashing, and repeated reads within the block ride the same fast
paths as document-backed nodes.
Document is the only object that owns C memory. Everything else
(Element, Text, Attr, NodeSet, …) is a borrowed handle that
is valid only while its Document is alive.
-
Free a document explicitly with
Document#free. After#free, any further method call on the document or its nodes raisesLeptris::XML::UseAfterFreeError. -
If you don’t call
#free, GC will — a finalizer captures the raw pointer address (not the Ruby wrapper) and callsleptris_document_freeexactly once. -
NodeSet`s holding XPath results own their own `LeptrisXPathResultand free it on GC. -
Don’t hold a
Nodereference past the lifetime of itsDocument. The C memory is gone; using the wrapper is undefined behaviour.
Two facts worth knowing for memory-sensitive workloads:
Finalizers drain asynchronously. When a Document becomes
unreachable, its C tree frees when Ruby runs the registered
finalizer — and MRI executes finalizers on its own scheduling, not
synchronously inside GC.start. A tight GC.start loop can leave
the last document’s wrappers (flagged uncollectible) alive
indefinitely; a short wall-clock yield drains the queue:
doc = Leptris::XML::Document.parse(xml)
doc.root.children # ...
doc = nil
GC.start
sleep 0.05 # yield to the finalizer queue
GC.start # now fully collected, C tree freedHeld documents carry the Ruby wrapper layer. A fully walked,
held document costs roughly +552 kB of wrapper objects on top of
the C tree’s +456 kB (about 1.57x Nokogiri for the same held
shape) — the measured cost of the FFI-only, no-compile-at-install
architecture: every wrapped node is a small Ruby object holding an
FFI::Pointer. Parse-and-discard workloads are unaffected (the
tree is the dominant cost and the wrappers die young). See
#147 for the
analysis and the open TypedData options.
All Leptris errors descend from Leptris::XML::Error:
ParseError
|
raised by |
XPathError
|
raised by |
UseAfterFreeError
|
raised when calling methods on a freed |
Error
|
generic (mutation precondition failures, etc.). |
begin
Leptris::XML.parse("<unclosed>")
rescue Leptris::XML::ParseError => e
warn "parse failed: #{e.message}"
endFor most read-only XPath use cases the swap is mechanical:
# Nokogiri
require "nokogiri"
doc = Nokogiri::XML(File.read("doc.xml"))
doc.xpath("//item[@id='1']").each { |n| puts n.text }
# Leptris
require "leptris"
doc = Leptris::XML.parse(File.read("doc.xml"))
doc.xpath("//item[@id='1']").each { |n| puts n.content }Notable differences:
-
Node#textexists but the canonical name is#content(Nokogiri uses both). -
Node#childrenincludes whitespace text nodes (same as Nokogiri); use#element_childrenor#first_element_childto skip them. -
CSS support is intentionally minimal — for advanced selectors, drop to
xpath.cssis receiver-relative (Nokogiri semantics): scoped to an element or fragment, document-wide from a Document. -
DocumentFragmentis searchable (fragment.xpath/at_xpath/css/at_css/search) — Nokogiri fragment parity. -
Expanded-name attribute access:
Element#attribute_ns(uri, local)/#has_attribute_ns?(uri, local)— XML Namespaces 1.0 semantics (cross-prefix match, nil URI matches no-namespace, xmlns invisible). -
Leptris::XML.parse(xml, recover: true)returns an empty document with the failure recorded on the thread-global last error instead of raising ParseError — libxml2XML_PARSE_RECOVERsemantics. The companionDocument#last_error_positionreturns[line, column]. -
Leptris::XML.parse(xml, readonly: true)(orDocument#readonly!) freezes the document for reading: mutations raise ReadOnlyError, read methods memoize aggressively. Faster steady state; no Nokogiri equivalent. -
Lifetime contract: a borrowed handle used after the owning document has been freed (or GC’d) raises
Leptris::XML::UseAfterFreeError. Nokogiri is silent on this — migrating code that holds Node references past Document disposal will see the error; silence-replace-UAF patterns from Nokogiri do not apply. -
No
Nokogiri::CSSparser. HTML parsing is supported (Leptris::HTML/Leptris::HTML4/Leptris::HTML5facades overLeptris::XML.parse_html, libleptris 1.9.75); Nokogiri’s HTML-specific node subclasses have no equivalent. -
RelaxNG and DTD validation ARE supported (
Leptris::XML::RelaxNG,Leptris::XML::DTD— parse once, validate many, structured errors). No W3C XML Schema and no schema caching (XSLT 1.0–3.0 and an XQuery 1.0 core are supported — see above). -
TruffleRuby and JRuby ARE supported out of the box (#160): the
ruby-platform variant vendors precompiled binaries and the binding is FFI-based (no MRI C extension required to load).
Head-to-head against published Nokogiri 1.19.4 on a 1.86 MB / 250k-event document (arm64-darwin, CPU totals, best-of-5):
| DOM parse |
10–12x faster |
| CSS search |
4.0–4.8x |
| XPath nodeset |
2.4–3.0x |
XPath scalar (string(//item[1]))
|
4.6x |
| Serialization |
2.2–2.3x |
at_css
|
1.4x |
| SAX text-only handler |
6x (119 ms all-events vs Nokogiri’s 131) |
| Memory held |
17.5 MB/doc vs Nokogiri’s 31.3 (1.8x lighter) |
Two rows sit at the Ruby allocation floor by documented choice: the
cold full-tree walk (children recursion) remains ~1.5x behind
Nokogiri’s C-extension node creation — Node#visit is the wrap-free
lever when that matters.
bundle install # install Ruby deps
bundle exec rspec # full test suite (229 specs)
bundle exec rspec spec/xml/xpath_spec.rb:42 # one example by line
bundle exec rubocop # lintCI pins libleptris to a released tag (currently v1.1.1) and builds it
from source on each runner; see .github/workflows/build.yml.
MIT — see LICENSE.