Skip to content

Repository files navigation

leptris-ruby

RubyGems Version CI

A Nokogiri-compatible Ruby binding for libleptris, a pure-C99 XML 1.0 parser with full XPath 1.0, XML Namespaces 1.0, SAX, and C14N (1.0 / 1.1 / Exclusive), plus XSLT 1.0–3.0 transforms, an XQuery 1.0 core face, and a growing XPath 2/3.1 expression subset available standalone from #xpath itself.

The C DOM is the single source of truth — Ruby objects are thin FFI handles over the C pointers, so every Ruby method maps to one FFI call. No tree hydration, no parallel Ruby-side model.

Installation

Add to your Gemfile:

gem "leptris"

Then bundle install.

Runtime requirement: libleptris

leptris shells out to the native libleptris shared library via FFI. You need libleptris.{dylib,so,dll} installed on the host. Options:

  1. Homebrew (macOS, easiest): brew install lutaml/tap/libleptris (if packaged) or build from source (see below).

  2. Build from source (Linux/macOS/Windows):

    git clone https://github.com/leptris/leptris.git
    cd leptris
    cmake -B build -S . \
      -DCMAKE_BUILD_TYPE=Release \
      -DLEPTRIS_BUILD_SHARED=ON \
      -DLEPTRIS_BUILD_STATIC=OFF \
      -DCMAKE_WINDOWS_EXPORT_ALL_SYMBOLS=ON
    cmake --build build -j
    sudo cmake --install build   # optional, system-wide
  3. Point Leptris at a specific path by setting LEPTRIS_LIB_PATH:

    export LEPTRIS_LIB_PATH=/usr/local/lib/libleptris.dylib

If Leptris can’t find the library at startup, every parse call raises LoadError.

Alternative engines: TruffleRuby and JRuby (#160) — zero setup

The ruby-platform variant (what TruffleRuby and JRuby resolve) vendors precompiled libleptris + libutf8proc binaries for the common engine platforms under lib/leptris/vendor/:

  • arm64-darwin, x86_64-darwin

  • x86_64-linux, aarch64-linux (glibc) and x86_64-linux-musl, aarch64-linux-musl (tried as fallback for Alpine hosts)

  • arm-linux, arm-linux-musl (32-bit ARM EABI5)

  • ppc64le-linux, s390x-linux, s390x-linux-musl (POWER8+ LE and IBM Z, qemu-built)

gem install leptris just works on both engines — the FFI layer selects the matching vendored binary at require time (the binding is FFI-based; no MRI C extension anywhere). LEPTRIS_LIB_PATH still overrides, and a system library remains the fallback for OSes the variant does not carry.

Parsing

The top-level entry point is Leptris::XML. Parse a string or an IO:

require "leptris"

doc = Leptris::XML.parse(<<~XML)
  <library xmlns="http://example.org/ns">
    <book id="b1" lang="en">
      <title>Refactoring</title>
      <author>Martin Fowler</author>
    </book>
    <book id="b2" lang="fr">
      <title>Programmer en Ruby</title>
    </book>
  </library>
XML

doc.root.name        # => "library"
doc.root.children.size   # => 5 (2 element children + 3 whitespace text nodes)

Or a file:

doc = Leptris::XML.parse_file("books.xml")

This is the direct Nokogiri equivalent of Nokogiri::XML(…​). The returned object is a Leptris::XML::Document.

Malformed input raises Leptris::XML::ParseError. Parse with recover: true to get libxml2’s XML_PARSE_RECOVER semantics instead — an empty document back, with the failure recorded on the thread-global last error:

doc = Leptris::XML.parse("<broken", recover: true)
doc.root                                # => nil
Leptris::XML::FFI.leptris_last_error    # => "..."

HTML parsing

Leptris::XML.parse_html(html) parses tolerant HTML4/5 into a standard Document — the same nodes, pool, serializer, and XPath/XSLT/XQuery machinery as XML (libleptris 1.9.75, the last Nokogiri capability gap). The Leptris::HTML, Leptris::HTML4, and Leptris::HTML5 module facades provide the Nokogiri-shaped entry points over the same engine (Leptris::HTML5.parse(html), Leptris::HTML(html)):

doc = Leptris::XML.parse_html(%(<ul><li>a<li>b</ul>))
doc.at_css("body").inner_html   # => "<ul><li>a</li><li>b</li></ul>"

Implied end tags (p/li/td/tr/…​), void elements, raw-text <script>/<style>, case-insensitive lowercased names, minimized/unquoted attributes, and the HTML named entities all parse; html/body are synthesized (an empty head is not) and tbody is never implied. Malformed markup degrades to text rather than raising.

Reading nodes

Document#root

root Element, or nil for an empty document.

Node#name

element name (e.g. "book").

Node#content (alias #text, #inner_text)

all descendant text concatenated.

Node#[] (alias #attr, #get_attribute)

attribute value by name.

Node#attributes

hash of {name ⇒ Attr}.

Node#key? (alias #has_attribute?)

attribute presence.

Node#children

NodeSet of all children (elements, text, comments, …).

Node#element_children

NodeSet of element children only.

Node#first_element_child, #last_element_child

first/last element child (skip text nodes).

Node#next_element, #previous_element

next/prev sibling element.

Node#parent, #next_sibling, #previous_sibling

tree navigation.

Node#line

1-based source line number.

Node#type (alias #node_type)

integer type code. Element predicates: #element?, #text?, #comment?, #cdata?, #processing_instruction?.

Example — walk all book titles:

doc.root.children.select(&:element?).each do |book|
  title = book.children.find { |c| c.element? && c.name == "title" }
  puts "#{book[:id]}: #{title&.content}"
end
# b1: Refactoring
# b2: Programmer en Ruby

Tree iteration

Node#visit walks the subtree with ONE C call and C-tracked depth — elements yield (node, entering, depth) enter/leave pairs, every other kind yields once; the leanest full-subtree iteration the binding offers:

doc.root.visit do |node, entering, depth|
  puts "#{'  ' * depth}#{node.name} #{entering ? 'enter' : 'leave'}"
end

Node#traverse walks the subtree in post-order via a single C-side callback (one FFI call for the whole traversal, not one per node):

doc.root.traverse do |node|
  case node
  when Leptris::XML::Element   then puts "E  #{node.name}"
  when Leptris::XML::Text      then puts "T  #{node.content.inspect}"
  when Leptris::XML::Comment   then puts "C  #{node.content.inspect}"
  end
end

The document chain

Document#node is the document-level navigation head (the libxml2 model): Document#children reads [prolog comments/PIs, the root, epilog comments/PIs] in document order, with whitespace kept per libxml2’s exact rule. Document-level PIs are first-class: Document#add_pi(target, data) appends one, Document#remove_pi(target_or_index) removes by target or index, and PI#target=/PI#data=/PI#unlink mutate and detach. Document#processing_instructions and Document#comments remain the memoized readers.

Readonly mode

For the dominant parse-query-serialize workload, parse with readonly: true (or call Document#readonly!, one-way):

doc = Leptris::XML.parse(xml, readonly: true)
doc.root.children.first["id"]        # reads are memoized
doc.root.name                        # plain ivar after first call
doc.root << doc.create_element("x")  # raises ReadOnlyError

Reads (name, content, children, attributes) memoize aggressively — they cannot go stale because mutation is forbidden. Every mutator raises Leptris::XML::ReadOnlyError. Detached factories (create_element and friends) still work: building a new tree against a readonly document is legal; mutating the frozen one is not.

Searching: XPath and CSS

Document, Element, and DocumentFragment (via Leptris::XML::Searchable) support:

#xpath(*exprs)

evaluate XPath; returns NodeSet, true/false, Float, or String depending on the expression.

#at_xpath(*exprs)

first match (or scalar), like xpath(*exprs).first.

#css(*selectors)

minimal CSS-to-XPath translation, then xpath. Receiver-relative (Nokogiri semantics): scoped to the element or fragment, document-wide from a Document.

#at_css(*selectors)

first match of css.

#search(*exprs)

dispatches on syntax — path-prefixed expressions (/, ., ..) go to xpath; everything else translates as CSS (comma unions included).

#at(*exprs)

first match of search.

doc.xpath("//book")                       # => NodeSet of both <book>
doc.xpath("count(//book)")                # => 2.0
doc.xpath("//book[@lang='fr']/title")     # => NodeSet[<title>Programmer en Ruby</title>]
doc.at_xpath("//book[@id='b1']")          # => <book id="b1" ...>
doc.at_xpath("string(//book[1]/@id)")     # => "b1"

doc.css("book[lang='en'] title")          # => NodeSet[<title>Refactoring</title>]
doc.at_css("book#b1 title")               # => <title>Refactoring</title>  (id selector)
doc.css("book:first-child")               # first <book>

# css is receiver-relative: scoped to the receiver, not the document
doc.root.at_css("book").css("title")      # titles under THAT book only
frag = doc.fragment("<a x='1'><n/></a>")
frag.css("a > n")                         # searches the fragment

XPath result type follows XPath 1.0 semantics: count(…​) → Float, boolean(…​) → true/false, string(…​) → String, otherwise a Leptris::XML::NodeSet.

Beyond XPath 1.0, the engine accepts a growing XPath 2/3.1 expression subset standalone — no stylesheet required:

doc.xpath("let $n := count(//book) return $n + 1")   # => 3.0
doc.xpath("for $b in //book return count($b/*)")     # sequence
doc.xpath("if (count(//book) = 2) then 'two' else '?'")
doc.xpath("count(1 to 4)")                           # => 4.0
doc.xpath("//book => count()")                       # 3.1 arrow
doc.xpath("//title ! string(.)")                     # 3.1 simple map
doc.xpath("'v' || count(//book)")                    # 3.1 concat
doc.xpath("'42' castable as xs:integer")             # 2.0 type ops
doc.xpath("1.9 cast as xs:integer")                  # => 1.0 (truncates)
doc.xpath("//title instance of node()+")             # 2.0 sequence types
doc.xpath("map { 'b': 'beta' }?b")                   # 3.1 maps
doc.xpath("[10, 20, 30]?2")                          # 3.1 arrays
doc.xpath(%q{serialize(parse-json('{"b":"beta"}'), map { 'method': 'json' })})
doc.xpath("function($x) { $x + 1 }(41)")             # 3.1 function items
doc.xpath("fold-left(1 to 4, 0, function($a, $b) { $a + $b })")
doc.xpath("count(//a) eq 2")                        # 2.0 value comparators
doc.xpath("some $a in //a satisfies $a/@v > 1")     # 2.0 quantifiers
doc.xpath("count(//a except //a[@v = 1])")          # 2.0 set algebra

Sequence and constructor items arrive as Leptris::XML::ResultText objects — #content serves the value directly:

doc.xpath("for $w in //a return string($w/@v)").map(&:content)  # => ["1", "2"]

Remaining grammar gaps (tracked upstream): 3.1 string templates, XQuery beyond the 1.0 core.

Supported CSS selectors

Minimal subset (translated to XPath via Leptris::XML::CssToXPath):

  • Type/universal: book, *

  • Class/ID: .highlight, #b1

  • Attribute presence: [lang]

  • Attribute value: [lang='en'], [lang~='en'], [lang^='en'], [lang$='en'], [lang*='en']

  • Combinators: descendant (space), child (>), comma (multi-selector)

  • Pseudo-classes: :first-child, :last-child, :only-child, :empty, :root, :not(…​)

For anything more sophisticated, drop down to xpath.

Building and mutating

Documents expose factory methods; elements expose mutation methods:

doc  = Leptris::XML.parse("<root/>")
book = doc.create_element("book")
book[:id] = "b3"
book.add_child(doc.create_element("title")).content = "New book"
doc.root.add_child(book)

puts doc.to_xml
# <?xml version="1.0"?>
# <root><book id="b3"><title>New book</title></book></root>
Document#create_element(name)

detached element owned by the document.

Document#create_text_node(str), #create_comment(str), #create_cdata(str)

text-class factories.

Document#create_processing_instruction(target, data)

PI factory.

Document#fragment(markup)

parse a markup fragment (multiple top-level children allowed).

Element#name=, #content=

rename / replace inner text.

Element#[]= (alias #set_attribute)

add/update an attribute.

Element#remove_attribute (alias #delete)

drop an attribute.

Element#add_child(node_or_markup) (alias #<<)

append a Node, or parse+append a markup String.

Element#prepend_child(node)

insert as the first child.

Element#add_next_sibling(node), #add_previous_sibling(node)

sibling insertion.

Element#remove_child(node)

detach (does not free).

Element#children=

replace all children.

Element#replace(node) / #swap(node)

replace in parent.

Element#wrap(node_or_markup)

wrap this element in a new one.

Node#unlink

detach from the tree.

Element#attribute_pairs

bulk read-only attribute listing — [[name, value], …​] in document order, duplicates included; one C crossing on the native layer (libleptris 1.9.210 era).

Document diffing (Leptris::XML.diff(a, b)) yields an op list with #ops, per-kind counts via #summary ({update_attr: 1, insert: 1}, {} when identical), and #to_json for the op list — the leptris diff --summary/--json CLI modes' library face.

Building from scratch (no parse)

Document.create starts an empty document; build up from there (document-level prolog parts included — PIs, comments, the XML declaration, and the DOCTYPE are all first-class):

doc   = Leptris::XML::Document.create
book  = doc.create_element("book")
book[:id] = "b3"
book << doc.create_element("title")
doc.root = doc.create_element("catalog")
doc.root << book
doc.add_comment("built programmatically")   # epilog comment
doc.set_doctype("catalog", system_id: "catalog.dtd")

puts doc.to_xml
Document.create

empty document (no root yet).

Document#root=

attach the root element (bottom-up construction).

Document#set_doctype(name, public_id:, system_id:)

programmatic DOCTYPE.

Document#doctype (alias #internal_subset)

the DocType reader (name/root_name/public_id/system_id/internal_subset).

Document#clear_declaration

un-set the XML declaration (serialize as if the input had none; idempotent).

Document#remove_doctype

un-set the DOCTYPE (true when removed; the DocType stays readable until #free).

Document#add_pi(target, data) / #remove_pi(target_or_index)

document-level processing instructions.

Document#add_comment(str)

epilog comment.

Document#create_entity_reference(name)

&name; node, serializes verbatim once attached.

Namespaces

Element#namespace

the element’s in-scope namespace as a Namespace (or nil).

Element#namespaces

all in-scope namespaces (inherited from ancestors) as a {prefix_or_xmlns ⇒ href} hash.

Element#namespace_definitions

only namespaces declared directly on this element.

Element#add_namespace_definition(prefix, href) (alias #add_namespace)

declare xmlns:prefix="href" on this element.

Element#default_namespace=(href)

declare/replace xmlns="href".

Element#namespace=(uri)

Nokogiri node.namespace= semantics (libleptris 1.9.76): nil detaches — prefix clears, xmlns="" blocks in-scope defaults; a URI rebinds to an in-scope declaration carrying it (adopting its prefix; raises when none is in scope — declare first).

Element#remove_namespace_definition(prefix)

drop a declaration.

Element#attribute_ns(uri, local)

attribute value by expanded name (URI + local); nil URI matches no-namespace attributes only.

Element#has_attribute_ns?(uri, local)

presence by expanded name.

Attr#prefix

the attribute’s prefix as written (nil when none).

Attr#namespace_uri

resolved through the owning element’s in-scope declarations at read time (xml prebound; nil for undeclared prefixes).

root = doc.root
root.add_namespace_definition("t", "https://example.org/types")
puts root.namespaces
# {"xmlns"=>"http://example.org/ns", "xmlns:t"=>"https://example.org/types"}

# XPath with prefixes is dispatched straight to libleptris, which resolves
# prefixes using the in-scope namespace declarations.
doc.xpath("//t:title")

Serialization and canonicalization

Document#to_xml(indent: 0, no_decl: false, encoding: nil, indent_text: false) (aliases #to_s, #serialize)

serialize the whole document.

Element#to_xml(…​)

serialize a subtree (also takes indent_text: — the unit string).

indent_text carries two meanings: a STRING is the indent unit with Nokogiri’s semantics (the unit replaces the default spaces, repeated indent times per depth level — byte-identical to Nokogiri’s output); true selects the display form, which also indents text and mixed content:

doc.to_xml(indent: 2, indent_text: "\t")   # tab-indented
doc.to_xml(indent: 2, indent_text: true)    # display form (documents only)

Element#inner_html serializes the children with correct escaping — well-formed by construction (a re-parse spec pins it).

Element#to_xml(expand_empty: true) emits <a></a> instead of <a/> for empty elements — libxml2’s XML_SAVE_NO_EMPTY_TAGS parity.

Node#digest(drop_ws: false) answers a content-defined 64-bit Merkle hash of the subtree: equal flags and equal digests imply structural equivalence (names, namespaces, sorted attributes, document-order children); inequality implies nothing — descend and decide. Stable across processes; drop_ws: true skips whitespace-only text nodes.

XSLT 1.0–3.0 transforms run through Leptris::XML::XSLT:

style = Leptris::XML::XSLT.parse(stylesheet_xml)   # or .parse_file
result = style.apply_to(doc)                       # => Document
style.serialize(doc)                               # => String

The engine dispatches on the stylesheet’s declared version. XSLT 1.0 is complete; the 3.0 instruction set has grown through libleptris 1.9.36 — grouping, xsl:accumulator, xsl:analyze-string, xsl:mode/@on-no-match dispositions, xsl:sequence, xsl:perform-sort, and fn:format-integer — with the XPath 3.1 core expressions available inside transforms (see the subset list under Searching: XPath and CSS; value comparators and sequence types are still out, tracked upstream). Known engine bug: the shallow-skip and text-only-copy dispositions drop unmatched subtrees in the built-in initial descent (#705) — the other four dispositions are spec-correct. XQuery has no entry point yet. == Descriptor materialization

Compile a schema descriptor once, then materialize a whole subtree against it in ONE native pass — no per-element Ruby calls (frameworks rebuilding typed models from XML; libleptris 1.9.162):

descriptor = Leptris::XML::Descriptor.build(
  name: "catalog",
  attributes: [{ name: "version", kind: :scalar }],
  children: [
    { name: "item", kind: :nested, plan: {
      name: "item",
      attributes: [{ name: "id", kind: :scalar }],
      children: [
        { name: "name", kind: :scalar },
        { name: "price", kind: :scalar },
        { name: "opt", kind: :collection },
      ] } },
  ])

tree = descriptor.walk(doc.root).to_ruby
# { kind: :element, attributes: { "version" => "2.0" }, children: [
#   { kind: :element, name: "item", attributes: { "id" => "1" },
#     children: ["first", "1.99", ["a", "b"]] }, ...] }

Row kinds: :scalar, :collection, :nested (via plan:), :raw (serialized subtree), :content (mixed-content text runs — pair with flags: [:mixed_content]), :callback (raw value + document byte offset + type_tag echo, matching Node#byte_offset). Namespace binding lives on plans: ns: is :none (default), :any, or { exact: "urn:…​" }; flags: also accepts :cdata, :ordered, :ns_lenient (#754 out-of-namespace adoption), and :order_spine (#1273 — unmatched sibling text/comment/PI runs emit as positioned scalars so ordered hosts rebuild element order without re-parsing). Children the plan does not describe are skipped; undescribed document order within a row is preserved. #walk returns a lazy PlanValue tree (the result outlives the document); #to_ruby materializes it.

Typed scalars and the fused entry (#230/#298)

Plan rows accept type: :integer | :float | :boolean — the walk parses the values natively (in-pass, libleptris #1269a) and #to_ruby / the bulk face return Integer / Float / true / false. Unparseable values fall back to the raw String; PlanValue#string_value stays the raw escape, and #int_value / #float_value / #bool_value expose the parsed accessors directly.

Descriptor#materialize(source) is the fused entry: source bytes → typed rows in ONE C call — parse, walk, and free happen natively (#1269b); the standalone PlanValue tree is all that surfaces.

descriptor = Leptris::XML::Descriptor.build(
  name: "iso",
  children: [{ name: "row", kind: :collection, plan: {
    name: "row",
    attributes: [{ name: "id", kind: :scalar, type: :integer }],
    children: [
      { name: "price", kind: :scalar, type: :float },
      { name: "active", kind: :scalar, type: :boolean },
    ] } },
  ])

rows = descriptor.materialize(xml_source)   # one C call
rows.as_kwarg_hash[:children]["row"]        # typed, hydrator-ready

Predicate rows (#1272) and order identity (#1273)

Same-wire-name rows can partition on attribute values with when: — AND across pairs, exclusive routing (first matching row wins per occurrence):

Leptris::XML::Descriptor.build(
  name: "r",
  children: [
    { name: "item", kind: :collection, when: { "kind" => "a" },
      plan: { name: "item", children: [{ name: "name", kind: :scalar }] } },
    { name: "item", kind: :collection, when: { "kind" => "b" },
      plan: { name: "item", children: [{ name: "name", kind: :scalar }] } },
  ])

PlanValue#node_kind and #order_index carry document-order identity for matched rows (and, under :order_spine, for the interleaved text runs) — ordered/mixed content reconstructs without fragment re-parsing. PlanValue#as_kwarg_hash is the bulk hydrator face: one Ruby call per node returns the typed attribute/children hash (the crossings floor on the 5k-row ISO fixture is ~4 FFI crossings per row, pinned by spec).

XQuery

XQuery 1.0 core through Leptris::XML::XQuery (compile once, evaluate many):

query = Leptris::XML::XQuery.parse(<<~XQ)
  declare variable $min := 3;
  for $i in //item
  where number($i/@qty) > $min
  order by $i/@qty descending
  return <big>{$i/name/text()}</big>
XQ
query.eval(doc)   # plain expressions keep their XPath result type;
                  # FLWOR results arrive as the sequence channel —
                  # read them through an aggregate until the engine
                  # materializes readable sequence items

Supported: the prolog (declare variable / declare namespace / declare function local:*), nested for with at positions, let, where, stable multi-key order by, group by, direct and computed constructors with attribute value templates, and plain XPath expression bodies. Known grammar gaps are tracked upstream.

Leptris::XML.buffer_has_nonstandard_entity?(string) is a one-pass C pre-scan (for adapter layers): true when the buffer contains a named entity outside the five predefined ones (or numeric) — ~3x faster than the equivalent Ruby regex and no false positives on bare &.

Document#save(path, **opts)

serialize to a file.

Document#canonicalize(version, inclusive_ns, with_comments:, exclusive:, mode:) (alias #c14n)

canonical XML.

Element#canonicalize(…​)

subtree canonicalization.

doc.to_xml                       # one-line, no indent
doc.to_xml(indent: 2)            # pretty-printed
doc.canonicalize                # C14N 1.0
doc.canonicalize(Leptris::XML::FFI::C14N_1_1)              # C14N 1.1
doc.canonicalize(exclusive: true)                         # Exclusive C14N
doc.canonicalize(with_comments: true)                     # keep comments
doc.canonicalize(exclusive: true, inclusive_namespaces: ["ds"])  # InclusiveNamespaces

SAX parsing

For very large documents, use the streaming SAX parser. Subclass Leptris::XML::SAX::Document and override the events you care about:

class Counter < Leptris::XML::SAX::Document
  attr_reader :elements, :depth
  def initialize
    @elements = 0
    @depth    = 0
  end

  def start_element(name, attrs = [])
    @elements += 1
    @depth    += 1
    puts "  " * (@depth - 1) + "<#{name}>"
  end

  def end_element(name)
    @depth -= 1
  end

  def characters(str)
    puts "  " * @depth + "text: #{str.inspect}" unless str.strip.empty?
  end
end

parser = Leptris::XML::SAX::Parser.new(Counter.new)
parser.parse(File.open("huge.xml"))   # streams in 4 KB chunks

SAX::Parser#parse accepts a String, an IO, or any object responding to #read. The handler callbacks are:

start_document, end_document

document boundaries.

xmldecl(version, encoding, standalone)

XML declaration.

start_element(name, attrs), end_element(name)

element events; attrs is an array of [name, value] pairs in source order.

characters(str), comment(str), cdata_block(str)

text-class events.

processing_instruction(name, content)

PI event.

start_prefix_mapping(prefix, uri), end_prefix_mapping(prefix)

namespace events.

warning(str), error(msg, line, col)

recoverable parser messages.

Transports: interest-proportional delivery

The parser picks its transport by what your handler overrides. Overriding one hot kind (say characters) attaches only that callback — the engine skips C-side emission for the rest entirely. Overriding several rides the bulk recorder (one C call stages the whole document, then a lean dispatch loop) — measured on a 1.9 MB document: text-only 21 ms, all-events 119 ms vs Nokogiri’s 130/131 ms. The handler cannot tell the transports apart; xmlns declarations ride the attribute pairs and prefix-mapping events fire alongside.

For raw bulk event streams, SAX::Recorder.parse(xml, kinds:) drains buffered events per chunk — unwanted kinds cost one array read — and Recorder#reset reuses one recorder across documents.

Pull parsing and iterparse

Leptris::XML::Pull is the StAX-style cursor: Parser#each delivers Event structs with types including :start_prefix/:end_prefix (the default namespace’s prefix is ""); Parser#each_batch(max) delivers events in bulk with a corruption guard that fails loudly rather than delivering garbage. Leptris::XML::Iterparse.parse(xml, mode:) yields completed subtrees top-level (:top_level) or every element post-order (:full_document) with bounded memory, plus #namespace_uri on the last yielded element and an #error channel for truncated input. Yielded elements carry an internal lifetime scope: #document answers nil (the documented contract) while memoization and liveness guards engage — using an element after the iteration raises UseAfterFreeError instead of crashing, and repeated reads within the block ride the same fast paths as document-backed nodes.

Memory model

Document is the only object that owns C memory. Everything else (Element, Text, Attr, NodeSet, …) is a borrowed handle that is valid only while its Document is alive.

  • Free a document explicitly with Document#free. After #free, any further method call on the document or its nodes raises Leptris::XML::UseAfterFreeError.

  • If you don’t call #free, GC will — a finalizer captures the raw pointer address (not the Ruby wrapper) and calls leptris_document_free exactly once.

  • NodeSet`s holding XPath results own their own `LeptrisXPathResult and free it on GC.

  • Don’t hold a Node reference past the lifetime of its Document. The C memory is gone; using the wrapper is undefined behaviour.

Memory behavior

Two facts worth knowing for memory-sensitive workloads:

Finalizers drain asynchronously. When a Document becomes unreachable, its C tree frees when Ruby runs the registered finalizer — and MRI executes finalizers on its own scheduling, not synchronously inside GC.start. A tight GC.start loop can leave the last document’s wrappers (flagged uncollectible) alive indefinitely; a short wall-clock yield drains the queue:

doc = Leptris::XML::Document.parse(xml)
doc.root.children   # ...
doc = nil
GC.start
sleep 0.05          # yield to the finalizer queue
GC.start            # now fully collected, C tree freed

Held documents carry the Ruby wrapper layer. A fully walked, held document costs roughly +552 kB of wrapper objects on top of the C tree’s +456 kB (about 1.57x Nokogiri for the same held shape) — the measured cost of the FFI-only, no-compile-at-install architecture: every wrapped node is a small Ruby object holding an FFI::Pointer. Parse-and-discard workloads are unaffected (the tree is the dominant cost and the wrappers die young). See #147 for the analysis and the open TypedData options.

Errors

All Leptris errors descend from Leptris::XML::Error:

ParseError

raised by parse / parse_file / SAX on malformed input.

XPathError

raised by xpath on malformed or unsupported expressions.

UseAfterFreeError

raised when calling methods on a freed Document.

Error

generic (mutation precondition failures, etc.).

begin
  Leptris::XML.parse("<unclosed>")
rescue Leptris::XML::ParseError => e
  warn "parse failed: #{e.message}"
end

Migrating from Nokogiri

For most read-only XPath use cases the swap is mechanical:

# Nokogiri
require "nokogiri"
doc = Nokogiri::XML(File.read("doc.xml"))
doc.xpath("//item[@id='1']").each { |n| puts n.text }

# Leptris
require "leptris"
doc = Leptris::XML.parse(File.read("doc.xml"))
doc.xpath("//item[@id='1']").each { |n| puts n.content }

Notable differences:

  • Node#text exists but the canonical name is #content (Nokogiri uses both).

  • Node#children includes whitespace text nodes (same as Nokogiri); use #element_children or #first_element_child to skip them.

  • CSS support is intentionally minimal — for advanced selectors, drop to xpath. css is receiver-relative (Nokogiri semantics): scoped to an element or fragment, document-wide from a Document.

  • DocumentFragment is searchable (fragment.xpath/at_xpath/css/at_css/search) — Nokogiri fragment parity.

  • Expanded-name attribute access: Element#attribute_ns(uri, local) / #has_attribute_ns?(uri, local) — XML Namespaces 1.0 semantics (cross-prefix match, nil URI matches no-namespace, xmlns invisible).

  • Leptris::XML.parse(xml, recover: true) returns an empty document with the failure recorded on the thread-global last error instead of raising ParseError — libxml2 XML_PARSE_RECOVER semantics. The companion Document#last_error_position returns [line, column].

  • Leptris::XML.parse(xml, readonly: true) (or Document#readonly!) freezes the document for reading: mutations raise ReadOnlyError, read methods memoize aggressively. Faster steady state; no Nokogiri equivalent.

  • Lifetime contract: a borrowed handle used after the owning document has been freed (or GC’d) raises Leptris::XML::UseAfterFreeError. Nokogiri is silent on this — migrating code that holds Node references past Document disposal will see the error; silence-replace-UAF patterns from Nokogiri do not apply.

  • No Nokogiri::CSS parser. HTML parsing is supported (Leptris::HTML / Leptris::HTML4 / Leptris::HTML5 facades over Leptris::XML.parse_html, libleptris 1.9.75); Nokogiri’s HTML-specific node subclasses have no equivalent.

  • RelaxNG and DTD validation ARE supported (Leptris::XML::RelaxNG, Leptris::XML::DTD — parse once, validate many, structured errors). No W3C XML Schema and no schema caching (XSLT 1.0–3.0 and an XQuery 1.0 core are supported — see above).

  • TruffleRuby and JRuby ARE supported out of the box (#160): the ruby-platform variant vendors precompiled binaries and the binding is FFI-based (no MRI C extension required to load).

Performance

Head-to-head against published Nokogiri 1.19.4 on a 1.86 MB / 250k-event document (arm64-darwin, CPU totals, best-of-5):

DOM parse

10–12x faster

CSS search

4.0–4.8x

XPath nodeset

2.4–3.0x

XPath scalar (string(//item[1]))

4.6x

Serialization

2.2–2.3x

at_css

1.4x

SAX text-only handler

6x (119 ms all-events vs Nokogiri’s 131)

Memory held

17.5 MB/doc vs Nokogiri’s 31.3 (1.8x lighter)

Two rows sit at the Ruby allocation floor by documented choice: the cold full-tree walk (children recursion) remains ~1.5x behind Nokogiri’s C-extension node creation — Node#visit is the wrap-free lever when that matters.

Development

bundle install                       # install Ruby deps
bundle exec rspec                    # full test suite (229 specs)
bundle exec rspec spec/xml/xpath_spec.rb:42   # one example by line
bundle exec rubocop                  # lint

CI pins libleptris to a released tag (currently v1.1.1) and builds it from source on each runner; see .github/workflows/build.yml.

License

MIT — see LICENSE.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages