From unfamiliar XML to analysis-ready R tables.
xmlrectr is an ambitious attempt at a universal XML
rectangler for R.
It is not trying to be a universal XML parser. Excellent XML parsers already exist. The problem addressed here is different: how do you turn hierarchical XML into a table or tibble that is genuinely useful for analysis with the tidyverse, base R, Arrow, and other tools built around tabular data?
That problem becomes especially difficult when you are simply handed
an XML file and know little or nothing about it. You may have no schema,
no documentation, no knowledge of the XML vocabulary, and no predefined
extraction rules. xmlrectr can inspect the actual document
structure and, in many cases, construct a workable analyst-friendly
tibble automatically.
When you do know the XML structure, have documentation or an XSD, or know exactly what you want to expose, the package becomes more explicit rather than less useful. You can review structural proposals, provide row and identifier choices, select and rename fields, use schema evidence, save reusable profiles, and obtain a rectangle closely aligned with your analytical purpose.
The package is deliberately generic. It contains no special rules for MARC, FHIR, EAD, ONIX, GPX, UBL, or any other XML vocabulary.
At present, xmlrectr can be installed from GitHub.
The recommended method is pak:
install.packages("pak")
pak::pak("larry77/xmlrectr")Alternatively:
install.packages("remotes")
remotes::install_github("larry77/xmlrectr")Because xmlrectr contains native C code, the GitHub
version is compiled from source and requires a working build toolchain
and libxml2 development files.
Windows is a first-class supported platform. Use the Rtools version matching your R installation and install it in the default location.
For the currently supported R series:
With a normal R/Rtools installation, manual PATH
configuration should not usually be necessary.
If compilation fails, first check the toolchain:
install.packages("pkgbuild")
pkgbuild::has_build_tools(debug = TRUE)Avoid mixing Rtools with unrelated MSYS2/MinGW/Strawberry Perl
toolchains on PATH, as this can produce
difficult-to-diagnose linking problems.
The GitHub Actions workflow compiles and tests the package on
windows-latest.
On Debian/Ubuntu:
sudo apt install libxml2-dev pkg-configThen install xmlrectr from R using pak or
remotes.
With Homebrew:
brew install libxml2 pkg-config
export PKG_CONFIG_PATH="$(brew --prefix libxml2)/lib/pkgconfig:$PKG_CONFIG_PATH"Then install xmlrectr from R.
For a local checkout:
pak::local_install(".")XML is hierarchical. Analytical data is usually rectangular.
A naive “flatten everything” strategy can easily:
There is also an unavoidable conceptual limit: an arbitrary XML document does not have one mathematically unique tabular interpretation. Domain knowledge can always improve a rectangle.
xmlrectr therefore does not claim to infer the intended
semantics of every XML document. Instead, it aims to make a useful,
conservative structural interpretation when knowledge is scarce, while
exposing progressively more control when the user knows more.
The core design is:
unknown XML
|
+--> automatic analyst-oriented rectangle
|
+--> structural proposal
|
v
human review
|
v
reusable profile
|
v
compiled rectangle
XSD information can contribute evidence, but it is advisory rather than blindly authoritative.
A major goal of xmlrectr is to automate both the
analytical problem and the computational problem, while keeping
both layers tunable.
For exploration, rectangle_xml_analyst() can start from
the XML file itself:
library(xmlrectr)
file <- system.file("extdata", "orders.xml", package = "xmlrectr")
analyst <- rectangle_xml_analyst(file)
analystThe result is one self-contained tibble with:
xml_* provenance and entity columns;This is the deliberately ambitious part of the package: starting from unfamiliar XML and attempting to produce something an R analyst can immediately inspect and work with.
For one-off exploration, this may be all you need.
For repeated or production use, once you understand the structure you will usually want to move to an explicit profile so that the intended rectangle becomes a stable, reviewable contract.
Once a rectangle is defined, execution has its own set of choices: sequential or parallel processing, worker count, chunk size, task size, memory ownership, and whether a large document should be streamed instead of fully materialised.
Those choices are intentionally separated from the analytical meaning of the rectangle.
For the normal rectangling APIs, one argument can ask
xmlrectr to choose whether parallel execution is
worthwhile:
out <- rectangle_xml(
file,
profile,
parallel = "auto"
)Automatic execution planning is based on the observed structural workload and available resources, not on XML vocabulary names, filenames, or rules tuned to the validation corpus. The package does not apply one fixed worker/chunk recipe to every document.
Advanced controls remain available, but they are optional. The point of the architecture is that you should not need to turn every screw before getting useful work done on a large, nested or unfamiliar XML file.
xmlrectr is designed to support a continuum of prior
knowledge.
Start with the automatic analyst table:
analyst <- rectangle_xml_analyst(file)This is the quickest route from an unfamiliar document to an R tibble.
Ask the package for a proposal:
proposal <- propose_xml_profile(file)
review_xml_proposal(proposal, "rows")
review_xml_proposal(proposal, "ids")
review_xml_proposal(proposal, "fields")The proposal is evidence, not an executable command. It lets the package inspect the XML and suggest plausible structural choices without pretending that software can know your analytical intent.
Define it explicitly:
profile <- xml_profile(
rows = "order",
id = "id"
)
out <- rectangle_xml(file, profile)A profile can also select and rename fields, specify types, provide namespace information, and choose the desired layout.
You can save the reviewed decision:
write_xml_profile(profile, "orders-profile.json")
profile2 <- read_xml_profile("orders-profile.json")and reuse it across files belonging to the same XML family.
For repeated processing, compile the profile once:
spec <- compile_xml_profile(profile, file)
out <- rectangle_xml(
file,
spec,
parallel = "auto"
)This is where domain knowledge pays off: when you know the XML structure and the analytical question, the package can produce a rectangle much more closely aligned with your desiderata than any fully automatic method could infer.
If an XSD is available, xmlrectr can use it as
additional structural evidence:
xml <- system.file("extdata", "types.xml", package = "xmlrectr")
xsd <- system.file("extdata", "types.xsd", package = "xmlrectr")
proposal <- propose_xml_profile(xml, xsd = xsd)
review_xml_proposal(proposal, "xsd")inspect_xsd() can expose useful declaration, occurrence,
required-attribute and scalar-type information.
The important design choice is that XSD is advisory. A schema describes valid document structure, but it does not necessarily tell an analyst what should constitute a row, which repeated structures should become separate entities, or which fields are relevant to a particular analysis.
So the workflow remains:
XML structure + optional XSD + user knowledge
|
v
reviewed profile
|
v
analytical rectangle
The automatic analyst representation is designed as one self-contained atomic table per XML document.
Its structural contract is conservative:
xml_entity and xml_entity_id identify
analytical entities;NA analyst columns are avoided;xml_* provenance columns;The underlying canonical representation is even more explicit. It records document order, node identity, parentage, depth, namespaces, attributes, text/CDATA and other retained XML node types.
That canonical layer is the loss-aware structural foundation on which higher-level rectangles are built.
The ambition to work with arbitrary XML is useful only if the implementation can cope with XML as it exists in practice: large documents, deep nesting, repeated records, namespaces, irregular branches and substantial structural overhead.
A large part of the development of xmlrectr has
therefore focused on algorithmic cost, memory behaviour,
workload decomposition and semantic parity.
The package deliberately separates two questions:
Changing the computational engine must never change the first answer.
Performance-critical canonical reading and structural/indexing
operations have native C implementations using libxml2.
The R implementation remains the semantic reference. Compiled code is used to accelerate structural bottlenecks, not to introduce a second set of rectangling semantics.
This distinction matters: generic XML rectangling repeatedly performs structural operations for which interpreted R can become expensive on large trees. Moving those bottlenecks into compiled code makes the generic design practical without introducing vocabulary-specific shortcuts.
For XML that should not be represented as one complete in-memory
canonical table, xmlrectr provides a streaming path based
on complete record subtrees.
batches <- list()
stats <- xml_stream_rectangle(
"large.xml",
spec,
callback = function(batch) {
batches[[length(batches) + 1L]] <<- batch
},
parallel = "auto"
)The streaming architecture has an important correctness boundary:
xml2/libxml external pointers are
never passed between processes;The point is not merely to “use less RAM”. The package tries to keep parser state, record ownership, output ordering and rectangling semantics cleanly separated.
Large results can be written without first collecting the entire rectangle into one R object:
rectangle_xml_csv(
"large.xml",
spec,
output = "large.csv",
parallel = "auto"
)For typed analytical output:
rectangle_xml_parquet(
"large.xml",
spec,
output_dir = "large-parquet",
parallel = "auto"
)Parquet support requires the optional arrow package.
The same high-level execution controls are used by the main in-memory, streaming, CSV and Parquet interfaces. Parallel execution is therefore part of the normal rectangling architecture rather than a separate workflow bolted onto one output format.
Parallelism is an execution choice of the same rectangling operation, not a separate family of user-facing functions.
# Exact sequential path
seq_out <- rectangle_xml(
file,
spec,
parallel = FALSE
)
# Request parallel execution with tuned defaults
par_out <- rectangle_xml(
file,
spec,
parallel = TRUE
)
# Let the engine decide whether parallel work is worthwhile
auto_out <- rectangle_xml(
file,
spec,
parallel = "auto"
)The sequential implementation is the semantic oracle. Parallel execution is required to preserve the same result:
identical(seq_out, par_out)The three modes have deliberately simple meanings:
parallel = FALSE uses the exact sequential path and
does not require the parallel stack;parallel = TRUE requests parallel execution using the
package’s tuned defaults;parallel = "auto" asks the engine to decide, from the
observed workload, whether parallel execution is justified.Lower-level *_parallel() functions exist for
compatibility, testing and diagnostics. They are not the intended
everyday API.
The parallel architecture follows a few strict rules:
xml2/libxml external pointers are never sent to
workers.These constraints are less flashy than a benchmark chart, but they are fundamental to making parallel execution a trustworthy implementation detail rather than a second semantics.
Parallel XML rectangling is not simply a matter of running:
workers <- parallel::detectCores()and dividing a file into equal pieces.
A useful execution plan depends on the actual structural work available:
For this reason, xmlrectr does not use filenames, XML
vocabulary names, or rules learned specifically from the validation
corpus to decide whether to parallelise.
The automatic policy is structural.
Benchmarks showed that increasing the number of workers beyond four can still reduce elapsed time, but efficiency falls and memory pressure increases.
The automatic/default worker policy therefore normally uses up to four workers, while preserving a core for the system where possible.
This is not a hard maximum.
Advanced users can explicitly request more workers:
rectangle_xml(
file,
spec,
parallel = TRUE,
workers = 8
)The default is intended to be a balanced choice, not a claim that four workers are universally optimal.
"auto" currently decides whether there is enough workThe current in-memory auto policy asks whether there is enough canonical record work per worker to justify process-level parallelism.
The policy includes a work floor of approximately:
25,000 canonical record nodes per worker
together with sufficient record/subtree structure, approximately:
at least 128 records per worker
or
a median record subtree of at least 1,000 canonical nodes
These are engineering defaults derived from broad structural benchmarking. They are not XML-vocabulary rules, and they should not be read as eternal constants or promises of a particular speedup.
There is also a cheap impossibility check. If the entire canonical table has fewer than roughly:
workers * 25,000
nodes, the workload cannot satisfy the per-worker work floor. Auto mode can then remain sequential immediately instead of performing a more expensive record-span analysis.
This fast path is important. An early version of automatic planning was semantically correct but could make small sequential jobs noticeably slower simply because planning repeated structural work that the sequential path would perform anyway. The current design reuses validation and rejects obviously too-small workloads before doing that extra work.
In other words, auto mode is designed not only to find parallel opportunities, but also to get out of the way when parallelism would be pointless.
xmlrectr retains two parallel execution strategies
because throughput and memory pressure are different optimisation
problems.
parallel_chunks:
throughput-orientedWith parallel_chunks, workers own independent vectorised
outer chunks.
Conceptually:
chunk 1 ---> worker 1
chunk 2 ---> worker 2
chunk 3 ---> worker 3
chunk 4 ---> worker 4
This strategy is designed primarily to:
For in-memory automatic execution, this is generally the preferred strategy.
A useful scheduling model is approximately one active owned outer chunk per worker.
shared_chunk:
memory-orientedWith shared_chunk, one bounded outer chunk is shared and
subdivided into multiple vectorised tasks coordinated through
mori.
Conceptually:
bounded shared outer chunk
/ | | \
task task task task
| | | |
workers process coarse ranges
Its purpose is to reduce input-memory duplication while still preserving parallel work.
This is particularly attractive for streaming, CSV and Parquet workflows, where bounded memory is part of the reason for choosing the execution mode in the first place.
Automatic streaming/output execution therefore prefers
shared_chunk when the required stack is available. If
mori is unavailable, automatic execution can fall back to
parallel_chunks.
The trade-off is intentional:
parallel_chunks is primarily throughput-oriented;shared_chunk is primarily memory-oriented.Neither strategy dominates the other on every machine and workload.
The advanced API exposes:
workers
strategy
chunk_records
task_records
The last two parameters solve different problems.
chunk_records:
working-set admissionchunk_records controls the size of the outer
batch admitted at once.
It therefore influences:
For the current streaming defaults:
parallel_chunks:
chunk_records = 512
For shared_chunk:
chunk_records = min(1024, max(256, workers * 256))
These are tuned defaults, not XML-specific rules.
task_records:
scheduling granularityWithin a shared outer chunk, task_records controls how
finely the work is subdivided for scheduling.
Benchmarks found a fairly broad useful plateau around 4 to 8 tasks per worker. The balanced automatic setting is therefore approximately 4 tasks per worker.
That gives workers enough independent work for load balancing without producing a large number of tiny tasks whose scheduling cost dominates useful computation.
So:
chunk_records -> controls admitted working-set size / memory
task_records -> controls scheduling granularity inside that work
Keeping these concepts separate is especially important for
shared_chunk.
Performance results are included here because the defaults were not chosen by intuition alone.
They should nevertheless be interpreted carefully:
Never compare raw elapsed times from different machines as though they belong to one benchmark series.
Absolute timings depend on processor, memory subsystem, operating system, R build, package versions and background load. The useful quantities are same-machine sequential/parallel speedup, worker efficiency, memory behaviour and exact semantic parity.
The following results are representative development measurements on
one machine, referred to during development as einstein.
They document why the current defaults exist; they are not runtime
guarantees.
A representative synthetic workload was approximately 16.266 MiB.
Sequential execution:
193.977 s
Parallel results:
| Workers | parallel_chunks |
Speedup | shared_chunk |
Speedup |
|---|---|---|---|---|
| 2 | 121.585 s | 1.60x | 107.070 s | 1.81x |
| 4 | 63.340 s | 3.06x | 68.869 s | 2.82x |
| 11 | 46.051 s | 4.21x | 55.525 s | 3.49x |
Several conclusions follow.
First, more than four workers can improve elapsed time. Four is therefore not a hard ceiling.
Second, scaling efficiency declines substantially at high worker counts. The extra processes are doing useful work, but the cost of coordination, memory traffic and finite task parallelism becomes increasingly important.
Third, four workers gave a strong compromise between speedup, efficiency and memory pressure. That is why automatic/default worker selection is normally capped there unless the user explicitly chooses otherwise.
On the same approximate 16.266 MiB workload with four workers:
| Strategy | Peak PSS |
|---|---|
parallel_chunks |
about 3618 MiB |
shared_chunk |
about 3142 MiB |
In that experiment, shared_chunk reduced peak
proportional set size by roughly 13% and private memory
by roughly 15%.
Depending on phase, median PSS could fall by considerably more.
That reduction is meaningful, even though shared_chunk
can be slower on some workloads. It is the empirical reason the
memory-oriented strategy remains part of the package rather than being
removed in favour of the single fastest throughput result.
These measurements do not imply:
shared_chunk always saves exactly 13% memory;The actual structural workload matters more than the byte size of the XML file.
The final automatic policy was also checked against real XML on the same development machine.
Representative P6.1 smoke timings were:
| XML | Sequential | Auto | Auto decision |
|---|---|---|---|
| UBL | 0.960 s | 0.874 s | sequential |
| Maven | 1.111 s | 1.126 s | sequential |
| EAD | 33.756 s | 30.973 s | parallel |
The important observation for UBL and Maven is not the tiny timing difference. It is that automatic planning did not impose a material penalty on small jobs that should remain sequential.
The EAD document crossed the structural threshold and was sent to the parallel engine.
That particular EAD run should not be used to argue either that parallelism is spectacular or that it is useless. It happened to lie relatively close to the crossover region on that machine. Larger synthetic workloads demonstrated much stronger same-machine speedups.
Most users should stop at:
rectangle_xml(file, spec, parallel = "auto")The following controls exist for benchmarking, unusually constrained machines, or specialist tuning:
rectangle_xml(
file,
spec,
parallel = TRUE,
workers = 4,
strategy = "shared_chunk",
chunk_records = 1024,
task_records = 64
)Useful questions for an advanced tuning exercise are:
Do not tune from filenames or XML vocabulary names.
Do not assume that tiny chunk/task values used in semantic stress tests are production recommendations.
And do not optimise one specific XML file at the expense of the generic structural rules.
For performance work, compare strategies on the same machine and the same XML.
A simple reproducible pattern is:
spec <- compile_xml_profile(profile, file)
t_seq <- system.time(
seq_out <- rectangle_xml(
file,
spec,
parallel = FALSE
)
)
t_auto <- system.time(
auto_out <- rectangle_xml(
file,
spec,
parallel = "auto"
)
)
t_forced <- system.time(
forced_out <- rectangle_xml(
file,
spec,
parallel = TRUE
)
)
stopifnot(
identical(seq_out, auto_out),
identical(seq_out, forced_out)
)
rbind(
sequential = t_seq,
auto = t_auto,
forced_parallel = t_forced
)For serious benchmarking:
Small XML documents may correctly be faster sequentially because process startup and scheduling have a cost. Larger, sufficiently coarse record workloads are where parallel execution can pay off.
This is why parallel = "auto" exists:
parallelism is a tool, not a goal in itself.
Some of the harshest parallel tests deliberately used settings that would make poor production defaults.
For example, the 30-file forced-parallel semantic stress run used approximately:
workers = 2
chunk_records = 64
task_records = 8
Those small values were chosen to force many scheduling boundaries and expose correctness problems.
They were not selected for speed.
The result was:
This distinction is important when reading the project’s benchmark history. Different experiments answer different questions:
Conclusions from one class should not be casually transferred to another.
Before conversion into the R package, the frozen validated engine was exercised against a deliberately diverse corpus of 30 real-world XML documents.
The validation included:
parallel = "auto"
interface;In the final 30-file automatic-policy run:
Crucially, the engine was not modified with vocabulary-specific rules, filename-specific exceptions or thresholds tuned to make these files pass.
The corpus covers very different XML domains and structures:
The purpose of this corpus is structural diversity, not optimisation for these particular vocabularies. General structural rules are preferred over corpus-specific special cases.
This validation does not mean that every arbitrary XML document has one objectively correct analyst table. It means that the package’s generic structural rules and execution engine have been exercised across a broad set of real-world XML shapes without resorting to vocabulary-specific parsers.
The engineering conclusions behind the current defaults can be summarised compactly:
parallel_chunks is the primary
throughput-oriented strategy.shared_chunk is the primary
memory-oriented strategy.shared_chunk has shown materially lower peak/private
memory in representative tests.chunk_records and inner task_records
solve different problems and are tuned separately.The public API is simple because this machinery sits underneath it, not because the machinery does not exist.
xmlrectr expects well-formed XML.
Malformed XML is detected and reported as an error or warning where
appropriate. Repairing broken XML is intentionally outside the scope of
the package: xmlrectr rectangles XML; it does not try to
guess how a malformed source document should be rewritten.
Likewise, XSD inspection is intended to provide useful schema
evidence for rectangling. xmlrectr is not a complete XSD
validation or repair framework.
This README is intended to be a self-contained
introduction. You should not need to install the package or
open a vignette merely to understand what xmlrectr is
trying to do.
For readers who want more detail, the repository also contains technical material that can be read directly on GitHub:
The technical articles are intentionally more detailed than this README. They are the right place for readers interested in scheduler design, bounded-memory execution, native acceleration, benchmark interpretation, worker/chunk/task tuning, validation boundaries and the engineering decisions behind the simple public interface.
A few principles define the project:
The intended experience is simple even though the implementation underneath is not:
Give
xmlrectran XML file. If you know nothing about it, start exploring immediately. If you know more, tell the package what you know. If the workload is large, let the execution engine do the heavy lifting.
xmlrectr aims to be a universal
rectangler, not a universal semantic interpreter.
It is designed to take arbitrary well-formed XML and produce useful R-oriented tabular representations without requiring a vocabulary-specific parser. It can exploit schema information and user knowledge when they exist, but it does not require them for exploratory use.
That is a deliberately ambitious target. The package cannot know the domain meaning of every XML vocabulary, but it can do a great deal of the structural and computational work required to move from hierarchical XML to an analyst-friendly table.
That is the problem xmlrectr is built to solve.