Tables are where document intelligence earns its keep: they carry the numbers, but they are exactly what naive text extraction mangles. doclingr uses Docling’s table-structure model to recover cells, then hands each table back as a tibble.
docling_tables() returns a list with one tibble per
detected table, in document order:
library(doclingr)
doc <- docling_convert("financials.pdf")
tables <- docling_tables(doc)
length(tables) # how many tables Docling found
tables[[1]] # the first table, as a tibbleEach tibble carries a page attribute recording where the
table came from:
The table model has two modes. The default "accurate"
recovers complex structure (spanning cells, nested headers) at some
cost; "fast" is quicker and often enough for clean
grids:
Because each table is a tibble, the whole tidyverse is available. For example, tag every table with its page and stack them into one long frame:
library(dplyr)
library(purrr)
all_tables <- docling_tables(doc) |>
imap(\(tbl, i) mutate(tbl, .table = i, .page = attr(tbl, "page"))) |>
list_rbind()
all_tablesOr write each table to its own CSV:
readr::type_convert() or
dplyr::mutate(across(...)) once you know each table’s
schema.docling_convert(..., ocr = TRUE), the default).docling_chunk() can repeat the header row – see
vignette("rag").
Need a high-speed mirror for your open-source project?
Contact our mirror admin team at info@clientvps.com.
This archive is provided as a free public service to the community.
Proudly supported by infrastructure from VPSPulse , RxServers , BuyNumber , UnitVPS , OffshoreName and secure payment technology by ArionPay.