A pipeline step can contain another pipeline: instead of computing a result directly, the step builds an inner pipeline and returns it. This allows to reuse a standard analysis pipeline inside a larger workflow.
We start with the same pipeline as in the previous “Split, map, and reduce” vignette, fitting a linear model and returning its coefficients. This will serve as our inner pipeline.
library(pipeflow)
inner <- pip_new("coefficients") |>
pip_add("data", \(data = NULL) data) |>
pip_add(
"fit",
\(data = ~data, xVar = "x", yVar = "y") {
lm(paste(yVar, "~", xVar), data = data)
}
) |>
pip_add("coefs", \(fit = ~fit) coefficients(fit))
inner
# <pipeflow> coefficients (3 steps)
# ---------------------------------
# step params depends state
# 1: data data new
# 2: fit data,xVar,yVar data new
# 3: coefs fit fit new
# ---------------------------------
# <ready> last run: neverThe outer pipeline splits the data into subsets and derives the model coefficients by running the inner pipeline for each split.
# Helper to run inner pipeline
run_inner_pip <- function(pip, name, data) {
pip$name <- sprintf("coefs for *%s*", name)
# Set data subset for inner pipeline and run it
pip_set_params(pip, list(data = data)) |> pip_run()
pip[["coefs", "out"]]
}
outer <- pip_new("full analysis") |>
pip_add("data", \(data = NULL) data) |>
pip_add(
"split_data", \(data = ~data, byVar = "by") {
split(data, f = data[[byVar]])
}
) |>
pip_add(
"inner_run",
\(dataList = ~split_data, xVar = "x", yVar = "y") {
p <- pip_clone(inner)
# Forward parameters to inner
pip_set_params(p, list(xVar = xVar, yVar = yVar))
Map(
f = run_inner_pip,
name = names(dataList),
data = dataList,
MoreArgs = list(pip = p)
)
}
) |>
pip_add(
"combine",
\(coefs = ~inner_run) as.data.frame(do.call(rbind, coefs))
)
outer
# <pipeflow> full analysis (4 steps)
# ----------------------------------
# step params depends state
# 1: data data new
# 2: split_data data,byVar data new
# 3: inner_run dataList,xVar,yVar split_data new
# 4: combine coefs inner_run new
# ----------------------------------
# <ready> last run: neverNote the inner_run step: its parameters
xVar and yVar are forwarded to the inner
pipeline.
Let’s now set the analysis parameters and run the full pipeline:
outer |>
pip_set_params(
list(
data = iris,
xVar = "Sepal.Length",
yVar = "Sepal.Width",
byVar = "Species"
)
) |>
pip_run()
# info [2026-09-27 18:21:16.780 UTC]: Starting run of pipeflow 'full analysis'
# info [2026-09-27 18:21:16.780 UTC]: Step 1/4 data
# info [2026-09-27 18:21:16.781 UTC]: Step 2/4 split_data
# info [2026-09-27 18:21:16.782 UTC]: Step 3/4 inner_run
# info [2026-09-27 18:21:16.785 UTC]: Starting run of pipeflow 'coefs for *setosa*'
# info [2026-09-27 18:21:16.785 UTC]: Step 1/3 data
# info [2026-09-27 18:21:16.786 UTC]: Step 2/3 fit
# info [2026-09-27 18:21:16.787 UTC]: Step 3/3 coefs
# info [2026-09-27 18:21:16.788 UTC]: Finished run of pipeflow 'coefs for *setosa*'
# info [2026-09-27 18:21:16.792 UTC]: Starting run of pipeflow 'coefs for *versicolor*'
# info [2026-09-27 18:21:16.792 UTC]: Step 1/3 data
# info [2026-09-27 18:21:16.792 UTC]: Step 2/3 fit
# info [2026-09-27 18:21:16.793 UTC]: Step 3/3 coefs
# info [2026-09-27 18:21:16.794 UTC]: Finished run of pipeflow 'coefs for *versicolor*'
# info [2026-09-27 18:21:16.796 UTC]: Starting run of pipeflow 'coefs for *virginica*'
# info [2026-09-27 18:21:16.796 UTC]: Step 1/3 data
# info [2026-09-27 18:21:16.796 UTC]: Step 2/3 fit
# info [2026-09-27 18:21:16.797 UTC]: Step 3/3 coefs
# info [2026-09-27 18:21:16.798 UTC]: Finished run of pipeflow 'coefs for *virginica*'
# info [2026-09-27 18:21:16.799 UTC]: Step 4/4 combine
# info [2026-09-27 18:21:16.800 UTC]: Finished run of pipeflow 'full analysis'The output of the inner_run step is a list of
coefficient vectors, one for each species,
outer[["inner_run", "out"]]
# $setosa
# (Intercept) Sepal.Length
# -0.5694327 0.7985283
#
# $versicolor
# (Intercept) Sepal.Length
# 0.8721460 0.3197193
#
# $virginica
# (Intercept) Sepal.Length
# 1.4463054 0.2318905and the combine step returns the expected combined
table.
outer[["combine", "out"]]
# (Intercept) Sepal.Length
# setosa -0.5694327 0.7985283
# versicolor 0.8721460 0.3197193
# virginica 1.4463054 0.2318905Now suppose we want to change one of the model settings, say use
Petal.Length instead of Sepal.Length as the
predictor.
pip_set_params(outer, params = list(xVar = "Petal.Length"))
outer
# <pipeflow> full analysis (4 steps)
# ----------------------------------
# step params depends state out
# 1: data data done <data.frame[150x5]>
# 2: split_data data,byVar data done <list[3]>
# 3: inner_run dataList,xVar,yVar split_data outdated <list[3]>
# 4: combine coefs inner_run outdated <data.frame[3x2]>
# ----------------------------------
# <ready> last run: 2026-09-27 20:21:16Since the xVar parameter is part of the
inner_run step’s function arguments, the
inner_run step’s state (and its downstream dependencies)
correctly has now been marked as “outdated”.
While forwarding the parameters manually to the inner pipeline is straight-forward in our toy example, trying to manually synchronize real-world parameter sets between the outer and inner pipeline quickly becomes a unfeasible and a source for bugs that are hard to detect.
For this reason, in practice, the following pattern should be used,
which basically just forwards the combined set of all existing
parameters. To do this, we replace the inner_run step as
follows:
outer |> pip_replace(
"inner_run",
\(dataList = ~split_data, ...) {
p <- pip_clone(inner)
# Forward all parameters (from outer and inner)
all_params <- .self$get_params()
pip_set_params(p, all_params)
Map(
f = run_inner_pip,
name = names(dataList),
data = dataList,
MoreArgs = list(pip = p)
)
},
params = pip_get_params(inner) # <--- default parameters of inner pipeline
)
outer
# <pipeflow> full analysis (4 steps)
# ----------------------------------
# step params depends state out
# 1: data data done <data.frame[150x5]>
# 2: split_data data,byVar data done <list[3]>
# 3: inner_run data,xVar,yVar,dataList split_data new [NULL]
# 4: combine coefs inner_run outdated <data.frame[3x2]>
# ----------------------------------
# <ready> last run: 2026-09-27 20:21:16In this version, the inner_run step no longer declares
xVar and yVar as its own arguments. Instead,
its parameters are seeded with the inner pipeline’s parameters via
params = pip_get_params(inner), and the step forwards the
combined parameter set at run time. Three aspects are worth spelling
out:
.self refers to the pipeline that is
currently being run and is available inside every step function
without having to declare it. It exposes the usual pipeline interface,
so .self$get_params() returns the current unbound
parameters of the outer pipeline, and things like
.self$name or .self$run() work as expected.
For more examples on this mechanism see the self-modifying pipelines
vignette.params = pip_get_params(inner)
registers the inner pipeline’s unbound parameters (data,
xVar, yVar) as parameters of the
inner_run step. They therefore show up in the step’s
params column, become part of the outer pipeline’s
parameter set (and can be updated with
pip_set_params(outer, ...)), and are passed to the step
function when it runs.params overlaps with the arguments of the step function,
the function’s default values take precedence. This is why the
replacement function above declares only dataList and
... — had it declared e.g. xVar = "x", that
default would win over the value coming from params, and
the forwarding would silently be ignored. As another example,
pip_add("s", \(x = 99, ...) x, params = list(x = 1)) stores
x = 99, not x = 1.Let’s re-run the full pipeline.
outer |>
pip_set_params(
list(
data = iris,
xVar = "Sepal.Length",
yVar = "Sepal.Width",
byVar = "Species"
)
) |>
pip_run()
# info [2026-09-27 18:21:16.869 UTC]: Starting run of pipeflow 'full analysis'
# info [2026-09-27 18:21:16.869 UTC]: Step 1/4 data
# info [2026-09-27 18:21:16.869 UTC]: Step 2/4 split_data
# info [2026-09-27 18:21:16.870 UTC]: Step 3/4 inner_run
# warn [2026-09-27 18:21:16.872 UTC]: Trying to set parameters not defined in the target: byVar
# Warning in pip_set_params(p, all_params): Trying to set parameters not defined in the target: byVar
# info [2026-09-27 18:21:16.874 UTC]: Starting run of pipeflow 'coefs for *setosa*'
# info [2026-09-27 18:21:16.874 UTC]: Step 1/3 data
# info [2026-09-27 18:21:16.875 UTC]: Step 2/3 fit
# info [2026-09-27 18:21:16.876 UTC]: Step 3/3 coefs
# info [2026-09-27 18:21:16.877 UTC]: Finished run of pipeflow 'coefs for *setosa*'
# info [2026-09-27 18:21:16.878 UTC]: Starting run of pipeflow 'coefs for *versicolor*'
# info [2026-09-27 18:21:16.878 UTC]: Step 1/3 data
# info [2026-09-27 18:21:16.879 UTC]: Step 2/3 fit
# info [2026-09-27 18:21:16.880 UTC]: Step 3/3 coefs
# info [2026-09-27 18:21:16.881 UTC]: Finished run of pipeflow 'coefs for *versicolor*'
# info [2026-09-27 18:21:16.882 UTC]: Starting run of pipeflow 'coefs for *virginica*'
# info [2026-09-27 18:21:16.882 UTC]: Step 1/3 data
# info [2026-09-27 18:21:16.883 UTC]: Step 2/3 fit
# info [2026-09-27 18:21:16.884 UTC]: Step 3/3 coefs
# info [2026-09-27 18:21:16.885 UTC]: Finished run of pipeflow 'coefs for *virginica*'
# info [2026-09-27 18:21:16.885 UTC]: Step 4/4 combine
# info [2026-09-27 18:21:16.886 UTC]: Finished run of pipeflow 'full analysis'Note that the warning in the log above is expected and harmless.
Basically, .self$get_params() returns the outer pipeline’s
entire parameter set, which includes byVar (from
the split_data step) while the inner pipeline only defines
data, xVar, and yVar. Since
pip_set_params() reports any parameters that are not
defined in the target and simply leaves them unset, the inner pipeline
still receives all parameters it knows and the result is unaffected.
If you need to omit the warning (e.g. in production code), just suppress it:
Alternatively, you could first restrict the forwarded parameters to those the inner pipeline actually knows. That has the same effect but adds code, and since forwarding the full parameter set is the whole point of this pattern, there is no need for it.
Again, changing one of the inner parameters will correctly outdate
the inner_run step plus downstream dependencies.
pip_set_params(outer, params = list(xVar = "Petal.Length"))
outer
# <pipeflow> full analysis (4 steps)
# ----------------------------------
# step params depends state out
# 1: data data done <data.frame[150x5]>
# 2: split_data data,byVar data done <list[3]>
# 3: inner_run data,xVar,yVar,dataList split_data outdated <list[3]>
# 4: combine coefs inner_run outdated <data.frame[3x2]>
# ----------------------------------
# <ready> last run: 2026-09-27 20:21:16With the above pattern, you can now change both the inner and outer pipeline, adding and/or removing any steps or parameters, without having to worry about parameter synchronization.
Both this and the previous
vignette solve the same “split, apply, and combine” problem. The
built-in execution modes
(exec = "split"/"reduce") are the recommended
default whenever they fit while nested pipelines can be considered the
more general tool. As a rule of thumb: