---
title: "Combine pipelines"
output:
  rmarkdown::html_vignette:
    toc: true
    toc_depth: 4
description: >
  How to combine different pipelines to a single pipeline.
vignette: >
  %\VignetteIndexEntry{Combine pipelines}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r knitr-setup, include = FALSE}
knitr::opts_chunk$set(
    comment = "#",
    prompt = FALSE,
    tidy = FALSE,
    cache = FALSE,
    collapse = TRUE
)

old <- options(width = 100L)
```

The possibility to combine pipelines basically allows to modularize the
pipeline creation process. This is especially useful when you have a set of
pipelines that are used in different contexts and you want to avoid code
duplication^[
    Note that code duplication is not bad per se and can even be
    preferable, since it reduces entanglement and improves local
    readability, or as Sandi Metz put it, "duplication is far cheaper than
    the wrong abstraction".
].

### Two pipelines

Let's define one pipeline that is used for data preprocessing and one that
does the modelling^[
    The step functions in these pipelines are minimal to keep the focus
    on the combine functionality.
].

```{r define-prepocessing-pipeline}
library(pipeflow)

pip1 <- pip_new("preprocess") |>
    pip_add("data", \(x = 1:5) x) |>
    pip_add("prep", \(x = ~data) x + 1) |>
    pip_add("standardize", \(x = ~prep, scale = 2) x * scale)
```

```{r define-modelling-pipeline}
pip2 <- pip_new("model") |>
    pip_add("data", \(x = 1:5) x) |>
    pip_add("fit", \(x = ~data, k = 2, b = 0) x * k + b) |>
    pip_add("predict", \(x = ~fit) paste0("pred: ", x))
```


### Combined pipeline

Next we combine the two pipelines using `rbind()`.

```{r}
pip <- rbind(pip1, pip2)

pip
```

Note that the `data` step of the second pipeline
has been renamed to `data2` in both the `step` and the `depends` columns
(see line 4 above). That is, when "rbinding" pipelines, {pipeflow}
automatically ensures that all step names stay unique by renaming
any duplicates accordingly.

As is also visible from the graphical representation of the pipeline,

```{r, eval = FALSE}
library(visNetwork)
do.call(visNetwork, args = pip_graph(pip)) |>
    visHierarchicalLayout(direction = "LR")
```

```{r, echo = FALSE}
library(visNetwork)
do.call(
    visNetwork,
    args = c(pip_graph(pip), list(height = 250, width = 500))
) |>
    visHierarchicalLayout(direction = "LR", sortMethod = "directed")
```

the two pipelines are not yet connected. To make sense of the combined
pipeline, we want to use the output of the `standardize` step as the input of
the `data2` step, which we can do by applying the `replace` function,
which was introduced in the previous vignette
[modify the pipeline](v02-modify-pipeline.html),
as follows:

```{r}
pip |> pip_replace("data2", \(x = ~standardize) x)

pip
```

The `data2` step points to the output of the `standardize` step,
so that both pipelines are now connected.

```{r, echo = FALSE}
do.call(
    visNetwork,
    args = c(pip_graph(pip), list(height = 100, width = 700))
) |>
    visHierarchicalLayout(direction = "LR")
```

#### Relative indexing

Since the name of the re-routed step might not always be known^[A typical
example would be appending several pipelines in a programmatic context.],
the {pipeflow} package also provides a relative position indexing mechanism,
which allows to rewrite the above command using a number (instead of the
step name `standardize`) while having the same effect as above.

```{r}
pip |> pip_replace("data2", \(x = ~ -1) x)

pip
```

The relative indexing mechanism allows to refer to steps positioned
above the current step. The index `~-1` can be interpreted as "go one step back", `~-2`
as "go two steps back", and so on.


### Combined pipeline results

Let's now run the combined pipeline and inspect the results of the final step.

```{r}
pip_run(pip)
```

```{r}
pip[["predict", "out"]]
```

As we can see, the outputs of the preprocessing pipeline flow into the
modelling pipeline. We can now go ahead and for example change the multiplier
of the `fit` step and rerun the pipeline.

```{r}
pip_set_params(pip, params = list(k = 3))
```

```{r, echo = FALSE}
library(visNetwork)
do.call(
    visNetwork,
    args = c(pip_graph(pip), list(height = 100, width = 700))
) |>
    visHierarchicalLayout(direction = "LR", sortMethod = "directed")
```

```{r}
pip_run(pip)
```

```{r}
pip[["predict", "out"]]
```

Whether you combine them or not, in practice pipelines can get long quickly.
The next vignette shows how you can easily focus on certain parts of your
entire analysis workflow by
[Pipeline views](v03b-pipeline-views.html).

```{r, include = FALSE}
options(old)
```
