fastPLS always contains a compiled CPU implementation.
Apple Metal and NVIDIA CUDA are optional backends selected while the
package is compiled. The package does not download CUDA, OpenBLAS, or
other native libraries after installation. This avoids changing an
installed package’s native dependencies and ensures that every linked
library matches the operating system and CPU architecture.
The default source-build policy is:
Backend availability is established at installation time. Requesting an accelerator that was not compiled and is not available produces an error; fastPLS does not silently run the operation on the CPU.
Install the released package from CRAN:
The development version can be installed from GitHub. Use a fresh R session when replacing an existing installation.
install.packages("remotes")
remotes::install_github(
"tkcaccia/fastPLS",
upgrade = "never",
force = TRUE,
build_vignettes = TRUE
)Binary packages contain only the capabilities compiled by their builder. To enable a locally installed CUDA Toolkit or Apple Metal, install fastPLS from source on that computer.
Install Apple’s command-line developer tools before compiling from source:
Apple Accelerate supplies the CPU BLAS/LAPACK implementation. A separate OpenBLAS installation is neither required nor recommended for ordinary macOS use. Metal is enabled automatically when its frameworks are available.
Sys.setenv(FASTPLS_USE_METAL = "1")
remotes::install_github(
"tkcaccia/fastPLS",
force = TRUE,
upgrade = "never",
build_vignettes = TRUE
)Setting FASTPLS_USE_METAL=1 makes a missing Metal
toolchain an installation error. Omit it to use automatic detection, or
set it to 0 for a CPU-only build.
Install the compiler toolchain and OpenBLAS development files before building fastPLS:
The following setting requires OpenBLAS and prevents an unnoticed fallback to R’s BLAS/LAPACK:
For performance work, requiring the OpenBLAS family is necessary but not sufficient. Record the exact OpenBLAS release and the CPU kernel selected at runtime. An older dynamic-architecture build may recognize the library as OpenBLAS while selecting a generic or legacy kernel on a newer processor, which can substantially change matrix-multiplication timings. Prefer a current build compiled for the target architecture.
Install the corresponding Fedora development packages:
Then install fastPLS from source with
FASTPLS_USE_OPENBLAS=1, as shown for Ubuntu.
Install the Rtools release matching the installed R version. A standard source installation works without a separate OpenBLAS installation and uses the BLAS/LAPACK supplied by R when OpenBLAS is unavailable.
For faster CPU execution, install an x86-64 OpenBLAS development archive. For example, OpenBLAS can be installed from an MSYS2 UCRT64 terminal:
Point the package configuration to that prefix:
Every linked static library must be compiled for ARM64. Do not set
OPENBLAS_ROOT to an x86-64 Rtools or MSYS2 directory.
fastPLS checks the target architecture and rejects incompatible OpenBLAS
archives rather than allowing the linker to combine x86-64 and ARM64
objects.
When a compatible ARM64 OpenBLAS installation is unavailable, retain the default automatic policy:
Sys.unsetenv(c("OPENBLAS_ROOT", "FASTPLS_USE_OPENBLAS"))
remotes::install_github(
"tkcaccia/fastPLS",
force = TRUE,
upgrade = "never",
build_vignettes = TRUE
)This builds the CPU backend with R’s BLAS/LAPACK and portable float32 kernels. The current Windows CUDA build requires the x86-64 NVIDIA toolkit and is not enabled for Windows ARM64.
A CUDA-capable NVIDIA driver alone is not sufficient to compile
fastPLS. The CUDA Toolkit, including nvcc, headers, cuBLAS,
cuSOLVER, cuRAND, and runtime libraries, must be installed before the R
package. cuDNN is not required. The driver and toolkit are separate.
fastPLS never installs, removes, upgrades, or downgrades the host NVIDIA
driver. Do not install Debian’s nvidia-cuda-toolkit package
merely to build fastPLS on a computer with a working vendor driver. Use
a compatible NVIDIA toolkit installed separately, or a controlled NVIDIA
CUDA container with host-GPU passthrough.
If CUDA is installed in a standard location, the default
auto policy detects it. For a reproducible source build,
provide the toolkit location and require CUDA explicitly:
Sys.setenv(
FASTPLS_USE_CUDA = "1",
FASTPLS_REQUIRE_CUDA = "1",
CUDA_HOME = "/usr/local/cuda"
)
install.packages("fastPLS", type = "source")On Windows x86-64, CUDA_ROOT typically resembles:
Use a toolkit version supported by the installed NVIDIA driver.
CUDA_ROOT, CUDA_HOME, and
CUDA_PATH are accepted. The configuration searches
include, lib, and lib64 below
that prefix as well as targets/*/include,
targets/*/lib, and targets/*/lib64. It
compiles and links a minimal program against CUDA Runtime, cuBLAS,
cuSOLVER, and cuRAND; nvcc being present is not sufficient.
The default CUDA host compiler is a system compiler under
/usr/bin, ahead of CUDA and Conda wrappers. Set
FASTPLS_CUDA_HOST_CXX to an absolute compatible system
compiler only when the default is unsuitable.
Restart R after installation, load fastPLS, and inspect the compiled CPU library and accelerator capabilities:
The backend field is "Accelerate",
"OpenBLAS", or "R BLAS/LAPACK". For OpenBLAS
builds, the report also contains the library version, full configuration
string, selected CPU core, parallel runtime, active thread count, and
resolved library path when available. Use
fastPLS_blas(details = FALSE) when a script requires only
the former scalar backend name.
Publication benchmarks should verify three items before timing:
The validation utility and campaign wrapper in fastPLS-extra
enforce these checks. A timing run made with a different or unverified
OpenBLAS build is a separate benchmark condition and should not be
combined with verified results.
Test the selected runtime backend with a small deterministic fit:
X <- as.matrix(iris[, 1:4])
y <- iris$Species
fit <- pls(X, y, ncomp = 1:2, backend = "cpu", seed = 11)
stopifnot(length(fit$Yfit) == 2L)
if (has_cuda()) {
fit_cuda <- pls(X, y, ncomp = 1:2, backend = "cuda", seed = 11)
}
if (has_metal()) {
fit_metal <- pls(X, y, ncomp = 1:2, backend = "metal", seed = 11)
}Installation determines which backends exist. A function argument selects a backend for one operation, while session options establish defaults:
options(backend = "cpu", n.cores = 4L)
fit <- pls(X, y, ncomp = 1:2)
fit_one_core <- pls(X, y, ncomp = 1:2, n.cores = 1L)An explicit backend or n.cores argument
overrides the corresponding option. The FASTPLS_BACKEND
environment variable is consulted only when no explicit argument or
options(backend=...) value is present. CPU remains the
default.
| Variable | Values | Purpose |
|---|---|---|
FASTPLS_USE_OPENBLAS |
auto, 0, 1 |
OpenBLAS policy. |
OPENBLAS_ROOT |
directory | OpenBLAS installation prefix. |
FASTPLS_USE_CUDA |
auto, 0, 1 |
Detect, disable, or request CUDA. |
FASTPLS_REQUIRE_CUDA |
0, 1 |
Require a successful CUDA compile/link probe. |
FASTPLS_CUDA_DIAGNOSTIC_ONLY |
0, 1 |
Build an explicitly labelled nonfunctional CUDA diagnostic package. |
FASTPLS_CUDA_HOST_CXX |
file | Absolute system host compiler for
nvcc. |
CUDA_ROOT, CUDA_HOME,
CUDA_PATH |
directory | CUDA Toolkit prefix. |
FASTPLS_USE_METAL |
auto, 0, 1 |
Metal policy. |
Set these variables before installing the package. Changing them afterward does not add a backend to an already compiled package.
Windows cannot replace a loaded DLL. Restart R without loading fastPLS, then reinstall it.
Remove stale OPENBLAS_ROOT, R_TOOLS_SOFT,
and CUDA settings, restart R, and install again. The configuration
output should report target architecture aarch64 and must
not name an x86-64 OpenBLAS archive.
GitHub source installations may omit built vignettes unless
requested. Install with build_vignettes=TRUE, then use:
Retain the package installation output and record the resulting environment:
For benchmark reports, retain the output of
fastPLS_blas() and also record compiler versions, the
resolved OpenBLAS library, OpenBLAS release and selected CPU core (or
the Accelerate version), CUDA toolkit and driver versions, requested CPU
thread count, operating system, and CPU/GPU hardware.
sessionInfo()
#> R version 4.6.0 (2026-04-24)
#> Platform: aarch64-apple-darwin23
#> Running under: macOS Sonoma 14.5
#>
#> Matrix products: default
#> BLAS: /Library/Frameworks/R.framework/Versions/4.6/Resources/lib/libRblas.0.dylib
#> LAPACK: /Library/Frameworks/R.framework/Versions/4.6/Resources/lib/libRlapack.dylib; LAPACK version 3.12.1
#>
#> locale:
#> [1] en_US.UTF-8/en_US.UTF-8/en_US.UTF-8/C/en_US.UTF-8/en_US.UTF-8
#>
#> time zone: Africa/Johannesburg
#> tzcode source: internal
#>
#> attached base packages:
#> [1] stats graphics grDevices utils datasets methods
#> [7] base
#>
#> other attached packages:
#> [1] fastPLS_0.3
#>
#> loaded via a namespace (and not attached):
#> [1] digest_0.6.39 R6_2.6.1 fastmap_1.2.0 xfun_0.60
#> [5] float_0.3-3 cachem_1.1.0 knitr_1.51 htmltools_0.5.9
#> [9] rmarkdown_2.31 lifecycle_1.0.5 cli_3.6.6 sass_0.4.10
#> [13] jquerylib_0.1.4 compiler_4.6.0 tools_4.6.0 evaluate_1.0.5
#> [17] bslib_0.12.0 yaml_2.3.12 otel_0.2.0 rlang_1.3.0
#> [21] jsonlite_2.0.0