fastPLS system requirements
===========================

fastPLS always builds the CPU backend. CUDA and Apple Metal are optional
acceleration backends and are not required to install the package. On macOS,
Metal is enabled automatically when the system frameworks are available.

CPU-only installation
---------------------

No external GPU software is required. macOS uses Apple Accelerate by default.
On Linux and Windows, fastPLS uses OpenBLAS when it is available and otherwise
uses the BLAS/LAPACK libraries supplied by R. For faster Linux and Windows CPU
execution, install the OpenBLAS development headers and libraries. On Debian or
Ubuntu, install them with:

  sudo apt-get install libopenblas-dev pkg-config

If OpenBLAS is installed outside the standard search path, set OPENBLAS_ROOT
to its prefix before installation. Set FASTPLS_USE_OPENBLAS=1 to require
OpenBLAS explicitly; installation then fails with an informative error if it
cannot be found. Set FASTPLS_USE_OPENBLAS=0 to use the R-supplied libraries.
For source builds intended for repeated large analyses, use a current OpenBLAS
release compiled for the target architecture rather than a generic portable
binary. After installation, use fastPLS::fastPLS_blas() to verify the selected
CPU library, its version, and runtime CPU kernel. Use
fastPLS::fastPLS_blas(details = FALSE) only when a scalar backend name is
required. Older dynamic-architecture builds can
identify themselves as OpenBLAS while selecting a generic or legacy kernel on
a newer processor.

On Windows ARM64, do not point OPENBLAS_ROOT at an x86-64 Rtools or MSYS2
installation. The configuration rejects a cross-architecture archive. If no
ARM64 OpenBLAS development archive is available, leave FASTPLS_USE_OPENBLAS at
its default auto setting to use the BLAS/LAPACK supplied by the ARM64 R build.

CUDA installation
-----------------

CUDA support requires an NVIDIA GPU, a compatible NVIDIA driver and a separate
CUDA Toolkit installation. fastPLS never installs, removes, upgrades or
downgrades the host NVIDIA driver. In particular, do not install Debian's
`nvidia-cuda-toolkit` package merely to build fastPLS on a machine whose vendor
driver already works. Install a compatible NVIDIA CUDA Toolkit separately or
build in a controlled CUDA container that receives the host GPU through the
NVIDIA container runtime.

To request a CUDA build from source, set FASTPLS_USE_CUDA=1. Set CUDA_ROOT,
CUDA_HOME or CUDA_PATH when the toolkit is not installed in a standard
location. The configuration searches include, lib and lib64 directly below the
toolkit prefix and under targets/*/. It then compiles and links a probe against
CUDA Runtime, cuBLAS, cuSOLVER and cuRAND; finding nvcc alone is insufficient.

Examples:

  FASTPLS_USE_CUDA=1 CUDA_ROOT=/usr/local/cuda R CMD INSTALL fastPLS

On Windows, CUDA_ROOT can point to a toolkit directory such as:

  C:/Program Files/NVIDIA GPU Computing Toolkit/CUDA/v12.6

If FASTPLS_USE_CUDA is unset, unavailable CUDA software is ignored and the
package builds in CPU-only mode. Set FASTPLS_REQUIRE_CUDA=1 for strict mode; a
missing header, library, compiler or failed link probe then stops installation.
FASTPLS_CUDA_HOST_CXX may name an absolute system C++ compiler accepted by nvcc.
The default prefers /usr/bin/g++ or /usr/bin/clang++ over CUDA or Conda wrappers.
FASTPLS_CUDA_DIAGNOSTIC_ONLY=1 deliberately builds without CUDA kernels while
marking the installation as diagnostic-only; it is intended for build-system
diagnosis and must not be reported as functional CUDA support.

Apple Metal installation
------------------------

Metal support is available only on macOS. It is detected automatically from the
macOS SDK or system frameworks. To require a Metal build explicitly, set:

  FASTPLS_USE_METAL=1

After installation, run `fastPLS::has_metal()`. It returns TRUE only when the
package was compiled with Metal and the current R process can access a Metal
device. Set FASTPLS_USE_METAL=0 only to disable Metal deliberately.

CRAN and hosted build services
------------------------------

Binary packages contain only the capabilities present on the service that built
them. A CPU-only CRAN build remains fully functional. Users who require CUDA,
Metal, or a particular OpenBLAS installation should compile the source package
on the target computer and verify the result with `fastPLS_blas()`,
`cuda_info()`, `has_cuda()` and `has_metal()`.
