---
title: "R0: Foundations of the sportsfeatures Analytical Framework"
author: "Mohammad Abbas"
date: "`r Sys.Date()`"
vignette: >
  %\VignetteEncoding{UTF-8}
  %\VignetteIndexEntry{R0: Foundations of the sportsfeatures Analytical Framework}
  %\VignetteEngine{quarto::html}
format:
  html:
    toc: true
    toc-depth: 3
    toc-location: left
    number-sections: false
    theme: flatly
    code-fold: false
  pdf:
    toc: true
    toc-depth: 3
    pdf-engine: xelatex
    fig-cap-location: top
    tbl-cap-location: top
    number-sections: true
    colorlinks: true
    include-in-header:
      text: |
        \usepackage{tocloft}
        \usepackage{float}
        \floatplacement{figure}{H}
  docx:
    toc: true
    toc-depth: 3
    number-sections: false
    fig-cap-location: bottom
editor: 
  markdown: 
    wrap: 72
---

```{r setup, include=FALSE}
#| warning: false
#| message: false
knitr::opts_chunk$set(echo = TRUE)
library(dplyr)
library(ggplot2)
library(lme4)
library(lmerTest)
library(glmnet)
library(mice)
library(ranger)
library(here)

#| message: false 
#| warning: false
#| 

data("sports_features_missing", package = "sportsfeatures")

df_sports <- sports_features_missing %>%
type.convert(as.is = TRUE)

```

## Overview

The `sportsfeatures analytical framework` introduces a rule‑based
simulation environment designed to meet the challenges of modern
statistical training. The primary purpose of `sportsfeatures` is to
provide a realistic and interpretable environment for learning modern
statistical modelling and data science techniques. The framework enables
users to explore hierarchical data structures, missing-data mechanisms,
feature engineering, predictive modelling, and latent structure
discovery within a coherent sporting context.

## Why sportsfeatures?

Modern statistical education often relies on small, highly curated
datasets that illustrate individual methods in isolation. While useful
for teaching specific techniques, such datasets rarely capture the
complexities encountered in real analytical settings. Hierarchical
structures, missing data, behavioural variability, feature engineering,
and latent patterns are frequently absent or heavily simplified. The
sportsfeatures framework was designed to address this gap by providing a
realistic yet controlled simulation environment for statistical
learning. Rather than focusing on a single analytical objective, the
framework integrates multiple challenges commonly encountered in applied
data science, including multilevel data structures, non constant
variance, engineered features, and missing-data mechanisms. Sport and
human performance provide a particularly useful context for this
purpose. Athletic data naturally contain repeated observations,
individual baselines, environmental influences, and physiological
variability. These characteristics create a rich setting for exploring
modern modelling techniques while maintaining an intuitive connection to
real-world processes. The goal of sportsfeatures is not merely to
generate synthetic observations but to create an extensible analytical
framework. The resulting environment allows users to progress from
foundational concepts through prediction, explanation, discovery, and
dimensional reduction, forming the basis of the R-series vignette
pathway presented throughout this package.

## Core Premise

Stable Subject Baselines: Each athlete maintains stable physiological
characteristics.

Dynamic Session Evolution: Training sessions vary across environmental,
physiological, and personal conditions.

Controlled Realism: Incorporates domain rules, controlled randomness,
and a deliberate Missing Not At Random (MNAR) mechanism tied to device
usage.

### Framework Structure

**Phase 1** – Simulation Pipeline

Conceptual idea and physiological hypothesis

Rule‑based simulation engine

Core dataset: 500 observations × 30 variables

**Phase 2** – Feature Engineering & Data Enhancement

Enhanced dataset: 500 × 42 variables

Key analytical measures: efficiency, workload normalisation, baselines,
variability

**Framework Core** – Three Analytical Pillars

**Pillar 1**: Variance Structure – within‑ vs between‑athlete
fluctuations **Pillar 2**: MNAR Missingness – non‑random dropout due to
fatigue/injury **Pillar 3**: Derived Variables – feature engineering
transforming metrics into insights

### Vignette Research Pathway

R0 – Foundation → R1 – Predict → R2 – Explain → R3 – Discover → R4 –
Reduce

![Conceptual Overview of the R0
AnalyticalFramework](./images/Figure-1_Infographics.png){#fig-core-workframe}
Figure 1.1 Conceptual Overview of the R0 Analytical Framework

As shown in Figure 1.1, the multi-phase structure of the
`sportsfeatures` environment illustrates how rule‑based simulation, data
enhancement, and analytical pillars integrate to form the foundation for
the R‑series vignettes.

## Pillar 1: Variance Structure & Hierarchical Non‑Constant Spread

Human physiology and athletic behaviour are inherently dynamic and
non‑linear. When performance metrics are analysed using traditional
linear models, non‑constant variance (heteroscedasticity) naturally
emerges — a reflection of biological complexity rather than a modelling
flaw.

### Baseline Illustration

A simple Ordinary Least Squares (OLS) model relating energy expenditure
to session volume serves as a baseline to demonstrate non-constant
variance.

$$\text{Calories Burned} = \beta_0 + \beta_1(\text{Distance KM}) + \varepsilon$$

```{r}
df_lm <- df_sports %>%
  filter(!is.na(calories_burned), !is.na(distance_km))

model_simple <- lm(calories_burned ~ distance_km, data = df_lm)

```

![Bivariate Scatterplot of Calories Burned vs
DistanceKM](./images/Figure-2_model_simple.png){#fig-scatterplot} Figure
1.2 Bivariate Scatterplot of Calories(kcal) Burned vs Distance(km)

Figure 1.2 demonstrates a relatively strong positive trend between total
energy expenditure (kcal) and distance covered (km). However, looking
closely at the spread of data points, it can be seen that the spread is
fanning out, suggesting the presence of non-constant variance
(heteroscedasticity).

Figure 1.3 below confirms the presence of non-constant variance. The
plot also reveals that as training volume increases, the spread of
observations widens — signalling the influence of physiological,
behavioural, and environmental factors.

![Residual Diagnostics for Simple OLS
Model](./images/Figure-3_diagnostic_plot.png){#fig-diagnostics} Figure
1.3 Residual Diagnostics for Simple OLS Model

### Table 1.1: Simple Linear Regression Model Coefficients

| **Characteristic** | **Beta** | **95% CI** | **p-value** |
|:-------------------|:---------|:-----------|:------------|
| (Intercept)        | -29      | -95, 38    | 0.4         |
| Distance (km)      | 59       | 53, 65     | \<0.001     |

**Adjusted R² 47.6%**

While the simple regression model calories_burned \~ distance_km
captures a substantial proportion of the variation in calorie
expenditure (Adjusted R² = 47.6%), considerable unexplained variation
remains. Importantly, the model assumes that all observations are
independent and that any remaining variability can be treated as random
error. In practice, however, repeated training sessions are recorded
from the same athletes. Individual differences in fitness, physiology,
training habits, and recovery status introduce structure into the
residual variation. Consequently, part of the unexplained variance
reflects meaningful athlete-level differences rather than pure noise. To
distinguish within-athlete fluctuations from between-athlete
differences, the `sportsfeatures` framework adopts a hierarchical
perspective. This transition motivates the mixed-effects modelling
approaches explored in later vignettes, where sources of variation can
be explicitly partitioned and interpreted.

## Hierarchical Variance Decomposition

Because observations are nested within athletes, mixed-effects models
provide a natural mechanism for separating within-athlete variability
from between-athlete differences.

In the `sportsfeatures` framework, this non‑constant spread arises from
nested sources of variability:

Within‑Athlete vs Between‑Athlete Variation: Repeated sessions within
individuals differ from those across athletes with distinct fitness
baselines and metabolic efficiencies.

Contextual & Environmental Drivers: Terrain, weather, fatigue,
hydration, and readiness states introduce additional layers of variance.

Rather than treating these fluctuations as random noise,
`sportsfeatures` partitions them through hierarchical and mixed‑effects
modelling (see R1–R2), enabling structured interpretation of nested
dependencies.

This aligns with Figure 1.1, where Variance Structure forms the apex of
the framework’s three analytical pillars.

### Modelling Implications

Non-constant variance in `sportsfeatures` is not a defect but a hallmark
of realism. Traditional approaches often suppress heteroscedasticity
through transformations; here, it is embraced as evidence of biological
and behavioural variability across athletes and training contexts.

Recognising this variability has important modelling implications.
Statistical methods that account for hierarchical structure, repeated
observations, and athlete-specific effects are often more appropriate
than approaches that assume complete independence among observations.
Consequently, the framework provides a natural setting for exploring
mixed-effects models, variance decomposition, and athlete-level
heterogeneity in subsequent vignettes.

Within the `sportsfeatures` framework, variability is treated not merely
as noise to be removed but as information to be understood. This
perspective underpins the analytical progression from prediction (R1) to
explanation (R2), discovery (R3), and dimensional reduction (R4).

## Pillar 2: Missing Not At Random (MNAR) & MICE Imputation

Missing data is an inherent feature of real‑world sports telemetry. The
`sportsfeatures` framework deliberately embeds a deterministic Missing
Not At Random (MNAR) mechanism to mirror realistic sensor dropout and
athlete fatigue patterns.

MICE was selected because it preserves multivariate relationships across
mixed physiological metrics while quantifying uncertainty through
multiple imputations. Unlike single-value imputation approaches, MICE
generates multiple plausible datasets, allowing uncertainty associated
with missing values to be propagated into subsequent analyses. This
makes it particularly well suited to the `sportsfeatures` framework,
where physiological variables are often correlated and influenced by
common underlying processes.

### Wearable Device Dropout Mechanics

Whenever `device_type == "none"`, key physiological streams —
`heart_rate_avg`,`hydration_status` and `exhaustion_level` — are absent
by design. This configuration replicates common wearable data gaps
caused by device failure, non‑compliance, or extreme fatigue.

Across the full dataset, this mechanism introduces 97 missing values (≈
19.4%) per affected variable, creating a controlled environment for
testing imputation strategies.

### Imputation Workflow

To preserve the multivariate coherence of athletic performance data,
missing entries are handled using Multivariate Imputation by Chained
Equations (MICE) with predictive mean matching (PMM):

```{r}
#| output: false
#| message: false
#| warning: false
mice_vars <- df_sports %>%
  select(calories_burned, distance_km, duration_min,
         heart_rate_avg, exhaustion_level, hydration_status, device_type)

mice_model <- mice(mice_vars, m = 5, maxit = 20,
                   method = "pmm", seed = 123, quiet = TRUE)
```

This iterative process models each variable conditionally, drawing
plausible replacements that maintain physiological realism and
statistical consistency.

### Diagnostic Verification: Iterative Chain Convergence

The convergence plot (Figure 1.4) illustrates how the imputation chains
stabilize across 20 iterations for key metrics — `calories_burned`,
`heart_rate_avg`, and `hydration_status`. Left panels show mean
trajectories; right panels show standard deviations. Consistent
crossover among independent chains indicates healthy mixing and
convergence toward a stationary imputation distribution.

![Convergence Diagnostics for Mice Imputation
Chains](./images/Figure-4_Mice_convergence.png){#fig-Mice_convergence}
Figure 1.4 MICE Convergence Diagnostics

### Interpretation

- Mean & Variance Stabilization: The imputed streams interweave
  smoothly, confirming robust chain mixing and an absence of systemic
  drift.

- Physiological Coherence: The completed data vectors remain
  statistically valid and biologically plausible, ensuring unbiased
  downstream modelling.

- Framework Integration: This pillar complements Variance Structure
  (Pillar 1) by addressing data incompleteness — reinforcing the realism
  of the `sportsfeatures` simulation pipeline.

## Pillar 3: Derived Variables & Physiological Constructs

Where Pillar 1 explains variability and Pillar 2 addresses
incompleteness, Pillar 3 transforms raw measurements into interpretable
physiological constructs.

Derived variables form the dynamic layer of the `sportsfeatures`
framework. While core variables describe observable training outcomes
such as distance covered, session duration, or energy expenditure, they
provide only a partial view of athlete performance. Derived variables
extend these measurements by capturing underlying physiological
processes, workload characteristics, and athlete-specific deviations
from baseline behaviour.

Within the `sportsfeatures` framework, feature engineering serves as a
bridge between raw observations and analytical insight. By transforming
primary measurements into meaningful analytical features, derived
variables support prediction, explanation, discovery, and dimensional
reduction across the R-series vignettes.

Consequently, derived variables do not simply increase the number of
predictors available for analysis; they provide mechanisms through which
variability can be understood, modelled, and interpreted.

### Table 1.2 summarises the principal derived variables within the framework, along with their hierarchical classification and intended role in modelling.

Some variables capture session-level physiological responses, whereas
others represent athlete-level baselines or deviations from those
baselines. Together, they provide the analytical bridge between raw
observations and higher-level statistical interpretation.

| Derived Variable | Level Classification | Role in Modelling Process |
|--------------------|------------------|----------------------------------|
| vo~2~\_utilisation | Level‑1 (Session) | Normalises aerobic demand relative to VO₂_max; quantifies session intensity as a fraction of physiological capacity. |
| efficiency_index | Level‑1 (Session) | Captures cardiovascular economy by relating oxygen uptake to heart rate; used to assess aerobic efficiency and readiness. |
| hr_reserve | Level‑1 (Session) | Represents net cardiovascular elevation above resting state; proxy for internal load and autonomic strain. |
| aerobic_efficiency | Level‑1 (Session) | Indicates aerobic contribution per unit heart rate; evaluates aerobic system utilisation and metabolic economy. |
| fatigue_per_km | Level‑1 (Session) | Normalises fatigue accumulation by distance; models workload‑adjusted fatigue dynamics. |
| athlete_mean_fatigue | Level‑2 (Between‑Athlete) | Chronic fatigue baseline; separates stable athlete‑level fatigue differences from session‑level variation. |
| fatigue_deviation | Level‑1 (Centered) | Within‑athlete centred fatigue; isolates session‑specific deviations from an athlete’s personal fatigue norm. |
| athlete_mean_distance | Level‑2 (Between‑Athlete) | Chronic volume baseline; captures habitual training load differences between athletes. |
| distance_centered | Level‑1 (Centered) | Within‑athlete centred distance; models session‑specific deviations from typical training volume. |
| pace_min_km | Level‑1 (Session) | Represents external speed demand; used to model performance intensity and pacing strategy. |
| calories_per_km | Level‑1 (Session) | Normalises metabolic cost by distance; quantifies energetic efficiency and workload economy. |
| calories_per_min | Level‑1 (Session) | Represents metabolic burn rate; models internal metabolic intensity independent of distance. |

### Notes on Multilevel Classification

::: callout-note
- **Level‑1 (Session level):** varies *within* athlete across sessions
- **Level‑2 (Athlete level):** stable *between* athletes; group‑level
  baselines
- **Level‑1 (Centered):** session‑level deviations from Level‑2
  baselines
:::

![Evolution from Simulation to Feature
Engineering](./images/Figure-5_evolution.png){#fig-evo} Figure 1.5
Evolution from Simulation to Feature Engineering

As illustrated in Figure 1.5, the `sportsfeatures` framework evolves
through three stages:

- Simulation Dataset (Stage 1): establishes baseline synthetic
  behaviour.

- Core Expansion (Stage 2): introduces physiological depth through
  additional subject-level characteristics.

- Feature Engineering (Stage 3): generates twelve derived variables that
  transform raw measurements into analytical constructs.

The derived variables generated during Stage 3 serve distinct analytical
purposes within the framework. Some variables quantify physiological
efficiency, others capture workload intensity, chronic athlete
baselines, or session-specific deviations from those baselines.
Collectively, they provide the mechanisms through which variability can
be explored, explained, and modelled.

![Derived variables
map](images/Figure-6_derived_variables.png){#fig-derived} Figure 1.6
Derived Variables Functional Map

These constructs — such as `Heart Rate Reserve`, `Fatigue Deviation`,
`VO₂ Utilisation`, and `Efficiency Index` — act as mechanism generators.
They reveal the underlying processes driving variability in human
performance, allowing the framework to decompose, model, and interpret
fluctuations rather than treat them as noise.

The derived variables support sportsfeatures' transition from a dataset
to a fully functional framework, thereby enabling hierarchical
modelling, interaction profiling, and advanced statistical exploration
across vignettes R1 to R4.

### What Next? From Framework to Analytical Discovery

The `sportsfeatures` framework establishes a structured foundation that
transforms raw athletic telemetry into biological signals. With
individual baseline variance partitioned, missingness appropriately
mapped, and 12 feature domains fully established, `sportsfeatures`
evolves from a dataset into an extensible analytical framework.

![Framework to analytical
discovery](images/./Figure-7_infographics.png){#fig-discovery} Figure
1.7 Framework to Analytical Discovery

This sets the stage for a four-part empirical progression across **R1
through R4**, as illustrated in Figure 1.7:

- **R1 (Predict):** Evaluating mechanical efficiency against athlete
  readiness using predictive multilevel frameworks.

- **R2 (Explain):** Deconstructing cardiorespiratory cost and
  partitioning within- versus between-athlete variance.

- **R3 (Discover):** Uncovering hidden physiological taxonomies through
  unsupervised profiling.

- **R4 (Reduce):** Isolating key latent dimensions across dynamic
  training variables using factor reduction.

### Concluding Remarks

R0 establishes the conceptual foundations of the `sportsfeatures`
framework by introducing its simulation architecture, variance
structure, missing-data mechanisms, and derived-variable ecosystem.
Together, these components create a coherent analytical environment that
supports increasingly sophisticated statistical investigations. The
subsequent vignettes build upon this foundation, progressing from
prediction and explanation to discovery and dimensional reduction within
a unified modelling framework.