---
title: "Design & Reporting Guide"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Design & Reporting Guide}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include=FALSE}
knitr::opts_chunk$set(collapse = TRUE, comment = "#>")
library(contentvalidR)
read_example <- function(name) {
  utils::read.csv(
    system.file("extdata", name, package = "contentvalidR"),
    stringsAsFactors = FALSE
  )
}
```

## Choosing a method

Use the design that matches the question posed to judges rather than selecting an
index after data collection.

- **Item sort:** judges assign each item to the construct definition that best
  represents it. Use `sort_validity()` for Psa/Csv, exact target-count screening,
  strongest-competitor diagnostics, and scale-level Colquitt norms.
- **Construct rating:** the same judges rate each item against every focal and
  orbiting definition. Use `rating_validity()` for HTC/HTD, repeated-measures
  inference, planned target-versus-orbiting contrasts, and scale-level norms.
- **Expert panel:** experts answer a relevance, essentiality, or congruence
  question. Use `expert_validity()` with the corresponding explicit `mode`.

The three designs can complement one another during scale development, but their
statistics are not interchangeable.

## Sample-size planning

For item sorts, plan judge N in relation to the exact retention rule and a
plausible true target-assignment probability. `sort_power()` provides exact
planning probabilities; avoid a universal judge-count rule of thumb.

For construct-rating studies, power depends on the number of judges, number of
construct definitions, within-judge target-orbiting separation, and missing
profiles. Report the effective complete-judge N itemwise.

For expert panels, panel-size sensitivity is part of the statistic. CVR exact
critical counts and common CVI review guidelines therefore need to be interpreted
with the actual effective N, not a nominal panel size that ignores missingness.

## Recommended reporting sequence

1. **Design:** judge population, recruitment, construct definitions, item pool,
   instructions, and response format.
2. **A priori rules:** alpha, relevance threshold, target mapping, multiplicity
   adjustment if used, and any planned norms.
3. **Item-level evidence:** correspondence, distinctiveness, exact/paired
   inference, strongest competitor, and review status as appropriate.
4. **Scale-level evidence:** target-scale averages or S-CVI summaries where the
   method defines them.
5. **Substantive decisions:** revisions, removals, retained domain-coverage
   items, and how qualitative comments informed those decisions.
6. **Reproducibility:** package version, analysis settings, anonymized data when
   permitted, and the script used to reproduce tables/figures.

The `reporting-examples` vignette expands this sequence into reusable methods and
results scaffolds. Treat those examples as reporting patterns rather than fixed
language that must be copied verbatim.

## Deterministic bundled examples

```{r example-data}
sort_dat <- read_example("sort_example.csv")
rating_dat <- read_example("rating_example.csv")
expert_rel <- read_example("expert_relevance_example.csv")

sort_fit <- sort_validity(sort_dat)
rating_fit <- rating_validity(rating_dat, scale_min = 1, scale_max = 5)
expert_fit <- expert_validity(
  as.matrix(expert_rel[setdiff(names(expert_rel), "expert")]),
  mode = "relevance", lo = 1, hi = 4
)
```

These files are synthetic and generated by `data-raw/build-example-data.R` in the
source repository. They deliberately include both supported and review-worthy
items so documentation exercises realistic output paths without depending on
random-number generation.

## Table templates

### Item sort

```{r sort-table}
sort_fit$results[c(
  "item", "target", "n", "n_target", "competitor",
  "psa", "csv", "p_value", "status", "recommendation"
)]
```

At the target-scale level, report mean Psa/Csv and the benchmark set actually
used. Do not convert Colquitt's scale-level norms into individual-item cutoffs.

### Construct rating

```{r rating-table}
rating_fit$results[c(
  "item", "target", "n_complete", "strongest_competitor",
  "htc", "htd", "p_value", "max_contrast_p", "status", "recommendation"
)]
rating_fit$scale_summary
```

Report the repeated-measures design and target-versus-orbiting contrasts. For a
review item, naming the strongest competitor is often more informative than a
standalone p value.

### Expert relevance

```{r expert-table}
expert_fit$results[c(
  "item", "N", "V", "ci_low", "ci_high", "I_CVI",
  "kappa_mod", "status", "recommendation"
)]
expert_fit$scale_summary
```

Essentiality and congruence require different expert tasks. Do not place CVR,
CVI, Aiken V, and IOC in one generic threshold table as if they answer the same
question.

## Diagnostics and later validation

`signal_detection()` and `reproducibility_phi()` remain auxiliary helpers. They
can be useful when researchers later compare pretest decisions with CFA/IRT
retention or an independent replication pretest, but they are not required parts
of the flagship workflows.

```{r diagnostics}
pretest_supported <- sort_fit$results$status == "Supported"
later_retained <- c(TRUE, TRUE, TRUE, FALSE, TRUE, FALSE)
signal_detection(pretest_supported, later_retained)

replication_supported <- c(TRUE, TRUE, TRUE, FALSE, TRUE, TRUE)
reproducibility_phi(pretest_supported, replication_supported)
```

## Power quick check

```{r power}
sort_power(N = c(20, 30, 40), true_p = c(.60, .70, .80))
```

## Good practices

Pre-register quantitative screening criteria when feasible, preserve qualitative
judge feedback, report effective N after missingness, and archive the construct
definitions and item wording used in the pretest. A content-validation statistic
is evidence from a designed judgment task; it is not a substitute for defining
and sampling the construct domain.
