---
title: "Item-Sort Content Validation: Anderson-Gerbing to Howard-Melloy to Colquitt"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Item-Sort Content Validation: Anderson-Gerbing to Howard-Melloy to Colquitt}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include=FALSE}
knitr::opts_chunk$set(collapse = TRUE, comment = "#>")
library(contentvalidR)
```

## What this workflow answers

Item-sort pretests ask judges to assign candidate items to construct definitions.
The workflow provides two related kinds of evidence:

1. **Definitional correspondence**: are items assigned to their intended construct?
2. **Definitional distinctiveness**: are items assigned to the intended construct more often than to a competing construct?

Anderson and Gerbing (1991) operationalized these ideas with **Psa** and **Csv**.
Howard and Melloy (2016) clarified exact inference for the target-assignment count,
particularly when more than two assignment alternatives are present. Colquitt et
al. (2019) later supplied empirical interpretation norms based on 112 published
scales.

These statistics do not establish the entire content-validity argument. In
particular, they do not establish that the item pool comprehensively covers the
construct domain.

## A reproducible example

```{r}
sort_dat <- data.frame(
  item = rep(c("A1", "A2", "A3", "B1", "B2", "B3"), each = 20),
  rater = rep(1:20, 6),
  target_construct = rep(c("A", "A", "A", "B", "B", "B"), each = 20),
  assigned_construct = c(
    rep("A", 18), rep("B", 2),
    rep("A", 16), rep("B", 4),
    rep("A", 13), rep("B", 7),
    rep("B", 18), rep("A", 2),
    rep("B", 17), rep("A", 3),
    rep("B", 14), rep("A", 6)
  )
)

fit <- sort_validity(sort_dat)
fit
```

The item table is intentionally diagnostic rather than merely numeric. A
`Review` flag is not a command to delete an item. The output reports the
strongest competing construct so that researchers can distinguish weak target
correspondence from specific construct overlap.

```{r}
summary(fit)
```

## Item-level inference: Howard-Melloy

The default exact test asks whether the target-assignment probability exceeds
`.50`. At `N = 20` and `alpha = .05`, an item needs 15 target assignments to
meet the one-sided exact criterion.

```{r}
csv_binom_test(n_c = 15, N = 20)
csv_binom_test(n_c = 14, N = 20)
```

`sort_validity()` therefore uses **Retain** to mean "meets this exact statistical
screening criterion" and **Review** to mean "does not meet it." Revision or
removal remains a substantive decision.

## Scale-level interpretation: Colquitt et al. (2019)

Colquitt et al. did not create their interpretation bands from individual item
values. They averaged Psa and Csv across the items in each of 112 scales and
then created empirical percentile bands. `sort_validity()` follows that design:
Howard-Melloy is used item by item, while Colquitt interpretation is reported for
each target scale's mean Psa and mean Csv.

The default uses the overall norms:

```{r}
colquitt_benchmarks("psa")
colquitt_benchmarks("csv")
```

The labels—Very Strong, Strong, Moderate, Weak, and Lack of—are empirical
normative standing, **not universal validity cutoffs**.

### Correlation-conditional norms

Colquitt et al. showed that Psa/Csv depend partly on how similar the focal scale
is to its orbiting scales. If substantive data provide an average focal-orbiting
correlation, supply it to the workflow. With multiple focal scales, use a named
vector.

```{r}
fit_normed <- sort_validity(
  sort_dat,
  orbiting_r = c(A = .42, B = .28)
)
fit_normed$scale_summary
```

The conditional panels are:

- `.34` or below: weaker focal-orbiting correlation;
- `.35` to `.50`: more moderate correlation;
- `.51` or above: stronger correlation.

A given Csv can be more impressive when the focal and orbiting constructs are
closely related, so the appropriate norm can change the descriptive category.

## Judge type matters

Anderson and Gerbing advocated naïve judges representative of the population of
interest, and Colquitt et al.'s norms were generated with that kind of judge.
Their criteria should not simply be transferred to expert panels.

```{r}
expert_fit <- sort_validity(sort_dat, judge_type = "expert")
expert_fit$scale_summary
```

The Psa/Csv statistics and item-level screening remain available, but Colquitt
normative labels are suppressed.

## Planning judge sample size

Use `sort_power()` to calculate the exact probability that an item will reach
the required target-assignment count under a plausible true target-assignment
probability.

```{r}
sort_power(N = c(20, 30, 40), true_p = c(.60, .70, .80))
```

This is preferable to treating a rule such as "20-40 judges" as a universal
sample-size requirement. Power depends on the assumed target-assignment
probability, `N`, the null probability, and alpha.

## Plotting item evidence

The one-index views remain available:

```{r fig.width=7, fig.height=4}
plot(fit, metric = "psa")
plot(fit, metric = "csv")
```

For diagnosis, the package also introduces a **correspondence-distinctiveness
evidence map**:

```{r fig.width=7, fig.height=5}
plot(fit, type = "map")
```

Psa and Csv are shown jointly, review items are labeled by default, and target-
scale averages are added as diamonds. This makes it easier to distinguish a
correspondence problem (low Psa) from a construct-overlap problem (low or
negative Csv). The map does not draw Colquitt cutoff regions across individual
items because those empirical norms were constructed from scale-level averages.

The exact planning object is also plottable:

```{r fig.width=7, fig.height=4}
plan <- sort_power(N = seq(10, 50, by = 5), true_p = c(.60, .70, .80))
plot(plan)
plot(plan, type = "critical")
```

No conventional target-power line is imposed unless the analyst supplies one.

## Why the Anderson-Gerbing legacy critical-Csv rule is not exposed

`contentvalidR` retains Anderson and Gerbing's Psa and Csv indices but does not
provide their legacy critical-Csv decision rule as a user-selectable alternative.
Howard and Melloy (2016) showed that the older rule is appropriate for the original
two-choice case but becomes miscalibrated when it is applied to sorts with more
than two construct choices. Their revised target-count procedure agrees with the
legacy logic in the two-choice case and is applicable to the broader designs now
used in practice. Exposing the obsolete rule would therefore add a reproducibility
option that is easy to misuse without adding a recommended analysis path.

## Reporting

A useful report should include:

- who the judges were and why they fit the intended design;
- the focal and orbiting constructs and their definitions;
- the number of judges and missing assignments;
- item-level Psa, Csv, target counts, strongest competitors, and exact decisions;
- target-scale mean Psa/Csv and the Colquitt norm set used;
- focal-orbiting correlations if conditional norms were used; and
- the substantive reasoning behind any revisions or removals.

## References

Anderson, J. C., & Gerbing, D. W. (1991). Predicting the performance of measures
in a confirmatory factor analysis with a pretest assessment of their substantive
validities. *Journal of Applied Psychology, 76*(5), 732-740. https://doi.org/10.1037/0021-9010.76.5.732

Howard, M. C., & Melloy, R. C. (2016). Evaluating item-sort task methods: The
presentation of a new statistical significance formula and methodological best
practices. *Journal of Business and Psychology, 31*(1), 173-186. https://doi.org/10.1007/s10869-015-9404-y

Colquitt, J. A., Sabey, T. B., Rodell, J. B., & Hill, E. T. (2019). Content
validation guidelines: Evaluation criteria for definitional correspondence and
definitional distinctiveness. *Journal of Applied Psychology, 104*(10),
1243-1265. https://doi.org/10.1037/apl0000406
