diff --git a/NEWS.md b/NEWS.md index 975720a..ea54c05 100644 --- a/NEWS.md +++ b/NEWS.md @@ -29,6 +29,21 @@ No computed value changes. `stability = "percent_change"`. The 1975 printing's pages could not be checked. +## The handoff contract, stated for readers + +Asked for by nomologR before handoff schema 1 becomes a 1.0 promise. The +schema and every handoff are unchanged. + +* `?content_handoff` has a new section, "Reading the decisions": take each + item's decision from `carried` and `status`, never re-derive it by + comparing `value` with `criterion` or by branching on the producer version. + The example is the nine-expert correction in 0.8.0, where an item's I-CVI + stayed at .778 while its decision changed. The section also restates that + `keying` of `NA` means unknown, never forward-worded. +* `?content_handoff` says an item sits in at most one scale, the one its + `item_evidence$scale` names. A test holds the sort, rating, and congruence + workflows to it, and checks that each refuses an item with two targets. + # contentvalidR 0.9.0 Ninth public release. v0.9.0 is about how the package presents itself before diff --git a/R/content_handoff.R b/R/content_handoff.R index 42b852a..798f986 100644 --- a/R/content_handoff.R +++ b/R/content_handoff.R @@ -498,7 +498,10 @@ #' \item{`scales`}{named list mapping each construct to its carried items, or #' `NULL` when the design has no construct mapping. Expert relevance and #' essentiality rate a single item set with no construct column, so they -#' produce `NULL`. Membership is one to one.} +#' produce `NULL`, as does a Delphi study. An item belongs to at most one +#' scale, the one its `item_evidence$scale` names: every workflow that maps +#' items to constructs requires exactly one target per item, and stops +#' otherwise.} #' \item{`item_evidence`}{data frame with one row per reviewed item: `item`, #' `scale` (`NA` without a construct mapping), `carried`, `status`, #' `recommendation`, `n_judges`, `rule`, and `round`. For a Delphi handoff @@ -530,6 +533,29 @@ #' Fields and columns added within schema version 1 are optional for a reader, #' which should check that they are present rather than assume it. #' +#' @section Reading the decisions: +#' Take each item's decision from the handoff, not from its statistics. +#' `carried` says whether the item travels forward, and `status` says why: it +#' is one of the four values `keep` accepts. `recommendation` words the same +#' decision for a person, and like `rule` it is prose. +#' +#' Don't re-derive a decision by comparing `value` with `criterion`, and don't +#' branch on `provenance$package_version`. A corrected rule can change a +#' decision between releases while the statistics stay the same, and the +#' handoff records the decision its producer made. contentvalidR 0.8.0 is the +#' example. For an item that 7 of 9 experts rated relevant, the I-CVI is .778 +#' in both releases. 0.7.0 compared it with a rounded .78 and held the item +#' back. 0.8.0 applies Lynn's (1986) 7 of 9, stores 7/9 as the criterion, and +#' carries the item. There the value equals the criterion, so a recomputed +#' `value >= criterion` would hang on floating-point rounding, where the +#' package itself compares counts of experts. Read `carried`, which holds what +#' each release decided. +#' +#' Keying is read the same way, from the field and not from a default: +#' `keying` is `1` for a forward-worded item, `-1` for a reverse-worded one, +#' and `NA` when nobody said, as described under "Instrument metadata". Treat +#' `NA` as unknown, never as forward-worded. +#' #' @section What version 1 freezes: #' Schema version 1 is frozen as of contentvalidR 0.7.0. Code that reads a #' handoff can rely on all of the following, in every release that reports @@ -726,6 +752,11 @@ #' empirical analysis tests, which is why the item set travels with its #' evidence rather than as a bare list of names. #' +#' @references +#' Lynn, M. R. (1986). Determination and quantification of content validity. +#' *Nursing Research, 35*(6), 382–385. +#' \doi{10.1097/00006199-198611000-00017} +#' #' @param fit A fitted `contentvalid_sort`, `contentvalid_rating`, #' `contentvalid_expert`, or `contentvalid_delphi` object. #' @param keep Statuses that travel forward, defaulting to `"Supported"`. Any of diff --git a/ROADMAP.md b/ROADMAP.md index 879d8c4..06cfc4d 100644 --- a/ROADMAP.md +++ b/ROADMAP.md @@ -810,8 +810,8 @@ reads no handoffs and takes no part in the joint 1.0. 0.7.0. - [x] schema version 1 is frozen and becomes a compatibility promise (0.7.0). - [ ] both packages are on CRAN. Waiting on CRAN's review of 0.4.0, which - has been in the `newbies` queue since 2026-09-18, and on nomologR's - first submission. + has been in the `newbies` queue since 2026-09-18, and of nomologR + 0.3.0, submitted on 2026-09-26. - [x] both public APIs are stable, under a written deprecation policy, as far as this package goes (0.7.0). `anova_content()`'s `posthoc_pass` column was removed in 0.8.0. `agreement_summary()`, the one function @@ -820,6 +820,32 @@ reads no handoffs and takes no part in the joint 1.0. - [x] there is a joint walkthrough from content review to empirical validation (0.7.0). - [ ] the two releases go out on the same day, each linking the other. +- The release-candidate checklist, proposed by nomologR and approved by the + maintainer on 2026-09-27. It adds to #53's criteria rather than replacing + them: + - [ ] a final fixture set from each release candidate, with a manifest of + md5 sums, that nomologR's tests pass before either package tags. + - [ ] the compatibility table in both READMEs: contentvalidR 1.x writes + handoff schema 1, and nomologR 1.x reads it. It goes in at the release + candidate, so neither README describes a 1.x before one exists. + - [ ] one handoff example shown in both READMEs, regenerated from + contentvalidR's release candidate. + - [ ] the joint walkthrough lives in one package, and the other links to + it. Recommended, and acceptable to nomologR: this package's + `vignette("one-item-set-both-stages")`, linking to nomologR's + guided-workflow article for the empirical side. The maintainer + decides. + - [ ] the release notes link each other. +- Three handoff clarifications nomologR asked for before schema 1 becomes a + 1.0 promise, approved on 2026-09-27. None changes the schema: + - [x] a fixture that records keying as checked with no item reversed + (`reverse_keyed = character(0)`, so `keying` is `1` throughout), + generated from v0.9.0. + - [x] `?content_handoff` says an item sits in at most one scale, and a test + holds every workflow to it. + - [x] `?content_handoff` says a reader takes each decision from `carried` + and `status`, never re-deriving it from the statistics or the + producer version. *** diff --git a/man/content_handoff.Rd b/man/content_handoff.Rd index 59e08bb..47fa7c2 100644 --- a/man/content_handoff.Rd +++ b/man/content_handoff.Rd @@ -68,7 +68,10 @@ consumer matches on \code{"cv_handoff"} and reads these fields: \item{\code{scales}}{named list mapping each construct to its carried items, or \code{NULL} when the design has no construct mapping. Expert relevance and essentiality rate a single item set with no construct column, so they -produce \code{NULL}. Membership is one to one.} +produce \code{NULL}, as does a Delphi study. An item belongs to at most one +scale, the one its \code{item_evidence$scale} names: every workflow that maps +items to constructs requires exactly one target per item, and stops +otherwise.} \item{\code{item_evidence}}{data frame with one row per reviewed item: \code{item}, \code{scale} (\code{NA} without a construct mapping), \code{carried}, \code{status}, \code{recommendation}, \code{n_judges}, \code{rule}, and \code{round}. For a Delphi handoff @@ -101,6 +104,31 @@ Fields and columns added within schema version 1 are optional for a reader, which should check that they are present rather than assume it. } +\section{Reading the decisions}{ + +Take each item's decision from the handoff, not from its statistics. +\code{carried} says whether the item travels forward, and \code{status} says why: it +is one of the four values \code{keep} accepts. \code{recommendation} words the same +decision for a person, and like \code{rule} it is prose. + +Don't re-derive a decision by comparing \code{value} with \code{criterion}, and don't +branch on \code{provenance$package_version}. A corrected rule can change a +decision between releases while the statistics stay the same, and the +handoff records the decision its producer made. contentvalidR 0.8.0 is the +example. For an item that 7 of 9 experts rated relevant, the I-CVI is .778 +in both releases. 0.7.0 compared it with a rounded .78 and held the item +back. 0.8.0 applies Lynn's (1986) 7 of 9, stores 7/9 as the criterion, and +carries the item. There the value equals the criterion, so a recomputed +\code{value >= criterion} would hang on floating-point rounding, where the +package itself compares counts of experts. Read \code{carried}, which holds what +each release decided. + +Keying is read the same way, from the field and not from a default: +\code{keying} is \code{1} for a forward-worded item, \code{-1} for a reverse-worded one, +and \code{NA} when nobody said, as described under "Instrument metadata". Treat +\code{NA} as unknown, never as forward-worded. +} + \section{What version 1 freezes}{ Schema version 1 is frozen as of contentvalidR 0.7.0. Code that reads a @@ -334,6 +362,11 @@ handoff$item_statistics # Carry items flagged for review as well, when the study protocol says so. content_handoff(fit, keep = c("Supported", "Review"))$items } +\references{ +Lynn, M. R. (1986). Determination and quantification of content validity. +\emph{Nursing Research, 35}(6), 382–385. +\doi{10.1097/00006199-198611000-00017} +} \seealso{ \code{\link[=as.data.frame.contentvalid_workflow]{as.data.frame.contentvalid_workflow()}} for the full results table, and \code{\link[=content_report]{content_report()}} for manuscript tables. diff --git a/tests/testthat/test-content-handoff.R b/tests/testthat/test-content-handoff.R index 82293be..bab6e79 100644 --- a/tests/testthat/test-content-handoff.R +++ b/tests/testthat/test-content-handoff.R @@ -155,6 +155,48 @@ test_that("statistics stack long with their criteria", { expect_setequal(unique(st$item), as.character(fit$results$item)) }) +test_that("an item sits in at most one scale, the one its evidence row names", { + fits <- list(sort = sort_fit(), rating = rating_fit(), + congruence = congruence_fit()) + for (workflow in names(fits)) { + h <- content_handoff(fits[[workflow]], keep = c("Supported", "Review")) + members <- unlist(h$scales, use.names = FALSE) + expect_equal(anyDuplicated(members), 0L, info = workflow) + carried <- h$item_evidence[h$item_evidence$carried, , drop = FALSE] + expect_true(nrow(carried) > 0, info = workflow) + for (i in seq_len(nrow(carried))) { + expect_true(carried$item[i] %in% h$scales[[carried$scale[i]]], + info = paste(workflow, carried$item[i])) + } + } + + # A handoff cannot place an item in two scales, because every workflow that + # maps items to constructs refuses an item with two targets. + sort_dat <- data.frame( + item = rep(c("A1", "B1"), each = 4), rater = rep(1:4, 2), + target_construct = c("A", "A", "A", "B", rep("B", 4)), + assigned_construct = c(rep("A", 4), rep("B", 4)), + stringsAsFactors = FALSE + ) + expect_error(sort_validity(sort_dat), "exactly one target construct") + + d <- expand.grid(item = c("A1", "B1"), rater = 1:4, construct = c("A", "B"), + stringsAsFactors = FALSE) + d$target_construct <- ifelse(d$item == "B1", "B", "A") + d$target_construct[d$item == "A1" & d$rater == 1] <- "B" + d$rating <- ifelse(d$construct == d$target_construct, 5, 2) + expect_error(rating_validity(d, scale_min = 1, scale_max = 5), + "exactly one target construct") + + con <- expand.grid(item = c("I1", "I2"), judge = 1:4, objective = c("A", "B"), + KEEP.OUT.ATTRS = FALSE, stringsAsFactors = FALSE) + con$score <- 1 + con$target_objective <- ifelse(con$item == "I1", "A", "B") + con$target_objective[con$item == "I1" & con$judge == 1] <- "B" + expect_error(expert_validity(con, mode = "congruence"), + "exactly one target objective") +}) + test_that("essentiality and congruence modes hand off", { ess <- content_handoff(expert_validity(c(10, 8, 6), mode = "essentiality", N = 12), keep = c("Supported", "Review")) @@ -721,3 +763,24 @@ test_that("instrument metadata covers held-back items too", { expect_false(anyNA(ev$keying)) expect_true(all(ev$response_max == 7L)) }) + +test_that("status is always one of the values keep accepts", { + valid <- .status_definitions()$status + handoffs <- list( + content_handoff(expert_fit()), + content_handoff(sort_fit()), + content_handoff(rating_fit()), + content_handoff(congruence_fit()), + content_handoff(congruence_fit(target = FALSE), keep = "Descriptive only"), + content_handoff(expert_validity(c(10, 8, 6), mode = "essentiality", N = 12)), + content_handoff(delphi_fit(B = 0)) + ) + for (h in handoffs) { + expect_true(all(h$item_evidence$status %in% valid), + info = h$provenance$workflow) + } + # `carried` is the decision; it follows `status` and `keep` exactly. + h <- content_handoff(sort_fit(), keep = c("Supported", "Review")) + expect_identical(h$item_evidence$carried, + h$item_evidence$status %in% c("Supported", "Review")) +})