Skip to content

Add figures for papers, posters, and teaching: content_evidence() and the distribution view - #82

Merged
JUhalt merged 2 commits into
masterfrom
feat/evidence-displays
Sep 28, 2026
Merged

JUhalt merged 2 commits into
masterfrom
feat/evidence-displays

Conversation

@JUhalt

@JUhalt JUhalt commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

The three displays the maintainer chose on 2026-09-28, prototyped on the walkthrough items and approved to go in before 1.0. No computed value changes. Each figure draws what the fits and handoffs already decided, and its help page says whether it follows a published form or is this package's own design. Graphics are the one area where the maintainer welcomes going beyond the literature.

content_evidence() (new)

It brings the handoffs from several review stages together, in the order they ran, such as a relevance panel and then an item sort. Each stage can be a handoff or a fitted workflow.

  • Print: it opens with the verdict ("10 of 12 items carried by every stage that reviewed them. Held back: EF5 (Relevance panel), TF5 (Item sort).").
    • It lists the stages.
    • One row per item gives each stage's decision beside the statistic its rule read, and a result column.
    • A key follows, which contentvalidR.show_key hides.
  • plot(x, type = "flow"): the item flow diagram, modelled on the PRISMA 2020 flow diagram (Page et al., 2021).
    • It shows what each stage reviewed, what it held back with the number behind each decision (for example "EF5 (Review): I-CVI .50, criterion .88"), and what went forward, by construct.
    • Stages can run in sequence or side by side. A stage that re-reviews an item an earlier stage already held back says so, rather than counting the item twice.
  • plot(x): the item evidence profile, this package's own design, laid out like a forest plot.
    • There is one panel per stage, with each item's statistic, its interval and its criterion.
    • A last column names the stages that held an item back.
    • It shows at a glance where methods agree and where they disagree. In the walkthrough, TF5 is relevant to every expert, and only the sort catches it.
  • The statistic each stage shows is the one its rule reads:
    • Psa for an item sort;
    • I-CVI for a relevance panel;
    • the share agreeing for a Delphi study;
    • CVR for an essentiality panel;
    • the IOC margin (or IOC) for congruence ratings;
    • HTC for construct ratings, with no criterion line, since that workflow decides on contrasts.
  • Decisions come from each handoff, never recomputed. A test tampers with a handoff's carried and checks that the display believes the handoff, per the reader contract in ?content_handoff.
  • Tier 1: it is added to the Tier 1 list in ?contentvalidR, so it becomes a stability promise at 1.0. as.data.frame() returns one row per item per stage.

The distribution view

  • plot(fit, type = "distribution") for relevance fits, and plot(fit, which = "distribution") for Delphi fits, with one bar per round.
  • Every rating is drawn as diverging stacked bars (Heiberger & Robbins, 2014), split at the cut the decision rule counts. The right-hand length is exactly the I-CVI, or the share agreeing, read against the dashed criterion. The symbol beside each bar is the decision the fit made.
  • It shows what the index can't: two items at an I-CVI of 1.00, one rated relevant with 4s and the other with 3s.
  • Relevance fits now keep their ratings in details$ratings. That's an addition; the handoff is unchanged. A fit from an earlier version gets a clear error asking for a refit.
  • labels sets the category names in the key.

apa = TRUE / FALSE

  • TRUE (the default) draws in gray, as an APA figure is printed. In the bars, darker means a higher rating.
  • FALSE uses a colorblind-safe brown–teal scheme for slides and posters: teal for evidence that met its criterion, brown for evidence under review.
  • Filled and open symbols carry the decision in both, so no reading depends on color.
  • The "item" views stay gray. The Delphi consensus and stability views keep coloring each item's line, as before.

Data and docs

  • walkthrough_relevance.csv: eight experts' ratings of the twelve walkthrough items, constructed by hand. It comes from its own script, data-raw/build-walkthrough-panel.R, which says what each item's ratings are built to show. The three walkthrough files nomologR ships are unchanged. Tests hold the file to what the script claims.
  • Vignettes:
    • vignette("reporting-examples") gains "Figures across review stages": both figures on the walkthrough, a color version, and a sample APA caption.
    • The expert-panel and Delphi vignettes show the distribution view.
    • Every new figure has alt text.
  • References: Heiberger & Robbins (2014) was checked against Crossref and the JSS article page (pp. 1–32). Page et al. (2021) was checked against Crossref and PubMed (volume 372, article n71). Both are in the README and REFERENCES.bib. The 26-author reference follows APA 7: the first 19 authors, an ellipsis, then the last. The vignette reference lists are updated.
  • In passing: the handoff printout and vignette("handoff-to-empirical-validation") now pass the whole handoff, nomo_screen(data, items = handoff), rather than handoff$items. That vignette no longer calls nomologR's handoff reader future work, and nomologR confirmed the calls.

The release gate, now that contentvalidR is on CRAN

contentvalidR 0.4.0 was published on CRAN on 2026-09-28. So R CMD check --as-cran now says "Days since last update" in its incoming-feasibility note where it said "New submission". The gate rejected that unfamiliar note, as it should, so the second commit changes the allow-list deliberately:

  • The incoming note passes only when every line in it is one we expect: the maintainer, "New submission", "Version contains large components", or "Days since last update".
  • Any other line still fails the stage. A test run confirmed that a misspelling or an invalid URL in that note is still caught.
  • When the last update was under 60 days ago, the stage prints a reminder that CRAN asks for updates no more often than every one to two months.
  • tools/README.md says the same.

Verification

  • All 3,060 tests pass, and spelling is clean.
  • The rendered figures, in gray and in color, were checked by eye: profile, flow in sequence and side by side, relevance distribution, and Delphi distribution.
  • tools/release-gate.R on this PR's head (2e66a4a), run in a detached worktree: build, smoke, and check all pass. R CMD check: 0 errors, 0 warnings, and only the two expected notes (CRAN incoming feasibility, now "Days since last update"; math rendering skipped without V8). The gate's new CRAN reminder printed as designed. An earlier run on 69225e3 also flagged content_structure's example at 5.02 s elapsed but 0.02 s CPU, a machine stall; it did not recur.

🤖 Generated with Claude Code

JUhalt and others added 2 commits September 28, 2026 11:50
… the distribution view

The maintainer asked for three displays before 1.0. Each draws what the
fits and handoffs already decided, and none computes anything new.

content_evidence() (new, Tier 1) brings the handoffs from several review
stages together. It prints one row per item, with each stage's decision
beside the statistic its rule read, and draws two figures:
- the item flow diagram, modelled on the PRISMA 2020 flow diagram (Page et
  al., 2021): what each stage reviewed, what it held back and the number
  behind each decision, and what went forward. Stages may run in sequence
  or side by side;
- the item evidence profile, the package's own design: one panel per stage
  with each statistic, its interval, and its criterion, and a column naming
  the stages that held an item back.
Decisions are read from each handoff, never recomputed, as the reader
contract in ?content_handoff asks.

plot(<relevance fit>, type = "distribution") and plot(<Delphi fit>,
which = "distribution") draw every rating as diverging stacked bars
(Heiberger & Robbins, 2014), split at the cut the decision rule counts, so
two items with the same I-CVI can be told apart. Relevance fits keep their
ratings in details$ratings to draw it.

All three take `apa`: grey by default, as an APA figure is printed, or a
colour-blind-safe scheme with teal for evidence that met its criterion and
brown for evidence under review. Symbols carry the decision either way.

walkthrough_relevance.csv gives the twelve walkthrough items a relevance
panel, from its own script, so the three files nomologR ships are
unchanged. vignette("reporting-examples") uses it to show both figures,
with a sample caption; the expert-panel and Delphi vignettes show the
distribution view.

In passing: the handoff printout and the handoff vignette now pass the
whole handoff to nomologR rather than handoff$items, and the vignette no
longer calls nomologR's handoff reader future work.

Heiberger & Robbins (2014) and Page et al. (2021) are checked against
Crossref, the journal pages, and PubMed, and added to the README and
REFERENCES.bib.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
contentvalidR 0.4.0 was published on CRAN on 2026-09-28. R CMD check
--as-cran now reports "Days since last update" in its incoming-feasibility
note where it reported "New submission", so the gate failed a clean
check on a note it had never seen.

The incoming note is now tolerated only when every line of it is
expected: the maintainer, "New submission", "Version contains large
components" (a development version), or "Days since last update". Any
other line in it, such as a misspelling or an invalid URL, still fails
the stage, as does any other note. When the last update was under 60 days
ago, the stage prints a reminder that CRAN asks for updates no more often
than every one to two months. tools/README.md says the same.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@JUhalt
JUhalt merged commit 3371656 into master Sep 28, 2026
8 checks passed
JUhalt added a commit that referenced this pull request Sep 28, 2026
Release contentvalidR 0.10.0

Stamps the version, release notes, citation metadata, README, and roadmap
for the tenth public release, the last planned before the joint 1.0
release candidate on 2026-10-17. It adds figures for papers, posters, and
teaching: content_evidence() with its item flow diagram and item evidence
profile, the rating-distribution view, and the apa switch (#82). It also
states the handoff contract nomologR reads (#79), makes this package's
walkthrough the joint one (#80, #81), separates the Delphi printout's
caveats from its teaching (#78), and records the release-candidate
timeline (#83).

No computed value changes: 25 analyses across every workflow return the
same values as v0.9.0. No breaking changes: the plot methods gain
arguments after their existing ones. agreement_summary() remains
deprecated and is removed in 1.0.0.

contentvalidR 0.4.0 has been on CRAN since 2026-09-28. CRAN asks for
updates no more often than every one to two months, so 0.10.0 is released
on GitHub and R-universe only.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant