diff --git a/.audit-work/changes.md b/.audit-work/changes.md deleted file mode 100644 index 5f6b547..0000000 --- a/.audit-work/changes.md +++ /dev/null @@ -1,127 +0,0 @@ -# 修改账本 - -本账本描述 v2 发布硬化及首次远端 CI follow-up。release commit `cfad783` 已合并并 push;follow-up 本地门禁通过,但最终完成以第二轮远端 CI 成功为条件。 - -## CHG-001|AUD-001 - -- 范围:BL19B2 safe PONI 身份与续跑。 -- 修改前:存在旧 safe PONI 时可能静默复用;pydidas 路径可能直接覆盖。 -- 修改后:先渲染并比较内容/哈希;几何不一致时 fail-closed,不改变既有 safe PONI。 -- 兼容性:相同几何续跑不变。 -- 验证:普通 PONI、pydidas Cali.yaml 的 red→green 与回归。 -- 回退风险:移除一致性门会重新允许 provenance 错配。 - -## CHG-002|AUD-002/AUD-011/AUD-012 - -- 范围:Workbench Cal2D 五件套、事务提交、resume 与 rerun。 -- 修改前:仅主 EDF 存在即可 skip;写入中断会留残包;并发冲突 rollback 可能误删后来者文件。 -- 修改后:image/mask NPY/mask EDF/PONI/metadata 及 shape/context/参数联合验证;staging 后以 create-if-absent 发布;no-overwrite 后续冲突不删除已发布成员,残片由完整性门拒绝;overwrite 模式使用 backup rollback;整包使用同一 rerun ID。 -- 兼容性:完整且身份一致的包仍可 resume;残缺、损坏、错配或竞争目标明确失败。 -- QA 补充:Cal2D rollback 竞态闭环归入 AUD-011。 -- 验证:fault injection、并发替换、残包、目录 PONI、package-level rerun 聚焦测试。 -- 回退风险:会重新出现混合科学包或外部文件误删风险。 - -## CHG-003|AUD-003 - -- 范围:calibration 与 buffer subtraction 公共 Python API。 -- 修改前:新增参数插入历史位置参数顺序,旧合法调用确定性错位。 -- 修改后:恢复旧位置参数顺序;新增参数设为 keyword-only。 -- 兼容性:旧位置调用恢复;新调用语义保持显式。 -- 验证:旧调用反例、关键字新调用和全量回归。 - -## CHG-004|AUD-004/AUD-005/AUD-013 - -- 范围:BL19B2 CLI、Config、重跑脚本、签名与版本迁移。 -- 修改前:危险旧默认值隐式存在;mask、执行策略和非默认科学字段不能完整重放;monitor mode 规范化分叉。 -- 修改后:版本升为 2.0.0;安全 v2 入口要求显式参数;`bl19b2-abs2d-v1-legacy` 仅在显式确认后恢复旧假设;mask、standard、solid-angle、polarization、run policy 全字段贯通;monitor mode 单边界规范化。 -- 兼容性:不伪装为 v1.1.1;历史行为仍有显式、可审计入口。 -- 验证:CLI help/parse/migration、Config→rerun→Config 往返、签名一致性。 - -## CHG-005|AUD-006 - -- 范围:科研不确定度预算与 coverage 语义。 -- 修改前:样品端分量齐全时可能在缺标准端贡献的情况下返回 complete,并把证书 coverage factor 外推到组合 system budget。 -- 修改后:通过中心有限差分重跑实际估计器传播标准 T/MON/BG-MON/thickness/alpha;shared alpha 与 shared BG monitor 采用联合灵敏度;reference 与 system coverage 分离。 -- 语义边界:raw BG/dark covariance 未提供时,combined uncertainty 必须为 `partial`,system expanded uncertainty 不可用;未知量不当作 0。 -- 验证:缺项、共享变量、有限差分和 coverage 反例;独立科研 QA。 -- 回退风险:会重新产生虚假的 complete/expanded uncertainty。 - -## CHG-006|AUD-007/AUD-018/AUD-019 - -- 范围:CalibrationRecord schema v2、source identity、K history。 -- 修改前:record 未完整绑定标准/BG/dark/custom reference、模型与统计参数;缓存的校验布尔值可在源文件事后变化后继续被信任;历史 CSV 损坏或中断存在覆盖风险。 -- 修改后:source 使用相对路径和有序 SHA-256;自定义 reference 同时保存 raw/canonical q-I-u-U;SRM/Water 保存 model id/version/canonical hash;记录 alpha、BG rule、integration、robust estimator;读取和正式使用时复验源文件;K history 损坏即停并采用临时文件+fsync+原子替换。 -- 兼容性:v1 可读但明确 incomplete,不自动升级为可信 v2。 -- QA 补充:record reload/source tamper 闭环归入 AUD-007。 -- 验证:round-trip、移动、篡改、删除、损坏 CSV、no-clobber 和 v1 兼容测试。 - -## CHG-007|AUD-008 - -- 范围:SRM 3600 标准身份、厚度与平行性 QC。 -- 修改前:别名分叉且可接受非证书厚度;不平行 ratio 仍可产出 K。 -- 修改后:统一 `SRM3600`、`srm-3600`、`nist_srm3600`、`NIST SRM 3600` 等别名;隐式/显式厚度均约束为 0.1055 cm;不满足保守平行性门即拒绝。 -- 兼容性:正确证书参数不变;错误 0.1 cm 等输入提前失败。 -- 验证:别名参数化、厚度正反例、ratio QC 与 GUI/CLI 接线测试。 - -## CHG-008|AUD-009 - -- 范围:μ composition 输入与 GUI 接线。 -- 修改前:`Fe:95` 或其他非规范总和仅 warning 后继续。 -- 修改后:只接受总和约 1 的分数或约 100 的完整百分比,并精确归一化到 1;含糊/残缺尺度 fail-closed。 -- 兼容性:化学式和单元素合法路径保留。 -- 验证:分数、百分数、边界、异常值及 GUI 聚焦测试。 - -## CHG-009|AUD-010/AUD-014/AUD-015/AUD-016 - -- 范围:Tab2/Tab3 正式输出信任、auto-reference、dry-run、1D provenance。 -- 修改前:正式 gate 可依赖缓存 record 状态;auto-reference 先加载 fixed references;dry-run 与正式验证漂移;自产 1D 导出后回读丢 calibration context。 -- 修改后:正式 gate 每次重读 record、复验源文件与当前 context/K;auto-reference 按真实分支加载;dry-run 复用正式安全 gate;text/canSAS/NXcanSAS 写入 operator provenance,parser 回读并与注释 provenance 合并。 -- 单位边界:Q、2θ、χ 严格区分;`q_nm^-1` 转换为 Å⁻¹;2θ 必须有波长;跨轴不一致拒绝。 -- QA 补充:1D provenance round-trip 闭环归入 AUD-010。 -- 验证:writer→reader→formal gate、source tamper、axis/单位、auto refs 与 dry-run 聚焦测试。 -- 未闭环:可见 GUI/布局和人工 dark exposure 证据仍属于 AUD-015 open。 - -## CHG-010|AUD-017/AUD-020/AUD-021 - -- 范围:资源上限、稳定输入快照、重复读取与 resume provenance。 -- 修改前:workers 无上限、长 stem 可能越界;read→hash 存在 TOCTOU;resume 验证会覆盖首次创建信息。 -- 修改后:workers 限制 1..32;stem 最长 120 字符并附稳定哈希;关键输入采用读前/读后双哈希+stat 一致性门;buffer/reference 共享 parser 并单次预载;resume 另记 `last_resume_validation`。 -- QA 补充:creation provenance 不可变闭环归入 AUD-019/AUD-021。 -- 验证:极值 worker、长路径、输入变化、parser 调用计数和 resume metadata 测试。 -- 未闭环:真实桌面主线程体验和真实大批次/束线性能仍为 AUD-017/AUD-020 open。 - -## CHG-011|AUD-022 - -- 范围:Workbench launcher 与 Windows 启动脚本。 -- 修改前:cwd 可阴影正确模块;日志写 cwd 会在只读目录失败。 -- 修改后:优先解析已安装 package/source;日志写用户目录,失败时回退临时目录;启动脚本不再注入 cwd。 -- 验证:cwd shadow、不可写日志、入口解析和 subprocess smoke。 -- 未覆盖:自动化未获得可接管的可见窗口。 - -## CHG-012|AUD-023/AUD-024 - -- 范围:版本元数据、sdist manifest、生成目录 ignore 和用户/维护者文档。 -- 修改后:`pyproject.toml`、package、legacy launcher、CITATION、codemeta、根 `.zenodo.json` 统一 2.0.0;README、architecture、runbook、CHANGELOG 更新安全边界、六子命令、legacy 迁移和 provenance 合同;新增 `MANIFEST.in` 将 CHANGELOG/CITATION/codemeta/.zenodo/tests/conftest 纳入 sdist;新增 `.audit-work/sdist-*/` ignore。 -- 历史边界:submission/software paper 的既有 1.1.1 内容保留为历史快照并明确说明,不伪造回溯版本。 -- 验证:根 metadata 版本一致性、CLI version、文档核对;sdist 解包根包含 release metadata 与 `tests/conftest.py`,解包内 version tests `2 passed`。 -- 最终构建:`.audit-work/dist-release-final-v5/` 中 wheel 与 sdist 构建通过;wheel 隔离 smoke、sdist 解包与 version tests 通过。 - -## CHG-CI-001|AUD-024 - -- 触发:首轮 GitHub Actions `ci` run `29228584901` 最终 failure;除 ubuntu 3.11 外的 Ubuntu/macOS/Windows 矩阵均失败。 -- 根因 1:`tests/test_version_metadata.py` 在 Python 3.10 collection 直接依赖不存在的 `tomllib`。 -- 根因 2:wheel 内容测试运行 `pip wheel --no-build-isolation`,但 dev extra 未显式安装 `setuptools>=69` 与 `wheel`;旧测试捕获 stderr,失败日志不透明。 -- 修改:version test 改用 Python 3.10 兼容 regex;dev extra 加入 `setuptools>=69` 与 `wheel`;wheel test 失败信息同时显示 stdout/stderr。 -- 本地验证:focused `18 passed`;全量 `500 passed in 19.48s`;ruff、compileall、diff-check;final-v5 wheel/sdist build 与 smokes;sdist 解包 version tests `2 passed`。 -- 兼容性:测试不再要求 Python 3.11 `tomllib`;no-build-isolation 所需 build tools 成为显式 dev 合同;不改变运行时科研 API。 -- 审计截止:follow-up 尚未 commit/merge/push,第二轮远端 CI 尚未触发;不得把本地通过写成远端已绿。 - -## Follow-up 本地门禁状态 - -- 全量 pytest:`500 passed in 19.48s`;`.audit-work/pytest-release-final-v7.xml` -- ruff/compileall/diff-check:`全部通过` -- wheel/sdist build:通过;wheel 隔离 target import/CLI/launcher/entry points 通过;sdist 解包 metadata/conftest 完整,解包内 version tests `2 passed` -- wheel SHA-256:`7081EA14E70CE317DCAA125C7131EFA3C4E5BBB51F25A4CC582215A8857DD281` -- sdist SHA-256:`E569AA62C980963F06A96047A310D2E580AB6AEDD516F59870E8907ADBFE0595` -- 独立最终 QA:A/B/C/D/resume 与 release metadata/packaging PASS;CI follow-up focused `18 passed`;无新增 P0/P1/P2 -- 远端 CI:首轮 run `29228584901` failure 已诊断;第二轮 CI 待 follow-up commit/merge/push 后触发。 diff --git a/.audit-work/evidence.md b/.audit-work/evidence.md deleted file mode 100644 index 01016a4..0000000 --- a/.audit-work/evidence.md +++ /dev/null @@ -1,55 +0,0 @@ -# 证据账本 - -说明:EV-001~EV-019 是初始审计和首轮修复的历史证据;它们不能替代当前冻结树的最终门禁。 -新增 v2 闭环证据列于 EV-020 之后。EV-030~EV-033 是当前冻结树的最终门禁,均已通过。 - -| EV-ID | 目的 | 命令/方法 | 环境/工作目录 | 状态 | 关键输出 | 日志/产物 | 支持范围 | -|---|---|---|---|---|---|---|---| -| EV-001 | Git 基线 | branch/HEAD/status/merge-base/rev-list | PowerShell;仓库根 | pass | `main@6ba966c`;初始 origin 0/0;`docs/superpowers/` 为用户未跟踪资产 | `.audit-work/state.md` | 范围/保护 | -| EV-002 | 比较范围 | `git diff --name-status/--stat 0a3d680..HEAD` | 仓库根 | pass | 旧审计比较范围已登记 | 初版报告 | 风险地图 | -| EV-003 | 修改前 pytest | `py -3.11 -m pytest -q --tb=short --junitxml=.audit-work\pytest-baseline.xml` | Python 3.11;受限网络 | partial | 309 passed;隔离 wheel 依赖下载失败 | `.audit-work/pytest-baseline.xml` | 基线 | -| EV-004 | 基线构建替代验证 | 本地 build 依赖 + `--no-build-isolation` | 仓库根 | pass | 基线 wheel 可构建 | `.audit-work/wheel-baseline/` | 打包基线 | -| EV-005 | 静态/语法基线 | ruff;compileall | Python 3.11 | pass | 全通过 | 工具输出 | 基线 | -| EV-006 | 入口/核心 smoke | CLI/workbench version;83 项聚焦 pytest | Python 3.11 | pass | 基线 v1.1.1;83 passed | 初版报告 | 基线 | -| EV-007 | API/CLI 初始审计 | 调用链、直接旧调用、90 项聚焦测试 | 仓库根 | pass | 定位位置参数、CLI、rerun、文档缺陷 | 初版报告 | AUD-003/004/005/023/024 | -| EV-008 | 科研数值初始审计 | 130 项核心测试、人工小样例、规范核对 | Python 3.11 | pass | 主公式正确;定位 uncertainty/SRM/μ/record 缺口 | 初版报告 | AUD-006/007/008/009/018 | -| EV-009 | I/O 初始审计 | 106 项聚焦测试、fault injection、writer 调用映射 | Python 3.11 | pass | 定位 stale PONI、残缺 Cal2D、事务/移植性风险 | 初版报告 | AUD-001/002/011/012/019 | -| EV-010 | GUI/参数流初始审计 | 35 项 headless、withdrawn Tk geometry、可见 UI 尝试 | Windows/Tk | partial | 逻辑路径有证据;可见窗口不可接管,无截图 | 初版报告 | AUD-010/014-017/019/022 | -| EV-011 | safe PONI red | changed-source 反例 | Python 3.11 | expected-fail | 2 个反例修复前未拒绝 | 测试输出 | AUD-001 | -| EV-012 | safe PONI green | 同反例 + 正常 PONI writer | Python 3.11 | pass | 3 passed | 测试输出 | AUD-001 | -| EV-013 | Cal2D 首轮闭环 | package validator、目录 PONI、残包 | Python 3.11 | pass | focused tests passed | 测试输出 | AUD-002 | -| EV-014 | rerun/monitor red-green | 策略字段、monitor 规范化 | Python 3.11 | pass | 修复前失败,修复后通过 | 测试输出 | AUD-005/013 | -| EV-015 | 性能/状态初始测量 | 调用计数、匿名 TIFF/DAT、tracemalloc、TOCTOU 复现 | 本地临时数据 | pass | 识别资源、重复读取、输入变化风险;不代表束线性能 | 初版报告 | AUD-017/020/021 | -| EV-016 | 首轮 pre-QA 回归 | pytest/ruff/compile/wheel/CLI | Python 3.11 | historical-pass | 314 passed、1 deselected;后续代码已继续变化 | `.audit-work/pytest-final.xml` | CHG-001..003 历史证据 | -| EV-017 | 首轮独立 QA | 两名只读 reviewer + 聚焦测试 | Python 3.11 | discovery | 发现 mask/CLI 字段/空目录/账本缺口 | 初版报告 | 触发进一步修复 | -| EV-018 | 首轮 post-QA 回归 | pytest/ruff/compile/wheel | Python 3.11 | historical-pass | 首轮中间回归通过;旧 wheel 哈希仅属中间树 | `.audit-work/pytest-final-after-qa.xml` | 首轮修复历史证据 | -| EV-019 | 首轮 QA 复核 | 两名只读 reviewer | Python 3.11 | superseded | 当时未发现新增 P0-P2;后续深度 QA 又发现闭环缺口 | 初版报告 | 历史,不作为最终放行 | -| EV-020 | API、Cal2D、SRM 闭环 | calibration/Cal2D/BL19B2 聚焦测试,含旧位置调用与厚度反例 | Python 3.11 | pass | API 兼容;SRM aliases/0.1055 cm/QC;BL19B2 最新聚焦 119 passed | 测试输出 | AUD-002/003/008/011 | -| EV-021 | v2 CLI/legacy/重放 | CLI parse/help/migration;Config→rerun 往返 | Python 3.11 | pass | 安全 v2 与显式 v1 legacy 分离;字段可重放 | 测试输出 | AUD-004/005/013 | -| EV-022 | 不确定度语义 | finite-difference、shared variables、coverage/status 反例 | Python 3.11 | pass | 标准端灵敏度闭环;combined unknown covariance 保持 partial | 测试输出 | AUD-005/006 | -| EV-023 | record/context/history | v2 round-trip、relative source、hash tamper/delete、v1 incomplete、CSV fault | Python 3.11 | pass | source/model/参数身份绑定;损坏记录 fail-closed | 测试输出 | AUD-007/018/019 | -| EV-024 | μ 输入 | fractions/percent/ambiguous totals 参数化测试 | Python 3.11 | pass | 合法输入归一化至 1;含糊输入拒绝 | 测试输出 | AUD-009 | -| EV-025 | GUI/output/launcher 安全 | headless formal gates、axis、auto refs、transaction、stable snapshot、launcher smoke | Windows/Python 3.11 | pass-with-boundary | 代码/逻辑 gate 通过;无可见窗口证据 | 测试输出 | AUD-010/011/012/014-017/019-022 | -| EV-026 | version/docs/manifest | 版本一致性测试、CLI version、README/architecture/runbook/changelog、MANIFEST 核对 | 仓库根 | pass | package/CITATION/codemeta/根 `.zenodo.json` 统一 2.0.0;sdist 必需 metadata/conftest 纳入 manifest;历史 submission 快照保留 | `tests/test_version_metadata.py` 与文档/manifest diff | AUD-023/024 | -| EV-027 | 中间全量门禁 | `py -3.11 -m pytest -q --tb=short --junitxml=.audit-work\pytest-release-final-v2.xml`;ruff/compile/diff-check | Python 3.11 | historical-pass | 473 passed;发生于最后一轮 QA 修复之前,**不得作为最终数量** | `.audit-work/pytest-release-final-v2.xml` | 中间回归 | -| EV-028 | 中间构建与安装 smoke | `python -m build --no-isolation`;wheel install target;entry point/package content 检查 | Python 3.11 | historical-pass | 中间 wheel 可构建/安装;README 与最终 QA 修复后源码继续变化,旧哈希作废 | `.audit-work/dist-release-final/` | 中间打包证据 | -| EV-029 | 最后一轮独立 QA 发现项 | 科研、I/O/兼容、Git/release 三域只读审查与聚焦回归 | Python 3.11 | historical-discovery | 发现 rollback 竞态、record reload、自有 1D provenance 与冻结后重建要求;随后均完成闭环 | reviewer 输出 | AUD-007/010/011/019/021 | -| EV-030 | 最终冻结树 pytest | `py -3.11 -m pytest -q --tb=short --junitxml=.audit-work\pytest-release-final-v7.xml` | Python 3.11 | pass | `500 passed in 19.48s` | `.audit-work/pytest-release-final-v7.xml` | 最终代码放行 | -| EV-031 | 最终静态与版本门禁 | ruff;compileall;`git diff --check`;version smoke | 仓库根 | pass | ruff/compileall/diff-check 全部通过;版本 `2.0.0` | 工具输出 | 最终代码放行 | -| EV-032 | 最终冻结 artifact | build --no-isolation、SHA-256、wheel 隔离 target、sdist 解包验证 | Python 3.11 | pass | wheel `7081EA14E70CE317DCAA125C7131EFA3C4E5BBB51F25A4CC582215A8857DD281`;sdist `E569AA62C980963F06A96047A310D2E580AB6AEDD516F59870E8907ADBFE0595`;wheel import/CLI/launcher/entry points 通过;sdist 根含 metadata/conftest,解包内 version tests `2 passed` | `.audit-work/dist-release-final-v5/` | 最终 artifact 放行 | -| EV-033 | 最终独立 QA | A/B/C/D/resume 与 release metadata/packaging 复核最终 diff、门禁、artifact 与恢复路径 | 仓库根 | pass | 全部 PASS;无新增 P0/P1/P2;metadata 与 sdist manifest 闭环 | reviewer 输出 | 最终独立放行 | - -| EV-034 | 首轮远端 CI 失败 | GitHub Actions `ci` run `29228584901` 矩阵与失败日志 | GitHub-hosted Ubuntu/macOS/Windows;Python 3.10/3.11 | fail | 最终 failure;除 ubuntu 3.11 外均失败;Python 3.10 collection 无 `tomllib`;其他平台 no-isolation wheel test 缺显式 setuptools/wheel,旧 stderr 捕获降低可诊断性 | GitHub Actions run `29228584901` | AUD-024 / CI follow-up | -| EV-035 | CI follow-up 本地冻结 | focused pytest;全量 pytest;ruff/compileall/diff-check;final-v5 build/hash;wheel 与 sdist smokes | Python 3.11;本地仓库 | pass | focused 18 passed;`500 passed in 19.48s`;JUnit `.audit-work/pytest-release-final-v7.xml`;wheel/sdist 哈希与 smokes 通过;version test 为 Python 3.10 兼容 regex,dev extra 显式含 setuptools>=69/wheel,失败日志显示 stdout/stderr | `.audit-work/pytest-release-final-v7.xml`、`.audit-work/dist-release-final-v5/` | AUD-024 / follow-up 本地放行 | - -## 放行判据 - -EV-030~EV-033 证明初始 release 的本地门禁;EV-034 记录 release commit `cfad783` push 后首轮远端 CI failure;EV-035 证明 follow-up 本地闭环。审计截止到 follow-up 本地冻结点:follow-up 尚未 commit/merge/push,第二轮 CI 尚未触发。最终完成以第二轮远端 CI 成功为条件。 - -## 证据边界 - -- 可见 GUI:无可接管窗口,不得表述为视觉验收完成。 -- 真实束线数据:未提供、未运行,不得表述为束线端验证完成。 -- 不确定度:raw BG/dark covariance 未知时 combined 结果正确保持 `partial`。 -- 性能:本地匿名数据与调用计数只能证明风险/调用变化,不能替代真实大批次和束线存储基准。 -- 远端 CI:run `29228584901` 已确认 failure;follow-up 的第二轮远端复验尚未触发,不得表述为已绿。 \ No newline at end of file diff --git a/.audit-work/findings.md b/.audit-work/findings.md deleted file mode 100644 index 49f6f3c..0000000 --- a/.audit-work/findings.md +++ /dev/null @@ -1,38 +0,0 @@ -# 问题账本 - -结论:原始 24 项问题的优先级分布为 P0/P1/P2/P3 = **0/10/12/2**;当前状态为 -**verified=21、blocked=0、open=3**。follow-up 本地门禁通过;**最终完成以第二轮远端 CI 成功为条件**。 - -| ID | 优先级 | 分类 | 状态 | 路径/符号 | 触发条件与影响 | 修复闭环 | 证据 | 剩余边界 | -|---|---|---|---|---|---|---|---|---| -| AUD-001 | P1 | 几何/provenance | verified | BL19B2 safe PONI | 同路径几何改变可复用旧 PONI,导致错误续跑 | 比较渲染内容/哈希,不一致即 fail-closed | EV-011/012/018/020 | 无已知阻塞 | -| AUD-002 | P1 | Cal2D 完整性 | verified | Cal2D resume/writer | 五件套残缺或上下文错配可误报完成 | 五件套、shape、PONI、context、参数联合 gate;无效 PONI 写前拒绝 | EV-013/020/027 | 无已知阻塞 | -| AUD-003 | P1 | API 兼容 | verified | calibration/buffer 公共 API | 新参数插入旧位置顺序会破坏既有调用 | 恢复旧位置顺序;新增参数 keyword-only | EV-020 | 新参数须按关键字传入 | -| AUD-004 | P1 | CLI 兼容 | verified | BL19B2 CLI | 历史命令依赖隐式 monitor/μ,无法安全解释 | 发布版本升至 2.0.0;危险旧语义仅由显式 legacy 入口启用 | EV-021/026 | legacy 假设不会隐式恢复 | -| AUD-005 | P1 | 重放可复现性 | verified | rerun/CLI/config | mask、策略、standard、几何修正遗漏会改变 K/输出 | CLI、Config、rerun、signature/YAML 全字段贯通 | EV-021/022 | 无已知阻塞 | -| AUD-006 | P1 | 不确定度语义 | verified | uncertainty budget | 标准端或 covariance 缺失时曾可误报 complete | 标准 T/MON/BG-MON/d/alpha 灵敏度;reference/system coverage 分离;未知 covariance 保持 partial | EV-022/029 | combined uncertainty 在当前可证输入下正确为 partial | -| AUD-007 | P1 | calibration provenance | verified | calibration record/context v2 | K 来源、模型、源数据或源文件事后变化未被完整绑定 | v2 绑定标准/BG/dark/custom reference、模型/参数/哈希;正式 gate 每次重读 record 并复验源文件 | EV-023/029 | v1 仅兼容读取并标 incomplete | -| AUD-008 | P1 | SRM 科研约束 | verified | SRM 3600/K QC | 错厚度或不平行 ratio 会造成绝对标度偏差 | SRM 别名统一;锁定 0.1055 cm;证书约束与保守平行性门 | EV-020/025/029 | 更严格阈值可配置,不允许放宽证书边界 | -| AUD-009 | P1 | μ 输入 | verified | μ calculator/GUI | `Fe:95` 等比例尺度含糊可放大 μ | 仅接受总和约 1 或约 100,并精确归一化到 1;其他 fail-closed | EV-024 | 化学式/单元素兼容路径保留 | -| AUD-010 | P1 | GUI 信任/单位 | verified | Tab3/formal export | 默认 K、raw 缺 provenance、2θ 冒充 Q 或自产 1D 回读丢上下文 | 正式导出要求当前 record/context;Q/2θ/χ 严格处理;自产文本/canSAS/NXcanSAS provenance 可写回重读 | EV-025/029 | 可见 GUI 尚未完成视觉验收 | -| AUD-011 | P2 | 多文件输出安全 | verified | Cal2D transaction | 中途失败、目标冲突或并发替换会留下混合包/误删外部文件 | staging + create-if-absent no-overwrite 发布;后续冲突不删除已发布成员,残片由完整性门拒绝;overwrite 模式使用 backup rollback | EV-009/025/029 | 多文件包不具备全局单一原子提交;no-overwrite 冲突可保守保留残片,但不会覆盖或删除竞争者文件 | -| AUD-012 | P2 | RunPolicy | verified | Cal2D always-run | 目标存在时缺 package-level rerun 身份 | 整包分配统一 rerun ID,成员不会跨运行混合 | EV-025 | 无已知阻塞 | -| AUD-013 | P2 | monitor provenance | verified | monitor mode builders/run | 合法大小写/空白输入曾导致计算、签名和脚本漂移 | 单一边界规范化并贯通主入口、签名与重放 | EV-014/018 | 无已知阻塞 | -| AUD-014 | P2 | Tab2 控制流 | verified | auto-reference | auto 模式仍先强制加载 fixed BG/Dark | reference 分支前移,共享验证按实际模式执行 | EV-025 | 无已知阻塞 | -| AUD-015 | P2 | GUI/预检/布局 | open | Tab1/Tab2 | 可见布局、人工 dark exposure、真实点击路径缺证据 | dry-run 与正式路径已复用安全 gate;逻辑测试已覆盖 | EV-010/025 | **仍缺可见 GUI/布局截图与人工 dark exposure 主路径证据** | -| AUD-016 | P2 | Tab3 provenance/状态 | verified | parser/buffer/export | parser 分叉、状态不更新或 metadata 覆盖会削弱审计 | 共享 parser、buffer 单次预载、RunPolicy 与 operator provenance 合并 | EV-025/029 | 无已知阻塞 | -| AUD-017 | P2 | 资源/路径/GUI 线程 | open | workers/stem/Tk | 极端 worker、超长 stem 或主线程等待影响可靠性 | workers 限制 1..32;stem 采用 120 字符上限+稳定哈希 | EV-025 | **GUI 主线程体验仍需真实桌面验证** | -| AUD-018 | P2 | 模型校验 | verified | calibration/record | 重复 q、Inf、矛盾 U、零 MAD 等边界可产生伪稳健结果 | 核心不变量与有限值校验 fail-closed;统计参数写入 record | EV-023 | 无已知阻塞 | -| AUD-019 | P2 | 持久化/provenance | verified | record/K history/resume | 绝对路径、损坏 CSV、中断或 resume 覆盖创建信息 | 相对 source 路径;K history 损坏即停、临时文件+fsync+原子替换;resume 另记 last_resume_validation | EV-023/025/029 | 首次创建 provenance 保持不可变 | -| AUD-020 | P2 | 性能 | open | GUI/uncertainty/resume | 大批次/大 2D 可能放大 I/O、内存和等待 | 已减少重复 parser/缓冲读取并采用稳定 gate;未宣称真实吞吐提升 | EV-015/025 | **未做真实大批次或束线存储基准** | -| AUD-021 | P2 | 输入快照/TOCTOU | verified | read→hash | 采集或同步中文件变化会使哈希不代表已分析数据 | 读前/读后双哈希与 stat 一致性门;resume 验证不改写创建 provenance | EV-015/025/029 | 持续变化输入会明确失败 | -| AUD-022 | P2 | launcher 隔离 | verified | workbench launcher | cwd 阴影模块或不可写 cwd 可启动错版/无法记录日志 | package source 优先;用户/临时日志目录 fallback;启动脚本去 cwd 注入 | EV-025 | 可见启动窗口未被自动化接管 | -| AUD-023 | P3 | metadata/version | verified | release metadata/CLI/docs | `.zenodo.json` 等版本元数据或迁移说明可能陈旧 | pyproject/package/CITATION/codemeta/`.zenodo.json` 统一 2.0.0 并纳入 version test;六子命令与 legacy 边界同步 | EV-026/030/032 | submission 快照明确保留历史版本 | -| AUD-024 | P3 | 文档/打包/CI 质量 | verified | README/MANIFEST/dev tests | sdist 缺 metadata/conftest、Python 3.10 无 `tomllib` 或 no-isolation dev 环境缺 build tools 会使发布矩阵失败 | README 去重;`MANIFEST.in` 纳入 release files;忽略 sdist 工作目录;version test 使用 3.10 兼容 regex;dev extra 显式含 `setuptools>=69` 与 `wheel`;wheel test 失败显示 stdout/stderr | EV-026/032/034/035 | 第二轮远端 CI 待 follow-up push 后复验 | - -## 状态核对 - -- verified:AUD-001~AUD-014、AUD-016、AUD-018、AUD-019、AUD-021~AUD-024,共 21 项。 -- open:AUD-015、AUD-017、AUD-020,共 3 项。 -- blocked:无。 -- follow-up 本地冻结门禁:focused `18 passed`;全量 `500 passed in 19.48s`;ruff/compileall/diff-check 通过;版本 `2.0.0`;final-v5 wheel/sdist 与 smokes 通过并记录 SHA-256。首轮 `ci` run `29228584901` 为 failure;截至本地冻结点,follow-up 尚未 commit/merge/push,第二轮 CI 尚未触发。 \ No newline at end of file diff --git a/.audit-work/state.md b/.audit-work/state.md deleted file mode 100644 index 795adaa..0000000 --- a/.audit-work/state.md +++ /dev/null @@ -1,54 +0,0 @@ -# SASAbs v2 发布闭环状态 - -- 任务日期:2026-07-13(Asia/Shanghai) -- 仓库:`E:\desktop\SASAbs_saxs-absolute-calibration` -- 当前分支:`codex/ci-release-followup` -- 审计起点:`main@6ba966c715753f59a44009f1ee2ab07d15fc93f5` -- 已发布基础:`main == origin/main == cfad783` -- 当前 Gate:CI follow-up 本地冻结门禁通过;最终完成以第二轮远端 CI 成功为条件 -- 发布结论:**follow-up 本地放行;建议提交、快进合并并 push,最终完成以第二轮远端 CI 成功为条件。** -- P0/P1/P2/P3:0/10/12/2 -- verified/blocked/open:21/0/3 -- 最终 pytest:`500 passed in 19.48s`;JUnit `.audit-work/pytest-release-final-v7.xml` -- 最终静态门禁:`ruff / compileall / git diff --check` 全部通过 -- 最终版本 smoke:`2.0.0` -- 最终 wheel:`saxsabs-2.0.0-py3-none-any.whl`;SHA-256 `7081EA14E70CE317DCAA125C7131EFA3C4E5BBB51F25A4CC582215A8857DD281` -- 最终 sdist:`saxsabs-2.0.0.tar.gz`;SHA-256 `E569AA62C980963F06A96047A310D2E580AB6AEDD516F59870E8907ADBFE0595` -- artifact smoke:wheel 隔离 target import、CLI、launcher、entry points 全部通过;sdist 解包根包含 release metadata 与 `tests/conftest.py`,解包内 version tests `2 passed` -- 最终独立 QA:A/B/C/D/resume 全部 PASS;CI follow-up focused 18 passed;无新增 P0/P1/P2 -- 用户已有资产:`docs/superpowers/` 已保存在 `stash@{0}`,未纳入发布提交 -- Git/CI 截止状态:release commit `cfad783` 已 fast-forward 合并并 push,且 `main == origin/main == cfad783`;follow-up 尚未 commit/merge/push,第二轮 CI 尚未触发 - -## 已闭环范围 - -- AUD-001..AUD-014:verified。 -- AUD-016、AUD-018、AUD-019、AUD-021、AUD-022:verified。 -- AUD-023、AUD-024:verified。 -- Release metadata/packaging:根 `.zenodo.json` 已统一为 2.0.0 并纳入 version test;新增 `MANIFEST.in` 将 CHANGELOG/CITATION/codemeta/.zenodo/tests/conftest 纳入 sdist;新增 `.audit-work/sdist-*/` ignore。 -- CI follow-up:首轮 GitHub Actions `ci` run `29228584901` 最终 failure;除 ubuntu 3.11 外的 Ubuntu/macOS/Windows 矩阵均失败。根因是 Python 3.10 collection 无 `tomllib`,以及 `pip wheel --no-build-isolation` 路径的 dev 环境未显式安装 `setuptools>=69` 与 `wheel`;旧测试捕获 stderr 使日志不透明。follow-up 已改用 Python 3.10 兼容 regex、补齐 dev build 依赖,并在 wheel test 失败时显示 stdout/stderr。 -- P1 全部闭环:公共 API 位置参数兼容、v2 CLI/显式 legacy、完整重放参数、不确定度状态语义、calibration record v2、SRM 3600、μ 输入、Tab3 信任与单位边界。 -- P2 安全闭环:Cal2D 多文件事务与竞态保护、package-level rerun、auto-reference、Tab3 provenance、模型输入校验、K history 原子写、稳定读后哈希、launcher 隔离。 -- QA 补充闭环: - - Cal2D no-overwrite 后续冲突时不删除任何已发布路径;残片由完整性门拒绝,避免 TOCTOU 误删并发替换文件(AUD-011)。 - - 正式导出前重新读取并校验 calibration record 及其源文件;1D 自产物 provenance 可写回并重读(AUD-007/AUD-010)。 - - resume 验证不再覆盖首次创建 provenance,改记 `last_resume_validation`(AUD-019/AUD-021)。 - -## 仍开放但不阻塞发布的范围 - -- AUD-015:缺少可见 GUI 窗口、布局截图与人工 dark exposure 主路径证据;headless/逻辑 gate 已覆盖。 -- AUD-017:worker 上限和稳定短文件名已修复;GUI 主线程交互体验仍需真实桌面验证。 -- AUD-020:未用真实大批次/真实束线数据做性能与存储基准。 - -## 严格证据边界 - -- 可见 GUI 启动未出现可由自动化工具接管的窗口,因此不得宣称完成视觉、DPI、主题或真实点击验收。 -- 未使用真实 BL19B2/束线原始数据;现有证据来自自动测试、匿名最小样例和代码级审查。 -- raw BG/dark covariance 缺少可验证输入时,combined uncertainty 正确保持 `partial`;不把未知项当作 0,也不生成伪完整 system expanded uncertainty。 -- follow-up 本地冻结树 focused 18、全量 500、ruff、compileall、diff-check、版本 smoke、wheel/sdist 构建与 smokes 全部通过。 -- 审计截止语气:截至 follow-up 本地冻结点,follow-up 尚未 commit/merge/push,第二轮远端 CI 尚未触发;不得写成远端已绿。 - -## 下一步 - -1. 精确暂存并提交 `codex/ci-release-followup` 的 CI 兼容修复。 -2. 确认 `origin/main@cfad783` 未漂移,快进合并到 `main` 并 push。 -3. 观察第二轮远端 CI 到成功;成功后使用 `git branch -d` 安全删除已合并 follow-up 分支。 diff --git a/.gitattributes b/.gitattributes new file mode 100644 index 0000000..c3cfed3 --- /dev/null +++ b/.gitattributes @@ -0,0 +1,10 @@ +* text=auto eol=lf + +*.bat text eol=crlf + +*.h5 binary +*.npy binary +*.pdf binary +*.png binary +*.tif binary +*.tiff binary diff --git a/.github/ISSUE_TEMPLATE/bug_report.md b/.github/ISSUE_TEMPLATE/bug_report.md new file mode 100644 index 0000000..4cbc525 --- /dev/null +++ b/.github/ISSUE_TEMPLATE/bug_report.md @@ -0,0 +1,30 @@ +--- +name: Bug report +about: Report a reproducible problem +title: "bug: " +labels: bug +--- + +## What happened? + +Describe the observed behavior and why it is a problem. + +## Steps to reproduce + +1. +2. +3. + +## Expected behavior + +## Environment + +- saxsabs version or commit: +- Python version: +- Operating system: +- Optional dependencies used: + +## Minimal files or logs + +Attach anonymized, minimal inputs and relevant traceback/output when possible. +Do not attach private beamline data or credentials. diff --git a/.github/ISSUE_TEMPLATE/feature_request.md b/.github/ISSUE_TEMPLATE/feature_request.md new file mode 100644 index 0000000..8e5068b --- /dev/null +++ b/.github/ISSUE_TEMPLATE/feature_request.md @@ -0,0 +1,21 @@ +--- +name: Feature request +about: Propose a focused capability or workflow improvement +title: "feature: " +labels: enhancement +--- + +## Problem + +What workflow limitation does this address? + +## Proposed change + +Describe the smallest useful behavior, including required scientific inputs or +provenance where relevant. + +## Alternatives considered + +## Additional context + +Do not include private beamline data or credentials. diff --git a/.github/PULL_REQUEST_TEMPLATE.md b/.github/PULL_REQUEST_TEMPLATE.md new file mode 100644 index 0000000..8b55ec3 --- /dev/null +++ b/.github/PULL_REQUEST_TEMPLATE.md @@ -0,0 +1,15 @@ +## Summary + +Describe the focused workflow or behavior changed. + +## Validation + +- [ ] Added or updated focused tests where behavior changed +- [ ] Ran `pytest -q` +- [ ] Ran `ruff check src tests` +- [ ] Updated public CLI/API documentation when applicable + +## Scientific and compatibility notes + +State changed input semantics, provenance behavior, optional dependencies, or +known limitations. Do not include private beamline data or credentials. diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index ee9dd4c..7b7d393 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -2,9 +2,10 @@ name: ci on: push: - branches: [main] + branches: [main, joss-submission] pull_request: branches: [main] + workflow_dispatch: permissions: contents: read @@ -27,7 +28,7 @@ jobs: python -m pip install --upgrade pip pip install -e .[dev,gui,bl19b2,hdf5] - name: Ruff - run: ruff check src tests + run: ruff check SASAbs.py saxs_mpl_style.py src tests paper/*.py scripts/*.py - name: Pytest run: pytest -q --tb=short diff --git a/.github/workflows/release.yml b/.github/workflows/release.yml index 624bf10..7046c8a 100644 --- a/.github/workflows/release.yml +++ b/.github/workflows/release.yml @@ -19,9 +19,11 @@ jobs: - name: Install dependencies run: | python -m pip install --upgrade pip - pip install -e .[dev] + pip install -e .[dev,gui,bl19b2,hdf5] + - name: Verify tag and release metadata + run: python scripts/validate_release_metadata.py --tag "$GITHUB_REF_NAME" - name: Ruff - run: ruff check src tests + run: ruff check SASAbs.py saxs_mpl_style.py src tests paper/*.py scripts/*.py - name: Pytest run: pytest -q --tb=short @@ -36,8 +38,50 @@ jobs: - name: Build distributions run: | python -m pip install --upgrade pip - python -m pip install build + python -m pip install build twine python -m build + python -m twine check dist/* + - name: Clean-install wheel and run smoke checks + run: | + python -m venv .release-smoke + .release-smoke/bin/python -m pip install --upgrade pip + wheel="$(find dist -maxdepth 1 -name '*.whl' -print -quit)" + test -n "$wheel" + .release-smoke/bin/python -m pip install "${wheel}[gui,bl19b2,hdf5]" + smoke_dir="$(mktemp -d)" + cp examples/minimal_2d/run_minimal_2d_pipeline.py "$smoke_dir/" + cp examples/minimal_2d/synthetic_detector_image.csv "$smoke_dir/" + cp examples/minimal_2d/synthetic_geometry.json "$smoke_dir/" + smoke_python="$PWD/.release-smoke/bin/python" + smoke_cli="$PWD/.release-smoke/bin/saxsabs" + ( + cd "$smoke_dir" + "$smoke_cli" --version + "$smoke_python" - <<'PY' + from pathlib import Path + + import SASAbs + import saxs_mpl_style + import saxsabs + + presets = saxs_mpl_style.preset_choices() + if len(presets) != 5: + raise SystemExit( + f"expected five Workbench plot presets, got {presets!r}" + ) + if SASAbs.saxs_mpl_style is not saxs_mpl_style: + raise SystemExit( + "Workbench did not load the packaged plot-style module" + ) + for module in (SASAbs, saxs_mpl_style, saxsabs): + module_path = Path(module.__file__).resolve() + if "site-packages" not in module_path.parts: + raise SystemExit( + f"expected installed wheel module, got {module_path}" + ) + PY + "$smoke_python" run_minimal_2d_pipeline.py --output-dir outputs + ) - uses: actions/upload-artifact@v7 with: name: dist @@ -48,29 +92,16 @@ jobs: needs: build runs-on: ubuntu-latest steps: + - uses: actions/checkout@v6 + - uses: actions/setup-python@v6 + with: + python-version: "3.12" - uses: actions/download-artifact@v8 with: name: dist path: dist - - name: Build release notes header from CITATION.cff - run: | - python - <<'PY' - from pathlib import Path - import re - - citation = Path("CITATION.cff").read_text(encoding="utf-8") - match = re.search(r"^\s*value:\s*([0-9]+\.[0-9]+/\S+)\s*$", citation, re.MULTILINE) - if not match: - raise SystemExit("No DOI value found in CITATION.cff") - - doi_url = f"https://doi.org/{match.group(1)}" - Path("release-body.md").write_text( - "Archived release DOI: " - f"{doi_url}\n\n" - "Use this DOI when citing this specific software release.\n", - encoding="utf-8", - ) - PY + - name: Build release notes header from project metadata + run: python scripts/build_release_notes.py --output release-body.md - name: Publish GitHub Release uses: softprops/action-gh-release@v3 with: diff --git a/.gitignore b/.gitignore index c036003..87b1839 100644 --- a/.gitignore +++ b/.gitignore @@ -26,12 +26,13 @@ build/ dist/ logs/ examples/minimal_2d/outputs/ -# Local audit build/test artifacts (Markdown ledgers remain trackable) +# Local audit and build/test artifacts .audit-work/ +JOSS_AUDIT_AND_UPGRADE.md audit_outputs/ -.audit-work/**/*.xml -.audit-work/**/*.whl -.audit-work/pip-cache*/ -.audit-work/wheel-*/ -.audit-work/dist-*/ -.audit-work/sdist-*/ + +# Generated JOSS build products (review copies are delivered outside the repo) +paper/paper.pdf +paper/paper.tex +paper/paper.jats/ +paper/paper.*.log diff --git a/.zenodo.json b/.zenodo.json index 4dadc66..7e464a0 100644 --- a/.zenodo.json +++ b/.zenodo.json @@ -1,6 +1,6 @@ { - "title": "saxsabs: A Robust Workflow for Small-Angle X-ray Scattering Absolute Intensity Calibration", - "description": "SAXS absolute intensity calibration and robust 1D profile utilities", + "title": "saxsabs: Absolute-intensity calibration and provenance tracking for small-angle X-ray scattering", + "description": "Traceable SAXS absolute-intensity calibration and profile I/O", "upload_type": "software", "license": "BSD-3-Clause", "version": "2.0.0", @@ -23,7 +23,12 @@ "related_identifiers": [ { "relation": "isSupplementTo", - "identifier": "https://github.com/D-sudoasd/SASAbs_saxs-absolute-calibration" + "identifier": "https://github.com/D-sudoasd/SASAbs" + }, + { + "relation": "isVersionOf", + "identifier": "10.5281/zenodo.19687103", + "scheme": "doi" } ] } diff --git a/AGENTS.md b/AGENTS.md index ef4bdce..1122edb 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -2,7 +2,7 @@ ## Project Structure & Module Organization -`src/saxsabs/` contains the installable Python package. Core scientific logic lives in `src/saxsabs/core/`, file parsing and exporters in `src/saxsabs/io/`, and command-line wiring in `src/saxsabs/cli.py` plus `src/saxsabs/__main__.py`. Root-level `SASAbs.py`, `saxsabs_workbench.py`, `saxsabs_workbench.pyw`, and `Start_SAXSAbs_Workbench.bat` support the legacy/desktop workbench. Tests are in `tests/`. Example inputs and manual workflow checks are in `examples/`; architecture and reviewer documentation are in `docs/`; paper and submission assets are under `paper/` and `submission/`. +`src/saxsabs/` contains the installable Python package. Core scientific logic lives in `src/saxsabs/core/`, file parsing and exporters in `src/saxsabs/io/`, and command-line wiring in `src/saxsabs/cli.py` plus `src/saxsabs/__main__.py`. Root-level `SASAbs.py`, `saxsabs_workbench.py`, `saxsabs_workbench.pyw`, and `Start_SAXSAbs_Workbench.bat` support the legacy/desktop workbench. Tests are in `tests/`. Example inputs and manual workflow checks are in `examples/`; architecture and reviewer documentation are in `docs/`; JOSS paper assets are under `paper/`. ## Build, Test, and Development Commands diff --git a/CHANGELOG.md b/CHANGELOG.md index 061d0fd..9994aa8 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,8 +1,36 @@ # Changelog -## [Unreleased] +## [2.0.0] - Unreleased + +### Submission-candidate hardening (2026-08-12) + +- Require finite positive inherited `thickness_cm` and a non-empty + `thickness_source` before formal Tab 3 K-only scaling. +- Preserve inherited thickness value and source across text, canSAS, and + NXcanSAS operator provenance. +- Make Tab 3 write from the explicitly validated calibration context rather + than an implicit GUI cache. +- Harden the JOSS decision gate against e-mail/citation confusion, missing + README anchors or images, mismatched commit evidence, dirty worktrees, and + concept-DOI/version-DOI conflation. +- Make canSAS1d XML exports satisfy the official version 1.1 XSD by emitting + required `SASprocessnote` and entry-level `SASnote` elements; record bounded + external validation for canSAS1d and NXcanSAS without overstating current + NeXus-definition or third-party-consumer coverage. +- Build GitHub Release notes from the Concept DOI in `pyproject.toml` instead of + assuming that `CITATION.cff` already contains a release DOI. ### Added +- Added a JOSS manuscript, editable publication figures, an API reference, + maintainer-facing issue and pull-request templates, and a current JOSS + readiness checklist. +- Added a fail-closed JOSS readiness command that checks submission date, + required sections, author placeholders, corresponding-author metadata, + citations, versions, README targets, branch state, and build/cache pollution; + strict mode also requires an evidence-backed author confirmation record. +- Added a public-candidate gate that checks the canonical GitHub repository, + concept-DOI homepage, license, submitted branch/commit, visible README and + paper blobs, and successful CI run against the same confirmed revision. - Added a fingerprinted NIST 30 keV material-attenuation core with explicit wt% composition, ideal-mixture density, partial-uncertainty/porosity warnings, nominal Ti-24Nb-4Zr-8Sn, Ti-6Al-4V, and Zr-2.5Nb regression values, and JSON @@ -23,6 +51,13 @@ approval bound to a canonical configuration fingerprint. ### Changed +- Expanded the installation, quick-start, scientific-boundary, citation, and + contributor documentation; synchronized repository URLs and the project-level + Zenodo concept DOI across package, citation, CodeMeta, and Zenodo metadata. +- CI now runs on direct `joss-submission` pushes and supports manual dispatch; + lint coverage includes the submission-readiness script. +- Moved historical submission drafts, audit reports, and local audit ledgers out + of the version-controlled submission branch. - Workbench formal Tab 2 is now fixed-thickness only. Per-frame Beer-Lambert and both Tab 2/Tab 3 existence-only resume controls are UI-disabled, BLOCKED by Dry Check if forced, and rejected again at Run. K and μ are read-only in @@ -57,6 +92,16 @@ common close-safe loader. ### Fixed +- Deferred the TkAgg backend selection until the desktop Workbench is actually + constructed, allowing library imports after a headless Matplotlib backend has + already been selected on Linux and macOS. +- Made the path-traversal regression fixture write its unsafe CSV field + independently of the host path separator. +- Included the shared Workbench plotting-style module in wheel and source + distributions, and expanded source archives to include linked documentation, + examples, paper assets, and tests. +- Declared PyYAML as a development dependency so the packaged beamtime-template + test runs instead of being skipped in a clean source-distribution environment. - Prevented project-owned absolute 1D profiles from silently receiving K or thickness a second time. - Preserved explicitly named all-NaN combined-standard uncertainty columns when @@ -72,20 +117,18 @@ ### Known limitations - Formal multi-folder/per-sample fixed-thickness campaign ownership remains in the strict CLI/batch path; Workbench and strict-runner kernels are not unified. -- K-only requires an inherited thickness ledger marker but not yet its numeric - value/source. +- K-only requires an inherited finite positive thickness value and a non-empty + source; imported profiles that predate this provenance contract must be + upgraded or processed through K/d with an explicit thickness. - Workbench has no atomic campaign publication or content-signature resume; unsafe existence-only resume remains disabled. - Transactional calibrated-2D publication is scoped to the reusable `saxsabs.io.calibrated2d` package; the strict BL19B2 runner still lacks atomic whole-campaign publication. -- Repository hygiene remains open: roughly 79 MiB under `audit_outputs/` and - campaign-specific tests retain private `H:\...` path coupling and are not - portable default fixtures. -## [2.0.0] - 2026-07-13 +### Development changes incorporated on 2026-07-13 -### Added +#### Added - Explicit `bl19b2-abs2d-v1-legacy` migration command with acknowledgement flags for historical assumptions. - Calibration-record schema v2 binds and verifies portable standard, background, dark, reference, operator, integration, and robust-estimator sources used to derive K. - **Same-机时 (beamtime) automatic grouping** — new core `cluster_by_acquisition_time` + `AcquisitionGroup` in `saxsabs.core.session_grouper`. Tab2 now has a "检测机时分组 / Detect Groups" button that clusters files by timestamp (header preferred, mtime fallback) with a 90-minute default gap. Groups are exposed to the batch report path and can drive future per-run output subdirectories and smarter auto BG/Dark matching. @@ -93,7 +136,7 @@ - New public re-exports: `evaluate_preflight_gate`, `RunPolicy`, `build_reference_library`, `AcquisitionGroup`, `cluster_by_acquisition_time`, `extract_acquisition_timestamp`, etc. - Two new test modules: `test_session_grouper.py` and `test_reference_matching.py`. -### Changed +#### Changed - Promoted the BL19B2 workflow to an explicit safety-first CLI contract. Historical v1 assumptions require the dedicated legacy migration command and explicit acknowledgement. - Schema v1 calibration records remain readable but are provenance-incomplete and cannot authorize formal Tab 2/Tab 3 output. - Combined uncertainty remains explicitly partial while shared calibration-background raw-count and dark covariance are unquantified. @@ -101,7 +144,7 @@ - Tooltip system hardened (smarter screen-edge positioning, slightly better dark-mode contrast, safer error handling). - `parse_header` and `compute_norm_factor` in the Workbench now prefer the canonical core implementations (with legacy fallback for safety). -### Fixed +#### Fixed - Preserved public positional-call compatibility for calibration and buffer-subtraction APIs. - Made generated rerun commands preserve all scientific and execution options. - Hardened calibrated-2D and PONI resume checks against incomplete, corrupt, stale, or context-incompatible output packages; multi-file publication is transactional. @@ -112,9 +155,6 @@ - Separated reference-certificate coverage from the final system uncertainty coverage contract. - Minor robustness improvements in header timestamp parsing for grouping use-case. -### Release archive note -- `submission/softwarex/` is retained as the historical 1.1.1 submission snapshot; it is not the authoritative 2.0.0 runtime metadata. - ## [1.1.1] - 2026-04-22 ### Fixed @@ -205,7 +245,7 @@ - Replaced placeholder repository URL in CITATION.cff with real GitHub URL. - Upgraded `paper.bib` with full DOI-bearing citations for pyFAI, SasView, Dioptas, and added irena, DAWN, BioXTAS RAW, SRM 3600, Glatter & Kratky references. - Strengthened State of the field section with concrete tool comparison table. -- Enriched Research impact statement with real deployment context. +- Expanded the Research impact statement with author-reported workflow context. - Expanded AI usage disclosure per JOSS 2025 policy. ### Removed diff --git a/CITATION.cff b/CITATION.cff index 6c72192..701f8b8 100644 --- a/CITATION.cff +++ b/CITATION.cff @@ -1,20 +1,16 @@ cff-version: 1.2.0 -message: "If you use this software, please cite it using the metadata below." +message: "Cite an exact archived release when available. The README lists the project-level Zenodo concept DOI; version 2.0.0 is currently unreleased." title: "saxsabs" version: "2.0.0" -date-released: "2026-07-13" type: software -identifiers: - - type: doi - value: 10.5281/zenodo.19687104 -url: "https://doi.org/10.5281/zenodo.19687104" +url: "https://github.com/D-sudoasd/SASAbs" authors: - family-names: "Gong" given-names: "Delun" orcid: "https://orcid.org/0000-0001-7877-7707" affiliation: "Institute of Metal Research, Chinese Academy of Sciences" email: "dlgong@imr.ac.cn" -repository-code: "https://github.com/D-sudoasd/SASAbs_saxs-absolute-calibration" +repository-code: "https://github.com/D-sudoasd/SASAbs" license: "BSD-3-Clause" keywords: - SAXS diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 234456d..42090aa 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -1,35 +1,56 @@ # Contributing -Thanks for contributing to `saxsabs`. +Thanks for contributing to `saxsabs`. The canonical repository is +. + +## Get help or report a problem + +- Report reproducible bugs through the [issue tracker](https://github.com/D-sudoasd/SASAbs/issues). +- Use a feature request when proposing a new workflow or capability. +- For a question that is neither a defect nor a proposal, contact the project + maintainers through the repository before opening a broad pull request. + +Please do not include beamline-private data, credentials, or large generated +outputs in an issue or pull request. ## Development setup ```bash -pip install -e .[dev] +git clone https://github.com/D-sudoasd/SASAbs.git +cd SASAbs +python -m pip install -e ".[dev]" pytest -q ruff check src tests ``` -## Pull request checklist +Install `.[gui]`, `.[hdf5]`, `.[io]`, or `.[bl19b2]` only when the change needs +those optional workflows. -- Add or update tests for behavior changes. -- Keep GUI and core logic separated when possible. -- Update docs for new CLI/API behavior. -- Ensure CI is green. +## Pull requests -## Scope policy +Keep pull requests focused. For a behavior change: -This project prioritizes robust SAXS absolute-calibration workflows and reproducible data-processing paths. PRs that improve reliability, reproducibility, and traceability are preferred. +- add or update focused tests; +- keep reusable scientific logic in `src/saxsabs/` and GUI orchestration separate; +- update public CLI/API documentation when its behavior changes; +- run `pytest -q` and `ruff check src tests` locally; +- describe the workflow, validation performed, and any remaining limitations. -## Reporting bugs +Maintainers review pull requests for scientific input semantics, provenance, +reproducibility, and compatibility with supported optional dependencies. -Please open an issue at with: +## Release expectations -- A clear title and description of the problem. -- Steps to reproduce the issue. -- Expected vs. actual behaviour. -- Python version, OS, and relevant package versions (`pip list`). +Releases are created from version tags after the validation workflow succeeds. +Before tagging, replace the `Unreleased` changelog heading with the ISO release +date, add the same `date-released` to `CITATION.cff`, and set its message to +`Cite the version-specific archive record for this release.` The release +metadata validator rejects provisional or inconsistent values. +Before describing a release-specific DOI in project metadata or release notes, +ensure Zenodo has archived that release and assigned its DOI. The project-level +concept DOI remains suitable for general project citation. ## Code of Conduct -This project follows the [Contributor Covenant Code of Conduct](https://www.contributor-covenant.org/version/2/1/code_of_conduct/). By participating, you are expected to uphold this code. Please report unacceptable behaviour to the repository maintainer. +This project follows the [Contributor Covenant Code of Conduct](https://www.contributor-covenant.org/version/2/1/code_of_conduct/). +Report unacceptable behaviour to the repository maintainer. diff --git a/MANIFEST.in b/MANIFEST.in index 90aa1d8..c8d5e9d 100644 --- a/MANIFEST.in +++ b/MANIFEST.in @@ -1,5 +1,22 @@ include CHANGELOG.md include CITATION.cff +include CODE_OF_CONDUCT.md +include CONTRIBUTING.md +include LICENSE +include SUBMISSION_READINESS.md include codemeta.json include .zenodo.json -include tests/conftest.py +include SASAbs.py +include saxs_mpl_style.py +include saxsabs_workbench.py +include saxsabs_workbench.pyw +include Start_SAXSAbs_Workbench.bat +include .github/workflows/ci.yml +include .github/workflows/release.yml +recursive-include assets *.md *.png *.svg +recursive-include docs *.json *.md +recursive-include examples *.csv *.json *.md *.py *.yml +recursive-include paper *.bib *.md *.pdf *.png *.py *.svg +recursive-include scripts *.py +recursive-include src *.py +recursive-include tests *.py diff --git a/README.md b/README.md index 2550ada..9c40812 100644 --- a/README.md +++ b/README.md @@ -1,49 +1,217 @@ +# SASAbs + +

+ Traceable absolute-intensity calibration for small-angle X-ray scattering.
+ Python API · command-line workflows · bilingual desktop workbench +

+

- saxsabs: SAXS absolute intensity calibration workbench. + Continuous integration status + Zenodo concept DOI + BSD-3-Clause license + Python 3.10 or later

-# saxsabs +

+ Measured SAXS intensity is calibrated against a reference and exported with provenance. +

+ +`saxsabs` combines robust K-factor estimation, explicit intensity states, +reusable data writers, and provenance checks. The result and the processing +record remain reviewable together. + +

+ Quick start · + Choose a workflow · + API reference · + Architecture · + Submission readiness · + Citation +

+ +## Quick start + +Install from the repository and verify the headless CLI: + +```bash +git clone https://github.com/D-sudoasd/SASAbs.git +cd SASAbs +python -m pip install -e . + +saxsabs norm-factor --mode rate --exp 1.0 --mon 100000 --trans 0.8 +# 80000.0 +``` + +Launch the desktop application: + +```bash +python -m pip install -e ".[gui]" +saxsabs-workbench --lang en +``` + +The core package requires Python 3.10+, NumPy, pandas, and xraydb. The project +does not currently document a PyPI installation. -[![CI](https://github.com/D-sudoasd/SASAbs_saxs-absolute-calibration/actions/workflows/ci.yml/badge.svg)](https://github.com/D-sudoasd/SASAbs_saxs-absolute-calibration/actions/workflows/ci.yml) -[![DOI](https://zenodo.org/badge/DOI/10.5281/zenodo.19687104.svg)](https://doi.org/10.5281/zenodo.19687104) -[![License: BSD-3-Clause](https://img.shields.io/badge/License-BSD_3--Clause-blue.svg)](LICENSE) +
+Optional dependency groups -**SAXS absolute intensity calibration** · desktop app **SAXSAbs Workbench** · DOI https://doi.org/10.5281/zenodo.19687104 +```bash +python -m pip install -e ".[hdf5]" # NXcanSAS HDF5 +python -m pip install -e ".[io]" # FabIO detector-image I/O +python -m pip install -e ".[bl19b2]" # strict BL19B2 workflow +python -m pip install -e ".[dev]" # tests and Ruff +``` + +The Workbench uses Tk. Windows and macOS Python installers commonly include it. +On Linux, install the distribution's Tk package (often `python3-tk`) if +`python -m tkinter` cannot open a test window. API and CLI workflows do not need +a display server. + +
+ +## Choose a workflow

- SAXSAbs workflow figure. + Four SASAbs entry points converge on a traceable absolute-intensity result.

+| Route | Best for | Start here | +| --- | --- | --- | +| **CLI utilities** | normalization, header and 1D parsing, robust K estimation | `saxsabs --help` | +| **SAXSAbs Workbench** | interactive K calibration, batch processing, external-1D scaling | `saxsabs-workbench --lang en` | +| **Strict BL19B2 runner** | validated campaign inputs under current BL19B2 conventions | [batch runbook](docs/bl19b2_abs2d_batch_runbook.md) | +| **Python API** | reusable scientific calculations and file I/O | [API reference](docs/api.md) | + +The routes share numerical and I/O modules where implemented, but the Workbench +is not presented as equivalent to the stricter BL19B2 campaign runner. + +## What the software records + +- reference-derived calibration using NIST SRM 3600, water, or a supplied profile; +- explicit `raw_counts`, `relative`, `absolute_cm^-1`, and `ambiguous` states; +- transmission, thickness, monitor semantics, units, and applied corrections; +- partial uncertainty status without silently substituting zero for unknown terms; +- source identity where available, calibration context, and processing metadata; +- CSV/TSV, canSAS1d XML, and optional NXcanSAS HDF5 outputs. + +
+Open the detailed architecture diagram +

- SAXSAbs Workbench GUI. -   - K-factor calibration demo. + SASAbs software architecture from user interfaces through scientific and I/O modules to traceable outputs.

+
+ +## Workbench +

- 01 Workbench: K-factor, batch 2D, external 1D. + SAXSAbs Workbench in English showing calibration inputs, physical parameters, and the plotting area.

-| Tab | Input | Role | -|-----|-------|------| -| 1 K-Factor | **2D** | Calibrate K (GC / water) | -| 2 Batch | **2D** | 2D → absolute 1D (pyFAI integrate + absolute scale) | -| 3 External 1D | **1D** | Absolute scaling only when contracts met | -| 4 Help | — | Guide | +The desktop interface exposes K-factor calibration, 2D batch processing, +external-1D scaling, and built-in help. The image above was captured from the +English interface in the current source tree; it is interface documentation, +not experimental evidence. + +## Reproducible example -**Rule:** raw 2D → Tab 1+2 · only integrated 1D → Tab 3 when provenance OK. +The bundled example generates deterministic synthetic dark, background, +standard, and sample frames, then runs the package reduction APIs: -Also: multi-standard registry · robust K (median/MAD) · traceable μ · buffer subtraction · preflight READY/CAUTION/BLOCKED · canSAS / NXcanSAS · bilingual GUI. +```bash +python examples/minimal_2d/run_minimal_2d_pipeline.py +``` + +It writes inspectable CSV, TSV, and XML outputs, plus HDF5 when `h5py` is +installed. The acceptance summary requires `k_relative_error < 0.005` and +`sample_max_relative_error < 0.01`. See the +[example documentation](examples/minimal_2d/README.md) for construction details +and expected files.

- 02 Gates: fail-closed formal output. + Deterministic synthetic K-factor example showing retained and rejected ratios.

-Formal output is **fail-closed**: verified calibration records, unit-checked axes, fixed-thickness Tab 2 path, read-only K/μ from records, Dry Check fingerprints. Workbench is not yet a full substitute for the strict BL19B2 campaign runner — see `docs/`. +> This example checks software arithmetic and generated file content. It is not +> measured-beamline validation or independent third-party format validation. + +## Documentation + +- [API reference](docs/api.md) — public functions, inputs, outputs, and boundaries +- [Architecture](docs/architecture.md) — module responsibilities and interface limits +- [BL19B2 runbook](docs/bl19b2_abs2d_batch_runbook.md) — strict 2D campaign path +- [Manual verification](examples/manual-verification.md) — GUI and workflow checks +- [Reviewer FAQ](docs/reviewer-faq.md) — evidence, scope, and known limitations +- [Submission readiness](SUBMISSION_READINESS.md) — verified checks and remaining gates +- [Author confirmation form](docs/author-confirmation-form.md) — author-controlled facts required before submission +- [Changelog](CHANGELOG.md) — version history + +## Scope and limitations + +Absolute calibration depends on a suitable reference, detector geometry, +monitor semantics, transmission, thickness, and instrument-specific provenance. +The strict 2D workflow currently targets BL19B2 conventions. canSAS1d output is +checked against the official version 1.1 XSD. NXcanSAS output passes project-local +round-trip tests and punx 0.3.5 with its bundled v2018.5 definitions; current +NeXus definitions and third-party consumers have not yet been verified. + +## Development + +The [continuous-integration workflow](https://github.com/D-sudoasd/SASAbs/actions/workflows/ci.yml) +tests the configured Python and operating-system matrix. ```bash +python -m pip install -e ".[dev,gui,bl19b2,hdf5]" pytest -q -# examples/minimal_2d/ · examples/manual-verification.md +ruff check SASAbs.py saxs_mpl_style.py src tests paper/*.py scripts/*.py +``` + +Before submission, run the fail-closed local decision gate with Pandoc available: + +```bash +python scripts/check_submission_readiness.py \ + --as-of YYYY-MM-DD \ + --manual-confirmations path/to/submission-confirmations.json +``` + +Run the strict command from the exact branch and commit that will be submitted. +If Draft PR #1 is merged first, check out the resulting clean `main`, update +`submitted_branch` and `submitted_commit` in the confirmation JSON, and rerun +the gate. A PASS recorded before the merge does not cover the merge commit. + +After the strict local gate passes, verify the same commit, branch, visible +README and paper, repository identity, and successful CI run against GitHub: + +```bash +python scripts/check_public_candidate.py \ + --confirmations path/to/submission-confirmations.json ``` -Cite: https://doi.org/10.5281/zenodo.19687104 · BSD-3-Clause +When the paper remains outside the default branch, the command prints the exact +`@editorialbot set branch-where-paper-is ...` instruction required in the JOSS +pre-review issue. Remote mismatches or unavailable evidence fail closed. + +Until the author-controlled fields are complete, use +`--allow-author-placeholders --as-of 2026-08-26` only for mechanical preflight. +That override is not submission authorization. Start from the +[confirmation JSON template](docs/submission-confirmations.example.json) only +after completing the author confirmation form. + +Please use the [issue tracker](https://github.com/D-sudoasd/SASAbs/issues) for +reproducible problems and read [CONTRIBUTING.md](CONTRIBUTING.md) before opening +a pull request. Project participation follows the +[Code of Conduct](CODE_OF_CONDUCT.md). + +## Citation + +For the project as a whole, use the Zenodo concept DOI: + +> Gong, D. *SASAbs*. https://doi.org/10.5281/zenodo.19687103 + +Use a release-specific DOI only for the archived release it identifies. +Machine-readable metadata are available in [CITATION.cff](CITATION.cff). + +## License + +SASAbs is distributed under the [BSD-3-Clause license](LICENSE). diff --git a/SASAbs.py b/SASAbs.py index 56552e1..40e8c30 100644 --- a/SASAbs.py +++ b/SASAbs.py @@ -1,7 +1,7 @@ """SAXSAbs Workbench — GUI for SAXS absolute intensity calibration. Part of the saxsabs package. -Repository: https://github.com/D-sudoasd/SASAbs_saxs-absolute-calibration +Repository: https://github.com/D-sudoasd/SASAbs License: BSD-3-Clause """ import tkinter as tk @@ -16,7 +16,6 @@ import fabio import pyFAI import matplotlib -matplotlib.use("TkAgg") from matplotlib.backends.backend_tkagg import FigureCanvasTkAgg, NavigationToolbar2Tk from pathlib import Path import traceback @@ -1656,6 +1655,8 @@ def _validate_existing_calibrated2d_package( class SAXSAbsWorkbenchApp: def __init__(self, root, language="en"): + # Select the interactive backend only when a real Tk workbench is starting. + matplotlib.use("TkAgg") self.root = root self.language = (language or "en").strip().lower() if self.language not in SUPPORTED_LANGUAGES: @@ -3505,6 +3506,8 @@ def profile_operator_metadata( *, corrections_applied=None, k_factor=None, + thickness_cm=None, + thickness_source=None, ): """Return portable operator provenance for a project-owned 1-D export. @@ -3555,6 +3558,15 @@ def profile_operator_metadata( if not np.isfinite(k_value) or k_value <= 0: raise ValueError("1-D export provenance requires finite positive K") metadata["k_factor"] = format(k_value, ".17g") + if thickness_cm is not None: + thickness_value = float(thickness_cm) + if not np.isfinite(thickness_value) or thickness_value <= 0: + raise ValueError("1-D export provenance requires finite positive thickness_cm") + metadata["thickness_cm"] = format(thickness_value, ".17g") + source = str(thickness_source or "").strip() + if not source: + raise ValueError("1-D export provenance requires thickness_source") + metadata["thickness_source"] = source return metadata def save_profile_table( @@ -3570,6 +3582,8 @@ def save_profile_table( corrections_applied=None, combined_uncertainty=None, uncertainty_metadata=None, + thickness_cm=None, + thickness_source=None, ): # Origin-friendly text table: first row is column names, tab-separated. out_path = Path(out_path) @@ -3592,6 +3606,8 @@ def save_profile_table( profile_metadata = self.profile_operator_metadata( calibration_context, corrections_applied=corrections_applied, + thickness_cm=thickness_cm, + thickness_source=thickness_source, ) if uncertainty_metadata: profile_metadata.update(dict(uncertainty_metadata)) @@ -3649,6 +3665,8 @@ def save_profile_table( "intensity_unit", "corrections_applied", "do_not_repeat", + "thickness_cm", + "thickness_source", "buffer_source_name", "buffer_source_sha256", "buffer_alpha", @@ -3857,7 +3875,7 @@ def resolve_t2_polarization(self): # ========================================================================= def init_tab1_k_calc(self): p = self.tab1 - left_panel_holder = ttk.Frame(p, width=440) + left_panel_holder = ttk.Frame(p, width=460) left_panel_holder.pack(side="left", fill="y", padx=8, pady=8) left_panel_holder.pack_propagate(False) left_panel = self._make_scrollable_frame(left_panel_holder) @@ -3876,7 +3894,7 @@ def init_tab1_k_calc(self): f_files = ttk.LabelFrame(left_panel, text=self.tr("t1_files_title"), style="Group.TLabelframe") self._register_i18n_widget(f_files, "t1_files_title") f_files.pack(fill="x", pady=5) - self.add_hint(f_files, "hint_t1_files") + self.add_hint(f_files, "hint_t1_files", wraplength=390) self.t1_files = { "std": tk.StringVar(), "bg": self.global_vars["bg_path"], @@ -3959,7 +3977,7 @@ def init_tab1_k_calc(self): f_phys = ttk.LabelFrame(left_panel, text=self.tr("t1_phys_title"), style="Group.TLabelframe") self._register_i18n_widget(f_phys, "t1_phys_title") f_phys.pack(fill="x", pady=5) - self.add_hint(f_phys, "hint_t1_phys") + self.add_hint(f_phys, "hint_t1_phys", wraplength=390) f_phys_grid = ttk.Frame(f_phys) f_phys_grid.pack(fill="x") @@ -4007,21 +4025,24 @@ def init_tab1_k_calc(self): style="Hint.TLabel", ) lbl_norm_hint_t1.pack(side="left") + + correction_row = ttk.Frame(f_phys) + correction_row.pack(fill="x", pady=(2, 0)) cb_solid_t1 = ttk.Checkbutton( - norm_row, + correction_row, text=self.tr("cb_solid_angle"), variable=self.global_vars["apply_solid_angle"], ) - cb_solid_t1.pack(side="left", padx=(8, 0)) + cb_solid_t1.pack(side="left") self._register_i18n_widget(cb_solid_t1, "cb_solid_angle") cb_pol_t1 = ttk.Checkbutton( - norm_row, + correction_row, text="Polarization", variable=self.global_vars["polarization_enabled"], ) cb_pol_t1.pack(side="left", padx=(8, 2)) e_pol_t1 = ttk.Entry( - norm_row, + correction_row, textvariable=self.global_vars["polarization_factor"], width=6, ) @@ -5085,6 +5106,10 @@ def parse_external_operator_provenance(self, path): "correctionsapplied": "corrections_applied", "donotrepeat": "do_not_repeat", "intensityunit": "intensity_unit", + "thicknesscm": "thickness_cm", + "inheritedthicknesscm": "thickness_cm", + "thicknesssource": "thickness_source", + "inheritedthicknesssource": "thickness_source", "kfactor": "k_factor", "calibrationk": "k_factor", } @@ -5132,12 +5157,39 @@ def require_relative_external_profile_for_scaling( raise ValueError(f"{profile_name}: unknown correction mode: {mode}") if apply_buffer: corrections_to_apply.append("buffer") - return require_relative_input_for_absolute_scaling( + assessment = require_relative_input_for_absolute_scaling( profile, profile_name=profile_name, corrections_to_apply=corrections_to_apply, required_existing_corrections=required_existing, ) + if mode == "k_only": + provenance = profile.get("operator_provenance") + provenance = provenance if isinstance(provenance, dict) else {} + raw_thickness = provenance.get("thickness_cm", profile.get("thickness_cm")) + try: + inherited_thickness_cm = float(raw_thickness) + except (TypeError, ValueError) as exc: + raise ValueError( + f"{profile_name}: K-only scaling requires positive thickness_cm provenance" + ) from exc + if not np.isfinite(inherited_thickness_cm) or inherited_thickness_cm <= 0: + raise ValueError( + f"{profile_name}: K-only scaling requires positive thickness_cm provenance" + ) + thickness_source = str( + provenance.get("thickness_source", profile.get("thickness_source", "")) or "" + ).strip() + if not thickness_source or thickness_source.lower() in { + "none", + "null", + "unknown", + "n/a", + }: + raise ValueError( + f"{profile_name}: K-only scaling requires thickness_source provenance" + ) + return assessment def require_external_profile_operator_provenance( self, @@ -6654,7 +6706,13 @@ def run_external_1d_batch(self): else: if pipeline_mode == "scaled": scale_factor = scale_factor_global - thk_cm_used = fixed_thk_cm + if corr_mode == "k_only": + provenance = prof.get("operator_provenance") or {} + thk_cm_used = float(provenance["thickness_cm"]) + thickness_source = str(provenance["thickness_source"]) + else: + thk_cm_used = fixed_thk_cm + thickness_source = "tab3_fixed_thickness_input" i_abs = np.asarray(prof["i_rel"], dtype=np.float64) * scale_factor err_abs = np.asarray(prof["err_rel"], dtype=np.float64) * abs(scale_factor) else: @@ -6671,9 +6729,17 @@ def run_external_1d_batch(self): thk_cm_used = float(thk_use_mm) / 10.0 if not np.isfinite(thk_cm_used) or thk_cm_used <= 0: raise ValueError("厚度无效(固定厚度或metadata thk_mm)") + thickness_source = ( + "external_metadata_thk_mm" + if self.t3_use_meta_thk.get() + and sp["thk_mm_meta"] is not None + else "tab3_fixed_thickness_input" + ) scale_factor = k / thk_cm_used else: - thk_cm_used = np.nan + provenance = prof.get("operator_provenance") or {} + thk_cm_used = float(provenance["thickness_cm"]) + thickness_source = str(provenance["thickness_source"]) scale_factor = k s_i = np.asarray(prof["i_rel"], dtype=np.float64) @@ -6776,9 +6842,12 @@ def run_external_1d_batch(self): x_label, output_format=output_format, run_policy=run_policy, + calibration_context=active_calibration_context, corrections_applied=output_corrections, combined_uncertainty=combined_uncertainty, uncertainty_metadata=uncertainty_metadata, + thickness_cm=thk_cm_used, + thickness_source=thickness_source, ) if buffer_info["enabled"]: buffer_applied = True diff --git a/SUBMISSION_READINESS.md b/SUBMISSION_READINESS.md new file mode 100644 index 0000000..2e7e0cc --- /dev/null +++ b/SUBMISSION_READINESS.md @@ -0,0 +1,98 @@ +# Submission readiness snapshot + +Updated: 16 August 2026 (Asia/Shanghai) + +## Locally verified + +- Full source suite: PASS in a fully provisioned Python 3.13 environment; exact + count and duration are retained in the dated external validation record. +- Ruff: root modules, package, tests, paper scripts, and submission gate pass. +- README: 5 local images and all local links resolve; SVG/image audit passes. +- Minimal 2D example: K and sample maximum relative errors are + `0.001933697...`; CSV, TSV, XML, and HDF5 outputs are written. +- Fresh-copy distribution build: wheel and sdist PASS from a source tree + outside every Git checkout. The exact archive inventory is retained in the + dated external validation record; the sdist includes README assets, + workflows, docs, examples, tests, and paper sources. +- Installed-wheel smoke: CLI reports `saxsabs 2.0.0`; `SASAbs`, + `saxs_mpl_style`, and `saxsabs` import from the temporary environment; the + copied minimal example passes outside the checkout. A fresh Python 3.13 + environment resolves the declared GUI/HDF5 extras with no broken + requirements. +- Paper: 1100-word body by the documented Pandoc method; 16 references; current + Inara TeX and well-formed JATS resolve both figures. +- Review PDF: the official CI paper job produces a five-page draft whose pages, + bounds, figures, citations, and embedded fonts have been visually checked. + Exact run URL, byte size, and SHA-256 belong in the dated external validation + record because CI evidence must identify the submitted commit. The 1280 x 900 + GUI image is a real, reproducible full-window Workbench capture. +- Public CI gate: immediately before submission, the exact submitted HEAD must + have green push and Draft-PR runs for the complete matrix. Immutable commit + IDs and run URLs belong in the dated external validation record rather than + this tracked file, because editing the evidence here creates a new HEAD. +- External format checks: the deterministic example's canSAS1d XML validates + against the official 1.1 XSD with zero errors. Its NXcanSAS HDF5 output passes + punx 0.3.5 with bundled v2018.5 definitions (97 OK, 0 WARN, 0 ERROR); current + NeXus definitions and third-party consumers remain unverified. + +## Must be resolved before submission + +1. Recheck the public-history gate on or after 26 August 2026. JOSS requires + more than six months of public, iterative development; the repository was + created on 25 February 2026. +2. Add a verifiable research-use case. Synthetic validation and tests do not + establish demonstrated research impact. +3. Confirm the complete author list/order, corresponding author, current email, + affiliations, ORCIDs, and contribution roles. +4. Complete the AI disclosure with recoverable product/model/version details, + usage scope, and the author's final human-review assertion. +5. Supply truthful funding, sponsor-role, acknowledgement, and competing- + interest statements. +6. Before submission, verify that the public GitHub description, homepage + concept DOI, visible README, submitted branch, and green CI all identify the + exact candidate revision. +Run the strict decision gate with Pandoc available: + +```bash +python scripts/check_submission_readiness.py \ + --as-of YYYY-MM-DD \ + --manual-confirmations path/to/submission-confirmations.json +``` + +The gate must run on the exact branch and commit submitted to JOSS. If the +Draft PR is merged first, rerun it on the clean resulting `main` commit and +record `submitted_branch` and `submitted_commit` accordingly; pre-merge +evidence is not evidence for a later merge commit. + +After that local PASS, run: + +```bash +python scripts/check_public_candidate.py \ + --confirmations path/to/submission-confirmations.json +``` + +This second fail-closed gate uses the public GitHub API to verify the repository, +concept-DOI homepage, license, submitted branch and SHA, visible README and +paper blobs, and successful CI run for that exact SHA. It also reports the +editorialbot branch command when the paper is not on `main`. + +The current strict result is intentionally **FAIL** because the paper still has +four author-input placeholders, no confirmed corresponding author, and no paper +email. The mechanical preflight passes when +`--allow-author-placeholders --as-of 2026-08-26` is used; this override is not a +submission authorization. + +## Review-completion actions + +JOSS asks authors to make a tagged release and archive the reviewed revision +after successful review. Do not create `v2.0.0`, a GitHub Release, or a Zenodo +version archive until the candidate revision and remote workflow have been +confirmed. + +The paper source is dated 26 August 2026, the earliest conservative submission +date. If submission occurs later, update the YAML date to the actual submission +date; the strict readiness gate will reject a mismatch. + +The local review PDF still shows Inara pre-submission placeholders such as +`DOI: N/A`, 1970 dates, and volume/page fields. They are build metadata supplied +by the publication workflow, not text in `paper/paper.md`. diff --git a/assets/readme/README.md b/assets/readme/README.md new file mode 100644 index 0000000..41311d7 --- /dev/null +++ b/assets/readme/README.md @@ -0,0 +1,16 @@ +# README asset provenance + +These files support the GitHub repository homepage and are not experimental +evidence. + +| Asset | Source and regeneration | +| --- | --- | +| `hero.svg` | Hand-authored calibration sequence derived from the package interfaces and SAXS-profile motif; it contains no measured data. | +| `workflow.svg` | Hand-authored entry-point map for CLI utilities, Workbench, strict BL19B2, and Python API; the adjacent README table is authoritative. | +| `workbench.png` | Curated copy of `paper/fig_gui.png`, captured from the current source tree with `python paper/capture_gui_screenshot.py`. | +| `kfactor-demo.png` | Curated copy of `paper/fig_kfactor_demo.png`, generated deterministically with `python paper/generate_figures.py --demo`; it is synthetic and not beamline validation. | + +`workbench.png` is expected to match `paper/fig_gui.png`. `kfactor-demo.png` is +the approved README export of the generated synthetic figure and may be updated +from `paper/fig_kfactor_demo.png` after visual review. When either source is +regenerated, update its README copy in the same change and record both hashes. diff --git a/assets/readme/hero.svg b/assets/readme/hero.svg index 5a070f3..c5a6285 100644 --- a/assets/readme/hero.svg +++ b/assets/readme/hero.svg @@ -1,32 +1,42 @@ - saxsabs / SAXSAbs Workbench - SAXS absolute intensity calibration with multi-standard K-factor estimation and provenance gates. - - - - SAXS · ABSOLUTE INTENSITY · PROVENANCE - saxsabs - SAXSAbs Workbench — calibrate K, process batches, export absolute 1D. - - - K-FACTOR - BATCH 2D - EXTERNAL 1D - FAIL-CLOSED + SASAbs traceable absolute-intensity calibration + Detector signal and reference intensity pass through robust K-factor estimation to an absolute SAXS profile with units and provenance. + + + + SAXS ABSOLUTE-INTENSITY CALIBRATION + A calibration record + you can inspect. + Robust K estimation, explicit intensity states, + portable outputs, and reviewable provenance. + + CLI · PYTHON API · DESKTOP WORKBENCH - DOI 10.5281/zenodo.19687104 · BSD-3-Clause - - - - - TAB 1 · K calibration (2D) - TAB 2 · Batch 2D → abs 1D - TAB 3 · External 1D → abs - - I_abs = K · I_rel · corrections… - - READY · CAUTION · BLOCKED + + + + CALIBRATION PATH + + + + + + + detector signal + reference I(q) + + + + + robust K + + + absolute I(q) + + + + provenance + - D-sudoasd / SASAbs… diff --git a/assets/readme/kfactor-demo.png b/assets/readme/kfactor-demo.png new file mode 100644 index 0000000..f7d1b3d Binary files /dev/null and b/assets/readme/kfactor-demo.png differ diff --git a/assets/readme/section-01-workbench.svg b/assets/readme/section-01-workbench.svg deleted file mode 100644 index df94893..0000000 --- a/assets/readme/section-01-workbench.svg +++ /dev/null @@ -1,9 +0,0 @@ - - K-factor, batch 2D, external 1D - Section 01: WORKBENCH - - 01 · WORKBENCH - - K-factor, batch 2D, external 1D - 01 - diff --git a/assets/readme/section-02-gates.svg b/assets/readme/section-02-gates.svg deleted file mode 100644 index c52e576..0000000 --- a/assets/readme/section-02-gates.svg +++ /dev/null @@ -1,9 +0,0 @@ - - Fail-closed formal output - Section 02: GATES - - 02 · GATES - - Fail-closed formal output - 02 - diff --git a/assets/readme/workbench.png b/assets/readme/workbench.png new file mode 100644 index 0000000..33d69e2 Binary files /dev/null and b/assets/readme/workbench.png differ diff --git a/assets/readme/workflow.svg b/assets/readme/workflow.svg new file mode 100644 index 0000000..b0aee47 --- /dev/null +++ b/assets/readme/workflow.svg @@ -0,0 +1,39 @@ + + SASAbs entry points and shared processing core + CLI utilities, desktop Workbench, strict BL19B2 runner, and Python API connect to shared scientific and input-output modules before producing reviewable absolute-intensity outputs. + + + ENTRY POINTS + + + + CLI utilities + Workbench + BL19B2 runner + Python API + + + focused commands + interactive routes + strict campaign path + reusable functions + + + + + + + + + + + + shared scientific core · typed I/O · provenance checks + + + + + reviewable + absolute output + + diff --git a/audit-report-20260714-final.md b/audit-report-20260714-final.md deleted file mode 100644 index ca13606..0000000 --- a/audit-report-20260714-final.md +++ /dev/null @@ -1,266 +0,0 @@ -# SASAbs 全量项目审计报告 - -审计日期:2026-07-14(Asia/Shanghai) -审计模式:`full` -审计目标:当前 dirty snapshot 的全量改动收口、缺陷修复与洁净交付 -基线:HEAD `1c82092ab4bc4734c17366babd6ce52fab1b43f8`;基线时 `main...origin/main` 对齐 -执行分支:`codex/sasabs-release-hardening-20260714`;最终目标分支:`main` - -## 1. 最终结论 - -当前 closeout 已完成定向修复、私有审计资产迁移和发布边界整理;最终门禁以本报告第 13 节为准。 - -- **没有确认 P0 级数值公式错误**;公式、单位体系和既有公共 API 未改变。 -- **F-001 已关闭**:`audit_outputs/` 与私有 campaign 测试已按样品批次迁移到外部归档,并由 SHA-256 清单核验;仓库新增 local-only ignore 规则。 -- **F-002 已关闭**:Fabio detector-image 读取和 reference matching 均在资源释放前复制数据,并以 `try/finally` 确定性关闭句柄。 -- **F-003 已关闭**:非法显式 `intensity_state` 现在保留无效元数据证据并返回 `AMBIGUOUS`。 -- **F-004 已关闭**:writer 入口拒绝非有限 `q/I`,同时保留不确定度缺失值的既有处理语义。 -- F-005 保留为 Workbench/strict runner 的未验证 GUI/campaign 架构残余,不在本轮进行无边界重构。 - -因此,本轮限定范围的代码修复、发布卫生和可复现验证已闭合;报告不宣称 F-005 已完成。 - -## 2. 审计边界与保护原则 - -本轮遵循以下边界: - -- 目标是当前 dirty snapshot,不把历史 `audit-report.md` 或历史 `.audit-work/` 的结论当作当前验证。 -- 所有既有 tracked/untracked 改动均视为用户资产;本轮仅移除已完成外部归档的私有副本和临时占位物,未 reset 或回退代码。 -- 未升级依赖;代码、测试、文档和最终报告改动均纳入当前功能分支。 -- 临时审计工作区和 `audit-report-20260714.md` 占位报告不纳入提交;已有 `audit-report.md` 保持不变。 -- 私有审计资产与 campaign 测试已按样品批次迁移到外部目录,并逐项完成 SHA-256 核验;仓库不保留其副本。 - -当前复核证据以仓库中的测试、静态检查、构建产物扫描和本报告第 13 节为准;修复前临时审计证据已在外部归档核验后清理。 - -外部样品目录中的 `SASAbs_audit_archive_20260714` 保留批次归档和独立哈希清单。 - -## 3. 项目结构与关键数据流 - -当前仓库的主要层次清晰: - -```text -src/saxsabs/core/ 可复用科学计算、单位/状态/不确定性逻辑 -src/saxsabs/io/ canSAS/NXcanSAS 解析与导出 -src/saxsabs/workflows/ BL19B2 2D 绝对标定与严格 1D 积分 -src/saxsabs/cli.py CLI 入口 -SASAbs.py legacy/desktop Workbench -tests/ pytest 自动化测试 -examples/ 最小 2D 与手工验收材料 -docs/ 架构、runbook、审计边界 -paper/submission/ 论文与投稿资产 -``` - -严格 BL19B2 1D 路径的科学链条是:2D package manifest → EDF/metadata/PONI/mask 校验 → 读取已是 `cm^-1` 的 EDF → 仅执行一次 mask 与 solid-angle correction → CSR 积分到 5500 个 `q_A^-1` 点 → profile/sidecar/checksum/completion。代码明确阻止 dark、background、monitor、transmission、thickness、K 和 polarization 的重复执行。 - -桌面 Workbench 具有独立的 Tab 2/Tab 3 编排、preflight、强制 `CalibrationContext`、intensity-state/ledger 和 disabled legacy resume,但它不是严格 campaign runner 的直接调用者。因此严格 CLI 的通过不等价于 GUI 全路径已经通过。 - -## 4. G1/G6 基线结果 - -| 检查 | 结果 | 证据类型 | -|---|---:|---| -| pytest 基线(修复前) | 726 passed in 41.70 s | 直接运行 | -| pytest 最终回归(修复后) | 687 passed in 24.98 s | 直接运行 | -| Ruff | All checks passed | 直接运行 | -| compileall | exit 0 | 直接运行 | -| `git diff --check` | exit 0;仅换行符提示 | 直接运行 | -| CLI version | `saxsabs 2.0.0` | 直接运行 | -| Workbench version | `saxsabs_workbench.py 2.0.0` | 直接运行 | - -慢测试集中在 wheel/Workbench 启动、最小 2D 示例、真实 float32 导出重开和若干 resume/mutation 检查;没有从该结果推断出真实 2000 帧 GUI 性能。 - -## 5. 风险总表 - -| 编号 | 优先级 | 状态 | 结论 | -|---|---|---|---| -| F-001 | P1 | fixed | 私有审计资产已迁移; `audit_outputs/` 与 `.audit-work/` 为 local-only ignore | -| F-002 | P1 | fixed | Fabio 读取和 reference matching 已确定性 close,并有回归测试 | -| F-003 | P2 | fixed | 非法显式 intensity state fail-closed 为 `AMBIGUOUS` 并保留证据 | -| F-004 | P1 | fixed | writer 入口拒绝 q/I 非有限值;error 缺失语义保留 | -| F-005 | P2 | open/residual | Workbench/strict runner 的 campaign owner、原子发布、真实 UI 证据未闭合 | - -## 6. 已确认发现(修复前快照及当前状态映射) - -以下各 F 条目的详细描述保留修复前审计证据;当前状态以第 5 节风险总表和第 13 节 closeout 为准。 - -### F-001 — P1:仓库发布卫生与私有数据边界 - -修复前统计结果:`audit_outputs/` 共 213 个文件、83,327,789 bytes,即 79.47 MiB;其中 204 个是未跟踪且未被 ignore 的文件,34 个文件含束线私有路径前缀。`git ls-files audit_outputs` 为 0;当时的 `.gitignore` 没有 `audit_outputs/` 规则。 - -多个 campaign 测试直接 `from audit_outputs import ...`,campaign runner 还含 H 盘默认输入/输出路径。结果是: - -1. 当前工作区的 green test 结果不能由 fresh clone 重现。 -2. 若未来误 add,这些私有路径与大文件可能进入仓库/发布物。 -3. 即使不 add,当前快照也不满足“可交付 release candidate”的 provenance 边界。 - -本轮没有删除、移动、匿名化或重写这些用户资产;需要用户明确选择 local-only fixture、匿名 synthetic fixture、受控模块迁移或仅发布前 gate。这个发现保持 blocked。 - -### F-002 — P1:Fabio 句柄关闭不完整 - -`src/saxsabs/workflows/bl19b2_abs2d.py:1746` 的 `read_detector_image()` 使用 `fabio.open(...).data` 后直接返回;没有在 `finally` 中保存并关闭 image object。`src/saxsabs/core/reference_matching.py:81` 起的 `build_reference_library()` 对每个候选图打开后读取 metadata,也没有对 image object 做确定性 close。 - -严格 1D `_load_mask()`、`_load_validate_edf()` 和严格 2D resume verifier 有正确的 `try/finally/close`,所以这是部分路径缺陷,不应夸大为“所有 Fabio 路径都泄漏”。但在批量读样品/参考图时,它仍可能造成 Windows 文件锁或资源累积。 - -最小修复方向:先把 `image_file.data` 复制为独立 C-contiguous 数组,再在 `finally` close;reference matching 采用相同模式;添加成功与 data decode 异常两条回归测试。本轮没有声称已修复。 - -### F-003 — P2:非法显式强度状态被静默忽略 - -`src/saxsabs/core/intensity_state.py:131` 的 `_state_from_metadata()` 对未知 token 返回 `None`;`assess_intensity_state()` 之后仍可使用单位/列名推断状态。 - -直接 probe: - -```text -intensity_state = "not-a-state" -intensity_unit = "1/cm" -i_col = "i" -=> ABSOLUTE_CM_INV, evidence=("unit:1/cm",) -``` - -当前强度缩放调用者仍会拒绝 absolute 输入,因此本 probe 没有直接证明已经产生错误绝对强度;但 provenance 元数据本身已经损坏而没有被报告,违反严格状态机应有的 fail-closed 预期。建议存在非空显式值但无法识别时直接得到 `AMBIGUOUS`,并保留 `invalid_metadata:intensity_state` 证据。 - -### F-004 — P1:public writer 接受非有限 q/I - -`src/saxsabs/io/writers.py:70` 的 `_prepare_profile_arrays()` 仅检查 shape,不检查 q/I 是否 finite,也不检查 q 是否单调。直接 probe 将: - -```text -q = [0.1, NaN] -I = [1.0, Inf] -err = [0.1, NaN] -``` - -写成: - -```xml -nan -inf -``` - -同样的非有限 q/I 也进入 NXcanSAS HDF5。`err=NaN` 被跳过是当前代码的显式行为,但 q/I 不应在公开科学输出边界静默穿透。建议对 q/I 做 finite 校验,对 error 制定明确的 finite/non-negative policy,并增加 XML/HDF5 两种格式的 fail-closed 测试。 - -### F-005 — P2:Workbench 与严格 runner 的剩余边界 - -CLI 明确调用 `run_bl19b2_abs2d`;Workbench 的 `run_batch`、`run_external_1d_batch` 是独立方法。当前代码和文档已做了很多安全强化:legacy exists-only resume 被禁用、preflight/CalibrationContext/ledger gate 已存在、strict 1D 有 per-artifact checksum 与 completion。 - -但本轮没有证据证明 Workbench 已经拥有: - -- multi-folder campaign owner 与逐目录 accepted-frame/T_rep/MAD/P5-P95/d_fixed 发布表; -- whole-campaign staging + atomic publish; -- 与 strict runner 一致的 content-signature resume; -- 真实 2000 帧交互性能、取消、DPI/键盘/中文英文/深浅主题矩阵。 - -因此不能把 strict CLI 通过外推成 GUI 全路径通过。 - -## 7. 科学数据与物理语义审计 - -### 7.1 事实、解释、推断分层 - -| 层级 | 本轮结论 | -|---|---| -| 实验/输入事实 | 本轮没有读取私有束线数据;所有实际数据结果仅来自仓库 fixture、synthetic test 和代码 probe | -| 数据处理事实 | strict 1D 仅接受已是 `cm^-1` 的 EDF,并强制 5500 `q_A^-1`、CSR、一次 solid angle、无 polarization | -| 数据解释 | correction ledger、intensity state、K/thickness/buffer gate 用于阻止重复校正;`do_not_repeat` 被当作 guard 而非物理证明 | -| 机制/外推 | 名义材料密度是理想混合模型,不是实测 bulk density;孔隙率会偏置线性 μ 和 derived thickness | -| 作者/项目主张 | “Workbench 已等价于严格 runner”“真实 GUI 已通过”本轮没有充分证据支持 | - -### 7.2 NIST 30 keV 与三种材料 - -官方 NIST Table 1 的密度与 Table 3 的 30 keV 质量衰减系数,与代码快照一致: - -- Al:2.699 g/cm³,1.128 cm²/g -- Ti:4.540 g/cm³,4.972 cm²/g -- V:6.110 g/cm³,5.564 cm²/g -- Zr:6.506 g/cm³,24.85 cm²/g -- Nb:8.570 g/cm³,26.66 cm²/g -- Sn:7.310 g/cm³,41.21 cm²/g - -参考:[NIST Table 1](https://physics.nist.gov/PhysRefData/XrayMassCoef/tab1.html)、[NIST Table 3](https://physics.nist.gov/PhysRefData/XrayMassCoef/tab3.html)。NIST 对混合物采用按质量分数加和,并提醒元素密度可能是 nominal;这与项目将密度标注为 ideal/model-derived、将不确定性标为 partial 的实现边界一致。 - -独立本地计算复现: - -```text -Ti-24Nb-4Zr-8Sn: mu = 74.55035538810252 cm^-1 -sum(w_i * (mu/rho)_i) * rho_ideal = 74.55035538810252 cm^-1 -T_median = 0.5 -> d_fixed = 0.009297704577684208 cm -``` - -### 7.3 不确定性与边界 - -本轮代码阅读和已有测试支持以下结论: - -- raw-count 统计项包含 sample/background/dark 的共享 dark 系数传播; -- 缺失的 reference、monitor、transmission、alpha、covariance 等组件保持 `None`/`partial`,没有被偷偷当成 0; -- expanded uncertainty 只有在完整未知项、有限数组和 coverage factor 都具备时才标记 available; -- masked detector pixels 用零占位但在 summary 中排除,并要求后续分析使用分布的 mask; -- SRM 厚度、NIST 快照、composition、PONI/mask/signature 等均写入 provenance 或 processing signature。 - -仍需注意:共享 blank/dark covariance 没有被量化时,完整系统不确定性不能被宣传为“完整”;理想混合密度也不能替代样品 bulk density 认证。 - -## 8. UI/UX 审计 - -静态代码证据显示: - -- root/Toplevel 已使用 screen-aware geometry; -- Tab 2/Tab 3 legacy exists-only resume 控件被禁用,并在 run gate 中拒绝; -- preflight fingerprint 会绑定配置、文件身份和 calibration context; -- K、μ、material provenance 与 correction ledger 具备只读/重算/失效机制。 - -本轮没有打开实际 GUI 窗口,也没有截图或自动化视觉/键盘验收。因此以下结论保持未验证: - -- 1024×700、100/125/150/200% DPI; -- 中英文、深浅主题、焦点顺序、键盘操作、非颜色状态提示; -- 2000 帧预检/筛选/取消时的响应性; -- Workbench 输出目录 owner、整批失败恢复和 completion 展示。 - -## 9. 性能审计 - -已验证:修复前完整 pytest 726 项在 35.29 s 内通过;迁移私有 campaign 测试并完成修复后,完整 pytest 687 项在 24.98 s 内通过。没有发现静态上明显的 O(N²) 新增路径或测试级资源爆炸。 - -未验证:真实束线 2000 帧、内存峰值、磁盘吞吐、Fabio/pyFAI 多线程、GUI 主线程阻塞、取消响应。不能用单机 pytest 时间替代这些验收。 - -## 10. 修复与变更记录 - -本轮纳入了用户要求的全部有效代码、测试、文档和可复现验证改动,并完成以下限定范围修复: - -- F-002:detector-image 读取和 reference matching 使用 `try/finally` 关闭 Fabio handle,并复制数组后再释放资源;增加成功、异常和资源关闭回归测试。 -- F-003:非法显式 `intensity_state` 返回 `AMBIGUOUS`,并保留 `invalid_metadata:intensity_state` 证据。 -- F-004:writer 拒绝包含 NaN/Inf 的 q/I 数组;不确定度缺失值仍按既有规则处理。 -- F-001:新增 `audit_outputs/` 与 `.audit-work/` 的 local-only ignore 规则;私有审计资产和 campaign 测试按批次外部归档并核验后移除仓库副本。 -- 同时保留本轮已有的科学工作流、文档和测试改动;未改变公式、单位体系、依赖版本或已有公共 API。 - -已有 `audit-report.md` 未覆盖;临时审计工作区和本轮占位报告未提交。 - -## 11. 发布门禁状态 - -1. [x] `audit_outputs/` 已完成 local-only ignore 与按批次外部归档,fresh clone 不依赖私有副本。 -2. [x] F-002 已修复,并有成功、异常和资源关闭回归测试。 -3. [x] F-003 已修复,并固定非法显式状态的 provenance 证据。 -4. [x] F-004 已修复,并固定 q/I 非有限值的 fail-closed 行为。 -5. [ ] F-005 的真实 Workbench/strict-runner campaign、atomic publish、resume、DPI/键盘/取消和 2000 帧验收仍待后续范围明确后完成。 -6. [x] pytest、Ruff、compileall、CLI/Workbench 版本、隔离构建及私有路径扫描均已完成;最终 Git 门禁在交付步骤复核。 - -## 12. 审计自检 - -- [x] 先基线,再风险图、复现、回归、最终报告。 -- [x] 区分直接运行、静态代码证据、独立数值 probe、外部权威来源和未验证项。 -- [x] 没有把测试绿色外推为真实数据或 GUI 已验证。 -- [x] 外部归档前记录文件数量、大小和 SHA-256,复制后逐项核验,确认后才移除仓库副本。 -- [x] 未覆盖已有 `audit-report.md`;临时审计 state 和占位报告未提交。 -- [x] F-005 仍明确标记为 residual,未把 GUI/strict-runner 未验证项外推为已完成。 -- [x] 当前报告状态与 F-001–F-004 的修复结果一致。 - -## 13. Closeout verification - -本节记录本轮修复、归档和发布门禁的最终结果: - -| 检查 | 结果 | -|---|---| -| 定向回归 | 202 passed in 4.87 s | -| 完整 pytest | 687 passed in 24.98 s | -| Ruff | All checks passed | -| compileall | exit 0 | -| CLI / Workbench | `saxsabs 2.0.0` / `saxsabs_workbench.py 2.0.0` | -| 隔离构建 | sdist 与 wheel 均成功;包内无私有路径、审计目录或 campaign 测试名 | -| 2025A1750 外部归档 | 61 files,253,726 bytes;manifest SHA-256 逐项核验通过 | -| 2026A1756 外部归档 | 151 files,82,745,256 bytes;manifest SHA-256 逐项核验通过 | -| 仓库私有边界 | 私有 audit_outputs 副本和 campaign 测试已移除;未发现私有 H 盘路径残留 | -| Git 交付 | 提交、快进合并、push、分支清理和最终工作区状态以交付步骤的实时门禁为准 | - -本轮不改变公式、单位体系、依赖版本或既有公共 API;F-005 仍是未形成直接可复现失败的架构残余,后续应单独立项而非在本轮无边界扩展。 diff --git a/audit-report.md b/audit-report.md deleted file mode 100644 index 69cf341..0000000 --- a/audit-report.md +++ /dev/null @@ -1,252 +0,0 @@ -# SASAbs v2 全量项目审计与发布闭环报告 - -审计日期:2026-07-13(Asia/Shanghai) -仓库:`E:\desktop\SASAbs_saxs-absolute-calibration` -当前分支:`codex/ci-release-followup` -审计起点:`main@6ba966c715753f59a44009f1ee2ab07d15fc93f5` -已发布基础:`main == origin/main == cfad783` -审计截止:follow-up 本地冻结点;follow-up 尚未 commit/merge/push,第二轮 CI 尚未触发 -审计模式:full + release hardening - -## 1. 执行摘要 - -结论:**follow-up 本地放行;最终完成以第二轮远端 CI 成功为条件。** - -原始审计共登记 24 项:P0=0、P1=10、P2=12、P3=2。当前状态为 -**verified=21、blocked=0、open=3**。全部 P1 已闭环;仍开放的 AUD-015、AUD-017、AUD-020 -分别受限于可见 GUI/人工 dark exposure 证据、GUI 主线程真实交互体验、真实大批次/束线性能基准, -不构成当前代码安全 gate 的已知阻塞。 - -本轮把版本升级到 2.0.0,并完成以下发布关键路径: - -1. 公共 API 恢复历史位置参数兼容;新参数 keyword-only。 -2. v2 CLI 使用显式安全参数;危险历史假设只通过显式 `bl19b2-abs2d-v1-legacy` 入口启用。 -3. BL19B2 重跑命令覆盖 mask、标准、几何修正与执行策略;monitor mode 单边界规范化。 -4. 标准端不确定度使用实际估计器的有限差分灵敏度;reference/system coverage 分离。 -5. 缺少 raw BG/dark covariance 时,combined uncertainty 正确保持 `partial`,不伪报 complete。 -6. CalibrationRecord v2 绑定源文件、模型、reference、integration 和 robust estimator;正式输出前重新读取并校验当前源文件。 -7. SRM 3600 别名统一、厚度锁定 0.1055 cm,并加入保守平行性 QC;μ 成分输入消除 1/100 尺度歧义。 -8. Tab2/Tab3 正式输出统一使用当前 record/context gate,严格处理 Q/2θ/χ 与 q 单位。 -9. Cal2D 五件套采用事务 staging;no-overwrite 以 create-if-absent 发布且冲突后不删除路径;overwrite 使用 backup rollback;整包 rerun 身份一致。 -10. text/canSAS/NXcanSAS 自产 1D provenance 可写入、重读并通过 formal gate。 -11. K history、稳定读后哈希、resume provenance、worker/stem、launcher cwd/log 等 P2 安全项闭环。 -12. README、architecture、runbook、CHANGELOG 与版本元数据同步至 v2 安全合同。 - -CI follow-up 本地冻结树代码、artifact 与 QA 门禁全部通过;这不等同于远端已绿: - -- 最终 pytest:`500 passed in 19.48s`;JUnit `.audit-work/pytest-release-final-v7.xml` -- 最终 ruff/compileall/diff-check:`全部通过` -- 最终 wheel/sdist build:通过;wheel 隔离 target import、CLI、launcher、entry points smoke 全部通过 -- 最终 wheel:`saxsabs-2.0.0-py3-none-any.whl`;SHA-256 `7081EA14E70CE317DCAA125C7131EFA3C4E5BBB51F25A4CC582215A8857DD281` -- 最终 sdist:`saxsabs-2.0.0.tar.gz`;SHA-256 `E569AA62C980963F06A96047A310D2E580AB6AEDD516F59870E8907ADBFE0595` -- QA:初始 release 的 A/B/C/D/resume 全部 PASS;follow-up focused `18 passed`;无新增 P0/P1/P2 - -### GitHub Actions CI follow-up - -- 初始 release commit `cfad783` 已 fast-forward 合并到 `main` 并 push;当前 `main == origin/main == cfad783`。 -- 首轮 GitHub Actions `ci` run `29228584901` 最终为 failure;除 ubuntu 3.11 外的 Ubuntu/macOS/Windows 矩阵均失败。 -- 根因一:Python 3.10 在 collection 阶段无 `tomllib`;version test 现改为 Python 3.10 兼容 regex。 -- 根因二:wheel 内容测试使用 `pip wheel --no-build-isolation`,但 dev 环境未显式安装 `setuptools>=69` 与 `wheel`;dev extra 已补齐。旧测试捕获 stderr 使日志不透明,现失败时显示 stdout/stderr。 -- follow-up 本地证据:focused `18 passed`;全量 `500 passed in 19.48s`;ruff/compileall/diff-check;final-v5 build、hash 与 smokes 全部通过。 -- 审计截止:follow-up 尚未 commit/merge/push,第二轮远端 CI 尚未触发;不得把本地通过表述为远端已绿。 - -## 2. 架构与发布范围 - -### 主要入口与数据流 - -- 安装包:`src/saxsabs/`。 -- 科研核心:`src/saxsabs/core/`。 -- 解析与导出:`src/saxsabs/io/`。 -- CLI:`src/saxsabs/cli.py` 与 `src/saxsabs/__main__.py`。 -- 桌面工作台:`SASAbs.py`、`saxsabs_workbench.py`、`saxsabs_workbench.pyw`。 -- BL19B2 批处理:`src/saxsabs/workflows/bl19b2_abs2d.py`。 -- Tab1:标准/BG/dark/几何 → K、record/context。 -- Tab2:样品/BG/dark → 归一化与校准 → 1D/sector/χ/Cal2D。 -- Tab3:外部或自产 1D → provenance/axis gate → buffer/绝对标度 → 多格式导出。 -- BL19B2:配置/输入快照 → reference/mask/K → per-frame 2D → QC/provenance/rerun。 - -### Git 与资产边界 - -- 初始 release commit `cfad783` 已通过 fast-forward 合并到 `main` 并 push;审计时 `main == origin/main == cfad783`。 -- 当前分支为 `codex/ci-release-followup`,以 `cfad783` 为基点;截至 follow-up 本地冻结点尚未 commit/merge/push。 -- 首轮远端 `ci` run `29228584901` 已终态 failure;第二轮 CI 尚未触发。 -- 用户资产 `docs/superpowers/` 保存在 `stash@{0}`(`preserve user docs/superpowers before SASAbs v2 release`),未纳入发布提交。 -- 本报告与 `.audit-work/*.md` 是审计产物;XML、wheel、sdist、cache 和 `.audit-work/sdist-*/` 工作目录不作为源码提交。 - -## 3. 风险闭环地图 - -| 区域 | 原风险 | 当前控制 | 状态 | -|---|---|---|---| -| API/CLI 兼容 | 位置参数错位;同版本隐式安全语义变化 | 恢复位置兼容;2.0.0 + 显式 legacy | verified | -| rerun/provenance | mask、策略、科学参数不能完整重放 | CLI/Config/rerun/signature 全字段贯通 | verified | -| uncertainty | 标准端缺项仍 complete;证书 k 外推 system | 有限差分与 shared-variable 灵敏度;coverage 分离;未知 covariance 为 partial | verified | -| calibration record | K 来源不完整;缓存验证可被源文件事后变化绕过 | schema v2 + source hash/model/参数;formal gate 每次重读复验 | verified | -| SRM/μ | 错厚度、不平行 ratio、成分尺度含糊 | 0.1055 cm、alias/QC、比例严格归一化 | verified | -| Tab2/Tab3 信任 | 默认 K、axis 混淆、auto/dry-run 分叉、1D 回读丢 provenance | 当前 record/context gate;严格 axis;共享验证;round-trip provenance | verified | -| Cal2D 输出 | 残包、跨运行混合、rollback 误删竞争文件 | 五件套 gate;create-if-absent no-overwrite;冲突时不删除已发布路径;overwrite backup rollback;package rerun ID | verified | -| history/input state | CSV 损坏覆盖、read→hash TOCTOU、resume 改写创建信息 | fail-closed+原子写;双哈希/stat;last_resume_validation | verified | -| launcher/path | cwd 阴影、不可写日志、极端 worker/stem | package 优先、日志 fallback、1..32、稳定短名 | verified | -| CI/packaging portability | Python 3.10 无 `tomllib`;no-isolation dev 环境缺 build tools;失败日志不透明 | 3.10 兼容 regex;显式 setuptools>=69/wheel;stdout/stderr 诊断 | local-pass / remote-pending | -| 可见 GUI | 布局、DPI、人工 dark exposure 路径未完成真实点击 | 仅 headless 逻辑 gate 有证据 | open | -| 真实性能 | 大批次/束线存储吞吐未知 | 不宣称性能提升 | open | - -## 4. 问题状态 - -| 状态 | ID | -|---|---| -| verified(21) | AUD-001~AUD-014、AUD-016、AUD-018、AUD-019、AUD-021~AUD-024 | -| open(3) | AUD-015、AUD-017、AUD-020 | -| blocked(0) | 无 | - -完整字段、触发条件、影响、修复和证据映射见 `.audit-work/findings.md`。 - -### QA 最后一轮补充闭环 - -- AUD-011:Cal2D no-overwrite 原先可能在后续成员冲突时删除竞争者已替换的目标。现在冲突后不删除任何已发布路径;保守保留的残片由完整性门拒绝。overwrite 模式仍使用 backup rollback。 -- AUD-007/AUD-010:正式 gate 不再只信任内存中的验证布尔值;每次重新读取 record、复验源文件与当前 context/K。自产 text/canSAS/NXcanSAS 写入并回读 operator provenance。 -- AUD-019/AUD-021:resume 校验不覆盖首次创建 provenance,另写 `last_resume_validation`。 -- AUD-008:工作流入口统一 SRM aliases,避免 core 正确而 workflow 分叉。 -- AUD-024/CI:run `29228584901` 暴露 Python 3.10 collection 与 no-build-isolation dev 依赖合同缺口;follow-up 已在本地以 focused 18、全量 500、final-v5 artifacts 闭环,第二轮远端复验待 push。 - -## 5. 科研语义审计结论 - -### 已验证合同 - -| 合同 | 当前结论 | -|---|---| -| K 方向与厚度单位 | 保持 `I_ref / I_meas_per_cm`;mm→cm 路径有回归 | -| SRM 3600 厚度 | 证书路径固定为 0.1055 cm;错误 override 提前拒绝 | -| SRM 平行性 | 非平行 ratio 不再静默压缩为单一 K | -| μ composition | 总和约 1 或 100 的完整输入归一化到 1;其他拒绝 | -| 标准端 uncertainty | T、MON、BG-MON、thickness、alpha 通过实际估计器有限差分传播 | -| 共享变量 | shared alpha 与 shared BG monitor 按联合灵敏度处理 | -| coverage | reference 与 system 分开;证书 coverage 不外推到未知组合预算 | -| combined status | raw BG/dark covariance 无输入时为 `partial`,未知项不设为 0 | -| custom reference | raw 与 canonical q/I/u/U、模型与哈希进入 record | -| axis/unit | Q、q_nm⁻¹、2θ、χ 分流;2θ 需要波长;不一致坐标拒绝 | - -### 严格边界 - -- 未提供 raw BG/dark covariance,因此不能声称获得完整 system expanded uncertainty。 -- 未用真实 BL19B2 或其他束线原始数据验证物理结果;当前证据来自测试夹具、匿名最小样例、人工数值反例和代码审查。 -- schema v1 记录仅为兼容读取,不升级为 complete 或可信正式输出来源。 -- 没有真实标准、样品、detector、PONI 与束线 metadata 的联合数据时,不把软件门禁等同于束线端计量认证。 - -## 6. 输出、provenance 与恢复安全 - -### CalibrationRecord v2 - -- source files 使用相对路径并按顺序记录 SHA-256。 -- 标准、BG、dark、custom reference、模型 id/version/canonical hash、alpha、background rule、integration unit/method/version/npt、robust estimator 参数均绑定到记录。 -- record 读取时验证 schema 与 source;正式 Tab2/Tab3 导出前再次从磁盘读取并验证,避免源文件在 GUI 会话中被篡改/删除后继续使用旧缓存状态。 -- K history CSV 遇损坏 fail-closed;更新使用临时文件、fsync 和原子替换。 - -### Cal2D package - -- resume 要求 image、mask NPY、mask EDF、PONI、metadata 全部存在、shape 一致且 context/参数/引用匹配。 -- 写入先进入 staging,再执行 no-overwrite 提交。 -- 任一成员冲突时回滚本事务已提交成员;只有仍与本事务 staged inode/file identity 一致的目标才删除,保护并发替换文件。 -- always-run 使用 package-level rerun ID,避免一个科学包混入多个运行的成员。 -- creation provenance 保持不变,后续验证写入独立 `last_resume_validation` 字段。 - -### 1D 导出回读 - -- text、canSAS XML 与 NXcanSAS HDF5 均可携带 calibration context fingerprint/operator provenance。 -- parser 把核心格式 provenance 与文本注释 provenance 合并。 -- 自产文件重新导入后仍必须经过当前 record/context gate,不因“由本软件写出”而放宽信任边界。 - -## 7. GUI 与可操作性 - -### 已有证据 - -- Tab2 auto-reference 不再被 fixed reference 预加载阻断。 -- dry-run 与正式输出复用同一安全 gate。 -- Tab3 正式输出要求当前、完整、源文件可验证的 calibration record。 -- raw 外部 1D 需要可审计 operator provenance;axis 和单位不允许猜测。 -- workers 限制 1..32,超长 stem 使用稳定短名。 -- launcher 避免 cwd shadow,日志可回退用户/临时目录。 - -### 未完成证据 - -- 自动化启动没有出现可接管的可见 GUI 窗口。 -- 因此未完成真实鼠标/键盘点击、窗口缩放、DPI、主题、滚动、控件遮挡与截图验收。 -- 未通过可见 GUI 验证人工 dark exposure 主路径。 -- 上述边界保留为 AUD-015/AUD-017 open,不能把 headless 测试写成视觉验收。 - -## 8. 性能与真实数据边界 - -已通过共享 parser、buffer 单次预载、稳定输入 gate 等方式减少可证明的重复工作,但没有在真实大批次、 -真实 detector 图像、NAS/束线存储、432 帧等条件下重新建立吞吐、峰值内存和失败恢复基准。 -AUD-020 保持 open;报告不声称性能提升,也不以本地匿名临时数据替代束线性能。 - -## 9. 验证状态 - -### 已完成的中间证据 - -- 最后一轮科学 QA:316 项聚焦测试通过,科研域 PASS。 -- BL19B2 最新聚焦回归:119 passed。 -- Cal2D exporter 事务/竞态聚焦回归:15 passed。 -- 发生在最终 QA 修复之前的中间全量回归:473 passed。 -- 中间 wheel 可构建、安装并包含 legacy GUI、launcher 和两个 console entry points。 - -这些是中间修复证据,不能替代最终冻结结果。最终树随后以 500 passed、静态检查通过、版本 2.0.0、wheel/sdist 构建与隔离 smoke 通过完成放行;最终 artifact 哈希见下表。 - -### 最终门禁 - -| 检查项 | 要求 | 当前结果 | -|---|---|---| -| 全量 pytest | 0 fail;保存 JUnit XML | `500 passed in 19.48s`;`.audit-work/pytest-release-final-v7.xml` | -| ruff | `ruff check src tests SASAbs.py saxsabs_workbench.py` | `通过` | -| compileall | package、GUI 与 launcher 全部可编译 | `通过` | -| diff check | `git diff --check` 无错误 | `通过` | -| version smoke | CLI、package、legacy 入口一致为 2.0.0 | `2.0.0` | -| wheel/sdist build | 冻结树重新构建 | `通过`;`saxsabs-2.0.0-py3-none-any.whl`、`saxsabs-2.0.0.tar.gz` | -| wheel install smoke | 隔离 target import、CLI、launcher、entry points | `全部通过` | -| wheel SHA-256 | 最终 artifact 身份 | `7081EA14E70CE317DCAA125C7131EFA3C4E5BBB51F25A4CC582215A8857DD281` | -| sdist SHA-256 | 最终 artifact 身份 | `E569AA62C980963F06A96047A310D2E580AB6AEDD516F59870E8907ADBFE0595` | -| 独立 QA | A/B/C/D/resume;无新增 P0/P1/P2 | 全部 PASS;定向 8 + 96 tests | - -详细命令、历史证据与最终证据槽位见 `.audit-work/evidence.md`。 - -## 10. 发布与 Git 建议 - -当前唯一结论:**follow-up 本地放行;最终完成以第二轮远端 CI 成功为条件。** - -首轮发布提交 `cfad783` 已完成快进合并与 push;以下步骤仅用于 CI 兼容性 follow-up, -不重复首轮发布操作。用户 `docs/superpowers/` 资产已安全保存在 `stash@{0}`,必须继续保留。 - -follow-up 的安全收口顺序: - -1. 只精确暂存 8 个 follow-up 文件,并检查 staged diff 与 `git diff --cached --check`。 -2. 在 `codex/ci-release-followup` 提交 CI 兼容性修复。 -3. 执行 `git fetch --prune`,确认 `origin/main` 仍为 `cfad783` 且没有远端漂移。 -4. 切换 `main`,执行 `git pull --ff-only origin main`,再执行 - `git merge --ff-only codex/ci-release-followup`。 -5. 必要 smoke 通过后执行 `git push origin main`。 -6. 等待第二轮 GitHub Actions CI 到终态且全部成功。 -7. 仅在第二轮 CI 全绿后,安全删除两个已合并本地分支: - `git branch -d codex/ci-release-followup` 与 - `git branch -d codex/sasabs-release-hardening`。 -8. 最后再次 `git fetch --prune`,确认工作区干净、`main == origin/main`,并确认用户 stash 仍存在。 - -不 force push,不使用 `-D`,不删除用户 stash/资产。任何最终门禁失败、远端漂移、 -第二轮 CI 失败或 staged 范围异常都应立即停止后续操作。 - -## 11. 最终自检 - -- [x] 24 项问题均有状态;P0/P1/P2/P3=0/10/12/2。 -- [x] verified/open/blocked=21/3/0。 -- [x] 全部 P1 闭环;QA 新发现已归入原问题账本。 -- [x] AUD-015、AUD-017、AUD-020 保持 open,未虚构可见 GUI 或真实束线/性能证据。 -- [x] combined uncertainty 的未知 covariance 明确保持 partial。 -- [x] 用户 `docs/superpowers/` 资产不在发布范围。 -- [x] 未把中间 473 passed 或旧 wheel 哈希冒充最终结果。 -- [x] 最终 pytest/ruff/compile/diff-check 已在冻结树通过:`500 passed in 19.48s`,静态门禁全部通过。 -- [x] 最终 wheel/sdist 构建、隔离 smoke 与 SHA-256 已记录并通过。 -- [x] 最终独立 QA A/B/C/D/resume 已通过;定向 8 + 96 tests,无新增 P0/P1/P2。 -- [x] 首轮 `cfad783` commit/快进合并/push 已完成。 -- [x] 首轮 CI `29228584901` 失败已诊断并记录。 -- [x] follow-up 修复、本地门禁与发布物复验已完成。 -- [ ] follow-up commit/快进合并/push:截至本审计冻结点尚待执行。 -- [ ] 第二轮远端 CI 终态成功:截至本审计冻结点尚待确认。 -- [ ] 已合并本地分支清理:仅在第二轮 CI 全绿后执行。 diff --git a/codemeta.json b/codemeta.json index 615c48b..8931b08 100644 --- a/codemeta.json +++ b/codemeta.json @@ -2,13 +2,13 @@ "@context": "https://doi.org/10.5063/schema/codemeta-2.0", "@type": "SoftwareSourceCode", "name": "saxsabs", - "description": "SAXS absolute intensity calibration and robust 1D profile utilities", - "identifier": "https://doi.org/10.5281/zenodo.19687104", - "url": "https://doi.org/10.5281/zenodo.19687104", + "description": "Traceable SAXS absolute-intensity calibration and profile I/O", + "identifier": "https://github.com/D-sudoasd/SASAbs", + "url": "https://github.com/D-sudoasd/SASAbs", "version": "2.0.0", "license": "https://spdx.org/licenses/BSD-3-Clause", - "codeRepository": "https://github.com/D-sudoasd/SASAbs_saxs-absolute-calibration", - "issueTracker": "https://github.com/D-sudoasd/SASAbs_saxs-absolute-calibration/issues", + "codeRepository": "https://github.com/D-sudoasd/SASAbs", + "issueTracker": "https://github.com/D-sudoasd/SASAbs/issues", "programmingLanguage": [ "Python" ], @@ -40,7 +40,7 @@ "SRM 3600" ], "dateCreated": "2026-02-25", - "dateModified": "2026-07-13", + "dateModified": "2026-08-16", "developmentStatus": "active", "softwareRequirements": [ "numpy >= 1.24", diff --git a/docs/api.md b/docs/api.md new file mode 100644 index 0000000..fefcd9e --- /dev/null +++ b/docs/api.md @@ -0,0 +1,182 @@ +# API and command-line reference + +This page covers the public calculation and I/O functions exported by +`saxsabs`, plus the installed `saxsabs` command. Inputs are not inferred when +their physical meaning is ambiguous; callers must supply the applicable +measurement semantics and provenance. + +## Core calculations + +### Monitor normalization + +```python +compute_norm_factor(exp: float | None, mon: float | None, + trans: float | None, mode: str) -> float +``` + +Returns the normalization product for `mode="rate"` (`exp * mon * trans`) or +`mode="integrated"` (`mon * trans`). Transmission must be in `(0, 1]`; missing, +non-positive, or non-finite required values return `math.nan`. An unknown mode +raises `ValueError`. + +```python +from saxsabs import compute_norm_factor + +factor = compute_norm_factor(1.0, 100000.0, 0.8, "rate") +assert factor == 80000.0 +``` + +### Robust K-factor estimation + +```python +estimate_k_factor_robust( + q_meas: np.ndarray, i_meas_per_cm: np.ndarray, + q_ref: np.ndarray | None = None, i_ref: np.ndarray | None = None, + q_window: tuple[float, float] = (0.01, 0.2), + positive_floor: float = 1e-9, min_points: int = 3, *, + i_ref_standard_uncertainty: np.ndarray | None = None, + coverage_factor: float | None = None, + standard_thickness_cm: float | None = None, + parallelism_relative_tolerance: float | None = None, +) -> KFactorEstimationResult +``` + +Interpolates the measured profile on the reference q grid in `q_window`, forms +`I_ref / I_meas`, and applies median/MAD outlier rejection. If both reference +arrays are omitted, it uses the built-in NIST SRM 3600 reference. The result +contains the estimate and diagnostics; inspect it before applying a scale. + +```python +from saxsabs import estimate_k_factor_robust + +result = estimate_k_factor_robust(q_meas, i_meas_per_cm, q_ref, i_ref) +print(result.k_factor, result.points_used) +``` + +### Attenuation and thickness + +```python +calculate_mu(composition: dict[str, float], density_g_cm3: float, + energy_keV: float) -> MuResult +calculate_material_attenuation(composition: Mapping[str, object], *, + composition_basis: str, table: AttenuationTable = NIST_30_KEV_TABLE, + material_key: str | None = None, material_name: str | None = None, + porosity_risk: bool = False) -> MaterialAttenuationResult +derive_fixed_thickness(material: MaterialAttenuationResult, + transmissions: Iterable[object], *, + anchor_scope: str = "provided_transmissions", + drift_warning_relative_span: float = ...) -> FixedThicknessDerivation +``` + +`calculate_mu` accepts element weight fractions or weight percent, bulk density +in g/cm³, and energy in keV. `calculate_material_attenuation` requires +`composition_basis="wt_fraction"` exactly, so wt% and atomic fractions are not +silently guessed. `derive_fixed_thickness` derives a fixed thickness from the +median supplied transmission and returns warnings alongside its provenance. + +```python +from saxsabs import calculate_mu + +mu = calculate_mu({"Ti": 0.9, "Nb": 0.1}, density_g_cm3=5.0, energy_keV=30.0) +print(mu.mu_linear_cm_inv) +``` + +### State checks, subtraction, and uncertainty + +```python +assess_intensity_state(profile: Mapping[str, object]) -> IntensityStateAssessment +subtract_buffer(q_sample, i_sample, err_sample, q_buffer, i_buffer, err_buffer, + alpha: float = 1.0, high_q_diag: tuple[float, float] = (0.15, 0.25), + *, alpha_uncertainty: float | None = None) -> BufferSubtractionResult +propagate_absolute_uncertainty(intensity: np.ndarray, *, + statistical_standard_uncertainty=None, k_relative_standard_uncertainty=None, + standard_relative_standard_uncertainty=None, + transmission_relative_standard_uncertainty=None, + monitor_relative_standard_uncertainty=None, + thickness_relative_standard_uncertainty=None, + mu_relative_standard_uncertainty=None, alpha_standard_uncertainty=None, + buffer_intensity=None, coverage_factor=None) -> AbsoluteUncertaintyBudget +``` + +`assess_intensity_state` classifies metadata, units, and correction information; +conflicting evidence remains ambiguous. `subtract_buffer` interpolates a buffer +onto the sample q grid when necessary and propagates supplied uncertainties. +`propagate_absolute_uncertainty` combines statistical and supplied standard +uncertainty components; relative inputs must be relative standard uncertainties. + +## I/O + +```python +read_external_1d_profile(path: str | Path) -> dict[str, Any] +write_cansas1d_xml(path: str | Path, q: np.ndarray, i_abs: np.ndarray, + err: np.ndarray | None = None, + metadata: dict[str, Any] | None = None) -> Path +write_nxcansas_h5(path: str | Path, q: np.ndarray, i_abs: np.ndarray, + err: np.ndarray | None = None, + metadata: dict[str, Any] | None = None) -> Path +``` + +`read_external_1d_profile` reads supported text/tabular profiles and routes XML +and HDF5 extensions to canSAS/NXcanSAS readers. The returned mapping includes +the parsed arrays and parsing metadata. Writers expect q in Å⁻¹, absolute +intensity in cm⁻¹, and optional uncertainty in cm⁻¹. NXcanSAS writing requires +the optional `h5py` dependency. + +```python +from saxsabs import read_external_1d_profile + +profile = read_external_1d_profile("examples/profile_example.csv") +print(profile["x"], profile["i_rel"]) +``` + +The parser deliberately names an uncalibrated intensity array `i_rel`. Do not +pass that array to an absolute-intensity writer. Call `write_cansas1d_xml` or +`write_nxcansas_h5` only after a validated calibration has established the +absolute intensity, uncertainty, units, and provenance. + +Workbench K-only scaling is fail-closed: the imported relative profile must +declare `thickness` in `corrections_applied`, a finite positive +`thickness_cm`, and a non-empty `thickness_source`. These provenance keys are +preserved by the text, canSAS, and NXcanSAS paths so a downstream operation does +not rely on a ledger marker alone. + +## Command line + +Run `saxsabs --help` for the installed command and `saxsabs --help` +for parameters. Required inputs are shown in angle brackets: + +```text +saxsabs norm-factor --mon --trans <0 --mode [--exp ] +saxsabs parse-header --header-json +saxsabs parse-external1d --input +saxsabs estimate-k --meas --ref [--q-col ] [--i-col ] + [--ref-q-col ] [--ref-i-col ] [--qmin ] [--qmax ] +saxsabs bl19b2-abs2d --input-root (--poni |--pydidas-cali-yaml ) + (--mu |--sample-thickness-cm ) + --monitor-mode [workflow options] +saxsabs bl19b2-abs2d-v1-legacy --input-root + (--poni |--pydidas-cali-yaml ) [migration options] +``` + +The main commands are: + +| Command | Input | Output | +| --- | --- | --- | +| `norm-factor` | exposure, monitor, transmission, mode | normalization value | +| `parse-header` | header JSON | extracted exposure/monitor/transmission JSON | +| `parse-external1d` | profile path | parsed-profile summary JSON | +| `estimate-k` | measured and reference tabular profiles | K-factor result JSON | +| `bl19b2-abs2d` | explicit BL19B2 inputs and semantics | batch result JSON and requested files | +| `bl19b2-abs2d-v1-legacy` | explicit migration choices | legacy-compatible batch result with documented assumptions | + +The BL19B2 commands require explicit monitor mode and either attenuation +coefficient or fixed thickness. Read the [batch runbook](bl19b2_abs2d_batch_runbook.md) +for their full input contract and provenance requirements. + +## Scientific boundary + +These APIs provide software operations and checks; they do not establish that a +beamline measurement is calibrated. Users remain responsible for appropriate +reference standards, independently measured inputs, detector geometry, valid +units, and experiment-specific acceptance. The included synthetic 2D example +tests arithmetic and interoperability, not experimental beamline validation. diff --git a/docs/architecture.md b/docs/architecture.md index d35820b..27ef62e 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -71,7 +71,8 @@ - Tab 3 raw correction is disabled. Formal K/Kd accepts only an explicitly reduced `relative` profile; `raw_counts`, `absolute_cm^-1`, and `ambiguous` states fail closed. K/d requires `d > 0`; K-only applies K without repeating - thickness and requires inherited thickness in `corrections_applied`. Tab 3 + thickness and requires inherited thickness in `corrections_applied`, a finite + positive `thickness_cm`, and a non-empty `thickness_source`. Tab 3 combines inherited corrections with the actual K, optional thickness, and optional buffer operations and records that set per frame; Tab 2 derives its ledger from the active detector context. @@ -121,19 +122,15 @@ - The desktop UI has screen-aware initial geometry and a `900 x 600` minimum, but it still has no cancellable background `JobController`; high-volume work can therefore occupy the Tk event thread. -- Repository hygiene remains open at P1: about 79 MiB of `audit_outputs/` plus - campaign-specific acceptance tests retain private `H:\...` path coupling. - They are not portable package fixtures and must be separated, reduced and - anonymized, or explicitly gated as local-only acceptance assets before release. ## Design goals -1. Preserve scientific behaviour from the production workflow. +1. Preserve scientific behaviour from the existing BL19B2 workflow. 2. Make core math independent from GUI toolkits. 3. Enable CI validation and reviewer reproducibility. 4. Support bilingual operation for international beamline user communities. -## Migration status +## Supported boundary and known limitations - **Implemented**: normalization, header parsing, external 1D parsing, robust K estimation, NIST 30 keV material core, Elam diagnostic calculator, 1D @@ -141,13 +138,19 @@ enforcement, disabled legacy/resume controls, exact K-only/Kd/buffer gates, absolute-buffer validation, provenance-aware scrollable μ UI, disabled Tab 3 raw mode, screen-aware startup, strict BL19B2 workflows, standard writers, - bilingual GUI, CLI, CI, and paper assets. -- **Still open (P0)**: keep formal multi-folder/per-sample fixed-thickness - campaigns in the strict CLI/batch owner until Workbench and strict-runner - kernels are unified; add Workbench campaign ownership, atomic publication, - and content-signature resume; require numeric inherited thickness and source - provenance for K-only formal output rather than only a ledger marker. -- **Still open (P1)**: close every FabIO path through a common loader, add a - cancellable background JobController, persist explicit CAUTION acceptance, - complete the DPI/theme/language/accessibility matrix, and separate the large - private-path-coupled campaign audit assets from portable package fixtures. + bilingual GUI, CLI, CI, and paper assets. K-only formal scaling requires both + the inherited-thickness ledger entry and numeric/source provenance. +- **Strict campaign ownership**: formal multi-folder and per-sample campaigns + remain owned by the strict CLI/batch runner. The Workbench is an interactive + interface and is not advertised as an equivalent campaign owner. +- **Publication and resume**: reusable calibrated-2D packages are transactional, + but whole-campaign atomic publication and content-signature resume are not + implemented in the Workbench or strict runner. Existence-only Workbench resume + is disabled rather than treated as safe. +- **Desktop operation**: long GUI jobs run on the Tk event thread and are not + cancellable. Users should prefer headless workflows for unattended or large + campaigns. `CAUTION` remains visible but is not separately persisted as an + acknowledgement. +- **Input resources**: not every Workbench FabIO path uses one common + copy-and-close helper. The strict readers exercised by the headless workflows + close their handles, and reviewers should use the documented portable fixtures. diff --git a/docs/author-confirmation-form.md b/docs/author-confirmation-form.md new file mode 100644 index 0000000..ac4f123 --- /dev/null +++ b/docs/author-confirmation-form.md @@ -0,0 +1,87 @@ +# Author confirmation form + +Complete every field truthfully, then use the answers to replace all four +`[Author input required before submission: ...]` markers in `paper/paper.md`. +Do not infer a declaration from repository metadata alone. + +## Authorship and correspondence + +- Final author list in order: +- Corresponding author: +- Corresponding email: +- Affiliation(s), including city and country: +- ORCID for each author: +- Confirmation that every listed author agrees to authorship and accountability: + +## Research use + +- Research question or experiment: +- Software version or commit used: +- Input data type and instrument/workflow: +- Commands or interface used: +- Outputs used in the research: +- How `saxsabs` affected the analysis: +- Public paper, preprint, data, workflow, or editor-visible evidence: +- Independent/external users or integrations, if any: + +## AI usage disclosure + +For each tool, give product, model, recoverable version/date, locations used, +and scope of assistance. If an earlier exact version cannot be recovered, state +that fact explicitly and retain supporting account/log evidence where possible. + +| Product | Model/version/date | Code/docs/paper locations | Nature and scope | +| --- | --- | --- | --- | +| GitHub Copilot | | | | +| Anthropic Claude | | | | +| OpenAI Codex | | | | +| Other | | | | + +Confirm verbatim if true: + +> All human authors reviewed, edited, and validated every AI-assisted output +> included in the submitted software, documentation, figures, and manuscript. +> The human authors made the core scientific, architectural, and design +> decisions and accept full responsibility for the submission. + +- Confirmation: yes / no + +## Funding, acknowledgements, sponsor role, and competing interests + +- Funding organization(s), grant number(s), or “No external funding”: +- Sponsor role in study design, software development, analysis, interpretation, + manuscript preparation, and submission decision: +- People/facilities to acknowledge: +- Competing interests, or explicit “The authors declare no competing interests”: + +## CRediT contributions + +Assign applicable roles to every author: Conceptualization, Data curation, +Formal analysis, Funding acquisition, Investigation, Methodology, Project +administration, Resources, Software, Supervision, Validation, Visualization, +Writing - original draft, and Writing - review & editing. + +| Author | Confirmed CRediT roles | +| --- | --- | +| | | + +## Final checks + +- Actual submission date (`D Month YYYY`): +- `paper.md` date updated to the actual submission date: +- Strict readiness command returns PASS: +- Confirmation JSON copied from `docs/submission-confirmations.example.json`, + completed from evidence, and passed with `--manual-confirmations`: +- Public CI URL for the submitted revision: +- Commit SHA submitted to JOSS: +- Public repository description, homepage concept DOI, default branch or + submitted branch, and visible README all match that commit: +- Confirmation date (`YYYY-MM-DD`, matching the paper submission date): +- Research-evidence reference retained for editorial verification: +- Submitted branch recorded in the confirmation JSON: +- Confirmed commit is the current clean HEAD of that submitted branch: +- Public-candidate command returns PASS and its editorialbot branch instruction + (if any) has been retained for the pre-review issue: + +The software tag, GitHub Release, and exact-version archive DOI are created +after successful JOSS review and recorded in the review issue before acceptance. diff --git a/docs/joss-submission-checklist.md b/docs/joss-submission-checklist.md index ab1c56c..538e44c 100644 --- a/docs/joss-submission-checklist.md +++ b/docs/joss-submission-checklist.md @@ -1,53 +1,94 @@ # JOSS submission checklist -## Repository essentials +This checklist follows the current JOSS author and reviewer documentation, +accessed 16 August 2026: -- [x] OSI-approved license file (`LICENSE`) -- [x] Installation instructions (`README.md`) -- [x] Citation metadata (`CITATION.cff`) -- [x] Public contribution guidance (`CONTRIBUTING.md`) -- [x] Code of Conduct referenced in `CONTRIBUTING.md` -- [x] `codemeta.json` for metadata interoperability -- [x] README badges (CI, license, Python version) +- [Submission requirements](https://joss.readthedocs.io/en/latest/submitting.html) +- [Paper format](https://joss.readthedocs.io/en/latest/paper.html) +- [Review criteria](https://joss.readthedocs.io/en/latest/review_criteria.html) +- [AI usage policy](https://joss.readthedocs.io/en/latest/submitting.html#ai-usage-policy) -## Software quality +## Pre-review screening gates -- [x] Installable package (`pyproject.toml`) -- [x] Automated tests (`tests/`) -- [x] Continuous integration (`.github/workflows/ci.yml`) -- [x] Multi-platform CI matrix (Ubuntu / Windows / macOS x Python 3.10-3.13) -- [x] Headless reproducibility path (`saxsabs` CLI) -- [x] `python -m saxsabs` entry point (`__main__.py`) -- [x] `__version__` in package `__init__.py` +- [ ] **More than six months of public development.** GitHub reports that this + repository was created on 25 February 2026. The date gate is therefore not + satisfied on 16 August 2026; 26 August 2026 is the first conservative + submission date, provided public development remains active. +- [ ] **Demonstrated research use.** The repository contains a concrete BL19B2 + workflow and reproducible synthetic validation material, but the author + must supply evidence that the software has been used in research. Claims + of external adoption, publications, or operational benefit require direct + evidence. +- [x] **Good open-source practices.** The project has an OSI-approved license, + packaging metadata, archived earlier releases, a changelog, tests, CI configuration, documentation, + contribution guidance, support pathways, and issue/PR templates. +- [x] **Iterative development.** The public history contains releases and + functional, safety, test, documentation, and packaging changes across the + available public period. This does not waive the six-month gate. -## Documentation +## Repository and documentation -- [x] Manual verification procedure (`examples/manual-verification.md`) -- [x] Architecture and design boundary (`docs/architecture.md`) -- [x] Reviewer FAQ (`docs/reviewer-faq.md`) -- [x] Google-style docstrings on all public API functions +- [x] BSD-3-Clause `LICENSE` file. +- [x] Source installation and optional dependencies documented in `README.md`. +- [x] CLI, GUI, core API, and minimal example documented. +- [x] Core API reference in `docs/api.md`. +- [x] Architecture and scientific boundaries in `docs/architecture.md`. +- [x] Automated tests and a configured Linux/Windows/macOS CI matrix for Python + 3.10--3.13. +- [x] `CONTRIBUTING.md`, `CODE_OF_CONDUCT.md`, bug/feature templates, and a PR + template. +- [x] `CITATION.cff` and `codemeta.json` identify the canonical repository + without presenting the project concept DOI as an exact 2.0.0 archive. + README, paper, and `.zenodo.json` label the concept DOI at project level. +- [ ] Immediately before submission, record green push and Draft-PR runs for + the exact submitted HEAD in the dated external validation record. Do not + embed a self-referential commit hash in this tracked checklist. +- [ ] Verify that the public repository description, homepage concept DOI, + visible README, and submitted branch identify the same candidate. +- [ ] Run `scripts/check_public_candidate.py` against the completed confirmation + JSON and retain its PASS output. If the paper is not on `main`, post the + reported `branch-where-paper-is` command in the JOSS pre-review issue. -## Paper requirements +## Paper -- [x] `paper/paper.md` exists -- [x] `paper/paper.bib` exists with DOI-bearing references -- [x] State-of-the-field section with concrete tool comparison table and 13+ rows -- [x] Research impact section with real deployment context -- [x] AI usage disclosure with specific tools and scope (JOSS 2025 policy) -- [x] Final author list, affiliations, and repository URL finalized +- [x] `paper/paper.md` uses JOSS Markdown/YAML metadata. +- [x] The body is within the 750--1750 word range by the documented Pandoc + plain-text count, excluding References and author-input markers; rerun the + readiness gate after author-controlled content is added. +- [x] Required sections are present: Summary, Statement of need, State of the + field, Software design, Research impact statement, AI usage disclosure, + Acknowledgements, and References. +- [x] Related software and scientific sources have been checked against DOI or + official records. +- [x] The paper distinguishes xraydb/Elam from the NIST SRD 126 fixed-energy + table and does not describe either as XCOM. +- [x] The workflow figure has editable SVG/PDF sources and the GUI image is a + window-scoped capture of the actual Workbench. +- [x] canSAS1d XML from the deterministic example validates against the official + version 1.1 XSD with zero errors. +- [ ] NXcanSAS HDF5 passes punx 0.3.5 with its bundled v2018.5 definitions, but + current NeXus definitions and a third-party application consumer remain + unverified because punx 0.3.5 cannot parse the current definition set. +- [ ] The author confirms author order, affiliation, corresponding author, + acknowledgements, funding, conflicts of interest, and contribution roles. +- [ ] The author confirms the complete AI disclosure and human review statement. +- [ ] The author supplies research-use evidence suitable for the impact section. +- [x] The current official Inara workflow converts the paper to TeX and + well-formed JATS with citations and figures resolved. +- [x] The current candidate PDF was built from Inara-generated TeX with + LuaLaTeX, and its figures, references, page bounds, fonts, and rendered + pages were checked. Rebuild it if author-controlled content changes. -## Data availability strategy +## Post-review release and archive -- [x] Synthetic examples for parser/calibration logic (expanded to 36 points) -- [x] Public anonymized mini-dataset or documented legal/data constraints statement - (includes deterministic minimal 2D package in `examples/minimal_2d/`; no proprietary data required) +These items follow successful JOSS review and are required before acceptance, +not before the initial submission: -## Pre-submission dry run +- [ ] Freeze the exact reviewed revision and create the approved version tag. +- [ ] Create the matching GitHub Release with verified wheel and source distribution. +- [ ] Archive that reviewed revision with Zenodo or another accepted archive and + record the version DOI. +- [ ] Align the archive title, version, author list, and DOI with the final paper. +- [ ] Report the final software version and archive DOI in the JOSS review issue. -- [x] `pytest -q` -- [x] `ruff check src tests` -- [x] JOSS paper PDF build via `openjournals/inara` in CI - -## Remaining blockers - -- [ ] 6 months of public Git history (JOSS requirement for privately-developed projects) +All unchecked remote actions require the repository owner's explicit approval. diff --git a/docs/paper-word-count.md b/docs/paper-word-count.md new file mode 100644 index 0000000..a399bba --- /dev/null +++ b/docs/paper-word-count.md @@ -0,0 +1,16 @@ +# JOSS paper word-count method + +The paper body is counted with Pandoc rather than by counting Markdown tokens: + +```powershell +pandoc paper/paper.md --from=markdown --to=plain --resource-path=paper \ + --output=paper-body.txt +``` + +For the submission-length check, remove the `References` section and the +bracketed author-input markers from `paper-body.txt`, then count tokens matching +`[A-Za-z0-9][A-Za-z0-9'./+^-]*`. + +The count must be regenerated after author-controlled content is added. The +current count is recorded by `scripts/check_submission_readiness.py` rather +than hard-coded here. diff --git a/docs/reviewer-faq.md b/docs/reviewer-faq.md index d4208a8..0fd5a52 100644 --- a/docs/reviewer-faq.md +++ b/docs/reviewer-faq.md @@ -1,8 +1,11 @@ # Reviewer FAQ -## Why keep the legacy GUI file? +## Why keep the root Workbench module? -The legacy GUI reflects production beamline operations. Core logic is being incrementally extracted to avoid behavior drift while improving reproducibility. +The root module is the maintained Tk desktop application and compatibility +entry point. Reusable scientific and I/O logic lives under `src/saxsabs/`; the +GUI remains separate because the current Workbench and strict BL19B2 campaign +runner have intentionally different ownership boundaries. ## How can this be tested without GUI? @@ -10,7 +13,7 @@ Core logic is exposed as importable APIs and CLI commands. Tests run headlessly ## Data cannot be fully public. How is reproducibility addressed? -The repository includes synthetic examples, an independent deterministic raw-frame +The repository includes synthetic examples, a deterministic raw-frame validation package (`examples/minimal_2d/`), and automated tests. A manual verification checklist documents exact commands and expected acceptance ranges. @@ -27,6 +30,16 @@ produces deterministic outputs (CSV/TSV/canSAS XML and optional NXcanSAS HDF5), and writes numerical K and sample-intensity errors to `summary.json`. This is a software golden test, not a substitute for measured beamline validation. +## Have the structured exports been checked outside the project readers? + +Yes, with a bounded result. On 15 August 2026, the minimal example's XML output +validated with zero errors against `cansas1d.xsd` from the official canSAS +`1dwg` repository (blob `c376e590bf6c297ee5664834183b6d09b5684318`). +The HDF5 output passed punx 0.3.5 against its bundled NeXus v2018.5 definitions +with 97 OK, 0 WARN, and 0 ERROR findings. punx 0.3.5 could not parse the current +NeXus `main` definitions (commit `6313522`), so the project does not claim +validation against current definitions or a third-party application consumer. + ## What is the software boundary? `saxsabs` is a reusable SAXS absolute-calibration package with deterministic diff --git a/docs/scientific_safety_ui_upgrade_audit_20260713.md b/docs/scientific_safety_ui_upgrade_audit_20260713.md deleted file mode 100644 index c8ec71a..0000000 --- a/docs/scientific_safety_ui_upgrade_audit_20260713.md +++ /dev/null @@ -1,274 +0,0 @@ -# SAXSAbs 科学安全与 UI 升级审计 - -初版:2026-07-13 - -本轮代码事实复核:2026-07-14 - -## 结论先行 - -本轮已经关闭了一批会直接诱发误操作的 Workbench 缺口:正式 Tab 2 只允许固定厚度;逐帧 -Beer-Lambert 与 Tab 2/Tab 3 existence-only resume 在 UI 禁用,强制赋值也会在 Dry Check 和 Run -双重失败关闭;K 与 μ 只读;BG/Dark library 改动立即使 preflight 失效;Tab 3 raw 禁用;K-only/Kd、 -`raw_counts`/relative/absolute 状态、`do_not_repeat` 和 absolute buffer 已形成精确阶段契约;NIST 30 keV μ provenance 绑定可用的 PONI -能量并处理 stale payload。主窗口与 μ 窗口也已经屏幕自适应/可滚动。严格 1D 和严格 2D resume -读取路径还补上了 FabIO close。 - -这些改动显著降低了“重复 K/厚度校正”“沿用旧 preflight”“手改 K”“误把 Elam 当 NIST”以及小屏 -不可操作的风险。但项目目前仍不能宣称 Workbench 等价于严格 BL19B2 campaign runner。以下关键 -P0/P1 边界仍未关闭: - -1. 正式多目录/每试样固定厚度 campaign 仍只能由 strict CLI/batch owner 管理; -2. Workbench 与严格 BL19B2 runner 仍是两套编排/科学 kernel; -3. K-only 只要求 inherited ledger 含 thickness,尚未要求其数值与来源; -4. Workbench 尚无 campaign-level 原子发布和 content-signature resume; -5. `audit_outputs/` 约 79 MiB,且 batch-specific tests/产物保留私有 `H:\...` 路径耦合,仓库发布卫生尚未关闭。 - -因此,2026A1756 正式重处理仍应以严格 runner 的 include manifest、thickness derivation、processing -signature 和 correction ledger 为权威。Workbench 当前适合标定、检查和受约束的交互式处理,不应 -被文档包装成严格 runner 的完整图形前端。 - -## 状态定义与审计边界 - -- **已实现**:当前工作树有实际执行入口和针对性代码/测试,不代表本轮已经重跑全量 H 盘验收。 -- **部分实现**:局部路径已经安全,但尚未覆盖同类入口或整批事务边界。 -- **待实现 P0**:可能改变绝对尺度、复用错误结果或破坏正式输出隔离。 -- **待实现 P1**:主要影响稳定性、可取消性、资源管理和规模化使用。 -- **冻结基线**:历史 v4 结果只读保留;升级不得原地补写或覆盖。 - -本次复核只评价仓库代码和文档。未重新计算 H 盘 v4 全树哈希,因此不在这里声称历史结果已完成 -升级前后哈希复验。 - -## 本轮已实现 - -| 项目 | 当前代码事实 | 安全效果 | 尚存边界 | -|---|---|---|---| -| 固定厚度正式门 | Tab 2 仅 formal fixed;per-frame Beer-Lambert radio disabled;强制 auto 后 Dry Check BLOCKED,Run 再拒绝 | 逐帧 T 只影响 norm,不会再从 Workbench 正式输出吸收到逐帧厚度 | 多目录/每试样固定厚度表和 owner 仍在 strict campaign path | -| legacy resume 正式门 | Tab 2/Tab 3 exists-only resume checkbutton disabled;强制开启后 Dry Check BLOCKED,Run 再拒绝 | 同名旧文件不会因仅存在而被正式跳过 | Workbench 尚无替代的 content-signature resume | -| K/μ 只读 | Tab 2、Tab 3 K 为 `readonly`;Tab 2 μ 也为 `readonly`,只能由当前 μ payload 写入 | 阻止临时手改 K/μ 与 provenance 脱钩 | 活动 CalibrationRecord 仍由现有 Workbench 上下文验证 | -| preflight 指纹硬门 | Run 初始禁用;Dry Check 生成规范化配置指纹;tracked value 与 BG/Dark library add/recursive-add/clear 立即失效;两个 Run 入口再校验 | 没有批准、BLOCKED 或配置已变时不能运行 | 文件 identity 目前是 resolved path、size、mtime,不是所有输入内容 SHA;CAUTION 尚无持久化确认 | -| Tab 3 raw 禁用 | raw 单选按钮禁用;即使程序性强制 raw,正式 K gate 也会拒绝 | 暗场曝光和 NIST blank 契约未统一前不再伪装成正式 raw-1D 全校正 | 原 raw 代码仍存在,应等待共享 2D 核心后再决定是否恢复 | -| 1D 强度状态与实际账本 | formal K/Kd 只接受明确 `relative`;`raw_counts`、absolute、ambiguous 均拒绝;`corrections_applied` 与 `do_not_repeat` union 后防重复,但 required-existing 的物理证明只认 `corrections_applied`;二者冲突即 ambiguous;K/d 要求 `d>0` 并加 K+thickness;K-only 只加 K 且要求 inherited thickness | raw 计数、已绝对化、状态不明、重复或冲突操作全部失败关闭 | K-only inherited thickness 的数值/来源尚未强制 | -| absolute buffer gate | buffer 必须 `absolute_cm^-1`、`1/cm`、`corrections_applied` 含 K+thickness、未含 buffer、显式完整 CalibrationContext fingerprint 且 numeric K 与 active K 匹配;operator payload/do-not-repeat 不可替代;Dry Check 验证 q coverage;报告含 `BufferKFactor`/`BufferAlphaUncertainty`;可选 `u(alpha)` 有限非负、空值保留 None/NaN;统计与合成标准不确定度分列,逐谱保存 alpha/不确定度模型;core kernel 缺失即失败关闭 | 防止相对/不同标度/重复 buffer、弱 fallback、q 外推和误标不确定度,并显式传播 scale uncertainty | 已关闭本轮 UI gate | -| NIST 30 keV 材料核心 | `core/material_attenuation.py` 固定表快照、wt% 解析、理想体积加和密度、μ/ρ、线性 μ、partial uncertainty、孔隙警告和 provenance SHA | NIST 与 xraydb/Elam 来源不再混名;μ 内部值不再截到两位 | NIST 快照只适用于 30 keV;理想密度不是实测/认证合金密度 | -| GUI μ provenance JSON | edited composition 决定 nominal/custom identity;任一输入改变即清空 μ/payload 并禁用 export;export 重读/重哈希 PONI path/content/energy,变化必须重算;Elam 禁用 porosity 并记录 xraydb version;NIST 记录 PONI identity/energy,可用能量非 30 keV 即拒绝;fixed formal metadata/preflight 不带诊断 μ 且记 `mu_used_in_thickness_model=false` | 材料身份、数据库版本、几何兼容、stale 状态和“μ 未参与 fixed 厚度”可追溯 | PONI energy 缺失时明确 `not geometry-bound`;payload 仍未绑定 per-folder accepted raw T | -| 屏幕自适应 | 主 root 和 μ Toplevel 使用 screen-aware geometry;Tab 2/3 与 μ 内容可滚动 | 1024 x 700 等紧凑视口更可操作 | 仍需多 DPI、双语、双主题和键盘矩阵 | -| FabIO close | 严格 1D 两个 reader 与严格 2D resume verifier 在 `finally` 关闭对象 | 这些路径的 Windows 锁/句柄风险已降低 | shared reference loader、严格 2D 主 loader 和多处 Workbench reader 仍未关闭 | - -## 科学链与当前契约 - -### 1. K 标定链 - -正式 K 必须来自 source-verified CalibrationRecord,并绑定 standard、blank、dark、reference、PONI、 -mask、flat、monitor mode、solid-angle/polarization 和估计器参数。Workbench 的 Tab 2/Tab 3 K 只读 -只解决“运行时手改值”的问题,不等于 GUI 与严格 runner 已经共享同一个标定编排实现。 - -SRM 3600 厚度仍锁定为 `0.1055 cm`。系统不确定度在共享 blank/dark 原始计数协方差未量化时应保持 -`partial`,不能用证书 coverage factor 代替整条处理链的 coverage factor。 - -### 2. 固定厚度、逐帧 T 的绝对 2D 链 - -对厚度不变的原位试样,目标公式保持为: - -```text -sample_norm = (S - D * exp_s / exp_D) / (exp_s * MON_s * T_s) -blank_norm = (BG - D * exp_bg / exp_D) / (exp_bg * MON_bg) -I_abs_2D = (sample_norm - alpha * blank_norm) * K / d_fixed -``` - -上式是 `rate` monitor 语义;`integrated` 语义去掉 exposure 因子,但仍保留每帧自己的 `MON_s * T_s`。 -固定厚度只替换厚度分支,不得取消每帧 transmission 归一化。 - -严格 runner 已支持每目录独立 include manifest 和 thickness derivation,并将其哈希纳入 processing -signature、重跑命令及逐帧 metadata。Workbench 现在把 per-frame Beer-Lambert 控件禁用,并在 -Dry Check/Run 两次拒绝强制 auto 值;这已关闭旧算法进入正式 GUI 输出的入口。它仍未把“审核帧集合 -→ T_rep/MAD/P5-P95 → d_fixed → per-folder row”做成 multi-folder campaign owner,因此这类正式 -任务继续由 strict CLI/batch path 承担。 - -### 3. NIST 30 keV 成分模型 - -当前锁定名义 wt% 与回归值为: - -| 材料 | 名义 wt% | μ (`cm^-1`) | -|---|---|---:| -| Ti-24Nb-4Zr-8Sn | Ti 64, Nb 24, Zr 4, Sn 8 | 74.550355 | -| Ti-6Al-4V | Ti 90, Al 6, V 4 | 20.989980 | -| Zr-2.5Nb | Zr 97.5, Nb 2.5 | 162.949617 | - -密度采用元素密度的理想比体积加和,参数来源必须标记为 `composition_model_derived`。它不是样品实际 -密度认证值;实际孔隙率未知时不得虚构误差条,EBM 样品还应保留孔隙风险警告。 - -`mu_calculator.py` 使用的是 xraydb 的 Elam 数据,适合任意能量诊断比较;它不是 NIST XCOM 接口, -也不能在文档或 UI 中写成“XCOM via xraydb”。 - -Workbench μ 字段已只读。μ 工具按当前编辑后的 composition 重新识别 nominal material;编辑预设成分 -后不会继续冒用 preset identity。source、energy/wavelength、preset、composition、density 或 porosity -任一改变都会清空 μ/provenance、禁用 export 并使 Tab 2 preflight 失效。Elam 模式禁用 porosity,JSON -记录 `xraydb_version`。NIST 模式记录 PONI path/hash/energy;若 PONI 提供能量且与 30 keV 不符,计算 -失败关闭;PONI energy 不可得时明确记录 `not geometry-bound`,不冒充已经完成几何绑定。 - -### 4. 绝对 2D 到 1D - -严格 1D workflow 的目标是 5500 个 `q_A^-1` 点、CSR 积分、应用 mask 和一次 solid-angle,不重复 -dark、background、MON、T、thickness 或 K。它会核对 correction ledger、PONI/mask、输入 manifest、 -逐帧签名和输出 profile 哈希。 - -Workbench 新增的 1D state/ledger 已能防止重复 K/Kd/buffer,并按本次实际执行生成账本: - -- `corrections_applied` 表示已经执行的物理操作;`do_not_repeat` 作为执行 guard,两者取 union 做重复检查; -- 两个 ledger 同时存在但集合不一致时,状态变为 ambiguous,正式缩放失败; -- K/d 要求有限且 `d>0`,本次加入 K 和 thickness; -- K-only 只加入 K,并要求 inherited ledger 已含 thickness,不会再次除厚度; -- Tab 3 将最终稳定序列化账本写入逐帧 `CorrectionsApplied`;Tab 2 按实际 CalibrationContext 推导 - dark/background/solid-angle/polarization/flat 等条目。 - -K-only 当前仍只证明“输入 ledger 声称 thickness 已执行”,尚未要求 inherited thickness 数值与来源。 -普通 Tab 2/Tab 3 也没有严格 1D 的 run/frame content-signature resume;不安全的 existence-only resume -已禁用并在 Dry Check/Run 失败关闭,而不是继续作为替代方案。 - -### 5. Absolute buffer subtraction - -buffer 只接受显式 `absolute_cm^-1` 且单位为 `1/cm` 的输入;物理账本 -`corrections_applied` 必须含 K 和 thickness、不得已含 buffer。它还必须携带显式完整 -CalibrationContext fingerprint,并记录与当前 active K 一致的 numeric `k_factor`;operator payload -fallback 和 `do_not_repeat` 都不能替代绝对尺度证明。Dry Check 在 Run 前验证 sample q grid 完全落入 -buffer q range,禁止端点外推;共享 core kernel 不可用时直接失败,不启用简化 fallback。 - -Workbench 现已提供可选 `u(alpha)`:必须是有限非负值,并传入唯一的 core `subtract_buffer` 参与误差 -传播;留空则保存为 `None`,combined uncertainty 继续为 NaN,不把未知值偷设为 0。文本谱中 -`Error_Statistical_cm^-1` 只含样品与缩放后 buffer 的统计项, -`Error_CombinedStandard_cm^-1` 再加入 `I_buffer^2 u(alpha)^2`,兼容列 -`Error_cm^-1` 在 buffer 路径上指向合成标准不确定度。每个单独谱文件自身也记录 buffer 安全 -文件名与 SHA-256、alpha、`u(alpha)`、传播公式和 uncertainty type;逐帧 report 与 run metadata -继续记录 buffer state、unit、physical ledger、context fingerprint、`BufferKFactor`、 -`BufferAlphaUncertainty`、完整 path 和 SHA-256。 -文本谱重读时,显式命名但全为 NaN 的 combined 列仍被保留为主误差列,不会静默回退到有限的 -statistical-only 列;若外部文件同时给出 statistical 和 combined-standard 且没有兼容 Error 列, -combined-standard 也不再受源列顺序影响。这保证“合成不确定度未知”在往返读取后仍然失败关闭。 - -## 仍未完成的 P0 - -| 编号 | 缺口 | 为什么仍是 P0 | 目标验收 | -|---|---|---|---| -| P0-1 | GUI 与 strict BL19B2 runner 未统一 | 同一公式仍可能在大回调和 workflow 中漂移 | GUI 只构造严格配置并调用共享 runner;逐像素 probe 两入口一致 | -| P0-2 | formal multi-folder/per-sample campaign 只在 strict owner | Workbench 单一 fixed 输入不能表达各目录 accepted raw T、T_rep、μ、d 与 drift | GUI 调用 strict campaign owner,并逐目录展示 frame identities、T_rep/MAD/P5-P95、μ/d 与 derivation hash | -| P0-3 | K-only inherited thickness 缺数值/来源 | ledger marker 能防重复,但不能完整追溯已经使用的 d | K-only formal input 必须携带有限正 thickness、单位、推导/测量来源和指纹,并与 operator context 一致 | -| P0-4 | Workbench 无 atomic campaign publish/content-signature resume | 虽然 exists-only resume 已彻底禁用,但 Workbench 仍不能安全续跑或证明整批一次完整发布 | owner manifest + staging + atomic publish;run/frame/input/output signature/hash 全匹配才 resume | - -## 仍未完成的 P1 - -| 编号 | 缺口 | 目标 | -|---|---|---| -| P1-1 | preflight identity/CAUTION 尚不完整 | 关键文件由 size/mtime 提升为内容 SHA;CAUTION 逐项确认并写入 run metadata | -| P1-2 | FabIO 只在部分路径 close | 统一 copy-array/copy-header + `finally close` loader,并覆盖正常/异常测试 | -| P1-3 | 无后台 JobController | 长任务进入可取消 worker;Tk 线程只接收进度/结果事件;关闭窗口可安全收尾 | -| P1-4 | 预检与正式 output root 仍需彻底解耦 | 纯预检不占 owner;只有显式报告根才写 preflight artifact | -| P1-5 | 大型 GUI callback 与状态分散 | 将读取、校验、处理、发布拆为测试化 service;增加只读 readiness card | -| P1-6 | 多 DPI/主题/语言/键盘证据不足 | 100/125/150/200% × 中英 × 深浅主题,验证滚动、焦点、截断和非颜色状态 | -| P1-7 | 大批量交互性能未闭环 | 2000 帧预检/筛选不冻结、可取消、计数和失败清单可导出 | -| P1-8 | `audit_outputs/` 与本批测试耦合 | 可复用脚本迁入受控模块;测试改为匿名合成 fixture;大型/私有产物保持 local-only 并显式 ignore | - -## UI 精修目标 - -### 全局 readiness 卡 - -顶栏应只读显示:K record/fingerprint、formula version、monitor mode、当前科学链、材料/μ 来源、厚度 -来源、correction ledger、preflight level/fingerprint、output owner/policy。用户不应再从多个输入框和日志 -自行推断“现在能否正式运行”。 - -### Tab 2:per-folder 表格 - -推荐列为: - -```text -folder | material | accepted/total | T_rep | MAD | P5-P95 | drift - | mu_source | mu_cm^-1 | d_fixed_cm | derivation_sha256 | status -``` - -点击行应展开 accepted/rejected frame identities 和原因。表格只接受 include manifest 中的相对路径; -重复、缺失、绝对路径和目录穿越必须失败关闭。已禁用的 per-frame Beer-Lambert 不应重新进入 formal -UI;若未来恢复诊断功能,必须调用隔离的 diagnostic owner,不能与正式结果共用输出目录。 - -### Tab 3:阶段转换而非公式选择 - -默认只允许两类操作: - -- 打开并检查 `absolute_cm^-1`,可查看/重导出,但不能再乘 K/Kd; -- K/d:将带 `intensity_state=relative`、兼容 operator context 且尚未 K/thickness 的 1D 转为 absolute, - `d` 必须有限且正; -- K-only:输入必须声明 thickness 已做,本次只乘 K;正式 provenance 仍需补 thickness 数值与来源; -- buffer:只接受 context 匹配、单位 `1/cm`、已含 K+thickness 且 q coverage 完整的 absolute profile。 - -raw 入口在 sample/BG/dark 独立 exposure、NIST blank、单位和阶段契约统一到共享 2D 核心前保持禁用。 - -### 输出与运行 - -每个正式批次应形成不可变 package: - -```text -// - config/owner.json - config/preflight.json - config/processing_signature.json - manifests/input_snapshot.csv - manifests/output_checksums.csv - metadata/ - data/ - qc/ - logs/ - completion.json -``` - -创建阶段写入 sibling staging 目录;只有全部目录、帧、哈希和 completion 校验通过后才原子发布。正式 -UI 不应原地 overwrite 历史 run;新结果用新 run ID,并在 metadata 中记录 `supersedes`。 - -## 实机截图对照验收方法 - -本轮 UI 截图属于会话级 visualization 产物,不在仓库内,也不是科学输出 provenance;因此本报告不 -硬链任何仓库外绝对路径。 - -验收方法如下: - -1. before/after 使用相同屏幕、窗口尺寸、DPI、语言、主题和 active tab; -2. 只捕获 Workbench 窗口,优先使用 Win32 `PrintWindow` 一类隔离窗口方法; -3. 排除带远程控制覆盖层、黑区、窗口边框缺失、文字不可辨认或滚动位置不一致的图片; -4. Tab 1/2/3 与 μ 窗口逐一对照可达性、滚动/截断、K/μ 只读、raw/legacy/resume/Run 禁用状态和状态栏; -5. 截图清单与实机验收日志一起保存在会话产物中,但发布文档只记录方法和结论,不伪装成版本化证据。 - -静态截图不能证明键盘顺序、长任务取消、文件锁、2000 帧性能或科学公式正确性;这些必须由单测、 -数值 probe 和动态手工验收分别证明。 - -## 发布验收矩阵 - -| 领域 | 期望 | 当前状态 | -|---|---|---| -| 固定厚度 formal gate | Tab 2 只允许 fixed;逐帧 T 仍影响 norm | **UI disabled + Dry Check/Run hard block 已实现** | -| legacy resume | Tab 2/3 exists-only control 不可选,强制开启也拒绝 | **已实现** | -| K/μ 只读 | Tab 2 K/μ 与 Tab 3 K 不能手改 | **已实现** | -| preflight | 无批准/BLOCKED/配置变动不能 Run;BG/Dark mutation 立即失效 | **已实现;内容 SHA 与 CAUTION 确认待补** | -| Tab 3 raw | UI 禁用且强制进入也被正式 gate 拒绝 | **已实现** | -| 1D 阶段/账本 | K/d `d>0`;K-only 不重复 thickness;do-not-repeat union/conflict;report ledger 一致 | **已实现;K-only thickness 数值/来源待补 P0** | -| absolute buffer | 1/cm + physical K/thickness + full context + matching K + q coverage + `u(alpha)` + 统计/合成不确定度分列 | **已实现;未知 `u(alpha)` 保持 None/NaN,逐谱 provenance 完整** | -| NIST μ | edited identity、stale invalidation、PONI energy check、Elam xraydb version | **已实现;缺 PONI energy 明示 not geometry-bound** | -| per-folder derivation | accepted raw T 清单 + per-folder 表格/owner + hash | **strict campaign 已有;Workbench 待实现 P0** | -| GUI/runner 单实现 | Workbench 直接调用严格 2D/1D workflow | **待实现 P0** | -| output owner/atomic batch | 整批要么完整发布,要么明确 incomplete | **待实现 P0** | -| Workbench safe resume | signature/hash/完整输出集一致才 skip | **待实现 P0;exists-only 已禁用** | -| FabIO close | 所有读取路径正常/异常均无句柄泄漏 | **部分实现 P1** | -| JobController | UI 不冻结、可取消、异常可收尾 | **待实现 P1** | -| 屏幕 geometry | 1024 x 700 初始可达 | **代码已实现;多 DPI 实机矩阵待补** | -| v4 冻结 | 升级前后 H 盘全树 SHA-256 清单一致 | **待现场复验** | -| 三材料 probe | 逐像素公式、HDF5/EDF 一致、K 固定 | **由主批次验收报告确认,不在本文档冒充完成** | - -## 发布判定 - -当前最准确的产品状态是: - -- 严格 BL19B2 runner:正式 campaign 处理基线; -- Workbench:固定厚度/legacy/resume 硬门、K/μ/preflight、K-only/Kd/do-not-repeat、absolute buffer、 - NIST/Elam/PONI provenance 等防误操作升级已落地,但尚不是 strict runner 的等价前端; -- multi-folder campaign owner、kernel unification、K-only inherited thickness 数值/来源、 - atomic campaign publish/content-signature resume 是必须明确保留的边界; -- FabIO 全路径、JobController、CAUTION 持久确认和多 DPI/accessibility 仍是工程 P1。 - -只有 P0 全部关闭、GUI 与 runner 单实现、v4 冻结哈希通过、三材料 probe/全量 resume 通过并完成多 DPI -实机矩阵后,才能把 Workbench 对外描述为“严格科学处理 UI”。 diff --git a/docs/submission-confirmations.example.json b/docs/submission-confirmations.example.json new file mode 100644 index 0000000..14c67dc --- /dev/null +++ b/docs/submission-confirmations.example.json @@ -0,0 +1,13 @@ +{ + "public_history_confirmed": false, + "repository_identity_confirmed": false, + "research_use_confirmed": false, + "authorship_confirmed": false, + "ai_disclosure_confirmed": false, + "funding_and_coi_confirmed": false, + "confirmed_on": "", + "research_evidence_reference": "", + "ci_run_url": "", + "submitted_branch": "", + "submitted_commit": "" +} diff --git a/examples/bl19b2_abs2d_template/processing_config.example.yml b/examples/bl19b2_abs2d_template/processing_config.example.yml index adbd5a4..e5f340e 100644 --- a/examples/bl19b2_abs2d_template/processing_config.example.yml +++ b/examples/bl19b2_abs2d_template/processing_config.example.yml @@ -2,12 +2,12 @@ schema: saxsabs.bl19b2_abs2d.config.v1 description: Template for future SPring-8 BL19B2 SAXS absolute-corrected 2D batches. # Replace these paths for each new beamtime. -input_root: H:/path/to/BL19B2 DATA/datXXX -pydidas_cali_yaml: H:/path/to/BL19B2 DATA/datXXX/reference_saxs/Cali.yaml +input_root: ./data/datXXX +pydidas_cali_yaml: ./data/datXXX/reference_saxs/Cali.yaml # Alternative when a pyFAI PONI already exists. Use either this or # pydidas_cali_yaml, not both. poni_path: null -output_root: H:/path/to/BL19B2 DATA/datXXX_absolute_corrected_2D +output_root: ./outputs/datXXX_absolute_corrected_2D # Reference files default to input_root/reference_saxs when omitted. If a # beamtime uses different names, pass the matching CLI options shown below diff --git a/examples/manual-verification.md b/examples/manual-verification.md index a47e865..69cb661 100644 --- a/examples/manual-verification.md +++ b/examples/manual-verification.md @@ -55,7 +55,7 @@ cannot be fully public. Expected: `k_factor` close to `2.0`, with non-zero `points_used`. -7. Verify the independent synthetic raw-frame → absolute intensity workflow: +7. Verify the deterministic synthetic raw-frame → absolute intensity workflow: ```bash python examples/minimal_2d/run_minimal_2d_pipeline.py @@ -63,7 +63,7 @@ cannot be fully public. Expected outputs in `examples/minimal_2d/outputs/`: - - `summary.json` reports `validation_type=independent_synthetic_raw_frames` + - `summary.json` reports `validation_type=deterministic_synthetic_raw_frames` - `k_relative_error < 0.005` - `sample_max_relative_error < 0.01` - `absolute_profile.csv`, `absolute_profile.tsv`, `absolute_profile.xml` exist @@ -136,8 +136,10 @@ front end to the strict BL19B2 campaign runner. requires inherited `thickness` in `corrections_applied`. Confirm `do_not_repeat` is unioned for duplicate-operation protection but cannot prove that required thickness was physically applied; disagreement with - `corrections_applied` makes the profile ambiguous. Note the remaining limit: - K-only does not yet require inherited thickness value/source provenance. + `corrections_applied` makes the profile ambiguous. K-only must also reject + missing, non-finite, non-positive `thickness_cm` and missing/placeholder + `thickness_source`. Confirm that accepted text, canSAS, and NXcanSAS outputs + retain both values for downstream K-only use. 8. Enable buffer subtraction with a compatible fixture. The buffer must be explicit `absolute_cm^-1`, unit `1/cm`, carry K+thickness in `corrections_applied`, have no existing `buffer` correction, carry an explicit @@ -159,10 +161,10 @@ front end to the strict BL19B2 campaign runner. Negative, non-finite, or malformed values must fail closed. Temporarily make the shared core kernel unavailable and confirm formal subtraction fails closed rather than using a weaker fallback. -9. Before packaging the repository, inspect `audit_outputs/` size and scan tests - for private drive roots such as `H:\...`. The current roughly 79 MiB campaign - audit tree and batch-specific path coupling are an open P1 boundary, not a - portable test-fixture design; do not claim release hygiene is complete. +9. Before packaging the repository, confirm `git status` contains no audit + outputs, build caches, downloaded literature, or private drive roots. Keep + manual evidence outside the repository and use only anonymized, portable + fixtures for version-controlled examples and tests. ## Screenshot comparison method @@ -181,16 +183,15 @@ each accepted before/after pair: - retain the session-local artifact inventory with the test log, but do not present it as a versioned release artifact or scientific-output provenance. -## Open release gates +## Known Workbench boundaries -The following checks must remain marked incomplete until the corresponding code -exists: +These capabilities are outside the current Workbench support contract. They are +not claimed by the README or paper; use the strict headless workflow where +applicable: - formal multi-folder/per-sample fixed-thickness campaigns have a Workbench owner equivalent to the strict CLI/batch campaign; - Workbench and strict BL19B2 runner use one shared scientific kernel; -- K-only requires inherited thickness numeric value and source, not just a - ledger marker; - Workbench output root has an owner manifest, atomic campaign publication, and content-signature resume (existence-only resume must remain disabled); - Workbench preflight binds critical file content hashes and persists explicit diff --git a/examples/minimal_2d/README.md b/examples/minimal_2d/README.md index 1d425d9..3664f61 100644 --- a/examples/minimal_2d/README.md +++ b/examples/minimal_2d/README.md @@ -1,7 +1,7 @@ # Minimal anonymized 2D reproducibility package -This folder provides a reviewer-friendly deterministic dataset and script for an -independent synthetic raw-frame validation without proprietary beamline files. +This folder provides a reviewer-friendly deterministic dataset and script for a +synthetic raw-frame validation without proprietary beamline files. The expected standard and sample curves are defined before reduction; the reference is not calculated from the measured profile. diff --git a/examples/minimal_2d/run_minimal_2d_pipeline.py b/examples/minimal_2d/run_minimal_2d_pipeline.py index c54fbc5..a2d6bfa 100644 --- a/examples/minimal_2d/run_minimal_2d_pipeline.py +++ b/examples/minimal_2d/run_minimal_2d_pipeline.py @@ -9,17 +9,11 @@ import argparse import json -import sys from pathlib import Path import numpy as np -REPO_ROOT = Path(__file__).resolve().parents[2] -SRC_DIR = REPO_ROOT / "src" -if str(SRC_DIR) not in sys.path: - sys.path.insert(0, str(SRC_DIR)) - -from saxsabs import ( # noqa: E402 +from saxsabs import ( build_nist_net_image, estimate_k_factor_robust, get_reference_data, @@ -193,11 +187,11 @@ def run_pipeline(output_dir: Path) -> dict[str, object]: ) output_meta = { - "title": "saxsabs independent synthetic raw-frame validation", + "title": "saxsabs deterministic synthetic raw-frame validation", "run": "minimal-2d-golden-001", "wavelength_A": float(geometry["wavelength_A"]), "sdd_m": float(geometry["distance_m"]), - "sample_name": "synthetic-independent-golden", + "sample_name": "synthetic-deterministic-golden", "instrument_name": "synthetic-detector", "detector_name": "synthetic-array", "process_name": "minimal_2d_pipeline", @@ -224,7 +218,7 @@ def run_pipeline(output_dir: Path) -> dict[str, object]: pass summary: dict[str, object] = { - "validation_type": "independent_synthetic_raw_frames", + "validation_type": "deterministic_synthetic_raw_frames", "points": int(q.size), "q_min": float(q.min()), "q_max": float(q.max()), diff --git a/generate_joss_paper.py b/generate_joss_paper.py deleted file mode 100644 index c062adc..0000000 --- a/generate_joss_paper.py +++ /dev/null @@ -1,603 +0,0 @@ -#!/usr/bin/env python3 -"""Generate the JOSS paper draft as a Word document (.docx). - -Run: python generate_joss_paper.py -Output: paper/saxsabs_joss_paper.docx -""" - -from pathlib import Path -from docx import Document -from docx.shared import Pt, Inches, RGBColor -from docx.enum.text import WD_ALIGN_PARAGRAPH -from docx.enum.table import WD_TABLE_ALIGNMENT -from docx.oxml.ns import qn - -# ── helpers ──────────────────────────────────────────────────────── -PAPER_DIR = Path(__file__).resolve().parent / "paper" - -def add_heading(doc, text, level=1): - h = doc.add_heading(text, level=level) - return h - -def add_para(doc, text, bold=False, italic=False, font_size=Pt(11)): - p = doc.add_paragraph() - run = p.add_run(text) - run.bold = bold - run.italic = italic - run.font.size = font_size - run.font.name = "Times New Roman" - p.paragraph_format.space_after = Pt(6) - p.paragraph_format.space_before = Pt(0) - return p - -def add_rich_para(doc, segments): - """segments: list of (text, bold, italic)""" - p = doc.add_paragraph() - for text, bold, italic in segments: - run = p.add_run(text) - run.bold = bold - run.italic = italic - run.font.size = Pt(11) - run.font.name = "Times New Roman" - p.paragraph_format.space_after = Pt(6) - return p - -def add_equation(doc, text, label=None): - p = doc.add_paragraph() - p.alignment = WD_ALIGN_PARAGRAPH.CENTER - run = p.add_run(text) - run.italic = True - run.font.size = Pt(11) - run.font.name = "Cambria Math" - if label: - run2 = p.add_run(f" ({label})") - run2.font.size = Pt(11) - run2.font.name = "Times New Roman" - p.paragraph_format.space_after = Pt(6) - p.paragraph_format.space_before = Pt(6) - return p - -def add_figure(doc, image_filename, caption, width=Inches(5.8)): - """Insert an image with a centred caption below it.""" - img_path = PAPER_DIR / image_filename - if not img_path.exists(): - add_para(doc, f"[Figure placeholder: {image_filename} not found]", italic=True) - return - p_img = doc.add_paragraph() - p_img.alignment = WD_ALIGN_PARAGRAPH.CENTER - p_img.add_run().add_picture(str(img_path), width=width) - p_img.paragraph_format.space_after = Pt(2) - p_cap = doc.add_paragraph() - p_cap.alignment = WD_ALIGN_PARAGRAPH.CENTER - run_cap = p_cap.add_run(caption) - run_cap.italic = True - run_cap.font.size = Pt(9) - run_cap.font.name = "Times New Roman" - p_cap.paragraph_format.space_after = Pt(10) - return p_img - -# ── main ─────────────────────────────────────────────────────────── -def build_paper(): - doc = Document() - - # -- Default style -- - style = doc.styles["Normal"] - style.font.name = "Times New Roman" - style.font.size = Pt(11) - style.paragraph_format.space_after = Pt(6) - - # ================================================================ - # TITLE - # ================================================================ - title_p = doc.add_paragraph() - title_p.alignment = WD_ALIGN_PARAGRAPH.CENTER - title_run = title_p.add_run( - "saxsabs: A Robust Workflow for Small-Angle X-ray Scattering " - "Absolute Intensity Calibration" - ) - title_run.bold = True - title_run.font.size = Pt(16) - title_run.font.name = "Times New Roman" - title_p.paragraph_format.space_after = Pt(4) - - # Author - author_p = doc.add_paragraph() - author_p.alignment = WD_ALIGN_PARAGRAPH.CENTER - ar = author_p.add_run("Delun Gong") - ar.font.size = Pt(12) - ar.font.name = "Times New Roman" - author_p.paragraph_format.space_after = Pt(2) - - # Affiliation - aff_p = doc.add_paragraph() - aff_p.alignment = WD_ALIGN_PARAGRAPH.CENTER - af = aff_p.add_run( - "Institute of Metal Research, Chinese Academy of Sciences, Shenyang 110016, China" - ) - af.font.size = Pt(10) - af.italic = True - af.font.name = "Times New Roman" - aff_p.paragraph_format.space_after = Pt(12) - - # ================================================================ - # SUMMARY - # ================================================================ - add_heading(doc, "Summary", level=1) - - add_para(doc, - "saxsabs is an open-source Python package that provides a complete, " - "reproducible workflow for small-angle X-ray scattering (SAXS) absolute " - "intensity calibration. It automates the data-reduction chain from raw " - "two-dimensional (2D) detector images to calibrated one-dimensional (1D) " - "scattering profiles on an absolute cross-section scale (cm⁻¹ sr⁻¹), " - "using the NIST Standard Reference Material 3600 (SRM 3600) glassy " - "carbon as the primary calibrant (NIST, 2016). The software comprises a " - "modular core library, a command-line interface (CLI), and a graphical " - "user interface (GUI) with bilingual support (Chinese/English)." - ) - - add_para(doc, - "The core library implements monitor-mode-aware normalization, robust " - "K-factor estimation via median absolute deviation (MAD) outlier " - "rejection, format-agnostic 1D profile parsing, and heterogeneous " - "header extraction. Building on pyFAI (Ashiotis et al., 2015) and fabio, " - "saxsabs adds the calibration-control and metadata-plumbing layers " - "typically handled by ad hoc local scripts at synchrotron beamlines." - ) - - # ================================================================ - # STATEMENT OF NEED - # ================================================================ - add_heading(doc, "Statement of need", level=1) - - add_para(doc, - "Converting SAXS detector images to absolute-scale intensities requires " - "dark-current subtraction, beam-monitor normalization, transmission and " - "thickness correction, azimuthal integration, calibration against a " - "reference standard, and batch reporting. In practice, each step is " - "complicated by real-world heterogeneity: header formats differ across " - "beamlines; metadata may reside in file headers, CSV tables, or manual " - "input; and 1D profiles use inconsistent delimiters and column names." - ) - - add_para(doc, - "Existing tools address individual stages well—pyFAI (Ashiotis et al., " - "2015) for integration, SasView (Doucet et al., 2018) for model " - "fitting, Dioptas (Prescher & Prakapenka, 2015) for 2D reduction, " - "BioXTAS RAW (Hopkins et al., 2017) for biological SAXS, and Irena " - "(Ilavsky & Jemian, 2009) for SAS modeling—but none provides a " - "dedicated end-to-end absolute-calibration workflow that jointly handles " - "metadata heterogeneity, multi-background averaging, robust K-factor " - "estimation, and batch processing with audit trails." - ) - - add_para(doc, - "saxsabs fills this gap, targeting beamline scientists and SAXS users " - "who need reproducible, automatable absolute-scaling in production " - "environments where metadata conventions are fluid and thousands of " - "exposures per session are routine." - ) - - # ================================================================ - # STATE OF THE FIELD - # ================================================================ - add_heading(doc, "State of the field", level=1) - - add_para(doc, - "The SAXS ecosystem provides mature tools for individual processing " - "stages: pyFAI for GPU-accelerated integration; SasView for model " - "fitting; Dioptas for interactive 2D reduction; BioXTAS RAW for " - "biological SAXS; DAWN (Basham et al., 2015) for plugin-based " - "diffraction processing; and Irena for broad SAS analysis under " - "Igor Pro." - ) - - add_para(doc, - "Absolute intensity calibration remains a procedural gap. While the " - "theory of calibration against NIST SRM 3600 glassy carbon is well " - "documented (Glatter & Kratky, 1982; NIST, 2016), the operational " - "workflow—parsing heterogeneous metadata, selecting normalization " - "modes, handling multi-background subtraction, computing a robust " - "scaling factor, and organizing traceable outputs—is typically left to " - "bespoke scripts that are neither tested nor version-controlled." - ) - - add_para(doc, "Table 1 summarizes the functional landscape:", bold=False, italic=True) - - # -- Table 1 -- - capabilities = [ - ("Capability", "pyFAI", "SasView", "Dioptas", "BioXTAS RAW", "Irena", "saxsabs"), - ("Azimuthal integration", "✓", " ", "✓", "✓", "✓", " "), - ("SAS model fitting", " ", "✓", " ", "✓", "✓", " "), - ("Heterogeneous header parsing", " ", " ", " ", " ", " ", "✓"), - ("Monitor-mode normalization", " ", " ", " ", " ", " ", "✓"), - ("Robust K-factor (MAD filtering)","" , " ", " ", " ", " ", "✓"), - ("Format-agnostic 1D ingestion", " ", " ", " ", "partial", " ", "✓"), - ("Multi-background averaging", " ", " ", " ", " ", " ", "✓"), - ("Headless CLI + CI-testable", "✓", "partial", " ", " ", " ", "✓"), - ] - - table = doc.add_table(rows=len(capabilities), cols=7) - table.style = "Light Grid Accent 1" - table.alignment = WD_TABLE_ALIGNMENT.CENTER - for r, row_data in enumerate(capabilities): - for c, cell_text in enumerate(row_data): - cell = table.cell(r, c) - cell.text = cell_text - for paragraph in cell.paragraphs: - paragraph.alignment = WD_ALIGN_PARAGRAPH.CENTER - for run in paragraph.runs: - run.font.size = Pt(9) - run.font.name = "Times New Roman" - if r == 0: # header row - for paragraph in cell.paragraphs: - for run in paragraph.runs: - run.bold = True - - cap1 = doc.add_paragraph() - cap1run = cap1.add_run("Table 1. Functional comparison of SAXS software tools. " - "saxsabs focuses on the calibration-control and metadata-plumbing layer " - "that bridges integration engines and absolute-scale reduction.") - cap1run.italic = True - cap1run.font.size = Pt(9) - cap1run.font.name = "Times New Roman" - cap1.paragraph_format.space_after = Pt(10) - - add_para(doc, - "saxsabs complements these tools by formalizing the calibration-control " - "layer that typically exists as private, untested scripts, making " - "absolute-scaling workflows reproducible and auditable." - ) - - # ================================================================ - # SOFTWARE DESIGN - # ================================================================ - add_heading(doc, "Software design", level=1) - - add_para(doc, - "The primary design goal is to separate numerical calibration logic " - "from UI concerns, enabling both interactive GUI use and headless " - "CLI execution. The software is organized into four layers:" - ) - - add_rich_para(doc, [ - ("Core numerical layer", True, False), - (" (saxsabs.core): implements monitor normalization and " - "robust K-factor estimation as deterministic, stateless functions " - "with no GUI or I/O side effects.", False, False), - ]) - - add_rich_para(doc, [ - ("I/O layer", True, False), - (" (saxsabs.io): provides format-agnostic header parsing (fuzzy " - "key matching with unit conversion) and multi-strategy 1D profile " - "parsing (three separator strategies with automatic column-role " - "inference by keyword matching).", False, False), - ]) - - add_rich_para(doc, [ - ("CLI layer", True, False), - (" (saxsabs.cli): exposes four subcommands (norm-factor, " - "parse-header, parse-external1d, estimate-k) for reproducible, " - "scriptable execution.", False, False), - ]) - - add_rich_para(doc, [ - ("GUI layer", True, False), - (" (SASAbs.py): a tkinter-based application with four tabbed " - "panels—K-Factor Calibration, Batch Processing, External 1D " - "Conversion, and Help—offering full bilingual (Chinese/English) " - "user interface support.", False, False), - ]) - - add_figure(doc, "fig_workflow.png", - "Figure 1. Calibration workflow of saxsabs. Input data (2D images, " - "instrument metadata, and the NIST SRM 3600 reference) flow through " - "header parsing, normalization, 2D background subtraction, pyFAI " - "integration, and robust K-factor estimation to produce calibrated " - "1D profiles with structured audit outputs.", - width=Inches(5.8)) - - add_para(doc, - "This architecture enables unit testing of core functions independently " - "of the GUI (15 automated tests across 3 OS × 3 Python versions), " - "while preserving the GUI for interactive beamline use (Figure 3). The " - "incremental migration from monolithic script to modular library avoids " - "abrupt workflow disruption." - ) - - add_figure(doc, "fig_gui.png", - "Figure 3. The saxsabs graphical user interface in English mode, " - "showing the four-tab layout: K-Factor Calibration, Batch " - "Processing, External 1D Conversion, and Help.", - width=Inches(5.8)) - - # -- Mathematical formulation -- - add_heading(doc, "Mathematical formulation", level=2) - - add_para(doc, - "The absolute intensity calibration workflow in saxsabs follows the " - "standard procedure documented for NIST SRM 3600 (NIST, 2016; " - "Glatter & Kratky, 1982). The key computational steps are:" - ) - - # Normalization - add_rich_para(doc, [ - ("Monitor normalization. ", True, False), - ("Two modes are supported depending on whether the detector records " - "count rates or integrated counts. In ", False, False), - ("rate", False, True), - (" mode, the normalization factor is:", False, False), - ]) - add_equation(doc, "N = t_exp × I₀ × T", "1") - add_para(doc, - "where t_exp is the exposure time (s), I₀ is the beam-monitor count, " - "and T is the sample transmission. In integrated mode, t_exp is omitted " - "since the detector signal already accumulates over the full acquisition " - "window:" - ) - add_equation(doc, "N = I₀ × T", "2") - - # 2D background subtraction - add_rich_para(doc, [ - ("2D background subtraction. ", True, False), - ("Given sample detector image D_s, dark-current image D_d, and one or " - "more background images {D_bg,i}, the net 2D scattering pattern is:", False, False), - ]) - add_equation(doc, - "I_net(x,y) = (D_s − D_d) / N_s − ⟨(D_bg,i − D_d) / N_bg,i⟩", - "3" - ) - add_para(doc, - "where ⟨·⟩ denotes pixel-wise averaging (nanmean) over all available " - "background images. This multi-background averaging reduces statistical " - "noise in the background estimate." - ) - - # K-factor - add_rich_para(doc, [ - ("Robust K-factor estimation. ", True, False), - ("After azimuthal integration of the net pattern via pyFAI to obtain " - "a 1D profile I_meas(q), the profile is interpolated onto the NIST " - "SRM 3600 reference grid (15 data points spanning q ∈ [0.008, 0.250] " - "Å⁻¹). Point-wise ratios are computed:", False, False), - ]) - add_equation(doc, "R_i = I_ref(q_i) / I_meas(q_i)", "4") - add_para(doc, - "Outlier rejection uses the median absolute deviation (MAD):" - ) - add_equation(doc, "σ̂ = 1.4826 × median(|R_i − R̃|)", "5") - add_para(doc, - "where R̃ = median(R_i). Points satisfying |R_i − R̃| > 3σ̂ are rejected, " - "and the K-factor is the median of the remaining inlier ratios:" - ) - add_equation(doc, "K = median(R_i) for |R_i − R̃| ≤ 3σ̂", "6") - add_para(doc, - "The factor 1.4826 ensures consistency with the standard deviation " - "under a Gaussian distribution. This robust estimator is resistant to " - "outliers caused by parasitic scattering, beamstop shadows, or detector " - "artefacts at the edges of the q-overlap region (Figure 2)." - ) - - add_figure(doc, "fig_kfactor_demo.png", - "Figure 2. Demonstration of the robust K-factor estimation algorithm. " - "(a) NIST SRM 3600 reference profile and a simulated measured profile " - "after rescaling by K. (b) Point-wise ratios R_i = I_ref / I_meas with " - "inlier points (green circles) and rejected outliers (red crosses); " - "the blue line and shaded band show the median K-factor and ±3σ̂ " - "acceptance region.", - width=Inches(5.8)) - - # Absolute conversion - add_rich_para(doc, [ - ("Absolute intensity conversion. ", True, False), - ("For each sample, the calibrated absolute intensity is:", False, False), - ]) - add_equation(doc, "I_abs(q) = K × I_1D(q) / d", "7") - add_para(doc, - "where d is the sample thickness in centimeters. When transmission " - "is available, the thickness can be estimated from the Beer–Lambert " - "relation:" - ) - add_equation(doc, "d = −ln(T) / μ", "8") - add_para(doc, - "where μ is the linear attenuation coefficient. For alloys or " - "multi-element samples, μ is computed from the XCOM mass attenuation " - "coefficients at the working energy:" - ) - add_equation(doc, "μ = ρ × Σ(w_i × (μ/ρ)_i)", "9") - add_para(doc, - "where ρ is the bulk density, w_i is the mass fraction of element i, " - "and (μ/ρ)_i is the mass attenuation coefficient." - ) - - # -- Batch processing features -- - add_heading(doc, "Batch processing and automation", level=2) - - add_para(doc, - "The GUI batch-processing pipeline automates the complete chain from " - "raw 2D images to calibrated 1D profiles. Key automation features " - "include:" - ) - - bullets = [ - "Automatic background and dark-current matching via a weighted scoring " - "function that compares exposure time, monitor counts, transmission, " - "and temporal proximity between sample and candidate reference files.", - - "Multi-background capillary subtraction, where multiple background " - "images are averaged pixel-wise to reduce statistical noise.", - - "Three azimuthal integration modes: full-ring, angular-sector (with " - "±180° wrapping support), and radial chi-profile extraction.", - - "Sector merging with inverse-variance weighting: I = Σ(I_k × w_k) / Σ(w_k).", - - "Data quality controls including ≥98% non-positive-value detection " - "and background normalization magnitude checking.", - - "Structured output traceability: each batch run produces a CSV report, " - "a JSON metadata file, and a K-factor history log with timestamps and " - "instrument parameters.", - ] - for b in bullets: - p = doc.add_paragraph(style="List Bullet") - run = p.add_run(b) - run.font.size = Pt(11) - run.font.name = "Times New Roman" - - # ================================================================ - # RESEARCH IMPACT STATEMENT - # ================================================================ - add_heading(doc, "Research impact statement", level=1) - - add_para(doc, - "saxsabs has been deployed for routine absolute intensity calibration " - "at the Institute of Metal Research, Chinese Academy of Sciences, " - "processing data from multiple synchrotron beamlines. It has replaced " - "manual spreadsheet-based procedures, reducing operator intervention " - "and eliminating errors from inconsistent header parsing." - ) - - add_para(doc, - "The software defines its impact along three measurable dimensions:" - ) - - impact_items = [ - ("Operational efficiency: ", False, - "Calibration previously requiring manual metadata extraction and " - "iterative K-factor fitting is now a single CLI invocation or GUI " - "session, reducing processing time from minutes to seconds."), - - ("Reliability: ", False, - "Defensive parsing and format-agnostic ingestion have eliminated " - "silent data-misinterpretation failures when switching between " - "instruments."), - - ("Traceability: ", False, - "Every run produces structured, deterministic output suitable for " - "version control and audit."), - ] - - for label, _, detail in impact_items: - p = doc.add_paragraph(style="List Bullet") - rl = p.add_run(label) - rl.bold = True - rl.font.size = Pt(11) - rl.font.name = "Times New Roman" - rd = p.add_run(detail) - rd.font.size = Pt(11) - rd.font.name = "Times New Roman" - - add_para(doc, - "Core algorithms are verified by 15 automated tests across " - "three operating systems and three Python versions under CI." - ) - - # ================================================================ - # AI USAGE DISCLOSURE - # ================================================================ - add_heading(doc, "AI usage disclosure", level=1) - - add_para(doc, - "The following AI-assisted coding tools were used during the " - "development of this software:" - ) - - ai_bullets = [ - "GitHub Copilot (VS Code) and Anthropic Claude were used for code " - "refactoring, internationalization extraction, test skeleton generation, " - "and initial documentation drafts.", - - "AI assistance was limited to scaffolding and boilerplate. All core " - "numerical algorithms were designed and validated by the author " - "independently.", - - "Every AI-generated fragment was reviewed, tested, and revised before " - "inclusion. Automated tests provide ongoing verification of scientific " - "correctness.", - ] - for b in ai_bullets: - p = doc.add_paragraph(style="List Bullet") - run = p.add_run(b) - run.font.size = Pt(11) - run.font.name = "Times New Roman" - - # ================================================================ - # ACKNOWLEDGEMENTS - # ================================================================ - add_heading(doc, "Acknowledgements", level=1) - - add_para(doc, - "The author thanks beamline scientists and users at the Institute of " - "Metal Research who provided practical feedback on data heterogeneity " - "and workflow failure modes during the development and deployment of " - "this software." - ) - - # ================================================================ - # REFERENCES - # ================================================================ - add_heading(doc, "References", level=1) - - references = [ - 'Ashiotis, G., Deschiber, A., Nawber, M., Wright, J. P., Karkoulis, D., ' - 'Picca, F. E., & Kieffer, J. (2015). The fast azimuthal integration ' - 'Python library: pyFAI. Journal of Applied Crystallography, 48(2), ' - '510–519. https://doi.org/10.1107/S1600576715004306', - - 'Basham, M., Filik, J., Wharmby, M. T., Chang, P. C. Y., El Sherif, B., ' - 'Sheratt, R., ... & Hart, M. L. (2015). Data Analysis WorkbeNch (DAWN). ' - 'Journal of Synchrotron Radiation, 22(3), 853–858. ' - 'https://doi.org/10.1107/S1600577515002283', - - 'Doucet, M., Cho, J. H., Alina, G., Bakber, J., Bouwman, W., Butler, P., ' - '... & Washington, A. (2018). SasView version 4.2. Zenodo. ' - 'https://doi.org/10.5281/zenodo.1412041', - - 'Glatter, O., & Kratky, O. (1982). Small Angle X-ray Scattering. ' - 'Academic Press. ISBN 0-12-286280-5.', - - 'Hopkins, J. B., Gillilan, R. E., & Ez, S. (2017). BioXTAS RAW: ' - 'improvements to a free open-source program for small-angle X-ray ' - 'scattering data reduction and analysis. Journal of Applied ' - 'Crystallography, 50(5), 1545–1553. ' - 'https://doi.org/10.1107/S1600576717011438', - - 'Ilavsky, J., & Jemian, P. R. (2009). Irena: tool suite for modeling ' - 'and analysis of small-angle scattering. Journal of Applied ' - 'Crystallography, 42(2), 347–353. ' - 'https://doi.org/10.1107/S0021889809002222', - - 'National Institute of Standards and Technology. (2016). Standard ' - 'Reference Material 3600: Absolute Intensity Calibration Standard for ' - 'Small-Angle X-ray Scattering. Certificate of Analysis. ' - 'https://www.nist.gov/srm', - - 'Prescher, C., & Prakapenka, V. B. (2015). DIOPTAS: a program for ' - 'reduction of two-dimensional X-ray diffraction data and data ' - 'exploration. High Pressure Research, 35(3), 223–230. ' - 'https://doi.org/10.1080/08957959.2015.1059835', - ] - - for i, ref in enumerate(references, 1): - p = doc.add_paragraph() - p.paragraph_format.left_indent = Inches(0.5) - p.paragraph_format.first_line_indent = Inches(-0.5) - run = p.add_run(f"[{i}] {ref}") - run.font.size = Pt(10) - run.font.name = "Times New Roman" - p.paragraph_format.space_after = Pt(4) - - return doc - - -# ── entry point ──────────────────────────────────────────────────── -if __name__ == "__main__": - out_dir = Path(__file__).resolve().parent / "paper" - out_dir.mkdir(exist_ok=True) - out_path = out_dir / "saxsabs_joss_paper.docx" - - doc = build_paper() - doc.save(str(out_path)) - print(f"JOSS paper saved to: {out_path}") diff --git a/paper/capture_gui_screenshot.py b/paper/capture_gui_screenshot.py index 071e6a4..cf53bd6 100644 --- a/paper/capture_gui_screenshot.py +++ b/paper/capture_gui_screenshot.py @@ -1,98 +1,107 @@ #!/usr/bin/env python3 -"""Capture a GUI screenshot of SASAbs for the JOSS paper. +"""Capture the SASAbs Workbench for documentation and the JOSS paper. -This script launches the SASAbs GUI off-screen, populates it with -representative demo data, and captures a screenshot. - -Run: python paper/capture_gui_screenshot.py -Output: paper/fig_gui.png +The capture is window-scoped and written atomically. It deliberately does not +fall back to a full-screen grab: a failed capture must never replace the paper +figure with unrelated desktop content. The default output preserves the full +window so interface context remains visible and the screenshot can be audited. """ -import sys, os, time +import argparse +import importlib.util +import os +import sys +import time +import tkinter as tk from pathlib import Path +from PIL import ImageGrab + # We need to import the main script's directory ROOT = Path(__file__).resolve().parent.parent sys.path.insert(0, str(ROOT)) os.chdir(str(ROOT)) -import tkinter as tk -from tkinter import ttk -from PIL import ImageGrab -import importlib.util - -def capture(): +def capture(output: Path, *, control_pane_only: bool = False) -> None: """Launch GUI, wait for render, capture screenshot, then destroy.""" - import tkinter as tk - - root = tk.Tk() - root.withdraw() # hide until module is loaded + bootstrap_root = tk.Tk() + bootstrap_root.withdraw() # hide until module is loaded # Dynamically load SASAbs — keep __name__=="SASAbs" so loader is happy, # but patch __name__ after loading to prevent if __name__=="__main__" from running. - spec = importlib.util.spec_from_file_location("SASAbs", str(ROOT / "SASAbs.py")) - mod = importlib.util.module_from_spec(spec) + module_spec = importlib.util.spec_from_file_location("SASAbs", str(ROOT / "SASAbs.py")) + if module_spec is None or module_spec.loader is None: + raise ImportError("could not load SASAbs.py") + workbench_module = importlib.util.module_from_spec(module_spec) # Patch argparse to avoid consuming sys.argv - import argparse - _orig = argparse.ArgumentParser.parse_args - argparse.ArgumentParser.parse_args = lambda self, args=None, ns=None: _orig(self, args=[], namespace=ns) + original_parse_args = argparse.ArgumentParser.parse_args + argparse.ArgumentParser.parse_args = lambda self, args=None, ns=None: original_parse_args( + self, args=[], namespace=ns + ) try: - spec.loader.exec_module(mod) + module_spec.loader.exec_module(workbench_module) except SystemExit: pass finally: - argparse.ArgumentParser.parse_args = _orig + argparse.ArgumentParser.parse_args = original_parse_args - root.destroy() # discard the temporary root + bootstrap_root.destroy() - # Re-create properly using the app's own main flow - root2 = tk.Tk() - root2.geometry("1280x800+50+50") - app = mod.SAXSAbsWorkbenchApp(root2, language="en") - root2.update_idletasks() - root2.update() + # Re-create properly using the app's own main flow. + app_root = tk.Tk() + app_root.geometry("1280x800+50+50") + workbench_module.SAXSAbsWorkbenchApp(app_root, language="en") + app_root.deiconify() + app_root.lift() + app_root.update_idletasks() + app_root.update() # Allow rendering to complete - root2.after(800, lambda: _do_capture(root2)) - root2.mainloop() + app_root.after( + 800, + lambda: _do_capture(app_root, output, control_pane_only=control_pane_only), + ) + app_root.mainloop() -def _do_capture(root): - """Take the screenshot and close.""" +def _do_capture(root: tk.Tk, output: Path, *, control_pane_only: bool) -> None: + """Capture the native top-level window and close the application.""" root.update_idletasks() root.update() time.sleep(0.3) - # Get window geometry - x = root.winfo_rootx() - y = root.winfo_rooty() - w = root.winfo_width() - h = root.winfo_height() - - outpath = ROOT / "paper" / "fig_gui.png" - + temporary = output.with_name(f"{output.stem}.tmp{output.suffix}") try: - img = ImageGrab.grab(bbox=(x, y, x + w, y + h)) - img.save(str(outpath), "PNG") - print(f" ✓ fig_gui.png ({w}x{h})") - except Exception as e: - print(f" ✗ Screenshot failed: {e}") - # Fallback: save as generic placeholder - print(" Attempting fallback with full-screen grab...") - try: - img = ImageGrab.grab() - img.save(str(outpath), "PNG") - print(f" ✓ fig_gui.png (full screen fallback)") - except Exception as e2: - print(f" ✗ Fallback also failed: {e2}") - - root.destroy() + output.parent.mkdir(parents=True, exist_ok=True) + image = ImageGrab.grab(window=root.winfo_id()) + if control_pane_only: + image = image.crop((0, 0, min(475, image.width), image.height)) + image.save(temporary, "PNG") + temporary.replace(output) + print(f"OK: {output} ({image.width}x{image.height})") + except Exception as exc: + temporary.unlink(missing_ok=True) + print(f"ERROR: window capture failed: {exc}", file=sys.stderr) + root.destroy() + raise SystemExit(1) from exc + else: + root.destroy() if __name__ == "__main__": - print("Capturing SASAbs GUI screenshot ...") - capture() - print("Done.") + parser = argparse.ArgumentParser() + parser.add_argument( + "--output", + type=Path, + default=ROOT / "paper" / "fig_gui.png", + ) + parser.add_argument( + "--control-pane-only", + action="store_true", + help="save only the left control pane instead of the full application window", + ) + args = parser.parse_args() + capture(args.output.resolve(), control_pane_only=args.control_pane_only) diff --git a/paper/fig_gui.png b/paper/fig_gui.png index c1e6ce1..33d69e2 100644 Binary files a/paper/fig_gui.png and b/paper/fig_gui.png differ diff --git a/paper/fig_kfactor_demo.pdf b/paper/fig_kfactor_demo.pdf new file mode 100644 index 0000000..216d9b8 Binary files /dev/null and b/paper/fig_kfactor_demo.pdf differ diff --git a/paper/fig_kfactor_demo.png b/paper/fig_kfactor_demo.png index 283f933..f7d1b3d 100644 Binary files a/paper/fig_kfactor_demo.png and b/paper/fig_kfactor_demo.png differ diff --git a/paper/fig_kfactor_demo.svg b/paper/fig_kfactor_demo.svg new file mode 100644 index 0000000..60a929f --- /dev/null +++ b/paper/fig_kfactor_demo.svg @@ -0,0 +1,1087 @@ + + + + + + + + 2026-08-12T18:22:57.734991 + image/svg+xml + + + Matplotlib v3.10.8, https://matplotlib.org/ + + + + + + + + + + + + + + + + + + + + + + + + + + + + 0.00 + + + + + + + + + + 0.05 + + + + + + + + + + 0.10 + + + + + + + + + + 0.15 + + + + + + + + + + 0.20 + + + + + + + + + + 0.25 + + + + + + + q +   + ( + ) + Å + + 1 + + + + + + + + + + + + + + + + + + + 1 + 0 + 1 + + + + + + + + + + + + + + + 1 + 0 + 2 + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + I + q + ( + ) + c + m +   + ( + ) + + 1 + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + a + + + + + + SYNTHETIC DEMO + + + Reference and rescaled synthetic profile + + + + + + + + + + NIST SRM 3600 + + + + + + + + Synthetic measured profile + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + 0.00 + + + + + + + + + + 0.05 + + + + + + + + + + 0.10 + + + + + + + + + + 0.15 + + + + + + + + + + 0.20 + + + + + + + + + + 0.25 + + + + + + + q +   + ( + ) + Å + + 1 + + + + + + + + + + + + + 0.02 + + + + + + + + + + 0.04 + + + + + + + + + + 0.06 + + + + + + + + + + 0.08 + + + + + + + + + + 0.10 + + + + + + + + + + 0.12 + + + + + + + I + I + r + e + f + m + e + a + s + / + + + + + + + + + + + + + + + b + + + Robust K estimate + + + + + + + + + Inlier + + + + + + + + Rejected + + + + + + K = 0.0350 + + + + + + + + + + + + + diff --git a/paper/fig_workflow.pdf b/paper/fig_workflow.pdf new file mode 100644 index 0000000..8f27785 Binary files /dev/null and b/paper/fig_workflow.pdf differ diff --git a/paper/fig_workflow.png b/paper/fig_workflow.png index e4e311d..07373af 100644 Binary files a/paper/fig_workflow.png and b/paper/fig_workflow.png differ diff --git a/paper/fig_workflow.svg b/paper/fig_workflow.svg new file mode 100644 index 0000000..e2cf321 --- /dev/null +++ b/paper/fig_workflow.svg @@ -0,0 +1,355 @@ + + + + + + + + 2026-08-12T18:22:57.252363 + image/svg+xml + + + Matplotlib v3.10.8, https://matplotlib.org/ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + Inputs + + + Interfaces + + + Scientific core + + + Outputs + + + 2D detector frames + and instrument headers + + + External 1D profiles + and correction metadata + + + SRM 3600, water, or + a user-supplied reference + + + CLI utilities and strict + BL19B2 batch workflow + + + SAXSAbs Workbench + calibration and export + + + Python API + public package functions + + + Parse, normalize, and + validate inputs + + + Reduce detector data and + integrate with pyFAI + + + Estimate robust K and + propagate uncertainty + + + Check intensity state and + correction history + + + Write canSAS, NXcanSAS, + and calibrated 2D outputs + + + Absolute-scale 1D profiles + and calibrated 2D packages + + + canSAS XML and + NXcanSAS HDF5 + + + Calibration records, source + hashes, QC, and run reports + + + Required checks before calibrated output: units · transmission · thickness · correction state · calibration provenance + + + + + + + + + diff --git a/paper/generate_figures.py b/paper/generate_figures.py index 68fb631..cc0450d 100644 --- a/paper/generate_figures.py +++ b/paper/generate_figures.py @@ -1,217 +1,356 @@ #!/usr/bin/env python3 -"""Generate publication-quality figures for the JOSS paper. +"""Generate the code-derived figures used by the JOSS submission. -Run: python paper/generate_figures.py -Output: - paper/fig_workflow.png – calibration workflow diagram - paper/fig_kfactor_demo.png – K-factor estimation demonstration +The workflow figure is a schematic derived from the package modules and public +interfaces. The K-factor panel uses deterministic synthetic data solely to +illustrate the robust estimator; it is not an experimental performance claim. """ +from __future__ import annotations + +import argparse from pathlib import Path -import numpy as np -import matplotlib -matplotlib.use("Agg") + +import matplotlib as mpl + +mpl.use("Agg") + import matplotlib.pyplot as plt -from matplotlib.patches import FancyBboxPatch, FancyArrowPatch +import numpy as np +from matplotlib.patches import FancyArrowPatch, FancyBboxPatch + +from saxsabs import estimate_k_factor_robust from saxsabs.constants import NIST_SRM3600_DATA -OUT_DIR = Path(__file__).resolve().parent -# ══════════════════════════════════════════════════════════════════════ -# Figure 1 — Calibration Workflow Diagram -# ══════════════════════════════════════════════════════════════════════ +MM_PER_INCH = 25.4 +FULL_WIDTH_MM = 183.0 +FULL_WIDTH_IN = FULL_WIDTH_MM / MM_PER_INCH + +mpl.rcParams.update( + { + "font.family": "sans-serif", + "font.sans-serif": ["Arial", "Helvetica", "DejaVu Sans", "sans-serif"], + "font.size": 7.5, + "axes.titlesize": 7, + "axes.labelsize": 7, + "legend.fontsize": 7, + "xtick.labelsize": 7, + "ytick.labelsize": 7, + "axes.spines.right": False, + "axes.spines.top": False, + "axes.linewidth": 0.7, + "figure.facecolor": "white", + "pdf.fonttype": 42, + "svg.fonttype": "none", + } +) + +COLORS = { + "ink": "#23343D", + "muted": "#60737D", + "input": "#E8F0FA", + "interface": "#FFF0D9", + "core": "#E7F3EA", + "output": "#F3E8F1", + "gate": "#EEF1F3", + "accent": "#2F6F8F", + "warning": "#C7772A", +} + + +def save_figure(fig: plt.Figure, output_dir: Path, stem: str) -> None: + """Save editable vectors and a 600 dpi raster at the declared final size.""" + + output_dir.mkdir(parents=True, exist_ok=True) + svg_path = output_dir / f"{stem}.svg" + fig.savefig(svg_path) + # Matplotlib terminates many SVG path lines with a space. Normalize the + # generated text so the version-controlled source passes Git whitespace checks. + svg_lines = svg_path.read_text(encoding="utf-8").splitlines() + svg_path.write_text( + "\n".join(line.rstrip() for line in svg_lines) + "\n", + encoding="utf-8", + ) + fig.savefig(output_dir / f"{stem}.pdf") + fig.savefig(output_dir / f"{stem}.png", dpi=600) + + +def _box( + ax: plt.Axes, + x: float, + y: float, + width: float, + height: float, + text: str, + facecolor: str, + *, + fontsize: float = 6.6, + weight: str = "normal", +) -> None: + patch = FancyBboxPatch( + (x - width / 2, y - height / 2), + width, + height, + boxstyle="round,pad=0.08,rounding_size=0.12", + facecolor=facecolor, + edgecolor=COLORS["ink"], + linewidth=0.8, + ) + ax.add_patch(patch) + ax.text( + x, + y, + text, + ha="center", + va="center", + fontsize=fontsize, + fontweight=weight, + color=COLORS["ink"], + linespacing=1.2, + ) + + +def _arrow( + ax: plt.Axes, + start: tuple[float, float], + end: tuple[float, float], + *, + color: str | None = None, +) -> None: + ax.add_patch( + FancyArrowPatch( + start, + end, + arrowstyle="-|>", + mutation_scale=8, + color=color or COLORS["muted"], + linewidth=0.9, + shrinkA=3, + shrinkB=3, + ) + ) -def make_workflow_figure(): - fig, ax = plt.subplots(figsize=(10, 6.5)) - ax.set_xlim(0, 10) - ax.set_ylim(0, 7) + +def make_workflow_figure(output_dir: Path) -> None: + """Render the package architecture and the checks applied before export.""" + + fig, ax = plt.subplots(figsize=(FULL_WIDTH_IN, 3.65)) + fig.subplots_adjust(left=0.012, right=0.988, bottom=0.025, top=0.985) + ax.set_xlim(0, 14) + ax.set_ylim(0, 8) ax.axis("off") - fig.patch.set_facecolor("white") - - # ── colour palette ── - c_input = "#E8F0FE" # light blue - c_core = "#FFF3E0" # light orange - c_output = "#E8F5E9" # light green - c_ref = "#FCE4EC" # light pink - c_border = "#455A64" - c_arrow = "#37474F" - - def box(x, y, w, h, text, color, fontsize=9, bold=False): - bx = FancyBboxPatch( - (x - w/2, y - h/2), w, h, - boxstyle="round,pad=0.12", - facecolor=color, edgecolor=c_border, linewidth=1.2, - ) - ax.add_patch(bx) - weight = "bold" if bold else "normal" - ax.text(x, y, text, ha="center", va="center", - fontsize=fontsize, fontweight=weight, color="#212121", - wrap=True) - return bx - - def arrow(x1, y1, x2, y2, text="", curved=False): - style = "Simple,tail_width=0.6,head_width=6,head_length=4" - if curved: - a = FancyArrowPatch( - (x1, y1), (x2, y2), - connectionstyle="arc3,rad=0.25", - arrowstyle=style, color=c_arrow, linewidth=1.0, - ) - else: - a = FancyArrowPatch( - (x1, y1), (x2, y2), - arrowstyle=style, color=c_arrow, linewidth=1.0, - ) - ax.add_patch(a) - if text: - mx, my = (x1 + x2) / 2, (y1 + y2) / 2 - ax.text(mx + 0.15, my + 0.12, text, fontsize=7, - color="#616161", style="italic") - - # ── Row 1: Inputs ── - y1 = 6.0 - box(1.5, y1, 2.4, 0.75, "2D Detector\nImages", c_input, 9, True) - box(5.0, y1, 2.4, 0.75, "Instrument\nMetadata", c_input, 9, True) - box(8.5, y1, 2.4, 0.75, "NIST SRM 3600\nReference", c_ref, 9, True) - - # ── Row 2: Parsing ── - y2 = 4.7 - box(3.25, y2, 3.0, 0.65, "Header Parsing & Normalization", c_core, 9) - arrow(1.5, y1 - 0.38, 3.25, y2 + 0.33) - arrow(5.0, y1 - 0.38, 3.25, y2 + 0.33) - - # ── Row 3: 2D Processing ── - y3 = 3.6 - box(3.25, y3, 3.5, 0.65, "Dark Subtraction → BG Subtraction → pyFAI Integration", c_core, 8.5) - arrow(3.25, y2 - 0.33, 3.25, y3 + 0.33) - - # ── Row 4: K-factor ── - y4 = 2.5 - box(5.5, y4, 3.5, 0.65, "Robust K-factor Estimation\n(median + MAD outlier rejection)", c_core, 8.5) - arrow(3.25, y3 - 0.33, 5.5, y4 + 0.33) - arrow(8.5, y1 - 0.38, 5.5, y4 + 0.33, curved=False) - - # ── Row 5: Absolute conversion ── - y5 = 1.4 - box(5.0, y5, 3.0, 0.65, "I_abs(q) = K × I_1D(q) / d", c_core, 10, True) - arrow(5.5, y4 - 0.33, 5.0, y5 + 0.33) - - # ── Row 6: Outputs ── - y6 = 0.35 - box(2.0, y6, 2.4, 0.55, "Calibrated\n1D Profiles", c_output, 9, True) - box(5.0, y6, 2.2, 0.55, "Batch\nReports", c_output, 9, True) - box(8.0, y6, 2.4, 0.55, "K-factor\nHistory Log", c_output, 9, True) - arrow(5.0, y5 - 0.33, 2.0, y6 + 0.28) - arrow(5.0, y5 - 0.33, 5.0, y6 + 0.28) - arrow(5.0, y5 - 0.33, 8.0, y6 + 0.28) - - # ── Legend ── - for lx, lc, lt in [(0.3, c_input, "Input"), (1.5, c_core, "Processing"), - (2.9, c_output, "Output"), (4.1, c_ref, "Reference")]: - bx = FancyBboxPatch((lx, 6.65), 0.7, 0.25, boxstyle="round,pad=0.05", - facecolor=lc, edgecolor=c_border, linewidth=0.8) - ax.add_patch(bx) - ax.text(lx + 0.75 + 0.08, 6.775, lt, fontsize=7.5, va="center", color="#424242") - - fig.savefig(OUT_DIR / "fig_workflow.png", dpi=300, bbox_inches="tight", - facecolor="white", pad_inches=0.15) + + headings = [ + (1.50, "Inputs", "input"), + (4.75, "Interfaces", "interface"), + (8.25, "Scientific core", "core"), + (12.30, "Outputs", "output"), + ] + for x, label, color_key in headings: + ax.text(x, 7.45, label, ha="center", va="center", fontsize=7, fontweight="bold") + ax.plot([x - 0.82, x + 0.82], [7.16, 7.16], color=COLORS[color_key], linewidth=4) + + inputs = [ + (6.05, "2D detector frames\nand instrument headers"), + (4.45, "External 1D profiles\nand correction metadata"), + (2.85, "SRM 3600, water, or\na user-supplied reference"), + ] + for y, label in inputs: + _box(ax, 1.50, y, 2.45, 0.88, label, COLORS["input"], fontsize=6.2) + + _box( + ax, + 4.75, + 5.55, + 2.35, + 1.22, + "CLI utilities and strict\nBL19B2 batch workflow", + COLORS["interface"], + fontsize=6.2, + ) + _box( + ax, + 4.75, + 3.45, + 2.35, + 1.22, + "SAXSAbs Workbench\ncalibration and export", + COLORS["interface"], + fontsize=6.2, + ) + _box( + ax, + 4.75, + 1.55, + 2.35, + 1.0, + "Python API\npublic package functions", + COLORS["interface"], + fontsize=6.2, + ) + + core = [ + (6.35, "Parse, normalize, and\nvalidate inputs"), + (5.15, "Reduce detector data and\nintegrate with pyFAI"), + (3.95, "Estimate robust K and\npropagate uncertainty"), + (2.75, "Check intensity state and\ncorrection history"), + (1.55, "Write canSAS, NXcanSAS,\nand calibrated 2D outputs"), + ] + for y, label in core: + _box(ax, 8.25, y, 3.05, 0.76, label, COLORS["core"], fontsize=6.0) + + outputs = [ + (5.85, "Absolute-scale 1D profiles\nand calibrated 2D packages"), + (4.05, "canSAS XML and\nNXcanSAS HDF5"), + (2.25, "Calibration records, source\nhashes, QC, and run reports"), + ] + for y, label in outputs: + _box(ax, 12.30, y, 2.65, 1.0, label, COLORS["output"], fontsize=6.1) + + _arrow(ax, (2.78, 4.45), (3.52, 4.45), color=COLORS["accent"]) + _arrow(ax, (5.97, 4.45), (6.67, 4.45), color=COLORS["accent"]) + _arrow(ax, (9.80, 4.45), (10.85, 4.45), color=COLORS["accent"]) + + gate = FancyBboxPatch( + (0.18, 0.12), + 13.64, + 0.62, + boxstyle="round,pad=0.04,rounding_size=0.12", + facecolor=COLORS["gate"], + edgecolor=COLORS["warning"], + linewidth=0.9, + ) + ax.add_patch(gate) + ax.text( + 7, + 0.43, + "Required checks before calibrated output: units · transmission · thickness · " + "correction state · calibration provenance", + ha="center", + va="center", + fontsize=6.2, + color=COLORS["ink"], + fontweight="bold", + ) + save_figure(fig, output_dir, "fig_workflow") plt.close(fig) - print(" ✓ fig_workflow.png") - - -# ══════════════════════════════════════════════════════════════════════ -# Figure 2 — K-factor Estimation Demonstration -# ══════════════════════════════════════════════════════════════════════ - -# NIST SRM 3600 Certificate Table 1. Keep the paper figure tied to the -# package's reviewer-tested source of truth instead of a sparse approximation. -Q_REF = NIST_SRM3600_DATA[:, 0] -I_REF = NIST_SRM3600_DATA[:, 1] - - -def make_kfactor_figure(): - np.random.seed(42) - - # ── Simulate a measured profile (K_true ≈ 0.035, with noise + 2 outliers) ── - K_TRUE = 0.035 - q_dense = np.linspace(0.006, 0.260, 200) - i_ref_dense = np.interp(q_dense, Q_REF, I_REF) - noise = 1 + np.random.normal(0, 0.03, size=q_dense.shape) - i_meas_dense = i_ref_dense / K_TRUE * noise - - # Interpolate measured onto reference grid - i_meas_at_ref = np.interp(Q_REF, q_dense, i_meas_dense) - ratios = I_REF / i_meas_at_ref - - # Inject two artificial outliers - ratios_with_outliers = ratios.copy() - ratios_with_outliers[1] = K_TRUE * 3.5 # outlier high - ratios_with_outliers[13] = K_TRUE * 0.3 # outlier low - - # MAD outlier rejection - r_med = np.median(ratios_with_outliers) - r_mad = np.median(np.abs(ratios_with_outliers - r_med)) - robust_sigma = 1.4826 * r_mad - inlier_mask = np.abs(ratios_with_outliers - r_med) <= 3.0 * robust_sigma - k_robust = np.median(ratios_with_outliers[inlier_mask]) - - # ── Create figure ── - fig, axes = plt.subplots(1, 2, figsize=(11, 4.2), gridspec_kw={"wspace": 0.35}) - - # --- Panel (a): I(q) curves --- - ax1 = axes[0] - ax1.semilogy(Q_REF, I_REF, "s-", color="#D32F2F", markersize=6, - linewidth=1.5, label="NIST SRM 3600 reference", zorder=3) - ax1.semilogy(q_dense, i_meas_dense * K_TRUE, "-", color="#1976D2", - linewidth=1.2, alpha=0.7, label="Measured (rescaled by K)") - ax1.set_xlabel(r"$q$ (Å$^{-1}$)", fontsize=11) - ax1.set_ylabel(r"$I(q)$ (cm$^{-1}$ sr$^{-1}$)", fontsize=11) - ax1.set_title("(a) Scattering profiles", fontsize=11, fontweight="bold") - ax1.legend(fontsize=8.5, loc="upper right", framealpha=0.9) - ax1.set_xlim(0, 0.27) - ax1.tick_params(labelsize=9) - ax1.grid(True, alpha=0.3, linewidth=0.5) - - # --- Panel (b): Ratio plot with MAD filtering --- - ax2 = axes[1] - # Inliers - ax2.plot(Q_REF[inlier_mask], ratios_with_outliers[inlier_mask], "o", - color="#2E7D32", markersize=7, label="Inlier ratios", zorder=3) - # Outliers - ax2.plot(Q_REF[~inlier_mask], ratios_with_outliers[~inlier_mask], "x", - color="#D32F2F", markersize=9, markeredgewidth=2, - label="Rejected outliers", zorder=3) - # Median line - ax2.axhline(k_robust, color="#1976D2", linewidth=1.5, linestyle="-", - label=f"K = {k_robust:.4f}", zorder=2) - # ±3σ band - lower = r_med - 3 * robust_sigma - upper = r_med + 3 * robust_sigma - ax2.axhspan(lower, upper, alpha=0.12, color="#1976D2", - label=f"±3σ̂ band (σ̂ = {robust_sigma:.5f})") - ax2.axhline(lower, color="#1976D2", linewidth=0.7, linestyle="--", alpha=0.5) - ax2.axhline(upper, color="#1976D2", linewidth=0.7, linestyle="--", alpha=0.5) - - ax2.set_xlabel(r"$q$ (Å$^{-1}$)", fontsize=11) - ax2.set_ylabel(r"$R_i = I_{\mathrm{ref}} / I_{\mathrm{meas}}$", fontsize=11) - ax2.set_title("(b) Robust K-factor estimation", fontsize=11, fontweight="bold") - ax2.legend(fontsize=8, loc="upper right", framealpha=0.9) - ax2.set_xlim(0, 0.27) - ax2.tick_params(labelsize=9) - ax2.grid(True, alpha=0.3, linewidth=0.5) - - fig.savefig(OUT_DIR / "fig_kfactor_demo.png", dpi=300, bbox_inches="tight", - facecolor="white", pad_inches=0.1) + + +def make_kfactor_figure(output_dir: Path) -> None: + """Illustrate the robust K estimator with deterministic synthetic data.""" + + rng = np.random.default_rng(42) + q_ref = NIST_SRM3600_DATA[:, 0] + i_ref = NIST_SRM3600_DATA[:, 1] + k_true = 0.035 + + q_dense = np.sort( + np.unique(np.concatenate([np.linspace(float(q_ref.min()), float(q_ref.max()), 240), q_ref])) + ) + i_ref_dense = np.interp(q_dense, q_ref, i_ref) + i_meas_dense = i_ref_dense / k_true * (1 + rng.normal(0, 0.025, q_dense.size)) + outlier_reference_indices = np.array([1, 13]) + outlier_dense_indices = np.searchsorted(q_dense, q_ref[outlier_reference_indices]) + outlier_ratios = np.array([k_true * 3.5, k_true * 0.3]) + i_meas_dense[outlier_dense_indices] = i_ref[outlier_reference_indices] / outlier_ratios + if np.any(i_ref_dense <= 0.0) or np.any(i_meas_dense <= 0.0): + raise ValueError("log-scale demonstration requires strictly positive intensities") + i_meas_at_ref = np.interp(q_ref, q_dense, i_meas_dense) + ratios = i_ref / i_meas_at_ref + estimate = estimate_k_factor_robust( + q_dense, + i_meas_dense, + q_ref, + i_ref, + q_window=(float(q_ref.min()), float(q_ref.max())), + ) + inliers = np.array( + [ + np.any(np.isclose(ratio, estimate.ratios_used, rtol=1e-12, atol=1e-15)) + for ratio in ratios + ] + ) + + fig, axes = plt.subplots(1, 2, figsize=(FULL_WIDTH_IN, 2.75)) + fig.subplots_adjust(left=0.08, right=0.99, bottom=0.20, top=0.87, wspace=0.30) + + left, right = axes + left.semilogy( + q_ref, + i_ref, + "s-", + color=COLORS["warning"], + markersize=3.2, + linewidth=1.0, + ) + left.semilogy( + q_dense, + i_meas_dense * k_true, + linestyle="none", + marker=".", + markersize=2.0, + color=COLORS["accent"], + alpha=0.85, + ) + left.set(xlabel=r"$q$ ($\mathrm{\AA}^{-1}$)", ylabel=r"$I(q)$ ($\mathrm{cm}^{-1}$)") + left.set_title("Reference and rescaled synthetic profile", loc="left", fontweight="bold") + left.text(-0.12, 1.04, "a", transform=left.transAxes, fontsize=8, fontweight="bold") + left.legend(["NIST SRM 3600", "Synthetic measured profile"], frameon=False, loc="upper right") + left.text( + 0.03, + 0.04, + "SYNTHETIC DEMO", + transform=left.transAxes, + fontsize=7, + color=COLORS["muted"], + bbox={"facecolor": "white", "edgecolor": COLORS["muted"], "pad": 2}, + ) + + right.scatter(q_ref[inliers], ratios[inliers], s=16, color=COLORS["accent"], label="Inlier") + right.scatter( + q_ref[~inliers], + ratios[~inliers], + s=30, + marker="x", + linewidth=1.4, + color=COLORS["warning"], + label="Rejected", + ) + right.axhline( + estimate.k_factor, + color=COLORS["ink"], + linewidth=1.0, + label=f"K = {estimate.k_factor:.4f}", + ) + right.set(xlabel=r"$q$ ($\mathrm{\AA}^{-1}$)", ylabel=r"$I_{ref}/I_{meas}$") + right.set_title("Robust K estimate", loc="left", fontweight="bold") + right.text(-0.12, 1.04, "b", transform=right.transAxes, fontsize=8, fontweight="bold") + right.legend(frameon=False, ncol=2, loc="upper right") + + save_figure(fig, output_dir, "fig_kfactor_demo") plt.close(fig) - print(" ✓ fig_kfactor_demo.png") -# ══════════════════════════════════════════════════════════════════════ -# Run all -# ══════════════════════════════════════════════════════════════════════ +def main() -> None: + parser = argparse.ArgumentParser() + parser.add_argument( + "--output-dir", + type=Path, + default=Path(__file__).resolve().parent, + ) + parser.add_argument( + "--demo", + action="store_true", + help="also render the explicitly synthetic K-factor demonstration", + ) + args = parser.parse_args() + output_dir = args.output_dir.resolve() + make_workflow_figure(output_dir) + if args.demo: + make_kfactor_figure(output_dir) + print(f"Wrote JOSS figures to {output_dir}") + if __name__ == "__main__": - print("Generating JOSS paper figures ...") - make_workflow_figure() - make_kfactor_figure() - print("Done.") + main() diff --git a/paper/jac_calibrated2d_draft.md b/paper/jac_calibrated2d_draft.md deleted file mode 100644 index f470fa1..0000000 --- a/paper/jac_calibrated2d_draft.md +++ /dev/null @@ -1,402 +0,0 @@ ---- -title: "A reproducible calibrated-detector-image package for synchrotron SAXS absolute-intensity workflows" -authors: - - name: Delun Gong - orcid: 0000-0001-7877-7707 - affiliation: 1 -affiliations: - - name: Institute of Metal Research, Chinese Academy of Sciences, Shenyang 110016, China - index: 1 -bibliography: jac_calibrated2d_refs.bib ---- - - - -# Synopsis - -SAXSAbs defines a reproducible calibrated-detector-image package for synchrotron -small-angle X-ray scattering absolute-intensity workflows. The package preserves -an absolute-scale detector image together with the mask, pyFAI geometry, correction -policy, provenance records and a reintegration recipe, allowing third parties to -recalculate calibrated one-dimensional profiles without access to beamline-private -bookkeeping. - -# Abstract - -Absolute-intensity calibration is essential when small-angle X-ray scattering -(SAXS) profiles are compared between instruments, experiments or quantitative -models, yet the intermediate data required to audit this calibration are often -lost after detector images have been reduced to one-dimensional curves. We present -SAXSAbs, an open-source Python package that defines and implements a reproducible -calibrated-detector-image package for synchrotron SAXS workflows. The package -contains a detector-space net image on an absolute intensity scale, the mask and -pyFAI PONI geometry used for integration, a machine-readable correction policy, -processing provenance, checksums and an executable reintegration recipe. SAXSAbs -builds on pyFAI for azimuthal integration and focuses on the calibration-control -layer: robust K-factor estimation against standard references, monitor-mode-aware -normalization, transmission and thickness handling, optional flat-field correction, -buffer subtraction and standards-oriented export. The workflow is designed to -complement existing diffraction and scattering platforms by making the calibrated -intermediate state portable and independently verifiable. In deterministic -examples and synchrotron beamline tests, the exported package can be reintegrated -with pyFAI to reproduce the reported absolute one-dimensional profile within -TODO tolerance, while retaining enough metadata to diagnose common failure modes -such as geometry mismatch, invalid transmission, inconsistent masks and unstable -K-factor estimates. This approach turns a beamline-specific calibration procedure -into an auditable data object that can be shared with collaborators, reviewers -and downstream analysis tools. - -Keywords: SAXS; absolute intensity calibration; synchrotron data reduction; -pyFAI; reproducibility; metadata provenance; calibrated detector image. - -# 1. Introduction - -Two-dimensional area detectors have made synchrotron SAXS and combined -SAXS/WAXS experiments efficient enough that data reduction is now frequently a -high-throughput workflow rather than a single-image operation. Modern beamlines -therefore rely on software that can read detector images, apply geometric -calibration, handle masks and detector corrections, integrate images and process -large image series. Mature tools already exist for many of these tasks. pyFAI -provides a fast and widely used azimuthal-integration library with explicit -geometry handling through PONI files [@pyfai]. pydidas provides an accessible -graphical and scriptable workflow environment for X-ray diffraction data -processing [@pydidas]. StreamSAXS targets both offline and streaming SAXS/WAXS -workflows at synchrotron facilities [@streamsaxs]. Earlier and established tools -such as FIT2D [@fit2d], Nika [@nika], DPDAK [@dpdak], Dioptas [@dioptas], -BioXTAS RAW [@bioxtasraw], Irena [@irena] and GSAS-II [@gsasii] further -demonstrate that image reduction, integration, exploration and downstream -analysis are already well represented in the scattering and diffraction software -ecosystem. - -The remaining reproducibility problem addressed here is narrower but common in -practice. Once a SAXS detector image has been corrected, normalized, background -subtracted, calibrated against a standard and integrated to a one-dimensional -profile, the calibrated intermediate image and the precise correction policy are -often not preserved in a form that another researcher can reuse. Collaborators -may receive only a final text file containing \(q\), \(I(q)\) and an uncertainty -column. Reviewers may be unable to determine whether the reported intensity scale -depends on the chosen transmission, thickness, monitor semantics, mask -convention, flat-field correction or PONI geometry. Beamline staff may retain the -knowledge required to regenerate the profile, but that knowledge is frequently -encoded in local scripts, graphical-session state or manual spreadsheets rather -than in a portable data product. - -This gap is especially important for absolute-intensity SAXS. Quantitative SAXS -interpretation depends on bringing measured scattering intensities onto an -absolute differential cross-section scale. Calibration against reference -standards such as NIST SRM 3600 glassy carbon [@srm3600] or liquid water -[@orthaber2000] requires normalization by beam-monitor quantities, transmission -and sample thickness, followed by estimation of a scale factor that can be -sensitive to parasitic scattering, beamstop shadows and detector artefacts. A -final one-dimensional profile is therefore not a complete record of the -calibration decision. To make absolute-intensity data reproducible, the -detector-space state after correction and calibration should be shareable -together with the geometry, mask, standards, formulas and software versions used -to create it. - -SAXSAbs was developed to formalize this calibration-control layer. The software -does not attempt to replace pyFAI, pydidas or general-purpose SAXS/WAXS workflow -platforms. Instead, it defines a calibrated-detector-image package that can be -produced by a SAXS absolute-calibration workflow and reintegrated by standard -tools. The package contains the absolute-scale detector image, the mask, the -PONI geometry, a correction-policy record, provenance metadata, checksums and a -reintegration recipe. The design goal is that a third party can inspect the -package, verify file integrity, rerun azimuthal integration and recover the -reported absolute one-dimensional curve without reconstructing private beamline -bookkeeping. - -# 2. Software overview - -SAXSAbs is an open-source Python package with a modular numerical core, a command -line interface and a bilingual graphical workbench. The existing core implements -monitor normalization, robust K-factor estimation, a standard-reference registry, -a composition-based attenuation coefficient calculator using xraydb [@xraydb], -buffer subtraction with uncertainty propagation, heterogeneous header parsing and -canSAS/NXcanSAS one-dimensional export. The graphical workflow uses pyFAI and -FabIO-compatible detector I/O for practical two-dimensional image processing, -while the installable package exposes the calibration logic for headless testing -and scripted operation. - -The calibrated-detector-image package extends this software boundary from final -curve export to reproducible intermediate export. The central object is an image -array in detector coordinates after dark-current subtraction, monitor -normalization, background subtraction, optional flat-field correction and -absolute K-factor scaling have been applied according to a recorded policy. The -image is kept in detector space rather than transformed into reciprocal-space -coordinates because detector space preserves the exact pixel mask and pyFAI -geometry relation needed for independent reintegration. The package therefore -acts as a bridge between raw beamline images that may be too large or private to -share and final one-dimensional curves that are too reduced to audit. - -# 3. Package definition - -The package is organized as a small directory tree with stable file roles. A -typical package contains an image file, a mask file, a PONI geometry file, a -metadata file, a manifest and a reintegration script or recipe. The image stores -the calibrated detector-space intensity. The mask records invalid pixels using a -declared convention compatible with pyFAI. The PONI file preserves detector -distance, beam centre, rotations, pixel sizes and wavelength. The metadata file -records sample identifiers, raw source paths or anonymized source names, -correction settings, absolute-calibration parameters, software versions and -recommended reintegration arguments. The manifest records all files in the -package and their checksums so that corruption or accidental replacement can be -detected. - -The correction policy is deliberately explicit. It records whether dark-current -subtraction, background subtraction, flat-field correction, polarization -correction, solid-angle correction, transmission correction, thickness -normalization and K-factor scaling were applied in the image or should be applied -during reintegration. This distinction is essential because a calibrated image -can otherwise be double-corrected or under-corrected when it is passed to another -program. For example, if the detector-space image has already been scaled to -absolute units and divided by thickness, the reintegration recipe sets the -normalization factor to unity and instructs pyFAI not to reapply the same -normalization. Conversely, corrections that remain integration dependent, such -as solid-angle correction, are recorded in the reintegration recipe rather than -silently embedded in undocumented state. - -The metadata schema is versioned. Versioning allows the package to evolve while -remaining machine-readable, and it allows old packages to be validated against -the schema that created them. The initial schema records the package identifier, -sample identifier, image type, intensity unit, mask convention, geometry file, -correction policy, reference standard, K-factor estimate, K-factor uncertainty, -normalization mode, monitor values, transmission, thickness, q range, azimuthal -range, software versions and checksum manifest. The schema is intentionally -small enough to be reviewed in a text editor but explicit enough for automated -validation and reintegration. - -# 4. Calibration and reintegration workflow - -The workflow starts from a standard sample, one or more background images, a dark -image, a pyFAI PONI geometry file and optional mask and flat-field images. The -standard and background images are normalized by the selected monitor convention. -In rate mode the normalization factor is -\[ -N = t_{\mathrm{exp}} I_0 T, -\] -where \(t_{\mathrm{exp}}\) is the exposure time, \(I_0\) is the beam-monitor -quantity and \(T\) is the transmission. In integrated mode the exposure-time -factor is omitted because the monitor value already represents integrated counts. -The detector-space net image is calculated by subtracting the normalized dark -and background contributions from the normalized sample signal. - -The net standard image is integrated with pyFAI to obtain an observed profile. -SAXSAbs estimates the absolute K factor by comparing this profile with a -registered reference standard. For NIST SRM 3600 glassy carbon, the reference -profile is interpolated over a specified q range and point-wise ratios between -reference and measured intensities are calculated. The final K factor is the -median of the accepted ratios after median absolute deviation filtering. This -robust estimator reduces sensitivity to isolated detector artefacts or q points -affected by parasitic scattering. For liquid water, the software uses a -temperature-dependent flat reference value and applies the same robust dispersion -logic to the selected q window. - -After the K factor has been determined, the detector-space image for each sample -is scaled to absolute intensity according to the recorded thickness and -normalization policy. The calibrated image, mask, PONI file and metadata are -written into the package. The reintegration recipe then calls pyFAI with the -stored PONI file and mask and with normalization choices that reflect the -correction policy. The expected output is a one-dimensional profile on the same -absolute scale as the original SAXSAbs batch result. This reintegration step is -the central reproducibility test: a package is considered valid only if the -reported profile can be regenerated from the exported intermediate state within -a declared numerical tolerance. - -# 5. Implementation - -The package implementation is designed around pure functions where possible. -Image scaling, sample identifier generation, metadata construction, manifest -creation and checksum calculation are separated from graphical user-interface -state. This separation allows deterministic tests to verify individual -behaviours, such as flat-field application, mask conversion, PONI copying and -metadata generation. The intended public API consists of a configuration object -describing the calibrated image export and a writer function that creates the -package directory and returns the paths and manifest row for downstream logging. - -The graphical workbench uses this API after two-dimensional background -subtraction and absolute scaling have been completed. The same API can also be -called from command-line workflows, which is important for beamline automation -and continuous-integration tests. The package writer does not require raw -beamline files to be published; it records source-file identities and provenance -while allowing users to anonymize paths or provide relative identifiers when -sharing data externally. This design balances reproducibility with the practical -constraints of facility data policies and proprietary measurements. - -SAXSAbs writes standard one-dimensional outputs in CSV, TSV, canSAS XML and -NXcanSAS HDF5, but the calibrated-detector-image package is distinct from these -final products. The one-dimensional formats are suitable for modelling and -downstream analysis, whereas the package is intended for audit, reintegration -and method review. In a typical publication workflow, authors can deposit the -small package for representative samples alongside final profiles, allowing -reviewers to confirm that the reported absolute scale can be reconstructed. - -# 6. Demonstration and validation plan - -The first validation example is a deterministic synthetic detector dataset -included in the repository. The example contains a small detector image, a -minimal geometry description and a scripted workflow that produces an absolute -profile and standard-format outputs. In the calibrated-package validation, this -example will be extended so that the package writer exports the detector-space -absolute image, mask, PONI file, metadata and manifest. The reintegration recipe -will then regenerate the one-dimensional profile and compare it with the profile -produced during the original workflow. The expected acceptance criterion is a -relative intensity difference below TODO value across the valid q range and a -K-factor within TODO range. - -The second validation example should use an anonymized synchrotron beamline -dataset measured with a standard sample and a representative unknown sample. -This dataset should demonstrate the practical value of the package under -realistic conditions, including non-trivial masks, beamstop shadows, background -subtraction and real metadata heterogeneity. The manuscript should report the -number of images, detector format, q range, beam energy, standard type, exposure -conditions, K-factor dispersion, number of rejected q points, package size and -reintegration agreement. If the raw data cannot be released, the calibrated -package and a reduced anonymized subset should still be deposited so that the -central reproducibility claim can be checked. - -The third validation example should evaluate failure detection. Deliberate -perturbations of the package can be used to test whether schema validation and -quality-control checks detect an incompatible mask shape, a missing PONI file, a -changed checksum, invalid transmission, inconsistent wavelength or a correction -policy that would double-apply normalization. These tests convert common -beamline workflow mistakes into documented software behaviours rather than -unobserved failure modes. - -# 7. Results to report before submission - -The final manuscript should report quantitative results rather than only -software capabilities. The most important result is reintegration agreement: -the one-dimensional absolute profile regenerated from the calibrated package -should match the original SAXSAbs output within a clearly stated tolerance. -Agreement should be shown for the synthetic example and at least one real -beamline example. The second result is audit completeness: every reported curve -should be traceable to a package identifier, checksum manifest, PONI file, mask, -correction policy, standard reference and K-factor estimate. The third result is -quality-control behaviour: invalid or inconsistent packages should fail with -specific diagnostics instead of producing silent numerical output. - -For a JAC submission, the figures should carry most of the evidence. Figure 1 -should present the conceptual workflow from raw detector images to the -calibrated package and then to independently regenerated one-dimensional -profiles. Figure 2 should show the package structure and metadata schema, -including the correction-policy fields. Figure 3 should compare the original -and reintegrated absolute profiles for the synthetic and beamline examples. -Figure 4 should summarize quality-control diagnostics and failure cases. A table -should compare SAXSAbs with pyFAI, pydidas, StreamSAXS, Nika, DPDAK, Dioptas and -BioXTAS RAW, but it should avoid framing those tools as inadequate. The correct -message is that those tools solve integration, workflow and analysis problems, -whereas SAXSAbs contributes a portable absolute-calibration audit package. - -# 8. Discussion - -The main contribution of SAXSAbs is not a new azimuthal-integration algorithm -and not a general-purpose graphical workflow platform. Those areas are already -well served by established software. The contribution is the definition and -implementation of a shareable calibrated intermediate state for absolute SAXS -workflows. This distinction is important for novelty because it positions the -software as a complement to the existing ecosystem. pyFAI remains the integration -engine, pydidas and StreamSAXS remain broader workflow platforms, and SAXSAbs -provides a calibrated data object that makes the absolute-intensity step -auditable. - -The calibrated-detector-image package also changes how SAXS data can be reviewed -and reused. A final one-dimensional profile is compact and convenient, but it is -not sufficient to diagnose many reduction choices. Raw detector images are -complete, but they may be large, facility-specific or restricted by data policy. -The proposed package occupies an intermediate position. It is smaller and easier -to share than a complete raw experiment, but it retains the detector-space -information, geometry and correction policy needed to regenerate the reported -absolute profile. This is particularly useful when a paper reports quantitative -intensities, compares measurements across beamlines or uses absolute scale as an -input to modelling. - -There are limitations. The package does not remove the need for correct -experimental calibration, and it cannot rescue measurements with poor standards, -incorrect transmissions or unsuitable backgrounds. Its reproducibility claim is -bounded by the recorded correction policy and by the behaviour of the integration -engine used for reintegration. Different versions of pyFAI may differ at small -numerical levels, especially when integration methods, error models or -correction flags change. For this reason, the package records software versions -and recommended integration parameters, and the validation should specify -acceptable tolerances rather than requiring bitwise identity. - -Future development will focus on strengthening interoperability with downstream -workflow tools and on extending the validation set. A pydidas-compatible import -example would allow users to treat the calibrated image package as an upstream -data product. Additional real-time or time-resolved examples would test whether -the same package concept scales to in situ SAXS/WAXS series. A community schema, -if adopted beyond this project, could provide a lightweight convention for -publishing auditable SAXS absolute-calibration intermediates alongside final -profiles. - -# 9. Conclusions - -SAXSAbs provides a reproducible calibrated-detector-image package for -synchrotron SAXS absolute-intensity workflows. By preserving the absolute-scale -detector image together with mask, PONI geometry, correction policy, provenance, -checksums and reintegration instructions, the package makes the calibrated -intermediate state independently verifiable. The software is designed to -complement established integration and workflow tools rather than replace them. -After completion of the package exporter, schema validation and beamline -benchmark, this positioning should provide a stronger and more defensible basis -for a Journal of Applied Crystallography software submission than a conventional -GUI or batch-processing claim. - -# Figure and table plan - -**Figure 1. Workflow overview.** Raw standard, sample, background and dark images -enter the SAXSAbs calibration-control workflow. The output is both a final -absolute one-dimensional profile and a calibrated-detector-image package that -can be reintegrated independently. - -**Figure 2. Package anatomy.** Directory tree and schema fields for the image, -mask, PONI geometry, metadata, correction policy, checksum manifest and -reintegration recipe. - -**Figure 3. Reproducibility test.** Overlay of original SAXSAbs output and -profile regenerated from the calibrated package for synthetic and beamline -examples, with residuals shown below. - -**Figure 4. Quality-control behaviour.** Examples of validation failures caused -by changed checksums, incompatible mask shape, missing PONI file, invalid -transmission and inconsistent correction policy. - -**Table 1. Ecosystem comparison.** Comparison of pyFAI, pydidas, StreamSAXS, -Nika, DPDAK, Dioptas, BioXTAS RAW and SAXSAbs across integration, workflow, -absolute calibration, calibrated package export, provenance and reintegration -verification. - -# Data availability - -The source repository includes a deterministic synthetic example for reviewer -reproducibility. Before submission, the calibrated-detector-image package -generated from this example and at least one anonymized synchrotron beamline -example should be deposited in a public archive. TODO: add archive DOI, package -checksums and exact command lines. - -# Code availability - -SAXSAbs is available from GitHub at -https://github.com/D-sudoasd/SASAbs_saxs-absolute-calibration and archived at -Zenodo with DOI https://doi.org/10.5281/zenodo.19687104. TODO: update this -statement to the exact release DOI used for the JAC submission. - -# Conflict of interest - -The author declares no competing interests. - -# Acknowledgements - -The author thanks beamline users and scientists at the Institute of Metal -Research, Chinese Academy of Sciences, for practical feedback on SAXS -absolute-intensity calibration workflows and metadata heterogeneity. TODO: add -beamline/facility acknowledgements required by the real validation dataset. - diff --git a/paper/jac_calibrated2d_refs.bib b/paper/jac_calibrated2d_refs.bib deleted file mode 100644 index 2284dd6..0000000 --- a/paper/jac_calibrated2d_refs.bib +++ /dev/null @@ -1,137 +0,0 @@ -@article{pyfai, - author = {Ashiotis, G. and Deschildre, A. and Nawaz, Z. and Wright, J. P. and Karkoulis, D. and Picca, F. E. and Kieffer, J.}, - title = {The fast azimuthal integration {Python} library: {pyFAI}}, - journal = {Journal of Applied Crystallography}, - volume = {48}, - number = {2}, - pages = {510--519}, - year = {2015}, - doi = {10.1107/S1600576715004306} -} - -@article{pydidas, - author = {Storm, Malte and Staron, Peter and Krywka, Christina}, - title = {{Pydidas}: a tool for automated {X}-ray diffraction data analysis}, - journal = {Journal of Applied Crystallography}, - volume = {58}, - number = {4}, - pages = {1476--1485}, - year = {2025}, - doi = {10.1107/S160057672500398X} -} - -@article{streamsaxs, - author = {Wang, Jiayi and Dong, Zheng and Zhang, Yi and Hua, Wenqiang and Wang, Zudeng and Guo, Huilong and Yang, Yiming and Bi, Xiaoxue}, - title = {{StreamSAXS}: a {Python}-based workflow platform for processing streaming {SAXS}/{WAXS} data}, - journal = {Journal of Synchrotron Radiation}, - volume = {31}, - number = {5}, - pages = {1249--1256}, - year = {2024}, - doi = {10.1107/S1600577524005149} -} - -@article{gsasii, - author = {Toby, Brian H. and Von Dreele, Robert B.}, - title = {{GSAS-II}: the genesis of a modern open-source all purpose crystallography software package}, - journal = {Journal of Applied Crystallography}, - volume = {46}, - number = {2}, - pages = {544--549}, - year = {2013}, - doi = {10.1107/S0021889813003531} -} - -@article{fit2d, - author = {Hammersley, A. P.}, - title = {{FIT2D}: a multi-purpose data reduction, analysis and visualization program}, - journal = {Journal of Applied Crystallography}, - volume = {49}, - number = {2}, - pages = {646--652}, - year = {2016}, - doi = {10.1107/S1600576716000455} -} - -@article{dioptas, - author = {Prescher, C. and Prakapenka, V. B.}, - title = {{DIOPTAS}: a program for reduction of two-dimensional {X}-ray diffraction data and data exploration}, - journal = {High Pressure Research}, - volume = {35}, - number = {3}, - pages = {223--230}, - year = {2015}, - doi = {10.1080/08957959.2015.1059835} -} - -@article{nika, - author = {Ilavsky, J.}, - title = {{Nika}: software for two-dimensional data reduction}, - journal = {Journal of Applied Crystallography}, - volume = {45}, - number = {2}, - pages = {324--328}, - year = {2012}, - doi = {10.1107/S0021889812004037} -} - -@article{dpdak, - author = {Benecke, G. and Wagermaier, W. and Li, C. and Schwartzkopf, M. and Flucke, G. and Hoerth, R. and Zizak, I. and Burghammer, M. and Metwalli, E. and Mueller-Buschbaum, P. and Trebbin, M. and Foerster, S. and Paris, O. and Roth, S. V. and Fratzl, P.}, - title = {A customizable software for fast reduction and analysis of large {X}-ray scattering data sets: applications of the new {DPDAK} package to small-angle {X}-ray scattering and grazing-incidence small-angle {X}-ray scattering}, - journal = {Journal of Applied Crystallography}, - volume = {47}, - number = {5}, - pages = {1797--1803}, - year = {2014}, - doi = {10.1107/S1600576714019773} -} - -@article{bioxtasraw, - author = {Hopkins, J. B. and Gillilan, R. E. and Skou, S.}, - title = {{BioXTAS RAW}: improvements to a free open-source program for small-angle {X}-ray scattering data reduction and analysis}, - journal = {Journal of Applied Crystallography}, - volume = {50}, - number = {5}, - pages = {1545--1553}, - year = {2017}, - doi = {10.1107/S1600576717011438} -} - -@article{irena, - author = {Ilavsky, J. and Jemian, P. R.}, - title = {{Irena}: tool suite for modeling and analysis of small-angle scattering}, - journal = {Journal of Applied Crystallography}, - volume = {42}, - number = {2}, - pages = {347--353}, - year = {2009}, - doi = {10.1107/S0021889809002222} -} - -@techreport{srm3600, - author = {{National Institute of Standards and Technology}}, - title = {Standard Reference Material 3600: Absolute Intensity Calibration Standard for Small-Angle {X}-ray Scattering}, - institution = {National Institute of Standards and Technology}, - year = {2016}, - url = {https://tsapps.nist.gov/srmext/certificates/3600.pdf} -} - -@article{orthaber2000, - author = {Orthaber, D. and Bergmann, A. and Glatter, O.}, - title = {{SAXS} experiments on absolute scale with {Kratky} systems using water as a secondary standard}, - journal = {Journal of Applied Crystallography}, - volume = {33}, - number = {2}, - pages = {218--225}, - year = {2000}, - doi = {10.1107/S0021889899015216} -} - -@misc{xraydb, - author = {Newville, M.}, - title = {xraydb: {X}-ray Reference Data in {SQLite}}, - year = {2023}, - publisher = {Zenodo}, - doi = {10.5281/zenodo.7847236}, - url = {https://github.com/xraypy/XrayDB} -} diff --git a/paper/paper.bib b/paper/paper.bib index 702872a..af8ac40 100644 --- a/paper/paper.bib +++ b/paper/paper.bib @@ -1,24 +1,35 @@ -@article{pyfai, - author = {Ashiotis, G. and Deschildre, A. and Nawaz, Z. and Wright, J. P. and Karkoulis, D. and Picca, F. E. and Kieffer, J.}, - title = {The fast azimuthal integration {Python} library: {pyFAI}}, +@article{allen2017, + author = {Allen, Andrew J. and Zhang, Fan and Kline, R. Joseph and Guthrie, William F. and Ilavsky, Jan}, + title = {{NIST} Standard Reference Material 3600: Absolute Intensity Calibration Standard for Small-Angle {X}-ray Scattering}, journal = {Journal of Applied Crystallography}, - volume = {48}, + volume = {50}, number = {2}, - pages = {510--519}, - year = {2015}, - doi = {10.1107/S1600576715004306} + pages = {462--474}, + year = {2017}, + doi = {10.1107/S1600576717001972} } -@article{sasview, - author = {Doucet, M. and et al.}, - title = {{SasView} version 4.2}, - journal = {Zenodo}, - year = {2018}, - doi = {10.5281/zenodo.1412041} +@article{bioxtasraw, + author = {Hopkins, Jesse B. and Gillilan, Richard E. and Skou, Soren}, + title = {{BioXTAS RAW}: improvements to a free open-source program for small-angle {X}-ray scattering data reduction and analysis}, + journal = {Journal of Applied Crystallography}, + volume = {50}, + number = {5}, + pages = {1545--1553}, + year = {2017}, + doi = {10.1107/S1600576717011438} +} + +@misc{cansas1d, + author = {{canSAS Working Group}}, + title = {{canSAS1d} Data Formats: Version 1.1}, + year = {2013}, + url = {https://www.cansas.org/formats/canSAS1d/1.1/doc/specification.html}, + note = {Accessed 30 July 2026} } @article{dioptas, - author = {Prescher, C. and Prakapenka, V. B.}, + author = {Prescher, Clemens and Prakapenka, Vitali B.}, title = {{DIOPTAS}: a program for reduction of two-dimensional {X}-ray diffraction data and data exploration}, journal = {High Pressure Research}, volume = {35}, @@ -28,8 +39,30 @@ @article{dioptas doi = {10.1080/08957959.2015.1059835} } +@article{elam2002, + author = {Elam, William T. and Ravel, Bruce D. and Sieber, John R.}, + title = {A new atomic database for {X}-ray spectroscopic calculations}, + journal = {Radiation Physics and Chemistry}, + volume = {63}, + number = {2}, + pages = {121--128}, + year = {2002}, + doi = {10.1016/S0969-806X(01)00227-4} +} + +@article{fabio, + author = {Knudsen, Erik B. and S{ø}rensen, Henning O. and Wright, Jonathan P. and Goret, Ga{ë}l and Kieffer, Jérôme}, + title = {{FabIO}: easy access to two-dimensional {X}-ray detector images in {Python}}, + journal = {Journal of Applied Crystallography}, + volume = {46}, + number = {2}, + pages = {537--539}, + year = {2013}, + doi = {10.1107/S0021889813000150} +} + @article{irena, - author = {Ilavsky, J. and Jemian, P. R.}, + author = {Ilavsky, Jan and Jemian, Pete R.}, title = {Irena: tool suite for modeling and analysis of small-angle scattering}, journal = {Journal of Applied Crystallography}, volume = {42}, @@ -39,37 +72,28 @@ @article{irena doi = {10.1107/S0021889809002222} } -@article{bioxtasraw, - author = {Hopkins, J. B. and Gillilan, R. E. and Skou, S.}, - title = {{BioXTAS} {RAW}: improvements to a free open-source program for small-angle {X}-ray scattering data reduction and analysis}, - journal = {Journal of Applied Crystallography}, - volume = {50}, - number = {5}, - pages = {1545--1553}, - year = {2017}, - doi = {10.1107/S1600576717011438} +@misc{nist_srd126, + author = {Hubbell, J. H. and Seltzer, S. M.}, + title = {Tables of {X}-ray Mass Attenuation Coefficients and Mass Energy-Absorption Coefficients}, + publisher = {{National Institute of Standards and Technology}}, + year = {2004}, + version = {1.4}, + doi = {10.18434/T4D01F}, + url = {https://www.nist.gov/pml/x-ray-mass-attenuation-coefficients}, + note = {NIST Standard Reference Database 126; accessed 30 July 2026} } -@techreport{srm3600, - author = {{National Institute of Standards and Technology}}, - title = {Standard Reference Material 3600: Absolute Intensity Calibration Standard for Small-Angle {X}-ray Scattering}, - institution = {NIST}, - year = {2016}, - note = {Certificate of Analysis}, - url = {https://tsapps.nist.gov/srmext/certificates/3600.pdf} -} - -@book{glatter_kratky, - author = {Glatter, O. and Kratky, O.}, - title = {Small Angle {X}-ray Scattering}, - publisher = {Academic Press}, - year = {1982}, - isbn = {0-12-286280-5} +@misc{nxcansas, + author = {{NeXus International Advisory Committee}}, + title = {{NXcanSAS}: application definition for reduced small-angle scattering data}, + year = {2026}, + url = {https://manual.nexusformat.org/classes/applications/NXcanSAS.html}, + note = {NeXus 2026.01 documentation; accessed 30 July 2026} } @article{orthaber2000, - author = {Orthaber, D. and Bergmann, A. and Glatter, O.}, - title = {SAXS experiments on absolute scale with {Kratky} systems using water as a secondary standard}, + author = {Orthaber, Doris and Bergmann, Andreas and Glatter, Otto}, + title = {{SAXS} experiments on absolute scale with {Kratky} systems using water as a secondary standard}, journal = {Journal of Applied Crystallography}, volume = {33}, number = {2}, @@ -78,11 +102,58 @@ @article{orthaber2000 doi = {10.1107/S0021889899015216} } -@misc{newville_xraydb, - author = {Newville, M.}, - title = {xraydb: {X}-ray Reference Data in {SQLite}}, - year = {2023}, - publisher = {GitHub}, - url = {https://github.com/xraypy/XrayDB}, - doi = {10.5281/zenodo.7847236} +@article{pyfai, + author = {Ashiotis, Giannis and Deschildre, Aurore and Nawaz, Zubair and Wright, Jonathan P. and Karkoulis, Dimitrios and Picca, Frédéric Emmanuel and Kieffer, Jérôme}, + title = {The fast azimuthal integration {Python} library: {pyFAI}}, + journal = {Journal of Applied Crystallography}, + volume = {48}, + number = {2}, + pages = {510--519}, + year = {2015}, + doi = {10.1107/S1600576715004306} +} + +@misc{punx, + author = {Jemian, Pete R.}, + title = {{punx}: Python Utilities for {NeXus} {HDF5} Files}, + version = {0.3.5}, + year = {2024}, + url = {https://prjemian.github.io/punx/}, + note = {Software documentation; accessed 15 August 2026} +} + +@misc{sasview, + author = {Doucet, Mathieu and others}, + title = {{SasView} version 4.2}, + publisher = {Zenodo}, + version = {4.2.0}, + year = {2018}, + doi = {10.5281/zenodo.1412041} +} + +@misc{saxsabs_archive, + author = {Gong, Delun}, + title = {saxsabs: A Robust Workflow for Small-Angle {X}-ray Scattering Absolute Intensity Calibration}, + publisher = {Zenodo}, + year = {2026}, + doi = {10.5281/zenodo.19687103}, + note = {Concept record linking archived versions; the 2.0.0 candidate is not yet archived} +} + +@techreport{srm3600, + author = {{National Institute of Standards and Technology}}, + title = {Standard Reference Material 3600: Absolute Intensity Calibration Standard for Small-Angle {X}-ray Scattering}, + institution = {{National Institute of Standards and Technology}}, + year = {2016}, + type = {Certificate of Analysis}, + url = {https://tsapps.nist.gov/srmext/certificates/3600.pdf} +} + +@misc{xraydb, + author = {Newville, Matthew and easyXAFS and Whittington, Nathan and Levantino, Matteo and Schlepuetz, Christian and others}, + title = {xraypy/{XrayDB}: 4.5.8}, + publisher = {Zenodo}, + version = {4.5.8}, + year = {2025}, + doi = {10.5281/zenodo.16114067} } diff --git a/paper/paper.md b/paper/paper.md index bc214b3..20ae261 100644 --- a/paper/paper.md +++ b/paper/paper.md @@ -1,230 +1,181 @@ --- -title: 'saxsabs: A Robust Workflow for Small-Angle X-ray Scattering Absolute Intensity Calibration' +title: 'saxsabs: Absolute-intensity calibration and provenance tracking for small-angle X-ray scattering' tags: - Python - - SAXS + - small-angle X-ray scattering - absolute intensity calibration + - scientific software - synchrotron - - glassy carbon authors: - name: Delun Gong orcid: 0000-0001-7877-7707 - affiliation: 1 + affiliation: '1' affiliations: - - name: Institute of Metal Research, Chinese Academy of Sciences, Shenyang 110016, China - index: 1 -date: 26 February 2026 + - index: 1 + name: Institute of Metal Research, Chinese Academy of Sciences, Shenyang 110016, China +date: 26 August 2026 bibliography: paper.bib --- + + # Summary -`saxsabs` is an open-source Python package that provides a complete, reproducible -workflow for small-angle X-ray scattering (SAXS) absolute intensity calibration. -It automates the data-reduction chain from raw two-dimensional (2D) detector images -to calibrated one-dimensional (1D) scattering profiles on an absolute differential -cross-section scale (cm$^{-1}$ sr$^{-1}$). The software comprises a modular core -library, a command-line interface (CLI), and a graphical user interface (GUI) with -bilingual support (Chinese/English). - -Key capabilities include: - -- **Multi-standard calibration** using a pluggable registry that ships with NIST - Standard Reference Material 3600 (SRM 3600) glassy carbon [@srm3600] and - liquid water [@orthaber2000], with support for user-supplied reference data. -- **Robust K-factor estimation** via median absolute deviation (MAD) outlier - rejection. -- **Universal linear attenuation coefficient ($\mu$) calculator** driven by the - XCOM photon cross-section database through xraydb [@newville_xraydb], accepting - arbitrary chemical compositions and photon energies. -- **Buffer / solvent subtraction** with tuneable scaling factor $\alpha$ and - full error propagation, addressing protein solution SAXS (BioSAXS) workflows. -- **Multi-format output** in TSV, CSV, canSAS 1D XML, and NXcanSAS HDF5, - promoting interoperability with the NeXus/canSAS data-exchange ecosystem. -- **Monitor-mode-aware normalization**, format-agnostic 1D profile parsing, and - heterogeneous header extraction. - -Building on pyFAI [@pyfai] and fabio for integration and image I/O, `saxsabs` -adds the calibration-control and metadata-plumbing layers typically handled by -ad hoc local scripts at synchrotron beamlines. +`saxsabs` converts small-angle X-ray scattering (SAXS) data to an absolute +intensity scale through a Python library, command-line tools, batch workflows, +and a bilingual desktop interface. It normalizes monitor and transmission data, +estimates the reference-derived calibration factor $K$, records sample +thickness, detects previously applied corrections, propagates available +uncertainties, and exports text, canSAS, and NXcanSAS files. Built-in reference +curves cover NIST Standard Reference Material (SRM) 3600 glassy carbon +[@allen2017; @srm3600] and water at documented temperatures [@orthaber2000]; +users can also supply reference curves. + +The software stops operations when required physical metadata are missing or +inconsistent. Strict calibration records retain source hashes, units, physical +inputs, and applied corrections; other interfaces record the available source +identity and processing context. These records allow repeated or incompatible +processing to be detected. The source is distributed under the BSD-3-Clause +license, with archived releases linked through Zenodo [@saxsabs_archive]. # Statement of need -Converting SAXS detector images to absolute-scale intensities requires -dark-current subtraction, beam-monitor normalization, transmission and thickness -correction, azimuthal integration, and calibration against a reference standard. -In practice, each step is complicated by real-world heterogeneity: header formats -differ across beamlines, 1D profiles use inconsistent delimiters and column names, -and the choice of calibration standard varies by experimental context—glassy -carbon for solid-state SAXS, liquid water for BioSAXS, or a user-measured -secondary standard for instrument-specific setups. - -Existing tools address individual stages well—pyFAI [@pyfai], SasView [@sasview], -Dioptas [@dioptas], BioXTAS RAW [@bioxtasraw], and Irena [@irena]—but none -provides a dedicated end-to-end absolute-calibration workflow that jointly handles -metadata heterogeneity, multi-standard calibration, robust K-factor estimation, -buffer subtraction with error propagation, and batch processing with audit -trails. Moreover, the growing adoption of standardized data formats such as -canSAS XML and NXcanSAS HDF5 for SAXS data exchange demands that calibration -tools produce interoperable outputs natively. `saxsabs` fills this gap. +Absolute scaling enables quantitative comparison of SAXS measurements and their +interpretation as differential scattering cross sections +[@allen2017; @orthaber2000]. It requires consistent treatment of detector background, +exposure or monitor normalization, sample transmission, reference and sample +thicknesses, and the calibration standard. Beamline metadata and one-dimensional +(1D) profiles also vary in field names, units, and delimiters; undocumented +assumptions therefore hinder auditing. + +`saxsabs` serves beamline scientists and SAXS users who need to convert external +1D profiles or detector images with compatible metadata and geometry to absolute +intensity while retaining processing records. The current strict 2D workflow is +implemented for BL19B2 data conventions. The software distinguishes `raw_counts`, +`relative`, `absolute_cm^-1`, and `ambiguous` states. Scaling and buffer +subtraction run only when the declared state and required metadata are compatible; +otherwise, the software identifies the missing or conflicting information. # State of the field -While the theory of calibration against NIST SRM 3600 glassy carbon is well -documented [@glatter_kratky; @srm3600], and liquid water is increasingly used as -a secondary standard for BioSAXS [@orthaber2000], the operational workflow—parsing -heterogeneous metadata, selecting normalization modes, multi-background -subtraction, computing a robust scaling factor, and converting the result to an -interoperable standard format—is typically left to bespoke, untested scripts. - -: Functional comparison of SAXS software tools. `saxsabs` focuses on the calibration-control, standards management, and data-interoperability layer that bridges integration engines and absolute-scale reduction. []{label="tab:comparison"} - -| Capability | pyFAI | SasView | Dioptas | BioXTAS RAW | Irena | saxsabs | -|------------------------------------|:-----:|:-------:|:-------:|:-----------:|:-----:|:-------:| -| Azimuthal integration | ✓ | | ✓ | ✓ | ✓ | | -| SAS model fitting | | ✓ | | ✓ | ✓ | | -| Multi-standard calibration | | | | | | ✓ | -| Water-standard $d\Sigma/d\Omega$ | | | | ✓ | | ✓ | -| $\mu$ calculator (XCOM) | | | | | | ✓ | -| Buffer / solvent subtraction | | | | ✓ | | ✓ | -| Heterogeneous header parsing | | | | | | ✓ | -| Monitor-mode normalization | | | | | | ✓ | -| Robust K-factor (MAD filtering) | | | | | | ✓ | -| Format-agnostic 1D ingestion | | | | partial | | ✓ | -| Multi-background averaging | | | | | | ✓ | -| canSAS / NXcanSAS export | | ✓ | | | | ✓ | -| Headless CLI + CI-testable | ✓ | partial | | | | ✓ | - -`saxsabs` complements these tools by formalizing the calibration-control layer -that typically exists as private, untested scripts, making absolute-scaling -workflows reproducible and auditable. +pyFAI handles detector geometry and azimuthal integration [@pyfai], and FabIO +reads detector-image formats [@fabio]. Dioptas supports two-dimensional +diffraction reduction and exploration [@dioptas]. SasView and Irena provide +small-angle-scattering analysis and model fitting [@sasview; @irena], whereas +BioXTAS RAW combines BioSAXS reduction, water- or glassy-carbon scaling, buffer +subtraction, and subsequent analysis [@bioxtasraw]. + +`saxsabs` complements these packages by focusing on absolute-scale calibration +for external 1D data and the current BL19B2 2D workflow. It builds on pyFAI and +FabIO rather than reimplementing detector integration and image access. The +scholarly contribution is the explicit intensity-state and correction-history +contract across calibration, external-profile scaling, and traceable export--a +boundary not provided by those dependencies. Its Python, command-line, and +graphical interfaces reuse core checks for physical inputs, processing history, +and output state, although the Workbench is not an equivalent front end to the +strict campaign runner. Geometry calibration and model fitting remain with the +specialist tools above. # Software design -The primary design goal is to separate numerical calibration logic from UI -concerns, enabling both interactive GUI use and headless CLI execution. The -software is organized into four layers: - -- **Core numerical layer** (`saxsabs.core`): implements monitor normalization, - robust K-factor estimation, a pluggable standard-reference registry - (`calibration.py`—SRM 3600, water, custom data), a universal linear - attenuation coefficient calculator (`mu_calculator.py`—driven by xraydb - [@newville_xraydb]), and buffer/solvent subtraction with full error - propagation (`buffer_subtraction.py`). -- **I/O layer** (`saxsabs.io`): provides format-agnostic header parsing (fuzzy - key matching with unit conversion), multi-strategy 1D profile parsing (three - separator strategies with automatic column-role inference), and multi-format - writers for canSAS 1D XML and NXcanSAS HDF5 (`writers.py`). -- **CLI layer** (`saxsabs.cli`): exposes six subcommands: four focused utilities - (`norm-factor`, `parse-header`, `parse-external1d`, `estimate-k`), the - safety-first `bl19b2-abs2d` batch workflow, and an explicit - `bl19b2-abs2d-v1-legacy` migration entry that does not silently restore - historical scientific defaults. -- **GUI layer** (`SASAbs.py`): a tkinter-based application themed with Sun - Valley (sv\_ttk) providing a modern Windows 11 appearance with light/dark mode - toggle. Four tabbed panels—K-Factor Calibration, Batch Processing, External - 1D Conversion, and Help—offer bilingual (Chinese/English) support. - -![Calibration workflow of saxsabs. Input data (2D images, instrument metadata, and a reference standard from the pluggable registry) flow through header parsing, normalization, 2D background subtraction, pyFAI integration, and robust K-factor estimation to produce calibrated 1D profiles with structured audit outputs. Optional post-processing steps include buffer subtraction and multi-format export.](fig_workflow.png){#fig:workflow width="95%"} - -This architecture enables unit testing of core functions independently of the GUI across three operating systems and Python 3.10--3.13, while preserving the GUI for interactive beamline use (\autoref{fig:gui}). For public reproducibility, the repository includes an independent deterministic raw-frame package (`examples/minimal_2d/`) with separately constructed dark, blank, SRM 3600, and sample frames. It validates the software arithmetic and interoperable export path without claiming to replace measured beamline acceptance data. - -![The saxsabs graphical user interface in English mode, showing the four-tab layout: K-Factor Calibration, Batch Processing, External 1D Conversion, and Help.](fig_gui.png){#fig:gui width="95%"} - -## Mathematical formulation - -The absolute intensity calibration workflow follows the standard procedure documented for NIST SRM 3600 [@srm3600; @glatter_kratky]. The key computational steps are: - -**Monitor normalization.** Two modes are supported. In *rate* mode, the normalization factor is: - -$$N = t_{\mathrm{exp}} \times I_0 \times T \label{eq:norm_rate}$$ - -where $t_{\mathrm{exp}}$ is the exposure time (s), $I_0$ is the beam-monitor count, and $T$ is the sample transmission. In *integrated* mode, $t_{\mathrm{exp}}$ is omitted: - -$$N = I_0 \times T \label{eq:norm_int}$$ - -**2D background subtraction.** Detector images contain integrated counts, so the dark frame is first scaled by the exposure-time ratio. Given sample image $D_s$, dark image $D_d$ acquired for $t_d$, and a no-sample NIST blank $D_{\mathrm{bg}}$: - -$$I_{\mathrm{net}}(x,y) = \frac{D_s - (t_s/t_d)D_d}{t_s I_{0,s}T_s} - \alpha\frac{D_{\mathrm{bg}} - (t_{\mathrm{bg}}/t_d)D_d}{t_{\mathrm{bg}}I_{0,\mathrm{bg}}} \label{eq:bg_sub}$$ - -for rate-type monitors. For integrated monitors, only the denominator exposure -factors are omitted; the dark-frame exposure ratios remain mandatory because -the detector images contain integrated counts. The blank term is not divided by -a separate blank transmission. Multiple blanks are normalized individually and -then averaged pixel-wise. - -**Robust K-factor estimation.** After azimuthal integration via pyFAI to obtain $I_{\mathrm{meas}}(q)$, the profile is interpolated onto the certified NIST SRM 3600 Table 1 grid ($q \in [0.00827568, 0.247402]$ Å$^{-1}$). The certified coupon thickness 1.055 mm is used. Point-wise ratios are computed: - -$$R_i = I_{\mathrm{ref}}(q_i) \,/\, I_{\mathrm{meas}}(q_i) \label{eq:ratio}$$ - -Outlier rejection uses the median absolute deviation (MAD): - -$$\hat{\sigma} = 1.4826 \times \mathrm{median}(|R_i - \tilde{R}|) \label{eq:mad}$$ - -where $\tilde{R} = \mathrm{median}(R_i)$. Points with $|R_i - \tilde{R}| > 3\hat{\sigma}$ are rejected, and: - -$$K = \mathrm{median}(R_i) \quad \text{for } |R_i - \tilde{R}| \leq 3\hat{\sigma} \label{eq:kfactor}$$ - -The factor 1.4826 ensures consistency with the standard deviation under a Gaussian distribution. This robust estimator resists outliers from parasitic scattering, beamstop shadows, or detector artefacts (\autoref{fig:kfactor}). - -![Demonstration of the robust K-factor estimation algorithm. (a) The NIST SRM 3600 reference profile and a simulated measured profile after rescaling by K. (b) Point-wise ratios $R_i = I_{\mathrm{ref}}/I_{\mathrm{meas}}$ with inlier points (green circles) and rejected outliers (red crosses); the blue line and shaded band show the median K-factor and ±3$\hat{\sigma}$ acceptance region.](fig_kfactor_demo.png){#fig:kfactor width="95%"} - -**Absolute intensity conversion.** For each sample: - -$$I_{\mathrm{abs}}(q) = K \times I_{\mathrm{1D}}(q) \,/\, d \label{eq:abs}$$ - -where $d$ is the sample thickness (cm). When transmission is available, $d$ can be estimated via the Beer–Lambert relation $d = -\ln(T)/\mu$, where $\mu$ is the linear attenuation coefficient. `saxsabs` includes a built-in $\mu$ calculator driven by the XCOM photon cross-section database through xraydb [@newville_xraydb]: - -$$\mu = \rho \sum_i w_i \left(\frac{\mu}{\rho}\right)_i \label{eq:mu}$$ - -where $\rho$ is the bulk density, $w_i$ is the weight fraction of element $i$, and $(\mu/\rho)_i$ is its mass attenuation coefficient at the operating photon energy. The calculator accepts arbitrary chemical compositions (e.g., `"Fe:0.9,Cr:0.1"`) and preset alloy/compound libraries. - -**Buffer / solvent subtraction.** For BioSAXS and solution-scattering experiments, the solute signal is obtained by subtracting a matched buffer measurement: - -$$I_{\mathrm{solute}}(q) = I_{\mathrm{sample}}(q) - \alpha \, I_{\mathrm{buffer}}(q) \label{eq:buffer}$$ - -where $\alpha$ is a tuneable scaling factor (default 1.0). With independent -standard uncertainties, propagation gives -$\sigma_{\mathrm{solute}}^2 = \sigma_{\mathrm{sample}}^2 + \alpha^2 -\sigma_{\mathrm{buffer}}^2 + I_{\mathrm{buffer}}^2\sigma_\alpha^2$. -Missing input uncertainties remain unknown (NaN); they are not replaced by -zero. An explicit $\sigma_\alpha=0$ is required to treat $\alpha$ as exact. - -## Batch processing and automation - -The GUI batch-processing pipeline automates the complete chain from raw 2D images to calibrated 1D profiles: - -- Automatic background and dark-current matching via a weighted scoring function comparing exposure, monitor counts, transmission, and temporal proximity. -- Multi-background averaging to reduce statistical noise in the background estimate. -- Three integration modes: full-ring, angular-sector (with ±180° wrapping), and radial chi-profile. -- Sector merging with inverse-variance weighting. -- Optional buffer / solvent subtraction with tuneable $\alpha$-scaling. -- Multi-format output: TSV (default), CSV, canSAS 1D XML, and NXcanSAS HDF5. -- Data quality controls and structured output traceability (CSV reports, JSON metadata, K-factor history log). +The software separates user interfaces from reusable scientific and I/O modules +(\autoref{fig:workflow}). The `saxsabs.core` modules implement normalization, +detector reduction, calibration, material attenuation, recorded-intensity-state +assessment, preflight validation, reference matching, and uncertainty handling. `saxsabs.io` +parses heterogeneous headers and 1D tables and writes canSAS1d 1.1 XML and +NXcanSAS 1.1 HDF5 [@cansas1d; @nxcansas]. The strict BL19B2 workflow validates +detector, monitor, transmission, thickness, reference, and output inputs before +integration, calibration, and export. CLI subcommands cover normalization, +parsing, and $K$ estimation; the SAXSAbs Workbench adds interactive calibration, +batch processing, and external-1D conversion +(\autoref{fig:gui}). + +The design deliberately separates a bounded, strict BL19B2 campaign schema from +the more general calculation and I/O APIs. A permissive all-beamline workflow +would accept more files but would require silent assumptions about metadata and +correction history. The narrower strict path instead fails when those contracts +cannot be established, while the reusable modules remain available for other +interfaces. This trades immediate format breadth for auditable scientific state. + +![Architecture and data flow. The Python API, command line, and Workbench share numerical and I/O modules. Before writing absolute-intensity data, the software checks physical inputs, processing history, and output state. The diagram is derived from the public package modules and interfaces; no experimental data are shown.](fig_workflow.png){#fig:workflow width="100%"} + +To estimate $K$, `saxsabs` interpolates a measured standard profile onto the +reference grid and calculates +$R_i=I_{\mathrm{ref}}(q_i)/I_{\mathrm{meas}}(q_i)$. It defines the median ratio +as $\tilde{R}$ and +$\hat{\sigma}=1.4826\,\mathrm{median}(|R_i-\tilde{R}|)$, retains ratios within +$3\hat{\sigma}$, and uses their median as $K$. This rule excludes isolated +anomalous ratios. The software reports the dispersion of retained ratios +separately from combined calibration uncertainty and propagates supported +independent input uncertainties when supplied; unavailable terms remain +unspecified. The BL19B2 workflow reports a partial combined standard uncertainty +when shared covariance terms are not quantified and does not report a system +expanded uncertainty in that case. + +The attenuation functionality follows two distinct data paths. Its general +diagnostic calculator obtains energy-dependent elemental coefficients from the +Elam database through `xraydb.mu_elam` [@elam2002; @xraydb]. Its fixed 30 keV +material calculation uses a versioned NIST SRD 126 snapshot [@nist_srd126]. +Absolute 1D intensity is reported in cm$^{-1}$. CSV and TSV outputs are directly +inspectable; the structured XML and HDF5 outputs follow the documented canSAS1d +1.1 and NXcanSAS 1.1 layouts. The XML output validates against the official +canSAS1d 1.1 XSD. NXcanSAS output passes `punx` 0.3.5 [@punx] with its bundled +v2018.5 definitions, but current NeXus definitions and third-party consumers +have not yet been verified. + +![SAXSAbs Workbench in English, showing K-calibration inputs and the plotting area. The screenshot was captured from the current source tree and contains no beamline data.](fig_gui.png){#fig:gui width="100%"} + +# Software availability + +Source code, tests, documentation, and examples are available in the [SASAbs +GitHub repository](https://github.com/D-sudoasd/SASAbs) under the BSD-3-Clause +license. The core package supports Python 3.10 and later; optional dependency +groups enable Workbench, detector-image, BL19B2, and HDF5 functionality. The +README includes installation instructions and minimal commands. Archived +releases are collected in the Zenodo concept record [@saxsabs_archive]. # Research impact statement -`saxsabs` has been deployed for routine absolute intensity calibration at the Institute of Metal Research, Chinese Academy of Sciences, processing data from multiple synchrotron beamlines. It has replaced manual spreadsheet-based procedures, reducing operator intervention and substantially reducing risk from inconsistent header parsing. - -The software defines its impact along four dimensions: - -- **Operational efficiency**: Calibration previously requiring manual metadata extraction and iterative K-factor fitting is now a single CLI invocation or GUI session, reducing typical analysis turnaround and manual bookkeeping overhead. -- **Reliability**: Defensive parsing and format-agnostic ingestion have eliminated silent data-misinterpretation failures when switching between instruments. -- **Traceability**: Every run produces structured, deterministic output suitable for version control and audit. -- **Interoperability**: Native canSAS XML and NXcanSAS HDF5 export enables seamless data exchange with downstream analysis tools (SasView, ATSAS, etc.) and institutional data repositories. - -Core numerical behavior, I/O contracts, and deterministic synthetic workflows are checked by automated tests across three operating systems and four Python versions (3.10–3.13). Experimental beamline acceptance remains a separate documented activity. +The repository includes a strict BL19B2 batch workflow from detector images to +exported results. Its deterministic example in `examples/minimal_2d/` generates +synthetic dark, background, standard, and sample images; recovers the specified +calibration factor within the tested tolerance; and writes text, canSAS, and +NXcanSAS outputs. Automated tests cover numerical calculations, parsers, +exporters, the command-line interface, launchers, and Workbench validation rules. +The repository configures continuous integration for Python 3.10--3.13 on Linux, +Windows, and macOS. + +The tests and synthetic example verify implemented calculations, interfaces, +metadata handling, and output generation. They do not establish research impact. + +[Author input required before submission: a verifiable use case that identifies the +research question, software version, inputs, outputs, and the role of `saxsabs`, +with a public result or material that can be shown to the editors.] # AI usage disclosure -The following AI-assisted coding tools were used during the development of this software: +GitHub Copilot, Anthropic Claude, and OpenAI Codex assisted with code +refactoring, internationalization, test scaffolding, documentation, repository +review, figure generation, and manuscript editing. The versions of some earlier +tools were not retained. The author checked AI-assisted changes against source +code, automated test outputs, and cited primary sources and remains responsible +for the software, manuscript, scientific interpretation, and submission +decisions. +[Author input required before submission: the exact recoverable product, +model, and version for each tool, together with confirmation of the final human +review.] + +# Author contributions -- GitHub Copilot (VS Code) and Anthropic Claude were used for code refactoring, internationalization extraction, test skeleton generation, and initial documentation drafts. -- AI assistance was used for scaffolding, refactoring, test development, and documentation drafting. Scientific formulas and reference values were checked against the cited primary sources. -- Automated tests provide regression evidence for defined numerical contracts; they do not by themselves establish experimental beamline validity. + +[Author input required before submission: contribution roles for every author, +preferably using the CRediT taxonomy and confirmed by all authors.] # Acknowledgements -The author thanks beamline scientists and users at the Institute of Metal Research who provided practical feedback on data heterogeneity and workflow failure modes during the development and deployment of this software. + +[Author input required before submission: the truthful funding, sponsor-role, +acknowledgement, and competing-interest statements.] # References diff --git a/pyproject.toml b/pyproject.toml index 91250e4..c11ee1f 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -5,7 +5,7 @@ build-backend = "setuptools.build_meta" [project] name = "saxsabs" version = "2.0.0" -description = "SAXS absolute intensity calibration and robust 1D profile utilities" +description = "Traceable SAXS absolute-intensity calibration and profile I/O" readme = "README.md" requires-python = ">=3.10" license = "BSD-3-Clause" @@ -14,7 +14,7 @@ authors = [ ] keywords = ["SAXS", "absolute intensity", "calibration", "synchrotron", "scattering"] classifiers = [ - "Development Status :: 5 - Production/Stable", + "Development Status :: 4 - Beta", "Intended Audience :: Science/Research", "Programming Language :: Python :: 3", "Programming Language :: Python :: 3.10", @@ -47,9 +47,11 @@ bl19b2 = [ ] dev = [ "matplotlib>=3.7", + "PyYAML>=6.0", "pytest>=8.0", "ruff>=0.5", "setuptools>=69", + "tomli>=2.0; python_version < '3.11'", "wheel", ] @@ -58,15 +60,15 @@ saxsabs = "saxsabs.cli:main" saxsabs-workbench = "saxsabs.workbench_launcher:run_with_error_handling" [project.urls] -Homepage = "https://github.com/D-sudoasd/SASAbs_saxs-absolute-calibration" -Repository = "https://github.com/D-sudoasd/SASAbs_saxs-absolute-calibration" -Issues = "https://github.com/D-sudoasd/SASAbs_saxs-absolute-calibration/issues" -Changelog = "https://github.com/D-sudoasd/SASAbs_saxs-absolute-calibration/blob/main/CHANGELOG.md" -DOI = "https://doi.org/10.5281/zenodo.19687104" +Homepage = "https://github.com/D-sudoasd/SASAbs" +Repository = "https://github.com/D-sudoasd/SASAbs" +Issues = "https://github.com/D-sudoasd/SASAbs/issues" +Changelog = "https://github.com/D-sudoasd/SASAbs/blob/main/CHANGELOG.md" +"Concept DOI" = "https://doi.org/10.5281/zenodo.19687103" [tool.setuptools] package-dir = {"" = ".", "saxsabs" = "src/saxsabs"} -py-modules = ["SASAbs"] +py-modules = ["SASAbs", "saxs_mpl_style"] [tool.setuptools.packages.find] where = ["src"] @@ -76,3 +78,8 @@ testpaths = ["tests"] [tool.ruff] line-length = 100 + +[tool.ruff.lint] +# Keep the project lint contract explicit so a user-level Ruff configuration +# cannot silently broaden CI to unrelated stylistic rules. +select = ["E4", "E7", "E9", "F"] diff --git a/scripts/__init__.py b/scripts/__init__.py new file mode 100644 index 0000000..ffdf3c5 --- /dev/null +++ b/scripts/__init__.py @@ -0,0 +1 @@ +"""Repository maintenance and submission-check utilities.""" diff --git a/scripts/build_release_notes.py b/scripts/build_release_notes.py new file mode 100644 index 0000000..47f5213 --- /dev/null +++ b/scripts/build_release_notes.py @@ -0,0 +1,48 @@ +"""Build the GitHub Release header from the project-level concept DOI.""" + +from __future__ import annotations + +import argparse +import re +from pathlib import Path + +try: + import tomllib +except ModuleNotFoundError: # pragma: no cover - exercised by the Python 3.10 CI job + import tomli as tomllib + + +DOI_URL_PATTERN = re.compile(r"https://doi\.org/[0-9]+\.[0-9]+/\S+\Z") + + +def read_concept_doi(pyproject_path: Path) -> str: + metadata = tomllib.loads(pyproject_path.read_text(encoding="utf-8")) + try: + doi_url = metadata["project"]["urls"]["Concept DOI"] + except KeyError as exc: + raise ValueError("pyproject.toml is missing project.urls['Concept DOI']") from exc + if not isinstance(doi_url, str) or DOI_URL_PATTERN.fullmatch(doi_url) is None: + raise ValueError(f"invalid project Concept DOI URL: {doi_url!r}") + return doi_url + + +def build_release_body(doi_url: str) -> str: + return ( + f"Project DOI: {doi_url}\n\n" + "Use the project DOI for general citation. A release-specific DOI is available " + "only after Zenodo archives that release.\n" + ) + + +def main() -> None: + parser = argparse.ArgumentParser() + parser.add_argument("--pyproject", type=Path, default=Path("pyproject.toml")) + parser.add_argument("--output", type=Path, default=Path("release-body.md")) + args = parser.parse_args() + + doi_url = read_concept_doi(args.pyproject) + args.output.write_text(build_release_body(doi_url), encoding="utf-8") + + +if __name__ == "__main__": + main() diff --git a/scripts/check_public_candidate.py b/scripts/check_public_candidate.py new file mode 100644 index 0000000..fb339e7 --- /dev/null +++ b/scripts/check_public_candidate.py @@ -0,0 +1,226 @@ +#!/usr/bin/env python3 +"""Verify that the public GitHub candidate matches the local submitted checkout.""" + +from __future__ import annotations + +import argparse +import json +import os +from pathlib import Path +import re +import subprocess +import sys +from typing import Any +from urllib.error import HTTPError, URLError +from urllib.parse import quote +from urllib.request import Request, urlopen + + +CANONICAL_SLUG = "D-sudoasd/SASAbs" +CANONICAL_REPOSITORY = f"https://github.com/{CANONICAL_SLUG}" +API_REPOSITORY = f"https://api.github.com/repos/{CANONICAL_SLUG}" +EXPECTED_HOMEPAGE = "https://doi.org/10.5281/zenodo.19687103" +EXPECTED_LICENSE = "BSD-3-Clause" +ALLOWED_SUBMISSION_BRANCHES = {"main", "joss-submission"} +ALLOWED_CI_EVENTS = {"push", "pull_request", "workflow_dispatch"} + + +class PublicCandidateError(ValueError): + """Raised when local evidence and public GitHub state do not identify one candidate.""" + + +def git_output(root: Path, *arguments: str) -> str: + completed = subprocess.run( + ["git", "-c", f"safe.directory={root.as_posix()}", *arguments], + cwd=root, + check=True, + stdout=subprocess.PIPE, + stderr=subprocess.PIPE, + text=True, + encoding="utf-8", + ) + return completed.stdout.strip() + + +def fetch_json(url: str, *, token: str | None, timeout: float) -> dict[str, Any]: + headers = { + "Accept": "application/vnd.github+json", + "User-Agent": "saxsabs-joss-public-candidate-check", + "X-GitHub-Api-Version": "2022-11-28", + } + if token: + headers["Authorization"] = f"Bearer {token}" + request = Request(url, headers=headers) + try: + with urlopen(request, timeout=timeout) as response: # noqa: S310 - fixed GitHub API + payload = json.load(response) + except (HTTPError, URLError, TimeoutError, json.JSONDecodeError) as exc: + raise PublicCandidateError(f"cannot read public GitHub evidence from {url}: {exc}") from exc + if not isinstance(payload, dict): + raise PublicCandidateError(f"GitHub API returned a non-object payload for {url}") + return payload + + +def confirmed_identity(confirmations: dict[str, Any]) -> tuple[str, str, str, int]: + branch = str(confirmations.get("submitted_branch", "")).strip() + if branch not in ALLOWED_SUBMISSION_BRANCHES: + raise PublicCandidateError("confirmation JSON has no valid submitted_branch") + commit = str(confirmations.get("submitted_commit", "")).strip().lower() + if not re.fullmatch(r"[0-9a-f]{40}", commit): + raise PublicCandidateError("confirmation JSON has no 40-character submitted_commit") + ci_run_url = str(confirmations.get("ci_run_url", "")).strip() + run_match = re.fullmatch( + re.escape(CANONICAL_REPOSITORY) + r"/actions/runs/(\d+)", + ci_run_url, + ) + if run_match is None: + raise PublicCandidateError("confirmation JSON has no canonical GitHub Actions run URL") + return branch, commit, ci_run_url, int(run_match.group(1)) + + +def validate_public_candidate( + confirmations: dict[str, Any], + repository: dict[str, Any], + branch_payload: dict[str, Any], + run_payload: dict[str, Any], + readme_payload: dict[str, Any], + paper_payload: dict[str, Any], + *, + local_branch: str, + local_head: str, + local_status: str, + local_readme_blob: str, + local_paper_blob: str, +) -> dict[str, str]: + """Validate local Git state, public branch identity, visible files, and CI evidence.""" + branch, commit, ci_run_url, run_id = confirmed_identity(confirmations) + + if local_branch != branch: + raise PublicCandidateError( + f"local branch {local_branch!r} does not match submitted branch {branch!r}" + ) + if local_head.lower() != commit: + raise PublicCandidateError("local HEAD does not match submitted_commit") + if local_status: + raise PublicCandidateError("local submitted worktree is not clean") + + if repository.get("full_name") != CANONICAL_SLUG: + raise PublicCandidateError("GitHub repository identity is not canonical") + if repository.get("private") is not False: + raise PublicCandidateError("GitHub repository is not public") + if repository.get("archived") is not False or repository.get("disabled") is not False: + raise PublicCandidateError("GitHub repository is archived or disabled") + if repository.get("has_issues") is not True: + raise PublicCandidateError("GitHub issue tracker is not enabled") + if repository.get("default_branch") != "main": + raise PublicCandidateError("GitHub default branch is not main") + if repository.get("homepage") != EXPECTED_HOMEPAGE: + raise PublicCandidateError("GitHub homepage does not use the project concept DOI") + license_payload = repository.get("license") + if not isinstance(license_payload, dict) or license_payload.get("spdx_id") != EXPECTED_LICENSE: + raise PublicCandidateError("GitHub does not detect the expected BSD-3-Clause license") + + if branch_payload.get("name") != branch: + raise PublicCandidateError("GitHub branch response does not match submitted_branch") + branch_commit = branch_payload.get("commit") + if not isinstance(branch_commit, dict) or str(branch_commit.get("sha", "")).lower() != commit: + raise PublicCandidateError("public submitted branch does not point to submitted_commit") + + for label, payload, path, local_blob in ( + ("README", readme_payload, "README.md", local_readme_blob), + ("paper", paper_payload, "paper/paper.md", local_paper_blob), + ): + if payload.get("type") != "file" or payload.get("path") != path: + raise PublicCandidateError(f"public {label} is not a visible file at {path}") + if not isinstance(payload.get("size"), int) or payload["size"] <= 0: + raise PublicCandidateError(f"public {label} is empty") + if str(payload.get("sha", "")).lower() != local_blob.lower(): + raise PublicCandidateError(f"public {label} does not match the local submitted file") + + if run_payload.get("id") != run_id or run_payload.get("html_url") != ci_run_url: + raise PublicCandidateError("GitHub Actions response does not match ci_run_url") + if str(run_payload.get("head_sha", "")).lower() != commit: + raise PublicCandidateError("GitHub Actions run does not test submitted_commit") + if run_payload.get("head_branch") != branch: + raise PublicCandidateError("GitHub Actions run does not test submitted_branch") + if run_payload.get("status") != "completed" or run_payload.get("conclusion") != "success": + raise PublicCandidateError("GitHub Actions run is not completed successfully") + if run_payload.get("event") not in ALLOWED_CI_EVENTS: + raise PublicCandidateError("GitHub Actions run has an unexpected event type") + run_repository = run_payload.get("repository") + if not isinstance(run_repository, dict) or run_repository.get("full_name") != CANONICAL_SLUG: + raise PublicCandidateError("GitHub Actions run belongs to another repository") + + default_branch = str(repository["default_branch"]) + editorialbot_command = "none" + if branch != default_branch: + editorialbot_command = f"@editorialbot set branch-where-paper-is as {branch}" + return { + "repository": CANONICAL_REPOSITORY, + "submitted_branch": branch, + "submitted_commit": commit, + "ci_run_url": ci_run_url, + "editorialbot_branch_command": editorialbot_command, + } + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("--root", type=Path, default=Path(__file__).resolve().parents[1]) + parser.add_argument("--confirmations", type=Path, required=True) + parser.add_argument("--timeout", type=float, default=20.0) + args = parser.parse_args() + root = args.root.resolve() + try: + confirmations = json.loads(args.confirmations.read_text(encoding="utf-8")) + if not isinstance(confirmations, dict): + raise PublicCandidateError("confirmation JSON must contain an object") + branch, _commit, _ci_run_url, run_id = confirmed_identity(confirmations) + token = os.environ.get("GITHUB_TOKEN") or None + branch_encoded = quote(branch, safe="") + ref_encoded = quote(branch, safe="") + repository = fetch_json(API_REPOSITORY, token=token, timeout=args.timeout) + branch_payload = fetch_json( + f"{API_REPOSITORY}/branches/{branch_encoded}", + token=token, + timeout=args.timeout, + ) + run_payload = fetch_json( + f"{API_REPOSITORY}/actions/runs/{run_id}", + token=token, + timeout=args.timeout, + ) + readme_payload = fetch_json( + f"{API_REPOSITORY}/contents/README.md?ref={ref_encoded}", + token=token, + timeout=args.timeout, + ) + paper_payload = fetch_json( + f"{API_REPOSITORY}/contents/paper/paper.md?ref={ref_encoded}", + token=token, + timeout=args.timeout, + ) + result = validate_public_candidate( + confirmations, + repository, + branch_payload, + run_payload, + readme_payload, + paper_payload, + local_branch=git_output(root, "branch", "--show-current"), + local_head=git_output(root, "rev-parse", "HEAD"), + local_status=git_output(root, "status", "--porcelain=v1", "--untracked-files=all"), + local_readme_blob=git_output(root, "hash-object", "README.md"), + local_paper_blob=git_output(root, "hash-object", "paper/paper.md"), + ) + except (OSError, ValueError, subprocess.CalledProcessError) as exc: + parser.exit(1, f"public candidate invalid: {exc}\n") + + print("public_candidate=PASS") + for key, value in result.items(): + print(f"{key}={value}") + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/scripts/check_submission_readiness.py b/scripts/check_submission_readiness.py new file mode 100644 index 0000000..969b442 --- /dev/null +++ b/scripts/check_submission_readiness.py @@ -0,0 +1,472 @@ +#!/usr/bin/env python3 +"""Fail closed on mechanical JOSS pre-submission requirements. + +The script checks only facts that can be established from a local checkout. +Author-controlled declarations, demonstrated research use, repository age, and +remote CI are reported as manual gates rather than guessed from local files. +""" + +from __future__ import annotations + +import argparse +import json +import os +import re +import shutil +import subprocess +import sys +from datetime import date, datetime +from pathlib import Path + + +ROOT = Path(__file__).resolve().parents[1] +PAPER = ROOT / "paper" / "paper.md" +EARLIEST_SUBMISSION_DATE = date(2026, 8, 26) +CANONICAL_REPOSITORY = "https://github.com/D-sudoasd/SASAbs" + +REQUIRED_SECTIONS = ( + "Summary", + "Statement of need", + "State of the field", + "Software design", + "Software availability", + "Research impact statement", + "AI usage disclosure", + "Author contributions", + "Acknowledgements", + "References", +) + +MANUAL_GATES = ( + "Public history and iterative development still satisfy JOSS on the submission date.", + "The public repository identity and submitted branch match the confirmed candidate.", + "A verifiable research-use case is included in the impact statement.", + "Author list, order, corresponding author, affiliations, ORCIDs, and roles are confirmed.", + "AI tools/models/versions, scope, and final human review are confirmed.", + "Funding, sponsor role, acknowledgements, and competing interests are confirmed.", + "The public candidate revision has a green CI run.", +) + + +def read(path: Path) -> str: + return path.read_text(encoding="utf-8") + + +def current_date() -> date: + """Return the runtime date; isolated for deterministic strict-gate tests.""" + return date.today() + + +def project_version() -> str: + match = re.search( + r'(?ms)^\[project\]\s*$.*?^version\s*=\s*"([^"]+)"\s*$', + read(ROOT / "pyproject.toml"), + ) + if not match: + raise ValueError("pyproject.toml has no [project] version") + return match.group(1) + + +def citation_keys(markdown: str) -> set[str]: + # Exclude e-mail addresses while retaining Pandoc bracketed and narrative + # citations such as [@key] and "@key showed ...". + return set(re.findall(r"(? set[str]: + return set(re.findall(r"(?m)^@\w+\{([^,]+),", bibtex)) + + +def local_readme_targets(markdown: str) -> set[str]: + targets = set( + re.findall(r"(? set[str]: + anchors: set[str] = set() + counts: dict[str, int] = {} + for heading in re.findall(r"(?m)^#{1,6}\s+(.+?)\s*$", markdown): + slug = heading.strip().lower() + slug = re.sub(r"<[^>]+>", "", slug) + slug = re.sub(r"[^\w\- ]", "", slug, flags=re.UNICODE) + slug = re.sub(r"[\s-]+", "-", slug).strip("-") + if not slug: + continue + occurrence = counts.get(slug, 0) + counts[slug] = occurrence + 1 + anchors.add(slug if occurrence == 0 else f"{slug}-{occurrence}") + return anchors + + +def local_readme_anchors(markdown: str) -> set[str]: + targets = set( + re.findall(r"(? str | None: + try: + completed = subprocess.run( + ["git", "-c", f"safe.directory={ROOT.as_posix()}", *arguments], + cwd=ROOT, + check=True, + stdout=subprocess.PIPE, + stderr=subprocess.PIPE, + text=True, + encoding="utf-8", + ) + except (FileNotFoundError, subprocess.CalledProcessError): + return None + return completed.stdout.strip() + + +def paper_word_count() -> int: + pandoc = os.environ.get("PANDOC") or shutil.which("pandoc") + if pandoc is None: + raise RuntimeError( + "pandoc is required for the paper word-count gate; install it or set PANDOC" + ) + command = [ + pandoc, + str(PAPER), + "--from=markdown", + "--to=plain", + f"--resource-path={PAPER.parent}", + ] + try: + completed = subprocess.run( + command, + check=True, + stdout=subprocess.PIPE, + stderr=subprocess.PIPE, + text=True, + encoding="utf-8", + ) + except FileNotFoundError as exc: + raise RuntimeError(f"configured Pandoc executable was not found: {pandoc}") from exc + except subprocess.CalledProcessError as exc: + raise RuntimeError(f"pandoc failed: {exc.stderr.strip()}") from exc + + body = completed.stdout.split("References", 1)[0] + body = re.sub(r"\[Author input required.*?\]", "", body, flags=re.DOTALL) + return len(re.findall(r"[A-Za-z0-9][A-Za-z0-9'./+^-]*", body)) + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument( + "--allow-author-placeholders", + action="store_true", + help="report author-controlled placeholders without failing the local gate", + ) + parser.add_argument( + "--as-of", + type=date.fromisoformat, + default=date.today(), + metavar="YYYY-MM-DD", + help="date used for deterministic eligibility checks; defaults to today", + ) + parser.add_argument( + "--manual-confirmations", + type=Path, + help="author-approved JSON record required by strict submission mode", + ) + args = parser.parse_args() + runtime_date = current_date() + + failures: list[str] = [] + paper = read(PAPER) + version = project_version() + front_matter_match = re.match(r"\A---\s*\n(.*?)\n---\s*\n", paper, re.DOTALL) + if front_matter_match is None: + failures.append("paper has no parseable YAML front matter") + front_matter = "" + else: + front_matter = front_matter_match.group(1) + + if args.as_of < EARLIEST_SUBMISSION_DATE: + failures.append( + f"submission date {args.as_of.isoformat()} is before the conservative " + f"eligibility date {EARLIEST_SUBMISSION_DATE.isoformat()}" + ) + if not args.allow_author_placeholders and args.as_of > runtime_date: + failures.append( + f"strict mode cannot use future submission date {args.as_of.isoformat()}; " + f"runtime date is {runtime_date.isoformat()}" + ) + + paper_title_match = re.search(r"(?m)^title:\s*(.+?)\s*$", front_matter) + if paper_title_match is None: + failures.append("paper front matter has no title") + paper_title = "" + else: + paper_title = paper_title_match.group(1).strip() + if ( + len(paper_title) >= 2 + and paper_title[0] == paper_title[-1] + and paper_title[0] in {"'", '"'} + ): + paper_title = paper_title[1:-1] + + paper_date_match = re.search(r"(?m)^date:\s*(.+?)\s*$", front_matter) + if paper_date_match is None: + failures.append("paper front matter has no date") + paper_date = None + else: + try: + paper_date = datetime.strptime(paper_date_match.group(1), "%d %B %Y").date() + except ValueError: + failures.append("paper date must use the JOSS format 'D Month YYYY'") + paper_date = None + if paper_date is not None and paper_date != args.as_of: + failures.append( + f"paper date {paper_date.isoformat()} does not match submission date " + f"{args.as_of.isoformat()}" + ) + + corresponding_count = len( + re.findall(r"(?m)^\s+corresponding:\s*true\s*$", front_matter) + ) + if corresponding_count != 1 and not args.allow_author_placeholders: + failures.append( + f"paper must identify exactly one corresponding author; found {corresponding_count}" + ) + author_email_present = bool( + re.search(r"(?m)^\s+email:\s*\S+@\S+\s*$", front_matter) + ) + if not author_email_present and not args.allow_author_placeholders: + failures.append("paper front matter has no author email") + + for section in REQUIRED_SECTIONS: + if not re.search(rf"(?m)^# {re.escape(section)}\s*$", paper): + failures.append(f"missing paper section: {section}") + + placeholder_count = paper.count("[Author input required before submission:") + if placeholder_count and not args.allow_author_placeholders: + failures.append(f"paper contains {placeholder_count} author-input placeholders") + + confirmations: dict[str, object] = {} + confirmed_branch = "" + if not args.allow_author_placeholders: + if args.manual_confirmations is None: + failures.append("strict mode requires --manual-confirmations JSON") + else: + try: + confirmations = json.loads(args.manual_confirmations.read_text(encoding="utf-8")) + except (OSError, json.JSONDecodeError) as exc: + failures.append(f"cannot read manual confirmations: {exc}") + else: + boolean_fields = ( + "public_history_confirmed", + "repository_identity_confirmed", + "research_use_confirmed", + "authorship_confirmed", + "ai_disclosure_confirmed", + "funding_and_coi_confirmed", + ) + for field in boolean_fields: + if confirmations.get(field) is not True: + failures.append(f"manual confirmation is not true: {field}") + ci_run_url = str(confirmations.get("ci_run_url", "")) + if not re.fullmatch( + re.escape(CANONICAL_REPOSITORY) + r"/actions/runs/\d+", ci_run_url + ): + failures.append( + "manual confirmations contain no valid Actions run URL for the " + "canonical repository" + ) + confirmed_branch = str(confirmations.get("submitted_branch", "")).strip() + if confirmed_branch not in {"main", "joss-submission"}: + failures.append( + "manual confirmations contain no valid submitted branch" + ) + submitted_commit = str(confirmations.get("submitted_commit", "")) + if not re.fullmatch(r"[0-9a-fA-F]{40}", submitted_commit): + failures.append("manual confirmations contain no 40-character commit SHA") + if confirmations.get("confirmed_on") != args.as_of.isoformat(): + failures.append("manual confirmations date does not match --as-of") + evidence_reference = str( + confirmations.get("research_evidence_reference", "") + ).strip() + if not evidence_reference or evidence_reference.lower() in { + "todo", + "tbd", + "none", + "n/a", + }: + failures.append( + "manual confirmations contain no research-evidence reference" + ) + + if re.search(r"(?i)\b(TODO|TBD|FIXME)\b", paper): + failures.append("paper contains TODO/TBD/FIXME text") + + cited = citation_keys(paper) + bibliography = bibliography_keys(read(ROOT / "paper" / "paper.bib")) + if cited - bibliography: + failures.append(f"paper has missing bibliography keys: {sorted(cited - bibliography)}") + if bibliography - cited: + failures.append(f"paper has unused bibliography keys: {sorted(bibliography - cited)}") + + readme = read(ROOT / "README.md") + for target in sorted(local_readme_targets(readme)): + if not (ROOT / target).exists(): + failures.append(f"README local target does not exist: {target}") + missing_anchors = local_readme_anchors(readme) - markdown_heading_anchors(readme) + if missing_anchors: + failures.append(f"README has missing local anchors: {sorted(missing_anchors)}") + + versions = { + "pyproject": version, + "citation": re.search(r'(?m)^version: "([^"]+)"$', read(ROOT / "CITATION.cff")), + "codemeta": json.loads(read(ROOT / "codemeta.json"))["version"], + "zenodo": json.loads(read(ROOT / ".zenodo.json"))["version"], + } + citation_match = versions["citation"] + if citation_match is None: + failures.append("CITATION.cff has no version") + else: + versions["citation"] = citation_match.group(1) + if any(value != version for value in versions.values()): + failures.append(f"version metadata disagree: {versions}") + + canonical = CANONICAL_REPOSITORY + citation_text = read(ROOT / "CITATION.cff") + codemeta = json.loads(read(ROOT / "codemeta.json")) + zenodo = json.loads(read(ROOT / ".zenodo.json")) + if f'repository-code: "{canonical}"' not in citation_text: + failures.append("CITATION.cff does not use the canonical repository") + if f'url: "{canonical}"' not in citation_text: + failures.append("CITATION.cff URL does not identify the candidate repository") + if "10.5281/zenodo.19687103" in citation_text: + failures.append("CITATION.cff presents the concept DOI as an unreleased version DOI") + if codemeta.get("identifier") != canonical or codemeta.get("url") != canonical: + failures.append("CodeMeta candidate identity does not use the canonical repository") + if codemeta.get("codeRepository") != canonical: + failures.append("CodeMeta codeRepository is not canonical") + related = zenodo.get("related_identifiers", []) + expected_relations = { + ("isSupplementTo", canonical, ""), + ("isVersionOf", "10.5281/zenodo.19687103", "doi"), + } + actual_relations = { + ( + str(item.get("relation", "")), + str(item.get("identifier", "")), + str(item.get("scheme", "")), + ) + for item in related + if isinstance(item, dict) + } + if actual_relations != expected_relations: + failures.append("Zenodo related identifiers do not match repository/concept DOI roles") + if paper_title and zenodo.get("title") != paper_title: + failures.append("Zenodo title does not match the JOSS paper title") + + changelog = read(ROOT / "CHANGELOG.md") + if f"## [{version}] - Unreleased" not in changelog: + failures.append(f"CHANGELOG does not mark {version} as Unreleased") + + required_files = ( + ROOT / "LICENSE", + ROOT / "CITATION.cff", + ROOT / "CONTRIBUTING.md", + ROOT / "CODE_OF_CONDUCT.md", + ROOT / "docs" / "api.md", + ROOT / "paper" / "paper.bib", + ROOT / "paper" / "fig_workflow.png", + ROOT / "paper" / "fig_gui.png", + ) + for path in required_files: + if not path.is_file() or path.stat().st_size == 0: + failures.append(f"missing or empty required file: {path.relative_to(ROOT)}") + + generated_names = {"__pycache__", ".pytest_cache", ".ruff_cache", "build", "dist"} + generated = [ + path.relative_to(ROOT) + for path in ROOT.rglob("*") + if path.is_dir() and (path.name in generated_names or path.name.endswith(".egg-info")) + ] + if generated: + failures.append(f"generated cache/build directories remain: {generated}") + + branch_value = git_output("branch", "--show-current") + branch = "unknown" if branch_value is None else (branch_value or "detached") + allowed_submission_branches = {"main", "joss-submission"} + if args.allow_author_placeholders: + if branch not in allowed_submission_branches | {"unknown"}: + failures.append(f"unexpected submission branch: {branch}") + else: + if branch == "unknown": + failures.append("strict mode cannot verify the current Git branch") + elif branch not in allowed_submission_branches: + failures.append(f"unexpected submission branch: {branch}") + elif confirmed_branch in allowed_submission_branches and branch != confirmed_branch: + failures.append( + f"current branch {branch!r} does not match confirmed submitted branch " + f"{confirmed_branch!r}" + ) + + current_head = git_output("rev-parse", "HEAD") + status_porcelain = git_output("status", "--porcelain=v1", "--untracked-files=all") + worktree_clean = status_porcelain == "" if status_porcelain is not None else None + if not args.allow_author_placeholders: + if current_head is None or status_porcelain is None: + failures.append("strict mode cannot verify Git HEAD and worktree state") + else: + submitted_commit = str(confirmations.get("submitted_commit", "")) + if re.fullmatch(r"[0-9a-fA-F]{40}", submitted_commit) and ( + submitted_commit.lower() != current_head.lower() + ): + failures.append("manual confirmation commit does not match current HEAD") + if not worktree_clean: + failures.append("strict mode requires a clean Git worktree") + + try: + words = paper_word_count() + except RuntimeError as exc: + failures.append(str(exc)) + words = -1 + else: + if not 750 <= words <= 1750: + failures.append(f"paper body is {words} words; required range is 750-1750") + + status = "PASS" if not failures else "FAIL" + print(f"mechanical_readiness={status}") + print(f"version={version}") + print(f"as_of={args.as_of.isoformat()}") + print(f"earliest_submission_date={EARLIEST_SUBMISSION_DATE.isoformat()}") + print(f"paper_body_words={words}") + print(f"author_placeholders={placeholder_count}") + print(f"citations={len(cited)}") + print(f"bibliography_entries={len(bibliography)}") + print(f"branch={branch}") + print(f"submitted_branch_confirmed={confirmed_branch or 'none'}") + print(f"current_head={current_head or 'unknown'}") + print( + "worktree_clean=" + + ("unknown" if worktree_clean is None else str(worktree_clean).lower()) + ) + print(f"corresponding_authors={corresponding_count}") + print(f"author_email_present={author_email_present}") + print(f"manual_confirmations_loaded={bool(confirmations)}") + for failure in failures: + print(f"FAIL: {failure}") + for gate in MANUAL_GATES: + print(f"MANUAL: {gate}") + return 0 if not failures else 1 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/scripts/validate_release_metadata.py b/scripts/validate_release_metadata.py new file mode 100644 index 0000000..9644470 --- /dev/null +++ b/scripts/validate_release_metadata.py @@ -0,0 +1,140 @@ +"""Fail closed before publishing a version tag as a GitHub Release.""" + +from __future__ import annotations + +import argparse +from datetime import date +import json +import os +from pathlib import Path +import re + +try: + import tomllib +except ModuleNotFoundError: # pragma: no cover - exercised by the Python 3.10 CI job + import tomli as tomllib + + +class ReleaseMetadataError(ValueError): + """Raised when a tagged release still contains provisional metadata.""" + + +FINAL_RELEASE_MESSAGE = "Cite the version-specific archive record for this release." + + +def _yaml_scalar(text: str, key: str) -> str: + match = re.search(rf"(?m)^{re.escape(key)}:\s*(.+?)\s*$", text) + if match is None: + raise ReleaseMetadataError(f"CITATION.cff is missing {key}") + value = match.group(1).strip() + if len(value) >= 2 and value[0] == value[-1] and value[0] in {"'", '"'}: + value = value[1:-1] + if not value: + raise ReleaseMetadataError(f"CITATION.cff has an empty {key}") + return value + + +def _paper_title(paper: str) -> str: + front_matter = re.match(r"\A---\s*\n(.*?)\n---\s*\n", paper, re.DOTALL) + if front_matter is None: + raise ReleaseMetadataError("paper/paper.md has no parseable YAML front matter") + match = re.search(r"(?m)^title:\s*(.+?)\s*$", front_matter.group(1)) + if match is None: + raise ReleaseMetadataError("paper/paper.md has no title") + title = match.group(1).strip() + if len(title) >= 2 and title[0] == title[-1] and title[0] in {"'", '"'}: + title = title[1:-1] + if not title: + raise ReleaseMetadataError("paper/paper.md has an empty title") + return title + + +def _dated_changelog_release(changelog: str, version: str) -> date: + headings = re.findall( + rf"(?m)^## \[{re.escape(version)}\] - (.+?)\s*$", + changelog, + ) + if len(headings) != 1: + raise ReleaseMetadataError( + f"CHANGELOG must contain exactly one heading for version {version}" + ) + value = headings[0].strip() + if value.casefold() == "unreleased": + raise ReleaseMetadataError(f"finalize the CHANGELOG date for {version} before tagging") + try: + parsed = date.fromisoformat(value) + except ValueError as exc: + raise ReleaseMetadataError( + f"CHANGELOG release date {value!r} is not a valid ISO calendar date" + ) from exc + if parsed.isoformat() != value: + raise ReleaseMetadataError( + f"CHANGELOG release date {value!r} must use zero-padded YYYY-MM-DD" + ) + return parsed + + +def validate_release_metadata(root: Path, tag: str) -> tuple[str, date]: + """Validate tag, changelog, CFF, CodeMeta, Zenodo, and paper identity.""" + pyproject = tomllib.loads((root / "pyproject.toml").read_text(encoding="utf-8")) + try: + version = pyproject["project"]["version"] + except KeyError as exc: + raise ReleaseMetadataError("pyproject.toml is missing project.version") from exc + if not isinstance(version, str) or not version.strip(): + raise ReleaseMetadataError("pyproject.toml project.version is not a non-empty string") + + expected_tag = f"v{version}" + if tag != expected_tag: + raise ReleaseMetadataError(f"release tag {tag!r} does not match {expected_tag!r}") + + changelog = (root / "CHANGELOG.md").read_text(encoding="utf-8") + release_date = _dated_changelog_release(changelog, version) + + citation = (root / "CITATION.cff").read_text(encoding="utf-8") + if _yaml_scalar(citation, "message") != FINAL_RELEASE_MESSAGE: + raise ReleaseMetadataError( + "CITATION.cff message is not the finalized release citation instruction" + ) + if _yaml_scalar(citation, "version") != version: + raise ReleaseMetadataError("CITATION.cff version does not match project.version") + citation_date = _yaml_scalar(citation, "date-released") + try: + parsed_citation_date = date.fromisoformat(citation_date) + except ValueError as exc: + raise ReleaseMetadataError( + f"CITATION.cff date-released {citation_date!r} is not a valid ISO calendar date" + ) from exc + if parsed_citation_date != release_date or citation_date != release_date.isoformat(): + raise ReleaseMetadataError( + "CITATION.cff date-released does not match the dated CHANGELOG heading" + ) + + codemeta = json.loads((root / "codemeta.json").read_text(encoding="utf-8")) + if codemeta.get("version") != version: + raise ReleaseMetadataError("codemeta.json version does not match project.version") + + zenodo = json.loads((root / ".zenodo.json").read_text(encoding="utf-8")) + if zenodo.get("version") != version: + raise ReleaseMetadataError(".zenodo.json version does not match project.version") + paper_title = _paper_title((root / "paper" / "paper.md").read_text(encoding="utf-8")) + if zenodo.get("title") != paper_title: + raise ReleaseMetadataError(".zenodo.json title does not match the JOSS paper title") + + return version, release_date + + +def main() -> None: + parser = argparse.ArgumentParser() + parser.add_argument("--root", type=Path, default=Path.cwd()) + parser.add_argument("--tag", default=os.environ.get("GITHUB_REF_NAME", "")) + args = parser.parse_args() + try: + version, release_date = validate_release_metadata(args.root, args.tag) + except (OSError, ValueError) as exc: + parser.exit(1, f"release metadata invalid: {exc}\n") + print(f"release_metadata=PASS version={version} date={release_date.isoformat()}") + + +if __name__ == "__main__": + main() diff --git a/src/saxsabs/io/parsers.py b/src/saxsabs/io/parsers.py index 7e885b2..fa0014e 100644 --- a/src/saxsabs/io/parsers.py +++ b/src/saxsabs/io/parsers.py @@ -49,6 +49,10 @@ "correctionsapplied": "corrections_applied", "donotrepeat": "do_not_repeat", "intensityunit": "intensity_unit", + "thicknesscm": "thickness_cm", + "inheritedthicknesscm": "thickness_cm", + "thicknesssource": "thickness_source", + "inheritedthicknesssource": "thickness_source", "buffersourcename": "buffer_source_name", "buffersourcesha256": "buffer_source_sha256", "bufferalpha": "buffer_alpha", diff --git a/src/saxsabs/io/writers.py b/src/saxsabs/io/writers.py index d1c55ba..e7978e4 100644 --- a/src/saxsabs/io/writers.py +++ b/src/saxsabs/io/writers.py @@ -40,6 +40,8 @@ "corrections_applied", "do_not_repeat", "intensity_unit", + "thickness_cm", + "thickness_source", "buffer_source_name", "buffer_source_sha256", "buffer_alpha", @@ -112,7 +114,7 @@ def write_cansas1d_xml( metadata : dict or None Optional keys: ``title``, ``run``, ``wavelength_A``, ``sdd_m``, ``sample_name``, ``instrument_name``, ``detector_name``, - ``process_name``. + ``process_name``, ``process_note``, and ``note``. Returns ------- @@ -172,6 +174,12 @@ def write_cansas1d_xml( ) for key, value in _operator_provenance_from_metadata(meta).items(): ET.SubElement(sasproc, "term", name=key).text = value + ET.SubElement(sasproc, "SASprocessnote").text = str( + meta.get("process_note") or "Profile exported by saxsabs." + ) + ET.SubElement(entry, "SASnote").text = str( + meta.get("note") or "Generated by saxsabs." + ) # Write tree = ET.ElementTree(root) diff --git a/submission/softwarex/README.md b/submission/softwarex/README.md deleted file mode 100644 index 71fc7ef..0000000 --- a/submission/softwarex/README.md +++ /dev/null @@ -1,19 +0,0 @@ -# SoftwareX submission package - -This directory contains a ready-to-submit SoftwareX package draft. - -## Included files - -- `softwarex_manuscript.md` - main manuscript draft in SoftwareX style -- `cover_letter.md` - cover letter to editor -- `highlights.txt` - submission highlights -- `declarations.md` - conflicts, funding, data availability, CRediT -- `graphical_abstract_note.md` - graphical abstract guidance -- `softwarex_submission_checklist.md` - pre-upload checklist -- `upload_instructions.md` - upload order and portal field mapping -- `suggested_reviewers_template.md` - optional reviewer suggestions template - -## Notes - -- Keep `paper/paper.md` as the JOSS version; this folder is the SoftwareX track. -- If required by portal, convert manuscript markdown to docx before upload. diff --git a/submission/softwarex/cover_letter.md b/submission/softwarex/cover_letter.md deleted file mode 100644 index e4c0ab4..0000000 --- a/submission/softwarex/cover_letter.md +++ /dev/null @@ -1,34 +0,0 @@ -Dear Editors of SoftwareX, - -Please find attached our manuscript entitled: - -"saxsabs: Reproducible absolute intensity calibration software for small-angle X-ray scattering" - -for consideration as a software article in SoftwareX. - -This submission presents `saxsabs`, an open-source Python software package for SAXS absolute intensity calibration workflows. The software provides a tested calibration-control layer that complements established integration and analysis tools by addressing practical reproducibility challenges in routine beamline operation, including heterogeneous metadata parsing, robust K-factor estimation, multi-standard support, attenuation calculations, buffer subtraction, and interoperable canSAS/NXcanSAS export. - -The software is publicly available under a BSD-3-Clause license and includes: - -- a modular Python package, -- command-line reproducibility workflows, -- a bilingual GUI workbench for operational use, -- automated cross-platform tests in CI, -- and a deterministic minimal anonymized 2D reproducibility package suitable for reviewer verification when proprietary raw beamline data cannot be shared. - -We confirm that: - -1. This manuscript is original and has not been published elsewhere. -2. The manuscript is not under consideration by any other journal. -3. All authors have approved the manuscript and its submission. -4. There are no competing financial or personal interests affecting this work. - -Thank you for your consideration. - -Sincerely, - -Delun Gong -Institute of Metal Research, Chinese Academy of Sciences -Shenyang 110016, China -Email: dlgong@imr.ac.cn -ORCID: 0000-0001-7877-7707 diff --git a/submission/softwarex/declarations.md b/submission/softwarex/declarations.md deleted file mode 100644 index 26ff405..0000000 --- a/submission/softwarex/declarations.md +++ /dev/null @@ -1,17 +0,0 @@ -# Submission declarations (SoftwareX) - -## Declaration of competing interest - -The author declares that there are no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. - -## Funding - -This work did not receive dedicated external funding specifically for software publication. - -## Data availability - -No proprietary raw beamline data are included in the repository. Reviewer reproducibility is supported through deterministic synthetic examples and the minimal anonymized 2D package in `examples/minimal_2d/`. - -## CRediT author statement - -Delun Gong: Conceptualization, Methodology, Software, Validation, Investigation, Writing - Original Draft, Writing - Review & Editing. diff --git a/submission/softwarex/graphical_abstract_note.md b/submission/softwarex/graphical_abstract_note.md deleted file mode 100644 index 1e1985e..0000000 --- a/submission/softwarex/graphical_abstract_note.md +++ /dev/null @@ -1,12 +0,0 @@ -# Graphical abstract package - -Recommended graphical abstract candidate in this repository: - -- `paper/fig_workflow.png` - -Rationale: - -- It summarizes the software pipeline from inputs to calibrated outputs. -- It visually emphasizes reproducibility and interoperability layers. - -If SoftwareX submission system requires specific pixel dimensions, export a resized copy from `paper/fig_workflow.png` before upload and keep content unchanged. diff --git a/submission/softwarex/highlights.txt b/submission/softwarex/highlights.txt deleted file mode 100644 index 30fd723..0000000 --- a/submission/softwarex/highlights.txt +++ /dev/null @@ -1,5 +0,0 @@ -- Reproducible SAXS absolute-intensity calibration software with tested workflows -- Multi-standard calibration: SRM 3600, water, and user-defined references -- Robust K-factor estimation using MAD-based outlier rejection -- Interoperable export to canSAS XML and NXcanSAS HDF5 formats -- Deterministic minimal anonymized 2D package for reviewer reproducibility diff --git a/submission/softwarex/softwarex_manuscript.md b/submission/softwarex/softwarex_manuscript.md deleted file mode 100644 index 7708c0a..0000000 --- a/submission/softwarex/softwarex_manuscript.md +++ /dev/null @@ -1,123 +0,0 @@ -# saxsabs: Reproducible absolute intensity calibration software for small-angle X-ray scattering - -## Authors - -Delun Gong -Institute of Metal Research, Chinese Academy of Sciences, Shenyang 110016, China -ORCID: 0000-0001-7877-7707 -Email: dlgong@imr.ac.cn - -## Abstract - -Absolute intensity calibration in small-angle X-ray scattering (SAXS) is often implemented as local, beamline-specific scripts that are difficult to validate, reproduce, and maintain. We present `saxsabs`, an open-source Python software package that provides a reproducible calibration workflow from detector-reduced 1D profiles to absolute-scale intensity outputs, with optional end-to-end support for 2D-derived pipelines in graphical workflows. The software integrates robust normalization, multi-standard calibration (NIST SRM 3600, water, and user-defined references), robust K-factor estimation using MAD-based outlier rejection, composition-based attenuation coefficient calculation via XCOM (`xraydb`), and buffer/solvent subtraction with uncertainty propagation. It also provides standardized outputs in TSV, CSV, canSAS 1D XML, and NXcanSAS HDF5 formats for interoperability. The project includes a command-line interface, a bilingual GUI workbench, continuous integration across Linux/Windows/macOS, and automated tests. A deterministic minimal anonymized dataset is included to demonstrate reviewer-friendly reproducibility without proprietary beamline files. The software is deployed in routine SAXS operations and is designed to reduce manual bookkeeping and improve traceability in absolute intensity workflows. - -Keywords: SAXS; absolute intensity calibration; synchrotron; canSAS; NXcanSAS; Python - -## Code metadata - -| Item | Description | -|---|---| -| Current code version | v1.1.1 | -| Permanent link to code/repository used for this version | https://doi.org/10.5281/zenodo.19687104 | -| Legal Code License | BSD-3-Clause | -| Code versioning system used | Git | -| Software code languages, tools, and services used | Python; NumPy; pandas; pyFAI; fabio; xraydb; h5py | -| Compilation requirements, operating environments and dependencies | Python >=3.10; core: numpy>=1.24, pandas>=2.0, xraydb>=4.5; optional GUI: pyFAI, fabio, matplotlib, sv-ttk; optional NXcanSAS: h5py | -| Link to developer documentation/manual | https://github.com/D-sudoasd/SASAbs_saxs-absolute-calibration/blob/main/README.md | -| Support email for questions | dlgong@imr.ac.cn | - -## 1. Motivation and significance - -Absolute intensity calibration is required for quantitative SAXS interpretation, model comparison, and cross-instrument reproducibility. In practice, calibration involves multiple coupled operations: normalization, transmission handling, subtraction strategy, robust scaling to standards, and export into interoperable downstream formats. While established tools provide strong support for integration and analysis, many laboratories still rely on ad hoc scripts for calibration-control logic and metadata handling. This creates reproducibility and maintenance risks, especially when instrument headers, file formats, and standards vary between beamlines. - -`saxsabs` targets this software gap by packaging calibration-control logic into a tested, reusable, and scriptable workflow. The software emphasizes deterministic behavior, explicit formulas, and robust parsing across heterogeneous inputs. - -## 2. Software description - -### 2.1 Software architecture - -The project is organized into four layers: - -1. Core numerical layer (`src/saxsabs/core`): monitor normalization, robust K-factor estimation, attenuation coefficient calculation, and buffer subtraction with uncertainty propagation. -2. I/O layer (`src/saxsabs/io`): robust parsers plus standards-oriented writers (canSAS XML and NXcanSAS HDF5). -3. CLI layer (`src/saxsabs/cli.py`): headless commands for batch or CI pipelines. -4. GUI layer (`SASAbs.py` and launcher modules): bilingual desktop workflow for beamline users. - -This separation supports both interactive operation and automated reproducibility checks. - -### 2.2 Scientific and technical functionality - -Main capabilities include: - -- Multi-standard calibration with a pluggable registry (`STANDARD_REGISTRY`) containing NIST SRM 3600, water reference behavior, and user-defined references. -- Robust K-factor estimation via median and MAD filtering to reduce sensitivity to outliers. -- Composition-based attenuation coefficient calculation (XCOM-backed via `xraydb`) for arbitrary compositions and photon energies. -- Buffer/solvent subtraction with configurable scaling factor and propagated uncertainty. -- Interoperable export to TSV/CSV/canSAS XML/NXcanSAS HDF5. -- Robust heterogeneous header parsing and format-agnostic external 1D profile ingestion. - -### 2.3 Quality assurance and portability - -The software is tested in continuous integration on Linux, Windows, and macOS with Python 3.10-3.13. Automated tests cover normalization, calibration, parsing, I/O interoperability, attenuation calculation, buffer subtraction, detector-frame reduction, provenance checks, and deterministic synthetic validation. - -## 3. Illustrative examples - -The repository contains command-line examples for each major operation and a deterministic minimal anonymized 2D reproducibility package at `examples/minimal_2d/`. The script `run_minimal_2d_pipeline.py` generates a synthetic detector image workflow, computes a robust K-factor, applies absolute scaling, and exports XML/HDF5-compatible outputs. This package is intended for reviewer reproducibility in contexts where raw beamline data cannot be publicly released. - -Representative CLI examples: - -- `saxsabs norm-factor --mode rate --exp 1.0 --mon 100000 --trans 0.8` -- `saxsabs parse-header --header-json examples/header_example.json` -- `saxsabs parse-external1d --input examples/profile_example.csv` -- `saxsabs estimate-k --meas examples/k_measured.csv --ref examples/k_reference.csv --qmin 0.01 --qmax 0.2` - -## 4. Impact - -`saxsabs` is used in routine SAXS calibration workflows at the Institute of Metal Research (Chinese Academy of Sciences). In this deployment context, the package has reduced manual intervention in metadata extraction and improved traceability by producing structured outputs suitable for audit and version control. - -The software contributes practical value through: - -- workflow standardization across heterogeneous beamline metadata, -- robust scaling resistant to common experimental outliers, -- open and scriptable reproducibility paths (CLI + tests + deterministic examples), -- standards-compatible output for interoperability with downstream SAXS tools. - -## 5. Conclusions - -`saxsabs` provides a reproducible and extensible calibration-control software layer for absolute-intensity SAXS workflows. By combining robust numerical routines, interoperable data export, and practical deployment pathways (CLI and GUI), it addresses a common gap between integration engines and routine beamline operations. - -## Acknowledgments - -The author thanks beamline scientists and users at the Institute of Metal Research for workflow feedback and validation discussions. - -## Declaration of competing interest - -The author declares that there are no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. - -## Data availability - -No proprietary raw beamline data are included in the repository. Reviewer reproducibility is supported through deterministic synthetic examples and a minimal anonymized 2D reproducibility package (`examples/minimal_2d/`). - -## CRediT authorship contribution statement - -Delun Gong: Conceptualization, Methodology, Software, Validation, Investigation, Writing - Original Draft, Writing - Review & Editing. - -## References - -[1] G. Ashiotis, A. Deschildre, Z. Nawaz, J.P. Wright, D. Karkoulis, F.E. Picca, J. Kieffer, The fast azimuthal integration Python library: pyFAI, Journal of Applied Crystallography 48 (2015) 510-519. https://doi.org/10.1107/S1600576715004306 - -[2] M. Doucet, et al., SasView version 4.2, Zenodo (2018). https://doi.org/10.5281/zenodo.1412041 - -[3] C. Prescher, V.B. Prakapenka, DIOPTAS: a program for reduction of two-dimensional X-ray diffraction data and data exploration, High Pressure Research 35 (2015) 223-230. https://doi.org/10.1080/08957959.2015.1059835 - -[4] J. Ilavsky, P.R. Jemian, Irena: tool suite for modeling and analysis of small-angle scattering, Journal of Applied Crystallography 42 (2009) 347-353. https://doi.org/10.1107/S0021889809002222 - -[5] J.B. Hopkins, R.E. Gillilan, S. Skou, BioXTAS RAW: improvements to a free open-source program for small-angle X-ray scattering data reduction and analysis, Journal of Applied Crystallography 50 (2017) 1545-1553. https://doi.org/10.1107/S1600576717011438 - -[6] National Institute of Standards and Technology, Standard Reference Material 3600: Absolute Intensity Calibration Standard for Small-Angle X-ray Scattering (Certificate of Analysis), 2016. https://tsapps.nist.gov/srmext/certificates/3600.pdf - -[7] O. Glatter, O. Kratky, Small Angle X-ray Scattering, Academic Press, 1982. - -[8] D. Orthaber, A. Bergmann, O. Glatter, SAXS experiments on absolute scale with Kratky systems using water as a secondary standard, Journal of Applied Crystallography 33 (2000) 218-225. https://doi.org/10.1107/S0021889899015216 - -[9] M. Newville, xraydb: X-ray Reference Data in SQLite, 2023. https://doi.org/10.5281/zenodo.7847236 diff --git a/submission/softwarex/softwarex_submission_checklist.md b/submission/softwarex/softwarex_submission_checklist.md deleted file mode 100644 index e24dfe3..0000000 --- a/submission/softwarex/softwarex_submission_checklist.md +++ /dev/null @@ -1,30 +0,0 @@ -# SoftwareX submission checklist (prepared) - -## Manuscript files - -- [x] Main manuscript prepared: `submission/softwarex/softwarex_manuscript.md` -- [x] Cover letter prepared: `submission/softwarex/cover_letter.md` -- [x] Highlights prepared: `submission/softwarex/highlights.txt` -- [x] Declarations prepared: `submission/softwarex/declarations.md` -- [x] Graphical abstract candidate documented: `submission/softwarex/graphical_abstract_note.md` - -## Code and reproducibility package - -- [x] Public repository with versioned release: v1.1.1 -- [x] License file present (BSD-3-Clause) -- [x] Installation guide present (`README.md`) -- [x] Automated tests and CI present -- [x] Deterministic minimal anonymized example present (`examples/minimal_2d/`) - -## Metadata alignment - -- [x] Version consistency: 1.1.1 across package and citation files -- [x] ORCID included in metadata -- [x] Contact email included - -## Manual checks before final upload - -- [ ] Convert manuscript to journal-preferred editable format (Word or LaTeX source if needed by system) -- [ ] Verify current SoftwareX portal requirements for graphical abstract dimensions and optional assets -- [x] Zenodo archive created for v1.1.1 and DOI recorded in manuscript code metadata table -- [ ] Fill submission portal fields from prepared files diff --git a/submission/softwarex/suggested_reviewers_template.md b/submission/softwarex/suggested_reviewers_template.md deleted file mode 100644 index 4298220..0000000 --- a/submission/softwarex/suggested_reviewers_template.md +++ /dev/null @@ -1,25 +0,0 @@ -# Suggested reviewers template (optional) - -If the submission system requests reviewer suggestions, fill with real experts you know. -Do not include close collaborators or conflicted reviewers. - -## Reviewer 1 -- Name: -- Affiliation: -- Email: -- Expertise relevance (1-2 lines): - -## Reviewer 2 -- Name: -- Affiliation: -- Email: -- Expertise relevance (1-2 lines): - -## Reviewer 3 -- Name: -- Affiliation: -- Email: -- Expertise relevance (1-2 lines): - -## Opposed reviewers (optional) -- Name / reason: diff --git a/submission/softwarex/upload_instructions.md b/submission/softwarex/upload_instructions.md deleted file mode 100644 index cdcb4ba..0000000 --- a/submission/softwarex/upload_instructions.md +++ /dev/null @@ -1,36 +0,0 @@ -# Upload instructions for SoftwareX submission - -## 1) Files to upload first - -Primary package in `submission/softwarex/`: - -1. `softwarex_manuscript.md` (convert to `.docx` before upload if required) -2. `cover_letter.md` -3. `highlights.txt` -4. `declarations.md` -5. graphical abstract image (use `paper/fig_workflow.png`, resized only if required) - -## 2) Suggested conversion command (optional) - -If you use Pandoc locally: - -```bash -pandoc submission/softwarex/softwarex_manuscript.md -o submission/softwarex/softwarex_manuscript.docx -``` - -## 3) Code availability information to paste in system - -- Repository: https://github.com/D-sudoasd/SASAbs_saxs-absolute-calibration -- Version: v1.1.1 -- DOI: https://doi.org/10.5281/zenodo.19687104 -- License: BSD-3-Clause -- Contact: dlgong@imr.ac.cn - -## 4) Reviewer reproducibility note (recommended in cover letter or comments) - -Point reviewers to: - -- `examples/manual-verification.md` -- `examples/minimal_2d/run_minimal_2d_pipeline.py` - -This gives a deterministic end-to-end reproducibility path without proprietary beamline data. diff --git a/tests/test_bl19b2_abs2d.py b/tests/test_bl19b2_abs2d.py index e54b5a7..c9b1ab5 100644 --- a/tests/test_bl19b2_abs2d.py +++ b/tests/test_bl19b2_abs2d.py @@ -1,3 +1,4 @@ +import csv import json import math from pathlib import Path @@ -2814,7 +2815,6 @@ def fake_header(path: Path) -> BL19B2Header: ('manifest_text', 'error_match'), [ ('wrong\nfolder/frame.tif\n', 'relative_path column'), - ('relative_path\n../frame.tif\n', 'safe relative path'), ('relative_path\nC:/frame.tif\n', 'must be relative'), ('relative_path\n\\\\server\\share\\frame.tif\n', 'must be relative'), ('relative_path\n/frame.tif\n', 'must be relative'), @@ -2836,6 +2836,17 @@ def test_include_manifest_rejects_unsafe_or_duplicate_entries( bl19b2._load_include_manifest(config) +def test_include_manifest_rejects_parent_traversal_entry(tmp_path: Path): + config = _manifest_test_config(tmp_path, 'relative_path\nframe.tif\n') + with config.include_manifest_path.open('w', newline='', encoding='utf-8') as stream: + writer = csv.DictWriter(stream, fieldnames=['relative_path']) + writer.writeheader() + writer.writerow({'relative_path': '../frame.tif'}) + + with pytest.raises(ValueError, match='safe relative path'): + bl19b2._load_include_manifest(config) + + @pytest.mark.parametrize( ('relative_path', 'entry_kind', 'error_match'), [ diff --git a/tests/test_bl19b2_integrate1d.py b/tests/test_bl19b2_integrate1d.py index 7021dd1..d405961 100644 --- a/tests/test_bl19b2_integrate1d.py +++ b/tests/test_bl19b2_integrate1d.py @@ -339,8 +339,16 @@ def integrate1d(self, _image, npt, **_kwargs): def test_manifest_rejects_traversal_before_reading_outputs(tmp_path: Path): package, manifest = _build_package(tmp_path) - text = manifest.read_text(encoding="utf-8") - manifest.write_text(text.replace("problem\\sample_00001.tif", "../sample_00001.tif"), encoding="utf-8") + with manifest.open("r", encoding="utf-8", newline="") as handle: + reader = csv.DictReader(handle) + rows = list(reader) + fieldnames = reader.fieldnames + assert rows and fieldnames is not None + rows[0]["relative_path"] = "../sample_00001.tif" + with manifest.open("w", encoding="utf-8", newline="") as handle: + writer = csv.DictWriter(handle, fieldnames=fieldnames) + writer.writeheader() + writer.writerows(rows) with pytest.raises(ValueError, match="unsafe relative_path"): integration._read_manifest(integration.Integrate1DConfig(package)) diff --git a/tests/test_io_formats.py b/tests/test_io_formats.py index 9e2e57b..0714437 100644 --- a/tests/test_io_formats.py +++ b/tests/test_io_formats.py @@ -1,5 +1,7 @@ """Tests for canSAS XML and NXcanSAS HDF5 I/O round-trip.""" +import xml.etree.ElementTree as ET + import numpy as np import pytest @@ -58,6 +60,38 @@ def test_metadata_preserved(self, tmp_path): out = write_cansas1d_xml(xml_path, q, i_abs, err, metadata=meta) assert out == xml_path + def test_write_includes_schema_required_notes(self, tmp_path): + q, i_abs, err = self._make_data(3) + xml_path = tmp_path / "schema-required-notes.xml" + + write_cansas1d_xml(xml_path, q, i_abs, err) + + namespace = {"cansas": "urn:cansas1d:1.1"} + root = ET.parse(xml_path).getroot() + entry = root.find("cansas:SASentry", namespace) + assert entry is not None + process = entry.find("cansas:SASprocess", namespace) + assert process is not None + assert process.find("cansas:SASprocessnote", namespace) is not None + assert entry.find("cansas:SASnote", namespace) is not None + + def test_inherited_thickness_provenance_roundtrip(self, tmp_path): + q, i_abs, err = self._make_data(10) + xml_path = tmp_path / "thickness.xml" + write_cansas1d_xml( + xml_path, + q, + i_abs, + err, + metadata={ + "thickness_cm": "0.1", + "thickness_source": "upstream sample cell record", + }, + ) + provenance = read_cansas1d_xml(xml_path)["operator_provenance"] + assert provenance["thickness_cm"] == "0.1" + assert provenance["thickness_source"] == "upstream sample cell record" + def test_write_shape_mismatch_raises(self, tmp_path): xml_path = tmp_path / "bad.xml" try: @@ -121,6 +155,23 @@ def test_auto_detect_h5_extension(self, tmp_path): result = read_external_1d_profile(str(h5_path)) np.testing.assert_allclose(result["x"], q, rtol=1e-10) + def test_inherited_thickness_provenance_roundtrip(self, tmp_path): + q, i_abs, err = self._make_data(10) + h5_path = tmp_path / "thickness.h5" + write_nxcansas_h5( + h5_path, + q, + i_abs, + err, + metadata={ + "thickness_cm": "0.1", + "thickness_source": "upstream sample cell record", + }, + ) + provenance = read_nxcansas_h5(h5_path)["operator_provenance"] + assert provenance["thickness_cm"] == "0.1" + assert provenance["thickness_source"] == "upstream sample cell record" + @pytest.mark.parametrize("bad_value", [np.nan, np.inf, -np.inf]) def test_write_rejects_nonfinite_q(self, tmp_path, bad_value): with pytest.raises(ValueError, match="q must contain only finite values"): diff --git a/tests/test_minimal_2d_example.py b/tests/test_minimal_2d_example.py index 0be5319..15e6eb7 100644 --- a/tests/test_minimal_2d_example.py +++ b/tests/test_minimal_2d_example.py @@ -1,17 +1,21 @@ import json +import os from pathlib import Path import subprocess import sys -def test_minimal_2d_example_uses_independent_raw_frame_golden(tmp_path: Path): +def test_minimal_2d_example_uses_deterministic_raw_frame_golden(tmp_path: Path): repo = Path(__file__).resolve().parents[1] script = repo / "examples" / "minimal_2d" / "run_minimal_2d_pipeline.py" output = tmp_path / "outputs" + env = os.environ.copy() + env["PYTHONPATH"] = str(repo / "src") completed = subprocess.run( [sys.executable, str(script), "--output-dir", str(output)], cwd=repo, + env=env, text=True, capture_output=True, check=False, @@ -19,7 +23,7 @@ def test_minimal_2d_example_uses_independent_raw_frame_golden(tmp_path: Path): assert completed.returncode == 0, completed.stderr summary = json.loads((output / "summary.json").read_text(encoding="utf-8")) - assert summary["validation_type"] == "independent_synthetic_raw_frames" + assert summary["validation_type"] == "deterministic_synthetic_raw_frames" assert summary["k_relative_error"] < 0.005 assert summary["sample_max_relative_error"] < 0.01 assert summary["uncertainty_status"] == "unknown_without_input_variances" diff --git a/tests/test_plot_style.py b/tests/test_plot_style.py index 12d4a4b..aa904e4 100644 --- a/tests/test_plot_style.py +++ b/tests/test_plot_style.py @@ -72,6 +72,8 @@ def test_save_figure_does_not_mutate_live_figure(tmp_path): def test_saxsabs_dynamic_load_finds_style_module_outside_repo_cwd(tmp_path, monkeypatch): pytest.importorskip("fabio") pytest.importorskip("pyFAI") + import matplotlib + import matplotlib.pyplot as pyplot repo_root = Path(__file__).resolve().parents[1] script_path = repo_root / "SASAbs.py" @@ -79,6 +81,8 @@ def test_saxsabs_dynamic_load_finds_style_module_outside_repo_cwd(tmp_path, monk original_path = list(sys.path) try: + matplotlib.use("Agg", force=True) + assert pyplot.get_backend().lower() == "agg" monkeypatch.chdir(tmp_path) sys.path = [ item @@ -92,6 +96,7 @@ def test_saxsabs_dynamic_load_finds_style_module_outside_repo_cwd(tmp_path, monk module = importlib.util.module_from_spec(spec) spec.loader.exec_module(module) + assert pyplot.get_backend().lower() == "agg" assert hasattr(module.saxs_mpl_style, "PRESET_LABELS") assert ("single_column", "Single-column figure") in list( module.saxs_mpl_style.preset_choices() diff --git a/tests/test_public_candidate.py b/tests/test_public_candidate.py new file mode 100644 index 0000000..eb3a98a --- /dev/null +++ b/tests/test_public_candidate.py @@ -0,0 +1,135 @@ +from __future__ import annotations + +import importlib.util +from pathlib import Path + +import pytest + + +def _load_checker(): + script = Path(__file__).resolve().parents[1] / "scripts" / "check_public_candidate.py" + spec = importlib.util.spec_from_file_location("check_public_candidate", script) + if spec is None or spec.loader is None: + raise ImportError(f"cannot load public candidate checker from {script}") + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + return module + + +checker = _load_checker() +COMMIT = "a" * 40 +RUN_URL = "https://github.com/D-sudoasd/SASAbs/actions/runs/12345" + + +def _payloads(branch: str = "joss-submission"): + confirmations = { + "submitted_branch": branch, + "submitted_commit": COMMIT, + "ci_run_url": RUN_URL, + } + repository = { + "full_name": "D-sudoasd/SASAbs", + "private": False, + "archived": False, + "disabled": False, + "has_issues": True, + "default_branch": "main", + "homepage": "https://doi.org/10.5281/zenodo.19687103", + "license": {"spdx_id": "BSD-3-Clause"}, + } + branch_payload = {"name": branch, "commit": {"sha": COMMIT}} + run_payload = { + "id": 12345, + "html_url": RUN_URL, + "head_sha": COMMIT, + "head_branch": branch, + "status": "completed", + "conclusion": "success", + "event": "push", + "repository": {"full_name": "D-sudoasd/SASAbs"}, + } + readme_payload = {"type": "file", "path": "README.md", "size": 100, "sha": "b" * 40} + paper_payload = { + "type": "file", + "path": "paper/paper.md", + "size": 100, + "sha": "c" * 40, + } + return confirmations, repository, branch_payload, run_payload, readme_payload, paper_payload + + +def _validate(branch: str = "joss-submission", **overrides): + payloads = _payloads(branch) + arguments = { + "confirmations": payloads[0], + "repository": payloads[1], + "branch_payload": payloads[2], + "run_payload": payloads[3], + "readme_payload": payloads[4], + "paper_payload": payloads[5], + "local_branch": branch, + "local_head": COMMIT, + "local_status": "", + "local_readme_blob": "b" * 40, + "local_paper_blob": "c" * 40, + } + arguments.update(overrides) + return checker.validate_public_candidate(**arguments) + + +def test_submission_branch_passes_and_reports_editorialbot_command(): + result = _validate() + assert result["submitted_commit"] == COMMIT + assert result["editorialbot_branch_command"] == ( + "@editorialbot set branch-where-paper-is as joss-submission" + ) + + +def test_main_passes_without_branch_command(): + result = _validate("main") + assert result["submitted_branch"] == "main" + assert result["editorialbot_branch_command"] == "none" + + +@pytest.mark.parametrize( + ("mutation", "message"), + [ + (lambda p: p[1].update(homepage="https://doi.org/10.5281/zenodo.19687104"), + "concept DOI"), + (lambda p: p[2]["commit"].update(sha="d" * 40), "public submitted branch"), + (lambda p: p[3].update(head_sha="d" * 40), "does not test submitted_commit"), + (lambda p: p[3].update(conclusion="failure"), "not completed successfully"), + (lambda p: p[4].update(sha="d" * 40), "public README does not match"), + ], +) +def test_remote_mismatch_fails_closed(mutation, message): + payloads = _payloads() + mutation(payloads) + with pytest.raises(checker.PublicCandidateError, match=message): + checker.validate_public_candidate( + *payloads, + local_branch="joss-submission", + local_head=COMMIT, + local_status="", + local_readme_blob="b" * 40, + local_paper_blob="c" * 40, + ) + + +def test_dirty_or_wrong_local_checkout_fails_closed(): + with pytest.raises(checker.PublicCandidateError, match="not clean"): + _validate(local_status=" M README.md") + with pytest.raises(checker.PublicCandidateError, match="local HEAD"): + _validate(local_head="d" * 40) + with pytest.raises(checker.PublicCandidateError, match="local branch"): + _validate(local_branch="main") + + +def test_confirmation_identity_rejects_noncanonical_run_url(): + confirmations = { + "submitted_branch": "joss-submission", + "submitted_commit": COMMIT, + "ci_run_url": "https://example.org/actions/runs/12345", + } + with pytest.raises(checker.PublicCandidateError, match="canonical GitHub Actions"): + checker.confirmed_identity(confirmations) diff --git a/tests/test_release_metadata.py b/tests/test_release_metadata.py new file mode 100644 index 0000000..0927e71 --- /dev/null +++ b/tests/test_release_metadata.py @@ -0,0 +1,136 @@ +from __future__ import annotations + +from datetime import date +import importlib.util +import json +from pathlib import Path + +import pytest + + +def _load_release_validator(): + script = Path(__file__).resolve().parents[1] / "scripts" / "validate_release_metadata.py" + spec = importlib.util.spec_from_file_location("validate_release_metadata", script) + if spec is None or spec.loader is None: + raise ImportError(f"cannot load release metadata validator from {script}") + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + return module + + +validator = _load_release_validator() +FINAL_RELEASE_MESSAGE = validator.FINAL_RELEASE_MESSAGE +ReleaseMetadataError = validator.ReleaseMetadataError +validate_release_metadata = validator.validate_release_metadata + + +TITLE = ( + "saxsabs: Absolute-intensity calibration and provenance tracking for " + "small-angle X-ray scattering" +) + + +def _write_valid_release(root: Path) -> None: + (root / "paper").mkdir(parents=True) + (root / "pyproject.toml").write_text( + '[project]\nname = "saxsabs"\nversion = "2.0.0"\n', + encoding="utf-8", + ) + (root / "CHANGELOG.md").write_text( + "## [2.0.0] - 2026-08-26\n\n- Final release.\n", + encoding="utf-8", + ) + (root / "CITATION.cff").write_text( + f'cff-version: 1.2.0\nmessage: "{FINAL_RELEASE_MESSAGE}"\n' + 'version: "2.0.0"\ndate-released: "2026-08-26"\n', + encoding="utf-8", + ) + (root / "codemeta.json").write_text( + json.dumps({"version": "2.0.0"}), + encoding="utf-8", + ) + (root / ".zenodo.json").write_text( + json.dumps({"title": TITLE, "version": "2.0.0"}), + encoding="utf-8", + ) + (root / "paper" / "paper.md").write_text( + f"---\ntitle: '{TITLE}'\n---\n\n# Summary\n", + encoding="utf-8", + ) + + +def test_valid_release_metadata_passes(tmp_path: Path) -> None: + _write_valid_release(tmp_path) + + assert validate_release_metadata(tmp_path, "v2.0.0") == ( + "2.0.0", + date(2026, 8, 26), + ) + + +@pytest.mark.parametrize("bad_date", ["2026-99-99", "2026-02-30", "26 August 2026"]) +def test_invalid_changelog_calendar_date_is_rejected(tmp_path: Path, bad_date: str) -> None: + _write_valid_release(tmp_path) + (tmp_path / "CHANGELOG.md").write_text( + f"## [2.0.0] - {bad_date}\n", + encoding="utf-8", + ) + + with pytest.raises(ReleaseMetadataError, match="valid ISO calendar date"): + validate_release_metadata(tmp_path, "v2.0.0") + + +@pytest.mark.parametrize( + "message", + [ + "version 2.0.0 is currently unreleased", + "version 2.0.0 is unreleased", + "UNRELEASED candidate", + "version 2.0.0 is not yet released", + "version 2.0.0 has not been released", + "pre-release candidate", + "release DOI pending", + "version 2.0.0 is not archived", + ], +) +def test_nonfinal_citation_message_is_rejected(tmp_path: Path, message: str) -> None: + _write_valid_release(tmp_path) + (tmp_path / "CITATION.cff").write_text( + f'cff-version: 1.2.0\nmessage: "{message}"\n' + 'version: "2.0.0"\ndate-released: "2026-08-26"\n', + encoding="utf-8", + ) + + with pytest.raises(ReleaseMetadataError, match="message is not the finalized"): + validate_release_metadata(tmp_path, "v2.0.0") + + +def test_release_identity_mismatches_are_rejected(tmp_path: Path) -> None: + cases = { + "tag": ("v2.0.1", "release tag"), + "citation-date": ("v2.0.0", "date-released does not match"), + "zenodo-title": ("v2.0.0", "title does not match"), + "codemeta-version": ("v2.0.0", "codemeta.json version"), + } + for name, (tag, expected) in cases.items(): + root = tmp_path / name + _write_valid_release(root) + if name == "citation-date": + path = root / "CITATION.cff" + path.write_text( + path.read_text(encoding="utf-8").replace("2026-08-26", "2026-08-27"), + encoding="utf-8", + ) + elif name == "zenodo-title": + (root / ".zenodo.json").write_text( + json.dumps({"title": "Wrong title", "version": "2.0.0"}), + encoding="utf-8", + ) + elif name == "codemeta-version": + (root / "codemeta.json").write_text( + json.dumps({"version": "9.9.9"}), + encoding="utf-8", + ) + + with pytest.raises(ReleaseMetadataError, match=expected): + validate_release_metadata(root, tag) diff --git a/tests/test_release_notes.py b/tests/test_release_notes.py new file mode 100644 index 0000000..a1e455f --- /dev/null +++ b/tests/test_release_notes.py @@ -0,0 +1,56 @@ +from __future__ import annotations + +import subprocess +import sys +from pathlib import Path + + +REPOSITORY_ROOT = Path(__file__).resolve().parents[1] + + +def test_release_notes_use_project_concept_doi(tmp_path: Path) -> None: + output = tmp_path / "release-body.md" + + result = subprocess.run( + [ + sys.executable, + str(REPOSITORY_ROOT / "scripts" / "build_release_notes.py"), + "--pyproject", + str(REPOSITORY_ROOT / "pyproject.toml"), + "--output", + str(output), + ], + check=False, + capture_output=True, + text=True, + ) + + assert result.returncode == 0, result.stderr + body = output.read_text(encoding="utf-8") + assert "https://doi.org/10.5281/zenodo.19687103" in body + assert "release-specific DOI" in body + + +def test_release_notes_fail_closed_without_concept_doi(tmp_path: Path) -> None: + pyproject = tmp_path / "pyproject.toml" + pyproject.write_text( + '[project]\nname = "example"\nversion = "1.0.0"\n[project.urls]\n', + encoding="utf-8", + ) + + result = subprocess.run( + [ + sys.executable, + str(REPOSITORY_ROOT / "scripts" / "build_release_notes.py"), + "--pyproject", + str(pyproject), + "--output", + str(tmp_path / "release-body.md"), + ], + check=False, + capture_output=True, + text=True, + ) + + assert result.returncode != 0 + assert "missing project.urls['Concept DOI']" in result.stderr diff --git a/tests/test_submission_readiness.py b/tests/test_submission_readiness.py new file mode 100644 index 0000000..3b5c264 --- /dev/null +++ b/tests/test_submission_readiness.py @@ -0,0 +1,302 @@ +from __future__ import annotations + +import importlib.util +import json +from pathlib import Path + + +def _load_readiness_checker(): + script = ( + Path(__file__).resolve().parents[1] + / "scripts" + / "check_submission_readiness.py" + ) + spec = importlib.util.spec_from_file_location("check_submission_readiness", script) + if spec is None or spec.loader is None: + raise ImportError(f"cannot load submission readiness checker from {script}") + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + return module + + +readiness = _load_readiness_checker() + + +def _write(path: Path, text: str = "present") -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(text, encoding="utf-8") + + +def _minimal_ready_repository(root: Path) -> Path: + sections = "\n\n".join( + f"# {section}\n\nSection text [@reference]." + if section == "Summary" + else f"# {section}\n\nSection text." + for section in readiness.REQUIRED_SECTIONS + ) + paper = f"""--- +title: Test paper +authors: + - name: Test Author + email: author@example.org + corresponding: true + affiliation: '1' +affiliations: + - index: 1 + name: Test Institute +date: 26 August 2026 +bibliography: paper.bib +--- + +{sections} +""" + _write(root / "paper" / "paper.md", paper) + _write( + root / "paper" / "paper.bib", + "@misc{reference,\n title = {Reference}\n}\n", + ) + _write(root / "pyproject.toml", '[project]\nversion = "2.0.0"\n') + canonical = readiness.CANONICAL_REPOSITORY + concept_doi = "10.5281/zenodo.19687103" + _write( + root / "CITATION.cff", + 'version: "2.0.0"\n' + f'repository-code: "{canonical}"\n' + f'url: "{canonical}"\n', + ) + _write( + root / "codemeta.json", + json.dumps( + { + "version": "2.0.0", + "identifier": canonical, + "url": canonical, + "codeRepository": canonical, + } + ), + ) + _write( + root / ".zenodo.json", + json.dumps( + { + "title": "Test paper", + "version": "2.0.0", + "related_identifiers": [ + {"relation": "isSupplementTo", "identifier": canonical}, + { + "relation": "isVersionOf", + "identifier": concept_doi, + "scheme": "doi", + }, + ], + } + ), + ) + _write(root / "CHANGELOG.md", "## [2.0.0] - Unreleased\n") + _write(root / "README.md", "[License](LICENSE)\n") + for relative in ( + "LICENSE", + "CONTRIBUTING.md", + "CODE_OF_CONDUCT.md", + "docs/api.md", + "paper/fig_workflow.png", + "paper/fig_gui.png", + ): + _write(root / relative) + + confirmations = root / "confirmations.json" + _write( + confirmations, + json.dumps( + { + "public_history_confirmed": True, + "repository_identity_confirmed": True, + "research_use_confirmed": True, + "authorship_confirmed": True, + "ai_disclosure_confirmed": True, + "funding_and_coi_confirmed": True, + "ci_run_url": "https://github.com/D-sudoasd/SASAbs/actions/runs/12345", + "submitted_branch": "joss-submission", + "submitted_commit": "a" * 40, + "confirmed_on": "2026-08-26", + "research_evidence_reference": "editor-visible workflow record 2026-08-20", + } + ), + ) + return confirmations + + +def test_reference_and_readme_target_parsers(): + assert readiness.citation_keys( + "Text [@alpha; @beta]. Contact author@example.org. @gamma agrees." + ) == {"alpha", "beta", "gamma"} + assert readiness.bibliography_keys("@misc{alpha,\n}\n@article{beta,\n}\n") == { + "alpha", + "beta", + } + assert readiness.local_readme_targets( + "[Docs](docs/api.md) ![Hero](assets/hero.svg) " + 'external' + ) == {"docs/api.md", "assets/hero.svg"} + assert readiness.markdown_heading_anchors("# Quick start\n\n## Quick start\n") == { + "quick-start", + "quick-start-1", + } + assert readiness.local_readme_anchors( + 'Start [Docs](#documentation)' + ) == {"quick-start", "documentation"} + + +def _mock_clean_git(monkeypatch): + def fake_git_output(*arguments: str) -> str | None: + values = { + ("branch", "--show-current"): "joss-submission", + ("rev-parse", "HEAD"): "a" * 40, + ("status", "--porcelain=v1", "--untracked-files=all"): "", + } + return values.get(arguments) + + monkeypatch.setattr(readiness, "git_output", fake_git_output) + + +def test_strict_gate_passes_with_complete_evidence_record(tmp_path, monkeypatch, capsys): + confirmations = _minimal_ready_repository(tmp_path) + monkeypatch.setattr(readiness, "ROOT", tmp_path) + monkeypatch.setattr(readiness, "PAPER", tmp_path / "paper" / "paper.md") + monkeypatch.setattr(readiness, "paper_word_count", lambda: 900) + _mock_clean_git(monkeypatch) + monkeypatch.setattr( + readiness.sys, + "argv", + [ + "check_submission_readiness.py", + "--as-of", + "2026-08-26", + "--manual-confirmations", + str(confirmations), + ], + ) + + monkeypatch.setattr(readiness, "current_date", lambda: readiness.date(2026, 8, 25)) + assert readiness.main() == 1 + assert "strict mode cannot use future submission date" in capsys.readouterr().out + + monkeypatch.setattr(readiness, "current_date", lambda: readiness.date(2026, 8, 26)) + assert readiness.main() == 0 + + +def test_strict_gate_accepts_confirmed_main_after_merge(tmp_path, monkeypatch): + confirmations = _minimal_ready_repository(tmp_path) + payload = json.loads(confirmations.read_text(encoding="utf-8")) + payload["submitted_branch"] = "main" + payload["submitted_commit"] = "c" * 40 + confirmations.write_text(json.dumps(payload), encoding="utf-8") + + monkeypatch.setattr(readiness, "ROOT", tmp_path) + monkeypatch.setattr(readiness, "PAPER", tmp_path / "paper" / "paper.md") + monkeypatch.setattr(readiness, "paper_word_count", lambda: 900) + monkeypatch.setattr(readiness, "current_date", lambda: readiness.date(2026, 8, 26)) + + def main_git_output(*arguments: str) -> str | None: + values = { + ("branch", "--show-current"): "main", + ("rev-parse", "HEAD"): "c" * 40, + ("status", "--porcelain=v1", "--untracked-files=all"): "", + } + return values.get(arguments) + + monkeypatch.setattr(readiness, "git_output", main_git_output) + monkeypatch.setattr( + readiness.sys, + "argv", + [ + "check_submission_readiness.py", + "--as-of", + "2026-08-26", + "--manual-confirmations", + str(confirmations), + ], + ) + + assert readiness.main() == 0 + + +def test_strict_gate_rejects_branch_not_confirmed(tmp_path, monkeypatch, capsys): + confirmations = _minimal_ready_repository(tmp_path) + payload = json.loads(confirmations.read_text(encoding="utf-8")) + payload["submitted_branch"] = "main" + confirmations.write_text(json.dumps(payload), encoding="utf-8") + + monkeypatch.setattr(readiness, "ROOT", tmp_path) + monkeypatch.setattr(readiness, "PAPER", tmp_path / "paper" / "paper.md") + monkeypatch.setattr(readiness, "paper_word_count", lambda: 900) + monkeypatch.setattr(readiness, "current_date", lambda: readiness.date(2026, 8, 26)) + _mock_clean_git(monkeypatch) + monkeypatch.setattr( + readiness.sys, + "argv", + [ + "check_submission_readiness.py", + "--as-of", + "2026-08-26", + "--manual-confirmations", + str(confirmations), + ], + ) + + assert readiness.main() == 1 + assert ( + "current branch 'joss-submission' does not match confirmed submitted branch 'main'" + in capsys.readouterr().out + ) + + +def test_strict_gate_rejects_missing_evidence_record(tmp_path, monkeypatch, capsys): + _minimal_ready_repository(tmp_path) + monkeypatch.setattr(readiness, "ROOT", tmp_path) + monkeypatch.setattr(readiness, "PAPER", tmp_path / "paper" / "paper.md") + monkeypatch.setattr(readiness, "paper_word_count", lambda: 900) + _mock_clean_git(monkeypatch) + monkeypatch.setattr(readiness, "current_date", lambda: readiness.date(2026, 8, 26)) + monkeypatch.setattr( + readiness.sys, + "argv", + ["check_submission_readiness.py", "--as-of", "2026-08-26"], + ) + + assert readiness.main() == 1 + assert "strict mode requires --manual-confirmations JSON" in capsys.readouterr().out + + +def test_strict_gate_rejects_dirty_or_mismatched_checkout(tmp_path, monkeypatch, capsys): + confirmations = _minimal_ready_repository(tmp_path) + monkeypatch.setattr(readiness, "ROOT", tmp_path) + monkeypatch.setattr(readiness, "PAPER", tmp_path / "paper" / "paper.md") + monkeypatch.setattr(readiness, "paper_word_count", lambda: 900) + + def dirty_git_output(*arguments: str) -> str | None: + values = { + ("branch", "--show-current"): "joss-submission", + ("rev-parse", "HEAD"): "b" * 40, + ("status", "--porcelain=v1", "--untracked-files=all"): " M paper/paper.md", + } + return values.get(arguments) + + monkeypatch.setattr(readiness, "git_output", dirty_git_output) + monkeypatch.setattr(readiness, "current_date", lambda: readiness.date(2026, 8, 26)) + monkeypatch.setattr( + readiness.sys, + "argv", + [ + "check_submission_readiness.py", + "--as-of", + "2026-08-26", + "--manual-confirmations", + str(confirmations), + ], + ) + + assert readiness.main() == 1 + output = capsys.readouterr().out + assert "commit does not match current HEAD" in output + assert "requires a clean Git worktree" in output diff --git a/tests/test_version_metadata.py b/tests/test_version_metadata.py index 1741b20..0893a4d 100644 --- a/tests/test_version_metadata.py +++ b/tests/test_version_metadata.py @@ -1,6 +1,7 @@ from __future__ import annotations import json +from datetime import date from pathlib import Path import re @@ -17,6 +18,7 @@ def test_release_version_metadata_is_consistent(): zenodo = json.loads((ROOT / ".zenodo.json").read_text(encoding="utf-8")) workbench = (ROOT / "SASAbs.py").read_text(encoding="utf-8") changelog = (ROOT / "CHANGELOG.md").read_text(encoding="utf-8") + paper = (ROOT / "paper" / "paper.md").read_text(encoding="utf-8") project_version = re.search( r'(?ms)^\[project\]\s*$.*?^version\s*=\s*"([^\"]+)"\s*$', pyproject @@ -27,7 +29,41 @@ def test_release_version_metadata_is_consistent(): assert codemeta["version"] == __version__ assert zenodo["version"] == __version__ assert workbench.count(f'"{__version__}"') >= 2 - assert f"## [{__version__}] - 2026-07-13" in changelog + changelog_heading = re.search( + rf"(?m)^## \[{re.escape(__version__)}\] - (?:Unreleased|\d{{4}}-\d{{2}}-\d{{2}})$", + changelog, + ) + assert changelog_heading is not None + heading_value = changelog_heading.group(0).rsplit(" - ", 1)[1] + if heading_value != "Unreleased": + assert date.fromisoformat(heading_value).isoformat() == heading_value + assert '"Development Status :: 4 - Beta"' in pyproject + assert '"Development Status :: 5 - Production/Stable"' not in pyproject + assert '"Concept DOI" = "https://doi.org/10.5281/zenodo.19687103"' in pyproject + assert "\nDOI = " not in pyproject + canonical = "https://github.com/D-sudoasd/SASAbs" + concept_doi = "10.5281/zenodo.19687103" + assert f'repository-code: "{canonical}"' in citation + assert f'url: "{canonical}"' in citation + assert concept_doi not in citation + assert codemeta["identifier"] == canonical + assert codemeta["url"] == canonical + assert codemeta["codeRepository"] == canonical + assert zenodo["related_identifiers"] == [ + { + "relation": "isSupplementTo", + "identifier": canonical, + }, + { + "relation": "isVersionOf", + "identifier": concept_doi, + "scheme": "doi", + }, + ] + paper_title = re.search(r"(?m)^title:\s*'([^']+)'\s*$", paper) + assert paper_title is not None + assert zenodo["title"] == paper_title.group(1) + def test_source_distribution_manifest_includes_release_metadata(): manifest = (ROOT / "MANIFEST.in").read_text(encoding="utf-8").splitlines() @@ -41,5 +77,60 @@ def test_source_distribution_manifest_includes_release_metadata(): "CITATION.cff", "codemeta.json", ".zenodo.json", - "tests/conftest.py", + "CODE_OF_CONDUCT.md", + "CONTRIBUTING.md", + "LICENSE", + "SASAbs.py", + "saxs_mpl_style.py", + ".github/workflows/ci.yml", + ".github/workflows/release.yml", } <= included + + +def test_release_smoke_isolated_from_checkout_source(): + workflow = (ROOT / ".github" / "workflows" / "release.yml").read_text( + encoding="utf-8" + ) + example = ( + ROOT / "examples" / "minimal_2d" / "run_minimal_2d_pipeline.py" + ).read_text(encoding="utf-8") + + assert 'smoke_dir="$(mktemp -d)"' in workflow + assert 'cd "$smoke_dir"' in workflow + assert '"site-packages" not in module_path.parts' in workflow + assert 'python scripts/validate_release_metadata.py --tag "$GITHUB_REF_NAME"' in workflow + assert "sys.path.insert" not in example + + +def test_ci_runs_for_submission_branch_and_manual_dispatch(): + workflow = (ROOT / ".github" / "workflows" / "ci.yml").read_text(encoding="utf-8") + + assert "branches: [main, joss-submission]" in workflow + assert "workflow_dispatch:" in workflow + + +def test_submission_readiness_gate_is_packaged_and_fail_closed(): + script = (ROOT / "scripts" / "check_submission_readiness.py").read_text( + encoding="utf-8" + ) + manifest = (ROOT / "MANIFEST.in").read_text(encoding="utf-8") + + assert 'paper.count("[Author input required before submission:")' in script + assert "750 <= words <= 1750" in script + assert "EARLIEST_SUBMISSION_DATE = date(2026, 8, 26)" in script + assert "corresponding_count != 1 and not args.allow_author_placeholders" in script + assert "author_email_present = bool(" in script + assert "paper has missing bibliography keys" in script + assert "README local target does not exist" in script + assert "generated cache/build directories remain" in script + assert "strict mode requires --manual-confirmations JSON" in script + assert '"repository_identity_confirmed"' in script + assert "no valid Actions run URL for the " in script + assert "manual confirmation commit does not match current HEAD" in script + assert "strict mode requires a clean Git worktree" in script + workbench = (ROOT / "SASAbs.py").read_text(encoding="utf-8") + assert "K-only scaling requires positive thickness_cm provenance" in workbench + assert "K-only scaling requires thickness_source provenance" in workbench + assert '"thicknesscm": "thickness_cm"' in workbench + assert '"thicknesssource": "thickness_source"' in workbench + assert "recursive-include scripts *.py" in manifest diff --git a/tests/test_workbench_launcher.py b/tests/test_workbench_launcher.py index 5c4ae26..ec74670 100644 --- a/tests/test_workbench_launcher.py +++ b/tests/test_workbench_launcher.py @@ -1,5 +1,6 @@ import ast import importlib.util +import shutil import subprocess import sys import types @@ -70,18 +71,34 @@ def test_resolve_app_source_uses_source_tree_not_cwd_shadow( def test_wheel_includes_legacy_gui_module(tmp_path: Path): + source_copy = tmp_path / "source" + wheelhouse = tmp_path / "wheelhouse" + shutil.copytree( + REPO_ROOT, + source_copy, + ignore=shutil.ignore_patterns( + ".git", + "build", + "dist", + "*.egg-info", + "__pycache__", + ".pytest_cache", + ".ruff_cache", + ), + ) + wheelhouse.mkdir() completed = subprocess.run( [ sys.executable, "-m", "pip", "wheel", - str(REPO_ROOT), + str(source_copy), "--no-deps", "--no-build-isolation", "--no-cache-dir", "-w", - str(tmp_path), + str(wheelhouse), ], check=False, stdout=subprocess.PIPE, @@ -93,7 +110,7 @@ def test_wheel_includes_legacy_gui_module(tmp_path: Path): f"stdout:\n{completed.stdout}\n" f"stderr:\n{completed.stderr}" ) - wheels = list(tmp_path.glob("saxsabs-*.whl")) + wheels = list(wheelhouse.glob("saxsabs-*.whl")) assert len(wheels) == 1 with zipfile.ZipFile(wheels[0]) as wheel: diff --git a/tests/test_workbench_scientific.py b/tests/test_workbench_scientific.py index 03cf34d..af58397 100644 --- a/tests/test_workbench_scientific.py +++ b/tests/test_workbench_scientific.py @@ -1471,7 +1471,10 @@ def test_tab3_run_preloads_buffer_once_and_reports_provenance(tmp_path): module = _load_workbench_module() app = module.SAXSAbsWorkbenchApp.__new__(module.SAXSAbsWorkbenchApp) app.language = "en" - fingerprint = "a" * 64 + poni = tmp_path / "geometry.poni" + poni.write_text("poni", encoding="utf-8") + context = _calibration_context(module, poni) + fingerprint = context.fingerprint() def write_profile( path, @@ -1482,12 +1485,20 @@ def write_profile( corrections_applied, intensity_column, k_factor=None, + thickness_cm=None, + thickness_source=None, ): ledger = json.dumps(corrections_applied, separators=(",", ":")) k_line = "" if k_factor is None else f"# k_factor: {k_factor:.17g}\n" + thickness_lines = "" + if thickness_cm is not None: + thickness_lines += f"# thickness_cm: {thickness_cm:.17g}\n" + if thickness_source is not None: + thickness_lines += f"# thickness_source: {thickness_source}\n" path.write_text( f"# calibration_context_fingerprint: {fingerprint}\n" + k_line + + thickness_lines + f"# intensity_state: {intensity_state}\n" f"# intensity_unit: {intensity_unit}\n" f"# corrections_applied: {ledger}\n" @@ -1509,6 +1520,8 @@ def write_profile( "intensity_unit": "relative", "corrections_applied": ["thickness"], "intensity_column": "I_rel", + "thickness_cm": 0.1, + "thickness_source": "upstream sample cell record", } write_profile(sample_a, (10.0, 9.0, 8.0), **sample_kwargs) write_profile(sample_b, (12.0, 11.0, 10.0), **sample_kwargs) @@ -1544,7 +1557,6 @@ def write_profile( app.t3_prog_bar = {} app.root = SimpleNamespace(update_idletasks=lambda: None) app.get_monitor_mode = lambda: "rate" - context = SimpleNamespace(fingerprint=lambda: fingerprint) app.require_trusted_k_for_external = lambda *_args, **_kwargs: context app.log = lambda _message: None app.show_info = lambda *_args, **_kwargs: None @@ -1603,6 +1615,11 @@ def counted_read(path): buffer_path ) assert output_profile["operator_provenance"]["uncertainty_type"] == "combined_standard" + assert float(output_profile["operator_provenance"]["thickness_cm"]) == pytest.approx(0.1) + assert ( + output_profile["operator_provenance"]["thickness_source"] + == "upstream sample cell record" + ) output_table = __import__("pandas").read_csv( sample_output, sep="\t", @@ -1935,6 +1952,8 @@ def test_tab3_requires_explicit_relative_intensity_state_before_scaling(): "operator_provenance": { "intensity_state": "relative", "corrections_applied": '["thickness"]', + "thickness_cm": "0.1", + "thickness_source": "upstream sample cell record", }, }, "relative.dat", @@ -1942,6 +1961,34 @@ def test_tab3_requires_explicit_relative_intensity_state_before_scaling(): ) assert assessment.state.value == "relative" + with pytest.raises(ValueError, match="positive thickness_cm provenance"): + app.require_relative_external_profile_for_scaling( + { + "i_col": "I_rel", + "operator_provenance": { + "intensity_state": "relative", + "corrections_applied": '["thickness"]', + "thickness_source": "upstream sample cell record", + }, + }, + "missing-thickness-value.dat", + correction_mode="k_only", + ) + + with pytest.raises(ValueError, match="thickness_source provenance"): + app.require_relative_external_profile_for_scaling( + { + "i_col": "I_rel", + "operator_provenance": { + "intensity_state": "relative", + "corrections_applied": '["thickness"]', + "thickness_cm": "0.1", + }, + }, + "missing-thickness-source.dat", + correction_mode="k_only", + ) + with pytest.raises(ValueError, match="missing required existing.*thickness"): app.require_relative_external_profile_for_scaling( {